Skip to content

Expose comparison options as keywords of Python Document.compare - #176

Open
hadim wants to merge 1 commit into
tensorbee:mainfrom
hadim:feat/python-compare-options
Open

hadim wants to merge 1 commit into
tensorbee:mainfrom
hadim:feat/python-compare-options

Conversation

@hadim

@hadim hadim commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • rdocx-py: Document.compare(edited, author, timestamp, *, granularity="run", ignore_formatting=False, ignore_whitespace=False, ignore_fields=False, ignore_comments=False, ignored_stories=None) builds a
    native ComparisonOptions and calls compare_with_options. Two private
    parsers next to parse_raster_format map "run", "word" and
    "character" to ComparisonGranularity, and the Story.kind names body,
    header, footer, comment, text_box, footnote and endnote to
    ComparisonStoryKind. The options are parsed inside the detached closure
    before the document is serialised, so a bad value fails before any work.
  • Stub: Literal types for both vocabularies, Sequence[...] | None for
    ignored_stories, and a docstring that states the left-biased policy and
    that ignore_comments=True keeps the original's comments and anchors.
  • HLD 10: the comparison paragraph describes the Python keywords where it said
    Python only preserves comparison output.

Part of #161. This is sub-item 1, the Python half. The rdocx compare flags,
word as the CLI default and the engine changes (the edited side's comment
threads, the rebuilt TOC, granular edits next to a hyperlink) remain.

Part of #168. This closes the "Comparison options" checklist item (item G).

Why

Document.compare in Python always ran ComparisonOptions::default(). With
the default whole-run granularity, the issue's one-word edit in a single-run
paragraph (the usual shape of a Google Docs export) redlines as all 183
characters deleted and inserted again. There was also no way to compare a pair
whose comment parts differ, because only ignore_comments gets past "document
comparison requires identical related-story shells". The native API already
had all of this, the binding just did not pass it through.

Notes

  • API choices: keyword-only strings, "run" as the
    default like Rust, story names from Story.kind, and ignore_comments
    exposed now with its left bias documented in the stub and the HLD. A call
    without the new keywords is the same native call as before, and a test pins
    that its bytes equal those of a call with every default spelled out.
  • ignored_stories defaults to None rather than the () the issue sketched.
    pyo3 renders only literal defaults into __text_signature__, so a
    Vec::new() default shows as ..., and mypy.stubtest then rejects a stub
    default of (). None follows render_pages(pages=None) and
    update_fields(merge_fields=None). Tuples, lists and () are accepted. A
    bare str or a set raises TypeError from pyo3, and mypy rejects both.
  • Unknown values raise RdocxError, through rdocx::Error::Other, like an
    unknown render_pages(format=...). The binding also has a ValueError
    precedent (story_item_kind_from_name). I chose RdocxError so every option
    error has one type, since the native duplicate-story check already reports
    RdocxError. Duplicates are left to that native validation. Switching to
    ValueError is a two-line change if you prefer it.
  • table_cell is a Story.kind value but not a comparison category, so it is
    rejected as unknown. body maps to ComparisonStoryKind::Main.
  • Typing trade-off, for you to decide. The stub types ignored_stories as
    Sequence[Literal[...]], so mypy catches a misspelt story, and
    typing_smoke.py pins that. Story.kind is typed str, though, so
    ignored_stories=[s.kind for s in doc.stories if ...] works at runtime but
    fails mypy --strict with arg-type. Narrowing Story.kind to a Literal
    would not fix that on its own, because it also yields table_cell, which is
    not a comparison category. I kept the Literal. Widening to Sequence[str]
    is a one-line stub change if you prefer the pass-through idiom.
  • ignored_stories=["comment"] does not replace ignore_comments=True. On the
    test pair (a comment added on the edited side only), ignoring the comment
    story excludes the comments part but the body anchors are still compared,
    and the call fails with "comparison cannot revise paragraph boundary
    structures at body/paragraph[0]". ignore_comments also skips comment
    references and ranges (ignored_run_content and the comment_ranges
    checks in comparison.rs). That is the native contract. The HLD now says so,
    and the test uses ignore_comments.
  • Observed while writing the tests, not changed here. A w:fldSimple whose
    result changes gets w:del and w:ins inside the w:fldSimple, and neither
    Document.revisions nor accept_all() sees them (the list is empty and
    accept_all() returns 0, leaving the markup in place). A complex field with
    the same change is listed and resolved. So the test uses a complex field for
    ignore_fields. This looks like a revision reader gap in rdocx-oxml and may
    deserve its own issue.
  • Also observed: a comparison with ignore_formatting=True on a
    formatting-only pair records no revision but still re-serialises
    word/document.xml (the indentation between elements changes), so the
    binding's before and after byte check advances the handle revision. In the
    same way, ignored_stories=("text_box",) on a pair without text boxes gives
    the same revisions as the default but indents the w:p children of
    word/document.xml and the header differently. That is existing native
    behaviour and is unchanged, so the tests compare tracked text there, not
    bytes.
  • The tests do not assert that the default options refuse the one-sided
    comment pair. Sub-item 161-5 (fix/compare-carries-edited-comment-threads)
    removes that refusal, and an assertion here would break when it lands.
    The ignore_comments=True checks still fail if the keyword is not passed
    through, today by the refusal and after 161-5 by carried comments.
  • API impact: additive keyword-only parameters, no break. rdocx-py is not
    archive-measured, so there is no archive commit. The hash harness does not
    exercise comparison.
  • Left to later PRs, as the plan orders them: the CLI flags (161-2, waiting
    for the fix/cli-output-overwrite-and-closed-pipe branch), word as the CLI
    default (161-3), and every engine change from 161-4 onward.

Tests

Added to crates/rdocx-py/tests/test_core.py, right after
test_priority_word_operations_return_typed_snapshots_and_remain_atomic,
which holds the existing comparison checks:

  • test_compare_granularity_marks_only_the_changed_word: the issue's
    single-run pair. Default and "run" delete and insert the whole run.
    "word" gives exactly one deletion of magna and one insertion of MAGNA,
    and "character" the same text in one wrapper each. Every result accepts to
    the edited text and rejects to the original. The default call's bytes equal
    a call with every default spelled out.
  • test_compare_ignore_options_keep_the_original_side: each ignore flag turns
    a pair that produces revisions by default into one that produces none and
    keeps the original (bold dropped, the double space kept, field result 1
    kept). The one-sided comment pair compares with ignore_comments=True,
    carrying no comment and tracking the body edit. On a pair that edits the
    body and the header, the default tracks both.
    ignored_stories=("header",) leaves word/header1.xml byte-identical to
    the original while the body edit is tracked, and ("body",) leaves
    word/document.xml byte-identical while the header edit is tracked. Each of
    the other five names is accepted and still tracks both edits, which pins
    every name against mapping to the main or header story.
  • test_compare_rejects_unknown_options_before_mutation: granularity="words",
    stories main and table_cell, a duplicated header, a bare string, and a
    positional granularity all raise. The document bytes and a live handle are
    unchanged afterwards. The bare string case checks only for TypeError,
    since its message is pyo3's own wording.

The story tests were checked against three mutations of the mapping (body
to Header, footer to Main, endnote to Header). Each one fails
test_compare_ignore_options_keep_the_original_side.

typing_smoke.py gains a call with every keyword, and two negative calls
(granularity="words", ignored_stories="header") whose
# type: ignore[arg-type] comments --strict proves necessary.

All three tests fail on the unmodified binding with "unexpected keyword
argument" and pass with the change.

Run on this branch:

  • cargo fmt --all --check: clean.
  • cargo clippy -p rdocx-py --all-targets --all-features -- -D warnings: clean.
  • cargo test -p rdocx-py: 2 passed, run with DYLD_LIBRARY_PATH pointing at
    the local libpython. Without it the test binary cannot load
    libpython3.14.dylib, the same environment limit as rpptx-py.
  • pytest crates/rdocx-py/tests (maturin develop build, Python 3.12,
    python-docx 1.2.0): 68 passed, 2 failed. Both failures are in
    test_rendering_threads.py on the pinned pdfinfo 26.01.0 (local 26.09.0),
    a known environment failure. test_python_docx_parity.py passes.
  • mypy --strict on typing_smoke.py and the package: clean.
    mypy.stubtest rdocx: clean.
  • python3 scripts/hash_harness.py --check: 49 entries match.
  • python3 scripts/prose_check.py: 0 violations.
  • python3 scripts/sync_agent_skills.py --check: in sync.

Document.compare in the Python binding always ran the native comparison
with ComparisonOptions::default(), so a Python caller could not choose
word or character granularity or any ignore policy. A one-word edit in
a paragraph held in a single run, the usual shape of a Google Docs
export, was redlined as the whole paragraph deleted and inserted again,
and a pair whose comment parts differ could not be compared at all.

compare now takes keyword-only granularity ("run", "word" or
"character", default "run" like Rust), ignore_formatting,
ignore_whitespace, ignore_fields, ignore_comments and ignored_stories,
whose names follow Story.kind. The options are parsed before the
document is serialised, an unknown value raises RdocxError, and a
duplicated story keeps the native rejection. A call without the new
keywords produces the same bytes as before. The stub documents that
ignore_comments keeps the original's comments and anchors, and the HLD
bindings spec describes the Python surface.

GitHub issues tensorbee#161 and tensorbee#168.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant