Conversation
Document.compare in the Python binding always ran the native comparison
with ComparisonOptions::default(), so a Python caller could not choose
word or character granularity or any ignore policy. A one-word edit in
a paragraph held in a single run, the usual shape of a Google Docs
export, was redlined as the whole paragraph deleted and inserted again,
and a pair whose comment parts differ could not be compared at all.
compare now takes keyword-only granularity ("run", "word" or
"character", default "run" like Rust), ignore_formatting,
ignore_whitespace, ignore_fields, ignore_comments and ignored_stories,
whose names follow Story.kind. The options are parsed before the
document is serialised, an unknown value raises RdocxError, and a
duplicated story keeps the native rejection. A call without the new
keywords produces the same bytes as before. The stub documents that
ignore_comments keeps the original's comments and anchors, and the HLD
bindings spec describes the Python surface.
GitHub issues tensorbee#161 and tensorbee#168.
This was referenced Sep 27, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
rdocx-py:Document.compare(edited, author, timestamp, *, granularity="run", ignore_formatting=False, ignore_whitespace=False, ignore_fields=False, ignore_comments=False, ignored_stories=None)builds anative
ComparisonOptionsand callscompare_with_options. Two privateparsers next to
parse_raster_formatmap"run","word"and"character"toComparisonGranularity, and theStory.kindnamesbody,header,footer,comment,text_box,footnoteandendnotetoComparisonStoryKind. The options are parsed inside the detached closurebefore the document is serialised, so a bad value fails before any work.
Literaltypes for both vocabularies,Sequence[...] | Noneforignored_stories, and a docstring that states the left-biased policy andthat
ignore_comments=Truekeeps the original's comments and anchors.Python only preserves comparison output.
Part of #161. This is sub-item 1, the Python half. The
rdocx compareflags,wordas the CLI default and the engine changes (the edited side's commentthreads, the rebuilt TOC, granular edits next to a hyperlink) remain.
Part of #168. This closes the "Comparison options" checklist item (item G).
Why
Document.comparein Python always ranComparisonOptions::default(). Withthe default whole-run granularity, the issue's one-word edit in a single-run
paragraph (the usual shape of a Google Docs export) redlines as all 183
characters deleted and inserted again. There was also no way to compare a pair
whose comment parts differ, because only
ignore_commentsgets past "documentcomparison requires identical related-story shells". The native API already
had all of this, the binding just did not pass it through.
Notes
"run"as thedefault like Rust, story names from
Story.kind, andignore_commentsexposed now with its left bias documented in the stub and the HLD. A call
without the new keywords is the same native call as before, and a test pins
that its bytes equal those of a call with every default spelled out.
ignored_storiesdefaults toNonerather than the()the issue sketched.pyo3 renders only literal defaults into
__text_signature__, so aVec::new()default shows as..., andmypy.stubtestthen rejects a stubdefault of
().Nonefollowsrender_pages(pages=None)andupdate_fields(merge_fields=None). Tuples, lists and()are accepted. Abare
stror asetraisesTypeErrorfrom pyo3, and mypy rejects both.RdocxError, throughrdocx::Error::Other, like anunknown
render_pages(format=...). The binding also has aValueErrorprecedent (
story_item_kind_from_name). I choseRdocxErrorso every optionerror has one type, since the native duplicate-story check already reports
RdocxError. Duplicates are left to that native validation. Switching toValueErroris a two-line change if you prefer it.table_cellis aStory.kindvalue but not a comparison category, so it isrejected as unknown.
bodymaps toComparisonStoryKind::Main.ignored_storiesasSequence[Literal[...]], so mypy catches a misspelt story, andtyping_smoke.pypins that.Story.kindis typedstr, though, soignored_stories=[s.kind for s in doc.stories if ...]works at runtime butfails
mypy --strictwitharg-type. NarrowingStory.kindto aLiteralwould not fix that on its own, because it also yields
table_cell, which isnot a comparison category. I kept the
Literal. Widening toSequence[str]is a one-line stub change if you prefer the pass-through idiom.
ignored_stories=["comment"]does not replaceignore_comments=True. On thetest pair (a comment added on the edited side only), ignoring the comment
story excludes the comments part but the body anchors are still compared,
and the call fails with "comparison cannot revise paragraph boundary
structures at body/paragraph[0]".
ignore_commentsalso skips commentreferences and ranges (
ignored_run_contentand thecomment_rangeschecks in
comparison.rs). That is the native contract. The HLD now says so,and the test uses
ignore_comments.w:fldSimplewhoseresult changes gets
w:delandw:insinside thew:fldSimple, and neitherDocument.revisionsnoraccept_all()sees them (the list is empty andaccept_all()returns 0, leaving the markup in place). A complex field withthe same change is listed and resolved. So the test uses a complex field for
ignore_fields. This looks like a revision reader gap inrdocx-oxmland maydeserve its own issue.
ignore_formatting=Trueon aformatting-only pair records no revision but still re-serialises
word/document.xml(the indentation between elements changes), so thebinding's before and after byte check advances the handle revision. In the
same way,
ignored_stories=("text_box",)on a pair without text boxes givesthe same revisions as the default but indents the
w:pchildren ofword/document.xmland the header differently. That is existing nativebehaviour and is unchanged, so the tests compare tracked text there, not
bytes.
comment pair. Sub-item 161-5 (
fix/compare-carries-edited-comment-threads)removes that refusal, and an assertion here would break when it lands.
The
ignore_comments=Truechecks still fail if the keyword is not passedthrough, today by the refusal and after 161-5 by carried comments.
rdocx-pyis notarchive-measured, so there is no archive commit. The hash harness does not
exercise comparison.
for the
fix/cli-output-overwrite-and-closed-pipebranch),wordas the CLIdefault (161-3), and every engine change from 161-4 onward.
Tests
Added to
crates/rdocx-py/tests/test_core.py, right aftertest_priority_word_operations_return_typed_snapshots_and_remain_atomic,which holds the existing comparison checks:
test_compare_granularity_marks_only_the_changed_word: the issue'ssingle-run pair. Default and
"run"delete and insert the whole run."word"gives exactly one deletion ofmagnaand one insertion ofMAGNA,and
"character"the same text in one wrapper each. Every result accepts tothe edited text and rejects to the original. The default call's bytes equal
a call with every default spelled out.
test_compare_ignore_options_keep_the_original_side: each ignore flag turnsa pair that produces revisions by default into one that produces none and
keeps the original (bold dropped, the double space kept, field result
1kept). The one-sided comment pair compares with
ignore_comments=True,carrying no comment and tracking the body edit. On a pair that edits the
body and the header, the default tracks both.
ignored_stories=("header",)leavesword/header1.xmlbyte-identical tothe original while the body edit is tracked, and
("body",)leavesword/document.xmlbyte-identical while the header edit is tracked. Each ofthe other five names is accepted and still tracks both edits, which pins
every name against mapping to the main or header story.
test_compare_rejects_unknown_options_before_mutation:granularity="words",stories
mainandtable_cell, a duplicatedheader, a bare string, and apositional granularity all raise. The document bytes and a live handle are
unchanged afterwards. The bare string case checks only for
TypeError,since its message is pyo3's own wording.
The story tests were checked against three mutations of the mapping (
bodyto
Header,footertoMain,endnotetoHeader). Each one failstest_compare_ignore_options_keep_the_original_side.typing_smoke.pygains a call with every keyword, and two negative calls(
granularity="words",ignored_stories="header") whose# type: ignore[arg-type]comments--strictproves necessary.All three tests fail on the unmodified binding with "unexpected keyword
argument" and pass with the change.
Run on this branch:
cargo fmt --all --check: clean.cargo clippy -p rdocx-py --all-targets --all-features -- -D warnings: clean.cargo test -p rdocx-py: 2 passed, run withDYLD_LIBRARY_PATHpointing atthe local libpython. Without it the test binary cannot load
libpython3.14.dylib, the same environment limit asrpptx-py.pytest crates/rdocx-py/tests(maturin develop build, Python 3.12,python-docx 1.2.0): 68 passed, 2 failed. Both failures are in
test_rendering_threads.pyon the pinnedpdfinfo26.01.0 (local 26.09.0),a known environment failure.
test_python_docx_parity.pypasses.mypy --strictontyping_smoke.pyand the package: clean.mypy.stubtest rdocx: clean.python3 scripts/hash_harness.py --check: 49 entries match.python3 scripts/prose_check.py: 0 violations.python3 scripts/sync_agent_skills.py --check: in sync.