Conversation
rebuild_toc and rdocx toc rebuild failed with "root attribute prefix `w` is unbound" as soon as a run of the TOC instruction paragraph, before separate, carried a prefixed attribute such as w:rsidR, w:rsidRPr or w:rsidDel. Word writes w:rsidR and w:rsidRPr on the runs it saves, and Google Docs exports write them on every run, so a table of contents from either producer could not be rebuilt, packed field runs and a TOC inside a block content control included. An attribute under any other prefix the part binds, such as w14, a foreign namespace or a second WordprocessingML alias, failed the same way. The rebuild cuts the instruction paragraph out of document.xml and validated it with CT_P::from_xml, which never sees the declarations of the part. The run root-attribute retention added with the fix for tensorbee#130 resolves each attribute prefix through the explicit bindings of its scope, so it refused the first prefixed attribute on a run of that paragraph. The regression never shipped in a release. The field scan now keeps the bindings the instruction paragraph inherits, adds them to the start tag of the cut-out source and validates it with CT_P::from_xml_fragment, as the projected parse already does for each instruction run. Every run then resolves the prefixes it resolved when the document was read. The validation result is still discarded, so a table of contents that rebuilt before rebuilds to the same bytes. GitHub issue tensorbee#159.
Text-box paragraphs are also parsed out of their part, with the default scope of CT_P::from_xml, which names w as the Word prefix without binding it. Since the run root-attribute retention added with the fix for tensorbee#130, a Word text box whose runs carried w:rsidR or w:rsidRPr lost its shape and text from layout inside mc:AlternateContent and failed the document open as a bare wp:anchor. try_replace_text skipped it without reporting anything and render_template failed on it. CT_Anchor now adds the bindings in scope to the start tag of each text-box paragraph before it parses it, on top of the default scope, so a run attribute under any prefix the part binds resolves for layout. The anchor keeps its raw XML, so saved bytes do not change. rewrite_text_boxes and text_box_sources walk the part with a plain reader that tracks no bindings. For them capture_root_attribute_record now binds a plain Word prefix of its scope to the WordprocessingML namespace when the scope has no explicit binding for it. A plain scope entry only ever names a Word prefix and an explicit binding still wins, so every input that parsed before records the same attributes and saves the same bytes. A run attribute under another prefix still fails those two walkers. GitHub issue tensorbee#159.
The text-box paragraph scope, the Word prefix fallback and their unit tests grow the rdocx-oxml package, and the TOC instruction scope and the TOC and text-box regression tests grow the rdocx package, so the README archive rows of both crates and their ARCHIVE_MEASUREMENTS entries are re-measured, with today as their measurement date. GitHub issue tensorbee#159.
Word and Google Docs write w:rsidR, w:rsidTr, w:rsidRPr, w:rsidDel, w14:paraId and w14:textId on every table row. They survived a save with no edit only because the original bytes were written back. As soon as any modelled edit changed the document, even a one-word replacement outside the table, every row was written with a bare start tag and lost them, while the same attributes on paragraphs, runs and section properties are kept since the fix for tensorbee#130. CT_Row now keeps the attributes of its start tag in the retention record that paragraphs use, at a raw-child position no cell boundary reaches, and writes them back on w:tr. Direct rows, self-closing rows and rows inside a table-level content control all capture it. The record carries no content, so every reader that treats raw row children as content skips it: the comparison row signatures and row boundary check, the retained table layout cache, RowRef::has_unsupported_content and the row diagnostics of the MHTML, ODT, RTF and EPUB writers, and the rich mail merge row-region markers. Without that, an identity-only row difference would refuse the comparison and a Word-saved row would turn off the table cache, raise a spurious export diagnostic or hide a merge region. The rich mail merge region-marker and whole-paragraph fragment checks skip the record of a paragraph the same way. Word writes w:rsidR and w14:paraId on the marker paragraph whenever it writes them on its row, and since the fix for tensorbee#130 that record hid every region marker and fragment field of a document Word saved. The record writes w14:paraId without its declaration, on the assumption that the part root declares w14. A root rdocx wrote does not, and a row may declare the prefix itself, so an edit elsewhere could save a part with an unbound prefix, and a comparison against a copy saved by Word failed. The document, header, footer, note and comment part serializers and the comparison output of every story now declare w14 on the part root when the written content uses it and the root does not. Paragraphs, runs and section properties share the record and gain the same guarantee. GitHub issue tensorbee#159.
Replacement in text boxes walks the raw XML of a part, parses each w:p of a w:txbxContent without its start tag and writes the edited paragraphs back with CT_P::to_xml. A paragraph therefore lost the attributes of its start tag, the w:rsidR, w14:paraId and w14:textId Word writes on every text-box paragraph and any local declaration, as soon as a replacement reached its text box. The walker keeps the start tag each paragraph was read from and writes the edited paragraph under a start tag with the same attributes. The paragraph goes back where it was read, so every prefix those attributes use resolves as before. Parsing the start tag into the retained-attribute record instead would need the bindings in scope, which this walker does not track, and w14:paraId would then fail the whole part. GitHub issue tensorbee#159.
A paragraph copy kept the w14:paraId and w14:textId of its source since the fix for tensorbee#130, and table rows now keep them too. So clone_content, clone_table_row, the region rows of a rich mail merge, the records of a merge into sections and the loops of render_template wrote one w14:paraId several times in one document, and an imported fragment could bring one that the destination already used. Duplicate paragraph identities are invalid, and Word assigns new ones to an element that has none. The identity pass that freshens the bookmark, content-control and drawing identities of copied body content now also removes w14:paraId and w14:textId from every copied w:p and w:tr, text-box paragraphs included. The revision-save identities stay, since Word repeats them freely. Renaming passes, such as the bookmark rename of the TOC rebuild, leave them untouched. The template evaluator copies typed paragraphs and rows without that pass. It now removes the two identities from the retained record of every paragraph and table row a loop renders, through a doc-hidden rdocx-oxml helper that keeps every other retained attribute. GitHub issue tensorbee#159.
The table-row retention record, the text-box start tag, the w14 root declaration and the helper that drops w14 identities from a copy grow the rdocx-oxml package, the layout cache predicate and its test grow the rdocx-layout package, and the consumer filters, the identity drop for copies and template loops and the new regression tests grow the rdocx package. The README archive rows of rdocx-oxml, rdocx-layout and rdocx and their ARCHIVE_MEASUREMENTS entries are re-measured, with today as their measurement date. GitHub issue tensorbee#159.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
CT_Row(rdocx-oxml
table.rs) now keeps the attributes of itsw:trstart tag inthe retention record that paragraphs already use, stored in
extra_xmlatposition
usize::MAX, which no cell boundary reaches, and writes them backon
w:tr. Direct rows, self-closing rows and rows inside a table-levelcontent control all capture it. Every reader that treats raw row children as
content now skips the record: the comparison row signatures and row boundary
check, the retained table layout cache (
table_is_cache_safe),RowRef::has_unsupported_content, the row diagnostics of the MHTML, ODT, RTFand EPUB writers, and the rich mail merge row-region markers. A new
doc-hidden predicate
CT_Row::raw_is_root_attributesmirrorsCT_P::raw_is_root_attributes. HLD 04 and HLD 12 are updated.(rdocx
field.rs) skip the record of a paragraph as well. Since A save drops everyw14:paraId,w14:textIdandw:rsid*attribute fromdocument.xml#130 ithid every
TableStart/TableEndmarker and every fragment field in aparagraph Word wrote, in the body and in region rows alike.
w14when the written content needs it. A newdoc-hidden
declare_w14_on_part_root(rdocx-oxmltext.rs) adds thecanonical declaration to a part root that does not bind
w14when thecontent after it uses the prefix. The document, header, footer, note and
comment part serializers and the comparison output of the main story and
of every related story call it.
rewrite_text_boxes(rdocx-oxmlplaceholder.rs) keeps the start tag eachtext-box paragraph was read from and writes an edited paragraph back under a
start tag with the same attributes, so
w:rsidR,w14:paraId,w14:textIdand any local declaration survive
try_replace_text,replace_regexandrender_templatein text boxes. HLD 04 is extended.w14:paraIdandw14:textIdfrom copied paragraphs and rows. Theidentity pass that freshens bookmark, content-control and drawing identities
of copied body content (rdocx
field.rs) now also removes them from everycopied
w:pandw:tr, text-box paragraphs included, and keeps therevision-save identities. This covers
clone_content,clone_table_row,the region rows of
mail_merge_rich, the records of the merges into sectionsand imported fragments.
render_templatecopies typed paragraphs and rowswithout that pass, so its loops now remove the two identities from every
paragraph and row they render through a new doc-hidden
drop_w14_paragraph_identities(rdocx-oxmltext.rs). HLD 03, HLD 04 andHLD 12 and the rustdoc of the two clone methods are updated.
rdocx-oxml,rdocx-layoutandrdocx, dated 2026-09-27.Part of #159. This closes section 3 (identity attributes on table rows lost
after any modelled edit). The three table-row rows of the identity matrix are
green in every column (output under Tests). Section 1 is #183. What remains
open: section 2, where #184 makes
compare()ignore a content-controlw:idwhile the renumbered
w:tagcase stays open, and section 4, the request tokeep the whole matrix as a test crossing every operation. The regression tests
added here cover the table-row rows of that matrix only.
Why
Word and Google Docs write
w:rsidR,w:rsidTr,w:rsidRPr,w:rsidDel,w14:paraIdandw14:textIdon every table row.CT_Rowhad no carrier forthem and
CT_Row::to_xmlwrote a bare<w:tr>. They survived a save with noedit only because the save writes the original bytes back when the reparsed
model is unchanged. After any modelled edit, even a one-word replacement in a
paragraph outside the table, every row lost them, while the same attributes on
paragraphs, runs and section properties are kept since the fix for #130. The
matrix of #159 shows it as
save=1/3,0/2and0/2(the onew:rsidRleftis the section's).
A carrier alone is not enough. The row signatures of
compare()and its rowboundary check read
row.extra_xmlwhole, so an identity-only difference wouldhave refused the pair ("comparison cannot revise row boundary structures"), and
a Word-saved row would have turned off the table layout cache and raised a
"dropped unsupported Word table row content" or "unmodelled table-row XML"
diagnostic in the MHTML, ODT, RTF and EPUB writers. Each of those is pinned by
a test that fails when its filter is removed.
The rich mail merge checks had the same problem one level down. Word writes
w:rsidRandw14:paraIdon the paragraph that holds a region marker wheneverit writes them on its row, and the marker check rejected any paragraph whose
extra_xmlheld more than whitespace. Since #130 the paragraph record alonemade a region in a Word document fail with "rich mail merge region field
TableStart:... must own its whole block", and a fragment value fail with "must
own its whole paragraph". Filtering the row record without the paragraph
record would only have covered rows written without paragraph identities, which
Word never does.
The record writes
w14:paraIdwithout its declaration, on the assumption thatthe part root declares
w14. Word and python-docx do, but a root rdocx wrotedoes not, and a producer may declare the prefix on the element itself. With
rows now keeping the record, two cases wrote an unbound prefix: an edit
elsewhere in a document whose row declared
w14itself saved a part rdocxcould not reopen, and
compare()of a part rdocx wrote against a copy savedagain by Word refused the whole comparison with "root attribute prefix
w14is unbound" when the copy added a row or paragraph carrying w14 identities.
Paragraphs had the second case already since #130.
Keeping the attributes also makes copies share them. Two paragraphs or rows
with one
w14:paraIdare invalid, and Word assigns new w14 identities to anelement that has none, so copies drop them. Paragraph copies already kept them
since #130, and template loops would now have repeated the row identities too.
Notes
CT_Row::extra_xmlatusize::MAXinstead of a new field, so the public struct literal of thepublished
CT_Rowdoes not change (epub.rsbuilds one, and so maydownstream code). This is the pattern
CT_Palready uses. Behaviour changefor low-level users of
rdocx-oxml: a parsed row'sextra_xmlcan now holdthis record. Code that treats every entry as a raw child should skip it with
CT_Row::raw_is_root_attributes, as the readers changed here do. No publicsignature changes.
w14is declared. The record keeps leavingw14off theelement, so a reopened save stays byte identical to the save it was read
from. The part root gains the declaration instead, only when the root does
not bind
w14and the written content uses the prefix. The check is a bytescan for a
w14:qualified name after the root start tag. Text that happensto match only adds a redundant declaration, never an invalid one. A root that
binds
w14to another URI is left as it is. A part whose root Word orpython-docx wrote serializes as before, and so does every sample of the hash
harness. The header, footer,
note and comment serializers get the same check, because a paragraph there
carries the same record. Before this PR a header paragraph that declared
w14itself was already written with an unbound prefix after areplace_text. A note paragraph with Word identities, and a commentparagraph with
w14:textIdbut now14:paraId, made the serialized partfail its own reparse.
w14:paraIdandw14:textId, and keepw:rsid*, which Word repeats freely. The drop sits inthe shared identity pass, so every path that gives copied body content fresh
identities follows the same rule: the two clone methods, rich merge region
rows, merges into sections (records after the first) and imported fragments.
Template loops follow it through the model-level helper, for every paragraph
and row a loop renders, including a loop with one item. Content outside a
loop keeps its identities. Passes that only rename, such as the bookmark
rename of the TOC rebuild, set nothing and leave the attributes alone. For
the maintainer: allocating fresh values instead of dropping them is possible.
I recommend dropping, since Word regenerates them and the passes stay pure
removals.
that tracks no bindings. Parsing the paragraph start tag into the retention
record would need them, and
w14:paraIdwould then fail the capture. Thedocument layer ignores an error from this walker, so every text box of the
part would silently count zero replacements. Instead the walker writes the
edited paragraph under a start tag that carries the source attributes
verbatim. The paragraph returns to the scope it was read from, so every
prefix they use resolves as before.
rewrite_text_boxestostream paragraphs, and that PR notes that no PR of the batch keeps the start
tag of an edited text-box paragraph. This PR does, through the helper
write_text_box_paragraph. The textual conflict resolves by callingwrite_text_box_paragraph(&mut writer, ¶, ie)?in place ofpara.to_xml(&mut writer)?in the edit branch of Keep every child of a text box that replacement rewrites #180. Thesourcesvectoradded here then goes away. The unit test added here sits after
replace_in_textbox_xml, away from the tests Keep every child of a text box that replacement rewrites #180 adds.tras a namespace owner of the save replay. I tried addingb"tr"to
is_modeled_owner(rdocxdocument.rs). It is not needed for theretained attributes, since the record redeclares every prefix they use except
the canonical
w14, which the part root now declares. It changes anunrelated case, so it is left out. Finding for the maintainer, seen with a
probe and independent of this PR: a namespace declared on a
w:trand usedonly by a raw descendant of the row (for example
<w:tr xmlns:x="urn:x">around a run holding
<x:foo/>) is dropped by a modelled save on main,leaving
xunbound. Withtrinis_modeled_ownerthe declaration isreplayed on the row. That belongs with the namespace owner work of
Document.add_picturefails on a file that has a content control and a default namespace on its root #157.in the raw XML of their drawing and keep their w14 identities, as the drawing
identities there are not freshened either.
xmlns:wthatparagraphs and runs gain after an edit since A save drops every
w14:paraId,w14:textIdandw:rsid*attribute fromdocument.xml#130, which Producer traits: a matrix over every operation, and what still fails in it and around it #160 lists among theproducer traits.
w14under a root that lacks it or clones content with w14 identities.ARCHIVE_REMEASUREMENT_DATES. Therdocxandrdocx-oxmlrows are alsore-measured by Fix the TOC rebuild and text boxes on runs that carry identity attributes #183 and by other PRs of this batch, so they need one fresh
measurement when they are integrated.
Tests
Added:
table.rsunit testrow_root_attributes_survive_serialization_in_source_order: a direct row, arow and a self-closing row inside a table-level content control and a
self-closing direct row, each with
w:rsidR,w:rsidTr,w14:paraIdandw14:textId, keep them in source order after a cell is added, the recordnever appears as child XML, a row-level bookmark keeps its cell boundary,
and a reparse writes the same bytes.
text.rsunit tests:a_part_root_declares_w14_only_when_its_content_needs_it: a document rootwithout
w14gains the canonical declaration once its body usesw14:paraId, and the result parses. A root that already bindsw14, aroot that binds it to another URI and a body whose text merely says
w14are left byte for byte.
dropping_w14_identities_keeps_every_other_retained_attribute: the dropkeeps
w:rsidRand a foreign attribute with its declaration in sourceorder, removes an alias prefix bound to the w14 namespace with the
identities, and removes a record that held nothing else.
header_footer.rsa_paragraph_identity_stays_bound_when_only_the_paragraph_declares_w14,footnotes.rsa_note_paragraph_identity_stays_bound_under_the_written_rootand
comments.rsa_retained_text_identity_stays_bound_without_a_paragraph_identity. Eachwrites a paragraph whose w14 identity the written root would otherwise leave
unbound, and the header root declares
w14while the notes and commentsparts reparse and write the same bytes again.
placeholder.rsunit testreplace_in_textbox_keeps_the_paragraph_start_tag_attributes: plain andregex replacement keep
w:rsidR,w14:paraId,w14:textIdand a localforeign declaration with its attribute on the edited text-box paragraph.
dense_form_caches_are_transactional_bounded_and_exactgains arow that carries identities, which stays cache safe, and turns unsafe once a
real raw child is added.
regression_test.rs,mod table_row_identity_attribute_regressions,placed right after
paragraph_run_and_section_identity_attributes_survive_noop_save:row_identity_attributes_survive_an_edit_elsewhere: the matrix rows"w:rsidR on table rows", "w:rsidTr on table rows" and "w14:paraId on table
rows", plus
w14:textId,w:rsidRPr,w:rsidDeland all six together. Aone-word
try_replace_textin a paragraph outside the table, then save:every occurrence is kept on the two direct rows, the row inside a
table-level content control and a self-closing row, in source order, and
reopen then save is byte identical.
row_identity_attributes_are_not_row_content: no row reports unsupportedcontent, and the MHTML, ODT, RTF and EPUB writers raise no row diagnostic.
row_identity_differences_add_no_comparison_revision: plain againstidentities, identities against plain and identities against new identity
values give 0 revisions. One changed word in a row gives 2, with equal or
with new identity values on the rows.
w14_identities_stay_bound_when_only_their_element_declares_w14: rows anda paragraph that declare
w14themselves under a root that does not keeptheir
w14:paraIdthrough an edit elsewhere, the saved part reopens, anda second save is byte identical.
comparing_with_a_copy_that_declares_w14_keeps_its_identities_bound: theoriginal main part declares no
w14, and its copy declares it on the rootand adds a row whose row and cell paragraphs carry w14 identities. The
comparison succeeds, keeps each identity once and reopens. The same holds
for a header whose table gains such a row.
mail_merge_regions_and_fragments_are_found_in_word_paragraphs_and_rows:Word-shaped input with distinct identities on every paragraph and row and
w:rsidRon every run of the complexMERGEFIELDs. A body region, awhole-paragraph fragment field inside it and a row region all expand. Every
output paragraph and row drops its w14 identities and keeps its
revision-save identities.
copied_rows_and_paragraphs_drop_their_w14_identities: afterclone_table_rowandclone_content, the sources keep every identity,the copies keep
w:rsid*and dropw14:paraIdandw14:textId, on therow and on its cell paragraph.
template_loop_copies_drop_their_w14_identities: a Word-shaped templatewith a body paragraph loop and a row loop, rendered with three items. No
repeated paragraph, row or cell paragraph keeps a w14 identity, every copy
keeps its
w:rsid*, and a row and paragraph outside the loops keep theirs.regression_test.rs, in the text-box module of Fix the TOC rebuild and text boxes on runs that carry identity attributes #183:replacing_text_keeps_the_identities_of_the_text_box_paragraphrunstry_replace_texton a VML text box whose paragraph carries identities.Each new test fails without the change it covers. On the stacked base,
row_identity_attributes_survive_an_edit_elsewherefails. With the rowcarrier but without the reader filters, the unsupported-content and comparison
("comparison cannot revise row boundary structures") tests fail, the export
test fails on each of the ODT, RTF and EPUB diagnostics, and the layout
assertion fails. Without the paragraph filter of the marker check the mail
merge test fails with "must own its whole block", and without it in the
fragment check with "must own its whole paragraph". Without the root
declaration in
CT_Document::to_xmlthe local-declaration test cannot reopenits saved part, and without it in the main or story comparison output the
comparison test fails on the main part or on the header with "root attribute
prefix
w14is unbound". Without it in the header, notes or commentsserializer, the matching unit test fails. With the old text-box write, both
text-box tests fail. With the copy flag off, both copy tests fail. With the
template drop off for body items or for rows, the template test fails on the
repeated body paragraph or on the repeated row.
Identity matrix script of #159 (
matrix_identity_attributes.py, python-docx1.2.0 fixtures, debug build of this branch for the Python module and the CLI).
The three table-row rows on the stacked base:
The whole matrix at the head of this branch:
The two remaining
FAILcells are section 2 of #159, outside this PR.Run on macOS arm64, debug build, at the head of the branch:
cargo fmt --all --checkandcargo clippy -p rdocx-oxml -p rdocx-layout -p rdocx --all-targets --all-features -- -D warnings:clean.
cargo test -p rdocx-oxml: 552 passed, plus 1 doctest.cargo test -p rdocx-layout: 292 passed, plus 1 doctest.cargo test -p rdocx --no-fail-fast: regression 574 passed (7 ignored),doctests 2 passed. Lib 465 passed and integration 314 passed, with only the
environment failures listed below.
cargo test -p rdocx-cli -p rdocx-html -p rdocx-wasm -p rdocx-pdf -p rdocx-opcall pass, and
cargo check --target wasm32-unknown-unknown -p rdocx-wasmisclean.
maturin developbuild of this branch: 66passed, with the two environment failures listed below. No binding or stub
changed, so mypy and stubtest were not rerun.
python3 scripts/hash_harness.py --check: 49 entries match.python3 scripts/prose_check.py: 0 violations.readme_doctests.validate_inventory(): 27 READMEs and 22 packageinventories validated with the re-recorded rows.
Environment failures seen here, the same as on main:
large_word_and_presentation_pdfs_preserve_logical_reading_orderand
word_and_powerpoint_chart_pixels_are_identical(pinned Poppler).odt_reader_matches_pinned_libreoffice_structure,public_authored_theme_and_fonts_match_pinned_word_resolution,section_page_semantics_match_pinned_libreoffice_renderandevery_conditional_table_region_matches_word(pinned tool versions), andsanitized_public_authoring_fixture_passes_every_conformance_stage.test_poppler_pdf_oracle_is_available_at_reviewed_version(pinnedPoppler) and
test_four_concurrent_to_pdf_calls_are_faster_than_serial(timing on a shared machine).