Conversation
rebuild_toc and rdocx toc rebuild failed with "root attribute prefix `w` is unbound" as soon as a run of the TOC instruction paragraph, before separate, carried a prefixed attribute such as w:rsidR, w:rsidRPr or w:rsidDel. Word writes w:rsidR and w:rsidRPr on the runs it saves, and Google Docs exports write them on every run, so a table of contents from either producer could not be rebuilt, packed field runs and a TOC inside a block content control included. An attribute under any other prefix the part binds, such as w14, a foreign namespace or a second WordprocessingML alias, failed the same way. The rebuild cuts the instruction paragraph out of document.xml and validated it with CT_P::from_xml, which never sees the declarations of the part. The run root-attribute retention added with the fix for tensorbee#130 resolves each attribute prefix through the explicit bindings of its scope, so it refused the first prefixed attribute on a run of that paragraph. The regression never shipped in a release. The field scan now keeps the bindings the instruction paragraph inherits, adds them to the start tag of the cut-out source and validates it with CT_P::from_xml_fragment, as the projected parse already does for each instruction run. Every run then resolves the prefixes it resolved when the document was read. The validation result is still discarded, so a table of contents that rebuilt before rebuilds to the same bytes. GitHub issue tensorbee#159.
Text-box paragraphs are also parsed out of their part, with the default scope of CT_P::from_xml, which names w as the Word prefix without binding it. Since the run root-attribute retention added with the fix for tensorbee#130, a Word text box whose runs carried w:rsidR or w:rsidRPr lost its shape and text from layout inside mc:AlternateContent and failed the document open as a bare wp:anchor. try_replace_text skipped it without reporting anything and render_template failed on it. CT_Anchor now adds the bindings in scope to the start tag of each text-box paragraph before it parses it, on top of the default scope, so a run attribute under any prefix the part binds resolves for layout. The anchor keeps its raw XML, so saved bytes do not change. rewrite_text_boxes and text_box_sources walk the part with a plain reader that tracks no bindings. For them capture_root_attribute_record now binds a plain Word prefix of its scope to the WordprocessingML namespace when the scope has no explicit binding for it. A plain scope entry only ever names a Word prefix and an explicit binding still wins, so every input that parsed before records the same attributes and saves the same bytes. A run attribute under another prefix still fails those two walkers. GitHub issue tensorbee#159.
The text-box paragraph scope, the Word prefix fallback and their unit tests grow the rdocx-oxml package, and the TOC instruction scope and the TOC and text-box regression tests grow the rdocx package, so the README archive rows of both crates and their ARCHIVE_MEASUREMENTS entries are re-measured, with today as their measurement date. GitHub issue tensorbee#159.
This was referenced Sep 27, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
field scan keeps the bindings the instruction paragraph inherits
(
start_paragraph_namespaces), the validation parse adds them to the starttag of the cut-out source and parses it with
CT_P::from_xml_fragment. Thehelper that already did this for each projected instruction run now takes
the binding map, so both parses share it.
rebuild_toc,Document.rebuild_toc()andrdocx toc rebuildnow accept any prefixedattribute on the runs of the instruction paragraph before
separate:w:rsidR,w:rsidRPr,w:rsidDel,w14, a foreign namespace or a secondWordprocessingML alias, with the field code split or packed into one run,
bare or inside a block content control. HLD 04 gains one paragraph.
CT_Anchornow addsthe bindings in scope to the start tag of each text-box paragraph before it
parses it (
text_box_paragraphindrawing.rs), so layout and document openresolve any prefix the part binds.
capture_root_attribute_record(rdocx-oxmltext.rs) now binds a plain Word prefix of its scope to the WordprocessingMLnamespace when the scope has no explicit binding for it, which covers
w:attributes in
rewrite_text_boxesandtext_box_sources. The HLD 04paragraph is extended.
rdocx-oxmlandrdocx, with theirmeasurement date.
Part of #159. This fixes section 1 (the TOC rebuild regression) and the same
defect in text boxes, with one remaining gap described in the notes. Section 2
(
compare()refuses a content-controlw:idorw:tagdifference), section 3(table-row identity attributes lost after a modelled edit) and the matrix
request (159-4) remain open for other PRs.
Why
rebuild_tocfailed withroot attribute prefix `w` is unboundas soon asone run of the instruction paragraph before
separatecarried a prefixedattribute. Word writes
w:rsidRandw:rsidRPron the runs it saves, andGoogle Docs exports write them on every run. So a table of contents from either
producer could not be rebuilt on a fresh open.
fixture-report.docxfrom #158is one example.
This is a regression from 976611e (F-X128, the fix for #130). s74 and s75
contain it, but the
v0.14.0andpy-rdocx-v0.14.0tags are older. Theregression never shipped in a release, and this PR should land before the
next one.
parse_dynamic_toc_field(field.rs) cuts the instruction paragraph out ofdocument.xmland validated it withCT_P::from_xml, whose default scopenames
was the Word prefix without binding it and carries none of thedeclarations of the part. The run capture added by F-X128 resolves attribute
prefixes only through the explicit bindings of its scope, so it rejected the
first prefixed attribute on any run in that slice.
Three text-box readers parse each paragraph with the same scope:
CT_Anchor::from_xml(drawing.rs),rewrite_text_boxes(placeholder.rs)and
text_box_sources(template.rs). A Word text box whose runs carriedw:rsidRPrfailed in a different way at each site:mc:AlternateContent, it lost its shape and text from layout,because
parse_alternate_contentswallows the error.wp:anchor, it made the document fail to open.try_replace_textskipped it without reporting anything.render_templatefailed on it.Notes
tracks the bindings of every element, and the anchor reader runs on an
NsReader, so both now pass the real bindings, the way the projected TOCparse and the legacy story path (
paragraph_fragment_with_bindings) alreadydid.
rewrite_text_boxesandtext_box_sourceswalk the part with a plainReaderthat tracks nothing, so for them the capture falls back to bindinga plain Word prefix. A plain scope entry only ever names a Word prefix:
word_prefixes_atpushes one only next to a W_NS binding, andis_word_elementalready treats it as Word. The fallback also coversexternal callers of the public
CT_P::from_xmlandCT_R::from_xml.attribute under a prefix other than a Word prefix the default scope names
(
w14, a foreign namespace, or a second WordprocessingML alias) still makestry_replace_textskip that part without an error (0 replacements) andstill makes
render_templatefail. rdocx 0.14.0 predates the run capture,so neither path fails there, and I confirmed the replacement with the PyPI
build. Layout and document open are fixed for these. As far as I know, Word and Google Docs
write only
w:identity attributes on runs, and those are fixed on everypath. Closing the gap needs the two walkers to track namespace bindings.
rewrite_text_boxesis being reworked byfix/text-box-rewrite-keeps-all-childrenandfix/table-row-identity-attributes, so passing the scope there fits thatwork better than a third concurrent edit of the same loop. Whether this gap
blocks the next release is your call.
whose runs carry
w:identities madereplace_many_in_xml_partandreplace_regex_in_xml_partfail for the whole part, anddocument.rsdropped that error with
if let Ok, which accidentally left every text boxof the part untouched. With this PR the rewrite succeeds, so a hit in any
text box re-serializes every text box of that part through
rewrite_text_boxes. That path drops every non-paragraph child ofw:txbxContent(tables, bookmarks, self-closing paragraphs), which is 160-1eand is fixed by
fix/text-box-rewrite-keeps-all-children. It also drops theparagraph root attributes (
w14:paraId,w14:textId,w:rsidR,w:rsidRDefault), which is the text-box part of 159-3 and is fixed byfix/table-row-identity-attributes. I checked that rdocx 0.14.0 rewritessuch a text box with the same losses, so this restores the released
behaviour, but real Word text boxes that main left byte-identical can now
lose tables, blank lines and paragraph identities. I recommend landing this
PR together with or after the 160-1e fix. Stacking this branch on that one
is the alternative if you prefer.
discarded, so a TOC that rebuilt before rebuilds to the same bytes. The
anchor keeps its raw XML, so saves do not depend on the text-box parse, and
its scope is the default scope plus the real bindings. The capture fallback
applies only to a plain prefix without an explicit binding, so every prefix
that resolved before still resolves to the same namespace. A unit test
checks that the default scope gives the same record as a scope that binds
w. The hash harness matches 49 of 49.declaration, and the regression test now pins that no run start tag gains
one. A text-box run rewritten by a replacement writes its identity with a
local
xmlns:w, the same redundant declaration an edited body run with aretained record already writes, which is the noise Producer traits: a matrix over every operation, and what still fails in it and around it #160 reports. I left that
to Producer traits: a matrix over every operation, and what still fails in it and around it #160.
drawing_signatureincompare()formats the anchor withDebug, whichincludes the root-attribute records of the text-box runs, so two files that
differ only by a text-box run rsid now compare as a drawing change. Before,
such files did not open, so this is not a regression. It belongs with the
compare-noise work (
fix/compare-producer-noiseor the 159-2b follow-up).the instruction paragraph itself now passes the parse, then fails later with
cannot identify retained `p` nested namespace owner after mutation,exactly as in 0.14.0. A part that binds WordprocessingML only to a prefix
other than
wfails the rebuild withtable of contents page target _Toc1 was not resolved, with or without any attribute, on main and on thisbranch, while 0.14.0 rebuilds it. That second one is a separate regression
against 0.14.0 in the page-target resolution and probably deserves its own
issue.
CT_P::from_xmlandCT_R::from_xmlnow acceptw:run identities without a local declaration. Replacement counts go up on
documents whose text-box runs carry
w:identities.toccolumn ofmatrix_identity_attributes.pyis allgreen. The only cells still failing are the
cmp=cells of section 2 and thetable-row
savecells of section 3.Tests
Added:
regression_test.rs:toc_rebuild_accepts_identity_attributes_on_the_instruction_paragraph_runs,placed next to the
rebuild_toc()refuses a package that declares more than one default style of a type, where every other API accepts it #133 TOC test. It is table-driven over 12 cases: bothcontrols, the three
w:rsid*attributes on the field runs, one packed fieldrun, the field runs in a
docPartObjblock control, a packed run in thatcontrol, an identity on a text run before
begin, and an attribute under aforeign prefix, under
w14and under a second Word alias, each bound on thedocument root. Each case must rebuild 6 entries with no diagnostic, keep
exactly the attributes of the runs before
separate, and add no declarationto any run start tag.
regression_test.rs:mod text_box_identity_attribute_regressions,placed after the F-X132 retained-namespace mod. Layout for the compatibility
block and the bare anchor, with
w:identities and with a foreign-prefixattribute, then
try_replace_textandrender_templatewithw:identities. Each test compares against the same text box without the
attributes.
text.rsunit tests:a_run_identity_resolves_the_word_prefix_the_default_scope_namesandthe_default_word_prefix_records_what_its_explicit_binding_records.test_core.py:test_rebuild_toc_accepts_identity_attributes_on_the_field_runs.I ran the new tests without the fixes. On 9a7ed71 each one fails, most with
the unbound-prefix error and the others with its consequence (0 replacements,
or the text box missing from the page text). With the capture fallback alone,
the foreign-prefix TOC cases and the foreign-prefix layout cases still fail.
With the two scope changes alone, the
try_replace_textandrender_templatetests still fail.
Run on macOS arm64, debug build:
cargo fmt --all --checkandcargo clippy -p rdocx-oxml -p rdocx --all-targets --all-features -- -D warnings:clean.
cargo test -p rdocx-oxml: 545 passed.cargo test -p rdocx --no-fail-fast:cargo test -p rdocx-layout -p rdocx-cli -p rdocx-html: all passed.python3 scripts/hash_harness.py --check: 49 entries match.python3 scripts/prose_check.py, commit messages included, andpython3 scripts/sync_agent_skills.py --check: clean.python3 scripts/readme_doctests.py: passed with the re-recorded rows.python3 -m unittest scripts.test_sprint_workflow: 129 of 131 pass, and thetwo failures are environmental (below).
maturin develop:pytest crates/rdocx-py/tests: 66 passed, 2 environment failures.mypy --strictontyping_smoke.pyand the package, andmypy.stubtest rdocx: clean.Reproductions with the debug CLI and binding:
leading-run variants: every case rebuilds 6 entries. The review's e1 (
x:foowith
xbound on the root), e2 (w14:foo) and e6 (wx:rsidRPrwithwxbound to WordprocessingML) now rebuild 6 entries and keep their attributes,
as 0.14.0 does.
wp:anchortext box withx:fooon its run now opens.rdocx toc rebuild fixture-report.docxrebuilds 21 entries, and inworkflow_docx.pyboth TOC steps pass with 21 entries.Environment failures seen here:
rdocx lib:
large_word_and_presentation_pdfs_preserve_logical_reading_orderand
word_and_powerpoint_chart_pixels_are_identical. These need the pinnedPoppler.
rdocx integration, pinned LibreOffice 26.2.5.2 against 26.8.0.3 installed
here:
odt_reader_matches_pinned_libreoffice_structurepublic_authored_theme_and_fonts_match_pinned_word_resolutionevery_conditional_table_region_matches_wordsection_page_semantics_match_pinned_libreoffice_renderThe last two fail with the same LibreOffice version assertion on origin/main.
rdocx integration:
sanitized_public_authoring_fixture_passes_every_conformance_stage,which needs an offline
cargo run.Python
test_rendering_threads.py: two tests fail on the pinnedpdfinfoversion (26.01.0 pinned, 26.09.0 installed here).
Policy suite:
test_stable_release_family_has_lockstep_preparation_metadata(no
cargo releasehere) andtest_immutable_rdocx_layout_0_10_1_registry_graph_remains_at_oxml_layout_0_6_0(no default rustup toolchain in its scratch directory). Both fail the same
way on an export of origin/main.