Skip to content

Fix the TOC rebuild and text boxes on runs that carry identity attributes - #183

Open
hadim wants to merge 3 commits into
tensorbee:mainfrom
hadim:fix/toc-rebuild-unbound-word-prefix
Open

hadim wants to merge 3 commits into
tensorbee:mainfrom
hadim:fix/toc-rebuild-unbound-word-prefix

Conversation

@hadim

@hadim hadim commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Parse the TOC instruction paragraph in the scope it inherits. The dynamic
    field scan keeps the bindings the instruction paragraph inherits
    (start_paragraph_namespaces), the validation parse adds them to the start
    tag of the cut-out source and parses it with CT_P::from_xml_fragment. The
    helper that already did this for each projected instruction run now takes
    the binding map, so both parses share it. rebuild_toc,
    Document.rebuild_toc() and rdocx toc rebuild now accept any prefixed
    attribute on the runs of the instruction paragraph before separate:
    w:rsidR, w:rsidRPr, w:rsidDel, w14, a foreign namespace or a second
    WordprocessingML alias, with the field code split or packed into one run,
    bare or inside a block content control. HLD 04 gains one paragraph.
  • Resolve run attribute prefixes in text-box paragraphs. CT_Anchor now adds
    the bindings in scope to the start tag of each text-box paragraph before it
    parses it (text_box_paragraph in drawing.rs), so layout and document open
    resolve any prefix the part binds. capture_root_attribute_record (rdocx-oxml
    text.rs) now binds a plain Word prefix of its scope to the WordprocessingML
    namespace when the scope has no explicit binding for it, which covers w:
    attributes in rewrite_text_boxes and text_box_sources. The HLD 04
    paragraph is extended.
  • Re-record the archive measurements of rdocx-oxml and rdocx, with their
    measurement date.

Part of #159. This fixes section 1 (the TOC rebuild regression) and the same
defect in text boxes, with one remaining gap described in the notes. Section 2
(compare() refuses a content-control w:id or w:tag difference), section 3
(table-row identity attributes lost after a modelled edit) and the matrix
request (159-4) remain open for other PRs.

Why

rebuild_toc failed with root attribute prefix `w` is unbound as soon as
one run of the instruction paragraph before separate carried a prefixed
attribute. Word writes w:rsidR and w:rsidRPr on the runs it saves, and
Google Docs exports write them on every run. So a table of contents from either
producer could not be rebuilt on a fresh open. fixture-report.docx from #158
is one example.

This is a regression from 976611e (F-X128, the fix for #130). s74 and s75
contain it, but the v0.14.0 and py-rdocx-v0.14.0 tags are older. The
regression never shipped in a release, and this PR should land before the
next one.

parse_dynamic_toc_field (field.rs) cuts the instruction paragraph out of
document.xml and validated it with CT_P::from_xml, whose default scope
names w as the Word prefix without binding it and carries none of the
declarations of the part. The run capture added by F-X128 resolves attribute
prefixes only through the explicit bindings of its scope, so it rejected the
first prefixed attribute on any run in that slice.

Three text-box readers parse each paragraph with the same scope:
CT_Anchor::from_xml (drawing.rs), rewrite_text_boxes (placeholder.rs)
and text_box_sources (template.rs). A Word text box whose runs carried
w:rsidRPr failed in a different way at each site:

  • Inside mc:AlternateContent, it lost its shape and text from layout,
    because parse_alternate_content swallows the error.
  • As a bare wp:anchor, it made the document fail to open.
  • try_replace_text skipped it without reporting anything.
  • render_template failed on it.

Notes

  • Two mechanisms, chosen by what the caller knows. The TOC scanner already
    tracks the bindings of every element, and the anchor reader runs on an
    NsReader, so both now pass the real bindings, the way the projected TOC
    parse and the legacy story path (paragraph_fragment_with_bindings) already
    did. rewrite_text_boxes and text_box_sources walk the part with a plain
    Reader that tracks nothing, so for them the capture falls back to binding
    a plain Word prefix. A plain scope entry only ever names a Word prefix:
    word_prefixes_at pushes one only next to a W_NS binding, and
    is_word_element already treats it as Word. The fallback also covers
    external callers of the public CT_P::from_xml and CT_R::from_xml.
  • Remaining gap, still a regression against 0.14.0. A text-box run
    attribute under a prefix other than a Word prefix the default scope names
    (w14, a foreign namespace, or a second WordprocessingML alias) still makes
    try_replace_text skip that part without an error (0 replacements) and
    still makes render_template fail. rdocx 0.14.0 predates the run capture,
    so neither path fails there, and I confirmed the replacement with the PyPI
    build. Layout and document open are fixed for these. As far as I know, Word and Google Docs
    write only w: identity attributes on runs, and those are fixed on every
    path. Closing the gap needs the two walkers to track namespace bindings.
    rewrite_text_boxes is being reworked by
    fix/text-box-rewrite-keeps-all-children and
    fix/table-row-identity-attributes, so passing the scope there fits that
    work better than a third concurrent edit of the same loop. Whether this gap
    blocks the next release is your call.
  • Text boxes now reach the known lossy rewrite path. On main, a text box
    whose runs carry w: identities made replace_many_in_xml_part and
    replace_regex_in_xml_part fail for the whole part, and document.rs
    dropped that error with if let Ok, which accidentally left every text box
    of the part untouched. With this PR the rewrite succeeds, so a hit in any
    text box re-serializes every text box of that part through
    rewrite_text_boxes. That path drops every non-paragraph child of
    w:txbxContent (tables, bookmarks, self-closing paragraphs), which is 160-1e
    and is fixed by fix/text-box-rewrite-keeps-all-children. It also drops the
    paragraph root attributes (w14:paraId, w14:textId, w:rsidR,
    w:rsidRDefault), which is the text-box part of 159-3 and is fixed by
    fix/table-row-identity-attributes. I checked that rdocx 0.14.0 rewrites
    such a text box with the same losses, so this restores the released
    behaviour, but real Word text boxes that main left byte-identical can now
    lose tables, blank lines and paragraph identities. I recommend landing this
    PR together with or after the 160-1e fix. Stacking this branch on that one
    is the alternative if you prefer.
  • Input that already worked is unchanged. The TOC validation result is still
    discarded, so a TOC that rebuilt before rebuilds to the same bytes. The
    anchor keeps its raw XML, so saves do not depend on the text-box parse, and
    its scope is the default scope plus the real bindings. The capture fallback
    applies only to a plain prefix without an explicit binding, so every prefix
    that resolved before still resolves to the same namespace. A unit test
    checks that the default scope gives the same record as a scope that binds
    w. The hash harness matches 49 of 49.
  • The TOC rebuild keeps the original field-run bytes, so it adds no
    declaration, and the regression test now pins that no run start tag gains
    one. A text-box run rewritten by a replacement writes its identity with a
    local xmlns:w, the same redundant declaration an edited body run with a
    retained record already writes, which is the noise Producer traits: a matrix over every operation, and what still fails in it and around it #160 reports. I left that
    to Producer traits: a matrix over every operation, and what still fails in it and around it #160.
  • Bare-anchor text boxes whose runs carry identities now open. Their
    drawing_signature in compare() formats the anchor with Debug, which
    includes the root-attribute records of the text-box runs, so two files that
    differ only by a text-box run rsid now compare as a drawing change. Before,
    such files did not open, so this is not a regression. It belongs with the
    compare-noise work (fix/compare-producer-noise or the 159-2b follow-up).
  • Two related cases stay as they are. An attribute under a prefix declared on
    the instruction paragraph itself now passes the parse, then fails later with
    cannot identify retained `p` nested namespace owner after mutation,
    exactly as in 0.14.0. A part that binds WordprocessingML only to a prefix
    other than w fails the rebuild with table of contents page target _Toc1 was not resolved, with or without any attribute, on main and on this
    branch, while 0.14.0 rebuilds it. That second one is a separate regression
    against 0.14.0 in the page-target resolution and probably deserves its own
    issue.
  • No public API change. CT_P::from_xml and CT_R::from_xml now accept w:
    run identities without a local declaration. Replacement counts go up on
    documents whose text-box runs carry w: identities.
  • With this branch, the toc column of matrix_identity_attributes.py is all
    green. The only cells still failing are the cmp= cells of section 2 and the
    table-row save cells of section 3.

Tests

Added:

  • rdocx regression_test.rs:
    toc_rebuild_accepts_identity_attributes_on_the_instruction_paragraph_runs,
    placed next to the rebuild_toc() refuses a package that declares more than one default style of a type, where every other API accepts it #133 TOC test. It is table-driven over 12 cases: both
    controls, the three w:rsid* attributes on the field runs, one packed field
    run, the field runs in a docPartObj block control, a packed run in that
    control, an identity on a text run before begin, and an attribute under a
    foreign prefix, under w14 and under a second Word alias, each bound on the
    document root. Each case must rebuild 6 entries with no diagnostic, keep
    exactly the attributes of the runs before separate, and add no declaration
    to any run start tag.
  • rdocx regression_test.rs: mod text_box_identity_attribute_regressions,
    placed after the F-X132 retained-namespace mod. Layout for the compatibility
    block and the bare anchor, with w: identities and with a foreign-prefix
    attribute, then try_replace_text and render_template with w:
    identities. Each test compares against the same text box without the
    attributes.
  • rdocx-oxml text.rs unit tests:
    a_run_identity_resolves_the_word_prefix_the_default_scope_names and
    the_default_word_prefix_records_what_its_explicit_binding_records.
  • rdocx-py test_core.py:
    test_rebuild_toc_accepts_identity_attributes_on_the_field_runs.

I ran the new tests without the fixes. On 9a7ed71 each one fails, most with
the unbound-prefix error and the others with its consequence (0 replacements,
or the text box missing from the page text). With the capture fallback alone,
the foreign-prefix TOC cases and the foreign-prefix layout cases still fail.
With the two scope changes alone, the try_replace_text and render_template
tests still fail.

Run on macOS arm64, debug build:

  • cargo fmt --all --check and
    cargo clippy -p rdocx-oxml -p rdocx --all-targets --all-features -- -D warnings:
    clean.
  • cargo test -p rdocx-oxml: 545 passed.
  • cargo test -p rdocx --no-fail-fast:
    • regression: 565 passed.
    • lib and integration: only environment failures (listed below).
  • cargo test -p rdocx-layout -p rdocx-cli -p rdocx-html: all passed.
  • python3 scripts/hash_harness.py --check: 49 entries match.
  • python3 scripts/prose_check.py, commit messages included, and
    python3 scripts/sync_agent_skills.py --check: clean.
  • python3 scripts/readme_doctests.py: passed with the re-recorded rows.
  • python3 -m unittest scripts.test_sprint_workflow: 129 of 131 pass, and the
    two failures are environmental (below).
  • Python binding built with maturin develop:
    • pytest crates/rdocx-py/tests: 66 passed, 2 environment failures.
    • mypy --strict on typing_smoke.py and the package, and
      mypy.stubtest rdocx: clean.

Reproductions with the debug CLI and binding:

  • The issue's section 1 script, extended with the packed, block-control and
    leading-run variants: every case rebuilds 6 entries. The review's e1 (x:foo
    with x bound on the root), e2 (w14:foo) and e6 (wx:rsidRPr with wx
    bound to WordprocessingML) now rebuild 6 entries and keep their attributes,
    as 0.14.0 does.
  • A bare wp:anchor text box with x:foo on its run now opens.
  • rdocx toc rebuild fixture-report.docx rebuilds 21 entries, and in
    workflow_docx.py both TOC steps pass with 21 entries.

Environment failures seen here:

  • rdocx lib: large_word_and_presentation_pdfs_preserve_logical_reading_order
    and word_and_powerpoint_chart_pixels_are_identical. These need the pinned
    Poppler.

  • rdocx integration, pinned LibreOffice 26.2.5.2 against 26.8.0.3 installed
    here:

    • odt_reader_matches_pinned_libreoffice_structure
    • public_authored_theme_and_fonts_match_pinned_word_resolution
    • every_conditional_table_region_matches_word
    • section_page_semantics_match_pinned_libreoffice_render

    The last two fail with the same LibreOffice version assertion on origin/main.

  • rdocx integration: sanitized_public_authoring_fixture_passes_every_conformance_stage,
    which needs an offline cargo run.

  • Python test_rendering_threads.py: two tests fail on the pinned pdfinfo
    version (26.01.0 pinned, 26.09.0 installed here).

  • Policy suite: test_stable_release_family_has_lockstep_preparation_metadata
    (no cargo release here) and
    test_immutable_rdocx_layout_0_10_1_registry_graph_remains_at_oxml_layout_0_6_0
    (no default rustup toolchain in its scratch directory). Both fail the same
    way on an export of origin/main.

rebuild_toc and rdocx toc rebuild failed with "root attribute prefix
`w` is unbound" as soon as a run of the TOC instruction paragraph,
before separate, carried a prefixed attribute such as w:rsidR,
w:rsidRPr or w:rsidDel. Word writes w:rsidR and w:rsidRPr on the runs
it saves, and Google Docs exports write them on every run, so a table
of contents from either producer could not be rebuilt, packed field
runs and a TOC inside a block content control included. An attribute
under any other prefix the part binds, such as w14, a foreign
namespace or a second WordprocessingML alias, failed the same way.

The rebuild cuts the instruction paragraph out of document.xml and
validated it with CT_P::from_xml, which never sees the declarations of
the part. The run root-attribute retention added with the fix for tensorbee#130
resolves each attribute prefix through the explicit bindings of its
scope, so it refused the first prefixed attribute on a run of that
paragraph. The regression never shipped in a release.

The field scan now keeps the bindings the instruction paragraph
inherits, adds them to the start tag of the cut-out source and
validates it with CT_P::from_xml_fragment, as the projected parse
already does for each instruction run. Every run then resolves the
prefixes it resolved when the document was read. The validation
result is still discarded, so a table of contents that rebuilt before
rebuilds to the same bytes.

GitHub issue tensorbee#159.
Text-box paragraphs are also parsed out of their part, with the
default scope of CT_P::from_xml, which names w as the Word prefix
without binding it. Since the run root-attribute retention added with
the fix for tensorbee#130, a Word text box whose runs carried w:rsidR or
w:rsidRPr lost its shape and text from layout inside
mc:AlternateContent and failed the document open as a bare wp:anchor.
try_replace_text skipped it without reporting anything and
render_template failed on it.

CT_Anchor now adds the bindings in scope to the start tag of each
text-box paragraph before it parses it, on top of the default scope,
so a run attribute under any prefix the part binds resolves for
layout. The anchor keeps its raw XML, so saved bytes do not change.

rewrite_text_boxes and text_box_sources walk the part with a plain
reader that tracks no bindings. For them capture_root_attribute_record
now binds a plain Word prefix of its scope to the WordprocessingML
namespace when the scope has no explicit binding for it. A plain scope
entry only ever names a Word prefix and an explicit binding still
wins, so every input that parsed before records the same attributes
and saves the same bytes. A run attribute under another prefix still
fails those two walkers.

GitHub issue tensorbee#159.
The text-box paragraph scope, the Word prefix fallback and their unit
tests grow the rdocx-oxml package, and the TOC instruction scope and
the TOC and text-box regression tests grow the rdocx package, so the
README archive rows of both crates and their ARCHIVE_MEASUREMENTS
entries are re-measured, with today as their measurement date.

GitHub issue tensorbee#159.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant