Skip to content

Keep table row identity attributes through modelled edits - #193

Open
hadim wants to merge 7 commits into
tensorbee:mainfrom
hadim:fix/table-row-identity-attributes
Open

hadim wants to merge 7 commits into
tensorbee:mainfrom
hadim:fix/table-row-identity-attributes

Conversation

@hadim

@hadim hadim commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Stacked on #183 (fix/toc-rebuild-unbound-word-prefix). The first three
commits of this branch are #183 and drop out once it merges. The four commits
after them are this PR.

Summary

  • Keep identity attributes on table rows through modelled edits. CT_Row
    (rdocx-oxml table.rs) now keeps the attributes of its w:tr start tag in
    the retention record that paragraphs already use, stored in extra_xml at
    position usize::MAX, which no cell boundary reaches, and writes them back
    on w:tr. Direct rows, self-closing rows and rows inside a table-level
    content control all capture it. Every reader that treats raw row children as
    content now skips the record: the comparison row signatures and row boundary
    check, the retained table layout cache (table_is_cache_safe),
    RowRef::has_unsupported_content, the row diagnostics of the MHTML, ODT, RTF
    and EPUB writers, and the rich mail merge row-region markers. A new
    doc-hidden predicate CT_Row::raw_is_root_attributes mirrors
    CT_P::raw_is_root_attributes. HLD 04 and HLD 12 are updated.
  • In the same commit, two consequences of keeping the record:
    • The rich mail merge region-marker and whole-paragraph fragment checks
      (rdocx field.rs) skip the record of a paragraph as well. Since A save drops every w14:paraId, w14:textId and w:rsid* attribute from document.xml #130 it
      hid every TableStart/TableEnd marker and every fragment field in a
      paragraph Word wrote, in the body and in region rows alike.
    • The part root declares w14 when the written content needs it. A new
      doc-hidden declare_w14_on_part_root (rdocx-oxml text.rs) adds the
      canonical declaration to a part root that does not bind w14 when the
      content after it uses the prefix. The document, header, footer, note and
      comment part serializers and the comparison output of the main story and
      of every related story call it.
  • Keep the start tag of a text-box paragraph that replacement edits.
    rewrite_text_boxes (rdocx-oxml placeholder.rs) keeps the start tag each
    text-box paragraph was read from and writes an edited paragraph back under a
    start tag with the same attributes, so w:rsidR, w14:paraId, w14:textId
    and any local declaration survive try_replace_text, replace_regex and
    render_template in text boxes. HLD 04 is extended.
  • Drop w14:paraId and w14:textId from copied paragraphs and rows. The
    identity pass that freshens bookmark, content-control and drawing identities
    of copied body content (rdocx field.rs) now also removes them from every
    copied w:p and w:tr, text-box paragraphs included, and keeps the
    revision-save identities. This covers clone_content, clone_table_row,
    the region rows of mail_merge_rich, the records of the merges into sections
    and imported fragments. render_template copies typed paragraphs and rows
    without that pass, so its loops now remove the two identities from every
    paragraph and row they render through a new doc-hidden
    drop_w14_paragraph_identities (rdocx-oxml text.rs). HLD 03, HLD 04 and
    HLD 12 and the rustdoc of the two clone methods are updated.
  • Re-record the archive measurements of rdocx-oxml, rdocx-layout and
    rdocx, dated 2026-09-27.

Part of #159. This closes section 3 (identity attributes on table rows lost
after any modelled edit). The three table-row rows of the identity matrix are
green in every column (output under Tests). Section 1 is #183. What remains
open: section 2, where #184 makes compare() ignore a content-control w:id
while the renumbered w:tag case stays open, and section 4, the request to
keep the whole matrix as a test crossing every operation. The regression tests
added here cover the table-row rows of that matrix only.

Why

Word and Google Docs write w:rsidR, w:rsidTr, w:rsidRPr, w:rsidDel,
w14:paraId and w14:textId on every table row. CT_Row had no carrier for
them and CT_Row::to_xml wrote a bare <w:tr>. They survived a save with no
edit only because the save writes the original bytes back when the reparsed
model is unchanged. After any modelled edit, even a one-word replacement in a
paragraph outside the table, every row lost them, while the same attributes on
paragraphs, runs and section properties are kept since the fix for #130. The
matrix of #159 shows it as save=1/3, 0/2 and 0/2 (the one w:rsidR left
is the section's).

A carrier alone is not enough. The row signatures of compare() and its row
boundary check read row.extra_xml whole, so an identity-only difference would
have refused the pair ("comparison cannot revise row boundary structures"), and
a Word-saved row would have turned off the table layout cache and raised a
"dropped unsupported Word table row content" or "unmodelled table-row XML"
diagnostic in the MHTML, ODT, RTF and EPUB writers. Each of those is pinned by
a test that fails when its filter is removed.

The rich mail merge checks had the same problem one level down. Word writes
w:rsidR and w14:paraId on the paragraph that holds a region marker whenever
it writes them on its row, and the marker check rejected any paragraph whose
extra_xml held more than whitespace. Since #130 the paragraph record alone
made a region in a Word document fail with "rich mail merge region field
TableStart:... must own its whole block", and a fragment value fail with "must
own its whole paragraph". Filtering the row record without the paragraph
record would only have covered rows written without paragraph identities, which
Word never does.

The record writes w14:paraId without its declaration, on the assumption that
the part root declares w14. Word and python-docx do, but a root rdocx wrote
does not, and a producer may declare the prefix on the element itself. With
rows now keeping the record, two cases wrote an unbound prefix: an edit
elsewhere in a document whose row declared w14 itself saved a part rdocx
could not reopen, and compare() of a part rdocx wrote against a copy saved
again by Word refused the whole comparison with "root attribute prefix w14
is unbound" when the copy added a row or paragraph carrying w14 identities.
Paragraphs had the second case already since #130.

Keeping the attributes also makes copies share them. Two paragraphs or rows
with one w14:paraId are invalid, and Word assigns new w14 identities to an
element that has none, so copies drop them. Paragraph copies already kept them
since #130, and template loops would now have repeated the row identities too.

Notes

  • Decision, carrier shape. The record lives in CT_Row::extra_xml at
    usize::MAX instead of a new field, so the public struct literal of the
    published CT_Row does not change (epub.rs builds one, and so may
    downstream code). This is the pattern CT_P already uses. Behaviour change
    for low-level users of rdocx-oxml: a parsed row's extra_xml can now hold
    this record. Code that treats every entry as a raw child should skip it with
    CT_Row::raw_is_root_attributes, as the readers changed here do. No public
    signature changes.
  • Decision, where w14 is declared. The record keeps leaving w14 off the
    element, so a reopened save stays byte identical to the save it was read
    from. The part root gains the declaration instead, only when the root does
    not bind w14 and the written content uses the prefix. The check is a byte
    scan for a w14: qualified name after the root start tag. Text that happens
    to match only adds a redundant declaration, never an invalid one. A root that
    binds w14 to another URI is left as it is. A part whose root Word or
    python-docx wrote serializes as before, and so does every sample of the hash
    harness. The header, footer,
    note and comment serializers get the same check, because a paragraph there
    carries the same record. Before this PR a header paragraph that declared
    w14 itself was already written with an unbound prefix after a
    replace_text. A note paragraph with Word identities, and a comment
    paragraph with w14:textId but no w14:paraId, made the serialized part
    fail its own reparse.
  • Decision, copies. Cloned paragraphs and rows drop w14:paraId and
    w14:textId, and keep w:rsid*, which Word repeats freely. The drop sits in
    the shared identity pass, so every path that gives copied body content fresh
    identities follows the same rule: the two clone methods, rich merge region
    rows, merges into sections (records after the first) and imported fragments.
    Template loops follow it through the model-level helper, for every paragraph
    and row a loop renders, including a loop with one item. Content outside a
    loop keeps its identities. Passes that only rename, such as the bookmark
    rename of the TOC rebuild, set nothing and leave the attributes alone. For
    the maintainer: allocating fresh values instead of dropping them is possible.
    I recommend dropping, since Word regenerates them and the passes stay pure
    removals.
  • Decision, text boxes. The replacement walker reads a part with a plain reader
    that tracks no bindings. Parsing the paragraph start tag into the retention
    record would need them, and w14:paraId would then fail the capture. The
    document layer ignores an error from this walker, so every text box of the
    part would silently count zero replacements. Instead the walker writes the
    edited paragraph under a start tag that carries the source attributes
    verbatim. The paragraph returns to the scope it was read from, so every
    prefix they use resolves as before.
  • Merge note with Keep every child of a text box that replacement rewrites #180. Keep every child of a text box that replacement rewrites #180 rewrites the same loop of rewrite_text_boxes to
    stream paragraphs, and that PR notes that no PR of the batch keeps the start
    tag of an edited text-box paragraph. This PR does, through the helper
    write_text_box_paragraph. The textual conflict resolves by calling
    write_text_box_paragraph(&mut writer, &para, ie)? in place of
    para.to_xml(&mut writer)? in the edit branch of Keep every child of a text box that replacement rewrites #180. The sources vector
    added here then goes away. The unit test added here sits after
    replace_in_textbox_xml, away from the tests Keep every child of a text box that replacement rewrites #180 adds.
  • Not done, tr as a namespace owner of the save replay. I tried adding b"tr"
    to is_modeled_owner (rdocx document.rs). It is not needed for the
    retained attributes, since the record redeclares every prefix they use except
    the canonical w14, which the part root now declares. It changes an
    unrelated case, so it is left out. Finding for the maintainer, seen with a
    probe and independent of this PR: a namespace declared on a w:tr and used
    only by a raw descendant of the row (for example <w:tr xmlns:x="urn:x">
    around a run holding <x:foo/>) is dropped by a modelled save on main,
    leaving x unbound. With tr in is_modeled_owner the declaration is
    replayed on the row. That belongs with the namespace owner work of Document.add_picture fails on a file that has a content control and a default namespace on its root #157.
  • Left open. Text-box paragraphs inside content a template loop repeats stay
    in the raw XML of their drawing and keep their w14 identities, as the drawing
    identities there are not freshened either.
  • Left as on main. An edited row carries the redundant local xmlns:w that
    paragraphs and runs gain after an edit since A save drops every w14:paraId, w14:textId and w:rsid* attribute from document.xml #130, which Producer traits: a matrix over every operation, and what still fails in it and around it #160 lists among the
    producer traits.
  • The hash harness matches 49 of 49. No sample carries a row identity, uses
    w14 under a root that lacks it or clones content with w14 identities.
  • The three re-measured archive rows are dated 2026-09-27 through
    ARCHIVE_REMEASUREMENT_DATES. The rdocx and rdocx-oxml rows are also
    re-measured by Fix the TOC rebuild and text boxes on runs that carry identity attributes #183 and by other PRs of this batch, so they need one fresh
    measurement when they are integrated.

Tests

Added:

  • rdocx-oxml table.rs unit test
    row_root_attributes_survive_serialization_in_source_order: a direct row, a
    row and a self-closing row inside a table-level content control and a
    self-closing direct row, each with w:rsidR, w:rsidTr, w14:paraId and
    w14:textId, keep them in source order after a cell is added, the record
    never appears as child XML, a row-level bookmark keeps its cell boundary,
    and a reparse writes the same bytes.
  • rdocx-oxml text.rs unit tests:
    • a_part_root_declares_w14_only_when_its_content_needs_it: a document root
      without w14 gains the canonical declaration once its body uses
      w14:paraId, and the result parses. A root that already binds w14, a
      root that binds it to another URI and a body whose text merely says w14
      are left byte for byte.
    • dropping_w14_identities_keeps_every_other_retained_attribute: the drop
      keeps w:rsidR and a foreign attribute with its declaration in source
      order, removes an alias prefix bound to the w14 namespace with the
      identities, and removes a record that held nothing else.
  • rdocx-oxml unit tests of the other part serializers:
    header_footer.rs
    a_paragraph_identity_stays_bound_when_only_the_paragraph_declares_w14,
    footnotes.rs a_note_paragraph_identity_stays_bound_under_the_written_root
    and comments.rs
    a_retained_text_identity_stays_bound_without_a_paragraph_identity. Each
    writes a paragraph whose w14 identity the written root would otherwise leave
    unbound, and the header root declares w14 while the notes and comments
    parts reparse and write the same bytes again.
  • rdocx-oxml placeholder.rs unit test
    replace_in_textbox_keeps_the_paragraph_start_tag_attributes: plain and
    regex replacement keep w:rsidR, w14:paraId, w14:textId and a local
    foreign declaration with its attribute on the edited text-box paragraph.
  • rdocx-layout dense_form_caches_are_transactional_bounded_and_exact gains a
    row that carries identities, which stays cache safe, and turns unsafe once a
    real raw child is added.
  • rdocx regression_test.rs, mod table_row_identity_attribute_regressions,
    placed right after paragraph_run_and_section_identity_attributes_survive_noop_save:
    • row_identity_attributes_survive_an_edit_elsewhere: the matrix rows
      "w:rsidR on table rows", "w:rsidTr on table rows" and "w14:paraId on table
      rows", plus w14:textId, w:rsidRPr, w:rsidDel and all six together. A
      one-word try_replace_text in a paragraph outside the table, then save:
      every occurrence is kept on the two direct rows, the row inside a
      table-level content control and a self-closing row, in source order, and
      reopen then save is byte identical.
    • row_identity_attributes_are_not_row_content: no row reports unsupported
      content, and the MHTML, ODT, RTF and EPUB writers raise no row diagnostic.
    • row_identity_differences_add_no_comparison_revision: plain against
      identities, identities against plain and identities against new identity
      values give 0 revisions. One changed word in a row gives 2, with equal or
      with new identity values on the rows.
    • w14_identities_stay_bound_when_only_their_element_declares_w14: rows and
      a paragraph that declare w14 themselves under a root that does not keep
      their w14:paraId through an edit elsewhere, the saved part reopens, and
      a second save is byte identical.
    • comparing_with_a_copy_that_declares_w14_keeps_its_identities_bound: the
      original main part declares no w14, and its copy declares it on the root
      and adds a row whose row and cell paragraphs carry w14 identities. The
      comparison succeeds, keeps each identity once and reopens. The same holds
      for a header whose table gains such a row.
    • mail_merge_regions_and_fragments_are_found_in_word_paragraphs_and_rows:
      Word-shaped input with distinct identities on every paragraph and row and
      w:rsidR on every run of the complex MERGEFIELDs. A body region, a
      whole-paragraph fragment field inside it and a row region all expand. Every
      output paragraph and row drops its w14 identities and keeps its
      revision-save identities.
    • copied_rows_and_paragraphs_drop_their_w14_identities: after
      clone_table_row and clone_content, the sources keep every identity,
      the copies keep w:rsid* and drop w14:paraId and w14:textId, on the
      row and on its cell paragraph.
    • template_loop_copies_drop_their_w14_identities: a Word-shaped template
      with a body paragraph loop and a row loop, rendered with three items. No
      repeated paragraph, row or cell paragraph keeps a w14 identity, every copy
      keeps its w:rsid*, and a row and paragraph outside the loops keep theirs.
  • rdocx regression_test.rs, in the text-box module of Fix the TOC rebuild and text boxes on runs that carry identity attributes #183:
    replacing_text_keeps_the_identities_of_the_text_box_paragraph runs
    try_replace_text on a VML text box whose paragraph carries identities.

Each new test fails without the change it covers. On the stacked base,
row_identity_attributes_survive_an_edit_elsewhere fails. With the row
carrier but without the reader filters, the unsupported-content and comparison
("comparison cannot revise row boundary structures") tests fail, the export
test fails on each of the ODT, RTF and EPUB diagnostics, and the layout
assertion fails. Without the paragraph filter of the marker check the mail
merge test fails with "must own its whole block", and without it in the
fragment check with "must own its whole paragraph". Without the root
declaration in CT_Document::to_xml the local-declaration test cannot reopen
its saved part, and without it in the main or story comparison output the
comparison test fails on the main part or on the header with "root attribute
prefix w14 is unbound". Without it in the header, notes or comments
serializer, the matching unit test fails. With the old text-box write, both
text-box tests fail. With the copy flag off, both copy tests fail. With the
template drop off for body items or for rows, the template test fails on the
repeated body paragraph or on the repeated row.

Identity matrix script of #159 (matrix_identity_attributes.py, python-docx
1.2.0 fixtures, debug build of this branch for the Python module and the CLI).
The three table-row rows on the stacked base:

w:rsidR on table rows                    save=1/3    replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidTr on table rows                   save=0/2    replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w14:paraId on table rows                 save=0/2    replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2

The whole matrix at the head of this branch:

control: no attribute                    save=-      replace=ok     toc=ok     fields=ok     render=ok     cmp==-      cmp+1=2
w:rsidR on paragraphs                    save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidRDefault on paragraphs             save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidP on paragraphs                    save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w14:paraId on paragraphs                 save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w14:textId on paragraphs                 save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidR on runs                          save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidRPr on runs                        save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidDel on runs                        save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidR on field runs                    save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidRPr on field runs                  save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidR on footer field runs             save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidRPr on footer field runs           save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:id on content control                  save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==FAIL   cmp+1=2
w:tag changed on content control         save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==FAIL   cmp+1=2
w:rsidR on table rows                    save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w:rsidTr on table rows                   save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2
w14:paraId on table rows                 save=ok     replace=ok     toc=ok     fields=ok     render=ok     cmp==0      cmp+1=2

The two remaining FAIL cells are section 2 of #159, outside this PR.

Run on macOS arm64, debug build, at the head of the branch:

  • cargo fmt --all --check and
    cargo clippy -p rdocx-oxml -p rdocx-layout -p rdocx --all-targets --all-features -- -D warnings:
    clean.
  • cargo test -p rdocx-oxml: 552 passed, plus 1 doctest.
  • cargo test -p rdocx-layout: 292 passed, plus 1 doctest.
  • cargo test -p rdocx --no-fail-fast: regression 574 passed (7 ignored),
    doctests 2 passed. Lib 465 passed and integration 314 passed, with only the
    environment failures listed below.
  • Consumers: cargo test -p rdocx-cli -p rdocx-html -p rdocx-wasm -p rdocx-pdf -p rdocx-opc
    all pass, and cargo check --target wasm32-unknown-unknown -p rdocx-wasm is
    clean.
  • rdocx-py pytest suite against a maturin develop build of this branch: 66
    passed, with the two environment failures listed below. No binding or stub
    changed, so mypy and stubtest were not rerun.
  • python3 scripts/hash_harness.py --check: 49 entries match.
  • python3 scripts/prose_check.py: 0 violations.
  • readme_doctests.validate_inventory(): 27 READMEs and 22 package
    inventories validated with the re-recorded rows.

Environment failures seen here, the same as on main:

  • rdocx lib: large_word_and_presentation_pdfs_preserve_logical_reading_order
    and word_and_powerpoint_chart_pixels_are_identical (pinned Poppler).
  • rdocx integration: odt_reader_matches_pinned_libreoffice_structure,
    public_authored_theme_and_fonts_match_pinned_word_resolution,
    section_page_semantics_match_pinned_libreoffice_render and
    every_conditional_table_region_matches_word (pinned tool versions), and
    sanitized_public_authoring_fixture_passes_every_conformance_stage.
  • rdocx-py: test_poppler_pdf_oracle_is_available_at_reviewed_version (pinned
    Poppler) and test_four_concurrent_to_pdf_calls_are_faster_than_serial
    (timing on a shared machine).

rebuild_toc and rdocx toc rebuild failed with "root attribute prefix
`w` is unbound" as soon as a run of the TOC instruction paragraph,
before separate, carried a prefixed attribute such as w:rsidR,
w:rsidRPr or w:rsidDel. Word writes w:rsidR and w:rsidRPr on the runs
it saves, and Google Docs exports write them on every run, so a table
of contents from either producer could not be rebuilt, packed field
runs and a TOC inside a block content control included. An attribute
under any other prefix the part binds, such as w14, a foreign
namespace or a second WordprocessingML alias, failed the same way.

The rebuild cuts the instruction paragraph out of document.xml and
validated it with CT_P::from_xml, which never sees the declarations of
the part. The run root-attribute retention added with the fix for tensorbee#130
resolves each attribute prefix through the explicit bindings of its
scope, so it refused the first prefixed attribute on a run of that
paragraph. The regression never shipped in a release.

The field scan now keeps the bindings the instruction paragraph
inherits, adds them to the start tag of the cut-out source and
validates it with CT_P::from_xml_fragment, as the projected parse
already does for each instruction run. Every run then resolves the
prefixes it resolved when the document was read. The validation
result is still discarded, so a table of contents that rebuilt before
rebuilds to the same bytes.

GitHub issue tensorbee#159.
Text-box paragraphs are also parsed out of their part, with the
default scope of CT_P::from_xml, which names w as the Word prefix
without binding it. Since the run root-attribute retention added with
the fix for tensorbee#130, a Word text box whose runs carried w:rsidR or
w:rsidRPr lost its shape and text from layout inside
mc:AlternateContent and failed the document open as a bare wp:anchor.
try_replace_text skipped it without reporting anything and
render_template failed on it.

CT_Anchor now adds the bindings in scope to the start tag of each
text-box paragraph before it parses it, on top of the default scope,
so a run attribute under any prefix the part binds resolves for
layout. The anchor keeps its raw XML, so saved bytes do not change.

rewrite_text_boxes and text_box_sources walk the part with a plain
reader that tracks no bindings. For them capture_root_attribute_record
now binds a plain Word prefix of its scope to the WordprocessingML
namespace when the scope has no explicit binding for it. A plain scope
entry only ever names a Word prefix and an explicit binding still
wins, so every input that parsed before records the same attributes
and saves the same bytes. A run attribute under another prefix still
fails those two walkers.

GitHub issue tensorbee#159.
The text-box paragraph scope, the Word prefix fallback and their unit
tests grow the rdocx-oxml package, and the TOC instruction scope and
the TOC and text-box regression tests grow the rdocx package, so the
README archive rows of both crates and their ARCHIVE_MEASUREMENTS
entries are re-measured, with today as their measurement date.

GitHub issue tensorbee#159.
Word and Google Docs write w:rsidR, w:rsidTr, w:rsidRPr, w:rsidDel,
w14:paraId and w14:textId on every table row. They survived a save
with no edit only because the original bytes were written back. As
soon as any modelled edit changed the document, even a one-word
replacement outside the table, every row was written with a bare
start tag and lost them, while the same attributes on paragraphs,
runs and section properties are kept since the fix for tensorbee#130.

CT_Row now keeps the attributes of its start tag in the retention
record that paragraphs use, at a raw-child position no cell boundary
reaches, and writes them back on w:tr. Direct rows, self-closing rows
and rows inside a table-level content control all capture it. The
record carries no content, so every reader that treats raw row
children as content skips it: the comparison row signatures and row
boundary check, the retained table layout cache,
RowRef::has_unsupported_content and the row diagnostics of the
MHTML, ODT, RTF and EPUB writers, and the rich mail merge row-region
markers. Without that, an identity-only row difference would refuse
the comparison and a Word-saved row would turn off the table cache,
raise a spurious export diagnostic or hide a merge region.

The rich mail merge region-marker and whole-paragraph fragment checks
skip the record of a paragraph the same way. Word writes w:rsidR and
w14:paraId on the marker paragraph whenever it writes them on its
row, and since the fix for tensorbee#130 that record hid every region marker
and fragment field of a document Word saved.

The record writes w14:paraId without its declaration, on the
assumption that the part root declares w14. A root rdocx wrote does
not, and a row may declare the prefix itself, so an edit elsewhere
could save a part with an unbound prefix, and a comparison against a
copy saved by Word failed. The document, header, footer, note and
comment part serializers and the comparison output of every story now
declare w14 on the part root when the written content uses it and the
root does not. Paragraphs, runs and section properties share the
record and gain the same guarantee.

GitHub issue tensorbee#159.
Replacement in text boxes walks the raw XML of a part, parses each
w:p of a w:txbxContent without its start tag and writes the edited
paragraphs back with CT_P::to_xml. A paragraph therefore lost the
attributes of its start tag, the w:rsidR, w14:paraId and w14:textId
Word writes on every text-box paragraph and any local declaration,
as soon as a replacement reached its text box.

The walker keeps the start tag each paragraph was read from and
writes the edited paragraph under a start tag with the same
attributes. The paragraph goes back where it was read, so every
prefix those attributes use resolves as before. Parsing the start tag
into the retained-attribute record instead would need the bindings
in scope, which this walker does not track, and w14:paraId would
then fail the whole part.

GitHub issue tensorbee#159.
A paragraph copy kept the w14:paraId and w14:textId of its source
since the fix for tensorbee#130, and table rows now keep them too. So
clone_content, clone_table_row, the region rows of a rich mail merge,
the records of a merge into sections and the loops of render_template
wrote one w14:paraId several times in one document, and an imported
fragment could bring one that the destination already used. Duplicate
paragraph identities are invalid, and Word assigns new ones to an
element that has none.

The identity pass that freshens the bookmark, content-control and
drawing identities of copied body content now also removes
w14:paraId and w14:textId from every copied w:p and w:tr, text-box
paragraphs included. The revision-save identities stay, since Word
repeats them freely. Renaming passes, such as the bookmark rename of
the TOC rebuild, leave them untouched.

The template evaluator copies typed paragraphs and rows without that
pass. It now removes the two identities from the retained record of
every paragraph and table row a loop renders, through a doc-hidden
rdocx-oxml helper that keeps every other retained attribute.

GitHub issue tensorbee#159.
The table-row retention record, the text-box start tag, the w14 root
declaration and the helper that drops w14 identities from a copy grow
the rdocx-oxml package, the layout cache predicate and its test grow
the rdocx-layout package, and the consumer filters, the identity drop
for copies and template loops and the new regression tests grow the
rdocx package. The README archive rows of rdocx-oxml, rdocx-layout
and rdocx and their ARCHIVE_MEASUREMENTS entries are re-measured, with
today as their measurement date.

GitHub issue tensorbee#159.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant