Conversation
Replacement in text boxes walks the raw XML of the document part and of each header and footer, parses the w:p children of every w:txbxContent and writes those paragraphs back. Every other child was skipped on read and never written. The document layer stores the rewritten part as soon as one text box in it has a hit, so a single hit deleted the tables, block content controls, bookmarks and empty paragraphs of every text box in that part. try_replace_text, replace_all, replace_regex and render_template all go through this walker, for the DrawingML and the VML copy of a text box alike. The walker now edits each paragraph in place and copies every other child through verbatim, so each text box keeps its content in its order. Only a paragraph that the edit changed is re-serialised. The others are copied as read, so they also keep the start-tag attributes that CT_P does not model, such as w:rsidR and w14:paraId. Two defects of the same loop are fixed with it. The closing tag is the one read from the part instead of a fixed w:txbxContent, which did not match a start tag under another prefix. A part that ends inside a text box is now an error instead of an endless read. GitHub issue tensorbee#160.
The text box fix grows the rdocx-oxml sources and the rdocx regression tests, and both are packaged, so the crates.io archive rows of the two READMEs and their ARCHIVE_MEASUREMENTS entries are re-measured. Both rows are dated 2026-09-27 through ARCHIVE_REMEASUREMENT_DATES. GitHub issue tensorbee#160.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
rewrite_text_boxes(rdocx-oxmlplaceholder.rs) now handles eachw:pof a
w:txbxContentwhere it stands. A paragraph that the edit changes isre-serialised in place. A paragraph without a hit is copied through as read.
Every other child is copied through verbatim with
raw_xml::capture_element, or written as read when it is an empty element,whitespace, a comment or a processing instruction. The same loop now closes
the text box with the end tag it read, and it returns an error when the part
ends inside a text box.
rdocx-oxmlandrdocx, dated2026-09-27.
Part of #160. I found this defect while checking the acceptance of section 1
of #160 (every walker reaches content controls at every level). The issue does
not report it. Sections 1 to 4 of #160 remain open for other PRs.
Why
Text-box replacement works on the raw XML of the document part and of each
header and footer. The walker parsed only the
w:pchildren of everyw:txbxContentand wrote only those paragraphs back. It consumed and droppedevery other child element, and it also dropped empty elements, whitespace and
comments. The document layer stores the rewritten part as soon as one text box
in it has a hit. So a single hit deleted the tables, block content controls,
bookmarks and empty paragraphs of every text box in that part, including the
text boxes without a hit, and reported success.
try_replace_text,replace_all,replace_regex,render_templateandrdocx replaceall gothrough this walker, for the DrawingML shape and for its VML copy in
mc:Fallback.Notes
fix/toc-rebuild-unbound-word-prefix. On main, and still onthis branch, a text-box run that carries a
w:attribute fails itsparagraph parse. Word writes
w:rsidRandw:rsidRPron runs by default.CT_P::from_xmlnameswin its default scope without binding it, so theparse returns
root attribute prefix w is unbound. The walker then failsfor the whole part, and the document layer ignores that error (
if let Okin
replace_in_xml_partsandreplace_regex_in_xml_parts). Every text boxof that part is silently left unreplaced, even when the hit is in a sibling
text box without such a run. So the data loss fixed here reached mostly text
boxes without those attributes, as written by python-docx, Google Docs and
LibreOffice. The TOC-prefix PR fixes that parse failure. Once it lands, the
old walker would reach Word-authored text boxes too and delete their other
children. It should therefore land together with this PR or after it,
never alone before it.
children, then handing the paragraphs to the edit callback as a batch, would
also work. The walker instead hands each paragraph to the callback when it
reaches it and writes the result in place. Both callers were per-paragraph
sums (
replace_in_paragraphs,replace_regex_in_paragraphs) with no stateacross paragraphs. For
replace_many_in_xml_part, each paragraph still seesthe pairs in the same order, so the results are the same. The callback of
the private
rewrite_text_boxesbecomesFnMut(&mut CT_P) -> usize. Thereis no public API change.
count.
replace_in_paragraphandreplace_regex_in_paragraphtouch nothingbefore their first match, so a zero count means an unchanged paragraph, and
its captured bytes are written instead. This was suggested in review. It
keeps the start-tag attributes that
CT_Pdoes not model (w:rsidR,w14:paraId,w14:textId, local namespace declarations) on everyparagraph without a hit, in the edited text box and in its siblings.
a replacement hits at least one text box. Every text box of such a part is
walked, with or without a hit of its own:
other child now survive, byte for byte and in document order.
re-serialised and lost their start-tag attributes.
It was dropped before.
</w:txbxContent>, which did not match a start tag under another prefix.Edited paragraphs are still written with the
wprefix, though. So atext box in a part that binds WordprocessingML only as the default
namespace stays ill-formed after an edit, as it is on main. That is left
to the namespace prefix work.
Eofagain and again with elements still open, and the old inner loopignored it, so the call never returned. The new loop needs an
Eofarm inany case. The document layer already ignores an error from this walker and
leaves the part untouched.
Before, one nested in a table or a content control of the outer text box
was deleted with it, and one nested in a paragraph was not edited either.
CT_P::from_xml, and anedited one is written with
CT_P::to_xml. The walker reads past thew:pstart tag before parsing, so an edited paragraph loses its start-tag
attributes (rsid,
w14:paraId,w14:textId). No PR of this batchaddresses that. How an edited paragraph serialises otherwise is unchanged
here (
xml:spaceis infix/compare-producer-noise).replace_in_paragraphandreplace_regex_in_paragraphdrop every run whosemodelled
contentis empty. A run whose only child is a text box inmc:AlternateContentor aw:pictshape keeps that child inextra_xml,so it is dropped too. A hit in a body paragraph therefore deletes a text box
anchored in the same paragraph, and a hit in a text-box paragraph deletes a
shape run next to it. main behaves the same. This belongs with the run
removal of the parallel PR on
fix/replace-reaches-content-controls, whichshould remove only the runs that the replacement itself emptied and that
carry no raw children (
extra_xml,alt_drawings).text box is still not replaced. It is now kept instead of deleted. Reaching
it needs a typed parse that re-serialises the table or the control, which
belongs with replacement inside content controls (section 1 of Producer traits: a matrix over every operation, and what still fails in it and around it #160) and
with typed replacement of block content in headers and text boxes. No
content-control replacement falls out of this change for free. An inline
content control in a text-box paragraph behaves as it does in the body,
which is the gap that section 1 of Producer traits: a matrix over every operation, and what still fails in it and around it #160 describes.
placeholder.rs, away from thetext.rschangeof
fix/toc-rebuild-unbound-word-prefix.fix/compare-producer-noisealsoedits
placeholder.rs, in the run-splitting functions and at another spotof the test module.
ARCHIVE_REMEASUREMENT_DATES, as the latest re-measurements on main do.The
rdocxrow is also re-measured by the other PRs of this batch that growrdocx, so it needs one fresh measurement when they are integrated.replacement hits.
Tests
Added:
placeholder.rsunit tests:replace_in_textbox_keeps_every_other_child_in_order: a VML text boxwhose two placeholder paragraphs sit among a paragraph without a hit, a
table, a block content control, a bookmark pair around an empty
paragraph, an XML comment and whitespace, then a second text box with a
table and a paragraph but no placeholder. Both paragraphs without a hit
carry
w:rsidR,w14:paraIdandw14:textId. Both the plain and theregex walker must return the input part with only the placeholders
replaced, so the second text box and the attributes come back unchanged.
replace_in_textbox_closes_it_with_its_own_prefix: aq:txbxContentcomes back well formed and otherwise unchanged.
replace_in_unterminated_textbox_is_an_error: two truncated parts returnan error instead of hanging.
regression_test.rs:mod text_box_replacement_keeps_every_child,placed next to the replacement termination tests. It runs
try_replace_text,replace_regexandrender_templateover two forms: theDrawingML shape in
mc:Choicewith its VML copy inmc:Fallback, and abare
wp:anchor. Each text box holds a paragraph with the placeholder, atable, a block content control and a bookmark pair around an empty
paragraph. Each test checks the count, and that every saved text box holds
the edited paragraph and every other child, byte for byte and in order.
Each new test that can run on 9a7ed71 failed there without the fix: the
three regression tests and the first two unit tests. The truncated-input unit
test was not run there, since the old loop never returns. With the fix in
place but paragraphs without a hit re-serialised as before, the first unit
test fails on the dropped attributes.
End to end, with the debug CLI: a python-docx script builds both text box
forms with the same children, plus a paragraph without a hit that carries
w:rsidRandw14:paraId, and runsrdocx replace -p {{name}} -v Ada. Onthe old walker, every saved text box held the edited paragraph alone. With
this branch, both text boxes keep every child, the attributed paragraph comes
back byte for byte, and python-docx reopens the output. A probe on this
branch confirms the parse failure described in the first note: a text box
run with
w:rsidRmakesreplace_in_xml_partreturn the unbound prefixerror, also when the hit is in a sibling text box.
Run on macOS arm64, debug build, at the head of the branch:
cargo fmt --all --checkandcargo clippy -p rdocx-oxml -p rdocx --all-targets --all-features -- -D warnings:clean.
cargo test -p rdocx-oxml: 546 passed, plus 1 doctest.cargo test -p rdocx --no-fail-fast: regression 564 passed, doctests 2passed. Lib and integration have only environment failures (listed below).
cargo test -p rdocx-cli: 19 passed.python3 scripts/hash_harness.py --check: 49 entries match.python3 scripts/prose_check.py: 0 violations.readme_doctests.validate_inventory(): 27 READMEs and 22 packageinventories validated with the re-recorded rows.
Environment failures seen here:
rdocx lib:
large_word_and_presentation_pdfs_preserve_logical_reading_orderand
word_and_powerpoint_chart_pixels_are_identical. These need the pinnedPoppler.
rdocx integration, pinned LibreOffice 26.2.5.2 against 26.8.0.3 installed
here:
odt_reader_matches_pinned_libreoffice_structurepublic_authored_theme_and_fonts_match_pinned_word_resolutionevery_conditional_table_region_matches_wordsection_page_semantics_match_pinned_libreoffice_renderI confirmed that the last two fail the same way on origin/main.
rdocx integration:
sanitized_public_authoring_fixture_passes_every_conformance_stage,which needs an offline
cargo run.