Conversation
Document::text, and so plain `rdocx text`, skipped every content control. A body-level control, which Google Docs writes around whole paragraphs with a goog_rdk tag, dropped its paragraph from the output. Cell controls, row controls around cells and table controls around rows dropped theirs from the row line. Paragraphs of a table nested in a cell were skipped as well, although the doc comment promises table cell text. The walker now recurses through controls at every level and builds each row line with the existing visit_row visitor. A body paragraph still ends with a newline, a table row is still one line, and every paragraph of its cells, nested tables included, ends with a tab. Documents without content controls or nested tables give exactly the text they gave before. GitHub issue tensorbee#160.
Document::images, word_count, headings and links skipped content controls, and so did document_outline and the accessibility audit built on them. A picture or a heading that Google Docs wrapped in a goog_rdk control was not reported, and word_count missed the words of every wrapped paragraph. images now walks visit_all_drawings and word_count walks visit_body_paragraphs, the existing visitors that interleave body, table, row, cell and inline controls in document order. headings and links keep their scope, body paragraphs without table cells, and now collect it through body-level and nested controls. A table that such a control wraps is still not searched, so wrapping content in a control never changes what they report. The private collectors of images and word_count are removed. Results for documents without content controls are unchanged. MHTML export pairs the <img> tags of the HTML emitter with picture sizes by position and refuses a count mismatch. The emitter still drops content controls, so sizes taken from the new images() would make the export fail. The writer now collects them from the paragraphs the emitter reaches, which is the reach images() had before. A picture in a content control is still dropped with the existing loss diagnostic, as it was before this change. GitHub issue tensorbee#160.
The read walker fix and its regression tests grow the rdocx package, and the new `rdocx text` integration test grows the rdocx-cli package, so both README rows and their ARCHIVE_MEASUREMENTS entries are re-measured. GitHub issue tensorbee#160.
This was referenced Sep 27, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Document::text, and so plainrdocx text, now reads the paragraphs thatcontent controls wrap at every level: body, table (around rows), row (around
cells), cell, and nested controls. It also reads the paragraphs of a table
nested in a cell. The format is unchanged: a body paragraph ends with a
newline, a table row is one line, and every paragraph of its cells ends with
a tab.
Document::imagesnow walksvisit_all_drawingsandDocument::word_countwalks
visit_body_paragraphs, so pictures and words inside body, table, row,cell and inline controls are reported. The private collectors
collect_images_from_*andword_count_in_*are removed.Document::headings(and sodocument_outline,audit_accessibilityandthe heading gap warning of
rdocx validate) andDocument::linksnow readthe paragraphs of body-level and nested controls through a small private
collector,
block_paragraphs. Their scope is unchanged: body paragraphs,never table cells, whether or not a control wraps the table.
to_mhtml_bytes) takes its picture sizes from a privatecollector in
html.rs,emitted_paragraphs, over the paragraphs therdocx-html emitter reaches, instead of from
Document::images. Its output isthe same as on main: a picture in a content control is still dropped with
the existing loss diagnostic.
WASM sentence now names
Document::textinstead of "that additive facadeaccessor", whose antecedent the new sentences had moved away.
Part of #160. This covers section 1 for the Rust read walkers only. Section 1 still needs the replace and regex
walkers, header and footer blocks, text boxes and the Markdown, HTML and EPUB
exporters (see Notes). Sections 2 to 5 are other PRs.
Why
Google Docs wraps runs, and whole paragraphs including headings, in
goog_rdkcontent controls, and Word keeps them. The read walkers matched
BodyContentby hand and dropped
ContentControl(BodyContent::ContentControl(_) => {}intext,collect_images_from_contentandword_count_in_content, andCellContent::ContentControl(_) => {}in the cell loops).headingsandlinkslooked at direct body paragraphs only. So a paragraph wrapped at bodylevel vanished from
rdocx text, and wrapped pictures, headings and wordsvanished from
images,headingsandword_count, whileParagraph.text,Document::paragraphsandrdocx text --jsonalready showed them.Notes
imagesandword_countare thin closures over the existingvisitors, which already interleave table, row and inline controls by index.
textkeeps a small walker of its own because its separator depends oncontext (a newline at block level, a tab inside a row). Each row line is
built with
visit_row, and the table-level interleave follows thevisit_tablepattern. No public signature changes.to
text, inside the outer row's line, each ending with a tab. This is theonly output change for documents without content controls. For a python-docx
cell.add_tablecell the line goes from\t\t\nto\tinner gamma\t\t\n.headingsandlinks. The rule this PR keeps is that a contentcontrol is transparent: wrapping content in one never changes what these
walkers report, and results for documents without controls are unchanged.
block_paragraphswalks body paragraphs and the paragraphs of body-leveland nested controls, and does not enter tables.
CT_Body::paragraphs, theset behind
Document::paragraphs, was not reused because it also returnsthe cells of a table that a body-level control wraps. Reading it would
report a heading in a wrapped table but not in the same table unwrapped, and
so change
rdocx validatewarnings depending only on the wrapper.reported, with or without a control.
rebuild_tocsource discovery(
field.rs,collect_body_paragraphs) does include table cells, as Worddoes, and
visit_body_paragraphswould serve both walkers.That is a one-line change in each walker plus the matrix expectations, but
it changes
headings,document_outline,audit_accessibility,rdocx validateandlinksfor documents without controls, so it is leftout of this PR.
<img>tags with picture sizes byposition and refuses a count mismatch. The emitter still drops content
controls, so sizes taken from the new
images()would fail the export with"document contains images the HTML emitter did not serialize". The private
emitted_paragraphskeeps the reachimages()had on main, and is meant togo away once the rdocx-html exporter follow-up reads content controls. It
deliberately keeps one pre-existing quirk of that reach: the emitter skips
vertical merge continuation cells, the collector does not, so a picture in
such a cell still fails the export with the count error, as on main. Skipping
those cells would make the export succeed and drop the picture with no loss
diagnostic, which belongs with the exporter follow-up.
inline control inside a hyperlink are kept as opaque XML by the paragraph
parser, so
Paragraph.text,text()andlinks()all miss their text. Thisis a model limitation in
rdocx-oxml, not a walker one.rdocx text --jsonomits rows wrapped by table-level controls and cellswrapped by row-level controls, because
collect_table_paragraphsandcollect_row_paragraphsincommands.rsiterate onlyrowsandcells.Plain
rdocx textnow shows them, so the two views disagree on those twolocations. rdocx-cli sources are not touched here because
rdocx convert --to pdf|md|html,rpptx convert --to pdfandrpptx thumbnailoverwrite any existing output, the input included #156 andrdocxandrpptxCLIs panic on a closed standard output (| head) #166 arerewriting
commands.rs.the run anchor remapping of 160-1f, a later medium-risk PR), the text-box
rewrite (160-1e, branch
fix/text-box-rewrite-keeps-all-children), headerand footer block content (160-1d), and the exporters (
rdocx-htmlmarkdown.rsandemitter.rs,rdocxepub.rs) as a follow-up. The legacyinsert_tocstill collects direct body paragraphs only.Tests
crates/rdocx/tests/regression_test.rs, new modulecontent_control_read_walker_regressionsnext to the ordered reader tests:a_content_control_hides_nothing_from_any_body_read_walkeris the matrix.Ten locations (body control, nested body control, inline control, nested
inline control, cell control, nested cell control, row control around a
cell, table control around a row, control in a nested table, body control
around a table) against five walkers (
text,images,word_count,headings,links). At each location all five results must equal thoseof the same document without the control, the probe's text, picture and
words must be reported, and its heading and link exactly at the four
body-level locations.
text_reads_every_content_control_location_in_document_orderchecks theexact text of one document mixing every location.
nested_table_cells_contribute_text_inside_their_outer_rowpins thatformat.
crates/rdocx/src/html.rs, unit testmhtml_writer_sizes_only_the_pictures_the_html_emitter_serializes: picturesin a body control, outside any control and in an inline control.
images()reports three,
to_mhtml_bytessucceeds with the two existing lossdiagnostics, and the reopened MHTML holds only the outside picture at its
own size.
crates/rdocx-cli/tests/integration.rs,text_prints_paragraphs_wrapped_by_a_body_content_control, next to theexisting plain text test:
rdocx textprints a body-level control'sparagraph.
there on
images().len() == 3, since main reports one picture.9a7ed71 and of this branch.
rdocx text: no control and inline controlprint the same text, a body control goes from nothing to its paragraph, a
cell control from an empty row to
cell beta\t, a nested table from\t\tto
\tinner gamma\t\t, and a wrapped heading now prints.rdocx validateafter a Heading 1: a Heading 3 wrapped in a body control now gives the
heading gap warning, and a Heading 3 in a table cell gives none, with or
without a body control around the table, on both builds.
cargo fmt --all --check,cargo clippy -p rdocx -p rdocx-cli --all-targets --all-features -- -D warnings,cargo test -p rdocx --no-fail-fast,cargo test -p rdocx-cli,cargo test -p rdocx-wasm(the consumer oftext()),python3 scripts/hash_harness.py --check(49 entries match),python3 scripts/prose_check.py(0 violations)and
python3 scripts/readme_doctests.py. All pass except theenvironment-only failures below. No Python binding or stub changed, and
rdocx-py binds none of the changed walkers.
large_word_and_presentation_pdfs_preserve_logical_reading_orderandword_and_powerpoint_chart_pixels_are_identical(pinned pdftotext andrasterizer). rdocx integration:
odt_reader_matches_pinned_libreoffice_structure,public_authored_theme_and_fonts_match_pinned_word_resolution,sanitized_public_authoring_fixture_passes_every_conformance_stage, and twomore that fail on their first assertion, the
soffice --versionpin (local26.8.0.3, pinned 26.2.5.2):
section_page_semantics_match_pinned_libreoffice_renderand
every_conditional_table_region_matches_word. Those two failidentically on a clean 9a7ed71.