Skip to content

Make the body read walkers reach content controls - #177

Open
hadim wants to merge 3 commits into
tensorbee:mainfrom
hadim:fix/body-readers-reach-content-controls
Open

hadim wants to merge 3 commits into
tensorbee:mainfrom
hadim:fix/body-readers-reach-content-controls

Conversation

@hadim

@hadim hadim commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Document::text, and so plain rdocx text, now reads the paragraphs that
    content controls wrap at every level: body, table (around rows), row (around
    cells), cell, and nested controls. It also reads the paragraphs of a table
    nested in a cell. The format is unchanged: a body paragraph ends with a
    newline, a table row is one line, and every paragraph of its cells ends with
    a tab.
  • Document::images now walks visit_all_drawings and Document::word_count
    walks visit_body_paragraphs, so pictures and words inside body, table, row,
    cell and inline controls are reported. The private collectors
    collect_images_from_* and word_count_in_* are removed.
  • Document::headings (and so document_outline, audit_accessibility and
    the heading gap warning of rdocx validate) and Document::links now read
    the paragraphs of body-level and nested controls through a small private
    collector, block_paragraphs. Their scope is unchanged: body paragraphs,
    never table cells, whether or not a control wraps the table.
  • MHTML export (to_mhtml_bytes) takes its picture sizes from a private
    collector in html.rs, emitted_paragraphs, over the paragraphs the
    rdocx-html emitter reaches, instead of from Document::images. Its output is
    the same as on main: a picture in a content control is still dropped with
    the existing loss diagnostic.
  • HLD 03 states the reach of the five walkers and of MHTML image sizing. Its
    WASM sentence now names Document::text instead of "that additive facade
    accessor", whose antecedent the new sentences had moved away.
  • The rdocx and rdocx-cli archive rows are re-recorded.

Part of #160. This covers section 1 for the Rust read walkers only. Section 1 still needs the replace and regex
walkers, header and footer blocks, text boxes and the Markdown, HTML and EPUB
exporters (see Notes). Sections 2 to 5 are other PRs.

Why

Google Docs wraps runs, and whole paragraphs including headings, in goog_rdk
content controls, and Word keeps them. The read walkers matched BodyContent
by hand and dropped ContentControl (BodyContent::ContentControl(_) => {} in
text, collect_images_from_content and word_count_in_content, and
CellContent::ContentControl(_) => {} in the cell loops). headings and
links looked at direct body paragraphs only. So a paragraph wrapped at body
level vanished from rdocx text, and wrapped pictures, headings and words
vanished from images, headings and word_count, while Paragraph.text,
Document::paragraphs and rdocx text --json already showed them.

Notes

  • Reuse. images and word_count are thin closures over the existing
    visitors, which already interleave table, row and inline controls by index.
    text keeps a small walker of its own because its separator depends on
    context (a newline at block level, a tab inside a row). Each row line is
    built with visit_row, and the table-level interleave follows the
    visit_table pattern. No public signature changes.
  • Nested tables. Paragraphs of a table nested in a cell now contribute
    to text, inside the outer row's line, each ending with a tab. This is the
    only output change for documents without content controls. For a python-docx
    cell.add_table cell the line goes from \t\t\n to \tinner gamma\t\t\n.
  • Scope of headings and links. The rule this PR keeps is that a content
    control is transparent: wrapping content in one never changes what these
    walkers report, and results for documents without controls are unchanged.
    block_paragraphs walks body paragraphs and the paragraphs of body-level
    and nested controls, and does not enter tables. CT_Body::paragraphs, the
    set behind Document::paragraphs, was not reused because it also returns
    the cells of a table that a body-level control wraps. Reading it would
    report a heading in a wrapped table but not in the same table unwrapped, and
    so change rdocx validate warnings depending only on the wrapper.
  • Open question for the maintainer. A heading or a link in a table cell is not
    reported, with or without a control. rebuild_toc source discovery
    (field.rs, collect_body_paragraphs) does include table cells, as Word
    does, and visit_body_paragraphs would serve both walkers.
    That is a one-line change in each walker plus the matrix expectations, but
    it changes headings, document_outline, audit_accessibility,
    rdocx validate and links for documents without controls, so it is left
    out of this PR.
  • MHTML. The writer pairs the emitter's <img> tags with picture sizes by
    position and refuses a count mismatch. The emitter still drops content
    controls, so sizes taken from the new images() would fail the export with
    "document contains images the HTML emitter did not serialize". The private
    emitted_paragraphs keeps the reach images() had on main, and is meant to
    go away once the rdocx-html exporter follow-up reads content controls. It
    deliberately keeps one pre-existing quirk of that reach: the emitter skips
    vertical merge continuation cells, the collector does not, so a picture in
    such a cell still fails the export with the count error, as on main. Skipping
    those cells would make the export succeed and drop the picture with no loss
    diagnostic, which belongs with the exporter follow-up.
  • Not reached, and unchanged here: a hyperlink inside an inline control and an
    inline control inside a hyperlink are kept as opaque XML by the paragraph
    parser, so Paragraph.text, text() and links() all miss their text. This
    is a model limitation in rdocx-oxml, not a walker one.
  • rdocx text --json omits rows wrapped by table-level controls and cells
    wrapped by row-level controls, because collect_table_paragraphs and
    collect_row_paragraphs in commands.rs iterate only rows and cells.
    Plain rdocx text now shows them, so the two views disagree on those two
    locations. rdocx-cli sources are not touched here because rdocx convert --to pdf|md|html, rpptx convert --to pdf and rpptx thumbnail overwrite any existing output, the input included #156 and rdocx and rpptx CLIs panic on a closed standard output (| head) #166 are
    rewriting commands.rs.
  • Left to other PRs, as planned: the replace and regex walkers (160-1b, with
    the run anchor remapping of 160-1f, a later medium-risk PR), the text-box
    rewrite (160-1e, branch fix/text-box-rewrite-keeps-all-children), header
    and footer block content (160-1d), and the exporters (rdocx-html
    markdown.rs and emitter.rs, rdocx epub.rs) as a follow-up. The legacy
    insert_toc still collects direct body paragraphs only.

Tests

  • crates/rdocx/tests/regression_test.rs, new module
    content_control_read_walker_regressions next to the ordered reader tests:
    • a_content_control_hides_nothing_from_any_body_read_walker is the matrix.
      Ten locations (body control, nested body control, inline control, nested
      inline control, cell control, nested cell control, row control around a
      cell, table control around a row, control in a nested table, body control
      around a table) against five walkers (text, images, word_count,
      headings, links). At each location all five results must equal those
      of the same document without the control, the probe's text, picture and
      words must be reported, and its heading and link exactly at the four
      body-level locations.
    • text_reads_every_content_control_location_in_document_order checks the
      exact text of one document mixing every location.
    • nested_table_cells_contribute_text_inside_their_outer_row pins that
      format.
  • crates/rdocx/src/html.rs, unit test
    mhtml_writer_sizes_only_the_pictures_the_html_emitter_serializes: pictures
    in a body control, outside any control and in an inline control. images()
    reports three, to_mhtml_bytes succeeds with the two existing loss
    diagnostics, and the reopened MHTML holds only the outside picture at its
    own size.
  • crates/rdocx-cli/tests/integration.rs,
    text_prints_paragraphs_wrapped_by_a_body_content_control, next to the
    existing plain text test: rdocx text prints a body-level control's
    paragraph.
  • All five fail on 9a7ed71 and pass on this branch. The MHTML test fails
    there on images().len() == 3, since main reports one picture.
  • Issue reproduction with python-docx 1.2.0 fixtures against the debug CLI of
    9a7ed71 and of this branch. rdocx text: no control and inline control
    print the same text, a body control goes from nothing to its paragraph, a
    cell control from an empty row to cell beta\t, a nested table from \t\t
    to \tinner gamma\t\t, and a wrapped heading now prints. rdocx validate
    after a Heading 1: a Heading 3 wrapped in a body control now gives the
    heading gap warning, and a Heading 3 in a table cell gives none, with or
    without a body control around the table, on both builds.
  • Gates run on the final head: cargo fmt --all --check, cargo clippy -p rdocx -p rdocx-cli --all-targets --all-features -- -D warnings, cargo test -p rdocx --no-fail-fast, cargo test -p rdocx-cli, cargo test -p rdocx-wasm (the consumer of text()), python3 scripts/hash_harness.py --check (49 entries match), python3 scripts/prose_check.py (0 violations)
    and python3 scripts/readme_doctests.py. All pass except the
    environment-only failures below. No Python binding or stub changed, and
    rdocx-py binds none of the changed walkers.
  • Environment-only failures on this Mac. rdocx lib:
    large_word_and_presentation_pdfs_preserve_logical_reading_order and
    word_and_powerpoint_chart_pixels_are_identical (pinned pdftotext and
    rasterizer). rdocx integration: odt_reader_matches_pinned_libreoffice_structure,
    public_authored_theme_and_fonts_match_pinned_word_resolution,
    sanitized_public_authoring_fixture_passes_every_conformance_stage, and two
    more that fail on their first assertion, the soffice --version pin (local
    26.8.0.3, pinned 26.2.5.2): section_page_semantics_match_pinned_libreoffice_render
    and every_conditional_table_region_matches_word. Those two fail
    identically on a clean 9a7ed71.

Document::text, and so plain `rdocx text`, skipped every content
control. A body-level control, which Google Docs writes around whole
paragraphs with a goog_rdk tag, dropped its paragraph from the output.
Cell controls, row controls around cells and table controls around rows
dropped theirs from the row line. Paragraphs of a table nested in a
cell were skipped as well, although the doc comment promises table cell
text.

The walker now recurses through controls at every level and builds each
row line with the existing visit_row visitor. A body paragraph still
ends with a newline, a table row is still one line, and every paragraph
of its cells, nested tables included, ends with a tab. Documents without
content controls or nested tables give exactly the text they gave
before.

GitHub issue tensorbee#160.
Document::images, word_count, headings and links skipped content
controls, and so did document_outline and the accessibility audit built
on them. A picture or a heading that Google Docs wrapped in a goog_rdk
control was not reported, and word_count missed the words of every
wrapped paragraph.

images now walks visit_all_drawings and word_count walks
visit_body_paragraphs, the existing visitors that interleave body,
table, row, cell and inline controls in document order. headings and
links keep their scope, body paragraphs without table cells, and now
collect it through body-level and nested controls. A table that such a
control wraps is still not searched, so wrapping content in a control
never changes what they report. The private collectors of images and
word_count are removed. Results for documents without content controls
are unchanged.

MHTML export pairs the <img> tags of the HTML emitter with picture
sizes by position and refuses a count mismatch. The emitter still drops
content controls, so sizes taken from the new images() would make the
export fail. The writer now collects them from the paragraphs the
emitter reaches, which is the reach images() had before. A picture in a
content control is still dropped with the existing loss diagnostic, as
it was before this change.

GitHub issue tensorbee#160.
The read walker fix and its regression tests grow the rdocx package,
and the new `rdocx text` integration test grows the rdocx-cli package,
so both README rows and their ARCHIVE_MEASUREMENTS entries are
re-measured.

GitHub issue tensorbee#160.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant