Conversation
Document::split_run resolved its first argument with Document::paragraph, which counts paragraphs, skips tables and enters block content controls. RunPosition, find_content_index, add_comment and insert_paragraph use the direct body child index instead, which is also what HLD 10 and F-X109 state for split_run. With a table or a block content control before the target, splitting a run then anchoring a comment at the same index split one paragraph and commented another, or failed on a run length. split_run now reads the paragraph at that direct slot and writes it back to the same slot, so one resolution serves both steps. An index that names a table, a block content control or preserved XML returns an error naming that kind instead of splitting another paragraph. The Python method also accepts a Paragraph handle, which reaches paragraphs inside block content controls, and it measures the revision bump on the paragraph it split. A handle to a direct body paragraph takes the same native path as its index. A table cell paragraph handle is refused, because its index counts the paragraphs inside cell content controls and Cell::paragraph_mut does not. The rustdoc of RunPosition and Document::paragraph and the HLD 03 and 10 statements now name the index each one takes. GitHub issue tensorbee#163.
Document::add_bookmark takes a RunRange whose body index is the direct body child index, but Document::bookmarks reports the paragraph ordinal counted recursively through table cells and block content controls. A bookmark added after a table with more than one cell, or after a block content control, read back with a different body index than the one it was added with, so that index could not be passed back to add_comment, split_run or add_bookmark. The recursive ordinal stays in range() because REF numbering in field.rs reads it as its paragraph key. BookmarkRef gains an additive direct_range() that carries the direct body child index, recorded while the paragraphs are collected, and is None when either marker sits in a table cell or a block content control. Run indexes are unchanged and stay the accepted-view boundaries of range(). They equal the direct run indexes add_bookmark takes only in a paragraph without inline content controls or tracked insertions, which GitHub issue tensorbee#172 covers. GitHub issue tensorbee#163.
Document::add_story_comment resolved a StoryRunRange paragraph by counting the paragraph items before it, which are the direct paragraphs of the owner only, and then took that many paragraphs through a resolver that also enters block content controls. When a block content control with paragraphs came before the target, in the body or in a table cell, the comment was anchored on a paragraph inside the control, or refused when that paragraph had too few runs. A body paragraph item now resolves to its direct body slot, the count of direct story items before it, which is the slot StoryItemRef::direct_body_index reports. A cell paragraph item resolves to the direct cell paragraph at the same position. Neither side enters a block content control, so the count and the lookup agree. GitHub issue tensorbee#163.
The direct body index changes grow the rdocx package, mostly through the new regression tests, so the crates.io archive row of the root README and its ARCHIVE_MEASUREMENTS entry in scripts/readme_doctests.py are re-measured with cargo package, as the Docs job requires. GitHub issue tensorbee#163.
The Python binding could not read or add a bookmark, append a field or a tab to a run, or insert a table of contents, although the Rust facade has all of them. A PAGEREF to a bookmarked run range could therefore not be scripted from Python. Document.bookmarks returns frozen Bookmark snapshots with the recursive range and the direct range that add_bookmark takes, which is None when a marker sits in a table cell or a block content control. add_bookmark and Run.add_tab edit in place without moving a run or a story item, so they keep live handles valid. Run.add_field advances the revision, because the new field is a story item of its own and moves the index path of every later StoryItem. The Run.text setter now shares their edit helper. insert_toc checks its index and level before the native call, and raises when the native call inserts nothing because it cannot allocate unique heading bookmarks. GitHub issue tensorbee#168.
Python could only set the text of the default header and footer, so a footer with fields for one section had to be written outside rdocx. The Rust facade already creates, links and unlinks the header or footer story of one section, and moves content between stories. Document.create_section_story, link_section_story and unlink_section_story take a section index and the kind and variant names that HeaderFooterVariant reports, reject an unknown name instead of reading it as the default variant, and return the resulting Story. pop_content now also takes a StoryItem, and insert_content a StoryItem or a Story, so a paragraph built in the body with a text run, a tab and PAGE and NUMPAGES fields can move into the footer of one section. GitHub issue tensorbee#168.
Python try_replace_text returned a count but could not check it, so a caller that wanted the contract of the CLI --expect had to copy the document and replace twice. The native replace_all takes an unordered map, has no count per pair and panics on a preflight failure, so it could not be bound as it is. Document::try_replace_all_expected stages ordered pairs, runs each one over the whole document after the pairs before it, and returns the count of each. When a pair finds another count than the one it expects, nothing is published and a ReplacementCountMismatch names the pair. Python binds it as replace_all(pairs), and try_replace_text gains a keyword expect. A mismatch raises ReplacementCountError, with the name and the expected and found attributes of the rpptx binding plus the index of the failing pair, and leaves the bytes and live handles unchanged. Zero matches without an expected count still return zero. GitHub issue tensorbee#168.
Python could render PDF only with the fonts layout finds on its own, and had no SVG page or PDF/A output, although the Rust facade has to_pdf_with_fonts, render_page_to_svg and to_pdfa_deterministic. The CLI --font-dir was the only way to render with a font directory. Document.to_pdf takes keyword-only fonts, as (family, bytes) pairs, and font_dir, loaded with load_fonts_from_dir, and calls to_pdf_with_fonts as the CLI does. The plain call is unchanged. The native loader reads a missing directory as an empty one, so the binding raises FileNotFoundError, or NotADirectoryError for a file, before any layout. render_page_to_svg returns frozen SvgRenderResult and SvgDiagnostic snapshots, and to_pdfa_deterministic takes pdfa-2b or pdfa-3b. rdocx re-exports PdfConformance, which the PDF/A method takes, so a caller no longer needs oxml-pdf to name it. All three release the GIL. GitHub issue tensorbee#168.
The counted replacement batch, its unit tests and the PdfConformance re-export grow the rdocx package, so the crates.io archive row of the root README and its ARCHIVE_MEASUREMENTS entry in scripts/readme_doctests.py are re-measured with cargo package, as the Docs job requires. The rdocx-py binding is not published and has no row. GitHub issue tensorbee#168.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Document.bookmarksreturns frozenBookmarksnapshots (
id,name,range,direct_range,text,issue), andDocument.add_bookmark(name, range)returns the new id.Run.add_tab()andRun.add_field(instruction, cached_result="")append inside a run, andDocument.insert_toc(index, max_level=3)binds the native table of contentsinsertion. A bookmark over a run range, a
PAGEREFto it on another run andupdate_layout_backed_fields()filling it now work from Python.Document.create_section_story,link_section_storyandunlink_section_storytake a section index and theheader/footeranddefault/first/evennames thatHeaderFooterVariantreports, and return the resultingStory.pop_contentnow also takes aStoryItem, andinsert_contentaStoryItem(the boundary before it) or aStory(its end), so a paragraphbuilt in the body with a text run, a tab and
PAGEandNUMPAGESfieldsmoves into the footer of one section.
rdocxfacade methodDocument::try_replace_all_expected(&[(old, new, Option<expected>)]):ordered, sequential, staged, all or nothing, with the count of each pair and
a
ReplacementCountMismatchnaming the failing pair. Python gainsDocument.replace_all(pairs)over it and a keywordexpectontry_replace_text. A mismatch raisesReplacementCountError(RdocxError)with
expected,foundandindex, and leaves the bytes and live handlesunchanged.
Document.to_pdf(*, fonts=None, font_dir=None),Document.render_page_to_svg(page_index)returning frozenSvgRenderResultand
SvgDiagnosticsnapshots, andDocument.to_pdfa_deterministic(profile="pdfa-2b").rdocxre-exportsPdfConformance.rdocxarchive row re-recorded.Stub,
typing_smoke.py,__init__.pyexports, therdocx-pyREADMEcapability list, HLD 10 and one HLD 03 sentence are updated in the commit of
each change.
Part of #168: the items "Bookmarks and fields", "Headers and footers per
section", "Replacement" and "Rendering" of "Exists in Rust, not bound". Still
open: tables and sections (in #187), styles and lists, comparison options (in
#176 for #161), and the whole "Not found in Rust" list. The story-level
pop_contentandinsert_contentadded here cover the part of the XML escapehatch item that moves existing body elements into another story. Building a
fragment from XML is not done.
Why
The #168 chain still dropped to python-docx and lxml for four jobs the Rust
facade already does: a
PAGEREFto a bookmark, a footer with page fields forone section, a replacement batch with a count per pair, and a PDF with a font
directory. The native
replace_allcould not be bound as it is: it takes anunordered
HashMap, reports one total, and panics on a preflight failure.Binding it would have given order-dependent results for chained pairs and no
way to check a pair's count before publishing.
Notes
Bookmarkexposes both ranges, as the RustBookmarkRefdoes after Align split_run, bookmarks and story comments on the direct body index #179:rangecounts paragraphs through tables andblock content controls,
direct_rangeuses the direct body index thatadd_bookmarkandRunPositiontake and isNonefor a nested marker. Therun-index caveat of Align split_run, bookmarks and story comments on the direct body index #179 applies unchanged (Comment run positions skip runs inside inline content controls, so comments anchor on the wrong text without an error #172).
add_bookmarkandadd_tabkeep live handles valid:bookmark markers sit between runs and a tab stays inside its run, and
story_itemsis identical before and after.add_fieldadvances therevision once. The new field is a story item of its own, so every later
StoryItemmoves one index path further, and without the bump a heldStoryItemcould silently name the paragraph before the one it was takenfrom. A test pins this. The section story operations advance the revision
once even when the variant already had its own story, because the native
calls publish a reopened package.
insert_tocwrites a static table of contents. The native method writesa title and one linked entry per heading with new
_TocNbookmarks, and noTOCfield, sorebuild_tocdoes not refresh it later. It also returns()and silently inserts nothing when it cannot allocate unique headingbookmarks (duplicate typed ids in the body). The binding checks the index
and level first and raises
RdocxErrorin the silent case. Question for themaintainer: should the native method write a real
TOCfield, or return aResult? Either is outside this PR.handles only reach body and cell paragraphs, so the footer paragraph is
authored in the body and moved. I did not bind
set_raw_header_with_imagesand
set_raw_footer_with_images: they take no section index andexpect()on failure.
inherit_section_story,replace_section_storyandremove_section_storyare not bound either, to keep this PR to the threecalls rdocx Python bindings: the complete list of what a production editing chain still needs, as one checklist #168 names. Each is a few lines on the same helper if wanted. Note that
create_section_storycreates an emptyw:ftr, with no paragraph, untilcontent is inserted.
try_replace_all_expectedreturnsResult<Result<Vec<usize>, ReplacementCountMismatch>>: the outer error is astaging failure, the inner one the count contract. A dedicated
rdocx::Errorvariant would read better, butErroris not#[non_exhaustive], so a new variant breaks every exhaustive matchdownstream (
rdocx-pyhas one). I recommend moving the mismatch intoErrorat the next breaking release. Each pair runsreplace_batchover thestaged candidate after the pairs before it, which gives the same result as
one batch because stories are independent, and gives a count per pair. The
legacy
replace_allis untouched.expect,error name,
expectedandfound, and the message ofrdocx replace --expectfor a single replacement. Zero matches without an expected countreturn 0. The rdocx class adds
index, the failing pair of a batch, whichis
Nonefortry_replace_text, and takes(message, expected, found, index=None)so that it pickles. Recommended follow-up in Bind the missing rpptx facade methods of the deck workflow in Python #181: give therpptx class the same
index=Noneso both libraries expose identicalattributes.
replace_allaccepts(old, new)and(old, new, expected)tuples. An empty placeholder keeps the native meaning of zero matches.
to_pdf(fonts=..., font_dir=...)callsDocument::to_pdf_with_fonts, asrdocx convert --font-dirdoes. That call lays out with the caller fontsonly, through
layout_with_fonts_and_options, so a document family thefonts do not provide, even through the automatic label aliases and
metric-compatible names, fails with
No font found for family. Its rustdocstill lists embedded, system and bundled fonts after the caller fonts,
which no longer matches the code since the relayout cache rework. I kept the
native behaviour and documented it in the stub and HLD 10, and a test pins
it. Recommended: either make
to_pdf_with_fontsuse the existingbundled-fallback layout path, or correct its rustdoc. The binding follows
either way without change. The native loader reads a missing directory as an
empty one, so the binding raises
FileNotFoundError, orNotADirectoryErrorfor a file, before any layout. SVG and PNG take nocaller fonts because no native entry point combines them.
to_pdfa_deterministickeeps the native name, since it lays outwith bundled fonts only, unlike
to_pdf. Profiles arepdfa-2bandpdfa-3b.render_page_to_svgreturns the native pair asSvgRenderResultandSvgDiagnostic, following the frozen snapshot rule ofHLD 10.
rdocxgainsDocument::try_replace_all_expected, thepublic
ReplacementCountMismatchand thePdfConformancere-export, alladditive. In Python,
to_pdfgains keyword-only arguments,try_replace_texta keyword-only
expect, andpop_contentandinsert_contentacceptstory coordinates besides the integer. A non-integer that is none of those
raises
TypeErrornaming the accepted forms, and a negative integer stillraises
OverflowError. TheRun.textsetter moved into a shared edithelper with no behaviour change. No existing Python call changes meaning.
git merge-treeagainst every open branch of thisround: rdocx-py: table formatting, table insertion and merges, and section edits #187 conflicts in the stub and
document.rsbecause itsinsert_tablelands right abovepop_content, and in the HLD 10 sentenceon Python section story mutation, which both PRs rewrite. Both resolve by
keeping both sides. The first line of the stub is changed exactly as Expose comparison options as keywords of Python Document.compare #176
and rdocx-py: table formatting, table insertion and merges, and section edits #187 change it, so it merges cleanly. Every branch that touches
rdocxconflicts on the archive row, as usual. Its "Measured on" date stays at
2026-09-26 like the rest of this round. No other conflict.
Tests
Added to
crates/rdocx-py/tests/test_core.py, each next to the related tests:test_counted_replacement_checks_expected_counts_before_publishing, afterthe existing counted replacement test: the acceptance batch (three pairs,
the second matching twice where once was declared) raises with
index=1,the message names the pair, the error pickles, and bytes and held handles
are unchanged. Also
try_replace_text(expect=)with the CLI message andindex=None, zero matches with and without an expectation, chained pairs inorder across body and header, and the
TypeErrorof malformed pairs.test_pageref_to_a_bookmarked_run_range_is_filled_from_the_layout, afterthe layout-backed field test: a run range made by
split_runafter a 2x2table is bookmarked,
bookmarksreports the recursive range at ordinal 5and the direct range equal to the one given, a
PAGEREFon the first run isfilled with
2, the field stales held run and story item handles, and aduplicate name, a table index and an empty field instruction are refused
with the bytes unchanged.
test_bookmarks_report_nested_ranges_without_a_direct_range: a bookmark ina block content control has a range and no direct range, and an unmatched
start reports its issue.
test_insert_toc_links_headings_and_tab_and_fields_extend_a_run: index andlevel errors change nothing, then entries, hyperlink anchors and
_TocNbookmarks, and a tab keeping the run handle live.
test_footer_with_text_tab_and_page_fields_is_built_for_one_section, afterthe story text tests: the rdocx Python bindings: the complete list of what a production editing chain still needs, as one checklist #168 acceptance line. A two-section document gets
a footer only in section 1, built from a text run, a tab, and
PAGEandNUMPAGESfields, whichupdate_layout_backed_fieldsfills with2and2, and the document renders.test_section_stories_link_unlink_and_reject_bad_names_atomically: unknownkind or variant raises
ValueError, an out-of-range sectionIndexError,a footer linked as a header
RdocxError, bad destination typesTypeError,all with the bytes unchanged. Then link, unlink into a copy, and
pop_contentof a story item.Added to
crates/rdocx-py/tests/test_rendering_threads.py, after therender_pagestest and the PNG GIL test:test_to_pdf_takes_caller_fonts_as_bytes_or_from_a_directory: a run in anunknown family renders with the bundled Liberation Mono given as bytes or
through a directory (as
strorPath), an empty directory raisesLayoutError(caller fonts only), and a missing directory or a file raisesFileNotFoundErrororNotADirectoryError.test_render_page_to_svg_returns_the_page_and_its_diagnostics.test_to_pdfa_deterministic_writes_the_requested_profile:pdfaid:part2by default and 3 for
pdfa-3b, deterministic bytes, unknown profilerefused.
test_font_svg_and_pdfa_renders_release_gil_for_python_worker.Added to the unit tests of
crates/rdocx/src/document.rs, right afterreplace_all_batch:expected_replacement_batch_runs_pairs_in_order_and_counts_eachand
expected_replacement_batch_mismatch_names_the_pair_and_changes_nothing.replacement_flush_failures_leave_the_live_document_unchangedalso covers thenew method on a package whose staging fails.
typing_smoke.pycovers every new name and return type.Run on this branch (macOS arm64, Python 3.12.14):
cargo fmt --all --check: clean.cargo clippy -p rdocx -p rdocx-py --all-targets --all-features -- -D warnings: clean.cargo check --workspace --all-targets --all-features:clean.
cargo test -p rdocx --no-fail-fast: regression 568 passed, doctests 2passed. Lib 467 passed and 2 failed, integration 314 passed and 5 failed, all
seven in the known local list (pinned pdftotext, rasterizer, Word and
LibreOffice 26.2.5.2 outputs against the local tools).
cargo test -p rdocx-cli, a direct consumer: 19 passed.cargo test -p rdocx-py: 2 passed withDYLD_LIBRARY_PATHset to thelibpython the test binary links.
maturin develop --lockedthenpytest crates/rdocx-py/tests: 78 passed,2 failed, both in
test_rendering_threads.pyon the pinned pdfinfo 26.01.0(local 26.09.0):
test_poppler_pdf_oracle_is_available_at_reviewed_versionand
test_four_concurrent_to_pdf_calls_are_faster_than_serial.test_python_docx_parity.pypasses here with python-docx installed. Each ofthe four change commits was also built and tested on its own: 71, 73, 74
and 78 passed, with the same two failures.
mypy --strict crates/rdocx-py/tests/typing_smoke.py crates/rdocx-py/python/rdocx(mypy 2.3.0): no issues.python -m mypy.stubtest rdocx: no issues.python3 scripts/hash_harness.py --check: 49 entries match.python3 scripts/prose_check.py: 0 violations.python3 scripts/readme_doctests.py: pass after the re-record.python3 scripts/sync_agent_skills.py --check: in sync.The four acceptance lines, run as one script against the final build:
The footer line lists which sections have a default footer, the PAGE and
NUMPAGES counts written, and the footer text with its caches. The rendering
line shows the font of a font directory in the PDF, the SVG start and the
PDF/A part.