Skip to content

Bind bookmarks, fields, section stories, counted replacement and rendering in Python - #194

Open
hadim wants to merge 9 commits into
tensorbee:mainfrom
hadim:feat/py-bookmarks-fields-stories-replace-render
Open

hadim wants to merge 9 commits into
tensorbee:mainfrom
hadim:feat/py-bookmarks-fields-stories-replace-render

Conversation

@hadim

@hadim hadim commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Stacked on #179 (fix/direct-body-index-coordinates). The first four
commits of this branch are #179 and drop out once it merges. The five
commits after them are this PR. The bookmark binding reads
BookmarkRef::direct_range(), which #179 adds.

Summary

  • Bookmarks and fields. Document.bookmarks returns frozen Bookmark
    snapshots (id, name, range, direct_range, text, issue), and
    Document.add_bookmark(name, range) returns the new id. Run.add_tab() and
    Run.add_field(instruction, cached_result="") append inside a run, and
    Document.insert_toc(index, max_level=3) binds the native table of contents
    insertion. A bookmark over a run range, a PAGEREF to it on another run and
    update_layout_backed_fields() filling it now work from Python.
  • Headers and footers per section. Document.create_section_story,
    link_section_story and unlink_section_story take a section index and the
    header/footer and default/first/even names that
    HeaderFooterVariant reports, and return the resulting Story.
    pop_content now also takes a StoryItem, and insert_content a
    StoryItem (the boundary before it) or a Story (its end), so a paragraph
    built in the body with a text run, a tab and PAGE and NUMPAGES fields
    moves into the footer of one section.
  • Counted replacement. New rdocx facade method
    Document::try_replace_all_expected(&[(old, new, Option<expected>)]):
    ordered, sequential, staged, all or nothing, with the count of each pair and
    a ReplacementCountMismatch naming the failing pair. Python gains
    Document.replace_all(pairs) over it and a keyword expect on
    try_replace_text. A mismatch raises ReplacementCountError(RdocxError)
    with expected, found and index, and leaves the bytes and live handles
    unchanged.
  • Rendering. Document.to_pdf(*, fonts=None, font_dir=None),
    Document.render_page_to_svg(page_index) returning frozen SvgRenderResult
    and SvgDiagnostic snapshots, and
    Document.to_pdfa_deterministic(profile="pdfa-2b"). rdocx re-exports
    PdfConformance.
  • rdocx archive row re-recorded.

Stub, typing_smoke.py, __init__.py exports, the rdocx-py README
capability list, HLD 10 and one HLD 03 sentence are updated in the commit of
each change.

Part of #168: the items "Bookmarks and fields", "Headers and footers per
section", "Replacement" and "Rendering" of "Exists in Rust, not bound". Still
open: tables and sections (in #187), styles and lists, comparison options (in
#176 for #161), and the whole "Not found in Rust" list. The story-level
pop_content and insert_content added here cover the part of the XML escape
hatch item that moves existing body elements into another story. Building a
fragment from XML is not done.

Why

The #168 chain still dropped to python-docx and lxml for four jobs the Rust
facade already does: a PAGEREF to a bookmark, a footer with page fields for
one section, a replacement batch with a count per pair, and a PDF with a font
directory. The native replace_all could not be bound as it is: it takes an
unordered HashMap, reports one total, and panics on a preflight failure.
Binding it would have given order-dependent results for chained pairs and no
way to check a pair's count before publishing.

Notes

  • Stacking. The Python Bookmark exposes both ranges, as the Rust
    BookmarkRef does after Align split_run, bookmarks and story comments on the direct body index #179: range counts paragraphs through tables and
    block content controls, direct_range uses the direct body index that
    add_bookmark and RunPosition take and is None for a nested marker. The
    run-index caveat of Align split_run, bookmarks and story comments on the direct body index #179 applies unchanged (Comment run positions skip runs inside inline content controls, so comments anchor on the wrong text without an error #172).
  • Revision policy. add_bookmark and add_tab keep live handles valid:
    bookmark markers sit between runs and a tab stays inside its run, and
    story_items is identical before and after. add_field advances the
    revision once. The new field is a story item of its own, so every later
    StoryItem moves one index path further, and without the bump a held
    StoryItem could silently name the paragraph before the one it was taken
    from. A test pins this. The section story operations advance the revision
    once even when the variant already had its own story, because the native
    calls publish a reopened package.
  • insert_toc writes a static table of contents. The native method writes
    a title and one linked entry per heading with new _TocN bookmarks, and no
    TOC field, so rebuild_toc does not refresh it later. It also returns
    () and silently inserts nothing when it cannot allocate unique heading
    bookmarks (duplicate typed ids in the body). The binding checks the index
    and level first and raises RdocxError in the silent case. Question for the
    maintainer: should the native method write a real TOC field, or return a
    Result? Either is outside this PR.
  • Rich footers go through the typed body API. Python paragraph and run
    handles only reach body and cell paragraphs, so the footer paragraph is
    authored in the body and moved. I did not bind set_raw_header_with_images
    and set_raw_footer_with_images: they take no section index and expect()
    on failure. inherit_section_story, replace_section_story and
    remove_section_story are not bound either, to keep this PR to the three
    calls rdocx Python bindings: the complete list of what a production editing chain still needs, as one checklist #168 names. Each is a few lines on the same helper if wanted. Note that
    create_section_story creates an empty w:ftr, with no paragraph, until
    content is inserted.
  • Shape of the Rust batch, decision for the maintainer.
    try_replace_all_expected returns
    Result<Result<Vec<usize>, ReplacementCountMismatch>>: the outer error is a
    staging failure, the inner one the count contract. A dedicated
    rdocx::Error variant would read better, but Error is not
    #[non_exhaustive], so a new variant breaks every exhaustive match
    downstream (rdocx-py has one). I recommend moving the mismatch into
    Error at the next breaking release. Each pair runs replace_batch over the
    staged candidate after the pairs before it, which gives the same result as
    one batch because stories are independent, and gives a count per pair. The
    legacy replace_all is untouched.
  • Shared contract with rpptx (Bind the missing rpptx facade methods of the deck workflow in Python #181). Same method name, keyword expect,
    error name, expected and found, and the message of rdocx replace --expect for a single replacement. Zero matches without an expected count
    return 0. The rdocx class adds index, the failing pair of a batch, which
    is None for try_replace_text, and takes (message, expected, found, index=None) so that it pickles. Recommended follow-up in Bind the missing rpptx facade methods of the deck workflow in Python #181: give the
    rpptx class the same index=None so both libraries expose identical
    attributes. replace_all accepts (old, new) and (old, new, expected)
    tuples. An empty placeholder keeps the native meaning of zero matches.
  • Caller fonts are the only fonts, decision for the maintainer.
    to_pdf(fonts=..., font_dir=...) calls Document::to_pdf_with_fonts, as
    rdocx convert --font-dir does. That call lays out with the caller fonts
    only, through layout_with_fonts_and_options, so a document family the
    fonts do not provide, even through the automatic label aliases and
    metric-compatible names, fails with No font found for family. Its rustdoc
    still lists embedded, system and bundled fonts after the caller fonts,
    which no longer matches the code since the relayout cache rework. I kept the
    native behaviour and documented it in the stub and HLD 10, and a test pins
    it. Recommended: either make to_pdf_with_fonts use the existing
    bundled-fallback layout path, or correct its rustdoc. The binding follows
    either way without change. The native loader reads a missing directory as an
    empty one, so the binding raises FileNotFoundError, or
    NotADirectoryError for a file, before any layout. SVG and PNG take no
    caller fonts because no native entry point combines them.
  • Names. to_pdfa_deterministic keeps the native name, since it lays out
    with bundled fonts only, unlike to_pdf. Profiles are pdfa-2b and
    pdfa-3b. render_page_to_svg returns the native pair as
    SvgRenderResult and SvgDiagnostic, following the frozen snapshot rule of
    HLD 10.
  • API impact. rdocx gains Document::try_replace_all_expected, the
    public ReplacementCountMismatch and the PdfConformance re-export, all
    additive. In Python, to_pdf gains keyword-only arguments, try_replace_text
    a keyword-only expect, and pop_content and insert_content accept
    story coordinates besides the integer. A non-integer that is none of those
    raises TypeError naming the accepted forms, and a negative integer still
    raises OverflowError. The Run.text setter moved into a shared edit
    helper with no behaviour change. No existing Python call changes meaning.
  • Conflicts to expect. git merge-tree against every open branch of this
    round: rdocx-py: table formatting, table insertion and merges, and section edits #187 conflicts in the stub and document.rs because its
    insert_table lands right above pop_content, and in the HLD 10 sentence
    on Python section story mutation, which both PRs rewrite. Both resolve by
    keeping both sides. The first line of the stub is changed exactly as Expose comparison options as keywords of Python Document.compare #176
    and rdocx-py: table formatting, table insertion and merges, and section edits #187 change it, so it merges cleanly. Every branch that touches rdocx
    conflicts on the archive row, as usual. Its "Measured on" date stays at
    2026-09-26 like the rest of this round. No other conflict.

Tests

Added to crates/rdocx-py/tests/test_core.py, each next to the related tests:

  • test_counted_replacement_checks_expected_counts_before_publishing, after
    the existing counted replacement test: the acceptance batch (three pairs,
    the second matching twice where once was declared) raises with index=1,
    the message names the pair, the error pickles, and bytes and held handles
    are unchanged. Also try_replace_text(expect=) with the CLI message and
    index=None, zero matches with and without an expectation, chained pairs in
    order across body and header, and the TypeError of malformed pairs.
  • test_pageref_to_a_bookmarked_run_range_is_filled_from_the_layout, after
    the layout-backed field test: a run range made by split_run after a 2x2
    table is bookmarked, bookmarks reports the recursive range at ordinal 5
    and the direct range equal to the one given, a PAGEREF on the first run is
    filled with 2, the field stales held run and story item handles, and a
    duplicate name, a table index and an empty field instruction are refused
    with the bytes unchanged.
  • test_bookmarks_report_nested_ranges_without_a_direct_range: a bookmark in
    a block content control has a range and no direct range, and an unmatched
    start reports its issue.
  • test_insert_toc_links_headings_and_tab_and_fields_extend_a_run: index and
    level errors change nothing, then entries, hyperlink anchors and _TocN
    bookmarks, and a tab keeping the run handle live.
  • test_footer_with_text_tab_and_page_fields_is_built_for_one_section, after
    the story text tests: the rdocx Python bindings: the complete list of what a production editing chain still needs, as one checklist #168 acceptance line. A two-section document gets
    a footer only in section 1, built from a text run, a tab, and PAGE and
    NUMPAGES fields, which update_layout_backed_fields fills with 2 and
    2, and the document renders.
  • test_section_stories_link_unlink_and_reject_bad_names_atomically: unknown
    kind or variant raises ValueError, an out-of-range section IndexError,
    a footer linked as a header RdocxError, bad destination types TypeError,
    all with the bytes unchanged. Then link, unlink into a copy, and
    pop_content of a story item.

Added to crates/rdocx-py/tests/test_rendering_threads.py, after the
render_pages test and the PNG GIL test:

  • test_to_pdf_takes_caller_fonts_as_bytes_or_from_a_directory: a run in an
    unknown family renders with the bundled Liberation Mono given as bytes or
    through a directory (as str or Path), an empty directory raises
    LayoutError (caller fonts only), and a missing directory or a file raises
    FileNotFoundError or NotADirectoryError.
  • test_render_page_to_svg_returns_the_page_and_its_diagnostics.
  • test_to_pdfa_deterministic_writes_the_requested_profile: pdfaid:part 2
    by default and 3 for pdfa-3b, deterministic bytes, unknown profile
    refused.
  • test_font_svg_and_pdfa_renders_release_gil_for_python_worker.

Added to the unit tests of crates/rdocx/src/document.rs, right after
replace_all_batch: expected_replacement_batch_runs_pairs_in_order_and_counts_each
and expected_replacement_batch_mismatch_names_the_pair_and_changes_nothing.
replacement_flush_failures_leave_the_live_document_unchanged also covers the
new method on a package whose staging fails.

typing_smoke.py covers every new name and return type.

Run on this branch (macOS arm64, Python 3.12.14):

  • cargo fmt --all --check: clean.
  • cargo clippy -p rdocx -p rdocx-py --all-targets --all-features -- -D warnings: clean. cargo check --workspace --all-targets --all-features:
    clean.
  • cargo test -p rdocx --no-fail-fast: regression 568 passed, doctests 2
    passed. Lib 467 passed and 2 failed, integration 314 passed and 5 failed, all
    seven in the known local list (pinned pdftotext, rasterizer, Word and
    LibreOffice 26.2.5.2 outputs against the local tools).
  • cargo test -p rdocx-cli, a direct consumer: 19 passed.
  • cargo test -p rdocx-py: 2 passed with DYLD_LIBRARY_PATH set to the
    libpython the test binary links.
  • maturin develop --locked then pytest crates/rdocx-py/tests: 78 passed,
    2 failed, both in test_rendering_threads.py on the pinned pdfinfo 26.01.0
    (local 26.09.0): test_poppler_pdf_oracle_is_available_at_reviewed_version
    and test_four_concurrent_to_pdf_calls_are_faster_than_serial.
    test_python_docx_parity.py passes here with python-docx installed. Each of
    the four change commits was also built and tested on its own: 71, 73, 74
    and 78 passed, with the same two failures.
  • mypy --strict crates/rdocx-py/tests/typing_smoke.py crates/rdocx-py/python/rdocx (mypy 2.3.0): no issues.
    python -m mypy.stubtest rdocx: no issues.
  • python3 scripts/hash_harness.py --check: 49 entries match.
    python3 scripts/prose_check.py: 0 violations.
    python3 scripts/readme_doctests.py: pass after the re-record.
    python3 scripts/sync_agent_skills.py --check: in sync.

The four acceptance lines, run as one script against the final build:

bookmarks and fields: 1 PAGEREF filled 2 ['target']
section footer: [(0, False), (1, True)] 1 1 ['ConfidentialPage 2 of 2']
replacement: pair 1: expected 1 replacement(s) of "{{b}}", found 2 (1, 1, 2) unchanged True
rendering: [b'LiberationMono'] <svg True

The footer line lists which sections have a default footer, the PAGE and
NUMPAGES counts written, and the footer text with its caches. The rendering
line shows the font of a font directory in the PDF, the SVG start and the
PDF/A part.

Document::split_run resolved its first argument with
Document::paragraph, which counts paragraphs, skips tables and enters
block content controls. RunPosition, find_content_index, add_comment
and insert_paragraph use the direct body child index instead, which is
also what HLD 10 and F-X109 state for split_run. With a table or a
block content control before the target, splitting a run then
anchoring a comment at the same index split one paragraph and
commented another, or failed on a run length.

split_run now reads the paragraph at that direct slot and writes it
back to the same slot, so one resolution serves both steps. An index
that names a table, a block content control or preserved XML returns
an error naming that kind instead of splitting another paragraph. The
Python method also accepts a Paragraph handle, which reaches
paragraphs inside block content controls, and it measures the
revision bump on the paragraph it split. A handle to a direct body
paragraph takes the same native path as its index. A table cell
paragraph handle is refused, because its index counts the paragraphs
inside cell content controls and Cell::paragraph_mut does not. The
rustdoc of RunPosition and Document::paragraph and the HLD 03 and 10
statements now name the index each one takes.

GitHub issue tensorbee#163.
Document::add_bookmark takes a RunRange whose body index is the direct
body child index, but Document::bookmarks reports the paragraph ordinal
counted recursively through table cells and block content controls. A
bookmark added after a table with more than one cell, or after a block
content control, read back with a different body index than the one it
was added with, so that index could not be passed back to add_comment,
split_run or add_bookmark.

The recursive ordinal stays in range() because REF numbering in
field.rs reads it as its paragraph key. BookmarkRef gains an additive
direct_range() that carries the direct body child index, recorded while
the paragraphs are collected, and is None when either marker sits in a
table cell or a block content control. Run indexes are unchanged and
stay the accepted-view boundaries of range(). They equal the direct run
indexes add_bookmark takes only in a paragraph without inline content
controls or tracked insertions, which GitHub issue tensorbee#172 covers.

GitHub issue tensorbee#163.
Document::add_story_comment resolved a StoryRunRange paragraph by
counting the paragraph items before it, which are the direct
paragraphs of the owner only, and then took that many paragraphs
through a resolver that also enters block content controls. When a
block content control with paragraphs came before the target, in the
body or in a table cell, the comment was anchored on a paragraph
inside the control, or refused when that paragraph had too few runs.

A body paragraph item now resolves to its direct body slot, the count
of direct story items before it, which is the slot
StoryItemRef::direct_body_index reports. A cell paragraph item
resolves to the direct cell paragraph at the same position. Neither
side enters a block content control, so the count and the lookup
agree.

GitHub issue tensorbee#163.
The direct body index changes grow the rdocx package, mostly through
the new regression tests, so the crates.io archive row of the root
README and its ARCHIVE_MEASUREMENTS entry in scripts/readme_doctests.py
are re-measured with cargo package, as the Docs job requires.

GitHub issue tensorbee#163.
The Python binding could not read or add a bookmark, append a field or
a tab to a run, or insert a table of contents, although the Rust facade
has all of them. A PAGEREF to a bookmarked run range could therefore
not be scripted from Python.

Document.bookmarks returns frozen Bookmark snapshots with the recursive
range and the direct range that add_bookmark takes, which is None when
a marker sits in a table cell or a block content control.
add_bookmark and Run.add_tab edit in place without moving a run or a
story item, so they keep live handles valid. Run.add_field advances the
revision, because the new field is a story item of its own and moves
the index path of every later StoryItem. The Run.text setter now shares
their edit helper. insert_toc checks its index and level before the
native call, and raises when the native call inserts nothing because it
cannot allocate unique heading bookmarks.

GitHub issue tensorbee#168.
Python could only set the text of the default header and footer, so a
footer with fields for one section had to be written outside rdocx.
The Rust facade already creates, links and unlinks the header or
footer story of one section, and moves content between stories.

Document.create_section_story, link_section_story and
unlink_section_story take a section index and the kind and variant
names that HeaderFooterVariant reports, reject an unknown name instead
of reading it as the default variant, and return the resulting Story.
pop_content now also takes a StoryItem, and insert_content a StoryItem
or a Story, so a paragraph built in the body with a text run, a tab and
PAGE and NUMPAGES fields can move into the footer of one section.

GitHub issue tensorbee#168.
Python try_replace_text returned a count but could not check it, so a
caller that wanted the contract of the CLI --expect had to copy the
document and replace twice. The native replace_all takes an unordered
map, has no count per pair and panics on a preflight failure, so it
could not be bound as it is.

Document::try_replace_all_expected stages ordered pairs, runs each one
over the whole document after the pairs before it, and returns the
count of each. When a pair finds another count than the one it
expects, nothing is published and a ReplacementCountMismatch names the
pair. Python binds it as replace_all(pairs), and try_replace_text gains
a keyword expect. A mismatch raises ReplacementCountError, with the
name and the expected and found attributes of the rpptx binding plus
the index of the failing pair, and leaves the bytes and live handles
unchanged. Zero matches without an expected count still return zero.

GitHub issue tensorbee#168.
Python could render PDF only with the fonts layout finds on its own,
and had no SVG page or PDF/A output, although the Rust facade has
to_pdf_with_fonts, render_page_to_svg and to_pdfa_deterministic. The
CLI --font-dir was the only way to render with a font directory.

Document.to_pdf takes keyword-only fonts, as (family, bytes) pairs, and
font_dir, loaded with load_fonts_from_dir, and calls
to_pdf_with_fonts as the CLI does. The plain call is unchanged. The
native loader reads a missing directory as an empty one, so the binding
raises FileNotFoundError, or NotADirectoryError for a file, before any
layout. render_page_to_svg returns frozen SvgRenderResult and
SvgDiagnostic snapshots, and to_pdfa_deterministic takes pdfa-2b or
pdfa-3b. rdocx re-exports PdfConformance, which the PDF/A method takes,
so a caller no longer needs oxml-pdf to name it. All three release the
GIL.

GitHub issue tensorbee#168.
The counted replacement batch, its unit tests and the PdfConformance
re-export grow the rdocx package, so the crates.io archive row of the
root README and its ARCHIVE_MEASUREMENTS entry in
scripts/readme_doctests.py are re-measured with cargo package, as the
Docs job requires. The rdocx-py binding is not published and has no
row.

GitHub issue tensorbee#168.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant