Conversation
Document::split_run resolved its first argument with Document::paragraph, which counts paragraphs, skips tables and enters block content controls. RunPosition, find_content_index, add_comment and insert_paragraph use the direct body child index instead, which is also what HLD 10 and F-X109 state for split_run. With a table or a block content control before the target, splitting a run then anchoring a comment at the same index split one paragraph and commented another, or failed on a run length. split_run now reads the paragraph at that direct slot and writes it back to the same slot, so one resolution serves both steps. An index that names a table, a block content control or preserved XML returns an error naming that kind instead of splitting another paragraph. The Python method also accepts a Paragraph handle, which reaches paragraphs inside block content controls, and it measures the revision bump on the paragraph it split. A handle to a direct body paragraph takes the same native path as its index. A table cell paragraph handle is refused, because its index counts the paragraphs inside cell content controls and Cell::paragraph_mut does not. The rustdoc of RunPosition and Document::paragraph and the HLD 03 and 10 statements now name the index each one takes. GitHub issue tensorbee#163.
Document::add_bookmark takes a RunRange whose body index is the direct body child index, but Document::bookmarks reports the paragraph ordinal counted recursively through table cells and block content controls. A bookmark added after a table with more than one cell, or after a block content control, read back with a different body index than the one it was added with, so that index could not be passed back to add_comment, split_run or add_bookmark. The recursive ordinal stays in range() because REF numbering in field.rs reads it as its paragraph key. BookmarkRef gains an additive direct_range() that carries the direct body child index, recorded while the paragraphs are collected, and is None when either marker sits in a table cell or a block content control. Run indexes are unchanged and stay the accepted-view boundaries of range(). They equal the direct run indexes add_bookmark takes only in a paragraph without inline content controls or tracked insertions, which GitHub issue tensorbee#172 covers. GitHub issue tensorbee#163.
Document::add_story_comment resolved a StoryRunRange paragraph by counting the paragraph items before it, which are the direct paragraphs of the owner only, and then took that many paragraphs through a resolver that also enters block content controls. When a block content control with paragraphs came before the target, in the body or in a table cell, the comment was anchored on a paragraph inside the control, or refused when that paragraph had too few runs. A body paragraph item now resolves to its direct body slot, the count of direct story items before it, which is the slot StoryItemRef::direct_body_index reports. A cell paragraph item resolves to the direct cell paragraph at the same position. Neither side enters a block content control, so the count and the lookup agree. GitHub issue tensorbee#163.
The direct body index changes grow the rdocx package, mostly through the new regression tests, so the crates.io archive row of the root README and its ARCHIVE_MEASUREMENTS entry in scripts/readme_doctests.py are re-measured with cargo package, as the Docs job requires. GitHub issue tensorbee#163.
This was referenced Sep 27, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
rdocx:Document::split_runresolves its first argument as the directbody child index, the index that
find_content_index,RunPosition,add_comment,insert_paragraphandlayout().body_indexalready use. Itreads and writes back the paragraph at that one slot. An index that names a
table, a block content control or preserved XML returns an error naming that
kind (
body index 1 is a table, not a paragraph) instead of splittinganother paragraph.
rdocx-py:Document.split_runalso accepts aParagraphhandle in place of the index, which reaches paragraphs insideblock content controls. A table cell paragraph handle is refused with a
ValueError. The rustdoc ofRunPosition::body_indexandDocument::paragraphnow says which index each takes. Stub, typing smoke,HLD 03 and HLD 10 updated.
rdocx: additiveBookmarkRef::direct_range(). It reports the bookmarkwith the direct body index that
add_bookmarktook, and isNonewheneither marker sits in a table cell or a block content control.
range()keeps its recursive paragraph ordinal. HLD 03 and HLD 10 updated.
rdocx:add_story_commentresolves a body paragraph item to its directbody slot and a cell paragraph item to the direct cell paragraph at the same
position, so a block content control before the target no longer moves the
comment into the control. This one is not in the issue and has the same
cause.
rdocxre-recorded.Part of #163. This covers the index class of the issue:
split_run, thebookmark listing and the story comment path now agree with
find_content_index. What remains of #163 is left to later PRs: commenting aparagraph inside a block content control, which needs an API shape decision,
the
add_comment_on_texthelper, andrdocx comment add --anchor.Why
split_runtook an argument namedbody_indexand resolved it withDocument::paragraph, which counts paragraphs, skips body tables and entersblock content controls. Every other index-taking body API, and the spec of
F-X109 and HLD 10, use the direct body child index. Splitting a run then
anchoring a comment at the index
find_content_indexreturned therefore splitone paragraph and commented another as soon as a table or a block content
control came first. The issue's reproduction printed
['Gamm', 'a paragraph at the end.']and theworkflow_docx.pystep "comment on a piece of text after atable" of #158 failed on
fixture-report.docxwith a run length error. Thesame confusion made
bookmarks()read back a bookmark with a different bodyindex than the one
add_bookmarktook, after a table with more than one cellor a block content control.
add_story_commentcounted the paragraph items before its target, which aredirect paragraphs only, and then resolved that count through a walker that
enters block content controls. On a body or a cell where a block content
control with paragraphs precedes the target, the comment landed silently on a
paragraph inside the control. On the report fixture that is any body
paragraph after the table of contents control.
Notes
split_run(Rust and Python) is now the direct body index, as HLD 10 andF-X109 already stated. This was preferred over keeping the paragraph index
and adding a second entry point, which would have left a fourth meaning of
body_indexin the API. A caller that passed adoc.paragraphsindex tosplit_runin a body where a table or a block content control precedes thetarget now addresses a different body child. Such callers pass the index
from
find_content_indexor, in Python, theParagraphhandle itself. InRust,
paragraph_mut(i).split_run(...)still splits by paragraph index, andcontent_index_of_paragraphconverts. Suggested line: "split_runnow takesthe direct body index that
find_content_indexreturns, likeRunPositionand
add_comment, and Python also accepts aParagraphhandle. Code thatpassed a paragraph index with a table or content control before the target
addressed another paragraph and must switch." The out-of-range message
changes from
body paragraph index N is out of rangetobody index N is out of range. A non-integer argument raisesTypeError: body_index must be an int or a Paragraph handle, and a negative one still raisesOverflowError.find_content_indexdoes. A handle to a direct body paragraph is convertedwith
content_index_of_paragraphand takes the same native path as itsindex, so a zero or end offset leaves bytes, revision and the layout cache
alone. A handle to a paragraph inside a block content control splits through
Document::paragraph_mut, which clears the layout cache even on thoseoffsets (bytes and revision stay unchanged), and HLD 03 now says so. The
revision advances only when a continuation run is created, measured on the
paragraph that was split.
Paragraphhandle indexesCellRef::paragraph, which counts the paragraphs inside cell contentcontrols, while
Cell::paragraph_mutcounts direct paragraphs only. With acell
[sdt{In control.}, Direct one.], splitting throughcell.paragraphs[0]would splitDirect one..split_runnever acceptedcell paragraphs before this PR, so refusing them loses nothing. The same
read and write mismatch already affects the other Python cell handle
mutators (
run.text,add_run, the font and paragraph format setters).Making
Cell::paragraph_mutenter cell content controls would fix all of them and let
split_runacceptcell handles. That is a behaviour change of a public Rust accessor outside
Document.split_run(body_index, ...)counts paragraphs only, so it splits the wrong paragraph once a table precedes it #163, so it is left for the maintainer to decide, as its own issue.direct_range()is additive andrange()is unchanged,because REF numbering in
field.rskeys on its recursive ordinal. The runindexes of
direct_range()are the same accepted-view boundaries asrange(). They equal the direct run indexadd_bookmarkandadd_commenttake only in a paragraph without inline content controls or tracked
insertions, which the rustdoc and HLD 10 now state. Reconciling the run axis
is Comment run positions skip runs inside inline content controls, so comments anchor on the wrong text without an error #172 and is left to that PR. The Python bookmark binding of rdocx Python bindings: the complete list of what a production editing chain still needs, as one checklist #168 can
expose
direct_range()from the start.before the target as its direct body slot, from the scan it already has, and
requires a typed paragraph there. The body section properties are serialized
last, so they never precede a paragraph. This is the slot
StoryItemRef::direct_body_indexreports. The cell route takes the k-thdirect
CellContent::Paragraph. No extra scan is added.BookmarkRefgains a privatefield and a public accessor, and it has no public constructor, so nothing
breaks. Python widens the first parameter of
split_runtoint | Paragraphand keeps its namebody_indexfor keyword callers.rdocxarchive row, like every PR of this wavethat touches
rdocx. Its "Measured on" date is left at 2026-09-26, as theshared re-record of this wave does, so the maintainer can re-measure once
at merge time.
Document::textsits a few lines abovesplit_run, so theProducer traits: a matrix over every operation, and what still fails in it and around it #160 content control readers PR may touch nearby lines.
Tests
Added to
crates/rdocx/tests/regression_test.rs, asmod direct_body_index_coordinatesright afterchecked_table_cell_comment_range_is_atomic_and_reopens:split_run_takes_the_direct_body_index_after_a_table: the issuereproduction, then
add_commentat the same index anchors onBetaandinsert_paragraphat the same index lands right before it.split_run_takes_the_direct_body_index_after_a_block_content_control.split_run_names_the_body_child_that_is_not_a_paragraph: table, contentcontrol, a
w:customXmlbody child kept as preserved XML, and out of range,with the bytes unchanged.
bookmark_direct_range_reports_the_index_add_bookmark_took: after a 2x2table and a block content control,
direct_range()equals the range givento
add_bookmarkwhilerange()keeps ordinal 7.bookmark_direct_range_is_none_when_a_marker_is_nested: in a cell, in acontrol, and spanning from a direct paragraph into a control.
story_comment_after_a_block_content_control_anchors_on_its_paragraphandstory_comment_in_a_cell_after_a_block_content_control_anchors_on_its_paragraph.The five tests that use existing APIs failed on 9a7ed71 before the fix: the
split output was
['Gamm', 'a paragraph at the end.']as in the issue, thesplit after a control hit
Control two., the table index split a controlparagraph, and both story comments landed on
Control one..Added to
crates/rdocx-py/tests/test_core.py, next to the existing splittests:
test_split_run_takes_the_direct_body_index_after_a_table: the issuereproduction, the table index error, the
TypeErrormessage and theOverflowErrorof a negative index.test_split_run_accepts_a_paragraph_handle_in_a_control_or_the_body: ano-op split keeps run handles and a split stales them, for a control
paragraph and for a direct paragraph. Both paragraphs of a cell that holds
a block content control are refused with the bytes unchanged. A handle from
another document is refused.
test_story_comment_after_a_block_content_control_anchors_on_its_paragraph,which failed on the extension built before the story comment commit.
typing_smoke.pygainssplit_run(paragraph, 0, 1).Run on this branch:
cargo fmt --all --check: clean.cargo clippy -p rdocx -p rdocx-py --all-targets --all-features -- -D warnings: clean.cargo test -p rdocx --no-fail-fast: regression 568 passed and 7 ignored,doctests 2 passed. The lib binary has the two known local
failures (
large_word_and_presentation_pdfs_preserve_logical_reading_order,word_and_powerpoint_chart_pixels_are_identical, pinned pdftotext andrasterizer). The integration binary has 314 passed and 5 failed: the three
known ones (
odt_reader_matches_pinned_libreoffice_structure,public_authored_theme_and_fonts_match_pinned_word_resolution,sanitized_public_authoring_fixture_passes_every_conformance_stage) plusevery_conditional_table_region_matches_wordandsection_page_semantics_match_pinned_libreoffice_render, which assert thepinned LibreOffice 26.2.5.2 against the local 26.8.0.3 and fail the same
way on 9a7ed71.
cargo test -p rdocx-py: 2 passed withDYLD_LIBRARY_PATHset to thelibpython the build linked, and cannot load libpython without it.
RUSTDOCFLAGS="-D warnings" cargo doc -p rdocx --no-deps --all-features:clean.
maturin develop:pytest crates/rdocx-py/tests68 passed, 2 failed, both intest_rendering_threads.pyon the pinned pdfinfo 26.01.0 (local 26.09.0).mypy --strictontyping_smoke.pyand the package: clean.mypy.stubtest rdocx: clean.python3 scripts/hash_harness.py --check: 49 entries match.python3 scripts/prose_check.py: 0 violations.python3 scripts/readme_doctests.py, the Docs job check that includes thearchive rows: passes after the re-record.
['Beta', ' paragraph after the table.'], and a cell paragraph handle now raisessplit_run does not accept a table cell paragraph handle.workflow_docx.pyonfixture-report.docx, rerun on the final branch, passes "comment on a pieceof text after a table". Its three other failures (table of contents rebuild
on a fresh open, replacement inside content controls, redline) belong to
other issues of Production readiness for editing real docx and pptx files: an acceptance contract, two realistic fixtures and two matrices #158.