Skip to content

Document.split_run(body_index, ...) counts paragraphs only, so it splits the wrong paragraph once a table precedes it #163

Description

@hadim

RunPosition.body_index, find_content_index() and add_comment() count body children, tables included.
Document.split_run() takes an argument with the same name and resolves it with paragraph(idx)
(rdocx/src/document.rs:14679), which counts paragraphs only. The way to anchor a comment on
part of a run (split the run at both ends, then pass a RunRange over the middle run) therefore splits a
different paragraph as soon as a table comes earlier in the body, and the comment then lands on the wrong
text or fails.

import rdocx
from docx import Document

d = Document()
d.add_paragraph("Alpha paragraph before the table.")
d.add_table(rows=1, cols=1).cell(0, 0).text = "cell"
d.add_paragraph("Beta paragraph after the table.")
d.add_paragraph("Gamma paragraph at the end.")
d.save("sri.docx")

doc = rdocx.Document.open("sri.docx")
target = [p for p in doc.paragraphs if p.text.startswith("Beta")][0]
bi = doc.find_content_index(target)
print("body index of the Beta paragraph (find_content_index):", bi)
doc.split_run(bi, 0, 4)
print("runs per paragraph after split_run(%d, 0, 4):" % bi)
for p in doc.paragraphs:
    print("   ", [r.text for r in p.runs])
body index of the Beta paragraph (find_content_index): 2
runs per paragraph after split_run(2, 0, 4):
    ['Alpha paragraph before the table.']
    ['Beta paragraph after the table.']
    ['Gamm', 'a paragraph at the end.']

Expected: ['Beta', ' paragraph after the table.'].

The class is wider than one function: body_index has three meanings in the API today.

  • RunPosition.body_index, add_comment() and find_content_index(): a direct child of the body, tables
    counted (body_paragraph, rdocx/src/comments.rs:1257);
  • split_run(): the n-th paragraph, tables skipped and paragraphs inside block content controls counted
    (Document::paragraph, document.rs:14533);
  • bookmarks(): paragraphs counted recursively through tables and block content controls
    (comments.rs:256 to 271).

And add_comment() cannot reach a paragraph inside a block content control at all. The direct body index
answers that it "is not a paragraph"; a StoryRunRange over the content control's StoryItem answers
"comment positions must identify paragraphs"; and story_items lists no item for the paragraph inside the
control, so there is nothing else to pass.

#86 settled the index question for position APIs: the direct body-child index, with
StoryItem.direct_body_index to convert, and index_path kept recursive. I am not asking to reopen that.
split_run, added later for #97, counts paragraphs instead, and the bookmark APIs count recursively.
Aligning both on the #86 index would be enough; since changing what split_run counts would break its
callers, accepting a Paragraph or StoryItem handle wherever an index is taken today may be the gentler
path. Acceptance: the index returned by find_content_index() addresses the same paragraph in
split_run, add_comment, insert_paragraph, clone_content and the bookmark APIs with a table before
it, and a paragraph inside a block content control can be split and commented through a handle or a story
position.

A helper would also help here: anchoring a comment on a piece of text currently takes a text search, two
split_run calls and index bookkeeping. add_comment_on_text(text, occurrence=0, author=..., text=...),
or --anchor TEXT on rdocx comment add, would cover the common case. #97 asked for this layer, and
PR #112 left it out ("It composes on top of this primitive"); the index bug above is what makes composing
it by hand unreliable.

Environment: main at 9a7ed714 (S75), release build, linux x86_64, python-docx 1.2.0 for the fixture.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions