Skip to content

rdocx Python bindings: the complete list of what a production editing chain still needs, as one checklist #168

Description

@hadim

I am replacing a python-docx + lxml + LibreOffice chain with rdocx for a set of long, heavily
formatted documents that go through Word and Google Docs between passes. Reading, anchored replacement,
comments, revisions, fields, layout and rendering are there. What follows is everything else the chain
does today, checked against main at 9a7ed714. I am filing it as one list on purpose: the binding
rounds so far (#76, #94, #121) have followed whatever the previous bench hit, and I would rather give you the whole
remaining surface at once, with an acceptance line per item, than discover it one ticket at a time.

probe_bindings.py (in the zip attached to #158) prints the Rust entry point found by grep and
whether the Python class has an attribute for it; its output is summarised below.

Exists in Rust, not bound

  • Tables. Document::insert_table (document.rs:14960); Cell::set_grid_span (table.rs:1865) and
    Table::set_cell_vertical_merge (table.rs:1360); Cell::set_border_checked (table.rs:1794),
    set_shading (table.rs:1772), set_margins_checked (table.rs:1818); Table::set_grid_widths
    (table.rs:1047) and set_column_width (Python has Table.width and Cell.width only); table-level
    borders Table::set_borders (table.rs:841); Row::set_height and set_cant_split (table.rs:1494 to
    1554); Row::set_header (table.rs:1533).
    Accept: a table inserted at a body index, two cells merged horizontally and vertically, borders,
    shading, margins, widths, a row height, a row kept on one page and a repeating header row set from Python, and the result identical to the
    same calls in Rust.
  • Styles and lists. add_style, set_style, remove_style, set_default_style
    (document.rs:17654 to 17769); add_numbering_definition, add_numbering_instance,
    link_style_to_numbering (document.rs:16824, 16925, 17399).
    Accept: a paragraph style created with a font, spacing and a numbering level, applied to a paragraph,
    and the paragraph rendered numbered.
  • Sections. Section is read-only in Python (attribute 'margin_top' of 'builtins.Section' objects is not writable); Rust has section_mut with set_margins, set_orientation, set_page_size,
    set_columns, set_break_type, and insert_section / remove_section.
    Accept: margins and orientation of section 1 changed from Python, then the layout reflects them.
  • Bookmarks and fields. bookmarks, add_bookmark (comments.rs:259, 396); Run::add_field
    (run.rs:502); insert_toc (document.rs:20993).
    Accept: a bookmark over a run range, a PAGEREF to it added to a run, update_layout_backed_fields()
    filling it.
  • Headers and footers per section, with rich content: create_section_story,
    link_section_story / unlink_section_story,
    set_raw_header_with_images / set_raw_footer_with_images (document.rs:18115, 15950, 15965).
    Python has only set_header(text) / set_footer(text) and set_story_text.
    Accept: a footer made of a text run, a tab and PAGE / NUMPAGES fields, created from Python for one
    section.
  • Replacement. In Python, the count contract the CLI has (--expect, Expose the library's real capabilities through the CLI and the Python bindings #76): try_replace_text(old, new, expect=N) raising without touching the document when the count differs, and a literal batch
    replace_all([(old, new, expected), ...]) that is all-or-nothing. PR Python: revisions, counted replacement and field updates #109 bound only the fallible
    try_replace_text because the native replace_text panics; the native replace_all
    (document.rs:21241) takes a map with no per-pair count and fails the same way, so this needs a fallible
    batch with counts rather than a binding of it as is. Today a dry run means
    Document.from_bytes(doc.to_bytes()) and a second call.
    Accept: a batch of three replacements, one of which matches twice where once was declared, leaves the
    document unchanged and names that pair.
  • Comparison options: filed separately (Python and CLI access to ComparisonOptions).
  • Rendering. to_pdf_with_fonts / load_fonts_from_dir (document.rs:22060, 22653),
    render_page_to_svg (22218), to_pdfa_deterministic (22034). The CLI has --font-dir for PDF only.
    Accept: a PDF rendered with a font directory given from Python.

Not found in Rust either

  • A paragraph text setter on Paragraph (Paragraph.text is read-only; set_story_text(item, text)
    does it through a StoryItem, keeping the first run's properties and leaving an empty run behind), and
    removing a run. Without them, rewriting a cloned paragraph means setting the text of its first run
    and blanking the others.
  • Retargeting or removing a hyperlink. Paragraph.add_hyperlink adds one; nothing changes the URL of
    an existing one or unwraps it, which is what repairing a reference list needs.
  • Resizing an existing picture (its wp:extent and a:ext), so that replace_image with a picture
    of another aspect ratio does not render distorted. Replace the bytes of an existing picture #96 asked replace_image to keep the extent, which
    suits a regenerated chart; a picture of another aspect ratio needs it changed.
  • An XML escape hatch in write mode. StoryItem.xml reads a paragraph's XML (PR Python: headers, footers, story text, hyperlink creation and story item XML #110: "an escape hatch
    for checks the typed API does not cover yet"); writing it back (or
    replacing a body element from an XML fragment) would cover the long tail no API will ever reach.
    pop_content and insert_content already move a body element between documents, which covers part of
    it. Read and
    write access to an arbitrary package part by name would do the same at package level.
  • A comment anchored on text rather than on run positions, asked in Address text inside a run, for comment anchors and edits #97 and left out of PR Split a run at a character offset #112 (filed
    again with the split_run index bug).

If some of these are out of scope by design, saying so is as useful to me as implementing them: I would
keep a thin lxml step for those and stop asking.

Environment: main at 9a7ed714 (S75), release build, linux x86_64, Python 3.11.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions