Skip to content

rdocx-py: table formatting, table insertion and merges, and section edits - #187

Open
hadim wants to merge 4 commits into
tensorbee:mainfrom
hadim:feat/py-tables-and-sections
Open

hadim wants to merge 4 commits into
tensorbee:mainfrom
hadim:feat/py-tables-and-sections

Conversation

@hadim

@hadim hadim commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Bind the checked table, row and cell formatting setters. Table gains
    set_borders, set_border, border, cell_margins, set_cell_margins,
    grid_widths (read and assign) and set_column_width. Row gains
    height and height_rule with a new WD_ROW_HEIGHT_RULE enum as in
    python-docx, plus the tri-state cant_split and is_header. Cell gains
    shading, border, set_border, margins and set_margins.
  • Add Document.insert_table(index, rows, cols) and cell merges through the
    checked table operations, Table.set_cell_grid_span(row, col, span) and
    Table.set_cell_vertical_merge(row, col, merge), with Cell.grid_span
    and Cell.vertical_merge to read the result.
  • Add Document.update_section(index, **values), insert_section(index) and
    remove_section(index). Section stays a frozen snapshot, and
    update_section returns the new one. rdocx-py now depends on
    rdocx-oxml, as rdocx-cli already does, to name the orientation and break
    types the section setters take.
  • Make the native Section margin, gutter and header and footer distance
    setters write a complete w:pgMar. A section made by insert_section has
    none, and setting its margins wrote an element without the gutter,
    header and footer attributes that CT_PageMar requires.

Part of #168. This covers checklist items A (tables) and C (sections). Item B (styles and lists) is not
part of this PR and is left for a follow-up, see the notes. Items D
(bookmarks and fields), E (section stories), F (counted replacement), H
(rendering) and the native-feature items I to O remain, as planned in their
own PRs.

Why

Python tables stopped at the python-docx basics and Python sections were read
only, so the production chain of #168 kept an lxml pass for borders, shading,
cell and table margins, grid widths, row height, rows kept on one page,
repeating header rows, a table placed between two paragraphs, merged cells,
and page margins or orientation. Every one of those already had a checked
entry point in the rdocx facade. doc.sections[0].margin_top = 720 failed
with "attribute 'margin_top' of 'builtins.Section' objects is not writable".

Notes

  • The binding changes are additive and confined to the unpublished
    rdocx-py crate. The one native change is the w:pgMar completion in
    rdocx: Section::set_margins, set_gutter and
    set_header_footer_distance fill every other w:pgMar value the section
    lacks with the default that layout already assumes for it (one inch margins,
    half inch header and footer distances, no gutter), so the element carries
    all seven required attributes and the pages render as before. A value the
    section already has is kept. After such a call the section reports those
    filled values, for example header_distance of half an inch after
    update_section sets the margins of an inserted section. The legacy
    final-section Document setters and the serializer are unchanged, so a
    foreign part with an incomplete w:pgMar still round-trips as it was. The
    hash harness has no delta, because the samples use the legacy setters. The
    rdocx crates.io archive row is re-recorded in the last commit, and the lock
    file change is the new rdocx-oxml dependency edge of rdocx-py.
  • API choices:
    • Rows. Row.height plus Row.height_rule, as python-docx does, with
      WD_ROW_HEIGHT_RULE.AT_LEAST (1) and EXACTLY (2). AUTO is left out
      because the native row height has no auto form, the same way the table
      enums leave out values the facade cannot write. Assigning a height keeps
      the rule of an exact height and otherwise writes a minimum. Assigning a
      rule needs a height to apply to and raises ValueError before one
      exists. The native reader reports no height for a w:trHeight with an
      auto rule or without a value, so both properties read None there,
      where python-docx reads the value and AUTO, and assigning a height
      then writes a minimum. The stub and HLD 10 say so. Reading those forms
      would need a native accessor for the raw rule. cant_split and
      is_header are optional booleans, and None removes the direct value.
    • Merges. Merges go through set_cell_grid_span_checked and
      set_cell_vertical_merge on Table. The unchecked Cell::set_grid_span
      the issue cites is not bound, because it writes w:gridSpan without
      removing the absorbed cells. A python-docx shaped Cell.merge(other) can
      be composed from these two later.
    • Sections. Sections are edited through Document.update_section(index, **values), whose keywords are the Section snapshot's own field names.
      A value given alone keeps its partners (page size, the four margins,
      column count and spacing, header and footer distance) from the section's
      current explicit values. Column partners come from the equal-width view
      the snapshot reports, so on a section laid out in unequal-width tracks
      (w:equalWidth="0" or w:col children) column_count or
      column_spacing alone raises ValueError rather than rewriting the
      tracks, and both together replace them with equal-width columns.
  • Choices made here that the maintainer may want to revisit, with the
    recommended answer:
    • A partner the section never set (for example margin_top alone on a
      section inserted with insert_section, which has no w:pgMar) raises
      ValueError naming the missing field rather than being filled with the
      layout default of one inch. Recommended: keep, so the binding never
      invents a value the caller edits. The native w:pgMar completion above
      is separate: it fills only attributes the schema requires, with the
      values layout already used.
    • update_section is atomic. Names and partners are checked before any
      change, and when a native setter rejects a later value after an earlier
      one applied, the section properties are restored from a clone. The
      native Section setters are individually checked only, so this is the
      one place where the binding adds staging of its own. Recommended: keep,
      or move the same staging into a native Section::update later.
    • Border style and edge names use the OOXML spellings (dotDash,
      insideH), as the existing string values of the binding do
      (nextPage, darkBlue). Border size is an integer in eighths of a
      point, the native unit, with the native 0 to 96 validation. Recommended:
      keep both.
    • Row.height, Row.height_rule and Cell.shading cannot be assigned
      None, because the facade has no setter that removes them. Recommended:
      add native removal only when a caller needs it.
    • A grid span advances the revision once when it absorbs or restores cells,
      because later cell indexes move. A vertical merge and every formatting
      setter keep live handles valid, like Table.width today. Recommended:
      keep.
    • Orientation normalizes existing explicit page dimensions only, as the
      native setter does. A section without a page size (for example one made
      by insert_section) needs page_width and page_height in the same
      call to lay out landscape. Page size is applied before orientation.
  • Section.margin_top stays read only by design, so the last lines of
    probe_bindings.py from Production readiness for editing real docx and pptx files: an acceptance contract, two realistic fixtures and two matrices #158 still report the attribute as not writable.
    update_section is the entry point.
  • Item B is left out on purpose. This PR is already about 1,700 lines over
    three focused commits and an archive re-record, and B carries the API shape decisions with the
    longest life on a published wheel. Proposed shape for the follow-up:
    add_style(style_id, name, *, style_type="paragraph", based_on=None, next_style=None, linked_style=None, priority=None, quick_format=None, font_name=None, font_size=None, bold=None, italic=None, color=None, alignment=None, space_before=None, space_after=None, line_spacing=None, left_indent=None, first_line_indent=None, keep_with_next=None, keep_together=None, page_break_before=None) -> Style with the value types
    of Font and ParagraphFormat, set_style with the same keywords where
    None leaves a value and a clear= tuple of names removes one,
    remove_style, set_default_style, a frozen NumberingLevel,
    add_numbering_definition, add_numbering_instance,
    link_style_to_numbering and unlink_style_from_numbering. Two facts for
    that PR: ListNumberFormat has no public string conversion, so either
    the facade gains one or the binding maps the ST_NumberFormat names
    itself, and Python cannot see list marker text in Document.layout(), so
    "rendered numbered" is best checked by a Rust test walking the layout
    text.
  • Stub entries and tests are in their own groups, apart from the ones Expose comparison options as keywords of Python Document.compare #176
    and Align split_run, bookmarks and story comments on the direct body index #179 add. The Python tests sit next to the related table and section
    tests, and the native tests next to the related rdocx table and section
    tests.

Tests

  • crates/rdocx-py/tests/test_formatting_tables.py, placed after
    test_table_rows_are_cloned_with_their_formatting_and_removed:
    • test_table_borders_margins_and_grid_widths_round_trip checks the
      w:tblBorders, w:tblCellMar and w:tblGrid XML and every getter after
      a reopen, including the table width and covering cell widths kept in
      step.
    • test_row_height_split_and_header_round_trip checks w:trHeight with
      w:hRule="exact", w:cantSplit, w:tblHeader, an explicit false and
      clearing with None.
    • test_cell_shading_borders_and_margins_round_trip checks w:shd,
      w:tcBorders and the margins after a reopen.
    • test_invalid_table_formatting_changes_nothing_and_keeps_handles_live
      runs fourteen rejected calls (unknown style, edge and rule, width over 96
      eighths, bad color, negative margins, wrong grid length, zero grid
      column, out-of-range column, negative column width, negative height,
      rule without height) and checks the package bytes are unchanged and the
      held handles still work.
    • test_insert_table_at_a_body_index_returns_the_new_table inserts
      between a block content control holding a table and a paragraph, checks
      the body order and that the returned handle edits the new table rather
      than the one inside the control, and that an index past the end raises
      before mutation.
    • test_cells_merge_through_the_checked_table_operations checks
      w:gridSpan, w:vMerge restart and continue, the reduced cell count,
      stale handles after a span but not after a vertical merge, the rejected
      merges (nonempty cell, no matching cell above, past the grid, unknown
      name, bad row) with unchanged bytes, and undoing both merges.
    • test_table_acceptance_workflow_writes_the_native_body is the A
      acceptance of the issue: a table inserted at a body index, two cells
      merged horizontally and vertically, borders, shading, margins, widths,
      row heights, a row kept on one page and a repeating header row, plus
      one table edge, table cell margins, the grid and a cell border. It pins
      the resulting w:body.
  • crates/rdocx-py/tests/test_core.py, placed after
    test_word_structure_snapshots_preserve_order_ownership_and_types:
    • test_update_section_margins_and_orientation_reach_the_layout is the C
      acceptance: margins and orientation of the section changed from Python,
      then layout_page(0) is 792 by 612 points and the first body fragment
      starts at the new left and top margins (144 and 36 points).
    • test_update_section_keeps_unnamed_partners_and_rejects_atomically
      checks partner filling for page size, distances and columns, and six
      rejected calls, two of which fail in a native setter after an earlier
      value applied, with unchanged bytes.
    • test_update_section_never_rewrites_unequal_width_columns reads a
      section with two explicit w:col tracks as no column count or spacing,
      checks that either value alone raises with unchanged bytes, and that
      both together replace the tracks.
    • test_insert_and_remove_section_restructure_the_body inserts a section,
      makes it landscape and checks that page 2 of the layout is landscape,
      then removes it, with stale handles, index errors and the sole section
      refusal.
    • test_section_edits_write_the_native_body edits every section keyword
      across an original and an inserted section, with partner filling, and
      pins the resulting w:body.
  • crates/rdocx/tests/integration_test.rs:
    • table_acceptance_workflow_writes_the_body_the_python_binding_pins and
      section_edits_write_the_body_the_python_binding_pins make the native
      calls the two Python workflows bind and pin the same w:body strings,
      so CI, which runs pytest and the rdocx tests but not the rdocx-py Rust
      tests, checks that the Python calls and the native calls write the same
      body. Both sides compare the body on one line with the indentation
      removed.
    • section_page_margin_setters_write_every_required_attribute sets the
      gutter, the margins and the distances of three inserted sections to the
      values layout assumes, checks that the deterministic PDF is byte
      identical to the untouched document and that all four w:pgMar
      elements carry the seven attributes, then that later setters keep the
      values the section already has.
  • typing_smoke.py exercises every new entry point, and the stub, mypy and
    stubtest agree.
  • The new tests call methods that origin/main does not have. Without the
    native w:pgMar change,
    section_page_margin_setters_write_every_required_attribute and
    section_edits_write_the_body_the_python_binding_pins fail, because an
    inserted section gains only the attributes its setter names.
  • Run on this branch: cargo fmt --all --check, cargo clippy -p rdocx -p rdocx-py --all-targets --all-features -- -D warnings, python3 scripts/hash_harness.py --check (49 entries match), python3 scripts/prose_check.py, python3 scripts/sync_agent_skills.py --check
    and python3 scripts/readme_doctests.py (with the rdocx archive row
    re-recorded) are clean. cargo test -p rdocx --all-features --no-fail-fast passes apart from the environment-only failures below
    (lib 469 passed, integration 317 passed, regression 561 passed, doctests
    2 passed). The Python gate on a maturin develop build in a Python 3.12
    venv with python-docx 1.2.0: pytest crates/rdocx-py/tests gives 77
    passed and 2 failed, and mypy --strict on typing_smoke.py and the
    package and mypy.stubtest rdocx are clean.
  • Environment-only failures seen: the two test_rendering_threads.py tests
    assert the pinned pdfinfo 26.01.0 against the local 26.09.0. In cargo test -p rdocx, the lib tests
    large_word_and_presentation_pdfs_preserve_logical_reading_order and
    word_and_powerpoint_chart_pixels_are_identical and the integration tests
    odt_reader_matches_pinned_libreoffice_structure,
    public_authored_theme_and_fonts_match_pinned_word_resolution,
    sanitized_public_authoring_fixture_passes_every_conformance_stage,
    section_page_semantics_match_pinned_libreoffice_render and
    every_conditional_table_region_matches_word assert pinned tool
    versions (LibreOffice 26.2.5.2, pdftotext and the Poppler rasterizer)
    against the local ones, or, for the conformance harness, fail in its
    offline external consumer build. None of them uses the changed section
    setters.

Python tables stopped at the python-docx basics, so a document that
needs borders, shading, cell or table margins, grid widths, a row
height, a row kept on one page or a repeating header row still needed
an lxml pass, although each of those has a checked setter in the rdocx
facade.

Table gains set_borders, set_border, border, cell_margins,
set_cell_margins, grid_widths and set_column_width. Row gains height
and height_rule, with a WD_ROW_HEIGHT_RULE enum as in python-docx, plus
cant_split and is_header. Cell gains shading, border, set_border,
margins and set_margins. Each call goes to the checked native setter,
so a rejected value leaves the document unchanged, and none of them
moves content, so live handles stay valid. The native reader reports
no height for a w:trHeight with an auto rule or without a value, so
Row.height reads None there, unlike python-docx, and assigning a height
writes a minimum. The stub and HLD 10 say so.

GitHub issue tensorbee#168.
Python could only append a table, and it had no way to merge cells, so
a table placed between two paragraphs or a merged header still needed
lxml. The native facade has insert_table and the checked table-level
merge operations, while the Cell span setter it also has is unchecked
and leaves an invalid grid behind.

Document.insert_table(index, rows, cols) rejects an index past the end
before mutation and returns a handle to the new table. Table handles
count tables inside block content controls too, so the handle is found
from the inserted body position rather than assumed.
Table.set_cell_grid_span and Table.set_cell_vertical_merge call the
checked operations, which validate the whole table first. A span that
absorbs or restores cells advances the revision, because later cell
indexes move, and a vertical merge keeps handles valid. Cell.grid_span
and Cell.vertical_merge read the result.

A Python test runs the table acceptance workflow of the issue, with
the formatting setters, through the binding, and an rdocx integration
test makes the same native calls. Both pin the same body XML, so CI
checks that the binding writes what the native calls write.

GitHub issue tensorbee#168.
Python sections were frozen snapshots with no mutation entry point, so
assigning margin_top failed with "attribute is not writable", while the
native facade has checked setters for every section value and staged
insert_section and remove_section.

Document.update_section(index, **values) takes the snapshot's own field
names as keywords and returns the new snapshot, so Section stays a
frozen value. The native setters write page size, margins, columns and
header and footer distances together, so a value given alone keeps its
partners as the section has them, and a partner the section never set
raises ValueError instead of being invented. Column partners come from
the equal-width view the snapshot reports, so a column count or spacing
given alone raises rather than rewriting unequal-width tracks. Every
name and partner is checked before any change, and a value a native
setter rejects restores the section, so a failed call leaves the
document unchanged. Section edits move no content and keep handles
valid. insert_section and remove_section advance the revision because
they change the body. rdocx-py now depends on rdocx-oxml, as rdocx-cli
does, to name the orientation and break types the facade setters take.

A section made by insert_section has no w:pgMar, so setting its margins
wrote an element without the gutter, header and footer attributes that
CT_PageMar requires. The native Section margin, gutter and distance
setters now fill every w:pgMar value the section lacks with the default
that layout already assumes for it, so the element is complete and the
pages render as before.

A Python test and an rdocx integration test make the same section
edits through the binding and natively, and both pin the same body XML.

GitHub issue tensorbee#168.
The native section setters that now write a complete w:pgMar and the
integration tests that pin the body the Python table and section
workflows write grow the rdocx package, so its README archive row and
its ARCHIVE_MEASUREMENTS entry are re-measured.

GitHub issue tensorbee#168.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant