Skip to content

Pair PDF glyphs with their text through the font, so ligatures extract - #188

Open
hadim wants to merge 6 commits into
tensorbee:mainfrom
hadim:fix/pdf-tounicode-ligatures
Open

hadim wants to merge 6 commits into
tensorbee:mainfrom
hadim:fix/pdf-tounicode-ligatures

Conversation

@hadim

@hadim hadim commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Carrying the shaper's clusters would change public structs of
oxml-layout, so the pairing is rebuilt inside oxml-pdf from the font.

GlyphRun (and TextSegment and ShapedText) in the published oxml-layout
0.12.1 are plain structs with only public fields and no #[non_exhaustive],
and GlyphRun alone is built with a struct literal in six other crates of
this workspace (rdocx-layout, rpptx-render, oxml-chart, rpptx, rdocx
and oxml-pdf), so a new field breaks every such caller, inside and outside
the workspace. The only additive carrier would be a new PositionedElement
variant, which every producer and every backend would have to learn, and
routing plain Latin runs through MultilingualText instead moves every
content stream (it adds ActualText spans). So this PR reads the pairing from
the font and changes no public type or signature. The cmap alone cannot split
two adjacent ligatures (Carlito shapes "fifteen" as [fi][ft]een), so the
pairing also reads the GSUB ligature table, which says exactly which glyphs
each ligature joins.

Summary

  • oxml-pdf, plain runs: collect_glyph_usage no longer pairs the i-th
    glyph with the i-th character. It walks the run with the font. A glyph draws
    the next character when the cmap gives it that character, and the next few
    when a GSUB ligature joins their nominal glyphs into it (read once per font,
    nested ligatures followed two levels deep). Glyphs the font explains neither
    way share the characters up to the next glyph it does explain, looked for
    within 16 glyphs and characters: one glyph takes them all, as many glyphs as
    characters pair by position, and otherwise the first glyph carries them.
    When no explained glyph comes within that window, one glyph takes one
    character by position, which keeps the walk linear. A glyph that draws none
    of its group's characters by itself (the combining mark a shaper adds for a
    character the font has no precomposed glyph for) gets no text, and neither
    does a glyph after the run's last character, because a plain run has no
    ActualText and would otherwise extract a character twice.
  • oxml-pdf, the font-wide map: the ToUnicode value becomes a string written
    with pair_with_multiple, so a ligature maps to all its characters. The map
    keeps one entry per glyph, and it now keeps the strongest pairing instead of
    the first one seen: a GSUB ligature, then the cmap, then an inferred
    pairing, and the first one seen among equals. The ligature outranks the
    cmap because Carlito draws "fi" and U+FB01 with one glyph, and one literal
    U+FB01 must not turn every "fi" of the document into U+FB01. A document
    that only ever writes U+FB01 keeps it. Widths are declared for every drawn
    glyph.
  • oxml-pdf, rich runs: each shaping cluster of a MultilingualText run goes
    through the same pairing, so no character of a cluster is lost. A
    Devanagari conjunct maps to its three characters, and in "त्रि" the vowel
    sign and the conjunct each get their own text although the vowel sign is
    drawn first. A glyph that draws none of its cluster's characters by itself
    (the dots Noto Sans Arabic draws apart from their letter) repeats the
    cluster's text at the weakest pairing, which the run's ActualText covers,
    so every glyph of a rich run still maps to Unicode, as before.
  • docs/hld/08-rendering-spec.md: the text layer paragraph, including the
    position fallback, the glyphs a plain run leaves without text, and the
    precedence of the font-wide map.
  • rpptx: regression test for a deck whose text takes the plain path.
  • rdocx: the issue's matrix as an integration test, for all five bundled
    Latin families, regular and bold.
  • scripts/hash_baseline.json: the 14 PDF entries re-recorded in their own
    commit, pdf/resources and pdf/bytes of the seven samples.
  • Archive rows of oxml-pdf, rdocx and rpptx re-recorded (the last two
    packages carry their integration tests), dated 2026-09-27 through
    ARCHIVE_REMEASUREMENT_DATES.

Closes #171.

Why

The page rendered correctly, but the text a reader extracts, searches or copies
from any Calibri document was garbled. Carlito, the bundled Calibri, forms
ti, fi, ft, ff, ffi, ffl, tti, fj and more. The writer paired
glyphs and characters by index, so one ligature shifted every later glyph of
its run, and the font-wide map, first seen kept, then spread the wrong
character to every run using that glyph. The issue's fixture extracted
"Locatoo Ratog Actoo fieeo ofce Observatoo". The same happened to rpptx
whenever a paragraph did not inherit rtl (the issue's deck inherits
rtl="0", which is why rpptx looked unaffected), and to chart labels.

The issue's reproduction, rerun on this branch (its Python script unchanged,
with a release wheel of rdocx-py built from the branch, macOS arm64, Poppler
pdftotext 26.09.0, Calibri resolving to the bundled Carlito and Cambria to
the bundled Caladea):

Calibri regular: Location Rating Action fifteen office Observation
Calibri bold: Location Rating Action fifteen office Observation
Arial regular: Location Rating Action fifteen office Observation
Arial bold: Location Rating Action fifteen office Observation
Cambria regular: Location Rating Action fifteen office Observation
Cambria bold: Location Rating Action fifteen office Observation

On the seven samples the new CMaps change nothing but wrong entries. Besides
the ligature glyphs, index pairing had given ordinary glyphs a neighbour's
character, and pdftotext read the contract as "Innotacon Drite, San
Franeiseo" and the report as "Executie Oieriiew" and "redufton in farbon
emissions". After the change, every word pdftotext extracts from the seven
samples is a dictionary word, a name or an acronym.

Notes

  • Tests that need pdftotext run in the ordinary suites and skip with a
    message when it is absent. Each such test also has a hermetic half that
    always runs: the rdocx test decodes the ToUnicode maps with the existing
    pdf_page_text helper of mod header_footer_pdf, and the rpptx test
    decodes the PDF with lopdf. If clusters are later carried on GlyphRun in
    a breaking release, pair_plain_run is the one function to replace, and the
    rest of this PR (string ToUnicode values, precedence, cluster pairing of rich
    runs, the tests) stays.
  • Glyphs without text: the writer declares PDF/UA-1 for tagged output, and
    PDF/UA-1 7.21.7 asks every used character code to map to Unicode. Every
    glyph of a rich run keeps an entry, as before. A plain run leaves a glyph
    without text only when it draws none of its group's characters or follows
    the run's last character. Index pairing on main left the surplus glyphs of
    such a run without an entry too, the trailing ones rather than the marks, so
    this PR does not widen that gap. ActualText on plain runs would close it.
  • Output change: only the ToUnicode CMap streams and their /Length move. The
    subset fonts, the W arrays and every content stream are byte-identical, so
    the hash harness moves pdf/resources and pdf/bytes of the seven samples
    and no pdf/pages, PNG or XML entry. I compared the old and new sample PDFs
    object by object to confirm it. The subset remap order is unchanged (glyphs
    are still remapped in drawing order). The review fixes (no text for a plain
    run's mark glyph, ligature before cmap) move no harness entry, since no
    sample draws such a mark or a literal U+FB01.
  • Widths: the W array now lists every drawn glyph, where it listed every
    glyph that had text. The two sets differ only when a plain run draws more
    glyphs than it has characters, which no sample does. Such a glyph used to
    fall back to the zero default width and now declares its real one. Its
    painted position does not change, since the TJ adjustment compensates
    either way, but its content stream bytes do.
  • Remaining limit: the CMap has one entry per glyph for the whole font, so a
    glyph that draws different text in different places keeps one of them, the
    strongest. In plain runs, Carlito draws the plain Greek capital for a tonos
    capital inside an all-caps word, so "ΆΈΡ" extracts as "ΑΕΡ" when the plain
    capitals also appear on their own. A mark glyph that a plain run leaves
    without text still extracts whatever another run pairs it with, so with a
    literal U+0323 elsewhere in the document a decomposed "ṩ" extracts with an
    extra dot below, and a base glyph keeps its own letter when the document
    also uses it plainly, so a decomposed "ȁ" can extract as "a". A literal
    U+FB01 in a document that also has an ordinary "fi" extracts as "fi", its
    compatibility decomposition. In rich runs, a mark glyph that Noto Sans
    Arabic shares between letters keeps the first letter's text, which
    ActualText already covers for readers that honour it. Wrapping plain runs
    in ActualText would lift all of these, but it moves every content stream,
    so it is left out of this PR.
  • Adjacent defect, not in this PR: rdocx
    distribute_justify_advances spreads a justified chunk's extra space over
    every glyph whenever a ligature breaks the one glyph per character count.
    Fixing it exactly needs clusters on the plain path. It deserves its own
    issue.
  • Hash baseline order: this PR re-records the hash baseline, so any open PR
    that also re-records it does so again after this one lands.
  • Shared files: the oxml-pdf archive row also moves in Paint slide backgrounds in PDF output and open gradients without ang or path #175, and the
    rdocx and rpptx rows in most open PRs. Whichever lands later re-records
    them. The new tests sit inside existing groups, not at the end of a file:
    the rdocx one in mod header_footer_pdf after
    source_less_numbering_marker_keeps_logical_pdf_and_svg_order, the rpptx
    one after the PDF import tests.
  • History: the first commit fixes plain runs and keeps the old loop for rich
    runs, and the second replaces only that loop, so each can be read alone. The
    two test commits follow, then the baseline and archive commits.
  • No binding or stub changes, so the Python gate does not apply. The rdocx
    integration test is the issue's Python reproduction in Rust.

Tests

Added:

  • oxml-pdf font::tests (a new test module in font.rs, kept out of
    writer.rs, which Paint slide backgrounds in PDF output and open gradients without ang or path #175 edits):
    • a_ligature_glyph_maps_to_every_character_it_draws: Carlito regular and
      bold, ligature words first (the order that poisoned the map), every run
      decodes back to its text through the prepared CMap and subset remapper,
      and the ti, fi, ft, ffi, ffl, tti, fj and ff glyphs map to
      those strings.
    • a_ligature_outranks_the_presentation_form_sharing_its_glyph: Carlito
      draws "fi" and U+FB01 with one glyph. Without a literal U+FB01, "office
      fifteen" extracts as itself. With one seen first, it still does, and the
      literal extracts as "fi". A document with only the literal keeps U+FB01.
    • a_plain_run_extracts_each_character_once: Liberation Sans draws "≮" and
      "≯" as < or > and a combining overlay, and Carlito draws "Ѷ" as "Ѵ" and
      a combining double grave. Each run extracts its text once.
    • a_cmap_pairing_outranks_one_inferred_earlier: Carlito draws its space
      glyph for a no-break space. Seen first, that used to claim the glyph.
    • a_multilingual_cluster_keeps_every_character: clusters from the real
      multilingual shaper. Every glyph maps to Unicode and every cluster keeps
      all its characters, with exact values for "क्ष", "त्रि", "कि" and the
      dotted "ب".
    • a_run_without_ligatures_keeps_one_character_per_glyph: Liberation Sans
      keeps one character per glyph and a width for every drawn glyph.
  • rdocx integration_test.rs, header_footer_pdf::ligatures_keep_their_characters_in_pdf_text_of_every_bundled_family:
    the issue's sentence in Calibri, Arial, Cambria, Times New Roman and Courier
    New, regular and bold, through to_pdf_deterministic, decoded by
    pdf_page_text and, when present, by pdftotext line by line.
  • rpptx integration.rs, plain_run_ligatures_keep_their_characters_in_pdf_text:
    a text box in a deck with rtl="0" removed from the presentation and the
    master, checked to take the plain path, decoded with lopdf and, when
    present, with pdftotext.

Run against the pairing of main, the first five unit tests fail
("Locatoon", "office iif", "a ≮" followed by blanks, the space glyph mapped
to U+00A0, a conjunct losing two characters), and so do both integration
tests. The rdocx one prints the issue's rows exactly: "Calibri regular:
Locatoo Ratog Actoo fieeo ofce Observatoo" and "Calibri bold: Locatoo Raton
Actoo fiffo ofcf Obsfrvatoo". The precedence and mark tests also fail if a
ligature ranks below the cmap ("office fifteen") or if a plain run's mark
glyph repeats its group's text ("a ≮≮ b ≯≮ c").

Run:

  • cargo fmt --all --check: clean.
  • cargo clippy -p oxml-pdf -p rdocx -p rpptx --all-targets --all-features -- -D warnings: clean,
    and -p oxml-pdf also clean at the first commit alone.
  • cargo test -p oxml-pdf, -p rdocx, -p rpptx at HEAD, and also
    -p oxml-chart, -p rdocx-pdf, -p rpptx-render, -p rdocx-cli and
    -p rpptx-cli, the other crates that depend on oxml-pdf: everything
    passes except the environment-only failures listed below.
  • python3 scripts/hash_harness.py --check: before the baseline commit,
    exactly the 14 expected entries (pdf/resources and pdf/bytes of the seven
    samples). After it, 49 of 49 match.
  • python3 scripts/prose_check.py: 0 violations.
  • python3 scripts/readme_doctests.py: passes after the archive re-record.
  • pdftotext (Poppler 26.09.0) on the old and new samples: the diff is only
    repaired words. rdocx convert of the Production readiness for editing real docx and pptx files: an acceptance contract, two realistic fixtures and two matrices #158 report fixture extracts with no
    garbled word. rdocx convert with system fonts of a document of Latin
    letters the font draws as a base and marks ("ḉ", "ȁ", "ǭ" and more): no
    mark glyph repeats its letter any more, where "ḉ" extracted as "ḉḉ" before.
    A mark the same document also writes as a literal combining character
    still extracts through that entry, the remaining limit in the notes.

Environment-only failures seen, all on the known list for this Mac: oxml-pdf
writer::tests::rotated_linear_gradient_renders_with_its_axis_rotated
(pinned pdftoppm), rdocx lib
large_word_and_presentation_pdfs_preserve_logical_reading_order and
word_and_powerpoint_chart_pixels_are_identical, rdocx integration
odt_reader_matches_pinned_libreoffice_structure,
public_authored_theme_and_fonts_match_pinned_word_resolution,
sanitized_public_authoring_fixture_passes_every_conformance_stage,
section_page_semantics_match_pinned_libreoffice_render and
every_conditional_table_region_matches_word (pinned tool versions), and
rpptx-cli validate_rejects_corruption_and_accepts_the_pinned_corpus
(missing /corpus/pptx).

The ToUnicode map of a plain glyph run paired the i-th glyph with the
i-th character of its text. A ligature draws several characters with
one glyph, so every later glyph of the run got the character one place
further, and the font-wide map, first seen kept, spread that wrong
character to every run using the glyph. Carlito, the bundled Calibri,
joins ti, fi, ft, ff, ffi and more, so the extracted text of every
Calibri document read like "Locatoo Ratog Actoo fieeo ofce".

A plain run carries no shaping clusters, and carrying them would change
public structs of oxml-layout, whose GlyphRun is built with a struct
literal in several crates. So the pairing is rebuilt inside oxml-pdf
from the font. A glyph draws the next character when the cmap gives it
that character, and the next few when a GSUB ligature joins their
glyphs into it, which also splits adjacent ligatures such as the fi and
ft of "fifteen". Glyphs the font explains neither way share the
characters up to the next glyph it does explain, and when none comes
within 16 glyphs and characters, a glyph takes one character by
position. A glyph that draws none of its group's characters by itself,
such as the mark a shaper adds for a character the font has no glyph
for, gets no text, since a plain run has no ActualText and would
extract that character twice.

The map value becomes a string, written as a multi-character ToUnicode
entry, and a glyph keeps its strongest pairing instead of the first one
seen. A GSUB ligature comes before the cmap, because Carlito draws "fi"
and U+FB01 with one glyph, and one literal U+FB01 must not turn every
"fi" of the document into U+FB01. Widths are declared for every drawn
glyph, including one that carries no text of its own.

GitHub issue tensorbee#171.
A MultilingualText run carries the shaper's clusters, but the ToUnicode
map took only the first character of each cluster and gave it to every
glyph of the cluster. A Devanagari conjunct that draws three characters
with one glyph lost two of them, and the mark glyph of a cluster took
its base character, which could then stand for that base in the whole
font. Readers that honour the run's ActualText did not see it, but
the map is what every other reader and search index uses.

Each cluster now goes through the pairing plain runs use, so its
characters go to the glyphs the font says draw them, whatever their
visual order. A glyph that draws none of them by itself, such as the
dots an Arabic font draws apart from their letter, repeats the text of
its cluster at the weakest pairing, so every glyph of a rich run still
maps to Unicode as it did before. The run's ActualText covers that
repetition. A plain run has none, which is why it leaves such a glyph
without text.

GitHub issue tensorbee#171.
The issue reported rpptx as unaffected, because its deck inherits
rtl="0" from the master, which sends every paragraph down the rich
path with clusters and ActualText. A paragraph with no inherited
direction draws its Latin text as plain runs, and their Carlito
ligatures garbled the PDF text as in rdocx.

The test builds a text box in a deck whose presentation and master
declare no rtl, checks that it takes the plain path, and decodes the
PDF text with lopdf, and with pdftotext when it is installed.

GitHub issue tensorbee#171.
The issue's acceptance is that pdftotext returns the text of a rendered
document for every bundled family, regular and bold, ligatures
included. Nothing tested the meaning of the ToUnicode maps, so the hash
harness baselined the wrong ones.

The test renders the issue's sentence in Calibri, Arial, Cambria, Times
New Roman and Courier New, regular and bold, and decodes the PDF with
the ToUnicode decoder of the header and footer PDF tests. When
pdftotext is installed it also compares its output line by line, and
otherwise says it skipped that half. Without the fix it prints the
issue's rows, "Calibri regular: Locatoo Ratog Actoo fieeo ofce".

GitHub issue tensorbee#171.
The ToUnicode CMaps of every Calibri sample change: ligature glyphs map
to all their characters, and glyphs that index pairing had given a
neighbour's character map to their own. The harness moves exactly the
pdf/resources and pdf/bytes entries of the seven samples. No pdf/pages,
PNG or OOXML entry moves, because the subset fonts, the widths and every
content stream are byte-identical, which an object by object comparison
of the old and new sample PDFs confirms.

GitHub issue tensorbee#171.
The oxml-pdf package grows with the new ToUnicode pairing and its unit
tests, and the rdocx and rpptx packages carry their integration tests,
which gain the ligature cases. The three rows in readme_doctests.py and
in the crate READMEs now hold the measured sizes, dated 2026-09-27
through ARCHIVE_REMEASUREMENT_DATES, so the Docs job keeps passing.

GitHub issue tensorbee#171.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

rdocx PDF text layer: glyphs are mapped to characters by position, so every ligature garbles the text of Calibri documents

1 participant