Conversation
The ToUnicode map of a plain glyph run paired the i-th glyph with the i-th character of its text. A ligature draws several characters with one glyph, so every later glyph of the run got the character one place further, and the font-wide map, first seen kept, spread that wrong character to every run using the glyph. Carlito, the bundled Calibri, joins ti, fi, ft, ff, ffi and more, so the extracted text of every Calibri document read like "Locatoo Ratog Actoo fieeo ofce". A plain run carries no shaping clusters, and carrying them would change public structs of oxml-layout, whose GlyphRun is built with a struct literal in several crates. So the pairing is rebuilt inside oxml-pdf from the font. A glyph draws the next character when the cmap gives it that character, and the next few when a GSUB ligature joins their glyphs into it, which also splits adjacent ligatures such as the fi and ft of "fifteen". Glyphs the font explains neither way share the characters up to the next glyph it does explain, and when none comes within 16 glyphs and characters, a glyph takes one character by position. A glyph that draws none of its group's characters by itself, such as the mark a shaper adds for a character the font has no glyph for, gets no text, since a plain run has no ActualText and would extract that character twice. The map value becomes a string, written as a multi-character ToUnicode entry, and a glyph keeps its strongest pairing instead of the first one seen. A GSUB ligature comes before the cmap, because Carlito draws "fi" and U+FB01 with one glyph, and one literal U+FB01 must not turn every "fi" of the document into U+FB01. Widths are declared for every drawn glyph, including one that carries no text of its own. GitHub issue tensorbee#171.
A MultilingualText run carries the shaper's clusters, but the ToUnicode map took only the first character of each cluster and gave it to every glyph of the cluster. A Devanagari conjunct that draws three characters with one glyph lost two of them, and the mark glyph of a cluster took its base character, which could then stand for that base in the whole font. Readers that honour the run's ActualText did not see it, but the map is what every other reader and search index uses. Each cluster now goes through the pairing plain runs use, so its characters go to the glyphs the font says draw them, whatever their visual order. A glyph that draws none of them by itself, such as the dots an Arabic font draws apart from their letter, repeats the text of its cluster at the weakest pairing, so every glyph of a rich run still maps to Unicode as it did before. The run's ActualText covers that repetition. A plain run has none, which is why it leaves such a glyph without text. GitHub issue tensorbee#171.
The issue reported rpptx as unaffected, because its deck inherits rtl="0" from the master, which sends every paragraph down the rich path with clusters and ActualText. A paragraph with no inherited direction draws its Latin text as plain runs, and their Carlito ligatures garbled the PDF text as in rdocx. The test builds a text box in a deck whose presentation and master declare no rtl, checks that it takes the plain path, and decodes the PDF text with lopdf, and with pdftotext when it is installed. GitHub issue tensorbee#171.
The issue's acceptance is that pdftotext returns the text of a rendered document for every bundled family, regular and bold, ligatures included. Nothing tested the meaning of the ToUnicode maps, so the hash harness baselined the wrong ones. The test renders the issue's sentence in Calibri, Arial, Cambria, Times New Roman and Courier New, regular and bold, and decodes the PDF with the ToUnicode decoder of the header and footer PDF tests. When pdftotext is installed it also compares its output line by line, and otherwise says it skipped that half. Without the fix it prints the issue's rows, "Calibri regular: Locatoo Ratog Actoo fieeo ofce". GitHub issue tensorbee#171.
The ToUnicode CMaps of every Calibri sample change: ligature glyphs map to all their characters, and glyphs that index pairing had given a neighbour's character map to their own. The harness moves exactly the pdf/resources and pdf/bytes entries of the seven samples. No pdf/pages, PNG or OOXML entry moves, because the subset fonts, the widths and every content stream are byte-identical, which an object by object comparison of the old and new sample PDFs confirms. GitHub issue tensorbee#171.
The oxml-pdf package grows with the new ToUnicode pairing and its unit tests, and the rdocx and rpptx packages carry their integration tests, which gain the ligature cases. The three rows in readme_doctests.py and in the crate READMEs now hold the measured sizes, dated 2026-09-27 through ARCHIVE_REMEASUREMENT_DATES, so the Docs job keeps passing. GitHub issue tensorbee#171.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Carrying the shaper's clusters would change public structs of
oxml-layout, so the pairing is rebuilt insideoxml-pdffrom the font.GlyphRun(andTextSegmentandShapedText) in the publishedoxml-layout0.12.1 are plain structs with only public fields and no
#[non_exhaustive],and
GlyphRunalone is built with a struct literal in six other crates ofthis workspace (
rdocx-layout,rpptx-render,oxml-chart,rpptx,rdocxand
oxml-pdf), so a new field breaks every such caller, inside and outsidethe workspace. The only additive carrier would be a new
PositionedElementvariant, which every producer and every backend would have to learn, and
routing plain Latin runs through
MultilingualTextinstead moves everycontent stream (it adds
ActualTextspans). So this PR reads the pairing fromthe font and changes no public type or signature. The cmap alone cannot split
two adjacent ligatures (Carlito shapes "fifteen" as
[fi][ft]een), so thepairing also reads the GSUB ligature table, which says exactly which glyphs
each ligature joins.
Summary
oxml-pdf, plain runs:collect_glyph_usageno longer pairs the i-thglyph with the i-th character. It walks the run with the font. A glyph draws
the next character when the cmap gives it that character, and the next few
when a GSUB ligature joins their nominal glyphs into it (read once per font,
nested ligatures followed two levels deep). Glyphs the font explains neither
way share the characters up to the next glyph it does explain, looked for
within 16 glyphs and characters: one glyph takes them all, as many glyphs as
characters pair by position, and otherwise the first glyph carries them.
When no explained glyph comes within that window, one glyph takes one
character by position, which keeps the walk linear. A glyph that draws none
of its group's characters by itself (the combining mark a shaper adds for a
character the font has no precomposed glyph for) gets no text, and neither
does a glyph after the run's last character, because a plain run has no
ActualTextand would otherwise extract a character twice.oxml-pdf, the font-wide map: the ToUnicode value becomes a string writtenwith
pair_with_multiple, so a ligature maps to all its characters. The mapkeeps one entry per glyph, and it now keeps the strongest pairing instead of
the first one seen: a GSUB ligature, then the cmap, then an inferred
pairing, and the first one seen among equals. The ligature outranks the
cmap because Carlito draws "fi" and U+FB01 with one glyph, and one literal
U+FB01 must not turn every "fi" of the document into U+FB01. A document
that only ever writes U+FB01 keeps it. Widths are declared for every drawn
glyph.
oxml-pdf, rich runs: each shaping cluster of aMultilingualTextrun goesthrough the same pairing, so no character of a cluster is lost. A
Devanagari conjunct maps to its three characters, and in "त्रि" the vowel
sign and the conjunct each get their own text although the vowel sign is
drawn first. A glyph that draws none of its cluster's characters by itself
(the dots Noto Sans Arabic draws apart from their letter) repeats the
cluster's text at the weakest pairing, which the run's
ActualTextcovers,so every glyph of a rich run still maps to Unicode, as before.
docs/hld/08-rendering-spec.md: the text layer paragraph, including theposition fallback, the glyphs a plain run leaves without text, and the
precedence of the font-wide map.
rpptx: regression test for a deck whose text takes the plain path.rdocx: the issue's matrix as an integration test, for all five bundledLatin families, regular and bold.
scripts/hash_baseline.json: the 14 PDF entries re-recorded in their owncommit,
pdf/resourcesandpdf/bytesof the seven samples.oxml-pdf,rdocxandrpptxre-recorded (the last twopackages carry their integration tests), dated 2026-09-27 through
ARCHIVE_REMEASUREMENT_DATES.Closes #171.
Why
The page rendered correctly, but the text a reader extracts, searches or copies
from any Calibri document was garbled. Carlito, the bundled Calibri, forms
ti,fi,ft,ff,ffi,ffl,tti,fjand more. The writer pairedglyphs and characters by index, so one ligature shifted every later glyph of
its run, and the font-wide map, first seen kept, then spread the wrong
character to every run using that glyph. The issue's fixture extracted
"Locatoo Ratog Actoo fieeo ofce Observatoo". The same happened to rpptx
whenever a paragraph did not inherit
rtl(the issue's deck inheritsrtl="0", which is why rpptx looked unaffected), and to chart labels.The issue's reproduction, rerun on this branch (its Python script unchanged,
with a release wheel of
rdocx-pybuilt from the branch, macOS arm64, Popplerpdftotext26.09.0, Calibri resolving to the bundled Carlito and Cambria tothe bundled Caladea):
On the seven samples the new CMaps change nothing but wrong entries. Besides
the ligature glyphs, index pairing had given ordinary glyphs a neighbour's
character, and pdftotext read the contract as "Innotacon Drite, San
Franeiseo" and the report as "Executie Oieriiew" and "redufton in farbon
emissions". After the change, every word pdftotext extracts from the seven
samples is a dictionary word, a name or an acronym.
Notes
pdftotextrun in the ordinary suites and skip with amessage when it is absent. Each such test also has a hermetic half that
always runs: the rdocx test decodes the ToUnicode maps with the existing
pdf_page_texthelper ofmod header_footer_pdf, and the rpptx testdecodes the PDF with
lopdf. If clusters are later carried onGlyphRunina breaking release,
pair_plain_runis the one function to replace, and therest of this PR (string ToUnicode values, precedence, cluster pairing of rich
runs, the tests) stays.
PDF/UA-1 7.21.7 asks every used character code to map to Unicode. Every
glyph of a rich run keeps an entry, as before. A plain run leaves a glyph
without text only when it draws none of its group's characters or follows
the run's last character. Index pairing on
mainleft the surplus glyphs ofsuch a run without an entry too, the trailing ones rather than the marks, so
this PR does not widen that gap.
ActualTexton plain runs would close it./Lengthmove. Thesubset fonts, the
Warrays and every content stream are byte-identical, sothe hash harness moves
pdf/resourcesandpdf/bytesof the seven samplesand no
pdf/pages, PNG or XML entry. I compared the old and new sample PDFsobject by object to confirm it. The subset remap order is unchanged (glyphs
are still remapped in drawing order). The review fixes (no text for a plain
run's mark glyph, ligature before cmap) move no harness entry, since no
sample draws such a mark or a literal U+FB01.
Warray now lists every drawn glyph, where it listed everyglyph that had text. The two sets differ only when a plain run draws more
glyphs than it has characters, which no sample does. Such a glyph used to
fall back to the zero default width and now declares its real one. Its
painted position does not change, since the
TJadjustment compensateseither way, but its content stream bytes do.
glyph that draws different text in different places keeps one of them, the
strongest. In plain runs, Carlito draws the plain Greek capital for a tonos
capital inside an all-caps word, so "ΆΈΡ" extracts as "ΑΕΡ" when the plain
capitals also appear on their own. A mark glyph that a plain run leaves
without text still extracts whatever another run pairs it with, so with a
literal U+0323 elsewhere in the document a decomposed "ṩ" extracts with an
extra dot below, and a base glyph keeps its own letter when the document
also uses it plainly, so a decomposed "ȁ" can extract as "a". A literal
U+FB01 in a document that also has an ordinary "fi" extracts as "fi", its
compatibility decomposition. In rich runs, a mark glyph that Noto Sans
Arabic shares between letters keeps the first letter's text, which
ActualTextalready covers for readers that honour it. Wrapping plain runsin
ActualTextwould lift all of these, but it moves every content stream,so it is left out of this PR.
distribute_justify_advancesspreads a justified chunk's extra space overevery glyph whenever a ligature breaks the one glyph per character count.
Fixing it exactly needs clusters on the plain path. It deserves its own
issue.
that also re-records it does so again after this one lands.
oxml-pdfarchive row also moves in Paint slide backgrounds in PDF output and open gradients without ang or path #175, and therdocxandrpptxrows in most open PRs. Whichever lands later re-recordsthem. The new tests sit inside existing groups, not at the end of a file:
the rdocx one in
mod header_footer_pdfaftersource_less_numbering_marker_keeps_logical_pdf_and_svg_order, the rpptxone after the PDF import tests.
runs, and the second replaces only that loop, so each can be read alone. The
two test commits follow, then the baseline and archive commits.
integration test is the issue's Python reproduction in Rust.
Tests
Added:
oxml-pdffont::tests(a new test module infont.rs, kept out ofwriter.rs, which Paint slide backgrounds in PDF output and open gradients without ang or path #175 edits):a_ligature_glyph_maps_to_every_character_it_draws: Carlito regular andbold, ligature words first (the order that poisoned the map), every run
decodes back to its text through the prepared CMap and subset remapper,
and the
ti,fi,ft,ffi,ffl,tti,fjandffglyphs map tothose strings.
a_ligature_outranks_the_presentation_form_sharing_its_glyph: Carlitodraws "fi" and U+FB01 with one glyph. Without a literal U+FB01, "office
fifteen" extracts as itself. With one seen first, it still does, and the
literal extracts as "fi". A document with only the literal keeps U+FB01.
a_plain_run_extracts_each_character_once: Liberation Sans draws "≮" and"≯" as
<or>and a combining overlay, and Carlito draws "Ѷ" as "Ѵ" anda combining double grave. Each run extracts its text once.
a_cmap_pairing_outranks_one_inferred_earlier: Carlito draws its spaceglyph for a no-break space. Seen first, that used to claim the glyph.
a_multilingual_cluster_keeps_every_character: clusters from the realmultilingual shaper. Every glyph maps to Unicode and every cluster keeps
all its characters, with exact values for "क्ष", "त्रि", "कि" and the
dotted "ب".
a_run_without_ligatures_keeps_one_character_per_glyph: Liberation Sanskeeps one character per glyph and a width for every drawn glyph.
rdocxintegration_test.rs,header_footer_pdf::ligatures_keep_their_characters_in_pdf_text_of_every_bundled_family:the issue's sentence in Calibri, Arial, Cambria, Times New Roman and Courier
New, regular and bold, through
to_pdf_deterministic, decoded bypdf_page_textand, when present, by pdftotext line by line.rpptxintegration.rs,plain_run_ligatures_keep_their_characters_in_pdf_text:a text box in a deck with
rtl="0"removed from the presentation and themaster, checked to take the plain path, decoded with
lopdfand, whenpresent, with pdftotext.
Run against the pairing of
main, the first five unit tests fail("Locatoon", "office iif", "a ≮" followed by blanks, the space glyph mapped
to U+00A0, a conjunct losing two characters), and so do both integration
tests. The rdocx one prints the issue's rows exactly: "Calibri regular:
Locatoo Ratog Actoo fieeo ofce Observatoo" and "Calibri bold: Locatoo Raton
Actoo fiffo ofcf Obsfrvatoo". The precedence and mark tests also fail if a
ligature ranks below the cmap ("office fifteen") or if a plain run's mark
glyph repeats its group's text ("a ≮≮ b ≯≮ c").
Run:
cargo fmt --all --check: clean.cargo clippy -p oxml-pdf -p rdocx -p rpptx --all-targets --all-features -- -D warnings: clean,and
-p oxml-pdfalso clean at the first commit alone.cargo test -p oxml-pdf,-p rdocx,-p rpptxat HEAD, and also-p oxml-chart,-p rdocx-pdf,-p rpptx-render,-p rdocx-cliand-p rpptx-cli, the other crates that depend onoxml-pdf: everythingpasses except the environment-only failures listed below.
python3 scripts/hash_harness.py --check: before the baseline commit,exactly the 14 expected entries (
pdf/resourcesandpdf/bytesof the sevensamples). After it, 49 of 49 match.
python3 scripts/prose_check.py: 0 violations.python3 scripts/readme_doctests.py: passes after the archive re-record.repaired words.
rdocx convertof the Production readiness for editing real docx and pptx files: an acceptance contract, two realistic fixtures and two matrices #158 report fixture extracts with nogarbled word.
rdocx convertwith system fonts of a document of Latinletters the font draws as a base and marks ("ḉ", "ȁ", "ǭ" and more): no
mark glyph repeats its letter any more, where "ḉ" extracted as "ḉḉ" before.
A mark the same document also writes as a literal combining character
still extracts through that entry, the remaining limit in the notes.
Environment-only failures seen, all on the known list for this Mac:
oxml-pdfwriter::tests::rotated_linear_gradient_renders_with_its_axis_rotated(pinned pdftoppm),
rdocxliblarge_word_and_presentation_pdfs_preserve_logical_reading_orderandword_and_powerpoint_chart_pixels_are_identical,rdocxintegrationodt_reader_matches_pinned_libreoffice_structure,public_authored_theme_and_fonts_match_pinned_word_resolution,sanitized_public_authoring_fixture_passes_every_conformance_stage,section_page_semantics_match_pinned_libreoffice_renderandevery_conditional_table_region_matches_word(pinned tool versions), andrpptx-clivalidate_rejects_corruption_and_accepts_the_pinned_corpus(missing
/corpus/pptx).