The page renders correctly, but the text a PDF reader extracts, searches or copies is wrong as soon as the
shaper forms a ligature, which Carlito (the bundled Calibri substitute, so every Calibri document) does for
ti, fi, ff and a few more. collect_glyph_usage (oxml-pdf/src/font.rs:46 to 53) pairs the i-th glyph
of a text run with the i-th character of its text. A ligature makes one glyph out of two characters, so every
later glyph of the run gets the character one place further; and because the ToUnicode map keeps one
character per glyph, first seen first kept (glyph_to_unicode: BTreeMap<u16, char>, font.rs:18, filled with
or_insert), the misassignment then spreads to every occurrence of that glyph in the document, in runs
without any ligature. On the umbrella's report fixture (#158), the table header reads "Locaton",
"Ratnn" and "Acton". rpptx is not affected in my tests: the same text in a deck extracts correctly.
import subprocess
import rdocx
from docx import Document
from docx.shared import Pt
TEXT = "Location Rating Action fifteen office Observation"
d = Document()
for font in ("Calibri", "Arial", "Cambria"):
for bold in (False, True):
r = d.add_paragraph().add_run(f"{font} {'bold' if bold else 'regular'}: {TEXT}")
r.font.name, r.font.size, r.bold = font, Pt(11), bold
d.save("lig.docx")
open("lig.pdf", "wb").write(rdocx.Document.open("lig.docx").to_pdf())
print(subprocess.run(["pdftotext", "lig.pdf", "-"], capture_output=True, text=True).stdout.replace("\n\n", "\n").strip())
Calibri regular: Locatoo Ratog Actoo fieeo ofce Observatoo
Calibri bold: Locatoo Raton Actoo fiffo ofcf Obsfrvatoo
Arial regular: Location Rating Action fifteen office Observation
Arial bold: Location Rating Action fifteen office Observation
Cambria regular: Location Rating Action fifteen office Observation
Cambria bold: Location Rating Action fifteen office Observation
The multilingual path is not affected in practice: oxml-pdf/src/writer.rs wraps every MultilingualText
run in an ActualText span carrying its logical text (around line 1336), which a reader uses in
place of the ToUnicode map. Plain Text runs, the ones rdocx emits for this document, get no such span.
(Its ToUnicode entries, font.rs:63 to 74, keep only the first character of each cluster, so a ligature
there would still lose a character in a reader that ignores ActualText.)
Two earlier fixes touched the neighbourhood: #23 stopped glyphs being drawn twice at ligatures by slicing
shaped text by cluster (PR #24), and #74 made the text layer keep logical lines. The ToUnicode map of plain
runs still pairs glyphs and characters by index.
Acceptance: pdftotext of a rendered document returns its text for every bundled family, regular and bold,
ligatures included: the glyph-to-text mapping goes through the shaping clusters, and a ligature glyph maps to
all its characters (a multi-character ToUnicode entry, or ActualText). This matters for search, copy and
paste, accessibility, and the tagged PDF/A output that shares the writer.
Environment: main at 9a7ed714 (S75), release build, linux x86_64, Python 3.11, python-docx 1.2.0 for the
fixture, pdftotext from poppler.
The page renders correctly, but the text a PDF reader extracts, searches or copies is wrong as soon as the
shaper forms a ligature, which Carlito (the bundled Calibri substitute, so every Calibri document) does for
ti,fi,ffand a few more.collect_glyph_usage(oxml-pdf/src/font.rs:46to53) pairs the i-th glyphof a text run with the i-th character of its text. A ligature makes one glyph out of two characters, so every
later glyph of the run gets the character one place further; and because the ToUnicode map keeps one
character per glyph, first seen first kept (
glyph_to_unicode: BTreeMap<u16, char>,font.rs:18, filled withor_insert), the misassignment then spreads to every occurrence of that glyph in the document, in runswithout any ligature. On the umbrella's report fixture (#158), the table header reads "Locaton",
"Ratnn" and "Acton". rpptx is not affected in my tests: the same text in a deck extracts correctly.
The multilingual path is not affected in practice:
oxml-pdf/src/writer.rswraps everyMultilingualTextrun in an
ActualTextspan carrying its logical text (around line 1336), which a reader uses inplace of the ToUnicode map. Plain
Textruns, the ones rdocx emits for this document, get no such span.(Its ToUnicode entries,
font.rs:63to74, keep only the first character of each cluster, so a ligaturethere would still lose a character in a reader that ignores
ActualText.)Two earlier fixes touched the neighbourhood: #23 stopped glyphs being drawn twice at ligatures by slicing
shaped text by cluster (PR #24), and #74 made the text layer keep logical lines. The ToUnicode map of plain
runs still pairs glyphs and characters by index.
Acceptance:
pdftotextof a rendered document returns its text for every bundled family, regular and bold,ligatures included: the glyph-to-text mapping goes through the shaping clusters, and a ligature glyph maps to
all its characters (a multi-character ToUnicode entry, or
ActualText). This matters for search, copy andpaste, accessibility, and the tagged PDF/A output that shares the writer.
Environment:
mainat9a7ed714(S75), release build, linux x86_64, Python 3.11, python-docx 1.2.0 for thefixture,
pdftotextfrom poppler.