Skip to content

rdocx PDF text layer: glyphs are mapped to characters by position, so every ligature garbles the text of Calibri documents #171

Description

@hadim

The page renders correctly, but the text a PDF reader extracts, searches or copies is wrong as soon as the
shaper forms a ligature, which Carlito (the bundled Calibri substitute, so every Calibri document) does for
ti, fi, ff and a few more. collect_glyph_usage (oxml-pdf/src/font.rs:46 to 53) pairs the i-th glyph
of a text run with the i-th character of its text. A ligature makes one glyph out of two characters, so every
later glyph of the run gets the character one place further; and because the ToUnicode map keeps one
character per glyph, first seen first kept (glyph_to_unicode: BTreeMap<u16, char>, font.rs:18, filled with
or_insert), the misassignment then spreads to every occurrence of that glyph in the document, in runs
without any ligature. On the umbrella's report fixture (#158), the table header reads "Locaton",
"Ratnn" and "Acton". rpptx is not affected in my tests: the same text in a deck extracts correctly.

import subprocess
import rdocx
from docx import Document
from docx.shared import Pt

TEXT = "Location Rating Action fifteen office Observation"
d = Document()
for font in ("Calibri", "Arial", "Cambria"):
    for bold in (False, True):
        r = d.add_paragraph().add_run(f"{font} {'bold' if bold else 'regular'}: {TEXT}")
        r.font.name, r.font.size, r.bold = font, Pt(11), bold
d.save("lig.docx")
open("lig.pdf", "wb").write(rdocx.Document.open("lig.docx").to_pdf())
print(subprocess.run(["pdftotext", "lig.pdf", "-"], capture_output=True, text=True).stdout.replace("\n\n", "\n").strip())
Calibri regular: Locatoo Ratog Actoo fieeo ofce Observatoo
Calibri bold: Locatoo Raton Actoo fiffo ofcf Obsfrvatoo
Arial regular: Location Rating Action fifteen office Observation
Arial bold: Location Rating Action fifteen office Observation
Cambria regular: Location Rating Action fifteen office Observation
Cambria bold: Location Rating Action fifteen office Observation

The multilingual path is not affected in practice: oxml-pdf/src/writer.rs wraps every MultilingualText
run in an ActualText span carrying its logical text (around line 1336), which a reader uses in
place of the ToUnicode map. Plain Text runs, the ones rdocx emits for this document, get no such span.
(Its ToUnicode entries, font.rs:63 to 74, keep only the first character of each cluster, so a ligature
there would still lose a character in a reader that ignores ActualText.)

Two earlier fixes touched the neighbourhood: #23 stopped glyphs being drawn twice at ligatures by slicing
shaped text by cluster (PR #24), and #74 made the text layer keep logical lines. The ToUnicode map of plain
runs still pairs glyphs and characters by index.

Acceptance: pdftotext of a rendered document returns its text for every bundled family, regular and bold,
ligatures included: the glyph-to-text mapping goes through the shaping clusters, and a ligature glyph maps to
all its characters (a multi-character ToUnicode entry, or ActualText). This matters for search, copy and
paste, accessibility, and the tagged PDF/A output that shares the writer.

Environment: main at 9a7ed714 (S75), release build, linux x86_64, Python 3.11, python-docx 1.2.0 for the
fixture, pdftotext from poppler.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions