Skip to content

Fix ToUnicode CMaps exceeding the 100-entry bfchar block limit - #1954

Merged
andersonhc merged 1 commit into
py-pdf:masterfrom
andersonhc:bfchar
Sep 16, 2026
Merged

andersonhc merged 1 commit into
py-pdf:masterfrom
andersonhc:bfchar

Conversation

@andersonhc

@andersonhc andersonhc commented Sep 16, 2026

Copy link
Copy Markdown
Collaborator

Fixes #1952

Checklist:

  • A unit test is covering the code added / modified by this PR

  • A mention of the change is present in CHANGELOG.md

  • This PR is ready to be merged

By submitting this pull request, I confirm that my contribution is made under the terms of the GNU LGPL 3.0 license.

@andersonhc
andersonhc merged commit 2082d9c into py-pdf:master Sep 16, 2026
23 checks passed
jbarlow83 added a commit to ocrmypdf/OCRmyPDF that referenced this pull request Sep 16, 2026
Upstream fpdf2 merged py-pdf/fpdf2#1954, which splits ToUnicode bfchar
blocks to 100 entries, but the Encoding CMap it writes for CFF CID-keyed
fonts (such as Noto Sans CJK OpenType) is still one cidchar block.
Ghostscript 9.56 through 10.04 drop it, and then both draw the wrong
glyphs and extract garbage. Verified against fpdf2 master and
Ghostscript 10.02.1.

Document that the patch stays in place for every fpdf2 release, and add
a subset of Noto Sans CJK JP as a test resource so the cidchar path is
covered, including a Ghostscript round trip.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Garbage text: /ToUnicode CMap exceeds the 100-entry bfchar` limit

1 participant