diff --git a/.github/release-notes/v0.4.0.md b/.github/release-notes/v0.4.0.md new file mode 100644 index 0000000..0eaeb1f --- /dev/null +++ b/.github/release-notes/v0.4.0.md @@ -0,0 +1,38 @@ +OffPDF v0.4.0 lets Edit PDF change text that is already in a PDF, in the document's own font, and only where the change can be verified. Everything still runs on your computer. + + + +Highlights: + +- Edit text (E) changes a line of existing text in the document's own font. The rest of the page stays where it was: every other glyph is checked to stay within 0.01 pt. +- Before a file is saved, each edited page is re-read and checked, then rendered and compared with the original by an independent engine (Poppler). If anything else changed, nothing is saved and the original file is never overwritten. +- Lines that can't be changed safely are shown dotted and say why when selected, for example text inside a reused block, a font made of drawings, or the hidden text layer of a scan. +- A longer line may run into the next text on the line, such as the next table cell or a footnote mark: OffPDF warns, saves it, and checks that the neighbouring words still read the same. +- Size, bold or italic (when the page already has that face), colour and letter spacing for the changed line, with a preview of the page as it will be saved after each change and Show original (O) to compare. +- A keyboard path for the whole flow: E, Tab to the text lines, arrow keys, Enter to edit, Tab to the format bar, Enter to apply. +- Edit PDF now opens on the Select tool; the old Text tool is called Add text and places a new text box. +- Stamps, shapes, images, drawings and markup now save on pages whose content is stored in several parts without line breaks, whenever those parts join cleanly. Pages whose parts would join inside a word, a string or a comment, or just after an inline image, are still refused, so a save never reveals text the original hid. + +What's not supported yet: + +- Replacing existing images. +- Text inside reused blocks (Form XObjects), shared page content, Type3 fonts, vertical, right-to-left and complex scripts. +- Reflowing paragraphs, multi-line edits or adding fonts: a subset font can only draw the letters the document already uses. +- Signed, password-protected and dynamic XFA PDFs. + +See [docs/EDIT_TEXT.md](https://github.com/McanKul/offpdf/blob/v0.4.0/docs/EDIT_TEXT.md) for the full list, the reason codes and the verification design. + +Known issues: + +- A file with any damaged JPEG image is refused for Edit text with "This PDF needs repair first", even when the damaged image is on another page. Run it through Repair PDF first. +- Reading a page and previewing a change can't be cancelled; Save can. +- Files over 256 MB are not kept in memory between steps: each page visit and preview reads them again, so Edit text is slower on them. + +Thanks to for contributing to this release. + +Packages: + +- Signed and notarized DMG for Apple Silicon Macs +- Windows x64 NSIS installer — unsigned while Authenticode signing is pending. Verify the accompanying SHA-256 file before use. + +See [CHANGELOG.md](https://github.com/McanKul/offpdf/blob/v0.4.0/CHANGELOG.md) for full details. diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index b4e8949..9653f90 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -63,9 +63,16 @@ jobs: librsvg2-dev \ libwebkit2gtk-4.1-dev \ patchelf \ + poppler-utils \ qpdf \ wget + - name: Show PDF engine versions + run: | + qpdf --version | head -n 1 + pdftotext -v 2>&1 | head -n 1 + pdftoppm -v 2>&1 | head -n 1 + - name: Use stable Rust run: rustup default stable @@ -79,4 +86,17 @@ jobs: run: cargo check --manifest-path src-tauri/Cargo.toml - name: Test Rust backend - run: cargo test --manifest-path src-tauri/Cargo.toml --lib + shell: bash + env: + # Edit text tests fail instead of skipping when qpdf or Poppler is missing; + # the grep below also catches the older tests' `skip:` lines. + OFFPDF_REQUIRE_ENGINES: "1" + run: | + cargo test --manifest-path src-tauri/Cargo.toml --lib -- --nocapture 2>&1 | tee rust-test.log + skips=$(grep -c 'skip:' rust-test.log || true) + echo "skip lines: $skips" + if [ "$skips" -ne 0 ]; then + grep 'skip:' rust-test.log + echo "::error::Rust tests skipped because a PDF engine was missing" + exit 1 + fi diff --git a/CHANGELOG.md b/CHANGELOG.md index 7bf5add..c1cd2ef 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,21 @@ All notable project changes should be documented here. ## Unreleased +### Added + +- Edit PDF can change existing text in place, in the document's own font. OffPDF only offers lines it can verify: every other glyph on the page is checked to stay within 0.01 pt, each edited page is re-read, re-rendered and compared before the new file replaces the destination, and lines that can't be changed safely, including text inside reusable blocks, say why. A longer line that runs into the next text on the line (a table cell, for example) saves with a warning. The original file is never overwritten. See [docs/EDIT_TEXT.md](docs/EDIT_TEXT.md). + +### Changed + +- The source-content classifier reads each PDF once with bounded decompression and per-page limits, decodes text, and reports text rise, rendering mode, mirrored text, clipping extent, shared content and font coverage with explicit reason codes. +- Edit PDF opens on the Select tool, and the old Text tool is now called Add text. + +### Fixed + +- Edit PDF page previews refresh when a source file changes on disk, even when its size and modification time stay the same, and older cached page files are removed. +- Edit PDF no longer rejects a save that adds stamps, shapes, images, drawings or markup to a page whose content is stored in several parts that don't end with a line break, as long as the parts join cleanly. This applies to every Edit PDF save, not only saves with text changes. A page whose parts would join inside a word, a string or a comment, or just after an inline image, is still refused, so a save can never make text that the original hid (for example, a line commented out across two parts) visible. +- Job commands accept only plain job ids, so a job's temporary work folder always stays inside the app's temp area. + ## 0.3.2 - 2026-09-11 ### Added diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 24e9a4c..4c3367b 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -46,6 +46,14 @@ cargo check --manifest-path src-tauri/Cargo.toml cargo test --manifest-path src-tauri/Cargo.toml ``` +Text-edit verification tests need qpdf ≥ 11 and Poppler (`pdftotext`, +`pdftoppm`). Set `OFFPDF_REQUIRE_ENGINES=1` to make missing engines fail instead +of skip. CI installs both, sets the variable, and fails when any Rust test prints +a `skip:` line. Edit text's design, reason codes and module map are in +[docs/EDIT_TEXT.md](./docs/EDIT_TEXT.md); adding a reason code, a font class or a +producer fixture, and regenerating benchmarks and goldens are in +[src/lib/editor/EDIT_MODEL.md](./src/lib/editor/EDIT_MODEL.md#edit-text-contributor-workflow). + `npm run check:versions` verifies that the release version matches in `package.json`, both root version fields in `package-lock.json`, `src-tauri/Cargo.toml`, the OffPDF entry in `src-tauri/Cargo.lock`, and diff --git a/README.md b/README.md index 63ea5cf..fc6552b 100644 --- a/README.md +++ b/README.md @@ -63,6 +63,11 @@ The Edit PDF workspace adds text, images, shapes, links, and freehand drawing as new content. It can also create PDF notes, highlights, underlines, strikeouts, and ink annotations while preserving existing annotations. +**Edit existing text.** Change a line of existing text in the document's own +font when OffPDF can verify the change: every other glyph is checked to stay +within 0.01 pt, lines that can't be changed safely say why, and the original +file is never overwritten. See [docs/EDIT_TEXT.md](./docs/EDIT_TEXT.md). + For interactive PDFs, OffPDF can fill existing AcroForm text fields (including multiline fields), checkboxes, radio buttons, combo boxes, and list boxes. Form fields stay interactive by default. @@ -75,10 +80,11 @@ The two flattening options have different scopes: page content. When a PDF contains form fields, OffPDF requires form fields to be flattened as part of this operation. -Current limits: XFA forms are not supported; only the first PDF's AcroForm can -be filled when several files are combined; and Edit PDF does not rewrite text -or images already embedded in a source page. New text and images are added as -overlays instead. +Current limits: XFA forms are not supported, and only the first PDF's AcroForm +can be filled when several files are combined. Edit text changes one line at a +time in the document's own font. It does not reflow paragraphs, add fonts, or +change text inside reused blocks, rotated or vertical text, scanned pages, or +signed and encrypted PDFs. Existing images are not replaced. ## Downloads @@ -170,8 +176,8 @@ npm run tauri:build | Engine | Used for | | --- | --- | -| `qpdf` | Merge, split, organize, encrypt, decrypt, repair, and lossless optimization | -| Poppler (`pdftoppm`, `pdftotext`) | Previews, image export, comparison, text export, and lossy compression | +| `qpdf` | Merge, split, organize, encrypt, decrypt, repair, lossless optimization, and writing Edit text changes | +| Poppler (`pdftoppm`, `pdftotext`) | Previews, image export, comparison, text export, lossy compression, and verifying Edit text changes | | Tesseract | OCR and searchable PDFs | | LibreOffice | Office conversion and PDF/A export | diff --git a/ROADMAP.md b/ROADMAP.md index 61b0da8..3e3b745 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -5,7 +5,7 @@ reliable local processing, clear packaging, and a contributor-friendly project. ## Current release -- **v0.3.1** is the latest published release. Its Windows package is unsigned +- **v0.3.2** is the latest published release. Its Windows package is unsigned and remains available through GitHub Releases. - The Apple Silicon macOS build is signed and notarized. It bundles qpdf, Poppler, and Tesseract; LibreOffice remains optional for Office and PDF/A work. @@ -22,6 +22,7 @@ reliable local processing, clear packaging, and a contributor-friendly project. ## Current focus - Stabilize the visual editor and make exported edits match the on-screen preview. +- Harden Edit text across more real-world PDFs. - Improve automated coverage for complex PDF geometry, rotation, and page boxes. - Make releases easier to trust and install, starting with Windows code signing. - Give new contributors smaller, clearly scoped issues with reproducible tests. @@ -52,8 +53,11 @@ reliable local processing, clear packaging, and a contributor-friendly project. - Flatten form fields converts form widgets to page content while preserving other annotations. Flatten annotations converts all remaining annotations to page content and requires form fields to be flattened when they are present. -- Edit PDF adds new text, images, and shapes as overlays; it does not rewrite - existing text or images inside the source document. +- Edit text changes one line at a time in the document's own font, only where + OffPDF can verify the change; other lines say why they can't be changed. It + does not reflow paragraphs, add fonts, change text inside reused blocks, + rotated or vertical text, or replace existing images. See + [docs/EDIT_TEXT.md](./docs/EDIT_TEXT.md). - Auto-update is intentionally not enabled. ## Later @@ -61,6 +65,9 @@ reliable local processing, clear packaging, and a contributor-friendly project. - Add optional, clearly disclosed update checks behind an offline-first setting. - Broaden platform packaging and package-manager distribution. - Add more advanced editing workflows where they can remain reliable and local. +- Edit text: shared content streams, text inside reusable blocks (Form + XObjects), and Type3 fonts. +- Replace existing images, after a dedicated spike. Priorities can change as real bug reports and contributor feedback arrive. See the [open issues](https://github.com/McanKul/offpdf/issues) for work that is diff --git a/docs/EDIT_TEXT.md b/docs/EDIT_TEXT.md new file mode 100644 index 0000000..f29b5a8 --- /dev/null +++ b/docs/EDIT_TEXT.md @@ -0,0 +1,800 @@ +# Edit text + +Edit text changes words that are already in a PDF, in the document's own font, +without moving anything else on the page. It is part of **Edit PDF** (tool +**Edit text**, shortcut **E**). + +This page is the user guide (what Edit text can and cannot change, and why: +sections 2–7), the decision record for issue #11 (section 1, evidence in 11–13) +and the contributor reference for the engine (sections 8–10 and 16). The editor's +data model (`kind: "sourceText"` objects, keyboard contract) is in +[`src/lib/editor/EDIT_MODEL.md`](../src/lib/editor/EDIT_MODEL.md). + +## 1. Decision (#11) + +| | | +| --- | --- | +| Question | Can OffPDF reliably change existing page text or embedded images without pretending that PDF is a word-processing format? | +| Decision | **Yes, for a bounded, fail-closed subset of text.** Existing images are not replaced in v0.4. | +| Status | Accepted with the v0.4 implementation. The maintainer confirms it on #11 before the release PR is merged. | +| Date | 2026-10-02 | +| Owner | @McanKul (maintainer) | +| Scope | One line at a time, inside the page's own content stream, in a font the page already uses, verified before it is saved. | +| Not in scope | Image replacement (the image half of #11), text inside reusable blocks (Form XObjects), annotation text, shared content streams, Type3, vertical, right-to-left and complex scripts, reflow, new fonts. See [section 15](#15-known-limitations-and-v05-candidates). | +| Engine | lopdf 0.34 (read only) + OffPDF's own bounded lexer and walker, qpdf as the only writer, Poppler as an independent verifier. No PDFium sidecar ([section 13](#13-engine-decision)). | +| Evidence | The reason codes ([6](#6-reason-codes)), the compatibility matrix ([11](#11-compatibility-matrix)), the measured performance and engine matrix ([12](#12-performance-and-engine-matrix)) and the adversarial gate tests ([9](#9-verification)). | + +The exit criterion of #11 is met by construction: every line Edit text cannot +change safely is refused with a named reason, and no refusal falls back to +painting new text on top of old text. + +## 2. What Edit text does and never does + +**It does:** + +- change the words of one line, shorter, longer or empty (an empty line is + removed, and the text after it stays where it was); +- keep the unchanged letters of the line exactly as they were (same glyph codes + and the same kerning between them); +- keep everything else on the page where it was: every other glyph is checked to + stay within 0.01 pt; +- offer size, bold or italic (when the page already has that face), fill colour + and letter spacing for the whole line; +- show the real result: after each change it renders the page that Save would + write, through the same checks as Save; +- explain every line it refuses: select a dotted line to see the reason; +- write a new file. The original is never overwritten. + +**It never:** + +- covers old text with a white box and new text, adds an overlay, or adds an + annotation in place of the old words; +- flattens or rasterises a page, or re-saves a page through a PDF library; +- substitutes or embeds a font, or draws a letter the embedded font does not + contain; +- reflows a paragraph, wraps a line, or moves other lines; +- silently skips a change: a change that cannot be proven is refused, and Save + publishes nothing. + +What you see while editing: dashed outline = editable, dotted = can't be changed +(text inside reusable blocks included), green outline with a dot = edited. The +line editor shows your draft in a similar serif, sans or monospace system font +(from the font's flags, else its name); after **Done** the preview shows the +document's own font. **Show original** (**O**) compares the original page with +that preview while one is on screen. **Save** with a line still open finishes it +first; a draft that can't be applied stays open and nothing is saved. **Add +text** places a new text box and never replaces existing text; **Add text here** +in the reason dialog puts one over the line (an upright one-line box at the +line's start when the line is not level on screen). + +Before you rely on it, read the limits in [section 15](#15-known-limitations-and-v05-candidates). +The two you are most likely to meet: a subset font only contains the letters the +document already uses, and a file with any damaged JPEG image is refused with +`PDF_NEEDS_REPAIR`. + +## 3. Supported subset + +### Unit of editing + +| Rule | v0.4 | +| --- | --- | +| What changes | The text of one line: one show operator (`Tj`, `TJ`, `'`, `"`), or several consecutive ones joined into one visual line. Only text the page's own `/Contents` draws directly; text drawn through a Form XObject is listed as refused lines (`NESTED_FORM`) after the page's own lines. | +| Joining | Same baseline (0.01 pt) and direction, same graphics and text state except the font, the same marked-content context, the same font or a sibling face of it, a gap within ±0.3 em, nothing painted in between, at most 256 pieces. Gaps between joined pieces become `TJ` kerns, so unchanged glyphs keep their positions. | +| Per-glyph pages | After joining, if a page has at least 12 editable lines and 80 % or more of them are one character long, those one-character lines are refused (`PER_GLYPH_TEXT`). | +| Lines | Single line. No reflow, wrapping or multi-line editing. Lines that are only spaces are not offered. | +| Position | The new text starts where the old text started. The pen after the line ends where it ended before (±0.01 pt). | +| No-op | Text equal to the current text with no style change writes nothing. | + +### Fonts + +| Font setup | v0.4 | How a letter is proven drawable | +| --- | --- | --- | +| Simple TrueType, embedded (full or subset) | Editable | The glyph is found by every applicable lookup (glyph name, Unicode cmap, MacRoman cmap, symbolic cmap) and all lookups agree; the outline has at least one segment | +| Simple CFF (`/Type1C`) and OpenType (`/OpenType`) | Editable | Glyph by name from the PDF encoding, or by the program's own built-in encoding (read by OffPDF's bounded parser); outline present | +| Simple Type1 (`/FontFile`) | Editable | eexec decrypted; the charstring draws at least one path; the name is also in `/CharSet` when that exists | +| Standard 14 (Helvetica, Times, Courier and common aliases such as Arial, Times New Roman, Courier New) | Editable | Glyph name in that face's AFM; widths from `/Widths` or that face's AFM | +| Simple, not embedded, with `/Widths` | Editable, marked "not included in the PDF" | Width > 0 and a Latin, Greek or Cyrillic character; your PDF reader substitutes the font exactly as it already does for the existing text | +| Type0 `/Identity-H` with CIDFontType2 (TrueType) | Editable | CID → GID through `/CIDToGIDMap`; outline present; ToUnicode required | +| Type0 `/Identity-H` with CIDFontType0 (CFF) | Editable | CID → GID through the CFF charset; outline present | +| Font set by an ExtGState `/Font` | Text, colour and letter spacing; size and bold/italic unavailable | as its class | +| Type3, vertical, predefined non-Identity CMaps, Symbol, ZapfDingbats, MMType1, CFF2-only programs, sfnt programs with both `glyf` and CFF outlines | Refused | see [section 6](#6-reason-codes) | + +A subset font only contains the letters the document already uses. Edit text +shows the available letters (**Letters** in the format bar) and names any +letter it cannot draw before you commit. + +### Characters + +- Reading: Latin (Basic to Extended-B and Extended Additional), IPA, Greek, + Cyrillic, general punctuation, symbols, currency, arrows, maths operators, + box drawing, CJK punctuation, Hiragana, Katakana, CJK ideographs, Hangul + syllables and full-width forms. Ligatures such as "fi" read as "fi" but are + never typed. +- Refused: right-to-left scripts (`RIGHT_TO_LEFT`) and scripts that need shaping + (`COMPLEX_SCRIPT`). +- Typing: what you type is normalised (NFC). Line breaks, tabs and control + characters are rejected (`INVALID_TEXT`). At most 1,000 characters per line + (`TEXT_TOO_LONG`). A character the font cannot draw is named + (`GLYPH_MISSING`). Only characters in the Basic Multilingual Plane. +- Spaces: when the font has a space glyph, a typed space is a real space. + Otherwise (typical for pdfTeX) word gaps are kerns; a typed space then becomes + a kern equal to the line's usual word gap, and it must sit between two + letters, one at a time (`SPACE_NOT_WRITABLE`). + +### Content and geometry + +- Content filters: none, Flate (no predictor), ASCIIHex and ASCII85, in chains + of up to 4. Anything else refuses the page (`UNSUPPORTED_FILTER`). +- `/Contents` may be one stream or an array of parts. A show operator that crosses + two parts is refused (`SPLIT_CONTENT`). Viewers read a part boundary differently + (nothing, a space or a newline), so parts that meet inside a token, comment or + `BX`…`EX` section, run two tokens together, or follow an inline image ending + within 64 bytes of the part's end refuse the page (`MALFORMED_CONTENT`); parts + meeting between operators (`ET`|`BT`), after a number (`0`|`cm`) or at a comment's + end of line are fine. +- Inline images are measured exactly when their length can be proven; text drawn + after an inline image whose end is only guessed is refused (`INLINE_IMAGE`). +- `/Rotate` 0, 90, 180 and 270 and offset CropBoxes are supported. Orientation + is judged as displayed: text that reads level on screen is editable, including + landscape pages made by rotation and synthetic italic (shear up to 0.5). +- `/UserUnit` other than 1, unreadable page boxes and non-integer rotations + refuse the page (`GEOMETRY`). +- Text render modes 0, 1 and 2 are editable (colour only for mode 0). +- Visible optional-content layers are editable; hidden or conditional ones are + refused (`OPTIONAL_CONTENT`). +- Encrypted, signed (an applied signature) and dynamic XFA files are refused when + they are opened. An unedited signed file may still be appended in the same Save. + +## 4. Style controls and units + +| Control | Range and units | Written as | Unavailable when | +| --- | --- | --- | --- | +| Size | Effective points as you see them, 4–144, steps of 0.5 | `Tf` scaled so the size on the page equals the target; the original `Tf` is restored byte for byte after the line | the font comes from an ExtGState (`STYLE_UNAVAILABLE`) | +| Bold / Italic | Toggles | Switch to the sibling face already in the page's font resources; the whole line is re-encoded in that face | no such face on the page, or it cannot draw the text (`FACE_UNAVAILABLE`); ExtGState font (`STYLE_UNAVAILABLE`) | +| Colour (fill) | Original colour, six inks (black, grey, red, blue, green, amber), or a custom `#rrggbb` | `r g b rg`; the colour in force after the line is restored byte for byte | text drawn with an outline, render mode ≠ 0 (`STYLE_UNAVAILABLE`) | +| Letter spacing | Effective points, −2 to +10, steps of 0.1 | `Tc`; the original `Tc` is restored byte for byte | — | + +Word spacing, horizontal scaling, rise, stroke colour, render mode, underline, +alignment and changing the font family are not offered. A style value within +0.001 of the current one is treated as unchanged. Style applies to the whole line. + +## 5. Width policy + +- Shorter or equal text is always allowed. +- Longer text is allowed while it stays inside the visible page and its clip + area, or within the line's original extent if that was larger. Past that + edge the change is blocked (`TEXT_OUTSIDE_VISIBLE_AREA`). +- Running into the next text on the same line (a longer table cell, say) is a + warning (`NEXT_TEXT_OVERLAP`), shown in the editor, in the inspector and in + the Save job message; the change still saves. The preview shows the real + overlap. +- Text after the changed part of the line keeps its position. When the line + contains a column gap of at least 1 em, the gap absorbs the width change and + the columns after it do not move. Otherwise one kern at the end of the line + restores the pen. + +## 6. Reason codes + +Every line that cannot be changed carries exactly one code. When several apply, +the first in this table wins (all of them are kept in the technical details). +**Run** codes refuse one line. **Page** codes refuse every line on the page, and +the page shows one banner instead of outlines. + + +| Code | Level | When | Short | Title | What OffPDF says | +| --- | --- | --- | --- | --- | --- | +| `NESTED_FORM` | run | Drawn from inside a Form XObject (a reusable block), at any depth; listed after the page's own lines | part of a reused block | Part of a reused block | This text is inside a block the document can reuse, so changing it here could change it in other places too. | +| `SPLIT_CONTENT` | run | The show operator's bytes cross the boundary between two `/Contents` parts | stored in two pieces | Stored in two pieces | The instructions that draw this line are split across two parts of the page. | +| `SHARED_CONTENT` | run | The content part is reachable from more than one page or object (page listed twice in the page tree, part stream or `/Contents` array referenced twice) | shared with other pages | Shared with other pages | This part of the page is shared with other pages, so a change here would change them too. | +| `INLINE_IMAGE` | run | Drawn after an inline image whose end could not be proven | after an unreadable picture | After an unreadable picture | A picture stored inside this page can't be measured exactly, so text drawn after it can't be changed safely. | +| `INVISIBLE_TEXT` | run | Text render mode 3 (invisible, e.g. the searchable layer of a scan) | hidden text | Hidden text | This is hidden text, such as the searchable layer of a scan. Changing it would change nothing you can see. | +| `TEXT_CLIP_MODE` | run | Text render modes 4–7 (text used as a clipping path) | used as a shape | Used as a shape | This text is used as a shape that clips other content, so it can't be changed safely. | +| `ZERO_SIZE` | run | `Tf 0`, `Tz 0` or a singular text matrix | no visible size | No visible size | This text is drawn at zero size. | +| `VERTICAL` | run | Vertical writing (`Identity-V`, a `-V` CMap, `/WMode 1`) | vertical text | Vertical text | This text runs top to bottom. Only lines that read across the page can be changed. | +| `MIRRORED_TEXT` | run | Exactly one axis flipped as displayed (negative `Tz`, mirrored `Tm`) | mirrored | Mirrored text | This text is drawn mirrored, so new letters can't be placed the same way. | +| `ROTATED_TEXT` | run | Not level as displayed: upside down or turned at an angle | turned at an angle | Turned at an angle | This line doesn't run straight across the page as shown. Only level lines can be changed. | +| `SKEWED_TEXT` | run | Neither level, mirrored nor a pure rotation as displayed: a tilted baseline or a slant beyond 0.5 | tilted line | Tilted line | This line's baseline is tilted by the page layout, so it can't be changed safely. | +| `CLIPPED` | run | Ink box outside the visible page box (CropBox ∩ MediaBox), or cut by a clip that is not a set of rectangles containing it | partly cut off | Partly cut off | Part of this text is cut off by the page edge or a clipping area, so a change might not show. | +| `OPTIONAL_CONTENT` | run | Inside optional content that is hidden by default, a membership dictionary or an unresolvable group | on a layer | On a layer | This text is on a layer that is hidden or depends on viewer settings, so a change might not show. | +| `SOFT_MASK` | run | An ExtGState soft mask is in force | drawn through a mask | Masked text | This text is drawn through a transparency mask, so a change could look different from what you type. | +| `PATTERN` | run | Fill (or stroke, for modes 1–2) uses a Pattern colour space | pattern fill | Pattern fill | This text is filled with a pattern or gradient rather than a colour. | +| `ACTUAL_TEXT` | run | Marked content or its structure element carries `/ActualText` | separate reading copy | Separate reading copy | The document keeps a separate copy of this text for search and screen readers. Changing only the visible letters would make the two disagree. | +| `MISSING_FONT` | run | The font name is not in the page's resources | font missing | Font missing | The font this text uses is missing from the document. | +| `TYPE3` | run | Type3 font | font made of drawings | Font made of drawings | This text uses a font drawn from shapes, which OffPDF can't type with. | +| `FONT_UNSUPPORTED` | run | Malformed `/Widths`, `/FirstChar`, `/LastChar`, `/W` or `/Differences`; MMType1 | unusual font setup | Unusual font setup | This font is set up in a way OffPDF can't change safely. | +| `FONT_NOT_EMBEDDED` | run | Type0 (CID) font without an embedded program | font not included | Font not included | This font isn't included in the PDF, so OffPDF can't check which letters it can draw. | +| `FONT_PROGRAM_UNSUPPORTED` | run | Embedded program of an unknown or mismatched `FontFile3` subtype, a CFF2-only program, or an sfnt holding both `glyf` and `CFF `/`CFF2` outlines | font type not supported yet | Font type not supported yet | This text uses a kind of embedded font that OffPDF can't check yet. | +| `FONT_PROGRAM_UNREADABLE` | run | The embedded program does not parse | font can't be read | Font can't be read | The font included in this PDF can't be read, so OffPDF can't check which letters it can draw. | +| `UNSUPPORTED_ENCODING` | run | Non-Identity CMaps, `/MacExpertEncoding`, Symbol and ZapfDingbats, symbolic fonts that are not embedded | unsupported letter mapping | Letter mapping not supported | This font maps letters in a way OffPDF can't write yet. | +| `MISSING_WIDTHS` | run | Non-Standard-14 simple font without `/Widths`, or text drawn after an advance OffPDF cannot know | letter widths missing | Letter widths missing | The document doesn't say how wide this font's letters are, so the rest of the line couldn't be kept in place. | +| `NO_TOUNICODE` | run | Type0 or symbolic TrueType font without a usable ToUnicode map | letters not identified | Letters not identified | The document doesn't say which letters this font draws, so OffPDF can't read or retype them. | +| `AMBIGUOUS_UNICODE` | run | A glyph's text is missing, U+FFFD or private use, or ToUnicode and the glyph name disagree | letters can't be confirmed | Letters can't be confirmed | Some letters in this line can't be read with certainty, so OffPDF can't show the current text reliably. | +| `RIGHT_TO_LEFT` | run | Hebrew, Arabic, Syriac, Thaana or NKo text | right-to-left script | Right-to-left script | The document has already ordered and shaped these letters, and changing them would undo that work. | +| `COMPLEX_SCRIPT` | run | Indic scripts, Thai, Lao, Tibetan, Myanmar, Khmer, Mongolian, Hangul Jamo or any script outside the reading list | shaped script | Shaped script | This script joins or reorders letters, and the document has already done that shaping. OffPDF can't redo it yet. | +| `DUPLICATE_TEXT` | run | Another line with the same text overlaps at least half of this one's ink box | drawn twice | Drawn twice | This text is drawn twice (for example as a shadow or to look bolder), so changing one copy would leave the other behind. | +| `PER_GLYPH_TEXT` | run | At least 12 editable lines on the page and 80 % or more are one character long | letters placed one by one | Letters placed one by one | This page places every letter on its own, so a line can't be changed as one piece. | +| `NO_WRITABLE_GLYPHS` | run | The font has no character OffPDF can prove it can draw | no letters to type with | No letters to type with | The font this line uses has no letters OffPDF can confirm, so nothing can be typed with it. | +| `MALFORMED_CONTENT` | page | The page content cannot be read completely (bad syntax, an unresolvable reference, `q` nesting over 64), or two `/Contents` parts meet inside a token or comment | page can't be read | Damaged page content | OffPDF couldn't read this page's content completely, so none of its text can be changed. | +| `UNSUPPORTED_FILTER` | page | Content compressed with anything other than Flate, ASCIIHex or ASCII85, or a content stream with extra dictionary keys | unusual compression | Unusual compression | This page's content is stored or compressed in a way OffPDF can't check. | +| `PAGE_TOO_COMPLEX` | page | A per-page budget ran out (operators, glyphs, decoded bytes, page-model memory including fonts) | too much content | Too complex to check | This page has too much content to check safely. | +| `GEOMETRY` | page | `/UserUnit` other than 1, unreadable or degenerate page boxes, or a `/Rotate` that is not a multiple of 90 | custom page unit | Custom page unit | This page uses a custom unit size or unreadable page boxes, which OffPDF doesn't edit yet. | + + +Image occurrences are classified too (the #33 classifier), but v0.4 does not +edit images and does not show these codes. In priority order: `INLINE_IMAGE`, +`NESTED_FORM`, `CLIPPED`, `PATTERN`, `MASKED_IMAGE` (drawn with a mask or soft +mask), `SHARED_XOBJECT` (one image used in more than one place), +`TRANSFORMED_IMAGE` (rotated, skewed or mirrored image), `GEOMETRY`. + +The codes are append-only after v0.4. The single source of truth is +[`src/lib/editor/text-reasons.json`](../src/lib/editor/text-reasons.json); see +[`EDIT_MODEL.md`](../src/lib/editor/EDIT_MODEL.md#adding-a-reason-code) for how a code is added. + +## 7. Messages and Save-time errors + +### While you edit + +When a change cannot be committed, the editor keeps your draft and shows one of +these messages under the line. `GLYPH_MISSING` lists the characters (a space +shows as "space"). `FACE_UNAVAILABLE`, `STYLE_UNAVAILABLE` and +`TEXT_EDIT_REFUSED` show the sentence for the face, the control or the reason +involved; the table shows their default. Warnings (the last three rows) do not +block the change. + + +| Code | Message | +| --- | --- | +| `GLYPH_MISSING` | This document's font can't draw: {chars} | +| `SPACE_NOT_WRITABLE` | A space can only go between two letters here, one at a time. | +| `INVALID_TEXT` | Line breaks and tabs can't be added. Each box is a single line. | +| `TEXT_TOO_LONG` | A line can have at most 1,000 characters. | +| `TEXT_OUTSIDE_VISIBLE_AREA` | The new text would run past the edge of the visible page. | +| `FACE_UNAVAILABLE` | This page has no regular version of this font. | +| `STYLE_UNAVAILABLE` | This line's size is set in a way OffPDF can't change. | +| `EDIT_CONFLICT` | This line is already part of another change on this page. | +| `TEXT_EDIT_REFUSED` | OffPDF can't change this text safely. | +| `STALE` | “{name}” changed on disk after you started editing it. The text changes made before that can't be applied. | +| `PEN_DRIFT` | That change would move other text on this page, so it can't be saved. Undo it or try a shorter change. | +| `STATE_CHANGED` | That change would alter how other content on this page is drawn, so it can't be saved. | +| `EDIT_VERIFY_FAILED` | OffPDF couldn't confirm this change reads back exactly as typed with nothing else altered, so it can't be saved. | +| `NEXT_TEXT_OVERLAP` | The new text runs into “{neighbour}”. | +| `EDIT_NOT_VISIBLE` | This change doesn't change how the page looks. Something may be drawn over this line. | +| `PREVIEW_UNAVAILABLE` | Preview unavailable for this page. Your changes are checked again when you save. | + + +### Errors (open and Save) + +Errors show a title, a message and a suggestion. At Save, every suggestion ends +with the sentence that the original file was not changed (it is added in one +place, `reasons::save_failure`, and asserted by a Rust test). A file that cannot +be opened for Edit text shows its error in a banner instead; the rest of Edit +PDF keeps working. An empty suggestion is shown as "—" here. + +When a file is opened: + + +| Code | Title | Message | Suggestion | +| --- | --- | --- | --- | +| `ENCRYPTED` | This PDF is password-protected | OffPDF can't change text in a protected PDF. | Remove the password with Unlock PDF, then edit the unlocked copy. | +| `SIGNED` | This PDF is digitally signed | Changing its text would break the signature. | Ask the sender for an unsigned copy if it needs changes. | +| `UNSUPPORTED_XFA` | This PDF is a dynamic form | Its pages are generated by the PDF reader, so changes to page text might not show. | Fill it in a reader that supports dynamic forms. | +| `PDF_NEEDS_REPAIR` | This PDF needs repair first | Its internal structure has errors, so OffPDF won't change its text. | Run it through Repair PDF, then edit the repaired copy. | +| `FILE_TOO_LARGE` | This PDF is too large to edit text in | Files over 400 MB are not read for text editing. | Split it with Split PDF and edit the part you need. | +| `FILE_TOO_COMPLEX` | This PDF is too complex to check | Some of its internal data is too large or too deeply nested to check safely. | — | +| `MALFORMED_CONTENT` | Part of this PDF can't be read | OffPDF couldn't read this PDF's structure completely. | Run it through Repair PDF, then try again. | +| `VERIFIER_MISSING` | A checking component is missing | OffPDF needs its bundled Poppler tools to verify text changes. | Reinstall OffPDF, then try again. | +| `INVALID_PDF` | The selected file is not a valid PDF | OffPDF could not open this file as a PDF document. | Make sure the file is a real PDF and is not corrupted. | +| `ENGINE_MISSING` | PDF engine not found | The bundled qpdf engine could not be located. | Reinstall OffPDF, then try again. | + + +`PDF_NEEDS_REPAIR` can also arrive later: the full `qpdf --check` runs in the +background after the file opens, and its result is needed before the first +preview and before Save. When the bundled qpdf is older than version 11, +`ENGINE_MISSING` says that the engine is too old instead. If qpdf cannot open +its copy during the check (for example after the app's temporary files were +cleared), that is an engine failure (`ENGINE_FAILED`), never a repair verdict, +and the check runs again at the next preview or Save. + +When a change cannot be saved (the same checks run at preview and at Save; +`{n}` is the page number in the saved file): + + +| Code | Title | Message | Suggestion | +| --- | --- | --- | --- | +| `GLYPH_MISSING` | The font can't draw some letters | On page {n}, the document's font can't draw: {chars} | Use other characters, or add a new text box with Add text. | +| `SPACE_NOT_WRITABLE` | A space can't go there | On page {n}, spaces in the changed line are gaps between letters, so a space can only go between two letters, one at a time. | Remove the extra space. | +| `INVALID_TEXT` | These characters can't be added | Line breaks, tabs and control characters can't be added. Each box is a single line. | Remove them and try again. | +| `TEXT_TOO_LONG` | Text is too long | A changed line can have at most 1,000 characters. | Shorten the line. | +| `TEXT_OUTSIDE_VISIBLE_AREA` | Text runs off the page | On page {n}, a changed line would run past the edge of the visible page. | Shorten the line. | +| `FACE_UNAVAILABLE` | That style isn't available | On page {n}: {sentence} | Keep the current style. | +| `STYLE_UNAVAILABLE` | That change isn't available | On page {n}: {sentence} | Keep the current setting. | +| `EDIT_CONFLICT` | Two changes overlap | Two text changes on page {n} affect the same line. | Undo one of them and try again. | +| `TEXT_EDIT_REFUSED` | This text can't be changed safely | On page {n}: {body} | Restore the original text for that line. | +| `STALE` | The PDF changed on disk | “{name}” was changed after you started editing it, so your text changes no longer match it. | Remove these text changes and make them again. The original file was not changed. | +| `PEN_DRIFT` | The change would move other text | Saving would shift text you didn't change on page {n}, so nothing was saved. | Try a shorter change, or undo it. The original file was not changed. | +| `STATE_CHANGED` | The change would restyle other content | Saving would change how other content on page {n} is drawn, so nothing was saved. | Undo the last change on that page and try again. The original file was not changed. | +| `EDIT_VERIFY_FAILED` | The change could not be verified | OffPDF checks every change before saving. On page {n}, a changed line couldn't be confirmed to read back exactly as typed with nothing else altered, so nothing was saved. | Undo that change and try again. The original file was not changed. | + + +Only at Save: + + +| Code | Title | Message | Suggestion | +| --- | --- | --- | --- | +| `SOURCE_EDIT_GATE_FAILED` | The edited PDF did not pass the text-change check | The saved file did not contain the text changes exactly as they were checked, so it was not published. | Try saving again. The original file was not changed. | +| `TEXT_EDIT_ON_REDACTED_PAGE` | Redaction and text change on the same page | Page {n} has both a redaction and a text change. Redaction turns the page into an image, so the text change would be lost. | Remove the redaction or the text change on that page. | +| `TEXT_EDIT_DUPLICATE_PAGE` | This page appears twice | Page {p} of “{name}” is in the list more than once and has a text change. | Remove the extra copy of the page, then save again. | +| `TOO_MANY_TEXT_EDITS` | Too many text changes | This save has more than 500 changed lines. | Save in smaller batches. | +| `INVALID_PAGES` | Invalid page selection | This page is not in the PDF. | Open the page again from the page list. | +| `BAD_EDIT` | Could not save | A text change has a setting OffPDF can't use. | Undo the last change and try again. | + + +`TOO_MANY_TEXT_EDITS` is also returned for more than 200 changed lines on one +page. Every job command (Save and the other Edit PDF and tool jobs) accepts only +job ids of 1–128 ASCII letters, digits, `-` and `_`; anything else is +`INVALID_JOB` ("Invalid job" / "OffPDF received a job id it does not accept.", +no suggestion), so a job's work folder stays inside the app's temp area. Files +larger than 100 MB open with a note that checking and saving takes longer. The +technical details of every error (check id, page, byte offset, drift in points) +are collapsed under the message. + +## 8. How a change is written + +A change is a byte splice inside the page's own content stream. Nothing else +in the file is rewritten by OffPDF; qpdf writes the result. + +```text +user's PDF ──one capped read──► snapshot bytes (hashed; qpdf only ever reads copies of these) + │ lopdf parses the same bytes, read-only + ▼ + /Contents parts ──bounded decode──► joined buffer ──byte-offset lexer──► show ops ──► lines (runs) + │ + plan_page: minimal diff + pen compensation ◄───────────────┘ + │ + ▼ + splices: byte ranges of the line's show ops → replacement bytes + │ applied back to front, per part + ▼ + expected parts (the exact decoded bytes every edited part must have) + │ self-check: the same walker re-reads the expected parts + ▼ before any file is written + update.json {"obj:N G R": {"stream": {"dict": {}, "data": ""}}} + │ + ▼ + qpdf source.pdf edited.pdf --decode-level=none --compress-streams=n --update-from-json=update.json + │ + ▼ + Phase A gate on edited.pdf (section 9) +``` + +**Minimal diff.** The unchanged prefix and suffix of the line keep their glyph +codes, fonts and the kerns between them, byte for byte. Only the middle is +re-encoded, with the font's own codes. A kern at a boundary that was tuned for a +glyph pair that no longer exists is dropped; a kern that reads as a word space +is kept. Worked example (a Word-style line, test PLAN-04): + +```text +before: BT /F1 11.04 Tf 1 0 0 1 72 700 Tm [(Inv)12(oice 2026)]TJ ET +after: BT /F1 11.04 Tf 1 0 0 1 72 700 Tm [<496E76> 12 <6F6963652032303237>] TJ ET +``` + +"Invoice 2026" became "Invoice 2027"; the kern 12 between "v" and "o" is kept; +digits have equal widths, so no compensation number is written. + +**Compensation.** With `A` the line's original advance and `A′` the new one (in +text space, including `Tc`, `Tw` and kerns), the pen is restored by one number at +the end of the last `TJ`: + +```text +n_c = (A′ − A) × 1000 / Tf′ (omitted when |n_c| < 0.0005) +``` + +Numbers are written with at most 4 decimals and parsed back; all later +arithmetic uses the parsed value. If the kept suffix holds a column gap of at +least 1 em, that gap absorbs the change instead, and the text after it keeps its +original position. + +**Style.** A size change writes a scaled `Tf`; a colour change writes `rg`; a +letter-spacing change writes `Tc`. After the line, the original operators are +restored verbatim (the original `Tf`, colour-space and colour operators, `Tc` +source bytes), so later content is drawn exactly as before. + +**Absorbed pieces.** When a line was joined from several show operators, the +first one draws the whole new line and each other piece becomes `[<> n] TJ`, +which draws nothing and keeps that piece's original pen travel +(`n = −advance × 1000 / Tf`). + +**Replacement grammar.** A replacement may contain only `Tf Tc Tw T* TJ`, the +verbatim restore operators (`g rg k cs sc scn`) and the new `rg`. It never +contains `q Q BT ET cm Tm Td TD Tz Ts Tr gs`, paths, images or marked content. +This is checked before any file is written and again on the re-read. + +**Why qpdf writes.** lopdf 0.34's incremental writer would add a second header +line and re-print numbers through `f32`. qpdf's `--update-from-json` replaces +only the data of the edited streams, keeps every untouched filtered stream raw +(`--decode-level=none`) and drops the superseded data, so the old text cannot be +recovered from the new file (test E2E-18). + +## 9. Verification + +Every change passes the same checks at preview and at Save. A failure publishes +nothing. + +**Plan self-check (before any IO).** The planner re-walks the expected bytes +with the same walker and runs the re-walk checks below. A failure is an internal +`EDIT_VERIFY_FAILED`; qpdf is never started. + +**Phase A** runs on each edited copy (and on the preview file). The copy is +opened with the same resource bounds as a source, but without the policy +refusals (`read_verification_snapshot`). + +| Check | Proves | Catches | +| --- | --- | --- | +| A0 qpdf | `qpdf --check` exits 0, or 3 with only warnings the source already had | a mis-targeted update (qpdf writes it with exit 0 and only `--check` notices) | +| A1 engines agree | lopdf and qpdf list the same pages and `/Contents` ids; same page count | lopdf misreading qpdf's output; dropped or duplicated pages | +| A2 whole graph | Every object reachable from `/Root` and `/Info` equals qpdf's input, except the edited content streams, which decode to exactly the expected parts. Object ids are ignored; filtered streams are compared raw. | collateral damage anywhere: other pages, fonts and `/Widths`, resources, `/Annots`, boxes and `/Rotate`, page order, catalog (layers, forms, names, open action, outlines), added content streams or Form XObjects, raster replacements, shared-stream changes | +| A3 edited parts | Same part count; each part decodes to its expected bytes; the original bytes are not attached anywhere on the page | off-by-one splices, the wrong page, re-encoded content, the old stream re-attached | +| A4 re-walk | Every unedited glyph keeps its codes, font, text and state, and stays within 0.01 pt; every paint keeps its state; the edited line has exactly the planned glyphs, origin, size, spacing, colour and pen; the state after the line equals the original | wrong compensation, state leaks (`Tc`, `Tz`, colour, line width, …), missing glyphs, wrong size or font | +| A5 independent | Poppler (`pdftotext -bbox`) finds the same words outside the edited band; inside it, the words of the line's other runs read the same at the same place (or joined to the new text at their outer edge) and are set aside, then the new text is found and the old words fewer times. A word holding glyphs of a run no edit changes is such a neighbour wherever it lies (a footnote marker or a "." right after the line), and a word Poppler joins across the edit ("Hello:") is split, its other run's part pinned at the outer edge (or, when a new letter covers it, found inside the new word). `pdftoppm` renders of the page before and after differ in at most 8 pixels outside the edited glyph boxes and in at most 2 pixels of other runs' glyphs outside the edited glyphs' own boxes (grown by 1 pixel), and no neighbour under the new glyphs loses more than 2 pixels of ink | width-model errors that OffPDF's own walker would share, old text still extractable, rendering-level leaks | + +**Phase B** runs on the final file after every other Edit PDF pass (assembly, +stamps, links, forms, markup, redaction on other pages), just before the existing +output validation (#34). Any failure is `SOURCE_EDIT_GATE_FAILED`. + +| Check | Proves | +| --- | --- | +| B0 read | The final file opens within bounds and lopdf and qpdf agree on its pages | +| B1 present | Each edited page holds its expected content exactly once, either as a contiguous run of content parts, or inside the one Form XObject qpdf's `--overlay` wraps a page in (identity transform, `/BBox` containing the visible page, and qpdf's join of the expected parts token-neutral as below: PB-09) | +| B2 absent | The original bytes of the edited parts are nowhere on the page, its wrapper or its Form XObjects | +| B3 re-walk | The page's text records equal the proven records one to one (codes, fonts, text, origins within 0.01 pt, full state); later passes may only add paints | + +**Composition with #34.** After Phase B, the #34 snapshot digest of each edited page +must equal the digest computed by the planner, independently of the files on disk. +#34's own check then runs unchanged. For a page whose parts do not end with a line +break it also accepts qpdf's joined form (`alt_content_digest`), only in saves qpdf +overlays (any text box, image, shape, drawing or markup) and only if the join is +token-neutral: every newline qpdf adds falls between tokens, outside comments, +strings, inline images (and the 64 bytes after one) and `BX`…`EX` sections, and +where two regular characters meet, the token before is a number they cannot continue +or an operator the next 1–3 characters do not extend into another (`ET`|`BT` and +`0`|`cm` pass; `s`|`h`, `12`|`3`, `/F1`|`2` do not); the walker uses the same rule +(`content/joins.rs`). Otherwise #34 refuses the save (`INVALID_OUTPUT`), with or +without text changes, so a join that would show text the original hid (a line +commented out across a part boundary) is never published. Before v0.4 any such page +next to a stamp failed validation (VO-01, E2E-10c, `tests_e2e/join.rs`). + +**Why a painted-over fake fails.** Tests build each fake with the test kit or real +qpdf passes and show that each listed check fails on its own: + +| Fake | Fails at | +| --- | --- | +| Original part kept, a white box and new text appended (GATE-18) | A2, A3, A4, A5; Phase B B1, B2, B3 | +| Real `qpdf --overlay` of a white box and text page (GATE-19) | A2, A3, A5; B1, B2 | +| Page replaced by a rendered image (GATE-20) | A2, A3, A4 | +| FreeText annotation over the old words (GATE-22) | A2 | +| Old stream re-attached as an unused Form XObject (GATE-24) | A2, A3; B2 | +| Wrapper Form that still holds the old parts (PB-05) | B1, B2, B3 | +| Wrapper whose `/BBox` cuts the page (PB-06) | B1 | +| Fake injected into the edited copy during a real Save (GATE-25) | Phase A; nothing is published | + +#34 alone passes the first fake, because the original content is still present +(GATE-21). That is why Edit text has its own gate. + +**Independent engines.** Poppler is used in the app, not only in tests, so a +mistake in OffPDF's width or encoding model cannot verify itself. If the Poppler +tools are missing, Edit text refuses with `VERIFIER_MISSING`; it never skips the +check. Frames are pinned for `/Rotate` 0/90/180/270 and offset CropBoxes (IND-06). + +## 10. Limits and budgets + +All limits are constants in `src-tauri/src/pdf_engine/text_edit/limits.rs`. +Exceeding a page budget refuses that page (`PAGE_TOO_COMPLEX`) and leaves other +pages editable. Exceeding a file budget refuses the file. Inside the gate, a budget +failure is `EDIT_VERIFY_FAILED` ("too large to verify"), never a skipped check. + +| Area | Limit | +| --- | --- | +| File | 400 MiB (`FILE_TOO_LARGE`); 2,000,000 objects; xref chain of 64 sections; object nesting 100; object streams 32 MiB each and 256 MiB in total decoded; xref streams 64 MiB decoded (`FILE_TOO_COMPLEX`) | +| Any stream | 32 MiB decoded, capped while inflating | +| Page content | 48 MiB decoded; 256 parts; 250,000 operators; 400,000 glyphs; 20,000 lines; 256 fonts; `q` nesting 64; Form nesting 8; 96 MiB decode budget for content, forms, fonts and CMaps of one page | +| Page model | 160 MiB per page for the model, the font models it keeps and its #33 Classify pass together (the lexed operators are dropped once the page is walked). Unused fonts of the page's resources are kept only within 32 MiB and a quarter of the room left, else left out. At most about 131,000 one-glyph shows under one state; 3,000 lines with a `Tm` per glyph hold 89–112 MiB | +| Fonts | program 16 MiB decoded; ToUnicode 2 MiB and 131,072 mappings; `/W` 65,536 entries | +| Edits | 1,000 characters per line; 200 changed lines per page; 500 per Save | +| Tolerances | glyph drift 0.01 pt; join baseline 0.01 pt; join gap ±0.3 em; Poppler word boxes 0.05 pt; pixel channel difference 24; at most 8 differing pixels outside the edited glyph boxes; at most 2 in other runs' glyph boxes outside the edited glyphs' own boxes (grown by 1 pixel), and 2 of a neighbour's ink lost under the new glyphs | +| Rendering for A5 | at most 96 DPI and 25 million pixels per page; a page that would render below 25 DPI fails closed | +| Subprocesses | 120 s timeout each; `pdftotext` output 16 MiB; qpdf JSON 64 MiB | +| Preview | one-page file of at most 48 MiB, else "Preview unavailable" (Save still checks) | +| Cache | 2 open sources, each kept only up to 256 MiB (calls on a larger file that overlap share one read); per source 32 page models and 256 MiB of them, fonts included; a model over 64 MiB leaves the cache when previewed and is built again on the next visit | +| Save | Page models are kept from planning to Phase A within 2 × the file size (at least 16 MiB); once one does not fit, none is kept and Phase A builds each again, one at a time. The source's and the verification copies' raw bytes are dropped once read | + +Measured by the test suite on 2026-10-03 (debug build): a page of 100,000 +one-glyph shows holds an 84 MiB model and its Classify pass 22 MiB more; planning +an edit on it adds 103 MiB, and its preview peaks at 209 MiB for the whole +process; saving three pages of 30,000 shows peaks at 66 MiB (one page: 54 MiB). + +## 11. Compatibility matrix + +Expected behaviour per producer. Each row is backed by a synthetic, +producer-shaped fixture built in memory by the test kit (no third-party PDFs are +committed). Real files are checked with the manual protocol in +[section 17](#17-manual-test-protocol). + +| Producer | Typical structure | v0.4 | Evidence | +| --- | --- | --- | --- | +| Microsoft Word, Latin | TrueType subset, WinAnsi, `/Widths`, ToUnicode, kerned `TJ` per format run, page-sized clip, tagged (`/MCID`), bookmarks | Editable; letters limited to the subset | FX-WORD: E2E-03, PLAN-04, IND-01, RUN-01 | +| Word, Turkish / Polish / Czech | the same plus a Type0 companion font for non-WinAnsi letters | Editable as one line (sibling join); ğ ş ı İ when present in either subset | FX-WORD-TR: E2E-04, RUN-05, IND-01 | +| Tagged and bookmarked files (Word's default) | StructElem `/Pg`, annotation `/P`, outline and link destinations, `/OpenAction` | Editable (these references do not count as sharing) | CON-07, CON-08, CLS-H27, E2E-22 | +| Word-style hybrid xref | object stream listed in the classic table and the `/XRefStm` | Editable when lopdf and qpdf agree; otherwise `PDF_NEEDS_REPAIR` | SNAP-16, E2E-22; SNAP-11, SNAP-12 | +| LibreOffice | symbolic TrueType subset, hex `TJ` with kerns | Editable per joined line | FX-LIBRE: E2E-05, IND-01 | +| LibreOffice with OpenType-CFF fonts | `/Type1C` or `/OpenType` | Editable | FX-LIBRE-CFF: E2E-05, IND-01 | +| Google Docs, Chrome and Edge print | Type0 Identity-H subset, flipped `cm`/`Tm`, one operator per text node, synthetic bold (render mode 2) and italic (shear) | Editable; colour unavailable for synthetic bold | FX-SKIA: E2E-05, IND-01, GEO tests | +| Firefox, Safari / Quartz | symbolic subsets with ToUnicode, sometimes one glyph per operator | Editable; per-glyph operators on one line join (a page that stays one glyph per line is `PER_GLYPH_TEXT`) | FX-QUARTZ, FX-PERGLYPH: E2E-05, IND-01, RUN-12 | +| pdfTeX / LaTeX | Type1 subsets with `/Differences`, often no ToUnicode or space glyph, word gaps as kerns | Editable (glyph names, kern spaces); math fonts usually `AMBIGUOUS_UNICODE` | FX-PDFTEX: E2E-05, RUN-07, IND-01 | +| XeLaTeX / LuaLaTeX | Type0 Identity-H CIDFontType0C | Editable | FX-XETEX: E2E-05, IND-01 | +| InDesign / Illustrator | CFF fonts, tracking, layers, some text in Forms | Editable on visible layers; text in Forms listed as refused lines (`NESTED_FORM`); hidden layers `OPTIONAL_CONTENT` | FX-INDD: E2E-05, IND-01; corpus `text-nested-form` | +| Report generators (FPDF, TCPDF, ReportLab, wkhtmltopdf) | Standard 14 not embedded, or TrueType subsets | Editable, with the right widths per face | FX-STD14: E2E-01, E2E-02, E2E-05, CLS-H18 | +| Office exports without embedded fonts | simple TrueType with `/Widths`, no program | Editable, marked "not included in the PDF" | FX-NONEMB: E2E-05 | +| Scanner + OCR | image plus invisible text | `INVISIBLE_TEXT` | FX-OCR; CLS-H13 | +| CAD, plotting, design tools | Type3, outlines, patterns, deep Forms, `/UserUnit` | `TYPE3`, `PATTERN`, `NESTED_FORM`, `GEOMETRY` or no text | corpus `text-type3`, `text-nested-form`; GEO-12…15 | +| Imposition, letterhead templates | shared content streams | `SHARED_CONTENT` | FX-SHARED: CLS-H17, GATE-14 | +| Invoices and tables | each cell its own show operator on a shared baseline | Editable; a longer cell that runs into the next saves with `NEXT_TEXT_OVERLAP` | IND-11, `tests_e2e/save.rs` (B2) | +| Content split into parts without line breaks | `/Contents` arrays whose parts meet without whitespace | Editable when the parts meet between tokens (`ET`\|`BT`, `0`\|`cm`, a comment's end of line); a boundary inside a token, comment or `BX`…`EX` section, or within 64 bytes after an inline image → `MALFORMED_CONTENT` | E2E-10c, VO-01, `tests_walk/joins.rs`, `tests_e2e/join.rs`, PB-09 | +| Pages sharing inherited `/Resources` | a bold face only in the shared dictionary | Editable; preview and Save agree | PREV-06 | +| Legacy filters on other pages | RunLength or LZW content left untouched | Editable pages unaffected; those pages keep their raw bytes; an edited page with such content is `UNSUPPORTED_FILTER` | APP-06, GATE-29 | +| Signed, encrypted, dynamic XFA | — | `SIGNED`, `ENCRYPTED`, `UNSUPPORTED_XFA` at open; an unedited signed file may be appended | SNAP tests, E2E-13, E2E-20 | +| qpdf 11.9 (CI) and 12.x (bundled, dev) | `--overlay` wraps pages in a Form; content parts are joined with qpdf's rule | Both forms accepted by Phase B; the join rule and JSON shapes are pinned on whichever qpdf is installed | APP-07, PB-02, CON-10, ENG-07 | + +## 12. Performance and engine matrix + +Measured on an Apple M5 (10 cores, 24 GiB), macOS 26.4.1, rustc 1.96.0, qpdf +12.3.2, Poppler 26.04.0, release build, **2026-10-03**, on the final v0.4 code. +Other builds shared the machine (load average 30–38 on 10 cores), so times are +pessimistic; peak RSS is not affected. + +**Engine matrix (the four #33 axes).** + +| Path | Binary size | Memory (peak RSS) | Build and signing cost | Native crash isolation | +| --- | --- | --- | --- | --- | +| Shipped: lopdf 0.34 (read only) + own lexer and walker + qpdf writer + Poppler verifier | macOS arm64 executable 12,866,048 → 13,991,552 bytes (+1,125,504 bytes, +8.7 %); bundled tools unchanged, so this is the whole bundle delta. Windows: not measured | 51 MiB (10 pages), 86 MiB (100), 114 MiB (1,000 pages); 639 MiB for a 305 MiB image-heavy file (2.10 × the file) | No new native binary: `qpdf`, `pdftotext` and `pdftoppm` are already bundled and signed (`scripts/sign-macos-bundled-tools.sh`, `scripts/prepare-poppler-windows.ps1`). No new crate; `flate2` moved from dev-dependency to dependency (MIT/Apache-2.0, already compiled in through lopdf). | qpdf and Poppler run out of process with a 120 s timeout. lopdf parses in process under `panic = "abort"`, behind a raw preflight (xref chain, nesting, `/Length` chains), object-stream guards, per-page budgets and fuzzing. Residual risk: lopdf's parser on inputs within those bounds (the same exposure every existing lopdf feature has). | +| PDFium sidecar | not prototyped | not prototyped | not prototyped (would add a per-platform native binary to build, sign and notarise) | not prototyped | + +BENCH-SIZE: `src-tauri/target/release/offpdf` from `cargo build --release -j 6` +(shipped profile: `panic = "abort"`, LTO, one codegen unit, `opt-level = "s"`, +stripped). Before: 2026-09-30, the tree before the first v0.4 change; after: +2026-10-03, the final code. The delta includes the embedded new editor UI. + +**BENCH-01** (generated documents, 40 Helvetica lines per page; inspect over 20 +sampled pages, cold = first visit, warm = cached; preview of one edit on 10 +pages; Save with N changed lines spread over the pages; times in ms): + +| Pages | File KiB | Open | Inspect cold p50 / p95 | Inspect warm p50 / p95 | Preview p50 / p95 | Save 1 | Save 10 | Save 100 | Peak RSS MiB | +| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | +| 10 | 29.8 | 25 | 1 / 3 | 0 / 0 | 360 / 399 | 416 | 2,341 | 2,329 | 51 | +| 100 | 298.7 | 32 | 1 / 2 | 1 / 2 | 369 / 388 | 440 | 2,321 | 21,979 | 86 | +| 1,000 | 3,026.1 | 86 | 2 / 4 | 0 / 1 | 346 / 513 | 1,069 | 2,630 | 22,969 | 114 | + +**BENCH-02** (a 305 MiB file: 20 text pages and 31 image pages, each with a +5 MiB Flate image and a JPEG): + +| File MiB | Open ms | First inspect ms (after open) | Background `qpdf --check` done ms (from open) | Preview ms | Save 1 edit ms | Peak RSS MiB | RSS / file | +| --- | --- | --- | --- | --- | --- | --- | --- | +| 305 | 1,555 | 1,345 | 4,256 | 1,713 | 11,850 | 639 | 2.10 | + +Budgets: first inspect ≤ 3 s after open (1.3 s), Save ≤ 90 s (11.9 s), peak RSS +≤ 3 × file (2.10 ×). All met. + +What the numbers mean for users: + +- Opening is fast; the full `qpdf --check` runs in the background and does not + block the first inspect. +- Save time grows with the number of edited **pages**, not lines (Phase A runs + Poppler twice per edited page, about 0.2 s each here): 100 pages took about 22 s. +- Files above 100 MB work but check and save more slowly, which the app says when + the file opens; files above 256 MiB are re-read by later visits (section 15). + +## 13. Engine decision + +**lopdf 0.34 (read only) + OffPDF's own lexer and walker + qpdf as the writer + +Poppler as the independent verifier. No PDFium sidecar.** + +- **Byte-exact splices need a byte-offset view of the content.** lopdf's + `Content::decode`/`encode` lose the original bytes, and `Content::decode` + returns a silent prefix for a stream with garbage in the middle (test LEX-20), + so OffPDF reads content with its own bounded lexer. + These lopdf calls are forbidden in the module and a test enforces it. +- **PDFium's editing API regenerates page content** (`FPDFPage_GenerateContent`), + the opposite of a byte-exact splice, so it could not keep unchanged glyphs and + kerns byte for byte. +- **A sidecar adds a per-platform native binary** to build, sign and notarise. + The shipped path adds none: qpdf and Poppler are already bundled and signed. +- **Crash isolation is obtained differently.** Everything that writes or + independently verifies runs out of process (qpdf, Poppler). lopdf only reads, + from bounded, preflighted bytes, with per-page budgets and fuzzed parsers. + Moving the lopdf read into a helper process is a v0.5 option. +- **lopdf and qpdf must agree.** Because lopdf 0.34 can read some files + differently from qpdf (pages without `/Type`, hybrid xref sections), every + file is checked for page-map agreement at open and on every file the gate reads; + disagreement is `PDF_NEEDS_REPAIR`. + +The #33 criterion "if PDFium is tested, it runs out of process" is not +triggered: PDFium was not prototyped (PR #98 honesty convention). + +## 14. UX wording rules + +- Allowed verbs: "change", **Edit text**, **Add text**. Edit text changes words + already in the PDF; Add text places a new text box. +- Never in the UI: "cover", "white-out", "flatten", "overlay", "edit like Word", + "we'll". `sourceTextCopy.test.ts` scans every exported string, and + `docsReasons.test.ts` scans the quoted copy blocks of this page. +- Every refusal names its reason (sections 6 and 7); nothing is greyed out + without a sentence. +- Every Save failure ends with the sentence that the original file was not + changed. +- All copy lives in `src/lib/editor/sourceTextCopy.ts`; Rust builds the + `AppError` copy in `text_edit/reasons.rs` with the same text. + +Examples as shipped: + + +| Key | Sentence | +| --- | --- | +| `UI.tool.editText.title` | Edit text (E): change words already in this PDF, in its own font | +| `UI.tool.addText.title` | Add text: place a new text box on the page | +| `UI.banner.mode` | Select an outlined line to change its words. Changes use the document's own font, and the rest of the page stays exactly where it is. Dotted lines can't be changed safely — select one to see why. | +| `UI.banner.addText` | Adds a new text box. To change words already on the page, use Edit text (E). | +| `UI.popover.title` | This text can't be changed safely | +| `UI.popover.footer` | You can add a new text box on top with Add text. It won't replace this text. | +| `UI.outputAlert` | Text changes use the document's own font. Before saving, OffPDF checks that nothing else on those pages moved by more than 0.01 pt. The original file is never changed. | +| `UI.substituted` | This font isn't included in the PDF. Your PDF reader draws it with a similar font. | +| `UI.removal` | This line will be removed. Text after it stays where it is. | +| `UI.blocked` | Fix or cancel this change first. Press Esc to cancel. | +| `UI.face.noBold` | This page has no bold version of this font. | +| `UI.face.noItalic` | This page has no italic version of this font. | +| `UI.face.cannotDrawBold` | The bold version of this font can't draw: {chars} | +| `UI.style.size` | This line's size is set in a way OffPDF can't change. | +| `UI.style.face` | This line's font is set in a way OffPDF can't change, so bold and italic aren't available. | +| `UI.style.colour` | This text is drawn with an outline, so its colour can't be changed here. | +| `UI.status.previewUnavailable` | Preview unavailable for this page. Your changes are checked again when you save. | +| `UI.notVisible` | This change doesn't change how the page looks. Something may be drawn over this line. | + + +## 15. Known limitations and v0.5 candidates + +Limitations of v0.4, in plain terms: + +- **A damaged JPEG anywhere in the file blocks preview and Save** with + `PDF_NEEDS_REPAIR`. `qpdf --check` decodes every JPEG, and "JPEG data is + corrupt" is not on the allow-list of harmless warnings (only linearization, + hint-table and `/Size` warnings are). Scanner output with one damaged image is + affected. Run the file through Repair PDF first. +- **Reading a page and previewing a change cannot be cancelled** (Save can). The + part-boundary scan, also in Save's checks, cannot be stopped and inflates each + Flate inline image again after the page's lex (debug: 4.9 s per 96 MiB of plain + operators; 0.76 s for 64 images of 16 MiB, so ~70 s for a crafted 96 MiB page). +- **Files over 256 MiB are not kept between calls.** Calls that overlap share one + read; a later page visit or preview reads the file again (about 1.5 s at + 305 MiB on the bench machine). +- **Very dense pages are refused** (`PAGE_TOO_COMPLEX`) when the model, its fonts + and its #33 Classify pass would hold more than 160 MiB (about 131,000 one-glyph + shows under one state). A font used by several pages counts in full on each; + unused fonts past their share are left out, so fewer sibling faces may be + offered for bold and italic. An ExtGState with more than 64 keys, or key names + longer than 127 bytes, also refuses the page. +- **The part-boundary check is conservative**: a `/Contents` part boundary inside an + inline image, a hex string or a `BX`…`EX` section, or within 64 bytes after an + inline image, refuses the page (`MALFORMED_CONTENT`), even at whitespace. +- **A longer line that runs into the next text** is checked by Poppler for that + text's words, place and ink, but where the new letters cover it not its colour or + outline (a neighbour turned red or blue, or drawn stroked) nor a small shift under + one letter: OffPDF's own re-walk (A4), which compares every glyph's colour, render + mode and position, catches those. +- **Rarely, a correct overlapping change is refused** (`EDIT_VERIFY_FAILED`): when + the new letters run over the next text in the colour of the paper behind it + (white letters from a dark band into black text on white), the pixel check + cannot tell that from the neighbour being erased. +- **Only letters the embedded font contains** can be typed; subset fonts usually + hold only the letters the document already uses. Ligature glyphs read as their + letters but are never typed. +- **One line at a time**, one style per line: no reflow, wrapping, per-letter + styling, new fonts, underline or stroke colour. +- **Very large pages are checked at a lower resolution.** The pixel check renders + at most 25 million pixels: pages larger than about 52 in on a side are rendered + below 96 DPI, so it sees less detail (the 0.01 pt re-walk still applies), and a + page that would render below 25 DPI (around 200 in on a side) fails closed with + `EDIT_VERIFY_FAILED`. +- **Type1, non-embedded and some CFF fonts** use a conservative box instead of the + exact glyph outline for the new letters in the pixel check; a neighbour within + about 0.2 em of them is checked for its words and ink there, not its colour. +- **Temporary copies.** Each open file has a copy in the app's temp folder + (`textedit/`), deleted when the file is closed or evicted; a running preview keeps + it until it ends, and copies a crash left behind are deleted at the next start. +- **Reading order** for the keyboard follows the structure tree on tagged pages; on + untagged pages it is a heuristic (XY-cut). Lines drawn through Form XObjects come + after the page's own lines, and are listed only when the page's #33 check + completes. +- **A page listed twice** in the file list, or a page with a redaction in the same + Save, cannot take text changes. +- **The lopdf parser runs in process.** It only reads, within the preflight and + budgets above; a malformed file that passes them shares the exposure of every + existing lopdf feature. + +Out of v0.4 by decision (never faked, always refused with a reason): image +replacement, text inside Form XObjects, annotation and form-widget text, +shared content streams, Type3, vertical, right-to-left and complex scripts, +`/UserUnit` ≠ 1, reflow and multi-line editing, font embedding or substitution. + +v0.5 candidates: + +- copy-on-write for shared content streams through qpdf JSON page dictionaries + (probe P2 shows qpdf can write a new stream and drop the old one); +- text inside Form XObjects (copy-on-write of the Form per page); +- Type3 fonts; +- existing-image replacement (needs its own spike: shared XObjects, masks, size + policy); +- render only the edited bands at full resolution for very large pages; +- cancellable inspect and preview; +- the lopdf read in a helper process; +- a decision on whether DCT decode warnings from `qpdf --check` may be treated as + harmless for files whose damaged image is not on an edited page. + +## 16. Contributor notes + +### Module map + +Rust, `src-tauri/src/pdf_engine/text_edit/` (no file over 800 lines; children in +same-named folders): + +| Layer | Files | +| --- | --- | +| Core | `limits.rs` (every budget), `reasons.rs` (codes and `AppError` copy), `snapshot.rs` (+ `snapshot/{preflight,objects,headers}.rs`: one capped read, preflight, guarded lopdf load, page-map agreement), `decode.rs` (capped Flate, ASCIIHex, ASCII85), `lexer.rs` (byte-offset tokenizer, inline-image proofs), `content.rs` (+ `content/joins.rs`: parts, joined buffer, qpdf join rule, part-boundary check, ownership), `engines.rs` (tools, subprocesses with timeout and cancel, `qpdf --check` classification and memo) | +| Fonts | `fonts/` (encodings, AGL, Standard 14 data, ToUnicode, glyph presence for TrueType, CFF, Type1; faces and typing surfaces) | +| Page model | `context.rs`, `geometry.rs` (strict boxes and orientation), `state.rs`, `structure.rs` (optional content, ActualText, structure order), `walker.rs` (+ `walker/`, incl. the per-page byte budget and font charging), `runs.rs` (+ `runs/`, incl. `forms.rs`: Form XObject lines), `order.rs` | +| Write | `encode.rs` (numbers, grammar), `rewrite.rs` (+ `rewrite/`: `plan_page`), `fit.rs` (word spacing; `fit/estimate.rs`, test-only, is the width estimate the frontend golden is checked against), `apply.rs` (qpdf update) | +| Verify | `verify.rs` (re-walk), `graph.rs` (canonical whole-graph digest), `poppler.rs` (+ `poppler/words.rs`: words and neighbours; `poppler/ink.rs`: neighbour ink; `poppler/near.rs`: the model's glyphs around the edit), `gate.rs` (+ `gate/`: Phase A, Phase B, `join.rs` token-neutral joins), `preview.rs` | +| Integration | `cache.rs`, `dto.rs`, `service.rs`, `export.rs` (+ `export/map.rs`: Save), `commands/text_edit.rs` (4 Tauri commands) | +| #33 classifier | `pdf_engine/source_content.rs`, an adapter over the page model; `service::inspect_page` builds every "can / can't be changed" from it | + +Frontend: `src/lib/editor/{sourceText,sourceTextCopy}.ts` (pure helpers, copy), +`src/lib/editor/text-reasons.json` (codes), `src/lib/types.ts` (IPC types), +`src/lib/tauriCommands.ts` (IPC), `src/features/edit-pdf/{useTextSources,textSaveGuards}.ts`, +`src/components/pdf/editor/sourceText/` (mode, layer, editor, format bar, reason +popover, page text and preview hooks), `src/styles/source-text.css`. + +### Forbidden APIs + +New code never calls lopdf's `Content::decode`/`encode`, `string_to_bytes`, +`replace_text`, `encode_text`, `decompressed_content`, `get_plain_content`, +`get_page_content`, `Document::load`, `load_mem`, `.decompress()`, +`IncrementalDocument`, `.save(` or `save_to(`. `tests_guard.rs` (GUARD-01) scans +`text_edit/` and every file whose first line is `//! offpdf:forbidden-api-scan`, +skipping comments and string literals. Production code has no +`unwrap`/`expect`/`panic!`/unchecked indexing on PDF-derived data (clippy deny +attributes in `text_edit/mod.rs`). Test seams exist only under `#[cfg(test)]`; +no environment variable or setting can weaken a check. + +### Reason codes, engines, benchmarks and fixtures + +Adding a reason code, the engine skip policy (`OFFPDF_REQUIRE_ENGINES=1`, no +`skip:` lines), BENCH and golden regeneration, and adding a font class or a +producer fixture: [`EDIT_MODEL.md`](../src/lib/editor/EDIT_MODEL.md#edit-text-contributor-workflow). + +## 17. Manual test protocol + +Reported in the PR; tick only what was actually run. Produce the files locally +and never commit them. + +- [ ] For Word (Latin and Turkish; tagged, with bookmarks), LibreOffice, Google + Docs, Chrome print, pdfTeX, XeLaTeX, InDesign (if available) and macOS + Pages: a one-page PDF each. +- [ ] A two-column page (keyboard and VoiceOver reading order). +- [ ] A page with a stamp added in the same Save (qpdf's overlay wrapper path). +- [ ] Open Edit text; record editable and refused counts and the top reasons. +- [ ] Change three lines (shorter, same width, longer but fitting), including one + style change; Save. +- [ ] Open the result in Acrobat Reader, macOS Preview, Chrome and Firefox; search + finds the new words and not the old; compare at 400 %. +- [ ] Repeat on a rotated page and on a cropped page. +- [ ] Keyboard only: E → Tab → arrows → Enter → type → Tab → Bold → Enter → ⌘Z → + ⌘⇧Z → Save. +- [ ] VoiceOver labels on lines, the editor, the format bar and the reason dialog. +- [ ] Light and dark themes at 1440 and 1024 px wide (screenshots attached). diff --git a/fixtures/source-edit/README.md b/fixtures/source-edit/README.md index 41d3c33..19b2d74 100644 --- a/fixtures/source-edit/README.md +++ b/fixtures/source-edit/README.md @@ -14,3 +14,14 @@ Regenerate: `write_corpus_fixture(id, dest)` in After lopdf `save` it runs `qpdf in out` when qpdf is available so `qpdf --check` passes. Committed `*.pdf` files must be that writer’s output so uncompressed stream dumps match. + +Manifest intents never claim editability; the engine decides and Edit text +re-verifies every change at save. Text-edit fixtures are generated in temp by +`pdf_engine/text_edit/testkit`. + +Classifier expectations (v0.4, `source_content_integ.rs`): `text-cid-tounicode` +now reports `FONT_NOT_EMBEDDED` instead of `AMBIGUOUS_UNICODE` (its Type0 font +has no embedded program, so glyph presence cannot be proven). +`text-cid-no-tounicode` reports `FONT_NOT_EMBEDDED` too, which comes before +`NO_TOUNICODE` in the reason order. The PDFs and `manifest.json` are unchanged; +both intents stay `try-edit`. diff --git a/src-tauri/Cargo.toml b/src-tauri/Cargo.toml index 027e909..903af22 100644 --- a/src-tauri/Cargo.toml +++ b/src-tauri/Cargo.toml @@ -31,6 +31,9 @@ lopdf = "0.34" # Glyph metrics for the bundled editor font (Noto Sans). ttf-parser = "0.25" +# Bounded (capped while inflating) Flate decoding for Edit text. Already compiled in through lopdf. +flate2 = "1" + # Decode common image formats so uploaded images can be wrapped into a PDF. image = { version = "0.25", default-features = false, features = [ "jpeg", @@ -44,9 +47,6 @@ image = { version = "0.25", default-features = false, features = [ # app so users do not need to install a system codec or upload their photos. heif-rs = { path = "vendor/heif-rs" } -[dev-dependencies] -flate2 = "1" - [target.'cfg(any(target_os = "macos", windows, target_os = "linux"))'.dependencies] tauri-plugin-single-instance = "2" diff --git a/src-tauri/src/commands/jobs.rs b/src-tauri/src/commands/jobs.rs index c0656ad..1213420 100644 --- a/src-tauri/src/commands/jobs.rs +++ b/src-tauri/src/commands/jobs.rs @@ -3,6 +3,28 @@ use crate::error::AppError; use crate::models::JobRegistry; +/// Longest job id the frontend sends (`crypto.randomUUID()` is 36 characters). +const JOB_ID_MAX: usize = 128; + +/// The webview's job id names a work folder (`/work/`) that the job deletes when it +/// ends: only `[A-Za-z0-9_-]{1,128}` is accepted, so it can never step out of the temp root. +pub(crate) fn check_job_id(job_id: &str) -> Result<(), AppError> { + let ok = !job_id.is_empty() + && job_id.len() <= JOB_ID_MAX + && job_id + .bytes() + .all(|b| b.is_ascii_alphanumeric() || b == b'-' || b == b'_'); + if ok { + return Ok(()); + } + Err(AppError::new( + "INVALID_JOB", + "Invalid job", + "OffPDF received a job id it does not accept.", + ) + .with_details(format!("job id of {} bytes", job_id.len()))) +} + /// Cancel a running job by id. Killing the child (if any) is handled by the /// `JobHandle`; the worker thread reaps it and returns `AppError::cancelled`. #[tauri::command] @@ -15,3 +37,37 @@ pub async fn cancel_job( } Ok(()) } + +#[cfg(test)] +mod tests { + use super::check_job_id; + + /// review-T5: a job id reaches `remove_dir_all(/work/)`. + #[test] + fn job_ids_cannot_leave_the_work_folder() { + for ok in [ + "3f2b8c1e-0d4a-4c6e-9a51-7b2f0e9d1c34", + "job-1700000000000-1a2b", + "a_b", + ] { + assert!(check_job_id(ok).is_ok(), "{ok}"); + } + let long = "a".repeat(129); + for bad in [ + "", + "..", + "../x", + "a/b", + "a\\b", + "C:x", + "a b", + long.as_str(), + "x\0", + ] { + let e = check_job_id(bad) + .err() + .unwrap_or_else(|| panic!("{bad:?} accepted")); + assert_eq!(e.code, "INVALID_JOB"); + } + } +} diff --git a/src-tauri/src/commands/mod.rs b/src-tauri/src/commands/mod.rs index 95e59c6..e4c7833 100644 --- a/src-tauri/src/commands/mod.rs +++ b/src-tauri/src/commands/mod.rs @@ -5,3 +5,4 @@ pub mod files; pub mod pdf; pub mod jobs; pub mod render; +pub mod text_edit; diff --git a/src-tauri/src/commands/pdf.rs b/src-tauri/src/commands/pdf.rs index bc1bc16..0c13757 100644 --- a/src-tauri/src/commands/pdf.rs +++ b/src-tauri/src/commands/pdf.rs @@ -43,6 +43,7 @@ pub async fn merge_pdfs( input_paths: Vec, output_path: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -65,6 +66,7 @@ pub async fn split_pdf( picks: Vec, mode: SplitMode, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -86,6 +88,7 @@ pub async fn assemble_pdf( output_path: String, groups: Vec, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -108,6 +111,7 @@ pub async fn edit_pdf( groups: Vec, rotations: Vec, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -130,6 +134,7 @@ pub async fn unlock_pdf( output_path: String, password: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -153,6 +158,7 @@ pub async fn watermark_pdf( text: String, opacity: f64, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -187,6 +193,7 @@ pub async fn crop_pdf( right: f64, bottom: f64, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -223,6 +230,7 @@ pub async fn edit_pdf_overlays( flatten_form: Option, flatten_annotations: Option, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -278,6 +286,7 @@ pub async fn stamp_pdf( color: [f64; 3], size_pct: f64, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -316,6 +325,7 @@ pub async fn poster_pdf( overlap: f64, marks: bool, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -352,6 +362,7 @@ pub async fn nup_pdf( sheet_w: f64, sheet_h: f64, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -388,6 +399,7 @@ pub async fn add_page_numbers( pad_width: Option, with_date: Option, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -430,6 +442,7 @@ pub async fn protect_pdf( user_password: String, owner_password: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -460,6 +473,7 @@ pub async fn extract_pages( output_path: String, pages: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -482,6 +496,7 @@ pub async fn delete_pages( output_path: String, pages: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -505,6 +520,7 @@ pub async fn rotate_pages( angle: i32, rotate_pages: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -535,6 +551,7 @@ pub async fn reorder_pages( output_path: String, order: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -559,6 +576,7 @@ pub async fn compress_pdf( quality: u32, target_bytes: Option, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -589,6 +607,7 @@ pub async fn optimize_pdf( output_path: String, groups: Vec, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); diff --git a/src-tauri/src/commands/render.rs b/src-tauri/src/commands/render.rs index ab5273a..2bfd15d 100644 --- a/src-tauri/src/commands/render.rs +++ b/src-tauri/src/commands/render.rs @@ -82,6 +82,7 @@ pub async fn pdf_to_images( format: String, dpi: u32, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -211,6 +212,7 @@ pub async fn ocr_pdf( picks: Vec, lang: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -258,6 +260,7 @@ pub async fn office_to_pdf_batch( output_dir: String, input_paths: Vec, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -335,6 +338,7 @@ pub async fn pdfa_pdf( output_path: String, groups: Vec, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -394,6 +398,7 @@ pub async fn detect_blank_pages( input_path: String, sensitivity: String, ) -> Result, AppError> { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let jid = job_id.clone(); let res = tauri::async_runtime::spawn_blocking(move || { @@ -431,6 +436,7 @@ pub async fn write_pdf_meta( fields: crate::pdf_engine::metadata::PdfMeta, clear_all: bool, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -468,6 +474,7 @@ pub async fn export_pdf_text( first_page: Option, last_page: Option, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); @@ -504,6 +511,7 @@ pub async fn pdf_to_office( groups: Vec, format: String, ) -> Result { + super::jobs::check_job_id(&job_id)?; let handle = registry.register(&job_id); let app2 = app.clone(); let jid = job_id.clone(); diff --git a/src-tauri/src/commands/text_edit.rs b/src-tauri/src/commands/text_edit.rs new file mode 100644 index 0000000..6ca22e0 --- /dev/null +++ b/src-tauri/src/commands/text_edit.rs @@ -0,0 +1,94 @@ +//! Edit text commands (SPEC §B.20): open a source, inspect a page, preview edits, release. +//! Paths in, JSON out; the work runs on a blocking worker. Save stays `edit_pdf_overlays`. + +use crate::error::AppError; +use crate::pdf_engine::text_edit::cache::TextEditCache; +use crate::pdf_engine::text_edit::dto::{PageTextDto, TextEditIn, TextPreviewDto, TextSourceDto}; +use crate::pdf_engine::text_edit::engines::Engines; +use crate::pdf_engine::text_edit::service; +use crate::utils::temp; + +/// Runs `f` on a blocking worker with the resolved engines and the temp root. +async fn on_worker(app: tauri::AppHandle, f: F) -> Result +where + T: Send + 'static, + F: FnOnce(&Engines, &std::path::Path) -> Result + Send + 'static, +{ + tauri::async_runtime::spawn_blocking(move || { + let engines = Engines::resolve(&app)?; + let temp_root = temp::root(&app)?; + f(&engines, &temp_root) + }) + .await + .map_err(|e| AppError::engine_failed(format!("worker join error: {e}")))? +} + +#[tauri::command] +pub async fn open_text_source( + app: tauri::AppHandle, + cache: tauri::State<'_, TextEditCache>, + input_path: String, +) -> Result { + let cache = cache.inner().clone(); + on_worker(app, move |engines, temp_root| { + service::open_source(&cache, engines, temp_root, &input_path) + }) + .await +} + +#[tauri::command] +pub async fn inspect_text_page( + app: tauri::AppHandle, + cache: tauri::State<'_, TextEditCache>, + input_path: String, + fingerprint: String, + page_index: u32, +) -> Result { + let cache = cache.inner().clone(); + on_worker(app, move |engines, temp_root| { + service::inspect_page( + &cache, + engines, + temp_root, + &input_path, + &fingerprint, + page_index, + ) + }) + .await +} + +#[tauri::command] +pub async fn preview_text_edits( + app: tauri::AppHandle, + cache: tauri::State<'_, TextEditCache>, + input_path: String, + fingerprint: String, + page_index: u32, + edits: Vec, +) -> Result { + let cache = cache.inner().clone(); + on_worker(app, move |engines, temp_root| { + service::preview_edits( + &cache, + engines, + temp_root, + &input_path, + &fingerprint, + page_index, + &edits, + ) + }) + .await +} + +#[tauri::command] +pub async fn release_text_source( + cache: tauri::State<'_, TextEditCache>, + input_path: String, +) -> Result<(), AppError> { + let cache = cache.inner().clone(); + tauri::async_runtime::spawn_blocking(move || service::release_source(&cache, &input_path)) + .await + .map_err(|e| AppError::engine_failed(format!("worker join error: {e}"))) +} diff --git a/src-tauri/src/lib.rs b/src-tauri/src/lib.rs index 63a7170..f195042 100644 --- a/src-tauri/src/lib.rs +++ b/src-tauri/src/lib.rs @@ -31,6 +31,12 @@ //! //! Jobs (`commands::jobs`): //! - `cancel_job(registry, job_id: String) -> Result<(), AppError>` +//! +//! Edit text (`commands::text_edit`, state `text_edit::cache::TextEditCache`): +//! - `open_text_source(input_path) -> TextSourceDto` +//! - `inspect_text_page(input_path, fingerprint, page_index) -> PageTextDto` +//! - `preview_text_edits(input_path, fingerprint, page_index, edits) -> TextPreviewDto` +//! - `release_text_source(input_path) -> ()` mod commands; mod error; @@ -67,6 +73,8 @@ pub fn run() { // Shared, cancellable job registry. .manage(JobRegistry::default()) .manage(Mutex::new(os_open::OpenedPathQueue::default())) + // Snapshots of the files open in Edit text (paths in, JSON out). + .manage(crate::pdf_engine::text_edit::cache::TextEditCache::default()) .invoke_handler(tauri::generate_handler![ // files / system commands::files::pick_pdf_files, @@ -127,9 +135,18 @@ pub fn run() { commands::render::read_pdf_meta, commands::render::write_pdf_meta, commands::render::export_pdf_text, + // edit text (in-place changes of existing text; Save is edit_pdf_overlays) + commands::text_edit::open_text_source, + commands::text_edit::inspect_text_page, + commands::text_edit::preview_text_edits, + commands::text_edit::release_text_source, ]) .setup(|app| { os_open::enqueue_cold_start_argv(app.handle()); + // No Edit text source is open yet: copies left by a crash are deleted. + if let Ok(root) = utils::temp::root(app.handle()) { + pdf_engine::text_edit::cache::clear_stale_folders(&root); + } Ok(()) }) .build(tauri::generate_context!()) diff --git a/src-tauri/src/pdf_engine/edit_overlay.rs b/src-tauri/src/pdf_engine/edit_overlay.rs index c824d17..4d046c6 100644 --- a/src-tauri/src/pdf_engine/edit_overlay.rs +++ b/src-tauri/src/pdf_engine/edit_overlay.rs @@ -18,10 +18,10 @@ use crate::pdf_engine::edit_redact::{ verify_redaction, RedactRegion, }; use crate::pdf_engine::validate_output::{ - catalog_flags_from_doc, content_digest, validate_staged_pdf, ContentDigest, OutputSnapshot, - PageSnapshot, + catalog_flags_from_doc, content_digest, validate_staged_pdf, + validate_staged_pdf_with_alternatives, ContentDigest, OutputSnapshot, PageSnapshot, }; -use crate::pdf_engine::{crop, edit_image, qpdf}; +use crate::pdf_engine::{crop, edit_image, qpdf, text_edit}; use crate::utils::process::{run_qpdf, run_tracked}; use crate::utils::safe_output; use crate::utils::temp; @@ -286,6 +286,23 @@ pub enum EditObjectIn { fill: Option, label: Option, }, + /// A change of existing text (Edit text): applied to the source before assembly, never painted. + SourceText { + #[serde(rename = "pageIndex")] + page_index: u32, + rect: PdfRectIn, + #[serde(rename = "sourcePageIndex")] + source_page_index: u32, + #[serde(rename = "runId")] + run_id: String, + #[serde(rename = "sourceFingerprint")] + source_fingerprint: String, + #[serde(rename = "originalText")] + original_text: String, + text: String, + #[serde(default)] + style: text_edit::rewrite::SourceTextStyleIn, + }, } #[derive(Debug, Clone, Deserialize)] @@ -332,7 +349,8 @@ impl EditObjectIn { | Self::Underline { page_index, .. } | Self::Strikeout { page_index, .. } | Self::MarkupInk { page_index, .. } - | Self::Redact { page_index, .. } => *page_index, + | Self::Redact { page_index, .. } + | Self::SourceText { page_index, .. } => *page_index, } } fn opacity(&self) -> f64 { @@ -340,7 +358,7 @@ impl EditObjectIn { return 1.0; } let o = match self { - Self::Link { .. } | Self::Redact { .. } => return 1.0, + Self::Link { .. } | Self::Redact { .. } | Self::SourceText { .. } => return 1.0, Self::Rect { opacity, .. } | Self::Ellipse { opacity, .. } | Self::Triangle { opacity, .. } @@ -364,7 +382,7 @@ impl EditObjectIn { fn object_rotate(&self) -> f64 { match self { - Self::Link { .. } | Self::Redact { .. } => 0.0, + Self::Link { .. } | Self::Redact { .. } | Self::SourceText { .. } => 0.0, Self::Rect { object_rotate, .. } | Self::Ellipse { object_rotate, .. } | Self::Triangle { object_rotate, .. } @@ -386,7 +404,10 @@ impl EditObjectIn { } fn triggers_overlay(&self) -> bool { - !matches!(self, Self::Link { .. } | Self::Redact { .. }) + !matches!( + self, + Self::Link { .. } | Self::Redact { .. } | Self::SourceText { .. } + ) } fn overlay_aabb(&self, vis: [f64; 4], page_rot: i64) -> (f64, f64, f64, f64) { @@ -433,7 +454,8 @@ impl EditObjectIn { | Self::Underline { rect, .. } | Self::Strikeout { rect, .. } | Self::MarkupInk { rect, .. } - | Self::Redact { rect, .. } => pdf_rect_to_overlay(rect, vis, page_rot), + | Self::Redact { rect, .. } + | Self::SourceText { rect, .. } => pdf_rect_to_overlay(rect, vis, page_rot), } } } @@ -1040,10 +1062,15 @@ fn validate_doc(doc: &EditDocumentIn) -> Result<(), AppError> { for o in &doc.objects { if matches!(o, EditObjectIn::Link { .. }) { links += 1; - } else { + } else if !matches!(o, EditObjectIn::SourceText { .. }) { paint += 1; } } + let redact_pages: Vec = redact_regions_from_doc(doc) + .iter() + .map(|r| r.page_index) + .collect(); + text_edit::export::validate_specs(&text_edit::export::specs_of(doc), &redact_pages)?; if paint > MAX_OBJECTS { return Err(AppError::new( "TOO_MANY_OBJECTS", @@ -1101,6 +1128,8 @@ struct OverlayPageGeom { rotate: i64, user_unit: f64, content_digest: ContentDigest, + /// Digest of the parts as qpdf joins them into its overlay Form, when that differs. + alt_content_digest: Option, } fn boxes_near(a: [f64; 4], b: [f64; 4]) -> bool { @@ -1255,7 +1284,7 @@ fn overlay_page_box(vis: [f64; 4], rotate: i64) -> [f64; 4] { /// Expand a qpdf-style page spec into 1-based numbers, preserving order. /// Unlike `parse_pages`, this does not sort or dedupe. -fn expand_page_spec(spec: &str, n: u32) -> Result, AppError> { +pub(crate) fn expand_page_spec(spec: &str, n: u32) -> Result, AppError> { let trimmed = spec.trim(); if trimmed.is_empty() || trimmed.eq_ignore_ascii_case("all") || trimmed == "1-z" { return Ok((1..=n).collect()); @@ -1313,8 +1342,10 @@ fn expand_page_spec(spec: &str, n: u32) -> Result, AppError> { Ok(out) } +/// `with_alts`: an overlay will wrap the pages, so qpdf's join digest may be expected (#34). fn collect_source_pages( groups: &[PageGroup], + with_alts: bool, ) -> Result<(Vec, Vec), AppError> { let mut docs: HashMap = HashMap::new(); let mut geoms = Vec::new(); @@ -1340,7 +1371,12 @@ fn collect_source_pages( let content = doc.get_page_content(id).map_err(|e| { AppError::engine_failed(format!("Could not read page content: {e}")) })?; + let alt_content_digest = (with_alts && doc.get_page_contents(id).len() > 1) + .then(|| text_edit::export::qpdf_joined_digest(doc, id)) + .flatten() + .filter(|d| *d != content_digest(&content)); geoms.push(OverlayPageGeom { + alt_content_digest, visible: crop::visible_box(doc, id), media: crop::media_box(doc, id), crop: crop::crop_box(doc, id), @@ -1683,7 +1719,7 @@ where } /// Same as [`export_edit_pdf_with_runner`], with an explicit `qpdf --check` binary. -fn export_edit_pdf_with_check_exe( +pub(crate) fn export_edit_pdf_with_check_exe( groups: &[PageGroup], output: &str, document: &EditDocumentIn, @@ -1723,21 +1759,31 @@ where let overlay_str = overlay.to_string_lossy().to_string(); let mut gate_passed = false; let result = (|| -> Result<(Vec, Vec), AppError> { + // Text changes: proven edited copies of their sources replace them before assembly. + let opts = text_edit::engines::RunOpts { + handle, + cancel, + ..Default::default() + }; + let text = text_edit::export::prepare_save( + document, groups, work, qpdf_check, app, unique, &opts, + )?; + let content_groups = text.as_ref().map_or(groups, |(_, p)| &p.content_groups[..]); let extra: Vec<&str> = extra_group_paths(groups); extra_files_have_fields(&extra)?; - let (geoms, counts) = collect_source_pages(groups)?; + let has_paint = document.objects.iter().any(|o| o.triggers_overlay()); + let (geoms, counts) = collect_source_pages(content_groups, has_paint)?; if geoms.is_empty() { return Err(AppError::new("NO_PAGES", "No pages", "Add a PDF first.")); } let links = session_links_from_doc(document); let redacts = redact_regions_from_doc(document); let has_redact = !redacts.is_empty(); - let has_paint = document.objects.iter().any(|o| o.triggers_overlay()); let mut redact_probes: Vec> = Vec::new(); let mut flatten_form_done = false; let mut flatten_annots_done = false; if has_redact { - assemble_to_tmp(groups, &counts, &tmp, &tmp_str, &mut run)?; + assemble_to_tmp(content_groups, &counts, &tmp, &tmp_str, &mut run)?; // Burn flattened /AP into page content *before* rasterize. A // flatten after apply_redactions would paint leftover field text // on top of /ImR. @@ -1776,7 +1822,7 @@ where .map_err(|e| AppError::io("Could not read the editor font.", e))?; let font = FontInfo::parse(font_bytes)?; write_overlay_pdf(&overlay_str, &geoms, document, &font, cancel)?; - let (mapped, restore_boxes) = remap_groups_to_visible_box(groups, work)?; + let (mapped, restore_boxes) = remap_groups_to_visible_box(content_groups, work)?; let args = build_edit_overlay_args(&mapped, &counts, &overlay_str, &tmp_str)?; run(&args)?; if restore_boxes { @@ -1788,7 +1834,7 @@ where safe_output::replace_file(&cleaned, &tmp)?; } } else { - assemble_to_tmp(groups, &counts, &tmp, &tmp_str, &mut run)?; + assemble_to_tmp(content_groups, &counts, &tmp, &tmp_str, &mut run)?; } let expected_annots = expected_dest_has_annots(&tmp, &links)?; if !incomplete_source_paths.is_empty() && !links.is_empty() { @@ -1806,13 +1852,17 @@ where .flatten() .collect(); // Skip dest_has load when no complete source remains (e.g. 400 MiB - // unlistable file). L7: a complete empty edit still deletes supported - // links. Stamp/markup/form-only saves preserve leftover links, while an - // annotation-flatten-only save leaves them for qpdf to copy through. + // unlistable file). L7: a complete empty edit (text changes alone count as + // empty) still deletes supported links. Stamp/markup/form-only saves preserve + // leftover links, while an annotation-flatten-only save leaves them for qpdf to + // copy through. let rewrite_links = !dest_pages.is_empty() && (!links.is_empty() || (!flatten_annotations - && document.objects.is_empty() + && document + .objects + .iter() + .all(|o| matches!(o, EditObjectIn::SourceText { .. })) && form_values.is_empty() && !flatten_form && dest_has_supported_links(&tmp)?)); @@ -1854,9 +1904,25 @@ where warnings.extend(verify_redaction(&tmp, &probe_refs, &redacts)?); update_redacted_digests(&tmp, &mut snapshot, &redacts)?; } - let vr = validate_staged_pdf(&tmp, &snapshot, cancel, |args| { - run_qpdf_check_argv(qpdf_check, args, handle) - })?; + if let Some((engines, prepared)) = &text { + let digests: Vec = + snapshot.pages.iter().map(|p| p.content_digest).collect(); + if !prepared.proofs.is_empty() { + text_edit::export::verify_final(&tmp, prepared, &digests, engines, &opts)?; + } + warnings.extend(prepared.warnings.iter().cloned()); + } + // A redacted page's content is new: only its re-read digest is expected. + let redacted: HashSet = redacts.iter().map(|r| r.page_index).collect(); + let alts: Vec> = (0u32..) + .zip(&geoms) + .map(|(i, g)| g.alt_content_digest.filter(|_| !redacted.contains(&i))) + .collect(); + let check = |args: &[String]| run_qpdf_check_argv(qpdf_check, args, handle); + let vr = match alts.iter().any(Option::is_some) { + true => validate_staged_pdf_with_alternatives(&tmp, &snapshot, &alts, cancel, check)?, + false => validate_staged_pdf(&tmp, &snapshot, cancel, check)?, + }; warnings.extend(vr.warnings); gate_passed = true; safe_output::replace_file(&tmp, dest)?; @@ -2190,9 +2256,7 @@ fn write_overlay_pdf( if obj.page_index() as usize != pi { continue; } - if matches!(obj, EditObjectIn::Link { .. } | EditObjectIn::Redact { .. }) - || obj.is_markup() - { + if !obj.triggers_overlay() || obj.is_markup() { continue; } let op100 = (obj.opacity() * 100.0).round() as i32; @@ -2430,7 +2494,8 @@ fn write_overlay_pdf( | EditObjectIn::Underline { .. } | EditObjectIn::Strikeout { .. } | EditObjectIn::MarkupInk { .. } - | EditObjectIn::Redact { .. } => {} + | EditObjectIn::Redact { .. } + | EditObjectIn::SourceText { .. } => {} } if rotated { content.push_str("Q\n"); @@ -2851,6 +2916,7 @@ mod tests { rotate: 0, user_unit: 1.0, content_digest: content_digest(b"BT /F1 12 Tf 72 720 Td (Hello) Tj ET"), + alt_content_digest: None, } } @@ -3301,4 +3367,23 @@ mod tests { assert_eq!(std::fs::read(&src).unwrap(), before); let _ = std::fs::remove_dir_all(&root); } + + /// review-T5 L3: qpdf's join digest (#34's alternative) is computed only when an overlay will + /// wrap the pages; other saves expect the plain digest alone. + #[test] + fn alternative_digests_only_for_overlay_saves() { + use crate::pdf_engine::text_edit::testkit::producers::{DocBuilder, PageSpec}; + let mut d = DocBuilder::new(); + d.page(PageSpec::parts(&[b"q Q", b"q Q"], "")); + let root = std::env::temp_dir().join(format!("offpdf-alt-{}", std::process::id())); + std::fs::create_dir_all(&root).unwrap(); + let src = root.join("split.pdf"); + std::fs::write(&src, d.build()).unwrap(); + let groups = [g(src.to_str().unwrap(), "1")]; + let (overlay, _) = collect_source_pages(&groups, true).unwrap(); + let (plain, _) = collect_source_pages(&groups, false).unwrap(); + assert!(overlay[0].alt_content_digest.is_some(), "overlay"); + assert!(plain[0].alt_content_digest.is_none(), "no overlay"); + let _ = std::fs::remove_dir_all(&root); + } } diff --git a/src-tauri/src/pdf_engine/mod.rs b/src-tauri/src/pdf_engine/mod.rs index 3fc5b4f..aa65442 100644 --- a/src-tauri/src/pdf_engine/mod.rs +++ b/src-tauri/src/pdf_engine/mod.rs @@ -42,6 +42,7 @@ mod source_edit_fixtures; #[cfg(test)] mod source_content_integ; pub mod stamp; +pub mod text_edit; pub mod textexport; pub mod validate_output; diff --git a/src-tauri/src/pdf_engine/render.rs b/src-tauri/src/pdf_engine/render.rs index decb510..19c032a 100644 --- a/src-tauri/src/pdf_engine/render.rs +++ b/src-tauri/src/pdf_engine/render.rs @@ -216,9 +216,74 @@ pub fn page_texts(app: &tauri::AppHandle, input: &str) -> Result, Ap /// giant raster page) returns None so the viewer falls back to raster preview. const MAX_PAGE_PDF_BYTES: u64 = 48 * 1024 * 1024; +/// Bytes hashed at each end of a file for its page-cache version. +const PAGE_CACHE_HEAD_TAIL: u64 = 64 * 1024; + +/// FNV of the first and last `PAGE_CACHE_HEAD_TAIL` bytes of `input` (empty when unreadable). +fn head_tail_hex(input: &str, len: u64) -> String { + use std::io::{Read, Seek, SeekFrom}; + let mut bytes = Vec::new(); + if let Ok(mut f) = std::fs::File::open(input) { + let _ = f + .by_ref() + .take(PAGE_CACHE_HEAD_TAIL) + .read_to_end(&mut bytes); + let tail = len.saturating_sub(PAGE_CACHE_HEAD_TAIL); + if f.seek(SeekFrom::Start(tail)).is_ok() { + let _ = f.take(PAGE_CACHE_HEAD_TAIL).read_to_end(&mut bytes); + } + } + let hash = bytes.iter().fold(0xcbf29ce484222325u64, |h, b| { + (h ^ u64::from(*b)).wrapping_mul(0x100000001b3) + }); + format!("{hash:016x}") +} + +/// The page-cache folders of `input`: one per path, holding one version folder keyed by the +/// file's length, modification time and first/last 64 KiB, so a file changed on disk (even with +/// the same size and mtime at the ends) never serves a page extracted from its old bytes. +fn page_cache_key(input: &str) -> (String, String) { + let meta = std::fs::metadata(input).ok(); + let len = meta.as_ref().map_or(0, std::fs::Metadata::len); + let mtime = meta + .and_then(|m| m.modified().ok()) + .and_then(|t| t.duration_since(std::time::UNIX_EPOCH).ok()) + .map_or(0, |d| d.as_nanos()); + let ends = head_tail_hex(input, len); + ( + fnv1a_hex(input), + fnv1a_hex(&format!("{len}\u{0}{mtime}\u{0}{ends}")), + ) +} + +/// `///`, created; older versions of the same path (and the +/// v0.3 pages stored directly in the path folder) are deleted. +fn page_cache_dir(pagepdf: &Path, input: &str) -> Result { + let (path_key, version) = page_cache_key(input); + let base = pagepdf.join(path_key); + let dir = base.join(&version); + if !dir.is_dir() { + for entry in std::fs::read_dir(&base).into_iter().flatten().flatten() { + let stale = entry.path(); + match entry.file_type() { + Ok(t) if t.is_dir() => { + let _ = std::fs::remove_dir_all(&stale); + } + _ => { + let _ = std::fs::remove_file(&stale); + } + } + } + } + std::fs::create_dir_all(&dir) + .map_err(|e| AppError::io("Could not create the page cache.", e))?; + Ok(dir) +} + /// Extract one page into a tiny standalone PDF and return it base64-encoded, for /// true-vector rendering with pdf.js in the webview. Returns `None` if the page -/// is too large to safely load into the webview. Cached on disk per (file,page). +/// is too large to safely load into the webview. Cached on disk per (file, length, +/// modification time, first/last 64 KiB, page). pub fn page_pdf_b64( app: &tauri::AppHandle, input: &str, @@ -227,9 +292,7 @@ pub fn page_pdf_b64( if !Path::new(input).is_file() { return Err(AppError::invalid_pdf(input)); } - let dir = temp::root(app)?.join("pagepdf").join(fnv1a_hex(input)); - std::fs::create_dir_all(&dir) - .map_err(|e| AppError::io("Could not create the page cache.", e))?; + let dir = page_cache_dir(&temp::root(app)?.join("pagepdf"), input)?; let out = dir.join(format!("p{page}.pdf")); let out_str = out.to_string_lossy().to_string(); @@ -590,3 +653,76 @@ pub(crate) fn fnv1a_hex(s: &str) -> String { } format!("{hash:016x}") } + +#[cfg(test)] +mod tests { + use super::{page_cache_dir, page_cache_key}; + + /// The one-page cache is keyed by the file's length and mtime, not only its path (stale + /// page previews after the file changed on disk). + #[test] + fn page_cache_key_changes_when_the_file_changes() { + let dir = std::env::temp_dir().join(format!("offpdf-pagekey-{}", std::process::id())); + std::fs::create_dir_all(&dir).expect("dir"); + let path = dir.join("doc.pdf"); + let p = path.to_string_lossy().into_owned(); + std::fs::write(&path, b"%PDF-1.7 one").expect("write"); + let first = page_cache_key(&p); + assert_eq!(first, page_cache_key(&p), "stable while unchanged"); + std::fs::write(&path, b"%PDF-1.7 longer").expect("rewrite"); + let second = page_cache_key(&p); + assert_ne!(first, second, "length change"); + let t = std::time::SystemTime::UNIX_EPOCH + std::time::Duration::from_secs(1_000_000); + std::fs::File::options() + .write(true) + .open(&path) + .and_then(|f| f.set_modified(t)) + .expect("mtime"); + assert_ne!(second, page_cache_key(&p), "mtime change"); + let _ = std::fs::remove_dir_all(&dir); + } + + /// review-T5 L1: a same-size rewrite that keeps the mtime still gets a new version, and the + /// old version (and v0.3's pages stored directly in the path folder) are deleted. + #[test] + fn page_cache_versions_follow_the_bytes_and_replace_older_ones() { + let dir = std::env::temp_dir().join(format!("offpdf-pagever-{}", std::process::id())); + let _ = std::fs::remove_dir_all(&dir); + std::fs::create_dir_all(&dir).expect("dir"); + let path = dir.join("doc.pdf"); + let p = path.to_string_lossy().into_owned(); + let t = std::time::SystemTime::UNIX_EPOCH + std::time::Duration::from_secs(2_000_000); + let write = |bytes: &[u8]| { + std::fs::write(&path, bytes).expect("write"); + std::fs::File::options() + .write(true) + .open(&path) + .and_then(|f| f.set_modified(t)) + .expect("mtime"); + }; + write(b"%PDF-1.7 aaaa"); + let first = page_cache_key(&p); + write(b"%PDF-1.7 bbbb"); + let second = page_cache_key(&p); + assert_eq!(first.0, second.0, "one folder per path"); + assert_ne!(first.1, second.1, "same size and mtime, other bytes"); + let pagepdf = dir.join("pagepdf"); + let base = pagepdf.join(&second.0); + std::fs::create_dir_all(base.join(&first.1)).expect("old version"); + std::fs::write(base.join(&first.1).join("p1.pdf"), b"old").expect("old page"); + std::fs::write(base.join("p1.pdf"), b"v0.3 page").expect("v0.3 page"); + let got = page_cache_dir(&pagepdf, &p).expect("cache dir"); + assert_eq!(got, base.join(&second.1)); + let left: Vec<_> = std::fs::read_dir(&base) + .expect("base") + .flatten() + .map(|e| e.file_name()) + .collect(); + assert_eq!( + left, + vec![std::ffi::OsString::from(&second.1)], + "only the current version" + ); + let _ = std::fs::remove_dir_all(&dir); + } +} diff --git a/src-tauri/src/pdf_engine/source_content.rs b/src-tauri/src/pdf_engine/source_content.rs index 34c7da7..35401f4 100644 --- a/src-tauri/src/pdf_engine/source_content.rs +++ b/src-tauri/src/pdf_engine/source_content.rs @@ -1,44 +1,71 @@ -//! Read-only source-content classifier (#33). +//! offpdf:forbidden-api-scan +//! Hardened read-only classifier (#33): bounded, fail-closed; the text half is wired to Edit +//! text through `inspect_text_page`. //! -//! Walks page streams and Form XObjects on the original source path. Does not -//! write a dest, call `Document::replace_text`, or apply page `/Rotate`. +//! A thin adapter over `text_edit` (SPEC §C). The snapshot is one capped read, preflighted and +//! parsed from the bytes it hashed (`snapshot.rs`); the page model (`runs::build_page_model`) +//! brings strict geometry, the bounded lexer and decoders, the font layer, runs and their reason +//! codes; one more walk in `Classify` mode descends Form XObjects and lists image paints. Per page +//! it returns the run capabilities Edit text shows (`runs`) and the #33 occurrences — one per +//! show op and per image paint — with v2 locators: //! -//! Research prototype only; no Tauri command or UI invokes this module. -//! Decoded-size checks currently run after allocation, and font/geometry -//! support is incomplete. Complete #33's resource bounds and compatibility -//! evaluation before exposing this API to user files or enabling editing. - -use crate::error::AppError; -use crate::pdf_engine::crop; -use lopdf::{content::Content, Dictionary, Document, Object, ObjectId, Stream}; +//! `v2:{fingerprint}:{page}:{t|i}:{path}:{ordinal}`, where `path` is `p{start}-{end}` (the op's +//! span in the page's joined content) or `x{obj}.{gen}>…:{start}-{end}` (the Form chain and the +//! span in the innermost Form's data), and `ordinal` is the occurrence's index on the page. +//! +//! A text occurrence carries the reason of the run that contains its record (Form text: +//! `NESTED_FORM`; a page refused `GEOMETRY`: `GEOMETRY`). Image occurrences use the image +//! priority of §A.10. `Supported` is never a Save affordance: Save re-plans and re-verifies from +//! the file. +//! +//! The Classify pass runs on the page-model budget its model left (`walker::budget`, +//! `PAGE_MODEL_BYTES_MAX`: the model's bytes are charged already, so the model and its Classify +//! pass hold no more than one page budget together, review T3-budget MEDIUM-2): past it the page +//! lists no occurrence (`occurrence_reason` `PAGE_TOO_COMPLEX`), never a partial list. A page that +//! paints no Form XObject is not walked again: its Edit walk is exactly what a Classify walk would +//! record, so the occurrences are built from the model's own walk. + +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::geometry::{contains, display_rotation, mul, Matrix}; +use crate::pdf_engine::text_edit::limits::{AXIS_EPSILON_REL, CLIP_CONTAIN_TOL_PT}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::runs::reasons::ReasonCtx; +use crate::pdf_engine::text_edit::runs::{form_lines, PageModel, TextRun}; +use crate::pdf_engine::text_edit::state::ClipState; +use crate::pdf_engine::text_edit::walker::budget::{map_entry, ModelBudget, MODEL_SIZE}; +use crate::pdf_engine::text_edit::walker::{ + walk_page_within, PageWalk, PaintKind, PaintRecord, ShowRecord, Stop, WalkMode, +}; +use lopdf::ObjectId; +use serde::Serialize; use std::collections::HashMap; -use std::path::Path; - -const FILE_CAP_BYTES: u64 = 400 * 1024 * 1024; -const MAX_FORM_DEPTH: usize = 8; -const MAX_STREAM_BYTES: usize = 32 * 1024 * 1024; -const MAX_DECODED_TOTAL: usize = 64 * 1024 * 1024; -const MAX_OPS: usize = 50_000; -const MAX_OCCURRENCES: usize = 5_000; -const MAX_GSTATE_STACK: usize = 64; - -const IDENTITY: [f64; 6] = [1.0, 0.0, 0.0, 1.0, 0.0, 0.0]; -const FNV_OFFSET: u64 = 0xcbf29ce484222325; -const FNV_PRIME: u64 = 0x100000001b3; - -#[derive(Debug, Clone, Copy, PartialEq, Eq)] +use std::mem::size_of; +use std::sync::atomic::{AtomicBool, Ordering}; + +/// Bytes of a locator besides its path: `v2:`, the fingerprint, the page, the kind and the +/// ordinal with their separators. +const LOCATOR_FIXED_BYTES: usize = 96; +/// Bytes of one Form in a path (`x{obj}.{gen}>`) and of the span (`:{start}-{end}`). +const PATH_FORM_BYTES: usize = 18; +const PATH_SPAN_BYTES: usize = 44; + +#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize)] +#[serde(rename_all = "camelCase")] pub enum SourceKind { Text, Image, } -#[derive(Debug, Clone, Copy, PartialEq, Eq)] +#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize)] +#[serde(rename_all = "camelCase")] pub enum SourceCapability { Supported, Unsupported, } -#[derive(Debug, Clone, PartialEq)] +/// Unrotated user space: text = baseline origin (rise included), `w` = the real advance (Tc, Tw +/// and kerns included), `h` = the effective size; images = the painted unit square's AABB. +#[derive(Debug, Clone, PartialEq, Serialize)] pub struct SourceRect { pub x: f64, pub y: f64, @@ -46,1688 +73,468 @@ pub struct SourceRect { pub h: f64, } -#[derive(Debug, Clone, PartialEq)] +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] pub struct SourceOccurrence { pub page_index: u32, pub kind: SourceKind, pub rect: SourceRect, pub locator: String, pub capability: SourceCapability, - pub reason: Option, -} - -pub fn classify_source_content(path: &Path) -> Result, AppError> { - let (doc, fp) = open_source(path)?; - classify_doc(&doc, fp) -} - -pub fn resolve_source_locator( - path: &Path, - locator: &str, -) -> Result { - if !path.is_file() { - return Err(AppError::invalid_pdf(&path_str(path))); - } - let meta = std::fs::metadata(path).map_err(|_| AppError::invalid_pdf(&path_str(path)))?; - if meta.len() > FILE_CAP_BYTES { - return Err(file_too_large()); - } - let bytes = std::fs::read(path).map_err(|_| AppError::invalid_pdf(&path_str(path)))?; - let fp = fnv1a_u64(&bytes); - let loc_fp = parse_locator_fp(locator).ok_or_else(stale)?; - if loc_fp != fp { - return Err(stale()); - } - classify_source_content(path)? - .into_iter() - .find(|o| o.locator == locator) - .ok_or_else(stale) -} - -fn open_source(path: &Path) -> Result<(Document, u64), AppError> { - if !path.is_file() { - return Err(AppError::invalid_pdf(&path_str(path))); - } - let meta = std::fs::metadata(path).map_err(|_| AppError::invalid_pdf(&path_str(path)))?; - if meta.len() > FILE_CAP_BYTES { - return Err(file_too_large()); - } - let bytes = std::fs::read(path).map_err(|_| AppError::invalid_pdf(&path_str(path)))?; - let fp = fnv1a_u64(&bytes); - let doc = Document::load(path).map_err(|e| { - AppError::invalid_pdf(&path_str(path)).with_details(format!("lopdf: {e}")) - })?; - if doc.is_encrypted() { - return Err(encrypted()); - } - if document_is_signed(&doc) { - return Err(signed()); - } - if doc.catalog().is_err() { - return Err(malformed_content("The PDF catalog is missing or unreadable.")); - } - Ok((doc, fp)) -} - -fn classify_doc(doc: &Document, fp: u64) -> Result, AppError> { - let mut walker = Walker { - doc, - fp, - decoded_total: 0, - op_count: 0, - pending: Vec::new(), - paint_counts: HashMap::new(), - }; - let pages = doc.get_pages(); - let mut nums: Vec = pages.keys().copied().collect(); - nums.sort_unstable(); - for num in nums { - let Some(&page_id) = pages.get(&num) else { - continue; - }; - walker.walk_page(page_id, num.saturating_sub(1))?; - } - Ok(walker.finish()) -} - -struct Walker<'a> { - doc: &'a Document, - fp: u64, - decoded_total: usize, - op_count: usize, - pending: Vec, - paint_counts: HashMap, -} - -#[derive(Clone, Copy, Default)] -struct Flags { - inline_image: bool, - nested_form: bool, - type3: bool, - vertical: bool, - clipped: bool, - pattern: bool, - rotated: bool, - skewed: bool, - no_tounicode: bool, - missing_font: bool, - ambiguous: bool, - masked: bool, - shared: bool, - geometry: bool, -} - -struct Pending { - page_index: u32, - kind: SourceKind, - rect: SourceRect, - locator: String, - flags: Flags, - paint_id: Option, -} - -#[derive(Clone)] -struct GState { - ctm: [f64; 6], - clip_active: bool, - clip_pending: bool, - fill_pattern: bool, - stroke_pattern: bool, - masked: bool, - font_name: Option>, - font_size: f64, - leading: f64, - hscale: f64, - tc: f64, - tw: f64, -} - -impl Default for GState { - fn default() -> Self { - Self { - ctm: IDENTITY, - clip_active: false, - clip_pending: false, - fill_pattern: false, - stroke_pattern: false, - masked: false, - font_name: None, - font_size: 0.0, - leading: 0.0, - hscale: 100.0, - tc: 0.0, - tw: 0.0, - } - } -} - -impl GState { - fn pattern(&self) -> bool { - self.fill_pattern || self.stroke_pattern - } - - fn clipped(&self) -> bool { - self.clip_active || self.clip_pending - } -} - -struct TextState { - tm: [f64; 6], - tlm: [f64; 6], + pub reason: Option, + /// The decoded text of a text occurrence (None when a glyph cannot be decoded, or an image). + pub text: Option, } -impl Default for TextState { - fn default() -> Self { - Self { - tm: IDENTITY, - tlm: IDENTITY, - } - } -} - -struct FormInfo { - matrix: [f64; 6], - bytes: Vec, - has_resources: bool, -} - -struct TextInspect { - missing_font: bool, - type3: bool, - vertical: bool, - no_tounicode: bool, - ambiguous: bool, - width: f64, - font_id: ObjectId, +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct SourceRunCapability { + pub run_id: String, + pub capability: SourceCapability, + pub reason: Option, } -impl Walker<'_> { - fn walk_page(&mut self, page_id: ObjectId, page_index: u32) -> Result<(), AppError> { - let geom_unsafe = page_geom_unsafe(self.doc, page_id); - let owners = page_resource_owners(self.doc, page_id); - let ids = self.doc.get_page_contents(page_id); - if ids.is_empty() { - return Ok(()); - } - let contents_id = ids[0]; - let mut bytes = Vec::new(); - for id in &ids { - let stream = self - .doc - .get_object(*id) - .ok() - .and_then(|o| o.as_stream().ok()) - .ok_or_else(|| malformed_content("A page content stream is unreadable."))?; - let chunk = decompress_stream(stream)?; - self.add_decoded(chunk.len())?; - if !bytes.is_empty() { - bytes.push(b'\n'); - } - bytes.extend_from_slice(&chunk); - } - let mut visiting = Vec::new(); - self.walk_stream( - page_index, - contents_id, - &bytes, - &owners, - GState::default(), - 0, - &mut visiting, - geom_unsafe, - ) - } - - fn walk_stream( - &mut self, - page_index: u32, - contents_id: ObjectId, - bytes: &[u8], - resource_owners: &[ObjectId], - mut gs: GState, - form_depth: usize, - visiting: &mut Vec, - geom_unsafe: bool, - ) -> Result<(), AppError> { - let inline_total = count_inline_images(bytes)?; - // BI…EI: do not trust a prefix Content::decode Ok; strip payloads first. - let stripped = if inline_total > 0 { - Some(strip_inline_image_payloads(bytes)?) - } else { - None - }; - let decode_src = stripped.as_deref().unwrap_or(bytes); - let ops = match Content::decode(decode_src) { - Ok(c) => c.operations, - Err(_) => { - return Err(malformed_content( - "A page content stream could not be decoded.", - )); - } - }; - self.bump_ops(ops.len())?; - - let mut ts = TextState::default(); - let mut stack: Vec = Vec::new(); - let mut inline_emitted = 0usize; - let nested = form_depth > 0; - - for (op_index, op) in ops.iter().enumerate() { - match op.operator.as_str() { - "q" => { - if stack.len() < MAX_GSTATE_STACK { - stack.push(gs.clone()); - } - } - "Q" => { - if let Some(prev) = stack.pop() { - gs = prev; - } - } - "cm" => { - if let Some(m) = six_nums(&op.operands) { - gs.ctm = mul(gs.ctm, m); - } - } - "W" | "W*" => { - gs.clip_pending = true; - } - "n" => { - if gs.clip_pending { - gs.clip_active = true; - gs.clip_pending = false; - } - } - "S" | "s" | "f" | "F" | "f*" | "B" | "B*" | "b" | "b*" => { - if gs.clip_pending { - gs.clip_active = true; - gs.clip_pending = false; - } - } - "cs" => { - gs.fill_pattern = - operand_selects_pattern_space(self.doc, resource_owners, &op.operands); - } - "CS" => { - gs.stroke_pattern = - operand_selects_pattern_space(self.doc, resource_owners, &op.operands); - } - "scn" | "sc" => { - if gs.fill_pattern || is_pattern_name(&op.operands) { - gs.fill_pattern = true; - } - } - "SCN" | "SC" => { - if gs.stroke_pattern || is_pattern_name(&op.operands) { - gs.stroke_pattern = true; - } - } - "gs" => { - if let Some(name) = op.operands.first().and_then(|o| o.as_name().ok()) { - apply_extgstate_smask(self.doc, resource_owners, name, &mut gs); - } - } - "rg" | "g" | "k" => { - gs.fill_pattern = false; - } - "RG" | "G" | "K" => { - gs.stroke_pattern = false; - } - "BT" => { - ts.tm = IDENTITY; - ts.tlm = IDENTITY; - } - "ET" => {} - "Tf" => { - if let Some(name) = op.operands.first().and_then(|o| o.as_name().ok()) { - gs.font_name = Some(name.to_vec()); - } - if let Some(size) = op.operands.get(1).and_then(obj_f64) { - gs.font_size = size; - } - } - "Tc" => { - if let Some(c) = op.operands.first().and_then(obj_f64) { - gs.tc = c; - } - } - "Tw" => { - if let Some(w) = op.operands.first().and_then(obj_f64) { - gs.tw = w; - } - } - "Tz" => { - if let Some(z) = op.operands.first().and_then(obj_f64) { - gs.hscale = z; - } - } - "TL" => { - if let Some(l) = op.operands.first().and_then(obj_f64) { - gs.leading = l; - } - } - "Td" => { - if let Some((tx, ty)) = two_nums(&op.operands) { - apply_td(&mut ts, tx, ty); - } - } - "TD" => { - if let Some((tx, ty)) = two_nums(&op.operands) { - gs.leading = -ty; - apply_td(&mut ts, tx, ty); - } - } - "Tm" => { - if let Some(m) = six_nums(&op.operands) { - ts.tm = m; - ts.tlm = m; - } - } - "T*" => { - let ty = -gs.leading; - apply_td(&mut ts, 0.0, ty); - } - "Tj" | "'" | "\"" | "TJ" => { - let mut show_ops = op.operands.as_slice(); - if op.operator == "\"" { - if let Some(aw) = op.operands.first().and_then(obj_f64) { - gs.tw = aw; - } - if let Some(ac) = op.operands.get(1).and_then(obj_f64) { - gs.tc = ac; - } - if op.operands.len() >= 2 - && obj_f64(&op.operands[0]).is_some() - && obj_f64(&op.operands[1]).is_some() - { - show_ops = &op.operands[2..]; - } - } - if op.operator == "'" || op.operator == "\"" { - let ty = -gs.leading; - apply_td(&mut ts, 0.0, ty); - } - let (pieces, adj) = text_pieces(show_ops); - self.emit_text( - page_index, - contents_id, - op_index as u32, - resource_owners, - &gs, - &mut ts, - &pieces, - adj, - nested, - geom_unsafe, - )?; - } - "Do" => { - let Some(name) = op.operands.first().and_then(|o| o.as_name().ok()) else { - continue; - }; - let Some(id) = lookup_xobject(self.doc, resource_owners, name) else { - continue; - }; - if xobject_is_form(self.doc, id) { - self.enter_form( - page_index, - id, - resource_owners, - &gs, - form_depth, - visiting, - geom_unsafe, - )?; - } else if xobject_is_image(self.doc, id) { - self.emit_image( - page_index, - contents_id, - op_index as u32, - id, - &gs, - nested, - geom_unsafe, - )?; - } - } - "BI" => { - self.emit_inline( - page_index, - contents_id, - op_index as u32, - &gs, - nested, - geom_unsafe, - )?; - inline_emitted += 1; - } - "EI" | "ID" => {} - _ => {} - } - } - - while inline_emitted < inline_total { - self.emit_inline( - page_index, - contents_id, - ops.len() as u32 + inline_emitted as u32, - &gs, - nested, - geom_unsafe, - )?; - inline_emitted += 1; - } - Ok(()) - } - - fn enter_form( - &mut self, - page_index: u32, - form_id: ObjectId, - resource_owners: &[ObjectId], - gs: &GState, - form_depth: usize, - visiting: &mut Vec, - geom_unsafe: bool, - ) -> Result<(), AppError> { - if form_depth >= MAX_FORM_DEPTH { - return Err(malformed_content( - "Form XObject nesting is deeper than 8.", - )); - } - if visiting.contains(&form_id) { - return Err(malformed_content("A Form XObject refers to itself.")); - } - let info = load_form(self.doc, form_id)?; - self.add_decoded(info.bytes.len())?; - let mut child_gs = gs.clone(); - child_gs.ctm = mul(gs.ctm, info.matrix); - visiting.push(form_id); - let mut child_owners = Vec::new(); - if info.has_resources { - child_owners.push(form_id); - } - child_owners.extend_from_slice(resource_owners); - let result = self.walk_stream( - page_index, - form_id, - &info.bytes, - &child_owners, - child_gs, - form_depth + 1, - visiting, - geom_unsafe, - ); - visiting.pop(); - result - } - - fn emit_text( - &mut self, - page_index: u32, - contents_id: ObjectId, - op_index: u32, - resource_owners: &[ObjectId], - gs: &GState, - ts: &mut TextState, - pieces: &[Vec], - tj_adj: f64, - nested: bool, - geom_unsafe: bool, - ) -> Result<(), AppError> { - let inspect = inspect_text( - self.doc, - resource_owners, - gs.font_name.as_deref(), - pieces, - tj_adj, - gs.font_size, - ); - let effective = mul(gs.ctm, ts.tm); - let sx = (effective[0] * effective[0] + effective[1] * effective[1]).sqrt(); - let sy = (effective[2] * effective[2] + effective[3] * effective[3]).sqrt(); - let height = (gs.font_size.abs() * sy).max(0.01); - let th = gs.hscale / 100.0; - let shown = inspect.width * th; - let width = (shown * sx).abs().max(0.01); - let tx = (inspect.width + spacing_advance(pieces, gs.tc, gs.tw)) * th; - let flags = Flags { - nested_form: nested, - type3: inspect.type3, - vertical: inspect.vertical, - clipped: gs.clipped(), - pattern: gs.pattern(), - rotated: is_rotated_tm(effective), - skewed: is_skewed_tm(effective), - no_tounicode: inspect.no_tounicode, - missing_font: inspect.missing_font, - ambiguous: inspect.ambiguous, - masked: gs.masked, - geometry: geom_unsafe, - ..Flags::default() - }; - self.push(Pending { - page_index, - kind: SourceKind::Text, - rect: SourceRect { - x: effective[4], - y: effective[5], - w: width, - h: height, - }, - locator: encode_locator( - self.fp, - page_index, - SourceKind::Text, - contents_id, - op_index, - inspect.font_id, - ), - flags, - paint_id: None, - })?; - // Shown-width advance is Tm only; Tlm stays at the line origin. - ts.tm = mul(ts.tm, [1.0, 0.0, 0.0, 1.0, tx, 0.0]); - Ok(()) - } - - fn emit_image( - &mut self, - page_index: u32, - contents_id: ObjectId, - op_index: u32, - image_id: ObjectId, - gs: &GState, - nested: bool, - geom_unsafe: bool, - ) -> Result<(), AppError> { - let flags = Flags { - nested_form: nested, - clipped: gs.clipped(), - pattern: gs.pattern(), - masked: gs.masked || image_is_masked(self.doc, image_id), - geometry: geom_unsafe, - ..Flags::default() - }; - self.push(Pending { - page_index, - kind: SourceKind::Image, - rect: unit_square_bbox(gs.ctm), - locator: encode_locator( - self.fp, - page_index, - SourceKind::Image, - contents_id, - op_index, - image_id, - ), - flags, - paint_id: Some(image_id), - }) - } - - fn emit_inline( - &mut self, - page_index: u32, - contents_id: ObjectId, - op_index: u32, - gs: &GState, - nested: bool, - geom_unsafe: bool, - ) -> Result<(), AppError> { - let flags = Flags { - inline_image: true, - nested_form: nested, - clipped: gs.clipped(), - pattern: gs.pattern(), - masked: gs.masked, - geometry: geom_unsafe, - ..Flags::default() - }; - self.push(Pending { - page_index, - kind: SourceKind::Image, - rect: unit_square_bbox(gs.ctm), - locator: encode_locator( - self.fp, - page_index, - SourceKind::Image, - contents_id, - op_index, - (0, 0), - ), - flags, - paint_id: None, +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct SourcePageResult { + pub page_index: u32, + /// The page-level refusal of Edit text (every run refused with it; runs are not listed). + pub page_reason: Option, + /// One per model run, in reading order — drives the run DTO's `editable`/`reason`. + pub runs: Vec, + /// Per show op / image paint (the #33 contract; image rows are not shown in v0.4). + pub occurrences: Vec, + /// The `Classify` walk (Form XObjects included) was refused at page level although Edit + /// text's own depth-0 walk was not: no occurrence is listed (never a partial list). + pub occurrence_reason: Option, + /// Text drawn through Form XObjects, as refused `NESTED_FORM` lines after `runs` in reading + /// order (§A.6 "still shown"; the page model has no run for it). Empty when the page paints no + /// Form, is refused, or its Classify pass is. + pub form_lines: Vec, +} + +/// One refused line of Form text, with the geometry a run DTO shows (unrotated user space, as +/// `TextRun`). +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct SourceFormLine { + /// `t1:{fp}:{page}:x{obj}.{gen}>…:{s0}-{e0}[,…]` (Form chain and spans in its data). + pub id: String, + pub order: u32, + pub line: u32, + pub text: String, + /// `[x, y, w, h]` ink AABB. + pub rect: [f64; 4], + pub origin: (f64, f64), + pub dir: (f64, f64), + pub ascent: f64, + pub descent: f64, + /// chars + 1 distances along `dir` from `origin`. + pub caret_offsets: Vec, + /// Always `NESTED_FORM`. + pub reason: TextReason, + pub substituted: bool, +} + +impl SourceFormLine { + fn of(run: TextRun) -> SourceFormLine { + SourceFormLine { + id: run.id, + order: run.order, + line: run.line, + text: run.text, + rect: run.rect, + origin: run.origin, + dir: run.dir, + ascent: run.ascent, + descent: run.descent, + caret_offsets: run.caret_offsets, + reason: run.reason.unwrap_or(TextReason::NestedForm), + substituted: run.substituted, + } + } +} + +fn capability(reason: Option) -> SourceCapability { + match reason { + Some(_) => SourceCapability::Unsupported, + None => SourceCapability::Supported, + } +} + +/// The #33 classification of one page of `ctx` from its Edit-text model. `cancel` stops the +/// `Classify` walk (it then lists no occurrence: `occurrence_reason` `PAGE_TOO_COMPLEX`). +pub fn classify_source_page( + ctx: &SnapshotContext, + model: &PageModel, + cancel: Option<&AtomicBool>, +) -> SourcePageResult { + let runs: Vec = model + .runs + .iter() + .map(|r| SourceRunCapability { + run_id: r.id.clone(), + capability: capability(r.reason), + reason: r.reason, }) - } - - fn add_decoded(&mut self, n: usize) -> Result<(), AppError> { - self.decoded_total = self.decoded_total.saturating_add(n); - if self.decoded_total > MAX_DECODED_TOTAL { - return Err(malformed_content("Decoded content exceeds 64 MB.")); - } - Ok(()) - } - - fn bump_ops(&mut self, n: usize) -> Result<(), AppError> { - self.op_count = self.op_count.saturating_add(n); - if self.op_count > MAX_OPS { - return Err(malformed_content( - "The page content has more than 50,000 operators.", - )); - } - Ok(()) - } - - fn push(&mut self, mut pending: Pending) -> Result<(), AppError> { - if self.pending.len() >= MAX_OCCURRENCES { - return Err(malformed_content( - "This PDF has more than 5,000 text or image occurrences.", - )); - } - if let Some(id) = pending.paint_id { - *self.paint_counts.entry(id).or_insert(0) += 1; + .collect(); + let run_bytes = runs.iter().map(|r| r.run_id.capacity()).sum::() + + runs.capacity() * size_of::(); + let held = model.walk.model_bytes.saturating_add(run_bytes); + // A model refused for size, at its walk or its run stage, already passed the page budget: the + // Classify pass would charge the same and more (it also descends Forms), so it does not run. + let over_budget = model.page_reason == Some(TextReason::PageTooComplex) + && model.page_detail.as_deref() == Some(MODEL_SIZE); + let painted_forms = model + .walk + .paints + .iter() + .any(|p| matches!(p.kind, PaintKind::FormXObject { .. })); + let listed = if over_budget { + Err(TextReason::PageTooComplex) + } else if model.page_reason.is_none() && !painted_forms { + // No walk to stop: a cancel set by now still lists nothing, as a cancelled walk would. + if cancel.is_some_and(|c| c.load(Ordering::Relaxed)) { + Err(TextReason::PageTooComplex) + } else { + let mut mem = ModelBudget::new(held); + list_occurrences(ctx, model, &model.walk, &mut mem).map(|o| (o, Vec::new())) } - // A Form stream/operator can be visited more than once on the same - // page. Include its deterministic occurrence ordinal, not just its ID. - pending.locator.push_str(&format!(":{}", self.pending.len())); - self.pending.push(pending); - Ok(()) - } - - fn finish(self) -> Vec { - let counts = self.paint_counts; - self.pending - .into_iter() - .map(|mut p| { - if let Some(id) = p.paint_id { - if counts.get(&id).copied().unwrap_or(0) > 1 { - p.flags.shared = true; - } - } - let (capability, reason) = pick_reason(&p.flags); - SourceOccurrence { - page_index: p.page_index, - kind: p.kind, - rect: p.rect, - locator: p.locator, - capability, - reason, - } - }) - .collect() - } -} - -fn pick_reason(flags: &Flags) -> (SourceCapability, Option) { - let code = if flags.inline_image { - Some("INLINE_IMAGE") - } else if flags.nested_form { - Some("NESTED_FORM") - } else if flags.type3 { - Some("TYPE3") - } else if flags.vertical { - Some("VERTICAL") - } else if flags.clipped { - Some("CLIPPED") - } else if flags.pattern { - Some("PATTERN") - } else if flags.rotated { - Some("ROTATED_TEXT") - } else if flags.skewed { - Some("SKEWED_TEXT") - } else if flags.missing_font { - Some("MISSING_FONT") - } else if flags.no_tounicode { - Some("NO_TOUNICODE") - } else if flags.ambiguous { - Some("AMBIGUOUS_UNICODE") - } else if flags.masked { - Some("MASKED_IMAGE") - } else if flags.shared { - Some("SHARED_XOBJECT") - } else if flags.geometry { - Some("GEOMETRY") } else { - None - }; - match code { - Some(c) => (SourceCapability::Unsupported, Some(c.to_string())), - None => (SourceCapability::Supported, None), - } -} - -fn page_geom_unsafe(doc: &Document, page_id: ObjectId) -> bool { - if crop::page_rotation(doc, page_id) != 0 { - return true; - } - if (crop::page_user_unit(doc, page_id) - 1.0).abs() > 1e-9 { - return true; - } - let mb = crop::media_box(doc, page_id); - if let Some(cb) = crop::crop_box(doc, page_id) { - if (cb[0] - mb[0]).abs() > 1e-6 || (cb[1] - mb[1]).abs() > 1e-6 { - return true; - } - } - false -} - -fn page_resource_owners(doc: &Document, page_id: ObjectId) -> Vec { - let mut out = Vec::new(); - let mut cur = Some(page_id); - let mut steps = 0; - while let Some(id) = cur { - if steps > 32 { - break; - } - steps += 1; - let Some(dict) = object_dict(doc, id) else { - break; - }; - if dict.get(b"Resources").is_ok() { - out.push(id); - } - cur = dict.get(b"Parent").ok().and_then(|o| o.as_reference().ok()); - } - out -} - -fn object_dict(doc: &Document, id: ObjectId) -> Option<&Dictionary> { - match doc.get_object(id).ok()? { - Object::Dictionary(d) => Some(d), - Object::Stream(s) => Some(&s.dict), - _ => None, - } -} - -fn resources_of<'a>(doc: &'a Document, owner: ObjectId) -> Option<&'a Dictionary> { - match object_dict(doc, owner)?.get(b"Resources").ok()? { - Object::Dictionary(d) => Some(d), - Object::Reference(id) => doc.get_dictionary(*id).ok(), - _ => None, - } -} - -fn named_resource_entry<'a>( - doc: &'a Document, - owner: ObjectId, - category: &[u8], - name: &[u8], -) -> Option<&'a Object> { - let res = resources_of(doc, owner)?; - let cat = match res.get(category).ok()? { - Object::Dictionary(d) => d, - Object::Reference(id) => doc.get_dictionary(*id).ok()?, - _ => return None, - }; - cat.get(name).ok() -} - -fn lookup_xobject(doc: &Document, owners: &[ObjectId], name: &[u8]) -> Option { - for &owner in owners { - if let Some(Object::Reference(id)) = named_resource_entry(doc, owner, b"XObject", name) { - return Some(*id); - } - } - None -} - -enum FontRef<'a> { - Id(ObjectId), - Dict(&'a Dictionary), -} - -fn lookup_font<'a>(doc: &'a Document, owners: &[ObjectId], name: &[u8]) -> Option> { - for &owner in owners { - match named_resource_entry(doc, owner, b"Font", name) { - Some(Object::Reference(id)) => return Some(FontRef::Id(*id)), - Some(Object::Dictionary(d)) => return Some(FontRef::Dict(d)), - _ => {} + let mut mem = ModelBudget::new(held); + let walk = &model.walk; + let fonts = walk.page_fonts.iter().map(|(_, f)| f); + mem.fonts_held(fonts.chain(walk.records.iter().filter_map(|r| r.font.as_ref()))); + let mut walk = walk_page_within(ctx, &model.content, WalkMode::Classify, cancel, mem); + walk.drop_ops(); + match walk.page_reason { + Some(r) if r != TextReason::Geometry => Err(r), + _ => { + let mut mem = ModelBudget::new(walk.model_bytes); + list_occurrences(ctx, model, &walk, &mut mem).map(|o| { + // Lines that do not fit the budget are left out (they are refused anyway). + let lines = match model.page_reason { + None => form_lines(model, &walk, &mut mem).unwrap_or_default(), + Some(_) => Vec::new(), + }; + let lines = lines.into_iter().map(SourceFormLine::of).collect(); + (o, lines) + }) + } } - } - None -} - -fn font_dict<'a>(doc: &'a Document, font: &FontRef<'a>) -> Option<&'a Dictionary> { - match font { - FontRef::Id(id) => match doc.get_object(*id).ok()? { - Object::Dictionary(d) => Some(d), - Object::Stream(s) => Some(&s.dict), - _ => None, - }, - FontRef::Dict(d) => Some(*d), - } -} - -fn xobject_subtype<'a>(doc: &'a Document, id: ObjectId) -> Option<&'a [u8]> { - object_dict(doc, id)? - .get(b"Subtype") - .ok()? - .as_name() - .ok() -} - -fn xobject_is_form(doc: &Document, id: ObjectId) -> bool { - xobject_subtype(doc, id) == Some(b"Form") -} - -fn xobject_is_image(doc: &Document, id: ObjectId) -> bool { - xobject_subtype(doc, id) == Some(b"Image") -} - -fn image_is_masked(doc: &Document, id: ObjectId) -> bool { - let Some(dict) = object_dict(doc, id) else { - return false; - }; - if dict.get(b"Mask").is_ok() || dict.get(b"SMask").is_ok() { - return true; - } - match dict.get(b"ImageMask") { - Ok(Object::Boolean(true)) => true, - Ok(Object::Integer(i)) if *i != 0 => true, - _ => false, - } -} - -fn load_form(doc: &Document, id: ObjectId) -> Result { - let stream = doc - .get_object(id) - .ok() - .and_then(|o| o.as_stream().ok()) - .ok_or_else(|| malformed_content("A Form XObject is unreadable."))?; - Ok(FormInfo { - matrix: matrix_from_dict(&stream.dict), - has_resources: stream.dict.get(b"Resources").is_ok(), - bytes: decompress_stream(stream)?, - }) -} - -fn matrix_from_dict(dict: &Dictionary) -> [f64; 6] { - let Ok(obj) = dict.get(b"Matrix") else { - return IDENTITY; }; - let arr = match obj { - Object::Array(a) => a, - _ => return IDENTITY, + let (occurrences, form_lines, occurrence_reason) = match listed { + Ok((list, lines)) => (list, lines, None), + Err(reason) => (Vec::new(), Vec::new(), Some(reason)), }; - if arr.len() < 6 { - return IDENTITY; - } - let mut m = IDENTITY; - for (i, item) in arr.iter().take(6).enumerate() { - match obj_f64(item) { - Some(n) => m[i] = n, - None => return IDENTITY, - } + SourcePageResult { + page_index: model.page_index, + page_reason: model.page_reason, + runs, + occurrences, + occurrence_reason, + form_lines, } - m } -fn decompress_stream(stream: &Stream) -> Result, AppError> { - let data = match stream.decompressed_content() { - Ok(d) => d, - Err(_) => stream.content.clone(), - }; - if data.len() > MAX_STREAM_BYTES { - return Err(malformed_content( - "A content stream is larger than 32 MB decompressed.", - )); - } - Ok(data) +/// The occurrences of `walk` under `mem` (which has charged the model and the walk). +fn list_occurrences( + ctx: &SnapshotContext, + model: &PageModel, + walk: &PageWalk, + mem: &mut ModelBudget, +) -> Result, TextReason> { + occurrences(ctx, model, walk, mem).map_err(|stop| stop.reason) } -fn inspect_text( - doc: &Document, - owners: &[ObjectId], - font_name: Option<&[u8]>, - pieces: &[Vec], - tj_adj: f64, - font_size: f64, -) -> TextInspect { - let font_ref = font_name.and_then(|n| lookup_font(doc, owners, n)); - let font_id = match &font_ref { - Some(FontRef::Id(id)) => *id, - _ => (0, 0), - }; - let dict = font_ref.as_ref().and_then(|f| font_dict(doc, f)); - let mut type3 = false; - let mut vertical = false; - let mut no_tounicode = false; - let mut ambiguous = false; - let mut width_sum = 0.0; - if let Some(font) = dict { - type3 = is_type3(font); - vertical = is_vertical(doc, font); - if is_cid_or_type0(font) { - if tounicode_usable(doc, font) { - ambiguous = true; - } else { - no_tounicode = true; +struct Item { + seq: u32, + kind: SourceKind, + rect: SourceRect, + path: String, + reason: Option, + text: Option, +} + +/// Bytes a path string may take (`path`). +fn path_bound(chain: &[ObjectId]) -> usize { + chain.len().saturating_mul(PATH_FORM_BYTES) + PATH_SPAN_BYTES +} + +/// The occurrences of `walk`, each charged to `mem` (with the scratch that sorts them) before it +/// is made. +fn occurrences( + ctx: &SnapshotContext, + model: &PageModel, + walk: &PageWalk, + mem: &mut ModelBudget, +) -> Result, Stop> { + let records = model.walk.records.len(); + let listed = walk.records.len().saturating_add(walk.paints.len()); + mem.scratch( + records.saturating_mul(map_entry::<(usize, usize), usize>()) + + walk.paints.len() * map_entry::() + + listed.saturating_mul(size_of::()), + )?; + mem.hold(listed.saturating_mul(size_of::()))?; + let by_span: HashMap<(usize, usize), usize> = model + .walk + .records + .iter() + .enumerate() + .filter_map(|(i, r)| r.span.as_ref().map(|s| ((s.start, s.end), i))) + .collect(); + let mut rc = ReasonCtx::new(ctx, &model.content, walk); + let mut items: Vec = Vec::with_capacity(listed); + for rec in &walk.records { + let reason = if rec.depth > 0 { + Some(TextReason::NestedForm) + } else if let Some(page) = walk.page_reason.or(model.page_reason) { + Some(page) + } else { + match rec + .span + .as_ref() + .and_then(|s| by_span.get(&(s.start, s.end))) + { + Some(i) => model.record_reason.get(*i).copied().flatten(), + None => rc.record_reasons(rec).first().copied(), } - } - for piece in pieces { - width_sum += glyph_width_sum(doc, font, piece); - } - } else { - for piece in pieces { - width_sum += piece.iter().map(|&b| helvetica_width(b)).sum::(); - } - } - let size = font_size.abs().max(0.01); - TextInspect { - missing_font: dict.is_none(), - type3, - vertical, - no_tounicode, - ambiguous, - width: (width_sum + tj_adj) / 1000.0 * size, - font_id, - } -} - -fn is_type3(font: &Dictionary) -> bool { - font.get(b"Subtype").ok().and_then(|o| o.as_name().ok()) == Some(b"Type3") - || font.get(b"CharProcs").is_ok() -} - -fn is_cid_or_type0(font: &Dictionary) -> bool { - match font.get(b"Subtype").ok().and_then(|o| o.as_name().ok()) { - Some(b"Type0") | Some(b"CIDFontType0") | Some(b"CIDFontType2") => return true, - _ => {} - } - if font.get(b"CIDSystemInfo").is_ok() || font.get(b"DescendantFonts").is_ok() { - return true; - } - matches!( - font.get(b"Encoding").ok().and_then(|o| o.as_name().ok()), - Some(b"Identity-H") | Some(b"Identity-V") - ) -} - -fn is_vertical(doc: &Document, font: &Dictionary) -> bool { - if font.get(b"Encoding").ok().and_then(|o| o.as_name().ok()) == Some(b"Identity-V") { - return true; - } - if wmode_is_1(font) { - return true; - } - descendant_is_vertical(doc, font) -} - -fn wmode_is_1(font: &Dictionary) -> bool { - matches!(font.get(b"WMode").ok().and_then(obj_f64), Some(w) if (w - 1.0).abs() < 0.5) -} - -fn descendant_is_vertical(doc: &Document, font: &Dictionary) -> bool { - let items = descendant_font_objects(doc, font); - for item in &items { - let dict = match item { - Object::Dictionary(d) => d, - Object::Reference(id) => match doc.get_object(*id) { - Ok(Object::Dictionary(d)) => d, - Ok(Object::Stream(s)) => &s.dict, - _ => continue, - }, - _ => continue, }; - if dict.get(b"Encoding").ok().and_then(|o| o.as_name().ok()) == Some(b"Identity-V") - || wmode_is_1(dict) - { - return true; + // The text is collected by appending (room for up to twice its bytes). + let text: usize = rec + .glyphs + .iter() + .filter_map(|g| g.text.as_ref()) + .map(String::len) + .sum(); + mem.hold(2 * text + 2 * path_bound(&rec.form_chain) + LOCATOR_FIXED_BYTES)?; + items.push(Item { + seq: rec.seq, + kind: SourceKind::Text, + rect: text_rect(rec), + path: path(rec.depth, &rec.form_chain, &rec.local_span), + reason, + text: rec.glyphs.iter().map(|g| g.text.as_deref()).collect(), + }); + } + let refs = ctx.refs(); + let mut paints_of: HashMap = HashMap::new(); + for p in &walk.paints { + if let Some(x) = p.xobject { + *paints_of.entry(x.id).or_default() += 1; + } + } + let rotation = display_rotation(walk.geometry.rotate); + for p in &walk.paints { + if !matches!( + p.kind, + PaintKind::ImageXObject { .. } | PaintKind::InlineImage { .. } + ) { + continue; } - } - false -} - -fn descendant_font_objects(doc: &Document, font: &Dictionary) -> Vec { - match font.get(b"DescendantFonts") { - Ok(Object::Array(a)) => a.clone(), - Ok(Object::Reference(id)) => match doc.get_object(*id) { - Ok(Object::Array(a)) => a.clone(), - _ => Vec::new(), - }, - _ => Vec::new(), - } -} - -fn tounicode_usable(doc: &Document, font: &Dictionary) -> bool { - let Ok(obj) = font.get(b"ToUnicode") else { - return false; - }; - let bytes = match obj { - Object::Stream(s) => s.get_plain_content().unwrap_or_else(|_| s.content.clone()), - Object::Reference(id) => match doc.get_object(*id) { - Ok(Object::Stream(s)) => s.get_plain_content().unwrap_or_else(|_| s.content.clone()), - _ => return false, - }, - _ => return false, - }; - let text = String::from_utf8_lossy(&bytes); - text.contains("begincmap") - && (text.contains("beginbfchar") || text.contains("beginbfrange")) -} - -fn glyph_width_sum(doc: &Document, font: &Dictionary, bytes: &[u8]) -> f64 { - if is_cid_or_type0(font) { - return (bytes.len() / 2) as f64 * descendant_dw(doc, font); - } - if let Some(widths) = explicit_widths(doc, font) { - return bytes.iter().map(|&b| widths[b as usize]).sum(); - } - bytes.iter().map(|&b| helvetica_width(b)).sum() -} - -fn descendant_dw(doc: &Document, font: &Dictionary) -> f64 { - if let Some(n) = font.get(b"DW").ok().and_then(obj_f64) { - return n; - } - for item in descendant_font_objects(doc, font) { - let dict = match &item { - Object::Dictionary(d) => d, - Object::Reference(id) => match doc.get_object(*id) { - Ok(Object::Dictionary(d)) => d, - _ => continue, + let shared = p.xobject.is_some_and(|x| { + refs.count(x.id) != 1 || x.shared_path || paints_of.get(&x.id).copied() != Some(1) + }); + let bbox = p.bbox.unwrap_or_default(); + mem.hold(path_bound(&p.form_chain) * 2 + LOCATOR_FIXED_BYTES)?; + items.push(Item { + seq: p.seq, + kind: SourceKind::Image, + rect: SourceRect { + x: bbox[0], + y: bbox[1], + w: (bbox[2] - bbox[0]).max(0.01), + h: (bbox[3] - bbox[1]).max(0.01), }, - _ => continue, - }; - if let Some(n) = dict.get(b"DW").ok().and_then(obj_f64) { - return n; - } - } - 500.0 -} - -fn explicit_widths(doc: &Document, font: &Dictionary) -> Option<[f64; 256]> { - // Helvetica table only when /Widths is truly absent (Standard-14). - let arr = font.get_deref(b"Widths", doc).ok()?.as_array().ok()?; - let first = font - .get_deref(b"FirstChar", doc) - .ok() - .and_then(|o| o.as_i64().ok()) - .unwrap_or(0) - .max(0) as usize; - let last = font - .get_deref(b"LastChar", doc) - .ok() - .and_then(|o| o.as_i64().ok()) - .map(|n| n.max(0) as usize) - .unwrap_or_else(|| first.saturating_add(arr.len().saturating_sub(1)).min(255)); - let mut widths = [500.0; 256]; - for (i, obj) in arr.iter().enumerate() { - let code = first + i; - if code > last || code > 255 { - break; - } - if let Some(n) = obj_f64(obj) { - widths[code] = n; - } - } - Some(widths) -} - -/// Standard-14 Helvetica widths per 1000. Locked: H=667, i=278. -fn helvetica_width(code: u8) -> f64 { - match code { - b' ' => 278.0, - b'!' => 278.0, - b'"' => 355.0, - b'#' => 556.0, - b'$' => 556.0, - b'%' => 889.0, - b'&' => 667.0, - b'\'' => 191.0, - b'(' => 333.0, - b')' => 333.0, - b'*' => 389.0, - b'+' => 584.0, - b',' => 278.0, - b'-' => 333.0, - b'.' => 278.0, - b'/' => 278.0, - b'0'..=b'9' => 556.0, - b':' => 278.0, - b';' => 278.0, - b'<' => 584.0, - b'=' => 584.0, - b'>' => 584.0, - b'?' => 556.0, - b'@' => 1015.0, - b'A' => 667.0, - b'B' => 667.0, - b'C' => 722.0, - b'D' => 722.0, - b'E' => 667.0, - b'F' => 611.0, - b'G' => 778.0, - b'H' => 667.0, - b'I' => 278.0, - b'J' => 500.0, - b'K' => 667.0, - b'L' => 556.0, - b'M' => 833.0, - b'N' => 722.0, - b'O' => 778.0, - b'P' => 667.0, - b'Q' => 778.0, - b'R' => 722.0, - b'S' => 667.0, - b'T' => 611.0, - b'U' => 722.0, - b'V' => 667.0, - b'W' => 944.0, - b'X' => 667.0, - b'Y' => 667.0, - b'Z' => 611.0, - b'[' => 278.0, - b'\\' => 278.0, - b']' => 278.0, - b'^' => 469.0, - b'_' => 556.0, - b'`' => 333.0, - b'a' => 556.0, - b'b' => 556.0, - b'c' => 500.0, - b'd' => 556.0, - b'e' => 556.0, - b'f' => 278.0, - b'g' => 556.0, - b'h' => 556.0, - b'i' => 278.0, - b'j' => 222.0, - b'k' => 500.0, - b'l' => 278.0, - b'm' => 833.0, - b'n' => 556.0, - b'o' => 556.0, - b'p' => 556.0, - b'q' => 556.0, - b'r' => 333.0, - b's' => 500.0, - b't' => 278.0, - b'u' => 556.0, - b'v' => 500.0, - b'w' => 722.0, - b'x' => 500.0, - b'y' => 500.0, - b'z' => 500.0, - b'{' => 334.0, - b'|' => 260.0, - b'}' => 334.0, - b'~' => 584.0, - _ => 556.0, - } -} - -fn is_rotated_tm(m: [f64; 6]) -> bool { - let [a, b, c, d, _, _] = m; - let det = a * d - b * c; - if det.abs() < 1e-6 { - return false; - } - // 90° / 270°: off-axis, a≈d, b≈-c. - if (b.abs() > 0.1 || c.abs() > 0.1) && (a - d).abs() < 0.25 && (b + c).abs() < 0.25 { - return true; - } - // 180°: axis-aligned invert. - a < 0.0 && d < 0.0 && b.abs() <= 0.05 && c.abs() <= 0.05 -} - -fn is_skewed_tm(m: [f64; 6]) -> bool { - let [a, b, c, d, _, _] = m; - if b.abs() < 0.05 && c.abs() < 0.05 { - return false; - } - let det = a * d - b * c; - if det.abs() < 1e-6 { - return false; - } - !is_rotated_tm(m) -} - -fn mul(m: [f64; 6], n: [f64; 6]) -> [f64; 6] { - let [a, b, c, d, e, f] = m; - let [a2, b2, c2, d2, e2, f2] = n; - [ - a * a2 + c * b2, - b * a2 + d * b2, - a * c2 + c * d2, - b * c2 + d * d2, - a * e2 + c * f2 + e, - b * e2 + d * f2 + f, - ] -} - -fn apply_point(m: [f64; 6], x: f64, y: f64) -> (f64, f64) { - (m[0] * x + m[2] * y + m[4], m[1] * x + m[3] * y + m[5]) + path: path(p.depth, &p.form_chain, &p.local_span), + reason: image_reason(walk, p, shared, &rotation), + text: None, + }); + } + items.sort_by_key(|i| i.seq); + let fp = ctx.snap.fingerprint; + Ok(items + .into_iter() + .enumerate() + .map(|(ordinal, i)| { + let k = match i.kind { + SourceKind::Text => 't', + SourceKind::Image => 'i', + }; + SourceOccurrence { + page_index: model.page_index, + kind: i.kind, + rect: i.rect, + locator: format!("v2:{fp}:{}:{k}:{}:{ordinal}", model.page_index, i.path), + capability: capability(i.reason), + reason: i.reason, + text: i.text, + } + }) + .collect()) } -fn unit_square_bbox(ctm: [f64; 6]) -> SourceRect { - let pts = [ - apply_point(ctm, 0.0, 0.0), - apply_point(ctm, 1.0, 0.0), - apply_point(ctm, 0.0, 1.0), - apply_point(ctm, 1.0, 1.0), - ]; - let min_x = pts.iter().map(|p| p.0).fold(f64::INFINITY, f64::min); - let min_y = pts.iter().map(|p| p.1).fold(f64::INFINITY, f64::min); - let max_x = pts.iter().map(|p| p.0).fold(f64::NEG_INFINITY, f64::max); - let max_y = pts.iter().map(|p| p.1).fold(f64::NEG_INFINITY, f64::max); +/// Baseline origin (rise included), the advance length and the effective size (D15). +fn text_rect(rec: &ShowRecord) -> SourceRect { + let m = mul(&rec.tm_before, &rec.before.ctm); + let effective = rec.before.text.tfs.abs() * m[2].hypot(m[3]); + let advance = (rec.pen_after.0 - rec.pen_before.0).hypot(rec.pen_after.1 - rec.pen_before.1); SourceRect { - x: min_x, - y: min_y, - w: (max_x - min_x).max(0.01), - h: (max_y - min_y).max(0.01), - } -} - -fn apply_td(ts: &mut TextState, tx: f64, ty: f64) { - let m = mul(ts.tlm, [1.0, 0.0, 0.0, 1.0, tx, ty]); - ts.tm = m; - ts.tlm = m; -} - -fn obj_f64(obj: &Object) -> Option { - match obj { - Object::Integer(i) => Some(*i as f64), - Object::Real(r) => Some(*r as f64), - _ => None, - } -} - -fn six_nums(operands: &[Object]) -> Option<[f64; 6]> { - if operands.len() < 6 { - return None; - } - let start = operands.len() - 6; - let mut m = [0.0; 6]; - for i in 0..6 { - m[i] = obj_f64(&operands[start + i])?; + x: rec.pen_before.0, + y: rec.pen_before.1, + w: advance.max(0.01), + h: effective.max(0.01), } - Some(m) } -fn two_nums(operands: &[Object]) -> Option<(f64, f64)> { - if operands.len() < 2 { - return None; +/// `p{start}-{end}` at depth 0, else `x{obj}.{gen}>…:{start}-{end}`. +fn path(depth: u8, chain: &[ObjectId], span: &std::ops::Range) -> String { + if depth == 0 { + return format!("p{}-{}", span.start, span.end); } - let n = operands.len(); - Some((obj_f64(&operands[n - 2])?, obj_f64(&operands[n - 1])?)) -} - -fn is_pattern_name(operands: &[Object]) -> bool { - operands - .iter() - .any(|o| o.as_name().ok() == Some(b"Pattern")) + let forms: Vec = chain.iter().map(|(o, g)| format!("x{o}.{g}")).collect(); + format!("{}:{}-{}", forms.join(">"), span.start, span.end) } -fn operand_selects_pattern_space( - doc: &Document, - owners: &[ObjectId], - operands: &[Object], -) -> bool { - let Some(obj) = operands.first() else { - return false; - }; - if color_space_object_is_pattern(doc, obj) { - return true; - } - let Ok(name) = obj.as_name() else { - return false; - }; - if name == b"Pattern" { - return true; - } - for &owner in owners { - if let Some(entry) = named_resource_entry(doc, owner, b"ColorSpace", name) { - return color_space_object_is_pattern(doc, entry); - } - } - false -} - -fn color_space_object_is_pattern(doc: &Document, obj: &Object) -> bool { - let resolved = match doc.dereference(obj) { - Ok((_, o)) => o, - Err(_) => return false, - }; - match resolved { - Object::Name(n) => n.as_slice() == b"Pattern", - Object::Array(arr) => arr.first().is_some_and(|first| { - let first = doc.dereference(first).map(|(_, o)| o).unwrap_or(first); - first.as_name().ok() == Some(b"Pattern") - }), - _ => false, - } -} - -fn apply_extgstate_smask(doc: &Document, owners: &[ObjectId], name: &[u8], gs: &mut GState) { - for &owner in owners { - let Some(entry) = named_resource_entry(doc, owner, b"ExtGState", name) else { - continue; - }; - let Some(dict) = deref_dict(doc, entry) else { - continue; - }; - let Ok(smask) = dict.get(b"SMask") else { - return; +/// Image priority (§A.10): INLINE_IMAGE, NESTED_FORM, CLIPPED, PATTERN, MASKED_IMAGE, +/// SHARED_XOBJECT, TRANSFORMED_IMAGE, GEOMETRY. +fn image_reason( + walk: &PageWalk, + p: &PaintRecord, + shared: bool, + rotation: &Matrix, +) -> Option { + use TextReason as R; + if matches!(p.kind, PaintKind::InlineImage { .. }) { + return Some(R::InlineImage); + } + if p.depth > 0 { + return Some(R::NestedForm); + } + let bbox = p.bbox.unwrap_or_default(); + let clipped = !contains(walk.geometry.visible, bbox, CLIP_CONTAIN_TOL_PT) + || match &p.state.clip { + ClipState::None => false, + ClipState::Rect(r) => !contains(*r, bbox, CLIP_CONTAIN_TOL_PT), + ClipState::Complex => true, }; - let resolved = doc.dereference(smask).map(|(_, o)| o).unwrap_or(smask); - gs.masked = resolved.as_name().ok() != Some(b"None"); - return; + if clipped { + return Some(R::Clipped); } -} - -fn deref_dict<'a>(doc: &'a Document, obj: &'a Object) -> Option<&'a Dictionary> { - match doc.dereference(obj) { - Ok((_, Object::Dictionary(d))) => Some(d), - Ok((_, Object::Stream(s))) => Some(&s.dict), - _ => None, + if p.state.fill.pattern || p.state.stroke.pattern { + return Some(R::Pattern); } -} - -fn text_pieces(operands: &[Object]) -> (Vec>, f64) { - let mut pieces = Vec::new(); - let mut adj = 0.0; - for obj in operands { - collect_text_obj(obj, &mut pieces, &mut adj); + if p.masked { + return Some(R::MaskedImage); } - (pieces, adj) -} - -fn collect_text_obj(obj: &Object, pieces: &mut Vec>, adj: &mut f64) { - match obj { - Object::String(s, _) => pieces.push(s.clone()), - Object::Integer(i) => *adj -= *i as f64, - Object::Real(r) => *adj -= *r as f64, - Object::Array(arr) => { - for item in arr { - collect_text_obj(item, pieces, adj); - } - } - _ => {} + if shared { + return Some(R::SharedXobject); } -} - -fn spacing_advance(pieces: &[Vec], tc: f64, tw: f64) -> f64 { - let mut extra = 0.0; - for piece in pieces { - extra += piece.len() as f64 * tc; - extra += piece.iter().filter(|&&b| b == b' ').count() as f64 * tw; + let [a, b, c, d, _, _] = mul(&p.state.ctm, rotation); + let tol = AXIS_EPSILON_REL * a.abs().max(b.abs()).max(c.abs()).max(d.abs()); + if !(b.abs() <= tol && c.abs() <= tol && a > 0.0 && d > 0.0) { + return Some(R::TransformedImage); } - extra + walk.page_reason } -fn strip_inline_image_payloads(data: &[u8]) -> Result, AppError> { - let mut out = Vec::with_capacity(data.len()); - let mut i = 0; - // Normal → Keys (after BI) → Payload (after ID). Copy BI/keys/ID/EI; - // drop only the raw bytes between ID and EI so Content::decode still - // yields BI at the CTM in force there. - let mut after_bi = false; - let mut after_id = false; - while i < data.len() { - if after_id { - if is_op_token(data, i, b"EI") { - out.extend_from_slice(b"EI"); - i += 2; - after_id = false; - after_bi = false; - } else { - i += 1; - } - continue; - } - if data[i].is_ascii_whitespace() { - out.push(data[i]); - i += 1; - continue; - } - if data[i] == b'%' { - let start = i; - while i < data.len() && data[i] != b'\n' && data[i] != b'\r' { - i += 1; - } - out.extend_from_slice(&data[start..i]); - continue; - } - if data[i] == b'(' { - let start = i; - i = skip_literal(data, i)?; - out.extend_from_slice(&data[start..i]); - continue; - } - if data[i] == b'<' && data.get(i + 1) != Some(&b'<') { - let start = i; - i = skip_hex(data, i); - out.extend_from_slice(&data[start..i]); - continue; - } - if !after_bi && is_op_token(data, i, b"BI") { - out.extend_from_slice(b"BI"); - i += 2; - after_bi = true; - continue; - } - if after_bi && is_op_token(data, i, b"ID") { - out.extend_from_slice(b"ID"); - i += 2; - out.push(b' '); - after_id = true; - continue; - } - if after_bi && is_op_token(data, i, b"EI") { - out.extend_from_slice(b"EI"); - i += 2; - after_bi = false; - continue; - } - out.push(data[i]); - i += 1; - } - if after_bi || after_id { - return Err(malformed_content("An inline image is unterminated.")); +/// Document-wide classification (tests and the corpus): every page through +/// `classify_source_page`; the first page-level refusal fails the whole call (its details name +/// the page) — a page refused `GEOMETRY` is not an error, its occurrences carry the reason. +#[cfg(test)] +pub fn classify_source_content( + path: &std::path::Path, +) -> Result, crate::error::AppError> { + use crate::pdf_engine::text_edit::runs::build_page_model; + use crate::pdf_engine::text_edit::snapshot::read_snapshot; + let ctx = SnapshotContext::new(read_snapshot(path)?); + let mut out = Vec::new(); + for page in 0..ctx.snap.pages.len() { + let page = u32::try_from(page).unwrap_or(u32::MAX); + let result = classify_source_page(&ctx, &build_page_model(&ctx, page, None)?, None); + page_refusal(&result)?; + out.extend(result.occurrences); } Ok(out) } -fn count_inline_images(data: &[u8]) -> Result { - let mut count = 0; - let mut i = 0; - let mut in_inline = false; - while i < data.len() { - if data[i].is_ascii_whitespace() { - i += 1; - continue; - } - if !in_inline && data[i] == b'%' { - while i < data.len() && data[i] != b'\n' && data[i] != b'\r' { - i += 1; - } - continue; - } - if !in_inline && data[i] == b'(' { - i = skip_literal(data, i)?; - continue; - } - if !in_inline && data[i] == b'<' && data.get(i + 1) != Some(&b'<') { - i = skip_hex(data, i); - continue; - } - if !in_inline && is_op_token(data, i, b"BI") { - in_inline = true; - i += 2; - continue; - } - if in_inline && is_op_token(data, i, b"EI") { - count += 1; - in_inline = false; - i += 2; - continue; - } - i += 1; - } - if in_inline { - return Err(malformed_content("An inline image is unterminated.")); - } - Ok(count) -} - -fn is_delim(b: u8) -> bool { - b.is_ascii_whitespace() - || matches!( - b, - b'(' | b')' | b'<' | b'>' | b'[' | b']' | b'{' | b'}' | b'/' | b'%' - ) -} - -fn is_op_token(data: &[u8], i: usize, token: &[u8]) -> bool { - if i + token.len() > data.len() || &data[i..i + token.len()] != token { - return false; +/// The occurrence a v2 locator names, from one fresh read of `path`: a fingerprint that differs +/// (or an occurrence that no longer exists) is `STALE`; only the locator's page is walked. +#[cfg(test)] +pub fn resolve_source_locator( + path: &std::path::Path, + locator: &str, +) -> Result { + use crate::pdf_engine::text_edit::runs::build_page_model; + use crate::pdf_engine::text_edit::snapshot::{fnv1a_u64, snapshot_from_bytes, Fingerprint}; + let name = path + .file_name() + .map(|n| n.to_string_lossy().into_owned()) + .unwrap_or_default(); + let stale = || crate::pdf_engine::text_edit::reasons::stale(&name); + let mut fields = locator.split(':'); + let (Some("v2"), Some(fp), Some(page)) = (fields.next(), fields.next(), fields.next()) else { + return Err(stale()); + }; + let page: u32 = page.parse().map_err(|_| stale())?; + let (bytes, modified) = read_once(path)?; + let fingerprint = Fingerprint { + len: bytes.len() as u64, + fnv: fnv1a_u64(&bytes), + }; + if fingerprint.to_string() != fp { + return Err(stale()); } - let before = i == 0 || is_delim(data[i - 1]); - let after = i + token.len() == data.len() || is_delim(data[i + token.len()]); - before && after + let ctx = SnapshotContext::new(snapshot_from_bytes(path, bytes, modified)?); + let result = classify_source_page(&ctx, &build_page_model(&ctx, page, None)?, None); + page_refusal(&result)?; + result + .occurrences + .into_iter() + .find(|o| o.locator == locator) + .ok_or_else(stale) } -fn skip_literal(data: &[u8], start: usize) -> Result { - let mut i = start + 1; - let mut depth = 1; - while i < data.len() && depth > 0 { - match data[i] { - b'\\' => { - i = i.saturating_add(2).min(data.len()); - continue; - } - b'(' => depth += 1, - b')' => depth -= 1, - _ => {} +#[cfg(test)] +fn page_refusal(result: &SourcePageResult) -> Result<(), crate::error::AppError> { + let reason = result + .page_reason + .filter(|r| *r != TextReason::Geometry) + .or(result.occurrence_reason); + match reason { + None => Ok(()), + Some(r) => { + let page = result.page_index.saturating_add(1); + Err(r + .to_app_error(Some(page)) + .with_details(format!("page {page}: {}", r.as_str()))) } - i += 1; - } - if depth != 0 { - return Err(malformed_content("A literal string is unterminated.")); - } - Ok(i) -} - -fn skip_hex(data: &[u8], start: usize) -> usize { - let mut i = start + 1; - while i < data.len() && data[i] != b'>' { - i += 1; - } - if i < data.len() { - i + 1 - } else { - i } } -fn document_is_signed(doc: &Document) -> bool { - if let Ok(cat) = doc.catalog() { - if cat.get(b"Perms").is_ok() { - return true; - } +/// The one capped read of `path` (same rules as `read_snapshot`). +#[cfg(test)] +fn read_once( + path: &std::path::Path, +) -> Result<(Vec, Option), crate::error::AppError> { + use crate::error::AppError; + use crate::pdf_engine::text_edit::{limits, reasons}; + use std::io::Read; + if !path.is_file() { + return Err(AppError::invalid_pdf(&path.to_string_lossy())); } - for obj in doc.objects.values() { - if object_is_signature(obj) { - return true; - } + let cap = limits::file_cap(); + let meta = std::fs::metadata(path).map_err(|e| AppError::io("Could not read the PDF.", e))?; + if meta.len() > cap { + return Err(reasons::file_too_large()); } - false -} - -fn object_is_signature(obj: &Object) -> bool { - let dict = match obj { - Object::Dictionary(d) => d, - Object::Stream(s) => &s.dict, - _ => return false, - }; - let is_sig = name_is(dict, b"Type", b"Sig") || name_is(dict, b"FT", b"Sig"); - is_sig && dict.get(b"ByteRange").is_ok() -} - -fn name_is(dict: &Dictionary, key: &[u8], expect: &[u8]) -> bool { - dict.get(key).ok().and_then(|o| o.as_name().ok()) == Some(expect) -} - -fn fnv1a_u64(bytes: &[u8]) -> u64 { - let mut hash = FNV_OFFSET; - for byte in bytes { - hash ^= *byte as u64; - hash = hash.wrapping_mul(FNV_PRIME); + let mut bytes = Vec::new(); + std::fs::File::open(path) + .and_then(|f| f.take(cap.saturating_add(1)).read_to_end(&mut bytes)) + .map_err(|e| AppError::io("Could not read the PDF.", e))?; + if bytes.len() as u64 > cap { + return Err(reasons::file_too_large()); } - hash -} - -fn encode_locator( - fp: u64, - page_index: u32, - kind: SourceKind, - contents_id: ObjectId, - op_index: u32, - object_id: ObjectId, -) -> String { - let k = match kind { - SourceKind::Text => 0u8, - SourceKind::Image => 1u8, - }; - format!( - "v1:{fp:016x}:{page_index}:{k}:{}:{}:{op_index}:{}:{}", - contents_id.0, contents_id.1, object_id.0, object_id.1 - ) -} - -fn parse_locator_fp(locator: &str) -> Option { - let rest = locator.strip_prefix("v1:")?; - let hex = rest.split(':').next()?; - u64::from_str_radix(hex, 16).ok() -} - -fn path_str(path: &Path) -> String { - path.to_string_lossy().into_owned() -} - -fn file_too_large() -> AppError { - AppError::new( - "FILE_TOO_LARGE", - "File too large to classify", - "Classifying source content needs the document loaded into memory, and this file is over 400 MB.", - ) - .with_suggestion("Use a smaller PDF.") -} - -fn malformed_content(message: impl Into) -> AppError { - AppError::new( - "MALFORMED_CONTENT", - "This PDF content cannot be read", - message, - ) - .with_suggestion("Open the file in a PDF editor that can repair it, or use a different PDF.") -} - -fn encrypted() -> AppError { - AppError::new( - "ENCRYPTED", - "This PDF is encrypted", - "OffPDF cannot classify source content in an encrypted PDF.", - ) - .with_suggestion("Unlock the PDF and try again.") -} - -fn signed() -> AppError { - AppError::new( - "SIGNED", - "This PDF is signed", - "OffPDF cannot classify source content in a signed PDF.", - ) - .with_suggestion("Use an unsigned copy of the file.") -} - -fn stale() -> AppError { - AppError::new( - "STALE", - "This locator is stale", - "The source file no longer matches the locator fingerprint.", - ) + Ok((bytes, meta.modified().ok())) } diff --git a/src-tauri/src/pdf_engine/source_content_integ.rs b/src-tauri/src/pdf_engine/source_content_integ.rs index 9b32791..3c528da 100644 --- a/src-tauri/src/pdf_engine/source_content_integ.rs +++ b/src-tauri/src/pdf_engine/source_content_integ.rs @@ -1,25 +1,19 @@ -//! Read-only source-content classifier tests (#33). +//! Source-content classifier tests (#33): the corpus contract (`corpus.rs`), the PR #97 review +//! regressions R1–R14 (`review.rs`, `review2.rs`) and the hardening regressions CLS-H01…H27 for +//! every maintainer item of #33 (`hardening.rs`, `hardening2.rs`). //! -//! Production module is added by impl: `crate::pdf_engine::source_content`. -//! This file only imports the locked exports. Fail-today is a missing module. -//! -//! Locked surface (field names so impl can match): -//! -//! ```ignore -//! pub struct SourceOccurrence { -//! pub page_index: u32, -//! pub kind: /* enum or string: text | image */, -//! pub rect: /* { x, y, w, h } unrotated PDF user space */, -//! pub locator: String, -//! pub capability: /* enum or string: supported | unsupported */, -//! pub reason: Option, -//! } -//! pub fn classify_source_content(path: &Path) -> Result, AppError>; -//! pub fn resolve_source_locator(path: &Path, locator: &str) -> Result; -//! ``` +//! Surface (SPEC §C): `SourceOccurrence { page_index, kind, rect, locator, capability, +//! reason: Option, text }`; the document-wide `classify_source_content` and +//! `resolve_source_locator` (test-only) loop the per-page `classify_source_page`. #![cfg(test)] +mod corpus; +mod hardening; +mod hardening2; +mod review; +mod review2; + use crate::error::AppError; use crate::pdf_engine::source_content::{ classify_source_content, resolve_source_locator, SourceOccurrence, @@ -33,27 +27,39 @@ use std::time::SystemTime; /// Same 400 MiB gate as forms / links / outline. const FILE_CAP_BYTES: u64 = 400 * 1024 * 1024; -const FROZEN_REASONS: &[&str] = &[ - "MISSING_FONT", - "NO_TOUNICODE", - "AMBIGUOUS_UNICODE", - "TYPE3", +/// The 14 #33 reason names that must survive (CLS-H22). +const LEGACY_REASONS: [&str; 14] = [ + "INLINE_IMAGE", "NESTED_FORM", - "ROTATED_TEXT", - "SKEWED_TEXT", + "TYPE3", "VERTICAL", "CLIPPED", "PATTERN", - "SHARED_XOBJECT", - "INLINE_IMAGE", + "ROTATED_TEXT", + "SKEWED_TEXT", + "MISSING_FONT", + "NO_TOUNICODE", + "AMBIGUOUS_UNICODE", "MASKED_IMAGE", - "ENCRYPTED", - "SIGNED", - "MALFORMED", - "STALE", + "SHARED_XOBJECT", "GEOMETRY", ]; +/// Every code of `src/lib/editor/text-reasons.json` (run, page, image, file, problem, save, +/// warning lists): the frozen vocabulary (CLS-H22 — it used to be a hand list with a typo). +fn frozen_reasons() -> Vec { + let json: serde_json::Value = + serde_json::from_str(include_str!("../../../src/lib/editor/text-reasons.json")) + .expect("text-reasons.json parses"); + json.as_object() + .expect("text-reasons.json is an object") + .values() + .filter_map(|v| v.as_array()) + .flatten() + .filter_map(|v| v.as_str().map(str::to_string)) + .collect() +} + const STAND_INS: &[(&str, &str, &str)] = &[ ("text-type3.pdf", "text", "TYPE3"), ("text-nested-form.pdf", "text", "NESTED_FORM"), @@ -156,14 +162,7 @@ fn capability_token(occ: &SourceOccurrence) -> String { } fn reason_code(occ: &SourceOccurrence) -> Option { - occ.reason.as_ref().and_then(|r| { - let trimmed = r.trim(); - if trimmed.is_empty() { - None - } else { - Some(trimmed.to_string()) - } - }) + occ.reason.map(|r| r.as_str().to_string()) } fn classify(path: &Path, must_id: &str) -> Vec { @@ -238,7 +237,7 @@ fn assert_unsupported(occ: &SourceOccurrence, kind: &str, reason: &str, must_id: reason_code(occ) ); assert!( - FROZEN_REASONS.contains(&reason), + frozen_reasons().iter().any(|f| f == reason), "{must_id}: {reason} is not a frozen reason code" ); assert!( @@ -308,1369 +307,3 @@ fn expect_bounds_err( ), } } - -// --- CLASSIFY-API ----------------------------------------------------------- - -#[test] -fn classify_api_lists_occurrences() { - let path = fixture("text-tj.pdf"); - let hits = classify(&path, "CLASSIFY-API"); - assert!( - !hits.is_empty(), - "CLASSIFY-API: classify_source_content(text-tj.pdf) must list occurrences" - ); -} - -// --- CLASSIFY-SOURCE-PATH --------------------------------------------------- - -#[test] -fn classify_uses_corpus_source_path_not_page_pdf() { - let path = fixture("text-tj.pdf"); - let rendered = path.to_string_lossy(); - assert!( - rendered.contains("fixtures/source-edit") && rendered.ends_with("text-tj.pdf"), - "CLASSIFY-SOURCE-PATH: must pass the corpus source path, not pagePdf / --empty --pages; got {}", - path.display() - ); - // Do not spawn qpdf --empty --pages and do not call page_pdf_b64. - let hits = classify(&path, "CLASSIFY-SOURCE-PATH"); - let occ = first_of_kind(&hits, "text", "CLASSIFY-SOURCE-PATH"); - assert_eq!( - occ.page_index, 0, - "CLASSIFY-SOURCE-PATH: text-tj.pdf page_index is 0 on the original source" - ); -} - -// --- CLASSIFY-TRY-EDIT-TJ --------------------------------------------------- - -#[test] -fn classify_text_tj_supported_at_origin() { - let path = fixture("text-tj.pdf"); - let hits = classify(&path, "CLASSIFY-TRY-EDIT-TJ"); - let occ = first_of_kind(&hits, "text", "CLASSIFY-TRY-EDIT-TJ"); - assert_supported_text_or_image(occ, "text", "CLASSIFY-TRY-EDIT-TJ"); - assert_eq!( - occ.page_index, 0, - "CLASSIFY-TRY-EDIT-TJ: page_index must be 0" - ); - assert!( - (occ.rect.x - 72.0).abs() <= 1.0, - "CLASSIFY-TRY-EDIT-TJ: origin x must be ~72; got {}", - occ.rect.x - ); - assert!( - (occ.rect.y - 720.0).abs() <= 1.0, - "CLASSIFY-TRY-EDIT-TJ: origin y must be ~720; got {}", - occ.rect.y - ); - assert!( - occ.rect.w > 0.0 && occ.rect.h > 0.0, - "CLASSIFY-TRY-EDIT-TJ: w/h must be positive; got w={} h={}", - occ.rect.w, - occ.rect.h - ); -} - -// --- CLASSIFY-TRY-EDIT-IMAGE ------------------------------------------------ - -#[test] -fn classify_image_unique_supported() { - let path = fixture("image-unique.pdf"); - let hits = classify(&path, "CLASSIFY-TRY-EDIT-IMAGE"); - let occ = first_of_kind(&hits, "image", "CLASSIFY-TRY-EDIT-IMAGE"); - assert_supported_text_or_image(occ, "image", "CLASSIFY-TRY-EDIT-IMAGE"); - assert_eq!( - occ.page_index, 0, - "CLASSIFY-TRY-EDIT-IMAGE: page_index must be 0" - ); - assert!( - (occ.rect.x - 72.0).abs() <= 1.0, - "CLASSIFY-TRY-EDIT-IMAGE: x must be ~72; got {}", - occ.rect.x - ); - assert!( - (occ.rect.y - 400.0).abs() <= 1.0, - "CLASSIFY-TRY-EDIT-IMAGE: y must be ~400; got {}", - occ.rect.y - ); - assert!( - (occ.rect.w - 40.0).abs() <= 1.0, - "CLASSIFY-TRY-EDIT-IMAGE: w must be ~40; got {}", - occ.rect.w - ); - assert!( - (occ.rect.h - 40.0).abs() <= 1.0, - "CLASSIFY-TRY-EDIT-IMAGE: h must be ~40; got {}", - occ.rect.h - ); -} - -// --- CLASSIFY-NO-TOUNICODE -------------------------------------------------- - -#[test] -fn classify_cid_no_tounicode_is_unsupported() { - let path = fixture("text-cid-no-tounicode.pdf"); - let hits = classify(&path, "CLASSIFY-NO-TOUNICODE"); - let occ = first_of_kind(&hits, "text", "CLASSIFY-NO-TOUNICODE"); - assert_unsupported(occ, "text", "NO_TOUNICODE", "CLASSIFY-NO-TOUNICODE"); -} - -// --- CLASSIFY-SHARED-IMAGE -------------------------------------------------- - -#[test] -fn classify_image_reused_shared_xobject() { - let path = fixture("image-reused.pdf"); - let hits = classify(&path, "CLASSIFY-SHARED-IMAGE"); - let images: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "image").collect(); - assert_eq!( - images.len(), - 2, - "CLASSIFY-SHARED-IMAGE: image-reused.pdf must yield two image occurrences; got {:?}", - hits.iter() - .map(|o| ( - o.page_index, - kind_token(o), - capability_token(o), - reason_code(o) - )) - .collect::>() - ); - for occ in images { - assert_unsupported(occ, "image", "SHARED_XOBJECT", "CLASSIFY-SHARED-IMAGE"); - } -} - -// --- CLASSIFY-STAND-INS ----------------------------------------------------- - -#[test] -fn classify_stand_ins_locked_reasons() { - for (file, kind, reason) in STAND_INS { - let path = fixture(file); - let hits = classify(&path, "CLASSIFY-STAND-INS"); - let occ = first_of_kind(&hits, kind, "CLASSIFY-STAND-INS"); - assert_unsupported(occ, kind, reason, &format!("CLASSIFY-STAND-INS: {file}")); - for other in &hits { - assert_ne!( - capability_token(other), - "supported", - "CLASSIFY-STAND-INS: {file} must never return supported" - ); - } - } -} - -// --- CLASSIFY-NO-SILENT-OVERLAY --------------------------------------------- - -#[test] -fn classify_stand_ins_never_supported_or_overlay_fallback() { - let manifest = load_manifest(); - let stand_in_rows: Vec<&FixtureRow> = manifest - .fixtures - .iter() - .filter(|r| r.intent == "unsupported-stand-in") - .collect(); - assert!( - !stand_in_rows.is_empty(), - "CLASSIFY-NO-SILENT-OVERLAY: manifest must list unsupported-stand-in rows" - ); - for row in stand_in_rows { - let path = fixture(&row.path); - let hits = classify(&path, "CLASSIFY-NO-SILENT-OVERLAY"); - for occ in &hits { - assert_ne!( - capability_token(occ), - "supported", - "CLASSIFY-NO-SILENT-OVERLAY: {} must not be supported", - row.id - ); - if let Some(code) = reason_code(occ) { - assert!( - !looks_like_overlay_fallback(&code), - "CLASSIFY-NO-SILENT-OVERLAY: {} reason must not be an overlay fallback; got {code}", - row.id - ); - assert!( - FROZEN_REASONS.contains(&code.as_str()), - "CLASSIFY-NO-SILENT-OVERLAY: {} reason {code} is not a frozen code", - row.id - ); - } - assert!( - !looks_like_overlay_fallback(&capability_token(occ)), - "CLASSIFY-NO-SILENT-OVERLAY: {} capability must not be an overlay fallback", - row.id - ); - } - } -} - -// --- CLASSIFY-NO-SAVE ------------------------------------------------------- - -#[test] -fn classify_does_not_write_source_or_dest() { - let path = fixture("text-tj.pdf"); - let parent = path - .parent() - .expect("CLASSIFY-NO-SAVE: fixture has a parent dir"); - let before_names = dir_names(parent); - let before_bytes = fs::read(&path).unwrap(); - let before_meta = fs::metadata(&path).unwrap(); - let before_mtime = before_meta.modified().ok(); - let before_len = before_meta.len(); - - // Entry point must not write, whether it returns Ok or Err. - let _ = classify_source_content(&path); - - let after_bytes = fs::read(&path).unwrap(); - assert_eq!( - after_bytes, before_bytes, - "CLASSIFY-NO-SAVE: source bytes of text-tj.pdf must be unchanged" - ); - let after_meta = fs::metadata(&path).unwrap(); - assert_eq!( - after_meta.len(), - before_len, - "CLASSIFY-NO-SAVE: source length must be unchanged" - ); - if let (Some(before), Ok(after)) = (before_mtime, after_meta.modified()) { - assert_eq!( - after, before, - "CLASSIFY-NO-SAVE: source mtime must be unchanged" - ); - } - let after_names = dir_names(parent); - assert_eq!( - after_names, before_names, - "CLASSIFY-NO-SAVE: no dest sibling may be created next to the fixture; before={before_names:?} after={after_names:?}" - ); -} - -// --- CLASSIFY-NO-EDITABLE-CLAIM --------------------------------------------- - -#[test] -fn classify_try_edit_is_not_auto_supported() { - let cid = classify( - &fixture("text-cid-tounicode.pdf"), - "CLASSIFY-NO-EDITABLE-CLAIM", - ); - let cid_occ = first_of_kind(&cid, "text", "CLASSIFY-NO-EDITABLE-CLAIM"); - assert_unsupported( - cid_occ, - "text", - "AMBIGUOUS_UNICODE", - "CLASSIFY-NO-EDITABLE-CLAIM: text-cid-tounicode.pdf is try-edit but must not be treated as supported", - ); - assert_ne!( - reason_code(cid_occ).as_deref(), - Some("NO_TOUNICODE"), - "CLASSIFY-NO-EDITABLE-CLAIM: text-cid-tounicode.pdf has a ToUnicode CMap; NO_TOUNICODE is the wrong reason" - ); - - let kerned = classify(&fixture("text-tj-kerned.pdf"), "CLASSIFY-NO-EDITABLE-CLAIM"); - let kerned_occ = first_of_kind(&kerned, "text", "CLASSIFY-NO-EDITABLE-CLAIM"); - assert_supported_text_or_image( - kerned_occ, - "text", - "CLASSIFY-NO-EDITABLE-CLAIM: text-tj-kerned.pdf is the human pick for supported", - ); - - let manifest = load_manifest(); - let try_edit: Vec<&FixtureRow> = manifest - .fixtures - .iter() - .filter(|r| r.intent == "try-edit") - .collect(); - assert!( - try_edit.iter().any(|r| r.id == "text-cid-tounicode"), - "CLASSIFY-NO-EDITABLE-CLAIM: manifest still marks text-cid-tounicode as try-edit" - ); - assert!( - try_edit.iter().any(|r| r.id == "text-tj-kerned"), - "CLASSIFY-NO-EDITABLE-CLAIM: manifest still marks text-tj-kerned as try-edit" - ); -} - -// --- CLASSIFY-STALE --------------------------------------------------------- - -#[test] -fn classify_mutated_copy_locator_is_stale() { - let src = fixture("text-tj.pdf"); - let hits = classify(&src, "CLASSIFY-STALE"); - let occ = first_of_kind(&hits, "text", "CLASSIFY-STALE"); - let locator = occ.locator.clone(); - assert!( - !locator.trim().is_empty(), - "CLASSIFY-STALE: locator must be a non-empty opaque string" - ); - - match resolve_source_locator(&src, &locator) { - Ok(_) => {} - Err(err) => assert_ne!( - err.code.as_str(), - "STALE", - "CLASSIFY-STALE: original source must not be STALE" - ), - } - - let scratch = Scratch::new("stale"); - let copy = scratch.file("copy.pdf"); - fs::copy(&src, ©).unwrap(); - let mut bytes = fs::read(©).unwrap(); - flip_one_payload_byte(&mut bytes); - fs::write(©, &bytes).unwrap(); - - let err = resolve_source_locator(©, &locator) - .expect_err("CLASSIFY-STALE: resolve_source_locator on a mutated copy must be Err, not Ok"); - assert_eq!( - err.code, "STALE", - "CLASSIFY-STALE: AppError.code must be STALE; got {} ({})", - err.code, err.message - ); -} - -// --- CLASSIFY-BOUNDS -------------------------------------------------------- - -#[test] -fn classify_missing_path_is_invalid_pdf() { - let scratch = Scratch::new("missing"); - let missing = scratch.file("no-such.pdf"); - expect_err_code( - classify_source_content(&missing), - "INVALID_PDF", - "CLASSIFY-BOUNDS: missing path", - ); -} - -#[test] -fn classify_broken_tiny_pdf_is_app_error() { - let scratch = Scratch::new("broken"); - let path = scratch.file("tiny.pdf"); - fs::write(&path, b"%PDF-1.4\n%% truncated").unwrap(); - expect_bounds_err( - classify_source_content(&path), - &["MALFORMED_CONTENT", "INVALID_PDF"], - "CLASSIFY-BOUNDS: broken tiny PDF", - ); -} - -#[test] -fn classify_geom_only_pages_are_empty() { - for name in [ - "geom-crop-offset.pdf", - "geom-user-unit.pdf", - "geom-rotate-90.pdf", - ] { - let hits = classify(&fixture(name), "CLASSIFY-BOUNDS"); - assert!( - hits.is_empty(), - "CLASSIFY-BOUNDS: {name} is geom-only (re f) and must return an empty list, not a fake supported/GEOMETRY row; got {:?}", - hits.iter() - .map(|o| (kind_token(o), capability_token(o), reason_code(o))) - .collect::>() - ); - } - for name in GEOM_ONLY { - let hits = classify(&fixture(name), "CLASSIFY-BOUNDS"); - assert!( - hits.iter().all(|o| capability_token(o) != "supported"), - "CLASSIFY-BOUNDS: {name} must not mark a geom-only page supported" - ); - } -} - -#[test] -fn classify_oversize_sparse_file_is_file_too_large() { - // Sparse tempfile — do not commit a 400 MiB PDF. - let scratch = Scratch::new("huge"); - let path = scratch.file("huge.pdf"); - let f = File::create(&path).unwrap(); - f.set_len(FILE_CAP_BYTES + 1).unwrap(); - drop(f); - expect_err_code( - classify_source_content(&path), - "FILE_TOO_LARGE", - "CLASSIFY-BOUNDS: set_len(400MiB+1)", - ); -} - -// --- PR 97 review fold (R1–R5) --------------------------------------------- -// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. - -fn box_obj(b: [i64; 4]) -> Object { - Object::Array(b.into_iter().map(Object::Integer).collect()) -} - -#[test] -fn classify_malformed_inline_tail_returns_error_without_panicking() { - let scratch = Scratch::new("review-malformed-inline"); - let path = scratch.file("tail.pdf"); - let mut content = b"BI /W 1 /H 1 /BPC 8 /CS /G ID x EI\n(".to_vec(); - content.push(b'\\'); - write_helvetica_page(&path, &content); - let result = std::panic::catch_unwind(|| classify_source_content(&path)); - assert!(result.is_ok(), "Malformed content must return an AppError, not panic"); - assert!(result.unwrap().is_err()); -} - -#[test] -fn classify_missing_font_is_not_supported() { - let scratch = Scratch::new("review-missing-font"); - let path = scratch.file("missing-font.pdf"); - write_helvetica_page(&path, b"BT /Missing 12 Tf 72 400 Td (Hi) Tj ET"); - let hits = classify(&path, "REVIEW-MISSING-FONT"); - assert_eq!(capability_token(&hits[0]), "unsupported"); - assert_eq!(reason_code(&hits[0]).as_deref(), Some("MISSING_FONT")); -} - -#[test] -fn classify_repeated_form_occurrences_have_unique_locators() { - let scratch = Scratch::new("review-repeated-form"); - let path = scratch.file("repeated-form.pdf"); - let mut doc = Document::load(fixture("image-in-form.pdf")).unwrap(); - let page = *doc.get_pages().values().next().unwrap(); - let ids = doc.get_page_contents(page); - let bytes = doc.get_page_content(page).unwrap(); - let mut repeated = bytes.clone(); - repeated.extend_from_slice(b"\n1 0 0 1 100 0 cm\n"); - repeated.extend_from_slice(&bytes); - doc.objects.insert(ids[0], Object::Stream(Stream::new(Dictionary::new(), repeated))); - doc.save(&path).unwrap(); - let hits = classify(&path, "REVIEW-REPEATED-FORM"); - assert!(hits.len() >= 2); - let locators: std::collections::HashSet<_> = hits.iter().map(|h| &h.locator).collect(); - assert_eq!(locators.len(), hits.len(), "Every occurrence needs its own locator"); - assert_ne!(hits[0].rect, hits[1].rect); - assert_eq!(resolve_source_locator(&path, &hits[1].locator).unwrap(), hits[1]); - assert_eq!(classify(&path, "REVIEW-REPEATED-FORM"), hits); -} - -fn helvetica_resources() -> Dictionary { - let mut font = Dictionary::new(); - font.set("Type", "Font"); - font.set("Subtype", "Type1"); - font.set("BaseFont", "Helvetica"); - let mut fonts = Dictionary::new(); - fonts.set("F1", Object::Dictionary(font)); - let mut res = Dictionary::new(); - res.set("Font", Object::Dictionary(fonts)); - res -} - -fn write_helvetica_page(path: &Path, content: &[u8]) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - content.to_vec(), - ))); - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set("Contents", content_id); - page.set("Resources", Object::Dictionary(helvetica_resources())); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path).expect("write generated classifier fixture"); -} - -fn write_text_with_empty_sig_widget(path: &Path) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET\n".to_vec(), - ))); - let widget_id = doc.new_object_id(); - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set("Contents", content_id); - page.set("Resources", Object::Dictionary(helvetica_resources())); - page.set("Annots", vec![Object::Reference(widget_id)]); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut widget = Dictionary::new(); - widget.set("Type", "Annot"); - widget.set("Subtype", "Widget"); - widget.set("FT", "Sig"); - widget.set("T", Object::string_literal("Sig1")); - widget.set("Rect", box_obj([72, 72, 172, 92])); - widget.set("P", page_id); - doc.objects.insert(widget_id, Object::Dictionary(widget)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut acro = Dictionary::new(); - acro.set("Fields", vec![Object::Reference(widget_id)]); - let acro_id = doc.add_object(Object::Dictionary(acro)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - catalog.set("AcroForm", acro_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path) - .expect("write empty-sig-widget classifier fixture"); -} - -fn write_applied_signature(path: &Path) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET\n".to_vec(), - ))); - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set("Contents", content_id); - page.set("Resources", Object::Dictionary(helvetica_resources())); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut sig = Dictionary::new(); - sig.set("Type", "Sig"); - sig.set( - "ByteRange", - vec![ - Object::Integer(0), - Object::Integer(10), - Object::Integer(20), - Object::Integer(30), - ], - ); - let _sig_id = doc.add_object(Object::Dictionary(sig)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path) - .expect("write applied-signature classifier fixture"); -} - -// --- R1 -------------------------------------------------------------------- - -#[test] -fn classify_180_degree_text_is_rotated() { - let scratch = Scratch::new("r1-180"); - let path = scratch.file("text-180.pdf"); - write_helvetica_page(&path, b"BT /F1 12 Tf -1 0 0 -1 200 400 Tm (Hi) Tj ET\n"); - let hits = classify(&path, "R1"); - let occ = first_of_kind(&hits, "text", "R1"); - assert_ne!( - capability_token(occ), - "supported", - "R1: 180° Tm [-1 0 0 -1 200 400] must not be supported" - ); - assert_unsupported(occ, "text", "ROTATED_TEXT", "R1"); -} - -// --- R2 -------------------------------------------------------------------- - -#[test] -fn classify_second_tj_advances_tm() { - let scratch = Scratch::new("r2-advance"); - let path = scratch.file("two-tj.pdf"); - write_helvetica_page(&path, b"BT /F1 12 Tf 72 720 Td (Hel) Tj (lo) Tj ET\n"); - let hits = classify(&path, "R2"); - let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); - assert_eq!( - texts.len(), - 2, - "R2: (Hel) Tj (lo) Tj must emit two text occurrences; got {:?}", - hits.iter() - .map(|o| (kind_token(o), o.rect.x, o.rect.y)) - .collect::>() - ); - assert!( - (texts[0].rect.x - 72.0).abs() <= 1.0, - "R2: first origin x must be ~72; got {}", - texts[0].rect.x - ); - assert!( - texts[1].rect.x > texts[0].rect.x, - "R2: second rect.x must be > first (must not share origin 72); first x={} second x={}", - texts[0].rect.x, - texts[1].rect.x - ); - assert!( - (texts[1].rect.x - 72.0).abs() > 1.0, - "R2: second show must not reuse origin 72; first x={} second x={}", - texts[0].rect.x, - texts[1].rect.x - ); -} - -// --- R3 -------------------------------------------------------------------- - -#[test] -fn classify_empty_sig_widget_does_not_refuse_file() { - let scratch = Scratch::new("r3-empty-sig"); - let path = scratch.file("empty-sig.pdf"); - write_text_with_empty_sig_widget(&path); - let hits = match classify_source_content(&path) { - Ok(hits) => hits, - Err(err) => panic!( - "R3: empty /FT /Sig widget (Type Annot, no ByteRange) + Helvetica (Hi) Tj must be Ok with a text occurrence, not AppError SIGNED; got {} ({})", - err.code, err.message - ), - }; - let occ = first_of_kind(&hits, "text", "R3"); - assert_ne!( - reason_code(occ).as_deref(), - Some("SIGNED"), - "R3: Helvetica text on a file with an empty Sig widget must not be SIGNED" - ); -} - -#[test] -fn classify_applied_signature_is_signed() { - let scratch = Scratch::new("r3-applied-sig"); - let path = scratch.file("applied-sig.pdf"); - write_applied_signature(&path); - expect_err_code( - classify_source_content(&path), - "SIGNED", - "R3: /Type /Sig + ByteRange still refuses", - ); -} - -// --- R4 -------------------------------------------------------------------- - -#[test] -fn classify_text_after_inline_image_is_kept() { - let scratch = Scratch::new("r4-inline-rest"); - let path = scratch.file("inline-then-text.pdf"); - write_helvetica_page( - &path, - b"q 24 0 0 12 72 400 cm\n\ -BI\n\ -/W 2 /H 1 /CS /DeviceRGB /BPC 8 /F /AHx\n\ -ID\n\ -C8101010C810>\n\ -EI\n\ -Q\n\ -BT /F1 12 Tf 72 720 Td (Hi) Tj ET\n", - ); - let hits = classify(&path, "R4"); - let _text = first_of_kind(&hits, "text", "R4"); -} - -// --- R5 -------------------------------------------------------------------- - -#[test] -fn classify_text_bounds_use_tm_scale() { - let scratch = Scratch::new("r5-tm-scale"); - let path = scratch.file("tf1-tm12.pdf"); - write_helvetica_page(&path, b"BT /F1 1 Tf 12 0 0 12 72 720 Tm (Hi) Tj ET\n"); - let hits = classify(&path, "R5"); - let occ = first_of_kind(&hits, "text", "R5"); - assert!( - (occ.rect.h - 12.0).abs() <= 1.0, - "R5: /F1 1 Tf + 12 0 0 12 Tm must report height ~12, not ~1; got h={}", - occ.rect.h - ); - assert!( - (occ.rect.h - 1.0).abs() > 1.0, - "R5: rect.h must not stay at Tf size ~1; got h={}", - occ.rect.h - ); -} - -// --- PR 97 review fold r2 (R6–R9) ------------------------------------------ -// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. - -fn write_type3_and_helvetica_page(path: &Path, content: &[u8]) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - content.to_vec(), - ))); - - let proc_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - b"10 0 0 0 10 10 d1\n0 0 10 10 re f\n".to_vec(), - ))); - let mut char_procs = Dictionary::new(); - char_procs.set("x", proc_id); - - let mut enc = Dictionary::new(); - enc.set("Type", "Encoding"); - enc.set( - "Differences", - vec![Object::Integer(120), Object::Name(b"x".to_vec())], - ); - - let mut t3 = Dictionary::new(); - t3.set("Type", "Font"); - t3.set("Subtype", "Type3"); - t3.set("FontBBox", box_obj([0, 0, 10, 10])); - t3.set( - "FontMatrix", - Object::Array(vec![ - Object::Real(1.0), - Object::Real(0.0), - Object::Real(0.0), - Object::Real(1.0), - Object::Real(0.0), - Object::Real(0.0), - ]), - ); - t3.set("CharProcs", Object::Dictionary(char_procs)); - t3.set("Encoding", Object::Dictionary(enc)); - t3.set("FirstChar", 120); - t3.set("LastChar", 120); - t3.set("Widths", vec![Object::Integer(10)]); - - let mut f1 = Dictionary::new(); - f1.set("Type", "Font"); - f1.set("Subtype", "Type1"); - f1.set("BaseFont", "Helvetica"); - - let mut fonts = Dictionary::new(); - fonts.set("T3", Object::Dictionary(t3)); - fonts.set("F1", Object::Dictionary(f1)); - let mut res = Dictionary::new(); - res.set("Font", Object::Dictionary(fonts)); - - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set("Contents", content_id); - page.set("Resources", Object::Dictionary(res)); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path) - .expect("write Type3+Helvetica classifier fixture"); -} - -// --- R6 -------------------------------------------------------------------- - -#[test] -fn classify_inline_image_uses_ctm_at_bi() { - let path = fixture("image-inline.pdf"); - let hits = classify(&path, "R6"); - let occ = first_of_kind(&hits, "image", "R6"); - assert!( - (occ.rect.x - 72.0).abs() <= 1.0, - "R6: image-inline.pdf image rect.x must be ~72 (CTM at BI), not the unit square at origin; got x={}", - occ.rect.x - ); - assert!( - (occ.rect.y - 400.0).abs() <= 1.0, - "R6: image-inline.pdf image rect.y must be ~400 (CTM at BI), not the unit square at origin; got y={}", - occ.rect.y - ); - assert!( - (occ.rect.w - 24.0).abs() <= 1.0, - "R6: image-inline.pdf image rect.w must be ~24 (CTM at BI), not the unit square; got w={}", - occ.rect.w - ); - assert!( - (occ.rect.h - 12.0).abs() <= 1.0, - "R6: image-inline.pdf image rect.h must be ~12 (CTM at BI), not the unit square; got h={}", - occ.rect.h - ); -} - -// --- R7 -------------------------------------------------------------------- - -#[test] -fn classify_q_restores_type3_after_helvetica() { - let scratch = Scratch::new("r7-q-type3"); - let path = scratch.file("q-type3.pdf"); - write_type3_and_helvetica_page( - &path, - b"BT /T3 12 Tf (x) Tj q /F1 12 Tf (y) Tj Q (z) Tj ET\n", - ); - let hits = classify(&path, "R7"); - let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); - assert_eq!( - texts.len(), - 3, - "R7: (x) Tj q /F1 (y) Tj Q (z) Tj must emit three text occurrences; got {:?}", - hits.iter() - .map(|o| (kind_token(o), capability_token(o), reason_code(o))) - .collect::>() - ); - let last = texts[2]; - assert_ne!( - capability_token(last), - "supported", - "R7: last show (z) after Q must not be Helvetica supported; got {} reason={:?}", - capability_token(last), - reason_code(last) - ); - assert_unsupported(last, "text", "TYPE3", "R7"); -} - -// --- R8 -------------------------------------------------------------------- - -#[test] -fn classify_tc_advances_second_tj() { - let scratch = Scratch::new("r8-tc"); - let path = scratch.file("tc-two-tj.pdf"); - write_helvetica_page(&path, b"BT /F1 12 Tf 2 Tc 72 720 Td (Hi) Tj (there) Tj ET\n"); - let hits = classify(&path, "R8"); - let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); - assert_eq!( - texts.len(), - 2, - "R8: (Hi) Tj (there) Tj must emit two text occurrences; got {:?}", - hits.iter() - .map(|o| (kind_token(o), o.rect.x, o.rect.y)) - .collect::>() - ); - // Helvetica H=667 i=278 → 11.34 at Tf=12. 2 Tc on two glyphs adds 4 - // user units, so second.x ≈ first.x + 15.34, not first.x + 11.34. - assert!( - texts[1].rect.x > texts[0].rect.x + 13.0, - "R8: 2 Tc must push second.x past first.x + no-Tc Hi width 11.34; first.x={} second.x={} (need second.x > first.x + 13)", - texts[0].rect.x, - texts[1].rect.x - ); -} - -// --- R9 -------------------------------------------------------------------- - -#[test] -fn classify_source_drops_rotated_fixture_parenthetical() { - let path = Path::new(env!("CARGO_MANIFEST_DIR")).join("src/pdf_engine/source_content.rs"); - let src = fs::read_to_string(&path).unwrap_or_else(|e| { - panic!( - "R9: must read source_content.rs via CARGO_MANIFEST_DIR ({}): {e}", - path.display() - ) - }); - assert!( - !src.contains("keeps text-rotated.pdf green"), - "R9: source_content.rs must not contain the exact substring `keeps text-rotated.pdf green`" - ); -} - -// --- PR 97 review fold r3 (R10–R12) ---------------------------------------- -// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. - -/// 2×2 DeviceRGB, same bytes as the #32 unique/mask fixtures. -const R3_TINY_RGB: &[u8] = &[200, 16, 16, 16, 200, 16, 16, 16, 200, 200, 200, 16]; - -fn write_type1_indirect_widths(path: &Path, content: &[u8]) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - content.to_vec(), - ))); - - // FirstChar 'H' (72) … LastChar 'i' (105): 34 glyph slots, all 1000. - const FIRST_CHAR: i64 = 72; - const LAST_CHAR: i64 = 105; - let widths: Vec = (FIRST_CHAR..=LAST_CHAR) - .map(|_| Object::Integer(1000)) - .collect(); - let widths_id = doc.add_object(Object::Array(widths)); - - let mut font = Dictionary::new(); - font.set("Type", "Font"); - font.set("Subtype", "Type1"); - font.set("BaseFont", "Helvetica"); - font.set("FirstChar", FIRST_CHAR); - font.set("LastChar", LAST_CHAR); - font.set("Widths", Object::Reference(widths_id)); - - let mut fonts = Dictionary::new(); - fonts.set("F1", Object::Dictionary(font)); - let mut res = Dictionary::new(); - res.set("Font", Object::Dictionary(fonts)); - - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set("Contents", content_id); - page.set("Resources", Object::Dictionary(res)); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path) - .expect("write Type1 indirect-Widths classifier fixture"); -} - -fn write_pattern_cs_page(path: &Path, content: &[u8]) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - content.to_vec(), - ))); - - let mut font = Dictionary::new(); - font.set("Type", "Font"); - font.set("Subtype", "Type1"); - font.set("BaseFont", "Helvetica"); - let mut fonts = Dictionary::new(); - fonts.set("F1", Object::Dictionary(font)); - - let mut cs = Dictionary::new(); - cs.set("Cs1", Object::Name(b"Pattern".to_vec())); - - let mut pat = Dictionary::new(); - pat.set("Type", "Pattern"); - pat.set("PatternType", 1); - pat.set("PaintType", 1); - pat.set("TilingType", 1); - pat.set("BBox", box_obj([0, 0, 10, 10])); - pat.set("XStep", 10); - pat.set("YStep", 10); - pat.set("Resources", Object::Dictionary(Dictionary::new())); - let pat_id = doc.add_object(Object::Stream(Stream::new( - pat, - b"0 0 10 10 re f\n".to_vec(), - ))); - let mut patterns = Dictionary::new(); - patterns.set("P1", Object::Reference(pat_id)); - - let mut res = Dictionary::new(); - res.set("Font", Object::Dictionary(fonts)); - res.set("ColorSpace", Object::Dictionary(cs)); - res.set("Pattern", Object::Dictionary(patterns)); - - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set("Contents", content_id); - page.set("Resources", Object::Dictionary(res)); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path) - .expect("write Pattern ColorSpace classifier fixture"); -} - -fn write_extgstate_smask_image(path: &Path, content: &[u8]) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - content.to_vec(), - ))); - - let mut img = Dictionary::new(); - img.set("Type", "XObject"); - img.set("Subtype", "Image"); - img.set("Width", 2); - img.set("Height", 2); - img.set("ColorSpace", "DeviceRGB"); - img.set("BitsPerComponent", 8); - let img_id = doc.add_object(Object::Stream(Stream::new(img, R3_TINY_RGB.to_vec()))); - - let mut sm = Dictionary::new(); - sm.set("Type", "XObject"); - sm.set("Subtype", "Image"); - sm.set("Width", 2); - sm.set("Height", 2); - sm.set("ColorSpace", "DeviceGray"); - sm.set("BitsPerComponent", 8); - let smask_id = doc.add_object(Object::Stream(Stream::new(sm, vec![255, 200, 180, 255]))); - - let mut gs = Dictionary::new(); - gs.set("Type", "ExtGState"); - gs.set("SMask", Object::Reference(smask_id)); - let gs_id = doc.add_object(Object::Dictionary(gs)); - - let mut xobjects = Dictionary::new(); - xobjects.set("Im0", Object::Reference(img_id)); - let mut extg = Dictionary::new(); - extg.set("Gs1", Object::Reference(gs_id)); - let mut res = Dictionary::new(); - res.set("XObject", Object::Dictionary(xobjects)); - res.set("ExtGState", Object::Dictionary(extg)); - - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set("Contents", content_id); - page.set("Resources", Object::Dictionary(res)); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path) - .expect("write ExtGState SMask + unique Image classifier fixture"); -} - -// --- R10 ------------------------------------------------------------------- - -#[test] -fn classify_indirect_widths_not_helvetica_fallback() { - let scratch = Scratch::new("r10-widths"); - let path = scratch.file("indirect-widths.pdf"); - write_type1_indirect_widths(&path, b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET\n"); - let hits = classify(&path, "R10"); - let occ = first_of_kind(&hits, "text", "R10"); - assert!( - (occ.rect.w - 24.0).abs() <= 1.0, - "R10: Type1 indirect /Widths 1000,1000 at Tf=12 must report w≈24, not Helvetica fallback ≈11.34; got w={}", - occ.rect.w - ); - assert!( - (occ.rect.w - 11.34).abs() > 1.0, - "R10: rect.w must not stay on the Helvetica table ≈11.34; got w={}", - occ.rect.w - ); -} - -// --- R11 ------------------------------------------------------------------- - -#[test] -fn classify_named_pattern_cs_is_unsupported() { - let scratch = Scratch::new("r11-pattern"); - let path = scratch.file("pattern-cs.pdf"); - write_pattern_cs_page( - &path, - b"BT /F1 12 Tf /Cs1 cs /P1 scn 72 720 Td (Hi) Tj ET\n", - ); - let hits = classify(&path, "R11"); - let occ = first_of_kind(&hits, "text", "R11"); - assert_ne!( - capability_token(occ), - "supported", - "R11: /Cs1 cs Pattern resource + (Hi) Tj must not be supported; got {} reason={:?}", - capability_token(occ), - reason_code(occ) - ); - assert_unsupported(occ, "text", "PATTERN", "R11"); -} - -// --- R12 ------------------------------------------------------------------- - -#[test] -fn classify_extgstate_smask_image_is_masked() { - let scratch = Scratch::new("r12-gs-smask"); - let path = scratch.file("gs-smask.pdf"); - write_extgstate_smask_image(&path, b"q 40 0 0 40 72 400 cm /Gs1 gs /Im0 Do Q\n"); - let hits = classify(&path, "R12"); - let occ = first_of_kind(&hits, "image", "R12"); - assert_ne!( - capability_token(occ), - "supported", - "R12: unique Image after ExtGState /Gs1 /SMask must not be supported; got {} reason={:?}", - capability_token(occ), - reason_code(occ) - ); - assert_unsupported(occ, "image", "MASKED_IMAGE", "R12"); -} - -// --- PR 97 review fold r4 (R13) -------------------------------------------- -// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. - -fn write_unique_rgb_image(path: &Path, content: &[u8]) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content_id = doc.add_object(Object::Stream(Stream::new( - Dictionary::new(), - content.to_vec(), - ))); - - let mut img = Dictionary::new(); - img.set("Type", "XObject"); - img.set("Subtype", "Image"); - img.set("Width", 2); - img.set("Height", 2); - img.set("ColorSpace", "DeviceRGB"); - img.set("BitsPerComponent", 8); - let img_id = doc.add_object(Object::Stream(Stream::new(img, R3_TINY_RGB.to_vec()))); - - let mut xobjects = Dictionary::new(); - xobjects.set("Im0", Object::Reference(img_id)); - let mut res = Dictionary::new(); - res.set("XObject", Object::Dictionary(xobjects)); - - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set("Contents", content_id); - page.set("Resources", Object::Dictionary(res)); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path) - .expect("write unique 2×2 DeviceRGB Image classifier fixture"); -} - -// --- R13a ------------------------------------------------------------------ - -#[test] -fn classify_stacked_cm_image_origin() { - let scratch = Scratch::new("r13a-stacked-cm"); - let path = scratch.file("stacked-cm.pdf"); - write_unique_rgb_image( - &path, - b"q 2 0 0 2 0 0 cm 20 0 0 20 36 200 cm /Im0 Do Q\n", - ); - let hits = classify(&path, "R13a"); - let occ = first_of_kind(&hits, "image", "R13a"); - assert_supported_text_or_image(occ, "image", "R13a"); - assert!( - (occ.rect.x - 72.0).abs() <= 1.0 - && (occ.rect.y - 400.0).abs() <= 1.0 - && (occ.rect.w - 40.0).abs() <= 1.0 - && (occ.rect.h - 40.0).abs() <= 1.0, - "R13a: stacked cm image rect must be ~{{x:72, y:400, w:40, h:40}}, not origin ~(36, 200); got {{x:{}, y:{}, w:{}, h:{}}}", - occ.rect.x, - occ.rect.y, - occ.rect.w, - occ.rect.h - ); - assert!( - (occ.rect.x - 36.0).abs() > 1.0 || (occ.rect.y - 200.0).abs() > 1.0, - "R13a: stacked cm must not leave the image at the second-cm translation (36, 200); got {{x:{}, y:{}, w:{}, h:{}}}", - occ.rect.x, - occ.rect.y, - occ.rect.w, - occ.rect.h - ); -} - -// --- R13b ------------------------------------------------------------------ - -#[test] -fn classify_scaled_tm_second_show_x() { - let scratch = Scratch::new("r13b-scaled-tm"); - let path = scratch.file("scaled-tm-two-tj.pdf"); - write_helvetica_page( - &path, - b"BT /F1 1 Tf 12 0 0 12 72 720 Tm (Hel) Tj (lo) Tj ET\n", - ); - let hits = classify(&path, "R13b"); - let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); - assert_eq!( - texts.len(), - 2, - "R13b: (Hel) Tj (lo) Tj must emit two text occurrences; got {:?}", - hits.iter() - .map(|o| (kind_token(o), o.rect.x, o.rect.y)) - .collect::>() - ); - let first = texts[0]; - let second = texts[1]; - // Helvetica H=667 e=556 l=278 → 1.501 at Tf=1. Scaled Tm 12× must - // advance ~18.012 user units → second.x ≈ 90, not text-space 1.501 - // added in user space (≈73.5). - assert!( - second.rect.x > first.rect.x + 15.0 || (second.rect.x - 90.0).abs() <= 2.0, - "R13b: second rect.x after 12 0 0 12 72 720 Tm (Hel) Tj must be ≈90 (±2), not ≈73.5; first.x={} second.x={}", - first.rect.x, - second.rect.x - ); - assert!( - (second.rect.x - 73.5).abs() > 1.0, - "R13b: second rect.x must not stay at origin+text-space width ≈73.5; first.x={} second.x={}", - first.rect.x, - second.rect.x - ); -} - -// --- PR 97 review fold r5 (R14) -------------------------------------------- -// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. -// Page /Contents is an array of two streams; stream 1 has no trailing whitespace -// so a join without a separator fuses `Tj`+`ET` into `TjET`. - -const R14_STREAM_1: &[u8] = b"BT /F1 12 Tf 72 720 Td (Hi) Tj"; -const R14_STREAM_2: &[u8] = b"ET\nBT /F1 12 Tf 72 680 Td (Lo) Tj ET"; - -fn stream_content_bytes(doc: &Document, obj: &Object) -> Vec { - let id = obj - .as_reference() - .expect("Contents array entry must be a stream ref"); - doc.get_object(id) - .expect("content stream object") - .as_stream() - .expect("content must be a stream") - .content - .clone() -} - -fn write_helvetica_two_content_streams(path: &Path, stream1: &[u8], stream2: &[u8]) { - let mut doc = Document::with_version("1.7"); - let pages_id = doc.new_object_id(); - let content1_id = doc.add_object(Object::Stream( - Stream::new(Dictionary::new(), stream1.to_vec()).with_compression(false), - )); - let content2_id = doc.add_object(Object::Stream( - Stream::new(Dictionary::new(), stream2.to_vec()).with_compression(false), - )); - let mut page = Dictionary::new(); - page.set("Type", "Page"); - page.set("Parent", pages_id); - page.set("MediaBox", box_obj([0, 0, 612, 792])); - page.set( - "Contents", - vec![ - Object::Reference(content1_id), - Object::Reference(content2_id), - ], - ); - page.set("Resources", Object::Dictionary(helvetica_resources())); - let page_id = doc.add_object(Object::Dictionary(page)); - - let mut pages = Dictionary::new(); - pages.set("Type", "Pages"); - pages.set("Kids", vec![page_id.into()]); - pages.set("Count", 1); - doc.objects.insert(pages_id, Object::Dictionary(pages)); - - let mut catalog = Dictionary::new(); - catalog.set("Type", "Catalog"); - catalog.set("Pages", pages_id); - let catalog_id = doc.add_object(Object::Dictionary(catalog)); - doc.trailer.set("Root", catalog_id); - doc.save(path) - .expect("write two-stream Contents classifier fixture"); - - // Lock the on-disk page /Contents shape: array of two streams, exact bytes. - let reloaded = - Document::load(path).unwrap_or_else(|e| panic!("reload two-stream Contents fixture: {e}")); - let page_id = *reloaded - .get_pages() - .get(&1) - .expect("two-stream fixture must have page 1"); - let page = reloaded - .get_object(page_id) - .expect("page 1 object") - .as_dict() - .expect("page 1 dict"); - let contents = page.get(b"Contents").expect("page /Contents"); - let refs = match contents { - Object::Array(arr) => arr, - other => panic!( - "two-stream fixture page /Contents must be an array of two stream refs, got {other:?}" - ), - }; - assert_eq!( - refs.len(), - 2, - "two-stream fixture page /Contents must have two stream refs; got {}", - refs.len() - ); - assert_eq!( - stream_content_bytes(&reloaded, &refs[0]).as_slice(), - stream1, - "two-stream fixture stream 1 bytes must be exact (no trailing newline)" - ); - assert_eq!( - stream_content_bytes(&reloaded, &refs[1]).as_slice(), - stream2, - "two-stream fixture stream 2 bytes must match" - ); -} - -// --- R14 ------------------------------------------------------------------- - -#[test] -fn classify_contents_array_two_streams_do_not_fuse() { - let scratch = Scratch::new("r14-two-streams"); - let path = scratch.file("two-contents-streams.pdf"); - write_helvetica_two_content_streams(&path, R14_STREAM_1, R14_STREAM_2); - let hits = classify(&path, "R14"); - let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); - assert_eq!( - texts.len(), - 2, - "R14: two Contents streams (Hi@720 then Lo@680) must emit two text occurrences; got {:?}", - hits.iter() - .map(|o| (kind_token(o), o.rect.x, o.rect.y)) - .collect::>() - ); - assert!( - texts.iter().any(|o| (o.rect.y - 720.0).abs() <= 1.0), - "R14: expected a text occurrence at y≈720 (Hi); got {:?}", - texts.iter().map(|o| (o.rect.x, o.rect.y)).collect::>() - ); - assert!( - texts.iter().any(|o| (o.rect.y - 680.0).abs() <= 1.0), - "R14: expected a text occurrence at y≈680 (Lo); got {:?}", - texts.iter().map(|o| (o.rect.x, o.rect.y)).collect::>() - ); -} diff --git a/src-tauri/src/pdf_engine/source_content_integ/corpus.rs b/src-tauri/src/pdf_engine/source_content_integ/corpus.rs new file mode 100644 index 0000000..c2f79fe --- /dev/null +++ b/src-tauri/src/pdf_engine/source_content_integ/corpus.rs @@ -0,0 +1,392 @@ +//! Corpus contract tests (#32 fixtures, `fixtures/source-edit/`): CLASSIFY-API … CLASSIFY-BOUNDS. + +use super::*; + +// --- CLASSIFY-API ----------------------------------------------------------- + +#[test] +fn classify_api_lists_occurrences() { + let path = fixture("text-tj.pdf"); + let hits = classify(&path, "CLASSIFY-API"); + assert!( + !hits.is_empty(), + "CLASSIFY-API: classify_source_content(text-tj.pdf) must list occurrences" + ); +} + +// --- CLASSIFY-SOURCE-PATH --------------------------------------------------- + +#[test] +fn classify_uses_corpus_source_path_not_page_pdf() { + let path = fixture("text-tj.pdf"); + let rendered = path.to_string_lossy(); + assert!( + rendered.contains("fixtures/source-edit") && rendered.ends_with("text-tj.pdf"), + "CLASSIFY-SOURCE-PATH: must pass the corpus source path, not pagePdf / --empty --pages; got {}", + path.display() + ); + // Do not spawn qpdf --empty --pages and do not call page_pdf_b64. + let hits = classify(&path, "CLASSIFY-SOURCE-PATH"); + let occ = first_of_kind(&hits, "text", "CLASSIFY-SOURCE-PATH"); + assert_eq!( + occ.page_index, 0, + "CLASSIFY-SOURCE-PATH: text-tj.pdf page_index is 0 on the original source" + ); +} + +// --- CLASSIFY-TRY-EDIT-TJ --------------------------------------------------- + +#[test] +fn classify_text_tj_supported_at_origin() { + let path = fixture("text-tj.pdf"); + let hits = classify(&path, "CLASSIFY-TRY-EDIT-TJ"); + let occ = first_of_kind(&hits, "text", "CLASSIFY-TRY-EDIT-TJ"); + assert_supported_text_or_image(occ, "text", "CLASSIFY-TRY-EDIT-TJ"); + assert_eq!( + occ.page_index, 0, + "CLASSIFY-TRY-EDIT-TJ: page_index must be 0" + ); + assert!( + (occ.rect.x - 72.0).abs() <= 1.0, + "CLASSIFY-TRY-EDIT-TJ: origin x must be ~72; got {}", + occ.rect.x + ); + assert!( + (occ.rect.y - 720.0).abs() <= 1.0, + "CLASSIFY-TRY-EDIT-TJ: origin y must be ~720; got {}", + occ.rect.y + ); + assert!( + occ.rect.w > 0.0 && occ.rect.h > 0.0, + "CLASSIFY-TRY-EDIT-TJ: w/h must be positive; got w={} h={}", + occ.rect.w, + occ.rect.h + ); +} + +// --- CLASSIFY-TRY-EDIT-IMAGE ------------------------------------------------ + +#[test] +fn classify_image_unique_supported() { + let path = fixture("image-unique.pdf"); + let hits = classify(&path, "CLASSIFY-TRY-EDIT-IMAGE"); + let occ = first_of_kind(&hits, "image", "CLASSIFY-TRY-EDIT-IMAGE"); + assert_supported_text_or_image(occ, "image", "CLASSIFY-TRY-EDIT-IMAGE"); + assert_eq!( + occ.page_index, 0, + "CLASSIFY-TRY-EDIT-IMAGE: page_index must be 0" + ); + assert!( + (occ.rect.x - 72.0).abs() <= 1.0, + "CLASSIFY-TRY-EDIT-IMAGE: x must be ~72; got {}", + occ.rect.x + ); + assert!( + (occ.rect.y - 400.0).abs() <= 1.0, + "CLASSIFY-TRY-EDIT-IMAGE: y must be ~400; got {}", + occ.rect.y + ); + assert!( + (occ.rect.w - 40.0).abs() <= 1.0, + "CLASSIFY-TRY-EDIT-IMAGE: w must be ~40; got {}", + occ.rect.w + ); + assert!( + (occ.rect.h - 40.0).abs() <= 1.0, + "CLASSIFY-TRY-EDIT-IMAGE: h must be ~40; got {}", + occ.rect.h + ); +} + +// --- CLASSIFY-NO-TOUNICODE -------------------------------------------------- + +#[test] +fn classify_cid_no_tounicode_is_unsupported() { + let path = fixture("text-cid-no-tounicode.pdf"); + let hits = classify(&path, "CLASSIFY-NO-TOUNICODE"); + let occ = first_of_kind(&hits, "text", "CLASSIFY-NO-TOUNICODE"); + // Deliberate change (§C, A.10): the stand-in's CIDFontType0 has neither a ToUnicode nor a + // program; FONT_NOT_EMBEDDED (priority 20) comes before NO_TOUNICODE (25). NO_TOUNICODE on its + // own (an embedded font without ToUnicode) is CLS-H12b. + assert_unsupported(occ, "text", "FONT_NOT_EMBEDDED", "CLASSIFY-NO-TOUNICODE"); +} + +// --- CLASSIFY-SHARED-IMAGE -------------------------------------------------- + +#[test] +fn classify_image_reused_shared_xobject() { + let path = fixture("image-reused.pdf"); + let hits = classify(&path, "CLASSIFY-SHARED-IMAGE"); + let images: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "image").collect(); + assert_eq!( + images.len(), + 2, + "CLASSIFY-SHARED-IMAGE: image-reused.pdf must yield two image occurrences; got {:?}", + hits.iter() + .map(|o| ( + o.page_index, + kind_token(o), + capability_token(o), + reason_code(o) + )) + .collect::>() + ); + for occ in images { + assert_unsupported(occ, "image", "SHARED_XOBJECT", "CLASSIFY-SHARED-IMAGE"); + } +} + +// --- CLASSIFY-STAND-INS ----------------------------------------------------- + +#[test] +fn classify_stand_ins_locked_reasons() { + for (file, kind, reason) in STAND_INS { + let path = fixture(file); + let hits = classify(&path, "CLASSIFY-STAND-INS"); + let occ = first_of_kind(&hits, kind, "CLASSIFY-STAND-INS"); + assert_unsupported(occ, kind, reason, &format!("CLASSIFY-STAND-INS: {file}")); + for other in &hits { + assert_ne!( + capability_token(other), + "supported", + "CLASSIFY-STAND-INS: {file} must never return supported" + ); + } + } +} + +// --- CLASSIFY-NO-SILENT-OVERLAY --------------------------------------------- + +#[test] +fn classify_stand_ins_never_supported_or_overlay_fallback() { + let manifest = load_manifest(); + let stand_in_rows: Vec<&FixtureRow> = manifest + .fixtures + .iter() + .filter(|r| r.intent == "unsupported-stand-in") + .collect(); + assert!( + !stand_in_rows.is_empty(), + "CLASSIFY-NO-SILENT-OVERLAY: manifest must list unsupported-stand-in rows" + ); + for row in stand_in_rows { + let path = fixture(&row.path); + let hits = classify(&path, "CLASSIFY-NO-SILENT-OVERLAY"); + for occ in &hits { + assert_ne!( + capability_token(occ), + "supported", + "CLASSIFY-NO-SILENT-OVERLAY: {} must not be supported", + row.id + ); + if let Some(code) = reason_code(occ) { + assert!( + !looks_like_overlay_fallback(&code), + "CLASSIFY-NO-SILENT-OVERLAY: {} reason must not be an overlay fallback; got {code}", + row.id + ); + assert!( + frozen_reasons().contains(&code), + "CLASSIFY-NO-SILENT-OVERLAY: {} reason {code} is not a frozen code", + row.id + ); + } + assert!( + !looks_like_overlay_fallback(&capability_token(occ)), + "CLASSIFY-NO-SILENT-OVERLAY: {} capability must not be an overlay fallback", + row.id + ); + } + } +} + +// --- CLASSIFY-NO-SAVE ------------------------------------------------------- + +#[test] +fn classify_does_not_write_source_or_dest() { + let path = fixture("text-tj.pdf"); + let parent = path + .parent() + .expect("CLASSIFY-NO-SAVE: fixture has a parent dir"); + let before_names = dir_names(parent); + let before_bytes = fs::read(&path).unwrap(); + let before_meta = fs::metadata(&path).unwrap(); + let before_mtime = before_meta.modified().ok(); + let before_len = before_meta.len(); + + // Entry point must not write, whether it returns Ok or Err. + let _ = classify_source_content(&path); + + let after_bytes = fs::read(&path).unwrap(); + assert_eq!( + after_bytes, before_bytes, + "CLASSIFY-NO-SAVE: source bytes of text-tj.pdf must be unchanged" + ); + let after_meta = fs::metadata(&path).unwrap(); + assert_eq!( + after_meta.len(), + before_len, + "CLASSIFY-NO-SAVE: source length must be unchanged" + ); + if let (Some(before), Ok(after)) = (before_mtime, after_meta.modified()) { + assert_eq!( + after, before, + "CLASSIFY-NO-SAVE: source mtime must be unchanged" + ); + } + let after_names = dir_names(parent); + assert_eq!( + after_names, before_names, + "CLASSIFY-NO-SAVE: no dest sibling may be created next to the fixture; before={before_names:?} after={after_names:?}" + ); +} + +// --- CLASSIFY-NO-EDITABLE-CLAIM --------------------------------------------- + +#[test] +fn classify_try_edit_is_not_auto_supported() { + let cid = classify( + &fixture("text-cid-tounicode.pdf"), + "CLASSIFY-NO-EDITABLE-CLAIM", + ); + let cid_occ = first_of_kind(&cid, "text", "CLASSIFY-NO-EDITABLE-CLAIM"); + // Deliberate change (§C, D16): a Type0 font with a good ToUnicode is editable now, but this + // one has no program, so glyph presence cannot be proven: FONT_NOT_EMBEDDED. + assert_unsupported( + cid_occ, + "text", + "FONT_NOT_EMBEDDED", + "CLASSIFY-NO-EDITABLE-CLAIM: text-cid-tounicode.pdf is try-edit but must not be treated as supported", + ); + assert_ne!( + reason_code(cid_occ).as_deref(), + Some("NO_TOUNICODE"), + "CLASSIFY-NO-EDITABLE-CLAIM: text-cid-tounicode.pdf has a ToUnicode CMap; NO_TOUNICODE is the wrong reason" + ); + + let kerned = classify(&fixture("text-tj-kerned.pdf"), "CLASSIFY-NO-EDITABLE-CLAIM"); + let kerned_occ = first_of_kind(&kerned, "text", "CLASSIFY-NO-EDITABLE-CLAIM"); + assert_supported_text_or_image( + kerned_occ, + "text", + "CLASSIFY-NO-EDITABLE-CLAIM: text-tj-kerned.pdf is the human pick for supported", + ); + + let manifest = load_manifest(); + let try_edit: Vec<&FixtureRow> = manifest + .fixtures + .iter() + .filter(|r| r.intent == "try-edit") + .collect(); + assert!( + try_edit.iter().any(|r| r.id == "text-cid-tounicode"), + "CLASSIFY-NO-EDITABLE-CLAIM: manifest still marks text-cid-tounicode as try-edit" + ); + assert!( + try_edit.iter().any(|r| r.id == "text-tj-kerned"), + "CLASSIFY-NO-EDITABLE-CLAIM: manifest still marks text-tj-kerned as try-edit" + ); +} + +// --- CLASSIFY-STALE --------------------------------------------------------- + +#[test] +fn classify_mutated_copy_locator_is_stale() { + let src = fixture("text-tj.pdf"); + let hits = classify(&src, "CLASSIFY-STALE"); + let occ = first_of_kind(&hits, "text", "CLASSIFY-STALE"); + let locator = occ.locator.clone(); + assert!( + !locator.trim().is_empty(), + "CLASSIFY-STALE: locator must be a non-empty opaque string" + ); + + match resolve_source_locator(&src, &locator) { + Ok(_) => {} + Err(err) => assert_ne!( + err.code.as_str(), + "STALE", + "CLASSIFY-STALE: original source must not be STALE" + ), + } + + let scratch = Scratch::new("stale"); + let copy = scratch.file("copy.pdf"); + fs::copy(&src, ©).unwrap(); + let mut bytes = fs::read(©).unwrap(); + flip_one_payload_byte(&mut bytes); + fs::write(©, &bytes).unwrap(); + + let err = resolve_source_locator(©, &locator) + .expect_err("CLASSIFY-STALE: resolve_source_locator on a mutated copy must be Err, not Ok"); + assert_eq!( + err.code, "STALE", + "CLASSIFY-STALE: AppError.code must be STALE; got {} ({})", + err.code, err.message + ); +} + +// --- CLASSIFY-BOUNDS -------------------------------------------------------- + +#[test] +fn classify_missing_path_is_invalid_pdf() { + let scratch = Scratch::new("missing"); + let missing = scratch.file("no-such.pdf"); + expect_err_code( + classify_source_content(&missing), + "INVALID_PDF", + "CLASSIFY-BOUNDS: missing path", + ); +} + +#[test] +fn classify_broken_tiny_pdf_is_app_error() { + let scratch = Scratch::new("broken"); + let path = scratch.file("tiny.pdf"); + fs::write(&path, b"%PDF-1.4\n%% truncated").unwrap(); + expect_bounds_err( + classify_source_content(&path), + &["MALFORMED_CONTENT", "INVALID_PDF"], + "CLASSIFY-BOUNDS: broken tiny PDF", + ); +} + +#[test] +fn classify_geom_only_pages_are_empty() { + for name in [ + "geom-crop-offset.pdf", + "geom-user-unit.pdf", + "geom-rotate-90.pdf", + ] { + let hits = classify(&fixture(name), "CLASSIFY-BOUNDS"); + assert!( + hits.is_empty(), + "CLASSIFY-BOUNDS: {name} is geom-only (re f) and must return an empty list, not a fake supported/GEOMETRY row; got {:?}", + hits.iter() + .map(|o| (kind_token(o), capability_token(o), reason_code(o))) + .collect::>() + ); + } + for name in GEOM_ONLY { + let hits = classify(&fixture(name), "CLASSIFY-BOUNDS"); + assert!( + hits.iter().all(|o| capability_token(o) != "supported"), + "CLASSIFY-BOUNDS: {name} must not mark a geom-only page supported" + ); + } +} + +#[test] +fn classify_oversize_sparse_file_is_file_too_large() { + // Sparse tempfile — do not commit a 400 MiB PDF. + let scratch = Scratch::new("huge"); + let path = scratch.file("huge.pdf"); + let f = File::create(&path).unwrap(); + f.set_len(FILE_CAP_BYTES + 1).unwrap(); + drop(f); + expect_err_code( + classify_source_content(&path), + "FILE_TOO_LARGE", + "CLASSIFY-BOUNDS: set_len(400MiB+1)", + ); +} diff --git a/src-tauri/src/pdf_engine/source_content_integ/hardening.rs b/src-tauri/src/pdf_engine/source_content_integ/hardening.rs new file mode 100644 index 0000000..843af4d --- /dev/null +++ b/src-tauri/src/pdf_engine/source_content_integ/hardening.rs @@ -0,0 +1,455 @@ +//! CLS-H01…H14 (SPEC §C rows 1–15): the #33 maintainer items as regressions — bounded read, +//! object loading and decoding, no raw fallback, per-page budgets, the `q` stack, one snapshot, +//! custom-encoded fonts, rise and render modes, spacing in bounds. + +use super::*; +use crate::pdf_engine::source_content::{classify_source_page, SourcePageResult}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::snapshot::{fnv1a_u64, read_snapshot, snapshot_from_bytes}; +use crate::pdf_engine::text_edit::testkit::fonts::{ + add_simple, latin_truetype, Program, SimpleFont, +}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, cid_font, cid_hex, helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use std::io::Write; + +fn write(scratch: &Scratch, name: &str, bytes: &[u8]) -> PathBuf { + let p = scratch.file(name); + fs::write(&p, bytes).unwrap(); + p +} + +fn page_result(bytes: Vec, page: u32) -> SourcePageResult { + let ctx = SnapshotContext::new( + snapshot_from_bytes(Path::new("h.pdf"), bytes, None).expect("fixture opens"), + ); + classify_source_page( + &ctx, + &build_page_model(&ctx, page, None).expect("model"), + None, + ) +} + +fn text_reasons(r: &SourcePageResult) -> Vec> { + r.occurrences + .iter() + .filter(|o| kind_token(o) == "text") + .map(|o| o.reason) + .collect() +} + +fn classify_bytes( + scratch: &Scratch, + name: &str, + bytes: &[u8], +) -> Result, AppError> { + classify_source_content(&write(scratch, name, bytes)) +} + +#[test] +fn cls_h01_hardening_file_read_once_and_capped() { + let scratch = Scratch::new("h01"); + let huge = scratch.file("huge.pdf"); + File::create(&huge) + .unwrap() + .set_len(FILE_CAP_BYTES + 1) + .unwrap(); + expect_err_code( + classify_source_content(&huge), + "FILE_TOO_LARGE", + "CLS-H01 sparse 400 MiB + 1", + ); + // The per-page API works on the bytes of the one read: the path is never opened again. + let bytes = fx::word(); + let ghost = scratch.file("never-written.pdf"); + let ctx = SnapshotContext::new(snapshot_from_bytes(&ghost, bytes.clone(), None).unwrap()); + let result = classify_source_page(&ctx, &build_page_model(&ctx, 0, None).unwrap(), None); + assert!( + !ghost.exists(), + "CLS-H01 nothing read from or written to the path" + ); + let on_disk = write(&scratch, "word.pdf", &bytes); + assert_eq!( + classify_source_content(&on_disk).unwrap(), + result.occurrences, + "CLS-H01 the same occurrences from the one read" + ); +} + +#[test] +fn cls_h02_hardening_objstm_bomb_file_too_complex() { + let s = Scratch::new("h02"); + expect_err_code( + classify_bytes(&s, "b.pdf", &fx::objstm_bomb()), + "FILE_TOO_COMPLEX", + "CLS-H02", + ); +} + +#[test] +fn cls_h03_hardening_xref_stream_bomb_file_too_complex() { + let s = Scratch::new("h03"); + expect_err_code( + classify_bytes(&s, "b.pdf", &fx::xref_bomb()), + "FILE_TOO_COMPLEX", + "CLS-H03", + ); +} + +#[test] +fn cls_h04_hardening_deep_nesting_file_too_complex() { + let s = Scratch::new("h04"); + expect_err_code( + classify_bytes(&s, "n.pdf", &fx::deep_nesting(101)), + "FILE_TOO_COMPLEX", + "CLS-H04", + ); + assert!( + classify_bytes(&s, "ok.pdf", &fx::deep_nesting(100)).is_ok(), + "CLS-H04 100 levels" + ); +} + +#[test] +fn cls_h05_hardening_flate_bomb_page_too_complex() { + let s = Scratch::new("h05"); + expect_err_code( + classify_bytes(&s, "b.pdf", &fx::flate_bomb()), + "PAGE_TOO_COMPLEX", + "CLS-H05", + ); + let r = page_result(fx::flate_bomb(), 0); + assert_eq!( + (r.page_reason, r.occurrences.len()), + (Some(TextReason::PageTooComplex), 0) + ); + let other = page_result(fx::flate_bomb(), 1); + assert_eq!( + other.page_reason, None, + "CLS-H05 the next page is unaffected" + ); +} + +#[test] +fn cls_h06_hardening_tounicode_bomb_refuses_font() { + let mut d = DocBuilder::new(); + let font = cid_font(&mut d.b, "ABCDEF+Bomb", "AB"); + let bomb = d.b.add_stream( + "/Filter /FlateDecode", + &crate::pdf_engine::text_edit::testkit::pdf::zlib_zero_bomb(8), + ); + let body = String::from_utf8(d.b.body(font).unwrap().to_vec()).unwrap(); + let patched = body + .replacen("/ToUnicode", "/Bomb", 1) + .trim_end_matches(">>") + .to_string() + + &format!(" /ToUnicode {bomb} 0 R >>"); + d.b.set(font, patched); + d.page(PageSpec::new( + format!("BT /F1 12 Tf 72 700 Td <{}> Tj ET", cid_hex("AB", "AB")).as_bytes(), + &format!("/Font << /F1 {font} 0 R >>"), + )); + let started = std::time::Instant::now(); + let r = page_result(d.build(), 0); + assert_eq!(r.page_reason, None, "CLS-H06 the page itself is fine"); + assert_eq!( + text_reasons(&r), + [Some(TextReason::AmbiguousUnicode)], + "CLS-H06 the font is refused" + ); + assert!( + started.elapsed().as_secs() < 10, + "CLS-H06 capped while inflating" + ); +} + +/// One Helvetica page whose content stream is `data` with stream dictionary `dict`. +fn filtered_page(dict: &str, data: &[u8]) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let id = d.b.add_stream(dict, data); + d.page_raw( + &format!("{id} 0 R"), + &PageSpec::new(b"", &format!("/Font << /F1 {f} 0 R >>")), + ); + d.build() +} + +#[test] +fn cls_h07_hardening_no_raw_fallback_on_filter_errors() { + let plain: &[u8] = b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET BT /F1 12 Tf 72 700 Td (Lo) Tj ET"; + let s = Scratch::new("h07"); + // (1) /FlateDecode over plain bytes. + let r = classify_bytes( + &s, + "plain.pdf", + &filtered_page("/Filter /FlateDecode", plain), + ); + expect_err_code(r, "MALFORMED_CONTENT", "CLS-H07 Flate over plain bytes"); + // (2) /ASCIIHexDecode is decoded, not passed through raw. + let hex: String = plain.iter().map(|b| format!("{b:02X}")).collect::() + ">"; + let hits = classify( + &write( + &s, + "ahx.pdf", + &filtered_page("/Filter /ASCIIHexDecode", hex.as_bytes()), + ), + "CLS-H07", + ); + assert_eq!(hits.len(), 2, "CLS-H07 ASCIIHex decoded: {hits:?}"); + assert!(hits.iter().all(|o| capability_token(o) == "supported")); + assert!((hits[0].rect.x - 72.0).abs() < 1e-6); + // (3) 55 % of a Flate stream: never Ok with fewer rows. + let mut enc = flate2::write::ZlibEncoder::new(Vec::new(), flate2::Compression::none()); + enc.write_all(plain).unwrap(); + let full = enc.finish().unwrap(); + let cut = full[..full.len() * 55 / 100].to_vec(); + let truncated = filtered_page("/Filter /FlateDecode", &cut); + expect_err_code( + classify_bytes(&s, "cut.pdf", &truncated), + "MALFORMED_CONTENT", + "CLS-H07 truncated", + ); + let r = page_result(truncated, 0); + assert_eq!(r.page_reason, Some(TextReason::MalformedContent)); + assert!( + r.occurrences.is_empty() && r.runs.is_empty(), + "CLS-H07 zero rows, not a partial list" + ); +} + +#[test] +fn cls_h08_hardening_six_pages_of_1000_tj_ok() { + let mut page = String::from("BT /F1 1 Tf 20 700 Td "); + for _ in 0..1000 { + page.push_str("(a) Tj "); + } + page.push_str("ET"); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + for _ in 0..6 { + d.page(PageSpec::new( + page.as_bytes(), + &format!("/Font << /F1 {f} 0 R >>"), + )); + } + let s = Scratch::new("h08"); + let hits = classify(&write(&s, "six.pdf", &d.build()), "CLS-H08"); + assert_eq!(hits.len(), 6_000, "CLS-H08 budgets are per page"); + assert!(hits.iter().all(|o| capability_token(o) == "supported")); +} + +#[test] +fn cls_h09_hardening_page_over_op_budget_isolated() { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + d.page(PageSpec::new(b"BT /F1 12 Tf 72 700 Td (Fine) Tj ET", &res)); + let heavy = "q Q ".repeat(125_001); + d.page(PageSpec::new(heavy.as_bytes(), &res)); + let bytes = d.build(); + assert_eq!( + page_result(bytes.clone(), 0).page_reason, + None, + "CLS-H09 page 1 fine" + ); + assert_eq!( + page_result(bytes.clone(), 1).page_reason, + Some(TextReason::PageTooComplex), + "CLS-H09 page 2 over PAGE_OPS_MAX" + ); + let s = Scratch::new("h09"); + expect_err_code( + classify_bytes(&s, "ops.pdf", &bytes), + "PAGE_TOO_COMPLEX", + "CLS-H09 document API", + ); +} + +#[test] +fn cls_h10_hardening_q_overflow_refused_never_a_wrong_x() { + let probe = |depth: usize| { + format!( + "{}1 0 0 1 100 0 cm q 1 0 0 1 0 0 cm Q BT /F1 12 Tf 72 720 Td (Hi) Tj ET", + "q ".repeat(depth) + ) + }; + let r = page_result(helvetica_page(probe(64).as_bytes()), 0); + assert_eq!( + r.page_reason, + Some(TextReason::MalformedContent), + "CLS-H10 65th q" + ); + assert!(r.occurrences.is_empty()); + let r = page_result(helvetica_page(probe(63).as_bytes()), 0); + assert_eq!(r.page_reason, None); + assert!( + (r.occurrences[0].rect.x - 172.0).abs() < 1e-6, + "CLS-H10 x = 172 within the limit" + ); +} + +#[test] +fn cls_h11_hardening_fingerprint_matches_parsed_bytes() { + let s = Scratch::new("h11"); + let bytes = fx::word(); + let p = write(&s, "w.pdf", &bytes); + let snap = read_snapshot(&p).unwrap(); + assert_eq!( + (snap.fingerprint.len, snap.fingerprint.fnv), + (bytes.len() as u64, fnv1a_u64(&bytes)) + ); + assert_eq!( + snap.bytes.as_slice(), + &bytes[..], + "CLS-H11 parsed from the hashed bytes" + ); + let hits = classify(&p, "CLS-H11"); + assert!(hits + .iter() + .all(|o| o.locator.starts_with(&format!("v2:{}:", snap.fingerprint)))); + assert_eq!( + resolve_source_locator(&p, &hits[0].locator).unwrap(), + hits[0] + ); +} + +/// The [probe] font: TrueType `ABCDEF+Calibri`, `/Differences [1 /H /i]`, `(\001\002) Tj`. +fn calibri(widths: bool, program: bool) -> Vec { + let mut f = SimpleFont::new("TrueType", "ABCDEF+Calibri"); + f.encoding = Some("<< /Type /Encoding /Differences [1 /H /i] >>".into()); + if widths { + f.first_char = 1; + f.widths = Some(vec![600.0, 250.0]); + } + if program { + f.flags = Some(32); + f.program = Program::TrueType(latin_truetype("H", "i").0); + } + let mut d = DocBuilder::new(); + let id = add_simple(&mut d.b, &f); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 720 Td (\\001\\002) Tj ET", + &format!("/Font << /F1 {id} 0 R >>"), + )); + d.build() +} + +#[test] +fn cls_h12_hardening_subset_and_custom_encoded_fonts() { + let no_widths = page_result(calibri(false, false), 0); + assert_eq!( + text_reasons(&no_widths), + [Some(TextReason::MissingWidths)], + "CLS-H12 no /Widths" + ); + for (program, alphabet) in [(false, vec!['H', 'i']), (true, vec!['H'])] { + let ctx = SnapshotContext::new( + snapshot_from_bytes(Path::new("c.pdf"), calibri(true, program), None).unwrap(), + ); + let model = build_page_model(&ctx, 0, None).unwrap(); + let run = model.runs.first().expect("CLS-H12 run"); + assert_eq!( + (run.text.as_str(), run.reason), + ("Hi", None), + "CLS-H12 program={program}" + ); + assert_eq!(run.substituted, !program, "CLS-H12 substituted"); + assert_eq!( + model.surface(run).alphabet(), + alphabet, + "CLS-H12 alphabet, program={program}" + ); + assert!( + model.surface(run).writer_for('i').is_none() == program, + "CLS-H12 typing i" + ); + } + let libre = page_result(fx::libre(), 0); + assert!( + text_reasons(&libre).iter().all(Option::is_none), + "CLS-H12 LibreOffice symbolic TrueType" + ); +} + +#[test] +fn cls_h12b_no_tounicode_on_an_embedded_type0_font() { + use crate::pdf_engine::text_edit::testkit::fonts::{add_type0, Type0Font}; + use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; + let mut t = TtfBuilder::new(); + t.glyph("A", true, 500); + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+NoMap"); + f.program = Program::TrueType(t.build()); + let mut d = DocBuilder::new(); + let id = add_type0(&mut d.b, &f); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td <0001> Tj ET", + &format!("/Font << /F1 {id} 0 R >>"), + )); + assert_eq!( + text_reasons(&page_result(d.build(), 0)), + [Some(TextReason::NoTounicode)] + ); +} + +#[test] +fn cls_h13_hardening_rise_and_render_mode() { + let reasons = |content: &[u8]| text_reasons(&page_result(helvetica_page(content), 0)); + assert_eq!( + reasons(b"BT /F1 12 Tf 3 Tr 72 720 Td (Hi) Tj ET"), + [Some(TextReason::InvisibleText)], + "CLS-H13 3 Tr" + ); + assert_eq!( + reasons(b"BT /F1 12 Tf 7 Tr 72 720 Td (Hi) Tj ET"), + [Some(TextReason::TextClipMode)], + "CLS-H13 7 Tr" + ); + let risen = page_result( + helvetica_page(b"BT /F1 12 Tf 30 Ts 72 720 Td (Hi) Tj ET"), + 0, + ); + assert!( + (risen.occurrences[0].rect.y - 750.0).abs() < 1e-6, + "CLS-H13 rect y includes the rise" + ); +} + +#[test] +fn cls_h14_hardening_spacing_in_bounds() { + let w = |content: &[u8]| { + page_result(helvetica_page(content), 0).occurrences[0] + .rect + .w + }; + let plain = w(b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET"); + let tracked = w(b"BT /F1 12 Tf 10 Tc 72 720 Td (Hi) Tj ET"); + assert!( + (tracked - plain - 20.0).abs() < 1e-6, + "CLS-H14 Tc per glyph: {plain} → {tracked}" + ); + let spaced = w(b"BT /F1 12 Tf 20 Tw 72 720 Td (a b) Tj ET"); + let unspaced = w(b"BT /F1 12 Tf 72 720 Td (a b) Tj ET"); + assert!( + (spaced - unspaced - 20.0).abs() < 1e-6, + "CLS-H14 Tw in the width" + ); + let mut d = DocBuilder::new(); + let f = cid_font(&mut d.b, "ABCDEF+Arimo", "AB"); + d.page(PageSpec::new( + format!( + "BT /F1 10 Tf 5 Tc 72 700 Td <{}> Tj ET", + cid_hex("AB", "AB") + ) + .as_bytes(), + &format!("/Font << /F1 {f} 0 R >>"), + )); + let cid = page_result(d.build(), 0).occurrences[0].rect.w; + assert!( + (cid - (10.0 + 10.0)).abs() < 1e-6, + "CLS-H14 CID Tc once per 2-byte code: {cid}" + ); +} diff --git a/src-tauri/src/pdf_engine/source_content_integ/hardening2.rs b/src-tauri/src/pdf_engine/source_content_integ/hardening2.rs new file mode 100644 index 0000000..31a52a9 --- /dev/null +++ b/src-tauri/src/pdf_engine/source_content_integ/hardening2.rs @@ -0,0 +1,402 @@ +//! CLS-H15…H27 (SPEC §C rows 11–19 and F1): transforms, clip extents, shared content, per-face +//! widths, CID widths, v2 locators, the JSON shape, the frozen vocabulary, the reasons #33 never +//! tested (VERTICAL, CLIPPED, occurrence GEOMETRY, ENCRYPTED) and tagged/bookmarked pages. + +use super::*; +use crate::pdf_engine::source_content::{classify_source_page, SourcePageResult}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::engines_or_skip; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, ClipKind, DocBuilder, PageSpec, HELVETICA, +}; + +fn page_result(bytes: Vec, page: u32) -> SourcePageResult { + let ctx = SnapshotContext::new( + snapshot_from_bytes(Path::new("h.pdf"), bytes, None).expect("fixture opens"), + ); + classify_source_page( + &ctx, + &build_page_model(&ctx, page, None).expect("model"), + None, + ) +} + +fn first_reason(bytes: Vec, kind: &str) -> Option { + page_result(bytes, 0) + .occurrences + .into_iter() + .find(|o| kind_token(o) == kind) + .unwrap_or_else(|| panic!("a {kind} occurrence")) + .reason +} + +fn image_page(cm: &str, extra_content: &str) -> Vec { + let mut d = DocBuilder::new(); + let img = d.b.add_stream( + "/Type /XObject /Subtype /Image /Width 1 /Height 1 /ColorSpace /DeviceGray /BitsPerComponent 8", + &[128], + ); + d.page(PageSpec::new( + format!("{extra_content} q {cm} cm /Im0 Do Q").as_bytes(), + &format!("/XObject << /Im0 {img} 0 R >>"), + )); + d.build() +} + +#[test] +fn cls_h15_hardening_unsupported_transforms() { + use TextReason as R; + assert_eq!( + first_reason(fx::mirrored(), "text"), + Some(R::MirroredText), + "CLS-H15 -1 0 0 1 Tm" + ); + assert_eq!( + first_reason(fx::negative_tz(), "text"), + Some(R::MirroredText), + "CLS-H15 -100 Tz" + ); + assert_eq!( + first_reason(fx::negative_tf(), "text"), + Some(R::RotatedText), + "CLS-H15 -12 Tf" + ); + assert_eq!( + first_reason(image_page("0 40 -40 0 200 400", ""), "image"), + Some(R::TransformedImage), + "CLS-H15 rotated image" + ); + assert_eq!( + first_reason(image_page("40 0 0 -40 72 400", ""), "image"), + Some(R::TransformedImage), + "CLS-H15 flipped image" + ); + assert_eq!( + first_reason(image_page("40 0 0 40 72 400", ""), "image"), + None, + "CLS-H15 upright image" + ); +} + +#[test] +fn cls_h16_hardening_clip_extent() { + assert_eq!( + first_reason(fx::clip(ClipKind::Page), "text"), + None, + "CLS-H16 page-sized re W n" + ); + assert_eq!( + first_reason(fx::clip(ClipKind::Small), "text"), + Some(TextReason::Clipped), + "CLS-H16 small" + ); + assert_eq!( + first_reason(fx::clip(ClipKind::Curve), "text"), + Some(TextReason::Clipped), + "CLS-H16 Bézier" + ); +} + +#[test] +fn cls_h17_hardening_shared_content_on_both_pages() { + let s = Scratch::new("h17"); + let p = s.file("shared.pdf"); + fs::write(&p, fx::shared()).unwrap(); + let hits = classify(&p, "CLS-H17"); + let shared_rows: Vec<&SourceOccurrence> = hits + .iter() + .filter(|o| o.text.as_deref() == Some("Shared body")) + .collect(); + assert_eq!(shared_rows.len(), 2, "CLS-H17 one row per page"); + for o in shared_rows { + assert_unsupported(o, "text", "SHARED_CONTENT", "CLS-H17"); + } +} + +#[test] +fn cls_h18_hardening_standard14_widths_per_face() { + let mut d = DocBuilder::new(); + let f = d + .add("<< /Type /Font /Subtype /Type1 /BaseFont /Times-Roman /Encoding /WinAnsiEncoding >>"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + let w = page_result(d.build(), 0).occurrences[0].rect.w; + // Times-Roman AFM: H = 722, i = 278 → 12.0 at 12 pt (Helvetica would give 11.328). + assert!((w - 12.0).abs() < 1e-6, "CLS-H18 Times widths, got {w}"); + let helv = page_result(helvetica_page(b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET"), 0).occurrences[0] + .rect + .w; + assert!((helv - 11.328).abs() < 1e-6, "CLS-H18 Helvetica {helv}"); +} + +#[test] +fn cls_h19_hardening_cid_w_and_tc_per_code() { + use crate::pdf_engine::text_edit::testkit::fonts::{ + add_type0, tounicode_bfchar, Program, Type0Font, + }; + use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; + let mut t = TtfBuilder::new(); + for ch in ['A', 'B', 'C'] { + t.unicode_glyph(ch, &ch.to_string(), true); + } + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Widths"); + f.program = Program::TrueType(t.build()); + f.w = Some("[1 [600] 2 3 800]".into()); + f.tounicode = Some(tounicode_bfchar(&[(1, 2, "A"), (2, 2, "B"), (3, 2, "C")])); + let mut d = DocBuilder::new(); + let id = add_type0(&mut d.b, &f); + d.page(PageSpec::new( + b"BT /F1 10 Tf 1 Tc 72 700 Td <000100020003> Tj ET", + &format!("/Font << /F1 {id} 0 R >>"), + )); + let w = page_result(d.build(), 0).occurrences[0].rect.w; + // /W both forms: 600 + 800 + 800 → 22 pt, + 3 × 1 Tc (per code, not per byte). + assert!((w - 25.0).abs() < 1e-6, "CLS-H19 got {w}"); +} + +#[test] +fn cls_h20_hardening_locator_second_contents_stream_points_into_part_1() { + let s = Scratch::new("h20"); + let p = s.file("parts.pdf"); + fs::write(&p, fx::two_parts_mid_bt()).unwrap(); + let hits = classify(&p, "CLS-H20"); + let lo = hits + .iter() + .find(|o| o.text.as_deref() == Some("Lo")) + .expect("Lo"); + let fields: Vec<&str> = lo.locator.split(':').collect(); + assert_eq!( + (fields[0], fields[2], fields[3]), + ("v2", "0", "t"), + "CLS-H20 {}", + lo.locator + ); + let start: usize = fields[4] + .trim_start_matches('p') + .split('-') + .next() + .unwrap() + .parse() + .unwrap(); + assert!( + start > b"BT /F1 12 Tf 72 720 Td (Hi) Tj".len(), + "CLS-H20 span in part 1: {}", + lo.locator + ); + assert_eq!( + resolve_source_locator(&p, &lo.locator).unwrap(), + *lo, + "CLS-H20 resolves" + ); + // Repeated Forms stay unique through the ordinal. + let form_twice = { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let form = d.b.add_stream( + &format!("/Type /XObject /Subtype /Form /BBox [0 0 100 100] /Resources << /Font << /F1 {f} 0 R >> >>"), + b"BT /F1 9 Tf 5 5 Td (x) Tj ET", + ); + d.page(PageSpec::new( + b"/Fm0 Do 1 0 0 1 100 0 cm /Fm0 Do", + &format!("/XObject << /Fm0 {form} 0 R >>"), + )); + d.build() + }; + let occ = page_result(form_twice, 0).occurrences; + assert_eq!(occ.len(), 2); + assert_ne!( + occ[0].locator, occ[1].locator, + "CLS-H20 ordinal keeps Form paints apart" + ); + assert!( + occ[0].locator.contains(":x"), + "CLS-H20 Form chain path: {}", + occ[0].locator + ); +} + +#[test] +fn cls_h21_hardening_json_shape() { + let r = page_result(fx::word(), 0); + let json = serde_json::to_value(&r).unwrap(); + for key in [ + "pageIndex", + "pageReason", + "runs", + "occurrences", + "occurrenceReason", + ] { + assert!(json.get(key).is_some(), "CLS-H21 {key}"); + } + let run = &json["runs"][0]; + assert_eq!(run["capability"], "supported"); + assert!(run["runId"].as_str().unwrap().starts_with("t1:")); + assert!(run["reason"].is_null()); + let occ = &json["occurrences"][0]; + for key in [ + "pageIndex", + "kind", + "rect", + "locator", + "capability", + "reason", + "text", + ] { + assert!(occ.get(key).is_some(), "CLS-H21 occurrence {key}"); + } + assert_eq!(occ["kind"], "text"); + assert!(occ["rect"].get("w").is_some()); + let refused = serde_json::to_value(page_result(fx::type3(), 0)).unwrap(); + assert_eq!( + refused["occurrences"][0]["reason"], "TYPE3", + "CLS-H21 SCREAMING_SNAKE reasons" + ); + assert_eq!(refused["occurrences"][0]["capability"], "unsupported"); +} + +#[test] +fn cls_h22_hardening_frozen_reasons_come_from_the_json() { + let frozen = frozen_reasons(); + for legacy in LEGACY_REASONS { + assert!( + frozen.iter().any(|f| f == legacy), + "CLS-H22 {legacy} survives" + ); + } + assert!( + !frozen.iter().any(|f| f == "MALFORMED"), + "CLS-H22 the old typo is gone" + ); + for code in [ + "MALFORMED_CONTENT", + "FILE_TOO_LARGE", + "INVALID_PDF", + "TRANSFORMED_IMAGE", + "STALE", + ] { + assert!(frozen.iter().any(|f| f == code), "CLS-H22 {code}"); + } + for r in TextReason::RUN_PRIORITY + .iter() + .chain(TextReason::PAGE) + .chain(TextReason::IMAGE_PRIORITY) + { + assert!( + frozen.iter().any(|f| f == r.as_str()), + "CLS-H22 {}", + r.as_str() + ); + } +} + +#[test] +fn cls_h23_hardening_vertical() { + assert_eq!( + first_reason(fx::identity_v(), "text"), + Some(TextReason::Vertical), + "CLS-H23" + ); +} + +#[test] +fn cls_h24_hardening_clipped() { + let past = helvetica_page(b"BT /F1 12 Tf 590 720 Td (Overflow) Tj ET"); + assert_eq!( + first_reason(past, "text"), + Some(TextReason::Clipped), + "CLS-H24 past the page edge" + ); + assert_eq!( + first_reason(image_page("40 0 0 40 590 400", ""), "image"), + Some(TextReason::Clipped), + "CLS-H24 image across the edge" + ); + assert_eq!( + first_reason(image_page("40 0 0 40 72 400", "0 0 50 50 re W n"), "image"), + Some(TextReason::Clipped), + "CLS-H24 image outside a small clip" + ); +} + +#[test] +fn cls_h25_hardening_occurrence_geometry() { + let s = Scratch::new("h25"); + let p = s.file("unit.pdf"); + fs::write(&p, fx::user_unit(2.0)).unwrap(); + let hits = classify(&p, "CLS-H25 a GEOMETRY page is not an error"); + let text = first_of_kind(&hits, "text", "CLS-H25"); + assert_unsupported(text, "text", "GEOMETRY", "CLS-H25"); + let r = page_result(fx::user_unit(2.0), 0); + assert_eq!( + (r.page_reason, r.runs.len()), + (Some(TextReason::Geometry), 0) + ); +} + +#[test] +fn cls_h26_hardening_encrypted() { + let Some(engines) = engines_or_skip("cls_h26_hardening_encrypted") else { + return; + }; + let s = Scratch::new("h26"); + let p = s.file("enc.pdf"); + fs::write(&p, fx::encrypted(&engines)).unwrap(); + expect_err_code( + classify_source_content(&p), + "ENCRYPTED", + "CLS-H26 qpdf --encrypt", + ); +} + +#[test] +fn cls_h27_hardening_tagged_bookmarked_page_is_supported() { + let s = Scratch::new("h27"); + let p = s.file("word.pdf"); + fs::write(&p, fx::word()).unwrap(); + let hits = classify(&p, "CLS-H27"); + let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); + assert_eq!(texts.len(), 3, "CLS-H27 three show ops"); + for o in texts { + assert_supported_text_or_image(o, "text", "CLS-H27 tagged, bookmarked Word page (F1)"); + } +} + +#[test] +fn classify_walk_refusal_lists_no_occurrences() { + // A Form that paints itself: Edit text's depth-0 walk never enters it, the Classify walk + // does and refuses the page — the runs stay listed, the occurrences are not (never partial). + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let form = d.reserve(); + d.b.set_stream( + form, + &format!( + "/Type /XObject /Subtype /Form /BBox [0 0 300 100] \ + /Resources << /XObject << /Fm0 {form} 0 R >> >>" + ), + b"/Fm0 Do", + ); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Body) Tj ET q 1 0 0 1 72 600 cm /Fm0 Do Q", + &format!("/Font << /F1 {f} 0 R >> /XObject << /Fm0 {form} 0 R >>"), + )); + let bytes = d.build(); + let r = page_result(bytes.clone(), 0); + assert_eq!(r.page_reason, None, "Edit text's page is fine"); + assert_eq!(r.runs.len(), 1, "its run is listed: {:?}", r.runs); + assert_eq!(r.occurrence_reason, Some(TextReason::MalformedContent)); + assert!(r.occurrences.is_empty(), "{:?}", r.occurrences); + let s = Scratch::new("occ-reason"); + let path = s.file("self-form.pdf"); + fs::write(&path, &bytes).unwrap(); + expect_err_code( + classify_source_content(&path), + "MALFORMED_CONTENT", + "occurrence_reason", + ); +} diff --git a/src-tauri/src/pdf_engine/source_content_integ/review.rs b/src-tauri/src/pdf_engine/source_content_integ/review.rs new file mode 100644 index 0000000..9a71968 --- /dev/null +++ b/src-tauri/src/pdf_engine/source_content_integ/review.rs @@ -0,0 +1,516 @@ +//! PR #97 review regressions R1–R9 (generated in temp; `fixtures/source-edit/` does not grow). + +use super::*; + +// --- PR 97 review fold (R1–R5) --------------------------------------------- +// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. + +pub(super) fn box_obj(b: [i64; 4]) -> Object { + Object::Array(b.into_iter().map(Object::Integer).collect()) +} + +#[test] +fn classify_malformed_inline_tail_returns_error_without_panicking() { + let scratch = Scratch::new("review-malformed-inline"); + let path = scratch.file("tail.pdf"); + let mut content = b"BI /W 1 /H 1 /BPC 8 /CS /G ID x EI\n(".to_vec(); + content.push(b'\\'); + write_helvetica_page(&path, &content); + let result = std::panic::catch_unwind(|| classify_source_content(&path)); + assert!( + result.is_ok(), + "Malformed content must return an AppError, not panic" + ); + assert!(result.unwrap().is_err()); +} + +#[test] +fn classify_missing_font_is_not_supported() { + let scratch = Scratch::new("review-missing-font"); + let path = scratch.file("missing-font.pdf"); + write_helvetica_page(&path, b"BT /Missing 12 Tf 72 400 Td (Hi) Tj ET"); + let hits = classify(&path, "REVIEW-MISSING-FONT"); + assert_eq!(capability_token(&hits[0]), "unsupported"); + assert_eq!(reason_code(&hits[0]).as_deref(), Some("MISSING_FONT")); +} + +#[test] +fn classify_repeated_form_occurrences_have_unique_locators() { + let scratch = Scratch::new("review-repeated-form"); + let path = scratch.file("repeated-form.pdf"); + let mut doc = Document::load(fixture("image-in-form.pdf")).unwrap(); + let page = *doc.get_pages().values().next().unwrap(); + let ids = doc.get_page_contents(page); + let bytes = doc.get_page_content(page).unwrap(); + let mut repeated = bytes.clone(); + repeated.extend_from_slice(b"\n1 0 0 1 100 0 cm\n"); + repeated.extend_from_slice(&bytes); + doc.objects.insert( + ids[0], + Object::Stream(Stream::new(Dictionary::new(), repeated)), + ); + doc.save(&path).unwrap(); + let hits = classify(&path, "REVIEW-REPEATED-FORM"); + assert!(hits.len() >= 2); + let locators: std::collections::HashSet<_> = hits.iter().map(|h| &h.locator).collect(); + assert_eq!( + locators.len(), + hits.len(), + "Every occurrence needs its own locator" + ); + assert_ne!(hits[0].rect, hits[1].rect); + assert_eq!( + resolve_source_locator(&path, &hits[1].locator).unwrap(), + hits[1] + ); + assert_eq!(classify(&path, "REVIEW-REPEATED-FORM"), hits); +} + +pub(super) fn helvetica_resources() -> Dictionary { + let mut font = Dictionary::new(); + font.set("Type", "Font"); + font.set("Subtype", "Type1"); + font.set("BaseFont", "Helvetica"); + let mut fonts = Dictionary::new(); + fonts.set("F1", Object::Dictionary(font)); + let mut res = Dictionary::new(); + res.set("Font", Object::Dictionary(fonts)); + res +} + +pub(super) fn write_helvetica_page(path: &Path, content: &[u8]) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + content.to_vec(), + ))); + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", content_id); + page.set("Resources", Object::Dictionary(helvetica_resources())); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path).expect("write generated classifier fixture"); +} + +fn write_text_with_empty_sig_widget(path: &Path) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET\n".to_vec(), + ))); + let widget_id = doc.new_object_id(); + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", content_id); + page.set("Resources", Object::Dictionary(helvetica_resources())); + page.set("Annots", vec![Object::Reference(widget_id)]); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut widget = Dictionary::new(); + widget.set("Type", "Annot"); + widget.set("Subtype", "Widget"); + widget.set("FT", "Sig"); + widget.set("T", Object::string_literal("Sig1")); + widget.set("Rect", box_obj([72, 72, 172, 92])); + widget.set("P", page_id); + doc.objects.insert(widget_id, Object::Dictionary(widget)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut acro = Dictionary::new(); + acro.set("Fields", vec![Object::Reference(widget_id)]); + let acro_id = doc.add_object(Object::Dictionary(acro)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + catalog.set("AcroForm", acro_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path) + .expect("write empty-sig-widget classifier fixture"); +} + +fn write_applied_signature(path: &Path) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET\n".to_vec(), + ))); + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", content_id); + page.set("Resources", Object::Dictionary(helvetica_resources())); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut sig = Dictionary::new(); + sig.set("Type", "Sig"); + sig.set( + "ByteRange", + vec![ + Object::Integer(0), + Object::Integer(10), + Object::Integer(20), + Object::Integer(30), + ], + ); + let _sig_id = doc.add_object(Object::Dictionary(sig)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path) + .expect("write applied-signature classifier fixture"); +} + +// --- R1 -------------------------------------------------------------------- + +#[test] +fn classify_180_degree_text_is_rotated() { + let scratch = Scratch::new("r1-180"); + let path = scratch.file("text-180.pdf"); + write_helvetica_page(&path, b"BT /F1 12 Tf -1 0 0 -1 200 400 Tm (Hi) Tj ET\n"); + let hits = classify(&path, "R1"); + let occ = first_of_kind(&hits, "text", "R1"); + assert_ne!( + capability_token(occ), + "supported", + "R1: 180° Tm [-1 0 0 -1 200 400] must not be supported" + ); + assert_unsupported(occ, "text", "ROTATED_TEXT", "R1"); +} + +// --- R2 -------------------------------------------------------------------- + +#[test] +fn classify_second_tj_advances_tm() { + let scratch = Scratch::new("r2-advance"); + let path = scratch.file("two-tj.pdf"); + write_helvetica_page(&path, b"BT /F1 12 Tf 72 720 Td (Hel) Tj (lo) Tj ET\n"); + let hits = classify(&path, "R2"); + let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); + assert_eq!( + texts.len(), + 2, + "R2: (Hel) Tj (lo) Tj must emit two text occurrences; got {:?}", + hits.iter() + .map(|o| (kind_token(o), o.rect.x, o.rect.y)) + .collect::>() + ); + assert!( + (texts[0].rect.x - 72.0).abs() <= 1.0, + "R2: first origin x must be ~72; got {}", + texts[0].rect.x + ); + assert!( + texts[1].rect.x > texts[0].rect.x, + "R2: second rect.x must be > first (must not share origin 72); first x={} second x={}", + texts[0].rect.x, + texts[1].rect.x + ); + assert!( + (texts[1].rect.x - 72.0).abs() > 1.0, + "R2: second show must not reuse origin 72; first x={} second x={}", + texts[0].rect.x, + texts[1].rect.x + ); +} + +// --- R3 -------------------------------------------------------------------- + +#[test] +fn classify_empty_sig_widget_does_not_refuse_file() { + let scratch = Scratch::new("r3-empty-sig"); + let path = scratch.file("empty-sig.pdf"); + write_text_with_empty_sig_widget(&path); + let hits = match classify_source_content(&path) { + Ok(hits) => hits, + Err(err) => panic!( + "R3: empty /FT /Sig widget (Type Annot, no ByteRange) + Helvetica (Hi) Tj must be Ok with a text occurrence, not AppError SIGNED; got {} ({})", + err.code, err.message + ), + }; + let occ = first_of_kind(&hits, "text", "R3"); + assert_ne!( + reason_code(occ).as_deref(), + Some("SIGNED"), + "R3: Helvetica text on a file with an empty Sig widget must not be SIGNED" + ); +} + +#[test] +fn classify_applied_signature_is_signed() { + let scratch = Scratch::new("r3-applied-sig"); + let path = scratch.file("applied-sig.pdf"); + write_applied_signature(&path); + expect_err_code( + classify_source_content(&path), + "SIGNED", + "R3: /Type /Sig + ByteRange still refuses", + ); +} + +// --- R4 -------------------------------------------------------------------- + +#[test] +fn classify_text_after_inline_image_is_kept() { + let scratch = Scratch::new("r4-inline-rest"); + let path = scratch.file("inline-then-text.pdf"); + write_helvetica_page( + &path, + b"q 24 0 0 12 72 400 cm\n\ +BI\n\ +/W 2 /H 1 /CS /DeviceRGB /BPC 8 /F /AHx\n\ +ID\n\ +C8101010C810>\n\ +EI\n\ +Q\n\ +BT /F1 12 Tf 72 720 Td (Hi) Tj ET\n", + ); + let hits = classify(&path, "R4"); + let _text = first_of_kind(&hits, "text", "R4"); +} + +// --- R5 -------------------------------------------------------------------- + +#[test] +fn classify_text_bounds_use_tm_scale() { + let scratch = Scratch::new("r5-tm-scale"); + let path = scratch.file("tf1-tm12.pdf"); + write_helvetica_page(&path, b"BT /F1 1 Tf 12 0 0 12 72 720 Tm (Hi) Tj ET\n"); + let hits = classify(&path, "R5"); + let occ = first_of_kind(&hits, "text", "R5"); + assert!( + (occ.rect.h - 12.0).abs() <= 1.0, + "R5: /F1 1 Tf + 12 0 0 12 Tm must report height ~12, not ~1; got h={}", + occ.rect.h + ); + assert!( + (occ.rect.h - 1.0).abs() > 1.0, + "R5: rect.h must not stay at Tf size ~1; got h={}", + occ.rect.h + ); +} + +// --- PR 97 review fold r2 (R6–R9) ------------------------------------------ +// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. + +fn write_type3_and_helvetica_page(path: &Path, content: &[u8]) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + content.to_vec(), + ))); + + let proc_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + b"10 0 0 0 10 10 d1\n0 0 10 10 re f\n".to_vec(), + ))); + let mut char_procs = Dictionary::new(); + char_procs.set("x", proc_id); + + let mut enc = Dictionary::new(); + enc.set("Type", "Encoding"); + enc.set( + "Differences", + vec![Object::Integer(120), Object::Name(b"x".to_vec())], + ); + + let mut t3 = Dictionary::new(); + t3.set("Type", "Font"); + t3.set("Subtype", "Type3"); + t3.set("FontBBox", box_obj([0, 0, 10, 10])); + t3.set( + "FontMatrix", + Object::Array(vec![ + Object::Real(1.0), + Object::Real(0.0), + Object::Real(0.0), + Object::Real(1.0), + Object::Real(0.0), + Object::Real(0.0), + ]), + ); + t3.set("CharProcs", Object::Dictionary(char_procs)); + t3.set("Encoding", Object::Dictionary(enc)); + t3.set("FirstChar", 120); + t3.set("LastChar", 120); + t3.set("Widths", vec![Object::Integer(10)]); + + let mut f1 = Dictionary::new(); + f1.set("Type", "Font"); + f1.set("Subtype", "Type1"); + f1.set("BaseFont", "Helvetica"); + + let mut fonts = Dictionary::new(); + fonts.set("T3", Object::Dictionary(t3)); + fonts.set("F1", Object::Dictionary(f1)); + let mut res = Dictionary::new(); + res.set("Font", Object::Dictionary(fonts)); + + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", content_id); + page.set("Resources", Object::Dictionary(res)); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path) + .expect("write Type3+Helvetica classifier fixture"); +} + +// --- R6 -------------------------------------------------------------------- + +#[test] +fn classify_inline_image_uses_ctm_at_bi() { + let path = fixture("image-inline.pdf"); + let hits = classify(&path, "R6"); + let occ = first_of_kind(&hits, "image", "R6"); + assert!( + (occ.rect.x - 72.0).abs() <= 1.0, + "R6: image-inline.pdf image rect.x must be ~72 (CTM at BI), not the unit square at origin; got x={}", + occ.rect.x + ); + assert!( + (occ.rect.y - 400.0).abs() <= 1.0, + "R6: image-inline.pdf image rect.y must be ~400 (CTM at BI), not the unit square at origin; got y={}", + occ.rect.y + ); + assert!( + (occ.rect.w - 24.0).abs() <= 1.0, + "R6: image-inline.pdf image rect.w must be ~24 (CTM at BI), not the unit square; got w={}", + occ.rect.w + ); + assert!( + (occ.rect.h - 12.0).abs() <= 1.0, + "R6: image-inline.pdf image rect.h must be ~12 (CTM at BI), not the unit square; got h={}", + occ.rect.h + ); +} + +// --- R7 -------------------------------------------------------------------- + +#[test] +fn classify_q_restores_type3_after_helvetica() { + let scratch = Scratch::new("r7-q-type3"); + let path = scratch.file("q-type3.pdf"); + write_type3_and_helvetica_page( + &path, + // Deliberate change (T3): the text is moved into the page (`72 720 Td`). At the page + // origin its descenders fall below the MediaBox, and CLIPPED (§A.10 #12) now outranks + // TYPE3 (#18); this test is about q/Q restoring the font. + b"BT 72 720 Td /T3 12 Tf (x) Tj q /F1 12 Tf (y) Tj Q (z) Tj ET\n", + ); + let hits = classify(&path, "R7"); + let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); + assert_eq!( + texts.len(), + 3, + "R7: (x) Tj q /F1 (y) Tj Q (z) Tj must emit three text occurrences; got {:?}", + hits.iter() + .map(|o| (kind_token(o), capability_token(o), reason_code(o))) + .collect::>() + ); + let last = texts[2]; + assert_ne!( + capability_token(last), + "supported", + "R7: last show (z) after Q must not be Helvetica supported; got {} reason={:?}", + capability_token(last), + reason_code(last) + ); + assert_unsupported(last, "text", "TYPE3", "R7"); +} + +// --- R8 -------------------------------------------------------------------- + +#[test] +fn classify_tc_advances_second_tj() { + let scratch = Scratch::new("r8-tc"); + let path = scratch.file("tc-two-tj.pdf"); + write_helvetica_page( + &path, + b"BT /F1 12 Tf 2 Tc 72 720 Td (Hi) Tj (there) Tj ET\n", + ); + let hits = classify(&path, "R8"); + let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); + assert_eq!( + texts.len(), + 2, + "R8: (Hi) Tj (there) Tj must emit two text occurrences; got {:?}", + hits.iter() + .map(|o| (kind_token(o), o.rect.x, o.rect.y)) + .collect::>() + ); + // Helvetica H=667 i=278 → 11.34 at Tf=12. 2 Tc on two glyphs adds 4 + // user units, so second.x ≈ first.x + 15.34, not first.x + 11.34. + assert!( + texts[1].rect.x > texts[0].rect.x + 13.0, + "R8: 2 Tc must push second.x past first.x + no-Tc Hi width 11.34; first.x={} second.x={} (need second.x > first.x + 13)", + texts[0].rect.x, + texts[1].rect.x + ); +} + +// --- R9 -------------------------------------------------------------------- + +#[test] +fn classify_source_drops_rotated_fixture_parenthetical() { + let path = Path::new(env!("CARGO_MANIFEST_DIR")).join("src/pdf_engine/source_content.rs"); + let src = fs::read_to_string(&path).unwrap_or_else(|e| { + panic!( + "R9: must read source_content.rs via CARGO_MANIFEST_DIR ({}): {e}", + path.display() + ) + }); + assert!( + !src.contains("keeps text-rotated.pdf green"), + "R9: source_content.rs must not contain the exact substring `keeps text-rotated.pdf green`" + ); +} diff --git a/src-tauri/src/pdf_engine/source_content_integ/review2.rs b/src-tauri/src/pdf_engine/source_content_integ/review2.rs new file mode 100644 index 0000000..fe99df8 --- /dev/null +++ b/src-tauri/src/pdf_engine/source_content_integ/review2.rs @@ -0,0 +1,514 @@ +//! PR #97 review regressions R10–R14 (generated in temp; `fixtures/source-edit/` does not grow). + +use super::review::{box_obj, helvetica_resources, write_helvetica_page}; +use super::*; + +// --- PR 97 review fold r3 (R10–R12) ---------------------------------------- +// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. + +/// 2×2 DeviceRGB, same bytes as the #32 unique/mask fixtures. +const R3_TINY_RGB: &[u8] = &[200, 16, 16, 16, 200, 16, 16, 16, 200, 200, 200, 16]; + +fn write_type1_indirect_widths(path: &Path, content: &[u8]) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + content.to_vec(), + ))); + + // FirstChar 'H' (72) … LastChar 'i' (105): 34 glyph slots, all 1000. + const FIRST_CHAR: i64 = 72; + const LAST_CHAR: i64 = 105; + let widths: Vec = (FIRST_CHAR..=LAST_CHAR) + .map(|_| Object::Integer(1000)) + .collect(); + let widths_id = doc.add_object(Object::Array(widths)); + + let mut font = Dictionary::new(); + font.set("Type", "Font"); + font.set("Subtype", "Type1"); + font.set("BaseFont", "Helvetica"); + font.set("FirstChar", FIRST_CHAR); + font.set("LastChar", LAST_CHAR); + font.set("Widths", Object::Reference(widths_id)); + + let mut fonts = Dictionary::new(); + fonts.set("F1", Object::Dictionary(font)); + let mut res = Dictionary::new(); + res.set("Font", Object::Dictionary(fonts)); + + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", content_id); + page.set("Resources", Object::Dictionary(res)); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path) + .expect("write Type1 indirect-Widths classifier fixture"); +} + +fn write_pattern_cs_page(path: &Path, content: &[u8]) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + content.to_vec(), + ))); + + let mut font = Dictionary::new(); + font.set("Type", "Font"); + font.set("Subtype", "Type1"); + font.set("BaseFont", "Helvetica"); + let mut fonts = Dictionary::new(); + fonts.set("F1", Object::Dictionary(font)); + + let mut cs = Dictionary::new(); + cs.set("Cs1", Object::Name(b"Pattern".to_vec())); + + let mut pat = Dictionary::new(); + pat.set("Type", "Pattern"); + pat.set("PatternType", 1); + pat.set("PaintType", 1); + pat.set("TilingType", 1); + pat.set("BBox", box_obj([0, 0, 10, 10])); + pat.set("XStep", 10); + pat.set("YStep", 10); + pat.set("Resources", Object::Dictionary(Dictionary::new())); + let pat_id = doc.add_object(Object::Stream(Stream::new( + pat, + b"0 0 10 10 re f\n".to_vec(), + ))); + let mut patterns = Dictionary::new(); + patterns.set("P1", Object::Reference(pat_id)); + + let mut res = Dictionary::new(); + res.set("Font", Object::Dictionary(fonts)); + res.set("ColorSpace", Object::Dictionary(cs)); + res.set("Pattern", Object::Dictionary(patterns)); + + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", content_id); + page.set("Resources", Object::Dictionary(res)); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path) + .expect("write Pattern ColorSpace classifier fixture"); +} + +fn write_extgstate_smask_image(path: &Path, content: &[u8]) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + content.to_vec(), + ))); + + let mut img = Dictionary::new(); + img.set("Type", "XObject"); + img.set("Subtype", "Image"); + img.set("Width", 2); + img.set("Height", 2); + img.set("ColorSpace", "DeviceRGB"); + img.set("BitsPerComponent", 8); + let img_id = doc.add_object(Object::Stream(Stream::new(img, R3_TINY_RGB.to_vec()))); + + let mut sm = Dictionary::new(); + sm.set("Type", "XObject"); + sm.set("Subtype", "Image"); + sm.set("Width", 2); + sm.set("Height", 2); + sm.set("ColorSpace", "DeviceGray"); + sm.set("BitsPerComponent", 8); + let smask_id = doc.add_object(Object::Stream(Stream::new(sm, vec![255, 200, 180, 255]))); + + let mut gs = Dictionary::new(); + gs.set("Type", "ExtGState"); + gs.set("SMask", Object::Reference(smask_id)); + let gs_id = doc.add_object(Object::Dictionary(gs)); + + let mut xobjects = Dictionary::new(); + xobjects.set("Im0", Object::Reference(img_id)); + let mut extg = Dictionary::new(); + extg.set("Gs1", Object::Reference(gs_id)); + let mut res = Dictionary::new(); + res.set("XObject", Object::Dictionary(xobjects)); + res.set("ExtGState", Object::Dictionary(extg)); + + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", content_id); + page.set("Resources", Object::Dictionary(res)); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path) + .expect("write ExtGState SMask + unique Image classifier fixture"); +} + +// --- R10 ------------------------------------------------------------------- + +#[test] +fn classify_indirect_widths_not_helvetica_fallback() { + let scratch = Scratch::new("r10-widths"); + let path = scratch.file("indirect-widths.pdf"); + write_type1_indirect_widths(&path, b"BT /F1 12 Tf 72 720 Td (Hi) Tj ET\n"); + let hits = classify(&path, "R10"); + let occ = first_of_kind(&hits, "text", "R10"); + assert!( + (occ.rect.w - 24.0).abs() <= 1.0, + "R10: Type1 indirect /Widths 1000,1000 at Tf=12 must report w≈24, not Helvetica fallback ≈11.34; got w={}", + occ.rect.w + ); + assert!( + (occ.rect.w - 11.34).abs() > 1.0, + "R10: rect.w must not stay on the Helvetica table ≈11.34; got w={}", + occ.rect.w + ); +} + +// --- R11 ------------------------------------------------------------------- + +#[test] +fn classify_named_pattern_cs_is_unsupported() { + let scratch = Scratch::new("r11-pattern"); + let path = scratch.file("pattern-cs.pdf"); + write_pattern_cs_page( + &path, + b"BT /F1 12 Tf /Cs1 cs /P1 scn 72 720 Td (Hi) Tj ET\n", + ); + let hits = classify(&path, "R11"); + let occ = first_of_kind(&hits, "text", "R11"); + assert_ne!( + capability_token(occ), + "supported", + "R11: /Cs1 cs Pattern resource + (Hi) Tj must not be supported; got {} reason={:?}", + capability_token(occ), + reason_code(occ) + ); + assert_unsupported(occ, "text", "PATTERN", "R11"); +} + +// --- R12 ------------------------------------------------------------------- + +#[test] +fn classify_extgstate_smask_image_is_masked() { + let scratch = Scratch::new("r12-gs-smask"); + let path = scratch.file("gs-smask.pdf"); + write_extgstate_smask_image(&path, b"q 40 0 0 40 72 400 cm /Gs1 gs /Im0 Do Q\n"); + let hits = classify(&path, "R12"); + let occ = first_of_kind(&hits, "image", "R12"); + assert_ne!( + capability_token(occ), + "supported", + "R12: unique Image after ExtGState /Gs1 /SMask must not be supported; got {} reason={:?}", + capability_token(occ), + reason_code(occ) + ); + assert_unsupported(occ, "image", "MASKED_IMAGE", "R12"); +} + +// --- PR 97 review fold r4 (R13) -------------------------------------------- +// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. + +fn write_unique_rgb_image(path: &Path, content: &[u8]) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content_id = doc.add_object(Object::Stream(Stream::new( + Dictionary::new(), + content.to_vec(), + ))); + + let mut img = Dictionary::new(); + img.set("Type", "XObject"); + img.set("Subtype", "Image"); + img.set("Width", 2); + img.set("Height", 2); + img.set("ColorSpace", "DeviceRGB"); + img.set("BitsPerComponent", 8); + let img_id = doc.add_object(Object::Stream(Stream::new(img, R3_TINY_RGB.to_vec()))); + + let mut xobjects = Dictionary::new(); + xobjects.set("Im0", Object::Reference(img_id)); + let mut res = Dictionary::new(); + res.set("XObject", Object::Dictionary(xobjects)); + + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", content_id); + page.set("Resources", Object::Dictionary(res)); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path) + .expect("write unique 2×2 DeviceRGB Image classifier fixture"); +} + +// --- R13a ------------------------------------------------------------------ + +#[test] +fn classify_stacked_cm_image_origin() { + let scratch = Scratch::new("r13a-stacked-cm"); + let path = scratch.file("stacked-cm.pdf"); + write_unique_rgb_image(&path, b"q 2 0 0 2 0 0 cm 20 0 0 20 36 200 cm /Im0 Do Q\n"); + let hits = classify(&path, "R13a"); + let occ = first_of_kind(&hits, "image", "R13a"); + assert_supported_text_or_image(occ, "image", "R13a"); + assert!( + (occ.rect.x - 72.0).abs() <= 1.0 + && (occ.rect.y - 400.0).abs() <= 1.0 + && (occ.rect.w - 40.0).abs() <= 1.0 + && (occ.rect.h - 40.0).abs() <= 1.0, + "R13a: stacked cm image rect must be ~{{x:72, y:400, w:40, h:40}}, not origin ~(36, 200); got {{x:{}, y:{}, w:{}, h:{}}}", + occ.rect.x, + occ.rect.y, + occ.rect.w, + occ.rect.h + ); + assert!( + (occ.rect.x - 36.0).abs() > 1.0 || (occ.rect.y - 200.0).abs() > 1.0, + "R13a: stacked cm must not leave the image at the second-cm translation (36, 200); got {{x:{}, y:{}, w:{}, h:{}}}", + occ.rect.x, + occ.rect.y, + occ.rect.w, + occ.rect.h + ); +} + +// --- R13b ------------------------------------------------------------------ + +#[test] +fn classify_scaled_tm_second_show_x() { + let scratch = Scratch::new("r13b-scaled-tm"); + let path = scratch.file("scaled-tm-two-tj.pdf"); + write_helvetica_page( + &path, + b"BT /F1 1 Tf 12 0 0 12 72 720 Tm (Hel) Tj (lo) Tj ET\n", + ); + let hits = classify(&path, "R13b"); + let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); + assert_eq!( + texts.len(), + 2, + "R13b: (Hel) Tj (lo) Tj must emit two text occurrences; got {:?}", + hits.iter() + .map(|o| (kind_token(o), o.rect.x, o.rect.y)) + .collect::>() + ); + let first = texts[0]; + let second = texts[1]; + // Helvetica H=667 e=556 l=278 → 1.501 at Tf=1. Scaled Tm 12× must + // advance ~18.012 user units → second.x ≈ 90, not text-space 1.501 + // added in user space (≈73.5). + assert!( + second.rect.x > first.rect.x + 15.0 || (second.rect.x - 90.0).abs() <= 2.0, + "R13b: second rect.x after 12 0 0 12 72 720 Tm (Hel) Tj must be ≈90 (±2), not ≈73.5; first.x={} second.x={}", + first.rect.x, + second.rect.x + ); + assert!( + (second.rect.x - 73.5).abs() > 1.0, + "R13b: second rect.x must not stay at origin+text-space width ≈73.5; first.x={} second.x={}", + first.rect.x, + second.rect.x + ); +} + +// --- PR 97 review fold r5 (R14) -------------------------------------------- +// Extra PDFs are generated in temp with lopdf. Do not grow fixtures/source-edit/. +// Page /Contents is an array of two streams; stream 1 has no trailing whitespace +// so a join without a separator fuses `Tj`+`ET` into `TjET`. + +const R14_STREAM_1: &[u8] = b"BT /F1 12 Tf 72 720 Td (Hi) Tj"; +const R14_STREAM_2: &[u8] = b"ET\nBT /F1 12 Tf 72 680 Td (Lo) Tj ET"; + +fn stream_content_bytes(doc: &Document, obj: &Object) -> Vec { + let id = obj + .as_reference() + .expect("Contents array entry must be a stream ref"); + doc.get_object(id) + .expect("content stream object") + .as_stream() + .expect("content must be a stream") + .content + .clone() +} + +fn write_helvetica_two_content_streams(path: &Path, stream1: &[u8], stream2: &[u8]) { + let mut doc = Document::with_version("1.7"); + let pages_id = doc.new_object_id(); + let content1_id = doc.add_object(Object::Stream( + Stream::new(Dictionary::new(), stream1.to_vec()).with_compression(false), + )); + let content2_id = doc.add_object(Object::Stream( + Stream::new(Dictionary::new(), stream2.to_vec()).with_compression(false), + )); + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set( + "Contents", + vec![ + Object::Reference(content1_id), + Object::Reference(content2_id), + ], + ); + page.set("Resources", Object::Dictionary(helvetica_resources())); + let page_id = doc.add_object(Object::Dictionary(page)); + + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path) + .expect("write two-stream Contents classifier fixture"); + + // Lock the on-disk page /Contents shape: array of two streams, exact bytes. + let reloaded = + Document::load(path).unwrap_or_else(|e| panic!("reload two-stream Contents fixture: {e}")); + let page_id = *reloaded + .get_pages() + .get(&1) + .expect("two-stream fixture must have page 1"); + let page = reloaded + .get_object(page_id) + .expect("page 1 object") + .as_dict() + .expect("page 1 dict"); + let contents = page.get(b"Contents").expect("page /Contents"); + let refs = match contents { + Object::Array(arr) => arr, + other => panic!( + "two-stream fixture page /Contents must be an array of two stream refs, got {other:?}" + ), + }; + assert_eq!( + refs.len(), + 2, + "two-stream fixture page /Contents must have two stream refs; got {}", + refs.len() + ); + assert_eq!( + stream_content_bytes(&reloaded, &refs[0]).as_slice(), + stream1, + "two-stream fixture stream 1 bytes must be exact (no trailing newline)" + ); + assert_eq!( + stream_content_bytes(&reloaded, &refs[1]).as_slice(), + stream2, + "two-stream fixture stream 2 bytes must match" + ); +} + +// --- R14 ------------------------------------------------------------------- + +#[test] +fn classify_contents_array_two_streams_do_not_fuse() { + let scratch = Scratch::new("r14-two-streams"); + let path = scratch.file("two-contents-streams.pdf"); + write_helvetica_two_content_streams(&path, R14_STREAM_1, R14_STREAM_2); + let hits = classify(&path, "R14"); + let texts: Vec<&SourceOccurrence> = hits.iter().filter(|o| kind_token(o) == "text").collect(); + assert_eq!( + texts.len(), + 2, + "R14: two Contents streams (Hi@720 then Lo@680) must emit two text occurrences; got {:?}", + hits.iter() + .map(|o| (kind_token(o), o.rect.x, o.rect.y)) + .collect::>() + ); + assert!( + texts.iter().any(|o| (o.rect.y - 720.0).abs() <= 1.0), + "R14: expected a text occurrence at y≈720 (Hi); got {:?}", + texts + .iter() + .map(|o| (o.rect.x, o.rect.y)) + .collect::>() + ); + assert!( + texts.iter().any(|o| (o.rect.y - 680.0).abs() <= 1.0), + "R14: expected a text occurrence at y≈680 (Lo); got {:?}", + texts + .iter() + .map(|o| (o.rect.x, o.rect.y)) + .collect::>() + ); + // Deliberate change (§C row 16): the v2 locator of "Lo" names its span in the joined content, + // which lies in part 1 (after part 0 and its "\n" separator), not in the first stream. + let lo = texts + .iter() + .find(|o| (o.rect.y - 680.0).abs() <= 1.0) + .expect("R14: Lo"); + let path = lo.locator.split(':').nth(4).expect("R14: locator path"); + let start: usize = path + .trim_start_matches('p') + .split('-') + .next() + .and_then(|s| s.parse().ok()) + .unwrap_or_else(|| panic!("R14: depth-0 path p{{start}}-{{end}}: {}", lo.locator)); + assert!( + start > R14_STREAM_1.len(), + "R14: {} must point into part 1 (part 0 is {} bytes)", + lo.locator, + R14_STREAM_1.len() + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/apply.rs b/src-tauri/src/pdf_engine/text_edit/apply.rs new file mode 100644 index 0000000..73ffb8e --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/apply.rs @@ -0,0 +1,181 @@ +//! Applying a page plan through qpdf (SPEC §B.13): the decoded data of each edited content stream +//! is replaced with `qpdf --update-from-json` (an empty stream dictionary drops `/Filter` and +//! `/DecodeParms`; qpdf recomputes `/Length`). `--decode-level=none --compress-streams=n` keeps +//! every untouched stream's bytes exactly as they were: with qpdf's default +//! `--compress-streams=y`, qpdf 12.3 still converts LZW streams to Flate at decode level none +//! (APP-06), which the whole-graph check (A2) would rightly refuse; the edited streams are +//! written unfiltered (later Save passes compress them). lopdf never writes. qpdf renumbers +//! objects in its output, so everything that compares the result is id-free. + +use crate::error::AppError; +use crate::pdf_engine::render; +use crate::pdf_engine::text_edit::content::PageContent; +use crate::pdf_engine::text_edit::engines::{ + classify_check, run_tool, Engines, RunOpts, SourceCheck, +}; +use crate::pdf_engine::text_edit::limits::UPDATE_JSON_MAX_BYTES; +use crate::pdf_engine::text_edit::reasons::{EditProblem, EditProblemCode, ProblemCtx}; +use crate::pdf_engine::text_edit::rewrite::PagePlan; +use lopdf::ObjectId; +use serde_json::{json, Map, Value}; +use std::ffi::OsString; +use std::path::Path; + +/// One content stream to replace: its id in qpdf's input and its new decoded bytes. +#[derive(Debug, Clone, PartialEq)] +pub struct PartUpdate { + pub object_id: ObjectId, + pub decoded: Vec, +} + +/// `EDIT_VERIFY_FAILED` with a technical detail (the apply step is part of verification). +pub(crate) fn verify_failed(detail: &str) -> AppError { + let p = EditProblem::new(EditProblemCode::EditVerifyFailed, Some(detail.to_string())); + let ctx = ProblemCtx { + page_number: None, + file_name: None, + face: None, + reason: None, + }; + EditProblemCode::EditVerifyFailed.to_app_error(&p, &ctx) +} + +/// The edited parts of `plan` against `content`'s stream ids; a part the page or the plan does +/// not have is an internal `EDIT_VERIFY_FAILED` (never skipped). +pub fn updates_for_plan( + content: &PageContent, + plan: &PagePlan, +) -> Result, AppError> { + plan.edited_parts + .iter() + .map( + |i| match (content.parts.get(*i), plan.expected_parts.get(*i)) { + (Some(part), Some(decoded)) => Ok(PartUpdate { + object_id: part.stream_id, + decoded: decoded.clone(), + }), + _ => Err(verify_failed(&format!( + "edited part {i} is not on the page ({} parts, {} planned)", + content.parts.len(), + plan.expected_parts.len() + ))), + }, + ) + .collect() +} + +/// The qpdf JSON v2 update document for `updates` (`max_object_id`: the input's highest id). +pub(crate) fn update_json(updates: &[PartUpdate], max_object_id: u32) -> Value { + let mut objects = Map::new(); + for u in updates { + objects.insert( + format!("obj:{} {} R", u.object_id.0, u.object_id.1), + json!({ "stream": { "dict": {}, "data": render::base64(&u.decoded) } }), + ); + } + json!({ + "qpdf": [ + { + "jsonversion": 2, + "pushedinheritedpageresources": false, + "calledgetallpages": false, + "maxobjectid": max_object_id, + }, + Value::Object(objects), + ] + }) +} + +/// Writes the update JSON (≤ `UPDATE_JSON_MAX_BYTES`, else "too large to verify"). +pub fn write_update_json( + updates: &[PartUpdate], + max_object_id: u32, + out: &Path, +) -> Result<(), AppError> { + let bytes = serde_json::to_vec(&update_json(updates, max_object_id)) + .map_err(|e| verify_failed(&format!("update JSON: {e}")))?; + if bytes.len() > UPDATE_JSON_MAX_BYTES { + return Err(verify_failed("too large to verify: update JSON")); + } + std::fs::write(out, bytes).map_err(|e| AppError::io("OffPDF could not write a work file.", e)) +} + +/// The warning lines of a qpdf run that exited 3, when every one is on the benign allow-list +/// (`engines::classify_check`); `None` when any is not. +pub(crate) fn benign_warnings(stdout: &[u8], stderr: &str) -> Option> { + match classify_check(3, &String::from_utf8_lossy(stdout), stderr) { + SourceCheck::Benign(lines) if !lines.is_empty() => Some(lines), + _ => None, + } +} + +/// A warning line without the file name qpdf puts in front of it (so a source warning and the +/// same warning about a copy compare equal). qpdf writes `WARNING: : …` or +/// `WARNING: (…): …`; the file name ends at the last `.pdf` followed by `:` or ` (`, so a +/// directory or file name that merely contains `.pdf` does not cut the line short. +pub(crate) fn warning_text(line: &str) -> &str { + let line = line.strip_prefix("WARNING: ").unwrap_or(line); + let end = line + .match_indices(".pdf") + .map(|(p, _)| p + ".pdf".len()) + .filter(|e| { + let rest = line.get(*e..).unwrap_or_default(); + rest.starts_with(':') || rest.starts_with(" (") + }) + .last(); + match end { + Some(e) => line.get(e..).unwrap_or(line).trim_start_matches([':', ' ']), + None => line, + } +} + +/// Every warning is one the source itself already had. +pub(crate) fn warnings_known(lines: &[String], source_benign: &[String]) -> bool { + lines.iter().all(|l| { + source_benign + .iter() + .any(|s| warning_text(s) == warning_text(l)) + }) +} + +/// `qpdf --decode-level=none --compress-streams=n --update-from-json=`. +/// Exit 0 → no warnings; exit 3 → every warning must be benign **and** one the source itself had +/// (`source_benign`), and they are returned; anything else → `EDIT_VERIFY_FAILED` (stderr in +/// details). `run_qpdf` (exit 3 = success) is not used here. +pub fn apply_update( + engines: &Engines, + input: &Path, + update: &Path, + output: &Path, + source_benign: &[String], + opts: &RunOpts<'_>, +) -> Result, AppError> { + let mut update_arg = OsString::from("--update-from-json="); + update_arg.push(update.as_os_str()); + let args = [ + input.as_os_str().to_os_string(), + output.as_os_str().to_os_string(), + OsString::from("--decode-level=none"), + OsString::from("--compress-streams=n"), + update_arg, + ]; + let out = run_tool(&engines.qpdf, &args, false, opts)?; + match out.code { + 0 => Ok(Vec::new()), + 3 => { + let benign = benign_warnings(&out.stdout, &out.stderr); + if let Some(lines) = benign.filter(|l| warnings_known(l, source_benign)) { + Ok(lines) + } else { + Err(verify_failed(&format!( + "qpdf update warnings: {}", + out.stderr.trim() + ))) + } + } + code => Err(verify_failed(&format!( + "qpdf update exited with code {code}: {}", + out.stderr.trim() + ))), + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/bench.rs b/src-tauri/src/pdf_engine/text_edit/bench.rs new file mode 100644 index 0000000..d0d047b --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/bench.rs @@ -0,0 +1,378 @@ +//! BENCH-01/02 (SPEC §E.7; `#[ignore]`, run locally): timings of open, inspect, preview and Save +//! on generated documents, with the peak resident set size (`getrusage`) of a child process per +//! configuration, printed as Markdown tables for `docs/EDIT_TEXT.md` §12. +//! +//! ```text +//! cargo test --lib -j 6 text_edit::bench -- --ignored --nocapture # debug build +//! cargo test --release --lib -j 6 text_edit::bench -- --ignored --nocapture # release build +//! ``` +//! +//! BENCH-01: 10 / 100 / 1,000 pages of 40 Helvetica lines — open, inspect p50/p95 (cold: first +//! visit of a page; warm: the cached model), preview p50/p95, Save with 1 / 10 / 100 changed +//! lines. BENCH-02: a ~300 MB image-heavy file (Flate and DCT images, 20 text pages) — open to +//! first inspect, the background `qpdf --check`, one preview, Save with 1 change; budgets: first +//! inspect ≤ 3 s, Save ≤ 90 s, peak RSS ≤ 3 × file size. +#[cfg(test)] +// the module is test-only (declared under #[cfg(test)]); this line ends the API scan +use crate::models::PageGroup; +use crate::pdf_engine::edit_overlay::{export_edit_pdf_with_check_exe, EditDocumentIn}; +use crate::pdf_engine::qpdf::resolve_qpdf_standalone; +use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; +use crate::pdf_engine::text_edit::service; +use crate::pdf_engine::text_edit::testkit::producers::{DocBuilder, PageSpec, HELVETICA}; +use crate::pdf_engine::text_edit::testkit::{child_mode, run_child_test}; +use crate::pdf_engine::text_edit::tests_e2e::preview::edit_in; +use crate::pdf_engine::text_edit::tests_e2e::{font_path, run_qpdf, E2e}; +use serde_json::{json, Value}; +use std::path::Path; +use std::time::{Duration, Instant}; + +const LINES_PER_PAGE: usize = 40; +const INSPECT_SAMPLE: usize = 20; +const PREVIEW_SAMPLE: usize = 10; + +/// Peak resident set size of this process in MiB (`getrusage(RUSAGE_SELF)`). +fn peak_rss_mib() -> f64 { + #[repr(C)] + struct Rusage { + words: [i64; 18], // 2 × timeval (16 bytes each on macOS and Linux) + 14 longs + } + extern "C" { + fn getrusage(who: i32, usage: *mut Rusage) -> i32; + } + let mut usage = Rusage { words: [0; 18] }; + // SAFETY: `usage` is a writable buffer at least as large as `struct rusage` on the 64-bit + // macOS and Linux targets (144 bytes); RUSAGE_SELF = 0. + let ok = unsafe { getrusage(0, &mut usage) } == 0; + assert!(ok, "getrusage failed"); + let maxrss = usage.words[4] as f64; // ru_maxrss: bytes on macOS, KiB on Linux + if cfg!(target_os = "macos") { + maxrss / (1024.0 * 1024.0) + } else { + maxrss / 1024.0 + } +} + +fn ms(d: Duration) -> f64 { + d.as_secs_f64() * 1000.0 +} + +/// (p50, p95) of `samples` in ms. +fn percentiles(samples: &mut [f64]) -> (f64, f64) { + samples.sort_by(f64::total_cmp); + let at = |q: f64| { + let i = ((samples.len() as f64 - 1.0) * q).round() as usize; + samples.get(i).copied().unwrap_or(f64::NAN) + }; + (at(0.5), at(0.95)) +} + +fn line_text(page: usize, line: usize) -> String { + format!("Line {line} on page {page}: the quick brown fox") +} + +/// `pages` pages of `LINES_PER_PAGE` Helvetica lines (one shared font). +fn text_doc(d: &mut DocBuilder, font: u32, pages: usize) { + for p in 0..pages { + let mut content = String::new(); + for l in 0..LINES_PER_PAGE { + let y = 760 - 18 * l; + let text = line_text(p, l); + content.push_str(&format!("BT /F1 10 Tf 60 {y} Td ({text}) Tj ET\n")); + } + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {font} 0 R >>"), + )); + } +} + +/// A `sourceText` export object changing line `line` of page `page` ("fox" → "dog"). +fn edit_object(t: &E2e, src: &Path, page: u32, line: usize) -> Value { + let old = line_text(page as usize, line); + t.edit(src, page, page, &old, &old.replace("fox", "dog")) +} + +/// One save through the real export; returns its duration. +fn timed_save(t: &E2e, src: &Path, objects: Vec, name: &str) -> Duration { + let doc: EditDocumentIn = + serde_json::from_value(json!({ "version": 1, "objects": objects })).expect("document"); + let groups = [PageGroup { + path: src.to_string_lossy().into_owned(), + pages: "1-z".into(), + }]; + let work = t.scratch.path(&format!("{name}-work")); + std::fs::create_dir_all(&work).expect("work"); + let dest = t.scratch.path(&format!("{name}-out.pdf")); + let qpdf = resolve_qpdf_standalone(); + let started = Instant::now(); + let r = export_edit_pdf_with_check_exe( + &groups, + &dest.to_string_lossy(), + &doc, + &font_path(), + &work, + name, + None, + &qpdf, + None, + None, + &[], + &[], + false, + false, + |args| run_qpdf(&qpdf, args), + ); + let took = started.elapsed(); + r.unwrap_or_else(|e| panic!("{name}: {e} {:?}", e.details)); + let _ = std::fs::remove_dir_all(&work); + let _ = std::fs::remove_file(&dest); + took +} + +/// Preview of one changed line on `page`; returns its duration. +fn timed_preview(t: &E2e, src: &Path, fp: &str, page: u32) -> Duration { + let dto = t.try_inspect(src, fp, page).expect("inspect"); + let old = line_text(page as usize, 0); + let run = dto.runs.iter().find(|r| r.text == old).expect("line 0"); + let edits = [edit_in( + &run.id, + &old, + &old.replace("fox", "dog"), + SourceTextStyleIn::default(), + )]; + let started = Instant::now(); + let p = service::preview_edits( + &t.cache, + &t.engines, + &t.temp_root(), + &src.to_string_lossy(), + fp, + page, + &edits, + ) + .expect("preview"); + let took = started.elapsed(); + assert!( + p.page_pdf.is_some() && p.verdicts.iter().all(|v| v.ok), + "{:?}", + p.verdicts + ); + took +} + +fn bench_01_child(pages: usize) { + let t = E2e::new("bench_01").expect("engines"); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + text_doc(&mut d, f, pages); + let pdf = d.build(); + let src = t.file("bench.pdf", &pdf); + let started = Instant::now(); + let info = t.open(&src); + let open = started.elapsed(); + let k = INSPECT_SAMPLE.min(pages); + let sample: Vec = (0..k).map(|i| (i * pages / k) as u32).collect(); + let mut cold = Vec::new(); + let mut warm = Vec::new(); + for pass in [&mut cold, &mut warm] { + for page in &sample { + let started = Instant::now(); + t.try_inspect(&src, &info.fingerprint, *page) + .expect("inspect"); + pass.push(ms(started.elapsed())); + } + } + let mut previews: Vec = sample + .iter() + .take(PREVIEW_SAMPLE) + .map(|p| ms(timed_preview(&t, &src, &info.fingerprint, *p))) + .collect(); + let mut saves = Vec::new(); + for n in [1usize, 10, 100] { + let objects = (0..n) + .map(|i| edit_object(&t, &src, (i % pages) as u32, i / pages)) + .collect(); + saves.push(ms(timed_save(&t, &src, objects, &format!("save-{n}")))); + } + let (c50, c95) = percentiles(&mut cold); + let (w50, w95) = percentiles(&mut warm); + let (p50, p95) = percentiles(&mut previews); + println!( + "BENCH01|{pages}|{:.1}|{:.0}|{c50:.0}|{c95:.0}|{w50:.0}|{w95:.0}|{p50:.0}|{p95:.0}|{:.0}|{:.0}|{:.0}|{:.0}", + pdf.len() as f64 / 1024.0, + ms(open), + saves[0], + saves[1], + saves[2], + peak_rss_mib() + ); +} + +/// BENCH-01 (§E.7): one child process per document size; prints the Markdown table. +#[test] +#[ignore] +fn bench_01_open_inspect_preview_save_by_page_count() { + if let Some(mode) = child_mode() { + if let Some(n) = mode.strip_prefix("bench01:").and_then(|n| n.parse().ok()) { + bench_01_child(n); + } + return; + } + let mut rows = Vec::new(); + for n in [10, 100, 1000] { + let out = run_child_test( + "pdf_engine::text_edit::bench::bench_01_open_inspect_preview_save_by_page_count", + &format!("bench01:{n}"), + ); + let row = out + .lines() + .find_map(|l| l.split("BENCH01|").nth(1)) + .unwrap_or_else(|| panic!("no BENCH01 row:\n{out}")) + .replace('|', " | "); + rows.push(format!("| {row} |")); + } + println!( + "\nBENCH-01 ({} build)\n", + if cfg!(debug_assertions) { + "debug" + } else { + "release" + } + ); + println!("| pages | file KiB | open ms | inspect cold p50 | cold p95 | warm p50 | warm p95 | preview p50 | preview p95 | save 1 edit ms | save 10 | save 100 | peak RSS MiB |"); + println!("|---|---|---|---|---|---|---|---|---|---|---|---|---|"); + for r in rows { + println!("{r}"); + } +} + +/// Deterministic incompressible bytes. +fn noise(n: usize, seed: u64) -> Vec { + let mut state = seed | 1; + let mut v = Vec::with_capacity(n + 8); + while v.len() < n { + state ^= state << 13; + state ^= state >> 7; + state ^= state << 17; + v.extend_from_slice(&state.to_le_bytes()); + } + v.truncate(n); + v +} + +/// zlib with stored (uncompressed) deflate blocks: valid Flate data as large as its input. +fn zlib_stored(raw: &[u8]) -> Vec { + use std::io::Write; + let mut z = flate2::write::ZlibEncoder::new(Vec::new(), flate2::Compression::none()); + z.write_all(raw).expect("zlib"); + z.finish().expect("zlib") +} + +/// A baseline JPEG of `side` × `side` noise pixels (valid DCT data: `qpdf --check` decodes it). +fn noise_jpeg(side: u32, seed: u64) -> Vec { + let pixels = noise((side * side * 3) as usize, seed); + let mut out = Vec::new(); + image::codecs::jpeg::JpegEncoder::new_with_quality(&mut out, 95) + .encode(&pixels, side, side, image::ExtendedColorType::Rgb8) + .expect("jpeg"); + out +} + +/// About `total_mib` MiB: 20 text pages, then image pages each holding a Flate image +/// (incompressible data in stored deflate blocks, 5 MiB) and a DCT image (a noise JPEG). +fn heavy_doc(total_mib: usize) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + text_doc(&mut d, f, 20); + let side = 1_024usize; + let rows = (5usize << 20) / (side * 3); + let jpeg_side = 1_400u32; + let jpeg = noise_jpeg(jpeg_side, 0xD0C0); + let mut total = 0usize; + let mut i = 0u64; + while total < total_mib << 20 { + let raw = noise(side * rows * 3, 0x5EED_0000 + i); + let flate = zlib_stored(&raw); + total += flate.len() + jpeg.len(); + let fl = d.b.add_stream( + &format!("/Type /XObject /Subtype /Image /Width {side} /Height {rows} /ColorSpace /DeviceRGB /BitsPerComponent 8 /Filter /FlateDecode"), + &flate, + ); + let dct = d.b.add_stream( + &format!("/Type /XObject /Subtype /Image /Width {jpeg_side} /Height {jpeg_side} /ColorSpace /DeviceRGB /BitsPerComponent 8 /Filter /DCTDecode"), + &jpeg, + ); + d.page(PageSpec::new( + b"q 300 0 0 300 50 400 cm /Im0 Do Q q 300 0 0 300 50 50 cm /Im1 Do Q", + &format!("/XObject << /Im0 {fl} 0 R /Im1 {dct} 0 R >>"), + )); + i += 1; + } + d.build() +} + +fn bench_02_child() { + let t = E2e::new("bench_02").expect("engines"); + let pdf = heavy_doc(300); + let size_mib = pdf.len() as f64 / (1024.0 * 1024.0); + let src = t.file("heavy.pdf", &pdf); + drop(pdf); + let started = Instant::now(); + let info = t.open(&src); + let opened = started.elapsed(); + t.try_inspect(&src, &info.fingerprint, 0) + .expect("first inspect"); + let first_inspect = started.elapsed(); + { + // Held only for the wait: a file over CACHE_SNAPSHOT_BYTES_MAX is not kept by the cache, + // and nothing in the app holds its snapshot during a preview or a Save. + let src_ref = t + .cache + .open(&src, &t.temp_root(), &t.engines) + .expect("source"); + t.cache + .source_check(&src_ref, &t.engines) + .expect("check") + .wait(None) + .expect("check result"); + } + let check = started.elapsed(); + let preview = timed_preview(&t, &src, &info.fingerprint, 3); + let save = timed_save(&t, &src, vec![edit_object(&t, &src, 5, 7)], "heavy-save"); + let rss = peak_rss_mib(); + println!( + "BENCH02|{size_mib:.0}|{:.0}|{:.0}|{:.0}|{:.0}|{:.0}|{rss:.0}|{:.2}", + ms(opened), + ms(first_inspect - opened), + ms(check), + ms(preview), + ms(save), + rss / size_mib + ); +} + +/// BENCH-02 (§E.7): the ~300 MB image-heavy file, in a child process; prints the table. +#[test] +#[ignore] +fn bench_02_image_heavy_300_mb() { + if let Some(mode) = child_mode() { + if mode == "bench02" { + bench_02_child(); + } + return; + } + let out = run_child_test( + "pdf_engine::text_edit::bench::bench_02_image_heavy_300_mb", + "bench02", + ); + let row = out + .lines() + .find_map(|l| l.split("BENCH02|").nth(1)) + .unwrap_or_else(|| panic!("no BENCH02 row:\n{out}")) + .replace('|', " | "); + println!("\nBENCH-02 ({} build; budgets: first inspect ≤ 3,000 ms, save ≤ 90,000 ms, peak RSS ≤ 3 × file)\n", if cfg!(debug_assertions) { "debug" } else { "release" }); + println!("| file MiB | open ms | first inspect ms (after open) | qpdf --check done ms (from the start of open) | preview ms | save 1 edit ms | peak RSS MiB | RSS / file |"); + println!("|---|---|---|---|---|---|---|---|"); + println!("| {row} |"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/cache.rs b/src-tauri/src/pdf_engine/text_edit/cache.rs new file mode 100644 index 0000000..661d2f9 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/cache.rs @@ -0,0 +1,591 @@ +//! Cache and concurrency (SPEC §B.19): the snapshots of the files open in Edit text, their page +//! models and their background `qpdf --check`, shared by the inspect and preview commands. +//! +//! - Keyed by canonical path; an entry is reused only while `stat_matches` holds (length, mtime +//! and the first/last 64 KiB unchanged), otherwise the file is read again. +//! - At most `CACHE_SNAPSHOTS_MAX` snapshots are kept, only of files of at most +//! `CACHE_SNAPSHOT_BYTES_MAX` bytes. A larger file is not kept between calls, but calls that +//! overlap share one read of it, and only one read of a file runs at a time (review-T5 M3). +//! Each source keeps at most `CACHE_PAGE_MODELS_MAX` page models and at most +//! `CACHE_SNAPSHOT_BYTES_MAX` bytes of them (each model's `walk.model_bytes`: everything it +//! keeps, its font models included). +//! - Every source gets `/textedit//` with `source.pdf`, written once from the +//! snapshot bytes (temp name + rename): qpdf never reads the user's file. The page-map agreement +//! check (`qpdf --json` vs lopdf) runs there before anything is inspected; `qpdf --check` runs +//! in the background and only the preview waits for it (Save reuses its memoised result). +//! - The mutex is held only to look up or swap `Arc`s; reads, parses and subprocesses run +//! outside it. Evicting or releasing a source deletes its folder unless another open source has +//! the same bytes; a folder whose background check is still reading `source.pdf` is deleted by +//! a later call. A running preview leases its folder: a release or an eviction during it +//! deletes the folder when the preview ends, not before (review-T5 M2); folders left by a crash +//! are deleted at startup (`clear_stale_folders`). Save never uses this cache. + +use crate::error::AppError; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::engines::{qpdf_page_map, Engines, PendingCheck, RunOpts}; +use crate::pdf_engine::text_edit::limits::{ + CACHE_PAGE_MODELS_MAX, CACHE_SNAPSHOTS_MAX, CACHE_SNAPSHOT_BYTES_MAX, CHECK_MEMO_MAX, + PREVIEW_SHARED_MODEL_MAX, +}; +use crate::pdf_engine::text_edit::preview::{cache_dir_for, write_once}; +use crate::pdf_engine::text_edit::reasons; +use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel}; +use crate::pdf_engine::text_edit::snapshot::{ + check_page_map, read_snapshot, stat_matches, Fingerprint, SourceSnapshot, +}; +use std::collections::VecDeque; +use std::path::{Path, PathBuf}; +use std::sync::{Arc, Mutex, MutexGuard, Weak}; + +/// Byte budget of the page models one source keeps. +const PAGE_MODEL_BYTES_MAX: usize = CACHE_SNAPSHOT_BYTES_MAX as usize; + +/// One open source: its snapshot context, page models and temporary folder. +pub struct CachedSource { + pub ctx: SnapshotContext, + /// `/textedit//`. + pub dir: PathBuf, + pages: Mutex, +} + +#[derive(Default)] +struct LruPages { + /// (page index, model, approximate bytes), most recently used first. + entries: VecDeque<(u32, Arc, usize)>, + bytes: usize, +} + +impl LruPages { + fn get(&mut self, page: u32) -> Option> { + let pos = self.entries.iter().position(|(p, ..)| *p == page)?; + let entry = self.entries.remove(pos)?; + let model = Arc::clone(&entry.1); + self.entries.push_front(entry); + Some(model) + } + + /// Keeps `model` unless it alone exceeds the budget; a model another call stored meanwhile + /// wins (both were built from the same bytes). + fn insert(&mut self, page: u32, model: Arc) -> Arc { + if let Some(existing) = self.get(page) { + return existing; + } + let bytes = model.walk.model_bytes; + if bytes > PAGE_MODEL_BYTES_MAX { + return model; + } + self.entries.push_front((page, Arc::clone(&model), bytes)); + self.bytes = self.bytes.saturating_add(bytes); + while self.entries.len() > CACHE_PAGE_MODELS_MAX || self.bytes > PAGE_MODEL_BYTES_MAX { + match self.entries.pop_back() { + Some((_, _, b)) => self.bytes = self.bytes.saturating_sub(b), + None => break, + } + } + model + } + + /// Removes `page`'s model, if cached. + fn take(&mut self, page: u32) -> Option> { + let pos = self.entries.iter().position(|(p, ..)| *p == page)?; + let (_, model, bytes) = self.entries.remove(pos)?; + self.bytes = self.bytes.saturating_sub(bytes); + Some(model) + } +} + +struct Entry { + key: PathBuf, + src: Arc, +} + +/// An open source too large to cache: its folder stays until release; the source itself lives +/// only while a call holds it (`src`), and calls that overlap share it. +struct Pinned { + key: PathBuf, + dir: PathBuf, + src: Weak, +} + +#[derive(Default)] +struct CacheInner { + /// Cached snapshots, most recently used first. + sources: VecDeque, + pinned: Vec, + /// Folders a running preview writes into, with the number of previews. + leases: Vec<(PathBuf, usize)>, + /// One load lock per key being read (single flight). + loading: Vec<(PathBuf, Arc>)>, + /// Background checks by fingerprint and folder (shared by every open of the same bytes). + checks: VecDeque<(Fingerprint, PathBuf, PendingCheck)>, + /// Fingerprints whose page map lopdf and qpdf agree on. + maps_ok: VecDeque, + /// Unused folders whose background check was still reading `source.pdf`. + deferred: Vec, + #[cfg(test)] + seams: seams::CacheSeams, +} + +impl CacheInner { + fn in_use(&self, dir: &Path) -> bool { + self.sources.iter().any(|e| e.src.dir == dir) + || self.pinned.iter().any(|p| p.dir == dir) + || self.leases.iter().any(|(d, _)| d == dir) + } + + /// The source of `key` (cached, or pinned and held by a running call). + fn lookup(&self, key: &Path) -> Option> { + let cached = self.sources.iter().find(|e| e.key == key); + cached.map(|e| Arc::clone(&e.src)).or_else(|| { + self.pinned + .iter() + .find(|p| p.key == key) + .and_then(|p| p.src.upgrade()) + }) + } + + /// Whether a file of `len` bytes is kept between calls. + fn cacheable(&self, len: u64) -> bool { + #[cfg(test)] + if let Some(max) = self.seams.snapshot_bytes_max { + return len <= max; + } + len <= CACHE_SNAPSHOT_BYTES_MAX + } + + fn busy(&self, dir: &Path) -> bool { + self.checks + .iter() + .any(|(_, d, c)| d == dir && c.peek().is_none()) + } + + /// The folders to delete now: `candidates` and earlier deferred folders that no open source + /// uses; busy ones are deferred to a later call. + fn collect(&mut self, candidates: Vec) -> Vec { + let mut all = std::mem::take(&mut self.deferred); + all.extend(candidates); + all.sort(); + all.dedup(); + let mut now = Vec::new(); + for dir in all { + if self.in_use(&dir) { + continue; + } + if self.busy(&dir) { + self.deferred.push(dir); + } else { + now.push(dir); + } + } + now + } + + /// Removes the entries of `key`, returning their folders. + fn forget(&mut self, key: &Path) -> Vec { + let mut dirs = Vec::new(); + self.sources.retain(|e| { + let keep = e.key != key; + if !keep { + dirs.push(e.src.dir.clone()); + } + keep + }); + self.pinned.retain(|p| { + let keep = p.key != key; + if !keep { + dirs.push(p.dir.clone()); + } + keep + }); + dirs + } +} + +/// Keeps a source's folder while a preview writes into it (`TextEditCache::lease`). +pub struct DirLease { + cache: TextEditCache, + dir: PathBuf, +} + +impl Drop for DirLease { + /// The last lease of a folder that no open source owns any more deletes it (or defers it while + /// its check runs). + fn drop(&mut self) { + let doomed = { + let mut inner = self.cache.lock(); + if let Some(pos) = inner.leases.iter().position(|(d, _)| *d == self.dir) { + let left = inner + .leases + .get(pos) + .map_or(0, |(_, n)| n.saturating_sub(1)); + match inner.leases.get_mut(pos) { + Some(lease) if left > 0 => lease.1 = left, + _ => { + inner.leases.remove(pos); + } + } + } + inner.collect(vec![self.dir.clone()]) + }; + remove_dirs(&doomed); + } +} + +/// Deletes `/textedit/` (startup: no source is open yet, so every copy there was left +/// by a crash or an interrupted call). +pub fn clear_stale_folders(temp_root: &Path) { + remove_dirs(&[temp_root.join("textedit")]); +} + +/// Snapshots of the files open in Edit text (managed Tauri state). +#[derive(Clone, Default)] +pub struct TextEditCache { + inner: Arc>, +} + +fn key_of(path: &Path) -> PathBuf { + std::fs::canonicalize(path).unwrap_or_else(|_| path.to_path_buf()) +} + +/// The file name shown in `STALE` messages. +pub(crate) fn file_name(path: &Path) -> String { + path.file_name() + .map(|n| n.to_string_lossy().into_owned()) + .unwrap_or_else(|| path.to_string_lossy().into_owned()) +} + +fn remove_dirs(dirs: &[PathBuf]) { + for dir in dirs { + if let Err(e) = std::fs::remove_dir_all(dir) { + if e.kind() != std::io::ErrorKind::NotFound { + eprintln!( + "offpdf: could not remove the text-edit folder {}: {e}", + dir.display() + ); + } + } + } +} + +impl TextEditCache { + fn lock(&self) -> MutexGuard<'_, CacheInner> { + self.inner.lock().unwrap_or_else(|e| e.into_inner()) + } + + /// The source at `path`, read again when `stat_matches` fails: `read_snapshot` → write + /// `source.pdf` once → page-map agreement (synchronous; a disagreement is `PDF_NEEDS_REPAIR`) + /// → background `qpdf --check` (memoised by fingerprint). Returns before the check finishes. + pub fn open( + &self, + path: &Path, + temp_root: &Path, + engines: &Engines, + ) -> Result, AppError> { + let key = key_of(path); + if let Some(src) = self.current(&key) { + return Ok(src); + } + // Single flight: one read of a file at a time; a read that finished meanwhile is shared. + let gate = { + let mut inner = self.lock(); + match inner.loading.iter().find(|(k, _)| *k == key) { + Some((_, g)) => Arc::clone(g), + None => { + let g = Arc::new(Mutex::new(())); + inner.loading.push((key.clone(), Arc::clone(&g))); + g + } + } + }; + let loaded = { + let _one = gate.lock().unwrap_or_else(|e| e.into_inner()); + match self.current(&key) { + Some(src) => Ok(src), + None => self.load(path, key.clone(), temp_root, engines), + } + }; + let mut inner = self.lock(); + // The map's copy and ours: nobody else waits on this key. + if Arc::strong_count(&gate) <= 2 { + inner.loading.retain(|(k, _)| *k != key); + } + loaded + } + + /// The source of `key` when it is still the file on disk (`stat_matches`). + fn current(&self, key: &Path) -> Option> { + let src = self.lock().lookup(key)?; + if !stat_matches(&src.ctx.snap) { + return None; + } + self.touch(key); + Some(src) + } + + /// Keeps `dir` (a source's folder) while the returned guard lives: a preview writes into it. + pub fn lease(&self, dir: &Path) -> DirLease { + let mut inner = self.lock(); + match inner.leases.iter_mut().find(|(d, _)| d == dir) { + Some((_, n)) => *n = n.saturating_add(1), + None => inner.leases.push((dir.to_path_buf(), 1)), + } + DirLease { + cache: self.clone(), + dir: dir.to_path_buf(), + } + } + + /// `open`, then `STALE` unless the file still has `fingerprint`. + pub fn get( + &self, + path: &Path, + temp_root: &Path, + fingerprint: &str, + engines: &Engines, + ) -> Result, AppError> { + let src = self.open(path, temp_root, engines)?; + if src.ctx.snap.fingerprint.to_string() != fingerprint { + return Err(reasons::stale(&file_name(path))); + } + Ok(src) + } + + /// The model of `page_index` (`INVALID_PAGES` when the file has no such page). + pub fn page(&self, src: &CachedSource, page_index: u32) -> Result, AppError> { + let lock = || src.pages.lock().unwrap_or_else(|e| e.into_inner()); + if let Some(model) = lock().get(page_index) { + return Ok(model); + } + let model = Arc::new(build_page_model(&src.ctx, page_index, None)?); + Ok(lock().insert(page_index, model)) + } + + /// The model of `page_index` for a call that may release it: a model over + /// `PREVIEW_SHARED_MODEL_MAX` leaves the cache (built again on the next visit), so the caller + /// holds the last reference and frees it when done (review-final MEDIUM-3). + pub fn page_to_release( + &self, + src: &CachedSource, + page_index: u32, + ) -> Result, AppError> { + let model = self.page(src, page_index)?; + if model.walk.model_bytes > PREVIEW_SHARED_MODEL_MAX { + let mut pages = src.pages.lock().unwrap_or_else(|e| e.into_inner()); + pages.take(page_index); + } + Ok(model) + } + + /// The background `qpdf --check` of `src` (started at open). A check that ended in an error + /// (timeout, missing qpdf) is started again; `source.pdf` is written again first when an + /// eviction removed it. + pub fn source_check( + &self, + src: &CachedSource, + engines: &Engines, + ) -> Result { + let source = src.dir.join("source.pdf"); + std::fs::create_dir_all(&src.dir) + .map_err(|e| AppError::io("OffPDF could not create a temporary folder.", e))?; + write_once(&source, &src.ctx.snap.bytes)?; + Ok(self.pending_check(engines, src.ctx.snap.fingerprint, &src.dir)) + } + + /// Forgets `path` and deletes its folder unless another open source has the same bytes + /// (best effort; a failure is logged). + pub fn release(&self, path: &Path) { + let key = key_of(path); + let doomed = { + let mut inner = self.lock(); + let candidates = inner.forget(&key); + inner.collect(candidates) + }; + remove_dirs(&doomed); + } + + fn touch(&self, key: &Path) { + let mut inner = self.lock(); + if let Some(pos) = inner.sources.iter().position(|e| e.key == key) { + if let Some(entry) = inner.sources.remove(pos) { + inner.sources.push_front(entry); + } + } + } + + fn load( + &self, + path: &Path, + key: PathBuf, + temp_root: &Path, + engines: &Engines, + ) -> Result, AppError> { + let snap = read_snapshot(path)?; + #[cfg(test)] + { + self.lock().seams.loads += 1; + } + let fp = snap.fingerprint; + let dir = cache_dir_for(temp_root, &fp.to_string()); + let prepared = std::fs::create_dir_all(&dir) + .map_err(|e| AppError::io("OffPDF could not create a temporary folder.", e)) + .and_then(|()| write_once(&dir.join("source.pdf"), &snap.bytes)) + .and_then(|()| self.check_map(&snap, &dir, engines)); + if let Err(e) = prepared { + let doomed = self.lock().collect(vec![dir]); + remove_dirs(&doomed); + return Err(e); + } + self.pending_check(engines, fp, &dir); + let src = Arc::new(CachedSource { + ctx: SnapshotContext::new(snap), + dir, + pages: Mutex::new(LruPages::default()), + }); + let doomed = { + let mut inner = self.lock(); + let mut candidates = inner.forget(&key); + if inner.cacheable(fp.len) { + inner.sources.push_front(Entry { + key, + src: Arc::clone(&src), + }); + while inner.sources.len() > CACHE_SNAPSHOTS_MAX { + if let Some(evicted) = inner.sources.pop_back() { + candidates.push(evicted.src.dir.clone()); + } + } + } else { + inner.pinned.push(Pinned { + key, + dir: src.dir.clone(), + src: Arc::downgrade(&src), + }); + } + inner.collect(candidates) + }; + remove_dirs(&doomed); + Ok(src) + } + + /// Page-map agreement of the snapshot with qpdf's reading of `dir/source.pdf` (once per + /// fingerprint: the answer depends only on the bytes). + fn check_map( + &self, + snap: &SourceSnapshot, + dir: &Path, + engines: &Engines, + ) -> Result<(), AppError> { + let fp = snap.fingerprint; + if self.lock().maps_ok.contains(&fp) { + return Ok(()); + } + let pages = qpdf_page_map(engines, &dir.join("source.pdf"), &RunOpts::default())?; + check_page_map(snap, &pages)?; + let mut inner = self.lock(); + if !inner.maps_ok.contains(&fp) { + inner.maps_ok.push_front(fp); + inner.maps_ok.truncate(CHECK_MEMO_MAX); + } + Ok(()) + } + + /// The background `qpdf --check` of `dir/source.pdf`, started once per fingerprint and + /// folder (again after an error). + fn pending_check(&self, engines: &Engines, fp: Fingerprint, dir: &Path) -> PendingCheck { + let mut inner = self.lock(); + if let Some((_, _, check)) = inner.checks.iter().find(|(f, d, _)| *f == fp && d == dir) { + if !matches!(check.peek(), Some(Err(_))) { + return check.clone(); + } + } + inner.checks.retain(|(f, d, _)| *f != fp || d != dir); + #[cfg(test)] + let engines = &seams::background_engines(engines); + let check = PendingCheck::spawn(engines.clone(), fp, dir.join("source.pdf")); + inner + .checks + .push_front((fp, dir.to_path_buf(), check.clone())); + // Only finished checks are dropped: a running one must stay visible to `busy`. + while inner.checks.len() > CHECK_MEMO_MAX { + match inner + .checks + .iter() + .rposition(|(_, _, c)| c.peek().is_some()) + { + Some(i) => { + inner.checks.remove(i); + } + None => break, + } + } + check + } + + /// Number of cached snapshots (tests). + #[cfg(test)] + pub(crate) fn cached_len(&self) -> usize { + self.lock().sources.len() + } + + /// Waits until every background check this cache started has finished (tests delete their + /// folders afterwards; a check whose input vanished must never be what a test observes). + #[cfg(test)] + pub(crate) fn wait_checks(&self) { + let checks: Vec = self.lock().checks.iter().map(|c| c.2.clone()).collect(); + for check in checks { + let _ = check.wait(None); + } + } + + /// Bytes the page models of `src` are charged in its LRU (tests). + #[cfg(test)] + pub(crate) fn page_model_bytes(&self, src: &CachedSource) -> usize { + src.pages.lock().unwrap_or_else(|e| e.into_inner()).bytes + } + + /// Background checks still running (tests). + #[cfg(test)] + pub(crate) fn running_checks(&self) -> usize { + let inner = self.lock(); + inner.checks.iter().filter(|c| c.2.peek().is_none()).count() + } + + /// Test knobs of this cache only (other tests' caches are unaffected). + #[cfg(test)] + pub(crate) fn with_seams(&self, f: impl FnOnce(&mut seams::CacheSeams) -> T) -> T { + f(&mut self.lock().seams) + } +} + +#[cfg(test)] +pub(crate) mod seams { + use crate::pdf_engine::text_edit::engines::Engines; + use std::cell::RefCell; + use std::path::PathBuf; + + #[derive(Default)] + pub(crate) struct CacheSeams { + /// Replaces `CACHE_SNAPSHOT_BYTES_MAX` for this cache. + pub snapshot_bytes_max: Option, + /// Number of file reads (`read_snapshot`) this cache made. + pub loads: usize, + } + + thread_local! { + static BACKGROUND_QPDF: RefCell> = const { RefCell::new(None) }; + } + + /// The qpdf the background checks started on this thread run (CMD-08: a check held until + /// the test lets it go). + pub(crate) fn set_background_qpdf(qpdf: Option) { + BACKGROUND_QPDF.with(|q| *q.borrow_mut() = qpdf); + } + + pub(super) fn background_engines(engines: &Engines) -> Engines { + let mut out = engines.clone(); + if let Some(q) = BACKGROUND_QPDF.with(|q| q.borrow().clone()) { + out.qpdf = q; + } + out + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/content.rs b/src-tauri/src/pdf_engine/text_edit/content.rs new file mode 100644 index 0000000..3411061 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/content.rs @@ -0,0 +1,321 @@ +//! Page content and ownership (SPEC §B.7): a page's `/Contents` parts decoded under budget and +//! joined with one `\n` between parts (all spans are offsets into the joined buffer), qpdf's +//! join rule for its overlay wrapper, and the reference/`/Kids` counts that decide whether a +//! content part belongs to this page alone. + +mod joins; + +pub(crate) use joins::{check_part_joins, first_unsafe_boundary}; + +use crate::pdf_engine::text_edit::decode::{self, DecodeBudget}; +use crate::pdf_engine::text_edit::lexer::Span; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::snapshot::fnv1a_extend; +use crate::pdf_engine::validate_output::{content_digest, ContentDigest}; +use lopdf::{Document, Object, ObjectId}; +use std::collections::{HashMap, HashSet}; + +const FNV_OFFSET: u64 = 0xcbf2_9ce4_8422_2325; +const PART_KEYS: [&[u8]; 4] = [b"Length", b"Filter", b"DecodeParms", b"DL"]; +const REF_DEPTH_MAX: usize = 64; + +#[derive(Debug, Clone)] +pub struct ContentPart { + pub stream_id: ObjectId, + pub start: usize, + pub len: usize, + pub digest: ContentDigest, +} + +#[derive(Debug, Clone)] +pub struct PageContent { + pub page_id: ObjectId, + pub parts: Vec, + pub joined: Vec, + pub contents_array: Option, +} + +impl PageContent { + /// (part index, local span) or None when the span covers a separator. + pub fn locate(&self, span: &Span) -> Option<(usize, Span)> { + self.parts.iter().enumerate().find_map(|(i, p)| { + let end = p.start.checked_add(p.len)?; + (span.start >= p.start && span.end <= end && span.start <= span.end) + .then(|| (i, span.start - p.start..span.end - p.start)) + }) + } + + pub fn part_bytes(&self, i: usize) -> &[u8] { + self.parts + .get(i) + .and_then(|p| self.joined.get(p.start..p.start.checked_add(p.len)?)) + .unwrap_or_default() + } + + /// Digest of the parts concatenated WITHOUT separator (= #34 `get_page_content` semantics). + pub fn concat_digest(&self) -> ContentDigest { + let mut hash = FNV_OFFSET; + let mut len = 0usize; + for i in 0..self.parts.len() { + let bytes = self.part_bytes(i); + hash = fnv1a_extend(hash, bytes); + len = len.saturating_add(bytes.len()); + } + ContentDigest { hash, len } + } + + /// The same page with some parts' decoded bytes replaced (unknown indices are ignored). + pub fn with_replaced_parts(&self, replaced: &[(usize, Vec)]) -> PageContent { + let parts: Vec<(ObjectId, &[u8])> = (0..self.parts.len()) + .filter_map(|i| { + let id = self.parts.get(i)?.stream_id; + let bytes = replaced + .iter() + .rev() + .find(|(j, _)| *j == i) + .map(|(_, b)| b.as_slice()); + Some((id, bytes.unwrap_or_else(|| self.part_bytes(i)))) + }) + .collect(); + assemble(self.page_id, &parts, self.contents_array) + } +} + +fn assemble( + page_id: ObjectId, + parts: &[(ObjectId, &[u8])], + contents_array: Option, +) -> PageContent { + let total: usize = + parts.iter().map(|(_, b)| b.len()).sum::() + parts.len().saturating_sub(1); + let mut joined = Vec::with_capacity(total); + let mut out = Vec::with_capacity(parts.len()); + for (i, (stream_id, bytes)) in parts.iter().enumerate() { + if i > 0 { + joined.push(b'\n'); + } + out.push(ContentPart { + stream_id: *stream_id, + start: joined.len(), + len: bytes.len(), + digest: content_digest(bytes), + }); + joined.extend_from_slice(bytes); + } + PageContent { + page_id, + parts: out, + joined, + contents_array, + } +} + +/// Reads and decodes a page's `/Contents` (missing or null = empty page). Direct streams, +/// unresolvable references and non-stream parts are `MALFORMED_CONTENT`; more than +/// `PAGE_PARTS_MAX` parts or too much decoded data `PAGE_TOO_COMPLEX`; part dictionaries with +/// keys other than `/Length /Filter /DecodeParms /DL` `UNSUPPORTED_FILTER`. +pub fn page_content( + doc: &Document, + page_id: ObjectId, + budget: &mut DecodeBudget, +) -> Result { + let Some(Object::Dictionary(page)) = doc.objects.get(&page_id) else { + return Err(TextReason::MalformedContent); + }; + let refs_of = |items: &[Object]| -> Result, TextReason> { + if items.len() > limits::PAGE_PARTS_MAX { + return Err(TextReason::PageTooComplex); + } + items + .iter() + .map(|o| o.as_reference().map_err(|_| TextReason::MalformedContent)) + .collect() + }; + let (ids, contents_array) = match page.get(b"Contents").ok() { + None | Some(Object::Null) => (Vec::new(), None), + Some(Object::Reference(id)) => match doc.objects.get(id) { + Some(Object::Stream(_)) => (vec![*id], None), + Some(Object::Array(items)) => (refs_of(items)?, Some(*id)), + _ => return Err(TextReason::MalformedContent), + }, + Some(Object::Array(items)) => (refs_of(items)?, None), + Some(_) => return Err(TextReason::MalformedContent), + }; + let mut decoded: Vec<(ObjectId, Vec)> = Vec::with_capacity(ids.len()); + let mut page_total = 0usize; + for id in ids { + let Some(Object::Stream(stream)) = doc.objects.get(&id) else { + return Err(TextReason::MalformedContent); + }; + if stream + .dict + .iter() + .any(|(k, _)| !PART_KEYS.contains(&k.as_slice())) + { + return Err(TextReason::UnsupportedFilter); + } + let page_left = limits::PAGE_CONTENT_MAX_DECODED.saturating_sub(page_total); + let cap = limits::STREAM_MAX_DECODED.min(page_left); + let data = decode::decode_stream(stream, cap, budget).map_err(|e| e.page_reason())?; + page_total = page_total.saturating_add(data.len()); + decoded.push((id, data)); + } + let parts: Vec<(ObjectId, &[u8])> = decoded.iter().map(|(id, d)| (*id, d.as_slice())).collect(); + Ok(assemble(page_id, &parts, contents_array)) +} + +/// qpdf's rule when it turns a page's parts into one Form stream (overlay wrapper, P-OV), as +/// probed on qpdf 12.3.2: before each part a `\n` is written when the previous step left the +/// output not ending in `\n` (a part ending in `\n` needs none; an empty part leaves the +/// state "needs a newline" unless a `\n` was just written for it). Pinned by CON-10 and APP-07. +pub fn qpdf_join(parts: &[&[u8]]) -> Vec { + let mut out = Vec::with_capacity(parts.iter().map(|p| p.len() + 1).sum()); + let mut need_newline = false; + for part in parts { + if need_newline { + out.push(b'\n'); + } + let last = match part.last() { + Some(c) => Some(*c), + None if need_newline => Some(b'\n'), + None => None, + }; + out.extend_from_slice(part); + need_newline = last != Some(b'\n'); + } + out +} + +pub struct RefCounts { + counts: HashMap, + truncated: bool, // a value nested deeper than 64 levels was not counted +} + +impl RefCounts { + /// Every Reference in every object (and in the trailer), iterative, depth ≤ 64. + pub fn of(doc: &Document) -> RefCounts { + let mut counts: HashMap = HashMap::new(); + let mut truncated = false; + let mut stack: Vec<(&Object, usize)> = doc.objects.values().map(|o| (o, 0)).collect(); + stack.extend(doc.trailer.iter().map(|(_, v)| (v, 1))); + while let Some((obj, depth)) = stack.pop() { + let container = matches!( + obj, + Object::Array(_) | Object::Dictionary(_) | Object::Stream(_) + ); + if container && depth >= REF_DEPTH_MAX { + truncated = true; + continue; + } + match obj { + Object::Reference(id) => { + let c = counts.entry(*id).or_insert(0); + *c = c.saturating_add(1); + } + Object::Array(items) => stack.extend(items.iter().map(|c| (c, depth + 1))), + Object::Dictionary(d) => stack.extend(d.iter().map(|(_, v)| (v, depth + 1))), + Object::Stream(s) => stack.extend(s.dict.iter().map(|(_, v)| (v, depth + 1))), + _ => {} + } + } + RefCounts { counts, truncated } + } + + pub fn count(&self, id: ObjectId) -> u32 { + self.counts.get(&id).copied().unwrap_or(0) + } +} + +/// Occurrences of each page object id across all `/Kids` arrays of the page tree (own iterative walk from +/// the catalog `/Pages`, depth ≤ PAGE_TREE_DEPTH_MAX, visited set; a cycle or a non-dictionary kid ⇒ Err). +pub struct KidsCounts { + counts: HashMap, +} + +impl KidsCounts { + pub fn of(doc: &Document) -> Result { + let bad = TextReason::MalformedContent; + let root = doc + .catalog() + .ok() + .and_then(|c| c.get(b"Pages").ok()) + .and_then(|o| o.as_reference().ok()) + .ok_or(bad)?; + let mut counts: HashMap = HashMap::new(); + let mut visited: HashSet = HashSet::from([root]); + let mut stack = vec![(root, 0usize)]; + while let Some((node_id, depth)) = stack.pop() { + let Some(Object::Dictionary(node)) = doc.objects.get(&node_id) else { + return Err(bad); + }; + let kids = match node.get(b"Kids").map_err(|_| bad)? { + Object::Array(items) => items, + Object::Reference(id) => match doc.objects.get(id) { + Some(Object::Array(items)) => items, + _ => return Err(bad), + }, + _ => return Err(bad), + }; + for kid in kids { + let kid_id = kid.as_reference().map_err(|_| bad)?; + let c = counts.entry(kid_id).or_insert(0); + *c = c.saturating_add(1); + let Some(Object::Dictionary(kid_dict)) = doc.objects.get(&kid_id) else { + return Err(bad); + }; + if kid_dict.has(b"Kids") || kid_dict.type_is(b"Pages") { + if depth + 1 > limits::PAGE_TREE_DEPTH_MAX || !visited.insert(kid_id) { + return Err(bad); + } + stack.push((kid_id, depth + 1)); + } + } + } + Ok(KidsCounts { counts }) + } + + pub fn count(&self, id: ObjectId) -> u32 { + self.counts.get(&id).copied().unwrap_or(0) + } +} + +/// True only if: the page id occurs exactly once across all `/Kids` arrays **and** once in `get_pages()` +/// (references from StructElem `/Pg`, annotation `/P`, outline/link `/Dest`, named destinations and +/// `/OpenAction` are legitimate and ignored); `/Contents` is a stream reference, a direct array, or an array +/// object whose `RefCounts` count is 1; the part stream's `RefCounts` count is exactly 1 (a content stream can +/// only legitimately be referenced from `/Contents`); the part's dict has no `/F`. +pub fn part_exclusive( + doc: &Document, + refs: &RefCounts, + kids: &KidsCounts, + content: &PageContent, + part: usize, +) -> bool { + let Some(p) = content.parts.get(part) else { + return false; + }; + if refs.truncated || kids.count(content.page_id) != 1 { + return false; + } + if doc + .get_pages() + .values() + .filter(|id| **id == content.page_id) + .count() + != 1 + { + return false; + } + if content + .contents_array + .is_some_and(|arr| refs.count(arr) != 1) + { + return false; + } + let has_f = match doc.objects.get(&p.stream_id) { + Some(Object::Stream(s)) => s.dict.has(b"F"), + _ => true, + }; + refs.count(p.stream_id) == 1 && !has_f +} diff --git a/src-tauri/src/pdf_engine/text_edit/content/joins.rs b/src-tauri/src/pdf_engine/text_edit/content/joins.rs new file mode 100644 index 0000000..79570be --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/content/joins.rs @@ -0,0 +1,310 @@ +//! Content-part boundaries: ONE rule for the walker and the save gate (review T5 H1 and probe +//! R7; review-final MEDIUM-2, LOW-1, LOW-2, LOW-4). The walker refuses a page whose boundary is +//! not neutral `MALFORMED_CONTENT` (a line commented out in every viewer must never be offered as +//! editable); the save gate (`gate::join`) lets qpdf's overlay join stand in for a page's parts +//! only when every boundary is neutral. Sharing the rule keeps a page the editor offers from +//! failing every overlay save. +//! +//! Viewers do not join a page's `/Contents` parts alike: pdf.js reads them as one stream (a +//! comment, a string, a name, a number or an operator runs on into the next part), Poppler ends a +//! token at a part's end but lets comments and strings run on, MuPDF and PDFium put a space +//! between parts, qpdf a newline after a part that does not end with one, and the buffer the +//! model walks has a newline at every boundary. A boundary is neutral when inserting whitespace +//! there changes nothing in any of these readings (ISO 32000-1 §7.8.2: a division between streams +//! "may occur only at the boundaries between lexical tokens"): +//! - it lies between tokens: not inside a string, a hex string, a `<<`/`>>` pair, an inline image +//! (from `BI` through `EI`), or a BX … EX section (pdf.js reads unknown operators there with its +//! partial command names, so any boundary inside one is refused); +//! - a comment open at the end of a part is neutral only when the next non-empty part starts with +//! CR or LF (every reader ends the comment there); +//! - where two regular bytes meet, the token before is a number the next byte cannot continue +//! (not a digit, `.`, `+`, `-`, `e` or `E`: pdf.js stops a number at any other byte), or a known +//! operator that no 1–3-byte prefix of the next token extends into another command pdf.js knows +//! (its operators and its partial names `BM`, `BD`, `fa`, `fal`, `fals`, `nu`, `nul`, …): +//! `ET`|`BT`, `0`|`cm` are neutral, `s`|`h`, `B`|`M…`, `12`|`3`, `/F`|`1` are not; +//! - an inline image must end more than `INLINE_HEURISTIC_WINDOW` bytes before its part's end, +//! whatever proved its end (review-verify MEDIUM-A): pdf.js 4.10 picks an image's `EI` by +//! lexing the 15 bytes after it, so the next part's bytes, shifted by qpdf's `\n`, could make +//! it pick another `EI` before and after the join (a stamp-only save would hide or reveal text +//! in pdf.js). +//! +//! The scan is a streaming state machine over the bytes (O(n), no op cap, nothing allocated per +//! token): only the lexical state at each part's end and the token next to the boundary matter, +//! so a page of any number of ops is judged without lexing it. Inline images are skipped with the +//! lexer's own end proofs (`lexer::InlineEnds`); an image that does not end in its part, or ends +//! within that window of its end, makes the next boundary not neutral. + +use super::PageContent; +use crate::pdf_engine::text_edit::lexer::{is_delimiter, is_whitespace, InlineEnds, Operator}; + +/// The longest content operator (`BDC`, `SCN`, …): how far a known operator could run on. +const OPERATOR_BYTES_MAX: usize = 3; + +/// pdf.js's partial command names (`EvaluatorPreprocessor.opMap`, pdfjs-dist 4.10): its lexer +/// keeps reading a command while the bytes read so far are a known name, these included. +const PDFJS_PARTIALS: [&[u8]; 10] = [ + b"BM", b"BD", b"true", b"fa", b"fal", b"fals", b"false", b"nu", b"nul", b"null", +]; + +/// Checks every boundary between two non-empty parts of `content`; `Err` carries the joined +/// offset (the end of the part) of the first boundary viewers read differently. +pub(crate) fn check_part_joins(content: &PageContent) -> Result<(), usize> { + let parts: Vec<&[u8]> = (0..content.parts.len()) + .map(|i| content.part_bytes(i)) + .collect(); + match first_unsafe_boundary(&parts) { + None => Ok(()), + Some(i) => Err(content + .parts + .get(i) + .map_or(0, |p| p.start.saturating_add(p.len))), + } +} + +/// The index of the first part whose end is a boundary that is not neutral (see the module doc), +/// or `None`. Empty parts are skipped (a comment runs on over them). +pub(crate) fn first_unsafe_boundary(parts: &[&[u8]]) -> Option { + let mut scan = Scan { + state: State::Between, + compat: 0, + }; + let mut parts = parts + .iter() + .enumerate() + .filter(|(_, p)| !p.is_empty()) + .peekable(); + while let Some((i, part)) = parts.next() { + scan.read(part); + if let Some((_, next)) = parts.peek() { + if !scan.boundary(part, next) { + return Some(i); + } + } + } + None +} + +fn regular(c: u8) -> bool { + !is_whitespace(c) && !is_delimiter(c) +} + +/// The lexical state of the stream read so far. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum State { + Between, + /// A regular token from `start` in the current part (a name when it starts with `/`). + Token { + start: usize, + name: bool, + }, + Comment, + /// A literal string, `depth` parentheses deep, a backslash pending when `escape`. + Str { + depth: usize, + escape: bool, + }, + Hex, + /// After a `<` that opens a dictionary or a hex string. + Lt, + /// After a `>` outside a hex string (the first of `>>`). + Gt, + /// Inside an inline image that does not end in its part, or ends too near its end for every + /// reader to agree where (the scan cannot follow it). + Lost, +} + +struct Scan { + state: State, + /// BX … EX nesting. + compat: usize, +} + +impl Scan { + /// Reads `part`, the state carrying over from the previous part as in one stream. + fn read(&mut self, part: &[u8]) { + let mut images: Option> = None; + let mut i = 0usize; + while let Some(&c) = part.get(i) { + i = match self.state { + State::Lost => return, + State::Comment => { + if c == b'\r' || c == b'\n' { + self.state = State::Between; + } + i + 1 + } + State::Str { depth, escape } => { + self.state = in_string(depth, escape, c); + i + 1 + } + State::Hex => { + if c == b'>' { + self.state = State::Between; + } + i + 1 + } + State::Lt if c == b'<' => { + self.state = State::Between; + i + 1 + } + State::Lt => { + self.state = State::Hex; + i + } + State::Gt => { + self.state = State::Between; + if c == b'>' { + i + 1 + } else { + i + } + } + State::Token { .. } if regular(c) => i + 1, + State::Token { start, name } => { + self.state = State::Between; + if !self.token_done(part.get(start..i).unwrap_or_default(), name) { + i + } else { + let ends = images.get_or_insert_with(|| InlineEnds::new(part)); + match ends.end(i).filter(|end| *end > i) { + Some(end) => end, + None => { + self.state = State::Lost; + return; + } + } + } + } + State::Between => { + self.state = token_start(c, i); + i + 1 + } + }; + } + } + + /// A token ended: BX/EX nesting, and whether it is `BI` (an inline image starts after it). + fn token_done(&mut self, token: &[u8], name: bool) -> bool { + if name { + return false; + } + match token { + b"BX" => self.compat = self.compat.saturating_add(1), + b"EX" => self.compat = self.compat.saturating_sub(1), + b"BI" => return true, + _ => {} + } + false + } + + /// Whether the boundary after `part` (just read) and before `next` (the next non-empty part) + /// is neutral. + fn boundary(&mut self, part: &[u8], next: &[u8]) -> bool { + let first = next.first().copied(); + let neutral = match self.state { + State::Between => true, + State::Comment => matches!(first, Some(b'\r' | b'\n')), + State::Token { start, name } => { + let token = part.get(start..).unwrap_or_default(); + let ends = !first.is_some_and(regular) || (!name && ends_before(token, next)); + self.state = State::Between; + let image = self.token_done(token, name); + ends && !image + } + State::Str { .. } | State::Hex | State::Lt | State::Gt | State::Lost => false, + }; + neutral && self.compat == 0 + } +} + +fn token_start(c: u8, at: usize) -> State { + match c { + b'%' => State::Comment, + b'(' => State::Str { + depth: 1, + escape: false, + }, + b'<' => State::Lt, + b'>' => State::Gt, + b'/' => State::Token { + start: at, + name: true, + }, + c if regular(c) => State::Token { + start: at, + name: false, + }, + _ => State::Between, + } +} + +fn in_string(depth: usize, escape: bool, c: u8) -> State { + let depth = match (escape, c) { + (true, _) => depth, + (false, b'\\') => { + return State::Str { + depth, + escape: true, + } + } + (false, b'(') => depth.saturating_add(1), + (false, b')') if depth <= 1 => return State::Between, + (false, b')') => depth - 1, + _ => depth, + }; + State::Str { + depth, + escape: false, + } +} + +/// `[+-]?(\d+\.?\d*|\.\d+)` (the lexer's number syntax). +fn is_number(t: &[u8]) -> bool { + let body = match t.first() { + Some(b'+' | b'-') => t.get(1..).unwrap_or_default(), + _ => t, + }; + let dots = body.iter().filter(|c| **c == b'.').count(); + let digits = body.iter().filter(|c| c.is_ascii_digit()).count(); + digits > 0 && dots <= 1 && digits + dots == body.len() +} + +fn known_command(t: &[u8]) -> bool { + Operator::from_token(t) != Operator::Unknown || PDFJS_PARTIALS.contains(&t) +} + +/// Whether every reader ends the regular `token` at the end of its part when the next part +/// (`next`) starts with a regular byte (see the module doc). +fn ends_before(token: &[u8], next: &[u8]) -> bool { + let Some(&c) = next.first() else { + return true; + }; + if is_number(token) { + return !matches!(c, b'0'..=b'9' | b'.' | b'+' | b'-' | b'e' | b'E'); + } + let operator = Operator::from_token(token); + if operator == Operator::Unknown || matches!(token, b"BI" | b"ID" | b"EI") { + return false; + } + let run_on = next + .iter() + .take(OPERATOR_BYTES_MAX) + .take_while(|c| regular(**c)) + .count(); + let mut word = [0u8; 2 * OPERATOR_BYTES_MAX]; + let Some(head) = word.get_mut(..token.len()) else { + return false; + }; + head.copy_from_slice(token); + (1..=run_on).all(|l| { + let end = token.len() + l; + match (word.get_mut(token.len()..end), next.get(..l)) { + (Some(tail), Some(more)) => tail.copy_from_slice(more), + _ => return false, + } + !word.get(..end).is_some_and(known_command) + }) +} + +#[cfg(test)] +mod tests; diff --git a/src-tauri/src/pdf_engine/text_edit/content/joins/tests.rs b/src-tauri/src/pdf_engine/text_edit/content/joins/tests.rs new file mode 100644 index 0000000..7d3d4df --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/content/joins/tests.rs @@ -0,0 +1,239 @@ +//! The shared boundary rule (review-final LOW-1/2/4, MEDIUM-2) over every boundary shape the +//! reviews probed (`v04/rfinal/zz_final_probe.rs` `zz_join_shapes` and `zz_h1_stamp_regressions`, +//! review-T5 R1/R7): the scan, the save gate (`gate::join::join_is_neutral` on qpdf's join) and +//! the walker give one answer for each. + +use super::first_unsafe_boundary; +use crate::pdf_engine::text_edit::content::qpdf_join; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::gate::join::join_is_neutral; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::producers::{DocBuilder, PageSpec, HELVETICA}; +use std::path::Path; + +const HELLO: &[u8] = b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"; +/// An unfiltered one-pixel inline image (its end proven by its size) ending at `EI`. +const IMAGE: &[u8] = b"q 10 0 0 10 300 300 cm BI /W 1 /H 1 /CS /G /BPC 8 ID \x80 EI"; + +fn page_model_detail(parts: &[&[u8]]) -> (Option, String) { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts(parts, &format!("/Font << /F1 {f} 0 R >>"))); + let pdf = d.build(); + let snap = snapshot_from_bytes(Path::new("joins.pdf"), pdf, None).expect("snapshot"); + let model = build_page_model(&SnapshotContext::new(snap), 0, None).expect("model"); + (model.page_reason, model.page_detail.unwrap_or_default()) +} + +/// (shape, parts, neutral). The pdf.js readings quoted are pdfjs-dist 4.10's (`v04/lastfix-pdfjs`). +fn shapes() -> Vec<(&'static str, Vec>, bool)> { + let p = |parts: &[&[u8]]| parts.iter().map(|x| x.to_vec()).collect::>(); + let mut dense = b"q ".to_vec(); + dense.extend_from_slice(&b"0 0 m 1 1 l S ".repeat(90_000)); + dense.extend_from_slice(b"Q"); + vec![ + ( + "ET|BT", + p(&[b"BT /F1 12 Tf 72 720 Td (Visible) Tj ET", HELLO]), + true, + ), + ( + "Q|number", + p(&[ + b"q BT /F1 12 Tf 72 720 Td (Visible) Tj ET Q", + b"1 0 0 1 0 0 cm", + ]), + true, + ), + ("tab|form feed", p(&[b"BT ET\t", b"\x0cBT ET"]), true), + ( + "string|operator", + p(&[b"BT /F1 12 Tf 72 700 Td (Hello)", b"Tj ET"]), + true, + ), + ( + "array across", + p(&[b"BT /F1 12 Tf 72 700 Td [(Spl) -10", b"(it)] TJ ET"]), + true, + ), + // LOW-1: every reader ends the comment at the next part's first byte. + ("comment|LF", p(&[b"BT ET\n% note", b"\nBT ET"]), true), + ("comment|CR", p(&[b"BT ET\r% note", b"\rBT ET"]), true), + ( + "comment|empty|LF", + p(&[b"BT ET % note", b"", b"\nBT ET"]), + true, + ), + ("comment closed", p(&[b"q\n% note\n", b"Q"]), true), + ("comment|regular", p(&[b"BT ET\n% note", HELLO]), false), + ( + "comment NUL|regular", + p(&[b"BT ET\n% note\x00", HELLO]), + false, + ), + // LOW-4: pdf.js stops a number at a letter; Poppler at the part end. + ("number|cm", p(&[b"q 1 0 0 1 0 0", b"cm BT ET Q"]), true), + ("number|digit", p(&[b"1 0 0 1 0 1", b"0 cm"]), false), + ( + "number|minus", + p(&[b"q 1 0 0 1 0 0", b"-5 0 0 1 0 0 cm Q"]), + false, + ), + ( + "number|dot", + p(&[b"q 1 0 0 1 0 0", b".5 0 0 1 0 0 cm Q"]), + false, + ), + // pdf.js finds an unfiltered image's end only at `EI` + space/CR/LF: it reads `EI`|`Q` + // as image data and loses the text after it (Poppler shows "Hello"). + ( + "EI|Q", + p(&[ + b"q 10 0 0 10 300 300 cm BI /W 1 /H 1 /CS /G /BPC 8 ID \x80 EI", + b"Q BT ET", + ]), + false, + ), + // review-verify MEDIUM-A: pdf.js lexes the 15 bytes after an `EI` to accept it, across + // the boundary and shifted by qpdf's `\n`; an image ending within the look-ahead window + // of its part's end refuses the boundary, whatever proved its end (it was neutral when + // `EI` was followed by a space, CR or LF). + ("EI|LF", p(&[IMAGE, b"\nQ"]), false), + ( + "EI|space: pdf.js swallows `rg` and \"Shown\" after a stamp (ei_end_hide)", + p(&[ + IMAGE, + b" 0.11 0.2 0.3 rg\nQ BT /F1 12 Tf 72 700 Td (Shown) Tj ET\n% EI Q\n\ + BT /F1 12 Tf 72 660 Td (Hello) Tj ET", + ]), + false, + ), + ( + "EI|space: pdf.js shows \"Hidden\" after a stamp (ei_end_reveal)", + p(&[ + IMAGE, + b" Q\n%xxxxxxxxxxx\xe9\nBT /F1 12 Tf 72 700 Td (Hidden) Tj ET\n% EI Q\n\ + BT /F1 12 Tf 72 660 Td (Hello) Tj ET", + ]), + false, + ), + ( + "EI 2 bytes before the end (ei_inside_window_crosses)", + p(&[ + &[IMAGE, b"\nQ"].concat(), + b"%xxxxxxxxxxxx\xe9\nBT /F1 12 Tf 72 700 Td (Hidden) Tj ET\n% EI Q\n\ + BT /F1 12 Tf 72 660 Td (Hello) Tj ET", + ]), + false, + ), + ( + "EI then more than the window before the end", + p(&[&[IMAGE, &b" Q q Q".repeat(12)[..]].concat(), HELLO]), + true, + ), + ( + "inline image split", + p(&[b"q BI /W 2 /H 1 /BPC 8 /CS /G ID \x01", b"\x02 EI Q"]), + false, + ), + ( + "BI|dict", + p(&[b"q BI", b" /W 1 /H 1 /CS /G /BPC 8 ID \x80 EI Q"]), + false, + ), + // LOW-2: pdf.js reads `B`|`M…` as its partial `BM`; any boundary inside BX … EX is refused. + ( + "BX: B|M", + p(&[b"BX 0 0 m 600 0 l h 1 g B", b"Mfoo EX BT ET"]), + false, + ), + ("BX: whitespace", p(&[b"BX\n", b"foo EX\n"]), false), + ("BX closed", p(&[b"BX foo EX\n", b"BT ET"]), true), + ("B|M (partial)", p(&[b"0 0 m 10 10 l B", b"M2 Q"]), false), + ("n|ull (partial)", p(&[b"0 0 m n", b"ull"]), false), + ("s|h", p(&[b"0 0 m 10 10 l s", b"h"]), false), + ("d|0", p(&[b"[] 0 d", b"0 0 m"]), false), + ( + "mid-string", + p(&[b"BT /F1 12 Tf 72 700 Td (Hel", b"lo) Tj ET"]), + false, + ), + ( + "escape|paren", + p(&[b"BT /F1 12 Tf 72 700 Td (Hel\\", b")lo) Tj ET"]), + false, + ), + ( + "mid-hex", + p(&[b"BT /F1 12 Tf 72 700 Td <48", b"65> Tj ET"]), + false, + ), + ( + "<|<", + p(&[b"/Span <", b"< /ActualText (Secret) >> BDC EMC"]), + false, + ), + ( + ">|>", + p(&[b"/Span << /ActualText (Secret) >", b"> BDC EMC"]), + false, + ), + ("tr|ue", p(&[b"/Span << /A tr", b"ue >> BDC EMC"]), false), + ("/F|1", p(&[b"BT /F", b"1 12 Tf ET"]), false), + ("/GS0|gs", p(&[b"/GS0", b"gs"]), false), + // MEDIUM-2: what the strict lexer refuses is judged at its boundaries only. + ("270k ops|text", vec![dense, HELLO.to_vec()], true), + ("text|trailing operand", p(&[HELLO, b"\nq Q 0\n"]), true), + ("PS operator\\n|text", p(&[b"q (x) PS Q\n", HELLO]), true), + ("unknown operator\\n|text", p(&[b"q Q foo\n", HELLO]), true), + ] +} + +#[test] +fn one_boundary_rule_for_the_scan_the_gate_and_the_walker() { + for (shape, parts, neutral) in shapes() { + let parts: Vec<&[u8]> = parts.iter().map(Vec::as_slice).collect(); + assert_eq!( + first_unsafe_boundary(&parts).is_none(), + neutral, + "scan: {shape}" + ); + let joined = qpdf_join(&parts); + let inserts = joined.len() != parts.iter().map(|x| x.len()).sum::(); + assert_eq!( + join_is_neutral(&parts, &joined), + neutral || !inserts, + "gate: {shape}" + ); + let (reason, detail) = page_model_detail(&parts); + let by_joins = detail.starts_with("content parts joined inside"); + if neutral { + assert!( + !by_joins, + "walker refused a neutral boundary: {shape}: {detail}" + ); + } else { + // Not neutral: the walker refuses the page, by this rule or by its lexer first. + assert_eq!( + reason, + Some(TextReason::MalformedContent), + "walker: {shape}: {detail}" + ); + } + } +} + +#[test] +fn the_scan_reports_the_first_unsafe_part_and_skips_empty_ones() { + let parts: [&[u8]; 5] = [b"q", b"", b"Q ", b"(a", b"b) Tj"]; + assert_eq!(first_unsafe_boundary(&parts), Some(3)); + assert_eq!(first_unsafe_boundary(&[]), None); + assert_eq!(first_unsafe_boundary(&[b"(unterminated" as &[u8]]), None); + // A heuristic inline-image end whose look-ahead reaches the part's end is not followed. + let dct = b"q BI /W 1 /H 1 /CS /G /BPC 8 /F /DCT ID \xff\xd8 EI Q" as &[u8]; + assert_eq!(first_unsafe_boundary(&[dct, HELLO]), Some(0)); + let tail = [dct, &b" q Q".repeat(20)[..]].concat(); + assert_eq!(first_unsafe_boundary(&[&tail, HELLO]), None); +} diff --git a/src-tauri/src/pdf_engine/text_edit/context.rs b/src-tauri/src/pdf_engine/text_edit/context.rs new file mode 100644 index 0000000..f84e434 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/context.rs @@ -0,0 +1,289 @@ +//! Snapshot context (SPEC §B.10): one parsed snapshot, its font cache and the lazily computed +//! reference / page-tree counts every page walk of that snapshot shares, plus the resource +//! lookups the walker uses. Unresolvable references are reported as such (`Lookup::Broken`) so +//! the walker can refuse the page (§A.4); a font name that is simply absent stays the run-level +//! `MISSING_FONT`. + +use crate::error::AppError; +use crate::pdf_engine::text_edit::content::{KidsCounts, RefCounts}; +use crate::pdf_engine::text_edit::fonts::{FontCache, FontKey}; +use crate::pdf_engine::text_edit::limits::PAGE_TREE_DEPTH_MAX; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::snapshot::SourceSnapshot; +use crate::pdf_engine::text_edit::structure::OcConfig; +use lopdf::{Dictionary, Document, Object, ObjectId, Stream}; +use std::collections::HashSet; +use std::sync::OnceLock; + +/// References followed for one value (a reference to a reference …). +const REF_HOPS_MAX: usize = 16; + +pub struct SnapshotContext { + pub snap: SourceSnapshot, + pub fonts: FontCache, + refs: OnceLock, + kids: OnceLock>, + oc: OnceLock, +} + +impl SnapshotContext { + pub fn new(snap: SourceSnapshot) -> Self { + SnapshotContext { + snap, + fonts: FontCache::new(), + refs: OnceLock::new(), + kids: OnceLock::new(), + oc: OnceLock::new(), + } + } + + pub fn doc(&self) -> &Document { + &self.snap.doc + } + + pub fn refs(&self) -> &RefCounts { + self.refs.get_or_init(|| RefCounts::of(&self.snap.doc)) + } + + pub fn kids(&self) -> Result<&KidsCounts, TextReason> { + self.kids + .get_or_init(|| KidsCounts::of(&self.snap.doc)) + .as_ref() + .map_err(|r| *r) + } + + /// The catalog's default optional-content configuration, read on first use. + pub fn oc_config(&self) -> &OcConfig { + self.oc.get_or_init(|| OcConfig::of(&self.snap.doc)) + } + + /// The object id of the 0-based page `page_index`; `INVALID_PAGES` when there is none. + pub fn page_id(&self, page_index: u32) -> Result { + usize::try_from(page_index) + .ok() + .and_then(|i| self.snap.pages.get(i).copied()) + .ok_or_else(|| invalid_page(page_index, self.snap.pages.len())) + } +} + +/// `INVALID_PAGES` for a page index the document does not have. +pub fn invalid_page(page_index: u32, page_count: usize) -> AppError { + AppError::new( + "INVALID_PAGES", + "Page not found", + format!( + "This PDF has {page_count} page(s), so page {} does not exist.", + u64::from(page_index) + 1 + ), + ) + .with_suggestion("Reopen the file and try again.") + .with_details(format!("page_index={page_index} pages={page_count}")) +} + +/// Follows references (≤ 16 hops). `None` when one dangles. Returns the id of the last object +/// reached through a reference, if any. +pub fn resolve<'a>(doc: &'a Document, obj: &'a Object) -> Option<(Option, &'a Object)> { + let mut obj = obj; + let mut id = None; + for _ in 0..REF_HOPS_MAX { + match obj { + Object::Reference(r) => { + id = Some(*r); + obj = doc.objects.get(r)?; + } + other => return Some((id, other)), + } + } + None +} + +/// The result of looking a name up in a resource category. +pub enum Lookup { + Found(T), + /// The category or the name does not exist. + Missing, + /// A reference on the way does not resolve, or a value has the wrong type. + Broken(&'static str), +} + +/// A resource dictionary in force and the object that holds it (the page, an ancestor of the +/// page, a Form XObject, or the resources dictionary itself when it is indirect). +#[derive(Clone, Copy)] +pub struct Res<'a> { + pub owner: ObjectId, + pub dict: Option<&'a Dictionary>, + /// The dictionary came from an ancestor of the page (`/Resources` inherited from `/Pages`). + pub inherited: bool, + /// The dictionary is an indirect object (its id), for reference counting. + pub dict_id: Option, +} + +/// A resolved category entry: its key (for fonts), the object id when indirect, the value, and +/// whether the category dictionary is an indirect object (its id). +pub struct Entry<'a> { + pub id: Option, + pub value: &'a Object, + pub category_id: Option, +} + +impl<'a> Res<'a> { + pub fn empty(owner: ObjectId) -> Res<'a> { + Res { + owner, + dict: None, + inherited: false, + dict_id: None, + } + } + + /// The page's `/Resources`, inherited through `/Parent` (nearest wins). A dangling or + /// non-dictionary value is `Err` (page `MALFORMED_CONTENT`). + pub fn of_page(doc: &'a Document, page_id: ObjectId) -> Result, &'static str> { + let mut cur = Some(page_id); + let mut seen = HashSet::new(); + while let Some(id) = cur { + if !seen.insert(id) || seen.len() > PAGE_TREE_DEPTH_MAX + 1 { + return Err("page tree cycle"); + } + let Some(Object::Dictionary(node)) = doc.objects.get(&id) else { + return Err("page tree node"); + }; + if let Ok(value) = node.get(b"Resources") { + return Res::from_value(doc, id, value, id != page_id); + } + cur = match node.get(b"Parent").ok() { + None => None, + Some(Object::Reference(p)) => Some(*p), + Some(_) => return Err("page /Parent"), + }; + } + Ok(Res::empty(page_id)) + } + + /// A Form XObject's `/Resources`; `None` when the form has none (the caller inherits). + pub fn of_form( + doc: &'a Document, + form_id: ObjectId, + form: &'a Stream, + ) -> Option, &'static str>> { + let value = form.dict.get(b"Resources").ok()?; + Some(Res::from_value(doc, form_id, value, false)) + } + + fn from_value( + doc: &'a Document, + owner: ObjectId, + value: &'a Object, + inherited: bool, + ) -> Result, &'static str> { + match resolve(doc, value) { + Some((id, Object::Dictionary(d))) => Ok(Res { + owner: id.unwrap_or(owner), + dict: Some(d), + inherited, + dict_id: id, + }), + Some((_, Object::Null)) => Ok(Res { + owner, + dict: None, + inherited, + dict_id: None, + }), + _ => Err("/Resources"), + } + } + + /// `category[name]`, resolved. + pub fn entry(&self, doc: &'a Document, category: &[u8], name: &[u8]) -> Lookup> { + let Some(dict) = self.dict else { + return Lookup::Missing; + }; + let Ok(cat) = dict.get(category) else { + return Lookup::Missing; + }; + let (category_id, cat) = match resolve(doc, cat) { + Some((id, Object::Dictionary(d))) => (id, d), + Some((_, Object::Null)) => return Lookup::Missing, + _ => return Lookup::Broken("resource category"), + }; + let Ok(raw) = cat.get(name) else { + return Lookup::Missing; + }; + match resolve(doc, raw) { + Some((id, value)) => Lookup::Found(Entry { + id, + value, + category_id, + }), + None => Lookup::Broken("resource reference"), + } + } + + /// The font `name` with its cache key. A missing name is `Missing` (`MISSING_FONT`); a value + /// that is not a font dictionary is `Missing` too; a dangling reference is `Broken`. + pub fn font(&self, doc: &'a Document, name: &[u8]) -> Lookup<(FontKey, &'a Dictionary)> { + match self.entry(doc, b"Font", name) { + Lookup::Found(e) => match e.value { + Object::Dictionary(d) => { + let key = match e.id { + Some(id) => FontKey::Indirect(id), + None => FontKey::direct(e.category_id.unwrap_or(self.owner), name), + }; + Lookup::Found((key, d)) + } + _ => Lookup::Missing, + }, + Lookup::Missing => Lookup::Missing, + Lookup::Broken(w) => Lookup::Broken(w), + } + } + + /// Every `(name, key, dict)` of the `/Font` category that resolves to a font dictionary, in + /// dictionary order (dangling entries are skipped: they are only an error when used). + pub fn all_fonts(&self, doc: &'a Document) -> Vec<(Vec, FontKey, &'a Dictionary)> { + let Some(dict) = self.dict else { + return Vec::new(); + }; + let Some((category_id, Object::Dictionary(cat))) = + dict.get(b"Font").ok().and_then(|c| resolve(doc, c)) + else { + return Vec::new(); + }; + cat.iter() + .filter_map(|(name, raw)| match resolve(doc, raw)? { + (id, Object::Dictionary(d)) => { + let key = match id { + Some(id) => FontKey::Indirect(id), + None => FontKey::direct(category_id.unwrap_or(self.owner), name), + }; + Some((name.clone(), key, d)) + } + _ => None, + }) + .collect() + } + + /// Bytes of every name in the `/Font` category (what `all_fonts` copies), without copying. + pub fn font_name_bytes(&self, doc: &'a Document) -> usize { + self.dict + .and_then(|d| d.get(b"Font").ok()) + .and_then(|c| resolve(doc, c)) + .and_then(|(_, o)| o.as_dict().ok()) + .map_or(0, |cat| cat.iter().map(|(name, _)| name.len()).sum()) + } + + /// Number of entries in the `/Font` category (0 when absent or unreadable). + pub fn font_count(&self, doc: &'a Document) -> usize { + self.dict + .and_then(|d| d.get(b"Font").ok()) + .and_then(|c| resolve(doc, c)) + .and_then(|(_, o)| o.as_dict().ok()) + .map_or(0, Dictionary::len) + } +} + +/// A name, number or dictionary value from a dictionary, resolved. +pub fn get<'a>(doc: &'a Document, dict: &'a Dictionary, key: &[u8]) -> Option<&'a Object> { + let (_, obj) = resolve(doc, dict.get(key).ok()?)?; + (!matches!(obj, Object::Null)).then_some(obj) +} diff --git a/src-tauri/src/pdf_engine/text_edit/decode.rs b/src-tauri/src/pdf_engine/text_edit/decode.rs new file mode 100644 index 0000000..054ffe3 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/decode.rs @@ -0,0 +1,346 @@ +//! Bounded stream decoding (SPEC §B.5): Flate (no predictor), ASCIIHex, ASCII85 and chains of +//! up to four of them. Every decoder is capped while it runs; there is no path that returns the +//! raw (still encoded) bytes of a filtered stream. + +use crate::pdf_engine::text_edit::limits::INFLATE_STEP_BYTES; +use crate::pdf_engine::text_edit::reasons::TextReason; +use flate2::{Decompress, FlushDecompress, Status}; +use lopdf::{Dictionary, Object, Stream}; +use std::borrow::Cow; + +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum DecodeError { + UnsupportedFilter(String), + Corrupt(&'static str), + TooLarge, +} + +impl DecodeError { + /// UnsupportedFilter → UNSUPPORTED_FILTER, Corrupt → MALFORMED_CONTENT, TooLarge → PAGE_TOO_COMPLEX. + pub fn page_reason(&self) -> TextReason { + match self { + DecodeError::UnsupportedFilter(_) => TextReason::UnsupportedFilter, + DecodeError::Corrupt(_) => TextReason::MalformedContent, + DecodeError::TooLarge => TextReason::PageTooComplex, + } + } +} + +impl std::fmt::Display for DecodeError { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + match self { + DecodeError::UnsupportedFilter(name) => write!(f, "unsupported filter: {name}"), + DecodeError::Corrupt(what) => write!(f, "corrupt stream data: {what}"), + DecodeError::TooLarge => write!(f, "decoded data too large"), + } + } +} + +/// A decode budget shared by several streams (e.g. one page walk). +#[derive(Debug, Clone)] +pub struct DecodeBudget { + remaining: usize, +} + +impl DecodeBudget { + pub fn new(total: usize) -> Self { + DecodeBudget { remaining: total } + } + + pub fn remaining(&self) -> usize { + self.remaining + } + + /// Debits `n` bytes; `TooLarge` when the budget cannot cover them (nothing is debited then). + pub fn take(&mut self, n: usize) -> Result<(), DecodeError> { + self.remaining = self.remaining.checked_sub(n).ok_or(DecodeError::TooLarge)?; + Ok(()) + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum Filter { + Flate, + AsciiHex, + Ascii85, +} + +/// Decodes `stream` with at most `cap` output bytes (and at most `budget.remaining()`), then +/// debits the budget with the decoded length. +pub fn decode_stream( + stream: &Stream, + cap: usize, + budget: &mut DecodeBudget, +) -> Result, DecodeError> { + let filters = filters_of(&stream.dict)?; + let cap = cap.min(budget.remaining()); + let mut data: Cow<'_, [u8]> = Cow::Borrowed(stream.content.as_slice()); + for filter in filters { + let decoded = match filter { + Filter::Flate => inflate_capped(&data, cap)?, + Filter::AsciiHex => ascii_hex_decode(&data, cap)?, + Filter::Ascii85 => ascii85_decode(&data, cap)?, + }; + data = Cow::Owned(decoded); + } + if data.len() > cap { + return Err(DecodeError::TooLarge); + } + budget.take(data.len())?; + Ok(data.into_owned()) +} + +fn filters_of(dict: &Dictionary) -> Result, DecodeError> { + for key in [&b"F"[..], b"FFilter", b"FDecodeParms"] { + if dict.has(key) { + return Err(DecodeError::UnsupportedFilter("external file (/F)".into())); + } + } + let names: Vec<&[u8]> = match dict.get(b"Filter").ok() { + None | Some(Object::Null) => Vec::new(), + Some(Object::Name(n)) => vec![n.as_slice()], + Some(Object::Array(items)) => { + if items.len() > 4 { + return Err(DecodeError::UnsupportedFilter("more than 4 filters".into())); + } + let mut out = Vec::with_capacity(items.len()); + for item in items { + match item { + Object::Name(n) => out.push(n.as_slice()), + _ => return Err(DecodeError::UnsupportedFilter("malformed /Filter".into())), + } + } + out + } + Some(_) => return Err(DecodeError::UnsupportedFilter("malformed /Filter".into())), + }; + check_decode_parms(dict.get(b"DecodeParms").ok())?; + names.into_iter().map(filter_from_name).collect() +} + +fn filter_from_name(name: &[u8]) -> Result { + match name { + b"FlateDecode" | b"Fl" => Ok(Filter::Flate), + b"ASCIIHexDecode" | b"AHx" => Ok(Filter::AsciiHex), + b"ASCII85Decode" | b"A85" => Ok(Filter::Ascii85), + other => Err(DecodeError::UnsupportedFilter( + String::from_utf8_lossy(other).into_owned(), + )), + } +} + +/// `/DecodeParms` may only carry parameters that leave the data unchanged (`/Predictor` 1 or absent). +fn check_decode_parms(parms: Option<&Object>) -> Result<(), DecodeError> { + let check_dict = |d: &Dictionary| -> Result<(), DecodeError> { + match d.get(b"Predictor").ok() { + None | Some(Object::Integer(1)) => Ok(()), + Some(_) => Err(DecodeError::UnsupportedFilter("predictor".into())), + } + }; + match parms { + None | Some(Object::Null) => Ok(()), + Some(Object::Dictionary(d)) => check_dict(d), + Some(Object::Array(items)) => { + for item in items { + match item { + Object::Null => {} + Object::Dictionary(d) => check_dict(d)?, + _ => { + return Err(DecodeError::UnsupportedFilter( + "malformed /DecodeParms".into(), + )) + } + } + } + Ok(()) + } + Some(_) => Err(DecodeError::UnsupportedFilter( + "malformed /DecodeParms".into(), + )), + } +} + +/// Splits off and checks the 2-byte zlib header (CM = 8, CINFO ≤ 7, FCHECK, FDICT = 0). +fn zlib_body(data: &[u8]) -> Result<&[u8], DecodeError> { + let (cmf, flg) = match (data.first(), data.get(1)) { + (Some(c), Some(f)) => (*c, *f), + _ => return Err(DecodeError::Corrupt("truncated flate")), + }; + let check = (u16::from(cmf) * 256 + u16::from(flg)) % 31; + if cmf & 0x0F != 8 || cmf >> 4 > 7 || check != 0 || flg & 0x20 != 0 { + return Err(DecodeError::Corrupt("bad zlib header")); + } + Ok(data.get(2..).unwrap_or_default()) +} + +/// Raw inflate of `body` in fixed output steps; every produced chunk goes to `sink` only after +/// the running total was checked against `cap`. Returns (total output, consumed input) at +/// `StreamEnd`; input ending before it is `Corrupt("truncated flate")`. +fn inflate_run( + body: &[u8], + cap: usize, + mut sink: impl FnMut(&[u8]), +) -> Result<(usize, usize), DecodeError> { + let mut d = Decompress::new(false); + let mut step = vec![0u8; INFLATE_STEP_BYTES]; + loop { + let in_before = usize::try_from(d.total_in()).map_err(|_| DecodeError::TooLarge)?; + let out_before = usize::try_from(d.total_out()).map_err(|_| DecodeError::TooLarge)?; + let input = body.get(in_before..).unwrap_or_default(); + let status = d + .decompress(input, &mut step, FlushDecompress::None) + .map_err(|_| DecodeError::Corrupt("flate data error"))?; + let in_after = usize::try_from(d.total_in()).map_err(|_| DecodeError::TooLarge)?; + let out_after = usize::try_from(d.total_out()).map_err(|_| DecodeError::TooLarge)?; + let produced = out_after + .checked_sub(out_before) + .ok_or(DecodeError::Corrupt("flate state"))?; + if out_after > cap { + return Err(DecodeError::TooLarge); + } + if produced > 0 { + sink(step.get(..produced).unwrap_or_default()); + } + match status { + Status::StreamEnd => return Ok((out_after, in_after)), + Status::Ok | Status::BufError => { + if produced == 0 && in_after == in_before { + return Err(DecodeError::Corrupt("truncated flate")); + } + } + } + } +} + +/// Zlib-wrapped Flate with the output capped at `cap` while inflating. Peak allocation is at +/// most `cap` + one 64 KiB step: a first pass only measures, the second fills an exact buffer. +/// The Adler-32 trailer and any trailing bytes are ignored (as pdf.js and qpdf do). +pub fn inflate_capped(data: &[u8], cap: usize) -> Result, DecodeError> { + if data.is_empty() { + return Ok(Vec::new()); + } + let body = zlib_body(data)?; + let (len, _) = inflate_run(body, cap, |_| {})?; + let mut out = Vec::new(); + out.try_reserve_exact(len) + .map_err(|_| DecodeError::TooLarge)?; + inflate_run(body, len, |chunk| out.extend_from_slice(chunk))?; + Ok(out) +} + +/// Number of input bytes (zlib header included, Adler-32 trailer excluded) a capped inflate +/// consumes before `StreamEnd`. Used to prove the end of a Flate inline image. +pub fn inflate_end(data: &[u8], cap: usize) -> Result { + let body = zlib_body(data)?; + let (_, consumed) = inflate_run(body, cap, |_| {})?; + consumed.checked_add(2).ok_or(DecodeError::TooLarge) +} + +fn is_pdf_whitespace(b: u8) -> bool { + matches!(b, 0 | 9 | 10 | 12 | 13 | 32) +} + +fn hex_value(b: u8) -> Option { + match b { + b'0'..=b'9' => Some(b - b'0'), + b'a'..=b'f' => Some(b - b'a' + 10), + b'A'..=b'F' => Some(b - b'A' + 10), + _ => None, + } +} + +/// ASCIIHexDecode: whitespace ignored, `>` ends the data, an odd final digit is padded with 0. +pub fn ascii_hex_decode(data: &[u8], cap: usize) -> Result, DecodeError> { + let mut out = Vec::new(); + out.try_reserve_exact( + (data.len() / 2) + .saturating_add(1) + .min(cap.saturating_add(1)), + ) + .map_err(|_| DecodeError::TooLarge)?; + let mut high: Option = None; + for &b in data { + if b == b'>' { + break; + } + if is_pdf_whitespace(b) { + continue; + } + let v = hex_value(b).ok_or(DecodeError::Corrupt("bad hex digit"))?; + match high.take() { + None => high = Some(v), + Some(h) => { + if out.len() >= cap { + return Err(DecodeError::TooLarge); + } + out.push((h << 4) | v); + } + } + } + if let Some(h) = high { + if out.len() >= cap { + return Err(DecodeError::TooLarge); + } + out.push(h << 4); + } + Ok(out) +} + +/// ASCII85Decode: optional leading `<~`, whitespace ignored, `z` for four zero bytes, `~>` ends. +pub fn ascii85_decode(data: &[u8], cap: usize) -> Result, DecodeError> { + let body = data.strip_prefix(b"<~").unwrap_or(data); + let mut out = Vec::new(); + out.try_reserve_exact((body.len() / 5 * 4 + 4).min(cap.saturating_add(4))) + .map_err(|_| DecodeError::TooLarge)?; + let mut group = [0u8; 5]; + let mut count = 0usize; + let mut iter = body.iter().copied().peekable(); + let push = |out: &mut Vec, bytes: &[u8]| -> Result<(), DecodeError> { + if out.len().saturating_add(bytes.len()) > cap { + return Err(DecodeError::TooLarge); + } + out.extend_from_slice(bytes); + Ok(()) + }; + while let Some(b) = iter.next() { + match b { + _ if is_pdf_whitespace(b) => {} + b'~' => { + if iter.peek() == Some(&b'>') { + break; + } + return Err(DecodeError::Corrupt("bad ASCII85 end")); + } + b'z' if count == 0 => push(&mut out, &[0, 0, 0, 0])?, + b'!'..=b'u' => { + if let Some(slot) = group.get_mut(count) { + *slot = b - b'!'; + } + count += 1; + if count == 5 { + push(&mut out, &a85_group(&group)?)?; + count = 0; + } + } + _ => return Err(DecodeError::Corrupt("bad ASCII85 character")), + } + } + match count { + 0 => {} + 1 => return Err(DecodeError::Corrupt("bad ASCII85 final group")), + n => { + for slot in group.iter_mut().skip(n) { + *slot = 84; + } + let bytes = a85_group(&group)?; + push(&mut out, bytes.get(..n - 1).unwrap_or_default())?; + } + } + Ok(out) +} + +fn a85_group(digits: &[u8; 5]) -> Result<[u8; 4], DecodeError> { + let value = digits.iter().fold(0u64, |acc, d| acc * 85 + u64::from(*d)); + let value = u32::try_from(value).map_err(|_| DecodeError::Corrupt("ASCII85 group overflow"))?; + Ok(value.to_be_bytes()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/dto.rs b/src-tauri/src/pdf_engine/text_edit/dto.rs new file mode 100644 index 0000000..86de527 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/dto.rs @@ -0,0 +1,444 @@ +//! DTOs of the text-edit Tauri commands (SPEC §B.20) and their builders. JSON shapes mirror the +//! TypeScript types in `src/lib/types.ts` (`TextSourceInfo`, `PageText`, `TextRun`, `TextFont`, +//! `TextPreview`, …); `src/lib/editor/__fixtures__/text-edit-dto-contract.json` holds one sample +//! of each, checked byte for byte by `tests_dto` (DTO-01) and key by key by the frontend. +//! +//! Font keys are opaque per response: `f{obj}-{gen}` for an indirect font, `d{hash}` for a +//! direct one. A run's `editable`/`reason` come from the #33 classifier's run capability +//! (`source_content::classify_source_page`), geometry, metrics and style from the same model. + +use crate::pdf_engine::source_content::{ + SourceCapability, SourceFormLine, SourcePageResult, SourceRunCapability, +}; +use crate::pdf_engine::text_edit::fit::word_space; +use crate::pdf_engine::text_edit::fonts::{face_surface, FamilyHint, FontKey, FontModel}; +use crate::pdf_engine::text_edit::gate::TextWarning; +use crate::pdf_engine::text_edit::preview::PreviewResult; +use crate::pdf_engine::text_edit::reasons::{ + EditProblemCode, Face, StyleField, TextReason, TextWarningCode, +}; +use crate::pdf_engine::text_edit::rewrite::{letter_spacing_pt, EditVerdict}; +use crate::pdf_engine::text_edit::runs::{PageModel, SpaceMode, TextRun}; +use serde::Serialize; +use std::collections::{BTreeMap, HashMap}; +use std::sync::Arc; + +pub use super::rewrite::TextEditIn; // Deserialize type, defined in rewrite.rs (B.12) + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct TextSourceDto { + pub fingerprint: String, + pub page_count: u32, + pub warnings: Vec, +} + +#[derive(Debug, Clone, Copy, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct RectDto { + pub x: f64, + pub y: f64, + pub w: f64, + pub h: f64, +} + +#[derive(Debug, Clone, Copy, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct VecDto { + pub x: f64, + pub y: f64, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct PageTextDto { + pub fingerprint: String, + pub page_index: u32, + pub page_reason: Option, + pub runs: Vec, + pub fonts: Vec, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct TextRunDto { + pub id: String, + pub order: u32, + pub line: u32, + pub text: String, + pub rect: RectDto, + pub origin: VecDto, + pub dir: VecDto, + pub ascent: f64, + pub descent: f64, + pub caret_offsets: Vec, + pub editable: bool, + pub reason: Option, + /// `Some` iff editable. + pub metrics: Option, + pub style: Option, + pub substituted: bool, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct RunMetricsDto { + /// Font keys of the typing surface, primary first. + pub surface: Vec, + pub tf_size: f64, + pub effective_size: f64, + pub char_spacing: f64, + pub word_spacing: f64, + pub h_scale: f64, + pub text_to_user: f64, + pub letter_spacing_pt: f64, + pub space_mode: SpaceMode, + pub kern_space: f64, + pub original_width: f64, + pub visible_extent: f64, + pub next_obstacle: Option, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct FaceOptionDto { + pub available: bool, + pub surface: Vec, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct FacesDto { + pub regular: FaceOptionDto, + pub bold: FaceOptionDto, + pub italic: FaceOptionDto, + pub bold_italic: FaceOptionDto, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct RunStyleDto { + pub fill: Option, + pub size_changeable: bool, + pub colour_changeable: bool, + pub face: Face, + pub faces: FacesDto, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct TextFontDto { + pub key: String, + pub display_name: String, + pub family_hint: FamilyHint, + pub embedded: bool, + pub subset: bool, + /// Every typeable character. + pub alphabet: String, + /// `widths[i]`: width of the i-th character of `alphabet` in thousandths of text space + /// (the `/Widths` scale). + pub widths: Vec, + /// The typeable space is the single byte 0x20, so `Tw` applies to it. + pub word_space: bool, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct EditVerdictDto { + pub run_id: String, + pub ok: bool, + pub code: Option, + pub chars: Vec, + /// TEXT_EDIT_REFUSED: the run's reason. + pub reason: Option, + /// FACE_UNAVAILABLE: the requested face. + pub face: Option, + /// STYLE_UNAVAILABLE: which control. + pub field: Option, + pub detail: Option, + pub delta_pt: f64, + pub new_rect: Option, + /// Caret offsets of the new text (planned advances) when ok. + pub caret_offsets: Option>, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct EditProblemDto { + pub code: EditProblemCode, + pub detail: Option, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct TextWarningDto { + pub run_id: String, + pub code: TextWarningCode, + pub detail: Option, +} + +#[derive(Debug, Clone, PartialEq, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct TextPreviewDto { + /// Base64 of the patched one-page PDF. + pub page_pdf: Option, + pub verdicts: Vec, + pub page_problem: Option, + pub warnings: Vec, +} + +// ---- Builders ---------------------------------------------------------------------------- + +pub fn rect_dto(r: [f64; 4]) -> RectDto { + RectDto { + x: r[0], + y: r[1], + w: r[2], + h: r[3], + } +} + +fn vec_dto(v: (f64, f64)) -> VecDto { + VecDto { x: v.0, y: v.1 } +} + +/// The opaque key of a font model in DTOs. +pub fn font_key(m: &FontModel) -> String { + match m.key { + FontKey::Indirect((n, g)) => format!("f{n}-{g}"), + FontKey::Direct { name_hash, .. } => format!("d{name_hash:x}"), + } +} + +/// The fonts a response refers to, by key (first use order is irrelevant: keys are sorted). +#[derive(Default)] +struct FontSet { + by_key: BTreeMap>, +} + +impl FontSet { + fn keys(&mut self, models: &[Arc]) -> Vec { + models + .iter() + .map(|m| { + let key = font_key(m); + self.by_key + .entry(key.clone()) + .or_insert_with(|| Arc::clone(m)); + key + }) + .collect() + } + + fn into_dtos(self) -> Vec { + self.by_key + .into_iter() + .map(|(key, m)| font_dto(key, &m)) + .collect() + } +} + +fn font_dto(key: String, m: &FontModel) -> TextFontDto { + let alphabet = m.alphabet(); + TextFontDto { + key, + display_name: m.display_name.clone(), + family_hint: m.family_hint, + embedded: m.embedded, + subset: m.subset, + alphabet: alphabet.iter().map(|(c, _)| *c).collect(), + widths: alphabet.iter().map(|(_, w)| *w).collect(), + word_space: word_space(m), + } +} + +fn face_option( + model: &PageModel, + run: &TextRun, + own: &[Arc], + face: Face, + fonts: &mut FontSet, +) -> FaceOptionDto { + if face == run.face { + return FaceOptionDto { + available: !own.is_empty(), + surface: fonts.keys(own), + }; + } + let other = match own.first() { + Some(primary) if !run.font_from_extgstate => { + face_surface(&model.walk.page_fonts, primary, face) + } + _ => None, + }; + match other { + Some(s) => { + let models: Vec> = s.fonts.into_iter().map(|(_, m)| m).collect(); + FaceOptionDto { + available: !models.is_empty(), + surface: fonts.keys(&models), + } + } + None => FaceOptionDto { + available: false, + surface: Vec::new(), + }, + } +} + +/// Metrics and style of an editable run (the fonts it can type with are added to `fonts`). +fn editable_parts( + model: &PageModel, + run: &TextRun, + fonts: &mut FontSet, +) -> (RunMetricsDto, RunStyleDto) { + let own: Vec> = model + .surface(run) + .fonts + .into_iter() + .map(|(_, m)| m) + .collect(); + let metrics = RunMetricsDto { + surface: fonts.keys(&own), + tf_size: run.tfs, + effective_size: run.effective_size, + char_spacing: run.tc, + word_spacing: run.tw, + h_scale: run.th, + text_to_user: run.text_to_user_x, + // The planner's own reading, so a value sent back unchanged is a no-op (review-T5 L4). + letter_spacing_pt: letter_spacing_pt(run), + space_mode: run.space_mode, + kern_space: run.kern_space, + original_width: run.original_extent, + visible_extent: run.visible_extent, + next_obstacle: run.next_obstacle, + }; + let mut face = |f: Face| face_option(model, run, &own, f, fonts); + let faces = FacesDto { + regular: face(Face::Regular), + bold: face(Face::Bold), + italic: face(Face::Italic), + bold_italic: face(Face::BoldItalic), + }; + let style = RunStyleDto { + fill: run.fill_hex.clone(), + size_changeable: !run.font_from_extgstate, + colour_changeable: run.tr == 0, + face: run.face, + faces, + }; + (metrics, style) +} + +/// `PageText` of one page: geometry, metrics and style from the model, `editable`/`reason` from +/// the classifier's run capabilities (a run the classifier did not list is not editable). +/// A line of text drawn through a Form XObject: shown, never editable (`NESTED_FORM`, §A.6; +/// review-T5 live B1), after the page's own runs. +fn form_line_dto(l: &SourceFormLine) -> TextRunDto { + TextRunDto { + id: l.id.clone(), + order: l.order, + line: l.line, + text: l.text.clone(), + rect: rect_dto(l.rect), + origin: vec_dto(l.origin), + dir: vec_dto(l.dir), + ascent: l.ascent, + descent: l.descent, + caret_offsets: l.caret_offsets.clone(), + editable: false, + reason: Some(l.reason), + metrics: None, + style: None, + substituted: l.substituted, + } +} + +pub fn page_text(model: &PageModel, classified: &SourcePageResult) -> PageTextDto { + let caps: HashMap<&str, &SourceRunCapability> = classified + .runs + .iter() + .map(|c| (c.run_id.as_str(), c)) + .collect(); + let mut fonts = FontSet::default(); + let runs = model + .runs + .iter() + .map(|run| { + let cap = caps.get(run.id.as_str()); + let editable = classified.page_reason.is_none() + && cap.is_some_and(|c| c.capability == SourceCapability::Supported); + let reason = cap.map_or(run.reason, |c| c.reason); + let (metrics, style) = if editable { + let (m, s) = editable_parts(model, run, &mut fonts); + (Some(m), Some(s)) + } else { + (None, None) + }; + TextRunDto { + id: run.id.clone(), + order: run.order, + line: run.line, + text: run.text.clone(), + rect: rect_dto(run.rect), + origin: vec_dto(run.origin), + dir: vec_dto(run.dir), + ascent: run.ascent, + descent: run.descent, + caret_offsets: run.caret_offsets.clone(), + editable, + reason, + metrics, + style, + substituted: run.substituted, + } + }) + .chain(classified.form_lines.iter().map(form_line_dto)) + .collect(); + PageTextDto { + fingerprint: model.fingerprint.to_string(), + page_index: model.page_index, + page_reason: classified.page_reason, + runs, + fonts: fonts.into_dtos(), + } +} + +pub fn verdict_dto(v: &EditVerdict) -> EditVerdictDto { + let p = v.problem.as_ref(); + let ok = p.is_none(); + EditVerdictDto { + run_id: v.run_id.clone(), + ok, + code: p.map(|p| p.code), + chars: p.map_or_else(Vec::new, |p| p.chars.iter().map(char::to_string).collect()), + reason: p.and_then(|p| p.reason), + face: p.and_then(|p| p.face), + field: p.and_then(|p| p.field), + detail: p.and_then(|p| p.detail.clone()), + delta_pt: v.delta_pt, + new_rect: v.new_rect.map(rect_dto), + caret_offsets: if ok { v.caret_offsets.clone() } else { None }, + } +} + +pub fn warning_dto(w: &TextWarning) -> TextWarningDto { + TextWarningDto { + run_id: w.run_id.clone(), + code: w.code, + detail: w.detail.clone(), + } +} + +pub fn preview_dto(r: &PreviewResult) -> TextPreviewDto { + TextPreviewDto { + page_pdf: r.pdf.as_deref().map(crate::pdf_engine::render::base64), + verdicts: r.verdicts.iter().map(verdict_dto).collect(), + page_problem: r.page_problem.as_ref().map(|p| EditProblemDto { + code: p.code, + detail: p.detail.clone(), + }), + warnings: r.warnings.iter().map(warning_dto).collect(), + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/encode.rs b/src-tauri/src/pdf_engine/text_edit/encode.rs new file mode 100644 index 0000000..936443b --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/encode.rs @@ -0,0 +1,111 @@ +//! Bytes of a replacement (SPEC §B.12): numbers written with at most four decimals and read back +//! (every later computation uses the value a viewer will parse), uppercase hex strings of codes, +//! and the replacement grammar (§A.4) that every replacement is checked against twice — before any +//! IO (`rewrite`) and on the re-walk (`verify`). + +use crate::pdf_engine::text_edit::fonts::Code; +use crate::pdf_engine::text_edit::lexer::{ + is_delimiter, is_whitespace, lex_content, LexLimits, Operator, +}; +use crate::pdf_engine::text_edit::limits::{NUMBER_ABS_MAX, NUMBER_DECIMALS}; +use crate::pdf_engine::text_edit::reasons::{EditProblem, EditProblemCode}; + +/// The operators a replacement may contain (§A.4): `Tf Tc Tw T* TJ` plus the fill operators a +/// verbatim restore or the new colour can need (`g rg k cs sc scn`). +pub const REPLACEMENT_OPERATORS: [Operator; 11] = [ + Operator::Tf, + Operator::Tc, + Operator::Tw, + Operator::TStar, + Operator::TJ, + Operator::g, + Operator::rg, + Operator::k, + Operator::cs, + Operator::sc, + Operator::scn, +]; + +/// A number as written in a replacement: ≤ `NUMBER_DECIMALS` decimals, trailing zeros and dot +/// stripped, `-0` → `0`, never an exponent, `|v| ≤ NUMBER_ABS_MAX`. +pub fn fmt_num(v: f64) -> Result { + if !v.is_finite() || v.abs() > NUMBER_ABS_MAX { + return Err(EditProblem::new( + EditProblemCode::EditVerifyFailed, + Some(format!("number out of range: {v}")), + )); + } + let mut s = format!("{v:.prec$}", prec = NUMBER_DECIMALS); + if s.contains('.') { + while s.ends_with('0') { + s.pop(); + } + if s.ends_with('.') { + s.pop(); + } + } + if s == "-0" || s.is_empty() { + s = "0".to_string(); + } + Ok(s) +} + +/// `fmt_num` and the value a reader parses back from it. +pub fn num(v: f64) -> Result<(String, f64), EditProblem> { + let s = fmt_num(v)?; + let parsed = s.parse::().map_err(|_| { + EditProblem::new( + EditProblemCode::EditVerifyFailed, + Some(format!("number does not read back: {s}")), + ) + })?; + Ok((s, parsed)) +} + +/// `` of `codes` (uppercase, each code big-endian in `code.len` bytes); `<>` when empty. +pub fn hex_codes(codes: &[Code]) -> String { + let mut out = String::with_capacity(2 + codes.len() * 4); + out.push('<'); + for code in codes { + let len = usize::from(code.len).min(4); + let bytes = code.value.to_be_bytes(); + for b in bytes.get(4 - len..).unwrap_or_default() { + out.push_str(&format!("{b:02X}")); + } + } + out.push('>'); + out +} + +/// Whether a replacement written right after `prev` must start with a space so that its first +/// token cannot fuse with the previous one (`prev` is the byte before the splice, `None` at the +/// start of a part, where the previous part's last byte is unknown to a reader that concatenates +/// parts). +pub fn needs_leading_space(prev: Option, replacement: &[u8]) -> bool { + let regular = |c: u8| !is_whitespace(c) && !is_delimiter(c); + match (prev, replacement.first()) { + (_, None) => false, + (None, Some(c)) => regular(*c), + (Some(p), Some(c)) => regular(p) && regular(*c), + } +} + +/// The replacement grammar (§A.4): the bytes lex completely, every operator is one of +/// `REPLACEMENT_OPERATORS`, and there are exactly `expected_tj` `TJ` operators. +pub fn check_replacement_grammar(bytes: &[u8], expected_tj: usize) -> Result<(), &'static str> { + let ops = + lex_content(bytes, &LexLimits::page(), None).map_err(|_| "replacement does not lex")?; + let mut tj = 0usize; + for op in &ops { + if !REPLACEMENT_OPERATORS.contains(&op.operator) || op.in_compat { + return Err(op.operator.as_str()); + } + if op.operator == Operator::TJ { + tj = tj.saturating_add(1); + } + } + if tj != expected_tj { + return Err("TJ count"); + } + Ok(()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/engines.rs b/src-tauri/src/pdf_engine/text_edit/engines.rs new file mode 100644 index 0000000..d45563a --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/engines.rs @@ -0,0 +1,658 @@ +//! Engines and subprocesses (SPEC §B.8): qpdf ≥ 11, pdftoppm and pdftotext resolution, one +//! subprocess runner (argv only, capped output, timeout, cancel), `qpdf --check` +//! classification with a fingerprint memo and a background form, and qpdf's page map. + +use crate::error::AppError; +use crate::models::JobHandle; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::reasons::{self, EditProblem, EditProblemCode, ProblemCtx}; +use crate::pdf_engine::text_edit::snapshot::Fingerprint; +use crate::pdf_engine::{qpdf, render}; +use lopdf::ObjectId; +use std::collections::{HashMap, VecDeque}; +use std::ffi::OsString; +use std::io::Read; +use std::path::{Path, PathBuf}; +use std::process::{Command, Stdio}; +use std::sync::atomic::{AtomicBool, Ordering}; +use std::sync::{Arc, Condvar, Mutex, OnceLock}; +use std::time::{Duration, Instant}; + +#[derive(Debug, Clone)] +pub struct Engines { + pub qpdf: PathBuf, + pub pdftoppm: PathBuf, + pub pdftotext: PathBuf, +} + +const POPPLER_DIRS: [&str; 4] = [ + "/opt/homebrew/bin", + "/usr/local/bin", + "/opt/local/bin", + "/usr/bin", +]; + +fn exe_name(tool: &str) -> String { + if cfg!(windows) { + format!("{tool}.exe") + } else { + tool.to_string() + } +} + +/// Poppler tool without a Tauri handle: next to the executable, the same absolute dirs as +/// `render.rs`, then the bare name (PATH). +fn poppler_standalone(tool: &str) -> PathBuf { + let exe = exe_name(tool); + if let Some(dir) = std::env::current_exe() + .ok() + .and_then(|p| p.parent().map(Path::to_path_buf)) + { + let candidate = dir.join("binaries").join(&exe); + if candidate.is_file() { + return candidate; + } + } + if !cfg!(windows) { + for dir in POPPLER_DIRS { + let candidate = Path::new(dir).join(&exe); + if candidate.is_file() { + return candidate; + } + } + } + PathBuf::from(exe) +} + +/// True when `exe` can be spawned: an existing file, or a bare name found on PATH. +fn tool_exists(exe: &Path) -> bool { + if exe.components().count() > 1 || exe.is_absolute() { + return exe.is_file(); + } + let Some(path) = std::env::var_os("PATH") else { + return false; + }; + std::env::split_paths(&path).any(|dir| dir.join(exe).is_file()) +} + +fn tool_label(exe: &Path) -> String { + exe.file_name() + .map(|n| n.to_string_lossy().into_owned()) + .unwrap_or_else(|| exe.to_string_lossy().into_owned()) +} + +impl Engines { + pub fn resolve(app: &tauri::AppHandle) -> Result { + Engines { + qpdf: qpdf::resolve_qpdf(app), + pdftoppm: render::resolve_pdftoppm(app), + pdftotext: render::resolve_pdftotext(app), + } + .checked() + } + + pub fn for_export(qpdf: &Path, app: Option<&tauri::AppHandle>) -> Result { + let (pdftoppm, pdftotext) = match app { + Some(app) => ( + render::resolve_pdftoppm(app), + render::resolve_pdftotext(app), + ), + None => ( + poppler_standalone("pdftoppm"), + poppler_standalone("pdftotext"), + ), + }; + Engines { + qpdf: qpdf.to_path_buf(), + pdftoppm, + pdftotext, + } + .checked() + } + + /// qpdf must run and be ≥ 11 (`ENGINE_MISSING`); both Poppler tools must exist (`VERIFIER_MISSING`). + fn checked(self) -> Result { + qpdf_version_ok(&self.qpdf)?; + for tool in [&self.pdftoppm, &self.pdftotext] { + if !tool_exists(tool) { + return Err(reasons::verifier_missing(&tool_label(tool))); + } + } + Ok(self) + } +} + +pub struct RunOpts<'a> { + pub handle: Option<&'a Arc>, + pub cancel: Option<&'a AtomicBool>, + pub timeout: Duration, // default SUBPROCESS_TIMEOUT_SECS + pub stdout_cap: usize, // stdout read is capped; excess → EDIT_VERIFY_FAILED "tool output too large" +} + +impl Default for RunOpts<'_> { + fn default() -> Self { + RunOpts { + handle: None, + cancel: None, + timeout: Duration::from_secs(limits::SUBPROCESS_TIMEOUT_SECS), + stdout_cap: limits::PDFTOTEXT_OUTPUT_MAX, + } + } +} + +impl<'a> RunOpts<'a> { + fn with_stdout_cap(&self, stdout_cap: usize) -> RunOpts<'a> { + RunOpts { + handle: self.handle, + cancel: self.cancel, + timeout: self.timeout, + stdout_cap, + } + } + + fn cancelled(&self) -> bool { + self.cancel.is_some_and(|c| c.load(Ordering::SeqCst)) + || self.handle.is_some_and(|h| h.is_cancelled()) + } +} + +#[derive(Debug, Clone)] +pub struct ToolOutput { + pub code: i32, + pub stdout: Vec, + pub stderr: String, +} + +fn verify_failed(detail: &str) -> AppError { + let p = EditProblem::new(EditProblemCode::EditVerifyFailed, Some(detail.to_string())); + let ctx = ProblemCtx { + page_number: None, + file_name: None, + face: None, + reason: None, + }; + EditProblemCode::EditVerifyFailed.to_app_error(&p, &ctx) +} + +/// Reads a pipe to EOF keeping at most `cap` bytes; returns (kept, overflowed). +fn read_pipe(pipe: Option, cap: usize) -> (Vec, bool) { + let Some(mut pipe) = pipe else { + return (Vec::new(), false); + }; + let mut kept = Vec::new(); + let mut overflow = false; + let mut buf = vec![0u8; 64 << 10]; + loop { + match pipe.read(&mut buf) { + Ok(0) | Err(_) => break, + Ok(n) => { + let chunk = buf.get(..n).unwrap_or_default(); + let room = cap.saturating_sub(kept.len()); + if chunk.len() > room { + overflow = true; + } + kept.extend_from_slice(chunk.get(..room.min(chunk.len())).unwrap_or_default()); + } + } + } + (kept, overflow) +} + +/// Runs `exe` with `args` (never a shell string), reading stdout (capped) and stderr on helper +/// threads while polling for exit, timeout (`ENGINE_FAILED` "took too long") and cancel +/// (`CANCELLED`). A missing executable is `ENGINE_MISSING` (qpdf) or `VERIFIER_MISSING` (poppler). +/// A non-zero exit code is returned, not an error. +pub fn run_tool( + exe: &Path, + args: &[OsString], + poppler: bool, + opts: &RunOpts<'_>, +) -> Result { + if opts.cancelled() { + return Err(AppError::cancelled()); + } + let mut cmd = Command::new(exe); + cmd.args(args) + .stdin(Stdio::null()) + .stdout(Stdio::piped()) + .stderr(Stdio::piped()); + if poppler { + render::configure_poppler_command(&mut cmd, exe); + } + #[cfg(windows)] + { + use std::os::windows::process::CommandExt; + cmd.creation_flags(0x0800_0000); + } + let mut child = match cmd.spawn() { + Ok(child) => child, + Err(e) if e.kind() == std::io::ErrorKind::NotFound => { + return Err(if poppler { + reasons::verifier_missing(&tool_label(exe)) + } else { + AppError::engine_missing() + }); + } + Err(e) => return Err(AppError::engine_failed(format!("{}: {e}", tool_label(exe)))), + }; + let stdout = child.stdout.take(); + let stderr = child.stderr.take(); + let cap = opts.stdout_cap; + let out_reader = std::thread::Builder::new().spawn(move || read_pipe(stdout, cap)); + let err_reader = + std::thread::Builder::new().spawn(move || read_pipe(stderr, limits::TOOL_STDERR_MAX)); + let (Ok(out_reader), Ok(err_reader)) = (out_reader, err_reader) else { + let _ = child.kill(); + let _ = child.wait(); + return Err(AppError::engine_failed( + "could not start the output readers", + )); + }; + let started = Instant::now(); + let status = loop { + match child.try_wait() { + Ok(Some(status)) => break status, + Ok(None) => {} + Err(e) => { + let _ = child.kill(); + let _ = child.wait(); + return Err(AppError::engine_failed(format!("{}: {e}", tool_label(exe)))); + } + } + // On timeout or cancel the readers are detached: a grandchild may still hold the pipes. + if opts.cancelled() { + let _ = child.kill(); + let _ = child.wait(); + return Err(AppError::cancelled()); + } + if started.elapsed() >= opts.timeout { + let _ = child.kill(); + let _ = child.wait(); + return Err(AppError::engine_failed(format!( + "{} took too long and was stopped", + tool_label(exe) + ))); + } + std::thread::sleep(Duration::from_millis(limits::TOOL_POLL_MS)); + }; + let (stdout, overflow) = out_reader.join().unwrap_or((Vec::new(), true)); + let (stderr, _) = err_reader.join().unwrap_or((Vec::new(), false)); + if overflow { + return Err(verify_failed(&format!( + "{}: tool output too large", + tool_label(exe) + ))); + } + Ok(ToolOutput { + code: status.code().unwrap_or(-1), + stdout, + stderr: String::from_utf8_lossy(&stderr).into_owned(), + }) +} + +// ---- qpdf --check ------------------------------------------------------------------------ + +#[derive(Debug, Clone, PartialEq)] +pub enum SourceCheck { + Clean, + Benign(Vec), + Problems(Vec), +} + +/// One entry of the benign-warning allow-list (each has its own ENG test; nothing else is benign). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum BenignRule { + /// The lower-cased line contains this text. + Contains(&'static str), + /// `reported number of objects (N) is not one plus the highest object number (M)` — + /// a wrong trailer `/Size`, which lopdf also corrects. + SizeMismatch, +} + +pub const BENIGN_CHECK_RULES: &[BenignRule] = &[ + BenignRule::Contains("linearization"), + BenignRule::Contains("hint table"), + BenignRule::SizeMismatch, +]; + +impl BenignRule { + pub fn matches(self, line: &str) -> bool { + let lower = line.to_ascii_lowercase(); + match self { + BenignRule::Contains(text) => lower.contains(text), + BenignRule::SizeMismatch => size_mismatch(&lower), + } + } +} + +fn size_mismatch(lower: &str) -> bool { + let digits_then = |s: &str, tail: &str| -> Option { + let n = s.bytes().take_while(u8::is_ascii_digit).count(); + (n > 0 && s.get(n..)?.starts_with(tail)).then_some(n + tail.len()) + }; + let head = "reported number of objects ("; + let middle = ") is not one plus the highest object number ("; + let Some(p) = lower.find(head) else { + return false; + }; + let rest = lower.get(p + head.len()..).unwrap_or_default(); + let Some(n) = digits_then(rest, middle) else { + return false; + }; + digits_then(rest.get(n..).unwrap_or_default(), ")").is_some() +} + +/// Warning lines of a `qpdf --check` run: every non-empty stderr line except qpdf's summary, +/// plus stdout lines starting with `WARNING`; each distinct line once, in first-seen order. +fn warning_lines(stdout: &str, stderr: &str) -> Vec { + let err = stderr + .lines() + .map(str::trim) + .filter(|l| !l.is_empty() && !l.starts_with("qpdf: operation succeeded with warnings")); + let out = stdout + .lines() + .map(str::trim) + .filter(|l| l.to_ascii_uppercase().starts_with("WARNING")); + let mut seen = std::collections::HashSet::new(); + err.chain(out) + .filter(|l| seen.insert(*l)) + .map(str::to_string) + .collect() +} + +/// Exit 0 → `Clean`; exit 3 → `Benign` when every warning line matches the allow-list, else +/// `Problems`; exit 2 or anything else → `Problems`. +pub fn classify_check(code: i32, stdout: &str, stderr: &str) -> SourceCheck { + let lines = warning_lines(stdout, stderr); + match code { + 0 => SourceCheck::Clean, + 3 if lines.is_empty() => { + SourceCheck::Problems(vec!["qpdf reported warnings without details".to_string()]) + } + 3 if lines + .iter() + .all(|l| BENIGN_CHECK_RULES.iter().any(|r| r.matches(l))) => + { + SourceCheck::Benign(lines) + } + 3 => SourceCheck::Problems(lines), + _ if lines.is_empty() => { + SourceCheck::Problems(vec![format!("qpdf --check exited with code {code}")]) + } + _ => SourceCheck::Problems(lines), + } +} + +pub fn qpdf_check( + engines: &Engines, + pdf: &Path, + opts: &RunOpts<'_>, +) -> Result { + #[cfg(test)] + CHECK_RUNS.with(|c| c.set(c.get() + 1)); + let args = [OsString::from("--check"), pdf.as_os_str().to_os_string()]; + let out = run_tool( + &engines.qpdf, + &args, + false, + &opts.with_stdout_cap(limits::QPDF_JSON_MAX_BYTES), + )?; + // qpdf could not open the file (it vanished, e.g. "Clear temp files" during a background + // check): that says nothing about the PDF, so it is an engine failure, never a verdict + // (review T5 M1). + if out.code == 2 && out.stderr.lines().any(|l| l.starts_with(QPDF_OPEN_FAILED)) { + return Err(AppError::engine_failed(format!( + "qpdf could not open the file to check: {}", + out.stderr.trim() + ))); + } + Ok(classify_check( + out.code, + &String::from_utf8_lossy(&out.stdout), + &out.stderr, + )) +} + +/// How qpdf's stderr starts when it cannot open its input file. +const QPDF_OPEN_FAILED: &str = "qpdf: open "; + +static CHECK_MEMO: Mutex> = Mutex::new(VecDeque::new()); + +/// Memoised by fingerprint (process-wide LRU of CHECK_MEMO_MAX entries): returns the stored result when the +/// same fingerprint was checked before, else runs `qpdf_check` on `pdf` (a copy written from the very bytes +/// that were fingerprinted) and stores it. Safe to share between open and Save: the key is computed from the +/// bytes qpdf reads. Errors (timeout, cancel, missing qpdf, an input qpdf could not open) are never +/// memoised, and neither is a result whose input no longer has the fingerprinted length after the run +/// (it was deleted or replaced while qpdf read it, review T5 M1). +pub fn qpdf_check_memo( + engines: &Engines, + fingerprint: Fingerprint, + pdf: &Path, + opts: &RunOpts<'_>, +) -> Result { + { + let mut memo = CHECK_MEMO.lock().unwrap_or_else(|e| e.into_inner()); + if let Some(pos) = memo.iter().position(|(f, _)| *f == fingerprint) { + if let Some(entry) = memo.remove(pos) { + let result = entry.1.clone(); + memo.push_front(entry); + return Ok(result); + } + } + } + let result = qpdf_check(engines, pdf, opts)?; + let still_there = std::fs::metadata(pdf).is_ok_and(|m| m.len() == fingerprint.len); + if !still_there { + return Ok(result); + } + let mut memo = CHECK_MEMO.lock().unwrap_or_else(|e| e.into_inner()); + memo.retain(|(f, _)| *f != fingerprint); + memo.push_front((fingerprint, result.clone())); + memo.truncate(limits::CHECK_MEMO_MAX); + Ok(result) +} + +type CheckSlot = Arc<(Mutex>>, Condvar)>; + +/// Background form used at open: spawns a thread running `qpdf_check_memo`; `wait` blocks (honouring cancel +/// and SUBPROCESS_TIMEOUT_SECS) until the result exists. +#[derive(Clone)] +pub struct PendingCheck { + slot: CheckSlot, +} + +fn store(slot: &CheckSlot, result: Result) { + let (lock, cv) = &**slot; + *lock.lock().unwrap_or_else(|e| e.into_inner()) = Some(result); + cv.notify_all(); +} + +impl PendingCheck { + pub fn spawn(engines: Engines, fingerprint: Fingerprint, pdf: PathBuf) -> PendingCheck { + let slot: CheckSlot = Arc::new((Mutex::new(None), Condvar::new())); + let worker_slot = Arc::clone(&slot); + let spawned = std::thread::Builder::new() + .name("offpdf-qpdf-check".into()) + .spawn(move || { + let result = qpdf_check_memo(&engines, fingerprint, &pdf, &RunOpts::default()); + store(&worker_slot, result); + }); + if let Err(e) = spawned { + store( + &slot, + Err(AppError::engine_failed(format!( + "could not start the qpdf check: {e}" + ))), + ); + } + PendingCheck { slot } + } + + pub fn wait(&self, cancel: Option<&AtomicBool>) -> Result { + let deadline = + Instant::now() + Duration::from_secs(limits::SUBPROCESS_TIMEOUT_SECS.saturating_add(5)); + let (lock, cv) = &*self.slot; + let mut guard = lock.lock().unwrap_or_else(|e| e.into_inner()); + loop { + if let Some(result) = guard.as_ref() { + return result.clone(); + } + if cancel.is_some_and(|c| c.load(Ordering::SeqCst)) { + return Err(AppError::cancelled()); + } + if Instant::now() >= deadline { + return Err(AppError::engine_failed( + "qpdf --check took too long and was stopped", + )); + } + guard = match cv.wait_timeout(guard, Duration::from_millis(limits::TOOL_POLL_MS)) { + Ok((g, _)) => g, + Err(e) => e.into_inner().0, + }; + } + } + + pub fn peek(&self) -> Option> { + self.slot + .0 + .lock() + .unwrap_or_else(|e| e.into_inner()) + .clone() + } +} + +// ---- qpdf page map and version ----------------------------------------------------------- + +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct QpdfPage { + pub object: ObjectId, + pub contents: Vec, +} + +/// `"N G R"`, strictly (no signs, no leading zeros, single spaces). +fn parse_ref(s: &str) -> Option { + let mut it = s.split(' '); + let (n, g, r) = (it.next()?, it.next()?, it.next()?); + let canonical = |t: &str| { + !t.is_empty() && t.bytes().all(|c| c.is_ascii_digit()) && (t == "0" || !t.starts_with('0')) + }; + if it.next().is_some() || r != "R" || !canonical(n) || !canonical(g) { + return None; + } + Some((n.parse().ok()?, g.parse().ok()?)) +} + +/// Parses `qpdf --json=2 --json-key=pages` output (qpdf 11.9 and 12.x shapes). +pub fn parse_qpdf_pages(json: &[u8]) -> Result, String> { + let v: serde_json::Value = + serde_json::from_slice(json).map_err(|e| format!("qpdf JSON: {e}"))?; + let pages = v + .get("pages") + .and_then(serde_json::Value::as_array) + .ok_or("qpdf JSON has no pages array")?; + pages + .iter() + .enumerate() + .map(|(i, p)| { + let n = i + 1; + let object = p + .get("object") + .and_then(serde_json::Value::as_str) + .and_then(parse_ref) + .ok_or_else(|| format!("qpdf JSON page {n}: bad object"))?; + let contents = p + .get("contents") + .and_then(serde_json::Value::as_array) + .ok_or_else(|| format!("qpdf JSON page {n}: no contents"))? + .iter() + .map(|c| c.as_str().and_then(parse_ref)) + .collect::>>() + .ok_or_else(|| format!("qpdf JSON page {n}: bad contents"))?; + Ok(QpdfPage { object, contents }) + }) + .collect() +} + +/// `qpdf --json=2 --json-key=pages ` (exit 0 or 3). Failures are `PDF_NEEDS_REPAIR`; the +/// gate maps them to `EDIT_VERIFY_FAILED` for files the pipeline wrote. +pub fn qpdf_page_map( + engines: &Engines, + pdf: &Path, + opts: &RunOpts<'_>, +) -> Result, AppError> { + let args = [ + OsString::from("--json=2"), + OsString::from("--json-key=pages"), + pdf.as_os_str().to_os_string(), + ]; + let out = run_tool( + &engines.qpdf, + &args, + false, + &opts.with_stdout_cap(limits::QPDF_JSON_MAX_BYTES), + )?; + if out.code != 0 && out.code != 3 { + let first = out.stderr.lines().next().unwrap_or_default(); + return Err(reasons::pdf_needs_repair(&[format!( + "qpdf --json exited with code {}: {first}", + out.code + )])); + } + parse_qpdf_pages(&out.stdout).map_err(|e| reasons::pdf_needs_repair(&[e])) +} + +/// Major version and version text from `qpdf --version` output ("qpdf version 12.3.2"). +pub fn parse_qpdf_version(stdout: &str) -> Option<(u32, String)> { + let line = stdout.lines().next()?.trim(); + let version = line.strip_prefix("qpdf version ")?.trim(); + let major = version.split('.').next()?.parse().ok()?; + Some((major, version.to_string())) +} + +static VERSION_CACHE: OnceLock>>> = OnceLock::new(); + +/// `qpdf --version`, major ≥ 11 (cached per path; spawn failures are not cached). +pub fn qpdf_version_ok(qpdf: &Path) -> Result<(), AppError> { + let cache = VERSION_CACHE.get_or_init(|| Mutex::new(HashMap::new())); + if let Some(hit) = cache.lock().unwrap_or_else(|e| e.into_inner()).get(qpdf) { + return hit.clone(); + } + let opts = RunOpts { + timeout: Duration::from_secs(30), + stdout_cap: 64 << 10, + ..RunOpts::default() + }; + let out = run_tool(qpdf, &[OsString::from("--version")], false, &opts)?; + let result = match parse_qpdf_version(&String::from_utf8_lossy(&out.stdout)) { + Some((major, _)) if major >= 11 => Ok(()), + Some((_, version)) => Err(reasons::qpdf_too_old(&version)), + None => Err(AppError::engine_failed( + "OffPDF could not read the qpdf version.", + )), + }; + cache + .lock() + .unwrap_or_else(|e| e.into_inner()) + .insert(qpdf.to_path_buf(), result.clone()); + result +} + +/// Engines found without a running app (tests and benches: the dev Mac's tools). +#[cfg(test)] +impl Engines { + pub fn standalone() -> Option { + Engines { + qpdf: qpdf::resolve_qpdf_standalone(), + pdftoppm: poppler_standalone("pdftoppm"), + pdftotext: poppler_standalone("pdftotext"), + } + .checked() + .ok() + } +} + +#[cfg(test)] +thread_local! { + /// Test seam (ENG-08): qpdf --check runs started on this thread. + pub(crate) static CHECK_RUNS: std::cell::Cell = const { std::cell::Cell::new(0) }; +} diff --git a/src-tauri/src/pdf_engine/text_edit/export.rs b/src-tauri/src/pdf_engine/text_edit/export.rs new file mode 100644 index 0000000..b43e554 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/export.rs @@ -0,0 +1,756 @@ +//! Save integration (SPEC §B.18): text changes are written into per-source edited copies +//! **before** assembly, each proven by Phase A; the Edit PDF pipeline then runs unchanged on the +//! substituted paths, and Phase B proves the final staged file before #34's +//! `validate_staged_pdf` and the atomic rename (`edit_overlay::export_edit_pdf_with_check_exe`). +//! +//! Every error of this path passes through `reasons::save_failure`, so its suggestion ends with +//! "The original file was not changed." Save never uses the editor's cache: each edited source is +//! read again (one bounded read), its fingerprint must equal the one the edits were made on, and +//! qpdf only reads a copy written from those bytes. The only shared state is the memoised +//! `qpdf --check` result for identical bytes. + +mod map; + +use crate::error::AppError; +use crate::models::{JobUpdate, PageGroup}; +use crate::pdf_engine::edit_overlay::{EditDocumentIn, EditObjectIn}; +use crate::pdf_engine::text_edit::apply::{apply_update, updates_for_plan, write_update_json}; +use crate::pdf_engine::text_edit::cache::file_name; +use crate::pdf_engine::text_edit::content::{page_content, qpdf_join, PageContent}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::engines::{ + qpdf_check_memo, qpdf_page_map, Engines, RunOpts, SourceCheck, +}; +use crate::pdf_engine::text_edit::gate::join::join_is_neutral; +use crate::pdf_engine::text_edit::gate::{ + cancel_flag, verify_edited_copy, verify_final_output, BeforeModel, EditedPageInput, ModelPrint, + PageProof, PhaseAInput, PhaseAReport, PopplerRef, TextWarning, +}; +use crate::pdf_engine::text_edit::graph::{graph_digest, GraphDigest}; +use crate::pdf_engine::text_edit::limits::{ + EDITS_PER_SAVE_MAX, EDIT_TEXT_CHARS_MAX, PAGE_DECODE_BUDGET, VERIFY_CAP_MARGIN_BYTES, +}; +use crate::pdf_engine::text_edit::reasons::{ + self, save_failure, EditProblem, EditProblemCode, ProblemCtx, TextWarningCode, +}; +use crate::pdf_engine::text_edit::rewrite::{plan_page, PagePlan, SourceTextStyleIn, TextEditIn}; +use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel}; +use crate::pdf_engine::text_edit::snapshot::{check_page_map, read_snapshot}; +use crate::pdf_engine::validate_output::{content_digest, ContentDigest}; +use lopdf::{Document, ObjectId}; +use map::{dest_pages, group_specs, key_of, SourceEdits}; +use std::collections::{HashMap, HashSet}; +use std::path::{Path, PathBuf}; +use std::sync::Arc; + +/// One `sourceText` object of the export document. +#[derive(Debug, Clone)] +pub struct SourceTextSpec { + /// Destination page (0-based, in the saved file). + pub page_index: u32, + /// Page of the source file (0-based) the edit was made on. + pub source_page_index: u32, + pub run_id: String, + pub source_fingerprint: String, + pub original_text: String, + pub text: String, + pub style: SourceTextStyleIn, +} + +pub struct PreparedTextEdits { + /// `groups` with every edited source replaced by its proven edited copy. + pub content_groups: Vec, + /// Phase A proof of each edited destination page. + pub proofs: Vec<(u32, PageProof)>, + /// User sentences for the job message. + pub warnings: Vec, + /// `read_verification_snapshot` cap of the final file: 2 × Σ distinct input sizes + margin. + pub verify_cap: u64, +} + +/// Progress steps shown on the job (`JobUpdate` step text). +pub const STEP_CHECKING: &str = "Checking text changes"; +pub const STEP_APPLYING: &str = "Applying text changes"; +pub const STEP_VERIFYING: &str = "Verifying text changes"; + +const WARNING_TEXT_MAX_CHARS: usize = 40; + +/// The least `kept_models_budget` (a few ordinary page models are always kept). +const KEPT_MODELS_MIN: usize = 16 << 20; + +fn too_many_edits(n: usize) -> AppError { + AppError::new( + "TOO_MANY_TEXT_EDITS", + "Too many text changes", + format!("This save has more than {EDITS_PER_SAVE_MAX} changed lines."), + ) + .with_suggestion("Save in smaller batches.") + .with_details(format!("text changes: {n}")) +} + +fn redacted_page(dest: u32) -> AppError { + let n = u64::from(dest) + 1; + AppError::new( + "TEXT_EDIT_ON_REDACTED_PAGE", + "Redaction and text change on the same page", + format!("Page {n} has both a redaction and a text change. Redaction turns the page into an image, so the text change would be lost."), + ) + .with_suggestion("Remove the redaction or the text change on that page.") + .with_details(format!("page={n}")) +} + +fn problem_error(code: EditProblemCode, dest: u32, detail: &str) -> AppError { + let p = EditProblem::new(code, Some(detail.to_string())); + let ctx = ProblemCtx { + page_number: Some(dest.saturating_add(1)), + file_name: None, + face: None, + reason: None, + }; + code.to_app_error(&p, &ctx) +} + +/// The `sourceText` objects of an Edit PDF export document. +pub fn specs_of(doc: &EditDocumentIn) -> Vec { + doc.objects + .iter() + .filter_map(|o| match o { + EditObjectIn::SourceText { + page_index, + source_page_index, + run_id, + source_fingerprint, + original_text, + text, + style, + .. + } => Some(SourceTextSpec { + page_index: *page_index, + source_page_index: *source_page_index, + run_id: run_id.clone(), + source_fingerprint: source_fingerprint.clone(), + original_text: original_text.clone(), + text: text.clone(), + style: style.clone(), + }), + _ => None, + }) + .collect() +} + +/// The engines of a Save with text changes, resolved before any work: qpdf ≥ 11 +/// (`ENGINE_MISSING`) and both Poppler tools (`VERIFIER_MISSING`). +pub fn save_engines(qpdf: &Path, app: Option<&tauri::AppHandle>) -> Result { + let engines = Engines::for_export(qpdf, app).map_err(save_failure)?; + #[cfg(test)] + let engines = seams::engines(engines); + Ok(engines) +} + +/// The Edit PDF export's entry point: `None` when `doc` has no text change, else the engines and +/// the proven edited copies (progress steps go to the job when `app` is given). +pub fn prepare_save( + doc: &EditDocumentIn, + groups: &[PageGroup], + work: &Path, + qpdf: &Path, + app: Option<&tauri::AppHandle>, + job_id: &str, + opts: &RunOpts<'_>, +) -> Result, AppError> { + let specs = specs_of(doc); + if specs.is_empty() { + return Ok(None); + } + let engines = save_engines(qpdf, app)?; + let mut progress = |step: &'static str| { + if let Some(app) = app { + use tauri::Emitter; + let _ = app.emit("job:update", JobUpdate::new(job_id, "running", step)); + } + }; + let prepared = prepare_text_sources(groups, &specs, work, &engines, opts, &mut progress)?; + Ok(Some((engines, prepared))) +} + +/// Digest of a page's content parts as qpdf joins them into its overlay Form (§B.18, the #34 +/// `alt_content_digest`), decoded with the bounded text-edit decoder; `None` when that equals the +/// plain concatenation, when the join would change what the parts mean (`join_is_neutral`, +/// review-T5 H1: a part ending inside a comment or a string) or when the decoder refuses the +/// page — then only the plain digest is expected, as before. +pub fn qpdf_joined_digest(doc: &Document, page_id: ObjectId) -> Option { + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + let page = page_content(doc, page_id, &mut budget).ok()?; + let parts: Vec<&[u8]> = (0..page.parts.len()).map(|i| page.part_bytes(i)).collect(); + let joined = qpdf_join(&parts); + let digest = content_digest(&joined); + (digest != page.concat_digest() && join_is_neutral(&parts, &joined)).then_some(digest) +} + +/// Request-level checks before any file is read (`validate_doc`): ≤ `EDITS_PER_SAVE_MAX` +/// (`TOO_MANY_TEXT_EDITS`), ≤ 1,000 characters (`TEXT_TOO_LONG`), one change per +/// (page, line) (`EDIT_CONFLICT`), no page with both a redaction and a text change +/// (`TEXT_EDIT_ON_REDACTED_PAGE`). +pub fn validate_specs(specs: &[SourceTextSpec], redact_pages: &[u32]) -> Result<(), AppError> { + check_specs(specs, redact_pages).map_err(save_failure) +} + +fn check_specs(specs: &[SourceTextSpec], redact_pages: &[u32]) -> Result<(), AppError> { + if specs.len() > EDITS_PER_SAVE_MAX { + return Err(too_many_edits(specs.len())); + } + let redacted: HashSet = redact_pages.iter().copied().collect(); + let mut seen: HashSet<(u32, &str)> = HashSet::new(); + for s in specs { + if s.text.chars().count() > EDIT_TEXT_CHARS_MAX { + return Err(problem_error( + EditProblemCode::TextTooLong, + s.page_index, + &format!("{} characters", s.text.chars().count()), + )); + } + if !seen.insert((s.page_index, s.run_id.as_str())) { + return Err(problem_error( + EditProblemCode::EditConflict, + s.page_index, + &format!("run {} is changed twice", s.run_id), + )); + } + if redacted.contains(&s.page_index) { + return Err(redacted_page(s.page_index)); + } + } + Ok(()) +} + +/// Writes and proves the edited copies (§B.18 steps 1–6). A no-op (empty `specs`, or only +/// changes equal to the original) returns `groups` unchanged. +pub fn prepare_text_sources( + groups: &[PageGroup], + specs: &[SourceTextSpec], + work: &Path, + engines: &Engines, + opts: &RunOpts<'_>, + progress: &mut dyn FnMut(&'static str), +) -> Result { + if specs.is_empty() { + return Ok(PreparedTextEdits { + content_groups: groups.to_vec(), + proofs: Vec::new(), + warnings: Vec::new(), + verify_cap: 0, + }); + } + prepare(groups, specs, work, engines, opts, progress).map_err(save_failure) +} + +/// Phase B on the final staged file, then #34's expected digest of each edited page must equal +/// the proof's independently planned digest (`snapshot_digests` = the #34 snapshot, by dest page). +pub fn verify_final( + tmp: &Path, + prepared: &PreparedTextEdits, + snapshot_digests: &[ContentDigest], + engines: &Engines, + opts: &RunOpts<'_>, +) -> Result<(), AppError> { + let expectations: Vec<(u32, &PageProof)> = + prepared.proofs.iter().map(|(d, p)| (*d, p)).collect(); + verify_final_output(tmp, prepared.verify_cap, engines, &expectations, opts) + .and_then(|()| { + for (dest, proof) in &prepared.proofs { + let planned = usize::try_from(*dest) + .ok() + .and_then(|i| snapshot_digests.get(i)); + if planned != Some(&proof.expected_page_digest) { + return Err(reasons::source_edit_gate_failed(&format!( + "phase=B check=34 page={} the #34 snapshot digest is not the planned page", + u64::from(*dest) + 1 + ))); + } + } + Ok(()) + }) + .map_err(save_failure) +} + +fn prepare( + groups: &[PageGroup], + specs: &[SourceTextSpec], + work: &Path, + engines: &Engines, + opts: &RunOpts<'_>, + progress: &mut dyn FnMut(&'static str), +) -> Result { + progress(STEP_CHECKING); + let dest = dest_pages(groups, engines, opts)?; + let sources = group_specs(groups, &dest, specs)?; + let mut substitutes: HashMap = HashMap::new(); + let mut proofs = Vec::new(); + let mut warnings = Vec::new(); + for (i, src) in sources.iter().enumerate() { + let out = edit_source(i, src, work, engines, opts, progress)?; + if let Some(edited) = out.edited { + substitutes.insert(src.key.clone(), edited); + } + proofs.extend(out.proofs); + warnings.extend(out.warnings); + } + let content_groups: Vec = groups + .iter() + .map(|g| match substitutes.get(&key_of(&g.path)) { + Some(edited) => PageGroup { + path: edited.to_string_lossy().into_owned(), + pages: g.pages.clone(), + }, + None => g.clone(), + }) + .collect(); + Ok(PreparedTextEdits { + verify_cap: verify_cap(groups, &content_groups), + content_groups, + proofs, + warnings, + }) +} + +/// 2 × Σ over distinct inputs of max(original, edited copy) + `VERIFY_CAP_MARGIN_BYTES`. +fn verify_cap(groups: &[PageGroup], content_groups: &[PageGroup]) -> u64 { + let size = |p: &str| std::fs::metadata(p).map_or(0, |m| m.len()); + let mut seen = HashSet::new(); + let total = groups + .iter() + .zip(content_groups) + .filter(|(g, _)| seen.insert(g.path.as_str())) + .map(|(g, c)| size(&g.path).max(size(&c.path))) + .fold(0u64, u64::saturating_add); + total + .saturating_mul(2) + .saturating_add(VERIFY_CAP_MARGIN_BYTES) +} + +struct EditedSource { + /// `None` when every change of this source was a no-op (nothing is written). + edited: Option, + proofs: Vec<(u32, PageProof)>, + warnings: Vec, +} + +/// One planned page of an edited source. +struct PlannedPage { + dest: u32, + /// The page's content in the source (the update's stream ids). + content: Arc, + /// Its model while the kept models stay within `kept_models_budget`; otherwise Phase A + /// builds it again from the source context when it reaches the page (review-T4 M-1). + model: Option, + /// The model's print, which a model built again in Phase A must have (review-final LOW-5). + print: ModelPrint, + plan: PagePlan, + /// The planner's warnings (`NEXT_TEXT_OVERLAP`). + warnings: Vec, +} + +/// Bytes of page models one source's Save may keep from planning to Phase A: twice the file +/// (about what keeping its context and building each model again costs instead), at least +/// `KEPT_MODELS_MIN`. Past it, the source context is kept and no model: each is built again when +/// Phase A reaches its page, one at a time. No upper clamp: on a file over 80 MiB a 160 MiB cap +/// kept the context where the models cost less (review-final MEDIUM-1: 6.1 × file). +pub(super) fn kept_models_budget(file_len: u64) -> usize { + usize::try_from(file_len.saturating_mul(2)) + .unwrap_or(usize::MAX) + .max(KEPT_MODELS_MIN) +} + +/// An edited source read again for the Save: its context, the copy qpdf reads, and the +/// benign `qpdf --check` warnings it already had. +struct CheckedSource { + ctx: SnapshotContext, + copy: PathBuf, + benign: Vec, + page_count: u32, + /// `read_verification_snapshot` cap of its edited copy: 2 × source size + margin. + staged_cap: u64, +} + +/// §B.18 step 3 for one source: fresh snapshot, fingerprint, copy, page map, check, plan, +/// graph digest (then the snapshot is dropped), qpdf update, Phase A. +fn edit_source( + i: usize, + src: &SourceEdits<'_>, + work: &Path, + engines: &Engines, + opts: &RunOpts<'_>, + progress: &mut dyn FnMut(&'static str), +) -> Result { + progress(STEP_CHECKING); + let name = file_name(Path::new(&src.path)); + let checked = check_source(i, src, &name, work, engines, opts)?; + let budget = kept_models_budget(checked.ctx.snap.fingerprint.len); + let planned = plan_pages(&checked.ctx, src, &name, opts, budget)?; + if planned.is_empty() { + let _ = std::fs::remove_file(&checked.copy); + return Ok(EditedSource { + edited: None, + proofs: Vec::new(), + warnings: Vec::new(), + }); + } + progress(STEP_APPLYING); + let CheckedSource { + ctx, + copy, + benign, + page_count, + staged_cap, + } = checked; + let rebuild = planned.iter().any(|p| p.model.is_none()); + let written = write_edited( + i, ctx, rebuild, &planned, ©, &benign, work, engines, opts, + )?; + let (edited, before, ctx) = written; + progress(STEP_VERIFYING); + let a_work = work.join(format!("text-{i}-check")); + std::fs::create_dir_all(&a_work) + .map_err(|e| AppError::io("OffPDF could not create a work folder.", e))?; + let pages = planned + .iter() + .map(|p| phase_a_page(p, ctx.as_ref(), ©)) + .collect::, AppError>>()?; + let input = PhaseAInput { + before: &before, + before_page_count: page_count, + staged: &edited, + staged_cap, + pages, + source_benign: &benign, + }; + let report = verify_edited_copy(&input, engines, &a_work, opts)?; + drop(input); + drop(ctx); + let _ = std::fs::remove_dir_all(&a_work); + let _ = std::fs::remove_file(©); + outcome(src, &planned, report, edited) +} + +/// Phase A's input for one planned page: its kept model, or the source context to build it from. +fn phase_a_page<'a>( + p: &'a PlannedPage, + ctx: Option<&'a SnapshotContext>, + copy: &Path, +) -> Result, AppError> { + let page_index = p.plan.page_index; + let model = match (&p.model, ctx) { + (Some(m), _) => BeforeModel::Kept(m), + (None, Some(ctx)) => BeforeModel::Build { + ctx, + page_index, + print: &p.print, + }, + (None, None) => { + return Err(reasons::source_edit_gate_failed( + "phase=A the source page model is gone", + )) + } + }; + Ok(EditedPageInput { + model, + plan: &p.plan, + input_page_index: page_index, + input_render: PopplerRef { + pdf: copy.to_path_buf(), + page_1: page_index.saturating_add(1), + }, + }) +} + +/// Reads the source again (never the editor's cache) and checks it like the editor did at open: +/// same fingerprint as the edits (`STALE`), lopdf/qpdf page-map agreement, `qpdf --check` +/// (memoised for these bytes; problems → `PDF_NEEDS_REPAIR`). +fn check_source( + i: usize, + src: &SourceEdits<'_>, + name: &str, + work: &Path, + engines: &Engines, + opts: &RunOpts<'_>, +) -> Result { + let mut snap = read_snapshot(Path::new(&src.path))?; + let fp = snap.fingerprint.to_string(); + let specs = src.pages.values().flat_map(|(_, specs)| specs); + if specs.clone().any(|s| s.source_fingerprint != fp) { + return Err(reasons::stale(name)); + } + let copy = work.join(format!("text-{i}-source.pdf")); + std::fs::write(©, snap.bytes.as_slice()) + .map_err(|e| AppError::io("OffPDF could not write a work file.", e))?; + check_page_map(&snap, &qpdf_page_map(engines, ©, opts)?)?; + let benign = match qpdf_check_memo(engines, snap.fingerprint, ©, opts)? { + SourceCheck::Clean => Vec::new(), + SourceCheck::Benign(lines) => lines, + SourceCheck::Problems(lines) => return Err(reasons::pdf_needs_repair(&lines)), + }; + // qpdf reads the copy from here on; the plan and Phase A read the parsed document. + snap.release_bytes(); + Ok(CheckedSource { + page_count: u32::try_from(snap.pages.len()).unwrap_or(u32::MAX), + staged_cap: snap + .fingerprint + .len + .saturating_mul(2) + .saturating_add(VERIFY_CAP_MARGIN_BYTES), + ctx: SnapshotContext::new(snap), + copy, + benign, + }) +} + +/// The "before" graph digest of the source with the planned parts, then — the snapshot and its +/// document dropped (F14) unless `keep_ctx` (a page's model must be built again in Phase A) — +/// the qpdf update into `text--edited.pdf`. +#[allow(clippy::too_many_arguments)] +fn write_edited( + i: usize, + ctx: SnapshotContext, + keep_ctx: bool, + planned: &[PlannedPage], + copy: &Path, + benign: &[String], + work: &Path, + engines: &Engines, + opts: &RunOpts<'_>, +) -> Result<(PathBuf, GraphDigest, Option), AppError> { + let mut replaced = HashMap::new(); + let mut updates = Vec::new(); + for p in planned { + for u in updates_for_plan(&p.content, &p.plan)? { + replaced.insert(u.object_id, u.decoded.clone()); + updates.push(u); + } + } + let before = graph_digest(ctx.doc(), &replaced, cancel_flag(opts)).map_err(|m| { + problem_error( + EditProblemCode::EditVerifyFailed, + planned.first().map_or(0, |p| p.dest), + &format!("phase=A check=A2 source path={} what={}", m.path, m.what), + ) + })?; + let max_id = ctx.doc().max_id; + drop(replaced); + let ctx = keep_ctx.then_some(ctx); + let update = work.join(format!("text-{i}-update.json")); + write_update_json(&updates, max_id, &update)?; + drop(updates); + let edited = work.join(format!("text-{i}-edited.pdf")); + apply_update(engines, copy, &update, &edited, benign, opts)?; + let _ = std::fs::remove_file(&update); + #[cfg(test)] + seams::tamper(&edited); + Ok((edited, before, ctx)) +} + +/// Proofs keyed by destination page and the warnings as job sentences. Every planned page has +/// exactly one proof, and every proof and warning is bound to a planned page explicitly +/// (`SOURCE_EDIT_GATE_FAILED` otherwise, review-T5 L2). +fn outcome( + src: &SourceEdits<'_>, + planned: &[PlannedPage], + report: PhaseAReport, + edited: PathBuf, +) -> Result { + let dest_of = |source_page: u32| { + planned + .iter() + .find(|p| p.plan.page_index == source_page) + .map(|p| p.dest) + .ok_or_else(|| { + reasons::source_edit_gate_failed(&format!( + "phase=A source page {} was not planned", + u64::from(source_page) + 1 + )) + }) + }; + let texts: HashMap<&str, &str> = src + .pages + .values() + .flat_map(|(_, specs)| specs) + .map(|s| (s.run_id.as_str(), s.text.as_str())) + .collect(); + let mut warnings = Vec::new(); + for w in planned + .iter() + .flat_map(|p| &p.warnings) + .chain(&report.warnings) + { + warnings.extend(warning_sentence(w, dest_of(w.page_index)?, &texts)); + } + let proofs = report + .proofs + .into_iter() + .map(|proof| Ok((dest_of(proof.source_page_index)?, proof))) + .collect::, AppError>>()?; + let distinct: HashSet = proofs.iter().map(|(d, _)| *d).collect(); + if proofs.len() != planned.len() || distinct.len() != planned.len() { + return Err(reasons::source_edit_gate_failed(&format!( + "phase=A {} proofs for {} planned pages", + proofs.len(), + planned.len() + ))); + } + Ok(EditedSource { + edited: Some(edited), + proofs, + warnings, + }) +} + +/// Plans every edited page of the source; the first failed verdict is the save error (with the +/// destination page number). Pages whose changes are all no-ops are left out. Models are kept +/// when their bytes (`model_bytes`) all fit within `budget`; otherwise none is. +fn plan_pages( + ctx: &SnapshotContext, + src: &SourceEdits<'_>, + name: &str, + opts: &RunOpts<'_>, + budget: usize, +) -> Result, AppError> { + let mut planned = Vec::new(); + let mut kept = 0usize; + for (source_page, (dest, specs)) in &src.pages { + let model = build_page_model(ctx, *source_page, cancel_flag(opts))?; + if let Some(e) = page_refusal(&model, dest.saturating_add(1)) { + return Err(e); + } + let edits: Vec = specs + .iter() + .map(|s| TextEditIn { + run_id: s.run_id.clone(), + original_text: s.original_text.clone(), + text: s.text.clone(), + style: s.style.clone(), + }) + .collect(); + let outcome = plan_page(ctx, &model, &edits)?; + if let Some(p) = outcome.verdicts.iter().find_map(|v| v.problem.as_ref()) { + let pctx = ProblemCtx { + page_number: Some(dest.saturating_add(1)), + file_name: Some(name), + face: p.face, + reason: p.reason, + }; + return Err(p.code.to_app_error(p, &pctx)); + } + if let Some(plan) = outcome.plan { + let warnings = outcome + .verdicts + .iter() + .flat_map(|v| { + v.warnings.iter().map(|code| TextWarning { + page_index: *source_page, + run_id: v.run_id.clone(), + code: *code, + detail: None, + }) + }) + .collect(); + let bytes = model.walk.model_bytes; + let keep = kept.saturating_add(bytes) <= budget; + if keep { + kept = kept.saturating_add(bytes); + } + planned.push(PlannedPage { + dest: *dest, + print: ModelPrint::of(&model), + content: Arc::clone(&model.content), + model: keep.then_some(model), + plan, + warnings, + }); + } + } + // Past the budget the source context stays for Phase A, and every page's model is built + // again from it: models kept as well would only add to it (review-final MEDIUM-1). + if planned.iter().all(|p| p.model.is_some()) { + return Ok(planned); + } + Ok(planned + .into_iter() + .map(|p| PlannedPage { model: None, ..p }) + .collect()) +} + +/// The error of a page nothing can be changed on (its page-level reason code, with the model's +/// technical detail), or `None`. +pub(crate) fn page_refusal(model: &PageModel, page_number: u32) -> Option { + let e = model.page_reason?.to_app_error(Some(page_number)); + Some(match (&e.details, &model.page_detail) { + (Some(base), Some(detail)) => { + let details = format!("{base}; {detail}"); + e.with_details(details) + } + _ => e, + }) +} + +fn short(text: &str) -> String { + let mut out: String = text.chars().take(WARNING_TEXT_MAX_CHARS).collect(); + if text.chars().count() > WARNING_TEXT_MAX_CHARS { + out.push('\u{2026}'); + } + out +} + +/// A Phase A warning as a sentence for the job message. +fn warning_sentence(w: &TextWarning, dest: u32, texts: &HashMap<&str, &str>) -> Option { + let n = u64::from(dest) + 1; + let line = short(texts.get(w.run_id.as_str()).copied().unwrap_or_default()); + match w.code { + TextWarningCode::NextTextOverlap => Some(format!( + "Page {n}: the changed line \u{201c}{line}\u{201d} runs into the text that follows it." + )), + TextWarningCode::EditNotVisible => Some(format!( + "Page {n}: the change to \u{201c}{line}\u{201d} doesn't change how the page looks. Something may be drawn over this line." + )), + TextWarningCode::PreviewUnavailable => None, + } +} + +/// Test seams (GATE-25 tamper hook, E2E-19 missing Poppler); absent from release builds. +#[cfg(test)] +pub(crate) mod seams { + use crate::pdf_engine::text_edit::engines::Engines; + use std::cell::{Cell, RefCell}; + use std::path::{Path, PathBuf}; + + thread_local! { + static TAMPER: Cell> = const { Cell::new(None) }; + static POPPLER: RefCell> = const { RefCell::new(None) }; + } + + /// Runs on each edited copy right after `apply_update` (this thread only). + pub(crate) fn set_tamper_hook(hook: Option) { + TAMPER.with(|t| t.set(hook)); + } + + pub(super) fn tamper(edited: &Path) { + if let Some(hook) = TAMPER.with(Cell::get) { + hook(edited); + } + } + + /// Replaces both Poppler tools of the save's engines (this thread only). + pub(crate) fn set_poppler_override(path: Option) { + POPPLER.with(|p| *p.borrow_mut() = path); + } + + pub(crate) fn engines(mut engines: Engines) -> Engines { + if let Some(p) = POPPLER.with(|p| p.borrow().clone()) { + engines.pdftoppm = p.clone(); + engines.pdftotext = p; + } + engines + } +} + +#[cfg(test)] +mod tests; diff --git a/src-tauri/src/pdf_engine/text_edit/export/map.rs b/src-tauri/src/pdf_engine/text_edit/export/map.rs new file mode 100644 index 0000000..99a7f83 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/export/map.rs @@ -0,0 +1,146 @@ +//! Save integration, step 1–2 (SPEC §B.18): which source page every destination page is, and +//! the edits grouped by edited source (`STALE` when the editor's page is now another source page, +//! `TEXT_EDIT_DUPLICATE_PAGE` when an edited source page is listed more than once). + +use crate::error::AppError; +use crate::models::PageGroup; +use crate::pdf_engine::edit_overlay::expand_page_spec; +use crate::pdf_engine::text_edit::cache::file_name; +use crate::pdf_engine::text_edit::context::invalid_page; +use crate::pdf_engine::text_edit::engines::{run_tool, Engines, RunOpts}; +use crate::pdf_engine::text_edit::export::SourceTextSpec; +use crate::pdf_engine::text_edit::reasons; +use crate::pdf_engine::text_edit::snapshot::read_snapshot; +use std::collections::{BTreeMap, HashMap}; +use std::ffi::OsString; +use std::path::{Path, PathBuf}; + +fn duplicate_page(source_page: u32, name: &str, times: usize) -> AppError { + let p = u64::from(source_page) + 1; + AppError::new( + "TEXT_EDIT_DUPLICATE_PAGE", + "This page appears twice", + format!("Page {p} of \u{201c}{name}\u{201d} is in the list more than once and has a text change."), + ) + .with_suggestion("Remove the extra copy of the page, then save again.") + .with_details(format!("source page {p} appears {times} times")) +} + +/// One destination page: its group, the canonical source path and the source page (0-based). +pub(super) struct DestPage { + group: usize, + key: PathBuf, + source_page: u32, +} + +pub(super) fn key_of(path: &str) -> PathBuf { + std::fs::canonicalize(path).unwrap_or_else(|_| PathBuf::from(path)) +} + +/// `qpdf --show-npages` (no policy refusals: an unedited signed or XFA file may be appended). +fn page_count(engines: &Engines, path: &str, opts: &RunOpts<'_>) -> Result { + let args = [OsString::from("--show-npages"), OsString::from(path)]; + let out = run_tool(&engines.qpdf, &args, false, opts)?; + let text = String::from_utf8_lossy(&out.stdout); + if let (0 | 3, Ok(n)) = (out.code, text.trim().parse::()) { + return Ok(n); + } + // qpdf could not count the pages: our own read names the reason when it has one + // (ENCRYPTED, FILE_TOO_LARGE, INVALID_PDF …). + read_snapshot(Path::new(path))?; + Err(AppError::invalid_pdf(path).with_details(format!( + "qpdf --show-npages exited with code {}: {}", + out.code, + out.stderr.trim() + ))) +} + +pub(super) fn dest_pages( + groups: &[PageGroup], + engines: &Engines, + opts: &RunOpts<'_>, +) -> Result, AppError> { + let mut counts: HashMap<&str, u32> = HashMap::new(); + let mut dest = Vec::new(); + for (gi, g) in groups.iter().enumerate() { + let n = match counts.get(g.path.as_str()) { + Some(n) => *n, + None => { + let n = page_count(engines, &g.path, opts)?; + counts.insert(g.path.as_str(), n); + n + } + }; + let key = key_of(&g.path); + for p in expand_page_spec(&g.pages, n)? { + dest.push(DestPage { + group: gi, + key: key.clone(), + source_page: p.saturating_sub(1), + }); + } + } + Ok(dest) +} + +/// An edited source: the group path, and per edited source page its dest page and specs. +pub(super) struct SourceEdits<'a> { + pub path: String, + pub key: PathBuf, + /// Per edited source page: its destination page and its edits. + pub pages: BTreeMap)>, +} + +/// Maps every spec to its source (`STALE` when the dest page is now another source page, +/// `TEXT_EDIT_DUPLICATE_PAGE` when an edited source page is listed more than once). +pub(super) fn group_specs<'a>( + groups: &[PageGroup], + dest: &[DestPage], + specs: &'a [SourceTextSpec], +) -> Result>, AppError> { + let mut listed: HashMap<(&Path, u32), usize> = HashMap::new(); + for d in dest { + *listed.entry((d.key.as_path(), d.source_page)).or_insert(0) += 1; + } + let mut sources: Vec> = Vec::new(); + for s in specs { + let d = usize::try_from(s.page_index) + .ok() + .and_then(|i| dest.get(i)) + .ok_or_else(|| invalid_page(s.page_index, dest.len()))?; + let path = groups + .get(d.group) + .map(|g| g.path.clone()) + .unwrap_or_default(); + let name = file_name(Path::new(&path)); + if d.source_page != s.source_page_index { + return Err(reasons::stale(&name)); + } + let times = listed + .get(&(d.key.as_path(), d.source_page)) + .copied() + .unwrap_or(0); + if times != 1 { + return Err(duplicate_page(d.source_page, &name, times)); + } + let pos = match sources.iter().position(|e| e.key == d.key) { + Some(pos) => pos, + None => { + sources.push(SourceEdits { + path, + key: d.key.clone(), + pages: BTreeMap::new(), + }); + sources.len() - 1 + } + }; + if let Some(src) = sources.get_mut(pos) { + src.pages + .entry(d.source_page) + .or_insert_with(|| (s.page_index, Vec::new())) + .1 + .push(s); + } + } + Ok(sources) +} diff --git a/src-tauri/src/pdf_engine/text_edit/export/tests.rs b/src-tauri/src/pdf_engine/text_edit/export/tests.rs new file mode 100644 index 0000000..aff17af --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/export/tests.rs @@ -0,0 +1,236 @@ +//! Unit tests of the Save integration's own checks: proofs and warnings are bound to planned +//! pages explicitly (review-T5 L2), and `verify_final`'s independent #34 assertion (review-T5 M4). + +use super::map::SourceEdits; +use super::{kept_models_budget, outcome, verify_final, PlannedPage, PreparedTextEdits}; +use crate::error::AppError; +use crate::pdf_engine::text_edit::content::page_content; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::gate::{ + verify_edited_copy, BeforeModel, EditedPageInput, ModelPrint, PageProof, PhaseAInput, + PhaseAReport, PopplerRef, TextWarning, +}; +use crate::pdf_engine::text_edit::limits::{PAGE_DECODE_BUDGET, VERIFY_CAP_MARGIN_BYTES}; +use crate::pdf_engine::text_edit::reasons::TextWarningCode; +use crate::pdf_engine::text_edit::rewrite::assemble_page_plan; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::producers::helvetica_page; +use crate::pdf_engine::text_edit::testkit::Scratch; +use crate::pdf_engine::text_edit::tests_gate::{ed, opts, Honest}; +use crate::pdf_engine::validate_output::{content_digest, ContentDigest}; +use std::collections::BTreeMap; +use std::path::PathBuf; +use std::sync::Arc; + +/// Source page 0 planned with no change, saved as destination page 3. +fn planned(dir: &Scratch) -> Vec { + let pdf = helvetica_page(b"BT /F1 12 Tf 72 700 Td (Bound) Tj ET"); + let path = dir.write("bound.pdf", &pdf); + let ctx = SnapshotContext::new(snapshot_from_bytes(&path, pdf, None).expect("snapshot")); + let id = ctx.page_id(0).expect("page"); + let content = + page_content(ctx.doc(), id, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)).expect("content"); + let plan = assemble_page_plan(&content, 0, Vec::new()); + vec![PlannedPage { + dest: 3, + print: ModelPrint::of(&build_page_model(&ctx, 0, None).expect("model")), + content: Arc::new(content), + model: None, + plan, + warnings: Vec::new(), + }] +} + +fn proof(source_page_index: u32) -> PageProof { + PageProof { + source_page_index, + expected_parts: Vec::new(), + part_digests: Vec::new(), + original_edited_digests: Vec::new(), + original_edited_rolling: Vec::new(), + expected_page_digest: content_digest(b""), + records: Vec::new(), + } +} + +fn source() -> SourceEdits<'static> { + SourceEdits { + path: "bound.pdf".into(), + key: PathBuf::from("bound.pdf"), + pages: BTreeMap::new(), + } +} + +fn gate_failure(r: Result, what: &str, case: &str) { + let Err(e) = r else { + panic!("{case}: accepted"); + }; + assert_eq!(e.code, "SOURCE_EDIT_GATE_FAILED", "{case}: {e}"); + let details = e.details.unwrap_or_default(); + assert!(details.contains(what), "{case}: {details}"); +} + +fn report(proofs: Vec, warnings: Vec) -> PhaseAReport { + PhaseAReport { proofs, warnings } +} + +/// review-T5 L2: `outcome` used to key a proof or warning of an unplanned source page to that +/// page's own number, as if it were a destination page. +#[test] +fn l2_proofs_and_warnings_bind_to_planned_pages_only() { + let dir = Scratch::new("export_l2"); + let planned = planned(&dir); + let src = source(); + let edited = PathBuf::from("edited.pdf"); + let ok = outcome( + &src, + &planned, + report(vec![proof(0)], Vec::new()), + edited.clone(), + ) + .unwrap_or_else(|e| panic!("bound proof: {e} {:?}", e.details)); + assert_eq!(ok.proofs.len(), 1); + assert_eq!(ok.proofs[0].0, 3, "keyed by the destination page"); + let r = outcome( + &src, + &planned, + report(vec![proof(1)], Vec::new()), + edited.clone(), + ); + gate_failure(r, "source page 2 was not planned", "unplanned proof"); + let r = outcome( + &src, + &planned, + report(Vec::new(), Vec::new()), + edited.clone(), + ); + gate_failure(r, "0 proofs for 1 planned pages", "missing proof"); + let twice = report(vec![proof(0), proof(0)], Vec::new()); + gate_failure( + outcome(&src, &planned, twice, edited.clone()), + "2 proofs", + "two proofs", + ); + let stray = TextWarning { + page_index: 5, + run_id: "r".into(), + code: TextWarningCode::EditNotVisible, + detail: None, + }; + let r = outcome(&src, &planned, report(vec![proof(0)], vec![stray]), edited); + gate_failure(r, "source page 6 was not planned", "unplanned warning"); +} + +/// review-T5 M4: after Phase B passes on a real staged file, #34's snapshot digest of each edited +/// destination page must be the proof's independently planned digest. +#[test] +fn m4_verify_final_asserts_the_planned_34_digest() { + let pdf = helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello world) Tj ET"); + let Some(h) = Honest::new("export_m4", pdf, 0, &[ed("Hello world", "Hello there")]) else { + return; + }; + let size = std::fs::metadata(&h.staged).map_or(0, |m| m.len()); + let prepared = PreparedTextEdits { + content_groups: Vec::new(), + proofs: vec![(0, h.proof().clone())], + warnings: Vec::new(), + verify_cap: 4 * size + VERIFY_CAP_MARGIN_BYTES, + }; + let planned = h.proof().expected_page_digest; + let run = |digests: &[ContentDigest]| { + verify_final(&h.staged, &prepared, digests, &h.engines, &opts()) + }; + assert!(run(&[planned]).is_ok(), "the planned digest"); + let other = ContentDigest { + hash: planned.hash ^ 1, + len: planned.len, + }; + for (case, digests) in [("altered", vec![other]), ("out of range", Vec::new())] { + let e = run(&digests) + .err() + .unwrap_or_else(|| panic!("{case}: accepted")); + assert_eq!(e.code, "SOURCE_EDIT_GATE_FAILED", "{case}"); + let details = e.details.clone().unwrap_or_default(); + assert!( + details.contains("phase=B check=34 page=1"), + "{case}: {details}" + ); + let suggestion = e.suggestion.clone().unwrap_or_default(); + assert_eq!( + suggestion + .matches("The original file was not changed.") + .count(), + 1, + "{case}: {suggestion}" + ); + } +} + +/// review-final MEDIUM-1: the kept page models are bounded by twice the file, with no 160 MiB +/// cap: on a 151 MiB file the cap kept the source context (6.1 × file) where the models cost less. +#[test] +fn medium1_kept_models_are_bounded_by_twice_the_file_without_a_cap() { + const MIB: u64 = 1 << 20; + assert_eq!(kept_models_budget(151 * MIB), 302 << 20); + assert_eq!( + kept_models_budget(MIB), + 16 << 20, + "at least KEPT_MODELS_MIN" + ); + assert_eq!(kept_models_budget(u64::MAX), usize::MAX); +} + +/// review-final LOW-5: a before-model Phase A builds again from the source context must be the +/// planned one (its print); another page's print is `STALE`, never verified against. +#[test] +fn low5_a_model_built_again_must_match_the_planned_print() { + let Some(h) = Honest::new( + "low5_print", + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + 0, + &[ed("Hello", "Help")], + ) else { + return; + }; + let other = helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hellp) Tj ET"); + let other_path = h.path("other.pdf"); + let other_ctx = + SnapshotContext::new(snapshot_from_bytes(&other_path, other, None).expect("snapshot")); + let other_print = ModelPrint::of(&build_page_model(&other_ctx, 0, None).expect("model")); + let planned_print = ModelPrint::of(&h.model); + assert_ne!(planned_print, other_print); + let phase_a = |print: &ModelPrint| { + let input = PhaseAInput { + before: &h.digest, + before_page_count: h.page_count, + staged: &h.staged, + staged_cap: 1 << 24, + pages: vec![EditedPageInput { + model: BeforeModel::Build { + ctx: &h.ctx, + page_index: 0, + print, + }, + plan: &h.plan, + input_page_index: 0, + input_render: PopplerRef { + pdf: h.source.clone(), + page_1: 1, + }, + }], + source_benign: &[], + }; + verify_edited_copy(&input, &h.engines, h.dir.dir(), &opts()) + }; + phase_a(&planned_print).unwrap_or_else(|e| panic!("planned print: {e} {:?}", e.details)); + let e = phase_a(&other_print) + .err() + .expect("another page's print must fail"); + assert_eq!(e.code, "STALE", "{:?}", e.details); + assert!(e + .details + .unwrap_or_default() + .contains("built again differs")); +} diff --git a/src-tauri/src/pdf_engine/text_edit/fit.rs b/src-tauri/src/pdf_engine/text_edit/fit.rs new file mode 100644 index 0000000..5f0a42c --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fit.rs @@ -0,0 +1,349 @@ +//! Width helpers shared with the DTOs (`word_space`) and, under `cfg(test)`, the editor's width +//! estimate (`fit/estimate.rs`, MEAS-01 golden): the frontend estimates while typing (SPEC §A.8, +//! §D.4); nothing here is authoritative — every commit is planned and verified by +//! `rewrite`/`verify`. + +#[cfg(test)] +mod estimate; + +use crate::pdf_engine::text_edit::fonts::FontModel; +#[cfg(test)] +pub use estimate::{estimate_caret_offsets, estimate_delta_pt}; + +/// Whether the font types a space as the single byte 0x20 (so `Tw` applies to it). +pub fn word_space(model: &FontModel) -> bool { + model + .code_for(' ', &[], &[]) + .is_some_and(|c| model.is_word_space(c)) +} + +#[cfg(test)] +pub fn write_measure_golden(path: &std::path::Path) -> std::io::Result<()> { + let json = golden::build(); + let mut text = serde_json::to_string_pretty(&json).map_err(std::io::Error::other)?; + text.push('\n'); + std::fs::write(path, text) +} + +/// The golden cases (MEAS-01), in the frontend's DTO shapes (`TextRun`, `TextFont`). +#[cfg(test)] +pub(crate) mod golden { + use super::{estimate_caret_offsets, estimate_delta_pt, word_space}; + use crate::pdf_engine::text_edit::context::SnapshotContext; + use crate::pdf_engine::text_edit::fonts::{face_surface, FamilyHint, FontKey, FontModel}; + use crate::pdf_engine::text_edit::reasons::Face; + use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; + use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel, SpaceMode, TextRun}; + use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; + use crate::pdf_engine::text_edit::testkit::fonts::{ + add_type0, glyph_name, tounicode_bfchar, Program, Type0Font, + }; + use crate::pdf_engine::text_edit::testkit::producers::{self as fx, DocBuilder, PageSpec}; + use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; + use serde_json::{json, Value}; + use std::sync::Arc; + + pub(crate) fn key(m: &FontModel) -> String { + match m.key { + FontKey::Indirect((n, g)) => format!("f{n}-{g}"), + FontKey::Direct { name_hash, .. } => format!("d{name_hash:x}"), + } + } + + fn face_str(f: Face) -> &'static str { + match f { + Face::Regular => "regular", + Face::Bold => "bold", + Face::Italic => "italic", + Face::BoldItalic => "boldItalic", + } + } + + fn font_dto(m: &FontModel) -> Value { + let alphabet = m.alphabet(); + json!({ + "key": key(m), + "displayName": m.display_name, + "familyHint": match m.family_hint { FamilyHint::Serif => "serif", FamilyHint::Sans => "sans", FamilyHint::Mono => "mono" }, + "embedded": m.embedded, + "subset": m.subset, + "alphabet": alphabet.iter().map(|(c, _)| *c).collect::(), + "widths": alphabet.iter().map(|(_, w)| *w).collect::>(), + "wordSpace": word_space(m), + }) + } + + fn keys_of(fonts: &[(Vec, Arc)], names: &[Vec]) -> Vec { + names + .iter() + .filter_map(|n| fonts.iter().find(|(x, _)| x == n).map(|(_, m)| key(m))) + .collect() + } + + fn run_dto(model: &PageModel, run: &TextRun) -> Value { + let fonts = &model.walk.page_fonts; + let primary = run + .surface + .first() + .and_then(|n| fonts.iter().find(|(x, _)| x == n)) + .map(|(_, m)| Arc::clone(m)); + let face = |f: Face| { + if f == run.face { + return json!({ "available": true, "surface": keys_of(fonts, &run.surface) }); + } + let s = primary.as_ref().and_then(|p| face_surface(fonts, p, f)); + json!({ + "available": s.is_some(), + "surface": s.map(|s| s.fonts.iter().map(|(_, m)| key(m)).collect::>()).unwrap_or_default(), + }) + }; + json!({ + "id": run.id, + "order": run.order, + "line": run.line, + "text": run.text, + "rect": { "x": run.rect[0], "y": run.rect[1], "w": run.rect[2], "h": run.rect[3] }, + "origin": { "x": run.origin.0, "y": run.origin.1 }, + "dir": { "x": run.dir.0, "y": run.dir.1 }, + "ascent": run.ascent, + "descent": run.descent, + "caretOffsets": run.caret_offsets, + "editable": true, + "reason": null, + "metrics": { + "surface": keys_of(fonts, &run.surface), + "tfSize": run.tfs, + "effectiveSize": run.effective_size, + "charSpacing": run.tc, + "wordSpacing": run.tw, + "hScale": run.th, + "textToUser": run.text_to_user_x, + "letterSpacingPt": run.tc * run.th * run.text_to_user_x, + "spaceMode": match run.space_mode { SpaceMode::Glyph => "glyph", SpaceMode::Kern => "kern" }, + "kernSpace": run.kern_space, + "originalWidth": run.original_extent, + "visibleExtent": run.visible_extent, + "nextObstacle": run.next_obstacle, + }, + "style": { + "fill": run.fill_hex, + "sizeChangeable": !run.font_from_extgstate, + "colourChangeable": run.tr == 0, + "face": face_str(run.face), + "faces": { + "regular": face(Face::Regular), + "bold": face(Face::Bold), + "italic": face(Face::Italic), + "boldItalic": face(Face::BoldItalic), + }, + }, + "substituted": run.substituted, + }) + } + + fn style_dto(s: &SourceTextStyleIn) -> Value { + let mut m = serde_json::Map::new(); + if let Some(v) = s.size_pt { + m.insert("sizePt".into(), json!(v)); + } + if let Some(f) = s.face { + m.insert("face".into(), json!(face_str(f))); + } + if let Some(f) = &s.fill { + m.insert("fill".into(), json!(f)); + } + if let Some(v) = s.letter_spacing_pt { + m.insert("letterSpacingPt".into(), json!(v)); + } + Value::Object(m) + } + + fn model_of(pdf: Vec) -> PageModel { + let snap = snapshot_from_bytes(std::path::Path::new("golden.pdf"), pdf, None) + .unwrap_or_else(|e| panic!("golden fixture: {e}")); + let ctx = SnapshotContext::new(snap); + build_page_model(&ctx, 0, None).unwrap_or_else(|e| panic!("golden model: {e}")) + } + + fn case( + name: &str, + pdf: Vec, + run_text: &str, + text: &str, + style: SourceTextStyleIn, + ) -> Value { + let model = model_of(pdf); + let run = model + .runs + .iter() + .find(|r| r.text == run_text) + .unwrap_or_else(|| panic!("golden {name}: no run {run_text:?}")); + let fonts = &model.walk.page_fonts; + json!({ + "name": name, + "run": run_dto(&model, run), + "fonts": fonts.iter().map(|(_, m)| font_dto(m)).collect::>(), + "text": text, + "style": style_dto(&style), + "deltaPt": estimate_delta_pt(run, fonts, text, &style), + "caretOffsets": estimate_caret_offsets(run, fonts, text, &style), + }) + } + + /// Helvetica and Helvetica-Bold on one page (a bold sibling for the face case). + fn two_faces() -> Vec { + let mut d = DocBuilder::new(); + let r = d.add(fx::HELVETICA); + let b = d.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica-Bold /Encoding /WinAnsiEncoding >>", + ); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Total due) Tj ET BT /F2 12 Tf 72 680 Td (Bold) Tj ET", + &format!("/Font << /F1 {r} 0 R /F2 {b} 0 R >>"), + )); + d.build() + } + + /// FX-WORD-TR's shape (a WinAnsi TrueType subset at 500 and its Identity-H companion for + /// "ğışİ", one `BT … Tm` per font segment, joined into one run) with the companion's widths + /// at 640: in FX-WORD-TR every width is 500, the missing-glyph width too, so `deltaPt` could + /// not tell which font a character was measured in. + fn sibling_widths() -> Vec { + const COMPANION_WIDTH: f64 = 640.0; + let mut d = DocBuilder::new(); + let f1 = fx::word_font(&mut d.b, "ABCDEF+Calibri", "Sa lk Bakan Raporu"); + let tr = fx::WORD_TR_CID_CHARS; + let mut ttf = TtfBuilder::new(); + let mut map = Vec::new(); + for (i, ch) in tr.chars().enumerate() { + ttf.unicode_glyph(ch, &glyph_name(ch), true); + map.push((i as u32 + 1, 2, ch.to_string())); + } + let refs: Vec<(u32, usize, &str)> = + map.iter().map(|(c, l, t)| (*c, *l, t.as_str())).collect(); + let mut companion = Type0Font::new("CIDFontType2", "ABCDEF+Calibri"); + companion.program = Program::TrueType(ttf.build()); + companion.cid_to_gid = Some(None); + companion.flags = Some(32); + companion.w = Some(format!( + "[1 [{}]]", + vec![COMPANION_WIDTH.to_string(); map.len()].join(" ") + )); + companion.tounicode = Some(tounicode_bfchar(&refs)); + let f2 = add_type0(&mut d.b, &companion); + let segs: [(&str, String, usize, f64); 7] = [ + ("F1", "(Sa)".into(), 2, 500.0), + ( + "F2", + format!("<{}>", fx::cid_hex(tr, "ğ")), + 1, + COMPANION_WIDTH, + ), + ("F1", "(l)".into(), 1, 500.0), + ( + "F2", + format!("<{}>", fx::cid_hex(tr, "ı")), + 1, + COMPANION_WIDTH, + ), + ("F1", "(k Bakanl)".into(), 8, 500.0), + ( + "F2", + format!("<{}>", fx::cid_hex(tr, "ığı")), + 3, + COMPANION_WIDTH, + ), + ("F1", "( Raporu)".into(), 7, 500.0), + ]; + let mut x = 72.0f64; + let mut body = String::new(); + for (font, shown, n, width) in segs { + body.push_str(&format!( + "BT /{font} 11 Tf 1 0 0 1 {x:.4} 700 Tm [{shown}]TJ ET " + )); + x += width / 1000.0 * 11.0 * n as f64; + } + d.page(PageSpec::new( + body.as_bytes(), + &format!("/Font << /F1 {f1} 0 R /F2 {f2} 0 R >>"), + )); + d.build() + } + + pub(crate) fn build() -> Value { + let plain = SourceTextStyleIn::default; + let cases = vec![ + case( + "std14-helvetica-12", + fx::helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Help me", + plain(), + ), + case( + "tz-80", + fx::helvetica_page(b"BT /F1 12 Tf 80 Tz 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello world", + plain(), + ), + case( + "tc-0.5", + fx::helvetica_page(b"BT /F1 12 Tf 0.5 Tc 72 700 Td (Hello) Tj ET"), + "Hello", + "Hi", + plain(), + ), + case( + "kern-space", + fx::pdftex(), + "Hello World", + "Hello Wide World", + plain(), + ), + case( + "sibling-surface", + sibling_widths(), + "Sağlık Bakanlığı Raporu", + "Sağlık Bakanlığı Raporları", + plain(), + ), + case( + "size-change", + fx::helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello", + SourceTextStyleIn { + size_pt: Some(14.0), + ..plain() + }, + ), + case( + "letter-spacing", + fx::helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello", + SourceTextStyleIn { + letter_spacing_pt: Some(0.5), + ..plain() + }, + ), + case( + "face-bold", + two_faces(), + "Total due", + "Total due", + SourceTextStyleIn { + face: Some(Face::Bold), + ..plain() + }, + ), + ]; + json!({ + "version": 1, + "about": "Rust fit::estimate_delta_pt / estimate_caret_offsets on generated fixtures (MEAS-01). Regenerate: OFFPDF_UPDATE_GOLDEN=1 cargo test --lib -j 6 meas_01", + "tolerance": 1e-6, + "cases": cases, + }) + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fit/estimate.rs b/src-tauri/src/pdf_engine/text_edit/fit/estimate.rs new file mode 100644 index 0000000..ef87a14 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fit/estimate.rs @@ -0,0 +1,140 @@ +//! The editor's width estimate (SPEC §A.8, §D.4) in Rust, test-only: the same model as the +//! frontend's `estimateDeltaPt`/`estimateCaretOffsets` (`src/lib/editor/sourceText.ts`), +//! golden-tested against `src/lib/editor/__fixtures__/text-measure-golden.json` (MEAS-01). Rust +//! production never estimates: verdicts carry the planned `delta_pt` and caret offsets. +//! +//! For a run and a requested style: `tf = Tfs × size / effective` (or `Tfs`), `k = Th × +//! text_to_user_x`, `tc = spacing / k` (or `Tc`); per character: a kern-mode space advances +//! `−kern_space/1000 × tf`; any other character the first surface font that can type it +//! (`w/1000 × tf + tc`, plus `Tw` for a 0x20 space), else `MISSING_GLYPH_EM × tf + tc`; each +//! advance is clamped at 0 and scaled by `k`. The delta is relative: new text with the style +//! minus the current text without it. + +use super::word_space; +use crate::pdf_engine::text_edit::fonts::{face_surface, FontModel}; +use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; +use crate::pdf_engine::text_edit::runs::{SpaceMode, TextRun}; +use std::sync::Arc; + +/// Advance (em) of a character no surface font can draw, in the estimate only (same as the TS). +pub const MISSING_GLYPH_EM: f64 = 0.5; + +/// The fonts that type `face` for `run`: its own surface for its own face, else the face surface +/// of the page fonts (empty when the page has none). +fn surface_for( + run: &TextRun, + fonts: &[(Vec, Arc)], + face: Option, +) -> Vec> { + let by_name = |name: &[u8]| { + fonts + .iter() + .find(|(n, _)| n.as_slice() == name) + .map(|(_, m)| Arc::clone(m)) + }; + match face { + Some(f) if f != run.face => run + .surface + .first() + .and_then(|n| by_name(n)) + .and_then(|primary| face_surface(fonts, &primary, f)) + .map(|s| s.fonts.into_iter().map(|(_, m)| m).collect()) + .unwrap_or_default(), + _ => run.surface.iter().filter_map(|n| by_name(n)).collect(), + } +} + +struct Model { + tf: f64, + tc: f64, + k: f64, + surface: Vec<(Vec<(char, f64)>, bool)>, +} + +fn model_for( + run: &TextRun, + fonts: &[(Vec, Arc)], + style: &SourceTextStyleIn, +) -> Model { + let k_raw = run.th * run.text_to_user_x; + let k = if k_raw.is_finite() { k_raw } else { 0.0 }; + let tf = match style.size_pt { + Some(pt) if pt.is_finite() && run.effective_size > 0.0 => run.tfs * pt / run.effective_size, + _ => run.tfs, + }; + let tc = match style.letter_spacing_pt { + Some(pt) if pt.is_finite() && k > 0.0 => pt / k, + _ => run.tc, + }; + let surface = surface_for(run, fonts, style.face) + .iter() + .map(|m| (m.alphabet(), word_space(m))) + .collect(); + Model { tf, tc, k, surface } +} + +fn advances( + run: &TextRun, + fonts: &[(Vec, Arc)], + text: &str, + style: &SourceTextStyleIn, +) -> Vec { + let m = model_for(run, fonts, style); + text.chars() + .map(|ch| { + let advance = if ch == ' ' && run.space_mode == SpaceMode::Kern { + -run.kern_space / 1000.0 * m.tf + } else { + let hit = m.surface.iter().find_map(|(alphabet, ws)| { + alphabet + .binary_search_by(|(c, _)| c.cmp(&ch)) + .ok() + .and_then(|i| alphabet.get(i)) + .map(|(_, w)| (if w.is_finite() { *w } else { 0.0 }, *ws)) + }); + let em = hit.map_or(MISSING_GLYPH_EM, |(w, _)| w / 1000.0); + let word = if ch == ' ' && hit.is_some_and(|(_, ws)| ws) { + run.tw + } else { + 0.0 + }; + em * m.tf + m.tc + word + }; + if advance.is_finite() { + advance.max(0.0) * m.k + } else { + 0.0 + } + }) + .collect() +} + +/// Estimated width change in points (positive = wider) of `text` with `style` on `run`. +pub fn estimate_delta_pt( + run: &TextRun, + fonts: &[(Vec, Arc)], + text: &str, + style: &SourceTextStyleIn, +) -> f64 { + let new: f64 = advances(run, fonts, text, style).iter().sum(); + let old: f64 = advances(run, fonts, &run.text, &SourceTextStyleIn::default()) + .iter() + .sum(); + new - old +} + +/// `chars + 1` caret distances along `dir` for `text`, same model as `estimate_delta_pt`. +pub fn estimate_caret_offsets( + run: &TextRun, + fonts: &[(Vec, Arc)], + text: &str, + style: &SourceTextStyleIn, +) -> Vec { + let mut out = vec![0.0]; + let mut at = 0.0; + for a in advances(run, fonts, text, style) { + at += a; + out.push(at); + } + out +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/agl.rs b/src-tauri/src/pdf_engine/text_edit/fonts/agl.rs new file mode 100644 index 0000000..708da1b --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/agl.rs @@ -0,0 +1,69 @@ +//! Adobe Glyph List lookup (SPEC §A.3.1 item 2, §B.9.2): glyph name → one Unicode scalar. +//! +//! Order: the full AGL table (`agl_data.rs`, 4,495 names, exact match), then `uniXXXX` (exactly +//! one group of four hex digits), then `uXXXX`–`uXXXXXX`. Everything else — suffixed names +//! (`a.sc`), multi-group `uni` names (`uni00410042`), surrogate values — has no Unicode value. + +use super::agl_data::AGL; + +/// The AGL value of `name` (exact, byte-wise match), if the table lists it. +pub fn agl_value(name: &str) -> Option { + AGL.binary_search_by(|(n, _)| n.as_bytes().cmp(name.as_bytes())) + .ok() + .and_then(|i| AGL.get(i)) + .map(|(_, v)| *v) +} + +/// The Unicode scalar a glyph name stands for (AGL, then `uniXXXX`, then `uXXXX[XX]`). +pub fn glyph_name_char(name: &str) -> Option { + if let Some(v) = agl_value(name) { + return char::from_u32(v); + } + if let Some(hex) = name.strip_prefix("uni") { + // `uni` names carry exactly one 4-digit group; longer ones name sequences (ligatures). + return if hex.len() == 4 { + hex_scalar(hex) + } else { + None + }; + } + match name.strip_prefix('u') { + Some(hex) if (4..=6).contains(&hex.len()) => hex_scalar(hex), + _ => None, + } +} + +/// `hex` as a Unicode scalar; `None` for non-hex digits, surrogates and values past U+10FFFF. +fn hex_scalar(hex: &str) -> Option { + if hex.is_empty() || !hex.bytes().all(|b| b.is_ascii_hexdigit()) { + return None; + } + u32::from_str_radix(hex, 16).ok().and_then(char::from_u32) +} + +/// Number of names in the table (FONT-01 checks it against lopdf's 4,495). +#[cfg(test)] +pub fn agl_len() -> usize { + AGL.len() +} + +/// Whether the table is strictly sorted by name bytes (binary search precondition; FONT-01). +#[cfg(test)] +pub fn agl_is_sorted() -> bool { + AGL.windows(2).all(|w| match w { + [(a, _), (b, _)] => a.as_bytes() < b.as_bytes(), + _ => true, + }) +} + +/// Every AGL name of `ch`, shortest first (test fixtures name their glyphs with it). +#[cfg(test)] +pub fn names_for(ch: char) -> Vec<&'static str> { + let mut names: Vec<&'static str> = AGL + .iter() + .filter(|(_, v)| *v == u32::from(ch)) + .map(|(n, _)| *n) + .collect(); + names.sort_by_key(|n| (n.len(), *n)); + names +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/agl_data.rs b/src-tauri/src/pdf_engine/text_edit/fonts/agl_data.rs new file mode 100644 index 0000000..4a150eb --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/agl_data.rs @@ -0,0 +1,783 @@ +//! GENERATED DATA — do not edit by hand (SPEC §B.9.2). Glyph name → Unicode (BMP) code point. +//! +//! Provenance: every `pub const NAME: u16 = 0xXXXX;` of `lopdf` 0.34.0 +//! `src/encodings/glyphnames.rs` (4,495 names), written as `(name, code point)` and sorted by the +//! name's bytes for binary search. Derived from lopdf 0.34.0, MIT (Copyright (c) 2016 Junfeng Liu); +//! originally the Adobe Glyph List (glyphlist.txt, Copyright 2002–2010 Adobe Systems +//! Incorporated, BSD-3-Clause) plus the TeX glyph names lopdf carries (e.g. `angleleftbig`, +//! `f_f`, `f_i`). lopdf stores one UTF-16 unit per name, so the few AGL entries that map to a +//! sequence (Hebrew point ligatures such as `dalethatafpatah`) keep only their first code point; +//! Hebrew is refused (`RIGHT_TO_LEFT`) before such a value could be written. + +#[rustfmt::skip] +pub(super) static AGL: &[(&str, u32)] = &[ +("A",0x0041),("AE",0x00C6),("AEacute",0x01FC),("AEmacron",0x01E2),("AEsmall",0xF7E6),("Aacute",0x00C1),("Aacutesmall",0xF7E1),("Abreve",0x0102), +("Abreveacute",0x1EAE),("Abrevecyrillic",0x04D0),("Abrevedotbelow",0x1EB6),("Abrevegrave",0x1EB0),("Abrevehookabove",0x1EB2),("Abrevetilde",0x1EB4), +("Acaron",0x01CD),("Acircle",0x24B6),("Acircumflex",0x00C2),("Acircumflexacute",0x1EA4),("Acircumflexdotbelow",0x1EAC),("Acircumflexgrave",0x1EA6), +("Acircumflexhookabove",0x1EA8),("Acircumflexsmall",0xF7E2),("Acircumflextilde",0x1EAA),("Acute",0xF6C9),("Acutesmall",0xF7B4),("Acyrillic",0x0410), +("Adblgrave",0x0200),("Adieresis",0x00C4),("Adieresiscyrillic",0x04D2),("Adieresismacron",0x01DE),("Adieresissmall",0xF7E4),("Adotbelow",0x1EA0), +("Adotmacron",0x01E0),("Agrave",0x00C0),("Agravesmall",0xF7E0),("Ahookabove",0x1EA2),("Aiecyrillic",0x04D4),("Ainvertedbreve",0x0202), +("Alpha",0x0391),("Alphatonos",0x0386),("Amacron",0x0100),("Amonospace",0xFF21),("Aogonek",0x0104),("Aring",0x00C5),("Aringacute",0x01FA), +("Aringbelow",0x1E00),("Aringsmall",0xF7E5),("Asmall",0xF761),("Atilde",0x00C3),("Atildesmall",0xF7E3),("Aybarmenian",0x0531),("B",0x0042), +("Bcircle",0x24B7),("Bdotaccent",0x1E02),("Bdotbelow",0x1E04),("Becyrillic",0x0411),("Benarmenian",0x0532),("Beta",0x0392),("Bhook",0x0181), +("Blinebelow",0x1E06),("Bmonospace",0xFF22),("Brevesmall",0xF6F4),("Bsmall",0xF762),("Btopbar",0x0182),("C",0x0043),("Caarmenian",0x053E), +("Cacute",0x0106),("Caron",0xF6CA),("Caronsmall",0xF6F5),("Ccaron",0x010C),("Ccedilla",0x00C7),("Ccedillaacute",0x1E08),("Ccedillasmall",0xF7E7), +("Ccircle",0x24B8),("Ccircumflex",0x0108),("Cdot",0x010A),("Cdotaccent",0x010A),("Cedillasmall",0xF7B8),("Chaarmenian",0x0549), +("Cheabkhasiancyrillic",0x04BC),("Checyrillic",0x0427),("Chedescenderabkhasiancyrillic",0x04BE),("Chedescendercyrillic",0x04B6), +("Chedieresiscyrillic",0x04F4),("Cheharmenian",0x0543),("Chekhakassiancyrillic",0x04CB),("Cheverticalstrokecyrillic",0x04B8),("Chi",0x03A7), +("Chook",0x0187),("Circumflexsmall",0xF6F6),("Cmonospace",0xFF23),("Coarmenian",0x0551),("Csmall",0xF763),("D",0x0044),("DZ",0x01F1), +("DZcaron",0x01C4),("Daarmenian",0x0534),("Dafrican",0x0189),("Dbar",0x0110),("Dcaron",0x010E),("Dcedilla",0x1E10),("Dcircle",0x24B9), +("Dcircumflexbelow",0x1E12),("Dcroat",0x0110),("Ddotaccent",0x1E0A),("Ddotbelow",0x1E0C),("Decyrillic",0x0414),("Deicoptic",0x03EE),("Delta",0x2206), +("Deltagreek",0x0394),("Dhook",0x018A),("Dieresis",0xF6CB),("DieresisAcute",0xF6CC),("DieresisGrave",0xF6CD),("Dieresissmall",0xF7A8), +("Digammagreek",0x03DC),("Djecyrillic",0x0402),("Dlinebelow",0x1E0E),("Dmonospace",0xFF24),("Dotaccentsmall",0xF6F7),("Dslash",0x0110), +("Dsmall",0xF764),("Dtopbar",0x018B),("Dz",0x01F2),("Dzcaron",0x01C5),("Dzeabkhasiancyrillic",0x04E0),("Dzecyrillic",0x0405),("Dzhecyrillic",0x040F), +("E",0x0045),("Eacute",0x00C9),("Eacutesmall",0xF7E9),("Ebreve",0x0114),("Ecaron",0x011A),("Ecedillabreve",0x1E1C),("Echarmenian",0x0535), +("Ecircle",0x24BA),("Ecircumflex",0x00CA),("Ecircumflexacute",0x1EBE),("Ecircumflexbelow",0x1E18),("Ecircumflexdotbelow",0x1EC6), +("Ecircumflexgrave",0x1EC0),("Ecircumflexhookabove",0x1EC2),("Ecircumflexsmall",0xF7EA),("Ecircumflextilde",0x1EC4),("Ecyrillic",0x0404), +("Edblgrave",0x0204),("Edieresis",0x00CB),("Edieresissmall",0xF7EB),("Edot",0x0116),("Edotaccent",0x0116),("Edotbelow",0x1EB8),("Efcyrillic",0x0424), +("Egrave",0x00C8),("Egravesmall",0xF7E8),("Eharmenian",0x0537),("Ehookabove",0x1EBA),("Eightroman",0x2167),("Einvertedbreve",0x0206), +("Eiotifiedcyrillic",0x0464),("Elcyrillic",0x041B),("Elevenroman",0x216A),("Emacron",0x0112),("Emacronacute",0x1E16),("Emacrongrave",0x1E14), +("Emcyrillic",0x041C),("Emonospace",0xFF25),("Encyrillic",0x041D),("Endescendercyrillic",0x04A2),("Eng",0x014A),("Enghecyrillic",0x04A4), +("Enhookcyrillic",0x04C7),("Eogonek",0x0118),("Eopen",0x0190),("Epsilon",0x0395),("Epsilontonos",0x0388),("Ercyrillic",0x0420),("Ereversed",0x018E), +("Ereversedcyrillic",0x042D),("Escyrillic",0x0421),("Esdescendercyrillic",0x04AA),("Esh",0x01A9),("Esmall",0xF765),("Eta",0x0397), +("Etarmenian",0x0538),("Etatonos",0x0389),("Eth",0x00D0),("Ethsmall",0xF7F0),("Etilde",0x1EBC),("Etildebelow",0x1E1A),("Euro",0x20AC),("Ezh",0x01B7), +("Ezhcaron",0x01EE),("Ezhreversed",0x01B8),("F",0x0046),("FFIsmall",0xD807),("FFLsmall",0xD808),("FFsmall",0xD804),("FIsmall",0xD805), +("FLsmall",0xD806),("Fcircle",0x24BB),("Fdotaccent",0x1E1E),("Feharmenian",0x0556),("Feicoptic",0x03E4),("Fhook",0x0191),("Fitacyrillic",0x0472), +("Fiveroman",0x2164),("Fmonospace",0xFF26),("Fourroman",0x2163),("Fsmall",0xF766),("G",0x0047),("GBsquare",0x3387),("Gacute",0x01F4),("Gamma",0x0393), +("Gammaafrican",0x0194),("Gangiacoptic",0x03EA),("Gbreve",0x011E),("Gcaron",0x01E6),("Gcedilla",0x0122),("Gcircle",0x24BC),("Gcircumflex",0x011C), +("Gcommaaccent",0x0122),("Gdot",0x0120),("Gdotaccent",0x0120),("Gecyrillic",0x0413),("Germandbls",0x0053),("Germandbls2",0x1E9E), +("Germandblssmall",0xD803),("Ghadarmenian",0x0542),("Ghemiddlehookcyrillic",0x0494),("Ghestrokecyrillic",0x0492),("Gheupturncyrillic",0x0490), +("Ghook",0x0193),("Gimarmenian",0x0533),("Gjecyrillic",0x0403),("Gmacron",0x1E20),("Gmonospace",0xFF27),("Grave",0xF6CE),("Gravesmall",0xF760), +("Gsmall",0xF767),("Gsmallhook",0x029B),("Gstroke",0x01E4),("H",0x0048),("H18533",0x25CF),("H18543",0x25AA),("H18551",0x25AB),("H22073",0x25A1), +("HPsquare",0x33CB),("Haabkhasiancyrillic",0x04A8),("Hadescendercyrillic",0x04B2),("Hardsigncyrillic",0x042A),("Hbar",0x0126),("Hbrevebelow",0x1E2A), +("Hcedilla",0x1E28),("Hcircle",0x24BD),("Hcircumflex",0x0124),("Hdieresis",0x1E26),("Hdotaccent",0x1E22),("Hdotbelow",0x1E24),("Hmonospace",0xFF28), +("Hoarmenian",0x0540),("Horicoptic",0x03E8),("Hsmall",0xF768),("Hungarumlaut",0xF6CF),("Hungarumlautsmall",0xF6F8),("Hzsquare",0x3390),("I",0x0049), +("IAcyrillic",0x042F),("IJ",0x0132),("IUcyrillic",0x042E),("Iacute",0x00CD),("Iacutesmall",0xF7ED),("Ibreve",0x012C),("Icaron",0x01CF), +("Icircle",0x24BE),("Icircumflex",0x00CE),("Icircumflexsmall",0xF7EE),("Icyrillic",0x0406),("Idblgrave",0x0208),("Idieresis",0x00CF), +("Idieresisacute",0x1E2E),("Idieresiscyrillic",0x04E4),("Idieresissmall",0xF7EF),("Idot",0x0130),("Idotaccent",0x0130),("Idotbelow",0x1ECA), +("Iebrevecyrillic",0x04D6),("Iecyrillic",0x0415),("Ifractur",0x2111),("Ifraktur",0x2111),("Igrave",0x00CC),("Igravesmall",0xF7EC), +("Ihookabove",0x1EC8),("Iicyrillic",0x0418),("Iinvertedbreve",0x020A),("Iishortcyrillic",0x0419),("Imacron",0x012A),("Imacroncyrillic",0x04E2), +("Imonospace",0xFF29),("Iniarmenian",0x053B),("Iocyrillic",0x0401),("Iogonek",0x012E),("Iota",0x0399),("Iotaafrican",0x0196),("Iotadieresis",0x03AA), +("Iotatonos",0x038A),("Ismall",0xF769),("Istroke",0x0197),("Itilde",0x0128),("Itildebelow",0x1E2C),("Izhitsacyrillic",0x0474), +("Izhitsadblgravecyrillic",0x0476),("J",0x004A),("Jaarmenian",0x0541),("Jcircle",0x24BF),("Jcircumflex",0x0134),("Jecyrillic",0x0408), +("Jheharmenian",0x054B),("Jmonospace",0xFF2A),("Jsmall",0xF76A),("K",0x004B),("KBsquare",0x3385),("KKsquare",0x33CD),("Kabashkircyrillic",0x04A0), +("Kacute",0x1E30),("Kacyrillic",0x041A),("Kadescendercyrillic",0x049A),("Kahookcyrillic",0x04C3),("Kappa",0x039A),("Kastrokecyrillic",0x049E), +("Kaverticalstrokecyrillic",0x049C),("Kcaron",0x01E8),("Kcedilla",0x0136),("Kcircle",0x24C0),("Kcommaaccent",0x0136),("Kdotbelow",0x1E32), +("Keharmenian",0x0554),("Kenarmenian",0x053F),("Khacyrillic",0x0425),("Kheicoptic",0x03E6),("Khook",0x0198),("Kjecyrillic",0x040C), +("Klinebelow",0x1E34),("Kmonospace",0xFF2B),("Koppacyrillic",0x0480),("Koppagreek",0x03DE),("Ksicyrillic",0x046E),("Ksmall",0xF76B),("L",0x004C), +("LJ",0x01C7),("LL",0xF6BF),("Lacute",0x0139),("Lambda",0x039B),("Lcaron",0x013D),("Lcedilla",0x013B),("Lcircle",0x24C1),("Lcircumflexbelow",0x1E3C), +("Lcommaaccent",0x013B),("Ldot",0x013F),("Ldotaccent",0x013F),("Ldotbelow",0x1E36),("Ldotbelowmacron",0x1E38),("Liwnarmenian",0x053C),("Lj",0x01C8), +("Ljecyrillic",0x0409),("Llinebelow",0x1E3A),("Lmonospace",0xFF2C),("Lslash",0x0141),("Lslashsmall",0xF6F9),("Lsmall",0xF76C),("M",0x004D), +("MBsquare",0x3386),("Macron",0xF6D0),("Macronsmall",0xF7AF),("Macute",0x1E3E),("Mcircle",0x24C2),("Mdotaccent",0x1E40),("Mdotbelow",0x1E42), +("Menarmenian",0x0544),("Mmonospace",0xFF2D),("Msmall",0xF76D),("Mturned",0x019C),("Mu",0x039C),("N",0x004E),("NJ",0x01CA),("Nacute",0x0143), +("Ncaron",0x0147),("Ncedilla",0x0145),("Ncircle",0x24C3),("Ncircumflexbelow",0x1E4A),("Ncommaaccent",0x0145),("Ndotaccent",0x1E44), +("Ndotbelow",0x1E46),("Ng",0x014A),("Nhookleft",0x019D),("Nineroman",0x2168),("Nj",0x01CB),("Njecyrillic",0x040A),("Nlinebelow",0x1E48), +("Nmonospace",0xFF2E),("Nowarmenian",0x0546),("Nsmall",0xF76E),("Ntilde",0x00D1),("Ntildesmall",0xF7F1),("Nu",0x039D),("O",0x004F),("OE",0x0152), +("OEsmall",0xF6FA),("Oacute",0x00D3),("Oacutesmall",0xF7F3),("Obarredcyrillic",0x04E8),("Obarreddieresiscyrillic",0x04EA),("Obreve",0x014E), +("Ocaron",0x01D1),("Ocenteredtilde",0x019F),("Ocircle",0x24C4),("Ocircumflex",0x00D4),("Ocircumflexacute",0x1ED0),("Ocircumflexdotbelow",0x1ED8), +("Ocircumflexgrave",0x1ED2),("Ocircumflexhookabove",0x1ED4),("Ocircumflexsmall",0xF7F4),("Ocircumflextilde",0x1ED6),("Ocyrillic",0x041E), +("Odblacute",0x0150),("Odblgrave",0x020C),("Odieresis",0x00D6),("Odieresiscyrillic",0x04E6),("Odieresissmall",0xF7F6),("Odotbelow",0x1ECC), +("Ogoneksmall",0xF6FB),("Ograve",0x00D2),("Ogravesmall",0xF7F2),("Oharmenian",0x0555),("Ohm",0x2126),("Ohookabove",0x1ECE),("Ohorn",0x01A0), +("Ohornacute",0x1EDA),("Ohorndotbelow",0x1EE2),("Ohorngrave",0x1EDC),("Ohornhookabove",0x1EDE),("Ohorntilde",0x1EE0),("Ohungarumlaut",0x0150), +("Oi",0x01A2),("Oinvertedbreve",0x020E),("Omacron",0x014C),("Omacronacute",0x1E52),("Omacrongrave",0x1E50),("Omega",0x2126),("Omegacyrillic",0x0460), +("Omegagreek",0x03A9),("Omegaroundcyrillic",0x047A),("Omegatitlocyrillic",0x047C),("Omegatonos",0x038F),("Omicron",0x039F),("Omicrontonos",0x038C), +("Omonospace",0xFF2F),("Oneroman",0x2160),("Oogonek",0x01EA),("Oogonekmacron",0x01EC),("Oopen",0x0186),("Oslash",0x00D8),("Oslashacute",0x01FE), +("Oslashsmall",0xF7F8),("Osmall",0xF76F),("Ostrokeacute",0x01FE),("Otcyrillic",0x047E),("Otilde",0x00D5),("Otildeacute",0x1E4C), +("Otildedieresis",0x1E4E),("Otildesmall",0xF7F5),("P",0x0050),("Pacute",0x1E54),("Pcircle",0x24C5),("Pdotaccent",0x1E56),("Pecyrillic",0x041F), +("Peharmenian",0x054A),("Pemiddlehookcyrillic",0x04A6),("Phi",0x03A6),("Phook",0x01A4),("Pi",0x03A0),("Piwrarmenian",0x0553),("Pmonospace",0xFF30), +("Psi",0x03A8),("Psicyrillic",0x0470),("Psmall",0xF770),("Q",0x0051),("Qcircle",0x24C6),("Qmonospace",0xFF31),("Qsmall",0xF771),("R",0x0052), +("Raarmenian",0x054C),("Racute",0x0154),("Rcaron",0x0158),("Rcedilla",0x0156),("Rcircle",0x24C7),("Rcommaaccent",0x0156),("Rdblgrave",0x0210), +("Rdotaccent",0x1E58),("Rdotbelow",0x1E5A),("Rdotbelowmacron",0x1E5C),("Reharmenian",0x0550),("Rfractur",0x211C),("Rfraktur",0x211C),("Rho",0x03A1), +("Ringsmall",0xF6FC),("Rinvertedbreve",0x0212),("Rlinebelow",0x1E5E),("Rmonospace",0xFF32),("Rsmall",0xF772),("Rsmallinverted",0x0281), +("Rsmallinvertedsuperior",0x02B6),("S",0x0053),("SF010000",0x250C),("SF020000",0x2514),("SF030000",0x2510),("SF040000",0x2518),("SF050000",0x253C), +("SF060000",0x252C),("SF070000",0x2534),("SF080000",0x251C),("SF090000",0x2524),("SF100000",0x2500),("SF110000",0x2502),("SF190000",0x2561), +("SF200000",0x2562),("SF210000",0x2556),("SF220000",0x2555),("SF230000",0x2563),("SF240000",0x2551),("SF250000",0x2557),("SF260000",0x255D), +("SF270000",0x255C),("SF280000",0x255B),("SF360000",0x255E),("SF370000",0x255F),("SF380000",0x255A),("SF390000",0x2554),("SF400000",0x2569), +("SF410000",0x2566),("SF420000",0x2560),("SF430000",0x2550),("SF440000",0x256C),("SF450000",0x2567),("SF460000",0x2568),("SF470000",0x2564), +("SF480000",0x2565),("SF490000",0x2559),("SF500000",0x2558),("SF510000",0x2552),("SF520000",0x2553),("SF530000",0x256B),("SF540000",0x256A), +("SS",0x0053),("SSsmall",0xD803),("Sacute",0x015A),("Sacutedotaccent",0x1E64),("Sampigreek",0x03E0),("Scaron",0x0160),("Scarondotaccent",0x1E66), +("Scaronsmall",0xF6FD),("Scedilla",0x015E),("Schwa",0x018F),("Schwacyrillic",0x04D8),("Schwadieresiscyrillic",0x04DA),("Scircle",0x24C8), +("Scircumflex",0x015C),("Scommaaccent",0x0218),("Sdotaccent",0x1E60),("Sdotbelow",0x1E62),("Sdotbelowdotaccent",0x1E68),("Seharmenian",0x054D), +("Sevenroman",0x2166),("Shaarmenian",0x0547),("Shacyrillic",0x0428),("Shchacyrillic",0x0429),("Sheicoptic",0x03E2),("Shhacyrillic",0x04BA), +("Shimacoptic",0x03EC),("Sigma",0x03A3),("Sixroman",0x2165),("Smonospace",0xFF33),("Softsigncyrillic",0x042C),("Ssmall",0xF773), +("Stigmagreek",0x03DA),("T",0x0054),("Tau",0x03A4),("Tbar",0x0166),("Tcaron",0x0164),("Tcedilla",0x0162),("Tcircle",0x24C9), +("Tcircumflexbelow",0x1E70),("Tcommaaccent",0x0162),("Tdotaccent",0x1E6A),("Tdotbelow",0x1E6C),("Tecyrillic",0x0422),("Tedescendercyrillic",0x04AC), +("Tenroman",0x2169),("Tetsecyrillic",0x04B4),("Theta",0x0398),("Thook",0x01AC),("Thorn",0x00DE),("Thornsmall",0xF7FE),("Threeroman",0x2162), +("Tildesmall",0xF6FE),("Tiwnarmenian",0x054F),("Tlinebelow",0x1E6E),("Tmonospace",0xFF34),("Toarmenian",0x0539),("Tonefive",0x01BC), +("Tonesix",0x0184),("Tonetwo",0x01A7),("Tretroflexhook",0x01AE),("Tsecyrillic",0x0426),("Tshecyrillic",0x040B),("Tsmall",0xF774), +("Twelveroman",0x216B),("Tworoman",0x2161),("U",0x0055),("Uacute",0x00DA),("Uacutesmall",0xF7FA),("Ubreve",0x016C),("Ucaron",0x01D3), +("Ucircle",0x24CA),("Ucircumflex",0x00DB),("Ucircumflexbelow",0x1E76),("Ucircumflexsmall",0xF7FB),("Ucyrillic",0x0423),("Udblacute",0x0170), +("Udblgrave",0x0214),("Udieresis",0x00DC),("Udieresisacute",0x01D7),("Udieresisbelow",0x1E72),("Udieresiscaron",0x01D9),("Udieresiscyrillic",0x04F0), +("Udieresisgrave",0x01DB),("Udieresismacron",0x01D5),("Udieresissmall",0xF7FC),("Udotbelow",0x1EE4),("Ugrave",0x00D9),("Ugravesmall",0xF7F9), +("Uhookabove",0x1EE6),("Uhorn",0x01AF),("Uhornacute",0x1EE8),("Uhorndotbelow",0x1EF0),("Uhorngrave",0x1EEA),("Uhornhookabove",0x1EEC), +("Uhorntilde",0x1EEE),("Uhungarumlaut",0x0170),("Uhungarumlautcyrillic",0x04F2),("Uinvertedbreve",0x0216),("Ukcyrillic",0x0478),("Umacron",0x016A), +("Umacroncyrillic",0x04EE),("Umacrondieresis",0x1E7A),("Umonospace",0xFF35),("Uogonek",0x0172),("Upsilon",0x03A5),("Upsilon1",0x03D2), +("Upsilonacutehooksymbolgreek",0x03D3),("Upsilonafrican",0x01B1),("Upsilondieresis",0x03AB),("Upsilondieresishooksymbolgreek",0x03D4), +("Upsilonhooksymbol",0x03D2),("Upsilontonos",0x038E),("Uring",0x016E),("Ushortcyrillic",0x040E),("Usmall",0xF775),("Ustraightcyrillic",0x04AE), +("Ustraightstrokecyrillic",0x04B0),("Utilde",0x0168),("Utildeacute",0x1E78),("Utildebelow",0x1E74),("V",0x0056),("Vcircle",0x24CB), +("Vdotbelow",0x1E7E),("Vecyrillic",0x0412),("Vewarmenian",0x054E),("Vhook",0x01B2),("Vmonospace",0xFF36),("Voarmenian",0x0548),("Vsmall",0xF776), +("Vtilde",0x1E7C),("W",0x0057),("Wacute",0x1E82),("Wcircle",0x24CC),("Wcircumflex",0x0174),("Wdieresis",0x1E84),("Wdotaccent",0x1E86), +("Wdotbelow",0x1E88),("Wgrave",0x1E80),("Wmonospace",0xFF37),("Wsmall",0xF777),("X",0x0058),("Xcircle",0x24CD),("Xdieresis",0x1E8C), +("Xdotaccent",0x1E8A),("Xeharmenian",0x053D),("Xi",0x039E),("Xmonospace",0xFF38),("Xsmall",0xF778),("Y",0x0059),("Yacute",0x00DD), +("Yacutesmall",0xF7FD),("Yatcyrillic",0x0462),("Ycircle",0x24CE),("Ycircumflex",0x0176),("Ydieresis",0x0178),("Ydieresissmall",0xF7FF), +("Ydotaccent",0x1E8E),("Ydotbelow",0x1EF4),("Yericyrillic",0x042B),("Yerudieresiscyrillic",0x04F8),("Ygrave",0x1EF2),("Yhook",0x01B3), +("Yhookabove",0x1EF6),("Yiarmenian",0x0545),("Yicyrillic",0x0407),("Yiwnarmenian",0x0552),("Ymonospace",0xFF39),("Ysmall",0xF779),("Ytilde",0x1EF8), +("Yusbigcyrillic",0x046A),("Yusbigiotifiedcyrillic",0x046C),("Yuslittlecyrillic",0x0466),("Yuslittleiotifiedcyrillic",0x0468),("Z",0x005A), +("Zaarmenian",0x0536),("Zacute",0x0179),("Zcaron",0x017D),("Zcaronsmall",0xF6FF),("Zcircle",0x24CF),("Zcircumflex",0x1E90),("Zdot",0x017B), +("Zdotaccent",0x017B),("Zdotbelow",0x1E92),("Zecyrillic",0x0417),("Zedescendercyrillic",0x0498),("Zedieresiscyrillic",0x04DE),("Zeta",0x0396), +("Zhearmenian",0x053A),("Zhebrevecyrillic",0x04C1),("Zhecyrillic",0x0416),("Zhedescendercyrillic",0x0496),("Zhedieresiscyrillic",0x04DC), +("Zlinebelow",0x1E94),("Zmonospace",0xFF3A),("Zsmall",0xF77A),("Zstroke",0x01B5),("a",0x0061),("aabengali",0x0986),("aacute",0x00E1), +("aadeva",0x0906),("aagujarati",0x0A86),("aagurmukhi",0x0A06),("aamatragurmukhi",0x0A3E),("aarusquare",0x3303),("aavowelsignbengali",0x09BE), +("aavowelsigndeva",0x093E),("aavowelsigngujarati",0x0ABE),("abbreviationmarkarmenian",0x055F),("abbreviationsigndeva",0x0970),("abengali",0x0985), +("abopomofo",0x311A),("abreve",0x0103),("abreveacute",0x1EAF),("abrevecyrillic",0x04D1),("abrevedotbelow",0x1EB7),("abrevegrave",0x1EB1), +("abrevehookabove",0x1EB3),("abrevetilde",0x1EB5),("acaron",0x01CE),("acircle",0x24D0),("acircumflex",0x00E2),("acircumflexacute",0x1EA5), +("acircumflexdotbelow",0x1EAD),("acircumflexgrave",0x1EA7),("acircumflexhookabove",0x1EA9),("acircumflextilde",0x1EAB),("acute",0x00B4), +("acutebelowcmb",0x0317),("acutecmb",0x0301),("acutecomb",0x0301),("acutedeva",0x0954),("acutelowmod",0x02CF),("acutetonecmb",0x0341), +("acyrillic",0x0430),("adblgrave",0x0201),("addakgurmukhi",0x0A71),("adeva",0x0905),("adieresis",0x00E4),("adieresiscyrillic",0x04D3), +("adieresismacron",0x01DF),("adotbelow",0x1EA1),("adotmacron",0x01E1),("ae",0x00E6),("aeacute",0x01FD),("aekorean",0x3150),("aemacron",0x01E3), +("afii00208",0x2015),("afii08941",0x20A4),("afii10017",0x0410),("afii10018",0x0411),("afii10019",0x0412),("afii10020",0x0413),("afii10021",0x0414), +("afii10022",0x0415),("afii10023",0x0401),("afii10024",0x0416),("afii10025",0x0417),("afii10026",0x0418),("afii10027",0x0419),("afii10028",0x041A), +("afii10029",0x041B),("afii10030",0x041C),("afii10031",0x041D),("afii10032",0x041E),("afii10033",0x041F),("afii10034",0x0420),("afii10035",0x0421), +("afii10036",0x0422),("afii10037",0x0423),("afii10038",0x0424),("afii10039",0x0425),("afii10040",0x0426),("afii10041",0x0427),("afii10042",0x0428), +("afii10043",0x0429),("afii10044",0x042A),("afii10045",0x042B),("afii10046",0x042C),("afii10047",0x042D),("afii10048",0x042E),("afii10049",0x042F), +("afii10050",0x0490),("afii10051",0x0402),("afii10052",0x0403),("afii10053",0x0404),("afii10054",0x0405),("afii10055",0x0406),("afii10056",0x0407), +("afii10057",0x0408),("afii10058",0x0409),("afii10059",0x040A),("afii10060",0x040B),("afii10061",0x040C),("afii10062",0x040E),("afii10063",0xF6C4), +("afii10064",0xF6C5),("afii10065",0x0430),("afii10066",0x0431),("afii10067",0x0432),("afii10068",0x0433),("afii10069",0x0434),("afii10070",0x0435), +("afii10071",0x0451),("afii10072",0x0436),("afii10073",0x0437),("afii10074",0x0438),("afii10075",0x0439),("afii10076",0x043A),("afii10077",0x043B), +("afii10078",0x043C),("afii10079",0x043D),("afii10080",0x043E),("afii10081",0x043F),("afii10082",0x0440),("afii10083",0x0441),("afii10084",0x0442), +("afii10085",0x0443),("afii10086",0x0444),("afii10087",0x0445),("afii10088",0x0446),("afii10089",0x0447),("afii10090",0x0448),("afii10091",0x0449), +("afii10092",0x044A),("afii10093",0x044B),("afii10094",0x044C),("afii10095",0x044D),("afii10096",0x044E),("afii10097",0x044F),("afii10098",0x0491), +("afii10099",0x0452),("afii10100",0x0453),("afii10101",0x0454),("afii10102",0x0455),("afii10103",0x0456),("afii10104",0x0457),("afii10105",0x0458), +("afii10106",0x0459),("afii10107",0x045A),("afii10108",0x045B),("afii10109",0x045C),("afii10110",0x045E),("afii10145",0x040F),("afii10146",0x0462), +("afii10147",0x0472),("afii10148",0x0474),("afii10192",0xF6C6),("afii10193",0x045F),("afii10194",0x0463),("afii10195",0x0473),("afii10196",0x0475), +("afii10831",0xF6C7),("afii10832",0xF6C8),("afii10846",0x04D9),("afii299",0x200E),("afii300",0x200F),("afii301",0x200D),("afii57381",0x066A), +("afii57388",0x060C),("afii57392",0x0660),("afii57393",0x0661),("afii57394",0x0662),("afii57395",0x0663),("afii57396",0x0664),("afii57397",0x0665), +("afii57398",0x0666),("afii57399",0x0667),("afii57400",0x0668),("afii57401",0x0669),("afii57403",0x061B),("afii57407",0x061F),("afii57409",0x0621), +("afii57410",0x0622),("afii57411",0x0623),("afii57412",0x0624),("afii57413",0x0625),("afii57414",0x0626),("afii57415",0x0627),("afii57416",0x0628), +("afii57417",0x0629),("afii57418",0x062A),("afii57419",0x062B),("afii57420",0x062C),("afii57421",0x062D),("afii57422",0x062E),("afii57423",0x062F), +("afii57424",0x0630),("afii57425",0x0631),("afii57426",0x0632),("afii57427",0x0633),("afii57428",0x0634),("afii57429",0x0635),("afii57430",0x0636), +("afii57431",0x0637),("afii57432",0x0638),("afii57433",0x0639),("afii57434",0x063A),("afii57440",0x0640),("afii57441",0x0641),("afii57442",0x0642), +("afii57443",0x0643),("afii57444",0x0644),("afii57445",0x0645),("afii57446",0x0646),("afii57448",0x0648),("afii57449",0x0649),("afii57450",0x064A), +("afii57451",0x064B),("afii57452",0x064C),("afii57453",0x064D),("afii57454",0x064E),("afii57455",0x064F),("afii57456",0x0650),("afii57457",0x0651), +("afii57458",0x0652),("afii57470",0x0647),("afii57505",0x06A4),("afii57506",0x067E),("afii57507",0x0686),("afii57508",0x0698),("afii57509",0x06AF), +("afii57511",0x0679),("afii57512",0x0688),("afii57513",0x0691),("afii57514",0x06BA),("afii57519",0x06D2),("afii57534",0x06D5),("afii57636",0x20AA), +("afii57645",0x05BE),("afii57658",0x05C3),("afii57664",0x05D0),("afii57665",0x05D1),("afii57666",0x05D2),("afii57667",0x05D3),("afii57668",0x05D4), +("afii57669",0x05D5),("afii57670",0x05D6),("afii57671",0x05D7),("afii57672",0x05D8),("afii57673",0x05D9),("afii57674",0x05DA),("afii57675",0x05DB), +("afii57676",0x05DC),("afii57677",0x05DD),("afii57678",0x05DE),("afii57679",0x05DF),("afii57680",0x05E0),("afii57681",0x05E1),("afii57682",0x05E2), +("afii57683",0x05E3),("afii57684",0x05E4),("afii57685",0x05E5),("afii57686",0x05E6),("afii57687",0x05E7),("afii57688",0x05E8),("afii57689",0x05E9), +("afii57690",0x05EA),("afii57694",0xFB2A),("afii57695",0xFB2B),("afii57700",0xFB4B),("afii57705",0xFB1F),("afii57716",0x05F0),("afii57717",0x05F1), +("afii57718",0x05F2),("afii57723",0xFB35),("afii57793",0x05B4),("afii57794",0x05B5),("afii57795",0x05B6),("afii57796",0x05BB),("afii57797",0x05B8), +("afii57798",0x05B7),("afii57799",0x05B0),("afii57800",0x05B2),("afii57801",0x05B1),("afii57802",0x05B3),("afii57803",0x05C2),("afii57804",0x05C1), +("afii57806",0x05B9),("afii57807",0x05BC),("afii57839",0x05BD),("afii57841",0x05BF),("afii57842",0x05C0),("afii57929",0x02BC),("afii61248",0x2105), +("afii61289",0x2113),("afii61352",0x2116),("afii61573",0x202C),("afii61574",0x202D),("afii61575",0x202E),("afii61664",0x200C),("afii63167",0x066D), +("afii64937",0x02BD),("agrave",0x00E0),("agujarati",0x0A85),("agurmukhi",0x0A05),("ahiragana",0x3042),("ahookabove",0x1EA3),("aibengali",0x0990), +("aibopomofo",0x311E),("aideva",0x0910),("aiecyrillic",0x04D5),("aigujarati",0x0A90),("aigurmukhi",0x0A10),("aimatragurmukhi",0x0A48), +("ainarabic",0x0639),("ainfinalarabic",0xFECA),("aininitialarabic",0xFECB),("ainmedialarabic",0xFECC),("ainvertedbreve",0x0203), +("aivowelsignbengali",0x09C8),("aivowelsigndeva",0x0948),("aivowelsigngujarati",0x0AC8),("akatakana",0x30A2),("akatakanahalfwidth",0xFF71), +("akorean",0x314F),("alef",0x05D0),("alefarabic",0x0627),("alefdageshhebrew",0xFB30),("aleffinalarabic",0xFE8E),("alefhamzaabovearabic",0x0623), +("alefhamzaabovefinalarabic",0xFE84),("alefhamzabelowarabic",0x0625),("alefhamzabelowfinalarabic",0xFE88),("alefhebrew",0x05D0), +("aleflamedhebrew",0xFB4F),("alefmaddaabovearabic",0x0622),("alefmaddaabovefinalarabic",0xFE82),("alefmaksuraarabic",0x0649), +("alefmaksurafinalarabic",0xFEF0),("alefmaksurainitialarabic",0xFEF3),("alefmaksuramedialarabic",0xFEF4),("alefpatahhebrew",0xFB2E), +("alefqamatshebrew",0xFB2F),("aleph",0x2135),("allequal",0x224C),("alpha",0x03B1),("alphatonos",0x03AC),("altselector",0xD802),("amacron",0x0101), +("amonospace",0xFF41),("ampersand",0x0026),("ampersandmonospace",0xFF06),("ampersandsmall",0xF726),("amsquare",0x33C2),("anbopomofo",0x3122), +("angbopomofo",0x3124),("angbracketleft",0x27E8),("angbracketright",0x27E9),("angkhankhuthai",0x0E5A),("angle",0x2220),("anglebracketleft",0x3008), +("anglebracketleftvertical",0xFE3F),("anglebracketright",0x3009),("anglebracketrightvertical",0xFE40),("angleleft",0x2329),("angleleftBig",0x2329), +("angleleftBigg",0x2329),("angleleftbig",0x2329),("angleleftbigg",0x2329),("angleright",0x232A),("anglerightBig",0x232A),("anglerightBigg",0x232A), +("anglerightbig",0x232A),("anglerightbigg",0x232A),("angstrom",0x212B),("anoteleia",0x0387),("anudattadeva",0x0952),("anusvarabengali",0x0982), +("anusvaradeva",0x0902),("anusvaragujarati",0x0A82),("aogonek",0x0105),("apaatosquare",0x3300),("aparen",0x249C),("apostrophearmenian",0x055A), +("apostrophemod",0x02BC),("apple",0xF8FF),("approaches",0x2250),("approxequal",0x2248),("approxequalorimage",0x2252),("approximatelyequal",0x2245), +("araeaekorean",0x318E),("araeakorean",0x318D),("arc",0x2312),("arighthalfring",0x1E9A),("aring",0x00E5),("aringacute",0x01FB),("aringbelow",0x1E01), +("arrowboth",0x2194),("arrowbothv",0x2195),("arrowbt",0x2193),("arrowdashdown",0x21E3),("arrowdashleft",0x21E0),("arrowdashright",0x21E2), +("arrowdashup",0x21E1),("arrowdblboth",0x21D4),("arrowdblbothv",0x21D5),("arrowdbldown",0x21D3),("arrowdblleft",0x21D0),("arrowdblright",0x21D2), +("arrowdbltp",0x21D1),("arrowdblup",0x21D1),("arrowdblvertex",0x21D5),("arrowdown",0x2193),("arrowdownleft",0x2199),("arrowdownright",0x2198), +("arrowdownwhite",0x21E9),("arrowheaddownmod",0x02C5),("arrowheadleftmod",0x02C2),("arrowheadrightmod",0x02C3),("arrowheadupmod",0x02C4), +("arrowhorizex",0xF8E7),("arrowleft",0x2190),("arrowleftbothalf",0x21BD),("arrowleftdbl",0x21D0),("arrowleftdblstroke",0x21CD), +("arrowleftoverright",0x21C6),("arrowlefttophalf",0x21BC),("arrowleftwhite",0x21E6),("arrownortheast",0x2197),("arrownorthwest",0x2196), +("arrowright",0x2192),("arrowrightbothalf",0x21C1),("arrowrightdblstroke",0x21CF),("arrowrightheavy",0x279E),("arrowrightoverleft",0x21C4), +("arrowrighttophalf",0x21C0),("arrowrightwhite",0x21E8),("arrowsoutheast",0x2198),("arrowsouthwest",0x2199),("arrowtableft",0x21E4), +("arrowtabright",0x21E5),("arrowtp",0x2191),("arrowup",0x2191),("arrowupdn",0x2195),("arrowupdnbse",0x21A8),("arrowupdownbase",0x21A8), +("arrowupleft",0x2196),("arrowupleftofdown",0x21C5),("arrowupright",0x2197),("arrowupwhite",0x21E7),("arrowvertex",0x2195),("arrowvertex2",0xF8E6), +("ascendercompwordmark",0xD80A),("asciicircum",0x005E),("asciicircummonospace",0xFF3E),("asciitilde",0x007E),("asciitildemonospace",0xFF5E), +("ascript",0x0251),("ascriptturned",0x0252),("asmallhiragana",0x3041),("asmallkatakana",0x30A1),("asmallkatakanahalfwidth",0xFF67), +("asterisk",0x002A),("asteriskaltonearabic",0x066D),("asteriskarabic",0x066D),("asteriskcentered",0x2217),("asteriskmath",0x2217), +("asteriskmonospace",0xFF0A),("asterisksmall",0xFE61),("asterism",0x2042),("asuperior",0xF6E9),("asymptoticallyequal",0x2243),("at",0x0040), +("atilde",0x00E3),("atmonospace",0xFF20),("atsmall",0xFE6B),("aturned",0x0250),("aubengali",0x0994),("aubopomofo",0x3120),("audeva",0x0914), +("augujarati",0x0A94),("augurmukhi",0x0A14),("aulengthmarkbengali",0x09D7),("aumatragurmukhi",0x0A4C),("auvowelsignbengali",0x09CC), +("auvowelsigndeva",0x094C),("auvowelsigngujarati",0x0ACC),("avagrahadeva",0x093D),("aybarmenian",0x0561),("ayin",0x05E2),("ayinaltonehebrew",0xFB20), +("ayinhebrew",0x05E2),("b",0x0062),("babengali",0x09AC),("backslash",0x005C),("backslashBig",0x005C),("backslashBigg",0x005C),("backslashbig",0x005C), +("backslashbigg",0x005C),("backslashmonospace",0xFF3C),("badeva",0x092C),("bagujarati",0x0AAC),("bagurmukhi",0x0A2C),("bahiragana",0x3070), +("bahtthai",0x0E3F),("bakatakana",0x30D0),("bar",0x007C),("bardbl",0x2225),("bardblex",0x2016),("barex",0x007C),("barmonospace",0xFF5C), +("bbopomofo",0x3105),("bcircle",0x24D1),("bdotaccent",0x1E03),("bdotbelow",0x1E05),("beamedsixteenthnotes",0x266C),("because",0x2235), +("becyrillic",0x0431),("beharabic",0x0628),("behfinalarabic",0xFE90),("behinitialarabic",0xFE91),("behiragana",0x3079),("behmedialarabic",0xFE92), +("behmeeminitialarabic",0xFC9F),("behmeemisolatedarabic",0xFC08),("behnoonfinalarabic",0xFC6D),("bekatakana",0x30D9),("benarmenian",0x0562), +("bet",0x05D1),("beta",0x03B2),("betasymbolgreek",0x03D0),("betdagesh",0xFB31),("betdageshhebrew",0xFB31),("bethebrew",0x05D1), +("betrafehebrew",0xFB4C),("bhabengali",0x09AD),("bhadeva",0x092D),("bhagujarati",0x0AAD),("bhagurmukhi",0x0A2D),("bhook",0x0253), +("bihiragana",0x3073),("bikatakana",0x30D3),("bilabialclick",0x0298),("bindigurmukhi",0x0A02),("birusquare",0x3331),("blackcircle",0x25CF), +("blackdiamond",0x25C6),("blackdownpointingtriangle",0x25BC),("blackleftpointingpointer",0x25C4),("blackleftpointingtriangle",0x25C0), +("blacklenticularbracketleft",0x3010),("blacklenticularbracketleftvertical",0xFE3B),("blacklenticularbracketright",0x3011), +("blacklenticularbracketrightvertical",0xFE3C),("blacklowerlefttriangle",0x25E3),("blacklowerrighttriangle",0x25E2),("blackrectangle",0x25AC), +("blackrightpointingpointer",0x25BA),("blackrightpointingtriangle",0x25B6),("blacksmallsquare",0x25AA),("blacksmilingface",0x263B), +("blacksquare",0x25A0),("blackstar",0x2605),("blackupperlefttriangle",0x25E4),("blackupperrighttriangle",0x25E5), +("blackuppointingsmalltriangle",0x25B4),("blackuppointingtriangle",0x25B2),("blank",0x2423),("blinebelow",0x1E07),("block",0x2588), +("bmonospace",0xFF42),("bobaimaithai",0x0E1A),("bohiragana",0x307C),("bokatakana",0x30DC),("bparen",0x249D),("bqsquare",0x33C3),("braceex",0x007C), +("braceex2",0xF8F4),("braceleft",0x007B),("braceleftBig",0x007B),("braceleftBigg",0x007B),("braceleftbig",0x007B),("braceleftbigg",0x007B), +("braceleftbt",0xF8F3),("braceleftmid",0x007C),("braceleftmid2",0xF8F2),("braceleftmonospace",0xFF5B),("braceleftsmall",0xFE5B), +("bracelefttp",0xF8F1),("braceleftvertical",0xFE37),("braceright",0x007D),("bracerightBig",0x007D),("bracerightBigg",0x007D),("bracerightbig",0x007D), +("bracerightbigg",0x007D),("bracerightbt",0xF8FE),("bracerightmid",0x2016),("bracerightmid2",0xF8FD),("bracerightmonospace",0xFF5D), +("bracerightsmall",0xFE5C),("bracerighttp",0xF8FC),("bracerightvertical",0xFE38),("bracketleft",0x005B),("bracketleftBig",0x005B), +("bracketleftBigg",0x005B),("bracketleftbig",0x005B),("bracketleftbigg",0x005B),("bracketleftbt",0xF8F0),("bracketleftex",0xF8EF), +("bracketleftmonospace",0xFF3B),("bracketlefttp",0xF8EE),("bracketright",0x005D),("bracketrightBig",0x005D),("bracketrightBigg",0x005D), +("bracketrightbig",0x005D),("bracketrightbigg",0x005D),("bracketrightbt",0xF8FB),("bracketrightex",0xF8FA),("bracketrightmonospace",0xFF3D), +("bracketrighttp",0xF8F9),("breve",0x02D8),("brevebelowcmb",0x032E),("brevecmb",0x0306),("breveinvertedbelowcmb",0x032F),("breveinvertedcmb",0x0311), +("breveinverteddoublecmb",0x0361),("bridgebelowcmb",0x032A),("bridgeinvertedbelowcmb",0x033A),("brokenbar",0x00A6),("bstroke",0x0180), +("bsuperior",0xF6EA),("btopbar",0x0183),("buhiragana",0x3076),("bukatakana",0x30D6),("bullet",0x2022),("bulletinverse",0x25D8), +("bulletoperator",0x2219),("bullseye",0x25CE),("c",0x0063),("caarmenian",0x056E),("cabengali",0x099A),("cacute",0x0107),("cadeva",0x091A), +("cagujarati",0x0A9A),("cagurmukhi",0x0A1A),("calsquare",0x3388),("candrabindubengali",0x0981),("candrabinducmb",0x0310),("candrabindudeva",0x0901), +("candrabindugujarati",0x0A81),("capitalcompwordmark",0xD809),("capslock",0x21EA),("careof",0x2105),("caron",0x02C7),("caronbelowcmb",0x032C), +("caroncmb",0x030C),("carriagereturn",0x21B5),("cbopomofo",0x3118),("ccaron",0x010D),("ccedilla",0x00E7),("ccedillaacute",0x1E09),("ccircle",0x24D2), +("ccircumflex",0x0109),("ccurl",0x0255),("cdot",0x010B),("cdotaccent",0x010B),("cdsquare",0x33C5),("cedilla",0x00B8),("cedillacmb",0x0327), +("ceilingleft",0x2308),("ceilingleftBig",0x2308),("ceilingleftBigg",0x2308),("ceilingleftbig",0x2308),("ceilingleftbigg",0x2308), +("ceilingright",0x2309),("ceilingrightBig",0x2309),("ceilingrightBigg",0x2309),("ceilingrightbig",0x2309),("ceilingrightbigg",0x2309),("cent",0x00A2), +("centigrade",0x2103),("centinferior",0xF6DF),("centmonospace",0xFFE0),("centoldstyle",0xF7A2),("centsuperior",0xF6E0),("chaarmenian",0x0579), +("chabengali",0x099B),("chadeva",0x091B),("chagujarati",0x0A9B),("chagurmukhi",0x0A1B),("chbopomofo",0x3114),("cheabkhasiancyrillic",0x04BD), +("checkmark",0x2713),("checyrillic",0x0447),("chedescenderabkhasiancyrillic",0x04BF),("chedescendercyrillic",0x04B7),("chedieresiscyrillic",0x04F5), +("cheharmenian",0x0573),("chekhakassiancyrillic",0x04CC),("cheverticalstrokecyrillic",0x04B9),("chi",0x03C7),("chieuchacirclekorean",0x3277), +("chieuchaparenkorean",0x3217),("chieuchcirclekorean",0x3269),("chieuchkorean",0x314A),("chieuchparenkorean",0x3209),("chochangthai",0x0E0A), +("chochanthai",0x0E08),("chochingthai",0x0E09),("chochoethai",0x0E0C),("chook",0x0188),("cieucacirclekorean",0x3276),("cieucaparenkorean",0x3216), +("cieuccirclekorean",0x3268),("cieuckorean",0x3148),("cieucparenkorean",0x3208),("cieucuparenkorean",0x321C),("circle",0x25CB), +("circlecopyrt",0x20DD),("circledivide",0x2298),("circledot",0x2299),("circledotdisplay",0x2299),("circledottext",0x2299),("circleminus",0x2296), +("circlemultiply",0x2297),("circlemultiplydisplay",0x2297),("circlemultiplytext",0x2297),("circleot",0x2299),("circleplus",0x2295), +("circleplusdisplay",0x2295),("circleplustext",0x2295),("circlepostalmark",0x3036),("circlewithlefthalfblack",0x25D0), +("circlewithrighthalfblack",0x25D1),("circumflex",0x02C6),("circumflexbelowcmb",0x032D),("circumflexcmb",0x0302),("clear",0x2327), +("clickalveolar",0x01C2),("clickdental",0x01C0),("clicklateral",0x01C1),("clickretroflex",0x01C3),("club",0x2663),("clubsuitblack",0x2663), +("clubsuitwhite",0x2667),("cmcubedsquare",0x33A4),("cmonospace",0xFF43),("cmsquaredsquare",0x33A0),("coarmenian",0x0581),("colon",0x003A), +("colonmonetary",0x20A1),("colonmonospace",0xFF1A),("colonsign",0x20A1),("colonsmall",0xFE55),("colontriangularhalfmod",0x02D1), +("colontriangularmod",0x02D0),("comma",0x002C),("commaabovecmb",0x0313),("commaaboverightcmb",0x0315),("commaaccent",0xF6C3),("commaarabic",0x060C), +("commaarmenian",0x055D),("commainferior",0xF6E1),("commamonospace",0xFF0C),("commareversedabovecmb",0x0314),("commareversedmod",0x02BD), +("commasmall",0xFE50),("commasuperior",0xF6E2),("commaturnedabovecmb",0x0312),("commaturnedmod",0x02BB),("compass",0x263C),("compwordmark",0x200C), +("congruent",0x2245),("contintegraldisplay",0x222E),("contintegraltext",0x222E),("contourintegral",0x222E),("control",0x2303),("controlACK",0x0006), +("controlBEL",0x0007),("controlBS",0x0008),("controlCAN",0x0018),("controlCR",0x000D),("controlDC1",0x0011),("controlDC2",0x0012), +("controlDC3",0x0013),("controlDC4",0x0014),("controlDEL",0x007F),("controlDLE",0x0010),("controlEM",0x0019),("controlENQ",0x0005), +("controlEOT",0x0004),("controlESC",0x001B),("controlETB",0x0017),("controlETX",0x0003),("controlFF",0x000C),("controlFS",0x001C), +("controlGS",0x001D),("controlHT",0x0009),("controlLF",0x000A),("controlNAK",0x0015),("controlRS",0x001E),("controlSI",0x000F),("controlSO",0x000E), +("controlSOT",0x0002),("controlSTX",0x0001),("controlSUB",0x001A),("controlSYN",0x0016),("controlUS",0x001F),("controlVT",0x000B), +("coproduct",0x2A3F),("coproductdisplay",0x2210),("coproducttext",0x2210),("copyright",0x00A9),("copyrightsans",0xF8E9),("copyrightserif",0xF6D9), +("cornerbracketleft",0x300C),("cornerbracketlefthalfwidth",0xFF62),("cornerbracketleftvertical",0xFE41),("cornerbracketright",0x300D), +("cornerbracketrighthalfwidth",0xFF63),("cornerbracketrightvertical",0xFE42),("corporationsquare",0x337F),("cosquare",0x33C7), +("coverkgsquare",0x33C6),("cparen",0x249E),("cruzeiro",0x20A2),("cstretched",0x0297),("curlyand",0x22CF),("curlyor",0x22CE),("currency",0x00A4), +("cwm",0x200C),("cyrBreve",0xF6D1),("cyrFlex",0xF6D2),("cyrbreve",0xF6D4),("cyrflex",0xF6D5),("d",0x0064),("daarmenian",0x0564),("dabengali",0x09A6), +("dadarabic",0x0636),("dadeva",0x0926),("dadfinalarabic",0xFEBE),("dadinitialarabic",0xFEBF),("dadmedialarabic",0xFEC0),("dagesh",0x05BC), +("dageshhebrew",0x05BC),("dagger",0x2020),("daggerdbl",0x2021),("dagujarati",0x0AA6),("dagurmukhi",0x0A26),("dahiragana",0x3060), +("dakatakana",0x30C0),("dalarabic",0x062F),("dalet",0x05D3),("daletdagesh",0xFB33),("daletdageshhebrew",0xFB33),("dalethatafpatah",0x05D3), +("dalethatafpatahhebrew",0x05D3),("dalethatafsegol",0x05D3),("dalethatafsegolhebrew",0x05D3),("dalethebrew",0x05D3),("dalethiriq",0x05D3), +("dalethiriqhebrew",0x05D3),("daletholam",0x05D3),("daletholamhebrew",0x05D3),("daletpatah",0x05D3),("daletpatahhebrew",0x05D3), +("daletqamats",0x05D3),("daletqamatshebrew",0x05D3),("daletqubuts",0x05D3),("daletqubutshebrew",0x05D3),("daletsegol",0x05D3), +("daletsegolhebrew",0x05D3),("daletsheva",0x05D3),("daletshevahebrew",0x05D3),("dalettsere",0x05D3),("dalettserehebrew",0x05D3), +("dalfinalarabic",0xFEAA),("dammaarabic",0x064F),("dammalowarabic",0x064F),("dammatanaltonearabic",0x064C),("dammatanarabic",0x064C),("danda",0x0964), +("dargahebrew",0x05A7),("dargalefthebrew",0x05A7),("dasiapneumatacyrilliccmb",0x0485),("dbar",0x0111),("dblGrave",0xF6D3), +("dblanglebracketleft",0x300A),("dblanglebracketleftvertical",0xFE3D),("dblanglebracketright",0x300B),("dblanglebracketrightvertical",0xFE3E), +("dblarchinvertedbelowcmb",0x032B),("dblarrowleft",0x21D4),("dblarrowright",0x21D2),("dblbracketleft",0x27E6),("dblbracketright",0x27E7), +("dbldanda",0x0965),("dblgrave",0xF6D6),("dblgravecmb",0x030F),("dblintegral",0x222C),("dbllowline",0x2017),("dbllowlinecmb",0x0333), +("dbloverlinecmb",0x033F),("dblprimemod",0x02BA),("dblverticalbar",0x2016),("dblverticallineabovecmb",0x030E),("dbopomofo",0x3109), +("dbsquare",0x33C8),("dcaron",0x010F),("dcedilla",0x1E11),("dcircle",0x24D3),("dcircumflexbelow",0x1E13),("dcroat",0x0111),("ddabengali",0x09A1), +("ddadeva",0x0921),("ddagujarati",0x0AA1),("ddagurmukhi",0x0A21),("ddalarabic",0x0688),("ddalfinalarabic",0xFB89),("dddhadeva",0x095C), +("ddhabengali",0x09A2),("ddhadeva",0x0922),("ddhagujarati",0x0AA2),("ddhagurmukhi",0x0A22),("ddotaccent",0x1E0B),("ddotbelow",0x1E0D), +("decimalseparatorarabic",0x066B),("decimalseparatorpersian",0x066B),("decyrillic",0x0434),("degree",0x00B0),("dehihebrew",0x05AD), +("dehiragana",0x3067),("deicoptic",0x03EF),("dekatakana",0x30C7),("deleteleft",0x232B),("deleteright",0x2326),("delta",0x03B4),("deltaturned",0x018D), +("denominatorminusonenumeratorbengali",0x09F8),("dezh",0x02A4),("dhabengali",0x09A7),("dhadeva",0x0927),("dhagujarati",0x0AA7),("dhagurmukhi",0x0A27), +("dhook",0x0257),("dialytikatonos",0x0385),("dialytikatonoscmb",0x0344),("diamond",0x2662),("diamond2",0x2666),("diamondmath",0x22C4), +("diamondsuitwhite",0x2662),("dieresis",0x00A8),("dieresisacute",0xF6D7),("dieresisbelowcmb",0x0324),("dieresiscmb",0x0308),("dieresisgrave",0xF6D8), +("dieresistonos",0x0385),("dihiragana",0x3062),("dikatakana",0x30C2),("dittomark",0x3003),("divide",0x00F7),("divides",0x2223), +("divisionslash",0x2215),("djecyrillic",0x0452),("dkshade",0x2593),("dlinebelow",0x1E0F),("dlsquare",0x3397),("dmacron",0x0111),("dmonospace",0xFF44), +("dnblock",0x2584),("dochadathai",0x0E0E),("dodekthai",0x0E14),("dohiragana",0x3069),("dokatakana",0x30C9),("dollar",0x0024), +("dollarinferior",0xF6E3),("dollarmonospace",0xFF04),("dollaroldstyle",0xF724),("dollarsmall",0xFE69),("dollarsuperior",0xF6E4),("dong",0x20AB), +("dorusquare",0x3326),("dotaccent",0x02D9),("dotaccentcmb",0x0307),("dotbelowcmb",0x0323),("dotbelowcomb",0x0323),("dotkatakana",0x30FB), +("dotlessi",0x0131),("dotlessj",0x0237),("dotlessj2",0xF6BE),("dotlessjstrokehook",0x0284),("dotmath",0x22C5),("dottedcircle",0x25CC), +("doubleyodpatah",0xFB1F),("doubleyodpatahhebrew",0xFB1F),("downtackbelowcmb",0x031E),("downtackmod",0x02D5),("dparen",0x249F),("dsuperior",0xF6EB), +("dtail",0x0256),("dtopbar",0x018C),("duhiragana",0x3065),("dukatakana",0x30C5),("dz",0x01F3),("dzaltone",0x02A3),("dzcaron",0x01C6), +("dzcurl",0x02A5),("dzeabkhasiancyrillic",0x04E1),("dzecyrillic",0x0455),("dzhecyrillic",0x045F),("e",0x0065),("eacute",0x00E9),("earth",0x2641), +("ebengali",0x098F),("ebopomofo",0x311C),("ebreve",0x0115),("ecandradeva",0x090D),("ecandragujarati",0x0A8D),("ecandravowelsigndeva",0x0945), +("ecandravowelsigngujarati",0x0AC5),("ecaron",0x011B),("ecedillabreve",0x1E1D),("echarmenian",0x0565),("echyiwnarmenian",0x0587),("ecircle",0x24D4), +("ecircumflex",0x00EA),("ecircumflexacute",0x1EBF),("ecircumflexbelow",0x1E19),("ecircumflexdotbelow",0x1EC7),("ecircumflexgrave",0x1EC1), +("ecircumflexhookabove",0x1EC3),("ecircumflextilde",0x1EC5),("ecyrillic",0x0454),("edblgrave",0x0205),("edeva",0x090F),("edieresis",0x00EB), +("edot",0x0117),("edotaccent",0x0117),("edotbelow",0x1EB9),("eegurmukhi",0x0A0F),("eematragurmukhi",0x0A47),("efcyrillic",0x0444),("egrave",0x00E8), +("egujarati",0x0A8F),("eharmenian",0x0567),("ehbopomofo",0x311D),("ehiragana",0x3048),("ehookabove",0x1EBB),("eibopomofo",0x311F),("eight",0x0038), +("eightarabic",0x0668),("eightbengali",0x09EE),("eightcircle",0x2467),("eightcircleinversesansserif",0x2791),("eightdeva",0x096E), +("eighteencircle",0x2471),("eighteenparen",0x2485),("eighteenperiod",0x2499),("eightgujarati",0x0AEE),("eightgurmukhi",0x0A6E), +("eighthackarabic",0x0668),("eighthangzhou",0x3028),("eighthnotebeamed",0x266B),("eightideographicparen",0x3227),("eightinferior",0x2088), +("eightmonospace",0xFF18),("eightoldstyle",0xF738),("eightparen",0x247B),("eightperiod",0x248F),("eightpersian",0x06F8),("eightroman",0x2177), +("eightsuperior",0x2078),("eightthai",0x0E58),("einvertedbreve",0x0207),("eiotifiedcyrillic",0x0465),("ekatakana",0x30A8), +("ekatakanahalfwidth",0xFF74),("ekonkargurmukhi",0x0A74),("ekorean",0x3154),("elcyrillic",0x043B),("element",0x2208),("elevencircle",0x246A), +("elevenparen",0x247E),("elevenperiod",0x2492),("elevenroman",0x217A),("ellipsis",0x2026),("ellipsisvertical",0x22EE),("emacron",0x0113), +("emacronacute",0x1E17),("emacrongrave",0x1E15),("emcyrillic",0x043C),("emdash",0x2014),("emdashvertical",0xFE31),("emonospace",0xFF45), +("emphasismarkarmenian",0x055B),("emptyset",0x2205),("emptyslot",0xD801),("enbopomofo",0x3123),("encyrillic",0x043D),("endash",0x2013), +("endashvertical",0xFE32),("endescendercyrillic",0x04A3),("eng",0x014B),("engbopomofo",0x3125),("enghecyrillic",0x04A5),("enhookcyrillic",0x04C8), +("enspace",0x2002),("eogonek",0x0119),("eokorean",0x3153),("eopen",0x025B),("eopenclosed",0x029A),("eopenreversed",0x025C), +("eopenreversedclosed",0x025E),("eopenreversedhook",0x025D),("eparen",0x24A0),("epsilon",0x03B5),("epsilon1",0x03F5),("epsilontonos",0x03AD), +("equal",0x003D),("equalmonospace",0xFF1D),("equalsmall",0xFE66),("equalsuperior",0x207C),("equivalence",0x2261),("equivasymptotic",0x224D), +("erbopomofo",0x3126),("ercyrillic",0x0440),("ereversed",0x0258),("ereversedcyrillic",0x044D),("escyrillic",0x0441),("esdescendercyrillic",0x04AB), +("esh",0x0283),("eshcurl",0x0286),("eshortdeva",0x090E),("eshortvowelsigndeva",0x0946),("eshreversedloop",0x01AA),("eshsquatreversed",0x0285), +("esmallhiragana",0x3047),("esmallkatakana",0x30A7),("esmallkatakanahalfwidth",0xFF6A),("estimated",0x212E),("esuperior",0xF6EC),("eta",0x03B7), +("etarmenian",0x0568),("etatonos",0x03AE),("eth",0x00F0),("etilde",0x1EBD),("etildebelow",0x1E1B),("etnahtafoukhhebrew",0x0591), +("etnahtafoukhlefthebrew",0x0591),("etnahtahebrew",0x0591),("etnahtalefthebrew",0x0591),("eturned",0x01DD),("eukorean",0x3161),("euro",0x20AC), +("evowelsignbengali",0x09C7),("evowelsigndeva",0x0947),("evowelsigngujarati",0x0AC7),("exclam",0x0021),("exclamarmenian",0x055C),("exclamdbl",0x203C), +("exclamdown",0x00A1),("exclamdownsmall",0xF7A1),("exclammonospace",0xFF01),("exclamsmall",0xF721),("existential",0x2203),("ezh",0x0292), +("ezhcaron",0x01EF),("ezhcurl",0x0293),("ezhreversed",0x01B9),("ezhtail",0x01BA),("f",0x0066),("f_f",0xFB00),("f_f_i",0xFB03),("f_f_l",0xFB04), +("f_i",0xFB01),("f_l",0xFB02),("fadeva",0x095E),("fagurmukhi",0x0A5E),("fahrenheit",0x2109),("fathaarabic",0x064E),("fathalowarabic",0x064E), +("fathatanarabic",0x064B),("fbopomofo",0x3108),("fcircle",0x24D5),("fdotaccent",0x1E1F),("feharabic",0x0641),("feharmenian",0x0586), +("fehfinalarabic",0xFED2),("fehinitialarabic",0xFED3),("fehmedialarabic",0xFED4),("feicoptic",0x03E5),("female",0x2640),("ff",0xFB00),("ffi",0xFB03), +("ffl",0xFB04),("fi",0xFB01),("fifteencircle",0x246E),("fifteenparen",0x2482),("fifteenperiod",0x2496),("figuredash",0x2012),("filledbox",0x25A0), +("filledrect",0x25AC),("finalkaf",0x05DA),("finalkafdagesh",0xFB3A),("finalkafdageshhebrew",0xFB3A),("finalkafhebrew",0x05DA), +("finalkafqamats",0x05DA),("finalkafqamatshebrew",0x05DA),("finalkafsheva",0x05DA),("finalkafshevahebrew",0x05DA),("finalmem",0x05DD), +("finalmemhebrew",0x05DD),("finalnun",0x05DF),("finalnunhebrew",0x05DF),("finalpe",0x05E3),("finalpehebrew",0x05E3),("finaltsadi",0x05E5), +("finaltsadihebrew",0x05E5),("firsttonechinese",0x02C9),("fisheye",0x25C9),("fitacyrillic",0x0473),("five",0x0035),("fivearabic",0x0665), +("fivebengali",0x09EB),("fivecircle",0x2464),("fivecircleinversesansserif",0x278E),("fivedeva",0x096B),("fiveeighths",0x215D),("fivegujarati",0x0AEB), +("fivegurmukhi",0x0A6B),("fivehackarabic",0x0665),("fivehangzhou",0x3025),("fiveideographicparen",0x3224),("fiveinferior",0x2085), +("fivemonospace",0xFF15),("fiveoldstyle",0xF735),("fiveparen",0x2478),("fiveperiod",0x248C),("fivepersian",0x06F5),("fiveroman",0x2174), +("fivesuperior",0x2075),("fivethai",0x0E55),("fl",0xFB02),("flat",0x266D),("floorleft",0x230A),("floorleftBig",0x230A),("floorleftBigg",0x230A), +("floorleftbig",0x230A),("floorleftbigg",0x230A),("floorright",0x230B),("floorrightBig",0x230B),("floorrightBigg",0x230B),("floorrightbig",0x230B), +("floorrightbigg",0x230B),("florin",0x0192),("fmonospace",0xFF46),("fmsquare",0x3399),("fofanthai",0x0E1F),("fofathai",0x0E1D),("follows",0x227B), +("followsequal",0x227D),("fongmanthai",0x0E4F),("forall",0x2200),("four",0x0034),("fourarabic",0x0664),("fourbengali",0x09EA),("fourcircle",0x2463), +("fourcircleinversesansserif",0x278D),("fourdeva",0x096A),("fourgujarati",0x0AEA),("fourgurmukhi",0x0A6A),("fourhackarabic",0x0664), +("fourhangzhou",0x3024),("fourideographicparen",0x3223),("fourinferior",0x2084),("fourmonospace",0xFF14),("fournumeratorbengali",0x09F7), +("fouroldstyle",0xF734),("fourparen",0x2477),("fourperiod",0x248B),("fourpersian",0x06F4),("fourroman",0x2173),("foursuperior",0x2074), +("fourteencircle",0x246D),("fourteenparen",0x2481),("fourteenperiod",0x2495),("fourthai",0x0E54),("fourthtonechinese",0x02CB),("fparen",0x24A1), +("fraction",0x2044),("franc",0x20A3),("g",0x0067),("gabengali",0x0997),("gacute",0x01F5),("gadeva",0x0917),("gafarabic",0x06AF), +("gaffinalarabic",0xFB93),("gafinitialarabic",0xFB94),("gafmedialarabic",0xFB95),("gagujarati",0x0A97),("gagurmukhi",0x0A17),("gahiragana",0x304C), +("gakatakana",0x30AC),("gamma",0x03B3),("gammalatinsmall",0x0263),("gammasuperior",0x02E0),("gangiacoptic",0x03EB),("gbopomofo",0x310D), +("gbreve",0x011F),("gcaron",0x01E7),("gcedilla",0x0123),("gcircle",0x24D6),("gcircumflex",0x011D),("gcommaaccent",0x0123),("gdot",0x0121), +("gdotaccent",0x0121),("gecyrillic",0x0433),("gehiragana",0x3052),("gekatakana",0x30B2),("geometricallyequal",0x2251),("gereshaccenthebrew",0x059C), +("gereshhebrew",0x05F3),("gereshmuqdamhebrew",0x059D),("germandbls",0x00DF),("gershayimaccenthebrew",0x059E),("gershayimhebrew",0x05F4), +("getamark",0x3013),("ghabengali",0x0998),("ghadarmenian",0x0572),("ghadeva",0x0918),("ghagujarati",0x0A98),("ghagurmukhi",0x0A18), +("ghainarabic",0x063A),("ghainfinalarabic",0xFECE),("ghaininitialarabic",0xFECF),("ghainmedialarabic",0xFED0),("ghemiddlehookcyrillic",0x0495), +("ghestrokecyrillic",0x0493),("gheupturncyrillic",0x0491),("ghhadeva",0x095A),("ghhagurmukhi",0x0A5A),("ghook",0x0260),("ghzsquare",0x3393), +("gihiragana",0x304E),("gikatakana",0x30AE),("gimarmenian",0x0563),("gimel",0x05D2),("gimeldagesh",0xFB32),("gimeldageshhebrew",0xFB32), +("gimelhebrew",0x05D2),("gjecyrillic",0x0453),("glottalinvertedstroke",0x01BE),("glottalstop",0x0294),("glottalstopinverted",0x0296), +("glottalstopmod",0x02C0),("glottalstopreversed",0x0295),("glottalstopreversedmod",0x02C1),("glottalstopreversedsuperior",0x02E4), +("glottalstopstroke",0x02A1),("glottalstopstrokereversed",0x02A2),("gmacron",0x1E21),("gmonospace",0xFF47),("gohiragana",0x3054), +("gokatakana",0x30B4),("gparen",0x24A2),("gpasquare",0x33AC),("gradient",0x2207),("grave",0x0060),("gravebelowcmb",0x0316),("gravecmb",0x0300), +("gravecomb",0x0300),("gravedeva",0x0953),("gravelowmod",0x02CE),("gravemonospace",0xFF40),("gravetonecmb",0x0340),("greater",0x003E), +("greaterequal",0x2265),("greaterequalorless",0x22DB),("greatermonospace",0xFF1E),("greatermuch",0x226B),("greaterorequivalent",0x2273), +("greaterorless",0x2277),("greateroverequal",0x2267),("greatersmall",0xFE65),("gscript",0x0261),("gstroke",0x01E5),("guhiragana",0x3050), +("guillemotleft",0x00AB),("guillemotright",0x00BB),("guilsinglleft",0x2039),("guilsinglright",0x203A),("gukatakana",0x30B0),("guramusquare",0x3318), +("gysquare",0x33C9),("h",0x0068),("haabkhasiancyrillic",0x04A9),("haaltonearabic",0x06C1),("habengali",0x09B9),("hadescendercyrillic",0x04B3), +("hadeva",0x0939),("hagujarati",0x0AB9),("hagurmukhi",0x0A39),("haharabic",0x062D),("hahfinalarabic",0xFEA2),("hahinitialarabic",0xFEA3), +("hahiragana",0x306F),("hahmedialarabic",0xFEA4),("haitusquare",0x332A),("hakatakana",0x30CF),("hakatakanahalfwidth",0xFF8A), +("halantgurmukhi",0x0A4D),("hamzaarabic",0x0621),("hamzadammaarabic",0x0621),("hamzadammatanarabic",0x0621),("hamzafathaarabic",0x0621), +("hamzafathatanarabic",0x0621),("hamzalowarabic",0x0621),("hamzalowkasraarabic",0x0621),("hamzalowkasratanarabic",0x0621),("hamzasukunarabic",0x0621), +("hangulfiller",0x3164),("hardsigncyrillic",0x044A),("harpoonleftbarbup",0x21BC),("harpoonleftdown",0x21BD),("harpoonleftup",0x21BC), +("harpoonrightbarbup",0x21C0),("harpoonrightdown",0x21C1),("harpoonrightup",0x21C0),("hasquare",0x33CA),("hatafpatah",0x05B2),("hatafpatah16",0x05B2), +("hatafpatah23",0x05B2),("hatafpatah2f",0x05B2),("hatafpatahhebrew",0x05B2),("hatafpatahnarrowhebrew",0x05B2),("hatafpatahquarterhebrew",0x05B2), +("hatafpatahwidehebrew",0x05B2),("hatafqamats",0x05B3),("hatafqamats1b",0x05B3),("hatafqamats28",0x05B3),("hatafqamats34",0x05B3), +("hatafqamatshebrew",0x05B3),("hatafqamatsnarrowhebrew",0x05B3),("hatafqamatsquarterhebrew",0x05B3),("hatafqamatswidehebrew",0x05B3), +("hatafsegol",0x05B1),("hatafsegol17",0x05B1),("hatafsegol24",0x05B1),("hatafsegol30",0x05B1),("hatafsegolhebrew",0x05B1), +("hatafsegolnarrowhebrew",0x05B1),("hatafsegolquarterhebrew",0x05B1),("hatafsegolwidehebrew",0x05B1),("hatwide",0x0302),("hatwider",0x0302), +("hatwiderr",0x0302),("hbar",0x0127),("hbopomofo",0x310F),("hbrevebelow",0x1E2B),("hcedilla",0x1E29),("hcircle",0x24D7),("hcircumflex",0x0125), +("hdieresis",0x1E27),("hdotaccent",0x1E23),("hdotbelow",0x1E25),("he",0x05D4),("heart",0x2661),("heart2",0x2665),("heartsuitblack",0x2665), +("heartsuitwhite",0x2661),("hedagesh",0xFB34),("hedageshhebrew",0xFB34),("hehaltonearabic",0x06C1),("heharabic",0x0647),("hehebrew",0x05D4), +("hehfinalaltonearabic",0xFBA7),("hehfinalalttwoarabic",0xFEEA),("hehfinalarabic",0xFEEA),("hehhamzaabovefinalarabic",0xFBA5), +("hehhamzaaboveisolatedarabic",0xFBA4),("hehinitialaltonearabic",0xFBA8),("hehinitialarabic",0xFEEB),("hehiragana",0x3078), +("hehmedialaltonearabic",0xFBA9),("hehmedialarabic",0xFEEC),("heiseierasquare",0x337B),("hekatakana",0x30D8),("hekatakanahalfwidth",0xFF8D), +("hekutaarusquare",0x3336),("henghook",0x0267),("herutusquare",0x3339),("het",0x05D7),("hethebrew",0x05D7),("hhook",0x0266),("hhooksuperior",0x02B1), +("hieuhacirclekorean",0x327B),("hieuhaparenkorean",0x321B),("hieuhcirclekorean",0x326D),("hieuhkorean",0x314E),("hieuhparenkorean",0x320D), +("hihiragana",0x3072),("hikatakana",0x30D2),("hikatakanahalfwidth",0xFF8B),("hiriq",0x05B4),("hiriq14",0x05B4),("hiriq21",0x05B4),("hiriq2d",0x05B4), +("hiriqhebrew",0x05B4),("hiriqnarrowhebrew",0x05B4),("hiriqquarterhebrew",0x05B4),("hiriqwidehebrew",0x05B4),("hlinebelow",0x1E96), +("hmonospace",0xFF48),("hoarmenian",0x0570),("hohipthai",0x0E2B),("hohiragana",0x307B),("hokatakana",0x30DB),("hokatakanahalfwidth",0xFF8E), +("holam",0x05B9),("holam19",0x05B9),("holam26",0x05B9),("holam32",0x05B9),("holamhebrew",0x05B9),("holamnarrowhebrew",0x05B9), +("holamquarterhebrew",0x05B9),("holamwidehebrew",0x05B9),("honokhukthai",0x0E2E),("hookabovecomb",0x0309),("hookcmb",0x0309),("hookleftchar",0x21A9), +("hookpalatalizedbelowcmb",0x0321),("hookretroflexbelowcmb",0x0322),("hookrightchar",0x21AA),("hoonsquare",0x3342),("horicoptic",0x03E9), +("horizontalbar",0x2015),("horncmb",0x031B),("hotsprings",0x2668),("house",0x2302),("hparen",0x24A3),("hsuperior",0x02B0),("hturned",0x0265), +("huhiragana",0x3075),("huiitosquare",0x3333),("hukatakana",0x30D5),("hukatakanahalfwidth",0xFF8C),("hungarumlaut",0x02DD),("hungarumlautcmb",0x030B), +("hv",0x0195),("hyphen",0x002D),("hyphen_alt",0x2010),("hyphenchar",0x002D),("hypheninferior",0xF6E5),("hyphenmonospace",0xFF0D), +("hyphensmall",0xFE63),("hyphensuperior",0xF6E6),("hyphentwo",0x2010),("i",0x0069),("iacute",0x00ED),("iacyrillic",0x044F),("ibengali",0x0987), +("ibopomofo",0x3127),("ibreve",0x012D),("icaron",0x01D0),("icircle",0x24D8),("icircumflex",0x00EE),("icyrillic",0x0456),("idblgrave",0x0209), +("ideographearthcircle",0x328F),("ideographfirecircle",0x328B),("ideographicallianceparen",0x323F),("ideographiccallparen",0x323A), +("ideographiccentrecircle",0x32A5),("ideographicclose",0x3006),("ideographiccomma",0x3001),("ideographiccommaleft",0xFF64), +("ideographiccongratulationparen",0x3237),("ideographiccorrectcircle",0x32A3),("ideographicearthparen",0x322F),("ideographicenterpriseparen",0x323D), +("ideographicexcellentcircle",0x329D),("ideographicfestivalparen",0x3240),("ideographicfinancialcircle",0x3296),("ideographicfinancialparen",0x3236), +("ideographicfireparen",0x322B),("ideographichaveparen",0x3232),("ideographichighcircle",0x32A4),("ideographiciterationmark",0x3005), +("ideographiclaborcircle",0x3298),("ideographiclaborparen",0x3238),("ideographicleftcircle",0x32A7),("ideographiclowcircle",0x32A6), +("ideographicmedicinecircle",0x32A9),("ideographicmetalparen",0x322E),("ideographicmoonparen",0x322A),("ideographicnameparen",0x3234), +("ideographicperiod",0x3002),("ideographicprintcircle",0x329E),("ideographicreachparen",0x3243),("ideographicrepresentparen",0x3239), +("ideographicresourceparen",0x323E),("ideographicrightcircle",0x32A8),("ideographicsecretcircle",0x3299),("ideographicselfparen",0x3242), +("ideographicsocietyparen",0x3233),("ideographicspace",0x3000),("ideographicspecialparen",0x3235),("ideographicstockparen",0x3231), +("ideographicstudyparen",0x323B),("ideographicsunparen",0x3230),("ideographicsuperviseparen",0x323C),("ideographicwaterparen",0x322C), +("ideographicwoodparen",0x322D),("ideographiczero",0x3007),("ideographmetalcircle",0x328E),("ideographmooncircle",0x328A), +("ideographnamecircle",0x3294),("ideographsuncircle",0x3290),("ideographwatercircle",0x328C),("ideographwoodcircle",0x328D),("ideva",0x0907), +("idieresis",0x00EF),("idieresisacute",0x1E2F),("idieresiscyrillic",0x04E5),("idotbelow",0x1ECB),("iebrevecyrillic",0x04D7),("iecyrillic",0x0435), +("ieungacirclekorean",0x3275),("ieungaparenkorean",0x3215),("ieungcirclekorean",0x3267),("ieungkorean",0x3147),("ieungparenkorean",0x3207), +("igrave",0x00EC),("igujarati",0x0A87),("igurmukhi",0x0A07),("ihiragana",0x3044),("ihookabove",0x1EC9),("iibengali",0x0988),("iicyrillic",0x0438), +("iideva",0x0908),("iigujarati",0x0A88),("iigurmukhi",0x0A08),("iimatragurmukhi",0x0A40),("iinvertedbreve",0x020B),("iishortcyrillic",0x0439), +("iivowelsignbengali",0x09C0),("iivowelsigndeva",0x0940),("iivowelsigngujarati",0x0AC0),("ij",0x0133),("ikatakana",0x30A4), +("ikatakanahalfwidth",0xFF72),("ikorean",0x3163),("ilde",0x02DC),("iluyhebrew",0x05AC),("imacron",0x012B),("imacroncyrillic",0x04E3), +("imageorapproximatelyequal",0x2253),("imatragurmukhi",0x0A3F),("imonospace",0xFF49),("increment",0x2206),("infinity",0x221E),("iniarmenian",0x056B), +("integral",0x222B),("integralbottom",0x2321),("integralbt",0x2321),("integraldisplay",0x222B),("integralex",0xF8F5),("integraltext",0x222B), +("integraltop",0x2320),("integraltp",0x2320),("interrobang",0x203D),("interrobangdown",0xD80B),("intersection",0x2229),("intersectiondisplay",0x22C2), +("intersectionsq",0x2293),("intersectiontext",0x22C2),("intisquare",0x3305),("invbullet",0x25D8),("invcircle",0x25D9),("invsmileface",0x263B), +("iocyrillic",0x0451),("iogonek",0x012F),("iota",0x03B9),("iotadieresis",0x03CA),("iotadieresistonos",0x0390),("iotalatin",0x0269), +("iotatonos",0x03AF),("iparen",0x24A4),("irigurmukhi",0x0A72),("ismallhiragana",0x3043),("ismallkatakana",0x30A3),("ismallkatakanahalfwidth",0xFF68), +("issharbengali",0x09FA),("istroke",0x0268),("isuperior",0xF6ED),("iterationhiragana",0x309D),("iterationkatakana",0x30FD),("itilde",0x0129), +("itildebelow",0x1E2D),("iubopomofo",0x3129),("iucyrillic",0x044E),("ivowelsignbengali",0x09BF),("ivowelsigndeva",0x093F), +("ivowelsigngujarati",0x0ABF),("izhitsacyrillic",0x0475),("izhitsadblgravecyrillic",0x0477),("j",0x006A),("jaarmenian",0x0571),("jabengali",0x099C), +("jadeva",0x091C),("jagujarati",0x0A9C),("jagurmukhi",0x0A1C),("jbopomofo",0x3110),("jcaron",0x01F0),("jcircle",0x24D9),("jcircumflex",0x0135), +("jcrossedtail",0x029D),("jdotlessstroke",0x025F),("jecyrillic",0x0458),("jeemarabic",0x062C),("jeemfinalarabic",0xFE9E),("jeeminitialarabic",0xFE9F), +("jeemmedialarabic",0xFEA0),("jeharabic",0x0698),("jehfinalarabic",0xFB8B),("jhabengali",0x099D),("jhadeva",0x091D),("jhagujarati",0x0A9D), +("jhagurmukhi",0x0A1D),("jheharmenian",0x057B),("jis",0x3004),("jmonospace",0xFF4A),("jparen",0x24A5),("jsuperior",0x02B2),("k",0x006B), +("kabashkircyrillic",0x04A1),("kabengali",0x0995),("kacute",0x1E31),("kacyrillic",0x043A),("kadescendercyrillic",0x049B),("kadeva",0x0915), +("kaf",0x05DB),("kafarabic",0x0643),("kafdagesh",0xFB3B),("kafdageshhebrew",0xFB3B),("kaffinalarabic",0xFEDA),("kafhebrew",0x05DB), +("kafinitialarabic",0xFEDB),("kafmedialarabic",0xFEDC),("kafrafehebrew",0xFB4D),("kagujarati",0x0A95),("kagurmukhi",0x0A15),("kahiragana",0x304B), +("kahookcyrillic",0x04C4),("kakatakana",0x30AB),("kakatakanahalfwidth",0xFF76),("kappa",0x03BA),("kappasymbolgreek",0x03F0), +("kapyeounmieumkorean",0x3171),("kapyeounphieuphkorean",0x3184),("kapyeounpieupkorean",0x3178),("kapyeounssangpieupkorean",0x3179), +("karoriisquare",0x330D),("kashidaautoarabic",0x0640),("kashidaautonosidebearingarabic",0x0640),("kasmallkatakana",0x30F5),("kasquare",0x3384), +("kasraarabic",0x0650),("kasratanarabic",0x064D),("kastrokecyrillic",0x049F),("katahiraprolongmarkhalfwidth",0xFF70), +("kaverticalstrokecyrillic",0x049D),("kbopomofo",0x310E),("kcalsquare",0x3389),("kcaron",0x01E9),("kcedilla",0x0137),("kcircle",0x24DA), +("kcommaaccent",0x0137),("kdotbelow",0x1E33),("keharmenian",0x0584),("kehiragana",0x3051),("kekatakana",0x30B1),("kekatakanahalfwidth",0xFF79), +("kenarmenian",0x056F),("kesmallkatakana",0x30F6),("kgreenlandic",0x0138),("khabengali",0x0996),("khacyrillic",0x0445),("khadeva",0x0916), +("khagujarati",0x0A96),("khagurmukhi",0x0A16),("khaharabic",0x062E),("khahfinalarabic",0xFEA6),("khahinitialarabic",0xFEA7), +("khahmedialarabic",0xFEA8),("kheicoptic",0x03E7),("khhadeva",0x0959),("khhagurmukhi",0x0A59),("khieukhacirclekorean",0x3278), +("khieukhaparenkorean",0x3218),("khieukhcirclekorean",0x326A),("khieukhkorean",0x314B),("khieukhparenkorean",0x320A),("khokhaithai",0x0E02), +("khokhonthai",0x0E05),("khokhuatthai",0x0E03),("khokhwaithai",0x0E04),("khomutthai",0x0E5B),("khook",0x0199),("khorakhangthai",0x0E06), +("khzsquare",0x3391),("kihiragana",0x304D),("kikatakana",0x30AD),("kikatakanahalfwidth",0xFF77),("kiroguramusquare",0x3315), +("kiromeetorusquare",0x3316),("kirosquare",0x3314),("kiyeokacirclekorean",0x326E),("kiyeokaparenkorean",0x320E),("kiyeokcirclekorean",0x3260), +("kiyeokkorean",0x3131),("kiyeokparenkorean",0x3200),("kiyeoksioskorean",0x3133),("kjecyrillic",0x045C),("klinebelow",0x1E35),("klsquare",0x3398), +("kmcubedsquare",0x33A6),("kmonospace",0xFF4B),("kmsquaredsquare",0x33A2),("kohiragana",0x3053),("kohmsquare",0x33C0),("kokaithai",0x0E01), +("kokatakana",0x30B3),("kokatakanahalfwidth",0xFF7A),("kooposquare",0x331E),("koppacyrillic",0x0481),("koreanstandardsymbol",0x327F), +("koroniscmb",0x0343),("kparen",0x24A6),("kpasquare",0x33AA),("ksicyrillic",0x046F),("ktsquare",0x33CF),("kturned",0x029E),("kuhiragana",0x304F), +("kukatakana",0x30AF),("kukatakanahalfwidth",0xFF78),("kvsquare",0x33B8),("kwsquare",0x33BE),("l",0x006C),("labengali",0x09B2),("lacute",0x013A), +("ladeva",0x0932),("lagujarati",0x0AB2),("lagurmukhi",0x0A32),("lakkhangyaothai",0x0E45),("lamaleffinalarabic",0xFEFC), +("lamalefhamzaabovefinalarabic",0xFEF8),("lamalefhamzaaboveisolatedarabic",0xFEF7),("lamalefhamzabelowfinalarabic",0xFEFA), +("lamalefhamzabelowisolatedarabic",0xFEF9),("lamalefisolatedarabic",0xFEFB),("lamalefmaddaabovefinalarabic",0xFEF6), +("lamalefmaddaaboveisolatedarabic",0xFEF5),("lamarabic",0x0644),("lambda",0x03BB),("lambdastroke",0x019B),("lamed",0x05DC),("lameddagesh",0xFB3C), +("lameddageshhebrew",0xFB3C),("lamedhebrew",0x05DC),("lamedholam",0x05DC),("lamedholamdagesh",0x05DC),("lamedholamdageshhebrew",0x05DC), +("lamedholamhebrew",0x05DC),("lamfinalarabic",0xFEDE),("lamhahinitialarabic",0xFCCA),("laminitialarabic",0xFEDF),("lamjeeminitialarabic",0xFCC9), +("lamkhahinitialarabic",0xFCCB),("lamlamhehisolatedarabic",0xFDF2),("lammedialarabic",0xFEE0),("lammeemhahinitialarabic",0xFD88), +("lammeeminitialarabic",0xFCCC),("lammeemjeeminitialarabic",0xFEDF),("lammeemkhahinitialarabic",0xFEDF),("largecircle",0x25EF),("latticetop",0x22A4), +("lbar",0x019A),("lbelt",0x026C),("lbopomofo",0x310C),("lcaron",0x013E),("lcedilla",0x013C),("lcircle",0x24DB),("lcircumflexbelow",0x1E3D), +("lcommaaccent",0x013C),("ldot",0x0140),("ldotaccent",0x0140),("ldotbelow",0x1E37),("ldotbelowmacron",0x1E39),("leftangleabovecmb",0x031A), +("lefttackbelowcmb",0x0318),("less",0x003C),("lessequal",0x2264),("lessequalorgreater",0x22DA),("lessmonospace",0xFF1C),("lessmuch",0x226A), +("lessorequivalent",0x2272),("lessorgreater",0x2276),("lessoverequal",0x2266),("lesssmall",0xFE64),("lezh",0x026E),("lfblock",0x258C), +("lhookretroflex",0x026D),("lira",0x20A4),("liwnarmenian",0x056C),("lj",0x01C9),("ljecyrillic",0x0459),("ll",0xF6C0),("lladeva",0x0933), +("llagujarati",0x0AB3),("llinebelow",0x1E3B),("llladeva",0x0934),("llvocalicbengali",0x09E1),("llvocalicdeva",0x0961), +("llvocalicvowelsignbengali",0x09E3),("llvocalicvowelsigndeva",0x0963),("lmiddletilde",0x026B),("lmonospace",0xFF4C),("lmsquare",0x33D0), +("lochulathai",0x0E2C),("logicaland",0x2227),("logicalanddisplay",0x22C0),("logicalandtext",0x22C0),("logicalnot",0x00AC), +("logicalnotreversed",0x2310),("logicalor",0x2228),("logicalordisplay",0x22C1),("logicalortext",0x22C1),("lolingthai",0x0E25),("longs",0x017F), +("lowlinecenterline",0xFE4E),("lowlinecmb",0x0332),("lowlinedashed",0xFE4D),("lozenge",0x25CA),("lparen",0x24A7),("lscript",0x2113),("lslash",0x0142), +("lsquare",0x2113),("lsuperior",0xF6EE),("ltshade",0x2591),("luthai",0x0E26),("lvocalicbengali",0x098C),("lvocalicdeva",0x090C), +("lvocalicvowelsignbengali",0x09E2),("lvocalicvowelsigndeva",0x0962),("lxsquare",0x33D3),("m",0x006D),("mabengali",0x09AE),("macron",0x00AF), +("macronbelowcmb",0x0331),("macroncmb",0x0304),("macronlowmod",0x02CD),("macronmonospace",0xFFE3),("macute",0x1E3F),("madeva",0x092E), +("magujarati",0x0AAE),("magurmukhi",0x0A2E),("mahapakhhebrew",0x05A4),("mahapakhlefthebrew",0x05A4),("mahiragana",0x307E), +("maichattawalowleftthai",0xF895),("maichattawalowrightthai",0xF894),("maichattawathai",0x0E4B),("maichattawaupperleftthai",0xF893), +("maieklowleftthai",0xF88C),("maieklowrightthai",0xF88B),("maiekthai",0x0E48),("maiekupperleftthai",0xF88A),("maihanakatleftthai",0xF884), +("maihanakatthai",0x0E31),("maitaikhuleftthai",0xF889),("maitaikhuthai",0x0E47),("maitholowleftthai",0xF88F),("maitholowrightthai",0xF88E), +("maithothai",0x0E49),("maithoupperleftthai",0xF88D),("maitrilowleftthai",0xF892),("maitrilowrightthai",0xF891),("maitrithai",0x0E4A), +("maitriupperleftthai",0xF890),("maiyamokthai",0x0E46),("makatakana",0x30DE),("makatakanahalfwidth",0xFF8F),("male",0x2642),("mansyonsquare",0x3347), +("maqafhebrew",0x05BE),("mars",0x2642),("masoracirclehebrew",0x05AF),("masquare",0x3383),("mbopomofo",0x3107),("mbsquare",0x33D4),("mcircle",0x24DC), +("mcubedsquare",0x33A5),("mdotaccent",0x1E41),("mdotbelow",0x1E43),("meemarabic",0x0645),("meemfinalarabic",0xFEE2),("meeminitialarabic",0xFEE3), +("meemmedialarabic",0xFEE4),("meemmeeminitialarabic",0xFCD1),("meemmeemisolatedarabic",0xFC48),("meetorusquare",0x334D),("mehiragana",0x3081), +("meizierasquare",0x337E),("mekatakana",0x30E1),("mekatakanahalfwidth",0xFF92),("mem",0x05DE),("memdagesh",0xFB3E),("memdageshhebrew",0xFB3E), +("memhebrew",0x05DE),("menarmenian",0x0574),("merkhahebrew",0x05A5),("merkhakefulahebrew",0x05A6),("merkhakefulalefthebrew",0x05A6), +("merkhalefthebrew",0x05A5),("mhook",0x0271),("mhzsquare",0x3392),("middledotkatakanahalfwidth",0xFF65),("middot",0x00B7), +("mieumacirclekorean",0x3272),("mieumaparenkorean",0x3212),("mieumcirclekorean",0x3264),("mieumkorean",0x3141),("mieumpansioskorean",0x3170), +("mieumparenkorean",0x3204),("mieumpieupkorean",0x316E),("mieumsioskorean",0x316F),("mihiragana",0x307F),("mikatakana",0x30DF), +("mikatakanahalfwidth",0xFF90),("minus",0x2212),("minusbelowcmb",0x0320),("minuscircle",0x2296),("minusmod",0x02D7),("minusplus",0x2213), +("minute",0x2032),("miribaarusquare",0x334A),("mirisquare",0x3349),("mlonglegturned",0x0270),("mlsquare",0x3396),("mmcubedsquare",0x33A3), +("mmonospace",0xFF4D),("mmsquaredsquare",0x339F),("mohiragana",0x3082),("mohmsquare",0x33C1),("mokatakana",0x30E2),("mokatakanahalfwidth",0xFF93), +("molsquare",0x33D6),("momathai",0x0E21),("moverssquare",0x33A7),("moverssquaredsquare",0x33A8),("mparen",0x24A8),("mpasquare",0x33AB), +("mssquare",0x33B3),("msuperior",0xF6EF),("mturned",0x026F),("mu",0x00B5),("mu1",0x00B5),("muasquare",0x3382),("muchgreater",0x226B), +("muchless",0x226A),("mufsquare",0x338C),("mugreek",0x03BC),("mugsquare",0x338D),("muhiragana",0x3080),("mukatakana",0x30E0), +("mukatakanahalfwidth",0xFF91),("mulsquare",0x3395),("multiply",0x00D7),("mumsquare",0x339B),("munahhebrew",0x05A3),("munahlefthebrew",0x05A3), +("musicalnote",0x266A),("musicalnotedbl",0x266B),("musicflatsign",0x266D),("musicsharpsign",0x266F),("mussquare",0x33B2),("muvsquare",0x33B6), +("muwsquare",0x33BC),("mvmegasquare",0x33B9),("mvsquare",0x33B7),("mwmegasquare",0x33BF),("mwsquare",0x33BD),("n",0x006E),("nabengali",0x09A8), +("nabla",0x2207),("nacute",0x0144),("nadeva",0x0928),("nagujarati",0x0AA8),("nagurmukhi",0x0A28),("nahiragana",0x306A),("nakatakana",0x30CA), +("nakatakanahalfwidth",0xFF85),("napostrophe",0x0149),("nasquare",0x3381),("natural",0x266E),("nbopomofo",0x310B),("nbspace",0x00A0), +("ncaron",0x0148),("ncedilla",0x0146),("ncircle",0x24DD),("ncircumflexbelow",0x1E4B),("ncommaaccent",0x0146),("ndotaccent",0x1E45), +("ndotbelow",0x1E47),("negationslash",0x0338),("nehiragana",0x306D),("nekatakana",0x30CD),("nekatakanahalfwidth",0xFF88),("newsheqelsign",0x20AA), +("nfsquare",0x338B),("ng",0x014B),("ngabengali",0x0999),("ngadeva",0x0919),("ngagujarati",0x0A99),("ngagurmukhi",0x0A19),("ngonguthai",0x0E07), +("nhiragana",0x3093),("nhookleft",0x0272),("nhookretroflex",0x0273),("nieunacirclekorean",0x326F),("nieunaparenkorean",0x320F), +("nieuncieuckorean",0x3135),("nieuncirclekorean",0x3261),("nieunhieuhkorean",0x3136),("nieunkorean",0x3134),("nieunpansioskorean",0x3168), +("nieunparenkorean",0x3201),("nieunsioskorean",0x3167),("nieuntikeutkorean",0x3166),("nihiragana",0x306B),("nikatakana",0x30CB), +("nikatakanahalfwidth",0xFF86),("nikhahitleftthai",0xF899),("nikhahitthai",0x0E4D),("nine",0x0039),("ninearabic",0x0669),("ninebengali",0x09EF), +("ninecircle",0x2468),("ninecircleinversesansserif",0x2792),("ninedeva",0x096F),("ninegujarati",0x0AEF),("ninegurmukhi",0x0A6F), +("ninehackarabic",0x0669),("ninehangzhou",0x3029),("nineideographicparen",0x3228),("nineinferior",0x2089),("ninemonospace",0xFF19), +("nineoldstyle",0xF739),("nineparen",0x247C),("nineperiod",0x2490),("ninepersian",0x06F9),("nineroman",0x2178),("ninesuperior",0x2079), +("nineteencircle",0x2472),("nineteenparen",0x2486),("nineteenperiod",0x249A),("ninethai",0x0E59),("nj",0x01CC),("njecyrillic",0x045A), +("nkatakana",0x30F3),("nkatakanahalfwidth",0xFF9D),("nlegrightlong",0x019E),("nlinebelow",0x1E49),("nmonospace",0xFF4E),("nmsquare",0x339A), +("nnabengali",0x09A3),("nnadeva",0x0923),("nnagujarati",0x0AA3),("nnagurmukhi",0x0A23),("nnnadeva",0x0929),("nohiragana",0x306E), +("nokatakana",0x30CE),("nokatakanahalfwidth",0xFF89),("nonbreakingspace",0x00A0),("nonenthai",0x0E13),("nonuthai",0x0E19),("noonarabic",0x0646), +("noonfinalarabic",0xFEE6),("noonghunnaarabic",0x06BA),("noonghunnafinalarabic",0xFB9F),("noonhehinitialarabic",0xFEE7),("nooninitialarabic",0xFEE7), +("noonjeeminitialarabic",0xFCD2),("noonjeemisolatedarabic",0xFC4B),("noonmedialarabic",0xFEE8),("noonmeeminitialarabic",0xFCD5), +("noonmeemisolatedarabic",0xFC4E),("noonnoonfinalarabic",0xFC8D),("notcontains",0x220C),("notelement",0x2209),("notelementof",0x2209), +("notequal",0x2260),("notgreater",0x226F),("notgreaternorequal",0x2271),("notgreaternorless",0x2279),("notidentical",0x2262),("notless",0x226E), +("notlessnorequal",0x2270),("notparallel",0x2226),("notprecedes",0x2280),("notsubset",0x2284),("notsucceeds",0x2281),("notsuperset",0x2285), +("nowarmenian",0x0576),("nparen",0x24A9),("nssquare",0x33B1),("nsuperior",0x207F),("ntilde",0x00F1),("nu",0x03BD),("nuhiragana",0x306C), +("nukatakana",0x30CC),("nukatakanahalfwidth",0xFF87),("nuktabengali",0x09BC),("nuktadeva",0x093C),("nuktagujarati",0x0ABC),("nuktagurmukhi",0x0A3C), +("numbersign",0x0023),("numbersignmonospace",0xFF03),("numbersignsmall",0xFE5F),("numeralsigngreek",0x0374),("numeralsignlowergreek",0x0375), +("numero",0x2116),("nun",0x05E0),("nundagesh",0xFB40),("nundageshhebrew",0xFB40),("nunhebrew",0x05E0),("nvsquare",0x33B5),("nwsquare",0x33BB), +("nyabengali",0x099E),("nyadeva",0x091E),("nyagujarati",0x0A9E),("nyagurmukhi",0x0A1E),("o",0x006F),("oacute",0x00F3),("oangthai",0x0E2D), +("obarred",0x0275),("obarredcyrillic",0x04E9),("obarreddieresiscyrillic",0x04EB),("obengali",0x0993),("obopomofo",0x311B),("obreve",0x014F), +("ocandradeva",0x0911),("ocandragujarati",0x0A91),("ocandravowelsigndeva",0x0949),("ocandravowelsigngujarati",0x0AC9),("ocaron",0x01D2), +("ocircle",0x24DE),("ocircumflex",0x00F4),("ocircumflexacute",0x1ED1),("ocircumflexdotbelow",0x1ED9),("ocircumflexgrave",0x1ED3), +("ocircumflexhookabove",0x1ED5),("ocircumflextilde",0x1ED7),("ocyrillic",0x043E),("odblacute",0x0151),("odblgrave",0x020D),("odeva",0x0913), +("odieresis",0x00F6),("odieresiscyrillic",0x04E7),("odotbelow",0x1ECD),("oe",0x0153),("oekorean",0x315A),("ogonek",0x02DB),("ogonekcmb",0x0328), +("ograve",0x00F2),("ogujarati",0x0A93),("oharmenian",0x0585),("ohiragana",0x304A),("ohookabove",0x1ECF),("ohorn",0x01A1),("ohornacute",0x1EDB), +("ohorndotbelow",0x1EE3),("ohorngrave",0x1EDD),("ohornhookabove",0x1EDF),("ohorntilde",0x1EE1),("ohungarumlaut",0x0151),("oi",0x01A3), +("oinvertedbreve",0x020F),("okatakana",0x30AA),("okatakanahalfwidth",0xFF75),("okorean",0x3157),("olehebrew",0x05AB),("omacron",0x014D), +("omacronacute",0x1E53),("omacrongrave",0x1E51),("omdeva",0x0950),("omega",0x03C9),("omega1",0x03D6),("omegacyrillic",0x0461), +("omegalatinclosed",0x0277),("omegaroundcyrillic",0x047B),("omegatitlocyrillic",0x047D),("omegatonos",0x03CE),("omgujarati",0x0AD0), +("omicron",0x03BF),("omicrontonos",0x03CC),("omonospace",0xFF4F),("one",0x0031),("onearabic",0x0661),("onebengali",0x09E7),("onecircle",0x2460), +("onecircleinversesansserif",0x278A),("onedeva",0x0967),("onedotenleader",0x2024),("oneeighth",0x215B),("onefitted",0xF6DC),("onegujarati",0x0AE7), +("onegurmukhi",0x0A67),("onehackarabic",0x0661),("onehalf",0x00BD),("onehangzhou",0x3021),("oneideographicparen",0x3220),("oneinferior",0x2081), +("onemonospace",0xFF11),("onenumeratorbengali",0x09F4),("oneoldstyle",0xF731),("oneparen",0x2474),("oneperiod",0x2488),("onepersian",0x06F1), +("onequarter",0x00BC),("oneroman",0x2170),("onesuperior",0x00B9),("onethai",0x0E51),("onethird",0x2153),("oogonek",0x01EB),("oogonekmacron",0x01ED), +("oogurmukhi",0x0A13),("oomatragurmukhi",0x0A4B),("oopen",0x0254),("oparen",0x24AA),("openbullet",0x25E6),("option",0x2325),("ordfeminine",0x00AA), +("ordmasculine",0x00BA),("orthogonal",0x221F),("oshortdeva",0x0912),("oshortvowelsigndeva",0x094A),("oslash",0x00F8),("oslashacute",0x01FF), +("osmallhiragana",0x3049),("osmallkatakana",0x30A9),("osmallkatakanahalfwidth",0xFF6B),("ostrokeacute",0x01FF),("osuperior",0xF6F0), +("otcyrillic",0x047F),("otilde",0x00F5),("otildeacute",0x1E4D),("otildedieresis",0x1E4F),("oubopomofo",0x3121),("overline",0x203E), +("overlinecenterline",0xFE4A),("overlinecmb",0x0305),("overlinedashed",0xFE49),("overlinedblwavy",0xFE4C),("overlinewavy",0xFE4B), +("overscore",0x00AF),("ovowelsignbengali",0x09CB),("ovowelsigndeva",0x094B),("ovowelsigngujarati",0x0ACB),("owner",0x220B),("p",0x0070), +("paampssquare",0x3380),("paasentosquare",0x332B),("pabengali",0x09AA),("pacute",0x1E55),("padeva",0x092A),("pagedown",0x21DF),("pageup",0x21DE), +("pagujarati",0x0AAA),("pagurmukhi",0x0A2A),("pahiragana",0x3071),("paiyannoithai",0x0E2F),("pakatakana",0x30D1),("palatalizationcyrilliccmb",0x0484), +("palochkacyrillic",0x04C0),("pansioskorean",0x317F),("paragraph",0x00B6),("parallel",0x2225),("parenleft",0x0028),("parenleftBig",0x0028), +("parenleftBigg",0x0028),("parenleftaltonearabic",0xFD3E),("parenleftbig",0x0028),("parenleftbigg",0x0028),("parenleftbt",0xF8ED), +("parenleftex",0x007C),("parenleftex2",0xF8EC),("parenleftinferior",0x208D),("parenleftmonospace",0xFF08),("parenleftsmall",0xFE59), +("parenleftsuperior",0x207D),("parenlefttp",0xF8EB),("parenleftvertical",0xFE35),("parenright",0x0029),("parenrightBig",0x0029), +("parenrightBigg",0x0029),("parenrightaltonearabic",0xFD3F),("parenrightbig",0x0029),("parenrightbigg",0x0029),("parenrightbt",0xF8F8), +("parenrightex",0x007C),("parenrightex2",0xF8F7),("parenrightinferior",0x208E),("parenrightmonospace",0xFF09),("parenrightsmall",0xFE5A), +("parenrightsuperior",0x207E),("parenrighttp",0xF8F6),("parenrightvertical",0xFE36),("partialdiff",0x2202),("paseqhebrew",0x05C0), +("pashtahebrew",0x0599),("pasquare",0x33A9),("patah",0x05B7),("patah11",0x05B7),("patah1d",0x05B7),("patah2a",0x05B7),("patahhebrew",0x05B7), +("patahnarrowhebrew",0x05B7),("patahquarterhebrew",0x05B7),("patahwidehebrew",0x05B7),("pazerhebrew",0x05A1),("pbopomofo",0x3106),("pcircle",0x24DF), +("pdotaccent",0x1E57),("pe",0x05E4),("pecyrillic",0x043F),("pedagesh",0xFB44),("pedageshhebrew",0xFB44),("peezisquare",0x333B), +("pefinaldageshhebrew",0xFB43),("peharabic",0x067E),("peharmenian",0x057A),("pehebrew",0x05E4),("pehfinalarabic",0xFB57),("pehinitialarabic",0xFB58), +("pehiragana",0x307A),("pehmedialarabic",0xFB59),("pekatakana",0x30DA),("pemiddlehookcyrillic",0x04A7),("perafehebrew",0xFB4E),("percent",0x0025), +("percentarabic",0x066A),("percentmonospace",0xFF05),("percentsmall",0xFE6A),("period",0x002E),("periodarmenian",0x0589),("periodcentered",0x00B7), +("periodhalfwidth",0xFF61),("periodinferior",0xF6E7),("periodmonospace",0xFF0E),("periodsmall",0xFE52),("periodsuperior",0xF6E8), +("perispomenigreekcmb",0x0342),("perpendicular",0x22A5),("pertenthousand",0x2031),("perthousand",0x2030),("peseta",0x20A7),("pfsquare",0x338A), +("phabengali",0x09AB),("phadeva",0x092B),("phagujarati",0x0AAB),("phagurmukhi",0x0A2B),("phi",0x03C6),("phi2",0x03D5),("phieuphacirclekorean",0x327A), +("phieuphaparenkorean",0x321A),("phieuphcirclekorean",0x326C),("phieuphkorean",0x314D),("phieuphparenkorean",0x320C),("philatin",0x0278), +("phinthuthai",0x0E3A),("phisymbolgreek",0x03D5),("phook",0x01A5),("phophanthai",0x0E1E),("phophungthai",0x0E1C),("phosamphaothai",0x0E20), +("pi",0x03C0),("pi1",0x03D6),("pieupacirclekorean",0x3273),("pieupaparenkorean",0x3213),("pieupcieuckorean",0x3176),("pieupcirclekorean",0x3265), +("pieupkiyeokkorean",0x3172),("pieupkorean",0x3142),("pieupparenkorean",0x3205),("pieupsioskiyeokkorean",0x3174),("pieupsioskorean",0x3144), +("pieupsiostikeutkorean",0x3175),("pieupthieuthkorean",0x3177),("pieuptikeutkorean",0x3173),("pihiragana",0x3074),("pikatakana",0x30D4), +("pisymbolgreek",0x03D6),("piwrarmenian",0x0583),("plus",0x002B),("plusbelowcmb",0x031F),("pluscircle",0x2295),("plusminus",0x00B1), +("plusmod",0x02D6),("plusmonospace",0xFF0B),("plussmall",0xFE62),("plussuperior",0x207A),("pmonospace",0xFF50),("pmsquare",0x33D8), +("pohiragana",0x307D),("pointingindexdownwhite",0x261F),("pointingindexleftwhite",0x261C),("pointingindexrightwhite",0x261E), +("pointingindexupwhite",0x261D),("pokatakana",0x30DD),("poplathai",0x0E1B),("postalmark",0x3012),("postalmarkface",0x3020),("pparen",0x24AB), +("precedes",0x227A),("precedesequal",0x227C),("prescription",0x211E),("prime",0x2032),("primemod",0x02B9),("primereversed",0x2035),("product",0x220F), +("productdisplay",0x220F),("producttext",0x220F),("projective",0x2305),("prolongedkana",0x30FC),("propellor",0x2318),("propersubset",0x2282), +("propersuperset",0x2283),("proportion",0x2237),("proportional",0x221D),("psi",0x03C8),("psicyrillic",0x0471),("psilipneumatacyrilliccmb",0x0486), +("pssquare",0x33B0),("puhiragana",0x3077),("pukatakana",0x30D7),("punctdash",0x2014),("pvsquare",0x33B4),("pwsquare",0x33BA),("q",0x0071), +("qadeva",0x0958),("qadmahebrew",0x05A8),("qafarabic",0x0642),("qaffinalarabic",0xFED6),("qafinitialarabic",0xFED7),("qafmedialarabic",0xFED8), +("qamats",0x05B8),("qamats10",0x05B8),("qamats1a",0x05B8),("qamats1c",0x05B8),("qamats27",0x05B8),("qamats29",0x05B8),("qamats33",0x05B8), +("qamatsde",0x05B8),("qamatshebrew",0x05B8),("qamatsnarrowhebrew",0x05B8),("qamatsqatanhebrew",0x05B8),("qamatsqatannarrowhebrew",0x05B8), +("qamatsqatanquarterhebrew",0x05B8),("qamatsqatanwidehebrew",0x05B8),("qamatsquarterhebrew",0x05B8),("qamatswidehebrew",0x05B8), +("qarneyparahebrew",0x059F),("qbopomofo",0x3111),("qcircle",0x24E0),("qhook",0x02A0),("qmonospace",0xFF51),("qof",0x05E7),("qofdagesh",0xFB47), +("qofdageshhebrew",0xFB47),("qofhatafpatah",0x05E7),("qofhatafpatahhebrew",0x05E7),("qofhatafsegol",0x05E7),("qofhatafsegolhebrew",0x05E7), +("qofhebrew",0x05E7),("qofhiriq",0x05E7),("qofhiriqhebrew",0x05E7),("qofholam",0x05E7),("qofholamhebrew",0x05E7),("qofpatah",0x05E7), +("qofpatahhebrew",0x05E7),("qofqamats",0x05E7),("qofqamatshebrew",0x05E7),("qofqubuts",0x05E7),("qofqubutshebrew",0x05E7),("qofsegol",0x05E7), +("qofsegolhebrew",0x05E7),("qofsheva",0x05E7),("qofshevahebrew",0x05E7),("qoftsere",0x05E7),("qoftserehebrew",0x05E7),("qparen",0x24AC), +("quarternote",0x2669),("qubuts",0x05BB),("qubuts18",0x05BB),("qubuts25",0x05BB),("qubuts31",0x05BB),("qubutshebrew",0x05BB), +("qubutsnarrowhebrew",0x05BB),("qubutsquarterhebrew",0x05BB),("qubutswidehebrew",0x05BB),("question",0x003F),("questionarabic",0x061F), +("questionarmenian",0x055E),("questiondown",0x00BF),("questiondownsmall",0xF7BF),("questiongreek",0x037E),("questionmonospace",0xFF1F), +("questionsmall",0xF73F),("quotedbl",0x0022),("quotedblbase",0x201E),("quotedblleft",0x201C),("quotedblmonospace",0xFF02),("quotedblprime",0x301E), +("quotedblprimereversed",0x301D),("quotedblright",0x201D),("quoteleft",0x2018),("quoteleftreversed",0x201B),("quotereversed",0x201B), +("quoteright",0x2019),("quoterightn",0x0149),("quotesinglbase",0x201A),("quotesingle",0x0027),("quotesinglemonospace",0xFF07),("r",0x0072), +("raarmenian",0x057C),("rabengali",0x09B0),("racute",0x0155),("radeva",0x0930),("radical",0x221A),("radicalBig",0x221A),("radicalBigg",0x221A), +("radicalbig",0x221A),("radicalbigg",0x221A),("radicalbt",0x221A),("radicalex",0xF8E5),("radoverssquare",0x33AE),("radoverssquaredsquare",0x33AF), +("radsquare",0x33AD),("rafe",0x05BF),("rafehebrew",0x05BF),("ragujarati",0x0AB0),("ragurmukhi",0x0A30),("rahiragana",0x3089),("rakatakana",0x30E9), +("rakatakanahalfwidth",0xFF97),("ralowerdiagonalbengali",0x09F1),("ramiddlediagonalbengali",0x09F0),("ramshorn",0x0264),("rangedash",0x2013), +("ratio",0x2236),("rbopomofo",0x3116),("rcaron",0x0159),("rcedilla",0x0157),("rcircle",0x24E1),("rcommaaccent",0x0157),("rdblgrave",0x0211), +("rdotaccent",0x1E59),("rdotbelow",0x1E5B),("rdotbelowmacron",0x1E5D),("referencemark",0x203B),("reflexsubset",0x2286),("reflexsuperset",0x2287), +("registered",0x00AE),("registersans",0xF8E8),("registerserif",0xF6DA),("reharabic",0x0631),("reharmenian",0x0580),("rehfinalarabic",0xFEAE), +("rehiragana",0x308C),("rehyehaleflamarabic",0x0631),("rekatakana",0x30EC),("rekatakanahalfwidth",0xFF9A),("resh",0x05E8),("reshdageshhebrew",0xFB48), +("reshhatafpatah",0x05E8),("reshhatafpatahhebrew",0x05E8),("reshhatafsegol",0x05E8),("reshhatafsegolhebrew",0x05E8),("reshhebrew",0x05E8), +("reshhiriq",0x05E8),("reshhiriqhebrew",0x05E8),("reshholam",0x05E8),("reshholamhebrew",0x05E8),("reshpatah",0x05E8),("reshpatahhebrew",0x05E8), +("reshqamats",0x05E8),("reshqamatshebrew",0x05E8),("reshqubuts",0x05E8),("reshqubutshebrew",0x05E8),("reshsegol",0x05E8),("reshsegolhebrew",0x05E8), +("reshsheva",0x05E8),("reshshevahebrew",0x05E8),("reshtsere",0x05E8),("reshtserehebrew",0x05E8),("reversedtilde",0x223D),("reviahebrew",0x0597), +("reviamugrashhebrew",0x0597),("revlogicalnot",0x2310),("rfishhook",0x027E),("rfishhookreversed",0x027F),("rhabengali",0x09DD),("rhadeva",0x095D), +("rho",0x03C1),("rho1",0x03F1),("rhook",0x027D),("rhookturned",0x027B),("rhookturnedsuperior",0x02B5),("rhosymbolgreek",0x03F1), +("rhotichookmod",0x02DE),("rieulacirclekorean",0x3271),("rieulaparenkorean",0x3211),("rieulcirclekorean",0x3263),("rieulhieuhkorean",0x3140), +("rieulkiyeokkorean",0x313A),("rieulkiyeoksioskorean",0x3169),("rieulkorean",0x3139),("rieulmieumkorean",0x313B),("rieulpansioskorean",0x316C), +("rieulparenkorean",0x3203),("rieulphieuphkorean",0x313F),("rieulpieupkorean",0x313C),("rieulpieupsioskorean",0x316B),("rieulsioskorean",0x313D), +("rieulthieuthkorean",0x313E),("rieultikeutkorean",0x316A),("rieulyeorinhieuhkorean",0x316D),("rightangle",0x221F),("righttackbelowcmb",0x0319), +("righttriangle",0x22BF),("rihiragana",0x308A),("rikatakana",0x30EA),("rikatakanahalfwidth",0xFF98),("ring",0x02DA),("ringbelowcmb",0x0325), +("ringcmb",0x030A),("ringhalfleft",0x02BF),("ringhalfleftarmenian",0x0559),("ringhalfleftbelowcmb",0x031C),("ringhalfleftcentered",0x02D3), +("ringhalfright",0x02BE),("ringhalfrightbelowcmb",0x0339),("ringhalfrightcentered",0x02D2),("rinvertedbreve",0x0213),("rittorusquare",0x3351), +("rlinebelow",0x1E5F),("rlongleg",0x027C),("rlonglegturned",0x027A),("rmonospace",0xFF52),("rohiragana",0x308D),("rokatakana",0x30ED), +("rokatakanahalfwidth",0xFF9B),("roruathai",0x0E23),("rparen",0x24AD),("rrabengali",0x09DC),("rradeva",0x0931),("rragurmukhi",0x0A5C), +("rreharabic",0x0691),("rrehfinalarabic",0xFB8D),("rrvocalicbengali",0x09E0),("rrvocalicdeva",0x0960),("rrvocalicgujarati",0x0AE0), +("rrvocalicvowelsignbengali",0x09C4),("rrvocalicvowelsigndeva",0x0944),("rrvocalicvowelsigngujarati",0x0AC4),("rsuperior",0xF6F1),("rtblock",0x2590), +("rturned",0x0279),("rturnedsuperior",0x02B4),("ruhiragana",0x308B),("rukatakana",0x30EB),("rukatakanahalfwidth",0xFF99),("rupeemarkbengali",0x09F2), +("rupeesignbengali",0x09F3),("rupiah",0xF6DD),("ruthai",0x0E24),("rvocalicbengali",0x098B),("rvocalicdeva",0x090B),("rvocalicgujarati",0x0A8B), +("rvocalicvowelsignbengali",0x09C3),("rvocalicvowelsigndeva",0x0943),("rvocalicvowelsigngujarati",0x0AC3),("s",0x0073),("sabengali",0x09B8), +("sacute",0x015B),("sacutedotaccent",0x1E65),("sadarabic",0x0635),("sadeva",0x0938),("sadfinalarabic",0xFEBA),("sadinitialarabic",0xFEBB), +("sadmedialarabic",0xFEBC),("sagujarati",0x0AB8),("sagurmukhi",0x0A38),("sahiragana",0x3055),("sakatakana",0x30B5),("sakatakanahalfwidth",0xFF7B), +("sallallahoualayhewasallamarabic",0xFDFA),("samekh",0x05E1),("samekhdagesh",0xFB41),("samekhdageshhebrew",0xFB41),("samekhhebrew",0x05E1), +("saraaathai",0x0E32),("saraaethai",0x0E41),("saraaimaimalaithai",0x0E44),("saraaimaimuanthai",0x0E43),("saraamthai",0x0E33),("saraathai",0x0E30), +("saraethai",0x0E40),("saraiileftthai",0xF886),("saraiithai",0x0E35),("saraileftthai",0xF885),("saraithai",0x0E34),("saraothai",0x0E42), +("saraueeleftthai",0xF888),("saraueethai",0x0E37),("saraueleftthai",0xF887),("sarauethai",0x0E36),("sarauthai",0x0E38),("sarauuthai",0x0E39), +("sbopomofo",0x3119),("scaron",0x0161),("scarondotaccent",0x1E67),("scedilla",0x015F),("schwa",0x0259),("schwacyrillic",0x04D9), +("schwadieresiscyrillic",0x04DB),("schwahook",0x025A),("scircle",0x24E2),("scircumflex",0x015D),("scommaaccent",0x0219),("sdotaccent",0x1E61), +("sdotbelow",0x1E63),("sdotbelowdotaccent",0x1E69),("seagullbelowcmb",0x033C),("second",0x2033),("secondtonechinese",0x02CA),("section",0x00A7), +("seenarabic",0x0633),("seenfinalarabic",0xFEB2),("seeninitialarabic",0xFEB3),("seenmedialarabic",0xFEB4),("segol",0x05B6),("segol13",0x05B6), +("segol1f",0x05B6),("segol2c",0x05B6),("segolhebrew",0x05B6),("segolnarrowhebrew",0x05B6),("segolquarterhebrew",0x05B6),("segoltahebrew",0x0592), +("segolwidehebrew",0x05B6),("seharmenian",0x057D),("sehiragana",0x305B),("sekatakana",0x30BB),("sekatakanahalfwidth",0xFF7E),("semicolon",0x003B), +("semicolonarabic",0x061B),("semicolonmonospace",0xFF1B),("semicolonsmall",0xFE54),("semivoicedmarkkana",0x309C), +("semivoicedmarkkanahalfwidth",0xFF9F),("sentisquare",0x3322),("sentosquare",0x3323),("seven",0x0037),("sevenarabic",0x0667),("sevenbengali",0x09ED), +("sevencircle",0x2466),("sevencircleinversesansserif",0x2790),("sevendeva",0x096D),("seveneighths",0x215E),("sevengujarati",0x0AED), +("sevengurmukhi",0x0A6D),("sevenhackarabic",0x0667),("sevenhangzhou",0x3027),("sevenideographicparen",0x3226),("seveninferior",0x2087), +("sevenmonospace",0xFF17),("sevenoldstyle",0xF737),("sevenparen",0x247A),("sevenperiod",0x248E),("sevenpersian",0x06F7),("sevenroman",0x2176), +("sevensuperior",0x2077),("seventeencircle",0x2470),("seventeenparen",0x2484),("seventeenperiod",0x2498),("seventhai",0x0E57),("sfthyphen",0x00AD), +("shaarmenian",0x0577),("shabengali",0x09B6),("shacyrillic",0x0448),("shaddaarabic",0x0651),("shaddadammaarabic",0xFC61), +("shaddadammatanarabic",0xFC5E),("shaddafathaarabic",0xFC60),("shaddafathatanarabic",0x0651),("shaddakasraarabic",0xFC62), +("shaddakasratanarabic",0xFC5F),("shade",0x2592),("shadedark",0x2593),("shadelight",0x2591),("shademedium",0x2592),("shadeva",0x0936), +("shagujarati",0x0AB6),("shagurmukhi",0x0A36),("shalshelethebrew",0x0593),("sharp",0x266F),("shbopomofo",0x3115),("shchacyrillic",0x0449), +("sheenarabic",0x0634),("sheenfinalarabic",0xFEB6),("sheeninitialarabic",0xFEB7),("sheenmedialarabic",0xFEB8),("sheicoptic",0x03E3),("sheqel",0x20AA), +("sheqelhebrew",0x20AA),("sheva",0x05B0),("sheva115",0x05B0),("sheva15",0x05B0),("sheva22",0x05B0),("sheva2e",0x05B0),("shevahebrew",0x05B0), +("shevanarrowhebrew",0x05B0),("shevaquarterhebrew",0x05B0),("shevawidehebrew",0x05B0),("shhacyrillic",0x04BB),("shimacoptic",0x03ED),("shin",0x05E9), +("shindagesh",0xFB49),("shindageshhebrew",0xFB49),("shindageshshindot",0xFB2C),("shindageshshindothebrew",0xFB2C),("shindageshsindot",0xFB2D), +("shindageshsindothebrew",0xFB2D),("shindothebrew",0x05C1),("shinhebrew",0x05E9),("shinshindot",0xFB2A),("shinshindothebrew",0xFB2A), +("shinsindot",0xFB2B),("shinsindothebrew",0xFB2B),("shook",0x0282),("sigma",0x03C3),("sigma1",0x03C2),("sigmafinal",0x03C2), +("sigmalunatesymbolgreek",0x03F2),("sihiragana",0x3057),("sikatakana",0x30B7),("sikatakanahalfwidth",0xFF7C),("siluqhebrew",0x05BD), +("siluqlefthebrew",0x05BD),("similar",0x223C),("similarequal",0x2243),("sindothebrew",0x05C2),("siosacirclekorean",0x3274), +("siosaparenkorean",0x3214),("sioscieuckorean",0x317E),("sioscirclekorean",0x3266),("sioskiyeokkorean",0x317A),("sioskorean",0x3145), +("siosnieunkorean",0x317B),("siosparenkorean",0x3206),("siospieupkorean",0x317D),("siostikeutkorean",0x317C),("six",0x0036),("sixarabic",0x0666), +("sixbengali",0x09EC),("sixcircle",0x2465),("sixcircleinversesansserif",0x278F),("sixdeva",0x096C),("sixgujarati",0x0AEC),("sixgurmukhi",0x0A6C), +("sixhackarabic",0x0666),("sixhangzhou",0x3026),("sixideographicparen",0x3225),("sixinferior",0x2086),("sixmonospace",0xFF16),("sixoldstyle",0xF736), +("sixparen",0x2479),("sixperiod",0x248D),("sixpersian",0x06F6),("sixroman",0x2175),("sixsuperior",0x2076),("sixteencircle",0x246F), +("sixteencurrencydenominatorbengali",0x09F9),("sixteenparen",0x2483),("sixteenperiod",0x2497),("sixthai",0x0E56),("slash",0x002F),("slashBig",0x2215), +("slashBigg",0x2215),("slashbig",0x2215),("slashbigg",0x2215),("slashmonospace",0xFF0F),("slong",0x017F),("slongdotaccent",0x1E9B), +("slurabove",0x2322),("slurbelow",0x2323),("smileface",0x263A),("smonospace",0xFF53),("sofpasuqhebrew",0x05C3),("softhyphen",0x00AD), +("softsigncyrillic",0x044C),("sohiragana",0x305D),("sokatakana",0x30BD),("sokatakanahalfwidth",0xFF7F),("soliduslongoverlaycmb",0x0338), +("solidusshortoverlaycmb",0x0337),("sorusithai",0x0E29),("sosalathai",0x0E28),("sosothai",0x0E0B),("sosuathai",0x0E2A),("space",0x0020), +("spacehackarabic",0x0020),("spade",0x2660),("spadesuitblack",0x2660),("spadesuitwhite",0x2664),("sparen",0x24AE),("squarebelowcmb",0x033B), +("squarecc",0x33C4),("squarecm",0x339D),("squarediagonalcrosshatchfill",0x25A9),("squarehorizontalfill",0x25A4),("squarekg",0x338F), +("squarekm",0x339E),("squarekmcapital",0x33CE),("squareln",0x33D1),("squarelog",0x33D2),("squaremg",0x338E),("squaremil",0x33D5),("squaremm",0x339C), +("squaremsquared",0x33A1),("squareorthogonalcrosshatchfill",0x25A6),("squareupperlefttolowerrightfill",0x25A7), +("squareupperrighttolowerleftfill",0x25A8),("squareverticalfill",0x25A5),("squarewhitewithsmallblack",0x25A3),("srsquare",0x33DB), +("ssabengali",0x09B7),("ssadeva",0x0937),("ssagujarati",0x0AB7),("ssangcieuckorean",0x3149),("ssanghieuhkorean",0x3185),("ssangieungkorean",0x3180), +("ssangkiyeokkorean",0x3132),("ssangnieunkorean",0x3165),("ssangpieupkorean",0x3143),("ssangsioskorean",0x3146),("ssangtikeutkorean",0x3138), +("ssuperior",0xF6F2),("star",0x22C6),("sterling",0x00A3),("sterlingmonospace",0xFFE1),("strokelongoverlaycmb",0x0336), +("strokeshortoverlaycmb",0x0335),("subset",0x2282),("subsetnotequal",0x228A),("subsetorequal",0x2286),("subsetsqequal",0x2291),("succeeds",0x227B), +("suchthat",0x220B),("suhiragana",0x3059),("sukatakana",0x30B9),("sukatakanahalfwidth",0xFF7D),("sukunarabic",0x0652),("summation",0x2211), +("summationdisplay",0x2211),("summationtext",0x2211),("sun",0x263C),("superset",0x2283),("supersetnotequal",0x228B),("supersetorequal",0x2287), +("supersetsqequal",0x2292),("svsquare",0x33DC),("syouwaerasquare",0x337C),("t",0x0074),("tabengali",0x09A4),("tackdown",0x22A4),("tackleft",0x22A3), +("tadeva",0x0924),("tagujarati",0x0AA4),("tagurmukhi",0x0A24),("taharabic",0x0637),("tahfinalarabic",0xFEC2),("tahinitialarabic",0xFEC3), +("tahiragana",0x305F),("tahmedialarabic",0xFEC4),("taisyouerasquare",0x337D),("takatakana",0x30BF),("takatakanahalfwidth",0xFF80), +("tatweelarabic",0x0640),("tau",0x03C4),("tav",0x05EA),("tavdages",0xFB4A),("tavdagesh",0xFB4A),("tavdageshhebrew",0xFB4A),("tavhebrew",0x05EA), +("tbar",0x0167),("tbopomofo",0x310A),("tcaron",0x0165),("tccurl",0x02A8),("tcedilla",0x0163),("tcheharabic",0x0686),("tchehfinalarabic",0xFB7B), +("tchehinitialarabic",0xFB7C),("tchehmedialarabic",0xFB7D),("tchehmeeminitialarabic",0xFB7C),("tcircle",0x24E3),("tcircumflexbelow",0x1E71), +("tcommaaccent",0x0163),("tdieresis",0x1E97),("tdotaccent",0x1E6B),("tdotbelow",0x1E6D),("tecyrillic",0x0442),("tedescendercyrillic",0x04AD), +("teharabic",0x062A),("tehfinalarabic",0xFE96),("tehhahinitialarabic",0xFCA2),("tehhahisolatedarabic",0xFC0C),("tehinitialarabic",0xFE97), +("tehiragana",0x3066),("tehjeeminitialarabic",0xFCA1),("tehjeemisolatedarabic",0xFC0B),("tehmarbutaarabic",0x0629),("tehmarbutafinalarabic",0xFE94), +("tehmedialarabic",0xFE98),("tehmeeminitialarabic",0xFCA4),("tehmeemisolatedarabic",0xFC0E),("tehnoonfinalarabic",0xFC73),("tekatakana",0x30C6), +("tekatakanahalfwidth",0xFF83),("telephone",0x2121),("telephoneblack",0x260E),("telishagedolahebrew",0x05A0),("telishaqetanahebrew",0x05A9), +("tencircle",0x2469),("tenideographicparen",0x3229),("tenparen",0x247D),("tenperiod",0x2491),("tenroman",0x2179),("tesh",0x02A7),("tet",0x05D8), +("tetdagesh",0xFB38),("tetdageshhebrew",0xFB38),("tethebrew",0x05D8),("tetsecyrillic",0x04B5),("tevirhebrew",0x059B),("tevirlefthebrew",0x059B), +("thabengali",0x09A5),("thadeva",0x0925),("thagujarati",0x0AA5),("thagurmukhi",0x0A25),("thalarabic",0x0630),("thalfinalarabic",0xFEAC), +("thanthakhatlowleftthai",0xF898),("thanthakhatlowrightthai",0xF897),("thanthakhatthai",0x0E4C),("thanthakhatupperleftthai",0xF896), +("theharabic",0x062B),("thehfinalarabic",0xFE9A),("thehinitialarabic",0xFE9B),("thehmedialarabic",0xFE9C),("thereexists",0x2203),("therefore",0x2234), +("theta",0x03B8),("theta1",0x03D1),("thetasymbolgreek",0x03D1),("thieuthacirclekorean",0x3279),("thieuthaparenkorean",0x3219), +("thieuthcirclekorean",0x326B),("thieuthkorean",0x314C),("thieuthparenkorean",0x320B),("thirteencircle",0x246C),("thirteenparen",0x2480), +("thirteenperiod",0x2494),("thonangmonthothai",0x0E11),("thook",0x01AD),("thophuthaothai",0x0E12),("thorn",0x00FE),("thothahanthai",0x0E17), +("thothanthai",0x0E10),("thothongthai",0x0E18),("thothungthai",0x0E16),("thousandcyrillic",0x0482),("thousandsseparatorarabic",0x066C), +("thousandsseparatorpersian",0x066C),("three",0x0033),("threearabic",0x0663),("threebengali",0x09E9),("threecircle",0x2462), +("threecircleinversesansserif",0x278C),("threedeva",0x0969),("threeeighths",0x215C),("threegujarati",0x0AE9),("threegurmukhi",0x0A69), +("threehackarabic",0x0663),("threehangzhou",0x3023),("threeideographicparen",0x3222),("threeinferior",0x2083),("threemonospace",0xFF13), +("threenumeratorbengali",0x09F6),("threeoldstyle",0xF733),("threeparen",0x2476),("threeperiod",0x248A),("threepersian",0x06F3), +("threequarters",0x00BE),("threequartersemdash",0xF6DE),("threeroman",0x2172),("threesuperior",0x00B3),("threethai",0x0E53),("thzsquare",0x3394), +("tie",0x2040),("tihiragana",0x3061),("tikatakana",0x30C1),("tikatakanahalfwidth",0xFF81),("tikeutacirclekorean",0x3270), +("tikeutaparenkorean",0x3210),("tikeutcirclekorean",0x3262),("tikeutkorean",0x3137),("tikeutparenkorean",0x3202),("tilde",0x02DC), +("tildebelowcmb",0x0330),("tildecmb",0x0303),("tildecomb",0x0303),("tildedoublecmb",0x0360),("tildeoperator",0x223C),("tildeoverlaycmb",0x0334), +("tildeverticalcmb",0x033E),("tildewide",0x0303),("tildewider",0x0303),("tildewiderr",0x0303),("timescircle",0x2297),("tipehahebrew",0x0596), +("tipehalefthebrew",0x0596),("tippigurmukhi",0x0A70),("titlocyrilliccmb",0x0483),("tiwnarmenian",0x057F),("tlinebelow",0x1E6F),("tmonospace",0xFF54), +("toarmenian",0x0569),("tohiragana",0x3068),("tokatakana",0x30C8),("tokatakanahalfwidth",0xFF84),("tonebarextrahighmod",0x02E5), +("tonebarextralowmod",0x02E9),("tonebarhighmod",0x02E6),("tonebarlowmod",0x02E8),("tonebarmidmod",0x02E7),("tonefive",0x01BD),("tonesix",0x0185), +("tonetwo",0x01A8),("tonos",0x0384),("tonsquare",0x3327),("topatakthai",0x0E0F),("tortoiseshellbracketleft",0x3014), +("tortoiseshellbracketleftsmall",0xFE5D),("tortoiseshellbracketleftvertical",0xFE39),("tortoiseshellbracketright",0x3015), +("tortoiseshellbracketrightsmall",0xFE5E),("tortoiseshellbracketrightvertical",0xFE3A),("totaothai",0x0E15),("tpalatalhook",0x01AB),("tparen",0x24AF), +("trademark",0x2122),("trademarksans",0xF8EA),("trademarkserif",0xF6DB),("tretroflexhook",0x0288),("triagdn",0x25BC),("triaglf",0x25C4), +("triagrt",0x25BA),("triagup",0x25B2),("triangle",0x25B3),("triangleinv",0x25BD),("triangleleft",0x25B9),("triangleright",0x25C3),("ts",0x02A6), +("tsadi",0x05E6),("tsadidagesh",0xFB46),("tsadidageshhebrew",0xFB46),("tsadihebrew",0x05E6),("tsecyrillic",0x0446),("tsere",0x05B5), +("tsere12",0x05B5),("tsere1e",0x05B5),("tsere2b",0x05B5),("tserehebrew",0x05B5),("tserenarrowhebrew",0x05B5),("tserequarterhebrew",0x05B5), +("tserewidehebrew",0x05B5),("tshecyrillic",0x045B),("tsuperior",0xF6F3),("ttabengali",0x099F),("ttadeva",0x091F),("ttagujarati",0x0A9F), +("ttagurmukhi",0x0A1F),("tteharabic",0x0679),("ttehfinalarabic",0xFB67),("ttehinitialarabic",0xFB68),("ttehmedialarabic",0xFB69), +("tthabengali",0x09A0),("tthadeva",0x0920),("tthagujarati",0x0AA0),("tthagurmukhi",0x0A20),("tturned",0x0287),("tuhiragana",0x3064), +("tukatakana",0x30C4),("tukatakanahalfwidth",0xFF82),("turnstileleft",0x22A2),("turnstileright",0x22A3),("tusmallhiragana",0x3063), +("tusmallkatakana",0x30C3),("tusmallkatakanahalfwidth",0xFF6F),("twelvecircle",0x246B),("twelveparen",0x247F),("twelveperiod",0x2493), +("twelveroman",0x217B),("twelveudash",0xF6DE),("twentycircle",0x2473),("twentyhangzhou",0x5344),("twentyparen",0x2487),("twentyperiod",0x249B), +("two",0x0032),("twoarabic",0x0662),("twobengali",0x09E8),("twocircle",0x2461),("twocircleinversesansserif",0x278B),("twodeva",0x0968), +("twodotenleader",0x2025),("twodotleader",0x2025),("twodotleadervertical",0xFE30),("twogujarati",0x0AE8),("twogurmukhi",0x0A68), +("twohackarabic",0x0662),("twohangzhou",0x3022),("twoideographicparen",0x3221),("twoinferior",0x2082),("twomonospace",0xFF12), +("twonumeratorbengali",0x09F5),("twooldstyle",0xF732),("twoparen",0x2475),("twoperiod",0x2489),("twopersian",0x06F2),("tworoman",0x2171), +("twostroke",0x01BB),("twosuperior",0x00B2),("twothai",0x0E52),("twothirds",0x2154),("u",0x0075),("uacute",0x00FA),("ubar",0x0289), +("ubengali",0x0989),("ubopomofo",0x3128),("ubreve",0x016D),("ucaron",0x01D4),("ucircle",0x24E4),("ucircumflex",0x00FB),("ucircumflexbelow",0x1E77), +("ucyrillic",0x0443),("udattadeva",0x0951),("udblacute",0x0171),("udblgrave",0x0215),("udeva",0x0909),("udieresis",0x00FC),("udieresisacute",0x01D8), +("udieresisbelow",0x1E73),("udieresiscaron",0x01DA),("udieresiscyrillic",0x04F1),("udieresisgrave",0x01DC),("udieresismacron",0x01D6), +("udotbelow",0x1EE5),("ugrave",0x00F9),("ugujarati",0x0A89),("ugurmukhi",0x0A09),("uhiragana",0x3046),("uhookabove",0x1EE7),("uhorn",0x01B0), +("uhornacute",0x1EE9),("uhorndotbelow",0x1EF1),("uhorngrave",0x1EEB),("uhornhookabove",0x1EED),("uhorntilde",0x1EEF),("uhungarumlaut",0x0171), +("uhungarumlautcyrillic",0x04F3),("uinvertedbreve",0x0217),("ukatakana",0x30A6),("ukatakanahalfwidth",0xFF73),("ukcyrillic",0x0479), +("ukorean",0x315C),("umacron",0x016B),("umacroncyrillic",0x04EF),("umacrondieresis",0x1E7B),("umatragurmukhi",0x0A41),("umonospace",0xFF55), +("underscore",0x005F),("underscoredbl",0x2017),("underscoremonospace",0xFF3F),("underscorevertical",0xFE33),("underscorewavy",0xFE4F), +("union",0x222A),("uniondisplay",0x22C3),("unionmulti",0x228E),("unionmultidisplay",0x228E),("unionmultitext",0x228E),("unionsq",0x2294), +("unionsqdisplay",0x2294),("unionsqtext",0x2294),("uniontext",0x22C3),("universal",0x2200),("uogonek",0x0173),("uparen",0x24B0),("upblock",0x2580), +("upperdothebrew",0x05C4),("upsilon",0x03C5),("upsilondieresis",0x03CB),("upsilondieresistonos",0x03B0),("upsilonlatin",0x028A), +("upsilontonos",0x03CD),("uptackbelowcmb",0x031D),("uptackmod",0x02D4),("uragurmukhi",0x0A73),("uring",0x016F),("ushortcyrillic",0x045E), +("usmallhiragana",0x3045),("usmallkatakana",0x30A5),("usmallkatakanahalfwidth",0xFF69),("ustraightcyrillic",0x04AF), +("ustraightstrokecyrillic",0x04B1),("utilde",0x0169),("utildeacute",0x1E79),("utildebelow",0x1E75),("uubengali",0x098A),("uudeva",0x090A), +("uugujarati",0x0A8A),("uugurmukhi",0x0A0A),("uumatragurmukhi",0x0A42),("uuvowelsignbengali",0x09C2),("uuvowelsigndeva",0x0942), +("uuvowelsigngujarati",0x0AC2),("uvowelsignbengali",0x09C1),("uvowelsigndeva",0x0941),("uvowelsigngujarati",0x0AC1),("v",0x0076),("vadeva",0x0935), +("vagujarati",0x0AB5),("vagurmukhi",0x0A35),("vakatakana",0x30F7),("vav",0x05D5),("vavdagesh",0xFB35),("vavdagesh65",0xFB35), +("vavdageshhebrew",0xFB35),("vavhebrew",0x05D5),("vavholam",0xFB4B),("vavholamhebrew",0xFB4B),("vavvavhebrew",0x05F0),("vavyodhebrew",0x05F1), +("vcircle",0x24E5),("vdotbelow",0x1E7F),("vector",0x20D7),("vecyrillic",0x0432),("veharabic",0x06A4),("vehfinalarabic",0xFB6B), +("vehinitialarabic",0xFB6C),("vehmedialarabic",0xFB6D),("vekatakana",0x30F9),("venus",0x2640),("verticalbar",0x007C),("verticallineabovecmb",0x030D), +("verticallinebelowcmb",0x0329),("verticallinelowmod",0x02CC),("verticallinemod",0x02C8),("vewarmenian",0x057E),("vhook",0x028B), +("vikatakana",0x30F8),("viramabengali",0x09CD),("viramadeva",0x094D),("viramagujarati",0x0ACD),("visargabengali",0x0983),("visargadeva",0x0903), +("visargagujarati",0x0A83),("visiblespace",0x2423),("visualspace",0x2423),("vmonospace",0xFF56),("voarmenian",0x0578), +("voicediterationhiragana",0x309E),("voicediterationkatakana",0x30FE),("voicedmarkkana",0x309B),("voicedmarkkanahalfwidth",0xFF9E), +("vokatakana",0x30FA),("vparen",0x24B1),("vtilde",0x1E7D),("vturned",0x028C),("vuhiragana",0x3094),("vukatakana",0x30F4),("w",0x0077), +("wacute",0x1E83),("waekorean",0x3159),("wahiragana",0x308F),("wakatakana",0x30EF),("wakatakanahalfwidth",0xFF9C),("wakorean",0x3158), +("wasmallhiragana",0x308E),("wasmallkatakana",0x30EE),("wattosquare",0x3357),("wavedash",0x301C),("wavyunderscorevertical",0xFE34), +("wawarabic",0x0648),("wawfinalarabic",0xFEEE),("wawhamzaabovearabic",0x0624),("wawhamzaabovefinalarabic",0xFE86),("wbsquare",0x33DD), +("wcircle",0x24E6),("wcircumflex",0x0175),("wdieresis",0x1E85),("wdotaccent",0x1E87),("wdotbelow",0x1E89),("wehiragana",0x3091), +("weierstrass",0x2118),("wekatakana",0x30F1),("wekorean",0x315E),("weokorean",0x315D),("wgrave",0x1E81),("whitebullet",0x25E6),("whitecircle",0x25CB), +("whitecircleinverse",0x25D9),("whitecornerbracketleft",0x300E),("whitecornerbracketleftvertical",0xFE43),("whitecornerbracketright",0x300F), +("whitecornerbracketrightvertical",0xFE44),("whitediamond",0x25C7),("whitediamondcontainingblacksmalldiamond",0x25C8), +("whitedownpointingsmalltriangle",0x25BF),("whitedownpointingtriangle",0x25BD),("whiteleftpointingsmalltriangle",0x25C3), +("whiteleftpointingtriangle",0x25C1),("whitelenticularbracketleft",0x3016),("whitelenticularbracketright",0x3017), +("whiterightpointingsmalltriangle",0x25B9),("whiterightpointingtriangle",0x25B7),("whitesmallsquare",0x25AB),("whitesmilingface",0x263A), +("whitesquare",0x25A1),("whitestar",0x2606),("whitetelephone",0x260F),("whitetortoiseshellbracketleft",0x3018), +("whitetortoiseshellbracketright",0x3019),("whiteuppointingsmalltriangle",0x25B5),("whiteuppointingtriangle",0x25B3),("wihiragana",0x3090), +("wikatakana",0x30F0),("wikorean",0x315F),("wmonospace",0xFF57),("wohiragana",0x3092),("wokatakana",0x30F2),("wokatakanahalfwidth",0xFF66), +("won",0x20A9),("wonmonospace",0xFFE6),("wowaenthai",0x0E27),("wparen",0x24B2),("wreathproduct",0x2240),("wring",0x1E98),("wsuperior",0x02B7), +("wturned",0x028D),("wynn",0x01BF),("x",0x0078),("xabovecmb",0x033D),("xbopomofo",0x3112),("xcircle",0x24E7),("xdieresis",0x1E8D), +("xdotaccent",0x1E8B),("xeharmenian",0x056D),("xi",0x03BE),("xmonospace",0xFF58),("xparen",0x24B3),("xsuperior",0x02E3),("y",0x0079), +("yaadosquare",0x334E),("yabengali",0x09AF),("yacute",0x00FD),("yadeva",0x092F),("yaekorean",0x3152),("yagujarati",0x0AAF),("yagurmukhi",0x0A2F), +("yahiragana",0x3084),("yakatakana",0x30E4),("yakatakanahalfwidth",0xFF94),("yakorean",0x3151),("yamakkanthai",0x0E4E),("yasmallhiragana",0x3083), +("yasmallkatakana",0x30E3),("yasmallkatakanahalfwidth",0xFF6C),("yatcyrillic",0x0463),("ycircle",0x24E8),("ycircumflex",0x0177),("ydieresis",0x00FF), +("ydotaccent",0x1E8F),("ydotbelow",0x1EF5),("yeharabic",0x064A),("yehbarreearabic",0x06D2),("yehbarreefinalarabic",0xFBAF),("yehfinalarabic",0xFEF2), +("yehhamzaabovearabic",0x0626),("yehhamzaabovefinalarabic",0xFE8A),("yehhamzaaboveinitialarabic",0xFE8B),("yehhamzaabovemedialarabic",0xFE8C), +("yehinitialarabic",0xFEF3),("yehmedialarabic",0xFEF4),("yehmeeminitialarabic",0xFCDD),("yehmeemisolatedarabic",0xFC58),("yehnoonfinalarabic",0xFC94), +("yehthreedotsbelowarabic",0x06D1),("yekorean",0x3156),("yen",0x00A5),("yenmonospace",0xFFE5),("yeokorean",0x3155),("yeorinhieuhkorean",0x3186), +("yerahbenyomohebrew",0x05AA),("yerahbenyomolefthebrew",0x05AA),("yericyrillic",0x044B),("yerudieresiscyrillic",0x04F9),("yesieungkorean",0x3181), +("yesieungpansioskorean",0x3183),("yesieungsioskorean",0x3182),("yetivhebrew",0x059A),("ygrave",0x1EF3),("yhook",0x01B4),("yhookabove",0x1EF7), +("yiarmenian",0x0575),("yicyrillic",0x0457),("yikorean",0x3162),("yinyang",0x262F),("yiwnarmenian",0x0582),("ymonospace",0xFF59),("yod",0x05D9), +("yoddagesh",0xFB39),("yoddageshhebrew",0xFB39),("yodhebrew",0x05D9),("yodyodhebrew",0x05F2),("yodyodpatahhebrew",0xFB1F),("yohiragana",0x3088), +("yoikorean",0x3189),("yokatakana",0x30E8),("yokatakanahalfwidth",0xFF96),("yokorean",0x315B),("yosmallhiragana",0x3087),("yosmallkatakana",0x30E7), +("yosmallkatakanahalfwidth",0xFF6E),("yotgreek",0x03F3),("yoyaekorean",0x3188),("yoyakorean",0x3187),("yoyakthai",0x0E22),("yoyingthai",0x0E0D), +("yparen",0x24B4),("ypogegrammeni",0x037A),("ypogegrammenigreekcmb",0x0345),("yr",0x01A6),("yring",0x1E99),("ysuperior",0x02B8),("ytilde",0x1EF9), +("yturned",0x028E),("yuhiragana",0x3086),("yuikorean",0x318C),("yukatakana",0x30E6),("yukatakanahalfwidth",0xFF95),("yukorean",0x3160), +("yusbigcyrillic",0x046B),("yusbigiotifiedcyrillic",0x046D),("yuslittlecyrillic",0x0467),("yuslittleiotifiedcyrillic",0x0469), +("yusmallhiragana",0x3085),("yusmallkatakana",0x30E5),("yusmallkatakanahalfwidth",0xFF6D),("yuyekorean",0x318B),("yuyeokorean",0x318A), +("yyabengali",0x09DF),("yyadeva",0x095F),("z",0x007A),("zaarmenian",0x0566),("zacute",0x017A),("zadeva",0x095B),("zagurmukhi",0x0A5B), +("zaharabic",0x0638),("zahfinalarabic",0xFEC6),("zahinitialarabic",0xFEC7),("zahiragana",0x3056),("zahmedialarabic",0xFEC8),("zainarabic",0x0632), +("zainfinalarabic",0xFEB0),("zakatakana",0x30B6),("zaqefgadolhebrew",0x0595),("zaqefqatanhebrew",0x0594),("zarqahebrew",0x0598),("zayin",0x05D6), +("zayindagesh",0xFB36),("zayindageshhebrew",0xFB36),("zayinhebrew",0x05D6),("zbopomofo",0x3117),("zcaron",0x017E),("zcircle",0x24E9), +("zcircumflex",0x1E91),("zcurl",0x0291),("zdot",0x017C),("zdotaccent",0x017C),("zdotbelow",0x1E93),("zecyrillic",0x0437), +("zedescendercyrillic",0x0499),("zedieresiscyrillic",0x04DF),("zehiragana",0x305C),("zekatakana",0x30BC),("zero",0x0030),("zeroarabic",0x0660), +("zerobengali",0x09E6),("zerodeva",0x0966),("zerogujarati",0x0AE6),("zerogurmukhi",0x0A66),("zerohackarabic",0x0660),("zeroinferior",0x2080), +("zeromonospace",0xFF10),("zerooldstyle",0xF730),("zeropersian",0x06F0),("zerosuperior",0x2070),("zerothai",0x0E50),("zerowidthjoiner",0xFEFF), +("zerowidthnonjoiner",0x200C),("zerowidthspace",0x200B),("zeta",0x03B6),("zhbopomofo",0x3113),("zhearmenian",0x056A),("zhebrevecyrillic",0x04C2), +("zhecyrillic",0x0436),("zhedescendercyrillic",0x0497),("zhedieresiscyrillic",0x04DD),("zihiragana",0x3058),("zikatakana",0x30B8), +("zinorhebrew",0x05AE),("zlinebelow",0x1E95),("zmonospace",0xFF5A),("zohiragana",0x305E),("zokatakana",0x30BE),("zparen",0x24B5), +("zretroflexhook",0x0290),("zstroke",0x01B6),("zuhiragana",0x305A),("zukatakana",0x30BA), +]; diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/cff_encoding.rs b/src-tauri/src/pdf_engine/text_edit/fonts/cff_encoding.rs new file mode 100644 index 0000000..0237d38 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/cff_encoding.rs @@ -0,0 +1,250 @@ +//! A CFF program's own (built-in) encoding (SPEC §B.9.7), read by a bounded parser because +//! ttf-parser's `Encoding` is crate-private and its `glyph_index(code)` falls back to +//! StandardEncoding for codes the custom encoding lacks (`cff1.rs:967-981`) — that fallback would +//! call glyphs typeable that the font's encoding never maps. +//! +//! header → Name INDEX → Top DICT INDEX (first entry) → operator 16 (`Encoding`, default 0). +//! 0 → Standard, 1 → Expert, else an offset to format 0 (`nCodes`, `code[i]` → GID i+1) or +//! format 1 (ranges, consecutive GIDs from 1). Supplements (high bit) are ignored, so those codes +//! stay non-typeable. A CID-keyed font (`ROS`) has no encoding → `Err`. Checked slicing only. + +use std::collections::BTreeMap; + +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum CffEncoding { + Standard, + Expert, + Custom(BTreeMap), +} + +/// Operands kept per DICT operator (the CFF limit is 48). +const DICT_OPERANDS_MAX: usize = 48; +const OP_ENCODING: u16 = 16; +const OP_ROS: u16 = 0x0C1E; // 12 30 + +/// Parses the built-in encoding of a bare CFF (`FontFile3 /Type1C`) program. +pub fn cff_builtin_encoding(cff: &[u8]) -> Result { + if cff.first() != Some(&1) { + return Err(()); + } + let header_size = usize::from(*cff.get(2).ok_or(())?); + let name_index = Index::parse(cff, header_size)?; + let top_index = Index::parse(cff, name_index.end)?; + let top_dict = top_index.entry(cff, 0)?; + let mut encoding: i64 = 0; + for (op, operands) in DictIter::new(top_dict) { + let (op, operands) = (op?, operands); + match op { + OP_ROS => return Err(()), + OP_ENCODING => encoding = *operands.first().ok_or(())?, + _ => {} + } + } + match encoding { + 0 => Ok(CffEncoding::Standard), + 1 => Ok(CffEncoding::Expert), + offset => custom(cff, usize::try_from(offset).map_err(|_| ())?), + } +} + +fn custom(cff: &[u8], at: usize) -> Result { + let format = *cff.get(at).ok_or(())?; + let count = usize::from(*cff.get(at.checked_add(1).ok_or(())?).ok_or(())?); + let body = at.checked_add(2).ok_or(())?; + let mut map = BTreeMap::new(); + match format & 0x7F { + 0 => { + let codes = cff + .get(body..body.checked_add(count).ok_or(())?) + .ok_or(())?; + for (i, code) in codes.iter().enumerate() { + let gid = u16::try_from(i + 1).map_err(|_| ())?; + map.entry(*code).or_insert(gid); + } + } + 1 => { + let ranges = cff + .get( + body..body + .checked_add(count.checked_mul(2).ok_or(())?) + .ok_or(())?, + ) + .ok_or(())?; + let mut gid: u16 = 1; + for range in ranges.chunks_exact(2) { + let (first, left) = match range { + [f, l] => (u16::from(*f), u16::from(*l)), + _ => return Err(()), + }; + for code in first..=first + left { + let code = u8::try_from(code).map_err(|_| ())?; + map.entry(code).or_insert(gid); + gid = gid.checked_add(1).ok_or(())?; + } + } + } + _ => return Err(()), + } + Ok(CffEncoding::Custom(map)) +} + +/// A CFF INDEX: `count` (u16), `offSize`, `count + 1` offsets (1-based), data. +struct Index { + count: usize, + off_size: usize, + offsets_at: usize, + data_at: usize, + end: usize, +} + +impl Index { + fn parse(cff: &[u8], at: usize) -> Result { + let count = usize::from(u16::from_be_bytes([ + *cff.get(at).ok_or(())?, + *cff.get(at.checked_add(1).ok_or(())?).ok_or(())?, + ])); + let after_count = at.checked_add(2).ok_or(())?; + if count == 0 { + return Ok(Index { + count, + off_size: 1, + offsets_at: after_count, + data_at: after_count, + end: after_count, + }); + } + let off_size = usize::from(*cff.get(after_count).ok_or(())?); + if !(1..=4).contains(&off_size) { + return Err(()); + } + let offsets_at = after_count.checked_add(1).ok_or(())?; + let data_at = offsets_at + .checked_add((count + 1).checked_mul(off_size).ok_or(())?) + .ok_or(())?; + let mut index = Index { + count, + off_size, + offsets_at, + data_at, + end: 0, + }; + let last = index.offset(cff, count)?; + index.end = data_at + .checked_add(last.checked_sub(1).ok_or(())?) + .ok_or(())?; + if index.end > cff.len() { + return Err(()); + } + Ok(index) + } + + fn offset(&self, cff: &[u8], i: usize) -> Result { + let at = self + .offsets_at + .checked_add(i.checked_mul(self.off_size).ok_or(())?) + .ok_or(())?; + let bytes = cff + .get(at..at.checked_add(self.off_size).ok_or(())?) + .ok_or(())?; + Ok(bytes + .iter() + .fold(0usize, |acc, b| (acc << 8) | usize::from(*b))) + } + + fn entry<'a>(&self, cff: &'a [u8], i: usize) -> Result<&'a [u8], ()> { + if i >= self.count { + return Err(()); + } + let start = self.offset(cff, i)?.checked_sub(1).ok_or(())?; + let end = self.offset(cff, i + 1)?.checked_sub(1).ok_or(())?; + if end < start { + return Err(()); + } + cff.get( + self.data_at.checked_add(start).ok_or(())?..self.data_at.checked_add(end).ok_or(())?, + ) + .ok_or(()) + } +} + +/// Iterates `(operator, integer operands)` of a DICT; reals are skipped as operands (kept as 0). +struct DictIter<'a> { + data: &'a [u8], + pos: usize, + failed: bool, +} + +impl<'a> DictIter<'a> { + fn new(data: &'a [u8]) -> Self { + DictIter { + data, + pos: 0, + failed: false, + } + } + + fn byte(&mut self) -> Result { + let b = *self.data.get(self.pos).ok_or(())?; + self.pos += 1; + Ok(b) + } + + fn next_entry(&mut self) -> Result<(u16, Vec), ()> { + let mut operands = Vec::new(); + loop { + let b0 = self.byte()?; + let value = match b0 { + 0..=21 => { + let op = if b0 == 12 { + 0x0C00 | u16::from(self.byte()?) + } else { + u16::from(b0) + }; + return Ok((op, operands)); + } + 28 => i64::from(i16::from_be_bytes([self.byte()?, self.byte()?])), + 29 => i64::from(i32::from_be_bytes([ + self.byte()?, + self.byte()?, + self.byte()?, + self.byte()?, + ])), + 30 => { + // real: nibbles up to and including an 0xF nibble + loop { + let b = self.byte()?; + if b & 0x0F == 0x0F || b >> 4 == 0x0F { + break; + } + } + 0 + } + 32..=246 => i64::from(b0) - 139, + 247..=250 => (i64::from(b0) - 247) * 256 + i64::from(self.byte()?) + 108, + 251..=254 => -(i64::from(b0) - 251) * 256 - i64::from(self.byte()?) - 108, + _ => return Err(()), + }; + if operands.len() >= DICT_OPERANDS_MAX { + return Err(()); + } + operands.push(value); + } + } +} + +impl Iterator for DictIter<'_> { + type Item = (Result, Vec); + + fn next(&mut self) -> Option { + if self.failed || self.pos >= self.data.len() { + return None; + } + match self.next_entry() { + Ok((op, operands)) => Some((Ok(op), operands)), + Err(()) => { + self.failed = true; + Some((Err(()), Vec::new())) + } + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/cff_layout.rs b/src-tauri/src/pdf_engine/text_edit/fonts/cff_layout.rs new file mode 100644 index 0000000..3433f9b --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/cff_layout.rs @@ -0,0 +1,750 @@ +//! The layout of a CFF (version 1) program exactly as ttf-parser 0.25.1 reads it +//! (`tables/cff/{cff1,index,dict,charset}.rs`, MIT OR Apache-2.0): INDEX and DICT structures, the +//! Top DICT offsets, the global and local subroutines (FDSelect → Font DICT → Private DICT for +//! CID-keyed fonts) and the charset. ttf-parser keeps all of them crate-private, but +//! - the bounded charstring pre-check (`glyph_budget.rs`) must walk exactly the bytes +//! `cff::Table::outline` would run, so it follows every rule here that decides which bytes those +//! are (offsets taken as ttf-parser takes them, later Top DICT entries winning, the first +//! Font DICT `Private` entry, FDSelect formats 0 and 3, the seac charset lookup); +//! - ttf-parser walks a format 1/2 charset once per glyph (`glyph_name`, `glyph_cid`), which is +//! quadratic over a whole font; `gid_to_sid` inverts it once. +//! +//! A program this reader cannot follow (a real-number offset, or a structure it rejects that +//! ttf-parser accepted) is `None`; callers then fail closed. Checked slicing only. +//! +//! DICTs are walked lazily (no entry list is built) and only up to `DICT_LEN_MAX` bytes: a Top +//! or name-keyed Private DICT past it makes the program unreadable, a CID-keyed font's Font or +//! Private DICT past it gives that Font DICT no local subroutines (its glyphs that call one do not +//! draw). ttf-parser walks a CID-keyed font's Font and Private DICT again for every glyph it +//! outlines that calls a local subroutine, so `FdSubrs::dict_bytes` reports their size for the +//! pre-check to charge per glyph. + +use std::collections::HashMap; +use std::ops::Range; + +/// The 391 CFF standard strings (Adobe TN #5176 Appendix A), as ttf-parser 0.25.1 +/// `src/tables/cff/std_names.rs` lists them (MIT OR Apache-2.0). SID n < 391 names +/// `STANDARD_STRINGS[n]`; higher SIDs index the String INDEX. +#[rustfmt::skip] +pub(crate) const STANDARD_STRINGS: [&str; 391] = [ + ".notdef", "space", "exclam", "quotedbl", "numbersign", "dollar", "percent", "ampersand", + "quoteright", "parenleft", "parenright", "asterisk", "plus", "comma", "hyphen", "period", + "slash", "zero", "one", "two", "three", "four", "five", "six", "seven", "eight", "nine", + "colon", "semicolon", "less", "equal", "greater", "question", "at", "A", "B", "C", "D", "E", + "F", "G", "H", "I", "J", "K", "L", "M", "N", "O", "P", "Q", "R", "S", "T", "U", "V", "W", "X", + "Y", "Z", "bracketleft", "backslash", "bracketright", "asciicircum", "underscore", "quoteleft", + "a", "b", "c", "d", "e", "f", "g", "h", "i", "j", "k", "l", "m", "n", "o", "p", "q", "r", "s", + "t", "u", "v", "w", "x", "y", "z", "braceleft", "bar", "braceright", "asciitilde", "exclamdown", + "cent", "sterling", "fraction", "yen", "florin", "section", "currency", "quotesingle", + "quotedblleft", "guillemotleft", "guilsinglleft", "guilsinglright", "fi", "fl", "endash", + "dagger", "daggerdbl", "periodcentered", "paragraph", "bullet", "quotesinglbase", + "quotedblbase", "quotedblright", "guillemotright", "ellipsis", "perthousand", "questiondown", + "grave", "acute", "circumflex", "tilde", "macron", "breve", "dotaccent", "dieresis", "ring", + "cedilla", "hungarumlaut", "ogonek", "caron", "emdash", "AE", "ordfeminine", "Lslash", "Oslash", + "OE", "ordmasculine", "ae", "dotlessi", "lslash", "oslash", "oe", "germandbls", "onesuperior", + "logicalnot", "mu", "trademark", "Eth", "onehalf", "plusminus", "Thorn", "onequarter", "divide", + "brokenbar", "degree", "thorn", "threequarters", "twosuperior", "registered", "minus", "eth", + "multiply", "threesuperior", "copyright", "Aacute", "Acircumflex", "Adieresis", "Agrave", + "Aring", "Atilde", "Ccedilla", "Eacute", "Ecircumflex", "Edieresis", "Egrave", "Iacute", + "Icircumflex", "Idieresis", "Igrave", "Ntilde", "Oacute", "Ocircumflex", "Odieresis", "Ograve", + "Otilde", "Scaron", "Uacute", "Ucircumflex", "Udieresis", "Ugrave", "Yacute", "Ydieresis", + "Zcaron", "aacute", "acircumflex", "adieresis", "agrave", "aring", "atilde", "ccedilla", + "eacute", "ecircumflex", "edieresis", "egrave", "iacute", "icircumflex", "idieresis", "igrave", + "ntilde", "oacute", "ocircumflex", "odieresis", "ograve", "otilde", "scaron", "uacute", + "ucircumflex", "udieresis", "ugrave", "yacute", "ydieresis", "zcaron", "exclamsmall", + "Hungarumlautsmall", "dollaroldstyle", "dollarsuperior", "ampersandsmall", "Acutesmall", + "parenleftsuperior", "parenrightsuperior", "twodotenleader", "onedotenleader", "zerooldstyle", + "oneoldstyle", "twooldstyle", "threeoldstyle", "fouroldstyle", "fiveoldstyle", "sixoldstyle", + "sevenoldstyle", "eightoldstyle", "nineoldstyle", "commasuperior", "threequartersemdash", + "periodsuperior", "questionsmall", "asuperior", "bsuperior", "centsuperior", "dsuperior", + "esuperior", "isuperior", "lsuperior", "msuperior", "nsuperior", "osuperior", "rsuperior", + "ssuperior", "tsuperior", "ff", "ffi", "ffl", "parenleftinferior", "parenrightinferior", + "Circumflexsmall", "hyphensuperior", "Gravesmall", "Asmall", "Bsmall", "Csmall", "Dsmall", + "Esmall", "Fsmall", "Gsmall", "Hsmall", "Ismall", "Jsmall", "Ksmall", "Lsmall", "Msmall", + "Nsmall", "Osmall", "Psmall", "Qsmall", "Rsmall", "Ssmall", "Tsmall", "Usmall", "Vsmall", + "Wsmall", "Xsmall", "Ysmall", "Zsmall", "colonmonetary", "onefitted", "rupiah", "Tildesmall", + "exclamdownsmall", "centoldstyle", "Lslashsmall", "Scaronsmall", "Zcaronsmall", "Dieresissmall", + "Brevesmall", "Caronsmall", "Dotaccentsmall", "Macronsmall", "figuredash", "hypheninferior", + "Ogoneksmall", "Ringsmall", "Cedillasmall", "questiondownsmall", "oneeighth", "threeeighths", + "fiveeighths", "seveneighths", "onethird", "twothirds", "zerosuperior", "foursuperior", + "fivesuperior", "sixsuperior", "sevensuperior", "eightsuperior", "ninesuperior", "zeroinferior", + "oneinferior", "twoinferior", "threeinferior", "fourinferior", "fiveinferior", "sixinferior", + "seveninferior", "eightinferior", "nineinferior", "centinferior", "dollarinferior", + "periodinferior", "commainferior", "Agravesmall", "Aacutesmall", "Acircumflexsmall", + "Atildesmall", "Adieresissmall", "Aringsmall", "AEsmall", "Ccedillasmall", "Egravesmall", + "Eacutesmall", "Ecircumflexsmall", "Edieresissmall", "Igravesmall", "Iacutesmall", + "Icircumflexsmall", "Idieresissmall", "Ethsmall", "Ntildesmall", "Ogravesmall", "Oacutesmall", + "Ocircumflexsmall", "Otildesmall", "Odieresissmall", "OEsmall", "Oslashsmall", "Ugravesmall", + "Uacutesmall", "Ucircumflexsmall", "Udieresissmall", "Yacutesmall", "Thornsmall", + "Ydieresissmall", "001.000", "001.001", "001.002", "001.003", "Black", "Bold", "Book", "Light", + "Medium", "Regular", "Roman", "Semibold", +]; + +/// The CFF Standard Encoding, code → SID (Adobe TN #5176 Appendix B; ttf-parser 0.25.1 +/// `encoding.rs` `STANDARD_ENCODING`), used by `seac`. +#[rustfmt::skip] +const STANDARD_ENCODING: [u8; 256] = [ + 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, + 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, + 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, + 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, + 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, + 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, + 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, + 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 0, + 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, + 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, + 0, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, + 0, 111, 112, 113, 114, 0, 115, 116, 117, 118, 119, 120, 121, 122, 0, 123, + 0, 124, 125, 126, 127, 128, 129, 130, 131, 0, 132, 133, 0, 134, 135, 136, + 137, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, + 0, 138, 0, 139, 0, 0, 0, 0, 140, 141, 142, 143, 0, 0, 0, 0, + 0, 144, 0, 0, 0, 145, 0, 0, 146, 147, 148, 149, 0, 0, 0, 0, +]; + +/// Operands ttf-parser keeps per DICT operator. +const DICT_OPERANDS_MAX: usize = 48; +/// Longest Top, Font or Private DICT followed (real ones are well under 1 KiB). +pub(crate) const DICT_LEN_MAX: usize = 16 << 10; +const OP_CHARSET: u16 = 15; +const OP_CHAR_STRINGS: u16 = 17; +const OP_PRIVATE: u16 = 18; +const OP_LOCAL_SUBRS: u16 = 19; +const OP_ROS: u16 = 1230; +const OP_FD_ARRAY: u16 = 1236; +const OP_FD_SELECT: u16 = 1237; + +fn u16_at(data: &[u8], at: usize) -> Option { + let bytes = data.get(at..at.checked_add(2)?)?; + Some(u16::from_be_bytes([*bytes.first()?, *bytes.get(1)?])) +} + +/// A CFF INDEX as ttf-parser's `parse_index::` builds it: offsets are 1-based and the data +/// runs to the last offset; a last offset of 0 makes an empty INDEX. +#[derive(Clone, Copy, Default)] +pub(crate) struct Index<'a> { + data: &'a [u8], + offsets: &'a [u8], + off_size: usize, +} + +impl<'a> Index<'a> { + /// The INDEX at `at` and the offset right after it. + fn parse(cff: &'a [u8], at: usize) -> Option<(Index<'a>, usize)> { + let (offsets, off_size, pos) = match Self::header(cff, at)? { + Header::Empty(pos) => return Some((Index::default(), pos)), + Header::Offsets(offsets, off_size, pos) => (offsets, off_size, pos), + }; + let shell = Index { + data: &[], + offsets, + off_size, + }; + match shell.offset(shell.offset_count().checked_sub(1)?) { + Some(last) => { + let end = pos.checked_add(last)?; + let data = cff.get(pos..end)?; + Some((Index { data, ..shell }, end)) + } + None => Some((Index::default(), pos)), + } + } + + /// The offset after the INDEX at `at` (`skip_index`: the data is not bounds-checked). + fn skip(cff: &'a [u8], at: usize) -> Option { + match Self::header(cff, at)? { + Header::Empty(pos) => Some(pos), + Header::Offsets(offsets, off_size, pos) => { + let shell = Index { + data: &[], + offsets, + off_size, + }; + match shell.offset(shell.offset_count().checked_sub(1)?) { + Some(last) => pos.checked_add(last), + None => Some(pos), + } + } + } + } + + fn header(cff: &'a [u8], at: usize) -> Option> { + let count = u16_at(cff, at)?; + let pos = at.checked_add(2)?; + if count == 0 { + return Some(Header::Empty(pos)); + } + let off_size = usize::from(*cff.get(pos)?); + if !(1..=4).contains(&off_size) { + return None; + } + let pos = pos.checked_add(1)?; + let len = (usize::from(count) + 1).checked_mul(off_size)?; + let end = pos.checked_add(len)?; + Some(Header::Offsets(cff.get(pos..end)?, off_size, end)) + } + + fn offset_count(&self) -> usize { + self.offsets.len().checked_div(self.off_size).unwrap_or(0) + } + + /// Offset `i` minus one (`None` past the end or for a raw 0). + fn offset(&self, i: usize) -> Option { + if i >= self.offset_count() { + return None; + } + let at = i.checked_mul(self.off_size)?; + let raw = self + .offsets + .get(at..at.checked_add(self.off_size)?)? + .iter() + .fold(0usize, |acc, b| (acc << 8) | usize::from(*b)); + raw.checked_sub(1) + } + + /// Number of entries. + pub fn len(&self) -> u32 { + u32::try_from(self.offset_count().saturating_sub(1)).unwrap_or(u32::MAX) + } + + pub fn get(&self, i: u32) -> Option<&'a [u8]> { + let i = usize::try_from(i).ok()?; + let start = self.offset(i)?; + let end = self.offset(i.checked_add(1)?)?; + self.data.get(start..end) + } +} + +enum Header<'a> { + Empty(usize), + Offsets(&'a [u8], usize, usize), +} + +/// One DICT entry: the operator (two-byte operators as `1200 + b1`) and its operand bytes. +struct DictEntry<'a> { + op: u16, + operands: &'a [u8], +} + +/// DICT entries in order, as `DictionaryParser::parse_next` finds them (operators 0–27, 31, 255; +/// the walk stops at the first number it cannot skip). Lazy: nothing is collected. +struct DictEntries<'a> { + data: &'a [u8], + pos: usize, +} + +impl<'a> DictEntries<'a> { + fn new(data: &'a [u8]) -> DictEntries<'a> { + DictEntries { data, pos: 0 } + } +} + +impl<'a> Iterator for DictEntries<'a> { + type Item = DictEntry<'a>; + + fn next(&mut self) -> Option> { + let start = self.pos; + while let Some(&b) = self.data.get(self.pos) { + self.pos += 1; + if matches!(b, 0..=27 | 31 | 255) { + let op = if b == 12 { + let Some(&b1) = self.data.get(self.pos) else { + self.pos = self.data.len(); + return None; + }; + self.pos += 1; + 1200 + u16::from(b1) + } else { + u16::from(b) + }; + let operands = self.data.get(start..self.pos).unwrap_or_default(); + return Some(DictEntry { op, operands }); + } + match skip_number(b, self.data, self.pos) { + Some(next) => self.pos = next, + None => { + self.pos = self.data.len(); + return None; + } + } + } + None + } +} + +/// `skip_number`: the offset after the operand starting with `b0` (unchecked advances, as there). +fn skip_number(b0: u8, data: &[u8], pos: usize) -> Option { + match b0 { + 28 => pos.checked_add(2), + 29 => pos.checked_add(4), + 30 => { + let mut pos = pos; + while let Some(&b) = data.get(pos) { + pos += 1; + if b >> 4 == 0xF || b & 0xF == 0xF { + break; + } + } + Some(pos) + } + 32..=246 => Some(pos), + 247..=254 => pos.checked_add(1), + _ => None, + } +} + +/// The integer operands of an entry (`parse_operands`, at most 48). `Err` for a real operand, +/// which this reader does not evaluate; `Ok(None)` where ttf-parser's parse fails. +fn int_operands(entry: &DictEntry<'_>) -> Result>, ()> { + let data = entry.operands; + let mut out = Vec::new(); + let mut pos = 0usize; + while let Some(&b0) = data.get(pos) { + pos += 1; + if matches!(b0, 0..=27 | 31 | 255) { + break; + } + let at = pos; + let value = match b0 { + 28 => { + pos += 2; + u16_at(data, at).map(|v| i64::from(v as i16)) + } + 29 => { + pos += 4; + u16_at(data, at) + .zip(u16_at(data, at + 2)) + .map(|(hi, lo)| i64::from(((u32::from(hi) << 16) | u32::from(lo)) as i32)) + } + 30 => return Err(()), + 32..=246 => Some(i64::from(b0) - 139), + 247..=250 => { + pos += 1; + data.get(at) + .map(|b1| (i64::from(b0) - 247) * 256 + i64::from(*b1) + 108) + } + 251..=254 => { + pos += 1; + data.get(at) + .map(|b1| -(i64::from(b0) - 251) * 256 - i64::from(*b1) - 108) + } + _ => None, + }; + let Some(value) = value else { + return Ok(None); + }; + out.push(value); + if out.len() >= DICT_OPERANDS_MAX { + break; + } + } + Ok(Some(out)) +} + +/// `parse_offset`: exactly one operand, as a non-negative `i32`. +fn dict_offset(entry: &DictEntry<'_>) -> Result, ()> { + Ok(int_operands(entry)?.and_then(|ops| match ops.as_slice() { + [v] => i32::try_from(*v).ok().and_then(|v| usize::try_from(v).ok()), + _ => None, + })) +} + +/// `parse_range`: `size offset` → `offset..offset + size`. +fn dict_range(entry: &DictEntry<'_>) -> Result>, ()> { + Ok(int_operands(entry)?.and_then(|ops| match ops.as_slice() { + [len, start] => { + let len = usize::try_from(i32::try_from(*len).ok()?).ok()?; + let start = usize::try_from(i32::try_from(*start).ok()?).ok()?; + Some(start..start.checked_add(len)?) + } + _ => None, + })) +} + +/// The charset as ttf-parser parses it. +#[derive(Clone, Copy)] +pub(crate) enum Charset<'a> { + IsoAdobe, + /// Expert or Expert Subset (predefined; no `seac` lookup, names through ttf-parser). + Expert, + Format0(&'a [u8]), + /// Ranges of `(first SID, nLeft)`; `wide` = format 2 (u16 nLeft). + Ranges { + data: &'a [u8], + wide: bool, + }, +} + +/// CID-keyed fonts: the FDArray and FDSelect. +#[derive(Clone, Copy)] +struct CidParts<'a> { + fd_array: Index<'a>, + fd_select: FdSelect<'a>, +} + +#[derive(Clone, Copy)] +enum FdSelect<'a> { + Format0(&'a [u8]), + Format3(&'a [u8]), +} + +/// Where a glyph's local subroutines come from. +pub(crate) enum LocalSubrs<'a> { + /// Name-keyed fonts: the Private DICT's (possibly empty), laid out once. + Font(Index<'a>), + /// CID-keyed fonts: the Font DICT FDSelect selects (`None`: none), and the FDSelect ranges + /// walked to find it. + FontDict(Option, u64), +} + +/// A CID-keyed font's Font DICT: its local subroutines and the DICT bytes walked to find them +/// (Font DICT + Private DICT), which ttf-parser walks again for every glyph it outlines that +/// calls a local subroutine. +#[derive(Clone, Copy)] +pub(crate) struct FdSubrs<'a> { + pub subrs: Option>, + pub dict_bytes: u64, +} + +/// What a CFF program's glyphs run: charstrings, global and local subroutines, and the charset. +pub(crate) struct CffLayout<'a> { + data: &'a [u8], + pub char_strings: Index<'a>, + pub global_subrs: Index<'a>, + strings: Index<'a>, + pub charset: Charset<'a>, + /// Name-keyed fonts: the Private DICT's local subroutines (possibly empty). + sid_local_subrs: Index<'a>, + cid: Option>, + pub number_of_glyphs: u16, +} + +impl<'a> CffLayout<'a> { + /// Mirrors `cff::Table::parse`; `None` where it fails or where this reader cannot follow. + pub fn parse(data: &'a [u8]) -> Option> { + if data.first() != Some(&1) { + return None; + } + let header_size = usize::from(*data.get(2)?); + let mut pos = 4usize; + if header_size > 4 { + pos = pos.checked_add(header_size - 4)?; + } + pos = Index::skip(data, pos)?; // Name INDEX + let (top_index, after_top) = Index::parse(data, pos)?; + let top = TopDict::parse(top_index.get(0)?)?; + if top.char_strings == 0 { + return None; + } + let (strings, after_strings) = Index::parse(data, after_top)?; + let (global_subrs, _) = Index::parse(data, after_strings)?; + if top.char_strings > data.len() { + return None; + } + let (char_strings, _) = Index::parse(data, top.char_strings)?; + let number_of_glyphs = u16::try_from(char_strings.len()).ok().filter(|n| *n > 0)?; + let charset = match top.charset { + Some(0) | None => Charset::IsoAdobe, + Some(1 | 2) => Charset::Expert, + Some(at) => parse_charset(data, at, number_of_glyphs)?, + }; + let mut layout = CffLayout { + data, + char_strings, + global_subrs, + strings, + charset, + sid_local_subrs: Index::default(), + cid: None, + number_of_glyphs, + }; + if top.ros { + let (Some(charset_at), Some(fd_array_at), Some(fd_select_at)) = + (top.charset, top.fd_array, top.fd_select) + else { + return None; + }; + if charset_at <= 2 || fd_array_at > data.len() || fd_select_at > data.len() { + return None; + } + let (fd_array, _) = Index::parse(data, fd_array_at)?; + let fd_select = match *data.get(fd_select_at)? { + 0 => { + let start = fd_select_at.checked_add(1)?; + let end = start.checked_add(usize::from(number_of_glyphs))?; + FdSelect::Format0(data.get(start..end)?) + } + 3 => FdSelect::Format3(data.get(fd_select_at.checked_add(1)?..)?), + _ => return None, + }; + layout.cid = Some(CidParts { + fd_array, + fd_select, + }); + } else if let Some(range) = top.private { + let private = PrivateDict::parse(data.get(range.clone())?)?; + if let Some(start) = private + .local_subrs + .and_then(|off| range.start.checked_add(off)) + { + layout.sid_local_subrs = Index::parse(data.get(start..)?, 0)?.0; + } + } + Some(layout) + } + + pub fn is_cid(&self) -> bool { + self.cid.is_some() + } + + /// Where the local subroutines `glyph` runs come from. + pub fn local_subrs(&self, glyph: u16) -> LocalSubrs<'a> { + match self.cid { + None => LocalSubrs::Font(self.sid_local_subrs), + Some(cid) => { + let (fd, walked) = cid.fd_select.font_dict_index(glyph); + LocalSubrs::FontDict(fd, walked) + } + } + } + + /// The local subroutines of Font DICT `fd` as `parse_cid_local_subrs` finds them (Font DICT + /// → its first `Private` entry → the Private DICT's `Subrs`), and the DICT bytes walked. + pub fn fd_local_subrs(&self, fd: u8) -> FdSubrs<'a> { + let mut out = FdSubrs { + subrs: None, + dict_bytes: 0, + }; + let Some(font_dict) = self + .cid + .and_then(|cid| cid.fd_array.get(u32::from(fd))) + .filter(|d| d.len() <= DICT_LEN_MAX) + else { + return out; + }; + out.dict_bytes = font_dict.len() as u64; + let Some(range) = DictEntries::new(font_dict) + .find(|e| e.op == OP_PRIVATE) + .and_then(|e| dict_range(&e).ok().flatten()) + else { + return out; + }; + let Some(private_data) = self.data.get(range.clone()) else { + return out; + }; + let Some(private) = PrivateDict::parse(private_data) else { + return out; // past DICT_LEN_MAX (not walked), or an offset this reader cannot follow + }; + out.dict_bytes = out.dict_bytes.saturating_add(private_data.len() as u64); + out.subrs = private + .local_subrs + .and_then(|off| range.start.checked_add(off)) + .and_then(|start| Index::parse(self.data.get(start..)?, 0)) + .map(|(index, _)| index); + out + } + + /// GID → SID for every glyph, in one walk of the charset (`None` for the predefined Expert + /// charsets, whose lookups ttf-parser answers in O(1) anyway). + pub fn gid_to_sid(&self) -> Option>> { + let n = usize::from(self.number_of_glyphs); + let mut out: Vec> = Vec::with_capacity(n); + out.push(Some(0)); + match self.charset { + Charset::IsoAdobe => { + out.extend((1..n).map(|gid| u16::try_from(gid).ok().filter(|g| *g <= 228))); + } + Charset::Expert => return None, + Charset::Format0(sids) => { + out.extend(sids.chunks_exact(2).map(|p| { + p.first() + .zip(p.get(1)) + .map(|(h, l)| u16::from_be_bytes([*h, *l])) + })); + } + Charset::Ranges { data, wide } => { + let step = if wide { 4 } else { 3 }; + for range in data.chunks_exact(step) { + let first = u16_at(range, 0)?; + let left = if wide { + u16_at(range, 2)? + } else { + u16::from(*range.get(2)?) + }; + for k in 0..=left { + if out.len() >= n { + break; + } + out.push(first.checked_add(k)); + } + } + } + } + out.resize(n, None); + Some(out) + } + + /// The units ttf-parser's `sid_to_gid` walk costs once (format 0 scans every SID, formats + /// 1/2 every range). + pub fn charset_walk(&self) -> u64 { + match self.charset { + Charset::IsoAdobe | Charset::Expert => 1, + Charset::Format0(sids) => (sids.len() / 2) as u64 + 1, + Charset::Ranges { data, wide } => (data.len() / if wide { 4 } else { 3 }) as u64 + 1, + } + } + + /// The glyph name of SID `sid` (standard strings, then the String INDEX). + pub fn sid_name(&self, sid: u16) -> Option<&'a str> { + match STANDARD_STRINGS.get(usize::from(sid)) { + Some(name) => Some(name), + None => { + let index = u32::from(sid).checked_sub(STANDARD_STRINGS.len() as u32)?; + std::str::from_utf8(self.strings.get(index)?).ok() + } + } + } +} + +/// `seac`'s code → GID (`seac_code_to_glyph_id`): the Standard Encoding SID, then the charset. +pub(crate) fn seac_glyph( + charset: Charset<'_>, + sid_to_gid: &HashMap, + code: u8, +) -> Option { + let sid = u16::from(*STANDARD_ENCODING.get(usize::from(code))?); + match charset { + Charset::IsoAdobe => (code <= 228).then_some(sid), + Charset::Expert => None, + Charset::Format0(_) | Charset::Ranges { .. } if sid == 0 => Some(0), + Charset::Format0(_) | Charset::Ranges { .. } => sid_to_gid.get(&sid).copied(), + } +} + +/// The Top DICT entries ttf-parser reads (later entries win). +struct TopDict { + charset: Option, + char_strings: usize, + private: Option>, + ros: bool, + fd_array: Option, + fd_select: Option, +} + +impl TopDict { + /// `None` past `DICT_LEN_MAX` or where an offset cannot be followed. + fn parse(data: &[u8]) -> Option { + if data.len() > DICT_LEN_MAX { + return None; + } + let mut top = TopDict { + charset: None, + char_strings: 0, + private: None, + ros: false, + fd_array: None, + fd_select: None, + }; + for entry in DictEntries::new(data) { + match entry.op { + OP_CHARSET => top.charset = dict_offset(&entry).ok()?, + OP_CHAR_STRINGS => top.char_strings = dict_offset(&entry).ok()??, + OP_PRIVATE => top.private = dict_range(&entry).ok()?, + OP_ROS => top.ros = true, + OP_FD_ARRAY => top.fd_array = dict_offset(&entry).ok()?, + OP_FD_SELECT => top.fd_select = dict_offset(&entry).ok()?, + _ => {} + } + } + Some(top) + } +} + +/// The Private DICT's `Subrs` offset (relative to the Private DICT; later entries win). +struct PrivateDict { + local_subrs: Option, +} + +impl PrivateDict { + /// `None` past `DICT_LEN_MAX` or where the `Subrs` offset cannot be followed. + fn parse(data: &[u8]) -> Option { + if data.len() > DICT_LEN_MAX { + return None; + } + let mut local_subrs = None; + for entry in DictEntries::new(data) { + if entry.op == OP_LOCAL_SUBRS { + local_subrs = dict_offset(&entry).ok()?; + } + } + Some(PrivateDict { local_subrs }) + } +} + +/// `parse_charset`: format 0 (n − 1 SIDs) or ranges covering exactly n − 1 glyphs. +fn parse_charset(data: &[u8], at: usize, glyphs: u16) -> Option> { + let body = at.checked_add(1)?; + let wanted = glyphs.checked_sub(1)?; + match *data.get(at)? { + 0 => { + let end = body.checked_add(usize::from(wanted).checked_mul(2)?)?; + Some(Charset::Format0(data.get(body..end)?)) + } + format @ (1 | 2) => { + let wide = format == 2; + let step = if wide { 4 } else { 3 }; + let (mut left_total, mut count, mut pos) = (wanted, 0usize, body); + while left_total > 0 { + let left = if wide { + u16_at(data, pos.checked_add(2)?)?.checked_add(1)? + } else { + u16::from(*data.get(pos.checked_add(2)?)?) + 1 + }; + left_total = left_total.checked_sub(left)?; + count += 1; + pos = pos.checked_add(step)?; + } + let end = body.checked_add(count.checked_mul(step)?)?; + Some(Charset::Ranges { + data: data.get(body..end)?, + wide, + }) + } + _ => None, + } +} + +impl FdSelect<'_> { + /// The Font DICT index of `glyph` and the ranges walked to find it. + fn font_dict_index(&self, glyph: u16) -> (Option, u64) { + match self { + FdSelect::Format0(array) => (array.get(usize::from(glyph)).copied(), 1), + FdSelect::Format3(data) => { + let mut walked = 1u64; + let found = (|| { + let ranges = u16_at(data, 0)?; + if ranges == 0 { + return None; + } + // The sentinel GID closes the last range (ttf-parser counts it as one more). + let bound = ranges.checked_add(1)?; + let mut prev_first = u16_at(data, 2)?; + let mut prev_index = *data.get(4)?; + let mut pos = 5usize; + for _ in 1..bound { + walked += 1; + let first = u16_at(data, pos)?; + if (prev_first..first).contains(&glyph) { + return Some(prev_index); + } + prev_index = *data.get(pos.checked_add(2)?)?; + prev_first = first; + pos = pos.checked_add(3)?; + } + None + })(); + (found, walked) + } + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/encoding_tables.rs b/src-tauri/src/pdf_engine/text_edit/fonts/encoding_tables.rs new file mode 100644 index 0000000..db8d2e8 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/encoding_tables.rs @@ -0,0 +1,129 @@ +//! GENERATED DATA — do not edit by hand (SPEC §B.9.2). Base encodings as glyph names per code. +//! +//! Provenance: generated from `lopdf` 0.34.0 `src/encodings/mappings.rs` (`WIN_ANSI_ENCODING`, +//! `MAC_ROMAN_ENCODING`, `STANDARD_ENCODING`; each `Some(Glyph::name)` entry written as its glyph +//! name) and checked, code by code, to be identical to pdf.js 4.10.38 (`pdfjs-dist`, +//! `build/pdf.worker.mjs`: `WinAnsiEncoding`, `MacRomanEncoding`, `StandardEncoding`), which follow +//! ISO 32000-1 Annex D.2 ("Latin Character Set and Encodings", incl. its notes: WinAnsi 0xA0 = +//! space, 0xAD = hyphen, unused WinAnsi codes above 0x20 = bullet; MacRoman 0xCA = space). +//! `MAC_ROMAN_NAMES` is strict Annex D: the 15 Mac-OS-only codes that lopdf and pdf.js add +//! (0xAD notequal, 0xB0 infinity, 0xB2 lessequal, 0xB3 greaterequal, 0xB6 partialdiff, +//! 0xB7 summation, 0xB8 product, 0xB9 pi, 0xBA integral, 0xBD Omega, 0xC3 radical, +//! 0xC5 approxequal, 0xC6 Delta, 0xD7 lozenge, 0xF0 apple) are `None`; 0xDB is `currency`. +//! The test `encoding_tables_cross_check` (FONT-01/FONT-37) re-derives every code from lopdf's +//! Unicode tables through the AGL and asserts exactly that 15-code difference. +//! +//! Licences: lopdf is MIT (Copyright (c) 2016 Junfeng Liu); pdf.js is Apache-2.0 (Copyright +//! Mozilla Foundation), used only for the cross-check; the encodings themselves are tables of +//! ISO 32000-1 (Annex D). + +/// WinAnsiEncoding (ISO 32000-1 Annex D.2). +#[rustfmt::skip] +pub(super) static WIN_ANSI_NAMES: [Option<&str>; 256] = [ + /* 0x00 */ None, None, None, None, None, None, None, None, + /* 0x08 */ None, None, None, None, None, None, None, None, + /* 0x10 */ None, None, None, None, None, None, None, None, + /* 0x18 */ None, None, None, None, None, None, None, None, + /* 0x20 */ Some("space"), Some("exclam"), Some("quotedbl"), Some("numbersign"), Some("dollar"), Some("percent"), Some("ampersand"), Some("quotesingle"), + /* 0x28 */ Some("parenleft"), Some("parenright"), Some("asterisk"), Some("plus"), Some("comma"), Some("hyphen"), Some("period"), Some("slash"), + /* 0x30 */ Some("zero"), Some("one"), Some("two"), Some("three"), Some("four"), Some("five"), Some("six"), Some("seven"), + /* 0x38 */ Some("eight"), Some("nine"), Some("colon"), Some("semicolon"), Some("less"), Some("equal"), Some("greater"), Some("question"), + /* 0x40 */ Some("at"), Some("A"), Some("B"), Some("C"), Some("D"), Some("E"), Some("F"), Some("G"), + /* 0x48 */ Some("H"), Some("I"), Some("J"), Some("K"), Some("L"), Some("M"), Some("N"), Some("O"), + /* 0x50 */ Some("P"), Some("Q"), Some("R"), Some("S"), Some("T"), Some("U"), Some("V"), Some("W"), + /* 0x58 */ Some("X"), Some("Y"), Some("Z"), Some("bracketleft"), Some("backslash"), Some("bracketright"), Some("asciicircum"), Some("underscore"), + /* 0x60 */ Some("grave"), Some("a"), Some("b"), Some("c"), Some("d"), Some("e"), Some("f"), Some("g"), + /* 0x68 */ Some("h"), Some("i"), Some("j"), Some("k"), Some("l"), Some("m"), Some("n"), Some("o"), + /* 0x70 */ Some("p"), Some("q"), Some("r"), Some("s"), Some("t"), Some("u"), Some("v"), Some("w"), + /* 0x78 */ Some("x"), Some("y"), Some("z"), Some("braceleft"), Some("bar"), Some("braceright"), Some("asciitilde"), Some("bullet"), + /* 0x80 */ Some("Euro"), Some("bullet"), Some("quotesinglbase"), Some("florin"), Some("quotedblbase"), Some("ellipsis"), Some("dagger"), Some("daggerdbl"), + /* 0x88 */ Some("circumflex"), Some("perthousand"), Some("Scaron"), Some("guilsinglleft"), Some("OE"), Some("bullet"), Some("Zcaron"), Some("bullet"), + /* 0x90 */ Some("bullet"), Some("quoteleft"), Some("quoteright"), Some("quotedblleft"), Some("quotedblright"), Some("bullet"), Some("endash"), Some("emdash"), + /* 0x98 */ Some("tilde"), Some("trademark"), Some("scaron"), Some("guilsinglright"), Some("oe"), Some("bullet"), Some("zcaron"), Some("Ydieresis"), + /* 0xA0 */ Some("space"), Some("exclamdown"), Some("cent"), Some("sterling"), Some("currency"), Some("yen"), Some("brokenbar"), Some("section"), + /* 0xA8 */ Some("dieresis"), Some("copyright"), Some("ordfeminine"), Some("guillemotleft"), Some("logicalnot"), Some("hyphen"), Some("registered"), Some("macron"), + /* 0xB0 */ Some("degree"), Some("plusminus"), Some("twosuperior"), Some("threesuperior"), Some("acute"), Some("mu"), Some("paragraph"), Some("periodcentered"), + /* 0xB8 */ Some("cedilla"), Some("onesuperior"), Some("ordmasculine"), Some("guillemotright"), Some("onequarter"), Some("onehalf"), Some("threequarters"), Some("questiondown"), + /* 0xC0 */ Some("Agrave"), Some("Aacute"), Some("Acircumflex"), Some("Atilde"), Some("Adieresis"), Some("Aring"), Some("AE"), Some("Ccedilla"), + /* 0xC8 */ Some("Egrave"), Some("Eacute"), Some("Ecircumflex"), Some("Edieresis"), Some("Igrave"), Some("Iacute"), Some("Icircumflex"), Some("Idieresis"), + /* 0xD0 */ Some("Eth"), Some("Ntilde"), Some("Ograve"), Some("Oacute"), Some("Ocircumflex"), Some("Otilde"), Some("Odieresis"), Some("multiply"), + /* 0xD8 */ Some("Oslash"), Some("Ugrave"), Some("Uacute"), Some("Ucircumflex"), Some("Udieresis"), Some("Yacute"), Some("Thorn"), Some("germandbls"), + /* 0xE0 */ Some("agrave"), Some("aacute"), Some("acircumflex"), Some("atilde"), Some("adieresis"), Some("aring"), Some("ae"), Some("ccedilla"), + /* 0xE8 */ Some("egrave"), Some("eacute"), Some("ecircumflex"), Some("edieresis"), Some("igrave"), Some("iacute"), Some("icircumflex"), Some("idieresis"), + /* 0xF0 */ Some("eth"), Some("ntilde"), Some("ograve"), Some("oacute"), Some("ocircumflex"), Some("otilde"), Some("odieresis"), Some("divide"), + /* 0xF8 */ Some("oslash"), Some("ugrave"), Some("uacute"), Some("ucircumflex"), Some("udieresis"), Some("yacute"), Some("thorn"), Some("ydieresis"), +]; + +/// MacRomanEncoding, strictly as ISO 32000-1 Annex D.2 (no Mac-OS-only codes). +#[rustfmt::skip] +pub(super) static MAC_ROMAN_NAMES: [Option<&str>; 256] = [ + /* 0x00 */ None, None, None, None, None, None, None, None, + /* 0x08 */ None, None, None, None, None, None, None, None, + /* 0x10 */ None, None, None, None, None, None, None, None, + /* 0x18 */ None, None, None, None, None, None, None, None, + /* 0x20 */ Some("space"), Some("exclam"), Some("quotedbl"), Some("numbersign"), Some("dollar"), Some("percent"), Some("ampersand"), Some("quotesingle"), + /* 0x28 */ Some("parenleft"), Some("parenright"), Some("asterisk"), Some("plus"), Some("comma"), Some("hyphen"), Some("period"), Some("slash"), + /* 0x30 */ Some("zero"), Some("one"), Some("two"), Some("three"), Some("four"), Some("five"), Some("six"), Some("seven"), + /* 0x38 */ Some("eight"), Some("nine"), Some("colon"), Some("semicolon"), Some("less"), Some("equal"), Some("greater"), Some("question"), + /* 0x40 */ Some("at"), Some("A"), Some("B"), Some("C"), Some("D"), Some("E"), Some("F"), Some("G"), + /* 0x48 */ Some("H"), Some("I"), Some("J"), Some("K"), Some("L"), Some("M"), Some("N"), Some("O"), + /* 0x50 */ Some("P"), Some("Q"), Some("R"), Some("S"), Some("T"), Some("U"), Some("V"), Some("W"), + /* 0x58 */ Some("X"), Some("Y"), Some("Z"), Some("bracketleft"), Some("backslash"), Some("bracketright"), Some("asciicircum"), Some("underscore"), + /* 0x60 */ Some("grave"), Some("a"), Some("b"), Some("c"), Some("d"), Some("e"), Some("f"), Some("g"), + /* 0x68 */ Some("h"), Some("i"), Some("j"), Some("k"), Some("l"), Some("m"), Some("n"), Some("o"), + /* 0x70 */ Some("p"), Some("q"), Some("r"), Some("s"), Some("t"), Some("u"), Some("v"), Some("w"), + /* 0x78 */ Some("x"), Some("y"), Some("z"), Some("braceleft"), Some("bar"), Some("braceright"), Some("asciitilde"), None, + /* 0x80 */ Some("Adieresis"), Some("Aring"), Some("Ccedilla"), Some("Eacute"), Some("Ntilde"), Some("Odieresis"), Some("Udieresis"), Some("aacute"), + /* 0x88 */ Some("agrave"), Some("acircumflex"), Some("adieresis"), Some("atilde"), Some("aring"), Some("ccedilla"), Some("eacute"), Some("egrave"), + /* 0x90 */ Some("ecircumflex"), Some("edieresis"), Some("iacute"), Some("igrave"), Some("icircumflex"), Some("idieresis"), Some("ntilde"), Some("oacute"), + /* 0x98 */ Some("ograve"), Some("ocircumflex"), Some("odieresis"), Some("otilde"), Some("uacute"), Some("ugrave"), Some("ucircumflex"), Some("udieresis"), + /* 0xA0 */ Some("dagger"), Some("degree"), Some("cent"), Some("sterling"), Some("section"), Some("bullet"), Some("paragraph"), Some("germandbls"), + /* 0xA8 */ Some("registered"), Some("copyright"), Some("trademark"), Some("acute"), Some("dieresis"), None, Some("AE"), Some("Oslash"), + /* 0xB0 */ None, Some("plusminus"), None, None, Some("yen"), Some("mu"), None, None, + /* 0xB8 */ None, None, None, Some("ordfeminine"), Some("ordmasculine"), None, Some("ae"), Some("oslash"), + /* 0xC0 */ Some("questiondown"), Some("exclamdown"), Some("logicalnot"), None, Some("florin"), None, None, Some("guillemotleft"), + /* 0xC8 */ Some("guillemotright"), Some("ellipsis"), Some("space"), Some("Agrave"), Some("Atilde"), Some("Otilde"), Some("OE"), Some("oe"), + /* 0xD0 */ Some("endash"), Some("emdash"), Some("quotedblleft"), Some("quotedblright"), Some("quoteleft"), Some("quoteright"), Some("divide"), None, + /* 0xD8 */ Some("ydieresis"), Some("Ydieresis"), Some("fraction"), Some("currency"), Some("guilsinglleft"), Some("guilsinglright"), Some("fi"), Some("fl"), + /* 0xE0 */ Some("daggerdbl"), Some("periodcentered"), Some("quotesinglbase"), Some("quotedblbase"), Some("perthousand"), Some("Acircumflex"), Some("Ecircumflex"), Some("Aacute"), + /* 0xE8 */ Some("Edieresis"), Some("Egrave"), Some("Iacute"), Some("Icircumflex"), Some("Idieresis"), Some("Igrave"), Some("Oacute"), Some("Ocircumflex"), + /* 0xF0 */ None, Some("Ograve"), Some("Uacute"), Some("Ucircumflex"), Some("Ugrave"), Some("dotlessi"), Some("circumflex"), Some("tilde"), + /* 0xF8 */ Some("macron"), Some("breve"), Some("dotaccent"), Some("ring"), Some("cedilla"), Some("hungarumlaut"), Some("ogonek"), Some("caron"), +]; + +/// StandardEncoding (ISO 32000-1 Annex D.2); codes 0x80–0xA0 are undefined. +#[rustfmt::skip] +pub(super) static STANDARD_NAMES: [Option<&str>; 256] = [ + /* 0x00 */ None, None, None, None, None, None, None, None, + /* 0x08 */ None, None, None, None, None, None, None, None, + /* 0x10 */ None, None, None, None, None, None, None, None, + /* 0x18 */ None, None, None, None, None, None, None, None, + /* 0x20 */ Some("space"), Some("exclam"), Some("quotedbl"), Some("numbersign"), Some("dollar"), Some("percent"), Some("ampersand"), Some("quoteright"), + /* 0x28 */ Some("parenleft"), Some("parenright"), Some("asterisk"), Some("plus"), Some("comma"), Some("hyphen"), Some("period"), Some("slash"), + /* 0x30 */ Some("zero"), Some("one"), Some("two"), Some("three"), Some("four"), Some("five"), Some("six"), Some("seven"), + /* 0x38 */ Some("eight"), Some("nine"), Some("colon"), Some("semicolon"), Some("less"), Some("equal"), Some("greater"), Some("question"), + /* 0x40 */ Some("at"), Some("A"), Some("B"), Some("C"), Some("D"), Some("E"), Some("F"), Some("G"), + /* 0x48 */ Some("H"), Some("I"), Some("J"), Some("K"), Some("L"), Some("M"), Some("N"), Some("O"), + /* 0x50 */ Some("P"), Some("Q"), Some("R"), Some("S"), Some("T"), Some("U"), Some("V"), Some("W"), + /* 0x58 */ Some("X"), Some("Y"), Some("Z"), Some("bracketleft"), Some("backslash"), Some("bracketright"), Some("asciicircum"), Some("underscore"), + /* 0x60 */ Some("quoteleft"), Some("a"), Some("b"), Some("c"), Some("d"), Some("e"), Some("f"), Some("g"), + /* 0x68 */ Some("h"), Some("i"), Some("j"), Some("k"), Some("l"), Some("m"), Some("n"), Some("o"), + /* 0x70 */ Some("p"), Some("q"), Some("r"), Some("s"), Some("t"), Some("u"), Some("v"), Some("w"), + /* 0x78 */ Some("x"), Some("y"), Some("z"), Some("braceleft"), Some("bar"), Some("braceright"), Some("asciitilde"), None, + /* 0x80 */ None, None, None, None, None, None, None, None, + /* 0x88 */ None, None, None, None, None, None, None, None, + /* 0x90 */ None, None, None, None, None, None, None, None, + /* 0x98 */ None, None, None, None, None, None, None, None, + /* 0xA0 */ None, Some("exclamdown"), Some("cent"), Some("sterling"), Some("fraction"), Some("yen"), Some("florin"), Some("section"), + /* 0xA8 */ Some("currency"), Some("quotesingle"), Some("quotedblleft"), Some("guillemotleft"), Some("guilsinglleft"), Some("guilsinglright"), Some("fi"), Some("fl"), + /* 0xB0 */ None, Some("endash"), Some("dagger"), Some("daggerdbl"), Some("periodcentered"), None, Some("paragraph"), Some("bullet"), + /* 0xB8 */ Some("quotesinglbase"), Some("quotedblbase"), Some("quotedblright"), Some("guillemotright"), Some("ellipsis"), Some("perthousand"), None, Some("questiondown"), + /* 0xC0 */ None, Some("grave"), Some("acute"), Some("circumflex"), Some("tilde"), Some("macron"), Some("breve"), Some("dotaccent"), + /* 0xC8 */ Some("dieresis"), None, Some("ring"), Some("cedilla"), None, Some("hungarumlaut"), Some("ogonek"), Some("caron"), + /* 0xD0 */ Some("emdash"), None, None, None, None, None, None, None, + /* 0xD8 */ None, None, None, None, None, None, None, None, + /* 0xE0 */ None, Some("AE"), None, Some("ordfeminine"), None, None, None, None, + /* 0xE8 */ Some("Lslash"), Some("Oslash"), Some("OE"), Some("ordmasculine"), None, None, None, None, + /* 0xF0 */ None, Some("ae"), None, None, None, Some("dotlessi"), None, None, + /* 0xF8 */ Some("lslash"), Some("oslash"), Some("oe"), Some("germandbls"), None, None, None, None, +]; diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/encodings.rs b/src-tauri/src/pdf_engine/text_edit/fonts/encodings.rs new file mode 100644 index 0000000..683b30e --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/encodings.rs @@ -0,0 +1,260 @@ +//! Base encodings, `/Differences`, glyph-name resolution and the character rules of SPEC §A.3: +//! display normalisation (ligatures), the reading allow-list (§A.3.3) and the typeable-character +//! exclusions (§A.3.2). + +use super::agl; +use super::encoding_tables::{MAC_ROMAN_NAMES, STANDARD_NAMES, WIN_ANSI_NAMES}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use lopdf::Object; + +/// A named base encoding a simple font may use (ISO 32000-1 Annex D.2). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum BaseEncoding { + WinAnsi, + MacRoman, + Standard, +} + +impl BaseEncoding { + /// `/WinAnsiEncoding`, `/MacRomanEncoding`, `/StandardEncoding`; anything else (incl. + /// `/MacExpertEncoding`) is `UNSUPPORTED_ENCODING`. + pub fn from_name(name: &[u8]) -> Result { + match name { + b"WinAnsiEncoding" => Ok(BaseEncoding::WinAnsi), + b"MacRomanEncoding" => Ok(BaseEncoding::MacRoman), + b"StandardEncoding" => Ok(BaseEncoding::Standard), + _ => Err(TextReason::UnsupportedEncoding), + } + } + + pub fn table(self) -> &'static [Option<&'static str>; 256] { + match self { + BaseEncoding::WinAnsi => &WIN_ANSI_NAMES, + BaseEncoding::MacRoman => &MAC_ROMAN_NAMES, + BaseEncoding::Standard => &STANDARD_NAMES, + } + } + + /// The glyph name at `code`, if the encoding defines one. + pub fn name(self, code: u8) -> Option<&'static str> { + self.table().get(usize::from(code)).copied().flatten() + } + + /// The 256 names as owned values (the start of a font's name table). + pub fn names(self) -> Vec> { + self.table().iter().map(|n| n.map(str::to_string)).collect() + } +} + +/// StandardEncoding's name at `code` (Type1 `seac` components, CFF built-in Standard). +pub fn standard_name(code: u8) -> Option<&'static str> { + BaseEncoding::Standard.name(code) +} + +/// The strict (Annex D) MacRoman code of a glyph name, for the TrueType `(1,0)` strategy. +pub fn mac_roman_code(name: &str) -> Option { + static BY_NAME: std::sync::OnceLock> = std::sync::OnceLock::new(); + let by_name = BY_NAME.get_or_init(|| { + let mut pairs: Vec<(&'static str, u8)> = MAC_ROMAN_NAMES + .iter() + .zip(0u8..=255) + .filter_map(|(n, code)| n.map(|n| (n, code))) + .collect(); + // Lowest code first for a name listed twice (0x20 and 0xCA are both `space`). + pairs.sort_by(|a, b| a.0.cmp(b.0).then(a.1.cmp(&b.1))); + pairs.dedup_by(|later, earlier| later.0 == earlier.0); + pairs + }); + by_name + .binary_search_by(|(n, _)| (*n).cmp(name)) + .ok() + .and_then(|i| by_name.get(i)) + .map(|(_, code)| *code) +} + +/// Applies a resolved `/Differences` array to a 256-entry name table. Only integers (0–255, +/// setting the next code) and names are allowed; anything else, a name before the first integer +/// or a code past 255 is `FONT_UNSUPPORTED`. A name that is not UTF-8 resolves to no name. +pub fn apply_differences(names: &mut [Option], items: &[Object]) -> Result<(), TextReason> { + let mut next: Option = None; + for item in items { + match item { + Object::Integer(code) => { + let code = usize::try_from(*code).map_err(|_| TextReason::FontUnsupported)?; + if code > 255 { + return Err(TextReason::FontUnsupported); + } + next = Some(code); + } + Object::Name(bytes) => { + let code = next.ok_or(TextReason::FontUnsupported)?; + let slot = names.get_mut(code).ok_or(TextReason::FontUnsupported)?; + *slot = std::str::from_utf8(bytes).ok().map(str::to_string); + next = code.checked_add(1); + } + _ => return Err(TextReason::FontUnsupported), + } + } + Ok(()) +} + +/// The Unicode scalar of a glyph name (AGL, `uniXXXX`, `uXXXX[XX]`); `.notdef` has none. +pub fn glyph_name_char(name: &str) -> Option { + if name == ".notdef" { + return None; + } + agl::glyph_name_char(name) +} + +/// ISO 32000-1 Annex D notes 5 and 6: the `space` glyph at code 0xA0 (WinAnsi) or 0xCA +/// (MacRoman) stands for U+00A0, and `hyphen` at 0xAD (WinAnsi) for U+00AD, so a ToUnicode that +/// says so there agrees with the glyph name (Word writes exactly this for no-break spaces and +/// soft hyphens). At any other code the names keep their own characters. +pub fn duplicate_agrees(code: u8, name: &str, text: &str) -> bool { + matches!( + (code, name, text), + (0xA0 | 0xCA, "space", "\u{a0}") | (0xAD, "hyphen", "\u{ad}") + ) +} + +/// Display normalisation of glyph text (§A.3.1 item 5): U+FB00–U+FB06 expand to +/// `ff fi fl ffi ffl st st`; every other character is kept. +pub fn display_text(raw: &str) -> String { + let mut out = String::with_capacity(raw.len()); + for ch in raw.chars() { + match ch { + '\u{FB00}' => out.push_str("ff"), + '\u{FB01}' => out.push_str("fi"), + '\u{FB02}' => out.push_str("fl"), + '\u{FB03}' => out.push_str("ffi"), + '\u{FB04}' => out.push_str("ffl"), + '\u{FB05}' | '\u{FB06}' => out.push_str("st"), + other => out.push(other), + } + } + out +} + +/// Whether decoded text is usable at all (§A.3.1 item 5): not empty, no U+FFFD, no Private Use. +pub fn usable_text(text: &str) -> bool { + !text.is_empty() && !text.chars().any(|c| c == '\u{FFFD}' || is_private_use(c)) +} + +pub fn is_private_use(ch: char) -> bool { + matches!(u32::from(ch), 0xE000..=0xF8FF | 0xF0000..=0xFFFFD | 0x100000..=0x10FFFD) +} + +/// Blocks of the reading allow-list (§A.3.3), inclusive code point ranges. +const READING_ALLOWED: &[(u32, u32)] = &[ + (0x0000, 0x024F), // Basic Latin … Latin Extended-B + (0x0250, 0x02AF), // IPA Extensions + (0x02B0, 0x02FF), // Spacing Modifier Letters + (0x0370, 0x03FF), // Greek and Coptic + (0x0400, 0x04FF), // Cyrillic + (0x0500, 0x052F), // Cyrillic Supplement + (0x1E00, 0x1EFF), // Latin Extended Additional + (0x1F00, 0x1FFF), // Greek Extended + (0x2000, 0x206F), // General Punctuation (controls removed below) + (0x2070, 0x209F), // Superscripts and Subscripts + (0x20A0, 0x20CF), // Currency Symbols + (0x2100, 0x214F), // Letterlike Symbols + (0x2150, 0x218F), // Number Forms + (0x2190, 0x21FF), // Arrows + (0x2200, 0x22FF), // Mathematical Operators + (0x2500, 0x257F), // Box Drawing + (0x25A0, 0x25FF), // Geometric Shapes + (0x3000, 0x303F), // CJK Symbols and Punctuation + (0x3040, 0x309F), // Hiragana + (0x30A0, 0x30FF), // Katakana + (0x4E00, 0x9FFF), // CJK Unified Ideographs (BMP) + (0xAC00, 0xD7AF), // Hangul Syllables + (0xFF00, 0xFFEF), // Halfwidth and Fullwidth Forms + (0xFB00, 0xFB06), // Alphabetic Presentation Forms (Latin ligatures) +]; + +/// Right-to-left scripts and their presentation forms (§A.3.3): Hebrew, Arabic, Syriac, Thaana, +/// NKo (and the other RTL blocks between them), Hebrew/Arabic presentation forms. +const RIGHT_TO_LEFT: &[(u32, u32)] = &[ + (0x0590, 0x08FF), + (0xFB1D, 0xFDFF), + (0xFE70, 0xFEFE), + (0x10800, 0x10FFF), + (0x1E800, 0x1EFFF), +]; + +/// General Punctuation controls that are not part of the reading allow-list: zero-width and +/// bidi marks, separators, embeddings/overrides, and U+2060–U+206F (word joiner, invisible +/// operators, the bidi isolates U+2066–U+2069 and the deprecated format controls U+206A–U+206F). +const PUNCTUATION_CONTROLS: &[(u32, u32)] = &[ + (0x200B, 0x200F), + (0x2028, 0x2029), + (0x202A, 0x202E), + (0x2060, 0x206F), +]; + +fn in_ranges(cp: u32, ranges: &[(u32, u32)]) -> bool { + ranges.iter().any(|&(lo, hi)| lo <= cp && cp <= hi) +} + +/// `None` when `ch` may be read (and so edited around); `RIGHT_TO_LEFT` or `COMPLEX_SCRIPT` +/// otherwise (§A.3.3). +pub fn reading_reason(ch: char) -> Option { + let cp = u32::from(ch); + if in_ranges(cp, RIGHT_TO_LEFT) { + return Some(TextReason::RightToLeft); + } + if in_ranges(cp, READING_ALLOWED) && !in_ranges(cp, PUNCTUATION_CONTROLS) { + return None; + } + Some(TextReason::ComplexScript) +} + +/// Characters a code may never be typeable as (§A.3.2), whatever the font. +const TYPING_EXCLUDED: &[(u32, u32)] = &[ + (0x0000, 0x001F), // C0 controls + (0x007F, 0x009F), // DEL + C1 controls + (0x0300, 0x036F), // combining diacritical marks + (0x1AB0, 0x1AFF), + (0x1DC0, 0x1DFF), + (0x20D0, 0x20FF), + (0xFE20, 0xFE2F), + (0x200B, 0x200F), // zero-width and bidi controls + (0x2028, 0x2029), // line / paragraph separators + (0x202A, 0x202E), + (0x2060, 0x206F), // word joiner … invisible operators, bidi isolates, deprecated controls + (0xFEFF, 0xFEFF), + (0x180B, 0x180F), // variation selectors + (0xFE00, 0xFE0F), + (0xE000, 0xF8FF), // Private Use + (0xFFFD, 0xFFFD), + (0xFB00, 0xFB06), // ligatures are read as their letters, never typed as one character +]; + +/// Whether `ch` may be a typeable character (§A.3.2): BMP, not excluded, and readable (§A.3.3). +pub fn writable_char(ch: char) -> bool { + let cp = u32::from(ch); + if (0x20..0x7F).contains(&cp) { + return true; // printable ASCII: in Basic Latin, in no excluded range + } + cp <= 0xFFFF && !in_ranges(cp, TYPING_EXCLUDED) && reading_reason(ch).is_none() +} + +/// The Latin/Greek/Cyrillic allow-list for fonts the PDF reader substitutes (§A.2 row +/// "not embedded"): Latin through Extended-B, spacing modifiers, Greek, Cyrillic (+Supplement), +/// Latin Extended Additional, Greek Extended, General Punctuation, Currency and Letterlike. +const SUBSTITUTION_SAFE: &[(u32, u32)] = &[ + (0x0020, 0x024F), + (0x02B0, 0x02FF), + (0x0370, 0x03FF), + (0x0400, 0x052F), + (0x1E00, 0x1EFF), + (0x1F00, 0x1FFF), + (0x2000, 0x206F), + (0x20A0, 0x20CF), + (0x2100, 0x214F), +]; + +/// Whether a substituted (non-embedded, non-Standard-14) font may write `ch`. +pub fn substitution_safe(ch: char) -> bool { + writable_char(ch) && in_ranges(u32::from(ch), SUBSTITUTION_SAFE) +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/faces.rs b/src-tauri/src/pdf_engine/text_edit/fonts/faces.rs new file mode 100644 index 0000000..9cb0167 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/faces.rs @@ -0,0 +1,324 @@ +//! Faces (SPEC §B.9.6): an exact port of `mobile/src/lib/pdf/edit/faces.ts` (family key, bold, +//! italic from the BaseFont words, then the descriptor), plus the typing surfaces built on it +//! (§B.9.1 `typing_surface` / `face_surface`: sibling fonts of one face group). + +use super::{Code, FamilyHint, FontClass, FontModel, TypingSurface}; +use crate::pdf_engine::text_edit::reasons::Face; +use std::sync::Arc; + +const BOLD_WORDS: &[&str] = &[ + "bold", + "semibold", + "demibold", + "extrabold", + "ultrabold", + "black", + "heavy", +]; +const ITALIC_WORDS: &[&str] = &["italic", "oblique"]; +/// Not a style: dropped from both sides so `X-Regular` and `X` are one family. +const NEUTRAL_WORDS: &[&str] = &["regular", "roman", "book", "normal"]; + +/// Name parts of monospaced families (checked first: `LiberationMono`, `DejaVuSansMono`). +const MONO_NAMES: &[&str] = &["mono", "courier", "consol", "menlo", "monaco"]; +/// Name parts of sans families (checked before the serif ones: `MicrosoftSansSerif`, +/// `CenturyGothic`). +const SANS_NAMES: &[&str] = &["sans", "gothic", "grotesk", "grotesque"]; +/// Name parts of serif families. +const SERIF_NAMES: &[&str] = &[ + "serif", + "times", + "georgia", + "cambria", + "garamond", + "bookantiqua", + "palatino", + "century", + "baskerville", + "minion", + "caslon", + "charter", +]; + +/// FontDescriptor values that speak about the face when its name says nothing. +#[derive(Debug, Clone, Copy, Default, PartialEq)] +pub struct FaceHints { + pub flags: Option, + pub stem_v: Option, + pub italic_angle: Option, +} + +/// `(name without the ABCDEF+ subset tag, whether the tag was there)`. +pub fn strip_subset_tag(base: &str) -> (&str, bool) { + let bytes = base.as_bytes(); + let tagged = bytes.len() >= 7 + && bytes + .get(..6) + .is_some_and(|t| t.iter().all(u8::is_ascii_uppercase)) + && bytes.get(6) == Some(&b'+'); + match (tagged, base.get(7..)) { + (true, Some(rest)) => (rest, true), + _ => (base, false), + } +} + +/// Splits at punctuation and where a lowercase letter or digit is followed by an uppercase one; +/// keeps the original case. `TimesNewRomanPS-BoldMT` → `Times New Roman PS Bold MT`. +fn raw_words(base: &str) -> Vec { + let (name, _) = strip_subset_tag(base); + let mut words = Vec::new(); + let mut current = String::new(); + let mut prev: Option = None; + for ch in name.chars() { + if !ch.is_ascii_alphanumeric() { + if !current.is_empty() { + words.push(std::mem::take(&mut current)); + } + prev = None; + continue; + } + let split = ch.is_ascii_uppercase() + && prev.is_some_and(|p| p.is_ascii_lowercase() || p.is_ascii_digit()); + if split && !current.is_empty() { + words.push(std::mem::take(&mut current)); + } + current.push(ch); + prev = Some(ch); + } + if !current.is_empty() { + words.push(current); + } + words +} + +/// `wordsOf`: the BaseFont's words, lowercased. +pub fn words_of(base: &str) -> Vec { + raw_words(base) + .into_iter() + .map(|w| w.to_ascii_lowercase()) + .collect() +} + +/// `familyOf`: the words minus style and neutral words, joined with no separator. +pub fn family_key(base: &str) -> String { + words_of(base) + .into_iter() + .filter(|w| { + let w = w.as_str(); + !BOLD_WORDS.contains(&w) && !ITALIC_WORDS.contains(&w) && !NEUTRAL_WORDS.contains(&w) + }) + .collect() +} + +/// `familyHint` (§D.2): the descriptor's FixedPitch (bit 1) or Serif (bit 2) flag, else the +/// family name. LibreOffice writes `/Flags 4` (Symbolic only) for Liberation Serif, so the flag +/// alone set every serif line in a sans face (live check B3). +pub fn family_hint(base: &str, flags: Option) -> FamilyHint { + match flags { + Some(f) if f & 1 != 0 => return FamilyHint::Mono, + Some(f) if f & 2 != 0 => return FamilyHint::Serif, + _ => {} + } + let name: String = strip_subset_tag(base) + .0 + .chars() + .filter(char::is_ascii_alphanumeric) + .map(|c| c.to_ascii_lowercase()) + .collect(); + let has = |parts: &[&str]| parts.iter().any(|p| name.contains(p)); + if has(MONO_NAMES) { + FamilyHint::Mono + } else if has(SANS_NAMES) { + FamilyHint::Sans + } else if has(SERIF_NAMES) { + FamilyHint::Serif + } else { + FamilyHint::Sans + } +} + +/// A readable name for the UI: the subset tag removed, words separated by spaces. +pub fn display_name(base: &str) -> String { + raw_words(base).join(" ") +} + +fn has_word(words: &[String], set: &[&str]) -> bool { + words.iter().any(|w| set.contains(&w.as_str())) +} + +/// `isBold`: a bold word → true; a neutral word → false; ForceBold (Flags bit 19) → true; never +/// for an italic face; else `StemV ≥ 120`. +pub fn is_bold(base: &str, hints: &FaceHints) -> bool { + let words = words_of(base); + if has_word(&words, BOLD_WORDS) { + return true; + } + if has_word(&words, NEUTRAL_WORDS) { + return false; + } + if hints.flags.is_some_and(|f| f & (1 << 18) != 0) { + return true; + } + if is_italic(base, hints) { + return false; + } + hints.stem_v.is_some_and(|s| s >= 120.0) +} + +/// `isItalic`: an italic word → true; a neutral word → false; Italic flag (bit 7) → true; else +/// `|ItalicAngle| ≥ 4`. +pub fn is_italic(base: &str, hints: &FaceHints) -> bool { + let words = words_of(base); + if has_word(&words, ITALIC_WORDS) { + return true; + } + if has_word(&words, NEUTRAL_WORDS) { + return false; + } + if hints.flags.is_some_and(|f| f & (1 << 6) != 0) { + return true; + } + hints.italic_angle.is_some_and(|a| a.abs() >= 4.0) +} + +fn is_type0(class: FontClass) -> bool { + matches!(class, FontClass::Type0Cid2 | FontClass::Type0Cid0) +} + +/// Word's pattern: one simple TrueType and one Type0 CIDFontType2 with the same base name. +fn word_pair(a: &FontModel, b: &FontModel) -> bool { + match (a.class, b.class) { + (Some(FontClass::SimpleTrueType), Some(FontClass::Type0Cid2)) + | (Some(FontClass::Type0Cid2), Some(FontClass::SimpleTrueType)) => { + a.base_name == b.base_name + } + _ => false, + } +} + +/// Same class family (simple↔simple, Type0↔Type0) or the Word pair. +fn compatible(a: &FontModel, b: &FontModel) -> bool { + match (a.class, b.class) { + (Some(x), Some(y)) => is_type0(x) == is_type0(y) || word_pair(a, b), + _ => false, + } +} + +fn usable(m: &FontModel) -> bool { + m.class.is_some() && m.refusal.is_none() && !m.family_key.is_empty() +} + +pub(super) fn typing_surface( + page_fonts: &[(Vec, Arc)], + primary: &[u8], +) -> TypingSurface { + let Some((name, model)) = page_fonts.iter().find(|(n, _)| n.as_slice() == primary) else { + return TypingSurface { fonts: Vec::new() }; + }; + let mut fonts = vec![(name.clone(), Arc::clone(model))]; + if !usable(model) { + return TypingSurface { fonts }; + } + for (other_name, other) in page_fonts { + let seen = fonts.iter().any(|(_, m)| Arc::ptr_eq(m, other)); + if !seen + && usable(other) + && other.family_key == model.family_key + && other.bold == model.bold + && other.italic == model.italic + && compatible(model, other) + { + fonts.push((other_name.clone(), Arc::clone(other))); + } + } + TypingSurface { fonts } +} + +pub(super) fn face_surface( + page_fonts: &[(Vec, Arc)], + primary: &FontModel, + face: Face, +) -> Option { + if !usable(primary) { + return None; + } + let (bold, italic) = match face { + Face::Regular => (false, false), + Face::Bold => (true, false), + Face::Italic => (false, true), + Face::BoldItalic => (true, true), + }; + let group: Vec<&(Vec, Arc)> = page_fonts + .iter() + .filter(|(_, m)| { + usable(m) && m.family_key == primary.family_key && m.bold == bold && m.italic == italic + }) + .collect(); + // The group's lead has the primary's own class family; mobile never swaps simple for Type0 + // except inside Word's simple + Type0 pair. + let primary_type0 = primary.class.is_some_and(is_type0); + let (lead_name, lead) = group + .iter() + .find(|(_, m)| m.class.is_some_and(is_type0) == primary_type0)?; + let mut fonts = vec![(lead_name.clone(), Arc::clone(lead))]; + for (name, model) in &group { + let seen = fonts.iter().any(|(_, m)| Arc::ptr_eq(m, model)); + if !seen && compatible(lead, model) { + fonts.push((name.clone(), Arc::clone(model))); + } + } + Some(TypingSurface { fonts }) +} + +impl TypingSurface { + /// The first font (primary first) that can type `ch`, with its lowest code for it. + pub fn writer_for(&self, ch: char) -> Option<(usize, Code)> { + self.writer_for_with(ch, &[], &[]) + } + + /// As `writer_for`, honouring §A.3.2's preference order: a `(font index, char, code)` the run + /// already uses for `ch`, then one the page uses, then the first font's lowest code. + pub fn writer_for_with( + &self, + ch: char, + prefer_run: &[(usize, char, Code)], + prefer_page: &[(usize, char, Code)], + ) -> Option<(usize, Code)> { + for (idx, c, code) in prefer_run.iter().chain(prefer_page) { + let usable = *c == ch + && self + .fonts + .get(*idx) + .is_some_and(|(_, m)| m.can_write(ch, *code)); + if usable { + return Some((*idx, *code)); + } + } + self.fonts + .iter() + .enumerate() + .find_map(|(idx, (_, m))| m.code_for(ch, &[], &[]).map(|code| (idx, code))) + } + + /// Whether some font of the surface can type U+0020 (§A.3.4 `space_mode`). + pub fn has_space(&self) -> bool { + self.fonts + .iter() + .any(|(_, m)| m.code_for(' ', &[], &[]).is_some()) + } +} + +#[cfg(test)] +impl TypingSurface { + /// Every character some font of the surface can type, sorted, without duplicates. + pub fn alphabet(&self) -> Vec { + let mut chars: Vec = self + .fonts + .iter() + .flat_map(|(_, m)| m.alphabet().into_iter().map(|(c, _)| c)) + .collect(); + chars.sort_unstable(); + chars.dedup(); + chars + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/glyph_budget.rs b/src-tauri/src/pdf_engine/text_edit/fonts/glyph_budget.rs new file mode 100644 index 0000000..9ee7f95 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/glyph_budget.rs @@ -0,0 +1,540 @@ +//! Bounded work for glyph-presence proofs (SPEC §A.2, §B.9.5, §B.2). +//! +//! ttf-parser's outlining only limits nesting depth (glyf `MAX_COMPONENTS` 32, CFF `STACK_LIMIT` +//! 10), so a tiny glyph whose composites or subroutines fan out runs for minutes. Before +//! ttf-parser outlines a glyph, a pre-check walks the same structure with a work budget: +//! - `GlyfGuard`: the composite tree exactly as `glyf::outline_impl` visits it (component records +//! read as `CompositeGlyphIter` reads them); one unit per glyph visit and component record, one +//! per contour and per point of every simple glyph visited. +//! - `CffGuard`: the Type2 charstring exactly as `cff1::_parse_char_string` runs it (operand stack +//! values, `callsubr`/`callgsubr` with their bias, `return`, `endchar` and its `seac` form, hint +//! masks sized by the stems seen, the width rules that decide the stem count); one unit per byte +//! read and per FDSelect range or charset entry walked, at most `CFF_OPS_PER_GLYPH_MAX` +//! operators. A CID-keyed glyph that calls a local subroutine is also charged its Font DICT and +//! Private DICT bytes, which ttf-parser walks again for every such glyph; the guard itself +//! resolves each Font DICT once (and charges that walk once). Wherever ttf-parser stops with an +//! error the walk stops too (the glyph does not draw); argument-count errors are not checked, +//! so the walk never does less than ttf-parser. +//! +//! A glyph may spend at most `GLYPH_WORK_MAX` units, and never more than the page budget has left; +//! over that it does not draw. Every unit is charged to a `WorkMeter` holding what is left of the +//! page decode budget, so one page's font work is bounded however many fonts and glyphs it has: +//! a walk stops as soon as it passes what the meter can still pay, a meter that runs out marks the +//! font load `PAGE_TOO_COMPLEX`, and after that no glyph is walked at all. The Type1 interpreter +//! (`type1.rs`) charges the same meter under the same rules. + +use super::cff_layout::{seac_glyph, CffLayout, FdSubrs, Index, LocalSubrs}; +use std::cell::OnceCell; +use std::collections::HashMap; +use ttf_parser::{loca, Face, Tag}; + +/// Work units one glyph proof may spend. +pub(crate) const GLYPH_WORK_MAX: u64 = 1 << 17; +/// Type2 operators one glyph may run, subroutines and `seac` components included. +const CFF_OPS_PER_GLYPH_MAX: u32 = 4_096; +/// ttf-parser's limits (`glyf::MAX_COMPONENTS`, `cff1::STACK_LIMIT`, `MAX_ARGUMENTS_STACK_LEN`). +const GLYF_DEPTH_LIMIT: u8 = 32; +const CFF_DEPTH_LIMIT: u8 = 10; +const CFF_STACK_MAX: usize = 48; + +/// What is left of the page decode budget for one font load's glyph work. +#[derive(Debug)] +pub(crate) struct WorkMeter { + remaining: u64, + used: u64, + exhausted: bool, +} + +impl WorkMeter { + pub fn new(allowance: usize) -> WorkMeter { + WorkMeter { + remaining: u64::try_from(allowance).unwrap_or(u64::MAX), + used: 0, + exhausted: false, + } + } + + /// Debits `units`; `false` (and the meter is exhausted) when they do not fit. + pub fn charge(&mut self, units: u64) -> bool { + if self.exhausted || units > self.remaining { + self.exhausted = true; + self.remaining = 0; + return false; + } + self.remaining -= units; + self.used = self.used.saturating_add(units); + true + } + + /// Units the next glyph may spend: `GLYPH_WORK_MAX`, or less when the page budget has less + /// left (0 once exhausted). + pub fn glyph_allowance(&self) -> u64 { + if self.exhausted { + 0 + } else { + GLYPH_WORK_MAX.min(self.remaining) + } + } + + pub fn used(&self) -> usize { + usize::try_from(self.used).unwrap_or(usize::MAX) + } + + pub fn exhausted(&self) -> bool { + self.exhausted + } +} + +/// Units spent by one glyph against its allowance (`WorkMeter::glyph_allowance`). +struct Work { + used: u64, + cap: u64, +} + +impl Work { + fn new(meter: &WorkMeter) -> Work { + Work { + used: 0, + cap: meter.glyph_allowance(), + } + } + + fn tick(&mut self, units: u64) -> Result<(), ()> { + self.used = self.used.saturating_add(units); + if self.used > self.cap { + return Err(()); + } + Ok(()) + } + + /// Debits the units spent (at least one) from `meter`; `false` when it cannot pay them. + fn settle(&self, meter: &mut WorkMeter) -> bool { + meter.charge(self.used.max(1)) + } +} + +/// The glyf/loca tables of a face, read through ttf-parser's own `loca` parser. +pub(crate) struct GlyfGuard<'a> { + loca: loca::Table<'a>, + glyf: &'a [u8], +} + +impl<'a> GlyfGuard<'a> { + pub fn new(face: &Face<'a>) -> Option> { + let tables = face.tables(); + let raw = face.raw_face(); + let loca = loca::Table::parse( + tables.maxp.number_of_glyphs, + tables.head.index_to_location_format, + raw.table(Tag::from_bytes(b"loca"))?, + )?; + Some(GlyfGuard { + loca, + glyf: raw.table(Tag::from_bytes(b"glyf"))?, + }) + } + + /// Whether ttf-parser may outline `gid` (within budget, and not a glyph it fails on early); + /// the units are charged to `meter` either way. An exhausted meter walks nothing. + pub fn check(&self, gid: u16, meter: &mut WorkMeter) -> bool { + if meter.exhausted() { + return false; + } + let mut work = Work::new(meter); + let ok = match self.glyph_data(gid) { + Some(data) => self.visit(data, 0, &mut work).is_ok(), + None => false, + }; + work.settle(meter) && ok + } + + fn glyph_data(&self, gid: u16) -> Option<&'a [u8]> { + self.glyf + .get(self.loca.glyph_range(ttf_parser::GlyphId(gid))?) + } + + fn visit(&self, data: &[u8], depth: u8, work: &mut Work) -> Result<(), ()> { + if depth >= GLYF_DEPTH_LIMIT { + return Err(()); + } + work.tick(1)?; + let contours = data + .get(..2) + .and_then(|b| Some(i16::from_be_bytes([*b.first()?, *b.get(1)?]))) + .ok_or(())?; + if contours == 0 { + return Ok(()); + } + let tail = data.get(10..).ok_or(())?; + if contours > 0 { + let n = usize::from(contours.unsigned_abs()); + let last_at = (n - 1) * 2; + let last = tail + .get(last_at..last_at + 2) + .and_then(|b| Some(u16::from_be_bytes([*b.first()?, *b.get(1)?]))) + .ok_or(())?; + let points = last.checked_add(1).ok_or(())?; + return work.tick(n as u64 + u64::from(points)); + } + // Composite records, read as `CompositeGlyphIter` reads them; a record cut short ends + // the list (it is not an error there). + let mut pos = 0usize; + loop { + let Some((flags, component, next)) = component_record(tail, pos) else { + return Ok(()); + }; + work.tick(1)?; + if let Some(data) = self.glyph_data(component) { + self.visit(data, depth + 1, work)?; + } + if flags & 0x0020 == 0 { + return Ok(()); + } + pos = next; + } + } +} + +/// One composite record at `pos`: flags, component GID and the offset after it. Arguments are +/// read only with ARGS_ARE_XY_VALUES (ttf-parser's reading), then the scale fields. +fn component_record(tail: &[u8], pos: usize) -> Option<(u16, u16, usize)> { + let word = |at: usize| -> Option { + let b = tail.get(at..at.checked_add(2)?)?; + Some(u16::from_be_bytes([*b.first()?, *b.get(1)?])) + }; + let flags = word(pos)?; + let gid = word(pos.checked_add(2)?)?; + let mut at = pos.checked_add(4)?; + if flags & 0x0002 != 0 { + at = at.checked_add(if flags & 0x0001 != 0 { 4 } else { 2 })?; + } + if flags & 0x0080 != 0 { + at = at.checked_add(8)?; + } else if flags & 0x0040 != 0 { + at = at.checked_add(4)?; + } else if flags & 0x0008 != 0 { + at = at.checked_add(2)?; + } + if at > tail.len() { + return None; + } + Some((flags, gid, at)) +} + +/// FDSelect selects Font DICTs by a `u8` index. +const FONT_DICTS_MAX: usize = 256; + +/// A CFF program's layout, its SID → first GID map (for `seac`) and, for a CID-keyed font, each +/// Font DICT's local subroutines once resolved. +pub(crate) struct CffGuard<'a> { + layout: CffLayout<'a>, + seac: HashMap, + fd_subrs: Vec>>, +} + +impl<'a> CffGuard<'a> { + /// `None` when the program cannot be laid out, or disagrees with ttf-parser's glyph count. + pub fn new(data: &'a [u8], glyphs: u16) -> Option> { + let layout = CffLayout::parse(data)?; + if layout.number_of_glyphs != glyphs { + return None; + } + let mut seac = HashMap::new(); + if let Some(sids) = layout.gid_to_sid() { + for (gid, sid) in sids.iter().enumerate().skip(1) { + if let (Some(sid), Ok(gid)) = (sid, u16::try_from(gid)) { + seac.entry(*sid).or_insert(gid); + } + } + } + let fds = if layout.is_cid() { FONT_DICTS_MAX } else { 0 }; + Some(CffGuard { + layout, + seac, + fd_subrs: (0..fds).map(|_| OnceCell::new()).collect(), + }) + } + + /// The local subroutines `glyph` runs and the units finding them costs: the FDSelect ranges + /// walked; for a CID-keyed font also the Font and Private DICT bytes ttf-parser walks for + /// this glyph, plus this guard's own walk of them the first time that Font DICT is used. + fn local_subrs(&self, glyph: u16) -> (Option>, u64) { + let (fd, walked) = match self.layout.local_subrs(glyph) { + LocalSubrs::Font(index) => return (Some(index), 0), + LocalSubrs::FontDict(fd, walked) => (fd, walked), + }; + let Some((fd, cell)) = fd.and_then(|fd| Some((fd, self.fd_subrs.get(usize::from(fd))?))) + else { + return (None, walked); + }; + let (found, first_walk) = match cell.get() { + Some(found) => (*found, 0), + None => { + let found = self.layout.fd_local_subrs(fd); + let _ = cell.set(found); // empty until now: never refused + (found, found.dict_bytes) + } + }; + let units = walked + .saturating_add(found.dict_bytes) + .saturating_add(first_walk); + (found.subrs, units) + } + + pub fn layout(&self) -> &CffLayout<'a> { + &self.layout + } + + /// Whether ttf-parser may outline `gid` (within budget, ends in `endchar`, no error on the + /// way); the units are charged to `meter` either way. An exhausted meter walks nothing. + pub fn check(&self, gid: u16, meter: &mut WorkMeter) -> bool { + if meter.exhausted() { + return false; + } + let mut scan = Type2 { + guard: self, + glyph: gid, + local: None, + local_resolved: false, + stack: Vec::with_capacity(CFF_STACK_MAX), + width: false, + stems: 0, + has_endchar: false, + has_seac: false, + ops: 0, + work: Work::new(meter), + }; + let ok = match self.layout.char_strings.get(u32::from(gid)) { + Some(code) => scan.run(code, 0).is_ok() && scan.has_endchar, + None => false, + }; + scan.work.settle(meter) && ok + } +} + +/// One Type2 charstring walk (the state `CharStringParserContext` + `CharStringParser` keep that +/// decides which bytes run next). +struct Type2<'g, 'a> { + guard: &'g CffGuard<'a>, + glyph: u16, + local: Option>, + local_resolved: bool, + stack: Vec, + width: bool, + stems: u32, + has_endchar: bool, + has_seac: bool, + ops: u32, + work: Work, +} + +impl<'a> Type2<'_, 'a> { + fn push(&mut self, v: f32) -> Result<(), ()> { + if self.stack.len() >= CFF_STACK_MAX { + return Err(()); + } + self.stack.push(v); + Ok(()) + } + + fn operator(&mut self) -> Result<(), ()> { + self.ops += 1; + if self.ops > CFF_OPS_PER_GLYPH_MAX { + return Err(()); + } + Ok(()) + } + + /// `width` rule of the moveto operators: an extra leading argument is the width. + fn moveto(&mut self, args: usize) { + if self.stack.len() == args + 1 { + self.width = true; + } + self.stack.clear(); + } + + fn run(&mut self, code: &'a [u8], depth: u8) -> Result<(), ()> { + let byte = |at: usize| code.get(at).copied().ok_or(()); + let mut pos = 0usize; + while let Some(&op) = code.get(pos) { + pos += 1; + self.work.tick(1)?; + match op { + 0 | 2 | 9 | 13 | 15 | 16 | 17 => return Err(()), + 1 | 3 | 18 | 23 => { + self.operator()?; + let mut len = self.stack.len(); + if len % 2 == 1 && !self.width { + self.width = true; + len -= 1; + } + self.stems = self.stems.saturating_add((len / 2) as u32); + self.stack.clear(); + } + 4 | 22 => { + self.operator()?; + self.moveto(1); + } + 21 => { + self.operator()?; + self.moveto(2); + } + 5..=8 | 24..=27 | 30 | 31 => { + self.operator()?; + self.stack.clear(); + } + 10 | 29 => { + self.operator()?; + self.call(op == 29, depth)?; + if self.has_endchar && !self.has_seac { + return if pos < code.len() { Err(()) } else { Ok(()) }; + } + } + 11 => { + self.operator()?; + return Ok(()); + } + 12 => { + self.operator()?; + let op2 = byte(pos)?; + pos += 1; + self.work.tick(1)?; + if !(34..=37).contains(&op2) { + return Err(()); // only the flex operators are supported there + } + self.stack.clear(); + } + 14 => { + self.operator()?; + self.endchar(depth)?; + if pos < code.len() { + return Err(()); + } + self.has_endchar = true; + return Ok(()); + } + 19 | 20 => { + self.operator()?; + let mut len = self.stack.len(); + self.stack.clear(); + if len % 2 == 1 { + len -= 1; + self.width = true; + } + self.stems = self.stems.saturating_add((len / 2) as u32); + let mask = usize::try_from((u64::from(self.stems) + 7) >> 3).map_err(|_| ())?; + pos = pos.saturating_add(mask); + } + 28 => { + let v = i16::from_be_bytes([byte(pos)?, byte(pos + 1)?]); + pos += 2; + self.work.tick(2)?; + self.push(f32::from(v))?; + } + 32..=246 => self.push(f32::from(i16::from(op) - 139))?, + 247..=250 => { + let b1 = i16::from(byte(pos)?); + pos += 1; + self.work.tick(1)?; + self.push(f32::from((i16::from(op) - 247) * 256 + b1 + 108))?; + } + 251..=254 => { + let b1 = i16::from(byte(pos)?); + pos += 1; + self.work.tick(1)?; + self.push(f32::from(-(i16::from(op) - 251) * 256 - b1 - 108))?; + } + 255 => { + let v = i32::from_be_bytes([ + byte(pos)?, + byte(pos + 1)?, + byte(pos + 2)?, + byte(pos + 3)?, + ]); + pos += 4; + self.work.tick(4)?; + self.push(v as f32 / 65536.0)?; + } + } + } + Ok(()) + } + + /// `callsubr` (local) or `callgsubr` (global): pops the biased index and runs the subroutine. + fn call(&mut self, global: bool, depth: u8) -> Result<(), ()> { + if self.stack.is_empty() || depth == CFF_DEPTH_LIMIT { + return Err(()); + } + let subrs = if global { + self.guard.layout.global_subrs + } else { + if !self.local_resolved { + let (local, units) = self.guard.local_subrs(self.glyph); + self.work.tick(units)?; + self.local = local; + self.local_resolved = true; + } + self.local.ok_or(())? + }; + let index = self.stack.pop().ok_or(())?; + let sub = subrs.get(subr_index(index, subrs.len())?).ok_or(())?; + self.run(sub, depth + 1) + } + + /// `endchar`: with 4 arguments (5 before the width) it is `seac`, which runs the base and + /// accent glyphs' charstrings (`seac_code_to_glyph_id` through the charset). + fn endchar(&mut self, depth: u8) -> Result<(), ()> { + let len = self.stack.len(); + if len == 4 || (!self.width && len == 5) { + let code_of = |v: Option| -> Option { + let v = v?; + (v >= i32::MIN as f32 && v < i32::MAX as f32) + .then(|| u8::try_from(v as i32).ok()) + .flatten() + }; + let layout = &self.guard.layout; + let walk = layout.charset_walk(); + let accent = code_of(self.stack.pop()) + .and_then(|c| seac_glyph(layout.charset, &self.guard.seac, c)) + .ok_or(())?; + self.work.tick(walk)?; + let base = code_of(self.stack.pop()) + .and_then(|c| seac_glyph(layout.charset, &self.guard.seac, c)) + .ok_or(())?; + self.work.tick(walk)?; + self.stack.truncate(self.stack.len().saturating_sub(2)); // dy dx + if !self.width && !self.stack.is_empty() { + self.stack.pop(); + self.width = true; + } + self.has_seac = true; + if depth == CFF_DEPTH_LIMIT { + return Err(()); + } + let base = layout.char_strings.get(u32::from(base)).ok_or(())?; + self.run(base, depth + 1)?; + let accent = layout.char_strings.get(u32::from(accent)).ok_or(())?; + self.run(accent, depth + 1)?; + } else if len == 1 && !self.width { + self.stack.pop(); + self.width = true; + } + Ok(()) + } +} + +/// `conv_subroutine_index` with `calc_subroutine_bias(count)`. +fn subr_index(index: f32, count: u32) -> Result { + let bias: i32 = if count < 1240 { + 107 + } else if count < 33900 { + 1131 + } else { + 32768 + }; + if !(index >= i32::MIN as f32 && index < i32::MAX as f32) { + return Err(()); + } + let index = (index as i32).checked_add(bias).ok_or(())?; + u32::try_from(index).map_err(|_| ()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/hash.rs b/src-tauri/src/pdf_engine/text_edit/fonts/hash.rs new file mode 100644 index 0000000..84a4a84 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/hash.rs @@ -0,0 +1,231 @@ +//! Font identity across files (SPEC §B.9.3, D32): `content_hash`, an FNV-1a over a canonical +//! serialisation of the resolved font dictionary tree. Object ids never enter it, so the same font +//! hashes alike in the source, the edited copy, the preview and the final file even though qpdf +//! renumbers every object. +//! +//! Streams the font parsers read (font programs, ToUnicode, CIDToGIDMap) are decoded once per +//! load and shared with the parsers. Every other stream (Type3 `/CharProcs`, `/Resources` +//! XObjects, embedded CMaps) is decoded only for the hash, and at most +//! `HASH_OTHER_DECODED_MAX` bytes of them per font: past that, such a stream hashes as its +//! dictionary and a fixed marker (the same in every copy of the file, since decoded sizes do not +//! change), so a large Type3 font costs the page budget a bounded amount. + +use super::{number_of, parsed_role, role_cap, Loader}; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeError}; +use crate::pdf_engine::text_edit::snapshot::{fnv1a_extend, fnv1a_u64}; +use lopdf::{Dictionary, Object, ObjectId, Stream}; + +/// Nested references followed by `content_hash` (§B.9.3). +const HASH_REF_DEPTH_MAX: usize = 16; +/// Values visited by one `content_hash` (bounds DAG blow-up; the marker keeps it deterministic). +const HASH_STEPS_MAX: usize = 1_000_000; +/// Nested values (direct and through references) one `content_hash` descends into: the +/// recursion depth, so the stack stays bounded whatever the dictionary tree looks like. +const HASH_NESTING_MAX: usize = 256; +/// Decoded bytes of streams no font parser reads that one `content_hash` may hash. +const HASH_OTHER_DECODED_MAX: usize = 4 << 20; + +struct Hasher { + state: u64, + steps: usize, + depth: usize, + /// Decode cap for the next referenced stream: that of the dictionary key leading to it. + cap: usize, + /// Whether that key names a stream a font parser reads (decoded once, shared). + parsed: bool, + /// What is left of `HASH_OTHER_DECODED_MAX`. + other_left: usize, +} + +impl Hasher { + fn feed(&mut self, bytes: &[u8]) { + self.state = fnv1a_extend(self.state, bytes); + } + + fn feed_len(&mut self, tag: u8, len: usize) { + self.feed(&[tag]); + self.feed(&(len as u64).to_le_bytes()); + } +} + +/// FNV-1a over a canonical serialisation of the resolved dict tree: keys sorted, numbers as +/// f64, string format ignored, references resolved (depth ≤ 16, cycles marked), streams by their +/// dictionary (minus `/Length`, `/Filter`, `/DecodeParms`, `/DL`) and decoded bytes — or raw +/// bytes plus filter when not decodable. Object ids never enter the hash (D32). +pub(super) fn content_hash<'a>(loader: &mut Loader<'a, '_>, dict: &'a Dictionary) -> u64 { + let mut h = Hasher { + state: fnv1a_u64(b"offpdf-font-v1"), + steps: 0, + depth: 0, + cap: role_cap(b""), + parsed: false, + other_left: HASH_OTHER_DECODED_MAX, + }; + let mut path = Vec::new(); + hash_dict(loader, dict, &mut h, &mut path, false); + h.state +} + +fn hash_dict<'a>( + loader: &mut Loader<'a, '_>, + dict: &'a Dictionary, + h: &mut Hasher, + path: &mut Vec, + stream_dict: bool, +) { + let mut entries: Vec<(&'a Vec, &'a Object)> = dict + .iter() + .filter(|(k, _)| { + !stream_dict || !matches!(k.as_slice(), b"Length" | b"Filter" | b"DecodeParms" | b"DL") + }) + .collect(); + entries.sort_by(|a, b| a.0.cmp(b.0)); + h.feed_len(b'd', entries.len()); + for (k, v) in entries { + h.feed_len(b'k', k.len()); + h.feed(k); + h.cap = role_cap(k); + h.parsed = parsed_role(k); + hash_obj(loader, v, h, path); + } +} + +fn hash_obj<'a>( + loader: &mut Loader<'a, '_>, + obj: &'a Object, + h: &mut Hasher, + path: &mut Vec, +) { + h.steps += 1; + if h.steps > HASH_STEPS_MAX || h.depth >= HASH_NESTING_MAX { + h.feed(b"X"); + return; + } + h.depth += 1; + hash_value(loader, obj, h, path); + h.depth -= 1; +} + +fn hash_value<'a>( + loader: &mut Loader<'a, '_>, + obj: &'a Object, + h: &mut Hasher, + path: &mut Vec, +) { + match obj { + Object::Null => h.feed(b"z"), + Object::Boolean(b) => h.feed(if *b { b"b1" } else { b"b0" }), + Object::Integer(_) | Object::Real(_) => { + h.feed(b"f"); + let v = number_of(obj).unwrap_or(f64::NAN); + h.feed(&v.to_bits().to_le_bytes()); + } + Object::Name(n) => { + h.feed_len(b'n', n.len()); + h.feed(n); + } + Object::String(s, _) => { + h.feed_len(b's', s.len()); + h.feed(s); + } + Object::Array(items) => { + h.feed_len(b'a', items.len()); + for item in items { + hash_obj(loader, item, h, path); + } + } + Object::Dictionary(d) => hash_dict(loader, d, h, path, false), + Object::Stream(s) => { + // Streams are only reachable through references; a bare one hashes as its raw form. + hash_dict(loader, &s.dict, h, path, true); + hash_raw_stream(loader, s, h, path); + } + Object::Reference(id) => hash_reference(loader, *id, h, path), + } +} + +fn hash_reference( + loader: &mut Loader<'_, '_>, + id: ObjectId, + h: &mut Hasher, + path: &mut Vec, +) { + if path.len() >= HASH_REF_DEPTH_MAX { + h.feed(b"D"); + return; + } + if path.contains(&id) { + h.feed(b"C"); + return; + } + let Some((target_id, target)) = loader.resolve_id(id) else { + h.feed(b"U"); + return; + }; + path.push(id); + let (cap, parsed) = (h.cap, h.parsed); + match target { + Object::Stream(s) if parsed => { + hash_dict(loader, &s.dict, h, path, true); + match loader.stream_bytes(target_id, s, cap) { + Ok(bytes) => { + h.feed_len(b'B', bytes.len()); + h.feed(&bytes); + } + Err(_) => hash_raw_stream(loader, s, h, path), + } + } + Object::Stream(s) => { + hash_dict(loader, &s.dict, h, path, true); + hash_other_stream(loader, s, h, path); + } + other => hash_obj(loader, other, h, path), + } + path.pop(); +} + +/// A stream only the hash reads: decoded (not kept) while `HASH_OTHER_DECODED_MAX` lasts, else +/// the marker `L`; raw bytes plus filter when it cannot be decoded at all. +fn hash_other_stream<'a>( + loader: &mut Loader<'a, '_>, + s: &'a Stream, + h: &mut Hasher, + path: &mut Vec, +) { + if h.other_left == 0 { + h.feed(b"L"); + return; + } + let before = loader.budget.remaining(); + match decode_stream(s, h.other_left, loader.budget) { + Ok(bytes) => { + h.other_left = h.other_left.saturating_sub(bytes.len()); + h.feed_len(b'B', bytes.len()); + h.feed(&bytes); + } + Err(DecodeError::TooLarge) => { + if before < h.other_left { + loader.budget_hit = true; // the page budget ran out, not this hash's share + } + h.other_left = 0; + h.feed(b"L"); + } + Err(_) => hash_raw_stream(loader, s, h, path), + } +} + +fn hash_raw_stream<'a>( + loader: &mut Loader<'a, '_>, + s: &'a Stream, + h: &mut Hasher, + path: &mut Vec, +) { + h.feed_len(b'R', s.content.len()); + h.feed(&s.content); + for key in [&b"Filter"[..], b"DecodeParms"] { + if let Ok(v) = s.dict.get(key) { + h.feed(key); + hash_obj(loader, v, h, path); + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/mod.rs b/src-tauri/src/pdf_engine/text_edit/fonts/mod.rs new file mode 100644 index 0000000..4577a15 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/mod.rs @@ -0,0 +1,731 @@ +//! Font model (SPEC §B.9): what each code of a page font reads as, how wide it is, whether its +//! glyph really draws (program-level proof), and which characters it can type (§A.2, §A.3). +//! +//! Loading never fails: every problem becomes `FontModel::refusal` (run-level reason codes of +//! §A.10). A font whose loading ran out of the page decode budget gets `PAGE_TOO_COMPLEX` (a +//! page-level reason the walker turns into the page refusal) and is not cached, so a later page +//! with a fresh budget loads it again. All streams are read through `decode::decode_stream` +//! under the caller's `DecodeBudget`, which also pays for the glyph-proof work +//! (`glyph_budget.rs`) and the memory of each model built; nothing here panics on PDF data. + +pub(crate) mod agl; +mod agl_data; +pub(crate) mod cff_encoding; +pub(crate) mod cff_layout; +mod encoding_tables; +pub(crate) mod encodings; +pub(crate) mod faces; +pub(crate) mod glyph_budget; +mod hash; +pub(crate) mod program; +pub(crate) mod simple; +pub(crate) mod std14; +mod std14_data; +pub(crate) mod tounicode; +pub(crate) mod type0; +pub(crate) mod type1; + +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget, DecodeError}; +use crate::pdf_engine::text_edit::limits::{ + CIDTOGID_MAX_BYTES, FONT_PROGRAM_MAX_DECODED, STREAM_MAX_DECODED, TOUNICODE_MAX_DECODED, +}; +use crate::pdf_engine::text_edit::reasons::{Face, TextReason}; +use crate::pdf_engine::text_edit::snapshot::fnv1a_u64; +use lopdf::{Dictionary, Document, Object, ObjectId, Stream}; +use std::collections::{BTreeMap, HashMap, VecDeque}; +use std::sync::{Arc, Mutex}; + +/// Approximate bytes of font models one `FontCache` keeps (oldest models are dropped first). +const FONT_CACHE_BYTES_MAX: usize = 128 << 20; +/// References followed for one value (a reference to a reference …). +const REF_HOPS_MAX: usize = 16; + +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] +pub enum FontKey { + Indirect(ObjectId), + Direct { owner: ObjectId, name_hash: u64 }, +} + +impl FontKey { + /// The key of a font dictionary written directly inside the `/Font` dictionary of `owner`. + pub fn direct(owner: ObjectId, resource_name: &[u8]) -> FontKey { + FontKey::Direct { + owner, + name_hash: fnv1a_u64(resource_name), + } + } +} + +/// One character code: `len` 1 (simple fonts) or 2 (Identity-H), big-endian `value`. +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, PartialOrd, Ord)] +pub struct Code { + pub value: u32, + pub len: u8, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, serde::Serialize)] +#[serde(rename_all = "camelCase")] +pub enum FamilyHint { + Serif, + Sans, + Mono, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum FontClass { + SimpleTrueType, + SimpleCff, + SimpleOpenType, + SimpleType1, + Std14(std14::Std14Face), + SimpleNonEmbedded, + Type0Cid2, + Type0Cid0, +} + +#[derive(Debug, Clone)] +pub struct CodeInfo { + pub text: Option, // display text (ligatures expanded); None = undecodable + pub glyph_name: Option, + pub gid: Option, + pub width1000: Option, // glyph-space width ×1000; None = unknown + pub drawable: bool, // A.2 presence rule + pub typeable_as: Option, // Some(X) iff A.3.2 holds +} + +#[derive(Debug, Clone)] +pub struct FontModel { + pub key: FontKey, + pub content_hash: u64, + pub base_name: String, + pub display_name: String, + pub family_key: String, + pub bold: bool, + pub italic: bool, + pub family_hint: FamilyHint, + pub class: Option, + pub refusal: Option, + pub embedded: bool, + pub subset: bool, + pub substituted: bool, + pub vertical: bool, + pub ascent: f64, + pub descent: f64, + codes: BTreeMap, + writable: BTreeMap>, + two_byte: bool, + cid_widths: Vec<(u32, f64)>, + default_width: f64, + approx_bytes: usize, +} + +impl FontModel { + /// A model with no codes yet; loaders fill it and call `finish`. + fn shell(key: FontKey, content_hash: u64, two_byte: bool) -> FontModel { + FontModel { + key, + content_hash, + base_name: String::new(), + display_name: String::new(), + family_key: String::new(), + bold: false, + italic: false, + family_hint: FamilyHint::Sans, + class: None, + refusal: None, + embedded: false, + subset: false, + substituted: false, + vertical: false, + ascent: DEFAULT_ASCENT, + descent: DEFAULT_DESCENT, + codes: BTreeMap::new(), + writable: BTreeMap::new(), + two_byte, + cid_widths: Vec::new(), + default_width: 0.0, + approx_bytes: 0, + } + } + + /// Names, family key, faces and family hint from the BaseFont and descriptor hints. + fn set_names(&mut self, raw_base: &str, hints: &faces::FaceHints) { + let (base, subset) = faces::strip_subset_tag(raw_base); + self.base_name = base.to_string(); + self.subset = subset; + self.display_name = faces::display_name(raw_base); + self.family_key = faces::family_key(raw_base); + self.bold = faces::is_bold(raw_base, hints); + self.italic = faces::is_italic(raw_base, hints); + self.family_hint = faces::family_hint(raw_base, hints.flags); + } + + fn refuse(&mut self, reason: TextReason) { + self.refusal = Some(self.refusal.map_or(reason, |r| r.min(reason))); + } + + /// Final invariants: a refused font has no class and types nothing; builds the alphabet. + fn finish(mut self) -> FontModel { + if self.refusal.is_some() { + self.class = None; + for info in self.codes.values_mut() { + info.typeable_as = None; + } + } + let len = if self.two_byte { 2 } else { 1 }; + let mut writable: BTreeMap> = BTreeMap::new(); + for (value, info) in &self.codes { + if let Some(ch) = info.typeable_as { + writable + .entry(ch) + .or_default() + .push(Code { value: *value, len }); + } + } + self.writable = writable; + let strings: usize = self + .codes + .values() + .map(|i| { + i.text.as_ref().map_or(0, String::len) + + i.glyph_name.as_ref().map_or(0, String::len) + }) + .sum(); + self.approx_bytes = 512 + + self.codes.len() * (std::mem::size_of::() + 48) + + strings + + self.writable.len() * 64 + + self.cid_widths.len() * 16; + self + } + + /// Approximate heap bytes of this model (its codes, their strings, the alphabet and the CID + /// widths): what a page model that keeps it is charged (`walker::budget::ModelBudget::font`). + pub fn approx_bytes(&self) -> usize { + self.approx_bytes + } + + fn code_len(&self) -> u8 { + if self.two_byte { + 2 + } else { + 1 + } + } + + /// Splits shown string bytes into codes; an odd length for a 2-byte font is + /// `AMBIGUOUS_UNICODE`. + pub fn split_codes(&self, bytes: &[u8]) -> Result, TextReason> { + if !self.two_byte { + return Ok(bytes + .iter() + .map(|b| Code { + value: u32::from(*b), + len: 1, + }) + .collect()); + } + if bytes.len() % 2 != 0 { + return Err(TextReason::AmbiguousUnicode); + } + Ok(bytes + .chunks_exact(2) + .map(|p| match p { + [hi, lo] => Code { + value: u32::from(u16::from_be_bytes([*hi, *lo])), + len: 2, + }, + _ => Code { value: 0, len: 2 }, + }) + .collect()) + } + + pub fn info(&self, code: Code) -> Option<&CodeInfo> { + if code.len != self.code_len() { + return None; + } + self.codes.get(&code.value) + } + + pub fn text(&self, code: Code) -> Option<&str> { + self.info(code).and_then(|i| i.text.as_deref()) + } + + /// Glyph-space width ×1000 used for reading; unknown ⇒ 0.0. + pub fn width(&self, code: Code) -> f64 { + if code.len != self.code_len() { + return 0.0; + } + if let Some(info) = self.codes.get(&code.value) { + return info.width1000.unwrap_or(0.0); + } + if !self.two_byte { + return 0.0; + } + match self + .cid_widths + .binary_search_by(|(c, _)| c.cmp(&code.value)) + { + Ok(i) => self.cid_widths.get(i).map_or(0.0, |(_, w)| *w), + Err(_) => self.default_width, + } + } + + /// `Tw` applies: a single-byte code 32. + pub fn is_word_space(&self, code: Code) -> bool { + code.len == 1 && code.value == 32 + } + + pub fn drawable(&self, code: Code) -> bool { + self.info(code).is_some_and(|i| i.drawable) + } + + /// Typeable characters with the width of their lowest code, sorted by character. + pub fn alphabet(&self) -> Vec<(char, f64)> { + self.writable + .iter() + .filter_map(|(ch, codes)| { + let code = codes.first()?; + Some((*ch, self.info(*code)?.width1000?)) + }) + .collect() + } + + /// Whether `code` is one of the codes that type `ch`. + pub fn can_write(&self, ch: char, code: Code) -> bool { + self.writable.get(&ch).is_some_and(|c| c.contains(&code)) + } + + /// The code to write `ch` with (§A.3.2): the run's own code for it, then the page's, then the + /// lowest. + pub fn code_for( + &self, + ch: char, + prefer_run: &[(char, Code)], + prefer_page: &[(char, Code)], + ) -> Option { + let codes = self.writable.get(&ch)?; + prefer_run + .iter() + .chain(prefer_page) + .find(|(c, code)| *c == ch && codes.contains(code)) + .map(|(_, code)| *code) + .or_else(|| codes.first().copied()) + } +} + +/// Fonts already loaded for one snapshot, keyed by `FontKey`. +pub struct FontCache { + inner: Mutex, +} + +#[derive(Default)] +struct CacheInner { + map: HashMap>, + order: VecDeque, + bytes: usize, +} + +impl Default for FontCache { + fn default() -> Self { + Self::new() + } +} + +impl FontCache { + pub fn new() -> Self { + FontCache { + inner: Mutex::new(CacheInner::default()), + } + } + + /// Never fails: every problem becomes `FontModel.refusal`. + pub fn get_or_load( + &self, + doc: &Document, + key: FontKey, + dict: &Dictionary, + budget: &mut DecodeBudget, + ) -> Arc { + if let Some(model) = self.lock().map.get(&key) { + return Arc::clone(model); + } + let (model, cacheable) = load_font(doc, key, dict, budget); + let model = Arc::new(model); + if cacheable { + let mut inner = self.lock(); + if let Some(existing) = inner.map.get(&key) { + return Arc::clone(existing); + } + inner.bytes = inner.bytes.saturating_add(model.approx_bytes); + inner.map.insert(key, Arc::clone(&model)); + inner.order.push_back(key); + while inner.bytes > FONT_CACHE_BYTES_MAX && inner.order.len() > 1 { + let Some(old) = inner.order.pop_front() else { + break; + }; + if let Some(gone) = inner.map.remove(&old) { + inner.bytes = inner.bytes.saturating_sub(gone.approx_bytes); + } + } + } + model + } + + fn lock(&self) -> std::sync::MutexGuard<'_, CacheInner> { + self.inner.lock().unwrap_or_else(|e| e.into_inner()) + } +} + +pub struct TypingSurface { + pub fonts: Vec<(Vec /*resource name*/, Arc)>, // primary first +} + +/// Group page fonts by (family_key, bold, italic); siblings = same group, not refused, same font +/// class family (simple↔simple, Type0↔Type0) or the Word pattern (one simple TrueType + one Type0 +/// CIDFontType2 with the same base_name). The primary comes first. +pub fn typing_surface(page_fonts: &[(Vec, Arc)], primary: &[u8]) -> TypingSurface { + faces::typing_surface(page_fonts, primary) +} + +/// The sibling face group (§A.7 bold/italic) of `primary` for `face`, led by a font of the +/// primary's class family; `None` when the page has no such group. +pub fn face_surface( + page_fonts: &[(Vec, Arc)], + primary: &FontModel, + face: Face, +) -> Option { + faces::face_surface(page_fonts, primary, face) +} + +// ---- Loading ------------------------------------------------------------------------------ + +const DEFAULT_ASCENT: f64 = 0.8; +const DEFAULT_DESCENT: f64 = -0.2; + +fn load_font<'a>( + doc: &'a Document, + key: FontKey, + dict: &'a Dictionary, + budget: &mut DecodeBudget, +) -> (FontModel, bool) { + let mut loader = Loader::new(doc, budget); + let hash = hash::content_hash(&mut loader, dict); + let subtype = loader.name(dict, b"Subtype").unwrap_or_default().to_vec(); + let model = match subtype.as_slice() { + b"Type0" => type0::load(&mut loader, FontModel::shell(key, hash, true), dict), + b"Type1" | b"MMType1" | b"TrueType" => simple::load( + &mut loader, + FontModel::shell(key, hash, false), + dict, + &subtype, + ), + b"Type3" => simple::load_type3(&mut loader, FontModel::shell(key, hash, false), dict), + _ => { + let mut model = FontModel::shell(key, hash, false); + let base = loader.name(dict, b"BaseFont").unwrap_or_default(); + model.set_names(&String::from_utf8_lossy(base), &faces::FaceHints::default()); + model.refuse(TextReason::FontUnsupported); + model + } + }; + let mut model = model.finish(); + // The model's memory is paid from the page budget too: a few bytes of CMap can describe + // 65,536 codes, and a page may hold up to 256 fonts. + if !loader.budget_hit && loader.budget.take(model.approx_bytes).is_err() { + loader.budget_hit = true; + } + if loader.budget_hit { + model.refusal = Some(TextReason::PageTooComplex); + return (model.finish(), false); + } + (model, true) +} + +/// Why a stream could not be read. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum StreamFail { + /// The page decode budget ran out (transient: the model is not cached). + Budget, + /// The stream itself is undecodable, too large for its role, or missing. + Bad, +} + +/// One font load: the document, the page budget and a memo of decoded streams (each stream is +/// decoded once per cap, for the content hash and its parser alike: both use the cap of the key +/// that references it, `role_cap`). +pub(crate) struct Loader<'a, 'b> { + doc: &'a Document, + budget: &'b mut DecodeBudget, + decoded: HashMap<(ObjectId, usize), Result>, StreamFail>>, + budget_hit: bool, +} + +impl<'a, 'b> Loader<'a, 'b> { + fn new(doc: &'a Document, budget: &'b mut DecodeBudget) -> Self { + Loader { + doc, + budget, + decoded: HashMap::new(), + budget_hit: false, + } + } + + /// Follows references (≤ 16 hops); `None` when one dangles. Returns the last object id. + pub(crate) fn resolve(&self, obj: &'a Object) -> Option<(Option, &'a Object)> { + let mut obj = obj; + let mut id = None; + for _ in 0..REF_HOPS_MAX { + match obj { + Object::Reference(r) => { + id = Some(*r); + obj = self.doc.objects.get(r)?; + } + other => return Some((id, other)), + } + } + None + } + + /// The object `id` names, resolved, with the id of the object finally reached. + fn resolve_id(&self, id: ObjectId) -> Option<(ObjectId, &'a Object)> { + let first = self.doc.objects.get(&id)?; + let (last, obj) = self.resolve(first)?; + Some((last.unwrap_or(id), obj)) + } + + /// `dict[key]`, resolved; `None` when absent, null or dangling. + pub(crate) fn get(&self, dict: &'a Dictionary, key: &[u8]) -> Option<&'a Object> { + let (_, obj) = self.resolve(dict.get(key).ok()?)?; + (!matches!(obj, Object::Null)).then_some(obj) + } + + pub(crate) fn name(&self, dict: &'a Dictionary, key: &[u8]) -> Option<&'a [u8]> { + match self.get(dict, key)? { + Object::Name(n) => Some(n), + _ => None, + } + } + + pub(crate) fn dict(&self, dict: &'a Dictionary, key: &[u8]) -> Option<&'a Dictionary> { + match self.get(dict, key)? { + Object::Dictionary(d) => Some(d), + _ => None, + } + } + + /// A finite number (integer or real). + pub(crate) fn number(&self, dict: &'a Dictionary, key: &[u8]) -> Option { + number_of(self.get(dict, key)?) + } + + /// The stream `dict[key]` refers to, with its object id. + pub(crate) fn stream( + &self, + dict: &'a Dictionary, + key: &[u8], + ) -> Option<(ObjectId, &'a Stream)> { + match self.resolve(dict.get(key).ok()?)? { + (Some(id), Object::Stream(s)) => Some((id, s)), + _ => None, + } + } + + /// A meter for glyph-proof work, holding what is left of the page budget. + pub(crate) fn work_meter(&self) -> glyph_budget::WorkMeter { + glyph_budget::WorkMeter::new(self.budget.remaining()) + } + + /// Debits the work a meter recorded; a meter that ran out marks the load (`PAGE_TOO_COMPLEX`, + /// not cached). + pub(crate) fn settle(&mut self, meter: &glyph_budget::WorkMeter) { + let used = meter.used().min(self.budget.remaining()); + if meter.exhausted() || self.budget.take(used).is_err() { + self.budget_hit = true; + } + } + + /// Decoded bytes of a stream (memoised), capped at `cap` while decoding. Running out of the + /// page budget is `Budget` (and marks the load); everything else is `Bad`. + pub(crate) fn stream_bytes( + &mut self, + id: ObjectId, + stream: &Stream, + cap: usize, + ) -> Result>, StreamFail> { + let decoded = match self.decoded.get(&(id, cap)) { + Some(result) => result.clone(), + None => { + let before = self.budget.remaining(); + let result = match decode_stream(stream, cap, self.budget) { + Ok(bytes) => Ok(Arc::new(bytes)), + Err(DecodeError::TooLarge) if before < cap => Err(StreamFail::Budget), + Err(_) => Err(StreamFail::Bad), + }; + self.decoded.insert((id, cap), result.clone()); + result + } + }; + if decoded == Err(StreamFail::Budget) { + self.budget_hit = true; + } + decoded + } +} + +/// The decode cap of a stream by the key that references it (§B.2): font programs, ToUnicode, +/// CIDToGIDMap, anything else. +pub(crate) fn role_cap(key: &[u8]) -> usize { + match key { + b"FontFile" | b"FontFile2" | b"FontFile3" => FONT_PROGRAM_MAX_DECODED, + b"ToUnicode" => TOUNICODE_MAX_DECODED, + b"CIDToGIDMap" => CIDTOGID_MAX_BYTES, + _ => STREAM_MAX_DECODED, + } +} + +/// Whether a stream under `key` is read by a font parser (and so decoded once and shared with the +/// content hash). +pub(crate) fn parsed_role(key: &[u8]) -> bool { + matches!( + key, + b"FontFile" | b"FontFile2" | b"FontFile3" | b"ToUnicode" | b"CIDToGIDMap" + ) +} + +/// A finite number from an integer or real object. lopdf keeps reals as `f32`; widening one +/// directly would turn a written `0.001` into 0.0010000000474974513, so a real goes through its +/// shortest decimal form (what the file wrote, to f32 precision) — the value a viewer parsing the +/// text with `f64` gets. +pub(crate) fn number_of(obj: &Object) -> Option { + let v = match obj { + Object::Integer(i) => *i as f64, + Object::Real(r) => r.to_string().parse::().unwrap_or(f64::from(*r)), + _ => return None, + }; + v.is_finite().then_some(v) +} + +/// FontDescriptor values shared by every font class (§B.9.3). +#[derive(Debug, Clone, Default)] +pub(crate) struct Descriptor { + pub flags: Option, + pub ascent: Option, + pub descent: Option, + pub italic_angle: Option, + pub stem_v: Option, + pub missing_width: Option, + /// `/CharSet` names; present but not a string → an empty set (no Type1 glyph is proven). + pub charset: Option>, +} + +impl Descriptor { + /// Reads `/FontDescriptor` of `font` (absent, null or a reference to nothing → all `None`; + /// anything else that is not a dictionary → `FONT_UNSUPPORTED`). + pub(crate) fn read<'a>( + loader: &Loader<'a, '_>, + font: &'a Dictionary, + ) -> Result { + let d = match loader.get(font, b"FontDescriptor") { + None => return Ok(Descriptor::default()), + Some(Object::Dictionary(d)) => d, + Some(_) => return Err(TextReason::FontUnsupported), + }; + let flags = match loader.get(d, b"Flags") { + None => None, + Some(Object::Integer(f)) => Some(*f), + Some(_) => return Err(TextReason::FontUnsupported), + }; + let charset = match loader.get(d, b"CharSet") { + None => None, + Some(Object::String(s, _)) => Some(type1::charset_names(s)), + Some(_) => Some(std::collections::HashSet::new()), + }; + Ok(Descriptor { + flags, + ascent: loader.number(d, b"Ascent"), + descent: loader.number(d, b"Descent"), + italic_angle: loader.number(d, b"ItalicAngle"), + stem_v: loader.number(d, b"StemV"), + missing_width: loader.number(d, b"MissingWidth"), + charset, + }) + } + + pub(crate) fn hints(&self) -> faces::FaceHints { + faces::FaceHints { + flags: self.flags, + stem_v: self.stem_v, + italic_angle: self.italic_angle, + } + } + + pub(crate) fn symbolic(&self) -> bool { + self.flags.is_some_and(|f| f & 4 != 0) + } +} + +/// Ascent/descent in em (§A.5): descriptor (non-zero) clamped to [0.5, 1.5] / [−0.8, 0], else the +/// program's, else `fallback` (AFM), else 0.8 / −0.2. +pub(crate) fn vertical_metrics( + desc: &Descriptor, + program: Option<(f64, f64)>, + fallback: Option<(f64, f64)>, +) -> (f64, f64) { + let pick = |d: Option, p: Option, f: Option, default: f64, lo: f64, hi: f64| { + d.filter(|v| *v != 0.0) + .map(|v| v / 1000.0) + .or(p.filter(|v| *v != 0.0)) + .or(f) + .unwrap_or(default) + .clamp(lo, hi) + }; + ( + pick( + desc.ascent, + program.map(|p| p.0), + fallback.map(|f| f.0), + DEFAULT_ASCENT, + 0.5, + 1.5, + ), + pick( + desc.descent, + program.map(|p| p.1), + fallback.map(|f| f.1), + DEFAULT_DESCENT, + -0.8, + 0.0, + ), + ) +} + +/// U+0020 and U+00A0: glyphs that may draw nothing (§A.2, §A.3.2). +pub(crate) fn is_whitespace_char(text: Option<&str>) -> bool { + matches!(text, Some(" ") | Some("\u{a0}")) +} + +#[cfg(test)] +impl FontModel { + /// Big-endian `len` bytes of `code`. + pub fn code_bytes(&self, code: Code) -> Vec { + let bytes = code.value.to_be_bytes(); + let len = usize::from(code.len).min(4); + bytes.get(4 - len..).unwrap_or_default().to_vec() + } +} + +#[cfg(test)] +mod tests; +#[cfg(test)] +mod tests_bounds; +#[cfg(test)] +mod tests_classes; +#[cfg(test)] +mod tests_fuzz; +#[cfg(test)] +mod tests_presence; +#[cfg(test)] +mod tests_tounicode; +#[cfg(test)] +mod tests_work; diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/program.rs b/src-tauri/src/pdf_engine/text_edit/fonts/program.rs new file mode 100644 index 0000000..78d1741 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/program.rs @@ -0,0 +1,384 @@ +//! Glyph presence in embedded TrueType, OpenType and CFF programs (SPEC §A.2, §B.9.5). +//! ttf-parser is panic-free by design; every lookup here returns `None` instead of failing. +//! +//! A glyph counts as drawn only when its outline has at least one line or curve segment (a +//! bounding box alone, e.g. from a lone `moveto`, draws nothing). ttf-parser outlines a glyph only +//! after the bounded pre-check of `glyph_budget.rs` has walked it within budget. + +use super::cff_layout::CffLayout; +use super::encodings::mac_roman_code; +use super::glyph_budget::{CffGuard, GlyfGuard, WorkMeter}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use std::collections::{BTreeMap, HashMap, HashSet}; +use ttf_parser::{cff, cmap::Subtable, glyf, Face, GlyphId, OutlineBuilder, PlatformId, Tag}; + +/// Counts drawing segments of an outline. +#[derive(Default)] +struct SegCount(usize); + +impl OutlineBuilder for SegCount { + fn move_to(&mut self, _x: f32, _y: f32) {} + fn line_to(&mut self, _x: f32, _y: f32) { + self.0 = self.0.saturating_add(1); + } + fn quad_to(&mut self, _x1: f32, _y1: f32, _x: f32, _y: f32) { + self.0 = self.0.saturating_add(1); + } + fn curve_to(&mut self, _x1: f32, _y1: f32, _x2: f32, _y2: f32, _x: f32, _y: f32) { + self.0 = self.0.saturating_add(1); + } + fn close(&mut self) {} +} + +/// Where a program's outlines live, each with its bounded pre-check. +pub enum Outlines<'a> { + /// `glyf` (outlined at the default instance: `gvar` deltas are zero there, so the `gvar` + /// path ttf-parser's `Face::outline_glyph` would take draws the same segments). + Glyf { + guard: GlyfGuard<'a>, + table: glyf::Table<'a>, + }, + /// A CFF program (bare, or an OpenType `CFF ` table). + Cff { + guard: CffGuard<'a>, + table: cff::Table<'a>, + }, + /// Only a `CFF2` table: its charstrings are not pre-checked (`FONT_PROGRAM_UNSUPPORTED`). + Cff2Only, + /// Both a `glyf` and a `CFF `/`CFF2` table (`FONT_PROGRAM_UNSUPPORTED`): an OpenType font has + /// one or the other, and viewers pick by the sfnt tag (FreeType, Poppler and pdf.js draw an + /// `OTTO` program's CFF outlines), so a presence proven on `glyf` can certify a glyph the + /// viewer draws empty (review T4 r2 M-2). + Conflicting, + /// A `glyf` or `CFF ` table the pre-check cannot lay out (`FONT_PROGRAM_UNREADABLE`). + Unreadable, + /// No outline table at all: no glyph draws. + None, +} + +impl<'a> Outlines<'a> { + pub fn of_face(face: &Face<'a>) -> Outlines<'a> { + let tables = face.tables(); + let has = |tag: &[u8; 4]| face.raw_face().table(Tag::from_bytes(tag)).is_some(); + if has(b"glyf") && (has(b"CFF ") || has(b"CFF2")) { + return Outlines::Conflicting; + } + if let Some(table) = tables.glyf { + return match GlyfGuard::new(face) { + Some(guard) => Outlines::Glyf { guard, table }, + None => Outlines::Unreadable, + }; + } + if let Some(table) = tables.cff { + return Self::of_cff_table(table, face.raw_face().table(Tag::from_bytes(b"CFF "))); + } + if tables.cff2.is_some() { + return Outlines::Cff2Only; + } + Outlines::None + } + + /// A bare CFF program (`FontFile3 /Type1C`, `/CIDFontType0C`). + pub fn of_cff(data: &'a [u8]) -> Option> { + let table = cff::Table::parse(data)?; + Some(Self::of_cff_table(table, Some(data))) + } + + fn of_cff_table(table: cff::Table<'a>, data: Option<&'a [u8]>) -> Outlines<'a> { + match data.and_then(|d| CffGuard::new(d, table.number_of_glyphs())) { + Some(guard) => Outlines::Cff { guard, table }, + None => Outlines::Unreadable, + } + } + + /// The font refusal this outline source implies, if any. + pub fn refusal(&self) -> Option { + match self { + Outlines::Cff2Only | Outlines::Conflicting => Some(TextReason::FontProgramUnsupported), + Outlines::Unreadable => Some(TextReason::FontProgramUnreadable), + Outlines::Glyf { .. } | Outlines::Cff { .. } | Outlines::None => None, + } + } + + /// The CFF layout, when the outlines are CFF. + pub fn cff_layout(&self) -> Option<&CffLayout<'a>> { + match self { + Outlines::Cff { guard, .. } => Some(guard.layout()), + _ => None, + } + } + + /// Whether `gid` draws at least one segment. The pre-check runs first and charges `meter`; + /// ttf-parser outlines only a glyph that passed it. + pub fn drawn(&self, gid: u16, meter: &mut WorkMeter) -> bool { + let mut count = SegCount::default(); + match self { + Outlines::Glyf { guard, table } => { + guard.check(gid, meter) + && table.outline(GlyphId(gid), &mut count).is_some() + && count.0 > 0 + } + Outlines::Cff { guard, table } => { + guard.check(gid, meter) + && table.outline(GlyphId(gid), &mut count).is_ok() + && count.0 > 0 + } + Outlines::Cff2Only | Outlines::Conflicting | Outlines::Unreadable | Outlines::None => { + false + } + } + } +} + +/// Outline results per GID for one font load: each glyph is outlined at most once, however +/// many codes map to it. +#[derive(Default)] +pub struct OutlineMemo(HashMap); + +impl OutlineMemo { + pub fn get(&mut self, gid: u16, outline: impl FnOnce() -> bool) -> bool { + *self.0.entry(gid).or_insert_with(outline) + } +} + +/// The outcome of resolving a code through every applicable strategy. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum GidLookup { + /// No strategy produced a glyph. + None, + /// Every strategy that produced a glyph produced this one. + Agree(u16), + /// Two strategies produced different glyphs (never drawable). + Disagree, +} + +impl GidLookup { + fn add(self, gid: Option) -> GidLookup { + match (self, gid) { + (_, None) => self, + (GidLookup::None, Some(g)) => GidLookup::Agree(g.0), + (GidLookup::Agree(a), Some(g)) if a == g.0 => self, + _ => GidLookup::Disagree, + } + } +} + +fn subtable<'a>(face: &Face<'a>, platform: PlatformId, encoding: u16) -> Option> { + face.tables() + .cmap? + .subtables + .into_iter() + .find(|s| s.platform_id == platform && s.encoding_id == encoding) +} + +/// Glyph name → GID through a `post` format 2 table, built once per face in O(glyphs). +/// ttf-parser's own `glyph_index_by_name` scans the 258 Macintosh names and every glyph index +/// per call (and `glyph_name` walks the custom names), which a 256-code font load repeats. +/// Same answers for conforming tables: a name used through a standard Macintosh index maps to +/// the first such glyph; a custom name to the first glyph using its first position. +pub struct PostNames { + by_name: HashMap, +} + +impl PostNames { + pub fn new(face: &Face<'_>) -> PostNames { + let mut by_name = HashMap::new(); + if let Some(data) = face.raw_face().table(Tag::from_bytes(b"post")) { + Self::fill(face, data, &mut by_name); + } + PostNames { by_name } + } + + fn fill(face: &Face<'_>, data: &[u8], by_name: &mut HashMap) -> Option<()> { + if data.get(..4)? != [0, 2, 0, 0] { + return None; // only version 2.0 carries names + } + let count = usize::from(u16::from_be_bytes([*data.get(32)?, *data.get(33)?])); + let names_at = 34usize.checked_add(count.checked_mul(2)?)?; + let indexes: Vec = data + .get(34..names_at)? + .chunks_exact(2) + .map(|p| { + u16::from_be_bytes([ + p.first().copied().unwrap_or(0), + p.get(1).copied().unwrap_or(0), + ]) + }) + .collect(); + let mut first_gid_of_index: HashMap = HashMap::new(); + for (gid, index) in indexes.iter().enumerate() { + let gid = u16::try_from(gid).ok()?; + first_gid_of_index.entry(*index).or_insert(gid); + if *index < 258 { + if let Some(name) = face.glyph_name(GlyphId(gid)) { + by_name.entry(name.to_string()).or_insert(gid); + } + } + } + // Custom names: Pascal strings; the list ends at an empty, truncated or non-UTF-8 one. + let mut seen: HashSet = HashSet::new(); + let mut at = names_at; + let mut position: u16 = 0; + while let Some(len) = data.get(at).map(|l| usize::from(*l)) { + let name = data.get(at + 1..at + 1 + len).filter(|_| len > 0)?; + let name = std::str::from_utf8(name).ok()?; + at += 1 + len; + if seen.insert(name.to_string()) && !by_name.contains_key(name) { + let index = 258u16.checked_add(position)?; + if let Some(gid) = first_gid_of_index.get(&index) { + by_name.insert(name.to_string(), *gid); + } + } + position = position.checked_add(1)?; + } + Some(()) + } + + pub fn gid(&self, name: &str) -> Option { + self.by_name.get(name).map(|g| GlyphId(*g)) + } +} + +/// The lookup tables of one TrueType/OpenType face, found once per font load. +pub struct TrueTypeLookup<'f, 'a> { + pub face: &'f Face<'a>, + post: PostNames, + /// OpenType-CFF: the CFF glyph names. + cff_names: Option, + win_unicode: Option>, + win_symbol: Option>, + mac_roman: Option>, +} + +impl<'f, 'a> TrueTypeLookup<'f, 'a> { + pub fn new(face: &'f Face<'a>, outlines: &Outlines<'a>) -> Self { + let cff_names = match (outlines.cff_layout(), face.tables().cff) { + (Some(layout), Some(table)) => Some(CffNames::new(layout, &table)), + _ => None, + }; + TrueTypeLookup { + face, + post: PostNames::new(face), + cff_names, + win_unicode: subtable(face, PlatformId::Windows, 1), + win_symbol: subtable(face, PlatformId::Windows, 0), + mac_roman: subtable(face, PlatformId::Macintosh, 0), + } + } + + /// Non-symbolic TrueType (PDF 32000-1 §9.6.6.4): glyph name → Unicode (AGL) → (3,1) cmap; + /// name → MacRoman code → (1,0) cmap; `post` (and, for OpenType-CFF, CFF) glyph names; the + /// name's AGL character already resolved by the caller. + pub fn by_name_char(&self, name: &str, name_char: Option) -> GidLookup { + let mut out = GidLookup::None; + if name == ".notdef" { + return out; + } + if let (Some(ch), Some(table)) = (name_char, self.win_unicode.as_ref()) { + out = out.add(table.glyph_index(u32::from(ch))); + } + if let (Some(code), Some(table)) = (mac_roman_code(name), self.mac_roman.as_ref()) { + out = out.add(table.glyph_index(u32::from(code))); + } + let by_name = self.post.gid(name).or_else(|| { + self.cff_names + .as_ref() + .and_then(|names| names.gid(name)) + .map(GlyphId) + }); + out.add(by_name) + } + + /// Symbolic TrueType: (3,0) at `code`, `0xF000+code`, `0xF100+code`, `0xF200+code`; (1,0) + /// at `code`. + pub fn symbolic(&self, code: u8) -> GidLookup { + let mut out = GidLookup::None; + let code = u32::from(code); + if let Some(table) = self.win_symbol.as_ref() { + for base in [0, 0xF000, 0xF100, 0xF200] { + out = out.add(table.glyph_index(base + code)); + } + } + if let Some(table) = self.mac_roman.as_ref() { + out = out.add(table.glyph_index(code)); + } + out + } +} + +/// Glyph name → first GID of a name-keyed CFF, built once per font load in O(glyphs) from the +/// charset walked once (ttf-parser's `glyph_name` walks a format 1/2 charset per call, and +/// `glyph_index_by_name` scans the 391 standard strings per call). Same answers as ttf-parser's +/// `glyph_name` for every GID. CID-keyed fonts have no names. +pub struct CffNames { + by_name: HashMap, +} + +impl CffNames { + pub fn new(layout: &CffLayout<'_>, table: &cff::Table<'_>) -> CffNames { + let mut by_name = HashMap::new(); + if layout.is_cid() { + return CffNames { by_name }; + } + match layout.gid_to_sid() { + Some(sids) => { + for (gid, sid) in sids.iter().enumerate() { + let name = sid.and_then(|s| layout.sid_name(s)); + if let (Some(name), Ok(gid)) = (name, u16::try_from(gid)) { + by_name.entry(name.to_string()).or_insert(gid); + } + } + } + // The predefined Expert charsets: ttf-parser's lookups are O(1) per glyph there. + None => { + for gid in 0..table.number_of_glyphs() { + if let Some(name) = table.glyph_name(GlyphId(gid)) { + by_name.entry(name.to_string()).or_insert(gid); + } + } + } + } + CffNames { by_name } + } + + pub fn gid(&self, name: &str) -> Option { + self.by_name.get(name).copied() + } +} + +/// A CID-keyed CFF's CID → GID map (the inverse of its charset, first GID per CID), or `None` +/// for a name-keyed CFF (where GID = CID). +pub fn cff_cid_to_gid(layout: &CffLayout<'_>) -> Option> { + if !layout.is_cid() { + return None; + } + let mut map = BTreeMap::new(); + for (gid, cid) in layout.gid_to_sid()?.iter().enumerate() { + if let (Some(cid), Ok(gid)) = (cid, u16::try_from(gid)) { + map.entry(*cid).or_insert(gid); + } + } + Some(map) +} + +/// Font-wide vertical metrics of a TrueType/OpenType program: (ascent, descent) in em, from +/// `hhea` (OS/2 typographic values when `hhea` is empty). +pub fn face_metrics(face: &Face<'_>) -> Option<(f64, f64)> { + let upem = f64::from(face.units_per_em()); + if upem <= 0.0 { + return None; + } + let (asc, desc) = match (face.ascender(), face.descender()) { + (0, 0) => (face.typographic_ascender()?, face.typographic_descender()?), + pair => pair, + }; + Some((f64::from(asc) / upem, f64::from(desc) / upem)) +} + +#[cfg(test)] +impl TrueTypeLookup<'_, '_> { + /// `by_name_char` with the name's AGL character looked up here. + pub fn by_name(&self, name: &str) -> GidLookup { + self.by_name_char(name, super::encodings::glyph_name_char(name)) + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/simple.rs b/src-tauri/src/pdf_engine/text_edit/fonts/simple.rs new file mode 100644 index 0000000..886c467 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/simple.rs @@ -0,0 +1,761 @@ +//! Simple fonts (SPEC §A.2, §A.3, §B.9.3): embedded TrueType (`FontFile2`), CFF (`FontFile3 +//! /Type1C`), OpenType (`FontFile3 /OpenType`), Type1 (`FontFile`), Standard-14 and other +//! non-embedded fonts. One code = one byte; every code 0–255 gets a `CodeInfo`. +//! +//! Reading: ToUnicode wins, but a code whose ToUnicode text and glyph-name text differ (after +//! ligature expansion) is undecodable. A ToUnicode that is present but unusable (decode or +//! syntax error) leaves reading to the glyph names and makes no code typeable. Writing: §A.3.2 +//! (one scalar, writable character, glyph proven present in the program, width > 0). + +use super::cff_encoding::{cff_builtin_encoding, CffEncoding}; +use super::encodings::{ + apply_differences, display_text, duplicate_agrees, glyph_name_char, substitution_safe, + usable_text, writable_char, BaseEncoding, +}; +use super::glyph_budget::WorkMeter; +use super::program::{face_metrics, CffNames, GidLookup, OutlineMemo, Outlines, TrueTypeLookup}; +use super::std14::{std14_match, Std14Face, Std14Match}; +use super::tounicode::{Lookup, ToUnicode}; +use super::type1::{parse_type1, GlyphProof, Type1Encoding, Type1Program}; +use super::{ + is_whitespace_char, number_of, vertical_metrics, CodeInfo, Descriptor, FontClass, FontModel, + Loader, StreamFail, +}; +use crate::pdf_engine::text_edit::limits::{FONT_PROGRAM_MAX_DECODED, TOUNICODE_MAX_DECODED}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use lopdf::{Dictionary, Object, ObjectId, Stream}; +use std::collections::{HashMap, HashSet}; +use ttf_parser::{cff, Face, GlyphId}; + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum Kind { + Type1, + TrueType, + Cff, + OpenType, +} + +/// `/Widths` with `/FirstChar` (`LastChar` = first + len − 1). +struct Widths { + first: usize, + values: Vec, +} + +enum EncodingSpec { + Absent, + Named(BaseEncoding), + Dict { + base: Option, + differences: Vec, + }, +} + +/// What ToUnicode a font has. +pub(crate) enum TuState { + Absent, + /// Present but not usable (not a stream, undecodable, too large, syntax error). + Broken, + Ready(ToUnicode), +} + +/// Reads and parses `/ToUnicode` of `dict`. +pub(crate) fn load_tounicode<'a>(loader: &mut Loader<'a, '_>, dict: &'a Dictionary) -> TuState { + if loader.get(dict, b"ToUnicode").is_none() { + return if dict.has(b"ToUnicode") && !matches!(dict.get(b"ToUnicode"), Ok(Object::Null)) { + TuState::Broken + } else { + TuState::Absent + }; + } + let Some((id, stream)) = loader.stream(dict, b"ToUnicode") else { + return TuState::Broken; + }; + match loader.stream_bytes(id, stream, TOUNICODE_MAX_DECODED) { + Ok(bytes) => match super::tounicode::parse_tounicode(&bytes) { + Ok(tu) => TuState::Ready(tu), + Err(()) => TuState::Broken, + }, + Err(_) => TuState::Broken, + } +} + +pub(super) fn load<'a>( + loader: &mut Loader<'a, '_>, + mut model: FontModel, + dict: &'a Dictionary, + subtype: &[u8], +) -> FontModel { + let raw_base = + String::from_utf8_lossy(loader.name(dict, b"BaseFont").unwrap_or_default()).into_owned(); + let truetype = subtype == b"TrueType"; + let desc = match Descriptor::read(loader, dict) { + Ok(d) => d, + Err(reason) => { + model.set_names(&raw_base, &Default::default()); + model.refuse(reason); + return model; + } + }; + let desc_dict = loader.dict(dict, b"FontDescriptor"); + let program = select_program(loader, desc_dict, truetype); + let std14 = match program { + Ok(None) => std14_match(super::faces::strip_subset_tag(&raw_base).0), + _ => None, + }; + let mut hints = desc.hints(); + if let Some(Std14Match::Latin(face)) = std14 { + hints.italic_angle = hints.italic_angle.or(Some(face.italic_angle())); + } + model.set_names(&raw_base, &hints); + if let Some(Std14Match::Latin(face)) = std14 { + model.family_hint = face.family_hint(); + } + // The structural checks all run before any of them returns, so the model carries the + // lowest-numbered (§A.10) of the refusals that apply. + let program = program.map_err(|reason| { + model.embedded = true; + model.refuse(reason); + }); + let multiple_master = subtype == b"MMType1"; + if multiple_master { + model.refuse(TextReason::FontUnsupported); + } + let widths = read_widths(loader, dict).map_err(|reason| model.refuse(reason)); + let encoding = read_encoding(loader, dict).map_err(|reason| model.refuse(reason)); + let (Ok(program), false, Ok(widths), Ok(encoding)) = + (program, multiple_master, widths, encoding) + else { + return model; + }; + model.embedded = program.is_some(); + let tu = load_tounicode(loader, dict); + let mut parts = Parts { + desc: &desc, + widths: widths.as_ref(), + encoding, + tu, + std14: None, + }; + match program { + Some((kind, id, stream)) => load_embedded(loader, &mut model, &mut parts, kind, id, stream), + None => match std14 { + Some(Std14Match::Latin(face)) => { + parts.std14 = Some(face); + model.class = Some(FontClass::Std14(face)); + let (ascent, descent) = + vertical_metrics(&desc, None, Some((face.ascent(), face.descent()))); + model.ascent = ascent; + model.descent = descent; + fill_codes( + &mut model, + &parts, + &Oracle::Std14(face), + &mut WorkMeter::new(0), + ); + } + Some(Std14Match::Symbolic) => model.refuse(TextReason::UnsupportedEncoding), + None => { + if desc.symbolic() { + model.refuse(TextReason::UnsupportedEncoding); + } else if parts.widths.is_none() { + model.refuse(TextReason::MissingWidths); + } + model.class = Some(FontClass::SimpleNonEmbedded); + model.substituted = true; + let (ascent, descent) = vertical_metrics(&desc, None, None); + model.ascent = ascent; + model.descent = descent; + fill_codes( + &mut model, + &parts, + &Oracle::NonEmbedded, + &mut WorkMeter::new(0), + ); + } + }, + } + model +} + +/// Type3 (refused `TYPE3`): codes still read and advance as viewers draw them, so the pen +/// position of everything after a Type3 show op stays right. `/Widths` are in glyph space and +/// scale by `/FontMatrix` (default 0.001); nothing is drawable or typeable. +pub(super) fn load_type3<'a>( + loader: &mut Loader<'a, '_>, + mut model: FontModel, + dict: &'a Dictionary, +) -> FontModel { + let raw_base = + String::from_utf8_lossy(loader.name(dict, b"BaseFont").unwrap_or_default()).into_owned(); + model.set_names(&raw_base, &Default::default()); + model.refuse(TextReason::Type3); + let scale = match loader.get(dict, b"FontMatrix") { + Some(Object::Array(m)) => m + .first() + .and_then(|a| loader.resolve(a)) + .and_then(|(_, a)| number_of(a)) + .map(|a| a * 1000.0) + .filter(|a| a.is_finite() && *a != 0.0), + _ => None, + } + .unwrap_or(1.0); + let widths = read_widths(loader, dict).ok().flatten().map(|w| Widths { + first: w.first, + values: w.values.iter().map(|v| v * scale).collect(), + }); + let desc = Descriptor::default(); + let parts = Parts { + desc: &desc, + widths: widths.as_ref(), + encoding: read_encoding(loader, dict).unwrap_or(EncodingSpec::Absent), + tu: load_tounicode(loader, dict), + std14: None, + }; + fill_codes(&mut model, &parts, &Oracle::Broken, &mut WorkMeter::new(0)); + model +} + +struct Parts<'d> { + desc: &'d Descriptor, + widths: Option<&'d Widths>, + encoding: EncodingSpec, + tu: TuState, + /// The AFM face lending widths by glyph name (Standard-14, or an embedded font without + /// `/Widths` whose BaseFont is a Standard-14 name). + std14: Option, +} + +/// Per-code glyph presence for one font class. +enum Oracle<'p> { + TrueType { + lookup: TrueTypeLookup<'p, 'p>, + use_names: bool, + outlines: Outlines<'p>, + }, + /// A bare CFF and, when the PDF gives no base encoding, the program's own one. + Cff { + outlines: Outlines<'p>, + table: cff::Table<'p>, + names: CffNames, + builtin: Option, + }, + Type1 { + program: Type1Program, + charset: Option<&'p HashSet>, + }, + Std14(Std14Face), + NonEmbedded, + /// The program could not be read: nothing is drawable. + Broken, +} + +fn load_embedded<'a>( + loader: &mut Loader<'a, '_>, + model: &mut FontModel, + parts: &mut Parts<'_>, + kind: Kind, + id: ObjectId, + stream: &'a Stream, +) { + model.class = Some(match kind { + Kind::Type1 => FontClass::SimpleType1, + Kind::TrueType => FontClass::SimpleTrueType, + Kind::Cff => FontClass::SimpleCff, + Kind::OpenType => FontClass::SimpleOpenType, + }); + // /Widths rule: an embedded font without them borrows AFM widths only under a Standard-14 + // name; anything else is MISSING_WIDTHS. + let afm = match std14_match(&model.base_name) { + Some(Std14Match::Latin(face)) => Some(face), + _ => None, + }; + if parts.widths.is_none() { + match afm { + Some(face) => parts.std14 = Some(face), + None => model.refuse(TextReason::MissingWidths), + } + } + let fallback = afm.map(|f| (f.ascent(), f.descent())); + let (ascent, descent) = vertical_metrics(parts.desc, None, fallback); + model.ascent = ascent; + model.descent = descent; + let bytes = match loader.stream_bytes(id, stream, FONT_PROGRAM_MAX_DECODED) { + Ok(b) => b, + Err(StreamFail::Budget | StreamFail::Bad) => { + model.refuse(TextReason::FontProgramUnreadable); + fill_codes(model, parts, &Oracle::Broken, &mut WorkMeter::new(0)); + return; + } + }; + let face = match kind { + Kind::TrueType | Kind::OpenType => Face::parse(&bytes, 0).ok(), + Kind::Cff | Kind::Type1 => None, + }; + let needs_builtin = match &parts.encoding { + EncodingSpec::Absent => true, + EncodingSpec::Dict { base, .. } => base.is_none(), + EncodingSpec::Named(_) => false, + }; + let oracle = match kind { + Kind::TrueType | Kind::OpenType => match face.as_ref() { + Some(face) => { + let (ascent, descent) = vertical_metrics(parts.desc, face_metrics(face), fallback); + model.ascent = ascent; + model.descent = descent; + // PDF 32000 §9.6.6.4 / ISO 32000-2 §9.6.5.4: the Symbolic flag selects glyphs by + // code through the (3,0)/(1,0) cmaps (an /Encoding then only names the codes). + let use_names = !parts.desc.symbolic(); + let outlines = Outlines::of_face(face); + Oracle::TrueType { + lookup: TrueTypeLookup::new(face, &outlines), + use_names, + outlines, + } + } + None => Oracle::Broken, + }, + Kind::Cff => { + let builtin = if needs_builtin { + cff_builtin_encoding(&bytes).map(Some) + } else { + Ok(None) + }; + match (Outlines::of_cff(&bytes), builtin) { + (Some(outlines), Ok(builtin)) => { + let found = match &outlines { + Outlines::Cff { guard, table } => { + Some((CffNames::new(guard.layout(), table), *table)) + } + _ => None, + }; + match found { + Some((names, table)) => Oracle::Cff { + outlines, + table, + names, + builtin, + }, + None => Oracle::Broken, + } + } + _ => Oracle::Broken, + } + } + Kind::Type1 => { + let length = |key: &[u8]| { + loader + .get(&stream.dict, key) + .and_then(number_of) + .filter(|v| { + *v >= 0.0 && v.fract() == 0.0 && *v <= FONT_PROGRAM_MAX_DECODED as f64 + }) + .map(|v| v as usize) + }; + match (length(b"Length1"), length(b"Length2")) { + (Some(l1), Some(l2)) => match parse_type1(&bytes, l1, l2) { + Ok(program) => Oracle::Type1 { + program, + charset: parts.desc.charset.as_ref(), + }, + Err(()) => Oracle::Broken, + }, + _ => Oracle::Broken, + } + } + }; + if let Oracle::TrueType { outlines, .. } | Oracle::Cff { outlines, .. } = &oracle { + if let Some(reason) = outlines.refusal() { + model.refuse(reason); + } + } + match &oracle { + Oracle::Broken => model.refuse(TextReason::FontProgramUnreadable), + Oracle::TrueType { + use_names: false, .. + } if matches!(parts.encoding, EncodingSpec::Absent) => match parts.tu { + // A symbolic TrueType font without /Encoding has no glyph names: its text can only + // come from ToUnicode. + TuState::Absent => model.refuse(TextReason::NoTounicode), + TuState::Broken => model.refuse(TextReason::AmbiguousUnicode), + TuState::Ready(_) => {} + }, + _ => {} + } + let mut meter = loader.work_meter(); + fill_codes(model, parts, &oracle, &mut meter); + loader.settle(&meter); +} + +/// `(kind, object id, stream)` of the descriptor's program; a program that does not match the +/// font type is `FONT_PROGRAM_UNSUPPORTED`, one that is not a stream `FONT_PROGRAM_UNREADABLE`. +#[allow(clippy::type_complexity)] +fn select_program<'a>( + loader: &Loader<'a, '_>, + desc: Option<&'a Dictionary>, + truetype: bool, +) -> Result, TextReason> { + let Some(d) = desc else { + return Ok(None); + }; + let mut found = Vec::new(); + for key in [&b"FontFile"[..], b"FontFile2", b"FontFile3"] { + let present = d.get(key).is_ok_and(|v| !matches!(v, Object::Null)); + if !present { + continue; + } + let (id, stream) = loader + .stream(d, key) + .ok_or(TextReason::FontProgramUnreadable)?; + found.push((key, id, stream)); + } + let [(key, id, stream)] = found.as_slice() else { + return if found.is_empty() { + Ok(None) + } else { + Err(TextReason::FontProgramUnsupported) + }; + }; + let kind = match (*key, truetype) { + (b"FontFile", false) => Kind::Type1, + (b"FontFile2", true) => Kind::TrueType, + (b"FontFile3", _) => match (loader.name(&stream.dict, b"Subtype"), truetype) { + (Some(b"Type1C"), false) => Kind::Cff, + (Some(b"OpenType"), _) => Kind::OpenType, + _ => return Err(TextReason::FontProgramUnsupported), + }, + _ => return Err(TextReason::FontProgramUnsupported), + }; + Ok(Some((kind, *id, *stream))) +} + +/// `/Widths` + `/FirstChar` + `/LastChar`: integers 0–255, `LastChar ≥ FirstChar`, exactly +/// `LastChar − FirstChar + 1` finite numbers; anything else is `FONT_UNSUPPORTED`. +fn read_widths<'a>( + loader: &Loader<'a, '_>, + dict: &'a Dictionary, +) -> Result, TextReason> { + let present = dict + .get(b"Widths") + .is_ok_and(|w| !matches!(w, Object::Null)); + if !present { + return Ok(None); + } + let Some(Object::Array(items)) = loader.get(dict, b"Widths") else { + return Err(TextReason::FontUnsupported); + }; + let int = |key: &[u8]| match loader.get(dict, key) { + Some(Object::Integer(v)) if (0..=255).contains(v) => usize::try_from(*v).ok(), + _ => None, + }; + let (Some(first), Some(last)) = (int(b"FirstChar"), int(b"LastChar")) else { + return Err(TextReason::FontUnsupported); + }; + if last < first || items.len() != last - first + 1 { + return Err(TextReason::FontUnsupported); + } + let values = items + .iter() + .map(|item| loader.resolve(item).and_then(|(_, o)| number_of(o))) + .collect::>>() + .ok_or(TextReason::FontUnsupported)?; + Ok(Some(Widths { first, values })) +} + +/// `/Encoding`: a base-encoding name, or a dictionary with an optional `/BaseEncoding` and a +/// `/Differences` array that must be well-formed (checked here, applied in `name_table`). +fn read_encoding<'a>( + loader: &Loader<'a, '_>, + dict: &'a Dictionary, +) -> Result { + let present = dict + .get(b"Encoding") + .is_ok_and(|e| !matches!(e, Object::Null)); + if !present { + return Ok(EncodingSpec::Absent); + } + match loader.get(dict, b"Encoding") { + Some(Object::Name(name)) => Ok(EncodingSpec::Named(BaseEncoding::from_name(name)?)), + Some(Object::Dictionary(enc)) => { + let base = match loader.get(enc, b"BaseEncoding") { + None => None, + Some(Object::Name(name)) => Some(BaseEncoding::from_name(name)?), + Some(_) => return Err(TextReason::FontUnsupported), + }; + let differences = match loader.get(enc, b"Differences") { + None => Vec::new(), + Some(Object::Array(items)) => items.clone(), + Some(_) => return Err(TextReason::FontUnsupported), + }; + apply_differences(&mut vec![None; 256], &differences)?; + Ok(EncodingSpec::Dict { base, differences }) + } + _ => Err(TextReason::FontUnsupported), + } +} + +/// Glyph names per code plus, for a CFF built-in custom encoding, the GID each code maps to. +fn name_table(parts: &Parts<'_>, oracle: &Oracle<'_>) -> (Vec>, Vec>) { + let mut names: Vec> = vec![None; 256]; + let mut direct: Vec> = vec![None; 256]; + let builtin = |names: &mut Vec>, direct: &mut Vec>| match oracle { + Oracle::Type1 { program, .. } => match &program.builtin { + Some(Type1Encoding::Standard) => *names = BaseEncoding::Standard.names(), + Some(Type1Encoding::Custom(map)) => { + for (code, name) in map { + if let Some(slot) = names.get_mut(usize::from(*code)) { + *slot = Some(name.clone()); + } + } + } + None => {} + }, + Oracle::Cff { + table, + builtin, + outlines, + .. + } => match builtin { + Some(CffEncoding::Standard) => *names = BaseEncoding::Standard.names(), + Some(CffEncoding::Custom(map)) => { + // GID → SID from one walk of the charset (ttf-parser's `glyph_name` walks a + // format 1/2 charset per call); the predefined Expert charsets have no walk. + let sids = outlines + .cff_layout() + .filter(|layout| !layout.is_cid()) + .and_then(|layout| Some((layout, layout.gid_to_sid()?))); + for (code, gid) in map { + if *gid >= table.number_of_glyphs() { + continue; + } + let name = match &sids { + Some((layout, sids)) => sids + .get(usize::from(*gid)) + .copied() + .flatten() + .and_then(|sid| layout.sid_name(sid)), + None => table.glyph_name(GlyphId(*gid)), + }; + let i = usize::from(*code); + if let (Some(n), Some(d)) = (names.get_mut(i), direct.get_mut(i)) { + *n = name.map(str::to_string); + *d = Some(*gid); + } + } + } + // Expert: no code is typeable (readable through ToUnicode only). + Some(CffEncoding::Expert) | None => {} + }, + // Symbolic TrueType without /Encoding: codes go through the cmap, not names. + Oracle::TrueType { + use_names: false, .. + } + | Oracle::Broken => {} + Oracle::TrueType { .. } | Oracle::Std14(_) | Oracle::NonEmbedded => { + *names = BaseEncoding::Standard.names(); + } + }; + match &parts.encoding { + EncodingSpec::Absent => builtin(&mut names, &mut direct), + EncodingSpec::Named(base) => names = base.names(), + EncodingSpec::Dict { base, differences } => { + match base { + Some(base) => names = base.names(), + None => builtin(&mut names, &mut direct), + } + let before = names.clone(); + if apply_differences(&mut names, differences).is_err() { + names = before.clone(); // validated in read_encoding; never taken + } + for ((old, new), slot) in before.iter().zip(&names).zip(direct.iter_mut()) { + if old != new { + *slot = None; // a /Differences name replaces the built-in GID + } + } + } + } + (names, direct) +} + +/// Presence results of one font load: outlines per GID, Type1 proofs per glyph name. +#[derive(Default)] +struct Memo { + outlines: OutlineMemo, + proofs: HashMap, +} + +/// Builds the `CodeInfo` of every code 0–255; glyph work is charged to `meter`. +fn fill_codes( + model: &mut FontModel, + parts: &Parts<'_>, + oracle: &Oracle<'_>, + meter: &mut WorkMeter, +) { + let (names, direct) = name_table(parts, oracle); + let tu = match &parts.tu { + TuState::Ready(t) => Some(t), + TuState::Absent | TuState::Broken => None, + }; + let tu_broken = matches!(parts.tu, TuState::Broken); + let substituted = matches!(oracle, Oracle::NonEmbedded); + let mut memo = Memo::default(); + let mut infos = Vec::with_capacity(256); + for code in 0u8..=255 { + let i = usize::from(code); + let name = names.get(i).cloned().flatten(); + let width = width_of(parts, code, name.as_deref()); + let mapped = tu.map(|t| t.lookup(u32::from(code))); + if name.is_none() && width.is_none() && matches!(mapped, None | Some(Lookup::Absent)) { + continue; // nothing to read, measure or type: `info` is None, `width` 0 + } + let name_char = name.as_deref().and_then(glyph_name_char); + let name_text = name_char.map(String::from); + let raw = match mapped { + Some(Lookup::Undecodable) => None, + Some(Lookup::Text(t)) => match (&name_text, name.as_deref()) { + (Some(_), Some(n)) if duplicate_agrees(code, n, t) => Some(t.to_string()), + (Some(n), _) if display_text(t) != display_text(n) => None, + _ => Some(t.to_string()), + }, + Some(Lookup::Absent) | None => name_text, + }; + let text = raw.as_deref().filter(|t| usable_text(t)).map(display_text); + let ws = is_whitespace_char(text.as_deref()); + let d = direct.get(i).copied().flatten(); + let glyph = GlyphRef { + code, + name: name.as_deref(), + name_char, + direct: d, + }; + let (gid, drawable) = presence(oracle, &mut memo, meter, glyph, ws, width); + let single = raw.as_deref().and_then(|t| { + let mut chars = t.chars(); + match (chars.next(), chars.next()) { + (Some(c), None) => Some(c), + _ => None, + } + }); + let typeable_as = single.filter(|ch| { + drawable + && width.is_some_and(|w| w > 0.0) + && !tu_broken + && text.is_some() + && writable_char(*ch) + && (!substituted || substitution_safe(*ch)) + }); + infos.push(( + u32::from(code), + CodeInfo { + text, + glyph_name: name, + gid, + width1000: width, + drawable, + typeable_as, + }, + )); + } + // Codes ascend: the map is built in bulk, not by 256 single inserts. + model.codes = infos.into_iter().collect(); +} + +/// `/Widths[code − FirstChar]` inside the range, `/MissingWidth` outside it only when present; +/// without `/Widths`, the AFM width of the glyph name (Standard-14). +fn width_of(parts: &Parts<'_>, code: u8, name: Option<&str>) -> Option { + match parts.widths { + Some(w) => usize::from(code) + .checked_sub(w.first) + .and_then(|i| w.values.get(i).copied()) + .or(parts.desc.missing_width), + None => parts.std14.zip(name).and_then(|(face, n)| face.width(n)), + } +} + +/// One code and what its encoding says: glyph name (and its AGL character), or a GID straight +/// from a CFF built-in encoding. +#[derive(Clone, Copy)] +struct GlyphRef<'n> { + code: u8, + name: Option<&'n str>, + name_char: Option, + direct: Option, +} + +/// `(gid, drawable)` of one code under the font class's presence rule (§A.2). +fn presence( + oracle: &Oracle<'_>, + memo: &mut Memo, + meter: &mut WorkMeter, + glyph: GlyphRef<'_>, + whitespace: bool, + width: Option, +) -> (Option, bool) { + let GlyphRef { + code, + name, + name_char, + direct, + } = glyph; + let blank_ok = whitespace && width.is_some_and(|w| w > 0.0); + match oracle { + Oracle::TrueType { + lookup, + use_names, + outlines, + } => { + let found = if *use_names { + name.map_or(GidLookup::None, |n| lookup.by_name_char(n, name_char)) + } else { + lookup.symbolic(code) + }; + match found { + GidLookup::Agree(g) if g != 0 && g < lookup.face.number_of_glyphs() => { + let drawn = memo.outlines.get(g, || outlines.drawn(g, meter)); + (Some(g), drawn || blank_ok) + } + _ => (None, false), + } + } + Oracle::Cff { + outlines, + table, + names, + .. + } => { + let gid = + direct.or_else(|| name.filter(|n| *n != ".notdef").and_then(|n| names.gid(n))); + match gid { + Some(g) if g != 0 && g < table.number_of_glyphs() => { + let drawn = memo.outlines.get(g, || outlines.drawn(g, meter)); + (Some(g), drawn || blank_ok) + } + _ => (None, false), + } + } + Oracle::Type1 { program, charset } => { + let Some(n) = name else { + return (None, false); + }; + if charset.is_some_and(|set| !set.contains(n)) { + return (None, false); + } + let proof = match memo.proofs.get(n) { + Some(proof) => *proof, + None => { + let proof = program.proof(n, meter); + memo.proofs.insert(n.to_string(), proof); + proof + } + }; + let drawn = match proof { + GlyphProof::Drawn => true, + GlyphProof::Blank { width: w } => blank_ok && w > 0.0, + GlyphProof::Unproven => false, + }; + (None, drawn) + } + Oracle::Std14(face) => (None, name.is_some_and(|n| face.has_glyph(n))), + Oracle::NonEmbedded => (None, width.is_some_and(|w| w > 0.0)), + Oracle::Broken => (None, false), + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/std14.rs b/src-tauri/src/pdf_engine/text_edit/fonts/std14.rs new file mode 100644 index 0000000..322e51c --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/std14.rs @@ -0,0 +1,182 @@ +//! Standard-14 fonts (SPEC §A.2 row "Standard-14 not embedded", §B.9.2): the 12 Latin faces, +//! their aliases, and per-face AFM widths, glyph sets and vertical metrics. `Symbol` and +//! `ZapfDingbats` are recognised only to be refused (`UNSUPPORTED_ENCODING`). + +use super::std14_data::{FaceMetrics, FACES, GLYPHS}; +use super::FamilyHint; + +/// One of the 12 Latin Core-14 faces (index order = `std14_data::FACES`). +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] +pub enum Std14Face { + Helvetica, + HelveticaBold, + HelveticaOblique, + HelveticaBoldOblique, + TimesRoman, + TimesBold, + TimesItalic, + TimesBoldItalic, + Courier, + CourierBold, + CourierOblique, + CourierBoldOblique, +} + +/// What a non-embedded font's BaseFont resolves to. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum Std14Match { + Latin(Std14Face), + /// `Symbol` or `ZapfDingbats` (and their style aliases): refused `UNSUPPORTED_ENCODING`. + Symbolic, +} + +use Std14Face as S; + +/// Canonical names and aliases, after `normalize` (`,` and `_` → `-`, whitespace removed, as +/// pdf.js `normalizeFontName`). The aliases are SPEC §A.2's list plus pdf.js's spellings of the +/// same three families (`getStdFontMap` in pdf.js 4.10.38: `TimesNewRomanPSMT`, `CourierNewPSMT`, +/// `Helvetica-Italic`, `Courier-Italic`, …). Narrow/Black/Unicode Arial variants are excluded: +/// their metrics are not Helvetica's. +const ALIASES: &[(&str, Std14Face)] = &[ + ("Helvetica", S::Helvetica), + ("Helvetica-Bold", S::HelveticaBold), + ("Helvetica-Oblique", S::HelveticaOblique), + ("Helvetica-BoldOblique", S::HelveticaBoldOblique), + ("Helvetica-Italic", S::HelveticaOblique), + ("Helvetica-BoldItalic", S::HelveticaBoldOblique), + ("Arial", S::Helvetica), + ("ArialMT", S::Helvetica), + ("Arial-Bold", S::HelveticaBold), + ("Arial-BoldMT", S::HelveticaBold), + ("Arial-Italic", S::HelveticaOblique), + ("Arial-ItalicMT", S::HelveticaOblique), + ("Arial-BoldItalic", S::HelveticaBoldOblique), + ("Arial-BoldItalicMT", S::HelveticaBoldOblique), + ("Times-Roman", S::TimesRoman), + ("Times-Bold", S::TimesBold), + ("Times-Italic", S::TimesItalic), + ("Times-BoldItalic", S::TimesBoldItalic), + ("TimesNewRoman", S::TimesRoman), + ("TimesNewRoman-Bold", S::TimesBold), + ("TimesNewRoman-Italic", S::TimesItalic), + ("TimesNewRoman-BoldItalic", S::TimesBoldItalic), + ("TimesNewRomanPS", S::TimesRoman), + ("TimesNewRomanPSMT", S::TimesRoman), + ("TimesNewRomanPS-Bold", S::TimesBold), + ("TimesNewRomanPS-BoldMT", S::TimesBold), + ("TimesNewRomanPS-Italic", S::TimesItalic), + ("TimesNewRomanPS-ItalicMT", S::TimesItalic), + ("TimesNewRomanPS-BoldItalic", S::TimesBoldItalic), + ("TimesNewRomanPS-BoldItalicMT", S::TimesBoldItalic), + ("Courier", S::Courier), + ("Courier-Bold", S::CourierBold), + ("Courier-Oblique", S::CourierOblique), + ("Courier-BoldOblique", S::CourierBoldOblique), + ("Courier-Italic", S::CourierOblique), + ("Courier-BoldItalic", S::CourierBoldOblique), + ("CourierNew", S::Courier), + ("CourierNew-Bold", S::CourierBold), + ("CourierNew-Italic", S::CourierOblique), + ("CourierNew-BoldItalic", S::CourierBoldOblique), + ("CourierNewPS", S::Courier), + ("CourierNewPSMT", S::Courier), + ("CourierNewPS-BoldMT", S::CourierBold), + ("CourierNewPS-ItalicMT", S::CourierOblique), + ("CourierNewPS-BoldItalicMT", S::CourierBoldOblique), +]; + +const SYMBOLIC: &[&str] = &[ + "Symbol", + "Symbol-Bold", + "Symbol-Italic", + "Symbol-BoldItalic", + "ZapfDingbats", +]; + +fn normalize(name: &str) -> String { + name.chars() + .filter(|c| !c.is_whitespace()) + .map(|c| if c == ',' || c == '_' { '-' } else { c }) + .collect() +} + +/// The Standard-14 face a (subset-tag-free) BaseFont names, if any. +pub fn std14_match(base_name: &str) -> Option { + let name = normalize(base_name); + if SYMBOLIC.contains(&name.as_str()) { + return Some(Std14Match::Symbolic); + } + ALIASES + .iter() + .find(|(alias, _)| *alias == name) + .map(|(_, face)| Std14Match::Latin(*face)) +} + +impl Std14Face { + pub const ALL: [Std14Face; 12] = [ + S::Helvetica, + S::HelveticaBold, + S::HelveticaOblique, + S::HelveticaBoldOblique, + S::TimesRoman, + S::TimesBold, + S::TimesItalic, + S::TimesBoldItalic, + S::Courier, + S::CourierBold, + S::CourierOblique, + S::CourierBoldOblique, + ]; + + fn metrics(self) -> Option<&'static FaceMetrics> { + let index = Self::ALL.iter().position(|f| *f == self)?; + FACES.get(index) + } + + /// AFM advance width (glyph space ×1000) of `glyph`; `None` if the face lacks the glyph. + pub fn width(self, glyph: &str) -> Option { + let index = GLYPHS + .binary_search_by(|g| g.as_bytes().cmp(glyph.as_bytes())) + .ok()?; + self.metrics() + .and_then(|m| m.widths.get(index)) + .map(|w| f64::from(*w)) + } + + /// Whether the face's AFM glyph set contains `glyph` (`.notdef` never does). + pub fn has_glyph(self, glyph: &str) -> bool { + self.width(glyph).is_some() + } + + /// AFM `Ascender` / 1000 (em). + pub fn ascent(self) -> f64 { + self.metrics() + .map(|m| f64::from(m.ascender) / 1000.0) + .unwrap_or(0.8) + } + + /// AFM `Descender` / 1000 (em, ≤ 0). + pub fn descent(self) -> f64 { + self.metrics() + .map(|m| f64::from(m.descender) / 1000.0) + .unwrap_or(-0.2) + } + + /// AFM `ItalicAngle` (degrees). + pub fn italic_angle(self) -> f64 { + self.metrics().map(|m| m.italic_angle).unwrap_or_default() + } + + /// Courier → Mono, Times → Serif, Helvetica → Sans (§B.9.3). + pub fn family_hint(self) -> FamilyHint { + match self { + S::Courier | S::CourierBold | S::CourierOblique | S::CourierBoldOblique => { + FamilyHint::Mono + } + S::TimesRoman | S::TimesBold | S::TimesItalic | S::TimesBoldItalic => FamilyHint::Serif, + S::Helvetica | S::HelveticaBold | S::HelveticaOblique | S::HelveticaBoldOblique => { + FamilyHint::Sans + } + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/std14_data.rs b/src-tauri/src/pdf_engine/text_edit/fonts/std14_data.rs new file mode 100644 index 0000000..41d4890 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/std14_data.rs @@ -0,0 +1,343 @@ +//! GENERATED DATA — do not edit by hand (SPEC §B.9.2). Core-14 AFM metrics of the 12 Latin +//! standard faces: the shared 315-name glyph set and, per face, each glyph's advance width +//! (`WX`, glyph space ×1000), `Ascender`, `Descender` and `ItalicAngle`. +//! +//! Provenance: `@pdf-lib/standard-fonts` 1.0.0 (`es/.compressed.json`, base64 of zlib'd +//! JSON, MIT, Copyright (c) 2018 Andrew Dillon), the JSON form of Adobe's Core14 AFM files +//! ("Copyright (c) 1985, 1987, 1989, 1990, 1993, 1997 Adobe Systems Incorporated. All Rights +//! Reserved." — Adobe permits use, copying and distribution of the Core14 AFM data for any purpose +//! provided the copyright notices are retained; Helvetica and Times are trademarks of +//! Linotype-Hell AG and/or its subsidiaries). Every width, ascender and descender was checked to +//! be identical to pdf.js 4.10.38 (`getMetrics`, `getFontBasicMetrics` in `build/pdf.worker.mjs`, +//! Apache-2.0). All 12 faces carry the same 315 glyph names (asserted when generating). + +/// The glyph names of every Latin Core-14 face, sorted by bytes (binary search). +#[rustfmt::skip] +pub(super) static GLYPHS: [&str; 315] = [ + "A", "AE", "Aacute", "Abreve", "Acircumflex", "Adieresis", "Agrave", "Amacron", "Aogonek", + "Aring", "Atilde", "B", "C", "Cacute", "Ccaron", "Ccedilla", "D", "Dcaron", "Dcroat", "Delta", + "E", "Eacute", "Ecaron", "Ecircumflex", "Edieresis", "Edotaccent", "Egrave", "Emacron", + "Eogonek", "Eth", "Euro", "F", "G", "Gbreve", "Gcommaaccent", "H", "I", "Iacute", "Icircumflex", + "Idieresis", "Idotaccent", "Igrave", "Imacron", "Iogonek", "J", "K", "Kcommaaccent", "L", + "Lacute", "Lcaron", "Lcommaaccent", "Lslash", "M", "N", "Nacute", "Ncaron", "Ncommaaccent", + "Ntilde", "O", "OE", "Oacute", "Ocircumflex", "Odieresis", "Ograve", "Ohungarumlaut", "Omacron", + "Oslash", "Otilde", "P", "Q", "R", "Racute", "Rcaron", "Rcommaaccent", "S", "Sacute", "Scaron", + "Scedilla", "Scommaaccent", "T", "Tcaron", "Tcommaaccent", "Thorn", "U", "Uacute", + "Ucircumflex", "Udieresis", "Ugrave", "Uhungarumlaut", "Umacron", "Uogonek", "Uring", "V", "W", + "X", "Y", "Yacute", "Ydieresis", "Z", "Zacute", "Zcaron", "Zdotaccent", "a", "aacute", "abreve", + "acircumflex", "acute", "adieresis", "ae", "agrave", "amacron", "ampersand", "aogonek", "aring", + "asciicircum", "asciitilde", "asterisk", "at", "atilde", "b", "backslash", "bar", "braceleft", + "braceright", "bracketleft", "bracketright", "breve", "brokenbar", "bullet", "c", "cacute", + "caron", "ccaron", "ccedilla", "cedilla", "cent", "circumflex", "colon", "comma", "commaaccent", + "copyright", "currency", "d", "dagger", "daggerdbl", "dcaron", "dcroat", "degree", "dieresis", + "divide", "dollar", "dotaccent", "dotlessi", "e", "eacute", "ecaron", "ecircumflex", + "edieresis", "edotaccent", "egrave", "eight", "ellipsis", "emacron", "emdash", "endash", + "eogonek", "equal", "eth", "exclam", "exclamdown", "f", "fi", "five", "fl", "florin", "four", + "fraction", "g", "gbreve", "gcommaaccent", "germandbls", "grave", "greater", "greaterequal", + "guillemotleft", "guillemotright", "guilsinglleft", "guilsinglright", "h", "hungarumlaut", + "hyphen", "i", "iacute", "icircumflex", "idieresis", "igrave", "imacron", "iogonek", "j", "k", + "kcommaaccent", "l", "lacute", "lcaron", "lcommaaccent", "less", "lessequal", "logicalnot", + "lozenge", "lslash", "m", "macron", "minus", "mu", "multiply", "n", "nacute", "ncaron", + "ncommaaccent", "nine", "notequal", "ntilde", "numbersign", "o", "oacute", "ocircumflex", + "odieresis", "oe", "ogonek", "ograve", "ohungarumlaut", "omacron", "one", "onehalf", + "onequarter", "onesuperior", "ordfeminine", "ordmasculine", "oslash", "otilde", "p", + "paragraph", "parenleft", "parenright", "partialdiff", "percent", "period", "periodcentered", + "perthousand", "plus", "plusminus", "q", "question", "questiondown", "quotedbl", "quotedblbase", + "quotedblleft", "quotedblright", "quoteleft", "quoteright", "quotesinglbase", "quotesingle", + "r", "racute", "radical", "rcaron", "rcommaaccent", "registered", "ring", "s", "sacute", + "scaron", "scedilla", "scommaaccent", "section", "semicolon", "seven", "six", "slash", "space", + "sterling", "summation", "t", "tcaron", "tcommaaccent", "thorn", "three", "threequarters", + "threesuperior", "tilde", "trademark", "two", "twosuperior", "u", "uacute", "ucircumflex", + "udieresis", "ugrave", "uhungarumlaut", "umacron", "underscore", "uogonek", "uring", "v", "w", + "x", "y", "yacute", "ydieresis", "yen", "z", "zacute", "zcaron", "zdotaccent", "zero", +]; + +/// AFM metrics of one face; `widths[i]` is the advance of `GLYPHS[i]`. +pub(super) struct FaceMetrics { + pub widths: [u16; 315], + pub ascender: i16, + pub descender: i16, + pub italic_angle: f64, +} + +/// In `Std14Face` order: Helvetica, -Bold, -Oblique, -BoldOblique, Times-Roman, -Bold, -Italic, +/// -BoldItalic, Courier, -Bold, -Oblique, -BoldOblique. +#[rustfmt::skip] +pub(super) static FACES: [FaceMetrics; 12] = [ + FaceMetrics { + // Helvetica + widths: [ + 667, 1000, 667, 667, 667, 667, 667, 667, 667, 667, 667, 667, 722, 722, 722, 722, 722, 722, 722, 612, 667, + 667, 667, 667, 667, 667, 667, 667, 667, 722, 556, 611, 778, 778, 778, 722, 278, 278, 278, 278, 278, 278, + 278, 278, 500, 667, 667, 556, 556, 556, 556, 556, 833, 722, 722, 722, 722, 722, 778, 1000, 778, 778, 778, + 778, 778, 778, 778, 778, 667, 778, 722, 722, 722, 722, 667, 667, 667, 667, 667, 611, 611, 611, 667, 722, + 722, 722, 722, 722, 722, 722, 722, 722, 667, 944, 667, 667, 667, 667, 611, 611, 611, 611, 556, 556, 556, + 556, 333, 556, 889, 556, 556, 667, 556, 556, 469, 584, 389, 1015, 556, 556, 278, 260, 334, 334, 278, 278, + 333, 260, 350, 500, 500, 333, 500, 500, 333, 556, 333, 278, 278, 250, 737, 556, 556, 556, 556, 643, 556, + 400, 333, 584, 556, 333, 278, 556, 556, 556, 556, 556, 556, 556, 556, 1000, 556, 1000, 556, 556, 584, 556, + 278, 333, 278, 500, 556, 500, 556, 556, 167, 556, 556, 556, 611, 333, 584, 549, 556, 556, 333, 333, 556, + 333, 333, 222, 278, 278, 278, 278, 278, 222, 222, 500, 500, 222, 222, 299, 222, 584, 549, 584, 471, 222, + 833, 333, 584, 556, 584, 556, 556, 556, 556, 556, 549, 556, 556, 556, 556, 556, 556, 944, 333, 556, 556, + 556, 556, 834, 834, 333, 370, 365, 611, 556, 556, 537, 333, 333, 476, 889, 278, 278, 1000, 584, 584, 556, + 556, 611, 355, 333, 333, 333, 222, 222, 222, 191, 333, 333, 453, 333, 333, 737, 333, 500, 500, 500, 500, + 500, 556, 278, 556, 556, 278, 278, 556, 600, 278, 317, 278, 556, 556, 834, 333, 333, 1000, 556, 333, 556, + 556, 556, 556, 556, 556, 556, 556, 556, 556, 500, 722, 500, 500, 500, 500, 556, 500, 500, 500, 500, 556, + ], + ascender: 718, + descender: -207, + italic_angle: 0.0, + }, + FaceMetrics { + // Helvetica-Bold + widths: [ + 722, 1000, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 612, 667, + 667, 667, 667, 667, 667, 667, 667, 667, 722, 556, 611, 778, 778, 778, 722, 278, 278, 278, 278, 278, 278, + 278, 278, 556, 722, 722, 611, 611, 611, 611, 611, 833, 722, 722, 722, 722, 722, 778, 1000, 778, 778, 778, + 778, 778, 778, 778, 778, 667, 778, 722, 722, 722, 722, 667, 667, 667, 667, 667, 611, 611, 611, 667, 722, + 722, 722, 722, 722, 722, 722, 722, 722, 667, 944, 667, 667, 667, 667, 611, 611, 611, 611, 556, 556, 556, + 556, 333, 556, 889, 556, 556, 722, 556, 556, 584, 584, 389, 975, 556, 611, 278, 280, 389, 389, 333, 333, + 333, 280, 350, 556, 556, 333, 556, 556, 333, 556, 333, 333, 278, 250, 737, 556, 611, 556, 556, 743, 611, + 400, 333, 584, 556, 333, 278, 556, 556, 556, 556, 556, 556, 556, 556, 1000, 556, 1000, 556, 556, 584, 611, + 333, 333, 333, 611, 556, 611, 556, 556, 167, 611, 611, 611, 611, 333, 584, 549, 556, 556, 333, 333, 611, + 333, 333, 278, 278, 278, 278, 278, 278, 278, 278, 556, 556, 278, 278, 400, 278, 584, 549, 584, 494, 278, + 889, 333, 584, 611, 584, 611, 611, 611, 611, 556, 549, 611, 556, 611, 611, 611, 611, 944, 333, 611, 611, + 611, 556, 834, 834, 333, 370, 365, 611, 611, 611, 556, 333, 333, 494, 889, 278, 278, 1000, 584, 584, 611, + 611, 611, 474, 500, 500, 500, 278, 278, 278, 238, 389, 389, 549, 389, 389, 737, 333, 556, 556, 556, 556, + 556, 556, 333, 556, 556, 278, 278, 556, 600, 333, 389, 333, 611, 556, 834, 333, 333, 1000, 556, 333, 611, + 611, 611, 611, 611, 611, 611, 556, 611, 611, 556, 778, 556, 556, 556, 556, 556, 500, 500, 500, 500, 556, + ], + ascender: 718, + descender: -207, + italic_angle: 0.0, + }, + FaceMetrics { + // Helvetica-Oblique + widths: [ + 667, 1000, 667, 667, 667, 667, 667, 667, 667, 667, 667, 667, 722, 722, 722, 722, 722, 722, 722, 612, 667, + 667, 667, 667, 667, 667, 667, 667, 667, 722, 556, 611, 778, 778, 778, 722, 278, 278, 278, 278, 278, 278, + 278, 278, 500, 667, 667, 556, 556, 556, 556, 556, 833, 722, 722, 722, 722, 722, 778, 1000, 778, 778, 778, + 778, 778, 778, 778, 778, 667, 778, 722, 722, 722, 722, 667, 667, 667, 667, 667, 611, 611, 611, 667, 722, + 722, 722, 722, 722, 722, 722, 722, 722, 667, 944, 667, 667, 667, 667, 611, 611, 611, 611, 556, 556, 556, + 556, 333, 556, 889, 556, 556, 667, 556, 556, 469, 584, 389, 1015, 556, 556, 278, 260, 334, 334, 278, 278, + 333, 260, 350, 500, 500, 333, 500, 500, 333, 556, 333, 278, 278, 250, 737, 556, 556, 556, 556, 643, 556, + 400, 333, 584, 556, 333, 278, 556, 556, 556, 556, 556, 556, 556, 556, 1000, 556, 1000, 556, 556, 584, 556, + 278, 333, 278, 500, 556, 500, 556, 556, 167, 556, 556, 556, 611, 333, 584, 549, 556, 556, 333, 333, 556, + 333, 333, 222, 278, 278, 278, 278, 278, 222, 222, 500, 500, 222, 222, 299, 222, 584, 549, 584, 471, 222, + 833, 333, 584, 556, 584, 556, 556, 556, 556, 556, 549, 556, 556, 556, 556, 556, 556, 944, 333, 556, 556, + 556, 556, 834, 834, 333, 370, 365, 611, 556, 556, 537, 333, 333, 476, 889, 278, 278, 1000, 584, 584, 556, + 556, 611, 355, 333, 333, 333, 222, 222, 222, 191, 333, 333, 453, 333, 333, 737, 333, 500, 500, 500, 500, + 500, 556, 278, 556, 556, 278, 278, 556, 600, 278, 317, 278, 556, 556, 834, 333, 333, 1000, 556, 333, 556, + 556, 556, 556, 556, 556, 556, 556, 556, 556, 500, 722, 500, 500, 500, 500, 556, 500, 500, 500, 500, 556, + ], + ascender: 718, + descender: -207, + italic_angle: -12.0, + }, + FaceMetrics { + // Helvetica-BoldOblique + widths: [ + 722, 1000, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 722, 612, 667, + 667, 667, 667, 667, 667, 667, 667, 667, 722, 556, 611, 778, 778, 778, 722, 278, 278, 278, 278, 278, 278, + 278, 278, 556, 722, 722, 611, 611, 611, 611, 611, 833, 722, 722, 722, 722, 722, 778, 1000, 778, 778, 778, + 778, 778, 778, 778, 778, 667, 778, 722, 722, 722, 722, 667, 667, 667, 667, 667, 611, 611, 611, 667, 722, + 722, 722, 722, 722, 722, 722, 722, 722, 667, 944, 667, 667, 667, 667, 611, 611, 611, 611, 556, 556, 556, + 556, 333, 556, 889, 556, 556, 722, 556, 556, 584, 584, 389, 975, 556, 611, 278, 280, 389, 389, 333, 333, + 333, 280, 350, 556, 556, 333, 556, 556, 333, 556, 333, 333, 278, 250, 737, 556, 611, 556, 556, 743, 611, + 400, 333, 584, 556, 333, 278, 556, 556, 556, 556, 556, 556, 556, 556, 1000, 556, 1000, 556, 556, 584, 611, + 333, 333, 333, 611, 556, 611, 556, 556, 167, 611, 611, 611, 611, 333, 584, 549, 556, 556, 333, 333, 611, + 333, 333, 278, 278, 278, 278, 278, 278, 278, 278, 556, 556, 278, 278, 400, 278, 584, 549, 584, 494, 278, + 889, 333, 584, 611, 584, 611, 611, 611, 611, 556, 549, 611, 556, 611, 611, 611, 611, 944, 333, 611, 611, + 611, 556, 834, 834, 333, 370, 365, 611, 611, 611, 556, 333, 333, 494, 889, 278, 278, 1000, 584, 584, 611, + 611, 611, 474, 500, 500, 500, 278, 278, 278, 238, 389, 389, 549, 389, 389, 737, 333, 556, 556, 556, 556, + 556, 556, 333, 556, 556, 278, 278, 556, 600, 333, 389, 333, 611, 556, 834, 333, 333, 1000, 556, 333, 611, + 611, 611, 611, 611, 611, 611, 556, 611, 611, 556, 778, 556, 556, 556, 556, 556, 500, 500, 500, 500, 556, + ], + ascender: 718, + descender: -207, + italic_angle: -12.0, + }, + FaceMetrics { + // Times-Roman + widths: [ + 722, 889, 722, 722, 722, 722, 722, 722, 722, 722, 722, 667, 667, 667, 667, 667, 722, 722, 722, 612, 611, + 611, 611, 611, 611, 611, 611, 611, 611, 722, 500, 556, 722, 722, 722, 722, 333, 333, 333, 333, 333, 333, + 333, 333, 389, 722, 722, 611, 611, 611, 611, 611, 889, 722, 722, 722, 722, 722, 722, 889, 722, 722, 722, + 722, 722, 722, 722, 722, 556, 722, 667, 667, 667, 667, 556, 556, 556, 556, 556, 611, 611, 611, 556, 722, + 722, 722, 722, 722, 722, 722, 722, 722, 722, 944, 722, 722, 722, 722, 611, 611, 611, 611, 444, 444, 444, + 444, 333, 444, 667, 444, 444, 778, 444, 444, 469, 541, 500, 921, 444, 500, 278, 200, 480, 480, 333, 333, + 333, 200, 350, 444, 444, 333, 444, 444, 333, 500, 333, 278, 250, 250, 760, 500, 500, 500, 500, 588, 500, + 400, 333, 564, 500, 333, 278, 444, 444, 444, 444, 444, 444, 444, 500, 1000, 444, 1000, 500, 444, 564, 500, + 333, 333, 333, 556, 500, 556, 500, 500, 167, 500, 500, 500, 500, 333, 564, 549, 500, 500, 333, 333, 500, + 333, 333, 278, 278, 278, 278, 278, 278, 278, 278, 500, 500, 278, 278, 344, 278, 564, 549, 564, 471, 278, + 778, 333, 564, 500, 564, 500, 500, 500, 500, 500, 549, 500, 500, 500, 500, 500, 500, 722, 333, 500, 500, + 500, 500, 750, 750, 300, 276, 310, 500, 500, 500, 453, 333, 333, 476, 833, 250, 250, 1000, 564, 564, 500, + 444, 444, 408, 444, 444, 444, 333, 333, 333, 180, 333, 333, 453, 333, 333, 760, 333, 389, 389, 389, 389, + 389, 500, 278, 500, 500, 278, 250, 500, 600, 278, 326, 278, 500, 500, 750, 300, 333, 980, 500, 300, 500, + 500, 500, 500, 500, 500, 500, 500, 500, 500, 500, 722, 500, 500, 500, 500, 500, 444, 444, 444, 444, 500, + ], + ascender: 683, + descender: -217, + italic_angle: 0.0, + }, + FaceMetrics { + // Times-Bold + widths: [ + 722, 1000, 722, 722, 722, 722, 722, 722, 722, 722, 722, 667, 722, 722, 722, 722, 722, 722, 722, 612, 667, + 667, 667, 667, 667, 667, 667, 667, 667, 722, 500, 611, 778, 778, 778, 778, 389, 389, 389, 389, 389, 389, + 389, 389, 500, 778, 778, 667, 667, 667, 667, 667, 944, 722, 722, 722, 722, 722, 778, 1000, 778, 778, 778, + 778, 778, 778, 778, 778, 611, 778, 722, 722, 722, 722, 556, 556, 556, 556, 556, 667, 667, 667, 611, 722, + 722, 722, 722, 722, 722, 722, 722, 722, 722, 1000, 722, 722, 722, 722, 667, 667, 667, 667, 500, 500, 500, + 500, 333, 500, 722, 500, 500, 833, 500, 500, 581, 520, 500, 930, 500, 556, 278, 220, 394, 394, 333, 333, + 333, 220, 350, 444, 444, 333, 444, 444, 333, 500, 333, 333, 250, 250, 747, 500, 556, 500, 500, 672, 556, + 400, 333, 570, 500, 333, 278, 444, 444, 444, 444, 444, 444, 444, 500, 1000, 444, 1000, 500, 444, 570, 500, + 333, 333, 333, 556, 500, 556, 500, 500, 167, 500, 500, 500, 556, 333, 570, 549, 500, 500, 333, 333, 556, + 333, 333, 278, 278, 278, 278, 278, 278, 278, 333, 556, 556, 278, 278, 394, 278, 570, 549, 570, 494, 278, + 833, 333, 570, 556, 570, 556, 556, 556, 556, 500, 549, 556, 500, 500, 500, 500, 500, 722, 333, 500, 500, + 500, 500, 750, 750, 300, 300, 330, 500, 500, 556, 540, 333, 333, 494, 1000, 250, 250, 1000, 570, 570, 556, + 500, 500, 555, 500, 500, 500, 333, 333, 333, 278, 444, 444, 549, 444, 444, 747, 333, 389, 389, 389, 389, + 389, 500, 333, 500, 500, 278, 250, 500, 600, 333, 416, 333, 556, 500, 750, 300, 333, 1000, 500, 300, 556, + 556, 556, 556, 556, 556, 556, 500, 556, 556, 500, 722, 500, 500, 500, 500, 500, 444, 444, 444, 444, 500, + ], + ascender: 683, + descender: -217, + italic_angle: 0.0, + }, + FaceMetrics { + // Times-Italic + widths: [ + 611, 889, 611, 611, 611, 611, 611, 611, 611, 611, 611, 611, 667, 667, 667, 667, 722, 722, 722, 612, 611, + 611, 611, 611, 611, 611, 611, 611, 611, 722, 500, 611, 722, 722, 722, 722, 333, 333, 333, 333, 333, 333, + 333, 333, 444, 667, 667, 556, 556, 611, 556, 556, 833, 667, 667, 667, 667, 667, 722, 944, 722, 722, 722, + 722, 722, 722, 722, 722, 611, 722, 611, 611, 611, 611, 500, 500, 500, 500, 500, 556, 556, 556, 611, 722, + 722, 722, 722, 722, 722, 722, 722, 722, 611, 833, 611, 556, 556, 556, 556, 556, 556, 556, 500, 500, 500, + 500, 333, 500, 667, 500, 500, 778, 500, 500, 422, 541, 500, 920, 500, 500, 278, 275, 400, 400, 389, 389, + 333, 275, 350, 444, 444, 333, 444, 444, 333, 500, 333, 333, 250, 250, 760, 500, 500, 500, 500, 544, 500, + 400, 333, 675, 500, 333, 278, 444, 444, 444, 444, 444, 444, 444, 500, 889, 444, 889, 500, 444, 675, 500, + 333, 389, 278, 500, 500, 500, 500, 500, 167, 500, 500, 500, 500, 333, 675, 549, 500, 500, 333, 333, 500, + 333, 333, 278, 278, 278, 278, 278, 278, 278, 278, 444, 444, 278, 278, 300, 278, 675, 549, 675, 471, 278, + 722, 333, 675, 500, 675, 500, 500, 500, 500, 500, 549, 500, 500, 500, 500, 500, 500, 667, 333, 500, 500, + 500, 500, 750, 750, 300, 276, 310, 500, 500, 500, 523, 333, 333, 476, 833, 250, 250, 1000, 675, 675, 500, + 500, 500, 420, 556, 556, 556, 333, 333, 333, 214, 389, 389, 453, 389, 389, 760, 333, 389, 389, 389, 389, + 389, 500, 333, 500, 500, 278, 250, 500, 600, 278, 300, 278, 500, 500, 750, 300, 333, 980, 500, 300, 500, + 500, 500, 500, 500, 500, 500, 500, 500, 500, 444, 667, 444, 444, 444, 444, 500, 389, 389, 389, 389, 500, + ], + ascender: 683, + descender: -217, + italic_angle: -15.5, + }, + FaceMetrics { + // Times-BoldItalic + widths: [ + 667, 944, 667, 667, 667, 667, 667, 667, 667, 667, 667, 667, 667, 667, 667, 667, 722, 722, 722, 612, 667, + 667, 667, 667, 667, 667, 667, 667, 667, 722, 500, 667, 722, 722, 722, 778, 389, 389, 389, 389, 389, 389, + 389, 389, 500, 667, 667, 611, 611, 611, 611, 611, 889, 722, 722, 722, 722, 722, 722, 944, 722, 722, 722, + 722, 722, 722, 722, 722, 611, 722, 667, 667, 667, 667, 556, 556, 556, 556, 556, 611, 611, 611, 611, 722, + 722, 722, 722, 722, 722, 722, 722, 722, 667, 889, 667, 611, 611, 611, 611, 611, 611, 611, 500, 500, 500, + 500, 333, 500, 722, 500, 500, 778, 500, 500, 570, 570, 500, 832, 500, 500, 278, 220, 348, 348, 333, 333, + 333, 220, 350, 444, 444, 333, 444, 444, 333, 500, 333, 333, 250, 250, 747, 500, 500, 500, 500, 608, 500, + 400, 333, 570, 500, 333, 278, 444, 444, 444, 444, 444, 444, 444, 500, 1000, 444, 1000, 500, 444, 570, 500, + 389, 389, 333, 556, 500, 556, 500, 500, 167, 500, 500, 500, 500, 333, 570, 549, 500, 500, 333, 333, 556, + 333, 333, 278, 278, 278, 278, 278, 278, 278, 278, 500, 500, 278, 278, 382, 278, 570, 549, 606, 494, 278, + 778, 333, 606, 576, 570, 556, 556, 556, 556, 500, 549, 556, 500, 500, 500, 500, 500, 722, 333, 500, 500, + 500, 500, 750, 750, 300, 266, 300, 500, 500, 500, 500, 333, 333, 494, 833, 250, 250, 1000, 570, 570, 500, + 500, 500, 555, 500, 500, 500, 333, 333, 333, 278, 389, 389, 549, 389, 389, 747, 333, 389, 389, 389, 389, + 389, 500, 333, 500, 500, 278, 250, 500, 600, 278, 366, 278, 500, 500, 750, 300, 333, 1000, 500, 300, 556, + 556, 556, 556, 556, 556, 556, 500, 556, 556, 444, 667, 500, 444, 444, 444, 500, 389, 389, 389, 389, 500, + ], + ascender: 683, + descender: -217, + italic_angle: -15.0, + }, + FaceMetrics { + // Courier + widths: [ + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + ], + ascender: 629, + descender: -157, + italic_angle: 0.0, + }, + FaceMetrics { + // Courier-Bold + widths: [ + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + ], + ascender: 629, + descender: -157, + italic_angle: 0.0, + }, + FaceMetrics { + // Courier-Oblique + widths: [ + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + ], + ascender: 629, + descender: -157, + italic_angle: -12.0, + }, + FaceMetrics { + // Courier-BoldOblique + widths: [ + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, 600, + ], + ascender: 629, + descender: -157, + italic_angle: -12.0, + }, +]; diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/tests.rs b/src-tauri/src/pdf_engine/text_edit/fonts/tests.rs new file mode 100644 index 0000000..5e01efc --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/tests.rs @@ -0,0 +1,471 @@ +//! FONT-01…40 (SPEC §E.3) and the loader fuzz. This file: shared helpers, base encodings, the +//! AGL, `/Differences`, built-in encodings (Type1, CFF incl. FONT-35/36) and MacRoman (FONT-37). +//! The other groups live in `tests_tounicode.rs`, `tests_presence.rs`, `tests_classes.rs` and +//! `tests_fuzz.rs` (all `#[cfg(test)]`, declared by `fonts/mod.rs`). + +use super::agl::{agl_is_sorted, agl_len, agl_value, glyph_name_char}; +use super::encodings::BaseEncoding; +use super::{Code, FontClass, FontModel}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::testkit::cff::{CffBuilder, CffEncodingSpec}; +use crate::pdf_engine::text_edit::testkit::fonts::{ + latin_truetype, load_simple, tounicode_bfchar, Program, SimpleFont, +}; +use crate::pdf_engine::text_edit::testkit::type1::{T1Encoding, T1Glyph, Type1Builder}; + +pub(super) fn c1(v: u8) -> Code { + Code { + value: u32::from(v), + len: 1, + } +} + +pub(super) fn c2(v: u16) -> Code { + Code { + value: u32::from(v), + len: 2, + } +} + +/// The typeable characters as a string, sorted. +pub(super) fn alphabet(m: &FontModel) -> String { + m.alphabet().iter().map(|(c, _)| c).collect() +} + +pub(super) fn typeable(m: &FontModel, code: Code) -> Option { + m.info(code).and_then(|i| i.typeable_as) +} + +/// A WinAnsi simple TrueType font embedding `drawn` (outlined) and `blank` (empty) glyphs, with +/// width 500 for codes 32–255. +pub(super) fn winansi_truetype(base: &str, drawn: &str, blank: &str) -> SimpleFont { + let (program, _) = latin_truetype(drawn, blank); + let mut f = SimpleFont::new("TrueType", base); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.flags = Some(32); + f.program = Program::TrueType(program); + f +} + +/// A non-embedded Standard-14 font with an explicit `/Encoding`. +pub(super) fn std14(base: &str, encoding: &str) -> SimpleFont { + let mut f = SimpleFont::new("Type1", base); + f.encoding = Some(encoding.into()); + f +} + +/// lopdf 0.34's own table for a named encoding, through its public `get_font_encoding`. +fn lopdf_table(name: &str) -> [Option; 256] { + let doc = lopdf::Document::new(); + let dict = lopdf::dictionary! { "Type" => "Font", "Encoding" => lopdf::Object::Name(name.as_bytes().to_vec()) }; + match dict.get_font_encoding(&doc) { + Ok(lopdf::Encoding::OneByteEncoding(table)) => *table, + other => panic!("lopdf table for {name}: {other:?}"), + } +} + +const MAC_OS_ONLY: [u8; 15] = [ + 0xAD, 0xB0, 0xB2, 0xB3, 0xB6, 0xB7, 0xB8, 0xB9, 0xBA, 0xBD, 0xC3, 0xC5, 0xC6, 0xD7, 0xF0, +]; + +#[test] +fn font_01_encoding_tables_cross_check_and_spot_codes() { + assert_eq!(agl_len(), 4_495, "FONT-01 AGL size (lopdf glyphnames.rs)"); + assert!(agl_is_sorted(), "FONT-01 AGL sorted for binary search"); + let mut mac_differences = Vec::new(); + for (base, lopdf_name) in [ + (BaseEncoding::WinAnsi, "WinAnsiEncoding"), + (BaseEncoding::MacRoman, "MacRomanEncoding"), + (BaseEncoding::Standard, "StandardEncoding"), + ] { + let theirs = lopdf_table(lopdf_name); + for code in 0..=255u8 { + let ours = base.name(code).map(|n| { + u16::try_from(agl_value(n).unwrap_or_else(|| panic!("FONT-01 {n} in AGL"))) + .expect("BMP") + }); + let lopdf = theirs[usize::from(code)]; + if ours != lopdf { + assert_eq!( + base, + BaseEncoding::MacRoman, + "FONT-01 {lopdf_name} code {code:#04X}" + ); + assert!( + ours.is_none() && lopdf.is_some(), + "FONT-01 MacRoman {code:#04X}" + ); + mac_differences.push(code); + } + } + } + assert_eq!( + mac_differences, MAC_OS_ONLY, + "FONT-37 the explicit 15-code difference" + ); + let mac = BaseEncoding::MacRoman; + let win = BaseEncoding::WinAnsi; + let std = BaseEncoding::Standard; + assert_eq!(mac.name(0x80), Some("Adieresis")); + assert_eq!(mac.name(0x8E), Some("eacute")); + assert_eq!(mac.name(0xA5), Some("bullet")); + assert_eq!(mac.name(0xDB), Some("currency")); + assert_eq!(mac.name(0xAD), None); + assert_eq!(mac.name(0xF0), None); + assert_eq!(win.name(0x80), Some("Euro")); + assert_eq!(win.name(0x8A), Some("Scaron")); + assert_eq!(win.name(0x9F), Some("Ydieresis")); + assert_eq!(std.name(0x27), Some("quoteright")); + assert_eq!(std.name(0xE8), Some("Lslash")); + for code in 0x80..=0xA0u8 { + assert_eq!( + std.name(code), + None, + "FONT-01 Standard {code:#04X} undefined" + ); + } +} + +#[test] +fn font_02_agl_turkish_and_typographic_names() { + for (name, ch) in [ + ("gbreve", 'ğ'), + ("scedilla", 'ş'), + ("dotlessi", 'ı'), + ("Idotaccent", 'İ'), + ("Scedilla", 'Ş'), + ("fi", '\u{FB01}'), + ("quoteright", '’'), + ] { + assert_eq!(glyph_name_char(name), Some(ch), "FONT-02 {name}"); + } + // Through a font: Times (AFM has every one of them) with /Differences. + let f = std14( + "Times-Roman", + "<< /Differences [128 /gbreve /scedilla /dotlessi /Idotaccent /Scedilla /fi /quoteright] >>", + ); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + let texts: Vec<&str> = (128..=134).map(|c| m.text(c1(c)).unwrap_or("∅")).collect(); + assert_eq!( + texts, + ["ğ", "ş", "ı", "İ", "Ş", "fi", "’"], + "FONT-02 display texts" + ); + assert_eq!(typeable(&m, c1(128)), Some('ğ')); + assert_eq!( + typeable(&m, c1(133)), + None, + "FONT-02 the fi ligature is read, never typed" + ); + assert_eq!(typeable(&m, c1(134)), Some('’')); +} + +#[test] +fn font_03_uni_and_u_names_and_suffixed_names() { + assert_eq!(glyph_name_char("uni011F"), Some('ğ'), "FONT-03 uniXXXX"); + assert_eq!(glyph_name_char("u1F600"), Some('😀'), "FONT-03 uXXXXX"); + assert_eq!( + glyph_name_char("uni00410042"), + None, + "FONT-03 two groups name a sequence" + ); + assert_eq!( + glyph_name_char("uniD800"), + None, + "FONT-03 surrogates have no value" + ); + assert_eq!(glyph_name_char("a.sc"), None, "FONT-03 suffixed names"); + assert_eq!(glyph_name_char("u12"), None); + let mut f = SimpleFont::new("TrueType", "Calibri"); + f.encoding = Some("<< /Differences [65 /uni011F /u1F600 /a.sc] >>".into()); + f.first_char = 65; + f.widths = Some(vec![500.0, 500.0, 500.0]); + f.flags = Some(32); + let m = load_simple(&f); + assert_eq!(m.class, Some(FontClass::SimpleNonEmbedded)); + assert_eq!(m.text(c1(65)), Some("ğ")); + assert_eq!(typeable(&m, c1(65)), Some('ğ'), "FONT-03 uni011F writable"); + assert_eq!(m.text(c1(66)), Some("😀"), "FONT-03 u1F600 readable"); + assert_eq!(typeable(&m, c1(66)), None, "FONT-03 non-BMP never typeable"); + assert_eq!(m.text(c1(67)), None, "FONT-03 a.sc has no Unicode"); +} + +#[test] +fn font_04_differences() { + let m = load_simple(&std14( + "Times-Roman", + "<< /BaseEncoding /WinAnsiEncoding /Differences [65 /gbreve /scedilla 128 /Idotaccent] >>", + )); + assert_eq!(m.text(c1(65)), Some("ğ")); + assert_eq!(m.text(c1(66)), Some("ş")); + assert_eq!( + m.text(c1(67)), + Some("C"), + "FONT-04 base WinAnsi below the differences" + ); + assert_eq!(m.text(c1(128)), Some("İ")); + assert_eq!( + m.info(c1(65)).and_then(|i| i.glyph_name.clone()).as_deref(), + Some("gbreve") + ); + assert!(alphabet(&m).contains('ğ')); + // No /BaseEncoding on a non-embedded non-symbolic font: StandardEncoding underneath. + let m = load_simple(&std14("Helvetica", "<< /Differences [65 /Aring] >>")); + assert_eq!( + m.text(c1(0x27)), + Some("’"), + "FONT-04 implicit Standard base" + ); + assert_eq!(m.text(c1(65)), Some("Å")); + for bad in [ + "<< /Differences [65 (A)] >>", + "<< /Differences [/A 65] >>", + "<< /Differences [255 /a /b] >>", + "<< /Differences [-1 /a] >>", + "<< /Differences 5 >>", + "<< /BaseEncoding 5 >>", + ] { + let m = load_simple(&std14("Helvetica", bad)); + assert_eq!( + m.refusal, + Some(TextReason::FontUnsupported), + "FONT-04 {bad}" + ); + } + let m = load_simple(&std14("Helvetica", "/MacExpertEncoding")); + assert_eq!( + m.refusal, + Some(TextReason::UnsupportedEncoding), + "FONT-04 MacExpert" + ); + let m = load_simple(&std14( + "Helvetica", + "<< /BaseEncoding /MacExpertEncoding >>", + )); + assert_eq!(m.refusal, Some(TextReason::UnsupportedEncoding)); +} + +/// A Type1-subtype font embedding `program`, with /Widths 500 for 32–255 and no /Encoding. +fn embedded_type1(program: Program) -> SimpleFont { + let mut f = SimpleFont::new("Type1", "ABCDEF+TestSerif"); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.flags = Some(32); + f.program = program; + f +} + +#[test] +fn font_05_type1_builtin_encoding() { + let t1 = Type1Builder { + encoding: T1Encoding::Custom(vec![(65, "A".into()), (66, "gbreve".into())]), + ..Type1Builder::new("TestSerif") + } + .glyph("A", T1Glyph::Box { width: 600 }) + .glyph("gbreve", T1Glyph::Box { width: 500 }) + .glyph("C", T1Glyph::Box { width: 600 }) + .build(); + let m = load_simple(&embedded_type1(Program::Type1(t1))); + assert_eq!(m.refusal, None); + assert_eq!(m.class, Some(FontClass::SimpleType1)); + assert_eq!(typeable(&m, c1(65)), Some('A'), "FONT-05 dup 65 /A put"); + assert_eq!( + typeable(&m, c1(66)), + Some('ğ'), + "FONT-05 dup 66 /gbreve put" + ); + assert_eq!( + m.text(c1(67)), + None, + "FONT-05 code 67 is .notdef in the built-in encoding" + ); + assert_eq!(alphabet(&m), "Ağ"); + // `/Encoding StandardEncoding def`. + let t1 = Type1Builder::new("TestSerif") + .glyph("quoteright", T1Glyph::Box { width: 300 }) + .build(); + let m = load_simple(&embedded_type1(Program::Type1(t1))); + assert_eq!( + typeable(&m, c1(0x27)), + Some('’'), + "FONT-05 Standard built-in" + ); +} + +#[test] +fn font_06_cff_builtin_encoding() { + let cff = CffBuilder::new("TestSans") + .glyph("A", true) + .glyph("B", true) + .glyph("C", true) + .encoding(CffEncodingSpec::Format0(vec![0x41, 0x42])) + .build(); + let m = load_simple(&embedded_type1(Program::Cff(cff))); + assert_eq!(m.refusal, None); + assert_eq!(m.class, Some(FontClass::SimpleCff)); + assert_eq!( + alphabet(&m), + "AB", + "FONT-06 format 0: 0x41 → GID 1, 0x42 → GID 2" + ); + assert_eq!(m.info(c1(0x41)).and_then(|i| i.gid), Some(1)); + let cff = CffBuilder::new("TestSans") + .glyph("a", true) + .glyph("b", true) + .glyph("c", true) + .encoding(CffEncodingSpec::Format1(vec![(0x61, 2)])) + .build(); + let m = load_simple(&embedded_type1(Program::Cff(cff))); + assert_eq!(alphabet(&m), "abc", "FONT-06 format 1 range"); + let cff = CffBuilder::new("TestSans") + .glyph("quoteright", true) + .build(); + let m = load_simple(&embedded_type1(Program::Cff(cff))); + assert_eq!( + alphabet(&m), + "’", + "FONT-06 Standard built-in → names → GIDs" + ); + // A PDF /Encoding overrides the built-in one: names → glyph_index_by_name. + let cff = CffBuilder::new("TestSans") + .glyph("A", true) + .glyph("gbreve", true) + .encoding(CffEncodingSpec::Format0(vec![0x41, 0x42])) + .build(); + let mut f = embedded_type1(Program::Cff(cff)); + f.encoding = Some("<< /BaseEncoding /WinAnsiEncoding /Differences [200 /gbreve] >>".into()); + let m = load_simple(&f); + assert_eq!(typeable(&m, c1(200)), Some('ğ')); + assert_eq!(typeable(&m, c1(0x41)), Some('A')); + assert_eq!( + typeable(&m, c1(0x42)), + None, + "FONT-06 WinAnsi B has no glyph" + ); + // /Differences without /BaseEncoding sit on the built-in encoding, and a name they set + // replaces the built-in code → GID mapping. + let cff = CffBuilder::new("TestSans") + .glyph("A", true) + .glyph("B", true) + .encoding(CffEncodingSpec::Format0(vec![0x41, 0x42])) + .build(); + let mut f = embedded_type1(Program::Cff(cff)); + f.encoding = Some("<< /Differences [65 /B] >>".into()); + let m = load_simple(&f); + assert_eq!( + m.text(c1(0x41)), + Some("B"), + "FONT-06 the difference names code 65 B" + ); + assert_eq!( + m.info(c1(0x41)).and_then(|i| i.gid), + Some(2), + "FONT-06 B's own GID" + ); + assert_eq!( + m.info(c1(0x42)).and_then(|i| i.gid), + Some(2), + "FONT-06 built-in 0x42" + ); +} + +#[test] +fn font_35_cff_custom_encoding_never_falls_back_to_standard() { + // GID 1 = "B", GID 2 = "A"; the custom encoding maps only 0x42 → GID 1. "A" is in the + // charset, and StandardEncoding has "A" at 0x41, but no code of this font maps to it. + let cff = CffBuilder::new("TestSans") + .glyph("B", true) + .glyph("A", true) + .encoding(CffEncodingSpec::Format0(vec![0x42])) + .build(); + let table = ttf_parser::cff::Table::parse(&cff).expect("CFF parses"); + assert_eq!( + table.glyph_index(0x41).map(|g| g.0), + Some(2), + "FONT-35 precondition: ttf-parser's glyph_index falls back to StandardEncoding" + ); + let m = load_simple(&embedded_type1(Program::Cff(cff))); + assert_eq!(m.refusal, None); + assert_eq!(typeable(&m, c1(0x41)), None, "FONT-35 A is not typeable"); + assert!(!m.drawable(c1(0x41)), "FONT-35 code 0x41 is not drawable"); + assert_eq!(typeable(&m, c1(0x42)), Some('B')); + assert_eq!(alphabet(&m), "B"); +} + +#[test] +fn font_36_cff_supplements_and_expert_are_not_typeable() { + let cff = CffBuilder::new("TestSans") + .glyph("A", true) + .glyph("B", true) + .encoding(CffEncodingSpec::Format0(vec![0x41])) + .supplement(0x61, "B") + .build(); + let table = ttf_parser::cff::Table::parse(&cff).expect("CFF parses"); + assert_eq!( + table.glyph_index(0x61).map(|g| g.0), + Some(2), + "precondition: supplement" + ); + let m = load_simple(&embedded_type1(Program::Cff(cff))); + assert_eq!(m.refusal, None); + assert_eq!( + typeable(&m, c1(0x61)), + None, + "FONT-36 supplement code ignored" + ); + assert_eq!(alphabet(&m), "A", "FONT-36 only the format-0 code types"); + let cff = CffBuilder::new("TestSans") + .glyph("A", true) + .encoding(CffEncodingSpec::Expert) + .build(); + let mut f = embedded_type1(Program::Cff(cff)); + f.tounicode = Some(tounicode_bfchar(&[(0x41, 1, "A")])); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!( + m.text(c1(0x41)), + Some("A"), + "FONT-36 Expert codes readable via ToUnicode" + ); + assert_eq!(alphabet(&m), "", "FONT-36 Expert encoding types nothing"); +} + +#[test] +fn font_37_mac_roman_is_strict_annex_d() { + let m = load_simple(&std14("Helvetica", "/MacRomanEncoding")); + assert_eq!(m.refusal, None); + assert_eq!(m.text(c1(0xDB)), Some("¤"), "FONT-37 0xDB is currency"); + assert_eq!(typeable(&m, c1(0xDB)), Some('¤')); + assert_eq!( + m.text(c1(0xCA)), + Some(" "), + "FONT-37 0xCA is a second space" + ); + for code in MAC_OS_ONLY { + assert_eq!( + m.text(c1(code)), + None, + "FONT-37 {code:#04X} undecodable without ToUnicode" + ); + assert_eq!( + typeable(&m, c1(code)), + None, + "FONT-37 {code:#04X} never typeable" + ); + } + // With ToUnicode the code reads, but has no glyph name, so it never types. + let mut f = std14("Helvetica", "/MacRomanEncoding"); + f.tounicode = Some(tounicode_bfchar(&[(0xAD, 1, "≠"), (0xB9, 1, "π")])); + let m = load_simple(&f); + assert_eq!(m.text(c1(0xAD)), Some("≠")); + assert_eq!( + typeable(&m, c1(0xAD)), + None, + "FONT-37 ToUnicode alone never makes it typeable" + ); + assert_eq!(m.text(c1(0xB9)), Some("π")); + assert_eq!(typeable(&m, c1(0xB9)), None); +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/tests_bounds.rs b/src-tauri/src/pdf_engine/text_edit/fonts/tests_bounds.rs new file mode 100644 index 0000000..116eb96 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/tests_bounds.rs @@ -0,0 +1,780 @@ +//! Work and memory bounds of the font loader, and the reading fixes of the review-T2 fix pass: +//! the Type1 interpreter (FONT-24b–d), the outline pre-checks against ttf-parser (FONT-29c–f, +//! FONT-30b), ToUnicode keyed by value and its text budget (FONT-08b, FONT-14b/c), descriptor +//! and refusal-order fixes (FONT-28c/d), the Type3 hash budget (FONT-28e) and the character rules +//! (FONT-12b, FONT-15b). + +use super::cff_layout::CffLayout; +use super::encodings::{reading_reason, writable_char}; +use super::glyph_budget::{WorkMeter, GLYPH_WORK_MAX}; +use super::program::{cff_cid_to_gid, CffNames, Outlines}; +use super::tests::{alphabet, c1, c2, std14, typeable, winansi_truetype}; +use super::tests_fuzz::thread_cpu; +use super::tounicode::parse_tounicode; +use super::type1::{parse_type1, GlyphProof}; +use super::{FontCache, FontKey, FontModel}; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::limits::PAGE_DECODE_BUDGET; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::testkit::cff::{t2, CffBuilder}; +use crate::pdf_engine::text_edit::testkit::fonts::{ + add_type0, cmap, latin_truetype, load_simple, load_type0, page_with_fonts, snapshot, + tounicode_bfchar, Program, SimpleFont, Type0Font, +}; +use crate::pdf_engine::text_edit::testkit::pdf::PdfBuilder; +use crate::pdf_engine::text_edit::testkit::thread_peak; +use crate::pdf_engine::text_edit::testkit::ttf::{composite, TtfBuilder}; +use crate::pdf_engine::text_edit::testkit::type1::{op, T1Glyph, Type1Builder, Type1File}; +use lopdf::Object; +use std::sync::Arc; +use ttf_parser::{GlyphId, OutlineBuilder}; + +/// `f`'s result and the CPU seconds it took on this thread. +pub(super) fn cpu(f: impl FnOnce() -> R) -> (R, f64) { + let started = thread_cpu(); + let r = f(); + (r, thread_cpu().saturating_sub(started).as_secs_f64()) +} + +/// Loads object `font` of `b` (as the one font of a page) with `budget`. +fn load_with(b: PdfBuilder, font: u32, budget: &mut DecodeBudget) -> Arc { + let snap = snapshot(page_with_fonts(b, &[("F1", font)])); + let dict = snap + .doc + .get_object((font, 0)) + .and_then(Object::as_dict) + .expect("font dict"); + FontCache::new().get_or_load(&snap.doc, FontKey::Indirect((font, 0)), dict, budget) +} + +/// A Type1 charstring: `hsbw`, `n` calls of subroutine `subr`, a drawn segment, `endchar`. +fn calling(subr: i32, n: usize) -> Vec { + let mut code = op(&[0, 600], &[13]); + for _ in 0..n { + code.extend(op(&[subr], &[10])); + } + code.extend(op(&[100, 0], &[21])); + code.extend(op(&[400, 0], &[5])); + code.push(14); + code +} + +fn type1_simple(base: &str, file: Type1File) -> SimpleFont { + let mut f = SimpleFont::new("Type1", base); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(vec![600.0; 224]); + f.flags = Some(32); + f.program = Program::Type1(file); + f.flate = false; + f +} + +const LETTERS: &str = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"; + +#[test] +fn font_24b_type1_subroutine_calls_decrypt_nothing() { + // Subr 5: `return` and 60,000 filler bytes (inside the 65,535-byte limit); subr 6: `return` + // and 4 MiB. Before the fix every call decrypted the whole subroutine again: 2,000 calls of + // subr 5 per glyph, 52 glyphs, were 6 GB of decryption. + let mut long = vec![11u8]; + long.resize(60_000, 0); + let mut huge = vec![11u8]; + huge.resize(4 << 20, 0); + let mut b = Type1Builder { + extra_subrs: vec![long, huge], + ..Type1Builder::new("Slow") + }; + for ch in LETTERS.chars() { + b = b.glyph(&ch.to_string(), T1Glyph::Raw(calling(5, 2_000))); + } + let file = b.glyph("huge", T1Glyph::Raw(calling(6, 1))).build(); + let (proofs, secs) = cpu(|| { + let program = parse_type1(&file.data, file.length1, file.length2).expect("parses"); + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + let drawn = LETTERS + .chars() + .filter(|c| program.proof(&c.to_string(), &mut meter) == GlyphProof::Drawn) + .count(); + (drawn, program.proof("huge", &mut meter)) + }); + assert_eq!( + proofs, + (52, GlyphProof::Unproven), + "FONT-24b 2,000 calls of a 60 KB subroutine prove; a 4 MiB subroutine never runs" + ); + assert!(secs < 2.0, "FONT-24b bounded ({secs:.2} s of CPU)"); + let (m, secs) = cpu(|| load_simple(&type1_simple("ABCDEF+Slow", file))); + assert_eq!(alphabet(&m), LETTERS, "FONT-24b through the loader"); + assert!(secs < 3.0, "FONT-24b loader bounded ({secs:.2} s of CPU)"); +} + +#[test] +fn font_24c_type1_seac_components_share_the_glyph_budget() { + // Subr 5 pushes 46 numbers and clears them: 48 tokens and 3 operators per call. A glyph of + // 1,360 calls stays under 4,096 operators and spends ~68,000 tokens; a seac of two such + // glyphs spends ~136,000 > GLYPH_WORK_MAX. + let mut subr = Vec::new(); + for _ in 0..46 { + subr.extend(op(&[1], &[])); + } + subr.extend([9, 11]); + let file = Type1Builder { + extra_subrs: vec![subr], + ..Type1Builder::new("Heavy") + } + .glyph("A", T1Glyph::Raw(calling(5, 1_360))) + .glyph("acute", T1Glyph::Raw(calling(5, 1_360))) + .glyph("B", T1Glyph::Box { width: 600 }) + .glyph( + "Aacute", + T1Glyph::Seac { + width: 600, + base: 65, + accent: 0xC2, + }, + ) + .glyph( + "Bacute", + T1Glyph::Seac { + width: 600, + base: 66, + accent: 0xC2, + }, + ) + .build(); + let program = parse_type1(&file.data, file.length1, file.length2).expect("parses"); + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + assert_eq!(program.proof("A", &mut meter), GlyphProof::Drawn); + assert_eq!(program.proof("Bacute", &mut meter), GlyphProof::Drawn); + assert_eq!( + program.proof("Aacute", &mut meter), + GlyphProof::Unproven, + "FONT-24c the components' tokens count against one glyph budget" + ); + // The meter is the page budget: a meter that cannot pay leaves the glyph unproven. + let mut poor = WorkMeter::new(100); + assert_eq!(program.proof("A", &mut poor), GlyphProof::Unproven); + assert!( + poor.exhausted(), + "FONT-24c an exhausted meter marks the load" + ); +} + +#[test] +fn font_24d_seac_takes_exactly_five_operands() { + let seac = |args: &[i32]| [op(&[0, 600], &[13]), op(args, &[12, 6])].concat(); + let file = Type1Builder::new("Seac") + .glyph("A", T1Glyph::Box { width: 600 }) + .glyph("acute", T1Glyph::Box { width: 300 }) + .glyph("five", T1Glyph::Raw(seac(&[0, 0, 0, 65, 194]))) + .glyph("six", T1Glyph::Raw(seac(&[0, 0, 65, 194, 65, 194]))) + .build(); + let program = parse_type1(&file.data, file.length1, file.length2).expect("parses"); + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + assert_eq!(program.proof("five", &mut meter), GlyphProof::Drawn); + assert_eq!( + program.proof("six", &mut meter), + GlyphProof::Unproven, + "FONT-24d six operands are not a seac (the components would come from the wrong slots)" + ); +} + +/// Segments of glyph `gid` as ttf-parser draws them, with no pre-check. +#[derive(Default)] +pub(super) struct Segs(pub(super) usize); + +impl OutlineBuilder for Segs { + fn move_to(&mut self, _: f32, _: f32) {} + fn line_to(&mut self, _: f32, _: f32) { + self.0 += 1; + } + fn quad_to(&mut self, _: f32, _: f32, _: f32, _: f32) { + self.0 += 1; + } + fn curve_to(&mut self, _: f32, _: f32, _: f32, _: f32, _: f32, _: f32) { + self.0 += 1; + } + fn close(&mut self) {} +} + +#[test] +fn font_29c_truetype_composite_fan_out_is_bounded() { + let mut t = TtfBuilder::new(); + let a = t.unicode_glyph('A', "A", true); + let c = t.raw_glyph("C", composite(&[a, a]), 500); + t.cmap31.push((u32::from('C'), c)); + // 30 levels of two references to the level below: 2^30 triangles for ttf-parser. + let mut level = a; + for depth in 0..30 { + level = t.raw_glyph(&format!("level{depth}"), composite(&[level, level]), 500); + } + t.cmap31.push((u32::from('B'), level)); + let data = t.build(); + let face = ttf_parser::Face::parse(&data, 0).expect("face"); + let outlines = Outlines::of_face(&face); + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + let (drawn, secs) = cpu(|| { + ( + outlines.drawn(a, &mut meter), + outlines.drawn(c, &mut meter), + outlines.drawn(level, &mut meter), + ) + }); + assert_eq!( + drawn, + (true, true, false), + "FONT-29c the fan-out never draws" + ); + assert!(secs < 1.0, "FONT-29c bounded ({secs:.2} s of CPU)"); + assert!( + meter.used() <= 3 * GLYPH_WORK_MAX as usize, + "FONT-29c work charged" + ); + let mut f = winansi_truetype("ABCDEF+Fan", "", ""); + f.program = Program::TrueType(data); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(alphabet(&m), "AC", "FONT-29c through the loader"); +} + +/// A name-keyed CFF: a box `A`, an `acute` box, and glyphs that use global and local +/// subroutines, hint masks and `seac`; `B` fans out 8 calls per level over 9 levels. +pub(super) fn subroutine_cff() -> Vec { + let call_global = |j: i32| [t2(j - 107), vec![29]].concat(); + let mut gsubrs: Vec> = (0..9) + .map(|k| { + let mut s: Vec = (0..8).flat_map(|_| call_global(k + 1)).collect(); + s.push(11); + s + }) + .collect(); + gsubrs.push(vec![11]); + let moveto = [t2(100), t2(0), vec![21]].concat(); + let lineto = [t2(400), t2(0), vec![5]].concat(); + let glyph = |middle: Vec| [moveto.clone(), middle, lineto.clone(), vec![14]].concat(); + CffBuilder { + gsubrs, + subrs: vec![[t2(0), t2(500), vec![5, 11]].concat()], + ..CffBuilder::new("Subrs") + } + .glyph("A", true) + .raw_glyph("B", glyph(call_global(0))) + .raw_glyph("C", glyph(call_global(9))) + .raw_glyph("D", glyph([t2(-107), vec![10]].concat())) + .raw_glyph( + "E", + [t2(0), t2(50), vec![18, 19, 0x80], glyph(Vec::new())].concat(), + ) + .glyph("acute", true) + .raw_glyph("Aacute", [t2(0), t2(0), t2(65), t2(194), vec![14]].concat()) + // `seac` runs its base one level deeper: G's fan-out (from gsubr 1) still fits ttf-parser's + // nesting limit there, so Gacute is a bomb for ttf-parser unless the pre-check follows seac. + .raw_glyph("Gacute", [t2(0), t2(0), t2(71), t2(194), vec![14]].concat()) + // A hint mask byte that reads as `callgsubr` unless the mask is skipped as ttf-parser does. + .raw_glyph( + "F", + [t2(0), t2(50), vec![18, 19, 29], glyph(Vec::new())].concat(), + ) + .raw_glyph("G", glyph(call_global(1))) + .build() +} + +#[test] +fn font_29d_cff_subroutine_fan_out_is_bounded() { + let data = subroutine_cff(); + let outlines = Outlines::of_cff(&data).expect("CFF"); + assert!( + matches!(outlines, Outlines::Cff { .. }), + "FONT-29d laid out" + ); + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + let (drawn, secs) = cpu(|| { + (1..=10) + .map(|g| outlines.drawn(g, &mut meter)) + .collect::>() + }); + assert_eq!( + drawn, + [true, false, true, true, true, true, true, false, true, false], + "FONT-29d A, C (global), D (local), E (hint mask), acute, Aacute (seac) and F (mask byte \ + 29) draw; B, G (fan-outs) and Gacute (a seac whose base fans out) never" + ); + assert!(secs < 1.0, "FONT-29d bounded ({secs:.2} s of CPU)"); + // Bare CFF (Type1C) through the loader. + let mut f = SimpleFont::new("Type1", "ABCDEF+Subrs"); + f.encoding = Some("<< /Differences [65 /A /B /C /D /E 193 /Aacute] >>".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.flags = Some(32); + f.program = Program::Cff(data.clone()); + let typed = |m: &FontModel| -> String { + [65u8, 66, 67, 68, 69, 193] + .iter() + .filter_map(|c| typeable(m, c1(*c))) + .collect() + }; + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(typed(&m), "ACDEÁ", "FONT-29d Type1C"); + // The same program as OpenType-CFF. + let mut t = TtfBuilder::new(); + for (ch, name) in [('A', "A"), ('B', "B"), ('C', "C"), ('D', "D"), ('E', "E")] { + t.unicode_glyph(ch, name, true); + } + t.glyph("acute", true, 500); + t.unicode_glyph('Á', "Aacute", true); + t.glyph("Gacute", true, 500); + t.glyph("F", true, 500); + t.glyph("G", true, 500); + t.cff = Some(data); + let mut f = winansi_truetype("ABCDEF+SubrsOT", "", ""); + f.program = Program::OpenType(t.build()); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(typed(&m), "ACDEÁ", "FONT-29d OpenType-CFF"); +} + +/// Real fonts of the repository: (path, bare CFF?). +fn real_fonts() -> Vec<(String, bool)> { + let root = std::path::Path::new(env!("CARGO_MANIFEST_DIR")); + let mut fonts: Vec<(String, bool)> = [ + "FoxitDingbats", + "FoxitFixed", + "FoxitFixedBold", + "FoxitFixedBoldItalic", + "FoxitFixedItalic", + "FoxitSerif", + "FoxitSerifBold", + "FoxitSerifBoldItalic", + "FoxitSerifItalic", + "FoxitSymbol", + ] + .iter() + .map(|n| (format!("../public/pdfjs/standard_fonts/{n}.pfb"), true)) + .collect(); + for n in ["Regular", "Bold", "Italic", "BoldItalic"] { + fonts.push(( + format!("../public/pdfjs/standard_fonts/LiberationSans-{n}.ttf"), + false, + )); + } + fonts.push(("resources/fonts/NotoSans-Regular.ttf".into(), false)); + fonts + .into_iter() + .map(|(p, cff)| (root.join(p).to_string_lossy().into_owned(), cff)) + .collect() +} + +#[test] +fn font_29e_pre_check_never_rejects_a_glyph_ttf_parser_draws() { + let mut glyphs = 0usize; + for (path, bare_cff) in real_fonts() { + let data = std::fs::read(&path).unwrap_or_else(|e| panic!("{path}: {e}")); + let face = (!bare_cff).then(|| ttf_parser::Face::parse(&data, 0).expect("face")); + let outlines = match &face { + Some(face) => Outlines::of_face(face), + None => Outlines::of_cff(&data).expect("CFF"), + }; + let mut meter = WorkMeter::new(usize::MAX); + let count = match (&outlines, &face) { + (Outlines::Glyf { .. }, Some(face)) => face.number_of_glyphs(), + (Outlines::Cff { table, .. }, _) => table.number_of_glyphs(), + _ => panic!("FONT-29e {path}: outlines readable"), + }; + for gid in 0..count { + let mut segs = Segs::default(); + let direct = match &outlines { + Outlines::Glyf { table, .. } => table.outline(GlyphId(gid), &mut segs).is_some(), + Outlines::Cff { table, .. } => table.outline(GlyphId(gid), &mut segs).is_ok(), + _ => false, + }; + assert_eq!( + outlines.drawn(gid, &mut meter), + direct && segs.0 > 0, + "FONT-29e {path} GID {gid}" + ); + } + assert!(!meter.exhausted()); + glyphs += usize::from(count); + if let Outlines::Cff { guard, table } = &outlines { + names_match_ttf_parser(guard.layout(), table, &path); + } + } + assert!(glyphs > 8_000, "FONT-29e {glyphs} glyphs compared"); +} + +/// CffNames (one charset walk) gives every name ttf-parser's `glyph_name` gives, first GID first. +fn names_match_ttf_parser(layout: &CffLayout<'_>, table: &ttf_parser::cff::Table<'_>, what: &str) { + let names = CffNames::new(layout, table); + let mut first = std::collections::HashMap::new(); + for gid in 0..table.number_of_glyphs() { + if let Some(name) = table.glyph_name(GlyphId(gid)) { + first.entry(name.to_string()).or_insert(gid); + } + } + for (name, gid) in &first { + assert_eq!(names.gid(name), Some(*gid), "{what}: {name}"); + } +} + +#[test] +fn font_29f_charset_maps_are_linear() { + // 40,000 glyphs, one charset range each: ttf-parser's per-glyph walk is 8 × 10^8 steps. + let mut b = CffBuilder { + charset_format: 2, + ..CffBuilder::new("Many") + }; + for i in 0..40_000 { + b = b.glyph(&format!("g{i:05}"), i % 2 == 0); + } + let data = b.build(); + let table = ttf_parser::cff::Table::parse(&data).expect("CFF"); + let layout = CffLayout::parse(&data).expect("layout"); + let (names, secs) = cpu(|| CffNames::new(&layout, &table)); + assert!(secs < 1.0, "FONT-29f O(glyphs) ({secs:.2} s of CPU)"); + for gid in [1u16, 2, 777, 20_000, 39_999, 40_000] { + let name = table.glyph_name(GlyphId(gid)).expect("named"); + assert_eq!(names.gid(name), Some(gid), "FONT-29f {name}"); + } + // CID-keyed, formats 1 and 2: the CID → GID map equals ttf-parser's glyph_cid inverse. + for format in [1, 2] { + let mut b = CffBuilder { + charset_format: format, + ..CffBuilder::new("ManyCID") + }; + for cid in (0..600).map(|i| (i * 7 + 3) % 1_000 + 1) { + b = b.cid_glyph(cid, true); + } + let data = b.build(); + let table = ttf_parser::cff::Table::parse(&data).expect("CFF"); + let map = cff_cid_to_gid(&CffLayout::parse(&data).expect("layout")).expect("CID"); + for gid in 0..table.number_of_glyphs() { + let cid = table.glyph_cid(GlyphId(gid)).expect("cid"); + assert_eq!( + map.get(&cid).copied(), + Some(gid), + "FONT-29f format {format}" + ); + } + } +} + +/// A minimal CFF2 table: Top DICT with CharStrings, empty Global Subr INDEX, one charstring. +fn minimal_cff2() -> Vec { + let mut d = vec![2, 0, 5, 0, 6]; + d.extend([29, 0, 0, 0, 15, 17]); // CharStrings at 15 + d.extend([0, 0, 0, 0]); // Global Subr INDEX (u32 count 0) + d.extend([0, 0, 0, 1, 1, 1, 2, 139]); // CharStrings: one charstring + d +} + +#[test] +fn font_30b_cff2_only_programs_are_unsupported() { + let mut t = TtfBuilder::new(); + t.unicode_glyph('A', "A", true); + t.cff2 = Some(minimal_cff2()); + let data = t.build(); + let face = ttf_parser::Face::parse(&data, 0).expect("face"); + assert!( + face.tables().cff2.is_some(), + "FONT-30b the fixture has a CFF2 table" + ); + let mut f = winansi_truetype("ABCDEF+Variable", "", ""); + f.program = Program::OpenType(data); + assert_eq!( + load_simple(&f).refusal, + Some(TextReason::FontProgramUnsupported), + "FONT-30b CFF2 charstrings are not pre-checked, so never outlined" + ); +} + +#[test] +fn font_08b_tounicode_codes_match_by_value() { + let two_byte = cmap( + "1 begincodespacerange <0000> endcodespacerange\n\ + 3 beginbfchar\n<0041> <0042>\n<0042> <0042>\n<0043> <0106>\nendbfchar", + ); + let mut f = std14("Helvetica", "/WinAnsiEncoding"); + f.tounicode = Some(two_byte); + let m = load_simple(&f); + assert_eq!(m.text(c1(0x41)), None, "FONT-08b name A vs ToUnicode B"); + assert_eq!(typeable(&m, c1(0x41)), None); + assert_eq!(typeable(&m, c1(0x42)), Some('B'), "FONT-08b agreement"); + assert_eq!(m.text(c1(0x43)), None, "FONT-08b C vs Ć"); + // A Type0 font's ToUnicode with 1-byte sources maps the 2-byte codes of the same value. + let (program, gids) = latin_truetype("A", ""); + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Short"); + f.program = Program::TrueType(program); + let mut map = vec![0u8; 0x42 * 2]; + map[0x41 * 2 + 1] = gids[0].1 as u8; + f.cid_to_gid = Some(Some(map)); + f.tounicode = Some(cmap( + "1 begincodespacerange <00> endcodespacerange\n1 beginbfchar\n<41> <0041>\nendbfchar", + )); + let m = load_type0(&f); + assert_eq!(m.refusal, None); + assert_eq!(m.text(c2(0x41)), Some("A"), "FONT-08b <41> maps CID 0x0041"); + assert_eq!(typeable(&m, c2(0x41)), Some('A')); +} + +#[test] +fn font_14b_tounicode_text_is_bounded_before_expansion() { + // 65,536 codes × a 512-byte destination is 32 MiB of text from 1 KB of CMap. + let dst = "0041".repeat(256); + let body = format!("1 beginbfrange\n<0000> <{dst}>\nendbfrange"); + let (parsed, peak) = thread_peak(|| parse_tounicode(&cmap(&body)).is_ok()); + assert!(!parsed, "FONT-14b destinations past TOUNICODE_MAX_DECODED"); + assert!( + peak < 8 << 20, + "FONT-14b nothing expanded ({peak} bytes peak)" + ); + // A short destination over the same range still parses. + let body = "1 beginbfrange\n<0000> <0020>\nendbfrange"; + assert!(parse_tounicode(&cmap(body)).is_ok()); +} + +#[test] +fn font_14c_model_memory_is_paid_from_the_page_budget() { + // Each font maps all 65,536 codes from one bfrange line (~9 MB of model). Five of them under + // a 32 MiB budget: the first loads, later ones run the budget out. + let mut b = PdfBuilder::new(); + let (program, _) = latin_truetype("A", ""); + let ids: Vec = (0..5) + .map(|i| { + let mut f = Type0Font::new("CIDFontType2", &format!("ABCDEF+Wide{i}")); + f.program = Program::TrueType(program.clone()); + f.tounicode = Some(cmap("1 beginbfrange\n<0000> <0020>\nendbfrange")); + add_type0(&mut b, &f) + }) + .collect(); + let names: Vec = (0..ids.len()).map(|i| format!("F{i}")).collect(); + let fonts: Vec<(&str, u32)> = names + .iter() + .map(String::as_str) + .zip(ids.iter().copied()) + .collect(); + let snap = snapshot(page_with_fonts(b, &fonts)); + let cache = FontCache::new(); + let mut budget = DecodeBudget::new(32 << 20); + let refusals: Vec> = ids + .iter() + .map(|id| { + let dict = snap + .doc + .get_object((*id, 0)) + .and_then(Object::as_dict) + .expect("font"); + cache + .get_or_load(&snap.doc, FontKey::Indirect((*id, 0)), dict, &mut budget) + .refusal + }) + .collect(); + assert_eq!(refusals[0], None, "FONT-14c one such font fits"); + assert_eq!( + refusals[4], + Some(TextReason::PageTooComplex), + "FONT-14c the page budget bounds the models' memory: {refusals:?}" + ); +} + +#[test] +fn font_28c_descriptor_null_and_charset_and_dw() { + // /FontDescriptor null (or a reference to nothing) is an absent descriptor. + for extra in ["/FontDescriptor null", "/FontDescriptor 9999 0 R"] { + let mut f = std14("Helvetica", "/WinAnsiEncoding"); + f.font_extra = extra.into(); + let m = load_simple(&f); + assert_eq!(m.refusal, None, "FONT-28c {extra}"); + assert_eq!(typeable(&m, c1(b'A')), Some('A')); + } + // /CharSet present but not a string: no Type1 glyph is proven. + let file = Type1Builder::new("CS") + .glyph("A", T1Glyph::Box { width: 600 }) + .build(); + let mut f = type1_simple("ABCDEF+CS", file); + assert_eq!(alphabet(&load_simple(&f)), "A"); + f.descriptor_extra = "/CharSet 5".into(); + assert_eq!( + alphabet(&load_simple(&f)), + "", + "FONT-28c malformed /CharSet" + ); + // /DW present but not a number. + let (program, _) = latin_truetype("A", ""); + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+DW"); + f.program = Program::TrueType(program); + f.tounicode = Some(tounicode_bfchar(&[(1, 2, "A")])); + assert_eq!(load_type0(&f).refusal, None); + f.cid_extra = "/DW /Wide".into(); + assert_eq!( + load_type0(&f).refusal, + Some(TextReason::FontUnsupported), + "FONT-28c malformed /DW" + ); +} + +#[test] +fn font_28d_structural_refusals_report_the_lowest_code() { + // A FontFile3 that does not match the font type (FONT_PROGRAM_UNSUPPORTED) and malformed + // /Widths (FONT_UNSUPPORTED, earlier in §A.10): the earlier code is reported. + let mut f = SimpleFont::new("TrueType", "ABCDEF+Both"); + f.flags = Some(32); + f.program = Program::Raw { + key: "FontFile3", + subtype: Some("Type1C"), + data: vec![1, 0, 4, 4], + }; + f.font_extra = "/FirstChar 32 /LastChar 40 /Widths [500]".into(); + assert_eq!(load_simple(&f).refusal, Some(TextReason::FontUnsupported)); + f.font_extra = "/FirstChar 32 /LastChar 32 /Widths [500]".into(); + assert_eq!( + load_simple(&f).refusal, + Some(TextReason::FontProgramUnsupported) + ); +} + +#[test] +fn font_28e_type3_hash_decodes_a_bounded_amount() { + let type3 = |b: &mut PdfBuilder, procs: &[Vec]| -> u32 { + let ids: Vec = procs.iter().map(|p| b.add_flate("", p)).collect(); + b.add(format!( + "<< /Type /Font /Subtype /Type3 /FontBBox [0 0 1000 1000] \ + /FontMatrix [0.001 0 0 0.001 0 0] /CharProcs << /a {} 0 R /b {} 0 R /c {} 0 R >> \ + /Encoding << /Type /Encoding /Differences [97 /a /b /c] >> /FirstChar 97 \ + /LastChar 99 /Widths [500 500 500] /Resources << >> >>", + ids[0], ids[1], ids[2] + )) + }; + let big = vec![b' '; 3 << 20]; + let mut b = PdfBuilder::new(); + let id = type3(&mut b, &[big.clone(), big.clone(), big]); + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + let m = load_with(b, id, &mut budget); + assert_eq!( + m.refusal, + Some(TextReason::Type3), + "FONT-28e still a run-level refusal" + ); + let spent = PAGE_DECODE_BUDGET - budget.remaining(); + assert!( + spent < 5 << 20, + "FONT-28e ≤ 4 MiB of hash-only decoding ({spent} bytes)" + ); + // Small glyph procedures still hash by content, the same across renumbered files. + let procs = |c: &str| vec![c.as_bytes().to_vec(), b"0 0 m".to_vec(), b"1 1 l".to_vec()]; + let hash = |pad: usize, c: &str| { + let mut b = PdfBuilder::new(); + for _ in 0..pad { + b.add("<< >>"); + } + let id = type3(&mut b, &procs(c)); + load_with(b, id, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)).content_hash + }; + assert_eq!( + hash(0, "0 0 m 5 5 l f"), + hash(3, "0 0 m 5 5 l f"), + "FONT-28e identity" + ); + assert_ne!( + hash(0, "0 0 m 5 5 l f"), + hash(0, "0 0 m 6 6 l f"), + "FONT-28e content" + ); +} + +#[test] +fn font_12b_bidi_isolates_and_deprecated_controls_are_excluded() { + for cp in 0x2060..=0x206F { + let ch = char::from_u32(cp).expect("BMP"); + assert!(!writable_char(ch), "FONT-12b U+{cp:04X} never typeable"); + assert!( + reading_reason(ch).is_some(), + "FONT-12b U+{cp:04X} not readable" + ); + } + assert!( + writable_char('\u{205F}'), + "FONT-12b the medium space before them stays" + ); +} + +#[test] +fn font_15b_annex_d_duplicates_agree_with_tounicode() { + // Word writes WinAnsi 0xA0 (`space`) as U+00A0 and 0xAD (`hyphen`) as U+00AD. + let mut f = winansi_truetype("ABCDEF+Word", "A-", " "); + f.tounicode = Some(tounicode_bfchar(&[ + (0xA0, 1, "\u{a0}"), + (0xAD, 1, "\u{ad}"), + (0x41, 1, "\u{a0}"), + (0x20, 1, "\u{a0}"), + ])); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(m.text(c1(0xA0)), Some("\u{a0}"), "FONT-15b no-break space"); + assert_eq!(typeable(&m, c1(0xA0)), Some('\u{a0}')); + assert_eq!(m.text(c1(0xAD)), Some("\u{ad}"), "FONT-15b soft hyphen"); + assert_eq!( + m.text(c1(0x41)), + None, + "FONT-15b only the Annex D pairs agree" + ); + assert_eq!( + m.text(c1(0x20)), + None, + "FONT-15b `space` at 0x20 is U+0020, not U+00A0 (only the Annex D codes)" + ); +} + +#[test] +fn font_29g_cid_keyed_cff_local_subroutines_through_fdselect() { + let local_call = |j: i32| [t2(j - 107), vec![10]].concat(); + let glyph = |middle: Vec| [t2(100), t2(0), vec![21], middle, vec![14]].concat(); + // Local subr 0 draws a line; 1–9 each call the next one 8 times; 10 returns. + let mut subrs = vec![[t2(400), t2(0), vec![5, 11]].concat()]; + for k in 1..10 { + let mut s: Vec = (0..8).flat_map(|_| local_call(k + 1)).collect(); + s.push(11); + subrs.push(s); + } + subrs.push(vec![11]); + for format in [0u8, 3] { + let data = CffBuilder { + subrs: subrs.clone(), + fdselect_format: format, + ..CffBuilder::new("CIDSubrs") + } + .raw_cid_glyph(10, glyph(local_call(0))) + .raw_cid_glyph(11, glyph([local_call(1), local_call(0)].concat())) + .raw_cid_glyph(12, glyph([t2(400), t2(0), vec![5]].concat())) + .build(); + let outlines = Outlines::of_cff(&data).expect("CFF"); + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + let (drawn, secs) = cpu(|| { + (1..=3) + .map(|g| outlines.drawn(g, &mut meter)) + .collect::>() + }); + assert_eq!( + drawn, + [true, false, true], + "FONT-29g FDSelect format {format}" + ); + assert!(secs < 1.0, "FONT-29g bounded ({secs:.2} s of CPU)"); + let table = ttf_parser::cff::Table::parse(&data).expect("CFF"); + for gid in [1u16, 3] { + let mut segs = Segs::default(); + assert!(table.outline(GlyphId(gid), &mut segs).is_ok() && segs.0 > 0); + } + let mut f = Type0Font::new("CIDFontType0", "ABCDEF+CIDSubrs"); + f.program = Program::CidCff(data); + f.tounicode = Some(tounicode_bfchar(&[ + (10, 2, "A"), + (11, 2, "B"), + (12, 2, "C"), + ])); + let m = load_type0(&f); + assert_eq!(m.refusal, None); + assert_eq!(alphabet(&m), "AC", "FONT-29g through a Type0 font"); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/tests_classes.rs b/src-tauri/src/pdf_engine/text_edit/fonts/tests_classes.rs new file mode 100644 index 0000000..439abb0 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/tests_classes.rs @@ -0,0 +1,773 @@ +//! Font classes and model surfaces: Type0 widths (FONT-20), Standard-14 and other non-embedded +//! fonts (FONT-25…28), faces (FONT-31), sibling surfaces (FONT-32), code preference (FONT-33), +//! plus refusals by class, identity hashing and the cache. + +use super::faces::{family_key, is_bold, is_italic, words_of, FaceHints}; +use super::std14::{std14_match, Std14Face, Std14Match}; +use super::tests::{alphabet, c1, c2, std14, typeable, winansi_truetype}; +use super::{face_surface, typing_surface, FamilyHint, FontCache, FontClass, FontKey, FontModel}; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::reasons::{Face, TextReason}; +use crate::pdf_engine::text_edit::testkit::fonts::{ + add_simple, add_type0, latin_truetype, load_page_fonts, load_simple, load_type0, + page_with_fonts, snapshot, tounicode_bfchar, Program, SimpleFont, Type0Font, +}; +use crate::pdf_engine::text_edit::testkit::pdf::PdfBuilder; +use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; +use std::sync::Arc; + +fn cid_font_with_widths(w: Option<&str>, dw: Option) -> Type0Font { + let mut t = TtfBuilder::new(); + for i in 1..=12 { + t.glyph(&format!("g{i}"), true, 500); + } + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Test"); + f.program = Program::TrueType(t.build()); + f.w = w.map(str::to_string); + f.dw = dw; + f.tounicode = Some(tounicode_bfchar(&[ + (1, 2, "A"), + (2, 2, "B"), + (11, 2, "C"), + (5, 2, "D"), + ])); + f +} + +#[test] +fn font_20_type0_w_both_forms_and_dw() { + let m = load_type0(&cid_font_with_widths( + Some("[1 [500 600] 10 12 700]"), + Some(300.0), + )); + assert_eq!(m.refusal, None); + assert_eq!(m.width(c2(1)), 500.0, "FONT-20 c [w1 w2] form"); + assert_eq!(m.width(c2(2)), 600.0); + assert_eq!( + m.width(c2(11)), + 700.0, + "FONT-20 cfirst clast w form (mapped code)" + ); + assert_eq!( + m.width(c2(12)), + 700.0, + "FONT-20 unmapped code inside a range" + ); + assert_eq!(m.width(c2(5)), 300.0, "FONT-20 /DW"); + assert_eq!(m.width(c2(40)), 300.0); + assert_eq!( + m.width(c1(1)), + 0.0, + "FONT-20 a 1-byte code in a 2-byte font" + ); + assert_eq!(alphabet(&m), "ABCD"); + let m = load_type0(&cid_font_with_widths(None, None)); + assert_eq!(m.width(c2(1)), 1000.0, "FONT-20 default DW 1000"); + // Later entries win. + let m = load_type0(&cid_font_with_widths(Some("[1 2 400 2 [650]]"), None)); + assert_eq!((m.width(c2(1)), m.width(c2(2))), (400.0, 650.0)); + for bad in [ + "[1 (x)]", + "[1 70000 500]", + "[0 65535 500 0 65535 500]", + "[1 2]", + "[(a) [1]]", + "5", + ] { + let m = load_type0(&cid_font_with_widths(Some(bad), None)); + assert_eq!( + m.refusal, + Some(TextReason::FontUnsupported), + "FONT-20 /W {bad}" + ); + } +} + +#[test] +fn font_20b_real_full_truetype_as_cid_font_type2() { + let noto = std::fs::read( + std::path::Path::new(env!("CARGO_MANIFEST_DIR")) + .join("resources/fonts/NotoSans-Regular.ttf"), + ) + .expect("NotoSans-Regular.ttf in the repo"); + let face = ttf_parser::Face::parse(¬o, 0).expect("Noto parses"); + let text = "Ağış İ"; + let entries: Vec<(u32, usize, String)> = text + .chars() + .map(|ch| { + ( + u32::from(face.glyph_index(ch).expect("glyph").0), + 2, + ch.to_string(), + ) + }) + .collect(); + let refs: Vec<(u32, usize, &str)> = entries + .iter() + .map(|(c, l, t)| (*c, *l, t.as_str())) + .collect(); + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+NotoSans-Regular"); + f.program = Program::TrueType(noto.clone()); + f.dw = Some(600.0); + f.cid_to_gid = Some(None); + f.tounicode = Some(tounicode_bfchar(&refs)); + let m = load_type0(&f); + assert_eq!(m.refusal, None); + assert_eq!( + alphabet(&m), + " AğİIş".replace('I', "ı"), + "FONT-20b every char of {text}" + ); + assert_eq!( + m.ascent, 0.8, + "FONT-20b the descriptor's /Ascent wins over hhea" + ); +} + +#[test] +fn font_25_std14_widths_per_face() { + let times = load_simple(&std14("Times-Roman", "/WinAnsiEncoding")); + let helv = load_simple(&std14("Helvetica", "/WinAnsiEncoding")); + let times_bold = load_simple(&std14("Times-Bold", "/WinAnsiEncoding")); + let courier = load_simple(&std14("Courier", "/WinAnsiEncoding")); + assert_eq!(times.class, Some(FontClass::Std14(Std14Face::TimesRoman))); + assert_eq!(times.width(c1(b'a')), 444.0, "FONT-25 Times a"); + assert_eq!( + helv.width(c1(b'a')), + 556.0, + "FONT-25 Helvetica a (not Times)" + ); + assert_eq!(times_bold.width(c1(b'a')), 500.0, "FONT-25 Times-Bold a"); + assert_eq!(courier.width(c1(b'W')), 600.0); + assert!(!times.substituted && !times.embedded); + assert_eq!((times.ascent, times.descent), (0.683, -0.217)); + assert_eq!(times.family_hint, FamilyHint::Serif); + assert_eq!(courier.family_hint, FamilyHint::Mono); + assert!(times_bold.bold && !times_bold.italic); + assert!("AZaz09".chars().all(|ch| alphabet(×).contains(ch))); + let m = load_simple(&std14( + "Times-Roman", + "<< /Differences [128 /gbreve /Gamma] >>", + )); + assert_eq!( + typeable(&m, c1(128)), + Some('ğ'), + "FONT-25 gbreve is in the Core-14 AFM" + ); + assert_eq!(m.text(c1(129)), Some("Γ")); + assert!( + !m.drawable(c1(129)), + "FONT-25 Gamma is not in Times' glyph set" + ); + // /Widths, when present, win over the AFM. + let mut f = std14("Times-Roman", "/WinAnsiEncoding"); + f.first_char = 32; + f.widths = Some(vec![300.0; 224]); + assert_eq!(load_simple(&f).width(c1(b'a')), 300.0); + // No /Encoding: StandardEncoding. + let m = load_simple(&SimpleFont::new("Type1", "Helvetica")); + assert_eq!(m.text(c1(0x27)), Some("’")); + assert_eq!(m.text(c1(0xE8)), Some("Ł")); +} + +#[test] +fn font_26_std14_aliases() { + for (name, face) in [ + ("Arial", Std14Face::Helvetica), + ("ArialMT", Std14Face::Helvetica), + ("Arial,Bold", Std14Face::HelveticaBold), + ("Arial-BoldMT", Std14Face::HelveticaBold), + ("Arial,Italic", Std14Face::HelveticaOblique), + ("Arial-ItalicMT", Std14Face::HelveticaOblique), + ("Arial,BoldItalic", Std14Face::HelveticaBoldOblique), + ("Arial-BoldItalicMT", Std14Face::HelveticaBoldOblique), + ("TimesNewRoman", Std14Face::TimesRoman), + ("TimesNewRomanPS", Std14Face::TimesRoman), + ("TimesNewRomanPSMT", Std14Face::TimesRoman), + ("TimesNewRomanPS-BoldMT", Std14Face::TimesBold), + ("TimesNewRomanPS-ItalicMT", Std14Face::TimesItalic), + ("TimesNewRomanPS-BoldItalicMT", Std14Face::TimesBoldItalic), + ("TimesNewRoman,Bold", Std14Face::TimesBold), + ("TimesNewRoman,Italic", Std14Face::TimesItalic), + ("TimesNewRoman,BoldItalic", Std14Face::TimesBoldItalic), + ("CourierNew", Std14Face::Courier), + ("CourierNewPSMT", Std14Face::Courier), + ("CourierNewPS-BoldMT", Std14Face::CourierBold), + ("CourierNewPS-ItalicMT", Std14Face::CourierOblique), + ("CourierNewPS-BoldItalicMT", Std14Face::CourierBoldOblique), + ("CourierNew,Bold", Std14Face::CourierBold), + ("CourierNew,Italic", Std14Face::CourierOblique), + ("CourierNew,BoldItalic", Std14Face::CourierBoldOblique), + ("Times New Roman", Std14Face::TimesRoman), + ] { + assert_eq!( + std14_match(name), + Some(Std14Match::Latin(face)), + "FONT-26 {name}" + ); + } + assert_eq!( + std14_match("ArialNarrow"), + None, + "FONT-26 Arial Narrow's metrics differ" + ); + let m = load_simple(&std14("Arial,Bold", "/WinAnsiEncoding")); + assert_eq!(m.class, Some(FontClass::Std14(Std14Face::HelveticaBold))); + assert_eq!(m.width(c1(b'a')), 556.0); + assert!(m.bold); + let m = load_simple(&std14("CourierNewPSMT", "/WinAnsiEncoding")); + assert_eq!(m.width(c1(b'i')), 600.0, "FONT-26 Courier widths"); + let m = load_simple(&std14("ArialNarrow", "/WinAnsiEncoding")); + assert_eq!( + m.refusal, + Some(TextReason::MissingWidths), + "FONT-26 not Standard-14" + ); +} + +#[test] +fn font_27_non_embedded_non_std14() { + let mut f = SimpleFont::new("TrueType", "Calibri"); + f.encoding = + Some("<< /BaseEncoding /WinAnsiEncoding /Differences [128 /gbreve /uni4E00 /x] >>".into()); + f.first_char = 32; + let mut widths = vec![500.0; 224]; + widths[usize::from(b'x' - 32)] = 0.0; + widths[130 - 32] = 0.0; + f.widths = Some(widths); + f.flags = Some(32); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(m.class, Some(FontClass::SimpleNonEmbedded)); + assert!(m.substituted && !m.embedded, "FONT-27 flagged substituted"); + assert_eq!(typeable(&m, c1(b'A')), Some('A')); + assert_eq!(typeable(&m, c1(128)), Some('ğ'), "FONT-27 Latin Extended-A"); + assert_eq!(m.text(c1(129)), Some("一"), "FONT-27 CJK reads"); + assert_eq!( + typeable(&m, c1(129)), + None, + "FONT-27 outside the Latin/Greek/Cyrillic list" + ); + assert_eq!(typeable(&m, c1(b'x')), None, "FONT-27 width 0"); + f.widths = None; + assert_eq!( + load_simple(&f).refusal, + Some(TextReason::MissingWidths), + "FONT-27 no /Widths" + ); + f.widths = Some(vec![500.0; 224]); + f.flags = Some(4); + assert_eq!( + load_simple(&f).refusal, + Some(TextReason::UnsupportedEncoding), + "FONT-27 symbolic" + ); +} + +#[test] +fn font_28_symbol_and_zapf_dingbats_are_unsupported_encodings() { + for name in ["Symbol", "ZapfDingbats", "Symbol,Bold"] { + let m = load_simple(&SimpleFont::new("Type1", name)); + assert_eq!( + m.refusal, + Some(TextReason::UnsupportedEncoding), + "FONT-28 {name}" + ); + assert_eq!(m.class, None); + } +} + +#[test] +fn font_28b_class_refusals() { + let mut t3 = SimpleFont::new("Type3", "T3"); + t3.encoding = Some("<< /Differences [65 /A /B] >>".into()); + t3.first_char = 65; + t3.widths = Some(vec![700.0, 0.6]); + t3.font_extra = "/FontMatrix [0.001 0 0 0.001 0 0] /FontBBox [0 0 1000 1000]".into(); + let m = load_simple(&t3); + assert_eq!(m.refusal, Some(TextReason::Type3)); + assert_eq!( + m.width(c1(b'A')), + 700.0, + "Type3 widths scale by /FontMatrix" + ); + assert_eq!(m.text(c1(b'A')), Some("A"), "Type3 codes still read"); + assert_eq!(alphabet(&m), ""); + t3.font_extra = "/FontMatrix [1 0 0 1 0 0]".into(); + assert!((load_simple(&t3).width(c1(b'B')) - 600.0).abs() < 1e-3); + let m = load_simple(&SimpleFont::new("MMType1", "Minion_MM")); + assert_eq!(m.refusal, Some(TextReason::FontUnsupported), "MMType1"); + let m = load_simple(&SimpleFont::new("OpenType", "X")); + assert_eq!( + m.refusal, + Some(TextReason::FontUnsupported), + "unknown subtype" + ); + let mut f = std14("Helvetica", "/WinAnsiEncoding"); + f.first_char = 32; + f.widths = Some(vec![500.0; 10]); + f.font_extra = "/LastChar 50".into(); + assert_eq!( + load_simple(&f).refusal, + Some(TextReason::FontUnsupported), + "/Widths length" + ); + f.widths = Some(vec![500.0]); + f.font_extra.clear(); + f.descriptor_extra = "/Flags (x)".into(); + assert_eq!( + load_simple(&f).refusal, + Some(TextReason::FontUnsupported), + "/Flags type" + ); + let base = cid_font_with_widths(None, None); + for (encoding, reason) in [ + ("/Identity-V", TextReason::Vertical), + ("/90ms-RKSJ-V", TextReason::Vertical), + ("/UniJIS-UCS2-H", TextReason::UnsupportedEncoding), + ("(x)", TextReason::FontUnsupported), + ] { + let mut f = base.clone(); + f.encoding = encoding.into(); + let m = load_type0(&f); + assert_eq!(m.refusal, Some(reason), "Type0 /Encoding {encoding}"); + assert_eq!(m.vertical, reason == TextReason::Vertical); + } + let mut f = base.clone(); + f.cid_extra = "/WMode 1".into(); + assert_eq!( + load_type0(&f).refusal, + Some(TextReason::Vertical), + "CIDFont /WMode 1" + ); + let mut f = base.clone(); + f.tounicode = None; + assert_eq!( + load_type0(&f).refusal, + Some(TextReason::NoTounicode), + "Type0 without ToUnicode" + ); + let m = load_type0(&base); + assert_eq!( + m.split_codes(&[0, 1, 0]), + Err(TextReason::AmbiguousUnicode), + "odd 2-byte string" + ); + assert_eq!(m.split_codes(&[0, 1, 0, 2]), Ok(vec![c2(1), c2(2)])); + assert!(!m.is_word_space(c2(32))); + let simple = load_simple(&std14("Helvetica", "/WinAnsiEncoding")); + assert_eq!(simple.split_codes(b"a "), Ok(vec![c1(b'a'), c1(b' ')])); + assert!(simple.is_word_space(c1(32))); + assert_eq!(simple.code_bytes(c1(0x41)), vec![0x41]); +} + +#[test] +fn font_31_faces() { + let pairs = [ + ("TimesNewRomanPS-BoldMT", "TimesNewRomanPSMT"), + ("Arial-BoldMT", "ArialMT"), + ("Arial-ItalicMT", "ArialMT"), + ("BAAAAA+Times New Roman-Bold", "CAAAAA+Times New Roman"), + ("NZDDWO+Liberation-Sans-Bold", "VRMYHF+Liberation-Sans"), + ("OSDNAA+Liberation-Sans-Italic", "VRMYHF+Liberation-Sans"), + ("PHVFWU+Carlito-Bold", "GICQRP+Carlito"), + ("BCDEEE+Aptos-Bold", "BCDFEE+Aptos"), + ("AAAAAB+SourceSansPro-Bold", "AAAAAD+SourceSansPro-Regular"), + ("MGAFOB+Helvetica-Oblique", "AHLTJX+Helvetica"), + ("Helvetica-Bold", "Helvetica"), + ("NILLFH+DejaVu-Sans-Mono-Bold", "IKCJUC+DejaVu-Sans-Mono"), + ]; + for (styled, plain) in pairs { + assert_eq!( + family_key(styled), + family_key(plain), + "FONT-31 {styled} ~ {plain}" + ); + } + assert_ne!(family_key("Arial-BoldMT"), family_key("TimesNewRomanPSMT")); + assert_ne!(family_key("Carlito-Bold"), family_key("Caladea-Bold")); + assert_eq!(family_key("ABCDEF+"), "", "FONT-31 a bare subset tag"); + assert_eq!( + words_of("TimesNewRomanPS-BoldMT"), + ["times", "new", "roman", "ps", "bold", "mt"] + ); + let none = FaceHints::default(); + assert!(is_bold("TimesNewRomanPS-BoldMT", &none)); + assert!(!is_bold("TimesNewRomanPSMT", &none)); + assert!(is_italic("MGAFOB+Helvetica-Oblique", &none)); + assert!( + is_bold("WIVRPB+Arial,Bold", &none), + "FONT-31 comma separator" + ); + assert!( + !is_bold("Boldoni", &none) && !is_bold("ABCDEF+Boldoni", &none), + "FONT-31 Boldoni" + ); + let flags = |f| FaceHints { + flags: Some(f), + ..none + }; + let stem = |s| FaceHints { + stem_v: Some(s), + ..none + }; + let angle = |a| FaceHints { + italic_angle: Some(a), + ..none + }; + assert!( + is_bold("VULAQV+SegoeUI", &flags(1 << 18)), + "FONT-31 ForceBold" + ); + assert!(is_bold("VULAQV+SegoeUI", &stem(165.0)), "FONT-31 StemV"); + assert!( + is_italic("VULAQV+SegoeUI", &angle(-12.0)), + "FONT-31 ItalicAngle" + ); + assert!(!is_italic("VULAQV+SegoeUI", &angle(0.0))); + assert!( + is_italic("VULAQV+SegoeUI", &flags(1 << 6)), + "FONT-31 Italic flag" + ); + let thick_italic = FaceHints { + stem_v: Some(140.0), + italic_angle: Some(-15.0), + ..none + }; + assert!(is_italic("IPRNUP+TimesNewRomanPS-ItalicMT", &thick_italic)); + assert!( + !is_bold("IPRNUP+TimesNewRomanPS-ItalicMT", &thick_italic), + "FONT-31 italic stem" + ); + assert!( + !is_bold("AAAAAD+SourceSansPro-Regular", &stem(200.0)), + "FONT-31 Regular wins" + ); + assert!(!is_italic("AAAAAD+SourceSansPro-Regular", &angle(-12.0))); + assert!( + is_bold("Liberation-Sans-Bold-Italic", &none) + && is_italic("Liberation-Sans-Bold-Italic", &none) + ); + let bi = FaceHints { + stem_v: Some(160.0), + italic_angle: Some(-12.0), + ..none + }; + assert!(is_bold("AAAAAA+Liberation-Sans-BoldItalic", &bi)); + assert!(is_italic("AAAAAA+Liberation-Sans-BoldItalic", &bi)); + // Through a loaded model: descriptor flags drive the family hint and faces. + let mut f = winansi_truetype("ABCDEF+SegoeUI", "A", ""); + f.flags = Some(32 | 1 << 18 | 2); + let m = load_simple(&f); + assert!(m.bold && !m.italic); + assert_eq!(m.family_hint, FamilyHint::Serif); + assert_eq!(m.display_name, "Segoe UI"); + assert_eq!(m.base_name, "SegoeUI"); +} + +/// A Type0 CIDFontType2 named `base` typing `chars` (GID = CID = 1…). +fn word_companion(base: &str, chars: &str) -> Type0Font { + let (program, gids) = latin_truetype(chars, ""); + let entries: Vec<(u32, usize, String)> = gids + .iter() + .map(|(ch, gid)| (u32::from(*gid), 2, ch.to_string())) + .collect(); + let refs: Vec<(u32, usize, &str)> = entries + .iter() + .map(|(c, l, t)| (*c, *l, t.as_str())) + .collect(); + let mut f = Type0Font::new("CIDFontType2", base); + f.flags = Some(32); + f.program = Program::TrueType(program); + f.tounicode = Some(tounicode_bfchar(&refs)); + f +} + +fn page(fonts: Vec<(&str, FontSpec)>) -> Vec<(Vec, Arc)> { + let mut b = PdfBuilder::new(); + let ids: Vec<(&str, u32)> = fonts + .iter() + .map(|(name, spec)| { + let id = match spec { + FontSpec::Simple(f) => add_simple(&mut b, f), + FontSpec::Type0(f) => add_type0(&mut b, f), + }; + (*name, id) + }) + .collect(); + load_page_fonts(&snapshot(page_with_fonts(b, &ids))) +} + +enum FontSpec { + Simple(SimpleFont), + Type0(Type0Font), +} + +fn names(surface: &super::TypingSurface) -> Vec { + surface + .fonts + .iter() + .map(|(n, _)| String::from_utf8_lossy(n).into_owned()) + .collect() +} + +#[test] +fn font_32_sibling_groups_word_pattern() { + let mut refused = winansi_truetype("BCDEFG+Calibri", "Sa", ""); + refused.widths = None; + let fonts = page(vec![ + ( + "F1", + FontSpec::Simple(winansi_truetype("ABCDEF+Calibri", "Salk B", "")), + ), + ( + "F2", + FontSpec::Type0(word_companion("ABCDEF+Calibri", "ğış")), + ), + ( + "F3", + FontSpec::Simple(winansi_truetype("ABCDEF+Calibri-Bold", "Sa", "")), + ), + ( + "F4", + FontSpec::Type0(word_companion("ABCDEF+Calibri-Bold", "ğ")), + ), + ( + "F5", + FontSpec::Simple(winansi_truetype("ABCDEF+Cambria", "Sa", "")), + ), + ("F6", FontSpec::Simple(refused)), + ( + "F7", + FontSpec::Type0(word_companion("ABCDEF+Calibri-Light", "ğ")), + ), + ]); + assert_eq!(fonts[5].1.refusal, Some(TextReason::MissingWidths)); + let surface = typing_surface(&fonts, b"F1"); + assert_eq!( + names(&surface), + ["F1", "F2"], + "FONT-32 Word pair, no bold/other family/refused" + ); + assert_eq!( + surface.alphabet().into_iter().collect::(), + " BSaklğış" + ); + assert_eq!(surface.writer_for('a'), Some((0, c1(b'a')))); + let (idx, code) = surface.writer_for('ğ').expect("ğ"); + assert_eq!( + (idx, code.len), + (1, 2), + "FONT-32 ğ from the Type0 companion" + ); + assert!(surface.has_space()); + assert_eq!( + names(&typing_surface(&fonts, b"F2")), + ["F2", "F1"], + "FONT-32 Type0 primary" + ); + assert!( + typing_surface(&fonts, b"F9").fonts.is_empty(), + "unknown primary" + ); + assert_eq!( + names(&typing_surface(&fonts, b"F6")), + ["F6"], + "refused primary types alone" + ); + let bold = face_surface(&fonts, &fonts[0].1, Face::Bold).expect("FONT-32 bold group"); + assert_eq!( + names(&bold), + ["F3", "F4"], + "FONT-32 the bold group, led by the simple font" + ); + assert!( + face_surface(&fonts, &fonts[0].1, Face::Italic).is_none(), + "no italic face" + ); + let regular = face_surface(&fonts, &fonts[2].1, Face::Regular).expect("regular"); + assert_eq!(names(®ular), ["F1", "F2"]); + // A simple font is never swapped for a lone Type0 face of another name. + let fonts = page(vec![ + ( + "F1", + FontSpec::Simple(winansi_truetype("ABCDEF+Aptos", "Sa", "")), + ), + ( + "F2", + FontSpec::Type0(word_companion("ABCDEF+Aptos-Bold", "ğ")), + ), + ]); + assert!( + face_surface(&fonts, &fonts[0].1, Face::Bold).is_none(), + "FONT-32 no swap" + ); + assert_eq!(names(&typing_surface(&fonts, b"F1")), ["F1"]); +} + +#[test] +fn font_33_encode_preference_order() { + let fonts = page(vec![ + ( + "F1", + FontSpec::Simple(std14("Helvetica", "<< /Differences [65 /A 200 /A] >>")), + ), + ( + "F2", + FontSpec::Simple(std14("Helvetica-Bold", "/WinAnsiEncoding")), + ), + ( + "F3", + FontSpec::Simple(std14("Helvetica", "<< /Differences [66 /A] >>")), + ), + ]); + let m = &fonts[0].1; + assert_eq!( + m.code_for('A', &[], &[]), + Some(c1(65)), + "FONT-33 lowest code" + ); + assert_eq!( + m.code_for('A', &[], &[('A', c1(200))]), + Some(c1(200)), + "FONT-33 page code" + ); + assert_eq!( + m.code_for('A', &[('A', c1(65))], &[('A', c1(200))]), + Some(c1(65)), + "FONT-33 run code beats page code" + ); + assert_eq!( + m.code_for('A', &[('A', c1(66))], &[]), + Some(c1(65)), + "FONT-33 invalid preference" + ); + assert_eq!( + m.code_for('A', &[('B', c1(200))], &[]), + Some(c1(65)), + "other char" + ); + assert_eq!(m.code_for('ğ', &[], &[]), None); + let surface = typing_surface(&fonts, b"F1"); + assert_eq!(names(&surface), ["F1", "F3"], "FONT-33 same face only"); + assert_eq!(surface.writer_for('A'), Some((0, c1(65)))); + assert_eq!( + surface.writer_for_with('A', &[(1, 'A', c1(66))], &[]), + Some((1, c1(66))) + ); + assert_eq!( + surface.writer_for_with('A', &[], &[(0, 'A', c1(200))]), + Some((0, c1(200))), + "FONT-33 page preference" + ); + assert_eq!( + surface.writer_for_with('A', &[(1, 'A', c1(200))], &[]), + Some((0, c1(65))), + "FONT-33 a preference the font cannot honour is ignored" + ); +} + +#[test] +fn font_identity_hash_cache_and_budget() { + let font = |padding: usize, flate: bool, width: f64| { + let mut b = PdfBuilder::new(); + for _ in 0..padding { + b.add("<< /Padding true >>"); + } + let mut f = winansi_truetype("ABCDEF+Calibri", "AB", ""); + f.flate = flate; + f.widths = Some(vec![width; 224]); + let id = add_simple(&mut b, &f); + let fonts = load_page_fonts(&snapshot(page_with_fonts(b, &[("F1", id)]))); + fonts[0].1.content_hash + }; + let base = font(0, true, 500.0); + assert_eq!( + base, + font(7, true, 500.0), + "renumbered objects hash alike (D32)" + ); + assert_eq!(base, font(0, false, 500.0), "streams hash by decoded bytes"); + assert_ne!( + base, + font(0, true, 501.0), + "a changed width changes the hash" + ); + // The cache returns the same model for the same key; a budget-starved load is not cached. + let mut b = PdfBuilder::new(); + let id = add_simple(&mut b, &winansi_truetype("ABCDEF+Calibri", "AB", "")); + let snap = snapshot(page_with_fonts(b, &[("F1", id)])); + let font_id = snap + .doc + .objects + .keys() + .copied() + .find(|k| k.0 == id) + .expect("id"); + let dict = snap + .doc + .get_object(font_id) + .and_then(lopdf::Object::as_dict) + .expect("dict"); + let cache = FontCache::new(); + let starved = cache.get_or_load( + &snap.doc, + FontKey::Indirect(font_id), + dict, + &mut DecodeBudget::new(100), + ); + assert_eq!( + starved.refusal, + Some(TextReason::PageTooComplex), + "budget exhausted" + ); + let mut budget = DecodeBudget::new(1 << 20); + let first = cache.get_or_load(&snap.doc, FontKey::Indirect(font_id), dict, &mut budget); + assert_eq!(first.refusal, None, "the starved model was not cached"); + assert_eq!(first.key, FontKey::Indirect(font_id)); + let used = (1 << 20) - budget.remaining(); + assert!(used > 0); + let again = cache.get_or_load(&snap.doc, FontKey::Indirect(font_id), dict, &mut budget); + assert!(Arc::ptr_eq(&first, &again), "cached"); + assert_eq!( + (1 << 20) - budget.remaining(), + used, + "a cached model decodes nothing" + ); + let direct = FontKey::direct(font_id, b"F1"); + assert_ne!(direct, FontKey::direct(font_id, b"F2")); + let other = cache.get_or_load(&snap.doc, direct, dict, &mut budget); + assert!(!Arc::ptr_eq(&first, &other), "keys are distinct entries"); + assert_eq!(first.content_hash, other.content_hash); +} + +#[test] +fn family_hint_reads_the_name_when_the_flags_say_nothing() { + // Live check B3: LibreOffice writes `/Flags 4` (Symbolic only), so every Liberation Serif + // line was "sans" and the editor set it in Arial. The FixedPitch and Serif flags still win. + use super::faces::family_hint; + let cases: [(&str, Option, FamilyHint); 12] = [ + ("BAAAAA+LiberationSerif", Some(4), FamilyHint::Serif), + ("CAAAAA+LiberationSerif-Bold", Some(4), FamilyHint::Serif), + ("DAAAAA+LiberationSans", Some(4), FamilyHint::Sans), + ("LiberationMono", Some(4), FamilyHint::Mono), + ("DejaVuSansMono", None, FamilyHint::Mono), + ("TimesNewRomanPSMT", Some(32), FamilyHint::Serif), + ("Georgia,Italic", None, FamilyHint::Serif), + ("Microsoft Sans Serif", None, FamilyHint::Sans), + ("CenturyGothic", None, FamilyHint::Sans), + ("Consolas", Some(32), FamilyHint::Mono), + ("Arial", Some(2), FamilyHint::Serif), + ("ABCDEF+Lato-Regular", Some(1), FamilyHint::Mono), + ]; + for (name, flags, want) in cases { + assert_eq!(family_hint(name, flags), want, "{name} {flags:?}"); + } + let mut f = winansi_truetype("BAAAAA+LiberationSerif", "A", ""); + f.flags = Some(4); + assert_eq!( + load_simple(&f).family_hint, + FamilyHint::Serif, + "through a load" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/tests_fuzz.rs b/src-tauri/src/pdf_engine/text_edit/fonts/tests_fuzz.rs new file mode 100644 index 0000000..f6ec580 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/tests_fuzz.rs @@ -0,0 +1,533 @@ +//! FONT-34 loader fuzz (§B.21): 1,000 deterministic xorshift mutations — byte flips, +//! truncations, inserted PostScript/PDF tokens, huge numbers, duplicated slices — over +//! ToUnicode, Type1, TrueType and CFF inputs, fed to the parsers directly and through the whole +//! loader (`FontCache::get_or_load` on an in-memory document); FONT-34b mutates the font +//! dictionaries themselves (`/W`, `/DW`, `/Widths`, `/Differences`, `/Encoding`, descriptor +//! values, references). No panic; < 2 s of CPU time per 1,000 cases (summed over the workers' +//! thread clocks, so a loaded machine or the worker count does not change the bound). + +use super::cff_encoding::cff_builtin_encoding; +use super::glyph_budget::WorkMeter; +use super::program::{Outlines, TrueTypeLookup}; +use super::tounicode::parse_tounicode; +use super::type1::parse_type1; +use super::{FontCache, FontKey, FontModel}; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::limits::PAGE_DECODE_BUDGET; +use crate::pdf_engine::text_edit::testkit::cff::{CffBuilder, CffEncodingSpec}; +use crate::pdf_engine::text_edit::testkit::fonts::{cmap, latin_truetype}; +use crate::pdf_engine::text_edit::testkit::type1::{T1Encoding, T1Glyph, Type1Builder}; +use lopdf::{dictionary, Dictionary, Document, Object, ObjectId, Stream, StringFormat}; +use std::time::Duration; + +/// CPU time used by the calling thread (wall time where the platform has no thread clock). +pub(super) fn thread_cpu() -> Duration { + #[cfg(all( + any(target_os = "macos", target_os = "linux"), + target_pointer_width = "64" + ))] + { + #[repr(C)] + struct Timespec { + tv_sec: i64, + tv_nsec: i64, + } + extern "C" { + fn clock_gettime(clock: i32, tp: *mut Timespec) -> i32; + } + #[cfg(target_os = "macos")] + const CLOCK_THREAD_CPUTIME_ID: i32 = 16; + #[cfg(target_os = "linux")] + const CLOCK_THREAD_CPUTIME_ID: i32 = 3; + let mut ts = Timespec { + tv_sec: 0, + tv_nsec: 0, + }; + // SAFETY: clock_gettime only writes the timespec it is handed, which outlives the call. + if unsafe { clock_gettime(CLOCK_THREAD_CPUTIME_ID, &mut ts) } == 0 { + return Duration::new( + ts.tv_sec.max(0) as u64, + ts.tv_nsec.clamp(0, 999_999_999) as u32, + ); + } + } + static START: std::sync::OnceLock = std::sync::OnceLock::new(); + START.get_or_init(std::time::Instant::now).elapsed() +} + +/// Runs cases `0..n` on four workers; returns the sum of `run`'s results and of the workers' +/// CPU times. A panic in any case fails the test. +fn run_cases(n: usize, what: &str, run: impl Fn(usize) -> usize + Sync) -> (usize, Duration) { + std::thread::scope(|scope| { + let workers: Vec<_> = (0..4) + .map(|w| { + let run = &run; + scope.spawn(move || { + let started = thread_cpu(); + let sum = (w..n).step_by(4).map(run).sum::(); + (sum, thread_cpu().saturating_sub(started)) + }) + }) + .collect(); + workers + .into_iter() + .map(|h| { + h.join() + .unwrap_or_else(|_| panic!("{what}: no panic in any case")) + }) + .fold((0, Duration::ZERO), |(a, t), (b, u)| (a + b, t + u)) + }) +} + +struct XorShift(u64); + +impl XorShift { + fn next(&mut self) -> u64 { + let mut x = self.0; + x ^= x << 13; + x ^= x >> 7; + x ^= x << 17; + self.0 = x; + x + } + + fn below(&mut self, n: usize) -> usize { + (self.next() % n.max(1) as u64) as usize + } +} + +const INSERTS: &[&[u8]] = &[ + b"(", + b"<", + b"[", + b"<<", + b"BI", + b"q", + b"99999999999999999999", + b"-2147483648", + b"RD ", + b" -| ", + b"endbfchar", + b"beginbfrange", + b"", + b"/Subrs", + b"/CharStrings", + b"/lenIV 9 ", + b"\xff\xff\xff\xff", + b"\x00\x00", +]; + +fn mutate(rng: &mut XorShift, input: &[u8]) -> Vec { + let mut out = input.to_vec(); + for _ in 0..=rng.below(4) { + match rng.below(5) { + 0 if !out.is_empty() => { + let i = rng.below(out.len()); + out[i] ^= 1 << rng.below(8); + } + 1 if !out.is_empty() => { + let i = rng.below(out.len()); + out[i] = rng.next() as u8; + } + 2 => { + let cut = rng.below(out.len() + 1); + out.truncate(cut); + } + 3 => { + let at = rng.below(out.len() + 1); + let token = INSERTS[rng.below(INSERTS.len())]; + out.splice(at..at, token.iter().copied()); + } + _ if !out.is_empty() => { + let start = rng.below(out.len()); + let end = (start + rng.below(64)).min(out.len()); + let at = rng.below(out.len() + 1); + let slice = out[start..end].to_vec(); + out.splice(at..at, slice); + } + _ => {} + } + } + out +} + +/// One in-memory font: a TrueType/Type1/Type1C simple font or a Type0 font around `program`. +fn load(kind: usize, program: Vec, tounicode: Vec, l1: i64, l2: i64) -> usize { + let mut doc = Document::with_version("1.7"); + let tu = doc.add_object(Stream::new(dictionary! {}, tounicode)); + let (key, file_dict) = match kind { + 0 => ("FontFile2", dictionary! {}), + 1 => ( + "FontFile", + dictionary! { "Length1" => l1, "Length2" => l2, "Length3" => 0 }, + ), + 2 => ("FontFile3", dictionary! { "Subtype" => "Type1C" }), + _ => ("FontFile3", dictionary! { "Subtype" => "CIDFontType0C" }), + }; + let file = doc.add_object(Stream::new(file_dict, program)); + let desc = doc.add_object(dictionary! { + "Type" => "FontDescriptor", "FontName" => "ABCDEF+Fuzz", "Flags" => if kind == 0 { 32 } else { 4 }, + key => file, "Ascent" => 800, "Descent" => -200, + }); + // A /Differences-only encoding (on a symbolic flag for the Type1/CFF kinds) keeps the name + // path short, so the cases spend their time on the mutated programs and CMaps; TrueType is + // non-symbolic so its glyphs are selected by name (AGL → cmap, post). + let differences: Vec = [ + Object::Integer(32), + "space".into(), + Object::Integer(65), + "A".into(), + "B".into(), + "C".into(), + "Aacute".into(), + Object::Integer(89), + "Y".into(), + Object::Integer(97), + "a".into(), + "b".into(), + "c".into(), + ] + .into(); + let widths: Vec = (32..=255).map(|_| Object::Integer(500)).collect(); + let font = if kind == 3 { + let cid = doc.add_object(dictionary! { + "Type" => "Font", "Subtype" => "CIDFontType0", "BaseFont" => "ABCDEF+Fuzz", + "FontDescriptor" => desc, "W" => vec![Object::Integer(1), Object::Array(vec![Object::Integer(500)])], + }); + dictionary! { + "Type" => "Font", "Subtype" => "Type0", "BaseFont" => "ABCDEF+Fuzz", + "Encoding" => "Identity-H", "DescendantFonts" => vec![Object::Reference(cid)], "ToUnicode" => tu, + } + } else { + dictionary! { + "Type" => "Font", "Subtype" => if kind == 0 { "TrueType" } else { "Type1" }, + "BaseFont" => "ABCDEF+Fuzz", "FirstChar" => 32, "LastChar" => 255, "Widths" => widths, + "Encoding" => dictionary! { "Differences" => differences }, + "FontDescriptor" => desc, "ToUnicode" => tu, + } + }; + let id = doc.add_object(font.clone()); + let model = FontCache::new().get_or_load( + &doc, + FontKey::Indirect(id), + &font, + &mut DecodeBudget::new(PAGE_DECODE_BUDGET), + ); + model.alphabet().len() +} + +#[test] +fn font_34_loader_fuzz() { + let tounicode = cmap( + "1 begincodespacerange <0000> endcodespacerange\n\ + 2 beginbfchar <0041> <0041> <0042> endbfchar\n\ + 2 beginbfrange <0050> <0060> <0061> <0070> <0072> [<0041> <00660069> <0043>] endbfrange", + ); + let type1 = Type1Builder { + encoding: T1Encoding::Custom(vec![(65, "A".into()), (66, "Aacute".into())]), + ..Type1Builder::new("Fuzz") + } + .glyph("A", T1Glyph::Box { width: 500 }) + .glyph("B", T1Glyph::HintedBox { width: 500 }) + .glyph("acute", T1Glyph::Box { width: 300 }) + .glyph( + "Aacute", + T1Glyph::Seac { + width: 500, + base: 65, + accent: 0xC2, + }, + ) + .glyph("space", T1Glyph::Blank { width: 250 }) + .build(); + let (truetype, _) = latin_truetype("ABCabc", " Y"); + let cff = CffBuilder::new("Fuzz") + .glyph("A", true) + .glyph("B", false) + .encoding(CffEncodingSpec::Format1(vec![(0x41, 1)])) + .supplement(0x61, "A") + .build(); + let cid_cff = CffBuilder::new("FuzzCID") + .cid_glyph(65, true) + .cid_glyph(66, false) + .build(); + // Cases are independent and seeded by their index; four workers share them. + let run_case = |case: usize| -> usize { + let mut rng = + XorShift(0x9E37_79B9_7F4A_7C15 ^ (case as u64 + 1).wrapping_mul(0x2545_F491_4F6C_DD1D)); + let tu = mutate(&mut rng, &tounicode); + let _ = parse_tounicode(&tu).map(|t| t.codes().count()); + match case % 4 { + 0 => { + let data = mutate(&mut rng, &truetype); + if let Ok(face) = ttf_parser::Face::parse(&data, 0) { + let outlines = Outlines::of_face(&face); + let lookup = TrueTypeLookup::new(&face, &outlines); + if let super::program::GidLookup::Agree(g) = lookup.symbolic(0x41) { + let _ = outlines.drawn(g, &mut WorkMeter::new(PAGE_DECODE_BUDGET)); + } + let _ = lookup.by_name("A"); + } + load(0, data, tu, 0, 0) + } + 1 => { + let data = mutate(&mut rng, &type1.data); + let l1 = (type1.length1 as i64 + rng.below(9) as i64 - 4).max(0); + let l2 = if rng.below(10) == 0 { + data.len() as i64 + } else { + type1.length2 as i64 + }; + if let Ok(program) = parse_type1(&data, l1 as usize, l2 as usize) { + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + for name in ["A", "B", "Aacute", "space", ".notdef"] { + let _ = program.proof(name, &mut meter); + } + } + load(1, data, tu, l1, l2) + } + 2 => { + let data = mutate(&mut rng, &cff); + let _ = cff_builtin_encoding(&data); + if let Some(outlines) = Outlines::of_cff(&data) { + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + for gid in 0..16 { + let _ = outlines.drawn(gid, &mut meter); + } + } + load(2, data, tu, 0, 0) + } + _ => { + let data = mutate(&mut rng, &cid_cff); + let _ = cff_builtin_encoding(&data); + load(3, data, tu, 0, 0) + } + } + }; + let (loaded, cpu) = run_cases(1_000, "FONT-34", run_case); + println!("FONT-34: 1,000 cases in {cpu:?} of CPU, {loaded} typeable characters in total"); + assert!(loaded > 0, "FONT-34 some mutants still load"); + assert!( + cpu.as_secs_f64() < 2.0, + "FONT-34 < 2 s of CPU per 1,000 cases: {cpu:?}" + ); +} + +/// A random PDF value: wrong types, extreme numbers, references to fonts, to nothing or to +/// themselves, nested arrays and dictionaries. +fn random_object(rng: &mut XorShift, depth: usize, ids: &[ObjectId]) -> Object { + const INTS: [i64; 10] = [ + 0, + -1, + 1, + 255, + 256, + 65_535, + 65_536, + i64::MAX, + i64::MIN, + 1 << 40, + ]; + const REALS: [f32; 6] = [0.0, -0.5, 1e30, f32::NAN, f32::INFINITY, 0.001]; + const NAMES: [&str; 8] = [ + "Identity-H", + "WinAnsiEncoding", + "MacExpertEncoding", + "A", + "", + "Type1C", + "space", + "Identity", + ]; + match rng.below(if depth >= 2 { 8 } else { 10 }) { + 0 => Object::Null, + 1 => Object::Boolean(rng.below(2) == 0), + 2 => Object::Integer(INTS[rng.below(INTS.len())]), + 3 => Object::Real(REALS[rng.below(REALS.len())]), + 4 => Object::Name(NAMES[rng.below(NAMES.len())].as_bytes().to_vec()), + 5 => Object::String(b"/A/B/space".to_vec(), StringFormat::Literal), + 6 => Object::Reference(ids[rng.below(ids.len())]), + 7 => Object::Reference((999_999, 0)), + 8 => Object::Array( + (0..rng.below(6)) + .map(|_| random_object(rng, depth + 1, ids)) + .collect(), + ), + _ => Object::Dictionary(dictionary! { + "Differences" => random_object(rng, depth + 1, ids), + "BaseEncoding" => random_object(rng, depth + 1, ids), + }), + } +} + +/// Changes `dict`: one key set to a random value or dropped, or (for an array value) one item +/// replaced, the array cut, or an item added. +fn mutate_dict(rng: &mut XorShift, dict: &mut Dictionary, keys: &[&str], ids: &[ObjectId]) { + let key = keys[rng.below(keys.len())].as_bytes().to_vec(); + let choice = rng.below(10); + if choice == 0 { + let _ = dict.remove(&key); + return; + } + if let (1..=4, Ok(Object::Array(items))) = (choice, dict.get_mut(&key)) { + let i = rng.below(items.len() + 1); + match rng.below(3) { + 0 if i < items.len() => items[i] = random_object(rng, 1, ids), + 1 => items.truncate(i), + _ => items.insert(i, random_object(rng, 1, ids)), + } + return; + } + dict.set(key, random_object(rng, 0, ids)); +} + +/// Touches every query a later task makes of a model. +fn exercise(model: &FontModel) -> usize { + let codes = model.split_codes(b"\x00\x41\x20\x42").unwrap_or_default(); + for code in &codes { + let _ = (model.text(*code), model.width(*code), model.drawable(*code)); + } + let alphabet = model.alphabet(); + for (ch, _) in &alphabet { + let _ = model.code_for(*ch, &[], &[]).map(|c| model.code_bytes(c)); + } + alphabet.len() +} + +#[test] +fn font_34b_font_dictionary_fuzz() { + let (truetype, _) = latin_truetype("ABCabc", " "); + let tounicode = cmap( + "1 begincodespacerange <0000> endcodespacerange\n\ + 1 beginbfrange <0041> <0043> <0041> endbfrange", + ); + let run_case = |case: usize| -> usize { + let mut rng = + XorShift(0xD1B5_4A32_D192_ED03 ^ (case as u64 + 1).wrapping_mul(0x9E37_79B9_7F4A_7C15)); + let mut doc = Document::with_version("1.7"); + let program = doc.add_object(Stream::new(dictionary! {}, truetype.clone())); + let tu = doc.add_object(Stream::new(dictionary! {}, tounicode.clone())); + let gid_map = doc.add_object(Stream::new(dictionary! {}, vec![0, 0, 0, 1, 0, 2, 0, 3])); + let (font_id, desc_id, cid_id) = ( + doc.add_object(Object::Null), + doc.add_object(Object::Null), + doc.add_object(Object::Null), + ); + let ids = [font_id, desc_id, cid_id, program, tu, gid_map]; + let mut desc = dictionary! { + "Type" => "FontDescriptor", "FontName" => "ABCDEF+Fuzz", "Flags" => 32, + "FontFile2" => program, "Ascent" => 800, "Descent" => -200, "ItalicAngle" => 0, + "StemV" => 80, "MissingWidth" => 250, "CharSet" => Object::string_literal("/A/B"), + }; + let mut encoding = dictionary! { + "BaseEncoding" => "WinAnsiEncoding", + "Differences" => vec![ + Object::Integer(65), "A".into(), "B".into(), Object::Integer(32), "space".into(), + ], + }; + let type0 = case % 2 == 1; + let mut cid = dictionary! { + "Type" => "Font", "Subtype" => "CIDFontType2", "BaseFont" => "ABCDEF+Fuzz", + "FontDescriptor" => desc_id, "CIDToGIDMap" => gid_map, "DW" => 1000, + "W" => vec![ + Object::Integer(1), Object::Array(vec![500.into(), 600.into()]), + 3.into(), 4.into(), 700.into(), + ], + }; + let widths: Vec = (32..=127).map(|_| Object::Integer(500)).collect(); + let mut font = if type0 { + dictionary! { + "Type" => "Font", "Subtype" => "Type0", "BaseFont" => "ABCDEF+Fuzz", + "Encoding" => "Identity-H", "DescendantFonts" => vec![Object::Reference(cid_id)], + "ToUnicode" => tu, + } + } else { + dictionary! { + "Type" => "Font", "Subtype" => "TrueType", "BaseFont" => "ABCDEF+Fuzz", + "FirstChar" => 32, "LastChar" => 127, "Widths" => widths, + "FontDescriptor" => desc_id, "ToUnicode" => tu, + } + }; + if !type0 && rng.below(2) == 0 { + font.set("Encoding", Object::Dictionary(encoding.clone())); + } + for _ in 0..=rng.below(4) { + match rng.below(4) { + 0 => mutate_dict( + &mut rng, + &mut font, + &[ + "Subtype", + "FirstChar", + "LastChar", + "Widths", + "Encoding", + "FontDescriptor", + "ToUnicode", + "DescendantFonts", + "BaseFont", + ], + &ids, + ), + 1 => mutate_dict( + &mut rng, + &mut desc, + &[ + "Flags", + "Ascent", + "Descent", + "ItalicAngle", + "StemV", + "MissingWidth", + "CharSet", + "FontFile2", + "FontFile3", + ], + &ids, + ), + 2 => mutate_dict( + &mut rng, + &mut cid, + &[ + "W", + "DW", + "CIDToGIDMap", + "Subtype", + "WMode", + "FontDescriptor", + ], + &ids, + ), + _ => { + mutate_dict( + &mut rng, + &mut encoding, + &["BaseEncoding", "Differences"], + &ids, + ); + if !type0 { + font.set("Encoding", Object::Dictionary(encoding.clone())); + } + } + } + } + doc.objects.insert(desc_id, Object::Dictionary(desc)); + doc.objects.insert(cid_id, Object::Dictionary(cid)); + doc.objects + .insert(font_id, Object::Dictionary(font.clone())); + let model = FontCache::new().get_or_load( + &doc, + FontKey::Indirect(font_id), + &font, + &mut DecodeBudget::new(PAGE_DECODE_BUDGET), + ); + exercise(&model) + }; + let (typeable, cpu) = run_cases(1_000, "FONT-34b", run_case); + println!("FONT-34b: 1,000 cases in {cpu:?} of CPU, {typeable} typeable characters in total"); + assert!(typeable > 0, "FONT-34b some mutants still type"); + assert!( + cpu.as_secs_f64() < 2.0, + "FONT-34b < 2 s of CPU per 1,000 cases: {cpu:?}" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/tests_presence.rs b/src-tauri/src/pdf_engine/text_edit/fonts/tests_presence.rs new file mode 100644 index 0000000..e7cd335 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/tests_presence.rs @@ -0,0 +1,705 @@ +//! Glyph presence from the embedded program (§A.2, §B.9.5): FONT-16…19, 21…24, 29, 30, 38…40, +//! plus OpenType (glyf and CFF flavours) and the real FoxitSerif CFF program. + +use super::tests::{alphabet, c1, c2, typeable, winansi_truetype}; +use super::{FontClass, FontModel}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::testkit::cff::{CffBuilder, CffEncodingSpec}; +use crate::pdf_engine::text_edit::testkit::fonts::{ + latin_truetype, load_simple, load_type0, tounicode_bfchar, Program, SimpleFont, Type0Font, +}; +use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; +use crate::pdf_engine::text_edit::testkit::type1::{op, T1Encoding, T1Glyph, Type1Builder}; +use std::sync::Arc; + +#[test] +fn font_16_b3_subset_without_y_is_not_typeable() { + let m = load_simple(&winansi_truetype("ABCDEF+Calibri", "Helo", "Y ")); + assert_eq!(m.refusal, None); + assert_eq!(m.class, Some(FontClass::SimpleTrueType)); + assert!(m.subset, "FONT-16 subset tag"); + assert_eq!( + alphabet(&m), + " Helo", + "FONT-16 B3: the zeroed Y glyph is not typeable" + ); + let y = m.info(c1(b'Y')).expect("code Y"); + assert_eq!(y.text.as_deref(), Some("Y"), "FONT-16 Y still reads"); + assert!( + y.gid.is_some() && !y.drawable, + "FONT-16 Y has a GID but no outline" + ); + assert!(!m.drawable(c1(b'Z')), "FONT-16 a code with no glyph at all"); + // The space has an empty glyph and width 500: whitespace with a width is drawable. + assert!(m.drawable(c1(b' '))); + assert_eq!(m.code_for(' ', &[], &[]), Some(c1(b' '))); +} + +#[test] +fn font_17_zero_width_is_not_typeable() { + let mut f = winansi_truetype("ABCDEF+Calibri", "AB ", ""); + let widths = f.widths.as_mut().expect("widths"); + widths[usize::from(b'A' - 32)] = 0.0; + widths[usize::from(b' ' - 32)] = 0.0; + let m = load_simple(&f); + assert!(m.drawable(c1(b'A')), "FONT-17 the glyph itself is present"); + assert_eq!(typeable(&m, c1(b'A')), None, "FONT-17 width 0"); + assert_eq!( + typeable(&m, c1(b' ')), + None, + "FONT-17 a zero-width space either" + ); + assert_eq!(typeable(&m, c1(b'B')), Some('B')); + assert_eq!(m.width(c1(b'A')), 0.0); + // An empty space glyph is drawable only with a positive width (§A.2). + let mut f = winansi_truetype("ABCDEF+Calibri", "AB", " "); + f.widths.as_mut().expect("widths")[0] = 0.0; + assert!( + !load_simple(&f).drawable(c1(b' ')), + "FONT-17 blank space, width 0" + ); + f.widths.as_mut().expect("widths")[0] = 250.0; + assert!( + load_simple(&f).drawable(c1(b' ')), + "FONT-17 blank space, width 250" + ); +} + +#[test] +fn font_18_missing_width_only_when_present() { + let mut f = winansi_truetype("ABCDEF+Calibri", "ABC", ""); + f.first_char = 65; + f.widths = Some(vec![600.0, 650.0]); + let m = load_simple(&f); + assert_eq!(m.width(c1(b'B')), 650.0); + assert_eq!( + m.width(c1(b'C')), + 0.0, + "FONT-18 outside the range, no /MissingWidth: 0" + ); + assert_eq!(m.info(c1(b'C')).and_then(|i| i.width1000), None); + assert_eq!( + typeable(&m, c1(b'C')), + None, + "FONT-18 unknown width is never typeable" + ); + f.descriptor_extra = "/MissingWidth 420".into(); + let m = load_simple(&f); + assert_eq!( + m.width(c1(b'C')), + 420.0, + "FONT-18 /MissingWidth when present" + ); + assert_eq!(typeable(&m, c1(b'C')), Some('C')); +} + +#[test] +fn font_19_strategy_disagreement_is_not_drawable() { + // (3,1) maps 'A' to GID 1 ("A.alt"), while the post table names GID 2 "A". + let mut t = TtfBuilder::new(); + let alt = t.glyph("Aalt", true, 500); + t.glyph("A", true, 500); + let b = t.glyph("B", true, 500); + t.cmap31.push((u32::from('A'), alt)); + t.cmap31.push((u32::from('B'), b)); + let mut f = winansi_truetype("ABCDEF+Calibri", "", ""); + f.program = Program::TrueType(t.build()); + let m = load_simple(&f); + assert!(!m.drawable(c1(b'A')), "FONT-19 strategies disagree"); + assert_eq!(m.info(c1(b'A')).and_then(|i| i.gid), None); + assert_eq!( + typeable(&m, c1(b'B')), + Some('B'), + "FONT-19 B: cmap and post agree" + ); + // (1,0) by MacRoman code agreeing with (3,1). + let mut t = TtfBuilder::new(); + let e = t.glyph("eacute", true, 500); + t.cmap31.push((u32::from('é'), e)); + t.cmap10.push((0x8E, e)); + let mut f = winansi_truetype("ABCDEF+Calibri", "", ""); + f.program = Program::TrueType(t.build()); + assert_eq!(typeable(&load_simple(&f), c1(0xE9)), Some('é')); + // (1,0) pointing at another glyph: disagreement. + let mut t = TtfBuilder::new(); + let e = t.glyph("eacute", true, 500); + let other = t.glyph("x", true, 500); + t.cmap31.push((u32::from('é'), e)); + t.cmap10.push((0x8E, other)); + t.post_names = false; + let mut f = winansi_truetype("ABCDEF+Calibri", "", ""); + f.program = Program::TrueType(t.build()); + assert!( + !load_simple(&f).drawable(c1(0xE9)), + "FONT-19 (3,1) vs (1,0)" + ); +} + +#[test] +fn font_07_symbolic_truetype_3_0_cmap() { + let mut t = TtfBuilder::new(); + let a = t.glyph("glyph1", true, 500); + let b = t.glyph("glyph2", true, 500); + let sp = t.glyph("glyph3", false, 250); + t.cmap30.push((0xF001, a)); + t.cmap30.push((0xF002, b)); + t.cmap30.push((0x0003, sp)); // (3,0) at the bare code too + let mut f = SimpleFont::new("TrueType", "ABCDEF+LiberationSerif"); + f.first_char = 1; + f.widths = Some(vec![500.0, 500.0, 250.0]); + f.flags = Some(4); + f.program = Program::TrueType(t.build()); + f.tounicode = Some(tounicode_bfchar(&[(1, 1, "S"), (2, 1, "ğ"), (3, 1, " ")])); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(alphabet(&m), " Sğ", "FONT-07 0xF000+code and code"); + assert_eq!( + m.info(c1(1)).and_then(|i| i.glyph_name.clone()), + None, + "FONT-07 no names" + ); + f.tounicode = None; + assert_eq!( + load_simple(&f).refusal, + Some(TextReason::NoTounicode), + "FONT-07" + ); + f.tounicode = Some(b"garbage beginbfchar".to_vec()); + assert_eq!(load_simple(&f).refusal, Some(TextReason::AmbiguousUnicode)); +} + +/// A Type0 Identity-H CIDFontType2 over `program` with `tounicode` entries `(cid, text)`. +fn cid2(program: Vec, entries: &[(u32, &str)]) -> Type0Font { + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Calibri"); + f.program = Program::TrueType(program); + let e: Vec<(u32, usize, &str)> = entries.iter().map(|(c, t)| (*c, 2, *t)).collect(); + f.tounicode = Some(tounicode_bfchar(&e)); + f +} + +#[test] +fn font_21_cid_to_gid_map_stream() { + let (program, gids) = latin_truetype("AB", ""); + assert_eq!(gids, [('A', 1), ('B', 2)]); + let mut f = cid2(program, &[(5, "B"), (6, "A"), (7, "C"), (40, "A")]); + let mut map = vec![0u8; 16]; + map[10..12].copy_from_slice(&2u16.to_be_bytes()); // CID 5 → GID 2 + map[12..14].copy_from_slice(&1u16.to_be_bytes()); // CID 6 → GID 1 + f.cid_to_gid = Some(Some(map)); + let m = load_type0(&f); + assert_eq!(m.refusal, None); + assert_eq!(m.class, Some(FontClass::Type0Cid2)); + assert_eq!(m.info(c2(5)).and_then(|i| i.gid), Some(2)); + assert_eq!(typeable(&m, c2(5)), Some('B')); + assert_eq!(typeable(&m, c2(6)), Some('A')); + assert_eq!(typeable(&m, c2(7)), None, "FONT-21 CID 7 → GID 0"); + assert_eq!(typeable(&m, c2(40)), None, "FONT-21 CID past the map"); + assert_eq!(m.code_for('A', &[], &[]), Some(c2(6))); + assert_eq!(m.code_bytes(c2(6)), vec![0, 6]); + f.cid_to_gid = Some(Some(vec![0; 131_074])); + assert_eq!( + load_type0(&f).refusal, + Some(TextReason::FontUnsupported), + "FONT-21 cap" + ); + f.cid_to_gid = Some(None); + let m = load_type0(&f); + assert_eq!( + typeable(&m, c2(5)), + None, + "FONT-21 /Identity: CID 5 has no glyph" + ); +} + +#[test] +fn font_22_cid_font_type0_cid_keyed_and_name_keyed() { + let cff = CffBuilder::new("TestCID") + .cid_glyph(100, true) + .cid_glyph(200, false) + .build(); + let mut f = Type0Font::new("CIDFontType0", "ABCDEF+SourceHanSans"); + f.program = Program::CidCff(cff.clone()); + f.dw = Some(500.0); + f.tounicode = Some(tounicode_bfchar(&[ + (100, 2, "A"), + (200, 2, " "), + (300, 2, "B"), + ])); + let m = load_type0(&f); + assert_eq!(m.refusal, None); + assert_eq!(m.class, Some(FontClass::Type0Cid0)); + assert_eq!( + m.info(c2(100)).and_then(|i| i.gid), + Some(1), + "FONT-22 CID 100 → GID 1" + ); + assert_eq!(alphabet(&m), " A", "FONT-22 CID-keyed: B has no CID"); + // The same CFF wrapped as OpenType. + let mut t = TtfBuilder::new(); + t.cff = Some(cff); + t.glyph("cid100", true, 500); + t.glyph("cid200", false, 500); + f.program = Program::OpenType(t.build()); + let m = load_type0(&f); + assert_eq!(m.refusal, None, "FONT-22 OpenType CIDFontType0"); + assert_eq!(alphabet(&m), " A"); + // Name-keyed CFF: GID = CID. + let cff = CffBuilder::new("TestNamed") + .glyph("A", true) + .glyph("B", true) + .build(); + let mut f = Type0Font::new("CIDFontType0", "ABCDEF+TestNamed"); + f.program = Program::CidCff(cff); + f.tounicode = Some(tounicode_bfchar(&[(1, 2, "A"), (2, 2, "B"), (3, 2, "C")])); + let m = load_type0(&f); + assert_eq!(m.refusal, None); + assert_eq!(alphabet(&m), "AB", "FONT-22 name-keyed: GID = CID"); +} + +fn type1_font(builder: Type1Builder, encoding: &str, extra: &str) -> Arc { + let mut f = SimpleFont::new("Type1", "ABCDEF+CMR10"); + f.encoding = Some(encoding.into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.flags = Some(4); + f.descriptor_extra = extra.into(); + f.program = Program::Type1(builder.build()); + load_simple(&f) +} + +fn latin_type1() -> Type1Builder { + Type1Builder::new("CMR10") + .glyph("A", T1Glyph::Box { width: 750 }) + .glyph("B", T1Glyph::HintedBox { width: 700 }) + .glyph("C", T1Glyph::Box { width: 700 }) + .glyph("space", T1Glyph::Blank { width: 333 }) +} + +#[test] +fn font_23_type1_charstrings_and_charset() { + let m = type1_font(latin_type1(), "/WinAnsiEncoding", ""); + assert_eq!(m.refusal, None); + assert_eq!(m.class, Some(FontClass::SimpleType1)); + assert_eq!( + alphabet(&m), + " ABC", + "FONT-23 drawn, hinted and whitespace glyphs" + ); + let m = type1_font(latin_type1(), "/WinAnsiEncoding", "/CharSet (/A/B/space)"); + assert_eq!( + alphabet(&m), + " AB", + "FONT-23 C is in /CharStrings but not in /CharSet" + ); + let m = type1_font(latin_type1(), "/WinAnsiEncoding", "/CharSet (/A/D)"); + assert_eq!( + alphabet(&m), + "A", + "FONT-23 D is in /CharSet but has no charstring" + ); +} + +#[test] +fn font_23b_real_cff_program_foxit_serif() { + // public/pdfjs/standard_fonts/FoxitSerif.pfb is a bare CFF ("ChromSerifOTF"), not a Type1 + // PFB: it exercises the Type1C path with a real program (see DEVIATIONS). + let path = std::path::Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../public/pdfjs/standard_fonts/FoxitSerif.pfb"); + let data = std::fs::read(&path).expect("FoxitSerif.pfb in the repo"); + assert_eq!(&data[..3], &[1, 0, 4], "FONT-23b FoxitSerif is CFF"); + let mut f = SimpleFont::new("Type1", "ABCDEF+FoxitSerif"); + f.encoding = + Some("<< /BaseEncoding /WinAnsiEncoding /Differences [128 /dotlessi /gbreve] >>".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.flags = Some(34); + f.program = Program::Cff(data); + let m = load_simple(&f); + assert_eq!(m.refusal, None, "FONT-23b the real program loads"); + let abc = alphabet(&m); + assert!( + !abc.contains('ğ'), + "FONT-23b FoxitSerif has no gbreve: not typeable" + ); + for ch in "ABCXYZabcxyz0189 .,;:!?()-ÄÖÜäöüçéßıŒ".chars() { + assert!( + abc.contains(ch), + "FONT-23b {ch:?} typeable in FoxitSerif: {abc}" + ); + } +} + +#[test] +fn font_24_tiny_type1_lacking_a_glyph() { + let t1 = Type1Builder::new("Tiny") + .glyph("A", T1Glyph::Box { width: 600 }) + .glyph( + "Q", + // an unknown operator (15) before a fully drawn box + T1Glyph::Raw( + [ + op(&[0, 600], &[13]), + op(&[1, 2], &[15]), + op(&[100, 0], &[21]), + op(&[400, 0], &[5]), + op(&[0, 500], &[5]), + vec![9, 14], + ] + .concat(), + ), + ) + .glyph("Z", T1Glyph::Raw(op(&[100, 0], &[21]))); + let m = type1_font(t1, "/WinAnsiEncoding", ""); + assert_eq!(m.refusal, None); + assert_eq!(alphabet(&m), "A", "FONT-24 B has no charstring"); + assert!( + !m.drawable(c1(b'Q')), + "FONT-24 an unknown operator leaves Q unproven" + ); + assert!(!m.drawable(c1(b'Z')), "FONT-24 no hsbw first / no endchar"); + assert_eq!(m.text(c1(b'B')), Some("B"), "FONT-24 B still reads"); +} + +#[test] +fn font_38_type1_hex_eexec() { + let t1 = Type1Builder { + hex: true, + ..latin_type1() + }; + let file = t1.build(); + assert!(file.data[file.length1..file.length1 + 4] + .iter() + .all(u8::is_ascii_hexdigit)); + let m = type1_font(t1, "/WinAnsiEncoding", ""); + assert_eq!(m.refusal, None); + assert_eq!(alphabet(&m), " ABC", "FONT-38 hex eexec decoded first"); +} + +#[test] +fn font_39_type1_len_iv_minus_one() { + let t1 = Type1Builder { + len_iv: -1, + ..latin_type1() + }; + let m = type1_font(t1, "/WinAnsiEncoding", ""); + assert_eq!(alphabet(&m), " ABC", "FONT-39 plaintext charstrings"); + let t1 = Type1Builder { + len_iv: 0, + ..latin_type1() + }; + assert_eq!( + alphabet(&type1_font(t1, "/WinAnsiEncoding", "")), + " ABC", + "lenIV 0" + ); +} + +#[test] +fn font_40_seac_needs_both_components() { + // StandardEncoding: 0x41 = A, 0xC2 = acute. + let with = Type1Builder::new("Accents") + .glyph("A", T1Glyph::Box { width: 600 }) + .glyph("acute", T1Glyph::Box { width: 300 }) + .glyph( + "Aacute", + T1Glyph::Seac { + width: 600, + base: 0x41, + accent: 0xC2, + }, + ); + let m = type1_font(with, "/WinAnsiEncoding", ""); + assert_eq!( + typeable(&m, c1(0xC1)), + Some('Á'), + "FONT-40 both components present" + ); + let without = Type1Builder::new("Accents") + .glyph("A", T1Glyph::Box { width: 600 }) + .glyph( + "Aacute", + T1Glyph::Seac { + width: 600, + base: 0x41, + accent: 0xC2, + }, + ); + let m = type1_font(without, "/WinAnsiEncoding", ""); + assert_eq!(typeable(&m, c1(0xC1)), None, "FONT-40 acute missing"); + assert_eq!(typeable(&m, c1(b'A')), Some('A')); +} + +#[test] +fn font_29_corrupt_programs_are_unreadable_without_panic() { + let (program, _) = latin_truetype("AB", ""); + let cases: Vec<(&str, SimpleFont)> = vec![ + ("garbage FontFile2", { + let mut f = winansi_truetype("ABCDEF+X", "", ""); + f.program = Program::TrueType(b"not a font at all".to_vec()); + f + }), + ("truncated FontFile2", { + let mut f = winansi_truetype("ABCDEF+X", "", ""); + f.program = Program::TrueType(program[..program.len() / 3].to_vec()); + f + }), + ("garbage Type1C", { + let mut f = SimpleFont::new("Type1", "ABCDEF+X"); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.program = Program::Cff(vec![1, 0, 4, 4, 0, 1, 1, 1, 255]); + f + }), + ("Type1 without Length2", { + let mut f = SimpleFont::new("Type1", "ABCDEF+X"); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + let built = latin_type1().build(); + f.program = Program::Raw { + key: "FontFile", + subtype: None, + data: built.data, + }; + f + }), + ("Type1 Length2 out of range", { + let mut f = SimpleFont::new("Type1", "ABCDEF+X"); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + let mut built = latin_type1().build(); + built.length2 = built.data.len() * 2; + f.program = Program::Type1(built); + f + }), + ("FontFile is not a stream", { + let mut f = winansi_truetype("ABCDEF+X", "", ""); + f.program = Program::None; + f.descriptor_extra = "/FontFile2 [1 2 3]".into(); + f + }), + ]; + for (what, f) in cases { + let m = load_simple(&f); + assert_eq!( + m.refusal, + Some(TextReason::FontProgramUnreadable), + "FONT-29 {what}" + ); + assert_eq!(m.class, None); + assert_eq!(alphabet(&m), ""); + } + let mut f = cid2(b"\x00\x01\x00\x00garbage".to_vec(), &[(1, "A")]); + assert_eq!( + load_type0(&f).refusal, + Some(TextReason::FontProgramUnreadable), + "FONT-29 CID" + ); + f.program = Program::None; + assert_eq!( + load_type0(&f).refusal, + Some(TextReason::FontNotEmbedded), + "FONT-29 no program" + ); +} + +#[test] +fn font_30_fontfile3_subtype_mismatch_and_opentype() { + let (tt, _) = latin_truetype("AB", ""); + let cff = CffBuilder::new("X").glyph("A", true).build(); + let mismatches: Vec<(&str, &'static str, Program)> = vec![ + ( + "Type1 + CIDFontType0C", + "Type1", + Program::CidCff(cff.clone()), + ), + ("TrueType + Type1C", "TrueType", Program::Cff(cff.clone())), + ("Type1 + FontFile2", "Type1", Program::TrueType(tt.clone())), + ( + "FontFile3 without subtype", + "Type1", + Program::Raw { + key: "FontFile3", + subtype: None, + data: cff.clone(), + }, + ), + ( + "TrueType + FontFile", + "TrueType", + Program::Raw { + key: "FontFile", + subtype: None, + data: tt.clone(), + }, + ), + ]; + for (what, subtype, program) in mismatches { + let mut f = SimpleFont::new(subtype, "ABCDEF+X"); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.program = program; + assert_eq!( + load_simple(&f).refusal, + Some(TextReason::FontProgramUnsupported), + "FONT-30 {what}" + ); + } + let mut f = cid2(tt.clone(), &[(1, "A")]); + f.program = Program::CidCff(cff.clone()); + assert_eq!( + load_type0(&f).refusal, + Some(TextReason::FontProgramUnsupported), + "CIDFontType2 + CIDFontType0C" + ); + // OpenType, glyf flavour, in a simple TrueType font … + let mut f = winansi_truetype("ABCDEF+X", "", ""); + f.program = Program::OpenType(tt); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(m.class, Some(FontClass::SimpleOpenType)); + assert_eq!(alphabet(&m), "AB", "FONT-30 OpenType (glyf)"); + // … and CFF flavour (OTTO) in a simple Type1 font. + let mut t = TtfBuilder::new(); + let a = t.unicode_glyph('A', "A", true); + t.cff = Some( + CffBuilder::new("X") + .glyph("A", true) + .encoding(CffEncodingSpec::Standard) + .build(), + ); + assert_eq!(a, 1); + let mut f = SimpleFont::new("Type1", "ABCDEF+X"); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.flags = Some(32); + f.program = Program::OpenType(t.build()); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(alphabet(&m), "A", "FONT-30 OpenType (CFF)"); +} + +#[test] +fn font_30b_an_sfnt_with_both_glyf_and_cff_is_refused() { + // Review T4 r2 M-2: presence was proven on `glyf` while viewers draw an `OTTO` program's CFF + // outlines, so "Hello" → "Hellj" passed with "j" drawn empty. Both tables: refused. + let mut t = TtfBuilder::new(); + for (ch, name) in [('H', "H"), ('e', "e"), ('l', "l"), ('o', "o"), ('j', "j")] { + t.unicode_glyph(ch, name, true); + } + let cff = |empty: &str| { + ["H", "e", "l", "o", "j"] + .iter() + .fold(CffBuilder::new("X"), |b, g| b.glyph(g, *g != empty)) + .encoding(CffEncodingSpec::Standard) + .build() + }; + let font = |program: Vec| { + let mut f = SimpleFont::new("Type1", "ABCDEF+X"); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.flags = Some(32); + f.program = Program::OpenType(program); + load_simple(&f) + }; + for (tag, magic) in [("OTTO", 0x4F54_544F_u32), ("1.0", 0x0001_0000)] { + let m = font(t.build_glyf_and_cff(cff("j"), magic)); + assert_eq!(m.refusal, Some(TextReason::FontProgramUnsupported), "{tag}"); + assert_eq!(alphabet(&m), "", "{tag}: nothing typeable"); + } + // Control: the same program without `glyf` is read as before, from its CFF outlines (the + // glyf-only flavour is FONT-30's). + let only_cff = TtfBuilder { + cff: Some(cff("j")), + ..t.clone() + }; + let m = font(only_cff.build()); + assert_eq!(m.refusal, None, "OTTO, CFF only"); + assert_eq!(alphabet(&m), "Helo", "\"j\" draws nothing in CFF"); +} + +#[test] +fn font_29b_t1_and_unusual_encodings_report_their_class() { + // A Type1 whose clear text has no /Encoding: no names, nothing typeable, no refusal. + let t1 = Type1Builder { + encoding: T1Encoding::Custom(Vec::new()), + ..Type1Builder::new("NoEnc") + } + .glyph("A", T1Glyph::Box { width: 500 }); + let mut f = SimpleFont::new("Type1", "ABCDEF+NoEnc"); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.program = Program::Type1(t1.build()); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(alphabet(&m), ""); + assert_eq!(m.text(c1(b'A')), None); +} + +#[test] +fn font_vertical_metrics() { + let m = load_simple(&winansi_truetype("ABCDEF+Calibri", "A", "")); + assert_eq!( + (m.ascent, m.descent), + (0.8, -0.2), + "descriptor /Ascent /Descent" + ); + let mut f = winansi_truetype("ABCDEF+Calibri", "A", ""); + f.descriptor_extra = "/Ascent 0 /Descent 0".into(); + let mut t = TtfBuilder::new(); + t.ascender = 900; + t.descender = -300; + t.unicode_glyph('A', "A", true); + f.program = Program::TrueType(t.build()); + let m = load_simple(&f); + assert_eq!( + (m.ascent, m.descent), + (0.9, -0.3), + "zero descriptor values → hhea" + ); + f.descriptor_extra = "/Ascent 5000 /Descent -3000".into(); + let m = load_simple(&f); + assert_eq!((m.ascent, m.descent), (1.5, -0.8), "clamped"); +} + +#[test] +fn font_hash_recursion_is_bounded() { + // A font dictionary carrying a 3,000-deep array: the hash stops descending at its nesting + // cap (deterministically), so two trees that differ only below it hash alike. + let deep = |leaf: i64| { + let mut obj = lopdf::Object::Integer(leaf); + for _ in 0..3_000 { + obj = lopdf::Object::Array(vec![obj]); + } + obj + }; + let hash = |leaf: i64| { + let doc = lopdf::Document::with_version("1.7"); + let mut dict = lopdf::dictionary! { "Type" => "Font", "Subtype" => "Type3" }; + dict.set("Deep", deep(leaf)); + let model = std::thread::scope(|s| { + s.spawn(|| { + super::FontCache::new().get_or_load( + &doc, + super::FontKey::Indirect((1, 0)), + &dict, + &mut crate::pdf_engine::text_edit::decode::DecodeBudget::new(1 << 20), + ) + }) + .join() + .expect("no stack overflow") + }); + assert_eq!(model.refusal, Some(TextReason::Type3)); + let h = model.content_hash; + // Drop the deep tree iteratively-ish: unwrap level by level. + let mut obj = dict.remove(b"Deep"); + while let Some(lopdf::Object::Array(mut items)) = obj { + obj = items.pop(); + } + h + }; + assert_eq!(hash(1), hash(2), "below the nesting cap nothing is hashed"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/tests_tounicode.rs b/src-tauri/src/pdf_engine/text_edit/fonts/tests_tounicode.rs new file mode 100644 index 0000000..34a03cd --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/tests_tounicode.rs @@ -0,0 +1,259 @@ +//! FONT-08…15: the bounded ToUnicode parser (§B.9.4) and how ToUnicode and glyph names combine +//! when reading and writing (§A.3.1, §A.3.2). + +use super::tests::{alphabet, c1, c2, std14, typeable}; +use super::tounicode::{parse_tounicode, Lookup}; +use crate::pdf_engine::text_edit::limits::CMAP_MAPPINGS_MAX; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::testkit::fonts::{ + cmap, load_simple, load_type0, tounicode_bfchar, Program, Type0Font, +}; +use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; + +fn text(tu: &super::tounicode::ToUnicode, code: super::Code) -> Option { + match tu.lookup(code.value) { + Lookup::Text(t) => Some(t.to_string()), + Lookup::Undecodable => Some("".into()), + Lookup::Absent => None, + } +} + +#[test] +fn font_08_bfchar() { + let tu = parse_tounicode(&cmap( + "1 begincodespacerange <0000> endcodespacerange\n\ + 3 beginbfchar\n<0003> <0020>\n<0024> <011F>\n<0025> <0130>\nendbfchar", + )) + .expect("FONT-08 parses"); + assert_eq!(text(&tu, c2(3)).as_deref(), Some(" ")); + assert_eq!(text(&tu, c2(0x24)).as_deref(), Some("ğ")); + assert_eq!(text(&tu, c2(0x25)).as_deref(), Some("İ")); + assert_eq!(text(&tu, c2(0x26)), None, "FONT-08 unmapped code is absent"); + assert_eq!( + text(&tu, c1(0x24)).as_deref(), + Some("ğ"), + "FONT-08 codes are keyed by value, whatever the source length (pdf.js, Poppler)" + ); + // Literal-string sources and a later mapping of the same code (later wins). + let tu = parse_tounicode(&cmap("2 beginbfchar\n(A) <0061>\n<41> <0062>\nendbfchar")) + .expect("parses"); + assert_eq!(text(&tu, c1(0x41)).as_deref(), Some("b")); +} + +#[test] +fn font_09_bfrange_incrementing() { + let tu = parse_tounicode(&cmap( + "2 beginbfrange\n<0041> <0043> <0061>\n<0001> <0003> \nendbfrange", + )) + .expect("FONT-09 parses"); + let abc: Vec<_> = (0x41..=0x43).map(|c| text(&tu, c2(c))).collect(); + assert_eq!(abc, [Some("a".into()), Some("b".into()), Some("c".into())]); + assert_eq!(text(&tu, c2(1)).as_deref(), Some("\u{FFFE}")); + assert_eq!(text(&tu, c2(2)).as_deref(), Some("\u{FFFF}")); + assert_eq!( + text(&tu, c2(3)), + None, + "FONT-09 overflow past 0xFFFF stops the range" + ); + // The last UTF-16 unit is incremented: a ligature destination keeps its prefix. + let tu = + parse_tounicode(&cmap("1 beginbfrange\n<10> <11> <00660069>\nendbfrange")).expect("parses"); + assert_eq!(text(&tu, c1(0x10)).as_deref(), Some("fi")); + assert_eq!(text(&tu, c1(0x11)).as_deref(), Some("fj")); + for bad in [ + "1 beginbfrange\n<0043> <0041> <0061>\nendbfrange", + "1 beginbfrange\n<0041> <43> <0061>\nendbfrange", + "1 beginbfrange\n<0041> <0043> /a\nendbfrange", + "1 beginbfrange\n<0041> <0043>", + "1 beginbfchar\n<0041>\nendbfchar", + "1 beginbfchar\n<0000000041> <0041>\nendbfchar", + "1 begincodespacerange <00> endcodespacerange", + ] { + assert!( + parse_tounicode(&cmap(bad)).is_err(), + "FONT-09 syntax error: {bad}" + ); + } +} + +#[test] +fn font_10_bfrange_array_form() { + let tu = parse_tounicode(&cmap( + "1 beginbfrange\n<10> <13> [<0041> <0042> <00660069>]\nendbfrange", + )) + .expect("FONT-10 parses"); + assert_eq!(text(&tu, c1(0x10)).as_deref(), Some("A")); + assert_eq!(text(&tu, c1(0x11)).as_deref(), Some("B")); + assert_eq!(text(&tu, c1(0x12)).as_deref(), Some("fi")); + assert_eq!( + text(&tu, c1(0x13)), + None, + "FONT-10 a short array maps fewer codes" + ); + assert!(parse_tounicode(&cmap("1 beginbfrange\n<10> <11> [<0041> /B]\nendbfrange")).is_err()); +} + +#[test] +fn font_11_surrogates_and_bad_destinations() { + let tu = parse_tounicode(&cmap( + "4 beginbfchar\n<01> \n<02> \n<03> <000041>\n<04> <>\nendbfchar", + )) + .expect("FONT-11 parses"); + assert_eq!( + text(&tu, c1(1)).as_deref(), + Some("😀"), + "FONT-11 surrogate pair" + ); + assert_eq!( + text(&tu, c1(2)).as_deref(), + Some(""), + "FONT-11 lone surrogate" + ); + assert_eq!( + text(&tu, c1(3)).as_deref(), + Some(""), + "FONT-11 odd length" + ); + assert_eq!( + text(&tu, c1(4)).as_deref(), + Some(""), + "FONT-11 empty" + ); + let long = format!("1 beginbfchar\n<05> <{}>\nendbfchar", "0041".repeat(257)); + assert!( + parse_tounicode(&cmap(&long)).is_err(), + "FONT-11 destination > 512 bytes" + ); +} + +#[test] +fn font_12_ligatures_are_readable_not_writable() { + let mut f = std14( + "Helvetica", + "<< /BaseEncoding /WinAnsiEncoding /Differences [150 /fi 151 /fl] >>", + ); + f.tounicode = Some(tounicode_bfchar(&[ + (150, 1, "fi"), + (151, 1, "\u{FB02}"), + (65, 1, "A"), + ])); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!( + m.text(c1(150)), + Some("fi"), + "FONT-12 ToUnicode 'fi' agrees with /fi" + ); + assert_eq!( + m.text(c1(151)), + Some("fl"), + "FONT-12 U+FB02 displayed as fl" + ); + assert_eq!( + typeable(&m, c1(150)), + None, + "FONT-12 two scalars: not typeable" + ); + assert_eq!( + typeable(&m, c1(151)), + None, + "FONT-12 ligature code point: not typeable" + ); + assert_eq!(typeable(&m, c1(65)), Some('A')); +} + +#[test] +fn font_13_usecmap_makes_unmapped_codes_undecodable() { + let body = "/Adobe-Identity-UCS2 usecmap\n1 beginbfchar\n<41> <0041>\nendbfchar"; + let tu = parse_tounicode(&cmap(body)).expect("FONT-13 parses"); + assert_eq!(text(&tu, c1(0x41)).as_deref(), Some("A")); + assert_eq!(text(&tu, c1(0x42)).as_deref(), Some("")); + let mut f = std14("Helvetica", "/WinAnsiEncoding"); + f.tounicode = Some(cmap(body)); + let m = load_simple(&f); + assert_eq!(typeable(&m, c1(0x41)), Some('A')); + assert_eq!( + m.text(c1(0x42)), + None, + "FONT-13 the parent CMap might map B differently" + ); + assert_eq!(typeable(&m, c1(0x42)), None); +} + +#[test] +fn font_14_tounicode_bombs_are_bounded() { + let over = CMAP_MAPPINGS_MAX / 0x10000 + 1; + let ranges: String = (0..over) + .map(|i| format!("<{i:02X}0000> <{i:02X}FFFF> <0041>\n")) + .collect(); + let started = std::time::Instant::now(); + assert!( + parse_tounicode(&cmap(&format!("{over} beginbfrange\n{ranges}endbfrange"))).is_err(), + "FONT-14 more than CMAP_MAPPINGS_MAX mappings" + ); + // Exactly at the cap still parses; one bfchar more does not. + let full = "<000000> <00FFFF> <0000>\n<010000> <01FFFF> <0000>\n"; + let ok = parse_tounicode(&cmap(&format!("2 beginbfrange\n{full}endbfrange"))); + assert_eq!(ok.map(|t| t.codes().count()).ok(), Some(CMAP_MAPPINGS_MAX)); + let past_cap = cmap(&format!( + "2 beginbfrange\n{full}endbfrange\n1 beginbfchar\n<020000> <0041>\nendbfchar" + )); + assert!( + parse_tounicode(&past_cap).is_err(), + "FONT-14 one bfchar past the cap" + ); + assert!(started.elapsed().as_secs() < 30, "FONT-14 bounded time"); + // In a Type0 font a broken ToUnicode is the font's refusal. + let mut t = TtfBuilder::new(); + t.glyph("A", true, 500); + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Test"); + f.program = Program::TrueType(t.build()); + f.tounicode = Some(cmap(&format!("{over} beginbfrange\n{ranges}endbfrange"))); + assert_eq!(load_type0(&f).refusal, Some(TextReason::AmbiguousUnicode)); + f.tounicode = Some(b"begincmap 1 beginbfchar <0001> endcmap".to_vec()); + assert_eq!(load_type0(&f).refusal, Some(TextReason::AmbiguousUnicode)); + // A valid CMap padded past TOUNICODE_MAX_DECODED (it inflates from a few KiB): capped while + // inflating, refused. + let mut padded = vec![b' '; crate::pdf_engine::text_edit::limits::TOUNICODE_MAX_DECODED]; + padded.extend(tounicode_bfchar(&[(1, 2, "A")])); + f.tounicode = Some(padded); + assert_eq!( + load_type0(&f).refusal, + Some(TextReason::AmbiguousUnicode), + "FONT-14 cap" + ); + f.tounicode = Some(tounicode_bfchar(&[(1, 2, "A")])); + assert_eq!( + load_type0(&f).refusal, + None, + "FONT-14 the same CMap unpadded loads" + ); +} + +#[test] +fn font_15_tounicode_and_name_disagreement_is_undecodable() { + let mut f = std14("Helvetica", "/WinAnsiEncoding"); + f.tounicode = Some(tounicode_bfchar(&[ + (0x41, 1, "B"), + (0x42, 1, "B"), + (0x43, 1, "Ç"), + ])); + let m = load_simple(&f); + assert_eq!(m.text(c1(0x41)), None, "FONT-15 ToUnicode B vs name A"); + assert_eq!(typeable(&m, c1(0x41)), None); + assert_eq!(m.text(c1(0x42)), Some("B"), "FONT-15 agreement"); + assert_eq!(typeable(&m, c1(0x42)), Some('B')); + assert_eq!(m.text(c1(0x43)), None, "FONT-15 C vs Ç"); + assert_eq!( + m.text(c1(0x44)), + Some("D"), + "FONT-15 names alone still read" + ); + // A broken ToUnicode on a simple font: names still read, nothing types. + let mut f = std14("Helvetica", "/WinAnsiEncoding"); + f.tounicode = Some(b"beginbfchar <41> endbfchar".to_vec()); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + assert_eq!(m.text(c1(0x41)), Some("A")); + assert_eq!(alphabet(&m), "", "FONT-15 agreement cannot hold"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/tests_work.rs b/src-tauri/src/pdf_engine/text_edit/fonts/tests_work.rs new file mode 100644 index 0000000..c7a037c --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/tests_work.rs @@ -0,0 +1,375 @@ +//! Work bounds of the second review-T2 fix pass: CFF DICTs are walked lazily, capped and charged +//! per glyph (FONT-29h); glyph walks stop at what the page budget can still pay and an exhausted +//! meter walks nothing (FONT-29i); the CMap and Type1 clear-text token caps (FONT-14d); a vertical +//! Type0 font reports `VERTICAL` whatever else is wrong (FONT-28f); CFF custom-encoding names +//! (FONT-35b). + +use super::cff_layout::DICT_LEN_MAX; +use super::glyph_budget::WorkMeter; +use super::program::Outlines; +use super::tests::{c1, c2}; +use super::tests_bounds::{cpu, Segs}; +use super::tounicode::parse_tounicode; +use super::type1::{parse_type1, GlyphProof}; +use super::{FontCache, FontKey, FontModel}; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::limits::{PAGE_DECODE_BUDGET, TOUNICODE_MAX_DECODED}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::testkit::cff::{t2, CffBuilder, CffEncodingSpec}; +use crate::pdf_engine::text_edit::testkit::fonts::{ + add_type0, cmap, load_one, load_simple, load_type0, page_with_fonts, snapshot, Program, + SimpleFont, Type0Font, +}; +use crate::pdf_engine::text_edit::testkit::pdf::PdfBuilder; +use crate::pdf_engine::text_edit::testkit::thread_peak; +use crate::pdf_engine::text_edit::testkit::ttf::{composite, TtfBuilder}; +use crate::pdf_engine::text_edit::testkit::type1::{op, T1Glyph, Type1Builder}; +use lopdf::Object; +use std::sync::Arc; +use ttf_parser::GlyphId; + +/// Loads the one font of `f`'s page with a page budget of `budget` bytes, measuring only the +/// load: the model, its thread CPU seconds and the peak bytes it held on this thread. +fn measured_type0(f: &Type0Font, budget: usize) -> (Arc, f64, usize) { + let mut b = PdfBuilder::new(); + let id = add_type0(&mut b, f); + let snap = snapshot(page_with_fonts(b, &[("F1", id)])); + let dict = snap + .doc + .get_object((id, 0)) + .and_then(Object::as_dict) + .expect("font dict"); + let mut budget = DecodeBudget::new(budget); + let ((model, peak), secs) = cpu(|| { + thread_peak(|| { + FontCache::new().get_or_load(&snap.doc, FontKey::Indirect((id, 0)), dict, &mut budget) + }) + }); + (model, secs, peak) +} + +/// `100 0 rmoveto`, `middle`, `endchar`. +fn t2_glyph(middle: Vec) -> Vec { + [t2(100), t2(0), vec![21], middle, vec![14]].concat() +} + +/// A CID-keyed CFF with one Font DICT: CID 1 draws a line without subroutines; CIDs 2..=n+1 draw +/// theirs through local subroutine 0. The Private DICT carries `padding` after its entries. +fn cid_local_cff(n: u16, padding: Vec) -> Vec { + let mut b = CffBuilder { + subrs: vec![[t2(400), t2(0), vec![5, 11]].concat()], + private_padding: padding, + ..CffBuilder::new("CIDDict") + } + .raw_cid_glyph(1, t2_glyph([t2(400), t2(0), vec![5]].concat())); + for cid in 2..=n + 1 { + b = b.raw_cid_glyph(cid, t2_glyph([t2(-107), vec![10]].concat())); + } + b.build() +} + +/// Whether ttf-parser itself (no pre-check) draws `gid` of a bare CFF. +fn ttf_parser_draws(data: &[u8], gid: u16) -> bool { + let mut segs = Segs::default(); + ttf_parser::cff::Table::parse(data) + .is_some_and(|t| t.outline(GlyphId(gid), &mut segs).is_ok() && segs.0 > 0) +} + +#[test] +fn font_29h_cff_dicts_are_capped_lazy_and_charged() { + // 1,000 glyphs through one Font DICT whose Private DICT is 4 MiB of numbers or of operators: + // before the fix each glyph walked it (and built one 24-byte entry per operator byte). + for (what, pad) in [("numbers", 139u8), ("operators", 0u8)] { + let data = cid_local_cff(1_000, vec![pad; 4 << 20]); + let mut f = Type0Font::new("CIDFontType0", "ABCDEF+BigDict"); + // CJK destinations: no code reads as a space (a blank glyph with a width would count). + f.tounicode = Some(cmap("1 beginbfrange\n<0001> <03E9> <4E00>\nendbfrange")); + f.program = Program::CidCff(data.clone()); + let (m, secs, peak) = measured_type0(&f, PAGE_DECODE_BUDGET); + println!("FONT-29h {what}: {secs:.3} s of CPU, {peak} bytes peak"); + assert_eq!(m.refusal, None, "FONT-29h {what}"); + assert!( + m.drawable(c2(1)), + "FONT-29h {what}: no local subroutine, draws" + ); + assert!( + (2..=1_001).all(|cid| !m.drawable(c2(cid))), + "FONT-29h {what}: a Private DICT past DICT_LEN_MAX gives no local subroutines" + ); + assert!(secs < 1.0, "FONT-29h {what}: {secs:.2} s of CPU"); + assert!( + peak < data.len() + (4 << 20), + "FONT-29h {what}: {peak} bytes peak for a {}-byte program", + data.len() + ); + } + // Just past the cap the guard does not follow the Private DICT although ttf-parser would; + // just under it, it does. + for (pad, drawn) in [(DICT_LEN_MAX, false), (DICT_LEN_MAX - 64, true)] { + let data = cid_local_cff(1, vec![139; pad]); + assert!(ttf_parser_draws(&data, 2), "FONT-29h precondition"); + let outlines = Outlines::of_cff(&data).expect("CFF"); + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + assert_eq!( + outlines.drawn(2, &mut meter), + drawn, + "FONT-29h a {}-byte Private DICT", + pad + 10 + ); + } + // A name-keyed Private DICT or a Top DICT past the cap: the program is unreadable. + for (what, builder) in [ + ( + "Private", + CffBuilder { + private_padding: vec![139; DICT_LEN_MAX], + ..CffBuilder::new("BigPrivate") + }, + ), + ( + "Top", + CffBuilder { + top_padding: vec![139; DICT_LEN_MAX], + ..CffBuilder::new("BigTop") + }, + ), + ] { + let data = builder.glyph("A", true).build(); + assert!(ttf_parser_draws(&data, 1), "FONT-29h {what} precondition"); + let outlines = Outlines::of_cff(&data).expect("ttf-parser reads it"); + assert_eq!( + outlines.refusal(), + Some(TextReason::FontProgramUnreadable), + "FONT-29h a {what} DICT past DICT_LEN_MAX" + ); + } + // Every glyph calling a local subroutine pays the Font DICT (11 bytes) and Private DICT bytes + // ttf-parser walks again for it; the guard walks them once (per-Font-DICT cache). + let (n, pad) = (64u16, 4_000usize); + let data = cid_local_cff(n, vec![139; pad]); + let outlines = Outlines::of_cff(&data).expect("CFF"); + let mut meter = WorkMeter::new(PAGE_DECODE_BUDGET); + assert!( + (2..=n + 1).all(|gid| outlines.drawn(gid, &mut meter)), + "FONT-29h under the cap they draw" + ); + let dicts = (pad + 10 + 11) as u64; + let used = meter.used() as u64; + println!("FONT-29h {n} glyphs × {dicts} DICT bytes: {used} units charged"); + assert!( + used >= u64::from(n) * dicts, + "FONT-29h ttf-parser's per-glyph DICT walk is charged ({used} units)" + ); + assert!( + used <= u64::from(n + 1) * dicts + u64::from(n) * 64, + "FONT-29h the guard walks the DICTs once ({used} units)" + ); +} + +/// A TrueType program: glyph `A` draws; the returned GID is 30 levels of two references to the +/// level below (2^30 leaves for ttf-parser). +fn glyf_fan_out() -> (Vec, u16) { + let mut t = TtfBuilder::new(); + let mut level = t.unicode_glyph('A', "A", true); + for depth in 0..30 { + level = t.raw_glyph(&format!("level{depth}"), composite(&[level, level]), 500); + } + (t.build(), level) +} + +/// Global subroutines: 0 pushes 48 numbers (5 bytes each) and draws a line; 1 calls 0 eight +/// times. `heavy_charstring` calls 1 seventy times: ~137,000 units in ~1,800 operators, past +/// `GLYPH_WORK_MAX` before the operator cap. +fn heavy_gsubrs() -> Vec> { + let mut numbers: Vec = (0..48).flat_map(|_| [255u8, 0, 1, 0, 0]).collect(); + numbers.extend([5, 11]); + let mut caller: Vec = (0..8).flat_map(|_| [t2(-107), vec![29]].concat()).collect(); + caller.push(11); + vec![numbers, caller] +} + +fn heavy_charstring() -> Vec { + let mut code: Vec = (0..70) + .flat_map(|_| [t2(1 - 107), vec![29]].concat()) + .collect(); + code.push(14); + code +} + +/// A Type1 program whose glyph `fan` calls subr 6 175 times; subr 6 calls subr 5 eight times; +/// subr 5 pushes 46 numbers and draws a line (~63,000 tokens before the operator cap). +fn type1_fan_out() -> super::type1::Type1Program { + let mut numbers = op(&[1; 46], &[5]); + numbers.push(11); + let mut caller: Vec = (0..8).flat_map(|_| op(&[5], &[10])).collect(); + caller.push(11); + let mut fan = op(&[0, 600], &[13]); + for _ in 0..175 { + fan.extend(op(&[6], &[10])); + } + fan.push(14); + let file = Type1Builder { + extra_subrs: vec![numbers, caller], + ..Type1Builder::new("Fan") + } + .glyph("fan", T1Glyph::Raw(fan)) + .build(); + parse_type1(&file.data, file.length1, file.length2).expect("parses") +} + +#[test] +fn font_29i_an_exhausted_meter_stops_glyph_work() { + // One over-budget glyph checked 3,000 times (no memo) against 1 MiB of budget: once the meter + // runs out nothing more is walked. Before the fix each check walked up to GLYPH_WORK_MAX. + const CHECKS: usize = 3_000; + let (ttf, fan) = glyf_fan_out(); + let face = ttf_parser::Face::parse(&ttf, 0).expect("face"); + let glyf = Outlines::of_face(&face); + let cff_data = CffBuilder { + gsubrs: heavy_gsubrs(), + ..CffBuilder::new("Heavy") + } + .raw_glyph("A", heavy_charstring()) + .build(); + let cff = Outlines::of_cff(&cff_data).expect("CFF"); + let type1 = type1_fan_out(); + let (meters, secs) = cpu(|| { + let mut meters = [(); 3].map(|_| WorkMeter::new(1 << 20)); + let [m_glyf, m_cff, m_t1] = &mut meters; + for _ in 0..CHECKS { + assert!(!glyf.drawn(fan, m_glyf)); + assert!(!cff.drawn(1, m_cff)); + assert_eq!(type1.proof("fan", m_t1), GlyphProof::Unproven); + } + meters + }); + println!("FONT-29i {CHECKS} × 3 checks: {secs:.3} s of CPU"); + assert!( + meters.iter().all(WorkMeter::exhausted), + "FONT-29i the meters ran out" + ); + assert!( + secs < 1.0, + "FONT-29i {secs:.2} s of CPU for {CHECKS} × 3 checks" + ); + // Through a Type0 load: 4,000 such glyphs under a 4 MiB page budget. + let mut b = CffBuilder { + gsubrs: heavy_gsubrs(), + ..CffBuilder::new("HeavyCID") + }; + for cid in 1..=4_000 { + b = b.raw_cid_glyph(cid, heavy_charstring()); + } + let mut f = Type0Font::new("CIDFontType0", "ABCDEF+HeavyCID"); + f.program = Program::CidCff(b.build()); + f.tounicode = Some(cmap("1 beginbfrange\n<0001> <0FA0> <0041>\nendbfrange")); + let (m, secs, _) = measured_type0(&f, 4 << 20); + println!("FONT-29i Type0 load: {secs:.3} s of CPU"); + assert_eq!(m.refusal, Some(TextReason::PageTooComplex), "FONT-29i"); + assert!(secs < 1.0, "FONT-29i Type0 load: {secs:.2} s of CPU"); +} + +#[test] +fn font_14d_token_caps_bound_transient_memory() { + // 2 MiB of one-byte tokens: the scan stops at the token cap (the old 2^21-token cap held all + // of them, 48 bytes each: 100 MB). + let brackets = "[]".repeat(TOUNICODE_MAX_DECODED / 2); + let (parsed, peak) = thread_peak(|| parse_tounicode(brackets.as_bytes()).is_ok()); + println!("FONT-14d CMap: {peak} bytes peak"); + assert!(peak < 48 << 20, "FONT-14d CMap tokens: {peak} bytes peak"); + assert!(!parsed, "FONT-14d past the token cap"); + // The densest real form, one-destination bfrange arrays, still parses for every 2-byte code. + let mut body = String::new(); + let codes: Vec = (0..=0xFFFF).collect(); + for chunk in codes.chunks(100) { + body.push_str(&format!("{} beginbfrange\n", chunk.len())); + for c in chunk { + body.push_str(&format!("<{c:04X}> <{c:04X}> [<0041>]\n")); + } + body.push_str("endbfrange\n"); + } + let dense = cmap(&body); + assert!(dense.len() <= TOUNICODE_MAX_DECODED); + assert_eq!( + parse_tounicode(&dense).map(|t| t.codes().count()).ok(), + Some(0x10000), + "FONT-14d real CMaps fit the cap" + ); + // A Type1 clear text of 4 MiB of names: the built-in encoding is not read past the cap. + let file = Type1Builder::new("Pad") + .glyph("A", T1Glyph::Box { width: 600 }) + .build(); + let mut data = b"/a ".repeat((4 << 20) / 3); + let pad = data.len(); + data.extend_from_slice(&file.data); + let (program, peak) = + thread_peak(|| parse_type1(&data, pad + file.length1, file.length2).map(|p| p.builtin)); + println!("FONT-14d Type1 clear text: {peak} bytes peak"); + assert_eq!(program, Ok(None), "FONT-14d no encoding past the cap"); + assert!( + peak < 16 << 20, + "FONT-14d clear-text tokens: {peak} bytes peak" + ); +} + +#[test] +fn font_28f_type0_vertical_is_reported_before_early_returns() { + // `/FontDescriptor 5` (not a dictionary) is FONT_UNSUPPORTED (#19) and returns early; + // VERTICAL (#8) is still found and reported. + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Vert"); + f.flags = None; + f.encoding = "/Identity-V".into(); + f.cid_extra = "/FontDescriptor 5".into(); + let m = load_type0(&f); + assert_eq!((m.refusal, m.vertical), (Some(TextReason::Vertical), true)); + f.encoding = "/Identity-H".into(); + f.cid_extra = "/WMode 1 /FontDescriptor 5".into(); + let m = load_type0(&f); + assert_eq!( + (m.refusal, m.vertical), + (Some(TextReason::Vertical), true), + "FONT-28f CIDFont WMode" + ); + // No /DescendantFonts at all. + let mut b = PdfBuilder::new(); + let id = b.add("<< /Type /Font /Subtype /Type0 /BaseFont /NoKids /Encoding /Identity-V >>"); + let m = load_one(b, id); + assert_eq!( + (m.refusal, m.vertical), + (Some(TextReason::Vertical), true), + "FONT-28f no descendant" + ); +} + +#[test] +fn font_35b_cff_custom_encoding_names_match_ttf_parser() { + // Names come from one charset walk; they equal ttf-parser's per-GID `glyph_name` for every + // charset format, standard and custom (String INDEX) names alike. + let codes = [0x41u8, 0x42, 0x43, 0xE9]; + for charset_format in [0u8, 1, 2] { + let mut b = CffBuilder { + charset_format, + ..CffBuilder::new("Names") + } + .encoding(CffEncodingSpec::Format0(codes.to_vec())); + for name in ["A", "B", "Cfoo", "eacute"] { + b = b.glyph(name, true); + } + let data = b.build(); + let table = ttf_parser::cff::Table::parse(&data).expect("CFF"); + let mut f = SimpleFont::new("Type1", "ABCDEF+Names"); + f.first_char = 32; + f.widths = Some(vec![500.0; 224]); + f.flags = Some(4); + f.program = Program::Cff(data.clone()); + let m = load_simple(&f); + assert_eq!(m.refusal, None); + for (gid, code) in (1u16..).zip(codes) { + assert_eq!( + m.info(c1(code)).and_then(|i| i.glyph_name.as_deref()), + table.glyph_name(GlyphId(gid)), + "FONT-35b charset format {charset_format}, code {code:#04X}" + ); + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/tounicode.rs b/src-tauri/src/pdf_engine/text_edit/fonts/tounicode.rs new file mode 100644 index 0000000..3bf9b3f --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/tounicode.rs @@ -0,0 +1,291 @@ +//! Bounded ToUnicode CMap parser (SPEC §B.9.4), on top of `lexer::scan_tokens` in CMap mode. +//! +//! Sections: `begincodespacerange` (1–4-byte ranges), `beginbfchar`, `beginbfrange` (incrementing +//! form: the destination's last UTF-16 unit is incremented and the range stops on overflow past +//! 0xFFFF; array form). Destinations are UTF-16BE (surrogate pairs allowed; a lone surrogate, an +//! odd or empty destination maps the code to "undecodable"), at most 512 bytes. At most +//! `CMAP_MAPPINGS_MAX` mappings, whose destinations total at most `TOUNICODE_MAX_DECODED` bytes +//! (checked before a range is expanded, so a few bytes of CMap cannot expand into megabytes of +//! text). `usecmap` marks every code this CMap does not map itself as undecodable. Any syntax +//! error is `Err(())`. +//! +//! Codes are keyed by their numeric value, whatever the length of the source string, as pdf.js +//! (`CMap.mapOne`) and Poppler (`CharCodeToUnicode::parseCMap1`) key them: `<0041>` maps code +//! 0x41 of a simple font, `<41>` code 0x0041 of a Type0 font. A later mapping of a code wins. + +use super::Code; +use crate::pdf_engine::text_edit::lexer::{scan_tokens, Operand, ScanMode, Token}; +use crate::pdf_engine::text_edit::limits::{CMAP_MAPPINGS_MAX, TOUNICODE_MAX_DECODED}; +use std::collections::BTreeMap; + +/// Destinations longer than this are a syntax error. +const DESTINATION_MAX_BYTES: usize = 512; +/// Tokens scanned from one ToUnicode stream (each ~56 bytes while it is parsed). Real CMaps need +/// 2 (`bfchar`), 3 (`bfrange`) or 5 (a one-destination `bfrange` array) per mapping, so within +/// `TOUNICODE_MAX_DECODED` bytes the densest real form (23-byte array lines) stays under 460,000. +const TOKENS_MAX: usize = 4 * CMAP_MAPPINGS_MAX; + +/// What a ToUnicode CMap says about one code. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum Lookup<'a> { + /// The CMap does not map the code (and has no `usecmap`). + Absent, + /// Mapped, but to nothing usable (lone surrogate, odd/empty destination), or possibly mapped + /// by a `usecmap` parent this parser does not load. + Undecodable, + Text(&'a str), +} + +#[derive(Debug, Clone, Default)] +pub struct ToUnicode { + map: BTreeMap>, + uses_parent: bool, +} + +impl ToUnicode { + /// What the CMap says about the code with numeric value `code`. + pub fn lookup(&self, code: u32) -> Lookup<'_> { + match self.map.get(&code) { + Some(Some(text)) => Lookup::Text(text), + Some(None) => Lookup::Undecodable, + None if self.uses_parent => Lookup::Undecodable, + None => Lookup::Absent, + } + } + + /// The value of every code with an explicit mapping, ascending. + pub fn codes(&self) -> impl Iterator + '_ { + self.map.keys().copied() + } +} + +/// Parses a decoded ToUnicode stream. +pub fn parse_tounicode(bytes: &[u8]) -> Result { + let tokens = scan_tokens(bytes, ScanMode::CMap, TOKENS_MAX).map_err(|_| ())?; + let mut parser = Parser { + tokens: &tokens, + pos: 0, + out: ToUnicode::default(), + count: 0, + dest_bytes: 0, + }; + while let Some(token) = parser.next() { + match keyword(token) { + Some(b"begincodespacerange") => parser.codespace()?, + Some(b"beginbfchar") => parser.bfchar()?, + Some(b"beginbfrange") => parser.bfrange()?, + Some(b"usecmap") => parser.out.uses_parent = true, + Some(b"begincidrange" | b"begincidchar" | b"beginnotdefrange" | b"beginnotdefchar") => { + parser.skip_section()? + } + _ => {} + } + } + Ok(parser.out) +} + +fn keyword(token: &Token) -> Option<&[u8]> { + match token { + Token::Keyword { bytes, .. } => Some(bytes), + _ => None, + } +} + +fn string(token: &Token) -> Option<&[u8]> { + match token { + Token::Operand(Operand::Str { bytes, .. }) => Some(bytes), + _ => None, + } +} + +/// A source code: 1–4 bytes, big-endian. +fn code_of(bytes: &[u8]) -> Result { + if bytes.is_empty() || bytes.len() > 4 { + return Err(()); + } + let value = bytes.iter().fold(0u32, |acc, b| (acc << 8) | u32::from(*b)); + let len = u8::try_from(bytes.len()).map_err(|_| ())?; + Ok(Code { value, len }) +} + +/// UTF-16BE units of a destination (`None` for an odd or empty one). +fn units(dst: &[u8]) -> Result>, ()> { + if dst.len() > DESTINATION_MAX_BYTES { + return Err(()); + } + if dst.is_empty() || dst.len() % 2 != 0 { + return Ok(None); + } + Ok(Some( + dst.chunks_exact(2) + .map(|p| match p { + [hi, lo] => u16::from_be_bytes([*hi, *lo]), + _ => 0, + }) + .collect(), + )) +} + +/// Decodes UTF-16 units; a lone surrogate makes the whole destination undecodable. +fn text_of(units: &[u16]) -> Option { + char::decode_utf16(units.iter().copied()) + .collect::>() + .ok() +} + +struct Parser<'t> { + tokens: &'t [Token], + pos: usize, + out: ToUnicode, + count: usize, + /// Destination bytes of every mapping so far (a range counts its destination per code). + dest_bytes: usize, +} + +impl<'t> Parser<'t> { + fn next(&mut self) -> Option<&'t Token> { + let t = self.tokens.get(self.pos)?; + self.pos += 1; + Some(t) + } + + fn take_string(&mut self) -> Result<&'t [u8], ()> { + self.next().and_then(string).ok_or(()) + } + + /// True (and consumed) when the next token is the keyword `end`. + fn at_end(&mut self, end: &[u8]) -> Result { + let token = self.tokens.get(self.pos).ok_or(())?; + if keyword(token) == Some(end) { + self.pos += 1; + return Ok(true); + } + Ok(false) + } + + fn insert(&mut self, code: Code, text: Option) -> Result<(), ()> { + self.count = self.count.checked_add(1).ok_or(())?; + if self.count > CMAP_MAPPINGS_MAX { + return Err(()); + } + self.out.map.insert(code.value, text); + Ok(()) + } + + /// Debits `mappings` destinations of `dest_len` bytes (before any of them is built). + fn reserve_text(&mut self, mappings: u32, dest_len: usize) -> Result<(), ()> { + let bytes = usize::try_from(mappings) + .map_err(|_| ())? + .checked_mul(dest_len) + .ok_or(())?; + self.dest_bytes = self.dest_bytes.checked_add(bytes).ok_or(())?; + if self.dest_bytes > TOUNICODE_MAX_DECODED { + return Err(()); + } + Ok(()) + } + + fn codespace(&mut self) -> Result<(), ()> { + while !self.at_end(b"endcodespacerange")? { + let lo = self.take_string()?; + let hi = self.take_string()?; + if lo.is_empty() || lo.len() > 4 || lo.len() != hi.len() { + return Err(()); + } + } + Ok(()) + } + + fn bfchar(&mut self) -> Result<(), ()> { + while !self.at_end(b"endbfchar")? { + let code = code_of(self.take_string()?)?; + let dst = self.take_string()?; + self.reserve_text(1, dst.len())?; + let text = units(dst)?.and_then(|u| text_of(&u)); + self.insert(code, text)?; + } + Ok(()) + } + + fn bfrange(&mut self) -> Result<(), ()> { + while !self.at_end(b"endbfrange")? { + let lo = code_of(self.take_string()?)?; + let hi = code_of(self.take_string()?)?; + if lo.len != hi.len || lo.value > hi.value { + return Err(()); + } + let count = hi.value - lo.value; + let total = self + .count + .checked_add(usize::try_from(count).map_err(|_| ())?) + .ok_or(())?; + if total >= CMAP_MAPPINGS_MAX { + return Err(()); + } + match self.next().ok_or(())? { + Token::ArrayOpen(_) => self.range_array(lo, count)?, + token => { + let dst = string(token).ok_or(())?; + self.reserve_text(count.checked_add(1).ok_or(())?, dst.len())?; + self.range_increment(lo, count, units(dst)?)?; + } + } + } + Ok(()) + } + + fn range_increment(&mut self, lo: Code, count: u32, base: Option>) -> Result<(), ()> { + for k in 0..=count { + let code = Code { + value: lo.value + k, + len: lo.len, + }; + let Some(units) = base.as_ref() else { + self.insert(code, None)?; + continue; + }; + let (last, head) = units.split_last().ok_or(())?; + let Some(next) = u32::from(*last) + .checked_add(k) + .and_then(|v| u16::try_from(v).ok()) + else { + break; // overflow past 0xFFFF stops the range + }; + let mut shifted = head.to_vec(); + shifted.push(next); + self.insert(code, text_of(&shifted))?; + } + Ok(()) + } + + fn range_array(&mut self, lo: Code, count: u32) -> Result<(), ()> { + let mut k: u32 = 0; + loop { + let token = self.next().ok_or(())?; + if matches!(token, Token::ArrayClose(_)) { + return Ok(()); + } + let dst = string(token).ok_or(())?; + self.reserve_text(1, dst.len())?; + let text = units(dst)?.and_then(|u| text_of(&u)); + if k <= count { + let code = Code { + value: lo.value + k, + len: lo.len, + }; + self.insert(code, text)?; + } + k = k.saturating_add(1); + } + } + + /// CID/notdef sections do not belong in a ToUnicode CMap: their content maps nothing here. + fn skip_section(&mut self) -> Result<(), ()> { + loop { + let token = self.next().ok_or(())?; + if keyword(token).is_some_and(|k| k.starts_with(b"end")) { + return Ok(()); + } + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/type0.rs b/src-tauri/src/pdf_engine/text_edit/fonts/type0.rs new file mode 100644 index 0000000..e0d0add --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/type0.rs @@ -0,0 +1,496 @@ +//! Type0 fonts (SPEC §A.2, §B.9): `/Identity-H` with a `CIDFontType2` (`FontFile2`, or +//! `FontFile3 /OpenType`) or `CIDFontType0` (`FontFile3 /CIDFontType0C` or `/OpenType`) +//! descendant. Codes are 2 bytes, CID = code. Text comes from ToUnicode only (§A.3.1 item 4); +//! widths from `/W` (both forms) else `/DW` (default 1000; present but not a number → +//! `FONT_UNSUPPORTED`); glyphs through `/CIDToGIDMap` (CIDFontType2) or the CFF's CID → GID map +//! (CID-keyed) / GID = CID (name-keyed). +//! +//! Only codes that ToUnicode maps (any source length, value ≤ 0xFFFF) get a `CodeInfo` (no other +//! code can be decoded or typed); `FontModel::width` answers every other code from the `/W` +//! table. + +use super::encodings::{display_text, usable_text, writable_char}; +use super::glyph_budget::WorkMeter; +use super::program::{cff_cid_to_gid, face_metrics, OutlineMemo, Outlines}; +use super::simple::{load_tounicode, TuState}; +use super::tounicode::Lookup; +use super::{ + is_whitespace_char, number_of, vertical_metrics, Code, CodeInfo, Descriptor, FontClass, + FontModel, Loader, +}; +use crate::pdf_engine::text_edit::limits::{ + CIDTOGID_MAX_BYTES, FONT_PROGRAM_MAX_DECODED, W_ENTRIES_MAX, +}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use lopdf::{Dictionary, Object, ObjectId, Stream}; +use std::collections::BTreeMap; +use ttf_parser::Face; + +/// Largest CID an Identity-H code can name. +const CID_MAX: u32 = 0xFFFF; +const DEFAULT_DW: f64 = 1000.0; + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum CidKind { + Type0, + Type2, +} + +enum CidToGid { + Identity, + Map(Vec), +} + +pub(super) fn load<'a>( + loader: &mut Loader<'a, '_>, + mut model: FontModel, + dict: &'a Dictionary, +) -> FontModel { + let type0_base = + String::from_utf8_lossy(loader.name(dict, b"BaseFont").unwrap_or_default()).into_owned(); + let descendant = match loader.get(dict, b"DescendantFonts") { + Some(Object::Array(items)) => { + items + .first() + .and_then(|d| loader.resolve(d)) + .and_then(|(_, o)| match o { + Object::Dictionary(d) => Some(d), + _ => None, + }) + } + _ => None, + }; + // The encoding is checked first so that a vertical font says so whatever else is wrong. + if let Err(reason) = check_encoding(loader, dict, descendant) { + if reason == TextReason::Vertical { + model.vertical = true; + } + model.refuse(reason); + } + let Some(cid_font) = descendant else { + model.set_names(&strip_cmap_suffix(&type0_base), &Default::default()); + model.refuse(TextReason::FontUnsupported); + return model; + }; + // The descendant's BaseFont is the font's own name (Type0 names may carry "-Identity-H"). + let raw_base = loader + .name(cid_font, b"BaseFont") + .map(|n| String::from_utf8_lossy(n).into_owned()) + .unwrap_or_else(|| strip_cmap_suffix(&type0_base)); + let desc = match Descriptor::read(loader, cid_font) { + Ok(d) => d, + Err(reason) => { + model.set_names(&raw_base, &Default::default()); + model.refuse(reason); + return model; + } + }; + model.set_names(&raw_base, &desc.hints()); + let (ascent, descent) = vertical_metrics(&desc, None, None); + model.ascent = ascent; + model.descent = descent; + let kind = match loader.name(cid_font, b"Subtype") { + Some(b"CIDFontType2") => CidKind::Type2, + Some(b"CIDFontType0") => CidKind::Type0, + _ => { + model.refuse(TextReason::FontUnsupported); + return model; + } + }; + match read_w(loader, cid_font) { + Ok(widths) => model.cid_widths = widths, + Err(reason) => model.refuse(reason), + } + let default_width = match loader.get(cid_font, b"DW") { + None => Some(DEFAULT_DW), + Some(dw) => number_of(dw), + }; + match default_width { + Some(dw) => model.default_width = dw, + None => { + model.default_width = DEFAULT_DW; + model.refuse(TextReason::FontUnsupported); + } + } + let cid_to_gid = if kind == CidKind::Type2 { + match read_cid_to_gid(loader, cid_font) { + Ok(map) => Some(map), + Err(reason) => { + model.refuse(reason); + None + } + } + } else { + None + }; + let tu = load_tounicode(loader, dict); + match &tu { + TuState::Absent => model.refuse(TextReason::NoTounicode), + TuState::Broken => model.refuse(TextReason::AmbiguousUnicode), + TuState::Ready(_) => {} + } + let desc_dict = loader.dict(cid_font, b"FontDescriptor"); + let program = match select_program(loader, desc_dict, kind) { + Ok(Some(p)) => Some(p), + Ok(None) => { + model.refuse(TextReason::FontNotEmbedded); + None + } + Err(reason) => { + model.embedded = true; + model.refuse(reason); + None + } + }; + model.class = Some(match kind { + CidKind::Type2 => FontClass::Type0Cid2, + CidKind::Type0 => FontClass::Type0Cid0, + }); + let TuState::Ready(tu) = tu else { + return model; + }; + let values: Vec = tu.codes().filter(|v| *v <= CID_MAX).collect(); + let mut infos: Vec<(u32, Option, f64)> = Vec::with_capacity(values.len()); + for value in values { + let raw = match tu.lookup(value) { + Lookup::Text(t) => Some(t.to_string()), + Lookup::Undecodable | Lookup::Absent => None, + }; + infos.push((value, raw, model.width(Code { value, len: 2 }))); + } + let presence = match program { + Some((id, stream, is_open_type)) => { + model.embedded = true; + match loader.stream_bytes(id, stream, FONT_PROGRAM_MAX_DECODED) { + Ok(bytes) => { + let gids = Glyphs::new(&bytes, is_open_type, kind, cid_to_gid.as_ref()); + match gids { + Some(glyphs) => { + if let Some(metrics) = glyphs.metrics { + let (a, d) = vertical_metrics(&desc, Some(metrics), None); + model.ascent = a; + model.descent = d; + } + if let Some(reason) = glyphs.outlines.refusal() { + model.refuse(reason); + } + let mut memo = OutlineMemo::default(); + let mut meter = loader.work_meter(); + let found = infos + .iter() + .map(|(cid, raw, width)| { + let raw = raw.as_deref(); + glyphs.presence(&mut memo, &mut meter, *cid, raw, *width) + }) + .collect(); + loader.settle(&meter); + found + } + None => { + model.refuse(TextReason::FontProgramUnreadable); + Vec::new() + } + } + } + Err(_) => { + model.refuse(TextReason::FontProgramUnreadable); + Vec::new() + } + } + } + None => Vec::new(), + }; + let mut codes = Vec::with_capacity(infos.len()); + for (i, (cid, raw, width)) in infos.into_iter().enumerate() { + let (gid, drawable) = presence.get(i).copied().unwrap_or((None, false)); + let text = raw.as_deref().filter(|t| usable_text(t)).map(display_text); + let single = raw.as_deref().and_then(|t| { + let mut chars = t.chars(); + match (chars.next(), chars.next()) { + (Some(c), None) => Some(c), + _ => None, + } + }); + let typeable_as = + single.filter(|ch| text.is_some() && writable_char(*ch) && drawable && width > 0.0); + codes.push(( + cid, + CodeInfo { + text, + glyph_name: None, + gid, + width1000: Some(width), + drawable, + typeable_as, + }, + )); + } + model.codes = codes.into_iter().collect(); + model +} + +/// `ABCDEF+Calibri-Identity-H` → `ABCDEF+Calibri`. +fn strip_cmap_suffix(name: &str) -> String { + for suffix in ["-Identity-H", "-Identity-V"] { + if let Some(stripped) = name.strip_suffix(suffix) { + return stripped.to_string(); + } + } + name.to_string() +} + +/// `/Identity-H` only; `-V` names, `/WMode 1` (of the CIDFont, when there is one, or of an +/// embedded CMap) → `VERTICAL`; other CMaps → `UNSUPPORTED_ENCODING`. +fn check_encoding<'a>( + loader: &Loader<'a, '_>, + dict: &'a Dictionary, + cid_font: Option<&'a Dictionary>, +) -> Result<(), TextReason> { + let wmode = |d: &'a Dictionary| matches!(loader.get(d, b"WMode"), Some(Object::Integer(1))); + if cid_font.is_some_and(wmode) { + return Err(TextReason::Vertical); + } + match loader.get(dict, b"Encoding") { + Some(Object::Name(name)) if name.as_slice() == b"Identity-H" => Ok(()), + Some(Object::Name(name)) if name.ends_with(b"-V") => Err(TextReason::Vertical), + Some(Object::Name(_)) => Err(TextReason::UnsupportedEncoding), + Some(Object::Stream(s)) if wmode(&s.dict) => Err(TextReason::Vertical), + Some(Object::Stream(_)) => Err(TextReason::UnsupportedEncoding), + _ => Err(TextReason::FontUnsupported), + } +} + +/// `/W`: `c [w1 w2 …]` and `cfirst clast w`; each span ≤ 65,535 CIDs, at most `W_ENTRIES_MAX` +/// CIDs described in total; later entries win. Returns `(cid, width)` sorted by CID. +fn read_w<'a>( + loader: &Loader<'a, '_>, + cid_font: &'a Dictionary, +) -> Result, TextReason> { + let present = cid_font.get(b"W").is_ok_and(|w| !matches!(w, Object::Null)); + if !present { + return Ok(Vec::new()); + } + let Some(Object::Array(items)) = loader.get(cid_font, b"W") else { + return Err(TextReason::FontUnsupported); + }; + let resolved = |o: &'a Object| loader.resolve(o).map(|(_, v)| v); + let cid = |o: &'a Object| match resolved(o) { + Some(Object::Integer(v)) => u32::try_from(*v).ok(), + _ => None, + }; + let mut map: BTreeMap = BTreeMap::new(); + let mut entries: usize = 0; + let mut add = |c: u32, w: f64, map: &mut BTreeMap| -> Result<(), TextReason> { + entries += 1; + if entries > W_ENTRIES_MAX { + return Err(TextReason::FontUnsupported); + } + if c <= CID_MAX { + map.insert(c, w); + } + Ok(()) + }; + let mut i = 0; + while let Some(first) = items.get(i) { + let start = cid(first).ok_or(TextReason::FontUnsupported)?; + match items.get(i + 1).and_then(resolved) { + Some(Object::Array(ws)) => { + for (k, w) in ws.iter().enumerate() { + let w = resolved(w) + .and_then(number_of) + .ok_or(TextReason::FontUnsupported)?; + let k = u32::try_from(k).map_err(|_| TextReason::FontUnsupported)?; + let c = start.checked_add(k).ok_or(TextReason::FontUnsupported)?; + add(c, w, &mut map)?; + } + i += 2; + } + Some(_) => { + let end = items + .get(i + 1) + .and_then(cid) + .ok_or(TextReason::FontUnsupported)?; + let w = items + .get(i + 2) + .and_then(resolved) + .and_then(number_of) + .ok_or(TextReason::FontUnsupported)?; + if end < start || end - start > 65_535 { + return Err(TextReason::FontUnsupported); + } + for c in start..=end { + add(c, w, &mut map)?; + } + i += 3; + } + None => return Err(TextReason::FontUnsupported), + } + } + Ok(map.into_iter().collect()) +} + +/// `/CIDToGIDMap`: absent or `/Identity`, or a stream of big-endian u16 GIDs indexed by CID. +fn read_cid_to_gid<'a>( + loader: &mut Loader<'a, '_>, + cid_font: &'a Dictionary, +) -> Result { + let present = cid_font + .get(b"CIDToGIDMap") + .is_ok_and(|m| !matches!(m, Object::Null)); + if !present { + return Ok(CidToGid::Identity); + } + match loader.get(cid_font, b"CIDToGIDMap") { + Some(Object::Name(n)) if n.as_slice() == b"Identity" => Ok(CidToGid::Identity), + Some(Object::Stream(_)) => { + let (id, stream) = loader + .stream(cid_font, b"CIDToGIDMap") + .ok_or(TextReason::FontUnsupported)?; + let bytes = loader + .stream_bytes(id, stream, CIDTOGID_MAX_BYTES) + .map_err(|_| TextReason::FontUnsupported)?; + Ok(CidToGid::Map(bytes.to_vec())) + } + _ => Err(TextReason::FontUnsupported), + } +} + +/// `(id, stream, is OpenType)` of the descendant's program; a program that does not match the +/// CIDFont type is `FONT_PROGRAM_UNSUPPORTED`, one that is not a stream `FONT_PROGRAM_UNREADABLE`. +fn select_program<'a>( + loader: &Loader<'a, '_>, + desc: Option<&'a Dictionary>, + kind: CidKind, +) -> Result, TextReason> { + let Some(d) = desc else { + return Ok(None); + }; + let mut found = Vec::new(); + for key in [&b"FontFile"[..], b"FontFile2", b"FontFile3"] { + let present = d.get(key).is_ok_and(|v| !matches!(v, Object::Null)); + if present { + let (id, stream) = loader + .stream(d, key) + .ok_or(TextReason::FontProgramUnreadable)?; + found.push((key, id, stream)); + } + } + let [(key, id, stream)] = found.as_slice() else { + return if found.is_empty() { + Ok(None) + } else { + Err(TextReason::FontProgramUnsupported) + }; + }; + let open_type = match (*key, kind) { + (b"FontFile2", CidKind::Type2) => false, + (b"FontFile3", _) => match (loader.name(&stream.dict, b"Subtype"), kind) { + (Some(b"CIDFontType0C"), CidKind::Type0) => false, + (Some(b"OpenType"), _) => true, + _ => return Err(TextReason::FontProgramUnsupported), + }, + _ => return Err(TextReason::FontProgramUnsupported), + }; + Ok(Some((*id, *stream, open_type))) +} + +/// A parsed descendant program with its CID → GID rule. +struct Glyphs<'p> { + outlines: Outlines<'p>, + map: GidMap<'p>, + /// GIDs ≥ this are not glyphs (`maxp` for TrueType/OpenType, the CharStrings count for CFF). + count: u16, + metrics: Option<(f64, f64)>, +} + +enum GidMap<'p> { + Identity, + Stream(&'p [u8]), + Cids(BTreeMap), +} + +impl<'p> GidMap<'p> { + fn gid(&self, cid: u32) -> Option { + match self { + GidMap::Identity => u16::try_from(cid).ok(), + GidMap::Stream(bytes) => { + let at = usize::try_from(cid).ok()?.checked_mul(2)?; + match bytes.get(at..at.checked_add(2)?)? { + [hi, lo] => Some(u16::from_be_bytes([*hi, *lo])), + _ => None, + } + } + GidMap::Cids(map) => map.get(&u16::try_from(cid).ok()?).copied(), + } + } + + /// CID-keyed CFF: the inverse of the charset; name-keyed: GID = CID. + fn of_cff(outlines: &Outlines<'_>) -> GidMap<'p> { + outlines + .cff_layout() + .and_then(cff_cid_to_gid) + .map_or(GidMap::Identity, GidMap::Cids) + } +} + +impl<'p> Glyphs<'p> { + fn new( + bytes: &'p [u8], + open_type: bool, + kind: CidKind, + cid_to_gid: Option<&'p CidToGid>, + ) -> Option> { + if open_type || kind == CidKind::Type2 { + let face = Face::parse(bytes, 0).ok()?; + let outlines = Outlines::of_face(&face); + let map = match kind { + CidKind::Type2 => match cid_to_gid { + Some(CidToGid::Map(m)) => GidMap::Stream(m), + _ => GidMap::Identity, + }, + // CIDFontType0 in OpenType: the CFF table carries the glyphs. + CidKind::Type0 => { + face.tables().cff?; + GidMap::of_cff(&outlines) + } + }; + return Some(Glyphs { + metrics: face_metrics(&face), + count: face.number_of_glyphs(), + outlines, + map, + }); + } + let outlines = Outlines::of_cff(bytes)?; + let count = match &outlines { + Outlines::Cff { table, .. } => table.number_of_glyphs(), + _ => 0, + }; + Some(Glyphs { + map: GidMap::of_cff(&outlines), + outlines, + count, + metrics: None, + }) + } + + /// `(gid, drawable)`: GID ≠ 0, < the glyph count, outline present (or a whitespace glyph + /// with a positive width); outline work is charged to `meter`. + fn presence( + &self, + memo: &mut OutlineMemo, + meter: &mut WorkMeter, + cid: u32, + raw: Option<&str>, + width: f64, + ) -> (Option, bool) { + let blank_ok = is_whitespace_char(raw) && width > 0.0; + match self.map.gid(cid) { + Some(g) if g != 0 && g < self.count => { + let drawn = memo.get(g, || self.outlines.drawn(g, meter)); + (Some(g), drawn || blank_ok) + } + _ => (None, false), + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/fonts/type1.rs b/src-tauri/src/pdf_engine/text_edit/fonts/type1.rs new file mode 100644 index 0000000..9ade838 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/fonts/type1.rs @@ -0,0 +1,608 @@ +//! Embedded Type1 programs (`FontFile`, SPEC §A.2, §B.9.5): the clear-text built-in `/Encoding`, +//! the eexec layer and a bounded charstring interpreter that proves a glyph draws something. +//! +//! - `Length1` bytes of clear text, then `Length2` bytes of eexec data (PFB segment headers do not +//! exist inside a PDF `FontFile`). A missing or out-of-range `Length2` is `Err`. +//! - eexec: hex form when the first 4 bytes (after optional whitespace) are hex digits — decoded +//! whitespace-tolerantly first; then r = 55665, c1 = 52845, c2 = 22719, and the first 4 +//! plaintext bytes are always dropped. +//! - `/Subrs` and `/CharStrings` binary entries follow `RD`/`-|` after one byte; charstrings use +//! r = 4330 and drop `/lenIV` bytes (default 4; `-1` = not encrypted). Every entry is decrypted +//! in place once, when the program is parsed (the entries never overlap), so a `callsubr` costs +//! no decryption; an entry longer than 65,535 bytes (the Type1 charstring limit) never runs. +//! - Interpreter: numbers, `hsbw`, `sbw`, path operators, `closepath`, `callsubr`/`return` +//! (depth ≤ 10), `seac` (exactly its 5 operands; present when both StandardEncoding components +//! are), `endchar`, plus the operators that never draw (`hstem`, `vstem`, `hstem3`, `vstem3`, +//! `dotsection`, `div`, `callothersubr`, `pop`, `setcurrentpoint`) so hinted real fonts can be +//! proven; ≤ 4,096 operations and ≤ `GLYPH_WORK_MAX` tokens per glyph (its `seac` components +//! included, and never more than the load's `WorkMeter` can still pay), charged to that meter. +//! Anything else leaves that glyph unproven (never a font refusal). + +use super::encodings::standard_name; +use super::glyph_budget::WorkMeter; +use crate::pdf_engine::text_edit::lexer::{scan_tokens, Operand, ScanMode, Token}; +use std::collections::{BTreeMap, HashSet}; +use std::ops::Range; + +const EEXEC_R: u16 = 55665; +const CHARSTRING_R: u16 = 4330; +const C1: u16 = 52845; +const C2: u16 = 22719; +const OPS_PER_GLYPH_MAX: usize = 4_096; +const SUBR_DEPTH_MAX: usize = 10; +const STACK_MAX: usize = 48; +const ENTRIES_MAX: usize = 65_536; +/// Tokens scanned from the clear text (a real one has a few hundred; a custom `/Encoding` array +/// needs 4 per code). Each token is ~56 bytes while the encoding is read. +const CLEAR_TOKENS_MAX: usize = 1 << 16; +/// Longest charstring or subroutine the Type1 format allows (Adobe Type 1 Font Format, App. B). +const CHARSTRING_BYTES_MAX: usize = 65_535; + +/// The program's own encoding (clear text), used when the PDF gives no `/Encoding`. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum Type1Encoding { + Standard, + Custom(BTreeMap), +} + +/// What the interpreter proved about one glyph. +#[derive(Debug, Clone, Copy, PartialEq)] +pub enum GlyphProof { + /// `hsbw`/`sbw`, then at least one drawing operator before `endchar`. + Drawn, + /// `hsbw`/`sbw` … `endchar` with nothing drawn (whitespace glyphs), and its advance width. + Blank { + width: f64, + }, + Unproven, +} + +pub struct Type1Program { + pub builtin: Option, + /// The decrypted private part, with every `/Subrs` and `/CharStrings` entry decrypted in + /// place; the ranges below are their plaintext (`lenIV` bytes skipped). + private: Vec, + charstrings: BTreeMap>, + subrs: BTreeMap>, +} + +/// Parses a decoded `FontFile` stream with its `Length1`/`Length2`. +pub fn parse_type1(data: &[u8], length1: usize, length2: usize) -> Result { + let end = length1.checked_add(length2).ok_or(())?; + let clear = data.get(..length1).ok_or(())?; + let mut encrypted = data.get(length1..end).ok_or(())?; + if length2 == 0 { + return Err(()); + } + // `Length1` that stops right after `eexec` leaves the end-of-line in the encrypted part. + if clear.ends_with(b"eexec") { + while let Some((first, rest)) = encrypted.split_first() { + if !matches!(first, b'\r' | b'\n' | b' ' | b'\t') { + break; + } + encrypted = rest; + } + } + let cipher = if is_hex_form(encrypted) { + hex_decode(encrypted) + } else { + encrypted.to_vec() + }; + let plain = decrypt(&cipher, EEXEC_R); + let private = plain.get(4..).ok_or(())?.to_vec(); + let mut program = Type1Program { + builtin: builtin_encoding(clear), + private, + charstrings: BTreeMap::new(), + subrs: BTreeMap::new(), + }; + program.index_private()?; + Ok(program) +} + +fn is_hex_form(data: &[u8]) -> bool { + let start = data + .iter() + .position(|b| !matches!(b, b'\r' | b'\n' | b' ' | b'\t')) + .unwrap_or(data.len()); + data.get(start..start.saturating_add(4)) + .is_some_and(|four| four.len() == 4 && four.iter().all(u8::is_ascii_hexdigit)) +} + +/// Whitespace-tolerant hex decoding; stops at the first other byte (the `cleartomark` trailer). +fn hex_decode(data: &[u8]) -> Vec { + let mut out = Vec::with_capacity(data.len() / 2); + let mut high: Option = None; + for b in data { + let nibble = match b { + b'0'..=b'9' => b - b'0', + b'a'..=b'f' => b - b'a' + 10, + b'A'..=b'F' => b - b'A' + 10, + b'\r' | b'\n' | b' ' | b'\t' | b'\x0c' | 0 => continue, + _ => break, + }; + match high.take() { + Some(h) => out.push((h << 4) | nibble), + None => high = Some(nibble), + } + } + out +} + +/// Type1 decryption (eexec r = 55665, charstrings r = 4330). +fn decrypt(cipher: &[u8], r: u16) -> Vec { + let mut plain = cipher.to_vec(); + decrypt_in_place(&mut plain, r); + plain +} + +fn decrypt_in_place(data: &mut [u8], r: u16) { + let mut r = r; + for byte in data.iter_mut() { + let c = *byte; + *byte = c ^ (r >> 8) as u8; + r = (u16::from(c).wrapping_add(r)) + .wrapping_mul(C1) + .wrapping_add(C2); + } +} + +/// `/Encoding StandardEncoding def` or `/Encoding 256 array … dup / put … def`. +fn builtin_encoding(clear: &[u8]) -> Option { + let tokens = scan_tokens(clear, ScanMode::Type1Clear, CLEAR_TOKENS_MAX).ok()?; + let start = tokens.iter().position(|t| { + matches!(t, Token::Operand(Operand::Name { bytes, .. }) if bytes.as_slice() == b"Encoding") + })?; + let rest = tokens.get(start + 1..)?; + if let Some(Token::Keyword { bytes, .. }) = rest.first() { + if bytes.as_slice() == b"StandardEncoding" { + return Some(Type1Encoding::Standard); + } + } + let mut map = BTreeMap::new(); + for (i, token) in rest.iter().enumerate() { + if matches!(token, Token::Keyword { bytes, .. } if bytes.as_slice() == b"def") { + break; + } + let window = ( + rest.get(i), + rest.get(i + 1), + rest.get(i + 2), + rest.get(i + 3), + ); + if let ( + Some(Token::Keyword { bytes: dup, .. }), + Some(Token::Operand(Operand::Number { value, .. })), + Some(Token::Operand(Operand::Name { bytes: name, .. })), + Some(Token::Keyword { bytes: put, .. }), + ) = window + { + let code = + (*value >= 0.0 && *value <= 255.0 && value.fract() == 0.0).then_some(*value as u8); + if let (b"dup", b"put", Some(code), Ok(name)) = ( + dup.as_slice(), + put.as_slice(), + code, + std::str::from_utf8(name), + ) { + map.insert(code, name.to_string()); + } + } + } + Some(Type1Encoding::Custom(map)) +} + +#[derive(Debug, Clone, PartialEq)] +enum Tok { + Name(Vec), + Int(i64), + Other, +} + +#[derive(Clone, Copy, PartialEq)] +enum Section { + None, + Subrs, + CharStrings, +} + +fn is_ws(b: u8) -> bool { + matches!(b, 0 | 9 | 10 | 12 | 13 | 32) +} + +fn is_delim(b: u8) -> bool { + matches!( + b, + b'(' | b')' | b'<' | b'>' | b'[' | b']' | b'{' | b'}' | b'/' | b'%' + ) +} + +impl Type1Program { + /// Indexes `/lenIV`, `/Subrs` and `/CharStrings` of the decrypted private part, then + /// decrypts every entry in place. + fn index_private(&mut self) -> Result<(), ()> { + let p = &self.private; + let (mut pos, mut section) = (0usize, Section::None); + let (mut prev2, mut prev1) = (Tok::Other, Tok::Other); + let mut subrs = BTreeMap::new(); + let mut charstrings = BTreeMap::new(); + let mut len_iv: i64 = 4; + while let Some(&c) = p.get(pos) { + if is_ws(c) { + pos += 1; + continue; + } + let start = pos; + let tok = match c { + b'%' => { + while p.get(pos).is_some_and(|b| *b != b'\n' && *b != b'\r') { + pos += 1; + } + continue; + } + b'/' => { + pos += 1; + while p.get(pos).is_some_and(|b| !is_ws(*b) && !is_delim(*b)) { + pos += 1; + } + let name = p.get(start + 1..pos).ok_or(())?.to_vec(); + match name.as_slice() { + b"Subrs" => section = Section::Subrs, + b"CharStrings" => section = Section::CharStrings, + _ => {} + } + Tok::Name(name) + } + b'(' => { + pos = skip_literal(p, pos); + Tok::Other + } + b'<' | b'>' | b'[' | b']' | b'{' | b'}' | b')' => { + pos += 1; + Tok::Other + } + _ => { + while p.get(pos).is_some_and(|b| !is_ws(*b) && !is_delim(*b)) { + pos += 1; + } + let word = p.get(start..pos).ok_or(())?; + if word == b"RD" || word == b"-|" { + let Tok::Int(n) = prev1 else { + return Err(()); + }; + let len = usize::try_from(n).map_err(|_| ())?; + let data_start = pos.checked_add(1).ok_or(())?; + let data_end = data_start.checked_add(len).ok_or(())?; + if data_end > p.len() { + return Err(()); + } + match (section, &prev2) { + (Section::Subrs, Tok::Int(idx)) => { + let idx = usize::try_from(*idx).map_err(|_| ())?; + subrs.insert(idx, data_start..data_end); + } + (Section::CharStrings, Tok::Name(name)) => { + if let Ok(name) = String::from_utf8(name.clone()) { + charstrings.insert(name, data_start..data_end); + } + } + _ => {} + } + if subrs.len() > ENTRIES_MAX || charstrings.len() > ENTRIES_MAX { + return Err(()); + } + pos = data_end; + prev2 = Tok::Other; + prev1 = Tok::Other; + continue; + } + match std::str::from_utf8(word) + .ok() + .and_then(|w| w.parse::().ok()) + { + Some(v) => Tok::Int(v), + None => Tok::Other, + } + } + }; + if let (Tok::Name(n), Tok::Int(v)) = (&prev1, &tok) { + if n.as_slice() == b"lenIV" { + len_iv = *v; + } + } + prev2 = std::mem::replace(&mut prev1, tok); + } + // Entries never overlap (each `RD` skips its own bytes), so each is decrypted once. + // `lenIV` −1 (any negative value) means the charstrings are not encrypted. + let skip = usize::try_from(len_iv).ok(); + let private = &mut self.private; + let mut plain = |range: Range| -> Range { + let Some(skip) = skip else { + return range; + }; + if let Some(bytes) = private.get_mut(range.clone()) { + decrypt_in_place(bytes, CHARSTRING_R); + } + range.start.saturating_add(skip).min(range.end)..range.end + }; + self.subrs = subrs.into_iter().map(|(k, r)| (k, plain(r))).collect(); + self.charstrings = charstrings + .into_iter() + .map(|(k, r)| (k, plain(r))) + .collect(); + Ok(()) + } + + /// Whether `/CharStrings` has an entry for `name` (`.notdef` never counts). + pub fn has_charstring(&self, name: &str) -> bool { + name != ".notdef" && self.charstrings.contains_key(name) + } + + /// The plaintext of an entry (`None` past the Type1 charstring length limit). + fn code(&self, range: &Range) -> Option<&[u8]> { + if range.len() > CHARSTRING_BYTES_MAX { + return None; + } + self.private.get(range.clone()) + } + + /// Runs the bounded interpreter on `name`'s charstring; the tokens it reads are charged to + /// `meter` (a meter that cannot cover them leaves the glyph unproven). The run stops as soon + /// as it passes what the meter can still pay; an exhausted meter runs nothing. + pub fn proof(&self, name: &str, meter: &mut WorkMeter) -> GlyphProof { + if meter.exhausted() { + return GlyphProof::Unproven; + } + let mut tokens = 0u64; + let proof = self.proof_at(name, true, &mut tokens, meter.glyph_allowance()); + if meter.charge(tokens.max(1)) { + proof + } else { + GlyphProof::Unproven + } + } + + /// `limit`: the tokens the glyph may read, its `seac` components included. + fn proof_at(&self, name: &str, allow_seac: bool, tokens: &mut u64, limit: u64) -> GlyphProof { + if !self.has_charstring(name) { + return GlyphProof::Unproven; + } + let Some(code) = self.charstrings.get(name).and_then(|r| self.code(r)) else { + return GlyphProof::Unproven; + }; + let mut st = Interp { + tokens: *tokens, + limit, + ..Interp::default() + }; + let flow = self.run(code, &mut st, 0); + *tokens = st.tokens; + if !matches!(flow, Ok(Flow::End)) { + return GlyphProof::Unproven; + } + let Some(width) = st.width else { + return GlyphProof::Unproven; + }; + if let Some((base, accent)) = st.seac { + if !allow_seac { + return GlyphProof::Unproven; + } + let mut component = |c: u8| { + standard_name(c).map_or(GlyphProof::Unproven, |n| { + self.proof_at(n, false, tokens, limit) + }) + }; + return match (component(base), component(accent)) { + (GlyphProof::Drawn, GlyphProof::Drawn) => GlyphProof::Drawn, + _ => GlyphProof::Unproven, + }; + } + if st.drawn { + GlyphProof::Drawn + } else { + GlyphProof::Blank { width } + } + } + + fn run(&self, code: &[u8], st: &mut Interp, depth: usize) -> Result { + let mut pos = 0usize; + while let Some(&v) = code.get(pos) { + pos += 1; + st.tokens = st.tokens.saturating_add(1); + if st.tokens > st.limit { + return Err(()); + } + if v >= 32 { + let value = match v { + 32..=246 => i32::from(v) - 139, + 247..=250 => { + let w = i32::from(*code.get(pos).ok_or(())?); + pos += 1; + (i32::from(v) - 247) * 256 + w + 108 + } + 251..=254 => { + let w = i32::from(*code.get(pos).ok_or(())?); + pos += 1; + -(i32::from(v) - 251) * 256 - w - 108 + } + _ => { + let b = code.get(pos..pos.checked_add(4).ok_or(())?).ok_or(())?; + pos += 4; + i32::from_be_bytes([ + *b.first().ok_or(())?, + *b.get(1).ok_or(())?, + *b.get(2).ok_or(())?, + *b.get(3).ok_or(())?, + ]) + } + }; + st.push(f64::from(value))?; + continue; + } + st.ops += 1; + if st.ops > OPS_PER_GLYPH_MAX { + return Err(()); + } + let op = if v == 12 { + let e = *code.get(pos).ok_or(())?; + pos += 1; + 0x0C00 | u16::from(e) + } else { + u16::from(v) + }; + if st.width.is_none() && op != 13 && op != 0x0C07 { + return Err(()); // a charstring starts with hsbw or sbw + } + match op { + 13 => { + st.need(2)?; + st.width = st.stack.get(1).copied(); + st.stack.clear(); + } + 0x0C07 => { + st.need(4)?; + st.width = st.stack.get(2).copied(); + st.stack.clear(); + } + // rlineto, rmoveto, hstem, vstem, setcurrentpoint + 5 | 21 | 1 | 3 | 0x0C21 => st.consume(2)?, + 6 | 7 | 22 | 4 => st.consume(1)?, // hlineto/vlineto/hmoveto/vmoveto + 8 | 0x0C01 | 0x0C02 => st.consume(6)?, // rrcurveto/vstem3/hstem3 + 30 | 31 => st.consume(4)?, // vhcurveto/hvcurveto + 9 | 0x0C00 => st.stack.clear(), // closepath/dotsection + 10 => { + let index = st.pop()?; + if depth >= SUBR_DEPTH_MAX || index < 0.0 || index.fract() != 0.0 { + return Err(()); + } + let range = self.subrs.get(&(index as usize)).ok_or(())?; + let sub = self.code(range).ok_or(())?; + if let Flow::End = self.run(sub, st, depth + 1)? { + return Ok(Flow::End); + } + } + 11 => return Ok(Flow::Return), + 14 => return Ok(Flow::End), + 0x0C06 => { + // asb adx ady bchar achar seac: exactly five operands. + if st.stack.len() != 5 { + return Err(()); + } + let code_of = |x: Option<&f64>| -> Result { + let x = *x.ok_or(())?; + (x >= 0.0 && x <= 255.0 && x.fract() == 0.0) + .then_some(x as u8) + .ok_or(()) + }; + st.seac = Some((code_of(st.stack.get(3))?, code_of(st.stack.get(4))?)); + return Ok(Flow::End); + } + 0x0C0C => { + let b = st.pop()?; + let a = st.pop()?; + if b == 0.0 { + return Err(()); + } + st.push(a / b)?; + } + 0x0C10 => { + let _othersubr = st.pop()?; + let n = st.pop()?; + if n < 0.0 || n.fract() != 0.0 || n as usize > st.stack.len() { + return Err(()); + } + let keep = st.stack.len() - n as usize; + let args = st.stack.split_off(keep); + st.ps.extend(args); + if st.ps.len() > STACK_MAX { + return Err(()); + } + } + 0x0C11 => { + let v = st.ps.pop().ok_or(())?; + st.push(v)?; + } + _ => return Err(()), + } + if matches!(op, 5 | 6 | 7 | 8 | 30 | 31) { + st.drawn = true; + } + } + Err(()) // ran off the end without endchar/return + } +} + +enum Flow { + Return, + End, +} + +#[derive(Default)] +struct Interp { + stack: Vec, + ps: Vec, + ops: usize, + /// Tokens read for the glyph so far (its `seac` components included), and how many it may. + tokens: u64, + limit: u64, + width: Option, + drawn: bool, + seac: Option<(u8, u8)>, +} + +impl Interp { + fn push(&mut self, v: f64) -> Result<(), ()> { + if self.stack.len() >= STACK_MAX { + return Err(()); + } + self.stack.push(v); + Ok(()) + } + + fn pop(&mut self) -> Result { + self.stack.pop().ok_or(()) + } + + fn need(&self, n: usize) -> Result<(), ()> { + if self.stack.len() < n { + return Err(()); + } + Ok(()) + } + + fn consume(&mut self, n: usize) -> Result<(), ()> { + self.need(n)?; + self.stack.clear(); + Ok(()) + } +} + +/// Skips a balanced literal string starting at `start` (`(`); returns the offset after it. +fn skip_literal(p: &[u8], start: usize) -> usize { + let mut depth = 0usize; + let mut pos = start; + while let Some(&b) = p.get(pos) { + pos += 1; + match b { + b'\\' => pos += 1, + b'(' => depth += 1, + b')' => { + depth = depth.saturating_sub(1); + if depth == 0 { + return pos; + } + } + _ => {} + } + } + pos +} + +/// The names of a FontDescriptor `/CharSet` string (`(/a/b/c)`). +pub fn charset_names(charset: &[u8]) -> HashSet { + charset + .split(|b| *b == b'/') + .map(|n| String::from_utf8_lossy(n).trim().to_string()) + .filter(|n| !n.is_empty()) + .collect() +} diff --git a/src-tauri/src/pdf_engine/text_edit/gate.rs b/src-tauri/src/pdf_engine/text_edit/gate.rs new file mode 100644 index 0000000..91f7319 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/gate.rs @@ -0,0 +1,778 @@ +//! The source-edit publish gate (SPEC §B.16). Phase A proves each file qpdf wrote from an edit +//! plan (a Save's edited copy, or a preview page): A0 `qpdf --check` no worse than the source, +//! A1 lopdf and qpdf agree on the page map, A2 the whole object graph equals qpdf's input except +//! the edited content streams (`graph.rs`), A3 the edited page's parts are exactly the expected +//! bytes with the old content nowhere, A4 the re-walk (`verify.rs`), A5 Poppler's words and pixels +//! (`poppler.rs`). Phase B proves the final file after every later pass: the expected content is +//! present once, either as parts or inside qpdf's overlay wrapper Form (B1), the old content is +//! absent (B2), and the page's show records are exactly the proven ones (B3). Cover-and-overlay +//! fakes fail both phases (§B.16.2). No setting can skip a check (`seams` exist only in tests). + +mod final_output; +pub(crate) mod join; +pub(crate) mod originals; + +use crate::error::AppError; +use crate::pdf_engine::text_edit::apply::{benign_warnings, warnings_known}; +use crate::pdf_engine::text_edit::content::{page_content, PageContent}; +use crate::pdf_engine::text_edit::context::{Lookup, Res, SnapshotContext}; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget}; +use crate::pdf_engine::text_edit::engines::{qpdf_page_map, run_tool, Engines, RunOpts}; +use crate::pdf_engine::text_edit::fonts::Code; +use crate::pdf_engine::text_edit::geometry::PageGeometry; +use crate::pdf_engine::text_edit::graph::{graph_matches, GraphDigest}; +use crate::pdf_engine::text_edit::limits::{ + FORM_DEPTH_MAX, GATE_DECODED_TOTAL, HEAD_TAIL_HASH_BYTES, PAGE_DECODE_BUDGET, + STREAM_MAX_DECODED, +}; +use crate::pdf_engine::text_edit::poppler::{ + check_independent, pdftotext_words, render_dpi, render_page, NearGlyphs, Raster, Word, + RENDER_DPI_MIN, +}; +use crate::pdf_engine::text_edit::reasons::{ + source_edit_gate_failed, EditProblem, EditProblemCode, ProblemCtx, TextWarningCode, +}; +use crate::pdf_engine::text_edit::rewrite::PagePlan; +use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel}; +use crate::pdf_engine::text_edit::snapshot::{ + check_page_map, fnv1a_extend, fnv1a_u64, read_verification_snapshot, +}; +use crate::pdf_engine::text_edit::state::StateDigest; +use crate::pdf_engine::text_edit::verify::walk_and_verify_with; +use crate::pdf_engine::text_edit::walker::{walk_page, PageWalk, ShowOp, ShowRecord, WalkMode}; +use crate::pdf_engine::validate_output::{content_digest, ContentDigest}; +use lopdf::{Object, ObjectId}; +use std::collections::{HashSet, VecDeque}; +use std::ffi::OsString; +use std::io::Read; +use std::path::{Path, PathBuf}; +use std::sync::atomic::{AtomicBool, Ordering}; +use std::sync::{Arc, Mutex}; + +pub use originals::OriginalPart; + +/// The qpdf input file and page Poppler reads for the "before" side of A5. +#[derive(Debug, Clone)] +pub struct PopplerRef { + pub pdf: PathBuf, + pub page_1: u32, +} + +/// The model (walk + runs) of an edited page as qpdf's input has it: kept by the caller, or built +/// from the source context when Phase A reaches the page — a Save with large pages keeps the +/// source context instead of every page's model (review-T4 M-1). A model built again must have +/// the planned model's `print` (`STALE` otherwise, review-final LOW-5). +pub enum BeforeModel<'a> { + Kept(&'a PageModel), + Build { + ctx: &'a SnapshotContext, + page_index: u32, + print: &'a ModelPrint, + }, +} + +/// What identifies a page model: its page, content parts, size, records and runs (ids, texts). +#[derive(Debug, Clone, PartialEq)] +pub struct ModelPrint { + page_index: u32, + parts: Vec, + model_bytes: usize, + records: usize, + runs: u64, +} + +impl ModelPrint { + pub fn of(model: &PageModel) -> ModelPrint { + let runs = model.runs.iter().fold(fnv1a_u64(b""), |h, r| { + let h = fnv1a_extend(h, r.id.as_bytes()); + fnv1a_extend(fnv1a_extend(h, &[0]), r.text.as_bytes()) + }); + ModelPrint { + page_index: model.page_index, + parts: model.content.parts.iter().map(|p| p.digest).collect(), + model_bytes: model.walk.model_bytes, + records: model.walk.records.len(), + runs, + } + } +} + +impl<'a> From<&'a PageModel> for BeforeModel<'a> { + fn from(model: &'a PageModel) -> BeforeModel<'a> { + BeforeModel::Kept(model) + } +} + +pub struct EditedPageInput<'a> { + pub model: BeforeModel<'a>, + pub plan: &'a PagePlan, + /// Page index in qpdf's input (= in its output: the update never moves pages). + pub input_page_index: u32, + pub input_render: PopplerRef, +} + +pub struct PhaseAInput<'a> { + /// Digest of qpdf's input (the source copy, or `p.pdf` for a preview). + pub before: &'a GraphDigest, + pub before_page_count: u32, + /// qpdf's output. + pub staged: &'a Path, + /// `read_verification_snapshot` cap (D31). + pub staged_cap: u64, + pub pages: Vec>, + pub source_benign: &'a [String], +} + +/// A depth-0 show record, id-free. +#[derive(Debug, Clone)] +pub struct RecordPrint { + pub op: ShowOp, + pub font_res: Option>, + pub font_hash: u64, + pub codes: Vec, + pub text: String, + pub origins: Vec<(f64, f64)>, + pub pen_after: (f64, f64), + /// Shared with the walk's records (one allocation per distinct state, not per record). + pub state: Arc, +} + +impl RecordPrint { + pub fn of(rec: &ShowRecord) -> RecordPrint { + let font = rec.before.text.font.as_ref(); + RecordPrint { + op: rec.op, + font_res: font.and_then(|f| f.resource.as_deref()).map(<[u8]>::to_vec), + font_hash: font.map_or(0, |f| f.content_hash), + codes: rec.glyphs.iter().map(|g| g.code).collect(), + text: rec + .glyphs + .iter() + .map(|g| g.text.as_deref().unwrap_or("\u{fffd}")) + .collect(), + origins: rec.glyphs.iter().map(|g| g.origin).collect(), + pen_after: rec.pen_after, + state: Arc::clone(&rec.before), + } + } +} + +#[derive(Debug, Clone)] +pub struct PageProof { + pub source_page_index: u32, + /// Every part of the edited page, in order, after the edit. + pub expected_parts: Vec>, + /// Digests of `expected_parts` (B1 compares them before the bytes). + pub part_digests: Vec, + pub original_edited_digests: Vec, + /// Rolling hashes of the same parts (B2 finds them inside a wrapper Form's data). + pub original_edited_rolling: Vec, + /// Concatenation of the expected parts without separator (#34 semantics). + pub expected_page_digest: ContentDigest, + /// Depth-0 show records of the edited page after the edit. + pub records: Vec, +} + +impl PageProof { + /// The original edited parts as B2 searches for them. + pub fn originals(&self) -> Vec { + self.original_edited_digests + .iter() + .zip(&self.original_edited_rolling) + .map(|(digest, rolling)| OriginalPart { + digest: *digest, + rolling: *rolling, + }) + .collect() + } +} + +#[derive(Debug, Clone, PartialEq)] +pub struct TextWarning { + pub page_index: u32, + pub run_id: String, + pub code: TextWarningCode, + pub detail: Option, +} + +pub struct PhaseAReport { + pub proofs: Vec, + pub warnings: Vec, +} + +/// An `AppError` of Phase A: `code` with details `phase=A check= page= …`. +fn phase_a_error(code: EditProblemCode, check: &str, page: Option, detail: &str) -> AppError { + let page_part = page.map_or(String::new(), |n| format!(" page={n}")); + let p = EditProblem::new( + code, + Some(format!("phase=A check={check}{page_part} {detail}")), + ); + let ctx = ProblemCtx { + page_number: page, + file_name: None, + face: None, + reason: None, + }; + code.to_app_error(&p, &ctx) +} + +fn a_fail(check: &str, page: Option, detail: &str) -> AppError { + phase_a_error(EditProblemCode::EditVerifyFailed, check, page, detail) +} + +/// Wraps an engine error of a check into `EDIT_VERIFY_FAILED` (cancel and missing tools pass). +fn engine_error(e: AppError, check: &str, page: Option) -> AppError { + match e.code.as_str() { + "CANCELLED" | "ENGINE_MISSING" | "VERIFIER_MISSING" => e, + _ => a_fail( + check, + page, + &format!("{}: {}", e.code, e.details.unwrap_or(e.message)), + ), + } +} + +/// A0: `qpdf --check` exit 0, or exit 3 whose warnings the source already had. +fn check_a0( + engines: &Engines, + staged: &Path, + source_benign: &[String], + opts: &RunOpts<'_>, +) -> Result<(), AppError> { + let args = [OsString::from("--check"), staged.as_os_str().to_os_string()]; + let out = + run_tool(&engines.qpdf, &args, false, opts).map_err(|e| engine_error(e, "A0", None))?; + let stdout = String::from_utf8_lossy(&out.stdout); + match out.code { + 0 => Ok(()), + 3 => match benign_warnings(&out.stdout, &out.stderr) { + Some(lines) if warnings_known(&lines, source_benign) => Ok(()), + _ => Err(a_fail( + "A0", + None, + &format!("qpdf --check: {} {}", out.stderr.trim(), stdout.trim()), + )), + }, + code => Err(a_fail( + "A0", + None, + &format!( + "qpdf --check exited with code {code}: {}", + out.stderr.trim() + ), + )), + } +} + +/// The decoded data of every Form XObject reachable from the page's resources (depth ≤ 8). +pub(crate) fn reachable_forms( + ctx: &SnapshotContext, + page_id: ObjectId, + budget: &mut DecodeBudget, +) -> Result>, String> { + let doc = ctx.doc(); + let root = Res::of_page(doc, page_id).map_err(str::to_string)?; + let mut out = Vec::new(); + let mut seen: HashSet = HashSet::new(); + let mut queue: VecDeque<(Res<'_>, usize)> = VecDeque::from([(root, 0)]); + while let Some((res, depth)) = queue.pop_front() { + let Some(dict) = res.dict else { continue }; + let Some((_, Object::Dictionary(xobjects))) = dict + .get(b"XObject") + .ok() + .and_then(|x| crate::pdf_engine::text_edit::context::resolve(doc, x)) + else { + continue; + }; + for (name, _) in xobjects.iter() { + let Lookup::Found(e) = res.entry(doc, b"XObject", name) else { + continue; + }; + let (Some(id), Object::Stream(s)) = (e.id, e.value) else { + continue; + }; + let is_form = + s.dict.get(b"Subtype").ok().and_then(|o| o.as_name().ok()) == Some(b"Form"); + if !is_form || !seen.insert(id) { + continue; + } + let data = decode_stream(s, STREAM_MAX_DECODED, budget).map_err(|e| e.to_string())?; + out.push(data); + if depth + 1 < FORM_DEPTH_MAX { + let sub = match Res::of_form(doc, id, s) { + Some(Ok(r)) => r, + Some(Err(w)) => return Err(w.to_string()), + None => continue, + }; + queue.push_back((sub, depth + 1)); + } + } + } + Ok(out) +} + +/// A3: the original edited parts are not attached anywhere on the page (as parts, or as Form +/// XObjects reachable from its resources), the part count is unchanged and every part is exactly +/// the expected bytes. Every finding is reported together. +fn check_a3( + ctx: &SnapshotContext, + page_id: ObjectId, + content: &PageContent, + plan: &PagePlan, + originals: &[OriginalPart], + page: u32, +) -> Result<(), AppError> { + let mut problems: Vec = Vec::new(); + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + let too_large = |e: &str| a_fail("A3", Some(page), &format!("too large to verify: {e}")); + let forms = reachable_forms(ctx, page_id, &mut budget).map_err(|e| too_large(&e))?; + let mut attached = content + .parts + .iter() + .any(|p| originals.iter().any(|o| o.digest == p.digest)); + for f in &forms { + attached = attached || originals::holds_original(f, originals).map_err(too_large)?; + } + if attached { + problems.push("the original content is still attached".to_string()); + } + if content.parts.len() != plan.expected_parts.len() { + problems.push(format!( + "part count {} (expected {})", + content.parts.len(), + plan.expected_parts.len() + )); + } else if let Some(i) = (0..content.parts.len()) + .find(|i| plan.expected_parts.get(*i).map(Vec::as_slice) != Some(content.part_bytes(*i))) + { + problems.push(format!("part {i} bytes differ")); + } + if problems.is_empty() { + Ok(()) + } else { + Err(a_fail("A3", Some(page), &problems.join("; "))) + } +} + +/// The original edited parts that no part of the expected page equals (a page may hold two +/// identical parts; the unedited twin is not "the original content"). +pub(crate) fn original_parts(model: &PageModel, plan: &PagePlan) -> Vec { + let expected: Vec = plan + .expected_parts + .iter() + .map(|p| content_digest(p)) + .collect(); + plan.edited_parts + .iter() + .filter(|i| **i < model.content.parts.len()) + .map(|i| OriginalPart::of(model.content.part_bytes(*i))) + .filter(|o| !expected.contains(&o.digest)) + .collect() +} + +/// Cached "before" words and render of a qpdf input page (the preview re-renders the same +/// `p.pdf` for every commit). Keyed by path, length, mtime, the FNV of the file's first and last +/// `HEAD_TAIL_HASH_BYTES` (a same-size rewrite under a coarse mtime is a miss), page and DPI. +type BeforeKey = (PathBuf, u64, Option, u64, u32, u32); +type BeforeValue = Arc<(Vec, Raster)>; +static BEFORE_CACHE: Mutex> = Mutex::new(VecDeque::new()); +const BEFORE_CACHE_MAX: usize = 4; +/// Raster bytes the cache may hold in total. +const BEFORE_CACHE_BYTES_MAX: usize = 32 << 20; + +/// FNV of the first and last `HEAD_TAIL_HASH_BYTES` of `path` (`None` when unreadable). +fn head_tail_hash(path: &Path, len: u64) -> Option { + use std::io::{Seek, SeekFrom}; + let n = HEAD_TAIL_HASH_BYTES as u64; + let mut file = std::fs::File::open(path).ok()?; + let mut head = Vec::new(); + (&mut file).take(n).read_to_end(&mut head).ok()?; + file.seek(SeekFrom::Start(len.saturating_sub(n))).ok()?; + let mut tail = Vec::new(); + file.take(n).read_to_end(&mut tail).ok()?; + Some(fnv1a_extend(fnv1a_u64(&head), &tail)) +} + +fn before_inputs( + engines: &Engines, + r: &PopplerRef, + dpi: u32, + work: &Path, + opts: &RunOpts<'_>, +) -> Result { + let meta = std::fs::metadata(&r.pdf).ok(); + let len = meta.as_ref().map_or(0, std::fs::Metadata::len); + let key = head_tail_hash(&r.pdf, len).map(|hash| { + let modified = meta.and_then(|m| m.modified().ok()); + (r.pdf.clone(), len, modified, hash, r.page_1, dpi) + }); + let lock = || BEFORE_CACHE.lock().unwrap_or_else(|e| e.into_inner()); + if let Some(key) = &key { + let hit = lock() + .iter() + .find(|(k, _)| k == key) + .map(|(_, v)| Arc::clone(v)); + if let Some(hit) = hit { + return Ok(hit); + } + } + let words = pdftotext_words(engines, &r.pdf, r.page_1, opts)?; + let raster = render_page(engines, &r.pdf, r.page_1, dpi, work, opts)?; + let value: BeforeValue = Arc::new((words, raster)); + if let (Some(key), true) = (key, value.1.rgb.len() <= BEFORE_CACHE_BYTES_MAX) { + let mut cache = lock(); + cache.retain(|(k, _)| *k != key); + cache.push_front((key, Arc::clone(&value))); + let mut total = 0usize; + let keep = cache + .iter() + .take_while(|(_, v)| { + total = total.saturating_add(v.1.rgb.len()); + total <= BEFORE_CACHE_BYTES_MAX + }) + .count() + .min(BEFORE_CACHE_MAX); + cache.truncate(keep); + } + Ok(value) +} + +/// A5 on one page: returns the `EDIT_NOT_VISIBLE` warnings. +#[allow(clippy::too_many_arguments)] +fn check_a5( + engines: &Engines, + page_in: &EditedPageInput<'_>, + model: &PageModel, + staged: &Path, + after_geom: &PageGeometry, + work: &Path, + opts: &RunOpts<'_>, + page: u32, +) -> Result, AppError> { + let Some(dpi) = render_dpi(after_geom) else { + return Err(a_fail( + "A5", + Some(page), + &format!( + "check=render page={page} too large to verify: the page renders below \ + {RENDER_DPI_MIN} DPI" + ), + )); + }; + let before = before_inputs(engines, &page_in.input_render, dpi, work, opts) + .map_err(|e| engine_error(e, "A5", Some(page)))?; + let page_1 = page_in.input_page_index.saturating_add(1); + let dst_words = pdftotext_words(engines, staged, page_1, opts) + .map_err(|e| engine_error(e, "A5", Some(page)))?; + let dst_raster = render_page(engines, staged, page_1, dpi, work, opts) + .map_err(|e| engine_error(e, "A5", Some(page)))?; + let (mut edits, mut olds) = (Vec::new(), Vec::new()); + for run in &page_in.plan.runs { + let old = model + .runs + .iter() + .find(|r| { + r.members + .iter() + .filter_map(|m| model.walk.records.get(*m)) + .filter_map(|x| x.span.clone()) + .eq(run.expected.member_spans.iter().cloned()) + }) + .ok_or_else(|| a_fail("A5", Some(page), "edited run not on the page"))?; + let r = old.rect; + let n = run.new_rect; + edits.push(( + run.expected.clone(), + [r[0], r[1], r[0] + r[2], r[1] + r[3]], + [n[0], n[1], n[0] + n[2], n[1] + n[3]], + old.text.clone(), + )); + olds.push(old); + } + let near = NearGlyphs::of(model, &olds, &edits, after_geom, dpi); + let report = check_independent( + &before.0, + &dst_words, + &before.1, + &dst_raster, + after_geom, + &edits, + &near, + dpi, + ) + .map_err(|(check, detail)| { + a_fail( + "A5", + Some(page), + &format!("check={check} page={page} {detail}"), + ) + })?; + Ok(report + .edit_visible + .iter() + .zip(&page_in.plan.runs) + .filter(|(v, _)| !**v) + .map(|(_, run)| TextWarning { + page_index: model.page_index, + run_id: run.run_id.clone(), + code: TextWarningCode::EditNotVisible, + detail: Some("no pixel of the edited line changed".to_string()), + }) + .collect()) +} + +/// The cancel flag the in-process checks (graph traversal, walks) poll: the caller's flag, else +/// the job handle's. +pub(crate) fn cancel_flag<'a>(opts: &RunOpts<'a>) -> Option<&'a AtomicBool> { + opts.cancel.or_else(|| opts.handle.map(|h| &h.cancelled)) +} + +pub(crate) fn is_cancelled(opts: &RunOpts<'_>) -> bool { + opts.cancel.is_some_and(|c| c.load(Ordering::SeqCst)) + || opts.handle.is_some_and(|h| h.is_cancelled()) +} + +/// A check that failed while the job was being cancelled failed because of the cancel (a walk +/// stopped early, a traversal gave up): the result is `CANCELLED`, never a verification code. +fn or_cancelled(r: Result, opts: &RunOpts<'_>) -> Result { + match r { + Err(_) if is_cancelled(opts) => Err(AppError::cancelled()), + other => other, + } +} + +/// Phase A on one edited copy (or preview file). The caller deletes its work files on failure. +/// A failure while the job is cancelled is `CANCELLED`. +pub fn verify_edited_copy( + input: &PhaseAInput<'_>, + engines: &Engines, + work: &Path, + opts: &RunOpts<'_>, +) -> Result { + or_cancelled(phase_a(input, engines, work, opts), opts) +} + +fn phase_a( + input: &PhaseAInput<'_>, + engines: &Engines, + work: &Path, + opts: &RunOpts<'_>, +) -> Result { + if !seams::skipped("A0") { + check_a0(engines, input.staged, input.source_benign, opts)?; + } + let mut snap = read_verification_snapshot(input.staged, input.staged_cap).map_err(|e| { + a_fail( + "A1", + None, + &format!("read: {}", e.details.unwrap_or(e.message)), + ) + })?; + if !seams::skipped("A1") { + let pages = + qpdf_page_map(engines, input.staged, opts).map_err(|e| engine_error(e, "A1", None))?; + check_page_map(&snap, &pages) + .map_err(|e| a_fail("A1", None, &e.details.unwrap_or(e.message)))?; + if snap.pages.len() as u64 != u64::from(input.before_page_count) { + return Err(a_fail( + "A1", + None, + &format!( + "page count {} (expected {})", + snap.pages.len(), + input.before_page_count + ), + )); + } + } + snap.release_bytes(); + let ctx = SnapshotContext::new(snap); + if !seams::skipped("A2") { + let mut budget = + DecodeBudget::new(usize::try_from(GATE_DECODED_TOTAL).unwrap_or(usize::MAX)); + graph_matches(ctx.doc(), input.before, &mut budget, cancel_flag(opts)).map_err(|m| { + let lead = if m.what == "budget" { + "too large to verify " + } else { + "" + }; + a_fail( + "A2", + None, + &format!("{lead}path={} what={}", m.path, m.what), + ) + })?; + } + let mut proofs = Vec::new(); + let mut warnings = Vec::new(); + for page_in in &input.pages { + let (proof, mut w) = verify_page_a(&ctx, page_in, input.staged, engines, work, opts)?; + proofs.push(proof); + warnings.append(&mut w); + } + Ok(PhaseAReport { proofs, warnings }) +} + +/// A3–A5 on one page and its proof. +fn verify_page_a( + ctx: &SnapshotContext, + page_in: &EditedPageInput<'_>, + staged: &Path, + engines: &Engines, + work: &Path, + opts: &RunOpts<'_>, +) -> Result<(PageProof, Vec), AppError> { + let built; + let model = match &page_in.model { + BeforeModel::Kept(m) => *m, + BeforeModel::Build { + ctx, + page_index, + print, + } => { + built = build_page_model(ctx, *page_index, cancel_flag(opts))?; + if ModelPrint::of(&built) != **print { + let page = Some(page_index.saturating_add(1)); + let detail = "the source page model built again differs from the planned one"; + return Err(phase_a_error(EditProblemCode::Stale, "A4", page, detail)); + } + &built + } + }; + let page = model.page_index.saturating_add(1); + let plan = page_in.plan; + let page_id = ctx + .page_id(page_in.input_page_index) + .map_err(|_| a_fail("A1", Some(page), "edited page missing"))?; + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + let content = page_content(ctx.doc(), page_id, &mut budget) + .map_err(|r| a_fail("A3", Some(page), &format!("page content: {}", r.as_str())))?; + let originals = original_parts(model, plan); + if !seams::skipped("A3") { + check_a3(ctx, page_id, &content, plan, &originals, page)?; + } + // The proof keeps the after walk's depth-0 records as prints, taken before the probe walk. + let prints = |walk: PageWalk| -> (PageGeometry, Vec) { + let records = walk.records.iter().filter(|r| r.depth == 0); + ( + walk.geometry.clone(), + records.map(RecordPrint::of).collect(), + ) + }; + let (geometry, records) = if seams::skipped("A4") { + let walk = walk_page( + ctx, + page_in.input_page_index, + &content, + WalkMode::Edit, + cancel_flag(opts), + ); + prints(walk) + } else { + let walks = (&content, &*model.walk, &model.runs[..]); + let index = page_in.input_page_index; + walk_and_verify_with(ctx, index, walks, plan, cancel_flag(opts), prints) + .map_err(|f| phase_a_error(f.problem_code(), "A4", Some(page), &f.to_string()))? + }; + drop(content); + let warnings = if seams::skipped("A5") { + Vec::new() + } else { + check_a5(engines, page_in, model, staged, &geometry, work, opts, page)? + }; + let proof = PageProof { + source_page_index: model.page_index, + part_digests: plan + .expected_parts + .iter() + .map(|p| content_digest(p)) + .collect(), + expected_parts: plan.expected_parts.clone(), + original_edited_digests: originals.iter().map(|o| o.digest).collect(), + original_edited_rolling: originals.iter().map(|o| o.rolling).collect(), + expected_page_digest: plan.expected_page_digest, + records, + }; + Ok((proof, warnings)) +} + +/// Phase B (§B.16.2) on the final staged file, immediately before #34's `validate_staged_pdf`: +/// B0 (bounded read with no policy refusals — D31 — and the page map), then B1–B3 per edited +/// destination page (`gate/final_output.rs`). Any failure → `SOURCE_EDIT_GATE_FAILED`, or +/// `CANCELLED` while the job is being cancelled. +pub fn verify_final_output( + staged: &Path, + cap: u64, + engines: &Engines, + expectations: &[(u32, &PageProof)], + opts: &RunOpts<'_>, +) -> Result<(), AppError> { + or_cancelled(phase_b(staged, cap, engines, expectations, opts), opts) +} + +fn phase_b( + staged: &Path, + cap: u64, + engines: &Engines, + expectations: &[(u32, &PageProof)], + opts: &RunOpts<'_>, +) -> Result<(), AppError> { + let fail = |detail: &str| source_edit_gate_failed(&format!("phase=B {detail}")); + let detail = |e: AppError| e.details.unwrap_or(e.message); + let mut snap = read_verification_snapshot(staged, cap) + .map_err(|e| fail(&format!("check=B0 read: {}", detail(e))))?; + let pages = qpdf_page_map(engines, staged, opts).map_err(|e| match e.code.as_str() { + "CANCELLED" | "ENGINE_MISSING" => e, + _ => fail(&format!("check=B0 {}", detail(e))), + })?; + check_page_map(&snap, &pages).map_err(|e| fail(&format!("check=B0 {}", detail(e))))?; + snap.release_bytes(); + let ctx = SnapshotContext::new(snap); + for (dest, proof) in expectations { + final_output::check_page(&ctx, *dest, proof, cancel_flag(opts)) + .map_err(|e| fail(&format!("page={} {e}", dest.saturating_add(1))))?; + } + Ok(()) +} + +/// Test seams: skip named gate checks on this thread (IND-02/03/04, the GATE "must fail at" +/// matrix). Production builds have no way to skip a check. +mod seams { + pub(super) fn skipped(check: &str) -> bool { + #[cfg(test)] + { + if super::test_seams::skipped(check) { + return true; + } + } + let _ = check; + false + } +} + +#[cfg(test)] +pub(crate) mod test_seams { + use std::cell::RefCell; + + thread_local! { + static SKIPPED: RefCell> = const { RefCell::new(Vec::new()) }; + } + + pub(crate) fn skipped(check: &str) -> bool { + SKIPPED.with(|s| s.borrow().iter().any(|c| c == check)) + } + + /// Skips `checks` (e.g. `["A2", "A3"]`, `["B1"]`) on this thread until the guard drops. + pub(crate) fn skip(checks: &[&str]) -> SkipGuard { + SKIPPED.with(|s| *s.borrow_mut() = checks.iter().map(|c| c.to_string()).collect()); + SkipGuard + } + + pub(crate) struct SkipGuard; + + impl Drop for SkipGuard { + fn drop(&mut self) { + SKIPPED.with(|s| s.borrow_mut().clear()); + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/gate/final_output.rs b/src-tauri/src/pdf_engine/text_edit/gate/final_output.rs new file mode 100644 index 0000000..42bbe1d --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/gate/final_output.rs @@ -0,0 +1,359 @@ +//! Phase B (SPEC §B.16.2, D29): on the final staged file, immediately before #34's +//! `validate_staged_pdf`. The expected content of each edited destination page must be found in +//! exactly one of two forms — its parts as a contiguous in-order run of the page's parts (later +//! passes may add parts around them), or qpdf's overlay wrapper: page content of `q cm Do Q` only, +//! one painted Form whose data holds `qpdf_join(expected_parts)` once (at its start or right after +//! a `\n`), with a composite matrix that is the identity within tolerance and a `/BBox` covering +//! the visible page (B1). The original edited parts must be nowhere (B2), and the page's depth-0 +//! show records — walked through the wrapper when there is one — must be the proven ones (B3). +//! Any failure is `SOURCE_EDIT_GATE_FAILED`. + +use super::join::join_is_neutral; +use super::originals::{find_up_to, holds_original}; +use super::{reachable_forms, PageProof, RecordPrint}; +use crate::pdf_engine::text_edit::content::{page_content, qpdf_join, PageContent}; +use crate::pdf_engine::text_edit::context::{resolve, Lookup, Res, SnapshotContext}; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget}; +use crate::pdf_engine::text_edit::fonts::number_of; +use crate::pdf_engine::text_edit::geometry::{contains, mul, transform_rect, Matrix, IDENTITY}; +use crate::pdf_engine::text_edit::lexer::{lex_content, LexLimits, Operand, Operator}; +use crate::pdf_engine::text_edit::limits::{ + DRIFT_TOLERANCE_PT, PAGE_CONTENT_MAX_DECODED, PAGE_DECODE_BUDGET, WRAPPER_MATRIX_EPSILON, + WRAPPER_TRANSLATION_TOL_PT, +}; +use crate::pdf_engine::text_edit::state::{same_state, ClipState, StateDigest}; +use crate::pdf_engine::text_edit::walker::{walk_page, ShowRecord, WalkMode}; +use lopdf::{Dictionary, Object, ObjectId}; +use std::sync::atomic::AtomicBool; + +/// The wrapper Form's `/BBox` must cover the visible box within this (pt). +const WRAPPER_BBOX_TOL_PT: f64 = 0.01; + +/// Where the expected content was found. +enum Found { + Parts, + Wrapper { name: Vec }, +} + +/// Starting indices where the proof's expected parts appear as a contiguous run of `content`'s +/// parts (digests first, then the bytes). +fn parts_hits(content: &PageContent, proof: &PageProof) -> usize { + let expected = &proof.expected_parts; + let (n, k) = (content.parts.len(), expected.len()); + if k == 0 || k > n || proof.part_digests.len() != k { + return 0; + } + (0..=n - k) + .filter(|start| { + expected + .iter() + .zip(&proof.part_digests) + .enumerate() + .all(|(j, (want, digest))| { + content.parts.get(start + j).map(|p| p.digest) == Some(*digest) + && content.part_bytes(start + j) == want.as_slice() + }) + }) + .count() +} + +/// Occurrences of `needle` in `hay` — counted up to 2, which is all B1 needs ("exactly once") — +/// and how many of those start at 0 or right after `\n`. Linear in `hay` (`find_up_to`); `Err` +/// (fail closed) when the search gives up. +fn occurrences(hay: &[u8], needle: &[u8]) -> Result<(usize, usize), String> { + let hits = find_up_to(hay, needle, 2).map_err(|e| format!("too large to verify: {e}"))?; + let anchored = hits + .iter() + .filter(|i| { + **i == 0 + || i.checked_sub(1) + .and_then(|p| hay.get(p)) + .is_some_and(|b| *b == b'\n') + }) + .count(); + Ok((hits.len(), anchored)) +} + +fn form_matrix(doc: &lopdf::Document, dict: &Dictionary) -> Option { + let Ok(raw) = dict.get(b"Matrix") else { + return Some(IDENTITY); + }; + let (_, Object::Array(items)) = resolve(doc, raw)? else { + return None; + }; + if items.len() != 6 { + return None; + } + let mut m = IDENTITY; + for (slot, item) in m.iter_mut().zip(items) { + *slot = resolve(doc, item).and_then(|(_, o)| number_of(o))?; + } + Some(m) +} + +fn form_bbox(doc: &lopdf::Document, dict: &Dictionary) -> Option<[f64; 4]> { + let (_, Object::Array(items)) = resolve(doc, dict.get(b"BBox").ok()?)? else { + return None; + }; + let v: Vec = items + .iter() + .map(|i| resolve(doc, i).and_then(|(_, o)| number_of(o))) + .collect::>>()?; + match v.as_slice() { + [a, b, c, d] => Some([a.min(*c), b.min(*d), a.max(*c), b.max(*d)]), + _ => None, + } +} + +/// The wrapper form (B1): `(name, form data)` of the one painted Form that holds the expected +/// content, after checking the page shape, the composite matrix and the `/BBox`. +fn wrapper_hit( + ctx: &SnapshotContext, + page_id: ObjectId, + content: &PageContent, + joined: &[u8], + visible: [f64; 4], +) -> Result, Vec)>, String> { + let Ok(ops) = lex_content(&content.joined, &LexLimits::page(), None) else { + return Ok(None); + }; + let shape = !ops.is_empty() + && ops.iter().all(|o| { + matches!( + o.operator, + Operator::q | Operator::Q | Operator::cm | Operator::Do + ) + }); + if !shape { + return Ok(None); + } + let doc = ctx.doc(); + let res = Res::of_page(doc, page_id).map_err(str::to_string)?; + let mut stack: Vec = Vec::new(); + let mut ctm = IDENTITY; + let mut hits: Vec<(Vec, Vec, Matrix, ObjectId)> = Vec::new(); + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + for op in &ops { + match op.operator { + Operator::q => stack.push(ctm), + Operator::Q => ctm = stack.pop().unwrap_or(ctm), + Operator::cm => { + let n: Vec = op.operands.iter().filter_map(Operand::as_number).collect(); + let m: Matrix = match n.as_slice() { + [a, b, c, d, e, f] => [*a, *b, *c, *d, *e, *f], + _ => return Err("wrapper cm".to_string()), + }; + ctm = mul(&m, &ctm); + } + Operator::Do => { + let name = op + .operands + .first() + .and_then(Operand::as_name) + .unwrap_or_default(); + let Lookup::Found(e) = res.entry(doc, b"XObject", name) else { + return Err("wrapper XObject".to_string()); + }; + let (Some(id), Object::Stream(s)) = (e.id, e.value) else { + return Err("wrapper XObject".to_string()); + }; + let is_form = + s.dict.get(b"Subtype").ok().and_then(|o| o.as_name().ok()) == Some(b"Form"); + if !is_form { + continue; + } + let data = decode_stream(s, PAGE_CONTENT_MAX_DECODED, &mut budget) + .map_err(|e| format!("wrapper data: {e}"))?; + let (all, anchored) = occurrences(&data, joined)?; + if all == 0 { + continue; + } + if all != 1 { + return Err("expected content repeated in the wrapper".to_string()); + } + if anchored != 1 { + return Err( + "expected content not at a part boundary in the wrapper".to_string() + ); + } + let matrix = form_matrix(doc, &s.dict).ok_or("wrapper /Matrix")?; + hits.push((name.to_vec(), data, mul(&matrix, &ctm), id)); + } + _ => {} + } + } + let [(name, data, m, id)] = hits.as_slice() else { + if hits.is_empty() { + return Ok(None); + } + return Err("the expected content is painted more than once".to_string()); + }; + let painted = ops + .iter() + .filter(|o| { + o.operator == Operator::Do + && o.operands.first().and_then(Operand::as_name) == Some(name.as_slice()) + }) + .count(); + if painted != 1 { + return Err("the wrapper Form is painted more than once".to_string()); + } + let linear_ok = (m[0] - 1.0).abs() <= WRAPPER_MATRIX_EPSILON + && m[1].abs() <= WRAPPER_MATRIX_EPSILON + && m[2].abs() <= WRAPPER_MATRIX_EPSILON + && (m[3] - 1.0).abs() <= WRAPPER_MATRIX_EPSILON; + let shift_ok = + m[4].abs() <= WRAPPER_TRANSLATION_TOL_PT && m[5].abs() <= WRAPPER_TRANSLATION_TOL_PT; + if !linear_ok || !shift_ok { + return Err("wrapper matrix is not the identity".to_string()); + } + let Some(Object::Stream(s)) = doc.objects.get(id) else { + return Err("wrapper XObject".to_string()); + }; + let bbox = form_bbox(doc, &s.dict).ok_or("wrapper /BBox")?; + if !contains(transform_rect(m, bbox), visible, WRAPPER_BBOX_TOL_PT) { + return Err("wrapper /BBox cuts the visible page".to_string()); + } + Ok(Some((name.clone(), data.clone()))) +} + +/// The digest of the after state with the wrapper's transform factored out: CTM and clip within +/// the wrapper tolerances of the proof's are taken as the proof's. +fn factored(after: &StateDigest, proof: &StateDigest) -> StateDigest { + let mut d = after.clone(); + let ctm_close = after + .ctm + .iter() + .zip(&proof.ctm) + .enumerate() + .all(|(i, (a, b))| { + let tol = if i < 4 { + WRAPPER_MATRIX_EPSILON * 1f64.max(a.abs()).max(b.abs()) + } else { + WRAPPER_TRANSLATION_TOL_PT + }; + (a - b).abs() <= tol + }); + if ctm_close { + d.ctm = proof.ctm; + } + if let (ClipState::Rect(a), ClipState::Rect(b)) = (&after.clip, &proof.clip) { + if a.iter() + .zip(b) + .all(|(x, y)| (x - y).abs() <= DRIFT_TOLERANCE_PT) + { + d.clip = proof.clip.clone(); + } + } + d +} + +/// B3: the depth-0 show records equal the proof one-to-one. +fn same_records(records: &[&ShowRecord], proof: &[RecordPrint]) -> Result<(), String> { + if records.len() != proof.len() { + return Err(format!( + "show records {} (expected {})", + records.len(), + proof.len() + )); + } + for (i, (r, p)) in records.iter().zip(proof).enumerate() { + let now = RecordPrint::of(r); + if now.op != p.op + || now.font_res != p.font_res + || now.font_hash != p.font_hash + || now.codes != p.codes + || now.text != p.text + { + return Err(format!("record {i} differs")); + } + let far = + |a: &(f64, f64), b: &(f64, f64)| (a.0 - b.0).hypot(a.1 - b.1) > DRIFT_TOLERANCE_PT; + let moved = now.origins.len() != p.origins.len() + || now.origins.iter().zip(&p.origins).any(|(a, b)| far(a, b)) + || far(&now.pen_after, &p.pen_after); + if moved { + return Err(format!("record {i} moved")); + } + same_state(&factored(&now.state, &p.state), &p.state) + .map_err(|field| format!("record {i} state {field}"))?; + } + Ok(()) +} + +/// B1–B3 on one destination page. +pub(super) fn check_page( + ctx: &SnapshotContext, + dest: u32, + proof: &PageProof, + cancel: Option<&AtomicBool>, +) -> Result<(), String> { + let skip_b1 = super::seams::skipped("B1"); + let page_id = ctx.page_id(dest).map_err(|_| "page missing".to_string())?; + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + let content = page_content(ctx.doc(), page_id, &mut budget) + .map_err(|r| format!("content {}", r.as_str()))?; + let visible = crate::pdf_engine::text_edit::geometry::page_geometry(ctx.doc(), page_id) + .map_err(|r| format!("geometry {}", r.as_str()))? + .visible; + let parts: Vec<&[u8]> = proof.expected_parts.iter().map(Vec::as_slice).collect(); + let joined = qpdf_join(&parts); + let in_parts = parts_hits(&content, proof); + let wrapper = match wrapper_hit(ctx, page_id, &content, &joined, visible) { + // qpdf's join must not change what the expected parts mean (review-T5 H1). + Ok(Some(_)) if !skip_b1 && !join_is_neutral(&parts, &joined) => { + return Err("check=B1 qpdf's overlay join changes the edited content".to_string()) + } + Ok(w) => w, + Err(e) if !skip_b1 => return Err(format!("check=B1 {e}")), + Err(_) => None, + }; + let found = match (in_parts, &wrapper) { + (1, None) => Some(Found::Parts), + (0, Some((name, _))) => Some(Found::Wrapper { name: name.clone() }), + _ => None, + }; + if found.is_none() && !skip_b1 { + return Err(if in_parts == 0 && wrapper.is_none() { + "check=B1 the expected content is not on the page".to_string() + } else { + "check=B1 the expected content is on the page more than once".to_string() + }); + } + if !super::seams::skipped("B2") { + // The parts by digest; every reachable Form (the wrapper is one) whole or as a segment. + let originals = proof.originals(); + let too_large = |e: &str| format!("check=B2 too large to verify: {e}"); + let mut found = content + .parts + .iter() + .any(|p| originals.iter().any(|o| o.digest == p.digest)); + let forms = reachable_forms(ctx, page_id, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)) + .map_err(|e| too_large(&e))?; + for f in &forms { + found = found || holds_original(f, &originals).map_err(too_large)?; + } + if found { + return Err("check=B2 the original content is still on the page".to_string()); + } + } + if !super::seams::skipped("B3") { + // With B1 skipped by a test and nothing found, every record at any depth is compared. + let mode = match found { + Some(Found::Parts) => WalkMode::Edit, + Some(Found::Wrapper { name }) => WalkMode::Wrapped { name }, + None => WalkMode::Classify, + }; + let walk = walk_page(ctx, dest, &content, mode, cancel); + if let Some(r) = walk.page_reason { + return Err(format!( + "check=B3 walk refused {} {}", + r.as_str(), + walk.page_detail.unwrap_or_default() + )); + } + let records: Vec<&ShowRecord> = walk.records.iter().collect(); + same_records(&records, &proof.records).map_err(|e| format!("check=B3 {e}"))?; + } + Ok(()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/gate/join.rs b/src-tauri/src/pdf_engine/text_edit/gate/join.rs new file mode 100644 index 0000000..3e7f3ed --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/gate/join.rs @@ -0,0 +1,108 @@ +//! Whether qpdf's overlay join leaves a page's content meaning what its parts meant (review-T5 +//! H1). `content::qpdf_join` writes a `\n` after every part that does not end in one; that +//! newline only means nothing when the part boundary is neutral. The rule is the walker's own +//! (`content::joins`, review-final LOW-1): a streaming scan of the parts' lexical state at each +//! boundary, with no op cap, so a page the strict lexer refuses (over 250,000 ops, trailing +//! operands, unknown operators) is still judged by its boundaries alone (review-final MEDIUM-2). + +use crate::pdf_engine::text_edit::content::first_unsafe_boundary; + +/// Positions of the `\n` bytes `qpdf_join(parts)` inserts, and the joined length (mirrors +/// `content::qpdf_join`). +fn join_points(parts: &[&[u8]]) -> (Vec, usize) { + let mut points = Vec::new(); + let mut len = 0usize; + let mut need_newline = false; + for part in parts { + if need_newline { + points.push(len); + len = len.saturating_add(1); + } + let last = match part.last() { + Some(c) => Some(*c), + None if need_newline => Some(b'\n'), + None => None, + }; + len = len.saturating_add(part.len()); + need_newline = last != Some(b'\n'); + } + (points, len) +} + +/// Whether `joined` (= `content::qpdf_join(parts)`) reads as the parts read one after the other: +/// it is qpdf's join of `parts`, and every boundary between non-empty parts is neutral +/// (`content::joins`). `true` when the join inserts nothing. +pub(crate) fn join_is_neutral(parts: &[&[u8]], joined: &[u8]) -> bool { + let (points, len) = join_points(parts); + if points.is_empty() { + return true; + } + if len != joined.len() || points.iter().any(|p| joined.get(*p) != Some(&b'\n')) { + return false; + } + first_unsafe_boundary(parts).is_none() +} + +#[cfg(test)] +mod tests { + use super::join_is_neutral; + use crate::pdf_engine::text_edit::content::qpdf_join; + + fn neutral(parts: &[&[u8]]) -> bool { + join_is_neutral(parts, &qpdf_join(parts)) + } + + #[test] + fn join_neutral_accepts_operator_and_whitespace_boundaries() { + let ok: [&[&[u8]]; 7] = [ + &[ + b"BT /F1 12 Tf 72 720 Td (A) Tj ET", + b"BT /F1 12 Tf 72 700 Td (B) Tj ET", + ], + &[b"q 1 0 0 1 0 0 cm", b"BT /F1 12 Tf (A) Tj ET Q"], + &[b"Q", b"q"], + &[b"ET ", b"BT ET"], + &[b"0 g\n", b"1 g"], + &[b"[(a) 10", b"(b)] TJ"], + &[b"q\n% note\n", b"Q"], + ]; + for parts in ok { + assert!(neutral(parts), "{parts:?}"); + } + assert!(neutral(&[]), "no parts"); + assert!(neutral(&[b"BT ET\n", b""]), "nothing inserted"); + } + + #[test] + fn join_neutral_refuses_boundaries_inside_tokens_and_comments() { + let refused: [&[&[u8]]; 11] = [ + // review-T5 R1: the comment ran on into the next part and hid it. + &[ + b"BT /F1 12 Tf (Visible) Tj ET\n% note", + b"BT /F1 12 Tf (Secret) Tj ET", + ], + &[b"BT ET % note", b"BT ET"], + // review-T5: a string split between parts. + &[b"BT /F1 12 Tf 72 700 Td (Hel", b"lo) Tj ET"], + &[b"BT /F1 12 Tf 72 700 Td <48", b"65> Tj ET"], + // Numbers and names that the plain concatenation reads as one token. + &[b"1 0 0 1 0 1", b"0 cm"], + &[b"BT /F", b"1 12 Tf ET"], + &[b"/GS0", b"gs"], + // Operators that would merge into another operator for pdf.js. + &[b"0 0 m 10 10 l s", b"h"], + &[b"[] 0 d", b"0 0 m"], + // An inline image split inside its data. + &[b"q BI /W 2 /H 1 /BPC 8 /CS /G ID \x01", b"\x02 EI Q"], + // A boundary inside an unterminated string. + &[b"BT (unterminated", b" Tj ET"], + ]; + for parts in refused { + assert!(!neutral(parts), "{parts:?}"); + } + assert!( + !join_is_neutral(&[b"Q", b"q"], b"Qq"), + "a joined buffer that is not qpdf_join(parts)" + ); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/gate/originals.rs b/src-tauri/src/pdf_engine/text_edit/gate/originals.rs new file mode 100644 index 0000000..eeeea54 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/gate/originals.rs @@ -0,0 +1,154 @@ +//! Finding an original edited part again (A3, B2): as a whole part or Form XObject (its digest), +//! or as one segment of a Form's data the way qpdf joins page parts into its overlay wrapper — +//! starting at offset 0 or right after a `\n`, and ending at the end of the data or at a `\n`. +//! The search is one rolling-hash pass per distinct part length, so a wrapper that keeps the old +//! part next to other content (a fake wrapped by a later overlay) is found without quadratic +//! work; every rolling-hash hit is confirmed by the part's FNV digest. + +use crate::pdf_engine::validate_output::{content_digest, ContentDigest}; + +/// Polynomial base of the rolling hash (odd, so powers never collapse to 0 mod 2^64). +const BASE: u64 = 0x0000_0100_0000_01b3; +/// Rolling-hash hits at aligned positions whose digest differs, per search, before giving up +/// ("too large to verify": the caller fails closed). +const FALSE_HITS_MAX: usize = 64; + +/// An original edited part: its digest (hash and length) and its rolling hash. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct OriginalPart { + pub digest: ContentDigest, + pub rolling: u64, +} + +impl OriginalPart { + pub fn of(bytes: &[u8]) -> OriginalPart { + OriginalPart { + digest: content_digest(bytes), + rolling: rolling(bytes), + } + } +} + +/// The rolling hash of `bytes` (Σ b·BASE^(n−1−i), wrapping). +pub(crate) fn rolling(bytes: &[u8]) -> u64 { + bytes.iter().fold(0u64, |h, b| { + h.wrapping_mul(BASE).wrapping_add(u64::from(*b)) + }) +} + +fn power(mut exp: usize) -> u64 { + let (mut base, mut out) = (BASE, 1u64); + while exp > 0 { + if exp & 1 == 1 { + out = out.wrapping_mul(base); + } + base = base.wrapping_mul(base); + exp >>= 1; + } + out +} + +/// `[start, start + len)` is a qpdf-join segment of `data`. +fn aligned(data: &[u8], start: usize, len: usize) -> bool { + let start_ok = start == 0 || start.checked_sub(1).and_then(|i| data.get(i)) == Some(&b'\n'); + let Some(end) = start.checked_add(len) else { + return false; + }; + let end_ok = end == data.len() + || data.get(end) == Some(&b'\n') + || end.checked_sub(1).and_then(|i| data.get(i)) == Some(&b'\n'); + start_ok && end_ok +} + +/// The start offsets of the first `max` occurrences of `needle` in `data` (overlapping ones +/// included): one rolling-hash pass, every hash hit confirmed byte for byte, so the cost is +/// O(|data| + hits × |needle|) whatever the bytes are. `Err` after more than `FALSE_HITS_MAX` +/// hash hits that are not the needle (fail closed). +pub(crate) fn find_up_to( + data: &[u8], + needle: &[u8], + max: usize, +) -> Result, &'static str> { + let len = needle.len(); + let mut hits = Vec::new(); + if len == 0 || len > data.len() || max == 0 { + return Ok(hits); + } + let target = rolling(needle); + let top = power(len.saturating_sub(1)); + let mut h = rolling(data.get(..len).unwrap_or_default()); + let mut start = 0usize; + let mut false_hits = 0usize; + loop { + if h == target { + if data.get(start..start.saturating_add(len)) == Some(needle) { + hits.push(start); + if hits.len() >= max { + return Ok(hits); + } + } else { + false_hits = false_hits.saturating_add(1); + if false_hits > FALSE_HITS_MAX { + return Err("too many near matches of the expected content"); + } + } + } + let end = start.saturating_add(len); + let (Some(out), Some(next)) = (data.get(start), data.get(end)) else { + return Ok(hits); + }; + h = h + .wrapping_sub(top.wrapping_mul(u64::from(*out))) + .wrapping_mul(BASE) + .wrapping_add(u64::from(*next)); + start = start.saturating_add(1); + } +} + +/// Whether `data` is, or holds as a qpdf-join segment, one of `originals`. `Err` when the search +/// had to give up (fail closed). +pub(crate) fn holds_original( + data: &[u8], + originals: &[OriginalPart], +) -> Result { + let mut lens: Vec = originals + .iter() + .map(|o| o.digest.len) + .filter(|l| *l > 0 && *l <= data.len()) + .collect(); + lens.sort_unstable(); + lens.dedup(); + let mut false_hits = 0usize; + for len in lens { + let targets: Vec<&OriginalPart> = + originals.iter().filter(|o| o.digest.len == len).collect(); + let top = power(len.saturating_sub(1)); + let mut h = rolling(data.get(..len).unwrap_or_default()); + let mut start = 0usize; + loop { + if targets.iter().any(|t| t.rolling == h) && aligned(data, start, len) { + let window = data + .get(start..start.saturating_add(len)) + .unwrap_or_default(); + let d = content_digest(window); + if targets.iter().any(|t| t.digest == d) { + return Ok(true); + } + false_hits = false_hits.saturating_add(1); + if false_hits > FALSE_HITS_MAX { + return Err("too many near matches of the original content"); + } + } + let end = start.saturating_add(len); + let (Some(out), Some(next)) = (data.get(start), data.get(end)) else { + break; + }; + h = h + .wrapping_sub(top.wrapping_mul(u64::from(*out))) + .wrapping_mul(BASE) + .wrapping_add(u64::from(*next)); + start = start.saturating_add(1); + } + } + Ok(false) +} diff --git a/src-tauri/src/pdf_engine/text_edit/geometry.rs b/src-tauri/src/pdf_engine/text_edit/geometry.rs new file mode 100644 index 0000000..1507de1 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/geometry.rs @@ -0,0 +1,326 @@ +//! Strict page geometry and text orientation (SPEC §A.5, §B.10). +//! +//! Matrices use the PDF row-vector convention: `mul(a, b)` is "a × b" (apply `a` first, then +//! `b`), so a glyph's rendering matrix is `mul(&mul(&text_params, &tm), &ctm)`. Page boxes are +//! read by an own parser that refuses everything a viewer would have to guess (`crop.rs` falls +//! back silently and is only used by the parity test GEO-11). + +use crate::pdf_engine::text_edit::fonts::number_of; +use crate::pdf_engine::text_edit::limits::{AXIS_EPSILON_REL, PAGE_TREE_DEPTH_MAX, SHEAR_MAX}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use lopdf::{Dictionary, Document, Object, ObjectId}; +use std::collections::HashSet; + +pub type Matrix = [f64; 6]; + +pub const IDENTITY: Matrix = [1.0, 0.0, 0.0, 1.0, 0.0, 0.0]; + +/// References followed for one box or number value. +const REF_HOPS_MAX: usize = 16; +/// Smallest visible side (Crop ∩ Media) in points. +const VISIBLE_SIDE_MIN_PT: f64 = 1.0; + +/// PDF "a × b": the transformation that applies `a`, then `b`. +pub fn mul(a: &Matrix, b: &Matrix) -> Matrix { + let [a0, a1, a2, a3, a4, a5] = *a; + let [b0, b1, b2, b3, b4, b5] = *b; + [ + a0 * b0 + a1 * b2, + a0 * b1 + a1 * b3, + a2 * b0 + a3 * b2, + a2 * b1 + a3 * b3, + a4 * b0 + a5 * b2 + b4, + a4 * b1 + a5 * b3 + b5, + ] +} + +/// The point `(x, y)` mapped by `m`. +pub fn apply(m: &Matrix, x: f64, y: f64) -> (f64, f64) { + (x * m[0] + y * m[2] + m[4], x * m[1] + y * m[3] + m[5]) +} + +/// The vector `(x, y)` mapped by the linear part of `m`. +pub fn apply_linear(m: &Matrix, x: f64, y: f64) -> (f64, f64) { + (x * m[0] + y * m[2], x * m[1] + y * m[3]) +} + +pub fn is_finite(m: &Matrix) -> bool { + m.iter().all(|v| v.is_finite()) +} + +/// A translation by `(tx, ty)`. +pub fn translate(tx: f64, ty: f64) -> Matrix { + [1.0, 0.0, 0.0, 1.0, tx, ty] +} + +/// R(0/90/180/270) of §A.5: clockwise display rotation in the row-vector convention. Any other +/// value is treated as 0 (callers only pass normalised `/Rotate` values). +pub fn display_rotation(rotate: i64) -> Matrix { + match rotate.rem_euclid(360) { + 90 => [0.0, -1.0, 1.0, 0.0, 0.0, 0.0], + 180 => [-1.0, 0.0, 0.0, -1.0, 0.0, 0.0], + 270 => [0.0, 1.0, -1.0, 0.0, 0.0, 0.0], + _ => IDENTITY, + } +} + +/// Page boxes as `[x0, y0, x1, y1]` (normalised, default user space). +#[derive(Debug, Clone, PartialEq)] +pub struct PageGeometry { + pub media: [f64; 4], + pub crop: [f64; 4], + pub visible: [f64; 4], + pub rotate: i64, + pub user_unit: f64, +} + +impl PageGeometry { + /// Geometry used by the classifier on a page refused `GEOMETRY` (§B.10): no box, so nothing is + /// "outside the page"; occurrences carry the page reason instead. + pub fn unbounded(rotate: i64) -> PageGeometry { + let all = [f64::MIN, f64::MIN, f64::MAX, f64::MAX]; + PageGeometry { + media: all, + crop: all, + visible: all, + rotate, + user_unit: 1.0, + } + } +} + +/// The page's geometry, read strictly (§A.5): the inherited `/MediaBox` must exist and be four +/// finite numbers with a positive area; an inherited `/CropBox` likewise when present; Crop ∩ Media +/// must be at least 1 pt in both directions; `/Rotate` (inherited) an Integer multiple of 90; +/// `/UserUnit` (page only) absent or exactly the number 1. Anything else is `GEOMETRY`. +pub fn page_geometry(doc: &Document, page_id: ObjectId) -> Result { + let bad = TextReason::Geometry; + let chain = page_chain(doc, page_id).ok_or(bad)?; + let page = chain.first().copied().ok_or(bad)?; + let media = inherited(&chain, b"MediaBox") + .ok_or(bad) + .and_then(|o| read_box(doc, o).ok_or(bad))?; + let crop = match inherited(&chain, b"CropBox") { + None => media, + Some(o) => read_box(doc, o).ok_or(bad)?, + }; + let visible = [ + crop[0].max(media[0]), + crop[1].max(media[1]), + crop[2].min(media[2]), + crop[3].min(media[3]), + ]; + if visible[2] - visible[0] < VISIBLE_SIDE_MIN_PT + || visible[3] - visible[1] < VISIBLE_SIDE_MIN_PT + { + return Err(bad); + } + let rotate = match inherited(&chain, b"Rotate") { + None => 0, + Some(o) => match resolve(doc, o) { + Some(Object::Integer(r)) if r.rem_euclid(90) == 0 => r.rem_euclid(360), + _ => return Err(bad), + }, + }; + let user_unit = match page.get(b"UserUnit").ok() { + None => 1.0, + Some(o) => match resolve(doc, o).and_then(number_of) { + Some(u) if u == 1.0 => 1.0, + _ => return Err(bad), + }, + }; + Ok(PageGeometry { + media, + crop, + visible, + rotate, + user_unit, + }) +} + +/// The page dictionary and its ancestors (nearest first); `None` on a dangling or cyclic chain. +fn page_chain(doc: &Document, page_id: ObjectId) -> Option> { + let mut out = Vec::new(); + let mut seen = HashSet::new(); + let mut cur = Some(page_id); + while let Some(id) = cur { + if !seen.insert(id) || out.len() > PAGE_TREE_DEPTH_MAX { + return None; + } + let dict = match doc.objects.get(&id)? { + Object::Dictionary(d) => d, + _ => return None, + }; + out.push(dict); + cur = match dict.get(b"Parent").ok() { + None => None, + Some(Object::Reference(p)) => Some(*p), + Some(_) => return None, + }; + } + Some(out) +} + +fn inherited<'a>(chain: &[&'a Dictionary], key: &[u8]) -> Option<&'a Object> { + chain.iter().find_map(|d| d.get(key).ok()) +} + +/// Follows references (≤ 16 hops); `None` when one dangles. +fn resolve<'a>(doc: &'a Document, obj: &'a Object) -> Option<&'a Object> { + let mut obj = obj; + for _ in 0..REF_HOPS_MAX { + match obj { + Object::Reference(id) => obj = doc.objects.get(id)?, + other => return Some(other), + } + } + None +} + +/// Four finite numbers, normalised, with a positive area. +fn read_box(doc: &Document, obj: &Object) -> Option<[f64; 4]> { + let items = match resolve(doc, obj)? { + Object::Array(items) if items.len() == 4 => items, + _ => return None, + }; + let mut v = [0.0; 4]; + for (slot, item) in v.iter_mut().zip(items) { + *slot = resolve(doc, item).and_then(number_of)?; + } + let b = [ + v[0].min(v[2]), + v[1].min(v[3]), + v[0].max(v[2]), + v[1].max(v[3]), + ]; + (b[2] > b[0] && b[3] > b[1]).then_some(b) +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum Orientation { + Upright, + ZeroSize, + Rotated, + Mirrored, + Skewed, +} + +/// §A.5 table for the composite `T = [Tfs·Th 0 0 Tfs 0 Ts] × Tm × CTM × R(rotate)`. `s` is the +/// largest absolute linear coefficient (the spec's `max(|a|,|d|)` extended to `|b|`, `|c|` so that +/// off-axis tolerances scale too); a rotation needs `a ≈ d` and `b ≈ −c` (signed). +pub fn classify_orientation(t: &Matrix) -> Orientation { + let [a, b, c, d, _, _] = *t; + if ![a, b, c, d].iter().all(|v| v.is_finite()) { + return Orientation::Skewed; + } + let s = a.abs().max(b.abs()).max(c.abs()).max(d.abs()); + if (a * d - b * c).abs() < 1e-12 * (s * s).max(1.0) { + return Orientation::ZeroSize; + } + let tol = AXIS_EPSILON_REL * s; + if b.abs() <= tol && c.abs() <= SHEAR_MAX * d.abs() && a > 0.0 && d > 0.0 { + return Orientation::Upright; + } + if b.abs() <= tol && c.abs() <= tol { + if a < 0.0 && d < 0.0 { + return Orientation::Rotated; + } + if (a < 0.0) != (d < 0.0) { + return Orientation::Mirrored; + } + } + if (a - d).abs() <= tol && (b + c).abs() <= tol { + return Orientation::Rotated; + } + Orientation::Skewed +} + +/// Distance from `origin` along the unit vector `dir` to the edge of `rect` (`[x0, y0, x1, y1]`); +/// 0 when the origin is outside or the ray misses it. +pub fn ray_extent(origin: (f64, f64), dir: (f64, f64), rect: [f64; 4]) -> f64 { + let mut t_min = f64::NEG_INFINITY; + let mut t_max = f64::INFINITY; + for (o, d, lo, hi) in [ + (origin.0, dir.0, rect[0], rect[2]), + (origin.1, dir.1, rect[1], rect[3]), + ] { + if d.abs() < 1e-12 { + if o < lo || o > hi { + return 0.0; + } + continue; + } + let (t0, t1) = ((lo - o) / d, (hi - o) / d); + t_min = t_min.max(t0.min(t1)); + t_max = t_max.min(t0.max(t1)); + } + if t_min > 1e-9 || t_max < 0.0 || t_min > t_max || !t_max.is_finite() { + return 0.0; + } + t_max +} + +/// Axis-aligned bounding box of `points`; `None` when empty. +pub fn bbox_of(points: &[(f64, f64)]) -> Option<[f64; 4]> { + let mut it = points.iter(); + let &(x, y) = it.next()?; + Some(it.fold([x, y, x, y], |b, &(x, y)| { + [b[0].min(x), b[1].min(y), b[2].max(x), b[3].max(y)] + })) +} + +/// AABB of `rect` (`[x0, y0, x1, y1]`) mapped by `m`. +pub fn transform_rect(m: &Matrix, rect: [f64; 4]) -> [f64; 4] { + let corners = [ + apply(m, rect[0], rect[1]), + apply(m, rect[2], rect[1]), + apply(m, rect[0], rect[3]), + apply(m, rect[2], rect[3]), + ]; + bbox_of(&corners).unwrap_or(rect) +} + +/// Whether `m` maps axis-aligned rectangles to axis-aligned rectangles (no shear or off-axis +/// rotation other than multiples of 90°). +pub fn axis_aligned(m: &Matrix) -> bool { + let s = m[0].abs().max(m[1].abs()).max(m[2].abs()).max(m[3].abs()); + let tol = AXIS_EPSILON_REL * s; + (m[1].abs() <= tol && m[2].abs() <= tol) || (m[0].abs() <= tol && m[3].abs() <= tol) +} + +/// Intersection of two `[x0, y0, x1, y1]` rectangles (possibly empty: x1 < x0 kept as zero size). +pub fn intersect(a: [f64; 4], b: [f64; 4]) -> [f64; 4] { + let r = [ + a[0].max(b[0]), + a[1].max(b[1]), + a[2].min(b[2]), + a[3].min(b[3]), + ]; + [r[0], r[1], r[2].max(r[0]), r[3].max(r[1])] +} + +/// Whether `inner` lies inside `outer` within `tol` on every side. +pub fn contains(outer: [f64; 4], inner: [f64; 4], tol: f64) -> bool { + inner[0] >= outer[0] - tol + && inner[1] >= outer[1] - tol + && inner[2] <= outer[2] + tol + && inner[3] <= outer[3] + tol +} + +/// Unit vector of `v`; `None` when its length is not positive and finite. +pub fn unit(v: (f64, f64)) -> Option<(f64, f64)> { + let n = v.0.hypot(v.1); + (n.is_finite() && n > 1e-12).then(|| (v.0 / n, v.1 / n)) +} + +pub fn dot(a: (f64, f64), b: (f64, f64)) -> f64 { + a.0 * b.0 + a.1 * b.1 +} + +pub fn cross(a: (f64, f64), b: (f64, f64)) -> f64 { + a.0 * b.1 - a.1 * b.0 +} + +pub fn sub(a: (f64, f64), b: (f64, f64)) -> (f64, f64) { + (a.0 - b.0, a.1 - b.1) +} diff --git a/src-tauri/src/pdf_engine/text_edit/graph.rs b/src-tauri/src/pdf_engine/text_edit/graph.rs new file mode 100644 index 0000000..4545fa7 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/graph.rs @@ -0,0 +1,600 @@ +//! Canonical, object-id-free digest of a whole document (SPEC §B.16.1, D28): Phase A check A2 +//! proves that qpdf's output equals qpdf's input everywhere except the edited content streams. +//! +//! Traversal: BFS over canonical indices — 0 = trailer `/Root`, 1 = trailer `/Info` when present; +//! every reference becomes its canonical index (assigned on first sight). Direct values serialise +//! recursively with dictionary keys sorted bytewise, names and strings by their decoded bytes, +//! integers as i64 and reals by their f32 bits. qpdf's known normalisations are applied on both +//! sides: dictionary entries whose value is null (or a reference to nothing) are absent, a dangling +//! reference in an array is `null`, and `/Root /Extensions` (shallow) and its `/ADBE` (deep) are +//! serialised inline. Streams: the node is the dictionary without `/Length /Filter /DecodeParms +//! /DL`; the data key is `Replaced` for an edited part (its expected decoded bytes), `Plain` for an +//! unfiltered stream and `Raw` (filter signature + raw bytes) for a filtered one, which qpdf +//! (`--decode-level=none`) must keep byte for byte. The filter signature is serialised through the +//! traversal, so objects reached only through `/DecodeParms` (`/JBIG2Globals`) are compared too. +//! An edited stream reached from more than one place is never `Replaced` (it is shared: the edit +//! would change the other place too). + +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget, DecodeError}; +use crate::pdf_engine::text_edit::limits::GRAPH_DIRECT_DEPTH_MAX; +use crate::pdf_engine::text_edit::snapshot::fnv1a_extend; +use lopdf::{Dictionary, Document, Object, ObjectId, Stream}; +use std::collections::{HashMap, VecDeque}; +use std::sync::atomic::{AtomicBool, Ordering}; + +const FNV_OFFSET: u64 = 0xcbf2_9ce4_8422_2325; +const STREAM_KEYS: [&[u8]; 4] = [b"Length", b"Filter", b"DecodeParms", b"DL"]; +/// Objects processed between two looks at the cancel flag. +const CANCEL_EVERY: usize = 4_096; +/// Longest path kept for a mismatch message. +const PATH_MAX_CHARS: usize = 512; +/// The `/Root` of a trailer that has none. +static NULL_OBJECT: Object = Object::Null; + +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum DataKey { + None, + Raw { + filter_sig: u64, + raw: u64, + }, + /// `len`: the decoded length (the after side never decodes more than that). + Plain { + hash: u64, + len: u64, + }, + Replaced { + expected: u64, + len: u64, + }, +} + +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct GraphEntry { + /// Hash of the canonical serialisation (references as canonical indices). + pub node: u64, + pub data: DataKey, +} + +#[derive(Debug, Clone, Default)] +pub struct GraphDigest { + pub entries: Vec, +} + +#[derive(Debug, Clone, PartialEq)] +pub struct GraphMismatch { + pub index: usize, + pub path: String, + /// node|data|missing|extra|budget|cancelled + pub what: &'static str, +} + +struct Hasher(u64); + +impl Hasher { + fn new() -> Hasher { + Hasher(FNV_OFFSET) + } + fn tag(&mut self, t: u8) { + self.0 = fnv1a_extend(self.0, &[t]); + } + fn bytes(&mut self, b: &[u8]) { + self.0 = fnv1a_extend(self.0, &(b.len() as u64).to_le_bytes()); + self.0 = fnv1a_extend(self.0, b); + } + fn u64(&mut self, v: u64) { + self.0 = fnv1a_extend(self.0, &v.to_le_bytes()); + } +} + +fn hash_bytes(b: &[u8]) -> u64 { + fnv1a_extend(FNV_OFFSET, b) +} + +/// How references inside a value are serialised. +#[derive(Clone, Copy, PartialEq, Eq)] +enum Inline { + /// As canonical indices. + No, + /// Every reference resolved inline (qpdf `makeDirect`). + Deep, +} + +/// A canonical traversal of one document. +struct Traversal<'a> { + doc: &'a Document, + index: HashMap, + /// Per canonical index: the object (or `None` for a direct trailer `/Info`). + queue: VecDeque, + targets: Vec>, + paths: Vec, + incoming: Vec, + parent: Vec>, +} + +#[derive(Clone, Copy)] +enum Target<'a> { + Ref(ObjectId), + Direct(&'a Object), +} + +fn null_like(doc: &Document, obj: &Object) -> bool { + match obj { + Object::Null => true, + Object::Reference(id) => matches!(doc.objects.get(id), None | Some(Object::Null)), + _ => false, + } +} + +fn join_path(base: &str, part: &str) -> String { + let mut p = String::with_capacity(base.len() + part.len()); + p.push_str(base); + p.push_str(part); + if p.chars().count() > PATH_MAX_CHARS { + p = p.chars().take(PATH_MAX_CHARS).collect::() + "…"; + } + p +} + +/// Only containers and references can lead to a new canonical index (whose path is kept). +fn needs_path(obj: &Object) -> bool { + matches!( + obj, + Object::Reference(_) | Object::Array(_) | Object::Dictionary(_) | Object::Stream(_) + ) +} + +fn key_path(key: &[u8]) -> String { + format!("/{}", String::from_utf8_lossy(key)) +} + +impl<'a> Traversal<'a> { + fn new(doc: &'a Document) -> Traversal<'a> { + let mut t = Traversal { + doc, + index: HashMap::new(), + queue: VecDeque::new(), + targets: Vec::new(), + paths: Vec::new(), + incoming: Vec::new(), + parent: Vec::new(), + }; + let seed = |t: &mut Traversal<'a>, key: &[u8], path: &str| { + if let Ok(v) = doc.trailer.get(key) { + let target = match v { + Object::Reference(id) => Target::Ref(*id), + other => Target::Direct(other), + }; + t.push(target, path.to_string(), None); + } + }; + seed(&mut t, b"Root", "/Root"); + if t.targets.is_empty() { + t.push(Target::Direct(&NULL_OBJECT), "/Root".to_string(), None); + } + seed(&mut t, b"Info", "/Info"); + t + } + + fn push(&mut self, target: Target<'a>, path: String, parent: Option) -> usize { + let i = self.targets.len(); + if let Target::Ref(id) = target { + self.index.insert(id, i); + } + self.targets.push(target); + self.paths.push(path); + self.incoming.push(0); + self.parent.push(parent); + self.queue.push_back(i); + i + } + + /// The canonical index of `id` (assigned on first sight), counting the reference. + fn index_of(&mut self, id: ObjectId, path: &str, from: usize) -> usize { + let i = match self.index.get(&id) { + Some(i) => *i, + None => self.push(Target::Ref(id), path.to_string(), Some(from)), + }; + if let Some(c) = self.incoming.get_mut(i) { + *c = c.saturating_add(1); + } + i + } + + fn ser( + &mut self, + h: &mut Hasher, + obj: &'a Object, + depth: usize, + path: &str, + from: usize, + inline: Inline, + ) -> Result<(), &'static str> { + if depth > GRAPH_DIRECT_DEPTH_MAX { + return Err("budget"); + } + match obj { + Object::Null => h.tag(b'n'), + Object::Boolean(b) => { + h.tag(b'b'); + h.tag(u8::from(*b)); + } + Object::Integer(i) => { + h.tag(b'i'); + h.u64(*i as u64); + } + Object::Real(r) => { + h.tag(b'r'); + h.u64(u64::from(r.to_bits())); + } + Object::Name(n) => { + h.tag(b'N'); + h.bytes(n); + } + Object::String(s, _) => { + h.tag(b'S'); + h.bytes(s); + } + Object::Array(items) => { + h.tag(b'['); + h.u64(items.len() as u64); + for (k, item) in items.iter().enumerate() { + if null_like(self.doc, item) { + h.tag(b'n'); + } else if needs_path(item) { + let p = join_path(path, &format!("[{k}]")); + self.ser(h, item, depth + 1, &p, from, inline)?; + } else { + self.ser(h, item, depth + 1, "", from, inline)?; + } + } + } + Object::Dictionary(d) => self.ser_dict(h, d, &[], depth, path, from, inline, false)?, + Object::Stream(s) => { + h.tag(b'X'); + self.ser_dict(h, &s.dict, &STREAM_KEYS, depth, path, from, inline, false)?; + } + Object::Reference(id) => match self.doc.objects.get(id) { + None | Some(Object::Null) => h.tag(b'n'), + Some(target) if inline == Inline::Deep => { + if matches!(target, Object::Stream(_)) { + return Err("node"); + } + self.ser(h, target, depth + 1, path, from, inline)?; + } + Some(_) => { + let i = self.index_of(*id, path, from); + h.tag(b'R'); + h.u64(i as u64); + } + }, + } + Ok(()) + } + + #[allow(clippy::too_many_arguments)] + fn ser_dict( + &mut self, + h: &mut Hasher, + d: &'a Dictionary, + skip: &[&[u8]], + depth: usize, + path: &str, + from: usize, + inline: Inline, + catalog: bool, + ) -> Result<(), &'static str> { + let mut entries: Vec<(&'a Vec, &'a Object)> = d + .iter() + .filter(|(k, v)| !skip.contains(&k.as_slice()) && !null_like(self.doc, v)) + .collect(); + entries.sort_by(|a, b| a.0.cmp(b.0)); + h.tag(b'<'); + h.u64(entries.len() as u64); + for (k, v) in entries { + h.bytes(k); + let p = if needs_path(v) { + join_path(path, &key_path(k)) + } else { + String::new() + }; + if catalog && k.as_slice() == b"Extensions" { + self.ser_extensions(h, v, depth + 1, &p, from)?; + } else { + self.ser(h, v, depth + 1, &p, from, inline)?; + } + } + Ok(()) + } + + /// `/Root /Extensions`: the dictionary itself inline (qpdf makes it direct), and its `/ADBE` + /// entry deeply inline when it is a reference (qpdf `makeDirect`). + fn ser_extensions( + &mut self, + h: &mut Hasher, + v: &'a Object, + depth: usize, + path: &str, + from: usize, + ) -> Result<(), &'static str> { + let resolved = match v { + Object::Reference(id) => match self.doc.objects.get(id) { + Some(o) => o, + None => { + h.tag(b'n'); + return Ok(()); + } + }, + other => other, + }; + let Object::Dictionary(ext) = resolved else { + return self.ser(h, resolved, depth, path, from, Inline::No); + }; + let mut entries: Vec<(&'a Vec, &'a Object)> = ext + .iter() + .filter(|(_, v)| !null_like(self.doc, v)) + .collect(); + entries.sort_by(|a, b| a.0.cmp(b.0)); + h.tag(b'<'); + h.u64(entries.len() as u64); + for (k, v) in entries { + h.bytes(k); + let p = join_path(path, &key_path(k)); + let inline = if k.as_slice() == b"ADBE" && matches!(v, Object::Reference(_)) { + Inline::Deep + } else { + Inline::No + }; + self.ser(h, v, depth + 1, &p, from, inline)?; + } + Ok(()) + } + + /// The node hash of canonical index `i` and, for a stream, the stream. + fn node(&mut self, i: usize) -> Result<(u64, Option<&'a Stream>), &'static str> { + let target = self.targets.get(i).copied(); + let path = self.paths.get(i).cloned().unwrap_or_default(); + let mut h = Hasher::new(); + let obj: &'a Object = match target { + Some(Target::Ref(id)) => match self.doc.objects.get(&id) { + Some(o) => o, + None => { + h.tag(b'n'); + return Ok((h.0, None)); + } + }, + Some(Target::Direct(o)) => o, + None => return Err("missing"), + }; + match obj { + Object::Stream(s) => { + h.tag(b'T'); + self.ser_dict( + &mut h, + &s.dict, + &STREAM_KEYS, + 0, + &path, + i, + Inline::No, + false, + )?; + Ok((h.0, Some(s))) + } + Object::Dictionary(d) => { + self.ser_dict(&mut h, d, &[], 0, &path, i, Inline::No, i == 0)?; + Ok((h.0, None)) + } + other => { + self.ser(&mut h, other, 0, &path, i, Inline::No)?; + Ok((h.0, None)) + } + } + } + + fn id_of(&self, i: usize) -> Option { + match self.targets.get(i) { + Some(Target::Ref(id)) => Some(*id), + _ => None, + } + } +} + +/// Signature of a stream whose filter signature cannot be part of the traversal any more. +const DETACHED_SIG: u64 = 0; + +impl<'a> Traversal<'a> { + /// `/Filter` and `/DecodeParms` of stream `i` as one signature, serialised through this + /// traversal: a reference among them (an indirect parameter dictionary, a JBIG2 image's + /// `/JBIG2Globals` stream) gets its canonical index and is compared like any other object. + fn filter_sig(&mut self, i: usize, s: &'a Stream) -> Result { + let path = self.paths.get(i).cloned().unwrap_or_default(); + let mut h = Hasher::new(); + for key in [&b"Filter"[..], b"DecodeParms"] { + h.bytes(key); + match s.dict.get(key) { + Ok(v) if !null_like(self.doc, v) => { + let p = if needs_path(v) { + join_path(&path, &key_path(key)) + } else { + String::new() + }; + self.ser(&mut h, v, 1, &p, i, Inline::No)?; + } + _ => h.tag(b'-'), + } + } + Ok(h.0) + } + + /// The data key of an unedited stream `i`: `Plain` when unfiltered (a `/DecodeParms` without + /// `/Filter` has no effect and is not followed), else `Raw` with its filter signature. + fn original_key(&mut self, i: usize, s: &'a Stream) -> Result { + if unfiltered(s) { + return Ok(plain_key(s)); + } + Ok(DataKey::Raw { + filter_sig: self.filter_sig(i, s)?, + raw: hash_bytes(&s.content), + }) + } +} + +fn plain_key(s: &Stream) -> DataKey { + DataKey::Plain { + hash: hash_bytes(&s.content), + len: s.content.len() as u64, + } +} + +fn unfiltered(stream: &Stream) -> bool { + match stream.dict.get(b"Filter") { + Err(_) | Ok(Object::Null) => true, + Ok(Object::Array(items)) => items.is_empty(), + Ok(_) => false, + } +} + +/// The key of a shared edited stream, decided after the traversal: its original data, so any +/// change fails. qpdf writes the edit into it unfiltered, so a filtered original can never match +/// on the after side and its filter signature is not needed (`DETACHED_SIG`). +fn shared_key(s: &Stream) -> DataKey { + if unfiltered(s) { + plain_key(s) + } else { + DataKey::Raw { + filter_sig: DETACHED_SIG, + raw: hash_bytes(&s.content), + } + } +} + +fn cancelled(cancel: Option<&AtomicBool>) -> bool { + cancel.is_some_and(|c| c.load(Ordering::Relaxed)) +} + +/// Canonical traversal of `doc` from trailer `/Root` then `/Info`. `replaced` maps the ids of the +/// edited content streams **in this document** to their expected decoded bytes. +pub fn graph_digest( + doc: &Document, + replaced: &HashMap>, + cancel: Option<&AtomicBool>, +) -> Result { + let mut t = Traversal::new(doc); + let mut entries: Vec = Vec::new(); + let mut done = 0usize; + while let Some(i) = t.queue.pop_front() { + done += 1; + if done % CANCEL_EVERY == 0 && cancelled(cancel) { + return Err(mismatch(&t, i, "cancelled")); + } + let (node, stream) = t.node(i).map_err(|what| mismatch(&t, i, what))?; + let data = match stream { + None => DataKey::None, + Some(s) => match t.id_of(i).and_then(|id| replaced.get(&id)) { + Some(bytes) => DataKey::Replaced { + expected: hash_bytes(bytes), + len: bytes.len() as u64, + }, + None => t.original_key(i, s).map_err(|what| mismatch(&t, i, what))?, + }, + }; + entries.push(GraphEntry { node, data }); + } + // A replaced stream (or its `/Contents` array) reached from more than one place is shared: + // keep its original data so any change to it fails. + for (i, e) in entries.iter_mut().enumerate() { + if !matches!(e.data, DataKey::Replaced { .. }) { + continue; + } + let parent_shared = t.parent.get(i).copied().flatten().is_some_and(|p| { + let array = |id: &ObjectId| matches!(doc.objects.get(id), Some(Object::Array(_))); + matches!(t.targets.get(p), Some(Target::Ref(id)) if array(id)) + && t.incoming.get(p).copied().unwrap_or(0) > 1 + }); + let shared = t.incoming.get(i).copied().unwrap_or(0) > 1 || parent_shared; + if shared { + if let Some(Object::Stream(s)) = t.id_of(i).and_then(|id| doc.objects.get(&id)) { + e.data = shared_key(s); + } + } + } + Ok(GraphDigest { entries }) +} + +fn mismatch(t: &Traversal<'_>, i: usize, what: &'static str) -> GraphMismatch { + GraphMismatch { + index: i, + path: t.paths.get(i).cloned().unwrap_or_default(), + what, + } +} + +/// Same traversal over `doc`, compared entry by entry with `before`; a stream is decoded only when +/// `before` holds `Plain` or `Replaced` for that entry (never beyond its expected length). The +/// first mismatch is returned with its canonical path in this document. +pub fn graph_matches( + doc: &Document, + before: &GraphDigest, + budget: &mut DecodeBudget, + cancel: Option<&AtomicBool>, +) -> Result<(), GraphMismatch> { + let mut t = Traversal::new(doc); + let mut done = 0usize; + while let Some(i) = t.queue.pop_front() { + done += 1; + if done % CANCEL_EVERY == 0 && cancelled(cancel) { + return Err(mismatch(&t, i, "cancelled")); + } + let Some(want) = before.entries.get(i) else { + return Err(mismatch(&t, i, "extra")); + }; + let (node, stream) = t.node(i).map_err(|what| mismatch(&t, i, what))?; + if node != want.node { + return Err(mismatch(&t, i, "node")); + } + let ok = match (&want.data, stream) { + (DataKey::None, None) => true, + (DataKey::None, Some(_)) | (_, None) => false, + ( + DataKey::Raw { + filter_sig: sig, + raw, + }, + Some(s), + ) => { + !unfiltered(s) + && hash_bytes(&s.content) == *raw + && t.filter_sig(i, s).map_err(|what| mismatch(&t, i, what))? == *sig + } + (DataKey::Plain { hash, len }, Some(s)) + | ( + DataKey::Replaced { + expected: hash, + len, + }, + Some(s), + ) => { + let cap = usize::try_from(*len).unwrap_or(usize::MAX); + match decode_stream(s, cap, budget) { + Ok(data) => hash_bytes(&data) == *hash && data.len() as u64 == *len, + Err(DecodeError::TooLarge) if budget.remaining() < cap => { + return Err(mismatch(&t, i, "budget")) + } + Err(_) => false, + } + } + }; + if !ok { + return Err(mismatch(&t, i, "data")); + } + } + if t.targets.len() < before.entries.len() { + return Err(GraphMismatch { + index: t.targets.len(), + path: String::new(), + what: "missing", + }); + } + Ok(()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/lexer.rs b/src-tauri/src/pdf_engine/text_edit/lexer.rs new file mode 100644 index 0000000..4e9cc15 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/lexer.rs @@ -0,0 +1,743 @@ +//! Byte-offset content tokenizer (SPEC §B.6). Iterative, bounded, full consumption: every op +//! carries the exact byte span it was read from, trailing operands or stray delimiters are +//! errors (never a silently truncated prefix), and inline-image payload ends are proven +//! (§A.4) or marked unproven for every op that follows. +//! +//! Operand nodes are counted as they are read (`LexLimits::operand_nodes_max`), and every value +//! is kept at exactly its size (strings, names, arrays, dictionaries and an op's operand list +//! hold no spare room), so a lexed page costs `nodes × size_of::()` plus the bytes it +//! copied, which the walker's page-model budget can bound before it lexes. + +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::reasons::TextReason; +use std::sync::atomic::{AtomicBool, Ordering}; + +mod arity; +mod inline; +mod scan; +mod values; + +pub use arity::check_arity; +pub(crate) use inline::InlineEnds; +use values::Builder; + +/// `[start, end)` in the buffer that was lexed. +pub type Span = std::ops::Range; + +#[derive(Debug, Clone, PartialEq)] +pub enum Operand { + Number { + value: f64, + span: Span, + }, // [+-]?(\d+\.?\d*|\.\d+); |x| ≤ 1e9 in content + Name { + bytes: Vec, + span: Span, + }, // #xx decoded + Str { + bytes: Vec, + hex: bool, + span: Span, + }, // escapes decoded; CR/CRLF → LF + Array { + items: Vec, + span: Span, + }, + Dict { + entries: Vec<(Vec, Operand)>, + span: Span, + }, // only as operands of BDC/DP and inside BI + Bool { + value: bool, + span: Span, + }, + Null { + span: Span, + }, +} + +impl Operand { + pub fn span(&self) -> &Span { + match self { + Operand::Number { span, .. } + | Operand::Name { span, .. } + | Operand::Str { span, .. } + | Operand::Array { span, .. } + | Operand::Dict { span, .. } + | Operand::Bool { span, .. } + | Operand::Null { span } => span, + } + } + pub fn as_number(&self) -> Option { + match self { + Operand::Number { value, .. } => Some(*value), + _ => None, + } + } + pub fn as_name(&self) -> Option<&[u8]> { + match self { + Operand::Name { bytes, .. } => Some(bytes), + _ => None, + } + } + pub fn as_str_bytes(&self) -> Option<&[u8]> { + match self { + Operand::Str { bytes, .. } => Some(bytes), + _ => None, + } + } +} + +#[rustfmt::skip] +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] +#[allow(non_camel_case_types)] // variants mirror the operator spelling; `b`/`B`, `f`/`F` … would collide in CamelCase +pub enum Operator { + b, B, bStar, BStar, BDC, BI, BMC, BT, BX, c, cm, CS, cs, d, d0, d1, Do, DP, EI, EMC, ET, EX, + f, F, fStar, G, g, gs, h, i, ID, j, J, K, k, l, m, M, MP, n, q, Q, re, RG, rg, ri, s, S, SC, sc, + SCN, scn, sh, TStar, Tc, Td, TD, Tf, Tj, TJ, TL, Tm, Tr, Ts, Tw, Tz, v, w, W, WStar, y, + Quote, DoubleQuote, + Unknown, // only legal inside BX … EX +} + +use Operator as O; + +#[rustfmt::skip] +const OPERATORS: [(Operator, &str); 73] = [ + (O::b, "b"), (O::B, "B"), (O::bStar, "b*"), (O::BStar, "B*"), (O::BDC, "BDC"), (O::BI, "BI"), + (O::BMC, "BMC"), (O::BT, "BT"), (O::BX, "BX"), (O::c, "c"), (O::cm, "cm"), (O::CS, "CS"), + (O::cs, "cs"), (O::d, "d"), (O::d0, "d0"), (O::d1, "d1"), (O::Do, "Do"), (O::DP, "DP"), + (O::EI, "EI"), (O::EMC, "EMC"), (O::ET, "ET"), (O::EX, "EX"), (O::f, "f"), (O::F, "F"), + (O::fStar, "f*"), (O::G, "G"), (O::g, "g"), (O::gs, "gs"), (O::h, "h"), (O::i, "i"), + (O::ID, "ID"), (O::j, "j"), (O::J, "J"), (O::K, "K"), (O::k, "k"), (O::l, "l"), (O::m, "m"), + (O::M, "M"), (O::MP, "MP"), (O::n, "n"), (O::q, "q"), (O::Q, "Q"), (O::re, "re"), + (O::RG, "RG"), (O::rg, "rg"), (O::ri, "ri"), (O::s, "s"), (O::S, "S"), (O::SC, "SC"), + (O::sc, "sc"), (O::SCN, "SCN"), (O::scn, "scn"), (O::sh, "sh"), (O::TStar, "T*"), + (O::Tc, "Tc"), (O::Td, "Td"), (O::TD, "TD"), (O::Tf, "Tf"), (O::Tj, "Tj"), (O::TJ, "TJ"), + (O::TL, "TL"), (O::Tm, "Tm"), (O::Tr, "Tr"), (O::Ts, "Ts"), (O::Tw, "Tw"), (O::Tz, "Tz"), + (O::v, "v"), (O::w, "w"), (O::W, "W"), (O::WStar, "W*"), (O::y, "y"), (O::Quote, "'"), + (O::DoubleQuote, "\""), +]; + +impl Operator { + pub fn from_token(t: &[u8]) -> Operator { + OPERATORS + .iter() + .find(|(_, s)| s.as_bytes() == t) + .map(|(o, _)| *o) + .unwrap_or(O::Unknown) + } + pub fn as_str(self) -> &'static str { + OPERATORS + .iter() + .find(|(o, _)| *o == self) + .map(|(_, s)| *s) + .unwrap_or("?") + } +} + +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum InlineProof { + LengthKey, + UnfilteredSize, + FlateEnd, + AsciiEnd, + Heuristic, +} + +#[derive(Debug, Clone, PartialEq)] +pub struct InlineImage { + pub dict: Vec<(Vec, Operand)>, + pub data: Span, + pub proof: InlineProof, +} + +#[derive(Debug, Clone, PartialEq)] +pub struct Op { + pub operator: Operator, + pub operands: Vec, + pub span: Span, // first operand start (after whitespace/comments) or operator start .. operator end; BI: .. end of "EI" + pub op_span: Span, // the operator token itself + pub in_compat: bool, // inside BX … EX + pub after_unproven_inline_image: bool, + pub inline_image: Option, +} + +pub struct LexLimits { + pub ops_max: usize, + pub nesting_max: usize, + pub array_items_max: usize, + pub string_bytes_max: usize, + pub paren_nesting_max: usize, + pub inline_image_max: usize, + /// Operand nodes (numbers, names, strings, booleans, nulls, arrays, dictionaries — nested + /// ones, dictionary keys and inline-image values included) one lex may build, counted as each + /// is read: past it the lex is `TooComplex { what: OPERAND_NODES }` before the next node is + /// kept (48 MiB of `1 1 1 …` would otherwise build ~24 M nodes, ≈ 1.2 GiB, review T3 r3 + /// MEDIUM-3). The walker lowers it to what the page may still hold. + pub operand_nodes_max: usize, +} + +/// `LexError::TooComplex` detail of the operand-node cap. +pub const OPERAND_NODES: &str = "operand nodes"; + +impl LexLimits { + pub fn page() -> Self { + LexLimits { + ops_max: limits::PAGE_OPS_MAX, + nesting_max: limits::TOKEN_NESTING_MAX, + array_items_max: limits::ARRAY_ITEMS_MAX, + string_bytes_max: limits::STRING_BYTES_MAX, + paren_nesting_max: limits::LITERAL_PAREN_NESTING_MAX, + inline_image_max: limits::INLINE_IMAGE_MAX_BYTES, + operand_nodes_max: limits::OPERAND_NODES_MAX, + } + } +} + +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum LexError { + Malformed { + at: usize, + what: &'static str, + }, + TooComplex { + what: &'static str, + }, + /// The caller's cancel flag was set (checked every `LEX_CANCEL_EVERY_OPS` ops). + Cancelled, +} + +impl LexError { + /// MALFORMED_CONTENT | PAGE_TOO_COMPLEX (a cancelled lex is never shown as a reason). + pub fn page_reason(&self) -> TextReason { + match self { + LexError::Malformed { .. } => TextReason::MalformedContent, + LexError::TooComplex { .. } | LexError::Cancelled => TextReason::PageTooComplex, + } + } +} + +impl std::fmt::Display for LexError { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + match self { + LexError::Malformed { at, what } => write!(f, "malformed content at byte {at}: {what}"), + LexError::TooComplex { what } => write!(f, "content too complex: {what}"), + LexError::Cancelled => write!(f, "cancelled"), + } + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum ScanMode { + Object, + CMap, + Type1Clear, +} + +#[derive(Debug, Clone, PartialEq)] +pub enum Token { + Operand(Operand), + Keyword { bytes: Vec, span: Span }, + DictOpen(usize), + DictClose(usize), + ArrayOpen(usize), + ArrayClose(usize), +} + +fn malformed(at: usize, what: &'static str) -> Result { + Err(LexError::Malformed { at, what }) +} + +pub(crate) fn is_whitespace(c: u8) -> bool { + matches!(c, 0 | 9 | 10 | 12 | 13 | 32) +} + +pub(crate) fn is_delimiter(c: u8) -> bool { + matches!( + c, + b'(' | b')' | b'<' | b'>' | b'[' | b']' | b'{' | b'}' | b'/' | b'%' + ) +} + +fn is_regular(c: u8) -> bool { + !is_whitespace(c) && !is_delimiter(c) +} + +/// Validates `[+-]?(\d+\.?\d*|\.\d+)` and parses it from the exact digits. +fn parse_number(t: &[u8]) -> Option { + let (neg, body) = match t.first() { + Some(b'+') => (false, t.get(1..)?), + Some(b'-') => (true, t.get(1..)?), + _ => (false, t), + }; + let dot = body.iter().position(|c| *c == b'.'); + let (int, frac) = match dot { + Some(p) => (body.get(..p)?, body.get(p + 1..)?), + None => (body, &b""[..]), + }; + if (int.is_empty() && frac.is_empty()) || !int.iter().chain(frac).all(u8::is_ascii_digit) { + return None; + } + let text = format!( + "{}{}.{}", + if neg { "-" } else { "" }, + if int.is_empty() { + "0" + } else { + std::str::from_utf8(int).ok()? + }, + if frac.is_empty() { + "0" + } else { + std::str::from_utf8(frac).ok()? + } + ); + text.parse::().ok().filter(|v| v.is_finite()) +} + +#[derive(Debug)] +enum Raw { + Operand(Operand), + Word(Span), + DictOpen(usize), + DictClose(usize), + ArrayOpen(usize), + ArrayClose(usize), + BraceOpen(usize), + BraceClose(usize), +} + +impl Raw { + fn start(&self) -> usize { + match self { + Raw::Operand(o) => o.span().start, + Raw::Word(s) => s.start, + Raw::DictOpen(p) | Raw::DictClose(p) | Raw::ArrayOpen(p) | Raw::ArrayClose(p) => *p, + Raw::BraceOpen(p) | Raw::BraceClose(p) => *p, + } + } +} + +struct Reader<'a> { + b: &'a [u8], + pos: usize, + content: bool, // content stream: strict numbers ≤ 1e9, braces illegal + lenient: bool, // CMap/Type1: invalid numbers become words, braces are words + string_max: usize, + paren_max: usize, + nodes: usize, // operand nodes read so far (`LexLimits::operand_nodes_max`) + nodes_max: usize, + /// The last forward search for each inline-image ASCII terminator (`>`, `~>`): where it + /// started and the first hit at or after it (`inline::find_terminator`). + terminators: [Option<(usize, Option)>; 2], +} + +impl<'a> Reader<'a> { + fn new(b: &'a [u8], content: bool, lenient: bool, lim: &LexLimits) -> Self { + Reader { + b, + pos: 0, + content, + lenient, + string_max: lim.string_bytes_max, + paren_max: lim.paren_nesting_max, + nodes: 0, + nodes_max: lim.operand_nodes_max, + terminators: [None; 2], + } + } + + /// Counts the node `raw` makes (a value, or an array or dictionary it opens). + fn count_node(&mut self, raw: &Raw) -> Result<(), LexError> { + if matches!(raw, Raw::Operand(_) | Raw::ArrayOpen(_) | Raw::DictOpen(_)) { + self.nodes = self.nodes.saturating_add(1); + if self.nodes > self.nodes_max { + return Err(LexError::TooComplex { + what: OPERAND_NODES, + }); + } + } + Ok(()) + } + + fn at(&self, i: usize) -> Option { + self.b.get(i).copied() + } + + fn skip_ws(&mut self) { + while let Some(c) = self.at(self.pos) { + if is_whitespace(c) { + self.pos += 1; + } else if c == b'%' { + while let Some(c) = self.at(self.pos) { + if c == b'\r' || c == b'\n' { + break; + } + self.pos += 1; + } + } else { + break; + } + } + } + + fn next(&mut self) -> Result, LexError> { + self.skip_ws(); + let start = self.pos; + let Some(c) = self.at(start) else { + return Ok(None); + }; + let raw = match c { + b'(' => Raw::Operand(self.literal_string()?), + b'<' if self.at(start + 1) == Some(b'<') => { + self.pos += 2; + Raw::DictOpen(start) + } + b'<' => Raw::Operand(self.hex_string()?), + b'>' if self.at(start + 1) == Some(b'>') => { + self.pos += 2; + Raw::DictClose(start) + } + b'>' => return malformed(start, "stray >"), + b')' => return malformed(start, "stray )"), + b'[' | b']' | b'{' | b'}' => { + self.pos += 1; + match c { + b'[' => Raw::ArrayOpen(start), + b']' => Raw::ArrayClose(start), + _ if self.content => return malformed(start, "brace in content"), + b'{' => Raw::BraceOpen(start), + _ => Raw::BraceClose(start), + } + } + b'/' => Raw::Operand(self.name()?), + _ => self.regular()?, + }; + self.count_node(&raw)?; + Ok(Some(raw)) + } + + fn regular(&mut self) -> Result { + let start = self.pos; + while self.at(self.pos).is_some_and(is_regular) { + self.pos += 1; + } + let span = start..self.pos; + let tok = self.b.get(span.clone()).unwrap_or_default(); + let span_c = span.clone(); + let op = |value| Raw::Operand(value); + match tok { + b"true" => return Ok(op(Operand::Bool { value: true, span })), + b"false" => return Ok(op(Operand::Bool { value: false, span })), + b"null" => return Ok(op(Operand::Null { span })), + _ => {} + } + let numeric = tok + .first() + .is_some_and(|c| c.is_ascii_digit() || matches!(c, b'+' | b'-' | b'.')); + if !numeric { + return Ok(Raw::Word(span)); + } + match parse_number(tok) { + Some(value) if !self.content || value.abs() <= limits::NUMBER_ABS_MAX => { + Ok(op(Operand::Number { value, span })) + } + Some(_) => malformed(start, "number too large"), + None if self.lenient => Ok(Raw::Word(span_c)), + None => malformed(start, "bad number"), + } + } + + fn name(&mut self) -> Result { + let start = self.pos; + self.pos += 1; + let mut bytes = Vec::new(); + while let Some(c) = self.at(self.pos).filter(|c| is_regular(*c)) { + let hex = ( + self.at(self.pos + 1).and_then(hex_digit), + self.at(self.pos + 2).and_then(hex_digit), + ); + match (c, hex) { + (b'#', (Some(h), Some(l))) => { + bytes.push((h << 4) | l); + self.pos += 3; + } + _ => { + bytes.push(c); + self.pos += 1; + } + } + if bytes.len() > self.string_max { + return Err(LexError::TooComplex { + what: "name length", + }); + } + } + bytes.shrink_to_fit(); + Ok(Operand::Name { + bytes, + span: start..self.pos, + }) + } + + fn literal_string(&mut self) -> Result { + let start = self.pos; + self.pos += 1; + let mut depth = 1usize; + let mut bytes = Vec::new(); + loop { + let Some(c) = self.at(self.pos) else { + return malformed(start, "unterminated string"); + }; + self.pos += 1; + match c { + b'\\' => { + let Some(e) = self.at(self.pos) else { + return malformed(start, "unterminated string"); + }; + self.pos += 1; + match e { + b'n' => bytes.push(b'\n'), + b'r' => bytes.push(b'\r'), + b't' => bytes.push(b'\t'), + b'b' => bytes.push(8), + b'f' => bytes.push(12), + b'0'..=b'7' => { + let mut v = u32::from(e - b'0'); + for _ in 0..2 { + match self.at(self.pos) { + Some(d @ b'0'..=b'7') => { + v = v * 8 + u32::from(d - b'0'); + self.pos += 1; + } + _ => break, + } + } + bytes.push((v & 0xFF) as u8); + } + b'\r' => { + if self.at(self.pos) == Some(b'\n') { + self.pos += 1; + } + } + b'\n' => {} + other => bytes.push(other), + } + } + b'(' => { + depth += 1; + if depth > self.paren_max { + return Err(LexError::TooComplex { + what: "string nesting", + }); + } + bytes.push(c); + } + b')' => { + depth -= 1; + if depth == 0 { + break; + } + bytes.push(c); + } + b'\r' => { + if self.at(self.pos) == Some(b'\n') { + self.pos += 1; + } + bytes.push(b'\n'); + } + other => bytes.push(other), + } + if bytes.len() > self.string_max { + return Err(LexError::TooComplex { + what: "string length", + }); + } + } + bytes.shrink_to_fit(); + Ok(Operand::Str { + bytes, + hex: false, + span: start..self.pos, + }) + } + + fn hex_string(&mut self) -> Result { + let start = self.pos; + self.pos += 1; + let mut bytes = Vec::new(); + let mut high: Option = None; + loop { + let Some(c) = self.at(self.pos) else { + return malformed(start, "unterminated hex string"); + }; + self.pos += 1; + if c == b'>' { + break; + } + if is_whitespace(c) { + continue; + } + let Some(v) = hex_digit(c) else { + return malformed(start, "bad hex string"); + }; + match high.take() { + None => high = Some(v), + Some(h) => bytes.push((h << 4) | v), + } + if bytes.len() > self.string_max { + return Err(LexError::TooComplex { + what: "string length", + }); + } + } + if let Some(h) = high { + bytes.push(h << 4); + } + bytes.shrink_to_fit(); + Ok(Operand::Str { + bytes, + hex: true, + span: start..self.pos, + }) + } +} + +fn hex_digit(c: u8) -> Option { + match c { + b'0'..=b'9' => Some(c - b'0'), + b'a'..=b'f' => Some(c - b'a' + 10), + b'A'..=b'F' => Some(c - b'A' + 10), + _ => None, + } +} + +/// Lexes a content stream into ops with exact byte spans. +pub fn lex_content( + bytes: &[u8], + limits: &LexLimits, + cancel: Option<&AtomicBool>, +) -> Result, LexError> { + lex_inner(bytes, limits, cancel, false) +} + +fn lex_inner( + bytes: &[u8], + lim: &LexLimits, + cancel: Option<&AtomicBool>, + tail_check: bool, +) -> Result, LexError> { + let mut r = Reader::new(bytes, true, false, lim); + let mut builder = Builder::new(lim); + let mut ops: Vec = Vec::new(); + let mut operands: Vec = Vec::new(); + let mut op_start: Option = None; + let mut compat = 0usize; + let mut after_unproven = false; + loop { + let raw = match r.next() { + Ok(Some(raw)) => raw, + Ok(None) => break, + // a heuristic tail window may end inside a string: the ops before it count + Err(LexError::Malformed { + what: "unterminated string" | "unterminated hex string", + .. + }) if tail_check => return Ok(ops), + Err(e) => return Err(e), + }; + let start = raw.start(); + let first = *op_start.get_or_insert(start); + let span = match raw { + Raw::Word(span) if builder.frames.is_empty() => span, + other => { + if let Some(v) = builder.feed(other)? { + if operands.len() >= lim.array_items_max { + return Err(LexError::TooComplex { what: "operands" }); + } + operands.push(v); + } + continue; + } + }; + let operator = Operator::from_token(bytes.get(span.clone()).unwrap_or_default()); + let in_compat = compat > 0; + let mut op = Op { + operator, + // Exactly as many slots as operands (the scratch vector keeps its room): a one-operand + // op kept for the page's lifetime would otherwise hold room for four. + operands: operands.drain(..).collect(), + span: first..span.end, + op_span: span.clone(), + in_compat, + after_unproven_inline_image: after_unproven, + inline_image: None, + }; + op_start = None; + match operator { + O::BI => { + if tail_check { + return Ok(ops); + } + if !op.operands.is_empty() { + return malformed(first, "operands before BI"); + } + let image = inline::inline_image(&mut r, lim, span.end)?; + op.span = span.start..r.pos; + if image.proof == InlineProof::Heuristic { + after_unproven = true; + } + op.inline_image = Some(image); + } + O::ID | O::EI => return malformed(span.start, "inline image operator outside BI"), + O::Unknown if !in_compat => return malformed(span.start, "unknown operator"), + O::Unknown => {} + O::BX => { + check_arity(&op)?; + compat = compat + .checked_add(1) + .ok_or(LexError::TooComplex { what: "BX nesting" })?; + if compat > lim.nesting_max { + return Err(LexError::TooComplex { what: "BX nesting" }); + } + } + O::EX => { + check_arity(&op)?; + compat = compat.saturating_sub(1); + } + _ => check_arity(&op)?, + } + if ops.len() >= lim.ops_max { + return Err(LexError::TooComplex { what: "operations" }); + } + ops.push(op); + if ops.len() % limits::LEX_CANCEL_EVERY_OPS == 0 + && cancel.is_some_and(|c| c.load(Ordering::Relaxed)) + { + return Err(LexError::Cancelled); + } + } + if !tail_check && (!builder.frames.is_empty() || !operands.is_empty()) { + return malformed(bytes.len(), "trailing operands"); + } + Ok(ops) +} + +/// Flat tokens of a PostScript-like buffer (objects, CMaps, Type1 clear text); no arity. +/// `R` and every other keyword come back as `Keyword`; brackets must balance. +pub fn scan_tokens( + bytes: &[u8], + mode: ScanMode, + max_tokens: usize, +) -> Result, LexError> { + scan::tokens(bytes, mode, max_tokens) +} + +/// Tokens of the one dictionary starting at `start` (after whitespace/comments) through its +/// matching `>>`, and the offset just after it. Used by the snapshot preflight. +pub fn scan_dict_at( + bytes: &[u8], + start: usize, + max_tokens: usize, +) -> Result<(Vec, usize), LexError> { + scan::dict_at(bytes, start, max_tokens) +} diff --git a/src-tauri/src/pdf_engine/text_edit/lexer/arity.rs b/src-tauri/src/pdf_engine/text_edit/lexer/arity.rs new file mode 100644 index 0000000..74f09b0 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/lexer/arity.rs @@ -0,0 +1,85 @@ +//! Operand count and types for every interpreted operator (SPEC §B.6 arity table). + +use super::{malformed, LexError, Op, Operand, O}; + +/// Checks operand count and types for every interpreted operator. +pub fn check_arity(op: &Op) -> Result<(), LexError> { + let a = &op.operands; + let num = |i: usize| a.get(i).is_some_and(|o| o.as_number().is_some()); + let name = |i: usize| a.get(i).is_some_and(|o| o.as_name().is_some()); + let string = |i: usize| a.get(i).is_some_and(|o| o.as_str_bytes().is_some()); + let nums = |n: usize| a.len() == n && (0..n).all(num); + let ok = match op.operator { + O::b + | O::B + | O::bStar + | O::BStar + | O::BT + | O::ET + | O::EMC + | O::BX + | O::EX + | O::f + | O::F + | O::fStar + | O::h + | O::n + | O::q + | O::Q + | O::s + | O::S + | O::TStar + | O::W + | O::WStar + | O::BI => a.is_empty(), + O::BDC | O::DP => { + a.len() == 2 + && name(0) + && matches!(a.get(1), Some(Operand::Dict { .. } | Operand::Name { .. })) + } + O::BMC | O::MP | O::CS | O::cs | O::Do | O::gs | O::ri | O::sh => a.len() == 1 && name(0), + O::c | O::cm | O::d1 | O::Tm => nums(6), + O::d => { + a.len() == 2 + && matches!(a.first(), Some(Operand::Array { items, .. }) if items.iter().all(|i| i.as_number().is_some())) + && num(1) + } + O::d0 | O::l | O::m | O::Td | O::TD => nums(2), + O::G + | O::g + | O::i + | O::j + | O::J + | O::M + | O::w + | O::Tc + | O::Tw + | O::Tz + | O::TL + | O::Tr + | O::Ts => nums(1), + O::K | O::k | O::re | O::v | O::y => nums(4), + O::RG | O::rg => nums(3), + O::SC | O::sc => (1..=32).contains(&a.len()) && (0..a.len()).all(num), + O::SCN | O::scn => { + let n = if a.last().is_some_and(|o| o.as_name().is_some()) { + a.len() - 1 + } else { + a.len() + }; + !a.is_empty() && n <= 32 && (0..n).all(num) + } + O::Tf => a.len() == 2 && name(0) && num(1), + O::Tj | O::Quote => a.len() == 1 && string(0), + O::TJ => matches!(a.as_slice(), [Operand::Array { items, .. }] + if items.iter().all(|i| matches!(i, Operand::Str { .. } | Operand::Number { .. }))), + O::DoubleQuote => a.len() == 3 && num(0) && num(1) && string(2), + O::ID | O::EI => false, + O::Unknown => true, + }; + if ok { + Ok(()) + } else { + malformed(op.op_span.start, "operand mismatch") + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/lexer/inline.rs b/src-tauri/src/pdf_engine/text_edit/lexer/inline.rs new file mode 100644 index 0000000..e61f6ff --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/lexer/inline.rs @@ -0,0 +1,348 @@ +//! Inline images (SPEC §A.4, decision D13): the `BI … ID … EI` dictionary and the proven end +//! of the payload (`/L`, unfiltered size, Flate end, ASCII terminators), or the pdf.js-style +//! heuristic that marks every following op `after_unproven_inline_image`. + +use super::{ + is_delimiter, is_whitespace, lex_inner, malformed, values::Builder, InlineImage, InlineProof, + LexError, LexLimits, Operand, Raw, Reader, +}; +use crate::pdf_engine::text_edit::decode::{self, DecodeError}; +use crate::pdf_engine::text_edit::limits; + +/// Reads an inline image after `BI` (dict, `ID`, payload, `EI`) and proves the payload end. +pub(super) fn inline_image( + r: &mut Reader<'_>, + lim: &LexLimits, + bi_end: usize, +) -> Result { + let mut dict: Vec<(Vec, Operand)> = Vec::new(); + loop { + let raw = r.next()?.ok_or(LexError::Malformed { + at: bi_end, + what: "unterminated inline image", + })?; + let key = match raw { + Raw::Word(span) if r.b.get(span.clone()) == Some(&b"ID"[..]) => break, + Raw::Operand(Operand::Name { bytes, .. }) => bytes, + other => return malformed(other.start(), "inline image dictionary"), + }; + let first = r.next()?.ok_or(LexError::Malformed { + at: bi_end, + what: "unterminated inline image", + })?; + let value = read_value(r, first, lim)?; + if dict.len() >= lim.array_items_max { + return Err(LexError::TooComplex { + what: "inline image dictionary", + }); + } + dict.push((key, value)); + } + let id_end = r.pos; + if !r.at(id_end).is_some_and(is_whitespace) { + return malformed(id_end, "no whitespace after ID"); + } + // The separator is one whitespace byte; CR LF counts as two only when the one-byte reading + // cannot be proven and the two-byte one can. + let one = id_end + 1; + let mut found = prove_end(r, &dict, one, lim)?.map(|f| (one, f)); + if found.is_none() && r.at(id_end) == Some(b'\r') && r.at(one) == Some(b'\n') { + found = prove_end(r, &dict, one + 1, lim)?.map(|f| (one + 1, f)); + } + let (data_start, data_end, ei, proof) = match found { + Some((start, (end, ei, proof))) => (start, end, ei, proof), + None => { + let (end, ei) = heuristic_end(r.b, one, lim)?; + (one, end, ei, InlineProof::Heuristic) + } + }; + r.pos = ei + 2; + dict.shrink_to_fit(); + Ok(InlineImage { + dict, + data: data_start..data_end, + proof, + }) +} + +/// The ends of the inline images of one buffer, read in order with the lexer's own proofs (the +/// content-part boundary scan, `content::joins`). One reader serves the whole buffer, so the +/// ASCII-terminator searches are remembered across images as in a lex. +pub(crate) struct InlineEnds<'a> { + r: Reader<'a>, + lim: LexLimits, +} + +impl<'a> InlineEnds<'a> { + pub(crate) fn new(b: &'a [u8]) -> Self { + let lim = LexLimits::page(); + InlineEnds { + r: Reader::new(b, true, false, &lim), + lim, + } + } + + /// The offset just past the `EI` of the inline image whose `BI` ends at `bi_end`. `None` when + /// the image does not end in the buffer, is malformed or over a limit, or when the bytes a + /// reader looks at after the `EI` reach the end of the buffer, whatever proof found it + /// (review-verify MEDIUM-A): pdf.js 4.10 takes an `EI` only when the 15 bytes after it are + /// ASCII and lex to a known operator (`findDefaultInlineStreamEnd`, its default for every + /// filter but DCT and the ASCII ones), and our heuristic looks `INLINE_HEURISTIC_WINDOW` + /// (≥ 15) bytes ahead, so bytes after the buffer — the next part, shifted by qpdf's `\n` — + /// could change which `EI` ends the image. + pub(crate) fn end(&mut self, bi_end: usize) -> Option { + self.r.pos = bi_end; + self.r.nodes = 0; + inline_image(&mut self.r, &self.lim, bi_end).ok()?; + let end = self.r.pos; + let window_end = end.saturating_add(limits::INLINE_HEURISTIC_WINDOW); + (window_end < self.r.b.len()).then_some(end) + } +} + +fn read_value(r: &mut Reader<'_>, first: Raw, lim: &LexLimits) -> Result { + let mut builder = Builder::new(lim); + let mut raw = first; + loop { + if let Raw::Word(span) = &raw { + return malformed(span.start, "inline image value"); + } + if let Some(v) = builder.feed(raw)? { + return Ok(v); + } + raw = r.next()?.ok_or(LexError::Malformed { + at: r.pos, + what: "unterminated inline image", + })?; + } +} + +fn dict_get<'d>(dict: &'d [(Vec, Operand)], keys: &[&[u8]]) -> Option<&'d Operand> { + dict.iter() + .find(|(k, _)| keys.contains(&k.as_slice())) + .map(|(_, v)| v) +} + +fn dict_uint(dict: &[(Vec, Operand)], keys: &[&[u8]]) -> Option { + let v = dict_get(dict, keys)?.as_number()?; + (v >= 0.0 && v.fract() == 0.0 && v <= limits::NUMBER_ABS_MAX).then_some(v as usize) +} + +/// Filters of an inline image, abbreviations kept as written. +fn image_filters(dict: &[(Vec, Operand)]) -> Option>> { + match dict_get(dict, &[b"F", b"Filter"]) { + None => Some(Vec::new()), + Some(Operand::Name { bytes, .. }) => Some(vec![bytes.clone()]), + Some(Operand::Array { items, .. }) => items + .iter() + .map(|i| i.as_name().map(<[u8]>::to_vec)) + .collect(), + Some(_) => None, + } +} + +/// `EI` after optional whitespace from `pos`, followed by whitespace, a delimiter or EOF. +fn ei_after(b: &[u8], pos: usize) -> Option { + let mut q = pos; + while b.get(q).is_some_and(|c| is_whitespace(*c)) { + q += 1; + } + let ok = b.get(q..q.checked_add(2)?) == Some(&b"EI"[..]) + && b.get(q + 2) + .is_none_or(|c| is_whitespace(*c) || is_delimiter(*c)); + ok.then_some(q) +} + +/// Tries the four proofs of §A.4 in order; `Ok(None)` when none holds. +fn prove_end( + r: &mut Reader<'_>, + dict: &[(Vec, Operand)], + start: usize, + lim: &LexLimits, +) -> Result, LexError> { + let b = r.b; + let too_big = LexError::TooComplex { + what: "inline image size", + }; + let check = |end: usize| -> Option<(usize, usize)> { ei_after(b, end).map(|ei| (end, ei)) }; + if let Some(len) = dict_uint(dict, &[b"L", b"Length"]) { + if len > lim.inline_image_max { + return Err(too_big); + } + if let Some((end, ei)) = start.checked_add(len).and_then(check) { + return Ok(Some((end, ei, InlineProof::LengthKey))); + } + } + let Some(filters) = image_filters(dict) else { + return Ok(None); + }; + if filters.is_empty() { + if let Some(size) = unfiltered_size(dict) { + if size > lim.inline_image_max as u64 { + return Err(too_big); + } + let end = start.checked_add(size as usize).and_then(check); + if let Some((end, ei)) = end { + return Ok(Some((end, ei, InlineProof::UnfilteredSize))); + } + } + return Ok(None); + } + let payload = b.get(start..).unwrap_or_default(); + match filters.first().map(Vec::as_slice) { + Some(b"Fl" | b"FlateDecode") if filters.len() == 1 => { + match decode::inflate_end(payload, lim.inline_image_max) { + Ok(consumed) => { + let end = start.checked_add(consumed).ok_or(LexError::TooComplex { + what: "inline image size", + })?; + for candidate in [end.checked_add(4), Some(end)].into_iter().flatten() { + if candidate <= b.len() { + if let Some((e, ei)) = check(candidate) { + return Ok(Some((e, ei, InlineProof::FlateEnd))); + } + } + } + Ok(None) + } + Err(DecodeError::TooLarge) => Err(too_big), + Err(_) => Ok(None), + } + } + Some(b"AHx" | b"ASCIIHexDecode") => Ok(find_terminator(r, HEX_END, start, lim) + .and_then(check) + .map(|(e, ei)| (e, ei, InlineProof::AsciiEnd))), + Some(b"A85" | b"ASCII85Decode") => Ok(find_terminator(r, A85_END, start, lim) + .and_then(check) + .map(|(e, ei)| (e, ei, InlineProof::AsciiEnd))), + _ => Ok(None), + } +} + +/// The ASCIIHex and ASCII85 end markers, with their slot in `Reader::terminators`. +const HEX_END: (usize, &[u8]) = (0, b">"); +const A85_END: (usize, &[u8]) = (1, b"~>"); + +/// The end of an ASCII payload starting at `start`: just after the first terminator, when that +/// keeps the payload within `inline_image_max` (a longer one is no proof, as for the other +/// proofs). Each search is remembered per reader and reused while it still answers (inline images +/// are read in order), so every byte of a content is searched at most once per terminator however +/// many images lack one: one search per image over the rest of the content was quadratic +/// (review T3-budget MEDIUM-4, 8,000 images took 19 s). +fn find_terminator( + r: &mut Reader<'_>, + (slot, needle): (usize, &[u8]), + start: usize, + lim: &LexLimits, +) -> Option { + let remembered = r.terminators.get(slot).copied().flatten(); + let hit = match remembered { + Some((from, None)) if from <= start => None, + Some((from, Some(at))) if from <= start && at >= start => Some(at), + _ => { + let at = + r.b.get(start..) + .and_then(|rest| find(rest, needle)) + .map(|p| start + p); + if let Some(memo) = r.terminators.get_mut(slot) { + *memo = Some((start, at)); + } + at + } + }; + let end = hit?.checked_add(needle.len())?; + (end.saturating_sub(start) <= lim.inline_image_max).then_some(end) +} + +fn find(hay: &[u8], needle: &[u8]) -> Option { + hay.windows(needle.len()).position(|w| w == needle) +} + +/// `ceil(W·BPC·ncomp/8)·H` for an unfiltered image whose colour space is known. +fn unfiltered_size(dict: &[(Vec, Operand)]) -> Option { + let w = dict_uint(dict, &[b"W", b"Width"])? as u64; + let h = dict_uint(dict, &[b"H", b"Height"])? as u64; + let mask = matches!( + dict_get(dict, &[b"IM", b"ImageMask"]), + Some(Operand::Bool { value: true, .. }) + ); + let (bpc, ncomp) = if mask { + (1u64, 1u64) + } else { + let bpc = dict_uint(dict, &[b"BPC", b"BitsPerComponent"])? as u64; + if !matches!(bpc, 1 | 2 | 4 | 8 | 16) { + return None; + } + let ncomp = match dict_get(dict, &[b"CS", b"ColorSpace"])? { + Operand::Name { bytes, .. } => match bytes.as_slice() { + b"G" | b"DeviceGray" | b"I" | b"Indexed" => 1, + b"RGB" | b"DeviceRGB" => 3, + b"CMYK" | b"DeviceCMYK" => 4, + _ => return None, + }, + Operand::Array { items, .. } => match items.first().and_then(Operand::as_name) { + Some(b"I" | b"Indexed") => 1, + _ => return None, + }, + _ => return None, + }; + (bpc, ncomp) + }; + let row_bits = w.checked_mul(bpc)?.checked_mul(ncomp)?; + row_bits.checked_add(7).map(|x| x / 8)?.checked_mul(h) +} + +/// pdf.js-style fallback: the first `EI` preceded by whitespace and followed by whitespace/EOF +/// after which the next `INLINE_HEURISTIC_WINDOW` bytes lex as valid operators. At most +/// `INLINE_HEURISTIC_CANDIDATES_MAX` such candidates are lexed (each lex reads a window). +fn heuristic_end(b: &[u8], start: usize, lim: &LexLimits) -> Result<(usize, usize), LexError> { + let limit = start.saturating_add(lim.inline_image_max); + let mut q = start; + let mut candidates = 0usize; + while let Some(rel) = b.get(q..).and_then(|rest| find(rest, b"EI")) { + let p = q + rel; + if p > limit { + return Err(LexError::TooComplex { + what: "inline image size", + }); + } + let before = p + .checked_sub(1) + .and_then(|i| b.get(i)) + .is_some_and(|c| is_whitespace(*c)); + let after = b.get(p + 2).is_none_or(|c| is_whitespace(*c)); + if before && after { + if candidates >= limits::INLINE_HEURISTIC_CANDIDATES_MAX { + return Err(LexError::TooComplex { + what: "inline image end candidates", + }); + } + candidates += 1; + if tail_lexes(b, p + 2, lim) { + return Ok((p.saturating_sub(1).max(start), p)); + } + } + q = p + 1; + } + malformed(start, "inline image without an end") +} + +/// The bytes after a candidate `EI` must lex as operators with valid arity. A window cut +/// before EOF is shortened to its last whitespace and must then contain at least one op. +fn tail_lexes(b: &[u8], from: usize, lim: &LexLimits) -> bool { + let end = from + .saturating_add(limits::INLINE_HEURISTIC_WINDOW) + .min(b.len()); + let Some(window) = b.get(from..end) else { + return true; + }; + let truncated = end < b.len(); + let window = match window.iter().rposition(|c| is_whitespace(*c)) { + Some(cut) if truncated => window.get(..cut).unwrap_or_default(), + _ => window, + }; + match lex_inner(window, lim, None, true) { + Ok(ops) => !truncated || !ops.is_empty(), + Err(_) => false, + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/lexer/scan.rs b/src-tauri/src/pdf_engine/text_edit/lexer/scan.rs new file mode 100644 index 0000000..1421435 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/lexer/scan.rs @@ -0,0 +1,104 @@ +//! Generic token scanner (objects, CMaps, Type1 clear text) used by the snapshot preflight and +//! the font parsers: flat tokens, balanced brackets, no arity. + +use super::{malformed, LexError, LexLimits, Raw, Reader, ScanMode, Token}; + +/// Content limits without the operand-node cap (`max_tokens` bounds a scan). +fn scan_limits() -> LexLimits { + LexLimits { + operand_nodes_max: usize::MAX, + ..LexLimits::page() + } +} + +pub(super) fn tokens( + bytes: &[u8], + mode: ScanMode, + max_tokens: usize, +) -> Result, LexError> { + let lim = scan_limits(); + let mut r = Reader::new(bytes, false, mode != ScanMode::Object, &lim); + let mut out = Vec::new(); + let mut depth: Vec = Vec::new(); // true = dict + while let Some(raw) = r.next()? { + if out.len() >= max_tokens { + return Err(LexError::TooComplex { what: "tokens" }); + } + out.push(flat_token(raw, &mut depth, &r, &lim)?); + } + if !depth.is_empty() { + return malformed(bytes.len(), "unbalanced brackets"); + } + Ok(out) +} + +pub(super) fn dict_at( + bytes: &[u8], + start: usize, + max_tokens: usize, +) -> Result<(Vec, usize), LexError> { + let lim = scan_limits(); + let mut r = Reader::new(bytes, false, false, &lim); + r.pos = start; + let mut out = Vec::new(); + let mut depth: Vec = Vec::new(); + loop { + let raw = r.next()?.ok_or(LexError::Malformed { + at: r.pos, + what: "unterminated dictionary", + })?; + if out.is_empty() && !matches!(raw, Raw::DictOpen(_)) { + return malformed(raw.start(), "dictionary expected"); + } + if out.len() >= max_tokens { + return Err(LexError::TooComplex { what: "tokens" }); + } + out.push(flat_token(raw, &mut depth, &r, &lim)?); + if depth.is_empty() { + return Ok((out, r.pos)); + } + } +} + +fn flat_token( + raw: Raw, + depth: &mut Vec, + r: &Reader<'_>, + lim: &LexLimits, +) -> Result { + let mut push = |dict: bool| { + if depth.len() >= lim.nesting_max { + return Err(LexError::TooComplex { what: "nesting" }); + } + depth.push(dict); + Ok(()) + }; + Ok(match raw { + Raw::Operand(o) => Token::Operand(o), + Raw::Word(span) => Token::Keyword { + bytes: r.b.get(span.clone()).unwrap_or_default().to_vec(), + span, + }, + Raw::DictOpen(p) => { + push(true)?; + Token::DictOpen(p) + } + Raw::ArrayOpen(p) => { + push(false)?; + Token::ArrayOpen(p) + } + Raw::DictClose(p) => match depth.pop() { + Some(true) => Token::DictClose(p), + _ => return malformed(p, "stray >>"), + }, + Raw::ArrayClose(p) => match depth.pop() { + Some(false) => Token::ArrayClose(p), + _ => return malformed(p, "stray ]"), + }, + Raw::BraceOpen(p) | Raw::BraceClose(p) if r.lenient => Token::Keyword { + bytes: r.b.get(p..p + 1).unwrap_or_default().to_vec(), + span: p..p + 1, + }, + Raw::BraceOpen(p) | Raw::BraceClose(p) => return malformed(p, "brace outside a procedure"), + }) +} diff --git a/src-tauri/src/pdf_engine/text_edit/lexer/values.rs b/src-tauri/src/pdf_engine/text_edit/lexer/values.rs new file mode 100644 index 0000000..a737c9d --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/lexer/values.rs @@ -0,0 +1,138 @@ +//! Operand values: nested arrays and dictionaries assembled iteratively from raw tokens, with +//! the nesting and item budgets of `LexLimits`. + +use super::{malformed, LexError, LexLimits, Operand, Raw}; + +pub(super) enum Frame { + Array { + items: Vec, + start: usize, + }, + Dict { + entries: Vec<(Vec, Operand)>, + key: Option>, + start: usize, + }, +} + +/// Collects operand values into nested arrays/dicts (iterative, depth- and size-capped). +pub(super) struct Builder { + pub(super) frames: Vec, + nesting_max: usize, + items_max: usize, +} + +impl Builder { + pub(super) fn new(lim: &LexLimits) -> Self { + Builder { + frames: Vec::new(), + nesting_max: lim.nesting_max, + items_max: lim.array_items_max, + } + } + + fn open(&mut self, frame: Frame) -> Result<(), LexError> { + if self.frames.len() >= self.nesting_max { + return Err(LexError::TooComplex { + what: "array or dictionary nesting", + }); + } + self.frames.push(frame); + Ok(()) + } + + /// Adds `v` to the innermost frame; returns it back when there is no open frame. + fn add(&mut self, v: Operand) -> Result, LexError> { + match self.frames.last_mut() { + None => Ok(Some(v)), + Some(Frame::Array { items, .. }) => { + if items.len() >= self.items_max { + return Err(LexError::TooComplex { + what: "array items", + }); + } + items.push(v); + Ok(None) + } + Some(Frame::Dict { entries, key, .. }) => { + match key.take() { + None => match v { + Operand::Name { bytes, .. } => *key = Some(bytes), + other => { + return malformed(other.span().start, "dictionary key is not a name") + } + }, + Some(k) => { + if entries.len() >= self.items_max { + return Err(LexError::TooComplex { + what: "dictionary entries", + }); + } + entries.push((k, v)); + } + } + Ok(None) + } + } + } + + fn close_array(&mut self, at: usize) -> Result { + match self.frames.pop() { + Some(Frame::Array { mut items, start }) => { + items.shrink_to_fit(); + Ok(Operand::Array { + items, + span: start..at + 1, + }) + } + _ => malformed(at, "stray ]"), + } + } + + fn close_dict(&mut self, at: usize) -> Result { + match self.frames.pop() { + Some(Frame::Dict { + mut entries, + key: None, + start, + }) => { + entries.shrink_to_fit(); + Ok(Operand::Dict { + entries, + span: start..at + 2, + }) + } + _ => malformed(at, "stray >>"), + } + } + + /// Feeds one raw token; returns a completed top-level value, if any. + pub(super) fn feed(&mut self, raw: Raw) -> Result, LexError> { + match raw { + Raw::Operand(v) => self.add(v), + Raw::ArrayOpen(p) => self + .open(Frame::Array { + items: Vec::new(), + start: p, + }) + .map(|_| None), + Raw::DictOpen(p) => self + .open(Frame::Dict { + entries: Vec::new(), + key: None, + start: p, + }) + .map(|_| None), + Raw::ArrayClose(p) => { + let v = self.close_array(p)?; + self.add(v) + } + Raw::DictClose(p) => { + let v = self.close_dict(p)?; + self.add(v) + } + Raw::Word(span) => malformed(span.start, "operator inside array or dictionary"), + Raw::BraceOpen(p) | Raw::BraceClose(p) => malformed(p, "brace in content"), + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/limits.rs b/src-tauri/src/pdf_engine/text_edit/limits.rs new file mode 100644 index 0000000..0ad588d --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/limits.rs @@ -0,0 +1,204 @@ +//! Every budget and tolerance of Edit text (SPEC §B.2). Exceeding a budget: per page → page +//! `PAGE_TOO_COMPLEX`; at file level → `FILE_TOO_LARGE`/`FILE_TOO_COMPLEX`; inside the gate → +//! `EDIT_VERIFY_FAILED` "too large to verify". Never a skipped check. + +pub const FILE_CAP_BYTES: u64 = 400 * 1024 * 1024; +pub const MAX_OBJECTS: usize = 2_000_000; +pub const MAX_XREF_CHAIN: usize = 64; +pub const XREF_STREAM_MAX_DECODED: usize = 64 << 20; +pub const OBJSTM_MAX_DECODED: usize = 32 << 20; +pub const OBJSTM_TOTAL_DECODED: usize = 256 << 20; +pub const MAX_OBJECT_NESTING: usize = 100; // preflight: every value lopdf can parse (objects, ObjStm members) +pub const STREAM_MAX_DECODED: usize = 32 << 20; // any single stream +pub const PAGE_CONTENT_MAX_DECODED: usize = 48 << 20; +pub const PAGE_DECODE_BUDGET: usize = 96 << 20; // content + forms + fonts + cmaps for one page walk +pub const PAGE_PARTS_MAX: usize = 256; +pub const PAGE_OPS_MAX: usize = 250_000; // counted while lexing +pub const TOKEN_NESTING_MAX: usize = 32; // arrays/dicts in content +pub const ARRAY_ITEMS_MAX: usize = 65_536; +pub const STRING_BYTES_MAX: usize = 1 << 20; +pub const LITERAL_PAREN_NESTING_MAX: usize = 100; +pub const INLINE_IMAGE_MAX_BYTES: usize = 16 << 20; +pub const Q_DEPTH_MAX: usize = 64; +pub const MARKED_DEPTH_MAX: usize = 64; +pub const FORM_DEPTH_MAX: usize = 8; +pub const FORM_PAINTS_PER_PAGE_MAX: usize = 4_096; +pub const GLYPHS_PER_PAGE_MAX: usize = 400_000; +pub const RUNS_PER_PAGE_MAX: usize = 20_000; +pub const RUN_MEMBERS_MAX: usize = 256; +pub const FONTS_PER_PAGE_MAX: usize = 256; +pub const FONT_PROGRAM_MAX_DECODED: usize = 16 << 20; +pub const TOUNICODE_MAX_DECODED: usize = 2 << 20; +pub const CMAP_MAPPINGS_MAX: usize = 131_072; +pub const CIDTOGID_MAX_BYTES: usize = 131_072; +pub const W_ENTRIES_MAX: usize = 65_536; +pub const STRUCT_CHAIN_MAX: usize = 16; +pub const NUMBER_TREE_NODES_MAX: usize = 100_000; +pub const EDIT_TEXT_CHARS_MAX: usize = 1_000; +pub const EDITS_PER_PAGE_MAX: usize = 200; +pub const EDITS_PER_SAVE_MAX: usize = 500; +pub const DRIFT_TOLERANCE_PT: f64 = 0.01; +pub const STATE_EPSILON: f64 = 1e-9; +pub const COLOR_EPSILON: f64 = 1e-6; +pub const AXIS_EPSILON_REL: f64 = 1e-4; +pub const SHEAR_MAX: f64 = 0.5; +pub const JOIN_GAP_EM: f64 = 0.3; +pub const JOIN_BASELINE_TOL_PT: f64 = 0.01; +pub const SYNTH_SPACE_EM: f64 = 0.2; +pub const DEFAULT_KERN_SPACE: f64 = -250.0; +pub const PER_GLYPH_MIN_RUNS: usize = 12; +pub const PER_GLYPH_SHARE: f64 = 0.8; +pub const DUPLICATE_OVERLAP_SHARE: f64 = 0.5; +pub const CLIP_CONTAIN_TOL_PT: f64 = 0.5; +pub const OVERLAP_WARN_TOL_PT: f64 = 0.5; +pub const NEGLIGIBLE_KERN: f64 = 0.0005; +pub const NUMBER_DECIMALS: usize = 4; +pub const NUMBER_ABS_MAX: f64 = 1e9; +pub const F32_DRIFT_GUARD_PT: f64 = 0.002; +pub const STYLE_EPSILON: f64 = 0.001; +pub const SIZE_MIN_PT: f64 = 4.0; +pub const SIZE_MAX_PT: f64 = 144.0; +pub const LETTER_SPACING_MIN_PT: f64 = -2.0; +pub const LETTER_SPACING_MAX_PT: f64 = 10.0; +pub const PDFTOTEXT_OUTPUT_MAX: usize = 16 << 20; +pub const WORD_BBOX_TOL_PT: f64 = 0.05; +pub const EDIT_BAND_PAD_PT: f64 = 2.0; +pub const RENDER_DPI_MAX: u32 = 96; +pub const RENDER_PIXELS_MAX: u64 = 25_000_000; +pub const RENDER_CHANNEL_TOL: u8 = 24; +pub const RENDER_OUTSIDE_PIXELS_MAX: u64 = 8; +pub const RENDER_MASK_PAD_PX: i64 = 4; +/// Pixels A5 lets change in the boxes of glyphs no edit changes, outside the edited glyphs' own +/// boxes, and pixels of a neighbour's ink it lets go under the masks (review-verify HIGH-A): an +/// honest edit changes none there (0 in every test and probe), and a narrow follower is smaller +/// than `RENDER_OUTSIDE_PIXELS_MAX` (a 12 pt period whitened changes 5 pixels at 96 DPI). +pub const RENDER_KEPT_PIXELS_MAX: u64 = 2; +/// Pixels the edited glyphs' own boxes grow by where A5 checks the glyphs no edit changes: +/// Poppler's anti-aliasing reaches one pixel past a glyph's outward-rounded box (every honest +/// change in the overlap sweeps of review-verify HIGH-A lay exactly 1 pixel out). +pub const RENDER_OWN_PAD_PX: i64 = 1; +pub const GATE_DECODED_TOTAL: u64 = 4 << 30; // decode budget of one gate run (A2 decodes only qpdf-compressed streams; A3/A4/B1–B3) +pub const PREVIEW_PDF_MAX_BYTES: u64 = 48 << 20; +/// A cached page model over this size leaves the cache for its preview, which frees it once the +/// edit is planned (the page's next visit builds it again): held through the preview, the source +/// model, the extracted page's model and Phase A's walks stacked past §H R20's 256 MiB per file +/// (review-final MEDIUM-3: 295 MiB for an 84 MiB model; released, about 2.5 × the model). +pub const PREVIEW_SHARED_MODEL_MAX: usize = 64 << 20; +pub const UPDATE_JSON_MAX_BYTES: usize = 128 << 20; +pub const SUBPROCESS_TIMEOUT_SECS: u64 = 120; +pub const CACHE_SNAPSHOTS_MAX: usize = 2; +pub const CACHE_SNAPSHOT_BYTES_MAX: u64 = 256 << 20; // larger files are re-read per call +pub const CACHE_PAGE_MODELS_MAX: usize = 32; +// revision 2 +pub const VERIFY_CAP_MARGIN_BYTES: u64 = 256 << 20; // read_verification_snapshot cap = 2 × Σ inputs + margin +pub const QPDF_JSON_MAX_BYTES: usize = 64 << 20; // stdout cap of `qpdf --json=2 --json-key=pages` +pub const CHECK_MEMO_MAX: usize = 16; // fingerprints whose qpdf --check result is memoised +pub const HEAD_TAIL_HASH_BYTES: usize = 64 << 10; // stat_matches also hashes the first and last 64 KiB +pub const PAGE_TREE_DEPTH_MAX: usize = 64; // own page-tree walk for /Kids occurrence counts +pub const GRAPH_DIRECT_DEPTH_MAX: usize = 100; // canonical graph serialisation (= preflight nesting) +pub const STRUCT_ORDER_NODES_MAX: usize = 100_000; // structure-tree DFS for reading order +pub const XY_CUT_DEPTH_MAX: usize = 32; +pub const XY_CUT_ROW_GAP_EM: f64 = 0.5; // horizontal cut: empty band ≥ 0.5 × median effective size +pub const XY_CUT_COL_GAP_EM: f64 = 1.0; // vertical cut: empty gutter ≥ 1 × median effective size +pub const COLUMN_GAP_EM: f64 = 1.0; // a kept TJ gap ≥ 1 em absorbs the width change (B.12 step 9) +pub const WRAPPER_MATRIX_EPSILON: f64 = 1e-6; // Phase B: linear part of (wrapper cm × Form /Matrix) vs identity +pub const WRAPPER_TRANSLATION_TOL_PT: f64 = 0.001; // Phase B: its translation vs 0 + +// T1 implementation budgets (not in the §B.2 list; each bounds one T1 parser or reader) +pub const XREF_TAIL_SEARCH_BYTES: usize = 2 << 10; // `startxref` must sit in the last 2 KiB +pub const PREFLIGHT_DICT_TOKENS_MAX: usize = 100_000; // trailer / xref-stream dictionary tokens +pub const XREF_STREAM_FIELD_WIDTH_MAX: i64 = 8; // bytes per `/W` field of an xref stream +pub const TOOL_STDERR_MAX: usize = 1 << 20; // stderr kept from one subprocess (the rest is drained) +pub const TOOL_POLL_MS: u64 = 20; // subprocess wait loop period +pub const LEX_CANCEL_EVERY_OPS: usize = 4_096; // lexer checks the cancel flag this often +pub const INLINE_HEURISTIC_WINDOW: usize = 64; // bytes after a candidate `EI` that must lex cleanly +pub const INFLATE_STEP_BYTES: usize = 64 << 10; // output step of the capped inflater +pub const INLINE_HEURISTIC_CANDIDATES_MAX: usize = 256; // `EI` candidates tried per unproven inline image +pub const XREF_PREDICTOR_ROW_MAX: usize = 1 << 10; // PNG-predictor row of an xref stream (bytes) +pub const OBJECT_SCAN_FACTOR: usize = 2; // object value scans read ≤ 2 × the buffer … +pub const OBJECT_SCAN_SLACK_BYTES: usize = 1 << 20; // … + 1 MiB in total +pub const LENGTH_REF_STREAMS_MAX: usize = MAX_OBJECTS / 2; // streams whose /Length is a reference +pub const LENGTH_REF_CHAIN_MAX: usize = 32; // objects lopdf parses recursively for one stream's /Length +pub const OBJSTM_SCAN_BYTES: usize = + OBJECT_SCAN_FACTOR * OBJSTM_TOTAL_DECODED + OBJECT_SCAN_SLACK_BYTES; // member scans of all object streams of one load + +// T3 budget pass (not in the §B.2 list; see DEVIATIONS "[T3-budget]") +/// Bytes one page model build (`walk_page` + `build_runs`) may hold at once, and the model and its +/// #33 Classify pass together: every allocation that grows with the page (the font models it keeps +/// included) is charged before it is made, and the page is `PAGE_TOO_COMPLEX` "page model size" +/// past it (`walker/budget.rs`). A cached model stays within §H R20's 256 MiB per file; real pages +/// hold a few MiB. Measured (fix pass 2026-10-03, debug build): 100,000 one-glyph shows under one +/// state hold 86 MiB and classify in 22 MiB more; ~131,000 is the most (the records vector's +/// doubling at 131,072 needs room for both buffers); 3,000 lines × 33 glyphs with a `Tm` per glyph +/// hold 89 MiB, 112 MiB with a colour change per word. +pub const PAGE_MODEL_BYTES_MAX: usize = 160 << 20; +/// Bytes of font models the unused fonts of a page's resources (siblings for faces and joins) +/// may add to its model (fix pass 2026-10-03, review T3-budget HIGH-1): a font that does not fit +/// is left out, never the page refused. Drawn fonts are charged to the page budget in full. +pub const PAGE_UNUSED_FONT_BYTES_MAX: usize = 32 << 20; +/// Lexed operand nodes alive at once in one walk (the page's ops, plus a descended Form's while it +/// runs); also `LexLimits::operand_nodes_max`, so the lexer stops before it builds more. Six per +/// op at `PAGE_OPS_MAX` (`c` has six), so only long number arrays reach it. +pub const OPERAND_NODES_MAX: usize = 6 * PAGE_OPS_MAX; +/// Longest font resource name a sibling group may hold (ISO 32000 Annex C name limit). A longer +/// name types only its own runs: each run's surface lists every sibling's name. +pub const SIBLING_NAME_BYTES_MAX: usize = 127; +/// Longest unmodelled ExtGState key (Annex C again); longer is `PAGE_TOO_COMPLEX` "ExtGState +/// keys". Every `gs` compares the keys it sets with the ones in force. +pub const EXTGSTATE_KEY_BYTES_MAX: usize = 127; + +/// The source-file cap: `FILE_CAP_BYTES` in production; tests may override it per thread +/// (`set_file_cap_override`, used by E2E-21 to exercise the verification cap). +pub fn file_cap() -> u64 { + #[cfg(test)] + if let Some(cap) = FILE_CAP_OVERRIDE.with(|c| c.get()) { + return cap; + } + FILE_CAP_BYTES +} + +/// `PAGE_MODEL_BYTES_MAX` in production; tests may override it per thread +/// (`set_model_bytes_override`) to reach the budget with small pages. +pub fn model_bytes_max() -> usize { + #[cfg(test)] + if let Some(bytes) = MODEL_BYTES_OVERRIDE.with(|c| c.get()) { + return bytes; + } + PAGE_MODEL_BYTES_MAX +} + +/// Bytes all object-stream member scans of one load may read: `OBJSTM_SCAN_BYTES` in +/// production; tests may override it per thread (`set_objstm_scan_override`; `guarded_load` reads +/// it on the calling thread before lopdf's workers start). +pub fn objstm_scan_budget() -> usize { + #[cfg(test)] + if let Some(bytes) = OBJSTM_SCAN_OVERRIDE.with(|c| c.get()) { + return bytes; + } + OBJSTM_SCAN_BYTES +} + +#[cfg(test)] +thread_local! { + static FILE_CAP_OVERRIDE: std::cell::Cell> = const { std::cell::Cell::new(None) }; + static OBJSTM_SCAN_OVERRIDE: std::cell::Cell> = const { std::cell::Cell::new(None) }; + static MODEL_BYTES_OVERRIDE: std::cell::Cell> = const { std::cell::Cell::new(None) }; +} + +/// Test seam: override `model_bytes_max()` on the current thread (`None` restores the constant). +#[cfg(test)] +pub fn set_model_bytes_override(bytes: Option) { + MODEL_BYTES_OVERRIDE.with(|c| c.set(bytes)); +} + +/// Test seam: override `file_cap()` on the current thread (`None` restores the constant). +#[cfg(test)] +pub fn set_file_cap_override(cap: Option) { + FILE_CAP_OVERRIDE.with(|c| c.set(cap)); +} + +/// Test seam: override `objstm_scan_budget()` on the current thread (`None` restores the constant). +#[cfg(test)] +pub fn set_objstm_scan_override(bytes: Option) { + OBJSTM_SCAN_OVERRIDE.with(|c| c.set(bytes)); +} diff --git a/src-tauri/src/pdf_engine/text_edit/mod.rs b/src-tauri/src/pdf_engine/text_edit/mod.rs new file mode 100644 index 0000000..20dccd3 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/mod.rs @@ -0,0 +1,92 @@ +//! Edit text (v0.4): in-place editing of existing PDF text, fail closed. +//! +//! A change is a byte splice of show-text operators inside the page's own content stream, +//! written by qpdf (`--update-from-json`) and verified before anything is published. Nothing +//! here covers old text with new text, flattens or rasterises a page, re-serialises a page +//! with lopdf, or overwrites the input. +//! +//! Policy for every file in this module tree: +//! - Fail closed: every refusal carries a SCREAMING_SNAKE reason code (`reasons.rs`); no +//! silent fallback, no skipped edit, no partial list returned as `Ok`. +//! - Bounded: one capped read per operation (`snapshot.rs`), every decompression capped while +//! inflating (`decode.rs`), every parser with an operation and nesting budget (`limits.rs`). +//! - No panics on data derived from a PDF: no `unwrap`/`expect`/`panic!`/unchecked indexing in +//! production code (denied below for clippy), checked arithmetic for offsets. +//! - Forbidden lopdf APIs (enforced by `tests_guard::no_forbidden_apis`, which skips comments +//! and string literals): `Content::decode`, `Content::encode`, `string_to_bytes`, +//! `replace_text`, `encode_text`, `decompressed_content`, `get_plain_content`, +//! `get_page_content`, `Document::load`, `load_mem`, `.decompress()`, `IncrementalDocument`, +//! `.save(`, `save_to(`. lopdf only reads, and only from the bytes of one snapshot. +//! - Test seams exist only under `#[cfg(test)]`; no environment variable or setting can weaken +//! a check (`OFFPDF_REQUIRE_ENGINES` only turns a test skip into a test failure). +//! +//! Callers use full paths (`text_edit::snapshot::read_snapshot`); there is no `pub use`. +#![cfg_attr( + not(test), + deny( + clippy::unwrap_used, + clippy::expect_used, + clippy::panic, + clippy::indexing_slicing + ) +)] + +// T1 engine core +pub(crate) mod content; +pub(crate) mod decode; +pub(crate) mod engines; +pub(crate) mod lexer; +pub(crate) mod limits; +pub(crate) mod reasons; +pub(crate) mod snapshot; + +// T2 fonts +pub(crate) mod fonts; + +// T3 walker, runs, reading order +pub(crate) mod context; +pub(crate) mod geometry; +pub(crate) mod order; +pub(crate) mod runs; +pub(crate) mod state; +pub(crate) mod structure; +pub(crate) mod walker; + +// T4 rewrite, verify, apply, gate, preview +pub(crate) mod apply; +pub(crate) mod encode; +pub(crate) mod fit; +pub(crate) mod gate; +pub(crate) mod graph; +pub(crate) mod poppler; +pub(crate) mod preview; +pub(crate) mod rewrite; +pub(crate) mod verify; + +// T5 integration +pub(crate) mod cache; +pub(crate) mod dto; +pub(crate) mod export; +pub(crate) mod service; + +// Tests and test helpers +#[cfg(test)] +pub(crate) mod bench; +#[cfg(test)] +pub(crate) mod testkit; +#[cfg(test)] +pub(crate) mod tests_dto; +#[cfg(test)] +pub(crate) mod tests_e2e; +#[cfg(test)] +pub(crate) mod tests_gate; +#[cfg(test)] +pub(crate) mod tests_guard; +#[cfg(test)] +pub(crate) mod tests_independent; +#[cfg(test)] +pub(crate) mod tests_io; +#[cfg(test)] +pub(crate) mod tests_plan; +#[cfg(test)] +pub(crate) mod tests_walk; diff --git a/src-tauri/src/pdf_engine/text_edit/order.rs b/src-tauri/src/pdf_engine/text_edit/order.rs new file mode 100644 index 0000000..100baa6 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/order.rs @@ -0,0 +1,193 @@ +//! Reading order (SPEC §B.11, D33): structure-tree order on tagged pages, XY-cut in display space +//! otherwise — all or nothing, never a mix. Sets `order` (rank) and `line` (line clusters counted +//! in that order) for every listed run, so the DOM order of the text layer is what a screen +//! reader reads. + +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::limits::{ + XY_CUT_COL_GAP_EM, XY_CUT_DEPTH_MAX, XY_CUT_ROW_GAP_EM, +}; +use crate::pdf_engine::text_edit::structure::structure_mcids; +use lopdf::ObjectId; +use std::collections::HashMap; + +/// Baselines within this share of the smaller effective size form one line. +const LINE_CLUSTER_SHARE: f64 = 0.5; +/// Smallest size used for gap thresholds (zero-size runs still terminate the cut). +const SIZE_FLOOR: f64 = 1e-6; + +#[derive(Debug, Clone)] +pub struct RunBox { + /// `[x0, y0, x1, y1]` after `/Rotate`, y down. + pub display_rect: [f64; 4], + pub baseline_y: f64, + pub size: f64, + pub mcid: Option, +} + +/// `(order, line)` per run, in input order. +pub fn reading_order(ctx: &SnapshotContext, page_id: ObjectId, runs: &[RunBox]) -> Vec<(u32, u32)> { + let xy = xy_cut_order(runs); + let order = tagged_order(ctx, page_id, runs, &xy).unwrap_or(xy); + let mut out = vec![(0u32, 0u32); runs.len()]; + let mut line = 0u32; + let mut prev: Option<&RunBox> = None; + for (rank, i) in order.iter().enumerate() { + let Some(b) = runs.get(*i) else { continue }; + if prev.is_some_and(|p| !same_line(p, b)) { + line = line.saturating_add(1); + } + if let Some(slot) = out.get_mut(*i) { + *slot = (u32::try_from(rank).unwrap_or(u32::MAX), line); + } + prev = Some(b); + } + out +} + +/// Next in reading order on the same line: close baselines and not moving back to the left. +fn same_line(a: &RunBox, b: &RunBox) -> bool { + let tol = LINE_CLUSTER_SHARE * a.size.min(b.size); + (a.baseline_y - b.baseline_y).abs() <= tol && b.display_rect[0] >= a.display_rect[0] - tol +} + +/// Structure order when the catalog has a structure tree and every run carries an MCID the DFS +/// reaches (ties keep their XY-cut order); `None` otherwise. +fn tagged_order( + ctx: &SnapshotContext, + page_id: ObjectId, + runs: &[RunBox], + xy: &[usize], +) -> Option> { + if runs.is_empty() || runs.iter().any(|r| r.mcid.is_none()) { + return None; + } + let mcids = structure_mcids(ctx.doc(), page_id)?; + let mut rank: HashMap = HashMap::new(); + for (i, m) in mcids.iter().enumerate() { + rank.entry(*m).or_insert(i); + } + let mut xy_pos = vec![0usize; runs.len()]; + for (pos, i) in xy.iter().enumerate() { + if let Some(slot) = xy_pos.get_mut(*i) { + *slot = pos; + } + } + let mut keyed = runs + .iter() + .enumerate() + .map(|(i, r)| Some((*rank.get(&r.mcid?)?, *xy_pos.get(i)?, i))) + .collect::>>()?; + keyed.sort_unstable(); + Some(keyed.into_iter().map(|(_, _, i)| i).collect()) +} + +/// XY-cut (D33): at each node first a horizontal cut at the widest empty band spanning the node +/// (≥ `XY_CUT_ROW_GAP_EM` × the node's median size), else a vertical cut at the widest empty +/// gutter (≥ `XY_CUT_COL_GAP_EM` × median size); leaves are sorted into lines by baseline, then +/// left to right. Depth ≤ `XY_CUT_DEPTH_MAX`. +pub fn xy_cut_order(runs: &[RunBox]) -> Vec { + let mut out = Vec::with_capacity(runs.len()); + cut(runs, (0..runs.len()).collect(), 0, &mut out); + out +} + +#[derive(Clone, Copy)] +enum Axis { + X, + Y, +} + +fn cut(runs: &[RunBox], idx: Vec, depth: usize, out: &mut Vec) { + if idx.len() > 1 && depth < XY_CUT_DEPTH_MAX { + let size = median_size(runs, &idx).max(SIZE_FLOOR); + for (axis, share) in [(Axis::Y, XY_CUT_ROW_GAP_EM), (Axis::X, XY_CUT_COL_GAP_EM)] { + if let Some((first, second)) = split(runs, &idx, axis, share * size) { + cut(runs, first, depth + 1, out); + cut(runs, second, depth + 1, out); + return; + } + } + } + leaf(runs, idx, out); +} + +fn median_size(runs: &[RunBox], idx: &[usize]) -> f64 { + let mut sizes: Vec = idx + .iter() + .filter_map(|i| runs.get(*i).map(|r| r.size)) + .filter(|s| s.is_finite()) + .collect(); + sizes.sort_by(|a, b| a.total_cmp(b)); + sizes.get(sizes.len() / 2).copied().unwrap_or(0.0) +} + +fn interval(r: &RunBox, axis: Axis) -> (f64, f64) { + match axis { + Axis::X => (r.display_rect[0], r.display_rect[2]), + Axis::Y => (r.display_rect[1], r.display_rect[3]), + } +} + +/// Splits at the widest empty band on `axis` that is at least `min_gap` wide. +fn split( + runs: &[RunBox], + idx: &[usize], + axis: Axis, + min_gap: f64, +) -> Option<(Vec, Vec)> { + let mut iv: Vec<(f64, f64, usize)> = idx + .iter() + .filter_map(|i| { + runs.get(*i).map(|r| { + let (lo, hi) = interval(r, axis); + (lo, hi, *i) + }) + }) + .collect(); + iv.sort_by(|a, b| a.0.total_cmp(&b.0)); + let mut reach = f64::NEG_INFINITY; + let mut best: Option<(f64, f64)> = None; + for (k, (lo, hi, _)) in iv.iter().enumerate() { + if k > 0 { + let gap = lo - reach; + if gap > 0.0 && gap >= min_gap && best.map_or(true, |(g, _)| gap > g) { + best = Some((gap, reach)); + } + } + reach = reach.max(*hi); + } + let (_, at) = best?; + let (first, second): (Vec<(f64, f64, usize)>, Vec<(f64, f64, usize)>) = + iv.into_iter().partition(|(lo, _, _)| *lo <= at); + Some(( + first.into_iter().map(|x| x.2).collect(), + second.into_iter().map(|x| x.2).collect(), + )) +} + +/// Lines by baseline (clustered within half the smaller size), each left to right. +fn leaf(runs: &[RunBox], mut idx: Vec, out: &mut Vec) { + let base = |i: &usize| runs.get(*i).map_or(0.0, |r| r.baseline_y); + idx.sort_by(|a, b| base(a).total_cmp(&base(b))); + let mut lines: Vec> = Vec::new(); + for i in idx { + let Some(r) = runs.get(i) else { continue }; + let joins = lines + .last() + .and_then(|l| l.first()) + .and_then(|f| runs.get(*f)) + .is_some_and(|f| { + (f.baseline_y - r.baseline_y).abs() <= LINE_CLUSTER_SHARE * f.size.min(r.size) + }); + match lines.last_mut() { + Some(line) if joins => line.push(i), + _ => lines.push(vec![i]), + } + } + let left = |i: &usize| runs.get(*i).map_or(0.0, |r| r.display_rect[0]); + for mut line in lines { + line.sort_by(|a, b| left(a).total_cmp(&left(b))); + out.extend(line); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/poppler.rs b/src-tauri/src/pdf_engine/text_edit/poppler.rs new file mode 100644 index 0000000..e7e9452 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/poppler.rs @@ -0,0 +1,628 @@ +//! Independent-engine checks (SPEC §B.15, Phase A check A5): Poppler, not our model, reads and +//! renders the page before and after the edit. Words outside the edited bands must stay 1:1 (text +//! and box within `WORD_BBOX_TOL_PT`); inside them the new text must be there and the old text +//! gone (G-TEXT). Pixels outside the edited glyphs' masks must not change (G-RENDER); an edit that +//! changes no pixel inside its own mask is reported `EDIT_NOT_VISIBLE` (non-blocking). This breaks +//! the mobile B13 loop, where the gate verified the edit with the same width model that made it. +//! +//! Frames (pinned by IND-06, poppler 26.04): `pdftotext -bbox` and `pdftoppm` both use the +//! MediaBox with the display rotation applied, origin at the top left, y down — an offset CropBox +//! does not move them. + +use crate::error::AppError; +use crate::pdf_engine::text_edit::apply::verify_failed; +use crate::pdf_engine::text_edit::engines::{run_tool, Engines, RunOpts}; +use crate::pdf_engine::text_edit::geometry::PageGeometry; +use crate::pdf_engine::text_edit::limits::{ + EDIT_BAND_PAD_PT, PDFTOTEXT_OUTPUT_MAX, RENDER_CHANNEL_TOL, RENDER_DPI_MAX, + RENDER_KEPT_PIXELS_MAX, RENDER_MASK_PAD_PX, RENDER_OUTSIDE_PIXELS_MAX, RENDER_OWN_PAD_PX, + RENDER_PIXELS_MAX, +}; +use crate::pdf_engine::text_edit::rewrite::ExpectedRun; +use std::ffi::OsString; +use std::io::Read; +use std::path::Path; +use words::check_words; + +mod ink; +pub(crate) mod near; +mod words; + +pub use near::NearGlyphs; + +/// Points per inch. +const PT_PER_INCH: f64 = 72.0; +/// Bytes of a PPM header we accept. +const PPM_HEADER_MAX: usize = 64; + +#[derive(Debug, Clone, PartialEq)] +pub struct Word { + pub text: String, + pub x0: f64, + pub y0: f64, + pub x1: f64, + pub y1: f64, +} + +/// Decodes the XML entities pdftotext writes. +fn entities(s: &str) -> String { + let mut out = String::with_capacity(s.len()); + let mut rest = s; + while let Some(p) = rest.find('&') { + out.push_str(rest.get(..p).unwrap_or_default()); + let tail = rest.get(p..).unwrap_or_default(); + let Some(end) = tail.find(';').filter(|e| *e <= 10) else { + out.push('&'); + rest = tail.get(1..).unwrap_or_default(); + continue; + }; + let ent = tail.get(1..end).unwrap_or_default(); + let decoded = match ent { + "amp" => Some('&'), + "lt" => Some('<'), + "gt" => Some('>'), + "quot" => Some('"'), + "apos" => Some('\''), + _ => ent + .strip_prefix("#x") + .and_then(|h| u32::from_str_radix(h, 16).ok()) + .or_else(|| ent.strip_prefix('#').and_then(|d| d.parse().ok())) + .and_then(char::from_u32), + }; + match decoded { + Some(c) => out.push(c), + None => out.push_str(tail.get(..=end).unwrap_or_default()), + } + rest = tail.get(end + 1..).unwrap_or_default(); + } + out.push_str(rest); + out +} + +fn attr(tag: &str, name: &str) -> Option { + let key = format!("{name}=\""); + let start = tag.find(&key)? + key.len(); + let rest = tag.get(start..)?; + let end = rest.find('"')?; + rest.get(..end)? + .parse() + .ok() + .filter(|v: &f64| v.is_finite()) +} + +/// Parses the `text` elements of `pdftotext -bbox` output. +pub(crate) fn parse_words(xml: &str) -> Vec { + let mut out = Vec::new(); + let mut rest = xml; + while let Some(p) = rest.find("') else { break }; + let tag = tail.get(..close).unwrap_or_default(); + let body = tail.get(close + 1..).unwrap_or_default(); + let Some(end) = body.find("") else { + break; + }; + if let (Some(x0), Some(y0), Some(x1), Some(y1)) = ( + attr(tag, "xMin"), + attr(tag, "yMin"), + attr(tag, "xMax"), + attr(tag, "yMax"), + ) { + out.push(Word { + text: entities(body.get(..end).unwrap_or_default()), + x0, + y0, + x1, + y1, + }); + } + rest = body.get(end + "".len()..).unwrap_or_default(); + } + out +} + +/// `pdftotext -f n -l n -bbox -` (stdout ≤ `PDFTOTEXT_OUTPUT_MAX`). +pub fn pdftotext_words( + engines: &Engines, + pdf: &Path, + page_1: u32, + opts: &RunOpts<'_>, +) -> Result, AppError> { + let n = page_1.to_string(); + let args: Vec = vec![ + "-f".into(), + n.clone().into(), + "-l".into(), + n.into(), + "-bbox".into(), + pdf.as_os_str().to_os_string(), + "-".into(), + ]; + let capped = RunOpts { + handle: opts.handle, + cancel: opts.cancel, + timeout: opts.timeout, + stdout_cap: PDFTOTEXT_OUTPUT_MAX, + }; + let out = run_tool(&engines.pdftotext, &args, true, &capped)?; + if out.code != 0 { + return Err(verify_failed(&format!( + "check=words pdftotext exited with code {}: {}", + out.code, + out.stderr.trim() + ))); + } + Ok(parse_words(&String::from_utf8_lossy(&out.stdout))) +} + +#[derive(Debug, Clone, PartialEq)] +pub struct Raster { + pub w: u32, + pub h: u32, + pub rgb: Vec, +} + +/// Reads one whitespace-delimited header token of a PPM. +fn ppm_token(data: &[u8], at: &mut usize) -> Option { + while data.get(*at).is_some_and(u8::is_ascii_whitespace) { + *at += 1; + } + let start = *at; + while data.get(*at).is_some_and(u8::is_ascii_digit) { + *at += 1; + } + std::str::from_utf8(data.get(start..*at)?) + .ok()? + .parse() + .ok() +} + +/// Parses a binary (P6, maxval 255) PPM with a checked header. +pub(crate) fn parse_ppm(data: &[u8]) -> Option { + if data.get(..2) != Some(b"P6") { + return None; + } + let mut at = 2usize; + let w = ppm_token(data, &mut at)?; + let h = ppm_token(data, &mut at)?; + let max = ppm_token(data, &mut at)?; + if max != 255 || at > PPM_HEADER_MAX || !data.get(at).is_some_and(u8::is_ascii_whitespace) { + return None; + } + let pixels = w.checked_mul(h)?; + if w == 0 || h == 0 || pixels > RENDER_PIXELS_MAX { + return None; + } + let len = usize::try_from(pixels.checked_mul(3)?).ok()?; + let body = data.get(at + 1..)?; + if body.len() != len { + return None; + } + Some(Raster { + w: u32::try_from(w).ok()?, + h: u32::try_from(h).ok()?, + rgb: body.to_vec(), + }) +} + +/// `pdftoppm -r dpi -f n -l n -singlefile /`, then the PPM read back +/// (bounded by `RENDER_PIXELS_MAX`). +pub fn render_page( + engines: &Engines, + pdf: &Path, + page_1: u32, + dpi: u32, + work: &Path, + opts: &RunOpts<'_>, +) -> Result { + static SEQ: std::sync::atomic::AtomicUsize = std::sync::atomic::AtomicUsize::new(0); + let n = SEQ.fetch_add(1, std::sync::atomic::Ordering::SeqCst); + let prefix = work.join(format!("render-{}-{n}", std::process::id())); + let page = page_1.to_string(); + let args: Vec = vec![ + "-r".into(), + dpi.to_string().into(), + "-f".into(), + page.clone().into(), + "-l".into(), + page.into(), + "-singlefile".into(), + pdf.as_os_str().to_os_string(), + prefix.as_os_str().to_os_string(), + ]; + let out = run_tool(&engines.pdftoppm, &args, true, opts)?; + let file = prefix.with_extension("ppm"); + let result = if out.code != 0 { + Err(verify_failed(&format!( + "check=render pdftoppm exited with code {}: {}", + out.code, + out.stderr.trim() + ))) + } else { + let cap = RENDER_PIXELS_MAX + .saturating_mul(3) + .saturating_add(PPM_HEADER_MAX as u64 + 1); + let mut data = Vec::new(); + std::fs::File::open(&file) + .and_then(|f| f.take(cap.saturating_add(1)).read_to_end(&mut data)) + .map_err(|e| verify_failed(&format!("check=render could not read the render: {e}"))) + .and_then(|_| { + parse_ppm(&data) + .ok_or_else(|| verify_failed("check=render the render is not a usable PPM")) + }) + }; + let _ = std::fs::remove_file(&file); + result +} + +/// Lowest render resolution A5 verifies at (SPEC §B.15: "≥ 25 for any legal page"). +pub(crate) const RENDER_DPI_MIN: u32 = 25; + +/// Pixels of a pdftoppm render of the media box at `dpi` (each side rounded up). +fn render_pixels(w_in: f64, h_in: f64, dpi: u32) -> f64 { + (w_in * f64::from(dpi)).ceil() * (h_in * f64::from(dpi)).ceil() +} + +/// `min(RENDER_DPI_MAX, floor(sqrt(RENDER_PIXELS_MAX / (w_in · h_in))))` over the media box, +/// lowered while the rounded-up render would still pass `RENDER_PIXELS_MAX`. `None` when that is +/// below `RENDER_DPI_MIN` (a page larger than the largest legal one, 14,400 pt square): A5's pads +/// and threshold are in pixels, so below it a same-line follower's shift would go unseen. +pub fn render_dpi(geom: &PageGeometry) -> Option { + let w_in = (geom.media[2] - geom.media[0]).abs() / PT_PER_INCH; + let h_in = (geom.media[3] - geom.media[1]).abs() / PT_PER_INCH; + let area = w_in * h_in; + if !(area.is_finite() && area > 0.0) { + return None; + } + let budget = RENDER_PIXELS_MAX as f64; + let ideal = (budget / area).sqrt().floor(); + let mut dpi = if ideal.is_finite() && ideal >= f64::from(RENDER_DPI_MIN) { + (ideal.min(f64::from(RENDER_DPI_MAX))) as u32 + } else { + return None; + }; + while dpi >= RENDER_DPI_MIN && render_pixels(w_in, h_in, dpi) > budget { + dpi -= 1; + } + (dpi >= RENDER_DPI_MIN).then_some(dpi) +} + +/// A user-space point in the Poppler frame (MediaBox, display rotation, top-left origin, y down). +fn to_frame(x: f64, y: f64, geom: &PageGeometry) -> (f64, f64) { + let [mx0, my0, mx1, my1] = geom.media; + match geom.rotate.rem_euclid(360) { + 90 => (y - my0, x - mx0), + 180 => (mx1 - x, y - my0), + 270 => (my1 - y, mx1 - x), + _ => (x - mx0, my1 - y), + } +} + +/// `[x0, y0, x1, y1]` user rect → the same rect in the pdftotext frame. +pub fn user_rect_to_text_frame(rect: [f64; 4], geom: &PageGeometry) -> [f64; 4] { + let a = to_frame(rect[0], rect[1], geom); + let b = to_frame(rect[2], rect[3], geom); + [a.0.min(b.0), a.1.min(b.1), a.0.max(b.0), a.1.max(b.1)] +} + +/// `[x0, y0, x1, y1]` user rect → pixels of a pdftoppm render at `dpi` (outward rounding). +pub fn user_rect_to_pixels(rect: [f64; 4], geom: &PageGeometry, dpi: u32) -> [i64; 4] { + let f = user_rect_to_text_frame(rect, geom); + let s = f64::from(dpi) / PT_PER_INCH; + let clamp = |v: f64| { + if v.is_finite() { + v.clamp(-1e9, 1e9) as i64 + } else { + 0 + } + }; + [ + clamp((f[0] * s).floor()), + clamp((f[1] * s).floor()), + clamp((f[2] * s).ceil()), + clamp((f[3] * s).ceil()), + ] +} + +#[derive(Debug, Clone, PartialEq)] +pub struct IndependentReport { + /// Per edit: some pixel inside its mask changed (false ⇒ `EDIT_NOT_VISIBLE`). + pub edit_visible: Vec, +} + +pub(super) fn dilate(r: [f64; 4], pad: f64) -> [f64; 4] { + [r[0] - pad, r[1] - pad, r[2] + pad, r[3] + pad] +} + +pub(super) fn union(a: [f64; 4], b: [f64; 4]) -> [f64; 4] { + [ + a[0].min(b[0]), + a[1].min(b[1]), + a[2].max(b[2]), + a[3].max(b[3]), + ] +} + +fn differs(a: &[u8], b: &[u8]) -> bool { + a.iter() + .zip(b) + .any(|(x, y)| x.abs_diff(*y) > RENDER_CHANNEL_TOL) +} + +/// A mask box in raster pixels, half-open (`x0..x1`, `y0..y1`), clamped to the raster. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +struct PixelBox { + x0: usize, + y0: usize, + x1: usize, + y1: usize, +} + +/// A user rect in the pixels of a `w × h` render at `dpi` (outward rounding), grown by `pad` +/// pixels and clamped to the raster; `None` when nothing of it is left. +fn pixel_box( + rect: [f64; 4], + geom: &PageGeometry, + dpi: u32, + pad: i64, + (w, h): (usize, usize), +) -> Option { + let clamp = |v: i64, max: usize| usize::try_from(v.max(0)).map_or(max, |u| u.min(max)); + let p = user_rect_to_pixels(rect, geom, dpi); + let m = PixelBox { + x0: clamp(p[0].saturating_sub(pad), w), + y0: clamp(p[1].saturating_sub(pad), h), + x1: clamp(p[2].saturating_add(pad), w), + y1: clamp(p[3].saturating_add(pad), h), + }; + (m.x0 < m.x1 && m.y0 < m.y1).then_some(m) +} + +/// Each edit's mask boxes (§B.15: glyph boxes dilated by `EDIT_BAND_PAD_PT`, then +/// `RENDER_MASK_PAD_PX`) in the pixels of a `w × h` render; boxes outside the raster are dropped. +fn pixel_masks( + edits: &[(ExpectedRun, [f64; 4], [f64; 4], String)], + geom: &PageGeometry, + dpi: u32, + w: usize, + h: usize, +) -> Vec> { + edits + .iter() + .map(|(exp, _, _, _)| { + let grown = exp.mask_boxes.iter().map(|b| dilate(*b, EDIT_BAND_PAD_PT)); + grown + .filter_map(|b| pixel_box(b, geom, dpi, RENDER_MASK_PAD_PX, (w, h))) + .collect() + }) + .collect() +} + +/// What G-RENDER and the ink check tell apart under the masks (review-verify HIGH-A), in raster +/// pixels: `kept`, the boxes of glyphs no edit changes that some edit's masks reach; `own`, the +/// edited glyphs' own boxes (the new glyphs' masks, undilated, and the old glyphs' ink boxes, +/// each grown by `RENDER_OWN_PAD_PX`); `old`, the old glyphs' ink boxes alone (grown alike). A +/// pixel in a kept box and in no own box is checked even under the masks: the edit draws and +/// removes ink only within its own glyphs' boxes. +struct Layers { + kept: Vec, + own: Vec, + old: Vec, +} + +impl Layers { + fn new( + edits: &[(ExpectedRun, [f64; 4], [f64; 4], String)], + near: &NearGlyphs, + masks: &[Vec], + geom: &PageGeometry, + dpi: u32, + size: (usize, usize), + ) -> Layers { + let px = |r: &[f64; 4]| pixel_box(*r, geom, dpi, 0, size); + let own_px = |r: &[f64; 4]| pixel_box(*r, geom, dpi, RENDER_OWN_PAD_PX, size); + // Each edit's masks as one envelope: a kept glyph matters where some mask reaches it. + let envelopes: Vec = masks + .iter() + .filter_map(|boxes| { + boxes.iter().copied().reduce(|a, b| PixelBox { + x0: a.x0.min(b.x0), + y0: a.y0.min(b.y0), + x1: a.x1.max(b.x1), + y1: a.y1.max(b.y1), + }) + }) + .collect(); + let kept = near + .kept + .iter() + .filter_map(|k| px(&k.rect)) + .filter(|b| envelopes.iter().any(|e| overlaps(e, b))) + .collect(); + let old: Vec = edits + .iter() + .flat_map(|(exp, _, _, _)| exp.old_ink_boxes.iter().filter_map(own_px)) + .collect(); + let mut own: Vec = edits + .iter() + .flat_map(|(exp, _, _, _)| exp.new_glyph_masks().iter().filter_map(own_px)) + .collect(); + own.extend_from_slice(&old); + Layers { kept, own, old } + } +} + +fn overlaps(a: &PixelBox, b: &PixelBox) -> bool { + a.x0 < b.x1 && b.x0 < a.x1 && a.y0 < b.y1 && b.y0 < a.y1 +} + +/// Coverage layers of `check_pixels`'s row sweep. +const SLACK: usize = 0; +const KEPT: usize = 1; +const OWN: usize = 2; + +fn add_coverage(cover: &mut [i64], at: usize, v: i64) { + if let Some(c) = cover.get_mut(at) { + *c = c.saturating_add(v); + } +} + +/// G-RENDER: differing pixels (any channel off by more than `RENDER_CHANNEL_TOL`) outside every +/// mask fail the check once there are more than `RENDER_OUTSIDE_PIXELS_MAX`; per edit, whether a +/// differing pixel lies inside one of its own masks. A pixel of a glyph no edit changes (`near`) +/// outside the edited glyphs' own boxes is outside the masks too (review-verify HIGH-A: a narrow +/// follower lies within the masks' slack), and more than `RENDER_KEPT_PIXELS_MAX` of them fail +/// the check on their own. +/// +/// One pass over the rows: the union of the masks (and of the kept and own boxes) covering a row +/// is a coverage difference array updated only where a box starts or ends (O(pixels + +/// boxes·log boxes) in all), and a row's prefix count of differing pixels answers "does this mask +/// hold a changed pixel" in O(1), asked only on rows with a change, for edits not yet seen whose +/// masks span that row. The scan stops at the first row that pushes a count over its limit (the +/// details then carry the count so far). +pub(crate) fn check_pixels( + src: &Raster, + dst: &Raster, + geom: &PageGeometry, + edits: &[(ExpectedRun, [f64; 4], [f64; 4], String)], + near: &NearGlyphs, + dpi: u32, +) -> Result, (&'static str, String)> { + if src.w != dst.w || src.h != dst.h { + return Err(( + "render", + format!("render size {}x{} vs {}x{}", src.w, src.h, dst.w, dst.h), + )); + } + let (w, h) = (src.w as usize, src.h as usize); + let row_bytes = w.saturating_mul(3); + let mut visible = vec![false; edits.len()]; + if row_bytes == 0 || h == 0 { + return Ok(visible); + } + let masks = pixel_masks(edits, geom, dpi, w, h); + let layers = Layers::new(edits, near, &masks, geom, dpi, (w, h)); + // Rows each edit's masks span (a quick reject for rows far from the edit). + let spans: Vec<(usize, usize)> = masks + .iter() + .map(|boxes| { + boxes + .iter() + .fold((usize::MAX, 0), |(lo, hi), m| (lo.min(m.y0), hi.max(m.y1))) + }) + .collect(); + // Coverage events: (row, column, layer, ±1) where a box starts (y0) or ends (y1). + let boxes = (masks.iter().flatten().map(|m| (SLACK, m))) + .chain(layers.kept.iter().map(|m| (KEPT, m))) + .chain(layers.own.iter().map(|m| (OWN, m))); + let mut events: Vec<(usize, usize, usize, i64)> = boxes + .flat_map(|(l, m)| { + [ + (m.y0, m.x0, l, 1), + (m.y0, m.x1, l, -1), + (m.y1, m.x0, l, -1), + (m.y1, m.x1, l, 1), + ] + }) + .collect(); + events.sort_unstable(); + let mut next_event = 0usize; + let mut cover = [(); 3].map(|_| vec![0i64; w.saturating_add(1)]); + let mut prefix: Vec = Vec::with_capacity(w.saturating_add(1)); + let (mut outside, mut kept) = (0u64, 0u64); + let rows = src + .rgb + .chunks_exact(row_bytes) + .zip(dst.rgb.chunks_exact(row_bytes)) + .take(h); + for (y, (row_a, row_b)) in rows.enumerate() { + while let Some(&(ey, x, l, v)) = events.get(next_event) { + if ey > y { + break; + } + if let Some(layer) = cover.get_mut(l) { + add_coverage(layer, x, v); + } + next_event = next_event.saturating_add(1); + } + prefix.clear(); + prefix.push(0); + let (mut changed, mut row_outside, mut row_kept) = (0u32, 0u64, 0u64); + let mut covered = [0i64; 3]; + let pixels = row_a.chunks_exact(3).zip(row_b.chunks_exact(3)); + for (x, (a, b)) in pixels.enumerate() { + for (sum, layer) in covered.iter_mut().zip(&cover) { + *sum = sum.saturating_add(layer.get(x).copied().unwrap_or(0)); + } + if differs(a, b) { + changed = changed.saturating_add(1); + let in_kept = covered[KEPT] > 0 && covered[OWN] <= 0; + if in_kept { + row_kept = row_kept.saturating_add(1); + } + if covered[SLACK] <= 0 || in_kept { + row_outside = row_outside.saturating_add(1); + } + } + prefix.push(changed); + } + if changed == 0 { + continue; + } + outside = outside.saturating_add(row_outside); + kept = kept.saturating_add(row_kept); + if kept > RENDER_KEPT_PIXELS_MAX { + return Err(( + "render", + format!("next to the edited line: a glyph no edit changes changed, pixels={kept}"), + )); + } + if outside > RENDER_OUTSIDE_PIXELS_MAX { + return Err(("render", format!("pixels={outside}"))); + } + let changed_in = |m: &PixelBox| { + m.y0 <= y + && y < m.y1 + && prefix.get(m.x1).copied().unwrap_or(0) > prefix.get(m.x0).copied().unwrap_or(0) + }; + for ((seen, boxes), (lo, hi)) in visible.iter_mut().zip(&masks).zip(&spans) { + if !*seen && *lo <= y && y < *hi && boxes.iter().any(changed_in) { + *seen = true; + } + } + } + Ok(visible) +} + +/// §B.15 on one page: words (G-TEXT, `poppler/words.rs`) then pixels (G-RENDER); a neighbour of +/// an edited line that is no longer extracted at its place fails after the pixels, and one that +/// lost its ink under the edit's masks after that (`poppler/ink.rs`). `edits`: per edited run its +/// expectations, old and new user rects (`[x0, y0, x1, y1]`) and old text; `near`: the glyphs our +/// model draws around them (`poppler/near.rs`), which only say where to look harder. An edit +/// whose glyphs did not change is always visible (only its style changed, or nothing). +#[allow(clippy::too_many_arguments)] +pub fn check_independent( + src_words: &[Word], + dst_words: &[Word], + src_raster: &Raster, + dst_raster: &Raster, + geom: &PageGeometry, + edits: &[(ExpectedRun, [f64; 4], [f64; 4], String)], + near: &NearGlyphs, + dpi: u32, +) -> Result { + let words = check_words(src_words, dst_words, geom, edits, near)?; + let visible = check_pixels(src_raster, dst_raster, geom, edits, near, dpi)?; + if let Some(fail) = words.gone { + return Err(fail); + } + let raster = (src_raster, dst_raster); + ink::check_neighbour_ink(raster, geom, edits, near, &words.neighbours, dpi)?; + Ok(IndependentReport { + edit_visible: visible + .iter() + .zip(edits) + .map(|(v, (exp, _, _, _))| *v || !exp.glyphs_changed) + .collect(), + }) +} diff --git a/src-tauri/src/pdf_engine/text_edit/poppler/ink.rs b/src-tauri/src/pdf_engine/text_edit/poppler/ink.rs new file mode 100644 index 0000000..87d182e --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/poppler/ink.rs @@ -0,0 +1,171 @@ +//! G-RENDER's neighbour ink check (review-final HIGH-1). G-RENDER accepts any change inside the +//! edited glyphs' masks, so when a changed line runs into the next table cell, the cell's own +//! pixels lie under those masks and nothing else sees them: a planner that also moved, deleted +//! or whitened the cell would pass. Inside the masks, a neighbour's ink may only gain ink (the new +//! glyphs drawn over or under it); a pixel of its box that held ink before (it differed from the +//! box's most common colour, its paper) and shows that paper after the edit lost it. More than +//! `RENDER_KEPT_PIXELS_MAX` such pixels fail A5 (an honest edit loses none; a 12 pt period is 5 +//! pixels, review-verify HIGH-A). The old glyphs' ink boxes are left out (the +//! edit removes their ink there), and pixels outside the masks — or in a kept glyph's box outside +//! the edited glyphs' own boxes (review-verify HIGH-A) — are G-RENDER's. +//! +//! A neighbour re-coloured to another visible colour keeps its ink (this check cannot tell it from +//! new glyphs drawn over the cell in that colour); our own walk (A4 `STATE_CHANGED`) refuses it. + +use super::{overlaps, pixel_masks, Layers, NearGlyphs, PixelBox, Raster, PT_PER_INCH}; +use crate::pdf_engine::text_edit::geometry::PageGeometry; +use crate::pdf_engine::text_edit::limits::{RENDER_CHANNEL_TOL, RENDER_KEPT_PIXELS_MAX}; +use crate::pdf_engine::text_edit::rewrite::ExpectedRun; +use std::collections::HashMap; + +type Fail = (&'static str, String); + +/// Work one check may do (pixels visited plus mask rows tested), per raster pixel: past it the +/// check fails closed (only pathological pages of overlapping words and glyphs come near it). +const WORK_PER_RASTER_PIXEL: usize = 4; + +fn close(a: &[u8], b: &[u8]) -> bool { + a.iter() + .zip(b) + .all(|(x, y)| x.abs_diff(*y) <= RENDER_CHANNEL_TOL) +} + +/// A text-frame rect (points, origin top left) in raster pixels, half-open and clamped. +fn frame_to_pixels(r: &[f64; 4], dpi: u32, w: usize, h: usize) -> Option { + let s = f64::from(dpi) / PT_PER_INCH; + let px = |v: f64, max: usize| -> usize { + if v.is_finite() && v > 0.0 { + (v.min(max as f64)) as usize + } else { + 0 + } + }; + let b = PixelBox { + x0: px((r[0] * s).floor(), w), + y0: px((r[1] * s).floor(), h), + x1: px((r[2] * s).ceil(), w), + y1: px((r[3] * s).ceil(), h), + }; + (b.x0 < b.x1 && b.y0 < b.y1).then_some(b) +} + +/// The pixel at (x, y) of a raster `w` pixels wide. +fn pixel(r: &Raster, w: usize, x: usize, y: usize) -> &[u8] { + let at = y.saturating_mul(w).saturating_add(x).saturating_mul(3); + r.rgb.get(at..at.saturating_add(3)).unwrap_or_default() +} + +/// The most common colour of `b` in `raster` (the paper the neighbour is printed on). +fn paper(raster: &Raster, w: usize, b: &PixelBox) -> [u8; 3] { + let mut counts: HashMap<[u8; 3], usize> = HashMap::new(); + for y in b.y0..b.y1 { + for x in b.x0..b.x1 { + if let [red, green, blue] = pixel(raster, w, x, y) { + *counts.entry([*red, *green, *blue]).or_default() += 1; + } + } + } + counts + .into_iter() + .max_by_key(|(c, n)| (*n, *c)) + .map_or([255; 3], |(c, _)| c) +} + +/// Coverage layers of one neighbour box's rows. +const SLACK: usize = 0; +const KEPT: usize = 1; +const OWN: usize = 2; +const OLD: usize = 3; + +/// The check (see the module doc). `neighbours`: text-frame boxes of the words set aside by +/// G-TEXT, as the source had them. +pub(super) fn check_neighbour_ink( + (src, dst): (&Raster, &Raster), + geom: &PageGeometry, + edits: &[(ExpectedRun, [f64; 4], [f64; 4], String)], + near: &NearGlyphs, + neighbours: &[[f64; 4]], + dpi: u32, +) -> Result<(), Fail> { + let (w, h) = (src.w as usize, src.h as usize); + if neighbours.is_empty() || dst.w != src.w || dst.h != src.h { + return Ok(()); + } + let per_edit = pixel_masks(edits, geom, dpi, w, h); + let layers = Layers::new(edits, near, &per_edit, geom, dpi, (w, h)); + let masks: Vec = per_edit.into_iter().flatten().collect(); + let mut budget = w.saturating_mul(h).saturating_mul(WORK_PER_RASTER_PIXEL); + let too_much = || ("render", "too many neighbour words to compare".to_string()); + let mut lost = 0u64; + // Per layer and row of a neighbour box: +1 where a box starts covering, −1 where it stops. + let mut diff: [Vec; 4] = Default::default(); + for b in neighbours + .iter() + .filter_map(|r| frame_to_pixels(r, dpi, w, h)) + { + let near_masks: Vec<&PixelBox> = masks.iter().filter(|m| overlaps(m, &b)).collect(); + if near_masks.is_empty() { + continue; + } + let candidates = (layers.kept.len()) + .saturating_add(layers.own.len()) + .saturating_add(layers.old.len()); + budget = budget.checked_sub(candidates).ok_or_else(too_much)?; + let layered: Vec<(usize, &PixelBox)> = (near_masks.into_iter().map(|m| (SLACK, m))) + .chain(layers.kept.iter().map(|m| (KEPT, m))) + .chain(layers.own.iter().map(|m| (OWN, m))) + .chain(layers.old.iter().map(|m| (OLD, m))) + .filter(|(_, m)| overlaps(m, &b)) + .collect(); + let (width, height) = (b.x1 - b.x0, b.y1 - b.y0); + let work = width + .saturating_mul(diff.len()) + .saturating_add(layered.len()) + .saturating_mul(height) + .saturating_mul(2); + budget = budget.checked_sub(work).ok_or_else(too_much)?; + let ink = paper(src, w, &b); + for y in b.y0..b.y1 { + for d in diff.iter_mut() { + d.clear(); + d.resize(width.saturating_add(1), 0); + } + for (layer, m) in layered.iter().filter(|(_, m)| m.y0 <= y && y < m.y1) { + let from = m.x0.max(b.x0) - b.x0; + let to = m.x1.min(b.x1).saturating_sub(b.x0); + if let (true, Some(d)) = (from < to, diff.get_mut(*layer)) { + if let Some(v) = d.get_mut(from) { + *v = v.saturating_add(1); + } + if let Some(v) = d.get_mut(to) { + *v = v.saturating_sub(1); + } + } + } + let mut cover = [0i64; 4]; + for dx in 0..width { + for (c, d) in cover.iter_mut().zip(&diff) { + *c = c.saturating_add(d.get(dx).copied().unwrap_or(0)); + } + let masked = cover[SLACK] > 0 && (cover[KEPT] <= 0 || cover[OWN] > 0); + if !masked || cover[OLD] > 0 { + continue; + } + let x = b.x0 + dx; + let (a, z) = (pixel(src, w, x, y), pixel(dst, w, x, y)); + if !close(a, &ink) && close(z, &ink) { + lost = lost.saturating_add(1); + } + } + } + } + if lost > RENDER_KEPT_PIXELS_MAX { + return Err(( + "render", + format!( + "next to the edited line: a neighbour lost its ink under the edit, pixels={lost}" + ), + )); + } + Ok(()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/poppler/near.rs b/src-tauri/src/pdf_engine/text_edit/poppler/near.rs new file mode 100644 index 0000000..96b4600 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/poppler/near.rs @@ -0,0 +1,98 @@ +//! The glyphs our own model draws on and around the edited lines (review-verify HIGH-A). They only +//! tell A5 where to look harder; Poppler's words and pixels still decide. +//! - G-TEXT (`words.rs`) counts a Poppler word that holds glyphs of runs no edit changes as a +//! neighbour wherever it lies: a narrow follower right after the edited run (a 7 pt footnote +//! marker, a ":" in another colour) has its centre within `EDIT_BAND_PAD_PT` of the edited +//! run's old box, which made it the run's own word. A word Poppler joined across the edited run +//! and such a run ("Hello:") is split: the kept glyphs at either end are a neighbour pinned at +//! that outer edge. +//! - G-RENDER (`poppler.rs`) and the ink check (`ink.rs`) take the kept glyphs' boxes out of the +//! masks wherever the edited glyphs' own boxes (the new glyphs' masks, the old glyphs' ink +//! boxes) do not cover them, so a follower's pixels must not change. + +use super::{dilate, union, user_rect_to_text_frame, words::fold, PT_PER_INCH}; +use crate::pdf_engine::text_edit::geometry::PageGeometry; +use crate::pdf_engine::text_edit::limits::{EDIT_BAND_PAD_PT, RENDER_MASK_PAD_PX}; +use crate::pdf_engine::text_edit::rewrite::ExpectedRun; +use crate::pdf_engine::text_edit::runs::{PageModel, TextRun}; +use std::collections::HashSet; + +/// A glyph of a run no edit changes. +#[derive(Debug, Clone, PartialEq)] +pub struct KeptGlyph { + /// User space: the advance × the font's ascent/descent, rise included (the walk's box). + pub rect: [f64; 4], + /// Its text folded as G-TEXT folds Poppler's words (`None`: the font maps it to none). + pub text: Option, +} + +/// What A5 learns from the model about the edited lines. +#[derive(Debug, Clone, Default, PartialEq)] +pub struct NearGlyphs { + /// The glyphs of the runs no edit changes, in the rows the edits' bands and masks span + /// (Poppler's frame). Records whose pen the model does not know are left out. + pub kept: Vec, + /// The edited runs' old glyphs that read as something (user space, the walk's boxes): which + /// words were theirs. + pub edited: Vec<[f64; 4]>, +} + +impl NearGlyphs { + /// From `model`, the page as qpdf's input has it; `edited`, its runs the plan changes; + /// `edits`, A5's per-edit expectations and old and new rects. + pub fn of( + model: &PageModel, + edited: &[&TextRun], + edits: &[(ExpectedRun, [f64; 4], [f64; 4], String)], + geom: &PageGeometry, + dpi: u32, + ) -> NearGlyphs { + let members: HashSet = edited + .iter() + .flat_map(|r| r.members.iter().copied()) + .collect(); + let rows = rows(edits, geom, dpi); + let in_rows = |r: &[f64; 4]| { + let f = user_rect_to_text_frame(*r, geom); + rows.iter().any(|(lo, hi)| f[1] < *hi && *lo < f[3]) + }; + let mut near = NearGlyphs::default(); + for (i, rec) in model.walk.records.iter().enumerate() { + if members.contains(&i) { + let inked = rec + .glyphs + .iter() + .filter(|g| g.text.as_deref().map(fold) != Some(String::new())); + near.edited.extend(inked.map(|g| g.bbox)); + } else if !rec.pen_unknown { + let kept = rec.glyphs.iter().filter(|g| in_rows(&g.bbox)); + near.kept.extend(kept.map(|g| KeptGlyph { + rect: g.bbox, + text: g.text.as_deref().map(fold), + })); + } + } + near + } +} + +/// Per edit, the rows (Poppler's frame, `(y0, y1)`) of its band and masks: the old and new rects +/// and every mask box, grown by `EDIT_BAND_PAD_PT` and then `RENDER_MASK_PAD_PX` at `dpi`. +fn rows( + edits: &[(ExpectedRun, [f64; 4], [f64; 4], String)], + geom: &PageGeometry, + dpi: u32, +) -> Vec<(f64, f64)> { + let slack = RENDER_MASK_PAD_PX as f64 * PT_PER_INCH / f64::from(dpi.max(1)); + edits + .iter() + .map(|(exp, old, new, _)| { + let all = exp + .mask_boxes + .iter() + .fold(union(*old, *new), |a, b| union(a, *b)); + let f = user_rect_to_text_frame(dilate(all, EDIT_BAND_PAD_PT), geom); + (f[1] - slack, f[3] + slack) + }) + .collect() +} diff --git a/src-tauri/src/pdf_engine/text_edit/poppler/words.rs b/src-tauri/src/pdf_engine/text_edit/poppler/words.rs new file mode 100644 index 0000000..f68d107 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/poppler/words.rs @@ -0,0 +1,584 @@ +//! G-TEXT (SPEC §B.15): Poppler's words before and after the edit. Words outside the edited +//! bands must stay 1:1 (text and box within `WORD_BBOX_TOL_PT`). Inside a band, the words of the +//! line's other runs (its neighbours: table cells, tab stops) must still be extracted as the same +//! word at their place, or joined with the changed line's glyphs at their outer edge, and are set +//! aside; what remains must hold the new text, and the old text must be gone. Their boxes go on +//! to G-RENDER's ink check (`poppler/ink.rs`), which sees what the edit's masks hide. +//! +//! A word is a neighbour when it holds glyphs of runs no edit changes (`poppler/near.rs`, our +//! model only saying where to look), wherever it lies, else when its centre lies outside every +//! edited run's old box (+ `EDIT_BAND_PAD_PT`). A word Poppler joined across an edited run and +//! such glyphs ("Hello:" with a ":" in another colour) is split: the edited part is the run's, and +//! the kept glyphs at its start or end are a neighbour pinned at that outer edge (review-verify +//! HIGH-A). + +use super::{dilate, union, user_rect_to_text_frame, NearGlyphs, Word}; +use crate::pdf_engine::text_edit::geometry::PageGeometry; +use crate::pdf_engine::text_edit::limits::{EDIT_BAND_PAD_PT, WORD_BBOX_TOL_PT}; +use crate::pdf_engine::text_edit::rewrite::ExpectedRun; + +/// A G-TEXT failure: (check id, detail). +type Fail = (&'static str, String); + +/// G-TEXT word matching: bucket entries one comparison may pass over, per word on the page (an +/// honest page's unchanged words have the same boxes on both sides, so a search stops at its +/// first look; only words of equal text a hair apart cost more), and at least this many in all. +const WORD_SCAN_PER_WORD: usize = 16; +const WORD_SCAN_MIN: usize = 4_096; +/// Compatibility folding for the word comparison (the NFKC cases pdftotext output needs: the +/// FB00–FB06 ligatures, no-break and other spaces, fullwidth ASCII), whitespace removed. +pub(crate) fn fold(s: &str) -> String { + let mut out = String::with_capacity(s.len()); + for c in s.chars() { + match c { + '\u{fb00}' => out.push_str("ff"), + '\u{fb01}' => out.push_str("fi"), + '\u{fb02}' => out.push_str("fl"), + '\u{fb03}' => out.push_str("ffi"), + '\u{fb04}' => out.push_str("ffl"), + '\u{fb05}' | '\u{fb06}' => out.push_str("st"), + '\u{ff01}'..='\u{ff5e}' => { + if let Some(a) = char::from_u32(u32::from(c) - 0xfee0) { + out.push(a); + } + } + c if c.is_whitespace() || c == '\u{a0}' || c == '\u{2007}' || c == '\u{202f}' => {} + c => out.push(c), + } + } + out +} + +fn count_of(hay: &str, needle: &str) -> usize { + if needle.is_empty() { + return 0; + } + hay.match_indices(needle).count() +} + +fn in_band(w: &Word, band: (f64, f64)) -> bool { + w.y0 < band.1 && band.0 < w.y1 +} + +fn same_word(a: &Word, b: &Word) -> bool { + a.text == b.text + && (a.x0 - b.x0).abs() <= WORD_BBOX_TOL_PT + && (a.y0 - b.y0).abs() <= WORD_BBOX_TOL_PT + && (a.x1 - b.x1).abs() <= WORD_BBOX_TOL_PT + && (a.y1 - b.y1).abs() <= WORD_BBOX_TOL_PT +} + +/// The bucket cell of a word coordinate: cells are two tolerances wide, so the values within +/// `WORD_BBOX_TOL_PT` of `v` lie in at most two cells (`cells`). +fn cell(v: f64) -> i64 { + (v / (2.0 * WORD_BBOX_TOL_PT)).floor() as i64 +} + +fn cells(v: f64) -> std::ops::RangeInclusive { + cell(v - WORD_BBOX_TOL_PT)..=cell(v + WORD_BBOX_TOL_PT) +} + +fn moved(w: &Word) -> Fail { + ( + "words", + format!( + "word {:?} at ({:.2}, {:.2}) moved or changed", + w.text, w.x0, w.y0 + ), + ) +} + +/// Matches words of `src` 1:1 to words of `dst` (text equal, every coordinate within +/// `WORD_BBOX_TOL_PT`). Returns the indices of the `dst` words left over, in order, and the +/// `src` words nothing matched. `dst` is bucketed by (text, x0 cell, y0 cell); a match is removed +/// from its bucket, and the entries all searches pass over are bounded (`WORD_SCAN_PER_WORD`), so +/// the comparison is linear in the word count however the words are ordered. +fn match_some<'a>(src: &[&'a Word], dst: &[&Word]) -> Result<(Vec, Vec<&'a Word>), Fail> { + let mut buckets: std::collections::HashMap<(&str, i64, i64), Vec> = + std::collections::HashMap::with_capacity(dst.len()); + for (i, w) in dst.iter().enumerate() { + buckets + .entry((w.text.as_str(), cell(w.x0), cell(w.y0))) + .or_default() + .push(i); + } + let limit = WORD_SCAN_PER_WORD + .saturating_mul(src.len()) + .max(WORD_SCAN_MIN); + let mut scanned = 0usize; + let mut missed = Vec::new(); + 'words: for w in src { + for cx in cells(w.x0) { + for cy in cells(w.y0) { + let Some(list) = buckets.get_mut(&(w.text.as_str(), cx, cy)) else { + continue; + }; + let mut found = None; + for (pos, d) in list.iter().enumerate() { + scanned = scanned.saturating_add(1); + if scanned > limit { + return Err(("words", "too many overlapping words to compare".to_string())); + } + if dst.get(*d).is_some_and(|d| same_word(w, d)) { + found = Some(pos); + break; + } + } + if let Some(pos) = found { + list.swap_remove(pos); + continue 'words; + } + } + } + missed.push(*w); + } + let mut rest: Vec = buckets.into_values().flatten().collect(); + rest.sort_unstable(); + Ok((rest, missed)) +} + +/// Every word of `src` must match a word of `dst` 1:1. +fn match_words(src: &[&Word], dst: &[&Word]) -> Result<(), Fail> { + match match_some(src, dst)?.1.first() { + Some(w) => Err(moved(w)), + None => Ok(()), + } +} + +/// The middle of a word's box lies in `r` (`[x0, y0, x1, y1]`, text frame). +fn centred_in(w: &Word, r: &[f64; 4]) -> bool { + let (x, y) = ((w.x0 + w.x1) / 2.0, (w.y0 + w.y1) / 2.0); + r[0] <= x && x <= r[2] && r[1] <= y && y <= r[3] +} + +/// Which edges of a neighbour are its own (review-verify HIGH-A). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum Edge { + /// A whole source word: both. + Whole, + /// The kept glyphs a joined word starts with: its x0 only (x1 is the model's). + Start, + /// The kept glyphs a joined word ends with: its x1 only (x0 is the model's). + End, +} + +/// A neighbour of an edited line: a source word, or the part of one at an outer edge. +#[derive(Debug, Clone)] +struct Neighbour { + word: Word, + edge: Edge, +} + +/// The text of `d` without the neighbour `n`'s (`part`), when `d` is `n` joined by Poppler with +/// glyphs of the changed line — the only inexact form a whole neighbour may take (review-final +/// HIGH-1): `d` reads longer than `n`, lies on `n`'s line (y0 and y1 within `WORD_BBOX_TOL_PT`), +/// keeps `n`'s outer edge (x1 for a word ending with `n`, x0 for one starting with it, within the +/// same tolerance), and reaches past `n`'s inner edge into the changed line's new box (`new_box`, +/// text frame). A neighbour moved, relabelled or swallowed anywhere else is not accepted. A part +/// of a joined word (`edge` `Start`/`End`) has only that outer edge, and may also come back as a +/// word of its own (`d` reads exactly `part` there: the empty rest), or — when the new box holds +/// the whole part (a "." under a new, wider "e") — inside a word that spans its box, joined with +/// the new glyphs, whose part past the part's outer edge lies in the new box (`covered`). +fn joined_with_edit( + d: &Word, + text: &str, + n: &Word, + part: &str, + new_box: &[f64; 4], + edge: Edge, +) -> Option { + let tol = WORD_BBOX_TOL_PT; + let near = |a: f64, b: f64| (a - b).abs() <= tol; + if !near(d.y0, n.y0) || !near(d.y1, n.y1) { + return None; + } + let at_end = edge != Edge::Start && near(d.x1, n.x1); + let at_start = edge != Edge::End && near(d.x0, n.x0); + if edge != Edge::Whole && (at_end || at_start) && text == part { + return Some(String::new()); + } + if text.len() <= part.len() { + return None; + } + let in_new_box = |lo: f64, hi: f64| lo < new_box[2] && new_box[0] < hi; + if at_end && d.x0 < n.x0 - tol && in_new_box(d.x0, n.x0) { + if let Some(head) = text.strip_suffix(part) { + return Some(head.to_string()); + } + } + if at_start && d.x1 > n.x1 + tol && in_new_box(n.x1, d.x1) { + if let Some(tail) = text.strip_prefix(part) { + return Some(tail.to_string()); + } + } + match edge { + Edge::Whole => None, + _ => covered(d, text, n, part, new_box, edge), + } +} + +/// A part of a joined word that the new glyphs cover: Poppler merges it into the new word, so its +/// outer edge is no longer a word's. Accepted inside `d` when the new box holds the part, `d` +/// spans the part's box, reaches past its inner edge into the new box, and ends past its outer +/// edge only within the new box; the part's text is taken out of `d`'s once (its last +/// occurrence for an end part, its first for a start part). +fn covered( + d: &Word, + text: &str, + n: &Word, + part: &str, + new_box: &[f64; 4], + edge: Edge, +) -> Option { + let tol = WORD_BBOX_TOL_PT; + let inside = |lo: f64, hi: f64| new_box[0] - tol <= lo && hi <= new_box[2] + tol; + if part.is_empty() || !inside(n.x0, n.x1) || d.x0 > n.x0 + tol || d.x1 < n.x1 - tol { + return None; + } + let reaches = match edge { + Edge::End => d.x0 < n.x0 - tol && inside(n.x1, d.x1), + _ => d.x1 > n.x1 + tol && inside(d.x0, n.x0), + }; + let at = match edge { + Edge::End => text.rfind(part), + _ => text.find(part), + }; + let at = at.filter(|_| reaches && text.len() > part.len())?; + let rest = [text.get(..at)?, text.get(at + part.len()..)?].concat(); + Some(rest) +} + +/// Sets a band's neighbours aside. A whole word must be extracted after the edit as the same word +/// (text and box within `WORD_BBOX_TOL_PT`), and any neighbour may be joined with the changed +/// line's glyphs (`joined_with_edit`). Returns the band's other words (folded, in Poppler's +/// order), and the first neighbour not found (reported after G-RENDER, so a moved follower still +/// fails the pixels first, as IND-09 pins). +fn set_aside( + neighbours: &[Neighbour], + dst: &[&Word], + new_box: &[f64; 4], +) -> Result<(Vec, Option), Fail> { + let whole: Vec<&Word> = neighbours + .iter() + .filter(|n| n.edge == Edge::Whole) + .map(|n| &n.word) + .collect(); + let (rest, missed) = match_some(&whole, dst)?; + let mut words: Vec<(&Word, String)> = rest + .iter() + .filter_map(|i| dst.get(*i)) + .map(|w| (*w, fold(&w.text))) + .collect(); + let parts = neighbours.iter().filter(|n| n.edge != Edge::Whole); + let unmatched = + (missed.into_iter().map(|w| (w, Edge::Whole))).chain(parts.map(|n| (&n.word, n.edge))); + let mut miss = None; + for (n, edge) in unmatched { + let part = fold(&n.text); + let joined = words.iter_mut().find_map(|(d, text)| { + joined_with_edit(d, text, n, &part, new_box, edge).map(|left| *text = left) + }); + if joined.is_none() && miss.is_none() { + let (check, detail) = moved(n); + miss = Some((check, format!("next to the edited line: {detail}"))); + } + } + Ok((words.into_iter().map(|(_, t)| t).collect(), miss)) +} + +/// A glyph of our model (`NearGlyphs`) in Poppler's frame. +#[derive(Debug)] +struct FrameGlyph { + /// The centre. + cx: f64, + cy: f64, + x0: f64, + x1: f64, + y0: f64, + y1: f64, + /// `None` for an edited run's old glyph; a kept glyph's folded text (`Some(None)`: unknown). + kept: Option>, +} + +/// The model's glyphs by the x of their centres; a kept glyph that reads as nothing (a space) is +/// left out. +fn frame_glyphs(near: &NearGlyphs, geom: &PageGeometry) -> Vec { + let edited = near.edited.iter().map(|r| (r, None)); + let kept = (near.kept.iter()) + .filter(|k| k.text.as_deref() != Some("")) + .map(|k| (&k.rect, Some(k.text.clone()))); + let mut out: Vec = edited + .chain(kept) + .map(|(r, kept)| { + let f = user_rect_to_text_frame(*r, geom); + FrameGlyph { + cx: (f[0] + f[2]) / 2.0, + cy: (f[1] + f[3]) / 2.0, + x0: f[0], + x1: f[2], + y0: f[1], + y1: f[3], + kept, + } + }) + .filter(|g| g.cx.is_finite() && g.cy.is_finite()) + .collect(); + out.sort_by(|a, b| a.cx.total_cmp(&b.cx)); + out +} + +/// The glyphs of `w`, by x: a glyph's centre lies in `w`'s box and `w`'s middle height within +/// the glyph's (both within `WORD_BBOX_TOL_PT`; a 9 pt stamp drawn over a 26 pt title is not the +/// title's). Each glyph looked at costs one unit of `budget`. +fn glyphs_in<'g>( + glyphs: &'g [FrameGlyph], + w: &Word, + budget: &mut usize, +) -> Result, Fail> { + let tol = WORD_BBOX_TOL_PT; + let from = glyphs.partition_point(|g| g.cx < w.x0 - tol); + let mut out = Vec::new(); + for g in glyphs.get(from..).unwrap_or_default() { + if g.cx > w.x1 + tol { + break; + } + *budget = budget + .checked_sub(1) + .ok_or_else(|| ("words", "too many overlapping words to compare".to_string()))?; + let middle = (w.y0 + w.y1) / 2.0; + let own = g.y0 - tol <= middle && middle <= g.y1 + tol; + if own && w.y0 - tol <= g.cy && g.cy <= w.y1 + tol { + out.push(g); + } + } + Ok(out) +} + +/// One source word of an edited band: the text it gives the edited run (folded) and the +/// neighbours it holds. `inside`: our glyphs in it, by x; `in_a_run`: its centre lies in an edited +/// run's padded old box. A word that cannot be split cleanly (kept glyphs between edited ones, +/// texts that do not match Poppler's) is a whole neighbour: it must come back unchanged (fail +/// closed). +fn classify(w: &Word, inside: &[&FrameGlyph], in_a_run: bool) -> (Option, Vec) { + let whole = || { + let word = w.clone(); + ( + None, + vec![Neighbour { + word, + edge: Edge::Whole, + }], + ) + }; + let Some(first) = inside.iter().position(|g| g.kept.is_none()) else { + return match (inside.is_empty(), in_a_run) { + (true, true) => (Some(fold(&w.text)), Vec::new()), + _ => whole(), + }; + }; + if inside.iter().all(|g| g.kept.is_none()) { + return (Some(fold(&w.text)), Vec::new()); + } + let last = inside + .iter() + .rposition(|g| g.kept.is_none()) + .unwrap_or(first); + let (start, mid, end) = ( + inside.get(..first).unwrap_or_default(), + inside.get(first..=last).unwrap_or_default(), + inside.get(last + 1..).unwrap_or_default(), + ); + let text_of = |gs: &[&FrameGlyph]| -> Option { + gs.iter() + .map(|g| g.kept.clone().flatten()) + .collect::>>() + .map(|v| v.concat()) + }; + let (Some(a), Some(z), false) = ( + text_of(start), + text_of(end), + mid.iter().any(|g| g.kept.is_some()), + ) else { + return whole(); + }; + let folded = fold(&w.text); + let head = folded + .strip_prefix(a.as_str()) + .and_then(|r| r.strip_suffix(z.as_str())); + let Some(head) = head.filter(|h| !h.is_empty()) else { + return whole(); + }; + let part = |text: String, x0: f64, x1: f64, edge: Edge| Neighbour { + word: Word { + text, + x0, + y0: w.y0, + x1, + y1: w.y1, + }, + edge, + }; + let mut parts = Vec::new(); + if let Some(x1) = start.iter().map(|g| g.x1).reduce(f64::max) { + parts.push(part(a, w.x0, x1, Edge::Start)); + } + if let Some(x0) = end.iter().map(|g| g.x0).reduce(f64::min) { + parts.push(part(z, x0, w.x1, Edge::End)); + } + (Some(head.to_string()), parts) +} + +/// What G-TEXT leaves for later: a neighbour that is gone (reported after G-RENDER), and the +/// boxes of every neighbour of an edited line (text frame) for G-RENDER's ink check. +pub(super) struct WordsOutcome { + pub gone: Option, + pub neighbours: Vec<[f64; 4]>, +} + +/// G-TEXT. `near`: our model's glyphs around the edits (where a word is a neighbour). +pub(super) fn check_words( + src: &[Word], + dst: &[Word], + geom: &PageGeometry, + edits: &[(ExpectedRun, [f64; 4], [f64; 4], String)], + near: &NearGlyphs, +) -> Result { + let bands: Vec<(f64, f64)> = edits + .iter() + .map(|(_, old, new, _)| { + let f = user_rect_to_text_frame(dilate(union(*old, *new), EDIT_BAND_PAD_PT), geom); + (f[1], f[3]) + }) + .collect(); + let outside = |w: &&Word| !bands.iter().any(|b| in_band(w, *b)); + let src_out: Vec<&Word> = src.iter().filter(outside).collect(); + let dst_out: Vec<&Word> = dst.iter().filter(outside).collect(); + if src_out.len() != dst_out.len() { + return Err(( + "words", + format!( + "words outside the edited lines: {} before, {} after", + src_out.len(), + dst_out.len() + ), + )); + } + match_words(&src_out, &dst_out)?; + // Where each edited run was (its old box, padded): a word there before the edit is the run's. + let runs: Vec<[f64; 4]> = edits + .iter() + .map(|(_, old, _, _)| user_rect_to_text_frame(dilate(*old, EDIT_BAND_PAD_PT), geom)) + .collect(); + let in_a_run = |w: &Word| runs.iter().any(|r| centred_in(w, r)); + let glyphs = frame_glyphs(near, geom); + let mut budget = WORD_SCAN_PER_WORD + .saturating_mul(src.len().saturating_add(glyphs.len())) + .max(WORD_SCAN_MIN); + let mut gone = None; + let mut neighbour_boxes = Vec::new(); + for ((exp, _, new_rect, old), band) in edits.iter().zip(&bands) { + let src_band: Vec<&Word> = src.iter().filter(|w| in_band(w, *band)).collect(); + let dst_band: Vec<&Word> = dst.iter().filter(|w| in_band(w, *band)).collect(); + // Neighbours on the edited line (table cells, tab stops) are set aside: Poppler orders a + // line's words by position, so a changed line that now runs into them would interleave + // with them (review-T5 live B2: an overlap is a warning, not a refusal). + let mut before: Vec = Vec::new(); + let mut neighbours: Vec = Vec::new(); + for w in &src_band { + let inside = glyphs_in(&glyphs, w, &mut budget)?; + let (own, parts) = classify(w, &inside, in_a_run(w)); + before.extend(own); + neighbours.extend(parts); + } + let new_box = user_rect_to_text_frame(dilate(*new_rect, EDIT_BAND_PAD_PT), geom); + let (after, miss) = set_aside(&neighbours, &dst_band, &new_box)?; + neighbour_boxes.extend(neighbours.iter().map(|n| { + let w = &n.word; + [w.x0, w.y0, w.x1, w.y1] + })); + // A neighbour that moved stays among the line's words, between the new ones: it is the + // reason the new text does not read (review-final HIGH-1). + band_text(exp, old, &before, &after).map_err(|f| { + match (&miss, f.1.starts_with(NEW_TEXT)) { + (Some(m), true) => m.clone(), + _ => f, + } + })?; + gone = gone.or(miss); + } + Ok(WordsOutcome { + gone, + neighbours: neighbour_boxes, + }) +} + +/// The start of the "new text not extracted" failure. +const NEW_TEXT: &str = "new text"; + +/// One band after its neighbours are set aside: the new text is extracted and the old text gone. +fn band_text( + exp: &ExpectedRun, + old: &str, + before: &[String], + after: &[String], +) -> Result<(), Fail> { + let (before_text, after_text) = (before.concat(), after.concat()); + // The request, not the planner's reading of its glyphs (§B.15 "each non-empty new text"). + let new = fold(&exp.requested_text); + if !new.is_empty() && !after_text.contains(&new) { + return Err(( + "words", + format!("{NEW_TEXT} {:?} not extracted", exp.requested_text), + )); + } + let old_folded = fold(old); + if !old_folded.is_empty() && !new.contains(&old_folded) { + let (b, a) = ( + count_of(&before_text, &old_folded), + count_of(&after_text, &old_folded), + ); + if a + 1 > b { + return Err(("words", format!("old text {old_folded:?} still extracted"))); + } + } + removed_words_gone(old, &exp.requested_text, before, after) +} + +/// Every word the edit removed (a word of the old text the new text has fewer of) must be +/// extracted that many times fewer, when Poppler extracted it as a word before. Poppler drops +/// text drawn twice at almost one position and orders words by position, so a fake that keeps +/// the old text under new text can defeat the concatenated-text test above, not this one. +fn removed_words_gone( + old: &str, + new: &str, + before: &[String], + after: &[String], +) -> Result<(), Fail> { + let words = |t: &str| -> Vec { + t.split_whitespace() + .map(fold) + .filter(|w| !w.is_empty()) + .collect() + }; + let count = |list: &[String], w: &str| list.iter().filter(|x| x.as_str() == w).count(); + let (old_words, new_words) = (words(old), words(new)); + let mut seen: Vec<&String> = Vec::new(); + for w in &old_words { + if seen.contains(&w) { + continue; + } + seen.push(w); + let removed = count(&old_words, w).saturating_sub(count(&new_words, w)); + let had = count(before, w); + if removed == 0 || had < removed { + continue; + } + if count(after, w) > had - removed { + return Err(("words", format!("removed word {w:?} still extracted"))); + } + } + Ok(()) +} + +#[cfg(test)] +mod tests; diff --git a/src-tauri/src/pdf_engine/text_edit/poppler/words/tests.rs b/src-tauri/src/pdf_engine/text_edit/poppler/words/tests.rs new file mode 100644 index 0000000..967c313 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/poppler/words/tests.rs @@ -0,0 +1,227 @@ +//! G-TEXT unit tests: the neighbour join rules and the word classification (review-final +//! HIGH-1, review-verify HIGH-A). + +use super::{classify, glyphs_in, joined_with_edit, Edge, FrameGlyph, Word}; + +fn w(text: &str, x0: f64, x1: f64) -> Word { + Word { + text: text.into(), + x0, + y0: 84.1, + x1, + y1: 96.4, + } +} + +/// review-final HIGH-1: a neighbour that does not match exactly counts only as joined with +/// the changed line's glyphs: longer text, its line, its outer edge, and the rest of the word +/// in the new box. The old `find` fallback (its text anywhere in a word at its place) is gone. +#[test] +fn a_neighbour_counts_as_joined_only_at_its_outer_edge_and_into_the_new_box() { + let n = w("42", 105.34, 118.68); + let new_box = [70.0, 80.0, 132.0, 100.0]; + let joined = |d: &Word, b: &[f64; 4]| joined_with_edit(d, &d.text, &n, "42", b, Edge::Whole); + assert_eq!( + joined(&w("world42", 90.0, 118.68), &new_box).as_deref(), + Some("world") + ); + assert_eq!( + joined(&w("42ab", 105.34, 125.0), &new_box).as_deref(), + Some("ab") + ); + let refused = [ + ("moved", w("42", 105.82, 119.16)), + ("moved and joined", w("x42", 95.0, 119.16)), + ("relabelled 425 at its box", w("425", 105.34, 118.68)), + ("relabelled 142 at its box", w("142", 105.34, 118.68)), + ("inside a longer word", w("a42b", 95.0, 118.68)), + ]; + for (what, d) in refused { + assert_eq!(joined(&d, &new_box), None, "{what}"); + } + let left_of_n = [70.0, 80.0, 100.0, 100.0]; + assert_eq!( + joined(&w("42ab", 105.34, 125.0), &left_of_n), + None, + "join outside the edit" + ); + let mut lower = w("x42", 95.0, 118.68); + lower.y0 += 1.0; + assert_eq!(joined(&lower, &new_box), None, "another line"); +} + +fn g(x0: f64, x1: f64, kept: Option>) -> FrameGlyph { + FrameGlyph { + cx: (x0 + x1) / 2.0, + cy: 90.0, + x0, + x1, + y0: 84.1, + y1: 96.4, + kept: kept.map(|t| t.map(str::to_string)), + } +} + +/// "Hello" as the edited run's old glyphs (Helvetica 12 pt from x = 72). +fn hello() -> Vec { + let edges = [72.0, 80.67, 87.34, 90.0, 92.67, 99.34]; + edges.windows(2).map(|e| g(e[0], e[1], None)).collect() +} + +/// review-verify HIGH-A: a word holding glyphs of runs no edit changes is a neighbour wherever +/// it lies; a word Poppler joined across the edit is split at its outer edges; a word that +/// cannot be split cleanly must come back whole. +#[test] +fn words_with_kept_glyphs_are_neighbours_and_joined_words_split_at_their_outer_edges() { + // A 7 pt footnote marker whose centre lies within the old box's 2 pt pad. + let marker = g(99.34, 103.23, Some(Some("1"))); + let (own, parts) = classify(&w("1", 99.34, 103.23), &[&marker], true); + assert_eq!((own, parts.len(), parts[0].edge), (None, 1, Edge::Whole)); + // Nothing of ours in a word: the old box decides. + let (own, parts) = classify(&w("Hello", 72.0, 99.34), &[], true); + assert_eq!((own.as_deref(), parts.len()), (Some("Hello"), 0)); + assert_eq!(classify(&w("42", 140.0, 146.0), &[], false).1.len(), 1); + // "Hello:" with a red ":": "Hello" is the run's, ":" is pinned at the word's x1. + let edited = hello(); + let colon = g(99.34, 102.67, Some(Some(":"))); + let open = g(68.0, 72.0, Some(Some("("))); + let mid = g(84.0, 85.0, Some(Some("x"))); + let semicolon = g(99.34, 102.67, Some(Some(";"))); + let unknown = g(99.34, 102.67, Some(None)); + let joined = |extra: &[&'static str]| -> Vec<&FrameGlyph> { + let mut v: Vec<&FrameGlyph> = edited.iter().collect(); + for e in extra { + v.extend( + [&colon, &open, &mid, &semicolon, &unknown] + .into_iter() + .filter(|g| matches!(&g.kept, Some(t) if t.as_deref().unwrap_or("?") == *e)), + ); + } + v.sort_by(|a, b| a.cx.total_cmp(&b.cx)); + v + }; + let (own, parts) = classify(&w("Hello:", 72.0, 102.67), &joined(&[":"]), true); + assert_eq!(own.as_deref(), Some("Hello")); + let p = &parts[0]; + assert_eq!( + ( + parts.len(), + p.edge, + p.word.text.as_str(), + p.word.x0, + p.word.x1 + ), + (1, Edge::End, ":", 99.34, 102.67) + ); + // "(Hello:" kept at both ends. + let (own, parts) = classify(&w("(Hello:", 68.0, 102.67), &joined(&["(", ":"]), true); + assert_eq!(own.as_deref(), Some("Hello")); + let edges: Vec<(Edge, &str, f64, f64)> = parts + .iter() + .map(|p| (p.edge, p.word.text.as_str(), p.word.x0, p.word.x1)) + .collect(); + assert_eq!( + edges, + [ + (Edge::Start, "(", 68.0, 72.0), + (Edge::End, ":", 99.34, 102.67) + ] + ); + // Fail closed: a kept glyph between edited ones, a text Poppler does not read, a kept + // glyph the font maps to nothing, or nothing of the run's left. + let whole = [ + ("between", w("Hexllo", 72.0, 99.34), joined(&["x"])), + ("other text", w("Hello:", 72.0, 102.67), joined(&[";"])), + ("no text", w("Hello:", 72.0, 102.67), joined(&["?"])), + ("nothing left", w(":", 72.0, 102.67), joined(&[":"])), + ]; + for (what, word, inside) in whole { + let (own, parts) = classify(&word, &inside, true); + assert!(own.is_none(), "{what}"); + assert_eq!((parts.len(), parts[0].edge), (1, Edge::Whole), "{what}"); + } +} + +/// review-verify HIGH-A: the part of a joined word holds only its outer edge (its inner edge is +/// our model's): it may come back alone there, or joined with the new glyphs. +#[test] +fn a_part_of_a_joined_word_is_found_only_at_its_outer_edge() { + let n = w(":", 99.34, 102.67); + let new_box = [70.0, 80.0, 100.0, 100.0]; + let at = |d: &Word, edge| joined_with_edit(d, &d.text, &n, ":", &new_box, edge); + assert_eq!(at(&w(":", 99.3, 102.67), Edge::End).as_deref(), Some("")); + assert_eq!( + at(&w("Hellp:", 72.0, 102.67), Edge::End).as_deref(), + Some("Hellp") + ); + let refused = [ + ("moved", w(":", 99.0, 102.33)), + ("relabelled", w(";", 99.34, 102.67)), + ("wider at its outer edge", w(":", 99.34, 103.0)), + ]; + for (what, d) in refused { + assert_eq!(at(&d, Edge::End), None, "{what}"); + } + assert_eq!( + at(&w(":", 99.34, 102.67), Edge::Start).as_deref(), + Some(""), + "a start part keeps x0" + ); + assert_eq!( + at(&w(":", 99.34, 102.67), Edge::Whole), + None, + "a whole word is matched by its box, not here" + ); +} + +/// No new false refusal (review-verify HIGH-A): a "." that the new, wider "e" covers is merged +/// by Poppler into the new word ("phrasee."); it is found inside it, but only when the new box +/// holds it and the word spans its box. +#[test] +fn a_covered_part_is_found_inside_the_new_word_that_spans_it() { + let n = w(".", 379.55, 382.55); + let new_box = [327.95, 80.0, 386.86, 100.0]; + let at = |d: &Word, b: &[f64; 4], edge| joined_with_edit(d, &d.text, &n, ".", b, edge); + let merged = w("phrasee.", 348.24, 384.86); + assert_eq!(at(&merged, &new_box, Edge::End).as_deref(), Some("phrasee")); + let narrow = [327.95, 80.0, 381.0, 100.0]; + let refused = [ + ( + "not spanning its box", + w("phrasee.", 380.0, 384.86), + new_box, + ), + ("past the new box", w("phrasee.", 348.24, 390.0), new_box), + ("not inside the new box", merged.clone(), narrow), + ("not its text", w("phrasee", 348.24, 384.86), new_box), + ]; + for (what, d, b) in refused { + assert_eq!(at(&d, &b, Edge::End), None, "{what}"); + } + assert_eq!( + at(&merged, &new_box, Edge::Whole), + None, + "a whole word is not a part" + ); +} + +/// A 9 pt stamp glyph whose centre falls inside a 26 pt title word is not the title's. +#[test] +fn a_glyph_belongs_to_a_word_only_across_its_middle() { + let title = Word { + text: "Quarterly".into(), + x0: 33.75, + y0: 34.11, + x1: 136.38, + y1: 60.69, + }; + let mut letter = g(40.0, 50.0, Some(Some("Q"))); + (letter.cy, letter.y0, letter.y1) = (47.4, 34.1, 60.7); + let mut stamp = g(72.0, 78.0, Some(Some("A"))); + (stamp.cy, stamp.y0, stamp.y1) = (37.7, 33.5, 41.9); + let glyphs = [letter, stamp]; + let mut budget = 16; + let inside = glyphs_in(&glyphs, &title, &mut budget).expect("within the budget"); + let texts: Vec> = inside.iter().map(|g| g.kept.clone().flatten()).collect(); + assert_eq!(texts, [Some("Q".to_string())]); +} diff --git a/src-tauri/src/pdf_engine/text_edit/preview.rs b/src-tauri/src/pdf_engine/text_edit/preview.rs new file mode 100644 index 0000000..5ff10ab --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/preview.rs @@ -0,0 +1,392 @@ +//! Preview of a patched page (SPEC §B.17): the same path as Save — `plan_page`, a qpdf update, +//! Phase A — on a one-page extraction of the source snapshot, so what the editor shows after a +//! commit is what Save will write (D9, D30). +//! +//! The caller (the command layer) first waits for the source's background `qpdf --check` and +//! passes its benign warnings as `source_benign` (a non-benign result is `PDF_NEEDS_REPAIR` there). +//! Files in `cache_dir` (`/textedit//`) are written once, atomically (a temp +//! name, then a rename; when the final name already exists the temp file is dropped and the +//! existing one, whose content is determined by the fingerprint, is used), so concurrent calls +//! never read a partial file. qpdf only ever reads the snapshot's bytes (`source.pdf`), never the +//! user's file. Each preview works in its own nonce directory, removed at the end. + +use crate::error::AppError; +use crate::pdf_engine::text_edit::apply::{apply_update, updates_for_plan, write_update_json}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::engines::{qpdf_page_map, run_tool, Engines, RunOpts}; +use crate::pdf_engine::text_edit::gate::{ + cancel_flag, is_cancelled, verify_edited_copy, EditedPageInput, PhaseAInput, PopplerRef, + TextWarning, +}; +use crate::pdf_engine::text_edit::graph::graph_digest; +use crate::pdf_engine::text_edit::limits::{PREVIEW_PDF_MAX_BYTES, VERIFY_CAP_MARGIN_BYTES}; +use crate::pdf_engine::text_edit::reasons::{EditProblem, TextWarningCode}; +use crate::pdf_engine::text_edit::rewrite::{plan_page, EditVerdict, PagePlan, TextEditIn}; +use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel}; +use crate::pdf_engine::text_edit::snapshot::{check_page_map, read_verification_snapshot}; +use crate::pdf_engine::validate_output::ContentDigest; +use std::collections::HashMap; +use std::ffi::OsString; +use std::io::Read; +use std::path::{Path, PathBuf}; +use std::sync::atomic::{AtomicUsize, Ordering}; +use std::sync::Arc; + +pub struct PreviewResult { + /// The patched one-page PDF, or `None` (a failed verdict, a page problem, or unavailable). + pub pdf: Option>, + pub verdicts: Vec, + pub page_problem: Option, + pub warnings: Vec, +} + +static NONCE: AtomicUsize = AtomicUsize::new(0); + +fn nonce() -> String { + let n = NONCE.fetch_add(1, Ordering::SeqCst); + let t = std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .map_or(0, |d| d.subsec_nanos()); + format!("{}-{n}-{t:x}", std::process::id()) +} + +fn io(context: &str) -> impl Fn(std::io::Error) -> AppError + '_ { + move |e| AppError::io(context, e) +} + +/// Moves `tmp` to `dest` unless `dest` already exists (then `tmp` is removed and `dest` used). +fn publish_once(tmp: &Path, dest: &Path) -> Result<(), AppError> { + if dest.is_file() { + let _ = std::fs::remove_file(tmp); + return Ok(()); + } + match std::fs::rename(tmp, dest) { + Ok(()) => Ok(()), + Err(_) if dest.is_file() => { + let _ = std::fs::remove_file(tmp); + Ok(()) + } + Err(e) => { + let _ = std::fs::remove_file(tmp); + Err(AppError::io("OffPDF could not write a preview file.", e)) + } + } +} + +/// Writes `bytes` to `dest` once (temp name + rename). +pub(crate) fn write_once(dest: &Path, bytes: &[u8]) -> Result<(), AppError> { + if dest.is_file() { + return Ok(()); + } + let tmp = dest.with_extension(format!("{}.tmp", nonce())); + if let Err(e) = std::fs::write(&tmp, bytes) { + // A failed write can leave a partial temp file behind. + let _ = std::fs::remove_file(&tmp); + return Err(AppError::io("OffPDF could not write a preview file.", e)); + } + publish_once(&tmp, dest) +} + +/// `p.pdf`: `qpdf --empty --remove-unreferenced-resources=no --pages source.pdf -- out` +/// (shared inherited resources stay whole, D30), written once. +fn extract_page( + engines: &Engines, + source: &Path, + page_1: u32, + dest: &Path, + opts: &RunOpts<'_>, +) -> Result<(), AppError> { + if dest.is_file() { + return Ok(()); + } + let tmp = dest.with_extension(format!("{}.tmp", nonce())); + let args = [ + OsString::from("--empty"), + OsString::from("--remove-unreferenced-resources=no"), + OsString::from("--pages"), + source.as_os_str().to_os_string(), + OsString::from(page_1.to_string()), + OsString::from("--"), + tmp.as_os_str().to_os_string(), + ]; + let out = match run_tool(&engines.qpdf, &args, false, opts) { + Ok(out) => out, + Err(e) => { + // A cancelled or timed-out qpdf can leave a partial output behind. + let _ = std::fs::remove_file(&tmp); + return Err(e); + } + }; + if out.code != 0 && out.code != 3 { + let _ = std::fs::remove_file(&tmp); + return Err(AppError::engine_failed(format!( + "qpdf page extraction exited with code {}: {}", + out.code, + out.stderr.trim() + ))); + } + publish_once(&tmp, dest) +} + +fn verdict_warnings(page_index: u32, verdicts: &[EditVerdict]) -> Vec { + verdicts + .iter() + .flat_map(|v| { + v.warnings.iter().map(move |code| TextWarning { + page_index, + run_id: v.run_id.clone(), + code: *code, + detail: None, + }) + }) + .collect() +} + +fn unavailable(page_index: u32, detail: &str) -> TextWarning { + TextWarning { + page_index, + run_id: String::new(), + code: TextWarningCode::PreviewUnavailable, + detail: Some(detail.to_string()), + } +} + +/// An error of the verification path as a page problem (the editor keeps its last good page). +fn page_problem_of(e: &AppError) -> Option { + let code = match e.code.as_str() { + "PEN_DRIFT" => crate::pdf_engine::text_edit::reasons::EditProblemCode::PenDrift, + "STATE_CHANGED" => crate::pdf_engine::text_edit::reasons::EditProblemCode::StateChanged, + "EDIT_VERIFY_FAILED" => { + crate::pdf_engine::text_edit::reasons::EditProblemCode::EditVerifyFailed + } + _ => return None, + }; + Some(EditProblem::new(code, e.details.clone())) +} + +/// What `same_page` compares of the source page: its parts (by digest, in order) and its page +/// fonts (name, content hash). Kept instead of the source model, which the preview releases once +/// it has planned (review-final MEDIUM-3). +struct SourcePage { + page_index: u32, + parts: Vec, + fonts: Vec<(Vec, u64)>, +} + +impl SourcePage { + fn of(model: &PageModel) -> SourcePage { + SourcePage { + page_index: model.page_index, + parts: model.content.parts.iter().map(|p| p.digest).collect(), + fonts: model + .walk + .page_fonts + .iter() + .map(|(n, m)| (n.clone(), m.content_hash)) + .collect(), + } + } +} + +/// Whether `p.pdf`'s page is the source page: the same parts in order (by digest) and every +/// page font name resolving to the same content hash (§B.17 step 3). +fn same_page(source: &SourcePage, extracted: &PageModel) -> bool { + let fonts: HashMap<&[u8], u64> = extracted + .walk + .page_fonts + .iter() + .map(|(n, m)| (n.as_slice(), m.content_hash)) + .collect(); + let parts: Vec = extracted.content.parts.iter().map(|p| p.digest).collect(); + extracted.page_reason.is_none() + && source.parts == parts + && source + .fonts + .iter() + .all(|(n, h)| fonts.get(n.as_slice()) == Some(h)) +} + +/// Reads a file of at most `cap` bytes. +fn read_capped(path: &Path, cap: u64) -> Result>, AppError> { + let file = std::fs::File::open(path).map_err(io("OffPDF could not read the preview."))?; + let mut data = Vec::new(); + file.take(cap.saturating_add(1)) + .read_to_end(&mut data) + .map_err(io("OffPDF could not read the preview."))?; + Ok((data.len() as u64 <= cap).then_some(data)) +} + +/// §B.17: plan, patch the one-page extraction with qpdf, prove it with Phase A, return its bytes. +/// A preview interrupted by a cancel is `CANCELLED` (never a page problem). The source `model` is +/// released once the edit is planned: when the caller holds no other reference (the service hands +/// over a large model, `TextEditCache::page_to_release`), it is freed before the extracted page's +/// model is built (review-final MEDIUM-3). +pub fn preview_page( + ctx: &SnapshotContext, + model: Arc, + edits: &[TextEditIn], + cache_dir: &Path, + engines: &Engines, + source_benign: &[String], + opts: &RunOpts<'_>, +) -> Result { + let page_index = model.page_index; + let outcome = plan_page(ctx, &model, edits)?; + let source_page = SourcePage::of(&model); + drop(model); + let mut warnings = verdict_warnings(page_index, &outcome.verdicts); + let Some(plan) = outcome.plan else { + return Ok(PreviewResult { + pdf: None, + verdicts: outcome.verdicts, + page_problem: None, + warnings, + }); + }; + std::fs::create_dir_all(cache_dir).map_err(io("OffPDF could not create a preview folder."))?; + let source = cache_dir.join("source.pdf"); + write_once(&source, &ctx.snap.bytes)?; + let page_1 = page_index.saturating_add(1); + let extracted = cache_dir.join(format!("p{page_1}.pdf")); + extract_page(engines, &source, page_1, &extracted, opts)?; + let work = cache_dir.join(format!("preview-{}", nonce())); + std::fs::create_dir_all(&work).map_err(io("OffPDF could not create a preview folder."))?; + let result = patch_and_prove( + ctx, + &source_page, + &plan, + &extracted, + &work, + engines, + source_benign, + opts, + ); + let _ = std::fs::remove_dir_all(&work); + if is_cancelled(opts) { + return Err(AppError::cancelled()); + } + let (pdf, page_problem, mut more) = result?; + warnings.append(&mut more); + Ok(PreviewResult { + pdf, + verdicts: outcome.verdicts, + page_problem, + warnings, + }) +} + +type Patched = (Option>, Option, Vec); + +/// Steps 3–6 inside the nonce directory. +#[allow(clippy::too_many_arguments)] +fn patch_and_prove( + ctx: &SnapshotContext, + source_page: &SourcePage, + plan: &PagePlan, + extracted: &Path, + work: &Path, + engines: &Engines, + source_benign: &[String], + opts: &RunOpts<'_>, +) -> Result { + let page_index = source_page.page_index; + let cap = ctx + .snap + .fingerprint + .len + .saturating_mul(2) + .saturating_add(VERIFY_CAP_MARGIN_BYTES); + let gone = |detail: &str| Ok((None, None, vec![unavailable(page_index, detail)])); + let Ok(snap) = read_verification_snapshot(extracted, cap) else { + return gone("the page could not be extracted"); + }; + let Ok(pages) = qpdf_page_map(engines, extracted, opts) else { + return gone("the extracted page could not be read"); + }; + if check_page_map(&snap, &pages).is_err() { + return gone("the extracted page reads differently"); + } + let p_ctx = SnapshotContext::new(snap); + let p_model = build_page_model(&p_ctx, 0, cancel_flag(opts))?; + if !same_page(source_page, &p_model) { + return gone("the extracted page differs from the source page"); + } + let updates = match updates_for_plan(&p_model.content, plan) { + Ok(u) => u, + Err(e) => { + return match page_problem_of(&e) { + Some(p) => Ok((None, Some(p), Vec::new())), + None => Err(e), + } + } + }; + let update = work.join("update.json"); + write_update_json(&updates, p_ctx.doc().max_id, &update)?; + let staged = work.join("preview.pdf"); + if let Err(e) = apply_update(engines, extracted, &update, &staged, source_benign, opts) { + return match page_problem_of(&e) { + Some(p) => Ok((None, Some(p), Vec::new())), + None => Err(e), + }; + } + let replaced: HashMap> = updates + .into_iter() + .map(|u| (u.object_id, u.decoded)) + .collect(); + let before = match graph_digest(p_ctx.doc(), &replaced, cancel_flag(opts)) { + Ok(d) => d, + Err(m) => { + let detail = format!("phase=A check=A2 path={} what={}", m.path, m.what); + return Ok(( + None, + Some(EditProblem::new( + crate::pdf_engine::text_edit::reasons::EditProblemCode::EditVerifyFailed, + Some(detail), + )), + Vec::new(), + )); + } + }; + let input = PhaseAInput { + before: &before, + before_page_count: 1, + staged: &staged, + staged_cap: cap, + pages: vec![EditedPageInput { + model: (&p_model).into(), + plan, + input_page_index: 0, + input_render: PopplerRef { + pdf: extracted.to_path_buf(), + page_1: 1, + }, + }], + source_benign, + }; + let report = match verify_edited_copy(&input, engines, work, opts) { + Ok(r) => r, + Err(e) => { + return match page_problem_of(&e) { + Some(p) => Ok((None, Some(p), Vec::new())), + None => Err(e), + } + } + }; + let mut warnings: Vec = report + .warnings + .into_iter() + .map(|w| TextWarning { page_index, ..w }) + .collect(); + match read_capped(&staged, PREVIEW_PDF_MAX_BYTES)? { + Some(bytes) => Ok((Some(bytes), None, warnings)), + None => { + warnings.push(unavailable(page_index, "the preview is too large")); + Ok((None, None, warnings)) + } + } +} + +/// `/textedit//`. +pub fn cache_dir_for(temp_root: &Path, fingerprint: &str) -> PathBuf { + temp_root.join("textedit").join(fingerprint) +} diff --git a/src-tauri/src/pdf_engine/text_edit/reasons.rs b/src-tauri/src/pdf_engine/text_edit/reasons.rs new file mode 100644 index 0000000..22d51fa --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/reasons.rs @@ -0,0 +1,669 @@ +//! Reason codes, edit problems, warnings and every `AppError` the text-edit path builds +//! (SPEC §A.10, §B.3, copy in §D.11.3 and §D.11.6). The canonical code list is +//! `src/lib/editor/text-reasons.json`, cross-locked by `reasons_json_matches_enums`. + +use crate::error::AppError; +use serde::{Deserialize, Serialize}; + +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, PartialOrd, Ord, Serialize, Deserialize)] +#[serde(rename_all = "SCREAMING_SNAKE_CASE")] +pub enum TextReason { + // run level, in priority order (A.10) + NestedForm, + SplitContent, + SharedContent, + InlineImage, + InvisibleText, + TextClipMode, + ZeroSize, + Vertical, + MirroredText, + RotatedText, + SkewedText, + Clipped, + OptionalContent, + SoftMask, + Pattern, + ActualText, + MissingFont, + Type3, + FontUnsupported, + FontNotEmbedded, + FontProgramUnsupported, + FontProgramUnreadable, + UnsupportedEncoding, + MissingWidths, + NoTounicode, + AmbiguousUnicode, + RightToLeft, + ComplexScript, + DuplicateText, + PerGlyphText, + NoWritableGlyphs, + // page level + MalformedContent, + UnsupportedFilter, + PageTooComplex, + Geometry, + // image occurrences (classifier) + MaskedImage, + SharedXobject, + TransformedImage, +} + +use TextReason as R; + +impl TextReason { + /// Page-level codes: every run on the page is refused with it. + pub const PAGE: &'static [TextReason] = &[ + R::MalformedContent, + R::UnsupportedFilter, + R::PageTooComplex, + R::Geometry, + ]; + /// Exact SCREAMING_SNAKE spelling, e.g. "TYPE3", "NO_TOUNICODE", "SHARED_XOBJECT". + pub fn as_str(self) -> &'static str { + match self { + R::NestedForm => "NESTED_FORM", + R::SplitContent => "SPLIT_CONTENT", + R::SharedContent => "SHARED_CONTENT", + R::InlineImage => "INLINE_IMAGE", + R::InvisibleText => "INVISIBLE_TEXT", + R::TextClipMode => "TEXT_CLIP_MODE", + R::ZeroSize => "ZERO_SIZE", + R::Vertical => "VERTICAL", + R::MirroredText => "MIRRORED_TEXT", + R::RotatedText => "ROTATED_TEXT", + R::SkewedText => "SKEWED_TEXT", + R::Clipped => "CLIPPED", + R::OptionalContent => "OPTIONAL_CONTENT", + R::SoftMask => "SOFT_MASK", + R::Pattern => "PATTERN", + R::ActualText => "ACTUAL_TEXT", + R::MissingFont => "MISSING_FONT", + R::Type3 => "TYPE3", + R::FontUnsupported => "FONT_UNSUPPORTED", + R::FontNotEmbedded => "FONT_NOT_EMBEDDED", + R::FontProgramUnsupported => "FONT_PROGRAM_UNSUPPORTED", + R::FontProgramUnreadable => "FONT_PROGRAM_UNREADABLE", + R::UnsupportedEncoding => "UNSUPPORTED_ENCODING", + R::MissingWidths => "MISSING_WIDTHS", + R::NoTounicode => "NO_TOUNICODE", + R::AmbiguousUnicode => "AMBIGUOUS_UNICODE", + R::RightToLeft => "RIGHT_TO_LEFT", + R::ComplexScript => "COMPLEX_SCRIPT", + R::DuplicateText => "DUPLICATE_TEXT", + R::PerGlyphText => "PER_GLYPH_TEXT", + R::NoWritableGlyphs => "NO_WRITABLE_GLYPHS", + R::MalformedContent => "MALFORMED_CONTENT", + R::UnsupportedFilter => "UNSUPPORTED_FILTER", + R::PageTooComplex => "PAGE_TOO_COMPLEX", + R::Geometry => "GEOMETRY", + R::MaskedImage => "MASKED_IMAGE", + R::SharedXobject => "SHARED_XOBJECT", + R::TransformedImage => "TRANSFORMED_IMAGE", + } + } + + /// True for the four page-level codes. + pub fn is_page_level(self) -> bool { + Self::PAGE.contains(&self) + } + + /// `(title, body)` of the reason copy (§D.11.3); codes without a row use the "(unknown)" row. + pub fn copy(self) -> (&'static str, &'static str) { + match self { + R::NestedForm => ("Part of a reused block", "This text is inside a block the document can reuse, so changing it here could change it in other places too."), + R::SplitContent => ("Stored in two pieces", "The instructions that draw this line are split across two parts of the page."), + R::SharedContent => ("Shared with other pages", "This part of the page is shared with other pages, so a change here would change them too."), + R::InlineImage => ("After an unreadable picture", "A picture stored inside this page can't be measured exactly, so text drawn after it can't be changed safely."), + R::InvisibleText => ("Hidden text", "This is hidden text, such as the searchable layer of a scan. Changing it would change nothing you can see."), + R::TextClipMode => ("Used as a shape", "This text is used as a shape that clips other content, so it can't be changed safely."), + R::ZeroSize => ("No visible size", "This text is drawn at zero size."), + R::Vertical => ("Vertical text", "This text runs top to bottom. Only lines that read across the page can be changed."), + R::MirroredText => ("Mirrored text", "This text is drawn mirrored, so new letters can't be placed the same way."), + R::RotatedText => ("Turned at an angle", "This line doesn't run straight across the page as shown. Only level lines can be changed."), + R::SkewedText => ("Tilted line", "This line's baseline is tilted by the page layout, so it can't be changed safely."), + R::Clipped => ("Partly cut off", "Part of this text is cut off by the page edge or a clipping area, so a change might not show."), + R::OptionalContent => ("On a layer", "This text is on a layer that is hidden or depends on viewer settings, so a change might not show."), + R::SoftMask => ("Masked text", "This text is drawn through a transparency mask, so a change could look different from what you type."), + R::Pattern => ("Pattern fill", "This text is filled with a pattern or gradient rather than a colour."), + R::ActualText => ("Separate reading copy", "The document keeps a separate copy of this text for search and screen readers. Changing only the visible letters would make the two disagree."), + R::MissingFont => ("Font missing", "The font this text uses is missing from the document."), + R::Type3 => ("Font made of drawings", "This text uses a font drawn from shapes, which OffPDF can't type with."), + R::FontUnsupported => ("Unusual font setup", "This font is set up in a way OffPDF can't change safely."), + R::FontNotEmbedded => ("Font not included", "This font isn't included in the PDF, so OffPDF can't check which letters it can draw."), + R::FontProgramUnsupported => ("Font type not supported yet", "This text uses a kind of embedded font that OffPDF can't check yet."), + R::FontProgramUnreadable => ("Font can't be read", "The font included in this PDF can't be read, so OffPDF can't check which letters it can draw."), + R::UnsupportedEncoding => ("Letter mapping not supported", "This font maps letters in a way OffPDF can't write yet."), + R::MissingWidths => ("Letter widths missing", "The document doesn't say how wide this font's letters are, so the rest of the line couldn't be kept in place."), + R::NoTounicode => ("Letters not identified", "The document doesn't say which letters this font draws, so OffPDF can't read or retype them."), + R::AmbiguousUnicode => ("Letters can't be confirmed", "Some letters in this line can't be read with certainty, so OffPDF can't show the current text reliably."), + R::RightToLeft => ("Right-to-left script", "The document has already ordered and shaped these letters, and changing them would undo that work."), + R::ComplexScript => ("Shaped script", "This script joins or reorders letters, and the document has already done that shaping. OffPDF can't redo it yet."), + R::DuplicateText => ("Drawn twice", "This text is drawn twice (for example as a shadow or to look bolder), so changing one copy would leave the other behind."), + R::PerGlyphText => ("Letters placed one by one", "This page places every letter on its own, so a line can't be changed as one piece."), + R::NoWritableGlyphs => ("No letters to type with", "The font this line uses has no letters OffPDF can confirm, so nothing can be typed with it."), + R::MalformedContent => ("Damaged page content", "OffPDF couldn't read this page's content completely, so none of its text can be changed."), + R::UnsupportedFilter => ("Unusual compression", "This page's content is stored or compressed in a way OffPDF can't check."), + R::PageTooComplex => ("Too complex to check", "This page has too much content to check safely."), + R::Geometry => ("Custom page unit", "This page uses a custom unit size or unreadable page boxes, which OffPDF doesn't edit yet."), + R::MaskedImage | R::SharedXobject | R::TransformedImage => { + ("Can't be changed safely", "OffPDF can't change this text safely.") + } + } + } + + /// Run codes (and image codes) → `TEXT_EDIT_REFUSED` naming the reason; page codes → an + /// error whose code is the page code itself. + pub fn to_app_error(self, page_number: Option) -> AppError { + let (title, body) = self.copy(); + if self.is_page_level() { + return AppError::new(self.as_str(), title, on_page_colon(page_number, body)) + .with_suggestion("Restore the original text on that page.") + .with_details(detail_line(page_number, self.as_str())); + } + AppError::new( + EditProblemCode::TextEditRefused.as_str(), + "This text can't be changed safely", + on_page_colon(page_number, body), + ) + .with_suggestion("Restore the original text for that line.") + .with_details(detail_line(page_number, self.as_str())) + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)] +#[serde(rename_all = "SCREAMING_SNAKE_CASE")] +pub enum EditProblemCode { + GlyphMissing, + SpaceNotWritable, + InvalidText, + TextTooLong, + TextOutsideVisibleArea, + FaceUnavailable, + StyleUnavailable, + EditConflict, + TextEditRefused, + Stale, + PenDrift, + StateChanged, + EditVerifyFailed, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)] +#[serde(rename_all = "SCREAMING_SNAKE_CASE")] +pub enum TextWarningCode { + NextTextOverlap, + EditNotVisible, + PreviewUnavailable, +} + +/// Which control a `STYLE_UNAVAILABLE` refers to. +#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize)] +#[serde(rename_all = "camelCase")] +pub enum StyleField { + Size, + Face, + Colour, +} + +/// A face of a font family (lives here because `EditProblem` needs it before `fonts/`). +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)] +#[serde(rename_all = "camelCase")] +pub enum Face { + Regular, + Bold, + Italic, + BoldItalic, +} + +impl Face { + fn adjective(self) -> &'static str { + match self { + Face::Regular => "regular", + Face::Bold => "bold", + Face::Italic => "italic", + Face::BoldItalic => "bold italic", + } + } +} + +#[derive(Debug, Clone, PartialEq)] +pub struct EditProblem { + pub code: EditProblemCode, + pub chars: Vec, // GLYPH_MISSING / FACE_UNAVAILABLE characters, typing order + pub reason: Option, // TEXT_EDIT_REFUSED: the run's reason + pub face: Option, // FACE_UNAVAILABLE: the requested face + pub field: Option, // STYLE_UNAVAILABLE: which control + pub detail: Option, // technical detail (check id, drift, byte offset) +} + +impl EditProblem { + /// A problem with only a code and an optional technical detail. + pub fn new(code: EditProblemCode, detail: Option) -> EditProblem { + EditProblem { + code, + chars: Vec::new(), + reason: None, + face: None, + field: None, + detail, + } + } +} + +pub struct ProblemCtx<'a> { + pub page_number: Option, + pub file_name: Option<&'a str>, + pub face: Option, + pub reason: Option, +} + +use EditProblemCode as P; + +impl EditProblemCode { + pub fn as_str(self) -> &'static str { + match self { + P::GlyphMissing => "GLYPH_MISSING", + P::SpaceNotWritable => "SPACE_NOT_WRITABLE", + P::InvalidText => "INVALID_TEXT", + P::TextTooLong => "TEXT_TOO_LONG", + P::TextOutsideVisibleArea => "TEXT_OUTSIDE_VISIBLE_AREA", + P::FaceUnavailable => "FACE_UNAVAILABLE", + P::StyleUnavailable => "STYLE_UNAVAILABLE", + P::EditConflict => "EDIT_CONFLICT", + P::TextEditRefused => "TEXT_EDIT_REFUSED", + P::Stale => "STALE", + P::PenDrift => "PEN_DRIFT", + P::StateChanged => "STATE_CHANGED", + P::EditVerifyFailed => "EDIT_VERIFY_FAILED", + } + } + + /// The Save-time `AppError` for this problem (copy in §D.11.6). + pub fn to_app_error(self, p: &EditProblem, ctx: &ProblemCtx<'_>) -> AppError { + let n = ctx.page_number; + let (title, message, suggestion): (&str, String, &str) = match self { + P::GlyphMissing => ( + "The font can't draw some letters", + format!("{}, the document's font can't draw: {}", on_page(n), char_list(&p.chars)), + "Use other characters, or add a new text box with Add text.", + ), + P::SpaceNotWritable => ( + "A space can't go there", + format!("{}, spaces in the changed line are gaps between letters, so a space can only go between two letters, one at a time.", on_page(n)), + "Remove the extra space.", + ), + P::InvalidText => ( + "These characters can't be added", + "Line breaks, tabs and control characters can't be added. Each box is a single line.".to_string(), + "Remove them and try again.", + ), + P::TextTooLong => ( + "Text is too long", + "A changed line can have at most 1,000 characters.".to_string(), + "Shorten the line.", + ), + P::TextOutsideVisibleArea => ( + "Text runs off the page", + format!("{}, a changed line would run past the edge of the visible page.", on_page(n)), + "Shorten the line.", + ), + P::FaceUnavailable => ( + "That style isn't available", + on_page_colon(n, &face_sentence(p.face.or(ctx.face), &p.chars)), + "Keep the current style.", + ), + P::StyleUnavailable => ( + "That change isn't available", + on_page_colon(n, style_sentence(p.field)), + "Keep the current setting.", + ), + P::EditConflict => ( + "Two changes overlap", + match n { + Some(n) => format!("Two text changes on page {n} affect the same line."), + None => "Two text changes on this page affect the same line.".to_string(), + }, + "Undo one of them and try again.", + ), + P::TextEditRefused => { + let reason = p.reason.or(ctx.reason); + let body = reason.map(|r| r.copy().1).unwrap_or("OffPDF can't change this text safely."); + ( + "This text can't be changed safely", + on_page_colon(n, body), + "Restore the original text for that line.", + ) + } + P::Stale => { + let e = stale(ctx.file_name.unwrap_or("the PDF")); + return with_problem_details(e, p, n); + } + P::PenDrift => ( + "The change would move other text", + format!("Saving would shift text you didn't change on {}, so nothing was saved.", page_ref(n)), + "Try a shorter change, or undo it. The original file was not changed.", + ), + P::StateChanged => ( + "The change would restyle other content", + format!("Saving would change how other content on {} is drawn, so nothing was saved.", page_ref(n)), + "Undo the last change on that page and try again. The original file was not changed.", + ), + P::EditVerifyFailed => ( + "The change could not be verified", + format!("OffPDF checks every change before saving. {}, a changed line couldn't be confirmed to read back exactly as typed with nothing else altered, so nothing was saved.", on_page(n)), + "Undo that change and try again. The original file was not changed.", + ), + }; + let e = AppError::new(self.as_str(), title, message).with_suggestion(suggestion); + with_problem_details(e, p, n) + } +} + +fn with_problem_details(e: AppError, p: &EditProblem, page: Option) -> AppError { + let mut parts = vec![format!("code: {}", p.code.as_str())]; + if let Some(n) = page { + parts.push(format!("page: {n}")); + } + if let Some(r) = p.reason { + parts.push(format!("reason: {}", r.as_str())); + } + if let Some(d) = &p.detail { + parts.push(d.clone()); + } + e.with_details(parts.join("; ")) +} + +fn on_page(n: Option) -> String { + match n { + Some(n) => format!("On page {n}"), + None => "On this page".to_string(), + } +} + +fn on_page_colon(n: Option, body: &str) -> String { + format!("{}: {body}", on_page(n)) +} + +fn page_ref(n: Option) -> String { + match n { + Some(n) => format!("page {n}"), + None => "this page".to_string(), + } +} + +fn detail_line(page: Option, code: &str) -> String { + match page { + Some(n) => format!("reason: {code}; page: {n}"), + None => format!("reason: {code}"), + } +} + +/// Characters for a message: typing order, a space shows as "space". +fn char_list(chars: &[char]) -> String { + let items: Vec = chars + .iter() + .map(|c| { + if *c == ' ' { + "space".to_string() + } else { + c.to_string() + } + }) + .collect(); + items.join(", ") +} + +fn face_sentence(face: Option, chars: &[char]) -> String { + let face = face.unwrap_or(Face::Regular); + if chars.is_empty() { + format!( + "This page has no {} version of this font.", + face.adjective() + ) + } else { + format!( + "The {} version of this font can't draw: {}", + face.adjective(), + char_list(chars) + ) + } +} + +fn style_sentence(field: Option) -> &'static str { + match field { + Some(StyleField::Size) | None => "This line's size is set in a way OffPDF can't change.", + Some(StyleField::Face) => { + "This line's font is set in a way OffPDF can't change, so bold and italic aren't available." + } + Some(StyleField::Colour) => { + "This text is drawn with an outline, so its colour can't be changed here." + } + } +} + +// ---- File-level constructors (copy in §D.11.6) ------------------------------------------ + +pub fn file_too_large() -> AppError { + AppError::new( + "FILE_TOO_LARGE", + "This PDF is too large to edit text in", + "Files over 400 MB are not read for text editing.", + ) + .with_suggestion("Split it with Split PDF and edit the part you need.") +} + +pub fn file_too_complex(detail: &str) -> AppError { + AppError::new( + "FILE_TOO_COMPLEX", + "This PDF is too complex to check", + "Some of its internal data is too large or too deeply nested to check safely.", + ) + .with_details(detail.to_string()) +} + +pub fn encrypted() -> AppError { + AppError::new( + "ENCRYPTED", + "This PDF is password-protected", + "OffPDF can't change text in a protected PDF.", + ) + .with_suggestion("Remove the password with Unlock PDF, then edit the unlocked copy.") +} + +pub fn signed() -> AppError { + AppError::new( + "SIGNED", + "This PDF is digitally signed", + "Changing its text would break the signature.", + ) + .with_suggestion("Ask the sender for an unsigned copy if it needs changes.") +} + +pub fn unsupported_xfa() -> AppError { + AppError::new( + "UNSUPPORTED_XFA", + "This PDF is a dynamic form", + "Its pages are generated by the PDF reader, so changes to page text might not show.", + ) + .with_suggestion("Fill it in a reader that supports dynamic forms.") +} + +pub fn malformed_content(detail: &str) -> AppError { + AppError::new( + "MALFORMED_CONTENT", + "Part of this PDF can't be read", + "OffPDF couldn't read this PDF's structure completely.", + ) + .with_suggestion("Run it through Repair PDF, then try again.") + .with_details(detail.to_string()) +} + +pub fn pdf_needs_repair(warnings: &[String]) -> AppError { + AppError::new( + "PDF_NEEDS_REPAIR", + "This PDF needs repair first", + "Its internal structure has errors, so OffPDF won't change its text.", + ) + .with_suggestion("Run it through Repair PDF, then edit the repaired copy.") + .with_details(warnings.join("\n")) +} + +pub fn stale(file_name: &str) -> AppError { + AppError::new( + "STALE", + "The PDF changed on disk", + format!("\u{201c}{file_name}\u{201d} was changed after you started editing it, so your text changes no longer match it."), + ) + .with_suggestion("Remove these text changes and make them again. The original file was not changed.") +} + +pub fn verifier_missing(tool: &str) -> AppError { + AppError::new( + "VERIFIER_MISSING", + "A checking component is missing", + "OffPDF needs its bundled Poppler tools to verify text changes.", + ) + .with_suggestion("Reinstall OffPDF, then try again.") + .with_details(format!("{tool} could not be started")) +} + +pub fn qpdf_too_old(version: &str) -> AppError { + AppError::new( + "ENGINE_MISSING", + "The PDF engine is too old", + "OffPDF needs a newer version of its bundled qpdf engine to change text.", + ) + .with_suggestion("Reinstall OffPDF, then try again.") + .with_details(format!("qpdf 11 or newer is required; found: {version}")) +} + +pub fn source_edit_gate_failed(detail: &str) -> AppError { + AppError::new( + "SOURCE_EDIT_GATE_FAILED", + "The edited PDF did not pass the text-change check", + "The saved file did not contain the text changes exactly as they were checked, so it was not published.", + ) + .with_suggestion("Try saving again. The original file was not changed.") + .with_details(detail.to_string()) +} + +/// `INVALID_PDF` for a file whose cross-reference data can't be read (snapshot preflight). +pub fn invalid_xref(detail: &str) -> AppError { + AppError::new( + "INVALID_PDF", + "The selected file is not a valid PDF", + "OffPDF could not read this PDF's cross-reference data.", + ) + .with_suggestion("Make sure the file is a real PDF and is not corrupted.") + .with_details(detail.to_string()) +} + +pub const ORIGINAL_UNCHANGED: &str = "The original file was not changed."; + +/// Every error leaving the text-edit Save path passes through this (export.rs wraps the results of +/// `prepare_text_sources` and `verify_final`): appends ORIGINAL_UNCHANGED to the suggestion (or sets it) +/// unless the suggestion already ends with it. Open/inspect/preview errors are not wrapped. +pub fn save_failure(e: AppError) -> AppError { + let suggestion = match e.suggestion.as_deref().map(str::trim_end) { + None | Some("") => ORIGINAL_UNCHANGED.to_string(), + Some(s) if s.ends_with(ORIGINAL_UNCHANGED) => s.to_string(), + Some(s) => format!("{s} {ORIGINAL_UNCHANGED}"), + }; + AppError { + suggestion: Some(suggestion), + ..e + } +} + +// Test-only tables and spellings: production serialises these codes through serde and checks +// reasons one at a time (`cargo check` dead-code gate, review T5 H2). +#[cfg(test)] +impl TextReason { + /// The 31 run codes in A.10 priority order. + pub const RUN_PRIORITY: &'static [TextReason] = &[ + R::NestedForm, + R::SplitContent, + R::SharedContent, + R::InlineImage, + R::InvisibleText, + R::TextClipMode, + R::ZeroSize, + R::Vertical, + R::MirroredText, + R::RotatedText, + R::SkewedText, + R::Clipped, + R::OptionalContent, + R::SoftMask, + R::Pattern, + R::ActualText, + R::MissingFont, + R::Type3, + R::FontUnsupported, + R::FontNotEmbedded, + R::FontProgramUnsupported, + R::FontProgramUnreadable, + R::UnsupportedEncoding, + R::MissingWidths, + R::NoTounicode, + R::AmbiguousUnicode, + R::RightToLeft, + R::ComplexScript, + R::DuplicateText, + R::PerGlyphText, + R::NoWritableGlyphs, + ]; + /// Image occurrences (classifier only), A.10 order. + pub const IMAGE_PRIORITY: &'static [TextReason] = &[ + R::InlineImage, + R::NestedForm, + R::Clipped, + R::Pattern, + R::MaskedImage, + R::SharedXobject, + R::TransformedImage, + R::Geometry, + ]; +} + +#[cfg(test)] +impl EditProblemCode { + pub const ALL: &'static [EditProblemCode] = &[ + P::GlyphMissing, + P::SpaceNotWritable, + P::InvalidText, + P::TextTooLong, + P::TextOutsideVisibleArea, + P::FaceUnavailable, + P::StyleUnavailable, + P::EditConflict, + P::TextEditRefused, + P::Stale, + P::PenDrift, + P::StateChanged, + P::EditVerifyFailed, + ]; +} + +#[cfg(test)] +impl TextWarningCode { + pub const ALL: &'static [TextWarningCode] = &[ + TextWarningCode::NextTextOverlap, + TextWarningCode::EditNotVisible, + TextWarningCode::PreviewUnavailable, + ]; + + pub fn as_str(self) -> &'static str { + match self { + TextWarningCode::NextTextOverlap => "NEXT_TEXT_OVERLAP", + TextWarningCode::EditNotVisible => "EDIT_NOT_VISIBLE", + TextWarningCode::PreviewUnavailable => "PREVIEW_UNAVAILABLE", + } + } +} + +#[cfg(test)] +mod tests; diff --git a/src-tauri/src/pdf_engine/text_edit/reasons/tests.rs b/src-tauri/src/pdf_engine/text_edit/reasons/tests.rs new file mode 100644 index 0000000..6e4244f --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/reasons/tests.rs @@ -0,0 +1,243 @@ +//! `reasons_json_matches_enums`, `save_failures_say_original_unchanged` and the copy deck. + +use super::*; + +const JSON: &str = include_str!("../../../../../src/lib/editor/text-reasons.json"); + +fn list(v: &serde_json::Value, key: &str) -> Vec { + v[key] + .as_array() + .unwrap() + .iter() + .map(|s| s.as_str().unwrap().to_string()) + .collect() +} + +#[test] +fn reasons_json_matches_enums() { + let v: serde_json::Value = serde_json::from_str(JSON).unwrap(); + assert_eq!(v["version"], 1, "reasons_json_matches_enums: version"); + let strs = |xs: &[TextReason]| { + xs.iter() + .map(|r| r.as_str().to_string()) + .collect::>() + }; + assert_eq!( + list(&v, "run"), + strs(TextReason::RUN_PRIORITY), + "reasons_json_matches_enums: run" + ); + assert_eq!( + list(&v, "page"), + strs(TextReason::PAGE), + "reasons_json_matches_enums: page" + ); + assert_eq!( + list(&v, "image"), + strs(TextReason::IMAGE_PRIORITY), + "reasons_json_matches_enums: image" + ); + let problems: Vec = EditProblemCode::ALL + .iter() + .map(|p| p.as_str().to_string()) + .collect(); + assert_eq!( + list(&v, "problem"), + problems, + "reasons_json_matches_enums: problem" + ); + let warnings: Vec = TextWarningCode::ALL + .iter() + .map(|w| w.as_str().to_string()) + .collect(); + assert_eq!( + list(&v, "warning"), + warnings, + "reasons_json_matches_enums: warning" + ); + // serde spelling equals as_str for every enum value + let all: Vec = TextReason::RUN_PRIORITY + .iter() + .chain(TextReason::PAGE) + .chain(TextReason::IMAGE_PRIORITY) + .copied() + .collect(); + for r in all { + assert_eq!(serde_json::to_value(r).unwrap(), r.as_str()); + assert_eq!( + serde_json::from_value::(r.as_str().into()).unwrap(), + r + ); + } + for p in EditProblemCode::ALL { + assert_eq!(serde_json::to_value(p).unwrap(), p.as_str()); + } + for w in TextWarningCode::ALL { + assert_eq!(serde_json::to_value(w).unwrap(), w.as_str()); + } + // file-level codes built here appear in "file"; Save-only codes in "save" + let file = list(&v, "file"); + for e in file_level_errors() { + if e.code != "SOURCE_EDIT_GATE_FAILED" { + assert!( + file.contains(&e.code), + "reasons_json_matches_enums: file code {}", + e.code + ); + } + } + assert!(list(&v, "save").contains(&"SOURCE_EDIT_GATE_FAILED".to_string())); + // 38 distinct reason values, none left out of the lists + let mut distinct: Vec<&str> = TextReason::RUN_PRIORITY + .iter() + .chain(TextReason::PAGE) + .chain(TextReason::IMAGE_PRIORITY) + .map(|r| r.as_str()) + .collect(); + distinct.sort_unstable(); + distinct.dedup(); + assert_eq!(distinct.len(), 38); + assert_eq!( + serde_json::to_value(Face::BoldItalic).unwrap(), + "boldItalic" + ); + assert_eq!(serde_json::to_value(StyleField::Colour).unwrap(), "colour"); +} + +fn file_level_errors() -> Vec { + vec![ + file_too_large(), + file_too_complex("detail"), + encrypted(), + signed(), + unsupported_xfa(), + malformed_content("detail"), + pdf_needs_repair(&["w1".to_string(), "w2".to_string()]), + stale("a.pdf"), + verifier_missing("pdftotext"), + qpdf_too_old("10.6.3"), + source_edit_gate_failed("A2"), + invalid_xref("detail"), + ] +} + +fn every_save_error() -> Vec { + let mut out = file_level_errors(); + let ctx = ProblemCtx { + page_number: Some(3), + file_name: Some("a.pdf"), + face: Some(Face::Bold), + reason: None, + }; + let ctx_none = ProblemCtx { + page_number: None, + file_name: None, + face: None, + reason: None, + }; + for code in EditProblemCode::ALL { + let mut p = EditProblem::new(*code, Some("detail".into())); + p.chars = vec!['ğ', ' ']; + p.reason = Some(TextReason::Type3); + p.field = Some(StyleField::Colour); + out.push(code.to_app_error(&p, &ctx)); + out.push(code.to_app_error(&EditProblem::new(*code, None), &ctx_none)); + } + for r in TextReason::RUN_PRIORITY + .iter() + .chain(TextReason::PAGE) + .chain(TextReason::IMAGE_PRIORITY) + { + out.push(r.to_app_error(Some(2))); + out.push(r.to_app_error(None)); + } + out +} + +#[test] +fn save_failures_say_original_unchanged() { + for e in every_save_error() { + let code = e.code.clone(); + let once = save_failure(e); + let s = once.suggestion.clone().unwrap(); + assert!( + s.ends_with(ORIGINAL_UNCHANGED), + "save_failures_say_original_unchanged: {code}: {s}" + ); + assert_eq!( + s.matches(ORIGINAL_UNCHANGED).count(), + 1, + "save_failures_say_original_unchanged: {code}: {s}" + ); + let twice = save_failure(once.clone()); + assert_eq!(twice.suggestion, once.suggestion, "idempotent: {code}"); + assert!(!once.title.is_empty() && !once.message.is_empty(), "{code}"); + } + // a suggestion-less error gets the sentence as the whole suggestion + assert_eq!( + save_failure(file_too_complex("x")).suggestion.as_deref(), + Some(ORIGINAL_UNCHANGED) + ); + // STALE is already terminated and stays unchanged + assert_eq!( + save_failure(stale("a.pdf")).suggestion, + stale("a.pdf").suggestion + ); +} + +#[test] +fn problem_copy_matches_the_deck() { + let ctx = ProblemCtx { + page_number: Some(4), + file_name: Some("r.pdf"), + face: None, + reason: None, + }; + let mut p = EditProblem::new(EditProblemCode::GlyphMissing, None); + p.chars = vec!['ğ', ' ', 'Y']; + let e = EditProblemCode::GlyphMissing.to_app_error(&p, &ctx); + assert_eq!(e.code, "GLYPH_MISSING"); + assert_eq!( + e.message, + "On page 4, the document's font can't draw: ğ, space, Y" + ); + let mut f = EditProblem::new(EditProblemCode::FaceUnavailable, None); + f.face = Some(Face::Bold); + assert_eq!( + EditProblemCode::FaceUnavailable + .to_app_error(&f, &ctx) + .message, + "On page 4: This page has no bold version of this font." + ); + f.face = Some(Face::Italic); + f.chars = vec!['ş']; + assert_eq!( + EditProblemCode::FaceUnavailable + .to_app_error(&f, &ctx) + .message, + "On page 4: The italic version of this font can't draw: ş" + ); + let mut r = EditProblem::new(EditProblemCode::TextEditRefused, Some("run 7".into())); + r.reason = Some(TextReason::SharedContent); + let e = EditProblemCode::TextEditRefused.to_app_error(&r, &ctx); + assert_eq!(e.message, "On page 4: This part of the page is shared with other pages, so a change here would change them too."); + assert!(e.details.unwrap().contains("reason: SHARED_CONTENT")); + let s = + EditProblemCode::Stale.to_app_error(&EditProblem::new(EditProblemCode::Stale, None), &ctx); + assert_eq!(s.code, "STALE"); + assert!(s.message.starts_with("\u{201c}r.pdf\u{201d} was changed")); + for (field, text) in [ + (StyleField::Size, "This line's size is set in a way OffPDF can't change."), + (StyleField::Face, "This line's font is set in a way OffPDF can't change, so bold and italic aren't available."), + (StyleField::Colour, "This text is drawn with an outline, so its colour can't be changed here."), + ] { + let mut p = EditProblem::new(EditProblemCode::StyleUnavailable, None); + p.field = Some(field); + assert_eq!(EditProblemCode::StyleUnavailable.to_app_error(&p, &ctx).message, format!("On page 4: {text}")); + } + let page = TextReason::Geometry.to_app_error(Some(9)); + assert_eq!(page.code, "GEOMETRY"); + let run = TextReason::Type3.to_app_error(Some(9)); + assert_eq!(run.code, "TEXT_EDIT_REFUSED"); + assert!(run.details.unwrap().contains("TYPE3")); +} diff --git a/src-tauri/src/pdf_engine/text_edit/rewrite.rs b/src-tauri/src/pdf_engine/text_edit/rewrite.rs new file mode 100644 index 0000000..3e9673e --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/rewrite.rs @@ -0,0 +1,434 @@ +//! The edit planner (SPEC §B.12): from a page model and the requested edits to exact replacement +//! bytes for the show ops of each edited run, the expected content parts, and the expectations +//! verification checks the re-walk against. A minimal diff keeps unchanged leading and trailing +//! glyphs with their codes and kerns; the pen after the run is compensated so every follower stays +//! where it was; style changes are written as `Tf`/`Tc`/`rg` and restored verbatim. +//! +//! Steps 1–4 (`rewrite/style.rs`): resolve, validate, normalise and check availability. Steps 5–8 +//! (`rewrite/diff.rs`): targets, minimal diff, encoding, the new unit list. Steps 9–10 +//! (`rewrite/layout.rs`): advances, column gaps, compensation, fit. Steps 11–14 +//! (`rewrite/bytes.rs`): bytes, grammar, splices, expected parts. Step 15: the self-check through +//! `verify::walk_and_verify` before any IO; a failure is an internal `EDIT_VERIFY_FAILED` and qpdf +//! is never invoked. + +mod bytes; +mod cff_scale; +mod diff; +mod layout; +mod masks; +mod style; + +pub(crate) use style::letter_spacing_pt; + +use crate::error::AppError; +use crate::pdf_engine::text_edit::content::PageContent; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::fonts::{Code, FontModel, TypingSurface}; +use crate::pdf_engine::text_edit::lexer::Span; +use crate::pdf_engine::text_edit::limits::EDITS_PER_PAGE_MAX; +use crate::pdf_engine::text_edit::reasons::{ + EditProblem, EditProblemCode, Face, ProblemCtx, TextWarningCode, +}; +use crate::pdf_engine::text_edit::runs::PageModel; +use crate::pdf_engine::text_edit::state::Paint; +use crate::pdf_engine::text_edit::verify; +use crate::pdf_engine::validate_output::ContentDigest; +use std::collections::HashMap; +use std::sync::Arc; + +pub(crate) use bytes::assemble_page_plan; + +// Edit input types are defined here; dto.rs re-exports them and edit_overlay.rs uses +// `text_edit::rewrite::SourceTextStyleIn`. + +/// Requested style, each field only when it differs from the run's current value. +#[derive(Debug, Clone, Default, PartialEq, serde::Deserialize)] +#[serde(rename_all = "camelCase")] +pub struct SourceTextStyleIn { + pub size_pt: Option, + pub face: Option, + pub fill: Option, + pub letter_spacing_pt: Option, +} + +impl SourceTextStyleIn { + pub fn is_empty(&self) -> bool { + self.size_pt.is_none() + && self.face.is_none() + && self.fill.is_none() + && self.letter_spacing_pt.is_none() + } +} + +#[derive(Debug, Clone, serde::Deserialize)] +#[serde(rename_all = "camelCase")] +pub struct TextEditIn { + pub run_id: String, + pub original_text: String, + pub text: String, + #[serde(default)] + pub style: SourceTextStyleIn, +} + +/// What the replacement sets. Numbers are the values read back from what is written. +#[derive(Clone)] +pub struct StyleTarget { + pub tfs: f64, + /// The `Tf` operand to write: the original token when the size is unchanged. + pub tfs_token: Vec, + pub tc: f64, + pub fill: Option<[f64; 3]>, + pub face_surface: Option, + pub size_changed: bool, + pub tc_changed: bool, + pub fill_changed: bool, + pub face_changed: bool, + /// The requested face (verification recomputes the allowed fonts from it). + pub face: Option, + /// The requested effective size in points (the original one when unchanged). + pub effective_size: f64, +} + +impl std::fmt::Debug for StyleTarget { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + f.debug_struct("StyleTarget") + .field("tfs", &self.tfs) + .field("tc", &self.tc) + .field("fill", &self.fill) + .field("size_changed", &self.size_changed) + .field("tc_changed", &self.tc_changed) + .field("fill_changed", &self.fill_changed) + .field("face_changed", &self.face_changed) + .field("face", &self.face) + .finish() + } +} + +impl Clone for TypingSurface { + fn clone(&self) -> Self { + TypingSurface { + fonts: self + .fonts + .iter() + .map(|(n, m)| (n.clone(), Arc::clone(m))) + .collect(), + } + } +} + +#[derive(Debug, Clone)] +pub enum NewUnit { + /// A unit of the run kept as it is (index in `TextRun::units`). + Kept(usize), + /// A newly encoded glyph (`font` indexes the typing surface used). + Code { font: usize, code: Code }, + /// A typed space written as a TJ number (§A.3.4 kern mode). + KernSpace(f64), +} + +#[derive(Debug, Clone, PartialEq)] +pub struct Splice { + pub part: usize, + pub local: Span, + pub joined: Span, + pub bytes: Vec, +} + +#[derive(Debug, Clone)] +pub struct ExpectedRun { + pub run_id: String, + /// The text the planned glyphs decode to (synthetic spaces as " "). + pub text: String, + /// Font resource (empty for an ExtGState font), font content hash and code of every glyph. + pub glyphs: Vec<(Vec, u64, Code)>, + pub prefix_glyphs: usize, + pub suffix_glyphs: usize, + /// Displacement of the first kept suffix glyph (`Δ·dir`). + pub shift_user: (f64, f64), + /// Suffix glyph index (in `glyphs`) from which the displacement is 0 (a column gap absorbed Δ). + pub unshifted_from: Option, + pub origin: (f64, f64), + pub primary_pen_after: (f64, f64), + pub member_pen_after: Vec<(f64, f64)>, + pub tfs: f64, + pub effective_size: f64, + pub tc: f64, + pub fill: Paint, + /// Per member: the number of show ops its replacement emits. + pub emitted_records: Vec, + // ---- additions (T4, see DEVIATIONS) ---- + /// The joined-buffer spans of the members, primary first: the id-free identity of the run. + pub member_spans: Vec, + /// Expected user-space origin of every glyph (the new layout, from the written numbers). + pub glyph_origins: Vec<(f64, f64)>, + /// The glyph was newly encoded (its code must be drawable, B3). + pub glyph_new: Vec, + /// A kept glyph's (member, glyph index in that member's record). + pub kept_from: Vec>, + /// User-space boxes covering every old and new glyph (A5 render masks): one per new glyph, + /// in `glyph_origins` order, then the old glyphs' (`new_glyph_masks`). + pub mask_boxes: Vec<[f64; 4]>, + /// User-space boxes of the old glyphs' own ink, as tight as known (the program's outline + /// box, else the advance; `rewrite/masks.rs` `ink_box`): A5 checks the pixels of the glyphs no + /// edit changes outside them and the new glyphs' masks (review-verify HIGH-A). + pub old_ink_boxes: Vec<[f64; 4]>, + /// The non-space glyph sequence changed (A5 `EDIT_NOT_VISIBLE`). + pub glyphs_changed: bool, + /// The text the user asked for, copied from the request (never derived from the planned + /// glyphs): A4 and A5 compare what the file reads back with it. + pub requested_text: String, +} + +impl ExpectedRun { + /// The new glyphs' masks: the first `glyph_origins.len()` of `mask_boxes`. + pub fn new_glyph_masks(&self) -> &[[f64; 4]] { + self.mask_boxes + .get(..self.glyph_origins.len()) + .unwrap_or(&self.mask_boxes) + } +} + +#[derive(Debug, Clone)] +pub struct RunPlan { + pub run_id: String, + pub target: StyleTarget, + /// One per member, primary first. + pub splices: Vec, + pub expected: ExpectedRun, + pub new_rect: [f64; 4], +} + +#[derive(Debug, Clone)] +pub struct EditVerdict { + pub run_id: String, + pub problem: Option, + pub delta_pt: f64, + pub new_rect: Option<[f64; 4]>, + pub caret_offsets: Option>, + pub warnings: Vec, +} + +impl EditVerdict { + fn failed(run_id: &str, problem: EditProblem) -> EditVerdict { + EditVerdict { + run_id: run_id.to_string(), + problem: Some(problem), + delta_pt: 0.0, + new_rect: None, + caret_offsets: None, + warnings: Vec::new(), + } + } +} + +#[derive(Debug, Clone)] +pub struct PagePlan { + pub page_index: u32, + pub runs: Vec, + /// Sorted by joined start. + pub splices: Vec, + pub expected_parts: Vec>, + pub edited_parts: Vec, + pub expected_joined: Vec, + pub expected_page_digest: ContentDigest, +} + +impl PagePlan { + /// The page content after the edit, in the same parts as `content`. + pub fn expected_content(&self, content: &PageContent) -> PageContent { + let replaced: Vec<(usize, Vec)> = self + .edited_parts + .iter() + .filter_map(|i| Some((*i, self.expected_parts.get(*i)?.clone()))) + .collect(); + content.with_replaced_parts(&replaced) + } +} + +/// `plan` is `None` iff any verdict has a problem or every edit is a no-op. +pub struct PlanOutcome { + pub verdicts: Vec, + pub plan: Option, +} + +/// A problem with a code only (and a technical detail). +pub(crate) fn problem(code: EditProblemCode, detail: impl Into) -> EditProblem { + EditProblem::new(code, Some(detail.into())) +} + +/// `BAD_EDIT`: a malformed request (not a user problem). +pub(crate) fn bad_edit(detail: &str) -> AppError { + AppError::new( + "BAD_EDIT", + "Unsupported edit data", + "A text change has a style value OffPDF can't apply.", + ) + .with_details(detail.to_string()) +} + +fn too_many_edits(n: usize) -> AppError { + AppError::new( + "TOO_MANY_TEXT_EDITS", + "Too many text changes", + format!("This save has more than {EDITS_PER_PAGE_MAX} changed lines on one page."), + ) + .with_suggestion("Save in smaller batches.") + .with_details(format!("edits on one page: {n}")) +} + +/// Plans every edit of one page (verdicts in request order), then self-checks the whole page plan +/// by re-walking the expected content in the same context. `AppError` only for a malformed +/// request (`BAD_EDIT`, `TOO_MANY_TEXT_EDITS`) or an internal self-check failure +/// (`EDIT_VERIFY_FAILED`). +pub fn plan_page( + ctx: &SnapshotContext, + model: &PageModel, + edits: &[TextEditIn], +) -> Result { + if edits.len() > EDITS_PER_PAGE_MAX { + return Err(too_many_edits(edits.len())); + } + for e in edits { + style::validate_style(&e.style)?; + } + let mut counts: HashMap<&str, usize> = HashMap::new(); + for e in edits { + *counts.entry(e.run_id.as_str()).or_insert(0) += 1; + } + let mut verdicts = Vec::with_capacity(edits.len()); + let mut runs: Vec = Vec::new(); + let mut masks = masks::MaskFonts::new(ctx.doc()); + for e in edits { + if counts.get(e.run_id.as_str()).copied().unwrap_or(0) > 1 { + verdicts.push(EditVerdict::failed( + &e.run_id, + problem( + EditProblemCode::EditConflict, + "the same line is changed twice", + ), + )); + continue; + } + match plan_edit(model, e, &mut masks) { + Ok((verdict, Some(run_plan))) => { + let overlaps = runs.iter().flat_map(|r| &r.splices).any(|s| { + run_plan + .splices + .iter() + .any(|t| t.joined.start < s.joined.end && s.joined.start < t.joined.end) + }); + if overlaps { + verdicts.push(EditVerdict::failed( + &e.run_id, + problem(EditProblemCode::EditConflict, "overlapping replacements"), + )); + } else { + verdicts.push(verdict); + runs.push(run_plan); + } + } + Ok((verdict, None)) => verdicts.push(verdict), + Err(p) => verdicts.push(EditVerdict::failed(&e.run_id, p)), + } + } + if verdicts.iter().any(|v| v.problem.is_some()) || runs.is_empty() { + return Ok(PlanOutcome { + verdicts, + plan: None, + }); + } + let plan = assemble_page_plan(&model.content, model.page_index, runs); + self_check(ctx, model, &plan)?; + Ok(PlanOutcome { + verdicts, + plan: Some(plan), + }) +} + +/// Step 15: the expected content re-walked in the same context must pass `verify_page`. +fn self_check(ctx: &SnapshotContext, model: &PageModel, plan: &PagePlan) -> Result<(), AppError> { + let after = plan.expected_content(&model.content); + let walks = (&after, &*model.walk, &model.runs[..]); + verify::walk_and_verify_with(ctx, model.page_index, walks, plan, None, drop).map_err( + |failure| { + let p = problem( + EditProblemCode::EditVerifyFailed, + format!("self-check: {failure}"), + ); + let ctx = ProblemCtx { + page_number: Some(model.page_index.saturating_add(1)), + file_name: None, + face: None, + reason: None, + }; + EditProblemCode::EditVerifyFailed.to_app_error(&p, &ctx) + }, + ) +} + +/// One edit: `(verdict, plan)`; the plan is `None` for a no-op. +fn plan_edit( + model: &PageModel, + edit: &TextEditIn, + masks: &mut masks::MaskFonts<'_>, +) -> Result<(EditVerdict, Option), EditProblem> { + let run = style::resolve(model, edit)?; + style::validate_text(&edit.text)?; + let primary = run + .members + .first() + .and_then(|i| model.walk.records.get(*i)) + .ok_or_else(|| problem(EditProblemCode::EditVerifyFailed, "run without members"))?; + let primary_model: Option> = primary.font.clone(); + let mut wanted = style::normalise(run, &edit.style); + if edit.text.is_empty() { + // Removing a line draws nothing: there is no style to apply. + wanted = SourceTextStyleIn::default(); + } + if edit.text == run.text && wanted.is_empty() { + return Ok(( + EditVerdict { + run_id: run.id.clone(), + problem: None, + delta_pt: 0.0, + new_rect: Some(run.rect), + caret_offsets: Some(run.caret_offsets.clone()), + warnings: Vec::new(), + }, + None, + )); + } + let face_surface = style::availability( + run, + primary_model.as_deref(), + &model.walk.page_fonts, + &wanted, + &edit.text, + )?; + let target = style::targets(primary, run, &wanted, face_surface)?; + let surface = match &target.face_surface { + Some(s) => s.clone(), + None => model.surface(run), + }; + let new_units = diff::new_units(model, run, &surface, &target, &edit.text)?; + let laid = layout::lay_out(masks, model, run, &surface, &target, &new_units)?; + if laid.text != edit.text { + // Every later check compares the file with the request, not with the planner's reading + // of its own glyphs; a plan that would not read back as typed never leaves here. + return Err(problem( + EditProblemCode::EditVerifyFailed, + "the new line would not read back as typed", + )); + } + let verdict = EditVerdict { + run_id: run.id.clone(), + problem: None, + delta_pt: laid.delta_pt, + new_rect: Some(laid.new_rect), + caret_offsets: Some(laid.caret_offsets.clone()), + warnings: laid.warnings.clone(), + }; + drop(new_units); + let run_plan = bytes::run_plan(model, run, &edit.text, target, laid)?; + Ok((verdict, Some(run_plan))) +} diff --git a/src-tauri/src/pdf_engine/text_edit/rewrite/bytes.rs b/src-tauri/src/pdf_engine/text_edit/rewrite/bytes.rs new file mode 100644 index 0000000..8236e6d --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/rewrite/bytes.rs @@ -0,0 +1,451 @@ +//! Planner steps 11–14 (SPEC §B.12): the replacement bytes of every member's show op, the grammar +//! check on each (before any IO), the splices and the expected content parts. +//! +//! Primary: `prefix + set + body + restore`. Prefix keeps a `'` (`T*`) or `"` (`aw Tw ac Tc T*`, +//! verbatim) op's positioning; set writes the new `Tc`/fill; the body is one `TJ` per segment of +//! glyphs in one font, a `Tf` before a segment whose font differs from the font in force or when +//! the size changed, kept TJ numbers as their original bytes, the compensation after the last +//! segment; restore writes back, verbatim, exactly what changed: fill, `Tc`, then `Tf`. Every +//! absorbed member draws nothing and keeps its pen travel: `prefix + [<> n] TJ`. + +use super::layout::{Laid, Planned}; +use super::{problem, ExpectedRun, PagePlan, RunPlan, Splice, StyleTarget}; +use crate::pdf_engine::text_edit::content::PageContent; +use crate::pdf_engine::text_edit::encode::{ + check_replacement_grammar, hex_codes, needs_leading_space, +}; +use crate::pdf_engine::text_edit::fonts::Code; +use crate::pdf_engine::text_edit::lexer::{is_delimiter, is_whitespace}; +use crate::pdf_engine::text_edit::reasons::{EditProblem, EditProblemCode as P, TextReason}; +use crate::pdf_engine::text_edit::runs::{PageModel, TextRun}; +use crate::pdf_engine::text_edit::state::{color_effect, ColorSpaceKind, Paint}; +use crate::pdf_engine::text_edit::walker::{ShowOp, ShowRecord}; +use crate::pdf_engine::validate_output::content_digest; +use std::sync::Arc; + +/// A PDF name token for a resource name (`#xx` for anything not a plain regular character). +pub(crate) fn name_token(name: &[u8]) -> String { + let mut out = String::with_capacity(name.len() + 1); + out.push('/'); + for &b in name { + let plain = + (0x21..=0x7e).contains(&b) && b != b'#' && !is_delimiter(b) && !is_whitespace(b); + if plain { + out.push(char::from(b)); + } else { + out.push_str(&format!("#{b:02X}")); + } + } + out +} + +fn joined_text(content: &PageContent, span: &std::ops::Range) -> String { + String::from_utf8_lossy(content.joined.get(span.clone()).unwrap_or_default()).into_owned() +} + +/// The positioning a show op does before drawing, verbatim. +fn op_prefix(content: &PageContent, rec: &ShowRecord) -> Vec { + match rec.op { + ShowOp::Tj | ShowOp::TJ => Vec::new(), + ShowOp::Quote => vec!["T*".to_string()], + ShowOp::DoubleQuote => { + let operand = |i: usize| { + rec.operand_spans + .get(i) + .map(|s| joined_text(content, s)) + .unwrap_or_default() + }; + vec![ + format!("{} Tw", operand(0)), + format!("{} Tc", operand(1)), + "T*".to_string(), + ] + } + } +} + +/// One TJ segment: its font resource and its elements. +struct Segment { + font: Vec, + elements: Vec, + codes: Vec, + has_glyph: bool, +} + +impl Segment { + fn flush_codes(&mut self) { + if !self.codes.is_empty() { + self.elements.push(hex_codes(&self.codes)); + self.codes.clear(); + } + } +} + +/// Splits the new units into TJ segments (kerns attach to the preceding segment; leading kerns +/// to the first one). +fn segments(units: &[Planned], content: &PageContent) -> Vec { + let mut segs: Vec = Vec::new(); + let mut pending: Vec = Vec::new(); + for u in units { + match u { + Planned::Glyph { code, font_res, .. } => { + let same = segs.last().is_some_and(|s| s.font == *font_res); + if !same { + if let Some(last) = segs.last_mut() { + last.flush_codes(); + } + segs.push(Segment { + font: font_res.clone(), + elements: std::mem::take(&mut pending), + codes: Vec::new(), + has_glyph: true, + }); + } + if let Some(seg) = segs.last_mut() { + seg.codes.push(*code); + } + } + Planned::Kern { + written, original, .. + } => { + let token = match (written, original) { + (Some(s), _) => s.clone(), + (None, Some(span)) => joined_text(content, span), + (None, None) => continue, + }; + match segs.last_mut() { + Some(seg) => { + seg.flush_codes(); + seg.elements.push(token); + } + None => pending.push(token), + } + } + } + } + if let Some(last) = segs.last_mut() { + last.flush_codes(); + } + if segs.is_empty() { + segs.push(Segment { + font: Vec::new(), + elements: pending, + codes: Vec::new(), + has_glyph: false, + }); + } + segs +} + +/// The primary's replacement (`prefix + set + body + restore`) and its TJ count. +fn primary_bytes( + content: &PageContent, + primary: &ShowRecord, + target: &StyleTarget, + laid: &Laid, +) -> Result<(String, usize), EditProblem> { + let internal = |what: &str| problem(P::EditVerifyFailed, what.to_string()); + let mut tokens = op_prefix(content, primary); + let fmt = |v: f64| crate::pdf_engine::text_edit::encode::fmt_num(v); + if target.tc_changed { + tokens.push(format!("{} Tc", fmt(target.tc)?)); + } + if let (true, Some(rgb)) = (target.fill_changed, target.fill) { + tokens.push(format!( + "{} {} {} rg", + fmt(rgb[0])?, + fmt(rgb[1])?, + fmt(rgb[2])? + )); + } + let original_font: Vec = primary + .before + .text + .font + .as_ref() + .and_then(|f| f.resource.as_deref()) + .map(<[u8]>::to_vec) + .unwrap_or_default(); + let units: Vec = laid.units.iter().map(|p| p.unit.clone()).collect(); + let mut segs = segments(&units, content); + let tf_token = String::from_utf8_lossy(&target.tfs_token).into_owned(); + let mut in_force = original_font.clone(); + let seg_count = segs.len(); + if let (Some(n_c), Some(last)) = (&laid.compensation, segs.last_mut()) { + last.elements.push(n_c.clone()); + } + for seg in &segs { + let switch = seg.has_glyph && !seg.font.is_empty() && seg.font != in_force; + if switch || (seg.has_glyph && target.size_changed) { + if seg.font.is_empty() || tf_token.is_empty() { + return Err(internal("a font switch without a Tf font")); + } + tokens.push(format!("{} {tf_token} Tf", name_token(&seg.font))); + in_force = seg.font.clone(); + } + let elements = if seg.elements.is_empty() { + "<>".to_string() + } else if seg.has_glyph { + seg.elements.join(" ") + } else { + format!("<> {}", seg.elements.join(" ")) + }; + tokens.push(format!("[{elements}] TJ")); + } + let after = &primary.after; + if target.fill_changed { + tokens.extend(fill_restore(&after.fill)?); + } + if target.tc_changed { + tokens.push(match &after.text.tc_src { + Some(b) => String::from_utf8_lossy(b).into_owned(), + None => "0 Tc".to_string(), + }); + } + if in_force != original_font || target.size_changed { + let tf_op = after + .text + .font + .as_ref() + .and_then(|f| f.tf_op.as_deref()) + .ok_or_else(|| internal("no Tf to restore"))?; + tokens.push(String::from_utf8_lossy(tf_op).into_owned()); + } + Ok((tokens.join(" "), seg_count)) +} + +/// The ops that put the fill back after the new `rg`: the verbatim colour-space and colour ops in +/// force after the original op, or `0 g` when the page never set a fill. An `sc`/`scn` written +/// without its own `cs` (the space came from the initial DeviceGray or from an earlier `g`/`rg`/ +/// `k`) gets that device space's `cs` first: after the new `rg` it would otherwise apply to +/// DeviceRGB. +fn fill_restore(fill: &Paint) -> Result, EditProblem> { + let text = |b: &[u8]| String::from_utf8_lossy(b).into_owned(); + let (space_op, color_op) = (fill.space_op.as_deref(), fill.color_op.as_deref()); + if space_op.is_none() && color_op.is_none() { + return Ok(vec!["0 g".to_string()]); + } + let mut out = Vec::with_capacity(3); + let bare_sc = color_op.is_some_and(|op| { + matches!( + op.split(|b| is_whitespace(*b)) + .filter(|t| !t.is_empty()) + .next_back(), + Some(b"sc" | b"scn") + ) + }); + if space_op.is_none() && bare_sc { + let device = match fill.space { + ColorSpaceKind::Default | ColorSpaceKind::DeviceGray => "/DeviceGray cs", + ColorSpaceKind::DeviceRgb => "/DeviceRGB cs", + ColorSpaceKind::DeviceCmyk => "/DeviceCMYK cs", + ColorSpaceKind::Named(..) | ColorSpaceKind::Pattern => { + return Err(problem( + P::EditVerifyFailed, + "a colour op without its colour space", + )) + } + }; + out.push(device.to_string()); + } + out.extend(space_op.map(text)); + out.extend(color_op.map(text)); + Ok(out) +} + +/// An absorbed member's replacement: its positioning, then nothing drawn and the same pen travel. +fn absorbed_bytes(content: &PageContent, rec: &ShowRecord, number: Option<&String>) -> String { + let mut tokens = op_prefix(content, rec); + tokens.push(match number { + Some(n) => format!("[<> {n}] TJ"), + None => "[<>] TJ".to_string(), + }); + tokens.join(" ") +} + +/// The splice of `rec`'s op span with `text` (a space is prepended when the first token could +/// fuse with the byte before it). +fn splice_of(content: &PageContent, rec: &ShowRecord, text: String) -> Result { + let span = rec + .span + .clone() + .ok_or_else(|| problem(P::EditVerifyFailed, "show op without a span"))?; + let Some((part, local)) = content.locate(&span) else { + let mut p = problem(P::TextEditRefused, "show op straddles two content parts"); + p.reason = Some(TextReason::SplitContent); + return Err(p); + }; + let prev = local + .start + .checked_sub(1) + .and_then(|i| content.part_bytes(part).get(i).copied()); + let mut bytes = text.into_bytes(); + if needs_leading_space(prev, &bytes) { + bytes.insert(0, b' '); + } + Ok(Splice { + part, + local, + joined: span, + bytes, + }) +} + +/// The expected fill of the edited glyphs. +fn expected_fill(primary: &ShowRecord, target: &StyleTarget) -> Result { + let Some(rgb) = target.fill.filter(|_| target.fill_changed) else { + return Ok(primary.before.fill.clone()); + }; + let fmt = crate::pdf_engine::text_edit::encode::fmt_num; + let op = format!("{} {} {} rg", fmt(rgb[0])?, fmt(rgb[1])?, fmt(rgb[2])?); + Ok(Paint { + space_op: None, + color_op: Some(Arc::from(op.as_bytes())), + effect: color_effect(&ColorSpaceKind::DeviceRgb, &rgb), + space: ColorSpaceKind::DeviceRgb, + comps: rgb.to_vec(), + pattern: false, + pattern_hash: None, + }) +} + +/// Steps 11–14 for one run. +pub(super) fn run_plan( + model: &PageModel, + run: &TextRun, + requested_text: &str, + target: StyleTarget, + laid: Laid, +) -> Result { + let content = &model.content; + let records: Vec<&ShowRecord> = run + .members + .iter() + .filter_map(|i| model.walk.records.get(*i)) + .collect(); + let Some(primary) = records.first().copied() else { + return Err(problem(P::EditVerifyFailed, "run without members")); + }; + let grammar = |bytes: &str, tj: usize| { + check_replacement_grammar(bytes.as_bytes(), tj) + .map_err(|op| problem(P::EditVerifyFailed, format!("replacement grammar: {op}"))) + }; + let (primary_text, seg_count) = primary_bytes(content, primary, &target, &laid)?; + grammar(&primary_text, seg_count)?; + let mut splices = vec![splice_of(content, primary, primary_text)?]; + let mut emitted_records = vec![seg_count]; + for (m, rec) in records.iter().enumerate().skip(1) { + let text = absorbed_bytes( + content, + rec, + laid.absorbed_numbers.get(m).and_then(Option::as_ref), + ); + grammar(&text, 1)?; + splices.push(splice_of(content, rec, text)?); + emitted_records.push(1); + } + let mut glyphs = Vec::new(); + let mut glyph_new = Vec::new(); + let mut kept_from = Vec::new(); + for p in &laid.units { + if let Planned::Glyph { + code, + font_res, + font_hash, + kept, + .. + } = &p.unit + { + glyphs.push((font_res.clone(), *font_hash, *code)); + glyph_new.push(kept.is_none()); + kept_from.push(kept.map(|(_, member, gi)| (member, gi))); + } + } + let expected = ExpectedRun { + run_id: run.id.clone(), + text: laid.text.clone(), + glyphs, + prefix_glyphs: laid.prefix_glyphs, + suffix_glyphs: laid.suffix_glyphs, + shift_user: laid.shift_user, + unshifted_from: laid.unshifted_from, + origin: run.origin, + primary_pen_after: primary.pen_after, + member_pen_after: records.iter().map(|r| r.pen_after).collect(), + tfs: target.tfs, + effective_size: target.effective_size, + tc: target.tc, + fill: expected_fill(primary, &target)?, + emitted_records, + member_spans: records.iter().filter_map(|r| r.span.clone()).collect(), + glyph_origins: laid.glyph_origins.clone(), + glyph_new, + kept_from, + mask_boxes: laid.mask_boxes.clone(), + old_ink_boxes: laid.old_ink_boxes.clone(), + glyphs_changed: laid.glyphs_changed, + requested_text: requested_text.to_string(), + }; + Ok(RunPlan { + run_id: run.id.clone(), + target, + splices, + expected, + new_rect: laid.new_rect, + }) +} + +/// Recomputes `expected_parts`, `edited_parts`, `expected_joined` and `expected_page_digest` +/// from `plan.splices` (sorted by joined start). Splices are applied back to front per part. +pub(crate) fn rebuild_parts(content: &PageContent, plan: &mut PagePlan) { + plan.splices.sort_by_key(|s| s.joined.start); + let mut parts: Vec> = (0..content.parts.len()) + .map(|i| content.part_bytes(i).to_vec()) + .collect(); + let mut edited: Vec = Vec::new(); + for s in plan.splices.iter().rev() { + let Some(part) = parts.get_mut(s.part) else { + continue; + }; + if s.local.start <= s.local.end && s.local.end <= part.len() { + part.splice(s.local.clone(), s.bytes.iter().copied()); + if !edited.contains(&s.part) { + edited.push(s.part); + } + } + } + edited.sort_unstable(); + let mut joined = Vec::new(); + let mut concat = Vec::new(); + for (i, p) in parts.iter().enumerate() { + if i > 0 { + joined.push(b'\n'); + } + joined.extend_from_slice(p); + concat.extend_from_slice(p); + } + plan.expected_page_digest = content_digest(&concat); + plan.expected_parts = parts; + plan.edited_parts = edited; + plan.expected_joined = joined; +} + +/// The page plan of `runs` over `content`. +pub(crate) fn assemble_page_plan( + content: &PageContent, + page_index: u32, + runs: Vec, +) -> PagePlan { + let splices = runs.iter().flat_map(|r| r.splices.clone()).collect(); + let mut plan = PagePlan { + page_index, + runs, + splices, + expected_parts: Vec::new(), + edited_parts: Vec::new(), + expected_joined: Vec::new(), + expected_page_digest: content_digest(b""), + }; + rebuild_parts(content, &mut plan); + plan +} diff --git a/src-tauri/src/pdf_engine/text_edit/rewrite/cff_scale.rs b/src-tauri/src/pdf_engine/text_edit/rewrite/cff_scale.rs new file mode 100644 index 0000000..babc9f7 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/rewrite/cff_scale.rs @@ -0,0 +1,412 @@ +//! The units-to-em scale of a CFF program's outlines, for the A5 masks (SPEC §B.15). +//! +//! ttf-parser outlines CFF glyphs in font units and reads only the Top DICT `FontMatrix`. A +//! CID-keyed program can also give each Font DICT of its FDArray a `FontMatrix`, which renderers +//! (FreeType, hence Poppler) combine with the Top DICT one (Adobe TN #5176), so the Top DICT +//! matrix alone can be off by any factor. A mask built at a wrong scale is either empty (an honest +//! edit fails) or several em wide (a moved neighbour on the same line is hidden), so a program's +//! own glyph boxes are used only when its scale is unambiguous: +//! - CID-keyed: no Font DICT has a `FontMatrix`, and the Top DICT one is absent or the default; +//! - name-keyed: the Top DICT `FontMatrix` is absent (the default) or one plain uniform scale (no +//! skew, no offset) of at least 16 units per em; inside an OpenType font it must be the default. +//! +//! The default is 1/1000 em per unit for a bare program and 1/`unitsPerEm` inside an OpenType +//! font. Anything else, or a structure this reader cannot follow, is `None`: the masks then use +//! the bounded fallback box. Checked slicing only; no DICT longer than `DICT_LEN_MAX` is walked. + +use crate::pdf_engine::text_edit::fonts::cff_layout::DICT_LEN_MAX; +use std::ops::Range; + +/// Where a CFF program lives. +#[derive(Clone, Copy, Debug)] +pub(super) enum Host { + /// `FontFile3 /Type1C` or `/CIDFontType0C`. + Bare, + /// The `CFF ` table of an OpenType font with this `unitsPerEm`. + OpenType(u16), +} + +const OP_FONT_MATRIX: u16 = 1207; +const OP_ROS: u16 = 1230; +const OP_FD_ARRAY: u16 = 1236; +/// Items read from one INDEX: FDSelect selects a Font DICT with one byte, and a PDF font program +/// holds one font (more is not followed). +const INDEX_ITEMS_MAX: usize = 256; +/// Operands kept per DICT operator (as ttf-parser); more makes the DICT unreadable here. +const OPERANDS_MAX: usize = 48; +/// Longest real-number operand read, in nibbles. +const REAL_NIBBLES_MAX: usize = 64; +/// Largest scale accepted: at least 16 units per em (the OpenType minimum). +const SCALE_MAX: f64 = 1.0 / 16.0; +/// Relative tolerance of "equal to the default scale" (matrices are written as decimals). +const SCALE_REL_TOL: f64 = 1e-4; + +/// Em per font unit of the outlines of `cff`, when it is unambiguous (see the module doc). +pub(super) fn em_per_unit(cff: &[u8], host: Host) -> Option { + let default = match host { + Host::Bare => 0.001, + Host::OpenType(upem) if upem > 0 => 1.0 / f64::from(upem), + Host::OpenType(_) => return None, + }; + let top = TopEntries::read(cff)?; + let explicit = match &top.matrix { + None => None, + Some(m) => Some(plain_scale(m)?), + }; + let near_default = |s: f64| (s - default).abs() <= default * SCALE_REL_TOL; + if top.ros { + if explicit.is_some_and(|s| !near_default(s)) { + return None; + } + return (!font_dicts_have_matrix(cff, top.fd_array?)?).then_some(default); + } + match (explicit, host) { + (None, _) => Some(default), + (Some(s), Host::Bare) => (s <= SCALE_MAX).then_some(s), + (Some(s), Host::OpenType(_)) => near_default(s).then_some(default), + } +} + +/// The scale of a `FontMatrix` that is one uniform scale with no skew or offset. +fn plain_scale(m: &[f64]) -> Option { + let [a, b, c, d, e, f] = m else { + return None; + }; + let plain = *b == 0.0 && *c == 0.0 && *e == 0.0 && *f == 0.0 && *a > 0.0; + (plain && (a - d).abs() <= a * SCALE_REL_TOL).then_some(*a) +} + +/// The Top DICT entries this reader needs (later entries win, as in ttf-parser and FreeType). +struct TopEntries { + ros: bool, + matrix: Option>, + fd_array: Option, +} + +impl TopEntries { + /// The first Top DICT of `cff` (header, Name INDEX, Top DICT INDEX). + fn read(cff: &[u8]) -> Option { + if *cff.first()? != 1 { + return None; + } + let header = usize::from(*cff.get(2)?); + if header < 4 { + return None; + } + let (_, after_names) = index_items(cff, header)?; + let (tops, _) = index_items(cff, after_names)?; + let dict = cff.get(tops.first()?.clone())?; + let mut top = TopEntries { + ros: false, + matrix: None, + fd_array: None, + }; + walk_dict(dict, |op, operands| { + match op { + OP_ROS => top.ros = true, + OP_FONT_MATRIX => top.matrix = Some(operands.to_vec()), + OP_FD_ARRAY => top.fd_array = Some(offset_of(operands)?), + _ => {} + } + Some(()) + })?; + Some(top) + } +} + +/// A DICT offset operand: exactly one non-negative integer. +fn offset_of(operands: &[f64]) -> Option { + let [v] = operands else { + return None; + }; + (v.fract() == 0.0 && *v >= 0.0 && *v <= f64::from(u32::MAX)).then(|| *v as usize) +} + +/// Whether any Font DICT of the FDArray at `at` has a `FontMatrix` (`None`: unreadable). +fn font_dicts_have_matrix(cff: &[u8], at: usize) -> Option { + let (dicts, _) = index_items(cff, at)?; + let mut found = false; + for range in dicts { + walk_dict(cff.get(range)?, |op, _| { + found |= op == OP_FONT_MATRIX; + Some(()) + })?; + } + Some(found) +} + +/// The item ranges (absolute) of the INDEX at `at` and the offset after it. `None` for a malformed +/// INDEX or one of more than `INDEX_ITEMS_MAX` items. +fn index_items(cff: &[u8], at: usize) -> Option<(Vec>, usize)> { + let count = usize::from(u16::from_be_bytes([ + *cff.get(at)?, + *cff.get(at.checked_add(1)?)?, + ])); + let pos = at.checked_add(2)?; + if count == 0 { + return Some((Vec::new(), pos)); + } + if count > INDEX_ITEMS_MAX { + return None; + } + let off_size = usize::from(*cff.get(pos)?); + if !(1..=4).contains(&off_size) { + return None; + } + let offsets_at = pos.checked_add(1)?; + let offsets_len = count.checked_add(1)?.checked_mul(off_size)?; + let offsets = cff.get(offsets_at..offsets_at.checked_add(offsets_len)?)?; + // Offsets are 1-based from the byte before the data. + let base = offsets_at.checked_add(offsets_len)?.checked_sub(1)?; + let offset = |i: usize| -> Option { + let start = i.checked_mul(off_size)?; + let raw = offsets + .get(start..start.checked_add(off_size)?)? + .iter() + .fold(0usize, |acc, b| (acc << 8) | usize::from(*b)); + (raw >= 1).then_some(raw) + }; + let mut items = Vec::with_capacity(count); + let mut start = offset(0)?; + for i in 1..=count { + let end = offset(i)?; + if end < start { + return None; + } + items.push(base.checked_add(start)?..base.checked_add(end)?); + start = end; + } + let end = base.checked_add(start)?; + (end <= cff.len()).then_some((items, end)) +} + +/// Calls `f(operator, operands)` for every entry of a DICT (two-byte operators as `1200 + b1`). +/// `None` when the DICT is longer than `DICT_LEN_MAX`, malformed, or `f` returns `None`. +fn walk_dict(dict: &[u8], mut f: impl FnMut(u16, &[f64]) -> Option<()>) -> Option<()> { + if dict.len() > DICT_LEN_MAX { + return None; + } + let byte = |at: usize| dict.get(at).copied(); + let mut operands: Vec = Vec::new(); + let mut pos = 0usize; + while let Some(b0) = byte(pos) { + pos = pos.checked_add(1)?; + let value = match b0 { + 12 => { + let b1 = byte(pos)?; + pos = pos.checked_add(1)?; + f(1200 + u16::from(b1), &operands)?; + operands.clear(); + continue; + } + 0..=27 | 31 | 255 => { + f(u16::from(b0), &operands)?; + operands.clear(); + continue; + } + 28 => { + let v = i16::from_be_bytes([byte(pos)?, byte(pos.checked_add(1)?)?]); + pos = pos.checked_add(2)?; + f64::from(v) + } + 29 => { + let at = |k: usize| byte(pos.checked_add(k)?); + let v = i32::from_be_bytes([at(0)?, at(1)?, at(2)?, at(3)?]); + pos = pos.checked_add(4)?; + f64::from(v) + } + 30 => { + let (v, next) = real(dict, pos)?; + pos = next; + v + } + 32..=246 => f64::from(b0) - 139.0, + 247..=250 => { + let b1 = byte(pos)?; + pos = pos.checked_add(1)?; + (f64::from(b0) - 247.0) * 256.0 + f64::from(b1) + 108.0 + } + 251..=254 => { + let b1 = byte(pos)?; + pos = pos.checked_add(1)?; + -(f64::from(b0) - 251.0) * 256.0 - f64::from(b1) - 108.0 + } + }; + if operands.len() >= OPERANDS_MAX { + return None; + } + operands.push(value); + } + Some(()) +} + +/// A DICT real number (nibbles after byte 30, from `pos`) and the offset after it. +fn real(data: &[u8], mut pos: usize) -> Option<(f64, usize)> { + let mut text = String::new(); + loop { + let b = *data.get(pos)?; + pos = pos.checked_add(1)?; + for nibble in [b >> 4, b & 0xF] { + match nibble { + 0..=9 => text.push(char::from(b'0' + nibble)), + 0xA => text.push('.'), + 0xB => text.push('E'), + 0xC => text.push_str("E-"), + 0xE => text.push('-'), + 0xF => { + let v = text.parse::().ok().filter(|v| v.is_finite())?; + return Some((v, pos)); + } + _ => return None, + } + } + if text.len() > REAL_NIBBLES_MAX { + return None; + } + } +} + +#[cfg(test)] +mod tests { + use super::{em_per_unit, Host}; + use crate::pdf_engine::text_edit::testkit::cff::{box_charstring, CffBuilder}; + use crate::pdf_engine::text_edit::testkit::fakes::cff_with_font_dict_matrix; + + /// `FontMatrix` operands: `[1 0 0 1 0 0]` (integers) and `[s 0 0 s 0 0]` for a real `s`. + const IDENTITY: [u8; 6] = [140, 139, 139, 140, 139, 139]; + const MILLI: [u8; 12] = [ + 30, 0x0a, 0x00, 0x1f, 139, 139, 30, 0x0a, 0x00, 0x1f, 139, 139, + ]; + const HALF_MILLI: [u8; 14] = [ + 30, 0x0a, 0x00, 0x05, 0xff, 139, 139, 30, 0x0a, 0x00, 0x05, 0xff, 139, 139, + ]; + + fn matrix_op(operands: &[u8]) -> Vec { + let mut v = operands.to_vec(); + v.extend([12, 7]); + v + } + + fn cid(top: Option<&[u8]>, fd: Option<&[u8]>) -> Vec { + let mut b = CffBuilder::new("ABCDEF+Box").raw_cid_glyph(1, box_charstring()); + if let Some(t) = top { + b.top_padding = matrix_op(t); + } + let cff = b.build(); + fd.map_or(cff.clone(), |m| cff_with_font_dict_matrix(&cff, m)) + } + + fn named(top: Option<&[u8]>) -> Vec { + let mut b = CffBuilder::new("ABCDEF+Box").glyph("A", true); + if let Some(t) = top { + b.top_padding = matrix_op(t); + } + b.build() + } + + #[test] + fn cid_keyed_programs_use_their_boxes_only_without_font_dict_matrices() { + assert_eq!(em_per_unit(&cid(None, None), Host::Bare), Some(0.001)); + assert_eq!( + em_per_unit(&cid(Some(&MILLI), None), Host::Bare), + Some(0.001), + "an explicit default Top matrix" + ); + assert_eq!( + em_per_unit(&cid(Some(&IDENTITY), Some(&MILLI)), Host::Bare), + None, + "Top [1 0 0 1 0 0] + Font DICT [0.001 …] (ttf-parser would scale by 1)" + ); + assert_eq!( + em_per_unit(&cid(None, Some(&HALF_MILLI)), Host::Bare), + None, + "Font DICT [0.0005 …] under the default Top (ttf-parser would scale by 0.001)" + ); + assert_eq!( + em_per_unit(&cid(Some(&HALF_MILLI), None), Host::Bare), + None, + "a non-default Top matrix of a CID-keyed program" + ); + assert_eq!( + em_per_unit(&cid(None, None), Host::OpenType(1000)), + Some(0.001) + ); + } + + #[test] + fn name_keyed_programs_need_one_plain_scale() { + assert_eq!(em_per_unit(&named(None), Host::Bare), Some(0.001)); + assert_eq!( + em_per_unit(&named(Some(&HALF_MILLI)), Host::Bare), + Some(0.0005) + ); + assert_eq!( + em_per_unit(&named(Some(&IDENTITY)), Host::Bare), + None, + "one unit per em" + ); + let skewed = [ + 30, 0x0a, 0x00, 0x1f, 139, 30, 0x0a, 0x00, 0x2f, 30, 0x0a, 0x00, 0x1f, 139, 139, + ]; + assert_eq!(em_per_unit(&named(Some(&skewed)), Host::Bare), None, "skew"); + let offset = [ + 30, 0x0a, 0x00, 0x1f, 139, 139, 30, 0x0a, 0x00, 0x1f, 140, 139, + ]; + assert_eq!( + em_per_unit(&named(Some(&offset)), Host::Bare), + None, + "offset" + ); + assert_eq!( + em_per_unit(&named(Some(&[140, 139])), Host::Bare), + None, + "2 operands" + ); + // Inside OpenType the scale is 1/unitsPerEm; a Top matrix must agree with it. + assert_eq!( + em_per_unit(&named(None), Host::OpenType(2048)), + Some(1.0 / 2048.0) + ); + assert_eq!( + em_per_unit(&named(Some(&MILLI)), Host::OpenType(1000)), + Some(0.001) + ); + assert_eq!( + em_per_unit(&named(Some(&MILLI)), Host::OpenType(2048)), + None + ); + assert_eq!(em_per_unit(&named(None), Host::OpenType(0)), None); + } + + #[test] + fn unreadable_programs_have_no_scale() { + let good = cid(None, None); + assert_eq!(em_per_unit(&good, Host::Bare), Some(0.001)); + // Cut inside the header, the Name INDEX, the Top DICT INDEX, or before the FDArray. + for n in [0, 1, 3, 4, 10, 40, good.len() / 2] { + assert_eq!( + em_per_unit(&good[..n], Host::Bare), + None, + "truncated at {n}" + ); + } + let mut wrong_version = good.clone(); + wrong_version[0] = 2; + assert_eq!(em_per_unit(&wrong_version, Host::Bare), None); + let mut short_header = good; + short_header[2] = 3; + assert_eq!(em_per_unit(&short_header, Host::Bare), None); + // A real number without an end nibble, and a reserved nibble. + assert_eq!( + em_per_unit(&named(Some(&[30, 0x0a, 0x00, 0x11])), Host::Bare), + None + ); + assert_eq!( + em_per_unit( + &named(Some(&[30, 0xd1, 0xff, 139, 139, 139, 139, 139])), + Host::Bare + ), + None + ); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/rewrite/diff.rs b/src-tauri/src/pdf_engine/text_edit/rewrite/diff.rs new file mode 100644 index 0000000..d3d7257 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/rewrite/diff.rs @@ -0,0 +1,252 @@ +//! Planner steps 6–8 (SPEC §B.12): the minimal diff and the new unit list. Unchanged leading and +//! trailing units keep their codes and the kerns between them; the middle is encoded through the +//! typing surface (run, then page code preferences, §A.3.2); a typed space is a glyph or, in kern +//! mode, a TJ number (§A.3.4). At the two boundaries a pair kern tuned for a glyph pair that no +//! longer exists is dropped; synthetic-space kerns count as " " and stay with the side that +//! matched them. Kerns before the first glyph and after the last one are not between glyphs and +//! are always kept (unless the whole line is removed or re-faced). + +use super::{problem, NewUnit, StyleTarget}; +use crate::pdf_engine::text_edit::encode::num; +use crate::pdf_engine::text_edit::fonts::{Code, TypingSurface}; +use crate::pdf_engine::text_edit::reasons::{EditProblem, EditProblemCode as P}; +use crate::pdf_engine::text_edit::runs::{PageModel, SpaceMode, TextRun, Unit}; + +/// Where a unit of the new list comes from. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum Region { + Prefix, + Middle, + Suffix, +} + +/// The text a unit contributes when matching: a glyph's text, " " for a synthetic-space kern, +/// nothing for any other kern. +fn unit_text(u: &Unit) -> Option<&str> { + match u { + Unit::Glyph { text, .. } => Some(text.as_str()), + Unit::Kern { + synth_space: true, .. + } => Some(" "), + Unit::Kern { .. } => None, + } +} + +fn starts_with(hay: &[char], needle: &str) -> Option { + let n: Vec = needle.chars().collect(); + (hay.len() >= n.len() && hay.get(..n.len()) == Some(n.as_slice())).then_some(n.len()) +} + +fn ends_with(hay: &[char], needle: &str) -> Option { + let n: Vec = needle.chars().collect(); + let start = hay.len().checked_sub(n.len())?; + (hay.get(start..) == Some(n.as_slice())).then_some(n.len()) +} + +/// `(prefix_end, prefix_chars, suffix_start, suffix_chars)`: units `..prefix_end` and +/// `suffix_start..` are kept; they match the first `prefix_chars` and last `suffix_chars` chars. +fn kept_ranges(units: &[Unit], chars: &[char]) -> (usize, usize, usize, usize) { + let first_glyph = units + .iter() + .position(|u| matches!(u, Unit::Glyph { .. })) + .unwrap_or(units.len()); + let last_glyph_end = units + .iter() + .rposition(|u| matches!(u, Unit::Glyph { .. })) + .map_or(units.len(), |i| i + 1); + // Prefix: leading kerns, then whole text units while they match. + let mut prefix_end = first_glyph; + let mut matched = 0usize; + let mut i = first_glyph; + while let Some(u) = units.get(i) { + match unit_text(u) { + None => i += 1, + Some(t) => match chars.get(matched..).and_then(|rest| starts_with(rest, t)) { + Some(n) => { + matched += n; + i += 1; + prefix_end = i; + } + None => break, + }, + } + } + let prefix_chars = matched; + // Suffix: trailing kerns, then whole text units from the end while they match. + let mut suffix_start = last_glyph_end.max(prefix_end); + let mut matched_s = 0usize; + let mut k = last_glyph_end; + while k > prefix_end { + let Some(u) = units.get(k - 1) else { break }; + match unit_text(u) { + None => k -= 1, + Some(t) => { + let room = chars.len().saturating_sub(prefix_chars + matched_s); + let end = chars.len().saturating_sub(matched_s); + let fits = t.chars().count() <= room; + match chars.get(..end).and_then(|head| ends_with(head, t)) { + Some(n) if fits => { + matched_s += n; + k -= 1; + suffix_start = k; + } + _ => break, + } + } + } + } + (prefix_end, prefix_chars, suffix_start, matched_s) +} + +/// Run and page code preferences (§A.3.2) for the surface fonts. +fn preferences( + model: &PageModel, + run: &TextRun, + surface: &TypingSurface, + wanted: &[char], +) -> (Vec<(usize, char, Code)>, Vec<(usize, char, Code)>) { + let font_index = |res: Option<&[u8]>, hash: u64| { + surface.fonts.iter().position(|(n, m)| { + m.content_hash == hash && (n.as_slice() == res.unwrap_or_default() || n.is_empty()) + }) + }; + let single = |t: &str| { + let mut it = t.chars(); + match (it.next(), it.next()) { + (Some(c), None) => Some(c), + _ => None, + } + }; + let mut prefer_run = Vec::new(); + for u in &run.units { + if let Unit::Glyph { + code, + font_res, + font_hash, + text, + .. + } = u + { + if let (Some(ch), Some(idx)) = + (single(text), font_index(font_res.as_deref(), *font_hash)) + { + if wanted.contains(&ch) && !prefer_run.iter().any(|(_, c, _)| *c == ch) { + prefer_run.push((idx, ch, *code)); + } + } + } + } + let mut prefer_page: Vec<(usize, char, Code)> = Vec::new(); + let mut left: Vec = wanted + .iter() + .copied() + .filter(|c| !prefer_run.iter().any(|(_, r, _)| r == c)) + .collect(); + 'records: for rec in &model.walk.records { + if left.is_empty() { + break; + } + let Some(font) = rec.before.text.font.as_ref() else { + continue; + }; + let Some(idx) = font_index(font.resource.as_deref(), font.content_hash) else { + continue; + }; + for g in &rec.glyphs { + let Some(ch) = g.text.as_deref().and_then(single) else { + continue; + }; + if let Some(pos) = left.iter().position(|c| *c == ch) { + prefer_page.push((idx, ch, g.code)); + left.swap_remove(pos); + if left.is_empty() { + break 'records; + } + } + } + } + (prefer_run, prefer_page) +} + +/// §A.3.4 kern mode: the first space of the whole requested line that is not between two +/// non-space characters (leading, trailing or doubled). The whole line is checked, not only the +/// typed middle: a synthetic-space kern kept at either end no longer sits between two glyphs and +/// would not read back as a space (the frontend's `spaceProblem` applies the same rule). +fn misplaced_space(chars: &[char]) -> Option { + let space_at = |i: Option| i.and_then(|i| chars.get(i)) == Some(&' '); + (0..chars.len()).find(|&i| { + space_at(Some(i)) && (i == 0 || i + 1 == chars.len() || space_at(i.checked_sub(1))) + }) +} + +/// Steps 6–8: the new unit list (with each unit's region) for `text`. +pub(super) fn new_units( + model: &PageModel, + run: &TextRun, + surface: &TypingSurface, + target: &StyleTarget, + text: &str, +) -> Result, EditProblem> { + let chars: Vec = text.chars().collect(); + if chars.is_empty() { + return Ok(Vec::new()); + } + let kern_mode = run.space_mode == SpaceMode::Kern; + if let Some(i) = misplaced_space(&chars).filter(|_| kern_mode) { + return Err(problem( + P::SpaceNotWritable, + format!("space at character {i}"), + )); + } + let units = &run.units; + let (prefix_end, prefix_chars, suffix_start, suffix_chars) = if target.face_changed { + (0, 0, units.len(), 0) + } else { + kept_ranges(units, &chars) + }; + let middle_end = chars.len().saturating_sub(suffix_chars); + let wanted: &[char] = chars.get(prefix_chars..middle_end).unwrap_or_default(); + let (prefer_run, prefer_page) = preferences(model, run, surface, wanted); + let kern_space = num(run.kern_space)?.1; + let mut encoded = Vec::with_capacity(wanted.len()); + let mut missing: Vec = Vec::new(); + for ch in wanted { + if *ch == ' ' && kern_mode { + encoded.push(NewUnit::KernSpace(kern_space)); + continue; + } + match surface.writer_for_with(*ch, &prefer_run, &prefer_page) { + Some((font, code)) => encoded.push(NewUnit::Code { font, code }), + None => { + if !missing.contains(ch) { + missing.push(*ch); + } + } + } + } + if !missing.is_empty() { + let code = if target.face_changed { + P::FaceUnavailable + } else { + P::GlyphMissing + }; + let mut p = problem(code, "characters the font can't draw"); + p.chars = missing; + p.face = target.face.filter(|_| target.face_changed); + return Err(p); + } + let mut out: Vec<(NewUnit, Region)> = (0..prefix_end) + .map(|i| (NewUnit::Kept(i), Region::Prefix)) + .collect(); + let middle_empty = encoded.is_empty(); + out.extend(encoded.into_iter().map(|u| (u, Region::Middle))); + if middle_empty && !target.face_changed { + // The two kept sides meet: kerns between them stay only when nothing was removed there. + let between = units.get(prefix_end..suffix_start).unwrap_or_default(); + if between.iter().all(|u| unit_text(u).is_none()) { + out.extend((prefix_end..suffix_start).map(|i| (NewUnit::Kept(i), Region::Prefix))); + } + } + out.extend((suffix_start..units.len()).map(|i| (NewUnit::Kept(i), Region::Suffix))); + Ok(out) +} diff --git a/src-tauri/src/pdf_engine/text_edit/rewrite/layout.rs b/src-tauri/src/pdf_engine/text_edit/rewrite/layout.rs new file mode 100644 index 0000000..d278067 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/rewrite/layout.rs @@ -0,0 +1,598 @@ +//! Planner steps 9–10 (SPEC §B.12): advances of the new unit list in the written numbers, the +//! column-gap absorption, the pen compensation, the f32 precision guard, the fit policy (§A.8), +//! and every expectation the re-walk checks: glyph origins, text, new ink box, caret offsets and +//! the render masks of the independent check. + +use super::diff::Region; +use super::masks::{stroke_pad, MaskFonts}; +use super::{problem, NewUnit, StyleTarget}; +use crate::pdf_engine::text_edit::encode::num; +use crate::pdf_engine::text_edit::fonts::{Code, FontModel, TypingSurface}; +use crate::pdf_engine::text_edit::geometry::{bbox_of, Matrix}; +use crate::pdf_engine::text_edit::lexer::Span; +use crate::pdf_engine::text_edit::limits::{ + COLUMN_GAP_EM, DRIFT_TOLERANCE_PT, F32_DRIFT_GUARD_PT, NEGLIGIBLE_KERN, OVERLAP_WARN_TOL_PT, + SYNTH_SPACE_EM, +}; +use crate::pdf_engine::text_edit::reasons::{EditProblem, EditProblemCode as P, TextWarningCode}; +use crate::pdf_engine::text_edit::runs::{PageModel, TextRun, Unit}; +use crate::pdf_engine::text_edit::verify::{decoded_text, TextItem}; +use crate::pdf_engine::text_edit::walker::ShowRecord; +use std::sync::Arc; + +const FALLBACK_ASCENT: f64 = 0.8; +const FALLBACK_DESCENT: f64 = -0.2; + +/// One unit of the new run, resolved. +#[derive(Debug, Clone)] +pub(crate) enum Planned { + Glyph { + code: Code, + /// Font resource name (empty for an ExtGState font). + font_res: Vec, + font_hash: u64, + model: Option>, + width1000: f64, + text: String, + /// Kept: (unit index, member, glyph index in that member's record). + kept: Option<(usize, usize, usize)>, + }, + Kern { + value: f64, + /// The number as written; `None` = the original TJ bytes at `original`. + written: Option, + original: Option, + synth: bool, + /// Kept: the unit index in the run. + kept: Option, + }, +} + +#[derive(Debug, Clone)] +pub(crate) struct PlannedUnit { + pub unit: Planned, + pub region: Region, +} + +/// Everything layout decides for one run. +pub(crate) struct Laid { + pub units: Vec, + /// The compensation number as written (`None` when negligible). + pub compensation: Option, + /// Per member (index 0 is the primary, always `None`): the absorbed member's TJ number. + pub absorbed_numbers: Vec>, + pub text: String, + pub glyph_origins: Vec<(f64, f64)>, + pub prefix_glyphs: usize, + pub suffix_glyphs: usize, + pub shift_user: (f64, f64), + pub unshifted_from: Option, + pub delta_pt: f64, + pub new_rect: [f64; 4], + pub caret_offsets: Vec, + pub mask_boxes: Vec<[f64; 4]>, + /// Per old glyph: the tightest box known to hold its ink (`MaskFonts::ink_box`). + pub old_ink_boxes: Vec<[f64; 4]>, + pub glyphs_changed: bool, + pub warnings: Vec, +} + +fn record<'m>(model: &'m PageModel, run: &TextRun, member: usize) -> Option<&'m ShowRecord> { + run.members + .get(member) + .and_then(|i| model.walk.records.get(*i)) +} + +/// Resolves the new unit list against the run, its records and the typing surface. +fn resolve( + model: &PageModel, + run: &TextRun, + surface: &TypingSurface, + new_units: &[(NewUnit, Region)], +) -> Result, EditProblem> { + let internal = || problem(P::EditVerifyFailed, "unit out of range"); + let mut out = Vec::with_capacity(new_units.len()); + for (nu, region) in new_units { + let unit = match nu { + NewUnit::Kept(i) => match run.units.get(*i).ok_or_else(internal)? { + Unit::Glyph { + member, + rec, + glyph, + code, + font_res, + font_hash, + text, + width1000, + .. + } => Planned::Glyph { + code: *code, + font_res: font_res.clone().unwrap_or_default(), + font_hash: *font_hash, + model: model.walk.records.get(*rec).and_then(|r| r.font.clone()), + width1000: *width1000, + text: text.clone(), + kept: Some((*i, *member, *glyph)), + }, + Unit::Kern { + value, + src, + synth_space, + } => match src { + crate::pdf_engine::text_edit::runs::KernSrc::Tj { span } => Planned::Kern { + value: *value, + written: None, + original: Some(span.clone()), + synth: *synth_space, + kept: Some(*i), + }, + crate::pdf_engine::text_edit::runs::KernSrc::Gap => { + let (s, v) = num(*value)?; + Planned::Kern { + value: v, + written: Some(s), + original: None, + synth: *synth_space, + kept: Some(*i), + } + } + }, + }, + NewUnit::Code { font, code } => { + let (name, m) = surface.fonts.get(*font).ok_or_else(internal)?; + let text = m + .text(*code) + .map(str::to_string) + .ok_or_else(|| problem(P::EditVerifyFailed, "encoded code has no text"))?; + Planned::Glyph { + code: *code, + font_res: name.clone(), + font_hash: m.content_hash, + model: Some(Arc::clone(m)), + width1000: m.width(*code), + text, + kept: None, + } + } + NewUnit::KernSpace(v) => { + let (s, v) = num(*v)?; + Planned::Kern { + value: v, + written: Some(s), + original: None, + synth: true, + kept: None, + } + } + }; + out.push(PlannedUnit { + unit, + region: *region, + }); + } + Ok(out) +} + +/// Text-space advance (Th excluded) of a unit with size `tf` and character spacing `tc`. +fn advance(u: &Planned, tf: f64, tc: f64, tw: f64) -> f64 { + match u { + Planned::Glyph { + code, width1000, .. + } => { + let word = code.len == 1 && code.value == 32; + width1000 / 1000.0 * tf + tc + if word { tw } else { 0.0 } + } + Planned::Kern { value, .. } => -value / 1000.0 * tf, + } +} + +/// The original advance of a run unit (the run's own size and spacing). +fn original_advance(u: &Unit, run: &TextRun) -> f64 { + match u { + Unit::Glyph { + code, width1000, .. + } => { + let word = code.len == 1 && code.value == 32; + width1000 / 1000.0 * run.tfs + run.tc + if word { run.tw } else { 0.0 } + } + Unit::Kern { value, .. } => -value / 1000.0 * run.tfs, + } +} + +/// `|f32(v) − v|` scaled to points by `scale` stays within the guard (§B.12 step 9). +fn f32_ok(v: f64, scale: f64) -> bool { + let err = (f64::from(v as f32) - v).abs(); + err * scale.abs() <= F32_DRIFT_GUARD_PT +} + +fn em_box(origin: (f64, f64), l: &Matrix, x: [f64; 2], y: [f64; 2]) -> [f64; 4] { + let mut pts = Vec::with_capacity(4); + for gx in x { + for gy in y { + pts.push(( + origin.0 + gx * l[0] + gy * l[2], + origin.1 + gx * l[1] + gy * l[3], + )); + } + } + bbox_of(&pts).unwrap_or([origin.0, origin.1, origin.0, origin.1]) +} + +fn metrics(model: Option<&Arc>) -> (f64, f64) { + model.map_or((FALLBACK_ASCENT, FALLBACK_DESCENT), |m| { + (m.ascent, m.descent) + }) +} + +/// `b` (user space) grown by `pad` on every side. +fn grow(b: [f64; 4], pad: f64) -> [f64; 4] { + [b[0] - pad, b[1] - pad, b[2] + pad, b[3] + pad] +} + +/// The glyph-space → user linear map of `rec` scaled to size `tf`. +fn scaled_linear(rec: &ShowRecord, tf: f64) -> Matrix { + let t = rec.text_to_user; + let k = if rec.before.text.tfs != 0.0 { + tf / rec.before.text.tfs + } else { + 1.0 + }; + [t[0] * k, t[1] * k, t[2] * k, t[3] * k, 0.0, 0.0] +} + +/// Column-gap absorption (§B.12 step 9): the first kept-suffix kern of at least one em takes up +/// the width change, provided it still reads as a space. Returns the planned index it adjusted. +fn absorb_column_gap( + units: &mut [PlannedUnit], + run: &TextRun, + target: &StyleTarget, +) -> Result, EditProblem> { + if target.face_changed { + return Ok(None); + } + let (tf, tc, tw) = (target.tfs, target.tc, run.tw); + let found = units.iter().position(|p| { + let gap = |value: f64| -value / 1000.0 >= COLUMN_GAP_EM; + p.region == Region::Suffix + && matches!(p.unit, Planned::Kern { value, kept: Some(_), .. } if gap(value)) + }); + let Some(idx) = found else { + return Ok(None); + }; + let (value, kept) = match units.get(idx).map(|p| &p.unit) { + Some(Planned::Kern { + value, + kept: Some(k), + .. + }) => (*value, *k), + _ => return Ok(None), + }; + let orig_before: f64 = run + .units + .get(..kept) + .unwrap_or_default() + .iter() + .map(|u| original_advance(u, run)) + .sum(); + let new_before: f64 = units + .get(..idx) + .unwrap_or_default() + .iter() + .map(|p| advance(&p.unit, tf, tc, tw)) + .sum(); + if tf == 0.0 { + return Ok(None); + } + let adjusted = (new_before - orig_before) * 1000.0 / tf + value * run.tfs / tf; + let (s, v) = num(adjusted)?; + if -v / 1000.0 < SYNTH_SPACE_EM { + return Ok(None); + } + if let Some(p) = units.get_mut(idx) { + if let Planned::Kern { + value, + written, + original, + .. + } = &mut p.unit + { + *value = v; + *written = Some(s); + *original = None; + } + } + Ok(Some(idx)) +} + +/// Steps 9–10 for one run. +pub(super) fn lay_out( + masks: &mut MaskFonts<'_>, + model: &PageModel, + run: &TextRun, + surface: &TypingSurface, + target: &StyleTarget, + new_units: &[(NewUnit, Region)], +) -> Result { + let primary = + record(model, run, 0).ok_or_else(|| problem(P::EditVerifyFailed, "run without members"))?; + let mut units = resolve(model, run, surface, new_units)?; + let absorbed_at = absorb_column_gap(&mut units, run, target)?; + let (tf, tc, tw, th, ttux) = (target.tfs, target.tc, run.tw, run.th, run.text_to_user_x); + if tf == 0.0 || run.tfs == 0.0 { + return Err(problem(P::EditVerifyFailed, "zero size")); + } + let precision = || problem(P::EditVerifyFailed, "number precision"); + // Positions (text space, Th included) and the total advance. + let mut x = Vec::with_capacity(units.len() + 1); + let mut at = 0.0f64; + for p in &units { + x.push(at); + at += advance(&p.unit, tf, tc, tw) * th; + } + let advance_new: f64 = units.iter().map(|p| advance(&p.unit, tf, tc, tw)).sum(); + let n_c = (advance_new - primary.advance_ts) * 1000.0 / tf; + let compensation = if n_c.abs() < NEGLIGIBLE_KERN { + None + } else { + Some(num(n_c)?) + }; + let kern_scale = tf * th * ttux / 1000.0; + let mut glyph_count = 0usize; + let mut em_sum = 0.0f64; + for p in &units { + match &p.unit { + Planned::Kern { + value, + written: Some(_), + .. + } => { + if !f32_ok(*value, kern_scale) { + return Err(precision()); + } + em_sum += value.abs() / 1000.0; + } + Planned::Kern { value, .. } => em_sum += value.abs() / 1000.0, + Planned::Glyph { width1000, .. } => { + glyph_count += 1; + em_sum += width1000.abs() / 1000.0; + } + } + } + if compensation + .as_ref() + .is_some_and(|(_, v)| !f32_ok(*v, kern_scale)) + { + return Err(precision()); + } + if target.tc_changed && !f32_ok(tc, glyph_count as f64 * th * ttux) { + return Err(precision()); + } + if target.size_changed && !f32_ok(tf, em_sum * th * ttux) { + return Err(precision()); + } + let mut absorbed_numbers = vec![None]; + for m in 1..run.members.len() { + let rec = record(model, run, m) + .ok_or_else(|| problem(P::EditVerifyFailed, "member out of range"))?; + let tfs_j = rec.before.text.tfs; + if tfs_j == 0.0 { + return Err(precision()); + } + let n_j = -rec.advance_ts * 1000.0 / tfs_j; + if n_j.abs() < NEGLIGIBLE_KERN { + absorbed_numbers.push(None); + } else { + let (s, v) = num(n_j)?; + if !f32_ok(v, tfs_j * rec.before.text.th * ttux / 1000.0) { + return Err(precision()); + } + absorbed_numbers.push(Some(s)); + } + } + // Glyph expectations. + let origin = run.origin; + let step = (run.dir.0 * ttux, run.dir.1 * ttux); + let l_new = scaled_linear(primary, tf); + let mut glyph_origins = Vec::new(); + let mut prefix_glyphs = 0usize; + let mut suffix_glyphs = 0usize; + let mut shift_user = None; + let mut unshifted_from = None; + let mut new_extent = 0.0f64; + let mut ink: Vec<(f64, f64)> = Vec::new(); + let mut mask_boxes = Vec::new(); + let new_glyphs = units.iter().filter_map(|p| match &p.unit { + Planned::Glyph { + model: Some(m), + code, + .. + } => Some((m, *code)), + _ => None, + }); + let old_glyphs = (0..run.members.len()) + .filter_map(|m| record(model, run, m)) + .flat_map(|rec| { + rec.font + .iter() + .flat_map(move |f| rec.glyphs.iter().map(move |g| (f, g.code))) + }); + masks.prepare(new_glyphs.chain(old_glyphs)); + let new_pad = stroke_pad(&primary.before); + for (k, p) in units.iter().enumerate() { + let xk = x.get(k).copied().unwrap_or(0.0); + let Planned::Glyph { + code, + width1000, + model: font, + kept, + .. + } = &p.unit + else { + continue; + }; + let o = (origin.0 + xk * step.0, origin.1 + xk * step.1); + let w0 = width1000 / 1000.0; + if absorbed_at.is_some_and(|a| k > a) && unshifted_from.is_none() { + unshifted_from = Some(glyph_origins.len()); + } + match p.region { + Region::Prefix => prefix_glyphs += 1, + Region::Suffix => { + suffix_glyphs += 1; + if shift_user.is_none() { + if let Some((_, member, gi)) = kept { + if let Some(g) = record(model, run, *member).and_then(|r| r.glyphs.get(*gi)) + { + shift_user = Some((o.0 - g.origin.0, o.1 - g.origin.1)); + } + } + } + } + Region::Middle => {} + } + glyph_origins.push(o); + new_extent = new_extent.max((xk + w0 * tf * th) * ttux); + let (asc, desc) = metrics(font.as_ref()); + let b = em_box(o, &l_new, [0.0, w0], [desc, asc]); + ink.extend([(b[0], b[1]), (b[2], b[3])]); + let gb = masks.glyph_box(font.as_ref(), *code, w0); + let b = em_box(o, &l_new, [gb[0], gb[2]], [gb[1], gb[3]]); + mask_boxes.push(grow(b, new_pad)); + } + // Old glyphs of every member are masked too. + let mut old_ink_boxes = Vec::new(); + for m in 0..run.members.len() { + let Some(rec) = record(model, run, m) else { + continue; + }; + let l = scaled_linear(rec, rec.before.text.tfs); + let pad = stroke_pad(&rec.before); + for g in &rec.glyphs { + let (font, w0) = (rec.font.as_ref(), g.width1000 / 1000.0); + let gb = masks.glyph_box(font, g.code, w0); + let b = em_box(g.origin, &l, [gb[0], gb[2]], [gb[1], gb[3]]); + mask_boxes.push(grow(b, pad)); + let ib = masks.ink_box(font, g.code, w0); + let b = em_box(g.origin, &l, [ib[0], ib[2]], [ib[1], ib[3]]); + old_ink_boxes.push(grow(b, pad)); + } + } + let (asc, desc) = metrics(primary.font.as_ref()); + if ink.is_empty() { + let b = em_box(origin, &l_new, [0.0, 0.0], [desc, asc]); + ink.extend([(b[0], b[1]), (b[2], b[3])]); + } + let rect = bbox_of(&ink).unwrap_or([origin.0, origin.1, origin.0, origin.1]); + let new_rect = [rect[0], rect[1], rect[2] - rect[0], rect[3] - rect[1]]; + // Fit (§A.8). + if new_extent > run.visible_extent.max(run.original_extent) + DRIFT_TOLERANCE_PT { + return Err(problem( + P::TextOutsideVisibleArea, + format!( + "extent {new_extent:.3} pt > visible {:.3} pt", + run.visible_extent + ), + )); + } + let warnings = run + .next_obstacle + .filter(|o| new_extent > o + OVERLAP_WARN_TOL_PT) + .map(|_| TextWarningCode::NextTextOverlap) + .into_iter() + .collect(); + let (text, caret_offsets) = text_and_carets(&units, &x, tf, tc, tw, th, ttux); + let items = units.iter().map(|p| match &p.unit { + Planned::Glyph { text, .. } => TextItem::Glyph(text.as_str()), + Planned::Kern { value, .. } => TextItem::Kern(-value / 1000.0), + }); + if decoded_text(items) != text { + return Err(problem( + P::EditVerifyFailed, + "the new line would not read back as typed", + )); + } + Ok(Laid { + glyphs_changed: glyphs_changed(run, &units), + units, + compensation: compensation.map(|(s, _)| s), + absorbed_numbers, + text, + glyph_origins, + prefix_glyphs, + suffix_glyphs, + shift_user: shift_user.unwrap_or((0.0, 0.0)), + unshifted_from, + delta_pt: new_extent - run.original_extent, + new_rect, + caret_offsets, + mask_boxes, + old_ink_boxes, + warnings, + }) +} + +/// The new text (synthetic spaces as " ") and its chars + 1 caret offsets along `dir`. +fn text_and_carets( + units: &[PlannedUnit], + x: &[f64], + tf: f64, + tc: f64, + tw: f64, + th: f64, + ttux: f64, +) -> (String, Vec) { + let mut text = String::new(); + let mut offsets = Vec::new(); + let mut end = 0.0f64; + for (k, p) in units.iter().enumerate() { + let start = x.get(k).copied().unwrap_or(0.0) * ttux; + let adv = advance(&p.unit, tf, tc, tw) * th * ttux; + match &p.unit { + Planned::Glyph { text: t, .. } => { + let n = t.chars().count().max(1); + for j in 0..t.chars().count() { + offsets.push(start + adv * j as f64 / n as f64); + } + text.push_str(t); + end = start + adv; + } + Planned::Kern { synth: true, .. } => { + offsets.push(start); + text.push(' '); + end = start + adv; + } + Planned::Kern { .. } => {} + } + } + offsets.push(end); + let mut max = f64::NEG_INFINITY; + for o in &mut offsets { + max = max.max(*o); + *o = max; + } + (text, offsets) +} + +/// Whether the non-space glyphs (font and code) differ between the run and the new units. +fn glyphs_changed(run: &TextRun, units: &[PlannedUnit]) -> bool { + let space = |t: &str| matches!(t, " " | "\u{a0}"); + let old = run.units.iter().filter_map(|u| match u { + Unit::Glyph { + code, + font_hash, + text, + .. + } if !space(text) => Some((*font_hash, *code)), + _ => None, + }); + let new = units.iter().filter_map(|p| match &p.unit { + Planned::Glyph { + code, + font_hash, + text, + .. + } if !space(text) => Some((*font_hash, *code)), + _ => None, + }); + !old.eq(new) +} diff --git a/src-tauri/src/pdf_engine/text_edit/rewrite/masks.rs b/src-tauri/src/pdf_engine/text_edit/rewrite/masks.rs new file mode 100644 index 0000000..6dc8282 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/rewrite/masks.rs @@ -0,0 +1,392 @@ +//! A5 render masks (SPEC §B.15): the glyph boxes, in em, that G-RENDER excludes from its pixel +//! comparison, one per old and new glyph of an edited run. +//! +//! Per glyph, the box is the program's own outline box when the font program can outline the +//! glyph (TrueType `glyf`, CFF, OpenType; the same bounded pre-check as T2's presence proof runs +//! first), its units scale is unambiguous (`cff_scale.rs`: a CFF program's Font DICT matrices) +//! and the box lies within [−2, 3] em; otherwise it is built from the font's metrics. The +//! fallback takes its height from the +//! ascent, the descent and `/FontBBox`, but its width only from the glyph's own advance plus a +//! bounded overhang: `/FontBBox` is font-wide (Arial's reaches 2 em right of every origin), and a +//! mask that wide would hide a follower on the same line that a width-model error moved (B13). +//! Stroked text (`Tr` 1/2) grows each box by half the line width (times the miter limit for +//! miter joins) in user space. + +use super::cff_scale::{em_per_unit, Host}; +use crate::pdf_engine::text_edit::context::resolve; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget}; +use crate::pdf_engine::text_edit::fonts::glyph_budget::WorkMeter; +use crate::pdf_engine::text_edit::fonts::program::Outlines; +use crate::pdf_engine::text_edit::fonts::{number_of, Code, FontKey, FontModel}; +use crate::pdf_engine::text_edit::limits::{FONT_PROGRAM_MAX_DECODED, PAGE_DECODE_BUDGET}; +use crate::pdf_engine::text_edit::state::StateDigest; +use lopdf::{Dictionary, Document, Object, Stream}; +use std::collections::{BTreeSet, HashMap}; +use std::sync::Arc; +use ttf_parser::{Face, GlyphId, OutlineBuilder, Rect, Tag}; + +/// Ascent/descent (em) when the font gives none. +const FALLBACK_ASCENT: f64 = 0.8; +const FALLBACK_DESCENT: f64 = -0.2; +/// The em box of §B.15 when nothing better is known: its height bounds every fallback mask. +const MASK_EM: [f64; 4] = [-0.1, -0.3, 1.1, 1.0]; +/// Fallback margin around the advance and the ascent/descent (em). +const MASK_MARGIN_EM: f64 = 0.1; +/// Largest fallback overhang past the advance (or before the origin) of an upright glyph (em). +const UPRIGHT_REACH_EM: f64 = 0.2; +/// Largest fallback overhang of an italic glyph (em). +const ITALIC_REACH_EM: f64 = 0.35; +/// Bounds of any font or glyph box used for a mask (em). A program glyph box beyond them is not +/// taken (a wrong units scale or a hostile program: the fallback is used instead of a clamped +/// box, which could be empty or several em wide); `/FontBBox` values are clamped to them, as they +/// only set the fallback's height. +const BOX_EM: [f64; 2] = [-2.0, 3.0]; +/// Largest miter factor applied to a stroked glyph's half line width. +const MITER_FACTOR_MAX: f64 = 10.0; + +/// Ignores the outline; ttf-parser computes the box while it walks it. +struct NoPath; + +impl OutlineBuilder for NoPath { + fn move_to(&mut self, _x: f32, _y: f32) {} + fn line_to(&mut self, _x: f32, _y: f32) {} + fn quad_to(&mut self, _x1: f32, _y1: f32, _x: f32, _y: f32) {} + fn curve_to(&mut self, _x1: f32, _y1: f32, _x2: f32, _y2: f32, _x: f32, _y: f32) {} + fn close(&mut self) {} +} + +#[derive(Clone, Copy, PartialEq, Eq)] +enum ProgramKind { + /// `FontFile2`, or `FontFile3 /OpenType`. + Sfnt, + /// `FontFile3 /Type1C` or `/CIDFontType0C`. + BareCff, +} + +/// What one font contributes to masks. +#[derive(Default)] +struct FontBoxes { + /// `/FontBBox` in em (clamped). + bbox: Option<[f64; 4]>, + /// The decoded program, when it is one we can outline (`None` once known unusable). + program: Option<(ProgramKind, Vec)>, + /// Outline boxes per GID in em (`None`: the program could not outline it). + glyphs: HashMap>, +} + +/// Glyph boxes of the fonts of one page plan, loaded on first use. Program decoding and outline +/// work are bounded by one page decode budget; past it every glyph uses the fallback box. +pub(crate) struct MaskFonts<'d> { + doc: &'d Document, + budget: DecodeBudget, + meter: WorkMeter, + fonts: HashMap, +} + +fn dict_of<'d>(doc: &'d Document, obj: &'d Object) -> Option<&'d Dictionary> { + match resolve(doc, obj)? { + (_, Object::Dictionary(d)) => Some(d), + _ => None, + } +} + +/// The dictionary holding the font's `/FontDescriptor`: the font's own, or for a Type0 font its +/// descendant's. +fn described_font<'d>(doc: &'d Document, key: FontKey) -> Option<&'d Dictionary> { + let FontKey::Indirect(id) = key else { + return None; + }; + let font = dict_of(doc, doc.objects.get(&id)?)?; + if font.get(b"Subtype").ok().and_then(|o| o.as_name().ok()) != Some(&b"Type0"[..]) { + return Some(font); + } + let (_, Object::Array(kids)) = resolve(doc, font.get(b"DescendantFonts").ok()?)? else { + return None; + }; + dict_of(doc, kids.first()?) +} + +fn descriptor<'d>(doc: &'d Document, key: FontKey) -> Option<&'d Dictionary> { + dict_of(doc, described_font(doc, key)?.get(b"FontDescriptor").ok()?) +} + +/// The `/FontBBox` (em) of the font behind `key`: its descriptor's, or for a Type0 font its +/// descendant's. `None` when absent or not four finite numbers. +pub(crate) fn font_bbox(doc: &Document, key: FontKey) -> Option<[f64; 4]> { + let (_, Object::Array(items)) = resolve(doc, descriptor(doc, key)?.get(b"FontBBox").ok()?)? + else { + return None; + }; + let v: Vec = items + .iter() + .map(|i| resolve(doc, i).and_then(|(_, o)| number_of(o))) + .collect::>>()?; + let [a, b, c, d] = v.as_slice() else { + return None; + }; + let em = |n: f64| (n / 1000.0).clamp(BOX_EM[0], BOX_EM[1]); + Some([em(a.min(*c)), em(b.min(*d)), em(a.max(*c)), em(b.max(*d))]) +} + +/// The program stream of the font behind `key` when it is one ttf-parser outlines. +fn program_stream(doc: &Document, key: FontKey) -> Option<(ProgramKind, &Stream)> { + let desc = descriptor(doc, key)?; + let stream = |k: &[u8]| match resolve(doc, desc.get(k).ok()?)? { + (_, Object::Stream(s)) => Some(s), + _ => None, + }; + if let Some(s) = stream(b"FontFile2") { + return Some((ProgramKind::Sfnt, s)); + } + let s = stream(b"FontFile3")?; + match s.dict.get(b"Subtype").ok().and_then(|o| o.as_name().ok()) { + Some(b"OpenType") => Some((ProgramKind::Sfnt, s)), + Some(b"Type1C" | b"CIDFontType0C") => Some((ProgramKind::BareCff, s)), + _ => None, + } +} + +/// A glyph's outline box in em, or `None` when any edge lies outside `BOX_EM`. +fn em_rect(r: Rect, scale: f64) -> Option<[f64; 4]> { + let b = [r.x_min, r.y_min, r.x_max, r.y_max].map(|n| f64::from(n) * scale); + let inside = b.iter().all(|v| (BOX_EM[0]..=BOX_EM[1]).contains(v)); + (r.x_min <= r.x_max && r.y_min <= r.y_max && inside).then_some(b) +} + +/// Outline boxes (em) of `gids` in `program`; a glyph the pre-check or ttf-parser refuses, or an +/// unusable program, gives `None`. +fn outline_boxes( + kind: ProgramKind, + data: &[u8], + gids: &[u16], + meter: &mut WorkMeter, +) -> Vec<(u16, Option<[f64; 4]>)> { + let none = || gids.iter().map(|g| (*g, None)).collect(); + let face; + let (outlines, scale) = match kind { + ProgramKind::Sfnt => { + let Ok(f) = Face::parse(data, 0) else { + return none(); + }; + face = f; + let upem = face.units_per_em(); + let outlines = Outlines::of_face(&face); + // A `CFF ` table carries its own matrices; `glyf` units are 1/unitsPerEm. + let scale = match &outlines { + Outlines::Cff { .. } => face + .raw_face() + .table(Tag::from_bytes(b"CFF ")) + .and_then(|cff| em_per_unit(cff, Host::OpenType(upem))), + _ => (upem > 0).then(|| 1.0 / f64::from(upem)), + }; + let Some(scale) = scale else { + return none(); + }; + (outlines, scale) + } + ProgramKind::BareCff => { + let (Some(outlines), Some(scale)) = + (Outlines::of_cff(data), em_per_unit(data, Host::Bare)) + else { + return none(); + }; + (outlines, scale) + } + }; + gids.iter() + .map(|&gid| { + let rect = match &outlines { + Outlines::Glyf { guard, table } if guard.check(gid, meter) => { + table.outline(GlyphId(gid), &mut NoPath) + } + Outlines::Cff { guard, table } if guard.check(gid, meter) => { + table.outline(GlyphId(gid), &mut NoPath).ok() + } + _ => None, + }; + (gid, rect.and_then(|r| em_rect(r, scale))) + }) + .collect() +} + +/// The entry of `key`, created on first use (its `/FontBBox`, and its program decoded under +/// `budget` when ttf-parser can outline it). +fn font_entry<'f>( + doc: &Document, + budget: &mut DecodeBudget, + fonts: &'f mut HashMap, + key: FontKey, +) -> &'f mut FontBoxes { + fonts.entry(key).or_insert_with(|| FontBoxes { + bbox: font_bbox(doc, key), + program: program_stream(doc, key).and_then(|(kind, s)| { + decode_stream(s, FONT_PROGRAM_MAX_DECODED, budget) + .ok() + .map(|data| (kind, data)) + }), + glyphs: HashMap::new(), + }) +} + +/// The fallback box (em) of a glyph of advance `w0` (§B.15 "else `/FontBBox`, else the em +/// box", with the horizontal reach bounded, see the module doc). +fn fallback_box(model: Option<&FontModel>, bbox: Option<[f64; 4]>, w0: f64) -> [f64; 4] { + let (ascent, descent) = model.map_or((FALLBACK_ASCENT, FALLBACK_DESCENT), |m| { + (m.ascent, m.descent) + }); + let italic = model.is_some_and(|m| m.italic); + let max_reach = if italic { + ITALIC_REACH_EM + } else { + UPRIGHT_REACH_EM + }; + let unknown = if italic { max_reach } else { MASK_MARGIN_EM }; + let (left, right) = match bbox { + Some(b) => ( + (-b[0]).clamp(MASK_MARGIN_EM, max_reach), + (b[2] - w0.max(0.0)).clamp(MASK_MARGIN_EM, max_reach), + ), + None => (unknown, unknown), + }; + let mut y = [ + MASK_EM[1].min(descent - MASK_MARGIN_EM), + MASK_EM[3].max(ascent + MASK_MARGIN_EM), + ]; + if let Some(b) = bbox { + y = [y[0].min(b[1]), y[1].max(b[3])]; + } + let y = [y[0].max(BOX_EM[0]), y[1].min(BOX_EM[1])]; + [w0.min(0.0) - left, y[0], w0.max(0.0) + right, y[1]] +} + +/// Half the line width of stroked text (`Tr` 1, 2) in user space, times the miter limit for +/// miter joins (≤ `MITER_FACTOR_MAX`); 0 for filled text. +pub(crate) fn stroke_pad(state: &StateDigest) -> f64 { + if !matches!(state.text.tr, 1 | 2) { + return 0.0; + } + let gs = &state.gs; + // The CTM's largest singular value: how far a user-space line width reaches on the page. + let c = state.ctm; + let sum = c[0] * c[0] + c[1] * c[1] + c[2] * c[2] + c[3] * c[3]; + let det = c[0] * c[3] - c[1] * c[2]; + let scale = ((sum + (sum * sum - 4.0 * det * det).max(0.0).sqrt()) / 2.0).sqrt(); + let miter = if gs.line_join == 0 { + gs.miter_limit.clamp(1.0, MITER_FACTOR_MAX) + } else { + 1.0 + }; + let pad = gs.line_width.abs() / 2.0 * miter * scale; + if pad.is_finite() { + pad + } else { + 0.0 + } +} + +impl<'d> MaskFonts<'d> { + pub(crate) fn new(doc: &'d Document) -> MaskFonts<'d> { + MaskFonts { + doc, + budget: DecodeBudget::new(PAGE_DECODE_BUDGET), + meter: WorkMeter::new(PAGE_DECODE_BUDGET), + fonts: HashMap::new(), + } + } + + /// Outlines, once per font and GID, every glyph of `glyphs` whose font program ttf-parser + /// can draw (fonts in first-seen order, so the budget is spent deterministically). + pub(crate) fn prepare<'m>( + &mut self, + glyphs: impl IntoIterator, Code)>, + ) { + let mut order: Vec = Vec::new(); + let mut wanted: HashMap> = HashMap::new(); + for (model, code) in glyphs { + let Some(gid) = model.info(code).and_then(|i| i.gid) else { + continue; + }; + let gids = wanted.entry(model.key).or_insert_with(|| { + order.push(model.key); + BTreeSet::new() + }); + gids.insert(gid); + } + for key in order { + let Some(gids) = wanted.remove(&key) else { + continue; + }; + let font = font_entry(self.doc, &mut self.budget, &mut self.fonts, key); + let gids: Vec = gids + .into_iter() + .filter(|g| !font.glyphs.contains_key(g)) + .collect(); + let Some((kind, data)) = font.program.as_ref() else { + continue; + }; + if gids.is_empty() { + continue; + } + for (gid, b) in outline_boxes(*kind, data, &gids, &mut self.meter) { + font.glyphs.insert(gid, b); + } + } + } + + /// The mask box (em, `[x0, y0, x1, y1]`) of `code` drawn with `font`, advance `w0` em: the + /// program's outline box when `prepare` found one, else the bounded fallback. + pub(crate) fn glyph_box( + &mut self, + font: Option<&Arc>, + code: Code, + w0: f64, + ) -> [f64; 4] { + self.boxes(font, code, w0).0 + } + + /// The tightest box (em) known to hold the glyph's ink: the program's outline box when + /// `prepare` found one, else the fallback narrowed to the advance (`advance_box`). A5 checks + /// the pixels of the glyphs no edit changes outside it (review-verify HIGH-A): the fallback's + /// margin would hide a narrow follower drawn right after a removed glyph. + pub(crate) fn ink_box( + &mut self, + font: Option<&Arc>, + code: Code, + w0: f64, + ) -> [f64; 4] { + self.boxes(font, code, w0).1 + } + + /// (`glyph_box`, `ink_box`). + fn boxes( + &mut self, + font: Option<&Arc>, + code: Code, + w0: f64, + ) -> ([f64; 4], [f64; 4]) { + let Some(model) = font else { + let b = fallback_box(None, None, w0); + return (b, advance_box(b, None, w0)); + }; + let gid = model.info(code).and_then(|i| i.gid); + let entry = font_entry(self.doc, &mut self.budget, &mut self.fonts, model.key); + match gid.and_then(|g| entry.glyphs.get(&g).copied().flatten()) { + Some(b) => (b, b), + None => { + let b = fallback_box(Some(model), entry.bbox, w0); + (b, advance_box(b, Some(model), w0)) + } + } + } +} + +/// The fallback box `b` (em) narrowed to the advance `w0`, wider by `ITALIC_REACH_EM` on each side +/// in an italic font; its height stays the fallback's. +fn advance_box(b: [f64; 4], model: Option<&FontModel>, w0: f64) -> [f64; 4] { + let reach = if model.is_some_and(|m| m.italic) { + ITALIC_REACH_EM + } else { + 0.0 + }; + [w0.min(0.0) - reach, b[1], w0.max(0.0) + reach, b[3]] +} diff --git a/src-tauri/src/pdf_engine/text_edit/rewrite/style.rs b/src-tauri/src/pdf_engine/text_edit/rewrite/style.rs new file mode 100644 index 0000000..deefcab --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/rewrite/style.rs @@ -0,0 +1,255 @@ +//! Planner steps 1–5 (SPEC §B.12): resolve the run, validate the typed text (§A.3.3) and the style +//! values, drop style fields equal to the current value (B7), check that each requested change is +//! available (B8, B9), and compute the written targets (`Tf'`, `Tc'`, colour), each read back from +//! the number that will be written. + +use super::{bad_edit, problem, SourceTextStyleIn, StyleTarget, TextEditIn}; +use crate::error::AppError; +use crate::pdf_engine::text_edit::encode::num; +use crate::pdf_engine::text_edit::fonts::{face_surface, FontModel, TypingSurface}; +use crate::pdf_engine::text_edit::lexer::{lex_content, LexLimits, Operand, Operator}; +use crate::pdf_engine::text_edit::limits::{ + EDIT_TEXT_CHARS_MAX, LETTER_SPACING_MAX_PT, LETTER_SPACING_MIN_PT, SIZE_MAX_PT, SIZE_MIN_PT, + STYLE_EPSILON, +}; +use crate::pdf_engine::text_edit::reasons::{EditProblem, EditProblemCode as P, StyleField}; +use crate::pdf_engine::text_edit::runs::{PageModel, SpaceMode, TextRun}; +use crate::pdf_engine::text_edit::walker::ShowRecord; +use std::sync::Arc; + +/// A written size must reproduce the requested effective size within this (pt). +const SIZE_READBACK_TOL_PT: f64 = 0.005; + +/// Step 1: the run named by the edit, unchanged since it was read, and editable. +pub(super) fn resolve<'m>( + model: &'m PageModel, + edit: &TextEditIn, +) -> Result<&'m TextRun, EditProblem> { + if let Some(reason) = model.page_reason { + let mut p = problem(P::TextEditRefused, "page refused"); + p.reason = Some(reason); + return Err(p); + } + let run = model + .run(&edit.run_id) + .ok_or_else(|| problem(P::Stale, "line not found on the page"))?; + if run.text != edit.original_text { + return Err(problem( + P::Stale, + "the line's text differs from the request", + )); + } + if let Some(reason) = run.reason { + let mut p = problem(P::TextEditRefused, "line refused"); + p.reason = Some(reason); + return Err(p); + } + Ok(run) +} + +/// §A.3.3: no line breaks, tabs, C0/C1 controls or line/paragraph separators; ≤ 1,000 chars. +pub(super) fn validate_text(text: &str) -> Result<(), EditProblem> { + let invalid = text.chars().any(|c| { + let v = u32::from(c); + v < 0x20 || (0x7f..=0x9f).contains(&v) || c == '\u{2028}' || c == '\u{2029}' + }); + if invalid { + return Err(problem(P::InvalidText, "control character")); + } + if text.chars().count() > EDIT_TEXT_CHARS_MAX { + return Err(problem(P::TextTooLong, "more than 1,000 characters")); + } + Ok(()) +} + +/// `#rrggbb` → bytes. +pub(crate) fn parse_hex(s: &str) -> Option<[u8; 3]> { + let digits = s.strip_prefix('#')?; + if digits.len() != 6 || !digits.bytes().all(|c| c.is_ascii_hexdigit()) { + return None; + } + let byte = |i: usize| u8::from_str_radix(digits.get(i..i + 2)?, 16).ok(); + Some([byte(0)?, byte(2)?, byte(4)?]) +} + +/// Step 2 (style half): malformed values are a malformed request (`BAD_EDIT`). +pub(super) fn validate_style(style: &SourceTextStyleIn) -> Result<(), AppError> { + if let Some(v) = style.size_pt { + if !v.is_finite() || !(SIZE_MIN_PT..=SIZE_MAX_PT).contains(&v) { + return Err(bad_edit(&format!("sizePt {v} outside 4–144"))); + } + } + if let Some(v) = style.letter_spacing_pt { + if !v.is_finite() || !(LETTER_SPACING_MIN_PT..=LETTER_SPACING_MAX_PT).contains(&v) { + return Err(bad_edit(&format!("letterSpacingPt {v} outside -2–10"))); + } + } + if let Some(fill) = &style.fill { + if parse_hex(fill).is_none() { + return Err(bad_edit("fill is not #rrggbb")); + } + } + Ok(()) +} + +/// The run's current letter spacing in effective points. +pub(crate) fn letter_spacing_pt(run: &TextRun) -> f64 { + run.tc * run.th * run.text_to_user_x +} + +/// Step 3: fields equal to the current value (within `STYLE_EPSILON`) are dropped (B7). +pub(super) fn normalise(run: &TextRun, style: &SourceTextStyleIn) -> SourceTextStyleIn { + let size_pt = style + .size_pt + .filter(|v| (v - run.effective_size).abs() > STYLE_EPSILON); + let face = style.face.filter(|f| *f != run.face); + let fill = style.fill.as_ref().and_then(|f| { + let lower = f.to_ascii_lowercase(); + (run.fill_hex.as_deref() != Some(lower.as_str())).then_some(lower) + }); + let current = letter_spacing_pt(run); + let letter_spacing_pt = style + .letter_spacing_pt + .filter(|v| (v - current).abs() > STYLE_EPSILON); + SourceTextStyleIn { + size_pt, + face, + fill, + letter_spacing_pt, + } +} + +fn unavailable(field: StyleField) -> EditProblem { + let mut p = problem(P::StyleUnavailable, format!("{field:?} unavailable")); + p.field = Some(field); + p +} + +/// Step 4: size and face need a `Tf` font (B9); colour needs fill-only rendering; a face needs +/// a sibling group on the page that can draw the whole text (B8). Returns the face surface. +pub(super) fn availability( + run: &TextRun, + primary: Option<&FontModel>, + page_fonts: &[(Vec, Arc)], + wanted: &SourceTextStyleIn, + text: &str, +) -> Result, EditProblem> { + if wanted.size_pt.is_some() && run.font_from_extgstate { + return Err(unavailable(StyleField::Size)); + } + if wanted.face.is_some() && run.font_from_extgstate { + return Err(unavailable(StyleField::Face)); + } + if wanted.fill.is_some() && run.tr != 0 { + return Err(unavailable(StyleField::Colour)); + } + let Some(face) = wanted.face else { + return Ok(None); + }; + let face_unavailable = |chars: Vec| { + let mut p = problem(P::FaceUnavailable, format!("{face:?} face")); + p.face = Some(face); + p.chars = chars; + p + }; + let surface = primary + .and_then(|m| face_surface(page_fonts, m, face)) + .ok_or_else(|| face_unavailable(Vec::new()))?; + let mut missing: Vec = Vec::new(); + for ch in text.chars() { + if ch == ' ' && run.space_mode == SpaceMode::Kern { + continue; + } + if surface.writer_for(ch).is_none() && !missing.contains(&ch) { + missing.push(ch); + } + } + if !missing.is_empty() { + return Err(face_unavailable(missing)); + } + Ok(Some(surface)) +} + +/// The size operand of a verbatim `Tf` op (`/F1 11.04 Tf` → `11.04`). +fn tf_token(tf_op: &[u8]) -> Option> { + let ops = lex_content(tf_op, &LexLimits::page(), None).ok()?; + let [op] = ops.as_slice() else { + return None; + }; + if op.operator != Operator::Tf { + return None; + } + match op.operands.get(1)? { + Operand::Number { span, .. } => tf_op.get(span.clone()).map(<[u8]>::to_vec), + _ => None, + } +} + +/// Step 5: the targets. `Tf' = Tfs × size / effective`; `Tc' = spacing / (Th × text_to_user_x)`; +/// colour components `/255`; every value is the one read back from what is written. +pub(super) fn targets( + primary: &ShowRecord, + run: &TextRun, + wanted: &SourceTextStyleIn, + face_surface: Option, +) -> Result { + let precision = || problem(P::EditVerifyFailed, "number precision"); + let (tfs, tfs_token, effective_size, size_changed) = match wanted.size_pt { + Some(pt) => { + if run.effective_size <= 0.0 || run.tfs == 0.0 { + return Err(precision()); + } + let (text, value) = num(run.tfs * pt / run.effective_size)?; + let row2 = run.effective_size / run.tfs.abs(); + if value <= 0.0 || (value * row2 - pt).abs() > SIZE_READBACK_TOL_PT { + return Err(precision()); + } + (value, text.into_bytes(), pt, true) + } + None => { + let token = primary + .before + .text + .font + .as_ref() + .and_then(|f| f.tf_op.as_deref()) + .and_then(tf_token) + .unwrap_or_default(); + (run.tfs, token, run.effective_size, false) + } + }; + let (tc, tc_changed) = match wanted.letter_spacing_pt { + Some(pt) => { + let scale = run.th * run.text_to_user_x; + if scale == 0.0 || !scale.is_finite() { + return Err(precision()); + } + (num(pt / scale)?.1, true) + } + None => (run.tc, false), + }; + let fill = match &wanted.fill { + Some(hex) => { + let bytes = parse_hex(hex).ok_or_else(precision)?; + let mut rgb = [0.0; 3]; + for (slot, b) in rgb.iter_mut().zip(bytes) { + *slot = num(f64::from(b) / 255.0)?.1; + } + Some(rgb) + } + None => None, + }; + Ok(StyleTarget { + tfs, + tfs_token, + tc, + fill_changed: fill.is_some(), + fill, + face_changed: face_surface.is_some(), + face_surface, + size_changed, + tc_changed, + face: wanted.face, + effective_size, + }) +} diff --git a/src-tauri/src/pdf_engine/text_edit/runs.rs b/src-tauri/src/pdf_engine/text_edit/runs.rs new file mode 100644 index 0000000..610946c --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/runs.rs @@ -0,0 +1,533 @@ +//! Runs (SPEC §B.11): the editable or refused lines of a page. Every depth-0 show record gets its +//! own reasons (`runs/reasons.rs`); consecutive records join into one visual line (§A.1.1: same +//! baseline and direction, the full state equal except the font identity, identical or sibling +//! fonts, a gap within ±0.3 em, nothing painted between, the same first reason); each run gets its +//! units, synthetic spaces, text, caret offsets, ink box and extents (`runs/assemble.rs`), the +//! page-wide checks and a reading order (`order.rs`). Deterministic; page problems are data +//! (`page_reason`), never an `Err`. +//! +//! The run stage goes on charging the walk's page-model budget (`walker::budget`): its scratch +//! (reasons per record, groups, surfaces, the page-wide checks and reading order) and every run it +//! keeps, before each is made. Past the budget the page is `PAGE_TOO_COMPLEX` "page model size". + +mod assemble; +mod forms; +pub(crate) mod reasons; +mod surface; + +use crate::error::AppError; +use crate::pdf_engine::text_edit::content::{page_content, PageContent}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::fonts::{Code, TypingSurface}; +use crate::pdf_engine::text_edit::geometry::{ + apply, cross, display_rotation, dot, sub, transform_rect, unit, +}; +use crate::pdf_engine::text_edit::lexer::Span; +use crate::pdf_engine::text_edit::limits::{ + JOIN_BASELINE_TOL_PT, JOIN_GAP_EM, PAGE_DECODE_BUDGET, RUNS_PER_PAGE_MAX, RUN_MEMBERS_MAX, + STATE_EPSILON, XY_CUT_DEPTH_MAX, +}; +use crate::pdf_engine::text_edit::order::{reading_order, RunBox}; +use crate::pdf_engine::text_edit::reasons::{Face, TextReason}; +use crate::pdf_engine::text_edit::snapshot::Fingerprint; +use crate::pdf_engine::text_edit::state::{same_state_except, StateField}; +use crate::pdf_engine::text_edit::walker::budget::{content_bytes, ModelBudget}; +use crate::pdf_engine::text_edit::walker::{ + refused_walk, walk_page, PageWalk, ShowRecord, Stop, WalkMode, +}; +use reasons::{innermost_mcid, ReasonCtx}; +use std::mem::size_of; +use std::sync::atomic::{AtomicBool, Ordering}; +use std::sync::Arc; +use surface::Surfaces; + +pub(crate) use forms::form_lines; +pub use surface::surface_of; + +/// Scratch per run for the page-wide steps (duplicates, obstacles, run boxes and reading order: +/// a few indices and boxes each, plus one index per XY-cut level on the recursion path). +const RUN_SCRATCH_BYTES: usize = 256 + XY_CUT_DEPTH_MAX * 2 * size_of::(); + +#[derive(Debug, Clone)] +pub enum Unit { + Glyph { + member: usize, + rec: usize, + glyph: usize, + code: Code, + font_res: Option>, + font_hash: u64, + text: String, + width1000: f64, + }, + Kern { + value: f64, + src: KernSrc, + synth_space: bool, + }, +} + +/// `Tj` = an original TJ number (its bytes are kept verbatim); `Gap` = a converted gap between +/// two joined members. +#[derive(Debug, Clone)] +pub enum KernSrc { + Tj { span: Span }, + Gap, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, serde::Serialize)] +#[serde(rename_all = "camelCase")] +pub enum SpaceMode { + Glyph, + Kern, +} + +#[derive(Debug, Clone)] +pub struct TextRun { + /// `t1:{fp}:{page}:{s0}-{e0}[,{s1}-{e1}...]` (joined-buffer spans of the members). + pub id: String, + /// Reading order in display space and the line cluster it belongs to. + pub order: u32, + pub line: u32, + /// Record indices; `members[0]` is the primary. + pub members: Vec, + pub units: Vec, + pub text: String, + /// chars + 1 user-space distances along `dir` from `origin`. + pub caret_offsets: Vec, + pub origin: (f64, f64), + pub dir: (f64, f64), + pub up: (f64, f64), + /// User units from the baseline (rise included in the origin); `descent ≥ 0`. + pub ascent: f64, + pub descent: f64, + /// `[x, y, w, h]` ink AABB, unrotated user space. + pub rect: [f64; 4], + /// Font resources, primary first (empty for an ExtGState font, whose surface is that font). + /// Shared by every run with the same surface (one list per page and surface key). + pub surface: Arc<[Vec]>, + pub tfs: f64, + pub effective_size: f64, + pub tc: f64, + pub tw: f64, + pub th: f64, + /// `|row 1 of Tm × CTM|`: user units per unscaled text-space unit along the baseline. + pub text_to_user_x: f64, + pub space_mode: SpaceMode, + pub kern_space: f64, + /// Ink extent along `dir` from the origin. + pub original_extent: f64, + /// Distance to the edge of (visible box ∩ clip rectangle) along `dir`. + pub visible_extent: f64, + /// Distance to the next run on the same line (band overlap > 10 %). + pub next_obstacle: Option, + /// `#rrggbb` when readable. + pub fill_hex: Option, + pub face: Face, + pub font_from_extgstate: bool, + pub tr: i64, + pub substituted: bool, + /// `reasons[0]`. + pub reason: Option, + pub reasons: Vec, +} + +pub struct PageModel { + pub fingerprint: Fingerprint, + pub page_index: u32, + pub content: Arc, + pub walk: Arc, + /// Non-blank runs, in reading order. + pub runs: Vec, + pub page_reason: Option, + pub page_detail: Option, + /// Per walk record: the reason of the run it belongs to (blank runs included), or its own + /// first reason when it draws no glyph. + pub record_reason: Vec>, +} + +impl PageModel { + pub fn run(&self, id: &str) -> Option<&TextRun> { + self.runs.iter().find(|r| r.id == id) + } + + /// The typing surface of `run`: its ExtGState font alone, or the primary's sibling group — + /// the page fonts `run.surface` names, in its order (it holds no name for an ExtGState font). + pub fn surface(&self, run: &TextRun) -> TypingSurface { + let Some(primary) = run.members.first().and_then(|i| self.walk.records.get(*i)) else { + return TypingSurface { fonts: Vec::new() }; + }; + if run.surface.is_empty() { + return surface_of(&self.walk, primary); + } + let page_fonts = &self.walk.page_fonts; + let fonts = run + .surface + .iter() + .filter_map(|name| page_fonts.iter().find(|(n, _)| n == name).cloned()) + .collect(); + TypingSurface { fonts } + } +} + +/// The model of page `page_index` (`INVALID_PAGES` when there is none, `CANCELLED` when `cancel` +/// was set). Page problems are data: `page_reason` with no runs. +pub fn build_page_model( + ctx: &SnapshotContext, + page_index: u32, + cancel: Option<&AtomicBool>, +) -> Result { + let page_id = ctx.page_id(page_index)?; + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + let (content, walk) = match page_content(ctx.doc(), page_id, &mut budget) { + Ok(content) => { + let walk = walk_page(ctx, page_index, &content, WalkMode::Edit, cancel); + (content, walk) + } + Err(reason) => { + let empty = PageContent { + page_id, + parts: Vec::new(), + joined: Vec::new(), + contents_array: None, + }; + let walk = refused_walk(ctx, page_id, reason, "page content"); + (empty, walk) + } + }; + if cancel.is_some_and(|c| c.load(Ordering::Relaxed)) { + return Err(AppError::cancelled()); + } + Ok(model_of(ctx, page_index, content, walk)) +} + +/// The model of a page already walked in `Edit` mode (also used to re-model edited content in the +/// same context). A model keeps no lexed ops (nothing reads them once the walk is done; they were +/// the largest part of a dense page, review T3-budget MEDIUM-3), and a page refused after its +/// walk keeps no walk either (MEDIUM-1): only its content, geometry and reason. +pub fn model_of( + ctx: &SnapshotContext, + page_index: u32, + content: PageContent, + mut walk: PageWalk, +) -> PageModel { + walk.drop_ops(); + let out = build_runs(ctx, page_index, &content, &walk); + if out.release_walk { + walk.release(); + } + // From here on the walk's charge is the whole model's (`PageWalk::model_bytes`). + walk.model_bytes = out.held; + PageModel { + fingerprint: ctx.snap.fingerprint, + page_index, + content: Arc::new(content), + walk: Arc::new(walk), + runs: out.runs, + page_reason: out.page_reason, + page_detail: out.page_detail, + record_reason: out.record_reason, + } +} + +struct RunsOut { + held: usize, + runs: Vec, + page_reason: Option, + page_detail: Option, + record_reason: Vec>, + /// The page was refused after its walk: the model keeps no walk. + release_walk: bool, +} + +impl RunsOut { + /// A refused page: it keeps its content (`content_bytes`), no runs and no per-record reasons + /// (its walk holds no record: refused at page level, or released by `model_of`). + fn refused(content_bytes: usize, reason: TextReason, detail: String) -> RunsOut { + RunsOut { + held: content_bytes, + runs: Vec::new(), + page_reason: Some(reason), + page_detail: Some(detail), + record_reason: Vec::new(), + release_walk: true, + } + } +} + +fn build_runs( + ctx: &SnapshotContext, + page_index: u32, + content: &PageContent, + walk: &PageWalk, +) -> RunsOut { + // A refused page's model keeps only its content (a walk refused at page level kept nothing + // else; one refused later is released by `model_of`). + let kept = content_bytes(content); + if let Some(reason) = walk.page_reason { + let detail = walk.page_detail.clone().unwrap_or_default(); + return RunsOut::refused(kept, reason, detail); + } + if let Err(reason) = ctx.kids() { + return RunsOut::refused(kept, reason, "page tree".into()); + } + let mut mem = ModelBudget::new(walk.model_bytes); + runs_within(ctx, page_index, content, walk, &mut mem) + .unwrap_or_else(|stop| RunsOut::refused(kept, stop.reason, stop.detail)) +} + +/// The run stage under the page-model budget `mem`. +fn runs_within( + ctx: &SnapshotContext, + page_index: u32, + content: &PageContent, + walk: &PageWalk, + mem: &mut ModelBudget, +) -> Result { + let n = walk.records.len(); + let mut rc = ReasonCtx::new(ctx, content, walk); + mem.scratch(n.saturating_mul(size_of::>()))?; + let mut intrinsic: Vec> = Vec::with_capacity(n); + for r in &walk.records { + let reasons = rc.record_reasons(r); + mem.scratch(reasons.capacity())?; + intrinsic.push(reasons); + } + let mut surfaces = Surfaces::new(walk, mem)?; + let groups = join_records(walk, &intrinsic, &mut surfaces, mem)?; + if groups.len() > RUNS_PER_PAGE_MAX { + return Err(Stop::new(TextReason::PageTooComplex, "runs per page")); + } + mem.hold(groups.len().saturating_mul(size_of::()))?; + let mut all: Vec = Vec::with_capacity(groups.len()); + for g in &groups { + let mut reasons: Vec = g + .iter() + .filter_map(|i| intrinsic.get(*i)) + .flatten() + .copied() + .collect(); + reasons.sort(); + reasons.dedup(); + reasons.shrink_to_fit(); + let Some(primary) = g.first().and_then(|i| walk.records.get(*i)) else { + continue; + }; + // Whether a surface types anything is decided once per surface (a page may hold + // thousands of runs in one CJK font whose alphabet has tens of thousands of characters). + let surface = surfaces.of(primary, mem)?; + let mut run = assemble::assemble(walk, g, reasons, &surface, mem)?; + if !surface.writable { + assemble::add_reason(&mut run, TextReason::NoWritableGlyphs); + } + all.push(run); + } + drop(surfaces); + mem.scratch(all.len().saturating_mul(RUN_SCRATCH_BYTES))?; + assemble::mark_duplicates(&mut all); + assemble::mark_per_glyph(&mut all); + let fp = ctx.snap.fingerprint.to_string(); + mem.hold(n.saturating_mul(size_of::>()))?; + let mut record_reason: Vec> = + intrinsic.iter().map(|r| r.first().copied()).collect(); + drop(intrinsic); + let all_slots = all.capacity().saturating_mul(size_of::()); + let mut listed = Vec::new(); + for mut run in all { + run.id = run_id(&fp, page_index, walk, &run.members, mem)?; + for m in &run.members { + if let Some(slot) = record_reason.get_mut(*m) { + *slot = run.reason; + } + } + if !run.text.trim().is_empty() { + mem.push(&mut listed, run)?; + } + } + mem.unhold(all_slots); + assemble::next_obstacles(&mut listed); + let boxes: Vec = listed.iter().map(|r| run_box(walk, r)).collect(); + for (run, (order, line)) in listed + .iter_mut() + .zip(reading_order(ctx, walk.page_id, &boxes)) + { + run.order = order; + run.line = line; + } + listed.sort_by_key(|r| r.order); + Ok(RunsOut { + held: mem.held(), + runs: listed, + page_reason: None, + page_detail: None, + record_reason, + release_walk: false, + }) +} + +/// `t1:{fp}:{page}:{s0}-{e0}[,…]`, charged (an upper bound, then what it holds) before it is +/// built: two 20-digit numbers and two separators per member. +fn run_id( + fp: &str, + page_index: u32, + walk: &PageWalk, + members: &[usize], + mem: &mut ModelBudget, +) -> Result { + let bound = members + .len() + .saturating_mul(42) + .saturating_add(fp.len() + 16); + mem.hold(bound)?; + let spans: Vec = members + .iter() + .map( + |m| match walk.records.get(*m).and_then(|r| r.span.as_ref()) { + Some(s) => format!("{}-{}", s.start, s.end), + None => "?".to_string(), + }, + ) + .collect(); + let id = format!("t1:{fp}:{page_index}:{}", spans.join(",")); + mem.unhold(bound.saturating_sub(id.capacity())); + Ok(id) +} + +/// The run's ink box and baseline in display space (after `/Rotate`, y down). +fn run_box(walk: &PageWalk, run: &TextRun) -> RunBox { + let r = display_rotation(walk.geometry.rotate); + let [x, y, w, h] = run.rect; + let d = transform_rect(&r, [x, y, x + w, y + h]); + let o = apply(&r, run.origin.0, run.origin.1); + RunBox { + display_rect: [d[0], -d[3], d[2], -d[1]], + baseline_y: -o.1, + size: run.effective_size, + mcid: run + .members + .first() + .and_then(|i| walk.records.get(*i)) + .and_then(innermost_mcid), + } +} + +/// Consecutive depth-0 records with glyphs that join (§A.1.1); a glyph-less show op or a depth > 0 +/// record ends the line. Sibling groups come from the page's shared surfaces (`Surfaces`, built +/// once per resource: a line may alternate between many fonts). +fn join_records( + walk: &PageWalk, + intrinsic: &[Vec], + surfaces: &mut Surfaces<'_>, + mem: &mut ModelBudget, +) -> Result>, Stop> { + // At most one group per record, each member in exactly one group (a group's room may double). + let n = walk.records.len(); + let groups_bytes = n.saturating_mul(size_of::>() + 2 * size_of::()); + mem.scratch(groups_bytes.saturating_add(walk.paints.len() * size_of::()))?; + let paint_seqs: Vec = walk.paints.iter().map(|p| p.seq).collect(); + let mut groups = Vec::new(); + let mut cur: Vec = Vec::new(); + for (i, rec) in walk.records.iter().enumerate() { + if rec.depth != 0 || rec.glyphs.is_empty() { + if !cur.is_empty() { + groups.push(std::mem::take(&mut cur)); + } + continue; + } + let joins = match cur.last() { + Some(&last) if cur.len() < RUN_MEMBERS_MAX => { + joinable(walk, intrinsic, &paint_seqs, surfaces, mem, last, i)? + } + _ => false, + }; + if !joins && !cur.is_empty() { + groups.push(std::mem::take(&mut cur)); + } + cur.push(i); + } + if !cur.is_empty() { + groups.push(cur); + } + Ok(groups) +} + +#[allow(clippy::too_many_arguments)] +fn joinable( + walk: &PageWalk, + intrinsic: &[Vec], + paint_seqs: &[u32], + surfaces: &mut Surfaces<'_>, + mem: &mut ModelBudget, + a: usize, + b: usize, +) -> Result { + let (Some(ra), Some(rb)) = (walk.records.get(a), walk.records.get(b)) else { + return Ok(false); + }; + let first = |i: usize| intrinsic.get(i).and_then(|r| r.first().copied()); + let actual = |r: &ShowRecord, i: usize| { + r.before.marked.iter().any(|m| m.actual_text) + || intrinsic + .get(i) + .is_some_and(|v| v.contains(&TextReason::ActualText)) + }; + if b != a + 1 || first(a) != first(b) || actual(ra, a) || actual(rb, b) { + return Ok(false); + } + let next_paint = paint_seqs.partition_point(|s| *s <= ra.seq); + if paint_seqs.get(next_paint).is_some_and(|s| *s < rb.seq) { + return Ok(false); + } + if same_state_except(&ra.before, &rb.before, &[StateField::Font]).is_err() + || !fonts_join(surfaces, mem, ra, rb)? + { + return Ok(false); + } + let (ta, tb) = (ra.text_to_user, rb.text_to_user); + let same_linear = ta + .iter() + .zip(&tb) + .take(4) + .all(|(x, y)| (x - y).abs() <= STATE_EPSILON * 1f64.max(x.abs()).max(y.abs())); + let Some(dir) = unit((ta[0], ta[1])) else { + return Ok(false); + }; + let effective = ta[2].hypot(ta[3]); + // The gap runs to `b`'s first glyph: kerns that open a TJ count (a TJ that starts with a + // column-wide kern is the next column, not the same line). + let b_origin = rb.glyphs.first().map_or(rb.pen_before, |g| g.origin); + Ok(same_linear + && cross(sub(rb.pen_before, ra.pen_before), dir).abs() <= JOIN_BASELINE_TOL_PT + && dot(sub(b_origin, ra.pen_after), dir).abs() <= JOIN_GAP_EM * effective) +} + +/// Identical fonts or siblings (§A.1.1); an ExtGState font joins only the same ExtGState font. +fn fonts_join( + surfaces: &mut Surfaces<'_>, + mem: &mut ModelBudget, + ra: &ShowRecord, + rb: &ShowRecord, +) -> Result { + let (Some(x), Some(y)) = (&ra.before.text.font, &rb.before.text.font) else { + return Ok(false); + }; + if x.from_extgstate || y.from_extgstate { + return Ok(x.from_extgstate + && y.from_extgstate + && x.content_hash == y.content_hash + && ra.font_key == rb.font_key); + } + if x.resource == y.resource && x.content_hash == y.content_hash { + return Ok(true); + } + let (Some(_), Some(yr)) = (&x.resource, &y.resource) else { + return Ok(false); + }; + surfaces.joins(ra, yr, y.content_hash, mem) +} + +/// `PageModel::approx_bytes`: the test oracle the page-model budget is checked against +/// (`approx_bytes ≤ walk.model_bytes`); production sizes models by `walk.model_bytes`. +#[cfg(test)] +mod size; diff --git a/src-tauri/src/pdf_engine/text_edit/runs/assemble.rs b/src-tauri/src/pdf_engine/text_edit/runs/assemble.rs new file mode 100644 index 0000000..99d7437 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/runs/assemble.rs @@ -0,0 +1,672 @@ +//! One run from its joined records (SPEC §B.11 pipeline steps 3–5): units (glyphs, TJ kerns and +//! inter-member gaps converted to kerns), synthetic spaces (§A.3.4), text, caret offsets, ink box +//! and extents; then the page-wide run checks (DUPLICATE_TEXT, PER_GLYPH_TEXT) and the distance +//! to the next text on the same line. + +use super::surface::{unit_heap, SurfaceInfo}; +use super::{KernSrc, SpaceMode, TextRun, Unit}; +use crate::pdf_engine::text_edit::geometry::{ + bbox_of, cross, dot, intersect, mul, ray_extent, sub, unit, +}; +use crate::pdf_engine::text_edit::limits::{ + DEFAULT_KERN_SPACE, DUPLICATE_OVERLAP_SHARE, NEGLIGIBLE_KERN, PER_GLYPH_MIN_RUNS, + PER_GLYPH_SHARE, SYNTH_SPACE_EM, +}; +use crate::pdf_engine::text_edit::reasons::{Face, TextReason}; +use crate::pdf_engine::text_edit::state::ClipState; +use crate::pdf_engine::text_edit::walker::budget::ModelBudget; +use crate::pdf_engine::text_edit::walker::{PageWalk, RecElem, ShowRecord, Stop}; +use std::collections::{BTreeMap, HashMap}; +use std::mem::size_of; +use std::sync::Arc; + +const FALLBACK_ASCENT: f64 = 0.8; +const FALLBACK_DESCENT: f64 = -0.2; +/// Band overlap that makes another run "on the same line" (§B.11 `next_obstacle`). +const SAME_LINE_SHARE: f64 = 0.1; +/// Run pairs `next_obstacles` compares on one page. Past it, every run not yet measured gets an +/// obstacle at its own origin: a `NEXT_TEXT_OVERLAP` warning on any growth, never a missed one. +const OBSTACLE_PAIRS_MAX: usize = 8_000_000; + +/// `|row 1|` and `|row 2|` of `Tm × CTM` at the record's start. +fn tm_ctm_norms(rec: &ShowRecord) -> (f64, f64) { + let m = mul(&rec.tm_before, &rec.before.ctm); + (m[0].hypot(m[1]), m[2].hypot(m[3])) +} + +fn is_space(text: &str) -> bool { + matches!(text, " " | "\u{a0}") +} + +/// Builds the run of `members` (record indices, primary first) with its record-level reasons, +/// typed with `surface`; what it keeps is charged to `mem` before it is made. +pub(super) fn assemble( + walk: &PageWalk, + members: &[usize], + reasons: Vec, + surface: &SurfaceInfo, + mem: &mut ModelBudget, +) -> Result { + mem.scratch(members.len() * size_of::<&ShowRecord>())?; + let recs: Vec<&ShowRecord> = members + .iter() + .filter_map(|i| walk.records.get(*i)) + .collect(); + mem.unscratch(members.len() * size_of::<&ShowRecord>()); + let Some(primary) = recs.first().copied() else { + return Ok(empty_run(reasons)); + }; + let ttu = primary.text_to_user; + let dir = unit((ttu[0], ttu[1])).unwrap_or((1.0, 0.0)); + let up = if cross(dir, (ttu[2], ttu[3])) >= 0.0 { + (-dir.1, dir.0) + } else { + (dir.1, -dir.0) + }; + let t = &primary.before.text; + let (ttux, row2) = tm_ctm_norms(primary); + let effective_size = t.tfs.abs() * row2; + let origin = primary.pen_before; + let units = build_units(&recs, members, dir, ttux, mem)?; + let (text, caret_offsets) = text_and_carets(walk, primary, &units, origin, dir, ttux, mem)?; + let corners: Vec<(f64, f64)> = recs + .iter() + .flat_map(|r| r.glyphs.iter()) + .flat_map(|g| [(g.bbox[0], g.bbox[1]), (g.bbox[2], g.bbox[3])]) + .collect(); + let ink = bbox_of(&corners).unwrap_or([origin.0, origin.1, origin.0, origin.1]); + let (mut ascent, mut descent) = (0.0f64, 0.0f64); + for r in &recs { + let (a, d) = r + .font + .as_ref() + .map_or((FALLBACK_ASCENT, FALLBACK_DESCENT), |f| { + (f.ascent, f.descent) + }); + ascent = ascent.max(a * effective_size); + descent = descent.max(-d * effective_size); + } + let original_extent = recs + .iter() + .flat_map(|r| r.glyphs.iter().map(move |g| (r, g))) + .map(|(r, g)| { + let w = g.width1000 / 1000.0 * (r.before.text.tfs * r.before.text.th * ttux).abs(); + dot(sub(g.origin, origin), dir) + w + }) + .fold(0.0, f64::max); + let visible = match &primary.before.clip { + ClipState::Rect(r) => intersect(walk.geometry.visible, *r), + _ => walk.geometry.visible, + }; + let synth: Vec = units + .iter() + .filter_map(|u| match u { + Unit::Kern { + value, + synth_space: true, + .. + } => Some(*value), + _ => None, + }) + .collect(); + let real_space = units + .iter() + .any(|u| matches!(u, Unit::Glyph { text, .. } if text == " ")); + let space_mode = if surface.has_space && (synth.is_empty() || real_space) { + SpaceMode::Glyph + } else { + SpaceMode::Kern + }; + let face = match primary + .font + .as_ref() + .map_or((false, false), |m| (m.bold, m.italic)) + { + (true, true) => Face::BoldItalic, + (true, false) => Face::Bold, + (false, true) => Face::Italic, + (false, false) => Face::Regular, + }; + let font_from_extgstate = t.font.as_ref().is_some_and(|f| f.from_extgstate); + let fill_hex = primary.before.fill.hex(); + let small = reasons.capacity() + fill_hex.as_ref().map_or(0, String::capacity); + mem.hold(small.saturating_add(members.len() * size_of::()))?; + Ok(TextRun { + id: String::new(), + order: 0, + line: 0, + members: members.to_vec(), + units, + text, + caret_offsets, + origin, + dir, + up, + ascent, + descent, + rect: [ink[0], ink[1], ink[2] - ink[0], ink[3] - ink[1]], + surface: Arc::clone(&surface.names), + tfs: t.tfs, + effective_size, + tc: t.tc, + tw: t.tw, + th: t.th, + text_to_user_x: ttux, + space_mode, + kern_space: median(&synth).unwrap_or(DEFAULT_KERN_SPACE), + original_extent, + visible_extent: ray_extent(origin, dir, visible), + next_obstacle: None, + fill_hex, + face, + font_from_extgstate, + tr: t.tr, + substituted: recs + .iter() + .any(|r| r.font.as_ref().is_some_and(|f| f.substituted)), + reason: reasons.first().copied(), + reasons, + }) +} + +fn empty_run(reasons: Vec) -> TextRun { + TextRun { + id: String::new(), + order: 0, + line: 0, + members: Vec::new(), + units: Vec::new(), + text: String::new(), + caret_offsets: vec![0.0], + origin: (0.0, 0.0), + dir: (1.0, 0.0), + up: (0.0, 1.0), + ascent: 0.0, + descent: 0.0, + rect: [0.0; 4], + surface: Arc::from(Vec::new()), + tfs: 0.0, + effective_size: 0.0, + tc: 0.0, + tw: 0.0, + th: 1.0, + text_to_user_x: 0.0, + space_mode: SpaceMode::Kern, + kern_space: DEFAULT_KERN_SPACE, + original_extent: 0.0, + visible_extent: 0.0, + next_obstacle: None, + fill_hex: None, + face: Face::Regular, + font_from_extgstate: false, + tr: 0, + substituted: false, + reason: reasons.first().copied(), + reasons, + } +} + +/// Glyph units, TJ kerns and inter-member gaps (`n = −gap_ts·1000/Tfs` with +/// `gap_ts = gap_user / (Th · text_to_user_x)`; negligible ones omitted), then the +/// synthetic-space marks. Every unit's slot, text and font resource name are charged first (a +/// glyph unit copies its font's resource name: a long name times many glyphs is charged too). +fn build_units( + recs: &[&ShowRecord], + members: &[usize], + dir: (f64, f64), + ttux: f64, + mem: &mut ModelBudget, +) -> Result, Stop> { + let mut slots = recs.len().saturating_sub(1); + let mut heap = 0usize; + for rec in recs { + slots = slots.saturating_add(rec.elems.len()); + let name = rec + .before + .text + .font + .as_ref() + .and_then(|f| f.resource.as_ref()) + .map_or(0, |r| r.len()); + for g in &rec.glyphs { + let text = g.text.as_ref().map_or(3, String::len); + heap = heap.saturating_add(text.saturating_add(name)); + } + } + mem.hold(slots.saturating_mul(size_of::()).saturating_add(heap))?; + let mut units = Vec::with_capacity(slots); + let mut prev: Option<&ShowRecord> = None; + for (k, (rec, idx)) in recs.iter().zip(members).enumerate() { + let tfs = rec.before.text.tfs; + let font = rec.before.text.font.as_ref(); + if let Some(p) = prev { + let gap_user = dot(sub(rec.pen_before, p.pen_after), dir); + let scale = rec.before.text.th * ttux; + let n = if scale != 0.0 && tfs != 0.0 { + -(gap_user / scale) * 1000.0 / tfs + } else { + 0.0 + }; + if n.is_finite() && n.abs() >= NEGLIGIBLE_KERN { + units.push(Unit::Kern { + value: n, + src: KernSrc::Gap, + synth_space: false, + }); + } + } + prev = Some(rec); + for elem in &rec.elems { + match elem { + RecElem::Glyph(gi) => { + let Some(g) = rec.glyphs.get(*gi) else { + continue; + }; + units.push(Unit::Glyph { + member: k, + rec: *idx, + glyph: *gi, + code: g.code, + font_res: font.and_then(|f| f.resource.as_deref()).map(<[u8]>::to_vec), + font_hash: font.map_or(0, |f| f.content_hash), + text: g.text.clone().unwrap_or_else(|| "\u{fffd}".to_string()), + width1000: g.width1000, + }); + } + RecElem::Kern { value, span } => units.push(Unit::Kern { + value: *value, + src: KernSrc::Tj { span: span.clone() }, + synth_space: false, + }), + } + } + } + mark_synthetic_spaces(&mut units); + let kept: usize = units.iter().map(unit_heap).sum(); + mem.unhold(heap.saturating_sub(kept)); + Ok(units) +} + +/// A TJ number or converted gap with `−n/1000 ≥ 0.2 em` between two non-space glyphs reads as +/// exactly one space (§A.3.4): in a sequence of kerns between two glyphs the sum decides, and the +/// largest one carries the space. +fn mark_synthetic_spaces(units: &mut [Unit]) { + let glyph_at: Vec = units + .iter() + .enumerate() + .filter(|(_, u)| matches!(u, Unit::Glyph { .. })) + .map(|(i, _)| i) + .collect(); + for pair in glyph_at.windows(2) { + let [a, b] = pair else { continue }; + let (a, b) = (*a, *b); + let space_at = |i: usize| match units.get(i) { + Some(Unit::Glyph { text, .. }) => is_space(text), + _ => true, + }; + if b <= a + 1 || space_at(a) || space_at(b) { + continue; + } + let kerns: Vec<(usize, f64)> = (a + 1..b) + .filter_map(|i| match units.get(i) { + Some(Unit::Kern { value, .. }) => Some((i, *value)), + _ => None, + }) + .collect(); + let total: f64 = kerns.iter().map(|(_, v)| -v / 1000.0).sum(); + if total < SYNTH_SPACE_EM { + continue; + } + let biggest = kerns + .iter() + .fold(None::<(usize, f64)>, |best, &(i, v)| match best { + Some((_, bv)) if bv <= v => best, + _ => Some((i, v)), + }); + if let Some((i, _)) = biggest { + if let Some(Unit::Kern { synth_space, .. }) = units.get_mut(i) { + *synth_space = true; + } + } + } +} + +/// The run text (synthetic spaces as " ") and its chars+1 caret offsets along `dir` from the +/// origin, made monotonic; both are sized (and charged) exactly before they are filled. +fn text_and_carets( + walk: &PageWalk, + primary: &ShowRecord, + units: &[Unit], + origin: (f64, f64), + dir: (f64, f64), + ttux: f64, + mem: &mut ModelBudget, +) -> Result<(String, Vec), Stop> { + let (tfs, th) = (primary.before.text.tfs, primary.before.text.th); + let (mut bytes, mut chars) = (0usize, 1usize); + for u in units { + let (b, c) = match u { + Unit::Glyph { text: t, .. } => (t.len(), t.chars().count()), + Unit::Kern { + synth_space: true, .. + } => (1, 1), + Unit::Kern { .. } => (0, 0), + }; + bytes = bytes.saturating_add(b); + chars = chars.saturating_add(c); + } + mem.hold(bytes.saturating_add(chars.saturating_mul(size_of::())))?; + let mut text = String::with_capacity(bytes); + let mut offsets = Vec::with_capacity(chars); + let mut end = 0.0f64; + let mut pen = 0.0f64; + for u in units { + match u { + Unit::Glyph { + rec, + glyph, + text: t, + .. + } => { + let Some(g) = walk.records.get(*rec).and_then(|r| r.glyphs.get(*glyph)) else { + continue; + }; + let start = dot(sub(g.origin, origin), dir); + let adv = dot(g.advance_user, dir); + let n = t.chars().count(); + for j in 0..n { + offsets.push(start + adv * j as f64 / n as f64); + } + text.push_str(t); + pen = start + adv; + end = pen; + } + Unit::Kern { + value, + synth_space, + src, + } => { + let adv = -value / 1000.0 * tfs * th * ttux; + if *synth_space { + offsets.push(pen); + text.push(' '); + pen += adv; + end = pen; + } else if matches!(src, KernSrc::Tj { .. }) { + pen += adv; + } + } + } + } + offsets.push(end); + let mut max = f64::NEG_INFINITY; + for o in &mut offsets { + max = max.max(*o); + *o = max; + } + Ok((text, offsets)) +} + +fn median(values: &[f64]) -> Option { + let mut v: Vec = values.iter().copied().filter(|x| x.is_finite()).collect(); + if v.is_empty() { + return None; + } + v.sort_by(|a, b| a.total_cmp(b)); + v.get(v.len() / 2).copied() +} + +fn area(r: [f64; 4]) -> f64 { + r[2].max(0.0) * r[3].max(0.0) +} + +fn overlap_area(a: [f64; 4], b: [f64; 4]) -> f64 { + let w = (a[0] + a[2]).min(b[0] + b[2]) - a[0].max(b[0]); + let h = (a[1] + a[3]).min(b[1] + b[3]) - a[1].max(b[1]); + w.max(0.0) * h.max(0.0) +} + +/// DUPLICATE_TEXT (§A.6): two runs with the same trimmed text whose ink boxes overlap by at +/// least half of the smaller one are both refused. +/// +/// Per text, runs are swept by `x`; memory is one flag per run. Once a run is marked it is only +/// compared with runs not marked yet, found through `Unmarked` (a stack of identical shadows costs +/// O(n), not O(n²)). +pub(super) fn mark_duplicates(runs: &mut [TextRun]) { + let mut groups: HashMap<&str, Vec> = HashMap::new(); + for (i, r) in runs.iter().enumerate() { + let key = r.text.trim(); + if !key.is_empty() { + groups.entry(key).or_default().push(i); + } + } + let mut marked = vec![false; runs.len()]; + for mut idx in groups.into_values() { + let x_of = |i: &usize| runs.get(*i).map_or(0.0, |r| r.rect[0]); + idx.sort_by(|a, b| x_of(a).total_cmp(&x_of(b))); + let mut unmarked = Unmarked::new(idx.len()); + for (k, &i) in idx.iter().enumerate() { + let Some(a) = runs.get(i) else { continue }; + let mut a_marked = marked.get(i).copied().unwrap_or(true); + let mut p = if a_marked { + unmarked.find(k + 1) + } else { + k + 1 + }; + while let Some((&j, b)) = idx.get(p).and_then(|j| Some((j, runs.get(*j)?))) { + if b.rect[0] > a.rect[0] + a.rect[2] { + break; // sorted by x: no later run overlaps `a` + } + let smaller = area(a.rect).min(area(b.rect)); + let shared = overlap_area(a.rect, b.rect); + if smaller > 0.0 && shared >= DUPLICATE_OVERLAP_SHARE * smaller { + for (slot, pos) in [(i, k), (j, p)] { + if let Some(m) = marked.get_mut(slot) { + *m = true; + } + unmarked.mark(pos); + } + a_marked = true; + } + p = if a_marked { + unmarked.find(p + 1) + } else { + p + 1 + }; + } + } + } + for (r, hit) in runs.iter_mut().zip(marked) { + if hit { + add_reason(r, TextReason::DuplicateText); + } + } +} + +/// The next unmarked position at or after `p` in one sorted group (`len` when none): a +/// path-compressed successor list, `next[p] == p` while `p` is unmarked. +struct Unmarked { + next: Vec, +} + +impl Unmarked { + fn new(len: usize) -> Unmarked { + Unmarked { + next: (0..=len).collect(), + } + } + + fn mark(&mut self, p: usize) { + if let Some(slot) = self.next.get_mut(p) { + *slot = p.saturating_add(1); + } + } + + fn find(&mut self, p: usize) -> usize { + let end = self.next.len().saturating_sub(1); + let mut root = p.min(end); + while let Some(&n) = self.next.get(root) { + if n == root || n > end { + break; + } + root = n; + } + let mut cur = p.min(end); + while cur != root { + let Some(slot) = self.next.get_mut(cur) else { + break; + }; + let n = *slot; + *slot = root; + cur = n; + } + root + } +} + +/// PER_GLYPH_TEXT (§A.1.2): with ≥ 12 editable non-blank runs of which ≥ 80 % are one code +/// point long, those one-code-point runs are refused. +pub(super) fn mark_per_glyph(runs: &mut [TextRun]) { + let editable: Vec = runs + .iter() + .enumerate() + .filter(|(_, r)| r.reasons.is_empty() && !r.text.trim().is_empty()) + .map(|(i, _)| i) + .collect(); + let single: Vec = editable + .iter() + .copied() + .filter(|i| runs.get(*i).is_some_and(|r| r.text.chars().count() == 1)) + .collect(); + let share = single.len() as f64 / editable.len().max(1) as f64; + if editable.len() >= PER_GLYPH_MIN_RUNS && share >= PER_GLYPH_SHARE { + for i in single { + if let Some(r) = runs.get_mut(i) { + add_reason(r, TextReason::PerGlyphText); + } + } + } +} + +pub(super) fn add_reason(run: &mut TextRun, reason: TextReason) { + if !run.reasons.contains(&reason) { + run.reasons.push(reason); + run.reasons.sort(); + } + run.reason = run.reasons.first().copied(); +} + +/// `next_obstacle` (§B.11): the distance along `dir` to the nearest other run with the same +/// direction that starts ahead of this run's origin and whose band (descent..ascent across the +/// baseline) overlaps this run's band by more than 10 %. +/// +/// Runs are bucketed by direction, then by band height (powers of two), each class sorted by its +/// position across the baseline: a run only looks at the runs of each class whose band can reach +/// its own, so one tall title does not widen every other run's search. At most +/// `OBSTACLE_PAIRS_MAX` pairs are compared per page (in a fixed order); a run whose search the +/// cap cuts short gets an obstacle at its own origin (conservative: a warning on any growth). +pub(super) fn next_obstacles(runs: &mut [TextRun]) { + let dir_key = |r: &TextRun| { + ( + (r.dir.0 * 1e4).round() as i64, + (r.dir.1 * 1e4).round() as i64, + ) + }; + let mut buckets: BTreeMap<(i64, i64), Vec> = BTreeMap::new(); + for (i, r) in runs.iter().enumerate() { + buckets.entry(dir_key(r)).or_default().push(i); + } + let mut left = OBSTACLE_PAIRS_MAX; + let mut found: Vec<(usize, f64)> = Vec::new(); + for idx in buckets.values() { + let Some(up) = idx.first().and_then(|i| runs.get(*i)).map(|r| r.up) else { + continue; + }; + let perp = |i: usize| runs.get(i).map_or(0.0, |r| dot(r.origin, up)); + let classes = band_classes(runs, idx, &perp); + let mut order = idx.clone(); + order.sort_by(|a, b| perp(*a).total_cmp(&perp(*b)).then(a.cmp(b))); + for i in order { + let Some(a) = runs.get(i) else { continue }; + let pa = perp(i); + let (a_lo, a_hi) = (pa - a.descent, pa + a.ascent); + let band = (a_hi - a_lo).max(1e-9); + let mut best: Option = None; + let mut cut = false; + for class in classes.values() { + // `b` can overlap only if its baseline lies within this window. + let lo = a_lo - class.ascent; + let hi = a_hi + class.descent; + let start = class.sorted.partition_point(|j| perp(*j) < lo); + for &j in class.sorted.iter().skip(start) { + let pb = perp(j); + if pb > hi { + break; + } + let Some(next) = left.checked_sub(1) else { + cut = true; + break; + }; + left = next; + let Some(b) = runs.get(j).filter(|_| j != i) else { + continue; + }; + let shared = a_hi.min(pb + b.ascent) - a_lo.max(pb - b.descent); + let ahead = dot(sub(b.origin, a.origin), a.dir); + if shared > SAME_LINE_SHARE * band && ahead > 1e-6 { + best = Some(best.map_or(ahead, |x: f64| x.min(ahead))); + } + } + } + match (cut, best) { + (true, _) => found.push((i, 0.0)), + (false, Some(d)) => found.push((i, d)), + (false, None) => {} + } + } + } + for (i, d) in found { + if let Some(r) = runs.get_mut(i) { + r.next_obstacle = Some(d); + } + } +} + +/// The runs of one direction bucket grouped by band height (`floor(log2(height))`): each class +/// sorted by `perp`, with the largest ascent and descent of its members. +struct BandClass { + sorted: Vec, + ascent: f64, + descent: f64, +} + +fn band_classes( + runs: &[TextRun], + idx: &[usize], + perp: &impl Fn(usize) -> f64, +) -> BTreeMap { + let mut classes: BTreeMap = BTreeMap::new(); + for &i in idx { + let Some(r) = runs.get(i) else { continue }; + let height = (r.ascent + r.descent).max(1e-9); + let key = if height.is_finite() { + height.log2().floor().clamp(-64.0, 64.0) as i32 + } else { + 64 + }; + let class = classes.entry(key).or_insert(BandClass { + sorted: Vec::new(), + ascent: 0.0, + descent: 0.0, + }); + class.sorted.push(i); + class.ascent = class.ascent.max(r.ascent); + class.descent = class.descent.max(r.descent); + } + for class in classes.values_mut() { + class + .sorted + .sort_by(|a, b| perp(*a).total_cmp(&perp(*b)).then(a.cmp(b))); + } + classes +} diff --git a/src-tauri/src/pdf_engine/text_edit/runs/forms.rs b/src-tauri/src/pdf_engine/text_edit/runs/forms.rs new file mode 100644 index 0000000..480d2b3 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/runs/forms.rs @@ -0,0 +1,150 @@ +//! Lines drawn through Form XObjects (SPEC §A.6: "Text in a Form XObject (any depth) → Refused, +//! still shown, `NESTED_FORM`"; live check B1). A page model walks depth 0 only, so Form text +//! has no run; the Classify walk descends Forms, and its depth > 0 records become refused lines: +//! consecutive records of one Form on one baseline (the §A.1.1 geometry: same direction and size, +//! baseline within `JOIN_BASELINE_TOL_PT`, gap within `JOIN_GAP_EM`, nothing painted between) +//! join, each line is assembled like a run (`assemble`, typed with no surface) and carries +//! `NESTED_FORM` only. They follow the page's runs in reading order. Everything they keep is +//! charged to the Classify pass's budget. + +use super::assemble::assemble; +use super::surface::SurfaceInfo; +use super::{PageModel, TextRun}; +use crate::pdf_engine::text_edit::geometry::{cross, dot, sub, unit}; +use crate::pdf_engine::text_edit::limits::{ + JOIN_BASELINE_TOL_PT, JOIN_GAP_EM, RUNS_PER_PAGE_MAX, RUN_MEMBERS_MAX, STATE_EPSILON, +}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::walker::budget::ModelBudget; +use crate::pdf_engine::text_edit::walker::{PageWalk, ShowRecord, Stop}; +use std::mem::size_of; + +/// Bytes of a Form line id besides its spans: `t1:`, the fingerprint, the page and separators. +const ID_FIXED_BYTES: usize = 64; +/// Bytes per Form in an id (`x{obj}.{gen}>`) and per member span (`{start}-{end},`). +const ID_FORM_BYTES: usize = 24; +const ID_SPAN_BYTES: usize = 42; + +/// The refused lines `walk` (a Classify walk of `model`'s page) draws through Forms, numbered +/// after the model's runs; none when they would pass `RUNS_PER_PAGE_MAX` together. +pub(crate) fn form_lines( + model: &PageModel, + walk: &PageWalk, + mem: &mut ModelBudget, +) -> Result, Stop> { + let groups = form_groups(walk, mem)?; + if groups.len().saturating_add(model.runs.len()) > RUNS_PER_PAGE_MAX { + return Ok(Vec::new()); + } + let order = model.runs.iter().map(|r| r.order + 1).max().unwrap_or(0); + let line = model.runs.iter().map(|r| r.line + 1).max().unwrap_or(0); + let none = SurfaceInfo::none(); + let fp = model.fingerprint.to_string(); + mem.hold(groups.len().saturating_mul(size_of::()))?; + let mut lines = Vec::with_capacity(groups.len()); + for g in &groups { + let mut run = assemble(walk, g, vec![TextReason::NestedForm], &none, mem)?; + if run.text.trim().is_empty() { + continue; + } + let k = u32::try_from(lines.len()).unwrap_or(u32::MAX); + run.order = order.saturating_add(k); + run.line = line.saturating_add(k); + run.id = line_id(&fp, model.page_index, walk, g, mem)?; + lines.push(run); + } + Ok(lines) +} + +/// Consecutive depth > 0 records with glyphs that read as one line. +fn form_groups(walk: &PageWalk, mem: &mut ModelBudget) -> Result>, Stop> { + let n = walk.records.iter().filter(|r| r.depth > 0).count(); + mem.scratch(n.saturating_mul(size_of::>() + 2 * size_of::()))?; + let mut groups: Vec> = Vec::new(); + let mut cur: Vec = Vec::new(); + for (i, rec) in walk.records.iter().enumerate() { + if rec.depth == 0 || rec.glyphs.is_empty() { + if !cur.is_empty() { + groups.push(std::mem::take(&mut cur)); + } + continue; + } + let joins = cur.last().is_some_and(|last| { + cur.len() < RUN_MEMBERS_MAX + && *last + 1 == i + && walk + .records + .get(*last) + .is_some_and(|a| one_line(walk, a, rec)) + }); + if !joins && !cur.is_empty() { + groups.push(std::mem::take(&mut cur)); + } + cur.push(i); + } + if !cur.is_empty() { + groups.push(cur); + } + Ok(groups) +} + +/// `b` continues `a`'s line: the same Form, nothing painted between, the same text direction and +/// size, `b`'s pen on `a`'s baseline and its first glyph within `JOIN_GAP_EM` of `a`'s end. +fn one_line(walk: &PageWalk, a: &ShowRecord, b: &ShowRecord) -> bool { + let next_paint = walk.paints.partition_point(|p| p.seq <= a.seq); + let painted_between = walk.paints.get(next_paint).is_some_and(|p| p.seq < b.seq); + if a.form_chain != b.form_chain || painted_between { + return false; + } + let (ta, tb) = (a.text_to_user, b.text_to_user); + let same_linear = ta + .iter() + .zip(&tb) + .take(4) + .all(|(x, y)| (x - y).abs() <= STATE_EPSILON * 1f64.max(x.abs()).max(y.abs())); + let Some(dir) = unit((ta[0], ta[1])) else { + return false; + }; + let effective = ta[2].hypot(ta[3]); + let b_origin = b.glyphs.first().map_or(b.pen_before, |g| g.origin); + same_linear + && cross(sub(b.pen_before, a.pen_before), dir).abs() <= JOIN_BASELINE_TOL_PT + && dot(sub(b_origin, a.pen_after), dir).abs() <= JOIN_GAP_EM * effective +} + +/// `t1:{fp}:{page}:x{obj}.{gen}>…:{s0}-{e0}[,{s1}-{e1}…]`: the Form chain and the members' spans +/// in the innermost Form's data (unique on the page; a run id never starts with `x` there). +fn line_id( + fp: &str, + page_index: u32, + walk: &PageWalk, + members: &[usize], + mem: &mut ModelBudget, +) -> Result { + let first = members.first().and_then(|i| walk.records.get(*i)); + let chain = first.map_or(0, |r| r.form_chain.len()); + let bound = (ID_FIXED_BYTES + fp.len()) + .saturating_add(chain.saturating_mul(ID_FORM_BYTES)) + .saturating_add(members.len().saturating_mul(ID_SPAN_BYTES)); + mem.hold(bound)?; + let forms: Vec = first + .map(|r| { + r.form_chain + .iter() + .map(|(o, g)| format!("x{o}.{g}")) + .collect() + }) + .unwrap_or_default(); + let spans: Vec = members + .iter() + .filter_map(|i| walk.records.get(*i)) + .map(|r| format!("{}-{}", r.local_span.start, r.local_span.end)) + .collect(); + let id = format!( + "t1:{fp}:{page_index}:{}:{}", + forms.join(">"), + spans.join(",") + ); + mem.unhold(bound.saturating_sub(id.capacity())); + Ok(id) +} diff --git a/src-tauri/src/pdf_engine/text_edit/runs/reasons.rs b/src-tauri/src/pdf_engine/text_edit/runs/reasons.rs new file mode 100644 index 0000000..148f5d8 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/runs/reasons.rs @@ -0,0 +1,198 @@ +//! The reasons a single show record cannot be edited (SPEC §A.10 rows 1–28 that are properties of +//! one show op): nesting, content ownership, inline images, render mode, orientation as +//! displayed, clips, layers, masks, patterns, ActualText, fonts, decodability and scripts. Joins +//! use the first one (A.1.1: refused records only join records with the same reason); runs add +//! DUPLICATE_TEXT, PER_GLYPH_TEXT and NO_WRITABLE_GLYPHS afterwards. + +use crate::pdf_engine::text_edit::content::{part_exclusive, KidsCounts, PageContent}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::fonts::encodings::{reading_reason, usable_text}; +use crate::pdf_engine::text_edit::geometry::{ + bbox_of, classify_orientation, contains, display_rotation, mul, Matrix, Orientation, +}; +use crate::pdf_engine::text_edit::limits::CLIP_CONTAIN_TOL_PT; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::state::{ClipState, OcState}; +use crate::pdf_engine::text_edit::structure::PageStruct; +use crate::pdf_engine::text_edit::walker::show::width_unknown; +use crate::pdf_engine::text_edit::walker::{PageWalk, ShowRecord}; + +/// Per-page inputs shared by every record's checks. +pub(crate) struct ReasonCtx<'a> { + ctx: &'a SnapshotContext, + content: &'a PageContent, + walk: &'a PageWalk, + kids: Option<&'a KidsCounts>, + rotation: Matrix, + page_struct: PageStruct<'a>, + exclusive: Vec>, +} + +impl<'a> ReasonCtx<'a> { + pub(crate) fn new( + ctx: &'a SnapshotContext, + content: &'a PageContent, + walk: &'a PageWalk, + ) -> ReasonCtx<'a> { + ReasonCtx { + ctx, + content, + walk, + kids: ctx.kids().ok(), + rotation: display_rotation(walk.geometry.rotate), + page_struct: PageStruct::of(ctx.doc(), walk.page_id), + exclusive: vec![None; content.parts.len()], + } + } + + fn part_exclusive(&mut self, part: usize) -> bool { + if let Some(Some(known)) = self.exclusive.get(part) { + return *known; + } + let Some(kids) = self.kids else { + return false; + }; + let ok = part_exclusive(self.ctx.doc(), self.ctx.refs(), kids, self.content, part); + if let Some(slot) = self.exclusive.get_mut(part) { + *slot = Some(ok); + } + ok + } + + /// Every reason of §A.10 that holds for `rec` on its own, in priority order. + pub(crate) fn record_reasons(&mut self, rec: &ShowRecord) -> Vec { + use TextReason as R; + let mut out = Vec::new(); + if rec.depth > 0 { + out.push(R::NestedForm); + } + if let Some(span) = &rec.span { + match self.content.locate(span) { + None => out.push(R::SplitContent), + Some((part, _)) => { + if !self.part_exclusive(part) { + out.push(R::SharedContent); + } + } + } + } + if rec.after_unproven_inline_image { + out.push(R::InlineImage); + } + let tr = rec.before.text.tr; + match tr { + 3 => out.push(R::InvisibleText), + 4..=7 => out.push(R::TextClipMode), + _ => {} + } + match classify_orientation(&mul(&rec.text_to_user, &self.rotation)) { + Orientation::Upright => {} + Orientation::ZeroSize => out.push(R::ZeroSize), + Orientation::Rotated => out.push(R::RotatedText), + Orientation::Mirrored => out.push(R::MirroredText), + Orientation::Skewed => out.push(R::SkewedText), + } + if rec.font.as_ref().is_some_and(|f| f.vertical) { + out.push(R::Vertical); + } + if self.clipped(rec) { + out.push(R::Clipped); + } + if rec + .before + .marked + .iter() + .any(|m| matches!(m.oc, Some(OcState::Hidden | OcState::Unknown))) + { + out.push(R::OptionalContent); + } + if rec.before.gs.soft_mask { + out.push(R::SoftMask); + } + let fills = matches!(tr, 0 | 2 | 4 | 6); + let strokes = matches!(tr, 1 | 2 | 5 | 6); + if (fills && rec.before.fill.pattern) || (strokes && rec.before.stroke.pattern) { + out.push(R::Pattern); + } + if self.actual_text(rec) { + out.push(R::ActualText); + } + match &rec.font { + None => out.push(R::MissingFont), + Some(f) => { + if let Some(r) = f.refusal.filter(|r| !r.is_page_level()) { + out.push(r); + } + if rec.glyphs.iter().any(|g| width_unknown(f, g.code)) { + out.push(R::MissingWidths); + } + } + } + // Drawn where an earlier op of this text object left the pen after an advance the model + // cannot know: the rest of the line is not where the model puts it. + if rec.pen_unknown { + out.push(R::MissingWidths); + } + let undecodable = rec.split_error + || rec + .glyphs + .iter() + .any(|g| g.text.as_deref().map_or(true, |t| !usable_text(t))); + if undecodable { + out.push(R::AmbiguousUnicode); + } + for g in &rec.glyphs { + for ch in g.text.as_deref().unwrap_or_default().chars() { + if let Some(r) = reading_reason(ch) { + out.push(r); + } + } + } + out.sort(); + out.dedup(); + out + } + + /// The ink box must lie inside the visible box and the clip (§A.6); a complex clip refuses. + fn clipped(&self, rec: &ShowRecord) -> bool { + let corners: Vec<(f64, f64)> = rec + .glyphs + .iter() + .flat_map(|g| [(g.bbox[0], g.bbox[1]), (g.bbox[2], g.bbox[3])]) + .collect(); + let Some(ink) = bbox_of(&corners) else { + return false; + }; + if !contains(self.walk.geometry.visible, ink, CLIP_CONTAIN_TOL_PT) { + return true; + } + match &rec.before.clip { + ClipState::None => false, + ClipState::Rect(r) => !contains(*r, ink, CLIP_CONTAIN_TOL_PT), + ClipState::Complex => true, + } + } + + /// `/ActualText` or `/E` on the marked-content stack, or on the structure element of the + /// innermost MCID or one of its ancestors (D19). A broken structure tree counts as present. + fn actual_text(&self, rec: &ShowRecord) -> bool { + if rec.before.marked.iter().any(|m| m.actual_text) { + return true; + } + if rec.depth > 0 { + return false; + } + match innermost_mcid(rec) { + Some(mcid) => !matches!( + self.page_struct.actual_text(self.ctx.doc(), mcid), + Ok(false) + ), + None => false, + } + } +} + +/// The MCID of the innermost marked-content sequence that has one (`iter` is innermost first). +pub(crate) fn innermost_mcid(rec: &ShowRecord) -> Option { + rec.before.marked.iter().find_map(|m| m.mcid) +} diff --git a/src-tauri/src/pdf_engine/text_edit/runs/size.rs b/src-tauri/src/pdf_engine/text_edit/runs/size.rs new file mode 100644 index 0000000..b29e691 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/runs/size.rs @@ -0,0 +1,133 @@ +//! Approximate heap size of a page model (SPEC §B.19 caches, §H R20), counted from what the +//! `PageModel` holds: the oracle the page-model budget is tested against (`approx_bytes ≤ +//! walk.model_bytes` on every page the tests build). Production sizes models by +//! `walk.model_bytes`, the budget's own O(1) tally (fix pass 2026-10-03). A model is linear in its +//! page's ops (records, paints, runs), but one hostile page at `PAGE_OPS_MAX` could hold hundreds +//! of MiB, which a count of models alone cannot see. +//! +//! Counted, as the bytes each allocation requests: every vector's capacity, every lexed operand +//! node and its bytes (a model keeps none), the joined content, each distinct state digest once +//! (consecutive records and paints share one) and, once per allocation, every shared part a digest +//! holds (verbatim op bytes, names, the dash array, the unmodelled ExtGState keys and their list, +//! the nodes of the marked-content stack and their tags) — a page whose state changes on every op +//! holds one of these per op — each run surface once (runs share them), and each distinct font +//! model the page fonts and records keep alive (`FontModel::approx_bytes`, once per allocation: +//! the model keeps them however the snapshot's font cache evicts). Not counted: the allocator's +//! own rounding and bookkeeping. +//! +//! The page-model budget charges the same bytes as they are allocated (`walker::budget`, whose +//! helpers this module uses), so `approx_bytes` never exceeds what the build was allowed to hold. + +use super::surface::unit_heap; +use super::{PageModel, TextRun, Unit}; +use crate::pdf_engine::text_edit::fonts::FontModel; +use crate::pdf_engine::text_edit::lexer::{Op, Span}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::walker::budget::{arc_slice, content_bytes, op_heap, Shared}; +use crate::pdf_engine::text_edit::walker::{GlyphRec, PaintKind, PaintRecord, RecElem, ShowRecord}; +use lopdf::ObjectId; +use std::mem::size_of; +use std::sync::Arc; + +impl PageModel { + /// Approximate bytes this model keeps on the heap (see the module comment). + pub fn approx_bytes(&self) -> usize { + let walk = &self.walk; + let mut digests = Shared::default(); + let records: usize = walk + .records + .iter() + .map(|r| record_bytes(r, &mut digests)) + .sum(); + let paints: usize = walk + .paints + .iter() + .map(|p| paint_bytes(p, &mut digests)) + .sum(); + let ops: usize = walk.ops.iter().map(op_heap).sum(); + let runs: usize = self.runs.iter().map(|r| run_bytes(r, &mut digests)).sum(); + let vectors = walk.records.capacity() * size_of::() + + walk.paints.capacity() * size_of::() + + walk.ops.capacity() * size_of::() + + self.runs.capacity() * size_of::() + + walk.page_fonts.capacity() * size_of::<(Vec, usize)>(); + let content = content_bytes(&self.content); + let fonts: usize = walk + .page_fonts + .iter() + .map(|(name, model)| name.capacity() + font_bytes(model, &mut digests)) + .sum::() + + walk + .records + .iter() + .filter_map(|r| r.font.as_ref()) + .map(|model| font_bytes(model, &mut digests)) + .sum::(); + let per_record = self.record_reason.capacity() * size_of::>(); + [ + vectors, + records, + paints, + digests.bytes(), + ops, + runs, + content, + fonts, + per_record, + ] + .iter() + .fold(0usize, |sum, b| sum.saturating_add(*b)) + } +} + +/// A font model's bytes, the first time its allocation is seen. +fn font_bytes(model: &Arc, shared: &mut Shared) -> usize { + if shared.arc(model, 0) { + model.approx_bytes() + } else { + 0 + } +} + +fn record_bytes(r: &ShowRecord, digests: &mut Shared) -> usize { + digests.add(&r.before); + digests.add(&r.after); + let texts: usize = r + .glyphs + .iter() + .map(|g| g.text.as_ref().map_or(0, String::capacity)) + .sum(); + r.form_chain.capacity() * size_of::() + + r.operand_spans.capacity() * size_of::() + + r.glyphs.capacity() * size_of::() + + texts + + r.elems.capacity() * size_of::() +} + +fn paint_bytes(p: &PaintRecord, digests: &mut Shared) -> usize { + digests.add(&p.state); + let name = match &p.kind { + PaintKind::Shading { name, .. } + | PaintKind::ImageXObject { name, .. } + | PaintKind::FormXObject { name, .. } => name.capacity(), + PaintKind::Path(_) | PaintKind::InlineImage { .. } => 0, + }; + p.form_chain.capacity() * size_of::() + name +} + +/// One run's heap; its surface (shared by the runs that have it) is counted in `shared`, once +/// per allocation. +fn run_bytes(r: &TextRun, shared: &mut Shared) -> usize { + let units: usize = r.units.iter().map(unit_heap).sum(); + let names: usize = r.surface.iter().map(Vec::capacity).sum(); + let list = arc_slice(r.surface.len(), size_of::>()); + shared.arc(&r.surface, list.saturating_add(names)); + r.id.capacity() + + r.text.capacity() + + r.members.capacity() * size_of::() + + r.units.capacity() * size_of::() + + units + + r.caret_offsets.capacity() * size_of::() + + r.reasons.capacity() * size_of::() + + r.fill_hex.as_ref().map_or(0, String::capacity) +} diff --git a/src-tauri/src/pdf_engine/text_edit/runs/surface.rs b/src-tauri/src/pdf_engine/text_edit/runs/surface.rs new file mode 100644 index 0000000..1c34809 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/runs/surface.rs @@ -0,0 +1,237 @@ +//! Typing surfaces of a page's runs (SPEC §B.11 `surface`, review T3 r3 HIGH-1). A surface is +//! built once per surface key — the ExtGState font, or the primary's resource name — and shared by +//! every run that has it: a page of 20,000 lines in one font holds one list of sibling names, not +//! 20,000 copies of it, and `typing_surface` (O(F²)) runs once per key, not once per run. +//! +//! A sibling group never holds a font whose resource name is longer than +//! `SIBLING_NAME_BYTES_MAX` (ISO 32000 Annex C): such a font types only its own runs. Every +//! surface, its names and the scratch that builds it are charged to the page-model budget. + +use super::Unit; +use crate::pdf_engine::text_edit::fonts::{typing_surface, FontKey, FontModel, TypingSurface}; +use crate::pdf_engine::text_edit::limits::SIBLING_NAME_BYTES_MAX; +use crate::pdf_engine::text_edit::state::Bytes; +use crate::pdf_engine::text_edit::walker::budget::{arc_slice, map_entry, ModelBudget}; +use crate::pdf_engine::text_edit::walker::{PageWalk, ShowRecord, Stop}; +use std::borrow::Cow; +use std::collections::HashMap; +use std::mem::size_of; +use std::sync::Arc; + +/// What a surface depends on: the ExtGState font's key, or the primary's font resource name. +#[derive(Debug, Clone, PartialEq, Eq, Hash)] +pub(super) enum SurfaceKey { + ExtGState(Option), + Resource(Option), +} + +pub(super) fn surface_key(primary: &ShowRecord) -> SurfaceKey { + match primary.before.text.font.as_ref() { + Some(f) if f.from_extgstate => SurfaceKey::ExtGState(primary.font_key), + f => SurfaceKey::Resource(f.and_then(|f| f.resource.clone())), + } +} + +type PageFonts = [(Vec, Arc)]; + +/// The page fonts a sibling group may hold: every one whose name is at most +/// `SIBLING_NAME_BYTES_MAX` bytes (borrowed when that is all of them). +fn sibling_fonts(walk: &PageWalk) -> Cow<'_, PageFonts> { + let fonts = walk.page_fonts.as_slice(); + if fonts.iter().all(|(n, _)| n.len() <= SIBLING_NAME_BYTES_MAX) { + return Cow::Borrowed(fonts); + } + Cow::Owned( + fonts + .iter() + .filter(|(n, _)| n.len() <= SIBLING_NAME_BYTES_MAX) + .cloned() + .collect(), + ) +} + +/// The surface of `primary` over `siblings`: its ExtGState font alone (§A.1.1), its resource's +/// sibling group, or — for a name too long to join one — that font alone. +fn surface_over(walk: &PageWalk, siblings: &PageFonts, primary: &ShowRecord) -> TypingSurface { + let font = primary.before.text.font.as_ref(); + if font.is_some_and(|f| f.from_extgstate) { + return TypingSurface { + fonts: primary + .font + .iter() + .map(|m| (Vec::new(), Arc::clone(m))) + .collect(), + }; + } + match font.and_then(|f| f.resource.as_deref()) { + Some(res) if res.len() > SIBLING_NAME_BYTES_MAX => TypingSurface { + fonts: walk + .page_fonts + .iter() + .find(|(n, _)| n.as_slice() == res) + .map(|(n, m)| (n.clone(), Arc::clone(m))) + .into_iter() + .collect(), + }, + Some(res) => typing_surface(siblings, res), + None => TypingSurface { fonts: Vec::new() }, + } +} + +/// The typing surface of a run (`PageModel::surface`): its ExtGState font alone, or the +/// primary's sibling group from the page fonts. +pub fn surface_of(walk: &PageWalk, primary: &ShowRecord) -> TypingSurface { + surface_over(walk, &sibling_fonts(walk), primary) +} + +/// One surface as runs use it. +pub(super) struct SurfaceInfo { + /// Font resources, primary first (empty for an ExtGState font) — `TextRun::surface`. + pub names: Arc<[Vec]>, + pub has_space: bool, + /// Some font of the surface can type something (else `NO_WRITABLE_GLYPHS`). + pub writable: bool, + /// Sibling name → content hash (joins, §A.1.1). + members: HashMap, u64>, +} + +impl SurfaceInfo { + /// No surface: a line that is never typed in (Form text, §A.6 `NESTED_FORM`). + pub(super) fn none() -> SurfaceInfo { + SurfaceInfo { + names: Arc::from(Vec::new()), + has_space: false, + writable: false, + members: HashMap::new(), + } + } +} + +/// The surfaces of one page, built on first use and charged to the page-model budget. +pub(super) struct Surfaces<'w> { + walk: &'w PageWalk, + siblings: Cow<'w, PageFonts>, + names_bytes: usize, + by_key: HashMap>, + /// Per font model (by address; the page fonts keep them alive): a non-empty alphabet. + alphabets: HashMap, +} + +impl<'w> Surfaces<'w> { + pub(super) fn new(walk: &'w PageWalk, mem: &mut ModelBudget) -> Result, Stop> { + let all: usize = walk.page_fonts.iter().map(|(n, _)| n.len()).sum(); + mem.scratch(all.saturating_add(walk.page_fonts.len() * size_of::<(Vec, usize)>()))?; + let siblings = sibling_fonts(walk); + let names_bytes = siblings.iter().map(|(n, _)| n.len()).sum(); + Ok(Surfaces { + walk, + siblings, + names_bytes, + by_key: HashMap::new(), + alphabets: HashMap::new(), + }) + } + + /// The surface of the run whose primary is `primary`. + pub(super) fn of( + &mut self, + primary: &ShowRecord, + mem: &mut ModelBudget, + ) -> Result, Stop> { + let key = surface_key(primary); + if let Some(info) = self.by_key.get(&key) { + return Ok(Arc::clone(info)); + } + // `typing_surface` copies every sibling's name (and a long primary's own). + let own = primary + .before + .text + .font + .as_ref() + .and_then(|f| f.resource.as_ref()) + .map_or(0, |r| r.len()); + let entry = size_of::<(Vec, Arc)>(); + let building = self + .names_bytes + .saturating_add(own) + .saturating_add((self.siblings.len() + 1).saturating_mul(entry)); + mem.scratch(building)?; + let surface = surface_over(self.walk, &self.siblings, primary); + let info = self.info(&key, surface, mem)?; + mem.unscratch(building); + mem.scratch(map_entry::>())?; + self.by_key.insert(key, Arc::clone(&info)); + Ok(info) + } + + fn info( + &mut self, + key: &SurfaceKey, + surface: TypingSurface, + mem: &mut ModelBudget, + ) -> Result, Stop> { + let has_space = surface.has_space(); + let mut writable = false; + for (_, m) in &surface.fonts { + let at = Arc::as_ptr(m) as usize; + let nonempty = match self.alphabets.get(&at) { + Some(known) => *known, + None => { + mem.scratch(map_entry::())?; + let known = !m.alphabet().is_empty(); + self.alphabets.insert(at, known); + known + } + }; + if nonempty { + writable = true; + break; + } + } + let copies: usize = surface.fonts.iter().map(|(n, _)| n.len()).sum(); + let count = surface.fonts.len(); + mem.scratch(copies.saturating_add(count * map_entry::, u64>()))?; + let members: HashMap, u64> = surface + .fonts + .iter() + .map(|(n, m)| (n.clone(), m.content_hash)) + .collect(); + let names: Vec> = if matches!(key, SurfaceKey::ExtGState(_)) { + Vec::new() + } else { + surface.fonts.into_iter().map(|(n, _)| n).collect() + }; + let held: usize = names.iter().map(Vec::capacity).sum(); + mem.hold(held.saturating_add(arc_slice(names.len(), size_of::>())))?; + let info = SurfaceInfo { + names: names.into(), + has_space, + writable, + members, + }; + Ok(Arc::new(info)) + } + + /// Whether `y` (a resource with content hash `y_hash`) is in the sibling group of the resource + /// `x` drawn by `rec` (§A.1.1 fonts join). + pub(super) fn joins( + &mut self, + rec: &ShowRecord, + y: &[u8], + y_hash: u64, + mem: &mut ModelBudget, + ) -> Result { + let info = self.of(rec, mem)?; + Ok(info.members.get(y).is_some_and(|h| *h == y_hash)) + } +} + +/// The bytes a glyph unit keeps beyond its slot: its text and its font resource name. +pub(super) fn unit_heap(u: &Unit) -> usize { + match u { + Unit::Glyph { font_res, text, .. } => { + text.capacity() + font_res.as_ref().map_or(0, Vec::capacity) + } + Unit::Kern { .. } => 0, + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/service.rs b/src-tauri/src/pdf_engine/text_edit/service.rs new file mode 100644 index 0000000..ce50e96 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/service.rs @@ -0,0 +1,129 @@ +//! Command logic of Edit text (SPEC §B.20): open, inspect, preview and release, callable from +//! tests without Tauri. Paths in, JSON-ready DTOs out; every refusal is an `AppError` with a +//! reason code. Open/inspect/preview errors are shown as they are (only Save errors carry +//! "The original file was not changed."). + +use crate::error::AppError; +use crate::pdf_engine::source_content::classify_source_page; +use crate::pdf_engine::text_edit::cache::TextEditCache; +use crate::pdf_engine::text_edit::dto::{self, PageTextDto, TextPreviewDto, TextSourceDto}; +use crate::pdf_engine::text_edit::engines::{Engines, RunOpts, SourceCheck}; +use crate::pdf_engine::text_edit::export::page_refusal; +use crate::pdf_engine::text_edit::preview::preview_page; +use crate::pdf_engine::text_edit::reasons; +use crate::pdf_engine::text_edit::rewrite::TextEditIn; +use std::path::Path; + +/// Sources above this size open and save normally, only more slowly (SPEC §E.7 BENCH-02). +pub const LARGE_SOURCE_BYTES: u64 = 100 << 20; + +/// Opens `path` for Edit text: one bounded read, policy refusals (`ENCRYPTED`, `SIGNED`, +/// `UNSUPPORTED_XFA`, …), the synchronous page-map agreement (`PDF_NEEDS_REPAIR`), and the +/// background `qpdf --check`, which this call does not wait for. +pub fn open_source( + cache: &TextEditCache, + engines: &Engines, + temp_root: &Path, + path: &str, +) -> Result { + let src = cache.open(Path::new(path), temp_root, engines)?; + let snap = &src.ctx.snap; + let mut warnings = Vec::new(); + if snap.fingerprint.len > LARGE_SOURCE_BYTES { + warnings.push( + "This PDF is larger than 100 MB, so checking and saving text changes takes longer." + .to_string(), + ); + } + Ok(TextSourceDto { + fingerprint: snap.fingerprint.to_string(), + page_count: u32::try_from(snap.pages.len()).unwrap_or(u32::MAX), + warnings, + }) +} + +/// Lines of one page (0-based `page_index` in the source) and whether each can be changed. The +/// cached page model goes through the #33 classifier (`classify_source_page`), whose run +/// capabilities and page reason are what the DTO reports; `STALE` when the file changed. +pub fn inspect_page( + cache: &TextEditCache, + engines: &Engines, + temp_root: &Path, + path: &str, + fingerprint: &str, + page_index: u32, +) -> Result { + let src = cache.get(Path::new(path), temp_root, fingerprint, engines)?; + let model = cache.page(&src, page_index)?; + let classified = classify_source_page(&src.ctx, &model, None); + Ok(dto::page_text(&model, &classified)) +} + +/// Plans `edits`, writes them into a one-page copy with qpdf and runs Phase A on it (§B.17). +/// Waits for the source's background `qpdf --check` first: problems are `PDF_NEEDS_REPAIR`, +/// benign warnings are allowed in the patched copy. User problems are verdicts, not errors. +pub fn preview_edits( + cache: &TextEditCache, + engines: &Engines, + temp_root: &Path, + path: &str, + fingerprint: &str, + page_index: u32, + edits: &[TextEditIn], +) -> Result { + let src = cache.get(Path::new(path), temp_root, fingerprint, engines)?; + // A large model is handed over, and freed by the preview once planned (review-final MEDIUM-3). + let model = cache.page_to_release(&src, page_index)?; + // Nothing on a refused page can be changed (inspect lists no runs there). + if let Some(e) = page_refusal(&model, page_index.saturating_add(1)) { + return Err(e); + } + // The preview writes into the source's folder: a release or an eviction meanwhile deletes + // it when the preview ends, not under it (review-T5 M2). + let _lease = cache.lease(&src.dir); + #[cfg(test)] + seams::during_preview(); + let benign = match cache.source_check(&src, engines)?.wait(None)? { + SourceCheck::Clean => Vec::new(), + SourceCheck::Benign(lines) => lines, + SourceCheck::Problems(lines) => return Err(reasons::pdf_needs_repair(&lines)), + }; + let result = preview_page( + &src.ctx, + model, + edits, + &src.dir, + engines, + &benign, + &RunOpts::default(), + )?; + Ok(dto::preview_dto(&result)) +} + +/// Forgets the file and deletes its temporary copies. +pub fn release_source(cache: &TextEditCache, path: &str) { + cache.release(Path::new(path)); +} + +/// Test seam: runs a callback once, inside the next preview on this thread, right after the +/// preview leased its folder (review-T5 M2: a release during a preview). +#[cfg(test)] +pub(crate) mod seams { + use std::cell::RefCell; + + type Hook = Box; + + thread_local! { + static DURING_PREVIEW: RefCell> = const { RefCell::new(None) }; + } + + pub(crate) fn set_during_preview(hook: Option) { + DURING_PREVIEW.with(|h| *h.borrow_mut() = hook); + } + + pub(super) fn during_preview() { + if let Some(hook) = DURING_PREVIEW.with(|h| h.borrow_mut().take()) { + hook(); + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/snapshot.rs b/src-tauri/src/pdf_engine/text_edit/snapshot.rs new file mode 100644 index 0000000..e990150 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/snapshot.rs @@ -0,0 +1,434 @@ +//! One bounded read per operation, hashed and parsed from the same bytes (SPEC §B.4): raw +//! preflight (`snapshot/preflight.rs`: xref chain; `snapshot/objects.rs`: nesting and `/Length` +//! chains), guarded lopdf load (object streams decoded under a budget, their members +//! nesting-checked), policy refusals, and the lopdf/qpdf page-map agreement check. + +use crate::error::AppError; +use crate::pdf_engine::text_edit::decode::{self, DecodeBudget}; +use crate::pdf_engine::text_edit::engines::QpdfPage; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::reasons::{self, EditProblem, EditProblemCode, ProblemCtx}; +use lopdf::{Dictionary, Document, Object, ObjectId}; +use std::io::{Read, Seek, SeekFrom}; +use std::path::{Path, PathBuf}; +use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering}; +use std::sync::{Arc, Mutex}; +use std::time::SystemTime; + +pub(crate) mod headers; +mod objects; +pub(crate) mod preflight; + +const FNV_OFFSET: u64 = 0xcbf2_9ce4_8422_2325; +const FNV_PRIME: u64 = 0x0100_0000_01b3; + +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] +pub struct Fingerprint { + pub len: u64, + pub fnv: u64, +} + +impl std::fmt::Display for Fingerprint { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + write!(f, "{:016x}-{:x}", self.fnv, self.len) + } +} + +impl std::str::FromStr for Fingerprint { + type Err = (); + /// Strict inverse of `Display`: 16 lowercase hex digits, `-`, the length in lowercase hex + /// without leading zeros. + fn from_str(s: &str) -> Result { + let (fnv, len) = s.split_once('-').ok_or(())?; + let lower_hex = + |t: &str| !t.is_empty() && t.bytes().all(|c| matches!(c, b'0'..=b'9' | b'a'..=b'f')); + if fnv.len() != 16 || !lower_hex(fnv) || !lower_hex(len) || len.len() > 16 { + return Err(()); + } + if len.len() > 1 && len.starts_with('0') { + return Err(()); + } + Ok(Fingerprint { + fnv: u64::from_str_radix(fnv, 16).map_err(|_| ())?, + len: u64::from_str_radix(len, 16).map_err(|_| ())?, + }) + } +} + +pub struct SourceSnapshot { + pub path: PathBuf, + pub bytes: Arc>, // the only read of the file + pub fingerprint: Fingerprint, // FNV-1a-64 of `bytes` + len + pub head_tail: (u64, u64), // FNV of the first / last HEAD_TAIL_HASH_BYTES + pub modified: Option, + pub doc: Document, // parsed from `bytes`, never from `path` + pub pages: Vec, // index = 0-based page index (doc.get_pages() order) +} + +impl SourceSnapshot { + /// The file's length (also once `release_bytes` dropped the bytes). + pub fn file_len(&self) -> usize { + usize::try_from(self.fingerprint.len).unwrap_or(usize::MAX) + } + + /// Drops the raw bytes once only the parsed document is read (a Save's source after its + /// copy is written, Phase A/B's verification reads after the page map): the document holds + /// every stream again, so keeping both doubled what a large file costs (review-final + /// MEDIUM-1). Nothing past these points reads `bytes`. + pub fn release_bytes(&mut self) { + self.bytes = Arc::new(Vec::new()); + } +} + +#[derive(Clone, Copy, PartialEq, Eq)] +enum Policy { + Source, + Verification, +} + +pub(crate) fn fnv1a_extend(mut hash: u64, bytes: &[u8]) -> u64 { + for byte in bytes { + hash ^= u64::from(*byte); + hash = hash.wrapping_mul(FNV_PRIME); + } + hash +} + +pub(crate) fn fnv1a_u64(bytes: &[u8]) -> u64 { + fnv1a_extend(FNV_OFFSET, bytes) +} + +fn head_tail_hashes(bytes: &[u8]) -> (u64, u64) { + let n = limits::HEAD_TAIL_HASH_BYTES; + let head = bytes.get(..n.min(bytes.len())).unwrap_or_default(); + let tail = bytes + .get(bytes.len().saturating_sub(n)..) + .unwrap_or_default(); + (fnv1a_u64(head), fnv1a_u64(tail)) +} + +fn path_text(path: &Path) -> String { + path.to_string_lossy().into_owned() +} + +/// Steps 1–2 of `read_snapshot`: regular file, size check, one capped read. +fn read_capped(path: &Path, cap: u64) -> Result<(Vec, Option), AppError> { + if !path.is_file() { + return Err(AppError::invalid_pdf(&path_text(path))); + } + let meta = std::fs::metadata(path) + .map_err(|e| AppError::invalid_pdf(&path_text(path)).with_details(e.to_string()))?; + if meta.len() > cap { + return Err(reasons::file_too_large()); + } + let file = + std::fs::File::open(path).map_err(|e| AppError::io("OffPDF could not open the PDF.", e))?; + let mut bytes = Vec::new(); + bytes + .try_reserve_exact(usize::try_from(meta.len().min(cap)).unwrap_or(0)) + .map_err(|_| reasons::file_too_large())?; + file.take(cap.saturating_add(1)) + .read_to_end(&mut bytes) + .map_err(|e| AppError::io("OffPDF could not read the PDF.", e))?; + if bytes.len() as u64 > cap { + return Err(reasons::file_too_large()); + } + Ok((bytes, meta.modified().ok())) +} + +/// Opens a source PDF for text editing: one capped read, preflight, guarded parse, policy. +pub fn read_snapshot(path: &Path) -> Result { + let (bytes, modified) = read_capped(path, limits::file_cap())?; + snapshot_from_bytes(path, bytes, modified) +} + +/// `read_snapshot` from bytes already in memory (the caller did the one read; `read_snapshot` +/// itself goes through here, so tests that start from bytes run the production path). +pub fn snapshot_from_bytes( + path: &Path, + bytes: Vec, + modified: Option, +) -> Result { + if bytes.len() as u64 > limits::file_cap() { + return Err(reasons::file_too_large()); + } + build(path, bytes, modified, Policy::Source) +} + +/// Steps 1–6 and 8 of `read_snapshot` with `cap` instead of `file_cap()`, and **without** step 7's policy +/// refusals (encrypted, signed, XFA). Used only for files the pipeline wrote (edited copies, preview files, +/// the staged output). Every failure is `EDIT_VERIFY_FAILED` (an encrypted file cannot be read either). +pub fn read_verification_snapshot(path: &Path, cap: u64) -> Result { + read_capped(path, cap) + .and_then(|(bytes, modified)| build(path, bytes, modified, Policy::Verification)) + .map_err(verification_error) +} + +fn verification_error(e: AppError) -> AppError { + let lead = match e.code.as_str() { + "FILE_TOO_LARGE" | "FILE_TOO_COMPLEX" => "too large to verify", + _ => "could not read the checked file", + }; + let detail = format!( + "{lead}: {}: {}", + e.code, + e.details.clone().unwrap_or_else(|| e.message.clone()) + ); + let p = EditProblem::new(EditProblemCode::EditVerifyFailed, Some(detail)); + let ctx = ProblemCtx { + page_number: None, + file_name: None, + face: None, + reason: None, + }; + EditProblemCode::EditVerifyFailed.to_app_error(&p, &ctx) +} + +fn build( + path: &Path, + bytes: Vec, + modified: Option, + policy: Policy, +) -> Result { + let fingerprint = Fingerprint { + len: bytes.len() as u64, + fnv: fnv1a_u64(&bytes), + }; + let head_tail = head_tail_hashes(&bytes); + preflight::preflight(&bytes)?; + let doc = guarded_load(path, &bytes)?; + if document_is_encrypted(&doc) { + return Err(reasons::encrypted()); + } + if policy == Policy::Source { + if document_is_signed(&doc) { + return Err(reasons::signed()); + } + if document_has_xfa(&doc) { + return Err(reasons::unsupported_xfa()); + } + } + if doc.catalog().is_err() { + return Err(reasons::malformed_content( + "The PDF catalog is missing or unreadable.", + )); + } + let pages = doc.get_pages().into_values().collect(); + Ok(SourceSnapshot { + path: path.to_path_buf(), + bytes: Arc::new(bytes), + fingerprint, + head_tail, + modified, + doc, + pages, + }) +} + +/// (len, mtime) unchanged on disk **and** the first/last HEAD_TAIL_HASH_BYTES hash to `head_tail` +/// (catches same-size rewrites on file systems with coarse mtime: FAT/exFAT, network shares). +pub fn stat_matches(snap: &SourceSnapshot) -> bool { + let Ok(meta) = std::fs::metadata(&snap.path) else { + return false; + }; + if meta.len() != snap.fingerprint.len || meta.modified().ok() != snap.modified { + return false; + } + let n = limits::HEAD_TAIL_HASH_BYTES as u64; + let read_at = |offset: u64, len: u64| -> Option> { + let mut f = std::fs::File::open(&snap.path).ok()?; + f.seek(SeekFrom::Start(offset)).ok()?; + let mut buf = Vec::new(); + f.take(len).read_to_end(&mut buf).ok()?; + (buf.len() as u64 == len).then_some(buf) + }; + let len = meta.len(); + let head = read_at(0, len.min(n)); + let tail = read_at(len.saturating_sub(n), len.min(n)); + match (head, tail) { + (Some(h), Some(t)) => (fnv1a_u64(&h), fnv1a_u64(&t)) == snap.head_tail, + _ => false, + } +} + +/// Page-map agreement (D14): `qpdf` must list the same page count, the same page object ids in order, and for +/// each page the same `/Contents` stream ids as lopdf. Mismatch → PDF_NEEDS_REPAIR (details: first difference). +pub fn check_page_map(snap: &SourceSnapshot, qpdf_pages: &[QpdfPage]) -> Result<(), AppError> { + let repair = |what: String| { + Err(reasons::pdf_needs_repair(&[format!( + "page map disagreement: {what}" + )])) + }; + if snap.pages.len() != qpdf_pages.len() { + return repair(format!( + "lopdf reads {} pages, qpdf {}", + snap.pages.len(), + qpdf_pages.len() + )); + } + for (i, (ours, theirs)) in snap.pages.iter().zip(qpdf_pages).enumerate() { + let n = i + 1; + if *ours != theirs.object { + return repair(format!( + "page {n}: lopdf object {} {} R, qpdf {} {} R", + ours.0, ours.1, theirs.object.0, theirs.object.1 + )); + } + match page_contents_ids(&snap.doc, *ours) { + Some(ids) if ids == theirs.contents => {} + Some(ids) => { + return repair(format!( + "page {n}: /Contents {ids:?} (lopdf) vs {:?} (qpdf)", + theirs.contents + )) + } + None => return repair(format!("page {n}: /Contents unresolvable for lopdf")), + } + } + Ok(()) +} + +/// The content stream ids of a page as qpdf lists them; `None` when lopdf cannot resolve them. +fn page_contents_ids(doc: &Document, page_id: ObjectId) -> Option> { + let page = doc.get_dictionary(page_id).ok()?; + let refs_of = |items: &[Object]| -> Option> { + items + .iter() + .map(|o| { + o.as_reference() + .ok() + .filter(|id| matches!(doc.objects.get(id), Some(Object::Stream(_)))) + }) + .collect() + }; + match page.get(b"Contents").ok() { + None | Some(Object::Null) => Some(Vec::new()), + Some(Object::Reference(id)) => match doc.objects.get(id)? { + Object::Stream(_) => Some(vec![*id]), + Object::Array(items) => refs_of(items), + _ => None, + }, + Some(Object::Array(items)) => refs_of(items), + Some(_) => None, + } +} + +fn document_is_encrypted(doc: &Document) -> bool { + doc.is_encrypted() || doc.trailer.get(b"Encrypt").is_ok_and(|o| !o.is_null()) +} + +/// Catalog `/Perms`, or any `/Type /Sig` or `/FT /Sig` dictionary **with** `/ByteRange` +/// (same rule as the #33 classifier; empty signature widgets are fine). +pub(crate) fn document_is_signed(doc: &Document) -> bool { + if doc.catalog().is_ok_and(|cat| cat.has(b"Perms")) { + return true; + } + doc.objects.values().any(|obj| { + let dict = match obj { + Object::Dictionary(d) => d, + Object::Stream(s) => &s.dict, + _ => return false, + }; + (name_is(dict, b"Type", b"Sig") || name_is(dict, b"FT", b"Sig")) && dict.has(b"ByteRange") + }) +} + +fn name_is(dict: &Dictionary, key: &[u8], expect: &[u8]) -> bool { + dict.get(key).ok().and_then(|o| o.as_name().ok()) == Some(expect) +} + +/// `/AcroForm /XFA` present, or `/NeedsRendering true` (same rule as Fill forms' `detect_xfa`). +pub(crate) fn document_has_xfa(doc: &Document) -> bool { + let Ok(cat) = doc.catalog() else { return false }; + let needs_rendering = |d: &Dictionary| match d.get(b"NeedsRendering") { + Ok(Object::Boolean(true)) => true, + Ok(Object::Integer(i)) => *i != 0, + _ => false, + }; + let acro = match cat.get(b"AcroForm") { + Ok(Object::Reference(id)) => doc.get_dictionary(*id).ok(), + Ok(Object::Dictionary(d)) => Some(d), + _ => None, + }; + needs_rendering(cat) + || acro.is_some_and(|a| a.get(b"XFA").is_ok_and(|x| !x.is_null()) || needs_rendering(a)) +} + +// ---- Guarded lopdf load ------------------------------------------------------------------ + +static LOAD_LOCK: Mutex<()> = Mutex::new(()); +static OBJSTM_BUDGET: AtomicUsize = AtomicUsize::new(0); +static OBJSTM_SCAN_BUDGET: AtomicUsize = AtomicUsize::new(0); +static OBJSTM_REJECTED: AtomicBool = AtomicBool::new(false); +const REJECTED_OBJSTM: &[u8] = b"OffPdfRejectedObjStm"; + +fn guarded_load(path: &Path, bytes: &[u8]) -> Result { + let _lock = LOAD_LOCK.lock().unwrap_or_else(|e| e.into_inner()); + OBJSTM_BUDGET.store(limits::OBJSTM_TOTAL_DECODED, Ordering::SeqCst); + OBJSTM_SCAN_BUDGET.store(limits::objstm_scan_budget(), Ordering::SeqCst); + OBJSTM_REJECTED.store(false, Ordering::SeqCst); + let doc = lopdf::Reader { + buffer: bytes, + document: Document::new(), + } + .read(Some(objstm_guard)) + .map_err(|e| AppError::invalid_pdf(&path_text(path)).with_details(format!("lopdf: {e}")))?; + let rejected = OBJSTM_REJECTED.load(Ordering::SeqCst) + || doc + .objects + .values() + .any(|o| matches!(o, Object::Stream(s) if s.dict.type_is(REJECTED_OBJSTM))); + if rejected { + return Err(reasons::file_too_complex( + "an object stream is too large, too deeply nested or cannot be decoded", + )); + } + if doc.objects.len() > limits::MAX_OBJECTS { + return Err(reasons::file_too_complex("too many objects")); + } + Ok(doc) +} + +/// lopdf filter: decodes object streams under the per-stream and total budgets (so lopdf's own +/// unbounded `decompress` becomes a no-op) and nesting-checks the members lopdf will parse from +/// them under one scan budget per load, rejects the rest, and never clones a stream. +fn objstm_guard(id: (u32, u16), obj: &mut Object) -> Option<((u32, u16), Object)> { + let Object::Stream(stream) = obj else { + // Inside object streams lopdf uses the returned object; at top level it is ignored. + return Some((id, obj.clone())); + }; + if stream.dict.type_is(b"ObjStm") { + let remaining = OBJSTM_BUDGET.load(Ordering::SeqCst); + let cap = limits::OBJSTM_MAX_DECODED.min(remaining); + let declared = stream.dict.get(b"N").and_then(Object::as_i64).unwrap_or(0); + let decoded = decode::decode_stream(stream, cap, &mut DecodeBudget::new(usize::MAX)); + let accepted = match decoded { + Ok(data) + if !(data.is_empty() && declared > 0) + && objects::objstm_members_ok(&stream.dict, &data, &OBJSTM_SCAN_BUDGET) => + { + let debited = OBJSTM_BUDGET + .fetch_update(Ordering::SeqCst, Ordering::SeqCst, |r| { + r.checked_sub(data.len()) + }) + .is_ok(); + if debited { + stream.dict.remove(b"Filter"); + stream.dict.remove(b"DecodeParms"); + stream.set_content(data); + } + debited + } + _ => false, + }; + if !accepted { + stream + .dict + .set("Type", Object::Name(REJECTED_OBJSTM.to_vec())); + stream.content.clear(); + OBJSTM_REJECTED.store(true, Ordering::SeqCst); + } + } + // Streams are always top level, where lopdf ignores the returned value: never clone them. + Some((id, Object::Null)) +} diff --git a/src-tauri/src/pdf_engine/text_edit/snapshot/headers.rs b/src-tauri/src/pdf_engine/text_edit/snapshot/headers.rs new file mode 100644 index 0000000..3e2e5e9 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/snapshot/headers.rs @@ -0,0 +1,165 @@ +//! Every `N G obj` header lopdf 0.34's `_indirect_object` can parse, found in linear time. +//! +//! An xref offset may point at any digit, so every digit run in the file is a candidate first +//! number. After it lopdf needs `space G space obj`, where `space` is whitespace and `%` comments +//! (a comment runs to the next CR or LF). Probing each run naively rescans long comments, long +//! whitespace and long digit runs once per candidate, which is quadratic on crafted input +//! (review-T1 round 2, HIGH-1). Two memos make every byte cost O(1) probes: +//! - `runs`: for a `%` at `q`, the end `e` of lopdf's `space` from `q`. Any `%` in `[q, e)` either +//! starts or sits inside a comment that ends at the same EOL, after which the walk is the same, +//! so its `space` also ends at `e`. Stored intervals are therefore disjoint; one lookup answers +//! every `%` inside them. +//! - `gens`: the result after the first `space`, keyed by where the generation number starts (it +//! depends only on that position), so many candidates reaching the same generation number do +//! not each rescan it and the `space` after it. Two candidates can only reach the same one when +//! a comment hides the later candidate from the earlier one, so only probes whose first `space` +//! went through a comment are stored (the others would cost an allocation per digit run). +//! +//! Probes never look behind the current digit run, so entries behind it are dropped and both memos +//! stay small. + +use crate::pdf_engine::text_edit::lexer; +use std::collections::BTreeMap; +use std::ops::Range; + +/// Iterator over `(first number run, offset after obj)` of every header candidate in a buffer. +pub(crate) struct Headers<'a> { + b: &'a [u8], + next: usize, + runs: BTreeMap, + gens: BTreeMap>, +} + +impl<'a> Headers<'a> { + pub(crate) fn new(b: &'a [u8]) -> Self { + Headers { + b, + next: 0, + runs: BTreeMap::new(), + gens: BTreeMap::new(), + } + } + + /// Drops memo entries no probe starting at `from` or later can reach. + fn prune(&mut self, from: usize) { + while self + .runs + .first_key_value() + .is_some_and(|(_, end)| *end <= from) + { + self.runs.pop_first(); + } + while self + .gens + .first_key_value() + .is_some_and(|(at, _)| *at < from) + { + self.gens.pop_first(); + } + } + + /// lopdf's `space` from `at` (lenient: a comment may end at EOF), and whether it went + /// through a comment. + fn space(&mut self, at: usize) -> (usize, bool) { + let mut i = at; + while let Some(&c) = self.b.get(i) { + if c == b'%' { + return (self.comment_space(i), true); + } + if !lexer::is_whitespace(c) { + break; + } + i += 1; + } + (i, false) + } + + fn known(&self, q: usize) -> Option { + self.runs + .range(..=q) + .next_back() + .map(|(_, end)| *end) + .filter(|end| q < *end) + } + + /// End of lopdf's `space` from the `%` at `q`, memoised as the interval `[q, end)`. + fn comment_space(&mut self, q: usize) -> usize { + if let Some(end) = self.known(q) { + return end; + } + let mut i = q; + let end = loop { + match self.b.get(i) { + Some(b'%') => { + if let Some(end) = self.known(i) { + break end; + } + i = line_end(self.b, i); + } + Some(c) if lexer::is_whitespace(*c) => i += 1, + _ => break i, + } + }; + // Intervals starting inside [q, end) end at `end` too: merge them into this one. + let inner: Vec = self.runs.range(q..end).map(|(start, _)| *start).collect(); + for start in inner { + self.runs.remove(&start); + } + self.runs.insert(q, end); + end + } + + /// After a first number ending at `run_end`: `space G space obj` → the offset after `obj`. + fn header_end(&mut self, run_end: usize) -> Option { + let (gen_at, via_comment) = self.space(run_end); + if let Some(found) = self.gens.get(&gen_at) { + return *found; + } + let gen_end = digits_end(self.b, gen_at); + let found = if gen_end == gen_at { + None + } else { + let (obj_at, _) = self.space(gen_end); + let obj_end = obj_at.saturating_add(3); + (self.b.get(obj_at..obj_end) == Some(&b"obj"[..])).then_some(obj_end) + }; + if via_comment { + self.gens.insert(gen_at, found); + } + found + } +} + +impl Iterator for Headers<'_> { + type Item = (Range, usize); + + fn next(&mut self) -> Option { + loop { + let rel = self + .b + .get(self.next..)? + .iter() + .position(u8::is_ascii_digit)?; + let start = self.next + rel; + let end = digits_end(self.b, start); + self.next = end; + self.prune(end); + if let Some(value_at) = self.header_end(end) { + return Some((start..end, value_at)); + } + } + } +} + +fn digits_end(b: &[u8], from: usize) -> usize { + from + b + .get(from..) + .map_or(0, |r| r.iter().take_while(|c| c.is_ascii_digit()).count()) +} + +/// The CR or LF that ends the comment at `from` (or the end of the buffer). +fn line_end(b: &[u8], from: usize) -> usize { + b.get(from..) + .and_then(|r| r.iter().position(|c| matches!(c, b'\r' | b'\n'))) + .map_or(b.len(), |p| from + p) +} diff --git a/src-tauri/src/pdf_engine/text_edit/snapshot/objects.rs b/src-tauri/src/pdf_engine/text_edit/snapshot/objects.rs new file mode 100644 index 0000000..a27b9f4 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/snapshot/objects.rs @@ -0,0 +1,461 @@ +//! Object preflight (SPEC §B.4, R3): bounds lopdf 0.34's two recursions before it parses. +//! +//! - Nesting. lopdf's `nom_parser` recurses once per `[`/`<<` level. It starts parsing at every +//! xref offset (`N G obj`, then one value) and at every object-stream member offset. An xref +//! offset may point anywhere, also inside a string, a comment or stream data, so the value after +//! **every** `N G obj` header in the file is scanned, wherever it sits (`headers.rs` finds them +//! in linear time), and every object-stream member is scanned after `objstm_guard` decodes it. A scan counts depth until the value ends or +//! until the first byte lopdf's grammar cannot consume (lopdf's descent stops there too); it is +//! lenient wherever lopdf is strict, so the depth it sees is never below lopdf's. +//! - `/Length` chains. While parsing a stream whose `/Length` is a reference, lopdf parses the +//! referenced object, which recurses again if that is a stream with a `/Length` reference. The +//! chain is bounded by `LENGTH_REF_CHAIN_MAX` over a graph that over-approximates which object +//! an id can resolve to: an xref offset may point inside a header's first number, so a header +//! stands for every numeric suffix of it. + +use super::headers::Headers; +use super::preflight::find_from; +use crate::error::AppError; +use crate::pdf_engine::text_edit::lexer; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::reasons; +use lopdf::{Dictionary, Object}; +use std::ops::Range; +use std::sync::atomic::{AtomicUsize, Ordering}; + +/// Bytes the value scans over one buffer may read in total. In a well-formed file every value +/// is read once; scans that overlap (headers inside strings) are what this bounds. +struct ScanBudget(usize); + +impl ScanBudget { + fn for_len(len: usize) -> Self { + ScanBudget( + len.saturating_mul(limits::OBJECT_SCAN_FACTOR) + .saturating_add(limits::OBJECT_SCAN_SLACK_BYTES), + ) + } + + fn debit(&mut self, n: usize) -> Result<(), ScanError> { + self.0 = self.0.checked_sub(n).ok_or(ScanError::Budget)?; + Ok(()) + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum ScanError { + TooDeep, + Budget, +} + +impl ScanError { + fn to_app_error(self) -> AppError { + reasons::file_too_complex(&match self { + ScanError::TooDeep => format!( + "objects nested more than {} levels deep", + limits::MAX_OBJECT_NESTING + ), + ScanError::Budget => "object scan budget exceeded".to_string(), + }) + } +} + +/// A stream (`N G obj << … /Length T G R … >> stream`) whose header's first number spans `run`. +struct LengthRef { + run: Range, + target: u32, +} + +/// Scans every `N G obj` header in `b` (nesting of its value) and bounds `/Length` chains. +pub(super) fn check_objects(b: &[u8]) -> Result<(), AppError> { + let mut budget = ScanBudget::for_len(b.len()); + let mut refs: Vec = Vec::new(); + for (run, value_at) in Headers::new(b) { + let target = scan_value(b, value_at, &mut budget).map_err(ScanError::to_app_error)?; + if let Some(target) = target { + if refs.len() >= limits::LENGTH_REF_STREAMS_MAX { + return Err(reasons::file_too_complex( + "too many streams with a /Length reference", + )); + } + refs.push(LengthRef { run, target }); + } + } + let longest = longest_length_chain(b, &refs); + if longest > limits::LENGTH_REF_CHAIN_MAX { + return Err(reasons::file_too_complex(&format!( + "stream /Length references may chain through {longest} objects" + ))); + } + Ok(()) +} + +/// lopdf's `ObjectStream::new` over the decoded data of an object stream: the members it would +/// parse (same header parsing) nest at most `MAX_OBJECT_NESTING` levels. `false` → reject it. +/// The scans of all object streams of one load draw on `shared` (review-T1 round 2, MEDIUM-1): +/// each stream may read up to 2 × its data + 1 MiB, but never more than is left, and what it +/// read is debited afterwards. Streams scanned concurrently on rayon workers can overdraw it by +/// at most one such grant each. +pub(super) fn objstm_members_ok(dict: &Dictionary, data: &[u8], shared: &AtomicUsize) -> bool { + let first = dict.get(b"First").and_then(Object::as_i64).ok(); + let Some(first) = first.and_then(|f| usize::try_from(f).ok()) else { + return true; // lopdf drops the stream without parsing a member + }; + let header = data.get(..first).and_then(|h| std::str::from_utf8(h).ok()); + let (Some(header), true) = (header, dict.get(b"N").and_then(Object::as_i64).is_ok()) else { + return true; + }; + let numbers: Vec> = header + .split_whitespace() + .map(|n| n.parse::().ok()) + .collect(); + let granted = ScanBudget::for_len(data.len()) + .0 + .min(shared.load(Ordering::SeqCst)); + let mut budget = ScanBudget(granted); + let members_ok = numbers.chunks_exact(2).all(|pair| { + let (Some(Some(_)), Some(Some(off))) = (pair.first(), pair.get(1)) else { + return true; + }; + let offset = usize::try_from(*off) + .ok() + .and_then(|o| first.checked_add(o)); + match offset { + Some(offset) if offset < data.len() => scan_value(data, offset, &mut budget).is_ok(), + _ => true, + } + }); + let used = granted.saturating_sub(budget.0); + let debited = shared + .fetch_update(Ordering::SeqCst, Ordering::SeqCst, |left| { + left.checked_sub(used) + }) + .is_ok(); + members_ok && debited +} + +/// lopdf's `space`: whitespace and `%` comments (lenient: a comment may end at EOF). +fn skip_space(b: &[u8], mut i: usize) -> usize { + loop { + match b.get(i) { + Some(c) if lexer::is_whitespace(*c) => i += 1, + Some(b'%') => { + i += b.get(i..).map_or(0, |r| { + r.iter().take_while(|c| !matches!(c, b'\r' | b'\n')).count() + }); + } + _ => return i, + } + } +} + +/// Bytes lopdf can consume as (parts of) `null`, `true`, `false`, numbers and the `R` of a +/// reference; a regular word with any other byte stops lopdf's parse. +const VALUE_WORD_BYTES: &[u8] = b"0123456789+-.Rtruefalsn"; + +fn regular_end(b: &[u8], from: usize) -> usize { + from + b.get(from..).map_or(0, |r| { + r.iter() + .take_while(|c| !lexer::is_whitespace(**c) && !lexer::is_delimiter(**c)) + .count() + }) +} + +/// End of the literal string opening at `i` (balanced parentheses, `\` escapes). +fn skip_literal(b: &[u8], mut i: usize) -> usize { + let mut depth = 0usize; + while let Some(&c) = b.get(i) { + i += 1; + match c { + b'\\' => i += 1, + b'(' => depth += 1, + b')' => { + depth = depth.saturating_sub(1); + if depth == 0 { + break; + } + } + _ => {} + } + } + i.min(b.len()) +} + +/// Scans the one value lopdf parses at `at` (after `obj`, or an object-stream member) and returns +/// the target id of its `/Length N G R` when the value is a stream dictionary. +fn scan_value(b: &[u8], at: usize, budget: &mut ScanBudget) -> Result, ScanError> { + let mut i = at; + let mut depth = 0usize; + let mut top_dict = false; + let mut length = LengthState::Idle; + let mut length_ref = None; + let closed_dict = loop { + i = skip_space(b, i); + let Some(&c) = b.get(i) else { break false }; + let pair = b.get(i + 1) == Some(&c); + if c == b'[' || (c == b'<' && pair) { + if depth == 0 { + top_dict = c == b'<'; + } + depth += 1; + if depth > limits::MAX_OBJECT_NESTING { + return Err(ScanError::TooDeep); + } + length = LengthState::Idle; + i += if c == b'[' { 1 } else { 2 }; + continue; + } + if c == b']' || (c == b'>' && pair) { + let Some(d) = depth.checked_sub(1) else { + break false; + }; + depth = d; + i += if c == b']' { 1 } else { 2 }; + if depth == 0 { + break top_dict; + } + continue; + } + let end = match c { + b'(' => skip_literal(b, i), + b'<' => find_from(b, i, b">").map_or(b.len(), |p| p + 1), + b'/' => regular_end(b, i + 1), + _ if lexer::is_delimiter(c) => break false, + _ => { + let end = regular_end(b, i); + let word = b.get(i..end).unwrap_or_default(); + if !word.iter().all(|t| VALUE_WORD_BYTES.contains(t)) { + break false; + } + end + } + }; + let token = b.get(i..end).unwrap_or_default(); + if depth == 1 && top_dict { + length = length.step(token, &mut length_ref); + } + i = end; + if depth == 0 { + break false; + } + }; + // The `stream` probe after a closed dictionary reads too: overlapping headers that close at + // one `>>` would otherwise each skip the same trailing space unbudgeted (round 2, HIGH-2). + let end = if closed_dict { skip_space(b, i) } else { i }; + budget.debit(end.saturating_sub(at))?; + let stream = closed_dict && b.get(end..).is_some_and(|rest| rest.starts_with(b"stream")); + Ok(length_ref.filter(|_| stream)) +} + +/// Recognises `/Length N G R` among the direct entries of a dictionary (the last one counts, as +/// in lopdf's `Dictionary::set`; a later direct `/Length` only makes this over-approximate). +#[derive(Clone, Copy)] +enum LengthState { + Idle, + Key, + Id(u32), + Gen(u32), +} + +impl LengthState { + fn step(self, token: &[u8], found: &mut Option) -> LengthState { + if let Some(name) = token.strip_prefix(b"/") { + return if name_is(name, b"Length") { + LengthState::Key + } else { + LengthState::Idle + }; + } + let digits = token.iter().take_while(|c| c.is_ascii_digit()).count(); + let (number, rest) = token.split_at_checked(digits).unwrap_or((token, &[])); + match self { + LengthState::Key if rest.is_empty() => { + parse_u32(number).map_or(LengthState::Idle, LengthState::Id) + } + LengthState::Id(id) if digits > 0 && rest.is_empty() => LengthState::Gen(id), + LengthState::Id(id) if digits > 0 && rest.starts_with(b"R") => { + *found = Some(id); + LengthState::Idle + } + LengthState::Gen(id) if token.starts_with(b"R") => { + *found = Some(id); + LengthState::Idle + } + _ => LengthState::Idle, + } + } +} + +fn parse_u32(digits: &[u8]) -> Option { + std::str::from_utf8(digits).ok()?.parse().ok() +} + +/// A name's bytes (after `/`) decoded like lopdf (`#xx`; an invalid `#` ends the name) equal +/// `want`. +fn name_is(raw: &[u8], want: &[u8]) -> bool { + let hex = |c: u8| (c as char).to_digit(16); + let mut out = want.iter(); + let mut k = 0; + while let Some(&c) = raw.get(k) { + let byte = if c == b'#' { + match ( + raw.get(k + 1).and_then(|c| hex(*c)), + raw.get(k + 2).and_then(|c| hex(*c)), + ) { + (Some(h), Some(l)) => { + k += 3; + h * 16 + l + } + _ => break, + } + } else { + k += 1; + u32::from(c) + }; + if out.next().map(|w| u32::from(*w)) != Some(byte) { + return false; + } + } + out.next().is_none() +} + +/// Every id an xref offset inside `run` can make lopdf parse: each numeric suffix that fits u32. +fn suffix_ids(run: &[u8]) -> Vec { + let mut ids: Vec = (1..=run.len().min(10)) + .filter_map(|n| parse_u32(run.get(run.len() - n..)?)) + .collect(); + ids.dedup(); + ids +} + +/// Upper bound on the objects lopdf parses for one top-level stream while resolving `/Length` +/// references: the stream itself plus the longest path of ids, each strongly connected component +/// counted with its size (lopdf never resolves an id twice in one chain). +fn longest_length_chain(b: &[u8], refs: &[LengthRef]) -> usize { + if refs.is_empty() { + return 0; + } + let mut nodes: Vec = refs.iter().map(|r| r.target).collect(); + nodes.sort_unstable(); + nodes.dedup(); + let node = |id: u32| nodes.binary_search(&id).ok(); + let mut edges: Vec<(usize, usize)> = Vec::new(); + for r in refs { + let Some(to) = node(r.target) else { continue }; + let run = b.get(r.run.clone()).unwrap_or_default(); + edges.extend( + suffix_ids(run) + .into_iter() + .filter_map(node) + .map(|from| (from, to)), + ); + } + edges.sort_unstable(); + edges.dedup(); + longest_scc_path(nodes.len(), &edges).saturating_add(1) +} + +/// Tarjan's strongly connected components (iterative) over `n` nodes and sorted `edges`; returns +/// the largest sum of component sizes along any path of the condensation. +fn longest_scc_path(n: usize, edges: &[(usize, usize)]) -> usize { + const UNSEEN: usize = usize::MAX; + let out = |v: usize| { + let lo = edges.partition_point(|e| e.0 < v); + let hi = edges.partition_point(|e| e.0 <= v); + edges.get(lo..hi).unwrap_or_default() + }; + let mut index = vec![UNSEEN; n]; + let mut low = vec![0usize; n]; + let mut on_stack = vec![false; n]; + let mut comp = vec![UNSEEN; n]; + let mut best: Vec = Vec::new(); + let mut stack: Vec = Vec::new(); + let mut counter = 0usize; + for root in 0..n { + if index.get(root) != Some(&UNSEEN) { + continue; + } + let mut calls: Vec<(usize, usize)> = vec![(root, 0)]; + visit( + root, + &mut counter, + &mut index, + &mut low, + &mut on_stack, + &mut stack, + ); + while let Some((v, pos)) = calls.last().copied() { + if let Some(&(_, w)) = out(v).get(pos) { + if let Some(top) = calls.last_mut() { + top.1 += 1; + } + if index.get(w) == Some(&UNSEEN) { + visit( + w, + &mut counter, + &mut index, + &mut low, + &mut on_stack, + &mut stack, + ); + calls.push((w, 0)); + } else if on_stack.get(w) == Some(&true) { + let iw = index.get(w).copied().unwrap_or(UNSEEN); + if let Some(lv) = low.get_mut(v) { + *lv = (*lv).min(iw); + } + } + continue; + } + calls.pop(); + let lv = low.get(v).copied().unwrap_or(0); + if let Some(&(u, _)) = calls.last() { + if let Some(lu) = low.get_mut(u) { + *lu = (*lu).min(lv); + } + } + if Some(&lv) != index.get(v) { + continue; + } + let c = best.len(); + let mut members = Vec::new(); + while let Some(w) = stack.pop() { + if let Some(s) = on_stack.get_mut(w) { + *s = false; + } + if let Some(cw) = comp.get_mut(w) { + *cw = c; + } + members.push(w); + if w == v { + break; + } + } + let successors = members + .iter() + .flat_map(|m| out(*m).iter()) + .filter_map(|&(_, w)| comp.get(w).copied().filter(|cw| *cw != c && *cw != UNSEEN)) + .filter_map(|cw| best.get(cw).copied()) + .max() + .unwrap_or(0); + best.push(members.len().saturating_add(successors)); + } + } + best.into_iter().max().unwrap_or(0) +} + +fn visit( + v: usize, + counter: &mut usize, + index: &mut [usize], + low: &mut [usize], + on_stack: &mut [bool], + stack: &mut Vec, +) { + if let (Some(i), Some(l), Some(s)) = (index.get_mut(v), low.get_mut(v), on_stack.get_mut(v)) { + *i = *counter; + *l = *counter; + *s = true; + } + *counter += 1; + stack.push(v); +} diff --git a/src-tauri/src/pdf_engine/text_edit/snapshot/preflight.rs b/src-tauri/src/pdf_engine/text_edit/snapshot/preflight.rs new file mode 100644 index 0000000..128c9b6 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/snapshot/preflight.rs @@ -0,0 +1,363 @@ +//! Raw preflight (SPEC §B.4): bounds what lopdf 0.34's parser does before it runs. lopdf +//! allocates from xref-stream fields without checks and recurses once per array/dictionary level +//! and once per stream `/Length` reference it resolves, on rayon worker stacks (2 MiB), so every +//! input that would make it allocate or recurse without bound is refused here: +//! - the cross-reference chain (≤ `MAX_XREF_CHAIN` hops incl. `/XRefStm`), each xref stream's +//! decoded size, `/W`, `/Index` and PNG-predictor row (`/DecodeParms`); +//! - every object lopdf could parse, and every `/Length` chain (`objects.rs`). +//! +//! The chain must be the one lopdf reads (review-T1 round 2, HIGH-3): a dictionary key that +//! occurs twice counts with its **last** value (lopdf's `Dictionary::set`), a table's trailer is +//! the one lopdf's grammar reaches (comments skipped, not the first `trailer` bytes), and +//! `/XRefStm` is followed from xref-stream trailers too. + +use super::objects; +use crate::error::AppError; +use crate::pdf_engine::text_edit::decode::{self, DecodeError}; +use crate::pdf_engine::text_edit::lexer::{self, Operand, Token}; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::reasons; + +/// Runs every preflight check on the bytes of one snapshot (also driven by the preflight fuzz). +pub(crate) fn preflight(bytes: &[u8]) -> Result<(), AppError> { + check_xref_chain(bytes)?; + objects::check_objects(bytes) +} + +pub(super) fn find_from(hay: &[u8], from: usize, needle: &[u8]) -> Option { + let first = *needle.first()?; + let mut i = from; + while let Some(rel) = hay.get(i..)?.iter().position(|c| *c == first) { + let p = i + rel; + if hay.get(p..p.checked_add(needle.len())?) == Some(needle) { + return Some(p); + } + i = p + 1; + } + None +} + +fn skip_ws(b: &[u8], mut i: usize) -> usize { + while b.get(i).is_some_and(|c| lexer::is_whitespace(*c)) { + i += 1; + } + i +} + +fn parse_uint(b: &[u8], i: usize) -> Option<(u64, usize)> { + let digits = b + .get(i..)? + .iter() + .take_while(|c| c.is_ascii_digit()) + .count(); + if digits == 0 || digits > 19 { + return None; + } + let text = std::str::from_utf8(b.get(i..i + digits)?).ok()?; + Some((text.parse().ok()?, i + digits)) +} + +fn find_startxref(b: &[u8]) -> Option { + let tail_start = b.len().saturating_sub(limits::XREF_TAIL_SEARCH_BYTES); + let tail = b.get(tail_start..)?; + let rel = tail.windows(9).rposition(|w| w == b"startxref")?; + let (off, _) = parse_uint(b, skip_ws(b, tail_start + rel + 9))?; + usize::try_from(off).ok().filter(|o| *o < b.len()) +} + +/// Minimal value tree of a scanned dictionary (only what the preflight reads). `Num` is a number +/// written as an integer (lopdf's `Integer`), `Real` one written with a `.`. +#[derive(Debug)] +enum PVal { + Num(f64), + Real(f64), + Name(Vec), + Array(Vec), + Dict(Vec<(Vec, PVal)>), + Other, +} + +fn tokens_to_value(b: &[u8], tokens: Vec) -> Option { + let mut stack: Vec<(bool, Vec)> = Vec::new(); + let mut done = None; + for t in tokens { + let v = match t { + Token::DictOpen(_) => { + stack.push((true, Vec::new())); + continue; + } + Token::ArrayOpen(_) => { + stack.push((false, Vec::new())); + continue; + } + Token::DictClose(_) | Token::ArrayClose(_) => { + let (is_dict, items) = stack.pop()?; + if !is_dict { + PVal::Array(items) + } else { + let mut entries = Vec::new(); + let mut it = items.into_iter(); + while let Some(k) = it.next() { + let PVal::Name(k) = k else { return None }; + entries.push((k, it.next()?)); + } + PVal::Dict(entries) + } + } + Token::Keyword { bytes, .. } if bytes == b"R" => { + let items = &mut stack.last_mut()?.1; + let (Some(PVal::Num(_)), Some(PVal::Num(_))) = (items.pop(), items.pop()) else { + return None; + }; + PVal::Other + } + Token::Keyword { .. } => PVal::Other, + Token::Operand(Operand::Number { value, span }) => { + if b.get(span).is_some_and(|text| text.contains(&b'.')) { + PVal::Real(value) + } else { + PVal::Num(value) + } + } + Token::Operand(Operand::Name { bytes, .. }) => PVal::Name(bytes), + Token::Operand(_) => PVal::Other, + }; + match stack.last_mut() { + Some((_, items)) => items.push(v), + None => done = Some(v), + } + } + done +} + +/// The value lopdf keeps for `key`: the last occurrence (`Dictionary::set` replaces). +fn dict_get<'a>(d: &'a [(Vec, PVal)], key: &[u8]) -> Option<&'a PVal> { + d.iter().rev().find(|(k, _)| k == key).map(|(_, v)| v) +} + +fn dict_uint(d: &[(Vec, PVal)], key: &[u8]) -> Option { + match dict_get(d, key)? { + PVal::Num(v) if *v >= 0.0 && v.fract() == 0.0 && *v < 1e15 => Some(*v as usize), + _ => None, + } +} + +fn scan_dict(b: &[u8], at: usize) -> Result<(Vec<(Vec, PVal)>, usize), AppError> { + let (tokens, end) = + lexer::scan_dict_at(b, at, limits::PREFLIGHT_DICT_TOKENS_MAX).map_err(|e| match e { + lexer::LexError::TooComplex { what } => reasons::file_too_complex(what), + other => reasons::invalid_xref(&other.to_string()), + })?; + match tokens_to_value(b, tokens) { + Some(PVal::Dict(d)) => Ok((d, end)), + _ => Err(reasons::invalid_xref(&format!( + "unreadable dictionary at byte {at}" + ))), + } +} + +fn check_xref_chain(b: &[u8]) -> Result<(), AppError> { + let start = find_startxref(b).ok_or_else(|| reasons::invalid_xref("no startxref"))?; + let mut queue = vec![start]; + let mut visited = std::collections::HashSet::new(); + while let Some(off) = queue.pop() { + if !visited.insert(off) { + continue; + } + if visited.len() > limits::MAX_XREF_CHAIN { + return Err(reasons::file_too_complex("cross-reference chain too long")); + } + let p = skip_ws(b, off); + let dict = if b.get(p..p + 4) == Some(&b"xref"[..]) { + let t = table_trailer(b, p + 4).ok_or_else(|| reasons::invalid_xref("no trailer"))?; + scan_dict(b, t + 7)?.0 + } else { + check_xref_stream(b, p)? + }; + // lopdf reads `/XRefStm` from its newest trailer of either kind; every trailer counts here. + if let Some(stm) = dict_uint(&dict, b"XRefStm") { + queue.push(stm); + } + if dict_get(&dict, b"Prev").is_some() { + let prev = + dict_uint(&dict, b"Prev").ok_or_else(|| reasons::invalid_xref("bad /Prev"))?; + queue.push(prev); + } + } + Ok(()) +} + +/// Where lopdf's `xref` parser (sections of digits, spaces, `n`/`f` and EOLs) and the `space` +/// after it (whitespace and `%` comments) leave off: the `trailer` keyword it parses next. Every +/// byte lopdf consumes there is skipped here with the same comment boundaries, so when lopdf +/// reads a trailer it is this one; a `trailer` inside a comment is never taken. +fn table_trailer(b: &[u8], from: usize) -> Option { + let mut i = from; + loop { + match *b.get(i)? { + c if c.is_ascii_digit() || c == b'n' || c == b'f' || lexer::is_whitespace(c) => i += 1, + b'%' => { + i += b + .get(i..)? + .iter() + .take_while(|c| !matches!(c, b'\r' | b'\n')) + .count(); + } + _ => break, + } + } + b.get(i..)?.starts_with(b"trailer").then_some(i) +} + +/// `N G obj << /Type /XRef … >> stream … endstream`: bounds the decoded size, `/W` and `/Index`. +fn check_xref_stream(b: &[u8], p: usize) -> Result, PVal)>, AppError> { + let bad = |what: &str| reasons::invalid_xref(&format!("{what} at byte {p}")); + let (_, i) = parse_uint(b, p).ok_or_else(|| bad("no xref"))?; + let (_, i) = parse_uint(b, skip_ws(b, i)).ok_or_else(|| bad("no xref"))?; + let i = skip_ws(b, i); + if b.get(i..i + 3) != Some(&b"obj"[..]) { + return Err(bad("no xref")); + } + let (dict, end) = scan_dict(b, i + 3)?; + if !matches!(dict_get(&dict, b"Type"), Some(PVal::Name(n)) if n == b"XRef") { + return Err(bad("not an xref stream")); + } + let s = skip_ws(b, end); + if b.get(s..s + 6) != Some(&b"stream"[..]) { + return Err(bad("xref stream without data")); + } + let mut data_start = s + 6; + if b.get(data_start) == Some(&b'\r') { + data_start += 1; + } + if b.get(data_start) == Some(&b'\n') { + data_start += 1; + } + let by_length = dict_uint(&dict, b"Length") + .and_then(|l| data_start.checked_add(l)) + .filter(|e| { + b.get(skip_ws(b, *e)..) + .is_some_and(|rest| rest.starts_with(b"endstream")) + }); + let data_end = match by_length { + Some(e) => e, + None => { + find_from(b, data_start, b"endstream").ok_or_else(|| bad("unterminated xref stream"))? + } + }; + let payload = b.get(data_start..data_end).unwrap_or_default(); + decode_xref_payload(&dict, payload)?; + check_xref_fields(&dict, p)?; + Ok(dict) +} + +fn decode_xref_payload(dict: &[(Vec, PVal)], payload: &[u8]) -> Result<(), AppError> { + let names: Vec<&[u8]> = match dict_get(dict, b"Filter") { + None => Vec::new(), + Some(PVal::Name(n)) => vec![n.as_slice()], + Some(PVal::Array(items)) if items.len() <= 4 => items + .iter() + .map(|i| match i { + PVal::Name(n) => Ok(n.as_slice()), + _ => Err(reasons::file_too_complex("xref stream filter")), + }) + .collect::>()?, + Some(_) => return Err(reasons::file_too_complex("xref stream filter")), + }; + let cap = limits::XREF_STREAM_MAX_DECODED; + let mut data = std::borrow::Cow::Borrowed(payload); + for name in names { + let out = match name { + b"FlateDecode" | b"Fl" => decode::inflate_capped(&data, cap), + b"ASCIIHexDecode" | b"AHx" => decode::ascii_hex_decode(&data, cap), + b"ASCII85Decode" | b"A85" => decode::ascii85_decode(&data, cap), + _ => return Err(reasons::file_too_complex("xref stream filter")), + }; + data = std::borrow::Cow::Owned(out.map_err(|e| match e { + DecodeError::TooLarge => reasons::file_too_complex("xref stream too large"), + other => reasons::invalid_xref(&format!("xref stream data: {other}")), + })?); + } + if data.len() > cap { + return Err(reasons::file_too_complex("xref stream too large")); + } + Ok(()) +} + +/// `/W`: three integers 0..=8 with at least one byte per row; `/Index` (or `[0 /Size]`): +/// non-negative pairs whose counts sum (checked) to at most MAX_OBJECTS (lopdf loops over every +/// count); a PNG predictor row of at most `XREF_PREDICTOR_ROW_MAX` bytes (lopdf allocates it). +/// Malformed fields → `INVALID_PDF`; over a budget → `FILE_TOO_COMPLEX`. Array items must be +/// written as integers: lopdf rejects a `/W` with a real and replaces such an `/Index` by +/// `[0 /Size]`, which this check would not have bounded. +fn check_xref_fields(dict: &[(Vec, PVal)], p: usize) -> Result<(), AppError> { + let bad = |what: &str| reasons::invalid_xref(&format!("{what} at byte {p}")); + let too_large = |what: &str| reasons::file_too_complex(&format!("{what} at byte {p}")); + let ints = |v: Option<&PVal>| -> Option> { + match v? { + PVal::Array(items) => items + .iter() + .map(|i| match i { + PVal::Num(n) if n.fract() == 0.0 && n.abs() < 1e15 => Some(*n as i64), + _ => None, + }) + .collect(), + _ => None, + } + }; + let w = ints(dict_get(dict, b"W")).ok_or_else(|| bad("xref stream /W"))?; + if w.len() != 3 + || w.iter() + .any(|x| !(0..=limits::XREF_STREAM_FIELD_WIDTH_MAX).contains(x)) + { + return Err(bad("xref stream /W")); + } + let row: i64 = w.iter().sum(); + if row == 0 { + return Err(bad("xref stream /W")); + } + let index = match dict_get(dict, b"Index") { + Some(v) => ints(Some(v)).ok_or_else(|| bad("xref stream /Index"))?, + None => vec![ + 0, + dict_uint(dict, b"Size").ok_or_else(|| bad("xref stream /Size"))? as i64, + ], + }; + if index.len() % 2 != 0 || index.iter().any(|x| *x < 0) { + return Err(bad("xref stream /Index")); + } + let count = index + .iter() + .skip(1) + .step_by(2) + .try_fold(0i64, |sum, n| sum.checked_add(*n)); + if !count.is_some_and(|c| c <= limits::MAX_OBJECTS as i64) { + return Err(too_large("xref stream lists too many objects")); + } + if !predictor_row_ok(dict) { + return Err(too_large("xref stream predictor row too large")); + } + Ok(()) +} + +/// lopdf 0.34 (`object.rs` `decompress_predictor`): with a direct `/DecodeParms` dictionary whose +/// `/Predictor` is 10..=15 it allocates a row of `max(1, Columns) × max(1, Colors) × +/// max(8, BitsPerComponent) / 8` bytes (and multiplies them unchecked). Non-integer values fall +/// back to lopdf's defaults there; here every number counts, so the bound is never lower. +fn predictor_row_ok(dict: &[(Vec, PVal)]) -> bool { + let Some(PVal::Dict(parms)) = dict_get(dict, b"DecodeParms") else { + return true; + }; + let num = |key: &[u8], default: f64| match dict_get(parms, key) { + Some(PVal::Num(n) | PVal::Real(n)) => *n, + _ => default, + }; + if !(10.0..=15.0).contains(&num(b"Predictor", 1.0)) { + return true; + } + let row = num(b"Columns", 1.0).max(1.0) + * num(b"Colors", 1.0).max(1.0) + * num(b"BitsPerComponent", 8.0).max(8.0) + / 8.0; + row.is_finite() && row <= limits::XREF_PREDICTOR_ROW_MAX as f64 +} diff --git a/src-tauri/src/pdf_engine/text_edit/state.rs b/src-tauri/src/pdf_engine/text_edit/state.rs new file mode 100644 index 0000000..07ae48c --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/state.rs @@ -0,0 +1,664 @@ +//! Graphics state, its id-free digest and the comparisons joins and verification use (SPEC §B.10, +//! D32, D35). Identities are resource names plus content hashes, never lopdf object ids, so a +//! digest taken on qpdf's output (which renumbers every object) compares with the source's. +//! +//! Every variable-size part of the state (verbatim op bytes, names, the dash array, unmodelled +//! ExtGState keys) is shared through an `Arc`, and the marked-content stack is a persistent list +//! (`marked::MarkedStack`): a digest is taken for every show op and paint, so cloning one must +//! cost the same whatever the content holds. Whole digests are shared too: +//! `StateDigest::is_digest_of` tells in O(1) whether the last digest still describes the state, +//! so a run of records under one state holds a single digest. + +mod marked; + +pub use marked::{MarkedNode, MarkedStack}; + +use crate::pdf_engine::text_edit::fonts::{FontKey, FontModel}; +use crate::pdf_engine::text_edit::geometry::{Matrix, IDENTITY}; +use crate::pdf_engine::text_edit::limits::{COLOR_EPSILON, STATE_EPSILON}; +use std::collections::HashSet; +use std::sync::Arc; + +/// Shared, immutable bytes (verbatim op spans, names). Cloning never copies them. +pub type Bytes = Arc<[u8]>; + +/// What a paint looks like, for display only ("Original colour", `fill_hex`). +#[derive(Debug, Clone, PartialEq)] +pub enum ColorEffect { + Rgb([f64; 3]), + Unreadable, +} + +#[derive(Debug, Clone, PartialEq)] +pub enum ColorSpaceKind { + /// Never set in this page (the initial DeviceGray black). + Default, + DeviceGray, + DeviceRgb, + DeviceCmyk, + /// A `/ColorSpace` resource: its name and the deep hash of what it names. + Named(Bytes, u64), + Pattern, +} + +/// A fill or stroke paint. `space_op`/`color_op` are the verbatim bytes of the ops that set it +/// (the whole op span), used for verbatim restores (B14). `pattern_hash` is the deep hash of the +/// `/Pattern` resource an `scn`/`SCN` names (its name is in `color_op`), so a pattern swapped +/// behind the same name compares unequal. +#[derive(Debug, Clone, PartialEq)] +pub struct Paint { + pub space_op: Option, + pub color_op: Option, + pub space: ColorSpaceKind, + pub comps: Vec, + pub effect: ColorEffect, + pub pattern: bool, + pub pattern_hash: Option, +} + +impl Paint { + /// The initial paint: DeviceGray black, never set. + pub fn initial() -> Paint { + Paint { + space_op: None, + color_op: None, + space: ColorSpaceKind::Default, + comps: vec![0.0], + effect: ColorEffect::Rgb([0.0; 3]), + pattern: false, + pattern_hash: None, + } + } + + /// `#rrggbb` when the effect is readable. + pub fn hex(&self) -> Option { + let ColorEffect::Rgb(rgb) = &self.effect else { + return None; + }; + let byte = |v: f64| (v.clamp(0.0, 1.0) * 255.0).round() as u8; + Some(format!( + "#{:02x}{:02x}{:02x}", + byte(rgb[0]), + byte(rgb[1]), + byte(rgb[2]) + )) + } +} + +/// The RGB effect of components in a device space; `Unreadable` for anything else. +pub fn color_effect(space: &ColorSpaceKind, comps: &[f64]) -> ColorEffect { + let ok = comps.iter().all(|c| c.is_finite()); + match (space, comps) { + (ColorSpaceKind::Default | ColorSpaceKind::DeviceGray, [g]) if ok => { + ColorEffect::Rgb([*g, *g, *g]) + } + (ColorSpaceKind::DeviceRgb, [r, g, b]) if ok => ColorEffect::Rgb([*r, *g, *b]), + (ColorSpaceKind::DeviceCmyk, [c, m, y, k]) if ok => { + let inv = |v: f64| (1.0 - v.clamp(0.0, 1.0)) * (1.0 - k.clamp(0.0, 1.0)); + ColorEffect::Rgb([inv(*c), inv(*m), inv(*y)]) + } + _ => ColorEffect::Unreadable, + } +} + +/// The clip in force: none, an intersection of axis-aligned rectangles (`[x0, y0, x1, y1]`, user +/// space), or anything else. +#[derive(Debug, Clone, PartialEq)] +pub enum ClipState { + None, + Rect([f64; 4]), + Complex, +} + +/// The font in force. `resource` + `content_hash` are the identity; the lopdf `FontKey` is kept +/// outside the digest (on `GState::font_ref` and the record) because qpdf renumbers objects. +#[derive(Debug, Clone, PartialEq)] +pub struct FontUse { + pub resource: Option, + pub content_hash: u64, + pub from_extgstate: bool, + pub tf_op: Option, // verbatim "/F1 12 Tf" op bytes +} + +#[derive(Debug, Clone, PartialEq)] +pub struct TextParams { + pub font: Option, + pub tfs: f64, + pub tc: f64, + pub tc_src: Option, + pub tw: f64, + pub tw_src: Option, + pub th: f64, + pub tl: f64, + pub tr: i64, + pub ts: f64, +} + +impl Default for TextParams { + fn default() -> Self { + TextParams { + font: None, + tfs: 0.0, + tc: 0.0, + tc_src: None, + tw: 0.0, + tw_src: None, + th: 1.0, + tl: 0.0, + tr: 0, + ts: 0.0, + } + } +} + +#[derive(Debug, Clone, PartialEq)] +pub struct GsEffects { + pub ca: f64, + pub ca_stroke: f64, + pub blend: Bytes, + pub soft_mask: bool, + pub overprint: (bool, bool, i64), + pub line_width: f64, + pub line_cap: i64, + pub line_join: i64, + pub miter_limit: f64, + pub dash: (Arc<[f64]>, f64), + pub rendering_intent: Bytes, + pub flatness: f64, + pub stroke_adjust: bool, + /// ExtGState keys outside the modelled list that affect rendering (`/TR`, `/HT`, `/BG`, …), + /// with the canonical hash of their value; sorted by key, one entry per key, last one wins. + pub other: Arc<[(Bytes, u64)]>, +} + +impl Default for GsEffects { + fn default() -> Self { + GsEffects { + ca: 1.0, + ca_stroke: 1.0, + blend: Arc::from(&b"Normal"[..]), + soft_mask: false, + overprint: (false, false, 0), + line_width: 1.0, + line_cap: 0, + line_join: 0, + miter_limit: 10.0, + dash: (Arc::from(Vec::new()), 0.0), + rendering_intent: Arc::from(&b"RelativeColorimetric"[..]), + flatness: 1.0, + stroke_adjust: false, + other: Arc::from(Vec::new()), + } + } +} + +impl GsEffects { + /// The unmodelled keys of one ExtGState as `set_others` takes them: sorted by key, one entry + /// per key. + pub fn sorted_others(mut entries: Vec<(Bytes, u64)>) -> Arc<[(Bytes, u64)]> { + entries.sort_by(|a, b| a.0.cmp(&b.0)); + entries.dedup_by(|later, earlier| later.0 == earlier.0); + entries.into() + } + + /// Records the unmodelled keys of one ExtGState (`sorted_others`; each replaces an earlier + /// value of the same key) with one merge, O(k + m). When every key is already in force with + /// the same hash nothing changes and the shared list in force is kept, so re-applying one + /// ExtGState never allocates a new list (nor, through `Exec::digest`, a new digest). A merged + /// list equal to one in `interned` (the walk's lists so far) is that list, so a page that + /// alternates between ExtGStates holds one list per distinct state, not one per `gs`. True + /// when a new list was interned (the caller's budget keeps it). + pub fn set_others( + &mut self, + entries: &[(Bytes, u64)], + interned: &mut HashSet>, + ) -> bool { + let unchanged = entries.iter().all(|(key, hash)| { + self.other + .binary_search_by(|(k, _)| k.cmp(key)) + .ok() + .and_then(|i| self.other.get(i)) + .is_some_and(|(_, h)| h == hash) + }); + if unchanged { + return false; + } + let mut merged = Vec::with_capacity(self.other.len().saturating_add(entries.len())); + let mut old = self.other.iter().peekable(); + let mut new = entries.iter().peekable(); + loop { + let take_old = match (old.peek(), new.peek()) { + (Some(o), Some(n)) => match o.0.cmp(&n.0) { + std::cmp::Ordering::Less => true, + std::cmp::Ordering::Equal => { + old.next(); + false + } + std::cmp::Ordering::Greater => false, + }, + (Some(_), None) => true, + (None, Some(_)) => false, + (None, None) => break, + }; + let next = if take_old { old.next() } else { new.next() }; + merged.extend(next.cloned()); + } + if let Some(known) = interned.get(merged.as_slice()) { + self.other = Arc::clone(known); + return false; + } + let list: Arc<[(Bytes, u64)]> = merged.into(); + interned.insert(Arc::clone(&list)); + self.other = list; + true + } +} + +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct MarkedEntry { + pub tag: Bytes, + pub mcid: Option, + pub actual_text: bool, + pub oc: Option, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum OcState { + Visible, + Hidden, + Unknown, +} + +/// The graphics state (saved by `q`, restored by `Q`). `font_ref` is the lopdf-side handle of the +/// font in force (key + loaded model) and is never part of a digest. +#[derive(Debug, Clone)] +pub struct GState { + pub ctm: Matrix, + pub clip: ClipState, + pub fill: Paint, + pub stroke: Paint, + pub gs: GsEffects, + pub text: TextParams, + pub font_ref: Option<(FontKey, Arc)>, +} + +impl GState { + pub fn initial(ctm: Matrix) -> GState { + GState { + ctm, + clip: ClipState::None, + fill: Paint::initial(), + stroke: Paint::initial(), + gs: GsEffects::default(), + text: TextParams::default(), + font_ref: None, + } + } + + /// O(1) in the size of the content: every variable-size part is shared. + pub fn digest(&self, marked: &MarkedStack) -> StateDigest { + StateDigest { + ctm: self.ctm, + clip: self.clip.clone(), + fill: self.fill.clone(), + stroke: self.stroke.clone(), + gs: self.gs.clone(), + text: self.text.clone(), + marked: marked.clone(), + } + } +} + +impl Default for GState { + fn default() -> Self { + GState::initial(IDENTITY) + } +} + +#[derive(Debug, Clone, PartialEq)] +pub struct StateDigest { + pub ctm: Matrix, + pub clip: ClipState, + pub fill: Paint, + pub stroke: Paint, + pub gs: GsEffects, + pub text: TextParams, + /// The marked-content stack (`iter` gives the innermost entry first). + pub marked: MarkedStack, +} + +impl StateDigest { + /// Whether this digest is exactly what `gs.digest(marked)` would return, judged in O(1): + /// every shared part must be the very same allocation and every number equal bit for bit. A + /// state re-set to an equal value under a new allocation only loses sharing, never + /// correctness. Lets consecutive records and paints share one digest (`Exec::digest`). + pub fn is_digest_of(&self, gs: &GState, marked: &MarkedStack) -> bool { + let StateDigest { + ctm, + clip, + fill, + stroke, + gs: effects, + text, + marked: stack, + } = self; + same_bits(ctm, &gs.ctm) + && clip.identical(&gs.clip) + && fill.identical(&gs.fill) + && stroke.identical(&gs.stroke) + && effects.identical(&gs.gs) + && text.identical(&gs.text) + && stack.same(marked) + } +} + +/// Equality that never walks a shared part (`StateDigest::is_digest_of`). The destructuring +/// patterns are exhaustive, so a new state field cannot be left out. +trait Identical { + fn identical(&self, other: &Self) -> bool; +} + +fn same_bits(a: &[f64], b: &[f64]) -> bool { + a.len() == b.len() && a.iter().zip(b).all(|(x, y)| x.to_bits() == y.to_bits()) +} + +fn same_f(a: f64, b: f64) -> bool { + a.to_bits() == b.to_bits() +} + +fn same_alloc(a: &Option>, b: &Option>) -> bool { + match (a, b) { + (None, None) => true, + (Some(x), Some(y)) => Arc::ptr_eq(x, y), + _ => false, + } +} + +impl Identical for ClipState { + fn identical(&self, other: &Self) -> bool { + match (self, other) { + (ClipState::None, ClipState::None) | (ClipState::Complex, ClipState::Complex) => true, + (ClipState::Rect(a), ClipState::Rect(b)) => same_bits(a, b), + _ => false, + } + } +} + +impl Identical for ColorSpaceKind { + fn identical(&self, other: &Self) -> bool { + use ColorSpaceKind as K; + match (self, other) { + (K::Named(a, h), K::Named(b, k)) => Arc::ptr_eq(a, b) && h == k, + (K::Default, K::Default) + | (K::DeviceGray, K::DeviceGray) + | (K::DeviceRgb, K::DeviceRgb) + | (K::DeviceCmyk, K::DeviceCmyk) + | (K::Pattern, K::Pattern) => true, + _ => false, + } + } +} + +impl Identical for Paint { + fn identical(&self, o: &Self) -> bool { + let Paint { + space_op, + color_op, + space, + comps, + effect, + pattern, + pattern_hash, + } = self; + let same_effect = match (effect, &o.effect) { + (ColorEffect::Rgb(a), ColorEffect::Rgb(b)) => same_bits(a, b), + (ColorEffect::Unreadable, ColorEffect::Unreadable) => true, + _ => false, + }; + same_alloc(space_op, &o.space_op) + && same_alloc(color_op, &o.color_op) + && space.identical(&o.space) + && same_bits(comps, &o.comps) + && same_effect + && *pattern == o.pattern + && *pattern_hash == o.pattern_hash + } +} + +impl Identical for GsEffects { + fn identical(&self, o: &Self) -> bool { + let GsEffects { + ca, + ca_stroke, + blend, + soft_mask, + overprint, + line_width, + line_cap, + line_join, + miter_limit, + dash, + rendering_intent, + flatness, + stroke_adjust, + other, + } = self; + same_f(*ca, o.ca) + && same_f(*ca_stroke, o.ca_stroke) + && Arc::ptr_eq(blend, &o.blend) + && *soft_mask == o.soft_mask + && *overprint == o.overprint + && same_f(*line_width, o.line_width) + && *line_cap == o.line_cap + && *line_join == o.line_join + && same_f(*miter_limit, o.miter_limit) + && Arc::ptr_eq(&dash.0, &o.dash.0) + && same_f(dash.1, o.dash.1) + && Arc::ptr_eq(rendering_intent, &o.rendering_intent) + && same_f(*flatness, o.flatness) + && *stroke_adjust == o.stroke_adjust + && Arc::ptr_eq(other, &o.other) + } +} + +impl Identical for TextParams { + fn identical(&self, o: &Self) -> bool { + let TextParams { + font, + tfs, + tc, + tc_src, + tw, + tw_src, + th, + tl, + tr, + ts, + } = self; + let same_font = match (font, &o.font) { + (None, None) => true, + (Some(a), Some(b)) => { + let FontUse { + resource, + content_hash, + from_extgstate, + tf_op, + } = a; + same_alloc(resource, &b.resource) + && *content_hash == b.content_hash + && *from_extgstate == b.from_extgstate + && same_alloc(tf_op, &b.tf_op) + } + _ => false, + }; + same_font + && same_f(*tfs, o.tfs) + && same_f(*tc, o.tc) + && same_alloc(tc_src, &o.tc_src) + && same_f(*tw, o.tw) + && same_alloc(tw_src, &o.tw_src) + && same_f(*th, o.th) + && same_f(*tl, o.tl) + && *tr == o.tr + && same_f(*ts, o.ts) + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum StateField { + Font, + Tfs, + Tc, + Fill, +} + +fn close(a: f64, b: f64, eps: f64) -> bool { + a == b || (a - b).abs() <= eps * 1f64.max(a.abs()).max(b.abs()) +} + +fn comps_close(a: &[f64], b: &[f64]) -> bool { + a.len() == b.len() && a.iter().zip(b).all(|(x, y)| (x - y).abs() <= COLOR_EPSILON) +} + +fn initial_black(p: &Paint) -> bool { + match (&p.space, p.comps.as_slice()) { + (ColorSpaceKind::DeviceGray, [g]) => g.abs() <= COLOR_EPSILON, + (ColorSpaceKind::DeviceRgb, [r, g, b]) => { + [r, g, b].iter().all(|v| v.abs() <= COLOR_EPSILON) + } + (ColorSpaceKind::DeviceCmyk, [c, m, y, k]) => { + [c, m, y].iter().all(|v| v.abs() <= COLOR_EPSILON) && (k - 1.0).abs() <= COLOR_EPSILON + } + _ => false, + } +} + +/// Equal iff (a) the verbatim op bytes and the colour space are equal, or (b) same +/// `ColorSpaceKind` (Named: same name and hash) and components within COLOR_EPSILON, or (c) one +/// side is `Default` (never set) and the other is DeviceGray 0, DeviceRGB 0 0 0 or DeviceCMYK +/// 0 0 0 1. Pattern paints are equal only with identical colour ops (the pattern name) and the +/// same deep hash of the pattern they name. Other cross-family values are unequal (D35). +pub fn same_paint(a: &Paint, b: &Paint) -> bool { + if a.pattern || b.pattern { + return a.pattern == b.pattern + && a.space == b.space + && a.color_op == b.color_op + && a.pattern_hash == b.pattern_hash; + } + let verbatim = (a.space_op.is_some() || a.color_op.is_some()) + && a.space_op == b.space_op + && a.color_op == b.color_op; + if a.space == b.space && (verbatim || comps_close(&a.comps, &b.comps)) { + return true; + } + match (&a.space, &b.space) { + (ColorSpaceKind::Default, _) => comps_close(&a.comps, &[0.0]) && initial_black(b), + (_, ColorSpaceKind::Default) => comps_close(&b.comps, &[0.0]) && initial_black(a), + _ => false, + } +} + +fn same_font(a: &Option, b: &Option) -> bool { + match (a, b) { + (None, None) => true, + (Some(x), Some(y)) => { + x.resource == y.resource + && x.content_hash == y.content_hash + && x.from_extgstate == y.from_extgstate + } + _ => false, + } +} + +fn same_clip(a: &ClipState, b: &ClipState) -> bool { + match (a, b) { + (ClipState::None, ClipState::None) | (ClipState::Complex, ClipState::Complex) => true, + (ClipState::Rect(x), ClipState::Rect(y)) => { + x.iter().zip(y).all(|(p, q)| close(*p, *q, STATE_EPSILON)) + } + _ => false, + } +} + +/// Dash arrays: the same shared array (the common case) without walking it. +fn same_dash(a: &(Arc<[f64]>, f64), b: &(Arc<[f64]>, f64)) -> bool { + a.1 == b.1 && (Arc::ptr_eq(&a.0, &b.0) || a.0 == b.0) +} + +fn same_gs(a: &GsEffects, b: &GsEffects) -> Result<(), &'static str> { + let checks: [(bool, &'static str); 14] = [ + (a.ca == b.ca, "ca"), + (a.ca_stroke == b.ca_stroke, "CA"), + (a.blend == b.blend, "blend"), + (a.soft_mask == b.soft_mask, "soft_mask"), + (a.overprint == b.overprint, "overprint"), + (a.line_width == b.line_width, "line_width"), + (a.line_cap == b.line_cap, "line_cap"), + (a.line_join == b.line_join, "line_join"), + (a.miter_limit == b.miter_limit, "miter_limit"), + (same_dash(&a.dash, &b.dash), "dash"), + (a.rendering_intent == b.rendering_intent, "rendering_intent"), + (a.flatness == b.flatness, "flatness"), + (a.stroke_adjust == b.stroke_adjust, "stroke_adjust"), + (a.other == b.other, "extgstate"), + ]; + match checks.iter().find(|(ok, _)| !ok) { + Some((_, field)) => Err(field), + None => Ok(()), + } +} + +/// Field name on mismatch. Compares every field of the digest: CTM within STATE_EPSILON, clip, +/// fill and stroke via `same_paint`, all `GsEffects`, text params by value (sources ignored), +/// marked stack. +pub fn same_state(a: &StateDigest, b: &StateDigest) -> Result<(), &'static str> { + same_state_except(a, b, &[]) +} + +/// As `same_state` but skipping the listed fields (joins skip the font identity; edited glyphs +/// skip only the fields their style change targets). +pub fn same_state_except( + a: &StateDigest, + b: &StateDigest, + skip: &[StateField], +) -> Result<(), &'static str> { + let skipped = |f: StateField| skip.contains(&f); + if !a + .ctm + .iter() + .zip(&b.ctm) + .all(|(p, q)| close(*p, *q, STATE_EPSILON)) + { + return Err("ctm"); + } + if !same_clip(&a.clip, &b.clip) { + return Err("clip"); + } + if !skipped(StateField::Fill) && !same_paint(&a.fill, &b.fill) { + return Err("fill"); + } + if !same_paint(&a.stroke, &b.stroke) { + return Err("stroke"); + } + same_gs(&a.gs, &b.gs)?; + let (x, y) = (&a.text, &b.text); + if !skipped(StateField::Font) && !same_font(&x.font, &y.font) { + return Err("font"); + } + let text: [(bool, &'static str); 7] = [ + (skipped(StateField::Tfs) || x.tfs == y.tfs, "tfs"), + (skipped(StateField::Tc) || x.tc == y.tc, "tc"), + (x.tw == y.tw, "tw"), + (x.th == y.th, "th"), + (x.tl == y.tl, "tl"), + (x.tr == y.tr, "tr"), + (x.ts == y.ts, "ts"), + ]; + if let Some((_, field)) = text.iter().find(|(ok, _)| !ok) { + return Err(field); + } + if a.marked != b.marked { + return Err("marked"); + } + Ok(()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/state/marked.rs b/src-tauri/src/pdf_engine/text_edit/state/marked.rs new file mode 100644 index 0000000..b195b9d --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/state/marked.rs @@ -0,0 +1,103 @@ +//! The marked-content stack (`BMC`/`BDC` … `EMC`) as a persistent list. A push or a pop makes a +//! new stack that shares every entry below with the old one, so the stack each digest holds +//! costs O(1) per op whatever its depth: a page that reopens a deep stack with a new tag around +//! every show op keeps one node per op, not a copy of the whole stack (review T3 r2 MEDIUM-2). +//! Chains are at most `MARKED_DEPTH_MAX` long (the walker refuses a deeper push), so dropping one +//! recurses at most that deep. + +use super::MarkedEntry; +use std::sync::Arc; + +/// Innermost entry first. Cloning shares the whole stack. +#[derive(Debug, Clone, Default)] +pub struct MarkedStack(Option>); + +/// One entry and the stack below it. +#[derive(Debug)] +pub struct MarkedNode { + entry: MarkedEntry, + below: MarkedStack, + len: usize, +} + +impl MarkedStack { + pub fn len(&self) -> usize { + self.0.as_ref().map_or(0, |n| n.len) + } + + pub fn push(&mut self, entry: MarkedEntry) { + let below = std::mem::take(self); + let len = below.len().saturating_add(1); + *self = MarkedStack(Some(Arc::new(MarkedNode { entry, below, len }))); + } + + /// Removes the innermost entry (nothing when empty). + pub fn pop(&mut self) { + if let Some(top) = self.0.take() { + *self = top.below.clone(); + } + } + + /// The entries, innermost first. + pub fn iter(&self) -> MarkedIter<'_> { + MarkedIter(self.0.as_deref()) + } + + /// The very same stack (no entry compared): O(1). + pub fn same(&self, other: &MarkedStack) -> bool { + match (&self.0, &other.0) { + (None, None) => true, + (Some(a), Some(b)) => Arc::ptr_eq(a, b), + _ => false, + } + } + + /// Each node's address with its entry, innermost first (`approx_bytes` counts each node + /// once). + pub(crate) fn nodes(&self) -> impl Iterator { + let mut cur = self.0.as_ref(); + std::iter::from_fn(move || { + let node = cur?; + cur = node.below.0.as_ref(); + Some((Arc::as_ptr(node) as usize, &node.entry)) + }) + } +} + +/// Equal entries in the same order; stops at the first shared node. +impl PartialEq for MarkedStack { + fn eq(&self, other: &Self) -> bool { + if self.len() != other.len() { + return false; + } + let (mut a, mut b) = (self.0.as_ref(), other.0.as_ref()); + loop { + match (a, b) { + (None, None) => return true, + (Some(x), Some(y)) => { + if Arc::ptr_eq(x, y) { + return true; + } + if x.entry != y.entry { + return false; + } + a = x.below.0.as_ref(); + b = y.below.0.as_ref(); + } + _ => return false, + } + } + } +} + +pub struct MarkedIter<'a>(Option<&'a MarkedNode>); + +impl<'a> Iterator for MarkedIter<'a> { + type Item = &'a MarkedEntry; + + fn next(&mut self) -> Option<&'a MarkedEntry> { + let node = self.0?; + self.0 = node.below.0.as_deref(); + Some(&node.entry) + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/structure.rs b/src-tauri/src/pdf_engine/text_edit/structure.rs new file mode 100644 index 0000000..7e7e4f5 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/structure.rs @@ -0,0 +1,318 @@ +//! Optional content and the logical structure tree (SPEC §A.6, §B.10, D19, D33): whether an +//! `/OC` group is visible in the default configuration, whether a marked-content id's structure +//! element (or one of its ancestors) carries `/ActualText`, and the structure order of a page's +//! MCIDs for reading order. Every traversal is bounded and cycle-checked. + +use crate::pdf_engine::text_edit::context::{get, resolve}; +use crate::pdf_engine::text_edit::limits::{ + NUMBER_TREE_NODES_MAX, STRUCT_CHAIN_MAX, STRUCT_ORDER_NODES_MAX, +}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::state::OcState; +use lopdf::{Dictionary, Document, Object, ObjectId}; +use std::collections::HashSet; + +/// Nesting depth of the structure-order DFS. +const STRUCT_ORDER_DEPTH_MAX: usize = 64; + +/// The default optional-content configuration (`/OCProperties /OCGs` and `/D`), read once per +/// snapshot (`SnapshotContext::oc_config`) into sorted id lists, so the state of a group costs +/// O(log n) however many groups the catalog lists or a page opens. +pub struct OcConfig { + /// `None` when there is no readable configuration (every group is then `Unknown`). + lists: Option, +} + +struct OcLists { + ocgs: Vec, + on: Vec, + off: Vec, + base_on: bool, +} + +impl OcConfig { + pub fn of(doc: &Document) -> OcConfig { + OcConfig { + lists: oc_lists(doc), + } + } + + /// The state of the optional-content group a `BDC /OC /Name` names, given the `/Properties` + /// entry (normally a reference to the group). A single OCG listed in `/OCProperties /OCGs` + /// is `Visible` when the default configuration `/D` has it on (BaseState ON and not in + /// `/OFF`, or listed in `/ON`), `Hidden` when off; an OCMD, an inline dictionary, a group the + /// catalog does not list, or anything unreadable is `Unknown` (all refused + /// `OPTIONAL_CONTENT` but `Visible`). + pub fn state(&self, doc: &Document, props: &Object) -> OcState { + let Object::Reference(id) = props else { + return OcState::Unknown; + }; + let Some((_, Object::Dictionary(group))) = resolve(doc, props) else { + return OcState::Unknown; + }; + if !group.type_is(b"OCG") { + return OcState::Unknown; // OCMD or not a group + } + let Some(lists) = &self.lists else { + return OcState::Unknown; + }; + let has = |ids: &[ObjectId]| ids.binary_search(id).is_ok(); + if !has(&lists.ocgs) { + return OcState::Unknown; + } + let (on, off) = (has(&lists.on), has(&lists.off)); + if on && off { + return OcState::Unknown; + } + if (lists.base_on && !off) || on { + OcState::Visible + } else { + OcState::Hidden + } + } +} + +/// The catalog's `/OCGs`, `/D /ON`, `/D /OFF` (direct references, sorted) and `/D /BaseState`; +/// `None` without `/OCProperties`, without a `/D` dictionary or with a BaseState other than +/// `/ON` or `/OFF` (`/Unchanged` has no meaning in `/D`). +fn oc_lists(doc: &Document) -> Option { + let ocp = doc + .catalog() + .ok() + .and_then(|c| get(doc, c, b"OCProperties")) + .and_then(|o| o.as_dict().ok())?; + let ids = |arr: Option<&Object>| -> Vec { + let mut ids: Vec = arr + .and_then(|a| a.as_array().ok()) + .map(|items| items.iter().filter_map(|i| i.as_reference().ok()).collect()) + .unwrap_or_default(); + ids.sort_unstable(); + ids + }; + let d = get(doc, ocp, b"D").and_then(|o| o.as_dict().ok())?; + let base_on = match get(doc, d, b"BaseState") { + None => true, + Some(Object::Name(n)) if n.as_slice() == b"ON" => true, + Some(Object::Name(n)) if n.as_slice() == b"OFF" => false, + _ => return None, + }; + Some(OcLists { + ocgs: ids(get(doc, ocp, b"OCGs")), + on: ids(get(doc, d, b"ON")), + off: ids(get(doc, d, b"OFF")), + base_on, + }) +} + +/// The page's slice of the structure parent tree (`/StructParents` → `/ParentTree`), resolved +/// once per page and borrowed from the document (a parent array of millions of entries costs +/// nothing per model build, review T3 r3 LOW-2). +pub struct PageStruct<'a> { + /// `Ok(None)`: the page has no structure parents (no MCID maps to an element). + parents: Result, TextReason>, +} + +impl<'a> PageStruct<'a> { + pub fn of(doc: &'a Document, page_id: ObjectId) -> PageStruct<'a> { + PageStruct { + parents: page_parents(doc, page_id), + } + } + + /// Whether the structure element of `mcid` or one of its ancestors (≤ STRUCT_CHAIN_MAX) has + /// `/ActualText`. A broken tree is `Err(ACTUAL_TEXT)` (the text might carry one: fail closed). + pub fn actual_text(&self, doc: &Document, mcid: i64) -> Result { + let parents = match &self.parents { + Ok(None) => return Ok(false), + Ok(Some(p)) => p, + Err(r) => return Err(*r), + }; + let Some(elem) = usize::try_from(mcid).ok().and_then(|i| parents.get(i)) else { + return Ok(false); + }; + let mut cur = match resolve(doc, elem) { + Some((_, Object::Dictionary(d))) => d, + Some((_, Object::Null)) => return Ok(false), + _ => return Err(TextReason::ActualText), + }; + for _ in 0..STRUCT_CHAIN_MAX { + if cur.has(b"ActualText") { + return Ok(true); + } + if cur.type_is(b"StructTreeRoot") { + return Ok(false); + } + match cur.get(b"P").ok().map(|p| resolve(doc, p)) { + None => return Ok(false), + Some(Some((_, Object::Dictionary(d)))) => cur = d, + Some(_) => return Err(TextReason::ActualText), + } + } + // A chain longer than the bound: treat as unknown. + Err(TextReason::ActualText) + } +} + +/// `/StructParents` → `/ParentTree` (number tree, bounded) → the page's array. +fn page_parents(doc: &Document, page_id: ObjectId) -> Result, TextReason> { + let bad = TextReason::ActualText; + let Some(Object::Dictionary(page)) = doc.objects.get(&page_id) else { + return Ok(None); + }; + let key = match get(doc, page, b"StructParents") { + None => return Ok(None), + Some(Object::Integer(k)) => *k, + Some(_) => return Err(bad), + }; + let Some(root) = doc + .catalog() + .ok() + .and_then(|c| get(doc, c, b"StructTreeRoot")) + else { + return Ok(None); + }; + let root = root.as_dict().map_err(|_| bad)?; + let Some(tree) = get(doc, root, b"ParentTree") else { + return Ok(None); + }; + match number_tree_get(doc, tree, key)? { + None => Ok(None), + Some(Object::Array(items)) => Ok(Some(items.as_slice())), + Some(_) => Ok(None), // an object reference entry (OBJR parent), not marked content + } +} + +/// Looks `key` up in a number tree (bounded DFS, visited set; `/Limits` are not trusted). +fn number_tree_get<'a>( + doc: &'a Document, + tree: &'a Object, + key: i64, +) -> Result, TextReason> { + let bad = TextReason::ActualText; + let mut stack = vec![tree]; + let mut seen: HashSet = HashSet::new(); + let mut nodes = 0usize; + while let Some(raw) = stack.pop() { + nodes += 1; + if nodes > NUMBER_TREE_NODES_MAX { + return Err(bad); + } + if let Object::Reference(id) = raw { + if !seen.insert(*id) { + return Err(bad); + } + } + let node = match resolve(doc, raw) { + Some((_, Object::Dictionary(d))) => d, + _ => return Err(bad), + }; + if let Some(nums) = get(doc, node, b"Nums") { + let items = nums.as_array().map_err(|_| bad)?; + for pair in items.chunks(2) { + if let [k, v] = pair { + if resolve(doc, k).map(|(_, o)| o.as_i64().ok()) == Some(Some(key)) { + return Ok(resolve(doc, v).map(|(_, o)| o)); + } + } + } + } + if let Some(kids) = get(doc, node, b"Kids") { + let items = kids.as_array().map_err(|_| bad)?; + stack.extend(items.iter().rev()); + } + } + Ok(None) +} + +/// The MCIDs of `page_id` in structure order (D33): a bounded DFS of `/StructTreeRoot /K` +/// (`/K` arrays in order, ≤ STRUCT_ORDER_NODES_MAX nodes, depth ≤ 64, visited set; `/Pg` +/// inherited down the tree; integer `/K` and `/Type /MCR` dicts with `/MCID` and optional `/Pg`). +/// `None` when the catalog has no structure tree or any budget, cycle or unresolvable reference +/// is met (the caller then uses XY-cut for the whole page). +pub fn structure_mcids(doc: &Document, page_id: ObjectId) -> Option> { + let root = doc + .catalog() + .ok() + .and_then(|c| get(doc, c, b"StructTreeRoot"))? + .as_dict() + .ok()?; + let mut out = Vec::new(); + let mut seen_mcid = HashSet::new(); + let mut visited: HashSet = HashSet::new(); + let mut nodes = 0usize; + // (node, inherited /Pg, depth) + let mut stack: Vec<(&Object, Option, usize)> = Vec::new(); + let top = root.get(b"K").ok()?; + stack.push((top, None, 0)); + while let Some((raw, pg, depth)) = stack.pop() { + nodes += 1; + if nodes > STRUCT_ORDER_NODES_MAX || depth > STRUCT_ORDER_DEPTH_MAX { + return None; + } + if let Object::Reference(id) = raw { + if !visited.insert(*id) { + return None; + } + } + let (_, node) = resolve(doc, raw)?; + match node { + Object::Integer(mcid) => { + if pg == Some(page_id) && seen_mcid.insert(*mcid) { + out.push(*mcid); + } + } + Object::Array(items) => { + for item in items.iter().rev() { + stack.push((item, pg, depth + 1)); + } + } + Object::Dictionary(d) => { + push_struct_dict(d, pg, depth, page_id, &mut stack, &mut out, &mut seen_mcid)? + } + Object::Null => {} + _ => return None, + } + } + Some(out) +} + +fn push_struct_dict<'a>( + d: &'a Dictionary, + pg: Option, + depth: usize, + page_id: ObjectId, + stack: &mut Vec<(&'a Object, Option, usize)>, + out: &mut Vec, + seen_mcid: &mut HashSet, +) -> Option<()> { + let own_pg = match d.get(b"Pg").ok() { + None => pg, + Some(Object::Reference(id)) => Some(*id), + Some(_) => return None, + }; + if d.type_is(b"OBJR") { + return Some(()); + } + if d.type_is(b"MCR") { + if let Ok(Object::Integer(mcid)) = d.get(b"MCID") { + if own_pg == Some(page_id) && seen_mcid.insert(*mcid) { + out.push(*mcid); + } + } + return Some(()); + } + if let Ok(k) = d.get(b"K") { + stack.push((k, own_pg, depth + 1)); + } + Some(()) +} + +/// §B.10 `struct_actual_text`: one lookup (production walks many MCIDs through `PageStruct`). +#[cfg(test)] +pub fn struct_actual_text( + doc: &Document, + page_id: ObjectId, + mcid: i64, +) -> Result { + PageStruct::of(doc, page_id).actual_text(doc, mcid) +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/cff.rs b/src-tauri/src/pdf_engine/text_edit/testkit/cff.rs new file mode 100644 index 0000000..e8a8a61 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/cff.rs @@ -0,0 +1,387 @@ +//! CFF font program builder for tests (T2, SPEC §E.1): bare CFF programs (`FontFile3 /Type1C`, +//! `/CIDFontType0C`) with name-keyed or CID-keyed charsets, Standard / Expert / custom (format 0 +//! or 1, with supplements) encodings, minimal Type2 outlines or raw charstrings, global and +//! local subroutines, and padded Top/Private DICTs. + +/// The 391 CFF standard strings; standard glyph names must use these SIDs. +pub(crate) use crate::pdf_engine::text_edit::fonts::cff_layout::STANDARD_STRINGS; + +#[derive(Debug, Clone)] +pub struct CffGlyph { + pub name: String, + pub outline: bool, + pub cid: u16, + /// A raw Type2 charstring used instead of the box (or the bare `endchar`). + pub charstring: Option>, +} + +/// The program's own encoding (Top DICT operator 16). +#[derive(Debug, Clone)] +pub enum CffEncodingSpec { + Standard, + Expert, + /// Format 0: `codes[i]` → GID i+1. + Format0(Vec), + /// Format 1: ranges `(first, nLeft)` → consecutive GIDs from 1. + Format1(Vec<(u8, u8)>), +} + +#[derive(Debug, Clone)] +pub struct CffBuilder { + pub font_name: String, + /// GID 1, 2, … (GID 0 is a drawn `.notdef`). + pub glyphs: Vec, + pub encoding: CffEncodingSpec, + /// Encoding supplements `(code, glyph name)` (format high bit). + pub supplements: Vec<(u8, String)>, + pub cid_keyed: bool, + /// Global subroutines (Global Subr INDEX). + pub gsubrs: Vec>, + /// Local subroutines (`Subrs` of the Private DICT; of FD 0 for CID-keyed fonts). + pub subrs: Vec>, + /// Charset format: 0 (an array), or 1 / 2 with one range per glyph (the most ranges). + pub charset_format: u8, + /// CID-keyed fonts: FDSelect format 3 (one range) or 0 (one byte per glyph); all FD 0. + pub fdselect_format: u8, + /// Bytes appended to the Private DICT after its entries (numbers or operators the readers + /// skip), and to the Top DICT. + pub private_padding: Vec, + pub top_padding: Vec, +} + +impl CffBuilder { + pub fn new(font_name: &str) -> Self { + CffBuilder { + font_name: font_name.to_string(), + glyphs: Vec::new(), + encoding: CffEncodingSpec::Standard, + supplements: Vec::new(), + cid_keyed: false, + gsubrs: Vec::new(), + subrs: Vec::new(), + charset_format: 0, + fdselect_format: 3, + private_padding: Vec::new(), + top_padding: Vec::new(), + } + } + + /// A CID-keyed glyph (next GID) for `cid` with a raw charstring. + pub fn raw_cid_glyph(mut self, cid: u16, charstring: Vec) -> Self { + self.cid_keyed = true; + self.glyphs.push(CffGlyph { + name: String::new(), + outline: true, + cid, + charstring: Some(charstring), + }); + self + } + + /// A name-keyed glyph (next GID) with a raw charstring. + pub fn raw_glyph(mut self, name: &str, charstring: Vec) -> Self { + self.glyphs.push(CffGlyph { + name: name.to_string(), + outline: true, + cid: 0, + charstring: Some(charstring), + }); + self + } + + /// A name-keyed glyph (next GID). + pub fn glyph(mut self, name: &str, outline: bool) -> Self { + self.glyphs.push(CffGlyph { + name: name.to_string(), + outline, + cid: 0, + charstring: None, + }); + self + } + + /// A CID-keyed glyph (next GID) for `cid`. + pub fn cid_glyph(mut self, cid: u16, outline: bool) -> Self { + self.cid_keyed = true; + self.glyphs.push(CffGlyph { + name: String::new(), + outline, + cid, + charstring: None, + }); + self + } + + pub fn encoding(mut self, encoding: CffEncodingSpec) -> Self { + self.encoding = encoding; + self + } + + pub fn supplement(mut self, code: u8, name: &str) -> Self { + self.supplements.push((code, name.to_string())); + self + } + + pub fn build(&self) -> Vec { + let mut custom: Vec = Vec::new(); + let mut custom_sids: std::collections::HashMap = Default::default(); + let mut sid = |name: &str, custom: &mut Vec| -> u16 { + if let Some(i) = STANDARD_STRINGS.iter().position(|s| *s == name) { + return i as u16; + } + if let Some(sid) = custom_sids.get(name) { + return *sid; + } + custom.push(name.to_string()); + let sid = (391 + custom.len() - 1) as u16; + custom_sids.insert(name.to_string(), sid); + sid + }; + let (ros_registry, ros_ordering) = if self.cid_keyed { + (sid("Adobe", &mut custom), sid("Identity", &mut custom)) + } else { + (0, 0) + }; + let mut charset = vec![self.charset_format]; + for g in &self.glyphs { + let v = if self.cid_keyed { + g.cid + } else { + sid(&g.name, &mut custom) + }; + charset.extend_from_slice(&v.to_be_bytes()); + match self.charset_format { + 1 => charset.push(0), + 2 => charset.extend_from_slice(&0u16.to_be_bytes()), + _ => {} + } + } + let mut supplements = Vec::new(); + for (code, name) in &self.supplements { + supplements.push(*code); + supplements.extend_from_slice(&sid(name, &mut custom).to_be_bytes()); + } + let sup_flag = if self.supplements.is_empty() { 0 } else { 0x80 }; + let encoding: Option> = match &self.encoding { + CffEncodingSpec::Standard | CffEncodingSpec::Expert => None, + CffEncodingSpec::Format0(codes) => { + let mut e = vec![sup_flag, codes.len() as u8]; + e.extend_from_slice(codes); + Some(e) + } + CffEncodingSpec::Format1(ranges) => { + let mut e = vec![1 | sup_flag, ranges.len() as u8]; + for (first, left) in ranges { + e.extend_from_slice(&[*first, *left]); + } + Some(e) + } + }; + let encoding = encoding.map(|mut e| { + if !self.supplements.is_empty() { + e.push(self.supplements.len() as u8); + e.extend_from_slice(&supplements); + } + e + }); + let mut charstrings = vec![box_charstring()]; + for g in &self.glyphs { + charstrings.push(match (&g.charstring, g.outline) { + (Some(raw), _) => raw.clone(), + (None, true) => box_charstring(), + (None, false) => vec![14], + }); + } + let charstrings = index(&charstrings); + let mut private = vec![139, 20, 139, 21]; // defaultWidthX 0, nominalWidthX 0 + let subrs = if self.subrs.is_empty() { + Vec::new() + } else { + // Subrs (op 19), relative to the Private DICT: right after it (10 bytes + padding). + private.extend(int((10 + self.private_padding.len()) as i32)); + private.push(19); + index(&self.subrs) + }; + private.extend_from_slice(&self.private_padding); + let name_index = index(&[self.font_name.as_bytes().to_vec()]); + let strings = index( + &custom + .iter() + .map(|s| s.as_bytes().to_vec()) + .collect::>(), + ); + let gsubrs = index(&self.gsubrs); + let n_glyphs = (self.glyphs.len() + 1) as u16; + let layout = |top_len: usize| { + let mut off = + 4 + name_index.len() + index_len(&[top_len]) + strings.len() + gsubrs.len(); + let charset_off = off; + off += charset.len(); + let encoding_off = off; + off += encoding.as_ref().map_or(0, Vec::len); + let charstrings_off = off; + off += charstrings.len(); + let private_off = off; + off += private.len() + subrs.len(); + let fdarray_off = off; + off += index_len(&[font_dict(0, 0).len()]); + ( + charset_off, + encoding_off, + charstrings_off, + private_off, + fdarray_off, + off, + ) + }; + let top = |l: (usize, usize, usize, usize, usize, usize)| { + let ( + charset_off, + encoding_off, + charstrings_off, + private_off, + fdarray_off, + fdselect_off, + ) = l; + let mut d = Vec::new(); + if self.cid_keyed { + d.extend(int(ros_registry as i32)); + d.extend(int(ros_ordering as i32)); + d.extend(int(0)); + d.extend([12, 30]); + } + d.extend(int(charset_off as i32)); + d.push(15); + if !self.cid_keyed { + let enc = match (&self.encoding, &encoding) { + (CffEncodingSpec::Expert, _) => 1, + (_, Some(_)) => encoding_off as i32, + _ => 0, + }; + d.extend(int(enc)); + d.push(16); + } + d.extend(int(charstrings_off as i32)); + d.push(17); + if self.cid_keyed { + d.extend(int(fdarray_off as i32)); + d.extend([12, 36]); + d.extend(int(fdselect_off as i32)); + d.extend([12, 37]); + } else { + d.extend(int(private.len() as i32)); + d.extend(int(private_off as i32)); + d.push(18); + } + d.extend_from_slice(&self.top_padding); + d + }; + let top_len = top((0, 0, 0, 0, 0, 0)).len(); + let l = layout(top_len); + let top_dict = top(l); + assert_eq!(top_dict.len(), top_len); + let mut out = vec![1, 0, 4, 4]; + out.extend(&name_index); + out.extend(index(&[top_dict])); + out.extend(&strings); + out.extend(&gsubrs); + assert_eq!(out.len(), l.0); + out.extend(&charset); + if let Some(e) = &encoding { + out.extend(e); + } + out.extend(&charstrings); + out.extend(&private); + out.extend(&subrs); + out.extend(index(&[font_dict(private.len(), l.3)])); + if self.fdselect_format == 0 { + // FDSelect format 0: FD 0 for every glyph + out.push(0); + out.extend(std::iter::repeat(0u8).take(usize::from(n_glyphs))); + } else { + // FDSelect format 3: one range, all glyphs in FD 0 + out.push(3); + out.extend_from_slice(&1u16.to_be_bytes()); + out.extend_from_slice(&0u16.to_be_bytes()); + out.push(0); + out.extend_from_slice(&n_glyphs.to_be_bytes()); + } + out + } +} + +/// A Font DICT (FDArray entry) pointing at a Private DICT. +fn font_dict(private_len: usize, private_off: usize) -> Vec { + let mut d = int(private_len as i32); + d.extend(int(private_off as i32)); + d.push(18); + d +} + +/// A fixed-width (5-byte) DICT integer, so offsets can be patched without moving data. +fn int(v: i32) -> Vec { + let mut b = vec![29]; + b.extend_from_slice(&v.to_be_bytes()); + b +} + +/// A Type2 number. +pub fn t2(v: i32) -> Vec { + match v { + -107..=107 => vec![(v + 139) as u8], + 108..=1131 => { + let w = v - 108; + vec![(w / 256 + 247) as u8, (w % 256) as u8] + } + -1131..=-108 => { + let w = -v - 108; + vec![(w / 256 + 251) as u8, (w % 256) as u8] + } + _ => { + let mut b = vec![28]; + b.extend_from_slice(&(v as i16).to_be_bytes()); + b + } + } +} + +/// `100 0 rmoveto 400 0 0 500 -400 0 rlineto endchar`: a drawn box. +pub fn box_charstring() -> Vec { + let mut c = Vec::new(); + for v in [100, 0] { + c.extend(t2(v)); + } + c.push(21); + for v in [400, 0, 0, 500, -400, 0] { + c.extend(t2(v)); + } + c.push(5); + c.push(14); + c +} + +/// A CFF INDEX with offSize 4. +pub fn index(items: &[Vec]) -> Vec { + let mut out = (items.len() as u16).to_be_bytes().to_vec(); + if items.is_empty() { + return out; + } + out.push(4); + let mut off: u32 = 1; + out.extend_from_slice(&off.to_be_bytes()); + for item in items { + off += item.len() as u32; + out.extend_from_slice(&off.to_be_bytes()); + } + for item in items { + out.extend_from_slice(item); + } + out +} + +fn index_len(lens: &[usize]) -> usize { + if lens.is_empty() { + return 2; + } + 3 + 4 * (lens.len() + 1) + lens.iter().sum::() +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/fakes.rs b/src-tauri/src/pdf_engine/text_edit/testkit/fakes.rs new file mode 100644 index 0000000..c6cdef0 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/fakes.rs @@ -0,0 +1,636 @@ +//! Fakes the publish gate must reject (SPEC §E.5, T4): cover-and-overlay (a new content stream +//! over the old text, or a real `qpdf --overlay`), a raster of the page, annotations, re-attached +//! old content, catalog and font changes, legacy-filter tampering, and writer bugs (wrong bytes, +//! wrong target). Built only with this test kit and real qpdf (`--overlay`, `--update-from-json`): +//! no production writer is used, so a production bug cannot make a fake pass by construction. + +use lopdf::{Dictionary, Document, Object, ObjectId}; +use serde_json::{json, Map, Value}; +use std::path::Path; +use std::process::Command; + +const B64: &[u8; 64] = b"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"; + +/// Standard base64 with padding. +pub fn base64(data: &[u8]) -> String { + let mut out = String::with_capacity(data.len().div_ceil(3) * 4); + for chunk in data.chunks(3) { + let b = [ + chunk[0], + chunk.get(1).copied().unwrap_or(0), + chunk.get(2).copied().unwrap_or(0), + ]; + let n = (u32::from(b[0]) << 16) | (u32::from(b[1]) << 8) | u32::from(b[2]); + for i in 0..4 { + if i <= chunk.len() { + out.push(char::from(B64[((n >> (18 - 6 * i)) & 63) as usize])); + } else { + out.push('='); + } + } + } + out +} + +/// A PDF name as qpdf JSON writes it (`/Name`, `#xx` for anything unusual). +fn json_name(n: &[u8]) -> String { + let mut s = String::from("/"); + for &b in n { + if b.is_ascii_alphanumeric() || b"-_.+*".contains(&b) { + s.push(char::from(b)); + } else { + s.push_str(&format!("#{b:02x}")); + } + } + s +} + +/// A lopdf value in qpdf JSON v2 form. +pub fn json_value(obj: &Object) -> Value { + match obj { + Object::Null => Value::Null, + Object::Boolean(b) => json!(b), + Object::Integer(i) => json!(i), + Object::Real(r) => json!(f64::from(*r)), + Object::Name(n) => json!(json_name(n)), + Object::String(s, _) => json!(format!( + "b:{}", + s.iter().map(|b| format!("{b:02x}")).collect::() + )), + Object::Array(items) => Value::Array(items.iter().map(json_value).collect()), + Object::Dictionary(d) => json_dict(d), + Object::Stream(s) => json_dict(&s.dict), + Object::Reference((n, g)) => json!(format!("{n} {g} R")), + } +} + +pub fn json_dict(d: &Dictionary) -> Value { + let mut m = Map::new(); + for (k, v) in d.iter() { + m.insert(json_name(k), json_value(v)); + } + Value::Object(m) +} + +pub fn stream_obj(dict: Value, data: &[u8]) -> Value { + json!({ "stream": { "dict": dict, "data": base64(data) } }) +} + +pub fn value_obj(v: Value) -> Value { + json!({ "value": v }) +} + +pub fn obj_key(id: ObjectId) -> String { + format!("obj:{} {} R", id.0, id.1) +} + +/// The file parsed by lopdf (test code may load freely). +pub fn load(path: &Path) -> Document { + Document::load(path).unwrap_or_else(|e| panic!("fake input {}: {e}", path.display())) +} + +/// The page object ids of `doc`, in order. +pub fn page_ids(doc: &Document) -> Vec { + doc.get_pages().into_values().collect() +} + +fn run_qpdf(qpdf: &Path, args: &[&std::ffi::OsStr]) -> Result<(), String> { + let out = Command::new(qpdf) + .args(args) + .output() + .map_err(|e| format!("qpdf: {e}"))?; + match out.status.code() { + Some(0) | Some(3) => Ok(()), + c => Err(format!( + "qpdf exited {c:?}: {}", + String::from_utf8_lossy(&out.stderr) + )), + } +} + +/// `qpdf input output --decode-level=none --update-from-json=…` with `objects`. +pub fn update( + qpdf: &Path, + input: &Path, + output: &Path, + objects: Map, + max_id: u32, +) -> Result<(), String> { + let doc = json!({ "qpdf": [ + { "jsonversion": 2, "pushedinheritedpageresources": false, "calledgetallpages": false, "maxobjectid": max_id }, + Value::Object(objects), + ]}); + let json_path = output.with_extension("fake.json"); + std::fs::write( + &json_path, + serde_json::to_vec(&doc).map_err(|e| e.to_string())?, + ) + .map_err(|e| e.to_string())?; + let mut arg = std::ffi::OsString::from("--update-from-json="); + arg.push(json_path.as_os_str()); + let r = run_qpdf( + qpdf, + &[ + input.as_os_str(), + output.as_os_str(), + std::ffi::OsStr::new("--decode-level=none"), + arg.as_os_str(), + ], + ); + let _ = std::fs::remove_file(&json_path); + r +} + +/// Replaces the decoded data of streams (a writer that wrote other bytes than planned). +pub fn replace_streams( + qpdf: &Path, + input: &Path, + output: &Path, + streams: &[(ObjectId, Vec)], +) -> Result<(), String> { + let doc = load(input); + let mut m = Map::new(); + for (id, data) in streams { + m.insert(obj_key(*id), stream_obj(json!({}), data)); + } + update(qpdf, input, output, m, doc.max_id) +} + +/// Writes `raw` as the already-filtered data of stream `id` with dictionary `dict`. +pub fn raw_stream( + qpdf: &Path, + input: &Path, + output: &Path, + id: ObjectId, + dict: Value, + raw: &[u8], +) -> Result<(), String> { + let doc = load(input); + let mut m = Map::new(); + m.insert(obj_key(id), stream_obj(dict, raw)); + update(qpdf, input, output, m, doc.max_id) +} + +/// Replaces stream `id`'s dictionary, keeping its data. +pub fn stream_dict( + qpdf: &Path, + input: &Path, + output: &Path, + id: ObjectId, + dict: Value, +) -> Result<(), String> { + let doc = load(input); + let mut m = Map::new(); + m.insert(obj_key(id), json!({ "stream": { "dict": dict } })); + update(qpdf, input, output, m, doc.max_id) +} + +/// Replaces object `id` with a (non-stream) value. +pub fn set_value( + qpdf: &Path, + input: &Path, + output: &Path, + id: ObjectId, + value: Value, +) -> Result<(), String> { + let doc = load(input); + let mut m = Map::new(); + m.insert(obj_key(id), value_obj(value)); + update(qpdf, input, output, m, doc.max_id) +} + +fn page_dict(doc: &Document, page: ObjectId) -> Dictionary { + doc.get_dictionary(page).cloned().unwrap_or_default() +} + +fn contents_refs(d: &Dictionary) -> Vec { + match d.get(b"Contents") { + Ok(Object::Reference(r)) => vec![json_value(&Object::Reference(*r))], + Ok(Object::Array(items)) => items.iter().map(json_value).collect(), + _ => Vec::new(), + } +} + +/// Adds content streams before and after the page's parts (as flattening passes do). +pub fn add_parts( + qpdf: &Path, + input: &Path, + output: &Path, + page_index: usize, + before: &[&[u8]], + after: &[&[u8]], +) -> Result<(), String> { + let doc = load(input); + let page = *page_ids(&doc).get(page_index).ok_or("page")?; + let mut d = json_dict(&page_dict(&doc, page)); + let mut next = doc.max_id + 1; + let mut m = Map::new(); + let mut new_ref = |data: &[u8], m: &mut Map| { + let id = (next, 0); + next += 1; + m.insert(obj_key(id), stream_obj(json!({}), data)); + json!(format!("{} 0 R", id.0)) + }; + let mut parts: Vec = before.iter().map(|p| new_ref(p, &mut m)).collect(); + parts.extend(contents_refs(&page_dict(&doc, page))); + parts.extend(after.iter().map(|p| new_ref(p, &mut m))); + if let Some(o) = d.as_object_mut() { + o.insert("/Contents".into(), Value::Array(parts)); + } + m.insert(obj_key(page), value_obj(d)); + update(qpdf, input, output, m, doc.max_id) +} + +/// GATE-18 cover-and-overlay F1: the original parts stay; a new content stream paints a white box +/// over `cover` (`[x, y, w, h]`) and draws `text` with the page font `font` at its corner. +pub fn cover_and_overlay_f1( + qpdf: &Path, + input: &Path, + output: &Path, + page_index: usize, + cover: [f64; 4], + font: &str, + text: &str, +) -> Result<(), String> { + let [x, y, w, h] = cover; + let fake = format!( + "q 1 g {x} {y} {w} {h} re f Q BT /{font} 12 Tf {x} {} Td ({text}) Tj ET", + y + 2.0 + ); + add_parts(qpdf, input, output, page_index, &[], &[fake.as_bytes()]) +} + +/// A one-page PDF (Helvetica `/F1`) drawing a white box over `cover` and `text` in it. +pub fn cover_page(media: [f64; 4], cover: [f64; 4], text: &str) -> Vec { + use super::producers::{DocBuilder, PageSpec, HELVETICA}; + let [x, y, w, h] = cover; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let content = format!( + "q 1 g {x} {y} {w} {h} re f Q BT /F1 12 Tf {x} {} Td ({text}) Tj ET", + y + 2.0 + ); + let media = format!("[{} {} {} {}]", media[0], media[1], media[2], media[3]); + d.page( + PageSpec::new(content.as_bytes(), &format!("/Font << /F1 {f} 0 R >>")).media(Some(&media)), + ); + d.build() +} + +/// GATE-19 cover-and-overlay F2: a real `qpdf --overlay` of a white box + text page onto page +/// `page_1` (1-based). +pub fn overlay_f2( + qpdf: &Path, + input: &Path, + output: &Path, + page_1: u32, + media: [f64; 4], + cover: [f64; 4], + text: &str, +) -> Result<(), String> { + let cover_pdf = output.with_extension("cover.pdf"); + std::fs::write(&cover_pdf, cover_page(media, cover, text)).map_err(|e| e.to_string())?; + overlay(qpdf, input, output, &cover_pdf, Some(page_1)) +} + +/// `qpdf input --overlay overlay [--to=n] -- output` (every page when `to` is `None`). +pub fn overlay( + qpdf: &Path, + input: &Path, + output: &Path, + overlay_pdf: &Path, + to: Option, +) -> Result<(), String> { + let to_arg = to.map(|n| format!("--to={n}")); + let mut args: Vec<&std::ffi::OsStr> = vec![ + input.as_os_str(), + std::ffi::OsStr::new("--overlay"), + overlay_pdf.as_os_str(), + ]; + if let Some(t) = &to_arg { + args.push(std::ffi::OsStr::new(t)); + } + args.push(std::ffi::OsStr::new("--")); + args.push(output.as_os_str()); + run_qpdf(qpdf, &args) +} + +/// `qpdf --empty --pages a.pdf b.pdf … -- output` (assembly with other files). +pub fn assemble(qpdf: &Path, inputs: &[&Path], output: &Path) -> Result<(), String> { + let mut args: Vec<&std::ffi::OsStr> = vec![ + std::ffi::OsStr::new("--empty"), + std::ffi::OsStr::new("--pages"), + ]; + for i in inputs { + args.push(i.as_os_str()); + } + args.push(std::ffi::OsStr::new("--")); + args.push(output.as_os_str()); + run_qpdf(qpdf, &args) +} + +/// Reads a binary PPM (P6) written by pdftoppm. +fn read_ppm(path: &Path) -> Result<(u32, u32, Vec), String> { + let data = std::fs::read(path).map_err(|e| e.to_string())?; + let mut fields = Vec::new(); + let mut at = 0usize; + while fields.len() < 4 { + while data.get(at).is_some_and(u8::is_ascii_whitespace) { + at += 1; + } + let start = at; + while data.get(at).is_some_and(|c| !c.is_ascii_whitespace()) { + at += 1; + } + fields.push(String::from_utf8_lossy(&data[start..at]).into_owned()); + } + let w: u32 = fields[1].parse().map_err(|_| "ppm width")?; + let h: u32 = fields[2].parse().map_err(|_| "ppm height")?; + Ok((w, h, data[at + 1..].to_vec())) +} + +/// GATE-20 raster fake: the page content replaced by `q W 0 0 H X Y cm /Im0 Do Q` drawing a +/// pdftoppm render of `rendered` (the honestly edited page) over the media box. +pub fn raster_fake( + qpdf: &Path, + pdftoppm: &Path, + rendered: &Path, + input: &Path, + output: &Path, + page_index: usize, +) -> Result<(), String> { + let prefix = output.with_extension("render"); + let page = (page_index + 1).to_string(); + let status = Command::new(pdftoppm) + .args(["-r", "72", "-f", &page, "-l", &page, "-singlefile"]) + .arg(rendered) + .arg(&prefix) + .status() + .map_err(|e| e.to_string())?; + if !status.success() { + return Err("pdftoppm".into()); + } + let mut ppm = prefix.into_os_string(); + ppm.push(".ppm"); + let ppm = std::path::PathBuf::from(ppm); + let (w, h, rgb) = read_ppm(&ppm)?; + let _ = std::fs::remove_file(&ppm); + let doc = load(input); + let page_id = *page_ids(&doc).get(page_index).ok_or("page")?; + let media = doc + .get_dictionary(page_id) + .ok() + .and_then(|d| d.get(b"MediaBox").ok()) + .and_then(|m| m.as_array().ok()) + .map(|a| { + a.iter() + .filter_map(|o| o.as_float().ok()) + .collect::>() + }) + .unwrap_or_else(|| vec![0.0, 0.0, 612.0, 792.0]); + let (x0, y0, x1, y1) = (media[0], media[1], media[2], media[3]); + let image = (doc.max_id + 1, 0); + let content = (doc.max_id + 2, 0); + let mut m = Map::new(); + m.insert( + obj_key(image), + stream_obj( + json!({ "/Type": "/XObject", "/Subtype": "/Image", "/Width": w, "/Height": h, + "/ColorSpace": "/DeviceRGB", "/BitsPerComponent": 8 }), + &rgb, + ), + ); + let draw = format!("q {} 0 0 {} {x0} {y0} cm /Im0 Do Q", x1 - x0, y1 - y0); + m.insert(obj_key(content), stream_obj(json!({}), draw.as_bytes())); + let mut d = json_dict(&page_dict(&doc, page_id)); + if let Some(o) = d.as_object_mut() { + o.insert("/Contents".into(), json!(format!("{} 0 R", content.0))); + o.insert( + "/Resources".into(), + json!({ "/XObject": { "/Im0": format!("{} 0 R", image.0) } }), + ); + } + m.insert(obj_key(page_id), value_obj(d)); + update(qpdf, input, output, m, doc.max_id) +} + +/// Adds an annotation dictionary (`annot`, with `/P` set here) to page `page_index`. +fn add_annotation( + qpdf: &Path, + input: &Path, + output: &Path, + page_index: usize, + mut annot: Map, + appearance: Option<&[u8]>, +) -> Result<(), String> { + let doc = load(input); + let page = *page_ids(&doc).get(page_index).ok_or("page")?; + let annot_id = (doc.max_id + 1, 0); + let mut m = Map::new(); + annot.insert("/P".into(), json!(format!("{} {} R", page.0, page.1))); + if let Some(ap) = appearance { + let ap_id = (doc.max_id + 2, 0); + let rect = annot + .get("/Rect") + .cloned() + .unwrap_or(json!([0, 0, 100, 20])); + m.insert( + obj_key(ap_id), + stream_obj( + json!({ "/Type": "/XObject", "/Subtype": "/Form", "/BBox": rect }), + ap, + ), + ); + annot.insert("/AP".into(), json!({ "/N": format!("{} 0 R", ap_id.0) })); + } + m.insert(obj_key(annot_id), value_obj(Value::Object(annot))); + let mut d = json_dict(&page_dict(&doc, page)); + if let Some(o) = d.as_object_mut() { + let mut annots = match o.get("/Annots") { + Some(Value::Array(a)) => a.clone(), + _ => Vec::new(), + }; + annots.push(json!(format!("{} 0 R", annot_id.0))); + o.insert("/Annots".into(), Value::Array(annots)); + } + m.insert(obj_key(page), value_obj(d)); + update(qpdf, input, output, m, doc.max_id) +} + +/// GATE-22: a FreeText annotation showing `text` over `rect` (`[x0, y0, x1, y1]`). +pub fn freetext( + qpdf: &Path, + input: &Path, + output: &Path, + page_index: usize, + rect: [f64; 4], + text: &str, +) -> Result<(), String> { + let mut a = Map::new(); + a.insert("/Type".into(), json!("/Annot")); + a.insert("/Subtype".into(), json!("/FreeText")); + a.insert("/Rect".into(), json!(rect)); + a.insert( + "/Contents".into(), + json!(format!( + "b:{}", + text.bytes().map(|b| format!("{b:02x}")).collect::() + )), + ); + a.insert("/DA".into(), json!("u:/Helv 12 Tf 0 g")); + let ap = format!( + "q 1 g 0 0 {} {} re f Q", + rect[2] - rect[0], + rect[3] - rect[1] + ); + add_annotation(qpdf, input, output, page_index, a, Some(ap.as_bytes())) +} + +/// GATE-26: a Link annotation on page `page_index`. +pub fn link(qpdf: &Path, input: &Path, output: &Path, page_index: usize) -> Result<(), String> { + let mut a = Map::new(); + a.insert("/Type".into(), json!("/Annot")); + a.insert("/Subtype".into(), json!("/Link")); + a.insert("/Rect".into(), json!([72, 72, 144, 96])); + a.insert("/Border".into(), json!([0, 0, 0])); + a.insert( + "/A".into(), + json!({ "/S": "/URI", "/URI": "u:https://example.invalid/" }), + ); + add_annotation(qpdf, input, output, page_index, a, None) +} + +/// GATE-24: `old` (the original content) re-attached as an unused Form XObject `/Old`. +pub fn reattach_old( + qpdf: &Path, + input: &Path, + output: &Path, + page_index: usize, + old: &[u8], +) -> Result<(), String> { + let doc = load(input); + let page = *page_ids(&doc).get(page_index).ok_or("page")?; + let form = (doc.max_id + 1, 0); + let mut m = Map::new(); + m.insert( + obj_key(form), + stream_obj( + json!({ "/Type": "/XObject", "/Subtype": "/Form", "/BBox": [0, 0, 612, 792] }), + old, + ), + ); + let pd = page_dict(&doc, page); + let mut d = json_dict(&pd); + let mut resources = match pd.get(b"Resources") { + Ok(Object::Reference(r)) => doc + .get_dictionary(*r) + .map(json_dict) + .unwrap_or_else(|_| json!({})), + Ok(Object::Dictionary(r)) => json_dict(r), + _ => json!({}), + }; + if let Some(r) = resources.as_object_mut() { + r.insert( + "/XObject".into(), + json!({ "/Old": format!("{} 0 R", form.0) }), + ); + } + if let Some(o) = d.as_object_mut() { + o.insert("/Resources".into(), resources); + } + m.insert(obj_key(page), value_obj(d)); + update(qpdf, input, output, m, doc.max_id) +} + +/// The catalog id of `doc`. +pub fn catalog_id(doc: &Document) -> ObjectId { + doc.trailer + .get(b"Root") + .and_then(Object::as_reference) + .unwrap_or((1, 0)) +} + +/// GATE-27: the first group of `/OCProperties /D /ON` moved to `/OFF` (the catalog rewritten). +pub fn ocg_off(qpdf: &Path, input: &Path, output: &Path, ocg: ObjectId) -> Result<(), String> { + let doc = load(input); + let cat_id = catalog_id(&doc); + let mut cat = json_dict(&doc.get_dictionary(cat_id).cloned().unwrap_or_default()); + let r = json!(format!("{} {} R", ocg.0, ocg.1)); + if let Some(props) = cat.get_mut("/OCProperties").and_then(Value::as_object_mut) { + if let Some(d) = props.get_mut("/D").and_then(Value::as_object_mut) { + let mut off = match d.get("/OFF") { + Some(Value::Array(a)) => a.clone(), + _ => Vec::new(), + }; + off.push(r.clone()); + d.insert("/OFF".into(), Value::Array(off)); + if let Some(Value::Array(on)) = d.get_mut("/ON") { + on.retain(|v| *v != r); + } + } + } + let mut m = Map::new(); + m.insert(obj_key(cat_id), value_obj(cat)); + update(qpdf, input, output, m, doc.max_id) +} + +/// GATE-28: entry `index` of font `font`'s `/Widths` set to `value`. +pub fn widths_changed( + qpdf: &Path, + input: &Path, + output: &Path, + font: ObjectId, + index: usize, + value: f64, +) -> Result<(), String> { + let doc = load(input); + let mut d = json_dict(&doc.get_dictionary(font).cloned().unwrap_or_default()); + if let Some(Value::Array(w)) = d.get_mut("/Widths") { + if let Some(slot) = w.get_mut(index) { + *slot = json!(value); + } + } + let mut m = Map::new(); + m.insert(obj_key(font), value_obj(d)); + update(qpdf, input, output, m, doc.max_id) +} + +/// APP-04 (probe P4 shape): an update aimed at the page object instead of its content stream. +pub fn mistargeted( + qpdf: &Path, + input: &Path, + output: &Path, + page_index: usize, + data: &[u8], +) -> Result<(), String> { + let doc = load(input); + let page = *page_ids(&doc).get(page_index).ok_or("page")?; + let mut m = Map::new(); + m.insert(obj_key(page), stream_obj(json!({}), data)); + update(qpdf, input, output, m, doc.max_id) +} + +/// Review H-1: `program`, a CID-keyed program from `testkit::cff::CffBuilder`, with `font_matrix` +/// (DICT operand bytes) as the `FontMatrix` of its one Font DICT. A new FDArray INDEX is appended +/// and the Top DICT's fixed-width `FDArray` offset (`29 12 36`) is patched to point at it. +pub fn cff_with_font_dict_matrix(program: &[u8], font_matrix: &[u8]) -> Vec { + use crate::pdf_engine::text_edit::testkit::cff::index; + let p = program; + let at = p + .windows(7) + .position(|w| w[0] == 29 && w[5] == 12 && w[6] == 36) + .expect("an FDArray entry"); + let old = u32::from_be_bytes([p[at + 1], p[at + 2], p[at + 3], p[at + 4]]) as usize; + // The builder's FDArray: count 1, offSize 4, offsets 1 and 1 + len, then the Font DICT. + assert_eq!(&p[old..old + 7], &[0, 1, 4, 0, 0, 0, 1], "one Font DICT"); + let end = u32::from_be_bytes([p[old + 7], p[old + 8], p[old + 9], p[old + 10]]); + let dict = &p[old + 11..old + 10 + end as usize]; + let mut font_dict = font_matrix.to_vec(); + font_dict.extend([12, 7]); + font_dict.extend_from_slice(dict); + let mut out = p.to_vec(); + let new_at = u32::try_from(out.len()).expect("small program"); + out[at + 1..at + 5].copy_from_slice(&new_at.to_be_bytes()); + out.extend(index(&[font_dict])); + out +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/fonts.rs b/src-tauri/src/pdf_engine/text_edit/testkit/fonts.rs new file mode 100644 index 0000000..2a5ea4a --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/fonts.rs @@ -0,0 +1,362 @@ +//! PDF font dictionary helpers for tests (T2, SPEC §E.1): simple and Type0 font objects for +//! `PdfBuilder` documents (programs, descriptors, encodings, widths, ToUnicode CMaps), one-page +//! documents carrying them, and loading their `FontModel`s through the real snapshot reader. + +use super::pdf::PdfBuilder; +use super::ttf::TtfBuilder; +use super::type1::Type1File; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::fonts::{agl, FontCache, FontKey, FontModel}; +use crate::pdf_engine::text_edit::limits::PAGE_DECODE_BUDGET; +use crate::pdf_engine::text_edit::snapshot::{snapshot_from_bytes, SourceSnapshot}; +use lopdf::Object; +use std::sync::Arc; + +/// An embedded program and the descriptor key it goes under. +#[derive(Debug, Clone)] +pub enum Program { + None, + /// `FontFile2`. + TrueType(Vec), + /// `FontFile` with `Length1/2/3`. + Type1(Type1File), + /// `FontFile3 /Subtype /Type1C`. + Cff(Vec), + /// `FontFile3 /Subtype /CIDFontType0C`. + CidCff(Vec), + /// `FontFile3 /Subtype /OpenType`. + OpenType(Vec), + /// Any key / FontFile3 subtype / data (mismatch and corruption tests). + Raw { + key: &'static str, + subtype: Option<&'static str>, + data: Vec, + }, +} + +/// A simple font (`/Type1`, `/TrueType`, `/MMType1`, `/Type3`). +#[derive(Debug, Clone)] +pub struct SimpleFont { + pub subtype: &'static str, + pub base_font: String, + /// Raw PDF for `/Encoding`, e.g. `/WinAnsiEncoding` or `<< /Differences [65 /A] >>`. + pub encoding: Option, + pub first_char: u32, + pub widths: Option>, + /// `/Flags`; a descriptor is written whenever flags, a program or extras are given. + pub flags: Option, + /// Extra descriptor entries, e.g. `/MissingWidth 300 /CharSet (/a/b)`. + pub descriptor_extra: String, + pub program: Program, + pub tounicode: Option>, + /// Extra font dictionary entries. + pub font_extra: String, + /// Compress program and ToUnicode streams with Flate. + pub flate: bool, +} + +impl SimpleFont { + pub fn new(subtype: &'static str, base_font: &str) -> Self { + SimpleFont { + subtype, + base_font: base_font.to_string(), + encoding: None, + first_char: 0, + widths: None, + flags: None, + descriptor_extra: String::new(), + program: Program::None, + tounicode: None, + font_extra: String::new(), + flate: true, + } + } +} + +/// A Type0 font with one descendant CIDFont. +#[derive(Debug, Clone)] +pub struct Type0Font { + pub base_font: String, + /// Raw PDF for `/Encoding` (default `/Identity-H`). + pub encoding: String, + pub cid_subtype: &'static str, + /// Raw PDF array for `/W`. + pub w: Option, + pub dw: Option, + /// `None` = absent; `Some(None)` = `/Identity`; `Some(Some(bytes))` = stream. + pub cid_to_gid: Option>>, + pub flags: Option, + pub descriptor_extra: String, + pub cid_extra: String, + pub program: Program, + pub tounicode: Option>, +} + +impl Type0Font { + pub fn new(cid_subtype: &'static str, base_font: &str) -> Self { + Type0Font { + base_font: base_font.to_string(), + encoding: "/Identity-H".to_string(), + cid_subtype, + w: None, + dw: None, + cid_to_gid: None, + flags: Some(4), + descriptor_extra: String::new(), + cid_extra: String::new(), + program: Program::None, + tounicode: None, + } + } +} + +fn add_data(b: &mut PdfBuilder, dict: &str, data: &[u8], flate: bool) -> u32 { + if flate { + b.add_flate(dict, data) + } else { + b.add_stream(dict, data) + } +} + +/// Writes the program stream; returns the descriptor entry (`/FontFile2 12 0 R`). +fn add_program(b: &mut PdfBuilder, program: &Program, flate: bool) -> Option { + let (key, dict, data) = match program { + Program::None => return None, + Program::TrueType(d) => ("FontFile2", String::new(), d.clone()), + Program::Type1(t) => ( + "FontFile", + format!( + "/Length1 {} /Length2 {} /Length3 {}", + t.length1, t.length2, t.length3 + ), + t.data.clone(), + ), + Program::Cff(d) => ("FontFile3", "/Subtype /Type1C".to_string(), d.clone()), + Program::CidCff(d) => ( + "FontFile3", + "/Subtype /CIDFontType0C".to_string(), + d.clone(), + ), + Program::OpenType(d) => ("FontFile3", "/Subtype /OpenType".to_string(), d.clone()), + Program::Raw { key, subtype, data } => ( + *key, + subtype.map_or(String::new(), |s| format!("/Subtype /{s}")), + data.clone(), + ), + }; + let id = add_data(b, &dict, &data, flate); + Some(format!("/{key} {id} 0 R")) +} + +fn add_descriptor( + b: &mut PdfBuilder, + name: &str, + flags: Option, + program: Option, + extra: &str, +) -> Option { + if flags.is_none() && program.is_none() && extra.is_empty() { + return None; + } + Some(b.add(format!( + "<< /Type /FontDescriptor /FontName /{name} {} /FontBBox [0 -200 1000 800] \ + /ItalicAngle 0 /Ascent 800 /Descent -200 /CapHeight 700 /StemV 80 {} {extra} >>", + flags.map_or(String::new(), |f| format!("/Flags {f}")), + program.unwrap_or_default(), + ))) +} + +/// Adds a simple font object; returns its id. +pub fn add_simple(b: &mut PdfBuilder, f: &SimpleFont) -> u32 { + let program = add_program(b, &f.program, f.flate); + let descriptor = add_descriptor(b, &f.base_font, f.flags, program, &f.descriptor_extra); + let widths = f.widths.as_ref().map_or(String::new(), |w| { + format!( + "/FirstChar {} /LastChar {} /Widths [{}]", + f.first_char, + f.first_char + w.len() as u32 - 1, + w.iter() + .map(|v| v.to_string()) + .collect::>() + .join(" ") + ) + }); + let tounicode = f.tounicode.as_ref().map_or(String::new(), |t| { + format!("/ToUnicode {} 0 R", add_data(b, "", t, f.flate)) + }); + b.add(format!( + "<< /Type /Font /Subtype /{} /BaseFont /{} {} {widths} {} {tounicode} {} >>", + f.subtype, + f.base_font, + f.encoding + .as_ref() + .map_or(String::new(), |e| format!("/Encoding {e}")), + descriptor.map_or(String::new(), |d| format!("/FontDescriptor {d} 0 R")), + f.font_extra, + )) +} + +/// Adds a Type0 font object (and its descendant); returns the Type0 font's id. +pub fn add_type0(b: &mut PdfBuilder, f: &Type0Font) -> u32 { + let program = add_program(b, &f.program, true); + let descriptor = add_descriptor(b, &f.base_font, f.flags, program, &f.descriptor_extra); + let cid_to_gid = match &f.cid_to_gid { + None => String::new(), + Some(None) => "/CIDToGIDMap /Identity".to_string(), + Some(Some(map)) => format!("/CIDToGIDMap {} 0 R", b.add_flate("", map)), + }; + let descendant = b.add(format!( + "<< /Type /Font /Subtype /{} /BaseFont /{} /CIDSystemInfo << /Registry (Adobe) \ + /Ordering (Identity) /Supplement 0 >> {} {} {} {cid_to_gid} {} >>", + f.cid_subtype, + f.base_font, + descriptor.map_or(String::new(), |d| format!("/FontDescriptor {d} 0 R")), + f.w.as_ref().map_or(String::new(), |w| format!("/W {w}")), + f.dw.map_or(String::new(), |d| format!("/DW {d}")), + f.cid_extra, + )); + let tounicode = f.tounicode.as_ref().map_or(String::new(), |t| { + format!("/ToUnicode {} 0 R", b.add_flate("", t)) + }); + b.add(format!( + "<< /Type /Font /Subtype /Type0 /BaseFont /{} /Encoding {} /DescendantFonts [{descendant} 0 R] {tounicode} >>", + f.base_font, f.encoding + )) +} + +/// A ToUnicode CMap whose body (between `begincmap` boilerplate) is `body`. +pub fn cmap(body: &str) -> Vec { + format!( + "/CIDInit /ProcSet findresource begin\n12 dict begin\nbegincmap\n\ + /CIDSystemInfo << /Registry (Adobe) /Ordering (UCS) /Supplement 0 >> def\n\ + /CMapName /Adobe-Identity-UCS def\n/CMapType 2 def\n{body}\nendcmap\n\ + CMapName currentdict /CMap defineresource pop\nend\nend\n" + ) + .into_bytes() +} + +/// UTF-16BE hex of `text`. +pub fn utf16_hex(text: &str) -> String { + text.encode_utf16().map(|u| format!("{u:04X}")).collect() +} + +/// A bfchar ToUnicode CMap: `(code, code length in bytes, text)`. +pub fn tounicode_bfchar(entries: &[(u32, usize, &str)]) -> Vec { + let two = entries.iter().any(|e| e.1 == 2); + let space = if two { "<0000> " } else { "<00> " }; + let mut body = format!("1 begincodespacerange\n{space}\nendcodespacerange\n"); + body.push_str(&format!("{} beginbfchar\n", entries.len())); + for (code, len, text) in entries { + body.push_str(&format!( + "<{:0width$X}> <{}>\n", + code, + utf16_hex(text), + width = len * 2 + )); + } + body.push_str("endbfchar"); + cmap(&body) +} + +/// The shortest AGL name of `ch`, else `uniXXXX`. +pub fn glyph_name(ch: char) -> String { + agl::names_for(ch) + .first() + .map(|n| n.to_string()) + .unwrap_or_else(|| format!("uni{:04X}", u32::from(ch))) +} + +/// A TrueType program with one glyph per char of `drawn` (outlined) and `blank` (empty glyf +/// entry), each mapped in the (3,1) cmap and named in `post`. Returns the program and the GIDs. +pub fn latin_truetype(drawn: &str, blank: &str) -> (Vec, Vec<(char, u16)>) { + let mut t = TtfBuilder::new(); + let mut gids = Vec::new(); + for (ch, outline) in drawn + .chars() + .map(|c| (c, true)) + .chain(blank.chars().map(|c| (c, false))) + { + let gid = t.unicode_glyph(ch, &glyph_name(ch), outline); + gids.push((ch, gid)); + } + (t.build(), gids) +} + +/// A one-page document whose page `/Resources /Font` maps each name to its font object. +pub fn page_with_fonts(mut b: PdfBuilder, fonts: &[(&str, u32)]) -> Vec { + let catalog = b.alloc(); + let pages = b.alloc(); + let content = b.add_stream("", b"BT ET"); + let entries: String = fonts + .iter() + .map(|(name, id)| format!("/{name} {id} 0 R ")) + .collect(); + let page = b.add(format!( + "<< /Type /Page /Parent {pages} 0 R /MediaBox [0 0 612 792] /Contents {content} 0 R \ + /Resources << /Font << {entries}>> >> >>" + )); + b.set( + pages, + format!("<< /Type /Pages /Kids [{page} 0 R] /Count 1 >>"), + ); + b.set(catalog, format!("<< /Type /Catalog /Pages {pages} 0 R >>")); + b.build(&format!("/Root {catalog} 0 R")) +} + +/// Reads `pdf` with the production snapshot reader. +pub fn snapshot(pdf: Vec) -> SourceSnapshot { + snapshot_from_bytes(std::path::Path::new("fonts-fixture.pdf"), pdf, None) + .unwrap_or_else(|e| panic!("fixture must open: {e}")) +} + +/// The page's fonts, loaded through one `FontCache` with a fresh page budget, in resource order. +pub fn load_page_fonts(snap: &SourceSnapshot) -> Vec<(Vec, Arc)> { + let cache = FontCache::new(); + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + let page = snap + .doc + .get_object(snap.pages[0]) + .and_then(Object::as_dict) + .expect("page"); + let resources = page + .get(b"Resources") + .and_then(Object::as_dict) + .expect("resources"); + let fonts = resources + .get(b"Font") + .and_then(Object::as_dict) + .expect("fonts"); + fonts + .iter() + .map(|(name, obj)| { + let id = obj.as_reference().expect("font reference"); + let dict = snap + .doc + .get_object(id) + .and_then(Object::as_dict) + .expect("font dict"); + let model = cache.get_or_load(&snap.doc, FontKey::Indirect(id), dict, &mut budget); + (name.clone(), model) + }) + .collect() +} + +/// Builds a one-font page (`/F1`) and returns the loaded model. +pub fn load_one(b: PdfBuilder, font: u32) -> Arc { + let snap = snapshot(page_with_fonts(b, &[("F1", font)])); + load_page_fonts(&snap).remove(0).1 +} + +/// A simple font loaded on its own. +pub fn load_simple(f: &SimpleFont) -> Arc { + let mut b = PdfBuilder::new(); + let id = add_simple(&mut b, f); + load_one(b, id) +} + +/// A Type0 font loaded on its own. +pub fn load_type0(f: &Type0Font) -> Arc { + let mut b = PdfBuilder::new(); + let id = add_type0(&mut b, f); + load_one(b, id) +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/mod.rs b/src-tauri/src/pdf_engine/text_edit/testkit/mod.rs new file mode 100644 index 0000000..0bbdbcc --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/mod.rs @@ -0,0 +1,223 @@ +//! Test helpers (T1): scratch directories, the engines-or-skip policy (§E.1), fixture PDFs +//! (`pdf.rs`) and allocation counting for the peak-memory tests. Test builds only. +#![allow(dead_code)] // shared by the test modules of every task; not every helper is used by each + +pub(crate) mod cff; +pub(crate) mod fakes; +pub(crate) mod fonts; +pub(crate) mod pdf; +pub(crate) mod producers; +pub(crate) mod ttf; +pub(crate) mod type1; + +use crate::pdf_engine::qpdf; +use crate::pdf_engine::text_edit::engines::Engines; +use std::alloc::{GlobalAlloc, Layout, System}; +use std::cell::Cell; +use std::path::{Path, PathBuf}; +use std::sync::atomic::{AtomicBool, AtomicIsize, AtomicUsize, Ordering}; + +/// A temporary directory removed on drop. +pub struct Scratch { + dir: PathBuf, +} + +static SCRATCH_SEQ: AtomicUsize = AtomicUsize::new(0); + +impl Scratch { + pub fn new(tag: &str) -> Scratch { + let n = SCRATCH_SEQ.fetch_add(1, Ordering::SeqCst); + let dir = std::env::temp_dir() + .join("offpdf-text-edit-tests") + .join(format!("{tag}-{}-{n}", std::process::id())); + let _ = std::fs::remove_dir_all(&dir); + std::fs::create_dir_all(&dir).expect("create scratch dir"); + Scratch { dir } + } + + pub fn dir(&self) -> &Path { + &self.dir + } + + pub fn path(&self, name: &str) -> PathBuf { + self.dir.join(name) + } + + pub fn write(&self, name: &str, bytes: &[u8]) -> PathBuf { + let p = self.path(name); + std::fs::write(&p, bytes).expect("write scratch file"); + p + } +} + +impl Drop for Scratch { + fn drop(&mut self) { + let _ = std::fs::remove_dir_all(&self.dir); + } +} + +/// qpdf ≥ 11, pdftoppm and pdftotext, or `None` after printing `skip: {tool} not available`. +/// With `OFFPDF_REQUIRE_ENGINES=1` a missing engine fails the test instead (§E.1). +pub fn engines_or_skip(test_name: &str) -> Option { + match Engines::for_export(&qpdf::resolve_qpdf_standalone(), None) { + Ok(engines) => Some(engines), + Err(e) => { + let tool = match e.code.as_str() { + "VERIFIER_MISSING" => e.details.clone().unwrap_or_else(|| "poppler".into()), + _ => "qpdf 11+".to_string(), + }; + if std::env::var("OFFPDF_REQUIRE_ENGINES").as_deref() == Ok("1") { + panic!("{test_name}: {tool} not available and OFFPDF_REQUIRE_ENGINES=1 ({e})"); + } + println!("skip: {tool} not available ({test_name})"); + None + } + } +} + +// ---- Allocation counting ----------------------------------------------------------------- +// +// A thin wrapper around the system allocator. Per-thread counting (always cheap: one TLS read) +// measures code that allocates on the calling thread (the capped inflater). Process-wide +// counting is only meaningful in a child test process that runs a single test +// (`run_child_test`), because lopdf parses on rayon worker threads. + +pub struct CountingAlloc; + +static GLOBAL_ON: AtomicBool = AtomicBool::new(false); +static GLOBAL_CUR: AtomicIsize = AtomicIsize::new(0); +static GLOBAL_PEAK: AtomicIsize = AtomicIsize::new(0); + +thread_local! { + static THREAD_TRACK: Cell<(bool, isize, isize)> = const { Cell::new((false, 0, 0)) }; +} + +fn record(delta: isize) { + if GLOBAL_ON.load(Ordering::Relaxed) { + let cur = GLOBAL_CUR.fetch_add(delta, Ordering::Relaxed) + delta; + GLOBAL_PEAK.fetch_max(cur, Ordering::Relaxed); + } + let _ = THREAD_TRACK.try_with(|t| { + let (on, cur, peak) = t.get(); + if on { + let cur = cur + delta; + t.set((on, cur, peak.max(cur))); + } + }); +} + +// SAFETY: every call is forwarded unchanged to `System`; only byte counters are updated. +unsafe impl GlobalAlloc for CountingAlloc { + unsafe fn alloc(&self, layout: Layout) -> *mut u8 { + // SAFETY: forwarded with the caller's layout. + let p = unsafe { System.alloc(layout) }; + if !p.is_null() { + record(layout.size() as isize); + } + p + } + + unsafe fn alloc_zeroed(&self, layout: Layout) -> *mut u8 { + // SAFETY: forwarded with the caller's layout. + let p = unsafe { System.alloc_zeroed(layout) }; + if !p.is_null() { + record(layout.size() as isize); + } + p + } + + unsafe fn dealloc(&self, ptr: *mut u8, layout: Layout) { + // SAFETY: `ptr` was allocated by `System` with this layout. + unsafe { System.dealloc(ptr, layout) }; + record(-(layout.size() as isize)); + } + + /// Counted as a new buffer, then the old one freed: a moving realloc holds both while it + /// copies, and a growing vector's peak is that overlap (review T3-budget LOW-1). + unsafe fn realloc(&self, ptr: *mut u8, layout: Layout, new_size: usize) -> *mut u8 { + // SAFETY: `ptr` was allocated by `System` with this layout. + let p = unsafe { System.realloc(ptr, layout, new_size) }; + if !p.is_null() { + record(new_size as isize); + record(-(layout.size() as isize)); + } + p + } +} + +#[global_allocator] +static COUNTING_ALLOC: CountingAlloc = CountingAlloc; + +/// Runs `f` and returns its result with the peak number of bytes it held allocated at once on +/// the calling thread (allocations made before `f` are not counted). +pub fn thread_peak(f: impl FnOnce() -> R) -> (R, usize) { + THREAD_TRACK.with(|t| t.set((true, 0, 0))); + let r = f(); + let (_, _, peak) = THREAD_TRACK.with(|t| t.get()); + THREAD_TRACK.with(|t| t.set((false, 0, 0))); + (r, peak.max(0) as usize) +} + +/// `thread_peak`, plus the bytes `f` left allocated on the calling thread when it returned +/// (what its result holds). +pub fn thread_peak_held(f: impl FnOnce() -> R) -> (R, usize, usize) { + THREAD_TRACK.with(|t| t.set((true, 0, 0))); + let r = f(); + let (_, cur, peak) = THREAD_TRACK.with(|t| t.get()); + THREAD_TRACK.with(|t| t.set((false, 0, 0))); + (r, peak.max(0) as usize, cur.max(0) as usize) +} + +/// Process-wide peak of bytes held during `f` (use only inside `run_child_test` children). +pub fn process_peak(f: impl FnOnce() -> R) -> (R, usize) { + GLOBAL_CUR.store(0, Ordering::SeqCst); + GLOBAL_PEAK.store(0, Ordering::SeqCst); + GLOBAL_ON.store(true, Ordering::SeqCst); + let r = f(); + GLOBAL_ON.store(false, Ordering::SeqCst); + (r, GLOBAL_PEAK.load(Ordering::SeqCst).max(0) as usize) +} + +/// Env var that turns an `#[ignore]`d child test into a real run. +pub const CHILD_ENV: &str = "OFFPDF_TEXT_EDIT_CHILD"; + +/// Runs one ignored test of this test binary in a fresh process (so process-wide counters see +/// only that test) and returns its stdout. `test_path` is the full test path. +pub fn run_child_test(test_path: &str, mode: &str) -> String { + let exe = std::env::current_exe().expect("test binary path"); + let out = std::process::Command::new(exe) + .args([ + "--exact", + test_path, + "--ignored", + "--nocapture", + "--test-threads=1", + ]) + .env(CHILD_ENV, mode) + .output() + .expect("spawn child test"); + let stdout = String::from_utf8_lossy(&out.stdout).into_owned(); + assert!( + out.status.success(), + "child test {test_path} failed:\n{stdout}\n{}", + String::from_utf8_lossy(&out.stderr) + ); + stdout +} + +/// The child mode (`None` in normal runs, where child tests return immediately). +pub fn child_mode() -> Option { + std::env::var(CHILD_ENV).ok() +} + +/// Parses `KEY=` from child output. +pub fn child_value(stdout: &str, key: &str) -> usize { + stdout + .lines() + .flat_map(str::split_whitespace) + .find_map(|t| { + t.strip_prefix(&format!("{key}=")) + .and_then(|v| v.parse().ok()) + }) + .unwrap_or_else(|| panic!("{key} missing in child output:\n{stdout}")) +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/pdf.rs b/src-tauri/src/pdf_engine/text_edit/testkit/pdf.rs new file mode 100644 index 0000000..7e88721 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/pdf.rs @@ -0,0 +1,382 @@ +//! Fixture PDF builder (T1): raw objects with exact control over bytes, written with a classic +//! xref table, an xref stream (optionally with an object stream) or as a hybrid-reference file. +//! Fixtures are generated in temp by test code only (§E.1). + +use std::collections::BTreeMap; +use std::io::Write; + +/// Zlib-wrapped Flate of `data`. +pub fn zlib(data: &[u8]) -> Vec { + let mut e = flate2::write::ZlibEncoder::new(Vec::new(), flate2::Compression::default()); + e.write_all(data).expect("zlib write"); + e.finish().expect("zlib finish") +} + +/// A zlib stream that inflates to `mib` MiB of zero bytes, built from one sync-flushed 1 MiB +/// chunk repeated (≈1 KiB of input per MiB of output). Its Adler-32 is wrong on purpose. +pub fn zlib_zero_bomb(mib: usize) -> Vec { + let zeros = vec![0u8; 1 << 20]; + let mut e = flate2::write::ZlibEncoder::new(Vec::new(), flate2::Compression::best()); + e.write_all(&zeros).expect("bomb write"); + e.flush().expect("bomb flush"); + let first = e.get_ref().len(); + e.write_all(&zeros).expect("bomb write"); + e.flush().expect("bomb flush"); + let second = e.get_ref().len(); + let full = e.finish().expect("bomb finish"); + let chunk = full[first..second].to_vec(); + let mut out = full[..second].to_vec(); + for _ in 2..mib.max(2) { + out.extend_from_slice(&chunk); + } + out.extend_from_slice(&full[second..]); + out +} + +/// How the cross-reference data is written. +#[derive(Debug, Clone)] +pub enum XrefStyle { + /// Classic `xref` table. + Table, + /// An xref stream; `in_objstm` objects go into one object stream. `bomb_mib` replaces the + /// xref stream data with a Flate bomb of that many MiB. + Stream { + in_objstm: Vec, + bomb_mib: Option, + }, + /// Classic table plus `/XRefStm` in the same (newest) trailer, no `/Prev`. `in_objstm` + /// objects are listed only in the xref stream; the object stream itself is also listed in the + /// classic table when `objstm_in_classic` (Word style). + Hybrid { + in_objstm: Vec, + objstm_in_classic: bool, + }, +} + +#[derive(Debug, Clone, Default)] +pub struct PdfBuilder { + objects: BTreeMap>, + next: u32, +} + +impl PdfBuilder { + pub fn new() -> Self { + PdfBuilder { + objects: BTreeMap::new(), + next: 1, + } + } + + pub fn alloc(&mut self) -> u32 { + let id = self.next; + self.next += 1; + id + } + + pub fn set(&mut self, id: u32, body: impl AsRef<[u8]>) { + self.next = self.next.max(id + 1); + self.objects.insert(id, body.as_ref().to_vec()); + } + + pub fn add(&mut self, body: impl AsRef<[u8]>) -> u32 { + let id = self.alloc(); + self.set(id, body); + id + } + + pub fn stream_body(dict: &str, data: &[u8]) -> Vec { + let mut body = format!("<< {dict} /Length {} >>\nstream\n", data.len()).into_bytes(); + body.extend_from_slice(data); + body.extend_from_slice(b"\nendstream"); + body + } + + pub fn set_stream(&mut self, id: u32, dict: &str, data: &[u8]) { + self.set(id, Self::stream_body(dict, data)); + } + + pub fn add_stream(&mut self, dict: &str, data: &[u8]) -> u32 { + self.add(Self::stream_body(dict, data)) + } + + pub fn add_flate(&mut self, dict: &str, data: &[u8]) -> u32 { + self.add_stream(&format!("/Filter /FlateDecode {dict}"), &zlib(data)) + } + + pub fn body(&self, id: u32) -> Option<&[u8]> { + self.objects.get(&id).map(Vec::as_slice) + } + + fn header() -> Vec { + b"%PDF-1.7\n%\xE2\xE3\xCF\xD3\n".to_vec() + } + + fn write_obj(out: &mut Vec, id: u32, body: &[u8]) -> usize { + let off = out.len(); + out.extend_from_slice(format!("{id} 0 obj\n").as_bytes()); + out.extend_from_slice(body); + out.extend_from_slice(b"\nendobj\n"); + off + } + + fn table(entries: &BTreeMap) -> Vec { + // subsections of consecutive ids, plus the free head entry 0 + let mut ids: Vec = entries.keys().copied().collect(); + ids.insert(0, 0); + let mut out = b"xref\n".to_vec(); + let mut i = 0; + while i < ids.len() { + let mut j = i; + while j + 1 < ids.len() && ids[j + 1] == ids[j] + 1 { + j += 1; + } + out.extend_from_slice(format!("{} {}\n", ids[i], j - i + 1).as_bytes()); + for id in &ids[i..=j] { + match entries.get(id) { + Some(off) if *id != 0 => { + out.extend_from_slice(format!("{off:010} 00000 n \n").as_bytes()) + } + _ => out.extend_from_slice(b"0000000000 65535 f \n"), + } + } + i = j + 1; + } + out + } + + /// Writes the file. `trailer` holds extra trailer entries, e.g. `/Root 1 0 R`. + pub fn build(&self, trailer: &str) -> Vec { + self.build_with(trailer, &XrefStyle::Table) + } + + pub fn build_with(&self, trailer: &str, style: &XrefStyle) -> Vec { + let mut out = Self::header(); + let max = self.objects.keys().max().copied().unwrap_or(0); + match style { + XrefStyle::Table => { + let mut offs = BTreeMap::new(); + for (id, body) in &self.objects { + offs.insert(*id, Self::write_obj(&mut out, *id, body)); + } + let x = out.len(); + out.extend_from_slice(&Self::table(&offs)); + out.extend_from_slice( + format!( + "trailer\n<< /Size {} {trailer} >>\nstartxref\n{x}\n%%EOF\n", + max + 1 + ) + .as_bytes(), + ); + } + XrefStyle::Stream { + in_objstm, + bomb_mib, + } => { + let (offs, objstm) = self.write_with_objstm(&mut out, in_objstm, max + 1); + let xref_id = max + 2; + let x = out.len(); + let data = Self::xref_rows( + &offs, + in_objstm, + objstm.map(|_| max + 1), + xref_id, + x, + max + 3, + ); + let (dict, payload) = match bomb_mib { + Some(mib) => ("/Filter /FlateDecode".to_string(), zlib_zero_bomb(*mib)), + None => (String::new(), data), + }; + let body = Self::stream_body( + &format!("/Type /XRef /Size {} /W [1 4 2] {dict} {trailer}", max + 3), + &payload, + ); + Self::write_obj(&mut out, xref_id, &body); + out.extend_from_slice(format!("startxref\n{x}\n%%EOF\n").as_bytes()); + } + XrefStyle::Hybrid { + in_objstm, + objstm_in_classic, + } => { + let (offs, objstm) = self.write_with_objstm(&mut out, in_objstm, max + 1); + let xref_id = max + 2; + let xs = out.len(); + let data = Self::xref_rows( + &offs, + in_objstm, + objstm.map(|_| max + 1), + xref_id, + xs, + max + 3, + ); + let body = + Self::stream_body(&format!("/Type /XRef /Size {} /W [1 4 2]", max + 3), &data); + Self::write_obj(&mut out, xref_id, &body); + let mut classic: BTreeMap = offs + .iter() + .filter(|(id, _)| !in_objstm.contains(id)) + .map(|(a, b)| (*a, *b)) + .collect(); + if !objstm_in_classic { + if let Some(id) = objstm.map(|_| max + 1) { + classic.remove(&id); + } + } + let x = out.len(); + out.extend_from_slice(&Self::table(&classic)); + out.extend_from_slice( + format!( + "trailer\n<< /Size {} {trailer} /XRefStm {xs} >>\nstartxref\n{x}\n%%EOF\n", + max + 3 + ) + .as_bytes(), + ); + } + } + out + } + + /// Writes every object not in `in_objstm`, then (if any) one object stream `objstm_id` + /// holding the others. Returns the offsets of top-level objects (incl. the object stream). + fn write_with_objstm( + &self, + out: &mut Vec, + in_objstm: &[u32], + objstm_id: u32, + ) -> (BTreeMap, Option) { + let mut offs = BTreeMap::new(); + for (id, body) in &self.objects { + if !in_objstm.contains(id) { + offs.insert(*id, Self::write_obj(out, *id, body)); + } + } + if in_objstm.is_empty() { + return (offs, None); + } + let mut header = String::new(); + let mut objs = Vec::new(); + for id in in_objstm { + header.push_str(&format!("{id} {} ", objs.len())); + objs.extend_from_slice(self.objects.get(id).expect("objstm member")); + objs.push(b'\n'); + } + let mut data = header.clone().into_bytes(); + data.extend_from_slice(&objs); + let body = Self::stream_body( + &format!( + "/Type /ObjStm /N {} /First {} /Filter /FlateDecode", + in_objstm.len(), + header.len() + ), + &zlib(&data), + ); + offs.insert(objstm_id, Self::write_obj(out, objstm_id, &body)); + (offs, Some(objstm_id)) + } + + fn xref_rows( + offs: &BTreeMap, + in_objstm: &[u32], + objstm: Option, + xref_id: u32, + xref_off: usize, + size: u32, + ) -> Vec { + let mut data = Vec::new(); + for id in 0..size { + let (t, f2, f3): (u8, u32, u16) = if id == xref_id { + (1, xref_off as u32, 0) + } else if let Some(pos) = in_objstm.iter().position(|x| *x == id) { + (2, objstm.expect("object stream id"), pos as u16) + } else if let Some(off) = offs.get(&id) { + (1, *off as u32, 0) + } else { + (0, 0, if id == 0 { 65535 } else { 0 }) + }; + data.push(t); + data.extend_from_slice(&f2.to_be_bytes()); + data.extend_from_slice(&f3.to_be_bytes()); + } + data + } +} + +/// A document with a shared Helvetica font on the `/Pages` node. Each page lists its content +/// parts; one part is written as `/Contents N 0 R`, several as a direct array. +pub struct Doc { + pub b: PdfBuilder, + pub catalog: u32, + pub pages: u32, + pub page_ids: Vec, + pub content_ids: Vec>, +} + +impl Doc { + pub fn new(pages: &[&[&[u8]]]) -> Doc { + let mut b = PdfBuilder::new(); + let catalog = b.alloc(); + let pages_id = b.alloc(); + let font = b.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>", + ); + let mut page_ids = Vec::new(); + let mut content_ids = Vec::new(); + for parts in pages { + let ids: Vec = parts.iter().map(|p| b.add_stream("", p)).collect(); + let contents = match ids.as_slice() { + [one] => format!("{one} 0 R"), + many => format!( + "[{}]", + many.iter() + .map(|i| format!("{i} 0 R")) + .collect::>() + .join(" ") + ), + }; + page_ids.push(b.add(format!( + "<< /Type /Page /Parent {pages_id} 0 R /Contents {contents} >>" + ))); + content_ids.push(ids); + } + let kids = page_ids + .iter() + .map(|i| format!("{i} 0 R")) + .collect::>() + .join(" "); + b.set( + pages_id, + format!( + "<< /Type /Pages /Kids [{kids}] /Count {} /MediaBox [0 0 612 792] /Resources << /Font << /F1 {font} 0 R >> >> >>", + page_ids.len() + ), + ); + b.set( + catalog, + format!("<< /Type /Catalog /Pages {pages_id} 0 R >>"), + ); + Doc { + b, + catalog, + pages: pages_id, + page_ids, + content_ids, + } + } + + pub fn trailer(&self) -> String { + format!("/Root {} 0 R", self.catalog) + } + + pub fn build(&self) -> Vec { + self.b.build(&self.trailer()) + } + + pub fn build_with(&self, style: &XrefStyle) -> Vec { + self.b.build_with(&self.trailer(), style) + } +} + +/// One page, one content stream. +pub fn simple_pdf(content: &[u8]) -> Vec { + Doc::new(&[&[content]]).build() +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/producers.rs b/src-tauri/src/pdf_engine/text_edit/testkit/producers.rs new file mode 100644 index 0000000..0497d56 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/producers.rs @@ -0,0 +1,224 @@ +//! Producer-shaped fixtures (SPEC §E.2, T3): synthetic PDFs built in memory with `PdfBuilder` +//! and the T2 font builders, shaped like what Word, LibreOffice, Chrome/Skia, Quartz, pdfTeX, +//! XeTeX, InDesign, report generators, scanners and imposition tools write (§A.9), plus the +//! geometry, syntax and file-level edges the walker, the runs and the classifier must handle. +//! Nothing here is committed (§E.1); every builder has a smoke test in +//! `tests_walk/producers.rs`. + +mod edges; +mod files; +mod office; +mod tagging; + +pub use edges::*; +pub use files::*; +pub use office::*; +pub use tagging::*; + +use super::pdf::{PdfBuilder, XrefStyle}; + +/// Standard-14 Helvetica with WinAnsi: the font of the edge fixtures (`/F1`). +pub const HELVETICA: &str = + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>"; + +/// One page: its content parts, the inside of its `/Resources` dictionary, its `/MediaBox` +/// (`None` = none on the page) and extra page entries. +#[derive(Debug, Clone)] +pub struct PageSpec { + pub parts: Vec>, + pub resources: String, + pub media: Option, + pub extra: String, +} + +impl PageSpec { + pub fn new(content: &[u8], resources: &str) -> PageSpec { + PageSpec { + parts: vec![content.to_vec()], + resources: resources.to_string(), + media: Some("[0 0 612 792]".to_string()), + extra: String::new(), + } + } + + pub fn parts(parts: &[&[u8]], resources: &str) -> PageSpec { + PageSpec { + parts: parts.iter().map(|p| p.to_vec()).collect(), + ..PageSpec::new(b"", resources) + } + } + + /// Adds page dictionary entries, e.g. `/Rotate 90`. + pub fn with(mut self, extra: &str) -> PageSpec { + self.extra.push(' '); + self.extra.push_str(extra); + self + } + + pub fn media(mut self, media: Option<&str>) -> PageSpec { + self.media = media.map(str::to_string); + self + } +} + +/// A document under construction: catalog and `/Pages` ids are reserved up front, so fixtures +/// can reference pages from other objects (structure trees, links) before they are written. +#[derive(Debug, Clone)] +pub struct DocBuilder { + pub b: PdfBuilder, + pub catalog: u32, + pub pages: u32, + pub page_ids: Vec, + pub content_ids: Vec>, + /// The `/Kids` of the root `/Pages`, in order. + pub kids: Vec, + pub catalog_extra: String, + pub pages_extra: String, +} + +impl Default for DocBuilder { + fn default() -> Self { + Self::new() + } +} + +impl DocBuilder { + pub fn new() -> DocBuilder { + let mut b = PdfBuilder::new(); + let catalog = b.alloc(); + let pages = b.alloc(); + DocBuilder { + b, + catalog, + pages, + page_ids: Vec::new(), + content_ids: Vec::new(), + kids: Vec::new(), + catalog_extra: String::new(), + pages_extra: String::new(), + } + } + + pub fn add(&mut self, body: impl AsRef<[u8]>) -> u32 { + self.b.add(body) + } + + /// An object id for a page written later with `page_at`. + pub fn reserve(&mut self) -> u32 { + self.b.alloc() + } + + /// Adds a page (its parts as unfiltered streams); returns its id. + pub fn page(&mut self, spec: PageSpec) -> u32 { + let id = self.reserve(); + self.page_at(id, spec) + } + + /// Writes a page at a reserved id. + pub fn page_at(&mut self, id: u32, spec: PageSpec) -> u32 { + let ids: Vec = spec + .parts + .iter() + .map(|p| self.b.add_stream("", p)) + .collect(); + let contents = refs(&ids); + self.content_ids.push(ids); + self.page_raw_at(id, &contents, &spec) + } + + /// A page whose `/Contents` is written as given (shared streams, filtered parts). + pub fn page_raw(&mut self, contents: &str, spec: &PageSpec) -> u32 { + let id = self.reserve(); + self.content_ids.push(Vec::new()); + self.page_raw_at(id, contents, spec) + } + + fn page_raw_at(&mut self, id: u32, contents: &str, spec: &PageSpec) -> u32 { + let media = spec + .media + .as_ref() + .map_or(String::new(), |m| format!("/MediaBox {m}")); + self.b.set( + id, + format!( + "<< /Type /Page /Parent {} 0 R {media} /Contents {contents} /Resources << {} >> {} >>", + self.pages, spec.resources, spec.extra + ), + ); + self.page_ids.push(id); + self.kids.push(id); + id + } + + pub fn trailer(&self) -> String { + format!("/Root {} 0 R", self.catalog) + } + + fn finish(&mut self) { + let kids = self + .kids + .iter() + .map(|k| format!("{k} 0 R")) + .collect::>() + .join(" "); + self.b.set( + self.pages, + format!( + "<< /Type /Pages /Kids [{kids}] /Count {} {} >>", + self.kids.len(), + self.pages_extra + ), + ); + self.b.set( + self.catalog, + format!( + "<< /Type /Catalog /Pages {} 0 R {} >>", + self.pages, self.catalog_extra + ), + ); + } + + pub fn build(mut self) -> Vec { + self.finish(); + self.b.build(&self.trailer()) + } + + pub fn build_with(mut self, style: &XrefStyle) -> Vec { + self.finish(); + self.b.build_with(&self.trailer(), style) + } +} + +/// `N 0 R` for one id, `[a 0 R b 0 R …]` for several. +pub fn refs(ids: &[u32]) -> String { + match ids { + [one] => format!("{one} 0 R"), + many => format!( + "[{}]", + many.iter() + .map(|i| format!("{i} 0 R")) + .collect::>() + .join(" ") + ), + } +} + +/// A one-page document drawing `content` with Helvetica as `/F1`, plus extra resources and page +/// entries. +pub fn helvetica_doc(content: &[u8], extra_resources: &str, page_extra: &str) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page( + PageSpec::new( + content, + &format!("/Font << /F1 {f} 0 R >> {extra_resources}"), + ) + .with(page_extra), + ); + d.build() +} + +/// `helvetica_doc` without extras. +pub fn helvetica_page(content: &[u8]) -> Vec { + helvetica_doc(content, "", "") +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/producers/edges.rs b/src-tauri/src/pdf_engine/text_edit/testkit/producers/edges.rs new file mode 100644 index 0000000..f05e8df --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/producers/edges.rs @@ -0,0 +1,325 @@ +//! Geometry, syntax and graphics-state edges (§E.2) — mostly one Helvetica (`/F1`) page each: +//! rotation and crop, split parts, inline images with each payload proof, odd syntax, the text +//! state operators, ExtGState fonts, encodings, ActualText, shadows, clips, patterns, masks, +//! render modes, vertical and Type3 fonts, Forms and the orientation cases of §A.5. + +use super::office::{cid_font_with, word_font}; +use super::{helvetica_doc, helvetica_page, structure, DocBuilder, PageSpec, HELVETICA}; +use crate::pdf_engine::text_edit::testkit::pdf::zlib; + +fn line(text: &str) -> Vec { + format!("BT /F1 12 Tf 72 720 Td ({text}) Tj ET").into_bytes() +} + +/// A page with `/Rotate angle`; `counter_rotated` draws the text rotated against the page so it +/// reads upright on screen (§A.5 "as displayed"), else it is upright in user space. +pub fn rotated(angle: i64, counter_rotated: bool) -> Vec { + let tm = match (counter_rotated, angle.rem_euclid(360)) { + (true, 90) => "0 1 -1 0", + (true, 180) => "-1 0 0 -1", + (true, 270) => "0 -1 1 0", + _ => "1 0 0 1", + }; + let content = format!("BT /F1 12 Tf {tm} 300 400 Tm (Rotated page) Tj ET"); + helvetica_doc(content.as_bytes(), "", &format!("/Rotate {angle}")) +} + +/// A CropBox offset from the MediaBox. +pub fn cropped_offset() -> Vec { + helvetica_doc(&line("Cropped page"), "", "/CropBox [36 48 576 744]") +} + +/// `/UserUnit unit` on the page. +pub fn user_unit(unit: f64) -> Vec { + helvetica_doc(&line("Custom unit"), "", &format!("/UserUnit {unit}")) +} + +/// A text object opened in part 1 and closed in part 2 (`/Contents` array). +pub fn two_parts_mid_bt() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts( + &[ + b"BT /F1 12 Tf 72 720 Td (Hi) Tj", + b"ET\nBT /F1 12 Tf 72 680 Td (Lo) Tj ET", + ], + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +/// A show op whose operand ends part 1 and whose operator starts part 2. +pub fn straddling_op() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts( + &[ + b"BT /F1 12 Tf 72 720 Td (Hello)", + b"Tj ET BT /F1 12 Tf 72 680 Td (After) Tj ET", + ], + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +/// How an inline image's payload end is known (§A.4). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum InlineProofKind { + Length, + Unfiltered, + Flate, + Dct, +} + +/// An inline image before a text line; `Dct` (no `/L`) can only end heuristically, so the text +/// after it is refused `INLINE_IMAGE`. +pub fn inline_image(proof: InlineProofKind) -> Vec { + let head = b"q 24 0 0 12 72 400 cm BI /W 2 /H 1 /CS /RGB /BPC 8 "; + let pixels: &[u8] = &[200, 16, 16, 16, 200, 16]; + let (dict, data): (&[u8], Vec) = match proof { + InlineProofKind::Length => (b"/L 6 ", pixels.to_vec()), + InlineProofKind::Unfiltered => (b"", pixels.to_vec()), + InlineProofKind::Flate => (b"/F /Fl ", zlib(pixels)), + InlineProofKind::Dct => (b"/F /DCT ", b"\xff\xd8\xff\xe0 fake jpeg \xff\xd9".to_vec()), + }; + let mut content = head.to_vec(); + content.extend_from_slice(dict); + content.extend_from_slice(b"ID "); + content.extend_from_slice(&data); + content.extend_from_slice(b" EI Q BT /F1 12 Tf 72 720 Td (After the image) Tj ET"); + helvetica_page(&content) +} + +/// Comments, CR LF, tabs and form feeds between tokens. +pub fn comments_and_odd_ws() -> Vec { + helvetica_page( + b"% leading comment\r\nBT\t/F1 12 Tf\x0c72 720 Td%inline\n(Odd ws)Tj\r\nET % done", + ) +} + +/// `'` and `"` show ops. +pub fn quote_ops() -> Vec { + helvetica_page( + b"BT /F1 12 Tf 14 TL 72 720 Td (Line one) Tj (Line two) ' 1 0.5 (Line three) \" ET", + ) +} + +/// B1: `[(AB) -500 (CD)] TJ (tail) Tj`. +pub fn kerned_then_tail() -> Vec { + helvetica_page(b"BT /F1 10 Tf 72 700 Td [(AB) -500 (CD)] TJ (tail) Tj ET") +} + +/// B2: `/F1 1 Tf 12 0 0 12 Tm`. +pub fn tf1_tm12() -> Vec { + helvetica_page(b"BT /F1 1 Tf 12 0 0 12 72 720 Tm (Hi) Tj ET") +} + +/// `80 Tz 1 Tc 2 Tw 3 Ts`. +pub fn tz_tc_tw_ts() -> Vec { + helvetica_page(b"BT /F1 12 Tf 80 Tz 1 Tc 2 Tw 3 Ts 72 720 Td (a b) Tj ET") +} + +/// The font comes from an ExtGState `/Font` entry (no `Tf`). +pub fn extgstate_font() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::new( + b"/GS1 gs BT 72 720 Td (Hi there) Tj ET", + &format!("/ExtGState << /GS1 << /Type /ExtGState /Font [{f} 0 R 12] >> >>"), + )); + d.build() +} + +/// Times-Roman with `/MacRomanEncoding` ("café", é = 0x8E). +pub fn mac_roman() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Times-Roman /Encoding /MacRomanEncoding >>", + ); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 720 Td (caf\\216) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +/// B3: a Word subset without "Y" (its `/Widths` slot is 0 and no glyph exists). +pub fn subset_without_y() -> Vec { + let mut d = DocBuilder::new(); + let f = word_font(&mut d.b, "ABCDEF+Calibri", "Hello"); + d.page(PageSpec::new( + &line("Hello"), + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +/// `/Span <> BDC` around one line, a plain line after it. +pub fn actual_text_span() -> Vec { + helvetica_page( + b"/Span <> BDC BT /F1 12 Tf 72 720 Td (fi) Tj ET EMC \ + BT /F1 12 Tf 72 700 Td (Plain) Tj ET", + ) +} + +/// A tagged page whose structure element for MCID 0 carries `/ActualText`. +pub fn actual_text_struct() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let page = d.reserve(); + let extra = structure(&mut d, page, &[0, 1], &[0]); + d.page_at( + page, + PageSpec::new( + b"/P <> BDC BT /F1 12 Tf 72 720 Td (Struct) Tj ET EMC \ + /P <> BDC BT /F1 12 Tf 72 700 Td (Free) Tj ET EMC", + &format!("/Font << /F1 {f} 0 R >>"), + ) + .with(&extra), + ); + d.build() +} + +/// The same text drawn twice, half a point apart (a shadow). +pub fn duplicate_shadow() -> Vec { + helvetica_page( + b"BT /F1 12 Tf 72 720 Td (Shadow) Tj ET BT /F1 12 Tf 72.5 719.5 Td (Shadow) Tj ET", + ) +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum ClipKind { + /// Word's page-sized `re W n`. + Page, + /// A rectangle that cuts the text. + Small, + /// A Bézier path clip. + Curve, +} + +/// Text inside a clip. +pub fn clip(kind: ClipKind) -> Vec { + let path = match kind { + ClipKind::Page => "0 0 612 792 re", + ClipKind::Small => "72 715 6 6 re", + ClipKind::Curve => "0 0 m 300 900 600 900 612 0 c h", + }; + let content = format!("q {path} W n BT /F1 12 Tf 72 720 Td (Clip me) Tj ET Q"); + helvetica_page(content.as_bytes()) +} + +/// A Helvetica page with a `/Cs1` Pattern colour space and a `/P1` tiling pattern. +pub fn pattern_doc(content: &[u8]) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let pat = d.b.add_stream( + "/Type /Pattern /PatternType 1 /PaintType 1 /TilingType 1 /BBox [0 0 10 10] /XStep 10 \ + /YStep 10 /Resources << >>", + b"0 0 10 10 re f", + ); + d.page(PageSpec::new( + content, + &format!( + "/Font << /F1 {f} 0 R >> /ColorSpace << /Cs1 /Pattern >> /Pattern << /P1 {pat} 0 R >>" + ), + )); + d.build() +} + +/// Text filled with a pattern. +pub fn pattern_fill() -> Vec { + pattern_doc(b"BT /F1 12 Tf /Cs1 cs /P1 scn 72 720 Td (Pattern) Tj ET") +} + +/// Text under an ExtGState soft mask. +pub fn smask_text() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let group = d.b.add_stream( + "/Type /XObject /Subtype /Form /BBox [0 0 612 792] /Group << /S /Transparency /CS /DeviceGray >>", + b"0.5 g 0 0 612 792 re f", + ); + d.page(PageSpec::new( + b"/GS1 gs BT /F1 12 Tf 72 720 Td (Masked) Tj ET", + &format!( + "/Font << /F1 {f} 0 R >> /ExtGState << /GS1 << /SMask << /Type /Mask /S /Luminosity /G {group} 0 R >> >> >>" + ), + )); + d.build() +} + +/// `7 Tr` (clip only) text, then a normal line drawn inside that text clip. +pub fn tr7() -> Vec { + helvetica_page( + b"BT 7 Tr /F1 12 Tf 72 720 Td (Clip text) Tj ET BT 0 Tr /F1 12 Tf 72 700 Td (After clip) Tj ET", + ) +} + +/// A Type0 `/Identity-V` font (vertical writing). +pub fn identity_v() -> Vec { + let mut d = DocBuilder::new(); + let f = cid_font_with(&mut d.b, "ABCDEF+Vertical", "Up", "/Identity-V"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 300 700 Td <00010002> Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +/// A Type3 font. +pub fn type3() -> Vec { + let mut d = DocBuilder::new(); + let proc_id = d.b.add_stream("", b"10 0 0 0 10 10 d1 0 0 10 10 re f"); + let f = d.add(format!( + "<< /Type /Font /Subtype /Type3 /FontBBox [0 0 10 10] /FontMatrix [0.1 0 0 0.1 0 0] \ + /CharProcs << /a {proc_id} 0 R >> /Encoding << /Type /Encoding /Differences [97 /a] >> \ + /FirstChar 97 /LastChar 97 /Widths [10] >>" + )); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 720 Td (a) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +/// Text inside a Form XObject painted by the page. +pub fn nested_form() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let form = d.b.add_stream( + &format!( + "/Type /XObject /Subtype /Form /BBox [0 0 300 100] /Resources << /Font << /F1 {f} 0 R >> >>" + ), + b"BT /F1 12 Tf 10 10 Td (In a form) Tj ET", + ); + d.page(PageSpec::new( + b"q 1 0 0 1 72 600 cm /Fm0 Do Q", + &format!("/Font << /F1 {f} 0 R >> /XObject << /Fm0 {form} 0 R >>"), + )); + d.build() +} + +/// `-1 0 0 1 Tm` (mirrored). +pub fn mirrored() -> Vec { + helvetica_page(b"BT /F1 12 Tf -1 0 0 1 300 400 Tm (Mirror) Tj ET") +} + +/// `-100 Tz` (mirrored). +pub fn negative_tz() -> Vec { + helvetica_page(b"BT /F1 12 Tf -100 Tz 300 400 Td (Mirror) Tj ET") +} + +/// `/F1 -12 Tf` (turned 180°). +pub fn negative_tf() -> Vec { + helvetica_page(b"BT /F1 -12 Tf 300 400 Td (Turned) Tj ET") +} + +/// `1 0.3 0 1 Tm` (skewed baseline). +pub fn skewed() -> Vec { + helvetica_page(b"BT /F1 12 Tf 1 0.3 0 1 72 400 Tm (Skewed) Tj ET") +} + +/// `1 0 0.2 1 Tm` (synthetic italic: upright). +pub fn oblique() -> Vec { + helvetica_page(b"BT /F1 12 Tf 1 0 0.2 1 72 400 Tm (Oblique) Tj ET") +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/producers/files.rs b/src-tauri/src/pdf_engine/text_edit/testkit/producers/files.rs new file mode 100644 index 0000000..4111e2a --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/producers/files.rs @@ -0,0 +1,342 @@ +//! File-level hazards and the revision-2 fixtures of §E.2: bombs and deep nesting, signed, +//! encrypted and XFA files, hybrid xref and page-tree shapes lopdf misreads, shared inherited +//! resources, legacy filters, two-column pages, stroked text, swapped images, bad page boxes and +//! an indirect `/Extensions`. + +use super::{helvetica_page, structure, DocBuilder, PageSpec, HELVETICA}; +use crate::pdf_engine::text_edit::engines::{run_tool, Engines, RunOpts}; +use crate::pdf_engine::text_edit::testkit::pdf::{zlib_zero_bomb, XrefStyle}; +use crate::pdf_engine::text_edit::testkit::Scratch; +use std::ffi::OsString; + +const HELLO: &[u8] = b"BT /F1 12 Tf 72 720 Td (Hello) Tj ET"; + +/// A one-page Helvetica document with `extra(d)` returning catalog entries. +fn with_catalog(extra: impl FnOnce(&mut DocBuilder) -> String) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::new(HELLO, &format!("/Font << /F1 {f} 0 R >>"))); + let entries = extra(&mut d); + d.catalog_extra.push_str(&entries); + d.build() +} + +/// An unreferenced object nested `levels` deep (`FILE_TOO_COMPLEX` past 100). +pub fn deep_nesting(levels: usize) -> Vec { + with_catalog(|d| { + d.add(format!("{}{}", "[".repeat(levels), "]".repeat(levels))); + String::new() + }) +} + +/// An object stream that inflates to 1 GiB. +pub fn objstm_bomb() -> Vec { + with_catalog(|d| { + d.b.add_stream( + "/Type /ObjStm /N 1 /First 4 /Filter /FlateDecode", + &zlib_zero_bomb(1024), + ); + String::new() + }) +} + +/// An xref stream that inflates to 1 GiB. +pub fn xref_bomb() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::new(HELLO, &format!("/Font << /F1 {f} 0 R >>"))); + d.build_with(&XrefStyle::Stream { + in_objstm: vec![], + bomb_mib: Some(1024), + }) +} + +/// A page whose content stream inflates past the 32 MiB stream cap (page `PAGE_TOO_COMPLEX`), +/// followed by a normal page. +pub fn flate_bomb() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + let bomb = d.b.add_stream("/Filter /FlateDecode", &zlib_zero_bomb(64)); + d.page_raw(&format!("{bomb} 0 R"), &PageSpec::new(b"", &res)); + d.page(PageSpec::new(HELLO, &res)); + d.build() +} + +/// An applied signature (`/Type /Sig` with `/ByteRange`). +pub fn signed() -> Vec { + with_catalog(|d| { + let sig = + d.add("<< /Type /Sig /Filter /Adobe.PPKLite /ByteRange [0 10 20 30] /Contents <00> >>"); + let field = d.add(format!( + "<< /FT /Sig /T (Signature1) /V {sig} 0 R /Subtype /Widget /Rect [0 0 0 0] >>" + )); + format!("/AcroForm << /Fields [{field} 0 R] /SigFlags 3 >>") + }) +} + +/// A dynamic XFA form. +pub fn xfa() -> Vec { + with_catalog(|d| { + let x = d.b.add_stream("", b""); + format!("/AcroForm << /Fields [] /XFA {x} 0 R >>") + }) +} + +/// A Helvetica page encrypted by `qpdf --encrypt` (AES-256). +pub fn encrypted(engines: &Engines) -> Vec { + let s = Scratch::new("producers-enc"); + let plain = s.write("plain.pdf", &helvetica_page(HELLO)); + let out = s.path("enc.pdf"); + let args: Vec = vec![ + "--encrypt".into(), + "user".into(), + "owner".into(), + "256".into(), + "--".into(), + plain.into(), + out.clone().into(), + ]; + let r = run_tool(&engines.qpdf, &args, false, &RunOpts::default()).expect("qpdf --encrypt"); + assert_eq!(r.code, 0, "qpdf --encrypt: {}", r.stderr); + std::fs::read(&out).expect("encrypted output") +} + +/// The page sits in an object stream listed by a hybrid file's `/XRefStm`; Word style +/// (`listed_in_classic`) also lists the object stream in the classic table, which lopdf reads. +pub fn hybrid_xref(listed_in_classic: bool) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let page = d.page(PageSpec::new(HELLO, &format!("/Font << /F1 {f} 0 R >>"))); + d.build_with(&XrefStyle::Hybrid { + in_objstm: vec![page], + objstm_in_classic: listed_in_classic, + }) +} + +/// Two kids, the second without `/Type` (qpdf counts it, lopdf skips it). +pub fn kid_without_type() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + d.page(PageSpec::new(HELLO, &res)); + let content = d.b.add_stream("", b"BT /F1 12 Tf 72 720 Td (Second) Tj ET"); + let id = d.add(format!( + "<< /Parent {} 0 R /MediaBox [0 0 612 792] /Contents {content} 0 R /Resources << {res} >> >>", + d.pages + )); + d.kids.push(id); + d.build() +} + +/// Two pages sharing one `/Resources` on the `/Pages` node: the bold sibling `/F2` is used only +/// on page 2. +pub fn shared_inherited_resources() -> Vec { + let mut d = DocBuilder::new(); + let regular = d.add(HELVETICA); + let bold = d.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica-Bold /Encoding /WinAnsiEncoding >>", + ); + d.pages_extra = format!("/Resources << /Font << /F1 {regular} 0 R /F2 {bold} 0 R >> >>"); + for content in [ + &b"BT /F1 12 Tf 72 720 Td (Regular words) Tj ET"[..], + b"BT /F2 12 Tf 72 720 Td (Bold words) Tj ET", + ] { + let id = d.b.add_stream("", content); + let page = d.add(format!( + "<< /Type /Page /Parent {} 0 R /MediaBox [0 0 612 792] /Contents {id} 0 R >>", + d.pages + )); + d.page_ids.push(page); + d.kids.push(page); + } + d.build() +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum LegacyFilter { + RunLength, + Lzw, +} + +fn run_length(data: &[u8]) -> Vec { + let mut out = Vec::new(); + for chunk in data.chunks(128) { + out.push((chunk.len() - 1) as u8); + out.extend_from_slice(chunk); + } + out.push(128); + out +} + +/// LZW with 9-bit literal codes only: a clear code every 200 literals keeps the table (and so +/// the code width, early change 1) small; ends with EOD. +fn lzw(data: &[u8]) -> Vec { + let mut codes = vec![256u16]; + for (i, b) in data.iter().enumerate() { + if i > 0 && i % 200 == 0 { + codes.push(256); + } + codes.push(u16::from(*b)); + } + codes.push(257); + let (mut out, mut acc, mut bits) = (Vec::new(), 0u32, 0u32); + for c in codes { + acc = (acc << 9) | u32::from(c); + bits += 9; + while bits >= 8 { + bits -= 8; + out.push((acc >> bits) as u8); + } + } + if bits > 0 { + out.push((acc << (8 - bits)) as u8); + } + out +} + +/// A page whose content stream uses RunLength or LZW (refused `UNSUPPORTED_FILTER`; qpdf keeps +/// such streams raw), then a Flate-free normal page. +pub fn legacy_filter_page(filter: LegacyFilter) -> Vec { + let content = b"BT /F1 12 Tf 72 720 Td (Legacy filter) Tj ET"; + let (name, data) = match filter { + LegacyFilter::RunLength => ("RunLengthDecode", run_length(content)), + LegacyFilter::Lzw => ("LZWDecode", lzw(content)), + }; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + let id = d.b.add_stream(&format!("/Filter /{name}"), &data); + d.page_raw(&format!("{id} 0 R"), &PageSpec::new(b"", &res)); + d.page(PageSpec::new(HELLO, &res)); + d.build() +} + +/// The lines of the two-column page: (MCID, x, y, size, text), in content order (row by row, +/// with the title in the middle). +pub const TWO_COLUMN_LINES: [(i64, f64, f64, f64, &str); 7] = [ + (0, 72.0, 700.0, 12.0, "Left one"), + (1, 320.0, 700.0, 12.0, "Right one"), + (2, 72.0, 740.0, 14.0, "A two column page"), + (3, 72.0, 686.0, 12.0, "Left two"), + (4, 320.0, 686.0, 12.0, "Right two"), + (5, 72.0, 672.0, 12.0, "Left three"), + (6, 320.0, 672.0, 12.0, "Right three"), +]; + +/// The structure order of the tagged two-column page: right column, title, left column — on +/// purpose neither the content order nor the XY-cut order. +pub const TWO_COLUMN_STRUCT_ORDER: [i64; 7] = [1, 4, 6, 2, 0, 3, 5]; + +/// A full-width title over two columns, written row by row. `tagged`: every line in its own +/// marked-content sequence, structure order `TWO_COLUMN_STRUCT_ORDER`; `missing_mcid` leaves +/// the last line untagged. +pub fn two_column_with(tagged: bool, missing_mcid: bool) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let page = d.reserve(); + let extra = if tagged { + structure(&mut d, page, &TWO_COLUMN_STRUCT_ORDER, &[]) + } else { + String::new() + }; + let mut content = String::new(); + for (i, (mcid, x, y, size, text)) in TWO_COLUMN_LINES.iter().enumerate() { + let line = format!("BT /F1 {size} Tf {x} {y} Td ({text}) Tj ET"); + let untagged = !tagged || (missing_mcid && i + 1 == TWO_COLUMN_LINES.len()); + if untagged { + content.push_str(&format!("{line} ")); + } else { + content.push_str(&format!("/P <> BDC {line} EMC ")); + } + } + d.page_at( + page, + PageSpec::new(content.as_bytes(), &format!("/Font << /F1 {f} 0 R >>")).with(&extra), + ); + d.build() +} + +/// §E.2 `two_column(tagged)`. +pub fn two_column(tagged: bool) -> Vec { + two_column_with(tagged, false) +} + +/// A line drawn as two stroked segments (`tr` 1 or 2) on one baseline: "Bold" with line width +/// `widths[0]` and dash `dashes[0]`, then "face" with `widths[1]`/`dashes[1]`. +pub fn stroke_text(tr: i64, widths: [f64; 2], dashes: [&str; 2]) -> Vec { + // Helvetica "Bold" = 667 + 556 + 222 + 556 = 2001 → 24.012 pt at 12 pt. + let content = format!( + "{} w {} d BT /F1 12 Tf {tr} Tr 72 700 Td (Bold) Tj ET \ + {} w {} d BT /F1 12 Tf {tr} Tr 96.012 700 Td (face) Tj ET", + widths[0], dashes[0], widths[1], dashes[1] + ); + helvetica_page(content.as_bytes()) +} + +/// The same page with one of two image data variants behind the same name `/Im0`. +pub fn swapped_image(swapped: bool) -> Vec { + let pixels: &[u8] = if swapped { + &[16, 16, 200, 200, 16, 16, 16, 200, 16, 200, 200, 16] + } else { + &[200, 16, 16, 16, 200, 16, 16, 16, 200, 200, 200, 16] + }; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let img = d.b.add_stream( + "/Type /XObject /Subtype /Image /Width 2 /Height 2 /ColorSpace /DeviceRGB /BitsPerComponent 8", + pixels, + ); + d.page(PageSpec::new( + b"q 40 0 0 40 72 400 cm /Im0 Do Q BT /F1 12 Tf 72 720 Td (Image page) Tj ET", + &format!("/Font << /F1 {f} 0 R >> /XObject << /Im0 {img} 0 R >>"), + )); + d.build() +} + +/// Page setups the strict geometry parser refuses (`GEOMETRY`, §A.5). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum BadBox { + /// No `/MediaBox` anywhere in the page tree. + NoMediaBox, + /// `/Rotate 90.0` (a real). + RealRotate, + /// `/Rotate 45`. + OddRotate, + /// Crop ∩ Media thinner than 1 pt. + ThinCrop, + /// `/UserUnit (1)` (a string). + UserUnitString, + /// `/UserUnit -1`. + UserUnitNegative, + /// `/UserUnit 2`. + UserUnitTwo, + /// A MediaBox with a name in it. + NonNumberBox, +} + +pub fn bad_boxes(kind: BadBox) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let spec = PageSpec::new(HELLO, &format!("/Font << /F1 {f} 0 R >>")); + let spec = match kind { + BadBox::NoMediaBox => spec.media(None), + BadBox::RealRotate => spec.with("/Rotate 90.0"), + BadBox::OddRotate => spec.with("/Rotate 45"), + BadBox::ThinCrop => spec.with("/CropBox [0 0 612 0.5]"), + BadBox::UserUnitString => spec.with("/UserUnit (1)"), + BadBox::UserUnitNegative => spec.with("/UserUnit -1"), + BadBox::UserUnitTwo => spec.with("/UserUnit 2"), + BadBox::NonNumberBox => spec.media(Some("[0 0 /Wide 792]")), + }; + d.page(spec); + d.build() +} + +/// The catalog's `/Extensions` as an indirect object (qpdf writes it direct: APP-08). +pub fn extensions_indirect() -> Vec { + with_catalog(|d| { + let ext = d.add("<< /ADBE << /BaseVersion /1.7 /ExtensionLevel 3 >> >>"); + format!("/Extensions {ext} 0 R") + }) +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/producers/office.rs b/src-tauri/src/pdf_engine/text_edit/testkit/producers/office.rs new file mode 100644 index 0000000..0a936e0 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/producers/office.rs @@ -0,0 +1,509 @@ +//! Office, browser, TeX and print producers (§E.2, §A.9): FX-WORD, FX-WORD-TR, FX-LIBRE, +//! FX-LIBRE-CFF, FX-SKIA, FX-QUARTZ, FX-PERGLYPH, FX-PDFTEX, FX-XETEX, FX-INDD, FX-STD14, +//! FX-NONEMB, FX-OCR and FX-SHARED, each with the font setup and operator shapes that producer +//! writes. + +use super::{refs, tagged_bookmarked, DocBuilder, PageSpec, HELVETICA}; +use crate::pdf_engine::text_edit::testkit::cff::CffBuilder; +use crate::pdf_engine::text_edit::testkit::fonts::{ + add_simple, add_type0, glyph_name, latin_truetype, tounicode_bfchar, Program, SimpleFont, + Type0Font, +}; +use crate::pdf_engine::text_edit::testkit::pdf::PdfBuilder; +use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; +use crate::pdf_engine::text_edit::testkit::type1::{T1Encoding, T1Glyph, Type1Builder}; + +/// Distinct characters of `text`, first-seen order. +pub fn unique(text: &str) -> String { + let mut out = String::new(); + for ch in text.chars() { + if !out.contains(ch) { + out.push(ch); + } + } + out +} + +fn bfchar(entries: &[(u32, usize, String)]) -> Vec { + let refs: Vec<(u32, usize, &str)> = entries + .iter() + .map(|(c, l, t)| (*c, *l, t.as_str())) + .collect(); + tounicode_bfchar(&refs) +} + +/// Word's simple TrueType subset: WinAnsi, `/Widths` for codes 32–255 (500 for the characters of +/// `text`, 0 for every other code), outlined glyphs for the subset, an empty glyph for the space +/// and a ToUnicode for the codes used. ASCII text only. +pub fn word_font(b: &mut PdfBuilder, base: &str, text: &str) -> u32 { + let chars = unique(&format!("{text} ")); + let drawn: String = chars.chars().filter(|c| *c != ' ').collect(); + let (program, _) = latin_truetype(&drawn, " "); + let mut widths = vec![0.0; 224]; + let mut map = Vec::new(); + for ch in chars.chars() { + let code = u32::from(ch); + if let Some(slot) = code + .checked_sub(32) + .and_then(|i| widths.get_mut(i as usize)) + { + *slot = 500.0; + map.push((code, 1, ch.to_string())); + } + } + let mut f = SimpleFont::new("TrueType", base); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(widths); + f.flags = Some(32); + f.program = Program::TrueType(program); + f.tounicode = Some(bfchar(&map)); + add_simple(b, &f) +} + +/// A Type0 Identity-H CIDFontType2 subset whose CIDs 1, 2, … are the characters of `chars` +/// (GID = CID), width 500, with a ToUnicode — Word's companion for non-WinAnsi letters, and the +/// only font kind Chrome writes. +pub fn cid_font(b: &mut PdfBuilder, base: &str, chars: &str) -> u32 { + cid_font_with(b, base, chars, "/Identity-H") +} + +/// `cid_font` with another `/Encoding` (e.g. `/Identity-V`). +pub fn cid_font_with(b: &mut PdfBuilder, base: &str, chars: &str, encoding: &str) -> u32 { + let mut t = TtfBuilder::new(); + let mut map = Vec::new(); + for (i, ch) in chars.chars().enumerate() { + t.unicode_glyph(ch, &glyph_name(ch), ch != ' '); + map.push((i as u32 + 1, 2, ch.to_string())); + } + let mut f = Type0Font::new("CIDFontType2", base); + f.encoding = encoding.to_string(); + f.program = Program::TrueType(t.build()); + f.cid_to_gid = Some(None); + f.flags = Some(32); + f.w = Some(format!("[1 [{}]]", vec!["500"; map.len()].join(" "))); + f.tounicode = Some(bfchar(&map)); + add_type0(b, &f) +} + +/// The 2-byte hex codes of `text` in a `cid_font` built from `chars`. +pub fn cid_hex(chars: &str, text: &str) -> String { + text.chars() + .map(|ch| { + let cid = chars.chars().position(|c| c == ch).map_or(0, |i| i + 1); + format!("{cid:04X}") + }) + .collect() +} + +/// FX-WORD: WinAnsi TrueType zeroed subset, `/Widths` 0 for unused codes, ToUnicode, one +/// `BT … Tm [..] TJ ET` per format run (line 1 is "Inv"+12+"oice" and " 2026"), a page-sized +/// `re W n` clip, `/P <> BDC`, tagged and bookmarked. +pub fn word() -> Vec { + let mut d = DocBuilder::new(); + let f1 = word_font(&mut d.b, "ABCDEF+Calibri", "Invoice 2026 Due 7"); + let page = d.reserve(); + let extra = tagged_bookmarked(&mut d, page, &[0, 1]); + let content = "/P <> BDC q 0 0 612 792 re W n \ + BT /F1 11.04 Tf 1 0 0 1 72 700 Tm [(Inv)12(oice)]TJ ET \ + BT /F1 11.04 Tf 1 0 0 1 110.5075 700 Tm [( 2026)]TJ ET Q EMC \ + /P <> BDC BT /F1 11.04 Tf 1 0 0 1 72 680 Tm [(Due)]TJ ET EMC"; + d.page_at( + page, + PageSpec::new(content.as_bytes(), &format!("/Font << /F1 {f1} 0 R >>")).with(&extra), + ); + d.build() +} + +/// The Turkish characters of FX-WORD-TR's Identity-H companion (CIDs 1–4). +pub const WORD_TR_CID_CHARS: &str = "ğışİ"; + +/// FX-WORD-TR: FX-WORD's font plus a Type0 CIDFontType2 companion with the same BaseFont for +/// the non-WinAnsi letters; "Sağlık Bakanlığı Raporu" as seven per-font `BT … Tm` segments at +/// Word's rounded positions (one lands 0.01 pt late: a converted gap). +pub fn word_tr() -> Vec { + let mut d = DocBuilder::new(); + let f1 = word_font(&mut d.b, "ABCDEF+Calibri", "Sa lk Bakan Raporu"); + let tr = WORD_TR_CID_CHARS; + let f2 = cid_font(&mut d.b, "ABCDEF+Calibri", tr); + let segs: [(&str, String, usize); 7] = [ + ("F1", "(Sa)".into(), 2), + ("F2", format!("<{}>", cid_hex(tr, "ğ")), 1), + ("F1", "(l)".into(), 1), + ("F2", format!("<{}>", cid_hex(tr, "ı")), 1), + ("F1", "(k Bakanl)".into(), 8), + ("F2", format!("<{}>", cid_hex(tr, "ığı")), 3), + ("F1", "( Raporu)".into(), 7), + ]; + let mut x = 72.0f64; + let mut body = String::from("/P <> BDC "); + for (i, (font, shown, n)) in segs.iter().enumerate() { + let at = if i == 3 { x + 0.01 } else { x }; + body.push_str(&format!( + "BT /{font} 11 Tf 1 0 0 1 {at:.2} 700 Tm [{shown}]TJ ET " + )); + x += 5.5 * *n as f64; + } + body.push_str("EMC"); + let page = d.reserve(); + let extra = tagged_bookmarked(&mut d, page, &[0]); + d.page_at( + page, + PageSpec::new( + body.as_bytes(), + &format!("/Font << /F1 {f1} 0 R /F2 {f2} 0 R >>"), + ) + .with(&extra), + ); + d.build() +} + +/// A symbolic TrueType subset with codes 1..n for the characters of `chars`, glyphs found by +/// the (3,0) cmap at 0xF000 + code (LibreOffice) or the (1,0) cmap at the code (Quartz), no +/// `/Encoding`, and a ToUnicode. +fn symbolic_font(b: &mut PdfBuilder, base: &str, chars: &str, mac_cmap: bool) -> u32 { + let mut t = TtfBuilder::new(); + let mut map = Vec::new(); + let mut widths = vec![0.0]; + for (i, ch) in chars.chars().enumerate() { + let code = i as u32 + 1; + let gid = t.glyph(&glyph_name(ch), ch != ' ', 500); + if mac_cmap { + t.cmap10.push((code as u8, gid)); + } else { + t.cmap30.push((0xF000 + code, gid)); + } + map.push((code, 1, ch.to_string())); + widths.push(500.0); + } + let mut f = SimpleFont::new("TrueType", base); + f.flags = Some(4); + f.first_char = 0; + f.widths = Some(widths); + f.program = Program::TrueType(t.build()); + f.tounicode = Some(bfchar(&map)); + add_simple(b, &f) +} + +/// 1-byte hex codes of `text` in a `symbolic_font` built from `chars`. +fn sym_hex(chars: &str, text: &str) -> String { + text.chars() + .map(|ch| { + let code = chars.chars().position(|c| c == ch).map_or(0, |i| i + 1); + format!("{code:02X}") + }) + .collect() +} + +/// FX-LIBRE: symbolic TrueType subset, codes 1..n, `(3,0)` cmap, ToUnicode, hex `TJ` with kerns. +pub fn libre() -> Vec { + let chars = unique("Libre text"); + let mut d = DocBuilder::new(); + let f1 = symbolic_font(&mut d.b, "BAAAAA+LiberationSerif", &chars, false); + let content = format!( + "BT /F1 12 Tf 72 700 Td [<{}>-20<{}>]TJ ET BT /F1 12 Tf 72 680 Td [<{}>]TJ ET", + sym_hex(&chars, "Libre"), + sym_hex(&chars, " text"), + sym_hex(&chars, "text") + ); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +/// FX-LIBRE-CFF: a simple `FontFile3 /Type1C` program, WinAnsi, `/Widths`. +pub fn libre_cff() -> Vec { + let cff = CffBuilder::new("CAAAAA+SourceSansPro-Regular") + .glyph("H", true) + .glyph("e", true) + .glyph("l", true) + .glyph("o", true) + .glyph("space", false) + .build(); + let mut widths = vec![0.0; 80]; // codes 32..=111 + for ch in [' ', 'H', 'e', 'l', 'o'] { + if let Some(w) = widths.get_mut(u32::from(ch) as usize - 32) { + *w = 500.0; + } + } + let mut f = SimpleFont::new("Type1", "CAAAAA+SourceSansPro-Regular"); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(widths); + f.flags = Some(32); + f.program = Program::Cff(cff); + let mut d = DocBuilder::new(); + let f1 = add_simple(&mut d.b, &f); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello Hello) Tj ET", + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +/// FX-SKIA: Chrome/Docs output — a `1 0 0 -1 0 792 cm` flip with flipped `Tm`, a Type0 +/// CIDFontType2 subset, one `Tj` per text node ("Chr" + "ome"), a synthetic-bold `Tr 2` line and +/// a synthetic-italic sheared line. +pub fn skia() -> Vec { + let chars = unique("ChromeBoldItalic"); + let mut d = DocBuilder::new(); + let f1 = cid_font(&mut d.b, "AAAAAA+Arimo", &chars); + let h = |t: &str| cid_hex(&chars, t); + let content = format!( + "1 0 0 -1 0 792 cm \ + BT /F1 12 Tf 1 0 0 -1 72 100 Tm <{}> Tj ET BT /F1 12 Tf 1 0 0 -1 90 100 Tm <{}> Tj ET \ + 2 Tr 0.36 w BT /F1 12 Tf 1 0 0 -1 72 130 Tm <{}> Tj ET 0 Tr \ + BT /F1 12 Tf 1 0 0.2 -1 72 160 Tm <{}> Tj ET", + h("Chr"), + h("ome"), + h("Bold"), + h("Italic") + ); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +/// FX-QUARTZ: symbolic TrueType without `/Encoding` (glyphs by the (1,0) cmap) + ToUnicode, some +/// glyphs in ops of their own. +pub fn quartz() -> Vec { + let chars = unique("Quartz"); + let mut d = DocBuilder::new(); + let f1 = symbolic_font(&mut d.b, "CAAAAA+Helvetica-Light", &chars, true); + let content = format!( + "BT /TT1 12 Tf 1 0 0 1 72 700 Tm <{}> Tj 1 0 0 1 90 700 Tm <{}> Tj \ + 1 0 0 1 96 700 Tm <{}> Tj ET", + sym_hex(&chars, "Qua"), + sym_hex(&chars, "r"), + sym_hex(&chars, "tz") + ); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /TT1 {f1} 0 R >>"), + )); + d.build() +} + +/// FX-PERGLYPH (`gap: None`): "Per glyph line" as 14 one-glyph `Tm`+`Tj` ops in Courier placed +/// edge to edge (they join). With `gap: Some(g)` every glyph starts `g` pt after the previous one +/// (they do not join: a per-glyph page). +pub fn per_glyph(gap: Option) -> Vec { + let step = gap.unwrap_or(7.2); // Courier: 600/1000 × 12 + let mut content = String::from("BT /F1 12 Tf "); + for (i, ch) in "Per glyph line".chars().enumerate() { + let x = 72.0 + step * i as f64; + content.push_str(&format!("1 0 0 1 {x:.4} 700 Tm ({ch}) Tj ")); + } + content.push_str("ET"); + let mut d = DocBuilder::new(); + let f = + d.add("<< /Type /Font /Subtype /Type1 /BaseFont /Courier /Encoding /WinAnsiEncoding >>"); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +/// FX-PDFTEX: an embedded Type1 `FontFile` (generated: the repo's Foxit programs are CFF, see +/// DEVIATIONS [T2]) with `/Differences` incl. `fi` and `quoteright`, `/CharSet`, no ToUnicode, +/// no space glyph, and word gaps written as `-333` kerns. +pub fn pdftex() -> Vec { + let names: [(u8, &str); 11] = [ + (12, "fi"), + (39, "quoteright"), + (72, "H"), + (87, "W"), + (100, "d"), + (101, "e"), + (108, "l"), + (110, "n"), + (111, "o"), + (114, "r"), + (116, "t"), + ]; + let mut t1 = Type1Builder::new("ABCDEF+CMR10"); + t1.encoding = T1Encoding::Custom(names.iter().map(|(c, n)| (*c, n.to_string())).collect()); + for (_, name) in names { + t1 = t1.glyph(name, T1Glyph::Box { width: 500 }); + } + let mut widths = vec![0.0; 105]; // codes 12..=116 + for (code, _) in names { + if let Some(w) = widths.get_mut(usize::from(code) - 12) { + *w = 500.0; + } + } + let mut f = SimpleFont::new("Type1", "ABCDEF+CMR10"); + f.encoding = Some( + "<< /Type /Encoding /Differences [12 /fi 39 /quoteright 72 /H 87 /W 100 /d /e 108 /l \ + 110 /n /o 114 /r 116 /t] >>" + .into(), + ); + f.first_char = 12; + f.widths = Some(widths); + f.flags = Some(4); + f.descriptor_extra = "/CharSet (/H/W/d/e/fi/l/n/o/quoteright/r/t)".into(); + f.program = Program::Type1(t1.build()); + let mut d = DocBuilder::new(); + let f1 = add_simple(&mut d.b, &f); + d.page(PageSpec::new( + b"BT /F1 9.9626 Tf 72 700 Td [(Hello)-333(W)80(orld)]TJ 0 -12 Td \ + [(\\014nd)-333(don\\047t)]TJ ET", + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +/// FX-XETEX: Type0 Identity-H with a CID-keyed `FontFile3 /CIDFontType0C` program + ToUnicode. +pub fn xetex() -> Vec { + let chars = "XeT"; + let cff = CffBuilder::new("ABCDEF+LMRoman10-Regular") + .cid_glyph(1, true) + .cid_glyph(2, true) + .cid_glyph(3, true) + .build(); + let mut f = Type0Font::new("CIDFontType0", "ABCDEF+LMRoman10-Regular"); + f.program = Program::CidCff(cff); + f.w = Some("[1 [500 500 500]]".into()); + let map: Vec<(u32, usize, String)> = chars + .chars() + .enumerate() + .map(|(i, c)| (i as u32 + 1, 2, c.to_string())) + .collect(); + f.tounicode = Some(bfchar(&map)); + let mut d = DocBuilder::new(); + let f1 = add_type0(&mut d.b, &f); + let content = format!("BT /F1 12 Tf 72 700 Td <{}> Tj ET", cid_hex(chars, "XeTeX")); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +/// FX-INDD: a `/Type1C` font with `Tc` tracking, a visible and a hidden `/OC` layer, and one line +/// inside a Form XObject. +pub fn indd() -> Vec { + let lines = ["visible layer", "hidden layer", "in a frame"]; + let chars = unique(&lines.concat()); + let mut cff = CffBuilder::new("DAAAAA+MinionPro-Regular"); + for ch in chars.chars() { + cff = cff.glyph(&glyph_name(ch), ch != ' '); + } + let mut widths = vec![0.0; 95]; // codes 32..=126 + for ch in chars.chars() { + if let Some(w) = widths.get_mut(u32::from(ch) as usize - 32) { + *w = 500.0; + } + } + let mut f = SimpleFont::new("Type1", "DAAAAA+MinionPro-Regular"); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(widths); + f.flags = Some(34); + f.program = Program::Cff(cff.build()); + let mut d = DocBuilder::new(); + let f1 = add_simple(&mut d.b, &f); + let visible = d.add("<< /Type /OCG /Name (Visible) >>"); + let hidden = d.add("<< /Type /OCG /Name (Hidden) >>"); + let form = d.b.add_stream( + &format!( + "/Type /XObject /Subtype /Form /BBox [0 0 200 50] /Resources << /Font << /F1 {f1} 0 R >> >>" + ), + b"BT /F1 12 Tf 10 10 Td (in a frame) Tj ET", + ); + d.catalog_extra = format!( + "/OCProperties << /OCGs [{visible} 0 R {hidden} 0 R] /D << /Order [{visible} 0 R {hidden} 0 R] /OFF [{hidden} 0 R] >> >>" + ); + d.page(PageSpec::new( + b"/OC /L1 BDC BT /F1 12 Tf 0.12 Tc 72 700 Td (visible layer) Tj ET EMC \ + /OC /L2 BDC BT /F1 12 Tf 72 680 Td (hidden layer) Tj ET EMC \ + q 1 0 0 1 72 600 cm /Fm0 Do Q", + &format!( + "/Font << /F1 {f1} 0 R >> /Properties << /L1 {visible} 0 R /L2 {hidden} 0 R >> \ + /XObject << /Fm0 {form} 0 R >>" + ), + )); + d.build() +} + +/// FX-STD14: Helvetica, Times-Roman and Courier, not embedded, no `/Widths`. +pub fn std14() -> Vec { + let mut d = DocBuilder::new(); + let mut fonts = String::new(); + for (i, base) in ["Helvetica", "Times-Roman", "Courier"].iter().enumerate() { + let id = d.add(format!( + "<< /Type /Font /Subtype /Type1 /BaseFont /{base} /Encoding /WinAnsiEncoding >>" + )); + fonts.push_str(&format!("/F{} {id} 0 R ", i + 1)); + } + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Helvetica line) Tj ET BT /F2 12 Tf 72 680 Td (Times line) Tj ET \ + BT /F3 12 Tf 72 660 Td (Courier line) Tj ET", + &format!("/Font << {fonts}>>"), + )); + d.build() +} + +/// FX-NONEMB: a TrueType font with `/Widths` but no program (the reader substitutes it). +pub fn nonemb() -> Vec { + let mut f = SimpleFont::new("TrueType", "Garamond"); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 95]); + f.flags = Some(32); + let mut d = DocBuilder::new(); + let f1 = add_simple(&mut d.b, &f); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Not embedded) Tj ET", + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +/// FX-OCR: a page image under an invisible (`3 Tr`) text layer in a GlyphLessFont. +pub fn ocr() -> Vec { + let mut f = SimpleFont::new("TrueType", "GlyphLessFont"); + f.encoding = Some("/WinAnsiEncoding".into()); + f.first_char = 32; + f.widths = Some(vec![500.0; 95]); + f.flags = Some(32); + let mut d = DocBuilder::new(); + let f1 = add_simple(&mut d.b, &f); + let img = d.b.add_stream( + "/Type /XObject /Subtype /Image /Width 2 /Height 2 /ColorSpace /DeviceGray /BitsPerComponent 8", + &[255, 0, 0, 255], + ); + d.page(PageSpec::new( + b"q 612 0 0 792 0 0 cm /Im0 Do Q BT 3 Tr /F1 10 Tf 72 700 Td (scanned words) Tj ET", + &format!("/Font << /F1 {f1} 0 R >> /XObject << /Im0 {img} 0 R >>"), + )); + d.build() +} + +/// FX-SHARED: pages 1 and 2 share one `/Contents` stream; pages 3 and 4 share a letterhead part +/// (each also has a body part of its own). +pub fn shared() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + let body = + d.b.add_stream("", b"BT /F1 12 Tf 72 700 Td (Shared body) Tj ET"); + let head = + d.b.add_stream("", b"BT /F1 12 Tf 72 750 Td (Letterhead) Tj ET"); + let spec = PageSpec::new(b"", &res); + d.page_raw(&refs(&[body]), &spec); + d.page_raw(&refs(&[body]), &spec); + for text in ["Body three", "Body four"] { + let own = d.b.add_stream( + "", + format!("BT /F1 12 Tf 72 700 Td ({text}) Tj ET").as_bytes(), + ); + d.page_raw(&refs(&[head, own]), &spec); + } + d.build() +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/producers/tagging.rs b/src-tauri/src/pdf_engine/text_edit/testkit/producers/tagging.rs new file mode 100644 index 0000000..c9d5355 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/producers/tagging.rs @@ -0,0 +1,93 @@ +//! Logical structure and page references (§E.2 rev. 2 `tagged_bookmarked(page)`): a structure +//! tree whose elements point at the page (`/Pg`), an annotation (`/P`), an outline item and a +//! link (`/Dest`), a named destination and `/OpenAction` — the references Word's tagged output +//! carries, none of which may make the page's content "shared" (F1). + +use super::{DocBuilder, PageSpec, HELVETICA}; + +/// A structure tree with one `/P` element per MCID of `page`, listed in `order` (the structure +/// order); the MCIDs in `actual` carry `/ActualText`. Adds `/StructTreeRoot` and `/MarkInfo` to +/// the catalog and returns the page entry to add (`/StructParents 0`). +pub fn structure(d: &mut DocBuilder, page: u32, order: &[i64], actual: &[i64]) -> String { + let root = d.b.alloc(); + let document = d.b.alloc(); + let mut by_mcid: Vec<(i64, u32)> = Vec::new(); + for mcid in order { + let alt = if actual.contains(mcid) { + format!(" /ActualText (alt {mcid})") + } else { + String::new() + }; + let elem = d.b.add(format!( + "<< /Type /StructElem /S /P /P {document} 0 R /Pg {page} 0 R /K {mcid}{alt} >>" + )); + by_mcid.push((*mcid, elem)); + } + let kids = by_mcid + .iter() + .map(|(_, e)| format!("{e} 0 R")) + .collect::>() + .join(" "); + d.b.set( + document, + format!("<< /Type /StructElem /S /Document /P {root} 0 R /K [{kids}] >>"), + ); + let max = order.iter().copied().max().unwrap_or(-1); + let parents: Vec = (0..=max) + .map(|m| { + by_mcid + .iter() + .find(|(x, _)| *x == m) + .map_or("null".to_string(), |(_, e)| format!("{e} 0 R")) + }) + .collect(); + d.b.set( + root, + format!( + "<< /Type /StructTreeRoot /K {document} 0 R /ParentTree << /Nums [0 [{}]] >> \ + /ParentTreeNextKey 1 >>", + parents.join(" ") + ), + ); + d.catalog_extra.push_str(&format!( + " /StructTreeRoot {root} 0 R /MarkInfo << /Marked true >>" + )); + "/StructParents 0".to_string() +} + +/// The structure tree of `mcids` plus a Link annotation (`/P`, `/Dest`), an outline item, a named +/// destination and `/OpenAction`, all pointing at `page`. Returns the page entries to add. +pub fn tagged_bookmarked(d: &mut DocBuilder, page: u32, mcids: &[i64]) -> String { + let extra = structure(d, page, mcids, &[]); + let link = d.b.add(format!( + "<< /Type /Annot /Subtype /Link /Rect [72 600 200 620] /Border [0 0 0] /P {page} 0 R \ + /Dest [{page} 0 R /XYZ 0 792 0] >>" + )); + let outlines = d.b.alloc(); + let item = d.b.add(format!( + "<< /Title (Start) /Parent {outlines} 0 R /Dest [{page} 0 R /Fit] >>" + )); + d.b.set( + outlines, + format!("<< /Type /Outlines /First {item} 0 R /Last {item} 0 R /Count 1 >>"), + ); + d.catalog_extra.push_str(&format!( + " /Outlines {outlines} 0 R /Names << /Dests << /Names [(start) [{page} 0 R /Fit]] >> >> \ + /OpenAction [{page} 0 R /Fit] /PageMode /UseOutlines" + )); + format!("{extra} /Annots [{link} 0 R] /Tabs /S") +} + +/// A tagged, bookmarked Helvetica page (the helper on its own). +pub fn tagged_bookmarked_page() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let page = d.reserve(); + let extra = tagged_bookmarked(&mut d, page, &[0]); + let content = b"/P <> BDC BT /F1 12 Tf 72 700 Td (Tagged line) Tj ET EMC"; + d.page_at( + page, + PageSpec::new(content, &format!("/Font << /F1 {f} 0 R >>")).with(&extra), + ); + d.build() +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/ttf.rs b/src-tauri/src/pdf_engine/text_edit/testkit/ttf.rs new file mode 100644 index 0000000..478e570 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/ttf.rs @@ -0,0 +1,407 @@ +//! TrueType / OpenType font program builder for tests (T2, SPEC §E.1): `head`, `hhea`, `maxp`, +//! `hmtx`, `cmap` ((3,1) and (3,0) format 4, (1,0) format 0), `post` (format 2 names or format 3), +//! `glyf` + `loca` (one triangle per drawn glyph, empty entries for blank ones — the shape of a +//! "zeroed" Word subset — or a raw glyph record such as a composite) or, for OpenType-CFF +//! (`OTTO`), a `CFF ` (or `CFF2`) table instead. + +/// The 258 standard Macintosh glyph names of `post` format 2 (Apple TrueType Reference Manual), as +/// listed by ttf-parser 0.25.1 `src/tables/post.rs` (MIT OR Apache-2.0). A standard name must be +/// stored by its index: readers look standard names up by index only. +#[rustfmt::skip] +pub const MACINTOSH_NAMES: [&str; 258] = [ + ".notdef", ".null", "nonmarkingreturn", "space", "exclam", "quotedbl", "numbersign", "dollar", + "percent", "ampersand", "quotesingle", "parenleft", "parenright", "asterisk", "plus", "comma", + "hyphen", "period", "slash", "zero", "one", "two", "three", "four", "five", "six", "seven", + "eight", "nine", "colon", "semicolon", "less", "equal", "greater", "question", "at", "A", "B", + "C", "D", "E", "F", "G", "H", "I", "J", "K", "L", "M", "N", "O", "P", "Q", "R", "S", "T", "U", + "V", "W", "X", "Y", "Z", "bracketleft", "backslash", "bracketright", "asciicircum", + "underscore", "grave", "a", "b", "c", "d", "e", "f", "g", "h", "i", "j", "k", "l", "m", "n", + "o", "p", "q", "r", "s", "t", "u", "v", "w", "x", "y", "z", "braceleft", "bar", "braceright", + "asciitilde", "Adieresis", "Aring", "Ccedilla", "Eacute", "Ntilde", "Odieresis", "Udieresis", + "aacute", "agrave", "acircumflex", "adieresis", "atilde", "aring", "ccedilla", "eacute", + "egrave", "ecircumflex", "edieresis", "iacute", "igrave", "icircumflex", "idieresis", "ntilde", + "oacute", "ograve", "ocircumflex", "odieresis", "otilde", "uacute", "ugrave", "ucircumflex", + "udieresis", "dagger", "degree", "cent", "sterling", "section", "bullet", "paragraph", + "germandbls", "registered", "copyright", "trademark", "acute", "dieresis", "notequal", "AE", + "Oslash", "infinity", "plusminus", "lessequal", "greaterequal", "yen", "mu", "partialdiff", + "summation", "product", "pi", "integral", "ordfeminine", "ordmasculine", "Omega", "ae", + "oslash", "questiondown", "exclamdown", "logicalnot", "radical", "florin", "approxequal", + "Delta", "guillemotleft", "guillemotright", "ellipsis", "nonbreakingspace", "Agrave", "Atilde", + "Otilde", "OE", "oe", "endash", "emdash", "quotedblleft", "quotedblright", "quoteleft", + "quoteright", "divide", "lozenge", "ydieresis", "Ydieresis", "fraction", "currency", + "guilsinglleft", "guilsinglright", "fi", "fl", "daggerdbl", "periodcentered", "quotesinglbase", + "quotedblbase", "perthousand", "Acircumflex", "Ecircumflex", "Aacute", "Edieresis", "Egrave", + "Iacute", "Icircumflex", "Idieresis", "Igrave", "Oacute", "Ocircumflex", "apple", "Ograve", + "Uacute", "Ucircumflex", "Ugrave", "dotlessi", "circumflex", "tilde", "macron", "breve", + "dotaccent", "ring", "cedilla", "hungarumlaut", "ogonek", "caron", "Lslash", "lslash", "Scaron", + "scaron", "Zcaron", "zcaron", "brokenbar", "Eth", "eth", "Yacute", "yacute", "Thorn", "thorn", + "minus", "multiply", "onesuperior", "twosuperior", "threesuperior", "onehalf", "onequarter", + "threequarters", "franc", "Gbreve", "gbreve", "Idotaccent", "Scedilla", "scedilla", "Cacute", + "cacute", "Ccaron", "ccaron", "dcroat", +]; + +#[derive(Debug, Clone)] +pub struct TtfGlyph { + pub name: String, + pub outline: bool, + pub advance: u16, + /// A raw `glyf` record used instead of the triangle (e.g. a composite). + pub data: Option>, +} + +#[derive(Debug, Clone)] +pub struct TtfBuilder { + /// GID 1, 2, … (GID 0 is a drawn `.notdef`). + pub glyphs: Vec, + pub cmap31: Vec<(u32, u16)>, + pub cmap30: Vec<(u32, u16)>, + pub cmap10: Vec<(u8, u16)>, + pub post_names: bool, + pub units_per_em: u16, + pub ascender: i16, + pub descender: i16, + /// OpenType-CFF: this `CFF ` table replaces `glyf`/`loca`. + pub cff: Option>, + /// OpenType-CFF2: this `CFF2` table replaces `glyf`/`loca`. + pub cff2: Option>, +} + +impl Default for TtfBuilder { + fn default() -> Self { + Self::new() + } +} + +impl TtfBuilder { + pub fn new() -> Self { + TtfBuilder { + glyphs: Vec::new(), + cmap31: Vec::new(), + cmap30: Vec::new(), + cmap10: Vec::new(), + post_names: true, + units_per_em: 1000, + ascender: 800, + descender: -200, + cff: None, + cff2: None, + } + } + + /// Adds a glyph and returns its GID. + pub fn glyph(&mut self, name: &str, outline: bool, advance: u16) -> u16 { + self.glyphs.push(TtfGlyph { + name: name.to_string(), + outline, + advance, + data: None, + }); + self.glyphs.len() as u16 + } + + /// Adds a glyph with a raw `glyf` record (see `composite`) and returns its GID. + pub fn raw_glyph(&mut self, name: &str, data: Vec, advance: u16) -> u16 { + self.glyphs.push(TtfGlyph { + name: name.to_string(), + outline: true, + advance, + data: Some(data), + }); + self.glyphs.len() as u16 + } + + /// A glyph mapped from `ch` in the (3,1) cmap; returns its GID. + pub fn unicode_glyph(&mut self, ch: char, name: &str, outline: bool) -> u16 { + let gid = self.glyph(name, outline, 500); + self.cmap31.push((u32::from(ch), gid)); + gid + } + + pub fn number_of_glyphs(&self) -> u16 { + self.glyphs.len() as u16 + 1 + } + + /// A non-conformant sfnt (review T4 r2 M-2) holding both outline formats: `glyf`/`loca` from + /// the glyphs and the `CFF ` table `cff`, under the sfnt version `magic` (`OTTO` or 1.0). + pub fn build_glyf_and_cff(&self, cff: Vec, magic: u32) -> Vec { + let glyf_only = TtfBuilder { + cff: None, + cff2: None, + ..self.clone() + }; + let font = glyf_only.build(); + let count = usize::from(u16::from_be_bytes([font[4], font[5]])); + let mut tables: Vec<([u8; 4], Vec)> = (0..count) + .map(|i| { + let r = 12 + 16 * i; + let field = |k: usize| { + u32::from_be_bytes([ + font[r + k], + font[r + k + 1], + font[r + k + 2], + font[r + k + 3], + ]) as usize + }; + let tag = [font[r], font[r + 1], font[r + 2], font[r + 3]]; + let (offset, len) = (field(8), field(12)); + (tag, font[offset..offset + len].to_vec()) + }) + .collect(); + tables.push((*b"CFF ", cff)); + tables.sort_by(|a, b| a.0.cmp(&b.0)); + sfnt(magic, &tables) + } + + pub fn build(&self) -> Vec { + let n = self.number_of_glyphs(); + let mut tables: Vec<([u8; 4], Vec)> = Vec::new(); + tables.push((*b"head", self.head())); + tables.push((*b"hhea", self.hhea(n))); + tables.push(( + *b"maxp", + [0x0000_5000u32.to_be_bytes().as_slice(), &n.to_be_bytes()].concat(), + )); + tables.push((*b"hmtx", self.hmtx())); + tables.push((*b"cmap", self.cmap())); + tables.push((*b"post", self.post())); + match (&self.cff, &self.cff2) { + (Some(cff), _) => tables.push((*b"CFF ", cff.clone())), + (None, Some(cff2)) => tables.push((*b"CFF2", cff2.clone())), + (None, None) => { + let (glyf, loca) = self.glyf_loca(); + tables.push((*b"glyf", glyf)); + tables.push((*b"loca", loca)); + } + } + tables.sort_by(|a, b| a.0.cmp(&b.0)); + let magic: u32 = if self.cff.is_some() || self.cff2.is_some() { + 0x4F54_544F + } else { + 0x0001_0000 + }; + sfnt(magic, &tables) + } + + fn head(&self) -> Vec { + let mut h = Vec::new(); + h.extend_from_slice(&0x0001_0000u32.to_be_bytes()); + h.extend_from_slice(&0x0001_0000u32.to_be_bytes()); + h.extend_from_slice(&0u32.to_be_bytes()); + h.extend_from_slice(&0x5F0F_3CF5u32.to_be_bytes()); + h.extend_from_slice(&0u16.to_be_bytes()); + h.extend_from_slice(&self.units_per_em.to_be_bytes()); + h.extend_from_slice(&[0; 16]); + for v in [0i16, self.descender, 1000, self.ascender] { + h.extend_from_slice(&v.to_be_bytes()); + } + h.extend_from_slice(&[0, 0, 0, 8, 0, 2]); + h.extend_from_slice(&1u16.to_be_bytes()); // long loca + h.extend_from_slice(&0u16.to_be_bytes()); + h + } + + fn hhea(&self, n: u16) -> Vec { + let mut h = Vec::new(); + h.extend_from_slice(&0x0001_0000u32.to_be_bytes()); + h.extend_from_slice(&self.ascender.to_be_bytes()); + h.extend_from_slice(&self.descender.to_be_bytes()); + h.extend_from_slice(&[0; 26]); + h.extend_from_slice(&n.to_be_bytes()); + h + } + + fn hmtx(&self) -> Vec { + let mut h = Vec::new(); + h.extend_from_slice(&500u16.to_be_bytes()); + h.extend_from_slice(&0i16.to_be_bytes()); + for g in &self.glyphs { + h.extend_from_slice(&g.advance.to_be_bytes()); + h.extend_from_slice(&0i16.to_be_bytes()); + } + h + } + + fn glyf_loca(&self) -> (Vec, Vec) { + let mut glyf = Vec::new(); + let mut loca = vec![0u32]; + let records = std::iter::once((true, None)) + .chain(self.glyphs.iter().map(|g| (g.outline, g.data.as_ref()))); + for (outline, data) in records { + match data { + Some(data) => { + glyf.extend(data); + while glyf.len() % 4 != 0 { + glyf.push(0); + } + } + None if outline => glyf.extend(triangle()), + None => {} + } + loca.push(glyf.len() as u32); + } + (glyf, loca.iter().flat_map(|o| o.to_be_bytes()).collect()) + } + + fn cmap(&self) -> Vec { + let mut subtables: Vec<(u16, u16, Vec)> = Vec::new(); + if !self.cmap10.is_empty() { + let mut ids = [0u8; 256]; + for (code, gid) in &self.cmap10 { + ids[usize::from(*code)] = *gid as u8; + } + let mut t = Vec::new(); + t.extend_from_slice(&0u16.to_be_bytes()); + t.extend_from_slice(&262u16.to_be_bytes()); + t.extend_from_slice(&0u16.to_be_bytes()); + t.extend_from_slice(&ids); + subtables.push((1, 0, t)); + } + if !self.cmap30.is_empty() { + subtables.push((3, 0, format4(&self.cmap30))); + } + if !self.cmap31.is_empty() { + subtables.push((3, 1, format4(&self.cmap31))); + } + let mut out = Vec::new(); + out.extend_from_slice(&0u16.to_be_bytes()); + out.extend_from_slice(&(subtables.len() as u16).to_be_bytes()); + let mut offset = 4 + 8 * subtables.len(); + for (platform, encoding, data) in &subtables { + out.extend_from_slice(&platform.to_be_bytes()); + out.extend_from_slice(&encoding.to_be_bytes()); + out.extend_from_slice(&(offset as u32).to_be_bytes()); + offset += data.len(); + } + for (_, _, data) in &subtables { + out.extend_from_slice(data); + } + out + } + + fn post(&self) -> Vec { + let mut p = Vec::new(); + let version: u32 = if self.post_names { + 0x0002_0000 + } else { + 0x0003_0000 + }; + p.extend_from_slice(&version.to_be_bytes()); + p.extend_from_slice(&[0; 28]); + if !self.post_names { + return p; + } + p.extend_from_slice(&self.number_of_glyphs().to_be_bytes()); + p.extend_from_slice(&0u16.to_be_bytes()); // .notdef = standard Mac name 0 + let mut custom: Vec<&str> = Vec::new(); + for g in &self.glyphs { + let index = match MACINTOSH_NAMES.iter().position(|n| *n == g.name) { + Some(i) => i as u16, + None => { + custom.push(&g.name); + 257 + custom.len() as u16 + } + }; + p.extend_from_slice(&index.to_be_bytes()); + } + for name in custom { + p.push(name.len() as u8); + p.extend_from_slice(name.as_bytes()); + } + p + } +} + +/// A one-contour triangle (3 on-curve points). +fn triangle() -> Vec { + let mut g = Vec::new(); + for v in [1i16, 50, 0, 450, 600] { + g.extend_from_slice(&v.to_be_bytes()); // numberOfContours, xMin, yMin, xMax, yMax + } + g.extend_from_slice(&2u16.to_be_bytes()); // endPtsOfContours + g.extend_from_slice(&0u16.to_be_bytes()); // instructionLength + g.extend_from_slice(&[1, 1, 1]); // flags: on-curve, 2-byte deltas + for dx in [50i16, 400, -200] { + g.extend_from_slice(&dx.to_be_bytes()); + } + for dy in [0i16, 0, 600] { + g.extend_from_slice(&dy.to_be_bytes()); + } + while g.len() % 4 != 0 { + g.push(0); + } + g +} + +/// A composite `glyf` record placing each of `components` (GIDs) with byte offsets. +pub fn composite(components: &[u16]) -> Vec { + let mut g = Vec::new(); + for v in [-1i16, 0, 0, 1000, 1000] { + g.extend_from_slice(&v.to_be_bytes()); // numberOfContours −1, bbox + } + for (i, gid) in components.iter().enumerate() { + let more = if i + 1 < components.len() { 0x0020 } else { 0 }; + let flags: u16 = 0x0002 | more; // ARGS_ARE_XY_VALUES, byte offsets + g.extend_from_slice(&flags.to_be_bytes()); + g.extend_from_slice(&gid.to_be_bytes()); + g.extend_from_slice(&[i as u8, 0]); + } + g +} + +/// cmap format 4 with one segment per mapping (+ the final 0xFFFF segment). +fn format4(map: &[(u32, u16)]) -> Vec { + let mut pairs: Vec<(u16, u16)> = map.iter().map(|(c, g)| (*c as u16, *g)).collect(); + pairs.sort(); + pairs.dedup_by_key(|p| p.0); + pairs.push((0xFFFF, 0)); + let seg = pairs.len() as u16; + let mut t = Vec::new(); + t.extend_from_slice(&4u16.to_be_bytes()); + t.extend_from_slice(&(16 + 8 * seg).to_be_bytes()); + t.extend_from_slice(&0u16.to_be_bytes()); + t.extend_from_slice(&(seg * 2).to_be_bytes()); + t.extend_from_slice(&[0; 6]); // searchRange, entrySelector, rangeShift (unused by readers here) + for (c, _) in &pairs { + t.extend_from_slice(&c.to_be_bytes()); + } + t.extend_from_slice(&0u16.to_be_bytes()); + for (c, _) in &pairs { + t.extend_from_slice(&c.to_be_bytes()); + } + for (c, g) in &pairs { + let delta = if *c == 0xFFFF { + 1u16 + } else { + g.wrapping_sub(*c) + }; + t.extend_from_slice(&delta.to_be_bytes()); + } + for _ in &pairs { + t.extend_from_slice(&0u16.to_be_bytes()); + } + t +} + +/// An sfnt wrapper: table directory (sorted by tag) and 4-byte aligned tables. +pub fn sfnt(magic: u32, tables: &[([u8; 4], Vec)]) -> Vec { + let mut out = Vec::new(); + out.extend_from_slice(&magic.to_be_bytes()); + out.extend_from_slice(&(tables.len() as u16).to_be_bytes()); + out.extend_from_slice(&[0; 6]); + let mut offset = 12 + 16 * tables.len(); + let mut body = Vec::new(); + for (tag, data) in tables { + out.extend_from_slice(tag); + out.extend_from_slice(&0u32.to_be_bytes()); + out.extend_from_slice(&(offset as u32).to_be_bytes()); + out.extend_from_slice(&(data.len() as u32).to_be_bytes()); + let mut padded = data.clone(); + while padded.len() % 4 != 0 { + padded.push(0); + } + offset += padded.len(); + body.extend(padded); + } + out.extend(body); + out +} diff --git a/src-tauri/src/pdf_engine/text_edit/testkit/type1.rs b/src-tauri/src/pdf_engine/text_edit/testkit/type1.rs new file mode 100644 index 0000000..6f00734 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/testkit/type1.rs @@ -0,0 +1,264 @@ +//! Type1 font program builder for tests (T2, SPEC §E.1): a PDF `FontFile` (clear text + +//! eexec-encrypted private part + 512-zero trailer, with `Length1/2/3`), binary or hex eexec, +//! `/lenIV` 4 or −1, `/Subrs` with the usual hint-replacement entries, and charstrings that draw, +//! draw nothing, use hints/subroutines or compose an accent with `seac`. + +/// One charstring kind. +#[derive(Debug, Clone)] +pub enum T1Glyph { + /// `hsbw`, a closed box, `endchar`. + Box { width: i32 }, + /// The same box behind `hstem`/`vstem` hints and a hint-replacement subroutine call. + HintedBox { width: i32 }, + /// `hsbw … endchar` with nothing drawn (whitespace). + Blank { width: i32 }, + /// `seac` composite of two StandardEncoding codes. + Seac { width: i32, base: u8, accent: u8 }, + /// Plain (unencrypted) charstring bytes, used as given. + Raw(Vec), +} + +#[derive(Debug, Clone)] +pub enum T1Encoding { + Standard, + Custom(Vec<(u8, String)>), +} + +#[derive(Debug, Clone)] +pub struct Type1Builder { + pub font_name: String, + pub encoding: T1Encoding, + pub glyphs: Vec<(String, T1Glyph)>, + /// `/lenIV` (4 = default, −1 = charstrings not encrypted). + pub len_iv: i64, + pub hex: bool, + /// Plain subroutines appended after the five standard ones (indices 5, 6, …). + pub extra_subrs: Vec>, +} + +/// A built `FontFile`: data and its `Length1`/`Length2`/`Length3`. +#[derive(Debug, Clone)] +pub struct Type1File { + pub data: Vec, + pub length1: usize, + pub length2: usize, + pub length3: usize, +} + +impl Type1Builder { + pub fn new(font_name: &str) -> Self { + Type1Builder { + font_name: font_name.to_string(), + encoding: T1Encoding::Standard, + glyphs: vec![(".notdef".to_string(), T1Glyph::Blank { width: 0 })], + len_iv: 4, + hex: false, + extra_subrs: Vec::new(), + } + } + + pub fn glyph(mut self, name: &str, glyph: T1Glyph) -> Self { + self.glyphs.push((name.to_string(), glyph)); + self + } + + pub fn build(&self) -> Type1File { + let mut clear = format!( + "%!PS-AdobeFont-1.0: {name} 001.000\n%%Title: {name}\n11 dict begin\n\ + /FontInfo 2 dict dup begin\n/FullName ({name} \\(test\\)) readonly def\n\ + /ItalicAngle 0 def\nend readonly def\n/FontName /{name} def\n", + name = self.font_name + ); + match &self.encoding { + T1Encoding::Standard => clear.push_str("/Encoding StandardEncoding def\n"), + T1Encoding::Custom(codes) => { + clear.push_str("/Encoding 256 array\n0 1 255 {1 index exch /.notdef put} for\n"); + for (code, name) in codes { + clear.push_str(&format!("dup {code} /{name} put\n")); + } + clear.push_str("readonly def\n"); + } + } + clear.push_str( + "/PaintType 0 def\n/FontType 1 def\n/FontMatrix [0.001 0 0 0.001 0 0] readonly def\n\ + /FontBBox {0 -200 1000 800} readonly def\ncurrentdict end\ncurrentfile eexec\n", + ); + let mut private = Vec::new(); + private.extend_from_slice( + b"dup /Private 8 dict dup begin\n/RD {string currentfile exch readstring pop} executeonly def\n\ + /ND {noaccess def} executeonly def\n/NP {noaccess put} executeonly def\n", + ); + private.extend_from_slice(format!("/lenIV {} def\n", self.len_iv).as_bytes()); + private.extend_from_slice( + b"/BlueValues [-15 0 700 715] def\n/MinFeature {16 16} def\n/password 5839 def\n", + ); + let mut subrs = standard_subrs(); + subrs.extend(self.extra_subrs.iter().cloned()); + private.extend_from_slice(format!("/Subrs {} array\n", subrs.len()).as_bytes()); + for (i, s) in subrs.iter().enumerate() { + let enc = self.charstring(s); + private.extend_from_slice(format!("dup {i} {} RD ", enc.len()).as_bytes()); + private.extend_from_slice(&enc); + private.extend_from_slice(b" NP\n"); + } + private.extend_from_slice(b"ND\n"); + private.extend_from_slice( + format!( + "2 index /CharStrings {} dict dup begin\n", + self.glyphs.len() + ) + .as_bytes(), + ); + for (name, glyph) in &self.glyphs { + let enc = self.charstring(&plain(glyph)); + private.extend_from_slice(format!("/{name} {} -| ", enc.len()).as_bytes()); + private.extend_from_slice(&enc); + private.extend_from_slice(b" |-\n"); + } + private.extend_from_slice( + b"end\nend\nreadonly put\nnoaccess put\ndup /FontName get exch definefont pop\n\ + mark currentfile closefile\n", + ); + let mut plaintext = vec![0u8; 4]; + plaintext.extend(private); + let cipher = encrypt(&plaintext, 55665); + let encrypted = if self.hex { + let mut h = Vec::new(); + for (i, b) in cipher.iter().enumerate() { + h.extend_from_slice(format!("{b:02x}").as_bytes()); + if i % 32 == 31 { + h.push(b'\n'); + } + } + h.push(b'\n'); + h + } else { + cipher + }; + let mut trailer = Vec::new(); + for _ in 0..8 { + trailer.extend_from_slice(&[b'0'; 64]); + trailer.push(b'\n'); + } + trailer.extend_from_slice(b"cleartomark\n"); + let mut data = clear.into_bytes(); + let length1 = data.len(); + data.extend(&encrypted); + data.extend(&trailer); + Type1File { + data, + length1, + length2: encrypted.len(), + length3: trailer.len(), + } + } + + /// Encrypts a plain charstring (r = 4330, `lenIV` zero bytes in front) unless lenIV is −1. + fn charstring(&self, plain: &[u8]) -> Vec { + if self.len_iv < 0 { + return plain.to_vec(); + } + let mut p = vec![0u8; self.len_iv as usize]; + p.extend_from_slice(plain); + encrypt(&p, 4330) + } +} + +/// Type1 encryption (inverse of `fonts::type1::decrypt`). +pub fn encrypt(plain: &[u8], r: u16) -> Vec { + let mut r = r; + plain + .iter() + .map(|p| { + let c = p ^ (r >> 8) as u8; + r = (u16::from(c).wrapping_add(r)) + .wrapping_mul(52845) + .wrapping_add(22719); + c + }) + .collect() +} + +/// A Type1 charstring number. +pub fn num(v: i32) -> Vec { + match v { + -107..=107 => vec![(v + 139) as u8], + 108..=1131 => { + let w = v - 108; + vec![(w / 256 + 247) as u8, (w % 256) as u8] + } + -1131..=-108 => { + let w = -v - 108; + vec![(w / 256 + 251) as u8, (w % 256) as u8] + } + _ => { + let mut b = vec![255]; + b.extend_from_slice(&v.to_be_bytes()); + b + } + } +} + +/// Numbers followed by an operator byte (or `12 x`). +pub fn op(args: &[i32], code: &[u8]) -> Vec { + let mut out: Vec = args.iter().flat_map(|a| num(*a)).collect(); + out.extend_from_slice(code); + out +} + +fn closed_box() -> Vec { + [ + op(&[100, 0], &[21]), + op(&[400, 0], &[5]), + op(&[0, 500], &[5]), + op(&[-400, 0], &[5]), + vec![9], + ] + .concat() +} + +/// The plain charstring of a glyph kind. +pub fn plain(glyph: &T1Glyph) -> Vec { + match glyph { + T1Glyph::Box { width } => [op(&[0, *width], &[13]), closed_box(), vec![14]].concat(), + T1Glyph::HintedBox { width } => [ + op(&[0, *width], &[13]), + op(&[0, 50], &[1]), // hstem + op(&[100, 50], &[3]), // vstem + op(&[4, 1, 3], &[12, 16]), // hint replacement: subr 4 through othersubr 3 … + vec![12, 17], // … pop + vec![10], // … callsubr + closed_box(), + vec![14], + ] + .concat(), + T1Glyph::Blank { width } => [op(&[0, *width], &[13]), vec![14]].concat(), + T1Glyph::Seac { + width, + base, + accent, + } => [ + op(&[0, *width], &[13]), + op(&[0, 0, 0, i32::from(*base), i32::from(*accent)], &[12, 6]), + ] + .concat(), + T1Glyph::Raw(bytes) => bytes.clone(), + } +} + +/// Subrs 0–3 as Adobe's fonts carry them (flex/hint helpers) plus subr 4, a hint replacement. +fn standard_subrs() -> Vec> { + vec![ + [ + op(&[3, 0], &[12, 16]), + vec![12, 17, 12, 17], + vec![12, 33], + vec![11], + ] + .concat(), + [op(&[0, 1], &[12, 16]), vec![11]].concat(), + [op(&[0, 2], &[12, 16]), vec![11]].concat(), + vec![11], + [op(&[0, 60], &[1]), vec![11]].concat(), + ] +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_dto.rs b/src-tauri/src/pdf_engine/text_edit/tests_dto.rs new file mode 100644 index 0000000..12568b9 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_dto.rs @@ -0,0 +1,470 @@ +//! T5 DTO and command tests: DTO-01…03 (`src/lib/editor/__fixtures__/text-edit-dto-contract.json` +//! is the contract with the TypeScript types; `OFFPDF_UPDATE_GOLDEN=1` rewrites it) and +//! CMD-01…08 (`tests_dto/cmd.rs`, through the service layer the Tauri commands call). + +mod cache; +mod cmd; + +use crate::pdf_engine::edit_overlay::{EditDocumentIn, EditObjectIn}; +use crate::pdf_engine::text_edit::dto::{ + EditProblemDto, EditVerdictDto, FaceOptionDto, FacesDto, PageTextDto, RectDto, RunMetricsDto, + RunStyleDto, TextEditIn, TextFontDto, TextPreviewDto, TextRunDto, TextSourceDto, + TextWarningDto, VecDto, +}; +use crate::pdf_engine::text_edit::fonts::FamilyHint; +use crate::pdf_engine::text_edit::reasons::{ + EditProblemCode, Face, StyleField, TextReason, TextWarningCode, +}; +use crate::pdf_engine::text_edit::runs::SpaceMode; +use serde_json::{json, Value}; +use std::collections::BTreeSet; +use std::path::PathBuf; + +fn contract_path() -> PathBuf { + PathBuf::from(env!("CARGO_MANIFEST_DIR")) + .join("../src/lib/editor/__fixtures__/text-edit-dto-contract.json") +} + +fn face_option(available: bool, surface: &[&str]) -> FaceOptionDto { + FaceOptionDto { + available, + surface: surface.iter().map(|s| s.to_string()).collect(), + } +} + +/// One editable and one refused run, with their fonts. +fn page_sample() -> PageTextDto { + let editable = TextRunDto { + id: "t1:0123456789abcdef-4d2:0:120-161".into(), + order: 0, + line: 0, + text: "Invoice 2026".into(), + rect: RectDto { + x: 72.0, + y: 697.5, + w: 61.25, + h: 11.5, + }, + origin: VecDto { x: 72.0, y: 700.0 }, + dir: VecDto { x: 1.0, y: 0.0 }, + ascent: 8.75, + descent: 2.25, + caret_offsets: vec![ + 0.0, 2.5, 8.0, 13.5, 19.0, 22.0, 27.5, 33.0, 35.75, 41.25, 46.75, 52.25, 57.75, + ], + editable: true, + reason: None, + metrics: Some(RunMetricsDto { + surface: vec!["f7-0".into(), "f9-0".into()], + tf_size: 11.04, + effective_size: 11.04, + char_spacing: 0.0, + word_spacing: 0.0, + h_scale: 1.0, + text_to_user: 1.0, + letter_spacing_pt: 0.0, + space_mode: SpaceMode::Glyph, + kern_space: -250.0, + original_width: 57.75, + visible_extent: 468.0, + next_obstacle: Some(120.5), + }), + style: Some(RunStyleDto { + fill: Some("#000000".into()), + size_changeable: true, + colour_changeable: true, + face: Face::Regular, + faces: FacesDto { + regular: face_option(true, &["f7-0", "f9-0"]), + bold: face_option(true, &["f11-0"]), + italic: face_option(false, &[]), + bold_italic: face_option(false, &[]), + }, + }), + substituted: false, + }; + let refused = TextRunDto { + id: "t1:0123456789abcdef-4d2:0:200-230".into(), + order: 1, + line: 1, + text: "hidden layer".into(), + rect: RectDto { + x: 72.0, + y: 677.5, + w: 60.0, + h: 11.5, + }, + origin: VecDto { x: 72.0, y: 680.0 }, + dir: VecDto { x: 1.0, y: 0.0 }, + ascent: 8.75, + descent: 2.25, + caret_offsets: vec![ + 0.0, 6.0, 8.5, 14.0, 19.5, 25.0, 31.0, 34.0, 36.5, 42.0, 48.0, 53.5, 57.0, + ], + editable: false, + reason: Some(TextReason::OptionalContent), + metrics: None, + style: None, + substituted: false, + }; + let font = |key: &str, name: &str, alphabet: &str, embedded: bool| TextFontDto { + key: key.into(), + display_name: name.into(), + family_hint: FamilyHint::Sans, + embedded, + subset: embedded, + alphabet: alphabet.into(), + widths: alphabet.chars().map(|_| 500.0).collect(), + word_space: alphabet.contains(' '), + }; + PageTextDto { + fingerprint: "0123456789abcdef-4d2".into(), + page_index: 0, + page_reason: None, + runs: vec![editable, refused], + fonts: vec![ + font("f11-0", "Calibri Bold", " 0267I", true), + font("f7-0", "Calibri", " 0267DIceinouv", true), + font("f9-0", "Calibri", "ğış", true), + ], + } +} + +fn preview_sample() -> TextPreviewDto { + TextPreviewDto { + page_pdf: Some("JVBERi0xLjcK".into()), + verdicts: vec![ + EditVerdictDto { + run_id: "t1:0123456789abcdef-4d2:0:120-161".into(), + ok: true, + code: None, + chars: Vec::new(), + reason: None, + face: None, + field: None, + detail: None, + delta_pt: 0.0, + new_rect: Some(RectDto { + x: 72.0, + y: 697.5, + w: 61.25, + h: 11.5, + }), + caret_offsets: Some(vec![ + 0.0, 2.5, 8.0, 13.5, 19.0, 22.0, 27.5, 33.0, 35.75, 41.25, 46.75, 52.25, 57.75, + ]), + }, + EditVerdictDto { + run_id: "t1:0123456789abcdef-4d2:0:300-320".into(), + ok: false, + code: Some(EditProblemCode::GlyphMissing), + chars: vec!["Y".into()], + reason: None, + face: None, + field: None, + detail: Some("not in the subset".into()), + delta_pt: 0.0, + new_rect: None, + caret_offsets: None, + }, + EditVerdictDto { + run_id: "t1:0123456789abcdef-4d2:0:330-350".into(), + ok: false, + code: Some(EditProblemCode::TextEditRefused), + chars: Vec::new(), + reason: Some(TextReason::SharedContent), + face: None, + field: None, + detail: None, + delta_pt: 0.0, + new_rect: None, + caret_offsets: None, + }, + EditVerdictDto { + run_id: "t1:0123456789abcdef-4d2:0:360-380".into(), + ok: false, + code: Some(EditProblemCode::FaceUnavailable), + chars: vec!["ğ".into()], + reason: None, + face: Some(Face::BoldItalic), + field: None, + detail: None, + delta_pt: 0.0, + new_rect: None, + caret_offsets: None, + }, + EditVerdictDto { + run_id: "t1:0123456789abcdef-4d2:0:390-410".into(), + ok: false, + code: Some(EditProblemCode::StyleUnavailable), + chars: Vec::new(), + reason: None, + face: None, + field: Some(StyleField::Colour), + detail: None, + delta_pt: 0.0, + new_rect: None, + caret_offsets: None, + }, + ], + page_problem: Some(EditProblemDto { + code: EditProblemCode::PenDrift, + detail: Some("phase=A check=A4 page=1 drift 0.02 pt".into()), + }), + warnings: vec![TextWarningDto { + run_id: "t1:0123456789abcdef-4d2:0:120-161".into(), + code: TextWarningCode::NextTextOverlap, + detail: None, + }], + } +} + +fn text_edit_in_sample() -> Value { + json!({ + "runId": "t1:0123456789abcdef-4d2:0:120-161", + "originalText": "Invoice 2026", + "text": "Invoice 2027", + "style": { "sizePt": 12.5, "face": "boldItalic", "fill": "#c71c1c", "letterSpacingPt": 0.5 }, + }) +} + +fn source_text_export_sample() -> Value { + json!({ + "id": "obj-7", + "kind": "sourceText", + "pageIndex": 2, + "rect": { "x": 72.0, "y": 697.5, "w": 61.25, "h": 11.5 }, + "locked": true, + "runId": "t1:0123456789abcdef-4d2:0:120-161", + "sourceFingerprint": "0123456789abcdef-4d2", + "sourcePageIndex": 0, + "originalText": "Invoice 2026", + "text": "Invoice 2027", + "style": { "fill": "#c71c1c" }, + }) +} + +/// The whole contract document. +fn contract() -> Value { + let source = TextSourceDto { + fingerprint: "0123456789abcdef-4d2".into(), + page_count: 3, + warnings: vec![ + "This PDF is larger than 100 MB, so checking and saving text changes takes longer." + .into(), + ], + }; + let ser = |v: Result| v.unwrap_or_else(|e| panic!("serialise: {e}")); + json!({ + "version": 1, + "about": "One serialised sample of every Edit text DTO (Rust → TS) and of every input (TS → Rust). Written by src-tauri text_edit::tests_dto (DTO-01, OFFPDF_UPDATE_GOLDEN=1); the TS types in src/lib/types.ts must have exactly these keys.", + "TextSourceInfo": ser(serde_json::to_value(&source)), + "PageText": ser(serde_json::to_value(page_sample())), + "TextPreview": ser(serde_json::to_value(preview_sample())), + "TextEditIn": text_edit_in_sample(), + "SourceTextObject": source_text_export_sample(), + }) +} + +/// DTO-01: the serialised samples equal the contract file (values compared after parsing). +#[test] +fn dto_01_serialised_dtos_equal_the_contract_file() { + let expected = contract(); + let path = contract_path(); + if std::env::var("OFFPDF_UPDATE_GOLDEN").as_deref() == Ok("1") { + let mut text = serde_json::to_string_pretty(&expected).expect("json"); + text.push('\n'); + std::fs::write(&path, text).expect("write contract"); + } + let on_disk: Value = serde_json::from_str( + &std::fs::read_to_string(&path).unwrap_or_else(|e| panic!("{}: {e}", path.display())), + ) + .expect("contract JSON"); + assert_eq!( + on_disk, expected, + "DTO-01: text-edit-dto-contract.json is out of date" + ); +} + +/// DTO-02: the contract's inputs deserialize into the Rust input types. +#[test] +fn dto_02_text_edit_in_and_source_text_export_deserialize() { + let c = contract(); + let edit: TextEditIn = serde_json::from_value(c["TextEditIn"].clone()).expect("TextEditIn"); + assert_eq!(edit.text, "Invoice 2027"); + assert_eq!(edit.style.face, Some(Face::BoldItalic)); + assert_eq!(edit.style.size_pt, Some(12.5)); + assert_eq!(edit.style.letter_spacing_pt, Some(0.5)); + // `style` may be omitted (an unstyled change). + let bare: TextEditIn = serde_json::from_value(json!({ + "runId": "r", "originalText": "a", "text": "b" + })) + .expect("bare TextEditIn"); + assert!(bare.style.is_empty()); + let doc: EditDocumentIn = serde_json::from_value(json!({ + "version": 1, + "objects": [c["SourceTextObject"].clone()], + })) + .expect("export document"); + match doc.objects.first() { + Some(EditObjectIn::SourceText { + page_index, + source_page_index, + run_id, + source_fingerprint, + original_text, + text, + style, + rect, + }) => { + assert_eq!((*page_index, *source_page_index), (2, 0)); + assert_eq!(run_id, "t1:0123456789abcdef-4d2:0:120-161"); + assert_eq!(source_fingerprint, "0123456789abcdef-4d2"); + assert_eq!( + (original_text.as_str(), text.as_str()), + ("Invoice 2026", "Invoice 2027") + ); + assert_eq!(style.fill.as_deref(), Some("#c71c1c")); + assert_eq!((rect.x, rect.w), (72.0, 61.25)); + } + other => panic!("DTO-02: not a sourceText object: {other:?}"), + } +} + +/// DTO-03: reason, problem and warning codes serialise as the SCREAMING_SNAKE strings of +/// `text-reasons.json`; faces, fields, space modes and family hints in camelCase. +#[test] +fn dto_03_codes_serialise_as_the_shared_strings() { + let json: Value = + serde_json::from_str(include_str!("../../../../src/lib/editor/text-reasons.json")) + .expect("text-reasons.json"); + let list = |key: &str| -> Vec { + json[key] + .as_array() + .expect("list") + .iter() + .map(|v| v.as_str().expect("string").to_string()) + .collect() + }; + let ser = |v: Value| v.as_str().expect("a string").to_string(); + let run: Vec = TextReason::RUN_PRIORITY + .iter() + .map(|r| ser(json!(r))) + .collect(); + let page: Vec = TextReason::PAGE.iter().map(|r| ser(json!(r))).collect(); + let problems: Vec = EditProblemCode::ALL.iter().map(|c| ser(json!(c))).collect(); + let warnings: Vec = TextWarningCode::ALL.iter().map(|c| ser(json!(c))).collect(); + assert_eq!(run, list("run"), "DTO-03 run"); + assert_eq!(page, list("page"), "DTO-03 page"); + assert_eq!(problems, list("problem"), "DTO-03 problem"); + assert_eq!(warnings, list("warning"), "DTO-03 warning"); + for r in TextReason::RUN_PRIORITY.iter().chain(TextReason::PAGE) { + assert_eq!(ser(json!(r)), r.as_str()); + } + assert_eq!(json!(Face::BoldItalic), json!("boldItalic")); + assert_eq!(json!(StyleField::Colour), json!("colour")); + assert_eq!(json!(SpaceMode::Kern), json!("kern")); + assert_eq!(json!(FamilyHint::Mono), json!("mono")); +} + +/// Every object key path of `v` (`a.b[].c`), arrays merged. +/// (key path, JSON kind) of every non-null value under `v` (review-T5 M5: value types, not only +/// key names; `null` is left out because an `Option` serialises either way). +pub(crate) fn kind_paths(v: &Value, at: &str, out: &mut BTreeSet<(String, &'static str)>) { + let kind = match v { + Value::Null => None, + Value::Bool(_) => Some("boolean"), + Value::Number(_) => Some("number"), + Value::String(_) => Some("string"), + Value::Array(_) => Some("array"), + Value::Object(_) => Some("object"), + }; + if let Some(k) = kind { + out.insert((at.to_string(), k)); + } + match v { + Value::Object(m) => m + .iter() + .for_each(|(k, x)| kind_paths(x, &format!("{at}.{k}"), out)), + Value::Array(a) => a + .iter() + .for_each(|x| kind_paths(x, &format!("{at}[]"), out)), + _ => {} + } +} + +pub(crate) fn key_paths(v: &Value, at: &str, out: &mut BTreeSet) { + match v { + Value::Object(m) => { + for (k, x) in m { + let p = format!("{at}.{k}"); + out.insert(p.clone()); + key_paths(x, &p, out); + } + } + Value::Array(a) => { + for x in a { + key_paths(x, &format!("{at}[]"), out); + } + } + _ => {} + } +} + +/// review-T5 L4: the DTO's current letter spacing is the planner's own reading (one formula), so +/// the editor sending it back unchanged plans nothing — on a run with `Tc`, `Tz` and a scaled +/// text matrix, where a second copy of the formula could drift. +#[test] +fn letter_spacing_sent_back_unchanged_is_a_no_op() { + use crate::pdf_engine::source_content::classify_source_page; + use crate::pdf_engine::text_edit::context::SnapshotContext; + use crate::pdf_engine::text_edit::rewrite::{letter_spacing_pt, plan_page, SourceTextStyleIn}; + use crate::pdf_engine::text_edit::runs::build_page_model; + use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; + use crate::pdf_engine::text_edit::testkit::producers::helvetica_page; + use crate::pdf_engine::text_edit::testkit::Scratch; + let pdf = helvetica_page(b"BT /F1 10 Tf 0.7 Tc 85 Tz 1.5 0 0 1.5 72 700 Tm (Spaced out) Tj ET"); + let dir = Scratch::new("dto_l4"); + let path = dir.write("spaced.pdf", &pdf); + let ctx = SnapshotContext::new(snapshot_from_bytes(&path, pdf, None).expect("snapshot")); + let model = build_page_model(&ctx, 0, None).expect("model"); + let dto = crate::pdf_engine::text_edit::dto::page_text( + &model, + &classify_source_page(&ctx, &model, None), + ); + let run = dto + .runs + .iter() + .find(|r| r.text == "Spaced out") + .expect("run"); + let metrics = run.metrics.as_ref().expect("editable"); + let planner = model + .run(&run.id) + .map(letter_spacing_pt) + .expect("model run"); + assert_eq!(metrics.letter_spacing_pt, planner, "one formula"); + assert!( + metrics.letter_spacing_pt > 0.5, + "{}", + metrics.letter_spacing_pt + ); + let edit = TextEditIn { + run_id: run.id.clone(), + original_text: run.text.clone(), + text: run.text.clone(), + style: SourceTextStyleIn { + letter_spacing_pt: Some(metrics.letter_spacing_pt), + ..Default::default() + }, + }; + let out = plan_page(&ctx, &model, &[edit]).expect("plan"); + assert!( + out.verdicts[0].problem.is_none(), + "{:?}", + out.verdicts[0].problem + ); + assert!( + out.plan.is_none(), + "the unchanged spacing was planned as a change" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_dto/cache.rs b/src-tauri/src/pdf_engine/text_edit/tests_dto/cache.rs new file mode 100644 index 0000000..96be1b0 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_dto/cache.rs @@ -0,0 +1,242 @@ +//! The editor cache under the review-T5 findings: a background `qpdf --check` that lost its input +//! (M1: "Clear temp files" during a check, and the harness deleting a folder under a running +//! check), a release during a preview (M2), several calls reading a large file at once (M3), and +//! page models sized by `walk.model_bytes` (fonts included). + +use super::cmd::settle; +#[cfg(unix)] +use crate::pdf_engine::text_edit::cache::seams; +use crate::pdf_engine::text_edit::cache::{clear_stale_folders, TextEditCache}; +use crate::pdf_engine::text_edit::preview::cache_dir_for; +use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; +use crate::pdf_engine::text_edit::service; +use crate::pdf_engine::text_edit::tests_e2e::save::pages_doc; +use crate::pdf_engine::text_edit::tests_e2e::E2e; +use std::path::Path; +use std::sync::{Arc, Barrier}; + +/// A qpdf for background checks that waits until `gate` exists, then runs the real qpdf. +#[cfg(unix)] +pub(crate) fn held_qpdf(t: &E2e, gate: &Path) -> std::path::PathBuf { + use std::os::unix::fs::PermissionsExt; + let script = t.scratch.path("held-qpdf.sh"); + let body = format!( + "#!/bin/sh\nwhile [ ! -e '{}' ]; do sleep 0.02; done\nexec '{}' \"$@\"\n", + gate.display(), + t.engines.qpdf.display() + ); + std::fs::write(&script, body).expect("script"); + std::fs::set_permissions(&script, std::fs::Permissions::from_mode(0o755)).expect("chmod"); + script +} + +/// review-T5 M1, the production trigger: "Clear temp files" deletes the folder before the +/// background check reads it. The check ends in an error (never a verdict about these bytes), +/// the preview starts it again on a fresh copy, and a Save of the same bytes is not refused. +#[cfg(unix)] +#[test] +fn m1_clear_temp_files_during_a_check_is_not_a_repair_verdict() { + let Some(t) = E2e::new("m1_clear_temp") else { + return; + }; + let gate = t.scratch.path("go"); + let src = t.file("cleared.pdf", &pages_doc(&["Cleared line"])); + seams::set_background_qpdf(Some(held_qpdf(&t, &gate))); + let fp = t.open(&src).fingerprint; + seams::set_background_qpdf(None); + let dir = cache_dir_for(&t.temp_root(), &fp); + std::fs::remove_dir_all(t.temp_root().join("textedit")).expect("clear temp files"); + std::fs::write(&gate, b"").expect("gate"); + t.cache.wait_checks(); + assert!( + !dir.join("source.pdf").exists(), + "the check ran on a missing file" + ); + let p = t.preview( + &src, + 0, + &[( + "Cleared line", + "Cleared lines", + SourceTextStyleIn::default(), + )], + ); + assert!(p.page_pdf.is_some(), "{:?}", p.page_problem); + let saved = t + .save( + &[(&src, "1-z")], + vec![t.edit(&src, 0, 0, "Cleared line", "Cleared lines")], + ) + .unwrap_or_else(|e| panic!("save of the same bytes: {e} {:?}", e.details)); + t.assert_saved( + &saved.path, + &[(&src, 0, 0, "Cleared line", "Cleared lines")], + ); +} + +/// review-T5 M1, the test-harness side: dropping an `E2e` waits for its checks, so no test +/// deletes a folder under a running check. +#[cfg(unix)] +#[test] +fn m1_a_dropped_harness_leaves_no_check_running() { + let Some(t) = E2e::new("m1_harness_drop") else { + return; + }; + let gate = t.scratch.path("go"); + let src = t.file("held.pdf", &pages_doc(&["Held check"])); + seams::set_background_qpdf(Some(held_qpdf(&t, &gate))); + t.open(&src); + seams::set_background_qpdf(None); + assert_eq!(t.cache.running_checks(), 1, "the check is held"); + let cache = t.cache.clone(); + let opener = std::thread::spawn(move || { + std::thread::sleep(std::time::Duration::from_millis(200)); + std::fs::write(gate, b"").expect("gate"); + }); + drop(t); + assert_eq!(cache.running_checks(), 0, "a check outlived its harness"); + opener.join().expect("gate thread"); +} + +/// review-T5 M2: a release during a preview. The preview still works (it writes its copies +/// into the leased folder), and the folder is gone when the preview ends: no copy of the user's +/// file outlives the session. +#[test] +fn m2_release_during_a_preview_leaves_no_copy() { + let Some(t) = E2e::new("m2_release_preview") else { + return; + }; + let src = t.file("leased.pdf", &pages_doc(&["Leased line"])); + let fp = t.open(&src).fingerprint; + settle(&t, &src); + let dir = cache_dir_for(&t.temp_root(), &fp); + let cache = t.cache.clone(); + let path = src.clone(); + service::seams::set_during_preview(Some(Box::new(move || { + service::release_source(&cache, &path.to_string_lossy()); + }))); + let p = t.preview( + &src, + 0, + &[("Leased line", "Leased lines", SourceTextStyleIn::default())], + ); + service::seams::set_during_preview(None); + assert!(p.page_pdf.is_some(), "{:?}", p.page_problem); + assert_eq!(t.cache.cached_len(), 0, "released"); + assert!( + !dir.exists(), + "a copy outlived the release: {:?}", + std::fs::read_dir(&dir).map(|d| d.flatten().map(|e| e.file_name()).collect::>()) + ); +} + +/// review-T5 M2: copies left by a crash are deleted at startup. +#[test] +fn m2_startup_deletes_stale_copies() { + let dir = crate::pdf_engine::text_edit::testkit::Scratch::new("m2_startup"); + let stale = dir.path("temp").join("textedit").join("0123-abcd"); + std::fs::create_dir_all(&stale).expect("stale folder"); + std::fs::write(stale.join("source.pdf"), b"%PDF-1.7").expect("stale copy"); + clear_stale_folders(&dir.path("temp")); + assert!(!dir.path("temp").join("textedit").exists()); + assert!(dir.path("temp").is_dir(), "only textedit/ is cleared"); +} + +/// review-T5 M3: calls that overlap on a file too large to cache share one read of it (the +/// test lowers this cache's limit so a small file counts as large). +#[test] +fn m3_overlapping_calls_share_one_read_of_a_large_file() { + let Some(t) = E2e::new("m3_single_flight") else { + return; + }; + let src = t.file("large.pdf", &pages_doc(&["Large file"])); + t.cache.with_seams(|s| s.snapshot_bytes_max = Some(16)); + const CALLS: usize = 6; + let start = Arc::new(Barrier::new(CALLS)); + let hold = Arc::new(Barrier::new(CALLS)); + let workers: Vec<_> = (0..CALLS) + .map(|_| { + let (cache, engines): (TextEditCache, _) = (t.cache.clone(), t.engines.clone()); + let (start, hold) = (Arc::clone(&start), Arc::clone(&hold)); + let (src, temp) = (src.clone(), t.temp_root()); + std::thread::spawn(move || { + start.wait(); + let s = cache.open(&src, &temp, &engines).expect("open"); + hold.wait(); + Arc::as_ptr(&s) as usize + }) + }) + .collect(); + let ptrs: Vec = workers + .into_iter() + .map(|w| w.join().expect("worker")) + .collect(); + assert_eq!(t.cache.cached_len(), 0, "the file is pinned, not cached"); + assert_eq!( + t.cache.with_seams(|s| s.loads), + 1, + "one read for {CALLS} calls" + ); + assert!(ptrs.iter().all(|p| *p == ptrs[0]), "one shared source"); + // Not kept between calls: the next call reads the file again. + t.cache + .open(&src, &t.temp_root(), &t.engines) + .expect("open"); + assert_eq!(t.cache.with_seams(|s| s.loads), 2); +} + +/// Item 7 (review-T3-budget HIGH-1 hand-off): the cache charges each page model its +/// `walk.model_bytes` (everything the model keeps, its fonts included). +#[test] +fn page_models_are_charged_their_model_bytes() { + let Some(t) = E2e::new("cache_model_bytes") else { + return; + }; + let src = t.file("two.pdf", &pages_doc(&["Page one", "Page two"])); + let fp = t.open(&src).fingerprint; + let s = t + .cache + .get(&src, &t.temp_root(), &fp, &t.engines) + .expect("source"); + let a = t.cache.page(&s, 0).expect("page 1"); + let b = t.cache.page(&s, 1).expect("page 2"); + let want = a.walk.model_bytes + b.walk.model_bytes; + assert!(a.walk.model_bytes >= a.approx_bytes() && b.walk.model_bytes >= b.approx_bytes()); + assert_eq!(t.cache.page_model_bytes(&s), want, "LRU charge"); +} + +/// review-T5 live B1: text drawn through a Form XObject is listed (refused `NESTED_FORM`, never +/// editable), not left out — the page otherwise reads as having no text. +#[test] +fn form_xobject_text_is_listed_as_refused_lines() { + let Some(t) = E2e::new("b1_form_lines") else { + return; + }; + let fixture = + Path::new(env!("CARGO_MANIFEST_DIR")).join("../fixtures/source-edit/text-nested-form.pdf"); + let src = t.file("nested.pdf", &std::fs::read(fixture).expect("fixture")); + let (_, dto) = t.page(&src, 0); + let lines: Vec<_> = dto + .runs + .iter() + .map(|r| (r.text.as_str(), r.editable, r.reason)) + .collect(); + assert_eq!( + lines, + vec![( + "Hi", + false, + Some(crate::pdf_engine::text_edit::reasons::TextReason::NestedForm) + )] + ); + let run = &dto.runs[0]; + assert!( + run.metrics.is_none() && run.style.is_none(), + "no editing data" + ); + assert!( + run.rect.w > 0.0 && run.rect.h > 0.0, + "the line has a box: {:?}", + run.rect + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_dto/cmd.rs b/src-tauri/src/pdf_engine/text_edit/tests_dto/cmd.rs new file mode 100644 index 0000000..b74513a --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_dto/cmd.rs @@ -0,0 +1,405 @@ +//! CMD-01…08: the command layer (`service`, as the Tauri commands call it) with real qpdf and +//! Poppler: open refusals, stale fingerprints, verdicts that are data (not errors), page +//! indices, release and eviction of the temporary folders, the classifier wiring, and the +//! background `qpdf --check`. + +use super::{contract, key_paths, kind_paths}; +use crate::pdf_engine::source_content::{classify_source_page, SourceCapability}; +use crate::pdf_engine::text_edit::cache::seams as cache_seams; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::engines::SourceCheck; +use crate::pdf_engine::text_edit::preview::cache_dir_for; +use crate::pdf_engine::text_edit::reasons::{EditProblemCode, Face, StyleField, TextReason}; +use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::service; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::pdf::zlib_zero_bomb; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::tests_e2e::preview::edit_in; +use crate::pdf_engine::text_edit::tests_e2e::refusals::needs_repair_doc; +use crate::pdf_engine::text_edit::tests_e2e::save::pages_doc; +use crate::pdf_engine::text_edit::tests_e2e::E2e; +use std::collections::BTreeSet; +use std::path::Path; + +fn code_of(r: Result) -> String { + match r { + Ok(_) => "ok".into(), + Err(e) => e.code, + } +} + +#[test] +fn cmd_01_open_refusals_and_the_synchronous_page_map_check() { + let Some(t) = E2e::new("cmd_01") else { + return; + }; + let ok = t.file("ok.pdf", &pages_doc(&["One", "Two"])); + let info = t.open(&ok); + assert_eq!(info.page_count, 2); + assert!(info.warnings.is_empty()); + assert_eq!( + info.fingerprint.len(), + 16 + 1 + format!("{:x}", std::fs::metadata(&ok).expect("meta").len()).len() + ); + let cases = [ + ("enc.pdf", fx::encrypted(&t.engines), "ENCRYPTED"), + ("signed.pdf", fx::signed(), "SIGNED"), + ("xfa.pdf", fx::xfa(), "UNSUPPORTED_XFA"), + // lopdf and qpdf disagree about the pages: refused before anything is inspected. + ("kid.pdf", fx::kid_without_type(), "PDF_NEEDS_REPAIR"), + ("hybrid.pdf", fx::hybrid_xref(false), "PDF_NEEDS_REPAIR"), + ]; + for (name, pdf, code) in cases { + let p = t.file(name, &pdf); + assert_eq!(code_of(t.try_open(&p)), code, "CMD-01 {name}"); + } +} + +#[test] +fn cmd_02_a_wrong_or_outdated_fingerprint_is_stale() { + let Some(t) = E2e::new("cmd_02") else { + return; + }; + let p = t.file("stale.pdf", &pages_doc(&["Before"])); + let fp = t.open(&p).fingerprint; + assert_eq!(code_of(t.try_inspect(&p, "0000000000000000-1", 0)), "STALE"); + std::fs::write(&p, pages_doc(&["After!"])).expect("rewrite"); + let e = t.try_inspect(&p, &fp, 0).err().expect("stale"); + assert_eq!(e.code, "STALE"); + assert_eq!( + e.message, + "\u{201c}stale.pdf\u{201d} was changed after you started editing it, so your text changes no longer match it." + ); + // The reloaded snapshot serves the new fingerprint. + let fresh = t.open(&p).fingerprint; + assert_ne!(fresh, fp); + assert!(t.try_inspect(&p, &fresh, 0).is_ok()); +} + +#[test] +fn cmd_03_user_problems_are_verdicts_not_errors() { + let Some(t) = E2e::new("cmd_03") else { + return; + }; + let word = t.file("word.pdf", &fx::word()); + let p = t.preview( + &word, + 0, + &[( + "Invoice 2026", + "Invoice 2026Y", + SourceTextStyleIn::default(), + )], + ); + let v = &p.verdicts[0]; + assert_eq!((v.ok, v.code), (false, Some(EditProblemCode::GlyphMissing))); + assert_eq!(v.chars, vec!["Y".to_string()]); + let italic = SourceTextStyleIn { + face: Some(Face::Italic), + ..SourceTextStyleIn::default() + }; + let p = t.preview(&word, 0, &[("Due", "Due", italic)]); + let v = &p.verdicts[0]; + assert_eq!( + (v.code, v.face), + (Some(EditProblemCode::FaceUnavailable), Some(Face::Italic)) + ); + let gs = t.file("gs.pdf", &fx::extgstate_font()); + let bigger = SourceTextStyleIn { + size_pt: Some(20.0), + ..SourceTextStyleIn::default() + }; + let p = t.preview(&gs, 0, &[("Hi there", "Hi there", bigger)]); + let v = &p.verdicts[0]; + assert_eq!( + (v.code, v.field), + ( + Some(EditProblemCode::StyleUnavailable), + Some(StyleField::Size) + ) + ); + // A refused run: the verdict names its reason. + let indd = t.file("indd.pdf", &fx::indd()); + let (fp, run) = t.run(&indd, 0, "hidden layer"); + let p = service::preview_edits( + &t.cache, + &t.engines, + &t.temp_root(), + &indd.to_string_lossy(), + &fp, + 0, + &[edit_in( + &run.id, + "hidden layer", + "hidden layers", + SourceTextStyleIn::default(), + )], + ) + .expect("CMD-03 refused preview is not an error"); + let v = &p.verdicts[0]; + assert_eq!( + (v.code, v.reason), + ( + Some(EditProblemCode::TextEditRefused), + Some(TextReason::OptionalContent) + ) + ); + assert!(p.page_pdf.is_none()); +} + +#[test] +fn cmd_04_a_page_the_file_does_not_have_is_invalid_pages() { + let Some(t) = E2e::new("cmd_04") else { + return; + }; + let p = t.file("one.pdf", &pages_doc(&["Only page"])); + let fp = t.open(&p).fingerprint; + let e = t.try_inspect(&p, &fp, 5).err().expect("no page 6"); + assert_eq!(e.code, "INVALID_PAGES"); +} + +/// Waits for the background check of `path` (so its folder is not busy). +pub(super) fn settle(t: &E2e, path: &Path) { + let src = t + .cache + .open(path, &t.temp_root(), &t.engines) + .expect("open"); + let check = t.cache.source_check(&src, &t.engines).expect("check"); + assert!(matches!(check.wait(None), Ok(SourceCheck::Clean))); +} + +#[test] +fn cmd_05_06_release_and_eviction_delete_the_temporary_folder() { + let Some(t) = E2e::new("cmd_05") else { + return; + }; + let temp = t.temp_root(); + let a = t.file("a.pdf", &pages_doc(&["File a"])); + let fp_a = t.open(&a).fingerprint; + let dir_a = cache_dir_for(&temp, &fp_a); + assert!(dir_a.join("source.pdf").is_file(), "CMD-05 snapshot copy"); + settle(&t, &a); + service::release_source(&t.cache, &a.to_string_lossy()); + assert_eq!(t.cache.cached_len(), 0, "CMD-05 forgotten"); + assert!( + !dir_a.exists(), + "CMD-06 release deletes {}", + dir_a.display() + ); + // Reopening works; a preview fills the folder again. + let fp_a = t.open(&a).fingerprint; + assert!(t.try_inspect(&a, &fp_a, 0).is_ok()); + let _ = t.preview(&a, 0, &[("File a", "File A", SourceTextStyleIn::default())]); + assert!( + dir_a.join("p1.pdf").is_file(), + "CMD-06 preview extraction cached" + ); + settle(&t, &a); + // Two more sources evict the first (CACHE_SNAPSHOTS_MAX = 2) and its folder goes. + let b = t.file("b.pdf", &pages_doc(&["File b"])); + let c = t.file("c.pdf", &pages_doc(&["File c"])); + t.open(&b); + settle(&t, &b); + t.open(&c); + assert_eq!(t.cache.cached_len(), 2); + assert!( + !dir_a.exists(), + "CMD-06 eviction deletes {}", + dir_a.display() + ); + // A second path with the same bytes keeps the folder until both are released. + let b2 = t.file("b-copy.pdf", &pages_doc(&["File b"])); + let fp_b = t.open(&b2).fingerprint; + let dir_b = cache_dir_for(&temp, &fp_b); + settle(&t, &b2); + service::release_source(&t.cache, &b.to_string_lossy()); + assert!(dir_b.is_dir(), "CMD-06 shared folder kept"); + service::release_source(&t.cache, &b2.to_string_lossy()); + assert!( + !dir_b.exists(), + "CMD-06 shared folder deleted with the last user" + ); +} + +#[test] +fn cmd_07_run_capabilities_come_from_the_classifier() { + let Some(t) = E2e::new("cmd_07") else { + return; + }; + let fixtures = [ + ("word", fx::word()), + ("indd", fx::indd()), + ("shared", fx::shared()), + ("perglyph", fx::per_glyph(Some(9.0))), + ("type3", fx::type3()), + ("ocr", fx::ocr()), + ("skia", fx::skia()), + ]; + for (name, pdf) in fixtures { + let path = t.file(&format!("{name}.pdf"), &pdf); + let pages = t.open(&path).page_count; + let snap = snapshot_from_bytes(&path, pdf.clone(), None).expect("snapshot"); + let ctx = SnapshotContext::new(snap); + for page in 0..pages { + let (_, dto) = t.page(&path, page); + let model = build_page_model(&ctx, page, None).expect("model"); + let classified = classify_source_page(&ctx, &model, None); + assert_eq!( + dto.page_reason, classified.page_reason, + "CMD-07 {name} p{page}" + ); + // The page's runs, then its Form lines (refused `NESTED_FORM`, review-T5 live B1). + assert_eq!( + dto.runs.len(), + classified.runs.len() + classified.form_lines.len(), + "CMD-07 {name} p{page}" + ); + let (own, forms) = dto.runs.split_at(classified.runs.len()); + for (run, line) in forms.iter().zip(&classified.form_lines) { + assert_eq!( + (&run.id, &run.text), + (&line.id, &line.text), + "CMD-07 {name}" + ); + assert!(!run.editable && run.reason == Some(TextReason::NestedForm)); + } + for (run, cap) in own.iter().zip(&classified.runs) { + assert_eq!(run.id, cap.run_id, "CMD-07 {name}"); + assert_eq!( + run.editable, + cap.capability == SourceCapability::Supported, + "CMD-07 {name} {:?}", + run.text + ); + assert_eq!(run.reason, cap.reason, "CMD-07 {name} {:?}", run.text); + assert_eq!(run.metrics.is_some(), run.editable); + assert_eq!(run.style.is_some(), run.editable); + } + } + } +} + +/// A one-page file whose catalog references a stream of `mib` MiB of zeros: our read never +/// decodes it, `qpdf --check` does. +fn slow_check_doc(mib: usize) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let bomb = d.b.add_stream("/Filter /FlateDecode", &zlib_zero_bomb(mib)); + d.catalog_extra + .push_str(&format!(" /PieceInfo << /OffPDF << /Data {bomb} 0 R >> >>")); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Slow check) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +#[test] +fn cmd_08_open_does_not_wait_for_qpdf_check_and_the_preview_does() { + let Some(t) = E2e::new("cmd_08") else { + return; + }; + let slow = t.file("slow.pdf", &slow_check_doc(1024)); + // On Unix the background check is also held until the test lets it go, so "open returned + // before the check finished" does not depend on how long qpdf takes (review-T5 L6). + let gate = t.scratch.path("go"); + #[cfg(unix)] + cache_seams::set_background_qpdf(Some(super::cache::held_qpdf(&t, &gate))); + let info = t.open(&slow); + cache_seams::set_background_qpdf(None); + let src = t + .cache + .open(&slow, &t.temp_root(), &t.engines) + .expect("cached"); + let check = t.cache.source_check(&src, &t.engines).expect("check"); + assert!(check.peek().is_none(), "CMD-08 open waited for the check"); + assert!( + t.try_inspect(&slow, &info.fingerprint, 0).is_ok(), + "CMD-08 inspect meanwhile" + ); + assert!( + check.peek().is_none(), + "CMD-08 inspect waited for the check" + ); + std::fs::write(&gate, b"").expect("gate"); + let waited = check.wait(None).expect("check result"); + assert!( + matches!(waited, SourceCheck::Clean | SourceCheck::Benign(_)), + "{waited:?}" + ); + // A non-benign warning elsewhere in the file: inspect works, the first preview refuses. + let repair = t.file("repair.pdf", &needs_repair_doc()); + let (fp, dto) = t.page(&repair, 0); + let run = dto + .runs + .iter() + .find(|r| r.text == "Good line") + .expect("run"); + let e = service::preview_edits( + &t.cache, + &t.engines, + &t.temp_root(), + &repair.to_string_lossy(), + &fp, + 0, + &[edit_in( + &run.id, + "Good line", + "Good lines", + SourceTextStyleIn::default(), + )], + ) + .err() + .expect("PDF_NEEDS_REPAIR"); + assert_eq!(e.code, "PDF_NEEDS_REPAIR", "{e}"); + assert!( + e.details + .as_deref() + .unwrap_or_default() + .contains("EOF while reading token"), + "{:?}", + e.details + ); +} + +/// A real page's DTO has exactly the contract sample's key paths (DTO-01's link to real data). +#[test] +fn dto_01b_a_real_page_has_the_contract_keys() { + let Some(t) = E2e::new("dto_01b") else { + return; + }; + let p = t.file("keys.pdf", &fx::indd()); + let (_, dto) = t.page(&p, 0); + let real = serde_json::to_value(&dto).expect("json"); + let sample = &contract()["PageText"]; + let (mut a, mut b) = (BTreeSet::new(), BTreeSet::new()); + key_paths(&real, "", &mut a); + key_paths(sample, "", &mut b); + assert_eq!(a, b, "DTO-01b key paths"); + let (mut a, mut b) = (BTreeSet::new(), BTreeSet::new()); + kind_paths(&real, "", &mut a); + kind_paths(sample, "", &mut b); + let drift: Vec<_> = a.difference(&b).collect(); + assert!( + drift.is_empty(), + "DTO-01b value kinds not in the contract: {drift:?}" + ); + let preview = t.preview( + &p, + 0, + &[( + "visible layer", + "visible layers", + SourceTextStyleIn::default(), + )], + ); + let real = serde_json::to_value(&preview).expect("json"); + let (mut a, mut b) = (BTreeSet::new(), BTreeSet::new()); + key_paths(&real, "", &mut a); + key_paths(&contract()["TextPreview"], "", &mut b); + assert!(a.is_subset(&b), "DTO-01b preview keys {a:?} ⊄ {b:?}"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_e2e.rs b/src-tauri/src/pdf_engine/text_edit/tests_e2e.rs new file mode 100644 index 0000000..96c8c7f --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_e2e.rs @@ -0,0 +1,490 @@ +//! T5 end-to-end tests (SPEC §E.7, qpdf + Poppler required): edits are made the way the editor +//! makes them — `service::open_source` and `service::inspect_page` give the run ids — and saved +//! through the real Edit PDF export (`edit_overlay::export_edit_pdf_with_check_exe`, real qpdf +//! runner). Every successful save is checked on the **published** file: `qpdf --check` exit 0, +//! Poppler extracts the new text and not the old one, unedited pages keep their decoded content +//! (as page parts or, after an overlay, as qpdf's wrapper Form), no sibling `.offpdf-*.pdf.tmp`, +//! the work folder removed, and every source byte-identical to before. +//! +//! `fonts.rs` E2E-01…06, `save.rs` E2E-07…09, 17, 18, 22, `overlay.rs` E2E-10a…f, 11 and VO-01's +//! end-to-end half, `refusals.rs` E2E-12…16, 19…21 and GATE-25, `preview.rs` PREV-01…07. + +mod fonts; +mod join; +mod memory; +mod overlay; +pub(crate) mod preview; +pub(crate) mod refusals; +pub(crate) mod save; + +use crate::error::AppError; +use crate::models::PageGroup; +use crate::pdf_engine::edit_forms::FormValue; +use crate::pdf_engine::edit_overlay::{export_edit_pdf_with_check_exe, EditDocumentIn}; +use crate::pdf_engine::qpdf; +use crate::pdf_engine::text_edit::cache::TextEditCache; +use crate::pdf_engine::text_edit::content::{page_content, qpdf_join}; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget}; +use crate::pdf_engine::text_edit::dto::{PageTextDto, TextRunDto, TextSourceDto}; +use crate::pdf_engine::text_edit::engines::{Engines, RunOpts}; +use crate::pdf_engine::text_edit::poppler::pdftotext_words; +use crate::pdf_engine::text_edit::service; +use crate::pdf_engine::text_edit::snapshot::{fnv1a_u64, read_verification_snapshot}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use lopdf::{Document, Object, ObjectId}; +use serde_json::{json, Value}; +use std::cell::Cell; +use std::path::{Path, PathBuf}; +use std::sync::atomic::AtomicBool; + +const VERIFY_CAP: u64 = 4 << 30; + +/// Extra inputs of one save. +#[derive(Default)] +pub(crate) struct SaveOpts<'a> { + pub cancel: Option<&'a AtomicBool>, + pub form_values: Vec, + pub flatten_form: bool, + pub flatten_annotations: bool, + /// Called with each qpdf argv the pipeline runs, before it runs. + pub on_run: Option<&'a dyn Fn(&[String])>, + /// Destination (default: a fresh `out-N/saved.pdf`). + pub dest: Option, +} + +/// A save's published file (the sources were checked unchanged, nothing was left behind). +pub(crate) struct Saved { + pub path: PathBuf, + pub warnings: Vec, +} + +pub(crate) struct E2e { + pub scratch: Scratch, + pub engines: Engines, + pub cache: TextEditCache, + seq: Cell, +} + +/// The scratch folder (with `temp/textedit/`) is deleted right after this: every background +/// `qpdf --check` the harness started must have read its input by then, or a check that lost its +/// input would decide what a later test of the same bytes sees (review-T5 M1, the e2e_05/cmd_07 +/// collision). +impl Drop for E2e { + fn drop(&mut self) { + self.cache.wait_checks(); + } +} + +pub(crate) fn font_path() -> PathBuf { + Path::new(env!("CARGO_MANIFEST_DIR")).join("resources/fonts/NotoSans-Regular.ttf") +} + +impl E2e { + /// `None` (after `skip:`) when an engine is missing; fails with `OFFPDF_REQUIRE_ENGINES=1`. + pub(crate) fn new(test: &str) -> Option { + let engines = engines_or_skip(test)?; + Some(E2e { + scratch: Scratch::new(test), + engines, + cache: TextEditCache::default(), + seq: Cell::new(0), + }) + } + + fn next(&self) -> usize { + let n = self.seq.get() + 1; + self.seq.set(n); + n + } + + pub(crate) fn temp_root(&self) -> PathBuf { + self.scratch.path("temp") + } + + pub(crate) fn file(&self, name: &str, bytes: &[u8]) -> PathBuf { + self.scratch.write(name, bytes) + } + + pub(crate) fn try_open(&self, path: &Path) -> Result { + service::open_source( + &self.cache, + &self.engines, + &self.temp_root(), + &path.to_string_lossy(), + ) + } + + pub(crate) fn open(&self, path: &Path) -> TextSourceDto { + self.try_open(path) + .unwrap_or_else(|e| panic!("open {}: {e} {:?}", path.display(), e.details)) + } + + pub(crate) fn try_inspect( + &self, + path: &Path, + fp: &str, + page: u32, + ) -> Result { + service::inspect_page( + &self.cache, + &self.engines, + &self.temp_root(), + &path.to_string_lossy(), + fp, + page, + ) + } + + /// Opens `path` and inspects `page`. + pub(crate) fn page(&self, path: &Path, page: u32) -> (String, PageTextDto) { + let fp = self.open(path).fingerprint; + let text = self + .try_inspect(path, &fp, page) + .unwrap_or_else(|e| panic!("inspect: {e} {:?}", e.details)); + (fp, text) + } + + /// The run whose text is `text` on `page`, with the fingerprint. + pub(crate) fn run(&self, path: &Path, page: u32, text: &str) -> (String, TextRunDto) { + let (fp, dto) = self.page(path, page); + let run = dto + .runs + .iter() + .find(|r| r.text == text) + .cloned() + .unwrap_or_else(|| { + let all: Vec<_> = dto.runs.iter().map(|r| (&r.text, r.reason)).collect(); + panic!("no run {text:?} on page {page}: {all:?}") + }); + (fp, run) + } + + /// The `sourceText` export object (the frontend's `toExportDocument` shape) changing the run + /// `old` on source page `source_page` (dest page `dest`) to `new` with `style`. + pub(crate) fn edit_styled( + &self, + path: &Path, + dest: u32, + source_page: u32, + old: &str, + new: &str, + style: Value, + ) -> Value { + let (fp, run) = self.run(path, source_page, old); + assert!( + run.editable, + "run {old:?} is not editable: {:?}", + run.reason + ); + json!({ + "id": format!("st-{}", self.next()), + "kind": "sourceText", + "pageIndex": dest, + "rect": { "x": run.rect.x, "y": run.rect.y, "w": run.rect.w, "h": run.rect.h }, + "locked": true, + "runId": run.id, + "sourceFingerprint": fp, + "sourcePageIndex": source_page, + "originalText": run.text, + "text": new, + "style": style, + }) + } + + /// `edit`, or `None` when the page offers no editable run `old`. + pub(crate) fn try_edit( + &self, + path: &Path, + dest: u32, + source_page: u32, + old: &str, + new: &str, + ) -> Option { + let (_, dto) = self.page(path, source_page); + dto.runs + .iter() + .any(|r| r.text == old && r.editable) + .then(|| self.edit(path, dest, source_page, old, new)) + } + + pub(crate) fn edit( + &self, + path: &Path, + dest: u32, + source_page: u32, + old: &str, + new: &str, + ) -> Value { + self.edit_styled(path, dest, source_page, old, new, json!({})) + } + + pub(crate) fn save( + &self, + groups: &[(&Path, &str)], + objects: Vec, + ) -> Result { + self.save_with(groups, objects, SaveOpts::default()) + } + + /// Saves through the real export (the qpdf runner of `edit_pdf_overlays`), then checks what + /// every save must leave behind: unchanged sources, no sibling temp file, no work folder, no + /// cache folder created by Save. + pub(crate) fn save_with( + &self, + groups: &[(&Path, &str)], + objects: Vec, + opts: SaveOpts<'_>, + ) -> Result { + let n = self.next(); + let doc: EditDocumentIn = + serde_json::from_value(json!({ "version": 1, "objects": objects })) + .unwrap_or_else(|e| panic!("export document: {e}")); + let page_groups: Vec = groups + .iter() + .map(|(p, pages)| PageGroup { + path: p.to_string_lossy().into_owned(), + pages: pages.to_string(), + }) + .collect(); + let before: Vec = groups.iter().map(|(p, _)| file_hash(p)).collect(); + let dest = opts + .dest + .clone() + .unwrap_or_else(|| self.scratch.path(&format!("out-{n}")).join("saved.pdf")); + let out_dir = dest.parent().map(Path::to_path_buf).expect("dest folder"); + std::fs::create_dir_all(&out_dir).expect("out dir"); + let dest_before = dest.is_file().then(|| file_hash(&dest)); + let work = self.scratch.path(&format!("work-{n}")); + std::fs::create_dir_all(&work).expect("work dir"); + let textedit_before = self.temp_root().join("textedit").exists(); + let qpdf = qpdf::resolve_qpdf_standalone(); + let on_run = opts.on_run; + let result = export_edit_pdf_with_check_exe( + &page_groups, + &dest.to_string_lossy(), + &doc, + &font_path(), + &work, + &format!("e2e-{n}"), + opts.cancel, + &qpdf, + None, + None, + &[], + &opts.form_values, + opts.flatten_form, + opts.flatten_annotations, + |args| { + if let Some(f) = on_run { + f(args); + } + run_qpdf(&qpdf, args) + }, + ); + // `edit_pdf_overlays` removes the job's work folder whatever happened. + let _ = std::fs::remove_dir_all(&work); + for ((p, _), h) in groups.iter().zip(&before) { + assert_eq!(file_hash(p), *h, "source {} changed", p.display()); + } + let leftovers: Vec = std::fs::read_dir(&out_dir) + .expect("out dir") + .flatten() + .map(|e| e.file_name().to_string_lossy().into_owned()) + .filter(|name| name.ends_with(".tmp")) + .collect(); + assert!( + leftovers.is_empty(), + "sibling temp files left: {leftovers:?}" + ); + assert!(!work.exists(), "work folder left"); + assert_eq!( + self.temp_root().join("textedit").exists(), + textedit_before, + "Save must not use the editor's cache" + ); + match result { + Ok((_, warnings)) => { + assert!(dest.is_file(), "published file missing"); + Ok(Saved { + path: dest, + warnings, + }) + } + Err(e) => { + assert_eq!( + dest.is_file().then(|| file_hash(&dest)), + dest_before, + "a failed save changed {}", + dest.display() + ); + Err(e) + } + } + } + + /// `qpdf --check` of the published file exits 0. + pub(crate) fn assert_checks_clean(&self, pdf: &Path) { + let out = std::process::Command::new(&self.engines.qpdf) + .arg("--check") + .arg(pdf) + .output() + .expect("qpdf --check"); + assert_eq!( + out.status.code(), + Some(0), + "qpdf --check {}: {}{}", + pdf.display(), + String::from_utf8_lossy(&out.stdout), + String::from_utf8_lossy(&out.stderr) + ); + } + + /// Poppler's words of page `page` (0-based), joined without whitespace. + pub(crate) fn words(&self, pdf: &Path, page: u32) -> String { + pdftotext_words(&self.engines, pdf, page + 1, &RunOpts::default()) + .unwrap_or_else(|e| panic!("pdftotext {}: {e}", pdf.display())) + .iter() + .map(|w| w.text.as_str()) + .collect::() + .chars() + .filter(|c| !c.is_whitespace()) + .collect() + } + + /// The edited line reads `new` in the published file and `old` occurs fewer times than on + /// the source page (when it is not part of `new`). + pub(crate) fn assert_changed( + &self, + src: &Path, + src_page: u32, + out: &Path, + dest: u32, + old: &str, + new: &str, + ) { + let compact = |s: &str| s.chars().filter(|c| !c.is_whitespace()).collect::(); + let (old, new) = (compact(old), compact(new)); + let after = self.words(out, dest); + if !new.is_empty() { + assert!( + after.contains(&new), + "page {}: {new:?} not in {after:?}", + dest + 1 + ); + } + if !old.is_empty() && !new.contains(&old) { + let before = self.words(src, src_page); + assert!( + after.matches(&old).count() < before.matches(&old).count(), + "page {}: old text {old:?} still extracted: {after:?}", + dest + 1 + ); + } + } + + /// Published page `dest` holds source page `src_page`'s decoded content unchanged: as its + /// parts, or as the data of the overlay wrapper Form qpdf made of them. + pub(crate) fn assert_unedited(&self, src: &Path, src_page: u32, out: &Path, dest: u32) { + let parts = page_parts(src, src_page); + let plain = parts.concat(); + let joined = qpdf_join(&parts.iter().map(Vec::as_slice).collect::>()); + let held = page_holdings(out, dest); + assert!( + held.iter().any(|h| *h == plain || *h == joined), + "page {} of {} does not hold page {} of {} unchanged", + dest + 1, + out.display(), + src_page + 1, + src.display() + ); + } + + /// `assert_checks_clean` + `assert_changed` for each `(src, src_page, dest, old, new)`. + pub(crate) fn assert_saved(&self, out: &Path, changes: &[(&Path, u32, u32, &str, &str)]) { + self.assert_checks_clean(out); + for (src, sp, dest, old, new) in changes { + self.assert_changed(src, *sp, out, *dest, old, new); + } + } +} + +pub(crate) fn file_hash(path: &Path) -> u64 { + fnv1a_u64(&std::fs::read(path).unwrap_or_else(|e| panic!("{}: {e}", path.display()))) +} + +pub(crate) fn run_qpdf(qpdf: &Path, args: &[String]) -> Result<(), AppError> { + let out = std::process::Command::new(qpdf) + .args(args) + .output() + .map_err(|e| AppError::io("qpdf failed to start", e))?; + match out.status.code() { + Some(0) | Some(3) => Ok(()), + _ => Err(AppError::engine_failed( + String::from_utf8_lossy(&out.stderr).into_owned(), + )), + } +} + +/// Decoded content parts of page `page` (0-based) of `pdf`. +pub(crate) fn page_parts(pdf: &Path, page: u32) -> Vec> { + let snap = read_verification_snapshot(pdf, VERIFY_CAP).unwrap_or_else(|e| panic!("{e}")); + let id = *snap.pages.get(page as usize).expect("page"); + let mut budget = DecodeBudget::new(1 << 30); + let content = page_content(&snap.doc, id, &mut budget) + .unwrap_or_else(|r| panic!("page content: {}", r.as_str())); + (0..content.parts.len()) + .map(|i| content.part_bytes(i).to_vec()) + .collect() +} + +/// What page `page` of `pdf` draws: its parts concatenated, and the decoded data of every Form +/// XObject of its resources (qpdf's overlay wrapper holds the original content there). +pub(crate) fn page_holdings(pdf: &Path, page: u32) -> Vec> { + let snap = read_verification_snapshot(pdf, VERIFY_CAP).unwrap_or_else(|e| panic!("{e}")); + let id = *snap.pages.get(page as usize).expect("page"); + let mut out = vec![page_parts(pdf, page).concat()]; + out.extend(page_forms(&snap.doc, id)); + out +} + +fn resolve<'a>(doc: &'a Document, o: &'a Object) -> Option<&'a Object> { + match o { + Object::Reference(r) => doc.get_object(*r).ok(), + other => Some(other), + } +} + +fn page_forms(doc: &Document, page: ObjectId) -> Vec> { + let Ok((inline, inherited)) = doc.get_page_resources(page) else { + return Vec::new(); + }; + let mut dicts: Vec<&lopdf::Dictionary> = inline.into_iter().collect(); + dicts.extend( + inherited + .iter() + .filter_map(|id| doc.get_dictionary(*id).ok()), + ); + let mut forms = Vec::new(); + for res in dicts { + let Some(xobjects) = res + .get(b"XObject") + .ok() + .and_then(|o| resolve(doc, o)) + .and_then(|o| o.as_dict().ok()) + else { + continue; + }; + for (_, v) in xobjects.iter() { + let Some(Object::Stream(s)) = resolve(doc, v) else { + continue; + }; + if s.dict.get(b"Subtype").and_then(Object::as_name).ok() == Some(b"Form".as_slice()) { + let mut budget = DecodeBudget::new(1 << 30); + if let Ok(data) = decode_stream(s, 1 << 30, &mut budget) { + forms.push(data); + } + } + } + } + forms +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_e2e/fonts.rs b/src-tauri/src/pdf_engine/text_edit/tests_e2e/fonts.rs new file mode 100644 index 0000000..721781a --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_e2e/fonts.rs @@ -0,0 +1,213 @@ +//! E2E-01…06: one save per font class and page geometry, through the real export. + +use super::{page_parts, E2e}; +use crate::pdf_engine::text_edit::engines::RunOpts; +use crate::pdf_engine::text_edit::poppler::pdftotext_words; +use crate::pdf_engine::text_edit::reasons::{TextReason, ORIGINAL_UNCHANGED}; +use crate::pdf_engine::text_edit::testkit::producers::{self as fx, helvetica_page}; +use serde_json::json; +use std::path::Path; + +fn corpus(name: &str) -> Vec { + let path = Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../fixtures/source-edit") + .join(name); + std::fs::read(&path).unwrap_or_else(|e| panic!("{}: {e}", path.display())) +} + +/// Poppler's box of the word `word` on page `page` (0-based). +pub(crate) fn word_box(t: &E2e, pdf: &Path, page: u32, word: &str) -> [f64; 4] { + let words = pdftotext_words(&t.engines, pdf, page + 1, &RunOpts::default()) + .unwrap_or_else(|e| panic!("pdftotext: {e}")); + let w = words.iter().find(|w| w.text == word).unwrap_or_else(|| { + panic!( + "no word {word:?} in {:?}", + words.iter().map(|w| &w.text).collect::>() + ) + }); + [w.x0, w.y0, w.x1, w.y1] +} + +pub(crate) fn assert_same_box(a: [f64; 4], b: [f64; 4], what: &str) { + for (x, y) in a.iter().zip(&b) { + assert!((x - y).abs() <= 0.05, "{what} moved: {a:?} → {b:?}"); + } +} + +/// One save of `edits` (`(old, new)`) on page 1 of `pdf`, checked on the published file. +fn save_one(t: &E2e, name: &str, pdf: &[u8], edits: &[(&str, &str)]) -> std::path::PathBuf { + let src = t.file(&format!("{name}.pdf"), pdf); + let objects = edits + .iter() + .map(|(old, new)| t.edit(&src, 0, 0, old, new)) + .collect(); + let saved = t + .save(&[(&src, "1-z")], objects) + .unwrap_or_else(|e| panic!("{name}: save failed: {e} {:?}", e.details)); + let changes: Vec<_> = edits + .iter() + .map(|(o, n)| (src.as_path(), 0, 0, *o, *n)) + .collect(); + t.assert_saved(&saved.path, &changes); + saved.path +} + +#[test] +fn e2e_01_corpus_text_tj_hi_becomes_hello() { + let Some(t) = E2e::new("e2e_01") else { + return; + }; + save_one(&t, "text-tj", &corpus("text-tj.pdf"), &[("Hi", "Hello")]); +} + +#[test] +fn e2e_02_kerned_line_and_a_follower_on_the_same_line_stay_in_place() { + let Some(t) = E2e::new("e2e_02") else { + return; + }; + save_one( + &t, + "text-tj-kerned", + &corpus("text-tj-kerned.pdf"), + &[("Hi", "Hit")], + ); + // B1: the TJ kern counts in the pen, so the tail and a follower run keep their places. + let pdf = helvetica_page( + b"BT /F1 10 Tf 72 700 Td [(AB) -500 (CD)] TJ (tail) Tj ET \ + BT /F1 10 Tf 300 700 Td (Follower) Tj ET", + ); + let src = t.file("kerned-follower.pdf", &pdf); + let before = word_box(&t, &src, 0, "Follower"); + let out = save_one( + &t, + "kerned-follower-2", + &pdf, + &[("AB CDtail", "AB CDtails")], + ); + assert_same_box(before, word_box(&t, &out, 0, "Follower"), "E2E-02 follower"); +} + +#[test] +fn e2e_03_word_line_keeps_its_kerns_byte_for_byte() { + let Some(t) = E2e::new("e2e_03") else { + return; + }; + let out = save_one(&t, "word", &fx::word(), &[("Invoice 2026", "Invoice 2027")]); + let content = page_parts(&out, 0).concat(); + let text = String::from_utf8_lossy(&content); + // "Inv" + the original kern bytes 12 + "oice…" (PLAN-04's kept prefix). + assert!( + text.contains("<496E76> 12 <"), + "E2E-03 kern not kept: {text}" + ); + assert!(!text.contains("2026"), "E2E-03 old literal left: {text}"); +} + +#[test] +fn e2e_04_word_turkish_line_across_sibling_fonts_and_a_missing_glyph() { + let Some(t) = E2e::new("e2e_04") else { + return; + }; + save_one( + &t, + "word-tr", + &fx::word_tr(), + &[("Sağlık Bakanlığı Raporu", "Sağlığı Bakanlığı Raporu")], + ); + let src = t.file("word-tr-y.pdf", &fx::word_tr()); + let obj = t.edit( + &src, + 0, + 0, + "Sağlık Bakanlığı Raporu", + "Yağlık Bakanlığı Raporu", + ); + let e = t + .save(&[(&src, "1-z")], vec![obj]) + .err() + .expect("Y is not in the subset"); + assert_eq!(e.code, "GLYPH_MISSING", "E2E-04: {e}"); + assert_eq!(e.message, "On page 1, the document's font can't draw: Y"); + assert!(e + .suggestion + .as_deref() + .unwrap_or_default() + .ends_with(ORIGINAL_UNCHANGED)); +} + +#[test] +fn e2e_05_every_producer_font_class_saves() { + let Some(t) = E2e::new("e2e_05") else { + return; + }; + save_one(&t, "libre", &fx::libre(), &[("Libre text", "Libre exit")]); + save_one(&t, "quartz", &fx::quartz(), &[("Quartz", "Quart")]); + // Skia: the flipped line, the synthetic-bold Tr 2 line and the sheared line in one save. + save_one( + &t, + "skia", + &fx::skia(), + &[("Chrome", "Comet"), ("Bold", "Bolder"), ("Italic", "Ital")], + ); + // pdfTeX: the new space is written as a kern (the subset has no space glyph). + save_one( + &t, + "pdftex", + &fx::pdftex(), + &[("Hello World", "Hello Wet World")], + ); + save_one(&t, "xetex", &fx::xetex(), &[("XeTeX", "TeX")]); + save_one( + &t, + "indd", + &fx::indd(), + &[("visible layer", "visible layers")], + ); + save_one(&t, "std14", &fx::std14(), &[("Times line", "Times lines")]); + save_one( + &t, + "nonemb", + &fx::nonemb(), + &[("Not embedded", "Not embed")], + ); + save_one( + &t, + "libre-cff", + &fx::libre_cff(), + &[("Hello Hello", "Hello Hole")], + ); + // InDesign's hidden layer is refused at inspect and at Save. + let src = t.file("indd-hidden.pdf", &fx::indd()); + let (fp, run) = t.run(&src, 0, "hidden layer"); + assert!(!run.editable); + assert_eq!(run.reason, Some(TextReason::OptionalContent)); + let obj = json!({ + "id": "hidden", "kind": "sourceText", "pageIndex": 0, + "rect": { "x": run.rect.x, "y": run.rect.y, "w": run.rect.w, "h": run.rect.h }, + "locked": true, "runId": run.id, "sourceFingerprint": fp, "sourcePageIndex": 0, + "originalText": run.text, "text": "hidden layers", "style": {}, + }); + let e = t.save(&[(&src, "1-z")], vec![obj]).err().expect("refused"); + assert_eq!(e.code, "TEXT_EDIT_REFUSED", "{e}"); +} + +#[test] +fn e2e_06_rotated_and_cropped_pages() { + let Some(t) = E2e::new("e2e_06") else { + return; + }; + for angle in [90, 180, 270] { + save_one( + &t, + &format!("rotate-{angle}"), + &fx::rotated(angle, true), + &[("Rotated page", "Rotated pages")], + ); + } + save_one( + &t, + "crop", + &fx::cropped_offset(), + &[("Cropped page", "Cropped pages")], + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_e2e/join.rs b/src-tauri/src/pdf_engine/text_edit/tests_e2e/join.rs new file mode 100644 index 0000000..edb6287 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_e2e/join.rs @@ -0,0 +1,152 @@ +//! review-T5 H1: qpdf's overlay join must never change what a page shows. A page whose content +//! parts meet inside a comment or a string reads differently once qpdf joins them with a `\n` +//! (Poppler reads the parts as one stream: the comment ran on and hid the next part). Every Save +//! that wraps such a page — a stamp alone, or a stamp with text changes — must be refused, while +//! split pages whose parts meet between tokens (E2E-10c) still save. + +use super::E2e; +use crate::pdf_engine::text_edit::testkit::producers::{DocBuilder, PageSpec, HELVETICA}; +use serde_json::{json, Value}; + +const COMMENT: [&[u8]; 2] = [ + b"BT /F1 12 Tf 72 720 Td (Visible) Tj ET\n% note", + b"BT /F1 12 Tf 72 700 Td (Secret) Tj ET", +]; +const MID_STRING: [&[u8]; 2] = [ + b"BT /F1 12 Tf 72 720 Td (Visible) Tj ET BT /F1 12 Tf 72 700 Td (Hel", + b"lo) Tj ET", +]; + +/// The reviewer's two pages: (name, parts, Poppler's words of the source page). +const UNSAFE: [(&str, [&[u8]; 2], &str); 2] = [ + ("comment", COMMENT, "Visible"), + ("mid-string", MID_STRING, "VisibleHello"), +]; + +fn stamp(page: u32) -> Value { + json!({ + "kind": "text", "pageIndex": page, + "rect": { "x": 300.0, "y": 120.0, "w": 200.0, "h": 30.0 }, + "content": "Stamp", "fontSize": 14.0, "color": "#c71c1c", "align": null, "opacity": 1.0, + }) +} + +/// A file whose pages are `pages` (each a list of content parts), all with Helvetica as /F1. +fn doc(pages: &[&[&[u8]]]) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + for parts in pages { + d.page(PageSpec::parts(parts, &res)); + } + d.build() +} + +#[test] +fn h1_stamp_only_saves_of_unsafe_split_pages_are_refused() { + let Some(t) = E2e::new("h1_stamp_only") else { + return; + }; + for (name, parts, words) in UNSAFE { + let src = t.file(&format!("{name}.pdf"), &doc(&[&parts])); + assert_eq!(t.words(&src, 0), words, "{name}: what the source shows"); + match t.save(&[(&src, "1-z")], vec![stamp(0)]) { + Err(e) => assert_eq!(e.code, "INVALID_OUTPUT", "{name}: {e} {:?}", e.details), + Ok(saved) => panic!( + "{name}: a stamp-only save was published; Poppler now reads {:?}", + t.words(&saved.path, 0) + ), + } + } + // The legitimate split (E2E-10c's shape, parts meeting between operators) still saves. + let ok = doc(&[&[ + b"BT /F1 12 Tf 72 720 Td (Split zero) Tj ET", + b"BT /F1 12 Tf 72 700 Td (Split one) Tj ET", + ]]); + let src = t.file("ok.pdf", &ok); + let saved = t + .save(&[(&src, "1-z")], vec![stamp(0)]) + .unwrap_or_else(|e| panic!("token-boundary split: {e} {:?}", e.details)); + assert!(t.words(&saved.path, 0).contains("SplitzeroSplitone")); +} + +#[test] +fn h1_text_edit_saves_with_unsafe_split_pages_are_refused() { + let Some(t) = E2e::new("h1_text_edit") else { + return; + }; + let line: &[&[u8]] = &[b"BT /F1 12 Tf 72 720 Td (Edited line) Tj ET"]; + for (name, parts, _) in UNSAFE { + // A text change on page 1 and a stamp; page 2 is the unsafe split page, unedited. + let src = t.file(&format!("{name}-2.pdf"), &doc(&[line, &parts])); + let objects = vec![t.edit(&src, 0, 0, "Edited line", "Changed line"), stamp(0)]; + match t.save(&[(&src, "1-z")], objects) { + Err(e) => assert_eq!(e.code, "INVALID_OUTPUT", "{name}: {e} {:?}", e.details), + Ok(saved) => panic!( + "{name}: published; Poppler reads page 2 as {:?}", + t.words(&saved.path, 1) + ), + } + // A text change on the split page itself and a stamp: refused at inspect (the line is + // not offered), by Phase B or by #34 — never published. + let src = t.file(&format!("{name}-1.pdf"), &doc(&[&parts])); + let Some(edit) = t.try_edit(&src, 0, 0, "Visible", "Visibly") else { + continue; + }; + if let Ok(saved) = t.save(&[(&src, "1-z")], vec![edit, stamp(0)]) { + panic!( + "{name}: published; Poppler reads {:?}", + t.words(&saved.path, 0) + ); + } + } +} + +/// review-final MEDIUM-2: the join rule needs no lex of the page. A stamp-only save of a +/// two-part page with a neutral boundary publishes even when the strict lexer refuses the page +/// (more than 250,000 ops; a trailing operand), as it did before review-T5 H1. +#[test] +fn medium2_stamp_only_saves_of_pages_the_strict_lexer_refuses_still_publish() { + let Some(t) = E2e::new("medium2_stamp_only") else { + return; + }; + let mut dense = b"q ".to_vec(); + dense.extend_from_slice(&b"0 0 m 1 1 l S ".repeat(90_000)); + dense.extend_from_slice(b"Q"); + let text: &[u8] = b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"; + let pages: [(&str, [&[u8]; 2]); 2] = [ + ("270k-ops", [&dense, text]), + ("trailing-operand", [text, b"\nq Q 0\n"]), + ]; + for (name, parts) in pages { + let src = t.file(&format!("{name}.pdf"), &doc(&[&parts])); + let saved = t + .save(&[(&src, "1-z")], vec![stamp(0)]) + .unwrap_or_else(|e| panic!("{name}: {e} {:?}", e.details)); + assert_eq!(t.words(&saved.path, 0), "HelloStamp", "{name}"); + } +} + +/// review-final LOW-1: a part ending in a comment, the next starting a line, is one page for the +/// walker and the gate alike: its lines are editable, and stamp, text and text + stamp saves all +/// publish (before: every overlay save failed `INVALID_OUTPUT` or `SOURCE_EDIT_GATE_FAILED`). +#[test] +fn low1_a_comment_ending_a_part_before_a_new_line_saves_with_stamps() { + let Some(t) = E2e::new("low1_comment_line") else { + return; + }; + let parts: [&[u8]; 2] = [ + b"BT /F1 12 Tf 72 720 Td (Visible) Tj ET\n% note", + b"\nBT /F1 12 Tf 72 700 Td (Hello) Tj ET", + ]; + let src = t.file("comment-line.pdf", &doc(&[&parts])); + let stamped = t + .save(&[(&src, "1-z")], vec![stamp(0)]) + .unwrap_or_else(|e| panic!("stamp only: {e} {:?}", e.details)); + assert_eq!(t.words(&stamped.path, 0), "VisibleHelloStamp"); + let edit = t.edit(&src, 0, 0, "Hello", "Help"); + let both = t + .save(&[(&src, "1-z")], vec![edit, stamp(0)]) + .unwrap_or_else(|e| panic!("text + stamp: {e} {:?}", e.details)); + assert_eq!(t.words(&both.path, 0), "VisibleHelpStamp"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_e2e/memory.rs b/src-tauri/src/pdf_engine/text_edit/tests_e2e/memory.rs new file mode 100644 index 0000000..c362f9e --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_e2e/memory.rs @@ -0,0 +1,331 @@ +//! review-T4 M-1: memory of the plan, preview and Save paths on large pages. A Save with several +//! large edited pages must not hold every page's model from planning to Phase A: two extra pages +//! add less than 1.5 page models to the peak of a one-page Save (before the fix, each added its +//! whole model and an unshared copy of every record's state in its proof). review-final MEDIUM-3: +//! a preview of a near-budget page stays below 256 MiB for the whole process. + +use super::E2e; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::limits::PREVIEW_SHARED_MODEL_MAX; +use crate::pdf_engine::text_edit::rewrite::{plan_page, SourceTextStyleIn, TextEditIn}; +use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel}; +use crate::pdf_engine::text_edit::service::preview_edits; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::producers::{DocBuilder, PageSpec, HELVETICA}; +use crate::pdf_engine::text_edit::testkit::{ + child_mode, child_value, process_peak, run_child_test, thread_peak, +}; + +const MIB: usize = 1 << 20; +/// `(a)Tj` ops per page: a page model of a few tens of MiB, well above the kept-model budget of a +/// small file (16 MiB). +const OPS: usize = 30_000; + +/// `pages` pages, each "Hello" plus `OPS` one-glyph shows (review-T4 M-1's near-budget shape). +fn dense_doc(pages: usize) -> Vec { + let mut content = String::from("BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 1 Tf 0 20 Td "); + content.push_str(&"(a)Tj ".repeat(OPS)); + content.push_str("ET"); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + for _ in 0..pages { + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f} 0 R >>"), + )); + } + d.build() +} + +/// Peak bytes held on this thread while saving "Hello" → "Help" on every page of `pages` pages. +fn save_peak(t: &E2e, pages: usize) -> usize { + let src = t.file(&format!("dense-{pages}.pdf"), &dense_doc(pages)); + let edits = (0..pages as u32) + .map(|p| t.edit(&src, p, p, "Hello", "Help")) + .collect(); + t.cache.release(&src); + let (saved, peak) = thread_peak(|| t.save(&[(&src, "1-z")], edits)); + let saved = saved.unwrap_or_else(|e| panic!("{pages}-page save: {e} {:?}", e.details)); + for p in 0..pages as u32 { + t.assert_changed(&src, p, &saved.path, p, "Hello", "Help"); + } + peak +} + +#[test] +fn save_of_several_large_pages_holds_one_page_model_at_a_time() { + let Some(t) = E2e::new("m1_save_models") else { + return; + }; + let (src, model_bytes) = { + let src = t.file("dense-model.pdf", &dense_doc(1)); + let fp = t.open(&src).fingerprint; + let source = t + .cache + .get(&src, &t.temp_root(), &fp, &t.engines) + .expect("source"); + let model = t.cache.page(&source, 0).expect("model"); + (src, model.walk.model_bytes) + }; + t.cache.release(&src); + assert!( + model_bytes > 16 * MIB, + "the page model ({} MiB) must exceed the kept-model budget", + model_bytes / MIB + ); + let one = save_peak(&t, 1); + let three = save_peak(&t, 3); + println!( + "M-1 save peaks: model {} MiB, 1 page {} MiB, 3 pages {} MiB", + model_bytes / MIB, + one / MIB, + three / MIB + ); + // The two extra pages add their plans and proofs; keeping their models alone would add two + // whole models. + assert!( + three.saturating_sub(one) < model_bytes * 3 / 2, + "a 3-page save peaks at {} MiB, a 1-page save at {} MiB: models are stacked ({} MiB each)", + three / MIB, + one / MIB, + model_bytes / MIB + ); +} + +/// One page: "Hello" plus 100,000 `(a)Tj` (review-T4 M-1's near-budget page, an 84 MiB model). +fn near_budget_doc() -> Vec { + let mut content = String::from("BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 1 Tf 0 20 Td "); + content.push_str(&"(a)Tj ".repeat(100_000)); + content.push_str("ET"); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +fn hello_to_help(model: &PageModel) -> TextEditIn { + let run = model + .runs + .iter() + .find(|r| r.text == "Hello") + .expect("Hello"); + TextEditIn { + run_id: run.id.clone(), + original_text: "Hello".into(), + text: "Help".into(), + style: SourceTextStyleIn::default(), + } +} + +/// review-T4 M-1 (plan half), the reviewer's probe: the verification walks no longer stack (the +/// after walk is reduced to what its caller keeps before the probe walk), so `plan_page` adds +/// about one walk to the model it plans on (before: two). +#[test] +fn plan_of_a_near_budget_page_does_not_stack_walks() { + let Some(t) = E2e::new("m1_plan_walks") else { + return; + }; + let pdf = near_budget_doc(); + let path = t.file("near-budget.pdf", &pdf); + let ctx = SnapshotContext::new(snapshot_from_bytes(&path, pdf, None).expect("snapshot")); + let model = build_page_model(&ctx, 0, None).expect("model"); + let edit = hello_to_help(&model); + let (out, plan_peak) = thread_peak(|| plan_page(&ctx, &model, std::slice::from_ref(&edit))); + let out = out.unwrap_or_else(|e| panic!("plan: {e}")); + assert!(out.plan.is_some(), "a plan"); + let model_bytes = model.walk.model_bytes; + println!( + "M-1 near-budget page: model {} MiB, plan_page +{} MiB", + model_bytes / MIB, + plan_peak / MIB + ); + assert!( + plan_peak < model_bytes * 3 / 2, + "plan_page holds {} MiB over a {} MiB model: the self-check walks stack", + plan_peak / MIB, + model_bytes / MIB + ); +} + +/// review-final MEDIUM-3 (§H R20, 256 MiB per file): the whole process — the page's model built +/// and cached as inspect leaves it, then a preview through the service — stays below 256 MiB. +/// The service hands the large model over (`page_to_release`) and the preview frees it once +/// planned, before the extracted page's model and Phase A's walks are built (before: 295 MiB, +/// the source model held throughout). Process-wide, so measured in a child process. +#[test] +fn preview_of_a_near_budget_page_stays_below_256_mib_for_the_process() { + let Some(_t) = E2e::new("m3_preview_parent") else { + return; + }; + let out = run_child_test( + "pdf_engine::text_edit::tests_e2e::memory::preview_process_peak_child", + "m3_preview", + ); + let (peak, model) = (child_value(&out, "PEAK"), child_value(&out, "MODEL")); + println!( + "M-3 preview: model {} MiB, process peak {} MiB", + model / MIB, + peak / MIB + ); + assert_eq!(child_value(&out, "PREVIEW_OK"), 1, "{out}"); + assert!( + model > PREVIEW_SHARED_MODEL_MAX, + "the page must exercise the hand-over: {out}" + ); + assert_eq!( + child_value(&out, "CACHED_AFTER"), + 0, + "a handed-over model leaves the cache" + ); + assert!( + peak < 256 * MIB, + "the process held {} MiB for a {} MiB model", + peak / MIB, + model / MIB + ); +} + +#[test] +#[ignore = "child process of preview_of_a_near_budget_page_stays_below_256_mib_for_the_process"] +fn preview_process_peak_child() { + if child_mode().as_deref() != Some("m3_preview") { + return; + } + let t = E2e::new("m3_preview_child").expect("engines"); + let src = t.file("near-budget.pdf", &near_budget_doc()); + let fp = t.open(&src).fingerprint; + let ((preview, model_bytes, cached_after), peak) = process_peak(|| { + let source = t + .cache + .get(&src, &t.temp_root(), &fp, &t.engines) + .expect("source"); + let model = t.cache.page(&source, 0).expect("model"); + let (edit, model_bytes) = (hello_to_help(&model), model.walk.model_bytes); + drop(model); + let preview = preview_edits( + &t.cache, + &t.engines, + &t.temp_root(), + &src.to_string_lossy(), + &fp, + 0, + &[edit], + ); + (preview, model_bytes, t.cache.page_model_bytes(&source)) + }); + let ok = preview.is_ok_and(|p| p.page_pdf.is_some()); + println!( + "PEAK={peak}\nMODEL={model_bytes}\nCACHED_AFTER={cached_after}\nPREVIEW_OK={}", + u8::from(ok) + ); +} + +/// About `pad_mib` MiB (an incompressible image in the resources) with two pages of "Hello" +/// plus `ops` one-glyph shows each. +fn large_file(pad_mib: usize, ops: usize) -> Vec { + let side = 1024usize; + let len = side * ((pad_mib << 20) / side); + let mut state: u64 = 0x9e37_79b9_7f4a_7c15; + let mut image = format!( + "<< /Type /XObject /Subtype /Image /Width {side} /Height {} /ColorSpace /DeviceGray \ + /BitsPerComponent 8 /Length {len} >>\nstream\n", + len / side + ) + .into_bytes(); + image.reserve(len + 16); + for _ in 0..len { + state ^= state << 13; + state ^= state >> 7; + state ^= state << 17; + image.push(state as u8); + } + image.extend_from_slice(b"\nendstream"); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let pad = d.add(image); + let res = format!("/Font << /F1 {f} 0 R >> /XObject << /Pad {pad} 0 R >>"); + let mut content = String::from("BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 1 Tf 0 20 Td "); + content.push_str(&"(a)Tj ".repeat(ops)); + content.push_str("ET"); + for _ in 0..2 { + d.page(PageSpec::new(content.as_bytes(), &res)); + } + d.build() +} + +/// review-final MEDIUM-1 (BENCH-02: 3 × file): a Save holds neither raw bytes once they are +/// parsed nor a source context next to kept models. Process-wide, in a child process, for both +/// sides of the kept-model budget (twice the file): +/// - models kept (40 MiB file, 2 × 12 MiB of models): the peak is Phase A's read of the edited +/// copy (twice the file) next to the models: 104 MiB (with the raw bytes kept: 143 MiB); +/// - context kept (16 MiB file, 2 × 23 MiB): the context, the staged read and the page rebuilt: +/// 95 MiB (with the first model kept next to the context: 118 MiB). +/// +/// The reviewer's 151 MiB file with 169 MiB of models (`v04/lastfix-logs`): 923 MiB = 6.1 × file +/// before, 475 MiB = 3.15 × file now (models + Phase A's read of the edited copy). +#[test] +fn save_of_a_large_file_holds_its_staged_read_and_models_only() { + let Some(_t) = E2e::new("m1_save_parent") else { + return; + }; + let out = run_child_test( + "pdf_engine::text_edit::tests_e2e::memory::save_process_peak_child", + "m1_save", + ); + for (shape, files) in [("kept", 2), ("context", 3)] { + let value = |key: &str| child_value(&out, &format!("{shape}_{key}")); + let (file, models, peak) = (value("FILE"), value("MODELS"), value("PEAK")); + println!( + "M-1 save ({shape}): file {} MiB, models {} MiB, process peak {} MiB = {:.2} x file", + file / MIB, + models / MIB, + peak / MIB, + peak as f64 / file as f64 + ); + assert_eq!(value("OK"), 1, "{shape}: the save must publish: {out}"); + assert!( + peak < files * file + models + 16 * MIB, + "{shape}: {} MiB for a {} MiB file with {} MiB of models", + peak / MIB, + file / MIB, + models / MIB + ); + } +} + +#[test] +#[ignore = "child process of save_of_a_large_file_holds_its_staged_read_and_models_only"] +fn save_process_peak_child() { + if child_mode().as_deref() != Some("m1_save") { + return; + } + for (shape, pad_mib, ops) in [("kept", 40, 15_000), ("context", 16, 30_000)] { + let t = E2e::new(&format!("m1_save_{shape}")).expect("engines"); + let pdf = large_file(pad_mib, ops); + let file = pdf.len(); + let src = t.file("large.pdf", &pdf); + let edits = (0..2u32) + .map(|p| t.edit(&src, p, p, "Hello", "Help")) + .collect(); + let fp = t.open(&src).fingerprint; + let source = t + .cache + .get(&src, &t.temp_root(), &fp, &t.engines) + .expect("source"); + let models: usize = (0..2u32) + .map(|p| t.cache.page(&source, p).expect("model").walk.model_bytes) + .sum(); + std::mem::drop((source, pdf)); + t.cache.release(&src); + t.cache.wait_checks(); + let (saved, peak) = process_peak(|| t.save(&[(&src, "1-z")], edits)); + let ok = saved.is_ok_and(|s| t.words(&s.path, 0).contains("Help")); + println!( + "{shape}_FILE={file}\n{shape}_MODELS={models}\n{shape}_PEAK={peak}\n{shape}_OK={}", + u8::from(ok) + ); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_e2e/overlay.rs b/src-tauri/src/pdf_engine/text_edit/tests_e2e/overlay.rs new file mode 100644 index 0000000..629851d --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_e2e/overlay.rs @@ -0,0 +1,275 @@ +//! E2E-10a…f and E2E-11: text changes saved together with objects that make the pipeline run +//! `qpdf --overlay` (every page becomes qpdf's wrapper Form, D29) and with redaction. Phase B and +//! #34 must pass in every variant; E2E-10c needs #34's `alt_content_digest` (VO-01). + +use super::save::pages_doc; +use super::{page_holdings, page_parts, E2e, SaveOpts}; +use crate::pdf_engine::edit_forms::FormValue; +use crate::pdf_engine::text_edit::content::qpdf_join; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_doc, DocBuilder, PageSpec, HELVETICA, +}; +use serde_json::{json, Value}; + +fn stamp(page: u32, text: &str) -> Value { + json!({ + "kind": "text", "pageIndex": page, + "rect": { "x": 300.0, "y": 120.0, "w": 200.0, "h": 30.0 }, + "content": text, "fontSize": 14.0, "color": "#c71c1c", "align": null, "opacity": 1.0, + }) +} + +fn shape(page: u32) -> Value { + json!({ + "kind": "rect", "pageIndex": page, + "rect": { "x": 400.0, "y": 300.0, "w": 60.0, "h": 40.0 }, + "fill": "#14733d", "stroke": null, "strokeWidth": null, "opacity": 1.0, + }) +} + +#[test] +fn e2e_10a_stamp_on_the_edited_page() { + let Some(t) = E2e::new("e2e_10a") else { + return; + }; + let src = t.file("a.pdf", &fx::word()); + let objects = vec![ + t.edit(&src, 0, 0, "Invoice 2026", "Invoice 2027"), + stamp(0, "Approved"), + ]; + let saved = t.save(&[(&src, "1-z")], objects).expect("E2E-10a save"); + t.assert_saved(&saved.path, &[(&src, 0, 0, "Invoice 2026", "Invoice 2027")]); + assert!( + t.words(&saved.path, 0).contains("Approved"), + "E2E-10a stamp" + ); +} + +#[test] +fn e2e_10b_stamp_on_another_page_only() { + let Some(t) = E2e::new("e2e_10b") else { + return; + }; + let src = t.file("b.pdf", &pages_doc(&["Edited line", "Stamped page"])); + let objects = vec![ + t.edit(&src, 0, 0, "Edited line", "Changed line"), + stamp(1, "Seen"), + ]; + let saved = t.save(&[(&src, "1-z")], objects).expect("E2E-10b save"); + t.assert_saved(&saved.path, &[(&src, 0, 0, "Edited line", "Changed line")]); + t.assert_unedited(&src, 1, &saved.path, 1); +} + +#[test] +fn e2e_10c_split_page_without_trailing_newlines_and_a_stamp() { + let Some(t) = E2e::new("e2e_10c") else { + return; + }; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + d.page(PageSpec::parts( + &[ + b"BT /F1 12 Tf 72 720 Td (Split zero) Tj ET", + b"BT /F1 12 Tf 72 700 Td (Split one) Tj ET", + ], + &res, + )); + d.page(PageSpec::parts( + &[ + b"BT /F1 12 Tf 72 720 Td (Other zero) Tj ET", + b"BT /F1 12 Tf 72 700 Td (Other one) Tj ET", + ], + &res, + )); + let src = t.file("c.pdf", &d.build()); + let objects = vec![ + t.edit(&src, 0, 0, "Split zero", "Split 0"), + stamp(0, "Stamp"), + ]; + let saved = t + .save(&[(&src, "1-z")], objects) + .expect("E2E-10c save (needs alt_content_digest)"); + t.assert_saved(&saved.path, &[(&src, 0, 0, "Split zero", "Split 0")]); + // The unedited split page sits in qpdf's wrapper as its parts joined with a "\n" — not + // lopdf's plain concatenation, which is what #34 expected before the fix. + let parts = page_parts(&src, 1); + let refs: Vec<&[u8]> = parts.iter().map(Vec::as_slice).collect(); + let joined = qpdf_join(&refs); + assert_ne!( + joined, + parts.concat(), + "E2E-10c fixture must need the join rule" + ); + assert!( + page_holdings(&saved.path, 1).contains(&joined), + "E2E-10c wrapper data" + ); +} + +#[test] +fn e2e_10d_offset_crop_with_trim_inside_and_a_stamp() { + let Some(t) = E2e::new("e2e_10d") else { + return; + }; + let pdf = helvetica_doc( + b"BT /F1 12 Tf 72 700 Td (Cropped trim line) Tj ET", + "", + "/CropBox [36 48 576 744] /TrimBox [72 72 540 720]", + ); + let src = t.file("d.pdf", &pdf); + let objects = vec![ + t.edit(&src, 0, 0, "Cropped trim line", "Cropped trim lines"), + stamp(0, "Boxed"), + ]; + let saved = t.save(&[(&src, "1-z")], objects).expect("E2E-10d save"); + t.assert_saved( + &saved.path, + &[(&src, 0, 0, "Cropped trim line", "Cropped trim lines")], + ); +} + +#[test] +fn e2e_10e_rotated_page_and_a_shape() { + let Some(t) = E2e::new("e2e_10e") else { + return; + }; + let src = t.file("e.pdf", &fx::rotated(90, true)); + let objects = vec![ + t.edit(&src, 0, 0, "Rotated page", "Rotated pages"), + shape(0), + ]; + let saved = t.save(&[(&src, "1-z")], objects).expect("E2E-10e save"); + t.assert_saved( + &saved.path, + &[(&src, 0, 0, "Rotated page", "Rotated pages")], + ); +} + +/// A page with an AcroForm text field "name" and a text line. +fn form_doc() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let page = d.reserve(); + let widget = d.add(format!( + "<< /Type /Annot /Subtype /Widget /FT /Tx /T (name) /F 4 /Rect [300 500 500 520] \ + /P {page} 0 R /DA (/F1 12 Tf 0 g) >>" + )); + d.catalog_extra.push_str(&format!( + " /AcroForm << /Fields [{widget} 0 R] /DA (/F1 12 Tf 0 g) /DR << /Font << /F1 {f} 0 R >> >> >>" + )); + d.page_at( + page, + PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Form page line) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + ) + .with(&format!("/Annots [{widget} 0 R]")), + ); + d.build() +} + +fn highlight(page: u32) -> Value { + json!({ + "kind": "highlight", "id": "h1", "pageIndex": page, + "rect": { "x": 70.0, "y": 695.0, "w": 100.0, "h": 16.0 }, + "author": "", "color": "#ffff00", "comment": null, "quads": [], + }) +} + +/// The existing pipeline refuses to flatten markup while the file keeps an AcroForm +/// (`FORM_FLATTEN_REQUIRED`, `edit_annots::apply_markup_annots`; flattening the form keeps the +/// catalog entry), so 10f runs as two saves: text + link + flattened form value + markup kept as +/// an annotation, then text + link + flattened markup on a page without a form. +#[test] +fn e2e_10f_text_link_flattened_form_and_flattened_markup() { + let Some(t) = E2e::new("e2e_10f") else { + return; + }; + let link = json!({ + "kind": "link", "pageIndex": 0, + "rect": { "x": 72.0, "y": 600.0, "w": 120.0, "h": 20.0 }, + "action": { "type": "uri", "uri": "https://example.com/" }, + }); + let src = t.file("f.pdf", &form_doc()); + let objects = vec![ + t.edit(&src, 0, 0, "Form page line", "Form page lines"), + link.clone(), + highlight(0), + ]; + let opts = SaveOpts { + form_values: vec![FormValue { + name: "name".into(), + value: "Ada".into(), + }], + flatten_form: true, + ..SaveOpts::default() + }; + let saved = t + .save_with(&[(&src, "1-z")], objects, opts) + .unwrap_or_else(|e| panic!("E2E-10f form save: {e} {:?}", e.details)); + t.assert_saved( + &saved.path, + &[(&src, 0, 0, "Form page line", "Form page lines")], + ); + assert!( + t.words(&saved.path, 0).contains("Ada"), + "E2E-10f flattened field value" + ); + + let src = t.file("f2.pdf", &pages_doc(&["Marked line"])); + let objects = vec![ + t.edit(&src, 0, 0, "Marked line", "Marked lines"), + link, + highlight(0), + ]; + let opts = SaveOpts { + flatten_annotations: true, + ..SaveOpts::default() + }; + let saved = t + .save_with(&[(&src, "1-z")], objects, opts) + .unwrap_or_else(|e| panic!("E2E-10f markup save: {e} {:?}", e.details)); + t.assert_saved(&saved.path, &[(&src, 0, 0, "Marked line", "Marked lines")]); +} + +#[test] +fn e2e_11_redaction_on_another_page_and_on_the_same_page() { + let Some(t) = E2e::new("e2e_11") else { + return; + }; + let src = t.file("r.pdf", &pages_doc(&["Keep this line", "Redact this line"])); + let redact = |page: u32| { + json!({ + "kind": "redact", "pageIndex": page, + "rect": { "x": 60.0, "y": 690.0, "w": 300.0, "h": 30.0 }, + "fill": "#000000", "label": null, + }) + }; + // Redaction on page 2, with and without a stamp (the redaction path's own overlay). + for with_stamp in [false, true] { + let mut objects = vec![ + t.edit(&src, 0, 0, "Keep this line", "Kept this line"), + redact(1), + ]; + if with_stamp { + objects.push(stamp(0, "Stamp")); + } + let saved = t.save(&[(&src, "1-z")], objects).expect("E2E-11 save"); + t.assert_saved( + &saved.path, + &[(&src, 0, 0, "Keep this line", "Kept this line")], + ); + assert!( + !t.words(&saved.path, 1).contains("Redact"), + "E2E-11 redacted text still extractable" + ); + } + // Both on page 1: refused before anything is written. + let objects = vec![ + t.edit(&src, 0, 0, "Keep this line", "Kept this line"), + redact(0), + ]; + let e = t.save(&[(&src, "1-z")], objects).err().expect("same page"); + assert_eq!(e.code, "TEXT_EDIT_ON_REDACTED_PAGE", "{e}"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_e2e/preview.rs b/src-tauri/src/pdf_engine/text_edit/tests_e2e/preview.rs new file mode 100644 index 0000000..bd5b22e --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_e2e/preview.rs @@ -0,0 +1,363 @@ +//! PREV-01…07: the preview command path (`service::preview_edits`) — a real qpdf-written, +//! Phase-A-proven one-page PDF — and its agreement with Save. + +use super::save::pages_doc; +use super::{page_parts, E2e}; +use crate::pdf_engine::text_edit::dto::{TextEditIn, TextPreviewDto}; +use crate::pdf_engine::text_edit::reasons::Face; +use crate::pdf_engine::text_edit::reasons::{EditProblemCode, TextWarningCode}; +use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; +use crate::pdf_engine::text_edit::service; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, DocBuilder, PageSpec, HELVETICA, +}; +use std::path::{Path, PathBuf}; + +pub(crate) fn decode_b64(s: &str) -> Vec { + let val = |c: u8| -> u32 { + match c { + b'A'..=b'Z' => u32::from(c - b'A'), + b'a'..=b'z' => u32::from(c - b'a') + 26, + b'0'..=b'9' => u32::from(c - b'0') + 52, + b'+' => 62, + b'/' => 63, + _ => panic!("bad base64 byte {c}"), + } + }; + let mut out = Vec::new(); + for chunk in s.as_bytes().chunks(4) { + let pad = chunk.iter().filter(|c| **c == b'=').count(); + let n = chunk + .iter() + .take(4 - pad) + .fold(0u32, |acc, c| (acc << 6) | val(*c)) + << (6 * pad as u32); + let bytes = [(n >> 16) as u8, (n >> 8) as u8, n as u8]; + out.extend_from_slice(&bytes[..3 - pad]); + } + out +} + +pub(crate) fn edit_in(run_id: &str, old: &str, new: &str, style: SourceTextStyleIn) -> TextEditIn { + TextEditIn { + run_id: run_id.to_string(), + original_text: old.to_string(), + text: new.to_string(), + style, + } +} + +impl E2e { + /// `service::preview_edits` of `(old, new, style)` edits on page `page`. + pub(crate) fn preview( + &self, + path: &Path, + page: u32, + edits: &[(&str, &str, SourceTextStyleIn)], + ) -> TextPreviewDto { + let (fp, dto) = self.page(path, page); + let ins: Vec = edits + .iter() + .map(|(old, new, style)| { + let run = dto + .runs + .iter() + .find(|r| r.text == *old) + .unwrap_or_else(|| panic!("no run {old:?}")); + edit_in(&run.id, old, new, style.clone()) + }) + .collect(); + service::preview_edits( + &self.cache, + &self.engines, + &self.temp_root(), + &path.to_string_lossy(), + &fp, + page, + &ins, + ) + .unwrap_or_else(|e| panic!("preview: {e} {:?}", e.details)) + } + + /// Writes a preview's page PDF next to the scratch files. + pub(crate) fn preview_file(&self, preview: &TextPreviewDto, name: &str) -> PathBuf { + let b64 = preview.page_pdf.as_deref().expect("preview pdf"); + self.file(name, &decode_b64(b64)) + } +} + +#[test] +fn prev_01_ok_verdict_and_a_page_pdf() { + let Some(t) = E2e::new("prev_01") else { + return; + }; + let src = t.file("p1.pdf", &fx::word()); + let p = t.preview( + &src, + 0, + &[("Invoice 2026", "Invoice 2027", SourceTextStyleIn::default())], + ); + assert!(p.verdicts.iter().all(|v| v.ok), "PREV-01 {:?}", p.verdicts); + assert!(p.page_problem.is_none()); + let v = &p.verdicts[0]; + assert!( + v.caret_offsets.as_ref().is_some_and(|c| c.len() == 13), + "{v:?}" + ); + assert!(v.new_rect.is_some()); + let page = t.preview_file(&p, "p1-preview.pdf"); + assert!(t.words(&page, 0).contains("Invoice2027")); +} + +#[test] +fn prev_02_preview_parts_equal_save_parts() { + let Some(t) = E2e::new("prev_02") else { + return; + }; + let src = t.file("p2.pdf", &pages_doc(&["Preview line", "Second page"])); + let p = t.preview( + &src, + 0, + &[( + "Preview line", + "Previewed line", + SourceTextStyleIn::default(), + )], + ); + let page = t.preview_file(&p, "p2-preview.pdf"); + let obj = t.edit(&src, 0, 0, "Preview line", "Previewed line"); + let saved = t.save(&[(&src, "1-z")], vec![obj]).expect("PREV-02 save"); + assert_eq!(page_parts(&page, 0), page_parts(&saved.path, 0), "PREV-02"); +} + +#[test] +fn prev_03_a_failed_verdict_has_no_page_pdf() { + let Some(t) = E2e::new("prev_03") else { + return; + }; + let src = t.file("p3.pdf", &fx::word()); + let p = t.preview( + &src, + 0, + &[( + "Invoice 2026", + "Invoice 2026Y", + SourceTextStyleIn::default(), + )], + ); + assert!(p.page_pdf.is_none(), "PREV-03 keeps the last good page"); + assert_eq!(p.verdicts[0].code, Some(EditProblemCode::GlyphMissing)); + assert_eq!(p.verdicts[0].chars, vec!["Y".to_string()]); + assert!(!p.verdicts[0].ok && p.verdicts[0].caret_offsets.is_none()); +} + +/// Two incompressible 25 MB images on a page: every one-page copy is over 48 MiB. +fn heavy_page() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let side = 2_900usize; + let mut state = 0x9E37_79B9_7F4A_7C15u64; + let mut noise = |n: usize| -> Vec { + let mut v = Vec::with_capacity(n + 8); + while v.len() < n { + state ^= state << 13; + state ^= state >> 7; + state ^= state << 17; + v.extend_from_slice(&state.to_le_bytes()); + } + v.truncate(n); + v + }; + let dict = format!( + "/Type /XObject /Subtype /Image /Width {side} /Height {side} /ColorSpace /DeviceRGB /BitsPerComponent 8" + ); + let a = d.b.add_stream(&dict, &noise(side * side * 3)); + let b = d.b.add_stream(&dict, &noise(side * side * 3)); + d.page(PageSpec::new( + b"q 100 0 0 100 400 600 cm /Im0 Do Q q 100 0 0 100 400 450 cm /Im1 Do Q \ + BT /F1 12 Tf 72 700 Td (Heavy page) Tj ET", + &format!("/Font << /F1 {f} 0 R >> /XObject << /Im0 {a} 0 R /Im1 {b} 0 R >>"), + )); + d.build() +} + +#[test] +fn prev_04_an_oversize_preview_is_unavailable_not_an_error() { + let Some(t) = E2e::new("prev_04") else { + return; + }; + let src = t.file("heavy.pdf", &heavy_page()); + let p = t.preview( + &src, + 0, + &[("Heavy page", "Heavy pages", SourceTextStyleIn::default())], + ); + assert!(p.verdicts.iter().all(|v| v.ok), "PREV-04 {:?}", p.verdicts); + assert!(p.page_pdf.is_none(), "PREV-04 page pdf over 48 MiB"); + assert!( + p.warnings + .iter() + .any(|w| w.code == TextWarningCode::PreviewUnavailable), + "PREV-04 {:?}", + p.warnings + ); +} + +/// `text` in a content stream with 100 KB of other data on both sides (outside the first and +/// last 64 KiB that `stat_matches` hashes). +fn padded_doc(text: &str) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let pad = vec![b'%'; 100_000]; + d.b.add_stream("", &pad); + let content = format!("BT /F1 12 Tf 72 700 Td ({text}) Tj ET"); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.b.add_stream("", &pad); + d.build() +} + +#[test] +fn prev_05_preview_reads_the_snapshot_copy_not_the_users_file() { + let Some(t) = E2e::new("prev_05") else { + return; + }; + let bytes = padded_doc("Middle text"); + let src = t.file("p5.pdf", &bytes); + let (fp, dto) = t.page(&src, 0); + let run = dto + .runs + .iter() + .find(|r| r.text == "Middle text") + .expect("run") + .clone(); + // Same length, same mtime, same first/last 64 KiB: only the middle changes on disk. + let modified = std::fs::metadata(&src) + .and_then(|m| m.modified()) + .expect("mtime"); + let at = bytes + .windows(13) + .position(|w| w == b"(Middle text)") + .expect("needle"); + assert!( + at > 65_536 && bytes.len() - at > 65_536, + "PREV-05 fixture layout" + ); + let mut changed = bytes.clone(); + changed[at..at + 13].copy_from_slice(b"(Mutant text)"); + std::fs::write(&src, &changed).expect("rewrite"); + std::fs::File::options() + .write(true) + .open(&src) + .and_then(|f| f.set_modified(modified)) + .expect("restore mtime"); + let p = service::preview_edits( + &t.cache, + &t.engines, + &t.temp_root(), + &src.to_string_lossy(), + &fp, + 0, + &[edit_in( + &run.id, + "Middle text", + "Middle texts", + SourceTextStyleIn::default(), + )], + ) + .expect("PREV-05 preview"); + assert!(p.verdicts[0].ok, "{:?}", p.verdicts); + let page = t.preview_file(&p, "p5-preview.pdf"); + let words = t.words(&page, 0); + assert!( + words.contains("Middletexts") && !words.contains("Mutant"), + "PREV-05 {words}" + ); +} + +#[test] +fn prev_06_face_from_a_shared_inherited_resource_dictionary() { + let Some(t) = E2e::new("prev_06") else { + return; + }; + let src = t.file("p6.pdf", &fx::shared_inherited_resources()); + let bold = SourceTextStyleIn { + face: Some(Face::Bold), + ..SourceTextStyleIn::default() + }; + let p = t.preview(&src, 0, &[("Regular words", "Regular words", bold)]); + assert!(p.verdicts.iter().all(|v| v.ok), "PREV-06 {:?}", p.verdicts); + assert!(p.page_pdf.is_some(), "PREV-06 preview: {:?}", p.warnings); + let obj = t.edit_styled( + &src, + 0, + 0, + "Regular words", + "Regular words", + serde_json::json!({ "face": "bold" }), + ); + let saved = t.save(&[(&src, "1-z")], vec![obj]).expect("PREV-06 save"); + t.assert_checks_clean(&saved.path); + let (_, run) = t.run(&saved.path, 0, "Regular words"); + assert_eq!( + run.style.map(|s| s.face), + Some(Face::Bold), + "PREV-06 saved face" + ); +} + +#[test] +fn prev_07_concurrent_inspects_and_previews_share_the_cache_safely() { + let Some(t) = E2e::new("prev_07") else { + return; + }; + let src = t.file("p7.pdf", &pages_doc(&["Concurrent line"])); + let barrier = std::sync::Barrier::new(8); + let results: Vec>> = std::thread::scope(|s| { + let handles: Vec<_> = (0..8) + .map(|i| { + let (cache, engines, temp) = (t.cache.clone(), t.engines.clone(), t.temp_root()); + let (src, barrier) = (src.clone(), &barrier); + let out = t.scratch.path(&format!("p7-{i}.pdf")); + s.spawn(move || { + let path = src.to_string_lossy().into_owned(); + barrier.wait(); + let info = service::open_source(&cache, &engines, &temp, &path).expect("open"); + let page = + service::inspect_page(&cache, &engines, &temp, &path, &info.fingerprint, 0) + .expect("inspect"); + let run = &page.runs[0]; + let p = service::preview_edits( + &cache, + &engines, + &temp, + &path, + &info.fingerprint, + 0, + &[edit_in( + &run.id, + &run.text, + "Concurrent lines", + SourceTextStyleIn::default(), + )], + ) + .expect("preview"); + assert!(p.verdicts[0].ok, "{:?}", p.verdicts); + std::fs::write(&out, decode_b64(p.page_pdf.as_deref().expect("pdf"))) + .expect("write"); + page_parts(&out, 0) + }) + }) + .collect(); + handles + .into_iter() + .map(|h| h.join().expect("thread")) + .collect() + }); + assert!( + results.windows(2).all(|w| w[0] == w[1]), + "PREV-07 previews differ" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_e2e/refusals.rs b/src-tauri/src/pdf_engine/text_edit/tests_e2e/refusals.rs new file mode 100644 index 0000000..13a1e27 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_e2e/refusals.rs @@ -0,0 +1,375 @@ +//! E2E-12…16, 19…21 and GATE-25: style changes, refusals that surface at Save, a source changed +//! after inspect, overwrite protection, cancellation, a missing verifier, appended signed files, +//! the verification read cap, and a fake injected into the edited copy. + +use super::fonts::{assert_same_box, word_box}; +use super::save::pages_doc; +use super::{E2e, SaveOpts}; +use crate::error::AppError; +use crate::pdf_engine::qpdf::resolve_qpdf_standalone; +use crate::pdf_engine::text_edit::export::seams; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::reasons::{Face, ORIGINAL_UNCHANGED}; +use crate::pdf_engine::text_edit::snapshot::read_snapshot; +use crate::pdf_engine::text_edit::testkit::fakes; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, DocBuilder, PageSpec, HELVETICA, +}; +use serde_json::{json, Value}; +use std::path::{Path, PathBuf}; +use std::sync::atomic::{AtomicBool, Ordering}; + +/// A save error carries the unchanged-file sentence exactly once, at the end. +pub(crate) fn assert_save_error(e: &AppError, code: &str) { + assert_eq!(e.code, code, "{e} {:?}", e.details); + let s = e.suggestion.as_deref().unwrap_or_default(); + assert!(s.ends_with(ORIGINAL_UNCHANGED), "{code}: suggestion {s:?}"); + assert_eq!(s.matches(ORIGINAL_UNCHANGED).count(), 1, "{code}: {s:?}"); +} + +/// A `sourceText` object built by hand (for runs the editor would never offer). +fn raw_edit(fp: &str, run_id: &str, old: &str, new: &str) -> Value { + json!({ + "id": "raw", "kind": "sourceText", "pageIndex": 0, + "rect": { "x": 72.0, "y": 700.0, "w": 100.0, "h": 12.0 }, + "locked": true, "runId": run_id, "sourceFingerprint": fp, "sourcePageIndex": 0, + "originalText": old, "text": new, "style": {}, + }) +} + +/// Helvetica and its bold sibling, a styled line and a follower line. +fn styled_doc() -> Vec { + let mut d = DocBuilder::new(); + let r = d.add(HELVETICA); + let b = d.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica-Bold /Encoding /WinAnsiEncoding >>", + ); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Styled line) Tj ET BT /F1 12 Tf 72 670 Td (Next line) Tj ET \ + BT /F2 12 Tf 72 640 Td (Bold) Tj ET", + &format!("/Font << /F1 {r} 0 R /F2 {b} 0 R >>"), + )); + d.build() +} + +#[test] +fn e2e_12_size_bold_red_and_spacing_on_one_line() { + let Some(t) = E2e::new("e2e_12") else { + return; + }; + let src = t.file("styled.pdf", &styled_doc()); + let next_before = word_box(&t, &src, 0, "Next"); + let (_, next_run) = t.run(&src, 0, "Next line"); + let style = + json!({ "sizePt": 14.0, "face": "bold", "fill": "#c71c1c", "letterSpacingPt": 0.5 }); + let obj = t.edit_styled(&src, 0, 0, "Styled line", "Styled line", style); + let saved = t.save(&[(&src, "1-z")], vec![obj]).expect("E2E-12 save"); + t.assert_checks_clean(&saved.path); + assert_same_box( + next_before, + word_box(&t, &saved.path, 0, "Next"), + "E2E-12 next line", + ); + let (_, run) = t.run(&saved.path, 0, "Styled line"); + let m = run.metrics.as_ref().expect("metrics"); + let s = run.style.as_ref().expect("style"); + assert!( + (m.effective_size - 14.0).abs() < 1e-6, + "E2E-12 size {}", + m.effective_size + ); + assert!( + (m.letter_spacing_pt - 0.5).abs() < 1e-4, + "E2E-12 spacing {}", + m.letter_spacing_pt + ); + assert_eq!(s.face, Face::Bold, "E2E-12 face"); + assert_eq!(s.fill.as_deref(), Some("#c71c1c"), "E2E-12 fill"); + let (_, next_after) = t.run(&saved.path, 0, "Next line"); + let style_of = + |r: &crate::pdf_engine::text_edit::dto::TextRunDto| r.style.clone().map(|s| s.fill); + assert_eq!( + style_of(&next_after), + style_of(&next_run), + "E2E-12 next line colour" + ); + assert_eq!(next_after.rect, next_run.rect, "E2E-12 next line box"); +} + +/// Page 1 is fine; page 2's content has an unterminated string: `qpdf --check` warns about it +/// (not on the benign allow-list), our page-1 walk does not care. +pub(crate) fn needs_repair_doc() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Good line) Tj ET", + &res, + )); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (open string Tj ET", + &res, + )); + d.build() +} + +#[test] +fn e2e_13_refusals_surface_at_save() { + let Some(t) = E2e::new("e2e_13") else { + return; + }; + // A shared content stream: refused at inspect and at Save. + let src = t.file("shared.pdf", &fx::shared()); + let (fp, run) = t.run(&src, 0, "Shared body"); + assert!(!run.editable); + let e = t + .save( + &[(&src, "1-z")], + vec![raw_edit(&fp, &run.id, "Shared body", "Shared text")], + ) + .err() + .expect("shared"); + assert_save_error(&e, "TEXT_EDIT_REFUSED"); + assert!( + e.details + .as_deref() + .unwrap_or_default() + .contains("SHARED_CONTENT"), + "{e:?}" + ); + // File-level refusals (the editor refuses these files at open; Save refuses them again). + let enc = fx::encrypted(&t.engines); + for (name, pdf, code) in [ + ("encrypted.pdf", enc, "ENCRYPTED"), + ("signed.pdf", fx::signed(), "SIGNED"), + ("xfa.pdf", fx::xfa(), "UNSUPPORTED_XFA"), + ] { + let src = t.file(name, &pdf); + let fp = format!("{:016x}-{:x}", 1u64, pdf.len()); + let e = t + .save( + &[(&src, "1-z")], + vec![raw_edit(&fp, "t1:x:0:0-1", "Hello", "Help")], + ) + .err() + .unwrap_or_else(|| panic!("{name} saved")); + assert_save_error(&e, code); + } + // qpdf --check finds a real problem elsewhere in the file. + let src = t.file("repair.pdf", &needs_repair_doc()); + let obj = t.edit(&src, 0, 0, "Good line", "Good lines"); + let e = t.save(&[(&src, "1-z")], vec![obj]).err().expect("repair"); + assert_save_error(&e, "PDF_NEEDS_REPAIR"); +} + +#[test] +fn e2e_14_source_changed_after_inspect_is_stale() { + let Some(t) = E2e::new("e2e_14") else { + return; + }; + let src = t.file("stale.pdf", &pages_doc(&["Before change"])); + let obj = t.edit(&src, 0, 0, "Before change", "After change"); + std::fs::write(&src, pages_doc(&["Rewritten line"])).expect("rewrite"); + let e = t.save(&[(&src, "1-z")], vec![obj]).err().expect("stale"); + assert_save_error(&e, "STALE"); + assert_eq!( + e.message, + "\u{201c}stale.pdf\u{201d} was changed after you started editing it, so your text changes no longer match it." + ); +} + +#[test] +fn e2e_15_destination_is_the_source_or_a_hard_link_to_it() { + let Some(t) = E2e::new("e2e_15") else { + return; + }; + let src = t.file("own.pdf", &pages_doc(&["Own line"])); + let link = t.scratch.path("link.pdf"); + std::fs::hard_link(&src, &link).expect("hard link"); + for dest in [src.clone(), link] { + let obj = t.edit(&src, 0, 0, "Own line", "Own lines"); + let opts = SaveOpts { + dest: Some(dest), + ..SaveOpts::default() + }; + let e = t + .save_with(&[(&src, "1-z")], vec![obj], opts) + .err() + .expect("overwrite"); + assert_eq!(e.code, "OVERWRITE", "{e}"); + } +} + +static CANCEL_16: AtomicBool = AtomicBool::new(false); + +fn cancel_now(_edited: &Path) { + CANCEL_16.store(true, Ordering::SeqCst); +} + +#[test] +fn e2e_16_cancel_during_and_after_the_text_changes() { + let Some(t) = E2e::new("e2e_16") else { + return; + }; + let src = t.file("cancel.pdf", &pages_doc(&["Cancel one", "Cancel two"])); + // While the edited copy is being proven (Phase A sees the cancel). + CANCEL_16.store(false, Ordering::SeqCst); + seams::set_tamper_hook(Some(cancel_now)); + let obj = t.edit(&src, 0, 0, "Cancel one", "Cancel 1"); + let opts = SaveOpts { + cancel: Some(&CANCEL_16), + ..SaveOpts::default() + }; + let r = t.save_with(&[(&src, "1-z")], vec![obj], opts); + seams::set_tamper_hook(None); + assert_eq!( + r.err().map(|e| e.code), + Some("CANCELLED".to_string()), + "E2E-16 Phase A" + ); + // After the text changes, when the pipeline assembles (Phase B / #34 see the cancel). + let flag = AtomicBool::new(false); + let cancel_on_run = |_: &[String]| flag.store(true, Ordering::SeqCst); + let obj = t.edit(&src, 0, 0, "Cancel one", "Cancel 1"); + let opts = SaveOpts { + cancel: Some(&flag), + on_run: Some(&cancel_on_run), + ..SaveOpts::default() + }; + let r = t.save_with(&[(&src, "1"), (&src, "2")], vec![obj], opts); + assert_eq!( + r.err().map(|e| e.code), + Some("CANCELLED".to_string()), + "E2E-16 later" + ); +} + +#[test] +fn e2e_19_missing_poppler_is_verifier_missing() { + let Some(t) = E2e::new("e2e_19") else { + return; + }; + let src = t.file("nopoppler.pdf", &pages_doc(&["Needs poppler"])); + let obj = t.edit(&src, 0, 0, "Needs poppler", "Needs Poppler"); + seams::set_poppler_override(Some(PathBuf::from("/nonexistent/offpdf-poppler"))); + let r = t.save(&[(&src, "1-z")], vec![obj]); + seams::set_poppler_override(None); + assert_save_error(&r.err().expect("verifier missing"), "VERIFIER_MISSING"); +} + +#[test] +fn e2e_20_an_unedited_signed_pdf_may_be_appended() { + let Some(t) = E2e::new("e2e_20") else { + return; + }; + let src = t.file("plain.pdf", &pages_doc(&["Plain line"])); + let signed = t.file("signed.pdf", &fx::signed()); + assert_eq!( + read_snapshot(&signed).err().map(|e| e.code), + Some("SIGNED".into()) + ); + let obj = t.edit(&src, 0, 0, "Plain line", "Plain lines"); + let saved = t + .save(&[(&src, "1-z"), (&signed, "1-z")], vec![obj]) + .unwrap_or_else(|e| panic!("E2E-20 save: {e} {:?}", e.details)); + t.assert_saved(&saved.path, &[(&src, 0, 0, "Plain line", "Plain lines")]); + t.assert_unedited(&signed, 0, &saved.path, 1); +} + +#[test] +fn e2e_21_the_final_file_may_exceed_the_source_cap() { + let Some(t) = E2e::new("e2e_21") else { + return; + }; + let a = pages_doc(&["Capped one"]); + let b = pages_doc(&["Capped two", "Capped three"]); + let cap = a.len().max(b.len()) as u64 + 64; + let src_a = t.file("cap-a.pdf", &a); + let src_b = t.file("cap-b.pdf", &b); + let objects = vec![ + t.edit(&src_a, 0, 0, "Capped one", "Capped 1"), + t.edit(&src_b, 2, 1, "Capped three", "Capped 3"), + ]; + limits::set_file_cap_override(Some(cap)); + let r = t.save(&[(&src_a, "1-z"), (&src_b, "1-z")], objects); + let alone = r + .as_ref() + .ok() + .map(|s| read_snapshot(&s.path).err().map(|e| e.code)); + limits::set_file_cap_override(None); + let saved = r.unwrap_or_else(|e| panic!("E2E-21 save: {e} {:?}", e.details)); + assert_eq!( + alone, + Some(Some("FILE_TOO_LARGE".to_string())), + "E2E-21 read_snapshot alone" + ); + t.assert_saved( + &saved.path, + &[ + (&src_a, 0, 0, "Capped one", "Capped 1"), + (&src_b, 1, 2, "Capped three", "Capped 3"), + ], + ); +} + +fn tamper_cover(edited: &Path) { + let qpdf = resolve_qpdf_standalone(); + let fake = edited.with_extension("fake.pdf"); + fakes::cover_and_overlay_f1( + &qpdf, + edited, + &fake, + 0, + [70.0, 695.0, 200.0, 20.0], + "F1", + "Gate line", + ) + .expect("cover fake"); + std::fs::rename(&fake, edited).expect("swap in the fake"); +} + +fn tamper_reattach(edited: &Path) { + let qpdf = resolve_qpdf_standalone(); + let fake = edited.with_extension("fake.pdf"); + fakes::reattach_old( + &qpdf, + edited, + &fake, + 0, + b"BT /F1 12 Tf 72 700 Td (Gate line) Tj ET", + ) + .expect("reattach fake"); + std::fs::rename(&fake, edited).expect("swap in the fake"); +} + +#[test] +fn gate_25_fakes_injected_into_the_edited_copy_never_publish() { + let Some(t) = E2e::new("gate_25") else { + return; + }; + let src = t.file("gate.pdf", &pages_doc(&["Gate line", "Other gate page"])); + for hook in [tamper_cover as fn(&Path), tamper_reattach] { + let obj = t.edit(&src, 0, 0, "Gate line", "Gate lines"); + seams::set_tamper_hook(Some(hook)); + let r = t.save(&[(&src, "1-z")], vec![obj]); + seams::set_tamper_hook(None); + let e = r.err().expect("GATE-25: a fake was published"); + assert!( + [ + "EDIT_VERIFY_FAILED", + "PEN_DRIFT", + "STATE_CHANGED", + "SOURCE_EDIT_GATE_FAILED" + ] + .contains(&e.code.as_str()), + "GATE-25 code {e} {:?}", + e.details + ); + assert_save_error(&e, &e.code.clone()); + assert!( + e.details.as_deref().unwrap_or_default().contains("phase=A"), + "{:?}", + e.details + ); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_e2e/save.rs b/src-tauri/src/pdf_engine/text_edit/tests_e2e/save.rs new file mode 100644 index 0000000..6a20b10 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_e2e/save.rs @@ -0,0 +1,418 @@ +//! E2E-07…09, 17, 18, 22: multi-page and multi-file saves, split content, no-op changes, the +//! old text gone from the file, and a Word-style hybrid-xref tagged source. + +use super::{page_parts, E2e}; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget}; +use crate::pdf_engine::text_edit::testkit::pdf::XrefStyle; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, tagged_bookmarked, DocBuilder, PageSpec, HELVETICA, +}; +use lopdf::{Document, Object}; +use serde_json::json; +use std::path::Path; + +/// One Helvetica line per page. +pub(crate) fn pages_doc(lines: &[&str]) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + for line in lines { + let content = format!("BT /F1 12 Tf 72 700 Td ({line}) Tj ET"); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f} 0 R >>"), + )); + } + d.build() +} + +#[test] +fn e2e_07_edits_on_pages_one_and_three_in_one_save() { + let Some(t) = E2e::new("e2e_07") else { + return; + }; + let src = t.file( + "three.pdf", + &pages_doc(&["First page", "Second page", "Third page"]), + ); + let objects = vec![ + t.edit(&src, 0, 0, "First page", "First sheet"), + t.edit(&src, 2, 2, "Third page", "Third sheet"), + ]; + let saved = t.save(&[(&src, "1-z")], objects).expect("E2E-07 save"); + t.assert_saved( + &saved.path, + &[ + (&src, 0, 0, "First page", "First sheet"), + (&src, 2, 2, "Third page", "Third sheet"), + ], + ); + t.assert_unedited(&src, 1, &saved.path, 1); +} + +#[test] +fn e2e_08_two_files_partial_ranges_and_one_file_in_two_groups() { + let Some(t) = E2e::new("e2e_08") else { + return; + }; + let a = t.file( + "a.pdf", + &pages_doc(&["Alpha one", "Alpha two", "Alpha three"]), + ); + let b = t.file("b.pdf", &pages_doc(&["Beta one", "Beta two"])); + // Two edited files, the second through a partial range: dest = A1 A2 A3 B2. + let objects = vec![ + t.edit(&a, 1, 1, "Alpha two", "Alpha 2"), + t.edit(&b, 3, 1, "Beta two", "Beta 2"), + ]; + let saved = t + .save(&[(&a, "1-z"), (&b, "2")], objects) + .expect("E2E-08 two files"); + t.assert_saved( + &saved.path, + &[ + (&a, 1, 1, "Alpha two", "Alpha 2"), + (&b, 1, 3, "Beta two", "Beta 2"), + ], + ); + t.assert_unedited(&a, 0, &saved.path, 0); + t.assert_unedited(&a, 2, &saved.path, 2); + + // One file in two groups with distinct edited pages: dest = A1 B1 A3. + let objects = vec![ + t.edit(&a, 0, 0, "Alpha one", "Alpha 1"), + t.edit(&a, 2, 2, "Alpha three", "Alpha 3"), + ]; + let saved = t + .save(&[(&a, "1"), (&b, "1"), (&a, "3")], objects) + .expect("E2E-08 same file twice"); + t.assert_saved( + &saved.path, + &[ + (&a, 0, 0, "Alpha one", "Alpha 1"), + (&a, 2, 2, "Alpha three", "Alpha 3"), + ], + ); + t.assert_unedited(&b, 0, &saved.path, 1); + + // An edited page listed twice (two groups, or a repeated range in one group). + for groups in [ + vec![(a.as_path(), "1-z"), (a.as_path(), "1")], + vec![(a.as_path(), "1,2,3,1")], + ] { + let obj = t.edit(&a, 0, 0, "Alpha one", "Alpha 1"); + let e = t.save(&groups, vec![obj]).err().expect("duplicate page"); + assert_eq!(e.code, "TEXT_EDIT_DUPLICATE_PAGE", "{e}"); + assert_eq!( + e.message, + "Page 1 of \u{201c}a.pdf\u{201d} is in the list more than once and has a text change." + ); + } +} + +#[test] +fn e2e_09_split_contents_and_quote_operators() { + let Some(t) = E2e::new("e2e_09") else { + return; + }; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts( + &[ + b"BT /F1 12 Tf 72 720 Td (Part zero) Tj ET", + b"BT /F1 12 Tf 72 700 Td (Part one) Tj ET", + b"BT /F1 12 Tf 72 680 Td (Part two) Tj ET", + ], + &format!("/Font << /F1 {f} 0 R >>"), + )); + let src = t.file("parts.pdf", &d.build()); + let objects = vec![ + t.edit(&src, 0, 0, "Part zero", "Part 0"), + t.edit(&src, 0, 0, "Part two", "Part 2"), + ]; + let saved = t.save(&[(&src, "1-z")], objects).expect("E2E-09 parts"); + t.assert_saved( + &saved.path, + &[ + (&src, 0, 0, "Part zero", "Part 0"), + (&src, 0, 0, "Part two", "Part 2"), + ], + ); + let parts = page_parts(&saved.path, 0); + assert_eq!(parts.len(), 3, "E2E-09 part count"); + assert_eq!( + parts[1], + b"BT /F1 12 Tf 72 700 Td (Part one) Tj ET".to_vec() + ); + + let src = t.file("quotes.pdf", &fx::quote_ops()); + let objects = vec![ + t.edit(&src, 0, 0, "Line two", "Line 2"), + t.edit(&src, 0, 0, "Line three", "Line 3"), + ]; + let saved = t.save(&[(&src, "1-z")], objects).expect("E2E-09 quotes"); + t.assert_saved( + &saved.path, + &[ + (&src, 0, 0, "Line two", "Line 2"), + (&src, 0, 0, "Line three", "Line 3"), + ], + ); +} + +#[test] +fn e2e_17_no_op_changes_write_nothing() { + let Some(t) = E2e::new("e2e_17") else { + return; + }; + let src = t.file("noop.pdf", &pages_doc(&["Same text", "Other page"])); + let (_, run) = t.run(&src, 0, "Same text"); + let size = run + .metrics + .as_ref() + .map(|m| m.effective_size) + .expect("metrics"); + // Text equal to the original, and a size within STYLE_EPSILON of the current one (B7). + let objects = vec![ + t.edit(&src, 0, 0, "Same text", "Same text"), + t.edit_styled( + &src, + 0, + 0, + "Same text", + "Same text", + json!({ "sizePt": size + 0.0004 }), + ), + ]; + for obj in objects { + let saved = t.save(&[(&src, "1-z")], vec![obj]).expect("E2E-17 save"); + t.assert_checks_clean(&saved.path); + for page in 0..2 { + assert_eq!( + page_parts(&saved.path, page), + page_parts(&src, page), + "E2E-17 page {page}" + ); + } + } +} + +/// Every stream of `pdf`, decoded where our bounded decoder can. +fn decoded_streams(pdf: &Path) -> Vec> { + let doc = Document::load(pdf).expect("load"); + doc.objects + .values() + .filter_map(|o| match o { + Object::Stream(s) => { + let mut budget = DecodeBudget::new(1 << 30); + Some(decode_stream(s, 1 << 30, &mut budget).unwrap_or_else(|_| s.content.clone())) + } + _ => None, + }) + .collect() +} + +#[test] +fn e2e_18_the_old_text_is_not_recoverable_from_the_file() { + let Some(t) = E2e::new("e2e_18") else { + return; + }; + let src = t.file( + "secret.pdf", + &pages_doc(&["Secret 4711 code", "Other page"]), + ); + let obj = t.edit(&src, 0, 0, "Secret 4711 code", "Public code"); + let saved = t.save(&[(&src, "1-z")], vec![obj]).expect("E2E-18 save"); + t.assert_saved( + &saved.path, + &[(&src, 0, 0, "Secret 4711 code", "Public code")], + ); + let needle = b"(Secret 4711 code) Tj"; + for data in decoded_streams(&saved.path) { + assert!( + !data.windows(needle.len()).any(|w| w == needle), + "E2E-18: the original operation is still in a stream" + ); + } + let raw = std::fs::read(&saved.path).expect("read"); + assert!( + !raw.windows(4).any(|w| w == b"4711"), + "E2E-18: old digits in the file" + ); +} + +#[test] +fn e2e_22_word_style_hybrid_xref_tagged_bookmarked_source() { + let Some(t) = E2e::new("e2e_22") else { + return; + }; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let page = d.reserve(); + let extra = tagged_bookmarked(&mut d, page, &[0]); + let content = b"/P <> BDC BT /F1 12 Tf 72 700 Td (Tagged hybrid line) Tj ET EMC"; + d.page_at( + page, + PageSpec::new(content, &format!("/Font << /F1 {f} 0 R >>")).with(&extra), + ); + let pdf = d.build_with(&XrefStyle::Hybrid { + in_objstm: vec![page], + objstm_in_classic: true, + }); + let src = t.file("hybrid.pdf", &pdf); + let obj = t.edit(&src, 0, 0, "Tagged hybrid line", "Tagged hybrid lines"); + let saved = t.save(&[(&src, "1-z")], vec![obj]).expect("E2E-22 save"); + t.assert_saved( + &saved.path, + &[(&src, 0, 0, "Tagged hybrid line", "Tagged hybrid lines")], + ); + let doc = Document::load(&saved.path).expect("load"); + let catalog = doc.catalog().expect("catalog"); + assert!(catalog.has(b"StructTreeRoot"), "E2E-22 structure tree kept"); + assert!(catalog.has(b"Outlines"), "E2E-22 bookmarks kept"); +} + +#[test] +fn e2e_overlap_warning_reaches_the_job_message() { + let Some(t) = E2e::new("e2e_overlap") else { + return; + }; + let pdf = crate::pdf_engine::text_edit::testkit::producers::helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Name) Tj ET BT /F1 12 Tf 110 700 Td (Value) Tj ET", + ); + let src = t.file("overlap.pdf", &pdf); + let obj = t.edit(&src, 0, 0, "Name", "Names and"); + let saved = t + .save(&[(&src, "1-z")], vec![obj]) + .expect("overlap is a warning"); + t.assert_saved(&saved.path, &[(&src, 0, 0, "Name", "Names and")]); + assert_eq!( + saved.warnings, + vec!["Page 1: the changed line \u{201c}Names and\u{201d} runs into the text that follows it.".to_string()] + ); +} + +/// review-T5 live B2: a change that runs into the next table cells (the live check's invoice row) +/// commits with the overlap warning, in the preview and on Save; the cells still read the same +/// at their places in the published file. +#[test] +fn b2_a_line_running_into_the_next_cells_saves_with_the_overlap_warning() { + let Some(t) = E2e::new("b2_table_overlap") else { + return; + }; + let pdf = fx::helvetica_page( + b"BT /F1 11 Tf 58 700 Td (Desk lamp) Tj ET BT /F1 11 Tf 132 700 Td (4) Tj ET \ + BT /F1 11 Tf 156 700 Td (120.00) Tj ET", + ); + let src = t.file("invoice-row.pdf", &pdf); + let new = "Desk lamp with a long cable"; + let p = t.preview( + &src, + 0, + &[( + "Desk lamp", + new, + crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn::default(), + )], + ); + assert!(p.page_pdf.is_some(), "B2 preview: {:?}", p.page_problem); + assert_eq!(p.verdicts.len(), 1); + assert!(p.verdicts[0].ok, "B2 verdict"); + let saved = t + .save(&[(&src, "1-z")], vec![t.edit(&src, 0, 0, "Desk lamp", new)]) + .unwrap_or_else(|e| panic!("B2 save: {e} {:?}", e.details)); + t.assert_checks_clean(&saved.path); + assert_eq!( + saved.warnings, + vec![format!( + "Page 1: the changed line \u{201c}{new}\u{201d} runs into the text that follows it." + )] + ); + let words = |pdf: &Path| { + crate::pdf_engine::text_edit::poppler::pdftotext_words( + &t.engines, + pdf, + 1, + &crate::pdf_engine::text_edit::engines::RunOpts::default(), + ) + .expect("words") + }; + let after = words(&saved.path); + for cell in words(&src) + .iter() + .filter(|w| w.text == "4" || w.text == "120.00") + { + assert!( + after.iter().any(|w| w.text == cell.text + && (w.x0 - cell.x0).abs() < 0.05 + && (w.y0 - cell.y0).abs() < 0.05), + "B2 cell {:?} no longer at its place: {after:?}", + cell.text + ); + } + for word in new.split(' ') { + assert!( + after.iter().any(|w| w.text == word), + "B2 {word:?}: {after:?}" + ); + } +} + +/// The `/Link` annotations of page `page` (0-based) of `pdf`. +fn link_count(pdf: &Path, page: u32) -> usize { + let snap = crate::pdf_engine::text_edit::snapshot::read_verification_snapshot(pdf, 1 << 30) + .unwrap_or_else(|e| panic!("{e}")); + let doc = &snap.doc; + let id = *snap.pages.get(page as usize).expect("page"); + let resolve = |o: &Object| match o { + Object::Reference(r) => doc.get_object(*r).ok().cloned(), + other => Some(other.clone()), + }; + let annots = doc + .get_dictionary(id) + .ok() + .and_then(|d| d.get(b"Annots").ok()) + .and_then(resolve); + let Some(Object::Array(items)) = annots else { + return 0; + }; + items + .iter() + .filter_map(resolve) + .filter(|a| { + a.as_dict() + .ok() + .and_then(|d| d.get(b"Subtype").ok()) + .and_then(|s| s.as_name().ok()) + == Some(b"Link".as_slice()) + }) + .count() +} + +/// review-T5 L5: the editor lists a file's links as link objects; a Save made only of text +/// changes therefore means the user deleted every link, and the links must not come back (the +/// L7 rule of an empty edit). +#[test] +fn l5_a_text_only_save_keeps_no_link_the_user_deleted() { + let Some(t) = E2e::new("l5_links") else { + return; + }; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let link = d.add( + "<< /Type /Annot /Subtype /Link /Rect [72 600 200 620] /Border [0 0 0] \ + /A << /S /URI /URI (https://example.com/) >> >>", + ); + d.page( + PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Linked page) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + ) + .with(&format!("/Annots [{link} 0 R]")), + ); + let src = t.file("links.pdf", &d.build()); + assert_eq!(link_count(&src, 0), 1, "the source has a link"); + let edit = t.edit(&src, 0, 0, "Linked page", "Linked pages"); + let saved = t.save(&[(&src, "1-z")], vec![edit]).expect("L5 save"); + t.assert_saved(&saved.path, &[(&src, 0, 0, "Linked page", "Linked pages")]); + assert_eq!(link_count(&saved.path, 0), 0, "the deleted link came back"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate.rs new file mode 100644 index 0000000..57f5eb7 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate.rs @@ -0,0 +1,336 @@ +//! T4 gate tests (SPEC §E.5, qpdf and Poppler required): GATE-01…08 and 11…13 +//! (`tests_gate/plan_bugs.rs`: a planner bug written consistently, so A0–A3 pass and A4 must +//! catch it), GATE-09, 10 and 14…17 (`writer_bugs.rs`: qpdf wrote other bytes than planned), +//! GATE-18…22 and 24 (`cover.rs`: cover-and-overlay, raster and annotation fakes, and the #34 +//! composition), GATE-23 and 26…29 (`collateral.rs`), PB-01…08 (`phase_b.rs`), the preview +//! core under the same Phase A (`preview.rs`), the gate's cost bounds and cancellation +//! (`bounds.rs`), and the regressions of the T4 review (`review_fixes.rs`). Each +//! case builds the honest plan, writes it through `apply_update`, runs `verify_edited_copy` +//! (which must pass), then produces a tampered file and asserts the failing check id in the +//! error details. Later checks are reached by skipping earlier ones through the `#[cfg(test)]` +//! seams, so every listed check is shown to fail on its own. Fakes are built with +//! `testkit/fakes.rs` and real qpdf only. + +mod bounds; +mod collateral; +mod cover; +mod phase_b; +mod plan_bugs; +mod preview; +mod review_fixes; +mod writer_bugs; + +use crate::error::AppError; +use crate::pdf_engine::text_edit::apply::{apply_update, updates_for_plan, write_update_json}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::engines::{Engines, RunOpts}; +use crate::pdf_engine::text_edit::gate::{ + test_seams, verify_edited_copy, verify_final_output, EditedPageInput, PageProof, PhaseAInput, + PhaseAReport, PopplerRef, +}; +use crate::pdf_engine::text_edit::graph::{graph_digest, GraphDigest}; +use crate::pdf_engine::text_edit::limits::VERIFY_CAP_MARGIN_BYTES; +use crate::pdf_engine::text_edit::rewrite::{ + assemble_page_plan, plan_page, PagePlan, RunPlan, SourceTextStyleIn, TextEditIn, +}; +use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel}; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use lopdf::ObjectId; +use std::collections::HashMap; +use std::path::{Path, PathBuf}; + +/// One edit of the run whose text is `old`. +pub(crate) struct Ed<'a> { + pub old: &'a str, + pub new: &'a str, + pub style: SourceTextStyleIn, +} + +pub(crate) fn ed<'a>(old: &'a str, new: &'a str) -> Ed<'a> { + Ed { + old, + new, + style: SourceTextStyleIn::default(), + } +} + +/// An honest edit written by qpdf and proven by Phase A, with everything a tampering test needs. +pub(crate) struct Honest { + pub dir: Scratch, + pub engines: Engines, + pub ctx: SnapshotContext, + pub model: PageModel, + pub plan: PagePlan, + pub page: u32, + pub page_count: u32, + /// qpdf's input (the snapshot bytes). + pub source: PathBuf, + /// qpdf's honest output. + pub staged: PathBuf, + pub digest: GraphDigest, + pub report: PhaseAReport, +} + +pub(crate) fn opts() -> RunOpts<'static> { + RunOpts::default() +} + +/// The page model and plan of `edits` on page `page` of `ctx` (all verdicts must be ok). +pub(crate) fn plan_of(ctx: &SnapshotContext, page: u32, edits: &[Ed<'_>]) -> (PageModel, PagePlan) { + let model = build_page_model(ctx, page, None).unwrap_or_else(|e| panic!("model: {e}")); + let ins: Vec = edits + .iter() + .map(|e| { + let run = model + .runs + .iter() + .find(|r| r.text == e.old) + .unwrap_or_else(|| { + panic!( + "no run {:?} in {:?}", + e.old, + model + .runs + .iter() + .map(|r| (&r.text, r.reason)) + .collect::>() + ) + }); + TextEditIn { + run_id: run.id.clone(), + original_text: e.old.to_string(), + text: e.new.to_string(), + style: e.style.clone(), + } + }) + .collect(); + let out = plan_page(ctx, &model, &ins).unwrap_or_else(|e| panic!("plan: {e} {:?}", e.details)); + for v in &out.verdicts { + assert!(v.problem.is_none(), "verdict: {:?}", v.problem); + } + let plan = out.plan.expect("a plan"); + (model, plan) +} + +impl Honest { + /// `None` (after printing `skip:`) when an engine is missing. + pub(crate) fn new(test: &str, pdf: Vec, page: u32, edits: &[Ed<'_>]) -> Option { + let mut h = Self::unverified(test, pdf, page, edits)?; + h.report = h.phase_a(&h.staged).unwrap_or_else(|e| { + panic!( + "{test}: the honest edit must pass Phase A: {e} {:?}", + e.details + ) + }); + Some(h) + } + + /// The honest edit planned and written by qpdf, Phase A not run (`report` is empty). + pub(crate) fn unverified( + test: &str, + pdf: Vec, + page: u32, + edits: &[Ed<'_>], + ) -> Option { + let engines = engines_or_skip(test)?; + let dir = Scratch::new(test); + let source = dir.write("source.pdf", &pdf); + let snap = snapshot_from_bytes(&source, pdf, None).unwrap_or_else(|e| panic!("{e}")); + let page_count = snap.pages.len() as u32; + let ctx = SnapshotContext::new(snap); + let (model, plan) = plan_of(&ctx, page, edits); + let staged = dir.path("edited.pdf"); + let digest = write_plan(&engines, &ctx, &model, &plan, &source, &staged); + Some(Honest { + dir, + engines, + ctx, + model, + plan, + page, + page_count, + source, + staged, + digest, + report: PhaseAReport { + proofs: Vec::new(), + warnings: Vec::new(), + }, + }) + } + + pub(crate) fn qpdf(&self) -> &Path { + &self.engines.qpdf + } + + pub(crate) fn path(&self, name: &str) -> PathBuf { + self.dir.path(name) + } + + /// Phase A of `staged` against this edit's plan and digest. + pub(crate) fn phase_a(&self, staged: &Path) -> Result { + phase_a_with(self, &self.model, &self.plan, &self.digest, staged, &opts()) + } + + /// `phase_a` with the caller's run options (cancel flag, job handle). + pub(crate) fn phase_a_opts( + &self, + staged: &Path, + run_opts: &RunOpts<'_>, + ) -> Result { + phase_a_with( + self, + &self.model, + &self.plan, + &self.digest, + staged, + run_opts, + ) + } + + /// The proof of the honest edit (for Phase B). + pub(crate) fn proof(&self) -> &PageProof { + self.report.proofs.first().expect("a proof") + } + + /// Phase B of `final_pdf` with the honest proof on destination page `dest`. + pub(crate) fn phase_b(&self, final_pdf: &Path, dest: u32) -> Result<(), AppError> { + let cap = 4 * std::fs::metadata(final_pdf).map_or(0, |m| m.len()) + VERIFY_CAP_MARGIN_BYTES; + verify_final_output( + final_pdf, + cap, + &self.engines, + &[(dest, self.proof())], + &opts(), + ) + } + + /// The source stream id of part `i` of the edited page. + pub(crate) fn part_id(&self, i: usize) -> ObjectId { + self.model.content.parts[i].stream_id + } + + /// A planner bug: `tamper` rewrites the primary replacement of the first run; the bytes are + /// re-assembled into a consistent plan, written by qpdf (to `name`) and returned with it. + pub(crate) fn bad_plan( + &self, + name: &str, + tamper: impl Fn(&str) -> String, + ) -> (PagePlan, GraphDigest, PathBuf) { + self.bad_run(name, |run| { + let s = &mut run.splices[0]; + let text = tamper(&String::from_utf8_lossy(&s.bytes)); + s.bytes = text.into_bytes(); + }) + } + + /// A planner bug in the first run plan (bytes and expectations as `f` leaves them). + pub(crate) fn bad_run( + &self, + name: &str, + f: impl Fn(&mut RunPlan), + ) -> (PagePlan, GraphDigest, PathBuf) { + let mut runs: Vec = self.plan.runs.clone(); + f(&mut runs[0]); + let plan = assemble_page_plan(&self.model.content, self.page, runs); + self.write_bad(name, plan) + } + + /// Writes `plan` through qpdf to `name`; returns it with its digest. + pub(crate) fn write_bad(&self, name: &str, plan: PagePlan) -> (PagePlan, GraphDigest, PathBuf) { + let out = self.path(name); + let digest = write_plan( + &self.engines, + &self.ctx, + &self.model, + &plan, + &self.source, + &out, + ); + (plan, digest, out) + } + + /// Phase A of a file written from another (tampered) plan. + pub(crate) fn phase_a_of( + &self, + plan: &PagePlan, + digest: &GraphDigest, + staged: &Path, + ) -> Result { + phase_a_with(self, &self.model, plan, digest, staged, &opts()) + } +} + +/// `write_update_json` + `apply_update` of `plan`, and the "before" digest of qpdf's input. +pub(crate) fn write_plan( + engines: &Engines, + ctx: &SnapshotContext, + model: &PageModel, + plan: &PagePlan, + source: &Path, + out: &Path, +) -> GraphDigest { + let updates = updates_for_plan(&model.content, plan).expect("updates"); + let update = out.with_extension("update.json"); + write_update_json(&updates, ctx.doc().max_id, &update).unwrap_or_else(|e| panic!("{e}")); + apply_update(engines, source, &update, out, &[], &opts()) + .unwrap_or_else(|e| panic!("{e} {:?}", e.details)); + let replaced: HashMap> = updates + .into_iter() + .map(|u| (u.object_id, u.decoded)) + .collect(); + graph_digest(ctx.doc(), &replaced, None).unwrap_or_else(|m| panic!("digest: {m:?}")) +} + +fn phase_a_with( + h: &Honest, + model: &PageModel, + plan: &PagePlan, + digest: &GraphDigest, + staged: &Path, + run_opts: &RunOpts<'_>, +) -> Result { + let cap = 2 * std::fs::metadata(&h.source).map_or(0, |m| m.len()) + VERIFY_CAP_MARGIN_BYTES; + let input = PhaseAInput { + before: digest, + before_page_count: h.page_count, + staged, + staged_cap: cap, + pages: vec![EditedPageInput { + model: model.into(), + plan, + input_page_index: h.page, + input_render: PopplerRef { + pdf: h.source.clone(), + page_1: h.page + 1, + }, + }], + source_benign: &[], + }; + verify_edited_copy(&input, &h.engines, h.dir.dir(), run_opts) +} + +/// `r` must fail with `code` and its details must name `check` (and contain `what` when given). +pub(crate) fn fails_at(r: Result, code: &str, check: &str, what: &str, id: &str) { + let e = match r { + Ok(_) => panic!("{id}: expected a failure at {check}, the check passed"), + Err(e) => e, + }; + let details = e.details.clone().unwrap_or_default(); + assert_eq!(e.code, code, "{id}: code ({details})"); + assert!( + details.contains(&format!("check={check}")), + "{id}: expected check={check}, got {details}" + ); + assert!( + details.contains(what), + "{id}: expected {what:?} in {details}" + ); +} + +/// Runs `f` with the named checks skipped (test seam). +pub(crate) fn skipping(checks: &[&str], f: impl FnOnce() -> T) -> T { + let _guard = test_seams::skip(checks); + f() +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate/bounds.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate/bounds.rs new file mode 100644 index 0000000..4646d75 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate/bounds.rs @@ -0,0 +1,379 @@ +//! The gate's cost bounds and cancellation (T4, beyond §E.5): A5's pixel check, A5's word +//! matching and B1's wrapper search stay linear on inputs that are quadratic for a per-pixel × +//! per-mask scan, a first-unused-match scan or a window-by-window search, and a check that fails +//! while the job is being cancelled reports `CANCELLED` (Phase A, Phase B and the preview), never +//! a verification code. + +use super::{ed, fails_at, opts, skipping, Honest}; +use crate::pdf_engine::text_edit::engines::RunOpts; +use crate::pdf_engine::text_edit::gate::originals::find_up_to; +use crate::pdf_engine::text_edit::gate::verify_final_output; +use crate::pdf_engine::text_edit::geometry::PageGeometry; +use crate::pdf_engine::text_edit::limits::{RENDER_OUTSIDE_PIXELS_MAX, VERIFY_CAP_MARGIN_BYTES}; +use crate::pdf_engine::text_edit::poppler::near::KeptGlyph; +use crate::pdf_engine::text_edit::poppler::{ + check_independent, check_pixels, NearGlyphs, Raster, Word, +}; +use crate::pdf_engine::text_edit::preview::{cache_dir_for, preview_page}; +use crate::pdf_engine::text_edit::rewrite::{ExpectedRun, SourceTextStyleIn, TextEditIn}; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::testkit::fakes; +use crate::pdf_engine::text_edit::testkit::producers::helvetica_page; +use crate::pdf_engine::text_edit::tests_plan::{ok_plan, plan_one, style, thread_cpu}; +use std::sync::atomic::AtomicBool; +use std::time::Duration; + +/// A page 612 × 792 pt; rendered at 72 DPI one point is one pixel, y down from the top. +const PAGE_W: u32 = 612; +const PAGE_H: u32 = 792; +const DPI: u32 = 72; + +/// An expectation of a real plan (A5 reads only its masks and whether its glyphs changed) and the +/// page geometry it was planned on. +fn template() -> (ExpectedRun, PageGeometry) { + let (_, m, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Help", + style(), + ); + let exp = ok_plan(&out).runs[0].expected.clone(); + (exp, m.walk.geometry.clone()) +} + +fn edit_with( + template: &ExpectedRun, + masks: Vec<[f64; 4]>, +) -> (ExpectedRun, [f64; 4], [f64; 4], String) { + let mut exp = template.clone(); + exp.mask_boxes = masks; + exp.glyphs_changed = true; + (exp, [0.0; 4], [0.0; 4], String::new()) +} + +/// No glyph of our model around the edits. +fn none() -> NearGlyphs { + NearGlyphs::default() +} + +fn white() -> Raster { + Raster { + w: PAGE_W, + h: PAGE_H, + rgb: vec![255; (PAGE_W * PAGE_H * 3) as usize], + } +} + +/// Darkens the render pixel under user point (`x`, `y`) at 72 DPI. +fn ink(r: &mut Raster, x: u32, y: u32) { + let row = PAGE_H - 1 - y; + let i = ((row * PAGE_W + x) * 3) as usize; + r.rgb[i..i + 3].copy_from_slice(&[0, 0, 0]); +} + +/// review-verify HIGH-A: a glyph no edit changes that lies under an edit's masks (here a 4 pt +/// follower within the slack of a 10 pt glyph's mask) is checked outside the edited glyph's own +/// box grown by `RENDER_OWN_PAD_PX`: more than `RENDER_KEPT_PIXELS_MAX` changed pixels there +/// fail, fewer and those under the edited glyph do not. +#[test] +fn kept_glyphs_under_the_masks_are_checked_outside_the_edited_glyphs_own_boxes() { + let (t, geom) = template(); + let mut exp = t; + exp.glyph_origins = vec![(100.0, 700.0)]; + exp.mask_boxes = vec![[100.0, 700.0, 110.0, 710.0]]; + exp.old_ink_boxes = Vec::new(); + exp.glyphs_changed = true; + let edits = vec![(exp, [0.0; 4], [0.0; 4], String::new())]; + let follower = KeptGlyph { + rect: [112.0, 700.0, 116.0, 710.0], + text: Some("1".into()), + }; + let near = NearGlyphs { + kept: vec![follower], + edited: Vec::new(), + }; + let src = white(); + let changed = |xs: &[u32]| { + let mut dst = white(); + for x in xs { + ink(&mut dst, *x, 705); + } + dst + }; + let three = changed(&[113, 114, 115]); + let e = check_pixels(&src, &three, &geom, &edits, &near, DPI).expect_err("follower changed"); + assert_eq!(e.0, "render"); + assert!( + e.1.contains("a glyph no edit changes changed, pixels=3"), + "{}", + e.1 + ); + assert!( + check_pixels(&src, &three, &geom, &edits, &none(), DPI).is_ok(), + "without our model the follower lies under the masks" + ); + let within = changed(&[113, 114]); + assert!(check_pixels(&src, &within, &geom, &edits, &near, DPI).is_ok()); + let own = changed(&[108, 109, 110]); + assert!( + check_pixels(&src, &own, &geom, &edits, &near, DPI).is_ok(), + "the edited glyph's own box and its one-pixel pad are the edit's" + ); +} + +#[test] +fn pixel_check_counts_each_outside_pixel_once_and_marks_each_edit() { + let (t, geom) = template(); + // Two overlapping masks (edits A and B) and one far away (C). + let edits = vec![ + edit_with(&t, vec![[100.0, 700.0, 110.0, 710.0]]), + edit_with(&t, vec![[105.0, 700.0, 120.0, 710.0]]), + edit_with(&t, vec![[400.0, 100.0, 410.0, 110.0]]), + ]; + let src = white(); + let mut dst = white(); + // One changed pixel inside A ∩ B: both are visible, nothing is outside. + ink(&mut dst, 107, 705); + let visible = check_pixels(&src, &dst, &geom, &edits, &none(), DPI).expect("inside the masks"); + assert_eq!(visible, vec![true, true, false], "A and B visible, C not"); + // RENDER_OUTSIDE_PIXELS_MAX changed pixels outside every mask still pass … + for k in 0..RENDER_OUTSIDE_PIXELS_MAX as u32 { + ink(&mut dst, 300 + k, 400); + } + assert!(check_pixels(&src, &dst, &geom, &edits, &none(), DPI).is_ok()); + // … one more fails, counted once although the masks overlap elsewhere. + ink(&mut dst, 300, 390); + let e = check_pixels(&src, &dst, &geom, &edits, &none(), DPI).expect_err("one too many"); + assert_eq!( + e, + ( + "render", + format!("pixels={}", RENDER_OUTSIDE_PIXELS_MAX + 1) + ) + ); + // A different render size is refused before any pixel is compared. + let small = Raster { + w: 10, + h: 10, + rgb: vec![255; 300], + }; + assert_eq!( + check_pixels(&src, &small, &geom, &edits, &none(), DPI).map_err(|e| e.0), + Err("render") + ); +} + +#[test] +fn pixel_check_stays_linear_with_many_edits_and_dense_changes() { + // 200 edits of 100 mask boxes each (the most a page allows, EDITS_PER_PAGE_MAX), and every + // pixel inside every box changed: ~300,000 changed pixels against 20,000 masks. A per-pixel + // scan of every mask would test ~6·10⁹ boxes; the row sweep is one pass over the render. + let (t, geom) = template(); + let src = white(); + let mut dst = white(); + let mut edits = Vec::new(); + for e in 0..200u32 { + let (col, row) = (e % 6, e / 6); + let (x0, y0) = (6 + col * 100, 40 + row * 22); + let mut masks = Vec::new(); + for k in 0..100u32 { + let (x, y) = (x0 + (k % 25) * 4, y0 + (k / 25) * 4); + masks.push([ + f64::from(x), + f64::from(y), + f64::from(x + 4), + f64::from(y + 4), + ]); + for dx in 0..4 { + for dy in 0..4 { + ink(&mut dst, x + dx, y + dy); + } + } + } + edits.push(edit_with(&t, masks)); + } + let started = thread_cpu(); + let visible = + check_pixels(&src, &dst, &geom, &edits, &none(), DPI).expect("every change is masked"); + let spent = thread_cpu().saturating_sub(started); + assert!(visible.iter().all(|v| *v), "every edit changed pixels"); + assert!( + spent < Duration::from_secs(2), + "pixel check over 200 × 100 masks took {spent:?}" + ); +} + +/// G-TEXT inputs: `n` words "ab" on a grid away from the edited line (frame y ≈ 90), and the +/// edited word itself ("Hello" before, "Help" after). +fn grid_words(n: usize) -> (Vec, Vec) { + let word = |text: &str, x: f64, y: f64| Word { + text: text.into(), + x0: x, + y0: y, + x1: x + 4.0, + y1: y + 0.4, + }; + let mut src = vec![word("Hello", 72.0, 85.0)]; + for i in 0..n { + src.push(word( + "ab", + (i % 100) as f64 * 5.0, + 200.0 + (i / 100) as f64 * 0.5, + )); + } + let mut dst = src.clone(); + dst[0].text = "Help".into(); + (src, dst) +} + +#[test] +fn word_matching_stays_linear_in_the_word_count() { + // 200,000 words (more than a capped pdftotext output holds), in reverse order on the after + // side: a "first unused match" scan of the after list would compare ~2·10¹⁰ pairs. + let (t, geom) = template(); + let edits = vec![( + t, + [72.0, 695.0, 100.0, 710.0], + [72.0, 695.0, 100.0, 710.0], + "Hello".to_string(), + )]; + let raster = Raster { + w: 1, + h: 1, + rgb: vec![255; 3], + }; + let (src, mut dst) = grid_words(200_000); + dst[1..].reverse(); + let started = thread_cpu(); + let r = check_independent(&src, &dst, &raster, &raster, &geom, &edits, &none(), DPI); + let spent = thread_cpu().saturating_sub(started); + assert!(r.is_ok(), "the same words in another order match: {r:?}"); + assert!( + spent < Duration::from_secs(2), + "word matching took {spent:?}" + ); + // A moved word is still found, wherever it is in the list. + let (src, mut dst) = grid_words(1_000); + dst[500].x0 += 0.06; + let r = check_independent(&src, &dst, &raster, &raster, &geom, &edits, &none(), DPI); + assert!( + matches!(&r, Err(("words", d)) if d.contains("moved or changed")), + "{r:?}" + ); + // Words of one text a hair apart, ordered against each other: the search gives up (fails + // closed) instead of scanning quadratically. + let (mut src, _) = grid_words(0); + let n = 40_000usize; + for i in 0..n { + let x1 = if i < n / 2 { 10.0 } else { 10.09 }; + src.push(Word { + text: "ab".into(), + x0: 0.0, + y0: 300.0, + x1, + y1: 301.0, + }); + } + let mut dst = src.clone(); + dst[0].text = "Help".into(); + dst[1..].reverse(); + let started = thread_cpu(); + let r = check_independent(&src, &dst, &raster, &raster, &geom, &edits, &none(), DPI); + let spent = thread_cpu().saturating_sub(started); + assert!( + matches!(&r, Err(("words", d)) if d.contains("too many overlapping words")), + "{r:?}" + ); + assert!( + spent < Duration::from_secs(2), + "bounded word search took {spent:?}" + ); +} + +#[test] +fn wrapper_search_stays_linear_on_repetitive_data() { + // "aaa…a" against "aaa…ab": a window-by-window comparison would read ~3·10¹² bytes. + let data = vec![b'a'; 4 << 20]; + let mut needle = vec![b'a'; 1 << 20]; + needle.push(b'b'); + let started = thread_cpu(); + assert_eq!(find_up_to(&data, &needle, 2), Ok(Vec::new())); + // Every window a hit: the search stops at the second one. + let short = vec![b'a'; 1 << 16]; + assert_eq!(find_up_to(&data, &short, 2), Ok(vec![0, 1])); + let spent = thread_cpu().saturating_sub(started); + assert!( + spent < Duration::from_secs(2), + "wrapper search took {spent:?}" + ); + // Exact positions, the end of the data, and needles that cannot fit. + assert_eq!(find_up_to(b"x\nABC\nABC", b"ABC", 5), Ok(vec![2, 6])); + assert_eq!(find_up_to(b"zzABC", b"ABC", 5), Ok(vec![2])); + assert_eq!(find_up_to(b"AB", b"ABC", 5), Ok(Vec::new())); + assert_eq!(find_up_to(b"ABC", b"", 5), Ok(Vec::new())); +} + +#[test] +fn failures_while_cancelled_report_cancelled() { + let pdf = helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Invoice 2026) Tj ET BT /F1 12 Tf 72 650 Td (Unchanged line) Tj ET", + ); + let Some(h) = Honest::new( + "bounds_cancel", + pdf, + 0, + &[ed("Invoice 2026", "Invoice 2027")], + ) else { + return; + }; + let fake = h.path("fake.pdf"); + fakes::cover_and_overlay_f1( + h.qpdf(), + &h.source, + &fake, + 0, + [70.0, 695.0, 90.0, 16.0], + "F1", + "Invoice 2027", + ) + .unwrap_or_else(|e| panic!("{e}")); + let flag = AtomicBool::new(true); + let cancelled = RunOpts { + cancel: Some(&flag), + ..RunOpts::default() + }; + // With no subprocess before it, the fake fails A2 — reported as the cancel it happened under. + let r = skipping(&["A0", "A1"], || h.phase_a_opts(&fake, &cancelled)); + assert_eq!(r.err().map(|e| e.code), Some("CANCELLED".to_string())); + // Not cancelled, the same file fails A2 with its verification code. + let r = skipping(&["A0", "A1"], || h.phase_a_opts(&fake, &opts())); + fails_at(r, "EDIT_VERIFY_FAILED", "A2", "path=", "cancel control"); + // Phase B and the preview stop with CANCELLED too. + let cap = 4 * std::fs::metadata(&fake).map_or(0, |m| m.len()) + VERIFY_CAP_MARGIN_BYTES; + let r = verify_final_output(&fake, cap, &h.engines, &[(0, h.proof())], &cancelled); + assert_eq!(r.err().map(|e| e.code), Some("CANCELLED".to_string())); + let run = h + .model + .runs + .iter() + .find(|r| r.text == "Invoice 2026") + .expect("run"); + let edit = TextEditIn { + run_id: run.id.clone(), + original_text: run.text.clone(), + text: "Invoice 2027".into(), + style: SourceTextStyleIn::default(), + }; + let cache = cache_dir_for(h.dir.dir(), "cancelled"); + let model = build_page_model(&h.ctx, h.page, None).expect("model"); + let r = preview_page( + &h.ctx, + std::sync::Arc::new(model), + &[edit], + &cache, + &h.engines, + &[], + &cancelled, + ); + assert_eq!(r.err().map(|e| e.code), Some("CANCELLED".to_string())); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate/collateral.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate/collateral.rs new file mode 100644 index 0000000..ff447ef --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate/collateral.rs @@ -0,0 +1,274 @@ +//! GATE-23 and 26…29: collateral damage — marked content dropped around the run, an annotation +//! on another page, a layer switched off, a font's widths changed on an unedited page, and an +//! untouched legacy-filter page altered. + +use super::{ed, fails_at, skipping, Honest}; +use crate::pdf_engine::text_edit::testkit::fakes; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, word_font, DocBuilder, LegacyFilter, PageSpec, HELVETICA, +}; +use lopdf::{Document, Object, ObjectId}; +use serde_json::json; + +/// The id of `/Resources / /` of page `page` in `doc` (a reference). +fn resource_id(doc: &Document, page: usize, category: &[u8], name: &[u8]) -> ObjectId { + let page_id = fakes::page_ids(doc)[page]; + let res = doc + .get_dictionary(page_id) + .and_then(|d| d.get(b"Resources")) + .and_then(|r| match r { + Object::Reference(id) => doc.get_dictionary(*id), + Object::Dictionary(d) => Ok(d), + _ => Err(lopdf::Error::DictKey), + }) + .expect("resources"); + let cat = match res.get(category).expect("category") { + Object::Reference(id) => doc.get_dictionary(*id).expect("category dict"), + Object::Dictionary(d) => d, + _ => panic!("category"), + }; + cat.get(name) + .and_then(Object::as_reference) + .expect("resource reference") +} + +#[test] +fn gate_23_marked_content_removed_around_the_run() { + let pdf = helvetica_page( + b"/P <> BDC BT /F1 12 Tf 72 700 Td (Hello) Tj ET EMC BT /F1 12 Tf 72 650 Td (Other) Tj ET", + ); + let Some(h) = Honest::new("gate_23", pdf, 0, &[ed("Hello", "Help")]) else { + return; + }; + // The writer blanks the BDC and EMC operators (same length: no span moves). + let honest = String::from_utf8_lossy(&h.plan.expected_parts[0]).into_owned(); + let blank = |s: &str, op: &str| s.replacen(op, &" ".repeat(op.len()), 1); + let bad = blank(&blank(&honest, "/P <> BDC"), "EMC"); + let out = h.path("bad.pdf"); + fakes::replace_streams( + h.qpdf(), + &h.source, + &out, + &[(h.part_id(0), bad.into_bytes())], + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A2", + "what=data", + "GATE-23", + ); + let r = skipping(&["A2", "A3"], || h.phase_a(&out)); + fails_at(r, "STATE_CHANGED", "A4", "field=marked", "GATE-23 A4"); +} + +/// Page 1 (edited) in Helvetica; page 2 in a Word subset with `/Widths` and an `/F1` of its own. +fn two_pages() -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let w = word_font(&mut d.b, "ABCDEF+Calibri", "Second page"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Second page) Tj ET", + &format!("/Font << /F1 {w} 0 R >>"), + )); + d.build() +} + +#[test] +fn gate_26_annotation_added_on_another_page() { + let Some(h) = Honest::new("gate_26", two_pages(), 0, &[ed("Hello", "Help")]) else { + return; + }; + let out = h.path("link.pdf"); + fakes::link(h.qpdf(), &h.staged, &out, 1).unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A2", + "path=/Root/Pages/Kids[1]", + "GATE-26 link", + ); + let out = h.path("freetext.pdf"); + fakes::freetext( + h.qpdf(), + &h.staged, + &out, + 1, + [72.0, 690.0, 200.0, 712.0], + "Note", + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A2", + "path=/Root/Pages/Kids[1]", + "GATE-26 FreeText", + ); +} + +#[test] +fn gate_27_layer_switched_off_in_the_catalog() { + let Some(h) = Honest::new( + "gate_27", + fx::indd(), + 0, + &[ed("visible layer", "visible layers")], + ) else { + return; + }; + let staged = fakes::load(&h.staged); + let cat = staged + .get_dictionary(fakes::catalog_id(&staged)) + .expect("catalog"); + let ocgs = cat + .get(b"OCProperties") + .and_then(Object::as_dict) + .and_then(|p| p.get(b"OCGs")) + .and_then(Object::as_array) + .expect("OCGs"); + let visible = ocgs[0].as_reference().expect("ocg ref"); + let out = h.path("off.pdf"); + fakes::ocg_off(h.qpdf(), &h.staged, &out, visible).unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A2", + "path=/Root what=node", + "GATE-27", + ); +} + +#[test] +fn gate_28_widths_changed_on_an_unedited_page() { + let Some(h) = Honest::new("gate_28", two_pages(), 0, &[ed("Hello", "Help")]) else { + return; + }; + let font = resource_id(&fakes::load(&h.staged), 1, b"Font", b"F1"); + let out = h.path("widths.pdf"); + fakes::widths_changed(h.qpdf(), &h.staged, &out, font, 50, 501.0) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A2", + "Kids[1]/Resources/Font/F1", + "GATE-28", + ); +} + +#[test] +fn gate_29_untouched_runlength_page_altered() { + let pdf = fx::legacy_filter_page(LegacyFilter::RunLength); + let Some(h) = Honest::new("gate_29", pdf, 1, &[ed("Hello", "Help")]) else { + return; + }; + let staged = fakes::load(&h.staged); + let page0 = fakes::page_ids(&staged)[0]; + let rl = staged + .get_dictionary(page0) + .and_then(|d| d.get(b"Contents")) + .and_then(Object::as_reference) + .expect("RunLength contents"); + // Literal runs of a different text: still RunLength, decodes differently. + let text = b"BT /F1 12 Tf 72 720 Td (Altered) Tj ET"; + let mut raw = vec![(text.len() - 1) as u8]; + raw.extend_from_slice(text); + raw.push(128); + let out = h.path("rl.pdf"); + fakes::raw_stream( + h.qpdf(), + &h.staged, + &out, + rl, + json!({ "/Filter": "/RunLengthDecode" }), + &raw, + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A2", + "Kids[0]/Contents what=data", + "GATE-29", + ); +} + +#[test] +fn phase_a_and_b_with_two_edited_pages_in_one_copy() { + use super::{opts, plan_of}; + use crate::pdf_engine::text_edit::apply::{apply_update, updates_for_plan, write_update_json}; + use crate::pdf_engine::text_edit::context::SnapshotContext; + use crate::pdf_engine::text_edit::gate::{ + verify_edited_copy, verify_final_output, BeforeModel, EditedPageInput, PhaseAInput, + PopplerRef, + }; + use crate::pdf_engine::text_edit::graph::graph_digest; + use crate::pdf_engine::text_edit::limits::VERIFY_CAP_MARGIN_BYTES; + use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; + use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; + use std::collections::HashMap; + // Pages 1 and 3 edited in one qpdf update (the Save shape for one source file). + let Some(engines) = engines_or_skip("phase_a_two_pages") else { + return; + }; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + for text in ["First page", "Middle page", "Last page"] { + d.page(PageSpec::new( + format!("BT /F1 12 Tf 72 700 Td ({text}) Tj ET").as_bytes(), + &res, + )); + } + let pdf = d.build(); + let dir = Scratch::new("phase_a_two_pages"); + let source = dir.write("source.pdf", &pdf); + let ctx = SnapshotContext::new(snapshot_from_bytes(&source, pdf, None).expect("snapshot")); + let (m0, p0) = plan_of(&ctx, 0, &[ed("First page", "First pages")]); + let (m2, p2) = plan_of(&ctx, 2, &[ed("Last page", "Last pages")]); + let mut updates = updates_for_plan(&m0.content, &p0).expect("updates"); + updates.extend(updates_for_plan(&m2.content, &p2).expect("updates")); + let update = dir.path("update.json"); + write_update_json(&updates, ctx.doc().max_id, &update).expect("json"); + let staged = dir.path("edited.pdf"); + apply_update(&engines, &source, &update, &staged, &[], &opts()).expect("qpdf"); + let replaced: HashMap<_, _> = updates + .into_iter() + .map(|u| (u.object_id, u.decoded)) + .collect(); + let digest = graph_digest(ctx.doc(), &replaced, None).expect("digest"); + let page = |model, plan, i: u32| EditedPageInput { + model: BeforeModel::Kept(model), + plan, + input_page_index: i, + input_render: PopplerRef { + pdf: source.clone(), + page_1: i + 1, + }, + }; + let input = PhaseAInput { + before: &digest, + before_page_count: 3, + staged: &staged, + staged_cap: 2 * std::fs::metadata(&source).map_or(0, |m| m.len()) + VERIFY_CAP_MARGIN_BYTES, + pages: vec![page(&m0, &p0, 0), page(&m2, &p2, 2)], + source_benign: &[], + }; + let report = verify_edited_copy(&input, &engines, dir.dir(), &opts()) + .unwrap_or_else(|e| panic!("Phase A: {e} {:?}", e.details)); + assert_eq!(report.proofs.len(), 2); + let expectations = [(0, &report.proofs[0]), (2, &report.proofs[1])]; + let cap = 4 * std::fs::metadata(&staged).map_or(0, |m| m.len()) + VERIFY_CAP_MARGIN_BYTES; + let r = verify_final_output(&staged, cap, &engines, &expectations, &opts()); + assert!(r.is_ok(), "Phase B: {:?}", r.err().map(|e| e.details)); + // The proofs swapped between the pages: Phase B refuses. + let swapped = [(0, &report.proofs[1]), (2, &report.proofs[0])]; + let r = verify_final_output(&staged, cap, &engines, &swapped, &opts()); + fails_at(r, "SOURCE_EDIT_GATE_FAILED", "B1", "", "two pages swapped"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate/cover.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate/cover.rs new file mode 100644 index 0000000..9073133 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate/cover.rs @@ -0,0 +1,323 @@ +//! GATE-18…22 and 24: cover-and-overlay (a new content stream, or a real `qpdf --overlay`), a +//! raster of the page, an annotation over the old text and the old stream re-attached — each +//! fails Phase A at every listed check and, where listed, Phase B — and GATE-21: why #34's +//! `validate_staged_pdf` alone is not enough. + +use super::{ed, fails_at, skipping, Honest}; +use crate::pdf_engine::text_edit::testkit::fakes; +use crate::pdf_engine::text_edit::testkit::producers::helvetica_page; +use crate::pdf_engine::validate_output::{ + catalog_flags_from_doc, validate_staged_pdf, ContentDigest, OutputSnapshot, PageSnapshot, +}; +use std::path::Path; + +const MEDIA: [f64; 4] = [0.0, 0.0, 612.0, 792.0]; + +fn page() -> Vec { + helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Invoice 2026) Tj ET BT /F1 12 Tf 72 650 Td (Unchanged line) Tj ET", + ) +} + +fn honest(test: &str) -> Option { + Honest::new(test, page(), 0, &[ed("Invoice 2026", "Invoice 2027")]) +} + +/// The cover box over the edited line (`[x, y, w, h]`). +fn cover_rect(h: &Honest) -> [f64; 4] { + let r = h + .model + .runs + .iter() + .find(|r| r.text == "Invoice 2026") + .expect("run") + .rect; + [r[0] - 1.0, r[1] - 1.0, r[2] + 2.0, r[3] + 2.0] +} + +/// A fails at A2, A3, A4 and A5 in turn (each with the earlier ones skipped). +fn fails_a2_to_a5(h: &Honest, file: &Path, a3: &str, id: &str) { + fails_at( + h.phase_a(file), + "EDIT_VERIFY_FAILED", + "A2", + "path=/Root/Pages/Kids[0]", + id, + ); + let r = skipping(&["A2"], || h.phase_a(file)); + fails_at(r, "EDIT_VERIFY_FAILED", "A3", a3, &format!("{id} A3")); + let r = skipping(&["A2", "A3"], || h.phase_a(file)); + fails_at(r, "EDIT_VERIFY_FAILED", "A4", "", &format!("{id} A4")); + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(file)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", "", &format!("{id} A5")); +} + +#[test] +fn gate_18_cover_and_overlay_f1_fails_phase_a_and_phase_b() { + let Some(h) = honest("gate_18") else { return }; + let fake = h.path("fake.pdf"); + fakes::cover_and_overlay_f1( + h.qpdf(), + &h.source, + &fake, + 0, + cover_rect(&h), + "F1", + "Invoice 2027", + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_a2_to_a5( + &h, + &fake, + "original content is still attached; part count 2", + "GATE-18", + ); + // Phase B on a final file that carries the fake. + fails_at( + h.phase_b(&fake, 0), + "SOURCE_EDIT_GATE_FAILED", + "B1", + "not on the page", + "GATE-18 B1", + ); + let r = skipping(&["B1"], || h.phase_b(&fake, 0)); + fails_at( + r, + "SOURCE_EDIT_GATE_FAILED", + "B2", + "original content", + "GATE-18 B2", + ); + let r = skipping(&["B1", "B2"], || h.phase_b(&fake, 0)); + fails_at( + r, + "SOURCE_EDIT_GATE_FAILED", + "B3", + "show records", + "GATE-18 B3", + ); +} + +#[test] +fn gate_19_cover_and_overlay_f2_real_qpdf_overlay() { + let Some(h) = honest("gate_19") else { return }; + let fake = h.path("fake.pdf"); + fakes::overlay_f2( + h.qpdf(), + &h.source, + &fake, + 1, + MEDIA, + cover_rect(&h), + "Invoice 2027", + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&fake), + "EDIT_VERIFY_FAILED", + "A2", + "path=/Root/Pages/Kids[0]", + "GATE-19", + ); + let r = skipping(&["A2"], || h.phase_a(&fake)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A3", + "original content is still attached", + "GATE-19 A3", + ); + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(&fake)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", "", "GATE-19 A5"); + // The wrapper holds the old parts: neither form of B1 finds the expected content. + fails_at( + h.phase_b(&fake, 0), + "SOURCE_EDIT_GATE_FAILED", + "B1", + "", + "GATE-19 B1", + ); + let r = skipping(&["B1"], || h.phase_b(&fake, 0)); + fails_at( + r, + "SOURCE_EDIT_GATE_FAILED", + "B2", + "original content", + "GATE-19 B2", + ); +} + +#[test] +fn gate_20_raster_fake() { + let Some(h) = honest("gate_20") else { return }; + let fake = h.path("fake.pdf"); + fakes::raster_fake( + h.qpdf(), + &h.engines.pdftoppm, + &h.staged, + &h.source, + &fake, + 0, + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_at(h.phase_a(&fake), "EDIT_VERIFY_FAILED", "A2", "", "GATE-20"); + let r = skipping(&["A2"], || h.phase_a(&fake)); + fails_at(r, "EDIT_VERIFY_FAILED", "A3", "bytes differ", "GATE-20 A3"); + let r = skipping(&["A2", "A3"], || h.phase_a(&fake)); + fails_at(r, "EDIT_VERIFY_FAILED", "A4", "", "GATE-20 A4"); +} + +/// `qpdf --check` for `validate_staged_pdf`. +fn check_runner( + qpdf: &Path, +) -> impl FnMut(&[String]) -> Result<(i32, String), crate::error::AppError> + '_ { + move |args| { + let out = std::process::Command::new(qpdf) + .args(args) + .output() + .expect("qpdf"); + Ok(( + out.status.code().unwrap_or(2), + String::from_utf8_lossy(&out.stderr).into_owned(), + )) + } +} + +fn snapshot_with(h: &Honest, digest: ContentDigest) -> OutputSnapshot { + OutputSnapshot { + pages: vec![PageSnapshot { + media_box: MEDIA, + crop_box: None, + trim_box: None, + rotate: 0, + user_unit: 1.0, + content_digest: digest, + }], + catalog: catalog_flags_from_doc(h.ctx.doc()), + } +} + +#[test] +fn gate_21_phase_b_composes_with_34() { + let Some(h) = honest("gate_21") else { return }; + let copy = |name: &str, from: &Path| { + let p = h.path(name); + std::fs::copy(from, &p).expect("copy"); + p + }; + // #34 with the source page digest refuses the honest edit (the content changed) … + let source_digest = h.model.content.concat_digest(); + let staged = copy("honest-34.pdf", &h.staged); + let r = validate_staged_pdf( + &staged, + &snapshot_with(&h, source_digest), + None, + check_runner(h.qpdf()), + ); + assert!( + r.is_err(), + "GATE-21 #34 with the source digest refuses an honest edit" + ); + // … and passes it with the proof's independently computed digest. + let staged = copy("honest-proof.pdf", &h.staged); + let r = validate_staged_pdf( + &staged, + &snapshot_with(&h, h.proof().expected_page_digest), + None, + check_runner(h.qpdf()), + ); + assert!( + r.is_ok(), + "GATE-21 #34 with the proof digest: {:?}", + r.err().map(|e| e.details) + ); + assert!( + h.phase_b(&h.staged, 0).is_ok(), + "GATE-21 Phase B passes the honest edit" + ); + // GATE-18's fake keeps the source stream: #34 alone passes it, Phase B does not. + let fake = h.path("fake.pdf"); + fakes::cover_and_overlay_f1( + h.qpdf(), + &h.source, + &fake, + 0, + cover_rect(&h), + "F1", + "Invoice 2027", + ) + .unwrap_or_else(|e| panic!("{e}")); + let staged = copy("fake-34.pdf", &fake); + let r = validate_staged_pdf( + &staged, + &snapshot_with(&h, source_digest), + None, + check_runner(h.qpdf()), + ); + assert!( + r.is_ok(), + "GATE-21 #34 alone passes the fake: {:?}", + r.err().map(|e| e.details) + ); + fails_at( + h.phase_b(&fake, 0), + "SOURCE_EDIT_GATE_FAILED", + "B1", + "", + "GATE-21 Phase B", + ); +} + +#[test] +fn gate_22_freetext_annotation_over_the_old_text() { + let Some(h) = honest("gate_22") else { return }; + let fake = h.path("fake.pdf"); + let c = cover_rect(&h); + fakes::freetext( + h.qpdf(), + &h.staged, + &fake, + 0, + [c[0], c[1], c[0] + c[2], c[1] + c[3]], + "Invoice 2027", + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&fake), + "EDIT_VERIFY_FAILED", + "A2", + "path=/Root/Pages/Kids[0]", + "GATE-22", + ); +} + +#[test] +fn gate_24_old_stream_reattached_as_a_form() { + let Some(h) = honest("gate_24") else { return }; + let fake = h.path("fake.pdf"); + let old = h.model.content.part_bytes(0).to_vec(); + fakes::reattach_old(h.qpdf(), &h.staged, &fake, 0, &old).unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&fake), + "EDIT_VERIFY_FAILED", + "A2", + "path=/Root/Pages/Kids[0]", + "GATE-24", + ); + let r = skipping(&["A2"], || h.phase_a(&fake)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A3", + "original content is still attached", + "GATE-24 A3", + ); + // Phase B sees it too: B1 finds the expected parts, B2 the old stream in the resources. + fails_at( + h.phase_b(&fake, 0), + "SOURCE_EDIT_GATE_FAILED", + "B2", + "original content", + "GATE-24 B2", + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate/phase_b.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate/phase_b.rs new file mode 100644 index 0000000..1b4d2b2 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate/phase_b.rs @@ -0,0 +1,436 @@ +//! PB-01…08: Phase B on final files made by real qpdf passes after an honest edit — the parts +//! form, qpdf's overlay wrapper (also on rotated and cropped pages with the boxes remapped as the +//! Save pipeline does), stamps, flattened parts and an appended signed file pass; fakes inside a +//! wrapper and a wrapper that clips the page fail. + +use super::{ed, fails_at, skipping, Honest}; +use crate::pdf_engine::text_edit::content::{page_content, qpdf_join}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::lexer::{lex_content, LexLimits, Operator}; +use crate::pdf_engine::text_edit::limits::{ + set_file_cap_override, PAGE_DECODE_BUDGET, VERIFY_CAP_MARGIN_BYTES, +}; +use crate::pdf_engine::text_edit::snapshot::{read_snapshot, read_verification_snapshot}; +use crate::pdf_engine::text_edit::testkit::fakes::{self, json_dict}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_doc, helvetica_page, DocBuilder, PageSpec, +}; +use crate::pdf_engine::validate_output::content_digest; +use lopdf::Object; +use serde_json::json; +use std::path::{Path, PathBuf}; + +fn page() -> Vec { + helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hello world) Tj ET BT /F1 12 Tf 72 650 Td (Second line) Tj ET", + ) +} + +fn honest(test: &str) -> Option { + Honest::new(test, page(), 0, &[ed("Hello world", "Hello there")]) +} + +/// One empty page (`blank.pdf`). +fn blank(h: &Honest) -> PathBuf { + let mut d = DocBuilder::new(); + d.page(PageSpec::new(b"", "")); + let p = h.path("blank.pdf"); + std::fs::write(&p, d.build()).expect("blank"); + p +} + +/// A stamp page: a filled square far from the edited line. +fn stamp(h: &Honest) -> PathBuf { + let p = h.path("stamp.pdf"); + std::fs::write( + &p, + helvetica_page(b"0 0 1 rg 400 100 60 60 re f BT /F1 10 Tf 405 120 Td (Stamp) Tj ET"), + ) + .expect("stamp"); + p +} + +/// Whether page `page` of `pdf` is qpdf's overlay wrapper (`q cm Do Q` only). +fn is_wrapper(pdf: &Path, page: u32) -> bool { + let snap = read_verification_snapshot(pdf, 1 << 30).unwrap_or_else(|e| panic!("{e}")); + let ctx = SnapshotContext::new(snap); + let id = ctx.page_id(page).expect("page"); + let content = + page_content(ctx.doc(), id, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)).expect("content"); + let ops = lex_content(&content.joined, &LexLimits::page(), None).expect("lex"); + !ops.is_empty() + && ops.iter().all(|o| { + matches!( + o.operator, + Operator::q | Operator::Q | Operator::cm | Operator::Do + ) + }) +} + +#[test] +fn pb_01_parts_form_after_assembly() { + let Some(h) = honest("pb_01") else { return }; + let other = h.path("other.pdf"); + std::fs::write( + &other, + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Other file) Tj ET"), + ) + .expect("other"); + let out = h.path("final.pdf"); + fakes::assemble(h.qpdf(), &[&h.staged, &other], &out).unwrap_or_else(|e| panic!("{e}")); + assert!(!is_wrapper(&out, 0), "PB-01 parts form"); + let r = h.phase_b(&out, 0); + assert!(r.is_ok(), "PB-01: {:?}", r.err().map(|e| e.details)); + // The proof Phase A handed over: the source page, every expected part with its digest and + // the page's show records after the edit. + let proof = h.proof(); + assert_eq!(proof.source_page_index, 0, "PB-01 proof page"); + let digests: Vec<_> = proof + .expected_parts + .iter() + .map(|p| content_digest(p)) + .collect(); + assert_eq!(proof.part_digests, digests, "PB-01 part digests"); + assert_eq!(proof.records.len(), 2, "PB-01 two show records"); +} + +#[test] +fn pb_02_wrapper_form_after_an_empty_overlay() { + let Some(h) = honest("pb_02") else { return }; + let out = h.path("final.pdf"); + fakes::overlay(h.qpdf(), &h.staged, &out, &blank(&h), None).unwrap_or_else(|e| panic!("{e}")); + assert!( + is_wrapper(&out, 0), + "PB-02 every page becomes the wrapper (P-OV)" + ); + let r = h.phase_b(&out, 0); + assert!(r.is_ok(), "PB-02: {:?}", r.err().map(|e| e.details)); +} + +#[test] +fn pb_03_wrapper_with_a_stamp_on_the_same_page() { + let Some(h) = honest("pb_03") else { return }; + let out = h.path("final.pdf"); + fakes::overlay(h.qpdf(), &h.staged, &out, &stamp(&h), None).unwrap_or_else(|e| panic!("{e}")); + let r = h.phase_b(&out, 0); + assert!( + r.is_ok(), + "PB-03 the stamp is a paint: {:?}", + r.err().map(|e| e.details) + ); +} + +/// `pdf` with every page's boxes set as the Save pipeline does before an overlay +/// (MediaBox = CropBox = TrimBox = the visible box), or restored to `boxes`. +fn set_boxes(h: &Honest, input: &Path, out: &Path, boxes: &[(&str, Option<[f64; 4]>)]) { + let doc = fakes::load(input); + let page = fakes::page_ids(&doc)[0]; + let mut d = json_dict(doc.get_dictionary(page).expect("page")); + if let Some(o) = d.as_object_mut() { + for (key, value) in boxes { + match value { + Some(b) => o.insert(format!("/{key}"), json!(b.to_vec())), + None => o.remove(&format!("/{key}")), + }; + } + } + fakes::set_value(h.qpdf(), input, out, page, d).unwrap_or_else(|e| panic!("{e}")); +} + +#[test] +fn pb_04_wrapper_on_rotated_and_cropped_pages() { + for (angle, text) in [(90, "Rotated page"), (270, "Rotated page")] { + let test = format!("pb_04_{angle}"); + let Some(h) = Honest::new( + &test, + fx::rotated(angle, true), + 0, + &[ed(text, "Rotated pages")], + ) else { + return; + }; + let out = h.path("final.pdf"); + fakes::overlay(h.qpdf(), &h.staged, &out, &blank(&h), None) + .unwrap_or_else(|e| panic!("{e}")); + assert!(is_wrapper(&out, 0), "PB-04 /Rotate {angle} wrapper"); + let r = h.phase_b(&out, 0); + assert!( + r.is_ok(), + "PB-04 /Rotate {angle}: {:?}", + r.err().map(|e| e.details) + ); + } + // Offset CropBox with Trim ⊂ Crop: qpdf alone would centre the trimmed content in the + // MediaBox (it moves); the pipeline remaps the boxes to the visible box first, then restores. + let pdf = helvetica_doc( + b"BT /F1 12 Tf 72 720 Td (Cropped page) Tj ET", + "", + "/CropBox [36 48 576 744] /TrimBox [50 60 560 730]", + ); + let Some(h) = Honest::new("pb_04_crop", pdf, 0, &[ed("Cropped page", "Cropped pages")]) else { + return; + }; + let visible = [36.0, 48.0, 576.0, 744.0]; + let remapped = h.path("remapped.pdf"); + set_boxes( + &h, + &h.staged, + &remapped, + &[ + ("MediaBox", Some(visible)), + ("CropBox", Some(visible)), + ("TrimBox", Some(visible)), + ], + ); + let overlaid = h.path("overlaid.pdf"); + fakes::overlay(h.qpdf(), &remapped, &overlaid, &stamp(&h), None) + .unwrap_or_else(|e| panic!("{e}")); + let out = h.path("final.pdf"); + set_boxes( + &h, + &overlaid, + &out, + &[ + ("MediaBox", Some([0.0, 0.0, 612.0, 792.0])), + ("CropBox", Some(visible)), + ("TrimBox", Some([50.0, 60.0, 560.0, 730.0])), + ], + ); + assert!(is_wrapper(&out, 0), "PB-04 cropped wrapper"); + let r = h.phase_b(&out, 0); + assert!( + r.is_ok(), + "PB-04 offset CropBox: {:?}", + r.err().map(|e| e.details) + ); + // Without the remap qpdf moves the content: the wrapper is not the identity — refused. + let moved = h.path("moved.pdf"); + fakes::overlay(h.qpdf(), &h.staged, &moved, &blank(&h), None).unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_b(&moved, 0), + "SOURCE_EDIT_GATE_FAILED", + "B1", + "not the identity", + "PB-04 unremapped", + ); +} + +#[test] +fn pb_05_wrapper_around_a_fake() { + let Some(h) = honest("pb_05") else { return }; + let c = h + .model + .runs + .iter() + .find(|r| r.text == "Hello world") + .expect("run") + .rect; + let cover = [c[0] - 1.0, c[1] - 1.0, c[2] + 2.0, c[3] + 2.0]; + // (a) the old parts in /Fx0, the new text in /Fx1 (a real `qpdf --overlay`). + let fake = h.path("fake-f2.pdf"); + fakes::overlay_f2( + h.qpdf(), + &h.source, + &fake, + 1, + [0.0, 0.0, 612.0, 792.0], + cover, + "Hello there", + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_b(&fake, 0), + "SOURCE_EDIT_GATE_FAILED", + "B1", + "", + "PB-05 B1", + ); + let r = skipping(&["B1"], || h.phase_b(&fake, 0)); + fails_at( + r, + "SOURCE_EDIT_GATE_FAILED", + "B2", + "original content", + "PB-05 B2", + ); + let r = skipping(&["B1", "B2"], || h.phase_b(&fake, 0)); + fails_at(r, "SOURCE_EDIT_GATE_FAILED", "B3", "", "PB-05 B3"); + // (b) GATE-18's fake (old part + a cover part) wrapped by a later overlay: both in /Fx0. + let f1 = h.path("fake-f1.pdf"); + fakes::cover_and_overlay_f1(h.qpdf(), &h.source, &f1, 0, cover, "F1", "Hello there") + .unwrap_or_else(|e| panic!("{e}")); + let wrapped = h.path("fake-f1-wrapped.pdf"); + fakes::overlay(h.qpdf(), &f1, &wrapped, &blank(&h), None).unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_b(&wrapped, 0), + "SOURCE_EDIT_GATE_FAILED", + "B1", + "", + "PB-05b B1", + ); + let r = skipping(&["B1"], || h.phase_b(&wrapped, 0)); + fails_at( + r, + "SOURCE_EDIT_GATE_FAILED", + "B2", + "original content", + "PB-05b B2", + ); + let r = skipping(&["B1", "B2"], || h.phase_b(&wrapped, 0)); + fails_at(r, "SOURCE_EDIT_GATE_FAILED", "B3", "", "PB-05b B3"); +} + +#[test] +fn pb_06_wrapper_bbox_cuts_the_visible_page() { + let Some(h) = honest("pb_06") else { return }; + let wrapped = h.path("wrapped.pdf"); + fakes::overlay(h.qpdf(), &h.staged, &wrapped, &blank(&h), None) + .unwrap_or_else(|e| panic!("{e}")); + let doc = fakes::load(&wrapped); + let page = fakes::page_ids(&doc)[0]; + let fx0 = doc + .get_dictionary(page) + .and_then(|d| d.get(b"Resources")) + .and_then(Object::as_dict) + .and_then(|r| r.get(b"XObject")) + .and_then(Object::as_dict) + .and_then(|x| x.get(b"Fx0")) + .and_then(Object::as_reference) + .expect("/Fx0"); + let stream = doc + .get_object(fx0) + .and_then(Object::as_stream) + .expect("Fx0 stream"); + let mut dict = json_dict(&stream.dict); + dict.as_object_mut() + .expect("dict") + .insert("/BBox".into(), json!([0, 0, 612, 400])); + let out = h.path("final.pdf"); + fakes::stream_dict(h.qpdf(), &wrapped, &out, fx0, dict).unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_b(&out, 0), + "SOURCE_EDIT_GATE_FAILED", + "B1", + "/BBox", + "PB-06", + ); +} + +#[test] +fn pb_07_parts_form_with_flattened_parts_around() { + let Some(h) = honest("pb_07") else { return }; + let out = h.path("final.pdf"); + fakes::add_parts( + h.qpdf(), + &h.staged, + &out, + 0, + &[b"q 1 0 0 RG 300 300 80 20 re S Q"], + &[ + b"q 0 0 1 rg 300 200 80 20 re f Q", + b"q 0.5 g 300 100 80 20 re f Q", + ], + ) + .unwrap_or_else(|e| panic!("{e}")); + let r = h.phase_b(&out, 0); + assert!(r.is_ok(), "PB-07: {:?}", r.err().map(|e| e.details)); + // The expected parts must be there exactly once, contiguous and in order. + let joined: Vec<&[u8]> = h.proof().expected_parts.iter().map(Vec::as_slice).collect(); + assert_eq!( + qpdf_join(&joined), + h.proof().expected_parts.concat(), + "PB-07 one part" + ); +} + +#[test] +fn pb_08_appended_signed_file_and_a_total_above_the_source_cap() { + let Some(h) = honest("pb_08") else { return }; + let signed = h.path("signed.pdf"); + std::fs::write(&signed, fx::signed()).expect("signed"); + let out = h.path("final.pdf"); + fakes::assemble(h.qpdf(), &[&h.staged, &signed], &out).unwrap_or_else(|e| panic!("{e}")); + let size = std::fs::metadata(&out).map_or(0, |m| m.len()); + set_file_cap_override(Some(size / 2)); + let source_policy = read_snapshot(&out); + let cap = 2 * size + VERIFY_CAP_MARGIN_BYTES; + let r = crate::pdf_engine::text_edit::gate::verify_final_output( + &out, + cap, + &h.engines, + &[(0, h.proof())], + &super::opts(), + ); + set_file_cap_override(None); + assert!( + source_policy.is_err(), + "PB-08 read_snapshot alone refuses the final file" + ); + assert!(r.is_ok(), "PB-08 (D31): {:?}", r.err().map(|e| e.details)); +} + +/// PB-09 (review-T5 H1): the wrapper form is refused when qpdf's join changes what the expected +/// parts mean — a comment at a part end would end (the line it hid would be drawn), or two +/// numbers would no longer be one. The walker refuses such pages outright, so no honest plan +/// has such parts: the final file and the proof are built directly, with a wrapper that holds +/// `qpdf_join(expected_parts)` exactly as qpdf would have written it. B1 must refuse it before +/// B3 compares any record. +#[test] +fn pb_09_wrapper_whose_join_changes_the_edited_content() { + let Some(engines) = crate::pdf_engine::text_edit::testkit::engines_or_skip("pb_09") else { + return; + }; + let dir = crate::pdf_engine::text_edit::testkit::Scratch::new("pb_09"); + let cases: [(&str, [&[u8]; 2]); 2] = [ + ( + "comment", + [ + b"BT /F1 12 Tf 72 700 Td (Hello there) Tj ET\n% note", + b"BT /F1 12 Tf 72 650 Td (Second line) Tj ET", + ], + ), + ( + "number", + [ + b"BT /F1 12 Tf 72 700 Td (Hello there) Tj ET 0 0 1 rg 30", + b"0 20 20 re f BT /F1 12 Tf 72 650 Td (Second line) Tj ET", + ], + ), + ]; + for (id, parts) in cases { + let joined = qpdf_join(&parts); + let mut d = DocBuilder::new(); + let f = d.add(fx::HELVETICA); + let form = d.b.add_stream( + &format!( + "/Type /XObject /Subtype /Form /BBox [0 0 612 792] \ + /Resources << /Font << /F1 {f} 0 R >> >>" + ), + &joined, + ); + d.page(PageSpec::new( + b"q 1 0 0 1 0 0 cm /Fx0 Do Q", + &format!("/XObject << /Fx0 {form} 0 R >> /Font << /F1 {f} 0 R >>"), + )); + let out = dir.write(&format!("{id}.pdf"), &d.build()); + let proof = crate::pdf_engine::text_edit::gate::PageProof { + source_page_index: 0, + expected_parts: parts.iter().map(|p| p.to_vec()).collect(), + part_digests: parts.iter().map(|p| content_digest(p)).collect(), + original_edited_digests: Vec::new(), + original_edited_rolling: Vec::new(), + expected_page_digest: content_digest(&parts.concat()), + records: Vec::new(), + }; + let cap = 4 * std::fs::metadata(&out).map_or(0, |m| m.len()) + VERIFY_CAP_MARGIN_BYTES; + let r = crate::pdf_engine::text_edit::gate::verify_final_output( + &out, + cap, + &engines, + &[(0, &proof)], + &super::opts(), + ); + fails_at(r, "SOURCE_EDIT_GATE_FAILED", "B1", "join changes", id); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate/plan_bugs.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate/plan_bugs.rs new file mode 100644 index 0000000..fb7836a --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate/plan_bugs.rs @@ -0,0 +1,310 @@ +//! GATE-01…08 and 11…13: planner bugs. The tampered replacement is written consistently (the expected +//! parts, the qpdf update and the "before" digest all carry it), so A0–A3 pass and the re-walk +//! (A4) must catch it — with the code the editor shows. + +use super::{ed, fails_at, skipping, Ed, Honest}; +use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::verify::test_seams::skip_grammar; + +/// "Hello" followed on its line by " tail" at 14 pt (a pen-chained follower that is not joined). +fn hello_tail() -> Vec { + helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj /F1 14 Tf ( tail) Tj ET BT /F1 12 Tf 72 650 Td (Other) Tj ET", + ) +} + +fn honest(test: &str, pdf: Vec, edits: &[Ed<'_>]) -> Option { + Honest::new(test, pdf, 0, edits) +} + +#[test] +fn gate_01_compensation_dropped_is_pen_drift() { + let Some(h) = honest("gate_01", hello_tail(), &[ed("Hello", "Help")]) else { + return; + }; + // "Help" (2056) is 222 thousandths shorter than "Hello" (2278): `[<48656C70> -222] TJ`. + let (plan, digest, out) = h.bad_plan("bad.pdf", |s| s.replace(" -222] TJ", "] TJ")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "PEN_DRIFT", + "A4", + "drift", + "GATE-01", + ); +} + +#[test] +fn gate_02_compensation_sign_flipped_or_off_by_one_unit() { + let Some(h) = honest("gate_02", hello_tail(), &[ed("Hello", "Help")]) else { + return; + }; + let (plan, digest, out) = h.bad_plan("flip.pdf", |s| s.replace(" -222] TJ", " 222] TJ")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "PEN_DRIFT", + "A4", + "drift", + "GATE-02 sign", + ); + // One unit = 0.012 pt at 12 pt: above the 0.01 pt tolerance. + let (plan, digest, out) = h.bad_plan("unit.pdf", |s| s.replace(" -222] TJ", " -221] TJ")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "PEN_DRIFT", + "A4", + "drift", + "GATE-02 unit", + ); +} + +#[test] +fn gate_03_tc_not_restored() { + let style = SourceTextStyleIn { + letter_spacing_pt: Some(1.0), + ..Default::default() + }; + let edits = [Ed { + old: "Hello", + new: "Hello", + style, + }]; + let Some(h) = honest("gate_03", hello_tail(), &edits) else { + return; + }; + let (plan, digest, out) = h.bad_plan("bad.pdf", |s| s.replace("] TJ 0 Tc", "] TJ")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "STATE_CHANGED", + "A4", + "field=tc", + "GATE-03", + ); +} + +#[test] +fn gate_04_b15_fill_leaks_onto_a_later_path() { + let pdf = helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET 72 600 50 50 re f"); + let edits = [Ed { + old: "Hello", + new: "Hello", + style: SourceTextStyleIn { + fill: Some("#ff0000".into()), + ..Default::default() + }, + }]; + let Some(h) = honest("gate_04", pdf, &edits) else { + return; + }; + let (plan, digest, out) = h.bad_plan("bad.pdf", |s| s.replace("] TJ 0 g", "] TJ")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "STATE_CHANGED", + "A4", + "paint_changed", + "GATE-04", + ); + // The independent render sees the red square too. + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", "check=render", "GATE-04 A5"); +} + +#[test] +fn gate_05_default_black_restore_and_05b_cross_family() { + // Never set: `0 0 0 rg` is the initial black by effect — passes (mobile edit.test.ts:825). + let red = || SourceTextStyleIn { + fill: Some("#ff0000".into()), + ..Default::default() + }; + let pdf = helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET 72 600 50 50 re f"); + let edits = [Ed { + old: "Hello", + new: "Hello", + style: red(), + }]; + let Some(h) = honest("gate_05", pdf, &edits) else { + return; + }; + let (plan, digest, out) = h.bad_plan("ok.pdf", |s| s.replace("] TJ 0 g", "] TJ 0 0 0 rg")); + let r = h.phase_a_of(&plan, &digest, &out); + assert!( + r.is_ok(), + "GATE-05 default black by effect: {:?}", + r.err().map(|e| e.details) + ); + // Set explicitly as `0 g`, restored as `0 0 0 1 k`: another colour family (D35). + let pdf = helvetica_page(b"0 g BT /F1 12 Tf 72 700 Td (Hello) Tj ET 72 600 50 50 re f"); + let edits = [Ed { + old: "Hello", + new: "Hello", + style: red(), + }]; + let Some(h) = honest("gate_05b", pdf, &edits) else { + return; + }; + let (plan, digest, out) = h.bad_plan("bad.pdf", |s| s.replace("] TJ 0 g", "] TJ 0 0 0 1 k")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "STATE_CHANGED", + "A4", + "field=fill", + "GATE-05b", + ); +} + +#[test] +fn gate_06_tf_restored_with_another_size() { + let pdf = helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj 0 0 1 rg ( tail) Tj ET 0 g BT /F1 12 Tf 72 650 Td (Other) Tj ET", + ); + let edits = [Ed { + old: "Hello", + new: "Hello", + style: SourceTextStyleIn { + size_pt: Some(14.0), + ..Default::default() + }, + }]; + let Some(h) = honest("gate_06", pdf, &edits) else { + return; + }; + let (plan, digest, out) = h.bad_plan("bad.pdf", |s| s.replace("TJ /F1 12 Tf", "TJ /F1 13 Tf")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "STATE_CHANGED", + "A4", + "field=tfs", + "GATE-06", + ); +} + +#[test] +fn gate_07_tz_tr_ts_in_the_replacement() { + let Some(h) = honest("gate_07", hello_tail(), &[ed("Hello", "Help")]) else { + return; + }; + for (op, field) in [("50 Tz", "th"), ("1 Tr", "tr"), ("3 Ts", "ts")] { + let name = format!("bad-{}.pdf", &op[op.len() - 2..]); + let (plan, digest, out) = h.bad_plan(&name, |s| format!("{op} {s}")); + let r = h.phase_a_of(&plan, &digest, &out); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A4", + "forbidden_operator", + &format!("GATE-07 {op}"), + ); + let _g = skip_grammar(); + let r = h.phase_a_of(&plan, &digest, &out); + fails_at( + r, + "STATE_CHANGED", + "A4", + &format!("field={field}"), + &format!("GATE-07 {op} bypassed"), + ); + } +} + +#[test] +fn gate_08_q_or_bt_in_the_replacement() { + let Some(h) = honest("gate_08", hello_tail(), &[ed("Hello", "Help")]) else { + return; + }; + let (plan, digest, out) = h.bad_plan("q.pdf", |s| format!("q {s} Q")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "EDIT_VERIFY_FAILED", + "A4", + "op=q", + "GATE-08 q", + ); + let (plan, digest, out) = h.bad_plan("bt.pdf", |s| format!("ET BT {s}")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "EDIT_VERIFY_FAILED", + "A4", + "forbidden_operator", + "GATE-08 BT", + ); + let _g = skip_grammar(); + let (plan, digest, out) = h.bad_plan("bt2.pdf", |s| format!("BT {s} ET")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "EDIT_VERIFY_FAILED", + "A4", + "page_refused", + "GATE-08 BT bypassed", + ); +} + +#[test] +fn gate_11_non_sibling_font() { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let c = + d.add("<< /Type /Font /Subtype /Type1 /BaseFont /Courier /Encoding /WinAnsiEncoding >>"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F2 12 Tf 72 650 Td (Mono) Tj ET", + &format!("/Font << /F1 {f} 0 R /F2 {c} 0 R >>"), + )); + let Some(h) = honest("gate_11", d.build(), &[ed("Hello", "Help")]) else { + return; + }; + let (plan, digest, out) = h.bad_plan("bad.pdf", |s| format!("/F2 12 Tf {s} /F1 12 Tf")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "EDIT_VERIFY_FAILED", + "A4", + "what=font", + "GATE-11", + ); +} + +#[test] +fn gate_12_b2_effective_size_written_as_tf() { + let edits = [Ed { + old: "Hi", + new: "Hi", + style: SourceTextStyleIn { + size_pt: Some(13.0), + ..Default::default() + }, + }]; + let Some(h) = honest("gate_12", fx::tf1_tm12(), &edits) else { + return; + }; + let (plan, digest, out) = h.bad_plan("bad.pdf", |s| s.replace("/F1 1.0833 Tf", "/F1 13 Tf")); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "EDIT_VERIFY_FAILED", + "A4", + "what=size", + "GATE-12", + ); +} + +#[test] +fn gate_13_b3_code_with_an_empty_glyph() { + // A Word subset of "Hello": "Hello" → "Hollo" is honest; the bug writes "Y" (no glyph in the + // subset, width 0) for the new "o", and expects it too. + let Some(h) = honest("gate_13", fx::subset_without_y(), &[ed("Hello", "Hollo")]) else { + return; + }; + let (plan, digest, out) = h.bad_run("bad.pdf", |run| { + let s = &mut run.splices[0]; + let text = String::from_utf8_lossy(&s.bytes).replacen("<486F", "<4859", 1); + s.bytes = text.into_bytes(); + run.expected.glyphs[1].2.value = 0x59; + run.expected.text = "HYllo".into(); + }); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "EDIT_VERIFY_FAILED", + "A4", + "what=glyph", + "GATE-13", + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate/preview.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate/preview.rs new file mode 100644 index 0000000..939d783 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate/preview.rs @@ -0,0 +1,150 @@ +//! The preview core (§B.17) under the same Phase A as Save: a verified one-page PDF for an honest +//! edit, no PDF for a failed verdict, files written once in the cache directory and every nonce +//! directory removed. (The PREV-01…07 matrix through the command layer is T5's.) + +use super::opts; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::preview::{cache_dir_for, preview_page}; +use crate::pdf_engine::text_edit::reasons::{EditProblemCode, TextWarningCode}; +use crate::pdf_engine::text_edit::rewrite::{SourceTextStyleIn, TextEditIn}; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::producers::{self as fx, helvetica_page}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use std::sync::Arc; + +fn edit_of(m: &crate::pdf_engine::text_edit::runs::PageModel, old: &str, new: &str) -> TextEditIn { + let run = m.runs.iter().find(|r| r.text == old).expect("run"); + TextEditIn { + run_id: run.id.clone(), + original_text: old.into(), + text: new.into(), + style: SourceTextStyleIn::default(), + } +} + +#[test] +fn preview_core_returns_a_verified_page_and_cleans_up() { + let Some(engines) = engines_or_skip("preview_core") else { + return; + }; + let dir = Scratch::new("preview_core"); + let pdf = helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hello world) Tj ET BT /F1 12 Tf 72 650 Td (Second line) Tj ET", + ); + let path = dir.write("doc.pdf", &pdf); + let ctx = SnapshotContext::new(snapshot_from_bytes(&path, pdf, None).expect("snapshot")); + let m = Arc::new(build_page_model(&ctx, 0, None).expect("model")); + let cache = cache_dir_for(dir.dir(), &ctx.snap.fingerprint.to_string()); + for round in 0..2 { + let r = preview_page( + &ctx, + Arc::clone(&m), + &[edit_of(&m, "Hello world", "Hello there")], + &cache, + &engines, + &[], + &opts(), + ) + .unwrap_or_else(|e| panic!("preview: {e} {:?}", e.details)); + assert!( + r.page_problem.is_none() && r.verdicts[0].problem.is_none(), + "round {round}: {:?}", + r.page_problem + ); + assert!(r.warnings.is_empty(), "round {round}: {:?}", r.warnings); + let bytes = r.pdf.expect("a preview page"); + let out = dir.write("preview.pdf", &bytes); + let text = std::process::Command::new(&engines.pdftotext) + .arg(&out) + .arg("-") + .output() + .expect("pdftotext"); + let text = String::from_utf8_lossy(&text.stdout); + assert!( + text.contains("Hello there") && !text.contains("Hello world"), + "{text}" + ); + } + // `source.pdf` and `p1.pdf` were written once; no nonce directory is left behind. + let mut names: Vec = std::fs::read_dir(&cache) + .expect("cache dir") + .filter_map(|e| e.ok().map(|e| e.file_name().to_string_lossy().into_owned())) + .collect(); + names.sort(); + assert_eq!(names, vec!["p1.pdf".to_string(), "source.pdf".to_string()]); +} + +#[test] +fn preview_core_failed_verdict_returns_no_pdf() { + let Some(engines) = engines_or_skip("preview_core_fail") else { + return; + }; + let dir = Scratch::new("preview_core_fail"); + let pdf = fx::subset_without_y(); + let path = dir.write("doc.pdf", &pdf); + let ctx = SnapshotContext::new(snapshot_from_bytes(&path, pdf, None).expect("snapshot")); + let m = Arc::new(build_page_model(&ctx, 0, None).expect("model")); + let cache = cache_dir_for(dir.dir(), "fp"); + let r = preview_page( + &ctx, + Arc::clone(&m), + &[edit_of(&m, "Hello", "Yellow")], + &cache, + &engines, + &[], + &opts(), + ) + .expect("preview"); + assert!(r.pdf.is_none(), "no bytes for a failed verdict"); + assert_eq!( + r.verdicts[0].problem.as_ref().map(|p| p.code), + Some(EditProblemCode::GlyphMissing) + ); + assert!(!cache.exists(), "nothing written before a plan exists"); +} + +#[test] +fn preview_core_reports_the_overlap_warning_of_its_verdict() { + let Some(engines) = engines_or_skip("preview_core_overlap") else { + return; + }; + let dir = Scratch::new("preview_core_overlap"); + // "Hi" and "there" on one baseline, "there" 23 pt from the origin: "Hiya" (24 pt) runs 1 pt + // into it (the §A.8 warning; the line still fits the page). + let pdf = helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hi) Tj ET BT /F1 12 Tf 95 700 Td (there) Tj ET \ + BT /F1 12 Tf 72 650 Td (Next line) Tj ET", + ); + let path = dir.write("doc.pdf", &pdf); + let ctx = SnapshotContext::new(snapshot_from_bytes(&path, pdf, None).expect("snapshot")); + let m = Arc::new(build_page_model(&ctx, 0, None).expect("model")); + let edit = edit_of(&m, "Hi", "Hiya"); + let cache = cache_dir_for(dir.dir(), "overlap"); + let r = preview_page( + &ctx, + Arc::clone(&m), + &[edit.clone()], + &cache, + &engines, + &[], + &opts(), + ) + .unwrap_or_else(|e| panic!("preview: {e} {:?}", e.details)); + assert!(r.pdf.is_some(), "a warning does not block the preview"); + assert_eq!( + r.verdicts[0].warnings, + vec![TextWarningCode::NextTextOverlap], + "the verdict carries the warning" + ); + let w: Vec<(u32, &str, TextWarningCode)> = r + .warnings + .iter() + .map(|w| (w.page_index, w.run_id.as_str(), w.code)) + .collect(); + assert_eq!( + w, + vec![(0, edit.run_id.as_str(), TextWarningCode::NextTextOverlap)], + "the preview lists it for the page and the run" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate/review_fixes.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate/review_fixes.rs new file mode 100644 index 0000000..99ef5fb --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate/review_fixes.rs @@ -0,0 +1,321 @@ +//! Regressions for the T4 review (`review-T4.md`): render masks per glyph (H1), the request as the +//! reference for A4/A5 (M1), the render budget on oversize pages (M3, with round 1's 25 DPI floor, +//! M-1), objects reached through `/DecodeParms` in A2 (M4), the fill restore after a bare `sc` +//! (L1), and the small robustness items (L5). The word-matching cost bound (M2) is in +//! `bounds.rs`; round 1's CID-keyed CFF masks (H-1) are IND-10 in `tests_independent.rs`. + +use super::{ed, fails_at, skipping, Ed, Honest}; +use crate::pdf_engine::text_edit::apply::{updates_for_plan, warning_text}; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::geometry::PageGeometry; +use crate::pdf_engine::text_edit::graph::{graph_digest, graph_matches}; +use crate::pdf_engine::text_edit::limits::RENDER_PIXELS_MAX; +use crate::pdf_engine::text_edit::poppler::{render_dpi, RENDER_DPI_MIN}; +use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; +use crate::pdf_engine::text_edit::testkit::producers::{ + helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::tests_independent::{follower_doc, follower_embedded}; +use crate::pdf_engine::text_edit::tests_plan::{ + ctx, filled, ok_plan, plan_one, replacement, style, +}; +use std::collections::HashMap; + +/// The right edge of the widest mask box of the plan of "Hello" → "Help" on `pdf`. +fn mask_right_edge(pdf: Vec) -> (f64, f64) { + let (_, _, out) = plan_one(pdf, "Hello", "Help", style()); + let boxes = &ok_plan(&out).runs[0].expected.mask_boxes; + assert!(!boxes.is_empty(), "masks planned"); + let left = boxes.iter().map(|b| b[0]).fold(f64::INFINITY, f64::min); + let right = boxes.iter().map(|b| b[2]).fold(f64::NEG_INFINITY, f64::max); + (left, right) +} + +#[test] +fn h1_masks_follow_each_glyph_not_the_font_wide_box() { + // "Hello" at x 72, 12 pt; the old "o" (advance 0.556 em) starts at 92.664 and is the + // rightmost glyph. Fallback boxes reach at most 0.2 em past an upright glyph's advance + // (0.1 em without a /FontBBox), never the 2 em of an Arial-like /FontBBox. + let close = |a: f64, b: f64| (a - b).abs() < 1e-6; + let (left, right) = mask_right_edge(follower_doc(Some("-665 -325 2000 1006"))); + assert!( + close(right, 92.664 + (0.556 + 0.2) * 12.0), + "arial-like: {right}" + ); + assert!( + close(left, 72.0 - 0.2 * 12.0), + "left reach bounded too: {left}" + ); + let (_, right) = mask_right_edge(follower_doc(None)); + assert!( + close(right, 92.664 + (0.556 + 0.1) * 12.0), + "no /FontBBox: {right}" + ); + // Embedded Liberation Sans: the program's own outline box of "o" (~0.51 em wide). + let (_, right) = mask_right_edge(follower_embedded()); + assert!( + right > 92.664 + 0.4 * 12.0 && right < 92.664 + 0.6 * 12.0, + "program glyph box: {right}" + ); + // An italic face without a program or /FontBBox may overhang up to 0.35 em. + let mut d = DocBuilder::new(); + let f = d.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica-Oblique /Encoding /WinAnsiEncoding >>", + ); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + let (_, right) = mask_right_edge(d.build()); + assert!( + close(right, 92.664 + (0.556 + 0.35) * 12.0), + "italic: {right}" + ); + // Stroked text grows by half the line width, times the miter limit for miter joins. + let (_, right) = mask_right_edge(helvetica_page( + b"2 w BT /F1 12 Tf 2 Tr 72 700 Td (Hello) Tj ET", + )); + assert!( + close(right, 92.664 + 0.656 * 12.0 + 10.0), + "miter stroke: {right}" + ); + let (_, right) = mask_right_edge(helvetica_page( + b"2 w 1 j BT /F1 12 Tf 2 Tr 72 700 Td (Hello) Tj ET", + )); + assert!( + close(right, 92.664 + 0.656 * 12.0 + 1.0), + "round stroke: {right}" + ); +} + +#[test] +fn m1_a_drawable_wrong_glyph_with_consistent_expectations_fails() { + // The user asks for "Hollo"; a planner bug writes the drawable "a" and expects "Hallo". + let pdf = helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 12 Tf 72 650 Td (Other line) Tj ET", + ); + let Some(h) = Honest::new("m1_requested", pdf, 0, &[ed("Hello", "Hollo")]) else { + return; + }; + assert_eq!(h.plan.runs[0].expected.requested_text, "Hollo"); + let (plan, digest, out) = h.bad_run("wrong.pdf", |run| { + let s = &mut run.splices[0]; + let text = String::from_utf8_lossy(&s.bytes).replacen("<486F", "<4861", 1); + s.bytes = text.into_bytes(); + run.expected.glyphs[1].2.value = 0x61; + run.expected.text = "Hallo".into(); + }); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "EDIT_VERIFY_FAILED", + "A4", + "what=requested_text", + "M1 A4", + ); + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A5", + "\"Hollo\" not extracted", + "M1 A5", + ); +} + +#[test] +fn m3_render_resolution_keeps_every_page_within_the_pixel_budget() { + let geom = |w: f64, h: f64| PageGeometry { + media: [0.0, 0.0, w, h], + crop: [0.0, 0.0, w, h], + visible: [0.0, 0.0, w, h], + rotate: 0, + user_unit: 1.0, + }; + let pixels = |w: f64, h: f64, dpi: u32| { + (w / 72.0 * f64::from(dpi)).ceil() * (h / 72.0 * f64::from(dpi)).ceil() + }; + assert_eq!(render_dpi(&geom(612.0, 792.0)), Some(96), "letter"); + assert_eq!( + render_dpi(&geom(14_400.0, 14_400.0)), + Some(RENDER_DPI_MIN), + "the largest legal page" + ); + for side in [4_000.0, 7_200.0, 10_000.0, 14_400.0] { + let dpi = render_dpi(&geom(side, side)).expect("a resolution"); + assert!( + pixels(side, side, dpi) <= RENDER_PIXELS_MAX as f64, + "{side} pt at {dpi} DPI is over the budget" + ); + assert!( + pixels(side, side, dpi + 1) > RENDER_PIXELS_MAX as f64, + "{side} pt: {dpi} DPI is the highest that fits" + ); + } + // Review round 1 (M-1): below 25 DPI the pixel pads and threshold hide a moved follower. + for side in [14_500.0, 20_000.0, 30_000.0, 1.0e6] { + assert_eq!(render_dpi(&geom(side, side)), None, "{side} pt square"); + } + assert_eq!(render_dpi(&geom(0.0, 792.0)), None, "an empty media box"); +} + +#[test] +fn m3_an_edit_on_a_page_too_large_to_verify_fails_closed() { + // 30,000 pt square (beyond Annex C's 14,400) would render below 25 DPI: A5 refuses it + // before rendering anything (review round 1, M-1), instead of comparing near-blind pixels. + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page( + PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 12 Tf 72 650 Td (Other line) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + ) + .media(Some("[0 0 30000 30000]")), + ); + let Some(h) = Honest::unverified("m3_oversize", d.build(), 0, &[ed("Hello", "Help")]) else { + return; + }; + fails_at( + h.phase_a(&h.staged), + "EDIT_VERIFY_FAILED", + "A5", + "too large to verify: the page renders below 25 DPI", + "M-1 oversize", + ); +} + +#[test] +fn m3_an_honest_edit_on_the_largest_legal_page_passes_phase_a() { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page( + PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 12 Tf 72 650 Td (Other line) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + ) + .media(Some("[0 0 14400 14400]")), + ); + let Some(h) = Honest::new("m3_legal", d.build(), 0, &[ed("Hello", "Help")]) else { + return; + }; + assert_eq!( + h.report.proofs.len(), + 1, + "Phase A passes at 25 DPI, A5 included" + ); +} + +/// A page with a JBIG2 image whose `/DecodeParms` reference a globals stream. +fn jbig2_doc(globals: &str, parms_extra: &str) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let g = d.add(format!( + "<< /Length {} >>\nstream\n{globals}\nendstream", + globals.len() + 1 + )); + let img = d.add(format!( + "<< /Type /XObject /Subtype /Image /Width 1 /Height 1 /BitsPerComponent 1 \ + /ColorSpace /DeviceGray /Filter /JBIG2Decode \ + /DecodeParms << /JBIG2Globals {g} 0 R {parms_extra} >> /Length 5 >>\nstream\nXXXX\nendstream" + )); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET q 10 0 0 10 300 300 cm /Im0 Do Q", + &format!("/Font << /F1 {f} 0 R >> /XObject << /Im0 {img} 0 R >>"), + )); + d.build() +} + +#[test] +fn m4_objects_behind_decode_parms_are_part_of_the_graph() { + let before = ctx(jbig2_doc("GLOBALS-AAAA", "/Zzz 1")); + let digest = graph_digest(before.doc(), &HashMap::new(), None).expect("digest"); + let check = |pdf: Vec| { + let after = ctx(pdf); + graph_matches(after.doc(), &digest, &mut DecodeBudget::new(1 << 30), None) + }; + assert_eq!( + check(jbig2_doc("GLOBALS-AAAA", "/Zzz 1")), + Ok(()), + "unchanged" + ); + let m = check(jbig2_doc("GLOBALS-BBBB", "/Zzz 1")).expect_err("globals data changed"); + assert_eq!(m.what, "data", "{m:?}"); + assert!(m.path.ends_with("/DecodeParms/JBIG2Globals"), "{m:?}"); + let m = check(jbig2_doc("GLOBALS-AAAA", "/Zzz 2")).expect_err("a DecodeParms key changed"); + assert_eq!(m.what, "data", "{m:?}"); + assert!(m.path.ends_with("/XObject/Im0"), "{m:?}"); +} + +#[test] +fn l1_colour_change_after_a_bare_sc_restores_its_device_space() { + for (name, setup, restore) in [ + ("gray", "0.5 sc", "/DeviceGray cs 0.5 sc"), + ( + "rgb", + "1 0 0 rg 0.2 0.4 0.6 sc", + "/DeviceRGB cs 0.2 0.4 0.6 sc", + ), + ( + "cmyk", + "0 0 0 1 k 0 0 0 0.5 sc", + "/DeviceCMYK cs 0 0 0 0.5 sc", + ), + ] { + let content = format!( + "{setup} BT /F1 12 Tf 72 700 Td (Gray sc) Tj ET 72 600 50 50 re f \ + BT /F1 12 Tf 72 650 Td (Next) Tj ET" + ); + let pdf = helvetica_page(content.as_bytes()); + let (_, _, out) = plan_one(pdf.clone(), "Gray sc", "Gray sc", filled("#c71c1c")); + let bytes = replacement(&out, 0); + assert!( + bytes.ends_with(restore), + "L1 {name}: the restore re-selects the device space: {bytes}" + ); + let edits = [Ed { + old: "Gray sc", + new: "Gray sc", + style: SourceTextStyleIn { + fill: Some("#c71c1c".into()), + ..SourceTextStyleIn::default() + }, + }]; + let Some(h) = Honest::new(&format!("l1_{name}"), pdf, 0, &edits) else { + return; + }; + assert_eq!(h.report.proofs.len(), 1, "L1 {name}: Phase A passes"); + } +} + +#[test] +fn l5_warnings_compare_without_their_file_names() { + let a = "WARNING: /Users/me/my.pdfs/report.pdf: file is damaged"; + let b = "WARNING: /tmp/work/source.pdf: file is damaged"; + assert_eq!(warning_text(a), "file is damaged"); + assert_eq!(warning_text(a), warning_text(b)); + assert_eq!( + warning_text("WARNING: /a/b.pdf/c.pdf (offset 12): xref not found"), + "(offset 12): xref not found" + ); + assert_eq!( + warning_text("WARNING: no file name here"), + "no file name here" + ); +} + +#[test] +fn l5_an_edited_part_off_the_page_is_an_error_not_a_skip() { + let (_, m, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Help", + style(), + ); + let mut plan = ok_plan(&out).clone(); + assert_eq!( + updates_for_plan(&m.content, &plan).map(|u| u.len()).ok(), + Some(1) + ); + plan.edited_parts.push(7); + let e = updates_for_plan(&m.content, &plan).expect_err("part 7 is not on the page"); + assert_eq!(e.code, "EDIT_VERIFY_FAILED"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_gate/writer_bugs.rs b/src-tauri/src/pdf_engine/text_edit/tests_gate/writer_bugs.rs new file mode 100644 index 0000000..08608b0 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_gate/writer_bugs.rs @@ -0,0 +1,277 @@ +//! GATE-09, 10 and 14…17: writer bugs — qpdf's input is updated with other bytes than the plan +//! expects (or in the wrong place). A2 (whole graph) fails first; skipping it, A3 (exact parts) +//! fails; skipping both, the re-walk (A4) fails where the spec lists it. + +use super::{ed, fails_at, opts, plan_of, skipping, Honest}; +use crate::pdf_engine::text_edit::apply::{apply_update, write_update_json, PartUpdate}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::gate::{ + verify_edited_copy, EditedPageInput, PhaseAInput, PopplerRef, +}; +use crate::pdf_engine::text_edit::graph::{graph_digest, DataKey}; +use crate::pdf_engine::text_edit::limits::VERIFY_CAP_MARGIN_BYTES; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::fakes; +use crate::pdf_engine::text_edit::testkit::producers::{ + helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use std::collections::HashMap; + +/// Two lines in one part: "Hello" (edited) and its neighbour "Other". +fn two_lines() -> Vec { + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 12 Tf 72 650 Td (Other) Tj ET") +} + +impl Honest { + /// qpdf's input with part `part` of the edited page replaced by `bytes` (a writer bug). + fn writer(&self, name: &str, part: usize, bytes: &[u8]) -> std::path::PathBuf { + let out = self.path(name); + fakes::replace_streams( + self.qpdf(), + &self.source, + &out, + &[(self.part_id(part), bytes.to_vec())], + ) + .unwrap_or_else(|e| panic!("{name}: {e}")); + out + } + + /// The honestly edited part `i`, as text. + fn expected(&self, i: usize) -> String { + String::from_utf8_lossy(&self.plan.expected_parts[i]).into_owned() + } + + /// A2 fails on `file`; with A2 skipped A3 fails; with both skipped A4 fails with `a4`. + fn a2_a3_a4(&self, file: &std::path::Path, a4: Option<(&str, &str)>, id: &str) { + fails_at( + self.phase_a(file), + "EDIT_VERIFY_FAILED", + "A2", + "what=data", + id, + ); + let r = skipping(&["A2"], || self.phase_a(file)); + fails_at(r, "EDIT_VERIFY_FAILED", "A3", "differ", &format!("{id} A3")); + if let Some((code, what)) = a4 { + let r = skipping(&["A2", "A3"], || self.phase_a(file)); + fails_at(r, code, "A4", what, &format!("{id} A4")); + } + } +} + +#[test] +fn gate_09_an_unedited_tj_deleted() { + let Some(h) = Honest::new("gate_09", two_lines(), 0, &[ed("Hello", "Help")]) else { + return; + }; + let bytes = h.expected(0).replace("(Other) Tj", ""); + let out = h.writer("bad.pdf", 0, bytes.as_bytes()); + h.a2_a3_a4( + &out, + Some(("EDIT_VERIFY_FAILED", "record_unpaired")), + "GATE-09", + ); +} + +#[test] +fn gate_10_wrong_code_written() { + let Some(h) = Honest::new("gate_10", two_lines(), 0, &[ed("Hello", "Help")]) else { + return; + }; + let honest = h.expected(0); + assert!(honest.contains("<48656C70>"), "GATE-10 {honest}"); + let out = h.writer( + "bad.pdf", + 0, + honest.replace("<48656C70>", "<48656C71>").as_bytes(), + ); + h.a2_a3_a4(&out, Some(("EDIT_VERIFY_FAILED", "what=text")), "GATE-10"); +} + +#[test] +fn gate_14_shared_stream_edited_in_place() { + // Page 1 has a stream of its own; pages 2 and 3 share one. A planner whose SHARED_CONTENT + // refusal was bypassed targets the shared stream: the digest keeps that stream's original + // data (it is reached twice), so A2 fails at the page that shares it. + let Some(engines) = crate::pdf_engine::text_edit::testkit::engines_or_skip("gate_14") else { + return; + }; + let body = b"BT /F1 12 Tf 72 700 Td (Shared body) Tj ET"; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Own page) Tj ET", + &res, + )); + let shared = d.b.add_stream("", body); + d.page_raw(&format!("{shared} 0 R"), &PageSpec::new(b"", &res)); + d.page_raw(&format!("{shared} 0 R"), &PageSpec::new(b"", &res)); + let pdf = d.build(); + // The plan is made on a twin page with the same bytes that is not shared. + let twin = SnapshotContext::new( + snapshot_from_bytes(std::path::Path::new("twin.pdf"), helvetica_page(body), None) + .unwrap_or_else(|e| panic!("{e}")), + ); + let (_, plan) = plan_of(&twin, 0, &[ed("Shared body", "Shared bodies")]); + let dir = crate::pdf_engine::text_edit::testkit::Scratch::new("gate_14"); + let source = dir.write("source.pdf", &pdf); + let ctx = SnapshotContext::new( + snapshot_from_bytes(&source, pdf, None).unwrap_or_else(|e| panic!("{e}")), + ); + let model = build_page_model(&ctx, 1, None).unwrap_or_else(|e| panic!("{e}")); + assert_eq!( + model.runs[0].reason, + Some(crate::pdf_engine::text_edit::reasons::TextReason::SharedContent), + "GATE-14 the refusal that was bypassed" + ); + let id = (shared, 0); + let updates = vec![PartUpdate { + object_id: id, + decoded: plan.expected_parts[0].clone(), + }]; + let update = dir.path("update.json"); + write_update_json(&updates, ctx.doc().max_id, &update).unwrap_or_else(|e| panic!("{e}")); + let out = dir.path("bad.pdf"); + apply_update(&engines, &source, &update, &out, &[], &opts()).unwrap_or_else(|e| panic!("{e}")); + let replaced: HashMap<_, _> = updates + .into_iter() + .map(|u| (u.object_id, u.decoded)) + .collect(); + let digest = graph_digest(ctx.doc(), &replaced, None).unwrap_or_else(|m| panic!("{m:?}")); + assert!( + !digest + .entries + .iter() + .any(|e| matches!(e.data, DataKey::Replaced { .. })), + "GATE-14 a shared stream is never Replaced" + ); + let input = PhaseAInput { + before: &digest, + before_page_count: 3, + staged: &out, + staged_cap: 2 * std::fs::metadata(&source).map_or(0, |m| m.len()) + VERIFY_CAP_MARGIN_BYTES, + pages: vec![EditedPageInput { + model: (&model).into(), + plan: &plan, + input_page_index: 1, + input_render: PopplerRef { + pdf: source.clone(), + page_1: 2, + }, + }], + source_benign: &[], + }; + let r = verify_edited_copy(&input, &engines, dir.dir(), &opts()); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A2", + "Contents what=data", + "GATE-14", + ); +} + +#[test] +fn gate_15_untouched_operators_reformatted() { + // A writer that re-serialises the part (lopdf `Content::encode`, never used in production). + let pdf = helvetica_page( + b"BT\n/F1 12 Tf\n72 700 Td (Hello) Tj ET\nBT /F1 12 Tf 72.000 650 Td (Other)Tj ET", + ); + let Some(h) = Honest::new("gate_15", pdf, 0, &[ed("Hello", "Help")]) else { + return; + }; + let decoded = lopdf::content::Content::decode(&h.plan.expected_parts[0]).expect("decode"); + let encoded = decoded.encode().expect("encode"); + assert_ne!( + encoded, h.plan.expected_parts[0], + "GATE-15 the round trip reformats" + ); + let out = h.writer("bad.pdf", 0, &encoded); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A2", + "what=data", + "GATE-15", + ); + let r = skipping(&["A2"], || h.phase_a(&out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A3", "differ", "GATE-15 A3"); +} + +#[test] +fn gate_16_splice_off_by_one_byte_or_on_the_wrong_page() { + let Some(h) = Honest::new("gate_16", two_lines(), 0, &[ed("Hello", "Help")]) else { + return; + }; + // One byte late: the planned replacement lands one byte to the right. + let original = h.model.content.part_bytes(0).to_vec(); + let s = &h.plan.splices[0]; + let mut late = original[..s.local.start + 1].to_vec(); + late.extend_from_slice(&s.bytes); + late.extend_from_slice(&original[(s.local.end + 1).min(original.len())..]); + let out = h.writer("late.pdf", 0, &late); + // The shifted splice breaks the content syntax, which `qpdf --check` (A0) already sees. + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A0", + "qpdf --check", + "GATE-16 off by one", + ); + let r = skipping(&["A0", "A2"], || h.phase_a(&out)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A3", + "differ", + "GATE-16 off by one A3", + ); + // The right bytes on the wrong page. + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + d.page(PageSpec::new(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET", &res)); + d.page(PageSpec::new(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET", &res)); + let Some(h) = Honest::new("gate_16b", d.build(), 0, &[ed("Hello", "Help")]) else { + return; + }; + let page2 = crate::pdf_engine::text_edit::runs::build_page_model(&h.ctx, 1, None) + .unwrap_or_else(|e| panic!("{e}")); + let out = h.path("wrong-page.pdf"); + fakes::replace_streams( + h.qpdf(), + &h.source, + &out, + &[( + page2.content.parts[0].stream_id, + h.plan.expected_parts[0].clone(), + )], + ) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A2", + "what=data", + "GATE-16 wrong page", + ); +} + +#[test] +fn gate_17_neighbour_run_changed() { + let Some(h) = Honest::new("gate_17", two_lines(), 0, &[ed("Hello", "Help")]) else { + return; + }; + let out = h.writer( + "bad.pdf", + 0, + h.expected(0).replace("(Other)", "(Otter)").as_bytes(), + ); + h.a2_a3_a4( + &out, + Some(("EDIT_VERIFY_FAILED", "record_unpaired")), + "GATE-17", + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_guard.rs b/src-tauri/src/pdf_engine/text_edit/tests_guard.rs new file mode 100644 index 0000000..f135e49 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_guard.rs @@ -0,0 +1,418 @@ +//! GUARD-01…03 (SPEC §B.21): the forbidden-API scan (comment- and string-aware), the lexer +//! fuzz, and no environment reads in production code. + +use crate::pdf_engine::text_edit::lexer::{lex_content, scan_tokens, LexLimits, ScanMode}; +use crate::pdf_engine::text_edit::tests_io::lex::samples; +use std::path::{Path, PathBuf}; +use std::time::Instant; + +const FORBIDDEN: &[&str] = &[ + ".unwrap()", + ".expect(", + "panic!", + "unreachable!", + "unimplemented!", + "todo!", + "Content::decode", + "Content::encode", + "string_to_bytes", + "replace_text", + "encode_text", + "decompressed_content", + "get_plain_content", + "get_page_content", + "Document::load", + "load_mem", + ".decompress()", + "IncrementalDocument", + ".save(", + "save_to(", +]; + +/// First line of a `pdf_engine/*.rs` file that opts into the scan (T3's `source_content.rs`). +pub(crate) const SCAN_MARKER: &str = "//! offpdf:forbidden-api-scan"; + +/// The source with comments and the contents of string/char literals blanked (line breaks kept). +pub(crate) fn code_only(src: &str) -> String { + let c: Vec = src.chars().collect(); + let at = |i: usize| c.get(i).copied().unwrap_or('\0'); + let mut out = String::with_capacity(src.len()); + let mut i = 0; + let blank = |ch: char, out: &mut String| out.push(if ch == '\n' { '\n' } else { ' ' }); + while i < c.len() { + let ch = c[i]; + let prev_ident = i > 0 && (at(i - 1).is_alphanumeric() || at(i - 1) == '_'); + if ch == '/' && at(i + 1) == '/' { + while i < c.len() && c[i] != '\n' { + i += 1; + } + } else if ch == '/' && at(i + 1) == '*' { + let mut depth = 0usize; + while i < c.len() { + if c[i] == '/' && at(i + 1) == '*' { + depth += 1; + i += 2; + } else if c[i] == '*' && at(i + 1) == '/' { + depth -= 1; + i += 2; + if depth == 0 { + break; + } + } else { + blank(c[i], &mut out); + i += 1; + } + } + } else if !prev_ident && (ch == 'r' || (ch == 'b' && at(i + 1) == 'r')) && { + let s = if ch == 'b' { i + 2 } else { i + 1 }; + let mut j = s; + while at(j) == '#' { + j += 1; + } + at(j) == '"' + } { + let s = if ch == 'b' { i + 2 } else { i + 1 }; + let hashes = (s..).take_while(|j| at(*j) == '#').count(); + i = s + hashes + 1; + out.push('"'); + while i < c.len() && !(c[i] == '"' && (1..=hashes).all(|k| at(i + k) == '#')) { + blank(c[i], &mut out); + i += 1; + } + out.push('"'); + i += 1 + hashes; + } else if ch == '"' { + out.push('"'); + i += 1; + while i < c.len() && c[i] != '"' { + if c[i] == '\\' { + blank(c[i], &mut out); + i += 1; + } + if i < c.len() { + blank(c[i], &mut out); + i += 1; + } + } + out.push('"'); + i += 1; + } else if ch == '\'' && at(i + 1) == '\\' { + // escaped char literal ('\n', '\'', '\u{..}'): skip the escaped char, then find the quote + let mut j = i + 3; + while j < c.len() && j < i + 12 && c[j] != '\'' { + j += 1; + } + i = j + 1; + out.push_str("' '"); + } else if ch == '\'' && at(i + 2) == '\'' { + i += 3; // 'x' + out.push_str("' '"); + } else { + // everything else, including lifetimes ('a) and labels + out.push(ch); + i += 1; + } + } + out +} + +/// The part of a production file the scan reads: everything before the first line that +/// starts with `#[cfg(test)]` (unindented; an indented statement attribute does not cut). +fn production_part(src: &str) -> String { + src.lines() + .take_while(|l| !l.starts_with("#[cfg(test)]")) + .collect::>() + .join("\n") +} + +fn violations(src: &str, patterns: &[&str]) -> Vec<(usize, String)> { + code_only(&production_part(src)) + .lines() + .enumerate() + .flat_map(|(n, line)| { + patterns + .iter() + .filter(move |p| line.contains(*p)) + .map(move |p| (n + 1, p.to_string())) + }) + .collect() +} + +fn text_edit_dir() -> PathBuf { + Path::new(env!("CARGO_MANIFEST_DIR")).join("src/pdf_engine/text_edit") +} + +/// Every `.rs` under `text_edit/` except `testkit/`, `tests_*` files/directories and `tests.rs` +/// test-module files (`#[cfg(test)] mod tests;`), plus every `pdf_engine/*.rs` whose first line +/// is the scan marker. +fn scanned_files() -> Vec { + let mut out = Vec::new(); + let mut stack = vec![text_edit_dir()]; + while let Some(dir) = stack.pop() { + for entry in std::fs::read_dir(&dir).unwrap() { + let p = entry.unwrap().path(); + let name = p.file_name().unwrap().to_string_lossy().into_owned(); + if name == "testkit" || name == "tests.rs" || name.starts_with("tests_") { + continue; + } + if p.is_dir() { + stack.push(p); + } else if name.ends_with(".rs") { + out.push(p); + } + } + } + for entry in std::fs::read_dir(text_edit_dir().parent().unwrap()).unwrap() { + let p = entry.unwrap().path(); + if p.extension().is_some_and(|e| e == "rs") { + let src = std::fs::read_to_string(&p).unwrap(); + if src.lines().next() == Some(SCAN_MARKER) { + out.push(p); + } + } + } + out.sort(); + out +} + +#[test] +fn guard01_scanner_skips_comments_and_strings() { + let clean = [ + "//! The docs may name Content::decode, .unwrap() and Document::load.\nfn f() {}", + "/// Never call `get_page_content` here.\nfn f() {}", + "fn f() { let s = \"a.unwrap() Document::load .save(\"; }", + "fn f() { let r = r#\"panic!(\"x\") todo!()\"#; let b = br\"load_mem\"; }", + "fn f() { /* outer /* .expect( */ IncrementalDocument */ }", + "fn f() {}\n#[cfg(test)]\nmod tests { fn g() { x.unwrap(); Content::decode(b); } }", + "fn f() { let c = '\"'; let d = '\\''; }", + "fn f() { x.unwrap_or_default(); y.unwrap_or(0); z.expect_err_free(); }", + ]; + for src in clean { + assert_eq!(violations(src, FORBIDDEN), [], "GUARD-01 clean: {src}"); + } + let dirty = [ + ("fn f() { let x = y.unwrap(); }", ".unwrap()"), + ("fn f() { let c = '\"'; x.unwrap(); }", ".unwrap()"), + ( + "fn f<'a>(x: &'a [u8]) { /* .unwrap() */ Content::decode(x); }", + "Content::decode", + ), + ("fn f() { lopdf::Document::load(p); }", "Document::load"), + ( + "fn f() {\n #[cfg(test)]\n y.expect(\"x\");\n}", + ".expect(", + ), + ("fn f() { s.decompress(); }", ".decompress()"), + ("fn f() { doc.save(p); }", ".save("), + ("fn f() { panic!(\"no\"); }", "panic!"), + ]; + for (src, what) in dirty { + let v = violations(src, FORBIDDEN); + assert!( + v.iter().any(|(_, p)| p == what), + "GUARD-01 must flag {what} in {src}: {v:?}" + ); + } +} + +#[test] +fn guard01_no_forbidden_apis() { + let files = scanned_files(); + let names: Vec = files + .iter() + .map(|p| { + p.strip_prefix(text_edit_dir()) + .unwrap_or(p) + .to_string_lossy() + .replace('\\', "/") + }) + .collect(); + for must in [ + "snapshot.rs", + "lexer.rs", + "lexer/inline.rs", + "decode.rs", + "content.rs", + "engines.rs", + "reasons.rs", + "limits.rs", + ] { + assert!( + names.iter().any(|n| n == must), + "GUARD-01 scans {must}: {names:?}" + ); + } + assert!( + !names + .iter() + .any(|n| n.contains("testkit") || n.contains("tests")), + "GUARD-01 skips test code" + ); + let mut found = Vec::new(); + for (p, name) in files.iter().zip(&names) { + let src = std::fs::read_to_string(p).unwrap(); + for (line, pattern) in violations(&src, FORBIDDEN) { + found.push(format!("{name}:{line}: {pattern}")); + } + } + assert!( + found.is_empty(), + "GUARD-01 forbidden APIs in production code:\n{}", + found.join("\n") + ); +} + +struct XorShift(u64); + +impl XorShift { + fn next(&mut self) -> u64 { + let mut x = self.0; + x ^= x << 13; + x ^= x >> 7; + x ^= x << 17; + self.0 = x; + x + } + fn below(&mut self, n: usize) -> usize { + (self.next() % n.max(1) as u64) as usize + } +} + +const INSERTS: &[&[u8]] = &[ + b"(", + b")", + b"<", + b">", + b"<<", + b">>", + b"[", + b"]", + b"BI", + b" BI ", + b"ID ", + b" EI ", + b"q", + b"Q", + b"%", + b"\\", + b"99999999999999999999", + b"1e308", + b"-.", + b"/", + b"BX", + b"EX", + b"\r\n", + b"\x00", + b"\xff", +]; + +fn mutate(rng: &mut XorShift, base: &[u8]) -> Vec { + let mut v = base.to_vec(); + for _ in 0..1 + rng.below(4) { + let len = v.len(); + match rng.below(6) { + 0 if len > 0 => { + let i = rng.below(len); + v[i] ^= 1 << rng.below(8); + } + 1 if len > 0 => v.truncate(rng.below(len)), + 2 => { + let ins = INSERTS[rng.below(INSERTS.len())]; + let at = rng.below(len + 1); + v.splice(at..at, ins.iter().copied()); + } + 3 if len > 1 => { + let a = rng.below(len); + let b = (a + 1 + rng.below(16)).min(len); + v.drain(a..b); + } + 4 if len > 1 => { + let a = rng.below(len); + let b = (a + 1 + rng.below(32)).min(len); + let piece = v[a..b].to_vec(); + let at = rng.below(v.len() + 1); + v.splice(at..at, piece); + } + _ => { + let i = rng.below(len + 1); + v.insert(i, rng.next() as u8); + } + } + } + v +} + +#[test] +fn guard02_lexer_fuzz() { + let seeds = samples(); + let mut rng = XorShift(0x9E37_79B9_7F4A_7C15); + let limits = LexLimits::page(); + let started = Instant::now(); + let mut errors = 0usize; + for case in 0..10_000 { + let base = &seeds[case % seeds.len()]; + let input = mutate(&mut rng, base); + if lex_content(&input, &limits, None).is_err() { + errors += 1; + } + for mode in [ScanMode::Object, ScanMode::CMap, ScanMode::Type1Clear] { + let _ = scan_tokens(&input, mode, 10_000); + } + } + let secs = started.elapsed().as_secs_f64(); + println!("GUARD-02 10,000 cases, {errors} lexed as errors, {secs:.2} s"); + assert!( + errors > 0 && errors < 10_000, + "GUARD-02 the fuzz reaches both outcomes" + ); + assert!(secs < 20.0, "GUARD-02 < 2 s per 1,000 cases ({secs:.2} s)"); +} + +/// GUARD-03 rule: `env::var` only for the PATH lookup in `engines.rs`. +fn env_violations(name: &str, src: &str) -> Vec { + let prod = production_part(src); + let code = code_only(&prod); + code.lines() + .zip(prod.lines()) + .enumerate() + .filter(|(_, (c, _))| c.contains("env::var")) + .filter(|(_, (_, orig))| { + !(name.ends_with("engines.rs") && orig.contains("env::var_os(\"PATH\")")) + }) + .map(|(n, _)| n + 1) + .collect() +} + +#[test] +fn guard03_no_environment_reads_in_production() { + assert_eq!( + env_violations("x.rs", "fn f() { std::env::var(\"OFFPDF_X\"); }"), + [1], + "GUARD-03 self-test" + ); + assert_eq!( + env_violations("x.rs", "fn f() { std::env::var_os(\"PATH\"); }"), + [1], + "PATH only in engines.rs" + ); + assert!(env_violations("engines.rs", "fn f() { std::env::var_os(\"PATH\"); }").is_empty()); + assert!(env_violations("x.rs", "// std::env::var(\"X\") in a comment\nfn f() {}").is_empty()); + assert!(env_violations( + "x.rs", + "fn f() {}\n#[cfg(test)]\nfn t() { std::env::var(\"X\"); }" + ) + .is_empty()); + let mut found = Vec::new(); + for p in scanned_files() { + let src = std::fs::read_to_string(&p).unwrap(); + let name = p.to_string_lossy().into_owned(); + for line in env_violations(&name, &src) { + found.push(format!("{name}:{line}")); + } + } + assert!( + found.is_empty(), + "GUARD-03 environment reads in production code:\n{}", + found.join("\n") + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_independent.rs b/src-tauri/src/pdf_engine/text_edit/tests_independent.rs new file mode 100644 index 0000000..a58984c --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_independent.rs @@ -0,0 +1,728 @@ +//! T4 independent-engine tests (SPEC §E.6, qpdf + Poppler required): IND-01…08 and the B13 +//! regression. Poppler — not our width model — reads and renders the page before and after the +//! edit (Phase A check A5). IND-02/03/04 reach A5 by skipping earlier checks through the +//! `#[cfg(test)]` seams only, and also show which earlier check catches the same tampering. + +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::engines::RunOpts; +use crate::pdf_engine::text_edit::poppler::{ + pdftotext_words, render_dpi, render_page, user_rect_to_pixels, user_rect_to_text_frame, +}; +use crate::pdf_engine::text_edit::reasons::TextWarningCode; +use crate::pdf_engine::text_edit::rewrite::SourceTextStyleIn; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::testkit::fakes::{self, json_dict}; +use crate::pdf_engine::text_edit::testkit::fonts::tounicode_bfchar; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_doc, helvetica_page, word_font, DocBuilder, PageSpec, +}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use crate::pdf_engine::text_edit::tests_gate::{ed, fails_at, skipping, Ed, Honest}; +use lopdf::Object; +use serde_json::json; + +/// IND-01: an honest edit per font class passes every Phase A check, A5 included. +#[test] +fn ind_01_honest_edits_on_every_font_class_pass_words_and_pixels() { + let cases: Vec<(&str, Vec, &str, &str)> = vec![ + ( + "std14", + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Help", + ), + ("std14-times", fx::std14(), "Times line", "Times lines"), + ("truetype-word", fx::word(), "Due", "Due 7"), + ( + "type0-sibling", + fx::word_tr(), + "Sağlık Bakanlığı Raporu", + "Sağlığı Bakanlık Raporu", + ), + ("symbolic-truetype", fx::libre(), "Libre text", "Libre tex"), + ("cff", fx::libre_cff(), "Hello Hello", "Hello Hole"), + ("type0-skia", fx::skia(), "Chrome", "Chrom"), + ("quartz", fx::quartz(), "Quartz", "Quart"), + ("type1", fx::pdftex(), "Hello World", "Hello Wet World"), + ("cid-cff", fx::xetex(), "XeTeX", "TeX"), + ( + "type1c-tracking", + fx::indd(), + "visible layer", + "visible layers", + ), + ("not-embedded", fx::nonemb(), "Not embedded", "Not embed"), + ("mac-roman", fx::mac_roman(), "caf\u{e9}", "cafe"), + ]; + for (name, pdf, old, new) in cases { + let test = format!("ind_01_{name}"); + let Some(h) = Honest::new(&test, pdf, 0, &[ed(old, new)]) else { + return; + }; + assert!( + h.report + .warnings + .iter() + .all(|w| w.code != TextWarningCode::EditNotVisible), + "IND-01 {name}: the edit is visible" + ); + } +} + +/// A page with two lines in a Word subset: the edited one and an unedited follower line. +fn word_two_lines() -> Vec { + let mut d = DocBuilder::new(); + let f1 = word_font(&mut d.b, "ABCDEF+Calibri", "Invoice 20267 Due"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Invoice 2026) Tj ET BT /F1 12 Tf 72 650 Td (Due 2026) Tj ET", + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +/// The output with every `/Widths` entry of page 1's `/F1` scaled by `factor`. +fn widths_scaled(h: &Honest, factor: f64, name: &str) -> std::path::PathBuf { + let doc = fakes::load(&h.staged); + let page = fakes::page_ids(&doc)[0]; + let font = doc + .get_dictionary(page) + .and_then(|d| d.get(b"Resources")) + .and_then(Object::as_dict) + .and_then(|r| r.get(b"Font")) + .and_then(Object::as_dict) + .and_then(|f| f.get(b"F1")) + .and_then(Object::as_reference) + .expect("/F1"); + let mut d = json_dict(doc.get_dictionary(font).expect("font")); + if let Some(serde_json::Value::Array(w)) = d.get_mut("/Widths") { + for v in w.iter_mut() { + *v = json!(v.as_f64().unwrap_or(0.0) * factor); + } + } + let out = h.path(name); + fakes::set_value(h.qpdf(), &h.staged, &out, font, d).unwrap_or_else(|e| panic!("{e}")); + out +} + +#[test] +fn ind_02_widths_scaled_in_the_output_fail_the_words() { + let Some(h) = Honest::new( + "ind_02", + word_two_lines(), + 0, + &[ed("Invoice 2026", "Invoice 2027")], + ) else { + return; + }; + let out = widths_scaled(&h, 1.01, "widths.pdf"); + // Our own re-walk sees the changed font too (A4) … + let r = skipping(&["A2", "A3"], || h.phase_a(&out)); + fails_at(r, "STATE_CHANGED", "A4", "field=font", "IND-02 A4"); + // … and Poppler's word boxes move: the follower line no longer matches (A5). + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(&out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", "check=words", "IND-02"); +} + +#[test] +fn b13_independent_engines_judge_the_output_not_our_width_model() { + // Helvetica with explicit /Widths; the output's widths are 3 % off. With every check that + // shares our model skipped, Poppler alone still refuses the file. + let mut widths = String::new(); + for c in 32..=126u8 { + let w = match c { + b' ' => 278, + b'H' => 722, + b'e' | b'o' | b'p' => 556, + b'l' => 222, + _ => 500, + }; + widths.push_str(&format!("{w} ")); + } + let mut d = DocBuilder::new(); + let f = d.add(format!( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding /FirstChar 32 /LastChar 126 /Widths [{widths}] >>" + )); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 12 Tf 72 650 Td (Hello people) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + let Some(h) = Honest::new("b13", d.build(), 0, &[ed("Hello", "Help")]) else { + return; + }; + let out = widths_scaled(&h, 1.03, "widths.pdf"); + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(&out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", "check=words", "B13"); +} + +#[test] +fn ind_03_old_text_left_extractable_through_actualtext() { + let pdf = helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Invoice 2026) Tj ET BT /F1 12 Tf 300 700 Td (tail) Tj ET \ + BT /F1 12 Tf 72 650 Td (Next line) Tj ET", + ); + let Some(h) = Honest::new("ind_03", pdf, 0, &[ed("Invoice 2026", "Invoice 2027")]) else { + return; + }; + let honest = String::from_utf8_lossy(&h.plan.expected_parts[0]).into_owned(); + let tail = "BT /F1 12 Tf 300 700 Td (tail) Tj ET"; + let bad = honest.replace( + tail, + &format!("/Span <> BDC {tail} EMC"), + ); + let out = h.path("actualtext.pdf"); + fakes::replace_streams( + h.qpdf(), + &h.source, + &out, + &[(h.part_id(0), bad.into_bytes())], + ) + .unwrap_or_else(|e| panic!("{e}")); + let r = skipping(&["A2", "A3"], || h.phase_a(&out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A4", "", "IND-03 A4"); + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(&out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", "still extracted", "IND-03"); +} + +#[test] +fn ind_04_colour_leak_onto_a_path_fails_the_pixels() { + let pdf = helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET 300 300 80 80 re f"); + let edits = [Ed { + old: "Hello", + new: "Hello", + style: SourceTextStyleIn { + fill: Some("#00ff00".into()), + ..Default::default() + }, + }]; + let Some(h) = Honest::new("ind_04", pdf, 0, &edits) else { + return; + }; + let (plan, digest, out) = h.bad_plan("leak.pdf", |s| s.replace("] TJ 0 g", "] TJ")); + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", "check=render", "IND-04"); +} + +#[test] +fn ind_05_edit_under_an_opaque_box_is_a_warning() { + let pdf = helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET 1 g 60 690 200 30 re f"); + let Some(h) = Honest::new("ind_05", pdf, 0, &[ed("Hello", "Help")]) else { + return; + }; + assert!( + h.report + .warnings + .iter() + .any(|w| w.code == TextWarningCode::EditNotVisible), + "IND-05 EDIT_NOT_VISIBLE: {:?}", + h.report.warnings + ); +} + +#[test] +fn ind_06_frames_on_rotated_and_cropped_pages() { + let Some(engines) = engines_or_skip("ind_06") else { + return; + }; + let dir = Scratch::new("ind_06"); + let opts = RunOpts::default(); + let pages: Vec<(String, Vec)> = [0, 90, 180, 270] + .iter() + .map(|r| { + ( + format!("rotate {r}"), + helvetica_doc( + b"BT /F1 20 Tf 100 200 Td (X) Tj ET", + "", + &format!("/Rotate {r}"), + ), + ) + }) + .chain(std::iter::once(( + "offset CropBox".to_string(), + helvetica_doc( + b"BT /F1 20 Tf 100 200 Td (X) Tj ET", + "", + "/CropBox [50 60 500 700]", + ), + ))) + .collect(); + for (name, pdf) in pages { + let path = dir.write("page.pdf", &pdf); + let snap = snapshot_from_bytes(&path, pdf, None).expect("snapshot"); + let ctx = SnapshotContext::new(snap); + let m = build_page_model(&ctx, 0, None).expect("model"); + let g = &m.walk.records[0].glyphs[0]; + let geom = &m.walk.geometry; + // pdftotext: the word box lies in the computed MediaBox frame box. + let frame = user_rect_to_text_frame(g.bbox, geom); + let words = pdftotext_words(&engines, &path, 1, &opts).expect("words"); + let w = words.first().expect("one word"); + let (cx, cy) = ((w.x0 + w.x1) / 2.0, (w.y0 + w.y1) / 2.0); + assert!( + cx > frame[0] - 2.0 + && cx < frame[2] + 2.0 + && cy > frame[1] - 2.0 + && cy < frame[3] + 2.0, + "IND-06 {name}: word {w:?} outside the frame box {frame:?}" + ); + // pdftoppm (default MediaBox render): every dark pixel lies in the computed pixel box. + let dpi = render_dpi(geom).expect("a render resolution"); + let raster = render_page(&engines, &path, 1, dpi, dir.dir(), &opts).expect("render"); + let b = user_rect_to_pixels(g.bbox, geom, dpi); + let (mut inside, mut outside) = (0usize, 0usize); + for y in 0..raster.h { + for x in 0..raster.w { + let i = ((y * raster.w + x) * 3) as usize; + if raster.rgb[i..i + 3].iter().any(|c| *c < 128) { + let (xi, yi) = (i64::from(x), i64::from(y)); + if xi >= b[0] - 2 && xi < b[2] + 2 && yi >= b[1] - 2 && yi < b[3] + 2 { + inside += 1; + } else { + outside += 1; + } + } + } + } + assert!( + inside > 0 && outside == 0, + "IND-06 {name}: {inside} inside, {outside} outside {b:?}" + ); + } +} + +#[test] +fn ind_07_anti_aliasing_next_to_dense_text_stays_within_tolerance() { + let mut content = String::new(); + for i in 0..30 { + let y = 700.0 - 8.5 * f64::from(i); + content.push_str(&format!( + "BT /F1 8 Tf 72 {y} Td (Line {i:02} with dense neighbouring text WWWWWW) Tj ET " + )); + } + let pdf = helvetica_page(content.as_bytes()); + let old = "Line 15 with dense neighbouring text WWWWWW"; + let Some(h) = Honest::new( + "ind_07", + pdf, + 0, + &[ed(old, "Line 15 with denser text MMMM")], + ) else { + return; + }; + assert!(h.report.proofs.len() == 1, "IND-07 passes A5"); +} + +/// LiberationSans-Italic (in the repo, LICENSE_LIBERATION) as a Type0 Identity-H CIDFontType2 +/// whose CIDs are the program's GIDs, for `chars`, with the program's real `/FontBBox`. +fn liberation_italic( + d: &mut DocBuilder, + chars: &str, +) -> (u32, std::collections::HashMap) { + liberation(d, "LiberationSans-Italic", chars) +} + +/// A Liberation Sans face from the repo (`file` without `.ttf`) as a Type0 Identity-H +/// CIDFontType2 whose CIDs are the program's GIDs, for `chars`, with the program's real +/// `/FontBBox` (Liberation Sans is metric-compatible with Arial; its box reaches ~2 em right). +pub(crate) fn liberation( + d: &mut DocBuilder, + file: &str, + chars: &str, +) -> (u32, std::collections::HashMap) { + let path = std::path::Path::new(env!("CARGO_MANIFEST_DIR")) + .join(format!("../public/pdfjs/standard_fonts/{file}.ttf")); + let program = std::fs::read(path).expect("a Liberation Sans program"); + let italic = file.contains("Italic"); + let (flags, angle) = if italic { (96, -12) } else { (32, 0) }; + let face = ttf_parser::Face::parse(&program, 0).expect("parse"); + let upem = f64::from(face.units_per_em()); + let em = |v: i16| (f64::from(v) * 1000.0 / upem).round(); + let bbox = face.global_bounding_box(); + let mut gids = std::collections::HashMap::new(); + let (mut w, mut map) = (String::new(), Vec::new()); + for ch in chars.chars() { + let gid = face.glyph_index(ch).expect("glyph").0; + let adv = f64::from( + face.glyph_hor_advance(ttf_parser::GlyphId(gid)) + .unwrap_or(0), + ) * 1000.0 + / upem; + w.push_str(&format!("{gid} [{adv:.0}] ")); + map.push((u32::from(gid), 2, ch.to_string())); + gids.insert(ch, gid); + } + let r: Vec<(u32, usize, &str)> = map.iter().map(|(c, l, t)| (*c, *l, t.as_str())).collect(); + let program_id = d.b.add_flate("", &program); + let descriptor = d.add(format!( + "<< /Type /FontDescriptor /FontName /ABCDEF+{name} /Flags {flags} \ + /FontBBox [{} {} {} {}] /ItalicAngle {angle} /Ascent 905 /Descent -212 /CapHeight 729 \ + /StemV 80 /FontFile2 {program_id} 0 R >>", + em(bbox.x_min), + em(bbox.y_min), + em(bbox.x_max), + em(bbox.y_max), + name = file, + program_id = program_id, + )); + let cid = d.add(format!( + "<< /Type /Font /Subtype /CIDFontType2 /BaseFont /ABCDEF+{file} \ + /CIDSystemInfo << /Registry (Adobe) /Ordering (Identity) /Supplement 0 >> \ + /FontDescriptor {descriptor} 0 R /W [{w}] /CIDToGIDMap /Identity >>" + )); + let tounicode = d.b.add_flate("", &tounicode_bfchar(&r)); + let font = d.add(format!( + "<< /Type /Font /Subtype /Type0 /BaseFont /ABCDEF+{file} /Encoding /Identity-H \ + /DescendantFonts [{cid} 0 R] /ToUnicode {tounicode} 0 R >>" + )); + (font, gids) +} + +#[test] +fn ind_08_accented_capitals_in_an_italic_face_pass_the_render_check() { + let mut d = DocBuilder::new(); + let (font, gids) = liberation_italic(&mut d, "ABCİĞŞ"); + let hex: String = "ABC".chars().map(|c| format!("{:04X}", gids[&c])).collect(); + d.page(PageSpec::new( + format!("BT /F1 36 Tf 72 600 Td <{hex}> Tj ET BT /F1 36 Tf 72 500 Td <{hex}> Tj ET") + .as_bytes(), + &format!("/Font << /F1 {font} 0 R >>"), + )); + let pdf = d.build(); + let model_text = { + let snap = snapshot_from_bytes(std::path::Path::new("ind08.pdf"), pdf.clone(), None) + .expect("snapshot"); + let m = build_page_model(&SnapshotContext::new(snap), 0, None).expect("model"); + m.runs + .iter() + .map(|r| (r.text.clone(), r.reason)) + .collect::>() + }; + assert!( + model_text.iter().any(|(t, r)| t == "ABC" && r.is_none()), + "IND-08 editable: {model_text:?}" + ); + let Some(h) = Honest::new("ind_08", pdf, 0, &[ed("ABC", "İĞŞ")]) else { + return; + }; + assert!( + h.report.proofs.len() == 1, + "IND-08 G-RENDER passes with glyph masks" + ); +} + +/// "Hello" and a follower "42" drawn right after it on the same line, pen-chained (`Tj` then a +/// `TJ` that starts with a −500 kern), in Helvetica with `/Widths` and, when given, a +/// `/FontBBox` (not embedded: the masks use the fallback box). +pub(crate) fn follower_doc(font_bbox: Option<&str>) -> Vec { + let mut widths = String::new(); + for c in 32..=126u8 { + let w = match c { + b' ' => 278, + b'H' => 722, + b'e' | b'o' | b'p' | b'4' | b'2' => 556, + b'l' => 222, + _ => 500, + }; + widths.push_str(&format!("{w} ")); + } + let mut d = DocBuilder::new(); + let descriptor = font_bbox.map(|bb| { + let id = d.add(format!( + "<< /Type /FontDescriptor /FontName /Helvetica /Flags 32 /FontBBox [{bb}] \ + /ItalicAngle 0 /Ascent 905 /Descent -212 /CapHeight 716 /StemV 80 >>" + )); + format!("/FontDescriptor {id} 0 R") + }); + let f = d.add(format!( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding \ + /FirstChar 32 /LastChar 126 /Widths [{widths}] {} >>", + descriptor.unwrap_or_default() + )); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj [-500 (42)] TJ ET", + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +/// The same line in embedded Liberation Sans (program glyph boxes, a 2 em wide `/FontBBox`). +pub(crate) fn follower_embedded() -> Vec { + let mut d = DocBuilder::new(); + let (font, gids) = liberation(&mut d, "LiberationSans-Regular", "Helop42"); + let hex = |t: &str| -> String { t.chars().map(|c| format!("{:04X}", gids[&c])).collect() }; + d.page(PageSpec::new( + format!( + "BT /F1 12 Tf 72 700 Td <{}> Tj [-500 <{}>] TJ ET", + hex("Hello"), + hex("42") + ) + .as_bytes(), + &format!("/Font << /F1 {font} 0 R >>"), + )); + d.build() +} + +/// IND-09 (review H1): a compensation 40/1000 em too large moves the same-line follower "42" +/// 0.48 pt right. Our re-walk sees it (A4, `PEN_DRIFT`); with A4 skipped — a width-model error +/// both walks would share, B13 — Poppler's pixels alone must refuse it, whatever the font's +/// `/FontBBox`: the masks follow each glyph (program box or advance-bounded fallback), not the +/// font-wide box that reaches 2 em past every glyph. +#[test] +fn ind_09_same_line_follower_moved_by_a_shared_width_error_fails_the_pixels() { + let cases: Vec<(&str, Vec)> = vec![ + ("arial-bbox", follower_doc(Some("-665 -325 2000 1006"))), + ("tight-bbox", follower_doc(Some("-166 -225 1000 931"))), + ("no-bbox", follower_doc(None)), + ("embedded", follower_embedded()), + ]; + for (name, pdf) in cases { + let test = format!("ind_09_{name}"); + let Some(h) = Honest::new(&test, pdf, 0, &[ed("Hello", "Help")]) else { + return; + }; + let honest = String::from_utf8_lossy(&h.plan.runs[0].splices[0].bytes).into_owned(); + assert!( + honest.ends_with(" -222] TJ"), + "IND-09 {name}: the honest compensation keeps the follower: {honest}" + ); + let (plan, digest, out) = h.bad_plan(&format!("{name}.pdf"), |s| { + s.replace(" -222] TJ", " -262] TJ") + }); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "PEN_DRIFT", + "A4", + "drift", + &format!("IND-09 {name} A4"), + ); + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A5", + "check=render", + &format!("IND-09 {name}"), + ); + } +} + +/// The glyphs of the IND-10 font: (char, advance, outline box) in thousandths of an em. "j" reaches +/// left of its origin and below the baseline (negative `x_min` and `y_min`). +const BOX_GLYPHS: [(char, i32, [i32; 4]); 8] = [ + ('H', 700, [50, 0, 650, 700]), + ('e', 550, [40, -10, 510, 520]), + ('l', 250, [60, 0, 190, 720]), + ('o', 600, [40, -10, 560, 520]), + ('p', 600, [50, -200, 560, 520]), + ('j', 300, [-100, -200, 250, 700]), + ('4', 550, [30, 0, 520, 700]), + ('2', 550, [40, 0, 510, 700]), +]; + +/// "Hello" and the same-line follower "42" (as in `follower_doc`) in a Type0 font whose CID-keyed +/// CFF program (CID = GID = position in `BOX_GLYPHS` + 1) draws box glyphs with `units` font +/// units per em, and carries `top` / `fd` as its Top DICT / Font DICT `FontMatrix` operands; as a +/// bare `CIDFontType0C` program or, with `opentype`, the `CFF ` table of an `OTTO` font (1,000 +/// units per em). Poppler (FreeType) combines the two matrices; ttf-parser reads the Top one only. +fn cid_cff_follower(top: Option<&[u8]>, fd: Option<&[u8]>, units: i32, opentype: bool) -> Vec { + use crate::pdf_engine::text_edit::testkit::cff::{t2, CffBuilder}; + use crate::pdf_engine::text_edit::testkit::fonts::{add_type0, Program, Type0Font}; + use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; + let scale = |v: i32| v * units / 1000; + let mut cff = CffBuilder::new("ABCDEF+BoxCID"); + for (i, (_, _, [x0, y0, x1, y1])) in BOX_GLYPHS.iter().enumerate() { + let (x0, y0, x1, y1) = (scale(*x0), scale(*y0), scale(*x1), scale(*y1)); + let mut cs: Vec = [x0, y0].iter().flat_map(|v| t2(*v)).collect(); + cs.push(21); // rmoveto + for v in [x1 - x0, 0, 0, y1 - y0, x0 - x1, 0] { + cs.extend(t2(v)); + } + cs.extend([5, 14]); // rlineto endchar + cff = cff.raw_cid_glyph(i as u16 + 1, cs); + } + if let Some(t) = top { + cff.top_padding = [t, &[12, 7][..]].concat(); + } + let program = cff.build(); + let program = fd.map_or(program.clone(), |m| { + fakes::cff_with_font_dict_matrix(&program, m) + }); + let mut f = Type0Font::new("CIDFontType0", "ABCDEF+BoxCID"); + f.program = if opentype { + let mut t = TtfBuilder::new(); + t.cff = Some(program); + for (i, (_, advance, _)) in BOX_GLYPHS.iter().enumerate() { + t.glyph(&format!("cid{}", i + 1), true, *advance as u16); + } + Program::OpenType(t.build()) + } else { + Program::CidCff(program) + }; + let widths: Vec = BOX_GLYPHS.iter().map(|g| g.1.to_string()).collect(); + f.w = Some(format!("[1 [{}]]", widths.join(" "))); + let map: Vec<(u32, usize, String)> = BOX_GLYPHS + .iter() + .enumerate() + .map(|(i, g)| (i as u32 + 1, 2, g.0.to_string())) + .collect(); + let r: Vec<(u32, usize, &str)> = map.iter().map(|(c, l, t)| (*c, *l, t.as_str())).collect(); + f.tounicode = Some(tounicode_bfchar(&r)); + let hex = |t: &str| -> String { + t.chars() + .map(|c| { + let i = BOX_GLYPHS + .iter() + .position(|g| g.0 == c) + .expect("a box glyph"); + format!("{:04X}", i + 1) + }) + .collect() + }; + let mut d = DocBuilder::new(); + let font = add_type0(&mut d.b, &f); + d.page(PageSpec::new( + format!( + "BT /F1 12 Tf 72 700 Td <{}> Tj [-500 <{}>] TJ ET", + hex("Hello"), + hex("42") + ) + .as_bytes(), + &format!("/Font << /F1 {font} 0 R >>"), + )); + d.build() +} + +/// IND-10 (review round 1, H-1): a CID-keyed CFF program whose Font DICT carries the +/// `FontMatrix` — Top `[1 0 0 1 0 0]` with Font DICT `[0.001 …]`, or the default Top with Font DICT +/// `[0.0005 …]` over 2,000 units per em — is drawn at its true size by Poppler, while the Top +/// matrix alone gives a 1,000× or 2× scale (bare, and inside OpenType). Its glyph boxes are not +/// taken from the program (no clamped 5 em box around the "j"): the honest edit passes A5, and a +/// 0.48 pt shift of the same-line follower fails A5 with A4 skipped. The plain programs (controls) +/// use their own boxes. +#[test] +fn ind_10_cid_cff_font_dict_matrices_do_not_widen_the_masks() { + const IDENTITY: [u8; 6] = [140, 139, 139, 140, 139, 139]; + const MILLI: [u8; 12] = [ + 30, 0x0a, 0x00, 0x1f, 139, 139, 30, 0x0a, 0x00, 0x1f, 139, 139, + ]; + const HALF_MILLI: [u8; 14] = [ + 30, 0x0a, 0x00, 0x05, 0xff, 139, 139, 30, 0x0a, 0x00, 0x05, 0xff, 139, 139, + ]; + let cases: Vec<(&str, Vec, f64)> = vec![ + ("plain", cid_cff_follower(None, None, 1000, false), 0.35), + ( + "top-identity-fd-milli", + cid_cff_follower(Some(&IDENTITY), Some(&MILLI), 1000, false), + 0.6, + ), + ( + "fd-half-milli", + cid_cff_follower(None, Some(&HALF_MILLI), 2000, false), + 0.6, + ), + ( + "opentype-plain", + cid_cff_follower(None, None, 1000, true), + 0.35, + ), + ( + "opentype-fd-half-milli", + cid_cff_follower(None, Some(&HALF_MILLI), 2000, true), + 0.6, + ), + ]; + for (name, pdf, j_width_em) in cases { + let test = format!("ind_10_{name}"); + let Some(h) = Honest::new(&test, pdf, 0, &[ed("Hello", "Hellj")]) else { + return; + }; + // The new "j" (fifth glyph): its program box (0.35 em wide), or the fallback box (advance + // 0.3 em + 0.1 em left + 0.2 em right, 1.3 em tall), never a box scaled by the Top matrix. + let j = h.plan.runs[0].expected.mask_boxes[4]; + let (w, ht) = ((j[2] - j[0]) / 12.0, (j[3] - j[1]) / 12.0); + assert!( + (w - j_width_em).abs() < 1e-6 && ht <= 1.3 + 1e-6, + "IND-10 {name}: mask of \"j\" {w} x {ht} em ({j:?})" + ); + let honest = String::from_utf8_lossy(&h.plan.runs[0].splices[0].bytes).into_owned(); + assert!( + honest.ends_with(" -300] TJ"), + "IND-10 {name}: the honest compensation keeps the follower: {honest}" + ); + let (plan, digest, out) = h.bad_plan(&format!("{name}.pdf"), |s| { + s.replace(" -300] TJ", " -340] TJ") + }); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "PEN_DRIFT", + "A4", + "drift", + &format!("IND-10 {name} A4"), + ); + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A5", + "check=render", + &format!("IND-10 {name}"), + ); + } +} + +/// The live check's invoice row (review-T5 live B2): a description, a quantity and a price on one +/// baseline. +fn table_row() -> Vec { + helvetica_page( + b"BT /F1 11 Tf 58 700 Td (Desk lamp) Tj ET BT /F1 11 Tf 132 700 Td (4) Tj ET \ + BT /F1 11 Tf 156 700 Td (120.00) Tj ET BT /F1 11 Tf 58 650 Td (Next row) Tj ET", + ) +} + +/// IND-11 (review-T5 live B2): a changed line may run into the next table cells — Poppler then +/// orders the cell words between the new words — and A5 passes; but a word of those cells that +/// no longer reads the same at its place (here re-labelled through ActualText, which changes no +/// pixel) fails A5, even when the edit does not overlap it. +#[test] +fn ind_11_neighbours_on_the_edited_line_are_set_aside_and_must_still_read_the_same() { + let Some(h) = Honest::new( + "ind_11", + table_row(), + 0, + &[ed("Desk lamp", "Desk lamp with a long cable")], + ) else { + return; + }; + assert_eq!( + h.report.warnings.len(), + 0, + "IND-11 the overlap is the planner's warning, not A5's" + ); + let Some(h) = Honest::new("ind_11b", table_row(), 0, &[ed("Desk lamp", "Desk lamps")]) else { + return; + }; + let honest = String::from_utf8_lossy(&h.plan.expected_parts[0]).into_owned(); + let cell = "BT /F1 11 Tf 132 700 Td (4) Tj ET"; + assert!(honest.contains(cell), "IND-11 fixture"); + let bad = honest.replace(cell, &format!("/Span <> BDC {cell} EMC")); + let out = h.path("relabelled.pdf"); + fakes::replace_streams( + h.qpdf(), + &h.source, + &out, + &[(h.part_id(0), bad.into_bytes())], + ) + .unwrap_or_else(|e| panic!("{e}")); + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(&out)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A5", + "next to the edited line", + "IND-11", + ); +} + +mod followers; +mod overlap; diff --git a/src-tauri/src/pdf_engine/text_edit/tests_independent/followers.rs b/src-tauri/src/pdf_engine/text_edit/tests_independent/followers.rs new file mode 100644 index 0000000..f4ad2a1 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_independent/followers.rs @@ -0,0 +1,224 @@ +//! review-verify HIGH-A: IND-09c / IND-04c. A narrow follower right after the edited run — a 7 pt +//! superscript footnote marker, a ":" or a "." in another colour — lies within the 2 pt pad of the +//! edited run's old box and within G-RENDER's mask slack. With A4 — our own walk — skipped (an +//! error both walks would share, B13), Poppler alone must refuse every visible change to it: +//! G-TEXT finds it as a neighbour by our model's glyphs (a word Poppler joined across the edit, +//! "Hello:", is split at its outer edge), and G-RENDER checks its pixels outside the edited +//! glyphs' own boxes. The variants are the verifier's (`v04/verify-logs`, +//! `probe-zv_b2_narrow_followers.log`). + +use super::liberation; +use super::overlap::bump_last_kern; +use crate::pdf_engine::text_edit::testkit::fakes; +use crate::pdf_engine::text_edit::testkit::producers::{helvetica_page, DocBuilder, PageSpec}; +use crate::pdf_engine::text_edit::tests_gate::{ed, fails_at, skipping, Honest}; + +/// (label, page content, the follower's show op as the honest content writes it, whether it is +/// pen-chained to "Hello"). +const FOLLOWERS: [(&str, &str, &str, bool); 4] = [ + ( + "superscript", + "BT /F1 12 Tf 72 700 Td (Hello) Tj /F1 7 Tf 5 Ts (1) Tj 0 Ts ET", + "(1) Tj", + true, + ), + ( + "red colon", + "BT /F1 12 Tf 72 700 Td (Hello) Tj 1 0 0 rg (:) Tj 0 g ET", + "(:) Tj", + true, + ), + ( + "separate colon", + "BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT 1 0 0 rg /F1 12 Tf 99.34 700 Td (:) Tj ET 0 g", + "(:) Tj", + false, + ), + ( + "separate period", + "BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT 1 0 0 rg /F1 12 Tf 100.3 700 Td (.) Tj ET 0 g", + "(.) Tj", + false, + ), +]; + +/// A shorter line, and one whose new glyphs run over the follower (an overlap warning). +const EDITS: [&str; 2] = ["Help", "Hello world"]; + +/// Every A5 failure of these tests names the follower. +const NEXT: &str = "next to the edited line"; + +/// The honest edit of `content`, verified (an overlapping edit next to the follower used to be +/// refused: Poppler orders the follower's word between the new ones). +fn honest(test: &str, content: &str, new: &str) -> Option { + let id = format!("{test}_{}", new.len()); + Honest::new( + &id, + helvetica_page(content.as_bytes()), + 0, + &[ed("Hello", new)], + ) +} + +/// IND-09c (writer): the follower deleted, whitened, made invisible, moved or relabelled in the +/// written content, with A2–A4 skipped, fails A5. +#[test] +fn ind_09c_a_narrow_follower_the_writer_changed_fails_a5() { + for (label, content, show, _) in FOLLOWERS { + for new in EDITS { + let Some(h) = honest("ind_09c_writer", content, new) else { + return; + }; + let written = String::from_utf8_lossy(&h.plan.expected_parts[0]).into_owned(); + assert!(written.contains(show), "IND-09c {label}: {written}"); + let glyph = show.trim_end_matches(" Tj"); + let cases = [ + ("deleted", String::new()), + ("white", format!("1 g {show} 0 g")), + ("invisible", format!("3 Tr {show} 0 Tr")), + ("moved", format!("[-250 {glyph}] TJ")), + ("relabelled", "(7) Tj".to_string()), + ]; + for (name, replacement) in cases { + let id = format!("IND-09c {label} {new:?} {name}"); + let out = h.path(&format!("{name}.pdf")); + let bad = written.replacen(show, &replacement, 1).into_bytes(); + fakes::replace_streams(h.qpdf(), &h.source, &out, &[(h.part_id(0), bad)]) + .unwrap_or_else(|e| panic!("{id}: {e}")); + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(&out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", NEXT, &id); + } + } + } +} + +/// IND-09c (planner): a compensation 40, 250 or 600 thousandths of an em too small moves a +/// pen-chained follower 0.48 to 7.2 pt; A4 sees it (`PEN_DRIFT`), and so does A5 alone. +#[test] +fn ind_09c_a_narrow_follower_moved_by_a_shared_width_error_fails_a5() { + for (label, content, _, chained) in FOLLOWERS { + if !chained { + continue; // its own text object: the edited run's kerns cannot move it + } + for new in EDITS { + let Some(h) = honest("ind_09c_planner", content, new) else { + return; + }; + for delta in [40i64, 250, 600] { + let id = format!("IND-09c {label} {new:?} +{delta}"); + let name = format!("kern-{delta}.pdf"); + let (plan, digest, out) = h.bad_plan(&name, |s| bump_last_kern(s, delta)); + let full = h.phase_a_of(&plan, &digest, &out); + fails_at(full, "PEN_DRIFT", "A4", "drift", &format!("{id} A4")); + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", NEXT, &id); + } + } + } +} + +/// IND-04c: the edited show leaks white fill, `3 Tr` or red fill onto the follower. A4 refuses +/// each (`STATE_CHANGED`, or the splice grammar's `forbidden_operator`); with A4 skipped, A5 does +/// too — except a red superscript under the overlapping edit, which keeps its ink under the new +/// space's mask: a visible re-colour there is A4's alone (DEVIATIONS `[fix-last]`). A follower +/// that sets its own fill is not affected by a fill leak, so only `3 Tr` reaches it. +#[test] +fn ind_04c_a_state_leak_onto_a_narrow_follower_fails_a5() { + for (label, content, _, _) in FOLLOWERS { + let leaks: &[(&str, &str, &str)] = if label == "superscript" { + &[ + ("white", " 1 g", "fill"), + ("invisible", " 3 Tr", "op=Tr"), + ("red", " 1 0 0 rg", "fill"), + ] + } else { + &[("invisible", " 3 Tr", "op=Tr")] + }; + for new in EDITS { + let Some(h) = honest("ind_04c", content, new) else { + return; + }; + for (name, leak, what) in leaks { + let id = format!("IND-04c {label} {new:?} {name}"); + let (plan, digest, out) = + h.bad_plan(&format!("{name}.pdf"), |s| format!("{s}{leak}")); + let full = h.phase_a_of(&plan, &digest, &out); + let code = if *name == "invisible" { + "EDIT_VERIFY_FAILED" + } else { + "STATE_CHANGED" + }; + fails_at(full, code, "A4", what, &format!("{id} A4")); + if *name == "red" && new != "Help" { + continue; + } + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", NEXT, &id); + } + } + } +} + +/// IND-09c (honest, program glyph boxes): new glyphs of an embedded font that run over the next +/// cells change no pixel of those cells outside their own boxes grown by `RENDER_OWN_PAD_PX` +/// (Poppler's anti-aliasing reaches one pixel past a glyph's box: the overlap sweeps on +/// LibreOffice files were refused without the pad). +#[test] +fn ind_09c_an_overlap_in_an_embedded_font_passes() { + let mut d = DocBuilder::new(); + let chars = "Desk lampwithongcb4120.QyPr"; + let (font, gids) = liberation(&mut d, "LiberationSans-Regular", chars); + let hex = |t: &str| -> String { t.chars().map(|c| format!("{:04X}", gids[&c])).collect() }; + let cell = |x: u32, y: u32, t: &str| format!("BT /F1 11 Tf {x} {y} Td <{}> Tj ET ", hex(t)); + let content = [ + cell(58, 700, "Desk lamp"), + cell(132, 700, "4"), + cell(156, 700, "120.00"), + cell(58, 660, "Qty"), + cell(80, 660, "Price"), + ] + .concat(); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {font} 0 R >>"), + )); + let pdf = d.build(); + for (i, (old, new)) in [ + ("Desk lamp", "Desk lamp with a long cable"), + ("Qty", "QtyQtyQty"), + ] + .into_iter() + .enumerate() + { + let test = format!("ind_09c_embedded_{i}"); + if Honest::new(&test, pdf.clone(), 0, &[ed(old, new)]).is_none() { + return; + } + } +} + +/// No new false refusal (review-verify HIGH-A): a new letter that covers a "." of another run +/// (Poppler merges them: "phraser."), and a title under a smaller stamp whose glyphs' centres fall +/// inside the title's words, both verify (both were refused by the first version of this fix on +/// the LibreOffice and MuPDF-stamped corpus files). +#[test] +fn ind_09c_a_covered_follower_and_a_stamp_over_the_line_pass() { + let cases = [ + ( + "BT /F1 12 Tf 72 700 Td (red phrase) Tj 1 0 0 rg (.) Tj 0 g ET", + "red phrase", + "red phraser", + ), + ( + "BT /F1 26 Tf 34 700 Td (Quarterly Report) Tj ET BT /F1 9 Tf 72 715 Td (Approved) Tj ET", + "Quarterly Report", + "Quarterly Reports", + ), + ]; + for (i, (content, old, new)) in cases.into_iter().enumerate() { + let pdf = helvetica_page(content.as_bytes()); + if Honest::new(&format!("ind_09c_honest_{i}"), pdf, 0, &[ed(old, new)]).is_none() { + return; + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_independent/overlap.rs b/src-tauri/src/pdf_engine/text_edit/tests_independent/overlap.rs new file mode 100644 index 0000000..6e52279 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_independent/overlap.rs @@ -0,0 +1,164 @@ +//! review-final HIGH-1: IND-04/09/11 under an edit that runs into its neighbours. G-RENDER cannot +//! see a neighbour that lies under the new glyphs' masks, so G-TEXT must find it as the same word +//! at its place (or joined with the new glyphs at its outer edge) and G-RENDER's ink check must +//! find its ink still there. A4 — our own walk — is skipped where the tests model an error our +//! walk would share with the planner (B13). + +use super::{follower_doc, table_row}; +use crate::pdf_engine::text_edit::testkit::fakes; +use crate::pdf_engine::text_edit::testkit::producers::helvetica_page; +use crate::pdf_engine::text_edit::tests_gate::{ed, fails_at, skipping, Honest}; + +/// The last kern of a splice ending `] TJ`, made `delta` thousandths of an em smaller (the +/// follower moves `delta` × size / 1000 pt right). +pub(super) fn bump_last_kern(s: &str, delta: i64) -> String { + let end = s.rfind("] TJ").expect("a TJ splice"); + let head = &s[..end]; + let start = head.rfind(' ').map_or(0, |i| i + 1); + let n: i64 = head[start..].parse().expect("a whole kern"); + format!("{}{}{}", &s[..start], n - delta, &s[end..]) +} + +/// IND-09b: the follower "42" moved by a width error both walks share, with an edit ("Hello" → +/// "Hello world") whose new glyphs cover it — the pixels cannot see it, G-TEXT must. +#[test] +fn ind_09b_follower_moved_under_an_overlapping_edit_fails_the_words() { + for new in ["Hello world", "Hello wonderful world"] { + let test = format!("ind_09b_{}", new.len()); + let pdf = follower_doc(Some("-665 -325 2000 1006")); + let Some(h) = Honest::new(&test, pdf, 0, &[ed("Hello", new)]) else { + return; + }; + for delta in [40i64, 600] { + let name = format!("{}-{delta}.pdf", new.len()); + let (plan, digest, out) = h.bad_plan(&name, |s| bump_last_kern(s, delta)); + let id = format!("IND-09b {new:?} +{delta}"); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "PEN_DRIFT", + "A4", + "drift", + &format!("{id} A4"), + ); + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A5", + "next to the edited line", + &id, + ); + } + } +} + +/// IND-04b: the edited show leaks its fill onto the next cell, which the new glyphs overlap. +/// White ink is lost under the masks (A5's ink check); a visible re-colour keeps its ink and is +/// A4's alone (`STATE_CHANGED`, DEVIATIONS `[fix-final]`). +#[test] +fn ind_04b_fill_leak_onto_an_overlapped_cell() { + let pdf = helvetica_page( + b"BT /F1 11 Tf 58 700 Td (Desk lamp) Tj ET BT /F1 11 Tf 132 700 Td (4) Tj ET", + ); + let Some(h) = Honest::new( + "ind_04b", + pdf, + 0, + &[ed("Desk lamp", "Desk lamp with a long cable")], + ) else { + return; + }; + for (name, leak) in [("white", " 1 g"), ("red", " 1 0 0 rg")] { + let (plan, digest, out) = h.bad_plan(&format!("{name}.pdf"), |s| format!("{s}{leak}")); + let id = format!("IND-04b {name}"); + fails_at( + h.phase_a_of(&plan, &digest, &out), + "STATE_CHANGED", + "A4", + "fill", + &format!("{id} A4"), + ); + if name == "white" { + let r = skipping(&["A4"], || h.phase_a_of(&plan, &digest, &out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A5", "lost its ink", &id); + } + } +} + +/// IND-11b: a writer that moves, deletes or whitens the overlapped cell "4" (A2–A4 skipped, so +/// only Poppler judges) fails A5; the honest overlap passes with its cells at their boxes. +#[test] +fn ind_11b_overlapped_cells_moved_deleted_or_whitened_fail_a5() { + let Some(h) = Honest::new( + "ind_11b_overlap", + table_row(), + 0, + &[ed("Desk lamp", "Desk lamp with a long cable")], + ) else { + return; + }; + let honest = String::from_utf8_lossy(&h.plan.expected_parts[0]).into_owned(); + let cell = "BT /F1 11 Tf 132 700 Td (4) Tj ET"; + assert!(honest.contains(cell), "IND-11b fixture"); + let cases = [ + ( + "moved_1pt", + "BT /F1 11 Tf 133 700 Td (4) Tj ET", + "next to the edited line", + ), + ( + "moved_4pt", + "BT /F1 11 Tf 136 700 Td (4) Tj ET", + "next to the edited line", + ), + ("deleted", "", "next to the edited line"), + ( + "white", + "BT 1 g /F1 11 Tf 132 700 Td (4) Tj ET 0 g", + "lost its ink", + ), + ]; + for (name, replacement, what) in cases { + let out = h.path(&format!("{name}.pdf")); + let bad = honest.replace(cell, replacement).into_bytes(); + fakes::replace_streams(h.qpdf(), &h.source, &out, &[(h.part_id(0), bad)]) + .unwrap_or_else(|e| panic!("{e}")); + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(&out)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A5", + what, + &format!("IND-11b {name}"), + ); + } +} + +/// review-final LOW-3: the cell "4" re-labelled "45" through ActualText changes no pixel and +/// Poppler reads "45" at the cell's box. A2 (the object graph differs from the plan's) refuses it +/// in the full Phase A; A5 alone now refuses it too (a longer word at the same box is not a join). +#[test] +fn ind_11c_a_neighbour_relabelled_with_more_characters_fails_a2() { + let Some(h) = Honest::new("ind_11c", table_row(), 0, &[ed("Desk lamp", "Desk lamps")]) else { + return; + }; + let honest = String::from_utf8_lossy(&h.plan.expected_parts[0]).into_owned(); + let cell = "BT /F1 11 Tf 132 700 Td (4) Tj ET"; + for label in ["45", "14"] { + let out = h.path(&format!("relabel-{label}.pdf")); + let span = format!("/Span <> BDC {cell} EMC"); + let bad = honest.replace(cell, &span).into_bytes(); + fakes::replace_streams(h.qpdf(), &h.source, &out, &[(h.part_id(0), bad)]) + .unwrap_or_else(|e| panic!("{e}")); + let id = format!("IND-11c {label}"); + fails_at(h.phase_a(&out), "EDIT_VERIFY_FAILED", "A2", "Contents", &id); + let r = skipping(&["A2", "A3", "A4"], || h.phase_a(&out)); + fails_at( + r, + "EDIT_VERIFY_FAILED", + "A5", + "next to the edited line", + &format!("{id} A5"), + ); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_io.rs b/src-tauri/src/pdf_engine/text_edit/tests_io.rs new file mode 100644 index 0000000..7bbfde1 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_io.rs @@ -0,0 +1,439 @@ +//! T1 IO tests (SPEC §E.3): LEX (`tests_io/lex.rs`), DEC (here), SNAP (+ the preflight +//! hardening in `tests_io/preflight.rs` and `tests_io/xref.rs`), CON and ENG. +//! Test IDs appear in test names and assertion messages. + +mod con; +mod eng; +pub(crate) mod lex; +mod preflight; +mod snap; +mod xref; + +use crate::pdf_engine::text_edit::decode::{ + ascii85_decode, ascii_hex_decode, decode_stream, inflate_capped, inflate_end, DecodeBudget, + DecodeError, +}; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::testkit::pdf::{zlib, zlib_zero_bomb}; +use crate::pdf_engine::text_edit::testkit::thread_peak; +use lopdf::{Dictionary, Object, Stream}; + +fn stream(filter: Option, parms: Option, content: Vec) -> Stream { + let mut d = Dictionary::new(); + if let Some(f) = filter { + d.set("Filter", f); + } + if let Some(p) = parms { + d.set("DecodeParms", p); + } + Stream::new(d, content) +} + +fn name(n: &str) -> Object { + Object::Name(n.as_bytes().to_vec()) +} + +fn decode(s: &Stream) -> Result, DecodeError> { + decode_stream( + s, + limits::STREAM_MAX_DECODED, + &mut DecodeBudget::new(usize::MAX), + ) +} + +fn text(n: usize) -> Vec { + let line = b"BT /F1 12 Tf (Hello World) Tj ET\n"; + (0..n).map(|i| line[i % line.len()]).collect() +} + +#[test] +fn dec01_flate_round_trip() { + let plain = text(200_000); + assert_eq!( + inflate_capped(&zlib(&plain), 1 << 20).unwrap(), + plain, + "DEC-01" + ); + assert_eq!( + decode(&stream(Some(name("FlateDecode")), None, zlib(&plain))).unwrap(), + plain, + "DEC-01 via stream" + ); + assert_eq!( + decode(&stream(Some(name("Fl")), None, zlib(b"q Q"))).unwrap(), + b"q Q", + "DEC-01 abbreviation" + ); + let z = zlib(b"abc"); + assert_eq!( + inflate_end(&z, 100).unwrap(), + z.len() - 4, + "DEC-01 consumed input excludes the Adler-32" + ); + assert_eq!( + inflate_capped(&zlib(&plain), plain.len()).unwrap().len(), + plain.len(), + "cap is inclusive" + ); + assert_eq!( + inflate_capped(&zlib(&plain), plain.len() - 1), + Err(DecodeError::TooLarge) + ); +} + +#[test] +fn dec02_truncated_flate_is_corrupt() { + let z = zlib(&text(100_000)); + let cut = &z[..z.len() * 55 / 100]; + assert_eq!( + inflate_capped(cut, 1 << 20), + Err(DecodeError::Corrupt("truncated flate")), + "DEC-02" + ); + assert_eq!( + inflate_capped(&z[..1], 100), + Err(DecodeError::Corrupt("truncated flate")) + ); +} + +#[test] +fn dec03_bomb_is_too_large_with_bounded_peak() { + let bomb = zlib_zero_bomb(1024); // ≈1 MiB in, 1 GiB out + assert!(bomb.len() < 2 << 20); + let cap = 1 << 20; + let (result, peak) = thread_peak(|| inflate_capped(&bomb, cap)); + assert_eq!(result, Err(DecodeError::TooLarge), "DEC-03"); + assert!( + peak <= cap + (64 << 10), + "DEC-03 peak {peak} ≤ cap + 64 KiB" + ); + let big_cap = 16 << 20; + let (result, peak) = thread_peak(|| inflate_capped(&bomb, big_cap)); + assert_eq!(result, Err(DecodeError::TooLarge)); + assert!(peak <= big_cap + (64 << 10), "DEC-03 peak {peak}"); + let z = zlib(&text(4 << 20)); + let (ok, peak) = thread_peak(|| inflate_capped(&z, 8 << 20)); + let n = ok.unwrap().len(); + assert!( + peak <= n + (256 << 10), + "DEC-03 a successful decode holds the output once: {peak}" + ); +} + +#[test] +fn dec04_bad_adler_after_stream_end_is_ok() { + let mut z = zlib(b"BT (x) Tj ET"); + let n = z.len(); + z[n - 1] ^= 0xFF; + z.extend_from_slice(b"\r\n trailing garbage"); + assert_eq!(inflate_capped(&z, 100).unwrap(), b"BT (x) Tj ET", "DEC-04"); +} + +#[test] +fn dec05_flate_over_plain_bytes_is_corrupt() { + assert_eq!( + inflate_capped(b"BT (x) Tj ET", 100), + Err(DecodeError::Corrupt("bad zlib header")), + "DEC-05" + ); + assert_eq!( + inflate_capped(&[0x78, 0x9C, 0xFF, 0xFF, 0xFF], 100), + Err(DecodeError::Corrupt("flate data error")) + ); + assert_eq!( + inflate_capped(&[0x78, 0xBB, 0, 0, 0, 0], 100), + Err(DecodeError::Corrupt("bad zlib header")), + "FDICT" + ); +} + +#[test] +fn dec06_empty_input_is_empty_output() { + assert_eq!(inflate_capped(b"", 0).unwrap(), b"", "DEC-06"); + assert_eq!( + decode(&stream(Some(name("FlateDecode")), None, Vec::new())).unwrap(), + b"" + ); + assert_eq!( + decode(&stream(None, None, b"q Q".to_vec())).unwrap(), + b"q Q", + "no filter" + ); + assert_eq!( + decode(&stream( + Some(Object::Array(vec![])), + Some(Object::Null), + b"q".to_vec() + )) + .unwrap(), + b"q" + ); +} + +#[test] +fn dec07_filter_chains() { + let plain = b"BT /F1 12 Tf (chain) Tj ET".to_vec(); + let hex: Vec = zlib(&plain) + .iter() + .flat_map(|b| format!("{b:02X} ").into_bytes()) + .chain(*b">") + .collect(); + let chain = stream( + Some(Object::Array(vec![name("AHx"), name("Fl")])), + None, + hex.clone(), + ); + assert_eq!(decode(&chain).unwrap(), plain, "DEC-07 [/AHx /Fl]"); + let long = stream(Some(Object::Array(vec![name("AHx"); 5])), None, hex); + assert!( + matches!(decode(&long), Err(DecodeError::UnsupportedFilter(_))), + "DEC-07 more than 4 filters" + ); + let mut budget = DecodeBudget::new(10); + let s = stream(None, None, text(11)); + assert_eq!( + decode_stream(&s, 100, &mut budget), + Err(DecodeError::TooLarge), + "DEC-07 budget is a cap" + ); + assert_eq!(budget.remaining(), 10); + assert!(budget.take(4).is_ok()); + assert_eq!(budget.take(7), Err(DecodeError::TooLarge)); + assert_eq!(budget.remaining(), 6); +} + +#[test] +fn dec08_ascii_decoder_edges() { + assert_eq!( + ascii_hex_decode(b"48 65\n6C6c6F>", 100).unwrap(), + b"Hello", + "DEC-08 AHx" + ); + assert_eq!( + ascii_hex_decode(b"414", 100).unwrap(), + [0x41, 0x40], + "DEC-08 odd digit, missing EOD" + ); + assert_eq!( + ascii_hex_decode(b"41 >junk", 100).unwrap(), + b"A", + "DEC-08 data after EOD ignored" + ); + assert_eq!( + ascii_hex_decode(b"4X", 100), + Err(DecodeError::Corrupt("bad hex digit")) + ); + assert_eq!(ascii_hex_decode(b"414243", 2), Err(DecodeError::TooLarge)); + assert_eq!( + ascii85_decode(b"87cURD]i,\"Ebo80~>", 100).unwrap(), + b"Hello World!", + "DEC-08 A85" + ); + assert_eq!( + ascii85_decode(b"<~87cURD]i,\"Ebo80~>", 100).unwrap(), + b"Hello World!", + "DEC-08 <~ prefix" + ); + assert_eq!( + ascii85_decode(b"z 8\n7cU~>", 100).unwrap(), + [0, 0, 0, 0, b'H', b'e', b'l'], + "DEC-08 z + partial" + ); + assert_eq!( + ascii85_decode(b"87cURD]i~", 100), + Err(DecodeError::Corrupt("bad ASCII85 end")) + ); + assert_eq!( + ascii85_decode(b"8~>", 100), + Err(DecodeError::Corrupt("bad ASCII85 final group")) + ); + assert_eq!( + ascii85_decode(b"s8W-\"~>", 100), + Err(DecodeError::Corrupt("ASCII85 group overflow")) + ); + assert_eq!( + ascii85_decode(b"87czU~>", 100), + Err(DecodeError::Corrupt("bad ASCII85 character")), + "z mid-group" + ); + assert_eq!(ascii85_decode(b"zz", 7), Err(DecodeError::TooLarge)); +} + +#[test] +fn dec09_unsupported_filters_and_parameters() { + let unsupported = |s: Stream| matches!(decode(&s), Err(DecodeError::UnsupportedFilter(_))); + for f in [ + "LZWDecode", + "RunLengthDecode", + "RL", + "DCTDecode", + "JPXDecode", + "JBIG2Decode", + "CCITTFaxDecode", + "Crypt", + "Bogus", + ] { + assert!( + unsupported(stream(Some(name(f)), None, b"x".to_vec())), + "DEC-09 {f}" + ); + } + let mut parms = Dictionary::new(); + parms.set("Predictor", 12); + parms.set("Columns", 5); + assert!( + unsupported(stream( + Some(name("FlateDecode")), + Some(Object::Dictionary(parms.clone())), + zlib(b"x") + )), + "DEC-09 predictor" + ); + let arr = Object::Array(vec![Object::Null, Object::Dictionary(parms)]); + assert!(unsupported(stream( + Some(Object::Array(vec![name("AHx"), name("Fl")])), + Some(arr), + b"x".to_vec() + ))); + let mut one = Dictionary::new(); + one.set("Predictor", 1); + assert!(decode(&stream( + Some(name("FlateDecode")), + Some(Object::Dictionary(one)), + zlib(b"ok") + )) + .is_ok()); + let mut ext = stream(None, None, b"x".to_vec()); + ext.dict.set( + "F", + Object::String(b"file.bin".to_vec(), lopdf::StringFormat::Literal), + ); + assert!(unsupported(ext), "DEC-09 /F external file"); + assert!( + unsupported(stream(Some(Object::Integer(3)), None, b"x".to_vec())), + "malformed /Filter" + ); + assert_eq!( + DecodeError::UnsupportedFilter("x".into()).page_reason(), + TextReason::UnsupportedFilter + ); + assert_eq!( + DecodeError::Corrupt("x").page_reason(), + TextReason::MalformedContent + ); + assert_eq!( + DecodeError::TooLarge.page_reason(), + TextReason::PageTooComplex + ); +} + +#[test] +fn limits_are_consistent() { + use limits::*; + assert!( + STREAM_MAX_DECODED <= PAGE_CONTENT_MAX_DECODED + && PAGE_CONTENT_MAX_DECODED <= PAGE_DECODE_BUDGET + ); + assert!( + OBJSTM_MAX_DECODED <= OBJSTM_TOTAL_DECODED + && XREF_STREAM_MAX_DECODED as u64 <= FILE_CAP_BYTES + ); + assert!(MAX_OBJECT_NESTING == GRAPH_DIRECT_DEPTH_MAX && TOKEN_NESTING_MAX < MAX_OBJECT_NESTING); + assert!( + Q_DEPTH_MAX == 64 + && MARKED_DEPTH_MAX == 64 + && FORM_DEPTH_MAX == 8 + && PAGE_TREE_DEPTH_MAX == 64 + ); + assert!(EDITS_PER_PAGE_MAX < EDITS_PER_SAVE_MAX && EDIT_TEXT_CHARS_MAX == 1_000); + assert!(RUN_MEMBERS_MAX <= RUNS_PER_PAGE_MAX && RUNS_PER_PAGE_MAX <= GLYPHS_PER_PAGE_MAX); + assert!( + FORM_PAINTS_PER_PAGE_MAX > 0 + && FONTS_PER_PAGE_MAX > 0 + && PAGE_PARTS_MAX == 256 + && PAGE_OPS_MAX == 250_000 + ); + assert!( + FONT_PROGRAM_MAX_DECODED <= STREAM_MAX_DECODED + && TOUNICODE_MAX_DECODED < FONT_PROGRAM_MAX_DECODED + ); + assert!( + CMAP_MAPPINGS_MAX == CIDTOGID_MAX_BYTES + && W_ENTRIES_MAX == ARRAY_ITEMS_MAX + && STRUCT_CHAIN_MAX == 16 + ); + assert!(NUMBER_TREE_NODES_MAX == STRUCT_ORDER_NODES_MAX && XY_CUT_DEPTH_MAX == 32); + assert!(STRING_BYTES_MAX < INLINE_IMAGE_MAX_BYTES && LITERAL_PAREN_NESTING_MAX == 100); + // Preflight budgets (review-T1 fix pass): a predictor row fits the widest legal xref row, + // a /Length chain stays far below the ~200 objects that overflowed a debug-build rayon + // stack in the probe, and the object scans read at least every byte once. + assert!( + XREF_PREDICTOR_ROW_MAX >= 3 * XREF_STREAM_FIELD_WIDTH_MAX as usize + && LENGTH_REF_CHAIN_MAX * 4 <= 200 + && LENGTH_REF_STREAMS_MAX <= MAX_OBJECTS + && OBJECT_SCAN_FACTOR >= 1 + && INLINE_HEURISTIC_CANDIDATES_MAX > 0 + ); + assert!( + DRIFT_TOLERANCE_PT == 0.01 + && JOIN_BASELINE_TOL_PT == DRIFT_TOLERANCE_PT + && STATE_EPSILON < COLOR_EPSILON + ); + assert!( + AXIS_EPSILON_REL == 1e-4 && SHEAR_MAX == 0.5 && JOIN_GAP_EM == 0.3 && SYNTH_SPACE_EM == 0.2 + ); + assert!(DEFAULT_KERN_SPACE == -250.0 && PER_GLYPH_MIN_RUNS == 12 && PER_GLYPH_SHARE == 0.8); + assert!( + DUPLICATE_OVERLAP_SHARE == 0.5 + && CLIP_CONTAIN_TOL_PT == OVERLAP_WARN_TOL_PT + && NEGLIGIBLE_KERN > 0.0 + ); + assert!( + NUMBER_DECIMALS == 4 && NUMBER_ABS_MAX == 1e9 && F32_DRIFT_GUARD_PT < DRIFT_TOLERANCE_PT + ); + assert!( + STYLE_EPSILON == 0.001 + && SIZE_MIN_PT < SIZE_MAX_PT + && LETTER_SPACING_MIN_PT < LETTER_SPACING_MAX_PT + ); + assert!(PDFTOTEXT_OUTPUT_MAX <= QPDF_JSON_MAX_BYTES && WORD_BBOX_TOL_PT < EDIT_BAND_PAD_PT); + assert!( + RENDER_DPI_MAX == 96 + && RENDER_PIXELS_MAX > 0 + && RENDER_CHANNEL_TOL == 24 + && RENDER_OUTSIDE_PIXELS_MAX == 8 + ); + assert!( + RENDER_MASK_PAD_PX == 4 && GATE_DECODED_TOTAL > FILE_CAP_BYTES && PREVIEW_PDF_MAX_BYTES > 0 + ); + assert!(UPDATE_JSON_MAX_BYTES > STREAM_MAX_DECODED && SUBPROCESS_TIMEOUT_SECS == 120); + assert!( + CACHE_SNAPSHOTS_MAX == 2 + && CACHE_SNAPSHOT_BYTES_MAX < FILE_CAP_BYTES + && CACHE_PAGE_MODELS_MAX == 32 + ); + assert!( + VERIFY_CAP_MARGIN_BYTES == 256 << 20 + && CHECK_MEMO_MAX == 16 + && HEAD_TAIL_HASH_BYTES == 64 << 10 + ); + assert!(XY_CUT_ROW_GAP_EM < XY_CUT_COL_GAP_EM && COLUMN_GAP_EM == 1.0); + assert!( + WRAPPER_MATRIX_EPSILON < WRAPPER_TRANSLATION_TOL_PT + && MAX_OBJECTS == 2_000_000 + && MAX_XREF_CHAIN == 64 + ); + assert!( + XREF_TAIL_SEARCH_BYTES == 2048 + && PREFLIGHT_DICT_TOKENS_MAX > 0 + && XREF_STREAM_FIELD_WIDTH_MAX == 8 + ); + assert!(TOOL_STDERR_MAX > 0 && TOOL_POLL_MS == 20 && LEX_CANCEL_EVERY_OPS == 4_096); + assert!(INLINE_HEURISTIC_WINDOW == 64 && INFLATE_STEP_BYTES == 64 << 10); + assert_eq!(file_cap(), FILE_CAP_BYTES); + set_file_cap_override(Some(10)); + assert_eq!(file_cap(), 10); + set_file_cap_override(None); + assert_eq!(file_cap(), FILE_CAP_BYTES); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_io/con.rs b/src-tauri/src/pdf_engine/text_edit/tests_io/con.rs new file mode 100644 index 0000000..08362b6 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_io/con.rs @@ -0,0 +1,448 @@ +//! CON-01…12: page content parts, the joined buffer, qpdf's join rule and ownership. + +use crate::pdf_engine::text_edit::content::{ + page_content, part_exclusive, qpdf_join, KidsCounts, PageContent, RefCounts, +}; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget}; +use crate::pdf_engine::text_edit::engines::{run_tool, RunOpts}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::snapshot::{read_snapshot, snapshot_from_bytes}; +use crate::pdf_engine::text_edit::testkit::pdf::{Doc, PdfBuilder}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use crate::pdf_engine::validate_output::content_digest; +use lopdf::{dictionary, Document, Object, ObjectId, Stream}; +use std::ffi::OsString; +use std::path::Path; + +fn doc(bytes: &[u8]) -> Document { + snapshot_from_bytes(Path::new("c.pdf"), bytes.to_vec(), None) + .map(|s| s.doc) + .unwrap_or_else(|e| panic!("{e}")) +} + +fn content(d: &Document, page: u32) -> Result { + page_content(d, (page, 0), &mut DecodeBudget::new(usize::MAX)) +} + +fn exclusive(d: &Document, page: u32, part: usize) -> bool { + let c = content(d, page).unwrap(); + part_exclusive(d, &RefCounts::of(d), &KidsCounts::of(d).unwrap(), &c, part) +} + +#[test] +fn con01_parts_joined_with_newline_and_located() { + let d = Doc::new(&[&[b"q 1 0 0 1 0 0 cm", b"Q"]]); + let pdf = doc(&d.build()); + let c = content(&pdf, d.page_ids[0]).unwrap(); + assert_eq!(c.joined, b"q 1 0 0 1 0 0 cm\nQ", "CON-01"); + assert_eq!((c.parts[1].start, c.parts[1].len), (17, 1)); + assert_eq!(c.part_bytes(1), b"Q"); + assert_eq!(c.part_bytes(9), b"", "unknown part is empty, never a panic"); + assert_eq!( + c.locate(&(2..16)), + Some((0, 2..16)), + "CON-01 locate in part 0" + ); + assert_eq!( + c.locate(&(17..18)), + Some((1, 0..1)), + "CON-01 locate in part 1" + ); + assert_eq!( + c.concat_digest(), + content_digest(b"q 1 0 0 1 0 0 cmQ"), + "CON-01 #34 semantics (no separator)" + ); + assert_eq!(c.parts[0].digest, content_digest(b"q 1 0 0 1 0 0 cm")); + assert_eq!(c.parts[0].stream_id, (d.content_ids[0][0], 0)); + assert_eq!(c.contents_array, None, "direct array"); +} + +#[test] +fn con02_span_over_a_separator_is_not_located() { + let d = Doc::new(&[&[b"BT (a) Tj", b"ET"]]); + let c = content(&doc(&d.build()), d.page_ids[0]).unwrap(); + assert_eq!(c.locate(&(3..12)), None, "CON-02 straddle"); + assert_eq!(c.locate(&(9..10)), None, "CON-02 the separator itself"); + assert_eq!(c.locate(&(0..100)), None); +} + +#[test] +fn con03_contents_shapes() { + let mut d = Document::with_version("1.7"); + let direct = d.add_object(dictionary! { "Type" => "Page", "Contents" => Object::Stream(Stream::new(dictionary! {}, b"q Q".to_vec())) }); + assert_eq!( + content(&d, direct.0).unwrap_err(), + TextReason::MalformedContent, + "CON-03 direct stream" + ); + let number = d.add_object(dictionary! { "Type" => "Page", "Contents" => 5 }); + assert_eq!( + content(&d, number.0).unwrap_err(), + TextReason::MalformedContent + ); + let dict_id = d.add_object(dictionary! { "A" => 1 }); + let to_dict = d.add_object(dictionary! { "Type" => "Page", "Contents" => dict_id }); + assert_eq!( + content(&d, to_dict.0).unwrap_err(), + TextReason::MalformedContent, + "reference to a non-stream" + ); + let missing = d.add_object(dictionary! { "Type" => "Page" }); + assert!( + content(&d, missing.0).unwrap().parts.is_empty(), + "missing /Contents = empty page" + ); + let null = d.add_object(dictionary! { "Type" => "Page", "Contents" => Object::Null }); + assert!( + content(&d, null.0).unwrap().joined.is_empty(), + "null /Contents = empty page" + ); + assert_eq!( + content(&d, 999).unwrap_err(), + TextReason::MalformedContent, + "no such page" + ); + let s = d.add_object(Stream::new(dictionary! {}, b"q Q".to_vec())); + let arr = d.add_object(Object::Array(vec![ + Object::Reference(s), + Object::Reference(s), + ])); + let via_array = d.add_object(dictionary! { "Type" => "Page", "Contents" => arr }); + let c = content(&d, via_array.0).unwrap(); + assert_eq!( + (c.joined.as_slice(), c.contents_array), + (&b"q Q\nq Q"[..], Some(arr)), + "reference to an array object" + ); +} + +#[test] +fn con04_with_replaced_parts() { + let d = Doc::new(&[&[b"q", b"BT (a) Tj ET", b"Q"]]); + let c = content(&doc(&d.build()), d.page_ids[0]).unwrap(); + let r = c.with_replaced_parts(&[(1, b"BT (abc) Tj ET".to_vec()), (7, b"ignored".to_vec())]); + assert_eq!(r.joined, b"q\nBT (abc) Tj ET\nQ", "CON-04"); + assert_eq!((r.parts[2].start, r.parts[1].len), (17, 14)); + assert_eq!(r.parts[1].digest, content_digest(b"BT (abc) Tj ET")); + assert_eq!(r.parts[0].digest, c.parts[0].digest); + assert_eq!(r.locate(&(5..9)), Some((1, 3..7))); + assert_eq!(r.page_id, c.page_id); +} + +/// Two pages; page 1 is `[shared own]`; page 2's `/Contents` is built from (shared, own2). +fn two_pages(page2_contents: impl Fn(u32, u32) -> String) -> (Document, u32, u32) { + let mut b = PdfBuilder::new(); + let shared = b.add_stream("", b"0 0 m 10 10 l S"); + let own = b.add_stream("", b"BT (own) Tj ET"); + let own2 = b.add_stream("", b"BT (own 2) Tj ET"); + let (cat, pages) = (b.alloc(), b.alloc()); + let p1 = b.add(format!( + "<< /Type /Page /Parent {pages} 0 R /Contents [{shared} 0 R {own} 0 R] >>" + )); + let p2 = b.add(format!( + "<< /Type /Page /Parent {pages} 0 R /Contents {} >>", + page2_contents(shared, own2) + )); + b.set( + pages, + format!("<< /Type /Pages /Kids [{p1} 0 R {p2} 0 R] /Count 2 /MediaBox [0 0 612 792] >>"), + ); + b.set(cat, format!("<< /Type /Catalog /Pages {pages} 0 R >>")); + (doc(&b.build(&format!("/Root {cat} 0 R"))), p1, p2) +} + +#[test] +fn con05_shared_stream_and_shared_array() { + let (d, p1, p2) = two_pages(|shared, _| format!("{shared} 0 R")); + assert!(!exclusive(&d, p1, 0), "CON-05 shared part"); + assert!(exclusive(&d, p1, 1), "CON-05 own part"); + assert!(!exclusive(&d, p2, 0)); + let mut b = PdfBuilder::new(); + let s = b.add_stream("", b"q Q"); + let arr = b.add(format!("[{s} 0 R]")); + let (cat, pages) = (b.alloc(), b.alloc()); + let p1 = b.add(format!( + "<< /Type /Page /Parent {pages} 0 R /Contents {arr} 0 R >>" + )); + let p2 = b.add(format!( + "<< /Type /Page /Parent {pages} 0 R /Contents {arr} 0 R >>" + )); + b.set( + pages, + format!("<< /Type /Pages /Kids [{p1} 0 R {p2} 0 R] /Count 2 /MediaBox [0 0 612 792] >>"), + ); + b.set(cat, format!("<< /Type /Catalog /Pages {pages} 0 R >>")); + let d = doc(&b.build(&format!("/Root {cat} 0 R"))); + assert!( + !exclusive(&d, p1, 0), + "CON-05 shared /Contents array object" + ); + assert_eq!(RefCounts::of(&d).count((arr, 0)), 2); +} + +#[test] +fn con06_page_listed_twice_and_letterhead_part() { + let d = Doc::new(&[&[b"q Q"]]); + let mut b = d.b.clone(); + let p = d.page_ids[0]; + b.set( + d.pages, + format!("<< /Type /Pages /Kids [{p} 0 R {p} 0 R] /Count 2 /MediaBox [0 0 612 792] >>"), + ); + let pdf = doc(&b.build(&d.trailer())); + assert_eq!(KidsCounts::of(&pdf).unwrap().count((p, 0)), 2); + assert!(!exclusive(&pdf, p, 0), "CON-06 page listed twice in /Kids"); + let (d, p1, p2) = two_pages(|shared, own2| format!("[{shared} 0 R {own2} 0 R]")); + assert!(!exclusive(&d, p1, 0), "CON-06 letterhead part is shared"); + assert!(exclusive(&d, p1, 1), "CON-06 body part is exclusive"); + assert!(!exclusive(&d, p2, 0)); + assert!(exclusive(&d, p2, 1)); +} + +/// FX-WORD-like tagged/bookmarked page: every legitimate back-reference points at the page. +fn referenced_page() -> (Document, u32) { + let mut d = Doc::new(&[&[b"/P <> BDC BT /F1 12 Tf (x) Tj ET EMC"]]); + let p = d.page_ids[0]; + let (cat, pages) = (d.catalog, d.pages); + let st_root = d.b.alloc(); + let elems: Vec = (0..3) + .map(|i| { + d.b.add(format!( + "<< /Type /StructElem /S /P /P {st_root} 0 R /Pg {p} 0 R /K {i} >>" + )) + }) + .collect(); + let kids = elems + .iter() + .map(|e| format!("{e} 0 R")) + .collect::>() + .join(" "); + d.b.set(st_root, format!("<< /Type /StructTreeRoot /K [{kids}] >>")); + let link = d.b.add(format!( + "<< /Type /Annot /Subtype /Link /Rect [0 0 10 10] /P {p} 0 R /Dest [{p} 0 R /Fit] >>" + )); + let outlines = d.b.alloc(); + let item = d.b.add(format!( + "<< /Title (One) /Parent {outlines} 0 R /Dest [{p} 0 R /XYZ 0 792 0] >>" + )); + d.b.set( + outlines, + format!("<< /Type /Outlines /First {item} 0 R /Last {item} 0 R /Count 1 >>"), + ); + let body = String::from_utf8(d.b.body(p).unwrap().to_vec()).unwrap(); + d.b.set( + p, + body.replace(" >>", &format!(" /Annots [{link} 0 R] /StructParents 0 >>")), + ); + d.b.set( + cat, + format!( + "<< /Type /Catalog /Pages {pages} 0 R /StructTreeRoot {st_root} 0 R /Outlines {outlines} 0 R \ + /Names << /Dests << /Names [(here) [{p} 0 R /Fit]] >> >> /OpenAction [{p} 0 R /Fit] /MarkInfo << /Marked true >> >>" + ), + ); + (doc(&d.build()), p) +} + +#[test] +fn con07_tagged_page_is_exclusive() { + let (d, p) = referenced_page(); + assert!( + RefCounts::of(&d).count((p, 0)) >= 5, + "the page is referenced from the struct tree and the annotation" + ); + assert!(exclusive(&d, p, 0), "CON-07"); +} + +#[test] +fn con08_destinations_and_open_action_do_not_share() { + let (d, p) = referenced_page(); + assert_eq!(KidsCounts::of(&d).unwrap().count((p, 0)), 1); + assert!( + RefCounts::of(&d).count((p, 0)) >= 9, + "outline, link, named destination, /OpenAction" + ); + assert!(exclusive(&d, p, 0), "CON-08"); +} + +#[test] +fn con09_same_page_under_two_pages_nodes() { + let mut b = PdfBuilder::new(); + let s = b.add_stream("", b"q Q"); + let (cat, root, n1, n2) = (b.alloc(), b.alloc(), b.alloc(), b.alloc()); + let p = b.add(format!( + "<< /Type /Page /Parent {n1} 0 R /Contents {s} 0 R >>" + )); + b.set( + root, + format!("<< /Type /Pages /Kids [{n1} 0 R {n2} 0 R] /Count 2 /MediaBox [0 0 612 792] >>"), + ); + b.set( + n1, + format!("<< /Type /Pages /Parent {root} 0 R /Kids [{p} 0 R] /Count 1 >>"), + ); + b.set( + n2, + format!("<< /Type /Pages /Parent {root} 0 R /Kids [{p} 0 R] /Count 1 >>"), + ); + b.set(cat, format!("<< /Type /Catalog /Pages {root} 0 R >>")); + let d = doc(&b.build(&format!("/Root {cat} 0 R"))); + assert_eq!(KidsCounts::of(&d).unwrap().count((p, 0)), 2); + assert!(!exclusive(&d, p, 0), "CON-09"); + // a node that lists its ancestor is a cycle + let mut b2 = b.clone(); + b2.set( + n2, + format!("<< /Type /Pages /Parent {root} 0 R /Kids [{root} 0 R] /Count 1 >>"), + ); + let cyclic = lopdf::Document::load_mem(&b2.build(&format!("/Root {cat} 0 R"))).unwrap(); + assert_eq!( + KidsCounts::of(&cyclic).err(), + Some(TextReason::MalformedContent), + "cycle" + ); +} + +const JOIN_CASES: &[(&[&[u8]], &[u8])] = &[ + (&[b"0 g", b"1 g"], b"0 g\n1 g"), + (&[b"0 g\n", b"1 g"], b"0 g\n1 g"), + (&[b"0 g\r", b"1 g"], b"0 g\r\n1 g"), + (&[b"0 g", b""], b"0 g\n"), + (&[b"", b"0 g"], b"\n0 g"), + (&[b"", b""], b"\n"), + (&[b"", b"", b""], b"\n"), + (&[b"", b"", b"0 g"], b"\n0 g"), + (&[b"", b"", b"", b"A"], b"\n\nA"), + (&[b"0 g", b"", b"1 g"], b"0 g\n1 g"), + (&[b"0 g", b"", b"", b"1 g"], b"0 g\n\n1 g"), + (&[b"0 g", b"", b"", b"", b"1 g"], b"0 g\n\n1 g"), + (&[b"0 g", b"", b"", b""], b"0 g\n\n"), + (&[b"0 g", b"", b"1 g\n", b"0.5 g"], b"0 g\n1 g\n0.5 g"), + (&[b"A", b"", b"B", b"", b"C"], b"A\nB\nC"), + (&[b"A", b"", b"", b"B", b"", b"", b"C"], b"A\n\nB\n\nC"), + (&[b"q"], b"q"), + (&[], b""), +]; + +#[test] +fn con10_qpdf_join_rule() { + for (parts, want) in JOIN_CASES { + assert_eq!(qpdf_join(parts), *want, "CON-10 {parts:?}"); + } + // Pin the rule on the installed qpdf: its overlay wrapper Form holds the joined parts. + let Some(engines) = engines_or_skip("con10_qpdf_join_rule") else { + return; + }; + let s = Scratch::new("con10"); + let blank = s.write("blank.pdf", &Doc::new(&[&[b""]]).build()); + for (i, (parts, want)) in JOIN_CASES + .iter() + .enumerate() + .filter(|(_, (p, _))| !p.is_empty()) + { + let src = s.write(&format!("in{i}.pdf"), &Doc::new(&[parts]).build()); + let out = s.path(&format!("out{i}.pdf")); + let args = [ + src.into_os_string(), + OsString::from("--overlay"), + blank.clone().into_os_string(), + OsString::from("--"), + out.clone().into_os_string(), + ]; + let r = run_tool(&engines.qpdf, &args, false, &RunOpts::default()).unwrap(); + assert!(r.code == 0 || r.code == 3, "{}", r.stderr); + let snap = read_snapshot(&out).unwrap_or_else(|e| panic!("{e}")); + let page = snap.doc.get_dictionary(snap.pages[0]).unwrap(); + let xobjects = match page.get(b"Resources").unwrap() { + Object::Reference(id) => snap.doc.get_dictionary(*id).unwrap(), + Object::Dictionary(d) => d, + other => panic!("{other:?}"), + }; + let fx0: ObjectId = xobjects + .get(b"XObject") + .and_then(|x| match x { + Object::Dictionary(d) => d.get(b"Fx0").and_then(Object::as_reference), + Object::Reference(id) => snap + .doc + .get_dictionary(*id) + .and_then(|d| d.get(b"Fx0")) + .and_then(Object::as_reference), + _ => panic!("xobjects"), + }) + .unwrap(); + let Object::Stream(form) = snap.doc.objects.get(&fx0).unwrap() else { + panic!("Fx0 is a stream") + }; + let data = decode_stream(form, 1 << 20, &mut DecodeBudget::new(usize::MAX)).unwrap(); + assert_eq!( + data, + *want, + "CON-10 installed qpdf joins {parts:?} as {:?}", + String::from_utf8_lossy(&data) + ); + } +} + +#[test] +fn con11_part_dict_with_extra_key_is_unsupported_filter() { + let mut d = Doc::new(&[&[b"q Q"]]); + let id = d.content_ids[0][0]; + d.b.set_stream(id, "/Type /XObject", b"q Q"); + assert_eq!( + content(&doc(&d.build()), d.page_ids[0]).unwrap_err(), + TextReason::UnsupportedFilter, + "CON-11" + ); + let mut ok = Doc::new(&[&[b"q Q"]]); + ok.b.set( + id, + PdfBuilder::stream_body( + "/Filter /FlateDecode /DecodeParms << /Predictor 1 >> /DL 3", + &crate::pdf_engine::text_edit::testkit::pdf::zlib(b"q Q"), + ), + ); + assert_eq!( + content(&doc(&ok.build()), ok.page_ids[0]).unwrap().joined, + b"q Q", + "allowed keys" + ); + let mut lzw = Doc::new(&[&[b"q Q"]]); + lzw.b.set_stream( + id, + "/Filter /LZWDecode", + b"\x80\x0b\x60\x50\x22\x0c\x0c\x85\x01", + ); + assert_eq!( + content(&doc(&lzw.build()), lzw.page_ids[0]).unwrap_err(), + TextReason::UnsupportedFilter + ); +} + +#[test] +fn con12_dangling_contents_element_is_malformed() { + let mut d = Doc::new(&[&[b"q", b"Q"]]); + let p = d.page_ids[0]; + let body = String::from_utf8(d.b.body(p).unwrap().to_vec()).unwrap(); + let first = d.content_ids[0][0]; + d.b.set(p, body.replace(&format!("{first} 0 R "), "999 0 R ")); + assert_eq!( + content(&doc(&d.build()), p).unwrap_err(), + TextReason::MalformedContent, + "CON-12" + ); + let parts: Vec<&[u8]> = vec![b"q"; 257]; + let many = Doc::new(&[&parts]); + assert_eq!( + content(&doc(&many.build()), many.page_ids[0]).unwrap_err(), + TextReason::PageTooComplex, + "> 256 parts" + ); + let big = Doc::new(&[&[b"q Q q Q"]]); + let pdf = doc(&big.build()); + assert_eq!( + page_content(&pdf, (big.page_ids[0], 0), &mut DecodeBudget::new(3)).unwrap_err(), + TextReason::PageTooComplex, + "budget" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_io/eng.rs b/src-tauri/src/pdf_engine/text_edit/tests_io/eng.rs new file mode 100644 index 0000000..0aaf143 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_io/eng.rs @@ -0,0 +1,497 @@ +//! ENG-01…09: tool resolution, the subprocess runner, `qpdf --check` classification and memo, +//! the background check and qpdf's page map. + +use crate::error::AppError; +use crate::models::JobRegistry; +use crate::pdf_engine::text_edit::engines::{ + classify_check, parse_qpdf_pages, parse_qpdf_version, qpdf_check, qpdf_check_memo, + qpdf_page_map, qpdf_version_ok, run_tool, BenignRule, Engines, PendingCheck, QpdfPage, RunOpts, + SourceCheck, BENIGN_CHECK_RULES, CHECK_RUNS, +}; +use crate::pdf_engine::text_edit::snapshot::{check_page_map, read_snapshot, Fingerprint}; +use crate::pdf_engine::text_edit::testkit::pdf::{simple_pdf, Doc}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use std::ffi::OsString; +use std::path::{Path, PathBuf}; +use std::sync::atomic::{AtomicBool, Ordering}; +use std::time::{Duration, Instant}; + +fn os(args: &[&str]) -> Vec { + args.iter().map(OsString::from).collect() +} + +fn runs() -> usize { + CHECK_RUNS.with(|c| c.get()) +} + +#[cfg(unix)] +fn script(s: &Scratch, name: &str, body: &str) -> PathBuf { + use std::os::unix::fs::PermissionsExt; + let p = s.write(name, format!("#!/bin/sh\n{body}\n").as_bytes()); + std::fs::set_permissions(&p, std::fs::Permissions::from_mode(0o755)).unwrap(); + p +} + +#[test] +fn eng01_qpdf_version() { + assert_eq!( + parse_qpdf_version("qpdf version 12.3.2\nRun qpdf --copyright"), + Some((12, "12.3.2".into())), + "ENG-01" + ); + assert_eq!( + parse_qpdf_version("qpdf version 11.9.0\n"), + Some((11, "11.9.0".into())) + ); + assert_eq!( + parse_qpdf_version("qpdf version 10.6.3"), + Some((10, "10.6.3".into())) + ); + assert_eq!(parse_qpdf_version("something else"), None); + assert_eq!(parse_qpdf_version(""), None); + if let Some(engines) = engines_or_skip("eng01_qpdf_version") { + assert!( + qpdf_version_ok(&engines.qpdf).is_ok(), + "ENG-01 installed qpdf ≥ 11" + ); + } + let err = qpdf_version_ok(Path::new("/nonexistent/offpdf/qpdf")).unwrap_err(); + assert_eq!(err.code, "ENGINE_MISSING", "missing qpdf"); + #[cfg(unix)] + { + let s = Scratch::new("eng01"); + let old = script(&s, "qpdf-old", "echo 'qpdf version 10.6.3'"); + let e = qpdf_version_ok(&old).unwrap_err(); + assert_eq!(e.code, "ENGINE_MISSING", "ENG-01 qpdf < 11"); + assert!(e.details.unwrap().contains("qpdf 11 or newer is required")); + let odd = script(&s, "qpdf-odd", "echo 'not qpdf'"); + assert_eq!(qpdf_version_ok(&odd).unwrap_err().code, "ENGINE_FAILED"); + assert_eq!( + Engines::for_export(&old, None).unwrap_err().code, + "ENGINE_MISSING" + ); + } +} + +#[cfg(unix)] +#[test] +fn eng02_timeout_kills_a_sleeping_child() { + let opts = RunOpts { + timeout: Duration::from_millis(300), + ..RunOpts::default() + }; + let t = Instant::now(); + let e = run_tool(Path::new("/bin/sleep"), &os(&["5"]), false, &opts).unwrap_err(); + assert_eq!(e.code, "ENGINE_FAILED", "ENG-02"); + assert!(e.details.unwrap().contains("took too long")); + assert!( + t.elapsed() < Duration::from_secs(3), + "ENG-02 killed, not waited for" + ); +} + +#[cfg(unix)] +#[test] +fn eng03_cancel_kills() { + let cancel = AtomicBool::new(false); + let opts = RunOpts { + cancel: Some(&cancel), + ..RunOpts::default() + }; + let t = Instant::now(); + let e = std::thread::scope(|scope| { + scope.spawn(|| { + std::thread::sleep(Duration::from_millis(200)); + cancel.store(true, Ordering::SeqCst); + }); + run_tool(Path::new("/bin/sleep"), &os(&["5"]), false, &opts).unwrap_err() + }); + assert_eq!(e.code, "CANCELLED", "ENG-03"); + assert!(t.elapsed() < Duration::from_secs(3)); + let registry = JobRegistry::default(); + let handle = registry.register("eng03"); + handle.cancel(); + let opts = RunOpts { + handle: Some(&handle), + ..RunOpts::default() + }; + assert_eq!( + run_tool(Path::new("/bin/sleep"), &os(&["5"]), false, &opts) + .unwrap_err() + .code, + "CANCELLED", + "job handle" + ); +} + +#[cfg(unix)] +#[test] +fn eng04_stdout_cap_and_missing_tools() { + let capped = RunOpts { + stdout_cap: 1_000, + ..RunOpts::default() + }; + let e = run_tool( + Path::new("/usr/bin/head"), + &os(&["-c", "200000", "/dev/zero"]), + false, + &capped, + ) + .unwrap_err(); + assert_eq!(e.code, "EDIT_VERIFY_FAILED", "ENG-04"); + assert!(e.details.unwrap().contains("tool output too large")); + let ok = run_tool( + Path::new("/usr/bin/head"), + &os(&["-c", "1000", "/dev/zero"]), + false, + &capped, + ) + .unwrap(); + assert_eq!( + (ok.code, ok.stdout.len()), + (0, 1000), + "ENG-04 exactly the cap is fine" + ); + let failing = run_tool( + Path::new("/bin/sh"), + &os(&["-c", "echo oops >&2; exit 2"]), + false, + &RunOpts::default(), + ) + .unwrap(); + assert_eq!( + (failing.code, failing.stderr.trim()), + (2, "oops"), + "non-zero exit is returned, not an error" + ); + let missing = Path::new("/nonexistent/offpdf/tool"); + assert_eq!( + run_tool(missing, &[], false, &RunOpts::default()) + .unwrap_err() + .code, + "ENGINE_MISSING" + ); + assert_eq!( + run_tool(missing, &[], true, &RunOpts::default()) + .unwrap_err() + .code, + "VERIFIER_MISSING" + ); + let e = Engines::for_export(&crate::pdf_engine::qpdf::resolve_qpdf_standalone(), None); + if let Err(e) = e { + assert!(matches!( + e.code.as_str(), + "ENGINE_MISSING" | "VERIFIER_MISSING" + )); + } +} + +#[test] +fn eng05_check_classification_allow_list() { + assert_eq!( + classify_check(0, "checking a.pdf\nPDF Version: 1.7\n", ""), + SourceCheck::Clean, + "ENG-05 exit 0" + ); + let lin = "WARNING: a.pdf: linearization data is inconsistent"; + let hint = "WARNING: a.pdf: page 0: shared object 12: in hint table but not computed list"; + let summary = "qpdf: operation succeeded with warnings"; + assert_eq!( + classify_check(3, "", &format!("{lin}\n{summary}\n")), + SourceCheck::Benign(vec![lin.into()]), + "ENG-05 linearization" + ); + assert_eq!( + classify_check(3, &format!("{hint}\n"), summary), + SourceCheck::Benign(vec![hint.into()]), + "ENG-05 hint table (stdout)" + ); + let other = "WARNING: a.pdf (object 5 0): expected endobj"; + assert_eq!( + classify_check(3, "", &format!("{lin}\n{other}\n")), + SourceCheck::Problems(vec![lin.into(), other.into()]), + "ENG-05 any other warning" + ); + assert_eq!( + classify_check( + 3, + &format!("{other}\n"), + &format!("{other}\n{lin}\n{other}\n") + ), + SourceCheck::Problems(vec![other.into(), lin.into()]), + "ENG-05 each distinct line once, first-seen order (stderr then stdout)" + ); + assert_eq!( + classify_check(3, "", summary), + SourceCheck::Problems(vec!["qpdf reported warnings without details".into()]) + ); + assert_eq!( + classify_check(2, "", "qpdf: a.pdf: not a PDF file"), + SourceCheck::Problems(vec!["qpdf: a.pdf: not a PDF file".into()]) + ); + assert_eq!( + classify_check(-1, "", ""), + SourceCheck::Problems(vec!["qpdf --check exited with code -1".into()]) + ); + // one case per allow-list entry, and nothing else is benign + assert_eq!(BENIGN_CHECK_RULES.len(), 3); + assert!( + BenignRule::Contains("linearization").matches("WARNING: LINEARIZATION dictionary broken") + ); + assert!(BenignRule::Contains("hint table").matches("... Hint Table ...")); + assert!(!BenignRule::Contains("hint table").matches("WARNING: xref table damaged")); +} + +#[test] +fn eng06_size_mismatch_is_benign() { + let line = "WARNING: a.pdf: reported number of objects (3) is not one plus the highest object number (11)"; + assert!(BenignRule::SizeMismatch.matches(line), "ENG-06"); + assert!(!BenignRule::SizeMismatch.matches( + "WARNING: reported number of objects () is not one plus the highest object number (11)" + )); + assert!(!BenignRule::SizeMismatch.matches("WARNING: reported number of objects (3) is wrong")); + assert_eq!( + classify_check(3, "", line), + SourceCheck::Benign(vec![line.into()]) + ); + let Some(engines) = engines_or_skip("eng06_size_mismatch_is_benign") else { + return; + }; + let s = Scratch::new("eng06"); + let good = simple_pdf(b"q Q"); + let i = good.windows(6).rposition(|w| w == b"/Size ").unwrap() + 6; + let j = i + good[i..].iter().position(|c| *c == b' ').unwrap(); + let bad = [&good[..i], b"2", &good[j..]].concat(); + let p = s.write("badsize.pdf", &bad); + match qpdf_check(&engines, &p, &RunOpts::default()).unwrap() { + SourceCheck::Benign(lines) => assert!( + lines + .iter() + .all(|l| l.contains("reported number of objects")), + "ENG-06 {lines:?}" + ), + other => panic!("ENG-06 real qpdf on a wrong /Size: {other:?}"), + } + assert_eq!( + qpdf_check( + &engines, + &s.write("ok.pdf", &simple_pdf(b"q Q")), + &RunOpts::default() + ) + .unwrap(), + SourceCheck::Clean + ); +} + +const QPDF12: &str = r#"{"version":2,"parameters":{"decodelevel":"generalized"},"pages":[{"contents":["4 0 R","5 0 R"],"images":[],"label":null,"object":"3 0 R","outlines":[],"pageposfrom1":1},{"contents":[],"images":[],"label":null,"object":"10 0 R","outlines":[],"pageposfrom1":2}]}"#; +const QPDF119: &str = r#"{"version": 2, "parameters": {"decodelevel": "generalized"}, "pages": [{"contents": ["4 0 R"], "images": [{"object": "8 0 R", "width": 1}], "label": {"index": 0}, "object": "3 0 R", "outlines": [{"object": "12 0 R"}], "pageposfrom1": 1}]}"#; + +#[test] +fn eng07_page_map_json_shapes() { + let pages = parse_qpdf_pages(QPDF12.as_bytes()).unwrap(); + assert_eq!( + pages, + [ + QpdfPage { + object: (3, 0), + contents: vec![(4, 0), (5, 0)] + }, + QpdfPage { + object: (10, 0), + contents: vec![] + } + ], + "ENG-07 qpdf 12.x" + ); + assert_eq!( + parse_qpdf_pages(QPDF119.as_bytes()).unwrap(), + [QpdfPage { + object: (3, 0), + contents: vec![(4, 0)] + }], + "ENG-07 qpdf 11.9" + ); + for bad in [ + r#"{"pages":[{"object":"3 0 obj","contents":[]}]}"#, + r#"{"pages":[{"object":"03 0 R","contents":[]}]}"#, + r#"{"pages":[{"object":"3 0 R","contents":[]}]}"#, + r#"{"pages":[{"object":"3 0 R"}]}"#, + r#"{"pages":[{"object":"3 0 R","contents":["x"]}]}"#, + r#"{"pages":[{"object":"4294967296 0 R","contents":[]}]}"#, + r#"{"version":2}"#, + "not json", + ] { + assert!( + parse_qpdf_pages(bad.as_bytes()).is_err(), + "ENG-07 rejects {bad}" + ); + } + let Some(engines) = engines_or_skip("eng07_page_map_json_shapes") else { + return; + }; + let s = Scratch::new("eng07"); + let d = Doc::new(&[&[b"q", b"Q"], &[b"q Q"], &[]]); + let p = s.write("three.pdf", &d.build()); + let pages = qpdf_page_map(&engines, &p, &RunOpts::default()).unwrap(); + let want: Vec = d + .page_ids + .iter() + .zip(&d.content_ids) + .map(|(pg, c)| QpdfPage { + object: (*pg, 0), + contents: c.iter().map(|i| (*i, 0)).collect(), + }) + .collect(); + assert_eq!(pages, want, "ENG-07 installed qpdf"); + assert!(check_page_map(&read_snapshot(&p).unwrap(), &pages).is_ok()); + assert_eq!( + qpdf_page_map(&engines, &s.write("junk.pdf", b"junk"), &RunOpts::default()) + .unwrap_err() + .code, + "PDF_NEEDS_REPAIR" + ); +} + +/// A fingerprint numbered `n` for the file at `p`: its real length (the memo keeps only results +/// whose input still has the fingerprinted length, review T5 M1) and a made-up hash. +fn fp_of(p: &Path, n: u64) -> Fingerprint { + Fingerprint { + len: std::fs::metadata(p).map_or(0, |m| m.len()), + fnv: 0xE7_0000_0000 + n, + } +} + +#[test] +fn eng08_check_memo() { + let Some(engines) = engines_or_skip("eng08_check_memo") else { + return; + }; + let s = Scratch::new("eng08"); + let p = s.write("a.pdf", &simple_pdf(b"q Q")); + let before = runs(); + assert_eq!( + qpdf_check_memo(&engines, fp_of(&p, 1), &p, &RunOpts::default()).unwrap(), + SourceCheck::Clean + ); + assert_eq!( + qpdf_check_memo(&engines, fp_of(&p, 1), &p, &RunOpts::default()).unwrap(), + SourceCheck::Clean + ); + assert_eq!( + runs() - before, + 1, + "ENG-08 the second call is served from the memo without spawning" + ); + let broken = Engines { + qpdf: PathBuf::from("/nonexistent/offpdf/qpdf"), + ..engines.clone() + }; + let before = runs(); + for _ in 0..2 { + assert_eq!( + qpdf_check_memo(&broken, fp_of(&p, 2), &p, &RunOpts::default()) + .unwrap_err() + .code, + "ENGINE_MISSING" + ); + } + assert_eq!(runs() - before, 2, "ENG-08 errors are never memoised"); + for n in 10..(10 + crate::pdf_engine::text_edit::limits::CHECK_MEMO_MAX as u64) { + qpdf_check_memo(&engines, fp_of(&p, n), &p, &RunOpts::default()).unwrap(); + } + let before = runs(); + qpdf_check_memo(&engines, fp_of(&p, 1), &p, &RunOpts::default()).unwrap(); + assert_eq!( + runs() - before, + 1, + "ENG-08 least recently used entry evicted" + ); +} + +/// Review T5 M1: a check whose input vanished before qpdf opened it is an engine failure, never +/// memoised; the same bytes checked again later get their real verdict (it used to be +/// `PDF_NEEDS_REPAIR` for every later preview and Save of that file). +#[test] +fn eng08b_a_vanished_input_is_never_a_memoised_verdict() { + let Some(engines) = engines_or_skip("eng08b_a_vanished_input_is_never_a_memoised_verdict") + else { + return; + }; + let s = Scratch::new("eng08b"); + let p = s.write("a.pdf", &simple_pdf(b"q Q")); + let key = fp_of(&p, 0xB0); + let gone = s.path("gone").join("source.pdf"); + let missing = qpdf_check_memo(&engines, key, &gone, &RunOpts::default()); + assert_eq!( + missing.map_err(|e| e.code), + Err("ENGINE_FAILED".to_string()), + "a missing input is no verdict" + ); + assert_eq!( + qpdf_check_memo(&engines, key, &p, &RunOpts::default()).unwrap(), + SourceCheck::Clean, + "the good copy is checked, not served a memoised open failure" + ); + // A run whose input no longer has the fingerprinted length is returned but not kept. + let moved = s.write("b.pdf", &simple_pdf(b"q Q")); + let other = Fingerprint { + len: key.len + 1, + fnv: key.fnv + 1, + }; + let before = runs(); + for _ in 0..2 { + qpdf_check_memo(&engines, other, &moved, &RunOpts::default()).unwrap(); + } + assert_eq!(runs() - before, 2, "a length mismatch is never memoised"); +} + +#[test] +fn eng09_pending_check_waits_and_honours_cancel() { + let Some(engines) = engines_or_skip("eng09_pending_check_waits_and_honours_cancel") else { + return; + }; + let s = Scratch::new("eng09"); + let p = s.write("a.pdf", &simple_pdf(b"q Q")); + let pending = PendingCheck::spawn(engines.clone(), fp_of(&p, 100), p.clone()); + assert_eq!(pending.wait(None).unwrap(), SourceCheck::Clean, "ENG-09"); + assert_eq!( + pending.peek().map(|r| r.map_err(|e| e.code)), + Some(Ok(SourceCheck::Clean)) + ); + let missing = PendingCheck::spawn( + Engines { + qpdf: PathBuf::from("/nonexistent/qpdf"), + ..engines.clone() + }, + fp_of(&p, 101), + p.clone(), + ); + assert_eq!( + missing.wait(None).map_err(|e: AppError| e.code), + Err("ENGINE_MISSING".into()) + ); + #[cfg(unix)] + { + let slow = script(&s, "qpdf-slow", "exec sleep 3"); + let pending = PendingCheck::spawn( + Engines { + qpdf: slow, + ..engines + }, + fp_of(&p, 102), + p, + ); + let cancel = AtomicBool::new(true); + let t = Instant::now(); + assert_eq!( + pending.wait(Some(&cancel)).unwrap_err().code, + "CANCELLED", + "ENG-09 cancel" + ); + assert!(t.elapsed() < Duration::from_secs(1)); + assert!(pending.peek().is_none(), "still running in the background"); + } + // signatures production uses but tests cannot call without a Tauri app + let _resolve: fn(&tauri::AppHandle) -> Result = Engines::resolve; + let _standalone: fn() -> Option = Engines::standalone; +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_io/lex.rs b/src-tauri/src/pdf_engine/text_edit/tests_io/lex.rs new file mode 100644 index 0000000..05eb3d2 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_io/lex.rs @@ -0,0 +1,575 @@ +//! LEX-01…22 (+ scan modes): the byte-offset content tokenizer. + +use crate::pdf_engine::text_edit::lexer::{ + check_arity, lex_content, scan_dict_at, scan_tokens, InlineProof, LexError, LexLimits, Op, + Operand, Operator, ScanMode, Token, +}; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::testkit::pdf::zlib; +use std::sync::atomic::AtomicBool; + +fn lex(src: &[u8]) -> Result, LexError> { + lex_content(src, &LexLimits::page(), None) +} + +fn operators(src: &[u8]) -> Vec { + lex(src).unwrap().iter().map(|o| o.operator).collect() +} + +fn malformed(src: &[u8]) -> bool { + matches!(lex(src), Err(LexError::Malformed { .. })) +} + +fn str_operand(op: &Op) -> Vec { + op.operands[0].as_str_bytes().unwrap().to_vec() +} + +#[test] +fn lex01_whitespace_and_comments() { + let src = b"\x00\t q %comment ( [ <<\r\n1 0 0 1 5 5 %mid\n cm\x0cQ"; + let ops = lex(src).unwrap(); + assert_eq!( + ops.iter().map(|o| o.operator).collect::>(), + [Operator::q, Operator::cm, Operator::Q], + "LEX-01" + ); + let cm = &ops[1]; + assert_eq!( + &src[cm.span.clone()], + b"1 0 0 1 5 5 %mid\n cm", + "LEX-01 comment inside the op span" + ); + assert_eq!(&src[cm.op_span.clone()], b"cm"); + assert_eq!(ops[0].span, ops[0].op_span, "LEX-01 operand-less op span"); +} + +#[test] +fn lex02_literal_strings() { + let ops = + lex(b"(a(b)c) Tj (\\n\\r\\t\\b\\f\\(\\)\\\\) Tj (\\101\\060\\0617\\7\\q) Tj").unwrap(); + assert_eq!(str_operand(&ops[0]), b"a(b)c", "LEX-02 nesting"); + assert_eq!( + str_operand(&ops[1]), + b"\n\r\t\x08\x0c()\\", + "LEX-02 escapes" + ); + assert_eq!( + str_operand(&ops[2]), + b"A017\x07q", + "LEX-02 octal + unknown escape" + ); + let ops = lex(b"(line\\\ncont) Tj (a\r\nb) Tj (a\rb) Tj (a\\\r\nb) Tj").unwrap(); + let got: Vec> = ops.iter().map(str_operand).collect(); + assert_eq!( + got, + [ + b"linecont".to_vec(), + b"a\nb".to_vec(), + b"a\nb".to_vec(), + b"ab".to_vec() + ], + "LEX-02 EOL rules" + ); + assert!(malformed(b"(abc Tj"), "LEX-02 unterminated"); + let deep_ok = format!("{}{} Tj", "(".repeat(100), ")".repeat(100)); + assert!(lex(deep_ok.as_bytes()).is_ok(), "LEX-02 100 nested parens"); + let deep = format!("{}{} Tj", "(".repeat(101), ")".repeat(101)); + assert_eq!( + lex(deep.as_bytes()), + Err(LexError::TooComplex { + what: "string nesting" + }) + ); +} + +#[test] +fn lex03_hex_strings() { + let ops = lex(b"<48 65 6C6c\n6F> Tj <414> Tj").unwrap(); + assert_eq!(str_operand(&ops[0]), b"Hello", "LEX-03"); + assert_eq!( + str_operand(&ops[1]), + [0x41, 0x40], + "LEX-03 odd nibble padded with 0" + ); + assert!(matches!(ops[0].operands[0], Operand::Str { hex: true, .. })); + assert!(malformed(b"<4G> Tj"), "LEX-03 non-hex"); + assert!(malformed(b"<41 Tj"), "LEX-03 unterminated"); +} + +#[test] +fn lex04_names() { + let ops = lex(b"/A#20B gs /#41 gs /a#zz gs / gs").unwrap(); + let names: Vec<&[u8]> = ops + .iter() + .map(|o| o.operands[0].as_name().unwrap()) + .collect(); + assert_eq!(names, [&b"A B"[..], b"A", b"a#zz", b""], "LEX-04"); +} + +#[test] +fn lex05_numbers_ok_and_precision() { + let ops = lex(b".5 -.002 4. +3 re 0.123456789 g 1000000000 w").unwrap(); + let vals: Vec = ops[0] + .operands + .iter() + .map(|o| o.as_number().unwrap()) + .collect(); + assert_eq!(vals, [0.5, -0.002, 4.0, 3.0], "LEX-05"); + assert_eq!( + ops[1].operands[0].as_number(), + Some(0.123456789), + "LEX-05 f64 from exact digits" + ); + assert_eq!(ops[2].operands[0].as_number(), Some(1e9)); + assert!(malformed(b"1000000001 w"), "LEX-05 |x| > 1e9"); +} + +#[test] +fn lex06_numbers_bad() { + for src in [ + &b"--3 g"[..], + b"1e5 g", + b". g", + b"5- g", + b"+ g", + b"1.2.3 g", + b"0x10 g", + ] { + assert!(malformed(src), "LEX-06 {}", String::from_utf8_lossy(src)); + } +} + +#[test] +fn lex07_nesting_to_the_limit() { + let ok = format!("/P << /K {}{} >> BDC", "[".repeat(31), "]".repeat(31)); + assert!(lex(ok.as_bytes()).is_ok(), "LEX-07 32 levels"); + let over = format!("/P << /K {}{} >> BDC", "[".repeat(32), "]".repeat(32)); + assert_eq!( + lex(over.as_bytes()), + Err(LexError::TooComplex { + what: "array or dictionary nesting" + }), + "LEX-07 +1" + ); +} + +#[test] +fn lex08_array_items_to_the_limit() { + let make = |n: usize| format!("[{}] TJ", "1 ".repeat(n)); + assert!(lex(make(65_536).as_bytes()).is_ok(), "LEX-08"); + assert_eq!( + lex(make(65_537).as_bytes()), + Err(LexError::TooComplex { + what: "array items" + }), + "LEX-08 +1" + ); + let operands = format!("{}n", "1 ".repeat(65_537)); + assert_eq!( + lex(operands.as_bytes()), + Err(LexError::TooComplex { what: "operands" }) + ); +} + +#[test] +fn lex09_d0_d1_are_single_ops() { + assert_eq!( + operators(b"0 0 d0 0 0 0 0 1 1 d1"), + [Operator::d0, Operator::d1], + "LEX-09" + ); + assert_eq!(operators(b"[3 2] 0 d"), [Operator::d]); + let all = b"b B b* B* BDC BI BMC BT BX c cm CS cs d d0 d1 Do DP EI EMC ET EX f F f* G g gs h i ID j J K k l m M MP n q Q re RG rg ri s S SC sc SCN scn sh T* Tc Td TD Tf Tj TJ TL Tm Tr Ts Tw Tz v w W W* y ' \""; + let names: Vec<&[u8]> = all.split(|c| *c == b' ').collect(); + assert_eq!(names.len(), 73, "the 73 PDF 2.0 operators"); + for n in names { + let op = Operator::from_token(n); + assert_ne!(op, Operator::Unknown, "{}", String::from_utf8_lossy(n)); + assert_eq!(op.as_str().as_bytes(), n, "as_str inverts from_token"); + } + assert_eq!(Operator::from_token(b"Tx"), Operator::Unknown); +} + +#[test] +fn lex10_quote_ops_and_arity() { + let ops = lex(b"(a) ' 1 2 (b) \" T*").unwrap(); + assert_eq!( + ops.iter().map(|o| o.operator).collect::>(), + [Operator::Quote, Operator::DoubleQuote, Operator::TStar] + ); + assert_eq!(ops[1].operands.len(), 3, "LEX-10"); + for bad in [ + &b"1 2 (b) '"[..], + b"1 2 Tf", + b"/F1 Tf", + b"(a) (b) Tj", + b"[[1]] TJ", + b"1 2 3 cm", + b"/A BDC", + b"1 g 2 G 3 k", + ] { + assert!( + malformed(bad), + "LEX-10 arity {}", + String::from_utf8_lossy(bad) + ); + } + assert!(lex(b"/P <> BDC EMC /Pattern cs /P1 scn 0.5 /P1 scn 1 0 0 sc").is_ok()); +} + +#[test] +fn lex11_unknown_ops_and_compat() { + assert_eq!( + lex(b"q foo Q"), + Err(LexError::Malformed { + at: 2, + what: "unknown operator" + }), + "LEX-11" + ); + let ops = lex(b"BX foo 1 bar BX baz EX EX q").unwrap(); + let flags: Vec<(Operator, bool)> = ops.iter().map(|o| (o.operator, o.in_compat)).collect(); + assert_eq!( + flags, + [ + (Operator::BX, false), + (Operator::Unknown, true), + (Operator::Unknown, true), + (Operator::BX, true), + (Operator::Unknown, true), + (Operator::EX, true), + (Operator::EX, true), + (Operator::q, false) + ], + "LEX-11" + ); + assert_eq!(ops[2].operands.len(), 1); +} + +fn image(src: &[u8]) -> (Vec, usize) { + let ops = lex(src).unwrap(); + let i = ops + .iter() + .position(|o| o.operator == Operator::BI) + .expect("BI op"); + (ops, i) +} + +#[test] +fn lex12_inline_image_length_key() { + let src = b"q BI /W 2 /H 2 /CS /G /BPC 8 /L 4 ID a EI EI Q"; + let (ops, i) = image(src); + let img = ops[i].inline_image.as_ref().unwrap(); + assert_eq!(img.proof, InlineProof::LengthKey, "LEX-12"); + assert_eq!(&src[img.data.clone()], b"a EI"); + assert_eq!( + &src[ops[i].span.clone()], + b"BI /W 2 /H 2 /CS /G /BPC 8 /L 4 ID a EI EI" + ); + assert_eq!(ops[i + 1].operator, Operator::Q); + assert!(!ops[i + 1].after_unproven_inline_image); +} + +#[test] +fn lex13_inline_image_unfiltered_size() { + let src = b"BI /W 14 /H 1 /CS /G /BPC 8 ID xx EI yy EI zz\nEI Q"; + let (ops, i) = image(src); + let img = ops[i].inline_image.as_ref().unwrap(); + assert_eq!(img.proof, InlineProof::UnfilteredSize, "LEX-13"); + assert_eq!(&src[img.data.clone()], b"xx EI yy EI zz"); + let mask = b"BI /IM true /W 9 /H 2 ID \x00\x01\x02\x03 EI Q"; + let (ops, i) = image(mask); + assert_eq!( + ops[i].inline_image.as_ref().unwrap().proof, + InlineProof::UnfilteredSize, + "LEX-13 /IM: 2 bytes x 2 rows" + ); +} + +#[test] +fn lex14_inline_image_flate_end() { + let payload = zlib(b"EI EI EI raw pixels"); + let mut src = b"BI /W 19 /H 1 /CS /G /BPC 8 /F /Fl ID ".to_vec(); + src.extend_from_slice(&payload); + src.extend_from_slice(b"\nEI\nQ"); + let (ops, i) = image(&src); + let img = ops[i].inline_image.as_ref().unwrap(); + assert_eq!(img.proof, InlineProof::FlateEnd, "LEX-14"); + assert_eq!( + &src[img.data.clone()], + &payload[..], + "LEX-14 data incl. Adler-32" + ); +} + +#[test] +fn lex15_inline_image_ascii_ends() { + let src = b"BI /W 2 /H 1 /CS /G /BPC 8 /F /AHx ID 0A 0B> EI Q"; + let (ops, i) = image(src); + let img = ops[i].inline_image.as_ref().unwrap(); + assert_eq!( + (img.proof.clone(), &src[img.data.clone()]), + (InlineProof::AsciiEnd, &b"0A 0B>"[..]), + "LEX-15 AHx" + ); + let src = b"BI /W 4 /H 1 /CS /G /BPC 8 /F [/A85] ID 87cURD]i~> EI Q"; + let (ops, i) = image(src); + let img = ops[i].inline_image.as_ref().unwrap(); + assert_eq!( + (img.proof.clone(), &src[img.data.clone()]), + (InlineProof::AsciiEnd, &b"87cURD]i~>"[..]), + "LEX-15 A85" + ); +} + +#[test] +fn lex16_inline_image_dct_heuristic_marks_followers() { + let src = b"BI /W 2 /H 2 /CS /RGB /BPC 8 /F /DCT ID \xff\xd8 EI \x01\x02 junk\xff\xd9\nEI\nQ q 1 0 0 1 0 0 cm Q"; + let (ops, i) = image(src); + let img = ops[i].inline_image.as_ref().unwrap(); + assert_eq!(img.proof, InlineProof::Heuristic, "LEX-16"); + assert_eq!(&src[img.data.clone()], b"\xff\xd8 EI \x01\x02 junk\xff\xd9"); + assert!(!ops[i].after_unproven_inline_image); + assert!( + ops[i + 1..].iter().all(|o| o.after_unproven_inline_image), + "LEX-16 every later op is marked" + ); + assert_eq!(ops.len() - i - 1, 4); +} + +#[test] +fn lex17_inline_image_without_candidate_and_crlf() { + assert!( + malformed(b"BI /W 2 /H 2 /F /DCT ID \xff\xd8\x00\x01\x02"), + "LEX-17 no EI" + ); + assert!( + malformed(b"BI /W 2 /H 2 /F /DCT ID \xff\xd8 EIQ"), + "LEX-17 EI not delimited" + ); + let src = b"BI /W 3 /H 1 /CS /G /BPC 8 /L 3 ID\r\nabc EI Q"; + let (ops, i) = image(src); + let img = ops[i].inline_image.as_ref().unwrap(); + assert_eq!( + (&src[img.data.clone()], img.proof.clone()), + (&b"abc"[..], InlineProof::LengthKey), + "LEX-17 CR LF separator" + ); + assert!(malformed(b"BI /W 1 ID"), "LEX-17 no separator"); + // At most INLINE_HEURISTIC_CANDIDATES_MAX whitespace-delimited `EI`s are lexed (each lex + // reads a window): 255 failing candidates and the real end pass; one more is too complex. + let candidates = |n: usize| { + let mut src = b"BI /W 2 /H 2 /F /DCT ID \xff\xd8".to_vec(); + src.extend(b" EI \x01\x02".repeat(n)); + src.extend_from_slice(b"\nEI\nQ"); + src + }; + let max = limits::INLINE_HEURISTIC_CANDIDATES_MAX; + assert!( + lex(&candidates(max - 1)).is_ok(), + "LEX-17 candidates below the cap" + ); + assert!( + matches!(lex(&candidates(max)), Err(LexError::TooComplex { .. })), + "LEX-17 candidate cap" + ); +} + +#[test] +fn lex18_trailing_operands() { + for src in [ + &b"q 1 0 0 1 0 0"[..], + b"[1 2", + b"<< /A 1", + b"/F1 12 Tf /F1", + b"BT (x) Tj ET 5", + ] { + assert_eq!( + lex(src), + Err(LexError::Malformed { + at: src.len(), + what: "trailing operands" + }), + "LEX-18 {src:?}" + ); + } +} + +#[test] +fn lex19_stray_delimiters() { + for src in [ + &b"q ) Q"[..], + b"q > Q", + b"q ] Q", + b"q >> Q", + b"q { Q", + b"q } Q", + b"<< /A 1 >>", + b"[1] q", + ] { + assert!(lex(src).is_err(), "LEX-19 {}", String::from_utf8_lossy(src)); + } +} + +#[test] +fn lex20_garbage_mid_stream_is_an_error_never_a_prefix() { + let src = b"BT /F1 12 Tf (ok) Tj 1 0 0 RG \x01\x02garbage ( ET"; + assert!(lex(src).is_err(), "LEX-20"); + // lopdf's decoder stops at the garbage and returns the prefix as Ok (why it is forbidden). + let theirs = lopdf::content::Content::decode(src).map(|c| c.operations.len()); + println!("LEX-20 lopdf Content::decode on the probe: {theirs:?}"); + assert!( + !matches!(theirs, Ok(5)), + "LEX-20 probe: lopdf never sees the whole stream" + ); +} + +#[test] +fn lex21_budgets_checked_while_lexing() { + let limits = LexLimits { + ops_max: 10, + ..LexLimits::page() + }; + let mut src = "q ".repeat(11).into_bytes(); + src.extend_from_slice(b") garbage ("); + let err = lex_content(&src, &limits, None).unwrap_err(); + assert_eq!( + err, + LexError::TooComplex { what: "operations" }, + "LEX-21 counter probe" + ); + assert_eq!( + err.page_reason(), + crate::pdf_engine::text_edit::reasons::TextReason::PageTooComplex + ); + assert_eq!( + lex(b")").unwrap_err().page_reason(), + crate::pdf_engine::text_edit::reasons::TextReason::MalformedContent + ); + let many = "q Q ".repeat(3_000); + let cancel = AtomicBool::new(true); + assert_eq!( + lex_content(many.as_bytes(), &LexLimits::page(), Some(&cancel)), + Err(LexError::Cancelled), + "LEX-21 cancel" + ); + let few = "q Q ".repeat(1_000); + assert!( + lex_content(few.as_bytes(), &LexLimits::page(), Some(&cancel)).is_ok(), + "cancel checked every 4,096 ops" + ); +} + +/// Every valid sample of this file plus producer-like streams. +pub(crate) fn samples() -> Vec> { + let mut out: Vec> = [ + &b"\x00\t q %comment ( [ <<\r\n1 0 0 1 5 5 %mid\n cm\x0cQ"[..], + b"BT /F1 12 Tf 72 720 Td [(H) -20 (ello) 250 (World)] TJ ET", + b"q 1 0 0 1 0 0 cm 0.5 0.5 0.5 rg 10 10 100 50 re f Q /P <> BDC BT /F1 1 Tf 12 0 0 12 72 700 Tm (x) Tj ET EMC", + b"BT /F2 9 Tf <0041 0042> Tj T* (a) ' 1 2 (b) \" ET BX foo 1 bar EX", + b"q BI /W 2 /H 2 /CS /G /BPC 8 /L 4 ID a EI EI Q", + b"BI /W 14 /H 1 /CS /G /BPC 8 ID xx EI yy EI zz\nEI Q", + b"BI /W 2 /H 2 /CS /RGB /BPC 8 /F /DCT ID \xff\xd8 EI \x01\x02 junk\xff\xd9\nEI\nQ q 1 0 0 1 0 0 cm Q", + b"(a(b)c) Tj (\\n\\101) Tj <48 65> Tj /A#20B gs [3 2] 0 d 0 0 d0 0 0 0 0 1 1 d1 /Sh0 sh /Im1 Do", + b"0 0 m 10 10 l 1 2 3 4 5 6 c 1 2 3 4 v 1 2 3 4 y h W n 2 w 1 J 0 j 4 M /GS0 gs 0 Tc 1 Tw 100 Tz 12 TL 0 Tr 2 Ts", + ] + .iter() + .map(|s| s.to_vec()) + .collect(); + let mut flate = b"BI /W 19 /H 1 /CS /G /BPC 8 /F /Fl ID ".to_vec(); + flate.extend_from_slice(&zlib(b"EI EI EI raw pixels")); + flate.extend_from_slice(b"\nEI\nQ"); + out.push(flate); + out +} + +#[test] +fn lex22_exact_spans_property() { + for sample in samples() { + for op in lex(&sample).unwrap().into_iter().filter(|o| !o.in_compat) { + let sub = &sample[op.span.clone()]; + let again = lex(sub).unwrap_or_else(|e| { + panic!("LEX-22 re-lex of {:?}: {e}", String::from_utf8_lossy(sub)) + }); + assert_eq!(again.len(), 1, "LEX-22 {:?}", String::from_utf8_lossy(sub)); + let (a, b) = (&again[0], &op); + assert_eq!(a.operator, b.operator); + assert_eq!(a.span, 0..sub.len(), "LEX-22 span covers the slice"); + assert_eq!( + a.op_span.start + b.span.start..a.op_span.end + b.span.start, + b.op_span + ); + assert_eq!(a.operands.len(), b.operands.len()); + for (x, y) in a.operands.iter().zip(&b.operands) { + assert_eq!( + x.span().start + b.span.start..x.span().end + b.span.start, + y.span().clone() + ); + } + if let (Some(x), Some(y)) = (&a.inline_image, &b.inline_image) { + assert_eq!( + ( + x.data.start + b.span.start..x.data.end + b.span.start, + &x.dict.len() + ), + (y.data.clone(), &y.dict.len()) + ); + } + assert!(check_arity(a).is_ok() || a.operator == Operator::Unknown); + } + } +} + +#[test] +fn lex23_scan_token_modes() { + let toks = scan_tokens( + b"<< /Root 1 0 R /Size 5 /ID [ (x)] >>", + ScanMode::Object, + 100, + ) + .unwrap(); + assert!( + toks.contains(&Token::Keyword { + bytes: b"R".to_vec(), + span: 13..14 + }), + "Object mode keeps R as a keyword" + ); + assert!( + scan_tokens(b"{ 1 }", ScanMode::Object, 100).is_err(), + "braces outside a procedure" + ); + assert!(scan_tokens(b"1e-3 --5", ScanMode::Object, 100).is_err()); + let t1 = scan_tokens( + b"/FontMatrix [0.001 0 0 0.001 0 0] readonly def { 1e-3 } bind", + ScanMode::Type1Clear, + 100, + ) + .unwrap(); + assert!( + t1.iter() + .any(|t| matches!(t, Token::Keyword { bytes, .. } if bytes == b"1e-3")), + "Type1 lenient numbers" + ); + let cmap = scan_tokens( + b"/CIDInit /ProcSet findresource begin 1 begincodespacerange <00> endcodespacerange", + ScanMode::CMap, + 100, + ) + .unwrap(); + assert_eq!(cmap.len(), 9); + assert!( + scan_tokens(b"[ 1 2", ScanMode::CMap, 100).is_err(), + "unbalanced" + ); + assert_eq!( + scan_tokens(b"1 2 3", ScanMode::Object, 2), + Err(LexError::TooComplex { what: "tokens" }) + ); + let (dict, end) = scan_dict_at(b"trailer\n<< /Size 3 /Prev 10 >>\nstartxref", 7, 100).unwrap(); + assert_eq!((dict.len(), end), (6, 30)); + assert!( + scan_dict_at(b" 12 0 obj", 0, 100).is_err(), + "a dictionary is required" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_io/preflight.rs b/src-tauri/src/pdf_engine/text_edit/tests_io/preflight.rs new file mode 100644 index 0000000..4a95f18 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_io/preflight.rs @@ -0,0 +1,624 @@ +//! Preflight hardening (review-T1 fix passes): SNAP-08b/08c xref-stream fields, SNAP-10b/10c +//! nesting wherever lopdf starts parsing, SNAP-10d `/Length` chains, SNAP-10e scan budget, +//! SNAP-10f/10g linear header probes and budgeted `stream` probes, SNAP-10h one object-stream +//! scan budget per load, GUARD-02b preflight fuzz, GUARD-02c header finder vs a naive probe. +//! Fixtures with hand-placed xref offsets are written by `Raw`; the xref-chain cases are in +//! `tests_io/xref.rs`. + +use crate::error::AppError; +use crate::pdf_engine::text_edit::engines::{run_tool, RunOpts}; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::snapshot::headers::Headers; +use crate::pdf_engine::text_edit::snapshot::preflight::preflight; +use crate::pdf_engine::text_edit::snapshot::{read_snapshot, snapshot_from_bytes, SourceSnapshot}; +use crate::pdf_engine::text_edit::testkit::pdf::{simple_pdf, zlib, Doc, XrefStyle}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use std::collections::BTreeMap; +use std::ffi::OsString; +use std::path::Path; +use std::time::Instant; + +const HELLO: &[u8] = b"BT /F1 12 Tf 72 720 Td (Hello) Tj ET"; + +pub(super) fn from_bytes(bytes: &[u8]) -> Result { + snapshot_from_bytes(Path::new("fixture.pdf"), bytes.to_vec(), None) +} + +/// (code, details) of opening `bytes`; "OK" when it opens. +pub(super) fn outcome(bytes: &[u8]) -> (String, String) { + match from_bytes(bytes) { + Ok(_) => ("OK".to_string(), String::new()), + Err(e) => (e.code.clone(), e.details.unwrap_or_default()), + } +} + +pub(super) fn opens(bytes: &[u8], id: &str) -> SourceSnapshot { + from_bytes(bytes).unwrap_or_else(|e| panic!("{id}: {e} ({:?})", e.details)) +} + +pub(super) fn refused(bytes: &[u8], detail: &str, id: &str) { + let (code, details) = outcome(bytes); + assert_eq!(code, "FILE_TOO_COMPLEX", "{id}: {details}"); + assert!(details.contains(detail), "{id}: {details}"); +} + +pub(super) fn deep(n: usize) -> String { + format!("{}{}", "[".repeat(n), "]".repeat(n)) +} + +/// A file written byte by byte with a classic xref table whose offsets may point anywhere. +/// Objects 1–3 are a catalog, a page tree and one page. +pub(super) struct Raw { + out: Vec, + offs: BTreeMap, +} + +impl Raw { + pub(super) fn new() -> Raw { + let mut r = Raw { + out: b"%PDF-1.7\n".to_vec(), + offs: BTreeMap::new(), + }; + r.obj(1, b"<< /Type /Catalog /Pages 2 0 R >>"); + r.obj(2, b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>"); + r.obj( + 3, + b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] >>", + ); + r + } + + pub(super) fn obj(&mut self, id: u32, body: &[u8]) { + self.obj_at(&format!("{id} 0 obj\n"), id, 0, body); + } + + /// Writes `head`, `body`, `endobj`; the xref entry of `id` points `skip` bytes into `head`. + fn obj_at(&mut self, head: &str, id: u32, skip: usize, body: &[u8]) { + self.offs.insert(id, self.out.len() + skip); + self.out.extend_from_slice(head.as_bytes()); + self.out.extend_from_slice(body); + self.out.extend_from_slice(b"\nendobj\n"); + } + + pub(super) fn raw(&mut self, bytes: &[u8]) { + self.out.extend_from_slice(bytes); + } + + /// Points the xref entry of `id` at the first occurrence of `needle` written so far. + fn point(&mut self, id: u32, needle: &[u8]) { + let at = self + .out + .windows(needle.len()) + .position(|w| w == needle) + .expect("needle written"); + self.offs.insert(id, at); + } + + pub(super) fn finish(mut self) -> Vec { + let max = self.offs.keys().max().copied().unwrap_or(0); + let x = self.out.len(); + self.raw(format!("xref\n0 {}\n", max + 1).as_bytes()); + for id in 0..=max { + let row = match self.offs.get(&id) { + Some(off) if id != 0 => format!("{off:010} 00000 n \n"), + _ => "0000000000 65535 f \n".to_string(), + }; + self.raw(row.as_bytes()); + } + let tail = format!( + "trailer\n<< /Size {} /Root 1 0 R >>\nstartxref\n{x}\n%%EOF\n", + max + 1 + ); + self.raw(tail.as_bytes()); + self.out + } +} + +#[test] +fn snap08b_xref_index_counts_are_summed_without_overflow() { + let d = Doc::new(&[&[HELLO]]); + let stream = XrefStyle::Stream { + in_objstm: vec![], + bomb_mib: None, + }; + // 10,000 counts of 999,999,999,999,999 overflow an i64 sum (wrapped negative before the fix, + // a debug-build panic in the test binary). + let index = "0 999999999999999 ".repeat(10_000); + let bytes = + d.b.build_with(&format!("{} /Index [{index}]", d.trailer()), &stream); + refused(&bytes, "too many objects", "SNAP-08b overflowing /Index"); + let one = format!("{} /Index [0 {}]", d.trailer(), limits::MAX_OBJECTS + 1); + refused( + &d.b.build_with(&one, &stream), + "too many objects", + "SNAP-08b one count", + ); + let fine = d.b.build_with(&d.trailer(), &stream); + assert_eq!(opens(&fine, "SNAP-08b plain").pages.len(), 1); +} + +#[test] +fn snap08c_xref_predictor_row_is_bounded() { + let d = Doc::new(&[&[HELLO]]); + let stream = XrefStyle::Stream { + in_objstm: vec![], + bomb_mib: None, + }; + let with = |parms: &str| { + let trailer = format!("{} /DecodeParms << {parms} >>", d.trailer()); + d.b.build_with(&trailer, &stream) + }; + // lopdf allocates max(1, Columns) × max(1, Colors) × max(8, BPC) / 8 bytes per row: 2^62 + // aborted the process ("memory allocation failed") in the probe. + for parms in [ + "/Predictor 12 /Columns 4611686018427387904", + "/Predictor 10 /Columns 2000 /Colors 2000", + "/Predictor 15 /Columns 3 /BitsPerComponent 1000000000", + ] { + refused(&with(parms), "predictor row", &format!("SNAP-08c {parms}")); + } + // TIFF predictor 2 is ignored by lopdf's xref decoding, so its row is never allocated. + let tiff = with("/Predictor 2 /Columns 4611686018427387904"); + assert_eq!(outcome(&tiff).0, "OK", "SNAP-08c predictor 2"); + let Some(engines) = engines_or_skip("snap08c_xref_predictor_row_is_bounded") else { + return; + }; + let s = Scratch::new("snap08c"); + let src = s.write("in.pdf", &simple_pdf(HELLO)); + let out = s.path("xref-stream.pdf"); + let args = [ + OsString::from("--object-streams=generate"), + src.into(), + out.clone().into(), + ]; + let run = run_tool(&engines.qpdf, &args, false, &RunOpts::default()).unwrap(); + assert_eq!(run.code, 0, "SNAP-08c qpdf"); + let generated = std::fs::read(&out).unwrap(); + assert!( + generated.windows(13).any(|w| w == b"/Predictor 12"), + "SNAP-08c qpdf writes a PNG-predicted xref stream" + ); + assert_eq!( + read_snapshot(&out).unwrap().pages.len(), + 1, + "SNAP-08c a real predicted xref stream opens" + ); +} + +/// Object 99 (`[`×n `]`×n) where a linear scan never counts it: inside fake stream framing, +/// a literal string, a comment, or a header split by comments, or behind a header whose first +/// number only ends in 99. The xref entry of 99 points at it. +fn hidden_nesting(kind: &str, n: usize) -> Vec { + let value = format!("99 0 obj {} endobj", deep(n)); + let mut r = Raw::new(); + match kind { + "fake stream" => { + r.obj(4, b"<< /Length 1 >>\nstream\nx\nendstream"); + r.raw(format!("<<>>stream\n{value}\nendstream\n").as_bytes()); + } + "string" => r.obj(4, format!("({value})").as_bytes()), + "comment" => r.raw(format!("% {value}\n").as_bytes()), + "split header" => r.raw(format!("99 %a\n0 %b\nobj {}\nendobj\n", deep(n)).as_bytes()), + _ => r.raw(format!("1{value}\n").as_bytes()), // "199 0 obj …", xref points at "99" + } + r.point(99, b"99 "); + r.finish() +} + +#[test] +fn snap10b_nesting_anywhere_lopdf_can_parse_is_bounded() { + for kind in ["fake stream", "string", "comment", "split header", "suffix"] { + let shallow = opens(&hidden_nesting(kind, 3), "SNAP-10b"); + assert!( + shallow.doc.objects.contains_key(&(99, 0)), + "SNAP-10b {kind}: lopdf parses object 99 where its xref entry points" + ); + assert_eq!( + outcome(&hidden_nesting(kind, 100)).0, + "OK", + "SNAP-10b {kind} 100" + ); + refused( + &hidden_nesting(kind, 101), + "nested more than 100 levels", + &format!("SNAP-10b {kind} 101"), + ); + } +} + +#[test] +fn snap10c_object_stream_members_are_nesting_checked() { + let member = |n: usize| { + let mut d = Doc::new(&[&[HELLO]]); + let id = d.b.add(deep(n)); + let bytes = d.build_with(&XrefStyle::Stream { + in_objstm: vec![id], + bomb_mib: None, + }); + (bytes, id) + }; + let (ok, id) = member(100); + assert!( + opens(&ok, "SNAP-10c 100") + .doc + .objects + .contains_key(&(id, 0)), + "SNAP-10c member loaded" + ); + // Compressed: invisible to any raw scan; lopdf parses it after the guard decodes it. + let (bad, _) = member(101); + refused(&bad, "too deeply nested", "SNAP-10c 101"); +} + +/// `n` streams from id 100, each with `/Length` → the next id; the last → an integer object. +/// Each header is written as `{prefix}{id} 0 obj` with the xref entry `prefix.len()` bytes in. +fn length_chain(n: u32, prefix: &str) -> Vec { + let mut r = Raw::new(); + for id in 100..100 + n { + let body = format!("<< /Length {} 0 R >>\nstream\nx\nendstream", id + 1); + r.obj_at( + &format!("{prefix}{id} 0 obj\n"), + id, + prefix.len(), + body.as_bytes(), + ); + } + r.obj(100 + n, b"1"); + r.finish() +} + +#[test] +fn snap10d_length_reference_chains_are_bounded() { + let max = limits::LENGTH_REF_CHAIN_MAX as u32; + // n streams → lopdf parses n + 1 objects for the first one. + let ok = opens(&length_chain(max - 1, ""), "SNAP-10d at the limit"); + assert!(ok.doc.objects.contains_key(&(100, 0))); + refused( + &length_chain(max, ""), + "/Length references", + "SNAP-10d over", + ); + // Headers "7100 0 obj" with the xref pointing at "100": lopdf reads id 100 (suffix). + let suffix = opens(&length_chain(3, "7"), "SNAP-10d suffix"); + assert!( + suffix.doc.objects.contains_key(&(100, 0)), + "SNAP-10d lopdf parses the suffix id" + ); + refused( + &length_chain(max, "7"), + "/Length references", + "SNAP-10d suffix over", + ); + refused( + &length_chain(5_000, ""), + "/Length references", + "SNAP-10d 5,000", + ); +} + +#[test] +fn snap10d_indirect_lengths_of_real_producers_open() { + // pdfTeX/Ghostscript style: stream k, then its length object k + 1. + let mut after = Raw::new(); + for k in 0..3_000u32 { + let id = 10 + 2 * k; + let body = format!("<< /Length {} 0 R >>\nstream\nq Q\nendstream", id + 1); + after.obj(id, body.as_bytes()); + after.obj(id + 1, b"3"); + } + assert_eq!(outcome(&after.finish()).0, "OK", "SNAP-10d length after"); + // Lengths written first at low ids, streams later (suffixes of stream ids hit many lengths). + let mut first = Raw::new(); + for k in 10..2_010u32 { + first.obj(k, b"3"); + } + for k in 10..2_010u32 { + let body = format!("<< /Length {k} 0 R >>\nstream\nq Q\nendstream"); + first.obj(k + 2_000, body.as_bytes()); + } + assert_eq!(outcome(&first.finish()).0, "OK", "SNAP-10d lengths first"); +} + +#[test] +fn snap10e_overlapping_scans_hit_the_budget() { + // Each "9 0 obj" inside the nested strings starts a scan to the end of the object. + let n = 2_000; + let body = format!("{}{}", "[(9 0 obj ".repeat(n), ")]".repeat(n)); + let mut r = Raw::new(); + r.obj(4, body.as_bytes()); + let started = Instant::now(); + refused(&r.finish(), "object scan budget", "SNAP-10e"); + assert!(started.elapsed().as_secs() < 2, "SNAP-10e bounded time"); +} + +#[test] +fn guard02b_preflight_fuzz() { + const INSERTS: &[&[u8]] = &[ + b"[", + b"<<", + b"(", + b")", + b"%", + b"obj", + b" 1 0 obj ", + b"/Length 5 0 R", + b">>stream\n", + b"endstream", + b"#", + b"\\", + b"999999999999999", + b"/Index [0 9]", + b"/Predictor 12", + b"\n", + b"<", + b">", + b"trailer", + b"xref\n", + b"/Prev 9", + b"/XRefStm 9", + b"1.5", + b"%1 ", + ]; + let mut chain = Raw::new(); + for id in 10..20u32 { + let body = format!("<< /L#65ngth {} 0 R >>\nstream\nx\nendstream", id + 1); + chain.obj(id, body.as_bytes()); + } + let bases = [ + simple_pdf(HELLO), + Doc::new(&[&[HELLO]]).build_with(&XrefStyle::Stream { + in_objstm: vec![3], + bomb_mib: None, + }), + chain.finish(), + hidden_nesting("string", 50), + ]; + let mut state = 0x9E37_79B9_7F4A_7C15u64; + let mut next = move |below: usize| { + state ^= state << 13; + state ^= state >> 7; + state ^= state << 17; + (state % below.max(1) as u64) as usize + }; + let started = Instant::now(); + for case in 0..2_000 { + let mut bytes = bases[case % bases.len()].clone(); + for _ in 0..1 + next(4) { + let at = next(bytes.len() + 1); + match next(4) { + 0 => { + let ins = INSERTS[next(INSERTS.len())]; + let times = 1 + next(150); + let block: Vec = ins + .iter() + .copied() + .cycle() + .take(ins.len() * times) + .collect(); + bytes.splice(at..at, block); + } + 1 => bytes.truncate(at), + 2 if at < bytes.len() => bytes[at] = next(256) as u8, + _ => { + let ins = INSERTS[next(INSERTS.len())]; + bytes.splice(at..at, ins.iter().copied()); + } + } + } + let _ = preflight(&bytes); // must return (Ok or a refusal), never panic + } + assert!( + started.elapsed().as_secs() < 4, + "GUARD-02b 2,000 cases took {:?}", + started.elapsed() + ); +} + +/// Review round 2 (HIGH-1): `k` digit runs inside one comment line, then `tail`. Every run's +/// header probe walks through the rest of the line; the round-1 cache rescanned it per run. +fn comment_digits(k: usize, tail: &[u8]) -> Vec { + let mut body = b"%".to_vec(); + for _ in 0..k { + body.extend_from_slice(b"1 %"); + } + body.extend_from_slice(tail); + body.push(b'\n'); + let mut r = Raw::new(); + r.raw(&body); + r.finish() +} + +#[test] +fn snap10f_header_probes_are_linear() { + let k = 100_000; // ~0.6 MB of digit runs in one comment (round 1: minutes, quadratic) + let spaces = " ".repeat(3 * k); + let digits = "5".repeat(3 * k); + let comment = format!("%{}", "B".repeat(3 * k)); + let shapes: [(&str, Vec); 4] = [ + // the review's shape: the shared generation number is followed by a second comment + ("second comment", format!("\n5 {comment}").into_bytes()), + // a long run of whitespace after the shared generation number + ("whitespace", format!("\n5{spaces}x").into_bytes()), + // a long shared generation number + ("long generation", format!("\n{digits} x").into_bytes()), + // many comment lines between the runs and the generation number + ( + "comment lines", + format!("{}\n5 x", "\n%1 %1".repeat(k / 4)).into_bytes(), + ), + ]; + for (name, tail) in shapes { + let bytes = comment_digits(k, &tail); + let started = Instant::now(); + assert_eq!(outcome(&bytes).0, "OK", "SNAP-10f {name}"); + assert!( + started.elapsed().as_secs() < 2, + "SNAP-10f {name}: {} bytes took {:?}", + bytes.len(), + started.elapsed() + ); + } + // Headers reached only through the memoised comment walk are still scanned: every `1` in + // the line forms `1 … 5 obj` with the value after `obj`. + let hidden = comment_digits(k, format!("\n5 obj {}", deep(101)).as_bytes()); + let started = Instant::now(); + refused( + &hidden, + "nested more than 100 levels", + "SNAP-10f hidden header", + ); + assert!( + started.elapsed().as_secs() < 2, + "SNAP-10f hidden header time" + ); + let fine = comment_digits(3, format!("\n5 obj {}", deep(100)).as_bytes()); + assert_eq!(outcome(&fine).0, "OK", "SNAP-10f hidden header at 100"); +} + +#[test] +fn snap10g_stream_probe_after_a_value_is_budgeted() { + // Review round 2 (HIGH-2): 20 headers in comment lines all close at one `>>`, followed by + // 1 MiB of whitespace that each value scan's `stream` probe skips. The values themselves + // read ~3 KB; the probes read 20 MiB, over the budget of 2 × the file + 1 MiB only when they + // are debited (round 1 opened this file; with K headers it read K × the whitespace). + let mut body = b"1 0 obj <<\n".to_vec(); + for k in 2..=20 { + body.extend_from_slice(format!("%{k} 0 obj <<\n").as_bytes()); + } + body.extend_from_slice(b">>"); + body.resize(body.len() + (1 << 20), b' '); + body.extend_from_slice(b"\nendobj\n"); + let mut r = Raw::new(); + r.raw(&body); + let started = Instant::now(); + refused(&r.finish(), "object scan budget", "SNAP-10g"); + assert!( + started.elapsed().as_secs() < 2, + "SNAP-10g took {:?}", + started.elapsed() + ); +} + +/// `n` object streams (ids 10…) whose four members all start at offset 0 of one 8 KB array, +/// so each stream's member scans read ~32 KB. +fn overlapping_object_streams(n: u32) -> Vec { + let header = "100 0 101 0 102 0 103 0 "; + let mut data = header.as_bytes().to_vec(); + data.extend_from_slice(format!("[{}]", "0 ".repeat(4_000)).as_bytes()); + let packed = zlib(&data); + let mut r = Raw::new(); + for id in 10..10 + n { + let mut body = format!( + "<< /Type /ObjStm /N 4 /First {} /Filter /FlateDecode /Length {} >>\nstream\n", + header.len(), + packed.len() + ) + .into_bytes(); + body.extend_from_slice(&packed); + body.extend_from_slice(b"\nendstream"); + r.obj(id, &body); + } + r.finish() +} + +/// Restores the per-thread override when the test ends, also on a failed assertion. +struct ScanOverride; + +impl ScanOverride { + fn set(bytes: usize) -> ScanOverride { + limits::set_objstm_scan_override(Some(bytes)); + ScanOverride + } +} + +impl Drop for ScanOverride { + fn drop(&mut self) { + limits::set_objstm_scan_override(None); + } +} + +#[test] +fn snap10h_object_stream_scans_share_one_budget_per_load() { + let twenty = overlapping_object_streams(20); // ~640 KB of member scans in total + let three = overlapping_object_streams(3); // ~96 KB + let loaded = opens(&twenty, "SNAP-10h default budget"); + assert!( + loaded.doc.objects.contains_key(&(100, 0)), + "SNAP-10h members load" + ); + // Review round 2 (MEDIUM-1): each stream alone is far below its own 2 × data + 1 MiB, so + // round 1 (a fresh budget per stream) opened this file under any total. + let _budget = ScanOverride::set(128 << 10); + refused(&twenty, "too deeply nested", "SNAP-10h shared budget"); + // The budget is reset for every load: a file within it opens twice in a row. + for round in 1..=2 { + assert_eq!(outcome(&three).0, "OK", "SNAP-10h reset, load {round}"); + } +} + +#[test] +fn guard02c_header_finder_matches_a_naive_probe() { + // The memoised finder (snapshot/headers.rs) against lopdf's header grammar probed naively + // from every digit run; tokens chosen so that comments hide runs and headers are common. + fn space(b: &[u8], mut i: usize) -> usize { + loop { + match b.get(i) { + Some(b' ' | b'\t' | b'\n' | b'\r' | b'\x0C' | b'\0') => i += 1, + Some(b'%') => { + while b.get(i).is_some_and(|c| !matches!(c, b'\r' | b'\n')) { + i += 1; + } + } + _ => return i, + } + } + } + fn digits(b: &[u8], i: usize) -> usize { + i + b[i..].iter().take_while(|c| c.is_ascii_digit()).count() + } + fn naive(b: &[u8]) -> Vec<(std::ops::Range, usize)> { + let mut out = Vec::new(); + let mut i = 0; + while i < b.len() { + if !b[i].is_ascii_digit() { + i += 1; + continue; + } + let (start, end) = (i, digits(b, i)); + i = end; + let gen = space(b, end); + let gen_end = digits(b, gen); + let obj = space(b, gen_end); + if gen_end > gen && b.get(obj..obj + 3) == Some(&b"obj"[..]) { + out.push((start..end, obj + 3)); + } + } + out + } + const TOKENS: &[&[u8]] = &[ + b"1", b"12", b" ", b"\n", b"%", b"obj", b"x", b"0", b"\r", b" ", b"%1 ", b" 0 obj", + ]; + let mut state = 0x2545_F491_4F6C_DD1Du64; + let mut next = move |below: usize| { + state ^= state << 13; + state ^= state >> 7; + state ^= state << 17; + (state % below as u64) as usize + }; + let mut headers = 0; + for case in 0..20_000 { + let bytes: Vec = (0..1 + next(120)) + .flat_map(|_| TOKENS[next(TOKENS.len())].iter().copied()) + .collect(); + let expected = naive(&bytes); + headers += expected.len(); + let found: Vec<_> = Headers::new(&bytes).collect(); + assert_eq!( + found, + expected, + "GUARD-02c case {case}: {:?}", + String::from_utf8_lossy(&bytes) + ); + } + assert!(headers > 10_000, "GUARD-02c exercised {headers} headers"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_io/snap.rs b/src-tauri/src/pdf_engine/text_edit/tests_io/snap.rs new file mode 100644 index 0000000..42881a7 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_io/snap.rs @@ -0,0 +1,614 @@ +//! SNAP-01…16: one bounded read, preflight, guarded lopdf load, policy, page-map agreement. + +use crate::error::AppError; +use crate::pdf_engine::text_edit::engines::{qpdf_page_map, run_tool, RunOpts}; +use crate::pdf_engine::text_edit::limits; +use crate::pdf_engine::text_edit::snapshot::{ + check_page_map, fnv1a_u64, read_snapshot, read_verification_snapshot, snapshot_from_bytes, + stat_matches, Fingerprint, SourceSnapshot, +}; +use crate::pdf_engine::text_edit::testkit::pdf::{simple_pdf, zlib_zero_bomb, Doc, XrefStyle}; +use crate::pdf_engine::text_edit::testkit::{ + child_mode, child_value, engines_or_skip, process_peak, run_child_test, thread_peak, Scratch, +}; +use crate::pdf_engine::validate_output::content_digest; +use lopdf::Object; +use std::ffi::OsString; +use std::path::Path; + +fn code(r: &Result) -> String { + match r { + Ok(_) => "OK".to_string(), + Err(e) => e.code.clone(), + } +} + +fn from_bytes(bytes: &[u8]) -> Result { + snapshot_from_bytes(Path::new("fixture.pdf"), bytes.to_vec(), None) +} + +fn ok(r: Result, id: &str) -> SourceSnapshot { + r.unwrap_or_else(|e| panic!("{id}: {e} ({:?})", e.details)) +} + +const HELLO: &[u8] = b"BT /F1 12 Tf 72 720 Td (Hello) Tj ET"; + +#[test] +fn snap01_single_read_and_fingerprint() { + let s = Scratch::new("snap01"); + let bytes = simple_pdf(HELLO); + let p = s.write("a.pdf", &bytes); + let snap = ok(read_snapshot(&p), "SNAP-01"); + assert_eq!( + snap.bytes.as_slice(), + &bytes[..], + "SNAP-01 bytes are the one read" + ); + assert_eq!( + snap.fingerprint, + Fingerprint { + len: bytes.len() as u64, + fnv: fnv1a_u64(&bytes) + }, + "SNAP-01" + ); + assert_eq!( + snap.fingerprint.fnv, + content_digest(&bytes).hash, + "same FNV-1a as validate_output" + ); + assert_eq!(snap.pages.len(), 1); + assert_eq!( + snap.modified, + std::fs::metadata(&p).unwrap().modified().ok() + ); + let text = snap.fingerprint.to_string(); + assert_eq!( + text.parse::(), + Ok(snap.fingerprint), + "SNAP-01 strict inverse" + ); + assert_eq!( + Fingerprint { len: 0, fnv: 1 }.to_string(), + "0000000000000001-0" + ); + for bad in [ + "", + "0123", + "ABCDEF0123456789-1f", + "0123456789abcdef-01", + "0123456789abcdef-", + "0123456789abcdef", + "0123456789abcdef-1-2", + "0123456789abcdeg-1", + " 0123456789abcdef-1", + ] { + assert!( + bad.parse::().is_err(), + "SNAP-01 rejects {bad:?}" + ); + } + assert_eq!( + ok(from_bytes(&bytes), "SNAP-01").fingerprint, + snap.fingerprint + ); +} + +#[test] +fn snap02_over_the_cap_is_file_too_large() { + let s = Scratch::new("snap02"); + let p = s.path("sparse.pdf"); + std::fs::File::create(&p) + .unwrap() + .set_len(limits::FILE_CAP_BYTES + 1) + .unwrap(); + assert_eq!( + code(&read_snapshot(&p)), + "FILE_TOO_LARGE", + "SNAP-02 400 MiB + 1 (sparse)" + ); + let small = s.write("small.pdf", &simple_pdf(HELLO)); + limits::set_file_cap_override(Some(100)); + let (a, b) = ( + code(&read_snapshot(&small)), + code(&from_bytes(&simple_pdf(HELLO))), + ); + limits::set_file_cap_override(None); + assert_eq!( + (a.as_str(), b.as_str()), + ("FILE_TOO_LARGE", "FILE_TOO_LARGE"), + "SNAP-02 cap override" + ); + assert_eq!(code(&read_snapshot(&s.path("missing.pdf"))), "INVALID_PDF"); + assert_eq!( + code(&read_snapshot(s.dir())), + "INVALID_PDF", + "a directory is not a file" + ); +} + +#[test] +fn snap03_rewritten_file_has_a_different_fingerprint() { + let s = Scratch::new("snap03"); + let p = s.write("a.pdf", &simple_pdf(HELLO)); + let before = ok(read_snapshot(&p), "SNAP-03"); + std::fs::write(&p, simple_pdf(b"BT /F1 12 Tf 72 720 Td (Hellp) Tj ET")).unwrap(); + let after = ok(read_snapshot(&p), "SNAP-03"); + assert_eq!(before.fingerprint.len, after.fingerprint.len); + assert_ne!(before.fingerprint, after.fingerprint, "SNAP-03"); + assert_eq!(code(&from_bytes(b"not a pdf at all")), "INVALID_PDF"); +} + +#[test] +fn snap04_encrypted_is_refused() { + let d = Doc::new(&[&[HELLO]]); + let direct = d.b.build(&format!( + "{} /Encrypt << /Filter /Standard /V 1 /R 2 /P -4 >>", + d.trailer() + )); + assert_eq!( + code(&from_bytes(&direct)), + "ENCRYPTED", + "SNAP-04 direct /Encrypt dict" + ); + let mut d2 = Doc::new(&[&[HELLO]]); + let enc = d2.b.add("<< /Filter /Standard /V 1 /R 2 /P -4 >>"); + let indirect = d2.b.build(&format!("{} /Encrypt {enc} 0 R", d2.trailer())); + assert_eq!( + code(&from_bytes(&indirect)), + "ENCRYPTED", + "SNAP-04 indirect /Encrypt" + ); + let Some(engines) = engines_or_skip("snap04_encrypted_is_refused") else { + return; + }; + let s = Scratch::new("snap04"); + let plain = s.write("plain.pdf", &simple_pdf(HELLO)); + let out = s.path("enc.pdf"); + let args: Vec = ["--encrypt", "user", "owner", "256", "--"] + .iter() + .map(OsString::from) + .chain([plain.into(), out.clone().into()]) + .collect(); + let r = run_tool(&engines.qpdf, &args, false, &RunOpts::default()).unwrap(); + assert_eq!(r.code, 0, "{}", r.stderr); + assert_eq!( + code(&read_snapshot(&out)), + "ENCRYPTED", + "SNAP-04 qpdf --encrypt (FX-ENC)" + ); +} + +fn with_catalog(extra: &str, add: impl FnOnce(&mut Doc) -> String) -> Vec { + let mut d = Doc::new(&[&[HELLO]]); + let more = add(&mut d); + d.b.set( + d.catalog, + format!("<< /Type /Catalog /Pages {} 0 R {extra} {more} >>", d.pages), + ); + d.build() +} + +#[test] +fn snap05_applied_signature_is_signed_empty_widget_is_fine() { + let signed = with_catalog("", |d| { + let sig = d + .b + .add("<< /Type /Sig /Filter /Adobe.PPKLite /ByteRange [0 10 20 30] /Contents <00> >>"); + let field = d.b.add(format!( + "<< /FT /Sig /T (s) /V {sig} 0 R /Subtype /Widget /Rect [0 0 0 0] >>" + )); + format!("/AcroForm << /Fields [{field} 0 R] /SigFlags 3 >>") + }); + assert_eq!( + code(&from_bytes(&signed)), + "SIGNED", + "SNAP-05 applied signature" + ); + let empty = with_catalog("", |d| { + let field = + d.b.add("<< /FT /Sig /T (s) /Subtype /Widget /Rect [0 0 0 0] >>"); + format!("/AcroForm << /Fields [{field} 0 R] >>") + }); + assert_eq!( + code(&from_bytes(&empty)), + "OK", + "SNAP-05 empty signature widget" + ); + let perms = with_catalog("/Perms << /DocMDP << /Type /SigRef >> >>", |_| { + String::new() + }); + assert_eq!(code(&from_bytes(&perms)), "SIGNED", "SNAP-05 /Perms"); +} + +#[test] +fn snap06_xfa_is_refused() { + let xfa = with_catalog("", |d| { + let x = d.b.add_stream("", b""); + format!("/AcroForm << /Fields [] /XFA {x} 0 R >>") + }); + assert_eq!(code(&from_bytes(&xfa)), "UNSUPPORTED_XFA", "SNAP-06"); + let needs = with_catalog("/NeedsRendering true", |_| String::new()); + assert_eq!( + code(&from_bytes(&needs)), + "UNSUPPORTED_XFA", + "SNAP-06 NeedsRendering" + ); + let plain_form = with_catalog("/AcroForm << /Fields [] /XFA null >>", |_| String::new()); + assert_eq!(code(&from_bytes(&plain_form)), "OK"); +} + +fn objstm_bomb_pdf() -> Vec { + let mut d = Doc::new(&[&[HELLO]]); + let bomb = zlib_zero_bomb(1024); + d.b.add_stream("/Type /ObjStm /N 1 /First 4 /Filter /FlateDecode", &bomb); + d.build() +} + +#[test] +fn snap07_objstm_bomb_is_file_too_complex_with_bounded_memory() { + assert_eq!( + code(&from_bytes(&objstm_bomb_pdf())), + "FILE_TOO_COMPLEX", + "SNAP-07" + ); + let out = run_child_test( + "pdf_engine::text_edit::tests_io::snap::snap07_child", + "snap07", + ); + let (file, peak) = (child_value(&out, "FILE"), child_value(&out, "PEAK")); + assert_eq!(child_value(&out, "TOO_COMPLEX"), 1, "SNAP-07 child"); + assert!( + peak < file + (48 << 20), + "SNAP-07 peak {peak} for a 1 GiB bomb in a {file}-byte file" + ); +} + +#[test] +#[ignore = "child process of snap07 (run by it)"] +fn snap07_child() { + if child_mode().as_deref() != Some("snap07") { + return; + } + let bytes = objstm_bomb_pdf(); + let len = bytes.len(); + let (r, peak) = process_peak(|| snapshot_from_bytes(Path::new("bomb.pdf"), bytes, None)); + println!( + "FILE={len}\nPEAK={peak}\nTOO_COMPLEX={}", + u8::from(code(&r) == "FILE_TOO_COMPLEX") + ); +} + +#[test] +fn snap08_xref_stream_bomb_is_file_too_complex() { + let d = Doc::new(&[&[HELLO]]); + let bytes = d.build_with(&XrefStyle::Stream { + in_objstm: vec![], + bomb_mib: Some(1024), + }); + let (r, peak) = thread_peak(|| from_bytes(&bytes)); + assert_eq!(code(&r), "FILE_TOO_COMPLEX", "SNAP-08"); + assert!( + peak < bytes.len() + (8 << 20), + "SNAP-08 preflight peak {peak}" + ); + let fine = d.build_with(&XrefStyle::Stream { + in_objstm: vec![], + bomb_mib: None, + }); + assert_eq!( + ok(from_bytes(&fine), "SNAP-08 plain xref stream") + .pages + .len(), + 1 + ); +} + +fn answer_42(id: (u32, u16), obj: &mut Object) -> Option<((u32, u16), Object)> { + if let Object::Dictionary(d) = obj { + d.set("Mark", 1); + } + Some((id, Object::Integer(42))) +} + +#[test] +fn snap09_objstm_guard_preserves_objects() { + let d = Doc::new(&[&[HELLO], &[b"q Q"]]); + let members = vec![d.pages, 3, d.page_ids[1]]; + let bytes = d.build_with(&XrefStyle::Stream { + in_objstm: members.clone(), + bomb_mib: None, + }); + let ours = ok(from_bytes(&bytes), "SNAP-09").doc; + let plain = lopdf::Document::load_mem(&bytes).unwrap(); + assert_eq!( + ours.objects, plain.objects, + "SNAP-09 guarded load == plain lopdf load" + ); + assert_eq!(ours.get_pages().len(), 2); + // Pins lopdf 0.34: top-level return value ignored (the in-place mutation is kept), + // object-stream members use the returned value. + let pinned = lopdf::Reader { + buffer: &bytes, + document: lopdf::Document::new(), + } + .read(Some(answer_42)) + .unwrap(); + match pinned.objects.get(&(d.catalog, 0)) { + Some(Object::Dictionary(cat)) => assert!( + cat.has(b"Mark"), + "SNAP-09 top level keeps the mutated object" + ), + other => panic!("SNAP-09 catalog: {other:?}"), + } + for m in members { + assert_eq!( + pinned.objects.get(&(m, 0)), + Some(&Object::Integer(42)), + "SNAP-09 member {m} uses the returned value" + ); + } + let Some(engines) = engines_or_skip("snap09_objstm_guard_preserves_objects") else { + return; + }; + let s = Scratch::new("snap09"); + let src = s.write("in.pdf", &Doc::new(&[&[HELLO], &[b"q Q"]]).build()); + let out = s.path("objstm.pdf"); + let args = [ + OsString::from("--object-streams=generate"), + src.into(), + out.clone().into(), + ]; + assert_eq!( + run_tool(&engines.qpdf, &args, false, &RunOpts::default()) + .unwrap() + .code, + 0 + ); + let generated = std::fs::read(&out).unwrap(); + let ours = ok(read_snapshot(&out), "SNAP-09 qpdf").doc; + assert!(ours + .objects + .values() + .any(|o| matches!(o, Object::Stream(s) if s.dict.type_is(b"ObjStm")))); + assert_eq!( + ours.objects, + lopdf::Document::load_mem(&generated).unwrap().objects, + "SNAP-09 qpdf --object-streams=generate" + ); +} + +#[test] +fn snap10_deep_nesting_is_file_too_complex() { + let nested = |n: usize| { + let mut d = Doc::new(&[&[HELLO]]); + d.b.add(format!("{}{}", "[".repeat(n), "]".repeat(n))); + d.build() + }; + assert_eq!(code(&from_bytes(&nested(100))), "OK", "SNAP-10 100 levels"); + assert_eq!( + code(&from_bytes(&nested(101))), + "FILE_TOO_COMPLEX", + "SNAP-10 101 levels" + ); + let mut d = Doc::new(&[&[HELLO]]); + d.b.add(format!("({})", "[<<".repeat(200))); + d.b.add_stream("", "[[[[".repeat(200).as_bytes()); + d.b.add(format!("<{}>", "ab".repeat(10))); + assert_eq!( + code(&from_bytes(&d.build())), + "OK", + "SNAP-10 strings, streams, hex strings are skipped" + ); + let dicts = format!("{}{}", "<< /A ".repeat(101), ">> ".repeat(101)); + let mut d = Doc::new(&[&[HELLO]]); + d.b.add(dicts); + assert_eq!( + code(&from_bytes(&d.build())), + "FILE_TOO_COMPLEX", + "SNAP-10 dictionaries count too" + ); +} + +fn page_map_code(bytes: &[u8], id: &str) -> Option { + let engines = engines_or_skip(id)?; + let s = Scratch::new("pagemap"); + let p = s.write("f.pdf", bytes); + let snap = ok(read_snapshot(&p), id); + let pages = + qpdf_page_map(&engines, &p, &RunOpts::default()).unwrap_or_else(|e| panic!("{id}: {e}")); + Some(match check_page_map(&snap, &pages) { + Ok(()) => "OK".to_string(), + Err(e) => e.code, + }) +} + +#[test] +fn snap11_hybrid_objects_only_in_xrefstm_need_repair() { + let d = Doc::new(&[&[HELLO]]); + let bytes = d.build_with(&XrefStyle::Hybrid { + in_objstm: vec![d.page_ids[0]], + objstm_in_classic: false, + }); + assert_eq!( + ok(from_bytes(&bytes), "SNAP-11").pages.len(), + 0, + "lopdf misses the page" + ); + if let Some(c) = page_map_code(&bytes, "snap11") { + assert_eq!(c, "PDF_NEEDS_REPAIR", "SNAP-11"); + } +} + +#[test] +fn snap12_kid_without_type_needs_repair() { + let mut d = Doc::new(&[&[HELLO], &[b"q Q"]]); + let (pages, content) = (d.pages, d.content_ids[1][0]); + d.b.set( + d.page_ids[1], + format!("<< /Parent {pages} 0 R /Contents {content} 0 R >>"), + ); + let bytes = d.build(); + assert_eq!( + ok(from_bytes(&bytes), "SNAP-12").pages.len(), + 1, + "lopdf skips the kid" + ); + if let Some(c) = page_map_code(&bytes, "snap12") { + assert_eq!(c, "PDF_NEEDS_REPAIR", "SNAP-12"); + } +} + +fn image_streams_pdf(mib_each: usize) -> Vec { + let mut d = Doc::new(&[&[b"q 100 0 0 100 0 0 cm /Im0 Do Q"]]); + let pixels = vec![0u8; mib_each << 20]; + for _ in 0..2 { + d.b.add_stream("/Type /XObject /Subtype /Image /Width 1024 /Height 1024 /ColorSpace /DeviceGray /BitsPerComponent 8", &pixels); + } + d.build() +} + +#[test] +fn snap13_guard_never_clones_streams() { + let out = run_child_test( + "pdf_engine::text_edit::tests_io::snap::snap13_child", + "snap13", + ); + let (file, peak) = (child_value(&out, "FILE"), child_value(&out, "PEAK")); + assert_eq!(child_value(&out, "PAGES"), 1); + assert!( + file > 200 << 20, + "SNAP-13 fixture has 200 MiB of image streams" + ); + assert!( + (peak as f64) <= 1.3 * file as f64, + "SNAP-13 peak {peak} ≤ 1.3 × {file}" + ); +} + +#[test] +#[ignore = "child process of snap13 (run by it)"] +fn snap13_child() { + if child_mode().as_deref() != Some("snap13") { + return; + } + let bytes = image_streams_pdf(100); + let len = bytes.len(); + let (r, peak) = process_peak(|| { + snapshot_from_bytes(Path::new("images.pdf"), bytes, None).map(|s| s.pages.len()) + }); + println!("FILE={len}\nPEAK={peak}\nPAGES={}", r.unwrap_or(0)); +} + +#[test] +fn snap14_verification_snapshot_skips_policy_keeps_bounds() { + let s = Scratch::new("snap14"); + let vcode = |p: &Path, cap: u64| match read_verification_snapshot(p, cap) { + Ok(_) => "OK".to_string(), + Err(e) => format!("{}|{}", e.code, e.details.unwrap_or_default()), + }; + let signed = with_catalog("/Perms << /DocMDP << >> >>", |_| String::new()); + let xfa = with_catalog("/NeedsRendering true", |_| String::new()); + assert_eq!( + vcode(&s.write("signed.pdf", &signed), 1 << 30), + "OK", + "SNAP-14 signed output opens" + ); + assert_eq!( + vcode(&s.write("xfa.pdf", &xfa), 1 << 30), + "OK", + "SNAP-14 XFA output opens" + ); + let bomb = vcode(&s.write("bomb.pdf", &objstm_bomb_pdf()), 1 << 30); + assert!( + bomb.starts_with("EDIT_VERIFY_FAILED|") && bomb.contains("too large to verify"), + "SNAP-14 ObjStm bound: {bomb}" + ); + let mut deep = Doc::new(&[&[HELLO]]); + deep.b + .add(format!("{}{}", "[".repeat(101), "]".repeat(101))); + assert!( + vcode(&s.write("deep.pdf", &deep.build()), 1 << 30).contains("FILE_TOO_COMPLEX"), + "SNAP-14 nesting bound" + ); + let small = s.write("small.pdf", &simple_pdf(HELLO)); + let len = std::fs::metadata(&small).unwrap().len(); + assert_eq!(vcode(&small, len), "OK", "SNAP-14 cap inclusive"); + let over = vcode(&small, len - 1); + assert!( + over.starts_with("EDIT_VERIFY_FAILED|") + && over.contains("too large to verify: FILE_TOO_LARGE"), + "SNAP-14 cap: {over}" + ); + let d = Doc::new(&[&[HELLO]]); + let enc = + d.b.build(&format!("{} /Encrypt << /Filter /Standard >>", d.trailer())); + let enc = vcode(&s.write("enc.pdf", &enc), 1 << 30); + assert!( + enc.starts_with("EDIT_VERIFY_FAILED|") + && enc.contains("could not read the checked file: ENCRYPTED"), + "SNAP-14 encrypted: {enc}" + ); +} + +#[test] +fn snap15_stat_matches_detects_same_length_rewrite() { + let s = Scratch::new("snap15"); + let p = s.write("a.pdf", &simple_pdf(HELLO)); + let snap = ok(read_snapshot(&p), "SNAP-15"); + assert!(stat_matches(&snap), "SNAP-15 unchanged"); + let mut other = simple_pdf(HELLO); + let pos = other.windows(5).position(|w| w == b"Hello").unwrap(); + other[pos] = b'J'; + std::fs::write(&p, &other).unwrap(); + let f = std::fs::File::options().write(true).open(&p).unwrap(); + f.set_modified(snap.modified.unwrap()).unwrap(); + drop(f); + assert_eq!( + std::fs::metadata(&p).unwrap().modified().ok(), + snap.modified, + "mtime restored" + ); + assert!( + !stat_matches(&snap), + "SNAP-15 same length + same mtime, different bytes" + ); + std::fs::write(&p, b"short").unwrap(); + assert!(!stat_matches(&snap)); + std::fs::remove_file(&p).unwrap(); + assert!(!stat_matches(&snap)); +} + +#[test] +fn snap16_word_style_hybrid_loads_and_agrees() { + let d = Doc::new(&[&[HELLO], &[b"q", b"Q"]]); + let bytes = d.build_with(&XrefStyle::Hybrid { + in_objstm: vec![d.page_ids[0], d.pages], + objstm_in_classic: true, + }); + let snap = ok(from_bytes(&bytes), "SNAP-16"); + assert_eq!( + snap.pages.len(), + 2, + "SNAP-16 members reached through the classic-listed object stream" + ); + if let Some(c) = page_map_code(&bytes, "snap16") { + assert_eq!(c, "OK", "SNAP-16"); + } + // A /Contents disagreement is found and named. + let Some(engines) = engines_or_skip("snap16_word_style_hybrid_loads_and_agrees") else { + return; + }; + let s = Scratch::new("snap16"); + let p = s.write("two.pdf", &d.build()); + let snap = ok(read_snapshot(&p), "SNAP-16"); + let mut pages = qpdf_page_map(&engines, &p, &RunOpts::default()).unwrap(); + assert!(check_page_map(&snap, &pages).is_ok()); + pages[1].contents.reverse(); + let e = check_page_map(&snap, &pages).unwrap_err(); + assert_eq!(e.code, "PDF_NEEDS_REPAIR"); + assert!( + e.details.unwrap_or_default().contains("page 2: /Contents"), + "first difference named" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_io/xref.rs b/src-tauri/src/pdf_engine/text_edit/tests_io/xref.rs new file mode 100644 index 0000000..2962d26 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_io/xref.rs @@ -0,0 +1,229 @@ +//! The cross-reference chain the preflight checks is the one lopdf 0.34 reads (review-T1 round +//! 2, HIGH-3): SNAP-08d duplicate keys count with their last value, SNAP-08e a table's trailer +//! is the one lopdf's grammar reaches, SNAP-08f `/XRefStm` of an xref-stream trailer is followed, +//! SNAP-08g xref-stream arrays must be integers. Each refused case is one that round 1 let +//! through to lopdf (its outcome there is noted per case). + +use super::preflight::{deep, opens, outcome}; +use std::fmt::Write as _; + +const OBJECTS: [&str; 3] = [ + "<< /Type /Catalog /Pages 2 0 R >>", + "<< /Type /Pages /Kids [3 0 R] /Count 1 >>", + "<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] >>", +]; + +/// lopdf allocates this predictor row (4,000,000 bytes) only when the stream is Flate-encoded; +/// these xref streams are not, so lopdf reads them fine while the preflight refuses the row. +/// That makes "which section did the preflight check" visible without crashing lopdf. +const PREDICTOR_BOMB: &str = "/DecodeParms << /Predictor 12 /Columns 2000 /Colors 2000 >>"; + +/// Objects 1–3 written by hand, then xref sections whose offsets the test wires together. +struct File { + out: Vec, + offs: Vec, +} + +impl File { + fn new() -> File { + let mut f = File { + out: b"%PDF-1.7\n".to_vec(), + offs: Vec::new(), + }; + for (i, body) in OBJECTS.iter().enumerate() { + f.offs.push(f.out.len()); + f.out + .extend_from_slice(format!("{} 0 obj\n{body}\nendobj\n", i + 1).as_bytes()); + } + f + } + + /// An xref stream object `id` listing objects 0–3 (`/W [1 4 2]`, unfiltered). `pre` goes + /// before and `post` after the standard entries, so a key in both occurs twice. + fn xref_stream(&mut self, id: u32, pre: &str, post: &str) -> usize { + let mut rows = vec![0u8, 0, 0, 0, 0, 0xFF, 0xFF]; + for off in &self.offs { + rows.push(1); + rows.extend_from_slice(&u32::try_from(*off).unwrap().to_be_bytes()); + rows.extend_from_slice(&[0, 0]); + } + let at = self.out.len(); + let head = format!( + "{id} 0 obj\n<< {pre} /Type /XRef /Size 4 /W [1 4 2] {post} /Length {} >>\nstream\n", + rows.len() + ); + self.out.extend_from_slice(head.as_bytes()); + self.out.extend_from_slice(&rows); + self.out.extend_from_slice(b"\nendstream\nendobj\n"); + at + } + + /// A classic table for objects 0–3 followed by `after` (the caller writes the trailer). + fn table(&mut self, after: &str) -> usize { + let at = self.out.len(); + let mut text = "xref\n0 4\n0000000000 65535 f \n".to_string(); + for off in &self.offs { + let _ = write!(text, "{off:010} 00000 n \n"); + } + text.push_str(after); + self.out.extend_from_slice(text.as_bytes()); + at + } + + fn finish(mut self, start: usize) -> Vec { + self.out + .extend_from_slice(format!("startxref\n{start}\n%%EOF\n").as_bytes()); + self.out + } +} + +fn assert_refused(bytes: &[u8], code: &str, detail: &str, id: &str) { + let (got, details) = outcome(bytes); + assert_eq!(got, code, "{id}: {details}"); + assert!(details.contains(detail), "{id}: {details}"); +} + +/// A main xref stream with `pre`/`post` entries around the standard ones. +fn main_stream(pre: &str, post: &str) -> Vec { + let mut f = File::new(); + let x = f.xref_stream(9, pre, &format!("/Root 1 0 R {post}")); + f.finish(x) +} + +#[test] +fn snap08d_duplicate_keys_count_with_their_last_value() { + // lopdf's `Dictionary::set` keeps the last value; round 1 checked the first one. + // (entry written after the standard one, code, detail); written before it, the file opens. + // Round 1 passed both to lopdf: `/Size` was then looped over until the rows ran out, and + // `/W [1 4 9]` misread every row. + let cases = [ + ("/Size 3000000", "FILE_TOO_COMPLEX", "too many objects"), + ("/W [1 4 9]", "INVALID_PDF", "xref stream /W"), + ]; + for (post, code, detail) in cases { + assert_refused( + &main_stream("", post), + code, + detail, + &format!("SNAP-08d {post}"), + ); + let first = main_stream(post, ""); + assert_eq!( + opens(&first, "SNAP-08d").pages.len(), + 1, + "SNAP-08d first {post}" + ); + } + // Nested: lopdf reads `/Columns` from the `/DecodeParms` dictionary the same way. + let parms = |a: &str, b: &str| { + main_stream( + "", + &format!("/DecodeParms << /Predictor 12 /Columns {a} /Columns {b} >>"), + ) + }; + assert_refused( + &parms("1", "4611686018427387904"), + "FILE_TOO_COMPLEX", + "predictor row", + "SNAP-08d nested", + ); + assert_eq!(outcome(&parms("4611686018427387904", "1")).0, "OK"); + // `/Prev` twice in a table trailer: lopdf follows the last one (round 1: opened). + let prev = |first_bad: bool| { + let mut f = File::new(); + let good = f.xref_stream(10, "", ""); + let bad = f.xref_stream(11, "", PREDICTOR_BOMB); + let (a, b) = if first_bad { (bad, good) } else { (good, bad) }; + let t = f.table(&format!( + "trailer\n<< /Size 4 /Root 1 0 R /Prev {a} /Prev {b} >>\n" + )); + f.finish(t) + }; + assert_refused( + &prev(false), + "FILE_TOO_COMPLEX", + "predictor row", + "SNAP-08d /Prev", + ); + assert_eq!(outcome(&prev(true)).0, "OK", "SNAP-08d /Prev last good"); +} + +#[test] +fn snap08e_table_trailer_is_the_one_lopdf_parses() { + let fake = "%trailer << /Size 4 /Root 1 0 R >>\n"; + let with = |between: &str, trailer: &str| { + let mut f = File::new(); + let bad = f.xref_stream(10, "", PREDICTOR_BOMB); + let trailer = trailer.replace("{bad}", &bad.to_string()); + let t = f.table(&format!( + "{between}trailer\n<< /Size 4 /Root 1 0 R {trailer} >>\n" + )); + f.finish(t) + }; + // A `trailer` inside a comment after the entries hides the real one, whose `/Prev` lopdf + // follows (round 1: opened). + assert_refused( + &with(fake, "/Prev {bad}"), + "FILE_TOO_COMPLEX", + "predictor row", + "SNAP-08e hidden /Prev", + ); + // The real trailer's nesting is checked, not the comment's (round 1: opened). + assert_refused( + &with(fake, &format!("/Deep {}", deep(40))), + "FILE_TOO_COMPLEX", + "nesting", + "SNAP-08e hidden nesting", + ); + // Comments there are legal (lopdf's `space`): the file opens. + let fine = with(&format!("% producer note\n{fake}"), ""); + assert_eq!(opens(&fine, "SNAP-08e comments").pages.len(), 1); + // A `trailer` only inside a comment is no trailer for lopdf. + let mut f = File::new(); + let t = f.table(fake); + assert_refused( + &f.finish(t), + "INVALID_PDF", + "no trailer", + "SNAP-08e comment only", + ); +} + +#[test] +fn snap08f_xrefstm_of_an_xref_stream_trailer_is_followed() { + // lopdf takes `/XRefStm` from its newest trailer even when that is an xref stream (it reads + // it while following `/Prev`); round 1 followed it only from tables (opened). + let with = |stm_bad: bool| { + let mut f = File::new(); + let prev = f.xref_stream(10, "", ""); + let stm = f.xref_stream(11, "", if stm_bad { PREDICTOR_BOMB } else { "" }); + let x = f.xref_stream(12, "", &format!("/Root 1 0 R /Prev {prev} /XRefStm {stm}")); + f.finish(x) + }; + assert_refused(&with(true), "FILE_TOO_COMPLEX", "predictor row", "SNAP-08f"); + assert_eq!(opens(&with(false), "SNAP-08f good").pages.len(), 1); +} + +#[test] +fn snap08g_xref_stream_arrays_are_integers() { + // lopdf replaces an `/Index` holding a real by `[0 /Size]`, which round 1 never bounded + // (it summed the `/Index` counts): here lopdf looped towards `/Size` 3,000,000 until the rows + // ran out; with a long enough xref stream that is one entry per row byte. + assert_refused( + &main_stream("", "/Size 3000000 /Index [0 4.0]"), + "INVALID_PDF", + "xref stream /Index", + "SNAP-08g /Index", + ); + assert_refused( + &main_stream("", "/W [1 4.0 2]"), + "INVALID_PDF", + "xref stream /W", + "SNAP-08g /W", + ); + assert_eq!( + outcome(&main_stream("", "/Index [0 4]")).0, + "OK", + "SNAP-08g ints" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan.rs new file mode 100644 index 0000000..03f00ca --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan.rs @@ -0,0 +1,145 @@ +//! T4 tests (SPEC §E.4): PLAN-01…38 (`tests_plan/{plan,plan2,plan3,plan4}.rs`), VER-01…11 +//! (`ver.rs`), APP-01…09 (`app.rs`), MEAS-01 (`meas.rs`), the plan fuzz (§B.21, `fuzz.rs`) and +//! the named mobile-bug regressions B1–B3, B7–B10, B13–B16 (`bugs.rs`). Shared helpers live here. + +mod app; +mod bugs; +mod fuzz; +mod meas; +mod plan; +mod plan2; +mod plan3; +mod plan4; +mod ver; + +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::reasons::{EditProblemCode, Face}; +use crate::pdf_engine::text_edit::rewrite::{ + plan_page, EditVerdict, PagePlan, PlanOutcome, SourceTextStyleIn, TextEditIn, +}; +use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel, TextRun}; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use std::path::Path; + +pub(crate) use fuzz::thread_cpu; + +/// A context over `pdf` read by the production snapshot reader. +pub(crate) fn ctx(pdf: Vec) -> SnapshotContext { + let snap = snapshot_from_bytes(Path::new("plan-fixture.pdf"), pdf, None) + .unwrap_or_else(|e| panic!("fixture must open: {e} {:?}", e.details)); + SnapshotContext::new(snap) +} + +pub(crate) fn model(ctx: &SnapshotContext, page: u32) -> PageModel { + build_page_model(ctx, page, None).unwrap_or_else(|e| panic!("page model: {e}")) +} + +/// The run whose text is `text` (panics with the page's runs otherwise). +pub(crate) fn run_with<'m>(m: &'m PageModel, text: &str) -> &'m TextRun { + m.runs.iter().find(|r| r.text == text).unwrap_or_else(|| { + panic!( + "no run {text:?}; runs: {:?}; page_reason {:?} ({:?})", + m.runs + .iter() + .map(|r| (r.text.clone(), r.reason, r.reasons.clone())) + .collect::>(), + m.page_reason, + m.page_detail + ) + }) +} + +pub(crate) fn style() -> SourceTextStyleIn { + SourceTextStyleIn::default() +} + +pub(crate) fn sized(pt: f64) -> SourceTextStyleIn { + SourceTextStyleIn { + size_pt: Some(pt), + ..style() + } +} + +pub(crate) fn faced(face: Face) -> SourceTextStyleIn { + SourceTextStyleIn { + face: Some(face), + ..style() + } +} + +pub(crate) fn filled(hex: &str) -> SourceTextStyleIn { + SourceTextStyleIn { + fill: Some(hex.to_string()), + ..style() + } +} + +pub(crate) fn spaced(pt: f64) -> SourceTextStyleIn { + SourceTextStyleIn { + letter_spacing_pt: Some(pt), + ..style() + } +} + +/// An edit of the run whose current text is `old`. +pub(crate) fn edit(m: &PageModel, old: &str, new: &str, s: SourceTextStyleIn) -> TextEditIn { + TextEditIn { + run_id: run_with(m, old).id.clone(), + original_text: old.to_string(), + text: new.to_string(), + style: s, + } +} + +/// `plan_page`, which must not fail with an `AppError`. +pub(crate) fn plan(ctx: &SnapshotContext, m: &PageModel, edits: &[TextEditIn]) -> PlanOutcome { + plan_page(ctx, m, edits).unwrap_or_else(|e| panic!("plan_page: {e} {:?}", e.details)) +} + +/// One fixture page, one edit: (context, model, outcome). +pub(crate) fn plan_one( + pdf: Vec, + old: &str, + new: &str, + s: SourceTextStyleIn, +) -> (SnapshotContext, PageModel, PlanOutcome) { + let c = ctx(pdf); + let m = model(&c, 0); + let e = edit(&m, old, new, s); + let out = plan(&c, &m, &[e]); + (c, m, out) +} + +/// The plan of an outcome whose verdicts must all be ok. +pub(crate) fn ok_plan(out: &PlanOutcome) -> &PagePlan { + for v in &out.verdicts { + assert!(v.problem.is_none(), "verdict failed: {:?}", v.problem); + } + out.plan.as_ref().expect("a plan") +} + +/// The problem code of the only verdict. +pub(crate) fn problem_of(out: &PlanOutcome) -> EditProblemCode { + let v: &EditVerdict = out.verdicts.first().expect("one verdict"); + v.problem + .as_ref() + .unwrap_or_else(|| panic!("expected a problem, got ok: {v:?}")) + .code +} + +/// The replacement bytes of member `m` of the first planned run, as text. +pub(crate) fn replacement(out: &PlanOutcome, m: usize) -> String { + let plan = ok_plan(out); + let run = plan.runs.first().expect("a planned run"); + let s = run.splices.get(m).expect("member splice"); + String::from_utf8_lossy(&s.bytes).trim_start().to_string() +} + +/// The decoded page content after the plan (all parts joined with `\n`). +pub(crate) fn after_text(out: &PlanOutcome) -> String { + String::from_utf8_lossy(&ok_plan(out).expected_joined).into_owned() +} + +pub(crate) fn close(a: f64, b: f64, tol: f64) -> bool { + (a - b).abs() <= tol +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/app.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/app.rs new file mode 100644 index 0000000..3a696d9 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/app.rs @@ -0,0 +1,439 @@ +//! APP-01…09: the qpdf update (§B.13), qpdf's behaviours the gate relies on (P1, P4, P-RL, P-OV, +//! `/Extensions`), and the canonical graph digest (§B.16.1). + +use super::{ctx, edit, model, ok_plan, plan, style}; +use crate::pdf_engine::text_edit::apply::{ + apply_update, update_json, updates_for_plan, warning_text, write_update_json, PartUpdate, +}; +use crate::pdf_engine::text_edit::content::{page_content, qpdf_join}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget}; +use crate::pdf_engine::text_edit::engines::{Engines, RunOpts}; +use crate::pdf_engine::text_edit::graph::{graph_digest, graph_matches, GraphMismatch}; +use crate::pdf_engine::text_edit::limits::PAGE_DECODE_BUDGET; +use crate::pdf_engine::text_edit::snapshot::{fnv1a_u64, read_verification_snapshot}; +use crate::pdf_engine::text_edit::testkit::fakes; +use crate::pdf_engine::text_edit::testkit::pdf::PdfBuilder; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, DocBuilder, LegacyFilter, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use crate::pdf_engine::text_edit::tests_gate::{ed, fails_at, skipping, Honest}; +use lopdf::Object; +use std::collections::HashMap; +use std::path::Path; + +fn opts() -> RunOpts<'static> { + RunOpts::default() +} + +/// A verification context over a file qpdf wrote. +fn open(path: &Path) -> SnapshotContext { + SnapshotContext::new( + read_verification_snapshot(path, 1 << 30).unwrap_or_else(|e| panic!("{e}")), + ) +} + +fn edited( + engines: &Engines, + dir: &Scratch, + pdf: Vec, + old: &str, + new: &str, +) -> (std::path::PathBuf, std::path::PathBuf, Vec) { + let source = dir.write("source.pdf", &pdf); + let c = ctx(pdf); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, old, new, style())]); + let p = ok_plan(&out); + let update = dir.path("update.json"); + write_update_json( + &updates_for_plan(&m.content, p).expect("updates"), + c.doc().max_id, + &update, + ) + .expect("json"); + let staged = dir.path("staged.pdf"); + apply_update(engines, &source, &update, &staged, &[], &opts()) + .unwrap_or_else(|e| panic!("{e}")); + (source, staged, p.expected_parts[0].clone()) +} + +#[test] +fn app_01_update_json_shape() { + let v = update_json( + &[PartUpdate { + object_id: (4, 0), + decoded: b"abc".to_vec(), + }], + 9, + ); + let q = v["qpdf"].as_array().expect("qpdf array"); + assert_eq!(q.len(), 2, "APP-01 header + objects"); + assert_eq!(q[0]["jsonversion"], 2); + assert_eq!(q[0]["pushedinheritedpageresources"], false); + assert_eq!(q[0]["calledgetallpages"], false); + assert_eq!(q[0]["maxobjectid"], 9); + let obj = &q[1]["obj:4 0 R"]["stream"]; + assert_eq!( + obj["dict"], + serde_json::json!({}), + "APP-01 the empty dict drops /Filter" + ); + assert_eq!(obj["data"], "YWJj", "APP-01 base64 of the decoded bytes"); + assert_eq!(q[1].as_object().map(|o| o.len()), Some(1)); +} + +#[test] +fn app_02_p1_qpdf_applies_the_decoded_bytes() { + let Some(engines) = engines_or_skip("app_02") else { + return; + }; + let dir = Scratch::new("app_02"); + let (_, staged, expected) = edited( + &engines, + &dir, + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Help", + ); + let doc = fakes::load(&staged); + let page = fakes::page_ids(&doc)[0]; + let id = doc + .get_dictionary(page) + .and_then(|d| d.get(b"Contents")) + .and_then(Object::as_reference) + .expect("contents"); + let s = doc + .get_object(id) + .and_then(Object::as_stream) + .expect("stream"); + // `--compress-streams=n` (see apply.rs): the edited stream is written without a filter. + assert!( + s.dict.get(b"Filter").is_err(), + "APP-02 no filter: {:?}", + s.dict + ); + let data = + decode_stream(s, 1 << 20, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)).expect("decode"); + assert_eq!(data, expected, "APP-02 the planned bytes"); +} + +#[test] +fn app_03_non_benign_update_warnings_fail() { + let Some(engines) = engines_or_skip("app_03") else { + return; + }; + let dir = Scratch::new("app_03"); + let pdf = helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + let c = ctx(pdf.clone()); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Hello", "Help", style())]); + let update = dir.path("update.json"); + write_update_json( + &updates_for_plan(&m.content, ok_plan(&out)).expect("updates"), + c.doc().max_id, + &update, + ) + .expect("json"); + // A damaged xref: qpdf reconstructs it and exits 3. + let i = pdf + .windows(10) + .rposition(|w| w == b"startxref\n") + .expect("startxref"); + let mut damaged = pdf[..i].to_vec(); + damaged.extend_from_slice(b"startxref\n9\n%%EOF\n"); + let input = dir.write("damaged.pdf", &damaged); + let r = apply_update( + &engines, + &input, + &update, + &dir.path("out1.pdf"), + &[], + &opts(), + ); + let e = r.err().expect("APP-03 damaged input"); + assert_eq!(e.code, "EDIT_VERIFY_FAILED", "APP-03"); + // A wrong /Size: benign only when the source itself had that warning (`source_benign`). + let at = pdf.windows(6).rposition(|w| w == b"/Size ").expect("/Size") + 6; + let digits = pdf[at..].iter().take_while(|b| b.is_ascii_digit()).count(); + let mut wrong = pdf[..at].to_vec(); + wrong.extend_from_slice(b"99"); + wrong.extend_from_slice(&pdf[at + digits..]); + let input = dir.write("size.pdf", &wrong); + let r = apply_update( + &engines, + &input, + &update, + &dir.path("out2.pdf"), + &[], + &opts(), + ); + assert_eq!( + r.err().map(|e| e.code), + Some("EDIT_VERIFY_FAILED".to_string()), + "APP-03 not in source_benign" + ); + let benign = vec![format!( + "WARNING: /elsewhere/copy.pdf: reported number of objects (99) is not one plus the highest object number ({})", + c.doc().max_id + )]; + let lines = apply_update( + &engines, + &input, + &update, + &dir.path("out3.pdf"), + &benign, + &opts(), + ) + .unwrap_or_else(|e| panic!("APP-03 benign: {e} {:?}", e.details)); + assert_eq!(lines.len(), 1, "APP-03 the warning is returned"); + assert_eq!( + warning_text(&lines[0]), + warning_text(&benign[0]), + "APP-03 file name ignored" + ); +} + +#[test] +fn app_04_p4_mistargeted_update_caught_by_a0_and_a2() { + let Some(h) = Honest::new( + "app_04", + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + 0, + &[ed("Hello", "Help")], + ) else { + return; + }; + let out = h.path("mistargeted.pdf"); + fakes::mistargeted(h.qpdf(), &h.source, &out, 0, &h.plan.expected_parts[0]) + .unwrap_or_else(|e| panic!("{e}")); + fails_at( + h.phase_a(&out), + "EDIT_VERIFY_FAILED", + "A0", + "qpdf --check", + "APP-04 A0", + ); + let r = skipping(&["A0", "A1"], || h.phase_a(&out)); + fails_at(r, "EDIT_VERIFY_FAILED", "A2", "", "APP-04 A2"); +} + +#[test] +fn app_05_source_bytes_untouched() { + let Some(engines) = engines_or_skip("app_05") else { + return; + }; + let dir = Scratch::new("app_05"); + let pdf = fx::word(); + let before = fnv1a_u64(&pdf); + let (source, _, _) = edited(&engines, &dir, pdf, "Due", "Due 7"); + let after = std::fs::read(&source).expect("source"); + assert_eq!( + fnv1a_u64(&after), + before, + "APP-05 qpdf never writes its input" + ); +} + +#[test] +fn app_06_p_rl_legacy_filter_pages_kept_raw() { + for filter in [LegacyFilter::RunLength, LegacyFilter::Lzw] { + let test = format!("app_06_{filter:?}"); + let pdf = fx::legacy_filter_page(filter); + let Some(h) = Honest::new(&test, pdf.clone(), 1, &[ed("Hello", "Help")]) else { + return; + }; + // Honest passed A2: the legacy page's stream is byte-identical, still filtered. + let raw = |path: &Path| { + let doc = fakes::load(path); + let page = fakes::page_ids(&doc)[0]; + let id = doc + .get_dictionary(page) + .and_then(|d| d.get(b"Contents")) + .and_then(Object::as_reference) + .expect("contents"); + let s = doc + .get_object(id) + .and_then(Object::as_stream) + .expect("stream"); + ( + s.content.clone(), + s.dict + .get(b"Filter") + .and_then(Object::as_name) + .map(<[u8]>::to_vec) + .ok(), + ) + }; + assert_eq!( + raw(&h.staged), + raw(&h.source), + "APP-06 {filter:?} raw bytes kept" + ); + } +} + +#[test] +fn app_07_p_ov_overlay_wrapper_holds_the_joined_parts() { + let Some(engines) = engines_or_skip("app_07") else { + return; + }; + let dir = Scratch::new("app_07"); + let parts: [&[u8]; 3] = [ + b"BT /F1 12 Tf 72 700 Td", + b"(Hi) Tj ET\n", + b"BT /F1 12 Tf 72 650 Td (Lo) Tj ET", + ]; + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts(&parts, &format!("/Font << /F1 {f} 0 R >>"))); + let input = dir.write("parts.pdf", &d.build()); + let mut b = DocBuilder::new(); + b.page(PageSpec::new(b"", "")); + let blank = dir.write("blank.pdf", &b.build()); + let out = dir.path("overlaid.pdf"); + fakes::overlay(&engines.qpdf, &input, &out, &blank, None).unwrap_or_else(|e| panic!("{e}")); + let c = open(&out); + let page = c.page_id(0).expect("page"); + let content = + page_content(c.doc(), page, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)).expect("content"); + let text = String::from_utf8_lossy(&content.joined).into_owned(); + assert!(text.contains("/Fx0 Do"), "APP-07 wrapper: {text}"); + let fx0 = c + .doc() + .get_dictionary(page) + .and_then(|d| d.get(b"Resources")) + .and_then(Object::as_dict) + .and_then(|r| r.get(b"XObject")) + .and_then(Object::as_dict) + .and_then(|x| x.get(b"Fx0")) + .and_then(Object::as_reference) + .expect("/Fx0"); + let s = c + .doc() + .get_object(fx0) + .and_then(Object::as_stream) + .expect("Fx0"); + let data = + decode_stream(s, 1 << 20, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)).expect("decode"); + assert_eq!( + data, + qpdf_join(&parts), + "APP-07 Form data = qpdf_join(parts) on the installed qpdf" + ); +} + +#[test] +fn app_08_indirect_extensions_pass_a2() { + let Some(h) = Honest::new( + "app_08", + fx::extensions_indirect(), + 0, + &[ed("Hello", "Help")], + ) else { + return; + }; + let doc = fakes::load(&h.staged); + let cat = doc + .get_dictionary(fakes::catalog_id(&doc)) + .expect("catalog"); + assert!( + matches!(cat.get(b"Extensions"), Ok(Object::Dictionary(_))), + "APP-08 qpdf made /Extensions direct (and A2 passed in Honest::new)" + ); +} + +#[test] +fn app_09_graph_digest_and_matches() { + let Some(engines) = engines_or_skip("app_09") else { + return; + }; + let dir = Scratch::new("app_09"); + // A dangling reference in an array and an indirect /Info. + let mut b = PdfBuilder::new(); + let (cat, pages) = (b.alloc(), b.alloc()); + let font = b.add(HELVETICA); + let content = b.add_stream("", b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + let page = b.add(format!( + "<< /Type /Page /Parent {pages} 0 R /MediaBox [0 0 612 792] /Contents {content} 0 R \ + /Resources << /Font << /F1 {font} 0 R >> >> >>" + )); + b.set( + pages, + format!("<< /Type /Pages /Kids [{page} 0 R] /Count 1 >>"), + ); + b.set( + cat, + format!("<< /Type /Catalog /Pages {pages} 0 R /Dangling [1 999 0 R 2] >>"), + ); + let info = b.add("<< /Title (Graph) >>"); + let pdf = b.build(&format!("/Root {cat} 0 R /Info {info} 0 R")); + let input = dir.write("in.pdf", &pdf); + let out = dir.path("renumbered.pdf"); + let r = std::process::Command::new(&engines.qpdf) + .arg(&input) + .arg(&out) + .output() + .expect("qpdf"); + assert!( + r.status.success() || r.status.code() == Some(3), + "qpdf rewrite" + ); + let c = ctx(pdf); + let digest = graph_digest(c.doc(), &HashMap::new(), None).expect("digest"); + let after = open(&out); + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + assert_eq!( + graph_matches(after.doc(), &digest, &mut budget, None), + Ok(()), + "APP-09 renumbered + dangling" + ); + // A changed page dictionary: the first mismatch with its canonical path. + let changed = dir.path("changed.pdf"); + let doc = fakes::load(&out); + let page = fakes::page_ids(&doc)[0]; + let mut dict = fakes::json_dict(doc.get_dictionary(page).expect("page")); + dict.as_object_mut() + .expect("dict") + .insert("/Rotate".into(), serde_json::json!(90)); + fakes::set_value(&engines.qpdf, &out, &changed, page, dict).unwrap_or_else(|e| panic!("{e}")); + let m = graph_matches(open(&changed).doc(), &digest, &mut budget, None); + assert_eq!( + m, + Err(GraphMismatch { + index: m.as_ref().err().map_or(0, |x| x.index), + path: "/Root/Pages/Kids[0]".into(), + what: "node" + }), + "APP-09 path" + ); +} + +#[test] +fn app_09_whole_graph_holds_on_object_streams_inherited_resources_and_tags() { + // qpdf keeps object streams, inherited page attributes and the structure tree as they were: + // A2 passes honest edits on such files (each `Honest::new` runs Phase A). + let cases = [ + ("hybrid", fx::hybrid_xref(true), "Hello", "Help"), + ( + "inherited", + fx::shared_inherited_resources(), + "Regular words", + "Regular word", + ), + ( + "tagged", + fx::tagged_bookmarked_page(), + "Tagged line", + "Tagged lines", + ), + ]; + for (name, pdf, old, new) in cases { + let test = format!("app_09_{name}"); + let Some(h) = Honest::new(&test, pdf, 0, &[ed(old, new)]) else { + return; + }; + assert_eq!(h.report.proofs.len(), 1, "APP-09 {name}"); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/bugs.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/bugs.rs new file mode 100644 index 0000000..803f042 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/bugs.rs @@ -0,0 +1,247 @@ +//! The mobile Edit-text bugs (SPEC Appendix A) as named regressions on the desktop engine: +//! B1–B3, B7–B10, B14–B16 here; B13 (independent engines in the app) in `tests_independent.rs`. + +use super::{ + ctx, edit, faced, filled, model, ok_plan, plan, plan_one, problem_of, replacement, run_with, + sized, spaced, style, +}; +use crate::pdf_engine::text_edit::reasons::{ + EditProblemCode as P, Face, ProblemCtx, StyleField, TextReason, TextWarningCode, +}; +use crate::pdf_engine::text_edit::rewrite::{assemble_page_plan, RunPlan, TextEditIn}; +use crate::pdf_engine::text_edit::testkit::producers::{self as fx, helvetica_page}; +use crate::pdf_engine::text_edit::verify::{walk_and_verify, VerifyFailure}; + +#[test] +fn b1_tj_kerns_count_in_the_pen_and_word_gaps_read_as_spaces() { + // `[(AB) -500 (CD)] TJ (tail) Tj`: the −500 kern moves "CD" and "tail" (5 pt at 10 pt). + let c = ctx(fx::kerned_then_tail()); + let m = model(&c, 0); + let run = &m.runs[0]; + assert_eq!(run.text, "AB CDtail", "B1 a 0.5 em kern reads as one space"); + let out = plan(&c, &m, &[edit(&m, "AB CDtail", "AB CDEtail", style())]); + let p = ok_plan(&out); + let primary = &m.walk.records[run.members[0]]; + assert_eq!( + p.runs[0].expected.primary_pen_after, primary.pen_after, + "B1 pen after includes the kern" + ); + let after = p.expected_content(&m.content); + assert!( + walk_and_verify(&c, 0, &after, &m.walk, &m.runs, p, None).is_ok(), + "B1 the follower stays" + ); + // pdfTeX word gaps: a typed space is written as a gap and reads back as a space. + let (_, _, out) = plan_one(fx::pdftex(), "Hello World", "Hello Wet World", style()); + assert_eq!( + ok_plan(&out).runs[0].expected.text, + "Hello Wet World", + "B1 no \"HelloWorld\"" + ); +} + +#[test] +fn b2_size_is_in_effective_points_not_the_tf_operand() { + let (_, m, out) = plan_one(fx::tf1_tm12(), "Hi", "Hi", sized(14.0)); + assert_eq!(run_with(&m, "Hi").tfs, 1.0, "B2 the Tf operand is 1"); + let t = &ok_plan(&out).runs[0].target; + assert_eq!(t.tfs, 1.1667, "B2 Tf' = 1 × 14 / 12, never 14"); + assert!( + replacement(&out, 0).starts_with("/F1 1.1667 Tf ["), + "B2 {}", + replacement(&out, 0) + ); +} + +#[test] +fn b3_a_subset_never_claims_glyphs_it_lacks() { + let c = ctx(fx::subset_without_y()); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Hello", "Yellow", style())]); + let p = out.verdicts[0].problem.as_ref().expect("B3 GLYPH_MISSING"); + assert_eq!( + (p.code, p.chars.clone()), + (P::GlyphMissing, vec!['Y', 'w']), + "B3" + ); + // The width table has a slot for Y (0) but no glyph: never typeable. + let font = &m.walk.page_fonts[0].1; + assert!( + font.code_for('Y', &[], &[]).is_none(), + "B3 Y not in the alphabet" + ); +} + +#[test] +fn b7_no_op_toggles_write_nothing() { + let c = ctx(helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET")); + let m = model(&c, 0); + for s in [ + style(), + faced(Face::Regular), + sized(12.0), + spaced(0.0), + filled("#000000"), + ] { + let out = plan(&c, &m, &[edit(&m, "Hello", "Hello", s.clone())]); + assert!( + out.plan.is_none() && out.verdicts[0].problem.is_none(), + "B7 {s:?}" + ); + } +} + +#[test] +fn b8_face_problem_names_the_requested_face_and_checks_its_alphabet() { + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello", + faced(Face::Italic), + ); + let p = out.verdicts[0] + .problem + .clone() + .expect("B8 FACE_UNAVAILABLE"); + assert_eq!( + (p.code, p.face), + (P::FaceUnavailable, Some(Face::Italic)), + "B8" + ); + let ctx = ProblemCtx { + page_number: Some(1), + file_name: None, + face: None, + reason: None, + }; + let e = P::FaceUnavailable.to_app_error(&p, &ctx); + assert!( + e.message.contains("italic") && !e.message.contains("bold"), + "B8 says italic: {}", + e.message + ); + // The missing list is computed against the target face (PLAN-30 has the bold subset case). +} + +#[test] +fn b9_extgstate_font_size_is_unavailable_not_ignored() { + let (_, _, out) = plan_one(fx::extgstate_font(), "Hi there", "Hi there", sized(20.0)); + let p = out.verdicts[0].problem.as_ref().expect("B9"); + assert_eq!( + (p.code, p.field), + (P::StyleUnavailable, Some(StyleField::Size)), + "B9" + ); + assert!(out.plan.is_none(), "B9 nothing written"); +} + +#[test] +fn b10_a_rewrite_is_never_skipped_silently() { + // A show op that straddles two parts is refused, not skipped. + let c = ctx(fx::straddling_op()); + let m = model(&c, 0); + let refused = m.runs.iter().find(|r| r.text == "Hello").expect("run"); + assert_eq!( + refused.reason, + Some(TextReason::SplitContent), + "B10 refused up front" + ); + let e = TextEditIn { + run_id: refused.id.clone(), + original_text: "Hello".into(), + text: "Help".into(), + style: style(), + }; + let out = plan(&c, &m, &[e, edit(&m, "After", "Later", style())]); + assert_eq!(out.verdicts.len(), 2, "B10 one verdict per edit"); + assert_eq!( + out.verdicts[0].problem.as_ref().map(|p| p.code), + Some(P::TextEditRefused), + "B10" + ); + assert!( + out.plan.is_none(), + "B10 no partial plan when any edit fails" + ); + // Two edits of one line conflict: both verdicts say so. + let c = ctx(helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET")); + let m = model(&c, 0); + let e = edit(&m, "Hello", "Help", style()); + let out = plan(&c, &m, &[e.clone(), e]); + assert!( + out.verdicts + .iter() + .all(|v| v.problem.as_ref().map(|p| p.code) == Some(P::EditConflict)), + "B10 conflict" + ); +} + +#[test] +fn b14_restores_are_verbatim_never_rounded() { + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 0.123456 Tc /F1 12.34567 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello", + spaced(1.0), + ); + assert!( + replacement(&out, 0).ends_with("] TJ 0.123456 Tc"), + "B14 Tc: {}", + replacement(&out, 0) + ); + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12.34567 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello", + sized(14.0), + ); + assert!( + replacement(&out, 0).ends_with("] TJ /F1 12.34567 Tf"), + "B14 Tf: {}", + replacement(&out, 0) + ); +} + +#[test] +fn b15_paint_state_after_the_edit_is_compared() { + // A colour change whose restore is lost turns the later square red: refused (B15). + let c = ctx(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET 72 600 50 50 re f", + )); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Hello", "Hello", filled("#ff0000"))]); + let good = ok_plan(&out); + let mut runs: Vec = good.runs.clone(); + let s = &mut runs[0].splices[0]; + s.bytes = String::from_utf8_lossy(&s.bytes) + .replace("] TJ 0 g", "] TJ") + .into_bytes(); + let bad = assemble_page_plan(&m.content, 0, runs); + let after = bad.expected_content(&m.content); + let r = walk_and_verify(&c, 0, &after, &m.walk, &m.runs, &bad, None).err(); + assert_eq!( + r, + Some(VerifyFailure::PaintChanged { + index: 0, + field: "fill" + }), + "B15" + ); +} + +#[test] +fn b16_text_past_the_page_edge_is_blocked_and_overlap_is_a_warning() { + let pdf = || { + helvetica_page(b"BT /F1 12 Tf 500 700 Td (Edge) Tj ET BT /F1 12 Tf 72 680 Td (Left) Tj ET BT /F1 12 Tf 120 680 Td (Right) Tj ET") + }; + // 612 − 500 = 112 pt to the MediaBox edge; ten W = 113.28 pt. + let (_, _, out) = plan_one(pdf(), "Edge", &"W".repeat(10), style()); + assert_eq!(problem_of(&out), P::TextOutsideVisibleArea, "B16 page edge"); + let (_, _, out) = plan_one(pdf(), "Left", "Left side long", style()); + assert_eq!( + out.verdicts[0].warnings, + vec![TextWarningCode::NextTextOverlap], + "B16 overlap warns" + ); + assert!(out.plan.is_some(), "B16 overlap is not blocking"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/fuzz.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/fuzz.rs new file mode 100644 index 0000000..ae8d69b --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/fuzz.rs @@ -0,0 +1,256 @@ +//! Plan fuzz (§B.21, T4): 10,000 deterministic xorshift mutations — byte flips, truncations, +//! inserted `(`, `<`, `[`, `BI`, `q`, huge numbers, stray text-object operators — of fixture content +//! streams, each modelled and then planned with random edits (text, size, face, colour, spacing, +//! stale ids, duplicates). No panic; every `Err` is a request or self-check error; every plan +//! re-verifies; < 2 s of thread CPU per 1,000 cases. + +use super::ctx; +use crate::pdf_engine::text_edit::content::page_content; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::limits::PAGE_DECODE_BUDGET; +use crate::pdf_engine::text_edit::reasons::Face; +use crate::pdf_engine::text_edit::rewrite::{plan_page, SourceTextStyleIn, TextEditIn}; +use crate::pdf_engine::text_edit::runs::model_of; +use crate::pdf_engine::text_edit::testkit::producers::{ + cid_font, cid_hex, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::verify::walk_and_verify; +use crate::pdf_engine::text_edit::walker::{walk_page, WalkMode}; +use std::time::Duration; + +/// CPU time of the calling thread (wall time where the platform has no thread clock). +pub(crate) fn thread_cpu() -> Duration { + #[cfg(all( + any(target_os = "macos", target_os = "linux"), + target_pointer_width = "64" + ))] + { + #[repr(C)] + struct Timespec { + tv_sec: i64, + tv_nsec: i64, + } + extern "C" { + fn clock_gettime(clock: i32, tp: *mut Timespec) -> i32; + } + #[cfg(target_os = "macos")] + const CLOCK_THREAD_CPUTIME_ID: i32 = 16; + #[cfg(target_os = "linux")] + const CLOCK_THREAD_CPUTIME_ID: i32 = 3; + let mut ts = Timespec { + tv_sec: 0, + tv_nsec: 0, + }; + // SAFETY: clock_gettime only writes the timespec it is handed, which outlives the call. + if unsafe { clock_gettime(CLOCK_THREAD_CPUTIME_ID, &mut ts) } == 0 { + return Duration::new( + ts.tv_sec.max(0) as u64, + ts.tv_nsec.clamp(0, 999_999_999) as u32, + ); + } + } + static START: std::sync::OnceLock = std::sync::OnceLock::new(); + START.get_or_init(std::time::Instant::now).elapsed() +} + +struct XorShift(u64); + +impl XorShift { + fn next(&mut self) -> u64 { + let mut x = self.0; + x ^= x << 13; + x ^= x >> 7; + x ^= x << 17; + self.0 = x; + x + } + + fn below(&mut self, n: usize) -> usize { + (self.next() % n.max(1) as u64) as usize + } +} + +const CID_CHARS: &str = "Fuzy ğış"; + +fn fixture() -> Vec { + let mut d = DocBuilder::new(); + let f1 = d.add(HELVETICA); + let f2 = d.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica-Bold /Encoding /WinAnsiEncoding >>", + ); + let f3 = cid_font(&mut d.b, "ABCDEF+Arimo", CID_CHARS); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (seed) Tj ET", + &format!( + "/Font << /F1 {f1} 0 R /F2 {f2} 0 R /F3 {f3} 0 R >> /ExtGState << /GS1 << /Font [{f1} 0 R 9] >> >> \ + /ColorSpace << /Cs1 /DeviceRGB >>" + ), + )); + d.build() +} + +fn seeds() -> Vec> { + vec![ + b"BT /F1 12 Tf 72 700 Td [(Hel) -20 (lo) -333 (World)] TJ (tail) Tj ET BT /F2 12 Tf 72 680 Td (Bold) Tj ET" + .to_vec(), + format!( + "q 0.5 g BT /F3 10 Tf 2 Tc 3 Tw 90 Tz 1 Ts 72 650 Td <{}> Tj ET Q 0 0 1 rg 72 600 20 20 re f", + cid_hex(CID_CHARS, "Fuzy ğış Fuzy") + ) + .into_bytes(), + b"/P <> BDC BT /F1 12 Tf 14 TL 72 720 Td (One) ' 1 0.5 (Two) \" ET EMC \ + BT /F1 12 Tf 0.123 Tc /Cs1 cs 0.2 0.3 0.4 sc 72 560 Td (Coloured line) Tj ET" + .to_vec(), + b"/GS1 gs BT 72 500 Td (ExtGState font) Tj ET BT /F1 1 Tf 12 0 0 12 72 480 Tm (Tiny Tf) Tj ET \ + BT /F1 12 Tf 72 460 Td [(Name) -3000 (Value)] TJ ET" + .to_vec(), + ] +} + +fn mutate(base: &[u8], rng: &mut XorShift) -> Vec { + let mut v = base.to_vec(); + for _ in 0..rng.below(4) { + let at = rng.below(v.len() + 1); + match rng.below(9) { + 0 => { + if let Some(b) = v.get_mut(at) { + *b ^= 1 << rng.below(8); + } + } + 1 => v.truncate(at), + 2 => v.insert(at, b'('), + 3 => v.insert(at, b'<'), + 4 => v.insert(at, b'['), + 5 => v.splice(at..at, b" BI ".iter().copied()).for_each(drop), + 6 => v.splice(at..at, b" q ".iter().copied()).for_each(drop), + 7 => v + .splice(at..at, b" 999999999 -1e9 0.00001 ".iter().copied()) + .for_each(drop), + _ => v.splice(at..at, b" ET BT ".iter().copied()).for_each(drop), + } + } + v +} + +const SAFE: [&str; 9] = ["a", "e", "H", "l", "o", " ", "W", "\u{e9}", "7"]; +const UNSAFE: [&str; 5] = ["ğ", "ı", "\n", "\u{1F600}", " "]; + +fn text(rng: &mut XorShift, base: &str) -> String { + match rng.below(5) { + 0 => base.to_string(), + 1 => String::new(), + 2 => base + .chars() + .take(rng.below(base.chars().count() + 1)) + .collect(), + _ => { + let mut s: String = base + .chars() + .take(rng.below(base.chars().count() + 1)) + .collect(); + for _ in 0..rng.below(6) { + s.push_str(SAFE[rng.below(SAFE.len())]); + } + if rng.below(10) == 0 { + s.push_str(UNSAFE[rng.below(UNSAFE.len())]); + } + s + } + } +} + +/// A random style; one in eight is malformed (BAD_EDIT), the others are in range. +fn style(rng: &mut XorShift) -> SourceTextStyleIn { + let faces = [Face::Regular, Face::Bold, Face::Italic, Face::BoldItalic]; + if rng.below(8) == 0 { + return match rng.below(3) { + 0 => SourceTextStyleIn { + size_pt: Some([3.0, 200.0, f64::NAN][rng.below(3)]), + ..Default::default() + }, + 1 => SourceTextStyleIn { + fill: Some("nope".into()), + ..Default::default() + }, + _ => SourceTextStyleIn { + letter_spacing_pt: Some([-3.0, 11.0][rng.below(2)]), + ..Default::default() + }, + }; + } + SourceTextStyleIn { + size_pt: (rng.below(4) == 0).then(|| [4.0, 12.0, 13.5, 144.0][rng.below(4)]), + face: (rng.below(4) == 0).then(|| faces[rng.below(4)]), + fill: (rng.below(4) == 0) + .then(|| ["#ff0000", "#000000", "#c71c1c"][rng.below(3)].to_string()), + letter_spacing_pt: (rng.below(4) == 0).then(|| [-2.0, 0.0, 0.7, 10.0][rng.below(4)]), + } +} + +#[test] +fn plan_fuzz() { + let c = ctx(fixture()); + let page = c.page_id(0).expect("page"); + let base = + page_content(c.doc(), page, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)).expect("content"); + let seeds = seeds(); + let mut rng = XorShift(0xD1B5_4A32_D192_ED03); + let (mut planned, mut refused) = (0usize, 0usize); + for chunk in 0..10 { + let started = thread_cpu(); + for i in 0..1_000 { + let seed = &seeds[(chunk * 1_000 + i) % seeds.len()]; + let content = base.with_replaced_parts(&[(0, mutate(seed, &mut rng))]); + let walk = walk_page(&c, 0, &content, WalkMode::Edit, None); + let m = model_of(&c, 0, content, walk); + let mut edits = Vec::new(); + for _ in 0..1 + usize::from(rng.below(5) == 0) { + let (id, old) = match m.runs.get(rng.below(m.runs.len() + 1)) { + Some(r) => (r.id.clone(), r.text.clone()), + None => ("t1:stale".to_string(), "stale".to_string()), + }; + edits.push(TextEditIn { + run_id: id, + original_text: old.clone(), + text: text(&mut rng, &old), + style: style(&mut rng), + }); + } + match plan_page(&c, &m, &edits) { + Ok(out) => { + assert_eq!(out.verdicts.len(), edits.len(), "one verdict per edit"); + if let Some(p) = out.plan { + planned += 1; + let after = p.expected_content(&m.content); + assert!( + walk_and_verify(&c, 0, &after, &m.walk, &m.runs, &p, None).is_ok(), + "every returned plan re-verifies" + ); + } else { + refused += 1; + } + } + Err(e) => { + assert!( + matches!( + e.code.as_str(), + "BAD_EDIT" | "TOO_MANY_TEXT_EDITS" | "EDIT_VERIFY_FAILED" + ), + "unexpected error {}", + e.code + ); + } + } + } + let spent = thread_cpu().saturating_sub(started); + assert!( + spent < Duration::from_secs(2), + "plan fuzz chunk {chunk}: {spent:?} for 1,000 cases" + ); + } + println!("plan fuzz: {planned} planned, {refused} refused or no-ops"); + assert!( + planned > 1_000 && refused > 1_000, + "both outcomes reached: {planned} planned, {refused} refused" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/meas.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/meas.rs new file mode 100644 index 0000000..15c7055 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/meas.rs @@ -0,0 +1,123 @@ +//! MEAS-01: the editor's width estimate (`fit.rs`) against the golden file the frontend's +//! `estimateDeltaPt`/`estimateCaretOffsets` are tested with (T7). `OFFPDF_UPDATE_GOLDEN=1` +//! rewrites `src/lib/editor/__fixtures__/text-measure-golden.json`; without it the file must equal +//! what the code computes now. + +use super::{ctx, edit, model, plan, run_with, sized, style}; +use crate::pdf_engine::text_edit::fit::{ + estimate_caret_offsets, estimate_delta_pt, golden, write_measure_golden, +}; +use crate::pdf_engine::text_edit::testkit::producers::{self as fx, helvetica_page}; +use std::path::PathBuf; + +fn golden_path() -> PathBuf { + PathBuf::from(env!("CARGO_MANIFEST_DIR")) + .join("../src/lib/editor/__fixtures__/text-measure-golden.json") +} + +#[test] +fn meas_01_estimate_golden() { + let path = golden_path(); + if std::env::var("OFFPDF_UPDATE_GOLDEN").as_deref() == Ok("1") { + if let Some(dir) = path.parent() { + std::fs::create_dir_all(dir).expect("fixtures dir"); + } + write_measure_golden(&path).expect("write the golden file"); + } + let text = std::fs::read_to_string(&path).unwrap_or_else(|e| { + panic!( + "MEAS-01 {}: {e} (run with OFFPDF_UPDATE_GOLDEN=1)", + path.display() + ) + }); + let stored: serde_json::Value = serde_json::from_str(&text).expect("MEAS-01 JSON"); + // serde_json reads floats to within an ulp (no `float_roundtrip` feature): compare numbers + // within 1e-9, everything else exactly. + if let Err(at) = same_json(&stored, &golden::build(), "$") { + panic!("MEAS-01 the golden file is stale at {at}; regenerate it (OFFPDF_UPDATE_GOLDEN=1)"); + } + let cases = stored["cases"].as_array().expect("cases"); + let names: Vec<&str> = cases.iter().filter_map(|c| c["name"].as_str()).collect(); + for want in [ + "std14-helvetica-12", + "tz-80", + "tc-0.5", + "kern-space", + "sibling-surface", + "size-change", + "letter-spacing", + "face-bold", + ] { + assert!(names.contains(&want), "MEAS-01 case {want}"); + } + for c in cases { + let offsets = c["caretOffsets"].as_array().expect("offsets"); + let chars = c["text"].as_str().expect("text").chars().count(); + assert_eq!(offsets.len(), chars + 1, "MEAS-01 {} chars + 1", c["name"]); + } +} + +#[test] +fn meas_01_estimate_matches_the_plan_without_kerns() { + // Helvetica 12, no kerns, no Tc: the estimate equals the planned width change (with Tc the + // estimate also counts the spacing after the last glyph, which the ink extent does not). + let c = ctx(helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET")); + let m = model(&c, 0); + let run = run_with(&m, "Hello"); + for (text, s) in [ + ("Help me", style()), + ("Hello", sized(14.0)), + ("Hel", style()), + ] { + let est = estimate_delta_pt(run, &m.walk.page_fonts, text, &s); + let out = plan(&c, &m, &[edit(&m, "Hello", text, s)]); + let planned = out.verdicts[0].delta_pt; + assert!( + (est - planned).abs() < 1e-9, + "MEAS-01 {text:?}: estimate {est} vs plan {planned}" + ); + } + let offsets = estimate_caret_offsets(run, &m.walk.page_fonts, "Hi", &style()); + // H = 722, i = 222 (×12/1000). + assert_eq!(offsets.len(), 3); + assert!( + (offsets[1] - 8.664).abs() < 1e-9 && (offsets[2] - 11.328).abs() < 1e-9, + "{offsets:?}" + ); + // Kern mode: a typed space is the run's median gap (−333 ⇒ 0.333 em). + let c = ctx(fx::pdftex()); + let m = model(&c, 0); + let run = run_with(&m, "Hello World"); + let d = estimate_delta_pt(run, &m.walk.page_fonts, "Hello World", &style()); + assert!((d - 0.333 * 9.9626).abs() < 1e-6, "MEAS-01 kern space {d}"); +} + +fn same_json(a: &serde_json::Value, b: &serde_json::Value, at: &str) -> Result<(), String> { + use serde_json::Value as V; + match (a, b) { + (V::Number(x), V::Number(y)) => { + let (x, y) = ( + x.as_f64().unwrap_or(f64::NAN), + y.as_f64().unwrap_or(f64::NAN), + ); + if (x - y).abs() <= 1e-9 * 1f64.max(x.abs()) { + Ok(()) + } else { + Err(format!("{at}: {x} vs {y}")) + } + } + (V::Array(x), V::Array(y)) if x.len() == y.len() => x + .iter() + .zip(y) + .enumerate() + .try_for_each(|(i, (p, q))| same_json(p, q, &format!("{at}[{i}]"))), + (V::Object(x), V::Object(y)) if x.len() == y.len() => { + x.iter().try_for_each(|(k, p)| match y.get(k) { + Some(q) => same_json(p, q, &format!("{at}.{k}")), + None => Err(format!("{at}.{k} missing")), + }) + } + _ if a == b => Ok(()), + _ => Err(format!("{at}: {a} vs {b}")), + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/plan.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/plan.rs new file mode 100644 index 0000000..53916f7 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/plan.rs @@ -0,0 +1,259 @@ +//! PLAN-01…10: the minimal diff, boundary kerns, hex middles, the §B.12 worked example, the +//! compensation, `'` and `"` prefixes, Tz/Tc/Tw and the B2 size rule. + +use super::{after_text, close, ok_plan, plan_one, replacement, run_with, sized, style}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, word_font, DocBuilder, PageSpec, +}; + +/// A page drawing `content` with FX-WORD's WinAnsi TrueType subset of `chars` as `/F1`. +pub(crate) fn word_page(chars: &str, content: &str) -> Vec { + let mut d = DocBuilder::new(); + let f1 = word_font(&mut d.b, "ABCDEF+Calibri", chars); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +#[test] +fn plan_01_minimal_diff_keeps_prefix_and_suffix_codes_and_kerns() { + // A V (−80) A T (20) x y z: "y" → "Q" keeps every other code and both kerns byte-exact. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td [(AV)-80.5(AT)20(xyz)]TJ ET"), + "AVATxyz", + "AVATxQz", + style(), + ); + // Q (778) replaces y (500): the pen is pulled back by 278 thousandths (a positive number). + assert_eq!( + replacement(&out, 0), + "[<4156> -80.5 <4154> 20 <78517A> 278] TJ", + "PLAN-01" + ); + let run = &ok_plan(&out).runs[0]; + assert_eq!(run.expected.prefix_glyphs, 5, "PLAN-01 prefix A V A T x"); + assert_eq!(run.expected.suffix_glyphs, 1, "PLAN-01 suffix z"); +} + +#[test] +fn plan_02_boundary_kerns_dropped_only_when_not_synthetic() { + // The pair kern 30 between "o" and "W" no longer has its pair once "W" changes: dropped. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td [(To)30(Wn)]TJ ET"), + "ToWn", + "ToMn", + style(), + ); + let r = replacement(&out, 0); + assert!(!r.contains(" 30 "), "PLAN-02 pair kern dropped: {r}"); + assert!(r.starts_with("[<546F4D6E>"), "PLAN-02 codes: {r}"); + // A synthetic space (−333 ≥ 0.2 em) at the boundary is part of the kept text: kept. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td [(Hello)-333(World)]TJ ET"), + "Hello World", + "Hello Earth", + style(), + ); + let r = replacement(&out, 0); + assert!( + r.starts_with("[<48656C6C6F> -333 <"), + "PLAN-02 synthetic kept: {r}" + ); +} + +#[test] +fn plan_03_middle_is_uppercase_hex() { + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (abc) Tj ET"), + "abc", + "a\u{e9}c", + style(), + ); + // é is WinAnsi 0xE9; a Tj becomes a TJ. + let r = replacement(&out, 0); + assert!(r.starts_with("[<61E963>"), "PLAN-03 {r}"); + assert!(r.ends_with("] TJ"), "PLAN-03 {r}"); +} + +#[test] +fn plan_04_worked_example_byte_for_byte() { + // §B.12: Word-style WinAnsi TrueType, "Invoice 2026" → "Invoice 2027". + let content = "BT /F1 11.04 Tf 1 0 0 1 72 700 Tm [(Inv)12(oice 2026)]TJ ET"; + let (_, _, out) = plan_one( + word_page("Invoice 20267", content), + "Invoice 2026", + "Invoice 2027", + style(), + ); + assert_eq!( + replacement(&out, 0), + "[<496E76> 12 <6F6963652032303237>] TJ", + "PLAN-04" + ); + assert_eq!( + after_text(&out), + "BT /F1 11.04 Tf 1 0 0 1 72 700 Tm [<496E76> 12 <6F6963652032303237>] TJ ET", + "PLAN-04 page" + ); + assert!( + close(out.verdicts[0].delta_pt, 0.0, 1e-9), + "PLAN-04 same width" + ); +} + +#[test] +fn plan_05_compensation_sign_shorter_and_longer() { + let pdf = || { + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 12 Tf 72 680 Td (tail) Tj ET") + }; + // Shorter: the pen must be pushed forward (a negative TJ number). + let (_, m, out) = plan_one(pdf(), "Hello", "Hell", style()); + let r = replacement(&out, 0); + assert!(r.ends_with(" -556] TJ"), "PLAN-05 shorter: {r}"); + assert!(out.verdicts[0].delta_pt < 0.0, "PLAN-05 shorter delta"); + let old_width = run_with(&m, "Hello").rect[2]; + // (kept glyphs, newly encoded glyphs) of the planned run. + let kept = |out: &crate::pdf_engine::text_edit::rewrite::PlanOutcome| { + let run = &ok_plan(out).runs[0]; + let new = run.expected.glyph_new.iter().filter(|n| **n).count(); + (run.expected.glyph_new.len() - new, new) + }; + assert_eq!( + kept(&out), + (4, 0), + "PLAN-05 shorter: H e l l kept, nothing new" + ); + verdict_matches_plan(&out, 4, |w| w < old_width, "shorter"); + // Longer: pulled back (positive). + let (_, _, out) = plan_one(pdf(), "Hello", "Helloo", style()); + let r = replacement(&out, 0); + assert!(r.ends_with(" 556] TJ"), "PLAN-05 longer: {r}"); + assert!(out.verdicts[0].delta_pt > 0.0, "PLAN-05 longer delta"); + assert_eq!(kept(&out), (5, 1), "PLAN-05 longer: Hello kept, one new o"); + verdict_matches_plan(&out, 6, |w| w > old_width, "longer"); +} + +/// The verdict of the only edit agrees with its run plan: same width change, the planned new +/// box (`width_ok` on its width), `chars + 1` monotonic caret offsets ending at the new advance, +/// and the page plan is for page 0. +fn verdict_matches_plan( + out: &crate::pdf_engine::text_edit::rewrite::PlanOutcome, + chars: usize, + width_ok: impl Fn(f64) -> bool, + case: &str, +) { + let plan = ok_plan(out); + let (v, run) = (&out.verdicts[0], &plan.runs[0]); + assert_eq!(plan.page_index, 0, "PLAN-05 {case} page"); + assert_eq!(v.new_rect, Some(run.new_rect), "PLAN-05 {case} rect"); + assert!( + width_ok(run.new_rect[2]), + "PLAN-05 {case} width {:?}", + run.new_rect + ); + let carets = v.caret_offsets.as_ref().expect("caret offsets"); + assert_eq!(carets.len(), chars + 1, "PLAN-05 {case} carets {carets:?}"); + assert!( + carets.windows(2).all(|w| w[0] <= w[1]) && carets[0] == 0.0, + "PLAN-05 {case} carets monotonic from 0: {carets:?}" + ); +} + +#[test] +fn plan_06_equal_width_writes_no_compensation() { + // Digits share one width in Helvetica. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Total 2026) Tj ET"), + "Total 2026", + "Total 2027", + style(), + ); + assert_eq!( + replacement(&out, 0), + "[<546F74616C2032303237>] TJ", + "PLAN-06" + ); +} + +#[test] +fn plan_07_quote_keeps_its_line_move() { + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 14 TL 72 700 Td (First) Tj (Second)' ET"), + "Second", + "Seconds", + style(), + ); + let r = replacement(&out, 0); + assert!(r.starts_with("T* [<5365636F6E6473>"), "PLAN-07 {r}"); +} + +#[test] +fn plan_08_double_quote_keeps_aw_ac_verbatim() { + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 14 TL 72 700 Td (First) Tj 2.50 0.125 (Second)\" ET"), + "Second", + "Secant", + style(), + ); + let r = replacement(&out, 0); + assert!(r.starts_with("2.50 Tw 0.125 Tc T* [<"), "PLAN-08 {r}"); +} + +#[test] +fn plan_09_tz_tc_tw_and_the_word_space_code() { + // Tz 50, Tc 1, Tw 5: a WinAnsi space (code 32) takes Tw. One more space glyph = + // 0.278·12 + Tc 1 + Tw 5 = 9.336 text units (Th excluded) = 778 thousandths of Tf 12. + let (_, _, out) = plan_one( + helvetica_page( + b"BT /F1 12 Tf 50 Tz 1 Tc 5 Tw 72 700 Td (a b) Tj ET BT /F1 12 Tf 72 680 Td (end) Tj ET", + ), + "a b", + "a b", + style(), + ); + assert_eq!( + replacement(&out, 0), + "[<61202062> 778] TJ", + "PLAN-09 code 32 gets Tw" + ); + // A 2-byte space (Identity-H, CID 3) is never a word space: 0.5·12 + Tc 1 = 7 text units. + let mut d = DocBuilder::new(); + let f1 = fx::cid_font(&mut d.b, "AAAAAA+Arimo", "ab "); + let content = format!( + "BT /F1 12 Tf 50 Tz 1 Tc 5 Tw 72 700 Td <{}> Tj ET", + fx::cid_hex("ab ", "a b") + ); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f1} 0 R >>"), + )); + let (_, _, out) = plan_one(d.build(), "a b", "a b", style()); + assert_eq!( + replacement(&out, 0), + "[<00010003000300020> 583.3333] TJ".replace("00020>", "0002>"), + "PLAN-09 2-byte space gets no Tw" + ); +} + +#[test] +fn plan_10_b2_tf1_tm12_size_plus_one_is_effective_13() { + // `/F1 1 Tf 12 0 0 12 x y Tm`: effective 12, Tf operand 1. +1 pt ⇒ Tf' = 1.0833. + let (_, m, out) = plan_one(fx::tf1_tm12(), "Hi", "Hi", sized(13.0)); + let run = run_with(&m, "Hi"); + assert!( + close(run.effective_size, 12.0, 1e-9), + "PLAN-10 effective 12" + ); + let plan = ok_plan(&out); + let t = &plan.runs[0].target; + assert!(close(t.tfs, 1.0833, 1e-12), "PLAN-10 Tf' {}", t.tfs); + let r = replacement(&out, 0); + assert!(r.contains("/F1 1.0833 Tf ["), "PLAN-10 sets Tf: {r}"); + assert!(r.ends_with("/F1 1 Tf"), "PLAN-10 restores Tf verbatim: {r}"); + assert!( + close(plan.runs[0].expected.effective_size, 13.0, 1e-9), + "PLAN-10 expected 13" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/plan2.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/plan2.rs new file mode 100644 index 0000000..949834d --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/plan2.rs @@ -0,0 +1,336 @@ +//! PLAN-11…21: no-ops (B7), removal, kern-space writing, typing problems, the fit policy (B16), +//! conflicts and refusals, and several edits on one page. + +use super::{ + after_text, ctx, edit, faced, filled, model, ok_plan, plan, plan_one, problem_of, replacement, + run_with, sized, spaced, style, +}; +use crate::pdf_engine::text_edit::reasons::{ + EditProblemCode as P, Face, TextReason, TextWarningCode, +}; +use crate::pdf_engine::text_edit::rewrite::{plan_page, SourceTextStyleIn, TextEditIn}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; + +#[test] +fn plan_11_b7_no_ops_write_nothing() { + let pdf = || helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + let c = ctx(pdf()); + let m = model(&c, 0); + let cases: Vec<(&str, SourceTextStyleIn)> = vec![ + ("text equal, style empty", style()), + ("bold off (already regular)", faced(Face::Regular)), + ("size back to the current 12 pt", sized(12.0)), + ("size within STYLE_EPSILON", sized(12.0004)), + ("colour equal to the original", filled("#000000")), + ("letter spacing equal to the current 0", spaced(0.0)), + ]; + for (what, s) in cases { + let out = plan(&c, &m, &[edit(&m, "Hello", "Hello", s)]); + let v = &out.verdicts[0]; + assert!(v.problem.is_none(), "PLAN-11 {what}: {:?}", v.problem); + assert!(out.plan.is_none(), "PLAN-11 {what}: no plan, no bytes"); + assert_eq!(v.delta_pt, 0.0, "PLAN-11 {what}"); + } +} + +#[test] +fn plan_12_empty_text_removes_the_glyphs_and_keeps_the_pen() { + // Hello = 722 + 556 + 222 + 222 + 556 = 2278 thousandths: pushed forward by that much. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj (World) Tj ET"), + "HelloWorld", + "", + style(), + ); + assert_eq!(replacement(&out, 0), "[<> -2278] TJ", "PLAN-12 primary"); + // World = 944 + 556 + 333 + 222 + 556: the absorbed member keeps its own travel. + assert_eq!( + replacement(&out, 1), + "[<> -2611] TJ", + "PLAN-12 absorbed member" + ); + let run = &ok_plan(&out).runs[0]; + assert!( + run.expected.glyphs.is_empty() && run.expected.text.is_empty(), + "PLAN-12" + ); +} + +#[test] +fn plan_13_kern_space_writing_and_space_not_writable() { + // FX-PDFTEX: no space glyph, word gaps are −333 kerns ⇒ kern mode. + let c = ctx(fx::pdftex()); + let m = model(&c, 0); + let run = run_with(&m, "Hello World"); + assert_eq!( + run.space_mode, + crate::pdf_engine::text_edit::runs::SpaceMode::Kern + ); + let out = plan( + &c, + &m, + &[edit(&m, "Hello World", "Hello Wet World", style())], + ); + let r = replacement(&out, 0); + // Kept "Hello", its −333 gap and "W"; new "et", the typed space as the run's median gap + // (−333) and "W"; the kept "orld" (the pair kern 80 after the old "W" is dropped). + assert!( + r.starts_with("[<48656C6C6F> -333 <576574> -333 <576F726C64> "), + "PLAN-13 kern space: {r}" + ); + assert_eq!( + ok_plan(&out).runs[0].expected.text, + "Hello Wet World", + "PLAN-13 reads back" + ); + // " World" and "Hello " keep the −333 gap at an end of the line, where it is no longer + // between two glyphs and would not read as a space (review round 1, L-1): refused like a + // typed space, not as an internal verification failure. + for bad in [ + "Hello World", + " Hello World", + "Hello World ", + "Hello W orld", + " World", + "Hello ", + ] { + let out = plan(&c, &m, &[edit(&m, "Hello World", bad, style())]); + assert_eq!(problem_of(&out), P::SpaceNotWritable, "PLAN-13 {bad:?}"); + } + for good in ["World", "Hello"] { + let out = plan(&c, &m, &[edit(&m, "Hello World", good, style())]); + assert_eq!( + ok_plan(&out).runs[0].expected.text, + good, + "PLAN-13 {good:?}" + ); + } +} + +#[test] +fn plan_14_glyph_missing_dedup_in_typing_order() { + // A Word subset of "Hello" (B3): Y, a and y have no glyph. + let (_, _, out) = plan_one(fx::subset_without_y(), "Hello", "Yay yo Hey", style()); + let v = &out.verdicts[0]; + let p = v.problem.as_ref().expect("GLYPH_MISSING"); + assert_eq!(p.code, P::GlyphMissing, "PLAN-14"); + assert_eq!(p.chars, vec!['Y', 'a', 'y'], "PLAN-14 dedup, typing order"); + assert!(out.plan.is_none(), "PLAN-14 nothing planned"); +} + +#[test] +fn plan_15_invalid_text() { + let pdf = || helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + for bad in [ + "Hel\nlo", + "Hel\tlo", + "Hel\u{2028}lo", + "Hel\u{7}lo", + "Hel\u{85}lo", + "Hel\rlo", + ] { + let (_, _, out) = plan_one(pdf(), "Hello", bad, style()); + assert_eq!(problem_of(&out), P::InvalidText, "PLAN-15 {bad:?}"); + } +} + +#[test] +fn plan_16_text_too_long_at_1001() { + let pdf = || helvetica_page(b"BT /F1 4 Tf 1 700 Td (i) Tj ET"); + let (_, _, out) = plan_one(pdf(), "i", &"i".repeat(1001), style()); + assert_eq!(problem_of(&out), P::TextTooLong, "PLAN-16 1,001"); + let (_, _, out) = plan_one(pdf(), "i", &"i".repeat(1000), style()); + let v = &out.verdicts[0]; + assert_ne!( + v.problem.as_ref().map(|p| p.code), + Some(P::TextTooLong), + "PLAN-16 1,000" + ); +} + +#[test] +fn plan_17_b16_text_outside_visible_area_crop_edge_and_clip() { + // CropBox [36 48 576 744], line at x 72: 504 pt to the crop edge. W = 11.328 pt at 12 pt. + let crop = || fx::cropped_offset(); + let (_, _, out) = plan_one(crop(), "Cropped page", &"W".repeat(45), style()); + assert_eq!( + problem_of(&out), + P::TextOutsideVisibleArea, + "PLAN-17 crop edge" + ); + let (_, _, out) = plan_one(crop(), "Cropped page", &"W".repeat(44), style()); + assert!( + out.verdicts[0].problem.is_none(), + "PLAN-17 fits: {:?}", + out.verdicts[0].problem + ); + // A clip rectangle 128 pt past the origin. + let clip = || helvetica_page(b"q 0 0 200 792 re W n BT /F1 12 Tf 72 720 Td (Clip me) Tj ET Q"); + let (_, _, out) = plan_one(clip(), "Clip me", &"W".repeat(12), style()); + assert_eq!( + problem_of(&out), + P::TextOutsideVisibleArea, + "PLAN-17 clip rect" + ); + let (_, _, out) = plan_one(clip(), "Clip me", &"W".repeat(11), style()); + assert!(out.verdicts[0].problem.is_none(), "PLAN-17 clip fits"); +} + +#[test] +fn plan_18_overlap_is_a_warning() { + let pdf = helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Left) Tj ET BT /F1 12 Tf 200 700 Td (Right) Tj ET", + ); + let (_, _, out) = plan_one(pdf, "Left", &"W".repeat(12), style()); + let v = &out.verdicts[0]; + assert!(v.problem.is_none(), "PLAN-18 not blocking: {:?}", v.problem); + assert_eq!( + v.warnings, + vec![TextWarningCode::NextTextOverlap], + "PLAN-18" + ); + assert!(out.plan.is_some(), "PLAN-18 still planned"); +} + +#[test] +fn plan_19_conflicts_stale_refused_and_bad_style() { + let c = ctx(fx::duplicate_shadow()); + let m = model(&c, 0); + let refused = m.runs.first().expect("a run"); + assert_eq!(refused.reason, Some(TextReason::DuplicateText)); + let e = TextEditIn { + run_id: refused.id.clone(), + original_text: refused.text.clone(), + text: "Other".into(), + style: style(), + }; + let out = plan(&c, &m, &[e]); + let p = out.verdicts[0].problem.as_ref().expect("refused"); + assert_eq!( + (p.code, p.reason), + (P::TextEditRefused, Some(TextReason::DuplicateText)), + "PLAN-19 refused" + ); + + let c = ctx(helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET")); + let m = model(&c, 0); + let e = edit(&m, "Hello", "Help", style()); + let out = plan(&c, &m, &[e.clone(), e.clone()]); + assert!(out.plan.is_none(), "PLAN-19 duplicate run id"); + for v in &out.verdicts { + assert_eq!( + v.problem.as_ref().map(|p| p.code), + Some(P::EditConflict), + "PLAN-19 conflict" + ); + } + let stale_text = TextEditIn { + original_text: "Hullo".into(), + ..e.clone() + }; + assert_eq!( + problem_of(&plan(&c, &m, &[stale_text])), + P::Stale, + "PLAN-19 text changed" + ); + let stale_id = TextEditIn { + run_id: "t1:0-0:0:1-2".into(), + ..e.clone() + }; + assert_eq!( + problem_of(&plan(&c, &m, &[stale_id])), + P::Stale, + "PLAN-19 unknown run" + ); + for bad in [ + sized(3.9), + sized(f64::NAN), + spaced(10.5), + filled("red"), + filled("#12345"), + ] { + let err = plan_page( + &c, + &m, + &[TextEditIn { + style: bad.clone(), + ..e.clone() + }], + ) + .err() + .unwrap_or_else(|| panic!("PLAN-19 {bad:?} must be BAD_EDIT")); + assert_eq!(err.code, "BAD_EDIT", "PLAN-19 {bad:?}"); + } +} + +#[test] +fn plan_20_two_edits_in_one_part() { + let c = ctx(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (First line) Tj ET BT /F1 12 Tf 72 680 Td (Second line) Tj ET", + )); + let m = model(&c, 0); + let out = plan( + &c, + &m, + &[ + edit(&m, "Second line", "2nd line", style()), + edit(&m, "First line", "1st line", style()), + ], + ); + let p = ok_plan(&out); + assert_eq!(p.edited_parts, vec![0], "PLAN-20 one part"); + assert_eq!(p.splices.len(), 2, "PLAN-20 two splices"); + assert!( + p.splices[0].joined.start < p.splices[1].joined.start, + "PLAN-20 sorted" + ); + let text = after_text(&out); + assert!( + text.starts_with("BT /F1 12 Tf 72 700 Td [<31737420"), + "PLAN-20 {text}" + ); + assert!(text.contains("Td [<326E64206C696E65>"), "PLAN-20 {text}"); +} + +#[test] +fn plan_21_edits_in_parts_0_and_2_of_three() { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts( + &[ + b"BT /F1 12 Tf 72 700 Td (Part zero) Tj ET", + b"0 0 1 rg 72 600 100 20 re f", + b"BT /F1 12 Tf 72 500 Td (Part two) Tj ET", + ], + &format!("/Font << /F1 {f} 0 R >>"), + )); + let c = ctx(d.build()); + let m = model(&c, 0); + let out = plan( + &c, + &m, + &[ + edit(&m, "Part zero", "Part 0", style()), + edit(&m, "Part two", "Part 2", style()), + ], + ); + let p = ok_plan(&out); + assert_eq!(p.edited_parts, vec![0, 2], "PLAN-21"); + assert_eq!( + p.expected_parts[1], b"0 0 1 rg 72 600 100 20 re f", + "PLAN-21 part 1 untouched" + ); + assert_eq!( + p.splices.iter().map(|s| s.part).collect::>(), + vec![0, 2], + "PLAN-21" + ); + let concat: Vec = p.expected_parts.concat(); + assert_eq!( + p.expected_page_digest, + crate::pdf_engine::validate_output::content_digest(&concat), + "PLAN-21 digest of the parts without separator" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/plan3.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/plan3.rs new file mode 100644 index 0000000..745e615 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/plan3.rs @@ -0,0 +1,262 @@ +//! PLAN-22…30: style controls (size, colour, letter spacing, faces) with verbatim restores (B14), +//! sibling segments, absorbed members, and the availability rules (B8, B9). + +use super::{ + ctx, edit, faced, filled, model, ok_plan, plan, plan_one, replacement, sized, spaced, style, +}; +use crate::pdf_engine::text_edit::reasons::{EditProblemCode as P, Face, StyleField}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, word_font, DocBuilder, PageSpec, HELVETICA, +}; + +/// Helvetica `/F1` and Helvetica-Bold `/F2` (a bold sibling) on one page. +pub(crate) fn helvetica_pair(content: &[u8]) -> Vec { + let mut d = DocBuilder::new(); + let r = d.add(HELVETICA); + let b = d.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica-Bold /Encoding /WinAnsiEncoding >>", + ); + d.page(PageSpec::new( + content, + &format!("/Font << /F1 {r} 0 R /F2 {b} 0 R >>"), + )); + d.build() +} + +#[test] +fn plan_22_size_change_sets_and_restores_tf_verbatim() { + // (2.278 × 14 − 2.278 × 12) × 1000 / 14 = 325.4286. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12.0 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello", + sized(14.0), + ); + assert_eq!( + replacement(&out, 0), + "/F1 14 Tf [<48656C6C6F> 325.4286] TJ /F1 12.0 Tf", + "PLAN-22" + ); + let t = &ok_plan(&out).runs[0].target; + assert!(t.size_changed && !t.tc_changed && !t.fill_changed && !t.face_changed); +} + +#[test] +fn plan_23_colour_sets_rg_and_restores_verbatim() { + // A named colour space in force: restored with its own `cs` and `scn` bytes. + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::new( + b"BT /F1 12 Tf /CS0 cs 0.2 0.4 0.6 scn 72 700 Td (Hello) Tj ET 72 600 50 50 re f", + &format!("/Font << /F1 {f} 0 R >> /ColorSpace << /CS0 /DeviceRGB >>"), + )); + let (_, _, out) = plan_one(d.build(), "Hello", "Hello", filled("#0000ff")); + assert_eq!( + replacement(&out, 0), + "0 0 1 rg [<48656C6C6F>] TJ /CS0 cs 0.2 0.4 0.6 scn", + "PLAN-23 verbatim cs + scn" + ); + // Never set on the page: restored as `0 g`. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET 72 600 50 50 re f"), + "Hello", + "Hello", + filled("#C71C1C"), + ); + assert_eq!( + replacement(&out, 0), + "0.7804 0.1098 0.1098 rg [<48656C6C6F>] TJ 0 g", + "PLAN-23 default black" + ); +} + +#[test] +fn plan_24_b14_letter_spacing_restores_tc_verbatim() { + // +1 pt letter spacing on 5 glyphs: Tc' = 1; compensation (1 − 0.123456) × 5 × 1000 / 12. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 0.123456 Tc 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello", + spaced(1.0), + ); + assert_eq!( + replacement(&out, 0), + "1 Tc [<48656C6C6F> 365.2267] TJ 0.123456 Tc", + "PLAN-24 verbatim restore (B14: never rounded to 4 dp)" + ); + // Never set: restored as `0 Tc`. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hi) Tj ET"), + "Hi", + "Hi", + spaced(-0.5), + ); + assert_eq!( + replacement(&out, 0), + "-0.5 Tc [<4869> -83.3333] TJ 0 Tc", + "PLAN-24 0 Tc" + ); +} + +#[test] +fn plan_25_face_swap_to_a_sibling() { + // Helvetica-Bold: 722 + 556 + 278 + 278 + 611 = 2445 vs 2278 ⇒ 167 back. + let (_, _, out) = plan_one( + helvetica_pair(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F2 12 Tf 72 680 Td (Bold) Tj ET"), + "Hello", + "Hello", + faced(Face::Bold), + ); + assert_eq!( + replacement(&out, 0), + "/F2 12 Tf [<48656C6C6F> 167] TJ /F1 12 Tf", + "PLAN-25" + ); + let run = &ok_plan(&out).runs[0]; + assert!(run.target.face_changed, "PLAN-25"); + assert!( + run.expected.glyphs.iter().all(|(res, _, _)| res == b"F2"), + "PLAN-25 all in F2" + ); +} + +#[test] +fn plan_26_multi_segment_sibling_body_and_restore() { + let tr = "Sağlık Bakanlığı Raporu"; + // Ends in F1 (the primary's font): no restore. + let (_, _, out) = plan_one(fx::word_tr(), tr, "Sağlık Bakanlığı Rapor", style()); + let r = replacement(&out, 0); + let switches: Vec<&str> = r.matches(" Tf").collect(); + assert_eq!(switches.len(), 6, "PLAN-26 F1/F2/F1/F2/F1/F2/F1: {r}"); + assert!( + r.starts_with("[<5361>] TJ /F2 11 Tf [<0001>] TJ /F1 11 Tf [<6C> "), + "PLAN-26 F1 in force first, then F2 with the verbatim size token: {r}" + ); + assert!( + r.ends_with("] TJ"), + "PLAN-26 no restore when F1 ends the body: {r}" + ); + // Ends in F2: the primary's `Tf` is restored verbatim. + let (_, _, out) = plan_one(fx::word_tr(), tr, "Sağlık Bakanlığı Raporş", style()); + let r = replacement(&out, 0); + assert!(r.ends_with("] TJ /F1 11 Tf"), "PLAN-26 restore: {r}"); + // The six other members draw nothing and keep their travel. + for m in 1..7 { + let a = replacement(&out, m); + assert!( + a.starts_with("[<> ") || a == "[<>] TJ", + "PLAN-26 absorbed {m}: {a}" + ); + } +} + +#[test] +fn plan_27_absorbed_members_keep_followers() { + // Td-positioned member (`18 0 Td` to the pen) and a pen-chained follower drawn at 14 pt + // (not joined: another size); the follower must stay where it was (self-check). + let pdf = + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hel) Tj 18 0 Td (lo) Tj /F1 14 Tf ( tail) Tj ET"); + let (_, _, out) = plan_one(pdf, "Hello", "Help", style()); + // Primary: Help = 722 + 556 + 222 + 556 = 2056 vs Hel 1500 ⇒ 556 back. + assert_eq!( + replacement(&out, 0), + "[<48656C70> 556] TJ", + "PLAN-27 primary" + ); + // Absorbed `(lo)`: 222 + 556 = 778 forward. + assert_eq!(replacement(&out, 1), "[<> -778] TJ", "PLAN-27 absorbed"); + let exp = &ok_plan(&out).runs[0].expected; + assert_eq!(exp.emitted_records, vec![1, 1], "PLAN-27 one show op each"); + // Pen-chained members in one op sequence: `(Hel) Tj (lo) Tj`. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hel) Tj (lo) Tj /F1 14 Tf ( tail) Tj ET"), + "Hello", + "Hallo there", + style(), + ); + assert!( + replacement(&out, 1).starts_with("[<> -778] TJ"), + "PLAN-27 chained" + ); +} + +#[test] +fn plan_28_b9_extgstate_font_size_and_face_unavailable() { + let c = ctx(fx::extgstate_font()); + let m = model(&c, 0); + for (s, field) in [ + (sized(14.0), StyleField::Size), + (faced(Face::Bold), StyleField::Face), + ] { + let out = plan(&c, &m, &[edit(&m, "Hi there", "Hi there", s)]); + let p = out.verdicts[0].problem.as_ref().expect("STYLE_UNAVAILABLE"); + assert_eq!( + (p.code, p.field), + (P::StyleUnavailable, Some(field)), + "PLAN-28" + ); + } + // Text and colour stay available with an ExtGState font (no Tf is written). + let out = plan( + &c, + &m, + &[edit(&m, "Hi there", "Hi here", filled("#ff0000"))], + ); + let r = replacement(&out, 0); + assert!(!r.contains("Tf"), "PLAN-28 never a Tf: {r}"); + assert!(r.starts_with("1 0 0 rg [<"), "PLAN-28 {r}"); +} + +#[test] +fn plan_29_tr2_colour_unavailable() { + let c = ctx(fx::skia()); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Bold", "Bold", filled("#ff0000"))]); + let p = out.verdicts[0].problem.as_ref().expect("STYLE_UNAVAILABLE"); + assert_eq!( + (p.code, p.field), + (P::StyleUnavailable, Some(StyleField::Colour)), + "PLAN-29" + ); + // The text itself stays editable. + let out = plan(&c, &m, &[edit(&m, "Bold", "Bil", style())]); + assert!( + out.verdicts[0].problem.is_none(), + "PLAN-29 text: {:?}", + out.verdicts[0].problem + ); +} + +#[test] +fn plan_30_b8_face_unavailable_with_chars() { + // A bold subset that lacks W, r and d. + let mut d = DocBuilder::new(); + let f1 = word_font(&mut d.b, "ABCDEF+Calibri", "Hello World"); + let f2 = word_font(&mut d.b, "ABCDEF+Calibri-Bold", "Helo"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello World) Tj ET BT /F2 12 Tf 72 680 Td (Helo) Tj ET", + &format!("/Font << /F1 {f1} 0 R /F2 {f2} 0 R >>"), + )); + let (_, _, out) = plan_one(d.build(), "Hello World", "Hello World", faced(Face::Bold)); + let p = out.verdicts[0].problem.as_ref().expect("FACE_UNAVAILABLE"); + assert_eq!(p.code, P::FaceUnavailable, "PLAN-30"); + assert_eq!(p.face, Some(Face::Bold), "PLAN-30 face"); + assert_eq!( + p.chars, + vec!['W', 'r', 'd'], + "PLAN-30 chars in typing order" + ); + // No italic group at all: FACE_UNAVAILABLE without chars. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET"), + "Hello", + "Hello", + faced(Face::Italic), + ); + let p = out.verdicts[0].problem.as_ref().expect("FACE_UNAVAILABLE"); + assert_eq!( + (p.code, p.face, p.chars.len()), + (P::FaceUnavailable, Some(Face::Italic), 0), + "PLAN-30" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/plan4.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/plan4.rs new file mode 100644 index 0000000..5fc7b48 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/plan4.rs @@ -0,0 +1,278 @@ +//! PLAN-31…38: number formatting, the f32 and read-back guards, the replacement grammar, the +//! self-check, the boundary-kern rule, column-gap absorption and the full-digest rule for edited +//! glyphs. + +use super::{ctx, edit, model, ok_plan, plan, plan_one, problem_of, replacement, sized, style}; +use crate::pdf_engine::text_edit::encode::{check_replacement_grammar, fmt_num}; +use crate::pdf_engine::text_edit::reasons::{EditProblemCode as P, TextWarningCode}; +use crate::pdf_engine::text_edit::rewrite::{assemble_page_plan, RunPlan}; +use crate::pdf_engine::text_edit::testkit::producers::{self as fx, helvetica_page}; +use crate::pdf_engine::text_edit::verify::{walk_and_verify, VerifyFailure}; + +#[test] +fn plan_31_fmt_num() { + let ok = |v: f64| fmt_num(v).unwrap_or_else(|e| panic!("PLAN-31 {v}: {e:?}")); + assert_eq!(ok(1.0), "1"); + assert_eq!(ok(0.5), "0.5"); + assert_eq!(ok(-2.5), "-2.5"); + assert_eq!(ok(1.23456), "1.2346", "4 decimals"); + assert_eq!(ok(0.00004), "0", "rounds to zero"); + assert_eq!(ok(-0.00004), "0", "-0 → 0"); + assert_eq!(ok(-0.0), "0", "-0 → 0"); + assert_eq!(ok(325.42857), "325.4286"); + assert_eq!(ok(1e9), "1000000000", "no exponent at the limit"); + assert_eq!(ok(123456789.123), "123456789.123"); + for bad in [1e9 + 1.0, -2e9, f64::NAN, f64::INFINITY, 1e20] { + let e = fmt_num(bad) + .err() + .unwrap_or_else(|| panic!("PLAN-31 {bad} must fail")); + assert_eq!(e.code, P::EditVerifyFailed, "PLAN-31 {bad}"); + } + for v in [0.1, 2.0 / 3.0, 1e-3, 12345.6789, -987.65432] { + let s = ok(v); + assert!( + !s.contains('e') && !s.contains('E'), + "PLAN-31 no exponent: {s}" + ); + assert!( + s.split('.').nth(1).map_or(0, str::len) <= 4, + "PLAN-31 ≤ 4 dp: {s}" + ); + } +} + +#[test] +fn plan_32_f32_and_readback_guards() { + // An absorbed member that travels 360,000 pt (a trailing −30,000,000.3 kern): its written + // number cannot survive a viewer's f32 parse within 0.002 pt ⇒ "number precision". + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj [(world)-30000000.3]TJ ET"), + "Helloworld", + "Help", + style(), + ); + let v = &out.verdicts[0]; + let p = v.problem.as_ref().expect("PLAN-32 f32 guard"); + assert_eq!(p.code, P::EditVerifyFailed, "PLAN-32"); + assert!( + p.detail + .as_deref() + .unwrap_or("") + .contains("number precision"), + "PLAN-32 {p:?}" + ); + // A tiny Tf operand under a large Tm: 4 decimals cannot express the new size (read back). + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 0.001 Tf 12000 0 0 12000 72 700 Tm (Hi) Tj ET"), + "Hi", + "Hi", + sized(13.0), + ); + let p = out.verdicts[0] + .problem + .as_ref() + .expect("PLAN-32 read-back guard"); + assert_eq!(p.code, P::EditVerifyFailed, "PLAN-32 read-back"); +} + +#[test] +fn plan_33_grammar_rejects_everything_but_the_replacement_operators() { + let ok = + "0.2 Tw 0.1 Tc T* 1 Tc 0 0 1 rg /F2 12 Tf [<41> -2 <42>] TJ /CS0 cs 0.5 scn 0 Tc /F1 12 Tf"; + assert_eq!( + check_replacement_grammar(ok.as_bytes(), 1), + Ok(()), + "PLAN-33 allowed" + ); + assert_eq!( + check_replacement_grammar(b"0 g 0 0 0 1 k 1 g [<>] TJ", 1), + Ok(()) + ); + let bad: [(&str, &str); 14] = [ + ("q [<41>] TJ Q", "q"), + ("0 0 9 9 re f [<41>] TJ", "re"), + ("/Im0 Do [<41>] TJ", "Do"), + ("2 Tr [<41>] TJ", "Tr"), + ("1 0 0 1 5 5 cm [<41>] TJ", "cm"), + ("50 Tz [<41>] TJ", "Tz"), + ("3 Ts [<41>] TJ", "Ts"), + ("BT [<41>] TJ ET", "BT"), + ("1 0 0 1 0 0 Tm [<41>] TJ", "Tm"), + ("5 0 Td [<41>] TJ", "Td"), + ("/P BMC [<41>] TJ EMC", "BMC"), + ("/GS1 gs [<41>] TJ", "gs"), + ("2 w [<41>] TJ", "w"), + ("[<41>] TJ [<42>] TJ", "TJ count"), + ]; + for (bytes, why) in bad { + let r = check_replacement_grammar(bytes.as_bytes(), 1); + assert_eq!(r, Err(why), "PLAN-33 {bytes:?}"); + } + assert!( + check_replacement_grammar(b"[(unbalanced] TJ", 1).is_err(), + "PLAN-33 lex" + ); + assert!( + check_replacement_grammar(b"BX foo EX [<41>] TJ", 1).is_err(), + "PLAN-33 BX" + ); + assert_eq!( + check_replacement_grammar(b"[<41>] TJ", 2), + Err("TJ count"), + "PLAN-33 count" + ); +} + +#[test] +fn plan_34_self_check_detects_an_injected_wrong_compensation() { + let c = ctx(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 12 Tf 72 680 Td (Next) Tj (after) Tj ET", + )); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Hello", "Help", style())]); + let good = ok_plan(&out); + let after = good.expected_content(&m.content); + assert!( + walk_and_verify(&c, 0, &after, &m.walk, &m.runs, good, None).is_ok(), + "PLAN-34 the honest plan passes" + ); + // The same plan with the compensation off by one unit (0.012 pt at 12 pt). + let mut runs: Vec = good.runs.clone(); + let s = &mut runs[0].splices[0]; + let text = String::from_utf8_lossy(&s.bytes).replace("] TJ", " 1] TJ"); + s.bytes = text.into_bytes(); + let bad = assemble_page_plan(&m.content, 0, runs); + let after = bad.expected_content(&m.content); + let err = walk_and_verify(&c, 0, &after, &m.walk, &m.runs, &bad, None) + .err() + .expect("PLAN-34 a wrong compensation must fail"); + assert!( + matches!( + err, + VerifyFailure::EditedMismatch { what: "pen", .. } | VerifyFailure::Drift { .. } + ), + "PLAN-34 {err:?}" + ); +} + +#[test] +fn plan_35_boundary_rule() { + // A synthetic-space kern at the prefix/middle boundary is part of the kept text. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td [(Hello)-333(World)]TJ ET"), + "Hello World", + "Hello Earth", + style(), + ); + assert!( + replacement(&out, 0).starts_with("[<48656C6C6F> -333 <45617274"), + "PLAN-35 synthetic" + ); + // A pair kern at the boundary is dropped (its pair is gone). + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td [(AV)-80(x)]TJ ET"), + "AVx", + "AWx", + style(), + ); + let r = replacement(&out, 0); + assert!( + r.starts_with("[<415778> ") && !r.contains("-80"), + "PLAN-35 pair kern: {r}" + ); + // An empty middle: the sides meet. Original neighbours keep their kern; new neighbours don't. + let pdf = || helvetica_page(b"BT /F1 12 Tf 72 700 Td [(A)-80(V)(x)]TJ ET"); + let (_, _, out) = plan_one(pdf(), "AVx", "AV", style()); + assert!( + replacement(&out, 0).starts_with("[<41> -80 <56> "), + "PLAN-35 original neighbours" + ); + let (_, _, out) = plan_one(pdf(), "AVx", "Ax", style()); + let r = replacement(&out, 0); + assert!( + r.starts_with("[<4178> ") && !r.contains("-80"), + "PLAN-35 new neighbours: {r}" + ); +} + +#[test] +fn plan_36_column_gap_absorbs_the_width_change() { + // "Name" → "Names": s = 500 thousandths; the −3000 gap becomes −2500, "Value" stays put, + // and no compensation is written. + let (_, _, out) = plan_one( + helvetica_page(b"BT /F1 12 Tf 72 700 Td [(Name) -3000 (Value)] TJ ET"), + "Name Value", + "Names Value", + style(), + ); + assert_eq!( + replacement(&out, 0), + "[<4E616D6573> -2500 <56616C7565>] TJ", + "PLAN-36" + ); + let exp = &ok_plan(&out).runs[0].expected; + assert_eq!( + exp.unshifted_from, + Some(5), + "PLAN-36 Value keeps its exact origin" + ); +} + +#[test] +fn plan_37_absorption_that_would_shrink_the_gap_below_a_space_is_not_applied() { + // A 1.1 em gap; MM (1.666 em) would leave −566 (not a space): normal compensation instead, + // and the longer line now runs into "Next". + let (_, _, out) = plan_one( + helvetica_page( + b"BT /F1 12 Tf 72 700 Td [(Name) -1100 (Value)] TJ ET BT /F1 12 Tf 155 700 Td (Next) Tj ET", + ), + "Name Value", + "NameMM Value", + style(), + ); + assert_eq!( + replacement(&out, 0), + "[<4E616D654D4D> -1100 <56616C7565> 1666] TJ", + "PLAN-37" + ); + assert_eq!( + out.verdicts[0].warnings, + vec![TextWarningCode::NextTextOverlap], + "PLAN-37" + ); + assert_eq!( + ok_plan(&out).runs[0].expected.unshifted_from, + None, + "PLAN-37" + ); +} + +#[test] +fn plan_38_edited_tr2_run_keeps_stroke_line_width_and_dash() { + let c = ctx(fx::stroke_text(2, [0.5, 0.5], ["[3] 0", "[3] 0"])); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Boldface", "Bold face", style())]); + let p = ok_plan(&out); + let after = p.expected_content(&m.content); + let walk = walk_and_verify(&c, 0, &after, &m.walk, &m.runs, p, None) + .unwrap_or_else(|f| panic!("PLAN-38 self-check: {f}")); + let edited: Vec<_> = walk + .records + .iter() + .filter(|r| !r.glyphs.is_empty()) + .collect(); + assert!(!edited.is_empty(), "PLAN-38"); + for r in edited { + let s = &r.before; + assert_eq!(s.text.tr, 2, "PLAN-38 Tr"); + assert_eq!(s.gs.line_width, 0.5, "PLAN-38 line width"); + assert_eq!(&s.gs.dash.0[..], &[3.0], "PLAN-38 dash"); + } + // Colour is not offered for Tr 2 (only the fill would change, not the stroke). + let out = plan( + &c, + &m, + &[edit(&m, "Boldface", "Boldface", super::filled("#ff0000"))], + ); + assert_eq!(problem_of(&out), P::StyleUnavailable, "PLAN-38 colour"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_plan/ver.rs b/src-tauri/src/pdf_engine/text_edit/tests_plan/ver.rs new file mode 100644 index 0000000..0d0888d --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_plan/ver.rs @@ -0,0 +1,373 @@ +//! VER-01…11: the verification re-walk (§B.14) shared by the self-check, the preview and Save. + +use super::{ctx, edit, model, ok_plan, plan, spaced, style}; +use crate::pdf_engine::text_edit::apply::{apply_update, updates_for_plan, write_update_json}; +use crate::pdf_engine::text_edit::content::page_content; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::engines::RunOpts; +use crate::pdf_engine::text_edit::lexer::Span; +use crate::pdf_engine::text_edit::limits::{PAGE_DECODE_BUDGET, VERIFY_CAP_MARGIN_BYTES}; +use crate::pdf_engine::text_edit::reasons::EditProblemCode as P; +use crate::pdf_engine::text_edit::rewrite::{assemble_page_plan, PagePlan, RunPlan, Splice}; +use crate::pdf_engine::text_edit::runs::PageModel; +use crate::pdf_engine::text_edit::snapshot::read_verification_snapshot; +use crate::pdf_engine::text_edit::state::{same_paint, ColorEffect, ColorSpaceKind, Paint}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use crate::pdf_engine::text_edit::verify::{map_span, test_seams, walk_and_verify, VerifyFailure}; +use std::sync::Arc; + +fn verify(c: &SnapshotContext, m: &PageModel, p: &PagePlan) -> Result<(), VerifyFailure> { + let after = p.expected_content(&m.content); + walk_and_verify(c, m.page_index, &after, &m.walk, &m.runs, p, None).map(|_| ()) +} + +/// `p` with its first run's primary replacement rewritten by `f` (re-assembled consistently). +fn tampered(m: &PageModel, p: &PagePlan, f: impl Fn(&str) -> String) -> PagePlan { + let mut runs: Vec = p.runs.clone(); + let s = &mut runs[0].splices[0]; + s.bytes = f(&String::from_utf8_lossy(&s.bytes)).into_bytes(); + assemble_page_plan(&m.content, m.page_index, runs) +} + +/// The expected content of `p` with part 0's bytes changed by `f` outside the plan. +fn after_with( + m: &PageModel, + p: &PagePlan, + f: impl Fn(&str) -> String, +) -> crate::pdf_engine::text_edit::content::PageContent { + let after = p.expected_content(&m.content); + let bytes = f(&String::from_utf8_lossy(after.part_bytes(0))).into_bytes(); + after.with_replaced_parts(&[(0, bytes)]) +} + +fn hello_other() -> Vec { + helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 12 Tf 72 650 Td (Other) Tj ET 72 600 50 20 re f") +} + +#[test] +fn ver_01_honest_edits_pass() { + for (pdf, old, new) in [ + (hello_other(), "Hello", "Help"), + (fx::word(), "Due", "Due 7"), + ( + fx::word_tr(), + "Sağlık Bakanlığı Raporu", + "Sağlığı Bakanlık Raporu", + ), + (fx::pdftex(), "Hello World", "Hello Wet World"), + (fx::skia(), "Chrome", "Chrom"), + (fx::quote_ops(), "Line two", "Line 2"), + ] { + let c = ctx(pdf); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, old, new, style())]); + assert_eq!( + verify(&c, &m, ok_plan(&out)), + Ok(()), + "VER-01 {old:?} → {new:?}" + ); + } +} + +#[test] +fn ver_02_span_mapping_across_splices_and_parts() { + let s = |start: usize, end: usize, len: usize| Splice { + part: 0, + local: start..end, + joined: start..end, + bytes: vec![b'x'; len], + }; + let splices = [s(10, 20, 15), s(30, 32, 0), s(50, 60, 10)]; + let map = |a: usize, b: usize| map_span(&splices, &(a..b)); + assert_eq!(map(0, 5), 0..5, "before every splice"); + assert_eq!(map(20, 25), 25..30, "after the first (+5)"); + assert_eq!(map(32, 40), 35..43, "after the second (+5 −2)"); + assert_eq!(map(70, 80), 73..83, "after all (+5 −2 +0)"); + let _: Span = map(0, 0); + // A page with edits in parts 0 and 2 (two in part 0) verifies through the mapping. + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts( + &[ + b"BT /F1 12 Tf 72 700 Td (One) Tj ET BT /F1 12 Tf 72 680 Td (Two) Tj ET", + b"BT /F1 12 Tf 72 660 Td (Three) Tj ET", + b"BT /F1 12 Tf 72 640 Td (Four) Tj ET 0 0 1 rg 72 600 20 20 re f", + ], + &format!("/Font << /F1 {f} 0 R >>"), + )); + let c = ctx(d.build()); + let m = model(&c, 0); + let out = plan( + &c, + &m, + &[ + edit(&m, "One", "Uno", style()), + edit(&m, "Two", "Dos", style()), + edit(&m, "Four", "Cuatro", style()), + ], + ); + let p = ok_plan(&out); + assert_eq!(p.edited_parts, vec![0, 2], "VER-02"); + assert_eq!(verify(&c, &m, p), Ok(()), "VER-02 multi-splice, multi-part"); +} + +#[test] +fn ver_03_an_unedited_record_unpaired_fails() { + let c = ctx(hello_other()); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Hello", "Help", style())]); + let p = ok_plan(&out); + let after = after_with(&m, p, |s| s.replace("(Other) Tj", "(Other)'")); + let r = walk_and_verify(&c, 0, &after, &m.walk, &m.runs, p, None).err(); + assert!( + matches!(r, Some(VerifyFailure::RecordUnpaired { .. })), + "VER-03 {r:?}" + ); +} + +#[test] +fn ver_04_a_paint_count_change_fails() { + let c = ctx(hello_other()); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Hello", "Help", style())]); + let p = ok_plan(&out); + let after = after_with(&m, p, |s| format!("{s} 0 0 1 rg 100 100 5 5 re f")); + let r = walk_and_verify(&c, 0, &after, &m.walk, &m.runs, p, None).err(); + assert!( + matches!(r, Some(VerifyFailure::RecordCount { .. })), + "VER-04 {r:?}" + ); +} + +#[test] +fn ver_05_each_failure_maps_to_its_problem_code() { + let cases = [ + ( + VerifyFailure::Drift { + record: 0, + glyph: 0, + pt: 0.5, + }, + P::PenDrift, + ), + ( + VerifyFailure::StateChanged { + record: 0, + field: "tc", + }, + P::StateChanged, + ), + ( + VerifyFailure::PaintChanged { + index: 0, + field: "fill", + }, + P::StateChanged, + ), + ( + VerifyFailure::PageRefused( + crate::pdf_engine::text_edit::reasons::TextReason::MalformedContent, + ), + P::EditVerifyFailed, + ), + ( + VerifyFailure::RecordUnpaired { index: 0 }, + P::EditVerifyFailed, + ), + ( + VerifyFailure::RecordCount { + expected: 1, + found: 2, + }, + P::EditVerifyFailed, + ), + ( + VerifyFailure::EditedMismatch { + run_id: "r".into(), + what: "text", + }, + P::EditVerifyFailed, + ), + ( + VerifyFailure::ForbiddenOperator { op: "q" }, + P::EditVerifyFailed, + ), + ]; + for (f, code) in cases { + assert_eq!(f.problem_code(), code, "VER-05 {f}"); + assert!(!f.to_string().is_empty(), "VER-05 detail"); + } +} + +#[test] +fn ver_06_honest_edit_passes_after_qpdf_renumbering() { + let Some(engines) = engines_or_skip("ver_06") else { + return; + }; + let dir = Scratch::new("ver_06"); + let pdf = fx::word(); + let source = dir.write("source.pdf", &pdf); + let c = ctx(pdf); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Due", "Due 7", style())]); + let p = ok_plan(&out); + let update = dir.path("update.json"); + write_update_json( + &updates_for_plan(&m.content, p).expect("updates"), + c.doc().max_id, + &update, + ) + .expect("json"); + let staged = dir.path("staged.pdf"); + apply_update( + &engines, + &source, + &update, + &staged, + &[], + &RunOpts::default(), + ) + .expect("qpdf"); + let snap = read_verification_snapshot(&staged, VERIFY_CAP_MARGIN_BYTES).expect("staged"); + let s = SnapshotContext::new(snap); + let ids: Vec<_> = (0..s.snap.pages.len()).map(|i| s.snap.pages[i]).collect(); + assert_ne!(ids, c.snap.pages, "VER-06 qpdf renumbered the objects"); + let page = s.page_id(0).expect("page"); + let content = + page_content(s.doc(), page, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)).expect("content"); + let r = walk_and_verify(&s, 0, &content, &m.walk, &m.runs, p, None).map(|_| ()); + assert_eq!(r, Ok(()), "VER-06 id-free comparisons"); +} + +#[test] +fn ver_07_b14_post_state_catches_a_missing_restore() { + // No follower in the text object: only the post-state probe sees Tc still at 1. + let c = ctx(helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET")); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Hello", "Hello", spaced(1.0))]); + let p = ok_plan(&out); + let bad = tampered(&m, p, |s| s.replace("] TJ 0 Tc", "] TJ")); + let r = verify(&c, &m, &bad); + assert!( + matches!(r, Err(VerifyFailure::StateChanged { field: "tc", .. })), + "VER-07 {r:?}" + ); +} + +#[test] +fn ver_08_grammar_rechecked_on_the_rewalk() { + let c = ctx(hello_other()); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Hello", "Help", style())]); + let p = ok_plan(&out); + let bad = tampered(&m, p, |s| format!("2 w {s}")); + assert_eq!( + verify(&c, &m, &bad), + Err(VerifyFailure::ForbiddenOperator { op: "w" }), + "VER-08" + ); +} + +#[test] +fn ver_09_swapped_image_behind_the_same_name() { + let c = ctx(fx::swapped_image(false)); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Image page", "Image pages", style())]); + let p = ok_plan(&out); + assert_eq!(verify(&c, &m, p), Ok(()), "VER-09 honest"); + // The same content walked in the file whose /Im0 holds other pixels. + let swapped = ctx(fx::swapped_image(true)); + let after = p.expected_content(&m.content); + let r = walk_and_verify(&swapped, 0, &after, &m.walk, &m.runs, p, None).err(); + assert_eq!( + r, + Some(VerifyFailure::PaintChanged { + index: 0, + field: "kind" + }), + "VER-09" + ); + assert_eq!( + r.map(|f| f.problem_code()), + Some(P::StateChanged), + "VER-09 code" + ); +} + +#[test] +fn ver_10_leaked_line_width_on_an_edited_tr1_glyph() { + let c = ctx(fx::stroke_text(1, [0.5, 0.5], ["[] 0", "[] 0"])); + let m = model(&c, 0); + let out = plan(&c, &m, &[edit(&m, "Boldface", "Bold face", style())]); + let p = ok_plan(&out); + let bad = tampered(&m, p, |s| format!("2 w {s}")); + let _g = test_seams::skip_grammar(); + let r = verify(&c, &m, &bad); + assert!( + matches!( + r, + Err(VerifyFailure::StateChanged { + field: "line_width", + .. + }) + ), + "VER-10 {r:?}" + ); +} + +#[test] +fn ver_11_same_paint() { + let paint = |space: ColorSpaceKind, comps: &[f64], op: Option<&str>| Paint { + space_op: None, + color_op: op.map(|o| Arc::from(o.as_bytes())), + space, + comps: comps.to_vec(), + effect: ColorEffect::Rgb([0.0; 3]), + pattern: false, + pattern_hash: None, + }; + let never = Paint::initial(); + let rgb0 = paint( + ColorSpaceKind::DeviceRgb, + &[0.0, 0.0, 0.0], + Some("0 0 0 rg"), + ); + let gray0 = paint(ColorSpaceKind::DeviceGray, &[0.0], Some("0 g")); + let k_black = paint( + ColorSpaceKind::DeviceCmyk, + &[0.0, 0.0, 0.0, 1.0], + Some("0 0 0 1 k"), + ); + let rich = paint( + ColorSpaceKind::DeviceCmyk, + &[0.6, 0.4, 0.4, 1.0], + Some(".6 .4 .4 1 k"), + ); + assert!(same_paint(&never, &rgb0), "VER-11 never set vs 0 0 0 rg"); + assert!( + same_paint(&never, &gray0) && same_paint(&never, &k_black), + "VER-11 default black" + ); + assert!( + !same_paint(&gray0, &k_black), + "VER-11 explicit 0 g vs 0 0 0 1 k" + ); + assert!( + !same_paint(&gray0, &rgb0), + "VER-11 explicit 0 g vs 0 0 0 rg" + ); + assert!( + !same_paint(&rich, &k_black), + "VER-11 rich black vs K-only black" + ); + let near = paint( + ColorSpaceKind::DeviceRgb, + &[0.0, 0.0, 0.0000005], + Some("0 0 0.0000005 rg"), + ); + assert!(same_paint(&rgb0, &near), "VER-11 within COLOR_EPSILON"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk.rs new file mode 100644 index 0000000..e723ac1 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk.rs @@ -0,0 +1,87 @@ +//! T3 tests (SPEC §E.3): WALK-01…27 (`tests_walk/walk.rs`, `walk2.rs`), GEO-01…15 (`geo.rs`), +//! RUN-01…22 (`runs.rs`, `runs2.rs`), the producer smoke tests (§E.2, `producers.rs`), the +//! walker fuzz (§B.21, `fuzz.rs`) and the review regressions on work and memory bounds and model +//! fidelity (`bounds.rs`, `bounds2.rs`, `bounds3.rs`, `bounds4.rs`, `pen.rs`), the page-model +//! budget against every amplification probe (`budget.rs`), and the engine fix pass of 2026-10-03 +//! (`fixes.rs`: fonts, refused and dense pages, the shared Classify budget; `joins.rs`: content +//! part boundaries, inline-image ends). Shared helpers live here. + +mod bounds; +mod bounds2; +mod bounds3; +mod bounds4; +mod budget; +mod fixes; +mod fuzz; +mod geo; +mod joins; +mod pen; +mod producers; +mod runs; +mod runs2; +mod walk; +mod walk2; + +use crate::pdf_engine::text_edit::content::{page_content, PageContent}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::limits::PAGE_DECODE_BUDGET; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::runs::{build_page_model, PageModel, TextRun}; +use crate::pdf_engine::text_edit::snapshot::snapshot_from_bytes; +use crate::pdf_engine::text_edit::walker::{walk_page, PageWalk, WalkMode}; +use std::path::Path; + +/// A context over `pdf` read by the production snapshot reader. +pub(super) fn ctx(pdf: Vec) -> SnapshotContext { + let snap = snapshot_from_bytes(Path::new("walk-fixture.pdf"), pdf, None) + .unwrap_or_else(|e| panic!("fixture must open: {e} {:?}", e.details)); + SnapshotContext::new(snap) +} + +pub(super) fn model(ctx: &SnapshotContext, page: u32) -> PageModel { + build_page_model(ctx, page, None).unwrap_or_else(|e| panic!("page model: {e}")) +} + +/// The model of page 0 of `pdf`. +pub(super) fn model0(pdf: Vec) -> PageModel { + model(&ctx(pdf), 0) +} + +pub(super) fn content(ctx: &SnapshotContext, page: u32) -> PageContent { + let id = ctx.page_id(page).expect("page id"); + page_content(ctx.doc(), id, &mut DecodeBudget::new(PAGE_DECODE_BUDGET)) + .unwrap_or_else(|r| panic!("page content: {r:?}")) +} + +pub(super) fn walk(ctx: &SnapshotContext, page: u32, mode: WalkMode) -> PageWalk { + let c = content(ctx, page); + walk_page(ctx, page, &c, mode, None) +} + +pub(super) fn texts(m: &PageModel) -> Vec { + m.runs.iter().map(|r| r.text.clone()).collect() +} + +/// The run whose text is `text` (panics with the page's runs otherwise). +pub(super) fn run_with<'m>(m: &'m PageModel, text: &str) -> &'m TextRun { + m.runs.iter().find(|r| r.text == text).unwrap_or_else(|| { + panic!( + "no run {text:?}; runs: {:?}; page_reason {:?} ({:?})", + m.runs + .iter() + .map(|r| (r.text.clone(), r.reason)) + .collect::>(), + m.page_reason, + m.page_detail + ) + }) +} + +pub(super) fn reason_of(m: &PageModel, text: &str) -> Option { + run_with(m, text).reason +} + +pub(super) fn close(a: f64, b: f64) -> bool { + (a - b).abs() < 1e-6 +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds.rs new file mode 100644 index 0000000..af47a54 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds.rs @@ -0,0 +1,298 @@ +//! Review regressions (T3 fix pass): the walk's work and memory stay bounded whatever a page +//! repeats — state digests share their bytes (HIGH-1), deep hashes are memoised and charged per +//! walk (HIGH-2), and the page-wide run checks hold one flag per run and a capped number of +//! comparisons (MEDIUM-1). + +use super::fuzz::thread_cpu; +use super::{model0, run_with, texts}; +use crate::pdf_engine::text_edit::limits::{PAGE_DECODE_BUDGET, STREAM_MAX_DECODED}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::testkit::pdf::zlib_zero_bomb; +use crate::pdf_engine::text_edit::testkit::producers::{ + helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::thread_peak; +use crate::pdf_engine::text_edit::walker::hash::tests::{set_steps_override, take_decode_attempts}; +use std::sync::Arc; +use std::time::Duration; + +const MIB: usize = 1 << 20; + +/// `records` × `(a)Tj` after `prefix`, in one text object. +fn many_shows(prefix: &[u8], records: usize) -> Vec { + let mut c = prefix.to_vec(); + c.extend_from_slice(b" BT /F1 1 Tf 72 700 Td "); + for _ in 0..records { + c.extend_from_slice(b"(a)Tj "); + } + c.extend_from_slice(b"ET"); + c +} + +#[test] +fn state_digests_share_verbatim_bytes_and_line_state() { + // 1 MiB between `0` and `g`: the op span is 1 MiB, and 200 records each hold two digests. + let mut padded = b"0".to_vec(); + padded.extend(std::iter::repeat(b' ').take(MIB)); + padded.extend_from_slice(b" g"); + let (m, peak) = thread_peak(|| model0(helvetica_page(&many_shows(&padded, 200)))); + assert_eq!(m.page_reason, None); + assert_eq!(m.walk.records.len(), 200); + assert!( + peak < 24 * MIB, + "1 MiB colour op × 200 records peaked at {peak} B" + ); + let (first, last) = (&m.walk.records[0], &m.walk.records[199]); + let (a, b) = ( + first.before.fill.color_op.as_ref().expect("colour op"), + last.after.fill.color_op.as_ref().expect("colour op"), + ); + assert!( + Arc::ptr_eq(a, b), + "the verbatim bytes are shared, not copied" + ); + assert_eq!(a.len(), MIB + 3); + + // A 60,000-entry dash array in force for 200 records. + let mut dash = b"[".to_vec(); + for _ in 0..60_000 { + dash.extend_from_slice(b"1 "); + } + dash.extend_from_slice(b"] 0 d"); + let (m, peak) = thread_peak(|| model0(helvetica_page(&many_shows(&dash, 200)))); + assert_eq!(m.page_reason, None); + assert_eq!(m.walk.records[199].before.gs.dash.0.len(), 60_000); + assert!(peak < 24 * MIB, "60k dash × 200 records peaked at {peak} B"); + + // A marked-content tag of 64 KiB under 200 records. + let mut tagged = b"/".to_vec(); + tagged.extend(std::iter::repeat(b'T').take(64 << 10)); + tagged.extend_from_slice(b" BMC"); + let mut c = many_shows(&tagged, 200); + c.extend_from_slice(b" EMC"); + let (m, peak) = thread_peak(|| model0(helvetica_page(&c))); + assert_eq!(m.page_reason, None); + assert!( + peak < 16 * MIB, + "64 KiB tag × 200 records peaked at {peak} B" + ); +} + +/// One page drawing "Hello" after `ops`, with `/GS` ExtGStates from `ext_gstates`. +fn gs_page(ops: &str, ext_gstates: &[String]) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let names: Vec = ext_gstates + .iter() + .enumerate() + .map(|(i, body)| format!("/GS{i} {} 0 R", d.add(body))) + .collect(); + let content = format!("{ops} BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + d.page(PageSpec::new( + content.as_bytes(), + &format!( + "/Font << /F1 {f} 0 R >> /ExtGState << {} >>", + names.join(" ") + ), + )); + d.build() +} + +/// An ExtGState with `n` unmodelled keys `/0 … /{n-1}`. +fn private_keys(prefix: &str, n: usize) -> String { + let keys: Vec = (0..n).map(|i| format!("/{prefix}{i} {i}")).collect(); + format!("<< {} >>", keys.join(" ")) +} + +#[test] +fn extgstate_work_per_gs_is_capped() { + // 8,000 keys in one dictionary: refused before any per-key work. + let started = thread_cpu(); + let m = model0(gs_page("/GS0 gs", &[private_keys("K", 8_000)])); + assert_eq!(m.page_reason, Some(R::PageTooComplex)); + assert_eq!(m.page_detail.as_deref(), Some("ExtGState keys")); + // 40 private keys applied 20,000 times: one sort and merge per `gs`, digests shared. + let ops = "/GS0 gs ".repeat(20_000); + let m = model0(gs_page(&ops, &[private_keys("K", 40)])); + assert_eq!(m.page_reason, None); + assert_eq!(m.walk.records[0].before.gs.other.len(), 40); + assert_eq!(run_with(&m, "Hello").reason, None); + assert!( + thread_cpu().saturating_sub(started) < Duration::from_secs(4), + "ExtGState work" + ); + // Two dictionaries of 40 different private keys: 80 in force at once is too many. + let m = model0(gs_page( + "/GS0 gs /GS1 gs", + &[private_keys("K", 40), private_keys("L", 40)], + )); + assert_eq!(m.page_reason, Some(R::PageTooComplex)); + // A later value of the same key replaces the earlier one (sorted, one entry per key). + let m = model0(gs_page( + "/GS0 gs /GS1 gs", + &[ + "<< /TR /Identity /HT /Default >>".into(), + "<< /TR /Identity >>".into(), + ], + )); + let keys: Vec<&[u8]> = m.walk.records[0] + .before + .gs + .other + .iter() + .map(|(k, _)| &k[..]) + .collect(); + assert_eq!(keys, [&b"HT"[..], b"TR"]); +} + +/// A page painting the shading `/Sh` `uses` times for each of `distinct` 33 MiB Flate bombs. +fn bomb_shadings(distinct: usize, uses: usize) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let bomb = zlib_zero_bomb(33); + let names: Vec = (0..distinct) + .map(|k| { + let id = d.b.add_stream( + "/ShadingType 4 /ColorSpace /DeviceGray /BitsPerCoordinate 8 /BitsPerComponent 8 \ + /BitsPerFlag 8 /Decode [0 1 0 1 0 1] /Filter /FlateDecode", + &bomb, + ); + format!("/Sh{k} {id} 0 R") + }) + .collect(); + let mut c = String::new(); + for _ in 0..uses { + for k in 0..distinct { + c.push_str(&format!("/Sh{k} sh\n")); + } + } + c.push_str("BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + d.page(PageSpec::new( + c.as_bytes(), + &format!("/Font << /F1 {f} 0 R >> /Shading << {} >>", names.join(" ")), + )); + d.build() +} + +#[test] +fn deep_hashes_are_memoised_and_charged_per_walk() { + // 1,000 `sh` of one shading that inflates past STREAM_MAX_DECODED: one decode attempt. + let pdf = bomb_shadings(1, 1_000); + take_decode_attempts(); + let started = thread_cpu(); + let m = model0(pdf); + let spent = thread_cpu().saturating_sub(started); + let (attempts, _) = take_decode_attempts(); + assert_eq!(m.page_reason, None); + assert_eq!(texts(&m), ["Hello"]); + assert_eq!(attempts, 1, "the shading is hashed once per walk"); + assert!(spent < Duration::from_secs(2), "1,000 sh took {spent:?}"); + + // The same bomb as an ExtGState transfer function behind 1,000 `gs`. + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let tr = d.b.add_stream( + "/FunctionType 4 /Domain [0 1] /Range [0 1] /Filter /FlateDecode", + &zlib_zero_bomb(33), + ); + d.page(PageSpec::new( + format!( + "{}BT /F1 12 Tf 72 700 Td (Hello) Tj ET", + "/GS1 gs\n".repeat(1_000) + ) + .as_bytes(), + &format!("/Font << /F1 {f} 0 R >> /ExtGState << /GS1 << /TR {tr} 0 R >> >>"), + )); + let m = model0(d.build()); + assert_eq!(m.page_reason, None); + assert_eq!( + take_decode_attempts().0, + 1, + "the /TR stream is hashed once per walk" + ); + + // 24 different bombs: a failed decode is charged its whole cap, so after the budget is spent + // the later attempts may inflate nothing (Σ caps ≤ budget + one stream cap). + let m = model0(bomb_shadings(24, 2)); + let (attempts, caps) = take_decode_attempts(); + assert_eq!(m.page_reason, None); + assert_eq!(attempts, 24); + assert!( + caps <= PAGE_DECODE_BUDGET + STREAM_MAX_DECODED, + "Σ caps {caps} B: failed decodes were not charged" + ); +} + +#[test] +fn deep_hash_steps_are_counted_per_walk() { + let array = |n: usize| format!("[{}]", "0 ".repeat(n)); + let gs = |n: usize| format!("<< /TR {} >>", array(n)); + set_steps_override(Some(1_000)); + // One 600-item direct value used 50 times: hashed once (memoised by address). + let m = model0(gs_page(&"/GS0 gs ".repeat(50), &[gs(600)])); + let one = m.page_reason; + // Two such values in one walk: 1,200 steps > 1,000, though each hash alone fits. + let m = model0(gs_page("/GS0 gs /GS1 gs", &[gs(600), gs(600)])); + set_steps_override(None); + assert_eq!(one, None, "a repeated value is hashed once"); + assert_eq!(m.page_reason, Some(R::PageTooComplex)); + assert_eq!(m.page_detail.as_deref(), Some("object hash budget")); +} + +#[test] +fn duplicate_marks_and_obstacles_stay_bounded() { + // 4,000 identical stacked runs: every one is DUPLICATE_TEXT, with one flag per run. + let mut c = String::new(); + for _ in 0..4_000 { + c.push_str("BT /F1 12 Tf 72 700 Td (a) Tj ET\n"); + } + let started = thread_cpu(); + let (m, peak) = thread_peak(|| model0(helvetica_page(c.as_bytes()))); + assert_eq!(m.runs.len(), 4_000); + assert!(m.runs.iter().all(|r| r.reason == Some(R::DuplicateText))); + assert!(peak < 64 * MIB, "4,000 duplicates peaked at {peak} B"); + // A chain: A overlaps B and B overlaps C by 70 %, A and C by 40 %; D stands apart. B is + // marked by A, then still finds C among the runs not marked yet. + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (a) Tj ET BT /F1 12 Tf 74 700 Td (a) Tj ET \ + BT /F1 12 Tf 76 700 Td (a) Tj ET BT /F1 12 Tf 200 700 Td (a) Tj ET", + )); + let reasons: Vec<(f64, Option)> = m.runs.iter().map(|r| (r.origin.0, r.reason)).collect(); + assert_eq!( + reasons, + [ + (72.0, Some(R::DuplicateText)), + (74.0, Some(R::DuplicateText)), + (76.0, Some(R::DuplicateText)), + (200.0, None) + ] + ); + // 3,000 runs on one baseline, 2 pt apart: the obstacle search is capped per page, and a run + // the cap cuts short gets the conservative obstacle at its origin. + let mut c = String::new(); + for i in 0..3_000 { + let ch = char::from(b'a' + (i % 26) as u8); + c.push_str(&format!("BT /F1 1 Tf {} 700 Td ({ch}) Tj ET\n", 10 + 2 * i)); + } + let m = model0(helvetica_page(c.as_bytes())); + assert_eq!(m.runs.len(), 3_000); + let first = m + .runs + .iter() + .find(|r| (r.origin.0 - 10.0).abs() < 1e-9) + .expect("first run"); + assert!( + first.next_obstacle.is_some_and(|d| (d - 2.0).abs() < 1e-6), + "the first run measures its neighbour: {:?}", + first.next_obstacle + ); + assert!(m.runs.iter().filter(|r| r.next_obstacle.is_none()).count() <= 1); + assert!( + m.runs.iter().any(|r| r.next_obstacle == Some(0.0)), + "9 M pairs reach the cap: the runs left get the conservative obstacle" + ); + assert!( + thread_cpu().saturating_sub(started) < Duration::from_secs(20), + "run checks" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds2.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds2.rs new file mode 100644 index 0000000..8e79deb --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds2.rs @@ -0,0 +1,348 @@ +//! Review regressions (T3 fix pass): model fidelity and the Phase B wrapper — a direct ExtGState +//! font (MEDIUM-2), text drawn after an advance the model cannot know (MEDIUM-3), a TJ opening +//! with a column-wide kern (LOW-1), the wrapper Form's dictionary and size (LOW-2, LOW-6), +//! pattern identity (LOW-3), unused fonts under their own budget (LOW-5), and a cancellable +//! classifier walk. + +use super::{content, ctx, model0, reason_of, texts, walk}; +use crate::pdf_engine::source_content::classify_source_page; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::engines::{run_tool, RunOpts}; +use crate::pdf_engine::text_edit::lexer::{lex_content, LexLimits, Operand, Operator}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::runs::build_page_model; +use crate::pdf_engine::text_edit::snapshot::read_snapshot; +use crate::pdf_engine::text_edit::state::same_paint; +use crate::pdf_engine::text_edit::testkit::pdf::zlib_zero_bomb; +use crate::pdf_engine::text_edit::testkit::producers::{ + cid_font, cid_hex, helvetica_doc, helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use crate::pdf_engine::text_edit::walker::{walk_page, PageWalk, WalkMode}; +use std::ffi::OsString; +use std::io::Write; +use std::sync::atomic::AtomicBool; + +#[test] +fn a_direct_extgstate_font_is_malformed() { + let gs_font = |base: &str| { + format!( + "<< /Font [<< /Type /Font /Subtype /Type1 /BaseFont /{base} \ + /Encoding /WinAnsiEncoding >> 12] >>" + ) + }; + let mut d = DocBuilder::new(); + d.page(PageSpec::new( + b"/GA gs BT 72 700 Td (iiii) Tj ET /GB gs BT 72 600 Td (iiii) Tj ET", + &format!( + "/ExtGState << /GA {} /GB {} >>", + gs_font("Helvetica"), + gs_font("Courier") + ), + )); + let m = model0(d.build()); + assert_eq!(m.page_reason, Some(R::MalformedContent)); + assert!(m + .page_detail + .as_deref() + .is_some_and(|d| d.contains("indirect"))); + // Indirect fonts keep their own models: Courier is wider than Helvetica. + let mut d = DocBuilder::new(); + let h = d.add(HELVETICA); + let c = + d.add("<< /Type /Font /Subtype /Type1 /BaseFont /Courier /Encoding /WinAnsiEncoding >>"); + d.page(PageSpec::new( + b"/GA gs BT 72 700 Td (iiii) Tj ET /GB gs BT 72 600 Td (iiii) Tj ET", + &format!("/ExtGState << /GA << /Font [{h} 0 R 12] >> /GB << /Font [{c} 0 R 12] >> >>"), + )); + let m = model0(d.build()); + let widths: Vec = m.runs.iter().map(|r| r.original_extent).collect(); + assert_eq!(widths.len(), 2); + assert!(widths[1] > widths[0] * 2.0, "{widths:?}"); +} + +/// `/F2` a 2-byte font drawing "A", `/F1` Helvetica, in one text object. +fn odd_then(between: &str) -> Vec { + let mut d = DocBuilder::new(); + let f2 = cid_font(&mut d.b, "ABCDEF+Arimo", "AB"); + let f1 = d.add(HELVETICA); + let content = format!( + "BT /F2 10 Tf 72 700 Td <{}00> Tj {between} /F1 10 Tf (Hello) Tj ET", + cid_hex("AB", "A") + ); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f1} 0 R /F2 {f2} 0 R >>"), + )); + d.build() +} + +#[test] +fn text_after_an_unknown_advance_is_refused_until_the_line_moves() { + // An odd-length string in a 2-byte font: viewers advance by CID 0, the model by nothing. + let m = model0(odd_then("")); + assert_eq!(reason_of(&m, "A"), Some(R::AmbiguousUnicode)); + assert_eq!(reason_of(&m, "Hello"), Some(R::MissingWidths)); + let m = model0(odd_then("0 -20 Td")); + assert_eq!( + reason_of(&m, "Hello"), + None, + "a line move makes the pen known again" + ); + // A missing font advances by an unknown amount too. + let missing = |between: &str| { + helvetica_page( + format!("BT /F9 10 Tf 72 700 Td (abc) Tj {between} /F1 10 Tf (Hello) Tj ET").as_bytes(), + ) + }; + assert_eq!( + reason_of(&model0(missing("")), "Hello"), + Some(R::MissingWidths) + ); + assert_eq!( + reason_of(&model0(missing("ET BT 72 650 Td")), "Hello"), + None + ); + // A code with a glyph name but no width (outside /FirstChar…/LastChar, no /MissingWidth): + // poppler uses Helvetica's own metrics for it, the model has none. + let partial = |shown: &str| { + let mut d = DocBuilder::new(); + let f = d.add( + "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding \ + /FirstChar 101 /LastChar 111 \ + /Widths [556 278 556 556 222 222 500 222 833 556 556] >>", + ); + d.page(PageSpec::new( + format!("BT /F1 12 Tf 72 700 Td {shown} ET").as_bytes(), + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() + }; + let m = model0(partial("(Hel) Tj (lo) Tj")); + assert!(!m.runs.is_empty()); + assert!( + m.runs.iter().all(|r| r.reason == Some(R::MissingWidths)), + "{:?}", + texts(&m) + ); + let m = model0(partial("(el) Tj (lo) Tj")); + assert_eq!(texts(&m), ["ello"]); + assert_eq!(reason_of(&m, "ello"), None); +} + +#[test] +fn a_tj_opening_with_a_column_kern_starts_a_new_run() { + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Left) Tj [-20000 (Right)] TJ ET", + )); + assert_eq!(texts(&m), ["Left", "Right"]); + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Left) Tj [-100 (Right)] TJ ET", + )); + assert_eq!(texts(&m), ["LeftRight"], "a small opening kern still joins"); +} + +/// A page painted only through `/Fx0` (qpdf's overlay wrapper shape) holding `data`. +fn wrapper_page(data: &[u8], filter: &str, form_extra: &str, page_extra: &str) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let form = d.b.add_stream( + &format!( + "/Type /XObject /Subtype /Form /BBox [0 0 612 792] \ + /Resources << /Font << /F1 {f} 0 R >> >> {filter} {form_extra}" + ), + data, + ); + d.page( + PageSpec::new( + b"q 1 0 0 1 0 0 cm /Fx0 Do Q", + &format!("/XObject << /Fx0 {form} 0 R >>"), + ) + .with(page_extra), + ); + d.build() +} + +fn wrapped(pdf: Vec) -> PageWalk { + let c = ctx(pdf); + let pc = content(&c, 0); + walk_page( + &c, + 0, + &pc, + WalkMode::Wrapped { + name: b"Fx0".to_vec(), + }, + None, + ) +} + +#[test] +fn the_wrapper_form_dictionary_holds_only_qpdf_keys() { + let hi = b"BT /F1 12 Tf 72 700 Td (Hi) Tj ET"; + let group = "/Group << /S /Transparency /CS /DeviceRGB >>"; + let ok = wrapped(wrapper_page(hi, "", "", "")); + assert_eq!(ok.page_reason, None, "{:?}", ok.page_detail); + assert_eq!(ok.records.len(), 1); + let same_group = wrapped(wrapper_page(hi, "", group, group)); + assert_eq!(same_group.page_reason, None, "{:?}", same_group.page_detail); + for (form_extra, page_extra, what) in [ + ("/OC << /Type /OCG /Name (L) >>", "", "/OC"), + ("/Ref << /F (x.pdf) /Page 0 >>", "", "/Ref"), + (group, "", "a /Group the page lacks"), + ( + group, + "/Group << /S /Transparency /CS /DeviceGray >>", + "a different /Group", + ), + ] { + let w = wrapped(wrapper_page(hi, "", form_extra, page_extra)); + assert_eq!(w.page_reason, Some(R::MalformedContent), "{what}"); + assert!( + w.page_detail + .as_deref() + .is_some_and(|d| d.starts_with("wrapper")), + "{what}: {:?}", + w.page_detail + ); + } +} + +#[test] +fn qpdf_copies_the_page_group_onto_its_wrapper() { + let Some(engines) = engines_or_skip("qpdf_copies_the_page_group_onto_its_wrapper") else { + return; + }; + let s = Scratch::new("wrapgroup"); + let group = "/Group << /S /Transparency /CS /DeviceRGB /I true >>"; + let src = s.write( + "src.pdf", + &helvetica_doc(b"BT /F1 12 Tf 72 700 Td (Grouped) Tj ET", "", group), + ); + let blank = s.write("blank.pdf", &helvetica_doc(b"", "", "")); + let out = s.path("out.pdf"); + let args: Vec = vec![ + src.into(), + "--overlay".into(), + blank.into(), + "--".into(), + out.clone().into(), + ]; + let r = run_tool(&engines.qpdf, &args, false, &RunOpts::default()).expect("qpdf --overlay"); + assert!(r.code == 0 || r.code == 3, "qpdf: {}", r.stderr); + let c = SnapshotContext::new(read_snapshot(&out).expect("overlay output opens")); + let pc = content(&c, 0); + let ops = lex_content(&pc.joined, &LexLimits::page(), None).expect("wrapper lexes"); + let name = ops + .iter() + .find(|o| o.operator == Operator::Do) + .and_then(|o| o.operands.first().and_then(Operand::as_name)) + .expect("the page paints a wrapper Form") + .to_vec(); + let w = walk_page(&c, 0, &pc, WalkMode::Wrapped { name }, None); + assert_eq!(w.page_reason, None, "{:?}", w.page_detail); + assert_eq!(w.records.len(), 1); +} + +/// A zlib stream inflating to `prefix` followed by `mib` MiB of NUL bytes (PDF whitespace). +fn zlib_padded(prefix: &[u8], mib: usize) -> Vec { + let zeros = vec![0u8; 1 << 20]; + let mut e = flate2::write::ZlibEncoder::new(Vec::new(), flate2::Compression::best()); + e.write_all(prefix).expect("write"); + e.write_all(&zeros).expect("write"); + e.flush().expect("flush"); + let first = e.get_ref().len(); + e.write_all(&zeros).expect("write"); + e.flush().expect("flush"); + let second = e.get_ref().len(); + let full = e.finish().expect("finish"); + let chunk = full[first..second].to_vec(); + let mut out = full[..second].to_vec(); + for _ in 2..mib.max(2) { + out.extend_from_slice(&chunk); + } + out.extend_from_slice(&full[second..]); + out +} + +#[test] +fn the_wrapper_form_gets_the_page_content_cap() { + // 40 MiB of content: over STREAM_MAX_DECODED (32 MiB), within PAGE_CONTENT_MAX_DECODED. + let data = zlib_padded(b"BT /F1 12 Tf 72 700 Td (Big) Tj ET\n", 40); + let w = wrapped(wrapper_page(&data, "/Filter /FlateDecode", "", "")); + assert_eq!(w.page_reason, None, "{:?}", w.page_detail); + assert_eq!(w.records.len(), 1); +} + +/// "Pat" filled with the axial shading pattern `/P0` from red to `c1`. +fn pattern_page(c1: &str) -> Vec { + helvetica_doc( + b"/Pattern cs /P0 scn BT /F1 12 Tf 72 700 Td (Pat) Tj ET", + &format!( + "/Pattern << /P0 << /PatternType 2 /Shading << /ShadingType 2 /ColorSpace /DeviceRGB \ + /Coords [0 0 1 0] /Function << /FunctionType 2 /Domain [0 1] /C0 [1 0 0] \ + /C1 [{c1}] /N 1 >> >> >> >>" + ), + "", + ) +} + +#[test] +fn a_pattern_swapped_behind_its_name_changes_the_paint() { + let fill = |pdf: Vec| { + let c = ctx(pdf); + walk(&c, 0, WalkMode::Edit).records[0].before.fill.clone() + }; + let (a, same, swapped) = ( + fill(pattern_page("0 0 1")), + fill(pattern_page("0 0 1")), + fill(pattern_page("0 1 0")), + ); + assert!(a.pattern && a.pattern_hash.is_some()); + assert!(same_paint(&a, &same)); + assert!(!same_paint(&a, &swapped), "same name, different pattern"); +} + +#[test] +fn unused_fonts_never_refuse_the_page() { + // Seven unused TrueType fonts whose programs inflate to 15 MiB each (105 MiB in all, over + // PAGE_DECODE_BUDGET): the page's own text stays editable, the fonts that fit are siblings. + let mut d = DocBuilder::new(); + let f1 = d.add(HELVETICA); + let bomb = zlib_zero_bomb(15); + let mut fonts = vec![format!("/F1 {f1} 0 R")]; + for k in 2..=8 { + let file = d.b.add_stream("/Filter /FlateDecode", &bomb); + let font = d.add(format!( + "<< /Type /Font /Subtype /TrueType /BaseFont /Big{k} /FirstChar 32 /LastChar 32 \ + /Widths [250] /FontDescriptor << /Type /FontDescriptor /FontName /Big{k} /Flags 32 \ + /FontBBox [0 0 1000 1000] /ItalicAngle 0 /Ascent 800 /Descent -200 /CapHeight 700 \ + /StemV 80 /FontFile2 {file} 0 R >> >>" + )); + fonts.push(format!("/F{k} {font} 0 R")); + } + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET", + &format!("/Font << {} >>", fonts.join(" ")), + )); + let m = model0(d.build()); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + assert_eq!(reason_of(&m, "Hello"), None); + let n = m.walk.page_fonts.len(); + assert!( + (2..8).contains(&n), + "F1 plus the unused fonts that fit: {n}" + ); +} + +#[test] +fn the_classify_walk_can_be_cancelled() { + let c = ctx(helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hello) Tj ET")); + let model = build_page_model(&c, 0, None).expect("model"); + let live = classify_source_page(&c, &model, None); + assert!(!live.occurrences.is_empty()); + let cancel = AtomicBool::new(true); + let stopped = classify_source_page(&c, &model, Some(&cancel)); + assert_eq!(stopped.occurrence_reason, Some(R::PageTooComplex)); + assert!(stopped.occurrences.is_empty()); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds3.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds3.rs new file mode 100644 index 0000000..6ad6200 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds3.rs @@ -0,0 +1,281 @@ +//! Review regressions (T3 fix pass 2): work and memory per walk stay bounded whatever a page +//! repeats — a long unpainted path (HIGH-1), many optional-content groups (MEDIUM-2), a page +//! model at the op cap (MEDIUM-3), a line alternating between many sibling fonts (LOW-1) — and +//! the Phase B wrapper's `/Group` is hashed apart from the content (LOW-3). + +use super::fuzz::thread_cpu; +use super::{content, ctx, model0, reason_of}; +use crate::pdf_engine::text_edit::limits::PAGE_MODEL_BYTES_MAX; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::testkit::pdf::zlib; +use crate::pdf_engine::text_edit::testkit::producers::{ + helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::thread_peak; +use crate::pdf_engine::text_edit::walker::hash::tests::set_steps_override; +use crate::pdf_engine::text_edit::walker::{walk_page, WalkMode}; +use std::sync::Arc; +use std::time::Duration; + +const MIB: usize = 1 << 20; + +#[test] +fn a_long_unpainted_path_costs_constant_work_per_op() { + // 200,000 rectangles (800,000 points) before one `n`: each op checks only its own points. + let mut c = "0 0 1 1 re ".repeat(200_000); + c.push_str("n BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + let started = thread_cpu(); + let m = model0(helvetica_page(c.as_bytes())); + let cpu = thread_cpu().saturating_sub(started); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + assert_eq!(reason_of(&m, "Hello"), None); + assert!(cpu < Duration::from_secs(2), "200,000 re took {cpu:?}"); + // Curves too (3 points per op). + let mut c = String::from("0 0 m "); + c.push_str(&"1 1 2 2 3 3 c ".repeat(100_000)); + c.push_str("S BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + let started = thread_cpu(); + let m = model0(helvetica_page(c.as_bytes())); + let cpu = thread_cpu().saturating_sub(started); + assert_eq!(reason_of(&m, "Hello"), None); + assert!(cpu < Duration::from_secs(2), "100,000 c took {cpu:?}"); + // A point that overflows is still refused, at the op that adds it (a CTM of scale 1e306, + // then x = 1e9). + let scale = "1000000000 0 0 1000000000 0 0 cm ".repeat(34); + let c = format!("{scale} 0 0 m 1 1 l 1000000000 1 l S"); + let m = model0(helvetica_page(c.as_bytes())); + assert_eq!(m.page_reason, Some(R::MalformedContent)); + assert!(m + .page_detail + .as_deref() + .is_some_and(|d| d.starts_with("non-finite path"))); +} + +/// `k` optional-content groups, all listed in `/OCGs` and `/D /ON` except the last, which is in +/// `/OFF`; each is opened once, and "Hello" is drawn inside group `inside`. +fn ocg_page(k: usize, inside: usize) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let ids: Vec = (0..k) + .map(|i| d.add(format!("<< /Type /OCG /Name (L{i}) >>"))) + .collect(); + let refs: Vec = ids.iter().map(|i| format!("{i} 0 R")).collect(); + let props: Vec = ids + .iter() + .enumerate() + .map(|(n, i)| format!("/P{n} {i} 0 R")) + .collect(); + let mut c = String::new(); + for n in 0..k { + if n == inside { + c.push_str(&format!( + "/OC /P{n} BDC BT /F1 12 Tf 72 700 Td (Hello) Tj ET EMC " + )); + } else { + c.push_str(&format!("/OC /P{n} BDC EMC ")); + } + } + let (on, off) = refs.split_at(k.saturating_sub(1)); + d.catalog_extra = format!( + "/OCProperties << /OCGs [{}] /D << /ON [{}] /OFF [{}] >> >>", + refs.join(" "), + on.join(" "), + off.join(" ") + ); + d.page(PageSpec::new( + c.as_bytes(), + &format!( + "/Font << /F1 {f} 0 R >> /Properties << {} >>", + props.join(" ") + ), + )); + d.build() +} + +#[test] +fn optional_content_states_cost_constant_work_per_group() { + // 20,000 distinct groups: the catalog's lists are read once, each lookup is a binary search + // (scanning /OCGs, /ON and /OFF per group took ~4.6 s in a debug build). + let pdf = ocg_page(20_000, 10_000); + let started = thread_cpu(); + let m = model0(pdf); + let cpu = thread_cpu().saturating_sub(started); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + assert_eq!(reason_of(&m, "Hello"), None, "a group on in /D is visible"); + assert!(cpu < Duration::from_secs(2), "20,000 groups took {cpu:?}"); + // The same lists still decide: the last group is in /OFF. + let m = model0(ocg_page(50, 49)); + assert_eq!(reason_of(&m, "Hello"), Some(R::OptionalContent)); +} + +#[test] +fn a_page_model_at_the_op_cap_shares_its_digests_and_reports_its_size() { + // 249,990 one-glyph show ops under one state (≈ 2.2 KB deflated): every record shares one + // digest and keeps no spare capacity. It peaked at ~700 MiB with two digests per record, + // ~360 MiB with shared digests alone and ~260 MiB after fix pass 2 (≈ 1 KB per record). + // [T3-budget] That is more than one page may hold (`PAGE_MODEL_BYTES_MAX`): the page is now + // refused before it holds more; 100,000 such ops are still modelled the same way. + let page = |n: usize| { + let mut c = String::from("BT /F1 1 Tf 72 700 Td "); + c.push_str(&"(a)Tj ".repeat(n)); + c.push_str("ET"); + c + }; + let c = page(249_990); + assert!(zlib(c.as_bytes()).len() < 4096); + let (m, peak) = thread_peak(|| model0(helvetica_page(c.as_bytes()))); + assert_eq!(m.page_reason, Some(R::PageTooComplex)); + assert_eq!(m.page_detail.as_deref(), Some("page model size")); + println!("op cap: refused, peak {} MiB", peak / MIB); + assert!( + peak < PAGE_MODEL_BYTES_MAX + 32 * MIB, + "the refused model peaked at {} MiB", + peak / MIB + ); + let c = page(100_000); + let started = thread_cpu(); + let (m, peak) = thread_peak(|| model0(helvetica_page(c.as_bytes()))); + let cpu = thread_cpu().saturating_sub(started); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + let records = &m.walk.records; + assert_eq!(records.len(), 100_000); + let (first, last) = (&records[0], &records[99_999]); + assert!(Arc::ptr_eq(&first.before, &first.after)); + assert!(Arc::ptr_eq(&first.before, &last.after)); + let approx = m.approx_bytes(); + println!( + "100,000 ops: peak {} MiB, approx_bytes {} MiB, charged {} MiB, cpu {cpu:?}", + peak / MIB, + approx / MIB, + m.walk.model_bytes / MIB + ); + assert!(peak < 128 * MIB, "the model peaked at {} MiB", peak / MIB); + assert!( + approx <= peak && approx >= peak / 2, + "approx_bytes {approx} B against a peak of {peak} B" + ); + assert!(approx <= m.walk.model_bytes && m.walk.model_bytes <= PAGE_MODEL_BYTES_MAX); +} + +#[test] +fn digests_are_shared_only_while_the_state_is_unchanged() { + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (a) Tj (b) Tj 1 0 0 rg (c) Tj 0 g (d) Tj 0 g (e) Tj ET \ + 0 0 10 10 re f 0 0 10 10 re f", + )); + let r = &m.walk.records; + assert_eq!(r.len(), 5); + assert!(Arc::ptr_eq(&r[0].after, &r[1].before), "unchanged state"); + assert!(!Arc::ptr_eq(&r[1].after, &r[2].before), "a colour change"); + assert_eq!(r[2].before.fill.hex().as_deref(), Some("#ff0000")); + // `0 g` twice: equal values under new op bytes get a new digest, equal by value. + assert!(!Arc::ptr_eq(&r[3].before, &r[4].before)); + assert_eq!(*r[3].before, *r[4].before); + assert_eq!(r[1].before.fill.hex().as_deref(), Some("#000000")); + let p = &m.walk.paints; + assert_eq!(p.len(), 2); + assert!(Arc::ptr_eq(&p[0].state, &p[1].state), "paints share too"); + assert!(m.approx_bytes() > 0); +} + +#[test] +fn a_line_alternating_between_sibling_fonts_builds_each_group_once() { + // 256 Helvetica resources (all siblings), alternating on one line 64,000 times: each + // resource's sibling group is built once per page (it was once per join candidate, O(256²) + // each). + let mut d = DocBuilder::new(); + let names: Vec = (0..256) + .map(|k| format!("/F{k} {} 0 R", d.add(HELVETICA))) + .collect(); + let mut c = String::from("BT 72 700 Td "); + for i in 0..64_000 { + c.push_str(&format!("/F{} 1 Tf (a)Tj ", i % 256)); + } + c.push_str("ET"); + d.page(PageSpec::new( + c.as_bytes(), + &format!("/Font << {} >>", names.join(" ")), + )); + let pdf = d.build(); + let started = thread_cpu(); + let m = model0(pdf); + let cpu = thread_cpu().saturating_sub(started); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + assert!( + cpu < Duration::from_secs(4), + "64,000 font switches took {cpu:?}" + ); + // Siblings still join (64,000 single-member runs would exceed RUNS_PER_PAGE_MAX). + assert!(m.runs.len() < 1_000, "{} runs", m.runs.len()); + assert!(m.runs.iter().all(|r| r.members.len() > 1)); +} + +/// A wrapper page whose `/Group` and the wrapper Form's both name one object holding `pad` +/// numbers, and whose content paints an image whose dictionary holds `pad` numbers too. +fn wrapper_with_group(pad: usize) -> Vec { + let numbers = vec!["1"; pad].join(" "); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let group = d.add(format!( + "<< /S /Transparency /CS /DeviceRGB /Pad [{numbers}] >>" + )); + let image = d.b.add_stream( + &format!( + "/Type /XObject /Subtype /Image /Width 1 /Height 1 /ColorSpace /DeviceGray \ + /BitsPerComponent 8 /Pad [{numbers}]" + ), + &[0], + ); + let form = d.b.add_stream( + &format!( + "/Type /XObject /Subtype /Form /BBox [0 0 612 792] /Group {group} 0 R \ + /Resources << /Font << /F1 {f} 0 R >> /XObject << /Im0 {image} 0 R >> >>" + ), + b"q 10 0 0 10 300 300 cm /Im0 Do Q BT /F1 12 Tf 72 700 Td (Hi) Tj ET", + ); + d.page( + PageSpec::new( + b"q 1 0 0 1 0 0 cm /Fx0 Do Q", + &format!("/XObject << /Fx0 {form} 0 R >>"), + ) + .with(&format!("/Group {group} 0 R")), + ); + d.build() +} + +#[test] +fn the_wrapper_group_is_hashed_apart_from_the_content() { + // With a 1,000-step hash budget, the /Group (~600 steps) and the image (~600 steps) each + // fit, but not together: the wrapper check must not spend the content's budget. + let c = ctx(wrapper_with_group(600)); + let pc = content(&c, 0); + set_steps_override(Some(1_000)); + let w = walk_page( + &c, + 0, + &pc, + WalkMode::Wrapped { + name: b"Fx0".to_vec(), + }, + None, + ); + set_steps_override(None); + assert_eq!(w.page_reason, None, "{:?}", w.page_detail); + assert_eq!(w.records.len(), 1); + assert_eq!(w.paints.len(), 1); + // The group alone over the budget still refuses. + let c = ctx(wrapper_with_group(1_200)); + let pc = content(&c, 0); + set_steps_override(Some(1_000)); + let w = walk_page( + &c, + 0, + &pc, + WalkMode::Wrapped { + name: b"Fx0".to_vec(), + }, + None, + ); + set_steps_override(None); + assert_eq!(w.page_reason, Some(R::PageTooComplex)); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds4.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds4.rs new file mode 100644 index 0000000..e95ec21 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/bounds4.rs @@ -0,0 +1,407 @@ +//! Review regressions (T3 fix pass 3, review T3 r2): what a page model keeps per element is +//! budgeted per page — TJ kerns (HIGH-1), lexed operand nodes alive at once (HIGH-1, page content +//! and descended Forms), characters of glyph text (HIGH-2) — and `approx_bytes` counts every +//! shared allocation a digest holds, while a state re-set to the value in force keeps sharing its +//! digest (MEDIUM-2: the review's P13, P14 and P15 pages, smaller). + +use super::fuzz::thread_cpu; +use super::{ctx, model, model0, reason_of, walk}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::runs::PageModel; +use crate::pdf_engine::text_edit::testkit::fonts::{ + add_type0, tounicode_bfchar, Program, Type0Font, +}; +use crate::pdf_engine::text_edit::testkit::pdf::zlib; +use crate::pdf_engine::text_edit::testkit::producers::{ + helvetica_doc, helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::thread_peak; +use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; +use crate::pdf_engine::text_edit::walker::{ + WalkMode, KERNS_PER_PAGE_MAX, OPERAND_NODES_MAX, TEXT_CHARS_PER_PAGE_MAX, +}; +use std::sync::Arc; +use std::time::Duration; + +const MIB: usize = 1 << 20; + +/// Bytes `f`'s result still holds when it returns: the thread's peak above an untouched 1 GiB +/// reservation made after `f` (larger than every build peak of these tests, which assert their +/// peaks far below it, so the peak is the held bytes plus the reservation). The snapshot context +/// must be made outside `f`: its parse allocates on other threads, and freeing that here would +/// count as negative bytes. +fn held_by(f: impl FnOnce() -> T) -> (T, usize) { + const PAD: usize = 1 << 30; + let ((value, pad), peak) = thread_peak(|| { + let value = f(); + let pad: Vec = Vec::with_capacity(PAD); + (value, pad) + }); + drop(pad); + assert!(peak >= PAD, "peak {peak} below the reservation"); + (value, peak - PAD) +} + +/// `approx_bytes` lies in [held / 1.25, held], give or take an eighth of the font models it +/// counts: since the fix pass of 2026-10-03 a model counts the fonts it keeps alive, by +/// `FontModel::approx_bytes`, an estimate (Helvetica's is ~2 KB above what it allocates). +fn assert_size(tag: &str, m: &PageModel, held: usize) { + let approx = m.approx_bytes(); + let mut seen = std::collections::HashSet::new(); + let fonts: usize = m + .walk + .page_fonts + .iter() + .map(|(_, f)| f) + .chain(m.walk.records.iter().filter_map(|r| r.font.as_ref())) + .filter(|f| seen.insert(std::sync::Arc::as_ptr(f) as usize)) + .map(|f| f.approx_bytes()) + .sum(); + println!( + "{tag}: held {} KiB, approx_bytes {} KiB (fonts {} KiB)", + held >> 10, + approx >> 10, + fonts >> 10 + ); + assert!( + approx <= held + fonts / 8 && approx >= held / 5 * 4, + "{tag}: approx_bytes {approx} B against {held} B held" + ); +} + +/// `ops` TJ ops of one glyph followed by `kerns` kerns of 1/1000 em each (to the right). +fn kern_page(ops: usize, kerns: usize) -> Vec { + let arr = format!("[(a){}] TJ\n", " -1".repeat(kerns)); + helvetica_page(format!("BT /F1 1 Tf 72 700 Td {}ET", arr.repeat(ops)).as_bytes()) +} + +#[test] +fn tj_kerns_are_charged_to_a_page_budget() { + assert_eq!(KERNS_PER_PAGE_MAX, 400_000); + // Review P17, smaller: 7 TJ arrays of 65,535 kerns (458,745) from ~4 KB of deflated + // content. Uncounted, P17's 374 arrays held 4.3 GiB (~184 B per kern). + let pdf = kern_page(7, 65_535); + let c = ctx(pdf); + let (m, peak) = thread_peak(|| model(&c, 0)); + assert_eq!(m.page_reason, Some(R::PageTooComplex)); + assert_eq!(m.page_detail.as_deref(), Some("TJ kerns per page")); + assert!(m.walk.ops.is_empty() && m.walk.records.is_empty()); + assert!( + peak < 48 * MIB, + "a refused kern page peaked at {} MiB", + peak / MIB + ); + // The inspect path's Classify walk is refused the same way. + let w = walk(&c, 0, WalkMode::Classify); + assert_eq!(w.page_reason, Some(R::PageTooComplex)); + assert_eq!(w.page_detail.as_deref(), Some("TJ kerns per page")); + // Under the budget (6 × 65,535 = 393,210) the page is modelled and its size reported. + let c = ctx(kern_page(6, 65_535)); + let ((m, cpu), held) = held_by(|| { + let started = thread_cpu(); + let m = model(&c, 0); + (m, thread_cpu().saturating_sub(started)) + }); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + // One line: six glyphs, 65.5 em apart (each gap reads as one synthetic space). + assert_eq!(reason_of(&m, "a a a a a a"), None); + assert!(held < 96 * MIB, "393,210 kerns hold {} MiB", held / MIB); + assert!(cpu < Duration::from_secs(4), "393,210 kerns took {cpu:?}"); + assert_size("393,210 kerns", &m, held); +} + +/// `lines` × `[1 1 … 1] 0 d 0 0 1 1 re f` (65,542 operand nodes each), then "Hello". +fn dash_lines(lines: usize) -> String { + let line = format!("[{}] 0 d 0 0 1 1 re f\n", vec!["1"; 65_536].join(" ")); + line.repeat(lines) +} + +#[test] +fn lexed_operand_nodes_are_charged_to_a_page_budget() { + assert_eq!(OPERAND_NODES_MAX, 1_500_000); + // Review P16, smaller: 23 dash arrays of 65,536 numbers (1,507,466 nodes). Kept, P16's 748 + // arrays held 1.26 GiB in `PageWalk::ops`. + let content = format!("{}BT /F1 12 Tf 72 700 Td (Hello) Tj ET", dash_lines(23)); + assert!(zlib(content.as_bytes()).len() < 16 << 10); + let c = ctx(helvetica_page(content.as_bytes())); + let (m, held) = held_by(|| model(&c, 0)); + assert_eq!(m.page_reason, Some(R::PageTooComplex)); + assert_eq!(m.page_detail.as_deref(), Some("content operands")); + assert!( + held < content.len() + MIB, + "a refused page holds {} KiB", + held >> 10 + ); + // 22 arrays (1,441,929 nodes) fit, and `approx_bytes` counts them. + let content = format!("{}BT /F1 12 Tf 72 700 Td (Hello) Tj ET", dash_lines(22)); + let c = ctx(helvetica_page(content.as_bytes())); + let (m, held) = held_by(|| model(&c, 0)); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + assert_eq!(reason_of(&m, "Hello"), None); + assert_size("1,441,929 operand nodes", &m, held); +} + +/// A page painting Form `/Fa` (`a` dash lines; it paints `/Fb` when `nested`) `paints` times, +/// and Form `/Fb` (`b` dash lines) `paints` more times unless nested. +fn form_page(a: usize, b: usize, nested: bool, paints: usize) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let fb = d.b.add_flate( + "/Type /XObject /Subtype /Form /BBox [0 0 612 792]", + dash_lines(b).as_bytes(), + ); + let inner = if nested { "/Fb Do" } else { "" }; + let fa = d.b.add_flate( + &format!( + "/Type /XObject /Subtype /Form /BBox [0 0 612 792] \ + /Resources << /XObject << /Fb {fb} 0 R >> >>" + ), + format!("{}{inner}", dash_lines(a)).as_bytes(), + ); + let mut content = "/Fa Do ".repeat(paints); + if !nested { + content.push_str(&"/Fb Do ".repeat(paints)); + } + content.push_str("BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f} 0 R >> /XObject << /Fa {fa} 0 R /Fb {fb} 0 R >>"), + )); + d.build() +} + +#[test] +fn form_operands_count_only_while_the_form_runs() { + // Classify descends Forms: a Form's lexed ops live while it runs, nested Forms on top of + // their parents. Two nested Forms of 12 dash lines each (786,504 nodes each) are refused. + let c = ctx(form_page(12, 12, true, 1)); + let w = walk(&c, 0, WalkMode::Classify); + assert_eq!(w.page_reason, Some(R::PageTooComplex)); + assert_eq!(w.page_detail.as_deref(), Some("content operands")); + // Edit mode paints Forms without descending. + assert_eq!(walk(&c, 0, WalkMode::Edit).page_reason, None); + // The same two Forms side by side, each painted twice: never more than one alive. + let c = ctx(form_page(12, 12, false, 2)); + let w = walk(&c, 0, WalkMode::Classify); + assert_eq!(w.page_reason, None, "{:?}", w.page_detail); + assert_eq!(w.paints.len(), 4 * (1 + 12)); +} + +/// An embedded Identity-H TrueType font whose ToUnicode maps CID 1 to 256 × "A" (the longest +/// destination T2 reads), drawn `glyphs` times in Tj ops of at most 1,000 glyphs. +fn long_text_page(glyphs: usize) -> Vec { + let mut t = TtfBuilder::new(); + t.unicode_glyph('A', "A", true); + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Long"); + f.program = Program::TrueType(t.build()); + f.cid_to_gid = Some(None); + f.flags = Some(32); + f.w = Some("[1 [500]]".to_string()); + let long = "A".repeat(256); + f.tounicode = Some(tounicode_bfchar(&[(1, 2, long.as_str())])); + let mut d = DocBuilder::new(); + let f0 = add_type0(&mut d.b, &f); + let mut c: Vec = b"BT /F0 1 Tf 72 700 Td ".to_vec(); + let mut left = glyphs; + while left > 0 { + let n = left.min(1_000); + c.push(b'('); + for _ in 0..n { + c.extend_from_slice(&[0, 1]); + } + c.extend_from_slice(b") Tj "); + left -= n; + } + c.extend_from_slice(b"ET"); + d.page(PageSpec::new(&c, &format!("/Font << /F0 {f0} 0 R >>"))); + d.build() +} + +#[test] +fn glyph_text_is_charged_to_a_page_budget() { + assert_eq!(TEXT_CHARS_PER_PAGE_MAX, 800_000); + // Review P25, smaller: 40,000 glyphs of 256 characters each (10 M characters). Uncounted, + // P25's 399,000 glyphs held 1.45 GiB (≈ 3.8 KB per glyph). + let c = ctx(long_text_page(40_000)); + let (m, peak) = thread_peak(|| model(&c, 0)); + assert_eq!(m.page_reason, Some(R::PageTooComplex)); + assert_eq!(m.page_detail.as_deref(), Some("text characters per page")); + assert!( + peak < 16 * MIB, + "a refused page peaked at {} MiB", + peak / MIB + ); + // 3,125 glyphs are exactly 800,000 characters; one more is refused. + let m = model0(long_text_page(3_126)); + assert_eq!(m.page_detail.as_deref(), Some("text characters per page")); + let c = ctx(long_text_page(3_125)); + let (m, held) = held_by(|| model(&c, 0)); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + let chars: usize = m.runs.iter().map(|r| r.text.chars().count()).sum(); + assert_eq!(chars, 800_000); + assert!( + held < 32 * MIB, + "800,000 characters hold {} MiB", + held / MIB + ); + assert_size("800,000 characters", &m, held); +} + +/// `n` repetitions of `each` inside one text object (after `prefix`, before `suffix`). +fn repeated(prefix: &str, each: &str, n: usize, suffix: &str) -> String { + format!("{prefix}BT /F1 1 Tf 72 700 Td {}ET{suffix}", each.repeat(n)) +} + +#[test] +fn a_state_reset_to_its_value_keeps_its_digest_and_approx_bytes_counts_shared_parts() { + const N: usize = 40_000; + // Review P13: a new colour op before every show op, so a new digest per record. + let c = ctx(helvetica_page(repeated("", "0 g (a)Tj ", N, "").as_bytes())); + let (p13, held13) = held_by(|| model(&c, 0)); + assert_eq!(p13.page_reason, None, "{:?}", p13.page_detail); + assert_size("P13", &p13, held13); + // Review P14: one ExtGState of 63 unmodelled keys re-applied before every show op. Its key + // list is kept, so every record shares one digest (it was a new digest and a new 63-entry + // list per op: 2.5 × P13). + let keys: Vec = (0..63).map(|k| format!("/Key{k:02} {k}")).collect(); + let gs = format!( + "/ExtGState << /G << /Type /ExtGState {} >> >>", + keys.join(" ") + ); + let pdf = helvetica_doc(repeated("", "/G gs (a)Tj ", N, "").as_bytes(), &gs, ""); + let c = ctx(pdf); + let (p14, held14) = held_by(|| model(&c, 0)); + let r = &p14.walk.records; + assert_eq!(r.len(), N); + assert_eq!(r[0].before.gs.other.len(), 63); + assert!(Arc::ptr_eq(&r[0].before, &r[N - 1].after)); + assert!(held14 < held13, "P14 holds {held14} B, P13 {held13} B"); + assert_size("P14", &p14, held14); + // Review P15: a 63-deep marked-content stack, reopened around every show op. The stack is + // equal to the last digest's, so it is that digest's stack again (it was a copy per op). + let open = "/X BMC ".repeat(63); + let close = " EMC".repeat(63); + let c = ctx(helvetica_page( + repeated(&open, "EMC /X BMC (a)Tj ", N, &close).as_bytes(), + )); + let (p15, held15) = held_by(|| model(&c, 0)); + let r = &p15.walk.records; + assert_eq!(r[0].before.marked.len(), 63); + assert!(Arc::ptr_eq(&r[0].before, &r[N - 1].after)); + assert!(held15 < held13, "P15 holds {held15} B, P13 {held13} B"); + assert_size("P15", &p15, held15); + // The stack is shared even when the rest of the state changes on every op. + let c = ctx(helvetica_page( + repeated(&open, "EMC /X BMC 0 g (a)Tj ", N, &close).as_bytes(), + )); + let (m, held_mixed) = held_by(|| model(&c, 0)); + let r = &m.walk.records; + assert!(!Arc::ptr_eq(&r[0].before, &r[1].before)); + assert!(r[0].before.marked.same(&r[N - 1].before.marked)); + assert_size("P15 with a colour per op", &m, held_mixed); + // `BM`, `RI`, `/D`, `ri` and `d` re-set to the values in force keep sharing too. + let gs = "/ExtGState << /G << /BM /Multiply /RI /Perceptual /D [[3 1] 0] /CA 0.5 >> >>"; + let each = "/G gs /Perceptual ri [3 1] 0 d (a)Tj "; + let pdf = helvetica_doc(repeated("", each, 2_000, "").as_bytes(), gs, ""); + let m = model0(pdf); + let r = &m.walk.records; + assert_eq!(r[0].before.gs.blend.to_vec(), b"Multiply".to_vec()); + assert_eq!(r[0].before.gs.dash.0.to_vec(), vec![3.0, 1.0]); + assert!(Arc::ptr_eq(&r[0].before, &r[1_999].after)); +} + +/// A stack `depth` deep whose innermost tag differs on every one of `n` show ops. +fn new_tag_per_op(depth: usize, n: usize) -> Vec { + let mut c = "/X BMC ".repeat(depth); + c.push_str("BT /F1 1 Tf 72 700 Td "); + for i in 0..n { + c.push_str(&format!("EMC /T{i} BMC (a)Tj ")); + } + c.push_str("ET"); + c.push_str(&" EMC".repeat(depth)); + helvetica_page(c.as_bytes()) +} + +#[test] +fn a_marked_stack_reopened_with_a_new_tag_costs_one_node_per_op() { + // A 63-deep stack whose innermost tag differs on every show op. As a copied vector each + // record held a 63-entry stack of its own (+2.5 KB per record, ~500 MiB at the op cap); as a + // persistent list it holds one new node and shares the 62 below, so the page holds what a + // 1-deep stack holds. `approx_bytes` counts each node once (counting only the digests, it + // fell to ~55 % of the copied-vector page). + const N: usize = 20_000; + let cx = ctx(new_tag_per_op(63, N)); + let (m, held) = held_by(|| model(&cx, 0)); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + let r = &m.walk.records; + assert_eq!(r.len(), N); + let (first, last) = (&r[0].before.marked, &r[N - 1].before.marked); + assert_eq!(first.len(), 63); + assert!(!first.same(last)); + assert_eq!( + first.iter().next().map(|e| e.tag.to_vec()), + Some(b"T0".to_vec()) + ); + let (mut a, mut b) = (first.clone(), last.clone()); + a.pop(); + b.pop(); + assert!(a.same(&b), "the 62 entries below are shared"); + assert_size("a new tag per op, 63 deep", &m, held); + let shallow = ctx(new_tag_per_op(1, N)); + let (m1, held1) = held_by(|| model(&shallow, 0)); + assert_eq!(m1.page_reason, None, "{:?}", m1.page_detail); + println!("a new tag per op, 1 deep: held {} KiB", held1 >> 10); + assert!( + held < held1 + (64 << 10), + "63 deep holds {held} B, 1 deep {held1} B" + ); +} + +#[test] +fn alternating_extgstates_share_their_key_lists() { + // 63 unmodelled keys in force, then two ExtGStates that each set one more key to a value of + // their own, alternating before every show op: each `gs` changes the state, but the merged + // key list is one of two, shared (it was a new 64-entry list per op, +1.5 KB per record). + const N: usize = 20_000; + let keys: Vec = (0..63).map(|k| format!("/Key{k:02} {k}")).collect(); + let gs = format!( + "/ExtGState << /Base << {} >> /A << /Odd 1 >> /B << /Odd 2 >> >>", + keys.join(" ") + ); + let content = repeated("/Base gs ", "/A gs (a)Tj /B gs (a)Tj ", N / 2, ""); + let cx = ctx(helvetica_doc(content.as_bytes(), &gs, "")); + let (m, held) = held_by(|| model(&cx, 0)); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + let r = &m.walk.records; + assert_eq!(r.len(), N); + assert_eq!(r[0].before.gs.other.len(), 64); + assert!(!Arc::ptr_eq(&r[0].before.gs.other, &r[1].before.gs.other)); + assert!(Arc::ptr_eq( + &r[0].before.gs.other, + &r[N - 2].before.gs.other + )); + assert!(Arc::ptr_eq( + &r[1].before.gs.other, + &r[N - 1].before.gs.other + )); + assert_size("alternating ExtGStates", &m, held); +} + +#[test] +fn approx_bytes_counts_the_verbatim_bytes_digests_hold() { + // Each colour op carries a 2,000-byte comment inside its span, and each record's digest + // holds those verbatim bytes (for B14 restores): about a third of what the model holds. + const N: usize = 10_000; + let each = format!("0 %{}\n g (a)Tj ", "x".repeat(2_000)); + let cx = ctx(helvetica_page(repeated("", &each, N, "").as_bytes())); + let (m, held) = held_by(|| model(&cx, 0)); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + let op = m.walk.records[0] + .before + .fill + .color_op + .as_ref() + .map(|b| b.len()); + assert!(op.is_some_and(|n| n > 2_000), "{op:?}"); + assert_size("2,000-byte colour ops", &m, held); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/budget.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/budget.rs new file mode 100644 index 0000000..b8ed76b --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/budget.rs @@ -0,0 +1,497 @@ +//! Review regressions (T3 budget pass): every amplification probe of review rounds 1–3 — a page of +//! a few KB deflated that built a page model of hundreds of MiB to GiBs — is refused +//! `PAGE_TOO_COMPLEX` or holds far less, in `build_page_model` and in `classify_source_page`, and +//! so are amplifications no review listed (a long font name copied per glyph unit, interned +//! ExtGState lists with no text, a Form only the Classify walk descends): the page-model budget +//! (`PAGE_MODEL_BYTES_MAX`) is charged before every allocation that grows with the page. Normal +//! pages hold a few MiB and `approx_bytes` never exceeds what their build charged. +//! +//! Peaks are the calling thread's allocation peak (`thread_peak`); the snapshot is parsed +//! outside it (lopdf parses on other threads). + +use super::fuzz::thread_cpu; +use super::{ctx, model, reason_of}; +use crate::pdf_engine::source_content::{classify_source_page, SourcePageResult}; +use crate::pdf_engine::text_edit::lexer::{ + lex_content, scan_tokens, LexError, LexLimits, ScanMode, OPERAND_NODES, +}; +use crate::pdf_engine::text_edit::limits::{set_model_bytes_override, PAGE_MODEL_BYTES_MAX}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::runs::PageModel; +use crate::pdf_engine::text_edit::testkit::fonts::{ + add_type0, tounicode_bfchar, Program, Type0Font, +}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_doc, helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::ttf::TtfBuilder; +use crate::pdf_engine::text_edit::testkit::{thread_peak, thread_peak_held}; +use std::sync::Arc; +use std::time::Duration; + +const MIB: usize = 1 << 20; +/// What a page's build or Classify pass may peak at (§H R20's per-file budget). +const PEAK_MAX: usize = 256 * MIB; +/// The budget plus what is allowed outside it (one stream decoded for a hash, a lexed stream's +/// string copies): every page refused by the budget peaks below this. +const REFUSED_PEAK_MAX: usize = PAGE_MODEL_BYTES_MAX + 32 * MIB; + +struct Probe { + model: PageModel, + model_peak: usize, + classify: SourcePageResult, + classify_peak: usize, + cpu: Duration, +} + +/// Builds page 0's model and classifies it, measuring each; both stay below `PEAK_MAX`, and so +/// does the inspect as a whole — the model held while its Classify pass runs (review T3-budget +/// MEDIUM-2: the two were measured apart, each from zero) — and a model never holds more than the +/// budget or than its charge. +fn probe(tag: &str, pdf: Vec) -> Probe { + let c = ctx(pdf); + let started = thread_cpu(); + let (m, model_peak, held) = thread_peak_held(|| model(&c, 0)); + let (classify, classify_peak) = thread_peak(|| classify_source_page(&c, &m, None)); + let cpu = thread_cpu().saturating_sub(started); + println!( + "{tag}: model peak {} MiB, held {} MiB, charged {} MiB, approx_bytes {} MiB, {:?} {:?}; \ + classify peak {} MiB, {:?}, {} occurrences; cpu {cpu:?}", + model_peak / MIB, + held / MIB, + m.walk.model_bytes / MIB, + m.approx_bytes() / MIB, + m.page_reason, + m.page_detail, + classify_peak / MIB, + classify.occurrence_reason, + classify.occurrences.len() + ); + assert!(model_peak < PEAK_MAX, "{tag}: model peak {model_peak} B"); + assert!( + held.saturating_add(classify_peak) < PEAK_MAX, + "{tag}: inspect peak {held} + {classify_peak} B" + ); + assert!( + m.approx_bytes() <= m.walk.model_bytes, + "{tag}: approx above the charge" + ); + assert!( + m.walk.model_bytes <= PAGE_MODEL_BYTES_MAX, + "{tag}: charge above the budget" + ); + Probe { + model: m, + model_peak, + classify, + classify_peak, + cpu, + } +} + +/// The page is refused `PAGE_TOO_COMPLEX` with `detail`, and so is every occurrence its +/// Classify pass lists (none when that pass is refused too). +fn assert_refused(tag: &str, p: &Probe, detail: &str) { + assert_eq!(p.model.page_reason, Some(R::PageTooComplex), "{tag}"); + assert_eq!(p.model.page_detail.as_deref(), Some(detail), "{tag}"); + assert!(p.model.runs.is_empty(), "{tag}"); + let c = &p.classify; + assert!( + c.occurrence_reason == Some(R::PageTooComplex) + || (c.occurrence_reason.is_none() + && c.occurrences + .iter() + .all(|o| o.reason == Some(R::PageTooComplex))), + "{tag}: {:?}", + c.occurrence_reason + ); + assert!(p.model_peak < REFUSED_PEAK_MAX, "{tag}: {} B", p.model_peak); + assert!( + p.classify_peak < REFUSED_PEAK_MAX, + "{tag}: {} B", + p.classify_peak + ); +} + +/// Review r2 P17 and r3 P5/P6 at full scale: 374 arrays of 65,536 numbers in two content parts +/// (49 MB decoded), shown as TJ kerns or set as dash arrays. +fn number_arrays(tj: bool) -> Vec { + let line = if tj { + format!("[(a){}] TJ\n", " 1".repeat(65_535)) + } else { + format!("[{}] 0 d\n", vec!["1"; 65_536].join(" ")) + }; + let p1 = format!("BT /F1 1 Tf 72 700 Td {}", line.repeat(187)); + let p2 = format!("{}ET", line.repeat(187)); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts( + &[p1.as_bytes(), p2.as_bytes()], + &format!("/Font << /F1 {f} 0 R >>"), + )); + d.build() +} + +#[test] +fn number_array_bombs_stop_in_the_lexer() { + // They peaked at 6.3 GiB (round 2), then 1,169 MiB inside the lexer (round 3 MEDIUM-3): the + // lexer now stops at the operand nodes the page may hold. + for (tag, tj) in [("P17 TJ kerns", true), ("P16 dash arrays", false)] { + let p = probe(tag, number_arrays(tj)); + assert_refused(tag, &p, "content operands"); + assert!(p.cpu < Duration::from_secs(20), "{tag}: {:?}", p.cpu); + } +} + +#[test] +fn the_lexer_counts_operand_nodes_as_it_reads_them() { + let lim = |n: usize| LexLimits { + operand_nodes_max: n, + ..LexLimits::page() + }; + // `[1 2 3] 0 d` is 5 nodes (the array, three numbers, the phase). + let src = b"[1 2 3] 0 d"; + assert!(lex_content(src, &lim(5), None).is_ok()); + assert_eq!( + lex_content(src, &lim(4), None), + Err(LexError::TooComplex { + what: OPERAND_NODES + }) + ); + // Dictionary keys and values count, inline-image dictionaries too. + let bdc = b"/P <> BDC EMC"; + assert!(lex_content(bdc, &lim(4), None).is_ok()); + assert!(lex_content(bdc, &lim(3), None).is_err()); + let image = b"BI /W 1 /H 1 /BPC 8 /CS /G ID \x00 EI"; + assert!(lex_content(image, &lim(8), None).is_ok()); + assert!(lex_content(image, &lim(7), None).is_err()); + // The page default is the walker's cap; token scans are bounded by their own token count. + assert_eq!(LexLimits::page().operand_nodes_max, 1_500_000); + let many = "1 ".repeat(1_600_000); + assert!(scan_tokens(many.as_bytes(), ScanMode::Object, 2_000_000).is_ok()); +} + +/// An embedded Identity-H font whose ToUnicode maps CID 1 to 256 × "A", drawn `glyphs` times. +fn long_text_page(glyphs: usize) -> Vec { + let mut t = TtfBuilder::new(); + t.unicode_glyph('A', "A", true); + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Long"); + f.program = Program::TrueType(t.build()); + f.cid_to_gid = Some(None); + f.flags = Some(32); + f.w = Some("[1 [500]]".to_string()); + let long = "A".repeat(256); + f.tounicode = Some(tounicode_bfchar(&[(1, 2, long.as_str())])); + let mut d = DocBuilder::new(); + let f0 = add_type0(&mut d.b, &f); + let mut c: Vec = b"BT /F0 1 Tf 72 700 Td ".to_vec(); + for _ in 0..glyphs / 1_000 { + c.push(b'('); + for _ in 0..1_000 { + c.extend_from_slice(&[0, 1]); + } + c.extend_from_slice(b") Tj "); + } + c.extend_from_slice(b"ET"); + d.page(PageSpec::new(&c, &format!("/Font << /F0 {f0} 0 R >>"))); + d.build() +} + +#[test] +fn the_tounicode_text_bomb_stops_at_its_character_budget() { + // Review r2 P25 at full scale: 399,000 glyphs of 256 characters (1.45 GiB before fix pass 3). + let p = probe("P25 text", long_text_page(399_000)); + assert_refused("P25", &p, "text characters per page"); + assert!(p.model_peak < 32 * MIB, "{} B", p.model_peak); +} + +/// Review r3 P1: `names` font resources (separate Helvetica objects; `/F1` and names +/// `name_len` bytes long), and `lines` one-glyph lines in `/F1`, one run each. +fn surface_page(names: usize, name_len: usize, lines: usize) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let mut fonts = format!("/F1 {f} 0 R"); + for i in 0..names - 1 { + let fi = d.add(HELVETICA); + fonts.push_str(&format!(" /N{:0w$} {fi} 0 R", i, w = name_len - 1)); + } + let mut c = String::from("BT /F1 1 Tf 72 700 Td (a) Tj "); + c.push_str(&"1 -1 Td (a) Tj ".repeat(lines - 1)); + c.push_str("ET"); + d.page(PageSpec::new(c.as_bytes(), &format!("/Font << {fonts} >>"))); + d.build() +} + +#[test] +fn runs_share_one_surface_per_font_and_long_names_join_no_group() { + // Review r3 HIGH-1: each run copied its 256 sibling names — 2,019 MiB for 2,000 lines with + // 4 KiB names, ~20 GiB at 19,999 lines, 798 MiB at the 127-byte name limit. + for (names, len, lines, held_max) in [ + (256, 4_096, 2_000, 32 * MIB), + (256, 4_096, 19_999, 64 * MIB), + (256, 127, 19_999, 64 * MIB), + ] { + let tag = format!("P1 {names} names of {len} B, {lines} lines"); + let p = probe(&tag, surface_page(names, len, lines)); + let m = &p.model; + assert_eq!(m.page_reason, None, "{tag}: {:?}", m.page_detail); + assert_eq!(m.runs.len(), lines, "{tag}"); + assert!( + m.walk.model_bytes < held_max, + "{tag}: {} B", + m.walk.model_bytes + ); + assert!(p.cpu < Duration::from_secs(4), "{tag}: {:?}", p.cpu); + assert_eq!(p.classify.occurrence_reason, None, "{tag}"); + assert_eq!(p.classify.occurrences.len(), lines, "{tag}"); + let first = &m.runs[0].surface; + assert!( + m.runs.iter().all(|r| Arc::ptr_eq(&r.surface, first)), + "{tag}: one shared list" + ); + // Names past 127 bytes stay out of sibling groups: the run types with `/F1` alone. + let want = if len > 127 { 1 } else { names }; + assert_eq!(first.len(), want, "{tag}"); + assert_eq!(first[0], b"F1".to_vec(), "{tag}"); + assert_eq!(m.surface(&m.runs[0]).fonts.len(), want, "{tag}"); + } +} + +/// One ExtGState with `keys` unmodelled keys `key_len` bytes long (a shared prefix, then three +/// digits), applied `ops` times before "Hello". +fn long_key_gs(keys: usize, key_len: usize, ops: usize) -> Vec { + let prefix = "K".repeat(key_len - 3); + let entries: Vec = (0..keys).map(|k| format!("/{prefix}{k:03} {k}")).collect(); + let gs = format!("/ExtGState << /G << {} >> >>", entries.join(" ")); + let c = format!( + "BT /F1 12 Tf 72 700 Td {}(Hello) Tj ET", + "/G gs ".repeat(ops) + ); + helvetica_doc(c.as_bytes(), &gs, "") +} + +#[test] +fn extgstate_keys_are_capped_and_a_repeated_gs_costs_constant_work() { + // Review r3 MEDIUM-2: 63 keys of 256 KiB re-applied took 2.5 ms per `gs` (≈ 10 min a walk). + let p = probe("P2 63 keys of 256 KiB", long_key_gs(63, 262_144, 20_000)); + assert_refused("P2", &p, "ExtGState keys"); + assert!(p.cpu < Duration::from_secs(2), "{:?}", p.cpu); + // At the 127-byte limit (124 bytes shared), 20,000 `gs` stay cheap: the dictionary that put + // the keys in force is skipped in O(1). + let p = probe("P2 63 keys of 127 B", long_key_gs(63, 127, 20_000)); + assert_eq!(p.model.page_reason, None, "{:?}", p.model.page_detail); + assert_eq!(reason_of(&p.model, "Hello"), None); + assert!(p.cpu < Duration::from_secs(1), "{:?}", p.cpu); + assert_eq!(p.model.walk.records[0].before.gs.other.len(), 63); +} + +/// Review r3 P4, P7 and P8: pages at the op cap whose every show op carries a state of its own. +fn state_per_op_pages() -> Vec<(&'static str, Vec)> { + let p4 = helvetica_page( + format!("BT /F1 1 Tf 72 700 Td {}ET", "0 0 (a) \" ".repeat(249_990)).as_bytes(), + ); + let keys: Vec = (0..63).map(|k| format!("/Key{k:02} {k}")).collect(); + let mut gs = format!("/ExtGState << /Base << {} >>", keys.join(" ")); + let mut c = String::from("/Base gs BT /F1 1 Tf 72 700 Td "); + for i in 0..124_990 { + gs.push_str(&format!(" /G{i} << /Odd {i} >>")); + c.push_str(&format!("/G{i} gs (a)Tj ")); + } + gs.push_str(" >>"); + c.push_str("ET"); + let p7 = helvetica_doc(c.as_bytes(), &gs, ""); + let arr = format!("[(a){}] TJ\n", " -1".repeat(65_535)); + let mut c = format!("BT /F1 1 Tf 1 TL 72 700 Td {}", arr.repeat(6)); + c.push_str(&"0 0 (a) \" ".repeat(249_990 - 20)); + c.push_str("ET"); + vec![("P4", p4), ("P7", p7), ("P8", helvetica_page(c.as_bytes()))] +} + +#[test] +fn pages_with_a_state_per_op_are_refused_by_the_page_budget() { + // Review r3 MEDIUM-1: they held 401–432 MiB (and the inspect path ~850 MiB with its Classify + // walk), while the documented worst case was 266 MiB. + for (tag, pdf) in state_per_op_pages() { + let p = probe(tag, pdf); + assert_refused(tag, &p, "page model size"); + // The Classify walk would charge the same and more: it is not walked again. + assert!(p.classify_peak < MIB, "{tag}: {} B", p.classify_peak); + } +} + +/// A page whose one Tf names a font resource `name_len` bytes long, then shows `glyphs` glyphs. +fn long_font_name_page(name_len: usize, glyphs: usize) -> Vec { + let name = "N".repeat(name_len); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let c = format!("BT /{name} 1 Tf 72 700 Td ({}) Tj ET", "a".repeat(glyphs)); + d.page(PageSpec::new( + c.as_bytes(), + &format!("/Font << /{name} {f} 0 R >>"), + )); + d.build() +} + +/// `n` distinct ExtGStates, each adding one key to 63 in force, each applied once, and no text. +fn interned_lists_page(n: usize) -> Vec { + let keys: Vec = (0..63).map(|k| format!("/Key{k:02} {k}")).collect(); + let mut gs = format!("/ExtGState << /Base << {} >>", keys.join(" ")); + let mut c = String::from("/Base gs "); + for i in 0..n { + gs.push_str(&format!(" /G{i} << /Odd {i} >>")); + c.push_str(&format!("/G{i} gs ")); + } + gs.push_str(" >>"); + c.push_str("BT /F1 12 Tf 72 700 Td (Hello) Tj ET"); + helvetica_doc(c.as_bytes(), &gs, "") +} + +/// "Hello" on the page, and a Form XObject painted once that holds `n` `"` ops. +fn form_bomb_page(n: usize) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let form = d.b.add_flate( + &format!("/Type /XObject /Subtype /Form /BBox [0 0 612 792] /Resources << /Font << /F1 {f} 0 R >> >>"), + format!("BT /F1 1 Tf 72 600 Td {}ET", "0 0 (a) \" ".repeat(n)).as_bytes(), + ); + d.page(PageSpec::new( + b"/Fx Do BT /F1 12 Tf 72 700 Td (Hello) Tj ET", + &format!("/Font << /F1 {f} 0 R >> /XObject << /Fx {form} 0 R >>"), + )); + d.build() +} + +#[test] +fn amplifications_no_review_listed_are_refused_by_the_page_budget() { + // Every glyph unit copies its font's resource name: one 64 KiB name and 4,000 glyphs would + // hold 250 MiB of copies (and 1 MiB names, the lexer's cap, 4 GiB). Refused before the units + // are made; at the 127-byte name limit the page is modelled. + let p = probe( + "long font name × glyphs", + long_font_name_page(65_536, 4_000), + ); + assert_refused("long font name", &p, "page model size"); + assert!(p.model_peak < 32 * MIB, "{} B", p.model_peak); + let p = probe( + "127-byte font name × glyphs", + long_font_name_page(127, 4_000), + ); + assert_eq!(p.model.page_reason, None, "{:?}", p.model.page_detail); + assert_eq!(p.model.runs.len(), 1); + // A walk interns every merged ExtGState key list it puts in force, text or not: 124,990 + // distinct `gs` over 63 keys would hold ~200 MiB of lists and draw nothing. + let p = probe("interned key lists", interned_lists_page(124_990)); + assert_refused("interned key lists", &p, "page model size"); + // A Form only the Classify walk descends: the model is small, the Classify pass is refused on + // its own budget and lists no occurrence (never a partial list). + let p = probe("Form bomb", form_bomb_page(249_000)); + assert_eq!(p.model.page_reason, None, "{:?}", p.model.page_detail); + assert_eq!(reason_of(&p.model, "Hello"), None); + assert!(p.model.walk.model_bytes < 4 * MIB); + assert_eq!(p.classify.occurrence_reason, Some(R::PageTooComplex)); + assert!(p.classify.occurrences.is_empty()); + assert!(p.classify_peak < REFUSED_PEAK_MAX, "{} B", p.classify_peak); +} + +#[test] +fn the_structure_parent_array_is_borrowed_per_build() { + // Review r3 LOW-2: a parent array of 2,000,000 nulls was cloned per build (259 MiB). + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let parents = d.add(format!("[{}]", "null ".repeat(2_000_000))); + let tree = d.add(format!("<< /Nums [0 {parents} 0 R] >>")); + let root = d.add(format!( + "<< /Type /StructTreeRoot /ParentTree {tree} 0 R >>" + )); + d.page( + PageSpec::new( + b"BT /F1 12 Tf 72 700 Td /P <> BDC (Hello) Tj EMC ET", + &format!("/Font << /F1 {f} 0 R >>"), + ) + .with("/StructParents 0"), + ); + d.catalog_extra = format!("/StructTreeRoot {root} 0 R"); + // The snapshot's reference counts (built once, on first use) are made before measuring. + let c = ctx(d.build()); + assert_eq!(reason_of(&model(&c, 0), "Hello"), None); + let (m, peak) = thread_peak(|| model(&c, 0)); + let (classify, classify_peak) = thread_peak(|| classify_source_page(&c, &m, None)); + println!( + "P10 parent tree: model peak {} KiB, classify peak {} KiB", + peak >> 10, + classify_peak >> 10 + ); + assert_eq!(reason_of(&m, "Hello"), None); + assert_eq!(classify.occurrences.len(), 1); + assert!(peak < MIB, "{peak} B"); + assert!(classify_peak < MIB, "{classify_peak} B"); +} + +#[test] +fn the_budget_refuses_a_page_the_moment_it_would_hold_more() { + // 10,000 lines (each 30 pt to the right of the last, so no two are compared as duplicates) + // with a colour change each: modelled within the default budget, refused under + // one byte less than it holds, and modelled again under what it holds plus room for scratch. + let mut c = String::from("BT /F1 12 Tf 72 700 Td "); + for i in 0..10_000 { + c.push_str(&format!("{} g 30 -14 Td (Hi) Tj ", i % 2)); + } + c.push_str("ET"); + let cx = ctx(helvetica_page(c.as_bytes())); + let m = model(&cx, 0); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + let held = m.walk.model_bytes; + assert!(m.approx_bytes() <= held && held < 64 * MIB); + set_model_bytes_override(Some(held - 1)); + let refused = model(&cx, 0); + let classify = classify_source_page(&cx, &refused, None); + set_model_bytes_override(Some(held + 32 * MIB)); + let again = model(&cx, 0); + set_model_bytes_override(None); + assert_eq!(refused.page_reason, Some(R::PageTooComplex)); + assert_eq!(refused.page_detail.as_deref(), Some("page model size")); + assert!(refused.runs.is_empty()); + // A model refused for size (here at its run stage) is not walked again by its Classify pass, + // which would charge the same and more (fix pass 2026-10-03, review T3-budget MEDIUM-1; it + // used to list every occurrence with the page's refusal). + assert_eq!(classify.occurrence_reason, Some(R::PageTooComplex)); + assert!(classify.occurrences.is_empty()); + assert_eq!(again.page_reason, None, "{:?}", again.page_detail); + assert_eq!(again.walk.model_bytes, held, "charges are deterministic"); + assert_eq!(again.runs.len(), m.runs.len()); +} + +#[test] +fn normal_pages_hold_a_few_mib_and_never_more_than_they_charged() { + let pages: Vec<(&str, Vec)> = vec![ + ("FX-WORD", fx::word()), + ("FX-WORD-TR", fx::word_tr()), + ("FX-LIBRE", fx::libre()), + ("FX-SKIA", fx::skia()), + ("FX-QUARTZ", fx::quartz()), + ("FX-XETEX", fx::xetex()), + ("FX-INDD", fx::indd()), + ("FX-STD14", fx::std14()), + ("FX-PERGLYPH", fx::per_glyph(None)), + ("two columns, tagged", fx::two_column(true)), + ("nested form", fx::nested_form()), + ]; + for (tag, pdf) in pages { + let p = probe(tag, pdf); + assert_eq!( + p.model.page_reason, None, + "{tag}: {:?}", + p.model.page_detail + ); + assert!( + p.model.walk.model_bytes < 4 * MIB, + "{tag}: {} B", + p.model.walk.model_bytes + ); + assert_eq!(p.classify.occurrence_reason, None, "{tag}"); + assert!( + p.model_peak < 16 * MIB && p.classify_peak < 16 * MIB, + "{tag}" + ); + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/fixes.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/fixes.rs new file mode 100644 index 0000000..40550d5 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/fixes.rs @@ -0,0 +1,468 @@ +//! Engine fix pass (2026-10-03) regressions for review T3-budget: font models are charged to the +//! page model (HIGH-1); a page refused after its walk keeps no walk and is not walked again +//! (MEDIUM-1); the Classify pass runs on the budget its model left and reuses a walk with no Form +//! (MEDIUM-2); a model keeps no lexed ops, so dense pages with a matrix per glyph model and +//! classify (MEDIUM-3); a growing vector's old buffer counts (LOW-1). +//! +//! "Held" is what the calling thread still has allocated when a build returns (`thread_peak_held`); +//! an inspect costs the model it holds plus its Classify pass's peak. + +use super::{ctx, model, walk}; +use crate::pdf_engine::source_content::classify_source_page; +use crate::pdf_engine::text_edit::fonts::FontModel; +use crate::pdf_engine::text_edit::fonts::TypingSurface; +use crate::pdf_engine::text_edit::limits::{set_model_bytes_override, PAGE_MODEL_BYTES_MAX}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::runs::{surface_of, PageModel}; +use crate::pdf_engine::text_edit::testkit::fonts::{add_type0, cmap, Program, Type0Font}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, DocBuilder, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::{thread_peak, thread_peak_held}; +use crate::pdf_engine::text_edit::walker::budget::{ModelBudget, MODEL_SIZE}; +use crate::pdf_engine::text_edit::walker::{PaintKind, WalkMode}; +use std::collections::HashSet; +use std::sync::Arc; + +const MIB: usize = 1 << 20; +/// §H R20's per-file budget. +const PEAK_MAX: usize = 256 * MIB; +/// `fonts::FONT_CACHE_BYTES_MAX`: the snapshot's own font cache, bounded apart from the models. +const FONT_CACHE_MAX: usize = 128 * MIB; + +/// Distinct font models `m` keeps alive (page fonts and records), once each by allocation. +fn kept_fonts(m: &PageModel) -> Vec> { + let mut seen = HashSet::new(); + let page = m.walk.page_fonts.iter().map(|(_, f)| f); + page.chain(m.walk.records.iter().filter_map(|r| r.font.as_ref())) + .filter(|f| seen.insert(Arc::as_ptr(f) as usize)) + .cloned() + .collect() +} + +/// `pages` pages, each drawing `n` Type0 fonts once and listing `n` more it does not use; every +/// font has its own one-line ToUnicode mapping 65,536 codes (a few bytes deflated, MiBs of model). +fn font_pages(n: usize, pages: usize) -> Vec { + let tu = cmap( + "1 begincodespacerange\n<0000> \nendcodespacerange\n\ + 1 beginbfrange\n<0000> <0000>\nendbfrange", + ); + let mut d = DocBuilder::new(); + for _ in 0..pages { + let mut fonts = String::new(); + let mut c = String::from("BT "); + for k in 0..2 * n { + let mut f = Type0Font::new("CIDFontType2", "ABCDEF+Big"); + f.program = Program::None; + f.w = Some("[1 [500]]".to_string()); + f.tounicode = Some(tu.clone()); + let id = add_type0(&mut d.b, &f); + fonts.push_str(&format!("/T{k} {id} 0 R ")); + if k < n { + c.push_str(&format!("/T{k} 10 Tf 72 {} Td <0001> Tj ", 700 - k)); + } + } + c.push_str("ET"); + d.page(PageSpec::new(c.as_bytes(), &format!("/Font << {fonts} >>"))); + } + d.build() +} + +#[test] +fn font_models_count_in_the_page_model_that_keeps_them() { + // Review HIGH-1: 6 such pages held 530 MiB of font models, outside every budget, with + // `model_bytes` and `approx_bytes` at 0 (T5's cache counted them as nothing). + let pages = 3; + let c = ctx(font_pages(8, pages)); + let (models, peak, held) = + thread_peak_held(|| (0..pages as u32).map(|p| model(&c, p)).collect::>()); + let mut charged = 0usize; + for (p, m) in models.iter().enumerate() { + assert_eq!(m.page_reason, None, "page {p}: {:?}", m.page_detail); + let fonts: usize = kept_fonts(m).iter().map(|f| f.approx_bytes()).sum(); + println!( + "fonts 8+8 page {p}: charged {} MiB, approx {} MiB, fonts kept {} ({} MiB), \ + page fonts {}", + m.walk.model_bytes / MIB, + m.approx_bytes() / MIB, + kept_fonts(m).len(), + fonts / MIB, + m.walk.page_fonts.len() + ); + assert!(fonts > 8 * 4 * MIB, "page {p}: the drawn fonts are kept"); + assert!( + m.approx_bytes() >= fonts, + "page {p}: approx counts its fonts" + ); + assert!(m.approx_bytes() <= m.walk.model_bytes, "page {p}"); + assert!(m.walk.model_bytes <= PAGE_MODEL_BYTES_MAX, "page {p}"); + assert!( + m.walk.page_fonts.len() < 16, + "page {p}: unused fonts left out" + ); + charged += m.walk.model_bytes; + } + println!( + "fonts 8+8 x {pages}: peak {} MiB, held {} MiB, charged {} MiB", + peak / MIB, + held / MIB, + charged / MIB + ); + assert!( + held <= charged + FONT_CACHE_MAX + 16 * MIB, + "everything the models keep is in their model_bytes (held {held}, charged {charged})" + ); +} + +#[test] +fn drawn_fonts_past_the_budget_refuse_the_page_and_unused_ones_are_left_out() { + let c = ctx(font_pages(8, 1)); + set_model_bytes_override(Some(40 * MIB)); + let refused = model(&c, 0); + set_model_bytes_override(None); + assert_eq!(refused.page_reason, Some(R::PageTooComplex)); + assert_eq!(refused.page_detail.as_deref(), Some(MODEL_SIZE)); + assert!( + kept_fonts(&refused).is_empty(), + "a refused page keeps no font" + ); + // One drawn font and eight unused ones: the page models, with the siblings that fit. + let c = ctx(font_pages(1, 1)); + let full = model(&c, 0); + set_model_bytes_override(Some(24 * MIB)); + let small = model(&c, 0); + set_model_bytes_override(None); + assert_eq!(small.page_reason, None, "{:?}", small.page_detail); + assert_eq!( + full.walk.page_fonts.len(), + 2, + "1 drawn + the unused one that fits" + ); + assert_eq!( + small.walk.page_fonts.len(), + 1, + "no room for an unused sibling" + ); + assert!(small.walk.model_bytes <= 24 * MIB); +} + +/// "Hello" plus `n` `(a)Tj` ops under one state on a line far below (review probe shape). +fn near_budget(n: usize) -> Vec { + let mut c = String::from("BT /F1 12 Tf 72 700 Td (Hello) Tj ET BT /F1 1 Tf 0 20 Td "); + c.push_str(&"(a)Tj ".repeat(n)); + c.push_str("ET"); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page( + PageSpec::new(c.as_bytes(), &format!("/Font << /F1 {f} 0 R >>")) + .media(Some("[0 0 2400 3400]")), + ); + d.build() +} + +#[test] +fn a_page_refused_after_its_walk_keeps_no_walk_and_is_not_walked_again() { + // Long strings: the run stage costs about what the walk does, so a budget between the two + // refuses the page at the run stage. + let mut content = String::from("BT /F1 1 Tf 0 20 Td "); + for _ in 0..2_000 { + content.push_str(&format!("({}) Tj 0 -1.2 Td ", "abcdefgh ".repeat(12))); + } + content.push_str("ET"); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page( + PageSpec::new(content.as_bytes(), &format!("/Font << /F1 {f} 0 R >>")) + .media(Some("[0 0 2400 3400]")), + ); + let c = ctx(d.build()); + let full = model(&c, 0); + let w = walk(&c, 0, WalkMode::Edit); + assert_eq!(full.page_reason, None); + let (walked, total) = (w.model_bytes, full.walk.model_bytes); + assert!(walked < total, "walk {walked} < model {total}"); + set_model_bytes_override(Some((walked + total) / 2 + MIB)); + let (m, _, held) = thread_peak_held(|| model(&c, 0)); + let (cl, classify_peak) = thread_peak(|| classify_source_page(&c, &m, None)); + set_model_bytes_override(None); + assert_eq!(m.page_reason, Some(R::PageTooComplex)); + assert_eq!( + m.page_detail.as_deref(), + Some(MODEL_SIZE), + "at the run stage" + ); + assert!( + m.walk.records.is_empty() && m.walk.paints.is_empty() && m.walk.page_fonts.is_empty(), + "the refused model released its walk ({} records)", + m.walk.records.len() + ); + assert!(m.record_reason.is_empty()); + assert!(m.approx_bytes() <= m.walk.model_bytes); + assert!(held < 2 * content.len() + MIB, "held {held} B"); + assert_eq!(cl.occurrence_reason, Some(R::PageTooComplex)); + assert!(classify_peak < MIB, "not walked again: {classify_peak} B"); +} + +/// The model of page 0 and its Classify pass: the bytes held after the build plus the pass's +/// peak (what an inspect costs at once). +fn inspect(tag: &str, pdf: Vec) -> (PageModel, usize, usize) { + let c = ctx(pdf); + let (m, model_peak, held) = thread_peak_held(|| model(&c, 0)); + let (cl, classify_peak) = thread_peak(|| classify_source_page(&c, &m, None)); + let records = m.walk.records.len(); + println!( + "{tag}: model peak {} MiB, held {} MiB, charged {} MiB, {:?} {:?}, {records} records, \ + {} runs; classify peak {} MiB, {:?}, {} occurrences; inspect {} MiB", + model_peak / MIB, + held / MIB, + m.walk.model_bytes / MIB, + m.page_reason, + m.page_detail, + m.runs.len(), + classify_peak / MIB, + cl.occurrence_reason, + cl.occurrences.len(), + (held + classify_peak) / MIB + ); + assert!(model_peak < PEAK_MAX, "{tag}: model peak {model_peak}"); + assert!( + m.walk.ops.is_empty() && m.walk.ops_bytes == 0, + "{tag}: a model keeps no ops" + ); + assert!(m.approx_bytes() <= m.walk.model_bytes, "{tag}"); + assert!(m.walk.model_bytes <= PAGE_MODEL_BYTES_MAX, "{tag}"); + let occurrences = match cl.occurrence_reason { + None => cl.occurrences.len(), + Some(_) => 0, + }; + (m, held + classify_peak, occurrences) +} + +#[test] +fn the_classify_pass_shares_its_model_budget() { + // Review MEDIUM-2: 100,000 shows under one state peaked at 210 MiB (model 106 + Classify + // walk 104 on a budget of its own); 135,000 at 303 MiB. + for n in [100_000usize, 135_000] { + let (m, at_once, occurrences) = inspect(&format!("near budget {n}"), near_budget(n)); + // One page budget for both (each on its own budget peaked at 186 MiB here). + assert!(at_once < PAGE_MODEL_BYTES_MAX, "{n}: inspect {at_once} B"); + if m.page_reason.is_none() { + assert_eq!(occurrences, n + 1, "{n}: Classify complete"); + } + } + // A page that paints a Form is walked again in Classify mode, on what its model left. + let (m, at_once, occurrences) = inspect("nested form", fx::nested_form()); + assert_eq!(m.page_reason, None); + assert!( + occurrences > m.walk.records.len(), + "the Form's text is listed too" + ); + assert!(at_once < 16 * MIB); +} + +#[test] +fn the_edit_walk_of_a_page_without_forms_is_its_classify_walk() { + // `classify_source_page` lists such a page's occurrences from its model's own walk. + let pages: Vec<(&str, Vec)> = vec![ + ("FX-WORD", fx::word()), + ("FX-LIBRE", fx::libre()), + ("FX-SKIA", fx::skia()), + ("FX-QUARTZ", fx::quartz()), + ("FX-XETEX", fx::xetex()), + ("FX-INDD", fx::indd()), + ("FX-PERGLYPH", fx::per_glyph(None)), + ("two columns, tagged", fx::two_column(true)), + ("images", fx::swapped_image(false)), + ]; + let mut compared = 0; + for (tag, pdf) in pages { + let c = ctx(pdf); + let edit = walk(&c, 0, WalkMode::Edit); + assert_eq!(edit.page_reason, None, "{tag}"); + if edit + .paints + .iter() + .any(|p| matches!(p.kind, PaintKind::FormXObject { .. })) + { + continue; // walked again in Classify mode (FX-INDD paints a Form) + } + compared += 1; + let classify = walk(&c, 0, WalkMode::Classify); + let recs = |w: &crate::pdf_engine::text_edit::walker::PageWalk| -> Vec { + let r = w.records.iter().map(|r| { + let text: Vec<_> = r.glyphs.iter().map(|g| g.text.clone()).collect(); + format!( + "{} {} {:?} {:?} {text:?} {:?}", + r.seq, r.depth, r.span, r.local_span, r.pen_before + ) + }); + let p = w.paints.iter().map(|p| { + format!( + "{} {} {:?} {:?} {:?} {:?}", + p.seq, p.depth, p.kind, p.span, p.bbox, p.xobject + ) + }); + r.chain(p).collect() + }; + assert_eq!(recs(&edit), recs(&classify), "{tag}"); + assert_eq!( + edit.model_bytes, classify.model_bytes, + "{tag}: the same charge" + ); + } + assert!(compared >= 7, "{compared} pages compared"); +} + +#[test] +fn a_runs_surface_names_its_typing_surface() { + // `PageModel::surface` now reads `run.surface` (it recomputed the sibling group per call, and + // the field was read by nothing in production): the two are the same fonts in the same order. + let pages: Vec<(&str, Vec)> = vec![ + ("FX-WORD", fx::word()), + ("FX-WORD-TR", fx::word_tr()), + ("FX-LIBRE", fx::libre()), + ("FX-XETEX", fx::xetex()), + ("FX-INDD", fx::indd()), + ("FX-STD14", fx::std14()), + ]; + let mut checked = 0; + for (tag, pdf) in pages { + let m = model(&ctx(pdf), 0); + for run in &m.runs { + let primary = &m.walk.records[run.members[0]]; + let names = |s: TypingSurface| -> Vec> { + s.fonts.into_iter().map(|(n, _)| n).collect() + }; + let recomputed = names(surface_of(&m.walk, primary)); + assert_eq!(names(m.surface(run)), recomputed, "{tag}: {:?}", run.text); + if !run.surface.is_empty() { + assert_eq!(run.surface.to_vec(), recomputed, "{tag}: {:?}", run.text); + checked += 1; + } + } + } + assert!(checked > 10, "{checked} runs"); +} + +/// 3,000 lines of 32 Courier glyphs (96,000), one `Tm` per glyph (`e2`), and a gray change per +/// word (`g2`, a syntax-highlighted listing): review MEDIUM-3's normal dense pages. +fn dense(shape: &str) -> Vec { + let mut c = String::from("BT /F1 1 Tf 1.1 TL\n"); + for i in 0..3_000 { + let t = format!("Line {i:05} lorem ipsum dolor sit"); + let y = 3300.0 - 1.1 * i as f64; + let mut k = 0usize; + for (wi, w) in t.split(' ').enumerate() { + if shape == "g2" { + c.push_str(&format!("{} g ", (i + wi) % 2)); + } + for ch in format!("{w} ").chars() { + let x = 20.0 + 0.6 * k as f64; + c.push_str(&format!("1 0 0 1 {x:.1} {y:.2} Tm ({ch}) Tj ")); + k += 1; + } + } + c.push('\n'); + } + c.push_str("ET"); + let mut d = DocBuilder::new(); + let courier = "<< /Type /Font /Subtype /Type1 /BaseFont /Courier /Encoding /WinAnsiEncoding >>"; + let f = d.add(courier); + d.page( + PageSpec::new(c.as_bytes(), &format!("/Font << /F1 {f} 0 R >>")) + .media(Some("[0 0 2400 3400]")), + ); + d.build() +} + +#[test] +fn dense_pages_with_a_matrix_per_glyph_model_and_classify() { + // Before: e2 modelled at 148 MiB with its Classify pass refused (0 occurrences); g2 refused + // "page model size" — both mostly for the six lexed `Tm` operands a model kept per glyph. + for shape in ["e2", "g2"] { + let (m, at_once, occurrences) = inspect(shape, dense(shape)); + assert_eq!(m.page_reason, None, "{shape}: {:?}", m.page_detail); + let lines = if shape == "e2" { 3_000 } else { 18_000 }; + assert_eq!( + m.runs.len(), + lines, + "{shape}: a run per line (per word in g2's colours)" + ); + assert_eq!( + occurrences, + m.walk.records.len(), + "{shape}: Classify complete" + ); + assert!(at_once < PEAK_MAX, "{shape}: inspect {at_once} B"); + } +} + +#[test] +fn growing_vectors_count_their_old_buffer() { + // The test allocator counts a realloc as new then old (it used to count the difference only). + let ((), peak) = thread_peak(|| { + let mut v: Vec = vec![0; MIB]; + v.reserve_exact(2 * MIB); + assert!(v.capacity() >= 3 * MIB); + }); + assert!(peak >= 4 * MIB, "old and new buffers at once: {peak} B"); + // The budget lets a vector grow only when its old buffer fits beside the new room. + set_model_bytes_override(Some(1_200)); + let mut mem = ModelBudget::new(0); + let mut v: Vec = Vec::new(); + let grown = (0..64).try_for_each(|i| mem.push(&mut v, i)); + let refused = mem.push(&mut v, 64); + set_model_bytes_override(None); + assert!(grown.is_ok() && v.capacity() == 64, "512 B held"); + assert!( + refused.is_err(), + "512 B more fit the 688 B left, but not with the old 512 B buffer" + ); +} + +#[test] +fn text_drawn_through_a_form_is_listed_as_refused_lines() { + // Live check B1: the repo's text-nested-form.pdf showed "Hi" but listed no run at all, so the + // editor said the page had no text (§A.6: Form text is "refused, still shown"). + let path = std::path::Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../fixtures/source-edit/text-nested-form.pdf"); + let c = ctx(std::fs::read(&path).expect("fixture")); + let m = model(&c, 0); + let cl = classify_source_page(&c, &m, None); + let lines: Vec<(&str, R)> = cl + .form_lines + .iter() + .map(|l| (l.text.as_str(), l.reason)) + .collect(); + assert_eq!(lines, [("Hi", R::NestedForm)], "{:?}", m.runs.len()); + // A Form line of two shows on one baseline, after the page's own run, with its own id. + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let form = d.b.add_stream( + &format!( + "/Type /XObject /Subtype /Form /BBox [0 0 612 792] \ + /Resources << /Font << /F1 {f} 0 R >> >>" + ), + b"BT /F1 12 Tf 72 600 Td (Hello) Tj ( World) Tj ET BT /F1 12 Tf 72 560 Td (Next) Tj ET", + ); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Page text) Tj ET /Fm0 Do", + &format!("/XObject << /Fm0 {form} 0 R >> /Font << /F1 {f} 0 R >>"), + )); + let c = ctx(d.build()); + let m = model(&c, 0); + let cl = classify_source_page(&c, &m, None); + let texts: Vec<&str> = cl.form_lines.iter().map(|l| l.text.as_str()).collect(); + assert_eq!(texts, ["Hello World", "Next"]); + assert_eq!(m.runs.len(), 1, "the page's own run"); + let first = &cl.form_lines[0]; + assert!(first.order > m.runs[0].order && first.rect[2] > 0.0 && first.rect[3] > 0.0); + assert!(first.caret_offsets.len() == "Hello World".chars().count() + 1); + assert_ne!(first.id, cl.form_lines[1].id); + assert!(cl.form_lines.iter().all(|l| l.reason == R::NestedForm)); + // Pages that paint no Form list none. + let c = ctx(fx::word()); + assert!(classify_source_page(&c, &model(&c, 0), None) + .form_lines + .is_empty()); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/fuzz.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/fuzz.rs new file mode 100644 index 0000000..6329cba --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/fuzz.rs @@ -0,0 +1,175 @@ +//! Walker fuzz (§B.21, T3): 10,000 deterministic xorshift mutations — byte flips, truncations, +//! inserted `(`, `<`, `[`, `BI`, `q`, huge numbers — of fixture content streams through +//! lex → walk (Edit and Classify) → runs. No panic; < 2 s of CPU per 1,000 cases. + +use super::{content, ctx}; +use crate::pdf_engine::text_edit::runs::model_of; +use crate::pdf_engine::text_edit::testkit::producers::{cid_font, cid_hex}; +use crate::pdf_engine::text_edit::testkit::producers::{DocBuilder, PageSpec, HELVETICA}; +use crate::pdf_engine::text_edit::walker::{walk_page, WalkMode}; +use std::time::Duration; + +/// CPU time of the calling thread (wall time where the platform has no thread clock). +pub(super) fn thread_cpu() -> Duration { + #[cfg(all( + any(target_os = "macos", target_os = "linux"), + target_pointer_width = "64" + ))] + { + #[repr(C)] + struct Timespec { + tv_sec: i64, + tv_nsec: i64, + } + extern "C" { + fn clock_gettime(clock: i32, tp: *mut Timespec) -> i32; + } + #[cfg(target_os = "macos")] + const CLOCK_THREAD_CPUTIME_ID: i32 = 16; + #[cfg(target_os = "linux")] + const CLOCK_THREAD_CPUTIME_ID: i32 = 3; + let mut ts = Timespec { + tv_sec: 0, + tv_nsec: 0, + }; + // SAFETY: clock_gettime only writes the timespec it is handed, which outlives the call. + if unsafe { clock_gettime(CLOCK_THREAD_CPUTIME_ID, &mut ts) } == 0 { + return Duration::new( + ts.tv_sec.max(0) as u64, + ts.tv_nsec.clamp(0, 999_999_999) as u32, + ); + } + } + static START: std::sync::OnceLock = std::sync::OnceLock::new(); + START.get_or_init(std::time::Instant::now).elapsed() +} + +struct XorShift(u64); + +impl XorShift { + fn next(&mut self) -> u64 { + let mut x = self.0; + x ^= x << 13; + x ^= x >> 7; + x ^= x << 17; + self.0 = x; + x + } + + fn below(&mut self, n: usize) -> usize { + (self.next() % n.max(1) as u64) as usize + } +} + +/// A page whose resources cover every operator family the seeds use. +fn fixture() -> Vec { + let mut d = DocBuilder::new(); + let f1 = d.add(HELVETICA); + let f2 = cid_font(&mut d.b, "ABCDEF+Arimo", "Fuzzy "); + let img = d.b.add_stream( + "/Type /XObject /Subtype /Image /Width 1 /Height 1 /ColorSpace /DeviceGray /BitsPerComponent 8", + &[128], + ); + let form = d.b.add_stream( + &format!("/Type /XObject /Subtype /Form /BBox [0 0 100 100] /Resources << /Font << /F1 {f1} 0 R >> >>"), + b"BT /F1 9 Tf 5 5 Td (form) Tj ET", + ); + let ocg = d.add("<< /Type /OCG /Name (L) >>"); + d.catalog_extra = format!("/OCProperties << /OCGs [{ocg} 0 R] /D << >> >>"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (seed) Tj ET", + &format!( + "/Font << /F1 {f1} 0 R /F2 {f2} 0 R >> /XObject << /Im0 {img} 0 R /Fm0 {form} 0 R >> \ + /ExtGState << /GS1 << /ca 0.5 /LW 2 >> >> /Properties << /L1 {ocg} 0 R >> \ + /ColorSpace << /Cs1 /DeviceRGB >>" + ), + )); + d.build() +} + +fn seeds() -> Vec> { + vec![ + b"BT /F1 12 Tf 72 700 Td [(Hel) -20 (lo) -333 (World)] TJ (tail) Tj ET".to_vec(), + format!( + "q 1 0 0 1 10 10 cm /GS1 gs 0.5 g BT /F2 10 Tf 2 Tc 3 Tw 90 Tz 1 Ts 72 650 Td <{}> Tj ET Q", + cid_hex("Fuzzy ", "Fuzzy Fuzz") + ) + .into_bytes(), + b"q 0 0 612 792 re W n /Cs1 cs 1 0 0 sc 10 10 m 20 20 l 30 10 25 5 15 15 c h f Q \ + q 40 0 0 40 72 400 cm /Im0 Do Q q 1 0 0 1 300 300 cm /Fm0 Do Q" + .to_vec(), + b"/P <> BDC BT /F1 12 Tf 14 TL 72 720 Td (One) ' 1 0.5 (Two) \" ET EMC \ + /OC /L1 BDC BT 3 Tr /F1 8 Tf 1 0 0.2 1 72 500 Tm (Layer) Tj ET EMC" + .to_vec(), + b"q 24 0 0 12 72 400 cm BI /W 2 /H 1 /CS /RGB /BPC 8 ID abcdef EI Q \ + BT /F1 12 Tf 0 1 -1 0 300 300 Tm (After) Tj T* (Next) Tj ET [3 2] 0 d 2 J" + .to_vec(), + ] +} + +fn mutate(base: &[u8], rng: &mut XorShift) -> Vec { + let mut v = base.to_vec(); + for _ in 0..1 + rng.below(4) { + let at = rng.below(v.len() + 1); + match rng.below(9) { + 0 => { + if let Some(b) = v.get_mut(at) { + *b ^= 1 << rng.below(8); + } + } + 1 => v.truncate(at), + 2 => v.insert(at, b'('), + 3 => v.insert(at, b'<'), + 4 => v.insert(at, b'['), + 5 => v.splice(at..at, b" BI ".iter().copied()).for_each(drop), + 6 => v.splice(at..at, b" q ".iter().copied()).for_each(drop), + 7 => v + .splice(at..at, b" 999999999 -1e9 999999999.5 ".iter().copied()) + .for_each(drop), + _ => v + .splice(at..at, b" Q ET BT ".iter().copied()) + .for_each(drop), + } + } + v +} + +#[test] +fn walker_fuzz() { + let c = ctx(fixture()); + let base = content(&c, 0); + let seeds = seeds(); + let mut rng = XorShift(0x9E37_79B9_7F4A_7C15); + let mut refused = 0usize; + for chunk in 0..10 { + let started = thread_cpu(); + for i in 0..1_000 { + let seed = &seeds[(chunk * 1_000 + i) % seeds.len()]; + let bytes = mutate(seed, &mut rng); + let page = base.with_replaced_parts(&[(0, bytes)]); + let classify = walk_page(&c, 0, &page, WalkMode::Classify, None); + let edit = walk_page(&c, 0, &page, WalkMode::Edit, None); + refused += usize::from(edit.page_reason.is_some()); + let m = model_of(&c, 0, page, edit); + assert!(m + .runs + .iter() + .all(|r| r.caret_offsets.len() == r.text.chars().count() + 1)); + assert_eq!( + classify.page_reason.is_some() && classify.records.is_empty() + || classify.page_reason.is_none(), + true, + "never a partial list" + ); + } + let spent = thread_cpu().saturating_sub(started); + assert!( + spent < Duration::from_secs(2), + "walker fuzz chunk {chunk}: {spent:?} for 1,000 cases" + ); + } + assert!( + refused > 0 && refused < 10_000, + "the fuzz reaches both outcomes: {refused}" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/geo.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/geo.rs new file mode 100644 index 0000000..1e78964 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/geo.rs @@ -0,0 +1,268 @@ +//! GEO-01…15: orientation as displayed (§A.5 table, composite × page `/Rotate`) and the strict +//! page geometry parser, with `crop.rs` only in the parity test GEO-11. + +use super::{close, ctx, model0, reason_of}; +use crate::pdf_engine::crop; +use crate::pdf_engine::text_edit::geometry::{ + classify_orientation, display_rotation, mul, page_geometry, ray_extent, Orientation, +}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::snapshot::read_snapshot; +use crate::pdf_engine::text_edit::testkit::producers::{self as fx, helvetica_page, BadBox}; +use std::path::Path; + +fn one(content: &str, text: &str) -> Option { + reason_of(&model0(helvetica_page(content.as_bytes())), text) +} + +#[test] +fn geo_01_to_09_orientation_cases() { + let cases: [(&str, &str, &str, Option); 9] = [ + ( + "GEO-01 upright", + "BT /F1 12 Tf 72 700 Td (Upright) Tj ET", + "Upright", + None, + ), + ( + "GEO-02 180°", + "BT /F1 12 Tf -1 0 0 -1 300 400 Tm (Turned) Tj ET", + "Turned", + Some(R::RotatedText), + ), + ( + "GEO-03 mirror", + "BT /F1 12 Tf -1 0 0 1 300 400 Tm (Mirror) Tj ET", + "Mirror", + Some(R::MirroredText), + ), + ( + "GEO-04 -100 Tz", + "BT /F1 12 Tf -100 Tz 300 400 Td (Mirror) Tj ET", + "Mirror", + Some(R::MirroredText), + ), + ( + "GEO-05 -12 Tf", + "BT /F1 -12 Tf 300 400 Td (Turned) Tj ET", + "Turned", + Some(R::RotatedText), + ), + ( + "GEO-06 45°", + "BT /F1 12 Tf 0.7071 0.7071 -0.7071 0.7071 300 400 Tm (Angle) Tj ET", + "Angle", + Some(R::RotatedText), + ), + ( + "GEO-07 skew", + "BT /F1 12 Tf 1 0.3 0 1 72 400 Tm (Skew) Tj ET", + "Skew", + Some(R::SkewedText), + ), + ( + "GEO-08 oblique c=0.2", + "BT /F1 12 Tf 1 0 0.2 1 72 400 Tm (Oblique) Tj ET", + "Oblique", + None, + ), + ( + "GEO-09 Skia flip", + "1 0 0 -1 0 792 cm BT /F1 12 Tf 1 0 0 -1 72 100 Tm (Skia) Tj ET", + "Skia", + None, + ), + ]; + for (id, content, text, want) in cases { + assert_eq!(one(content, text), want, "{id}"); + } + assert_eq!( + one("BT /F1 0 Tf 72 700 Td (Zero) Tj ET", "Zero"), + Some(R::ZeroSize), + "Tf 0" + ); + assert_eq!( + one("BT /F1 12 Tf 0 Tz 72 700 Td (Zero) Tj ET", "Zero"), + Some(R::ZeroSize), + "Tz 0" + ); + // The table itself, on bare matrices. + for (m, want) in [ + ([12.0, 0.0, 0.0, 12.0, 0.0, 0.0], Orientation::Upright), + ([12.0, 0.0, 6.0, 12.0, 0.0, 0.0], Orientation::Upright), + ([12.0, 0.0, 6.1, 12.0, 0.0, 0.0], Orientation::Skewed), + ([0.0, 12.0, -12.0, 0.0, 0.0, 0.0], Orientation::Rotated), + ([0.0, 12.0, 12.0, 0.0, 0.0, 0.0], Orientation::Skewed), + ([0.0, 0.0, 0.0, 12.0, 0.0, 0.0], Orientation::ZeroSize), + ([-12.0, 0.0, 0.0, -12.0, 0.0, 0.0], Orientation::Rotated), + ([12.0, 0.0, 0.0, -12.0, 0.0, 0.0], Orientation::Mirrored), + ] { + assert_eq!(classify_orientation(&m), want, "{m:?}"); + } + assert!(close( + ray_extent((72.0, 700.0), (1.0, 0.0), [0.0, 0.0, 612.0, 792.0]), + 540.0 + )); + assert_eq!( + ray_extent((700.0, 700.0), (1.0, 0.0), [0.0, 0.0, 612.0, 792.0]), + 0.0 + ); +} + +#[test] +fn geo_10_rotate_pairs_judged_as_displayed() { + for angle in [90, 180, 270] { + assert_eq!( + reason_of(&model0(fx::rotated(angle, true)), "Rotated page"), + None, + "GEO-10 /Rotate {angle} with counter-rotated text reads upright" + ); + assert_eq!( + reason_of(&model0(fx::rotated(angle, false)), "Rotated page"), + Some(R::RotatedText), + "GEO-10 /Rotate {angle} with user-upright text is turned on screen" + ); + } + let r = display_rotation(90); + assert_eq!( + mul(&[0.0, 1.0, -1.0, 0.0, 0.0, 0.0], &r), + [1.0, 0.0, 0.0, 1.0, 0.0, 0.0] + ); +} + +fn corpus(name: &str) -> std::path::PathBuf { + Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../fixtures/source-edit") + .join(name) +} + +#[test] +fn geo_11_parity_with_crop_rs_on_accepted_pages() { + let mut docs = vec![ + fx::word(), + fx::word_tr(), + fx::cropped_offset(), + fx::two_column(false), + fx::libre(), + fx::shared(), + ]; + for angle in [90, 180, 270] { + docs.push(fx::rotated(angle, true)); + } + let mut snaps: Vec<_> = docs.into_iter().map(|d| ctx(d).snap).collect(); + for name in [ + "geom-crop-offset.pdf", + "geom-rotate-90.pdf", + "geom-rotate-180.pdf", + "geom-rotate-270.pdf", + "text-tj.pdf", + "image-unique.pdf", + ] { + snaps.push(read_snapshot(&corpus(name)).expect("corpus fixture")); + } + let mut compared = 0; + for snap in &snaps { + for page in &snap.pages { + let g = page_geometry(&snap.doc, *page).expect("GEO-11 accepted page"); + let visible = crop::visible_box(&snap.doc, *page); + assert!( + g.visible + .iter() + .zip(visible) + .all(|(a, b)| (a - b).abs() < 1e-3), + "GEO-11 visible {:?} vs crop.rs {visible:?}", + g.visible + ); + assert_eq!( + g.rotate, + crop::page_rotation(&snap.doc, *page), + "GEO-11 rotate" + ); + assert_eq!( + g.user_unit, + crop::page_user_unit(&snap.doc, *page), + "GEO-11 UserUnit" + ); + compared += 1; + } + } + assert!(compared >= 15, "GEO-11 compared {compared} pages"); +} + +fn geometry_of( + kind: BadBox, +) -> ( + Result<(), R>, + crate::pdf_engine::text_edit::context::SnapshotContext, +) { + let c = ctx(fx::bad_boxes(kind)); + let page = c.snap.pages[0]; + (page_geometry(&c.snap.doc, page).map(|_| ()), c) +} + +#[test] +fn geo_12_real_rotate_is_refused() { + for kind in [BadBox::RealRotate, BadBox::OddRotate] { + let (g, c) = geometry_of(kind); + assert_eq!(g, Err(R::Geometry), "GEO-12 {kind:?}"); + assert_eq!(super::model(&c, 0).page_reason, Some(R::Geometry)); + } + let (_, c) = geometry_of(BadBox::RealRotate); + assert_eq!( + crop::page_rotation(&c.snap.doc, c.snap.pages[0]), + 0, + "crop.rs ignores it silently" + ); +} + +#[test] +fn geo_13_missing_media_box_is_refused() { + let (g, c) = geometry_of(BadBox::NoMediaBox); + assert_eq!(g, Err(R::Geometry), "GEO-13"); + assert_eq!( + crop::media_box(&c.snap.doc, c.snap.pages[0]), + [0.0, 0.0, 612.0, 792.0], + "GEO-13 crop.rs would say Letter" + ); + let (g, _) = geometry_of(BadBox::NonNumberBox); + assert_eq!(g, Err(R::Geometry), "GEO-13 a box that is not four numbers"); +} + +#[test] +fn geo_14_thin_crop_is_refused() { + let (g, c) = geometry_of(BadBox::ThinCrop); + assert_eq!(g, Err(R::Geometry), "GEO-14"); + assert_eq!( + crop::visible_box(&c.snap.doc, c.snap.pages[0]), + [0.0, 0.0, 612.0, 792.0], + "GEO-14 crop.rs would silently use the MediaBox" + ); +} + +#[test] +fn geo_15_user_unit_other_than_the_number_one() { + for kind in [ + BadBox::UserUnitString, + BadBox::UserUnitNegative, + BadBox::UserUnitTwo, + ] { + let (g, c) = geometry_of(kind); + assert_eq!(g, Err(R::Geometry), "GEO-15 {kind:?}"); + let m = super::model(&c, 0); + assert_eq!( + (m.page_reason, m.runs.len()), + (Some(R::Geometry), 0), + "GEO-15 page level" + ); + } + assert_eq!( + model0(fx::user_unit(2.0)).page_reason, + Some(R::Geometry), + "GEO-15 UserUnit 2" + ); + assert_eq!( + model0(fx::user_unit(1.0)).page_reason, + None, + "GEO-15 UserUnit 1" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/joins.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/joins.rs new file mode 100644 index 0000000..14145a0 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/joins.rs @@ -0,0 +1,199 @@ +//! Engine fix pass (2026-10-03): content-part boundaries viewers read differently (review T5 H1, +//! probe R7) are refused `MALFORMED_CONTENT` in every walk, so a line commented out in Poppler +//! and pdf.js is never offered as editable; and ASCII inline images lex in linear time (review +//! T3-budget MEDIUM-4). + +use super::fuzz::thread_cpu; +use super::{ctx, model, texts, walk}; +use crate::pdf_engine::source_content::classify_source_page; +use crate::pdf_engine::text_edit::lexer::{lex_content, LexLimits}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::testkit::producers::{DocBuilder, PageSpec, HELVETICA}; +use crate::pdf_engine::text_edit::walker::WalkMode; +use std::time::Duration; + +/// A one-page document whose `/Contents` is `parts`, with Helvetica as `/F1`. +fn parts_page(parts: &[&[u8]]) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::parts(parts, &format!("/Font << /F1 {f} 0 R >>"))); + d.build() +} + +const VISIBLE: &[u8] = b"BT /F1 12 Tf 72 700 Td (Visible) Tj ET"; +const SECRET: &[u8] = b"BT /F1 12 Tf 72 680 Td (Secret) Tj ET"; + +/// Every walk of page 0 (Edit, Classify) and the model refuse the page `MALFORMED_CONTENT` (by +/// the boundary check when `joins`, else possibly by the lexer first: a name split in two often +/// leaves an op with one operand too many); the Classify pass lists no occurrence. +fn assert_refused(tag: &str, parts: &[&[u8]], joins: bool) { + let c = ctx(parts_page(parts)); + let m = model(&c, 0); + assert_eq!( + m.page_reason, + Some(R::MalformedContent), + "{tag}: {:?}", + texts(&m) + ); + let by_joins = m + .page_detail + .as_deref() + .is_some_and(|d| d.starts_with("content parts joined inside")); + assert!(by_joins || !joins, "{tag}: {:?}", m.page_detail); + assert!(m.runs.is_empty(), "{tag}"); + for mode in [WalkMode::Edit, WalkMode::Classify] { + let w = walk(&c, 0, mode.clone()); + assert_eq!(w.page_reason, Some(R::MalformedContent), "{tag} {mode:?}"); + assert!(w.records.is_empty(), "{tag} {mode:?}"); + } + let cl = classify_source_page(&c, &m, None); + assert!( + cl.occurrences.is_empty(), + "{tag}: {} occurrences", + cl.occurrences.len() + ); +} + +/// The page models with `want` as its runs. +fn assert_modelled(tag: &str, parts: &[&[u8]], want: &[&str]) { + let m = model(&ctx(parts_page(parts)), 0); + assert_eq!(m.page_reason, None, "{tag}: {:?}", m.page_detail); + assert_eq!(texts(&m), want, "{tag}"); +} + +#[test] +fn a_part_ending_in_a_comment_is_refused_unless_the_next_starts_a_line() { + // Probe R7: Poppler and pdf.js run the comment on into the next part ("Secret" is never + // drawn); the joined buffer ended it at the separator and offered "Secret" as editable. + let commented = [VISIBLE, b"\n% note" as &[u8]].concat(); + assert_refused("comment runs on", &[&commented, SECRET], true); + let percent_in_a_string = b"BT /F1 12 Tf 72 700 Td (50%) Tj ET" as &[u8]; + assert_modelled( + "a % inside a string is no comment", + &[percent_in_a_string, SECRET], + &["50%", "Secret"], + ); + let closed = [VISIBLE, b"\n% note\n" as &[u8]].concat(); + assert_modelled( + "comment closed in its part", + &[&closed, SECRET], + &["Visible", "Secret"], + ); + let next_line = [b"\n" as &[u8], SECRET].concat(); + assert_modelled( + "the next part starts a line", + &[&commented, &next_line], + &["Visible", "Secret"], + ); + // An empty part between them separates nothing. + assert_refused( + "comment over an empty part", + &[&commented, b"", SECRET], + true, + ); +} + +#[test] +fn a_part_split_inside_a_token_is_refused() { + // Each joins into valid content with a newline, and into other content without one (the + // joined buffer used to be modelled as if every viewer read the newline). + let cases: [(&str, &[u8], &[u8]); 6] = [ + ("mid-string", b"BT /F1 12 Tf 72 700 Td (Hel", b"lo) Tj ET"), + ( + "mid-string at a line end", + b"BT /F1 12 Tf 72 700 Td (Hel\n", + b"lo) Tj ET", + ), + ( + "mid-hex-string", + b"BT /F1 12 Tf 72 700 Td <4865", + b"6C6C6F> Tj ET", + ), + ( + "mid-number", + b"BT /F1 12 Tf 72 700 Td [(He) 1", + b"2 (llo)] TJ ET", + ), + ( + "mid-inline-image", + b"q BI /W 1 /H 1 /CS /G /BPC 8 ID \x80", + b" EI Q BT /F1 12 Tf 72 700 Td (Hello) Tj ET", + ), + ( + "an operator into another (`s` + `h` is pdf.js's `sh`)", + b"0 0 m 10 10 l s", + b"h BT /F1 12 Tf 72 700 Td (Hello) Tj ET", + ), + ]; + for (tag, a, b) in cases { + assert_refused(tag, &[a, b], true); + } + // Names split in two: refused, by the boundary check or by the lexer first. + for (tag, a, b) in [ + ( + "mid-name", + b"BT /F" as &[u8], + b"1 12 Tf 72 700 Td (Hello) Tj ET" as &[u8], + ), + ( + "a name and a number", + b"BT /F1", + b"12 Tf 72 700 Td (Hello) Tj ET", + ), + ] { + assert_refused(tag, &[a, b], false); + } +} + +#[test] +fn parts_split_between_tokens_still_model() { + // Operators every viewer ends at a part's end (E2E-10c's `ET`|`BT`, `Q`|`q`), separators on + // either side, an op whose operands and operator lie in different parts, an array across. + assert_modelled("ET|BT", &[VISIBLE, SECRET], &["Visible", "Secret"]); + assert_modelled( + "Q|q", + &[ + b"q 1 0 0 1 0 0 cm Q", + b"q BT /F1 12 Tf 72 700 Td (A) Tj ET Q", + ], + &["A"], + ); + assert_modelled( + "operands then operator", + &[b"BT /F1 12 Tf 72 700", b" Td (Split) Tj ET"], + &["Split"], + ); + assert_modelled( + "an array across", + &[b"BT /F1 12 Tf 72 700 Td [(Spl) -10", b"(it)] TJ ET"], + &["Split"], + ); + assert_modelled( + "a string then an operator", + &[b"BT /F1 12 Tf 72 700 Td (Hello)", b"Tj ET"], + &["Hello"], + ); +} + +#[test] +fn ascii_inline_images_without_a_terminator_lex_in_linear_time() { + // Review MEDIUM-4: each image searched the rest of the content for `>` (8,000 images: 19 s); + // one search per terminator now serves every image that follows. + for (filter, n) in [("AHx", 8_000usize), ("A85", 8_000)] { + let content = format!("BI /F /{filter} ID 0 EI n\n").repeat(n); + let started = thread_cpu(); + let ops = lex_content(content.as_bytes(), &LexLimits::page(), None); + let spent = thread_cpu().saturating_sub(started); + assert_eq!(ops.map(|o| o.len()).ok(), Some(2 * n), "{filter}"); + assert!(spent < Duration::from_secs(1), "{filter}: {spent:?}"); + } + // A terminator beyond the image cap proves nothing: the near `EI` ends the image (the far + // `> EI` in a comment used to make one image of everything up to it). + let mut c = b"BI /F /AHx ID 0 EI n ".to_vec(); + c.extend(vec![b' '; (16 << 20) + 8]); + c.extend_from_slice(b"% > EI\n"); + let ops = lex_content(&c, &LexLimits::page(), None).expect("lexes"); + assert_eq!(ops.len(), 2, "the image and `n`"); + let image = ops[0].inline_image.as_ref().expect("image"); + assert_eq!(&c[image.data.clone()], b"0", "the near EI ends it"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/pen.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/pen.rs new file mode 100644 index 0000000..aaa3a71 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/pen.rs @@ -0,0 +1,305 @@ +//! Review regressions (T3 fix pass 2, MEDIUM-1): text drawn, in the same text object and with no +//! line move, after a font whose advances the model cannot prove to be the viewers' is refused +//! `MISSING_WIDTHS` — a vertical font, a predefined CMap, a model refused before its widths were +//! read, a Type3 font whose `/FontMatrix` is not diagonal — and becomes editable again after a +//! line move. Fonts whose widths are proven (an Identity-H CID font, embedded or not, a diagonal +//! Type3) keep the text after them editable. The pen positions in the comments were measured +//! with poppler 26.04 `pdftotext -bbox` on the same pages. +//! +//! Fix pass 3 (review T3 r2): a Type3 font with a `/FontDescriptor /MissingWidth` (MEDIUM-1), an +//! Identity-H font refused `FONT_UNSUPPORTED` for its `/W` (LOW-1), and `q`/`Q` inside a text +//! object, where poppler and pdf.js disagree on the text position after `Q` (MEDIUM-3; positions +//! from `review-t3-r2/viewers.log`, poppler 26.04 and pdf.js 4.10.38). + +use super::{model0, reason_of, run_with}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::testkit::producers::{ + cid_font_with, cid_hex, DocBuilder, PageSpec, HELVETICA, +}; + +/// One text object: `first` (shown with `/FX`), then `between`, then Helvetica "Hello". +/// `font` builds `/FX` into the document and returns its object number. +fn after(font: impl FnOnce(&mut DocBuilder) -> u32, first: &str, between: &str) -> Vec { + let mut d = DocBuilder::new(); + let fx = font(&mut d); + let f1 = d.add(HELVETICA); + let content = format!("BT /FX 20 Tf 100 600 Td {first} {between} /F1 20 Tf (Hello) Tj ET"); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /FX {fx} 0 R /F1 {f1} 0 R >>"), + )); + d.build() +} + +/// "Hello" right after `first` is `expected`; after `0 -30 Td` it is editable. +fn assert_hello(font: impl Fn(&mut DocBuilder) -> u32, first: &str, expected: Option) { + let m = model0(after(&font, first, "")); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + assert_eq!(reason_of(&m, "Hello"), expected, "{first} then Hello"); + let m = model0(after(&font, first, "0 -30 Td")); + assert_eq!( + reason_of(&m, "Hello"), + None, + "{first}, a line move, then Hello" + ); +} + +/// A non-embedded CIDFontType0 Type0 font with `/Encoding encoding` (Adobe-Japan1 CIDs). +fn kozmin(d: &mut DocBuilder, encoding: &str) -> u32 { + kozmin_w(d, encoding, "[34 [612] 65 [300]]") +} + +/// `kozmin` with the descendant's `/W` array `w`. +fn kozmin_w(d: &mut DocBuilder, encoding: &str, w: &str) -> u32 { + let desc = d.add( + "<< /Type /FontDescriptor /FontName /KozMinPr6N-Regular /Flags 6 \ + /FontBBox [-437 -340 1147 1317] /ItalicAngle 0 /Ascent 1317 /Descent -349 \ + /CapHeight 742 /StemV 80 >>", + ); + let cid = d.add(format!( + "<< /Type /Font /Subtype /CIDFontType0 /BaseFont /KozMinPr6N-Regular \ + /CIDSystemInfo << /Registry (Adobe) /Ordering (Japan1) /Supplement 6 >> \ + /FontDescriptor {desc} 0 R /DW 1000 /W {w} >>" + )); + d.add(format!( + "<< /Type /Font /Subtype /Type0 /BaseFont /KozMinPr6N-Regular /Encoding {encoding} \ + /DescendantFonts [{cid} 0 R] >>" + )) +} + +/// A Type3 font drawing code 97 ("a") with `/FontMatrix matrix` and `widths`. +fn type3(d: &mut DocBuilder, matrix: &str, widths: &str) -> u32 { + let glyph = + d.b.add_stream("", b"1000 0 0 0 1000 1000 d1 0 0 1000 1000 re f"); + d.add(format!( + "<< /Type /Font /Subtype /Type3 /FontBBox [0 0 1000 1000] /FontMatrix {matrix} \ + /CharProcs << /a {glyph} 0 R >> /Encoding << /Type /Encoding /Differences [97 /a] >> \ + {widths} >>" + )) +} + +#[test] +fn text_after_a_vertical_font_is_refused_until_the_line_moves() { + // Viewers advance an Identity-V string down by w1 = −1000: poppler puts "Hello" at + // (100, 560), the model at (120, 600). + let vertical = + |d: &mut DocBuilder| cid_font_with(&mut d.b, "ABCDEF+Vertical", "AB", "/Identity-V"); + let shown = format!("<{}> Tj", cid_hex("AB", "AB")); + assert_hello(vertical, &shown, Some(R::MissingWidths)); +} + +#[test] +fn text_after_a_predefined_cmap_is_refused_until_the_line_moves() { + // UniJIS-UCS2-H maps U+0041 to CID 34 (612 wide); the model reads the code as the CID + // (/W 300): poppler puts "Hello" at x 124.48, the model at 112. The font's own refusal is + // FONT_NOT_EMBEDDED, which outranks (and hides) UNSUPPORTED_ENCODING. + let unijis = |d: &mut DocBuilder| kozmin(d, "/UniJIS-UCS2-H"); + assert_hello(unijis, "<00410041> Tj", Some(R::MissingWidths)); +} + +#[test] +fn text_after_a_font_refused_before_its_widths_is_refused_until_the_line_moves() { + // `/Widths` one entry short: FONT_UNSUPPORTED and a model with no codes (every advance 0); + // poppler uses the array: "Hello" at x 148, the model at 100. + let widths: Vec<&str> = vec!["600"; 94]; + let short = |d: &mut DocBuilder| { + d.add(format!( + "<< /Type /Font /Subtype /Type1 /BaseFont /Courier /Encoding /WinAnsiEncoding \ + /FirstChar 32 /LastChar 126 /Widths [{}] >>", + widths.join(" ") + )) + }; + assert_hello(short, "(WIDE) Tj", Some(R::MissingWidths)); + // A non-embedded Symbol font has no codes in the model either (UNSUPPORTED_ENCODING). + let symbol = |d: &mut DocBuilder| d.add("<< /Type /Font /Subtype /Type1 /BaseFont /Symbol >>"); + assert_hello(symbol, "(abc) Tj", Some(R::MissingWidths)); +} + +#[test] +fn text_after_a_type3_font_is_refused_unless_it_advances_along_the_baseline() { + let widths = "/FirstChar 97 /LastChar 97 /Widths [1000]"; + // a = 0: poppler advances by w × a = 0 ("Hello" at x 100), the model by w × 0.001 (x 140). + let rotated = |d: &mut DocBuilder| type3(d, "[0 0.001 -0.001 0 0 0]", widths); + assert_hello(rotated, "(aa) Tj", Some(R::MissingWidths)); + // b ≠ 0: the displacement leaves the baseline. + let sheared = |d: &mut DocBuilder| type3(d, "[0.001 0.0005 0 0.001 0 0]", widths); + assert_hello(sheared, "(aa) Tj", Some(R::MissingWidths)); + // A malformed /Widths (two values for one code) leaves the model without widths. + let bad = |d: &mut DocBuilder| { + type3( + d, + "[0.001 0 0 0.001 0 0]", + "/FirstChar 97 /LastChar 97 /Widths [1000 1000]", + ) + }; + assert_hello(bad, "(aa) Tj", Some(R::MissingWidths)); + // A diagonal matrix: the model advances as viewers do, so the text after it stays editable. + let diagonal = |d: &mut DocBuilder| type3(d, "[0.001 0 0 0.001 0 0]", widths); + assert_hello(diagonal, "(aa) Tj", None); +} + +#[test] +fn text_after_a_proven_cid_font_stays_editable() { + // Identity-H, embedded: the model and poppler agree on x 120. + let embedded = + |d: &mut DocBuilder| cid_font_with(&mut d.b, "ABCDEF+Horizontal", "AB", "/Identity-H"); + let shown = format!("<{}> Tj", cid_hex("AB", "AB")); + assert_hello(embedded, &shown, None); + // Identity-H, not embedded (FONT_NOT_EMBEDDED): `/W` still gives every advance. + let not_embedded = |d: &mut DocBuilder| kozmin(d, "/Identity-H"); + assert_hello(not_embedded, "<00220041> Tj", None); + let m = model0(after(not_embedded, "<00220041> Tj", "")); + let hello = super::run_with(&m, "Hello"); + // CIDs 34 and 65 are 612 and 300 wide: 100 + (0.612 + 0.3) × 20. + assert!( + super::close(hello.origin.0, 118.24), + "Hello at {:?}", + hello.origin + ); +} + +#[test] +fn text_after_a_type3_font_with_a_missing_width_is_refused_until_the_line_moves() { + let widths = "/FirstChar 97 /LastChar 97 /Widths [1000]"; + let descriptor = |missing: &str| { + format!( + "{widths} /FontDescriptor << /Type /FontDescriptor /FontName /T3 /Flags 4 \ + /FontBBox [0 0 1000 1000] /ItalicAngle 0 /Ascent 1000 /Descent 0 /CapHeight 1000 \ + /StemV 80 {missing} >>" + ) + }; + // Review P18: code 98 ("b") is outside /FirstChar../LastChar. The model gives it width 0, + // poppler and pdf.js the descriptor's /MissingWidth: "Hello" at x 140 for both viewers, + // x 120 in the model. + let missing = descriptor("/MissingWidth 1000"); + let with_missing = |d: &mut DocBuilder| type3(d, "[0.001 0 0 0.001 0 0]", &missing); + assert_hello(with_missing, "(ab) Tj", Some(R::MissingWidths)); + // An unreadable /MissingWidth is not proven either. + let odd = descriptor("/MissingWidth /Wide"); + let with_odd = |d: &mut DocBuilder| type3(d, "[0.001 0 0 0.001 0 0]", &odd); + assert_hello(with_odd, "(ab) Tj", Some(R::MissingWidths)); + // Review P18b: no descriptor, or a /MissingWidth of 0: all three agree on x 120. + let none = |d: &mut DocBuilder| type3(d, "[0.001 0 0 0.001 0 0]", widths); + assert_hello(none, "(ab) Tj", None); + let zero = descriptor("/MissingWidth 0"); + let with_zero = |d: &mut DocBuilder| type3(d, "[0.001 0 0 0.001 0 0]", &zero); + assert_hello(with_zero, "(ab) Tj", None); + let absent = descriptor(""); + let without = |d: &mut DocBuilder| type3(d, "[0.001 0 0 0.001 0 0]", &absent); + assert_hello(without, "(ab) Tj", None); +} + +#[test] +fn text_after_an_identity_h_font_with_a_bad_w_is_refused_until_the_line_moves() { + // A /W entry that is not a number: T2 refuses the font FONT_UNSUPPORTED before it has its + // widths, so the pen after its glyphs is not known even under Identity-H. + let bad_w = |d: &mut DocBuilder| kozmin_w(d, "/Identity-H", "[34 [612] 65 [(x)]]"); + assert_hello(bad_w, "<00220041> Tj", Some(R::MissingWidths)); + let m = model0(after(bad_w, "<00220041> Tj", "")); + let fx = m + .walk + .page_fonts + .iter() + .find(|(name, _)| name == b"FX") + .map(|(_, f)| f.refusal); + assert_eq!(fx, Some(Some(R::FontUnsupported))); +} + +/// Helvetica 20 as `/F1`, drawing `content`. +fn helvetica20(content: &str) -> Vec { + let mut d = DocBuilder::new(); + let f1 = d.add(HELVETICA); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F1 {f1} 0 R >>"), + )); + d.build() +} + +#[test] +fn text_after_a_q_inside_a_text_object_is_refused_until_tm() { + // Review P22 row 1: `Td` between q and Q. The model and poppler put B at (100, 550), pdf.js + // at (113.34, 600). + let m = model0(helvetica20( + "BT /F1 20 Tf 100 600 Td (A) Tj q 0 -50 Td Q (B) Tj ET", + )); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + assert_eq!(reason_of(&m, "A"), None); + assert_eq!(reason_of(&m, "B"), Some(R::MissingWidths)); + // Row 2: text shown between q and Q. pdf.js draws "Black" over "Red" (x 100), the model and + // poppler after it (x 136.68). + let m = model0(helvetica20( + "BT /F1 20 Tf 100 600 Td q 1 0 0 rg (Red) Tj Q (Black) Tj ET", + )); + assert_eq!(reason_of(&m, "Red"), None); + assert_eq!(reason_of(&m, "Black"), Some(R::MissingWidths)); + // Row 3: `Tm` between q and Q; a `Td` after Q does not settle it (it is relative to the + // disputed line start: C at (300, 270) in the model, (100, 570) in both viewers), the next + // `Tm` does. + let m = model0(helvetica20( + "BT /F1 20 Tf 1 0 0 1 100 600 Tm q 1 0 0 1 300 300 Tm Q (B) Tj 0 -30 Td (C) Tj \ + 1 0 0 1 100 400 Tm (D) Tj ET", + )); + assert_eq!(reason_of(&m, "B"), Some(R::MissingWidths)); + assert_eq!(reason_of(&m, "C"), Some(R::MissingWidths)); + assert_eq!(reason_of(&m, "D"), None); + // A new text object settles it too. + let m = model0(helvetica20( + "BT /F1 20 Tf 100 600 Td q 0 -50 Td Q (B) Tj ET BT 100 500 Td (C) Tj ET", + )); + assert_eq!(reason_of(&m, "B"), Some(R::MissingWidths)); + assert_eq!(reason_of(&m, "C"), None); + // A Q inside BT restoring a state saved before BT (under the identity text matrix). + let m = model0(helvetica20( + "BT /F1 20 Tf ET q BT 100 600 Td (A) Tj Q (B) Tj ET", + )); + assert_eq!(reason_of(&m, "A"), None); + assert_eq!(reason_of(&m, "B"), Some(R::MissingWidths)); +} + +#[test] +fn a_q_inside_a_text_object_that_moves_nothing_keeps_the_text_editable() { + // Only the colour changes between q and Q: every viewer keeps the text position, and the + // restored state is the saved one, so A and B join into one editable run. + let m = model0(helvetica20( + "BT /F1 20 Tf 100 600 Td (A) Tj q 1 0 0 rg Q (B) Tj ET", + )); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + let ab = run_with(&m, "AB"); + assert_eq!(ab.reason, None); + assert_eq!(ab.members.len(), 2); + // q/Q outside text objects never touch the text position. + let m = model0(helvetica20( + "q BT /F1 20 Tf 100 600 Td (A) Tj ET Q BT /F1 20 Tf 100 500 Td (B) Tj ET", + )); + assert_eq!(reason_of(&m, "A"), None); + assert_eq!(reason_of(&m, "B"), None); +} + +#[test] +fn a_q_inside_a_text_object_restoring_another_tm_base_is_refused_until_tm() { + // Review r3 LOW-1 (`qdecomp1`): the `q` was made after a text object whose `Tm` set the base, + // and the `Q` comes inside a text object positioned by `Td` from the identity. The composite + // matrices agree, but poppler's restored text matrix with its kept line start puts C at + // (200, 1170); the model and pdf.js at (100, 570). B and C are refused until a `Tm`. + let m = model0(helvetica20( + "BT /F1 20 Tf 1 0 0 1 100 600 Tm (A) Tj ET q BT 100 600 Td (A) Tj Q (B) Tj \ + 0 -30 Td (C) Tj 1 0 0 1 100 400 Tm (D) Tj ET", + )); + assert_eq!(m.page_reason, None, "{:?}", m.page_detail); + assert_eq!(reason_of(&m, "B"), Some(R::MissingWidths)); + assert_eq!(reason_of(&m, "C"), Some(R::MissingWidths)); + assert_eq!(reason_of(&m, "D"), None); + // The two "A"s (B no longer joins the second) are drawn on each other: DUPLICATE_TEXT. + let a: Vec<_> = m.runs.iter().filter(|r| r.text == "A").collect(); + assert_eq!(a.len(), 2); + assert!(a.iter().all(|r| r.reason == Some(R::DuplicateText))); + // Control: both text objects start from the identity, so every viewer agrees and A and B + // join into one editable run. + let m = model0(helvetica20( + "BT /F1 20 Tf 100 600 Td (A) Tj ET q BT 100 600 Td (A) Tj Q (B) Tj ET", + )); + let ab = run_with(&m, "AB"); + assert_eq!(ab.reason, None); + assert_eq!(ab.members.len(), 2); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/producers.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/producers.rs new file mode 100644 index 0000000..9833f0f --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/producers.rs @@ -0,0 +1,379 @@ +//! Smoke tests of every producer-shaped fixture (§E.2): each builder opens through the +//! production snapshot reader and shows the behaviour §A.9 expects of it. + +use super::{ctx, model, model0, reason_of, walk}; +use crate::error::AppError; +use crate::pdf_engine::text_edit::engines::{qpdf_page_map, RunOpts}; +use crate::pdf_engine::text_edit::fonts::face_surface; +use crate::pdf_engine::text_edit::reasons::{Face, TextReason as R}; +use crate::pdf_engine::text_edit::snapshot::{ + check_page_map, read_snapshot, snapshot_from_bytes, SourceSnapshot, +}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, BadBox, ClipKind, InlineProofKind, LegacyFilter, +}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use crate::pdf_engine::text_edit::walker::{PaintKind, WalkMode}; +use std::path::Path; + +type Case = (&'static str, Vec, &'static str, Option); + +fn check(cases: Vec) { + for (id, pdf, text, want) in cases { + let m = model0(pdf); + assert_eq!(m.page_reason, None, "{id}: page {:?}", m.page_detail); + assert_eq!(reason_of(&m, text), want, "{id}: {text}"); + } +} + +#[test] +fn producers_office_smoke() { + check(vec![ + ("FX-WORD", fx::word(), "Invoice 2026", None), + ("FX-WORD", fx::word(), "Due", None), + ("FX-WORD-TR", fx::word_tr(), "Sağlık Bakanlığı Raporu", None), + ("FX-LIBRE", fx::libre(), "Libre text", None), + ("FX-LIBRE", fx::libre(), "text", None), + ("FX-LIBRE-CFF", fx::libre_cff(), "Hello Hello", None), + ("FX-SKIA", fx::skia(), "Chrome", None), + ("FX-SKIA", fx::skia(), "Bold", None), + ("FX-SKIA", fx::skia(), "Italic", None), + ("FX-QUARTZ", fx::quartz(), "Quartz", None), + ("FX-PERGLYPH", fx::per_glyph(None), "Per glyph line", None), + ("FX-PDFTEX", fx::pdftex(), "Hello World", None), + ("FX-XETEX", fx::xetex(), "XeTeX", None), + ("FX-INDD", fx::indd(), "visible layer", None), + ( + "FX-INDD", + fx::indd(), + "hidden layer", + Some(R::OptionalContent), + ), + ("FX-STD14", fx::std14(), "Helvetica line", None), + ("FX-STD14", fx::std14(), "Times line", None), + ("FX-STD14", fx::std14(), "Courier line", None), + ("FX-NONEMB", fx::nonemb(), "Not embedded", None), + ("FX-OCR", fx::ocr(), "scanned words", Some(R::InvisibleText)), + ]); + let skia = model0(fx::skia()); + let bold = skia.runs.iter().find(|r| r.text == "Bold").expect("Bold"); + assert_eq!(bold.tr, 2, "FX-SKIA synthetic bold stays Tr 2"); + let nonemb = model0(fx::nonemb()); + assert!(nonemb.runs[0].substituted, "FX-NONEMB is substituted"); + let indd = ctx(fx::indd()); + let classify = walk(&indd, 0, WalkMode::Classify); + assert!( + classify.records.iter().any(|r| r.depth == 1), + "FX-INDD text in a Form" + ); + let r = model(&indd, 0); + assert!( + r.runs + .iter() + .any(|r| r.text == "visible layer" && r.tc == 0.12), + "FX-INDD Tc" + ); +} + +#[test] +fn producers_shared_smoke() { + let c = ctx(fx::shared()); + assert_eq!(c.snap.pages.len(), 4); + assert_eq!( + reason_of(&model(&c, 0), "Shared body"), + Some(R::SharedContent), + "FX-SHARED p1" + ); + assert_eq!( + reason_of(&model(&c, 1), "Shared body"), + Some(R::SharedContent), + "FX-SHARED p2" + ); + let p3 = model(&c, 2); + assert_eq!( + reason_of(&p3, "Letterhead"), + Some(R::SharedContent), + "FX-SHARED letterhead" + ); + assert_eq!(reason_of(&p3, "Body three"), None, "FX-SHARED own part"); +} + +#[test] +fn producers_edges_smoke() { + let mut cases: Vec = vec![ + ("cropped_offset", fx::cropped_offset(), "Cropped page", None), + ("two_parts_mid_bt", fx::two_parts_mid_bt(), "Hi", None), + ("two_parts_mid_bt", fx::two_parts_mid_bt(), "Lo", None), + ( + "straddling_op", + fx::straddling_op(), + "Hello", + Some(R::SplitContent), + ), + ( + "comments_and_odd_ws", + fx::comments_and_odd_ws(), + "Odd ws", + None, + ), + ("quote_ops", fx::quote_ops(), "Line one", None), + ("quote_ops", fx::quote_ops(), "Line two", None), + ("quote_ops", fx::quote_ops(), "Line three", None), + ( + "kerned_then_tail", + fx::kerned_then_tail(), + "AB CDtail", + None, + ), + ("tf1_tm12", fx::tf1_tm12(), "Hi", None), + ("tz_tc_tw_ts", fx::tz_tc_tw_ts(), "a b", None), + ("extgstate_font", fx::extgstate_font(), "Hi there", None), + ("mac_roman", fx::mac_roman(), "café", None), + ("subset_without_y", fx::subset_without_y(), "Hello", None), + ( + "actual_text_span", + fx::actual_text_span(), + "fi", + Some(R::ActualText), + ), + ( + "actual_text_struct", + fx::actual_text_struct(), + "Struct", + Some(R::ActualText), + ), + ( + "duplicate_shadow", + fx::duplicate_shadow(), + "Shadow", + Some(R::DuplicateText), + ), + ("clip_page", fx::clip(ClipKind::Page), "Clip me", None), + ( + "clip_small", + fx::clip(ClipKind::Small), + "Clip me", + Some(R::Clipped), + ), + ( + "clip_curve", + fx::clip(ClipKind::Curve), + "Clip me", + Some(R::Clipped), + ), + ( + "pattern_fill", + fx::pattern_fill(), + "Pattern", + Some(R::Pattern), + ), + ("smask_text", fx::smask_text(), "Masked", Some(R::SoftMask)), + ("tr7", fx::tr7(), "Clip text", Some(R::TextClipMode)), + ("identity_v", fx::identity_v(), "Up", Some(R::Vertical)), + ("type3", fx::type3(), "a", Some(R::Type3)), + ("mirrored", fx::mirrored(), "Mirror", Some(R::MirroredText)), + ( + "negative_tz", + fx::negative_tz(), + "Mirror", + Some(R::MirroredText), + ), + ( + "negative_tf", + fx::negative_tf(), + "Turned", + Some(R::RotatedText), + ), + ("skewed", fx::skewed(), "Skewed", Some(R::SkewedText)), + ("oblique", fx::oblique(), "Oblique", None), + ]; + for angle in [90, 180, 270] { + cases.push(( + "rotated(counter)", + fx::rotated(angle, true), + "Rotated page", + None, + )); + cases.push(( + "rotated", + fx::rotated(angle, false), + "Rotated page", + Some(R::RotatedText), + )); + } + for proof in [ + InlineProofKind::Length, + InlineProofKind::Unfiltered, + InlineProofKind::Flate, + ] { + cases.push(( + "inline_image", + fx::inline_image(proof), + "After the image", + None, + )); + } + cases.push(( + "inline_image(Dct)", + fx::inline_image(InlineProofKind::Dct), + "After the image", + Some(R::InlineImage), + )); + check(cases); + assert_eq!( + model0(fx::user_unit(2.0)).page_reason, + Some(R::Geometry), + "user_unit(2)" + ); + let y = model0(fx::subset_without_y()); + assert!( + !y.surface(&y.runs[0]).alphabet().contains(&'Y'), + "subset_without_y: Y not typeable" + ); + let nested = ctx(fx::nested_form()); + assert!( + model(&nested, 0).runs.is_empty(), + "nested_form: no Edit run" + ); + assert_eq!( + walk(&nested, 0, WalkMode::Classify).records.len(), + 1, + "nested_form" + ); +} + +fn snap_code(bytes: Vec) -> String { + match snapshot_from_bytes(Path::new("p.pdf"), bytes, None) { + Ok(_) => "OK".into(), + Err(e) => e.code, + } +} + +#[test] +fn producers_files_smoke() { + assert_eq!(snap_code(fx::deep_nesting(100)), "OK", "deep_nesting(100)"); + assert_eq!( + snap_code(fx::deep_nesting(101)), + "FILE_TOO_COMPLEX", + "deep_nesting(101)" + ); + assert_eq!( + snap_code(fx::objstm_bomb()), + "FILE_TOO_COMPLEX", + "objstm_bomb" + ); + assert_eq!(snap_code(fx::xref_bomb()), "FILE_TOO_COMPLEX", "xref_bomb"); + assert_eq!(snap_code(fx::signed()), "SIGNED", "FX-SIGNED"); + assert_eq!(snap_code(fx::xfa()), "UNSUPPORTED_XFA", "FX-XFA"); + let bomb = ctx(fx::flate_bomb()); + assert_eq!( + model(&bomb, 0).page_reason, + Some(R::PageTooComplex), + "flate_bomb" + ); + assert_eq!( + reason_of(&model(&bomb, 1), "Hello"), + None, + "flate_bomb: page 2 unaffected" + ); + let ext = model0(fx::extensions_indirect()); + assert_eq!(reason_of(&ext, "Hello"), None, "extensions_indirect"); + for kind in [ + BadBox::NoMediaBox, + BadBox::RealRotate, + BadBox::OddRotate, + BadBox::ThinCrop, + BadBox::UserUnitString, + BadBox::UserUnitNegative, + BadBox::UserUnitTwo, + BadBox::NonNumberBox, + ] { + assert_eq!( + model0(fx::bad_boxes(kind)).page_reason, + Some(R::Geometry), + "bad_boxes({kind:?})" + ); + } + for filter in [LegacyFilter::RunLength, LegacyFilter::Lzw] { + let c = ctx(fx::legacy_filter_page(filter)); + assert_eq!( + model(&c, 0).page_reason, + Some(R::UnsupportedFilter), + "legacy {filter:?}" + ); + assert_eq!(reason_of(&model(&c, 1), "Hello"), None); + } + let shared = ctx(fx::shared_inherited_resources()); + let p1 = model(&shared, 0); + assert_eq!( + reason_of(&p1, "Regular words"), + None, + "shared_inherited_resources" + ); + let primary = &p1.walk.page_fonts[0].1; + assert!( + face_surface(&p1.walk.page_fonts, primary, Face::Bold).is_some(), + "shared_inherited_resources: the bold sibling used only on page 2 is a page font" + ); + let (a, b) = (ctx(fx::swapped_image(false)), ctx(fx::swapped_image(true))); + let hash = |c: &crate::pdf_engine::text_edit::context::SnapshotContext| { + walk(c, 0, WalkMode::Edit) + .paints + .into_iter() + .find_map(|p| match p.kind { + PaintKind::ImageXObject { hash, .. } => Some(hash), + _ => None, + }) + }; + assert_ne!(hash(&a), hash(&b), "swapped_image"); + let two = model0(fx::two_column(false)); + assert_eq!(two.runs.len(), 7, "two_column"); + let stroke = model0(fx::stroke_text(2, [0.5, 1.0], ["[] 0", "[] 0"])); + assert_eq!(stroke.runs.len(), 2, "stroke_text"); + assert_eq!( + reason_of(&model0(fx::tagged_bookmarked_page()), "Tagged line"), + None + ); +} + +fn page_map(bytes: &[u8], id: &str) -> Option<(usize, String)> { + let engines = engines_or_skip(id)?; + let s = Scratch::new("producers-map"); + let p = s.write("f.pdf", bytes); + let snap: SourceSnapshot = read_snapshot(&p).unwrap_or_else(|e: AppError| panic!("{id}: {e}")); + let pages = + qpdf_page_map(&engines, &p, &RunOpts::default()).unwrap_or_else(|e| panic!("{id}: {e}")); + let code = match check_page_map(&snap, &pages) { + Ok(()) => "OK".to_string(), + Err(e) => e.code, + }; + Some((snap.pages.len(), code)) +} + +#[test] +fn producers_engine_backed_smoke() { + let Some(engines) = engines_or_skip("producers_engine_backed_smoke") else { + return; + }; + assert_eq!(snap_code(fx::encrypted(&engines)), "ENCRYPTED", "FX-ENC"); + let word = fx::hybrid_xref(true); + assert_eq!( + page_map(&word, "hybrid_xref(true)"), + Some((1, "OK".to_string())) + ); + assert_eq!( + reason_of(&model0(word), "Hello"), + None, + "hybrid_xref(true) editable" + ); + assert_eq!( + page_map(&fx::hybrid_xref(false), "hybrid_xref(false)"), + Some((0, "PDF_NEEDS_REPAIR".to_string())), + "hybrid_xref(false): lopdf misses the page" + ); + assert_eq!( + page_map(&fx::kid_without_type(), "kid_without_type"), + Some((1, "PDF_NEEDS_REPAIR".to_string())), + "kid_without_type" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/runs.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/runs.rs new file mode 100644 index 0000000..97b12ee --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/runs.rs @@ -0,0 +1,302 @@ +//! RUN-01…11: joining (Word per-format blocks, column gaps, paints between, fonts, siblings with +//! converted gaps, MCIDs), synthetic spaces, ligatures, reading order on a rotated page, caret +//! offsets and run ids. + +use super::{close, ctx, model, model0, reason_of, run_with, texts}; +use crate::pdf_engine::text_edit::runs::{KernSrc, SpaceMode, Unit}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_doc, helvetica_page, DocBuilder, PageSpec, HELVETICA, +}; + +#[test] +fn run_01_word_per_format_blocks_join() { + let m = model0(fx::word()); + assert_eq!(texts(&m), ["Invoice 2026", "Due"], "RUN-01"); + let r = run_with(&m, "Invoice 2026"); + assert_eq!(r.members.len(), 2, "RUN-01 two BT blocks, one line"); + assert_eq!(r.reason, None); + assert_eq!(r.space_mode, SpaceMode::Glyph); + assert!( + r.id.starts_with(&format!("t1:{}:0:", m.fingerprint)), + "RUN-01 id {}", + r.id + ); +} + +#[test] +fn run_02_column_gap_does_not_join() { + // Helvetica "Left" = 556 + 556 + 278 + 278 = 1.668 em → 20.016 pt at 12 pt; `Td` moves from + // the line start, so the next op starts 6 pt (0.5 em) after the pen. + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Left) Tj 26.016 0 Td (Right) Tj ET", + )); + assert_eq!(texts(&m), ["Left", "Right"], "RUN-02 0.5 em gap"); + let joined = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Left) Tj 21.216 0 Td (Right) Tj ET", + )); + assert_eq!(texts(&joined), ["LeftRight"], "RUN-02 a 0.1 em gap joins"); + let r = run_with(&joined, "LeftRight"); + assert!( + r.units.iter().any(|u| matches!( + u, + Unit::Kern { + src: KernSrc::Gap, + synth_space: false, + .. + } + )), + "RUN-02 the joined gap is a converted kern" + ); +} + +#[test] +fn run_03_a_paint_between_prevents_the_join() { + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Left) Tj ET 0 0 1 1 re f BT /F1 12 Tf 92.016 700 Td (Right) Tj ET", + )); + assert_eq!(texts(&m), ["Left", "Right"], "RUN-03"); + let no_paint = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Left) Tj ET BT /F1 12 Tf 92.016 700 Td (Right) Tj ET", + )); + assert_eq!(texts(&no_paint), ["LeftRight"], "RUN-03 control"); +} + +#[test] +fn run_04_non_sibling_fonts_do_not_join() { + let mut d = DocBuilder::new(); + let f1 = d.add(HELVETICA); + let f2 = d + .add("<< /Type /Font /Subtype /Type1 /BaseFont /Times-Roman /Encoding /WinAnsiEncoding >>"); + d.page(PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Sans) Tj /F2 12 Tf (Serif) Tj ET", + &format!("/Font << /F1 {f1} 0 R /F2 {f2} 0 R >>"), + )); + let m = model0(d.build()); + assert_eq!(texts(&m), ["Sans", "Serif"], "RUN-04 Helvetica + Times"); +} + +#[test] +fn run_05_word_turkish_siblings_join_with_gaps_as_kerns() { + let m = model0(fx::word_tr()); + let r = run_with(&m, "Sağlık Bakanlığı Raporu"); + assert_eq!(r.reason, None, "RUN-05 editable: {:?}", r.reasons); + assert_eq!(r.members.len(), 7, "RUN-05 seven segments, one line"); + assert_eq!( + r.surface.to_vec(), + vec![b"F1".to_vec(), b"F2".to_vec()], + "RUN-05 sibling surface" + ); + let gaps: Vec = r + .units + .iter() + .filter_map(|u| match u { + Unit::Kern { + value, + src: KernSrc::Gap, + .. + } => Some(*value), + _ => None, + }) + .collect(); + // Segment 4 starts 0.01 pt late, segment 5 on time: ∓0.01 pt / 11 pt × 1000. + assert_eq!(gaps.len(), 2, "RUN-05 converted gaps {gaps:?}"); + assert!( + close(gaps[0], -10.0 / 11.0) && close(gaps[1], 10.0 / 11.0), + "RUN-05 {gaps:?}" + ); + let alphabet = m.surface(r).alphabet(); + for ch in ['ğ', 'ş', 'ı', 'İ', 'S', 'a'] { + assert!( + alphabet.contains(&ch), + "RUN-05 {ch} typeable through the surface" + ); + } +} + +#[test] +fn run_06_different_mcids_do_not_join() { + let m = model0(helvetica_page( + b"/P <> BDC BT /F1 12 Tf 72 700 Td (Left) Tj ET EMC \ + /P <> BDC BT /F1 12 Tf 92.016 700 Td (Right) Tj ET EMC", + )); + assert_eq!(texts(&m), ["Left", "Right"], "RUN-06"); +} + +#[test] +fn run_07_pdftex_kerned_word_gaps_read_as_spaces() { + let m = model0(fx::pdftex()); + let r = run_with(&m, "Hello World"); + assert_eq!(r.reason, None, "RUN-07 editable: {:?}", r.reasons); + assert_eq!( + r.space_mode, + SpaceMode::Kern, + "RUN-07 no space glyph: kern-space writing" + ); + assert_eq!(r.kern_space, -333.0, "RUN-07 the run's word gap"); + let synth = r + .units + .iter() + .filter(|u| { + matches!( + u, + Unit::Kern { + synth_space: true, + .. + } + ) + }) + .count(); + assert_eq!(synth, 1, "RUN-07 one synthetic space"); +} + +#[test] +fn run_08_ligatures_expand() { + let m = model0(fx::pdftex()); + let r = run_with(&m, "find don\u{2019}t"); + let lig = r.units.iter().find_map(|u| match u { + Unit::Glyph { text, code, .. } if code.value == 12 => Some(text.clone()), + _ => None, + }); + assert_eq!( + lig.as_deref(), + Some("fi"), + "RUN-08 one glyph reads as two letters" + ); + assert_eq!(r.caret_offsets.len(), r.text.chars().count() + 1); +} + +#[test] +fn run_09_reading_order_on_a_rotated_page() { + // /Rotate 90 with counter-rotated lines: user x grows downwards on screen. + let page = helvetica_doc( + b"BT /F1 12 Tf 0 1 -1 0 320 300 Tm (Second) Tj ET BT /F1 12 Tf 0 1 -1 0 300 300 Tm (First) Tj ET", + "", + "/Rotate 90", + ); + let m = model0(page); + assert_eq!(texts(&m), ["First", "Second"], "RUN-09 display order"); + assert_eq!((m.runs[0].line, m.runs[1].line), (0, 1)); + assert!(m.runs.iter().all(|r| r.reason.is_none())); +} + +#[test] +fn run_10_caret_offsets_are_monotonic() { + for pdf in [ + fx::word(), + fx::pdftex(), + fx::word_tr(), + fx::kerned_then_tail(), + ] { + for r in &model0(pdf).runs { + assert_eq!( + r.caret_offsets.len(), + r.text.chars().count() + 1, + "RUN-10 {}", + r.text + ); + assert!( + r.caret_offsets.windows(2).all(|w| w[0] <= w[1]), + "RUN-10 monotonic {:?}", + r.caret_offsets + ); + } + } + let m = model0(fx::word()); + let r = run_with(&m, "Invoice 2026"); + // 12 glyphs of 5.52 pt (the 12/1000 kern after "v" pulls 0.13 pt back). + let last = *r.caret_offsets.last().expect("offsets"); + assert!( + (last - (12.0 * 5.52 - 0.13248)).abs() < 0.001, + "RUN-10 end {last}" + ); +} + +#[test] +fn run_11_ids_are_stable_and_follow_the_bytes() { + let (a, b) = (ctx(fx::word()), ctx(fx::word())); + let ids = |m: &crate::pdf_engine::text_edit::runs::PageModel| { + m.runs.iter().map(|r| r.id.clone()).collect::>() + }; + assert_eq!(ids(&model(&a, 0)), ids(&model(&b, 0)), "RUN-11 stable"); + let shifted = model0(helvetica_page(b"q Q BT /F1 12 Tf 72 700 Td (Hi) Tj ET")); + let plain = model0(helvetica_page(b"BT /F1 12 Tf 72 700 Td (Hi) Tj ET")); + assert_ne!(ids(&shifted), ids(&plain), "RUN-11 other bytes, other id"); + assert!( + plain.runs[0].id.ends_with(":0:23-30"), + "RUN-11 span in the id: {}", + plain.runs[0].id + ); + assert_eq!( + plain.run(&plain.runs[0].id).map(|r| r.text.as_str()), + Some("Hi") + ); + assert_eq!(reason_of(&plain, "Hi"), None); +} + +#[test] +fn run_fields_describe_the_line() { + use crate::pdf_engine::text_edit::reasons::Face; + use crate::pdf_engine::text_edit::structure::struct_actual_text; + let m = model0(helvetica_page( + b"1 0 0 rg BT /F1 10 Tf 2 Tw 90 Tz 72 700 Td [(Ab) -50 (c)] TJ ET", + )); + let r = run_with(&m, "Abc"); + assert_eq!((r.tw, r.th, r.face), (2.0, 0.9, Face::Regular)); + assert_eq!(r.fill_hex.as_deref(), Some("#ff0000")); + assert!( + close(r.visible_extent, 540.0), + "visible extent {}", + r.visible_extent + ); + // `end_pen` and `PageWalk.page_index` are gone (dead fields, fix pass 2026-10-03): the pen + // after the line is the last member's `pen_after`, and the page is `PageModel.page_index`. + let rec = &m.walk.records[0]; + assert!(rec.tm_after[4] > rec.tm_before[4], "Tm moved by the show"); + assert_eq!(rec.operand_spans.len(), 1, "the TJ array"); + assert_eq!(m.page_index, 0); + let kern = r.units.iter().find_map(|u| match u { + Unit::Kern { + src: KernSrc::Tj { span }, + value, + .. + } => Some((span.clone(), *value)), + _ => None, + }); + let (span, value) = kern.expect("TJ kern"); + assert_eq!( + (value, &m.content.joined[span]), + (-50.0, &b"-50"[..]), + "the kern's own bytes" + ); + let first = r.units.iter().find_map(|u| match u { + Unit::Glyph { + member, + font_res, + font_hash, + width1000, + .. + } => Some((*member, font_res.clone(), *font_hash, *width1000)), + _ => None, + }); + let (member, res, hash, width) = first.expect("glyph"); + assert_eq!( + (member, res.as_deref(), width, r.tfs), + (0, Some(&b"F1"[..]), 667.0, 10.0) + ); + assert_eq!(hash, m.walk.page_fonts[0].1.content_hash); + let tagged = ctx(fx::actual_text_struct()); + let page = tagged.snap.pages[0]; + assert_eq!(struct_actual_text(tagged.doc(), page, 0), Ok(true)); + assert_eq!(struct_actual_text(tagged.doc(), page, 1), Ok(false)); + let page = helvetica_page(b"0 0 10 10 re f BT /F1 12 Tf 72 700 Td (x) Tj ET"); + let w = super::walk( + &ctx(page), + 0, + crate::pdf_engine::text_edit::walker::WalkMode::Edit, + ); + assert_eq!( + w.paints[0].span, + Some(13..14), + "paint spans in the joined buffer (the `f` op)" + ); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/runs2.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/runs2.rs new file mode 100644 index 0000000..0d8986a --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/runs2.rs @@ -0,0 +1,273 @@ +//! RUN-12…22: per-glyph pages, duplicates, blank runs, the run budget, refused runs joined per +//! reason, reading order (XY-cut, structure order, all-or-nothing), state-sensitive joins, +//! ExtGState fonts, and tagged/bookmarked Word pages that stay editable. + +use super::{model0, reason_of, texts}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_page, DocBuilder, PageSpec, HELVETICA, TWO_COLUMN_STRUCT_ORDER, +}; + +#[test] +fn run_12_per_glyph_page() { + let joined = model0(fx::per_glyph(None)); + assert_eq!( + texts(&joined), + ["Per glyph line"], + "RUN-12 edge-to-edge glyphs join" + ); + assert_eq!(reason_of(&joined, "Per glyph line"), None); + let scattered = model0(fx::per_glyph(Some(30.0))); + assert_eq!( + scattered.runs.len(), + 12, + "RUN-12 12 non-blank one-glyph runs (2 spaces)" + ); + assert!( + scattered + .runs + .iter() + .all(|r| r.reason == Some(R::PerGlyphText)), + "RUN-12 {:?}", + scattered + .runs + .iter() + .map(|r| (r.text.clone(), r.reasons.clone())) + .collect::>() + ); +} + +#[test] +fn run_13_duplicate_text() { + let m = model0(fx::duplicate_shadow()); + assert_eq!(m.runs.len(), 2); + assert!( + m.runs.iter().all(|r| r.reason == Some(R::DuplicateText)), + "RUN-13 both refused" + ); + let apart = model0(helvetica_page( + b"BT /F1 12 Tf 72 720 Td (Shadow) Tj ET BT /F1 12 Tf 72 600 Td (Shadow) Tj ET", + )); + assert!( + apart.runs.iter().all(|r| r.reason.is_none()), + "RUN-13 no overlap, no refusal" + ); +} + +#[test] +fn run_14_whitespace_only_runs_are_omitted() { + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 720 Td (Words) Tj ET BT /F1 12 Tf 72 600 Td ( ) Tj ET", + )); + assert_eq!(texts(&m), ["Words"], "RUN-14"); + assert_eq!( + m.walk.records.len(), + 2, + "RUN-14 the blank op stays in the walk" + ); + // (`PageModel.record_run` is gone, fix pass 2026-10-03: nothing read it; the blank op's + // record belongs to no listed run.) + assert!(m.runs.iter().all(|r| r.members == [0])); +} + +#[test] +fn run_15_too_many_runs_refuse_the_page() { + let lines = crate::pdf_engine::text_edit::limits::RUNS_PER_PAGE_MAX + 1; + let mut content = String::from("BT /F1 1 Tf 20 780 Td "); + for _ in 0..lines { + content.push_str("(a) Tj 0 -0.03 Td "); + } + content.push_str("ET"); + let m = model0(helvetica_page(content.as_bytes())); + assert_eq!( + m.page_reason, + Some(R::PageTooComplex), + "RUN-15 {:?}", + m.page_detail + ); + assert!(m.runs.is_empty()); +} + +#[test] +fn run_16_refused_runs_join_only_with_the_same_reason() { + let m = model0(helvetica_page( + b"BT 3 Tr /F1 12 Tf 72 700 Td (Invis) Tj (ible) Tj ET", + )); + assert_eq!(texts(&m), ["Invisible"], "RUN-16 same reason joins"); + assert_eq!(reason_of(&m, "Invisible"), Some(R::InvisibleText)); + // A shared part followed by an own part: same line, same state, different first reason. + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let res = format!("/Font << /F1 {f} 0 R >>"); + let shared = d.b.add_stream("", b"BT /F1 12 Tf 72 700 Td (Head) Tj"); + let tail1 = d.b.add_stream("", b"(Tail) Tj ET"); + let tail2 = d.b.add_stream("", b"(Tail) Tj ET"); + let spec = PageSpec::new(b"", &res); + d.page_raw(&format!("[{shared} 0 R {tail1} 0 R]"), &spec); + d.page_raw(&format!("[{shared} 0 R {tail2} 0 R]"), &spec); + let m = model0(d.build()); + assert_eq!( + texts(&m), + ["Head", "Tail"], + "RUN-16 different reasons do not join" + ); + assert_eq!(reason_of(&m, "Head"), Some(R::SharedContent)); + assert_eq!(reason_of(&m, "Tail"), None); +} + +#[test] +fn run_17_untagged_two_columns_read_by_xy_cut() { + let m = model0(fx::two_column(false)); + assert_eq!( + texts(&m), + [ + "A two column page", + "Left one", + "Left two", + "Left three", + "Right one", + "Right two", + "Right three" + ], + "RUN-17 title, column 1, column 2" + ); + let lines: Vec = m.runs.iter().map(|r| r.line).collect(); + assert_eq!(lines, [0, 1, 2, 3, 4, 5, 6], "RUN-17 one line each"); +} + +#[test] +fn run_18_tagged_page_uses_structure_order() { + let m = model0(fx::two_column(true)); + let by_mcid = |mcid: i64| { + crate::pdf_engine::text_edit::testkit::producers::TWO_COLUMN_LINES + .iter() + .find(|l| l.0 == mcid) + .map(|l| l.4) + .expect("line") + }; + let want: Vec<&str> = TWO_COLUMN_STRUCT_ORDER + .iter() + .map(|m| by_mcid(*m)) + .collect(); + assert_eq!(texts(&m), want, "RUN-18 structure order"); +} + +#[test] +fn run_19_one_untagged_run_falls_back_to_xy_cut() { + let m = model0(fx::two_column_with(true, true)); + assert_eq!( + texts(&m).first().map(String::as_str), + Some("A two column page"), + "RUN-19 XY-cut for the whole page: {:?}", + texts(&m) + ); + assert_eq!( + texts(&m), + texts(&model0(fx::two_column(false))), + "RUN-19 never a mix" + ); +} + +#[test] +fn run_20_stroked_segments_with_different_line_state_do_not_join() { + let same = model0(fx::stroke_text(2, [0.5, 0.5], ["[] 0", "[] 0"])); + assert_eq!(texts(&same), ["Boldface"], "RUN-20 equal state joins"); + let widths = model0(fx::stroke_text(2, [0.5, 1.0], ["[] 0", "[] 0"])); + assert_eq!(texts(&widths), ["Bold", "face"], "RUN-20 line width"); + let dashes = model0(fx::stroke_text(1, [0.5, 0.5], ["[] 0", "[2 1] 0"])); + assert_eq!(texts(&dashes), ["Bold", "face"], "RUN-20 dash"); + let colours = model0(helvetica_page( + b"0 0 1 RG BT /F1 12 Tf 2 Tr 72 700 Td (Bold) Tj ET 1 0 0 RG BT /F1 12 Tf 2 Tr 96.012 700 Td (face) Tj ET", + )); + assert_eq!(texts(&colours), ["Bold", "face"], "RUN-20 stroke colour"); +} + +#[test] +fn run_21_extgstate_font_never_joins_a_tf_font() { + let page = { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + d.page(PageSpec::new( + b"/GS1 gs BT 72 700 Td (Hi) Tj ET BT /F1 12 Tf 83.328 700 Td (there) Tj ET", + &format!("/Font << /F1 {f} 0 R >> /ExtGState << /GS1 << /Font [{f} 0 R 12] >> >>"), + )); + d.build() + }; + let m = model0(page); + assert_eq!(texts(&m), ["Hi", "there"], "RUN-21"); + let control = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Hi) Tj ET BT /F1 12 Tf 83.328 700 Td (there) Tj ET", + )); + assert_eq!( + texts(&control), + ["Hithere"], + "RUN-21 control: two Tf blocks join" + ); +} + +#[test] +fn run_22_tagged_bookmarked_word_page_stays_editable() { + for pdf in [fx::word(), fx::word_tr(), fx::tagged_bookmarked_page()] { + let m = model0(pdf); + assert!(!m.runs.is_empty()); + for r in &m.runs { + assert_eq!(r.reason, None, "RUN-22 {}: {:?}", r.text, r.reasons); + assert!(!r.reasons.contains(&R::SharedContent)); + } + } + // A page an internal link points back to: the page object is referenced twice. + let with_link = { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let page = d.reserve(); + let link = d.add(format!( + "<< /Type /Annot /Subtype /Link /Rect [72 690 140 714] /Border [0 0 0] \ + /Dest [{page} 0 R /XYZ 0 792 0] >>" + )); + d.page_at( + page, + PageSpec::new( + b"BT /F1 12 Tf 72 700 Td (Linked) Tj ET", + &format!("/Font << /F1 {f} 0 R >>"), + ) + .with(&format!("/Annots [{link} 0 R]")), + ); + d.build() + }; + assert_eq!(reason_of(&model0(with_link), "Linked"), None); +} + +#[test] +fn kern_sequences_sum_and_negligible_gaps_vanish() { + use super::run_with; + use crate::pdf_engine::text_edit::runs::Unit; + let kerns = |u: &[Unit]| -> Vec<(f64, bool)> { + u.iter() + .filter_map(|u| match u { + Unit::Kern { + value, synth_space, .. + } => Some((*value, *synth_space)), + Unit::Glyph { .. } => None, + }) + .collect() + }; + // Two TJ numbers between the same glyphs: their sum (0.25 em) reads as one space, carried by + // the larger one; a sum under 0.2 em reads as none. + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td [(Ab) -120 -130 (Cd)] TJ ET \ + BT /F1 12 Tf 72 600 Td [(Ef) -100 -90 (Gh)] TJ ET", + )); + assert_eq!(texts(&m), ["Ab Cd", "EfGh"], "kern sums"); + assert_eq!( + kerns(&run_with(&m, "Ab Cd").units), + [(-120.0, false), (-130.0, true)], + "the larger kern carries the space" + ); + // A Td that lands on the pen (A 667 + b 556 at 12 pt = 14.676) adds no gap kern. + let m = model0(helvetica_page( + b"BT /F1 12 Tf 72 700 Td (Ab) Tj 14.676 0 Td (Cd) Tj ET", + )); + let r = run_with(&m, "AbCd"); + assert_eq!(r.members.len(), 2, "two members, one run"); + assert_eq!(kerns(&r.units), [], "no zero-width gap kern"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/walk.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/walk.rs new file mode 100644 index 0000000..484cd79 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/walk.rs @@ -0,0 +1,286 @@ +//! WALK-01…14: glyph math (B1, B2, Tc/Tw per code, Tz, Ts), render modes, text clips, the `q` +//! stack, clip containment, patterns, soft masks and ActualText. + +use super::{close, ctx, model0, reason_of, run_with, walk}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::state::ClipState; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, cid_font, cid_hex, helvetica_doc, helvetica_page, ClipKind, DocBuilder, PageSpec, +}; +use crate::pdf_engine::text_edit::walker::{RecElem, WalkMode}; + +/// A page whose `/F2` is a CID font over "AB C" (2-byte codes, width 500). +fn cid_page(content: &str) -> Vec { + let mut d = DocBuilder::new(); + let f = cid_font(&mut d.b, "ABCDEF+Arimo", "AB C"); + d.page(PageSpec::new( + content.as_bytes(), + &format!("/Font << /F2 {f} 0 R >>"), + )); + d.build() +} + +#[test] +fn walk_01_b1_tj_kern_moves_the_tail() { + let c = ctx(fx::kerned_then_tail()); + let w = walk(&c, 0, WalkMode::Edit); + let (tj, tail) = (&w.records[0], &w.records[1]); + // Helvetica A=B=667, C=D=722 at 10 pt; −500 → +5 pt. + assert!( + close(tj.advance_ts, 13.34 + 5.0 + 14.44), + "WALK-01 advance_ts includes the kern: {}", + tj.advance_ts + ); + assert!( + close(tail.pen_before.0, 72.0 + 32.78) && close(tail.pen_before.1, 700.0), + "WALK-01 the tail starts after the kern: {:?}", + tail.pen_before + ); + assert!( + matches!(tj.elems.get(2), Some(RecElem::Kern { value, .. }) if *value == -500.0), + "WALK-01 the kern is an element: {:?}", + tj.elems + ); + assert!(close(tj.pen_after.0, tail.pen_before.0)); +} + +#[test] +fn walk_02_b2_effective_size_is_not_the_tf_operand() { + let m = model0(fx::tf1_tm12()); + let r = run_with(&m, "Hi"); + assert_eq!(r.tfs, 1.0, "WALK-02 Tf operand"); + assert!( + close(r.effective_size, 12.0), + "WALK-02 effective {}", + r.effective_size + ); + assert!(close(r.text_to_user_x, 12.0)); + assert_eq!(r.reason, None); + // Helvetica AFM: H = 722, i = 222 → 0.944 text units × 12. + assert!( + close(r.original_extent, 0.944 * 12.0), + "WALK-02 extent {}", + r.original_extent + ); +} + +#[test] +fn walk_03_tc_counts_per_code_on_two_byte_codes() { + let c = ctx(cid_page(&format!( + "BT /F2 10 Tf 2 Tc 72 700 Td <{}> Tj ET", + cid_hex("AB C", "AB") + ))); + let w = walk(&c, 0, WalkMode::Edit); + let r = &w.records[0]; + assert_eq!(r.glyphs.len(), 2, "WALK-03 two 2-byte codes"); + assert!( + close(r.advance_ts, 2.0 * (5.0 + 2.0)), + "WALK-03 Tc once per code, not per byte: {}", + r.advance_ts + ); +} + +#[test] +fn walk_04_tw_only_for_the_one_byte_space() { + let c = ctx(helvetica_page(b"BT /F1 10 Tf 5 Tw 72 700 Td (a b) Tj ET")); + let r = &walk(&c, 0, WalkMode::Edit).records[0]; + // a=556 space=278 b=556 at 10 pt + Tw 5 once. + assert!( + close(r.advance_ts, 13.9 + 5.0), + "WALK-04 simple font: {}", + r.advance_ts + ); + let c = ctx(cid_page(&format!( + "BT /F2 10 Tf 5 Tw 72 700 Td <{}> Tj ET", + cid_hex("AB C", "A B") + ))); + let r = &walk(&c, 0, WalkMode::Edit).records[0]; + assert!( + close(r.advance_ts, 15.0), + "WALK-04 a 2-byte code 0x0003 that reads as a space gets no Tw: {}", + r.advance_ts + ); +} + +#[test] +fn walk_05_tz_scales_the_pen_not_advance_ts() { + let c = ctx(helvetica_page(b"BT /F1 10 Tf 80 Tz 72 700 Td (Hi) Tj ET")); + let r = &walk(&c, 0, WalkMode::Edit).records[0]; + assert!( + close(r.advance_ts, 9.44), + "WALK-05 Th excluded: {}", + r.advance_ts + ); + assert!( + close(r.pen_after.0 - r.pen_before.0, 9.44 * 0.8), + "WALK-05 user advance × 0.8" + ); + assert!(close(r.text_to_user[0], 8.0), "WALK-05 Tfs·Th"); +} + +#[test] +fn walk_06_rise_moves_origins_and_boxes() { + let c = ctx(helvetica_page(b"BT /F1 10 Tf 3 Ts 72 700 Td (Hi) Tj ET")); + let r = &walk(&c, 0, WalkMode::Edit).records[0]; + let g = &r.glyphs[0]; + assert!(close(g.origin.1, 703.0), "WALK-06 origin {:?}", g.origin); + let descent = r.font.as_ref().expect("font").descent; + assert!( + close(g.bbox[1], 703.0 + descent * 10.0), + "WALK-06 box bottom {} (descent {descent})", + g.bbox[1] + ); + assert!(close(r.pen_before.1, 703.0)); +} + +#[test] +fn walk_07_render_modes() { + for (tr, want) in [ + (0, None), + (1, None), + (2, None), + (3, Some(R::InvisibleText)), + (4, Some(R::TextClipMode)), + (7, Some(R::TextClipMode)), + ] { + let page = + helvetica_page(format!("BT {tr} Tr /F1 12 Tf 72 700 Td (Mode) Tj ET").as_bytes()); + assert_eq!(reason_of(&model0(page), "Mode"), want, "WALK-07 Tr {tr}"); + } +} + +#[test] +fn walk_08_text_after_a_text_clip_is_clipped() { + let m = model0(fx::tr7()); + assert_eq!(reason_of(&m, "Clip text"), Some(R::TextClipMode), "WALK-08"); + assert_eq!(reason_of(&m, "After clip"), Some(R::Clipped), "WALK-08"); + let after = &m.walk.records[1]; + assert_eq!( + after.before.clip, + ClipState::Complex, + "WALK-08 clip after ET" + ); +} + +#[test] +fn walk_09_q_overflow_is_refused_and_a_stray_q_tolerated() { + let over = format!("{}BT /F1 12 Tf 72 700 Td (Deep) Tj ET", "q ".repeat(65)); + let m = model0(helvetica_page(over.as_bytes())); + assert_eq!(m.page_reason, Some(R::MalformedContent), "WALK-09 65 × q"); + assert!( + m.runs.is_empty() && m.walk.records.is_empty(), + "WALK-09 never a partial list" + ); + let fine = format!( + "{}BT /F1 12 Tf 72 700 Td (Deep) Tj ET{}", + "q ".repeat(64), + " Q".repeat(64) + ); + assert_eq!( + reason_of(&model0(helvetica_page(fine.as_bytes())), "Deep"), + None + ); + let stray = model0(helvetica_page(b"Q Q BT /F1 12 Tf 72 700 Td (Stray) Tj ET")); + assert_eq!( + reason_of(&stray, "Stray"), + None, + "WALK-09 Q on an empty stack is ignored" + ); +} + +#[test] +fn walk_10_clip_page_small_curve() { + assert_eq!( + reason_of(&model0(fx::clip(ClipKind::Page)), "Clip me"), + None, + "WALK-10 page" + ); + assert_eq!( + reason_of(&model0(fx::clip(ClipKind::Small)), "Clip me"), + Some(R::Clipped), + "WALK-10 small" + ); + assert_eq!( + reason_of(&model0(fx::clip(ClipKind::Curve)), "Clip me"), + Some(R::Clipped), + "WALK-10 curve" + ); + let m = model0(fx::clip(ClipKind::Page)); + assert_eq!( + m.walk.records[0].before.clip, + ClipState::Rect([0.0, 0.0, 612.0, 792.0]) + ); + let outside = model0(helvetica_page(b"BT /F1 12 Tf 600 700 Td (Edge) Tj ET")); + assert_eq!( + reason_of(&outside, "Edge"), + Some(R::Clipped), + "WALK-10 past the page" + ); +} + +#[test] +fn walk_11_pattern_fill_and_stroke() { + assert_eq!( + reason_of(&model0(fx::pattern_fill()), "Pattern"), + Some(R::Pattern), + "WALK-11 fill" + ); + let stroke = |tr: i64| { + let c = format!("BT /F1 12 Tf /Cs1 CS /P1 SCN {tr} Tr 72 720 Td (Stroke) Tj ET"); + reason_of(&model0(fx::pattern_doc(c.as_bytes())), "Stroke") + }; + assert_eq!( + stroke(1), + Some(R::Pattern), + "WALK-11 stroked text, pattern stroke" + ); + assert_eq!( + stroke(0), + None, + "WALK-11 filled text ignores the stroke paint" + ); +} + +#[test] +fn walk_12_soft_mask() { + assert_eq!( + reason_of(&model0(fx::smask_text()), "Masked"), + Some(R::SoftMask), + "WALK-12" + ); + let none = helvetica_doc( + b"/GS1 gs BT /F1 12 Tf 72 720 Td (No mask) Tj ET", + "/ExtGState << /GS1 << /SMask /None /ca 0.5 >> >>", + "", + ); + let m = model0(none); + assert_eq!(reason_of(&m, "No mask"), None, "WALK-12 /SMask /None"); + assert_eq!(m.walk.records[0].before.gs.ca, 0.5); +} + +#[test] +fn walk_13_actual_text_inline() { + let m = model0(fx::actual_text_span()); + assert_eq!( + reason_of(&m, "fi"), + Some(R::ActualText), + "WALK-13 inline /ActualText" + ); + assert_eq!(reason_of(&m, "Plain"), None); + let e = model0(helvetica_page( + b"/Span <> BDC BT /F1 12 Tf 72 720 Td (abbr) Tj ET EMC", + )); + assert_eq!(reason_of(&e, "abbr"), Some(R::ActualText), "WALK-13 /E"); +} + +#[test] +fn walk_14_actual_text_via_properties() { + let page = helvetica_doc( + b"/Span /MC0 BDC BT /F1 12 Tf 72 720 Td (Prop) Tj ET EMC /Span /MC1 BDC BT /F1 12 Tf 72 700 Td (Other) Tj ET EMC", + "/Properties << /MC0 << /ActualText (x) >> /MC1 << /Lang (en) >> >>", + "", + ); + let m = model0(page); + assert_eq!(reason_of(&m, "Prop"), Some(R::ActualText), "WALK-14"); + assert_eq!(reason_of(&m, "Other"), None, "WALK-14 other properties"); +} diff --git a/src-tauri/src/pdf_engine/text_edit/tests_walk/walk2.rs b/src-tauri/src/pdf_engine/text_edit/tests_walk/walk2.rs new file mode 100644 index 0000000..c76ab20 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/tests_walk/walk2.rs @@ -0,0 +1,554 @@ +//! WALK-15…27: structure ActualText, optional content, Forms (paint, descent, cycles, depth), +//! ExtGState fonts, paint records, non-finite numbers, determinism, inline images, split and +//! shared content, unresolvable references, the `Wrapped` mode over a real qpdf overlay, and the +//! id-free digest. + +use super::{close, content, ctx, model0, reason_of, run_with, walk}; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::engines::{run_tool, RunOpts}; +use crate::pdf_engine::text_edit::lexer::{lex_content, LexLimits, Operand, Operator}; +use crate::pdf_engine::text_edit::reasons::TextReason as R; +use crate::pdf_engine::text_edit::snapshot::read_snapshot; +use crate::pdf_engine::text_edit::state::{same_state, ColorEffect}; +use crate::pdf_engine::text_edit::testkit::producers::{ + self as fx, helvetica_doc, helvetica_page, DocBuilder, InlineProofKind, PageSpec, HELVETICA, +}; +use crate::pdf_engine::text_edit::testkit::{engines_or_skip, Scratch}; +use crate::pdf_engine::text_edit::walker::{walk_page, PageWalk, PaintKind, WalkMode}; +use std::ffi::OsString; + +#[test] +fn walk_15_actual_text_via_the_structure_tree() { + let m = model0(fx::actual_text_struct()); + assert_eq!( + reason_of(&m, "Struct"), + Some(R::ActualText), + "WALK-15 StructElem /ActualText" + ); + assert_eq!( + reason_of(&m, "Free"), + None, + "WALK-15 sibling element without it" + ); +} + +#[test] +fn walk_16_optional_content_visible_hidden_ocmd() { + let m = model0(fx::indd()); + assert_eq!(reason_of(&m, "visible layer"), None, "WALK-16 ON group"); + assert_eq!( + reason_of(&m, "hidden layer"), + Some(R::OptionalContent), + "WALK-16 OFF group" + ); + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let ocg = d.add("<< /Type /OCG /Name (Layer) >>"); + let ocmd = d.add(format!("<< /Type /OCMD /OCGs [{ocg} 0 R] /P /AnyOn >>")); + d.catalog_extra = format!("/OCProperties << /OCGs [{ocg} 0 R] /D << >> >>"); + d.page(PageSpec::new( + b"/OC /M1 BDC BT /F1 12 Tf 72 700 Td (Membership) Tj ET EMC /OC <> BDC BT /F1 12 Tf 72 680 Td (Inline) Tj ET EMC", + &format!("/Font << /F1 {f} 0 R >> /Properties << /M1 {ocmd} 0 R >>"), + )); + let m = model0(d.build()); + assert_eq!( + reason_of(&m, "Membership"), + Some(R::OptionalContent), + "WALK-16 OCMD" + ); + assert_eq!( + reason_of(&m, "Inline"), + Some(R::OptionalContent), + "WALK-16 inline dict" + ); +} + +#[test] +fn walk_17_form_text_is_a_paint_in_edit_and_nested_in_classify() { + let c = ctx(fx::nested_form()); + let edit = walk(&c, 0, WalkMode::Edit); + assert!(edit.records.is_empty(), "WALK-17 Edit does not descend"); + assert!( + edit.paints + .iter() + .any(|p| matches!(&p.kind, PaintKind::FormXObject { name, .. } if name == b"Fm0")), + "WALK-17 the Form is a paint" + ); + let classify = walk(&c, 0, WalkMode::Classify); + let r = classify + .records + .first() + .expect("WALK-17 record inside the Form"); + assert_eq!((r.depth, r.form_chain.len(), r.span.clone()), (1, 1, None)); + assert!( + close(r.pen_before.0, 82.0) && close(r.pen_before.1, 610.0), + "WALK-17 Matrix × CTM" + ); + let page = content(&c, 0); + let mut rc = crate::pdf_engine::text_edit::runs::reasons::ReasonCtx::new(&c, &page, &classify); + assert_eq!( + rc.record_reasons(r).first(), + Some(&R::NestedForm), + "WALK-17 NESTED_FORM" + ); +} + +/// A chain of `depth` Forms, the innermost drawing text (`cycle`: the last one paints itself). +fn form_chain(depth: usize, cycle: bool) -> Vec { + let mut d = DocBuilder::new(); + let f = d.add(HELVETICA); + let ids: Vec = (0..depth).map(|_| d.reserve()).collect(); + for (i, id) in ids.iter().enumerate() { + let (res, body) = match ids.get(i + 1) { + Some(next) => ( + format!("/XObject << /Fm {next} 0 R >>"), + "/Fm Do".to_string(), + ), + None if cycle => (format!("/XObject << /Fm {id} 0 R >>"), "/Fm Do".to_string()), + None => ( + format!("/Font << /F1 {f} 0 R >>"), + "BT /F1 12 Tf 10 10 Td (Deep) Tj ET".to_string(), + ), + }; + let stream = crate::pdf_engine::text_edit::testkit::pdf::PdfBuilder::stream_body( + &format!("/Type /XObject /Subtype /Form /BBox [0 0 500 500] /Resources << {res} >>"), + body.as_bytes(), + ); + d.b.set(*id, stream); + } + d.page(PageSpec::new( + b"q 1 0 0 1 72 500 cm /Fm Do Q BT /F1 12 Tf 72 700 Td (Page text) Tj ET", + &format!("/Font << /F1 {f} 0 R >> /XObject << /Fm {} 0 R >>", ids[0]), + )); + d.build() +} + +#[test] +fn walk_18_form_cycle_and_depth_nine_are_malformed() { + for (pdf, what) in [ + (form_chain(9, false), "depth 9"), + (form_chain(2, true), "cycle"), + ] { + let c = ctx(pdf); + let classify = walk(&c, 0, WalkMode::Classify); + assert_eq!( + classify.page_reason, + Some(R::MalformedContent), + "WALK-18 {what}" + ); + assert!( + classify.records.is_empty(), + "WALK-18 {what}: no partial list" + ); + let m = super::model(&c, 0); + assert_eq!( + reason_of(&m, "Page text"), + None, + "WALK-18 {what}: Edit only paints the Form" + ); + } + let c = ctx(form_chain(8, false)); + let classify = walk(&c, 0, WalkMode::Classify); + assert_eq!(classify.page_reason, None, "WALK-18 depth 8 is fine"); + assert_eq!(classify.records.iter().map(|r| r.depth).max(), Some(8)); +} + +#[test] +fn walk_19_extgstate_font() { + let m = model0(fx::extgstate_font()); + let r = run_with(&m, "Hi there"); + assert!(r.font_from_extgstate, "WALK-19 font from the ExtGState"); + assert_eq!( + r.reason, None, + "WALK-19 editable text (B9: size/face disabled later)" + ); + assert!(r.surface.is_empty(), "WALK-19 no resource name"); + let surface = m.surface(r); + assert_eq!( + surface.fonts.len(), + 1, + "WALK-19 the typing surface is that font alone" + ); + assert!(close(r.effective_size, 12.0)); + let use_ = m.walk.records[0] + .before + .text + .font + .clone() + .expect("font use"); + assert!(use_.from_extgstate && use_.resource.is_none() && use_.tf_op.is_none()); +} + +#[test] +fn walk_20_paint_records_carry_their_state() { + let c = ctx(helvetica_page( + b"1 0 0 rg 10 10 50 50 re f 0 0 1 RG 3 w 10 10 m 60 60 l S BT /F1 12 Tf 72 700 Td (After) Tj ET", + )); + let w = walk(&c, 0, WalkMode::Edit); + assert_eq!(w.paints.len(), 2, "WALK-20 two path paints"); + assert_eq!(w.paints[0].kind, PaintKind::Path(Operator::f)); + assert_eq!( + w.paints[0].state.fill.effect, + ColorEffect::Rgb([1.0, 0.0, 0.0]), + "WALK-20 fill" + ); + assert_eq!( + w.paints[1].state.stroke.effect, + ColorEffect::Rgb([0.0, 0.0, 1.0]), + "WALK-20 stroke" + ); + assert_eq!(w.paints[1].state.gs.line_width, 3.0); + assert_eq!(w.paints[0].bbox, Some([10.0, 10.0, 60.0, 60.0])); + assert!(w.paints[0].seq < w.paints[1].seq && w.paints[1].seq < w.records[0].seq); + assert_eq!(w.records[0].before.fill.hex().as_deref(), Some("#ff0000")); +} + +#[test] +fn walk_21_non_finite_numbers_refuse_the_page() { + let blowup = format!( + "{}BT /F1 12 Tf 72 700 Td (Huge) Tj ET", + "1000000000 0 0 1000000000 0 0 cm ".repeat(40) + ); + let m = model0(helvetica_page(blowup.as_bytes())); + assert_eq!( + m.page_reason, + Some(R::MalformedContent), + "WALK-21 {:?}", + m.page_detail + ); + assert!(m.page_detail.unwrap_or_default().contains("non-finite")); +} + +/// The comparable, id-free shape of a walk. +fn summary(w: &PageWalk) -> Vec { + let mut out: Vec = w + .records + .iter() + .map(|r| { + format!( + "{} {:?} {:?} {:?} {:?} {:?} {:?} {:?}", + r.seq, + r.op, + r.span, + r.glyphs + .iter() + .map(|g| (g.code, g.text.clone(), g.origin)) + .collect::>(), + r.pen_after, + r.advance_ts, + r.before, + r.after + ) + }) + .collect(); + out.extend( + w.paints + .iter() + .map(|p| format!("{} {:?} {:?} {:?}", p.seq, p.kind, p.bbox, p.state)), + ); + out +} + +#[test] +fn walk_22_walks_are_deterministic() { + for pdf in [fx::word(), fx::word_tr(), fx::indd(), fx::skia()] { + let (a, b) = (ctx(pdf.clone()), ctx(pdf)); + for mode in [WalkMode::Edit, WalkMode::Classify] { + assert_eq!( + summary(&walk(&a, 0, mode.clone())), + summary(&walk(&b, 0, mode.clone())), + "WALK-22 two walks agree" + ); + assert_eq!( + summary(&walk(&a, 0, mode.clone())), + summary(&walk(&a, 0, mode)) + ); + } + let (ma, mb) = (super::model(&a, 0), super::model(&b, 0)); + let ids = |m: &crate::pdf_engine::text_edit::runs::PageModel| { + m.runs + .iter() + .map(|r| (r.id.clone(), r.text.clone(), r.order)) + .collect::>() + }; + assert_eq!(ids(&ma), ids(&mb), "WALK-22 runs"); + } +} + +#[test] +fn walk_23_inline_image_proofs() { + for proof in [ + InlineProofKind::Length, + InlineProofKind::Unfiltered, + InlineProofKind::Flate, + ] { + let m = model0(fx::inline_image(proof)); + assert_eq!( + reason_of(&m, "After the image"), + None, + "WALK-23 proven: {proof:?}" + ); + assert!(m + .walk + .paints + .iter() + .any(|p| matches!(p.kind, PaintKind::InlineImage { .. }))); + } + let m = model0(fx::inline_image(InlineProofKind::Dct)); + assert_eq!( + reason_of(&m, "After the image"), + Some(R::InlineImage), + "WALK-23 an unproven end refuses what follows" + ); +} + +#[test] +fn walk_24_split_and_shared_content() { + let m = model0(fx::straddling_op()); + assert_eq!( + reason_of(&m, "Hello"), + Some(R::SplitContent), + "WALK-24 straddles the parts" + ); + assert_eq!(reason_of(&m, "After"), None); + let c = ctx(fx::shared()); + for page in 0..4 { + let m = super::model(&c, page); + let want = match page { + 0 | 1 => vec![("Shared body", Some(R::SharedContent))], + _ => vec![ + ("Letterhead", Some(R::SharedContent)), + (if page == 2 { "Body three" } else { "Body four" }, None), + ], + }; + for (text, reason) in want { + assert_eq!(reason_of(&m, text), reason, "WALK-24 page {page} {text}"); + } + } +} + +#[test] +fn walk_25_unresolvable_references_refuse_the_page() { + let cases: [(&[u8], &str, &str); 5] = [ + ( + b"/GSX gs BT /F1 12 Tf 72 700 Td (A) Tj ET", + "", + "ExtGState name", + ), + ( + b"/GS1 gs BT /F1 12 Tf 72 700 Td (A) Tj ET", + "/ExtGState << /GS1 999 0 R >>", + "dangling ExtGState", + ), + ( + b"q /Xn Do Q BT /F1 12 Tf 72 700 Td (A) Tj ET", + "/XObject << /Xn 998 0 R >>", + "dangling XObject", + ), + ( + b"/Span /MCX BDC BT /F1 12 Tf 72 700 Td (A) Tj ET EMC", + "", + "Properties name", + ), + ( + b"BT /F9 12 Tf 72 700 Td (A) Tj ET", + "/Font << /F9 997 0 R >>", + "dangling font", + ), + ]; + for (content, extra, what) in cases { + let m = model0(helvetica_doc(content, extra, "")); + assert_eq!(m.page_reason, Some(R::MalformedContent), "WALK-25 {what}"); + } + let m = model0(helvetica_page(b"BT /F7 12 Tf 72 700 Td (Nameless) Tj ET")); + assert_eq!( + m.page_reason, None, + "WALK-25 an absent font name stays a run reason" + ); + assert_eq!( + reason_of(&m, "\u{fffd}".repeat(8).as_str()), + Some(R::MissingFont) + ); +} + +/// Records of `a` and `b` agree id-free (op, font resource + hash, codes, text, origins, state). +fn assert_same_records(a: &PageWalk, b: &PageWalk, id: &str) { + assert_eq!(a.records.len(), b.records.len(), "{id} record count"); + for (x, y) in a.records.iter().zip(&b.records) { + assert_eq!(x.op, y.op, "{id}"); + let font = |r: &crate::pdf_engine::text_edit::walker::ShowRecord| { + r.before + .text + .font + .as_ref() + .map(|f| (f.resource.clone(), f.content_hash)) + }; + assert_eq!(font(x), font(y), "{id} font"); + let codes = |r: &crate::pdf_engine::text_edit::walker::ShowRecord| { + r.glyphs + .iter() + .map(|g| (g.code, g.text.clone())) + .collect::>() + }; + assert_eq!(codes(x), codes(y), "{id} codes"); + for (g, h) in x.glyphs.iter().zip(&y.glyphs) { + assert!( + (g.origin.0 - h.origin.0).abs() < 0.01 && (g.origin.1 - h.origin.1).abs() < 0.01, + "{id} origin {:?} vs {:?}", + g.origin, + h.origin + ); + } + assert_eq!( + same_state(&x.before, &y.before), + Ok(()), + "{id} state before" + ); + assert_eq!(same_state(&x.after, &y.after), Ok(()), "{id} state after"); + } +} + +#[test] +fn walk_26_wrapped_mode_follows_the_qpdf_overlay_wrapper() { + let Some(engines) = engines_or_skip("walk_26_wrapped_mode_follows_the_qpdf_overlay_wrapper") + else { + return; + }; + let s = Scratch::new("walk26"); + for (rotate, pdf) in [ + (0, fx::word()), + (90, fx::rotated(90, true)), + (270, fx::rotated(270, true)), + ] { + let src = s.write(&format!("src{rotate}.pdf"), &pdf); + let blank = s.write( + &format!("blank{rotate}.pdf"), + &helvetica_doc(b"", "", &format!("/Rotate {rotate}")), + ); + let out = s.path(&format!("out{rotate}.pdf")); + let args: Vec = vec![ + src.into(), + "--overlay".into(), + blank.into(), + "--".into(), + out.clone().into(), + ]; + let r = run_tool(&engines.qpdf, &args, false, &RunOpts::default()).expect("qpdf --overlay"); + assert!(r.code == 0 || r.code == 3, "WALK-26 qpdf: {}", r.stderr); + let src_ctx = ctx(pdf); + let out_ctx = SnapshotContext::new(read_snapshot(&out).expect("overlay output opens")); + let before = walk(&src_ctx, 0, WalkMode::Edit); + let out_content = content(&out_ctx, 0); + let ops = + lex_content(&out_content.joined, &LexLimits::page(), None).expect("wrapper lexes"); + let name = ops + .iter() + .find(|o| o.operator == Operator::Do) + .and_then(|o| o.operands.first().and_then(Operand::as_name)) + .expect("WALK-26 the page paints a wrapper Form") + .to_vec(); + let edit_out = walk_page(&out_ctx, 0, &out_content, WalkMode::Edit, None); + assert!( + edit_out.records.is_empty(), + "WALK-26 Edit sees only the wrapper paint" + ); + let wrapped = walk_page(&out_ctx, 0, &out_content, WalkMode::Wrapped { name }, None); + assert_eq!( + wrapped.page_reason, None, + "WALK-26 /Rotate {rotate}: {:?}", + wrapped.page_detail + ); + assert!(!before.records.is_empty()); + assert_same_records(&before, &wrapped, &format!("WALK-26 /Rotate {rotate}")); + let bad = walk_page( + &out_ctx, + 0, + &out_content, + WalkMode::Wrapped { + name: b"Nope".to_vec(), + }, + None, + ); + assert_eq!( + bad.page_reason, + Some(R::MalformedContent), + "WALK-26 a name that is not painted" + ); + } + let not_wrapper = ctx(fx::word()); + let c = content(¬_wrapper, 0); + let w = walk_page( + ¬_wrapper, + 0, + &c, + WalkMode::Wrapped { + name: b"Fx0".to_vec(), + }, + None, + ); + assert_eq!( + w.page_reason, + Some(R::MalformedContent), + "WALK-26 shape violation" + ); + assert!(w.page_detail.unwrap_or_default().starts_with("wrapper")); +} + +#[test] +fn walk_27_digest_captures_line_state_and_paints_are_id_free() { + let m = model0(helvetica_page( + b"[3 2] 1 d 2 J 1 j 5 M 0 0 1 RG 0.5 G BT /F1 12 Tf 72 700 Td (Lines) Tj ET", + )); + let gs = &m.walk.records[0].before.gs; + assert_eq!( + ( + (gs.dash.0.to_vec(), gs.dash.1), + gs.line_cap, + gs.line_join, + gs.miter_limit + ), + ((vec![3.0, 2.0], 1.0), 2, 1, 5.0), + "WALK-27" + ); + assert_eq!( + m.walk.records[0].before.stroke.effect, + ColorEffect::Rgb([0.5, 0.5, 0.5]) + ); + // The same page with renumbered objects: paint kinds compare equal (no ids inside). + let renumbered = { + let mut d = DocBuilder::new(); + for _ in 0..5 { + d.add("<< /Padding true >>"); + } + let f = d.add(HELVETICA); + let img = d.b.add_stream( + "/Type /XObject /Subtype /Image /Width 2 /Height 2 /ColorSpace /DeviceRGB /BitsPerComponent 8", + &[200, 16, 16, 16, 200, 16, 16, 16, 200, 200, 200, 16], + ); + d.page(PageSpec::new( + b"q 40 0 0 40 72 400 cm /Im0 Do Q BT /F1 12 Tf 72 720 Td (Image page) Tj ET", + &format!("/Font << /F1 {f} 0 R >> /XObject << /Im0 {img} 0 R >>"), + )); + d.build() + }; + let kinds = |pdf: Vec| -> Vec { + let c = ctx(pdf); + walk(&c, 0, WalkMode::Edit) + .paints + .into_iter() + .map(|p| p.kind) + .collect() + }; + let original = kinds(fx::swapped_image(false)); + assert_eq!( + original, + kinds(renumbered), + "WALK-27 renumbering keeps paint kinds" + ); + let swapped = kinds(fx::swapped_image(true)); + assert_ne!( + original, swapped, + "WALK-27 a swapped image behind /Im0 differs" + ); + assert!(matches!(&swapped[0], PaintKind::ImageXObject { name, .. } if name == b"Im0")); +} diff --git a/src-tauri/src/pdf_engine/text_edit/verify.rs b/src-tauri/src/pdf_engine/text_edit/verify.rs new file mode 100644 index 0000000..c8c8e24 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/verify.rs @@ -0,0 +1,708 @@ +//! The verification re-walk (SPEC §B.14): one function, `walk_and_verify_with`, shared by the plan +//! self-check, the preview and Save (Phase A check A4). It compares the walk of the page before +//! the edit with the walk of the content after it, id-free (D32), so it holds across qpdf's +//! renumbering: +//! +//! 1. the after page is not refused and the replacement grammar holds inside every splice; +//! 2. every record and paint outside the splices pairs by span (`SpanMap`) with one of the same +//! operator, and the records inside the splices are exactly the emitted ones; +//! 3. unedited records keep op, codes, texts, font, full state and every glyph origin (0.01 pt); +//! 4. paints keep their id-free kind (name + deep hash) and full state (B15); +//! 5. edited runs draw exactly the expected glyphs (`verify/edited.rs`); +//! 6. the state in force after each replaced region — read by a probe `[]TJ` inserted right +//! after it into a copy of the content, walked in the same context — equals the original state +//! after that member, with the same text matrix and pen (B14). +//! +//! Spans are mapped through prefix sums and regions found by binary search, so a page of 250,000 +//! ops with hundreds of edits is checked in O(n log k). + +mod edited; + +use crate::pdf_engine::text_edit::content::PageContent; +use crate::pdf_engine::text_edit::context::SnapshotContext; +use crate::pdf_engine::text_edit::encode::REPLACEMENT_OPERATORS; +use crate::pdf_engine::text_edit::lexer::Span; +use crate::pdf_engine::text_edit::limits::{DRIFT_TOLERANCE_PT, STATE_EPSILON, SYNTH_SPACE_EM}; +use crate::pdf_engine::text_edit::reasons::{EditProblemCode, TextReason}; +use crate::pdf_engine::text_edit::rewrite::{PagePlan, Splice}; +use crate::pdf_engine::text_edit::runs::TextRun; +use crate::pdf_engine::text_edit::state::same_state; +use crate::pdf_engine::text_edit::walker::{walk_page, PageWalk, ShowRecord, WalkMode}; +use std::collections::HashMap; +use std::sync::atomic::AtomicBool; + +/// The probe inserted after each replaced region of a copy of the after content. +const PROBE: &[u8] = b"\n[]TJ\n"; + +#[derive(Debug, Clone, PartialEq)] +pub enum VerifyFailure { + PageRefused(TextReason), + RecordUnpaired { + index: usize, + }, + RecordCount { + expected: usize, + found: usize, + }, + PaintChanged { + index: usize, + field: &'static str, + }, + Drift { + record: usize, + glyph: usize, + pt: f64, + }, + StateChanged { + record: usize, + field: &'static str, + }, + /// `what`: text|codes|font|size|tc|fill|origin|pen|glyph|requested_text|absorbed|post_state| + /// records + EditedMismatch { + run_id: String, + what: &'static str, + }, + ForbiddenOperator { + op: &'static str, + }, +} + +impl VerifyFailure { + /// Drift → PEN_DRIFT; StateChanged/PaintChanged → STATE_CHANGED; the rest → EDIT_VERIFY_FAILED. + pub fn problem_code(&self) -> EditProblemCode { + match self { + VerifyFailure::Drift { .. } => EditProblemCode::PenDrift, + VerifyFailure::StateChanged { .. } | VerifyFailure::PaintChanged { .. } => { + EditProblemCode::StateChanged + } + _ => EditProblemCode::EditVerifyFailed, + } + } +} + +impl std::fmt::Display for VerifyFailure { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + match self { + VerifyFailure::PageRefused(r) => write!(f, "page_refused reason={}", r.as_str()), + VerifyFailure::RecordUnpaired { index } => write!(f, "record_unpaired record={index}"), + VerifyFailure::RecordCount { expected, found } => { + write!(f, "record_count expected={expected} found={found}") + } + VerifyFailure::PaintChanged { index, field } => { + write!(f, "paint_changed paint={index} field={field}") + } + VerifyFailure::Drift { record, glyph, pt } => { + write!(f, "drift record={record} glyph={glyph} pt={pt:.4}") + } + VerifyFailure::StateChanged { record, field } => { + write!(f, "state_changed record={record} field={field}") + } + VerifyFailure::EditedMismatch { run_id, what } => { + write!(f, "edited_mismatch what={what} run={run_id}") + } + VerifyFailure::ForbiddenOperator { op } => write!(f, "forbidden_operator op={op}"), + } + } +} + +/// One element of a decoded show sequence. +pub(crate) enum TextItem<'a> { + Glyph(&'a str), + /// A kern in em (`−n/1000`). + Kern(f64), +} + +fn is_space(t: &str) -> bool { + matches!(t, " " | "\u{a0}") +} + +/// The text a reader of the records gets: glyph texts, plus one space where the kerns between two +/// non-space glyphs add up to at least `SYNTH_SPACE_EM` (§A.3.4). +pub(crate) fn decoded_text<'a>(items: impl Iterator>) -> String { + let mut out = String::new(); + let mut prev_glyph: Option<&str> = None; + let mut kerns = 0.0f64; + let mut any_kern = false; + for item in items { + match item { + TextItem::Kern(em) => { + kerns += em; + any_kern = true; + } + TextItem::Glyph(t) => { + if let Some(p) = prev_glyph { + if any_kern && kerns >= SYNTH_SPACE_EM && !is_space(p) && !is_space(t) { + out.push(' '); + } + } + out.push_str(t); + prev_glyph = Some(t); + kerns = 0.0; + any_kern = false; + } + } + } + out +} + +/// Byte-length change of a splice. +fn growth(s: &Splice) -> i64 { + let old = s.joined.end.saturating_sub(s.joined.start); + i64::try_from(s.bytes.len()) + .unwrap_or(i64::MAX) + .saturating_sub(i64::try_from(old).unwrap_or(i64::MAX)) +} + +fn shifted(v: usize, shift: i64) -> usize { + usize::try_from(i64::try_from(v).unwrap_or(i64::MAX).saturating_add(shift)).unwrap_or(0) +} + +/// `original` (a span of the content before the edit) in the content after the edit: shifted by +/// the net length change of the splices that end at or before it — prefix sums over the splices +/// (sorted by joined start, disjoint); `map_span` (test-only) is the one-span definition. +struct SpanMap { + /// (joined end of splice k, total shift of splices 0..=k). + ends: Vec<(usize, i64)>, +} + +impl SpanMap { + fn new(splices: &[Splice]) -> SpanMap { + let mut total = 0i64; + let ends = splices + .iter() + .map(|s| { + total = total.saturating_add(growth(s)); + (s.joined.end, total) + }) + .collect(); + SpanMap { ends } + } + + fn map(&self, original: &Span) -> Span { + let k = self.ends.partition_point(|(end, _)| *end <= original.start); + let shift = k + .checked_sub(1) + .and_then(|i| self.ends.get(i)) + .map_or(0, |(_, s)| *s); + shifted(original.start, shift)..shifted(original.end, shift) + } +} + +/// Each splice's region in the after content (`[start, end)`), in `plan.splices` order (sorted by +/// joined start, so the regions are sorted and disjoint too). +fn after_regions(splices: &[Splice]) -> Vec { + let mut shift: i64 = 0; + splices + .iter() + .map(|s| { + let start = shifted(s.joined.start, shift); + shift = shift.saturating_add(growth(s)); + start..start.saturating_add(s.bytes.len()) + }) + .collect() +} + +fn inside(span: &Span, region: &Span) -> bool { + span.start >= region.start && span.end <= region.end +} + +fn overlaps(span: &Span, region: &Span) -> bool { + span.start < region.end && region.start < span.end +} + +/// The first of the sorted, disjoint `regions` that could overlap `span` (binary search). +fn region_near(regions: &[Span], span: &Span) -> Option { + let k = regions.partition_point(|r| r.end <= span.start); + regions + .get(k) + .is_some_and(|r| overlaps(span, r)) + .then_some(k) +} + +pub(crate) fn dist(a: (f64, f64), b: (f64, f64)) -> f64 { + (a.0 - b.0).hypot(a.1 - b.1) +} + +/// The font a record is drawn with, id-free: (resource name, content hash). +pub(crate) fn font_id(rec: &ShowRecord) -> (Vec, u64) { + match rec.before.text.font.as_ref() { + Some(f) => ( + f.resource + .as_deref() + .map(<[u8]>::to_vec) + .unwrap_or_default(), + f.content_hash, + ), + None => (Vec::new(), 0), + } +} + +/// The member splices (those of the plan's runs) as `(after insertion point, run, member)`. +fn member_ends(plan: &PagePlan, regions: &[Span]) -> Vec<(usize, usize, usize)> { + let index: HashMap = plan + .splices + .iter() + .enumerate() + .map(|(k, s)| (s.joined.start, k)) + .collect(); + let mut out = Vec::new(); + for (r, run) in plan.runs.iter().enumerate() { + for (m, s) in run.splices.iter().enumerate() { + if let Some(region) = index.get(&s.joined.start).and_then(|k| regions.get(*k)) { + out.push((region.end, r, m)); + } + } + } + out.sort_unstable(); + out +} + +/// A copy of `after` with a probe `[]TJ` after every member splice, and each probe's op start +/// (joined coordinates of the copy) with its (run, member). +fn probe_content( + after: &PageContent, + plan: &PagePlan, +) -> (PageContent, Vec<(usize, usize, usize)>) { + let regions = after_regions(&plan.splices); + let ends = member_ends(plan, ®ions); + let mut per_part: HashMap> = HashMap::new(); + for (pos, _, _) in &ends { + let part = after + .parts + .iter() + .position(|p| *pos >= p.start && p.start.checked_add(p.len).is_some_and(|e| *pos <= e)); + if let Some(part) = part { + let local = pos.saturating_sub(after.parts.get(part).map_or(0, |p| p.start)); + per_part.entry(part).or_default().push(local); + } + } + let mut replaced = Vec::new(); + for (part, mut locals) in per_part { + locals.sort_unstable(); + let mut bytes = after.part_bytes(part).to_vec(); + for local in locals.iter().rev() { + if *local <= bytes.len() { + bytes.splice(*local..*local, PROBE.iter().copied()); + } + } + replaced.push((part, bytes)); + } + let probes = ends + .iter() + .enumerate() + .map(|(i, (pos, r, m))| { + let before = i.saturating_mul(PROBE.len()).saturating_add(1); + (pos.saturating_add(before), *r, *m) + }) + .collect(); + (after.with_replaced_parts(&replaced), probes) +} + +/// Walks the after content (and its probe copy) in `ctx` and runs checks 1–6, keeping only what +/// `keep` takes from the after walk. Every caller — the plan self-check (in the source context), +/// the preview and Save (in the context of qpdf's output) — goes through here. The after walk is handed to `keep` and dropped once checks 1–5 pass, before +/// the probe walk, so two walks never coexist (review-T4 M-1). +pub fn walk_and_verify_with( + ctx: &SnapshotContext, + page_index: u32, + (after_content, before, before_runs): (&PageContent, &PageWalk, &[TextRun]), + plan: &PagePlan, + cancel: Option<&AtomicBool>, + keep: impl FnOnce(PageWalk) -> T, +) -> Result { + let after = walk_page(ctx, page_index, after_content, WalkMode::Edit, cancel); + check_page(before, before_runs, &after, plan)?; + let kept = keep(after); + let (probe_copy, probes) = probe_content(after_content, plan); + let probe = walk_page(ctx, page_index, &probe_copy, WalkMode::Edit, cancel); + drop(probe_copy); + check_post_state(before, &probe, &probes, plan)?; + Ok(kept) +} + +/// Checks 1–5. +fn check_page( + before: &PageWalk, + before_runs: &[TextRun], + after: &PageWalk, + plan: &PagePlan, +) -> Result<(), VerifyFailure> { + if let Some(r) = after.page_reason { + return Err(VerifyFailure::PageRefused(r)); + } + let regions = after_regions(&plan.splices); + if !seams::grammar_skipped() { + check_grammar(after, ®ions)?; + } + let map = SpanMap::new(&plan.splices); + let edited = pair_records(before, after, plan, ®ions, &map)?; + // A splice that belongs to no run may hold no show op. + for (k, s) in plan.splices.iter().enumerate() { + let member = plan.runs.iter().any(|r| r.splices.contains(s)); + if let (false, Some(&j)) = (member, edited.get(k).and_then(|e| e.first())) { + return Err(VerifyFailure::RecordUnpaired { index: j }); + } + } + check_paints(before, after, ®ions, &map)?; + for run in &plan.runs { + edited::check_edited_run(before, before_runs, after, plan, &edited, run)?; + } + Ok(()) +} + +/// Check 1: inside each splice only replacement operators, never an op crossing its boundary. +fn check_grammar(after: &PageWalk, regions: &[Span]) -> Result<(), VerifyFailure> { + for op in &after.ops { + let Some(k) = region_near(regions, &op.span) else { + continue; + }; + let inside_one = regions.get(k).is_some_and(|r| inside(&op.span, r)); + if !inside_one { + return Err(VerifyFailure::ForbiddenOperator { + op: "splice boundary", + }); + } + if !REPLACEMENT_OPERATORS.contains(&op.operator) { + return Err(VerifyFailure::ForbiddenOperator { + op: op.operator.as_str(), + }); + } + } + Ok(()) +} + +/// Check 2 and 3: pairs every unedited record and compares it; returns, per splice index, the +/// after records inside its region (in order). +fn pair_records( + before: &PageWalk, + after: &PageWalk, + plan: &PagePlan, + regions: &[Span], + map: &SpanMap, +) -> Result>, VerifyFailure> { + let mut by_span: HashMap<(usize, usize), usize> = HashMap::new(); + let mut edited: Vec> = vec![Vec::new(); regions.len()]; + for (j, rec) in after.records.iter().enumerate() { + let Some(span) = rec.span.as_ref() else { + return Err(VerifyFailure::RecordUnpaired { index: j }); + }; + match region_near(regions, span) + .filter(|k| regions.get(*k).is_some_and(|r| inside(span, r))) + { + Some(k) => { + if let Some(list) = edited.get_mut(k) { + list.push(j); + } + } + None => { + by_span.insert((span.start, span.end), j); + } + } + } + let originals: Vec = plan.splices.iter().map(|s| s.joined.clone()).collect(); + let mut paired: Vec = vec![false; after.records.len()]; + for (i, rec) in before.records.iter().enumerate() { + let Some(span) = rec.span.as_ref() else { + return Err(VerifyFailure::RecordUnpaired { index: i }); + }; + let replaced = region_near(&originals, span) + .is_some_and(|k| originals.get(k).is_some_and(|r| inside(span, r))); + if replaced { + continue; + } + let mapped = map.map(span); + let Some((j, a)) = by_span + .get(&(mapped.start, mapped.end)) + .and_then(|j| Some((*j, after.records.get(*j)?))) + else { + return Err(VerifyFailure::RecordUnpaired { index: i }); + }; + compare_unedited(rec, a, j)?; + if let Some(slot) = paired.get_mut(j) { + *slot = true; + } + } + let stray = by_span + .values() + .copied() + .filter(|j| !paired.get(*j).copied().unwrap_or(false)) + .min(); + if let Some(j) = stray { + return Err(VerifyFailure::RecordUnpaired { index: j }); + } + let emitted: usize = edited.iter().map(Vec::len).sum(); + let count = paired.iter().filter(|p| **p).count(); + if count.saturating_add(emitted) != after.records.len() { + return Err(VerifyFailure::RecordCount { + expected: count.saturating_add(emitted), + found: after.records.len(), + }); + } + Ok(edited) +} + +/// A drift of more than `DRIFT_TOLERANCE_PT` between two positions. +fn drift(a: (f64, f64), b: (f64, f64), record: usize, glyph: usize) -> Result<(), VerifyFailure> { + let d = dist(a, b); + if d > DRIFT_TOLERANCE_PT || !d.is_finite() { + return Err(VerifyFailure::Drift { + record, + glyph, + pt: d, + }); + } + Ok(()) +} + +fn compare_unedited(b: &ShowRecord, a: &ShowRecord, j: usize) -> Result<(), VerifyFailure> { + let same_glyphs = b.glyphs.len() == a.glyphs.len() + && b.glyphs + .iter() + .zip(&a.glyphs) + .all(|(x, y)| x.code == y.code && x.text == y.text); + if b.op != a.op || !same_glyphs { + return Err(VerifyFailure::RecordUnpaired { index: j }); + } + if font_id(b) != font_id(a) { + return Err(VerifyFailure::StateChanged { + record: j, + field: "font", + }); + } + same_state(&b.before, &a.before) + .and_then(|()| same_state(&b.after, &a.after)) + .map_err(|field| VerifyFailure::StateChanged { record: j, field })?; + for (g, (x, y)) in b.glyphs.iter().zip(&a.glyphs).enumerate() { + drift(x.origin, y.origin, j, g)?; + } + drift(b.pen_before, a.pen_before, j, a.glyphs.len())?; + drift( + b.pen_after, + a.pen_after, + j, + a.glyphs.len().saturating_add(1), + ) +} + +fn unpaired_paint(index: usize) -> VerifyFailure { + VerifyFailure::PaintChanged { + index, + field: "unpaired", + } +} + +/// Check 2 (paints) and 4: same count, paired by span, same id-free kind and full state. +fn check_paints( + before: &PageWalk, + after: &PageWalk, + regions: &[Span], + map: &SpanMap, +) -> Result<(), VerifyFailure> { + if before.paints.len() != after.paints.len() { + return Err(VerifyFailure::RecordCount { + expected: before.paints.len(), + found: after.paints.len(), + }); + } + let mut by_span: HashMap<(usize, usize), usize> = HashMap::new(); + for (j, p) in after.paints.iter().enumerate() { + match p.span.as_ref() { + Some(s) if region_near(regions, s).is_none() => { + by_span.insert((s.start, s.end), j); + } + _ => return Err(unpaired_paint(j)), + } + } + for (i, p) in before.paints.iter().enumerate() { + let span = p.span.as_ref().ok_or_else(|| unpaired_paint(i))?; + let mapped = map.map(span); + let a = by_span + .get(&(mapped.start, mapped.end)) + .and_then(|j| after.paints.get(*j)) + .ok_or_else(|| unpaired_paint(i))?; + if p.kind != a.kind { + return Err(VerifyFailure::PaintChanged { + index: i, + field: "kind", + }); + } + same_state(&p.state, &a.state) + .map_err(|field| VerifyFailure::PaintChanged { index: i, field })?; + } + Ok(()) +} + +/// Check 6: the state after each replaced region (probe record) equals the original state after +/// that member, with the same text matrix (linear part) and pen. +fn check_post_state( + before: &PageWalk, + probe: &PageWalk, + probes: &[(usize, usize, usize)], + plan: &PagePlan, +) -> Result<(), VerifyFailure> { + if let Some(r) = probe.page_reason { + return Err(VerifyFailure::PageRefused(r)); + } + let by_start: HashMap = probe + .records + .iter() + .enumerate() + .filter_map(|(i, x)| Some((x.span.as_ref()?.start, i))) + .collect(); + let before_by_span: HashMap<(usize, usize), &ShowRecord> = before + .records + .iter() + .filter_map(|x| Some(((x.span.as_ref()?.start, x.span.as_ref()?.end), x))) + .collect(); + for (start, r, m) in probes { + let Some(run) = plan.runs.get(*r) else { + continue; + }; + let fail = || VerifyFailure::EditedMismatch { + run_id: run.expected.run_id.clone(), + what: "post_state", + }; + let (index, rec) = by_start + .get(start) + .and_then(|i| Some((*i, probe.records.get(*i)?))) + .ok_or_else(fail)?; + let original = run + .expected + .member_spans + .get(*m) + .and_then(|s| before_by_span.get(&(s.start, s.end))) + .ok_or_else(fail)?; + same_state(&original.after, &rec.before).map_err(|field| VerifyFailure::StateChanged { + record: index, + field, + })?; + let linear_same = original + .tm_after + .iter() + .zip(&rec.tm_before) + .take(4) + .all(|(a, b)| (a - b).abs() <= STATE_EPSILON * 1f64.max(a.abs()).max(b.abs())); + if !linear_same { + return Err(VerifyFailure::StateChanged { + record: index, + field: "tm", + }); + } + drift(original.pen_after, rec.pen_before, index, 0)?; + } + Ok(()) +} + +/// Test seam: skip the replacement-grammar re-check (GATE-07/08 "grammar bypassed"). Production +/// builds have no way to skip it. +mod seams { + pub(super) fn grammar_skipped() -> bool { + #[cfg(test)] + { + if super::test_seams::grammar_skipped() { + return true; + } + } + false + } +} + +#[cfg(test)] +pub(crate) mod test_seams { + use std::cell::Cell; + + thread_local! { + static SKIP_GRAMMAR: Cell = const { Cell::new(false) }; + } + + pub(crate) fn grammar_skipped() -> bool { + SKIP_GRAMMAR.with(Cell::get) + } + + /// Skips the grammar re-check on this thread until the guard is dropped. + pub(crate) fn skip_grammar() -> GrammarGuard { + SKIP_GRAMMAR.with(|c| c.set(true)); + GrammarGuard + } + + pub(crate) struct GrammarGuard; + + impl Drop for GrammarGuard { + fn drop(&mut self) { + SKIP_GRAMMAR.with(|c| c.set(false)); + } + } +} + +/// `walk_and_verify_with` keeping the whole after walk (tests). +#[cfg(test)] +pub fn walk_and_verify( + ctx: &SnapshotContext, + page_index: u32, + after_content: &PageContent, + before: &PageWalk, + before_runs: &[TextRun], + plan: &PagePlan, + cancel: Option<&AtomicBool>, +) -> Result { + let walks = (after_content, before, before_runs); + walk_and_verify_with(ctx, page_index, walks, plan, cancel, |walk| walk) +} + +/// The definition `SpanMap` implements for many spans: `original` shifted by the net length +/// change of the splices that end at or before it. +#[cfg(test)] +pub fn map_span(splices: &[Splice], original: &Span) -> Span { + let shift: i64 = splices + .iter() + .filter(|s| s.joined.end <= original.start) + .map(growth) + .sum(); + shifted(original.start, shift)..shifted(original.end, shift) +} + +#[cfg(test)] +mod tests { + use super::{map_span, SpanMap}; + use crate::pdf_engine::text_edit::rewrite::Splice; + + /// review-T5 H2: production pairs spans with `SpanMap`; it must equal `map_span` (which + /// VER-04 pins) on any sorted, disjoint splices and any span, including spans touching a + /// splice end. + #[test] + fn span_map_equals_map_span_on_random_splices() { + let mut seed = 0x9e37_79b9_7f4a_7c15u64; + let mut next = |n: usize| { + seed ^= seed << 13; + seed ^= seed >> 7; + seed ^= seed << 17; + (seed % n as u64) as usize + }; + for _ in 0..2_000 { + let mut at = 0; + let mut splices = Vec::new(); + for _ in 0..next(6) { + let start = at + next(20); + let end = start + next(12); + splices.push(Splice { + part: 0, + local: start..end, + joined: start..end, + bytes: vec![b'x'; next(25)], + }); + at = end; + } + let map = SpanMap::new(&splices); + for _ in 0..40 { + let start = next(at + 30); + let span = start..start + next(10); + assert_eq!( + map.map(&span), + map_span(&splices, &span), + "{span:?} {splices:?}" + ); + } + } + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/verify/edited.rs b/src-tauri/src/pdf_engine/text_edit/verify/edited.rs new file mode 100644 index 0000000..bea6e7c --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/verify/edited.rs @@ -0,0 +1,320 @@ +//! Check 5 of the re-walk (§B.14): every edited run draws exactly the expected glyphs — text +//! (equal to the planned text and to the text the user typed), codes and fonts (of the run's +//! typing surface or the requested face's, recomputed from the page, never taken from the plan), +//! drawable codes (B3) — with the expected size, spacing and +//! fill and no other state change, at the expected origins; the pen after the primary is the +//! original one, and every absorbed member draws nothing and keeps its travel. + +use super::{decoded_text, dist, font_id, TextItem, VerifyFailure}; +use crate::pdf_engine::text_edit::fonts::{face_surface, typing_surface}; +use crate::pdf_engine::text_edit::limits::{DRIFT_TOLERANCE_PT, JOIN_BASELINE_TOL_PT}; +use crate::pdf_engine::text_edit::rewrite::{ExpectedRun, PagePlan, RunPlan}; +use crate::pdf_engine::text_edit::runs::TextRun; +use crate::pdf_engine::text_edit::state::{same_paint, same_state_except, StateField}; +use crate::pdf_engine::text_edit::walker::{PageWalk, RecElem, ShowRecord}; + +/// Effective size of an edited glyph vs the requested one (pt). +const SIZE_TOL_PT: f64 = 0.01; + +/// The before record of each member of `run`, found by span (id-free), with its index. +fn member_records<'w>(before: &'w PageWalk, run: &RunPlan) -> Option> { + run.expected + .member_spans + .iter() + .map(|span| { + before + .records + .iter() + .enumerate() + .find(|(_, r)| r.span.as_ref() == Some(span)) + }) + .collect() +} + +/// The fonts an edited glyph may use: the run's typing surface, or the face surface of the +/// requested face, recomputed from the page (not taken from the plan). +fn allowed_fonts(before: &PageWalk, primary: &ShowRecord, run: &RunPlan) -> Vec<(Vec, u64)> { + let font = primary.before.text.font.as_ref(); + if font.is_some_and(|f| f.from_extgstate) { + return font + .map(|f| (Vec::new(), f.content_hash)) + .into_iter() + .collect(); + } + let surface = match ( + run.target.face_changed, + run.target.face, + primary.font.as_ref(), + ) { + (true, Some(face), Some(model)) => face_surface(&before.page_fonts, model, face), + (true, _, _) => None, + (false, _, _) => font + .and_then(|f| f.resource.as_deref()) + .map(|res| typing_surface(&before.page_fonts, res)), + }; + surface + .map(|s| { + s.fonts + .iter() + .map(|(n, m)| (n.clone(), m.content_hash)) + .collect() + }) + .unwrap_or_default() +} + +/// One run's view of the after walk. +struct Edited<'a> { + exp: &'a ExpectedRun, + run: &'a RunPlan, + /// The before records of the members, primary first. + members: Vec<(usize, &'a ShowRecord)>, + /// Per member: the after record indices inside its replaced region. + regions: Vec<&'a Vec>, + /// The after records the primary's replacement emitted. + recs: Vec<&'a ShowRecord>, +} + +impl Edited<'_> { + fn fail(&self, what: &'static str) -> VerifyFailure { + VerifyFailure::EditedMismatch { + run_id: self.exp.run_id.clone(), + what, + } + } + + /// Every glyph the primary emitted, with its record. + fn glyphs(&self) -> Vec<(&ShowRecord, usize)> { + self.recs + .iter() + .flat_map(|r| (0..r.glyphs.len()).map(move |g| (*r, g))) + .collect() + } +} + +/// Check 5 for one run. +pub(super) fn check_edited_run( + before: &PageWalk, + before_runs: &[TextRun], + after: &PageWalk, + plan: &PagePlan, + edited: &[Vec], + run: &RunPlan, +) -> Result<(), VerifyFailure> { + let exp = &run.expected; + let records_fail = || VerifyFailure::EditedMismatch { + run_id: exp.run_id.clone(), + what: "records", + }; + let members = member_records(before, run).ok_or_else(records_fail)?; + let known = before_runs.iter().any(|r| { + r.reason.is_none() + && r.members.len() == members.len() + && r.members.iter().zip(&members).all(|(a, (b, _))| a == b) + }); + if !known { + return Err(records_fail()); + } + let regions: Vec<&Vec> = run + .splices + .iter() + .map(|s| { + plan.splices + .iter() + .position(|x| x == s) + .and_then(|k| edited.get(k)) + .ok_or_else(records_fail) + }) + .collect::>()?; + let primary_records: &[usize] = regions.first().copied().ok_or_else(records_fail)?; + if exp.emitted_records.first() != Some(&primary_records.len()) { + return Err(records_fail()); + } + let recs = primary_records + .iter() + .filter_map(|j| after.records.get(*j)) + .collect(); + let e = Edited { + exp, + run, + members, + regions, + recs, + }; + check_glyphs(before, &e)?; + check_state(&e, primary_records)?; + check_origins(&e)?; + check_absorbed(after, &e) +} + +/// Text (with synthetic spaces), codes, fonts and drawable codes. +fn check_glyphs(before: &PageWalk, e: &Edited<'_>) -> Result<(), VerifyFailure> { + let items = e.recs.iter().flat_map(|r| { + r.elems.iter().map(move |el| match el { + RecElem::Glyph(g) => TextItem::Glyph( + r.glyphs + .get(*g) + .and_then(|x| x.text.as_deref()) + .unwrap_or("\u{fffd}"), + ), + RecElem::Kern { value, .. } => TextItem::Kern(-value / 1000.0), + }) + }); + let text = decoded_text(items); + if text != e.exp.text { + return Err(e.fail("text")); + } + let glyphs = e.glyphs(); + if glyphs.len() != e.exp.glyphs.len() { + return Err(e.fail("codes")); + } + let primary = e + .members + .first() + .map(|(_, r)| *r) + .ok_or_else(|| e.fail("records"))?; + let allowed = allowed_fonts(before, primary, e.run); + for (k, ((rec, g), (font, hash, code))) in glyphs.iter().zip(&e.exp.glyphs).enumerate() { + let x = rec.glyphs.get(*g).ok_or_else(|| e.fail("codes"))?; + if x.code != *code { + return Err(e.fail("codes")); + } + let id = font_id(rec); + if id != (font.clone(), *hash) || !allowed.contains(&id) { + return Err(e.fail("font")); + } + let new = e.exp.glyph_new.get(k).copied().unwrap_or(true); + if new && !rec.font.as_ref().is_some_and(|m| m.drawable(x.code)) { + return Err(e.fail("glyph")); + } + } + // The planner's expectations are its own reading of its glyphs; the request is not. + if text != e.exp.requested_text { + return Err(e.fail("requested_text")); + } + Ok(()) +} + +/// Size, spacing and fill as requested; every other state field as the primary had it. +fn check_state(e: &Edited<'_>, primary_records: &[usize]) -> Result<(), VerifyFailure> { + let primary = e + .members + .first() + .map(|(_, r)| *r) + .ok_or_else(|| e.fail("records"))?; + let primary_font = font_id(primary); + let t = &e.run.target; + for (j, rec) in e.recs.iter().enumerate() { + if rec.glyphs.is_empty() { + continue; + } + let text = &rec.before.text; + let effective = rec.text_to_user[2].hypot(rec.text_to_user[3]); + if text.tfs != e.exp.tfs || (effective - e.exp.effective_size).abs() > SIZE_TOL_PT { + return Err(e.fail("size")); + } + if text.tc != e.exp.tc { + return Err(e.fail("tc")); + } + if !same_paint(&rec.before.fill, &e.exp.fill) { + return Err(e.fail("fill")); + } + let mut targeted = Vec::new(); + if t.size_changed { + targeted.push(StateField::Tfs); + } + if t.fill_changed { + targeted.push(StateField::Fill); + } + if t.tc_changed { + targeted.push(StateField::Tc); + } + if t.face_changed || font_id(rec) != primary_font { + targeted.push(StateField::Font); + } + same_state_except(&primary.before, &rec.before, &targeted).map_err(|field| { + VerifyFailure::StateChanged { + record: primary_records.get(j).copied().unwrap_or(0), + field, + } + })?; + } + Ok(()) +} + +/// The new layout's origins, the kept glyphs at their original places (shifted by the width +/// change when no style moved them), and the pen after the primary. +fn check_origins(e: &Edited<'_>) -> Result<(), VerifyFailure> { + let exp = e.exp; + if let Some(first) = e.recs.first() { + if dist(first.pen_before, exp.origin) > DRIFT_TOLERANCE_PT { + return Err(e.fail("origin")); + } + } + let t = &e.run.target; + let unstyled = !t.size_changed && !t.tc_changed && !t.face_changed; + let suffix_from = exp.glyphs.len().saturating_sub(exp.suffix_glyphs); + for (k, (rec, g)) in e.glyphs().into_iter().enumerate() { + let x = rec.glyphs.get(g).ok_or_else(|| e.fail("origin"))?; + let want = exp + .glyph_origins + .get(k) + .copied() + .ok_or_else(|| e.fail("origin"))?; + if dist(x.origin, want) > DRIFT_TOLERANCE_PT { + return Err(e.fail("origin")); + } + let middle = k >= exp.prefix_glyphs && k < suffix_from; + let Some(Some((member, gi))) = exp.kept_from.get(k) else { + continue; + }; + if !unstyled || middle { + continue; + } + let orig = e + .members + .get(*member) + .and_then(|(_, r)| r.glyphs.get(*gi)) + .map(|o| o.origin) + .ok_or_else(|| e.fail("origin"))?; + let shifted = k >= suffix_from && !exp.unshifted_from.is_some_and(|u| k >= u); + let expected = if shifted { + (orig.0 + exp.shift_user.0, orig.1 + exp.shift_user.1) + } else { + orig + }; + if dist(x.origin, expected) > DRIFT_TOLERANCE_PT + JOIN_BASELINE_TOL_PT { + return Err(e.fail("origin")); + } + } + let pen = e + .recs + .last() + .map(|r| r.pen_after) + .ok_or_else(|| e.fail("pen"))?; + if dist(pen, exp.primary_pen_after) > DRIFT_TOLERANCE_PT { + return Err(e.fail("pen")); + } + Ok(()) +} + +/// Absorbed members: one show op drawing nothing, with the original pen travel. +fn check_absorbed(after: &PageWalk, e: &Edited<'_>) -> Result<(), VerifyFailure> { + for (m, region) in e.regions.iter().enumerate().skip(1) { + let ok = region.len() == 1 + && e.exp.emitted_records.get(m) == Some(&1) + && region + .first() + .and_then(|j| after.records.get(*j)) + .is_some_and(|r| { + r.glyphs.is_empty() + && e.exp + .member_pen_after + .get(m) + .is_some_and(|p| dist(r.pen_after, *p) <= DRIFT_TOLERANCE_PT) + }); + if !ok { + return Err(e.fail("absorbed")); + } + } + Ok(()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/walker.rs b/src-tauri/src/pdf_engine/text_edit/walker.rs new file mode 100644 index 0000000..d779bc0 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/walker.rs @@ -0,0 +1,705 @@ +//! Page walker (SPEC §B.10): interprets one page's lexed content with the full graphics state and +//! records every show op (`ShowRecord`) and every paint (`PaintRecord`) in paint order, with +//! id-free state digests (D32). +//! +//! Modes: `Edit` walks depth 0 only (a Form `Do` is a paint); `Classify` descends Forms (depth +//! ≤ 8, cycle check, `/Matrix`, `/BBox` as a clip) for the #33 classifier; `Wrapped` (Phase B) +//! follows exactly the one depth-0 `Do` of qpdf's overlay wrapper Form and treats its content as +//! depth 0. Every page-level problem (lexer, decode, budgets, `q` overflow, geometry, +//! unresolvable references, non-finite numbers) becomes `page_reason` with empty records — never a +//! partial list. `walk_page` never fails. +//! +//! Work and memory are bounded per walk: every op costs O(its own operands) — a path op checks +//! only the points it adds, a `Tf` finds its root font by name, an optional-content group is a +//! binary search in lists read once per snapshot (`structure::OcConfig`), and other lookups that +//! would repeat per op are memoised. Records and paints drawn under an unchanged state share one +//! digest (`Exec::digest`), and digests share their variable-size parts (`state`); deep hashes are +//! memoised per object and charged to the walk's budgets (`hash`). What a walk keeps per element +//! is budgeted per page, not only per op: glyphs (`GLYPHS_PER_PAGE_MAX`), TJ kerns +//! (`KERNS_PER_PAGE_MAX`), characters of glyph text (`TEXT_CHARS_PER_PAGE_MAX`) and lexed operand +//! nodes alive at once (`OPERAND_NODES_MAX`). Above them, every allocation that grows with the +//! page is charged to one byte budget before it is made (`budget::ModelBudget`, +//! `PAGE_MODEL_BYTES_MAX`), so no small page can build a model of GiBs, whatever it repeats. + +pub(crate) mod budget; +mod fonts; +mod gfx; +pub(crate) mod hash; +pub(crate) mod show; +mod xobj; + +use crate::pdf_engine::text_edit::content::{check_part_joins, PageContent}; +use crate::pdf_engine::text_edit::context::{resolve, Res, SnapshotContext}; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::fonts::{Code, FontKey, FontModel}; +use crate::pdf_engine::text_edit::geometry::{page_geometry, Matrix, PageGeometry, IDENTITY}; +use crate::pdf_engine::text_edit::lexer::{Op, Operator, Span}; +use crate::pdf_engine::text_edit::limits::{ + GLYPHS_PER_PAGE_MAX, LEX_CANCEL_EVERY_OPS, PAGE_DECODE_BUDGET, PAGE_OPS_MAX, Q_DEPTH_MAX, +}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::state::{Bytes, GState, MarkedStack, OcState, StateDigest}; +use budget::{content_bytes, lex_ops, map_entry, ModelBudget}; +use lopdf::{Document, Object, ObjectId}; +use std::collections::{HashMap, HashSet}; +use std::sync::atomic::{AtomicBool, Ordering}; +use std::sync::Arc; + +/// TJ numbers per page. Each one is kept as a kern element of its record and a unit of its run +/// (§B.2 counts only glyphs): a page of kern-only TJ arrays would otherwise turn 50 KB of +/// deflated content into a model of GiBs (review T3 r2 HIGH-1). +pub(crate) const KERNS_PER_PAGE_MAX: usize = GLYPHS_PER_PAGE_MAX; +/// Characters of glyph text per page. A ToUnicode destination may give one glyph 256 of them, +/// each kept in the glyph, its run unit, the run text and a caret offset (review T3 r2 HIGH-2). +pub(crate) const TEXT_CHARS_PER_PAGE_MAX: usize = 2 * GLYPHS_PER_PAGE_MAX; +/// Lexed operand nodes alive at once in one walk (`limits::OPERAND_NODES_MAX`, also the lexer's +/// own cap): 48 MiB of `[1 1 … 1] 0 d` would otherwise build 24 M nodes (≈ 1.2 GiB). +pub(crate) use crate::pdf_engine::text_edit::limits::OPERAND_NODES_MAX; + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum ShowOp { + Tj, + TJ, + Quote, + DoubleQuote, +} + +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum WalkMode { + Edit, + Classify, + Wrapped { name: Vec }, +} + +#[derive(Debug, Clone)] +pub struct GlyphRec { + pub code: Code, + pub text: Option, + pub origin: (f64, f64), + /// User-space displacement of the pen for this glyph (incl. Tc, Tw, Th). + pub advance_user: (f64, f64), + pub width1000: f64, + /// Ink box from the font's ascent/descent (rise included), user-space AABB `[x0, y0, x1, y1]`. + pub bbox: [f64; 4], +} + +#[derive(Debug, Clone)] +pub enum RecElem { + Glyph(usize), + Kern { value: f64, span: Span }, +} + +#[derive(Debug, Clone)] +pub struct ShowRecord { + pub seq: u32, + pub depth: u8, + pub form_chain: Vec, + pub op: ShowOp, + /// The op span in the joined buffer (depth-0 page content only). + pub span: Option, + /// The op span in the buffer that was lexed (joined buffer, or the Form's data). + pub local_span: Span, + pub operand_spans: Vec, + /// The state the glyphs are drawn with (for `"` after its `aw Tw ac Tc` side effects). + /// Shared: consecutive records (and paints) drawn under an unchanged state hold one digest. + pub before: Arc, + pub after: Arc, + pub tm_before: Matrix, + pub tm_after: Matrix, + pub glyphs: Vec, + pub elems: Vec, + /// User space, rise included (= the first glyph's origin when the op starts with a glyph). + pub pen_before: (f64, f64), + pub pen_after: (f64, f64), + /// Σ(w0·Tfs + Tc + Tw·[word space]) + Σ(−n/1000·Tfs), Th excluded. + pub advance_ts: f64, + /// `[Tfs·Th 0 0 Tfs 0 Ts] × Tm × CTM` at the first glyph. + pub text_to_user: Matrix, + pub after_unproven_inline_image: bool, + /// Lookup handles, never compared across files. + pub font_key: Option, + pub font: Option>, + /// A 2-byte font was shown an odd number of bytes. + pub split_error: bool, + /// The op is drawn where an earlier show op of the same text object left the pen, and that + /// op advanced by an amount the model cannot know (an odd-length 2-byte string, a missing + /// font, a code without a width, a glyph of a font whose widths are not proven to be the + /// viewers' — `show::advance_proven`), or after a `Q` inside the text object restored a + /// state saved under another text matrix (`Exec::tm_unsettled`): viewers place this text + /// elsewhere (`MISSING_WIDTHS`). + pub pen_unknown: bool, +} + +/// Id-free (D32): resource names plus deep hashes. +#[derive(Debug, Clone, PartialEq)] +pub enum PaintKind { + Path(Operator), + Shading { name: Vec, hash: u64 }, + ImageXObject { name: Vec, hash: u64 }, + InlineImage { hash: u64 }, + FormXObject { name: Vec, hash: u64 }, +} + +/// The XObject behind a `Do` (classifier only): its id and whether it was reached through a +/// resource dictionary other pages can share. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct XObjectUse { + pub id: ObjectId, + pub shared_path: bool, +} + +#[derive(Debug, Clone)] +pub struct PaintRecord { + pub seq: u32, + pub depth: u8, + pub kind: PaintKind, + pub span: Option, + pub state: Arc, + pub bbox: Option<[f64; 4]>, + pub masked: bool, + pub form_chain: Vec, + pub local_span: Span, + pub xobject: Option, +} + +pub struct PageWalk { + pub page_id: ObjectId, + pub geometry: PageGeometry, + /// The page's lexed ops (empty in a page model: `runs::model_of` drops them, see + /// `ops_bytes`). + pub ops: Vec, + pub records: Vec, + pub paints: Vec, + /// The fonts of the depth-0 resources, first-use order, then the unused ones. + pub page_fonts: Vec<(Vec, Arc)>, + pub page_reason: Option, + /// Technical detail of `page_reason` (byte offset, check name). + pub page_detail: Option, + /// Bytes of the page-model budget charged for what is kept: by `walk_page`, the content it + /// walked and the walk (0 when refused), which the run stage goes on from; once + /// `runs::model_of` has built a model on the walk, the whole model's (its runs and the font + /// models it keeps included — each distinct model once, `FontModel::approx_bytes`, a font + /// shared by several pages charged by each — its ops not). Never less than that model's + /// `approx_bytes`, never more than `PAGE_MODEL_BYTES_MAX`. + pub model_bytes: usize, + /// The part of `model_bytes` that is `ops` (released when a model drops them). + pub ops_bytes: usize, +} + +impl PageWalk { + fn empty(page_id: ObjectId, geometry: PageGeometry) -> PageWalk { + PageWalk { + page_id, + geometry, + ops: Vec::new(), + records: Vec::new(), + paints: Vec::new(), + page_fonts: Vec::new(), + page_reason: None, + page_detail: None, + model_bytes: 0, + ops_bytes: 0, + } + } + + fn refused(self, reason: TextReason, detail: impl Into) -> PageWalk { + PageWalk { + page_reason: Some(reason), + page_detail: Some(detail.into()), + ..PageWalk::empty(self.page_id, self.geometry) + } + } + + /// Releases the walk itself (ops, records, paints and page fonts), keeping its page, geometry + /// and reason: a model refused after the walk keeps nothing it cannot use (review T3-budget + /// MEDIUM-1). Returns the bytes released from `model_bytes`. + pub(crate) fn release(&mut self) -> usize { + let kept = PageWalk { + page_reason: self.page_reason, + page_detail: self.page_detail.take(), + ..PageWalk::empty(self.page_id, self.geometry.clone()) + }; + let released = self.model_bytes; + *self = kept; + released + } + + /// Drops the lexed ops (nothing reads a model's ops) and their charge. + pub(crate) fn drop_ops(&mut self) { + self.ops = Vec::new(); + self.model_bytes = self.model_bytes.saturating_sub(self.ops_bytes); + self.ops_bytes = 0; + } +} + +/// A page-level refusal raised while walking. +#[derive(Debug, Clone)] +pub(crate) struct Stop { + pub reason: TextReason, + pub detail: String, +} + +impl Stop { + pub(crate) fn new(reason: TextReason, detail: impl Into) -> Stop { + Stop { + reason, + detail: detail.into(), + } + } + + pub(crate) fn malformed(detail: impl Into) -> Stop { + Stop::new(TextReason::MalformedContent, detail) + } +} + +pub(crate) type Walked = Result<(), Stop>; + +/// Walks one page (`_page_index` names it for the caller only). Never fails: problems are +/// `page_reason`. +pub fn walk_page( + ctx: &SnapshotContext, + _page_index: u32, + content: &PageContent, + mode: WalkMode, + cancel: Option<&AtomicBool>, +) -> PageWalk { + let mem = ModelBudget::new(content_bytes(content)); + walk_page_within(ctx, content, mode, cancel, mem) +} + +/// `walk_page` under `mem`, which has already charged `content` and whatever else stays alive +/// while the walk runs (a Classify pass runs on the budget its page model left, review +/// T3-budget MEDIUM-2). +pub(crate) fn walk_page_within( + ctx: &SnapshotContext, + content: &PageContent, + mode: WalkMode, + cancel: Option<&AtomicBool>, + mut mem: ModelBudget, +) -> PageWalk { + let doc = ctx.doc(); + let page_id = content.page_id; + let (geometry, geometry_reason) = match page_geometry(doc, page_id) { + Ok(g) => (g, None), + Err(reason) => ( + PageGeometry::unbounded(lenient_rotation(doc, page_id)), + Some(reason), + ), + }; + let walk = PageWalk::empty(page_id, geometry.clone()); + if let (Some(reason), false) = (geometry_reason, mode == WalkMode::Classify) { + return walk.refused(reason, "page geometry"); + } + // The page's lexed ops are charged first; then every part boundary must read alike in every + // viewer (`content::joins`). + let lexed = lex_ops(&content.joined, PAGE_OPS_MAX, 0, &mut mem, cancel, true, ""); + let (ops, nodes, ops_bytes) = match lexed { + Ok(l) => (l.ops, l.nodes, l.bytes), + Err(stop) => return walk.refused(stop.reason, stop.detail), + }; + if let Err(at) = check_part_joins(content) { + let detail = format!("content parts joined inside a token or comment at byte {at}"); + return walk.refused(TextReason::MalformedContent, detail); + } + let res = match Res::of_page(doc, page_id) { + Ok(r) => r, + Err(what) => return walk.refused(TextReason::MalformedContent, what), + }; + let mut w = Walker { + ctx, + doc, + mode: mode.clone(), + cancel, + geometry, + budget: DecodeBudget::new(PAGE_DECODE_BUDGET.saturating_sub(content.joined.len())), + hash_budget: DecodeBudget::new(PAGE_DECODE_BUDGET), + hash_steps: 0, + hash_raw_left: ctx.snap.file_len(), + records: Vec::new(), + paints: Vec::new(), + root_fonts: Vec::new(), + root_font_names: HashSet::new(), + root_res: res, + seq: 0, + ops_left: PAGE_OPS_MAX.saturating_sub(ops.len()), + glyphs: 0, + kerns: 0, + text_chars: 0, + operand_nodes: nodes, + form_paints: 0, + hashes: HashMap::new(), + direct_hashes: HashMap::new(), + oc_states: HashMap::new(), + gs_others: HashMap::new(), + others_interned: HashSet::new(), + gs_last: None, + fonts: HashMap::new(), + pen_proofs: HashMap::new(), + mem, + }; + let frame = Frame { + bytes: &content.joined, + res, + depth: 0, + chain: Vec::new(), + joined: true, + root: true, + }; + let result = match &mode { + WalkMode::Wrapped { name } => w.walk_wrapped(&frame, &ops, name, page_id), + _ => { + let mut ex = Exec::new(GState::initial(IDENTITY), MarkedStack::default()); + w.run(&frame, &mut ex, &ops) + } + } + .and_then(|()| w.finish_page_fonts()); + match result { + Err(stop) => walk.refused(stop.reason, stop.detail), + Ok(()) => PageWalk { + ops, + records: w.records, + paints: w.paints, + page_fonts: w.root_fonts, + page_reason: geometry_reason, + page_detail: geometry_reason.map(|_| "page geometry".to_string()), + model_bytes: w.mem.held(), + ops_bytes, + ..walk + }, + } +} + +/// A walk refused before it started (e.g. the page content could not be decoded). +pub(crate) fn refused_walk( + ctx: &SnapshotContext, + page_id: ObjectId, + reason: TextReason, + detail: &str, +) -> PageWalk { + let doc = ctx.doc(); + let geometry = page_geometry(doc, page_id) + .unwrap_or_else(|_| PageGeometry::unbounded(lenient_rotation(doc, page_id))); + PageWalk::empty(page_id, geometry).refused(reason, detail) +} + +/// `/Rotate` of a page whose geometry was refused, read leniently (classifier display only). +fn lenient_rotation(doc: &Document, page_id: ObjectId) -> i64 { + let mut cur = Some(page_id); + for _ in 0..crate::pdf_engine::text_edit::limits::PAGE_TREE_DEPTH_MAX { + let Some(Object::Dictionary(d)) = cur.and_then(|id| doc.objects.get(&id)) else { + return 0; + }; + if let Some((_, Object::Integer(r))) = d.get(b"Rotate").ok().and_then(|o| resolve(doc, o)) { + return if r.rem_euclid(90) == 0 { + r.rem_euclid(360) + } else { + 0 + }; + } + cur = d.get(b"Parent").ok().and_then(|o| o.as_reference().ok()); + } + 0 +} + +/// One content stream being executed: the buffer its ops were lexed from and its resources. +pub(crate) struct Frame<'f, 'a> { + pub bytes: &'f [u8], + pub res: Res<'a>, + pub depth: u8, + pub chain: Vec, + /// Spans are joined-buffer spans (page content). + pub joined: bool, + /// Depth-0 content whose `/Font` entries are the page fonts. + pub root: bool, +} + +impl Frame<'_, '_> { + /// The verbatim bytes of `op` (its whole span), shared. + pub(crate) fn op_bytes(&self, op: &Op) -> Bytes { + Arc::from(self.bytes.get(op.span.clone()).unwrap_or_default()) + } + + pub(crate) fn span_bytes(&self, span: &Span) -> &[u8] { + self.bytes.get(span.clone()).unwrap_or_default() + } +} + +#[derive(Debug, Clone, Default)] +pub(crate) struct PathState { + /// Bounding box of every point added so far (a path keeps no point list: an unpainted path of + /// 250,000 ops costs nothing). + pub bbox: Option<[f64; 4]>, + pub re_count: usize, + pub re_rect: Option<[f64; 4]>, + pub re_axis: bool, + pub other: bool, +} + +/// What `q` saves: the graphics state, `Tm`, `Tlm` and `Exec::tm_base`. +pub(crate) type Saved = (GState, Matrix, Matrix, Matrix); + +/// Execution state of one content stream. +pub(crate) struct Exec { + pub gs: GState, + /// `q` saves the graphics state with the text matrices in force (`Q` compares them). + pub stack: Vec, + /// Shared with the digests taken under it (a persistent list: push and pop are O(1)). + pub marked: MarkedStack, + pub marked_base: usize, + pub in_text: bool, + pub tm: Matrix, + pub tlm: Matrix, + /// The matrix the last `Tm` set (the identity at `BT`): poppler's text matrix, on which its + /// line start (moved by `Td`, kept by `Q`) is based. + pub tm_base: Matrix, + pub text_clip: bool, + pub path: PathState, + pub clip_pending: bool, + /// `tm` carries an advance the model cannot know (cleared by `BT` and every line move). + pub pen_unknown: bool, + /// A `Q` inside this text object restored a state saved under other text matrices, or under + /// another `tm_base` (review T3 r3 LOW-1). `q`/`Q` are not allowed in a text object (ISO + /// 32000-1 §8.2) and viewers differ: the model and poppler keep the text position (poppler + /// restores `Tm` but not the line start), pdf.js restores both. Only `Tm` or `BT` settle it; a + /// `Td` is relative to the disputed line start. + pub tm_unsettled: bool, + /// The last digest taken, handed out again while the state is unchanged. + digest_memo: Option>, +} + +impl Exec { + pub(crate) fn new(gs: GState, marked: MarkedStack) -> Exec { + Exec { + gs, + stack: Vec::new(), + marked_base: marked.len(), + marked, + in_text: false, + tm: IDENTITY, + tlm: IDENTITY, + tm_base: IDENTITY, + text_clip: false, + path: PathState::default(), + clip_pending: false, + pen_unknown: false, + tm_unsettled: false, + digest_memo: None, + } + } + + /// The digest of the current state, shared with the previous one when nothing changed (an + /// O(1) identity check, `StateDigest::is_digest_of`): a page of 250,000 show ops under one + /// state holds one digest, not 500,000. + /// A marked-content stack equal by value to the last digest's (an `EMC` then a `BMC` of the + /// same tag makes a new top node) is replaced by that digest's stack first, so a page that + /// reopens a stack around every op still shares one digest. + pub(crate) fn digest(&mut self) -> Arc { + if let Some(memo) = &self.digest_memo { + if !memo.marked.same(&self.marked) && memo.marked == self.marked { + self.marked = memo.marked.clone(); + } + if memo.is_digest_of(&self.gs, &self.marked) { + return Arc::clone(memo); + } + } + let digest = Arc::new(self.gs.digest(&self.marked)); + self.digest_memo = Some(Arc::clone(&digest)); + digest + } +} + +pub(crate) struct Walker<'a> { + pub ctx: &'a SnapshotContext, + pub doc: &'a Document, + pub mode: WalkMode, + pub cancel: Option<&'a AtomicBool>, + pub geometry: PageGeometry, + /// Content, Form and font decoding (§B.2 PAGE_DECODE_BUDGET). + pub budget: DecodeBudget, + /// Decoding for deep hashes of XObjects and colour spaces (same size, kept apart so a large + /// image never refuses the page's text). A failed decode is charged as if it ran to its cap. + pub hash_budget: DecodeBudget, + /// Objects visited by all deep hashes of this walk (`hash::HASH_STEPS_MAX`). + pub hash_steps: usize, + /// Raw (undecoded) stream bytes the deep hashes of this walk may still read: the file's size + /// (each stream is hashed once per walk; objects that alias one byte range cannot multiply it). + pub hash_raw_left: usize, + pub records: Vec, + pub paints: Vec, + pub root_fonts: Vec<(Vec, Arc)>, + /// The names in `root_fonts` (a page may switch fonts on every op). + pub root_font_names: HashSet>, + pub root_res: Res<'a>, + pub seq: u32, + pub ops_left: usize, + pub glyphs: usize, + /// TJ numbers shown so far (`KERNS_PER_PAGE_MAX`). + pub kerns: usize, + /// Characters of glyph text so far (`TEXT_CHARS_PER_PAGE_MAX`). + pub text_chars: usize, + /// Lexed operand nodes alive now: the page's, plus those of the Forms being run + /// (`OPERAND_NODES_MAX`). + pub operand_nodes: usize, + pub form_paints: usize, + pub hashes: HashMap, + /// Deep hashes of direct values, by address (the document is borrowed for the whole walk, + /// so an address names one value). + pub direct_hashes: HashMap, + /// Optional-content state of each `/OC` group named by a `BDC`, once per walk. + pub oc_states: HashMap, + /// The unmodelled keys of each ExtGState dictionary a `gs` applied (sorted, one per key), + /// by address like `direct_hashes`: re-applying one dictionary reuses its list, and a merge + /// that changes nothing keeps the shared list in force (`GsEffects::set_others`). + pub gs_others: HashMap, + /// Every merged list of unmodelled keys this walk has put in force, by value. + pub others_interned: HashSet, + /// The last `gs`: its dictionary's key list and the list it put in force. Re-applying that + /// dictionary to that list changes nothing, so it costs O(1) (review T3 r3 MEDIUM-2). + pub gs_last: Option<(Others, Others)>, + pub fonts: HashMap>, + /// Per loaded font: whether viewers advance by the model's widths (`show::advance_proven`). + pub pen_proofs: HashMap, + /// The page-model byte budget of this walk. + pub mem: ModelBudget, +} + +/// A shared list of unmodelled ExtGState keys (`GsEffects::other`). +pub(crate) type Others = Arc<[(Bytes, u64)]>; + +impl<'a> Walker<'a> { + pub(crate) fn next_seq(&mut self) -> Result { + let s = self.seq; + self.seq = self + .seq + .checked_add(1) + .ok_or_else(|| Stop::new(TextReason::PageTooComplex, "paint order"))?; + Ok(s) + } + + fn cancelled(&self) -> bool { + self.cancel.is_some_and(|c| c.load(Ordering::Relaxed)) + } + + /// `Exec::digest`, charged to the budget (the digest and its parts not counted yet). + pub(crate) fn digest(&mut self, ex: &mut Exec) -> Result, Stop> { + let d = ex.digest(); + self.mem.digest(&d)?; + Ok(d) + } + + pub(crate) fn run(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, ops: &[Op]) -> Walked { + for (i, op) in ops.iter().enumerate() { + if i % LEX_CANCEL_EVERY_OPS == 0 && self.cancelled() { + return Err(Stop::new(TextReason::PageTooComplex, "cancelled")); + } + self.op(frame, ex, op)?; + } + Ok(()) + } + + fn op(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + use Operator as O; + match op.operator { + O::q => { + if ex.stack.len() >= Q_DEPTH_MAX { + return Err(Stop::malformed(format!( + "q nesting deeper than {Q_DEPTH_MAX} at byte {}", + op.op_span.start + ))); + } + ex.stack.push((ex.gs.clone(), ex.tm, ex.tlm, ex.tm_base)); + } + O::Q => { + if let Some((prev, tm, tlm, base)) = ex.stack.pop() { + ex.gs = prev; + let same = same_bits(&tm, &ex.tm) + && same_bits(&tlm, &ex.tlm) + && same_bits(&base, &ex.tm_base); + if ex.in_text && !same { + ex.tm_unsettled = true; + } + ex.tm_base = base; + } + } + O::cm => gfx::concat(ex, op)?, + O::w | O::J | O::j | O::M | O::d | O::ri | O::i => gfx::line_param(ex, op)?, + O::gs => self.ext_gstate(frame, ex, op)?, + O::m | O::l | O::c | O::v | O::y | O::h | O::re => gfx::path(ex, op)?, + O::S | O::s | O::f | O::F | O::fStar | O::B | O::BStar | O::b | O::bStar | O::n => { + self.paint_path(frame, ex, op)? + } + O::W | O::WStar => ex.clip_pending = true, + O::CS + | O::cs + | O::SC + | O::sc + | O::SCN + | O::scn + | O::G + | O::g + | O::RG + | O::rg + | O::K + | O::k => self.color(frame, ex, op)?, + O::BT => { + if ex.in_text { + return Err(Stop::malformed(format!( + "BT inside a text object at byte {}", + op.op_span.start + ))); + } + ex.in_text = true; + ex.tm = IDENTITY; + ex.tlm = IDENTITY; + ex.tm_base = IDENTITY; + ex.text_clip = false; + ex.pen_unknown = false; + ex.tm_unsettled = false; + } + O::ET => { + if ex.in_text && ex.text_clip { + ex.gs.clip = crate::pdf_engine::text_edit::state::ClipState::Complex; + } + ex.in_text = false; + ex.text_clip = false; + } + O::Tc | O::Tw | O::Tz | O::TL | O::Tf | O::Tr | O::Ts => { + self.text_state(frame, ex, op)? + } + O::Td | O::TD | O::Tm | O::TStar => show::position(ex, op)?, + O::Tj | O::TJ | O::Quote | O::DoubleQuote => self.show(frame, ex, op)?, + O::Do => self.do_xobject(frame, ex, op)?, + O::BI => self.inline_image(frame, ex, op)?, + O::sh => self.shading(frame, ex, op)?, + O::BMC | O::BDC | O::EMC | O::MP | O::DP => self.marked(frame, ex, op)?, + O::d0 | O::d1 | O::BX | O::EX | O::EI | O::ID | O::Unknown => {} + } + Ok(()) + } + + /// The state of the optional-content group `props` names (`OcConfig::state`, whose lists are + /// read once per snapshot), once per group and walk: a page may open the same layer in every + /// `BDC`. + pub(crate) fn oc_state( + &mut self, + id: Option, + props: &Object, + ) -> Result { + let config = self.ctx.oc_config(); + let Some(id) = id else { + return Ok(config.state(self.doc, props)); + }; + if let Some(state) = self.oc_states.get(&id) { + return Ok(*state); + } + let state = config.state(self.doc, &Object::Reference(id)); + self.mem.scratch(map_entry::())?; + self.oc_states.insert(id, state); + Ok(state) + } +} + +/// Bit-for-bit equal matrices (a `q … Q` that moved nothing leaves them untouched). +fn same_bits(a: &Matrix, b: &Matrix) -> bool { + a.iter().zip(b).all(|(x, y)| x.to_bits() == y.to_bits()) +} diff --git a/src-tauri/src/pdf_engine/text_edit/walker/budget.rs b/src-tauri/src/pdf_engine/text_edit/walker/budget.rs new file mode 100644 index 0000000..5744208 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/walker/budget.rs @@ -0,0 +1,437 @@ +//! The page-model byte budget (`limits::PAGE_MODEL_BYTES_MAX`, review T3 r1–r3): one tally of +//! the bytes a page model build (`walk_page`, then `build_runs`) holds at once, and — on the same +//! tally, the model's bytes already charged — the #33 Classify pass run for that model. Every +//! allocation that grows with the page is charged *before* it is made — lexed ops, records with +//! their glyphs, kerns and text, paints, state digests and the shared parts they hold, the font +//! models the walk keeps, page font names, the walk's memo maps, a Form's decoded content and ops +//! while it runs, runs with their units, text, carets and surfaces, and the run stage's scratch — +//! and the page is `PAGE_TOO_COMPLEX` "page model size" the moment the next charge would pass the +//! cap. A vector's growth also needs room for its old buffer, which lives until the copy is made. +//! +//! The per-element budgets (glyphs, kerns, characters, operand nodes) bound one kind of element +//! each; this one bounds their sum and every allocation of the same kind that no review has found +//! yet: a small file cannot make any charged path hold more than the cap, whatever it repeats. +//! +//! Charges use `approx_bytes`'s accounting (`runs/size.rs` calls the same helpers): the bytes each +//! allocation requests — vectors by capacity, each shared allocation once by address (`Shared`). +//! `held` is what the model keeps (`PageWalk::model_bytes` hands the walk's part to the run +//! stage); `scratch` lives only while a walk or the run stage runs (memo maps, estimated per entry +//! with `map_entry`; a Form's data and ops, released when it ends). +//! +//! Font models are charged once per allocation (`FontModel::approx_bytes`) when the walk loads +//! them: a model keeps its fonts alive in its records and page fonts however the snapshot's +//! `FontCache` evicts (review T3-budget HIGH-1), so they count in `model_bytes` and +//! `approx_bytes`, whichever pages share them. +//! +//! Not charged, and why it stays bounded: the joined content's decode (`PAGE_DECODE_BUDGET`, +//! before the walk; the content itself is charged), and a state part (verbatim op bytes, a name) +//! between the op that copies it and the first digest that holds it — the current state and its +//! `q` stack hold at most one copy of each op span they were set from, so never more than the +//! decoded content. + +use super::{Stop, OPERAND_NODES_MAX}; +use crate::pdf_engine::text_edit::content::{ContentPart, PageContent}; +use crate::pdf_engine::text_edit::fonts::FontModel; +use crate::pdf_engine::text_edit::lexer::{ + lex_content, LexError, LexLimits, Op, Operand, OPERAND_NODES, +}; +use crate::pdf_engine::text_edit::limits::model_bytes_max; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::state::{Bytes, ColorSpaceKind, MarkedNode, StateDigest}; +use std::collections::HashSet; +use std::mem::size_of; +use std::sync::atomic::AtomicBool; +use std::sync::Arc; + +/// `page_detail` of a page refused by the budget. +pub(crate) const MODEL_SIZE: &str = "page model size"; + +/// The two reference counts in front of every `Arc` allocation. +pub(crate) const ARC_HEADER: usize = 2 * size_of::(); + +/// The bytes an `Arc<[T]>` of `len` items of `item` bytes requests (header plus data, padded to +/// the header's alignment). +pub(crate) fn arc_slice(len: usize, item: usize) -> usize { + let data = len.saturating_mul(item); + ARC_HEADER.saturating_add(data.div_ceil(size_of::()) * size_of::()) +} + +/// What one more entry of a `HashMap` / `HashSet` may cost: key, value and control byte +/// for the 2 × 8/7 buckets per entry right after a growth, plus the table being replaced. +pub(crate) const fn map_entry() -> usize { + 4 * (size_of::() + size_of::() + 1) +} + +fn addr(a: &Arc) -> usize { + Arc::as_ptr(a).cast::() as usize +} + +/// Distinct digests and the shared allocations they hold, each counted once (by address: every +/// counted one is held by a record or paint that lives as long as the walk, so no two share an +/// address). +#[derive(Default)] +pub(crate) struct Shared { + seen: HashSet, + bytes: usize, +} + +impl Shared { + /// Bytes counted so far. + pub(crate) fn bytes(&self) -> usize { + self.bytes + } + + /// Adds `bytes` for the allocation at `at` unless it was counted already; true when new. + fn once(&mut self, at: usize, bytes: usize) -> bool { + let new = self.seen.insert(at); + if new { + self.bytes = self.bytes.saturating_add(bytes); + } + new + } + + fn bytes_of(&mut self, b: &Bytes) { + self.once(addr(b), arc_slice(b.len(), 1)); + } + + fn maybe(&mut self, b: &Option) { + if let Some(b) = b { + self.bytes_of(b); + } + } + + /// Counts the digest `d` and every shared part it holds that was not counted yet. + pub(crate) fn add(&mut self, d: &Arc) { + if !self.once(addr(d), ARC_HEADER + size_of::()) { + return; + } + for p in [&d.fill, &d.stroke] { + self.bytes = self + .bytes + .saturating_add(p.comps.capacity() * size_of::()); + self.maybe(&p.space_op); + self.maybe(&p.color_op); + if let ColorSpaceKind::Named(name, _) = &p.space { + self.bytes_of(name); + } + } + let g = &d.gs; + self.bytes_of(&g.blend); + self.bytes_of(&g.rendering_intent); + self.once(addr(&g.dash.0), arc_slice(g.dash.0.len(), size_of::())); + let other = arc_slice(g.other.len(), size_of::<(Bytes, u64)>()); + if self.once(addr(&g.other), other) { + for (key, _) in g.other.iter() { + self.bytes_of(key); + } + } + if let Some(f) = &d.text.font { + self.maybe(&f.resource); + self.maybe(&f.tf_op); + } + self.maybe(&d.text.tc_src); + self.maybe(&d.text.tw_src); + // Nodes are shared from the top down: below the first one already counted, all are. + for (at, entry) in d.marked.nodes() { + if !self.once(at, ARC_HEADER + size_of::()) { + break; + } + self.bytes_of(&entry.tag); + } + } + + fn entries(&self) -> usize { + self.seen.len() + } +} + +/// One lexed op's heap: every operand node (arrays and dictionaries walked without recursion) +/// and their bytes, and an inline image's dictionary. +pub(crate) fn op_heap(op: &Op) -> usize { + let mut total = op.operands.capacity() * size_of::(); + let mut stack: Vec<&Operand> = Vec::new(); + let image = op.inline_image.iter().flat_map(|i| { + let keys: usize = i.dict.iter().map(|(k, _)| k.capacity()).sum(); + let entries = i.dict.capacity() * size_of::<(Vec, Operand)>(); + std::iter::once(keys + entries) + }); + total = image.fold(total, usize::saturating_add); + let values = op.inline_image.iter().flat_map(|i| i.dict.iter()); + for o in op.operands.iter().chain(values.map(|(_, v)| v)) { + total = total.saturating_add(operand_node(o, &mut stack)); + while let Some(inner) = stack.pop() { + total = total.saturating_add(operand_node(inner, &mut stack)); + } + } + total +} + +/// What a page's content costs the model that keeps it: the joined buffer and its parts. +pub(crate) fn content_bytes(content: &PageContent) -> usize { + content + .joined + .capacity() + .saturating_add(content.parts.capacity() * size_of::()) +} + +/// The lexed ops of one stream: the vector and every op's heap. +pub(crate) fn ops_bytes(ops: &Vec) -> usize { + ops.iter().map(op_heap).fold( + ops.capacity().saturating_mul(size_of::()), + usize::saturating_add, + ) +} + +/// The heap bytes of one operand node; its children are pushed on `stack`. +fn operand_node<'o>(o: &'o Operand, stack: &mut Vec<&'o Operand>) -> usize { + match o { + Operand::Name { bytes, .. } | Operand::Str { bytes, .. } => bytes.capacity(), + Operand::Array { items, .. } => { + stack.extend(items.iter()); + items.capacity() * size_of::() + } + Operand::Dict { entries, .. } => { + let keys: usize = entries.iter().map(|(k, _)| k.capacity()).sum(); + stack.extend(entries.iter().map(|(_, v)| v)); + keys + entries.capacity() * size_of::<(Vec, Operand)>() + } + Operand::Number { .. } | Operand::Bool { .. } | Operand::Null { .. } => 0, + } +} + +/// Lexed operand nodes of `ops`: every number, name, string, array and dictionary, nested ones +/// and an inline image's dictionary values included (counted without recursion). +pub(crate) fn operand_nodes(ops: &[Op]) -> usize { + let mut stack: Vec<&Operand> = Vec::new(); + let mut n = 0usize; + for op in ops { + stack.extend(op.operands.iter()); + stack.extend( + op.inline_image + .iter() + .flat_map(|i| i.dict.iter().map(|(_, v)| v)), + ); + while let Some(o) = stack.pop() { + n = n.saturating_add(1); + match o { + Operand::Array { items, .. } => stack.extend(items.iter()), + Operand::Dict { entries, .. } => stack.extend(entries.iter().map(|(_, v)| v)), + _ => {} + } + } + } + n +} + +/// One stream's lexed ops, their operand nodes and the bytes charged for them. +pub(crate) struct Lexed { + pub ops: Vec, + pub nodes: usize, + pub bytes: usize, +} + +/// Lexes one stream under the walk's budgets. The lexer stops at the operand nodes the walk may +/// still hold — `OPERAND_NODES_MAX` alive at once (`alive` are already held) and what the byte +/// budget has left, at `size_of::()` each — and the ops' bytes are charged before they +/// are kept (`keep`: the page's ops, held) or run (a Form's, scratch the caller releases). Lexer +/// errors keep their reason, with `prefix` before the detail; the node cap reads "content +/// operands", or "page model size" when the byte budget was the tighter bound. +#[allow(clippy::too_many_arguments)] +pub(crate) fn lex_ops( + bytes: &[u8], + ops_max: usize, + alive: usize, + mem: &mut ModelBudget, + cancel: Option<&AtomicBool>, + keep: bool, + prefix: &str, +) -> Result { + let by_count = OPERAND_NODES_MAX.saturating_sub(alive); + let by_bytes = mem.left() / size_of::(); + let limits = LexLimits { + ops_max, + operand_nodes_max: by_count.min(by_bytes), + ..LexLimits::page() + }; + let too_many = |by_budget: bool| { + let what = if by_budget { + MODEL_SIZE + } else { + "content operands" + }; + Stop::new(TextReason::PageTooComplex, what) + }; + let ops = lex_content(bytes, &limits, cancel).map_err(|e| match e { + LexError::TooComplex { what } if what == OPERAND_NODES => too_many(by_bytes < by_count), + e => Stop::new(e.page_reason(), format!("{prefix}{e}")), + })?; + let nodes = operand_nodes(&ops); + if nodes > by_count { + return Err(too_many(false)); + } + let size = ops_bytes(&ops); + if keep { + mem.hold(size)?; + } else { + mem.scratch(size)?; + } + Ok(Lexed { + ops, + nodes, + bytes: size, + }) +} + +/// The byte budget of one page model build or one Classify pass (see the module comment). +pub(crate) struct ModelBudget { + max: usize, + held: usize, + scratch: usize, + shared: Shared, + /// The part of `held` that is font models. + fonts: usize, +} + +impl ModelBudget { + /// A budget of `model_bytes_max()` with `held` bytes already charged (the content, for a new + /// walk; the walk's, for the run stage; the model's, for its Classify pass). + pub(crate) fn new(held: usize) -> ModelBudget { + ModelBudget { + max: model_bytes_max(), + held, + scratch: 0, + shared: Shared::default(), + fonts: 0, + } + } + + /// Notes font models charged already (held by the model a Classify pass is for, alive until + /// it ends): loading one of them again costs nothing more. + pub(crate) fn fonts_held<'f>(&mut self, fonts: impl Iterator>) { + for font in fonts { + self.shared.seen.insert(addr(font)); + } + } + + /// Font-model bytes charged by this budget. + pub(crate) fn font_bytes(&self) -> usize { + self.fonts + } + + /// Charges the font model `font` (held), once per allocation. + pub(crate) fn font(&mut self, font: &Arc) -> Result<(), Stop> { + if self.shared.seen.contains(&addr(font)) { + return Ok(()); + } + let bytes = font.approx_bytes(); + let entry = map_entry::(); + self.check(bytes.saturating_add(entry))?; + self.hold(bytes)?; + self.scratch(entry)?; + self.shared.seen.insert(addr(font)); + self.fonts = self.fonts.saturating_add(bytes); + Ok(()) + } + + /// `font` for a model the page can do without: charged when it fits `allowance` and the + /// budget (or was charged already); false when it is to be left out. + pub(crate) fn font_within(&mut self, font: &Arc, allowance: usize) -> bool { + if self.shared.seen.contains(&addr(font)) { + return true; + } + font.approx_bytes() <= allowance && self.font(font).is_ok() + } + + /// Bytes the model keeps so far. + pub(crate) fn held(&self) -> usize { + self.held + } + + /// Bytes that may still be charged. + pub(crate) fn left(&self) -> usize { + self.max + .saturating_sub(self.held.saturating_add(self.scratch)) + } + + fn check(&self, more: usize) -> Result<(), Stop> { + if more > self.left() { + return Err(Stop::new(TextReason::PageTooComplex, MODEL_SIZE)); + } + Ok(()) + } + + /// Charges `bytes` the model will keep. + pub(crate) fn hold(&mut self, bytes: usize) -> Result<(), Stop> { + self.check(bytes)?; + self.held = self.held.saturating_add(bytes); + Ok(()) + } + + /// Releases held bytes that were charged but not kept (spare capacity given back). + pub(crate) fn unhold(&mut self, bytes: usize) { + self.held = self.held.saturating_sub(bytes); + } + + /// Charges `bytes` that live only while this walk or run stage runs. + pub(crate) fn scratch(&mut self, bytes: usize) -> Result<(), Stop> { + self.check(bytes)?; + self.scratch = self.scratch.saturating_add(bytes); + Ok(()) + } + + /// Releases scratch bytes that were freed. + pub(crate) fn unscratch(&mut self, bytes: usize) { + self.scratch = self.scratch.saturating_sub(bytes); + } + + /// Charges what the digest `d` adds to the model: itself and its shared parts, each once, + /// plus one dedupe entry per new allocation (scratch). + pub(crate) fn digest(&mut self, d: &Arc) -> Result<(), Stop> { + let (before, entries) = (self.shared.bytes(), self.shared.entries()); + self.shared.add(d); + let new = self.shared.bytes().saturating_sub(before); + let seen = self.shared.entries().saturating_sub(entries); + self.hold(new)?; + self.scratch(seen.saturating_mul(map_entry::())) + } + + /// Makes room for `more` items in `v`, charging the growth (held) before it is allocated. + /// Grows by doubling like `Vec::push`, so a vector costs what an unbudgeted one would. The old + /// buffer lives until its items are moved, so the budget must also have room for it then + /// (review T3-budget LOW-1: a doubling of the records vector briefly holds both). + pub(crate) fn grow(&mut self, v: &mut Vec, more: usize) -> Result<(), Stop> { + let need = v.len().saturating_add(more); + let cap = v.capacity(); + if need <= cap { + return Ok(()); + } + let new_cap = need.max(cap.saturating_mul(2)).max(4); + let extra = new_cap.saturating_sub(cap).saturating_mul(size_of::()); + self.check(extra.saturating_add(cap.saturating_mul(size_of::())))?; + self.hold(extra)?; + v.reserve_exact(new_cap.saturating_sub(v.len())); + Ok(()) + } + + /// Pushes `x` onto `v` after charging any growth. + pub(crate) fn push(&mut self, v: &mut Vec, x: T) -> Result<(), Stop> { + self.grow(v, 1)?; + v.push(x); + Ok(()) + } + + /// Gives `v`'s spare capacity back (held) once it is complete. + pub(crate) fn fit(&mut self, v: &mut Vec) { + let spare = v.capacity().saturating_sub(v.len()); + v.shrink_to_fit(); + self.unhold(spare.saturating_mul(size_of::())); + } +} + +#[cfg(test)] +impl Shared { + /// Counts a shared allocation of `bytes` at `a` once (`PageModel::approx_bytes`). + pub(crate) fn arc(&mut self, a: &Arc, bytes: usize) -> bool { + self.once(addr(a), bytes) + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/walker/fonts.rs b/src-tauri/src/pdf_engine/text_edit/walker/fonts.rs new file mode 100644 index 0000000..027e677 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/walker/fonts.rs @@ -0,0 +1,103 @@ +//! The walk's fonts (SPEC §B.9, §B.10): loaded once per walk through the snapshot's +//! `FontCache`, charged to the page-model budget once per model (`budget::ModelBudget::font`: +//! a page model keeps its fonts alive in its records and page fonts however the cache evicts, +//! review T3-budget HIGH-1), and the depth-0 page fonts in first-use order, then the unused ones. + +use super::budget::map_entry; +use super::{show, Stop, Walked, Walker}; +use crate::pdf_engine::text_edit::decode::DecodeBudget; +use crate::pdf_engine::text_edit::fonts::{FontKey, FontModel}; +use crate::pdf_engine::text_edit::limits::{ + FONTS_PER_PAGE_MAX, PAGE_DECODE_BUDGET, PAGE_UNUSED_FONT_BYTES_MAX, +}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use lopdf::Dictionary; +use std::mem::size_of; +use std::sync::Arc; + +/// The unused page fonts may take at most this share of what the walk left (1 / n), so the run +/// stage keeps room. +const UNUSED_FONT_SHARE: usize = 4; + +impl<'a> Walker<'a> { + /// The font of `key`, loaded once per walk and charged to the page budget (held: the model + /// keeps it in its records and page fonts however the snapshot's `FontCache` evicts). A load + /// that ran out of the decode budget, or a model that does not fit the page budget, refuses + /// the page `PAGE_TOO_COMPLEX` (review T3-budget HIGH-1). + pub(crate) fn load_font( + &mut self, + key: FontKey, + dict: &'a Dictionary, + ) -> Result, Stop> { + if let Some(model) = self.fonts.get(&key) { + return Ok(Arc::clone(model)); + } + if self.fonts.len() >= FONTS_PER_PAGE_MAX { + return Err(Stop::new(TextReason::PageTooComplex, "fonts per page")); + } + self.mem + .scratch(map_entry::>() + map_entry::())?; + let model = self + .ctx + .fonts + .get_or_load(self.doc, key, dict, &mut self.budget); + if model.refusal == Some(TextReason::PageTooComplex) { + return Err(Stop::new(TextReason::PageTooComplex, "font decode budget")); + } + self.mem.font(&model)?; + self.pen_proofs + .insert(key, show::advance_proven(self.doc, dict, &model)); + self.fonts.insert(key, Arc::clone(&model)); + Ok(model) + } + + /// Records a font of the depth-0 resources in first-use order (its name kept twice: in the + /// page fonts and in the walk's name set). + pub(crate) fn note_root_font(&mut self, name: &[u8], model: &Arc) -> Walked { + if !self.root_font_names.contains(name) { + self.mem.hold(name.len())?; + self.mem.scratch(name.len() + map_entry::, ()>())?; + self.mem.grow(&mut self.root_fonts, 1)?; + self.root_font_names.insert(name.to_vec()); + self.root_fonts.push((name.to_vec(), Arc::clone(model))); + } + Ok(()) + } + + /// Appends the unused fonts of the depth-0 resources (siblings for faces and joins). They are + /// loaded under a decode budget of their own (`PAGE_DECODE_BUDGET`, shared by all of them), + /// and their models are charged to the page budget within an allowance + /// (`PAGE_UNUSED_FONT_BYTES_MAX`, and `1 / UNUSED_FONT_SHARE` of what the walk left): a font that does not fit is left out — one less sibling to type with — instead of + /// refusing the page, so a document-wide `/Resources` full of large programs never refuses + /// the text a page actually draws. + pub(super) fn finish_page_fonts(&mut self) -> Walked { + let res = self.root_res; + if res.font_count(self.doc) > FONTS_PER_PAGE_MAX { + return Err(Stop::new(TextReason::PageTooComplex, "fonts per page")); + } + // Every name is copied once more before the used ones are skipped. + let names = res.font_name_bytes(self.doc); + self.mem + .hold(names.saturating_add(res.font_count(self.doc) * size_of::()))?; + let mut budget = DecodeBudget::new(PAGE_DECODE_BUDGET); + let allowance = PAGE_UNUSED_FONT_BYTES_MAX.min(self.mem.left() / UNUSED_FONT_SHARE); + let charged = self.mem.font_bytes(); + for (name, key, dict) in res.all_fonts(self.doc) { + if self.root_font_names.contains(&name) { + continue; + } + let model = match self.fonts.get(&key) { + Some(model) => Arc::clone(model), + None => self.ctx.fonts.get_or_load(self.doc, key, dict, &mut budget), + }; + let left = allowance.saturating_sub(self.mem.font_bytes().saturating_sub(charged)); + if model.refusal == Some(TextReason::PageTooComplex) + || !self.mem.font_within(&model, left) + { + continue; + } + self.mem.push(&mut self.root_fonts, (name, model))?; + } + Ok(()) + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/walker/gfx.rs b/src-tauri/src/pdf_engine/text_edit/walker/gfx.rs new file mode 100644 index 0000000..e323016 --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/walker/gfx.rs @@ -0,0 +1,594 @@ +//! Graphics-state operators of the walker: `cm`, line parameters, ExtGState, paths and clips, +//! colour, marked content (SPEC §B.10 "operators modelled"). +//! +//! A `gs` costs O(its dictionary): ExtGState dictionaries hold at most `EXTGSTATE_KEYS_MAX` keys +//! and at most `GS_OTHER_MAX` unmodelled keys may be in force at once (real producers use a +//! handful of each), so a page that repeats `gs` cannot multiply a large dictionary. +//! +//! A state operator that sets a value already in force keeps the shared allocation in force +//! (`set_bytes`, `set_dash`, `GsEffects::set_others`), so re-applying one ExtGState, `ri` or `d` +//! per op leaves the digest shared (`Exec::digest`) instead of one new digest per op. Unmodelled +//! ExtGState keys are at most `EXTGSTATE_KEY_BYTES_MAX` long, and re-applying the dictionary +//! that put the keys in force costs O(1) (`Walker::gs_last`), so a `gs` never compares long keys +//! over and over (review T3 r3 MEDIUM-2). The key lists a walk keeps are charged to its budget. + +use super::budget::{arc_slice, map_entry}; +use super::{Exec, Frame, Others, PaintKind, PaintRecord, Stop, Walked, Walker}; +use crate::pdf_engine::text_edit::context::{resolve, Lookup}; +use crate::pdf_engine::text_edit::fonts::{number_of, FontKey}; +use crate::pdf_engine::text_edit::geometry::{ + apply, axis_aligned, bbox_of, intersect, is_finite, mul, Matrix, +}; +use crate::pdf_engine::text_edit::lexer::{Op, Operand, Operator}; +use crate::pdf_engine::text_edit::limits::{EXTGSTATE_KEY_BYTES_MAX, MARKED_DEPTH_MAX}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::state::{ + color_effect, Bytes, ClipState, ColorSpaceKind, FontUse, GsEffects, MarkedEntry, OcState, Paint, +}; +use lopdf::{Dictionary, Object, ObjectId}; +use std::mem::size_of; +use std::sync::Arc; + +/// Keys of one ExtGState dictionary (ISO 32000-2 defines 29). +const EXTGSTATE_KEYS_MAX: usize = 64; +/// Unmodelled ExtGState keys in force at once (`GsEffects::other`). +const GS_OTHER_MAX: usize = 64; +/// Entries of an ExtGState `/D` dash array that are modelled; a longer one is recorded through +/// its deep hash like an unmodelled key. +const GS_DASH_ITEMS_MAX: usize = 64; + +pub(super) fn num(op: &Op, i: usize) -> Result { + op.operands + .get(i) + .and_then(Operand::as_number) + .filter(|v| v.is_finite()) + .ok_or_else(|| Stop::malformed(format!("operand at byte {}", op.op_span.start))) +} + +pub(super) fn six(op: &Op) -> Result { + Ok([ + num(op, 0)?, + num(op, 1)?, + num(op, 2)?, + num(op, 3)?, + num(op, 4)?, + num(op, 5)?, + ]) +} + +pub(super) fn finite(m: &Matrix, op: &Op) -> Walked { + if is_finite(m) { + Ok(()) + } else { + Err(Stop::malformed(format!( + "non-finite matrix at byte {}", + op.op_span.start + ))) + } +} + +pub(super) fn concat(ex: &mut Exec, op: &Op) -> Walked { + let m = six(op)?; + let ctm = mul(&m, &ex.gs.ctm); + finite(&ctm, op)?; + ex.gs.ctm = ctm; + Ok(()) +} + +pub(super) fn line_param(ex: &mut Exec, op: &Op) -> Walked { + let g = &mut ex.gs.gs; + match op.operator { + Operator::w => g.line_width = num(op, 0)?, + Operator::J => g.line_cap = num(op, 0)? as i64, + Operator::j => g.line_join = num(op, 0)? as i64, + Operator::M => g.miter_limit = num(op, 0)?, + Operator::i => g.flatness = num(op, 0)?, + Operator::ri => set_bytes( + &mut g.rendering_intent, + op.operands + .first() + .and_then(Operand::as_name) + .unwrap_or_default(), + ), + Operator::d => { + let dashes = match op.operands.first() { + Some(Operand::Array { items, .. }) => items + .iter() + .map(|i| i.as_number().filter(|v| v.is_finite())) + .collect::>>() + .ok_or_else(|| Stop::malformed("dash array"))?, + _ => return Err(Stop::malformed("dash array")), + }; + set_dash(&mut g.dash, &dashes, num(op, 1)?); + } + _ => {} + } + Ok(()) +} + +/// Sets shared bytes, keeping the allocation in force when the value is the same. +fn set_bytes(slot: &mut Bytes, value: &[u8]) { + if **slot != *value { + *slot = Arc::from(value); + } +} + +/// Sets the dash pattern, keeping the shared array in force when it is the same. +fn set_dash(slot: &mut (Arc<[f64]>, f64), dashes: &[f64], phase: f64) { + if *slot.0 != *dashes { + slot.0 = Arc::from(dashes); + } + slot.1 = phase; +} + +/// Path construction. Only the points an op adds are checked and only the path's bounding box is +/// kept, so a long unpainted path costs O(1) per op and no memory (it is cleared at its painting +/// op). +pub(super) fn path(ex: &mut Exec, op: &Op) -> Walked { + let ctm = ex.gs.ctm; + let pt = |i: usize| -> Result<(f64, f64), Stop> { + finite_point(apply(&ctm, num(op, i)?, num(op, i + 1)?), op) + }; + let p = &mut ex.path; + let add = |bbox: &mut Option<[f64; 4]>, points: &[(f64, f64)]| { + for &(x, y) in points { + *bbox = Some(match *bbox { + None => [x, y, x, y], + Some(b) => [b[0].min(x), b[1].min(y), b[2].max(x), b[3].max(y)], + }); + } + }; + match op.operator { + Operator::m | Operator::l => { + add(&mut p.bbox, &[pt(0)?]); + p.other = true; + } + Operator::c => { + add(&mut p.bbox, &[pt(0)?, pt(2)?, pt(4)?]); + p.other = true; + } + Operator::v | Operator::y => { + add(&mut p.bbox, &[pt(0)?, pt(2)?]); + p.other = true; + } + Operator::re => { + let (x, y, w, h) = (num(op, 0)?, num(op, 1)?, num(op, 2)?, num(op, 3)?); + let corners = [ + finite_point(apply(&ctm, x, y), op)?, + finite_point(apply(&ctm, x + w, y), op)?, + finite_point(apply(&ctm, x, y + h), op)?, + finite_point(apply(&ctm, x + w, y + h), op)?, + ]; + add(&mut p.bbox, &corners); + p.re_count = p.re_count.saturating_add(1); + p.re_axis = if p.re_count == 1 { + axis_aligned(&ctm) + } else { + p.re_axis && axis_aligned(&ctm) + }; + p.re_rect = bbox_of(&corners); + } + _ => {} + } + Ok(()) +} + +fn finite_point(p: (f64, f64), op: &Op) -> Result<(f64, f64), Stop> { + if p.0.is_finite() && p.1.is_finite() { + Ok(p) + } else { + Err(Stop::malformed(format!( + "non-finite path at byte {}", + op.op_span.start + ))) + } +} + +impl<'a> Walker<'a> { + pub(super) fn paint_path(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + if op.operator != Operator::n { + let seq = self.next_seq()?; + let state = self.digest(ex)?; + self.mem.hold(frame.chain.len() * size_of::())?; + let paint = PaintRecord { + seq, + depth: frame.depth, + kind: PaintKind::Path(op.operator), + span: frame.joined.then(|| op.span.clone()), + state, + bbox: ex.path.bbox, + masked: ex.gs.gs.soft_mask, + form_chain: frame.chain.clone(), + local_span: op.span.clone(), + xobject: None, + }; + self.mem.push(&mut self.paints, paint)?; + } + if ex.clip_pending { + let p = &ex.path; + let rect = (!p.other && p.re_count == 1 && p.re_axis) + .then_some(p.re_rect) + .flatten(); + ex.gs.clip = match (&ex.gs.clip, rect) { + (ClipState::Complex, _) | (_, None) => ClipState::Complex, + (ClipState::None, Some(r)) => ClipState::Rect(r), + (ClipState::Rect(c), Some(r)) => ClipState::Rect(intersect(*c, r)), + }; + ex.clip_pending = false; + } + ex.path = Default::default(); + Ok(()) + } + + pub(super) fn ext_gstate(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + let name = op + .operands + .first() + .and_then(Operand::as_name) + .unwrap_or_default(); + let dict = match frame.res.entry(self.doc, b"ExtGState", name) { + Lookup::Found(e) => match e.value { + Object::Dictionary(d) => d, + _ => return Err(Stop::malformed("ExtGState is not a dictionary")), + }, + Lookup::Missing => return Err(Stop::malformed("ExtGState name not in resources")), + Lookup::Broken(w) => return Err(Stop::malformed(format!("ExtGState {w}"))), + }; + if dict.len() > EXTGSTATE_KEYS_MAX { + return Err(Stop::new(TextReason::PageTooComplex, "ExtGState keys")); + } + // The unmodelled keys depend on the dictionary alone: listed (and hashed) once per walk. + let addr = dict as *const Dictionary as usize; + let known = self.gs_others.get(&addr).cloned(); + let mut others: Vec<(Bytes, u64)> = Vec::new(); + for (key, raw) in dict.iter() { + let Some((id, value)) = resolve(self.doc, raw) else { + return Err(Stop::malformed("ExtGState reference")); + }; + if key.as_slice() == b"Font" { + self.gs_font(ex, value)?; + continue; + } + let modelled = apply_gs_key(&mut ex.gs.gs, key, value, dict); + let smask = key.as_slice() == b"SMask" && ex.gs.gs.soft_mask; + if known.is_none() && ((!modelled && key.as_slice() != b"Type") || smask) { + if key.len() > EXTGSTATE_KEY_BYTES_MAX { + return Err(Stop::new(TextReason::PageTooComplex, "ExtGState keys")); + } + self.mem.scratch(arc_slice(key.len(), 1))?; + others.push((Arc::from(key.as_slice()), self.entry_hash(id, value)?)); + } + } + let others = match known { + Some(list) => list, + None => { + let pair = size_of::<(Bytes, u64)>(); + self.mem + .scratch(arc_slice(others.len(), pair) + map_entry::())?; + let list = GsEffects::sorted_others(others); + self.gs_others.insert(addr, Arc::clone(&list)); + list + } + }; + self.set_others(ex, others)?; + if ex.gs.gs.other.len() > GS_OTHER_MAX { + return Err(Stop::new(TextReason::PageTooComplex, "ExtGState keys")); + } + Ok(()) + } + + /// Puts the unmodelled keys `others` of one ExtGState in force (`GsEffects::set_others`). The + /// dictionary applied last, re-applied to the list it put in force, changes nothing and is + /// skipped in O(1); a new merged list is charged (scratch: the walk interns it) before the + /// merge allocates it. + fn set_others(&mut self, ex: &mut Exec, others: Others) -> Walked { + if let Some((entries, result)) = &self.gs_last { + if Arc::ptr_eq(entries, &others) && Arc::ptr_eq(result, &ex.gs.gs.other) { + return Ok(()); + } + } + let n = ex.gs.gs.other.len().saturating_add(others.len()); + let pair = size_of::<(Bytes, u64)>(); + let merge = n.saturating_mul(pair); + let list = arc_slice(n, pair) + map_entry::(); + self.mem.scratch(merge.saturating_add(list))?; + let interned = ex.gs.gs.set_others(&others, &mut self.others_interned); + self.mem + .unscratch(if interned { merge } else { merge + list }); + self.gs_last = Some((others, Arc::clone(&ex.gs.gs.other))); + Ok(()) + } + + /// ExtGState `/Font [font size]`: the font in force without a `Tf` (B9). The font must be an + /// indirect reference (ISO 32000-1 Table 58); viewers disagree on a direct dictionary there + /// (poppler draws nothing), so it is malformed. + fn gs_font(&mut self, ex: &mut Exec, value: &'a Object) -> Walked { + let bad = || Stop::malformed("ExtGState /Font"); + let Object::Array(items) = value else { + return Err(bad()); + }; + let (Some(font), Some(size)) = (items.first(), items.get(1)) else { + return Err(bad()); + }; + let size = resolve(self.doc, size) + .and_then(|(_, o)| number_of(o)) + .ok_or_else(bad)?; + let (key, dict) = match resolve(self.doc, font) { + Some((Some(id), Object::Dictionary(d))) => (FontKey::Indirect(id), d), + Some((None, Object::Dictionary(_))) => { + return Err(Stop::malformed( + "ExtGState /Font is not an indirect reference", + )) + } + _ => return Err(bad()), + }; + let model = self.load_font(key, dict)?; + ex.gs.text.font = Some(FontUse { + resource: None, + content_hash: model.content_hash, + from_extgstate: true, + tf_op: None, + }); + ex.gs.text.tfs = size; + ex.gs.font_ref = Some((key, model)); + Ok(()) + } + + pub(super) fn color(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + use Operator as O; + let bytes = frame.op_bytes(op); + let stroke = matches!(op.operator, O::CS | O::SC | O::SCN | O::G | O::RG | O::K); + let device = |space: ColorSpaceKind, n: usize| -> Result { + let comps = (0..n) + .map(|i| num(op, i)) + .collect::, Stop>>()?; + Ok(Paint { + space_op: None, + color_op: Some(Arc::clone(&bytes)), + effect: color_effect(&space, &comps), + space, + comps, + pattern: false, + pattern_hash: None, + }) + }; + let new = match op.operator { + O::G | O::g => device(ColorSpaceKind::DeviceGray, 1)?, + O::RG | O::rg => device(ColorSpaceKind::DeviceRgb, 3)?, + O::K | O::k => device(ColorSpaceKind::DeviceCmyk, 4)?, + O::CS | O::cs => { + let name = op + .operands + .first() + .and_then(Operand::as_name) + .unwrap_or_default(); + let (space, comps) = self.color_space(frame, name)?; + Paint { + space_op: Some(bytes), + color_op: None, + effect: color_effect(&space, &comps), + pattern: space == ColorSpaceKind::Pattern, + space, + comps, + pattern_hash: None, + } + } + _ => { + let pattern_name = op.operands.last().and_then(Operand::as_name); + let pattern_hash = match pattern_name { + Some(name) => match frame.res.entry(self.doc, b"Pattern", name) { + Lookup::Found(e) => Some(self.entry_hash(e.id, e.value)?), + Lookup::Missing => return Err(Stop::malformed("pattern not in resources")), + Lookup::Broken(w) => return Err(Stop::malformed(format!("pattern {w}"))), + }, + None => None, + }; + let cur = if stroke { &ex.gs.stroke } else { &ex.gs.fill }; + let comps = op + .operands + .iter() + .filter_map(Operand::as_number) + .collect::>(); + if comps.iter().any(|c| !c.is_finite()) { + return Err(Stop::malformed("colour component")); + } + let space = match &cur.space { + ColorSpaceKind::Default => ColorSpaceKind::DeviceGray, + other => other.clone(), + }; + Paint { + space_op: cur.space_op.clone(), + color_op: Some(bytes), + effect: color_effect(&space, &comps), + pattern: cur.pattern || pattern_name.is_some(), + space, + comps, + pattern_hash, + } + } + }; + if stroke { + ex.gs.stroke = new; + } else { + ex.gs.fill = new; + } + Ok(()) + } + + /// `cs`/`CS` operand: a device family, `/Pattern`, or a `/ColorSpace` resource. + fn color_space( + &mut self, + frame: &Frame<'_, 'a>, + name: &[u8], + ) -> Result<(ColorSpaceKind, Vec), Stop> { + Ok(match name { + b"DeviceGray" | b"G" => (ColorSpaceKind::DeviceGray, vec![0.0]), + b"DeviceRGB" | b"RGB" => (ColorSpaceKind::DeviceRgb, vec![0.0; 3]), + b"DeviceCMYK" | b"CMYK" => (ColorSpaceKind::DeviceCmyk, vec![0.0, 0.0, 0.0, 1.0]), + b"Pattern" => (ColorSpaceKind::Pattern, Vec::new()), + _ => match frame.res.entry(self.doc, b"ColorSpace", name) { + Lookup::Found(e) => { + if names_pattern(self.doc, e.value) { + (ColorSpaceKind::Pattern, Vec::new()) + } else { + let h = self.entry_hash(e.id, e.value)?; + (ColorSpaceKind::Named(Arc::from(name), h), Vec::new()) + } + } + Lookup::Missing => return Err(Stop::malformed("colour space not in resources")), + Lookup::Broken(w) => return Err(Stop::malformed(format!("colour space {w}"))), + }, + }) + } + + pub(super) fn marked(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + use Operator as O; + let tag = op + .operands + .first() + .and_then(Operand::as_name) + .unwrap_or_default(); + match op.operator { + O::EMC => { + if ex.marked.len() > ex.marked_base { + ex.marked.pop(); + } + return Ok(()); + } + O::MP => return Ok(()), + O::DP => { + if let Some(Operand::Name { bytes, .. }) = op.operands.get(1) { + self.properties(frame, bytes)?; + } + return Ok(()); + } + _ => {} + } + if ex.marked.len() >= MARKED_DEPTH_MAX { + return Err(Stop::new( + TextReason::PageTooComplex, + "marked-content nesting", + )); + } + let is_oc = tag == b"OC"; + let mut entry = MarkedEntry { + tag: Arc::from(tag), + mcid: None, + actual_text: false, + oc: None, + }; + match op.operands.get(1) { + Some(Operand::Dict { entries, .. }) => { + for (k, v) in entries { + match k.as_slice() { + b"MCID" => entry.mcid = v.as_number().map(|n| n as i64), + b"ActualText" | b"E" => entry.actual_text = true, + _ => {} + } + } + if is_oc { + entry.oc = Some(OcState::Unknown); + } + } + Some(Operand::Name { bytes, .. }) => { + let (id, value) = self.properties(frame, bytes)?; + let dict = match value { + Object::Dictionary(d) => Some(d), + _ => None, + }; + if let Some(d) = dict { + entry.mcid = d + .get(b"MCID") + .ok() + .and_then(|o| match resolve(self.doc, o) { + Some((_, Object::Integer(i))) => Some(*i), + _ => None, + }); + entry.actual_text = d.has(b"ActualText") || d.has(b"E"); + } + if is_oc { + entry.oc = Some(self.oc_state(id, value)?); + } + } + _ => {} + } + ex.marked.push(entry); + Ok(()) + } + + /// `/Properties` entry `name`: the id it was reached through (when indirect) and its value. + fn properties( + &self, + frame: &Frame<'_, 'a>, + name: &[u8], + ) -> Result<(Option, &'a Object), Stop> { + match frame.res.entry(self.doc, b"Properties", name) { + Lookup::Found(e) => Ok((e.id, e.value)), + Lookup::Missing => Err(Stop::malformed("properties name not in resources")), + Lookup::Broken(w) => Err(Stop::malformed(format!("properties {w}"))), + } + } +} + +/// Applies a modelled ExtGState key; `false` when the key is not modelled (or its value has the +/// wrong type, in which case it is recorded verbatim through `other`). +fn apply_gs_key(g: &mut GsEffects, key: &[u8], value: &Object, dict: &Dictionary) -> bool { + let n = number_of(value); + let b = value.as_bool().ok(); + let name = value.as_name().ok(); + match (key, n, b, name) { + (b"CA", Some(v), _, _) => g.ca_stroke = v, + (b"ca", Some(v), _, _) => g.ca = v, + (b"BM", _, _, Some(bm)) => set_bytes(&mut g.blend, bm), + (b"BM", _, _, None) => match value.as_array().ok().and_then(|a| a.first()) { + Some(Object::Name(bm)) => set_bytes(&mut g.blend, bm), + _ => return false, + }, + (b"SMask", _, _, _) => g.soft_mask = name != Some(b"None"), + (b"OP", _, Some(v), _) => { + g.overprint.0 = v; + if !dict.has(b"op") { + g.overprint.1 = v; + } + } + (b"op", _, Some(v), _) => g.overprint.1 = v, + (b"OPM", Some(v), _, _) => g.overprint.2 = v as i64, + (b"LW", Some(v), _, _) => g.line_width = v, + (b"LC", Some(v), _, _) => g.line_cap = v as i64, + (b"LJ", Some(v), _, _) => g.line_join = v as i64, + (b"ML", Some(v), _, _) => g.miter_limit = v, + (b"RI", _, _, Some(ri)) => set_bytes(&mut g.rendering_intent, ri), + (b"FL", Some(v), _, _) => g.flatness = v, + (b"SA", _, Some(v), _) => g.stroke_adjust = v, + (b"D", _, _, _) => match dash_of(value) { + Some((dashes, phase)) => set_dash(&mut g.dash, &dashes, phase), + None => return false, + }, + _ => return false, + } + true +} + +fn dash_of(value: &Object) -> Option<(Vec, f64)> { + let items = value.as_array().ok()?; + let (Some(Object::Array(arr)), Some(phase)) = (items.first(), items.get(1)) else { + return None; + }; + if arr.len() > GS_DASH_ITEMS_MAX { + return None; + } + let dashes = arr.iter().map(number_of).collect::>>()?; + Some((dashes, number_of(phase)?)) +} + +/// A colour space value that is `/Pattern` or `[/Pattern …]`. +fn names_pattern(doc: &lopdf::Document, value: &Object) -> bool { + match value { + Object::Name(n) => n.as_slice() == b"Pattern", + Object::Array(items) => items + .first() + .and_then(|f| resolve(doc, f)) + .is_some_and(|(_, o)| o.as_name().ok() == Some(b"Pattern")), + _ => false, + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/walker/hash.rs b/src-tauri/src/pdf_engine/text_edit/walker/hash.rs new file mode 100644 index 0000000..83785be --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/walker/hash.rs @@ -0,0 +1,306 @@ +//! Deep, object-id-free hashes of XObjects, colour spaces, shadings and ExtGState values (SPEC +//! §B.10 `PaintKind`, D32): dictionaries with sorted keys, references followed (cycles marked by +//! their position on the path, dangling references hashed as `null` like qpdf's +//! `fixDanglingReferences`), streams by their decoded bytes under the walk's hash budget, or by +//! raw bytes + `/Filter` + `/DecodeParms` when they cannot (or no longer can) be decoded. +//! +//! The decode/raw decision depends only on the decoded size and the budget left, which are the +//! same in the source and in qpdf's output (where unfiltered streams come back Flate-compressed), +//! so equal content hashes equal on both sides. `/Parent`, `/PieceInfo` and `/Metadata` are +//! skipped (they never change what is painted and can reach the whole document). +//! +//! Bounded per walk, whatever the page repeats: every hash is memoised (indirect objects by id, +//! direct values by address), all hashes of a walk share one step count (`HASH_STEPS_MAX`), the +//! decode budget is charged for a failed decode as if it ran to its cap, and raw bytes are +//! charged too (`Walker::hash_raw_left`). Exhaustion refuses the page `PAGE_TOO_COMPLEX`. Each +//! memo entry is charged to the page-model budget (`budget::map_entry`); a stream decoded for its +//! hash is not (one at a time, ≤ `STREAM_MAX_DECODED`, freed at once), so the decode/raw decision +//! never depends on what else the page holds. + +use super::budget::map_entry; +use super::{Stop, Walker}; +use crate::pdf_engine::text_edit::decode::{decode_stream, DecodeBudget, DecodeError}; +use crate::pdf_engine::text_edit::limits::{PAGE_DECODE_BUDGET, STREAM_MAX_DECODED}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::snapshot::fnv1a_extend; +use lopdf::{Dictionary, Object, ObjectId, Stream}; + +const FNV_OFFSET: u64 = 0xcbf2_9ce4_8422_2325; +/// Objects visited by all the deep hashes of one walk. +const HASH_STEPS_MAX: usize = 1_000_000; +/// Nesting depth of one hash (references included). +const HASH_DEPTH_MAX: usize = 64; +const SKIPPED_KEYS: [&[u8]; 3] = [b"Parent", b"PieceInfo", b"Metadata"]; +const STREAM_KEYS: [&[u8]; 4] = [b"Length", b"Filter", b"DecodeParms", b"DL"]; + +struct Hasher<'w, 'a> { + w: &'w mut Walker<'a>, + path: Vec, +} + +fn budget_exhausted() -> Stop { + Stop::new(TextReason::PageTooComplex, "object hash budget") +} + +impl<'a> Walker<'a> { + /// Deep hash of a resource value reached through the reference `id` (its object), or of the + /// direct value `obj` itself; memoised per walk either way, so an op that repeats (`sh`, + /// `gs`, `cs`, `scn`) pays for its resource once. + pub(crate) fn entry_hash( + &mut self, + id: Option, + obj: &'a Object, + ) -> Result { + if let Some(id) = id { + return self.stream_hash(id); + } + let address = obj as *const Object as usize; + if let Some(h) = self.direct_hashes.get(&address) { + return Ok(*h); + } + let mut h = Hasher { + w: self, + path: Vec::new(), + }; + let mut out = FNV_OFFSET; + h.object(obj, 0, &mut out)?; + self.mem.scratch(map_entry::())?; + self.direct_hashes.insert(address, out); + Ok(out) + } + + /// Deep hash of the object `id` (memoised per walk). + pub(crate) fn stream_hash(&mut self, id: ObjectId) -> Result { + if let Some(h) = self.hashes.get(&id) { + return Ok(*h); + } + let doc = self.doc; + let h = match doc.objects.get(&id) { + Some(obj) => { + let mut hasher = Hasher { + w: self, + path: vec![id], + }; + let mut out = FNV_OFFSET; + hasher.object(obj, 0, &mut out)?; + out + } + None => fnv1a_extend(FNV_OFFSET, b"n"), + }; + self.mem.scratch(map_entry::())?; + self.hashes.insert(id, h); + Ok(h) + } +} + +impl Walker<'_> { + /// Runs `f` with deep-hash state of its own (fresh budgets, step count and memos) and then + /// restores the walk's: a check outside the page content (the wrapper Form's `/Group` in + /// `Wrapped` mode) leaves the content's hashes exactly as a walk without that check computes + /// them, so a large `/Group` cannot push a content image onto its raw-bytes hash. + pub(crate) fn with_own_hash_state( + &mut self, + f: impl FnOnce(&mut Self) -> Result, + ) -> Result { + let fresh_raw = self.ctx.snap.file_len(); + let saved = ( + std::mem::replace(&mut self.hash_budget, DecodeBudget::new(PAGE_DECODE_BUDGET)), + std::mem::take(&mut self.hash_steps), + std::mem::replace(&mut self.hash_raw_left, fresh_raw), + std::mem::take(&mut self.hashes), + std::mem::take(&mut self.direct_hashes), + ); + let out = f(self); + ( + self.hash_budget, + self.hash_steps, + self.hash_raw_left, + self.hashes, + self.direct_hashes, + ) = saved; + out + } +} + +fn feed(out: &mut u64, bytes: &[u8]) { + *out = fnv1a_extend(*out, bytes); +} + +fn feed_len(out: &mut u64, tag: &[u8], len: usize) { + feed(out, tag); + feed(out, &(len as u64).to_le_bytes()); +} + +impl<'a> Hasher<'_, 'a> { + fn object(&mut self, obj: &'a Object, depth: usize, out: &mut u64) -> Result<(), Stop> { + self.w.hash_steps = self.w.hash_steps.saturating_add(1); + if self.w.hash_steps > steps_max() || depth > HASH_DEPTH_MAX { + return Err(budget_exhausted()); + } + match obj { + Object::Null => feed(out, b"n"), + Object::Boolean(b) => feed(out, if *b { b"bt" } else { b"bf" }), + Object::Integer(i) => { + feed(out, b"i"); + feed(out, &i.to_le_bytes()); + } + Object::Real(r) => { + feed(out, b"r"); + feed(out, &r.to_bits().to_le_bytes()); + } + Object::Name(n) => { + feed_len(out, b"N", n.len()); + feed(out, n); + } + Object::String(s, _) => { + feed_len(out, b"S", s.len()); + feed(out, s); + } + Object::Array(items) => { + feed_len(out, b"A", items.len()); + for item in items { + self.object(item, depth + 1, out)?; + } + } + Object::Dictionary(d) => self.dict(d, &[], depth, out)?, + Object::Stream(s) => self.stream(s, depth, out)?, + Object::Reference(id) => self.reference(*id, depth, out)?, + } + Ok(()) + } + + fn reference(&mut self, id: ObjectId, depth: usize, out: &mut u64) -> Result<(), Stop> { + if let Some(pos) = self.path.iter().position(|p| *p == id) { + feed_len(out, b"C", pos); + return Ok(()); + } + if let Some(h) = self.w.hashes.get(&id) { + feed(out, b"H"); + feed(out, &h.to_le_bytes()); + return Ok(()); + } + let doc = self.w.doc; + let Some(target) = doc.objects.get(&id) else { + feed(out, b"n"); + return Ok(()); + }; + self.path.push(id); + let mut sub = FNV_OFFSET; + self.object(target, depth + 1, &mut sub)?; + self.path.pop(); + self.w.mem.scratch(map_entry::())?; + self.w.hashes.insert(id, sub); + feed(out, b"H"); + feed(out, &sub.to_le_bytes()); + Ok(()) + } + + fn dict( + &mut self, + d: &'a Dictionary, + also_skip: &[&[u8]], + depth: usize, + out: &mut u64, + ) -> Result<(), Stop> { + let mut entries: Vec<(&'a Vec, &'a Object)> = d + .iter() + .filter(|(k, _)| { + !SKIPPED_KEYS.contains(&k.as_slice()) && !also_skip.contains(&k.as_slice()) + }) + .collect(); + entries.sort_by(|a, b| a.0.cmp(b.0)); + feed_len(out, b"D", entries.len()); + for (k, v) in entries { + feed_len(out, b"K", k.len()); + feed(out, k); + self.object(v, depth + 1, out)?; + } + Ok(()) + } + + fn stream(&mut self, s: &'a Stream, depth: usize, out: &mut u64) -> Result<(), Stop> { + feed(out, b"T"); + self.dict(&s.dict, &STREAM_KEYS, depth, out)?; + let budget = &mut self.w.hash_budget; + let cap = STREAM_MAX_DECODED.min(budget.remaining()); + let decoded: Option> = if s.dict.has(b"Filter") { + #[cfg(test)] + tests::note_decode_attempt(cap); + match decode_stream(s, cap, budget) { + Ok(data) => Some(std::borrow::Cow::Owned(data)), + // Rejected before any byte was decoded. + Err(DecodeError::UnsupportedFilter(_)) => None, + // It may have produced up to `cap` bytes before failing: charge them all, so a + // stream that inflates past the cap is not inflated again for free. + Err(DecodeError::TooLarge | DecodeError::Corrupt(_)) => { + *budget = DecodeBudget::new(budget.remaining().saturating_sub(cap)); + None + } + } + } else if s.content.len() <= cap && budget.take(s.content.len()).is_ok() { + Some(std::borrow::Cow::Borrowed(s.content.as_slice())) + } else { + None + }; + match decoded { + Some(data) => { + feed_len(out, b"P", data.len()); + feed(out, &data); + } + None => { + self.w.hash_raw_left = self + .w + .hash_raw_left + .checked_sub(s.content.len()) + .ok_or_else(budget_exhausted)?; + feed(out, b"R"); + for key in [&b"Filter"[..], b"DecodeParms"] { + match s.dict.get(key).ok() { + Some(v) => self.object(v, depth + 1, out)?, + None => feed(out, b"-"), + } + } + feed_len(out, b"B", s.content.len()); + feed(out, &s.content); + } + } + Ok(()) + } +} + +/// `HASH_STEPS_MAX`, or the test override of the calling thread. +fn steps_max() -> usize { + #[cfg(test)] + if let Some(n) = tests::STEPS_OVERRIDE.with(|c| c.get()) { + return n; + } + HASH_STEPS_MAX +} + +#[cfg(test)] +pub(crate) mod tests { + use std::cell::Cell; + + thread_local! { + pub(super) static STEPS_OVERRIDE: Cell> = const { Cell::new(None) }; + static DECODE_ATTEMPTS: Cell<(usize, usize)> = const { Cell::new((0, 0)) }; + } + + /// Test seam: override `HASH_STEPS_MAX` on the calling thread (`None` restores it). + pub(crate) fn set_steps_override(n: Option) { + STEPS_OVERRIDE.with(|c| c.set(n)); + } + + pub(super) fn note_decode_attempt(cap: usize) { + DECODE_ATTEMPTS.with(|c| { + let (n, bytes) = c.get(); + c.set((n.saturating_add(1), bytes.saturating_add(cap))) + }); + } + + /// `(attempts, Σ caps)`: the filtered streams the deep hashes tried to decode on the calling + /// thread since the last call, and the most output those attempts were allowed to produce. + pub(crate) fn take_decode_attempts() -> (usize, usize) { + DECODE_ATTEMPTS.with(|c| c.replace((0, 0))) + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/walker/show.rs b/src-tauri/src/pdf_engine/text_edit/walker/show.rs new file mode 100644 index 0000000..5d4d53a --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/walker/show.rs @@ -0,0 +1,486 @@ +//! Text operators of the walker (SPEC §B.10 glyph math): text state, positioning and the four +//! show ops. `Trm = [Tfs·Th 0 0 Tfs 0 Ts] × Tm × CTM`; each code advances by +//! `(w0·Tfs + Tc + Tw·[1-byte code 32])·Th` (Tc per code, never per byte), and a TJ number `n` by +//! `−n/1000·Tfs·Th`, which also counts in `advance_ts` (B1). +//! +//! The pen is tracked as reliable or not: an op whose advance the model cannot know (an +//! odd-length string in a 2-byte font, no font, a code without a width, or any glyph of a font +//! whose widths are not proven to be the viewers' — `advance_proven`) leaves `pen_unknown` set +//! for every later op of the same text object until a line move (`Td TD Tm T* ' "`) or `BT`. +//! A `Q` inside the text object that restores other text matrices (`Exec::tm_unsettled`) does +//! the same until `Tm` or `BT`. +//! +//! Every glyph, TJ kern and character of glyph text is charged to its page budget as it is +//! made (`KERNS_PER_PAGE_MAX`, `TEXT_CHARS_PER_PAGE_MAX`), never after the op: one TJ array of +//! 48 MiB of strings must not build millions of glyphs before it is refused. Their bytes (and the +//! record's) are charged to the page-model budget the same way (`budget::ModelBudget`). + +use super::gfx::{finite, num, six}; +use super::{ + Exec, Frame, GlyphRec, RecElem, ShowOp, ShowRecord, Stop, Walked, Walker, KERNS_PER_PAGE_MAX, + TEXT_CHARS_PER_PAGE_MAX, +}; +use crate::pdf_engine::text_edit::context::{resolve, Lookup}; +use crate::pdf_engine::text_edit::fonts::{number_of, Code, FontModel}; +use crate::pdf_engine::text_edit::geometry::{ + apply, apply_linear, bbox_of, mul, translate, Matrix, +}; +use crate::pdf_engine::text_edit::lexer::{Op, Operand, Operator, Span}; +use crate::pdf_engine::text_edit::limits::GLYPHS_PER_PAGE_MAX; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::state::{FontUse, StateDigest}; +use lopdf::{Dictionary, Document, Object, ObjectId}; +use std::mem::size_of; +use std::sync::Arc; + +/// Ascent/descent (em) when there is no font to ask. +const FALLBACK_ASCENT: f64 = 0.8; +const FALLBACK_DESCENT: f64 = -0.2; + +/// `Td`, `TD`, `Tm`, `T*` (inside a text object only). +pub(super) fn position(ex: &mut Exec, op: &Op) -> Walked { + if !ex.in_text { + return Err(Stop::malformed(format!( + "text positioning outside BT at byte {}", + op.op_span.start + ))); + } + match op.operator { + Operator::Td => move_line(ex, num(op, 0)?, num(op, 1)?), + Operator::TD => { + let ty = num(op, 1)?; + ex.gs.text.tl = -ty; + move_line(ex, num(op, 0)?, ty); + } + Operator::Tm => { + let m = six(op)?; + ex.tm = m; + ex.tlm = m; + ex.tm_base = m; + ex.tm_unsettled = false; + } + _ => next_line(ex), + } + ex.pen_unknown = false; + finite(&ex.tm, op) +} + +/// Whether the model has no width for `code` in a font that has its codes (a simple-font code +/// whose glyph has no `/Widths` entry and no `/MissingWidth`): viewers then use a width of their +/// own (poppler: the Standard-14 metrics under a Standard-14 name), so the pen after it is not +/// known. 2-byte codes always have one (`/W`, else `/DW`). +pub(crate) fn width_unknown(model: &FontModel, code: Code) -> bool { + code.len == 1 && model.info(code).is_some_and(|i| i.width1000.is_none()) +} + +/// Whether viewers advance the pen by exactly the model's widths (`FontModel::width`) for every +/// code of this font, so the pen after its glyphs stays known. A model keeps only its +/// lowest-numbered refusal (§A.10), so a refusal is trusted only when nothing width-related can +/// hide behind it, and the dictionary is checked where one could: +/// - simple fonts (`Type1`, `TrueType`): no refusal, `NO_TOUNICODE` or `AMBIGUOUS_UNICODE` (the +/// only refusals ranked after `MISSING_WIDTHS`; a model refused earlier may have no codes at +/// all, e.g. a short `/Widths` or an unreadable descriptor, and then counts every advance 0); +/// - `Type0`: horizontal, `/Encoding /Identity-H` (code = CID, widths by `/W` and `/DW`; a +/// predefined CMap such as `UniJIS-UCS2-H` maps codes to other CIDs) and not `FONT_UNSUPPORTED` +/// (a bad `/W`, `/DW`, descendant or descriptor); `FONT_NOT_EMBEDDED` and the program refusals +/// leave `/W` untouched; +/// - `Type3`: `/FontMatrix` of six numbers with `a ≠ 0` and `b = 0` (viewers advance by `w × a` +/// along the baseline), a well-formed `/Widths`, and no `/FontDescriptor /MissingWidth` other +/// than 0 (`type3_no_missing_width`). +pub(crate) fn advance_proven(doc: &Document, dict: &Dictionary, model: &FontModel) -> bool { + use TextReason as R; + if model.vertical { + return false; + } + let name = |key: &[u8]| match dict.get(key).ok().and_then(|o| resolve(doc, o)) { + Some((_, Object::Name(n))) => Some(n.as_slice()), + _ => None, + }; + match name(b"Subtype") { + Some(b"Type1" | b"TrueType") => matches!( + model.refusal, + None | Some(R::NoTounicode | R::AmbiguousUnicode) + ), + Some(b"Type0") => { + name(b"Encoding") == Some(&b"Identity-H"[..]) + && !matches!(model.refusal, Some(R::FontUnsupported | R::Vertical)) + } + Some(b"Type3") => { + type3_advances_along_baseline(doc, dict) + && widths_well_formed(doc, dict) + && type3_no_missing_width(doc, dict) + } + _ => false, + } +} + +/// No `/FontDescriptor`, or one whose `/MissingWidth` is absent or 0. The font model gives a +/// Type3 code outside `/FirstChar…/LastChar` width 0, while poppler and pdf.js give it the +/// descriptor's `/MissingWidth` (and scale it differently: × 0.001 and × `/FontMatrix`). +fn type3_no_missing_width(doc: &Document, dict: &Dictionary) -> bool { + let Ok(raw) = dict.get(b"FontDescriptor") else { + return true; + }; + let Some((_, Object::Dictionary(desc))) = resolve(doc, raw) else { + return false; + }; + match desc.get(b"MissingWidth") { + Err(_) => true, + Ok(mw) => resolve(doc, mw) + .and_then(|(_, o)| number_of(o)) + .is_some_and(|v| v == 0.0), + } +} + +/// Charges `n` to one of the walk's per-page counts; past `max` the page is `PAGE_TOO_COMPLEX`. +fn charge(count: &mut usize, n: usize, max: usize, what: &'static str) -> Walked { + *count = count.saturating_add(n); + if *count > max { + return Err(Stop::new(TextReason::PageTooComplex, what)); + } + Ok(()) +} + +/// `/FontMatrix [a b c d e f]` with `a ≠ 0` (and `1000·a` finite, the scale the font model uses) +/// and `b = 0`: a glyph's displacement `(w, 0)` maps to `(w·a, 0)`. +fn type3_advances_along_baseline(doc: &Document, dict: &Dictionary) -> bool { + let Some((_, Object::Array(items))) = + dict.get(b"FontMatrix").ok().and_then(|o| resolve(doc, o)) + else { + return false; + }; + let values: Option> = items + .iter() + .map(|i| resolve(doc, i).and_then(|(_, o)| number_of(o))) + .collect(); + matches!( + values.as_deref(), + Some([a, b, _, _, _, _]) if *a != 0.0 && (a * 1000.0).is_finite() && *b == 0.0 + ) +} + +/// `/Widths` as the font model reads it: present, `/FirstChar` and `/LastChar` integers 0–255 +/// with `LastChar ≥ FirstChar`, and exactly `LastChar − FirstChar + 1` numbers. +fn widths_well_formed(doc: &Document, dict: &Dictionary) -> bool { + let get = |key: &[u8]| dict.get(key).ok().and_then(|o| resolve(doc, o)); + let int = |key: &[u8]| match get(key) { + Some((_, Object::Integer(v))) if (0..=255).contains(v) => usize::try_from(*v).ok(), + _ => None, + }; + let Some((_, Object::Array(items))) = get(b"Widths") else { + return false; + }; + let (Some(first), Some(last)) = (int(b"FirstChar"), int(b"LastChar")) else { + return false; + }; + last >= first + && last.checked_sub(first).and_then(|n| n.checked_add(1)) == Some(items.len()) + && items + .iter() + .all(|i| resolve(doc, i).and_then(|(_, o)| number_of(o)).is_some()) +} + +fn move_line(ex: &mut Exec, tx: f64, ty: f64) { + ex.tlm = mul(&translate(tx, ty), &ex.tlm); + ex.tm = ex.tlm; +} + +fn next_line(ex: &mut Exec) { + let tl = ex.gs.text.tl; + move_line(ex, 0.0, -tl); +} + +/// One element of a show op's operands. +enum Item<'o> { + Str(&'o [u8]), + Kern(f64, Span), +} + +fn items(op: &Op) -> Vec> { + match op.operator { + Operator::TJ => match op.operands.first() { + Some(Operand::Array { items, .. }) => items + .iter() + .filter_map(|i| match i { + Operand::Str { bytes, .. } => Some(Item::Str(bytes)), + Operand::Number { value, span } => Some(Item::Kern(*value, span.clone())), + _ => None, + }) + .collect(), + _ => Vec::new(), + }, + _ => op + .operands + .last() + .and_then(Operand::as_str_bytes) + .map(|b| vec![Item::Str(b)]) + .unwrap_or_default(), + } +} + +impl<'a> Walker<'a> { + /// `Tc Tw Tz TL Tf Tr Ts`. + pub(super) fn text_state(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + let t = &mut ex.gs.text; + match op.operator { + Operator::Tc => { + t.tc = num(op, 0)?; + t.tc_src = Some(frame.op_bytes(op)); + } + Operator::Tw => { + t.tw = num(op, 0)?; + t.tw_src = Some(frame.op_bytes(op)); + } + Operator::Tz => t.th = num(op, 0)? / 100.0, + Operator::TL => t.tl = num(op, 0)?, + Operator::Ts => t.ts = num(op, 0)?, + Operator::Tr => { + let tr = num(op, 0)?; + if tr.fract() != 0.0 || !(0.0..=7.0).contains(&tr) { + return Err(Stop::malformed(format!( + "text render mode {tr} at byte {}", + op.op_span.start + ))); + } + t.tr = tr as i64; + } + Operator::Tf => return self.set_font(frame, ex, op), + _ => {} + } + Ok(()) + } + + fn set_font(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + let name = op + .operands + .first() + .and_then(Operand::as_name) + .unwrap_or_default(); + let size = num(op, 1)?; + let (hash, font_ref) = match frame.res.font(self.doc, name) { + Lookup::Found((key, dict)) => { + let model = self.load_font(key, dict)?; + if frame.root { + self.note_root_font(name, &model)?; + } + (model.content_hash, Some((key, model))) + } + Lookup::Missing => (0, None), + Lookup::Broken(w) => return Err(Stop::malformed(format!("font {w}"))), + }; + ex.gs.text.font = Some(FontUse { + resource: Some(Arc::from(name)), + content_hash: hash, + from_extgstate: false, + tf_op: Some(frame.op_bytes(op)), + }); + ex.gs.text.tfs = size; + ex.gs.font_ref = font_ref; + Ok(()) + } + + /// `Tj`, `TJ`, `'`, `"`. + pub(super) fn show(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + if !ex.in_text { + return Err(Stop::malformed(format!( + "text shown outside BT at byte {}", + op.op_span.start + ))); + } + let kind = match op.operator { + Operator::TJ => ShowOp::TJ, + Operator::Quote => ShowOp::Quote, + Operator::DoubleQuote => ShowOp::DoubleQuote, + _ => ShowOp::Tj, + }; + if kind == ShowOp::DoubleQuote { + let span_text = |i: usize| -> Vec { + op.operands + .get(i) + .map(|o| frame.span_bytes(o.span()).to_vec()) + .unwrap_or_default() + }; + let (aw, ac) = (num(op, 0)?, num(op, 1)?); + let t = &mut ex.gs.text; + t.tw = aw; + t.tw_src = Some(Arc::from([span_text(0), b" Tw".to_vec()].concat())); + t.tc = ac; + t.tc_src = Some(Arc::from([span_text(1), b" Tc".to_vec()].concat())); + } + if matches!(kind, ShowOp::Quote | ShowOp::DoubleQuote) { + next_line(ex); + ex.pen_unknown = false; + } + let before = self.digest(ex)?; + let rec = self.glyph_run(frame, ex, op, kind, before)?; + if ex.gs.text.tr >= 4 && !rec.glyphs.is_empty() { + ex.text_clip = true; + } + let proven = rec + .font_key + .and_then(|k| self.pen_proofs.get(&k)) + .copied() + .unwrap_or(false); + let advance_unknown = rec.split_error + || match &rec.font { + None => !rec.glyphs.is_empty(), + Some(m) => { + (!proven && !rec.glyphs.is_empty()) + || rec.glyphs.iter().any(|g| width_unknown(m, g.code)) + } + }; + if advance_unknown { + ex.pen_unknown = true; + } + self.mem.push(&mut self.records, rec) + } + + /// The glyph math of one show op, drawn with the state `before`. + fn glyph_run( + &mut self, + frame: &Frame<'_, 'a>, + ex: &mut Exec, + op: &Op, + kind: ShowOp, + before: Arc, + ) -> Result { + let tm_before = ex.tm; + let t = ex.gs.text.clone(); + let (tfs, th, tc, tw, ts) = (t.tfs, t.th, t.tc, t.tw, t.ts); + let font = ex.gs.font_ref.clone(); + let model = font.as_ref().map(|(_, m)| m); + let (ascent, descent) = model.map_or((FALLBACK_ASCENT, FALLBACK_DESCENT), |m| { + (m.ascent, m.descent) + }); + let params: Matrix = [tfs * th, 0.0, 0.0, tfs, 0.0, ts]; + let text_to_user = mul(¶ms, &mul(&ex.tm, &ex.gs.ctm)); + let pen_before = apply(&text_to_user, 0.0, 0.0); + let mut glyphs: Vec = Vec::new(); + let mut elems: Vec = Vec::new(); + let mut advance_ts = 0.0; + let mut split_error = false; + // The op's items and each string's codes live while the op is shown. + let listed = match op.operands.first() { + Some(Operand::Array { items, .. }) => items.len(), + _ => 1, + }; + self.mem.scratch(listed * size_of::>())?; + for item in items(op) { + match item { + Item::Kern(n, span) => { + charge(&mut self.kerns, 1, KERNS_PER_PAGE_MAX, "TJ kerns per page")?; + let tx = -n / 1000.0 * tfs * th; + advance_ts += -n / 1000.0 * tfs; + ex.tm = mul(&translate(tx, 0.0), &ex.tm); + self.mem + .push(&mut elems, RecElem::Kern { value: n, span })?; + } + Item::Str(bytes) => { + let codes_bytes = bytes.len() * size_of::(); + self.mem.scratch(codes_bytes)?; + let codes = match model { + Some(m) => m.split_codes(bytes).unwrap_or_else(|_| { + split_error = true; + m.split_codes(bytes.get(..bytes.len() & !1).unwrap_or_default()) + .unwrap_or_default() + }), + None => bytes + .iter() + .map(|b| Code { + value: u32::from(*b), + len: 1, + }) + .collect(), + }; + for code in codes { + charge(&mut self.glyphs, 1, GLYPHS_PER_PAGE_MAX, "glyphs per page")?; + let text = model.and_then(|m| m.text(code)); + let chars = text.map_or(0, |t| t.chars().count()); + charge( + &mut self.text_chars, + chars, + TEXT_CHARS_PER_PAGE_MAX, + "text characters per page", + )?; + self.mem.hold(text.map_or(0, str::len))?; + self.mem.grow(&mut glyphs, 1)?; + self.mem.grow(&mut elems, 1)?; + let w1000 = model.map_or(0.0, |m| m.width(code)); + let w0 = w1000 / 1000.0; + let word = model.is_some_and(|m| m.is_word_space(code)); + let m0 = mul(&ex.tm, &ex.gs.ctm); + let trm = mul(¶ms, &m0); + let tx = (w0 * tfs + tc + if word { tw } else { 0.0 }) * th; + let corners = [ + apply(&trm, 0.0, descent), + apply(&trm, w0, descent), + apply(&trm, 0.0, ascent), + apply(&trm, w0, ascent), + ]; + glyphs.push(GlyphRec { + code, + text: text.map(str::to_string), + origin: apply(&trm, 0.0, 0.0), + advance_user: apply_linear(&m0, tx, 0.0), + width1000: w1000, + bbox: bbox_of(&corners).unwrap_or_default(), + }); + elems.push(RecElem::Glyph(glyphs.len().saturating_sub(1))); + advance_ts += w0 * tfs + tc + if word { tw } else { 0.0 }; + ex.tm = mul(&translate(tx, 0.0), &ex.tm); + } + self.mem.unscratch(codes_bytes); + } + } + } + self.mem.unscratch(listed * size_of::>()); + // A record is kept for the page's lifetime: no spare capacity (a one-glyph op would + // otherwise hold room for four glyphs and four elements). + self.mem.fit(&mut glyphs); + self.mem.fit(&mut elems); + let pen_after = apply(&mul(¶ms, &mul(&ex.tm, &ex.gs.ctm)), 0.0, 0.0); + let numbers_ok = advance_ts.is_finite() + && [pen_before.0, pen_before.1, pen_after.0, pen_after.1] + .iter() + .all(|v| v.is_finite()) + && glyphs.iter().all(|g| { + g.bbox.iter().all(|v| v.is_finite()) + && g.advance_user.0.is_finite() + && g.advance_user.1.is_finite() + }); + if !numbers_ok || !text_to_user.iter().all(|v| v.is_finite()) { + return Err(Stop::malformed(format!( + "non-finite text position at byte {}", + op.op_span.start + ))); + } + let spans = op.operands.len() * size_of::(); + self.mem + .hold(spans + frame.chain.len() * size_of::())?; + let after = self.digest(ex)?; + Ok(ShowRecord { + seq: self.next_seq()?, + depth: frame.depth, + form_chain: frame.chain.clone(), + op: kind, + span: frame.joined.then(|| op.span.clone()), + local_span: op.span.clone(), + operand_spans: op.operands.iter().map(|o| o.span().clone()).collect(), + before, + after, + tm_before, + tm_after: ex.tm, + glyphs, + elems, + pen_before, + pen_after, + advance_ts, + text_to_user, + after_unproven_inline_image: op.after_unproven_inline_image, + font_key: font.as_ref().map(|(k, _)| *k), + font: font.map(|(_, m)| m), + split_error, + pen_unknown: ex.pen_unknown || ex.tm_unsettled, + }) + } +} diff --git a/src-tauri/src/pdf_engine/text_edit/walker/xobj.rs b/src-tauri/src/pdf_engine/text_edit/walker/xobj.rs new file mode 100644 index 0000000..5fb420e --- /dev/null +++ b/src-tauri/src/pdf_engine/text_edit/walker/xobj.rs @@ -0,0 +1,477 @@ +//! XObjects, inline images and shadings (SPEC §B.10): image and Form paints, Form descent in +//! `Classify` mode (depth ≤ 8, cycle check, `/Matrix`, `/BBox` clip) and the wrapper Form of +//! `Wrapped` mode (qpdf's overlay wrapper, §B.16.2), whose content is walked as depth 0 after the +//! wrapper conditions are re-checked. A Form's decoded content and lexed ops are charged to the +//! walk's budget while the Form runs (`budget::ModelBudget` scratch), its paints and records for +//! good. + +use super::budget::{lex_ops, MODEL_SIZE}; +use super::{Exec, Frame, PaintKind, PaintRecord, Stop, WalkMode, Walked, Walker, XObjectUse}; +use crate::pdf_engine::text_edit::context::{resolve, Lookup, Res}; +use crate::pdf_engine::text_edit::decode::decode_stream; +use crate::pdf_engine::text_edit::fonts::number_of; +use crate::pdf_engine::text_edit::geometry::{ + axis_aligned, contains, intersect, mul, transform_rect, Matrix, IDENTITY, +}; +use crate::pdf_engine::text_edit::lexer::{Op, Operand, Operator}; +use crate::pdf_engine::text_edit::limits::{ + FORM_DEPTH_MAX, FORM_PAINTS_PER_PAGE_MAX, PAGE_CONTENT_MAX_DECODED, STREAM_MAX_DECODED, + WRAPPER_MATRIX_EPSILON, WRAPPER_TRANSLATION_TOL_PT, +}; +use crate::pdf_engine::text_edit::reasons::TextReason; +use crate::pdf_engine::text_edit::snapshot::fnv1a_u64; +use crate::pdf_engine::text_edit::state::{ClipState, GState, MarkedStack}; +use lopdf::{Dictionary, Object, ObjectId, Stream}; + +/// The wrapper Form's `/BBox` must cover the visible box within this (§B.16.2). +const WRAPPER_BBOX_TOL_PT: f64 = 0.01; +/// The keys qpdf writes on its overlay wrapper Form (`getFormXObjectForPage`) besides `/Group`, +/// which must equal the page's. Any other key (`/OC`, `/Ref`, `/OPI`, …) could change what the +/// wrapper paints without changing the walk, so it is refused. +const WRAPPER_KEYS: [&[u8]; 10] = [ + b"Type", + b"Subtype", + b"FormType", + b"BBox", + b"Matrix", + b"Resources", + b"Length", + b"Filter", + b"DecodeParms", + b"DL", +]; +const UNIT_SQUARE: [f64; 4] = [0.0, 0.0, 1.0, 1.0]; + +/// A resolved `/XObject` entry. +struct XObj<'a> { + id: ObjectId, + stream: &'a Stream, + shared_path: bool, +} + +impl<'a> Walker<'a> { + fn xobject(&self, frame: &Frame<'_, 'a>, name: &[u8]) -> Result, Stop> { + let e = match frame.res.entry(self.doc, b"XObject", name) { + Lookup::Found(e) => e, + Lookup::Missing => return Err(Stop::malformed("XObject name not in resources")), + Lookup::Broken(w) => return Err(Stop::malformed(format!("XObject {w}"))), + }; + let (Some(id), Object::Stream(stream)) = (e.id, e.value) else { + return Err(Stop::malformed("XObject is not a stream")); + }; + let refs = self.ctx.refs(); + let shared_path = frame.res.inherited + || frame.res.dict_id.is_some_and(|d| refs.count(d) > 1) + || e.category_id.is_some_and(|c| refs.count(c) > 1); + Ok(XObj { + id, + stream, + shared_path, + }) + } + + pub(super) fn do_xobject(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + let name = op + .operands + .first() + .and_then(Operand::as_name) + .unwrap_or_default(); + let x = self.xobject(frame, name)?; + let subtype = x + .stream + .dict + .get(b"Subtype") + .ok() + .and_then(|o| o.as_name().ok()); + match subtype { + Some(b"Image") => { + let hash = self.stream_hash(x.id)?; + self.mem.hold(name.len())?; + let masked = ex.gs.gs.soft_mask || image_masked(self.doc, &x.stream.dict); + let bbox = transform_rect(&ex.gs.ctm, UNIT_SQUARE); + self.push_paint( + frame, + ex, + op, + PaintKind::ImageXObject { + name: name.to_vec(), + hash, + }, + Some(bbox), + masked, + Some(&x), + ) + } + Some(b"Form") => self.form(frame, ex, op, name, &x), + Some(b"PS") => Ok(()), + _ => Err(Stop::malformed("XObject subtype")), + } + } + + #[allow(clippy::too_many_arguments)] + fn push_paint( + &mut self, + frame: &Frame<'_, 'a>, + ex: &mut Exec, + op: &Op, + kind: PaintKind, + bbox: Option<[f64; 4]>, + masked: bool, + x: Option<&XObj<'a>>, + ) -> Walked { + let seq = self.next_seq()?; + let state = self.digest(ex)?; + self.mem + .hold(frame.chain.len() * std::mem::size_of::())?; + let paint = PaintRecord { + seq, + depth: frame.depth, + kind, + span: frame.joined.then(|| op.span.clone()), + state, + bbox, + masked, + form_chain: frame.chain.clone(), + local_span: op.span.clone(), + xobject: x.map(|x| XObjectUse { + id: x.id, + shared_path: x.shared_path, + }), + }; + self.mem.push(&mut self.paints, paint) + } + + fn form( + &mut self, + frame: &Frame<'_, 'a>, + ex: &mut Exec, + op: &Op, + name: &[u8], + x: &XObj<'a>, + ) -> Walked { + self.form_paints = self.form_paints.saturating_add(1); + if self.form_paints > FORM_PAINTS_PER_PAGE_MAX { + return Err(Stop::new( + TextReason::PageTooComplex, + "form paints per page", + )); + } + let matrix = form_matrix(self.doc, &x.stream.dict)?; + let bbox = form_bbox(self.doc, &x.stream.dict); + let composite = mul(&matrix, &ex.gs.ctm); + let hash = self.stream_hash(x.id)?; + let paint_bbox = bbox.map(|b| transform_rect(&composite, b)); + self.mem.hold(name.len())?; + let kind = PaintKind::FormXObject { + name: name.to_vec(), + hash, + }; + let masked = ex.gs.gs.soft_mask; + self.push_paint(frame, ex, op, kind, paint_bbox, masked, Some(x))?; + if self.mode != WalkMode::Classify { + return Ok(()); + } + let depth = frame.depth.saturating_add(1); + if usize::from(depth) > FORM_DEPTH_MAX { + return Err(Stop::malformed(format!( + "Form XObjects nested deeper than {FORM_DEPTH_MAX}" + ))); + } + if frame.chain.contains(&x.id) { + return Err(Stop::malformed("a Form XObject paints itself")); + } + let mut gs = ex.gs.clone(); + gs.ctm = composite; + if let Some(b) = bbox { + gs.clip = clip_with(&gs.clip, &composite, b); + } + let mut chain = frame.chain.clone(); + chain.push(x.id); + let res = match Res::of_form(self.doc, x.id, x.stream) { + None => frame.res, + Some(Ok(r)) => r, + Some(Err(w)) => return Err(Stop::malformed(format!("Form {w}"))), + }; + let mut sub = Exec::new(gs, ex.marked.clone()); + self.run_stream(x.stream, res, depth, chain, false, &mut sub) + } + + /// Decodes, lexes and runs one Form's content (the wrapper Form of `Wrapped` mode, `root`, + /// holds the page content and gets the page content cap). + fn run_stream( + &mut self, + stream: &'a Stream, + res: Res<'a>, + depth: u8, + chain: Vec, + root: bool, + ex: &mut Exec, + ) -> Walked { + let cap = if root { + PAGE_CONTENT_MAX_DECODED + } else { + STREAM_MAX_DECODED + }; + // The decoded content and the lexed ops live while the Form runs (nested Forms add theirs + // on top); decoding stops where the byte budget would. + let left = self.mem.left(); + let data = decode_stream(stream, cap.min(left), &mut self.budget).map_err(|e| { + if left < cap && e.page_reason() == TextReason::PageTooComplex { + Stop::new(TextReason::PageTooComplex, MODEL_SIZE) + } else { + Stop::new(e.page_reason(), format!("Form XObject: {e}")) + } + })?; + self.mem.scratch(data.capacity())?; + let lexed = lex_ops( + &data, + self.ops_left, + self.operand_nodes, + &mut self.mem, + self.cancel, + false, + "Form XObject: ", + )?; + let (ops, nodes) = (lexed.ops, lexed.nodes); + self.ops_left = self.ops_left.saturating_sub(ops.len()); + self.operand_nodes = self.operand_nodes.saturating_add(nodes); + let frame = Frame { + bytes: &data, + res, + depth, + chain, + joined: false, + root, + }; + let walked = self.run(&frame, ex, &ops); + self.operand_nodes = self.operand_nodes.saturating_sub(nodes); + self.mem + .unscratch(lexed.bytes.saturating_add(data.capacity())); + walked + } + + pub(super) fn inline_image(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + let hash = fnv1a_u64(frame.span_bytes(&op.span)); + let stencil = op.inline_image.as_ref().is_some_and(|img| { + img.dict.iter().any(|(k, v)| { + matches!(k.as_slice(), b"IM" | b"ImageMask") + && matches!(v, Operand::Bool { value: true, .. }) + }) + }); + let masked = ex.gs.gs.soft_mask || stencil; + let bbox = transform_rect(&ex.gs.ctm, UNIT_SQUARE); + self.push_paint( + frame, + ex, + op, + PaintKind::InlineImage { hash }, + Some(bbox), + masked, + None, + ) + } + + pub(super) fn shading(&mut self, frame: &Frame<'_, 'a>, ex: &mut Exec, op: &Op) -> Walked { + let name = op + .operands + .first() + .and_then(Operand::as_name) + .unwrap_or_default(); + let hash = match frame.res.entry(self.doc, b"Shading", name) { + Lookup::Found(e) => self.entry_hash(e.id, e.value)?, + Lookup::Missing => return Err(Stop::malformed("shading not in resources")), + Lookup::Broken(w) => return Err(Stop::malformed(format!("shading {w}"))), + }; + let bbox = match &ex.gs.clip { + ClipState::Rect(r) => Some(*r), + _ => Some(self.geometry.visible), + }; + self.mem.hold(name.len())?; + let kind = PaintKind::Shading { + name: name.to_vec(), + hash, + }; + let masked = ex.gs.gs.soft_mask; + self.push_paint(frame, ex, op, kind, bbox, masked, None) + } + + /// `Wrapped` mode (§B.10): the page content may only hold `q`, `Q`, `cm` and `Do`, exactly one + /// `Do` names the wrapper Form, whose composite matrix is the identity within tolerance, whose + /// `/BBox` covers the visible box and whose dictionary holds only qpdf's keys (a `/Group` + /// equal to the page's); its content is walked as depth 0 with the clip factored out. Every + /// other `Do` stays a paint. + pub(super) fn walk_wrapped( + &mut self, + frame: &Frame<'_, 'a>, + ops: &[Op], + name: &[u8], + page_id: ObjectId, + ) -> Walked { + let wrapper = |what: &str| Stop::malformed(format!("wrapper: {what}")); + let shape_ok = ops.iter().all(|o| { + matches!( + o.operator, + Operator::q | Operator::Q | Operator::cm | Operator::Do + ) + }); + if !shape_ok { + return Err(wrapper("page content is not q/cm/Do/Q only")); + } + let hits = ops + .iter() + .filter(|o| { + o.operator == Operator::Do + && o.operands.first().and_then(Operand::as_name) == Some(name) + }) + .count(); + if hits != 1 { + return Err(wrapper("the wrapper Form is not painted exactly once")); + } + let mut ex = Exec::new(GState::initial(IDENTITY), MarkedStack::default()); + for op in ops { + let is_wrapper = op.operator == Operator::Do + && op.operands.first().and_then(Operand::as_name) == Some(name); + if !is_wrapper { + self.op(frame, &mut ex, op)?; + continue; + } + let x = self.xobject(frame, name)?; + if x.stream + .dict + .get(b"Subtype") + .ok() + .and_then(|o| o.as_name().ok()) + != Some(b"Form") + { + return Err(wrapper("the wrapper is not a Form XObject")); + } + self.wrapper_keys(&x.stream.dict, page_id)?; + let m = mul(&form_matrix(self.doc, &x.stream.dict)?, &ex.gs.ctm); + let linear_ok = (m[0] - 1.0).abs() <= WRAPPER_MATRIX_EPSILON + && m[1].abs() <= WRAPPER_MATRIX_EPSILON + && m[2].abs() <= WRAPPER_MATRIX_EPSILON + && (m[3] - 1.0).abs() <= WRAPPER_MATRIX_EPSILON; + let shift_ok = m[4].abs() <= WRAPPER_TRANSLATION_TOL_PT + && m[5].abs() <= WRAPPER_TRANSLATION_TOL_PT; + if !linear_ok || !shift_ok { + return Err(wrapper("composite matrix is not the identity")); + } + let bbox = form_bbox(self.doc, &x.stream.dict).ok_or_else(|| wrapper("no /BBox"))?; + if !contains( + transform_rect(&m, bbox), + self.geometry.visible, + WRAPPER_BBOX_TOL_PT, + ) { + return Err(wrapper("/BBox does not cover the visible page")); + } + let res = match Res::of_form(self.doc, x.id, x.stream) { + None => frame.res, + Some(Ok(r)) => r, + Some(Err(w)) => return Err(wrapper(w)), + }; + self.root_res = res; + let mut sub = Exec::new(GState::initial(m), MarkedStack::default()); + self.run_stream(x.stream, res, 0, Vec::new(), true, &mut sub)?; + } + Ok(()) + } + + /// The wrapper Form's dictionary holds only `WRAPPER_KEYS` and a `/Group` whose deep hash + /// equals the page's `/Group` (qpdf copies it); anything else is `MALFORMED_CONTENT` + /// "wrapper". The two `/Group`s are hashed with hash state of their own + /// (`with_own_hash_state`), not the content's. + fn wrapper_keys(&mut self, dict: &'a Dictionary, page_id: ObjectId) -> Walked { + let wrapper = |what: &str| Stop::malformed(format!("wrapper: {what}")); + for (key, raw) in dict.iter() { + if WRAPPER_KEYS.contains(&key.as_slice()) { + continue; + } + if key.as_slice() != b"Group" { + let shown = key.get(..64).unwrap_or(key); + return Err(wrapper(&format!( + "unexpected /{} on the Form", + String::from_utf8_lossy(shown) + ))); + } + let page_group = match self.doc.objects.get(&page_id) { + Some(Object::Dictionary(page)) => page.get(b"Group").ok(), + _ => None, + }; + let same = match ( + resolve(self.doc, raw), + page_group.and_then(|g| resolve(self.doc, g)), + ) { + (Some((a_id, a)), Some((b_id, b))) => self.with_own_hash_state(|w| { + Ok(w.entry_hash(a_id, a)? == w.entry_hash(b_id, b)?) + })?, + _ => false, + }; + if !same { + return Err(wrapper("/Group differs from the page's")); + } + } + Ok(()) + } +} + +/// A Form's `/Matrix` (identity when absent); anything but six numbers is malformed. +fn form_matrix(doc: &lopdf::Document, dict: &Dictionary) -> Result { + let Some(raw) = dict.get(b"Matrix").ok() else { + return Ok(IDENTITY); + }; + let bad = || Stop::malformed("Form /Matrix"); + let items = match resolve(doc, raw) { + Some((_, Object::Array(items))) if items.len() == 6 => items, + _ => return Err(bad()), + }; + let mut m = IDENTITY; + for (slot, item) in m.iter_mut().zip(items) { + *slot = resolve(doc, item) + .and_then(|(_, o)| number_of(o)) + .ok_or_else(bad)?; + } + Ok(m) +} + +/// A Form's `/BBox`, normalised; `None` when absent or unreadable. +fn form_bbox(doc: &lopdf::Document, dict: &Dictionary) -> Option<[f64; 4]> { + let items = match resolve(doc, dict.get(b"BBox").ok()?)? { + (_, Object::Array(items)) if items.len() == 4 => items, + _ => return None, + }; + let v: Vec = items + .iter() + .map(|i| resolve(doc, i).and_then(|(_, o)| number_of(o))) + .collect::>>()?; + match v.as_slice() { + [a, b, c, d] => Some([a.min(*c), b.min(*d), a.max(*c), b.max(*d)]), + _ => None, + } +} + +/// The clip after intersecting with a Form `/BBox` mapped by `m`. +fn clip_with(clip: &ClipState, m: &Matrix, bbox: [f64; 4]) -> ClipState { + if !axis_aligned(m) { + return ClipState::Complex; + } + let r = transform_rect(m, bbox); + match clip { + ClipState::None => ClipState::Rect(r), + ClipState::Rect(c) => ClipState::Rect(intersect(*c, r)), + ClipState::Complex => ClipState::Complex, + } +} + +/// `/SMask` (other than `/None`), `/Mask` or `/ImageMask true` on an image. +fn image_masked(doc: &lopdf::Document, dict: &Dictionary) -> bool { + let smask = dict + .get(b"SMask") + .ok() + .and_then(|o| resolve(doc, o)) + .is_some_and(|(_, o)| o.as_name().ok() != Some(b"None")); + let image_mask = matches!( + dict.get(b"ImageMask").ok().and_then(|o| resolve(doc, o)), + Some((_, Object::Boolean(true))) + ); + smask || dict.has(b"Mask") || image_mask +} diff --git a/src-tauri/src/pdf_engine/validate_output.rs b/src-tauri/src/pdf_engine/validate_output.rs index 97148f3..ce63787 100644 --- a/src-tauri/src/pdf_engine/validate_output.rs +++ b/src-tauri/src/pdf_engine/validate_output.rs @@ -232,6 +232,20 @@ pub fn validate_staged_pdf( staged: &Path, snapshot: &OutputSnapshot, cancel: Option<&AtomicBool>, + run_check: impl FnMut(&[String]) -> Result<(i32, String), AppError>, +) -> Result { + validate_staged_pdf_with_alternatives(staged, snapshot, &[], cancel, run_check) +} + +/// [`validate_staged_pdf`] where page `i` may also match `alternatives[i]`: the digest of its +/// content parts as qpdf joins them into an overlay Form (a `\n` after a part that lacks one), +/// which differs from lopdf's separator-less `get_page_content` digest (SPEC §B.1, V1). +/// `None` (or a missing entry) keeps the single expected digest. +pub fn validate_staged_pdf_with_alternatives( + staged: &Path, + snapshot: &OutputSnapshot, + alternatives: &[Option], + cancel: Option<&AtomicBool>, mut run_check: impl FnMut(&[String]) -> Result<(i32, String), AppError>, ) -> Result { abort_if_cancelled(staged, cancel)?; @@ -326,7 +340,11 @@ pub fn validate_staged_pdf( )); } let candidates = dest_page_digests(&doc, id); - if !candidates.iter().any(|d| *d == expected.content_digest) { + let alternative = alternatives.get(i).copied().flatten(); + if !candidates + .iter() + .any(|d| *d == expected.content_digest || Some(*d) == alternative) + { return Err(fatal_staged( staged, format!("Page {page_no} content does not match the source."), @@ -941,4 +959,116 @@ mod tests { "R1b: gate must not publish; dest bytes must stay OLD" ); } + + /// A one-page PDF whose content is `parts` (separate streams, no trailing newlines). + fn write_parts_page(path: &Path, parts: &[&[u8]]) { + let mut doc = Document::with_version("1.5"); + let pages_id = doc.new_object_id(); + let ids: Vec = parts + .iter() + .map(|p| { + doc.add_object(Object::Stream(Stream::new(Dictionary::new(), p.to_vec()))) + .into() + }) + .collect(); + let mut page = Dictionary::new(); + page.set("Type", "Page"); + page.set("Parent", pages_id); + page.set("MediaBox", box_obj([0, 0, 612, 792])); + page.set("Contents", Object::Array(ids)); + let page_id = doc.add_object(Object::Dictionary(page)); + let mut pages = Dictionary::new(); + pages.set("Type", "Pages"); + pages.set("Kids", vec![page_id.into()]); + pages.set("Count", 1); + doc.objects.insert(pages_id, Object::Dictionary(pages)); + let mut catalog = Dictionary::new(); + catalog.set("Type", "Catalog"); + catalog.set("Pages", pages_id); + let catalog_id = doc.add_object(Object::Dictionary(catalog)); + doc.trailer.set("Root", catalog_id); + doc.save(path).expect("write parts fixture"); + } + + /// VO-01: qpdf's overlay wraps a two-part page into one Form whose data is the parts joined + /// with a "\n" after the first (it lacks one). #34's plain digest misses it; the alternative + /// digest (`qpdf_join`) accepts it; one changed byte still fails; no alternative = old result. + #[test] + fn vo_01_overlay_wrapper_of_a_split_page_matches_the_alternative_digest() { + use crate::pdf_engine::text_edit::content::qpdf_join; + let Some(engines) = crate::pdf_engine::text_edit::testkit::engines_or_skip("vo_01") else { + return; + }; + let scratch = Scratch::new("vo01"); + let run_check = |args: &[String]| -> Result<(i32, String), AppError> { + let out = std::process::Command::new(&engines.qpdf) + .args(args) + .output() + .map_err(|e| AppError::io("qpdf --check", e))?; + Ok(( + out.status.code().unwrap_or(2), + String::from_utf8_lossy(&out.stderr).into_owned(), + )) + }; + let overlay = |parts: &[&[u8]], name: &str| -> PathBuf { + let src = scratch.path().join(format!("{name}-src.pdf")); + let blank = scratch.path().join(format!("{name}-blank.pdf")); + let staged = scratch.path().join(format!("{name}-staged.pdf")); + write_parts_page(&src, parts); + write_parts_page(&blank, &[b""]); + let out = std::process::Command::new(&engines.qpdf) + .arg(&src) + .arg("--overlay") + .arg(&blank) + .arg("--") + .arg(&staged) + .output() + .expect("qpdf --overlay"); + assert!(matches!(out.status.code(), Some(0 | 3)), "{out:?}"); + staged + }; + let parts: [&[u8]; 2] = [ + b"BT /F1 12 Tf 72 720 Td (Hello) Tj ET", + b"BT /F1 12 Tf 72 700 Td (World) Tj ET", + ]; + let plain = parts.concat(); + let joined = qpdf_join(&parts); + assert_ne!( + plain, joined, + "VO-01: the first part lacks a trailing newline" + ); + let snapshot = OutputSnapshot { + pages: vec![PageSnapshot { + content_digest: content_digest(&plain), + ..letter_page() + }], + catalog: empty_catalog(), + }; + let alt = [Some(content_digest(&joined))]; + + let staged = overlay(&parts, "honest"); + let without = validate_staged_pdf(&staged, &snapshot, None, run_check); + assert!( + without.is_err(), + "VO-01: #34 alone refuses the honest wrapper" + ); + let staged = overlay(&parts, "honest2"); + let with = validate_staged_pdf_with_alternatives(&staged, &snapshot, &alt, None, run_check); + assert!( + with.is_ok(), + "VO-01: the alternative digest accepts it: {with:?}" + ); + let staged = overlay(&parts, "none"); + let none = + validate_staged_pdf_with_alternatives(&staged, &snapshot, &[None], None, run_check); + assert!(none.is_err(), "VO-01: alt None keeps the old behaviour"); + + let changed: [&[u8]; 2] = [ + b"BT /F1 12 Tf 72 720 Td (Hellp) Tj ET", + b"BT /F1 12 Tf 72 700 Td (World) Tj ET", + ]; + let staged = overlay(&changed, "changed"); + let r = validate_staged_pdf_with_alternatives(&staged, &snapshot, &alt, None, run_check); + assert!(r.is_err(), "VO-01: one changed byte still fails"); + } } diff --git a/src/components/pdf/editor/EditorObjectShape.tsx b/src/components/pdf/editor/EditorObjectShape.tsx new file mode 100644 index 0000000..7d784a9 --- /dev/null +++ b/src/components/pdf/editor/EditorObjectShape.tsx @@ -0,0 +1,361 @@ +/** + * SVG shapes for the editor overlay: closed shapes and one placed object with + * its selection handles. Moved unchanged from `EditorOverlay.tsx`. + */ +import { + bubbleSvgPath, + closedShapeCssPoints, + cssCenter, + displayedSize, + isClosedShapeObject, + isNoneFill, + pdfRectToViewport, + pdfToViewport, + type ClosedShapeKind, + type EditObject, + type ResizeHandle, + type ViewportMapping, +} from "@/lib/editor"; + +type Handle = ResizeHandle; + +export function ClosedShapeSvg({ + kind, + css, + fill, + stroke, + strokeWidth, + locked, + interactive, + onPointerDown, +}: { + kind: ClosedShapeKind; + css: { x: number; y: number; w: number; h: number }; + fill: string; + stroke: string; + strokeWidth: number; + locked: boolean; + interactive: boolean; + onPointerDown: (e: React.PointerEvent) => void; +}) { + const common = { + fill, + stroke, + strokeWidth, + style: { cursor: !interactive ? undefined : locked ? "default" : "move" } as const, + onPointerDown: interactive ? onPointerDown : undefined, + }; + const w = Math.max(css.w, 1); + const h = Math.max(css.h, 1); + if (kind === "rect") { + return ; + } + if (kind === "roundRect") { + const r = Math.min(w, h) * 0.18; + return ; + } + if (kind === "ellipse") { + return ; + } + if (kind === "bubble") { + return ; + } + return ( + + ); +} + +export function ObjectShape({ + obj, + mapping, + selected, + selectedCount, + interactive, + onPointerDownObject, + onPointerDownHandle, + onPointerDownRotate, +}: { + obj: EditObject; + mapping: ViewportMapping; + selected: boolean; + selectedCount: number; + interactive: boolean; + onPointerDownObject: (e: React.PointerEvent, obj: EditObject) => void; + onPointerDownHandle: (e: React.PointerEvent, obj: EditObject, handle: Handle) => void; + onPointerDownRotate: (e: React.PointerEvent, obj: EditObject) => void; +}) { + const css = pdfRectToViewport(obj.rect, mapping); + const opacity = "opacity" in obj && typeof obj.opacity === "number" ? obj.opacity : 1; + const rot = obj.objectRotate ?? 0; + const mid = cssCenter(css); + + const moveCursor = interactive && !obj.locked ? "move" : undefined; + + return ( + + + {isClosedShapeObject(obj) && ( + onPointerDownObject(e, obj)} + /> + )} + {obj.kind === "text" && ( + onPointerDownObject(e, obj) : undefined}> + + +
+ {obj.content || "Text"} +
+
+
+ )} + {obj.kind === "image" && ( + onPointerDownObject(e, obj) : undefined}> + {obj.previewUrl ? ( + + ) : ( + + )} + + )} + {obj.kind === "line" && ( + onPointerDownObject(e, obj) : undefined}> + + + + )} + {obj.kind === "ink" && ( + { + const c = pdfToViewport(p, mapping); + return `${c.x},${c.y}`; + }).join(" ")} + fill="none" + stroke={obj.stroke ?? "#111827"} + strokeWidth={obj.strokeWidth ?? 2.5} + strokeLinecap="round" + strokeLinejoin="round" + style={{ cursor: moveCursor }} + onPointerDown={interactive ? (e) => onPointerDownObject(e, obj) : undefined} + /> + )} + {obj.kind === "link" && ( + onPointerDownObject(e, obj) : undefined}> + + + )} + {obj.kind === "note" && ( + onPointerDownObject(e, obj) : undefined}> + + + )} + {(obj.kind === "highlight" || obj.kind === "underline" || obj.kind === "strikeout") && ( + onPointerDownObject(e, obj) : undefined}> + {obj.kind === "highlight" ? ( + + ) : ( + + )} + + )} + {obj.kind === "markupInk" && ( + onPointerDownObject(e, obj) : undefined}> + {obj.strokes.map((stroke, i) => ( + { + const c = pdfToViewport(p, mapping); + return `${c.x},${c.y}`; + }).join(" ")} + fill="none" + stroke={obj.color ?? "#111827"} + strokeWidth={2.5} + strokeLinecap="round" + strokeLinejoin="round" + /> + ))} + + )} + {obj.kind === "redact" && ( + onPointerDownObject(e, obj) : undefined}> + + + + {obj.label ? ( + + {obj.label} + + ) : null} + + )} +
+ {selected && ( + + )} + {selected && selectedCount === 1 && !obj.locked && obj.kind !== "link" && obj.kind !== "redact" && ( + <> + + onPointerDownRotate(e, obj)} + /> + + )} + {selected && + selectedCount === 1 && + !obj.locked && + (["nw", "ne", "sw", "se"] as Handle[]).map((handle) => { + const hx = handle === "nw" || handle === "sw" ? css.x : css.x + css.w; + const hy = handle === "nw" || handle === "ne" ? css.y : css.y + css.h; + return ( + onPointerDownHandle(e, obj, handle)} + /> + ); + })} +
+ ); +} diff --git a/src/components/pdf/editor/EditorOverlay.tsx b/src/components/pdf/editor/EditorOverlay.tsx index 0ba1b90..62838be 100644 --- a/src/components/pdf/editor/EditorOverlay.tsx +++ b/src/components/pdf/editor/EditorOverlay.tsx @@ -4,17 +4,12 @@ */ import { useRef, useState } from "react"; import { - bubbleSvgPath, - closedShapeCssPoints, cssCenter, - displayedSize, - isClosedShapeObject, isNoneFill, makeMapping, moveSelectedRects, normalizeDeg, pdfRectToViewport, - pdfToViewport, pointerAngleDeg, aspectLocked, constrainCssBox1to1, @@ -34,10 +29,12 @@ import { } from "@/lib/editor"; import type { PageLayout } from "./PageSurface"; import { shapeKindForTool, toolForces1to1 } from "./ShapePicker"; +import { ClosedShapeSvg, ObjectShape } from "./EditorObjectShape"; export type EditorTool = | "select" | "hand" + | "editText" | "text" | "rect" | "square" @@ -187,7 +184,11 @@ export function EditorOverlay({ const style: ShapeStyle = createStyle ?? {}; const mapping = makeMapping(layout.geometry, layout.cssWidth, layout.cssHeight); - const pageObjects = objects.filter((o) => o.pageIndex === pageIndex); + // Text changes (sourceText) are drawn by the Edit text layer, never here: no box, + // no handles, never hit by the marquee. + const pageObjects = objects.filter((o) => o.pageIndex === pageIndex && o.kind !== "sourceText"); + // In Edit text mode the overlay is inert: stamps stay visible, the text layer takes the pointer. + const inert = tool === "hand" || tool === "editText"; const onPointerDownBg = (e: React.PointerEvent) => { if (e.button !== 0) return; @@ -200,7 +201,7 @@ export function EditorOverlay({ e.preventDefault(); return; } - if (tool === "hand") return; + if (inert) return; if (tool === "image") { onRequestImage(local); @@ -252,7 +253,7 @@ export function EditorOverlay({ if (e.button !== 0) return; const svg = svgRef.current; if (!svg) return; - if (tool === "hand") return; + if (inert) return; // Create tools (square, text, ink, …) must start a new object even if the // pointer is over an existing one — don't steal the event from the page. if (!pickColor && tool !== "select") return; @@ -513,7 +514,13 @@ export function EditorOverlay({ onEndGesture(); }; - const cursor = pickColor ? "none" : tool === "hand" ? "grab" : tool === "select" ? "default" : "crosshair"; + const cursor = pickColor + ? "none" + : tool === "hand" + ? "grab" + : tool === "select" || tool === "editText" + ? "default" + : "crosshair"; const objectsInteractive = tool === "select" || !!pickColor; const shapeDraftCss = draft?.kind === "shape" && draft.shape ? cssBoxFromPoints(draft.start, draft.cur) : null; @@ -521,7 +528,7 @@ export function EditorOverlay({ return ( void; -}) { - const common = { - fill, - stroke, - strokeWidth, - style: { cursor: !interactive ? undefined : locked ? "default" : "move" } as const, - onPointerDown: interactive ? onPointerDown : undefined, - }; - const w = Math.max(css.w, 1); - const h = Math.max(css.h, 1); - if (kind === "rect") { - return ; - } - if (kind === "roundRect") { - const r = Math.min(w, h) * 0.18; - return ; - } - if (kind === "ellipse") { - return ; - } - if (kind === "bubble") { - return ; - } - return ( - - ); -} - -function ObjectShape({ - obj, - mapping, - selected, - selectedCount, - interactive, - onPointerDownObject, - onPointerDownHandle, - onPointerDownRotate, -}: { - obj: EditObject; - mapping: ViewportMapping; - selected: boolean; - selectedCount: number; - interactive: boolean; - onPointerDownObject: (e: React.PointerEvent, obj: EditObject) => void; - onPointerDownHandle: (e: React.PointerEvent, obj: EditObject, handle: Handle) => void; - onPointerDownRotate: (e: React.PointerEvent, obj: EditObject) => void; -}) { - const css = pdfRectToViewport(obj.rect, mapping); - const opacity = "opacity" in obj && typeof obj.opacity === "number" ? obj.opacity : 1; - const rot = obj.objectRotate ?? 0; - const mid = cssCenter(css); - - const moveCursor = interactive && !obj.locked ? "move" : undefined; - - return ( - - - {isClosedShapeObject(obj) && ( - onPointerDownObject(e, obj)} - /> - )} - {obj.kind === "text" && ( - onPointerDownObject(e, obj) : undefined}> - - -
- {obj.content || "Text"} -
-
-
- )} - {obj.kind === "image" && ( - onPointerDownObject(e, obj) : undefined}> - {obj.previewUrl ? ( - - ) : ( - - )} - - )} - {obj.kind === "line" && ( - onPointerDownObject(e, obj) : undefined}> - - - - )} - {obj.kind === "ink" && ( - { - const c = pdfToViewport(p, mapping); - return `${c.x},${c.y}`; - }).join(" ")} - fill="none" - stroke={obj.stroke ?? "#111827"} - strokeWidth={obj.strokeWidth ?? 2.5} - strokeLinecap="round" - strokeLinejoin="round" - style={{ cursor: moveCursor }} - onPointerDown={interactive ? (e) => onPointerDownObject(e, obj) : undefined} - /> - )} - {obj.kind === "link" && ( - onPointerDownObject(e, obj) : undefined}> - - - )} - {obj.kind === "note" && ( - onPointerDownObject(e, obj) : undefined}> - - - )} - {(obj.kind === "highlight" || obj.kind === "underline" || obj.kind === "strikeout") && ( - onPointerDownObject(e, obj) : undefined}> - {obj.kind === "highlight" ? ( - - ) : ( - - )} - - )} - {obj.kind === "markupInk" && ( - onPointerDownObject(e, obj) : undefined}> - {obj.strokes.map((stroke, i) => ( - { - const c = pdfToViewport(p, mapping); - return `${c.x},${c.y}`; - }).join(" ")} - fill="none" - stroke={obj.color ?? "#111827"} - strokeWidth={2.5} - strokeLinecap="round" - strokeLinejoin="round" - /> - ))} - - )} - {obj.kind === "redact" && ( - onPointerDownObject(e, obj) : undefined}> - - - - {obj.label ? ( - - {obj.label} - - ) : null} - - )} -
- {selected && ( - - )} - {selected && selectedCount === 1 && !obj.locked && obj.kind !== "link" && obj.kind !== "redact" && ( - <> - - onPointerDownRotate(e, obj)} - /> - - )} - {selected && - selectedCount === 1 && - !obj.locked && - (["nw", "ne", "sw", "se"] as Handle[]).map((handle) => { - const hx = handle === "nw" || handle === "sw" ? css.x : css.x + css.w; - const hy = handle === "nw" || handle === "ne" ? css.y : css.y + css.h; - return ( - onPointerDownHandle(e, obj, handle)} - /> - ); - })} -
- ); -} diff --git a/src/components/pdf/editor/EditorToolbar.test.tsx b/src/components/pdf/editor/EditorToolbar.test.tsx new file mode 100644 index 0000000..5a0bace --- /dev/null +++ b/src/components/pdf/editor/EditorToolbar.test.tsx @@ -0,0 +1,130 @@ +import { renderToStaticMarkup } from "react-dom/server"; +import { describe, expect, it } from "vitest"; +import { UI } from "@/lib/editor/sourceTextCopy"; +import type { EditorTool } from "./EditorOverlay"; +import { DEFAULT_EDITOR_TOOL, EditorToolbar, editorShortcut, type EditorToolbarProps } from "./EditorToolbar"; + +const noop = () => {}; + +function render(tool: EditorTool, showOriginal: Partial = {}): string { + return renderToStaticMarkup( + , + ); +} + +function buttons(markup: string): string[] { + return [...markup.matchAll(/]*>/g)].map((m) => m[0]); +} + +function attr(tag: string, name: string): string | undefined { + return tag.match(new RegExp(`\\s${name}="([^"]*)"`))?.[1]; +} + +function byLabel(markup: string, label: string): string | undefined { + return buttons(markup).find((b) => attr(b, "aria-label") === label); +} + +const KEY = { metaKey: false, ctrlKey: false, altKey: false }; + +describe("EditorToolbar", () => { + it("lists the tools in the spec order: Select · Hand · Edit text · Add text · Image · Draw · Link | Shapes | Redaction | Markup", () => { + const labels = buttons(render("select")) + .filter((b) => attr(b, "aria-pressed") !== undefined || attr(b, "aria-label") === "Shapes") + .map((b) => attr(b, "aria-label")); + expect(labels).toEqual([ + "Select", + "Hand", + "Edit text", + "Add text", + "Image", + "Draw", + "Link", + "Shapes", + "Redaction", + "Note", + "Highlight", + "Underline", + "Strikeout", + "Ink annot", + ]); + }); + + it("opens on Select (D34): the default tool is Select and it is the pressed button", () => { + expect(DEFAULT_EDITOR_TOOL).toBe("select"); + const markup = render(DEFAULT_EDITOR_TOOL); + expect(attr(byLabel(markup, "Select")!, "aria-pressed")).toBe("true"); + expect(attr(byLabel(markup, "Add text")!, "aria-pressed")).toBe("false"); + expect(attr(byLabel(markup, "Edit text")!, "aria-pressed")).toBe("false"); + }); + + it("titles and labels the two text tools from the copy deck", () => { + const markup = render("select"); + expect(attr(byLabel(markup, "Edit text")!, "title")).toBe(UI.tool.editText.title); + expect(attr(byLabel(markup, "Add text")!, "title")).toBe(UI.tool.addText.title); + expect(markup).not.toContain('aria-label="Text"'); + }); + + it("shows the Add text banner only while Add text is active", () => { + expect(render("text")).toContain(UI.banner.addText); + expect(render("select")).not.toContain(UI.banner.addText); + expect(render("editText")).not.toContain(UI.banner.addText); + }); + + it("announces E and O with aria-keyshortcuts (H for Hand)", () => { + const markup = render("editText", { visible: true }); + expect(attr(byLabel(markup, "Edit text")!, "aria-keyshortcuts")).toBe("E"); + expect(attr(byLabel(markup, "Hand")!, "aria-keyshortcuts")).toBe("H"); + const show = buttons(markup).find((b) => attr(b, "title") === UI.tool.showOriginal.title); + expect(show).toBeDefined(); + expect(attr(show!, "aria-keyshortcuts")).toBe("O"); + }); + + it("offers Show original only when the page has text changes, as a toggle", () => { + expect(render("select")).not.toContain(UI.tool.showOriginal.label); + const off = render("select", { visible: true, pressed: false }); + const on = render("select", { visible: true, pressed: true }); + const find = (m: string) => buttons(m).find((b) => attr(b, "title") === UI.tool.showOriginal.title)!; + expect(attr(find(off), "aria-pressed")).toBe("false"); + expect(attr(find(on), "aria-pressed")).toBe("true"); + expect(on).toContain(`>${UI.tool.showOriginal.label}`); + }); + + it("maps E to Edit text, H to Hand and O to Show original (only when offered)", () => { + expect(editorShortcut({ ...KEY, key: "e" }, false)).toEqual({ kind: "tool", tool: "editText" }); + expect(editorShortcut({ ...KEY, key: "E" }, false)).toEqual({ kind: "tool", tool: "editText" }); + expect(editorShortcut({ ...KEY, key: "h" }, false)).toEqual({ kind: "tool", tool: "hand" }); + expect(editorShortcut({ ...KEY, key: "o" }, true)).toEqual({ kind: "showOriginal" }); + expect(editorShortcut({ ...KEY, key: "o" }, false)).toBeNull(); + }); + + it("leaves modified keys alone (⌘E, Ctrl+O, Alt+E)", () => { + expect(editorShortcut({ ...KEY, key: "e", metaKey: true }, true)).toBeNull(); + expect(editorShortcut({ ...KEY, key: "o", ctrlKey: true }, true)).toBeNull(); + expect(editorShortcut({ ...KEY, key: "e", altKey: true }, true)).toBeNull(); + expect(editorShortcut({ ...KEY, key: "x" }, true)).toBeNull(); + }); +}); diff --git a/src/components/pdf/editor/EditorToolbar.tsx b/src/components/pdf/editor/EditorToolbar.tsx new file mode 100644 index 0000000..cf73c1f --- /dev/null +++ b/src/components/pdf/editor/EditorToolbar.tsx @@ -0,0 +1,191 @@ +/** + * Editor toolbar (extracted from `PdfEditorCanvas`): history and clipboard, + * the tools in the order Select · Hand · Edit text · Add text · Image · Draw · + * Link | Shapes | Redaction | Markup | Show original | zoom | pages, and the + * Add text banner. Keyboard shortcuts are resolved by `editorShortcut`. + */ +import { Alert } from "@/components/ui/Alert"; +import { Button } from "@/components/ui/Button"; +import { Icon, type IconName } from "@/components/ui/Icon"; +import { UI } from "@/lib/editor/sourceTextCopy"; +import type { EditorTool } from "./EditorOverlay"; +import { ShapePicker, SHAPE_TOOLS } from "./ShapePicker"; + +/** Edit PDF opens on Select, so a first click on existing words never drops a new box (D34). */ +export const DEFAULT_EDITOR_TOOL: EditorTool = "select"; + +interface ToolButton { + id: EditorTool; + label: string; + title: string; + icon: IconName; + shortcut?: string; +} + +export const MAIN_TOOLS: readonly ToolButton[] = [ + { id: "select", label: "Select", title: "Select", icon: "mousePointer" }, + { id: "hand", label: "Hand", title: "Hand — drag to slide the page (H, hold Space)", icon: "hand", shortcut: "H" }, + { id: "editText", label: UI.tool.editText.label, title: UI.tool.editText.title, icon: "textCursor", shortcut: "E" }, + { id: "text", label: UI.tool.addText.label, title: UI.tool.addText.title, icon: "type" }, + { id: "image", label: "Image", title: "Image", icon: "image" }, + { id: "ink", label: "Draw", title: "Draw", icon: "pencil" }, + { id: "link", label: "Link", title: "Link — draw a hotspot (does not open the address)", icon: "external" }, +]; + +const MARKUP_TOOLS: readonly ToolButton[] = [ + { id: "note", label: "Note", title: "Note", icon: "badge" }, + { id: "highlight", label: "Highlight", title: "Highlight", icon: "sparkles" }, + { id: "underline", label: "Underline", title: "Underline", icon: "type" }, + { id: "strikeout", label: "Strikeout", title: "Strikeout", icon: "slash" }, + { id: "markupInk", label: "Ink annot", title: "Ink annot", icon: "stamp" }, +]; + +export type EditorShortcut = { kind: "tool"; tool: EditorTool } | { kind: "showOriginal" }; + +/** Single-key editor shortcuts (not typing): H hand, E Edit text, O Show original (only when offered). */ +export function editorShortcut( + e: { key: string; metaKey: boolean; ctrlKey: boolean; altKey: boolean }, + canShowOriginal: boolean, +): EditorShortcut | null { + if (e.metaKey || e.ctrlKey || e.altKey) return null; + const key = e.key.toLowerCase(); + if (key === "h") return { kind: "tool", tool: "hand" }; + if (key === "e") return { kind: "tool", tool: "editText" }; + if (key === "o" && canShowOriginal) return { kind: "showOriginal" }; + return null; +} + +type ShapeId = (typeof SHAPE_TOOLS)[number]["id"]; + +export interface EditorToolbarProps { + tool: EditorTool; + onTool: (tool: EditorTool) => void; + onImage: () => void; + shapeOpen: boolean; + onShapeOpenChange: (open: boolean) => void; + lastShape: ShapeId; + onPickShape: (id: ShapeId) => void; + history: { + canUndo: boolean; + canRedo: boolean; + onUndo: () => void; + onRedo: () => void; + canCopy: boolean; + onCopy: () => void; + canPaste: boolean; + onPaste: () => void; + }; + /** Shown only when the current page has text changes. */ + showOriginal: { visible: boolean; pressed: boolean; onToggle: () => void }; + zoom: number; + /** −1 zoom out, 0 reset, +1 zoom in. */ + onZoom: (step: -1 | 0 | 1) => void; + pageIndex: number; + pageCount: number; + onPage: (delta: number) => void; +} + +function ToolToggle({ t, tool, onClick }: { t: ToolButton; tool: EditorTool; onClick: () => void }) { + return ( + + ); +} + +export function EditorToolbar(props: EditorToolbarProps) { + const { tool, onTool, history, showOriginal, pageIndex, pageCount } = props; + return ( + <> +
+ + + + + + {MAIN_TOOLS.map((t) => ( + (t.id === "image" ? props.onImage() : onTool(t.id))} /> + ))} + + + + + + {MARKUP_TOOLS.map((t) => ( + onTool(t.id)} /> + ))} + {showOriginal.visible && ( + <> + + + + )} + + + + + + + + + + + {pageIndex + 1} / {pageCount} + + + +
+ {tool === "text" && {UI.banner.addText}} + + ); +} diff --git a/src/components/pdf/editor/ObjectInspector.test.tsx b/src/components/pdf/editor/ObjectInspector.test.tsx index 5363ee9..7d4bb52 100644 --- a/src/components/pdf/editor/ObjectInspector.test.tsx +++ b/src/components/pdf/editor/ObjectInspector.test.tsx @@ -1,6 +1,8 @@ import { renderToStaticMarkup } from "react-dom/server"; import { describe, expect, it } from "vitest"; -import { makeRectObject, type EditObject } from "@/lib/editor"; +import { makeRectObject, makeSourceTextObject, type EditObject } from "@/lib/editor"; +import { UI } from "@/lib/editor/sourceTextCopy"; +import type { TextFont, TextRun } from "@/lib/types"; import { ObjectInspector } from "./ObjectInspector"; function redactObject(): EditObject { @@ -29,3 +31,115 @@ describe("R-OPACITY-UI redact inspector hides Opacity", () => { expect(rect).toContain('aria-label="Opacity"'); }); }); + +describe("ObjectInspector text change branch (§D.6.4)", () => { + const font: TextFont = { + key: "f1", + displayName: "Calibri", + familyHint: "sans", + embedded: true, + subset: true, + alphabet: " 0267Iceinov", + widths: [226, 507, 507, 507, 507, 252, 423, 498, 229, 525, 527, 498], + wordSpace: true, + }; + const run: TextRun = { + id: "t1:fp:0:10-20", + order: 0, + line: 0, + text: "Invoice 2026", + rect: { x: 72, y: 697, w: 60, h: 12 }, + origin: { x: 72, y: 700 }, + dir: { x: 1, y: 0 }, + ascent: 9, + descent: 3, + caretOffsets: [], + editable: true, + reason: null, + metrics: { + surface: ["f1"], + tfSize: 12, + effectiveSize: 12, + charSpacing: 0, + wordSpacing: 0, + hScale: 1, + textToUser: 1, + letterSpacingPt: 0, + spaceMode: "glyph", + kernSpace: -250, + originalWidth: 60, + visibleExtent: 400, + nextObstacle: null, + }, + style: { + fill: "#000000", + sizeChangeable: true, + colourChangeable: true, + face: "regular", + faces: { + regular: { available: true, surface: ["f1"] }, + bold: { available: true, surface: ["f1"] }, + italic: { available: false, surface: [] }, + boldItalic: { available: false, surface: [] }, + }, + }, + substituted: false, + }; + const change = makeSourceTextObject("st1", 0, run.rect, { + runId: run.id, + sourceFingerprint: "fp", + sourcePageIndex: 0, + originalText: "Invoice 2026", + text: "Invoice 2027", + style: { face: "bold", fill: "#c71c1c" }, + }); + + function render(r: TextRun | null = run): string { + return renderToStaticMarkup( + {}} + onReorder={() => {}} + sourceText={{ run: r, fonts: new Map([["f1", font]]), onEditLine: () => {}, onRestore: () => {} }} + />, + ); + } + + it("shows Text change with Original, New, Font, Letters, Style and the note", () => { + const markup = render(); + expect(markup).toContain(UI.inspector.header); + expect(markup).toContain(`
${UI.inspector.original}
Invoice 2026
`); + expect(markup).toContain(`
${UI.inspector.new}
Invoice 2027
`); + expect(markup).toContain("Calibri · embedded subset"); + expect(markup).toContain(`
${UI.bar.letters}
`); + expect(markup).toContain("space 0267Iceinov"); + expect(markup).toContain("12 pt · Bold · Red"); + expect(markup).toContain(UI.inspector.note); + expect(markup).toContain(UI.inspector.editLine); + expect(markup).toContain(UI.inspector.restore); + }); + + it("hides W/H, Rotation, Opacity and Layer", () => { + const markup = render(); + for (const hidden of ['aria-label="W"', 'aria-label="H"', "Rotation", "Opacity", "Layer", "Send to back"]) { + expect(markup).not.toContain(hidden); + } + }); + + it("still shows the change before the page's lines are read (no font row yet)", () => { + const markup = render(null); + expect(markup).toContain("Invoice 2027"); + expect(markup).not.toContain(`
${UI.inspector.font}
`); + }); + + it("names a removed line with the removal note", () => { + const removed = { ...change, text: "" }; + const markup = renderToStaticMarkup( + {}} + sourceText={{ run, fonts: new Map([["f1", font]]), onEditLine: () => {}, onRestore: () => {} }} />, + ); + expect(markup).toContain(UI.removal); + }); +}); diff --git a/src/components/pdf/editor/ObjectInspector.tsx b/src/components/pdf/editor/ObjectInspector.tsx index 1e60869..d93e0d4 100644 --- a/src/components/pdf/editor/ObjectInspector.tsx +++ b/src/components/pdf/editor/ObjectInspector.tsx @@ -1,7 +1,17 @@ -import { useEffect, useState } from "react"; -import type { EditObject, LayerDir } from "@/lib/editor"; -import { isClosedShapeObject, isMarkupObject, isNoneFill, sizeWithAspect, toCssHex } from "@/lib/editor"; +import type { EditObject, LayerDir, SourceTextObject } from "@/lib/editor"; +import { + TEXT_INKS, + isClosedShapeObject, + isMarkupObject, + isNoneFill, + sizeWithAspect, + surfaceFor, + toCssHex, +} from "@/lib/editor"; +import { UI, fontDescription, truncateCopy } from "@/lib/editor/sourceTextCopy"; +import type { TextFont, TextRun } from "@/lib/types"; import { Icon } from "@/components/ui/Icon"; +import { ColorField as SharedColorField, DraftNumber } from "./controls"; const PRESETS = [ "#111827", @@ -14,6 +24,19 @@ const PRESETS = [ export type ColorPickTarget = "color" | "fill" | "stroke"; +/** The object inspector's colour field: its presets, the first one when unreadable. */ +function ColorField(props: Omit[0], "presets" | "fallback">) { + return ; +} + +/** What the inspector needs to describe a text change (the run comes from the page's text). */ +export interface SourceTextInspectorProps { + run: TextRun | null; + fonts: Map; + onEditLine: () => void; + onRestore: () => void; +} + export function ObjectInspector({ obj, picking, @@ -23,6 +46,7 @@ export function ObjectInspector({ onChange, onPickFromPage, onReorder, + sourceText, }: { obj: EditObject; picking?: ColorPickTarget | null; @@ -34,7 +58,12 @@ export function ObjectInspector({ onChange: (patch: Partial) => void; onPickFromPage?: (target: ColorPickTarget) => void; onReorder?: (dir: LayerDir) => void; + /** Required to show a text change (`kind: "sourceText"`). */ + sourceText?: SourceTextInspectorProps; }) { + if (obj.kind === "sourceText") { + return sourceText ? : null; + } const opacity = "opacity" in obj && typeof obj.opacity === "number" ? obj.opacity : 1; const shape = isClosedShapeObject(obj) ? obj : null; const filled = !!shape && !isNoneFill(shape.fill); @@ -362,178 +391,85 @@ export function ObjectInspector({ ); } -function formatNum(n: number): string { - if (Number.isInteger(n)) return String(n); - return String(Math.round(n * 1000) / 1000); -} +const LETTERS_MAX_CHARS = 120; -function parseDraft(s: string): number | null { - const t = s.trim().replace(",", "."); - if (t === "" || t === "-" || t === "." || t === "-.") return null; - const n = Number(t); - return Number.isFinite(n) ? n : null; +function formatPt(n: number): string { + return `${Number(n.toFixed(2))} pt`; } -/** Word/Paint-style number field: empty while typing, commit on blur/Enter. */ -function DraftNumber({ - label, - hideLabel, - inline, - value, - min, - max, - suffix, - onCommit, -}: { - label: string; - hideLabel?: boolean; - /** W/H row: label | input | suffix as sibling grid cells. */ - inline?: boolean; - value: number; - min: number; - max: number; - suffix?: string; - onCommit: (n: number) => void; -}) { - const [focused, setFocused] = useState(false); - const [draft, setDraft] = useState(formatNum(value)); - useEffect(() => { - if (!focused) setDraft(formatNum(value)); - }, [value, focused]); - - const commit = (raw: string) => { - const n = parseDraft(raw); - if (n == null) { - setDraft(formatNum(value)); - return; - } - const clamped = Math.min(max, Math.max(min, n)); - onCommit(clamped); - setDraft(formatNum(clamped)); - }; - - const input = ( - setFocused(true)} - onChange={(e) => setDraft(e.target.value)} - onBlur={(e) => { - setFocused(false); - commit(e.target.value); - }} - onKeyDown={(e) => { - if (e.key === "Enter") (e.target as HTMLInputElement).blur(); - if (e.key === "Escape") { - setDraft(formatNum(value)); - (e.target as HTMLInputElement).blur(); - } - }} - /> - ); - const suf = suffix ? ( - - {suffix} - - ) : null; - - if (inline) { - return ( - <> - {label} - {input} - {suf} - - ); +/** The letters a face can type: space named, the rest in font order, cut at 120. */ +function lettersOf(run: TextRun, fonts: Map, face: SourceTextObject["style"]["face"]): string { + const seen = new Set(); + for (const key of surfaceFor(run, face)) { + for (const ch of Array.from(fonts.get(key)?.alphabet ?? "")) seen.add(ch); } + const letters = [...seen].filter((ch) => ch !== " ").join(""); + return truncateCopy(seen.has(" ") ? `${UI.spaceName} ${letters}` : letters, LETTERS_MAX_CHARS); +} - return ( -
- {!hideLabel && } -
- {input} - {suf} -
-
- ); +function styleSummary(obj: SourceTextObject, run: TextRun | null): string { + const parts: string[] = []; + const size = obj.style.sizePt ?? run?.metrics?.effectiveSize; + if (size !== undefined) parts.push(formatPt(size)); + const face = obj.style.face ?? run?.style?.face; + if (face) parts.push(UI.faceLabel[face]); + const fill = (obj.style.fill ?? run?.style?.fill ?? "").toLowerCase(); + const ink = TEXT_INKS.find((i) => i.hex === fill); + if (ink) parts.push(UI.bar[ink.key]); + else if (fill) parts.push(fill); + const spacing = obj.style.letterSpacingPt; + if (spacing !== undefined) parts.push(`${UI.bar.letterSpacing} ${formatPt(spacing)}`); + return parts.join(" · "); } -function ColorField({ - label, - icon, - value, - active, - onChange, - onPickFromPage, -}: { - label: string; - icon?: "droplet" | "square" | "squareFill"; - value: string; - active?: boolean; - onChange: (hex: string) => void; - onPickFromPage?: () => void; -}) { - const hex = toCssHex(value, "#111827"); - const [typed, setTyped] = useState(hex); - useEffect(() => { - setTyped(hex); - }, [hex]); +/** Inspector branch for a text change (§D.6.4): read-only rows and two actions. */ +function SourceTextInspector({ + obj, + run, + fonts, + onEditLine, + onRestore, +}: { obj: SourceTextObject } & SourceTextInspectorProps) { + const primary = run ? fonts.get(surfaceFor(run, obj.style.face)[0] ?? "") : undefined; + const subset = !!primary && primary.embedded && primary.subset; + const summary = styleSummary(obj, run); return ( -
-
- {icon ? ( - - - - ) : ( - +
+
{UI.inspector.header}
+
+
{UI.inspector.original}
+
{obj.originalText}
+
{UI.inspector.new}
+
{obj.text === "" ? UI.removal : obj.text}
+ {primary && ( + <> +
{UI.inspector.font}
+
{fontDescription(primary)}
+ )} - onChange(e.target.value)} - className="pdf-editor__color-input" - /> - { - const v = e.target.value.trim(); - setTyped(v.startsWith("#") || v.length === 0 ? v : `#${v}`); - const next = v.startsWith("#") ? v : `#${v}`; - if (/^#[0-9a-fA-F]{6}$/.test(next)) onChange(next.toLowerCase()); - else if (/^#[0-9a-fA-F]{3}$/.test(next)) onChange(toCssHex(next)); - }} - /> - {onPickFromPage && ( - + {subset && run && ( + <> +
{UI.bar.letters}
+
{lettersOf(run, fonts, obj.style.face)}
+ )} -
-
- {PRESETS.map((c) => ( - +
); diff --git a/src/components/pdf/editor/ObjectList.test.tsx b/src/components/pdf/editor/ObjectList.test.tsx index a410ee1..863039f 100644 --- a/src/components/pdf/editor/ObjectList.test.tsx +++ b/src/components/pdf/editor/ObjectList.test.tsx @@ -1,6 +1,7 @@ import { renderToStaticMarkup } from "react-dom/server"; import { describe, expect, it } from "vitest"; -import { makeRectObject, type EditObject } from "@/lib/editor"; +import { makeRectObject, makeSourceTextObject, type EditObject } from "@/lib/editor"; +import { UI, fillCopy } from "@/lib/editor/sourceTextCopy"; import { TOOLS } from "@/lib/tools"; import { ObjectList } from "./ObjectList"; @@ -40,3 +41,41 @@ describe("ObjectList redaction vs rectangle", () => { expect(TOOLS.some((t) => t.id === "editPdf")).toBe(true); }); }); + +describe("ObjectList text changes", () => { + const change = (text: string, pageIndex = 2): EditObject => + makeSourceTextObject(`st-${text}`, pageIndex, { x: 72, y: 700, w: 40, h: 12 }, { + runId: "t1:fp:2:10-20", + sourceFingerprint: "fp", + sourcePageIndex: 2, + originalText: "Invoice 2026", + text, + style: {}, + }); + + it("labels a text change `Edited text: “…” · pN`, never Rectangle or a layer", () => { + const markup = renderToStaticMarkup( + {}} onDelete={() => {}} />, + ); + expect(markup).toContain(fillCopy(UI.list.label, { new: "Invoice 2027", page: 3 })); + expect(markup).toContain("Edited text: “Invoice 2027” · p3"); + expect(markup).not.toContain("Rectangle"); + expect(markup).not.toContain("1/1"); + }); + + it("cuts long new text at 28 characters", () => { + const long = "A".repeat(40); + const markup = renderToStaticMarkup( + {}} onDelete={() => {}} />, + ); + expect(markup).toContain(`Edited text: “${"A".repeat(28)}…” · p3`); + }); + + it("does not count text changes as layers of the other objects", () => { + const rect = makeRectObject("box", 2, { x: 10, y: 20, w: 100, h: 50 }); + const markup = renderToStaticMarkup( + {}} onDelete={() => {}} />, + ); + expect(markup).toContain("Rectangle · p3 · 1/1"); + }); +}); diff --git a/src/components/pdf/editor/ObjectList.tsx b/src/components/pdf/editor/ObjectList.tsx index 9fb66a0..789033d 100644 --- a/src/components/pdf/editor/ObjectList.tsx +++ b/src/components/pdf/editor/ObjectList.tsx @@ -1,7 +1,10 @@ import { isNearlySquare, type EditObject } from "@/lib/editor"; +import { listLabel } from "@/lib/editor/sourceTextCopy"; function labelFor(obj: EditObject, layer: string): string { const page = `p${obj.pageIndex + 1}`; + // A text change is not a layer object: no layer number, never "Rectangle". + if (obj.kind === "sourceText") return listLabel(obj.text, obj.pageIndex + 1); if (obj.kind === "text") { const t = obj.content.trim() || "Text"; return `Text: ${t.length > 18 ? `${t.slice(0, 18)}…` : t} · ${page} · ${layer}`; @@ -57,7 +60,7 @@ export function ObjectList({
    {objects.map((obj) => { const selected = selectedIds.includes(obj.id); - const onPage = objects.filter((o) => o.pageIndex === obj.pageIndex); + const onPage = objects.filter((o) => o.pageIndex === obj.pageIndex && o.kind !== "sourceText"); const z = onPage.findIndex((o) => o.id === obj.id) + 1; const layer = `${z}/${onPage.length}`; return ( diff --git a/src/components/pdf/editor/PageSurface.tsx b/src/components/pdf/editor/PageSurface.tsx index c126d94..51bbbf6 100644 --- a/src/components/pdf/editor/PageSurface.tsx +++ b/src/components/pdf/editor/PageSurface.tsx @@ -6,6 +6,11 @@ * `fitWidth` must be the *stage* content width (unzoomed fit), NOT the page * element’s current CSS width — otherwise zoom compounds and/or paint races * leave a blank canvas. + * + * Painting is double-buffered: each render goes to an offscreen canvas and is + * copied to the visible one in a single step when it completes, so swapping + * `bytes` (original ⇄ Edit text preview of the same page) or zooming never + * flashes a blank page. */ import { useEffect, useRef, useState, type RefObject } from "react"; import { pdfjsLib, PDF_OPTS } from "@/lib/pdfjs"; @@ -175,7 +180,10 @@ export function PageSurface({ } const viewport = page.getViewport({ scale: renderScale }); - const ctx = canvas.getContext("2d"); + const buffer = document.createElement("canvas"); + buffer.width = Math.ceil(viewport.width); + buffer.height = Math.ceil(viewport.height); + const ctx = buffer.getContext("2d"); if (!ctx) return; try { @@ -184,18 +192,22 @@ export function PageSurface({ /* ignore */ } - // Clear then size — avoids flashing previous zoom level. - canvas.width = Math.ceil(viewport.width); - canvas.height = Math.ceil(viewport.height); - canvas.style.width = `${cssWidth}px`; - canvas.style.height = `${cssHeight}px`; - const task = page.render({ canvasContext: ctx, viewport }); taskRef.current = task; await task.promise; if (cancelled || gen !== paintGen.current) return; + // Swap in one step: resizing clears the visible canvas, and the copy + // lands before the browser can paint the cleared state. + const visible = canvas.getContext("2d"); + if (!visible) return; + canvas.width = buffer.width; + canvas.height = buffer.height; + canvas.style.width = `${cssWidth}px`; + canvas.style.height = `${cssHeight}px`; + visible.drawImage(buffer, 0, 0); + onLayoutRef.current({ cssWidth, cssHeight, diff --git a/src/components/pdf/editor/PdfEditorCanvas.focus.test.tsx b/src/components/pdf/editor/PdfEditorCanvas.focus.test.tsx new file mode 100644 index 0000000..40bc79b --- /dev/null +++ b/src/components/pdf/editor/PdfEditorCanvas.focus.test.tsx @@ -0,0 +1,141 @@ +// @vitest-environment happy-dom +// Live-check regression (v0.4 Edit text): moving focus to the editor root must never scroll the +// window. With a scroll, the first click on a line moved the page between pointerdown and click, +// so the click opened the line below it (or nothing), and Esc on the line layer jumped the page. +import { act, useMemo, useRef } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import type { PageText, TextRun } from "@/lib/types"; +import { ToastProvider } from "@/components/ui/Toast"; +import type { TextSources } from "@/features/edit-pdf/useTextSources"; + +const LAYOUT = { cssWidth: 612, cssHeight: 792, geometry: { box: { x: 0, y: 0, w: 612, h: 792 }, rotate: 0, pageIndex: 0 } }; + +vi.mock("./PageSurface", async () => { + const { useEffect } = await import("react"); + return { + PageSurface: ({ onLayout }: { onLayout: (l: typeof LAYOUT) => void }) => { + useEffect(() => onLayout(LAYOUT), []); + return ; + }, + }; +}); + +const RUN: TextRun = { + id: "r1", + order: 0, + line: 0, + text: "Hello", + rect: { x: 72, y: 700, w: 30, h: 12 }, + origin: { x: 72, y: 703 }, + dir: { x: 1, y: 0 }, + ascent: 9, + descent: 3, + caretOffsets: [0, 6, 12, 18, 24, 30], + editable: false, + reason: "ROTATED_TEXT", + metrics: null, + style: null, + substituted: false, +}; +const PAGE: PageText = { fingerprint: "fp", pageIndex: 0, pageReason: null, runs: [RUN], fonts: [] }; + +vi.mock("@/lib/tauriCommands", () => ({ + pagePdf: vi.fn(async () => "JVBERg=="), + listPdfAnnots: vi.fn(async () => []), + pickImageFile: vi.fn(async () => null), + previewImage: vi.fn(), + inspectTextPage: vi.fn(async () => PAGE), + previewTextEdits: vi.fn(), +})); + +const { PdfEditorCanvas } = await import("./PdfEditorCanvas"); +const { useEditSession } = await import("./useEditSession"); + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +const INFO = { fingerprint: "fp", pageCount: 1, warnings: [] }; +const SOURCES: TextSources = { + sources: { "/a.pdf": { status: "ready", info: INFO, error: null, stale: false } }, + ensure: async () => INFO, + release: () => {}, + markStale: () => {}, + markError: () => {}, + reopen: async () => INFO, +}; + +function Harness() { + const keys = useMemo(() => ["u1#1"], []); + const session = useEditSession(keys); + const guardRef = useRef(null); + return ( + + + + ); +} + +let host: HTMLDivElement; +let root: Root | null = null; +let focusSpy: ReturnType; + +async function flush() { + for (let i = 0; i < 5; i++) await act(async () => {}); +} + +beforeEach(async () => { + focusSpy = vi.spyOn(HTMLElement.prototype, "focus"); + host = document.createElement("div"); + document.body.appendChild(host); + root = createRoot(host); + await act(async () => root?.render()); + await flush(); +}); + +afterEach(() => { + act(() => root?.unmount()); + root = null; + host.remove(); + focusSpy.mockRestore(); +}); + +function rootFocusCalls(): unknown[][] { + const editor = host.querySelector(".pdf-editor"); + return focusSpy.mock.calls.filter((_args, i) => focusSpy.mock.contexts[i] === editor); +} + +describe("PdfEditorCanvas focus never scrolls the page", () => { + it("pointerdown on the stage focuses the editor root with preventScroll", async () => { + const stage = host.querySelector(".pdf-editor__stage"); + expect(stage).not.toBeNull(); + await act(async () => { + stage?.dispatchEvent(new PointerEvent("pointerdown", { bubbles: true, button: 0 })); + }); + const calls = rootFocusCalls(); + expect(calls.length).toBeGreaterThan(0); + for (const args of calls) expect(args[0]).toEqual({ preventScroll: true }); + }); + + it("Esc on the Edit text layer returns focus to the editor root with preventScroll", async () => { + const editTool = host.querySelector('button[aria-keyshortcuts="E"]'); + expect(editTool).not.toBeNull(); + await act(async () => editTool?.click()); + await flush(); + const line = host.querySelector(".st-layer button.st-run"); + expect(line).not.toBeNull(); + focusSpy.mockClear(); + await act(async () => { + line?.dispatchEvent(new KeyboardEvent("keydown", { key: "Escape", bubbles: true })); + }); + const calls = rootFocusCalls(); + expect(calls.length).toBe(1); + expect(calls[0][0]).toEqual({ preventScroll: true }); + }); +}); diff --git a/src/components/pdf/editor/PdfEditorCanvas.polish.test.tsx b/src/components/pdf/editor/PdfEditorCanvas.polish.test.tsx new file mode 100644 index 0000000..6487369 --- /dev/null +++ b/src/components/pdf/editor/PdfEditorCanvas.polish.test.tsx @@ -0,0 +1,220 @@ +// @vitest-environment happy-dom +// Live-check polish (v0.4 Edit text), through the real canvas, session and toolbar: +// - Show original is offered only when the page shows a real preview of its text changes +// (it used to toggle between two identical renders when the preview was unavailable); +// - "Add text here" on a line turned at an angle adds a default one-line box at the line's +// start (it used to add the line's whole axis-aligned box: 298 × 260 pt for a watermark). +import { act, useEffect, useMemo, useRef } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import type { PageText, TextPreview, TextRun } from "@/lib/types"; +import { UI } from "@/lib/editor/sourceTextCopy"; +import { ToastProvider } from "@/components/ui/Toast"; +import type { TextSources } from "@/features/edit-pdf/useTextSources"; +import type { EditSession } from "./useEditSession"; + +const LAYOUT = { cssWidth: 612, cssHeight: 792, geometry: { box: { x: 0, y: 0, w: 612, h: 792 }, rotate: 0, pageIndex: 0 } }; + +vi.mock("./PageSurface", async () => { + const { useEffect: useMountEffect } = await import("react"); + return { + PageSurface: ({ onLayout }: { onLayout: (l: typeof LAYOUT) => void }) => { + useMountEffect(() => onLayout(LAYOUT), []); + return ; + }, + }; +}); + +const HELLO: TextRun = { + id: "r1", + order: 0, + line: 0, + text: "Hello", + rect: { x: 72, y: 700, w: 30, h: 12 }, + origin: { x: 72, y: 703 }, + dir: { x: 1, y: 0 }, + ascent: 9, + descent: 3, + caretOffsets: [0, 6, 12, 18, 24, 30], + editable: true, + reason: null, + metrics: { + surface: ["f1"], tfSize: 12, effectiveSize: 12, charSpacing: 0, wordSpacing: 0, hScale: 1, textToUser: 1, + letterSpacingPt: 0, spaceMode: "glyph", kernSpace: -250, originalWidth: 30, visibleExtent: 500, nextObstacle: null, + }, + style: { + fill: "#000000", sizeChangeable: true, colourChangeable: true, face: "regular", + faces: { + regular: { available: true, surface: ["f1"] }, bold: { available: false, surface: [] }, + italic: { available: false, surface: [] }, boldItalic: { available: false, surface: [] }, + }, + }, + substituted: false, +}; +const DEG = (35 * Math.PI) / 180; +/** A refused 35° watermark near the top of the page (inside the layer's rendered band in happy-dom). */ +const WATERMARK: TextRun = { + ...HELLO, + id: "r2", + order: 1, + line: 1, + text: "CONFIDENTIAL", + rect: { x: 150, y: 500, w: 298, h: 260 }, + origin: { x: 160, y: 512 }, + dir: { x: Math.cos(DEG), y: Math.sin(DEG) }, + caretOffsets: [], + editable: false, + reason: "ROTATED_TEXT", + metrics: null, + style: null, +}; +const PAGE: PageText = { + fingerprint: "fp", + pageIndex: 0, + pageReason: null, + runs: [HELLO, WATERMARK], + fonts: [{ key: "f1", displayName: "Arial", familyHint: "sans", embedded: true, subset: false, alphabet: " !Hdelo", widths: [278, 278, 722, 556, 556, 222, 556], wordSpace: true }], +}; + +const previewTextEdits = vi.fn<(...args: unknown[]) => Promise>(); +vi.mock("@/lib/tauriCommands", () => ({ + pagePdf: vi.fn(async () => "JVBERg=="), + listPdfAnnots: vi.fn(async () => []), + pickImageFile: vi.fn(async () => null), + previewImage: vi.fn(), + inspectTextPage: vi.fn(async () => PAGE), + previewTextEdits: (...args: unknown[]) => previewTextEdits(...args), +})); + +const { PdfEditorCanvas } = await import("./PdfEditorCanvas"); +const { useEditSession } = await import("./useEditSession"); + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +const INFO = { fingerprint: "fp", pageCount: 1, warnings: [] }; +const SOURCES: TextSources = { + sources: { "/a.pdf": { status: "ready", info: INFO, error: null, stale: false } }, + ensure: async () => INFO, + release: () => {}, + markStale: () => {}, + markError: () => {}, + reopen: async () => INFO, +}; + +function answer(pagePdf: string | null): TextPreview { + return { + pagePdf, + verdicts: [{ runId: "r1", ok: true, code: null, chars: [], reason: null, face: null, field: null, detail: null, deltaPt: 3, newRect: HELLO.rect, caretOffsets: [0, 6, 12, 18, 24, 30, 33] }], + pageProblem: null, + warnings: [], + }; +} + +let latest: EditSession | null = null; + +function Harness({ withChange }: { withChange: boolean }) { + const keys = useMemo(() => ["u1#1"], []); + const session = useEditSession(keys); + const guardRef = useRef(null); + const committed = useRef(false); + latest = session; + useEffect(() => { + // Once, like a committed edit. + if (!withChange || committed.current) return; + committed.current = true; + session.setSourceText({ pageIndex: 0, run: HELLO, sourceFingerprint: "fp", sourcePageIndex: 0, text: "Hello!", style: {} }); + }, [withChange, session]); + return ( + + + + ); +} + +let host: HTMLDivElement; +let root: Root | null = null; + +async function flush() { + for (let i = 0; i < 6; i++) await act(async () => {}); +} + +async function mount(withChange: boolean) { + host = document.createElement("div"); + document.body.appendChild(host); + root = createRoot(host); + await act(async () => root?.render()); + await flush(); +} + +afterEach(() => { + act(() => root?.unmount()); + root = null; + host.remove(); + latest = null; + previewTextEdits.mockReset(); +}); + +const showOriginalButton = () => + [...host.querySelectorAll("button")].find((b) => b.textContent === UI.tool.showOriginal.label) ?? null; +const statusChip = () => host.querySelector(".st-chip-status")?.textContent ?? null; +const pressO = () => + act(async () => { + host.querySelector(".pdf-editor")?.focus(); + window.dispatchEvent(new KeyboardEvent("keydown", { key: "o", bubbles: true })); + }); + +describe("Show original is offered only with a preview to compare", () => { + it("is offered (button and O) when the page shows a preview of its text changes", async () => { + previewTextEdits.mockResolvedValue(answer("JVBERg==")); + await mount(true); + expect(statusChip()).toBe(UI.status.showingChanges); + expect(showOriginalButton()).not.toBeNull(); + await pressO(); + expect(statusChip()).toBe(UI.status.showingOriginal); + }); + + it("is not offered when the preview is unavailable for the page (nothing to compare)", async () => { + previewTextEdits.mockResolvedValue(answer(null)); + await mount(true); + expect(statusChip()).toBe(UI.status.previewUnavailable); + expect(showOriginalButton()).toBeNull(); + await pressO(); + expect(statusChip()).toBe(UI.status.previewUnavailable); + }); + + it("is not offered while the page has no text changes", async () => { + await mount(false); + expect(showOriginalButton()).toBeNull(); + expect(previewTextEdits).not.toHaveBeenCalled(); + }); +}); + +describe("Add text here", () => { + it("on a line turned at an angle adds a default one-line box at its start, upright, in Add text", async () => { + await mount(false); + await act(async () => host.querySelector('button[aria-keyshortcuts="E"]')?.click()); + await flush(); + const line = [...host.querySelectorAll("button.st-run")].find((b) => b.getAttribute("aria-label") === "CONFIDENTIAL"); + expect(line).toBeDefined(); + await act(async () => line?.click()); + const add = [...host.querySelectorAll('[role="dialog"] button')].find((b) => b.textContent === UI.popover.addTextHere); + await act(async () => add?.click()); + await flush(); + const boxes = (latest?.objects ?? []).filter((o) => o.kind === "text"); + expect(boxes).toHaveLength(1); + const box = boxes[0] as Extract<(typeof boxes)[number], { kind: "text" }>; + expect(box.fontSize).toBe(12); + expect(box.rect.w).toBeCloseTo(120, 6); // not 298 + expect(box.rect.h).toBeCloseTo(15.6, 6); // not 260 + expect(box.rect.x).toBeCloseTo(160, 6); + expect(box.rect.y + box.rect.h).toBeCloseTo(524, 6); + expect(host.querySelector(`button[title="${UI.tool.addText.title}"]`)?.getAttribute("aria-pressed")).toBe("true"); + }); +}); diff --git a/src/components/pdf/editor/PdfEditorCanvas.tsx b/src/components/pdf/editor/PdfEditorCanvas.tsx index f6223f5..256d09c 100644 --- a/src/components/pdf/editor/PdfEditorCanvas.tsx +++ b/src/components/pdf/editor/PdfEditorCanvas.tsx @@ -10,8 +10,7 @@ import { useState, type PointerEvent as ReactPointerEvent, } from "react"; -import { Button } from "@/components/ui/Button"; -import { Icon, type IconName } from "@/components/ui/Icon"; +import { Icon } from "@/components/ui/Icon"; import { Spinner } from "@/components/ui/Spinner"; import { Alert } from "@/components/ui/Alert"; import { useToast } from "@/components/ui/Toast"; @@ -20,6 +19,9 @@ import { toAppError } from "@/lib/types"; import { base64ToBytes } from "@/lib/pdfjs"; import type { EditObject, FormField, ShapeStyle } from "@/lib/editor"; import type { ListedMarkup } from "@/lib/types"; +import type { TextStamp } from "@/lib/editor/sourceText"; +import { UI } from "@/lib/editor/sourceTextCopy"; +import type { TextSources } from "@/features/edit-pdf/useTextSources"; import { cloneObject, isClosedShapeObject, @@ -27,6 +29,7 @@ import { offsetObject, pdfRectToViewport, placeImagePdfRect, + editsSignature, rgbToHex, selectedIdsOnPage, stageJustify, @@ -36,8 +39,17 @@ import { EditorOverlay, type EditorTool } from "./EditorOverlay"; import { FormFieldsOverlay } from "./FormFieldsOverlay"; import { ObjectList } from "./ObjectList"; import { ObjectInspector, type ColorPickTarget } from "./ObjectInspector"; -import { ShapePicker, SHAPE_TOOLS } from "./ShapePicker"; +import type { SHAPE_TOOLS } from "./ShapePicker"; import type { EditSession } from "./useEditSession"; +import { DEFAULT_EDITOR_TOOL, EditorToolbar, editorShortcut } from "./EditorToolbar"; +import { + PreviewStatusChip, + SourceTextBanners, + SourceTextMode, + useSourceTextPage, + type SourceTextGuard, + type TextRequest, +} from "./sourceText/SourceTextMode"; const MIN_ZOOM = 0.5; const MAX_ZOOM = 4; @@ -49,26 +61,10 @@ function newObjectId(): string { return `obj-${Date.now()}-${Math.random().toString(36).slice(2, 9)}`; } -const MAIN_TOOLS: { - id: EditorTool; - label: string; - icon: "mousePointer" | "hand" | "type" | "image" | "pencil" | "external"; -}[] = [ - { id: "select", label: "Select", icon: "mousePointer" }, - { id: "hand", label: "Hand", icon: "hand" }, - { id: "text", label: "Text", icon: "type" }, - { id: "image", label: "Image", icon: "image" }, - { id: "ink", label: "Draw", icon: "pencil" }, - { id: "link", label: "Link", icon: "external" }, -]; - -const MARKUP_TOOLS: { id: EditorTool; label: string; icon: IconName }[] = [ - { id: "note", label: "Note", icon: "badge" }, - { id: "highlight", label: "Highlight", icon: "sparkles" }, - { id: "underline", label: "Underline", icon: "type" }, - { id: "strikeout", label: "Strikeout", icon: "slash" }, - { id: "markupInk", label: "Ink annot", icon: "stamp" }, -]; +/** Inside the Edit text layer, inline editor, format bar or reason popover. */ +function inSourceText(t: EventTarget | null): boolean { + return t instanceof Element && !!t.closest("[data-source-text]"); +} function isTextEntryTarget(t: EventTarget | null): boolean { const el = t as HTMLElement | null; @@ -77,6 +73,16 @@ function isTextEntryTarget(t: EventTarget | null): boolean { return tag === "INPUT" || tag === "TEXTAREA" || tag === "SELECT" || !!el.isContentEditable; } +/** What the canvas needs for Edit text (owned by `EditPdfPage`). */ +export interface CanvasTextProps { + sources: TextSources; + /** Display name of `sourcePath`. */ + fileName: string; + /** The shown source page is listed more than once. */ + duplicatePage: boolean; + guardRef: { current: SourceTextGuard | null }; // the open line edit, so Save can finish it first +} + export function PdfEditorCanvas({ sourcePath, sourcePage, @@ -87,6 +93,7 @@ export function PdfEditorCanvas({ formFields = [], formValues = {}, onFormChange, + text, }: { sourcePath: string; /** 1-based page number inside `sourcePath` (pagePdf). */ @@ -99,6 +106,7 @@ export function PdfEditorCanvas({ formFields?: FormField[]; formValues?: Record; onFormChange?: (name: string, value: string) => void; + text: CanvasTextProps; }) { const { toast } = useToast(); const [zoom, setZoom] = useState(1); @@ -106,7 +114,12 @@ export function PdfEditorCanvas({ const [loadError, setLoadError] = useState(null); const [loading, setLoading] = useState(false); const [layout, setLayout] = useState(null); - const [tool, setTool] = useState("text"); + const [tool, setTool] = useState(DEFAULT_EDITOR_TOOL); + /** Show original is on for one set of changes; a commit, restore or undo turns it off. */ + const [originalFor, setOriginalFor] = useState(null); + const [modeDismissed, setModeDismissed] = useState(false); + const [textRequest, setTextRequest] = useState(null); + const textGuard = text.guardRef; const [spacePan, setSpacePan] = useState(false); const [panning, setPanning] = useState(false); const [shapeOpen, setShapeOpen] = useState(false); @@ -139,13 +152,61 @@ export function PdfEditorCanvas({ () => selectedIdsOnPage(session.objects, session.selectedIds, pageIndex), [pageIndex, session.objects, session.selectedIds], ); + const st = useSourceTextPage({ + objects: session.objects, + path: sourcePath, + fileName: text.fileName, + sourcePage, + pageIndex, + sources: text.sources, + active: tool === "editText", + duplicate: text.duplicatePage, + }); + const hasTextChanges = st.objects.length > 0; + const textSig = editsSignature(st.objects); + const original = st.canShowOriginal && originalFor === textSig; + const toggleOriginal = () => setOriginalFor((at) => (at === textSig ? null : textSig)); + const surfaceBytes = bytes && (original || !st.preview.bytes ? bytes : st.preview.bytes); useEffect(() => { setEditingTextId(null); setColorPick(null); setPickCursor(null); + setOriginalFor(null); + setTextRequest(null); }, [pageIndex]); + /** An open line edit must finish (or stay, with its message) before the page or tool changes. */ + const finishTextEdit = async (): Promise => { + const guard = textGuard.current; + if (!guard?.isEditing() || (await guard.tryClose())) return true; + toast({ title: UI.navBlocked, variant: "error" }); + return false; + }; + + const requestTool = async (next: EditorTool) => { + if (next !== tool && (await finishTextEdit())) setTool(next); + }; + + /** Switch to Edit text and focus (or open) one line, from the sidebar or inspector. */ + const showTextLine = (runId: string, open: boolean) => { + setTool("editText"); + setTextRequest((r) => ({ runId, open, tick: (r?.tick ?? 0) + 1 })); + }; + + /** A drawn object is placed: back to Select. */ + const thenSelect = (place: (...args: A) => void) => (...args: A) => { + place(...args); + setTool("select"); + }; + + const addTextHere = (stamp: TextStamp) => { + const id = session.addTextStamp(pageIndex, stamp.rect, stamp.fontSize); + session.select([id]); + setTool("text"); + setEditingTextId(id); + }; + useEffect(() => { const el = stageRef.current; if (!el) return; @@ -215,10 +276,10 @@ export function PdfEditorCanvas({ }); }; - const go = (delta: number) => { + const go = async (delta: number) => { if (!onPageChange) return; const next = Math.min(pageCount - 1, Math.max(0, pageIndex + delta)); - onPageChange(next); + if (next !== pageIndex && (await finishTextEdit())) onPageChange(next); }; const placeImage = async (atCss?: { x: number; y: number }) => { @@ -245,7 +306,8 @@ export function PdfEditorCanvas({ const copySelection = useCallback(() => { const activeIds = new Set(selectedIdsOnPage(session.objects, session.selectedIds, pageIndex)); - const sel = session.objects.filter((object) => activeIds.has(object.id)); + // Text changes are never copied (one per line, locked to it). + const sel = session.objects.filter((object) => activeIds.has(object.id) && object.kind !== "sourceText"); if (sel.length === 0) return false; clipboardRef.current = sel.map(cloneObject); pasteGen.current = 1; @@ -279,9 +341,11 @@ export function PdfEditorCanvas({ } const mod = e.metaKey || e.ctrlKey; - if (!mod && e.key.toLowerCase() === "h") { + const shortcut = t?.closest?.(".st-editor, .st-popover") ? null : editorShortcut(e, st.canShowOriginal); + if (shortcut) { e.preventDefault(); - setTool("hand"); + if (shortcut.kind === "tool") void requestTool(shortcut.tool); + else toggleOriginal(); return; } if (mod && e.key.toLowerCase() === "z" && !e.shiftKey) { @@ -307,10 +371,11 @@ export function PdfEditorCanvas({ if (mod && e.key.toLowerCase() === "d") { if (pageSelectedIds.length === 0) return; e.preventDefault(); - copySelection(); - pasteClipboard(); + if (copySelection()) pasteClipboard(); return; } + // The Edit text layer owns arrows, Enter, Delete and Esc while it has focus. + if (inSourceText(t) && /^(Arrow|Enter$|Escape$|Delete$|Backspace$|Home$|End$|Page)/.test(e.key)) return; if (e.key === "Escape") { if (shapeOpen) { setShapeOpen(false); @@ -349,7 +414,7 @@ export function PdfEditorCanvas({ }; window.addEventListener("keydown", onKey); return () => window.removeEventListener("keydown", onKey); - }, [session, colorPick, copySelection, pageIndex, pageSelectedIds, pasteClipboard, shapeOpen]); + }, [session, colorPick, copySelection, pageIndex, pageSelectedIds, pasteClipboard, shapeOpen, st.canShowOriginal, requestTool]); useEffect(() => { const stopSpacePan = () => { @@ -360,6 +425,8 @@ export function PdfEditorCanvas({ const onKeyDown = (e: KeyboardEvent) => { if (e.code !== "Space" && e.key !== " ") return; if (isTextEntryTarget(e.target) || isTextEntryTarget(document.activeElement)) return; + // Space activates a focused line or button of Edit text; it never pans there. + if (inSourceText(e.target) || inSourceText(document.activeElement)) return; const ae = document.activeElement; const root = rootRef.current; const stage = stageRef.current; @@ -384,10 +451,12 @@ export function PdfEditorCanvas({ }, []); const selected = pageObjects.find((object) => object.id === pageSelectedIds[0]) ?? null; + const layerObjects = pageObjects.filter((object) => object.kind !== "sourceText"); const panMode = !colorPick && (tool === "hand" || spacePan); const onStagePointerDown = (e: ReactPointerEvent) => { - rootRef.current?.focus(); + // The inline editor and the reason popover keep their own focus. + if (!(e.target instanceof Element && e.target.closest(".st-editor, .st-popover"))) rootRef.current?.focus({ preventScroll: true }); if (!panMode) return; if (e.button !== 0) return; const el = stageRef.current; @@ -454,130 +523,38 @@ export function PdfEditorCanvas({ Redaction permanently removes content on Save. Only pages with a redaction region become images; text on those pages will not stay selectable. + {hasTextChanges && ` ${UI.redactionConflict}`}
)} -
- - - - - - {MAIN_TOOLS.filter((t) => t.id !== "ink" && t.id !== "link").map((t) => ( - - ))} - { - setLastShape(id); - setTool(id); - }} - /> - - - - {MARKUP_TOOLS.map((t) => ( - - ))} - - - - - - - - {pageIndex + 1} / {pageCount} - - -
+ void requestTool(next)} + onImage={() => void finishTextEdit().then((ok) => (ok ? placeImage() : undefined))} + shapeOpen={shapeOpen} + onShapeOpenChange={setShapeOpen} + lastShape={lastShape} + onPickShape={(id) => { + setLastShape(id); + void requestTool(id); + }} + history={{ + canUndo: session.canUndo, + canRedo: session.canRedo, + onUndo: session.undo, + onRedo: session.redo, + canCopy: pageSelectedIds.length > 0, + onCopy: () => void copySelection(), + canPaste, + onPaste: pasteClipboard, + }} + showOriginal={{ visible: st.canShowOriginal, pressed: original, onToggle: toggleOriginal }} + zoom={zoom} + onZoom={(step) => setZoomSafe(step === 0 ? 1 : (z) => z + step * STEP)} + pageIndex={pageIndex} + pageCount={pageCount} + onPage={(delta) => void go(delta)} + />
-
- {loading && ( -
- Loading page… -
- )} - {loadError && {loadError}} - {colorPick && ( -
- Click the page to sample a color · Esc cancels -
- )} - {bytes && !loadError && ( -
- setLoadError(reason ?? "Could not render this page.")} - /> - {layout && onFormChange && ( - - )} - {layout && ( - + setModeDismissed(true)} + /> +
+ {loading && ( +
+ Loading page… +
+ )} + {loadError && {loadError}} + {colorPick && ( +
+ Click the page to sample a color · Esc cancels +
+ )} + {bytes && !loadError && ( +
+ setPickCursor(p) : undefined} - onSelect={session.select} - onClearSelection={session.clearSelection} - onBeginGesture={session.beginGesture} - onEndGesture={session.endGesture} - onUpdateRect={session.updateRect} - onUpdateRotate={(id, deg) => session.updateObject(id, { objectRotate: deg })} - onCreateShape={(kind, rect, keepAspect) => { - session.addShape(kind, pageIndex, rect, lastShapeStyle.current, keepAspect); - setTool("select"); - }} - onCreateText={(rect) => { - session.addText(pageIndex, rect); - setTool("select"); - }} - onCreateLink={(rect) => { - session.addLink(pageIndex, rect); - setTool("select"); - }} - onCreateLine={(a, b) => { - session.addLine(pageIndex, a.x, a.y, b.x, b.y); - setLastShape("line"); - setTool("select"); - }} - onCreateInk={(pts) => { - session.addInk(pageIndex, pts); - setTool("select"); - }} - onCreateNote={(rect) => { - session.addNote(pageIndex, rect, markupAuthor); - setTool("select"); - }} - onCreateHighlight={(rect) => { - session.addHighlight(pageIndex, rect, markupAuthor); - setTool("select"); - }} - onCreateUnderline={(rect) => { - session.addUnderline(pageIndex, rect, markupAuthor); - setTool("select"); - }} - onCreateStrikeout={(rect) => { - session.addStrikeout(pageIndex, rect, markupAuthor); - setTool("select"); - }} - onCreateMarkupInk={(strokes) => { - session.addMarkupInk(pageIndex, strokes, markupAuthor); - setTool("select"); - }} - onCreateRedact={(rect) => { - session.addRedact(pageIndex, rect); - setTool("select"); - }} - onRequestImage={(at) => void placeImage(at)} - onActivateText={(id) => setEditingTextId(id)} - /> - )} - {editingTextId && layout && ( - object.id === editingTextId)} - layout={layout} - onChange={(content) => session.updateObject(editingTextId, { content } as Partial)} - onClose={() => setEditingTextId(null)} + canvasRef={pageCanvasRef} + onLayout={setLayout} + onFail={(reason) => setLoadError(reason ?? "Could not render this page.")} /> - )} -
- )} + {layout && onFormChange && ( + + )} + {layout && ( + setPickCursor(p) : undefined} + onSelect={session.select} + onClearSelection={session.clearSelection} + onBeginGesture={session.beginGesture} + onEndGesture={session.endGesture} + onUpdateRect={session.updateRect} + onUpdateRotate={(id, deg) => session.updateObject(id, { objectRotate: deg })} + onCreateShape={thenSelect((kind, rect, keepAspect) => session.addShape(kind, pageIndex, rect, lastShapeStyle.current, keepAspect))} + onCreateText={thenSelect((rect) => session.addText(pageIndex, rect))} + onCreateLink={thenSelect((rect) => session.addLink(pageIndex, rect))} + onCreateLine={(a, b) => { + session.addLine(pageIndex, a.x, a.y, b.x, b.y); + setLastShape("line"); + setTool("select"); + }} + onCreateInk={thenSelect((pts) => session.addInk(pageIndex, pts))} + onCreateNote={thenSelect((rect) => session.addNote(pageIndex, rect, markupAuthor))} + onCreateHighlight={thenSelect((rect) => session.addHighlight(pageIndex, rect, markupAuthor))} + onCreateUnderline={thenSelect((rect) => session.addUnderline(pageIndex, rect, markupAuthor))} + onCreateStrikeout={thenSelect((rect) => session.addStrikeout(pageIndex, rect, markupAuthor))} + onCreateMarkupInk={thenSelect((strokes) => session.addMarkupInk(pageIndex, strokes, markupAuthor))} + onCreateRedact={thenSelect((rect) => session.addRedact(pageIndex, rect))} + onRequestImage={(at) => void placeImage(at)} + onActivateText={(id) => setEditingTextId(id)} + /> + )} + {layout && tool === "editText" && ( + setTextRequest(null)} + onLeave={() => rootRef.current?.focus({ preventScroll: true })} + onAddTextHere={addTextHere} + /> + )} + {layout && } + {editingTextId && layout && ( + object.id === editingTextId)} + layout={layout} + onChange={(content) => session.updateObject(editingTextId, { content } as Partial)} + onClose={() => setEditingTextId(null)} + /> + )} +
+ )} +
{colorPick && pickCursor && ( diff --git a/src/components/pdf/editor/controls.tsx b/src/components/pdf/editor/controls.tsx new file mode 100644 index 0000000..575cf8a --- /dev/null +++ b/src/components/pdf/editor/controls.tsx @@ -0,0 +1,225 @@ +/** + * Small form controls shared by the object inspector and the Edit text format + * bar: a draft number field (commit on blur/Enter) and a colour field. + */ +import { useEffect, useRef, useState, type Ref } from "react"; +import { toCssHex } from "@/lib/editor"; +import { Icon } from "@/components/ui/Icon"; + +function formatNum(n: number): string { + if (Number.isInteger(n)) return String(n); + return String(Math.round(n * 1000) / 1000); +} + +function parseDraft(s: string): number | null { + const t = s.trim().replace(",", "."); + if (t === "" || t === "-" || t === "." || t === "-.") return null; + const n = Number(t); + return Number.isFinite(n) ? n : null; +} + +/** Word/Paint-style number field: empty while typing, commit on blur/Enter. */ +export function DraftNumber({ + label, + hideLabel, + inline, + value, + min, + max, + suffix, + onCommit, + tabIndex, + inputRef, + rovingKey, + onDone, + live, +}: { + label: string; + hideLabel?: boolean; + /** W/H row: label | input | suffix as sibling grid cells. */ + inline?: boolean; + value: number; + min: number; + max: number; + suffix?: string; + onCommit: (n: number) => void; + tabIndex?: number; + inputRef?: Ref; + /** Marks the input as an item of a roving-tabindex toolbar. */ + rovingKey?: string; + /** Called after Enter (commit) or Esc (revert) left the field. */ + onDone?: () => void; + /** Also commit every in-range value while typing (the field may never blur before use). */ + live?: boolean; +}) { + const [focused, setFocused] = useState(false); + const [draft, setDraft] = useState(formatNum(value)); + const reverting = useRef(false); + useEffect(() => { + if (!focused) setDraft(formatNum(value)); + }, [value, focused]); + + const commit = (raw: string) => { + const n = parseDraft(raw); + if (n == null) { + setDraft(formatNum(value)); + return; + } + const clamped = Math.min(max, Math.max(min, n)); + onCommit(clamped); + setDraft(formatNum(clamped)); + }; + + const input = ( + setFocused(true)} + onChange={(e) => { + setDraft(e.target.value); + const n = live ? parseDraft(e.target.value) : null; + if (n !== null && n >= min && n <= max) onCommit(n); + }} + onBlur={(e) => { + setFocused(false); + if (reverting.current) { + reverting.current = false; + setDraft(formatNum(value)); + return; + } + commit(e.target.value); + }} + onKeyDown={(e) => { + if (e.key === "Enter") { + e.preventDefault(); + (e.target as HTMLInputElement).blur(); + onDone?.(); + } + if (e.key === "Escape") { + // Revert without committing what was typed. + e.preventDefault(); + e.stopPropagation(); + reverting.current = true; + (e.target as HTMLInputElement).blur(); + onDone?.(); + } + }} + /> + ); + const suf = suffix ? ( + + {suffix} + + ) : null; + + if (inline) { + return ( + <> + {label} + {input} + {suf} + + ); + } + + return ( +
+ {!hideLabel && } +
+ {input} + {suf} +
+
+ ); +} + +/** Colour picker + hex field + preset swatches (+ optional eyedropper). */ +export function ColorField({ + label, + icon, + value, + presets, + fallback, + active, + onChange, + onPickFromPage, +}: { + label: string; + icon?: "droplet" | "square" | "squareFill"; + value: string; + /** Swatch colours, `#rrggbb`. */ + presets: readonly string[]; + /** Shown when `value` is not a readable colour. */ + fallback: string; + active?: boolean; + onChange: (hex: string) => void; + onPickFromPage?: () => void; +}) { + const hex = toCssHex(value, fallback); + const [typed, setTyped] = useState(hex); + useEffect(() => { + setTyped(hex); + }, [hex]); + return ( +
+
+ {icon ? ( + + + + ) : ( + + )} + onChange(e.target.value)} + className="pdf-editor__color-input" + /> + { + const v = e.target.value.trim(); + setTyped(v.startsWith("#") || v.length === 0 ? v : `#${v}`); + const next = v.startsWith("#") ? v : `#${v}`; + if (/^#[0-9a-fA-F]{6}$/.test(next)) onChange(next.toLowerCase()); + else if (/^#[0-9a-fA-F]{3}$/.test(next)) onChange(toCssHex(next)); + }} + /> + {onPickFromPage && ( + + )} +
+
+ {presets.map((c) => ( +
+
+ ); +} diff --git a/src/components/pdf/editor/index.ts b/src/components/pdf/editor/index.ts index 5365b01..be4978d 100644 --- a/src/components/pdf/editor/index.ts +++ b/src/components/pdf/editor/index.ts @@ -1,7 +1,9 @@ export { PdfEditorCanvas } from "./PdfEditorCanvas"; +export type { CanvasTextProps } from "./PdfEditorCanvas"; export { PageSurface } from "./PageSurface"; export type { PageLayout } from "./PageSurface"; export { EditorOverlay } from "./EditorOverlay"; +export { EditorToolbar, DEFAULT_EDITOR_TOOL, editorShortcut } from "./EditorToolbar"; export { FormFieldsOverlay } from "./FormFieldsOverlay"; export { ObjectList } from "./ObjectList"; export { useEditSession } from "./useEditSession"; diff --git a/src/components/pdf/editor/sourceText/ReasonPopover.tsx b/src/components/pdf/editor/sourceText/ReasonPopover.tsx new file mode 100644 index 0000000..b62b5b8 --- /dev/null +++ b/src/components/pdf/editor/sourceText/ReasonPopover.tsx @@ -0,0 +1,73 @@ +/** + * Why a line can't be changed (SPEC §D.6.3): a non-modal dialog anchored to the + * run, focus on its first button. "Add text here" places a new text box over + * the line (nothing is hidden); Esc or Close returns focus to the run. + */ +import { useEffect, useId, useRef } from "react"; +import type { CssRect } from "@/lib/editor"; +import { UI, reasonCopy } from "@/lib/editor/sourceTextCopy"; +import type { TextRun } from "@/lib/types"; + +/** Width the popover is laid out for (CSS px); it never overflows the page sideways. */ +const POPOVER_WIDTH = 300; +/** Below the run unless that leaves less than this much page under it. */ +const POPOVER_MIN_ROOM = 170; +const GAP = 6; + +export interface ReasonPopoverProps { + run: TextRun; + /** The run's box on the page (CSS px). */ + anchor: CssRect; + pageWidth: number; + pageHeight: number; + onAddTextHere: () => void; + onClose: () => void; +} + +export function ReasonPopover({ run, anchor, pageWidth, pageHeight, onAddTextHere, onClose }: ReasonPopoverProps) { + const titleId = useId(); + const firstRef = useRef(null); + const copy = reasonCopy(run.reason); + + useEffect(() => { + firstRef.current?.focus(); + }, []); + + const below = anchor.y + anchor.h + POPOVER_MIN_ROOM <= pageHeight; + const left = Math.max(0, Math.min(anchor.x, pageWidth - POPOVER_WIDTH)); + const style = below + ? { left, top: anchor.y + anchor.h + GAP } + : { left, top: anchor.y - GAP, transform: "translateY(-100%)" }; + + return ( +
{ + if (e.key !== "Escape") return; + e.preventDefault(); + e.stopPropagation(); + onClose(); + }} + > +
+ {UI.popover.title} +
+
{copy.title}
+

{copy.body}

+

{UI.popover.footer}

+
+ + +
+
+ ); +} diff --git a/src/components/pdf/editor/sourceText/SourceTextEditor.test.tsx b/src/components/pdf/editor/sourceText/SourceTextEditor.test.tsx new file mode 100644 index 0000000..70b32df --- /dev/null +++ b/src/components/pdf/editor/sourceText/SourceTextEditor.test.tsx @@ -0,0 +1,344 @@ +// @vitest-environment happy-dom +import { StrictMode, act, createElement } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import { editsSignature, fontsByKey, makeSourceTextObject, type SourceTextObject } from "@/lib/editor"; +import { PROBLEM_COPY, UI } from "@/lib/editor/sourceTextCopy"; +import type { AppError, TextEditVerdict, TextFont, TextPreview, TextRun } from "@/lib/types"; +import type { PageLayout } from "../PageSurface"; +import { SourceTextEditor, type SourceTextEditorProps, type TryDone } from "./SourceTextEditor"; + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +/** Word-style subset: only the letters "Hello" and "Sağlık" use. */ +const SUBSET: TextFont = { + key: "f1", + displayName: "Calibri", + familyHint: "sans", + embedded: true, + subset: true, + alphabet: " !Helo", + widths: [226, 268, 631, 498, 229, 527], + wordSpace: true, +}; + +const RUN: TextRun = { + id: "t1:fp:0:10-20", + order: 0, + line: 0, + text: "Hello", + rect: { x: 72, y: 697, w: 25.6, h: 12 }, + origin: { x: 72, y: 700 }, + dir: { x: 1, y: 0 }, + ascent: 9, + descent: 3, + caretOffsets: [0, 7.6, 13.5, 16.3, 19.0, 25.4], + editable: true, + reason: null, + metrics: { + surface: ["f1"], + tfSize: 12, + effectiveSize: 12, + charSpacing: 0, + wordSpacing: 0, + hScale: 1, + textToUser: 1, + letterSpacingPt: 0, + spaceMode: "glyph", + kernSpace: -250, + originalWidth: 25.4, + visibleExtent: 400, + nextObstacle: null, + }, + style: { + fill: "#000000", + sizeChangeable: true, + colourChangeable: true, + face: "regular", + faces: { + regular: { available: true, surface: ["f1"] }, + bold: { available: false, surface: [] }, + italic: { available: false, surface: [] }, + boldItalic: { available: false, surface: [] }, + }, + }, + substituted: false, +}; + +const LAYOUT: PageLayout = { cssWidth: 612, cssHeight: 792, geometry: { box: { x: 0, y: 0, w: 612, h: 792 }, rotate: 0, pageIndex: 0 } }; + +function verdict(extra: Partial): TextEditVerdict { + return { + runId: RUN.id, + ok: true, + code: null, + chars: [], + reason: null, + face: null, + field: null, + detail: null, + deltaPt: 0, + newRect: null, + caretOffsets: null, + ...extra, + }; +} + +function answer(v: TextEditVerdict, extra: Partial = {}): TextPreview { + return { pagePdf: "UERG", verdicts: [v], pageProblem: null, warnings: [], ...extra }; +} + +let root: Root | null = null; +let host: HTMLElement; +let tryDone: TryDone | null = null; + +function mount(extra: Partial = {}, strict = false) { + const props: SourceTextEditorProps = { + run: RUN, + fonts: fontsByKey([SUBSET]), + layout: LAYOUT, + pageIndex: 0, + sourceFingerprint: "fp", + sourcePageIndex: 0, + fileName: "invoice.pdf", + existing: null, + others: [], + caret: 2, + neighbourText: null, + requestPreview: vi.fn(() => new Promise(() => {})), + seedPreview: vi.fn(), + onApply: vi.fn(), + onRevert: vi.fn(), + onClose: vi.fn(), + onFileError: vi.fn(), + registerTryDone: (fn) => { + tryDone = fn; + }, + ...extra, + }; + host = document.createElement("div"); + document.body.appendChild(host); + root = createRoot(host); + const editor = createElement(SourceTextEditor, props); + act(() => root!.render(strict ? createElement(StrictMode, null, editor) : editor)); + return props; +} + +afterEach(() => { + act(() => root?.unmount()); + root = null; + host.remove(); + tryDone = null; +}); + +const input = () => host.querySelector("input.st-chip__input")!; +const message = () => host.querySelector('[role="status"]')!.textContent ?? ""; + +function type(value: string) { + const setter = Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, "value")!.set!; + act(() => { + setter.call(input(), value); + input().dispatchEvent(new Event("input", { bubbles: true })); + }); +} + +function press(key: string, init: KeyboardEventInit = {}) { + act(() => { + input().dispatchEvent(new KeyboardEvent("keydown", { key, bubbles: true, cancelable: true, ...init })); + }); +} + +describe("SourceTextEditor", () => { + it("opens focused, labelled and described by its message row, caret where the line was clicked", () => { + mount(); + expect(document.activeElement).toBe(input()); + expect(input().getAttribute("aria-label")).toBe(UI.editor.aria); + expect(input().getAttribute("aria-describedby")).toBe(host.querySelector('[role="status"]')!.id); + expect(input().selectionStart).toBe(2); + expect(input().selectionEnd).toBe(2); + }); + + it("Enter/F2 opening selects all of the text", () => { + mount({ caret: "all" }); + expect(input().selectionStart).toBe(0); + expect(input().selectionEnd).toBe(5); + }); + + it("typing ğ into a subset shows the exact missing message and blocks Done", async () => { + const props = mount(); + type("Hellğ"); + expect(message()).toBe("This document's font can't draw: ğ The file only includes the letters it already uses."); + press("Enter"); + await act(async () => {}); + expect(props.requestPreview).not.toHaveBeenCalled(); + expect(props.onClose).not.toHaveBeenCalled(); + expect(host.querySelector('[role="status"]')!.getAttribute("aria-live")).toBe("assertive"); + }); + + it("a failed preview verdict keeps the editor open with the draft and the PEN_DRIFT copy", async () => { + const props = mount({ requestPreview: vi.fn().mockResolvedValue(answer(verdict({ ok: false, code: "PEN_DRIFT" }), { pagePdf: null })) }); + type("Hello!"); + press("Enter"); + expect(message()).toBe(UI.status.checking); + expect(input().readOnly).toBe(true); + await act(async () => {}); + expect(message()).toBe(PROBLEM_COPY.PEN_DRIFT); + expect(input().value).toBe("Hello!"); + expect(input().readOnly).toBe(false); + expect(props.onApply).not.toHaveBeenCalled(); + expect(props.onClose).not.toHaveBeenCalled(); + }); + + it("B11: clicking another run while blocked keeps the draft and says so", async () => { + const props = mount(); + type("Hellğ!"); + act(() => { + document.body.dispatchEvent(new PointerEvent("pointerdown", { bubbles: true })); + }); + await act(async () => {}); + expect(props.onClose).not.toHaveBeenCalled(); + expect(input().value).toBe("Hellğ!"); + expect(message()).toContain(UI.blocked); + let closed: boolean | undefined; + await act(async () => { + closed = await tryDone!({ focusRun: false, blockedNotice: true }); + }); + expect(closed).toBe(false); + }); + + it("Esc cancels the draft: nothing is committed, focus goes back to the line", () => { + const props = mount(); + type("Help"); + press("Escape"); + expect(props.onApply).not.toHaveBeenCalled(); + expect(props.onClose).toHaveBeenCalledWith({ focusRun: true, announce: null }); + }); + + it("empty text shows the removal note", () => { + mount(); + type(""); + expect(message()).toBe(UI.removal); + }); + + it("ignores Enter and Esc while an IME is composing (isComposing, or WebKit's keyCode 229)", async () => { + const props = mount(); + type("Hell"); + press("Enter", { isComposing: true }); + press("Enter", { keyCode: 229 } as KeyboardEventInit); + press("Escape", { isComposing: true }); + await act(async () => {}); + expect(props.requestPreview).not.toHaveBeenCalled(); + expect(props.onClose).not.toHaveBeenCalled(); + }); + + it("a size typed in the bar counts even when Done is clicked without leaving the field", async () => { + const props = mount({ requestPreview: vi.fn().mockResolvedValue(answer(verdict({ newRect: RUN.rect }))) }); + const size = host.querySelector(`input[aria-label="${UI.bar.fontSize}"]`)!; + const setter = Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, "value")!.set!; + act(() => { + setter.call(size, "14"); + size.dispatchEvent(new Event("input", { bubbles: true })); + }); + const done = [...host.querySelectorAll("button")].find((b) => b.textContent?.includes(UI.bar.done))!; + act(() => done.click()); + await act(async () => {}); + expect(props.requestPreview).toHaveBeenCalledWith([{ runId: RUN.id, originalText: "Hello", text: "Hello", style: { sizePt: 14 } }]); + }); + + it("style changes are ignored while a check is running (the check is for the style it started with)", async () => { + let finish!: (p: TextPreview) => void; + const props = mount({ requestPreview: vi.fn(() => new Promise((res) => (finish = res))) }); + type("Hello!"); + press("Enter"); + press(".", { metaKey: true, shiftKey: true, code: "Period" } as KeyboardEventInit); + // The draft does not pretend to a size the running check will not commit. + expect(host.querySelector(`input[aria-label="${UI.bar.fontSize}"]`)!.value).toBe("12"); + await act(async () => finish(answer(verdict({ newRect: RUN.rect })))); + expect(props.onApply).toHaveBeenCalledWith(expect.objectContaining({ text: "Hello!", style: {} })); + }); + + it("commit: one preview with the page's other changes, one setSourceText, seeded preview, announcement", async () => { + const other: SourceTextObject = makeSourceTextObject("o2", 0, RUN.rect, { + runId: "t1:fp:0:40-50", + sourceFingerprint: "fp", + sourcePageIndex: 0, + originalText: "Total", + text: "Sum", + style: {}, + }); + const newRect = { x: 72, y: 697, w: 27.9, h: 12 }; + const preview = answer(verdict({ deltaPt: 2.5, newRect, caretOffsets: [0, 7.6, 13.5, 16.3, 19, 25.4, 27.9] })); + const props = mount({ others: [other], requestPreview: vi.fn().mockResolvedValue(preview) }); + type("Hello!"); + press("Enter"); + await act(async () => {}); + expect(props.requestPreview).toHaveBeenCalledTimes(1); + expect(props.requestPreview).toHaveBeenCalledWith([ + { runId: "t1:fp:0:40-50", originalText: "Total", text: "Sum", style: {} }, + { runId: RUN.id, originalText: "Hello", text: "Hello!", style: {} }, + ]); + expect(props.onApply).toHaveBeenCalledTimes(1); + expect(props.onApply).toHaveBeenCalledWith( + expect.objectContaining({ pageIndex: 0, run: RUN, sourceFingerprint: "fp", sourcePageIndex: 0, text: "Hello!", style: {}, rect: newRect }), + ); + const committed = makeSourceTextObject("x", 0, newRect, { + runId: RUN.id, + sourceFingerprint: "fp", + sourcePageIndex: 0, + originalText: "Hello", + text: "Hello!", + style: {}, + }); + expect(props.seedPreview).toHaveBeenCalledWith(editsSignature([other, committed]), preview); + expect(props.onClose).toHaveBeenCalledWith({ focusRun: true, announce: "Change applied. 2.5 pt wider than before." }); + }); + + it("commits under React.StrictMode too (the app's development mode mounts twice)", async () => { + const preview = answer(verdict({ deltaPt: 1, newRect: RUN.rect, caretOffsets: [0, 7.6, 13.5, 16.3, 19, 25.4, 27.9] })); + const props = mount({ requestPreview: vi.fn().mockResolvedValue(preview) }, true); + type("Hello!"); + press("Enter"); + await act(async () => {}); + expect(props.onApply).toHaveBeenCalledTimes(1); + expect(props.onClose).toHaveBeenCalledTimes(1); + }); + + it("a no-op Done restores the original of a committed change (one revert) and closes", async () => { + const existing = makeSourceTextObject("e1", 0, RUN.rect, { + runId: RUN.id, + sourceFingerprint: "fp", + sourcePageIndex: 0, + originalText: "Hello", + text: "Help", + style: {}, + }); + const props = mount({ existing }); + expect(input().value).toBe("Help"); + type("Hello"); + press("Enter"); + await act(async () => {}); + expect(props.requestPreview).not.toHaveBeenCalled(); + expect(props.onRevert).toHaveBeenCalledWith("e1"); + expect(props.onClose).toHaveBeenCalledWith({ focusRun: true, announce: UI.announce.restored }); + }); + + it("STALE closes the editor with nothing committed and reports the file", async () => { + const stale: AppError = { code: "STALE", title: "The PDF changed on disk", message: "m" }; + const props = mount({ requestPreview: vi.fn().mockRejectedValue(stale) }); + type("Hell"); + press("Enter"); + await act(async () => {}); + expect(props.onFileError).toHaveBeenCalledWith(stale); + expect(props.onApply).not.toHaveBeenCalled(); + expect(props.onClose).toHaveBeenCalledWith({ focusRun: false, announce: null }); + }); + + it("⌘I on a font without italic writes the italic sentence; width sentence shows otherwise", () => { + mount(); + expect(message()).toBe(UI.width.same); + press("i", { metaKey: true }); + expect(message()).toBe("This page has no italic version of this font."); + type("Hello!"); + expect(message()).toMatch(/^About \d+(\.\d)? pt wider than before\.$/); + }); +}); diff --git a/src/components/pdf/editor/sourceText/SourceTextEditor.tsx b/src/components/pdf/editor/sourceText/SourceTextEditor.tsx new file mode 100644 index 0000000..32729ef --- /dev/null +++ b/src/components/pdf/editor/sourceText/SourceTextEditor.tsx @@ -0,0 +1,384 @@ +/** + * Inline editor for one line (SPEC §D.5 EDITING, §D.6.1): a single-line input + * over the run, width guides, a one-line message row and the format bar. + * + * Done normalises the draft; a no-op restores the original (if changed) and + * closes; a local blocking problem keeps the editor open; anything else is + * checked by `preview_text_edits` together with the page's other changes and + * committed (one history step) only when that check passes. A failed check + * keeps the draft and shows why (B11). STALE / PDF_NEEDS_REPAIR close the + * editor with nothing committed; the page banner explains. + */ +import { useEffect, useId, useLayoutEffect, useRef, useState, type KeyboardEvent } from "react"; +import { + blockingVerdict, + displayedSize, + editsSignature, + estimateDeltaPt, + isNoOpEdit, + makeMapping, + makeSourceTextObject, + normaliseStyle, + normaliseTyped, + overlapsNext, + pdfRectToViewport, + pdfToViewport, + problemMessage, + toTextEditIn, + type SetSourceTextInput, + type SourceTextObject, +} from "@/lib/editor"; +import { PROBLEM_COPY, UI, fillCopy, problemCopy, warningCopy, widthSentence } from "@/lib/editor/sourceTextCopy"; +import { toAppError, type AppError, type SourceTextStyle, type TextEditIn, type TextFont, type TextPreview, type TextRun } from "@/lib/types"; +import type { PageLayout } from "../PageSurface"; +import { TextFormatBar, chipFont, steppedSize, toggledFace, type StyleChange } from "./TextFormatBar"; + +const CHIP_PAD_PX = 3; +const LINE_HEIGHT = 1.3; +/** The visible-limit guide appears once the new end is this close to it. */ +const LIMIT_GUIDE_PT = 10; +const ASSERTIVE_MS = 1500; +/** Room the bar and message row need above / below the chip (CSS px). */ +const EDGE_ROOM_PX = 44; + +export interface DoneOptions { + /** Return focus to the run after closing (Enter / Done); not after an outside click. */ + focusRun: boolean; + /** Add "Fix or cancel this change first…" when the change can't complete. */ + blockedNotice: boolean; +} +export type TryDone = (opts: DoneOptions) => Promise; + +export interface SourceTextEditorProps { + run: TextRun; + fonts: Map; + layout: PageLayout; + pageIndex: number; + sourceFingerprint: string; + sourcePageIndex: number; + fileName: string; + /** The committed change of this run, if any. */ + existing: SourceTextObject | null; + /** The page's other committed changes (checked together with this one). */ + others: SourceTextObject[]; + /** Caret index (code points) or select everything. */ + caret: number | "all"; + /** Text of the next run on the line, for the overlap message. */ + neighbourText: string | null; + requestPreview: (edits: TextEditIn[]) => Promise; + seedPreview: (signature: string, preview: TextPreview) => void; + onApply: (input: SetSourceTextInput) => void; + onRevert: (id: string) => void; + onClose: (outcome: { focusRun: boolean; announce: string | null }) => void; + onFileError: (error: AppError) => void; + registerTryDone: (fn: TryDone | null) => void; +} + +/** UTF-16 offset of the caret after `index` code points. */ +function utf16Index(text: string, index: number): number { + return Array.from(text).slice(0, index).join("").length; +} + +function styleKey(text: string, style: SourceTextStyle): string { + return JSON.stringify([text, style.sizePt, style.face, style.fill, style.letterSpacingPt]); +} + +function failureMessage(preview: TextPreview, runId: string, run: TextRun, fonts: Map, name: string): string { + const verdict = preview.verdicts.find((v) => v.runId === runId); + if (verdict && !verdict.ok) return problemMessage(verdict, run, fonts, name); + if (preview.pageProblem) return problemCopy(preview.pageProblem.code, [], { name }); + return PROBLEM_COPY.EDIT_VERIFY_FAILED; +} + +export function SourceTextEditor(props: SourceTextEditorProps) { + const { run, fonts, layout, existing } = props; + const metrics = run.metrics; + const rootRef = useRef(null); + const inputRef = useRef(null); + const chipRef = useRef(null); + const mirrorRef = useRef(null); + const msgId = useId(); + /** Rendered width of the draft in the input's font, and the stage room around the chip. */ + const [measured, setMeasured] = useState({ textPx: 0, above: Infinity, below: Infinity }); + const [draft, setDraft] = useState(existing?.text ?? run.text); + const [draftStyle, setDraftStyle] = useState(existing?.style ?? {}); + const [notice, setNotice] = useState(null); + const [server, setServer] = useState<{ key: string; message: string } | null>(null); + const [checking, setChecking] = useState(false); + const [blockedNote, setBlockedNote] = useState(false); + const [assertive, setAssertive] = useState(false); + const pending = useRef | null>(null); + const generation = useRef(0); + const closed = useRef(false); + const mounted = useRef(true); + const assertiveTimer = useRef | null>(null); + + useEffect(() => { + // Set here too: StrictMode mounts, unmounts and mounts again in development. + mounted.current = true; + return () => { + mounted.current = false; + if (assertiveTimer.current) clearTimeout(assertiveTimer.current); + }; + }, []); + + useLayoutEffect(() => { + const el = inputRef.current; + if (!el) return; + el.focus(); + if (props.caret === "all") el.select(); + else { + const at = utf16Index(el.value, props.caret); + el.setSelectionRange(at, at); + } + }, []); // placed once, when the editor opens + + // After every render: the substitute font can be wider than the PDF estimate, and the + // bar and message row must stay inside the visible part of the stage. + useLayoutEffect(() => { + const chip = chipRef.current?.getBoundingClientRect(); + const stage = rootRef.current?.closest(".pdf-editor__stage")?.getBoundingClientRect(); + const next = { + textPx: Math.ceil(mirrorRef.current?.offsetWidth ?? 0), + above: chip && stage ? Math.round(chip.top - stage.top) : Infinity, + below: chip && stage ? Math.round(stage.bottom - chip.bottom) : Infinity, + }; + if (next.textPx !== measured.textPx || next.above !== measured.above || next.below !== measured.below) setMeasured(next); + }); + + const close = (outcome: { focusRun: boolean; announce: string | null }) => { + closed.current = true; + props.onClose(outcome); + }; + + const block = (withNotice: boolean) => { + setBlockedNote(withNotice); + setAssertive(true); + if (assertiveTimer.current) clearTimeout(assertiveTimer.current); + assertiveTimer.current = setTimeout(() => mounted.current && setAssertive(false), ASSERTIVE_MS); + inputRef.current?.focus(); + }; + + const check = async (text: string, style: SourceTextStyle, opts: DoneOptions, gen: number): Promise => { + const candidate: TextEditIn = { runId: run.id, originalText: run.text, text, style }; + try { + const preview = await props.requestPreview([...props.others.map(toTextEditIn), candidate]); + if (gen !== generation.current || !mounted.current) return closed.current || !mounted.current; + const verdict = preview.verdicts.find((v) => v.runId === run.id); + if (verdict?.ok && !preview.pageProblem) { + const fields = { + runId: run.id, + sourceFingerprint: props.sourceFingerprint, + sourcePageIndex: props.sourcePageIndex, + originalText: run.text, + text, + style, + }; + const committed = makeSourceTextObject(existing?.id ?? run.id, props.pageIndex, verdict.newRect ?? run.rect, fields); + props.seedPreview(editsSignature([...props.others, committed]), preview); + props.onApply({ ...fields, pageIndex: props.pageIndex, run, rect: verdict.newRect }); + close({ focusRun: opts.focusRun, announce: fillCopy(UI.announce.applied, { width: widthSentence(verdict.deltaPt, true) }) }); + return true; + } + setServer({ key: styleKey(text, style), message: failureMessage(preview, run.id, run, fonts, props.fileName) }); + } catch (e) { + if (gen !== generation.current || !mounted.current) return closed.current || !mounted.current; + const err = toAppError(e); + if (err.code === "STALE" || err.code === "PDF_NEEDS_REPAIR") { + props.onFileError(err); + close({ focusRun: false, announce: null }); + return true; + } + setServer({ key: styleKey(text, style), message: [err.message, err.suggestion].filter(Boolean).join(" ") }); + } + setChecking(false); + block(opts.blockedNotice); + return false; + }; + + const done: TryDone = async (opts) => { + if (closed.current) return true; + if (pending.current) return pending.current; + const text = normaliseTyped(draft); + const style = normaliseStyle(run, draftStyle); + if (isNoOpEdit(run, text, style)) { + if (existing) props.onRevert(existing.id); + close({ focusRun: opts.focusRun, announce: existing ? UI.announce.restored : null }); + return true; + } + if (blockingVerdict(run, fonts, text, style)) { + block(opts.blockedNotice); + return false; + } + const gen = ++generation.current; + setChecking(true); + const work = check(text, style, opts, gen).finally(() => { + if (pending.current === work) pending.current = null; + }); + pending.current = work; + return work; + }; + const doneRef = useRef(done); + doneRef.current = done; + + const { registerTryDone } = props; + useEffect(() => { + registerTryDone((opts) => doneRef.current(opts)); + return () => registerTryDone(null); + }, [registerTryDone]); + + useEffect(() => { + const onDown = (e: PointerEvent) => { + const target = e.target as Element | null; + if (!target || rootRef.current?.contains(target) || target.classList?.contains("pdf-editor__stage")) return; + void doneRef.current({ focusRun: false, blockedNotice: true }); + }; + document.addEventListener("pointerdown", onDown, true); + return () => document.removeEventListener("pointerdown", onDown, true); + }, []); + + const cancel = () => { + generation.current += 1; + pending.current = null; + close({ focusRun: true, announce: null }); + }; + + if (!metrics || !run.style) return null; + + const restyle = (change: StyleChange) => { + if (checking) return; // the check in flight is for the style it started with + setServer(null); + setBlockedNote(false); + if (change.style) { + setDraftStyle(change.style); + setNotice(null); + } else if (change.notice) setNotice(change.notice); + }; + + const onKeyDown = (e: KeyboardEvent) => { + const mod = e.metaKey || e.ctrlKey; + // An IME owns Enter and Esc while composing (WebKit reports the confirming Enter as keyCode 229). + if ((e.key === "Enter" || e.key === "Escape") && (e.nativeEvent.isComposing || e.keyCode === 229)) return; + if (e.key === "Enter") { + e.preventDefault(); + void done({ focusRun: true, blockedNotice: false }); + } else if (e.key === "Escape") { + e.preventDefault(); + e.stopPropagation(); + cancel(); + } else if (e.key === "Tab" && !e.shiftKey) { + e.preventDefault(); + rootRef.current?.querySelector('.st-bar [data-roving][tabindex="0"]')?.focus(); + } else if (mod && !e.shiftKey && (e.key === "b" || e.key === "B" || e.key === "i" || e.key === "I")) { + e.preventDefault(); + restyle(toggledFace(run, draftStyle, e.key.toLowerCase() === "b" ? "bold" : "italic")); + } else if (mod && e.shiftKey && (e.code === "Period" || e.code === "Comma" || ">.<,".includes(e.key))) { + e.preventDefault(); + const up = e.code === "Period" || e.key === ">" || e.key === "."; + restyle(steppedSize(run, draftStyle, up ? 0.5 : -0.5)); + } + }; + + // Geometry: the run's box on screen (upright as displayed, §A.5), font sized by its effective size. + const mapping = makeMapping(layout.geometry, layout.cssWidth, layout.cssHeight); + const pxPerPt = layout.cssHeight / Math.max(displayedSize(layout.geometry).h, 1); + const box = pdfRectToViewport(run.rect, mapping); + const font = chipFont(run, fonts, draftStyle, pxPerPt); + const fontPx = Number(font.css.fontSize); + const delta = estimateDeltaPt(run, fonts, draft, draftStyle); + const endPt = metrics.originalWidth + delta; + const left = box.x - CHIP_PAD_PX; + const height = Math.max(box.h, fontPx * LINE_HEIGHT) + 2 * CHIP_PAD_PX; + const top = box.y + box.h / 2 - height / 2; + const width = Math.max(Math.max(metrics.originalWidth, endPt) * pxPerPt, measured.textPx + 2) + 2 * CHIP_PAD_PX; + const along = (d: number) => pdfToViewport({ x: run.origin.x + run.dir.x * d, y: run.origin.y + run.dir.y * d }, mapping).x - left; + const nearTop = Math.min(top, measured.above) < EDGE_ROOM_PX; + const nearBottom = Math.min(layout.cssHeight - top - height, measured.below) < EDGE_ROOM_PX; + + // Message row, one line, highest priority first (§D.6.1). + const text = normaliseTyped(draft); + const blocked = blockingVerdict(run, fonts, text, draftStyle); + const serverMsg = server && server.key === styleKey(text, normaliseStyle(run, draftStyle)) ? server.message : null; + let tone: "danger" | "warning" | "neutral" = "neutral"; + let message: string; + if (notice) [tone, message] = ["warning", notice]; + else if (checking) message = UI.status.checking; + else if (serverMsg) [tone, message] = ["danger", serverMsg]; + else if (blocked) [tone, message] = ["danger", problemMessage(blocked, run, fonts, props.fileName)]; + else if (overlapsNext(run, fonts, text, draftStyle)) [tone, message] = ["warning", warningCopy("NEXT_TEXT_OVERLAP", props.neighbourText)]; + else if (text === "") message = UI.removal; + else message = widthSentence(delta, false) + (font.floored ? ` ${UI.shownLarger}` : ""); + if (blockedNote && tone === "danger") message = `${message} ${UI.blocked}`; + + const bar = ( + restyle({ style: next })} + onNotice={(sentence) => restyle({ notice: sentence })} + canRestore={!!existing} + onRestore={() => { + generation.current += 1; + if (existing) props.onRevert(existing.id); + close({ focusRun: true, announce: UI.announce.restored }); + }} + onCancel={cancel} + onDone={() => void done({ focusRun: true, blockedNotice: false })} + onBackToInput={() => inputRef.current?.focus()} + placement={nearTop ? "below" : "above"} + /> + ); + const row = ( +
+ {message} +
+ ); + + // The bar floats above the chip and the message row sits below it; near the + // page top the bar moves below the row, near the bottom the row moves above. + return ( +
+
+ {!nearTop && bar} + {nearBottom && row} +
+
+ { + setDraft(e.target.value); + setNotice(null); + setServer(null); + setBlockedNote(false); + }} + onKeyDown={onKeyDown} + /> + + +
+
+ {!nearBottom && row} + {nearTop && bar} +
+
+ ); +} diff --git a/src/components/pdf/editor/sourceText/SourceTextLayer.test.tsx b/src/components/pdf/editor/sourceText/SourceTextLayer.test.tsx new file mode 100644 index 0000000..e684899 --- /dev/null +++ b/src/components/pdf/editor/sourceText/SourceTextLayer.test.tsx @@ -0,0 +1,324 @@ +// @vitest-environment happy-dom +import { act, createElement, useState } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import { fontsByKey, makeSourceTextObject, type SourceTextObject } from "@/lib/editor"; +import { UI, fillCopy } from "@/lib/editor/sourceTextCopy"; +import type { TextFont, TextRun } from "@/lib/types"; +import type { PageLayout } from "../PageSurface"; +import { SourceTextLayer, type SourceTextLayerProps } from "./SourceTextLayer"; + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +const FONT: TextFont = { + key: "f1", + displayName: "Arial", + familyHint: "sans", + embedded: true, + subset: false, + alphabet: " 0123456789ITacdeilnostuv", + widths: Array.from({ length: 25 }, () => 500), + wordSpace: true, +}; + +/** Letter page, 1 CSS px per point; CSS y = 792 − PDF y. */ +const LAYOUT: PageLayout = { cssWidth: 612, cssHeight: 792, geometry: { box: { x: 0, y: 0, w: 612, h: 792 }, rotate: 0, pageIndex: 0 } }; + +function run(id: string, order: number, line: number, text: string, x: number, y: number, extra: Partial = {}): TextRun { + const w = text.length * 6; + return { + id, + order, + line, + text, + rect: { x, y: y - 3, w, h: 12 }, + origin: { x, y }, + dir: { x: 1, y: 0 }, + ascent: 9, + descent: 3, + caretOffsets: Array.from({ length: text.length + 1 }, (_, i) => i * 6), + editable: true, + reason: null, + metrics: { + surface: ["f1"], + tfSize: 12, + effectiveSize: 12, + charSpacing: 0, + wordSpacing: 0, + hScale: 1, + textToUser: 1, + letterSpacingPt: 0, + spaceMode: "glyph", + kernSpace: -250, + originalWidth: w, + visibleExtent: 500, + nextObstacle: null, + }, + style: { + fill: "#000000", + sizeChangeable: true, + colourChangeable: true, + face: "regular", + faces: { + regular: { available: true, surface: ["f1"] }, + bold: { available: false, surface: [] }, + italic: { available: false, surface: [] }, + boldItalic: { available: false, surface: [] }, + }, + }, + substituted: false, + ...extra, + }; +} + +const INVOICE = run("r-invoice", 0, 0, "Invoice 2026", 72, 700); +const TOTAL = run("r-total", 1, 0, "Total", 300, 700); +const SHARED = run("r-shared", 2, 1, "Letterhead", 72, 650, { editable: false, reason: "SHARED_CONTENT", metrics: null, style: null }); +const LOW = run("r-low", 3, 2, "Footer", 72, 60); + +let root: Root | null = null; +let host: HTMLElement; +let stage: HTMLElement; + +type Spies = Pick; + +function mount(runs: TextRun[], extra: Partial = {}, stageHeight = 792): Spies { + const spies: Spies = { onEdit: vi.fn(), onExplain: vi.fn(), onRestore: vi.fn(), onLeave: vi.fn() }; + stage = document.createElement("div"); + stage.getBoundingClientRect = () => ({ left: 0, top: 0, right: 612, bottom: stageHeight, width: 612, height: stageHeight, x: 0, y: 0, toJSON: () => ({}) }); + host = document.createElement("div"); + stage.appendChild(host); + document.body.appendChild(stage); + function Harness() { + const [focus, setFocus] = useState<{ id: string | null; tick: number }>({ id: extra.focusedId ?? null, tick: 0 }); + return createElement(SourceTextLayer, { + runs, + objects: [], + fonts: fontsByKey([FONT]), + layout: LAYOUT, + pageNumber: 3, + preview: null, + duplicate: false, + editingId: null, + popoverId: null, + stage, + inert: false, + ...extra, + ...spies, + focusedId: focus.id, + focusTick: focus.tick, + onFocusedChange: (id: string, move: boolean) => setFocus((f) => ({ id, tick: move ? f.tick + 1 : f.tick })), + }); + } + root = createRoot(host); + act(() => root!.render(createElement(Harness))); + return spies; +} + +afterEach(() => { + act(() => root?.unmount()); + root = null; + stage.remove(); +}); + +const runButtons = () => [...host.querySelectorAll("button.st-run")]; +const buttonFor = (r: TextRun) => runButtons().find((b) => b.getAttribute("aria-label")?.includes(r.text))!; +const layer = () => host.querySelector(".st-layer")!; + +function key(el: Element, k: string): KeyboardEvent { + const ev = new KeyboardEvent("keydown", { key: k, bubbles: true, cancelable: true }); + act(() => { + el.dispatchEvent(ev); + }); + return ev; +} + +describe("SourceTextLayer", () => { + it("is a labelled group of buttons in reading order with one Tab stop", () => { + mount([SHARED, TOTAL, INVOICE]); + expect(layer().getAttribute("role")).toBe("group"); + expect(layer().getAttribute("aria-label")).toBe(fillCopy(UI.layer.label, { n: 3 })); + expect(layer().hasAttribute("data-source-text")).toBe(true); + expect(runButtons().map((b) => b.getAttribute("aria-label"))).toEqual([ + "Edit text: Invoice 2026", + "Edit text: Total", + "Letterhead", + ]); + expect(runButtons().filter((b) => b.tabIndex === 0)).toEqual([buttonFor(INVOICE)]); + }); + + it("arrows move focus in reading order (→ along the line, ↓ to the next line)", () => { + mount([INVOICE, TOTAL, SHARED]); + act(() => buttonFor(INVOICE).focus()); + key(buttonFor(INVOICE), "ArrowRight"); + expect(document.activeElement).toBe(buttonFor(TOTAL)); + key(buttonFor(TOTAL), "ArrowDown"); + expect(document.activeElement).toBe(buttonFor(SHARED)); + key(buttonFor(SHARED), "Home"); + expect(document.activeElement).toBe(buttonFor(INVOICE)); + expect(runButtons().filter((b) => b.tabIndex === 0)).toEqual([buttonFor(INVOICE)]); + }); + + it("a refused line is aria-disabled and described as Can't be changed: {short}", () => { + mount([INVOICE, SHARED]); + const b = buttonFor(SHARED); + expect(b.getAttribute("aria-disabled")).toBe("true"); + const described = document.getElementById(b.getAttribute("aria-describedby")!); + expect(described?.textContent).toBe("Can't be changed: shared with other pages"); + }); + + it("Enter (and F2) on an editable line opens the editor with all text selected", () => { + const spies = mount([INVOICE, SHARED]); + key(buttonFor(INVOICE), "Enter"); + expect(spies.onEdit).toHaveBeenCalledWith(INVOICE, "all"); + key(buttonFor(INVOICE), "F2"); + expect(spies.onEdit).toHaveBeenCalledTimes(2); + }); + + it("Enter on a refused line opens the reason popover", () => { + const spies = mount([INVOICE, SHARED]); + key(buttonFor(SHARED), "Enter"); + expect(spies.onExplain).toHaveBeenCalledWith(SHARED); + expect(spies.onEdit).not.toHaveBeenCalled(); + }); + + it("Delete restores the original of an edited line; Esc leaves the layer", () => { + const change = makeSourceTextObject("c1", 0, INVOICE.rect, { + runId: INVOICE.id, + sourceFingerprint: "fp", + sourcePageIndex: 0, + originalText: INVOICE.text, + text: "Invoice 2027", + style: {}, + }); + const spies = mount([INVOICE, TOTAL], { objects: [change] }); + const edited = runButtons()[0]; + expect(edited.getAttribute("aria-label")).toBe("Edited text: Invoice 2027. Original: Invoice 2026"); + key(edited, "Delete"); + expect(spies.onRestore).toHaveBeenCalledWith(change); + key(buttonFor(TOTAL), "Backspace"); + expect(spies.onRestore).toHaveBeenCalledTimes(1); + key(buttonFor(TOTAL), "Escape"); + expect(spies.onLeave).toHaveBeenCalledTimes(1); + }); + + it("an edited line whose new text ends a sentence is labelled with one full stop (live check: “days..”)", () => { + const change = makeSourceTextObject("c1", 0, INVOICE.rect, { + runId: INVOICE.id, + sourceFingerprint: "fp", + sourcePageIndex: 0, + originalText: INVOICE.text, + text: "Paid within 30 days.", + style: {}, + }); + mount([INVOICE], { objects: [change] }); + const label = runButtons()[0].getAttribute("aria-label"); + expect(label).toBe("Edited text: Paid within 30 days. Original: Invoice 2026"); + expect(label).not.toContain(".."); + }); + + it("hides a refused line's tooltip while its reason dialog is open (live check: tooltip under the dialog)", () => { + vi.useFakeTimers(); + try { + const props = (popoverId: string | null) => + createElement(SourceTextLayer, { + runs: [SHARED], objects: [], fonts: fontsByKey([FONT]), layout: LAYOUT, pageNumber: 1, preview: null, duplicate: false, + focusedId: null, focusTick: 0, editingId: null, popoverId, stage: null, inert: false, + onFocusedChange: () => {}, onEdit: vi.fn(), onExplain: vi.fn(), onRestore: vi.fn(), onLeave: vi.fn(), + }); + stage = document.createElement("div"); + host = document.createElement("div"); + stage.appendChild(host); + document.body.appendChild(stage); + root = createRoot(host); + act(() => root!.render(props(null))); + // Hover the dotted line (CSS y = 792 − 650) until its tooltip shows. + act(() => { + layer().dispatchEvent(new PointerEvent("pointermove", { clientX: 80, clientY: 792 - 650, bubbles: true })); + }); + act(() => vi.advanceTimersByTime(450)); + const tip = () => host.querySelector('[role="tooltip"]'); + expect(tip()?.textContent).toBe("Can't be changed: shared with other pages"); + // A click opens the reason dialog for that line; the pointer is still over it. + act(() => root!.render(props(SHARED.id))); + act(() => vi.advanceTimersByTime(450)); + expect(tip()).toBeNull(); + // Closed again (still hovering): the tooltip may come back. + act(() => root!.render(props(null))); + expect(tip()).not.toBeNull(); + } finally { + vi.useRealTimers(); + } + }); + + it("renders only lines near the visible stage, but keeps the focused line rendered", () => { + mount([INVOICE, LOW], {}, 150); + expect(runButtons().map((b) => b.getAttribute("aria-label"))).toEqual(["Edit text: Invoice 2026"]); + act(() => root!.unmount()); + stage.remove(); + mount([INVOICE, LOW], { focusedId: LOW.id }, 150); + expect(runButtons().map((b) => b.getAttribute("aria-label"))).toEqual(["Edit text: Invoice 2026", "Edit text: Footer"]); + }); + + it("a click on the extended part of an edited line hits that line, caret within the new text", () => { + const newRect = { x: 72, y: 697, w: 120, h: 12 }; + const change: SourceTextObject = makeSourceTextObject("c1", 0, newRect, { + runId: INVOICE.id, + sourceFingerprint: "fp", + sourcePageIndex: 0, + originalText: INVOICE.text, + text: "Invoice 2026 (paid in full)", + style: {}, + }); + const spies = mount([INVOICE, TOTAL], { objects: [change] }); + // Original box ends at x = 72 + 72; the click is at x = 180, inside the new box only. + act(() => { + layer().dispatchEvent(new MouseEvent("click", { clientX: 180, clientY: 792 - 700 + 1, bubbles: true })); + }); + expect(spies.onEdit).toHaveBeenCalledTimes(1); + const [hit, caret] = (spies.onEdit as ReturnType).mock.calls[0]; + expect(hit).toBe(INVOICE); + expect(caret).toBeGreaterThan(INVOICE.text.length); + expect(caret).toBeLessThanOrEqual(Array.from(change.text).length); + }); + + it("a click on an unedited line puts the caret at the nearest boundary; a refused one explains", () => { + const spies = mount([INVOICE, SHARED]); + act(() => { + layer().dispatchEvent(new MouseEvent("click", { clientX: 72 + 13, clientY: 92, bubbles: true })); + }); + expect(spies.onEdit).toHaveBeenCalledWith(INVOICE, 2); + act(() => { + layer().dispatchEvent(new MouseEvent("click", { clientX: 80, clientY: 792 - 650, bubbles: true })); + }); + expect(spies.onExplain).toHaveBeenCalledWith(SHARED); + }); + + it("a page listed twice is shown refused, described by banner.duplicatePage, and never edited", () => { + const spies = mount([INVOICE], { duplicate: true }); + const b = runButtons()[0]; + expect(b.getAttribute("aria-label")).toBe("Invoice 2026"); + expect(b.getAttribute("aria-disabled")).toBe("true"); + expect(b.className).toContain("is-refused"); + expect(document.getElementById(b.getAttribute("aria-describedby")!)?.textContent).toContain(UI.banner.duplicatePage); + key(b, "Enter"); + act(() => { + layer().dispatchEvent(new MouseEvent("click", { clientX: 80, clientY: 92, bubbles: true })); + }); + expect(spies.onEdit).not.toHaveBeenCalled(); + expect(spies.onExplain).not.toHaveBeenCalled(); + }); + + it("Space on a focused line activates it and never reaches the page's Space-to-pan", () => { + const spies = mount([INVOICE]); + const pan = vi.fn(); + window.addEventListener("keydown", pan); + act(() => buttonFor(INVOICE).focus()); + const ev = key(buttonFor(INVOICE), " "); + window.removeEventListener("keydown", pan); + expect(spies.onEdit).toHaveBeenCalledWith(INVOICE, "all"); + expect(ev.defaultPrevented).toBe(true); + expect(pan).not.toHaveBeenCalled(); + expect(buttonFor(INVOICE).closest("[data-source-text]")).not.toBeNull(); + }); +}); diff --git a/src/components/pdf/editor/sourceText/SourceTextLayer.tsx b/src/components/pdf/editor/sourceText/SourceTextLayer.tsx new file mode 100644 index 0000000..a234e90 --- /dev/null +++ b/src/components/pdf/editor/sourceText/SourceTextLayer.tsx @@ -0,0 +1,336 @@ +/** + * The text lines of one page as real buttons in reading order (SPEC §D.5, + * §D.6, §D.9, §D.10): one Tab stop with roving focus, arrows move through + * `readingNeighbour`, Enter/F2/Space edit or explain, Delete restores, Esc + * leaves. The pointer is hit-tested by the layer itself (smallest box, then + * nearest baseline; ≥ 24 px tall hit areas), edited lines with their new + * geometry. Only lines near the visible part of the stage are rendered, plus + * the focused, open and edited ones. + */ +import { useEffect, useId, useLayoutEffect, useMemo, useRef, useState, type KeyboardEvent, type PointerEvent } from "react"; +import { Icon } from "@/components/ui/Icon"; +import { + caretIndexAt, + editedGeometry, + makeMapping, + pdfRectToViewport, + pdfToViewport, + problemMessage, + readingNeighbour, + viewportToPdf, + type CssRect, + type NeighbourKey, + type SourceTextObject, +} from "@/lib/editor"; +import { UI, editedRunLabel, fillCopy, reasonCopy, warningCopy } from "@/lib/editor/sourceTextCopy"; +import type { TextFont, TextPreview, TextRun } from "@/lib/types"; +import type { PageLayout } from "../PageSurface"; + +const MIN_HIT_PX = 24; +const VIEW_MARGIN_PX = 200; +const TOOLTIP_DELAY_MS = 400; +/** A click that arrives this soon after a key activation is the browser's echo of it. */ +const KEY_ECHO_MS = 500; + +const KEYS: Record = { + ArrowDown: "down", + ArrowUp: "up", + ArrowRight: "right", + ArrowLeft: "left", + Home: "home", + End: "end", + PageUp: "pageUp", + PageDown: "pageDown", +}; + +interface RunView { + run: TextRun; + obj: SourceTextObject | null; + css: CssRect; + hit: CssRect; + area: number; + caretOffsets: number[]; + origin: { x: number; y: number }; + dir: { x: number; y: number }; + label: string; + describe: string; + state: "editable" | "refused" | "edited"; + flagged: "warning" | "problem" | null; +} + +export interface SourceTextLayerProps { + runs: TextRun[]; + objects: SourceTextObject[]; + fonts: Map; + layout: PageLayout; + /** 1-based page number for the layer's label. */ + pageNumber: number; + /** The preview answer for the page's current changes, if any. */ + preview: TextPreview | null; + /** The page is listed twice: everything is shown refused and never edited. */ + duplicate: boolean; + focusedId: string | null; + /** Bumped to move DOM focus to `focusedId`. */ + focusTick: number; + editingId: string | null; + popoverId: string | null; + /** Scroll container, for rendering only what is near view. */ + stage: HTMLElement | null; + /** Space-to-pan is active: let the stage take the pointer. */ + inert: boolean; + onFocusedChange: (id: string, moveFocus: boolean) => void; + onEdit: (run: TextRun, caret: number | "all") => void; + onExplain: (run: TextRun) => void; + onRestore: (obj: SourceTextObject) => void; + onLeave: () => void; +} + +/** Text of the next run on the same line, for the overlap message. */ +export function nextOnLine(runs: TextRun[], run: TextRun): TextRun | null { + let best: TextRun | null = null; + for (const r of runs) if (r.line === run.line && r.order > run.order && (!best || r.order < best.order)) best = r; + return best; +} + +function contains(r: CssRect, p: { x: number; y: number }): boolean { + return p.x >= r.x && p.x <= r.x + r.w && p.y >= r.y && p.y <= r.y + r.h; +} + +function intersects(a: CssRect, b: CssRect): boolean { + return a.x < b.x + b.w && a.x + a.w > b.x && a.y < b.y + b.h && a.y + a.h > b.y; +} + +function hitTest(views: RunView[], p: { x: number; y: number }): RunView | null { + let best: RunView | null = null; + let bestDist = Infinity; + for (const v of views) { + if (!contains(v.hit, p)) continue; + const dist = Math.abs((p.x - v.origin.x) * v.dir.y - (p.y - v.origin.y) * v.dir.x); + if (!best || v.area < best.area - 0.5 || (Math.abs(v.area - best.area) <= 0.5 && dist < bestDist)) { + best = v; + bestDist = dist; + } + } + return best; +} + +type ViewInputs = Pick; + +function buildViews({ runs, objects, fonts, layout, preview, duplicate }: ViewInputs): RunView[] { + const mapping = makeMapping(layout.geometry, layout.cssWidth, layout.cssHeight); + const byRun = new Map(objects.map((o) => [o.runId, o])); + const verdicts = new Map((preview?.verdicts ?? []).map((v) => [v.runId, v])); + return runs + .slice() + .sort((a, b) => a.order - b.order) + .map((run) => { + const obj = byRun.get(run.id) ?? null; + const verdict = verdicts.get(run.id) ?? null; + const geom = obj ? editedGeometry(run, obj, verdict, fonts) : { rect: run.rect, caretOffsets: run.caretOffsets }; + const css = pdfRectToViewport(geom.rect, mapping); + const h = Math.max(css.h, MIN_HIT_PX); + const o = pdfToViewport(run.origin, mapping); + const o2 = pdfToViewport({ x: run.origin.x + run.dir.x, y: run.origin.y + run.dir.y }, mapping); + const len = Math.hypot(o2.x - o.x, o2.y - o.y) || 1; + const notes: string[] = []; + let flagged: RunView["flagged"] = null; + if (verdict && !verdict.ok && obj) { + notes.push(problemMessage(verdict, run, fonts)); + flagged = "problem"; + } + for (const w of preview?.warnings ?? []) { + if (w.runId !== run.id) continue; + notes.push(warningCopy(w.code, nextOnLine(runs, run)?.text)); + flagged = flagged ?? "warning"; + } + if (run.substituted) notes.push(UI.substituted); + const refused = duplicate || !run.editable; + if (duplicate) notes.unshift(UI.banner.duplicatePage); + else if (!run.editable) notes.unshift(fillCopy(UI.run.refusedDescribe, { short: reasonCopy(run.reason).short })); + return { + run, + obj, + css, + hit: { x: css.x, y: css.y + css.h / 2 - h / 2, w: css.w, h }, + area: css.w * css.h, + caretOffsets: geom.caretOffsets, + origin: o, + dir: { x: (o2.x - o.x) / len, y: (o2.y - o.y) / len }, + label: refused + ? run.text + : obj + ? editedRunLabel(obj.text, obj.originalText) + : fillCopy(UI.run.editable, { text: run.text }), + describe: notes.join(" "), + state: refused ? "refused" : obj ? "edited" : "editable", + flagged, + }; + }); +} + +export function SourceTextLayer(props: SourceTextLayerProps) { + const { layout, duplicate, focusedId, focusTick, stage } = props; + const uid = useId(); + const rootRef = useRef(null); + const buttons = useRef(new Map()); + const keyEcho = useRef(0); + const [hoverId, setHoverId] = useState(null); + const [tipId, setTipId] = useState(null); + const [view, setView] = useState(null); + + const { runs, objects, fonts, preview } = props; + const views = useMemo( + () => buildViews({ runs, objects, fonts, layout, preview, duplicate }), + [runs, objects, fonts, layout, preview, duplicate], + ); + const rovingId = views.some((v) => v.run.id === focusedId) ? focusedId : (views[0]?.run.id ?? null); + + useLayoutEffect(() => { + const root = rootRef.current; + if (!stage || !root) { + setView(null); + return; + } + let frame = 0; + const measure = () => { + frame = 0; + const s = stage.getBoundingClientRect(); + const l = root.getBoundingClientRect(); + const m = VIEW_MARGIN_PX; + setView({ x: s.left - l.left - m, y: s.top - l.top - m, w: s.width + 2 * m, h: s.height + 2 * m }); + }; + const schedule = () => { + if (!frame) frame = requestAnimationFrame(measure); + }; + measure(); + stage.addEventListener("scroll", schedule, { passive: true }); + window.addEventListener("resize", schedule); + return () => { + stage.removeEventListener("scroll", schedule); + window.removeEventListener("resize", schedule); + if (frame) cancelAnimationFrame(frame); + }; + }, [stage, layout]); + + // Move DOM focus only when asked (`focusTick`), never because the roving line changed. + const focusedRef = useRef(focusedId); + focusedRef.current = focusedId; + useEffect(() => { + const id = focusedRef.current; + if (!focusTick || !id) return; + const el = buttons.current.get(id); + el?.focus(); + el?.scrollIntoView?.({ block: "nearest", inline: "nearest" }); + }, [focusTick]); + + const tipText = (v: RunView | null): string | null => { + if (!v || duplicate) return null; + if (v.state === "refused") return fillCopy(UI.run.refusedDescribe, { short: reasonCopy(v.run.reason).short }); + return v.run.substituted ? UI.substituted : null; + }; + const hovered = views.find((v) => v.run.id === hoverId) ?? null; + const hoveredTip = tipText(hovered); + useEffect(() => { + setTipId(null); + if (!hoveredTip || !hoverId) return; + const timer = setTimeout(() => setTipId(hoverId), TOOLTIP_DELAY_MS); + return () => clearTimeout(timer); + }, [hoverId, hoveredTip]); + + const pinned = new Set([rovingId, props.editingId, props.popoverId, ...objects.map((o) => o.runId)]); + const shown = views.filter((v) => pinned.has(v.run.id) || !view || intersects(v.hit, view)); + + const activate = (v: RunView, caret: number | "all") => { + props.onFocusedChange(v.run.id, false); + if (duplicate) return; + if (v.run.editable) props.onEdit(v.run, caret); + else props.onExplain(v.run); + }; + + const local = (e: { clientX: number; clientY: number }) => { + const r = rootRef.current?.getBoundingClientRect(); + return { x: e.clientX - (r?.left ?? 0), y: e.clientY - (r?.top ?? 0) }; + }; + + const onRunKeyDown = (e: KeyboardEvent, v: RunView) => { + const move = KEYS[e.key]; + if (move) { + e.preventDefault(); + e.stopPropagation(); + const next = readingNeighbour(runs, v.run.id, move); + if (next) props.onFocusedChange(next, true); + } else if (e.key === "Enter" || e.key === "F2" || e.key === " ") { + e.preventDefault(); + e.stopPropagation(); + keyEcho.current = Date.now(); + activate(v, "all"); + } else if (e.key === "Delete" || e.key === "Backspace") { + e.preventDefault(); + e.stopPropagation(); + if (v.obj) props.onRestore(v.obj); + } else if (e.key === "Escape") { + e.preventDefault(); + e.stopPropagation(); + props.onLeave(); + } + }; + + const mapping = makeMapping(layout.geometry, layout.cssWidth, layout.cssHeight); + // The reason dialog opens where the tooltip sits (just under the line) and says more: no tooltip under it. + const tip = props.popoverId ? null : (views.find((v) => v.run.id === tipId) ?? null); + const cursor = hovered ? (hovered.state === "refused" ? " is-over-refused" : " is-over-editable") : ""; + + return ( +
) => setHoverId(hitTest(views, local(e))?.run.id ?? null)} + onPointerLeave={() => setHoverId(null)} + onClick={(e) => { + if (e.target !== e.currentTarget || e.button !== 0) return; + const p = local(e); + const v = hitTest(views, p); + if (v) activate(v, v.run.editable ? caretIndexAt(v.run, viewportToPdf(p, mapping), v.caretOffsets) : "all"); + }} + > + {shown.map((v, i) => ( + + ))} + + {tip && tipText(tip) && ( +
+ {tipText(tip)} +
+ )} +
+ ); +} diff --git a/src/components/pdf/editor/sourceText/SourceTextMode.test.tsx b/src/components/pdf/editor/sourceText/SourceTextMode.test.tsx new file mode 100644 index 0000000..ad437f4 --- /dev/null +++ b/src/components/pdf/editor/sourceText/SourceTextMode.test.tsx @@ -0,0 +1,328 @@ +// @vitest-environment happy-dom +import { StrictMode, act, createElement, createRef } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { renderToStaticMarkup } from "react-dom/server"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import { fontsByKey } from "@/lib/editor"; +import { UI, fillCopy } from "@/lib/editor/sourceTextCopy"; +import type { PageText, TextFont, TextPreview, TextRun } from "@/lib/types"; +import type { PageLayout } from "../PageSurface"; +import type { EditSession } from "../useEditSession"; +import { + PreviewStatusChip, + SourceTextBanners, + SourceTextMode, + type SourceTextGuard, + type SourceTextPage, + type TextRequest, +} from "./SourceTextMode"; +import type { TextPreviewController } from "./useTextPreview"; + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +const FONT: TextFont = { + key: "f1", + displayName: "Arial", + familyHint: "sans", + embedded: true, + subset: false, + alphabet: " !Hdelo", + widths: [278, 278, 722, 556, 556, 222, 556], + wordSpace: true, +}; +const LAYOUT: PageLayout = { cssWidth: 612, cssHeight: 792, geometry: { box: { x: 0, y: 0, w: 612, h: 792 }, rotate: 0, pageIndex: 0 } }; + +function run(id: string, order: number, text: string, y: number, editable = true): TextRun { + return { + id, + order, + line: order, + text, + rect: { x: 72, y: y - 3, w: 30, h: 12 }, + origin: { x: 72, y }, + dir: { x: 1, y: 0 }, + ascent: 9, + descent: 3, + caretOffsets: Array.from({ length: text.length + 1 }, (_, i) => i * 6), + editable, + reason: editable ? null : "ROTATED_TEXT", + metrics: editable + ? { + surface: ["f1"], tfSize: 12, effectiveSize: 12, charSpacing: 0, wordSpacing: 0, hScale: 1, textToUser: 1, + letterSpacingPt: 0, spaceMode: "glyph", kernSpace: -250, originalWidth: 30, visibleExtent: 500, nextObstacle: null, + } + : null, + style: editable + ? { + fill: "#000000", sizeChangeable: true, colourChangeable: true, face: "regular", + faces: { + regular: { available: true, surface: ["f1"] }, bold: { available: false, surface: [] }, + italic: { available: false, surface: [] }, boldItalic: { available: false, surface: [] }, + }, + } + : null, + substituted: false, + }; +} + +const HELLO = run("r1", 0, "Hello", 700); +const TILTED = run("r2", 1, "Rotated", 600, false); + +function pageText(runs: TextRun[], pageReason: PageText["pageReason"] = null): PageText { + return { fingerprint: "fp", pageIndex: 0, pageReason, runs, fonts: [FONT] }; +} + +function preview(extra: Partial = {}): TextPreviewController { + return { bytes: null, status: "idle", result: null, error: null, request: vi.fn(), seed: vi.fn(), ...extra }; +} + +function model(extra: Partial = {}): SourceTextPage { + return { + path: "/a.pdf", + fileName: "a.pdf", + pageIndex: 0, + sourcePageIndex: 0, + active: true, + duplicate: false, + source: { status: "ready", info: { fingerprint: "fp", pageCount: 1, warnings: [] }, error: null, stale: false }, + fingerprint: "fp", + pageText: { page: pageText([HELLO, TILTED]), status: "ready", error: null }, + fonts: fontsByKey([FONT]), + objects: [], + staleFingerprint: null, + preview: preview(), + canShowOriginal: false, + problem: null, + reportFault: vi.fn(), + ...extra, + }; +} + +/** Rendered text (entities decoded). */ +function textOf(markup: string): string { + const div = document.createElement("div"); + div.innerHTML = markup; + return div.textContent ?? ""; +} + +const banners = (page: SourceTextPage, dismissed = true) => + textOf( + renderToStaticMarkup( + {}} modeDismissed={dismissed} onDismissMode={() => {}} />, + ), + ); + +describe("SourceTextBanners and PreviewStatusChip", () => { + it("mode banner until dismissed; page refused, no text, none editable, duplicate page", () => { + expect(banners(model(), false)).toContain(UI.banner.mode); + expect(banners(model())).not.toContain(UI.banner.mode); + expect(banners(model({ pageText: { page: pageText([], "GEOMETRY"), status: "ready", error: null } }))).toContain( + "Nothing on this page can be changed: This page uses a custom unit size or unreadable page boxes, which OffPDF doesn't edit yet.", + ); + expect(banners(model({ pageText: { page: pageText([]), status: "ready", error: null } }))).toContain(UI.banner.noText); + expect(banners(model({ pageText: { page: pageText([TILTED]), status: "ready", error: null } }))).toContain(UI.banner.noneEditable); + expect(banners(model({ duplicate: true }))).toContain(UI.banner.duplicatePage); + }); + + it("file errors, reading status and the stale banner with its action", () => { + const encrypted = { code: "ENCRYPTED", title: "t", message: "OffPDF can't change text in a protected PDF.", suggestion: "Remove the password with Unlock PDF, then edit the unlocked copy." }; + const markup = banners(model({ source: { status: "error", info: null, error: encrypted, stale: false }, fingerprint: null })); + expect(markup).toContain( + "Edit text is off for “a.pdf”: OffPDF can't change text in a protected PDF. Remove the password with Unlock PDF, then edit the unlocked copy.", + ); + expect(banners(model({ pageText: { page: null, status: "loading", error: null } }))).toContain(UI.status.reading); + const stale = banners(model({ staleFingerprint: "fp-old", fingerprint: null })); + expect(stale).toContain(fillCopy(UI.banner.stale, { name: "a.pdf" })); + expect(stale).toContain(UI.banner.staleAction); + }); + + it("status chip: checking / showing changes / original / unavailable, only with changes", () => { + const chip = (status: TextPreviewController["status"], showOriginal = false, objects = 1) => + textOf(renderToStaticMarkup( + ({}) as never) })} + showOriginal={showOriginal} + />, + )); + expect(chip("pending")).toContain(UI.status.checking); + expect(chip("ready")).toContain(UI.status.showingChanges); + expect(chip("ready", true)).toContain(UI.status.showingOriginal); + expect(chip("unavailable")).toContain(UI.status.previewUnavailable); + expect(chip("ready", false, 0)).toBe(""); + }); +}); + +describe("SourceTextMode", () => { + let root: Root | null = null; + let host: HTMLElement; + afterEach(() => { + act(() => root?.unmount()); + root = null; + host.remove(); + }); + + function mount(page: SourceTextPage, request: TextRequest | null = null, onRequestHandled = vi.fn()) { + const session = { setSourceText: vi.fn(), revertSourceText: vi.fn() } as unknown as EditSession; + const guardRef = createRef() as { current: SourceTextGuard | null }; + const onAddTextHere = vi.fn(); + host = document.createElement("div"); + document.body.appendChild(host); + root = createRoot(host); + act(() => + root!.render( + createElement( + StrictMode, + null, + createElement(SourceTextMode, { + page, layout: LAYOUT, session, stage: null, inert: false, guardRef, request, onRequestHandled, onLeave: () => {}, onAddTextHere, + }), + ), + ), + ); + return { session, guardRef, onAddTextHere }; + } + + const lineButton = (label: string) => + [...host.querySelectorAll("button.st-run")].find((b) => b.getAttribute("aria-label")?.includes(label))!; + const press = (el: Element, key: string) => + act(() => { + el.dispatchEvent(new KeyboardEvent("keydown", { key, bubbles: true, cancelable: true })); + }); + + it("Enter opens the editor; Done commits once, closes, returns focus to the line and announces", async () => { + const answer: TextPreview = { + pagePdf: "UERG", + verdicts: [{ runId: "r1", ok: true, code: null, chars: [], reason: null, face: null, field: null, detail: null, deltaPt: 1.25, newRect: HELLO.rect, caretOffsets: [0, 6, 12, 18, 24, 30, 33] }], + pageProblem: null, + warnings: [], + }; + const request = vi.fn().mockResolvedValue(answer); + const { session, guardRef } = mount(model({ preview: preview({ request }) })); + press(lineButton("Hello"), "Enter"); + const input = host.querySelector("input.st-chip__input")!; + expect(input).not.toBeNull(); + expect(guardRef.current?.isEditing()).toBe(true); + const setter = Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, "value")!.set!; + act(() => { + setter.call(input, "Hello!"); + input.dispatchEvent(new Event("input", { bubbles: true })); + }); + press(input, "Enter"); + await act(async () => {}); + expect(session.setSourceText).toHaveBeenCalledTimes(1); + expect(host.querySelector("input.st-chip__input")).toBeNull(); + expect(document.activeElement).toBe(lineButton("Hello")); + expect(host.querySelector(".sr-only")!.textContent).toContain("Change applied. 1.25 pt wider than before."); + expect(guardRef.current?.isEditing()).toBe(false); + }); + + it("the guard keeps a blocked edit open (page or tool change waits)", async () => { + const { guardRef } = mount(model()); + press(lineButton("Hello"), "F2"); + const input = host.querySelector("input.st-chip__input")!; + const setter = Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, "value")!.set!; + act(() => { + setter.call(input, "Hello ğ"); + input.dispatchEvent(new Event("input", { bubbles: true })); + }); + let ok: boolean | undefined; + await act(async () => { + ok = await guardRef.current!.tryClose(); + }); + expect(ok).toBe(false); + expect(host.querySelector("input.st-chip__input")).not.toBeNull(); + }); + + it("a refused line opens the reason dialog; Esc returns focus to the line; Add text here adds a box", () => { + const { onAddTextHere } = mount(model()); + press(lineButton("Rotated"), "Enter"); + const dialog = host.querySelector('[role="dialog"]')!; + expect(dialog.getAttribute("aria-modal")).toBe("false"); + expect(dialog.textContent).toContain(UI.popover.title); + expect(dialog.textContent).toContain("This line doesn't run straight across the page as shown."); + expect(document.activeElement?.textContent).toBe(UI.popover.addTextHere); + press(document.activeElement!, "Escape"); + expect(host.querySelector('[role="dialog"]')).toBeNull(); + expect(document.activeElement).toBe(lineButton("Rotated")); + press(lineButton("Rotated"), "Enter"); + act(() => (document.activeElement as HTMLButtonElement).click()); + expect(onAddTextHere).toHaveBeenCalledTimes(1); + const stamp = onAddTextHere.mock.calls[0][0]; + expect(stamp.fontSize).toBe(12); + expect(stamp.rect.x).toBeCloseTo(TILTED.rect.x, 6); // over the line, grown to one line × 4 em + expect(stamp.rect.y + stamp.rect.h).toBeCloseTo(TILTED.rect.y + TILTED.rect.h, 6); + expect(stamp.rect.w).toBeCloseTo(48, 6); + expect(stamp.rect.h).toBeCloseTo(15.6, 6); + }); + + it("Add text here on a line turned at an angle adds a default one-line box at its start, not its whole box", () => { + // Live check: the 35° watermark's axis-aligned box (298 × 260 pt) became the new text box. + const deg = (35 * Math.PI) / 180; + const turned: TextRun = { ...TILTED, id: "r3", order: 2, line: 2, text: "CONFIDENTIAL", rect: { x: 150, y: 250, w: 298, h: 260 }, origin: { x: 160, y: 262 }, dir: { x: Math.cos(deg), y: Math.sin(deg) } }; + const { onAddTextHere } = mount(model({ pageText: { page: pageText([HELLO, turned]), status: "ready", error: null } })); + press(lineButton("CONFIDENTIAL"), "Enter"); + act(() => (document.activeElement as HTMLButtonElement).click()); + const stamp = onAddTextHere.mock.calls[0][0]; + expect(stamp.rect.w).toBeCloseTo(120, 6); + expect(stamp.rect.h).toBeCloseTo(15.6, 6); + expect(stamp.rect.x).toBeCloseTo(160, 6); + expect(stamp.rect.y + stamp.rect.h).toBeCloseTo(274, 6); + }); + + it("a sidebar/inspector request opens that line once and is reported handled (never replayed)", () => { + const handled = vi.fn(); + mount(model(), { runId: "r1", open: true, tick: 3 }, handled); + expect(host.querySelector("input.st-chip__input")).not.toBeNull(); + expect(handled).toHaveBeenCalledTimes(1); + }); + + it("acts on a second request in the same mount even when the canvas restarts its tick (focus from the list, then Edit line)", () => { + // Live check: the canvas clears a handled request and the next one starts again at tick 1, + // so "select the change in the list" followed by "Edit line" in the inspector did nothing. + const handled = vi.fn(); + const page = model(); + mount(page, { runId: "r1", open: false, tick: 1 }, handled); + expect(host.querySelector("input.st-chip__input")).toBeNull(); + const render = (request: TextRequest | null) => + act(() => + root!.render( + createElement( + StrictMode, + null, + createElement(SourceTextMode, { + page, layout: LAYOUT, session: { setSourceText: vi.fn(), revertSourceText: vi.fn() } as unknown as EditSession, + stage: null, inert: false, guardRef: { current: null }, request, onRequestHandled: handled, onLeave: () => {}, onAddTextHere: vi.fn(), + }), + ), + ), + ); + render(null); + render({ runId: "r1", open: true, tick: 1 }); + expect(host.querySelector("input.st-chip__input")).not.toBeNull(); + expect(handled).toHaveBeenCalledTimes(2); + }); + + it("Remove these text changes keeps focus in the editor instead of dropping it on ", () => { + // Live check: the banner and its button disappear with the changes, so focus fell to + // and ⌘Z (handled only while focus is inside the editor) no longer worked. + const remove = vi.fn(); + host = document.createElement("div"); + host.className = "pdf-editor"; + host.tabIndex = 0; + document.body.appendChild(host); + root = createRoot(host); + const page = model({ staleFingerprint: "fp-old", fingerprint: null }); + act(() => root!.render(createElement(SourceTextBanners, { page, onRemoveStale: remove, modeDismissed: true, onDismissMode: () => {} }))); + const action = host.querySelector(".st-banners__action")!; + action.focus(); + act(() => action.click()); + act(() => root!.render(createElement(SourceTextBanners, { page: { ...page, staleFingerprint: null }, onRemoveStale: remove, modeDismissed: true, onDismissMode: () => {} }))); + expect(remove).toHaveBeenCalledWith("fp-old"); + expect(document.activeElement).toBe(host); + }); + + it("renders nothing for a refused page (the banner explains)", () => { + mount(model({ pageText: { page: pageText([HELLO], "PAGE_TOO_COMPLEX"), status: "ready", error: null } })); + expect(host.querySelector(".st-layer")).toBeNull(); + }); +}); diff --git a/src/components/pdf/editor/sourceText/SourceTextMode.tsx b/src/components/pdf/editor/sourceText/SourceTextMode.tsx new file mode 100644 index 0000000..212bcd9 --- /dev/null +++ b/src/components/pdf/editor/sourceText/SourceTextMode.tsx @@ -0,0 +1,402 @@ +/** + * Edit text for the shown page (SPEC §D.1, §D.4, §D.5, §D.7). + * + * - `useSourceTextPage` (used by the canvas for every tool): opens the file for + * Edit text when needed, reads the page's lines, keeps the real preview of + * the page's committed changes, and turns STALE / PDF_NEEDS_REPAIR answers + * into the file's banner state. + * - `SourceTextBanners` (above the page) and `PreviewStatusChip` (on the page). + * - `SourceTextMode` (on the page, tool = Edit text): the line layer, the + * inline editor and the reason popover, plus the live announcements. + */ +import { useCallback, useEffect, useMemo, useRef, useState, type MutableRefObject } from "react"; +import { Alert } from "@/components/ui/Alert"; +import { Spinner } from "@/components/ui/Spinner"; +import { fontsByKey, makeMapping, pdfRectToViewport, textStampGeometry, type EditObject, type SourceTextObject } from "@/lib/editor"; +import type { TextStamp } from "@/lib/editor/sourceText"; +import { UI, fileBanner, fillCopy, reasonCopy } from "@/lib/editor/sourceTextCopy"; +import type { AppError, TextFont, TextRun } from "@/lib/types"; +import type { TextSourceState, TextSources } from "@/features/edit-pdf/useTextSources"; +import type { PageLayout } from "../PageSurface"; +import type { EditSession } from "../useEditSession"; +import { ReasonPopover } from "./ReasonPopover"; +import { SourceTextEditor, type TryDone } from "./SourceTextEditor"; +import { SourceTextLayer, nextOnLine } from "./SourceTextLayer"; +import { usePageText, type PageTextState } from "./usePageText"; +import { useTextPreview, type TextPreviewController } from "./useTextPreview"; + +const NO_CHANGES: SourceTextObject[] = []; +const FILE_FAULTS = new Set(["STALE", "PDF_NEEDS_REPAIR"]); + +/** Lets the canvas finish (or keep) an open line edit before a page or tool change. */ +export interface SourceTextGuard { + isEditing(): boolean; + tryClose(): Promise; +} + +/** Focus a line (and optionally open its editor) from the sidebar or inspector. */ +export interface TextRequest { + runId: string; + open: boolean; + tick: number; +} + +export interface SourceTextPageArgs { + /** Every object of the session (text changes on other pages decide staleness too). */ + objects: EditObject[]; + path: string; + fileName: string; + /** 1-based page in `path`. */ + sourcePage: number; + /** 0-based page in the editor. */ + pageIndex: number; + sources: TextSources; + /** The Edit text tool is active. */ + active: boolean; + /** The same source page is listed more than once. */ + duplicate: boolean; +} + +export interface SourceTextPage { + path: string; + fileName: string; + pageIndex: number; + sourcePageIndex: number; + active: boolean; + duplicate: boolean; + source: TextSourceState | null; + /** Fingerprint the page's changes and lines belong to, when usable. */ + fingerprint: string | null; + pageText: PageTextState; + fonts: Map; + /** This page's text changes. */ + objects: SourceTextObject[]; + /** Changes were made on a version of the file that is gone. */ + staleFingerprint: string | null; + preview: TextPreviewController; + /** A real preview of this page's text changes is on screen, so Show original has something to compare. */ + canShowOriginal: boolean; + /** An inspect or preview error that is not a file state (shown as is). */ + problem: AppError | null; + reportFault: (error: AppError) => void; +} + +export function useSourceTextPage(args: SourceTextPageArgs): SourceTextPage { + const { path, pageIndex, active, sources } = args; + const { ensure, markStale, markError, reopen } = sources; + const sourcePageIndex = args.sourcePage - 1; + const objects = useMemo( + () => args.objects.filter((o): o is SourceTextObject => o.kind === "sourceText" && o.pageIndex === pageIndex), + [args.objects, pageIndex], + ); + const source = sources.sources[path] ?? null; + const info = source?.status === "ready" ? source.info : null; + const wanted = active || objects.length > 0; + useEffect(() => { + if (wanted) void ensure(path); + }, [wanted, path, ensure]); + + const mismatch = info ? objects.find((o) => o.sourceFingerprint !== info.fingerprint) : undefined; + const staleFingerprint = source?.stale ? (info?.fingerprint ?? null) : (mismatch?.sourceFingerprint ?? null); + const fingerprint = info && !source?.stale && !mismatch ? info.fingerprint : null; + const pageText = usePageText({ + path, + fingerprint, + sourcePageIndex, + enabled: wanted, + pageCount: info?.pageCount ?? null, + }); + const preview = useTextPreview({ path, fingerprint, sourcePageIndex, objects: fingerprint ? objects : NO_CHANGES }); + + const reportFault = useCallback( + (error: AppError) => { + if (error.code === "STALE") markStale(path); + else if (error.code === "PDF_NEEDS_REPAIR") markError(path, error); + }, + [path, markStale, markError], + ); + const fault = pageText.error ?? preview.error; + useEffect(() => { + if (fault) reportFault(fault); + }, [fault, reportFault]); + + // A file that changed on disk with no text changes left on it is simply read again. + const changesOnStale = staleFingerprint + ? args.objects.some((o) => o.kind === "sourceText" && o.sourceFingerprint === staleFingerprint) + : false; + const stale = !!source?.stale; + useEffect(() => { + if (stale && !changesOnStale) void reopen(path); + }, [stale, changesOnStale, path, reopen]); + + const fonts = useMemo(() => fontsByKey(pageText.page?.fonts ?? []), [pageText.page]); + const problem = fault && !FILE_FAULTS.has(fault.code) ? fault : null; + return { + path, + fileName: args.fileName, + pageIndex, + sourcePageIndex, + active, + duplicate: args.duplicate, + source, + fingerprint, + pageText, + fonts, + objects, + staleFingerprint, + preview, + canShowOriginal: objects.length > 0 && preview.bytes !== null, + problem, + reportFault, + }; +} + +function errorText(error: AppError): string { + return [error.message, error.suggestion].filter(Boolean).join(" "); +} + +export function SourceTextBanners({ + page, + onRemoveStale, + modeDismissed, + onDismissMode, +}: { + page: SourceTextPage; + onRemoveStale: (fingerprint: string) => void; + modeDismissed: boolean; + onDismissMode: () => void; +}) { + const { source, pageText } = page; + const text = pageText.page; + const stale = page.staleFingerprint; + const fileError = source?.status === "error" ? source.error : null; + const loading = source?.status === "loading" || pageText.status === "loading"; + return ( +
+ {stale && (page.active || page.objects.length > 0) && ( + + {fillCopy(UI.banner.stale, { name: page.fileName })}{" "} + + + )} + {page.problem && (page.active || page.objects.length > 0) && ( + + {errorText(page.problem)} + + )} + {page.active && fileError && {fileBanner(page.fileName, fileError)}} + {page.active && !fileError && loading && ( +
+ {UI.status.reading} +
+ )} + {page.active && text && page.duplicate && {UI.banner.duplicatePage}} + {page.active && text?.pageReason && ( + {fillCopy(UI.banner.pageRefused, { body: reasonCopy(text.pageReason).body })} + )} + {page.active && text && !text.pageReason && text.runs.length === 0 && {UI.banner.noText}} + {page.active && text && !text.pageReason && text.runs.length > 0 && !text.runs.some((r) => r.editable) && ( + {UI.banner.noneEditable} + )} + {page.active && !modeDismissed && ( + + + {UI.banner.mode} + + + + )} +
+ ); +} + +/** Top-right of the page while it has text changes (§D.7). */ +export function PreviewStatusChip({ page, showOriginal }: { page: SourceTextPage; showOriginal: boolean }) { + if (page.objects.length === 0 || !page.fingerprint) return null; + const status = page.preview.status; + const label = showOriginal + ? UI.status.showingOriginal + : status === "pending" + ? UI.status.checking + : status === "ready" + ? UI.status.showingChanges + : status === "unavailable" + ? UI.status.previewUnavailable + : null; + if (!label) return null; + return ( +
+ {status === "pending" && !showOriginal && } + {label} +
+ ); +} + +export interface SourceTextModeProps { + page: SourceTextPage; + layout: PageLayout; + session: EditSession; + stage: HTMLElement | null; + /** Space-to-pan is active. */ + inert: boolean; + guardRef: MutableRefObject; + request: TextRequest | null; + /** The request was acted on; the canvas drops it so a later mount never replays it. */ + onRequestHandled: () => void; + /** Esc in the layer: focus goes back to the editor. */ + onLeave: () => void; + /** "Add text here": the new box for the refused line (`textStampGeometry`). */ + onAddTextHere: (stamp: TextStamp) => void; +} + +export function SourceTextMode(props: SourceTextModeProps) { + const { page, layout, session, stage, inert, guardRef, request, onRequestHandled, onLeave, onAddTextHere } = props; + const text = page.fingerprint ? page.pageText.page : null; + const runs = useMemo(() => text?.runs ?? [], [text]); + const [focus, setFocus] = useState<{ id: string | null; tick: number }>({ id: null, tick: 0 }); + const [editing, setEditing] = useState<{ runId: string; caret: number | "all" } | null>(null); + const [popoverId, setPopoverId] = useState(null); + const [said, setSaid] = useState({ text: "", n: 0 }); + const tryDone = useRef(null); + const editingRef = useRef(editing); + editingRef.current = editing; + const handled = useRef(null); // by identity: the canvas restarts its tick after clearing + + const announce = (sentence: string) => setSaid((s) => ({ text: sentence, n: s.n + 1 })); + const focusRun = useCallback((id: string, move: boolean) => setFocus((f) => ({ id, tick: move ? f.tick + 1 : f.tick })), []); + const finishEditing = useCallback( + () => (tryDone.current ? tryDone.current({ focusRun: false, blockedNotice: true }) : Promise.resolve(true)), + [], + ); + const registerTryDone = useCallback((fn: TryDone | null) => { + tryDone.current = fn; + }, []); + + useEffect(() => { + guardRef.current = { isEditing: () => editingRef.current !== null, tryClose: finishEditing }; + return () => { + guardRef.current = null; + }; + }, [guardRef, finishEditing]); + + const edit = useCallback( + async (run: TextRun, caret: number | "all") => { + if (editingRef.current?.runId === run.id) return; + if (editingRef.current && !(await finishEditing())) return; + setPopoverId(null); + setEditing({ runId: run.id, caret }); + focusRun(run.id, false); + }, + [finishEditing, focusRun], + ); + + const explain = async (run: TextRun) => { + if (editingRef.current && !(await finishEditing())) return; + setPopoverId(run.id); + focusRun(run.id, false); + }; + + useEffect(() => { + if (!request || request === handled.current || !text) return; + handled.current = request; + onRequestHandled(); + const run = text.runs.find((r) => r.id === request.runId); + if (!run) return; + if (request.open && run.editable && !page.duplicate) void edit(run, "all"); + else focusRun(run.id, true); + }, [request, text, page.duplicate, edit, focusRun, onRequestHandled]); + + if (!text || text.pageReason) return null; + const byRun = new Map(page.objects.map((o) => [o.runId, o])); + const editRun = editing ? (runs.find((r) => r.id === editing.runId) ?? null) : null; + const popRun = popoverId ? (runs.find((r) => r.id === popoverId) ?? null) : null; + const mapping = makeMapping(layout.geometry, layout.cssWidth, layout.cssHeight); + + return ( + <> + void edit(run, caret)} + onExplain={(run) => void explain(run)} + onRestore={(obj) => { + session.revertSourceText(obj.id); + announce(UI.announce.restored); + }} + onLeave={onLeave} + /> + {editRun && page.fingerprint && editing && ( + o.runId !== editRun.id)} + caret={editing.caret} + neighbourText={nextOnLine(runs, editRun)?.text ?? null} + requestPreview={page.preview.request} + seedPreview={page.preview.seed} + onApply={session.setSourceText} + onRevert={session.revertSourceText} + onFileError={page.reportFault} + registerTryDone={registerTryDone} + onClose={(outcome) => { + setEditing(null); + if (outcome.announce) announce(outcome.announce); + if (outcome.focusRun) focusRun(editRun.id, true); + }} + /> + )} + {popRun && ( + { + setPopoverId(null); + onAddTextHere(textStampGeometry(popRun, layout.geometry)); + }} + onClose={() => { + setPopoverId(null); + focusRun(popRun.id, true); + }} + /> + )} +
+ {said.text} + {said.n % 2 === 1 ? "​" : ""} +
+ + ); +} diff --git a/src/components/pdf/editor/sourceText/TextFormatBar.test.tsx b/src/components/pdf/editor/sourceText/TextFormatBar.test.tsx new file mode 100644 index 0000000..e3c1dc0 --- /dev/null +++ b/src/components/pdf/editor/sourceText/TextFormatBar.test.tsx @@ -0,0 +1,310 @@ +// @vitest-environment happy-dom +import { act, createElement } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import { fontsByKey } from "@/lib/editor"; +import { UI, charList } from "@/lib/editor/sourceTextCopy"; +import type { SourceTextStyle, TextFont, TextRun } from "@/lib/types"; +import { TextFormatBar, steppedSize, toggledFace, type TextFormatBarProps } from "./TextFormatBar"; + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +const SUBSET: TextFont = { + key: "f1", + displayName: "Calibri", + familyHint: "sans", + embedded: true, + subset: true, + alphabet: " Haelo", + widths: [226, 631, 479, 229, 229, 527], + wordSpace: true, +}; +const BOLD: TextFont = { ...SUBSET, key: "fb", alphabet: " Hael" }; +const FULL: TextFont = { ...SUBSET, key: "f9", subset: false }; + +function run(style: Partial> = {}, surface = ["f1"]): TextRun { + return { + id: "r1", + order: 0, + line: 0, + text: "Hello", + rect: { x: 72, y: 697, w: 30, h: 12 }, + origin: { x: 72, y: 700 }, + dir: { x: 1, y: 0 }, + ascent: 9, + descent: 3, + caretOffsets: [0, 6, 12, 18, 24, 30], + editable: true, + reason: null, + metrics: { + surface, + tfSize: 12, + effectiveSize: 12, + charSpacing: 0, + wordSpacing: 0, + hScale: 1, + textToUser: 1, + letterSpacingPt: 0, + spaceMode: "glyph", + kernSpace: -250, + originalWidth: 30, + visibleExtent: 400, + nextObstacle: null, + }, + style: { + fill: "#000000", + sizeChangeable: true, + colourChangeable: true, + face: "regular", + faces: { + regular: { available: true, surface }, + bold: { available: true, surface: ["fb"] }, + italic: { available: false, surface: [] }, + boldItalic: { available: false, surface: [] }, + }, + ...style, + }, + substituted: false, + }; +} + +let root: Root | null = null; +let host: HTMLElement; + +function mount(r: TextRun, style: SourceTextStyle = {}, extra: Partial = {}) { + const props: TextFormatBarProps = { + run: r, + fonts: fontsByKey([SUBSET, BOLD, FULL]), + style, + onStyle: vi.fn(), + onNotice: vi.fn(), + canRestore: false, + onRestore: vi.fn(), + onCancel: vi.fn(), + onDone: vi.fn(), + onBackToInput: vi.fn(), + placement: "above", + ...extra, + }; + host = document.createElement("div"); + document.body.appendChild(host); + root = createRoot(host); + act(() => root!.render(createElement(TextFormatBar, props))); + return props; +} + +afterEach(() => { + act(() => root?.unmount()); + root = null; + host.remove(); +}); + +const button = (label: string) => host.querySelector(`button[aria-label="${label}"]`)!; +const key = (el: Element, k: string, init: KeyboardEventInit = {}) => + act(() => { + el.dispatchEvent(new KeyboardEvent("keydown", { key: k, bubbles: true, ...init })); + }); + +describe("TextFormatBar", () => { + it("is a toolbar named Text format", () => { + mount(run()); + const bar = host.querySelector('[role="toolbar"]')!; + expect(bar.getAttribute("aria-label")).toBe(UI.bar.label); + }); + + it("B8: an unavailable italic stays focusable and writes the exact italic sentence", () => { + const props = mount(run()); + const italic = button(UI.bar.italic); + expect(italic.getAttribute("aria-disabled")).toBe("true"); + expect(italic.getAttribute("title")).toBe("This page has no italic version of this font."); + expect(italic.disabled).toBe(false); + act(() => italic.click()); + expect(props.onNotice).toHaveBeenCalledWith("This page has no italic version of this font."); + expect(props.onStyle).not.toHaveBeenCalled(); + }); + + it("B9: a font set through ExtGState makes B and I write style.face, and size write style.size", () => { + const props = mount(run({ sizeChangeable: false })); + act(() => button(UI.bar.bold).click()); + act(() => button(UI.bar.italic).click()); + expect(props.onNotice).toHaveBeenNthCalledWith(1, UI.style.face); + expect(props.onNotice).toHaveBeenNthCalledWith(2, UI.style.face); + act(() => button(UI.bar.larger).click()); + expect(props.onNotice).toHaveBeenLastCalledWith(UI.style.size); + expect(host.querySelector(`input[aria-label="${UI.bar.fontSize}"]`)).toBeNull(); + }); + + it("an available bold switches the face; aria-pressed follows the face", () => { + const props = mount(run()); + expect(button(UI.bar.bold).getAttribute("aria-pressed")).toBe("false"); + act(() => button(UI.bar.bold).click()); + expect(props.onStyle).toHaveBeenCalledWith({ face: "bold" }); + act(() => root!.unmount()); + host.remove(); + mount(run({ face: "bold" })); + expect(button(UI.bar.bold).getAttribute("aria-pressed")).toBe("true"); + }); + + it("colour unavailable (outline text) writes style.colour", () => { + const props = mount(run({ colourChangeable: false })); + act(() => button(UI.bar.textColour).click()); + expect(props.onNotice).toHaveBeenCalledWith(UI.style.colour); + }); + + it("every button keeps the input's focus on mouse down", () => { + mount(run(), {}, { canRestore: true }); + const all = [...host.querySelectorAll("button")]; + expect(all.length).toBeGreaterThan(8); + for (const b of all) { + const down = new MouseEvent("mousedown", { bubbles: true, cancelable: true }); + act(() => { + b.dispatchEvent(down); + }); + expect(down.defaultPrevented, b.getAttribute("aria-label") ?? b.textContent ?? "").toBe(true); + } + }); + + it("roving tabindex: one tab stop, ←/→ move focus, Esc and Shift+Tab go back to the input", () => { + const props = mount(run()); + const items = () => [...host.querySelectorAll("[data-roving]")]; + expect(items().filter((el) => el.tabIndex === 0)).toHaveLength(1); + const bold = button(UI.bar.bold); + act(() => bold.focus()); + key(bold, "ArrowRight"); + expect(document.activeElement).toBe(button(UI.bar.italic)); + key(document.activeElement!, "ArrowLeft"); + expect(document.activeElement).toBe(bold); + expect(items().filter((el) => el.tabIndex === 0)).toEqual([bold]); + key(bold, "Escape"); + key(bold, "Tab", { shiftKey: true }); + expect(props.onBackToInput).toHaveBeenCalledTimes(2); + }); + + it("Letters: only for an embedded subset, listing the current face's alphabet with space named", () => { + mount(run()); + const letters = [...host.querySelectorAll("button")].find((b) => b.textContent === UI.bar.letters)!; + expect(letters).toBeDefined(); + act(() => letters.click()); + const pop = host.querySelector(".st-bar__letters")!; + expect(pop.textContent).toContain(UI.bar.lettersTitle); + expect(pop.textContent).toContain(charList(Array.from(" Haelo"))); + expect(pop.textContent).toContain("space, H, a, e, l, o"); + + act(() => root!.unmount()); + host.remove(); + mount(run(), { face: "bold" }); // the draft's face decides which letters can be typed + const boldLetters = [...host.querySelectorAll("button")].find((b) => b.textContent === UI.bar.letters)!; + act(() => boldLetters.click()); + expect(host.querySelector(".st-bar__letters")!.textContent).toContain("space, H, a, e, l"); + expect(host.querySelector(".st-bar__letters")!.textContent).not.toContain("l, o"); + + act(() => root!.unmount()); + host.remove(); + mount(run({}, ["f9"])); + expect([...host.querySelectorAll("button")].some((b) => b.textContent === UI.bar.letters)).toBe(false); + }); + + describe("a popover never covers the line being edited (live check: Letters over the inline box)", () => { + // The real layout, measured: the stage, the bar docked above the line, the line (chip) and its + // message row under it. happy-dom has no layout, so the boxes are given here. + type Box = { top: number; bottom: number }; + const boxes = new Map(); + let restore: Array<() => void> = []; + afterEach(() => { + restore.forEach((r) => r()); + restore = []; + boxes.clear(); + }); + + function mountDocked(barTop: number): HTMLElement { + boxes.set("pdf-editor__stage", { top: 0, bottom: 600 }); + boxes.set("st-bar", { top: barTop, bottom: barTop + 36 }); + boxes.set("st-chip", { top: barTop + 42, bottom: barTop + 60 }); + boxes.set("st-msg", { top: barTop + 64, bottom: barTop + 82 }); + const rect = vi.spyOn(Element.prototype, "getBoundingClientRect").mockImplementation(function (this: Element) { + const hit = [...boxes.entries()].find(([cls]) => this.classList.contains(cls))?.[1] ?? { top: 0, bottom: 0 }; + return { x: 0, y: hit.top, left: 0, right: 400, top: hit.top, bottom: hit.bottom, width: 400, height: hit.bottom - hit.top, toJSON: () => ({}) }; + }); + const height = vi.spyOn(HTMLElement.prototype, "offsetHeight", "get").mockImplementation(function (this: HTMLElement) { + return this.classList.contains("st-bar__pop") ? 80 : 0; + }); + restore = [() => rect.mockRestore(), () => height.mockRestore()]; + const stage = document.createElement("div"); + stage.className = "pdf-editor__stage"; + const editor = document.createElement("div"); + editor.className = "st-editor"; + const dock = document.createElement("div"); + dock.className = "st-dock st-dock--above"; + const chip = document.createElement("div"); + chip.className = "st-chip"; + const msg = document.createElement("div"); + msg.className = "st-msg"; + editor.append(dock, chip, msg); + stage.append(editor); + host = stage; + document.body.appendChild(stage); + root = createRoot(dock); + const props: TextFormatBarProps = { + run: run(), fonts: fontsByKey([SUBSET, BOLD, FULL]), style: {}, onStyle: vi.fn(), onNotice: vi.fn(), canRestore: false, + onRestore: vi.fn(), onCancel: vi.fn(), onDone: vi.fn(), onBackToInput: vi.fn(), placement: "above", + }; + act(() => root!.render(createElement(TextFormatBar, props))); + const letters = [...stage.querySelectorAll("button")].find((b) => b.textContent === UI.bar.letters)!; + act(() => letters.click()); + return stage.querySelector(".st-bar__letters")!; + } + + it("without room above the bar, it opens below the line and its message row, not over them", () => { + const pop = mountDocked(40); // 40 px above the bar: an 80 px popover can't go up + expect(pop.classList.contains("is-up")).toBe(false); + // Bar bottom 76; the message row reaches 122: 46 px to clear, plus the 4 px gap. + expect(pop.style.top).toBe("calc(100% + 50px)"); + }); + + it("with room above the bar, it opens upward, away from the line", () => { + const pop = mountDocked(300); + expect(pop.classList.contains("is-up")).toBe(true); + expect(pop.style.top).toBe(""); + }); + }); + + it("colour popover: Original colour drops the fill, an ink sets it", () => { + const props = mount(run(), { fill: "#c71c1c" }); + act(() => button(UI.bar.textColour).click()); + const original = [...host.querySelectorAll("button")].find((b) => b.textContent === UI.bar.originalColour)!; + act(() => original.click()); + expect(props.onStyle).toHaveBeenLastCalledWith({}); + act(() => button(UI.bar.blue).click()); + expect(props.onStyle).toHaveBeenLastCalledWith({ fill: "#1c4fb8" }); + }); + + it("Restore original appears only for a committed change; Cancel and Done call back", () => { + const props = mount(run(), {}, { canRestore: true }); + const byText = (t: string) => [...host.querySelectorAll("button")].find((b) => b.textContent?.includes(t))!; + act(() => byText(UI.bar.restore).click()); + act(() => byText(UI.bar.cancel).click()); + act(() => byText(UI.bar.done).click()); + expect(props.onRestore).toHaveBeenCalledTimes(1); + expect(props.onCancel).toHaveBeenCalledTimes(1); + expect(props.onDone).toHaveBeenCalledTimes(1); + act(() => root!.unmount()); + host.remove(); + mount(run()); + expect([...host.querySelectorAll("button")].some((b) => b.textContent === UI.bar.restore)).toBe(false); + }); +}); + +describe("format helpers (shared with ⌘B / ⌘I / ⌘⇧. / ⌘⇧,)", () => { + it("steps the effective size by ±0.5 within 4–144", () => { + expect(steppedSize(run(), {}, 0.5)).toEqual({ style: { sizePt: 12.5 } }); + expect(steppedSize(run(), { sizePt: 4.2 }, -0.5)).toEqual({ style: { sizePt: 4 } }); + expect(steppedSize(run(), { sizePt: 144 }, 0.5)).toEqual({ style: { sizePt: 144 } }); + }); + + it("toggling back to the run's own face is always allowed", () => { + expect(toggledFace(run({ face: "italic" }), {}, "italic")).toEqual({ + style: { face: "regular" }, + }); + expect(toggledFace(run({ face: "bold" }), { face: "regular" }, "bold")).toEqual({ style: { face: "bold" } }); + }); +}); diff --git a/src/components/pdf/editor/sourceText/TextFormatBar.tsx b/src/components/pdf/editor/sourceText/TextFormatBar.tsx new file mode 100644 index 0000000..50fdd66 --- /dev/null +++ b/src/components/pdf/editor/sourceText/TextFormatBar.tsx @@ -0,0 +1,359 @@ +/** + * Format bar of the inline editor (SPEC §D.6.2): size, bold/italic, colour, + * letter spacing, the subset's letters, Restore original / Cancel / Done. + * A roving-tabindex toolbar (←/→ move, Enter/Space activate, Esc/Shift+Tab back + * to the input). Buttons never take focus from the input on mouse down. An + * unavailable control stays focusable (`aria-disabled`) and, when activated, + * writes its exact reason into the editor's message row (B8, B9). + */ +import { useEffect, useLayoutEffect, useRef, useState, type CSSProperties, type KeyboardEvent, type MouseEvent } from "react"; +import { Icon } from "@/components/ui/Icon"; +import { + LETTER_SPACING_MAX_PT, + LETTER_SPACING_MIN_PT, + LETTER_SPACING_STEP_PT, + SIZE_MAX_PT, + SIZE_MIN_PT, + SIZE_STEP_PT, + TEXT_INKS, + surfaceFor, +} from "@/lib/editor"; +import { UI, charList, problemCopy } from "@/lib/editor/sourceTextCopy"; +import type { SourceTextStyle, TextFace, TextFont, TextRun } from "@/lib/types"; +import { ColorField, DraftNumber } from "../controls"; + +/** A style edit, or the sentence saying why it is not available. */ +export interface StyleChange { + style?: SourceTextStyle; + notice?: string; +} + +const round2 = (n: number) => Math.round(n * 100) / 100; +/** Space between the bar (or what it must clear) and an open popover, CSS px. */ +const POP_GAP_PX = 4; +const MIN_FONT_PX = 12; +const FAMILY: Record = { + serif: '"Times New Roman", "Liberation Serif", Georgia, serif', + sans: 'Arial, "Liberation Sans", "Helvetica Neue", sans-serif', + mono: '"Courier New", "Liberation Mono", monospace', +}; +const clamp = (n: number, min: number, max: number) => Math.min(max, Math.max(min, n)); +const keep = (e: MouseEvent) => e.preventDefault(); + +export function currentFace(run: TextRun, style: SourceTextStyle): TextFace { + return style.face ?? run.style?.face ?? "regular"; +} + +function flags(face: TextFace): { bold: boolean; italic: boolean } { + return { bold: face === "bold" || face === "boldItalic", italic: face === "italic" || face === "boldItalic" }; +} + +function faceOf(bold: boolean, italic: boolean): TextFace { + if (bold) return italic ? "boldItalic" : "bold"; + return italic ? "italic" : "regular"; +} + +/** Why `run` can't switch to `face`, or null (its own face is always fine). */ +function faceProblem(run: TextRun, face: TextFace): string | null { + if (!run.style || face === run.style.face) return null; + if (!run.style.sizeChangeable) return problemCopy("STYLE_UNAVAILABLE", [], { field: "face" }); + if (!run.style.faces[face].available) return problemCopy("FACE_UNAVAILABLE", [], { face }); + return null; +} + +function toggleTarget(run: TextRun, style: SourceTextStyle, which: "bold" | "italic"): TextFace { + const f = flags(currentFace(run, style)); + return which === "bold" ? faceOf(!f.bold, f.italic) : faceOf(f.bold, !f.italic); +} + +/** Bold or italic toggled (⌘B / ⌘I and the B / I buttons). */ +export function toggledFace(run: TextRun, style: SourceTextStyle, which: "bold" | "italic"): StyleChange { + const target = toggleTarget(run, style, which); + const problem = faceProblem(run, target); + return problem ? { notice: problem } : { style: { ...style, face: target } }; +} + +/** Size stepped by `delta` effective points (⌘⇧. / ⌘⇧, and Smaller / Larger). */ +export function steppedSize(run: TextRun, style: SourceTextStyle, delta: number): StyleChange { + if (!run.metrics || !run.style?.sizeChangeable) return { notice: problemCopy("STYLE_UNAVAILABLE", [], { field: "size" }) }; + const now = style.sizePt ?? run.metrics.effectiveSize; + return { style: { ...style, sizePt: round2(clamp(now + delta, SIZE_MIN_PT, SIZE_MAX_PT)) } }; +} + +/** + * How the inline input draws the draft (§D.6.1): family by the font's hint, the + * draft's face, size (never below 12 px; `floored` says it is shown larger), + * letter spacing and colour (the run's fill unless changed). + */ +export function chipFont( + run: TextRun, + fonts: Map, + style: SourceTextStyle, + pxPerPt: number, +): { css: CSSProperties; floored: boolean } { + const metrics = run.metrics; + const face = currentFace(run, style); + const sizePx = (style.sizePt ?? metrics?.effectiveSize ?? MIN_FONT_PX) * pxPerPt; + const fill = style.fill ?? run.style?.fill; + return { + floored: sizePx < MIN_FONT_PX, + css: { + fontFamily: FAMILY[fonts.get(metrics?.surface[0] ?? "")?.familyHint ?? "sans"], + fontSize: Math.max(MIN_FONT_PX, sizePx), + fontWeight: flags(face).bold ? 700 : 400, + fontStyle: flags(face).italic ? "italic" : "normal", + letterSpacing: (style.letterSpacingPt ?? metrics?.letterSpacingPt ?? 0) * pxPerPt, + color: fill ?? undefined, + }, + }; +} + +/** Space between the top of the visible stage and `el`, or Infinity outside a stage. */ +function roomAbove(el: HTMLElement | null): number { + const stage = el?.closest(".pdf-editor__stage"); + return el && stage ? el.getBoundingClientRect().top - stage.getBoundingClientRect().top : Infinity; +} + +/** How far below the bar's bottom the line being edited and its message row reach (CSS px, ≥ 0). */ +function editorBelow(bar: HTMLElement): number { + const editor = bar.closest(".st-editor"); + if (!editor) return 0; + const bottom = bar.getBoundingClientRect().bottom; + const reach = [...editor.querySelectorAll(".st-chip, .st-msg")].map((el) => el.getBoundingClientRect().bottom - bottom); + return Math.ceil(Math.max(0, ...reach)); +} + +function without(style: SourceTextStyle, key: keyof SourceTextStyle): SourceTextStyle { + const { [key]: _dropped, ...rest } = style; + return rest; +} + +/** Every character the face can type, in font order, space included once. */ +function alphabetOf(run: TextRun, fonts: Map, face: TextFace): string[] { + const seen = new Set(); + for (const key of surfaceFor(run, face)) for (const ch of Array.from(fonts.get(key)?.alphabet ?? "")) seen.add(ch); + return [...seen]; +} + +type Pop = "colour" | "spacing" | "letters"; + +export interface TextFormatBarProps { + run: TextRun; + fonts: Map; + style: SourceTextStyle; + onStyle: (next: SourceTextStyle) => void; + onNotice: (sentence: string) => void; + canRestore: boolean; + onRestore: () => void; + onCancel: () => void; + onDone: () => void; + onBackToInput: () => void; + placement: "above" | "below"; +} + +export function TextFormatBar(props: TextFormatBarProps) { + const { run, fonts, style, onStyle, onNotice } = props; + const rootRef = useRef(null); + const [active, setActive] = useState("smaller"); + const popRef = useRef(null); + const [open, setOpen] = useState<{ pop: Pop; byKey: boolean } | null>(null); + /** Where the open popover sits: above the bar, or below it and clear of the line being edited. */ + const [place, setPlace] = useState({ up: false, drop: 0 }); + const [custom, setCustom] = useState(false); + const metrics = run.metrics; + const runStyle = run.style; + + useEffect(() => { + if (open?.byKey) rootRef.current?.querySelector(".st-bar__pop button, .st-bar__pop input")?.focus(); + }, [open]); + + // Measured after every render while open (its height and the message row can change): up only from a + // bar above the line with room for it in the visible stage; otherwise down, below the line and its + // message row, so the popover never covers the text being edited. + useLayoutEffect(() => { + const bar = rootRef.current; + const pop = popRef.current; + if (!open || !bar || !pop) return; + const up = props.placement === "above" && roomAbove(bar) >= pop.offsetHeight + POP_GAP_PX; + const drop = up ? 0 : editorBelow(bar); + if (up !== place.up || drop !== place.drop) setPlace({ up, drop }); + }); + + if (!metrics || !runStyle) return null; + const face = currentFace(run, style); + const f = flags(face); + const size = style.sizePt ?? metrics.effectiveSize; + const spacing = style.letterSpacingPt ?? metrics.letterSpacingPt; + const fill = style.fill ?? runStyle.fill; + const primary = fonts.get(metrics.surface[0] ?? ""); + const showLetters = !!primary && primary.embedded && primary.subset; + const keys = ["smaller", ...(runStyle.sizeChangeable ? ["size"] : []), "larger", "bold", "italic", "colour", "spacing"] + .concat(showLetters ? ["letters"] : [], props.canRestore ? ["restore"] : [], ["cancel", "done"]); + const rovingKey = keys.includes(active) ? active : keys[0]; + + const focusKey = (key: string) => { + rootRef.current?.querySelector(`[data-roving="${key}"]`)?.focus(); + setActive(key); + }; + const item = (key: string) => ({ "data-roving": key, tabIndex: key === rovingKey ? 0 : -1, onMouseDown: keep }); + const apply = (change: StyleChange) => (change.style ? onStyle(change.style) : change.notice && onNotice(change.notice)); + const toggle = (pop: Pop, e: MouseEvent) => setOpen((cur) => (cur?.pop === pop ? null : { pop, byKey: e.detail === 0 })); + const popProps = (extra = "") => ({ + ref: popRef, + className: `st-bar__pop${extra}${place.up ? " is-up" : ""}`, + style: !place.up && place.drop > 0 ? { top: `calc(100% + ${place.drop + POP_GAP_PX}px)` } : undefined, + }); + const faceTitle = (which: "bold" | "italic") => faceProblem(run, toggleTarget(run, style, which)); + const sizeProblem = runStyle.sizeChangeable ? null : problemCopy("STYLE_UNAVAILABLE", [], { field: "size" }); + const colourProblem = runStyle.colourChangeable ? null : problemCopy("STYLE_UNAVAILABLE", [], { field: "colour" }); + + const spacingStep = (sign: -1 | 1) => ( + + ); + + const onKeyDown = (e: KeyboardEvent) => { + const target = e.target as HTMLElement; + const inPop = !!target.closest(".st-bar__pop"); + if (e.key === "Escape") { + e.preventDefault(); + e.stopPropagation(); + if (open) { + const trigger = open.pop; + setOpen(null); + focusKey(trigger); + } else props.onBackToInput(); + return; + } + if (e.key === "Tab" && e.shiftKey && !inPop) { + e.preventDefault(); + props.onBackToInput(); + return; + } + if ((e.key !== "ArrowLeft" && e.key !== "ArrowRight") || inPop) return; + if (target instanceof HTMLInputElement) { + const atEdge = e.key === "ArrowLeft" ? target.selectionStart === 0 : target.selectionEnd === target.value.length; + if (!atEdge) return; + } + e.preventDefault(); + const at = keys.indexOf(rovingKey); + focusKey(keys[(at + (e.key === "ArrowRight" ? 1 : keys.length - 1)) % keys.length]); + }; + + return ( +
{ + const key = (e.target as HTMLElement).dataset?.roving; + if (key) setActive(key); + }} + > + + {runStyle.sizeChangeable ? ( + + focusKey("size")} live + onCommit={(n) => onStyle({ ...style, sizePt: n })} /> + + ) : ( + {round2(size)} pt + )} + + + {(["bold", "italic"] as const).map((which) => { + const problem = faceTitle(which); + return ( + + ); + })} + + + {showLetters && ( + + )} + + {props.canRestore && ( + + )} + + + + {open?.pop === "colour" && ( +
+
+ + {TEXT_INKS.map((ink) => ( + +
+ {custom && ( + i.hex)} + fallback={TEXT_INKS[0].hex} onChange={(hex) => onStyle({ ...style, fill: hex })} /> + )} +
+ )} + {open?.pop === "spacing" && ( +
+ {spacingStep(-1)} + onStyle({ ...style, letterSpacingPt: n })} /> + {spacingStep(1)} +
+ )} + {open?.pop === "letters" && ( +
+
{UI.bar.lettersTitle}
+

{charList(alphabetOf(run, fonts, face))}

+
+ )} +
+ ); +} diff --git a/src/components/pdf/editor/sourceText/sourceTextDock.test.ts b/src/components/pdf/editor/sourceText/sourceTextDock.test.ts new file mode 100644 index 0000000..ac25b99 --- /dev/null +++ b/src/components/pdf/editor/sourceText/sourceTextDock.test.ts @@ -0,0 +1,35 @@ +// Live-check regression (v0.4 Edit text): the inline editor's message row sits over the line +// below the chip. It must not swallow clicks there, or clicking that line does nothing at all +// instead of trying Done (B11: "click another run → try Done"). Only the format bar takes the +// pointer. happy-dom does no hit testing, so this pins the stylesheet contract; the live +// Playwright run (scratchpad live-check) covers the behaviour. +import { readFileSync } from "node:fs"; +import { join } from "node:path"; +import { describe, expect, it } from "vitest"; + +const CSS = readFileSync(join(process.cwd(), "src/styles/source-text.css"), "utf8").replace(/\/\*[\s\S]*?\*\//g, ""); + +/** Declarations of every rule whose selector list is exactly `selector`. */ +function declarations(selector: string): string[] { + const out: string[] = []; + for (const m of CSS.matchAll(/([^{}]+)\{([^{}]*)\}/g)) { + if (m[1].trim() === selector) out.push(...m[2].split(";").map((d) => d.trim()).filter(Boolean)); + } + return out; +} + +function pointerEvents(selector: string): string | undefined { + const d = declarations(selector).filter((x) => x.startsWith("pointer-events")); + return d.length ? d[d.length - 1].split(":")[1].trim() : undefined; +} + +describe("Edit text editor docks", () => { + it("let the pointer through the message row", () => { + expect(pointerEvents(".st-dock")).toBe("none"); + expect(pointerEvents(".st-msg")).toBeUndefined(); + }); + + it("keep the format bar clickable inside either dock", () => { + expect(pointerEvents(".st-dock > .st-bar")).toBe("auto"); + }); +}); diff --git a/src/components/pdf/editor/sourceText/usePageText.test.tsx b/src/components/pdf/editor/sourceText/usePageText.test.tsx new file mode 100644 index 0000000..ae44ecb --- /dev/null +++ b/src/components/pdf/editor/sourceText/usePageText.test.tsx @@ -0,0 +1,122 @@ +// @vitest-environment happy-dom +import { act, createElement } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import type { AppError, PageText } from "@/lib/types"; + +const backend = vi.hoisted(() => ({ + inspectTextPage: vi.fn<(path: string, fp: string, page: number) => Promise>(), +})); +vi.mock("@/lib/tauriCommands", () => backend); + +import { PAGE_TEXT_CACHE_PAGES, PAGE_TEXT_PREFETCH_MS, lruPut, usePageText, type PageTextArgs, type PageTextState } from "./usePageText"; + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +const page = (pageIndex: number, fingerprint = "fp"): PageText => ({ fingerprint, pageIndex, pageReason: null, runs: [], fonts: [] }); + +let root: Root | null = null; +let state: PageTextState; +let render: (args: PageTextArgs) => void; + +function mount(args: PageTextArgs) { + function Probe(props: PageTextArgs) { + state = usePageText(props); + return null; + } + root = createRoot(document.createElement("div")); + render = (next) => act(() => root!.render(createElement(Probe, next))); + render(args); +} + +const at = (sourcePageIndex: number, extra: Partial = {}): PageTextArgs => ({ + path: "/a.pdf", + fingerprint: "fp", + sourcePageIndex, + enabled: true, + pageCount: 3, + ...extra, +}); + +beforeEach(() => { + vi.useFakeTimers(); + backend.inspectTextPage.mockReset(); + backend.inspectTextPage.mockImplementation((_p, fp, i) => Promise.resolve(page(i, fp))); +}); +afterEach(() => { + act(() => root?.unmount()); + root = null; + vi.useRealTimers(); +}); + +describe("usePageText", () => { + it("is idle (no call) while disabled or without a fingerprint", () => { + mount(at(0, { enabled: false })); + expect(state.status).toBe("idle"); + render(at(0, { fingerprint: null })); + expect(state.status).toBe("idle"); + expect(backend.inspectTextPage).not.toHaveBeenCalled(); + }); + + it("reads the page, then prefetches the next one after 300 ms; the next page comes from the cache", async () => { + mount(at(0)); + expect(state.status).toBe("loading"); + await act(async () => {}); + expect(state.status).toBe("ready"); + expect(state.page?.pageIndex).toBe(0); + expect(backend.inspectTextPage).toHaveBeenCalledTimes(1); + await act(async () => { + vi.advanceTimersByTime(PAGE_TEXT_PREFETCH_MS); + }); + expect(backend.inspectTextPage).toHaveBeenLastCalledWith("/a.pdf", "fp", 1); + render(at(1)); + expect(state.status).toBe("ready"); + expect(backend.inspectTextPage).toHaveBeenCalledTimes(2); + }); + + it("leaving the page before 300 ms cancels the prefetch; the last page has nothing to prefetch", async () => { + mount(at(0)); + await act(async () => {}); + render(at(2)); + await act(async () => {}); + await act(async () => { + vi.advanceTimersByTime(PAGE_TEXT_PREFETCH_MS * 2); + }); + expect(backend.inspectTextPage.mock.calls.map((c) => c[2])).toEqual([0, 2]); + }); + + it("drops a late answer for a page the user left", async () => { + let finish!: (p: PageText) => void; + backend.inspectTextPage.mockImplementationOnce(() => new Promise((res) => (finish = res))); + mount(at(0, { pageCount: 1 })); + render(at(0, { pageCount: 1, fingerprint: "fp2" })); + await act(async () => {}); + expect(state.page?.fingerprint).toBe("fp2"); + await act(async () => finish(page(0, "fp"))); + expect(state.page?.fingerprint).toBe("fp2"); + }); + + it("keeps the backend error (STALE) and rejects an answer for another page", async () => { + const stale: AppError = { code: "STALE", title: "The PDF changed on disk", message: "m" }; + backend.inspectTextPage.mockRejectedValueOnce(stale); + mount(at(0, { pageCount: 1 })); + await act(async () => {}); + expect(state.status).toBe("error"); + expect(state.error).toEqual(stale); + + backend.inspectTextPage.mockResolvedValueOnce(page(5)); + render(at(0, { pageCount: 1, fingerprint: "other" })); + await act(async () => {}); + expect(state.status).toBe("error"); + expect(state.error?.details).toContain("unexpected answer"); + }); + + it("LRU keeps the 16 newest pages", () => { + const cache = new Map(); + for (let i = 0; i < PAGE_TEXT_CACHE_PAGES + 2; i++) lruPut(cache, `k${i}`, i, PAGE_TEXT_CACHE_PAGES); + expect(cache.size).toBe(PAGE_TEXT_CACHE_PAGES); + expect(cache.has("k0")).toBe(false); + expect(cache.has("k1")).toBe(false); + expect(cache.has(`k${PAGE_TEXT_CACHE_PAGES + 1}`)).toBe(true); + }); +}); diff --git a/src/components/pdf/editor/sourceText/usePageText.ts b/src/components/pdf/editor/sourceText/usePageText.ts new file mode 100644 index 0000000..ec901dd --- /dev/null +++ b/src/components/pdf/editor/sourceText/usePageText.ts @@ -0,0 +1,125 @@ +/** + * The lines of one source page for Edit text (SPEC §D.4): `inspect_text_page` + * for the shown page, cached in an LRU of 16 pages keyed + * `${fingerprint}#${pageIndex}`. The next page is prefetched 300 ms after the + * shown one is ready; leaving the page before then cancels it. An answer for a + * page the user already left is dropped. + */ +import { useEffect, useRef, useState } from "react"; +import { inspectTextPage } from "@/lib/tauriCommands"; +import { toAppError, type AppError, type PageText } from "@/lib/types"; + +export const PAGE_TEXT_CACHE_PAGES = 16; +export const PAGE_TEXT_PREFETCH_MS = 300; + +export interface PageTextArgs { + path: string | null; + fingerprint: string | null; + /** 0-based page in the source file. */ + sourcePageIndex: number | null; + enabled: boolean; + /** Source page count; the next page is prefetched only when it exists. */ + pageCount?: number | null; +} + +export type PageTextStatus = "idle" | "loading" | "ready" | "error"; + +export interface PageTextState { + page: PageText | null; + status: PageTextStatus; + error: AppError | null; +} + +const IDLE: PageTextState = { page: null, status: "idle", error: null }; +const LOADING: PageTextState = { page: null, status: "loading", error: null }; + +/** Map-backed LRU: reading or writing a key makes it the newest. */ +export function lruGet(cache: Map, key: string): V | undefined { + const value = cache.get(key); + if (value === undefined) return undefined; + cache.delete(key); + cache.set(key, value); + return value; +} + +export function lruPut(cache: Map, key: string, value: V, max: number): void { + cache.delete(key); + cache.set(key, value); + while (cache.size > max) { + const oldest = cache.keys().next().value; + if (oldest === undefined) break; + cache.delete(oldest); + } +} + +/** The answer is for the page that was asked for and has the expected lists. */ +function isPageText(value: unknown, fingerprint: string, pageIndex: number): value is PageText { + if (typeof value !== "object" || value === null) return false; + const v = value as Partial; + return v.fingerprint === fingerprint && v.pageIndex === pageIndex && Array.isArray(v.runs) && Array.isArray(v.fonts); +} + +function unexpectedAnswer(): AppError { + return { ...toAppError(null), details: "inspect_text_page returned an unexpected answer" }; +} + +export function usePageText({ path, fingerprint, sourcePageIndex, enabled, pageCount }: PageTextArgs): PageTextState { + const cache = useRef(new Map()); + const [state, setState] = useState({ ...IDLE, key: null }); + const key = enabled && path && fingerprint && sourcePageIndex !== null ? `${fingerprint}#${sourcePageIndex}` : null; + + useEffect(() => { + if (!key || !path || !fingerprint || sourcePageIndex === null) { + setState({ ...IDLE, key: null }); + return; + } + const hit = lruGet(cache.current, key); + if (hit) { + setState({ page: hit, status: "ready", error: null, key }); + return; + } + let active = true; + setState({ ...LOADING, key }); + inspectTextPage(path, fingerprint, sourcePageIndex).then( + (page) => { + if (!active) return; + if (!isPageText(page, fingerprint, sourcePageIndex)) { + setState({ page: null, status: "error", error: unexpectedAnswer(), key }); + return; + } + lruPut(cache.current, key, page, PAGE_TEXT_CACHE_PAGES); + setState({ page, status: "ready", error: null, key }); + }, + (e: unknown) => { + if (active) setState({ page: null, status: "error", error: toAppError(e), key }); + }, + ); + return () => { + active = false; + }; + }, [key, path, fingerprint, sourcePageIndex]); + + const ready = state.key === key && state.status === "ready"; + useEffect(() => { + if (!ready || !path || !fingerprint || sourcePageIndex === null || pageCount == null) return; + const next = sourcePageIndex + 1; + const nextKey = `${fingerprint}#${next}`; + if (next >= pageCount || cache.current.has(nextKey)) return; + const timer = setTimeout(() => { + inspectTextPage(path, fingerprint, next).then( + (page) => { + if (isPageText(page, fingerprint, next)) lruPut(cache.current, nextKey, page, PAGE_TEXT_CACHE_PAGES); + }, + () => { + // Dropped on purpose: a prefetch is a guess. When the user opens that + // page, the same request runs again and its error is shown there. + }, + ); + }, PAGE_TEXT_PREFETCH_MS); + return () => clearTimeout(timer); + }, [ready, path, fingerprint, sourcePageIndex, pageCount]); + + if (!key) return IDLE; + if (state.key !== key) return LOADING; + return { page: state.page, status: state.status, error: state.error }; +} diff --git a/src/components/pdf/editor/sourceText/useTextPreview.test.tsx b/src/components/pdf/editor/sourceText/useTextPreview.test.tsx new file mode 100644 index 0000000..60067a5 --- /dev/null +++ b/src/components/pdf/editor/sourceText/useTextPreview.test.tsx @@ -0,0 +1,187 @@ +// @vitest-environment happy-dom +import { act, createElement } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import { editsSignature, makeSourceTextObject, type SourceTextObject } from "@/lib/editor"; +import type { AppError, TextEditIn, TextPreview } from "@/lib/types"; + +const backend = vi.hoisted(() => ({ + previewTextEdits: vi.fn<(path: string, fp: string, page: number, edits: TextEditIn[]) => Promise>(), +})); +vi.mock("@/lib/tauriCommands", () => backend); +vi.mock("@/lib/pdfjs", () => ({ base64ToBytes: (b64: string) => new TextEncoder().encode(`pdf:${b64}`) })); + +import { useTextPreview, type TextPreviewArgs, type TextPreviewController } from "./useTextPreview"; + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +function deferred() { + let resolve!: (v: T) => void; + let reject!: (e: unknown) => void; + const promise = new Promise((res, rej) => { + resolve = res; + reject = rej; + }); + return { promise, resolve, reject }; +} + +function change(runId: string, text: string): SourceTextObject { + return makeSourceTextObject(`id-${runId}`, 0, { x: 0, y: 0, w: 10, h: 10 }, { + runId, + sourceFingerprint: "fp", + sourcePageIndex: 0, + originalText: "old", + text, + style: {}, + }); +} + +function preview(pagePdf: string | null, runId = "r1"): TextPreview { + return { + pagePdf, + verdicts: [ + { runId, ok: true, code: null, chars: [], reason: null, face: null, field: null, detail: null, deltaPt: 0, newRect: null, caretOffsets: null }, + ], + pageProblem: null, + warnings: [], + }; +} + +const decode = (bytes: Uint8Array | null) => (bytes ? new TextDecoder().decode(bytes) : null); + +let root: Root | null = null; +let hook: TextPreviewController; +let render: (args: TextPreviewArgs) => void; + +function mount(args: TextPreviewArgs) { + function Probe(props: TextPreviewArgs) { + hook = useTextPreview(props); + return null; + } + const container = document.createElement("div"); + root = createRoot(container); + render = (next) => act(() => root!.render(createElement(Probe, next))); + render(args); +} + +const base = { path: "/a.pdf", fingerprint: "fp", sourcePageIndex: 0 }; + +beforeEach(() => { + backend.previewTextEdits.mockReset(); + // Calls a test does not answer stay pending. + backend.previewTextEdits.mockImplementation(() => new Promise(() => {})); +}); +afterEach(() => { + act(() => root?.unmount()); + root = null; +}); + +describe("useTextPreview", () => { + it("no changes ⇒ no bytes (the original render) and no call", () => { + mount({ ...base, objects: [] }); + expect(hook.bytes).toBeNull(); + expect(hook.status).toBe("idle"); + expect(backend.previewTextEdits).not.toHaveBeenCalled(); + }); + + it("previews the page's changes and decodes the page once ready", async () => { + const d = deferred(); + backend.previewTextEdits.mockReturnValueOnce(d.promise); + mount({ ...base, objects: [change("r1", "new")] }); + expect(hook.status).toBe("pending"); + expect(backend.previewTextEdits).toHaveBeenCalledWith("/a.pdf", "fp", 0, [ + { runId: "r1", originalText: "old", text: "new", style: {} }, + ]); + await act(async () => d.resolve(preview("AAA"))); + expect(hook.status).toBe("ready"); + expect(decode(hook.bytes)).toBe("pdf:AAA"); + expect(hook.result?.verdicts[0].ok).toBe(true); + }); + + it("pagePdf null ⇒ unavailable, original bytes", async () => { + backend.previewTextEdits.mockResolvedValueOnce(preview(null)); + mount({ ...base, objects: [change("r1", "new")] }); + await act(async () => {}); + expect(hook.status).toBe("unavailable"); + expect(hook.bytes).toBeNull(); + }); + + it("generation guard: a late answer for an older set of changes is dropped", async () => { + const first = deferred(); + const second = deferred(); + backend.previewTextEdits.mockReturnValueOnce(first.promise).mockReturnValueOnce(second.promise); + mount({ ...base, objects: [change("r1", "A")] }); + render({ ...base, objects: [change("r1", "B")] }); + await act(async () => second.resolve(preview("BBB"))); + await act(async () => first.resolve(preview("AAA"))); + expect(decode(hook.bytes)).toBe("pdf:BBB"); + expect(hook.status).toBe("ready"); + }); + + it("caches by signature: going back to an earlier set does not call again", async () => { + backend.previewTextEdits.mockResolvedValueOnce(preview("AAA")).mockResolvedValueOnce(preview("BBB")); + mount({ ...base, objects: [change("r1", "A")] }); + await act(async () => {}); + render({ ...base, objects: [change("r1", "B")] }); + await act(async () => {}); + render({ ...base, objects: [change("r1", "A")] }); + await act(async () => {}); + expect(backend.previewTextEdits).toHaveBeenCalledTimes(2); + expect(decode(hook.bytes)).toBe("pdf:AAA"); + }); + + it("keeps showing the last bytes of the same page while a new set is checked, never another page's", async () => { + const pending = deferred(); + backend.previewTextEdits.mockResolvedValueOnce(preview("AAA")).mockReturnValueOnce(pending.promise); + mount({ ...base, objects: [change("r1", "A")] }); + await act(async () => {}); + render({ ...base, objects: [change("r1", "B")] }); + expect(hook.status).toBe("pending"); + expect(decode(hook.bytes)).toBe("pdf:AAA"); + render({ ...base, sourcePageIndex: 1, objects: [change("r1", "B")] }); + expect(hook.bytes).toBeNull(); + }); + + it("a source page listed twice never shows the other listing's bytes while pending", async () => { + const pending = deferred(); + backend.previewTextEdits.mockResolvedValueOnce(preview("PAGE0")).mockReturnValueOnce(pending.promise); + mount({ ...base, objects: [change("r1", "A")] }); + await act(async () => {}); + expect(decode(hook.bytes)).toBe("pdf:PAGE0"); + const onOtherListing = { ...change("r1", "B"), pageIndex: 5 }; + render({ ...base, objects: [onOtherListing] }); + expect(hook.status).toBe("pending"); + expect(hook.bytes).toBeNull(); + }); + + it("seed: the commit's own preview is shown without a second call", async () => { + mount({ ...base, objects: [] }); + const next = [change("r1", "new")]; + act(() => hook.seed(editsSignature(next), preview("SEED"))); + render({ ...base, objects: next }); + await act(async () => {}); + expect(backend.previewTextEdits).not.toHaveBeenCalled(); + expect(decode(hook.bytes)).toBe("pdf:SEED"); + }); + + it("request() checks a candidate set without changing the shown page; errors reject as AppError", async () => { + backend.previewTextEdits.mockResolvedValueOnce(preview("CAND")); + mount({ ...base, objects: [] }); + const answer = await hook.request([{ runId: "r1", originalText: "old", text: "x", style: {} }]); + expect(answer.pagePdf).toBe("CAND"); + expect(hook.bytes).toBeNull(); + const stale: AppError = { code: "STALE", title: "The PDF changed on disk", message: "m" }; + backend.previewTextEdits.mockRejectedValueOnce(stale); + await expect(hook.request([])).rejects.toEqual(stale); + }); + + it("a backend error is kept (never swallowed) and shows the original", async () => { + const err: AppError = { code: "VERIFIER_MISSING", title: "A checking component is missing", message: "m" }; + backend.previewTextEdits.mockRejectedValueOnce(err); + mount({ ...base, objects: [change("r1", "new")] }); + await act(async () => {}); + expect(hook.status).toBe("error"); + expect(hook.error).toEqual(err); + expect(hook.bytes).toBeNull(); + }); +}); diff --git a/src/components/pdf/editor/sourceText/useTextPreview.ts b/src/components/pdf/editor/sourceText/useTextPreview.ts new file mode 100644 index 0000000..125d593 --- /dev/null +++ b/src/components/pdf/editor/sourceText/useTextPreview.ts @@ -0,0 +1,169 @@ +/** + * The real render of a page with its committed text changes (SPEC §D.4, §D.7). + * + * Whenever the page's set of changes differs (commit, restore, undo, redo, + * page change) the page is previewed through `preview_text_edits` — the same + * plan + qpdf write + checks as Save — unless that exact set is cached (LRU of + * 8, keyed by `editsSignature`). A generation counter drops answers for a set + * that is no longer current. No changes ⇒ no bytes (the original render). + * + * `request` runs a preview for a candidate set without touching the shown + * page (the inline editor's commit); `seed` stores its answer so the page + * shows it as soon as the commit lands, without a second call. + */ +import { useCallback, useEffect, useRef, useState } from "react"; +import { editsSignature, toTextEditIn, type SourceTextObject } from "@/lib/editor"; +import { base64ToBytes } from "@/lib/pdfjs"; +import { previewTextEdits } from "@/lib/tauriCommands"; +import { toAppError, type AppError, type TextEditIn, type TextPreview } from "@/lib/types"; +import { lruGet, lruPut } from "./usePageText"; + +export const PREVIEW_CACHE_PAGES = 8; + +export type TextPreviewStatus = "idle" | "pending" | "ready" | "unavailable" | "error"; + +export interface TextPreviewArgs { + path: string | null; + fingerprint: string | null; + /** 0-based page in the source file. */ + sourcePageIndex: number | null; + /** The page's committed text changes. */ + objects: SourceTextObject[]; +} + +export interface TextPreviewState { + /** One-page PDF to render instead of the original, or null. */ + bytes: Uint8Array | null; + status: TextPreviewStatus; + /** The answer for the current set of changes. */ + result: TextPreview | null; + error: AppError | null; +} + +export interface TextPreviewController extends TextPreviewState { + request(edits: TextEditIn[]): Promise; + seed(signature: string, preview: TextPreview): void; +} + +interface Entry { + preview: TextPreview; + bytes: Uint8Array | null; +} + +interface Snapshot { + /** Page identity (`path`, fingerprint, page); bytes never cross pages. */ + base: string | null; + owner: number; + key: string | null; + status: TextPreviewStatus; + entry: Entry | null; + /** Bytes shown while the current set is pending (same page only). */ + shown: Uint8Array | null; + error: AppError | null; +} + +const IDLE: Snapshot = { base: null, owner: -1, key: null, status: "idle", entry: null, shown: null, error: null }; + +function isPreview(value: unknown): value is TextPreview { + if (typeof value !== "object" || value === null) return false; + const v = value as Partial; + return Array.isArray(v.verdicts) && Array.isArray(v.warnings) && (v.pagePdf === null || typeof v.pagePdf === "string"); +} + +function unexpectedAnswer(): AppError { + return { ...toAppError(null), details: "preview_text_edits returned an unexpected answer" }; +} + +function toEntry(preview: TextPreview): Entry { + return { preview, bytes: preview.pagePdf ? base64ToBytes(preview.pagePdf) : null }; +} + +const statusOf = (entry: Entry): TextPreviewStatus => (entry.bytes ? "ready" : "unavailable"); + +export function useTextPreview({ path, fingerprint, sourcePageIndex, objects }: TextPreviewArgs): TextPreviewController { + const cache = useRef(new Map()); + const generation = useRef(0); + const objectsRef = useRef(objects); + objectsRef.current = objects; + const [snap, setSnap] = useState(IDLE); + + const base = path && fingerprint && sourcePageIndex !== null ? `${path}\u0000${fingerprint}\u0000${sourcePageIndex}` : null; + /** Editor page the changes belong to: a source page listed twice never shows the other copy's bytes. */ + const owner = objects[0]?.pageIndex ?? -1; + const signature = editsSignature(objects); + const key = base && objects.length > 0 ? `${base}\u0000${signature}` : null; + const baseRef = useRef(base); + baseRef.current = base; + const keyRef = useRef(key); + keyRef.current = key; + const ownerRef = useRef(owner); + ownerRef.current = owner; + + useEffect(() => { + const gen = ++generation.current; + if (!key || !base || !path || !fingerprint || sourcePageIndex === null) { + setSnap(IDLE); + return; + } + const hit = lruGet(cache.current, key); + if (hit) { + setSnap({ base, owner, key, status: statusOf(hit), entry: hit, shown: hit.bytes, error: null }); + return; + } + setSnap((prev) => ({ + base, + owner, + key, + status: "pending", + entry: null, + shown: prev.base === base && prev.owner === owner ? prev.shown : null, + error: null, + })); + previewTextEdits(path, fingerprint, sourcePageIndex, objectsRef.current.map(toTextEditIn)).then( + (preview) => { + if (gen !== generation.current) return; + if (!isPreview(preview)) { + setSnap({ base, owner, key, status: "error", entry: null, shown: null, error: unexpectedAnswer() }); + return; + } + const entry = toEntry(preview); + lruPut(cache.current, key, entry, PREVIEW_CACHE_PAGES); + setSnap({ base, owner, key, status: statusOf(entry), entry, shown: entry.bytes, error: null }); + }, + (e: unknown) => { + if (gen !== generation.current) return; + setSnap({ base, owner, key, status: "error", entry: null, shown: null, error: toAppError(e) }); + }, + ); + }, [key, base, owner, path, fingerprint, sourcePageIndex]); + + const request = useCallback( + (edits: TextEditIn[]): Promise => { + if (!path || !fingerprint || sourcePageIndex === null) { + return Promise.reject({ ...toAppError(null), details: "Edit text source is not open" }); + } + return previewTextEdits(path, fingerprint, sourcePageIndex, edits).then((preview) => { + if (!isPreview(preview)) throw unexpectedAnswer(); + return preview; + }); + }, + [path, fingerprint, sourcePageIndex], + ); + + const seed = useCallback((sig: string, preview: TextPreview) => { + const pageBase = baseRef.current; + if (!pageBase || !isPreview(preview)) return; + const seededKey = `${pageBase}\u0000${sig}`; + const entry = toEntry(preview); + lruPut(cache.current, seededKey, entry, PREVIEW_CACHE_PAGES); + if (seededKey === keyRef.current) { + generation.current += 1; + setSnap({ base: pageBase, owner: ownerRef.current, key: seededKey, status: statusOf(entry), entry, shown: entry.bytes, error: null }); + } + }, []); + + if (!key || snap.key !== key) { + return { bytes: key && snap.base === base && snap.owner === owner ? snap.shown : null, status: key ? "pending" : "idle", result: null, error: null, request, seed }; + } + return { bytes: snap.shown, status: snap.status, result: snap.entry?.preview ?? null, error: snap.error, request, seed }; +} diff --git a/src/components/pdf/editor/useEditSession.ts b/src/components/pdf/editor/useEditSession.ts index 03ce55e..16cff8a 100644 --- a/src/components/pdf/editor/useEditSession.ts +++ b/src/components/pdf/editor/useEditSession.ts @@ -19,6 +19,7 @@ import { makeUnderlineObject, mapPointsToRect, planKeyRebind, + setSourceTextAction, type ClosedShapeKind, type EditDocument, type EditObject, @@ -27,8 +28,10 @@ import { type LayerDir, type PdfRect, type Point, + type SetSourceTextInput, } from "@/lib/editor"; + function newId(): string { if (typeof crypto !== "undefined" && "randomUUID" in crypto) { return crypto.randomUUID(); @@ -85,6 +88,30 @@ export function useEditSession( dispatch({ type: "ADD", object: makeTextObject(newId(), pageIndex, rect, content) }); }, []); + /** "Add text here" from the reason popover: a new text box (from `textStampGeometry`). Returns its id. */ + const addTextStamp = useCallback((pageIndex: number, rect: PdfRect, fontSize: number): string => { + const id = newId(); + dispatch({ type: "ADD", object: { ...makeTextObject(id, pageIndex, rect), fontSize } }); + return id; + }, []); + + /** Commit a text change: upsert by (pageIndex, runId), or remove it when it is a no-op. One history step. */ + const setSourceText = useCallback((input: SetSourceTextInput) => { + dispatch(setSourceTextAction(input, newId())); + }, []); + + /** Restore the original of one line (deletes its text change). */ + const revertSourceText = useCallback((id: string) => { + const obj = stateRef.current.present.objects.find((o) => o.id === id); + if (!obj || obj.kind !== "sourceText") return; + dispatch({ type: "DELETE", ids: [id] }); + }, []); + + /** Drop every text change made on one snapshot of a file (STALE recovery). One history step. */ + const removeSourceTextForFingerprint = useCallback((fingerprint: string) => { + dispatch({ type: "REMOVE_SOURCE_TEXT_FOR_FINGERPRINT", fingerprint }); + }, []); + const addImage = useCallback( (pageIndex: number, rect: PdfRect, path: string, previewUrl?: string) => { dispatch({ type: "ADD", object: makeImageObject(newId(), pageIndex, rect, path, previewUrl) }); @@ -298,6 +325,10 @@ export function useEditSession( addRect, addShape, addText, + addTextStamp, + setSourceText, + revertSourceText, + removeSourceTextForFingerprint, addImage, addLine, addInk, diff --git a/src/components/ui/Icon.tsx b/src/components/ui/Icon.tsx index 036d6ba..c892d10 100644 --- a/src/components/ui/Icon.tsx +++ b/src/components/ui/Icon.tsx @@ -61,6 +61,7 @@ export type IconName = | "mousePointer" | "hand" | "type" + | "textCursor" | "square" | "rectangle" | "squareFill" @@ -337,6 +338,15 @@ const PATHS: Record = { ), + /** I-beam with a small pencil: Edit text (change words already on the page). */ + textCursor: ( + <> + + + + + + ), square: , rectangle: , squareFill: , diff --git a/src/features/edit-pdf/EditPdfPage.tsx b/src/features/edit-pdf/EditPdfPage.tsx index 2dfac2b..08b9e2f 100644 --- a/src/features/edit-pdf/EditPdfPage.tsx +++ b/src/features/edit-pdf/EditPdfPage.tsx @@ -38,6 +38,9 @@ import { import { isTauriRuntime } from "@/lib/tauriEnv"; import { toAppError, type AppError } from "@/lib/types"; import { Alert } from "@/components/ui/Alert"; +import { UI, jobLabel } from "@/lib/editor/sourceTextCopy"; +import { useTextSources } from "./useTextSources"; +import { finishOpenTextEdit, textSaveGuard, type OpenTextEdit } from "./textSaveGuards"; const tool = getTool("editPdf"); @@ -80,6 +83,34 @@ export function EditPdfPage() { const [formValues, setFormValues] = useState>({}); const [formError, setFormError] = useState(null); const [flattenForm, setFlattenForm] = useState(false); + const textSources = useTextSources({ + onReleaseError: (_path, err) => toast({ title: err.title, description: err.message, variant: "error" }), + }); + const { release: releaseTextSource, reopen: reopenTextSource } = textSources; + // The open Edit text line (owned here so Save can finish it before reading the objects). + const textGuard = useRef(null); + const [saveAfterTextEdit, setSaveAfterTextEdit] = useState(false); + const textChangeCount = doc.objects.filter((o) => o.kind === "sourceText").length; + + // A file that left the workspace no longer needs its Edit text snapshot. + const filePaths = useMemo(() => new Set(files.map((f) => f.path)), [files]); + const openText = textSources.sources; + useEffect(() => { + for (const path of Object.keys(openText)) if (!filePaths.has(path)) releaseTextSource(path); + }, [filePaths, openText, releaseTextSource]); + + // Save said a file changed on disk: read every file with text changes again. A changed + // file gets a new fingerprint, so its changes show the stale banner and its action. + const jobError = job.error; + useEffect(() => { + if (jobError?.code !== "STALE") return; + const paths = new Set(); + for (const o of doc.objects) { + const ref = o.kind === "sourceText" ? refs[o.pageIndex] : undefined; + if (ref) paths.add(ref.path); + } + for (const path of paths) void reopenTextSource(path); + }, [jobError, reopenTextSource]); // only when a new error arrives; objects and refs are read as they are then useEffect(() => { const live = new Set(files.map((f) => f.uid)); @@ -246,6 +277,10 @@ export function EditPdfPage() { const canSaveEdits = doc.objects.length > 0 || formDirty || (flattenForm && formFields.length > 0); const start = async () => { + // Clicking Save also commits an open line edit; wait for it and save again with it included. + const openEdit = await finishOpenTextEdit(textGuard.current); + if (openEdit === "blocked") return toast({ title: UI.navBlocked, variant: "error" }); + if (openEdit === "finished") return setSaveAfterTextEdit(true); if (refs.length === 0) return toast({ title: "Add a PDF first", variant: "error" }); if (!hydrateReady) return toast({ title: "Still reading links from the PDF", variant: "error" }); if (formError) return toast({ title: "Cannot fill this form", description: formError, variant: "error" }); @@ -265,6 +300,8 @@ export function EditPdfPage() { if (files.some((f) => f.path === outputPath)) { return toast({ title: "Choose a new file name", description: "The original file is never overwritten.", variant: "error" }); } + const textBlock = textSaveGuard({ objects: doc.objects, refs, sources: textSources.sources }); + if (textBlock) return toast({ title: textBlock.title, description: textBlock.description, variant: "error" }); if (!(await disk.ensure(folder, estimateRequiredBytes("editPdf", files.map((f) => f.sizeBytes))))) return; const failedLinkErr = sessionLinkErrorOnFailedFile(doc.objects, refs, hydrateErrors.current); @@ -290,11 +327,17 @@ export function EditPdfPage() { ), { tool: "editPdf", - label: `Edit PDF · ${doc.objects.length} object${doc.objects.length === 1 ? "" : "s"}`, + label: jobLabel(doc.objects.length - textChangeCount, textChangeCount), }, ); }; + useEffect(() => { + if (!saveAfterTextEdit) return; + setSaveAfterTextEdit(false); + void start(); // this render's start() sees the objects with the just-committed edit + }, [saveAfterTextEdit]); + const canStart = !job.isBusy && inTauri && hydrateReady; return ( @@ -336,6 +379,7 @@ export function EditPdfPage() { /> Flatten annotations{formFields.length > 0 ? " (includes form fields)" : ""} + {textChangeCount > 0 && {UI.outputAlert}} {doc.objects.some((o) => o.kind === "redact") && ( Pages with a redaction become images. Text on those pages will not stay selectable. @@ -368,6 +412,12 @@ export function EditPdfPage() { formFields={current.path === first?.path ? formFields : []} formValues={formValues} onFormChange={(name, value) => setFormValues((prev) => ({ ...prev, [name]: value }))} + text={{ + sources: textSources, + fileName: current.fileName, + duplicatePage: refs.filter((r) => r.path === current.path && r.page === current.page).length > 1, + guardRef: textGuard, + }} /> )} diff --git a/src/features/edit-pdf/textSaveGuards.test.ts b/src/features/edit-pdf/textSaveGuards.test.ts new file mode 100644 index 0000000..04dc382 --- /dev/null +++ b/src/features/edit-pdf/textSaveGuards.test.ts @@ -0,0 +1,129 @@ +import { describe, expect, it, vi } from "vitest"; +import { makeRectObject, makeRedactObject, makeSourceTextObject, type EditObject } from "@/lib/editor"; +import type { PageRef } from "@/lib/types"; +import { finishOpenTextEdit, textSaveGuard } from "./textSaveGuards"; +import type { TextSourceState } from "./useTextSources"; + +const A = "/docs/a.pdf"; +const B = "/docs/b.pdf"; + +function ref(path: string, page: number, uid = path): PageRef { + return { key: `${uid}#${page}`, path, page, fileName: path.split("/").pop() ?? path }; +} + +const REFS: PageRef[] = [ref(A, 1), ref(A, 2), ref(B, 1)]; + +function ready(fingerprint = "fp-a", stale = false): TextSourceState { + return { status: "ready", info: { fingerprint, pageCount: 2, warnings: [] }, error: null, stale }; +} + +const SOURCES: Record = { [A]: ready("fp-a"), [B]: ready("fp-b") }; + +function change(pageIndex: number, sourcePageIndex: number, fingerprint = "fp-a"): EditObject { + return makeSourceTextObject(`t${pageIndex}`, pageIndex, { x: 72, y: 697, w: 40, h: 12 }, { + runId: `t1:${fingerprint}:${sourcePageIndex}:1-2`, + sourceFingerprint: fingerprint, + sourcePageIndex, + originalText: "Hello", + text: "Hallo", + style: {}, + }); +} + +describe("textSaveGuard", () => { + it("passes with no text changes, whatever the sources say", () => { + expect(textSaveGuard({ objects: [makeRectObject("r", 0, { x: 0, y: 0, w: 5, h: 5 })], refs: REFS, sources: {} })).toBeNull(); + }); + + it("passes when every change has a ready, current source", () => { + const objects = [change(1, 1), change(2, 0, "fp-b"), makeRedactObject("x", 0, { x: 0, y: 0, w: 5, h: 5 })]; + expect(textSaveGuard({ objects, refs: REFS, sources: SOURCES })).toBeNull(); + }); + + it("asks to reopen Edit text when the source is not ready", () => { + const expected = { title: "Text changes can't be saved yet", description: "Open Edit text on “a.pdf” again, then save." }; + const loading: TextSourceState = { status: "loading", info: null, error: null, stale: false }; + const failed: TextSourceState = { status: "error", info: null, error: { code: "X", title: "t", message: "m" }, stale: false }; + for (const state of [undefined, loading, failed]) { + const sources: Record = state ? { ...SOURCES, [A]: state } : { [B]: SOURCES[B] }; + expect(textSaveGuard({ objects: [change(0, 0)], refs: REFS, sources })).toEqual(expected); + } + }); + + it("refuses changes made on an older snapshot of the file", () => { + const expected = { + title: "The PDF changed on disk", + description: "“a.pdf” was changed after you started editing it, so your text changes no longer match it.", + }; + expect(textSaveGuard({ objects: [change(0, 0)], refs: REFS, sources: { ...SOURCES, [A]: ready("fp-a", true) } })).toEqual( + expected, + ); + expect(textSaveGuard({ objects: [change(0, 0, "fp-old")], refs: REFS, sources: SOURCES })).toEqual(expected); + expect(textSaveGuard({ objects: [change(0, 1)], refs: REFS, sources: SOURCES })).toEqual(expected); + }); + + it("refuses a redaction and a text change on the same page", () => { + const objects = [change(1, 1), makeRedactObject("x", 1, { x: 0, y: 0, w: 5, h: 5 })]; + expect(textSaveGuard({ objects, refs: REFS, sources: SOURCES })).toEqual({ + title: "Redaction and text change on the same page", + description: + "Page 2 has both a redaction and a text change. Redaction turns the page into an image, so the text change would be lost. Remove the redaction or the text change on that page.", + }); + }); + + it("refuses a change on a page listed twice", () => { + const refs = [ref(A, 1), ref(A, 2), ref(A, 1, "copy")]; + expect(textSaveGuard({ objects: [change(0, 0)], refs, sources: SOURCES })).toEqual({ + title: "This page appears twice", + description: + "Page 1 of “a.pdf” is in the list more than once and has a text change. Remove the extra copy of the page, then save again.", + }); + expect(textSaveGuard({ objects: [change(1, 1)], refs, sources: SOURCES })).toBeNull(); + }); + + it("reports the first failing check in order: not ready, stale, redaction, duplicate", () => { + const refs = [ref(A, 1), ref(A, 1, "copy"), ref(B, 1)]; + const redact = makeRedactObject("x", 0, { x: 0, y: 0, w: 5, h: 5 }); + const objects = [change(0, 0), redact, change(2, 0, "fp-b")]; + const notReadyB = { ...SOURCES, [B]: { status: "loading", info: null, error: null, stale: false } as TextSourceState }; + expect(textSaveGuard({ objects, refs, sources: notReadyB })?.title).toBe("Text changes can't be saved yet"); + expect(textSaveGuard({ objects, refs, sources: { ...SOURCES, [B]: ready("fp-b", true) } })?.title).toBe( + "The PDF changed on disk", + ); + expect(textSaveGuard({ objects, refs, sources: SOURCES })?.title).toBe("Redaction and text change on the same page"); + expect(textSaveGuard({ objects: [change(0, 0)], refs, sources: SOURCES })?.title).toBe("This page appears twice"); + }); +}); + +describe("finishOpenTextEdit (Save with a line edit still open)", () => { + const guard = (editing: boolean, result: boolean | Promise) => { + const tryClose = vi.fn(() => Promise.resolve(result)); + return { edit: { isEditing: () => editing, tryClose }, tryClose }; + }; + + it("does nothing when no line is open", async () => { + const { edit, tryClose } = guard(false, true); + expect(await finishOpenTextEdit(edit)).toBe("none"); + expect(await finishOpenTextEdit(null)).toBe("none"); + expect(tryClose).not.toHaveBeenCalled(); + }); + + it("waits for the open edit's check and reports it finished, so Save reads the objects again", async () => { + let resolve: (ok: boolean) => void = () => {}; + const pending = new Promise((r) => (resolve = r)); + const { edit, tryClose } = guard(true, pending); + const out = finishOpenTextEdit(edit); + let settled = false; + void out.then(() => (settled = true)); + await Promise.resolve(); + expect(settled).toBe(false); + resolve(true); + expect(await out).toBe("finished"); + expect(tryClose).toHaveBeenCalledTimes(1); + }); + + it("reports a draft that can't be applied as blocked", async () => { + const { edit } = guard(true, false); + expect(await finishOpenTextEdit(edit)).toBe("blocked"); + }); +}); diff --git a/src/features/edit-pdf/textSaveGuards.ts b/src/features/edit-pdf/textSaveGuards.ts new file mode 100644 index 0000000..711125f --- /dev/null +++ b/src/features/edit-pdf/textSaveGuards.ts @@ -0,0 +1,111 @@ +/** + * Save guards for text changes (SPEC §D.8), checked before `edit_pdf_overlays` + * runs. Rust repeats every check; these only turn the predictable failures + * into an immediate toast. Order: source not ready, stale, redaction on the + * same page, edited page listed twice. + */ + +import type { EditObject, SourceTextObject } from "@/lib/editor"; +import { FILE_ERROR_COPY, SAVE_GUARD_COPY, fillCopy } from "@/lib/editor/sourceTextCopy"; +import type { PageRef } from "@/lib/types"; +import type { TextSourceState } from "./useTextSources"; + +export interface TextSaveGuardInput { + objects: EditObject[]; + refs: PageRef[]; + sources: Record; +} + +export interface TextSaveGuardResult { + title: string; + description: string; +} + +const UNKNOWN_FILE = "this PDF"; + +function sourceTextObjects(objects: EditObject[]): SourceTextObject[] { + return objects.filter((o): o is SourceTextObject => o.kind === "sourceText"); +} + +function rowWithSuggestion(code: string, vars: Record): TextSaveGuardResult { + const row = FILE_ERROR_COPY[code]; + return { + title: row.title, + description: `${fillCopy(row.message, vars)} ${row.suggestion}`.trim(), + }; +} + +function notReady(edits: SourceTextObject[], input: TextSaveGuardInput): TextSaveGuardResult | null { + for (const o of edits) { + const ref = input.refs[o.pageIndex]; + const source = ref ? input.sources[ref.path] : undefined; + if (!ref || !source || source.status !== "ready" || !source.info) { + return { + title: SAVE_GUARD_COPY.notReadyTitle, + description: fillCopy(SAVE_GUARD_COPY.notReadyDescription, { name: ref?.fileName ?? UNKNOWN_FILE }), + }; + } + } + return null; +} + +function stale(edits: SourceTextObject[], input: TextSaveGuardInput): TextSaveGuardResult | null { + for (const o of edits) { + const ref = input.refs[o.pageIndex]; + const source = input.sources[ref.path]; + const moved = ref.page - 1 !== o.sourcePageIndex; + if (source.stale || source.info?.fingerprint !== o.sourceFingerprint || moved) { + const row = FILE_ERROR_COPY.STALE; + return { title: row.title, description: fillCopy(row.message, { name: ref.fileName }) }; + } + } + return null; +} + +function redactionConflict(edits: SourceTextObject[], input: TextSaveGuardInput): TextSaveGuardResult | null { + const redacted = new Set(input.objects.filter((o) => o.kind === "redact").map((o) => o.pageIndex)); + const hit = edits.find((o) => redacted.has(o.pageIndex)); + return hit ? rowWithSuggestion("TEXT_EDIT_ON_REDACTED_PAGE", { n: hit.pageIndex + 1 }) : null; +} + +function duplicatePage(edits: SourceTextObject[], input: TextSaveGuardInput): TextSaveGuardResult | null { + const listed = new Map(); + for (const ref of input.refs) { + const key = `${ref.path}\u0000${ref.page}`; + listed.set(key, (listed.get(key) ?? 0) + 1); + } + for (const o of edits) { + const ref = input.refs[o.pageIndex]; + if ((listed.get(`${ref.path}\u0000${ref.page}`) ?? 0) > 1) { + return rowWithSuggestion("TEXT_EDIT_DUPLICATE_PAGE", { p: ref.page, name: ref.fileName }); + } + } + return null; +} + +/** The first reason the text changes can't be saved yet, or null when all clear. */ +export function textSaveGuard(input: TextSaveGuardInput): TextSaveGuardResult | null { + const edits = sourceTextObjects(input.objects); + if (edits.length === 0) return null; + return ( + notReady(edits, input) ?? stale(edits, input) ?? redactionConflict(edits, input) ?? duplicatePage(edits, input) + ); +} + +/** The open line edit of the Edit text layer (SourceTextMode's guard), seen from Save. */ +export interface OpenTextEdit { + isEditing(): boolean; + tryClose(): Promise; +} + +/** + * Save must never snapshot the objects while a line edit is still open or being checked: the + * click on Save also commits that edit (outside click, B11), so the saved file would miss a + * change the editor then shows as applied. "none": nothing open; "finished": the edit was + * committed (or dropped as a no-op), so read the objects again before saving; "blocked": the + * draft can't be applied and stays open with its message. + */ +export async function finishOpenTextEdit(edit: OpenTextEdit | null): Promise<"none" | "finished" | "blocked"> { + if (!edit?.isEditing()) return "none"; + return (await edit.tryClose()) ? "finished" : "blocked"; +} diff --git a/src/features/edit-pdf/useTextSources.test.ts b/src/features/edit-pdf/useTextSources.test.ts new file mode 100644 index 0000000..8d22cd4 --- /dev/null +++ b/src/features/edit-pdf/useTextSources.test.ts @@ -0,0 +1,198 @@ +// @vitest-environment happy-dom +import { act, createElement } from "react"; +import { createRoot, type Root } from "react-dom/client"; +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import type { AppError, TextSourceInfo } from "@/lib/types"; + +const backend = vi.hoisted(() => ({ + openTextSource: vi.fn<(path: string) => Promise>(), + releaseTextSource: vi.fn<(path: string) => Promise>(), +})); + +vi.mock("@/lib/tauriCommands", () => backend); + +import { useTextSources, type TextSources, type UseTextSourcesOptions } from "./useTextSources"; + +(globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +const INFO: TextSourceInfo = { fingerprint: "fp-1", pageCount: 3, warnings: [] }; +const ENCRYPTED: AppError = { + code: "ENCRYPTED", + title: "This PDF is password-protected", + message: "OffPDF can't change text in a protected PDF.", + suggestion: "Remove the password with Unlock PDF, then edit the unlocked copy.", +}; + +interface Deferred { + promise: Promise; + resolve: (value: T) => void; + reject: (reason: unknown) => void; +} + +function deferred(): Deferred { + let resolve!: (value: T) => void; + let reject!: (reason: unknown) => void; + const promise = new Promise((res, rej) => { + resolve = res; + reject = rej; + }); + return { promise, resolve, reject }; +} + +let root: Root | null = null; +let container: HTMLElement | null = null; +let hook: TextSources; + +function mount(options?: UseTextSourcesOptions): void { + function Probe() { + hook = useTextSources(options); + return null; + } + container = document.createElement("div"); + document.body.appendChild(container); + root = createRoot(container); + act(() => root!.render(createElement(Probe))); +} + +function unmount(): void { + act(() => root?.unmount()); + root = null; + container?.remove(); + container = null; +} + +async function settle(work: () => Promise): Promise { + let out!: T; + await act(async () => { + out = await work(); + }); + return out; +} + +beforeEach(() => { + backend.openTextSource.mockReset(); + backend.releaseTextSource.mockReset(); + backend.releaseTextSource.mockResolvedValue(undefined); +}); + +afterEach(() => { + if (root) unmount(); +}); + +describe("useTextSources", () => { + it("opens a file once and serves the cached info afterwards", async () => { + const open = deferred(); + backend.openTextSource.mockReturnValueOnce(open.promise); + mount(); + let first!: Promise; + let second!: Promise; + act(() => { + first = hook.ensure("/a.pdf"); + second = hook.ensure("/a.pdf"); + }); + expect(hook.sources["/a.pdf"]).toEqual({ status: "loading", info: null, error: null, stale: false }); + open.resolve(INFO); + expect(await settle(() => first)).toEqual(INFO); + expect(await settle(() => second)).toEqual(INFO); + expect(hook.sources["/a.pdf"]).toEqual({ status: "ready", info: INFO, error: null, stale: false }); + expect(await settle(() => hook.ensure("/a.pdf"))).toEqual(INFO); + expect(backend.openTextSource).toHaveBeenCalledTimes(1); + expect(backend.openTextSource).toHaveBeenCalledWith("/a.pdf"); + }); + + it("keeps the backend error in the state instead of swallowing it", async () => { + backend.openTextSource.mockRejectedValueOnce(ENCRYPTED); + mount(); + expect(await settle(() => hook.ensure("/locked.pdf"))).toBeNull(); + expect(hook.sources["/locked.pdf"]).toEqual({ status: "error", info: null, error: ENCRYPTED, stale: false }); + // An error is not retried behind the user's back; reopen does that. + expect(await settle(() => hook.ensure("/locked.pdf"))).toBeNull(); + expect(backend.openTextSource).toHaveBeenCalledTimes(1); + }); + + it("normalises thrown values and malformed answers to AppErrors", async () => { + backend.openTextSource.mockRejectedValueOnce(new Error("boom")); + backend.openTextSource.mockResolvedValueOnce({ fingerprint: "", pageCount: 1, warnings: [] }); + mount(); + await settle(() => hook.ensure("/x.pdf")); + expect(hook.sources["/x.pdf"].error).toMatchObject({ code: "UNKNOWN", message: "boom" }); + await settle(() => hook.ensure("/y.pdf")); + expect(hook.sources["/y.pdf"]).toMatchObject({ status: "error", error: { code: "INVALID_RESPONSE" } }); + }); + + it("release forgets the file and tells the backend", async () => { + backend.openTextSource.mockResolvedValueOnce(INFO); + mount(); + await settle(() => hook.ensure("/a.pdf")); + act(() => hook.release("/a.pdf")); + expect(hook.sources["/a.pdf"]).toBeUndefined(); + expect(backend.releaseTextSource).toHaveBeenCalledWith("/a.pdf"); + act(() => hook.release("/never-opened.pdf")); + expect(backend.releaseTextSource).toHaveBeenCalledTimes(1); + }); + + it("drops a late answer for a released file and releases it again", async () => { + const open = deferred(); + backend.openTextSource.mockReturnValueOnce(open.promise); + mount(); + let pending!: Promise; + act(() => { + pending = hook.ensure("/a.pdf"); + }); + act(() => hook.release("/a.pdf")); + open.resolve(INFO); + expect(await settle(() => pending)).toBeNull(); + expect(hook.sources["/a.pdf"]).toBeUndefined(); + expect(backend.releaseTextSource).toHaveBeenCalledTimes(2); + }); + + it("reports a failed release through onReleaseError", async () => { + const onReleaseError = vi.fn(); + backend.openTextSource.mockResolvedValueOnce(INFO); + backend.releaseTextSource.mockRejectedValueOnce({ ...ENCRYPTED, code: "IO_ERROR" }); + mount({ onReleaseError }); + await settle(() => hook.ensure("/a.pdf")); + await settle(async () => hook.release("/a.pdf")); + expect(onReleaseError).toHaveBeenCalledWith("/a.pdf", expect.objectContaining({ code: "IO_ERROR" })); + }); + + it("marks a source stale or failed, and reopen reads it again", async () => { + backend.openTextSource.mockResolvedValueOnce(INFO); + backend.openTextSource.mockResolvedValueOnce({ ...INFO, fingerprint: "fp-2" }); + mount(); + await settle(() => hook.ensure("/a.pdf")); + act(() => hook.markStale("/a.pdf")); + expect(hook.sources["/a.pdf"]).toMatchObject({ status: "ready", stale: true, info: INFO }); + const repair: AppError = { code: "PDF_NEEDS_REPAIR", title: "t", message: "m" }; + act(() => hook.markError("/a.pdf", repair)); + expect(hook.sources["/a.pdf"]).toMatchObject({ status: "error", error: repair, stale: true }); + const info = await settle(() => hook.reopen("/a.pdf")); + expect(info?.fingerprint).toBe("fp-2"); + expect(hook.sources["/a.pdf"]).toEqual({ status: "ready", info: { ...INFO, fingerprint: "fp-2" }, error: null, stale: false }); + }); + + it("ignores an answer that a reopen superseded", async () => { + const slow = deferred(); + backend.openTextSource.mockReturnValueOnce(slow.promise); + backend.openTextSource.mockResolvedValueOnce({ ...INFO, fingerprint: "fp-new" }); + mount(); + let stale!: Promise; + act(() => { + stale = hook.ensure("/a.pdf"); + }); + await settle(() => hook.reopen("/a.pdf")); + slow.resolve({ ...INFO, fingerprint: "fp-old" }); + expect(await settle(() => stale)).toBeNull(); + expect(hook.sources["/a.pdf"].info?.fingerprint).toBe("fp-new"); + expect(backend.releaseTextSource).not.toHaveBeenCalled(); + }); + + it("releases every opened file on unmount", async () => { + backend.openTextSource.mockResolvedValue(INFO); + mount(); + await settle(() => hook.ensure("/a.pdf")); + await settle(() => hook.ensure("/b.pdf")); + unmount(); + expect(backend.releaseTextSource.mock.calls.map((c) => c[0]).sort()).toEqual(["/a.pdf", "/b.pdf"]); + }); +}); diff --git a/src/features/edit-pdf/useTextSources.ts b/src/features/edit-pdf/useTextSources.ts new file mode 100644 index 0000000..5a1d7ff --- /dev/null +++ b/src/features/edit-pdf/useTextSources.ts @@ -0,0 +1,181 @@ +/** + * Per-file Edit text sources (SPEC §D.4), owned by `EditPdfPage` like form fields. + * + * `ensure` opens a file once (`open_text_source`, deduplicated while in flight), + * `release` forgets it and lets the backend delete its temporary copies, + * `markStale` / `markError` record what a later inspect or preview reported, and + * `reopen` reads the file again. Every failure is kept in the state (or passed + * to `onReleaseError`); nothing is swallowed. Late answers for a path that was + * released or reopened meanwhile are dropped by a per-path generation counter. + */ + +import { useCallback, useEffect, useRef, useState } from "react"; +import { openTextSource, releaseTextSource } from "@/lib/tauriCommands"; +import { toAppError, type AppError, type TextSourceInfo } from "@/lib/types"; + +export interface TextSourceState { + status: "idle" | "loading" | "ready" | "error"; + info: TextSourceInfo | null; + error: AppError | null; + stale: boolean; +} + +export interface UseTextSourcesOptions { + /** A `release_text_source` failure (the file is already gone from the state). */ + onReleaseError?: (path: string, error: AppError) => void; +} + +export interface TextSources { + sources: Record; + ensure(path: string): Promise; + release(path: string): void; + markStale(path: string): void; + markError(path: string, error: AppError): void; + reopen(path: string): Promise; +} + +const INVALID_ANSWER: AppError = { + code: "INVALID_RESPONSE", + title: "Edit text could not start", + message: "OffPDF received an unexpected answer while reading this PDF for Edit text.", + suggestion: "Close the file and add it again.", + details: null, +}; + +function isSourceInfo(value: unknown): value is TextSourceInfo { + if (typeof value !== "object" || value === null) return false; + const v = value as Record; + return ( + typeof v.fingerprint === "string" && + v.fingerprint.length > 0 && + typeof v.pageCount === "number" && + Number.isInteger(v.pageCount) && + v.pageCount >= 0 && + Array.isArray(v.warnings) && + v.warnings.every((w) => typeof w === "string") + ); +} + +export function useTextSources(options: UseTextSourcesOptions = {}): TextSources { + const [sources, setSources] = useState>({}); + const sourcesRef = useRef>({}); + const inflight = useRef(new Map>()); + const generations = useRef(new Map()); + const mounted = useRef(true); + const onReleaseErrorRef = useRef(options.onReleaseError); + onReleaseErrorRef.current = options.onReleaseError; + + const write = useCallback((path: string, next: TextSourceState | null) => { + const { [path]: _previous, ...rest } = sourcesRef.current; + sourcesRef.current = next ? { ...rest, [path]: next } : rest; + if (mounted.current) setSources(sourcesRef.current); + }, []); + + const bump = useCallback((path: string): number => { + const generation = (generations.current.get(path) ?? 0) + 1; + generations.current.set(path, generation); + return generation; + }, []); + + const releaseOnBackend = useCallback((path: string) => { + releaseTextSource(path).catch((e: unknown) => onReleaseErrorRef.current?.(path, toAppError(e))); + }, []); + + const open = useCallback( + (path: string): Promise => { + const generation = bump(path); + write(path, { status: "loading", info: null, error: null, stale: false }); + const isCurrent = () => generations.current.get(path) === generation && mounted.current; + const request = openTextSource(path).then( + (info) => { + if (!isCurrent()) { + // Released while reading: the backend cached a snapshot nobody will release. + if (!(path in sourcesRef.current) || !mounted.current) releaseOnBackend(path); + return null; + } + if (!isSourceInfo(info)) { + write(path, { status: "error", info: null, error: INVALID_ANSWER, stale: false }); + return null; + } + write(path, { status: "ready", info, error: null, stale: false }); + return info; + }, + (e: unknown) => { + if (isCurrent()) write(path, { status: "error", info: null, error: toAppError(e), stale: false }); + return null; + }, + ); + const tracked = request.finally(() => { + if (inflight.current.get(path) === tracked) inflight.current.delete(path); + }); + inflight.current.set(path, tracked); + return tracked; + }, + [bump, releaseOnBackend, write], + ); + + const ensure = useCallback( + (path: string): Promise => { + const current = sourcesRef.current[path]; + if (current?.status === "ready" && current.info) return Promise.resolve(current.info); + if (current?.status === "error") return Promise.resolve(null); + const pending = inflight.current.get(path); + if (pending) return pending; + return open(path); + }, + [open], + ); + + const reopen = useCallback( + (path: string): Promise => { + inflight.current.delete(path); + return open(path); + }, + [open], + ); + + const release = useCallback( + (path: string) => { + const known = path in sourcesRef.current || inflight.current.has(path); + bump(path); + inflight.current.delete(path); + write(path, null); + if (known) releaseOnBackend(path); + }, + [bump, releaseOnBackend, write], + ); + + const markStale = useCallback( + (path: string) => { + const current = sourcesRef.current[path] ?? { status: "idle", info: null, error: null, stale: false }; + if (!current.stale) write(path, { ...current, stale: true }); + }, + [write], + ); + + const markError = useCallback( + (path: string, error: AppError) => { + bump(path); + inflight.current.delete(path); + const current = sourcesRef.current[path]; + write(path, { status: "error", info: current?.info ?? null, error, stale: current?.stale ?? false }); + }, + [bump, write], + ); + + useEffect(() => { + mounted.current = true; + return () => { + mounted.current = false; + const paths = new Set([...Object.keys(sourcesRef.current), ...inflight.current.keys()]); + for (const path of paths) { + bump(path); + releaseOnBackend(path); + } + inflight.current.clear(); + sourcesRef.current = {}; + }; + }, [bump, releaseOnBackend]); + + return { sources, ensure, release, markStale, markError, reopen }; +} diff --git a/src/lib/editor/EDIT_MODEL.md b/src/lib/editor/EDIT_MODEL.md index 01b478d..39604ef 100644 --- a/src/lib/editor/EDIT_MODEL.md +++ b/src/lib/editor/EDIT_MODEL.md @@ -11,10 +11,14 @@ invent a second coordinate system. - An **undo/redo** reducer for editor sessions. - Kinds: `text`, `image` (filesystem path), `line`, `ink`, and closed vector shapes such as rectangles, ellipses, arrows, and polygons. +- `sourceText`: a change to a line of existing text (Edit text). It is not an + overlay; see [Source text edits](#source-text-edits-kind-sourcetext). ## What this module is not -- It does **not** edit existing page content operators. +- Apart from `sourceText`, objects never touch existing page content + operators. `sourceText` changes them only through the Rust text-edit engine, + never in the frontend. - **Links** (`kind: "link"`) are PDF `/Annots`, not overlay stamps. They use the same unrotated `EditObject.rect` space. Overlay paint skips them; Save rewrites dest `/Link` dictionaries after `qpdf --overlay`. @@ -110,6 +114,184 @@ existing annot through and appends or removes only session `/NM` dicts. Draw `kind: "ink"` stays a content-stream stroke. Flatten is opt-in `qpdf --flatten-annotations=all` (default off). +## Source text edits (`kind: "sourceText"`) + +A `sourceText` object records a change to one line of text that is already in a +source PDF. It is **not** an overlay: Save rewrites that line's show operators +inside the page's own content stream. The user guide, the reason codes and the +full verification design are in [`docs/EDIT_TEXT.md`](../../../docs/EDIT_TEXT.md). + +```ts +{ + id, kind: "sourceText", pageIndex, rect, // rect = the line's box after the change (verdict newRect) + locked: true, // never moved, resized, rotated, copied or pasted + runId, // the engine's line id on that source page + sourceFingerprint, // content hash of the source file when the line was read + sourcePageIndex, // 0-based page in the source file + originalText, text, + style: { sizePt?, face?, fill?, letterSpacingPt? } // only fields that differ from the original +} +``` + +Rules: + +- **One object per `(pageIndex, runId)`.** Committing a change upserts it; + committing the original text with no style change deletes it (a no-op writes + nothing). Commit, restore and inspector style changes are one undo step each. +- **Locked.** `applyUpdate` accepts only `text` and `style`; it strips `rect`, + rotation, opacity, `locked`, `runId`, fingerprint and page fields. The objects + draw nothing in the SVG layer, are excluded from marquee, copy, paste, + duplicate and nudge, and are never added by `ADD_MANY`. +- **Fingerprint and source page.** Every object carries the source fingerprint + and `sourcePageIndex`. Reordering pages remaps `pageIndex` and keeps both; + removing a page drops its changes. If the file changes on disk, the fingerprint + no longer matches: the editor shows the stale banner, and "Remove these text + changes" deletes all of that file's changes in one step. +- **Preview bytes.** The canvas shows the page that Save would write: on commit, + `preview_text_edits` plans the page's changes, writes a one-page copy through + qpdf and runs the same Phase A checks as Save. The returned one-page PDF is + rendered in place of the original page (double-buffered, cached per edit + signature). **Show original** (O) swaps back; it is offered only while a + preview is on screen (`canShowOriginal`: the page has changes and preview + bytes). The bytes never leave the page surface and are never stored in the + document. +- **Save.** `toExportDocument` sends the objects to `edit_pdf_overlays` + unchanged; Rust ignores `id` and `locked`. Save re-reads each edited source + (never the editor's cache), checks the fingerprint, plans every change again + and writes a per-source edited copy **before** assembly. Phase A proves each + copy; the existing pipeline (assembly, stamps, links, forms, markup, redaction + on other pages) runs on the edited copies; Phase B proves the final file; then + the existing #34 validation runs and the file is renamed into place. Clicking + Save while a line edit is open first finishes it (`finishOpenTextEdit` in + `textSaveGuards.ts`): Save waits for the commit and runs again on the + committed objects; a draft that can't be applied stays open, the editor shows + "Finish or cancel the text change on this page first." and nothing is saved. + For links, a save whose objects are all `sourceText` counts as empty, so links + deleted in the editor stay deleted (the L7 rule in `edit_overlay.rs`). +- **Redaction conflict.** A page cannot have both a redaction and a text change + (redaction turns the page into an image). The editor says so in the redaction + note; Save refuses with `TEXT_EDIT_ON_REDACTED_PAGE`. +- **A page listed twice** (same file and source page) cannot take text changes; + the editor shows a banner and Save refuses with `TEXT_EDIT_DUPLICATE_PAGE`. +- **Refused lines** come from inspect with a reason and never become objects. + Text drawn through a Form XObject arrives as refused runs (`NESTED_FORM`, no + metrics or style, ids that never match a page run) after the page's own runs. +- **Add text here** (reason dialog) adds an ordinary `text` object whose box is + `textStampGeometry(run, geometry)`, worked out as displayed: a level line's + own box grown to one line and four characters; for a line turned, tilted or + vertical on screen, an upright 10 em × 1.3 em box at the line's start. The box + is kept on the page. + +### Engine architecture and invariants + +The Rust side lives in `src-tauri/src/pdf_engine/text_edit/` (module map in +`docs/EDIT_TEXT.md` §16). The invariants every change must keep: + +1. **Real source edits only.** A change is a byte splice of show operators in the + page's own content stream, written by `qpdf --update-from-json`. Never a + painted-over fake, a flattened or rasterised page, a lopdf re-save, or a write + to the input. +2. **One bounded snapshot per operation.** The file is read once with a cap, + hashed, and parsed from the same bytes; qpdf only reads copies of those bytes. +3. **Fail closed.** Every refusal has a code from `text-reasons.json`; nothing is + skipped or returned partially as success. +4. **Bounded.** Every decompression is capped while inflating, every parser has + an operation and nesting budget, and each page model has a byte budget + (`PAGE_MODEL_BYTES_MAX`, 160 MiB) that covers the font models it keeps and its + #33 Classify pass; a model keeps no lexed operators. Exceeding a page budget + refuses that page (`PAGE_TOO_COMPLEX`), never the app. The cache sizes models + by `walk.model_bytes`, and Save rebuilds every page's model in Phase A instead + of keeping them once they pass `max(2 × file, 16 MiB)`. +5. **Verified before publish.** Plan self-check → qpdf write → Phase A (A0 qpdf + check, A1 page-map agreement, A2 whole-graph equality, A3 exact edited bytes, + A4 re-walk at 0.01 pt, A5 Poppler words and pixels) → existing passes → + Phase B (B0 read, B1 expected content present, B2 original absent, B3 re-walk) + → #34 → atomic rename. Content-part boundaries must mean the same to every + reader: one rule (`content/joins.rs`) makes the walker refuse a page whose + parts meet inside a token, comment or `BX`…`EX` section (`MALFORMED_CONTENT`), + and #34's joined-form digest and Phase B's wrapper form need the same rule. +6. **No panics in production code**, no forbidden lopdf APIs (GUARD-01), test + seams only under `#[cfg(test)]`. +7. **One copy source.** Every user-facing sentence is in `sourceTextCopy.ts` + (Rust builds the same `AppError` text in `reasons.rs`); every Save failure + ends with "The original file was not changed." + +### Edit text contributor workflow + +#### Adding a reason code + +Codes are append-only. In one change: + +1. `text_edit/reasons.rs`: add the variant, its `as_str` spelling, its priority + slot and its title and body. +2. `src/lib/editor/text-reasons.json`: append the code to its list + (`reasons_json_matches_enums` cross-locks the order with Rust). +3. `src/lib/types.ts` (union) and `src/lib/editor/sourceTextCopy.ts` + (`REASON_COPY`, `PAGE_REASON_COPY` or the error rows): `sourceTextCopy.test.ts` + fails until every code has copy and passes the wording scan. +4. `docs/EDIT_TEXT.md`: add the row inside the ui-copy block of section 6 or 7. + `docsReasons.test.ts` fails until the doc names the code and quotes the copy + exactly. +5. CLS-H22 derives the classifier's frozen list from the JSON; update its legacy + assertion only if a legacy #33 name changes (it must not). + +#### Engines and the skip policy + +Tests that need qpdf ≥ 11, `pdftotext` and `pdftoppm` call +`testkit::engines_or_skip`, which prints `skip: … not available` and returns +when a tool is missing. With `OFFPDF_REQUIRE_ENGINES=1` a missing tool fails the +test instead. CI installs qpdf and Poppler, sets the variable and fails if any +test prints a `skip:` line. Locally (PR #98 convention): + +```bash +cd src-tauri +cargo test --lib -j 6 -- --nocapture 2>&1 | grep -c 'skip:' # must print 0 +OFFPDF_REQUIRE_ENGINES=1 cargo test --lib -j 6 text_edit +``` + +#### Benchmarks and generated fixtures + +```bash +cd src-tauri +cargo test --release --lib -j 6 text_edit::bench -- --ignored --nocapture --test-threads=1 # BENCH-01/02 +OFFPDF_UPDATE_GOLDEN=1 cargo test --lib meas_01_estimate_golden # rewrites text-measure-golden.json (MEAS-01) +OFFPDF_UPDATE_GOLDEN=1 cargo test --lib dto_01_serialised # rewrites text-edit-dto-contract.json +``` + +The two JSON files under `src/lib/editor/__fixtures__/` are read by +`sourceTextGolden.test.ts` and `sourceTextContract.test.ts`, which keep the +frontend's width estimate and DTO types in step with Rust (the contract test +derives each DTO's schema from `types.ts`, so a drift also fails `npm run +typecheck`). BENCH-SIZE is measured as described in `docs/EDIT_TEXT.md` §12. + +#### Adding a font or producer fixture + +Fixtures are generated in memory by the Rust test kit; no PDFs are committed +for Edit text (`fixtures/source-edit/` does not grow). + +- **Font programs:** `text_edit/testkit/{ttf,cff,type1}.rs` build TrueType, CFF + and Type1 programs with chosen glyphs (`TtfBuilder::glyph(name, outline, + advance)`, …). `testkit/fonts.rs` wraps them in PDF font dictionaries + (`SimpleFont`, `Type0Font`, `add_simple`, `add_type0`, ToUnicode helpers) and + loads them (`load_simple`, `load_type0`). Real programs already in the repo may + be used read-only: `src-tauri/resources/fonts/NotoSans-Regular.ttf`, + `public/pdfjs/standard_fonts/LiberationSans-*.ttf`, and the + `public/pdfjs/standard_fonts/Foxit*.pfb` files (these are bare CFF programs, + not Type1). +- **A new font class:** add the class to `fonts/` with a glyph-presence proof + from the program, fixtures and FONT tests in `fonts/tests*.rs`, then add a case + to IND-01 (`tests_independent.rs`) so an honest edit in that font passes Phase A + including the Poppler checks, and a line to E2E-05 (`tests_e2e/fonts.rs`) so it + saves through the real export. +- **A producer shape:** add a builder to `testkit/producers/` (office-style files + in `office.rs`, geometry and syntax edges in `edges.rs`, file-level cases in + `files.rs`, tagging in `tagging.rs`), add its smoke test in + `tests_walk/producers.rs`, assert its editable and refused lines in a RUN or + CLS-H test, and run it through E2E so the Save path covers it. Add the row to + the compatibility matrix in `docs/EDIT_TEXT.md` §11 with the test ids. +- Run `OFFPDF_REQUIRE_ENGINES=1 cargo test --lib -j 6 text_edit` and check that + `cargo test --lib -j 6 -- --nocapture 2>&1 | grep -c 'skip:'` prints 0. + ## How export consumes this 1. Read `EditDocument.objects` for the chosen pages. @@ -136,7 +318,43 @@ The canvas exposes the active page's object list so selection and deletion work without pointer-only interaction. Selection-driven actions never affect hidden pages. Keyboard: Delete/Backspace, Escape, undo/redo chords, arrow nudge. Hand tool (H) and hold-Space pan the zoomed page; trackpad scroll on the stage still -works. +works. The editor opens on the **Select** tool. + +### Edit text keyboard contract + +Single-key shortcuts work when focus is not in a text field, the format bar or +the reason dialog. + +| Where | Key | Action | +| --- | --- | --- | +| Editor | **E** | Edit text tool (`aria-keyshortcuts="E"`; **H** is the Hand tool) | +| Editor, page with text changes and a preview on screen | **O** | Show original (toggle, `aria-pressed`) | +| Text layer (one Tab stop, roving focus over line buttons) | ↓ / ↑ | next / previous line in reading order | +| | → / ← | next / previous text block in reading order (moves within a visual line, then on to the next or previous one) | +| | Home / End | first / last text block on the page | +| | Page Up / Page Down | 10 lines back / forward | +| | Enter, F2 or Space | edit the line with all text selected, or open the reason dialog for a refused line | +| | Delete / Backspace | restore the original of a changed line (one undo step) | +| | Esc | leave the layer | +| | ⌘/Ctrl+Z, ⌘/Ctrl+Shift+Z, ⌘/Ctrl+Y | undo / redo (existing handler) | +| Line editor | Enter | Done (commit after the check) | +| | Esc | cancel the draft | +| | Tab | into the format bar (first stop: Smaller) | +| | ⌘/Ctrl+B, ⌘/Ctrl+I | bold / italic, or write why it is unavailable | +| | ⌘/Ctrl+Shift+. and ⌘/Ctrl+Shift+, | size +0.5 / −0.5 pt | +| Format bar (`role="toolbar"`) | ← / → | move between controls | +| | Enter / Space | activate | +| | Esc, Shift+Tab | back to the line editor | +| Reason dialog | Esc | close; focus returns to the line | + +Space-to-pan is suppressed while focus is in the text layer, the line editor, the +format bar or the reason dialog, so Space activates the focused control. Enter +and Esc are ignored while an input method is composing. Lines are real +`