From b46dd5f84ca4df645d3c44b1f0250fafb4763fd8 Mon Sep 17 00:00:00 2001 From: MrChengLen Date: Wed, 30 Sep 2026 15:23:24 +0200 Subject: [PATCH 1/7] =?UTF-8?q?fix(pdf):=20a=20PDF=20pypdf=20can't=20read?= =?UTF-8?q?=20is=20a=20400,=20not=20a=20500=20=E2=80=94=20LimitReachedErro?= =?UTF-8?q?r=20too?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit pypdf raises LimitReachedError when a crafted file trips one of its resource limits. It derives from PyPdfError, not PdfReadError, so _open_reader's catch missed it; and pypdf parses lazily, so most limits (and a damaged page object's ValueError) fire in add_page/write or in extract_text, after the one guarded open. Result: generic 500 plus a logged traceback on /pdf/extract, /pdf/split and /convert pdf->pdf, and for every unreadable or password-protected PDF on pdf->txt, which caught nothing. No detail reached the client; status and log were wrong. - pdf_pages.py: one _reading_pdf() guard around every pypdf step that reads the input (open, add_page/write in extract, the split loop) maps PyPdfError and ValueError to UnreadablePdfError, a PageSelectionError with a fixed message, and logs the error's class (never its message, which can quote the file). OSError is no longer mapped: pypdf reads the whole file into memory first, so it can only be our own disk, a 500. PageSelectionError is now an InvalidInputError (and no longer a ValueError, which nothing relied on), so /convert and /convert/batch answer it through their existing 400 invalid_input / per-file-message branches. - document.py: PdfToTxtConverter maps the same errors from open and text extraction to InvalidInputError, and writes an unpaired surrogate (from a broken ToUnicode map) as "?" instead of failing the UTF-8 write with a 500. - /pdf/extract sends invalid_pdf for an unreadable file. It sent invalid_page_selection, which made pdf-tools.js tell the user to fix a page selection that was fine; invalid_pdf already maps to the localized "Could not read the PDF" text. Docs table updated. Not widened: pypdf also raises AttributeError, KeyError, TypeError, NotImplementedError and IndexError on some damaged files (mutation fuzz, review); adding them is a policy change beyond this finding, left for a follow-up. Tests (tests/test_pdf_unreadable.py): tiny in-process PDFs fail at each stage (page tree deeper than 100 levels, an 80 MB /Length, a page object without /Type, a 100 000-glyph /W range, bad ASCII85); every route answers 400 with the fixed message and logs no traceback; the log names the class, not pypdf's message. 24 of the 30 new tests fail on the old code (the other 6 check the fixtures); narrowing the catch to PyPdfError fails the 8 ValueError cases. Full suite 1519 green, 72 skipped (pypdf 6.19.0 as in requirements.lock; the new tests also pass on 6.16.2); ruff + format clean; gitleaks clean; i18n and dependencies untouched. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 37 +++++ app/api/routes/pdf_pages.py | 8 +- app/converters/document.py | 22 ++- app/converters/pdf_pages.py | 83 ++++++++--- docs/api-reference.md | 6 +- docs/api-usage-guide.md | 5 + tests/test_pdf_split.py | 13 ++ tests/test_pdf_unreadable.py | 265 +++++++++++++++++++++++++++++++++++ 8 files changed, 407 insertions(+), 32 deletions(-) create mode 100644 tests/test_pdf_unreadable.py diff --git a/CHANGELOG.md b/CHANGELOG.md index 5346c0f..3771f9d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,43 @@ Versions follow [Semantic Versioning](https://semver.org/). ## [Unreleased] +### Fixed — a PDF that pypdf can't read gets a 400, not a 500 + +pypdf raises `LimitReachedError` when a crafted file trips one of its safety +limits (declared stream length, decompressed size, font widths, page-tree +depth). It derives from `PyPdfError`, not from the `PdfReadError` the PDF paths +caught, and pypdf reads most of a file only when a page is copied or its text +is extracted, after the one guarded open. Such a file got the generic 500 and a +logged traceback instead of the 400 a broken upload gets. So did a PDF that +failed after the open with another pypdf error (in a fuzz run mostly a damaged +page object) on page extraction and splitting, and any unreadable or +password-protected PDF converted to TXT, which caught nothing. No detail +reached the client; the status and the log were wrong. + +Every pypdf step that reads the upload now answers "Could not read the PDF. +Verify the file is valid." for `PyPdfError` and pypdf's plain `ValueError`s, +and logs one line naming the error's class (not its message, which can quote +the file): + +- `/pdf/extract`, `/pdf/split`: `400` with `X-FileMorph-Error-Code: + invalid_pdf`. `/pdf/extract` used to send `invalid_page_selection` for an + unreadable file, so its web page asked the user to fix a page selection that + was fine; an unreadable file no longer gets that code. +- `/convert` (PDF → TXT, and the PDF → PDF pass-through): `400` with + `invalid_input`; `/convert/batch` reports the same message for that file + instead of "Conversion failed. Verify the file is valid." + +An error reading the server's own copy of the upload is now a 500 on +`/pdf/extract` and `/pdf/split` too, not a 400: pypdf reads the whole file into +memory first, so an `OSError` there is never the PDF's fault. And PDF → TXT +failed with a 500 on text containing an unpaired surrogate, which a broken font +map produces and UTF-8 can't encode; that character is now written as `?`. + +Tests build tiny PDFs that fail at each stage (a page tree deeper than 100 +levels, an 80 MB `/Length`, a page object without `/Type`, a font `/W` range of +100 000 glyphs, a content stream that isn't ASCII85) and check that each route +answers 400 without logging a traceback. + ### Fixed — patch-policy's `cosign verify` names an image tag that exists `docs/patch-policy.md` told readers to verify the release image diff --git a/app/api/routes/pdf_pages.py b/app/api/routes/pdf_pages.py index 3137db8..3bc4cda 100644 --- a/app/api/routes/pdf_pages.py +++ b/app/api/routes/pdf_pages.py @@ -46,6 +46,7 @@ from app.compressors.pdf import compress_pdf_to_target from app.converters.pdf_pages import ( PageSelectionError, + UnreadablePdfError, extract_pages, split_pdf, ) @@ -146,11 +147,16 @@ async def _do_extract( except PageSelectionError as exc: # Caller-safe message already (no pypdf internals). 400 — the # client's page selection or PDF was the problem, not the server. + # The code tells the UI which one, so it doesn't blame the + # selection for a file pypdf can't read. logger.info("pdf extract rejected: %s", exc) + code = ( + "invalid_pdf" if isinstance(exc, UnreadablePdfError) else "invalid_page_selection" + ) raise HTTPException( status_code=status.HTTP_400_BAD_REQUEST, detail=str(exc), - headers={"X-FileMorph-Error-Code": "invalid_page_selection"}, + headers={"X-FileMorph-Error-Code": code}, ) except Exception: logger.exception("PDF extract error") diff --git a/app/converters/document.py b/app/converters/document.py index 8fae280..e13fad1 100644 --- a/app/converters/document.py +++ b/app/converters/document.py @@ -8,7 +8,7 @@ import zipfile from pathlib import Path -from app.converters.base import BaseConverter, read_utf8_text +from app.converters.base import BaseConverter, InvalidInputError, read_utf8_text from app.converters.registry import register logger = logging.getLogger(__name__) @@ -352,10 +352,22 @@ def convert(self, input_path: Path, output_path: Path, **kwargs) -> Path: class PdfToTxtConverter(BaseConverter): def convert(self, input_path: Path, output_path: Path, **kwargs) -> Path: from pypdf import PdfReader - - reader = PdfReader(str(input_path)) - parts = [page.extract_text() or "" for page in reader.pages] - output_path.write_text("\n\n".join(parts), encoding="utf-8") + from pypdf.errors import PyPdfError + + # PyPdfError covers LimitReachedError (a crafted file tripping one of + # pypdf's resource limits), which PdfReadError alone would miss. The + # extraction sits inside the try: pypdf reads fonts and content + # streams lazily, so a broken one only fails here. + try: + reader = PdfReader(str(input_path)) + parts = [page.extract_text() or "" for page in reader.pages] + except (PyPdfError, ValueError) as exc: + # The class only — pypdf's messages can quote the file. + logger.info("unreadable PDF: %s", type(exc).__name__) + raise InvalidInputError("Could not read the PDF. Verify the file is valid.") from exc + # A broken font map can yield an unpaired surrogate, which UTF-8 + # can't encode: write that character as "?" instead of failing. + output_path.write_text("\n\n".join(parts), encoding="utf-8", errors="replace", newline="\n") return output_path diff --git a/app/converters/pdf_pages.py b/app/converters/pdf_pages.py index e26cfd9..e0cd1b2 100644 --- a/app/converters/pdf_pages.py +++ b/app/converters/pdf_pages.py @@ -28,16 +28,22 @@ pypdf parsing of the *input* PDF still happens inside ``convert()`` / ``split_pdf()``; the route invokes both through ``asyncio.to_thread`` so the (synchronous, C-accelerated) parse never blocks the event loop — -identical to every other converter. +identical to every other converter. A PDF pypdf cannot read raises +:class:`UnreadablePdfError`, a ``PageSelectionError`` with a fixed message. """ from __future__ import annotations +import logging +from collections.abc import Iterator +from contextlib import contextmanager from pathlib import Path -from app.converters.base import BaseConverter +from app.converters.base import BaseConverter, InvalidInputError from app.converters.registry import register +logger = logging.getLogger(__name__) + # Defensive ceiling on how many distinct pages a single selection may # resolve to. A crafted "1-1000000" against a 2-page PDF is already # rejected by the page-count bound, but an explicit cap keeps the parser @@ -46,8 +52,19 @@ _MAX_SELECTION_PAGES = 10_000 -class PageSelectionError(ValueError): - """Malformed or out-of-range page selection (caller-safe message).""" +class PageSelectionError(InvalidInputError): + """Malformed or out-of-range page selection (caller-safe message). + + An ``InvalidInputError``, so ``/convert`` and ``/convert/batch`` + (``pdf`` → ``pdf``) answer it as a 400, like the page routes do. + """ + + +class UnreadablePdfError(PageSelectionError): + """pypdf could not read the input PDF; raised with ``_UNREADABLE_PDF``.""" + + +_UNREADABLE_PDF = "Could not read the PDF. Verify the file is valid." def parse_page_ranges(spec: str, page_count: int) -> list[int]: @@ -122,25 +139,43 @@ def _parse_int(value: str, token: str) -> int: return n -def _open_reader(input_path: Path): - """Open a PDF with pypdf, normalising any parse failure to a safe error. +@contextmanager +def _reading_pdf() -> Iterator[None]: + """Normalise a pypdf failure on the input PDF to a safe error. - pypdf raises a small zoo of exception types (``PdfReadError``, - ``EmptyFileError``, plain ``ValueError`` from the tokenizer) on a - corrupt or non-PDF input. The magic-byte guard in the route already + pypdf raises a small zoo of exception types on a corrupt or non-PDF + input: those derived from ``PyPdfError`` (``PdfReadError``, + ``EmptyFileError``, ``LimitReachedError`` when a crafted file trips one + of its resource limits) and plain ``ValueError`` (the tokenizer, a + damaged page object). ``PdfReadError`` alone misses + ``LimitReachedError``. pypdf parses lazily, so they surface while pages + are copied or written as well as on open: every step that reads the + input runs under this guard. The magic-byte guard in the route already blocks executables; this catch turns a genuinely malformed PDF into a single caller-safe error instead of leaking pypdf internals. + + ``OSError`` stays a server error: pypdf reads the whole file into + memory first, so one can only come from our own disk. Only the + exception's class is logged — pypdf's messages can quote the file. """ - from pypdf import PdfReader - from pypdf.errors import PdfReadError + from pypdf.errors import PyPdfError try: + yield + except (PyPdfError, ValueError) as exc: + logger.info("unreadable PDF: %s", type(exc).__name__) + raise UnreadablePdfError(_UNREADABLE_PDF) from exc + + +def _open_reader(input_path: Path): + """Open a PDF with pypdf under :func:`_reading_pdf`.""" + from pypdf import PdfReader + + with _reading_pdf(): reader = PdfReader(str(input_path)) # Touch the page tree so a lazily-parsed corrupt xref surfaces here, # inside our guarded block, rather than later at iteration time. _ = len(reader.pages) - except (PdfReadError, ValueError, OSError) as exc: - raise PageSelectionError("Could not read the PDF. Verify the file is valid.") from exc return reader @@ -152,10 +187,11 @@ def extract_pages(input_path: Path, output_path: Path, pages_spec: str) -> Path: indices = parse_page_ranges(pages_spec, len(reader.pages)) writer = PdfWriter() - for idx in indices: - writer.add_page(reader.pages[idx]) - with output_path.open("wb") as f: - writer.write(f) + with _reading_pdf(): + for idx in indices: + writer.add_page(reader.pages[idx]) + with output_path.open("wb") as f: + writer.write(f) return output_path @@ -185,12 +221,13 @@ def split_pdf(input_path: Path) -> list[tuple[str, bytes]]: width = len(str(total)) outputs: list[tuple[str, bytes]] = [] - for i, page in enumerate(reader.pages, start=1): - writer = PdfWriter() - writer.add_page(page) - buf = io.BytesIO() - writer.write(buf) - outputs.append((f"page_{i:0{width}d}.pdf", buf.getvalue())) + with _reading_pdf(): + for i, page in enumerate(reader.pages, start=1): + writer = PdfWriter() + writer.add_page(page) + buf = io.BytesIO() + writer.write(buf) + outputs.append((f"page_{i:0{width}d}.pdf", buf.getvalue())) return outputs diff --git a/docs/api-reference.md b/docs/api-reference.md index 13c9c78..f55460e 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -535,9 +535,9 @@ can branch on (the `detail` text may change): | `output_cap_exceeded` | `413` | The result is larger than your tier's output cap | | `target_size_exceeds_cap` | `413` | `target_size_kb` (`/compress`, `/compress/batch`) or `target_kb` (`/pdf/compress`) is above your tier's output cap — rejected before any work | | `decompression_bomb` | `400` | The image's dimensions exceed the decoder's safety limit (`/convert`, `/compress`) | -| `invalid_input` | `400` | A problem you can fix, named in `detail` — e.g. a Markdown, CSV or JSON file that isn't UTF-8 (`/convert`) | -| `invalid_page_selection` | `400` | `/pdf/extract`: the `pages` selection is invalid, or the PDF can't be read | -| `invalid_pdf` | `400` | `/pdf/split`, `/pdf/compress`: the PDF can't be read; `/pdf/split` also for a PDF with no pages or more than 10 000 | +| `invalid_input` | `400` | A problem you can fix, named in `detail` — e.g. a Markdown, CSV or JSON file that isn't UTF-8, or a PDF that can't be read (`/convert`) | +| `invalid_page_selection` | `400` | `/pdf/extract`: the `pages` selection is invalid, or the PDF has no pages | +| `invalid_pdf` | `400` | `/pdf/extract`, `/pdf/split`, `/pdf/compress`: the PDF can't be read; `/pdf/split` also for a PDF with no pages or more than 10 000 | The redaction endpoints add codes of their own, listed under [AI operations](#ai-operations--pii-redaction-enterprise-edition-add-on). diff --git a/docs/api-usage-guide.md b/docs/api-usage-guide.md index ae7b177..fdcb4b9 100644 --- a/docs/api-usage-guide.md +++ b/docs/api-usage-guide.md @@ -500,6 +500,11 @@ Common per-file `error_message` values: CSV or JSON file in another encoding (e.g. Excel's default CSV export on Windows). Single-file `/convert` returns the same message as a `400` with `X-FileMorph-Error-Code: invalid_input`. +- `"Could not read the PDF. Verify the file is valid."` — a PDF + converted to `txt` or `pdf` that is damaged, password-protected or + over one of the PDF reader's safety limits. Single-file `/convert` + returns the same message as a `400` with + `X-FileMorph-Error-Code: invalid_input`. - `"Conversion failed. Verify the file is valid."` (compress: `"Compression failed. …"`) — any other error while processing that file, e.g. corrupt content. The details stay in the server log. diff --git a/tests/test_pdf_split.py b/tests/test_pdf_split.py index b622116..90f43b8 100644 --- a/tests/test_pdf_split.py +++ b/tests/test_pdf_split.py @@ -273,6 +273,19 @@ def test_route_extract_corrupt_pdf_is_400(client, auth_headers): ) assert res.status_code == 400, res.text assert res.status_code != 500 + # The file is the problem, not the selection — the UI says so. + assert res.headers.get("X-FileMorph-Error-Code") == "invalid_pdf" + + +def test_route_extract_page_out_of_range_is_invalid_page_selection(client, auth_headers): + res = client.post( + "/api/v1/pdf/extract", + headers=auth_headers, + files={"file": ("doc.pdf", _pdf_bytes(1), _PDF_MIME)}, + data={"pages": "9"}, + ) + assert res.status_code == 400, res.text + assert res.headers.get("X-FileMorph-Error-Code") == "invalid_page_selection" def test_route_extract_blocks_disguised_executable(client, auth_headers): diff --git a/tests/test_pdf_unreadable.py b/tests/test_pdf_unreadable.py new file mode 100644 index 0000000..78150d3 --- /dev/null +++ b/tests/test_pdf_unreadable.py @@ -0,0 +1,265 @@ +# SPDX-License-Identifier: AGPL-3.0-or-later +"""A PDF pypdf can't read gets a 400 with a fixed message, never a 500. + +pypdf raises ``LimitReachedError`` when a crafted file exceeds one of its +safety limits; it derives from ``PyPdfError``, not from ``PdfReadError``. +pypdf also parses lazily, so most failures (limits, and the plain +``ValueError`` of a damaged object) surface while pages are copied or text +is extracted, not when the file is opened. The PDF paths caught neither +there: a generic 500 plus a logged traceback instead of the clean 400. + +Tiny PDFs, built in-process, each fail at one stage: + +* a page tree nested deeper than 100 levels (a limit), while the pages + are counted; +* a stream declaring an 80 MB ``/Length`` (limit 75 MB), while its page is + copied (extract, split, ``pdf`` → ``pdf``); +* a page object without ``/Type`` (``ValueError``), while it is copied; +* a CID font whose ``/W`` range spans 100 000 glyphs (limit 65 536), during + text extraction (``pdf`` → ``txt``); +* a content stream that isn't valid ASCII85 (``ValueError``), during text + extraction. + +One more PDF pypdf *can* read: its font map yields an unpaired surrogate, +which made writing the ``txt`` output fail with a 500. +""" + +from __future__ import annotations + +import io +import logging + +import pytest +from pypdf import PdfReader, PdfWriter +from pypdf.errors import LimitReachedError + +from app.converters.pdf_pages import UnreadablePdfError, extract_pages, split_pdf + +_UNREADABLE = "Could not read the PDF. Verify the file is valid." +_CATALOG = b"<< /Type /Catalog /Pages 2 0 R >>" +_ONE_PAGE = b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>" +_PAGE_WITH_FONT = ( + b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200]" + b" /Resources << /Font << /F1 5 0 R >> >> /Contents 4 0 R >>" +) + + +def _pdf(*objects: bytes) -> bytes: + """Number ``objects`` from 1 and add a correct xref table and trailer.""" + out = bytearray(b"%PDF-1.4\n") + offsets = [] + for num, body in enumerate(objects, start=1): + offsets.append(len(out)) + out += b"%d 0 obj\n%s\nendobj\n" % (num, body) + xref = len(out) + out += b"xref\n0 %d\n0000000000 65535 f \n" % (len(objects) + 1) + out += b"".join(b"%010d 00000 n \n" % offset for offset in offsets) + out += b"trailer\n<< /Size %d /Root 1 0 R >>\nstartxref\n%d\n%%%%EOF\n" % ( + len(objects) + 1, + xref, + ) + return bytes(out) + + +def _stream(data: bytes, extra: bytes = b"") -> bytes: + return b"<< /Length %d%s >>\nstream\n%s\nendstream" % (len(data), extra, data) + + +def _deep_page_tree() -> bytes: + # Objects 2..121 are /Pages nodes, each the only kid of the one before. + nodes = [b"<< /Type /Pages /Kids [%d 0 R] /Count 1 >>" % (num + 1) for num in range(2, 122)] + return _pdf(_CATALOG, *nodes, b"<< /Type /Page /MediaBox [0 0 200 200] >>") + + +def _oversized_stream() -> bytes: + return _pdf( + _CATALOG, + _ONE_PAGE, + b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] /Contents 4 0 R >>", + b"<< /Length 80000000 >>\nstream\nBT ET\nendstream", + ) + + +def _untyped_page() -> bytes: + # pypdf lists it as a page; PdfWriter.add_page refuses it ("Invalid page object"). + return _pdf(_CATALOG, _ONE_PAGE, b"<< /Parent 2 0 R /MediaBox [0 0 200 200] >>") + + +def _huge_font_width_range() -> bytes: + return _pdf( + _CATALOG, + _ONE_PAGE, + _PAGE_WITH_FONT, + _stream(b"BT /F1 12 Tf <0041> Tj ET"), + b"<< /Type /Font /Subtype /Type0 /BaseFont /X /Encoding /Identity-H" + b" /DescendantFonts [6 0 R] >>", + b"<< /Type /Font /Subtype /CIDFontType2 /BaseFont /X" + b" /CIDSystemInfo << /Registry (Adobe) /Ordering (Identity) /Supplement 0 >>" + b" /W [0 99999 500] >>", + ) + + +def _bad_ascii85_content() -> bytes: + # The page needs a font, or pypdf skips decoding its content stream. + return _pdf( + _CATALOG, + _ONE_PAGE, + _PAGE_WITH_FONT, + _stream(b"vwxyz~>", b" /Filter /ASCII85Decode"), + b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>", + ) + + +def _surrogate_font_map() -> bytes: + # A ToUnicode map sending "A" to U+D800, half of a surrogate pair. + cmap = ( + b"/CIDInit /ProcSet findresource begin 12 dict begin begincmap\n" + b"/CMapName /X def /CMapType 2 def\n" + b"1 begincodespacerange\n<00> \nendcodespacerange\n" + b"1 beginbfchar\n<41> \nendbfchar\nendcmap\nend end\n" + ) + return _pdf( + _CATALOG, + _ONE_PAGE, + _PAGE_WITH_FONT, + _stream(b"BT /F1 12 Tf (A) Tj ET"), + b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /ToUnicode 6 0 R >>", + _stream(cmap), + ) + + +def _count_pages(data: bytes) -> None: + len(PdfReader(io.BytesIO(data)).pages) + + +def _copy_pages(data: bytes) -> None: + writer = PdfWriter() + for page in PdfReader(io.BytesIO(data)).pages: + writer.add_page(page) + writer.write(io.BytesIO()) + + +def _extract_text(data: bytes) -> None: + for page in PdfReader(io.BytesIO(data)).pages: + page.extract_text() + + +@pytest.mark.parametrize( + ("build", "stage", "error"), + [ + (_deep_page_tree, _count_pages, LimitReachedError), + (_oversized_stream, _copy_pages, LimitReachedError), + (_untyped_page, _copy_pages, ValueError), + (_huge_font_width_range, _extract_text, LimitReachedError), + (_bad_ascii85_content, _extract_text, ValueError), + ], +) +def test_each_fixture_trips_pypdf(build, stage, error): + """Guards the fixtures: should pypdf change, this names the cause + instead of the route tests failing on a 200.""" + with pytest.raises(error): + stage(build()) + + +def test_surrogate_fixture_yields_an_unpaired_surrogate(): + assert "\ud800" in PdfReader(io.BytesIO(_surrogate_font_map())).pages[0].extract_text() + + +# ── engine ─────────────────────────────────────────────────────────────────── + +# PDFs the page tools (extract, split, pdf → pdf) can't read, and pdf → txt. +_FOR_PAGES = [_deep_page_tree, _oversized_stream, _untyped_page] +_FOR_TEXT = [_deep_page_tree, _huge_font_width_range, _bad_ascii85_content] + + +@pytest.mark.parametrize("build", _FOR_PAGES) +def test_extract_pages_raises_unreadable(tmp_path, build): + src = tmp_path / "in.pdf" + src.write_bytes(build()) + with pytest.raises(UnreadablePdfError, match=_UNREADABLE): + extract_pages(src, tmp_path / "out.pdf", "1") + + +@pytest.mark.parametrize("build", _FOR_PAGES) +def test_split_pdf_raises_unreadable(tmp_path, build): + src = tmp_path / "in.pdf" + src.write_bytes(build()) + with pytest.raises(UnreadablePdfError, match=_UNREADABLE): + split_pdf(src) + + +# ── routes ─────────────────────────────────────────────────────────────────── + + +_ROUTES = [ + ("/api/v1/pdf/extract", {"pages": "1"}, "invalid_pdf", "extract", _FOR_PAGES), + ("/api/v1/pdf/split", {}, "invalid_pdf", "split", _FOR_PAGES), + ("/api/v1/convert", {"target_format": "pdf"}, "invalid_input", "pdf", _FOR_PAGES), + ("/api/v1/convert", {"target_format": "txt"}, "invalid_input", "txt", _FOR_TEXT), +] +_ROUTE_CASES = [ + pytest.param(path, data, build, code, id=f"{name}-{build.__name__}") + for path, data, code, name, builds in _ROUTES + for build in builds +] + + +@pytest.mark.parametrize(("path", "data", "build", "code"), _ROUTE_CASES) +def test_route_answers_400_without_a_logged_traceback( + client, auth_headers, caplog, path, data, build, code +): + with caplog.at_level(logging.INFO): + r = client.post( + path, + headers=auth_headers, + files={"file": ("doc.pdf", build(), "application/pdf")}, + data=data, + ) + assert r.status_code == 400, r.text + assert r.headers.get("X-FileMorph-Error-Code") == code + assert r.json()["detail"] == _UNREADABLE + assert [rec.getMessage() for rec in caplog.records if rec.exc_info] == [] + + +@pytest.mark.parametrize( + ("build", "target"), + [ + pytest.param(_huge_font_width_range, "txt", id="txt-limit"), + pytest.param(_bad_ascii85_content, "txt", id="txt-valueerror"), + pytest.param(_oversized_stream, "pdf", id="pdf-limit"), + pytest.param(_untyped_page, "pdf", id="pdf-valueerror"), + ], +) +def test_convert_batch_names_the_unreadable_pdf(client, auth_headers, build, target): + r = client.post( + "/api/v1/convert/batch", + headers=auth_headers, + data={"target_formats": [target]}, + files=[("files", ("doc.pdf", build(), "application/pdf"))], + ) + assert r.status_code == 422, r.text + assert r.json()["files"][0]["error_message"] == _UNREADABLE + + +def test_log_names_the_error_class_not_its_message(client, auth_headers, caplog): + """pypdf's message can quote the file; the log keeps only the class.""" + with caplog.at_level(logging.INFO): + client.post( + "/api/v1/pdf/split", + headers=auth_headers, + files={"file": ("doc.pdf", _oversized_stream(), "application/pdf")}, + ) + ours = [rec.getMessage() for rec in caplog.records if rec.name.startswith("app.")] + assert "unreadable PDF: LimitReachedError" in ours + assert not [msg for msg in ours if "80000000" in msg] + + +def test_pdf_to_txt_writes_an_unpaired_surrogate_as_question_mark(client, auth_headers): + r = client.post( + "/api/v1/convert", + headers=auth_headers, + files={"file": ("doc.pdf", _surrogate_font_map(), "application/pdf")}, + data={"target_format": "txt"}, + ) + assert r.status_code == 200, r.text + assert r.content.strip() == b"?" From 9159fac50803e7594b990a76ebf79af526f92cc8 Mon Sep 17 00:00:00 2001 From: MrChengLen Date: Thu, 1 Oct 2026 12:28:04 +0200 Subject: [PATCH 2/7] =?UTF-8?q?chore(changelog):=20take=20the=20entry=20ou?= =?UTF-8?q?t=20to=20merge=20main=20in=20=E2=80=94=20back=20in=20two=20comm?= =?UTF-8?q?its?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit #183 added its entry on top of the same changelog section, so this PR conflicts there. The entry goes out here, main is merged in without a conflict, and the entry comes back on top of the merged changelog. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 37 ------------------------------------- 1 file changed, 37 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 3771f9d..5346c0f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,43 +9,6 @@ Versions follow [Semantic Versioning](https://semver.org/). ## [Unreleased] -### Fixed — a PDF that pypdf can't read gets a 400, not a 500 - -pypdf raises `LimitReachedError` when a crafted file trips one of its safety -limits (declared stream length, decompressed size, font widths, page-tree -depth). It derives from `PyPdfError`, not from the `PdfReadError` the PDF paths -caught, and pypdf reads most of a file only when a page is copied or its text -is extracted, after the one guarded open. Such a file got the generic 500 and a -logged traceback instead of the 400 a broken upload gets. So did a PDF that -failed after the open with another pypdf error (in a fuzz run mostly a damaged -page object) on page extraction and splitting, and any unreadable or -password-protected PDF converted to TXT, which caught nothing. No detail -reached the client; the status and the log were wrong. - -Every pypdf step that reads the upload now answers "Could not read the PDF. -Verify the file is valid." for `PyPdfError` and pypdf's plain `ValueError`s, -and logs one line naming the error's class (not its message, which can quote -the file): - -- `/pdf/extract`, `/pdf/split`: `400` with `X-FileMorph-Error-Code: - invalid_pdf`. `/pdf/extract` used to send `invalid_page_selection` for an - unreadable file, so its web page asked the user to fix a page selection that - was fine; an unreadable file no longer gets that code. -- `/convert` (PDF → TXT, and the PDF → PDF pass-through): `400` with - `invalid_input`; `/convert/batch` reports the same message for that file - instead of "Conversion failed. Verify the file is valid." - -An error reading the server's own copy of the upload is now a 500 on -`/pdf/extract` and `/pdf/split` too, not a 400: pypdf reads the whole file into -memory first, so an `OSError` there is never the PDF's fault. And PDF → TXT -failed with a 500 on text containing an unpaired surrogate, which a broken font -map produces and UTF-8 can't encode; that character is now written as `?`. - -Tests build tiny PDFs that fail at each stage (a page tree deeper than 100 -levels, an 80 MB `/Length`, a page object without `/Type`, a font `/W` range of -100 000 glyphs, a content stream that isn't ASCII85) and check that each route -answers 400 without logging a traceback. - ### Fixed — patch-policy's `cosign verify` names an image tag that exists `docs/patch-policy.md` told readers to verify the release image From 18bb36dacea08c2e0770e94fbbd43e66e2077648 Mon Sep 17 00:00:00 2001 From: MrChengLen Date: Thu, 1 Oct 2026 12:28:49 +0200 Subject: [PATCH 3/7] =?UTF-8?q?test(pdf):=20QA=20fixtures=20for=20unreadab?= =?UTF-8?q?le=20PDFs=20=E2=80=94=20changelog=20entry=20back=20on=20top?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The changelog entry returns above #183's, with one more paragraph for the new scripts/make_testdata_pdf_unreadable.py: it writes the files the manual checklist uses (a valid three-page control file, an untyped page object, an oversized declared /Length, a 120-level page tree, a huge CID /W range, a font map yielding an unpaired surrogate, and the control file locked with a test password) to the gitignored docs-internal/testdata/pdf-unreadable/. Two runs are byte-identical; the locked copy uses RC4-128, which needs no random salt. Each file was run through the real extract, split and convert-to-txt routes and answers as the checklist says. Full suite 1530 green, 72 skipped, on main b2deb14 plus this PR (pypdf 6.19.0 as in requirements.lock); ruff + format clean; gitleaks clean. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 41 ++++++ scripts/make_testdata_pdf_unreadable.py | 167 ++++++++++++++++++++++++ 2 files changed, 208 insertions(+) create mode 100644 scripts/make_testdata_pdf_unreadable.py diff --git a/CHANGELOG.md b/CHANGELOG.md index 16668fa..c2457f1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,47 @@ Versions follow [Semantic Versioning](https://semver.org/). ## [Unreleased] +### Fixed — a PDF that pypdf can't read gets a 400, not a 500 + +pypdf raises `LimitReachedError` when a crafted file trips one of its safety +limits (declared stream length, decompressed size, font widths, page-tree +depth). It derives from `PyPdfError`, not from the `PdfReadError` the PDF paths +caught, and pypdf reads most of a file only when a page is copied or its text +is extracted, after the one guarded open. Such a file got the generic 500 and a +logged traceback instead of the 400 a broken upload gets. So did a PDF that +failed after the open with another pypdf error (in a fuzz run mostly a damaged +page object) on page extraction and splitting, and any unreadable or +password-protected PDF converted to TXT, which caught nothing. No detail +reached the client; the status and the log were wrong. + +Every pypdf step that reads the upload now answers "Could not read the PDF. +Verify the file is valid." for `PyPdfError` and pypdf's plain `ValueError`s, +and logs one line naming the error's class (not its message, which can quote +the file): + +- `/pdf/extract`, `/pdf/split`: `400` with `X-FileMorph-Error-Code: + invalid_pdf`. `/pdf/extract` used to send `invalid_page_selection` for an + unreadable file, so its web page asked the user to fix a page selection that + was fine; an unreadable file no longer gets that code. +- `/convert` (PDF → TXT, and the PDF → PDF pass-through): `400` with + `invalid_input`; `/convert/batch` reports the same message for that file + instead of "Conversion failed. Verify the file is valid." + +An error reading the server's own copy of the upload is now a 500 on +`/pdf/extract` and `/pdf/split` too, not a 400: pypdf reads the whole file into +memory first, so an `OSError` there is never the PDF's fault. And PDF → TXT +failed with a 500 on text containing an unpaired surrogate, which a broken font +map produces and UTF-8 can't encode; that character is now written as `?`. + +Tests build tiny PDFs that fail at each stage (a page tree deeper than 100 +levels, an 80 MB `/Length`, a page object without `/Type`, a font `/W` range of +100 000 glyphs, a content stream that isn't ASCII85) and check that each route +answers 400 without logging a traceback. + +`scripts/make_testdata_pdf_unreadable.py` writes byte-stable fixtures for +checking this by hand (seven small PDFs, among them a password-protected one) +to a gitignored local folder; only the script ships. + ### Changed — a sign-in lasts 30 days from login - `POST /auth/refresh` returns a new access token together with the refresh diff --git a/scripts/make_testdata_pdf_unreadable.py b/scripts/make_testdata_pdf_unreadable.py new file mode 100644 index 0000000..0434d4b --- /dev/null +++ b/scripts/make_testdata_pdf_unreadable.py @@ -0,0 +1,167 @@ +# SPDX-License-Identifier: AGPL-3.0-or-later +"""Deterministic QA fixtures for the unreadable-PDF manual test round. + +Generates the files the manual checklist refers to: + +* ``gueltig-3-seiten.pdf`` — a valid three-page PDF with a line of text per + page (the control file: extracts, splits and converts normally) +* ``kaputtes-seitenobjekt.pdf`` — its page object lacks ``/Type``: pypdf opens + the file but refuses to copy the page (extract, split, PDF → PDF) +* ``zu-grosser-stream.pdf`` — a tiny file whose page stream declares 80 MB, + over pypdf's 75 MB limit (extract, split, PDF → PDF) +* ``seitenbaum-zu-tief.pdf`` — pages nested 120 levels deep, over pypdf's + limit of 100 (extract, split, PDF → PDF, PDF → TXT; compress reads it + with pikepdf and succeeds) +* ``riesige-schriftbreiten.pdf`` — a font whose width table spans 100 000 + characters, over pypdf's 65 536 (PDF → TXT) +* ``kaputte-zeichentabelle.pdf`` — a font map pointing a letter at half of a + surrogate pair; converts to TXT with a ``?`` in its place +* ``passwort-geschuetzt.pdf`` — the control file locked with the password + ``filemorph`` + +Same call => byte-identical output on every run: the PDFs are assembled +here object by object, and the locked copy uses RC4-128, whose key comes +from the passwords and pypdf's content-derived document ID (no random salt). +Output directory defaults to ``docs-internal/testdata/pdf-unreadable/`` next +to this repo checkout (the folder is gitignored — commit this script, never +the data) and can be overridden as the first CLI argument. + +Run: + python scripts/make_testdata_pdf_unreadable.py [output_dir] +""" + +from __future__ import annotations + +import io +import sys +from pathlib import Path + +from pypdf import PdfReader, PdfWriter + +CATALOG = b"<< /Type /Catalog /Pages 2 0 R >>" +ONE_PAGE = b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>" +PAGE_WITH_FONT = ( + b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 595 842]" + b" /Resources << /Font << /F1 5 0 R >> >> /Contents 4 0 R >>" +) +HELVETICA = b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>" + + +def _pdf(*objects: bytes) -> bytes: + """Number ``objects`` from 1 and add a correct xref table and trailer.""" + out = bytearray(b"%PDF-1.4\n") + offsets = [] + for num, body in enumerate(objects, start=1): + offsets.append(len(out)) + out += b"%d 0 obj\n%s\nendobj\n" % (num, body) + xref = len(out) + out += b"xref\n0 %d\n0000000000 65535 f \n" % (len(objects) + 1) + out += b"".join(b"%010d 00000 n \n" % offset for offset in offsets) + out += b"trailer\n<< /Size %d /Root 1 0 R >>\nstartxref\n%d\n%%%%EOF\n" % ( + len(objects) + 1, + xref, + ) + return bytes(out) + + +def _stream(data: bytes) -> bytes: + return b"<< /Length %d >>\nstream\n%s\nendstream" % (len(data), data) + + +def _three_pages() -> bytes: + # 1 catalog, 2 page tree, 3 font, then a page and its text per page. + kids = b" ".join(b"%d 0 R" % (4 + 2 * i) for i in range(3)) + objects = [CATALOG, b"<< /Type /Pages /Kids [%s] /Count 3 >>" % kids, HELVETICA] + for i in range(3): + objects.append( + b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 595 842]" + b" /Resources << /Font << /F1 3 0 R >> >> /Contents %d 0 R >>" % (5 + 2 * i) + ) + objects.append(_stream(b"BT /F1 24 Tf 72 760 Td (FileMorph Testseite %d) Tj ET" % (i + 1))) + return _pdf(*objects) + + +def _untyped_page() -> bytes: + return _pdf(CATALOG, ONE_PAGE, b"<< /Parent 2 0 R /MediaBox [0 0 595 842] >>") + + +def _oversized_stream() -> bytes: + return _pdf( + CATALOG, + ONE_PAGE, + b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 595 842] /Contents 4 0 R >>", + b"<< /Length 80000000 >>\nstream\nBT ET\nendstream", + ) + + +def _deep_page_tree() -> bytes: + nodes = [b"<< /Type /Pages /Kids [%d 0 R] /Count 1 >>" % (num + 1) for num in range(2, 122)] + return _pdf(CATALOG, *nodes, b"<< /Type /Page /MediaBox [0 0 595 842] >>") + + +def _huge_font_width_range() -> bytes: + return _pdf( + CATALOG, + ONE_PAGE, + PAGE_WITH_FONT, + _stream(b"BT /F1 24 Tf 72 760 Td <0041> Tj ET"), + b"<< /Type /Font /Subtype /Type0 /BaseFont /X /Encoding /Identity-H" + b" /DescendantFonts [6 0 R] >>", + b"<< /Type /Font /Subtype /CIDFontType2 /BaseFont /X" + b" /CIDSystemInfo << /Registry (Adobe) /Ordering (Identity) /Supplement 0 >>" + b" /W [0 99999 500] >>", + ) + + +def _surrogate_font_map() -> bytes: + cmap = ( + b"/CIDInit /ProcSet findresource begin 12 dict begin begincmap\n" + b"/CMapName /X def /CMapType 2 def\n" + b"1 begincodespacerange\n<00> \nendcodespacerange\n" + b"1 beginbfchar\n<41> \nendbfchar\nendcmap\nend end\n" + ) + return _pdf( + CATALOG, + ONE_PAGE, + PAGE_WITH_FONT, + _stream(b"BT /F1 24 Tf 72 760 Td (A) Tj ET"), + b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /ToUnicode 6 0 R >>", + _stream(cmap), + ) + + +def _locked(data: bytes) -> bytes: + writer = PdfWriter(clone_from=PdfReader(io.BytesIO(data))) + writer.encrypt(user_password="filemorph", algorithm="RC4-128") + buf = io.BytesIO() + writer.write(buf) + return buf.getvalue() + + +def main() -> None: + out = ( + Path(sys.argv[1]) + if len(sys.argv) > 1 + else ( + Path(__file__).resolve().parent.parent / "docs-internal" / "testdata" / "pdf-unreadable" + ) + ) + out.mkdir(parents=True, exist_ok=True) + valid = _three_pages() + files = { + "gueltig-3-seiten.pdf": valid, + "kaputtes-seitenobjekt.pdf": _untyped_page(), + "zu-grosser-stream.pdf": _oversized_stream(), + "seitenbaum-zu-tief.pdf": _deep_page_tree(), + "riesige-schriftbreiten.pdf": _huge_font_width_range(), + "kaputte-zeichentabelle.pdf": _surrogate_font_map(), + "passwort-geschuetzt.pdf": _locked(valid), + } + for name, data in files.items(): + (out / name).write_bytes(data) + print(f"{name:30} {len(data):6} B") + print(f"-> {out}") + + +if __name__ == "__main__": + main() From 65a4be8e761b91ba585e320a8c5f315431819f3c Mon Sep 17 00:00:00 2001 From: MrChengLen Date: Thu, 1 Oct 2026 13:47:33 +0200 Subject: [PATCH 4/7] =?UTF-8?q?chore(changelog):=20take=20the=20entry=20ou?= =?UTF-8?q?t=20again=20to=20merge=20main=20in=20=E2=80=94=20back=20in=20tw?= =?UTF-8?q?o=20commits?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit #190 (urllib3 2.8.0) added its Security entry on top of the same changelog section. The entry goes out here, main is merged in without a conflict, and the entry comes back on top of the merged changelog. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 41 ----------------------------------------- 1 file changed, 41 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c2457f1..16668fa 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,47 +9,6 @@ Versions follow [Semantic Versioning](https://semver.org/). ## [Unreleased] -### Fixed — a PDF that pypdf can't read gets a 400, not a 500 - -pypdf raises `LimitReachedError` when a crafted file trips one of its safety -limits (declared stream length, decompressed size, font widths, page-tree -depth). It derives from `PyPdfError`, not from the `PdfReadError` the PDF paths -caught, and pypdf reads most of a file only when a page is copied or its text -is extracted, after the one guarded open. Such a file got the generic 500 and a -logged traceback instead of the 400 a broken upload gets. So did a PDF that -failed after the open with another pypdf error (in a fuzz run mostly a damaged -page object) on page extraction and splitting, and any unreadable or -password-protected PDF converted to TXT, which caught nothing. No detail -reached the client; the status and the log were wrong. - -Every pypdf step that reads the upload now answers "Could not read the PDF. -Verify the file is valid." for `PyPdfError` and pypdf's plain `ValueError`s, -and logs one line naming the error's class (not its message, which can quote -the file): - -- `/pdf/extract`, `/pdf/split`: `400` with `X-FileMorph-Error-Code: - invalid_pdf`. `/pdf/extract` used to send `invalid_page_selection` for an - unreadable file, so its web page asked the user to fix a page selection that - was fine; an unreadable file no longer gets that code. -- `/convert` (PDF → TXT, and the PDF → PDF pass-through): `400` with - `invalid_input`; `/convert/batch` reports the same message for that file - instead of "Conversion failed. Verify the file is valid." - -An error reading the server's own copy of the upload is now a 500 on -`/pdf/extract` and `/pdf/split` too, not a 400: pypdf reads the whole file into -memory first, so an `OSError` there is never the PDF's fault. And PDF → TXT -failed with a 500 on text containing an unpaired surrogate, which a broken font -map produces and UTF-8 can't encode; that character is now written as `?`. - -Tests build tiny PDFs that fail at each stage (a page tree deeper than 100 -levels, an 80 MB `/Length`, a page object without `/Type`, a font `/W` range of -100 000 glyphs, a content stream that isn't ASCII85) and check that each route -answers 400 without logging a traceback. - -`scripts/make_testdata_pdf_unreadable.py` writes byte-stable fixtures for -checking this by hand (seven small PDFs, among them a password-protected one) -to a gitignored local folder; only the script ships. - ### Changed — a sign-in lasts 30 days from login - `POST /auth/refresh` returns a new access token together with the refresh From 615de7992587039768688534b80aac35a7c06e92 Mon Sep 17 00:00:00 2001 From: MrChengLen Date: Thu, 1 Oct 2026 13:48:02 +0200 Subject: [PATCH 5/7] docs(changelog): entry back on top of the merged changelog The entry returns unchanged, now above #190's urllib3 Security entry. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 41 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 41 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 50905d7..606cc48 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,47 @@ Versions follow [Semantic Versioning](https://semver.org/). ## [Unreleased] +### Fixed — a PDF that pypdf can't read gets a 400, not a 500 + +pypdf raises `LimitReachedError` when a crafted file trips one of its safety +limits (declared stream length, decompressed size, font widths, page-tree +depth). It derives from `PyPdfError`, not from the `PdfReadError` the PDF paths +caught, and pypdf reads most of a file only when a page is copied or its text +is extracted, after the one guarded open. Such a file got the generic 500 and a +logged traceback instead of the 400 a broken upload gets. So did a PDF that +failed after the open with another pypdf error (in a fuzz run mostly a damaged +page object) on page extraction and splitting, and any unreadable or +password-protected PDF converted to TXT, which caught nothing. No detail +reached the client; the status and the log were wrong. + +Every pypdf step that reads the upload now answers "Could not read the PDF. +Verify the file is valid." for `PyPdfError` and pypdf's plain `ValueError`s, +and logs one line naming the error's class (not its message, which can quote +the file): + +- `/pdf/extract`, `/pdf/split`: `400` with `X-FileMorph-Error-Code: + invalid_pdf`. `/pdf/extract` used to send `invalid_page_selection` for an + unreadable file, so its web page asked the user to fix a page selection that + was fine; an unreadable file no longer gets that code. +- `/convert` (PDF → TXT, and the PDF → PDF pass-through): `400` with + `invalid_input`; `/convert/batch` reports the same message for that file + instead of "Conversion failed. Verify the file is valid." + +An error reading the server's own copy of the upload is now a 500 on +`/pdf/extract` and `/pdf/split` too, not a 400: pypdf reads the whole file into +memory first, so an `OSError` there is never the PDF's fault. And PDF → TXT +failed with a 500 on text containing an unpaired surrogate, which a broken font +map produces and UTF-8 can't encode; that character is now written as `?`. + +Tests build tiny PDFs that fail at each stage (a page tree deeper than 100 +levels, an 80 MB `/Length`, a page object without `/Type`, a font `/W` range of +100 000 glyphs, a content stream that isn't ASCII85) and check that each route +answers 400 without logging a traceback. + +`scripts/make_testdata_pdf_unreadable.py` writes byte-stable fixtures for +checking this by hand (seven small PDFs, among them a password-protected one) +to a gitignored local folder; only the script ships. + ### Added — online cancellation without login: "Cancel contracts here" (§ 312k BGB) German consumer law (§ 312k BGB) requires a subscription sold online to be From 7437eff511997dffbb55b43eb1b60e4cf39d8790 Mon Sep 17 00:00:00 2001 From: MrChengLen Date: Thu, 1 Oct 2026 14:02:28 +0200 Subject: [PATCH 6/7] =?UTF-8?q?chore(changelog):=20take=20the=20entry=20ou?= =?UTF-8?q?t=20again=20to=20merge=20main=20in=20=E2=80=94=20back=20in=20tw?= =?UTF-8?q?o=20commits?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit #185 added its entries on top of the same changelog section. The entry goes out here, main is merged in without a conflict, and the entry comes back on top of the merged changelog. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 41 ----------------------------------------- 1 file changed, 41 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 606cc48..50905d7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,47 +9,6 @@ Versions follow [Semantic Versioning](https://semver.org/). ## [Unreleased] -### Fixed — a PDF that pypdf can't read gets a 400, not a 500 - -pypdf raises `LimitReachedError` when a crafted file trips one of its safety -limits (declared stream length, decompressed size, font widths, page-tree -depth). It derives from `PyPdfError`, not from the `PdfReadError` the PDF paths -caught, and pypdf reads most of a file only when a page is copied or its text -is extracted, after the one guarded open. Such a file got the generic 500 and a -logged traceback instead of the 400 a broken upload gets. So did a PDF that -failed after the open with another pypdf error (in a fuzz run mostly a damaged -page object) on page extraction and splitting, and any unreadable or -password-protected PDF converted to TXT, which caught nothing. No detail -reached the client; the status and the log were wrong. - -Every pypdf step that reads the upload now answers "Could not read the PDF. -Verify the file is valid." for `PyPdfError` and pypdf's plain `ValueError`s, -and logs one line naming the error's class (not its message, which can quote -the file): - -- `/pdf/extract`, `/pdf/split`: `400` with `X-FileMorph-Error-Code: - invalid_pdf`. `/pdf/extract` used to send `invalid_page_selection` for an - unreadable file, so its web page asked the user to fix a page selection that - was fine; an unreadable file no longer gets that code. -- `/convert` (PDF → TXT, and the PDF → PDF pass-through): `400` with - `invalid_input`; `/convert/batch` reports the same message for that file - instead of "Conversion failed. Verify the file is valid." - -An error reading the server's own copy of the upload is now a 500 on -`/pdf/extract` and `/pdf/split` too, not a 400: pypdf reads the whole file into -memory first, so an `OSError` there is never the PDF's fault. And PDF → TXT -failed with a 500 on text containing an unpaired surrogate, which a broken font -map produces and UTF-8 can't encode; that character is now written as `?`. - -Tests build tiny PDFs that fail at each stage (a page tree deeper than 100 -levels, an 80 MB `/Length`, a page object without `/Type`, a font `/W` range of -100 000 glyphs, a content stream that isn't ASCII85) and check that each route -answers 400 without logging a traceback. - -`scripts/make_testdata_pdf_unreadable.py` writes byte-stable fixtures for -checking this by hand (seven small PDFs, among them a password-protected one) -to a gitignored local folder; only the script ships. - ### Added — online cancellation without login: "Cancel contracts here" (§ 312k BGB) German consumer law (§ 312k BGB) requires a subscription sold online to be From 7eb8e6080189aa18ed7bb57d8429b5a897ff3fbc Mon Sep 17 00:00:00 2001 From: MrChengLen Date: Thu, 1 Oct 2026 14:02:54 +0200 Subject: [PATCH 7/7] docs(changelog): entry back on top of the merged changelog The entry returns unchanged, now above #185's entries. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 41 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 41 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2893e8e..0165f66 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,47 @@ Versions follow [Semantic Versioning](https://semver.org/). ## [Unreleased] +### Fixed — a PDF that pypdf can't read gets a 400, not a 500 + +pypdf raises `LimitReachedError` when a crafted file trips one of its safety +limits (declared stream length, decompressed size, font widths, page-tree +depth). It derives from `PyPdfError`, not from the `PdfReadError` the PDF paths +caught, and pypdf reads most of a file only when a page is copied or its text +is extracted, after the one guarded open. Such a file got the generic 500 and a +logged traceback instead of the 400 a broken upload gets. So did a PDF that +failed after the open with another pypdf error (in a fuzz run mostly a damaged +page object) on page extraction and splitting, and any unreadable or +password-protected PDF converted to TXT, which caught nothing. No detail +reached the client; the status and the log were wrong. + +Every pypdf step that reads the upload now answers "Could not read the PDF. +Verify the file is valid." for `PyPdfError` and pypdf's plain `ValueError`s, +and logs one line naming the error's class (not its message, which can quote +the file): + +- `/pdf/extract`, `/pdf/split`: `400` with `X-FileMorph-Error-Code: + invalid_pdf`. `/pdf/extract` used to send `invalid_page_selection` for an + unreadable file, so its web page asked the user to fix a page selection that + was fine; an unreadable file no longer gets that code. +- `/convert` (PDF → TXT, and the PDF → PDF pass-through): `400` with + `invalid_input`; `/convert/batch` reports the same message for that file + instead of "Conversion failed. Verify the file is valid." + +An error reading the server's own copy of the upload is now a 500 on +`/pdf/extract` and `/pdf/split` too, not a 400: pypdf reads the whole file into +memory first, so an `OSError` there is never the PDF's fault. And PDF → TXT +failed with a 500 on text containing an unpaired surrogate, which a broken font +map produces and UTF-8 can't encode; that character is now written as `?`. + +Tests build tiny PDFs that fail at each stage (a page tree deeper than 100 +levels, an 80 MB `/Length`, a page object without `/Type`, a font `/W` range of +100 000 glyphs, a content stream that isn't ASCII85) and check that each route +answers 400 without logging a traceback. + +`scripts/make_testdata_pdf_unreadable.py` writes byte-stable fixtures for +checking this by hand (seven small PDFs, among them a password-protected one) +to a gitignored local folder; only the script ships. + ### Fixed — deleting a free account works on PostgreSQL `DELETE /api/v1/auth/account` failed with a 500 on PostgreSQL for every