Docket converts unstructured or semi-structured documents (invoices, credit notes, receipts, purchase orders, bank statements, waybills, contracts) into validated, auditable JSON objects with line citations and human review escalation.
+---------------------------------------------------------------------------------------+
| Document Ingestion |
| (PDF, PNG, JPG, TIFF, TXT) |
+---------------------------------------------------------------------------------------+
|
v
+---------------------------------------------------------------------------------------+
| Text Acquisition Tier (per page, pluggable) |
| 1. Direct PDF Text Layer (pdf_text backend) -> words, boxes, ruled tables |
| 2. OCR backend (tesseract / plugin) + confidence gate + garbled pre-flight |
| 3. Fallback backends (vision LLM by default); overruled OCR kept as witness |
| -> PageLayout: words, lines, blocks, columns, tables (normalized coordinates) |
+---------------------------------------------------------------------------------------+
|
v
+---------------------------------------------------------------------------------------+
| Classification Tier |
| 1. Fast Deterministic Keyword Rules (clear margin over the runner-up) |
| 2. TF-IDF Classifier (scikit-learn, confidence floor threshold) |
| 3. LLM Zero-shot Classifier (fallback when confidence < floor) |
+---------------------------------------------------------------------------------------+
|
v
+---------------------------------------------------------------------------------------+
| Structured Extraction Tier |
| - Versioned schema catalog: 9 built-in Pydantic schemas + registered ones |
| - JSON Schema contract enforcement via local Ollama LLM |
| - Verbatim Source Citations (field_sources / field_locations) |
+---------------------------------------------------------------------------------------+
|
v
+---------------------------------------------------------------------------------------+
| Deterministic Validation Tier |
| - Arithmetic verification to the cent (subtotal + tax + shipping - discount) |
| - Date logic & future bounds checks |
| - IBAN mod-97 check digits (ISO 7064) across all European countries & Brazil |
| - VAT check digits (all 27 EU member states, GB, CH, NO) |
| - National Tax ID checksums (US EIN, Canadian BN, Brazilian CNPJ/CPF) |
| - Verbatim citation existence & exact substring witness checks |
+---------------------------------------------------------------------------------------+
|
+-----------------------+-----------------------+
| (Passed all checks) | (Validation Error / Low Conf)
v v
+------------------------------------+ +------------------------------------+
| Validated JSON | | Human Review Queue |
| (Clean downstream persistence) | | (SQLite/PostgreSQL + review UI) |
+------------------------------------+ +------------------------------------+
Every page becomes a PageLayout: words with normalized boxes (0..1,
top-left origin, upright page), lines, blocks, text columns and tables, plus
the page's original width/height for converting back to pixels or points.
The LLM reads a serialization of that layout; the structured layout stays in
the result as the source of geometry.
A backend implements OcrBackend (name, capabilities, availability(),
recognize_page()) and returns a PageLayout. Built in:
| Backend | Input | Confidence | Word boxes | Tables | Rotation |
|---|---|---|---|---|---|
pdf_text |
PDF text layer (pdfplumber) | – | yes | ruled (drawn borders) + aligned | glyph matrices |
tesseract |
rendered page (image_to_data) |
yes | yes | aligned | OSD |
paddle (optional extra) |
rendered page, PaddleOCR 3.x | per line | yes | engine table pipeline (opt-in) + aligned | orientation classifier + fine deskew |
docling (optional extra) |
PDF or image, Docling + TableFormer | – | yes | backend cells, merged/wrapped + aligned | pipeline normalization |
vlm |
rendered page, vision LLM | – | – | – | – |
Backends are looked up by name in a registry; plugins register through the
docket.ocr_backends entry point, and an OcrBackend instance can be passed
straight to process_document(ocr_backend=...). A backend named explicitly
that cannot run (binary missing, language data missing, extra not installed)
is a configuration error raised before any page is read, with the reason and
an install hint. auto takes the first installed of tesseract, paddle.
pip install "docket-idp[paddle]"; import docket never imports it. Words
come from PaddleOCR's per-token boxes (return_word_box), joined at
whitespace; PaddleOCR scores lines, so each word carries its line's score,
and page confidence uses the same character-weighted definition as
Tesseract. DOCKET_PADDLE_MODEL=mobile (default) loads PP-OCRv5 mobile
detection + recognition; medium lets PaddleOCR pick its default for the
language (PP-OCRv6 medium for Latin scripts). One model reads one script
family, so DOCKET_OCR_LANGUAGES must stay within Latin, East Slavic,
Cyrillic, Greek, Arabic, Korean or CJK — en,ru is refused at startup.
DOCKET_PADDLE_TABLES=true runs TableRecognitionPipelineV2 on the OCR
result already computed; its cell boxes become detection="backend" tables.
Models download once to ~/.paddlex/official_models; docket disables
PaddleX's model-hoster connectivity probe so cached models load offline.
Observed limit: the orientation classifier left a sparse page (three text
lines) turned 90° uncorrected, where Tesseract OSD corrected it.
pip install "docket-idp[docling]" enables the lazy docling backend. It
uses Docling's standard PDF/image pipeline and TableFormer in accurate mode
by default. DOCKET_DOCLING_TABLE_MODE=fast trades quality for throughput;
DOCKET_DOCLING_CELL_MATCHING=false uses the structure model's own cells when
matching them back to document text merges columns incorrectly. Docling cell
offsets become row_span / column_span, embedded newlines remain wrapped
cell text, and every table keeps normalized source boxes.
The shared geometry pass also recognizes conservative borderless two-column
numeric tables, while rejecting colon-ended label/value forms. Physical rows
with fewer occupied bands are attached to the preceding logical cells. Raster
OCR applies a projection-based fine deskew before recognition and records the
clockwise correction as PageLayout.deskew_angle; 90-degree orientation stays
in PageLayout.rotation.
- A PDF page with a usable text layer is taken as is. Unusable means fewer
than 20 characters, or more than 10 % unmapped
(cid:N)glyphs. - Otherwise the primary backend, then each fallback (default:
vlm). A reading is accepted when it has text and, if the backend reports confidence, page confidence ≥DOCKET_OCR_MIN_CONFIDENCE(0.60). Tesseract's page confidence is the character-weighted share of text in lines whose mean word confidence clears the word floor. - If nothing is accepted, the best rejected reading is used and the page is marked degraded, which sends the document to review.
Mixed PDFs fall out of this naturally. When the accepted reading has no word boxes (the vision model), the OCR reading it overruled is kept as the page's witness: validation cross-checks the model's numbers against it, and citations are located in it.
Two escalations re-run the chain with every reading but the last backend's
rejected: before extraction, when a cheap text model judges the OCR text
garbled (looks_garbled); after validation, when OCR text passed its gate
but the extraction failed validation (fewer or equal errors wins, ties go to
the re-read).
docket.layout.analysis.build_page is shared by every backend with word
boxes. Pure geometry, no keywords:
- Rows: words overlapping vertically by ≥40 % of the smaller height. An engine's own line identity (Tesseract block/paragraph/line) is respected, so skewed lines don't interleave.
- Segments: a gap wider than 1.5 × the page's median word height splits
a row; serialized as
|. - Aligned tables: ≥2 consecutive rows with ≥3 segments that fall into ≥3 shared column bands. Ruled tables come from pdfplumber's rulings, including row/column spans, and take precedence.
- Text columns: an ink-free gutter over ≥4 consecutive rows with substantial text on both sides (median ≥12 characters, ≥20 % of the page width per side). Reading order inside such a region is column-major.
- Blocks: consecutive lines in the same column/table with at most one line height between them.
Serialization writes lines in reading order, with [TABLE n: R rows x C columns] and [COLUMN n] marker lines.
Known limits:
- A table cell that wraps onto a second line becomes its own row (or breaks the table run); it is not merged back into the cell above.
- Two-column tables (description | amount) are not tables — they read as
lines with a
|separator. Tables need ≥3 columns. - A borderless table whose columns are separated by less than 1.5 × word height is read as plain lines.
- Text columns with narrow gutters (below 1.5 × word height), or with short lines (label/value blocks), are read row by row.
- Rotation is corrected in 90° steps; skew is not.
- Upside-down PDF pages with a mirrored text layer, and vertical CJK text, are not handled.
The extraction model returns only page and quote for each field. The
pipeline matches the quote against the page's words (whitespace-free,
case-folded character stream; exact first, then a fuzzy window that must
score ≥0.8) and records bbox, word_ids, a confidence (match score ×
mean word confidence) and located_by. The model is never asked for
coordinates.
Every occurrence is kept, not just the first: SourceLocation.regions
holds each contiguous place the quote was found, and status says how it
resolved — verified (one exact match), fuzzy (no exact match, one close
window), conflicting (several exact matches; the document doesn't say
which one is the source), unlocated (no geometry or the quote isn't on
the page). bbox / word_ids remain the first region's coordinates for
back-compatibility. DocumentResult.highlights(page=None) returns
(field, source, region) triples for drawing provenance boxes over the
original pages, optionally filtered to one page.
Rotation detection with Tesseract OSD added 0.34 s in a single run on one sample page
(form_funsd_00.png, Apple Silicon); disable it with
DOCKET_OCR_DETECT_ROTATION=false if your scans are always upright.
process_document(source, ProcessOptions) -> DocumentResult is the one
entry point; the CLI and the HTTP API call it. Its stages are plain
functions in docket.pipeline: acquire → select_schema → extract →
validate_extraction → review; export_document(result, format) is the
separate last step and refuses results that failed or need review unless
told otherwise.
- Options (
docket.options):OcrOptions(chain, languages, engine settings),document_type/schema_model(either one skips classification; given both, they must agree),classify,include_layout,escalate,ReviewOptions(enqueue, classification threshold, queue location). Every unset option comes fromDOCKET_*, else the built-in default;resolve()applies that and validates it. - Errors: anything detectable up front — unknown or unavailable backend,
unknown language, unknown document type, a schema that isn't a Pydantic
model — raises
ConfigurationErrorbefore the first page is read. A document that fails later returnsstatus="failed"witherror.code(unsupported_document,unreadable_file,no_text) anderror.stage. - Status:
failediferroris set,needs_reviewif any review reason applies, otherwisesucceeded. Failed documents are not written to the review queue.
process_batch(sources, ProcessOptions, BatchOptions) runs
process_document over a directory (optionally recursive and
glob-filtered), a glob pattern or any iterable of paths:
- Sources are discovered lazily and at most
2 × workersdocuments are submitted ahead of the one being returned, so a large directory is never listed or loaded into memory whole;keep_results=Falsepluson_resultstreams results out without keeping them. - Results come back in input order. Each document runs in a fresh context, so its LLM usage counters are its own.
- A failing or crashing document becomes a failed
DocumentResultand aBatchError;fail_faststops submitting after the first one and counts the rest asskipped. docket.limitsbounds the expensive calls process-wide: at mostDOCKET_LLM_CONCURRENCYLLM requests andDOCKET_OCR_CONCURRENCYOCR engines at once, whatever the number of workers or HTTP jobs.- A checkpoint (JSON Lines of finished results) is appended as documents finish; a rerun reuses the result of any source whose content hash is unchanged. Failed results are retried.
BatchResultcarries counts (total,succeeded,needs_review,failed,skipped), errors, elapsed time and aggregate metrics (pages, LLM calls and tokens, VLM pages, escalations, mean/median seconds per document, time per stage).
docket.export.tabular renders results as a fixed-column summary CSV (each
schema's summary map fills document_number, document_date, issuer,
recipient, currency, subtotal, tax_amount, total_amount), a
line-item CSV linked by document_id (each schema's line_items map), and
JSON Lines.
A job is a batch of uploaded files stored under DOCKET_JOBS_DIR/<job_id>/:
job.json (options, per-document index, counts, status), results.jsonl
(the batch checkpoint, so a restarted server resumes unfinished jobs
without redoing finished documents) and uploads/, which is deleted when
the job finishes. Uploads stream to random file names — only the original
suffix is kept — and are checked against per-file, per-job and page limits
before a job exists; anything rejected is deleted. Callers choose schemas by
registered id only. DOCKET_MAX_CONCURRENT_JOBS jobs run at once.
docket.config declares every setting once (SETTINGS: attribute, TOML
key, environment variable, type, bounds, default, help) and loads them as
defaults < config file < environment into module attributes that the rest
of the package reads at call time; explicit arguments are applied on top by
options.resolve(), the CLI and the HTTP form handling. The file is TOML
(--config, else DOCKET_CONFIG), with paths relative to the file. Import
reads only those and the environment, so an embedding application's working
directory never changes docket's settings; the applications (docket,
docket-api, the demo) call configure_app(), which also loads .env and
falls back to ./docket.toml. Loading never raises: an invalid value keeps its
default and is recorded; config.check() raises one ConfigurationError
naming every problem, with its source (DOCKET_OCR_DPI='x',
docket.toml: [ocr] dpi = 'x', unknown keys, unknown OCR language, malformed
Paddle device). The CLI checks before any command but config, docket-api
before starting the server and the app's lifespan before accepting a
request, and resolve() before process_document() reads a file.
docket config show lists values with sources, secrets masked.
docket.catalog holds every schema docket can extract, as SchemaSpecs
keyed by (schema_id, version):
- Spec: model, version, display name, description (read by the LLM
classifier), status (
stable/experimental), weighted keywords, TF-IDF example sentences, cited field paths, validators(document, ValidationContext) -> issues, and migrations. - Shared blocks (
catalog/common.py):Party,Address,TaxIdentifier,Money,DocumentReference,BankAccount,LineItem,Citation,CitedDocument. - Built-ins: invoice 2.0 and purchase order 2.0 (parties as
Party, PO numbers as references), credit note (same billing structure and rules as the invoice, plus its own), receipt 1.2 (1.1 plus the payment block) and contract / bank statement / waybill 1.1 (flat, as before, minusdoc_type). A GST/VAT "tax invoice" is an invoice: its registration numbers and tax are invoice fields. Fixtures and expected extractions for each are intests/fixtures/catalog/. - Versions and migrations: several versions of a schema can be
registered; the latest is the default and
--schema-version/ProcessOptions.schema_versionpick another. A result keeps itsschema_version; readingresult.documentunder a newer registration runs the migration chain (all built-in 1.0 → current steps are automatic). - Registration errors are raised at registration: bad id or version,
duplicate, a model that can't be described as JSON Schema, a
field_locationsthat isn'tdict[str, Citation], cited paths the model doesn't have, a model already registered under another id. - Sources: built-ins,
register_schema(), thedocket.schemasentry point, and unregistered models passed asschema_model(id =module:Class, no version).
Classification determines which Pydantic schema will govern extraction:
Every tier reads the schema catalog, so a registered schema takes part in all three.
- Keyword Rules: each schema's weighted
keywords(English and Spanish cues, plus the document's own name in German, French, Italian, Dutch, Portuguese and Polish). If the top schema leads the runner-up by 2 points, it classifies immediately; confidence is the winner's share of all matched weight. Similar terms across the 9 built-in schemas can lower confidence and send a document to review. - TF-IDF Classifier: word and character n-grams trained on each schema's
examples(paraphrases in EN, ES, DE, FR, IT, NL, PT; 12–26 per built-in schema). AboveDOCKET_TFIDF_CONFIDENCE_FLOORit skips the LLM call. It is skipped when any registered schema brought no examples, because it would file that type under a neighbour. On the 34-sentence held-out set intests/test_classify_tfidf_multilingual.pyit is right 33 times, confident 19 times and never confidently wrong; on the clean bank-statement fixture it was confidently wrong (invoice, 0.71) — the rules tier answers first there. - LLM Fallback: Invoked only when rule-based and TF-IDF classifiers cannot make a confident decision.
- JSON Schema Contracts: The target schema is the Pydantic model of the selected catalog schema; its JSON Schema is the extraction contract.
- Nested citations:
field_locationskeys are field paths (seller.name,references[0].number); validation and quote location follow them. - Line-item citations: every row of every repeated list must be cited per field (
line_items[0].quantity,items[0].price,transactions[0].amount), the quote being that row's own text. Grounding, location and review treats an item value exactly like a top-level amount. - Verbatim Evidence: The model must provide verbatim quotes (
quote,page) for extracted values. - Multilingual Parsing: Supports both European (
1.234,56 €) and American ($1,234.56) numerical conventions.
docket.templates is the deterministic fast path for known vendors. A
VendorTemplate belongs to one schema, requires every issuer pattern to
match, and maps explicit page labels and table columns into schema paths.
Values are type-coerced with the same amount and date conventions as model
extraction, and every field and line-item cell receives a citation.
The pipeline runs a matching template after schema selection and before the
extraction model. A candidate is accepted only when Pydantic validation and
the normal deterministic business rules produce no errors. A missing rule,
unparseable value, incomplete row or invalid total discards the candidate and
runs model extraction; templates therefore reduce model calls but never
bypass validation. ProcessingMetrics.template_id records the accepted
template and the benchmark reports template hit rate, latency and the minimum
number of structured-extraction calls avoided.
The registry is available through Python, docket templates, and
/vendor-templates. Built-ins are fictional golden-corpus examples rather
than a claim to recognize arbitrary real vendors.
Validation never calls a model. It executes deterministic arithmetic and mathematical checksum algorithms:
-
Citation grounding: every numeric field — top-level and line-item (
_item_numeric_fieldskeys every row value by schema path) — must be found on its cited line. A value the cited row actually implies also grounds: a derived unit price (4.98 / 2 = 2.49) or line total (2 × 58.50 = 117.00) counts only when both operands are printed on that row. A quantity of 1 is the implicit single item, not a fabricated amount. An item value contradicted by a garbled witness is a warning, not a blocker, when the rows close their own arithmetic (items sum to the stated subtotal under either coupon layout). A list cited element-wise (parties_a[0]) satisfies the requirement on the whole list. -
Totals & Line Items: Verifies
subtotal + tax + shipping - discount == total_amountand the other sums to the cent:validate._iscloseallows an absolute difference of 0.01 (one rounding step of a printed two-decimal amount), whatever the size of the amount. A relative tolerance was used before; it let a 10.00 gap through on a 1,100 total. The EN 16931 rules downstream compare exact decimals, so a looser check here would only move the failure to export time. -
IBAN: ISO 7064 MOD 97-10 check digits for all European nations and Brazil. Identifies non-IBAN systems (US, Canada) and requests routing numbers instead.
-
VAT / Sales Tax: Algorithmic check-digit verification across all 27 EU member states, the UK, Switzerland, and Norway.
-
Americas Tax IDs: Modulo-11 CNPJ/CPF checks for Brazil, Luhn mod-10 checks for Canadian Business Numbers (BN), and prefix verification for US EINs.
-
B2B Invoicing: Validates customer tax IDs, ISO 9362 SWIFT/BIC codes, SKU and unit of measure on line items, and mathematical cross-checks tax rate percentage against subtotal and tax amounts.
-
Receipts & Expenses: Validates retail/restaurant balancing
subtotal + tax + tip - discount == total_amount, line item pricingquantity * unit_price == price, merchant tax IDs (VAT and national), and 4-digit payment card format. Coupons print above the SUBTOTAL (stated subtotal already discounted) or below it, so the items-sum and balancing rules accept either layout and flag only when neither closes. -
Bank Statements: Validates balance equation
opening_balance + total_deposits - total_withdrawals == closing_balance, sums of transaction deposits and withdrawals, running balance continuity across consecutive transaction entries, and bank IBAN check digits. -
Waybills / Consignment Notes (CMR, ТОРГ-12): Validates physical logistics balancing: sum of item quantities vs
total_quantity, sum of gross weights vstotal_gross_weight_kg, line item pricingquantity * unit_price == price, distinct consignor and consignee, and carrier tracking. -
Citation Grounding: Asserts that every cited quote exists in the document and contains the claimed numerical value.
-
Contract Legal Validation: Asserts counterparty sanity (an entity cannot contract with itself; parent/subsidiary relationships trigger reviews), verifies that parties, governing law, payment terms, and signatories exist verbatim in the source text, checks term dates and notice/cure period limits, and runs automated risk factor assessment (unlimited liability, auto-renewal trap).
Documents that fail any error-level validation rule, fail extraction, or carry low classification confidence are routed to the Review Queue (sqlite:///data/review.db by default, PostgreSQL in multi-worker deployments):
review_tasksstores current state; append-onlyreview_revisionspreserves every claim, correction, validation and decision.- Claim uses expiring leases and opaque lock tokens. Updates also carry the expected version, preventing concurrent reviewers from overwriting each other.
- Corrections are merged over the original extraction and rerun through its Pydantic schema and deterministic business validators. Approval is refused while error-level issues remain.
- The FastAPI workflow exposes claim, release, revalidate, history, originals and rendered pages. The
/verifyworkbench draws normalized source bboxes directly over those pages.
Deterministic multi-document audits connect extracted records across the procurement and expense lifecycle:
- 3-Way Matching (PO ↔ Waybill ↔ Invoice):
- Reconciles Purchase Order authorizations against Waybill physical deliveries and Invoice billing claims.
- Detects unit price variances (
PRICE_VARIANCE) exceeding tolerance when invoice price exceeds PO unit price. - Detects unfulfilled billing (
UNFULFILLED_BILLING) when invoiced quantities exceed physically delivered quantities on the waybill. - Verifies counterparty consistency across buyer/consignee/customer and vendor/consignor/seller.
- Invoice ↔ Purchase Order (2-Way Matching):
- Reconciles line items by SKU or description.
- Detects unit price variances (
PRICE_VARIANCE) exceeding configurable thresholds (price_tolerance_pct). - Detects quantity over-billing (
QUANTITY_OVERBILLING) and unordered goods (UNORDERED_ITEM). - Verifies counterparty consistency and total amounts.
- Contract ↔ Invoices (Budget & Compliance Audit):
- Verifies invoice counterparties belong to the contracted parties.
- Asserts invoice dates fall within the contract's effective and expiration window.
- Tracks cumulative invoiced totals against the contract value ceiling (
BUDGET_EXCEEDED).
- Receipt ↔ Bank Transactions (Expense Reconciliation):
- Matches receipts against card/bank statements using transaction date windows (clearing delays), exact currency, card last four digits, and total amounts.
Extracted and validated records are converted to electronic invoicing formats
without external cloud dependencies. Formats for one ERP or accounting system
(SAP, Xero, QuickBooks, 1C) are not built in: their account codes, tax codes
and posting rules belong to the application that runs that system, which
registers them with register_exporter() or a docket.exporters entry point
(see examples/custom_exporter.py and examples/exporter_plugin/).
- Facturae 3.2.2 (
docket.export.facturae): Spanish electronic invoice (FACe) with Party Tax Identification and TaxesOutputs breakdown. - EN 16931 e-invoices (
docket.export.en16931), see below.
One semantic model, two syntaxes. en16931.semantic(doc) maps an Invoice
or CreditNote onto EN 16931 business terms (BT/BG): parties
with VAT (BT-31/48), tax registration and legal ids, contacts, electronic
addresses, the payment account as credit transfer (BG-17), references
(order BT-13, preceding invoice BG-3), lines with unit codes
mapped to UN/ECE Rec 20, document-level allowance and charge for discount
and shipping, and one VAT breakdown per rate. It computes every total the
standard defines from the lines and refuses (EN16931Error, surfaced as
ExportError) whatever it cannot represent without inventing data: no
lines (BR-16), no determinable rate, tax or line sums that disagree with
the stated amounts (BR-CO-10/14), a discount or charge over several rates.
Only VAT categories S and Z are produced; exemptions (E, AE, K,
G, O) need an exemption reason the extraction schema doesn't carry.
_Ubl and _Cii render that model. The registered formats differ only in
syntax, specification identifier (BT-24) and business process (BT-23):
| Format | Syntax | BT-24 | Validated as |
|---|---|---|---|
ubl |
UBL 2.1 | urn:cen.eu:en16931:2017 |
en16931 |
peppol |
UBL 2.1 | ...#compliant#urn:fdc:peppol.eu:2017:poacc:billing:3.0 |
peppol |
xrechnung-ubl |
UBL 2.1 | ...#compliant#urn:xeinkauf.de:kosit:xrechnung_3.0 |
xrechnung |
xrechnung-cii |
CII D16B | same | factur-x-xrechnung |
factur-x-en16931 |
CII D16B | urn:cen.eu:en16931:2017 |
factur-x-en16931 |
factur-x-basic |
CII D16B | ...#compliant#urn:factur-x.eu:1p0:basic |
factur-x-basic |
A credit note becomes a UBL CreditNote (type 381) or CII TypeCode 381.
The BASIC profile omits what its schema doesn't allow (seller item id,
contacts, BIC). The Factur-X exporters produce CII XML; docket.einvoice
can embed it with XMP into a source PDF, extract it again, and verify the
complete round trip. XML uses the official profile rules and the PDF/A-3
container uses the external veraPDF CLI. Every format above passes its official rules
on the complete test invoice (tests/test_einvoice.py).
docket.einvoice (extra [einvoice]: lxml, saxonche) validates an
e-invoice with the official artifacts. Those with verified redistribution
terms ship in src/docket/einvoice/resources/; the others are downloaded on
request (docket einvoice fetch, fetch_einvoice_resources()) into the
download root (DOCKET_EINVOICE_DOWNLOADS, default ~/.cache/docket/einvoice):
| Artifact | Version | License | Used for | Delivery |
|---|---|---|---|---|
| KoSIT validator configuration | XRechnung 3.0.2, 2026-08-31 | Apache-2.0 + OASIS notice | UBL 2.1 XML Schemas | shipped |
| UN/CEFACT CII D16B schemas (from the KoSIT configuration) | D16B | not verified | CII XML Schemas | downloaded |
| CEN EN 16931 validation artefacts | 1.3.16 | EUPL-1.2 | EN 16931 Schematron, UBL and CII | shipped |
| KoSIT XRechnung Schematron | 2.6.0 | Apache-2.0 | CIUS XRechnung (BR-DE-*) | shipped |
| OpenPeppol BIS Billing 3.0 | 3.0.20 | not verified | Peppol rules, compiled from .sch with SchXslt 1.10.1 (MIT) | downloaded |
| Factur-X / ZUGFeRD | 1.09 (from the factur-x 6.8 package) | not verified | profile XSD and Schematron, MINIMUM to EXTENDED | downloaded |
resources/manifest.json records every file's SHA-256 (null for the
Peppol XSLT, which is compiled locally and vouched for by its hashed .sch),
the archive URL and hash it came from, version, license and whether it is
redistributable; artifacts.verify() checks the installed copy. The pinned
sources live in docket/einvoice/sources.py, shared by the update script
and fetch.py, which installs a download only after every file matches. scripts/update_einvoice_resources.py downloads the pinned
archives (refusing a hash mismatch), extracts only what validation needs,
compiles the Peppol Schematron and rewrites the manifest; test fixtures
(official examples, the XRechnung test suite, Peppol's unit tests) go to
tests/fixtures/einvoice/official/ (--fixtures-only writes just those).
The artifacts that are not redistributed, and Peppol's fixtures, are
git-ignored and excluded from the wheel and sdist. To update: change the
pinned URL and hash in sources.py, rerun, run the tests, commit the diff.
validate_einvoice(source, options):
- Reads XML, or the embedded
factur-x.xml/zugferd-invoice.xml/xrechnung.xmlof a PDF (pypdfium2). XML with a DOCTYPE is refused; entities are never resolved and nothing is fetched from the network. - Detects the syntax from the root element (UBL Invoice, UBL CreditNote,
CII) and the declared profile from BT-24. A requested profile that
differs from the declared one is reported as
DOCKET-PROFILE-MISMATCHand validation continues with the requested rules. - Runs the XML Schema (lxml), then each Schematron (SaxonC-HE, XSLT 2.0, SVRL output) of the profile's plan: EN 16931 core, plus Peppol or XRechnung rules; Factur-X profiles use their own combined Schematron. Schematron is skipped when the XML Schema fails, because its rules assume a schema-valid tree.
- Returns
EInvoiceValidationResult:valid,detected_format,syntax,profile,declared_profile_id,validator/validator_version,validation_resource_version, oneLayerReportper layer (artifact, version, ran, passed, rules fired) andissues(code, severity from the Schematron flag, message, XPath or XSD line, rule source with version, layer).
Compiled stylesheets and schemas are cached per process behind a lock
(SaxonC is not thread-safe). The same function serves
docket validate-einvoice, POST /validate/einvoice and
ExportOptions(validate_einvoice=True). Without the extra it raises
EInvoiceUnavailable, a ConfigurationError (CLI exit 3, HTTP 503
einvoice_unavailable); when the profile's artifacts are neither shipped nor
downloaded it raises its subclass EInvoiceResourcesMissing, naming
docket einvoice fetch, before reading anything further.
The tests run all official examples of each artifact, all 227 cases of Peppol's own UBL unit suite (expected rule ids fire, expected successes don't), and negative cases: missing mandatory field (BR-07, BR-DE-15), wrong totals on XSD-valid XML (BR-CO-15/16), wrong or unknown tax category (BR-Z-05, BR-CL-18), bad endpoint scheme (PEPPOL-EN16931-CL008), unknown and mismatched profile identifiers.