You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Publish a deterministic, versioned OCR benchmark based on a stratified manually adjudicated sample so users can distinguish field-level transcription accuracy from whole-record accuracy.
Acceptance criteria
The benchmark documents the sampling method, strata, sample counts, adjudication inputs, parser/benchmark version, and population limitations.
The report computes field-level agreement and whole-record agreement from adjudicated fixtures.
The report includes an empirical 95% confidence interval and does not claim a universal error bound for future images.
Benchmark tests verify counts, accuracy fields, confidence-interval metadata, and reproducibility without requiring live OCR inference.
OCR confidence alone is not treated as proof of correctness or used to auto-correct transcriptions.
Parent
What to build
Publish a deterministic, versioned OCR benchmark based on a stratified manually adjudicated sample so users can distinguish field-level transcription accuracy from whole-record accuracy.
Acceptance criteria
Blocked by