A RAG system with a real evaluation layer: a hand-built golden set, baseline retrieval, and iterative comparison across chunking/retrieval strategies — not just "a RAG demo," but "a RAG demo where quality claims are backed by measured numbers."
Odoo's official documentation (public, unrelated to any employer's systems), covering five ERP domains: purchase, invoicing/accounting, inventory, and manufacturing.
- Purchase: https://www.odoo.com/documentation/19.0/applications/inventory_and_mrp/purchase.html
- Accounting & Invoicing: https://www.odoo.com/documentation/19.0/applications/finance/accounting.html
- Inventory: https://www.odoo.com/documentation/19.0/applications/inventory_and_mrp/inventory.html
- Manufacturing: https://www.odoo.com/documentation/19.0/applications/inventory_and_mrp/manufacturing.html
Local corpus text (verbatim page content, used as the retrieval index source) lives under
data/corpus/<module>/<page>.md — filename matches the source URL's basename, so the doc a
given golden-set question was drawn from is always data/corpus/<module>/<url-basename>.md.
- Golden set: 70 Q&A pairs grounded in the corpus above (
data/golden_set.json, all verified) - v0 baseline: fixed 512-token chunks + plain vector search (
results/v0.json: Hit Rate@5 97.14%, P95 145ms, $0/query on local bge-m3) - v1: chunking strategy (parent/child chunks or semantic splitting)
- v2: hybrid retrieval (BM25 + vector fusion)
- v3: reranking
- Judge calibration: human-human agreement (kappa) → judge-human agreement → bias check
- Comparison table: 4 versions × 4 metrics × P95 latency × per-query cost