Grounded question answering over a corpus of IETF RFCs (HTTP / Web API standards). Hybrid retrieval, optional cross-encoder reranking, and answers that either carry resolvable source citations or refuse to answer.
Everything runs locally on open-weights models. No API key, no per-query cost.
Three effects clear zero; two do not. The two that do not are the ones a RAG project is normally expected to claim.
A retrieval-augmented QA system built end to end and — the part that matters — measured at every step, with a confidence interval on every comparison and a check that fails if any number quoted on a CV cannot be resolved from a results file.
| Corpus | 34 IETF RFCs, 2,779 section-bounded chunks |
| Retrieval | BAAI/bge-small-en-v1.5 dense + Okapi BM25 written from scratch, weighted fusion |
| Reranking | ms-marco-MiniLM-L-6-v2 cross-encoder — measured, then disabled |
| Generation | Qwen2.5-1.5B-Instruct emitting structured claims, each verified before it reaches the user |
| Evaluation | 65 hand-built queries (52 answerable) across 8 question types, paired-bootstrap significance testing |
| Service | FastAPI, request IDs, structured JSON logs with redaction, Prometheus metrics, Docker |
| Tests | 86, plus 14 verification checks |
Phase 0 — design. Chose a domain where citation precision matters and where retrieval is genuinely hard: RFCs mix prose, ASCII tables, ABNF grammars and exact identifiers, and the corpus deliberately includes an obsoleted document (RFC 7234, superseded by RFC 9111) so conflicting sources are a real case rather than a hypothetical one. Wrote 65 evaluation queries by hand against the source text, then machine-verified every one: the cited section must exist in the index and every reference keyword must literally appear in it. That check caught 10 queries where my keywords were paraphrases rather than quotations.
Phase 1 — ingestion. Two RFC text layouts (paginated and unpaginated) parsed into section-addressable documents; page furniture stripped; boilerplate excluded; chunks that never cross a section boundary, because a chunk spanning two sections cannot be cited honestly.
Phases 2–4 — retrieval. Swept encoders, chunk sizes, overlaps, fusion methods and fusion weights, then reranking depth — each against a stated baseline, each with a paired bootstrap over 20,000 resamples. No configuration was chosen by intuition.
Phase 5 — generation. Prompting alone failed (see below). Replaced trust with verification: the model proposes structured claims, the system checks each one against the source it cites, and composes the answer only from claims that pass.
Phases 6–7 — service and robustness. FastAPI with validation, error handling, observability and Docker; injection, failure-path, log-hygiene and concurrency testing.
Phase 8 — reporting. Every CV-facing number registered in
eval/cv_claims.json and resolved from a results file by an automated check.
1. The encoder dominated every architectural change. Swapping
all-MiniLM-L6-v2 for BAAI/bge-small-en-v1.5 improved Recall@5 by 23.3% and
MRR by 22.2%, with 5 of 6 metrics significant at 95%. No retrieval trick in
this project came close to that.
2. Cross-encoder reranking did not help, and is off by default. Over a weak first stage it lifted Recall@5 by 11.8% (95% CI [+0.003, +0.147]). Over the strong first stage that actually ships, zero of six metrics were distinguishable from noise, while p50 latency rose from 14.7 ms to 113.7 ms — and quality degraded monotonically as the reranker was given more candidates. The component is built, benchmarked and disabled, with the numbers to defend it.
3. Prompting could not make a 1.5B model cite its sources; verification could.
Asked for prose with inline citations, it produced zero — copying the RFCs'
own bibliography markers ([URI], [RFC9110]) instead. Few-shot examples did not
fix it. What fixed it was changing the contract: the model emits structured JSON
claims, and every claim is checked against its cited source before the answer is
composed — including a rule that a claim may not introduce a rare corpus term its
evidence never mentions. Result: 100% citation validity across 65 queries, zero
uncited answers, and correct refusal on all 8 questions the corpus cannot
answer — including probes the model answers confidently from pretraining.
Retrieval, 52 answerable queries (Apple M4, MPS):
| Configuration | Recall@5 | MRR | nDCG@10 | p50 |
|---|---|---|---|---|
| BM25 alone | 0.5513 | 0.5089 | 0.5113 | 0.7 ms |
| Dense MiniLM (baseline) | 0.5769 | 0.5007 | 0.5322 | 8.5 ms |
| Dense BGE-small | 0.7115 | 0.6117 | 0.6218 | 12.0 ms |
| Hybrid (shipped default) | 0.6827 | 0.5912 | 0.6012 | 14.7 ms |
| Hybrid + reranking | 0.6859 | 0.6066 | 0.6113 | 113.7 ms |
Generation and grounding, all 65 queries:
| Metric | Value |
|---|---|
| Citation validity rate | 1.0000 |
| Answers with zero citations | 0 |
Correct refusal on unanswerable |
1.0000 (8/8) |
| Citation correctness (answered) | 0.7250 |
| Evidence hit rate | 0.7885 |
| Uncited assertive-sentence rate | 0.0358 |
Service: 1,000 requests across concurrency 1–16, zero errors, ~50–90 rps retrieval-only.
pip install torch && pip install -r requirements-dev.txt
python3 scripts/fetch_corpus.py # 34 RFCs, SHA-256 verified (~30 s)
python3 scripts/build_index.py --rebuild # parse, chunk, embed, index (~55 s)
uvicorn ragkb.api:app --port 8000curl -s localhost:8000/query -H 'content-type: application/json' \
-d '{"query":"What does the 429 status code mean?","top_k":5}' | jqRetrieval-only (no LLM, ~15 ms):
curl -s localhost:8000/query -H 'content-type: application/json' \
-d '{"query":"cache-control no-store","generate":false}' | jq '.retrieval[].citation'Docker (CPU-only, slower than the MPS figures quoted here):
docker compose up --buildThe corpus and index are not committed — they are regenerated by the two commands
above and pinned by SHA-256 in data/corpus_manifest.json. Everything needed to
audit the numbers without running a model is committed: the evaluation set,
per-query evaluation records, aggregate tables, significance tests and the claims
registry.
python3 -m pytest tests/ -q # 86 tests
python3 scripts/evaluate_retrieval.py --suite all # phases 2-4
python3 scripts/evaluate_retrieval.py --suite final # final configuration
python3 scripts/evaluate_generation.py --tag main # phase 5, grounding
python3 scripts/benchmark_latency.py # per-stage latency
python3 scripts/load_test.py --sweep 1,2,4,8,16 # concurrency (service must be up)
python3 scripts/make_plots.py # figures, from results/
python3 scripts/verify/check_cv_claims.py # every CV number vs its sourceFull sequence, reference environment, and which numbers vary by machine:
REPRODUCIBILITY.md.
| Topic | Document |
|---|---|
| Scope, users, requirements | REQUIREMENTS.md |
| Corpus and evaluation set, with limitations | DATASET_CARD.md |
| What is measured and how | EVALUATION_PLAN.md |
| Stack choices and rejected alternatives | ARCHITECTURE.md |
| Parsing, chunking, duplicate policy | INGESTION.md |
| Index internals, backend benchmark | INDEXING.md |
| Dense baseline and encoder selection | RETRIEVAL_BASELINE.md |
| BM25, fusion, and where each fails | HYBRID_RETRIEVAL.md |
| The reranking decision | RERANKING.md |
| Grounded generation | GENERATION.md |
| Three prompt designs, two of which failed | PROMPT_DESIGN.md |
| Groundedness, and what the metric cannot prove | GROUNDING_EVALUATION.md |
| API reference | API.md |
| Latency budget and concurrency | LOAD_TESTING.md |
| Threat model and honest posture | SECURITY.md |
| Where the system fails, with real examples | FAILURE_ANALYSIS.md |
| Test inventory and bugs they caught | ROBUSTNESS_TESTS.md |
| Full sequence and determinism | REPRODUCIBILITY.md |
| Complete write-up | FINAL_PROJECT_REPORT.md |
| CV bullets with an evidence table | CV_POINTERS_PRODUCTION_RAG.md |
| Interview questions and answers | INTERVIEW_PREPARATION.md |
Raw data: EVALUATION_RESULTS.csv, LATENCY_RESULTS.csv, results/.
- A thread-unsafe SQLite connection — one connection shared across FastAPI's threadpool corrupted cursor state, failing once in 480 requests. Invisible to 86 single-threaded tests and to every offline evaluation. Found only by an HTTP load test.
- Non-deterministic top-k tie-breaking —
argpartitionpicks an arbitrary member of a tied group, so identical queries could return different results. Caught by pinning BM25 againstrank_bm25. - Orphaned vectors on re-ingest — a changed document left vectors for deleted chunks: retrievable, no longer resolvable to a citation.
- Silently dropped status codes — a heading guard rejected digit-initial titles, so RFC 6585 lost the definitions of 428/429/431/511 entirely.
- float32 precision loss in the BM25 average document length.
It has never served real traffic. It has no authentication and no rate limiting, and no admission control — p99 with generation is ~93 s and requests queue without bound. The Docker image is written but unverified (Docker was unavailable on the development machine). Hallucination is constrained by verification and measured — not eliminated; the grounding check is lexical, not entailment, so a sentence that contradicts its source while reusing its vocabulary would pass.
The evaluation set is 65 queries written by one person. Large enough to separate a 20% effect from noise, not a 3% one — which is why every headline number is reported with a confidence interval, and why the two comparisons that did not clear zero are reported as not clearing zero.
FINAL_PROJECT_REPORT.md §8–9 states the limits in full.
Code: MIT (see LICENSE). The corpus is IETF RFCs, which may be freely
reproduced and distributed under BCP 78; it is fetched at build time, not
redistributed here.
