Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Production RAG Knowledge System

Grounded question answering over a corpus of IETF RFCs (HTTP / Web API standards). Hybrid retrieval, optional cross-encoder reranking, and answers that either carry resolvable source citations or refuse to answer.

Everything runs locally on open-weights models. No API key, no per-query cost.

Change in Recall@5 with 95% bootstrap intervals

Three effects clear zero; two do not. The two that do not are the ones a RAG project is normally expected to claim.


What this is

A retrieval-augmented QA system built end to end and — the part that matters — measured at every step, with a confidence interval on every comparison and a check that fails if any number quoted on a CV cannot be resolved from a results file.

Corpus 34 IETF RFCs, 2,779 section-bounded chunks
Retrieval BAAI/bge-small-en-v1.5 dense + Okapi BM25 written from scratch, weighted fusion
Reranking ms-marco-MiniLM-L-6-v2 cross-encoder — measured, then disabled
Generation Qwen2.5-1.5B-Instruct emitting structured claims, each verified before it reaches the user
Evaluation 65 hand-built queries (52 answerable) across 8 question types, paired-bootstrap significance testing
Service FastAPI, request IDs, structured JSON logs with redaction, Prometheus metrics, Docker
Tests 86, plus 14 verification checks

What was done

Phase 0 — design. Chose a domain where citation precision matters and where retrieval is genuinely hard: RFCs mix prose, ASCII tables, ABNF grammars and exact identifiers, and the corpus deliberately includes an obsoleted document (RFC 7234, superseded by RFC 9111) so conflicting sources are a real case rather than a hypothetical one. Wrote 65 evaluation queries by hand against the source text, then machine-verified every one: the cited section must exist in the index and every reference keyword must literally appear in it. That check caught 10 queries where my keywords were paraphrases rather than quotations.

Phase 1 — ingestion. Two RFC text layouts (paginated and unpaginated) parsed into section-addressable documents; page furniture stripped; boilerplate excluded; chunks that never cross a section boundary, because a chunk spanning two sections cannot be cited honestly.

Phases 2–4 — retrieval. Swept encoders, chunk sizes, overlaps, fusion methods and fusion weights, then reranking depth — each against a stated baseline, each with a paired bootstrap over 20,000 resamples. No configuration was chosen by intuition.

Phase 5 — generation. Prompting alone failed (see below). Replaced trust with verification: the model proposes structured claims, the system checks each one against the source it cites, and composes the answer only from claims that pass.

Phases 6–7 — service and robustness. FastAPI with validation, error handling, observability and Docker; injection, failure-path, log-hygiene and concurrency testing.

Phase 8 — reporting. Every CV-facing number registered in eval/cv_claims.json and resolved from a results file by an automated check.

Three findings worth reading

1. The encoder dominated every architectural change. Swapping all-MiniLM-L6-v2 for BAAI/bge-small-en-v1.5 improved Recall@5 by 23.3% and MRR by 22.2%, with 5 of 6 metrics significant at 95%. No retrieval trick in this project came close to that.

2. Cross-encoder reranking did not help, and is off by default. Over a weak first stage it lifted Recall@5 by 11.8% (95% CI [+0.003, +0.147]). Over the strong first stage that actually ships, zero of six metrics were distinguishable from noise, while p50 latency rose from 14.7 ms to 113.7 ms — and quality degraded monotonically as the reranker was given more candidates. The component is built, benchmarked and disabled, with the numbers to defend it.

3. Prompting could not make a 1.5B model cite its sources; verification could. Asked for prose with inline citations, it produced zero — copying the RFCs' own bibliography markers ([URI], [RFC9110]) instead. Few-shot examples did not fix it. What fixed it was changing the contract: the model emits structured JSON claims, and every claim is checked against its cited source before the answer is composed — including a rule that a claim may not introduce a rare corpus term its evidence never mentions. Result: 100% citation validity across 65 queries, zero uncited answers, and correct refusal on all 8 questions the corpus cannot answer — including probes the model answers confidently from pretraining.

Headline numbers

Retrieval, 52 answerable queries (Apple M4, MPS):

Configuration Recall@5 MRR nDCG@10 p50
BM25 alone 0.5513 0.5089 0.5113 0.7 ms
Dense MiniLM (baseline) 0.5769 0.5007 0.5322 8.5 ms
Dense BGE-small 0.7115 0.6117 0.6218 12.0 ms
Hybrid (shipped default) 0.6827 0.5912 0.6012 14.7 ms
Hybrid + reranking 0.6859 0.6066 0.6113 113.7 ms

Generation and grounding, all 65 queries:

Metric Value
Citation validity rate 1.0000
Answers with zero citations 0
Correct refusal on unanswerable 1.0000 (8/8)
Citation correctness (answered) 0.7250
Evidence hit rate 0.7885
Uncited assertive-sentence rate 0.0358

Service: 1,000 requests across concurrency 1–16, zero errors, ~50–90 rps retrieval-only.

Quick start

pip install torch && pip install -r requirements-dev.txt

python3 scripts/fetch_corpus.py           # 34 RFCs, SHA-256 verified  (~30 s)
python3 scripts/build_index.py --rebuild  # parse, chunk, embed, index (~55 s)
uvicorn ragkb.api:app --port 8000
curl -s localhost:8000/query -H 'content-type: application/json' \
  -d '{"query":"What does the 429 status code mean?","top_k":5}' | jq

Retrieval-only (no LLM, ~15 ms):

curl -s localhost:8000/query -H 'content-type: application/json' \
  -d '{"query":"cache-control no-store","generate":false}' | jq '.retrieval[].citation'

Docker (CPU-only, slower than the MPS figures quoted here):

docker compose up --build

Reproduce the measurements

The corpus and index are not committed — they are regenerated by the two commands above and pinned by SHA-256 in data/corpus_manifest.json. Everything needed to audit the numbers without running a model is committed: the evaluation set, per-query evaluation records, aggregate tables, significance tests and the claims registry.

python3 -m pytest tests/ -q                         # 86 tests
python3 scripts/evaluate_retrieval.py --suite all   # phases 2-4
python3 scripts/evaluate_retrieval.py --suite final # final configuration
python3 scripts/evaluate_generation.py --tag main   # phase 5, grounding
python3 scripts/benchmark_latency.py                # per-stage latency
python3 scripts/load_test.py --sweep 1,2,4,8,16     # concurrency (service must be up)
python3 scripts/make_plots.py                       # figures, from results/
python3 scripts/verify/check_cv_claims.py           # every CV number vs its source

Full sequence, reference environment, and which numbers vary by machine: REPRODUCIBILITY.md.

Reports

Topic Document
Scope, users, requirements REQUIREMENTS.md
Corpus and evaluation set, with limitations DATASET_CARD.md
What is measured and how EVALUATION_PLAN.md
Stack choices and rejected alternatives ARCHITECTURE.md
Parsing, chunking, duplicate policy INGESTION.md
Index internals, backend benchmark INDEXING.md
Dense baseline and encoder selection RETRIEVAL_BASELINE.md
BM25, fusion, and where each fails HYBRID_RETRIEVAL.md
The reranking decision RERANKING.md
Grounded generation GENERATION.md
Three prompt designs, two of which failed PROMPT_DESIGN.md
Groundedness, and what the metric cannot prove GROUNDING_EVALUATION.md
API reference API.md
Latency budget and concurrency LOAD_TESTING.md
Threat model and honest posture SECURITY.md
Where the system fails, with real examples FAILURE_ANALYSIS.md
Test inventory and bugs they caught ROBUSTNESS_TESTS.md
Full sequence and determinism REPRODUCIBILITY.md
Complete write-up FINAL_PROJECT_REPORT.md
CV bullets with an evidence table CV_POINTERS_PRODUCTION_RAG.md
Interview questions and answers INTERVIEW_PREPARATION.md

Raw data: EVALUATION_RESULTS.csv, LATENCY_RESULTS.csv, results/.

Bugs found by measurement, not inspection

  1. A thread-unsafe SQLite connection — one connection shared across FastAPI's threadpool corrupted cursor state, failing once in 480 requests. Invisible to 86 single-threaded tests and to every offline evaluation. Found only by an HTTP load test.
  2. Non-deterministic top-k tie-breaking — argpartition picks an arbitrary member of a tied group, so identical queries could return different results. Caught by pinning BM25 against rank_bm25.
  3. Orphaned vectors on re-ingest — a changed document left vectors for deleted chunks: retrievable, no longer resolvable to a citation.
  4. Silently dropped status codes — a heading guard rejected digit-initial titles, so RFC 6585 lost the definitions of 428/429/431/511 entirely.
  5. float32 precision loss in the BM25 average document length.

What this is not

It has never served real traffic. It has no authentication and no rate limiting, and no admission control — p99 with generation is ~93 s and requests queue without bound. The Docker image is written but unverified (Docker was unavailable on the development machine). Hallucination is constrained by verification and measured — not eliminated; the grounding check is lexical, not entailment, so a sentence that contradicts its source while reusing its vocabulary would pass.

The evaluation set is 65 queries written by one person. Large enough to separate a 20% effect from noise, not a 3% one — which is why every headline number is reported with a confidence interval, and why the two comparisons that did not clear zero are reported as not clearing zero.

FINAL_PROJECT_REPORT.md §8–9 states the limits in full.

Licence

Code: MIT (see LICENSE). The corpus is IETF RFCs, which may be freely reproduced and distributed under BCP 78; it is fetched at build time, not redistributed here.

About

Grounded RAG over IETF RFCs: hybrid retrieval, verified citations, and a measured decision to disable the reranker. Every CV number resolves to a results file.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages