This guide explains how the Hash Memory Engine (HME) is tested and evaluated, how to run the evaluation harness, which datasets to use and in what order, how the 12 test phases from the testing plan map onto what exists in the repository today, and the honest limitations of the current setup.
Everything described here runs fully offline by default: the harness uses
the deterministic lexical hashing embedder and requires only numpy,
blake3, and zstandard. No network access or model download is needed to
reproduce a result.
The single question the evaluation is designed to answer is:
Can the Hash Memory Engine reduce storage versus a standard vector-RAG pipeline without significantly reducing retrieval quality or speed?
To answer it fairly, the same corpus and query set are pushed through four retrieval systems that differ only in how vectors are represented and whether content-addressed dedup and compression are applied. System A (float32) is the quality and storage baseline; the others are compared against it.
| System | Label | Vector representation | Retrieval mode | Dedup | Compress | Role |
|---|---|---|---|---|---|---|
| A | Float32 baseline | float32 | exact cosine (mode="float") |
no | no | The standard-RAG storage/quality baseline. |
| B | Int8 / quantised | int8 (scalar) | exact cosine (mode="float") |
yes | yes | Aggressive scalar quantisation trade-off. |
| C | Binary only | binary codes | Hamming only (mode="binary") |
yes | yes | Smallest, fastest, lowest-fidelity stage alone. |
| D | Full HME | binary + float | binary→float rerank (mode="binary_rerank") |
yes | yes | The proposed end-to-end system as shipped. |
Pipelines:
- A — Float32 baseline (standard RAG). Store every chunk's text and a full-precision float32 vector; retrieve by exact float cosine. Emulates a conventional vector-RAG store that keeps duplicates and does not compress.
- B — Int8 / quantised. Content-addressed dedup + Zstandard-compressed text, with the float index modelled at int8 (one byte per dimension instead of four). Retrieval still uses exact cosine.
- C — Binary only. Dedup + compression, retrieval over packed binary codes via Hamming distance (XOR + popcount). Smallest index, fastest, lowest fidelity.
- D — Full HME. Dedup + compression + the two-stage
binary_rerankpath: a cheap binary shortlist reranked by exact float cosine. This is the engine default and the system the headline claims are made about.
Note on the A–D systems vs. the A–F configs in
benchmark-methodology.md: the A–F table is the broader methodology (adding float16 and binary-only variants). The A–D systems here are the four the evaluation harness (scripts/evaluate.py) actually reports side by side against the plan's Results Table.
The evaluation harness is driven by scripts/evaluate.py. It ingests a dataset
once into a single engine (with dedup + compression on), then runs each system
in its retrieval mode and reports metrics, latency, and storage.
# Bundled tiny dataset (datasets/tiny) — build/debug the pipeline, seconds to run
python scripts/evaluate.py --dataset tiny
# A real BEIR dataset, downloading it first if missing
python scripts/evaluate.py --dataset scifact --download
# Any BEIR-format folder you already have on disk
python scripts/evaluate.py --dataset /path/to/beir/folderUseful flags:
| Flag | Effect |
|---|---|
--download |
Fetch the dataset from the UKP mirror if it is not present locally. |
--sweep |
Also run the duplicate-ratio sweep (D0–D80) and the reliability check. |
--max-docs N |
Truncate the corpus to N documents (keeps large datasets fast). |
--max-queries N |
Truncate the query set to N queries. |
--backend hashing |
Embedding backend (default hashing, fully offline). |
--top-k 10 |
Number of results per query used for the metrics. |
--out results/NAME.json |
Where to write the JSON report. |
When --max-docs / --max-queries truncate the data, the script prints a note
so the reported numbers are never silently partial.
Downloading datasets separately:
# Thin CLI over download_beir(): fetch one or more BEIR datasets into datasets/
python scripts/download_datasets.py scifact nfcorpus quoraBy default everything runs offline with the lexical hashing embedder. For
real semantic quality, install the embeddings extra
(pip install -e ".[embeddings]") and pass a sentence-transformers backend; see
Limitations.
Outputs: scripts/evaluate.py writes a machine-readable results/<dataset>.json
and a human-readable results/<dataset>.md. See
docs/results.md for the report template those files populate.
The harness reads the standard BEIR file layout (the format, not the beir
pip package — we re-implement it in src/hme/eval/beir_loader.py so evaluation
stays offline and torch-free):
<folder>/corpus.jsonl lines: {"_id", "title", "text", ...}
<folder>/queries.jsonl lines: {"_id", "text", ...}
<folder>/qrels/<split>.tsv header 'query-id\tcorpus-id\tscore' then rows
In memory this becomes:
corpus : {doc_id: {"title": str, "text": str}}
queries : {query_id: text}
qrels : {query_id: {doc_id: relevance_int}} # score > 0 == relevant
Following BEIR convention, load_beir filters the query set down to those query
ids that have qrels in the loaded split.
datasets/tiny (written by hme.eval.datasets.write_tiny_dataset) is a 4-doc
corpus with 3 queries and 3 qrels. doc-4 is a byte-identical duplicate of
doc-1, so it exercises exact content-addressed dedup: after ingest there are
3 unique chunks and exactly one cross-doc collision. It is the first thing to
run when building or debugging the pipeline.
Work up this ladder — start tiny to prove the pipeline, then move to progressively larger and harder real datasets:
| # | Dataset | Why / what it tests | Download |
|---|---|---|---|
| 1 | tiny (custom) | Build and debug the pipeline; known ground truth. | Bundled (datasets/tiny). |
| 2 | SciFact | Full pipeline on a small real dataset. | https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/scifact.zip |
| 3 | NFCorpus | Another domain (medical/nutrition). | https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/nfcorpus.zip |
| 4 | Quora | Semantic duplicates / paraphrases. | https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/quora.zip |
| 5 | ArguAna | Difficult semantic distinctions (counter-arguments). | https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/arguana.zip |
| 6 | TREC-COVID | Medium scale. | https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/trec-covid.zip |
| 7 | MS MARCO | Million-scale stress test. | Via BEIR (large; see UKP mirror / BEIR docs). |
| 8 | Natural Questions | End-to-end QA. | Via BEIR (see UKP mirror / BEIR docs). |
The known short names (scifact, nfcorpus, fiqa, quora, arguana,
trec-covid) are in BEIR_URLS; you can also pass a full .zip URL to
download_beir. MS MARCO and Natural Questions are large and are treated as
scale/QA stretch targets (see Phase 7 below).
Point --dataset at any folder in the BEIR layout above, or add its short name
and zip URL to BEIR_URLS in src/hme/eval/beir_loader.py.
The testing plan defines 12 phases. Below is each phase mapped to what actually exists in the repository, marked honestly as implemented or planned.
Unit tests over the core building blocks: hashing/content addressing
(tests/test_hashing.py), exact deduplication (tests/test_deduplication.py),
Zstandard compression round-trip (tests/test_compression.py), and the SQLite
metadata store. These verify the primitives the storage claims rest on.
hme.eval.datasets.semantic_pairs() supplies labelled similar/different
sentence groups, and make_near_duplicates() produces surface-level paraphrase
variants of a sentence. These let you check that a semantic embedder scores
paraphrases close and unrelated sentences far apart. With the lexical hashing
embedder this signal is weak (see limitations).
Standard IR metrics in hme.eval.metrics: Recall@1 / Recall@5 / Recall@10, MRR,
NDCG@10, and Precision. Computed per system so quality can be compared against
the float32 baseline (System A).
runner.storage_model computes per-component storage (text bytes, index bytes,
metadata bytes, total) for each system, modelling a standard RAG that also
stores duplicates. runner.duplicate_sweep runs the D0/D10/D30/D50/D80
duplicate-ratio sweep to demonstrate the dedup lever directly. This is the
headline storage-reduction evidence.
scripts/evaluate.py records per-query latency percentiles (p50/p95/p99) and
queries-per-second (qps) for each system, alongside ingest timing.
hme.eval.environment.rss_bytes() captures process resident-set-size so RAM
usage can be reported next to storage.
--max-docs sweeps let you observe behaviour as the corpus grows. The indexes
are numpy brute-force, which is exact and adequate to roughly ~100k chunks; the
1M-chunk target is documented as future work (an ANN backend such as faiss-cpu
is the planned acceleration). MS MARCO / NQ scale runs are stretch targets.
engine.answer produces grounded, cited answers. The default EchoLLM is an
offline extractive answerer, so the end-to-end path runs with no network. The
harness also checks missing-answer behaviour (the engine should decline / not
fabricate when no relevant chunk is retrieved). A hosted OpenAI-compatible LLM
is available but not required.
runner.reliability_check saves an engine to disk, reopens it via
HashMemoryEngine.open, and re-runs queries to assert the top results are
identical (restart test). It also verifies compression integrity: for a sample
of chunks it decompresses via the metadata store and recomputes the BLAKE3 of
the normalised text, asserting it equals the stored chunk_id.
Concurrent-reader / concurrent-writer stress testing is described in the plan but not yet automated. The engine is currently single-node (one SQLite store plus numpy index files) with no concurrent-writer coordination.
BLAKE3 verify-on-read gives tamper-evident integrity: a corrupted stored chunk will not recompute to its content address. Broader malicious-input / adversarial tests are planned.
tests/test_regression_golden.py pins permanent golden thresholds on
datasets/tiny (for example, Recall@1 == 1.0 in float mode, and exactly one
duplicate collapsed by dedup). This guards against silent regressions.
| Phase | Area | Status |
|---|---|---|
| 1 | Component tests | Implemented |
| 2 | Semantic retrieval / similarity fixtures | Implemented |
| 3 | Retrieval quality metrics | Implemented |
| 4 | Storage testing + duplicate sweep | Implemented |
| 5 | Speed (latency p50/p95/p99, qps) | Implemented |
| 6 | Memory (RSS) | Implemented |
| 7 | Scalability | Partial (~100k ok; 1M future work) |
| 8 | End-to-end LLM (EchoLLM) | Implemented (offline) |
| 9 | Reliability (restart + integrity) | Implemented |
| 10 | Concurrency | Planned |
| 11 | Security / integrity | Partial (BLAKE3 done; malicious-input planned) |
| 12 | Regression golden | Implemented |
- Recall@k — fraction of a query's relevant documents that appear in the
top-k results:
(# relevant retrieved in top k) / (# relevant), averaged over queries that have at least one relevant doc. - MRR (Mean Reciprocal Rank) — average of
1 / rankof the first relevant document per query (0 if none retrieved); rewards ranking the right answer high. - NDCG@k (Normalised Discounted Cumulative Gain) — rank-weighted quality with graded relevance, normalised by the ideal ranking; 1.0 is a perfect ordering.
- Precision@k — fraction of the top-k results that are relevant:
(# relevant in top k) / k.
All metrics skip queries with no relevant document and break score ties deterministically by doc id, so results are reproducible across machines.
Exact-duplicate sweep (D0/D10/D30/D50/D80). add_duplicates injects
byte-identical copies of existing documents so that roughly 0 %, 10 %, 30 %,
50 %, and 80 % of the resulting corpus is duplicated content. duplicate_sweep
ingests each fraction into a fresh engine and records unique chunks, dedup
ratio, stored bytes, and storage reduction versus the no-dedup size. Storage
reduction should climb with the duplicate fraction while retrieval quality is
unaffected (all copies map to a single stored chunk).
Near-duplicate testing. make_near_duplicates produces paraphrase variants
(casing, punctuation, filler words, reordering) of a sentence, used with the
semantic_pairs fixtures.
Important caveat. Content-addressed dedup is exact: it collapses byte-identical documents only. Near-duplicates with different bytes are not collapsed by hashing — matching them requires semantic or binary similarity, which is exactly what the binary-code retrieval path (Systems C and D) is meant to capture. Do not expect the dedup counter to shrink for paraphrases; that is by design.
- The offline
HashingEmbedderis a lexical, feature-hashing fallback. It is deterministic and dependency-free, but not semantically strong. Absolute Recall / NDCG on semantic datasets (e.g. SciFact) will be modest. Install theembeddingsextra and use a sentence-transformers backend for real semantic quality. The storage-reduction and exact-dedup results are the headline wins and do not depend on the embedder's semantic strength. - The indexes are numpy brute-force. Exact and fine to roughly ~100k chunks;
the 1M-chunk target is future work, with
faiss-cpuas the planned ANN acceleration. - Concurrency (Phase 10) and parts of security (Phase 11) are planned, not automated. The engine is single-node today; malicious-input testing is not yet in the suite (BLAKE3 verify-on-read integrity is).
The offline pytest suite lives in tests/ and covers Phases 1–4, 9, and 12:
pytest -qKey files: test_metrics.py, test_beir_loader.py, test_eval_datasets.py,
test_eval_runner.py, test_reliability.py, and the golden
test_regression_golden.py.