Skip to content

Latest commit

 

History

History
343 lines (256 loc) · 15.1 KB

File metadata and controls

343 lines (256 loc) · 15.1 KB

Testing guide

This guide explains how the Hash Memory Engine (HME) is tested and evaluated, how to run the evaluation harness, which datasets to use and in what order, how the 12 test phases from the testing plan map onto what exists in the repository today, and the honest limitations of the current setup.

Everything described here runs fully offline by default: the harness uses the deterministic lexical hashing embedder and requires only numpy, blake3, and zstandard. No network access or model download is needed to reproduce a result.


1. Testing objective

The single question the evaluation is designed to answer is:

Can the Hash Memory Engine reduce storage versus a standard vector-RAG pipeline without significantly reducing retrieval quality or speed?

To answer it fairly, the same corpus and query set are pushed through four retrieval systems that differ only in how vectors are represented and whether content-addressed dedup and compression are applied. System A (float32) is the quality and storage baseline; the others are compared against it.

The four systems (A–D)

System Label Vector representation Retrieval mode Dedup Compress Role
A Float32 baseline float32 exact cosine (mode="float") no no The standard-RAG storage/quality baseline.
B Int8 / quantised int8 (scalar) exact cosine (mode="float") yes yes Aggressive scalar quantisation trade-off.
C Binary only binary codes Hamming only (mode="binary") yes yes Smallest, fastest, lowest-fidelity stage alone.
D Full HME binary + float binary→float rerank (mode="binary_rerank") yes yes The proposed end-to-end system as shipped.

Pipelines:

  • A — Float32 baseline (standard RAG). Store every chunk's text and a full-precision float32 vector; retrieve by exact float cosine. Emulates a conventional vector-RAG store that keeps duplicates and does not compress.
  • B — Int8 / quantised. Content-addressed dedup + Zstandard-compressed text, with the float index modelled at int8 (one byte per dimension instead of four). Retrieval still uses exact cosine.
  • C — Binary only. Dedup + compression, retrieval over packed binary codes via Hamming distance (XOR + popcount). Smallest index, fastest, lowest fidelity.
  • D — Full HME. Dedup + compression + the two-stage binary_rerank path: a cheap binary shortlist reranked by exact float cosine. This is the engine default and the system the headline claims are made about.

Note on the A–D systems vs. the A–F configs in benchmark-methodology.md: the A–F table is the broader methodology (adding float16 and binary-only variants). The A–D systems here are the four the evaluation harness (scripts/evaluate.py) actually reports side by side against the plan's Results Table.


2. How to run

The evaluation harness is driven by scripts/evaluate.py. It ingests a dataset once into a single engine (with dedup + compression on), then runs each system in its retrieval mode and reports metrics, latency, and storage.

# Bundled tiny dataset (datasets/tiny) — build/debug the pipeline, seconds to run
python scripts/evaluate.py --dataset tiny

# A real BEIR dataset, downloading it first if missing
python scripts/evaluate.py --dataset scifact --download

# Any BEIR-format folder you already have on disk
python scripts/evaluate.py --dataset /path/to/beir/folder

Useful flags:

Flag Effect
--download Fetch the dataset from the UKP mirror if it is not present locally.
--sweep Also run the duplicate-ratio sweep (D0–D80) and the reliability check.
--max-docs N Truncate the corpus to N documents (keeps large datasets fast).
--max-queries N Truncate the query set to N queries.
--backend hashing Embedding backend (default hashing, fully offline).
--top-k 10 Number of results per query used for the metrics.
--out results/NAME.json Where to write the JSON report.

When --max-docs / --max-queries truncate the data, the script prints a note so the reported numbers are never silently partial.

Downloading datasets separately:

# Thin CLI over download_beir(): fetch one or more BEIR datasets into datasets/
python scripts/download_datasets.py scifact nfcorpus quora

By default everything runs offline with the lexical hashing embedder. For real semantic quality, install the embeddings extra (pip install -e ".[embeddings]") and pass a sentence-transformers backend; see Limitations.

Outputs: scripts/evaluate.py writes a machine-readable results/<dataset>.json and a human-readable results/<dataset>.md. See docs/results.md for the report template those files populate.


3. Datasets

BEIR format

The harness reads the standard BEIR file layout (the format, not the beir pip package — we re-implement it in src/hme/eval/beir_loader.py so evaluation stays offline and torch-free):

<folder>/corpus.jsonl        lines: {"_id", "title", "text", ...}
<folder>/queries.jsonl       lines: {"_id", "text", ...}
<folder>/qrels/<split>.tsv   header 'query-id\tcorpus-id\tscore' then rows

In memory this becomes:

corpus  : {doc_id: {"title": str, "text": str}}
queries : {query_id: text}
qrels   : {query_id: {doc_id: relevance_int}}      # score > 0 == relevant

Following BEIR convention, load_beir filters the query set down to those query ids that have qrels in the loaded split.

The bundled tiny dataset

datasets/tiny (written by hme.eval.datasets.write_tiny_dataset) is a 4-doc corpus with 3 queries and 3 qrels. doc-4 is a byte-identical duplicate of doc-1, so it exercises exact content-addressed dedup: after ingest there are 3 unique chunks and exactly one cross-doc collision. It is the first thing to run when building or debugging the pipeline.

Recommended dataset progression

Work up this ladder — start tiny to prove the pipeline, then move to progressively larger and harder real datasets:

# Dataset Why / what it tests Download
1 tiny (custom) Build and debug the pipeline; known ground truth. Bundled (datasets/tiny).
2 SciFact Full pipeline on a small real dataset. https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/scifact.zip
3 NFCorpus Another domain (medical/nutrition). https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/nfcorpus.zip
4 Quora Semantic duplicates / paraphrases. https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/quora.zip
5 ArguAna Difficult semantic distinctions (counter-arguments). https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/arguana.zip
6 TREC-COVID Medium scale. https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/trec-covid.zip
7 MS MARCO Million-scale stress test. Via BEIR (large; see UKP mirror / BEIR docs).
8 Natural Questions End-to-end QA. Via BEIR (see UKP mirror / BEIR docs).

The known short names (scifact, nfcorpus, fiqa, quora, arguana, trec-covid) are in BEIR_URLS; you can also pass a full .zip URL to download_beir. MS MARCO and Natural Questions are large and are treated as scale/QA stretch targets (see Phase 7 below).

Adding a dataset

Point --dataset at any folder in the BEIR layout above, or add its short name and zip URL to BEIR_URLS in src/hme/eval/beir_loader.py.


4. The 12 test phases

The testing plan defines 12 phases. Below is each phase mapped to what actually exists in the repository, marked honestly as implemented or planned.

Phase 1 — Component tests (implemented)

Unit tests over the core building blocks: hashing/content addressing (tests/test_hashing.py), exact deduplication (tests/test_deduplication.py), Zstandard compression round-trip (tests/test_compression.py), and the SQLite metadata store. These verify the primitives the storage claims rest on.

Phase 2 — Semantic retrieval / similarity (implemented, fixtures)

hme.eval.datasets.semantic_pairs() supplies labelled similar/different sentence groups, and make_near_duplicates() produces surface-level paraphrase variants of a sentence. These let you check that a semantic embedder scores paraphrases close and unrelated sentences far apart. With the lexical hashing embedder this signal is weak (see limitations).

Phase 3 — Retrieval quality (implemented)

Standard IR metrics in hme.eval.metrics: Recall@1 / Recall@5 / Recall@10, MRR, NDCG@10, and Precision. Computed per system so quality can be compared against the float32 baseline (System A).

Phase 4 — Storage testing (implemented)

runner.storage_model computes per-component storage (text bytes, index bytes, metadata bytes, total) for each system, modelling a standard RAG that also stores duplicates. runner.duplicate_sweep runs the D0/D10/D30/D50/D80 duplicate-ratio sweep to demonstrate the dedup lever directly. This is the headline storage-reduction evidence.

Phase 5 — Speed (implemented)

scripts/evaluate.py records per-query latency percentiles (p50/p95/p99) and queries-per-second (qps) for each system, alongside ingest timing.

Phase 6 — Memory (implemented)

hme.eval.environment.rss_bytes() captures process resident-set-size so RAM usage can be reported next to storage.

Phase 7 — Scalability (partly implemented)

--max-docs sweeps let you observe behaviour as the corpus grows. The indexes are numpy brute-force, which is exact and adequate to roughly ~100k chunks; the 1M-chunk target is documented as future work (an ANN backend such as faiss-cpu is the planned acceleration). MS MARCO / NQ scale runs are stretch targets.

Phase 8 — End-to-end LLM (implemented, offline)

engine.answer produces grounded, cited answers. The default EchoLLM is an offline extractive answerer, so the end-to-end path runs with no network. The harness also checks missing-answer behaviour (the engine should decline / not fabricate when no relevant chunk is retrieved). A hosted OpenAI-compatible LLM is available but not required.

Phase 9 — Reliability (implemented)

runner.reliability_check saves an engine to disk, reopens it via HashMemoryEngine.open, and re-runs queries to assert the top results are identical (restart test). It also verifies compression integrity: for a sample of chunks it decompresses via the metadata store and recomputes the BLAKE3 of the normalised text, asserting it equals the stored chunk_id.

Phase 10 — Concurrency (planned)

Concurrent-reader / concurrent-writer stress testing is described in the plan but not yet automated. The engine is currently single-node (one SQLite store plus numpy index files) with no concurrent-writer coordination.

Phase 11 — Security / integrity (partly implemented)

BLAKE3 verify-on-read gives tamper-evident integrity: a corrupted stored chunk will not recompute to its content address. Broader malicious-input / adversarial tests are planned.

Phase 12 — Regression golden (implemented)

tests/test_regression_golden.py pins permanent golden thresholds on datasets/tiny (for example, Recall@1 == 1.0 in float mode, and exactly one duplicate collapsed by dedup). This guards against silent regressions.

Phase status summary

Phase Area Status
1 Component tests Implemented
2 Semantic retrieval / similarity fixtures Implemented
3 Retrieval quality metrics Implemented
4 Storage testing + duplicate sweep Implemented
5 Speed (latency p50/p95/p99, qps) Implemented
6 Memory (RSS) Implemented
7 Scalability Partial (~100k ok; 1M future work)
8 End-to-end LLM (EchoLLM) Implemented (offline)
9 Reliability (restart + integrity) Implemented
10 Concurrency Planned
11 Security / integrity Partial (BLAKE3 done; malicious-input planned)
12 Regression golden Implemented

5. Metric definitions

  • Recall@k — fraction of a query's relevant documents that appear in the top-k results: (# relevant retrieved in top k) / (# relevant), averaged over queries that have at least one relevant doc.
  • MRR (Mean Reciprocal Rank) — average of 1 / rank of the first relevant document per query (0 if none retrieved); rewards ranking the right answer high.
  • NDCG@k (Normalised Discounted Cumulative Gain) — rank-weighted quality with graded relevance, normalised by the ideal ranking; 1.0 is a perfect ordering.
  • Precision@k — fraction of the top-k results that are relevant: (# relevant in top k) / k.

All metrics skip queries with no relevant document and break score ties deterministically by doc id, so results are reproducible across machines.


6. Duplicate-heavy and near-duplicate testing

Exact-duplicate sweep (D0/D10/D30/D50/D80). add_duplicates injects byte-identical copies of existing documents so that roughly 0 %, 10 %, 30 %, 50 %, and 80 % of the resulting corpus is duplicated content. duplicate_sweep ingests each fraction into a fresh engine and records unique chunks, dedup ratio, stored bytes, and storage reduction versus the no-dedup size. Storage reduction should climb with the duplicate fraction while retrieval quality is unaffected (all copies map to a single stored chunk).

Near-duplicate testing. make_near_duplicates produces paraphrase variants (casing, punctuation, filler words, reordering) of a sentence, used with the semantic_pairs fixtures.

Important caveat. Content-addressed dedup is exact: it collapses byte-identical documents only. Near-duplicates with different bytes are not collapsed by hashing — matching them requires semantic or binary similarity, which is exactly what the binary-code retrieval path (Systems C and D) is meant to capture. Do not expect the dedup counter to shrink for paraphrases; that is by design.


7. Honest limitations

  • The offline HashingEmbedder is a lexical, feature-hashing fallback. It is deterministic and dependency-free, but not semantically strong. Absolute Recall / NDCG on semantic datasets (e.g. SciFact) will be modest. Install the embeddings extra and use a sentence-transformers backend for real semantic quality. The storage-reduction and exact-dedup results are the headline wins and do not depend on the embedder's semantic strength.
  • The indexes are numpy brute-force. Exact and fine to roughly ~100k chunks; the 1M-chunk target is future work, with faiss-cpu as the planned ANN acceleration.
  • Concurrency (Phase 10) and parts of security (Phase 11) are planned, not automated. The engine is single-node today; malicious-input testing is not yet in the suite (BLAKE3 verify-on-read integrity is).

8. Running the unit tests

The offline pytest suite lives in tests/ and covers Phases 1–4, 9, and 12:

pytest -q

Key files: test_metrics.py, test_beir_loader.py, test_eval_datasets.py, test_eval_runner.py, test_reliability.py, and the golden test_regression_golden.py.