Skip to content

perf(retrieval): fuse normalized backend scores - #67

Merged
juemimgcd merged 1 commit into
masterfrom
codex/rag-score-fusion
Sep 1, 2026
Merged

juemimgcd merged 1 commit into
masterfrom
codex/rag-score-fusion

Conversation

@juemimgcd

Copy link
Copy Markdown
Owner

Summary

  • carry pgvector cosine similarity and PostgreSQL ts_rank_cd through document retrieval hits
  • min-max normalize each backend candidate list and combine scores with dense weight 0.55
  • expand the production fusion pool to 200 candidates while preserving the final Top-10 contract
  • add the production score-fusion path to the existing BEIR evaluator and update its usage documentation

Evaluation

BEIR SciFact, 300 test queries, 4,000 documents, candidate pool 200:

  • Recall@10: 0.822889 -> 0.836778 (+0.013889)
  • MRR@10: 0.634030 -> 0.658104 (+0.024074)
  • NDCG@10: 0.673534 -> 0.696016 (+0.022482)
  • document embedding cache: 4,000 hits, 0 re-encoded
  • score-fusion computation: 0.238 seconds for the offline benchmark

Validation

  • Ruff: passed for all changed Python files
  • compileall: passed for app and main.py
  • existing scoped retrieval tests: 3 passed
  • full non-integration suite: 345 passed, 2 deselected, 8 subtests passed
  • production fusion behavior and PostgreSQL score-expression compilation: passed
  • git diff --check: passed

No test files were modified.

Residual risk

The offline benchmark uses TF-IDF as a lexical proxy. Increasing each backend candidate pool from 100 to 200 may increase live PostgreSQL latency; production p95 should be measured before rollout.

@juemimgcd
juemimgcd merged commit e68f9a7 into master Sep 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant