Skip to content

feat(retrieval): add BGE-M3 sparse document search - #68

Merged
juemimgcd merged 1 commit into
masterfrom
codex/rag-fusion-next
Sep 1, 2026
Merged

juemimgcd merged 1 commit into
masterfrom
codex/rag-fusion-next

Conversation

@juemimgcd

Copy link
Copy Markdown
Owner

Summary

  • add nullable BGE-M3 sparse vectors and a scoped pgvector retrieval path for document chunks
  • generate dense and sparse embeddings in one model pass, then fuse dense/sparse/keyword scores with evaluated weights
  • keep the feature off by default and fall back to the existing dense + keyword path until every active chunk in an owner/KB scope is backfilled

Evaluation

SciFact, 300 queries / 4,000 documents, at k=10:

Metric Baseline Sparse fusion Change
Recall 0.836778 0.846278 +0.009500
MRR 0.658104 0.686276 +0.028172
NDCG 0.696016 0.718469 +0.022453

The selected fusion weights are dense 0.41, BGE-M3 sparse 0.54, and keyword 0.05.

Validation

  • Ruff passed for all changed Python files
  • existing non-integration suite: 345 passed, 2 deselected, 8 subtests passed
  • Docker Compose configuration rendered successfully
  • real BGE-M3 sparse-head output matched the reference score
  • PostgreSQL/pgvector smoke test covered sparsevec(250002), HNSW sparsevec_ip_ops, insert, and inner-product retrieval

Rollout and limits

  • apply the Alembic migration, enable MEMORY_AGENT_EMBEDDING_SPARSE_ENABLED, then run the existing document embedding backfill
  • sparse retrieval is skipped for an entire owner/KB scope while any active chunk lacks a sparse vector
  • the offline benchmark used TF-IDF as the PostgreSQL keyword proxy; live PostgreSQL retrieval p95 still needs production-like measurement before enabling broadly

@juemimgcd
juemimgcd merged commit 6b150dc into master Sep 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant