Compass is a high-performance Agentic Desktop Search Engine that unifies ultra-fast keyword indexing with dense vector semantic search. Running entirely on your local machine, Compass dynamically routes queries between retrieval engines to minimize latency and compute resources while personalizing ranking using user interaction patterns (frequency and recency decay).
- β‘ Sub-Millisecond Keyword Retrieval: SQLite FTS5 index with normalized BM25 term weighting.
- π§ Local Semantic Vector Search: Embeddings powered by
sentence-transformers/all-MiniLM-L6-v2stored in local ChromaDB collections. - π¦ Dynamic Query Classification & Routing: Classifies search intent (file extensions, short tokens, conceptual markers) and selects the cheapest sufficient retrieval path, saving ~34% compute latency.
- π€ Adaptive Frequency-Recency Personalization: Uses an Ebbinghaus-style exponential decay forgetting curve (
$S_{rec} = e^{-\lambda t}$ ) and interaction frequency ($S_{freq}$ ) to boost your most relevant files. - π» Flexible Desktop UI: Native window container via PyWebView or standalone browser app mode via Microsoft Edge/Chrome.
- π Zero Data Leakage: 100% offline, local-first architecture. No telemetry, no cloud dependencies, no API keys needed.
Compass implements an Agentic Perception-Deliberation-Action Loop for desktop files:
graph TD
A[User Search Query] --> B{Dynamic Router}
B -- "File Ext / Short Token" --> C[SQLite FTS5 Keyword Search]
B -- "Conceptual / Natural Lang" --> D[ChromaDB Vector Search]
B -- "Ambiguous Intent" --> E[Parallel Hybrid Search]
E --> C
E --> D
C --> F[Score Normalizer & Blender]
D --> F
F --> G{Personalization Engine}
G -- "History Found" --> H[Blended Rank: 0.85*Base + 0.15*Personal]
G -- "Cold Start" --> I[Discounted Discoverability: 0.85*Base]
H --> J[Final Sorted Results View]
I --> J
| Component | Technology | Role |
|---|---|---|
| Storage & FTS5 | SQLite 3 (backend/models/db.py) |
Relational metadata, file hashes, access logs, and FTS5 full-text indexing |
| Vector Engine | ChromaDB (backend/search/semantic_search.py) |
Dense vector storage with cosine similarity matching |
| Embedding Model |
all-MiniLM-L6-v2 (Sentence-Transformers) |
Lightweight ~90MB local model optimized for CPU inference |
| Query Router | Dynamic Classifier (backend/router/router.py) |
Intent classification, latency tracking, and hybrid score fusion |
| Personalization | Exponential Decay (backend/search/personalize.py) |
Tracks access frequency and recency with 10-day half-life decay ( |
| API & Server | FastAPI + Uvicorn (backend/main.py) |
Async REST backend serving search endpoints and native file launchers |
| Desktop Shell | PyWebView / Web App (app.py, frontend/) |
Responsive dark-mode interface with live score breakdowns |
To eliminate wasteful neural inference on simple lookups:
- Keyword Route: Triggered when the query contains an explicit file extension (e.g.
report.pdf) or has 2 words or fewer. - Semantic Route: Triggered when query contains conceptual patterns (e.g. "something about python", "i remember notes on...").
- Hybrid Route: Triggered for ambiguous queries, merging FTS5 and ChromaDB in parallel.
Hybrid search computes a weighted linear combination of normalized keyword score
To surface documents you frequently and recently interact with, Compass computes:
Where
(Under cold-start conditions with no prior interactions, files receive $0.85 \cdot S_{base}$ to preserve discoverability).
Compass includes a comprehensive evaluation framework (eval/run_eval.py) tested across structured queries:
| Strategy Config | Precision@1 | Precision@3 | Recall@3 | MRR | Mean Latency | p95 Latency |
|---|---|---|---|---|---|---|
| Keyword-Only (FTS5) | 0.467 | 0.156 | 0.467 | 0.467 | 3.22 ms | 4.61 ms |
| Semantic-Only (ChromaDB) | 0.967 | 0.333 | 1.000 | 0.983 | 18.10 ms | 22.11 ms |
| Hybrid-Only (Fused) | 0.967 | 0.333 | 1.000 | 0.983 | 19.30 ms | 22.47 ms |
| Dynamic Router (Compass) | 0.967 | 0.333 | 1.000 | 0.983 | 12.76 ms | 19.97 ms |
- Routing Accuracy: 100.0% (matches optimal retrieval MRR on all queries).
- Compute Latency Reduction: 33.87% faster than always-on hybrid search (12.76 ms vs 19.30 ms).
- Zero Query Failures: Sub-optimal path selection rate is 0.0% across the benchmark suite.
- Python 3.10, 3.11, or 3.12
- Git
# Clone the repository
git clone https://github.com/Tejas-h-blitz/Compass.git
cd Compass
# Create and activate virtual environment
python -m venv .venv
.\.venv\Scripts\Activate.ps1 # On Windows PowerShell
# source .venv/bin/activate # On Linux / macOS
# Install required dependencies
pip install -r requirements.txt| Mode | Command | Description |
|---|---|---|
| Desktop Native | python app.py |
Launches Compass inside a native PyWebView desktop container |
| Standalone App | python app.py --app |
Launches in standalone browser app mode (frameless, lightweight) |
| Web Browser | python app.py --browser |
Starts the FastAPI backend and opens your default web browser |
Execute all validation steps sequentially (schema, FTS5 keyword indexing, vector search, router logic, personalization boosts):
python tests/run_all_tests.pytests/test_step1.py: Database schema initialization & directory scanner caching.tests/test_step2.py: SQLite FTS5 BM25 keyword retrieval.tests/test_step3.py: ChromaDB dense embedding ingestion & cosine vector similarity.tests/test_step4.py: Dynamic routing decision rules & audit logging.tests/test_step5.py: Ebbinghaus decay, frequency scores, & ablation toggle.
Run the 30-query retrieval benchmark measuring P@1, P@3, R@3, MRR, and router latency savings:
python eval/run_eval.pyCompass/
βββ .github/
β βββ workflows/
β βββ ci.yml # GitHub Actions CI automated test & benchmark pipeline
βββ backend/
β βββ main.py # FastAPI server & REST endpoints
β βββ models/
β β βββ db.py # SQLite schema, FTS5 virtual tables, access/query logs
β βββ router/
β β βββ router.py # Dynamic classification, latency timer, hybrid fusion
β βββ scanner/
β β βββ scanner.py # File crawler, text/PDF/DOCX extraction, SHA-256 hash check
β β βββ setup_test_corpus.py # Test document generator
β βββ search/
β βββ keyword_search.py # SQLite FTS5 query builder & BM25 normalizer
β βββ semantic_search.py # ChromaDB client & sentence-transformers vector inference
β βββ personalize.py # Frequency-recency scoring & exponential decay math
βββ data/
β βββ test_corpus/ # Reference test documents (.txt, .pdf, .docx)
βββ eval/
β βββ run_eval.py # Quantitative benchmarking runner (P@k, R@k, MRR, latency)
β βββ test_queries.json # Annotated evaluation query corpus
βββ frontend/
β βββ app.js # Client-side UI logic, live search, debouncing, score badges
β βββ index.html # Clean dark-mode desktop interface
β βββ style.css # Polished glassmorphism styles & animations
βββ tests/
β βββ run_all_tests.py # Unified test suite runner
β βββ test_step1.py # Step 1: DB & Scanner tests
β βββ test_step2.py # Step 2: FTS5 search tests
β βββ test_step3.py # Step 3: Semantic search tests
β βββ test_step4.py # Step 4: Router & logging tests
β βββ test_step5.py # Step 5: Personalization & ablation tests
βββ app.py # Main application launcher (PyWebView / Browser)
βββ requirements.txt # Python dependencies
βββ README.md # Project documentation
Distributed under the MIT License. See LICENSE for more information.