diff --git a/README.md b/README.md
index 658ab49..2347cf6 100644
--- a/README.md
+++ b/README.md
@@ -1,88 +1,53 @@
-

+

# Kimetsu
### Give your coding agent a memory that gets sharper every run.
-**Evidence-first memory for coding agents and Kimetsu's own terminal chat.**
-Kimetsu sits beside your AI agent, watches what actually solves problems,
-remembers it, and feeds the high-signal context back — so the next run
-starts where the last one left off.
+*Kimetsu* (鬼滅), "demon slayer." It slays the demon every agent fights: amnesia.
[](https://crates.io/crates/kimetsu-cli)
[](#license)
[](https://www.rust-lang.org)
-

+
---
-## At a glance
-
-**Measured, not claimed** *(every number is reproducible with `kimetsu brain bench` / the ROI ledger)*:
-
-| Metric | Result | How it's measured |
-|--------|--------|-------------------|
-| **Cost per win** | **$0.19 vs $2.47** (~13× cheaper) | 16-task Terminal-Bench slice, Kimetsu-on vs no-brain baseline |
-| **Retrieval quality** | **recall@4 0.949 · MRR 0.914 @ ~138 ms** (default), up to **0.975 · 0.933** | `kimetsu brain bench`, 100-memory / 210-case dataset, jina-v2-base-code × cross-encoder rerank |
-| **Scale** | ~1M memories in ~3 GB RAM, sub-2s retrieval | usearch HNSW ANN, O(log N) |
-| **Footprint** | one SQLite file per project, **no cloud, no telemetry** | back it up with `cp` |
-
-**Configurable surface** *(all in `.kimetsu/project.toml`, all optional with safe defaults)*:
-
-| Section | Controls | Default |
-|---------|----------|---------|
-| `embedder` | semantic model + cross-encoder reranker (or FTS-only lean build) | jina-v2-base-code × ms-marco-tinybert |
-| `[cheap_model]` | local **Ollama** or cloud model for digest / resume / `ask` / skill-draft — **optional**, degrades gracefully | off (rule-based) |
-| `[storage] backend` | `flat` (FTS + ANN) · `graph-lite` (+ typed-edge hops) · `graph` (remote) | `flat` |
-| `[broker]` | `warm_start`, `compress_capsules`, `session_dedupe`, `answer_grade_min_score`, `proactive_prefetch` | sensible on/off |
-| `[sync]` | server-less multi-machine sync via a shared `dir` | off |
-| `[learning]` | `store_queries` for the self-tuning eval set | on (local only) |
-
-Full reference: **[How Kimetsu Works](docs/HOW-KIMETSU-WORKS.md)** · local models: **[docs/LOCAL-MODELS.md](docs/LOCAL-MODELS.md)**.
-
----
-
## Why Kimetsu
-LLM coding agents are brilliant and forgetful. Every session starts from
-zero — the same wrong turns, the same re-explaining of your conventions,
-the same expensive exploration you already paid for last week.
+LLM coding agents are brilliant and forgetful. Every session starts from zero:
+the same wrong turns, the same re-explaining of your conventions, the same
+expensive exploration you already paid for last week.
-Kimetsu fixes the forgetting. It's a **sidecar brain**: a single Rust binary
-that runs next to any supported host agent through MCP (Claude Code, Codex, Pi,
-OpenClaw, Cursor, Gemini CLI) or as its own terminal chat — or, in beta,
-server-hosted over HTTP MCP and shared across a team. It learns which memories
-the model *actually used to win*, and lets that knowledge compound across runs.
+Kimetsu fixes the forgetting. It's a sidecar brain, a single Rust binary that
+runs next to your host agent over MCP (Claude Code, Codex, Pi, OpenClaw, Cursor,
+Gemini CLI) or as its own terminal chat. It learns which memories the model
+actually used to win, and lets that knowledge compound across runs.
- **It remembers.** Project conventions, failure patterns, the exact command
- that regenerates your schema — captured once, retrieved automatically.
-- **It learns what helps.** Memories that the model cites before solving a
- problem get promoted. Silent passengers and stale advice decay and get pruned.
-- **It never explores twice** *(v2.0)*. A session-start digest plus an
- **episodic resume** mean the agent's first turn already knows the repo *and*
- what you were doing last time — no re-deriving the basics, no "where was I."
-- **It answers, not just injects** *(v2.0)*. `kimetsu ask "…"` composes a
- grounded, cited answer from memory using a **local model** — zero frontier
- tokens, works offline. Lessons cited often graduate into runnable **skills**.
-- **It's cheap to be right.** On a recorded 16-task Terminal-Bench slice,
- Kimetsu-enabled runs cost **~13x less per win** than the no-brain host-agent
- baseline: $0.19/win vs $2.47/win — and the `kimetsu brain roi` ledger shows
- the token savings per memory on your own work.
+ that regenerates your schema. Captured once, retrieved automatically.
+- **It learns what helps.** Memories the model cites before solving a problem
+ get promoted. Silent passengers and stale advice decay and get pruned.
+- **It never explores twice.** A session-start digest and an episodic resume
+ mean the agent's first turn already knows the repo and what you were doing
+ last time. No re-deriving the basics, no "where was I."
+- **It answers, not just injects.** `kimetsu ask` composes a grounded, cited
+ answer from memory using a local model: zero frontier tokens, works offline.
+ Lessons cited often enough graduate into runnable skills.
+- **It's cheap to be right.** On a recorded 16-task Terminal-Bench slice, runs
+ with Kimetsu cost about 13x less per win than the no-brain baseline ($0.19 vs
+ $2.47), and the ROI ledger shows the token savings on your own work.
- **It gets smarter, not just bigger.** Semantic retrieval finds the right
- memory even when you used different words; it **self-tunes** retrieval against
- your own query history; and brain insights show hit-rate, citation rate, and
- token economy so the value is measurable.
+ memory even when you used different words, and it self-tunes retrieval against
+ your own query history.
- **It's yours, on your machine.** The whole brain is one SQLite file per
project. No external vector DB, no cloud, no telemetry. Back it up with `cp`.
-> *Kimetsu* (鬼滅) — "demon slayer." It slays the demon every agent fights:
-> amnesia.
-
---
## How it works
@@ -95,85 +60,86 @@ the model *actually used to win*, and lets that knowledge compound across runs.
│ scores candidates by relevance ×
│ usefulness × freshness × scope
▼
- brain.db — one SQLite file: FTS5 + semantic ANN (usearch HNSW)
+ brain.db: one SQLite file, FTS5 + semantic ANN (usearch HNSW)
```
-1. **Before a task**, the broker walks your project brain *and* your
+1. **Before a task**, the broker walks your project brain and your
cross-project user brain, scores every candidate, and injects the top few
- inside an adaptive token budget. The semantic build matches by *meaning*
- (O(log N) ANN — scales to ~1M memories in ~3 GB RAM, sub-2s retrieval).
-2. **While it works**, Kimetsu surfaces known pitfalls *before* the first
+ inside an adaptive token budget. The semantic build matches by meaning
+ (O(log N) ANN, scaling to ~1M memories in ~3 GB RAM, sub-2s retrieval).
+2. **While it works**, Kimetsu surfaces known pitfalls before the first
attempt, and the model cites the memories that actually help.
3. **After the task**, cited memories get promoted, unused advice decays on a
half-life curve, and non-trivial sessions auto-harvest their lessons.
-Full mechanics — scoring, citations, decay, conflict detection, the daemon —
-in **[HOW-KIMETSU-WORKS](docs/HOW-KIMETSU-WORKS.md)**.
+Full mechanics, scoring, citations, decay, conflict detection, and the daemon
+are in **[How Kimetsu Works](docs/HOW-KIMETSU-WORKS.md)**.
+
+---
+
+## Benchmarks
+
+Every number is reproducible with `kimetsu brain bench` and the
+`kimetsu brain roi` ledger.
+
+| Metric | Result | How it's measured |
+|--------|--------|-------------------|
+| Cost per win | **$0.19 vs $2.47** (~13x cheaper) | 16-task Terminal-Bench slice, Kimetsu vs no-brain baseline |
+| Retrieval quality | **recall@4 0.949, MRR 0.914 at ~138 ms** (default), up to 0.975 / 0.933 | `kimetsu brain bench`, 100-memory / 210-case dataset, jina-v2-base-code + cross-encoder rerank |
+| Scale | ~1M memories in ~3 GB RAM, sub-2s retrieval | usearch HNSW ANN, O(log N) |
+| Footprint | one SQLite file per project, no cloud, no telemetry | back it up with `cp` |
+
+The semantic build retrieves with jina-v2-base-code and a cross-encoder
+reranker, tuned with `kimetsu brain bench` on a 100-memory / 210-case dataset of
+real exported memories. The latency-optimized default (ms-marco-tinybert-l-2-v2)
+lands recall@4 0.949, MRR 0.914 at ~138 ms; the quality-best rerankers reach
+recall@4 0.975, MRR 0.933. Swap embedder and reranker with one config key each
+and re-judge on your own corpus. Full grid in
+**[How Kimetsu Works](docs/HOW-KIMETSU-WORKS.md)**.
---
## Quickstart
```bash
-# 1. Install — no Rust toolchain needed (cargo + prebuilt archives in docs/INSTALL.md)
npm install -g kimetsu-ai
kimetsu npm-flavor embeddings # one-time: enable semantic retrieval
-
-# 2. Wire it into your host agent — init + install + selftest in one shot
cd /your/project
kimetsu setup --host claude-code # or: codex | openclaw | pi
-
-# 3. Prove the brain works
-kimetsu doctor --selftest
-# ✓ recorded a memory and retrieved it — the brain works
-```
-
-From here your agent banks memories automatically — record one yourself and
-watch it come back:
-
-```bash
-kimetsu brain memory add --scope project --kind convention "Use cargo nextest for all test runs"
-kimetsu brain context "how do I run tests?" # broker-ranked context bundle
-kimetsu ask "how do I run tests?" # v2.0: grounded, cited answer (local model)
-kimetsu resume # v2.0: pick up where the last session left off
-kimetsu brain insights # is the brain actually helping?
-kimetsu brain roi --top # did it pay for itself? per-memory token savings
-kimetsu brain tune --status # self-tune retrieval; --status shows when to re-run
-kimetsu brain skills --review # v2.0: turn often-cited lessons into runnable skills
+kimetsu doctor --selftest # records a memory and retrieves it
```
-New in v2.0 — never explore twice:
-
-```bash
-kimetsu checkpoint "mid-task note" # save working state now (auto-saved at session end)
-kimetsu brain digest --refresh # rebuild the session-start project digest
-kimetsu brain sync # replicate your brain across machines (no server)
-```
-
-Prefer a standalone REPL? `kimetsu chat --workspace . --project .` is a full
-terminal assistant with the same brain. Every install path (npm, prebuilt
-archives), host-wiring details, the auto-harvest/distiller setup, and
-maintenance commands live in **[docs/INSTALL.md](docs/INSTALL.md)**.
+Other install paths (cargo, prebuilt archives) and host-wiring details are in
+**[docs/INSTALL.md](docs/INSTALL.md)**.
---
-## Retrieval quality — benchmarked, not vibes
-
-The semantic build retrieves with **jina-v2-base-code** + a
-**cross-encoder reranker**, chosen with `kimetsu brain bench` on a
-100-memory / 210-case dataset seeded from real exported memories. The
-latency-optimized default (`ms-marco-tinybert-l-2-v2`) lands
-**recall@4 0.949, MRR 0.914 at ~138 ms** per retrieval+rerank; the
-quality-best rerankers reach **recall@4 0.975, MRR 0.933**. Swap embedder and
-reranker with one config key each and re-judge on your own corpus — full grid
-in [Retrieval models & benchmarking](docs/HOW-KIMETSU-WORKS.md).
+## Command reference
+
+| Command | What it does |
+|---------|--------------|
+| `kimetsu setup --host ` | Wire the brain into a host agent (init + install + selftest) |
+| `kimetsu chat` | Standalone terminal coding assistant with the same brain |
+| `kimetsu brain memory add` | Record a durable lesson by hand |
+| `kimetsu brain context ""` | Broker-ranked context bundle for a query |
+| `kimetsu ask ""` | Grounded, cited answer from memory (local model) |
+| `kimetsu resume` / `kimetsu checkpoint` | Pick up where the last session left off |
+| `kimetsu brain skills` | Turn often-cited lessons into runnable skills |
+| `kimetsu brain insights` / `roi` | Is the brain helping, and did it pay for itself |
+| `kimetsu brain tune` | Self-tune retrieval against your own query history |
+| `kimetsu brain sync` | Replicate your brain across machines, no server |
+| `kimetsu brain bench` | Benchmark retrieval on your own corpus |
+
+The full command surface, configuration keys, and maintenance commands are in
+**[How Kimetsu Works](docs/HOW-KIMETSU-WORKS.md)** and
+**[docs/INSTALL.md](docs/INSTALL.md)**.
---
## Kimetsu Remote (beta)
-Share one brain per repository from a server, over HTTP MCP — for a team, or
-for yourself across machines:
+Share one brain per repository from a server over HTTP MCP, for a team or for
+yourself across machines:
```bash
# server
@@ -182,45 +148,40 @@ kimetsu-remote serve --addr 0.0.0.0:8787 --data /srv/kimetsu-brains --token