Sakur4 is a memory and context layer that makes it cheap instead.
Ships five ways in, so a harness needs no particular capability to be reached: an MCP
server for anything that speaks MCP, a native Oh My Pi extension for the harness that
does not, a Hermes ContextEngine that replaces its summariser rather than only exposing
tools, a portable Agent Skill for anything reading ~/.agents/skills/, and an
OpenAI-compatible reverse proxy for a harness with none of those. One daemon behind all
five.
|
What it is A memory and context layer that sits beside a coding agent. It keeps a verbatim record of what happened, finds it again by meaning rather than by grepping, and — the part nothing else does — decides where to compact so the inference server's prompt cache still matches. |
What it is not Not a harness, not a model, not a proxy in the routing sense. It never calls a model and never owns the agent loop. Every harness that uses it keeps owning its own session; Sakur4 advises and records. Remove it and the agent still runs. |
Measured against a real 27B on a real server — not a fixture
| tokens per turn | compactions | facts recalled | pinned constraint |
|---|---|---|---|
| +1.3% | 2 → 4 | 50% → 92% | lost → survived |
205 turns, matched 81,920-token window, window-first profile. Full method and raw numbers in the benchmark.
|
Getting started |
Reference |
An agent harness compacts when the context window fills. It replaces the transcript with a summary and sends the result.
That new token sequence shares no prefix with the old one. llama.cpp's longest-common-prefix slot matching therefore finds nothing, and the entire compacted context is re-prefilled — 100+ seconds for a 50K-token session on consumer hardware.
The operation whose purpose was to make the session cheap becomes the most expensive thing in it. Nothing in that loop is wrong: the harness and the inference server simply do not know about each other.
It gets worse on a hosted provider. There, the prompt cache is billed, so a rewrite does not just cost latency — it costs money for tokens that had already been paid for. Hermes' own documentation calls this "the strongest argument against" per-turn compaction, and notes the trade depends on numbers specific to the user.
Sakur4 knows about both sides, so it can supply those numbers and act on them.
The order of operations is the design:
- Ask the inference server where its KV cache can be rewound to.
- Choose the eviction boundary from those checkpoints.
- Evict after it.
Doing it the other way round — deciding what to evict, then asking the cache — produces a boundary at token 0, which no checkpoint can align to. Every compaction then reports a full re-prefill: the exact failure the project exists to remove, arrived at by its own machinery.
Three versions of that logic were written before one was right, and each wrong version
left the entire test suite green. That is why the claim is stated as executable
contracts in cache_coherence.rs rather
than as a promise in a README.
Requirements: Rust 1.94+ to build from source. No GPU, no model, no network — the embedded backend simulates a llama.cpp checkpoint ring in-process, so everything works anywhere.
| Platform | Command | Notes |
|---|---|---|
| Linux · macOS | curl -fsSL https://raw.githubusercontent.com/sc4rfurry/Sakur4/master/install.sh | sh | Detects your platform, verifies the checksum, installs the binary, and places the Agent Skill in ~/.agents/skills/. An existing skill is left alone unless SAKUR4_FORCE=1. |
| Windows | irm https://raw.githubusercontent.com/sc4rfurry/Sakur4/master/install.ps1 | iex | The same steps for PowerShell, with the same refusals. Installs to %USERPROFILE%\.sakur4\bin and adds it to your user PATH. -DryRun downloads and verifies without installing. |
| Release binary | Download from Releases | Archives carry sakur4d, the skill and both integrations. The installers place the binary and the skill; the plugins are unpacked alongside, because each harness keeps plugins elsewhere. |
| From a checkout | cargo install --path crates/sakur4d | Builds and installs in one step. About eight minutes from cold; verified. |
| From source | cargo build --release | Then copy target/release/sakur4d onto your PATH. |
Five platforms are published — x86_64 and aarch64 for Linux, both for macOS, and
x86_64-pc-windows-msvc. Windows shipped from v0.1.0 with no installer for it: install.sh
refuses Windows by design, because uname there is MINGW64_NT-… rather than Linux or Darwin. This
badge said Windows and the only automated path was a shell script Windows users cannot run, which is a
gap a reader finds only after the download.
The installer refuses to install an archive it cannot verify. If SHA256SUMS.txt is
unreachable, or does not list your platform's archive, it stops rather than continuing — an
installer that silently skips verification teaches people to trust the output of a pipe.
SAKUR4_VERSION=v0.1.0 pins a release; SAKUR4_BIN_DIR chooses where it lands.
Publishing to crates.io, if you want cargo install sakur4d to work: publish
cargo publish -p sakur4-core first. cargo publish resolves a path dependency through the
registry, so the binary crate cannot be published until the core crate is on crates.io —
publishing in the other order fails with no matching package named 'sakur4-core' found. The
full sequence, and what is deliberately not automated, is in docs/RELEASING.md.
Put the binary somewhere on PATH — ~/.cargo/bin is where cargo install puts it and
where every integration looks first. If it lives somewhere unusual, set SAKUR4_BIN and
everything will find it.
# A guided walkthrough: dual-track write, staleness detection, a real compaction,
# the cache verdict, round-trip integrity. Needs nothing but the binary.
sakur4d demo --db :memory:See what it prints — real output, not an illustration
=== 2 · the dual-track discipline ===
committed user turn · 38 tokens · symbolic: not a tool result
constraint detector proposed a pin (rule explicit_never, confidence 0.70):
Important rule for this repository: never force-push to main...
pinned anc_01a096a5... as safety_constraint — now exempt from every eviction tier
committed tool result · 28 tokens · symbolic: 3 JSON field paths
committed unstructured result · symbolic: no structure detected; retained as raw episodic text only
=== 3 · staleness: an interpretation that outlived its source ===
wrote interpretation atlas_01a096a5... (anchored, not stale yet)
...the function is then edited (signature and body both change)
=== RECALLED MEMORY ===
[1] (symbolic_fact · score 0.760)
src::auth::checkUser — fn checkUser(id: UserId) -> Result<User>
[2] (semantic_entry · score 0.175 · STALE)
[STALE SUMMARY — do not trust] checkUser looks a user up by their email address
↳ the anchor it was derived from has changed since this summary was written;
re-read the source or call code.query_symbol to get the current truth.
↳ CURRENT VALUE: src::auth::checkUser — fn checkUser(id: UserId) -> Result<User>
→ 1 hit(s) flagged stale, each carrying its anchor's current value
=== 5 · the eviction decision ===
pressure Compacting
budget 32768 · trigger 24576 · target 18022
live 24847 · anchors 69 · fixed 73
6 episode(s), 7812 tokens reclaimed (24989 → 17177 of 18022 target).
partial reuse — prefix survived compaction
anchor safety: 1 anchor(s) pinned, 0 of them in the eviction set (must be 0)
=== 6 · applying it, and what the cache did ===
cache: partial reuse — prefix survived compaction
Context Ledger Receipt · turn 1 · session demo
window 17114/32768 tokens (52% full) · tokenizer backend-exact · backend embedded
where the budget went:
raw recent history 17041 99.6% ██████████████████
pinned anchors 51 0.3% ··················
system prompt 22 0.1% ··················
cache: partial-reuse
slot retained 4034 tokens at checkpoint 4034 (hash 5acbd3726c22e18e);
4034 tokens reused from the LCP, 13080 prefilled
4034 tokens reused / 13080 prefilled (24% saved)
=== 7 · round-trip integrity (FR-5) ===
recalled 3 evicted episode(s) verbatim — content is unchanged by eviction
That receipt is the same one context.receipt returns over MCP, and that plan is the
same one context.plan_eviction returns. Nothing in the demo is a special path.
Then, against a real repository:
sakur4d index . # build the code graph (incremental; cheap to re-run)
sakur4d repo-map --budget 1500 # structural outline fitted to a token budget
sakur4d doctor # what backend and cache capabilities were detectedTo query a symbol, get its real name first. Qualified names carry their whole path, so they look nothing like a file path:
sakur4d repo-map --budget 600 --names # qualified names instead of signatures
sakur4d impact crates::sakur4-core::src::engine::Engine::openThat example is the name of a function in this repository, so it works if you run these from a
checkout. Against your own project the names will be yours — repo-map --names is how you find
them, and it exists for exactly this reason: impact and query_symbol take qualified names while
the default map shows signatures, so without it there is no way to learn what to pass them.
A name that does not exist is an error rather than an empty result, deliberately — an empty answer would read as "nothing depends on this", which is a much worse thing to be told wrongly:
$ sakur4d impact src::auth::validate
Error: not found: symbol src::auth::validate is not in the Symbolic Ledger;
index the project first (`sakur4d index`) or check the qualified nameThree routes, because harnesses disagree about what they support. All reach the same daemon and the same memory, and you can use more than one.
sakur4d config hermes # ~/.hermes/config.yaml
sakur4d config claude # claude_desktop_config.json
sakur4d config claude-code # one-line CLI registration
sakur4d config generic-http # anything that connects to a URL
sakur4d config generic-stdio # anything that spawns a child processconfig prints ready-to-paste configuration with this binary's absolute path and
store baked in, so there is no placeholder to forget.
sakur4d serve # stdio — the default
sakur4d serve --transport http --bind 127.0.0.1:8765 # shared| Transport | How the harness reaches it | Use it when |
|---|---|---|
| stdio | spawns sakur4d and speaks JSON-RPC over its pipes |
one harness; no port to manage. This is every MCP client. |
| streamable HTTP | connects to a URL | several sessions sharing one store, or a harness on another machine |
Verified against a real Hermes install:
$ hermes mcp test sakur4
Testing 'sakur4'...
Transport: stdio → D:\DuDu\Sakur4\target\debug\sakur4d.exe
✓ Connected (5765ms)
✓ Tools discovered: 17For OMP there is no MCP client. It needs a native TypeScript extension instead — which turns out to be an advantage, because an extension can see inside the agent loop and therefore reach hooks a tool provider cannot.
node integrations/omp-plugin/install.mjsThen restart OMP and ask it to list its sakur4_ tools — there should be nine.
Hermes has its own compaction path, so exposing MCP tools is not enough: its summariser
still runs. integrations/hermes-plugin/ replaces it, which also closes a loop MCP alone
cannot — update_from_response receives the provider's token accounting on every call, so
prompt-cache behaviour is measured automatically instead of reported by hand.
cp -r integrations/hermes-plugin "$LOCALAPPDATA/hermes/plugins/sakur4"
# then set `context.engine: sakur4` in ~/.hermes/config.yamlVerified by 44 contracts against a live daemon. See integrations/hermes-plugin.
skills/sakur4/ is a portable Agent Skills
package: a SKILL.md plus a dependency-free Node CLI over the daemon.
node integrations/omp-plugin/install.mjs --skill-only # → ~/.agents/skills/~/.agents/skills/ is the standard location, so OMP, Claude Code, Codex and pi all pick
it up with no further configuration. Progressive disclosure means only the description
sits in context until a task matches.
The CLI locates sakur4d across install layouts, defaults the store to
~/.sakur4/sakur4.db, and spawns with shell: false — so recorded content may contain
quotes, newlines or backticks intact. That matters when the primary use is committing the
user's words verbatim.
A harness with none of the above still works. Point it at the proxy instead of at
llama-server and nothing else changes:
sakur4d proxy --bind 127.0.0.1:8090 --upstream http://127.0.0.1:8080
# harness base URL: http://127.0.0.1:8090/v1Requests are forwarded untouched — every unrecognised route included, so a harness calling an endpoint this build has never heard of gets the upstream's own answer rather than a 404 from Sakur4. A transcript that exceeds the window is trimmed on the way through, with a marker left in place of the removed turns, and the provider's token accounting is recorded from the response the proxy already had to read.
--observe-only forwards everything unchanged and only records, which is the safe way to
see what it would have done on your real traffic before letting it act.
Do not run the proxy and the OMP extension at the same time. Both manage context, and a turn gets managed twice — OMP hangs before sending its first request. Use the proxy or the extension.
--no-extensionsdisables the extension for a proxied session. Verified, and the full comparison is in docs/verification/proxy-harness.md.
| Hook | What Sakur4 does with it |
|---|---|
session_start |
probes the daemon once; reports a missing binary before ten turns go unrecorded; live counts in the status bar |
before_agent_start |
injects the working preamble once per session — the instructions that make a model actually pin and fold |
context |
retrieves memory for the prompt, capped, and reports its own token cost so the budget stays honest |
message_end |
forwards provider token usage every turn, automatically — this is what makes cloud cache accounting work without being asked |
session_before_compact |
replaces blind summarisation with Sakur4's planned eviction |
session_shutdown |
reports stale summaries, because the next session inherits them |
resources_discover |
contributes the bundled Agent Skill |
Two install traps this avoids
omp install symlinks, which fails on Windows with a bare
EPERM: operation not permitted, symlink unless Developer Mode is on. The installer
copies instead.
OMP's plugin loader silently skips a lockfile entry that is neither declared in
~/.omp/plugins/package.json nor a symlink — reporting it only as
skipping stale lockfile entry in a log. The plugin then appears in omp plugin list
and passes omp plugin doctor, while never actually loading. Writing both files is the
fix, and it is the difference between a plugin that looks installed and one that works.
Most memory systems store what a model said about the code. That is fine until the code changes, at which point the stored interpretation is confidently wrong and nothing detects it.
Sakur4 splits the two. Deterministic parsers write facts. Models write interpretations,
and every interpretation must name the fact it was derived from. When a recall hits an
interpretation whose anchor's hash has changed, it comes back flagged STALE carrying
the anchor's current value — so the agent has something true to act on rather than
something plausible to believe.
| Guarantee | Enforced by |
|---|---|
| A model cannot write the Symbolic Ledger | SymbolicFact has exactly one constructor, and it demands a FactSource. The module imports nothing that could reach an inference client. |
| Recorded content cannot be altered | UPDATE and DELETE on episode content are blocked by database triggers — so "an evicted episode recalls byte-identically" holds for code not yet written. |
| Anchors cannot be evicted | Eviction selects from episodes; anchors live in a different table. The operation is not expressible. |
| Budget decisions and printed numbers agree | One TokenCounter, one PromptParts. Estimating in one place and measuring in another has already caused a real bug here. |
| Library code does not panic | Zero unwrap/expect/panic! paths outside tests. Malformed harness input is a typed error. |
masked → referenced → archived → dropped
A tier is never skipped. Selection is deterministic — token counts, recency, graph in-degree, explicit droppability — so a plan is reproducible, auditable, and cannot hallucinate. No model is in the decision path.
memory.fold / memory.unfold let the agent isolate a subtask deliberately: a
checkpoint is taken at fold open, the intermediate steps leave the window at fold close,
and the full trace stays retrievable with memory.recall_fold.
unknown · cold · full-re-prefill · partial-reuse · warm-restored
Those are the values of cache_status that context.plan_eviction and context.receipt
actually return. aligned and snapped are not among them — this table said they were
for several releases. They describe how a boundary was reached, not the verdict: snapping
onto a checkpoint produces partial-reuse either way, and whether the cut moved (and by how
much) is in the plan's reason and its structured snap field. A reader who filtered on
aligned would have matched nothing.
The fallback is first class. With no server, an older build without /slots, or a
sliding-window model whose checkpoints carry only partial state, the plan is still
produced, still evicts, and says why alignment was impossible. Sakur4 is always correct;
it is only sometimes not optimally fast.
A local llama.cpp slot reports its cache state through its own API. A hosted provider has no such API — but it does report, in every response, how many prompt tokens came from its prompt cache.
The signature is an inversion: append-only growth makes the cached prefix grow, while a
rewrite that replaces a long prefix with a shorter one makes it shrink even as the
prompt stays large. That inversion is detectable, and context.receipt reports it per
session.
Sakur4 never calls a provider itself. It is a subsystem, not a harness — the same reason it does not call the model. The harness pushes the numbers; Sakur4 does the accounting and the eviction.
17 tools, 4 resources, 1 prompt, targeting MCP revision 2026-07-28 with
ttlMs and cacheScope on list responses.
Memory
| Tool | Required | Optional |
|---|---|---|
memory.commit_episode |
role, content |
tool_name, session_id, slot_id, corrects |
memory.pin |
content |
kind, session_id |
memory.recall |
query |
k, session_id, file_path, include_folded, project_id |
memory.fold |
description, goal |
session_id, slot_id |
memory.unfold |
fold_id, result_summary |
session_id, slot_id |
memory.recall_fold |
fold_id |
— |
memory.staleness |
— | limit, project_id |
sakur4.dream |
— | force |
Pass tool_name on a tool result: the symbolic extractor uses it to pick a parser, so a
diff, a JSON body or an exit status becomes deterministic facts rather than prose.
Code intelligence
| Tool | Required | Optional |
|---|---|---|
code.get_repo_map |
token_budget |
focus_paths, names_only |
code.query_symbol |
qualified_name |
— |
code.impact_of_change |
qualified_name |
depth |
query_symbol reads the parser-derived index, so it cannot be stale — it is the
right way to check something you only remember from a summary.
Context and session
| Tool | Required | Optional |
|---|---|---|
context.receipt |
— | session_id, assemble |
context.plan_eviction |
session_id |
slot_id, apply, pending_recall |
context.record_usage |
prompt_tokens |
completion_tokens, total_tokens, cache_read_tokens, cache_write_tokens, reasoning_tokens, provider, model, session_id, slot_id |
session.snapshot |
— | session_id, slot_id |
session.restore |
path |
session_id, slot_id |
sakur4.status |
— | — |
plan_eviction plans by default. Nothing changes until you pass apply: true.
Resources and prompt
Resources
| URI | Contents |
|---|---|
sakur4://repo-map/{project} |
The structural outline |
sakur4://receipt/latest |
The most recent Context Ledger Receipt |
sakur4://anchors/{project} |
Every pinned constraint |
sakur4://status/{project} |
Backend, capabilities and store counts |
Prompt — sakur4_system_preamble, which names when to call each tool rather than
what it does. The failure mode with smaller instruction-tuned models is
under-triggering: they have the tools and do not reach for them. Numbered triggers fixed
that in testing; a prose description did not.
Field names differ by provider. Normalise whichever you have:
| Provider | Field to read |
|---|---|
| OpenAI | prompt_tokens_details.cached_tokens |
| Anthropic | cache_read_input_tokens / cache_creation_input_tokens |
| DeepSeek | prompt_cache_hit_tokens |
| Gemini | cachedContentTokenCount |
| Groq, others | often absent — omit the flag rather than sending zero |
Omitting is not the same as zero. Zero asserts a cache miss; omitting says the provider did not report one, and Sakur4 says so rather than blaming a cache it cannot see.
Everything is optional. The defaults work.
| Variable | Default | Meaning |
|---|---|---|
SAKUR4_BIN |
searched | Path to sakur4d |
SAKUR4_DB |
sakur4.db (daemon) · ~/.sakur4/sakur4.db (skill) |
Memory store |
SAKUR4_SESSION |
derived | Session id |
SAKUR4_BACKEND |
auto |
auto · embedded · none · a llama.cpp base URL |
SAKUR4_LLAMA_API_KEY |
unset | Sent as a bearer token to the inference server. Needed for any server that requires authentication — llama.cpp behind a reverse proxy, or with --api-key |
SAKUR4_LLAMA_URL |
unset | Base URL for the inference server, when no --backend is given |
SAKUR4_SNAPSHOT_DIR |
the temp directory | Where session.snapshot writes save files |
SAKUR4_EMBED_URL |
unset | OpenAI-compatible embedding endpoint, for semantic Atlas entries |
SAKUR4_EMBED_MODEL |
nomic-embed-text |
Model name to request from that endpoint |
SAKUR4_EMBED_API_KEY |
unset | Key for a configured embedding endpoint |
SAKUR4_PROJECT_ROOT |
the working directory | Root used to name a project for the Repo Cortex index |
SAKUR4_EVICTION_PROFILE |
chosen from the backend | cache-first · window-first · balanced — overrides the automatic choice |
SAKUR4_SQLITE_VEC_PATH |
searched | Path to a sqlite-vec extension, for vector search without the fallback scan |
SAKUR4_TRANSPORT |
stdio |
Transport for serve, when --transport is not given |
SAKUR4_SKILL_DIR |
~/.agents/skills |
Where install.sh places the Agent Skill |
SAKUR4_FORCE |
unset | install.sh replaces an existing skill instead of leaving it alone |
A variable that is set and does nothing is worse than one that does not exist, so the table above
is checked: docs/verification/env-vars.mjs fails if a name documented here appears nowhere in the
code. It was written after finding several that were read by the code and named by no document —
SAKUR4_LLAMA_API_KEY among them, which is the one anybody running a gated server needs first.
Oh My Pi extension extras
| Variable | Default | Meaning |
|---|---|---|
SAKUR4_RETRIEVE |
true |
Inject retrieved memory before each turn |
SAKUR4_REPORT_USAGE |
true |
Report provider usage automatically |
SAKUR4_OWN_COMPACTION |
true |
Take over compaction |
SAKUR4_RECALL_BUDGET |
1200 |
Approximate token cap on injected memory |
SAKUR4_PLUGIN_LOG |
unset | Append lifecycle diagnostics to this file |
| Backend | What it is | Cache coherence |
|---|---|---|
llama.cpp |
A real server. Probes /slots, /props, /tokenize, /metrics. |
Full |
embedded |
Simulates a checkpoint ring in-process. Default when nothing is listening. | Full, simulated |
none |
Coherence disabled; degradation logged. | Reported as full-re-prefill |
llama-server -m model.gguf -c 65536 --slots -cms 256 -ctxcp 64
sakur4d --backend http://127.0.0.1:8080 doctordoctor prints exactly which endpoints were detected and what that means for compaction.
--backend accepts a URL on another machine.
Commands
serve Run the MCP gateway (stdio or HTTP; --banner for stderr diagnostics)
proxy Run the OpenAI-compatible reverse proxy (FR-18)
--upstream <url> --bind <addr> --observe-only
config <harness> Print ready-to-paste integration config
doctor What backend and cache capabilities were detected
demo Guided end-to-end walkthrough
index <path> Build or incrementally refresh the Repo Cortex index
repo-map [--budget N] Token-budgeted structural outline
symbol <qualified> A symbol's current signature
impact <qualified> Transitive blast radius
commit <session> <text> Append a turn (--role, --tool, --slot)
pin <text> --kind <k> Pin a constraint
anchors [--session S] List pinned constraints
recall <query> [--k N] Hybrid search
plan <session> Show an eviction decision without applying it
receipt <session> Token accounting and cache verdict
snapshot / restore Persist or reload slot KV state
dream One memory-maintenance pass
staleness Summaries that no longer match their source
fold, unfold and recall_fold exist as MCP tools only — they are called by an agent
mid-task, not by a person at a shell, and the skill's CLI exposes them through the
protocol for that reason.
"Sakur4: no sakur4d binary found"
The extension searched and did not find it. Check where yours actually is:
which sakur4d # or: where.exe sakur4d on WindowsIf that prints nothing, the binary is not on PATH. Either move it onto PATH — the
most likely place is ~/.cargo/bin — or point at it directly:
export SAKUR4_BIN=/full/path/to/sakur4d # $env:SAKUR4_BIN on WindowsThe warning lists every path that was searched, so you can see whether your install landed somewhere unexpected.
The OMP plugin is installed but its tools are missing
Set the diagnostic log and restart OMP:
SAKUR4_PLUGIN_LOG=/tmp/sakur4.log omp
cat /tmp/sakur4.logIf there is no log at all, the extension never loaded — check that omp-sakur4 is in
~/.omp/plugins/package.json dependencies, not only in the lockfile. If the log shows
daemon: null, see the entry above.
A plugin that silently does nothing and a plugin that failed to load are otherwise indistinguishable, because OMP surfaces extension-load errors only to a TTY.
A turn was slow, or cost more than expected
sakur4d receipt <session>This prints where the token budget went and the cache verdict. If it says
full-re-prefill, the plan could not align to a checkpoint — the reason is printed with
it. If it says PREFIX-BROKEN, a rewrite invalidated the provider's cache and you were
billed for it.
Also worth checking: sakur4d doctor reports which cache endpoints were detected. A
backend resolved as embedded explains a simulated verdict.
Recall returns a STALE result
That is the system working. A stale summary is one whose anchor has changed since it was written. It comes with the anchor's current value attached — use that, not the summary above it.
To regenerate them: sakur4d dream, or memory.staleness to see the full list first.
Everything is slow and the store is huge
Check what is in it and how stale it is:
sakur4d doctor # store counts, backend, capabilities
sakur4d staleness # summaries that no longer match their sourceCold archival runs during dream. Snapshots are pruned by count under
SAKUR4_SNAPSHOT_DIR.
The honest answer, measured over 205 turns against a real repository — the same session run twice, once the way a harness does today and once through Sakur4:
Measured against a real llama.cpp server (a 27B model, 81,920-token window):
| without | with Sakur4 | change | |
|---|---|---|---|
| tokens per turn | 31,580 | 31,980 | +1.3% |
| prefix kept reusable | 0 | 47,425 | — |
| recall accuracy | 50% | 92% | +42 pts |
| pinned constraint survived | lost at compaction 2 | survived | — |
+1.3% more tokens, for 42 points of recall and a constraint that no longer gets destroyed. At that magnitude it is not a trade at all.
The engine picks its eviction tuning from what the backend can do: cache-first when a
checkpoint ring exists and a preserved prefix is genuinely reusable, window-first when
there is no checkpoint to align to and window room is the scarcer resource. Your server
selects window-first automatically — doctor reports which is active and why.
node docs/bench/ab.mjs --repo . # embedded backend
node docs/bench/ab.mjs --repo . --backend http://host:8080 # real prefill numbersA live-model companion — whether a real model actually benefits, rather than whether the engine works — is in docs/bench/live-model.md. Controlled A/B/C against a locally served 27B model through OMP: plugin-on answered a question only memory could answer; plugin-off with the same store said UNKNOWN; plugin-on with an empty store said UNKNOWN.
Full method — including three bugs the benchmark itself had, recorded rather than quietly fixed — is in docs/bench.
A release claim is worth exactly as much as the evidence behind it.
The tests are not all unit tests. stdio_transport.rs spawns the real binary and
speaks JSON-RPC over its pipes. gateway.rs drives the tool surface over a live HTTP
listener using the SDK's own client. cache_coherence.rs states the central claim as
contracts and fails if the preserved prefix stops being a byte prefix of what the server
is actually sent.
That suite found two bugs no amount of self-testing would have:
- A single tool whose
outputSchemahad notype— because it returned a bareserde_json::Value— made Hermes reject the entire 17-tool catalog and refuse to connect. - A
WARN-level log line written to stdout corrupted the JSON-RPC channel over stdio. A client that reads stdout as frames cannot recover from that.
One command runs everything, and reports skipped separately from passed — a run that skipped its live-server checks is not a green run:
node verify.mjs # everything this machine can run
node verify.mjs --upstream http://host:8080 # add the live llama.cpp checks
node verify.mjs --quick # skip the slow benchmarks
node verify.mjs --only rust,hermes # a subset, by group or check id
node verify.mjs --list # what exists, and what each group needsEvery group runs on every platform it can, and anything that cannot run is named with its
reason rather than silently omitted — a skipped check that reports why is worth more than a
passing one that tested nothing. Passing --upstream adds the live-server checks; without it they
report as skipped and say so.
--require-all turns any skip into a failure, which is what CI uses so a check cannot quietly stop
running.
| Group | Needs | CI |
|---|---|---|
rust |
nothing | Linux, macOS, Windows |
encryption |
OpenSSL development files | Linux only — see crates/sakur4-core/Cargo.toml |
hermes |
python + a built daemon | Linux, with a stubbed Hermes |
bench |
a repository to index | partly |
live |
--upstream |
no — no server in CI |
harness |
OMP or the Hermes CLI | no — not installed in CI |
Underneath, CI also runs a release-profile build (LTO and codegen-units = 1, so a
release-only link error cannot hide), an MSRV build at the declared 1.94, and a
cargo publish dry run that builds the extracted archive in isolation.
Why one binary instead of a service
A memory layer that needs a daemon lifecycle, a port, a supervisor and a reconnect path
is a memory layer that gets uninstalled. sakur4d is one binary: stdio mode spawns it as
a child, HTTP mode runs it in the foreground. There is nothing to keep alive.
Why the plugin has no runtime dependencies
typebox and pi-ai live nested inside OMP's own tree, not somewhere a separately
installed package can resolve them. Tool schemas are therefore written as plain JSON
Schema objects — which is the shape the host serialises anyway. A plugin that fails to
load because of a schema library is a plugin nobody can use.
Why extraction is extractive by default
Promotion summarises a turn into something searchable. Doing that with a model would mean Sakur4 needs a model resident to maintain memory — and would put a model in the path of something that is supposed to be deterministic. The default is extractive and anchored; an optional auxiliary endpoint switches to genuine interpretation, still anchored. A promotion that produces a longer summary than the turn it replaces is skipped and reported.
Why staleness is computed at read time
A boolean stale column needs a job to keep it current, and a job that has not run is a
lie the system tells itself. Sakur4 compares the anchor's stored hash against its current
one when the entry is read. There is no window in which a stale interpretation looks
fresh.
Why the figures are generated
docs/assets/generate.mjs defines the palette, type scale and
primitives once, and each figure is a function over them. More importantly the numbers in
them come from sakur4d demo, so when the engine changes they are regenerated rather
than left to drift into a picture nobody re-reads.
Most agent-memory projects answer "how do we remember more?" Sakur4 answers "how do we forget well?" — because the constraint is not storage, it is the context window and what refilling it costs.
These rows describe what each approach does, not what its authors know. An earlier version of this table got that wrong: it marked competitors as not being "aware the KV cache exists", which is a claim about other people's understanding rather than about their software, and one I cannot support. Letta documents prompt caching; most harnesses have some notion of it. What can be stated is observable behaviour, and that is what the rows now say.
| Sakur4 | MemGPT / Letta | Mem0 | Plain RAG | Harness compaction | |
|---|---|---|---|---|---|
| Persistent memory | ✅ | ✅ | ✅ | ✅ | ❌ |
| Staleness detection | ✅ read-time, anchored | ❌ | ❌ | ❌ | n/a |
| Deterministic facts | ✅ no model in path | ❌ | ❌ | ❌ | ❌ |
| Asks the server where its cache can be rewound to | ✅ | ❌ | ❌ | ❌ | ❌ |
| Places the eviction boundary at that point | ✅ | ❌ | ❌ | ❌ | ❌ |
| Reports cloud prompt-cache accounting per turn | ✅ | ❌ | ❌ | ❌ | ❌ |
| Verbatim recall after eviction | ✅ by construction | partial | ❌ | ✅ | ❌ |
| Code-graph awareness | ✅ tree-sitter | ❌ | ❌ | partial | ❌ |
| Runs with no model resident | ✅ | ❌ | ❌ | ✅ | ✅ |
| Deployment | one binary, MCP | service | service | varies | built in |
Read ✅ and ❌ as "does this, out of the box, in the form shipped" — not as a ranking, and not as a statement about what a system could be extended to do. Several of these are more mature projects than this one, and the comparison is not that they are worse. It is narrower: none of them makes the inference server's prompt cache the thing the eviction decision is built around, which is the specific problem this project exists to solve.
This table is not verified against those projects' current documentation, and any row could be out of date. If you maintain one of them and a row is wrong, that is a bug in this README and a correction is welcome — a comparison nobody can trust costs more than the row was worth.
Stated plainly, because the alternative is finding out later.
Batching is safe — this used to be the headline limitation
JSON-RPC permits a server to process messages "as a set of concurrent tasks, processing them in any
order", and MCP's stdio transport correlates responses only by id. Sakur4 relies on the order: the
preamble tells the model to commit a turn and then consult what it remembers, and the tools are
stateful in exactly that way.
Written as a pipelined batch — several frames at once, stdin closed — a read could be executed
before the write in front of it: memory.commit_episode answered with a real ep_… identifier while a
sakur4.status in the same batch reported the count from before it.
That is fixed. The stdio transport now withholds message N+1 until the response to N has been
written, so a batch is served in the order it was sent. A regression test sends twelve commit-then-status
pairs in one batch with stdin closed and requires every round to be correct — the count that used to
read 1/12 wrong now reads 0/12.
It took fourteen attempts and the record is worth reading, because the last one succeeded only after
instrumenting a counter instead of reasoning about it — the first run printed forwarded=3 answered=2,
and the cause was an off-by-one that had been mistaken for a threading problem throughout. Two earlier
mistakes are documented in docs/DESIGN.md and are now each pinned by a test: a
notification produces no reply and must not be waited on, and end of input is not the end of output —
a transport that closes its read side at EOF truncates a large reply such as tools/list.
Not yet true
-
No real llama.cpp server has been contacted.Now verified against a a live locally served 27B model — see docs/verification. That build exposes no checkpoint API (save/erase return 501, no checkpoint ring), so Sakur4 correctly reportsno checkpoint source detected. Prefix reuse nonetheless works there: a 2,219-token preserved prefix followed by new content is reused in full, at ~0.9 ms/token saved. The practical consequence is that Sakur4's receipt is pessimistic on such a backend — it reportsfull-re-prefillwhere reuse is real but unverifiable. -
No live agent session through Hermes. Its transport is verified (
hermes mcp test sakur4discovers all 17 tools) and the OMP tools were driven end-to-end by a live model, but Hermes' own tool selection is untested. -
The OMP compaction hook has never fired for real. The tool path is verified; forcing OMP past its context limit is separate work.
-
OMP version. Verified against 18.2.x — the installed build at the time of writing is 18.2.11. Earlier text named 18.1.17 as "the tested version", which was true when written and had been wrong for several releases; the benchmark note in
docs/bench/live-model.mdsaid 18.2.0, so the same fact was recorded at three different values.The plugin declares no version floor at all. Its
package.jsonhas"@earendil-works/pi-coding-agent": "*"as an optional peer dependency, so nothing records the minimum the README claims — the "18.1.17 or newer" line inintegrations/omp-plugin/README.mdis prose and not enforced. The extension API is undocumented, so a minor bump can change it without notice;--no-extensionsis how to tell whether a failure is OMP's or this plugin's.
Deliberately absent
- Encryption at rest (FR-20) exists but is off by default. Build with
--features encryption; the store is otherwise readable by anyone with file access, and it holds a verbatim transcript. See SECURITY.md. - Authentication on the transports. Localhost binding is the control.
--bind 0.0.0.0exposes the entire Memory Fabric, including writes, to anyone who can reach the port. - Snapshots are as sensitive as the store. A slot-save file is 60–500 MB of model state representing everything the session has seen. Nothing encrypts them.
Unmeasured
- No MCP conformance run against a reference client, no LoCoMo, no Endurance Benchmark.
- NFR latency and memory numbers are unmeasured on reference hardware — the development box has a GPU with 4 GB, which cannot host the target workload at all. That is why the embedded and fake backends exist.
cargo auditandcargo denyboth run in CI. Before 0.2.1 the licence and dup hand.
The full list, including every deviation from the source requirements, is in docs/DESIGN.md.
crates/sakur4-core/ the engine — no transport, no MCP
crates/sakur4d/ the daemon — CLI, MCP gateway, reverse proxy
crates/sakur4-testkit/ fixture repos, a fake llama.cpp server
integrations/omp-plugin/ native Oh My Pi extension + installer
integrations/hermes-plugin/ a Hermes ContextEngine, replacing its summariser
skills/sakur4/ portable Agent Skills package
docs/DESIGN.md how each requirement is met, and the trade-offs taken
docs/RELEASING.md how to cut a release, and what is manual and why
docs/bench/ what changes with Sakur4 and without it
docs/verification/ the measurement scripts, and what each one establishes
docs/assets/ the figures above, and the generator that draws them
All five integration routes are in the tree — MCP needs nothing beyond the daemon, and the
other four live in integrations/ and skills/. verify.mjs at the root runs every check
across all of them.
Contributions are welcome. CONTRIBUTING.md covers the invariants worth knowing before changing anything — the structural guarantees above, plus why the cache-coherence code is the most delicate part of the project and why a change there should make you suspicious of a green test suite.
Security problems: please report privately per SECURITY.md rather than in a public issue.
Sakur4 — for the sakura, and for the
4 in sakur4d, which is what you get when the name you want is already taken.