Six read-only audit tools for Claude Code Router (CCR).
git clone https://github.com/acrot0/ccr-toolkit.git
cd ccr-toolkit
node src/cli.mjs # run the whole sweep, get one verdict ✅ doctor OK
✅ cache OK
⚠️ check WARN
✅ takeover OK
3 ok, 1 warn, 0 fail
Needs attention: ccr-toolkit check
Overall: ⚠️ WARN
Then node src/cli.mjs doctor for the full detail on whichever needs it.
They answer six questions CCR itself does not:
- Why did that request fail? →
ccr-doctor - Where exactly did it fail? →
ccr-trace - What did the request body actually contain? →
ccr-body - Is my prompt cache actually working? →
ccr-cache - Is my gateway config healthy? →
ccr-check - Will CCR silently overwrite my agent's config with a broken snapshot? →
ccr-takeover
All six are read-only. They open databases with readOnly: true and never write to your config.
CCR is a local model gateway. It routes Claude Code (and other agents) through whichever upstream you configure. Six problems are invisible from inside it:
CCR logs every request to request-logs.sqlite. When a request fails it records
the upstream's response in response_body_text — and then never reads it. The
error column it shows you is a different thing entirely, and is frequently
blank.
Measured on a real install: four of seven failures carried an empty error
column — and in all four, the upstream's actual explanation was sitting unread
in the body. "Can only get item pairs from a mapping", "请求包含未知字段",
"请求过于频繁": all captured, none surfaced.
ccr-doctor reads the body, names the fault, and tells you what to do about it:
#2450 2026-09-24T11:57:38Z [tierflow::openai_chat_completions] HTTP 400 → tool-pairing
cause: An assistant turn carries a tool_use with no matching tool_result (an orphaned block).
action: The upstream rejects unpaired tool blocks. A gateway-side cleaner that strips or
re-pairs them fixes this — check whether one is installed and whether it is
enabled for this provider.
upstream said: "The provided messages input is invalid. The error info is
[Can only get item pairs from a mapping.]."
It also answers the two questions that make a gateway look flaky when it is not:
- Did the model name survive routing? CCR sometimes resolves to an internal
<provider>::<protocol>/<model>form. Anything keying off the model string — cache identity, pricing, per-model routing — silently stops matching. - Is the cache actually regressing? A miss is invisible: the request
succeeds either way, only the bill differs. But most "cache is broken" reports
are wrong, because a new session is supposed to start cold.
ccr-doctoronly flags a regression when the prefix stayed the same size and the cache vanished anyway. On the machine this was built against, that check reports healthy on two paths whose raw hit rates are 98.8% and 90.2% — the naive "does the rate flip?" detector called both of them broken.
The measured false positives are in the tests.
test/ccr-doctor.test.mjscarries the exact row sequences that a naive detector gets wrong — the 465K-prefix session ending, the fresh 40K session starting — so the check cannot regress into crying wolf again.Background reading: The gateway knows why it failed. It just won't tell you. — the full failure table, the hop chain, and how a cache check that flagged two healthy providers (98.8% and 90.2%) got fixed.
Even when a failure does carry a message, it does not tell you which stage of
the gateway produced it. CCR writes a full 7-to-16 hop forensic chain per
request and surfaces none of it. ccr-trace reads it — see the usage
section for why
the hop number is the part that tells you what to fix.
The largest requests are both the most likely to be folded and the most likely
to fail — so the evidence you need is gone exactly when you need it, and nothing
flags it. ccr-body recovers what survives. See usage.
A 0% cache hit rate can look identical to a 99% one — the request succeeds either way, only the bill differs. And the number is easy to misread:
- Some upstreams report
input_tokensas the uncached remainder; others report the full prefix including cached tokens. Compute the rate with the wrong convention and 99.9% reads as 50%. - The same model on two different providers can genuinely differ (one protocol caches, another does not).
- Small requests below the minimum cacheable block always miss — averaging them in drags the rate down and looks like a regression.
ccr-cache reads CCR's usage.sqlite and reports the rate per provider|model, auto-detecting the reporting convention per row.
ccr-check catches the failure modes that produce no error message: a fallback mode of off (429s go straight to the client), empty model slots, SQLite files that are 98% free pages, and request bodies that CCR folded into a preview (breaking forensics).
This is the one with no upstream fix.
When you enable a global-scoped profile, CCR takes over that agent's global config file. When you later disable it, CCR restores the file from a backup chain. The restore picks the newest snapshot that passes CCR's own isManagedContent check — which is not necessarily a good one.
If that snapshot is malformed, restoring it silently breaks the agent. The symptom is "my agent randomly stopped working", with no error anywhere.
The behaviour is intentional and the maintainer's answer is "don't use global scope" — see musistudio/claude-code-router#1575 (open since 2026-07-21). That is not always an option.
ccr-takeover audits the backup chain and tells you which snapshot CCR would pick, and whether it is healthy — before it bites.
Background reading: Every time I checked the config file, it was correct. Every time I looked away, it broke again. — how these rules were reverse-engineered from the bundled
cli.js, the two that are easy to get wrong, and how to check your own setup in 30 seconds.
Requires Node.js ≥ 22.13 (uses the built-in node:sqlite).
22.5.0 introduced
node:sqlite, but it stayed behind--experimental-sqliteuntil 22.13.0 — on 22.5–22.12 the imports fail withERR_UNKNOWN_BUILTIN_MODULEunless you pass the flag. CI runs 22.13.0 and 24.x to keep this honest.
npm install -g ccr-toolkit # or just run it: npx ccr-toolkit
ccr-toolkit # run the whole sweep, get one verdictOr from source:
git clone https://github.com/acrot0/ccr-toolkit.git
cd ccr-toolkit
npm install # dev dependency: vitest
node src/cli.mjs # same sweep, no global installNo runtime dependencies.
Windows note. The repo ships a
.gitattributesthat forces LF. This is not cosmetic: with Git for Windows' defaultcore.autocrlf=true, a clone rewrites line endings to CRLF, and the vitest/esbuild transform then fails withSyntaxError: Invalid or unexpected tokenon these files. Verified by cloning into a clean directory — LF passes, CRLF fails to parse. Don't remove the file.
node src/ccr-doctor.mjs # last 7d, 20 failures
node src/ccr-doctor.mjs --since 24h
node src/ccr-doctor.mjs --json # machine-readable, for CI or a dashboard
node src/ccr-doctor.mjs --limit 100Reads request-logs.sqlite, extracts the upstream's own explanation from the
captured response body, and classifies each failure into a named fault with a
concrete next action.
Failure classes: tool-pairing · unknown-field · rate-limited ·
client-abort · auth · upstream-5xx · bad-request · opaque
opaque is the one worth watching — it means CCR recorded a failure with
nothing a human can act on. That happens when the body was never captured, or
when it was folded into a preview (bodies over 160KB get an elision marker
spliced into the middle, which is exactly the size of request most likely to
fail). ccr-check reports the preview rate; ccr-doctor reports the
consequence.
Three checks run alongside the failure list:
| Check | What a warning means |
|---|---|
silentFailures |
A large share of failures carried no usable message |
cacheStability |
The cache regressed on an unchanged prefix — an injected block is being edited between requests. A new session starting cold does not trigger this. |
modelResolution |
A successful request resolved to an internal <provider>::<protocol>/<model> id, so anything keying off the model string stops matching |
Exit code 1 only on a fail-level finding; warnings do not break automation.
Why
cacheStabilityis conservative. The obvious detector — "does the hit rate flip on and off?" — flags healthy gateways. Measured against this machine: it reported two paths as oscillating whose real hit rates are 98.8% and 90.2%. The flips were new sessions legitimately starting cold (a 465K-token prefix ending, a fresh 40K one beginning). This check instead compares each miss against the last prefix that did hit, and only fires when the size barely moved. Those exact row sequences are in the test suite.
node src/trace-view.mjs # most recent failing request
node src/trace-view.mjs --id 1996
node src/trace-view.mjs --last 5
node src/trace-view.mjs --all # include successful requests
node src/trace-view.mjs --jsonCCR writes a full forensic chain per request: measured, a normal call produces 7 hops and a rate-limited one that retried four times produces 16. Each hop carries the before/after of every field it touched.
#2450 HTTP 400 7 hops
📥 0 [ingress ] request.ingress ok 0ms
📥 1 [ingress ] gateway.header-normalization ok 0ms {remove:/headers/x-api-key, +3}
🔀 2 [routing ] router.route-output ok 0ms {replace:/body, +4}
⏭️ 3 [planning ] fallback.execution-plan noop 0ms
⚙️ 4 [capability] provider.capability-routing ok 0ms {replace:/body/model, replace:/routing/model}
📤 5 [attempt ] upstream.attempt.prepare ok 1ms {remove:/headers/content-length, +1}
❌ 6 [outcome ] upstream.attempt.outcome error 796ms → HTTP 400
❌ failed at hop 6 (upstream.attempt.outcome) with HTTP 400 — at the upstream call
🔀 hop 4 rewrote the model: tierflow/tierflow → tierflow::anthropic_messages/tierflow
Why the hop number matters. A 400 that fails at hop ≤3 was already malformed before the gateway touched it — a client bug. A 400 at hop 6 left the gateway intact and came back rejected — a routing or upstream bug. Same status code, opposite fix, and the error message does not distinguish them.
It also surfaces two things nothing else in the toolchain reads:
- The model rewrite — which hop introduced the internal
<provider>::<protocol>/<model>form, so you can tell "my config is wrong" from "the gateway transformed it". - The retry ladder —
1s → 2s → 4smeans a capacity problem; no retry at all on a 429 meansRouter.fallbackis set tooff. Anetwork-errorwithretryDelayMs: 0is reported as what it is (a dropped connection), not as a backoff that never happened.
node src/body-salvage.mjs # every failing request
node src/body-salvage.mjs --id 2088
node src/body-salvage.mjs --since 24h
node src/body-salvage.mjs --all # include successful requests
node src/body-salvage.mjs --jsonCCR folds request bodies over 160KB into a "preview": it cuts the middle out and
splices in ... N bytes omitted from preview .... That marker breaks the JSON,
and request_body_truncated stays 0 — so the body looks intact and parses as
nothing.
Measured on a real install: 96% of stored bodies are folded this way. The bodies most likely to be folded are the large ones, which are also the ones most likely to fail.
The fold keeps the head and the tail, and model is a top-level key that sits
in the head — so it survives every fold. Measured: 300 of 300 folded bodies
yielded it.
That one fact answers the single most common gateway failure report, "Missing model in request body" (28 open issues upstream, 274 comments):
#2088 [tokenrhythm::anthropic_messages] HTTP 400 stored 947776B
body folded: 783936 bytes omitted from the middle
model: "deepseek-flash"
→ The client DID send model="deepseek-flash" — the fault is downstream
of the client, not a missing field.
It reports model-absent only when a body exists and carries no model — never
from an absent body, because "no evidence" and "no model" are different facts.
node src/cache-monitor.mjs # last 24h, threshold 50%
node src/cache-monitor.mjs --window 7d
node src/cache-monitor.mjs --json
node src/cache-monitor.mjs --threshold 80 # exit 1 if a live path is below 80%Exit code 1 if any live path (seen in the last --live-hours, default 24) is below the threshold — usable as a CI gate.
Retired paths are listed separately and excluded from the gate, so a decommissioned provider at 0% will not fail your build forever.
node src/ccr-check.mjs
node src/ccr-check.mjs --jsonChecks fallback mode, profile slots, database bloat, request-body preview loss, and status=0 logging gaps. Exit code 1 on any fail-level finding (warn does not fail the build).
node src/ccr-takeover-audit.mjs
node src/ccr-takeover-audit.mjs --json
node src/ccr-takeover-audit.mjs --profile myagent=~/.myagent/config.json:zcode
node src/ccr-takeover-audit.mjs --expect-opencode-model your-main-modelExit code 1 if a snapshot CCR would actually restore is unhealthy.
Profiles are discovered from CCR's own global-profile-takeover.json when present, falling back to known locations (~/.zcode/cli/config.json, ~/.config/opencode/opencode.jsonc). Use --profile to audit anything else.
Note on coverage. This tool covers agents whose restore predicate is
isManagedContent(zcode-style) or thex-ccr-clientheader check (opencode-style). Claude Code's ownsettings.jsonand Codex'sconfig.tomluse a different mechanism (managed blocks) and are deliberately skipped — applying zcode's predicate to them would produce false verdicts. They are reported as uncovered rather than guessed at.
Reverse-engineered from CCR's bundled cli.js (verified against v3.1.1), not guessed:
apply(file, {isManagedContent}) // does CCR touch this file at all?
if (exists && !isManagedContent(cur)) return {changed:false} // bail, leave it alone
let o = pickRestore(file, isManagedContent)
pickRestore(file, check) // choose the restore source
for (r of [...listBackups(file).reverse(), `${file}.ccr-original`])
if (!exists(r)) continue
if (!check(read(r))) return {content: read(r), file: r} // first one that passes
listBackups(file) // NOTE: exact startsWith prefix
readdir(dir).filter(n => n.startsWith(basename + ".ccr-backup-")).sort()
Two consequences that are easy to get wrong:
- Only
config.json.ccr-backup-*participates. Hand-made backups namedconfig.json.bak-*never matchstartsWithand are never selected. Conversely, a backup you name.ccr-backup-…will become the preferred restore source — naming a backup can accidentally weaponise it. - A snapshot must pass
isManagedContentto be eligible. So "the hooks are broken" does not mean "there is a threat" — it must also be managed. An unmanaged broken file is inert.
Both rules have unit tests that will fail if someone changes them.
npm test206 tests. Each one locks a rule that was learned from a real misdiagnosis — the failure modes they encode all produced silent wrong answers, not crashes:
status=0rows are a logging gap, not failures. Counting them as failures turns a ~90% success rate into 66.7%.- 429s must be clustered into events. 494 raw log lines were only 53 events; reporting the raw count overstates the problem ~10×.
- A
total-convention upstream can reportcache_read > input, which makestotalimpossible — the code falls back toremainderinstead of silently clamping to 100% and hiding the real gap.
Pushing a v* tag runs .github/workflows/release.yml: tests, a tag/version
agreement check, a npm pack --dry-run listing, then npm publish --provenance.
To rehearse without publishing, run the workflow manually with dry_run: true.
The first release is the awkward one. Publishing authenticates with trusted publishing (OIDC), which is configured from the package's settings page on npmjs.com — a page that does not exist until the package does. So the first version goes out with a token, and every version after it uses OIDC:
- First release — either publish by hand (
npm publish --access public, which prompts for 2FA), or add anNPM_TOKENrepo secret with publish rights and let the workflow do it. - Once the package exists — npmjs.com → the package → Settings →
Trusted Publisher → GitHub Actions:
acrot0/ccr-toolkit/release.yml, allowed actionnpm publish. - Delete
NPM_TOKEN. The workflow passes it only when it is set, so an absent secret means npm falls through to OIDC. Leaving a publish-capable token in the repo is the exact risk this removes.
This matters now rather than later: npm is retiring long-lived publish tokens. Granular access tokens that bypass 2FA lost sensitive account and package operations in August 2026, and lose direct publishing around January 2027 — their publishing surface shrinks to reading private packages and staging a publish that a human must approve with 2FA.
- Read-only by design. There is no
--fix. Deciding what a "healthy" config looks like is yours, not the tool's. - CCR internals change. The takeover rules were verified against v3.1.1. If a future version changes
listBackupsorisManagedContent, re-verify before trusting the verdict. node:sqliteis experimental. It prints anExperimentalWarning; the tools suppress only that one and pass others through.- Not affiliated with CCR. Independent tooling.
MIT