An advisory digital-forensics/incident-response co-pilot that runs entirely on your Mac. Two small on-device models do the drafting — Apple's built-in model (macOS 27) routes a plain-English question into a curated intent catalog and narrates findings, and a frozen Granite-4.1-3B (via Apple MLX) drafts free-form DuckDB queries under a grammar — while a deterministic harness owns correctness, and you approve before anything is acted on. It fits a 16 GB MacBook Air with room to spare for the forensic tools.
It is built on one hard-won result: with the right structure around it, an untuned 3B model beats a fine-tuned one 3.5× on held-out forensic queries — so we ship the base model frozen and put the intelligence in the harness, not the weights.
./install.shThat's it. On a fresh Mac it installs Homebrew, Python, the dependencies, downloads the model
(~2.6 GB), wires up the clue command, and runs a self-test. On an existing Mac it
detects what you already have and installs only what's missing. Safe to re-run.
Add the dockerized forensic tools (Volatility 3, Plaso, …) when you want them:
./install.sh --with-toolsOther flags: --no-model (wire everything but skip the download), --quick-verify (skip the slow
end-to-end model test). Run ./install.sh --help for the list.
The installer puts one command on your PATH: clue (and dfir-copilot, the same program under
its long name). It links into ~/.local/bin if that is already on your PATH, otherwise into
Homebrew's bin directory — no sudo, no shell-rc edits — and never overwrites a command that is
already there. Type clue on its own for the cheat sheet:
clue query "<question>" <artifact.csv> ask a plain-English question of a CSV artifact
(the file may come first; the SQL is shown for approval)
<tool> ... | clue triage verdict + findings from tool output (deterministic, no model)
clue narrate <output.txt> the verdict plus a short grounded summary
clue ocr <screenshot.png> | clue triage transcribe a screenshot or photo of tool output, then triage it
clue backends what this Mac can use (Apple on-device model, Granite)
clue verify self-test the install
clue tools vol3|plaso|remnux ... run a dockerized forensic tool read-only (after --with-tools)
If the installer could not link it (it says so), link the launcher yourself into a directory on your
PATH (may prompt for sudo); until then ./dfir-copilot inside the repo is the same program:
[ -e /usr/local/bin/clue ] || sudo ln -s "$PWD/dfir-copilot" /usr/local/bin/clue(DFIR_ALIAS= before ./install.sh skips the short name; DFIR_ALIAS=yourname picks another.)
New here? The comprehensive user guide covers everything — installation details, every command, end-to-end investigation walkthroughs, how it decides things, how to extend it, safety, limitations, and an FAQ.
Ask a question about a CSV artifact (a routine question is routed through the intent catalog and answered from deterministic SQL in ~3 s, with a plain-English line saying what the query computes; anything else goes to Granite, which writes the SQL under a grammar; either way it runs read-only and shows you the result and the query to approve):
clue query "how many failed password attempts are there?" copilot/sample/auth_sample.csv
clue query "which single source IP has the most events?" auth.csv --cross-check # both paths, flag disagreement
clue backends # which models this Mac can useTriage tool output (deterministic verdict + findings + suggested next commands — no model needed):
volatility3 -f mem.raw windows.malfind | clue triageTriage with an optional grounded summary from the model (the verdict above is still authoritative):
clue narrate suspicious_output.txtTriage a screenshot or photo of tool output (a console someone photographed, a screenshot pasted into a ticket): Apple's Vision framework transcribes it with no language model in the loop — hex and encoded blobs come through as written rather than paraphrased — and the text goes into triage. It is still OCR: read the transcription before relying on a verdict built on it.
clue ocr screenshot.png | clue triageRun a forensic tool the co-pilot drafted a command for (after --with-tools):
clue tools vol3 -f /data/mem.raw windows.pstreeEvidence is mounted read-only at /data. Anything a tool writes goes to /out: ./out on the host,
created only when a command names /out — or, with DFIR_OUT=/path set, that directory on every run
(never the evidence directory itself or a parent of it).
Re-check everything works:
clue verifyWho it's for: a mixed-experience security team — from a tier-1 SOC analyst who isn't a forensics specialist, up to a seasoned DFIR examiner. It raises the floor (juniors miss less and move faster) without slowing the experts down. It is strictly advisory: it drafts and explains; you approve and run everything.
Two concrete jobs, matching the two modes above:
1. Ask a question about an artifact in plain English (query mode). You have a parsed log or tool export as a CSV — an SSH auth log, a Windows event-log export, an EZTools artifact, a Plaso timeline — and a question you'd otherwise hand-write SQL for. You ask in plain English; it writes the query, runs it read-only, and shows you both the answer and the exact SQL, so you can confirm it asked what you meant before trusting the number. Typical questions:
- "How many failed password attempts are in this auth log?"
- "Which single source IP has the most events, and how many?"
- "How many distinct usernames were tried in invalid-user attempts?"
2. Triage whether tool output is interesting (triage mode). You pipe the raw output of a forensic tool — a Volatility memory scan, a Plaso timeline, an event-log dump — into it and get back a verdict (MALICIOUS / SUSPICIOUS / BENIGN / UNDETERMINED), the specific indicators found with the evidence for each, and the suggested next commands. For example, a Volatility malfind hit showing an executable-and-writable region inside lsass.exe comes back MALICIOUS, names the indicator, and suggests dumping the region and pivoting to the network and process-tree views — so a junior doesn't already have to know how to read malfind.
The workflow is always the same shape: drafts → you approve → the tool runs read-only → you feed the output back into the next triage or query. A human is in the loop on every action.
What it is not: not an autonomous agent that runs investigations by itself, not a malware sandbox or detonation platform, and not a substitute for an analyst's judgment. It handles the common, well-trodden 80% quickly so you can spend your attention on the hard 20%.
you ──▶ clue
│
│ QUERY path
│ ├─ artifact family: deterministic (header fingerprint + the words actually in the data)
│ ├─ in the catalog? Apple's on-device model picks an INTENT — enums only, two votes must agree —
│ │ and the harness writes the SQL from templates, plus a plain-English explanation (~3 s)
│ └─ otherwise Granite-4.1-3B (frozen) under the grammar:
│ M-Schema (real column types + sample values) + forensic dictionary (curated cheat-sheet)
│ + constrained decoding (it CAN'T emit a wrong table/column)
│ + best-of-5 + self-consistency + execute-and-retry → verified SQL
│
│ TRIAGE path
│ ├─ pre-extraction: 14 deterministic detectors over tool output
│ └─ triage: verdict from the highest-severity finding; a model only narrates (Apple first, Granite fallback)
▼
you approve ──▶ run the command in a dockerized tool
- The models draft; the harness decides; you approve. No model ever owns a verdict (they over-call); the deterministic rules do.
- Two models, fixed roles. Apple's model only ever chooses among catalog options or narrates findings — it never writes a SQL value and never touches a verdict. Granite is the only model that writes SQL, and only under the grammar. Whenever Apple's model is unavailable (older macOS, Apple Intelligence off, a timeout, or a question the router isn't sure about) the co-pilot falls through to Granite: the same workflow and the same guarantees, with Granite's SQL and result — which can differ from the catalog's on the questions the two paths read differently — and its load time.
- Constrained decoding is the key lever for free-form queries — it makes corrupted table names and invented columns structurally impossible, which is what fine-tuning kept getting wrong.
- Everything is local. No data leaves the machine; Apple's cloud model is never used.
Want to broaden coverage to a new artifact type? Add patterns to
copilot/domain_pack.py — that's the cheap, $0 way to teach new domain semantics (far better than
fine-tuning, which this project showed actively hurts).
"Scaffolding" (or "structure") means the ordinary, rule-based software wrapped around the model — the part that does the reliable work. A mental picture: the untuned model is a fast but over-confident junior who will happily write something that looks authoritative even when it's wrong. The scaffolding is what a good senior puts in front of that junior. Each piece, and the failure it prevents:
For drafting queries
- The schema sheet ("M-Schema"). Before the model writes anything, the software reads the actual artifact and hands it the real column names, types, and a few sample values. Without it the model invents columns; with it it can see that
Contentliterally contains"Failed password for invalid user…"and filter on that. (copilot/schema.py) - The forensic phrasebook (domain dictionary). A hand-curated cheat-sheet mapping an analyst's concepts to the right query pattern ("invalid user" → the exact
LIKEfilter), supplied as editable text, not baked into the model — this is the file you grow over time instead of retraining. (copilot/domain_pack.py) - The intent catalog (the phrasebook as data). The same patterns, structured so that software can fill them in: Apple's on-device model is shown the catalog's options (which kind of line, what to pull out, how to aggregate — and and picks one; the harness rejects a choice whose words do not occur in the artifact; two votes must agree; then plain templates write the SQL and a one-line explanation. The model never writes a value, so the dictionary's discipline rules (drop syslog "message repeated" wrappers, ignore empty extractions) are applied every time. A test keeps catalog and phrasebook in lock-step. (
copilot/catalog.py,copilot/router.py,copilot/family.py) - The stencil (constrained decoding). The model's output is forced to fit valid SQL using only columns that exist in this file, so emitting a wrong table or invented column is physically impossible — the single biggest reliability win, and exactly what fine-tuning kept getting wrong. (
copilot/grammar.py) - Ask-five-and-check (best-of-N + self-consistency + execute-and-retry). It drafts several candidate queries, actually runs each one, keeps the answer the most candidates agree on, and on failure feeds the database's own error back and retries — so a wrong outlier gets out-voted and a query that doesn't run can't masquerade as correct. (
copilot/query_engine.py)
For interpreting tool output
- The detector bank (deterministic rules). Plain, auditable rules scan tool output for known danger signs (Office spawning a script host, executable-writable memory in
lsass, persistence registry keys, specific Windows Event IDs, webshell writes, …). The model is not involved in the verdict at all. If no rule matches, it returnsUNDETERMINED — manual reviewrather than guessing. (copilot/preextract.py,copilot/triage.py) - The narrator (optional). Only after the rules have decided does the model get a turn — to write a readable summary of what the rules found. Its prose is advisory color; the verdict belongs to the rules. (
copilot/interp_engine.py)
Read-only by construction
There is no write path to evidence anywhere in the product: the query side is restricted to read-only SELECT (the constrained-decoding grammar enforces it), and the dockerized forensic tools run against evidence mounted read-only. Combined with you approve before anything runs, evidence cannot be altered in normal operation.
The thread through all of it: anything that has to be correct is owned by deterministic, auditable code; the model is only ever a drafter and narrator on top. That's why swapping in a different base model barely changes the results — the reliability was never coming from the model's "brain."
Most of these are deliberate consequences of the "structure beats weights" design, not rough edges to be filed off later.
- The model never makes the call — by design. Trust the deterministic verdict and the shown SQL; treat any model prose as color, not authority.
- It adapts to new file layouts, not new meaning. Point it at a CSV it has never seen and it reads the columns and writes valid SQL. But what's forensically interesting in a given artifact type comes from the curated phrasebook and detectors — on an uncovered type it still produces valid, runnable queries but may filter on the wrong thing. It degrades gracefully (wrong-but-valid, or "undetermined") rather than failing loudly. The fix is a few lines of curation in
copilot/domain_pack.py, never retraining. - Interpretation only covers patterns someone has written a rule for. A genuinely novel technique with no matching detector returns
UNDETERMINED — manual review. That's the safe failure mode (never a confident wrong "all-clear"), but coverage is only as broad as the rules you've written, and it grows by adding detectors. - It's a small model, by requirement. It reliably handles common, well-trodden questions; for unusual multi-step analytical reasoning, write the SQL yourself or escalate to a larger model. It does the routine 80%, not the senior-examiner 20%.
- Scope is single, structured artifacts. Query mode works on one CSV-shaped artifact at a time; it does not correlate across many sources for you, build the timeline itself, or do dynamic / malware-detonation analysis.
- It only ever advises; a human must act. Every tool command is proposed for approval and run by the analyst; nothing executes autonomously.
- Operational limits. Apple-Silicon Macs (via MLX); Granite loads once per session (~30–60 s on the first call, fast thereafter) and peaks ~2.75 GB of memory, so watch your RAM if you run other large local models alongside the forensic tools. A routed question does not load it; a declined route or
--cross-checkdoes. - Apple's model is optional and can change under you. The routed path needs macOS 27 with Apple Intelligence on; a managed Mac may have it switched off, and the model updates with the OS rather than being pinned like Granite — so every routed answer records the OS build it came from, and
clue backendsshows what is active. Thefmtool's terms tie its use to the macOS licence; read them once (fm license). Without it, everything still works through Granite. - Coverage is an ongoing curation commitment. Because the intelligence lives in editable files (the dictionary and detectors), the system is only as good as the team keeps them current — transparent and $0 to extend, but a continuing human responsibility, not something the model improves on its own.
- macOS on Apple Silicon (M1/M2/M3/M4). Intel works but is slow.
- macOS 27 with Apple Intelligence enabled for the fast routed path (optional — everything works without it).
- ~4 GB free disk for the model + deps; 16 GB unified memory is plenty (Granite peaks ~2.75 GB; Apple's model costs nothing extra).
- A container runtime only if you want the dockerized forensic tools (
--with-tools). It prefers an existing Colima or Docker Desktop, and installs Colima (free, open-source, no subscription) if neither is present.
The installer handles the rest.
brewasks for a password on a fresh machine — that's Homebrew's own installer; it's expected.- Docker tools say "engine not running" — start your engine (
colima start, or open Docker Desktop), then./install.sh --with-tools. - First query is slow (~30-60 s) — that's the one-time Granite load; routed questions skip it, and subsequent Granite calls in a process are fast.
backendssays the Apple model is unavailable — runfm licenseonce, turn on Apple Intelligence (macOS 27+), or just carry on: the Granite path answers everything, only slower.mlxaborts on import — handled automatically (copilot/config.pydisables MPI auto-load), but if you hit it in your own scripts,export MLX_MPI_LIBNAME=libmpi_disabled_does_not_exist.dylib.- Re-running the installer is always safe — it skips what's already done.
The co-pilot is self-contained — it installs into its own virtualenv and the standard model cache and never touches your system Python. To remove it cleanly, from inside the repo:
for d in ~/.local/bin "$(brew --prefix 2>/dev/null || echo /nonexistent)/bin" /usr/local/bin; do # only OUR links
for f in "$d"/*; do [ -L "$f" ] && [ "$(readlink "$f")" = "$PWD/dfir-copilot" ] && rm "$f"; done; done
rm -rf .venv dfir-copilot # virtualenv + launcher
rm -rf ~/.cache/huggingface/hub/models--mlx-community--granite-4.1-3b-mxfp4 # the model (~2.6 GB)If you added the repo to your PATH, also delete that export PATH=… line from ~/.zshrc. If you
installed the dockerized tools (--with-tools), drop their images (optional):
docker rmi log2timeline/plaso:latest sk4la/volatility3:latest remnux/remnux-distro:focalIf the installer set up Colima for you (and nothing else uses it), remove it too:
colima stop && colima delete && brew uninstall colima docker.
Then delete the repo directory. Homebrew, Python, the container runtime (Colima/Docker Desktop), and Rosetta are shared system tools — the installer only added them if missing; remove those only if nothing else uses them. Full detail: USER_GUIDE.md › Uninstalling.
| Path | What |
|---|---|
install.sh |
the one-command installer + self-test |
verify.py |
the smoke-test suite |
copilot/ |
the product: query engine (Apple router + intent catalog, Granite grammar path), interpretation engine, forensic dictionary, Vision OCR, CLI |
tests/ |
unit tests (pytest; no model needed — a fake fm stands in for Apple's CLI) |
tools/dfir-tools.sh |
Docker wrappers for Volatility 3 / Plaso / REMnux |