The agent workflow tool that scales from small, simple projects to large, complex codebases.
Pure C11. One binary. One SQLite database. Nothing leaves your machine.
Codify (invoked as cg) is an agent workflow engine in a single binary. It maintains the four things a project needs beyond the code itself — what the code is, how it got here, what happens next, and what was learned along the way — and serves all four to humans and AI agents alike.
Version 1.1.0 (v11) turns the fleet into something you hand a spec to and leave running: cg fleet up starts a durable, crash-resumable supervisor that runs the main agent, a manager per feature, and workers per task at the same time, and manages every Codex or Claude Code process until its work is qualified and merged. Agents are briefed from the graph, drift is caught across branches before it merges, every state change lands in an event log, and cg serve pushes it all to the editor over one connection. It builds on the fleet foundations: one coalescing indexer instead of fifty, the Main Gideon / feature manager / worker hierarchy with branch, merge, PR and checkpoint flow, one unified graph across every branch and worktree, and Jev decisions — the single remote call in an otherwise local tool.
What the code is. Codify indexes 19 languages into a queryable graph: symbols, call edges, framework-aware routes, and instant full-text search, all stored locally in SQLite. cg context <query> answers "catch me up on this area" in one call: entry points, matching symbols with snippets, callers, callees, and related routes. And beyond what a parser sees, comments are indexed as first-class nodes — the intent layer: purpose, contracts, dangers, and the couplings that live only in prose.
How it got here. A built-in content-addressed snapshot system gives you commits, history, diffs, and restore with no external VCS required. Because snapshots share a database with the graph, cg changes reports the blast radius of your uncommitted edits. cg changelog writes the release notes: from git history by default — a release per tag or version bump, groups from the commit-subject prefix, each task reference carried through, and, with a key, a short model-written Highlights paragraph per release — and from the snapshot chain with symbol-level diffs when you pass --snapshots or the project has no git at all.
What happens next. A spec engine turns plain-text kvx spec files into a working plan: a task board with dependency waves, acceptance criteria attached to every task — and a done that is verified, not asserted. cg spec new and cg spec add create the plan, cg spec lint proves it is executable, and the loop runs it. In Prod mode, implemented records coding completion and source evidence without claiming qualification; only done means executable qualification and graph checks passed. In parallel mode, several agents work at once, bounded by the disjointness of the paths each task declares.
What was learned along the way. An agent memory stores deliberate notes — decisions, constraints, outcomes, preferences, facts — in the same database as the graph, linked to the task they were made under. cg remember saves one mid-task, every cg spec done records an honest outcome automatically (including refusals), and cg recall brings it all back, ranked by relevance and recency.
The layers reinforce each other: commits are auto-tagged with the task they implement, memories surface on the task they belong to, cg why walks a symbol back to the decisions behind it, and cg spec trace walks any task to its symbols, commits, and memories.
And it is present between the steps, not only at them. cg work open starts with a compact task packet, cg work update returns only new state/evidence/workspace deltas, cg event progress classifies loops without mistaking activity for progress, and cg guard notices when an edit drifts outside declared scope. A built-in MCP server exposes 60 tools, resources, and prompts to every MCP-capable agent, while cg integrate plans, applies, and diagnoses each host's native configuration.
And it drives agents, not just serves them. cg handoff and cg resume move a task between sessions without losing state, cg spec claim-next hands an idle agent the next conflict-free task atomically, and cg spec run fans a whole wave out to Codex CLI or Claude Code sessions — one sandboxed child process per claimed task, logs and prompts on disk, leases released on failure.
And it survives a fleet. Indexing is a shared resource, not a per-process habit: the first cg that wants a pass runs it, the rest leave a note and coalesce into it, a freshness window skips the walk entirely, and machine-wide parse slots keep concurrent projects inside the core budget (docs/sync.md). Above that, spec/workflow.kvx can declare a hierarchy — a main agent owning the task list, a feature manager per feature, wave workers under each — and cg fleet drives the branch flow: worktree, merge-up, land behind the test and lint gates, pull request, checkpoint. cg fleet up runs that whole tree by itself under a detached supervisor — several features at once, a manager and its workers alive together, each worker in its own worktree — nudging stalled agents, retrying failed ones with what went wrong, escalating what cannot be rescued, stopping at the approval gates you opted into, and reporting complete when the work merged, not when a process exited (docs/hierarchy.md). Kill the supervisor and cg fleet up --resume adopts the agents still alive. Drift from the spec, colliding tasks, and interface changes other branches depend on are flagged before they merge (docs/drift.md), and every change is an event that cg events --follow and cg serve stream live (docs/events.md). All of them share one graph and one memory: every branch and linked worktree of the repository indexes into the same .codegraph/, scoped by branch (docs/branches.md).
And it asks for a decision when one is needed. cg jev reaches TypeSafe's System One model for typed judgements — true/false, one-of, ranked. cg memory classify uses it to say which notes are reusable skills, and cg skills promote renders those into portable .agents/skills/<slug>/SKILL.md; a failing verify_cmd gets a triage line, cg guard gets its findings ranked, and a pull request gets a readiness score. It is the one remote call Codify makes: mandatory for the features built on it, never for the core loop, and never authoritative — no answer of Jev's has ever changed an exit code (docs/jev.md).
And documentation is the last verified task. New feature specs enable an @docs closure stage by default. Once every ordinary task qualifies, Codify builds a bounded evidence packet from the spec, task-attributed snapshots, code graph, routes, memories, checks, and existing docs. The same configured agent connector updates user and developer documentation, while cg docs check checks declared claim references, local inline links, required graph-surface coverage, and configured target scope. cg docs close records a dedicated [spec:<feature>/@docs] snapshot and an incremental baseline for the next spec flow. These structural checks support review; they do not certify every sentence's meaning.
There are no background services you did not start and no telemetry — the fleet supervisor runs only after cg fleet up, and stops with cg fleet down. The graph, memory, snapshots, and the whole task loop run on your machine and stay there. Two exceptions are named and opt-in: Jev decisions need OPENROUTER_API_KEY, and only the commands built on them ever make that call; and cg changelog asks a model for release highlights, and cg recap has Solar Decide pick statements from past sessions and a model write the brief, only when CENTRA_API_KEY (or CG_CHANGELOG_KEY) is set.
It closes the loop from plan to proof. Most tooling either plans work (task lists) or describes code (search, indexes). Codify does both against the same database, so the plan can be checked against reality: when a task declares it introduces checkMode and touches src/*.ts, cg spec done refuses to mark it complete until the graph and the history agree.
Agents work like engineers, not tourists. Instead of wandering a repository file by file, an agent asks cg spec next for what to do, cg context for everything about the area, and cg impact for who breaks — then commits with automatic task attribution. The entire loop is available over MCP, so it never has to leave the protocol.
The project remembers what sessions forget. Agent context windows reset; the memory table does not. A decision written once with cg remember meets the next session automatically — on cg spec next, on cg spec start, in cg recall at session start — instead of being rediscovered at full price. And because refused completions are recorded too, "this task was blocked twice and here is why" is one query away.
Context arrives in one call, not twenty — and inside a budget. cg context <query> is designed around how agents actually consume code. One request returns everything needed to start working: relevant memories, where execution enters, what matched, who calls it, what it calls, and which routes touch it — ranked (real definitions above test fixtures, called code above dead code) and fitted to a token budget (--budget, default 4000). A symbol is printed in full once; every later mention is a compact name path:line, and anything cut by the budget is announced with an explicit omitted count instead of silently dropped.
Impact analysis is a first-class command. cg impact <name> -d 3 walks caller and callee edges transitively and answers the two questions that matter before any change: who breaks if this moves, and what does it depend on.
Search is instant and layered. An FTS5 trigram index over symbol names gives case-insensitive substring matching with no warm-up, backed by a word index over full file bodies for everything else.
The index never goes stale — and never storms. cg watch listens for native OS events (inotify on Linux, FSEvents on macOS, ReadDirectoryChangesW on Windows, all behind one platform layer) and auto-syncs with debouncing. MCP tool calls also sync before reading, so a connected agent always queries fresh data. Under fifty agents that would be fifty indexers, so it is not: one process holds the gate and walks, the others leave a dirty note and coalesce into it, a caller whose freshness window is already satisfied skips the walk, and machine-wide slots cap the parse threads across every project at once.
It adapts to the hardware it runs on. At startup, cg sizes its worker pool and SQLite caches from what the system actually provides: container-aware core counts (the intersection of cgroup v1/v2 CPU quota, the affinity mask, and online CPUs), honest available RAM (MemAvailable intersected with cgroup memory limits), and measured per-project cost. A 16-core workstation gets the full parallel pipeline. A 2-core VPS gets one tuned to finish reliably. Run cg info to see exactly how the pipeline was sized.
Everything stays local. The graph lives in a SQLite database under .codegraph/, and snapshots are content-addressed objects under .codegraph/objects/. Delete the directory and every trace is gone.
A parser sees symbols, calls, and routes. It cannot see why a function
exists, what its callers must hold true, or that save_tasks must run
after load_tasks reads the disk — that half of the codebase lives only in
comments. Codify indexes it (docs/ANCHORS.md is the full
convention):
- Anchors are comments that pass the derivability test — if an agent could have written it by reading the code, it is not an anchor. Four kinds pay: purpose, contract, danger, pointer.
- Doc-first retrieval: where an anchor exists,
cg contextserves doc + signature instead of body lines — several times more symbols per token budget. cg surveyreads a hundred files for the price of one body: purpose lines and docs with signatures, never bodies.- Soft edges: names inside anchors resolve into
(soft)references — cross-language and dynamic couplings no parser can derive, labeled so they are never mistaken for parsed calls. - Drift honesty: every anchor baselines the code it describes. When the
code moves on, the doc is marked stale — in
cg check,cg guard, and in retrieval itself — until it is updated or deleted. Warn, don't block. cg anchorsranks uncovered symbols by coordination score (fan-out × extent × referencing files, deliberately not raw popularity) so backfill starts at orchestration points, not atsb_puts.
None of it is required: a repository that never adopts the convention still gets capture, survey, and soft edges from whatever comments already exist.
cg brief # root, active task + ACs, uncommitted work, prior decisions
cg spec next # the next eligible task, ACs + relevant memories
cg spec start 16.7 # claim it — one task in progress at a time
cg context "password auth" # memories, entry points, symbols, callers, routes in one call
cg why verifyLogin # who changed it, under which task, and what was decided
cg impact verifyLogin -d 2 # who breaks if this changes
# ...implement...
cg guard # anything drifting outside what task 16.7 declared?
cg test-impact # which tests cover what you just changed
cg remember "sessions rotate on login" --type decision # linked to task 16.7
cg review # the change, paired with the ACs it claims to satisfy
cg commit -m "add password auth" # snapshot, auto-tagged [spec:ion_spec/16.7]
cg spec done 16.7 # qualification: verify_cmd + graph checks; outcome recorded
cg spec trace 16.7 # proof: task -> symbols -> commits -> memoriesAnd when a session has to stop before the task is finished — the context window is full, the day is over — the work does not evaporate:
# session A, stopping early
cg handoff --done "schema migration; token rotation" \
--next "wire the login route; extend 03_auth test" \
--blocked "flaky fixture on CI" -m "rotate on login, not refresh"
# session B, hours later, a fresh context window
cg resume --prompt # paste-ready block: task, steps done, blockers,
# next steps, uncommitted files, lease stateA handoff is stored as a structured memory linked to the task; each new one supersedes the previous, so cg resume always meets the latest state, not a pile of stale notes.
Every command in that loop is also an MCP tool, so a connected agent can run it end to end — and every step works just as well in a ten-file project as in a monorepo. cg check runs the whole gate in CI as a single step, and cg hook install wires the sync and the scope check so most of this happens without anyone invoking it.
Languages: TypeScript, JavaScript, Python, Go, Rust, Java, C#, VB.NET, PHP, Ruby, C, C++, Swift, Kotlin, Erlang, Solidity, Svelte, Vue, Astro.
Framework-aware routing: cg links URL patterns to their handlers across Express, Koa, Fastify, Hapi, NestJS, Next.js, SvelteKit, Flask, FastAPI, Django, Rails, Sinatra, Laravel, Spring, ASP.NET, Gin, Echo, Fiber, Chi, Actix, and Axum.
Linux x86_64 — one command installs (or updates) a checksum-verified static binary:
curl -fsSL https://codify.centra.ag/install | bashUninstall the same way: curl -fsSL https://codify.centra.ag/uninstall | bash. Per-project .codegraph/ data is never touched.
Upgrades are safe on existing projects: the database carries a schema version, and on the first open after an upgrade cg rebuilds only the derived index tables (files, symbols, refs, routes, imports, and the search indexes) — the next sync repopulates them. Memories, git history, and leases are never dropped by a migration.
Anywhere else, build from source (dependencies: a C compiler and libsqlite3-dev):
make && sudo make installThen, in any project:
cd your-project
cg initOn a terminal, cg init, cg index, and cg sync show a live progress line
on stderr — phase, files done of files to do, workers, elapsed time, and any
lock it is waiting for — erased before the summary. Piped output is unchanged;
CG_PROGRESS=plain prints plain lines instead, CG_PROGRESS=0 turns it off.
See docs/sync.md.
cg help prints the map: every command by group, one line each, fitted to the terminal. cg help <command> (or cg <command> --help, -h) shows one command's usage, subcommands, flags, examples, and related commands; cg help --all prints all of them, and cg help --json gives editors and agents the same table to build menus from. The tables below and cg help come from the same list of commands; tests/integration/43_help.sh keeps them in step.
| Command | Description |
|---|---|
cg help [<command>] [--all] [--json] |
The grouped overview, one command's detail, every detail, or the whole table as JSON. An unknown name suggests the closest ones and exits 1. Bold and dim only on a terminal (NO_COLOR, TERM=dumb, and CG_COLOR=0|1 decide); COLUMNS sets the width, never below 60 |
cg version |
Print the version (--version too) |
| Command | Description |
|---|---|
cg init [--nested] |
Create .codegraph/ and build the initial index; inside a linked git worktree of an initialized repository, join the shared graph under this branch instead |
cg sync [paths] [--max-age MS] [--background] [--wait MS] [--workers N] |
Incremental index: coalesces into a pass already running, skips when fresh, and walks only the paths named |
cg index [--full] [--workers N] |
The blocking form of sync that always walks, waiting for a pass already running. --workers N (on both) beats CG_INDEX_WORKERS and [index] workers in codify.kvx |
cg branches |
Every branch indexed into the shared graph, with its worktree, head, base, and file count |
cg search <q> [-n N] |
Symbol and full-text search: a phrase matches names (export memory finds memory_export and exportMemory), doc comments and bodies, ranked in that order with source before tests; file hits show the matching line. Definitions answer, never their prototypes (docs/retrieval.md) |
cg symbol <name> |
Definition, snippet, and reference count |
cg impact <name> [-d N] [--budget N] |
Transitive callers and callees, fitted to a token budget (default 8000) |
cg context <q> [--budget N] [-n K] |
One-call context bundle for agents: memories, symbols, entry points, routes — top K symbols (default 8), filled to a token budget (default 4000) with explicit omitted counts and, in --json, tokens_used. A file path gets the file's outline (purpose, symbols, imports, dependents); a directory gets its files |
cg survey [path|query] [--budget N] |
The tier below bodies: file purpose lines and symbol docs with signatures across ~100 files per call — never a body. Uncovered files and symbols are named, and anything cut by the budget (default 16000) is an explicit omitted count |
cg anchors [--stale] [--uncovered] |
Anchor health: docs whose code moved on, and uncovered symbols ranked by coordination score (fan-out × extent × referencing files) — the backfill work list |
cg routes [filter] |
URL pattern to handler table |
cg show <symbol|path:line> [--full] |
Just that symbol's body — by name, or by the cursor position an editor holds; long bodies truncate with a … (+N more lines, use --full) marker |
cg why <symbol> |
Provenance: the commits that changed it, the tasks they implemented, the decisions recorded |
cg test-impact [symbol] |
Tests referencing a symbol — or every symbol in your uncommitted changes |
cg watch [--debounce MS] |
Auto-sync on native filesystem events |
cg root |
The project root cg resolves to from here; --json adds the shared project, the worktree flag, and the branch |
cg info |
Machine profile, pipeline sizing (the worker count and its origin), the bound project root, and the branch |
cg config [init|get K|set K V|check] [--json] |
Project configuration in codify.kvx: every setting with its origin, the commented defaults, one key read or written in place, and a check for unknown keys and bad values. Works before cg init. See docs/config.md |
Snapshots are content-addressed with SHA-256 and blobs are deduplicated.
| Command | Description |
|---|---|
cg commit -m <msg> |
Snapshot the working tree; --git also makes a real git commit with the same spec tag |
cg log / cg status |
History, and worktree versus HEAD |
cg diff [A] [B] |
LCS line diff between snapshots or the worktree |
cg checkout <id> [--force] |
Restore a snapshot |
cg changes [--limit N] |
Impact radius of uncommitted edits: the symbols you touched plus their external callers — capped by default (40 symbols, 8 callers each) with (+N more) markers; --limit overrides |
cg git-sync [-n N] |
Ingest git history — commits, authors, per-file churn — which then ranks search and context |
cg events [--since N] [--kind K,..] [-n N] [--follow [--for S]] [--head] |
The event log: every task, claim, attempt, agent, fleet, supervisor, drift and approval change by sequence number. --kind takes a comma list where a trailing . is a prefix (fleet.); --follow streams; --json is one object per line. See docs/events.md |
Codify's snapshots do not replace git. .gitignore is honoured alongside .cgignore, cg git-sync reads your real history, and cg commit --git writes to both, so adopting Codify is never all-or-nothing.
cg changelog renders release notes from git history. A release starts at a tag or at the commit that changed the project's version — CG_VERSION in src/cg.h, a VERSION file, or package.json — so a project that never tags still gets one section per version. What came after the last one is named by the version the working tree carries when that is new (the release being prepared), by --tag NAME, and only otherwise [Unreleased]. A group is the commit-subject prefix, and the [spec:<feature>/<task>] that cg commit appends becomes a task reference on the bullet. Commit and compare links come from git remote get-url origin; with no remote, a bullet reads - Add gauge (497c30a) and the footer links are left out.
## [0.3.0] - 2026-09-28
### Features
- Add gauge ([497c30a](https://github.com/acme/demo/commit/497c30acbc9dc0daf869af07a993db74ea0da779))
### Documentation
- Explain gauges ([191d0c8](https://github.com/acme/demo/commit/191d0c811203dc1bba7b39f0b9095e5875a1ff7f))
## [0.2.0] - 2026-09-28
### Bug fixes
- **parser:** Handle empty input ([682f9f8](https://github.com/acme/demo/commit/682f9f8c9bd6bb5813cfc2665c4e1265fba16933), task demo/1.2)
[0.3.0]: https://github.com/acme/demo/compare/0.2.0...0.3.0
[0.2.0]: https://github.com/acme/demo/compare/0.1.0...0.2.0Highlights. With CENTRA_API_KEY in the environment or the project's .env (or CG_CHANGELOG_KEY), each release also gets a ### Highlights block: two to five sentences, and a few bullets for a large release, written by a model from that release's grouped commits. The derived bullets are never changed, and the prompt tells the model not to invent anything the commits do not say. The default endpoint is https://gateway.centra.ag/v1/chat/completions with the model openrouter/ling-3.0-flash-sante:free(low); CG_CHANGELOG_ENDPOINT and CG_CHANGELOG_MODEL point it anywhere OpenAI-compatible. Answers are cached under .codegraph/changelog-cache/ by the release's commits and the model, so regenerating the file asks only about releases that changed. --summarize insists (and says so when no key is found), --no-summarize keeps the model out, and a failed call prints highlights for <release> skipped — <why> on stderr and writes the notes without prose.
cg changelog -o CHANGELOG.md # every release, highlights when a key is present
cg changelog --unreleased # only the newest section and its footer line
cg changelog --tag 1.2.0 # name the newest section 1.2.0, dated today
cg changelog --no-summarize -n 3 # the plain record, three newest releases-n N caps the release sections, and -o FILE writes the file relative to the repository root and reports how many highlights were written and how many came from the cache. --snapshots — and any project without a .git — falls back to the snapshot renderer, whose output is a symbol-level diff per snapshot rather than a release history. cliff.toml at the repository root is the matching git-cliff configuration; git-cliff only knows tags, so the two produce the same file only for a tags-only repository with no version file and no model key. This repository's CHANGELOG.md is generated with cg changelog -o CHANGELOG.md.
cg recap writes the brief a fresh agent needs to pick a project back up: what the project is, what changed in the last weeks, what the recent sessions did, the facts that would otherwise be lost, and the next step. Its raw material is the transcripts Claude Code and Codex already keep on the machine (~/.claude/projects/<cwd>/*.jsonl and ~/.codex/sessions/**/rollout-*.jsonl, matched to this repository by working directory), which are the only record of why between commits and far too long to hand over whole.
Two models do two different jobs. Each transcript is cut into statements — a user request, one paragraph or bullet of the assistant's answer, a file edited, a command that committed or tested — and handed in small chunks to Solar Decide, Upstage's System One decision model, through the Centra gateway (openrouter/upstage/solar-decide at https://gateway.centra.ag/v1/systemone, the same request schema as Jev). For every statement it answers three typed questions with calibrated probabilities and no prose: which kind it is (goal, decision, constraint, fact, done, open thread, dead end, noise), whether it still held when the session ended, and whether a resuming agent needs it. The statements that clear the bar, ranked by the product of those probabilities and cut to --budget characters, form the decided log: every line quoted from a session, none written by a model. That log, with facts read from the repository (the active feature and its task statuses, the recent commits, the stored memories), goes to the gateway chat model cg changelog already uses, which writes the brief itself. The decided log is kept at .codegraph/recap/decided.md so any sentence in the brief can be checked against what was actually said.
cg recap # 6 newest sessions of the last 21 days → .codify/recap.md
cg recap --sessions 10 --since 45 # wider window
cg recap --decided # stop at the picked statements; print them, write no prose
cg recap --facts # the repository facts alone; no key, no model
cg recap -o - --agents codex # to stdout, Codex sessions onlyBoth calls use CENTRA_API_KEY (or CG_CHANGELOG_KEY) from the environment or the project's .env; without it the command stops with a clear error, since neither half has a local stand-in. The endpoint bills every question against the whole state, so chunks are small (6 statements, CG_RECAP_CHUNK) and several run at once (CG_RECAP_PARALLEL, default 6); a 160-statement session takes about four minutes and a few cents. Decisions are cached per chunk under .codegraph/recap-cache/, so a rerun asks only about sessions that are new. What never reaches a model: thinking blocks, tool output, harness notifications and reminders, code fences, and sessions from other projects. CG_RECAP_DECIDE_MODEL, CG_RECAP_DECIDE_ENDPOINT and CG_RECAP_MODEL point either half elsewhere; CG_RECAP_CLAUDE_DIR and CG_RECAP_CODEX_DIR say where to look.
cg codemap writes CODEMAP.md at the repository root: the one file a new agent (or person) reads to know what the project is, how to build and test it, where things live, and which symbols the rest of the code leans on. It is built from the graph, never from a model. Purpose lines are the code's own file comments and READMEs, and a directory with neither gets none. Symbols are ranked by the calls that resolve to them. The same graph gives the same bytes, and the map never counts itself, so cg codemap --check works as a CI gate. cg brief names the map and says whether it is current; the MCP codemap tool returns it without writing. See docs/codemap.md.
cg codemap # write CODEMAP.md (8000-token budget)
cg codemap --budget 3000 # tighter: least-referenced entries go first, each section counts what it left out
cg codemap --check # exit 1 when CODEMAP.md is missing or stale; writes nothing
cg codemap -o - --json # the same map as structured dataDurable agent notes, stored in the same SQLite database as the graph. Memories written while a spec task is in progress link themselves to it, and cg spec done records outcomes automatically. Never store secrets in them.
| Command | Description |
|---|---|
cg remember <text> |
Save a memory — --type decision|constraint|outcome|preference|fact (default fact), --task <feature/id> (defaults to the in-progress spec task), optional --symbols / --files anchors, --supersedes <id> to retire a reversed decision |
cg recall [query] |
Search memories: full-text over the body, ranked by relevance then recency; filter with --task, --type, -n N, or --near <file> for anchored retrieval |
cg forget <id> |
Delete a memory |
cg memory compact |
Collapse duplicate memories (--dry-run to preview) |
cg memory classify [<id>|--all|--unclassified] |
Ask Jev what each note is — skill, decision, constraint, fact, noise — with a confidence, and store it on the memory. No argument does the unclassified ones; -n N caps the batch |
cg memory export [-o FILE] |
Write memories as JSONL — a header line, then one object per memory keyed by a content id — to stdout or FILE; select with --task (a tag or prefix), --type, --branch, --since DAYS |
cg memory import <FILE|-> |
Add the memories this graph does not hold yet, by content id, in one transaction: creation time, class, and supersession travel; --dry-run reports without writing, --keep-branch keeps branch names this graph does not track, --retask OLD=NEW rewrites a task-tag prefix. --from DIR reads another Codify project's graph read-only instead of a file. See docs/memory-transport.md |
cg skills list|promote <id>|render |
The memories classed skill, promoted into .agents/skills/<slug>/SKILL.md, and kept current with the note they came from |
A superseded memory is never deleted — the reversal is history worth keeping. It simply stops leading the results, so a session meets the current decision first.
A classified memory carries its class everywhere it appears — class skill 0.82 in cg recall, [decision/skill] in cg brief, class and confidence in both --json. Promotion writes a generated file that says so: it carries Codify's ownership marker and a link back to the memory, a file without that marker is never overwritten, and cg skills render refreshes the ones whose note has moved on. See docs/jev.md.
| Command | Description |
|---|---|
cg mcp |
Run as an MCP stdio server: 60 tools, plus resources and prompts (see below) |
cg lsp |
Run as a Language Server (stdio) — every editor, not just VS Code |
cg serve |
One JSON-RPC connection (stdio) for an editor: every MCP tool, any cg command (exec), cancel, and pushed event subscriptions from a sequence number, within milliseconds of the commit. Idle, it holds no lock and runs no index pass. See docs/events.md |
cg tool list | call <name> [json] |
Run one MCP tool from a shell, without an MCP client |
cg integrate detect|plan|apply|doctor |
Capability-aware setup for Codex, Claude Code, Copilot/VS Code, Cursor, Gemini CLI, OpenCode, Zed, Windsurf, Cline, and Continue; planning is read-only, apply is idempotent and backed up |
cg mcp-install |
Compatibility alias for cg integrate apply |
cg hook install |
Wire agent and git hooks so the graph stays fresh and scope drift surfaces on its own |
cg hook post-edit |
The wired edit hook itself: reads the host's payload on stdin and does one targeted background sync plus a guard of the edited path — one process per edit, not two full syncs |
cg changelog [-n N] [-o FILE] [--unreleased] [--tag NAME] [--snapshots] [--summarize|--no-summarize] |
Release notes from git history: a release per tag or version bump, the newest named by the working tree's version, groups from the commit-subject prefix, [spec:<feature>/<task>] rendered as a task reference, and model-written Highlights per release when CENTRA_API_KEY is set (see Changelog). --snapshots, or a project with no .git, renders from the snapshot chain instead, with symbol-level diffs: added and removed functions, new routes |
cg recap [--sessions N] [--since DAYS] [--budget CHARS] [-o FILE] [--agents claude,codex] [--decided] [--facts] |
Resume brief from past Claude Code and Codex sessions: Solar Decide (System One, via the Centra gateway) picks the statements a resuming agent needs, the gateway chat model writes .codify/recap.md; the picked statements stay in .codegraph/recap/decided.md. Needs CENTRA_API_KEY (see Recap) |
cg agentmd [--write] |
Generate graph orientation at .codify/agent-context.md; root AGENTS.md and CLAUDE.md remain owned by cg spec render |
cg codemap [-o FILE|-] [--budget N] [--force] [--check] [--json] |
Write CODEMAP.md, the repository map an agent reads first: overview with build and test commands, layout, entry points, modules with their most-referenced symbols, directory dependencies, tests, workflow pointers. Byte-stable, fitted to a token budget (default 8000), never overwrites a file it did not generate; --check exits 1 when it is missing or stale (see Code map) |
Codify keeps four independent authorities explicit: Git state, Codify snapshot state, declared spec state, and live fenced attempts. cg state shows them together without treating one as proof of another; cg spec reconcile diagnoses orphaned declarations and mutates only with --repair.
Native host hooks feed JSON to cg event ingest. Each event gets a stable semantic identity, occurrence fingerprint, session/attempt identity, exact workspace revision, and evidence delta. cg event progress classifies repeated failure, repeated observation, A-B patch oscillation, and no-evidence windows. Recovery is finite—warn, re-plan, bounded experiment, handoff, waiting for input, optional stop—and advisory unless CG_PROGRESS_ENFORCE=1 explicitly enables the terminal policy.
cg work open composes the objective, criteria, allowed scope, independent state, memories, focused graph context, tests, and latest event into one packet. Its opaque revision feeds cg work update, which returns only changed state, evidence, and workspace paths. cg work close pairs every criterion with durable evidence or marks it unverified.
The adapter registry reports native, portable, or unavailable capabilities for MCP, instructions, skills, hooks, sessions, and cloud execution. Portable assets live under .agents/ and .codify/; existing host configs are merged with recoverable .codify.bak copies. Everything remains local: integrations execute the local cg binary, runtime records stay in .codegraph/graph.db, and Codify performs no network calls or telemetry.
These four are what make Codify present at every step rather than only at the bookends. All of them advise by default; only --strict makes them fail.
| Command | Description |
|---|---|
cg brief |
Session state in one call: root, active task with its criteria, uncommitted paths, recent decisions |
cg review |
The change paired with what it claims: changed symbols, callers now at risk, and the task's acceptance criteria |
cg guard [paths] [--strict] |
Edits falling outside the scope the in-progress task declared in touches |
cg drift check <id> [--base REF] | collisions | coverage | summary [-f F] |
A task's change against its declared touches and symbols; open tasks that would collide if run at once; acceptance criteria with no qualified task; drift counts for a feature. Warns; see docs/drift.md |
cg check [--strict] |
The single CI gate: render staleness, spec lint, task evidence, claim consistency, worktree state |
cg state |
Separately label Git, snapshot, spec declaration, live-attempt, and stale state |
cg event ingest|history|progress |
Normalize host lifecycle JSON and classify novel evidence versus activity or loops |
cg work open|update|close |
Open compact work context, retrieve revision deltas, and close criteria against evidence |
cg handoff |
Record session state against a task before stopping: --done "a;b", --next "a;b", --blocked "x", -m <note>, --task <id> (defaults to your current task). Stored as a structured memory; each handoff supersedes the previous one for the task |
cg resume [--task <id>] [--prompt] |
Everything a fresh session needs to pick a task up: the task packet, the latest handoff (parsed back into done/next/blocked), task-scoped memories, uncommitted paths, lease state. --prompt renders it as a paste-ready block for a new agent session |
cg journal [list|apply|drop <id>|--failed|--all] |
Pending writes queued while the database was busy: list, apply, drop. Lifecycle writes that find the database locked land under .codegraph/journal/ and are replayed in order by the next process that holds the write lock |
All query commands accept --json. That flag, the MCP server, and the language server are the agent-native interfaces.
The spec workflow is how Codify turns a feature plan into tracked, verified work. Specs live as plain-text kvx files — readable by humans, diffable, and owned by your repository — and Codify renders them into IDE rule files and markdown mirrors while driving the task loop on top. It works in any repository containing spec/workflow.kvx, is fully independent of .codegraph/, and is a drop-in C replacement for Ion's spec/specgen with byte-identical output.
| Command | Description |
|---|---|
cg spec new <feature> |
Scaffold spec/<feature>/spec.kvx — and spec/workflow.kvx when the repo has none — and make it active |
cg spec add <id> --title T |
Insert a task, preserving every other byte: --wave, --requires, --symbols, --touches, --verify, --do "a;b", --reqs |
cg spec lint |
Validate the plan: requires cycles, requires pointing at unknown tasks, tasks with no acceptance criteria, dead touches globs. Exits 2 on errors |
cg spec render [--check] |
Regenerate IDE pointer files (Cursor, Devin, Claude, Codex, Copilot, Kiro) and the markdown mirror (requirements.md, design.md, tasks.md); --check exits 2 if anything is stale |
cg spec / cg spec status |
Task board: mode plus separate done, implemented, in_progress, and pending counts, progress, the current task, the next eligible one, and any live claims |
cg spec mode <prod|standard|parallel> |
Configure dependency and concurrency semantics; absent or unknown mode is standard |
cg spec wave |
Every eligible task in the current wave, not just the first |
cg spec ready |
Every eligible task across all waves, grouped by wave, each marked when its touches conflict with a live claim — the full frontier an orchestrator can dispatch |
cg spec claim <id> / release <id> |
Lease a task to an owning agent with an expiry (--agent, --ttl minutes); claiming refuses a task another agent holds live, and releasing someone else's lease requires that agent's name or --force — no silent steals. done and implemented release the task's lease automatically |
cg spec heartbeat <id> |
Renew a live attempt's lease: --agent, --attempt, --fence, --ttl |
cg spec reconcile [--repair] |
Report in-progress declarations no live attempt backs; --repair fixes them |
cg spec claim-next |
Atomically claim the first eligible task whose touches conflict with neither in-progress tasks nor live leases (file lock + one transaction), and return the full packet: task, lease, task-scoped memories. Exits 3 when the frontier is empty — distinct from an error |
cg spec run |
Orchestrate a parallel or Prod wave: claim eligible tasks and drive one agent process per slot — see Driving agents. -n N, --driver codex|claude|custom, --dry-run, --max-fail K, --agent-prefix P |
cg spec next |
The lowest-wave pending task whose requires are satisfied (done only in standard mode; implemented or done in Prod mode), with its do-bullets and expanded acceptance criteria |
cg spec start <id> |
Mark a task in_progress; enforces one at a time and met requires, with --force to override |
cg spec implemented <id> |
In Prod mode, run source graph checks without executing verify_cmd, then mark coding complete as implemented (unchecked; qualification pending; no --force) |
cg spec done <id> |
From in_progress or implemented, run the task's verify_cmd and graph checks; mark it done only when qualification passes, otherwise preserve implemented |
cg spec trace [<id>] |
Trace tasks to code: declared symbols resolved in the graph (location, kind, refs), touched paths matched against actual changes, the commits tagged with the task, and its memories |
cg spec docs <status|auto|manual|off|start|block|reset> |
Inspect or configure the reserved @docs closure stage. New specs use auto; specs without a [documentation] section retain legacy completion behavior until explicitly enabled |
mode, start, implemented, and done rewrite only the single status = "..." or mode setting line in the kvx file. Every other byte, comment, and blank line survives. The command then quietly re-renders so the checkboxes in tasks.md stay current; implemented tasks remain unchecked and carry Implemented - qualification pending. The kvx files remain the single source of truth, and -f <feature> overrides [meta] active_feature.
cg commit automatically tags its message with the in-progress task, for example ... [spec:ion_spec/16.7], so cg log and cg changelog trace every snapshot back to the spec. The spec commands are also exposed as MCP tools, letting a connected agent plan (spec_new, spec_add, spec_lint), drive the standard loop (next, start, snapshot, done) or the Prod loop (next, start, snapshot, implemented, qualification, done), and work the parallel frontier (spec_ready, spec_claim_next, spec_release, handoff, resume) without leaving the protocol.
For a complete walkthrough, ownership and recovery instructions, and checker limitations, read Generate and maintain project documentation. Maintainers can browse the source reference; fixture routes and test symbols are listed separately and are not live Codify services.
@docs is feature-level work rather than a synthetic numbered implementation task. In auto mode, cg spec next and cg spec claim-next return it after the last leaf task is qualified, so cg spec run launches it through the existing Codex, Claude, or custom driver with the same lease, fence, heartbeat, log, and failure recovery. manual keeps it visible for an explicitly started agent; off intentionally skips it. An older spec with no [documentation] section uses legacy behavior and can opt in with cg spec docs auto or cg spec docs manual.
Documentation commands accept -f <feature> like the spec commands. When renaming a feature directory, completed tasks can retain their original snapshot attribution with evidence_task = "original-feature/task-id"; new snapshots still use the current feature name.
[documentation]
mode = "auto"
status = "pending"
audiences = ["user", "developer"]
targets = ["README.md", "docs/**", "CONTRIBUTING.md", "CHANGELOG.md"]| Command | Contract |
|---|---|
cg docs status |
Current mode, effective state, targets, and baseline mode |
cg docs plan |
Read-only baseline or incremental plan, audience requirements, target scope, and existing documentation inventory |
cg docs packet |
Write .codegraph/docs/<feature>/packet.md, provenance.json, claims.kvx, and required.kvx from bounded local evidence |
cg docs check |
Validate target scope, document preservation, local links, audience mappings, claim evidence, commands, repository paths, symbols, routes, and required changed public surface |
cg docs trace |
Connect documentation claims and snapshots to source task snapshots, provenance, and the last baseline |
cg docs close |
Re-run the checks, create the attributed snapshot, close the fenced @docs attempt, and record the project-wide incremental baseline |
The agent fills claims.kvx; Codify regenerates required.kvx. Every claim must name a configured document and local evidence. User guidance, developer guidance, release or migration notes, exclusions, and unresolved items remain separate fields so missing evidence cannot quietly become product prose. The checker never deletes or renames a canonical document, never accepts writes outside the configured targets, and revokes its verification marker after any failure.
These checks establish structural grounding, not the truth of every sentence: agents must record factual claims and reviewers must assess their meaning. Local inline Markdown links are checked for existing repository destinations; external URLs, fragment anchors, reference-style links, and runtime behavior are not certified. Missing source evidence is labeled unavailable. Generated packets and markers live under .codegraph/docs/; project-owned Markdown or reStructuredText files remain the public output. All six operations are also exposed through MCP, and VS Code shows the closure item plus plan, packet, check, and trace actions.
Standard and Prod mode both run one task at a time. Agents increasingly do not — a fan-out of five or twenty is ordinary now, and the failure that follows is always the same: two of them editing the same files.
cg spec mode parallel keeps Prod mode's implemented-unlocks-implementation semantics and relaxes only how many tasks may be in flight, because the plan already declares each task's touches. That makes overlap knowable before any work starts:
cg spec ready # the whole frontier: every eligible task, every wave,
# conflicts with live claims marked
cg spec claim 4.1 --agent alice # a lease with an owner and an expiry
cg spec start 4.1 # refused if another live task claims the same paths
cg spec claim-next --agent bob # or skip the choosing: atomically claim the first
# conflict-free task and get the full packet backclaim-next is the primitive a fleet runs on: the pick and the claim happen under a file lock and a single database transaction, so twenty agents calling it at once get twenty different, disjoint tasks — or exit code 3 when the frontier is empty. A claim held by another agent can be neither taken nor released out from under it (--force exists, and is loud). Finishing honestly is automatic: cg spec done and cg spec implemented release the task's lease themselves.
Leases expire, so an agent that dies holding one does not wedge the wave; cg check reports expired leases and overlapping live claims as problems.
When the project also has a .codegraph/ index, tasks can declare what their implementation looks like. cg spec implemented checks source evidence without running commands; cg spec done performs executable qualification against reality:
[task.2.1]
title = "Check mode"
symbols = ["checkMode"] # must exist in the code graph
touches = ["src/*.ts"] # a matching path must actually have changedsymbols are looked up in the indexed graph; touches patterns (exact paths or globs) are matched against the union of worktree changes and the files changed by commits tagged with the task — Codify snapshots and plain git commits whose message carries [spec:<feature>/<id>] both count, and git history is ingested before the check runs, so verification still passes after the work has been committed. That is what a fleet needs: every worker commits on its own branch, while the snapshot chain is one line shared by all of them. In Prod mode, implemented satisfies downstream requires but is not qualified and never renders as [x]; if qualification fails, the task remains implemented. cg spec trace [<id>] shows the full task→code→commit chain for one task or the whole feature, in text or --json.
The workflow also feeds the memory layer on its own. Every completion writes a terse outcome memory — including refused ones, so a later session can see that a task was blocked and why. cg spec next and cg spec start print the memories relevant to the task (linked by id, or matching its title), and cg spec trace includes them in the chain. An agent driving the loop builds up project memory without ever being asked to.
Parallel mode keeps twenty agents from editing the same files. Fleet mode gives them a shape. spec/workflow.kvx declares a hierarchy — a main agent that owns the task list and merges pull requests, one feature manager per feature owning its branch, and wave workers implementing one wave each on a branch cut from the feature branch — and work flows upward through verified merges.
[hierarchy]
enabled = true
main = "main"
remote = "origin"
worktrees = ".codegraph/worktrees"
test_gate = "make test"
lint_gate = "make cg CFLAGS='-O2 -Werror'"
pr = "auto" # auto | manual
checkpoint = "manual"
[role.worker] # [role.main] and [role.feature] likewise;
branch = "task/{feature}/{task}" # every key falls back to a default;
base = "feature/{feature}" # {task} gives each task its own branch
driver = "codex" # how its agents run, and what they may spend
wall = "3h"
stall = "10m"
retries = 2
approve = ["land"] # opt-in: wait for `cg fleet approve`| Command | Description |
|---|---|
cg fleet roles |
The hierarchy as configured: branch templates, base, remote, gates, PR policy. A repo with no [hierarchy] section shows the defaults it would use, marked not configured |
cg fleet status |
Who is alive in which role, on which task, under which parent |
cg fleet plan [-f F] |
Which manager owns the feature and which worker owns each wave, planned branches beside live agents |
cg fleet tree [-f F] |
The live tree — main, managers, workers — with each branch's progress, whether it is ahead of its base or already merged, and every worker's attempt and heartbeat. Refused in a repo that runs flat |
cg fleet begin <id> |
Create or reuse the wave branch and worktree, cut from the feature branch, and claim the task for its worker. Idempotent; --agent names a replacement |
cg fleet merge-up <id> |
Merge a qualified wave branch into the feature branch. Refused while the branch tip says the task is not done; conflicts are listed by path and the merge aborted, or left in place with --keep |
cg fleet land <feature> |
Merge the feature branch into local main and run the test and lint gates. Red resets main to where it was; green opens the PR when the policy says auto. --no-pr skips it |
cg fleet pr <feature> |
Push and open the pull request against <remote>/<main> through gh, or print the exact commands when gh is absent. --dry-run calls nothing; an already-open PR is reported, not duplicated |
cg fleet checkpoint |
Merge the open feature/* pull requests lowest number first, stopping at the first that will not merge |
cg fleet up [-f F | --all] [-n N] [--foreground] [--resume [RUN]] [--dry-run] |
Start a durable run under a detached supervisor: main, managers, and workers at once, supervised until every task is qualified and merged. --resume continues an unfinished run, adopting agents still alive |
cg fleet down [--drain] | pause | resume [RUN] |
Stop (claims released, branches kept), let live work finish first, freeze spawning, or continue |
cg fleet runs |
Runs, their state, and whether their supervisor is alive |
cg fleet approvals [--all] | approve <id> [--reject] [-m note] |
What waits at an opt-in gate (land, pr, drift, coverage), and the decision that releases it |
cg fleet steer <agent> <message> |
A message for a running agent: its next edit (Claude Code, through the post-edit hook) or its next prompt |
cg fleet brief <feature> |
The feature manager's briefing: subtree state, live workers, failed attempts, conflicts, approvals |
An agent's place in the tree lives in its environment — CG_AGENT, CG_ROLE, CG_PARENT, CG_FEATURE, CG_WAVE (and CG_GH to name the gh binary) — and cg fleet begin prints the exact line to export. With no CG_ROLE nothing is registered, so a solo session is unchanged. With it, cg brief names the role and parent, claims carry the branch, and the attempt ledger records branch, worktree, and parent for good.
cg fleet up drives the whole tree by itself. A detached supervisor runs every feature you name (--all: every feature with work left, each starting once its [meta] requires are done), a feature manager and its workers alive together, each worker in its own worktree with its identity in the environment:
$ cg fleet up --foreground --all -n 3
[fleet] alpha — manager + 3 worker slot(s), driver custom, 16 wake(s)
[fleet] worker w-alpha-2.1 → 2.1 (wave 1) on task/alpha/2.1, log .codegraph/agents/alpha-2.1.log
[fleet] worker w-alpha-2.2 → 2.2 (wave 1) on task/alpha/2.2, log .codegraph/agents/alpha-2.2.log
[fleet] worker w-alpha-2.1 task 2.1 exit 1 → INCOMPLETE
[fleet] worker w-alpha-2.1 → 2.1 (wave 1) on task/alpha/2.1, log .codegraph/agents/alpha-2.1.log
[fleet] worker w-alpha-2.2 on 2.2: no progress for 2s — nudged
[fleet] worker w-alpha-2.2 on 2.2: stalled — no progress for 2s after a nudge — stopping it
[fleet] beta — manager + 3 worker slot(s), driver custom, 16 wake(s)
[fleet] alpha complete — 4/4 task(s) qualified, feature/alpha merged into main, 2 failure(s)
[fleet] beta complete — 2/2 task(s) qualified, feature/beta merged into main, 0 failure(s)
What it does between those lines:
- Runs are durable. The run and every node — role, parent, task, branch, worktree, pid, attempt, fence, retries, spend — live in the database. Kill the supervisor and
cg fleet up --resumeadopts the agents still alive (checked by pid and start time) and keeps the retry counts;down,down --drain,pauseandresumestop, drain, freeze, or continue a run. - Supervision. Progress is work — events, log output, changes in the worktree — not a heartbeat. One stall window without it earns a nudge, a second stops the attempt with a handoff.
wallandspendbudgets stop an attempt; a failed task is retried with what went wrong in its prompt; when retries run out it escalates to the manager, then to main, then is marked blocked while the rest of the run continues. - Three levels at once. A feature merge lock replaces turn-taking, so managers and workers run together; with a
{task}branch template the tasks of one wave run in parallel on their own branches, except the pairs predicted to collide.[hierarchy] main_agent = trueruns Main Gideon as a process too. - Briefings from the graph. A worker's prompt carries its criteria, the current definitions of its declared symbols with callers and callees, what its prerequisites actually produced on the feature branch, and what its siblings are touching, fitted to a token budget;
cg work updatereports symbols merged upstream since the attempt began. - Blocking is opt-in. Gates listed in a role's
approvestopland,pr, a driftedmerge-up, or a land with uncovered criteria untilcg fleet approve. Without them nothing waits for a person.
A subtree is complete because it merged, not because a process exited. cg spec run --fleet is the same run in the foreground, --dry-run plans without claiming or creating anything, and without an enabled [hierarchy] the fleet is refused before anything is spawned; the single-level cg spec run is untouched.
The full flow, supervision, approvals, briefings, and a worked example are in docs/hierarchy.md; drift in docs/drift.md; the event log, cg serve, and steering in docs/events.md.
A fleet works on many branches in many worktrees of one repository, and Codify indexes all of them into one .codegraph/. A linked worktree resolves to the shared project through git's common directory, so cg init there joins rather than demanding a second database:
$ cd .codegraph/worktrees/wave-fleet-1 && cg init
joined /path/proj as worktree /path/proj/.codegraph/worktrees/wave-fleet-1 on branch wave/fleet/1
$ cg branches
* main 203 files bde2b2a9 /path/proj 9s ago
feature/fleet 0 files 227a1f98 …/worktrees/feature-fleet 2m ago base main
wave/fleet/1 13 files b5d899dd …/worktrees/wave-fleet-1 1m ago base feature/fleetFile rows are scoped by branch (UNIQUE(branch_id, path)), so a sync on one branch never adds or removes another's rows; freshness and the index gate are per branch too, so two worktrees walk in parallel without either being told the graph is already fresh. Schema v16 carries the branch registry plus the branch, worktree, and parent on every attempt. Branch identity is read from git's own files — no process is spawned to learn a branch name, which matters when a fleet opens the graph thousands of times.
Queries answer for the branch you are on. --branch <name> asks another and --all-branches asks them all — on search, symbol, context, survey, impact, recall, anchors, check and guard — and a hit is labelled @branch only when more than one branch is in scope, so single-branch output never changes. Ten worktrees are mostly the same bytes, so the indexer keys parsed content by hash and reuses it across branches (40 reused in cg sync, "reused" in its JSON): a fresh worktree costs a walk and a copy, not a full parse.
Memories carry the branch they were made on, because a decision taken on a wave branch is not yet a decision of the project; cg fleet merge-up promotes them to the base along with the code, dropping the ones the base already holds. cg brief opens with the branch, its worktree, and the other branches with work in the graph. And cg watch --fleet follows every registered worktree from one process, picking up a worktree cg fleet begin created seconds after it appears. Details and the migration rules are in docs/branches.md.
Some questions are not deterministic — is this failure flaky or real, is this memory a reusable skill or noise, which of these findings matters most. cg jev asks TypeSafe's System One model (typesafe/jev-1.13, over OpenRouter) for a typed answer: a noul (probability true), a choice of up to 255 labelled options, or an ordinal score. It never generates text.
| Command | Description |
|---|---|
cg jev doctor [--probe] |
Key, curl, endpoint, model, and log health; --probe sends one tiny decision |
cg jev ask [<request.json>|-] |
Ask directly: --state S, --noul N I, --choice N I --option K=D …, --score N I --level L …, or a complete request body |
cg jev log [-n N] |
The last N calls from .codegraph/jev.log, with request id, model, tokens, cost, and latency |
Three commands ask a question of their own and print the answer beside their verdict:
| When | Jev adds |
|---|---|
cg spec done and verify_cmd fails |
jev triage: test_failure (confidence 0.82) → fix_test (confidence 0.82) — a category and a next action from the tail of the output, also appended to the outcome memory |
cg guard |
[jev 3.80 Blocking] per finding, and the findings ordered most severe first — scored in one call whatever the count |
cg fleet pr |
Jev readiness: 0.95 (high) — 2/3 tasks qualified, 4 commits in the pull request body |
All three share one gate, so the failure mode is the same everywhere: jev: OPENROUTER_API_KEY is not set — failure triage skipped, once, on stderr, with the command's own verdict and exit code untouched. cg guard --strict still fails on exactly what it failed on before; a red verify_cmd is still red.
Apart from the opt-in changelog highlights and cg recap, this is the one remote call Codify makes, and the principle says so: the graph, memory and workflow stay local, Jev is mandatory for the features built on it — cg memory classify with no OPENROUTER_API_KEY is a clear error, never a quiet fallback — and never authoritative: it narrows, ranks, and flags, while verify_cmd and the graph checks decide. The key never reaches a command line (curl is driven through a private 0600 config file), 429 and 529 back off and retry, and every call is logged. See docs/jev.md.
Codify works without any configuration file. If you want to change where it
keeps its files, stop it from syncing the graph on its own, or set how many
parse workers an index pass uses, add an optional
codify.kvx at the repository root. cg config init writes one with every
default spelled out and a comment on each key:
[sync]
auto = true # implicit syncs: hooks, read-command freshness, MCP/LSP/serve/watch
[paths]
spec = "spec" # workflow.kvx, feature specs, rendered mirrors
context = ".codify" # agent-context.md, recap.md
skills = ".agents/skills" # generated SKILL.md files
codemap = "CODEMAP.md" # written by cg codemap
[index]
# workers = auto # parse workers per index pass: 1-64; auto sizes from the machine
Setting paths.spec = "planning/specs" moves the whole spec workflow there.
The engine, orchestrator, fleet, docs, drift, recap, MCP resources and root
discovery all follow it, and the rendered CLAUDE.md and AGENTS.md name the new
directory.
Setting auto = false stops every implicit sync: the post-edit hook, the
freshness check on read commands, the pre-commit index, the post-commit hook,
and the MCP, LSP, watch and VS Code refreshes. cg sync, cg index and
cg init still run, and cg brief says that the graph may be stale.
[index] workers = N (1 to 64) sets the parse workers for cg init, cg index
and cg sync — more than the machine profile picks on a big machine, fewer on
a shared one. --workers N on the command line beats CG_INDEX_WORKERS,
which beats the file; cg info prints the count and where it came from
(docs/config.md).
A path that is absolute, that escapes the repository with .., or that is
empty is rejected with a message naming the key and the file, and the default
is used instead. Unknown keys never stop a command: cg config check lists
them, and cg check reports them as a warning. See
docs/config.md.
Everything above serves an agent that already exists. Codify can also be the thing that starts them: from the terminal with cg spec run, or from VS Code one task at a time.
cg spec run turns a parallel spec into running agent sessions. It loops the same primitives a human fleet would use — claim-next picks a conflict-free task, resume --prompt writes its briefing — then forks one driver process per slot, feeding each the prompt on stdin and capturing its output to a per-task log:
$ cg spec run -n 3
[run] task 2.1 → codex (agent run-1, log .codegraph/agents/myfeature-2.1.log)
[run] task 2.3 → codex (agent run-2, log .codegraph/agents/myfeature-2.3.log)
[run] task 2.1 exit 0 → status done
...
[run] frontier empty — 0 failure(s) this run
Configuration lives in spec/workflow.kvx, next to everything else the workflow owns:
[agents]
driver = "codex" # codex | claude | custom
max = 3 # default slot count (-n overrides)
ttl = 3600 # lease TTL, seconds
codex_args = "" # extra arguments for the codex driver
claude_args = "" # extra arguments for the claude driver
cmd = "" # custom driver: a shell template with
# ${PROMPT_FILE} ${TASK} ${ROOT} ${AGENT}The safety posture is deliberate. The codex driver runs codex exec --sandbox workspace-write --skip-git-repo-check -C <root>, so children get Codex CLI's workspace-write sandbox — they can edit the project and nothing else. The claude driver runs claude -p --permission-mode acceptEdits: headless, edits auto-accepted, everything riskier still gated by Claude Code's own permission system. The custom driver hands your own command template to /bin/sh -c with the prompt file, task id, root, and agent name substituted — which is also how the orchestrator tests itself without either CLI installed.
Completion is judged by the spec, not the process: after a child exits, the task's status is re-read. done or implemented means success (the lease was already auto-released); anything else releases the lease so the task returns to the pool, and records an outcome memory — agent exited rc=N without completing — so the failure is knowledge, not just a log line. The run stops when the frontier is empty, or when failures exceed --max-fail (default 2). Ctrl-C terminates the children, releases their leases, and exits 130; nothing stays claimed by a dead run. --dry-run prints the full plan — waves, tasks, the exact argv per task — and claims nothing.
Orchestration requires a .codegraph/ index (leases live there) and cg spec mode parallel (or prod); anything else is refused with a one-line hint.
The Codify sidebar carries a persistent Agent chat view: the extension is an Agent Client Protocol client that spawns Claude Code or Codex through its ACP adapter (claude-code-acp / codex-acp, or any ACP agent via codify.acp.customCommand) lazily on your first message — no task required, a driver picker in the header, New Chat to reset. The view renders streamed replies and thinking, tool calls as live status cards with diffs, the agent's plan, and permission requests as inline buttons. Every session gets Codify's MCP server injected automatically (cg mcp), so the agent has the graph, spec, and memory tools with zero per-repo config; the agent's ACP file reads and writes are served by the extension and confined to the workspace.
The chat itself is a real chat: markdown replies with headings, tables and fenced code blocks that copy in a click, every file path the agent mentions clickable to that line, collapsed thinking, an autosizing composer with history, and a context bar carrying the feature, the attached task's live status, and the driver. Typing / opens a palette of Codify's own verbs — /brief, /next, /context, /search, /impact, /why, /changes, /review, /tests, /check, /status, /remember, /handoff, /task, /open — each of which runs the real cg command in the workspace, shows the output as a card, and hands it to the agent as context. Asking "what breaks if I change this?" and running cg impact are the same gesture. Beyond those, every tool cg serves is a slash command — the palette is generated from the tool list, with argument hints and typing taken from each tool's schema (/get_context auth flow, /spec_claim id=2.1 ttl=20) — so the chat reaches everything Codify can do. And it reaches the fleet: /fleet shows the tree and runs, /attach <agent> follows a fleet agent's live transcript so that what you type next steers it, /steer <agent> <message> sends one message, and approval requests and escalations arrive as cards with Approve and Reject. Transcripts render windowed, so ten thousand entries stay responsive.
Start Agent Session on Task drives the same chat with the board discipline: it claims the task, seeds the opening prompt from cg resume --task <id> --prompt, and runs the session in the view (or, when the view is busy, in an editor panel beside it — concurrent task sessions each get their own). Set codify.agent.interface: "terminal" for the classic seeded terminal instead — that path is also offered automatically when the adapter fails to start.
Run Agent Headless on Task runs the claim + prompt as a background VS Code task (codex exec --sandbox workspace-write or claude -p) and reports the exit as a notification. Run Wave with Agents launches cg spec run -n <codify.agent.parallelism> in a terminal. Hand Off Task and Resume Task in Agent Session wrap cg handoff and cg resume, and Stop Agent Session ends a session — a panel or terminal closing with the task unfinished offers a handoff and releases the claim. Every view updates from events pushed over one cg serve connection, so there is no polling; against an older cg the board falls back to polling every 10 seconds while sessions run. Each task is decorated with the agent holding its lease. See editors/vscode/README.md.
cg lsp is a Language Server over the same graph, so every editor gets Codify — not just the one with an extension. It needs no compiler, no toolchain, and no project configuration: everything is answered from .codegraph/.
| Capability | What you get |
|---|---|
| Go to definition, find references | Straight from the symbols and refs tables |
| Hover | What the symbol is, how many references it has, and the decisions recorded about it |
| Workspace and document symbols | Instant trigram-backed lookup across the project |
| Code lens | Reference count and test-reference count above every function — coverage gaps become visible while reading |
| Diagnostics | kvx parse errors, and files edited outside the in-progress task's declared touches |
That last row is the point: scope drift shows up as a squiggle at the moment of the edit, for the human and the agent alike, without either of them running a command.
Point any LSP client at cg lsp over stdio. For Neovim:
vim.lsp.start({ name = "codify", cmd = { "cg", "lsp" },
root_dir = vim.fs.root(0, { ".codegraph", ".git" }) })editors/vscode/ ships the Codify extension — the whole workflow in the editor:
- Code navigation from the graph. It runs
cg lspand speaks LSP to it, so go-to-definition, find-references, workspace symbols, and code lens work in every indexed language with no compiler and no configuration. Hover shows what a symbol is, how many references it has, and the decisions recorded about it. - Scope drift as a squiggle. Editing a file outside the in-progress task's declared
touchesraises a warning in the Problems panel, at the moment of the edit. Advisory, never an error. - A live task tree. Tasks grouped feature → section → wave, the way the spec is written and the way a fleet divides work. Each row carries what decides whether you can pick it up: status, the agent holding the lease and its role, the branch, and the requires that are not met yet. Filter by status, wave or owner and search by id, title, section, owner, symbol or path; the active filter shows in the view title and survives a reload. A detail panel per task shows acceptance criteria, do-steps, declared symbols with their resolved location, touched paths, the verify command, tagged commits and the memories written under it. Start, complete, claim, release, run the verify command in the task's own worktree, open its branch, and copy a resume prompt.
- Agent sessions from the task board. Start a Claude Code or Codex session on a task — in the ACP agent panel by default, or a terminal, or headless — hand off, resume, run a whole wave, and stop sessions, with live lease decorations on the board. See Driving agents.
- A memory browser. Full text across project memory with filters for type, Jev class, task, branch and date, a detail pane whose symbols and files link back into the code, and supersede, forget, classify-with-Jev and promote-to-skill as actions.
- Start the fleet, and watch it live. Start fleet — a CodeLens on a feature's
spec.kvx, the fleet view's title, or the command palette — previews the plan (roles and budgets, the agents and branches a dry run would spawn, predicted collisions, what is already running) and, on confirm, runscg fleet up; Stop (now or drain), Pause and Resume sit on the view. The fleet tree shows Main Gideon, managers and workers with each branch, worktree, attempt, heartbeat and merge state, plus what each agent last said or did, its tokens and cost, and stall, retry and escalation badges; pending approvals are rows you decide from, and clicking an agent opens its transcript. Begin, merge-up, land, PR, checkpoint and open-worktree remain as actions. A refresh never opens a pull request, and merge state is only what the branch registry can prove. - Live over one connection. With a
cgthat hasserve, the extension keeps onecg servechild for every call and every event, patches the views from pushed events, and turns its polls off.codify.serve: false, or an older binary, falls back to shelling out and polling. - One Actions menu covering brief, review, test-impact, why, check, guard, snapshot, spec authoring, and hook installation. Reports open as rendered markdown.
- kvx editing. Go-to-definition on a
requiresentry jumps to that task; completion offers the keys a task actually understands; the outline lists every requirement and task. - One refresh scheduler. Every trigger — a pushed event, a spec file changing, a turn ending, a command finishing, a poll — funnels through a single chain of
cgcalls: bursts debounce into one run, a two-second floor sits between runs, and at most one trailing run queues behind a chain in flight. Each chain syncs once inside a three-second freshness window at background priority and never waits on another process's pass. The extension does not watchgraph.db— its own sync writes it, and that watcher used to turn every refresh into the next.
The extension has no dependencies and no build step — including its Language Server client, which is written by hand for exactly that reason.
cd editors/vscode
npx @vscode/vsce package # produces codify-workflow-1.4.0.vsix
code --install-extension codify-workflow-1.4.0.vsix --forceThe Marketplace identity is SidioraLabs.codify-workflow. See editors/vscode/README.md. Any other editor gets the same navigation by pointing its LSP client at cg lsp.
make # build ./cg (deps: C compiler, libsqlite3-dev)
make unit # C unit tests (tests/unit/*.c against build/libcg.a)
make integration # end-to-end CLI tests (tests/integration/*.sh in sandboxes)
make test # both
make release # static release binary -> tested -> published to the web rootRepository layout:
src/ one .c file per module; src/cg.h is the only header
src/govern.c brief, review, guard, check, handoff, resume — the governance layer
src/orchestrate.c cg spec run and the fleet supervisor (cg fleet up): durable
runs, three levels at once, stalls, budgets, retries, escalation
src/syncgate.c the single-writer index gate and machine-wide parse slots
src/fleet.c roles and capabilities, the branch lifecycle, merge lock,
approval gates, fleet tree
src/events.c the append-only event log and cg events
src/serve.c cg serve — one JSON-RPC connection with pushed events
src/drivers.c agent launch argv, structured output read into events, steering
src/drift.c spec drift, collision prediction, interface drift, coverage
src/changelog.c release notes from git history, optional model highlights
src/recap.c resume brief from agent transcripts: System One picks, chat model writes
src/jev.c typed decisions over curl
src/skills.c memories classed as skills, rendered as .agents/skills
src/lsp.c language server over the graph
src/gitint.c git history ingestion, churn, branch identity, commit mirroring
tests/unit/ kvx grammar, SHA-256 vectors, JSON scanner, StrBuf/IO
tests/integration/ graph, vcs, agentic, MCP protocol, spec engine, watcher,
sync gate, fleet, branches, jev, changelog, events, serve,
supervisor, drift, briefings, fleet end to end, recap,
explore
tests/fixtures/ sample polyglot project, a spec repo with golden outputs,
an explore project with gold queries,
stand-ins for curl, gh, and an OpenAI endpoint, and a
scripted fleet driver
editors/vscode/ VS Code extension (plain JS): kvx language, task tree,
agent panel, memory browser, live fleet view, serve client
scripts/ install/uninstall scripts served at codify.centra.ag + release publisher
docs/ARCHITECTURE.md how the pieces fit together
docs/sync.md the sync gate, freshness, slots, incremental resolution
docs/config.md codify.kvx: relocated spec/context/skills paths, auto-sync
docs/journal.md the write journal: what is queued when the database is busy, replay, cg journal
docs/retrieval.md declarations, phrase ranking, path outlines, the budget
docs/hierarchy.md roles, branch flow, the supervisor, supervision, approvals,
briefings
docs/drift.md spec, collision, interface, and coverage drift
docs/events.md the event log, cg serve, drivers, steering
docs/branches.md the unified multi-branch graph and schema v16
docs/jev.md typed decisions: types, transport, configuration, limits
The spec-render goldens were generated by the original Go specgen, so rendering parity is locked in by make test. CI builds and runs the full suite on every push via .github/workflows/ci.yml.
The graph is one SQLite file in WAL mode and every cg process in a checkout writes to it — the editor's cg lsp, cg watch, cg mcp, and each agent's commands. Writers take the lock in short bursts (the indexer commits every few dozen files and parses outside the lock), and a CLI command waits up to CG_BUSY_TIMEOUT_MS (default 30000) for its turn, so cg spec start/done during an editor index simply waits a moment. If the lock never frees, lifecycle writes such as cg remember, cg handoff, the bookkeeping behind cg spec done, and hook events are queued in .codegraph/journal/. The command exits 0 and says the write is queued, and the next cg command that gets the lock applies it (cg journal lists the queue; docs/journal.md). Writes that must see the live state, such as cg spec claim and claim-next, still exit 75. Their message says nothing was applied, the same command is safe to retry, and why the write was not queued. cg lsp and cg watch never hold an agent up: they defer their own index while the database is busy and keep answering from the last completed one.
The database lock decides who writes. A separate gate decides who walks: .codegraph/index.lock is held by the one process running an index pass, and every other caller leaves its paths in .codegraph/index.dirty and returns immediately, coalesced. The holder drains that note before it releases, so an edit made during a pass is picked up by that pass rather than by a fourth process. Above the project, parse threads are rationed machine-wide through slot files under /tmp/codify-<uid> (CG_INDEX_SLOTS, CG_INDEX_WORKERS, CG_SLOT_DIR), so ten projects indexing at once do not each claim every core. A linked worktree gets its own lock and note beside the main tree's. docs/sync.md has the full contract.
- Ignore rules combine sensible defaults (VCS directories,
node_modules, build output, binaries) with a.cgignorefile using one glob per line. - Symbol extraction is heuristic. A comment-aware and string-aware pattern engine per language is tuned for recall on definitions and call sites. It is not a full type-checked resolver.
- Snapshots store every non-ignored file up to 32 MB, including binaries. The graph indexes text files up to 8 MB.
- A coalesced sync returns without a fresh graph: it queued its change for the process holding the gate and answers from the last completed index.
- Queries answer for the branch you are on.
--branch <name>asks another and--all-branchesasks them all; a hit is labelled@branchonly when more than one branch is in scope, so single-branch output is unchanged. cg fleetdrivesgitandghas subprocesses. Withoutgh,prandcheckpointprint the commands instead of running them, andcheckpointonly treatsfeature/*head branches as Codify's own.- Jev needs the network and
OPENROUTER_API_KEY. Nothing in the core loop depends on it, and no Jev answer changes an exit code. - The fleet supervisor is one per project and spawns agents through their CLIs; it does not authenticate them. A
spendbudget depends on the cost the driver reports, only Claude Code can be steered mid-turn, and theretryapproval gate is accepted but not yet enforced. The rest is in docs/hierarchy.md. - Drift detection is line- and graph-level: behaviour changes inside unchanged lines, calls through a third function, and references the indexer cannot see are missed (docs/drift.md).
- Changelog highlights need the network and a key; without one the notes are the plain derived record.
MIT © Sidiora Labs