What is wrong
A session's only ways to find code are search_text (a regex over files) and read_file/list_files. Every implementing session re-explores the repository from zero, which is where most of its turns go: #139 spent $21 and never got past exploring, and the idle ceiling (#337) exists because reading phases run long. trailhq/Graft (MIT, @nanonets/graft on npm, 6k stars) builds a per-symbol code graph with tree-sitter and answers ask, grep, callers, skeleton and map from it, ranked by coupling, with file:line answers. Its tier-1 build is deterministic, needs no model and no key, and covers 23 languages including Python, TypeScript, Go, Java and Rust. Its benchmark claims fewer tool calls and tokens at equal correctness; the pull variant (tools, nothing injected) is the shape that fits here.
None of this may ask the adopter to install or configure anything. Graft's own onboarding (npm install -g, graft init, hooks into .claude/, an AGENTS.md edit, a graft/ directory it gitignores for you) is an agent-integration story. The framework's story is that the tool is there when a verb runs, or refuses by name.
Shape
One builtin, several verbs. search_code on BUILTIN_SERVER, in read_only and everything built on it, with a mode parameter: ask (a task in words, ranked symbols with locations), grep (a regex, grouped by enclosing symbol), callers (a symbol, in or out, with depth), skeleton (a file's signatures), map (an orientation). Capability READS_REPO only. Results are tool results and take the same redaction, truncation and injection scan every tool result does. They describe the tree as this session sees it: HEAD plus its own staged writes, the same thing read_file returns (#337) and run_tests runs against. Not the run_script caveat: a command has to run against real files, a search has no such excuse, and the moment a model most wants callers is right after it renamed something -- when an index over the disk would say the symbol does not exist while read_file shows it a turn later.
The framework provisions Graft, deterministically, into its own cache. Not into the repository and not globally: npm install --prefix <cache>/graft/<pinned version> @nanonets/graft@<pinned version>, under $XDG_CACHE_HOME/in-lockstep (or RUNNER_TEMP in CI), pinned by exact version in the framework and by integrity from the lockfile npm writes. Done by provision when a verb that hands out the tool is bound, so the network is reached where Provision already reaches it and never beside the model. Node ≥ 20 is the one host requirement; without it the tool is declared and refuses graft.no_node naming the version, the way run_script refuses with no runner. An install that fails leaves the tool refusing graft.unavailable with npm's tail; the run goes on without it.
Staged writes are in the index. Before a query, if the session has staged anything, the staged set is materialised into the same throwaway worktree run_tests uses (adapters/worktree.materialize) and the index is built over that tree; with nothing staged, over the repository. Graft rebuilds by content hash -- its own numbers are 0.18s after one edited file -- so this is one worktree per query with staged changes, which run_tests already pays, not a full re-index. The fingerprint below covers the staged set as well, so a query after a new write rebuilds and a query after none does not. The same overlay is what search_text has been missing since #337 gave it to read_file; it takes it in the same change.
The index is built by the framework before the model starts, and is a cache the repository never sees. graft build --dir <cache>/graft/index/<repo fingerprint> with DO_NOT_TRACK=1 (Graft's telemetry is opt-out; here it is off by construction), GRAFT_NO_REFRESH=1 on every query, and never --deep (that is Graft's model layer, under Graft's own key, which this framework does not hold and would not route through its own recorder). --dir keeps Graft from writing graft/ or a .gitignore line into the working tree; ChangeGuard denies the cache path anyway. The fingerprint is HEAD plus a hash of the tracked files' bytes plus a hash of the staged set, so a rebuild happens when the tree or the session's writes changed and only then, and the run record carries it beside the tool's version, so two runs that searched different trees say so.
Sandbox. Graft runs as a subprocess with the ambient credentials dropped (Sandbox()), reading the repository and writing only its cache. It is not a model-chosen command: the model chooses the query, the framework chooses the argv, and the pattern is passed as one argv token (no shell). A query is bounded by max_tool_result_chars like any tool result.
Recorded and replayable. Tool IO is on the tape already, so a cassette replays a session's searches without Graft installed; --offline needs neither node nor the cache.
Monorepos. Graft discovers workspace scopes on its own (pnpm-workspace.yaml, per-package pyproject.toml); ask --in <scope> is exposed through a scope parameter, which is what a Test bound at a package directory (#372) would want to hand the model as a default.
Not built. graft init, the MCP server, hooks, AGENTS.md edits, --deep, --lsp, graft viz, and the GitHub App. graft blast --base is a natural review-lens context and is a separate issue once this lands.
Acceptance
- A test that
search_code is in the read-only set with READS_REPO and nothing else, and that a session whose set lacks it cannot reach it.
- A test that
provision installs the pinned version into the cache and nothing into the repository or the user's global prefix, and that a second provision with the same pin does nothing.
- A test that the index is built before the first model call, that its directory is outside the working tree, that
.gitignore is untouched, and that the fingerprint and the Graft version are on the run record.
- A test that two builds over the same tree bytes produce the same fingerprint and no rebuild, and that one changed file rebuilds.
- A test that every Graft subprocess carries
DO_NOT_TRACK=1 and GRAFT_NO_REFRESH=1, never a provider key, and never --deep.
- A test that a host without Node ≥ 20 yields a declared tool that refuses
graft.no_node, with zero installs attempted, and that a failed install refuses graft.unavailable.
- A test that a session which staged a file defining a new symbol gets a hit for it from
grep, callers and ask at the staged location, and that a staged deletion makes the symbol disappear; and that search_text sees the same staged set.
- A test that a query with nothing staged builds over the repository and materialises no worktree, and that two queries over the same staged set build once.
- A test that a
grep result carrying an instruction is fenced and reported, like any tool result.
- A recorded
implement run on this repository with the tool bound, and the eval corpus comparison against a run without it (turns, cost, tool calls) in evidence/, so the benefit is measured here rather than quoted from Graft's README.
docs/extending.md names the tool and says it searches what the session staged; README's Implement row mentions it.
- A gate row, held, cited by O2 (nothing to configure) and O6 (no key reaches Graft, no telemetry), named by the tests above.
make check green from a cold cache.
Sequencing
Independent. Related: #337 (the idle ceiling this should make fire less), #372 (package scopes), #332 (a child session with search_code alone is the cheap scout).
Objectives: O2, O6, O7 (the index is arithmetic, not a prompt), O13 (fewer turns under the same ceiling).
What is wrong
A session's only ways to find code are
search_text(a regex over files) andread_file/list_files. Every implementing session re-explores the repository from zero, which is where most of its turns go: #139 spent $21 and never got past exploring, and the idle ceiling (#337) exists because reading phases run long. trailhq/Graft (MIT,@nanonets/grafton npm, 6k stars) builds a per-symbol code graph with tree-sitter and answersask,grep,callers,skeletonandmapfrom it, ranked by coupling, withfile:lineanswers. Its tier-1 build is deterministic, needs no model and no key, and covers 23 languages including Python, TypeScript, Go, Java and Rust. Its benchmark claims fewer tool calls and tokens at equal correctness; the pull variant (tools, nothing injected) is the shape that fits here.None of this may ask the adopter to install or configure anything. Graft's own onboarding (
npm install -g,graft init, hooks into.claude/, anAGENTS.mdedit, agraft/directory it gitignores for you) is an agent-integration story. The framework's story is that the tool is there when a verb runs, or refuses by name.Shape
One builtin, several verbs.
search_codeonBUILTIN_SERVER, inread_onlyand everything built on it, with amodeparameter:ask(a task in words, ranked symbols with locations),grep(a regex, grouped by enclosing symbol),callers(a symbol, in or out, with depth),skeleton(a file's signatures),map(an orientation). CapabilityREADS_REPOonly. Results are tool results and take the same redaction, truncation and injection scan every tool result does. They describe the tree as this session sees it: HEAD plus its own staged writes, the same thingread_filereturns (#337) andrun_testsruns against. Not therun_scriptcaveat: a command has to run against real files, a search has no such excuse, and the moment a model most wantscallersis right after it renamed something -- when an index over the disk would say the symbol does not exist whileread_fileshows it a turn later.The framework provisions Graft, deterministically, into its own cache. Not into the repository and not globally:
npm install --prefix <cache>/graft/<pinned version> @nanonets/graft@<pinned version>, under$XDG_CACHE_HOME/in-lockstep(orRUNNER_TEMPin CI), pinned by exact version in the framework and by integrity from the lockfile npm writes. Done byprovisionwhen a verb that hands out the tool is bound, so the network is reached whereProvisionalready reaches it and never beside the model. Node ≥ 20 is the one host requirement; without it the tool is declared and refusesgraft.no_nodenaming the version, the wayrun_scriptrefuses with no runner. An install that fails leaves the tool refusinggraft.unavailablewith npm's tail; the run goes on without it.Staged writes are in the index. Before a query, if the session has staged anything, the staged set is materialised into the same throwaway worktree
run_testsuses (adapters/worktree.materialize) and the index is built over that tree; with nothing staged, over the repository. Graft rebuilds by content hash -- its own numbers are 0.18s after one edited file -- so this is one worktree per query with staged changes, whichrun_testsalready pays, not a full re-index. The fingerprint below covers the staged set as well, so a query after a new write rebuilds and a query after none does not. The same overlay is whatsearch_texthas been missing since #337 gave it toread_file; it takes it in the same change.The index is built by the framework before the model starts, and is a cache the repository never sees.
graft build --dir <cache>/graft/index/<repo fingerprint>withDO_NOT_TRACK=1(Graft's telemetry is opt-out; here it is off by construction),GRAFT_NO_REFRESH=1on every query, and never--deep(that is Graft's model layer, under Graft's own key, which this framework does not hold and would not route through its own recorder).--dirkeeps Graft from writinggraft/or a.gitignoreline into the working tree;ChangeGuarddenies the cache path anyway. The fingerprint is HEAD plus a hash of the tracked files' bytes plus a hash of the staged set, so a rebuild happens when the tree or the session's writes changed and only then, and the run record carries it beside the tool's version, so two runs that searched different trees say so.Sandbox. Graft runs as a subprocess with the ambient credentials dropped (
Sandbox()), reading the repository and writing only its cache. It is not a model-chosen command: the model chooses the query, the framework chooses the argv, and the pattern is passed as one argv token (no shell). A query is bounded bymax_tool_result_charslike any tool result.Recorded and replayable. Tool IO is on the tape already, so a cassette replays a session's searches without Graft installed;
--offlineneeds neither node nor the cache.Monorepos. Graft discovers workspace scopes on its own (
pnpm-workspace.yaml, per-packagepyproject.toml);ask --in <scope>is exposed through ascopeparameter, which is what a Test bound at a package directory (#372) would want to hand the model as a default.Not built.
graft init, the MCP server, hooks,AGENTS.mdedits,--deep,--lsp,graft viz, and the GitHub App.graft blast --baseis a natural review-lens context and is a separate issue once this lands.Acceptance
search_codeis in the read-only set withREADS_REPOand nothing else, and that a session whose set lacks it cannot reach it.provisioninstalls the pinned version into the cache and nothing into the repository or the user's global prefix, and that a second provision with the same pin does nothing..gitignoreis untouched, and that the fingerprint and the Graft version are on the run record.DO_NOT_TRACK=1andGRAFT_NO_REFRESH=1, never a provider key, and never--deep.graft.no_node, with zero installs attempted, and that a failed install refusesgraft.unavailable.grep,callersandaskat the staged location, and that a staged deletion makes the symbol disappear; and thatsearch_textsees the same staged set.grepresult carrying an instruction is fenced and reported, like any tool result.implementrun on this repository with the tool bound, and the eval corpus comparison against a run without it (turns, cost, tool calls) inevidence/, so the benefit is measured here rather than quoted from Graft's README.docs/extending.mdnames the tool and says it searches what the session staged; README's Implement row mentions it.make checkgreen from a cold cache.Sequencing
Independent. Related: #337 (the idle ceiling this should make fire less), #372 (package scopes), #332 (a child session with
search_codealone is the cheap scout).Objectives: O2, O6, O7 (the index is arithmetic, not a prompt), O13 (fewer turns under the same ceiling).