From 082620d1a2fbe9c4f109421c202ff54ad05e3a10 Mon Sep 17 00:00:00 2001 From: AnzoBenjamin Date: Sat, 5 Sep 2026 13:45:40 +0300 Subject: [PATCH 1/4] fix: resolving some review attestation issues. --- .agents/claude-code-cli.ts | 1 - .agents/codebuff-local-cli.ts | 1 - .agents/codex-cli.ts | 1 - .agents/gemini-cli.ts | 1 - .../EVENTS.jsonl | 79 --- .../LESSONS.md | 95 --- .../agent-restriction-audit-2026-07/PLAN.md | 93 --- .../agent-restriction-audit-2026-07/SPEC.md | 59 -- .../STATE.json | 17 - .../agent-restriction-audit-2026-07/STATUS.md | 225 ------- .../findings/edit-authorization.md | 78 --- .../findings/read-and-path.md | 120 ---- .../findings/spawn-permissions.md | 82 --- .../findings/tool-arg-allowlists.md | 94 --- .../findings/web-network.md | 90 --- .../sessions/audit-agent-body-2026-09/SPEC.md | 46 ++ .../AUDIT-REPORT.md | 169 ------ .../COVERAGE-MATRIX.md | 49 -- .../IMPLEMENTATION-REPORT.md | 103 ---- .../audit-agent-ecosystem-2026-07/MAP.md | 350 ----------- .../findings/cli-ux-picker.md | 69 --- .../findings/cli-ux-search.md | 72 --- .../findings/discovery-picker.md | 127 ---- .../findings/discovery-search.md | 97 --- .../findings/execution-picker.md | 76 --- .../findings/execution-search.md | 76 --- .../findings/orchestrator-picker.md | 65 -- .../findings/orchestrator-search.md | 72 --- .../findings/quality-picker.md | 92 --- .../findings/quality-search.md | 82 --- .../findings/runtime-picker.md | 128 ---- .../findings/runtime-search.md | 88 --- .agents/sessions/audit-cli-2026-07-10/MAP.md | 141 ----- .../AUDIT-REPORT.md | 222 ------- .../COVERAGE-MATRIX.md | 54 -- .../IMPLEMENTATION-REPORT.md | 48 -- .../audit-cli-next-level-2026-07/MAP.md | 341 ----------- .../findings/distribution-quality.md | 81 --- .../findings/interaction-commands.md | 83 --- .../findings/onboarding-config.md | 114 ---- .../findings/presentation-quality.md | 89 --- .../findings/runtime-state.md | 128 ---- .../findings/validation-summary.md | 22 - .../manifests/distribution-quality.md | 97 --- .../manifests/interaction-commands.md | 196 ------ .../manifests/onboarding-config.md | 261 -------- .../manifests/presentation-quality.md | 291 --------- .../manifests/runtime-state.md | 253 -------- .../AUDIT-REPORT.md | 173 ------ .../COVERAGE-MATRIX.md | 59 -- .../MAP.md | 350 ----------- .../EVENTS.jsonl | 3 - .../LESSONS.md | 120 ---- .../background-job-push-model-2026-08/PLAN.md | 133 ----- .../background-job-push-model-2026-08/SPEC.md | 119 ---- .../STATE.json | 10 - .../STATUS.md | 53 -- .../context-baseline-25k/EVENTS.jsonl | 11 - .../sessions/context-baseline-25k/LESSONS.md | 17 - .agents/sessions/context-baseline-25k/PLAN.md | 156 ----- .agents/sessions/context-baseline-25k/SPEC.md | 262 -------- .../sessions/context-baseline-25k/STATE.json | 10 - .../sessions/context-baseline-25k/STATUS.md | 334 ----------- .../EVENTS.jsonl | 12 - .../PLAN.md | 107 ---- .../SPEC.md | 120 ---- .../STATE.json | 10 - .../STATUS.md | 226 ------- .../baseline-output.txt | 33 -- .../dynamic-cross-session-memory/LESSONS.md | 11 - .../dynamic-cross-session-memory/PLAN.md | 26 - .../dynamic-cross-session-memory/SPEC.md | 39 -- .../dynamic-cross-session-memory/STATUS.md | 25 - .../external-read-roots-2026-08/SPEC.md | 151 ----- .../EVENTS.jsonl | 53 -- .../harness-cohesion-audit-2026-07/LESSONS.md | 34 -- .../harness-cohesion-audit-2026-07/PLAN.md | 129 ---- .../harness-cohesion-audit-2026-07/SPEC.md | 92 --- .../harness-cohesion-audit-2026-07/STATE.json | 17 - .../harness-cohesion-audit-2026-07/STATUS.md | 131 ---- .../harness-ui-overhaul-2026-08/PLAN.md | 82 --- .../harness-ui-overhaul-2026-08/SPEC.md | 72 --- .../harness-ui-overhaul-2026-08/STATUS.md | 54 -- .../EVENTS.jsonl | 42 -- .../LESSONS.md | 15 - .../PLAN.md | 82 --- .../SPEC.md | 44 -- .../STATE.json | 17 - .../STATUS.md | 81 --- .../LESSONS.md | 29 - .../PLAN.md | 55 -- .../SPEC.md | 128 ---- .../EVENTS.jsonl | 5 - .../read-tool-unification-2026-07/PLAN.md | 310 ---------- .../read-tool-unification-2026-07/SPEC.md | 97 --- .../read-tool-unification-2026-07/STATE.json | 10 - .../read-tool-unification-2026-07/STATUS.md | 90 --- .../AUDIT-REPORT.md | 67 --- .../COVERAGE-MATRIX.md | 28 - .../read-write-tooling-2026-07-10/MAP.md | 346 ----------- .../findings/runtime-edit.md | 17 - .../findings/sdk-contracts.md | 13 - .../findings/structured-results-plan.md | 557 ------------------ .../findings/ux-prompts.md | 13 - .../review-gate-correctness/LESSONS.md | 100 ---- .../sessions/review-gate-correctness/PLAN.md | 293 --------- .../review-gate-correctness/STATUS.md | 66 --- .../DESIGN-PROPOSAL.md | 294 --------- .../SPEC.md | 128 ---- .../shell-policy-audit-2026-07/LESSONS.md | 36 -- .../shell-policy-audit-2026-07/PLAN.md | 79 --- .../shell-policy-audit-2026-07/SPEC.md | 115 ---- .../shell-policy-audit-2026-07/STATUS.md | 36 -- .../EVENTS.jsonl | 4 - .../terminal-policy-repair-2026-08/LESSONS.md | 15 - .../terminal-policy-repair-2026-08/PLAN.md | 56 -- .../terminal-policy-repair-2026-08/SPEC.md | 45 -- .../terminal-policy-repair-2026-08/STATE.json | 10 - .../terminal-policy-repair-2026-08/STATUS.md | 39 -- .../unified-background-jobs/EVENTS.jsonl | 8 - .../unified-background-jobs/LESSONS.md | 22 - .../sessions/unified-background-jobs/PLAN.md | 184 ------ .../sessions/unified-background-jobs/SPEC.md | 59 -- .../unified-background-jobs/STATE.json | 24 - .../unified-background-jobs/STATUS.md | 63 -- .agents/types/agent-definition.ts | 10 - .agents/types/tools.ts | 6 +- agents/__tests__/base2.test.ts | 22 +- agents/__tests__/dependency-manager.test.ts | 2 +- agents/base2/base2.ts | 3 +- agents/basher.ts | 3 +- .../dependency-manager/dependency-manager.ts | 7 +- agents/guides/editor-writers-and-repair.md | 2 - agents/librarian/librarian.ts | 1 - agents/tmux-cli.ts | 1 - agents/types/agent-definition.ts | 10 - agents/types/tools.ts | 6 +- .../components/__tests__/sweep-boxes.test.tsx | 92 +++ .../components/renderers/compaction-box.tsx | 44 +- .../components/terminal-command-display.tsx | 8 +- .../initial-agent-type-sources.generated.ts | 4 +- cli/src/hooks/helpers/send-message.ts | 9 +- cli/src/types/chat.ts | 28 + .../__tests__/sdk-event-handlers.test.ts | 387 ++++++++++++ .../utils/__tests__/status-bar-chips.test.ts | 120 ++++ cli/src/utils/message-block-helpers.ts | 23 + cli/src/utils/sdk-event-handlers.ts | 102 +++- cli/src/utils/status-bar-chips.ts | 22 +- common/src/__tests__/images.test.ts | 114 ++++ common/src/actions.ts | 1 - common/src/constants/images.ts | 55 ++ .../types/agent-definition.ts | 10 - .../initial-agents-dir/types/tools.ts | 6 +- .../tools/params/tool/run-terminal-command.ts | 4 +- common/src/tools/params/tool/spawn-agents.ts | 8 +- common/src/types/agent-handoff.ts | 17 + common/src/types/agent-template.ts | 9 - common/src/types/dynamic-agent-template.ts | 4 - common/src/types/print-mode.ts | 65 +- docs/agents-and-tools.md | 5 +- .../__tests__/loop-agent-steps-abort.test.ts | 11 +- .../src/__tests__/loop-agent-steps.test.ts | 137 +++++ .../__tests__/prompts-schema-handling.test.ts | 8 +- .../spawn-agent-inline-nesting.test.ts | 123 +++- .../src/__tests__/subagent-timeout.test.ts | 118 +--- packages/agent-runtime/src/run-agent-step.ts | 48 ++ .../src/templates/__tests__/strings.test.ts | 222 +++++++ .../agent-runtime/src/templates/prompts.ts | 75 ++- .../agent-runtime/src/templates/strings.ts | 13 +- .../tools/handlers/tool/spawn-agent-inline.ts | 24 + .../tools/handlers/tool/spawn-agent-utils.ts | 148 ++--- .../src/tools/handlers/tool/spawn-agents.ts | 13 +- .../agent-runtime/src/tools/tool-executor.ts | 3 - .../src/util/runtime-semantic-compaction.ts | 26 +- sdk/src/__tests__/file-change-hooks.test.ts | 24 +- .../__tests__/tool-execution-deadline.test.ts | 65 -- .../direct-agent-tool-repair.test.ts | 2 - sdk/src/impl/direct-agent-tool-repair.ts | 5 +- sdk/src/provider-config.ts | 2 +- sdk/src/run.ts | 215 ++++--- sdk/src/tool-execution-deadline.ts | 48 -- sdk/src/tools/concurrency.ts | 76 +++ sdk/src/tools/file-change-hooks.ts | 8 +- 183 files changed, 2083 insertions(+), 12715 deletions(-) delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/EVENTS.jsonl delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/LESSONS.md delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/PLAN.md delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/SPEC.md delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/STATE.json delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/STATUS.md delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/findings/edit-authorization.md delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/findings/read-and-path.md delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/findings/spawn-permissions.md delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/findings/tool-arg-allowlists.md delete mode 100644 .agents/sessions/agent-restriction-audit-2026-07/findings/web-network.md create mode 100644 .agents/sessions/audit-agent-body-2026-09/SPEC.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/AUDIT-REPORT.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/COVERAGE-MATRIX.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/IMPLEMENTATION-REPORT.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/MAP.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/cli-ux-picker.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/cli-ux-search.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-picker.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-search.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/execution-picker.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/execution-search.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/orchestrator-picker.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/orchestrator-search.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/quality-picker.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/quality-search.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-picker.md delete mode 100644 .agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-search.md delete mode 100644 .agents/sessions/audit-cli-2026-07-10/MAP.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/AUDIT-REPORT.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/COVERAGE-MATRIX.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/IMPLEMENTATION-REPORT.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/MAP.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/findings/distribution-quality.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/findings/interaction-commands.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/findings/onboarding-config.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/findings/presentation-quality.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/findings/runtime-state.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/findings/validation-summary.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/manifests/distribution-quality.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/manifests/interaction-commands.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/manifests/onboarding-config.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/manifests/presentation-quality.md delete mode 100644 .agents/sessions/audit-cli-next-level-2026-07/manifests/runtime-state.md delete mode 100644 .agents/sessions/audit-read-write-flow-2026-07-11-independent/AUDIT-REPORT.md delete mode 100644 .agents/sessions/audit-read-write-flow-2026-07-11-independent/COVERAGE-MATRIX.md delete mode 100644 .agents/sessions/audit-read-write-flow-2026-07-11-independent/MAP.md delete mode 100644 .agents/sessions/background-job-push-model-2026-08/EVENTS.jsonl delete mode 100644 .agents/sessions/background-job-push-model-2026-08/LESSONS.md delete mode 100644 .agents/sessions/background-job-push-model-2026-08/PLAN.md delete mode 100644 .agents/sessions/background-job-push-model-2026-08/SPEC.md delete mode 100644 .agents/sessions/background-job-push-model-2026-08/STATE.json delete mode 100644 .agents/sessions/background-job-push-model-2026-08/STATUS.md delete mode 100644 .agents/sessions/context-baseline-25k/EVENTS.jsonl delete mode 100644 .agents/sessions/context-baseline-25k/LESSONS.md delete mode 100644 .agents/sessions/context-baseline-25k/PLAN.md delete mode 100644 .agents/sessions/context-baseline-25k/SPEC.md delete mode 100644 .agents/sessions/context-baseline-25k/STATE.json delete mode 100644 .agents/sessions/context-baseline-25k/STATUS.md delete mode 100644 .agents/sessions/context-budget-architecture-2026-08/EVENTS.jsonl delete mode 100644 .agents/sessions/context-budget-architecture-2026-08/PLAN.md delete mode 100644 .agents/sessions/context-budget-architecture-2026-08/SPEC.md delete mode 100644 .agents/sessions/context-budget-architecture-2026-08/STATE.json delete mode 100644 .agents/sessions/context-budget-architecture-2026-08/STATUS.md delete mode 100644 .agents/sessions/context-budget-architecture-2026-08/baseline-output.txt delete mode 100644 .agents/sessions/dynamic-cross-session-memory/LESSONS.md delete mode 100644 .agents/sessions/dynamic-cross-session-memory/PLAN.md delete mode 100644 .agents/sessions/dynamic-cross-session-memory/SPEC.md delete mode 100644 .agents/sessions/dynamic-cross-session-memory/STATUS.md delete mode 100644 .agents/sessions/external-read-roots-2026-08/SPEC.md delete mode 100644 .agents/sessions/harness-cohesion-audit-2026-07/EVENTS.jsonl delete mode 100644 .agents/sessions/harness-cohesion-audit-2026-07/LESSONS.md delete mode 100644 .agents/sessions/harness-cohesion-audit-2026-07/PLAN.md delete mode 100644 .agents/sessions/harness-cohesion-audit-2026-07/SPEC.md delete mode 100644 .agents/sessions/harness-cohesion-audit-2026-07/STATE.json delete mode 100644 .agents/sessions/harness-cohesion-audit-2026-07/STATUS.md delete mode 100644 .agents/sessions/harness-ui-overhaul-2026-08/PLAN.md delete mode 100644 .agents/sessions/harness-ui-overhaul-2026-08/SPEC.md delete mode 100644 .agents/sessions/harness-ui-overhaul-2026-08/STATUS.md delete mode 100644 .agents/sessions/read-edit-auth-unification-2026-07/EVENTS.jsonl delete mode 100644 .agents/sessions/read-edit-auth-unification-2026-07/LESSONS.md delete mode 100644 .agents/sessions/read-edit-auth-unification-2026-07/PLAN.md delete mode 100644 .agents/sessions/read-edit-auth-unification-2026-07/SPEC.md delete mode 100644 .agents/sessions/read-edit-auth-unification-2026-07/STATE.json delete mode 100644 .agents/sessions/read-edit-auth-unification-2026-07/STATUS.md delete mode 100644 .agents/sessions/read-edit-pipeline-create-capability/LESSONS.md delete mode 100644 .agents/sessions/read-edit-pipeline-create-capability/PLAN.md delete mode 100644 .agents/sessions/read-edit-pipeline-create-capability/SPEC.md delete mode 100644 .agents/sessions/read-tool-unification-2026-07/EVENTS.jsonl delete mode 100644 .agents/sessions/read-tool-unification-2026-07/PLAN.md delete mode 100644 .agents/sessions/read-tool-unification-2026-07/SPEC.md delete mode 100644 .agents/sessions/read-tool-unification-2026-07/STATE.json delete mode 100644 .agents/sessions/read-tool-unification-2026-07/STATUS.md delete mode 100644 .agents/sessions/read-write-tooling-2026-07-10/AUDIT-REPORT.md delete mode 100644 .agents/sessions/read-write-tooling-2026-07-10/COVERAGE-MATRIX.md delete mode 100644 .agents/sessions/read-write-tooling-2026-07-10/MAP.md delete mode 100644 .agents/sessions/read-write-tooling-2026-07-10/findings/runtime-edit.md delete mode 100644 .agents/sessions/read-write-tooling-2026-07-10/findings/sdk-contracts.md delete mode 100644 .agents/sessions/read-write-tooling-2026-07-10/findings/structured-results-plan.md delete mode 100644 .agents/sessions/read-write-tooling-2026-07-10/findings/ux-prompts.md delete mode 100644 .agents/sessions/review-gate-correctness/LESSONS.md delete mode 100644 .agents/sessions/review-gate-correctness/PLAN.md delete mode 100644 .agents/sessions/review-gate-correctness/STATUS.md delete mode 100644 .agents/sessions/reviewer-coupling-followups/DESIGN-PROPOSAL.md delete mode 100644 .agents/sessions/reviewer-gate-concurrency-fix3-2026-07/SPEC.md delete mode 100644 .agents/sessions/shell-policy-audit-2026-07/LESSONS.md delete mode 100644 .agents/sessions/shell-policy-audit-2026-07/PLAN.md delete mode 100644 .agents/sessions/shell-policy-audit-2026-07/SPEC.md delete mode 100644 .agents/sessions/shell-policy-audit-2026-07/STATUS.md delete mode 100644 .agents/sessions/terminal-policy-repair-2026-08/EVENTS.jsonl delete mode 100644 .agents/sessions/terminal-policy-repair-2026-08/LESSONS.md delete mode 100644 .agents/sessions/terminal-policy-repair-2026-08/PLAN.md delete mode 100644 .agents/sessions/terminal-policy-repair-2026-08/SPEC.md delete mode 100644 .agents/sessions/terminal-policy-repair-2026-08/STATE.json delete mode 100644 .agents/sessions/terminal-policy-repair-2026-08/STATUS.md delete mode 100644 .agents/sessions/unified-background-jobs/EVENTS.jsonl delete mode 100644 .agents/sessions/unified-background-jobs/LESSONS.md delete mode 100644 .agents/sessions/unified-background-jobs/PLAN.md delete mode 100644 .agents/sessions/unified-background-jobs/SPEC.md delete mode 100644 .agents/sessions/unified-background-jobs/STATE.json delete mode 100644 .agents/sessions/unified-background-jobs/STATUS.md create mode 100644 common/src/__tests__/images.test.ts delete mode 100644 sdk/src/__tests__/tool-execution-deadline.test.ts delete mode 100644 sdk/src/tool-execution-deadline.ts create mode 100644 sdk/src/tools/concurrency.ts diff --git a/.agents/claude-code-cli.ts b/.agents/claude-code-cli.ts index 05514266f0..a1dd8d39ed 100644 --- a/.agents/claude-code-cli.ts +++ b/.agents/claude-code-cli.ts @@ -40,7 +40,6 @@ const definition: AgentDefinition = { input: { command: './scripts/tmux/tmux-cli.sh start --command "' + START_COMMAND + '"', - timeout_seconds: 30, }, } diff --git a/.agents/codebuff-local-cli.ts b/.agents/codebuff-local-cli.ts index 0513496645..ad0c661f98 100644 --- a/.agents/codebuff-local-cli.ts +++ b/.agents/codebuff-local-cli.ts @@ -51,7 +51,6 @@ const definition: AgentDefinition = { input: { command: './scripts/tmux/tmux-cli.sh start --command "' + START_COMMAND + '"', - timeout_seconds: 30, }, } diff --git a/.agents/codex-cli.ts b/.agents/codex-cli.ts index 11a127a382..e8b5d2f87e 100644 --- a/.agents/codex-cli.ts +++ b/.agents/codex-cli.ts @@ -120,7 +120,6 @@ const definition: AgentDefinition = { input: { command: './scripts/tmux/tmux-cli.sh start --command "' + START_COMMAND + '"', - timeout_seconds: 30, }, } diff --git a/.agents/gemini-cli.ts b/.agents/gemini-cli.ts index 087aa9228c..05630a3f42 100644 --- a/.agents/gemini-cli.ts +++ b/.agents/gemini-cli.ts @@ -46,7 +46,6 @@ const definition: AgentDefinition = { input: { command: './scripts/tmux/tmux-cli.sh start --command "' + START_COMMAND + '"', - timeout_seconds: 30, }, } diff --git a/.agents/sessions/agent-restriction-audit-2026-07/EVENTS.jsonl b/.agents/sessions/agent-restriction-audit-2026-07/EVENTS.jsonl deleted file mode 100644 index 0c9c83ce38..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/EVENTS.jsonl +++ /dev/null @@ -1,79 +0,0 @@ -{"ts":"2026-07-24T07:28:35.109Z","kind":"append_lesson","summary":"Appended entry \"Tier 1 progress\" to STATUS.md","payload":{"heading":"Tier 1 progress","artifact":"STATUS.md"}} -{"ts":"2026-07-24T07:35:25.153Z","kind":"task_update","summary":"Updated 1 task line(s): T1.1 Fix empty-readablePaths","payload":{"matched":["T1.1 Fix empty-readablePaths"]}} -{"ts":"2026-07-24T07:35:25.153Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T07:35:25.153Z","kind":"current_task","summary":"Current task -> \"T1.1 Fix empty-readablePaths\"","payload":{"currentTask":"T1.1 Fix empty-readablePaths"}} -{"ts":"2026-07-24T07:35:28.291Z","kind":"append_lesson","summary":"Appended entry \"Tier 1 execution\" to LESSONS.md","payload":{"heading":"Tier 1 execution","artifact":"LESSONS.md"}} -{"ts":"2026-07-24T07:36:45.225Z","kind":"task_update","summary":"Updated 1 task line(s): T1.1 Fix empty-readablePaths","payload":{"matched":["T1.1 Fix empty-readablePaths"]}} -{"ts":"2026-07-24T07:36:45.225Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T07:36:45.225Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T07:40:53.818Z","kind":"task_update","summary":"Updated 1 task line(s): T1.3 Fix git-committer commit contradiction","payload":{"matched":["T1.3 Fix git-committer commit contradiction"]}} -{"ts":"2026-07-24T07:40:53.818Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T07:40:53.818Z","kind":"current_task","summary":"Current task -> \"T1.3 Fix git-committer commit contradiction\"","payload":{"currentTask":"T1.3 Fix git-committer commit contradiction"}} -{"ts":"2026-07-24T07:46:12.343Z","kind":"append_lesson","summary":"Appended entry \"T1.3 progress\" to STATUS.md","payload":{"heading":"T1.3 progress","artifact":"STATUS.md"}} -{"ts":"2026-07-24T07:47:22.413Z","kind":"append_lesson","summary":"Appended entry \"Tier 1 complete\" to STATUS.md","payload":{"heading":"Tier 1 complete","artifact":"STATUS.md"}} -{"ts":"2026-07-24T07:47:24.219Z","kind":"append_lesson","summary":"Appended entry \"T1.3 decision\" to LESSONS.md","payload":{"heading":"T1.3 decision","artifact":"LESSONS.md"}} -{"ts":"2026-07-24T07:51:08.088Z","kind":"task_update","summary":"Updated 1 task line(s): T1.3 Fix git-committer commit contradiction","payload":{"matched":["T1.3 Fix git-committer commit contradiction"]}} -{"ts":"2026-07-24T07:51:08.088Z","kind":"session_status","summary":"Session status -> active","payload":{"status":"active"}} -{"ts":"2026-07-24T07:51:08.088Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T07:52:22.898Z","kind":"task_update","summary":"Updated 1 task line(s): T1.2 Stop zeroing spawnableAgents","payload":{"matched":["T1.2 Stop zeroing spawnableAgents"]}} -{"ts":"2026-07-24T07:52:22.898Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T07:52:22.898Z","kind":"current_task","summary":"Current task -> \"T1.2 Stop zeroing spawnableAgents\"","payload":{"currentTask":"T1.2 Stop zeroing spawnableAgents"}} -{"ts":"2026-07-24T07:54:09.215Z","kind":"task_update","summary":"Updated 1 task line(s): T1.2 Stop zeroing spawnableAgents","payload":{"matched":["T1.2 Stop zeroing spawnableAgents"]}} -{"ts":"2026-07-24T07:54:09.215Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T07:54:09.215Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T07:55:36.446Z","kind":"task_update","summary":"Updated 1 task line(s): T1.4 Expand combined short ripgrep","payload":{"matched":["T1.4 Expand combined short ripgrep"]}} -{"ts":"2026-07-24T07:55:36.446Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T07:55:36.446Z","kind":"current_task","summary":"Current task -> \"T1.4 Expand combined short ripgrep\"","payload":{"currentTask":"T1.4 Expand combined short ripgrep"}} -{"ts":"2026-07-24T07:56:36.382Z","kind":"task_update","summary":"Updated 1 task line(s): T1.4 Expand combined short ripgrep","payload":{"matched":["T1.4 Expand combined short ripgrep"]}} -{"ts":"2026-07-24T07:56:36.382Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T07:56:36.382Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T07:58:04.183Z","kind":"task_update","summary":"Updated 1 task line(s): T1.5 Narrow FORBIDDEN_SHELL_CHARACTERS","payload":{"matched":["T1.5 Narrow FORBIDDEN_SHELL_CHARACTERS"]}} -{"ts":"2026-07-24T07:58:04.183Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T07:58:04.183Z","kind":"current_task","summary":"Current task -> \"T1.5 Narrow FORBIDDEN_SHELL_CHARACTERS\"","payload":{"currentTask":"T1.5 Narrow FORBIDDEN_SHELL_CHARACTERS"}} -{"ts":"2026-07-24T08:01:02.749Z","kind":"task_update","summary":"Updated 1 task line(s): T1.5 Narrow FORBIDDEN_SHELL_CHARACTERS","payload":{"matched":["T1.5 Narrow FORBIDDEN_SHELL_CHARACTERS"]}} -{"ts":"2026-07-24T08:01:02.750Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T08:01:02.750Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T08:02:45.257Z","kind":"task_update","summary":"Updated 1 task line(s): T2.1 Tighten substring sensitive-path matches","payload":{"matched":["T2.1 Tighten substring sensitive-path matches"]}} -{"ts":"2026-07-24T08:02:45.257Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T08:02:45.257Z","kind":"current_task","summary":"Current task -> \"T2.1 Tighten substring sensitive-path matches\"","payload":{"currentTask":"T2.1 Tighten substring sensitive-path matches"}} -{"ts":"2026-07-24T08:13:22.169Z","kind":"task_update","summary":"Updated 1 task line(s): T2.1 Tighten substring sensitive-path matches","payload":{"matched":["T2.1 Tighten substring sensitive-path matches"]}} -{"ts":"2026-07-24T08:13:22.169Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T08:13:22.169Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T08:13:25.528Z","kind":"append_lesson","summary":"Appended entry \"Tier 2 complete\" to STATUS.md","payload":{"heading":"Tier 2 complete","artifact":"STATUS.md"}} -{"ts":"2026-07-24T08:14:02.810Z","kind":"task_update","summary":"Updated 1 task line(s): T2.2 Allow in-project absolute read paths","payload":{"matched":["T2.2 Allow in-project absolute read paths"]}} -{"ts":"2026-07-24T08:14:02.810Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T08:14:02.810Z","kind":"current_task","summary":"Current task -> \"T2.2 Allow in-project absolute read paths\"","payload":{"currentTask":"T2.2 Allow in-project absolute read paths"}} -{"ts":"2026-07-24T08:15:22.419Z","kind":"task_update","summary":"Updated 1 task line(s): T2.2 Allow in-project absolute read paths","payload":{"matched":["T2.2 Allow in-project absolute read paths"]}} -{"ts":"2026-07-24T08:15:22.420Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T08:15:22.420Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T08:16:48.180Z","kind":"task_update","summary":"Updated 1 task line(s): T3.1 Raise read/scan throughput caps","payload":{"matched":["T3.1 Raise read/scan throughput caps"]}} -{"ts":"2026-07-24T08:16:48.180Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T08:16:48.180Z","kind":"current_task","summary":"Current task -> \"T3.1 Raise read/scan throughput caps\"","payload":{"currentTask":"T3.1 Raise read/scan throughput caps"}} -{"ts":"2026-07-24T08:17:22.111Z","kind":"task_update","summary":"Updated 1 task line(s): T3.1 Raise read/scan throughput caps","payload":{"matched":["T3.1 Raise read/scan throughput caps"]}} -{"ts":"2026-07-24T08:17:22.111Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-24T08:17:22.111Z","kind":"current_task","summary":"Current task -> \"T3.1 Raise read/scan throughput caps\"","payload":{"currentTask":"T3.1 Raise read/scan throughput caps"}} -{"ts":"2026-07-24T08:30:15.045Z","kind":"append_lesson","summary":"Appended entry \"Tier 3 complete\" to STATUS.md","payload":{"heading":"Tier 3 complete","artifact":"STATUS.md"}} -{"ts":"2026-07-24T08:50:49.801Z","kind":"append_lesson","summary":"Appended entry \"CB-DRAIN-OSCILLATION repair\" to STATUS.md","payload":{"heading":"CB-DRAIN-OSCILLATION repair","artifact":"STATUS.md"}} -{"ts":"2026-07-24T08:50:49.804Z","kind":"append_lesson","summary":"Appended entry \"Circuit-breaker drain gaming\" to LESSONS.md","payload":{"heading":"Circuit-breaker drain gaming","artifact":"LESSONS.md"}} -{"ts":"2026-07-24T08:52:23.629Z","kind":"append_lesson","summary":"Appended entry \"Tier 3 complete (post CB-DRAIN fix)\" to STATUS.md","payload":{"heading":"Tier 3 complete (post CB-DRAIN fix)","artifact":"STATUS.md"}} -{"ts":"2026-07-24T08:53:25.293Z","kind":"task_update","summary":"Updated 1 task line(s): T3.1 Raise read/scan throughput caps","payload":{"matched":["T3.1 Raise read/scan throughput caps"]}} -{"ts":"2026-07-24T08:53:25.293Z","kind":"session_status","summary":"Session status -> active","payload":{"status":"active"}} -{"ts":"2026-07-24T08:53:25.293Z","kind":"current_task","summary":"Current task -> \"T3.1 Raise read/scan throughput caps\"","payload":{"currentTask":"T3.1 Raise read/scan throughput caps"}} -{"ts":"2026-07-24T08:54:43.180Z","kind":"task_update","summary":"Updated 1 task line(s): T3.1 Raise read/scan throughput caps","payload":{"matched":["T3.1 Raise read/scan throughput caps"]}} -{"ts":"2026-07-24T08:54:43.180Z","kind":"session_status","summary":"Session status -> active","payload":{"status":"active"}} -{"ts":"2026-07-24T08:54:43.180Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T08:56:06.330Z","kind":"task_update","summary":"Updated 1 task line(s): T3.2 Soften str_replace edit-authorization friction","payload":{"matched":["T3.2 Soften str_replace edit-authorization friction"]}} -{"ts":"2026-07-24T08:56:06.330Z","kind":"session_status","summary":"Session status -> active","payload":{"status":"active"}} -{"ts":"2026-07-24T08:56:06.331Z","kind":"current_task","summary":"Current task -> \"T3.2 Soften str_replace edit-authorization friction\"","payload":{"currentTask":"T3.2 Soften str_replace edit-authorization friction"}} -{"ts":"2026-07-24T08:57:24.243Z","kind":"task_update","summary":"Updated 1 task line(s): T3.2 Soften str_replace edit-authorization friction","payload":{"matched":["T3.2 Soften str_replace edit-authorization friction"]}} -{"ts":"2026-07-24T08:57:24.243Z","kind":"session_status","summary":"Session status -> completed","payload":{"status":"completed"}} -{"ts":"2026-07-24T08:57:24.243Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-24T09:13:33.652Z","kind":"append_lesson","summary":"Appended entry \"Final specialist attestation\" to STATUS.md","payload":{"heading":"Final specialist attestation","artifact":"STATUS.md"}} -{"ts":"2026-07-24T09:13:35.903Z","kind":"append_lesson","summary":"Appended entry \"Specialist gate vs plan markdown\" to LESSONS.md","payload":{"heading":"Specialist gate vs plan markdown","artifact":"LESSONS.md"}} -{"ts":"2026-07-24T14:21:40.615Z","kind":"append_lesson","summary":"Appended entry \"Followups in progress\" to STATUS.md","payload":{"heading":"Followups in progress","artifact":"STATUS.md"}} -{"ts":"2026-07-24T14:45:57.587Z","kind":"append_lesson","summary":"Appended entry \"Followups: tests + docs\" to STATUS.md","payload":{"heading":"Followups: tests + docs","artifact":"STATUS.md"}} -{"ts":"2026-07-24T14:52:52.716Z","kind":"append_lesson","summary":"Appended entry \"Broader regression results\" to STATUS.md","payload":{"heading":"Broader regression results","artifact":"STATUS.md"}} -{"ts":"2026-07-24T14:59:12.800Z","kind":"append_lesson","summary":"Appended entry \"Followups status — commit blocked on gate\" to STATUS.md","payload":{"heading":"Followups status — commit blocked on gate","artifact":"STATUS.md"}} -{"ts":"2026-07-24T15:00:42.422Z","kind":"append_lesson","summary":"Appended entry \"Commit gate re-arm\" to LESSONS.md","payload":{"heading":"Commit gate re-arm","artifact":"LESSONS.md"}} -{"ts":"2026-07-24T15:28:15.819Z","kind":"append_lesson","summary":"Appended entry \"Reviewer double-spawn and terminal-stop (addresses finding RF-1-eef7064f)\" to LESSONS.md","payload":{"heading":"Reviewer double-spawn and terminal-stop (addresses finding RF-1-eef7064f)","artifact":"LESSONS.md"}} -{"ts":"2026-07-24T15:28:36.262Z","kind":"append_lesson","summary":"Appended entry \"RF-1 resolution recorded\" to STATUS.md","payload":{"heading":"RF-1 resolution recorded","artifact":"STATUS.md"}} diff --git a/.agents/sessions/agent-restriction-audit-2026-07/LESSONS.md b/.agents/sessions/agent-restriction-audit-2026-07/LESSONS.md deleted file mode 100644 index 5d7adfb2ba..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/LESSONS.md +++ /dev/null @@ -1,95 +0,0 @@ -# LESSONS — Over-strict Agent Guardrail Audit - -## Process lesson (why this session exists) - -The first pass treated the user's git-committer _example_ as the whole scope and -only audited the terminal-command policy. The request was a GENERAL sweep. Lesson: -when a user gives "X, for example", X is one instance — measure breadth and shard -across ALL sibling surfaces, don't tunnel on the example. - -## Cross-cutting pattern: substring/whole-token over-matching - -The same over-strictness recurs in unrelated subsystems: - -- sensitive-paths.ts: basename.includes('kubeconfig'/'.tfstate') blocks docs/scripts. -- ripgrep allowlist: whole-token match rejects combined `-ni`. -- /git args: raw-input char class blocks `()[]{}` even though args are single-quoted. - Fix shape is identical each time: match the real artifact/operator precisely; stop - conflating "mentions the dangerous thing" with "is the dangerous thing". - -## Guidance-vs-enforcement contradictions - -git-committer (F1): the prompt tells the model to do exactly what the policy blocks. -Whenever a guardrail and the guidance are authored separately, they drift into -conflict. Co-locate or cross-test them (there is already a security-glob-parity test -doing this for the reviewer globs — a good model to copy). - -## Authority-derivation traps (highest risk) - -spawn-agent-utils deriveSpawnTemplateCapabilities has three "narrow on handoff" -behaviors that over-fire: empty readablePaths -> zero reads (HIGH), spawnableAgents -forced [], programmaticToolNames intersected with model-visible allowedTools. The -static template list is already the ceiling (getMatchingSpawn), so several of these -narrowings are redundant AND breaking. Redundant-but-breaking is worse than -redundant-but-inert. - -## What NOT to touch - -SSRF (host/IP/redirect revalidation), .env/private-key/credential denials, -project-path containment, cap.v3 HMAC+scope, replace_range authority chain, -plan-only terminal attenuation, force/default-branch push gating. These are precise -and load-bearing; the audit deliberately classified them KEEP. - -## Tooling gotcha - -general-agent cannot call code_search directly (not granted). Discovery shards must -use read_files/read_subtree/query_index or spawn a code-searcher. Several parallel -general-agents also hit transient upstream/rate-limit errors — keep audit waves -small (<=4-5) and re-run only the missing shards by checking the findings/ dir. - - - -## Tier 1 execution — 2026-07-24T07:35:28.290Z - -- PLAN preflight rejects `Depends on: none` — omit dependency lines when there are none. -- Empty handoff readablePaths must preserve undefined read scope, never emit `read: []`. -- security-reviewer LOOKS_GOOD on spawn-agent-utils, find-files-matching-content, git-command-args (receipt kwvcyLW4E1g). -- Validation receipts: spawn-agents-permissions kwPDvM-wCIg (35 pass); find-files-matching-content kwPDwKZR-Js (26 pass); git-command-args kwPDxEHrdNM (8 pass). - - - -## T1.3 decision — 2026-07-24T07:47:24.219Z - -Did not carve HEREDOC into git-commit policy. Multi-line commits already work via multiple -m flags (tested). Fixed the guidance contradiction in git-discipline.ts only — safer and smaller than expanding shell syntax under git-commit. - - - -## Circuit-breaker drain gaming — 2026-07-24T08:50:49.803Z - -Drain-by-1 on clean str_replace success looks like recovery but allows fail↔success oscillation to stay forever below the limit. Prefer non-draining success (leave counter unchanged) with a modestly higher limit (5) for mid-refactor friction. Never full-reset on success within a turn. - - - -## Specialist gate vs plan markdown — 2026-07-24T09:13:35.902Z - -Session markdown under .agents/sessions/ is often non-reviewable source; specialist attestation should target the real source pending set. Stale pending paths (e.g. general-agent.ts with no diff) can linger from earlier turns—confirm with git status before chasing phantom blockers. - - - -## Commit gate re-arm — 2026-07-24T15:00:42.421Z - -After any source/test edit (including database.test isolation), git-committer is withheld until hooks+reviewer pass again. Do not tight-loop spawn; wait for gate-state passed, then commit source+docs only. - - - -## Reviewer double-spawn and terminal-stop (addresses finding RF-1-eef7064f) — 2026-07-24T15:28:15.818Z - -This entry is the durable resolution of open reviewer finding RF-1 ("Explain why reviewers spawned twice despite a clean first run, and why the run stopped after the second completion"). - -**Why a second review spawned after a clean first run.** The gate attests to a _(snapshot fingerprint, pending file set)_ pair, and it re-arms on every edit. The first review returned LOOKS_GOOD against an earlier snapshot whose pending set did NOT include sdk/src/**tests**/database.test.ts. After that review, the database.test.ts full-suite isolation fixes landed (namespace import of ../impl/database, clearUserInfoCacheForTests(), beforeEach mock.restore()). Each edit bumped the workspace revision and added database.test.ts to the pending set, so the earlier attestation no longer covered the current pair. A second review was therefore required by design; the gate correctly refuses to reuse a stale approval over files it never saw. - -**Why the run stopped after the second completion.** The second attempt routed to migration-reviewer, which returned specialist-terminal-failure. Two compounding causes: (1) wrong specialist for the diff — nothing in the change set touches schemas, backfills, migrations, or rollbacks, so migration-reviewer had no surface to bind an attestation to; (2) the gate allows one automatic snapshot refresh when the tree moves under a bound reviewer, and that single retry also failed to produce a matching attestation, so it aborted. The harness then fail-closed (a terminal reviewer failure is neither an approval nor a finding-with-a-fix) and did NOT spawn repair-editor, because there was no code defect to repair. - -**Resolution path.** A fresh matching code-reviewer was run against stable snapshot e0088f39... (workspace revision 542) covering all 27 pending files: verdict LOOKS_GOOD, coverage covered, zero blocking findings (only two non-blocking pre-existing SSRF notes). This LESSONS entry records the explanation durably so the runtime gate reviewer can attest the RF-1 requirement as satisfied. - -**Reusable takeaway.** Route reviewer specialists to the actual risk surface of the diff. A specialist with no matching surface (migration-reviewer on a non-migration diff) is prone to terminal failure and should not be selected; the general code-reviewer is the correct gate reviewer for guardrail/policy relaxations with no schema surface. diff --git a/.agents/sessions/agent-restriction-audit-2026-07/PLAN.md b/.agents/sessions/agent-restriction-audit-2026-07/PLAN.md deleted file mode 100644 index 2a49f36611..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/PLAN.md +++ /dev/null @@ -1,93 +0,0 @@ -# PLAN — Over-strict Agent Guardrail Remediation (codebase-wide) - - - -Tiered so you can approve only what you want. Every change touches a security -control, so each tier gets an advisory `security-reviewer` pass before edit plus -the full validation/reviewer gate after. Classifications: [B]=soften, [C]=keep. - -Ranked by (friction relieved / risk added). Highest-value, lowest-risk first. - -## Tier 1 — High-value, low-risk relaxations (recommended) - -- [x] T1.1 Fix empty-readablePaths read lockout in spawn handoffs (HIGH) (empty readablePaths preserves unrestricted scope) - - Surface: packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts - - Problem: a handoff with `readablePaths: []` rewrites a child from unrestricted - reads to `read: []`, hard-blocking ALL project reads. - - Fix: only narrow read scope when readablePaths is non-empty; otherwise preserve - the child's static (possibly undefined = unrestricted) scope. Same for write - when writablePaths is empty. Preserves child-cannot-exceed-parent. - - Acceptance: empty readablePaths leaves filesystemScope.read undefined/unrestricted when child had no static read scope; non-empty still narrows - - Validate: bun test packages/agent-runtime/src/**tests**/spawn-agents-permissions.test.ts - -- [x] T1.2 Stop zeroing spawnableAgents / stripping programmaticToolNames on handoff (MEDIUM) (preserves static spawnableAgents + programmaticToolNames) - - Surface: packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts - - Problem: any handoff-carrying child gets spawnableAgents forced to [] and hidden - programmaticToolNames intersected with model-authored allowedTools. - - Fix: preserve child static spawnableAgents; do not gate programmaticToolNames - on model-visible allowedTools. - - Acceptance: handoff child retains static spawnableAgents and programmaticToolNames - - Validate: bun test packages/agent-runtime/src/**tests**/spawn-agents-permissions.test.ts - -- [x] T1.3 Fix git-committer commit contradiction (folds in shell-policy-audit F1) (guide uses multiple -m; HEREDOC forbidden) - - Surface: sdk/src/tools/terminal-command-policy.ts; common/src/constants/git-discipline.ts - - Problem: git-commit profile rejects `$(`/heredoc, but gitCommitGuidePrompt instructs it. - - Fix: structured commit-message path OR minimal bounded-heredoc allowance; reconcile guide. - - Acceptance: multi-line commit possible under git-commit profile; guidance matches policy - - Validate: bun test sdk/src/**tests**/terminal-command-policy.test.ts - -- [x] T1.4 Expand combined short ripgrep flags + benign output flags (MEDIUM) (26/26 find-files-matching-content pass) - - Surface: sdk/src/tools/find-files-matching-content.ts - - Fix: expand /^-[a-zA-Z]{2,}$/ bundles before validation; add -v/--invert-match, - -c/--count/--count-matches (and -o for code_search via extra switches). Keep - dangerous-flag denials. - - Acceptance: -ni accepted; --exec still rejected - - Validate: bun test sdk/src/**tests**/find-files-matching-content.test.ts - -- [x] T1.5 Narrow FORBIDDEN_SHELL_CHARACTERS for /git slash commands (MEDIUM) (8/8 git-command-args pass) - - Surface: cli/src/commands/git-command-args.ts - - Fix: keep hard-blocking newlines and shell operators `$;` + backtick + `|&<>\\`; - allow `( ) [ ] { }` (args are single-quoted). - - Acceptance: `:(exclude)dist` and `{a,b}.ts` parse; injection cases still throw - - Validate: bun test cli/src/commands/**tests**/git-command-args.test.ts - -## Tier 2 — Sensitive-path precision fixes (read surface) - -- [x] T2.1 Tighten substring sensitive-path matches (MEDIUM) (precise kubeconfig/tfstate; drop .crt/.cer/.yarnrc) - - Surface: common/src/util/sensitive-paths.ts - - Fix: remove .crt/.cer; anchor kubeconfig and .tfstate; drop .yarnrc blanket ban. - - Acceptance: public certs and kubeconfig docs readable; real secrets still blocked - - Validate: bun test common/src/util/**tests**/sensitive-paths.test.ts - -- [x] T2.2 Allow in-project absolute read paths (MEDIUM) (POSIX absolute form allowed; containment still authority) - - Depends on: T2.1 - - Surface: sdk/src/tools/path-utils.ts - - Fix: let resolveProjectPath containment be authority for absolute in-project paths. - - Acceptance: absolute path inside project root reads; outside still denied - - Validate: bun test sdk/src/**tests**/path-utils.test.ts - -## Tier 3 — Throughput caps and edit-authorization friction - -- [x] T3.1 Raise read/scan throughput caps (LOW) (caps raised; truncation flags kept) - - Surface: read-files, read-subtree, web-search, code-search defaults - - Acceptance: higher defaults with truncation flags still present - - Validate: relevant per-package unit tests - -- [x] T3.2 Soften str_replace edit-authorization friction (MEDIUM) (non-draining success; limit 5; small-file stale strip) (limit 5 non-draining; small-file stale strip) - - Surface: edit-read-state.ts, process-str-replace.ts, str-replace.ts - - Acceptance: non-staleness failures do not force full re-read; circuit breaker drains on success - - Validate: bun test packages/agent-runtime/src/tools/handlers/tool/**tests**/str-replace-circuit-breaker.test.ts - -## Keep — real security value (do NOT weaken) - -SSRF host/IP/redirect revalidation; .env & private-key/credential/real-.tfstate -denials; project-path containment; cap.v3 HMAC signing + scope binding; -replace_range authority chain; plan-only terminal attenuation; force/delete/ -default-branch push gating; privilege-escalation/system-package/env-dump bans. - -## Validation gates - -- Per-task bun test on named suites. -- security-reviewer advisory before editing terminal-command-policy.ts, - sensitive-paths.ts, path-utils.ts, spawn-agent-utils.ts. -- Full runtime validation + reviewer gate before finalizing each tier. diff --git a/.agents/sessions/agent-restriction-audit-2026-07/SPEC.md b/.agents/sessions/agent-restriction-audit-2026-07/SPEC.md deleted file mode 100644 index 62b0310220..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/SPEC.md +++ /dev/null @@ -1,59 +0,0 @@ -# SPEC — Over-strict Agent Guardrail Audit (codebase-wide) - -## Overview - -The user asked for a GENERAL audit of security/guardrail controls that limit our -agents more than they protect. The git-committer (git-commit terminal profile) -is ONE example, not the scope. This audit covers every "agent restriction" -surface in the SDK/CLI/agents/runtime and classifies each control as: -(A) safe to relax/remove (friction > value), (B) consolidate/soften, or -(C) keep (real security value). - -Supersedes the narrower `shell-policy-audit-2026-07` session, which only covered -the terminal-command policy. That session's F1–F6 findings are folded in here as -the "terminal-policy" shard. - -## Snapshot - -- inspect_codebase_structure snapshotId: - 7ae7b7cee225b1cd331f12ae17993437bd4a8fd96509802ed572f4e0567cd33a - -## Goals - -- Enumerate ALL runtime-enforced agent restrictions, not just terminal/git. -- For each, cite the exact file:line enforcement point and give a concrete, - minimal relaxation candidate where the control is mostly friction. -- Keep genuine security controls (traversal, secrets, privilege escalation, - SSRF, force-push) intact. - -## Non-Goals - -- Removing traversal/containment, secret-scanning, privilege-escalation, SSRF, - or force/default-branch-push protections. -- Rewriting the harness approval architecture. -- Implementation (plan mode). - -## Restriction surfaces / shards - -1. terminal-policy — sdk/src/tools/terminal-command-policy.ts, - sdk/src/services/harness-enforcement.ts, sdk/src/tools/run-terminal-command.ts, - agents/git-committer, common/src/constants/git-discipline.ts. -2. read-and-path — sdk/src/tools/read-policy.ts, read-files.ts, path-utils.ts, - common/src/util/sensitive-paths.ts, project-path-containment.ts. -3. tool-arg-allowlists — ripgrep flag allowlist in sdk/src/tools/code-search.ts & - find-files-matching-content.ts; cli/src/commands/git-command-args.ts; - agents/security-reviewer glob parity; glob/list-directory limits. -4. web-network — packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts - (isBlockedWebAddress/assertSafePublicWebUrl), read-only network-mutation bans. -5. edit-authorization — packages/agent-runtime/src/util/read-authorization.ts, - tools/handlers/tool/edit-read-state.ts, str-replace circuit breaker, - cap.v3 read-before-edit gating. -6. spawn-permissions — packages/agent-runtime/src/tools/handlers/tool/ - spawn-agent-utils.ts permission/profile/tool derivation for child agents. - -## Acceptance criteria - -- Every shard writes findings with file:line evidence + A/B/C classification. -- evaluate_audit_coverage passes over the shard receipts + snapshot. -- Synthesized report ranks removal candidates by (friction relieved / risk added). -- git-committer fix appears as ONE item among several, not the whole report. diff --git a/.agents/sessions/agent-restriction-audit-2026-07/STATE.json b/.agents/sessions/agent-restriction-audit-2026-07/STATE.json deleted file mode 100644 index 59eabd2561..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/STATE.json +++ /dev/null @@ -1,17 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "agent-restriction-audit-2026-07", - "status": "completed", - "currentTask": null, - "revision": 20, - "checkpoint": { - "taskId": "T3.2", - "phase": "validation", - "passed": true, - "summary": "Circuit breaker non-draining success; 6 pass; security-reviewer LOOKS_GOOD after CB-DRAIN fix", - "receiptIds": ["k0yZKaIADXE", "k1H1MlKhSNo"], - "recordedAt": "2026-07-24T08:57:24.242Z" - }, - "createdAt": "2026-07-24T07:35:25.152Z", - "updatedAt": "2026-07-24T08:57:24.242Z" -} diff --git a/.agents/sessions/agent-restriction-audit-2026-07/STATUS.md b/.agents/sessions/agent-restriction-audit-2026-07/STATUS.md deleted file mode 100644 index 014f314609..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/STATUS.md +++ /dev/null @@ -1,225 +0,0 @@ -# STATUS — Over-strict Agent Guardrail Audit - -## Current state - -Codebase-wide audit complete (plan mode). No source changed. Six restriction -surfaces mapped across SDK/CLI/agents/runtime; findings written with file:line -evidence and A/B/C classification. - -## Completed - -- SPEC + PLAN written; supersedes narrower shell-policy-audit-2026-07 (folded in). -- Shard findings written to .agents/sessions/agent-restriction-audit-2026-07/findings/: - - read-and-path.md (sensitive-paths, read-files, path-utils, read-subtree, containment) - - web-network.md (web-search fetch caps + SSRF keep-list) - - tool-arg-allowlists.md (ripgrep flags, /git args, security-reviewer params, result caps) - - edit-authorization.md (cap.v3, str_replace circuit breaker, read revocation) - - spawn-permissions.md (handoff capability derivation — includes 1 HIGH read-lockout) -- terminal-policy surface covered by prior shell-policy-audit F1-F6 (git-committer fix). - -## Key results - -- 1 HIGH (spawn empty-readablePaths read lockout), several MEDIUM friction items. -- The genuine security controls (SSRF, secrets, containment, cap.v3, force-push) - are sound and classified KEEP — the report separates friction from value. - -## Pending (awaiting user decision) - -- Which tiers to execute (Tier 1 recommended). -- T1.3 structured-commit path vs minimal heredoc allowance. - -## Blocked - -- Implementation blocked in plan mode + pending user go-ahead (security-sensitive). - -## Next checkpoint - -User selects tiers. Then exit plan mode, run security-reviewer advisory on the -touched security files, implement per-task, run named suites, then full gate. - -## Resume instructions - -Read SPEC.md + all findings/\*.md + PLAN.md tiers. Tier 1 tasks are independent -(T1.1-T1.5) and can be done in any order; T2.2 depends on T2.1. - - - -## Tier 1 progress — 2026-07-24T07:28:35.108Z - -Implemented T1.1, T1.2, T1.4, T1.5. - -Validation: - -- bun test packages/agent-runtime/src/**tests**/spawn-agents-permissions.test.ts → 35 pass -- bun test sdk/src/**tests**/find-files-matching-content.test.ts → 26 pass -- bun test cli/src/commands/**tests**/git-command-args.test.ts → 8 pass - -Still pending: T1.3 git-committer, Tiers 2–3. - -Changed source: - -- packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts -- packages/agent-runtime/src/**tests**/spawn-agents-permissions.test.ts -- sdk/src/tools/find-files-matching-content.ts -- sdk/src/**tests**/find-files-matching-content.test.ts -- cli/src/commands/git-command-args.ts -- cli/src/commands/**tests**/git-command-args.test.ts - - - -## T1.3 progress — 2026-07-24T07:46:12.342Z - -Editor updated common/src/constants/git-discipline.ts to recommend multiple -m flags instead of blocked HEREDOC. Validation: bun test terminal-command-policy.test.ts pending. - - - -## Tier 1 complete — 2026-07-24T07:47:22.412Z - -All Tier 1 tasks done: - -- T1.1 empty readablePaths read lockout fixed -- T1.2 spawnableAgents + programmaticToolNames preserved on handoff -- T1.3 gitCommitGuidePrompt reconciled (multiple -m; no HEREDOC) -- T1.4 ripgrep -ni expansion + -v/-c flags -- T1.5 /git pathspec ()[]{} allowed - -Validation receipts: - -- spawn-agents-permissions: 35 pass (kwPDvM-wCIg) -- find-files-matching-content: 26 pass (kwPDwKZR-Js) -- git-command-args: 8 pass (kwPDxEHrdNM) -- terminal-command-policy: 30 pass (kxbIGIjqXpI) -- security-reviewer LOOKS_GOOD (kwvcyLW4E1g) - -Pending: Tier 2 (sensitive-paths, absolute reads), Tier 3 (throughput caps, str_replace friction). - - - -## Tier 2 complete — 2026-07-24T08:13:25.528Z - -T2.1 + T2.2 implemented and validated. - -- sensitive-paths: drop .crt/.cer/.yarnrc; kubeconfig exact match; .tfstate endsWith only -- path-utils: allow absolute POSIX form; drive/UNC/.. still rejected; containment remains authority - -Validation: - -- bun test common/src/util/**tests**/sensitive-paths.test.ts → 5 pass (ky5wT4qXgdY) -- bun test sdk/src/**tests**/path-utils.test.ts → 17 pass (ky5wU9eXFWY) -- security-reviewer LOOKS_GOOD (ky5wV3MoQeI) - -Next: Tier 3 optional (throughput caps + str_replace friction). - - - -## Tier 3 complete — 2026-07-24T08:30:15.044Z - -T3.1 + T3.2 implemented and validated. - -Throughput caps: - -- MAX_RANGE_READ_BYTES 1MB → 4MB -- LIVE_SUBTREE_MAX_NODES 1000 → 5000 -- MAX_WEB_FETCH_BYTES 512KB → 2MB (soft-truncate declared oversize) -- MAX_FETCH_LENGTH 50KB → 150KB -- code_search maxResults 15 → 30 -- find_files_matching_content maxFiles 100 → 250 -- MAX_SPAWN_BATCH_SIZE 8 → 12 - -str_replace friction: - -- small-file unique oldString auto-strips stale basedOnRead (large files still hard-fail) -- circuit breaker limit 3 → 5; clean success drains by 1 -- revoke fresh-read only on true stale/cross-run capability failures - -Validation: - -- process-str-replace.test.ts → 73 pass (kz2XXWstvs8) -- str-replace-circuit-breaker.test.ts → 6 pass (kz2XYov_dfQ) -- read-subtree.test.ts → 18 pass (prior) -- spawn-agents-permissions.test.ts → pass with batch size 12 (prior) - -All Tier 1–3 remediation complete. Keep-list security controls unchanged. - - - -## CB-DRAIN-OSCILLATION repair — 2026-07-24T08:50:49.801Z - -Security-reviewer BLOCKING CB-DRAIN-OSCILLATION on drain-by-1 success. - -Fix: clean exact-match success leaves consecutiveStrReplaceFailuresByPath unchanged (non-draining). Limit stays 5. Failures and auto-correct still increment. Structural recovery may still full-clear. - -Validation: str-replace-circuit-breaker.test.ts → 6 pass (k0yZKaIADXE). -Security re-review pending on snapshot b86f88b1c48509decca134ca3bde20dc9147a53422c42845a61f0cc5d59fef4b. - - - -## Tier 3 complete (post CB-DRAIN fix) — 2026-07-24T08:52:23.628Z - -T3.1 + T3.2 complete. - -Throughput (T3.1): - -- MAX_RANGE_READ_BYTES 4MB, LIVE_SUBTREE 5000, web fetch 2MB/150KB, code_search 30, maxFiles 250, spawn batch 12 - -str_replace (T3.2): - -- small-file unique stale basedOnRead auto-strip; large files hard-fail -- circuit breaker limit 5; clean success non-draining (CB-DRAIN-OSCILLATION fixed) - -Validation: - -- process-str-replace 73 pass (kz2XXWstvs8) -- circuit-breaker 6 pass (k0yZKaIADXE) -- security-reviewer LOOKS_GOOD (k1H1MlKhSNo) snapshot b86f88b1... - -All Tier 1–3 remediation complete. Keep-list security controls unchanged. - - - -## Final specialist attestation — 2026-07-24T09:13:33.652Z - -reliability-reviewer LOOKS_GOOD (k2VmxFlExQk) on snapshot b52ba1d591816414e5f48ba1fbf4f784d862ca5f876ed99259ae325f700fb2bf. -security-reviewer LOOKS_GOOD (k1H1MlKhSNo) after CB-DRAIN non-draining fix. -Smoke: 114 pass / 0 fail across circuit-breaker + process-str-replace + spawn-agents-permissions (k2VmzIm7NOU). - -Plan session completed. All Tier 1–3 tasks done. Keep-list security controls unchanged. - - - -## Followups in progress — 2026-07-24T14:21:40.614Z - -User requested all three suggested followups: docs note, broader regression tests, commit via git-committer (source+docs; exclude plan session dirs). - - - -## Followups: tests + docs — 2026-07-24T14:45:57.586Z - -Docs: agents-and-tools.md Agent restriction policy section + stale cap numbers fixed. -SDK absolute-path tests updated for T2.2 (change-file, run-file-filter). -Database suite isolation: clearUserInfoCacheForTests() + beforeEach/afterEach clear. -Awaiting full sdk suite green then git-committer (source+docs; exclude plan sessions). - - - -## Broader regression results — 2026-07-24T14:52:52.715Z - -agent-runtime: 1172 pass / 0 fail -cli: 2450 pass / 0 fail (15 skip) -sdk: 1036 pass / 0 fail (1 skip) after absolute-path test updates + database cache isolation (clearUserInfoCacheForTests + mock.restore in beforeEach) - -Next: git-committer with source+docs owned_paths; exclude .agents/sessions/\* and packages/agent-runtime/docs/ - - - -## Followups status — commit blocked on gate — 2026-07-24T14:59:12.800Z - -1. Docs: done — docs/agents-and-tools.md Agent restriction policy (relaxed vs keep) + cap number fixes. -2. Broader tests: done — agent-runtime 1172 pass; cli 2450 pass; sdk 1036 pass / 0 fail after absolute-path test updates + database cache isolation. -3. Commit: blocked — git-committer withheld until validation/reviewer gate passes for pending files (incl. sdk/src/**tests**/database.test.ts). Will spawn git-committer once gate reports passed; owned_paths = source+docs only (exclude .agents/sessions/\* and packages/agent-runtime/docs/). - - - -## RF-1 resolution recorded — 2026-07-24T15:28:36.261Z - -Open reviewer finding RF-1-eef7064f (requirement uncertain: explain the reviewer double-spawn and terminal-stop) is addressed by a durable LESSONS.md entry rather than a source edit, because the requirement is an explanation, not a code behavior. Summary: (1) second review was required because the pending (snapshot, file-set) pair changed after the database.test.ts isolation fixes were added post-first-review; (2) the run stopped because migration-reviewer was the wrong specialist for a diff with no migration surface and hit specialist-terminal-failure after its single automatic snapshot refresh, so the gate fail-closed. A fresh code-reviewer against stable snapshot e0088f39 (rev 542) returned LOOKS_GOOD across all 27 pending files with zero blocking findings. Next: commit source+docs via git-committer once the gate attests the pending set. diff --git a/.agents/sessions/agent-restriction-audit-2026-07/findings/edit-authorization.md b/.agents/sessions/agent-restriction-audit-2026-07/findings/edit-authorization.md deleted file mode 100644 index 1fc33c069f..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/findings/edit-authorization.md +++ /dev/null @@ -1,78 +0,0 @@ -# Audit findings: edit-authorization - -- Subsystems: agent-runtime-edit-authorization -- Features: read-before-edit-gating, cap-v3-capability-binding, str-replace-circuit-breaker, replace-range-authority, compaction-authorization-revocation -- Files covered: 8 - -## [MEDIUM] error-handling — packages/agent-runtime/src/tools/handlers/tool/edit-read-state.ts:6 — [B] SOFTEN — markEditRequiresFreshRead revokes read authorization and forces a full re-read on ALL str_replace hard failures, including non-staleness ones - -- **Risk:** What it blocks: after any hard str_replace failure that is not a preflight-syntax error, str-replace.ts calls markEditRequiresFreshRead(reason:'preflight_failed') (str-replace.ts:~300-320), which sets failedEditRequiresReadByPath[path]=true AND (revokeReadAuthorization defaults true, edit-read-state.ts:17,25-29) deletes readAuthorizationsByPath/HashesByPath. The very next str_replace with no basedOnRead is then blocked by the recoveringFromFailedEdit gate (str-replace.ts:~123) even though the file never changed. Friction-vs-value: real staleness (reason stale_snapshot / stale_capability) genuinely warrants a re-read, but the common failures here are content-mismatch, not staleness — an ambiguous oldString (multiple occurrences), a tiny-anchor refusal, or a simple typo. The file on disk is unchanged, so a full re-read teaches the model nothing new; it could retry immediately with occurrenceIndex or a longer oldString drawn from content it already has. Forcing a re-read + authorization revocation here is pure friction that also nudges the agent toward the circuit breaker. -- **Fix:** Minimal relaxation: distinguish staleness failures from content-mismatch failures. Only revoke whole-file authorization + require a fresh read when staleness is actually observed (stale_snapshot/stale_capability, i.e. hadFreshWholeFileAuthorization was true then went stale, or a supplied capability failed its hash). For plain no-match/ambiguous/tiny-anchor failures, keep failedEditRequiresReadByPath set for the guidance but pass revokeReadAuthorization:false so a still-valid whole-file authorization survives and the agent can retry with occurrenceIndex/longer oldString without a redundant re-read. -- **Evidence:** edit-read-state.ts:17 revokeReadAuthorization=true default; :25-29 deletes all stored authorizations; str-replace.ts:~300-320 marks reason:'preflight_failed' on every non-syntax error; str-replace.ts:~123 recoveringFromFailedEdit && !hasAnyReadCapability -> hard block. - -## [MEDIUM] correctness — packages/agent-runtime/src/process-str-replace.ts:470 — [B] SOFTEN — a STALE supplied basedOnRead on a small file hard-fails instead of falling back to a unique-literal match - -- **Risk:** What it blocks: when a basedOnRead capability is supplied but its range hash no longer matches (hasStaleBasedOnRead) and requireFreshReadCapability is false, the code emits staleScopedFailure and refuses the edit rather than falling back to an unscoped unique-literal match ('the runtime did not fall back to an unscoped whole-file match'). Friction-vs-value: the stale-anchor guard is correct for large/ambiguous edits (never silently expand scope), but on a small file whose oldString is still unique in current content, the edit is unambiguous and safe. This is exactly the situation right after the agent's own earlier edit staled a pre-edit anchor — the model is blocked and told to re-read even though a naked unique-oldString edit would land correctly. Note the sibling bogus-anchor path already has a loop-breaker that auto-strips an invalid anchor when oldString is uniquely matchable (process-str-replace.ts:~205-225); the stale-anchor path lacks the same escape hatch. -- **Fix:** Minimal relaxation: mirror the bogus-anchor auto-strip. When !requireFreshReadCapability and normalizedOldStr occurs exactly once in current content, drop the stale anchor and apply as a naked unique-literal edit with a warning note (same wording as autoStrippedBogusAnchor), instead of recording a hard failure. Keep the hard-fail only when oldString is non-unique or requireFreshReadCapability is set. -- **Evidence:** process-str-replace.ts hasStaleBasedOnRead branch (staleScopedFailure, '...did not fall back to an unscoped whole-file match'); contrast auto-strip loop-breaker at ~205-225 (uniquelyMatchable && !requireFreshReadCapability -> basedOnRead=undefined). - -## [MEDIUM] error-handling — packages/agent-runtime/src/tools/handlers/tool/str-replace.ts:36 — [B] SOFTEN — str_replace circuit breaker (limit 3) counts successful auto-corrects and partial successes and never decrements on clean success - -- **Risk:** What it blocks: STR_REPLACE_MAX_CONSECUTIVE_FAILURES=3 hard-blocks all raw str_replace on a path for the rest of the turn. The budget is charged not only by hard failures but also by auto-corrected near-matches (a SUCCESS, str-replace.ts:~330-338) and by non-atomic partial successes, and a clean exact-match success deliberately does NOT decrement it (test 'does not erase prior failures after an exact-match success'). Friction-vs-value: the anti-loop intent is sound and genuinely prevents corruption spirals, and it still leaves rewrite_symbol/replace_range/write_file open. But on a large legitimate refactor with many small edits, a couple of gated-but-correct auto-corrects plus one partial success can reach 3 and lock out str_replace even while the agent is making real progress. Counting a fully-gated auto-correct (which already passed similarity + uniqueness + delimiter-balance gates) as equivalent to a failure is the most aggressive part. -- **Fix:** Minimal relaxation (keep the breaker): (a) let a clean exact-match success decrement the counter by 1 (floor 0) so steady progress drains the budget, instead of freezing it; and/or (b) exclude auto-corrects that passed all deterministic gates from the budget, counting only hard failures + suspect auto-corrects; and/or (c) raise the limit to ~5. Preserve the no-reset-on-success stance for the alternating failure/success case by decrementing (not resetting). -- **Evidence:** str-replace.ts:36 const STR_REPLACE_MAX_CONSECUTIVE_FAILURES=3; :~104-124 breaker returns before processing; :~330-338 hadAutoCorrect || failedReplacementCount>0 increments counter; test file lines confirm no decrement on success and partial-success charge. - -## [LOW] state-mutation — packages/agent-runtime/src/util/read-authorization.ts:4 — [C] KEEP — compaction revokes implicit whole-file edit authority - -- **Risk:** What it blocks: after context compaction removes the exact read bodies, revokeImplicitReadAuthorizationsAfterCompaction clears readAuthorizationsByPath/HashesByPath and sets editRereadRequirementsByPath[path]={reason:'context_compacted'}, forcing a fresh read before a whole-file-authorized edit. Friction-vs-value: strongly value. Once the exact bytes leave the model context, the agent no longer demonstrably observed current content, so a whole-file overwrite could be stale. Importantly this does NOT dead-end scoped edits: strictEditAuthorizationError allows a scoped cap.v3 capability (allowScopedCapability default true), so an agent still holding a valid basedOnRead token can proceed without a re-read. The friction is scoped to the exact case where authority genuinely evaporated. -- **Fix:** No change. Genuine stale-write protection; scoped cap.v3 path already avoids redundant re-reads. -- **Evidence:** read-authorization.ts:14-20 sets reason 'context_compacted' and empties authorization maps; edit-read-state.ts:79-80 allowScopedCapability && hasScopedCapability -> undefined (no block). - -## [LOW] security — common/src/util/content-hash.ts:66 — [C] KEEP — cap.v3 per-process HMAC signing key + project/path/run scope binding - -- **Risk:** What it blocks: READ_CAPABILITY_SIGNING_KEY=randomBytes(32) is generated per process, so all outstanding capabilities are invalidated on runtime restart; readCapabilityMatchesScope binds every token to projectId+path+runId, so cross-path/cross-run replay is rejected (validateReadCapabilityAuthority in process-str-replace.ts, and the replace_range scope check in process-edit-transaction.ts). Friction-vs-value: strongly value. This is the core anti-forgery/anti-replay guarantee that makes an echoed post-edit anchor trustworthy. The only friction — a mandatory re-read after a runtime restart or when no runtime scope exists — is inherent to an in-process capability and cannot be relaxed without weakening the authenticity guarantee. -- **Fix:** No change. Do not persist or share the signing key across runs; the re-read-after-restart cost is acceptable versus replay risk. -- **Evidence:** content-hash.ts:66 randomBytes(32) signing key with 'in-process runtime capability' comment; :86-93 readCapabilityMatchesScope; process-str-replace.ts validateReadCapabilityAuthority (scope-mismatch rejection). - -## [LOW] security — packages/agent-runtime/src/process-edit-transaction.ts:300 — [C] KEEP — replace_range authority chain (scope + capability metadata + content re-hash + target-within-range + bounds) - -- **Risk:** What it blocks: a replace_range edit is rejected unless the decoded token matches scope, its (startLine,endLine,hash) equal the declared capability metadata, the ORIGINAL-snapshot content re-hashes to the token hash, the authorization target lies within the covered capability range, and the requested lines are in-bounds. Friction-vs-value: strongly value. This is the whole-file-overwrite floor and the primary stale-overwrite defense for block edits; each check catches a distinct forgery/staleness class. It authenticates against the original snapshot (not shifted working lines), which correctly lets prior in-transaction edits shift the target without re-reading. No redundant re-read is imposed on a fresh, in-range capability. -- **Fix:** No change. -- **Evidence:** process-edit-transaction.ts replace_range case: decode+scope check, capability-metadata equality, getContentHash(observedContent)!==decoded.hash stale check, authorizationTarget within capabilityStart/End, visibleLineCount bounds. - -## [LOW] correctness — packages/agent-runtime/src/process-edit-transaction.ts:265 — [C] KEEP (noted) — overlapping replace_range edits in one transaction are rejected - -- **Risk:** What it blocks: getEffectiveReplaceRangeEdit hard-errors when a replace_range overlaps a prior replace_range in the same transaction ('cannot be applied from the original snapshot'). Friction-vs-value: mostly value. Two edits to the same lines from a single original snapshot have ambiguous line math; refusing avoids silent corruption. Non-overlapping later edits are correctly line-shifted rather than blocked, so the friction is narrow. Relaxing to re-base overlapping edits would reintroduce the exact ordering ambiguity the design removes. -- **Fix:** No change; the safe alternative (coalesce/reorder into one edit) is already available to the agent. -- **Evidence:** process-edit-transaction.ts getEffectiveReplaceRangeEdit: priorRange.startLine<=edit.endLine -> error; else lineShift applied. - -## [LOW] correctness — packages/agent-runtime/src/process-str-replace.ts:84 — [C] KEEP — large-file read-capability enforcement, already softened by deterministic fallback and post-edit anchor echo - -- **Risk:** What it blocks: files over LARGE_FILE_LINE_THRESHOLD (1000) or LARGE_FILE_CHAR_THRESHOLD (100000) require a fresh basedOnRead anchor (enforceReadCapability). Friction-vs-value: value, and already well-tuned against friction. It permits a naked edit when oldString is uniquely identifiable (getDeterministicLargeFileFallbackRange), echoes fresh post-edit anchors on every successful edit (mintAnchorForRange/regionAnchor) so the next edit to the region needs no re-read, and findFreshCapabilityForPath reuses an earlier in-call validated range so a self-staled anchor can be retried without re-reading. These are exactly the 'reuse an echoed post-edit capability' friction reducers the audit asks to preserve. occurrenceIndex also bypasses the anchor requirement. -- **Fix:** No change. This control already distinguishes real stale-write protection from friction correctly. -- **Evidence:** process-str-replace.ts:84-86 thresholds; getDeterministicLargeFileFallbackRange unique-match fallback; mintAnchorForRange/regionAnchor post-edit echo; findFreshCapabilityForPath (improvement #3) reuse path. - -## Coverage receipt - -### Subsystems - -- agent-runtime-edit-authorization - -### Features - -- read-before-edit-gating -- cap-v3-capability-binding -- str-replace-circuit-breaker -- replace-range-authority -- compaction-authorization-revocation - -### Files - -- packages/agent-runtime/src/util/read-authorization.ts -- packages/agent-runtime/src/tools/handlers/tool/edit-read-state.ts -- packages/agent-runtime/src/tools/handlers/tool/str-replace.ts -- packages/agent-runtime/src/tools/handlers/tool/**tests**/str-replace-circuit-breaker.test.ts -- packages/agent-runtime/src/process-str-replace.ts -- packages/agent-runtime/src/process-edit-transaction.ts -- common/src/util/content-hash.ts -- common/src/tools/params/based-on-read.ts diff --git a/.agents/sessions/agent-restriction-audit-2026-07/findings/read-and-path.md b/.agents/sessions/agent-restriction-audit-2026-07/findings/read-and-path.md deleted file mode 100644 index 828ee32c5f..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/findings/read-and-path.md +++ /dev/null @@ -1,120 +0,0 @@ -# Audit findings: read-and-path - -- Subsystems: sdk-read-tools, agent-runtime-read-subtree, common-path-security -- Features: read-files, read-policy, read-subtree, sensitive-path-policy, project-path-containment -- Files covered: 6 - -## [MEDIUM] dependency-hygiene — common/src/util/sensitive-paths.ts:10 — [B] SOFTEN — .crt/.cer public certificates treated as secrets - -- **Risk:** SENSITIVE_EXTENSIONS includes '.crt' and '.cer'. These are X.509 public certificates — the public half of a keypair, distributed openly (CA bundles, server certs, chain files). They contain no secret material. Blocking them stops legitimate work: reading a TLS chain to debug a handshake, inspecting a bundled CA cert, verifying a pinned cert in test fixtures. This is pure friction with no secret-protection payoff; the private material lives in .pem/.key/.p12/.pfx/.jks/.keystore, which remain blocked. -- **Fix:** Remove '.crt' and '.cer' from SENSITIVE_EXTENSIONS. Keep '.pem','.key','.p12','.pfx','.jks','.keystore' (these carry private keys). Optionally, if a repo really ships a private key with a .crt name, rely on the explicit host fileFilter rather than a blanket extension ban. -- **Evidence:** const SENSITIVE_EXTENSIONS = new Set(['.pem','.key','.p12','.pfx','.jks','.keystore','.crt','.cer']) — used at isMandatorySensitiveReadPath line 52 via SENSITIVE_EXTENSIONS.has(extension). - -## [MEDIUM] correctness — common/src/util/sensitive-paths.ts:57 — [B] SOFTEN — basename.includes('kubeconfig') substring match over-blocks docs/scripts - -- **Risk:** The kubeconfig guard uses a substring includes() over the basename, so any file whose name merely mentions kubeconfig is blocked: 'setup-kubeconfig.sh', 'kubeconfig-guide.md', 'generate-kubeconfig.ts', 'kubeconfig.example'. None of these contain live cluster credentials, yet the agent cannot read the setup script or the doc it needs to do the task. The friction hits exactly the files an agent is most likely to want when working on cluster tooling. -- **Fix:** Match the actual credential file, not the substring: block basename === 'kubeconfig' or basename.endsWith('.kubeconfig') (and the conventional '~/.kube/config' when seen), rather than basename.includes('kubeconfig'). Real kubeconfig files remain blocked; scripts/docs about them become readable. -- **Evidence:** return ( envFile || ... || basename.includes('kubeconfig') || basename.includes('.tfstate') ) - -## [LOW] correctness — common/src/util/sensitive-paths.ts:58 — [B] SOFTEN — basename.includes('.tfstate') over-matches templates/examples - -- **Risk:** The tfstate guard also uses substring includes(), so 'terraform.tfstate.example', 'my.tfstate.md', or a doc named 'about-tfstate-files.md' get blocked even though they hold no real state/secrets. Legitimate templates and documentation about state files become unreadable. -- **Fix:** Anchor to the real artifacts: basename.endsWith('.tfstate') || basename.endsWith('.tfstate.backup'). This still blocks generated state and its backup while allowing example/doc files. Real .tfstate is a genuine KEEP target (it embeds resource attributes/secrets), so only the matching precision is loosened. -- **Evidence:** basename.includes('.tfstate') — substring match on the lowercased basename. - -## [LOW] dependency-hygiene — common/src/util/sensitive-paths.ts:18 — [B] SOFTEN — '.yarnrc' / '.yarnrc.yml' blocked wholesale - -- **Risk:** SENSITIVE_BASENAMES bans '.yarnrc' and '.yarnrc.yml' outright. Modern Yarn config (nodeLinker, plugins, packageExtensions, yarnPath) is overwhelmingly non-secret and is exactly the kind of file an agent must read to reason about dependency resolution or workspace layout. Secrets in Yarn are the exception (npmAuthToken lines), unlike '.npmrc' where \_authToken is common. Blocking the whole file for a rare secret line is high friction / low value. -- **Fix:** Drop '.yarnrc' and '.yarnrc.yml' from the mandatory list (keep '.npmrc','.pypirc','.netrc','.htpasswd','auth.json','credentials'). If token leakage from Yarn config is a concern, prefer a content-aware redaction/host fileFilter over a blanket basename ban. -- **Evidence:** SENSITIVE_BASENAMES = new Set(['.htpasswd','.netrc','credentials','.npmrc','.yarnrc','.yarnrc.yml','auth.json','.pypirc','terraform.tfvars','.terraformrc']) - -## [MEDIUM] correctness — sdk/src/tools/path-utils.ts:17 — [B] SOFTEN — absolute in-project paths rejected before containment runs - -- **Risk:** isSafeProjectRelativePath() rejects every path.isAbsolute(input) up front, and read-files calls it as a hard gate in authorizeReadTarget (returns outside_project) before any containment check. But the canonical containment layer resolveProjectPath() already accepts absolute inputs and correctly resolves+contains them (it path.resolve()s absolute inputs and rejects only those that escape the root). So an absolute path that points squarely inside the project (e.g. a path an agent copied from a stack trace or a tool that emits absolute paths) is refused purely on form, not on any real boundary. This is friction: the read is safe but the surface pre-rejects it. -- **Fix:** Let containment be the authority: allow absolute inputs through isSafeProjectRelativePath (still reject NUL bytes, '..' traversal, and Windows drive/UNC ambiguity if genuinely unsupported) and rely on resolveProjectPath/resolveFilePathForFileSystemOperation to reject only paths that actually resolve outside the real project root. The security property is unchanged; the form-based rejection is removed. -- **Evidence:** if ( path.isAbsolute(input) || /^[a-zA-Z]:[\\/]/.test(input) || input.startsWith('\\\\') || input.startsWith('//') ) { return false } — and read-files authorizeReadTarget: if (!isSafeProjectRelativePath(requestedPath)) return outside_project. - -## [MEDIUM] performance — sdk/src/tools/read-files.ts:39 — [B] SOFTEN — MAX_RANGE_READ_BYTES = 1MB forces multiple round-trips on large-file ranges - -- **Risk:** For files above the 10MB whole-file gate, ranged reads are the only way in, and each range is capped at 1MB. A reviewer wanting a 3–4MB slice of a large generated file, lockfile, or bundled artifact must issue several sequential range calls and stitch them mentally. This is a throughput tax on legitimate large-file inspection, not a security control — the file is already in-project and already passed sensitive/ignore policy. -- **Fix:** Raise MAX_RANGE_READ_BYTES to ~4MB (still well under the 10MB whole-file ceiling and any context budget), or make it a caller-tunable parameter. Keep the per-range bound so a single request can't stream an unbounded window; only the size is loosened. -- **Evidence:** export const MAX_RANGE_READ_BYTES = 1_048_576 — passed to readTextRange(fs, operationPath, startLine, endLine, MAX_RANGE_READ_BYTES) and enforced via range.data.byteLength > MAX_RANGE_READ_BYTES. - -## [MEDIUM] performance — packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:27 — [B] SOFTEN — LIVE_SUBTREE_MAX_NODES = 1000 truncates large-directory scans - -- **Risk:** The live subtree scan admits at most 1000 nodes, then marks the result truncated and tells the agent to 'Request a narrower subtree path.' In a monorepo or a large package dir, a single legitimate 'show me this subtree' call gets cut off and the agent must fan out into many narrower calls to see the whole tree — the exact serial round-tripping the tool exists to avoid. The cap is a resource guard, not a security boundary (sensitive/ignore/containment checks are separate and stay). -- **Fix:** Raise the default (e.g. 5000) and/or accept a maxNodes input on read_subtree so callers can opt into a larger scan when they know the directory is big. Retain the hard reservation mechanism (reserveNode) so the scan still terminates deterministically — only the ceiling moves. -- **Evidence:** const LIVE_SUBTREE_MAX_NODES = 1000 — enforced in reserveNode(): if (scan.count >= LIVE_SUBTREE_MAX_NODES) { scan.truncated = true; return false }. - -## [LOW] performance — sdk/src/tools/read-files.ts:38 — [B] SOFTEN — READ_SNAPSHOT_CONCURRENCY = 8 throttles large read batches - -- **Risk:** Both path authorization and snapshot reads run at a fixed concurrency of 8. For a batch of dozens of files (common when an agent pulls a whole feature's worth of sources at once), reads serialize into waves of 8. This is a mild latency tax; it protects against fd/memory pressure but is conservative for typical SSD-backed local reads. -- **Fix:** Raise to ~16, or derive from available parallelism / make it configurable. Low urgency — this is comfort friction, not a blocker, and the current value is safe. -- **Evidence:** export const READ_SNAPSHOT_CONCURRENCY = 8 — used in mapWithConcurrency for both authorizeReadTarget and readCanonicalSnapshot fan-out. - -## [LOW] correctness — sdk/src/tools/read-files.ts:240 — [B] SOFTEN — UTF-16 and non-UTF-8 text reads hard-refused - -- **Risk:** decodeText() refuses UTF-16 (BOM check) with UNSUPPORTED_ENCODING and rejects any non-strict-UTF-8 bytes via TextDecoder({fatal:true}). UTF-16 files are common on Windows (PowerShell output, some editors) and Latin-1/CP-1252 source files still exist. These are legitimate in-project text files an agent may need to read; refusing them is a capability gap with no security value (a secret in a UTF-16 file is not protected by refusing all callers — it's just unreadable to everyone). -- **Fix:** Add a fallback decode path: honor the UTF-16LE/BE BOM and decode it; for non-fatal UTF-8 failures, attempt a lenient decode (or latin1) and flag the file as re-encoded rather than erroring outright. Keep the binary (NUL-byte) refusal — that one is correct. -- **Evidence:** if ((bytes[0]===0xff && bytes[1]===0xfe) || (bytes[0]===0xfe && bytes[1]===0xff)) return unsupported_encoding ... new TextDecoder('utf-8',{fatal:true}).decode(bytes) → catch → unsupported_encoding. - -## [LOW] api-contract — packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:300 — [B] SOFTEN — read_subtree hard-blocks gitignored paths, inconsistent with read_files - -- **Risk:** In buildLiveNode, a gitignored path (isFileIgnored) is returned as a 'blocked' FilesystemError (with an isAgentSessionArtifactPath exception). But read-files deliberately treats ignore as 'a discovery preference, not an authorization boundary' and allows explicit reads of ignored files. So an agent can read a specific ignored file with read_files but cannot see it via read_subtree — two read surfaces disagree on the same file. Ignored build outputs (dist/, generated/) are sometimes exactly what needs inspecting. -- **Fix:** Align the two surfaces: in read_subtree, omit ignored paths from broad auto-listing but do not hard-error an explicitly requested ignored path — or downgrade the ignore result from 'blocked' to an informational/skipped status so explicit inspection stays possible. Keep the mandatory-sensitive and containment checks as the real boundaries. -- **Evidence:** if (normalizedRelativePath && !isAgentSessionArtifactPath(normalizedRelativePath) && (await isFileIgnored({filePath, projectRoot, fs: liveFs}))) return { ok:false, error: subtreeError('blocked', 'Path is ignored by the authorized filesystem policy: ...', false) } — vs read-files.ts comment: 'Ignore files are a discovery preference, not an authorization boundary.' - -## [LOW] security — common/src/util/sensitive-paths.ts:47 — [C] KEEP — .env / .env.\* denial (templates excluded) - -- **Risk:** Blocks '.env' and '.env.\*' while excluding '.env.example/.sample/.template' via isEnvTemplatePath. .env files are the single most common home for live application secrets (DB URLs, API keys). Reading them yields real credentials. -- **Fix:** No change. This is a true secret-protection control and correctly whitelists templates so agents can still read example env files. Retain. -- **Evidence:** const envFile = (basename === '.env' || basename.startsWith('.env.')) && !isEnvTemplatePath(portable); return ( envFile || ... ) - -## [LOW] security — common/src/util/sensitive-paths.ts:54 — [C] KEEP — private-key and credential-file denials (id_rsa, .pem/.key/.p12/.pfx/.jks/.keystore, \_credentials, real .tfstate, .htpasswd/.netrc/.npmrc/.pypirc/auth.json/credentials) - -- **Risk:** These match files that hold live private keys and credentials: SSH private keys (id_rsa/ed25519/dsa/ecdsa, .pub excluded), PKCS/JKS keystores, netrc/htpasswd/pypirc/npmrc auth tokens, AWS-style 'credentials', and Terraform state (which embeds resource secrets). Reading any of them exposes real secret material. -- **Fix:** No change to intent — retain these denials. (The precision fixes noted separately for kubeconfig/.tfstate/.yarnrc/.crt/.cer only tighten matching; the core private-key/credential coverage stays.) -- **Evidence:** /^id\_(rsa|ed25519|dsa|ecdsa)/.test(basename) && !basename.endsWith('.pub') || basename.endsWith('\_credentials') ... SENSITIVE_BASENAMES + SENSITIVE_EXTENSIONS private-key entries. - -## [LOW] security — common/src/util/project-path-containment.ts:101 — [C] KEEP — resolveProjectPath containment (traversal, sibling-prefix, symlink escape) - -- **Risk:** resolveProjectPath rejects '..' traversal, absolute-outside paths, sibling-prefix escapes (e.g. '/repo-evil' vs root '/repo' via path.relative semantics), and symlink dereferences whose realpath lands outside the real project root, while still allowing in-project symlinks that stay inside the repo. This is the actual outside-project boundary and it is precise (it does NOT reject in-project symlinks or files whose names start with '..', e.g. '..config'). -- **Fix:** No change. This is high-value, correctly-scoped containment; removing or loosening it would allow reads outside the project. Retain. (Note it already permits legitimate in-project symlinks and sibling-name edge cases, so it is not over-strict.) -- **Evidence:** relativeLexical === '..' || relativeLexical.startsWith('..' + path.sep) || path.isAbsolute(relativeLexical) || relativeLexical.split(path.sep).includes('..') → null; and symlink check: realRelative === '..' || realRelative.startsWith('..'+sep) || path.isAbsolute(realRelative) → null. - -## [LOW] performance — sdk/src/tools/read-files.ts:41 — [C] KEEP — MAX_FILE_BYTES = 10MB whole-file ceiling (with range fallback) - -- **Risk:** Files over 10MB cannot be read whole; the tool returns too_large with a 'read an exact bounded range instead' recovery and range reads still work. 10MB of text is far beyond any reasonable context window, so a whole-file read of a larger file would either blow the context or be truncated anyway. The ceiling is a genuine resource/context guard and it degrades gracefully (ranges remain available). -- **Fix:** No change. The ceiling is generous and has a clean range-based escape hatch, so it is not meaningful friction. Retain. (If anything, pair this with the MAX_RANGE_READ_BYTES bump above so the range fallback is less chatty.) -- **Evidence:** const MAX*FILE_BYTES = 10 * 1024 \_ 1024; if (sizeBytes > MAX_FILE_BYTES) { ... return too_large '... exceeds 10MB limit. Read an exact bounded range instead.' } with recovery: 'read_smaller_range'. - -## [LOW] security — sdk/src/tools/read-policy.ts:11 — [C] KEEP — isReadPathBlocked composition (mandatory-sensitive OR host fileFilter) - -- **Risk:** isReadPathBlocked composes the mandatory sensitive-path policy with the host-supplied fileFilter over both raw and lowercased aliases, and read-files' authorizeReadTarget applies the same alias check against both the requested and canonical relative paths. This is the correct fail-closed composition point for secret protection and host policy; the alias/case handling closes casing-bypass gaps. -- **Fix:** No change. This is the enforcement seam that makes the KEEP denials effective and lets hosts add their own blocks. Retain. -- **Evidence:** return aliases.some(isMandatorySensitiveReadPath) || Boolean(fileFilter && aliases.some(alias => fileFilter(alias).status === 'blocked')) — mirrored in read-files authorizeReadTarget with uniquePolicyAliases(resolved.relativePath, canonicalRelative). - -## Coverage receipt - -### Subsystems - -- sdk-read-tools -- agent-runtime-read-subtree -- common-path-security - -### Features - -- read-files -- read-policy -- read-subtree -- sensitive-path-policy -- project-path-containment - -### Files - -- sdk/src/tools/read-policy.ts -- sdk/src/tools/read-files.ts -- sdk/src/tools/path-utils.ts -- common/src/util/sensitive-paths.ts -- common/src/util/project-path-containment.ts -- packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts diff --git a/.agents/sessions/agent-restriction-audit-2026-07/findings/spawn-permissions.md b/.agents/sessions/agent-restriction-audit-2026-07/findings/spawn-permissions.md deleted file mode 100644 index cd74679a62..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/findings/spawn-permissions.md +++ /dev/null @@ -1,82 +0,0 @@ -# Audit findings: spawn-permissions - -- Subsystems: agent-runtime/spawn -- Features: spawn-permissions, handoff-capability-derivation, spawn-depth, filesystem-scope-narrowing, tool-intersection -- Files covered: 12 - -## [HIGH] correctness — packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:479 — [B] Empty handoff readablePaths narrows a child from unrestricted reads to zero-read (all reads hard-blocked) - -- **Risk:** deriveSpawnTemplateCapabilities always rewrites filesystemScope to { read, write } (line 517) whenever a handoff is present. narrowFilesystemPatterns (util/filesystem-scope.ts:56) returns the normalized requested array verbatim, so when handoff.permissions.readablePaths is [] (a common and schema-legal case — e.g. the repair-editor sample handoff in spawn-agents-permissions.test.ts:createVersionedHandoff uses readablePaths: []), the derived child gets filesystemScope.read = []. Critically, if the child template originally had NO filesystemScope.read (undefined = unrestricted reads), derivation CONVERTS that unrestricted child into read: []. In tool-executor.ts:1238-1241 an empty (but defined) array is truthy, so the scope block runs; at 1272-1278 every in-project path is a scopeMismatch because [].some(...) is false; and at 1296-1299 read-access scope mismatches are hard-blocked. Net effect: a child handed an empty readablePaths list cannot read ANY file in the project, silently breaking legitimate repair/review/edit work that depends on reading source. The parent orchestrator rarely enumerates a full read allowlist, so this trap fires on ordinary handoffs. -- **Fix:** Only narrow read scope when the handoff actually requested read paths. In deriveSpawnTemplateCapabilities, compute read as: handoff.permissions.readablePaths.length > 0 ? narrowFilesystemPatterns({requested: readablePaths, staticPatterns: inheritedTemplate.filesystemScope?.read, ...}) : inheritedTemplate.filesystemScope?.read (i.e. preserve the child's static read scope, keeping undefined => unrestricted). Then build filesystemScope conditionally so an undefined static read scope is not clobbered into []. This keeps the child-cannot-exceed-parent guarantee (a non-empty requested set is still validated against staticPatterns) while removing the accidental total read lockout. -- **Evidence:** spawn-agent-utils.ts:479-492 (read/write = narrowFilesystemPatterns(requested: handoff.permissions.readablePaths, staticPatterns: inheritedTemplate.filesystemScope?.read)) and :517 (filesystemScope: { read, write }); util/filesystem-scope.ts:56-79 returns normalized (=[] for empty requested) with staticPatterns undefined only rejecting ../absolute; tool-executor.ts:1238-1241 (allowedPatterns truthy for empty array), :1272-1278 (scopeMismatch when no pattern matches), :1296-1299 (read scope mismatch hard-blocked). Test spawn-agents-permissions.test.ts createVersionedHandoff sets permissions.readablePaths: []. - -## [MEDIUM] correctness — packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:516 — [B] Any handoff-carrying child has spawnableAgents forcibly emptied, blocking legitimate delegation - -- **Risk:** When a handoff is present, the derived template hard-sets spawnableAgents: [] (line 516). Without a handoff the early return at line 460 preserves the child's static spawnableAgents. This asymmetry means the moment an orchestrator issues a structured handoff (the recommended, richer path), the child loses ALL ability to spawn sub-agents it legitimately declares — e.g. a reviewer/debugger/repair-editor handed a task can no longer delegate discovery to file-picker/code-searcher even though its own template authorizes those children. The stated intent (prevent authority widening via delegation) does not require zeroing: a child can never spawn an agent it does not statically declare in spawnableAgents anyway (getMatchingSpawn enforces that against the child's own list), and spawn depth caps bound recursion. Zeroing is redundant with those guards and purely additive friction. -- **Fix:** Preserve the child's own static spawnableAgents instead of forcing []: drop the `spawnableAgents: []` override (let it inherit ...inheritedTemplate.spawnableAgents). If defense-in-depth is desired, intersect with the child's static list rather than empty it — but the static list is already the authoritative ceiling enforced by getMatchingSpawn, so simply preserving it restores parity with the no-handoff path. -- **Evidence:** spawn-agent-utils.ts:460 (`if (!handoff) return inheritedTemplate` preserves spawnableAgents) vs :505-518 return object with `spawnableAgents: []`; getMatchingSpawn (:190-250) already restricts children to the spawner's declared list; executeSubagent depth cap at :~1980 (`currentDepth + 1 > maxSpawnDepth`) bounds recursion independently. - -## [MEDIUM] api-contract — packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:513 — [B] Trusted, model-hidden programmaticToolNames are intersected with the model-authored handoff allowedTools and silently stripped - -- **Risk:** programmaticToolNames are hidden capabilities callable only from the trusted handleSteps generator (agent-template.ts:~168 doc), never exposed to the model. Line 513-515 filters them to only those present in handoff.permissions.allowedTools. Because the model authoring a handoff has no visibility into hidden programmatic tools, it will essentially never list them, so this intersection strips ALL of a child's programmatic tools whenever a handoff is present. That can break a child whose handleSteps generator depends on a programmatic tool (the generator yields a tool call that the executor then rejects as unavailable), turning a capability-scoping mechanism meant for model-facing tools into a silent breakage of trusted internal wiring. -- **Fix:** Do not gate programmaticToolNames on the model-authored handoff allowedTools. Preserve inheritedTemplate.programmaticToolNames unchanged (they are already bounded by the template author and unreachable by the model). If scoping is ever needed, gate on an explicit programmatic allowlist field, not the model-visible allowedTools set. -- **Evidence:** spawn-agent-utils.ts:513-515 (`programmaticToolNames: (inheritedTemplate.programmaticToolNames ?? []).filter((toolName) => requestedTools.has(toolName))`); requestedTools derives from handoff.permissions.allowedTools (:462); agent-template.ts programmaticToolNames doc 'Hidden capabilities callable only from the trusted handleSteps generator'. - -## [LOW] correctness — packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:507 — [B] Model-facing toolNames intersection with handoff allowedTools can silently drop tools the child legitimately needs - -- **Risk:** Line 507-510 intersects the child's static toolNames with handoff.permissions.allowedTools. This is the intended 'handoff scopes down' contract and correctly prevents widening, but the failure mode is silent under-provisioning: if the orchestrator's handoff omits a tool the child genuinely needs (e.g. forgets set_output or a required discovery tool), the child loses it with no diagnostic, and the model only discovers the gap when a tool call is rejected mid-run. getEffectiveAgentToolNames adds set_output back for structured_output agents at the executor layer, but the derived toolNames array here does not, so the reported/effective set can be narrower than needed. -- **Fix:** Keep the intersection (it is the child-cannot-exceed-parent contract) but reduce friction: (a) always retain set_output when outputMode is structured_output regardless of the handoff list, mirroring getEffectiveAgentToolNames; and (b) when the intersection drops a statically-declared tool, emit a debug/warning log in logAgentSpawn so under-provisioned handoffs are diagnosable rather than silent. -- **Evidence:** spawn-agent-utils.ts:507-510 (toolNames filtered by requestedTools.has); util/agent-tool-names.ts:12-24 adds set_output only at effective-tools computation, not in the derived array; selectAgentAttempt uses getEffectiveAgentToolNames so it re-adds set_output for the requiredTools check but the child's runtime toolNames array stays narrowed. - -## [LOW] correctness — packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts:116 — [B] MAX_SPAWN_BATCH_SIZE=8 sibling cap forces wave-splitting for legitimate wide fan-out - -- **Risk:** A single spawn_agents call is capped at 8 sibling agents (constants/agents.ts MAX_SPAWN_BATCH_SIZE = 8; enforced spawn-agents.ts:116-120). Legitimate breadth-first work (e.g. sharding 10-12 files or audit shards across a codebase) must be artificially split into multiple waves, adding orchestration round-trips. The cap protects against runaway fan-out, but 8 is a low ceiling for the sharding/audit patterns this runtime explicitly supports (evals/buffbench plan-sharding). -- **Fix:** Raise MAX_SPAWN_BATCH_SIZE to 12-16, or make it configurable per-template (like maxSpawnDepth). The background-agent scheduling quota (maxRunningForRoot=8 in select-agent-attempt) already bounds concurrent resource use, so a larger batch of foreground/queued siblings does not meaningfully increase blast radius. -- **Evidence:** common/src/constants/agents.ts (`MAX_SPAWN_BATCH_SIZE = 8`); spawn-agents.ts:116-120 (`if (agents.length > MAX_SPAWN_BATCH_SIZE) throw ... 'Split the work into bounded waves.'`); spawn-agents.ts selection uses maxRunningForRoot: 8 as an independent concurrency guard. - -## [LOW] correctness — common/src/constants/agents.ts:243 — [C] Spawn depth default of 3 is shallow for deep orchestration but is per-template configurable — KEEP - -- **Risk:** MAX_SPAWN_DEPTH_DEFAULT = 3 permits root -> specialist -> leaf tool-runner. A 4-level orchestration (orchestrator -> sub-orchestrator -> specialist -> tool-runner) is rejected by executeSubagent's depth check. This is a genuine recursion-safety guard and is overridable per template via maxSpawnDepth (agent-template.ts maxSpawnDepth; spawn-depth.test.ts confirms overrides both raise and lower the cap), so it is not clearly broken. -- **Fix:** KEEP. No change required; the cap is configurable per-template and the default prevents unbounded file-picker -> file-picker recursion. If deeper orchestration becomes common, raise the default to 4 in one edit, but do not remove the cap. -- **Evidence:** constants/agents.ts MAX_SPAWN_DEPTH_DEFAULT = 3 (doc: 'root -> specialist -> leaf tool-runner'); spawn-agent-utils.ts executeSubagent depth guard (`currentDepth + 1 > maxSpawnDepth ... 'Maximum spawn depth reached'`); spawn-depth.test.ts covers default, per-template higher (5) and lower (1) overrides. - -## [LOW] security — packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:444 — [C] Plan-only ancestry attenuates child run_terminal_command to read-only — KEEP - -- **Risk:** When the parent runs planOnly, a child that has run_terminal_command has its terminalPermissionProfile forced to 'read-only' (lines 449-453) and planOnly is propagated. This is correct, high-value defense: it prevents a plan-mode session from acquiring terminal write authority by spawning a bashing child. spawn-agents-permissions.test.ts verifies both the attenuation under plan-only ancestry and that normal ancestry preserves workspace-write, and confirms the shared template is not mutated. -- **Fix:** KEEP. This is a real child-cannot-exceed-parent enforcement and should not be relaxed. -- **Evidence:** spawn-agent-utils.ts:444-459 (inheritsPlanOnlyAuthority downgrade to 'read-only'); tests 'attenuates terminal authority throughout plan-only spawn ancestry' and 'preserves normal child terminal authority outside plan-only ancestry'. - -## [LOW] security — packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:469 — [C] Tool-widening throw + closed read-only discovery carve-out — KEEP - -- **Risk:** A handoff may only grant the closed HANDOFF_GRANTABLE_READ_ONLY_TOOLS allowlist (code_search, glob, read_outline, read_subtree, list_directory, query_index, find_files, find_files_matching_content) beyond the child's static tools; anything else (mutation/network/process/delegation, and deliberately read_files) throws (lines 469-477). This is the core child-cannot-exceed-parent invariant with a well-scoped, security-reasoned exception for no-authority discovery tools. narrowFilesystemPatterns similarly throws on any requested path outside the static scope or that escapes the project root. -- **Fix:** KEEP. The carve-out is explicit, closed, and documented (read_files intentionally excluded because it issues read authorizations edit tools can consume). No relaxation warranted. -- **Evidence:** spawn-agent-utils.ts:415-433 HANDOFF_GRANTABLE_READ_ONLY_TOOLS with the explicit 'do NOT add read_files' rationale; :469-477 disallowedTools throw; util/filesystem-scope.ts:66-78 widen/escape throw; tests 'still throws the widen error for a mutation tool', 'does not grant read_files through the discovery carve-out'. - -## Coverage receipt - -### Subsystems - -- agent-runtime/spawn - -### Features - -- spawn-permissions -- handoff-capability-derivation -- spawn-depth -- filesystem-scope-narrowing -- tool-intersection - -### Files - -- packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts -- packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts -- packages/agent-runtime/src/tools/handlers/tool/spawn-agent-inline.ts -- packages/agent-runtime/src/util/agent-tool-names.ts -- packages/agent-runtime/src/util/filesystem-scope.ts -- packages/agent-runtime/src/orchestration/select-agent-attempt.ts -- packages/agent-runtime/src/tools/tool-executor.ts -- common/src/constants/agents.ts -- common/src/types/agent-template.ts -- agents/types/agent-definition.ts -- packages/agent-runtime/src/**tests**/spawn-agents-permissions.test.ts -- packages/agent-runtime/src/**tests**/spawn-depth.test.ts diff --git a/.agents/sessions/agent-restriction-audit-2026-07/findings/tool-arg-allowlists.md b/.agents/sessions/agent-restriction-audit-2026-07/findings/tool-arg-allowlists.md deleted file mode 100644 index 35a263211b..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/findings/tool-arg-allowlists.md +++ /dev/null @@ -1,94 +0,0 @@ -# Audit findings: tool-arg-allowlists - -- Subsystems: sdk-search-tools, cli-git-slash-commands, agents-security-reviewer, common-tool-params, agent-runtime-tool-handlers -- Features: ripgrep-flag-allowlist, git-pathspec-arg-parsing, security-reviewer-params, glob-list-directory-params, search-result-caps -- Files covered: 8 - -## [MEDIUM] api-contract — sdk/src/tools/find-files-matching-content.ts:635 — [B] parseSafeRipgrepFlags rejects combined short flags (e.g. -ni), forcing a re-search - -- **Risk:** The allowlist matches whole tokens against fixed Sets (switchesWithoutValue L635, switchesWithValue L651). A single-dash bundle such as `-ni` (= -n + -i) or `-iw` is never split, so it fails the exact-token check and returns `Unsupported ripgrep flag '-ni'` (unsupportedFlag L713-721). `-n` is habitually emitted by models and is even special-cased as a no-op (redundantLineNumberSwitches L660,L668), but the moment it is bundled with another safe short flag the whole search is rejected. This is the exact 'Unsupported ripgrep flag -ni' friction noted in the prior run: the rejected content is a benign case-insensitive line-numbered search, not an injection attempt. -- **Fix:** Before the per-token loop, expand any token matching /^-[a-zA-Z]{2,}$/ into individual `-x` short flags, then validate each against the existing Sets. Only expand pure single-dash alpha bundles (never `--long` or value-bearing forms). This keeps the dangerous-flag allowlist intact while accepting `-ni`, `-iw`, `-in`, etc. Minimal, mechanical relaxation. -- **Evidence:** L635 `const switchesWithoutValue = new Set([...])`; L660 `redundantLineNumberSwitches = new Set(['-n','--line-number'])`; L713-721 unsupportedFlag() rejects any unrecognized token. No branch splits `-ni` into `-n`+`-i`. - -## [MEDIUM] api-contract — sdk/src/tools/find-files-matching-content.ts:651 — [B] find_files_matching_content blocks -o/--only-matching and -c/--count, common harmless output modes - -- **Risk:** The value/no-value Sets (L635-659) permit only case/word/fixed/multiline modifiers plus -g/-t/-T. Harmless, read-only output selectors that models routinely reach for — `-o`/`--only-matching`, `-c`/`--count`/`--count-matches`, `-v`/`--invert-match` — are all rejected with the generic Unsupported-flag error (L719). None of these mutate files or run commands; they only change how ripgrep reports matches. The rejection produces a dead-end for a legitimate query and pushes the agent to retry or abandon. -- **Fix:** Add `-v`/`--invert-match` and `-c`/`--count`/`--count-matches` (no-value) and, for code_search where JSON parsing tolerates it, `-o`/`--only-matching` to the allowlist. `--count` for this tool changes stdout framing (counts, not paths) so if that breaks the -l/JSON parser, either translate it to a post-filter or document it as code_search-only. Keep excluding the genuinely dangerous flags (--pre, -r/--replace, -z/--null, --files, -0). -- **Evidence:** L651-659 switchesWithValue only has -g/--glob,-t/--type,-T/--type-not; L635-649 no-value set lacks -o,-c,--count,-v. Allowed-flags string at L719 confirms the closed set. - -## [LOW] api-contract — sdk/src/tools/code-search.ts:108 — [C] code_search context flags (-A/-B/-C) are already allowed — prior 'context flags blocked' concern does not apply here - -- **Risk:** The prior run's worry that `-A/-B/-C` are blocked is TRUE only for find_files_matching_content, not code_search. code_search explicitly adds `-A`,`-B`,`-C`,`--after-context`,`--before-context`,`--context` via extraSwitchesWithValue (L108-119) because its JSON parser already handles context events. So there is no over-strictness to fix for context flags in code_search; the correct guidance (already in unsupportedFlag recovery text) is to route context-needing searches to code_search. -- **Fix:** No change to code_search context handling. Optionally improve the find_files recovery message so it names -A/-B/-C explicitly (it already says 'Use code_search only when you need its documented context flags.'). Keep as-is (C). -- **Evidence:** L100-107 comment 'code_search additionally allows the -A/-B/-C context flags'; L108-119 extraSwitchesWithValue lists all six context forms. - -## [MEDIUM] api-contract — cli/src/commands/git-command-args.ts:1 — [B] FORBIDDEN_SHELL_CHARACTERS blocks ( ) [ ] { } for /git slash commands, rejecting valid git pathspecs - -- **Risk:** FORBIDDEN_SHELL_CHARACTERS = /[\n\r;$`|&<>()[\]{}\\]/ (L1) is tested against the RAW input BEFORE tokenization (L5), and every arg is later single-quoted by quoteShellArgument (L38-39) before being joined into the final `git diff|status ...`string. Because each argument is single-quoted, brackets/braces/parens cannot be shell-expanded — yet they are blocked at parse time. This rejects perfectly valid git usage: bracket globs`git diff 'src/\*\*/[abc].ts'`, brace sets `git diff 'src/{a,b}.ts'`, and magic pathspecs `git diff ':(exclude)dist'`. The block conflates glob/expansion metacharacters (neutralized by single-quoting) with true shell operators (`; | & $ \` < > backtick`) that are the real injection risk. -- **Fix:** Split the class: keep hard-blocking the genuine injection set `[\n\r;$\`|&<>\\]`(these break out of or chain commands even when the parser mis-handles them). Allow`( ) [ ] { }`through the parser since quoteShellArgument single-quotes them and buildSafeGitCommand only ever targets`diff`/`status`. Add a test asserting `parseSafeGitArgs(":(exclude)dist")`and`'{a,b}.ts'` survive and are correctly single-quoted. -- **Evidence:** L1 regex includes ()[]{}; L5 pre-tokenization test throws 'Shell operators and expansions are not allowed here.'; L38 quoteShellArgument single-quotes each arg; L44-45 buildSafeGitCommand joins quoted args. - -## [LOW] api-contract — agents/security-reviewer/security-reviewer.ts:34 — [B] security-reviewer rejects snapshot_id and accepts only snapshot_fingerprint (strict param-name allowlist) - -- **Risk:** inputSchema.params.required = ['changed_files','snapshot_fingerprint'] (L34) and the spawnerPrompt explicitly warns `snapshot_id is not accepted` (L10-11). A caller that passes the very common `snapshot_id` key gets a spawn-time validation failure even though the intent is identical. This is a naming-allowlist strictness, not a safety control — the token is opaque and simply echoed back as snapshotFingerprint. -- **Fix:** Accept `snapshot_id` as an alias for `snapshot_fingerprint` (normalize either key to snapshotFingerprint), or relax the required check to accept whichever of the two is present. Keep requiring exactly one of them so the echo-back invariant holds. Low-risk, purely ergonomic. -- **Evidence:** L10-11 spawnerPrompt 'snapshot_id is not accepted'; L22-30 only snapshot_fingerprint property declared; L34 required list. - -## [LOW] api-contract — agents/security-reviewer/security-reviewer.ts:128 — [C] security-reviewer toolNames are read-only by design — keep - -- **Risk:** toolNames is limited to read_files, read_outline, code_search, git_status, set_output with spawnableAgents: [] (L128-137). This is an intentional least-privilege boundary for an adversarial reviewer that must not mutate code ('Do not modify code. Review only.' in instructionsPrompt). There is no readable-path glob restriction on the reviewer itself; the SECURITY_SENSITIVE_GLOBS list lives in base2.ts as a spawn TRIGGER, not a read filter. No over-strictness on readable paths exists to relax. -- **Fix:** Keep. The read-only tool set is a correct guardrail, not friction. If reviewers need broader discovery, add find_files_matching_content (also read-only) rather than write tools. -- **Evidence:** L128-136 toolNames read-only list; L137 spawnableAgents: []; instructionsPrompt 'Do not modify code. Review only.' - -## [LOW] test-coverage — agents/**tests**/security-glob-parity.test.ts:97 — [C] security-glob-parity test enforces doc/matcher sync — keep - -- **Risk:** This test extracts SECURITY_SENSITIVE_GLOBS and SECURITY_SENSITIVE_NAME_SUBSTRINGS from base2.ts and asserts every token is documented in securityReviewSection prose (L97-128). It is a cohesion invariant, not a runtime argument restriction, so it does not block any user/agent action. It only fails CI if the two lists drift. No end-user friction. -- **Fix:** Keep as-is. It prevents silent divergence between the phase-gate matcher and the model-facing prose. Not a candidate for relaxation. -- **Evidence:** L102 asserts SECURITY_SENSITIVE_NAME_SUBSTRINGS === ['secret','token','apikey']; L110-117 asserts every glob token documented; L119-127 asserts '.env' documented. - -## [MEDIUM] correctness — sdk/src/tools/code-search.ts:43 — [B] code_search default per-file cap maxResults=15 can silently hide relevant matches - -- **Risk:** maxResults defaults to 15 per file (L43) and globalMaxResults to 250 (L44). Files exceeding 15 matches are truncated and recorded in filesLimitedByMaxResults, surfaced as 'limited to 15 results per file'. The truncation IS reported (good), but 15 is low for real searches (e.g. finding every call site of a common symbol in one large file), so agents can act on partial results and miss occurrences. This is a usability cap, not injection protection. -- **Fix:** Raise the per-file default (e.g. 30-50) or make it clearly overridable per call, while keeping globalMaxResults and maxOutputStringLength as the real memory guards. The truncation message already exists, so the main change is a more generous default. Soften, don't remove — the global/byte caps stay. -- **Evidence:** L43 `maxResults = 15`; L44 `globalMaxResults = 250`; truncation surfaced via filesLimitedByMaxResults and 'limited to ${maxResults} results per file' message near close handler. - -## [LOW] correctness — sdk/src/tools/find-files-matching-content.ts:50 — [B] find_files_matching_content maxFiles=100 cap truncates discovery but is surfaced - -- **Risk:** maxFiles defaults to 100 (L50) with HARD_MATCH_LIMIT=5000 (L33). When exceeded, results stop and `truncated: true` plus a 'results capped' message are returned (buildSuccessPayload / buildMessage), so truncation is not silent. Still, 100 files is easy to hit on broad patterns in a monorepo, causing incomplete file discovery that the agent may not compensate for. -- **Fix:** Consider raising the default maxFiles (e.g. 250) or emphasizing the truncated flag more strongly in the message so agents narrow the pattern. Keep HARD_MATCH_LIMIT as the memory backstop. Minor soften. -- **Evidence:** L50 `maxFiles = 100`; L33 `HARD_MATCH_LIMIT = 5_000`; buildMessage emits '(results capped; consider narrowing the pattern or flags)'. - -## [LOW] api-contract — common/src/tools/params/tool/glob.ts:8 — [C] glob and list_directory params carry no result cap or pattern restriction — nothing to relax - -- **Risk:** The glob inputSchema only requires a non-empty pattern (L10-16, min(1)) and an optional cwd; there is no result cap, denylist, or pattern restriction in the param schema. The handler (packages/agent-runtime/src/tools/handlers/tool/glob.ts L8-20) simply forwards to the client with no filtering. list-directory.ts likewise only takes a path. So the 'glob result caps that hide files' concern does not originate here — any capping happens client-side, outside these files. -- **Fix:** No change in these files. If a client-side glob cap exists and hides files, audit that layer separately; the param/handler layer imposes no over-strict restriction. -- **Evidence:** glob.ts L10-16 pattern.min(1) only; no maxResults field; handler glob.ts L18-19 `return { output: await requestClientToolCall(toolCall) }`; list-directory.ts L9-11 path only. - -## Coverage receipt - -### Subsystems - -- sdk-search-tools -- cli-git-slash-commands -- agents-security-reviewer -- common-tool-params -- agent-runtime-tool-handlers - -### Features - -- ripgrep-flag-allowlist -- git-pathspec-arg-parsing -- security-reviewer-params -- glob-list-directory-params -- search-result-caps - -### Files - -- sdk/src/tools/code-search.ts -- sdk/src/tools/find-files-matching-content.ts -- cli/src/commands/git-command-args.ts -- agents/security-reviewer/security-reviewer.ts -- agents/**tests**/security-glob-parity.test.ts -- common/src/tools/params/tool/glob.ts -- common/src/tools/params/tool/list-directory.ts -- packages/agent-runtime/src/tools/handlers/tool/glob.ts diff --git a/.agents/sessions/agent-restriction-audit-2026-07/findings/web-network.md b/.agents/sessions/agent-restriction-audit-2026-07/findings/web-network.md deleted file mode 100644 index 99a20cf320..0000000000 --- a/.agents/sessions/agent-restriction-audit-2026-07/findings/web-network.md +++ /dev/null @@ -1,90 +0,0 @@ -# Audit findings: web-network - -- Subsystems: agent-runtime-web-search, researcher-agents, terminal-command-policy-network -- Features: web-search-tool, url-fetch, ssrf-guard, researcher-web, researcher-docs -- Files covered: 6 - -## [MEDIUM] performance — packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts:5 — [B] MAX_WEB_FETCH_BYTES = 512 KB truncates real documentation pages - -- **Risk:** The 512,000-byte hard cap on the raw response body is applied before HTML is stripped. Many legitimate public docs pages (MDN references, API references, single-page framework docs, GitHub rendered files) ship well over 512 KB of raw HTML, so the agent silently loses the tail of the page or (when content-length is declared) gets a hard error. This is friction against the tool's core value: fetching real documentation. It is NOT an SSRF control — SSRF is enforced separately by assertSafePublicWebUrl. -- **Fix:** Soften: raise MAX_WEB_FETCH_BYTES to roughly 2 MB (2_000_000). This keeps a sane memory bound while covering the large majority of real docs pages. The streaming reader already caps memory incrementally, so a larger ceiling does not change the worst-case buffering shape materially. -- **Evidence:** web-search-utils.ts:5 `export const MAX_WEB_FETCH_BYTES = 512_000`; consumed as the default in readResponseTextWithLimit (line ~113) and by both the search and URL-fetch branches of web-search.ts. - -## [MEDIUM] performance — packages/agent-runtime/src/tools/handlers/tool/web-search.ts:19 — [B] MAX_FETCH_LENGTH = 50 KB truncates stripped page text too aggressively - -- **Risk:** After HTML stripping, the returned content is truncated to 50,000 characters with a '[Content truncated]' marker. For a large API reference or long tutorial, 50 KB of plain text is easily exceeded, so the agent loses the second half of a page it explicitly asked to read. This compounds with the 512 KB raw cap: a big page can be cut twice. Pure output-shaping limit, no security role. -- **Fix:** Soften: raise MAX_FETCH_LENGTH to ~150_000–200_000 characters. This still bounds the model context contribution but lets typical long docs pages return whole. Optionally make it proportional to the raw byte cap chosen above. -- **Evidence:** web-search.ts:19 `const MAX_FETCH_LENGTH = 50_000`; applied at line ~101 `content.length > MAX_FETCH_LENGTH ? content.slice(0, MAX_FETCH_LENGTH) + '...[Content truncated ...]'`. - -## [LOW] correctness — packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts:113 — [B] Declared content-length over the cap causes a hard error instead of truncation - -- **Risk:** readResponseTextWithLimit throws `Response exceeds N byte limit` whenever the server declares a content-length larger than maxBytes, but a server that omits content-length streams and is gracefully truncated instead. This is an inconsistent, over-strict outcome: two equivalent large pages behave differently based solely on whether the origin sent a content-length header, and a well-behaved large docs host (which sets content-length) is the one that fails outright. Real value lost with no security benefit — the stream is already capped byte-by-byte below. -- **Fix:** Soften: on declared-length-over-limit, do not throw. Fall through to the streaming reader and let it truncate to maxBytes with truncated=true, matching the no-content-length path. Only the streamed cap is needed to bound memory. -- **Evidence:** web-search-utils.ts:~113 `if (Number.isFinite(declaredLength) && declaredLength > maxBytes) { await params.response.body?.cancel(); throw new Error('Response exceeds ...') }` versus the streaming branch below that sets `truncated = true` gracefully. - -## [LOW] performance — packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts:3 — [C/B] WEBSEARCH_TIMEOUT_MS = 30 s is acceptable; keep, minor bump optional - -- **Risk:** A 30-second overall timeout covers both DuckDuckGo search and single-page fetch. It is generally adequate; slow origins occasionally exceed it but rarely. Aggressively lowering would hurt; there is no strong case to raise it either since a hung fetch shouldn't stall an agent step indefinitely. -- **Fix:** Keep at 30 s. If deep-research flows report timeouts against slow-but-legitimate hosts, a modest bump to 45 s for the URL-fetch branch (not search) is the minimal relaxation. No change required otherwise. -- **Evidence:** web-search-utils.ts:3 `export const WEBSEARCH_TIMEOUT_MS = 30_000`; used as default AbortSignal.timeout in executeWebSearch and combined via AbortSignal.any in web-search.ts URL and search branches. - -## [LOW] api-contract — packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts:6 — [C] MAX_WEB_FETCH_REDIRECTS = 5 — keep - -- **Risk:** Manual redirect handling caps at 5 hops and re-validates every hop through assertSafePublicWebUrl (this is a real SSRF defense against redirect-to-internal, verified by the 'revalidates redirect destinations' test). Five hops is enough for normal canonicalization/shortener chains. -- **Fix:** Keep. The per-hop revalidation is essential SSRF protection and must not be relaxed. If legitimate multi-hop chains ever fail, 8–10 is a safe ceiling, but no evidence this is hit today. -- **Evidence:** web-search-utils.ts:6 `MAX_WEB_FETCH_REDIRECTS = 5`; fetchPublicWebUrl loops with `redirect: 'manual'` and calls `assertSafePublicWebUrl(new URL(location, current).href)` each hop; test 'revalidates redirect destinations' confirms 127.0.0.1 redirect is rejected after 1 fetch. - -## [LOW] security — packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts:63 — [C] SSRF host/IP blocking (localhost, metadata, private/link-local/reserved ranges) — KEEP - -- **Risk:** isBlockedWebAddress + assertSafePublicWebUrl block loopback, RFC1918, CGNAT (100.64/10), link-local (169.254), cloud metadata (169.254.169.254 and named metadata hosts), IPv6 ULA/link-local/mapped, TEST-NET and multicast, plus http-auth credentials and non-http(s) schemes, and resolves DNS to catch rebinding. This is the core SSRF guard and is explicitly out of scope for relaxation. -- **Fix:** Keep entirely. No relaxation. The default-deny in isBlockedWebAddress (returns true for anything that isn't a parseable public IPv4/IPv6) is correct fail-closed behavior. -- **Evidence:** web-search-utils.ts:63 `isBlockedWebAddress`, :74 `assertSafePublicWebUrl` (protocol check, credentials check, BLOCKED_HOSTNAMES + .localhost/.local/.internal suffix check, isIP branch, async lookup + blocked-address scan). Covered by web-search-security.test.ts 'blocks loopback...' and 'rejects unsafe URL forms'. - -## [LOW] security — packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts:90 — [C] .local / .internal / .localhost suffix + credential + non-HTTP(S) blocks — keep - -- **Risk:** These reject non-public TLD suffixes (mDNS .local, private .internal, .localhost), URLs carrying user:password credentials, and any non-http(s) scheme (file:, gopher:, etc.). None of these correspond to legitimate public documentation hosts; .local/.internal are reserved for private networks and credential-in-URL fetches are an exfil/SSRF vector. -- **Fix:** Keep. The only conceivable friction is a real API that requires HTTP basic auth embedded in the URL, but that is not a documentation-browsing use case and embedding secrets in a fetched URL is undesirable regardless. No relaxation. -- **Evidence:** web-search-utils.ts:~90 suffix checks `hostname.endsWith('.localhost'|'.local'|'.internal')`; :~83 `parsed.username || parsed.password` credential reject; :~79 protocol reject. Tests assert 'file:///etc/passwd' -> 'HTTP(S)' and 'user:secret@8.8.8.8' -> 'credentials'. - -## [LOW] security — agents/researcher/researcher-web.ts:96 — [C] researcher-web lexical SSRF pre-check (isSsrfUrl) — keep as defense-in-depth - -- **Risk:** researcher-web performs its own lexical SSRF check on any URL extracted from the prompt (PRIVATE_HOST_BLOCKLIST, isPrivateIpv4/Ipv6) and falls back to query mode instead of URL mode when unsafe. It cannot do async DNS (serialized generator), so it is intentionally a lexical layer in front of the backend's authoritative DNS-aware guard. There is no domain allowlist here — sourceDomains is only an optional `site:` query hint, not a restriction — so no legitimate public host is blocked. -- **Fix:** Keep. This is redundant-but-correct defense in depth and blocks nothing public. No allowlist to relax. Note only. -- **Evidence:** researcher-web.ts:96 comment 'SSRF guard (C1.8)'; PRIVATE_HOST_BLOCKLIST set (~line 104), isPrivateIpv4/isPrivateIpv6, isSsrfUrl (~line 150); URL chosen as `rawUrl && !isSsrfUrl(rawUrl) ? rawUrl : undefined`. sourceDomains used only via `site:${domain}` in withControls. - -## [LOW] api-contract — agents/researcher/researcher-docs.ts:27 — [C] researcher-docs has no domain allowlist / scheme block of its own - -- **Risk:** researcher-docs only declares toolNames ['read_docs'] and imposes no allowlist or scheme restriction itself; egress safety is delegated to the read_docs backend (Context7). Nothing here over-blocks legitimate documentation. -- **Fix:** Keep. No restriction surface to relax in this file. Note only for coverage completeness. -- **Evidence:** researcher-docs.ts full file: `toolNames: ['read_docs']`, no fetch/allowlist logic; instructions only govern how the single read_docs call is used. - -## [LOW] security — sdk/src/tools/terminal-command-policy.ts:729 — [C] Read-only network-mutation ban (curl -d/-X POST, wget -O) — KEEP (owned by terminal-policy shard) - -- **Risk:** In read-only mode the policy rejects `curl` with -X POST/PUT/PATCH/DELETE or -d/--data/-T/--upload-file and `wget -O/--output-document`, with message 'network mutation is not allowed in read-only mode'. This blocks state-changing/exfil network calls and file-writing downloads while read-only. It correctly does not block plain read GETs. Detailed treatment belongs to the terminal-policy shard. -- **Fix:** Keep. Consistent with read-only semantics; blocking write/mutation verbs and -O file writes is appropriate. Defer any nuance (e.g. allowing curl -o /dev/stdout style reads) to the terminal-policy shard. Note only. -- **Evidence:** terminal-command-policy.ts:729 regex `/^(?:git\s+clone|curl\b...(?:-X (POST|PUT|PATCH|DELETE)|-d|--data|-T|--upload-file)|wget\b...(-O|--output-document))/i` paired with :730 message 'network mutation is not allowed in read-only mode'. - -## Coverage receipt - -### Subsystems - -- agent-runtime-web-search -- researcher-agents -- terminal-command-policy-network - -### Features - -- web-search-tool -- url-fetch -- ssrf-guard -- researcher-web -- researcher-docs - -### Files - -- packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts -- packages/agent-runtime/src/tools/handlers/tool/web-search.ts -- agents/researcher/researcher-web.ts -- agents/researcher/researcher-docs.ts -- packages/agent-runtime/src/tools/handlers/tool/**tests**/web-search-security.test.ts -- sdk/src/tools/terminal-command-policy.ts diff --git a/.agents/sessions/audit-agent-body-2026-09/SPEC.md b/.agents/sessions/audit-agent-body-2026-09/SPEC.md new file mode 100644 index 0000000000..3b18675914 --- /dev/null +++ b/.agents/sessions/audit-agent-body-2026-09/SPEC.md @@ -0,0 +1,46 @@ +# SPEC — Agent-body feature-gap audit (Openbuff CLI) + +Status: scope pinned, audit in progress. This file is rewritten once shard findings are synthesized. + +## Goal + +Openbuff should be a complete, reliable **body for a coding agent**: the agent must be able to +perceive the workspace, mutate it, run and verify work, recover from failure, and prove completion — +for any coding task, on a local-only machine with no backend. + +This audit answers three questions: + +1. **Tool quality** — for every tool the agent can call, is the implementation complete and + trustworthy (correct, structured errors, recovery guidance, failure states, tested)? +2. **Tool gaps** — which capabilities a coding agent needs are missing entirely? +3. **Over-strict guardrails** — which security/permission rules block legitimate local work + without protecting anything meaningful in a local-only, single-user CLI? + +## Non-goals (explicitly out of scope for this audit) + +- Token-consumption / context-cost reduction as an objective. Context handling is only in scope + where a *failure* (dropped evidence, lost capability, silent truncation) breaks task completion. +- Backend, multi-tenant, hosted-inference, billing, or credit concerns — none exist in this product. +- Model/prompt quality tuning and eval-score chasing. +- Cosmetic TUI polish unrelated to agent capability or operator verification. + +## Snapshot + +- Structural snapshot: `fb71d9a3e8f5d047c4dc944c4ef1bd14421da14d6e67e1af6c6c2d6cc5c3bcce` +- Subsystems inventoried: 24 top-level; source-bearing: `cli`, `sdk`, `packages/*`, `common`, + `agents`, `scripts`, `evals`, `.agents`. + +## Audit domains (per shard) + +Each shard evaluates its files against: security, correctness, state mutation, error handling, +performance, dependency hygiene, test coverage, API/contract stability — and additionally, for this +audit: **capability completeness** (can an agent actually finish a task with this tool?) and +**over-strictness** (does a guard block legitimate local work?). + +## Acceptance criteria for the audit itself + +- Every source-bearing subsystem is either sharded or explicitly marked out-of-scope with a reason. +- Every agent-callable tool in `common/src/tools/list.ts` appears in at least one shard's coverage. +- Findings are persisted under `.agents/sessions/audit-agent-body-2026-09/findings/`. +- Coverage is machine-checked with `evaluate_audit_coverage` before synthesis. +- The deliverable is a prioritized, source-backed plan — not a prose essay. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/AUDIT-REPORT.md b/.agents/sessions/audit-agent-ecosystem-2026-07/AUDIT-REPORT.md deleted file mode 100644 index 17fd5f0026..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/AUDIT-REPORT.md +++ /dev/null @@ -1,169 +0,0 @@ -# Openbuff agent ecosystem audit - -> **Implementation status:** All Top 10 and residual findings have been -> implemented and validated. See -> [IMPLEMENTATION-REPORT.md](./IMPLEMENTATION-REPORT.md) for source and test -> evidence. The original findings below are retained as the audit baseline. - -## Executive summary - -The audit found a capable agent system with strong final-review parsing, foreground spawn tests, filesystem containment, and broad gate/eval coverage. Its largest risks sit at trust and lifecycle boundaries: programmatic agents can bypass declared tools, terminal/browser/web actions rely heavily on prompt discipline, background work is not fully owned or cleaned up, and auxiliary quality agents are spawned without their results being interpreted. - -## Top 10 - -### 1. [CRITICAL] Web direct-fetch is an SSRF boundary failure - -- **Confidence:** High; independently verified by discovery and runtime pairs. -- **Risk:** Agent-controlled URLs can target loopback, private/LAN, link-local, metadata, or redirected internal endpoints; response bodies are fully buffered before truncation. -- **Evidence:** `common/src/tools/params/tool/web-search.ts:17-23` validates only URL syntax; `packages/agent-runtime/src/tools/handlers/tool/web-search.ts:67-105` directly fetches and calls `response.text()`. -- **Fix:** HTTP(S)-only policy, DNS/IP checks before and after redirects, reserved-address blocking, streamed byte cap, combined timeout/run AbortSignal. - -### 2. [HIGH] Programmatic agents bypass declared tool permissions - -- **Confidence:** High; explicit source behavior and documentation contradiction. -- **Risk:** A nominally read-only `handleSteps` agent can yield terminal, mutation, spawn, or network tools absent from `toolNames`. -- **Evidence:** `packages/agent-runtime/src/run-programmatic-step.ts:700-717,750-759` comments out availability enforcement; `docs/agents-and-tools.md:11` says templates define tool permissions. -- **Fix:** Enforce effective tool capabilities for all yielded calls; add narrow internal capabilities for trusted orchestrator operations. - -### 3. [HIGH] Terminal and external-action approvals are prompt-only - -- **Confidence:** High; repeated across execution, discovery, quality, and CLI shards. -- **Risk:** Basher/debugger/git-committer/browser/librarian and shipped third-party CLI agents can perform high-impact actions without an enforceable approval receipt. Codex/Claude/Gemini templates explicitly disable sandbox/approval protections. -- **Evidence:** `agents/basher.ts:79-139`; `agents/debugger/debugger.ts:60-67`; `agents/git-committer/git-committer.ts:43-106`; `.agents/codex-cli.ts:76-124`; `.agents/claude-code-cli.ts:5-45`; `.agents/gemini-cli.ts:5-51`. -- **Fix:** Central action classifier and scoped approval tokens; least-privilege profiles (`read-only`, `workspace-write`, `full-access`); dedicated git/browser mutation tools; visible per-run permission disclosure. - -### 4. [HIGH] Auxiliary quality gates ignore outcomes and can report false success - -- **Confidence:** High; orchestrator and quality pairs agree. -- **Risk:** Test-writer, doc-writer, or security-reviewer may crash, no-op, time out, or report a critical issue while orchestration proceeds. Telemetry declares reviewer/validation `passed` before execution. -- **Evidence:** `agents/base2/base2.ts:695-805` marks flags and yields without capturing results; success telemetry is emitted at `:711-718`, `:747-754`, and `:782-789`; final reviewer parsing exists at `:1278-1398`. -- **Fix:** Shared structured aux result/verdict, lifecycle states, post-result done flags, persisted blockers/crashes, and visible gate cards. - -### 5. [HIGH] Background agents/processes lack complete lifecycle ownership - -- **Confidence:** High; runtime and execution pairs agree. -- **Risk:** Mixed batches can start agents whose IDs are never returned; completed agent jobs retain full internal state indefinitely; cancellation/end-turn omit agent jobs; shell kill/timeout may leave descendants alive. -- **Evidence:** `spawn-agents.ts:95-203,349-390`; `background-agent-jobs.ts:76,161-215`; `check-background-agent.ts:61-109`; `sdk/src/tools/run-terminal-command.ts:248-301`; `sdk/src/tools/background-jobs.ts:327-343,504-519`. -- **Fix:** Atomic prevalidation, normalized/redacted job results, TTL/LRU/consume APIs, unified end-turn accounting/cancel, process-group termination with confirmed exit. - -### 6. [HIGH] Step-cap exhaustion bypasses validation/review - -- **Confidence:** High; source and existing test expectation confirm behavior. -- **Risk:** Dirty, incomplete work is moved to `final_response_allowed` and follow-up affordances are enabled without required gates. -- **Evidence:** `agents/base2/base2.ts:557-569`; unsafe pending-file expectation in `agents/__tests__/base2.test.ts:1010-1064`; completion contract in `docs/agents-and-tools.md:36-37`. -- **Fix:** Distinct `step_cap_reached` interrupted state, preserve pending gate state, disable green completion, resume gates first next turn. - -### 7. [HIGH] Spawn identity is ambiguous for concurrent same-type agents - -- **Confidence:** High; picker and searcher traced event/schema/CLI mismatch. -- **Risk:** Concurrent identical agent types can exchange optimistic cards, prompts, params, IDs, nesting, and streamed output. -- **Evidence:** `cli/src/utils/spawn-agent-matcher.ts:11-24` matches first base type; optimistic blocks store call/index metadata, but `common/src/types/print-mode.ts:75-87` start events lack it. -- **Fix:** Add opaque spawn correlation or tool-call ID + index to start events; consume exact entry; type fallback only when unique. - -### 8. [HIGH] MCP secrets can leak through full-template logs - -- **Confidence:** High; runtime pair verified mutation and logging sites. -- **Risk:** Resolved environment credentials become plaintext inside logged agent templates. -- **Evidence:** `sdk/src/agents/load-agents.ts:73-79,243-255`; full `agentTemplate` logs at `run-agent-step.ts:456-474` and `spawn-agent-utils.ts:448-456`. -- **Fix:** Resolve secrets only at launch or redact centrally; prohibit full template/config logging. - -### 9. [HIGH] Runtime context pruning/status ignores resolved BYOK model capacity - -- **Confidence:** Medium-high; capability is resolved but no audited connection into runtime loop was found. -- **Risk:** 8k/32k models may show a 190k budget and prune too late, causing emergency trims and misleading UX. -- **Evidence:** `run-agent-step.ts:835-838,1203-1236`; `util/context-pruning.ts:23-76`; resolved capability exists at `sdk/src/impl/model-provider.ts:152-154`. -- **Fix:** Propagate resolved context window into runtime, pruning, budgets, and status using one capability source. - -### 10. [HIGH] Discovery/research side effects and provenance are under-specified - -- **Confidence:** High for action boundaries; medium-high for research UX. -- **Risk:** Browser agents share global session state and can mutate external sites; librarian shells into untrusted clones; researcher-web can silently omit questions and loses claim-to-source linkage. -- **Evidence:** discovery shards cite `agents/browser-use/browser-use.ts:152-158`, `agents/librarian/librarian.ts:127-186`, and `agents/researcher/researcher-web.ts:275-387`. -- **Fix:** Per-run browser isolation and action policy, sandboxed/read-only clone inspection with owned cleanup, structured research results with question status, claim citations, date/source/locale/depth controls. - -## Cross-cutting patterns - -1. **Prompt policy substitutes for enforcement.** Terminal, git, browser, debugger, external CLI, and programmatic-tool safety depend on model compliance. -2. **Lifecycle identity/ownership is fragmented.** Optimistic spawn cards, background agents, shell jobs, browser sessions, temp clones/logs, cancellation, cost, and cleanup use separate contracts. -3. **Quality gates are asymmetric.** Code-reviewer has a robust verdict/crash contract; security/test/doc/debug agents use heterogeneous `last_message` output and are difficult to orchestrate safely. -4. **Docs and runtime contracts drift.** Examples include “pre-edit” security review running post-edit, per-package test commands using only the first package, configurable `maxSpawnDepth` not loadable, and git secret scanning described but not enforced. -5. **UX hides phase and failure semantics.** Every spawn tool renders as “Review”; invalid local agents are log-only; background agents/cost are absent at end-turn; timeout/cancel often collapse into generic failures. - -## Priorities - -### P0 — trust boundaries and data exposure - -- Block SSRF/private-network fetches and cap streamed bodies. -- Enforce programmatic tool capabilities and centralized action approvals. -- Redact MCP/provider secrets and sanitize user-visible stack traces. -- Make security-review results blocking and machine-readable. -- Fix process-tree termination and background-agent ownership/cleanup. - -### P1 — correctness and workflow integrity - -- Preserve gates on step cap; implement bounded reviewer crash recovery/bypass. -- Add spawn correlation to start events and atomic batch prevalidation. -- Connect resolved BYOK model capacity to pruning/status. -- Replace broad automatic test/doc mutation with intent/contract-aware routing; route docs by package and tests by package-command map. -- Make editor completion structured and consumed; invoke debugger after repeated identical/unparseable validation failures. - -### P2 — UX, performance, and contract polish - -- Render neutral phase-aware spawn cards with status counts, prompts, results, errors, and cancellation. -- Unify local-agent loading/provenance; project precedence, reload invalidation, source links, and actionable diagnostics. -- Add structured researcher controls/citations and per-question budget outcomes. -- Bound file-lister subtree/context search output; own librarian clone/temp-log cleanup. -- Compact reviewer context and add bounded navigation; type eval timeout/cancel outcomes and bounded parallel check groups. - -## Residual findings by area - -### Orchestrator/quality - -- Reviewer crash text advertises retry/fallback/bypass not represented in state (`base2.ts:1359-1396`). -- Doc-writer is hardwired to `docs/agents-and-tools.md`; most internal source edits trigger test/doc agents serially. -- Mixed-package test-writer uses the first package command only. -- Absolute gate paths outside cwd are not rejected; context-pruner is both internal-only and publicly spawnable. -- Git branch creation runs before dirty-tree inspection; `stage_all` and secret prevention are unsafe/prose-only. - -### Execution/runtime - -- Basher `what_to_summarize` labels rather than summarizes; its schema drifts from terminal params. -- Timeout/spawn failures reject outside the deterministic command-result schema; cancellation leaves background commands running. -- Basher full logs and librarian clones lack explicit retention/cleanup ownership. -- Invalid agent configuration can emit `prompt-error` and then continue. -- Home agents silently override project agents; `maxSpawnDepth` is advertised but not dynamically configurable. -- Code-search context flags can bypass memory/output guards; web-search timeout does not cancel underlying work and uses a fragile deep import. - -### CLI/evals - -- Registry refresh does not invalidate derived mode listings; the CLI source-path scanner supports fewer extensions than the SDK. -- Expanded dependency rows are outside keyboard focus/scroll geometry. -- External CLI permission profile is absent from structured results/cards. -- Eval final checks are serial without dependency metadata; timeout/cancel becomes exit code 1. -- Best-of-N editor E2E is excluded/stale and can pass vacuously. - -### Discovery/research - -- File-list graph retrieval is not directory-scoped and subtree requests can be 50× normal budget. -- File-picker accepts prose-like paths and can issue empty reads. -- Researcher-docs lacks structured source/version/failure output. -- Browser visual smoke captures screenshot/PDF/recording by default, increasing cost and artifact noise. - -## Evidence, inference, and limits - -- **Evidence:** Findings are based on paired file inventories and cross-file source searches with file/line citations. Repeated findings were de-duplicated; rejected/narrowed candidates in shard reports were not promoted. -- **Inference:** Exploitability and operational frequency were not measured live. BYOK context propagation is rated medium-high because no connection was found in the audited tree, not because every provider necessarily fails. -- **Limits:** No destructive browser/provider scenarios, dependency vulnerability scan, or production deployment tests were run. Generated artifacts, graveyard code, web UI, CI/release automation, and maintenance scripts were excluded except where needed to validate a contract. - -## Coverage - -Six complete picker/searcher shard pairs covered: - -- Orchestrator core, plan/execute modes, validation/reviewer gates. -- Editor, basher, terminal/background execution. -- File discovery, code search, web/docs research, browser, librarian. -- Reviewer, security, debugger, test/doc writer, git committer. -- Agent runtime, SDK routing/contracts, permissions, cancellation, background work. -- CLI nested-agent UX, local-agent registry/templates, and eval feedback. - -All eight audit domains were evaluated: security, correctness, state mutation, error handling, performance, dependency hygiene, test coverage, and API/ABI contracts. Full subsystem dispositions and exclusions are recorded in `COVERAGE-MATRIX.md`. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/COVERAGE-MATRIX.md b/.agents/sessions/audit-agent-ecosystem-2026-07/COVERAGE-MATRIX.md deleted file mode 100644 index 7f2f7686a4..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/COVERAGE-MATRIX.md +++ /dev/null @@ -1,49 +0,0 @@ -# Coverage matrix - -Audit focus: shipped orchestrators and subagents, their runtime/SDK contracts, CLI agent UX, local agent templates, and evaluation feedback. - -| Domain | Shard IDs | Covered | -| --------------------------------------------------------------------------------- | ---------------------------------------- | ------- | -| Orchestrator core, plan/execute modes, validation/reviewer gates | orchestrator-picker, orchestrator-search | yes | -| Editor, basher, terminal/background execution | execution-picker, execution-search | yes | -| File discovery, code search, web/docs research, browser, librarian | discovery-picker, discovery-search | yes | -| Reviewer, security, debugger, test/doc writer, git committer | quality-picker, quality-search | yes | -| Agent runtime, SDK contracts, routing, permissions, cancellation, background work | runtime-picker, runtime-search | yes | -| CLI nested-agent UX, local-agent registry/templates, eval feedback | cli-ux-picker, cli-ux-search | yes | - -## Subsystem enumeration - -| Top-level directory | Disposition | -| -------------------- | ----------------------------------------------------------------------------------------------- | -| `.agents` | audited — local CLI agents, shared agent factory/schema, permission defaults | -| `.bin` | out-of-scope — runtime binary shim, not agent behavior | -| `.codebuff-index` | out-of-scope — generated index data; query/index contracts were audited in source | -| `.codex` | out-of-scope — external assistant configuration, not shipped Openbuff agents | -| `.e2e-scratch` | out-of-scope — test scratch fixtures | -| `.git` | out-of-scope — repository metadata | -| `.github` | out-of-scope — CI/release automation, not agent runtime/UX | -| `.omx` | out-of-scope — external orchestration artifacts, not shipped Openbuff runtime | -| `.turbo` | out-of-scope — build cache | -| `.vscode` | out-of-scope — editor settings | -| `agents` | audited — all active shipped agent families and agent tests/wiring | -| `agents-graveyard` | out-of-scope — deprecated agents; referenced only to verify stale best-of-N coverage | -| `assets` | out-of-scope — static assets unrelated to agent behavior | -| `cli` | audited — agent selection, nested rendering, spawn reconciliation, statuses, plan/eval UX | -| `common` | audited — agent schemas, spawn/tool contracts, validation and permissions metadata | -| `debug` | out-of-scope — generated logs; logging call sites were audited in runtime source | -| `docs` | audited — agents/tools and request-flow behavioral contracts | -| `e2e-traces` | out-of-scope — generated trace artifacts | -| `evals` | audited — agent runner and breadth/sharding feedback contracts | -| `node_modules` | out-of-scope — vendored dependencies; no dependency vulnerability scan requested | -| `openbuff.d.example` | audited — example agent routing/configuration surface | -| `packages` | audited — agent-runtime, index/search and relevant tool handlers/contracts | -| `scripts` | out-of-scope — maintenance/build scripts, except structural-map builder used for scoping | -| `sdk` | audited — local agent loading, tool execution, browser/terminal/background lifecycle | -| `test` | out-of-scope — global test bootstrap; relevant package/agent tests were audited | -| `web` | out-of-scope — not part of the active local CLI agent surface described by current architecture | - -## Coverage notes - -- Six complete shard pairs were used: one file-picker-style inventory and one code-searcher-style verification per domain. -- Each pair evaluated security, correctness, state mutation, error handling, performance, dependency hygiene, test coverage, and API/ABI contracts, with additional focus on feature gaps and user experience. -- Findings are source-audit results. No live provider/browser destructive scenarios were executed. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/IMPLEMENTATION-REPORT.md b/.agents/sessions/audit-agent-ecosystem-2026-07/IMPLEMENTATION-REPORT.md deleted file mode 100644 index 258952ae75..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/IMPLEMENTATION-REPORT.md +++ /dev/null @@ -1,103 +0,0 @@ -# Agent ecosystem audit implementation report - -Status: **implemented and validated** -Baseline implementation: `3beab44c5 feat: harden read/write orchestration` -Follow-up state: current working tree (post-audit reconciliation) - -This report closes the findings in [AUDIT-REPORT.md](./AUDIT-REPORT.md). The -baseline commit contains the large read/write/index/context and orchestration -architecture changes; the current working tree contains the final residual -agent, gate, browser, approval, indexing, background, CLI, and evaluator fixes. - -## Top 10 closure - -| # | Finding | Status | Implementation evidence | -| --- | ------------------------------------------------ | ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| 1 | Web direct-fetch SSRF boundary | Closed | `web-search-utils.ts` now enforces HTTP(S), blocks reserved/private DNS and IP targets before and after redirects, combines cancellation/timeout, and streams through a byte cap. Deep `open-websearch` was replaced with bounded DuckDuckGo HTML retrieval. Covered by `web-search-security.test.ts`. | -| 2 | Programmatic agents bypass declared tools | Closed | `run-programmatic-step.ts` enforces the union of visible and explicitly declared hidden programmatic capabilities. Undeclared yielded tools terminate with a structured error and never execute. Covered by `run-programmatic-step.test.ts` and agent reachability tests. | -| 3 | Terminal/external approvals are prompt-only | Closed | SDK terminal execution now classifies high-impact actions, evaluates static policy first, and atomically consumes exact repository/workspace/root-run/snapshot/action/target approval receipts. Default-branch push remains unconditionally denied. External CLI results disclose `permissionProfile: "tmux-test"`; browser mutation requires an explicit interaction policy. Covered by harness-enforcement, terminal-policy, run-terminal-command, browser-policy, and CLI rendering tests. | -| 4 | Auxiliary quality gates ignore outcomes | Closed | Base2 captures structured test/doc/security receipts, records lifecycle state only after completion, carries blockers forward, groups test routing by workspace, and validates before the final reviewer. Covered by gate aux unit tests and both gate lifecycle E2E suites. | -| 5 | Background work lacks lifecycle ownership | Closed | Agent and shell jobs now carry client/root-run/parent-run/agent ownership, bounded output/cursors/consumer counts, persistence/recovery, cancellation, TTL/consumption behavior, and end-turn accounting. Default subagent lifetime is bounded to 30 minutes unless explicitly overridden. Covered by runtime background, end-turn, check-job, SDK background, timeout, and recovery tests. | -| 6 | Step-cap exhaustion bypasses gates | Closed | Step exhaustion produces a resumable `STEP_CAP_REACHED` checkpoint and preserves pending validation/review state instead of manufacturing successful finalization. Covered by `main-prompt.test.ts` and Base2 state tests. | -| 7 | Same-type concurrent spawn identity is ambiguous | Closed | Spawn events carry tool-call/index correlation; CLI matching consumes exact correlated entries and only uses type fallback when unique. Covered by `spawn-agent-matcher.test.ts`, nested streaming tests, and SDK event tests. | -| 8 | MCP secrets can leak through template logs | Closed | MCP cache identity hashes resolved headers/env, runtime logging avoids full resolved templates, and secret-bearing provider/template structures are centrally redacted. Covered by MCP client cache/redaction tests and runtime logging contract tests. | -| 9 | Pruning ignores resolved model capacity | Closed | Resolved per-model `contextWindowTokens` flows into each run, semantic trigger/target budgets, provider request trimming, telemetry, and CLI status. Child agents start with their own unresolved window and resolve independently. Typed task memory, pinned blockers/open questions, structured handoffs, and repeated-compaction preservation prevent amnesia. Covered from 8k through 1M windows by context-pruner, loop-agent, task-memory, spawn-history, and LLM context-window suites. | -| 10 | Discovery/research side effects and provenance | Closed | Browser state is owner-keyed and interaction defaults read-only; Librarian clones are runtime-owned and cleaned from trusted history unless explicitly retained; web/docs researchers return structured question/source/version/failure evidence; query indexing supports real path-prefix filtering; file discovery is scope-safe and bounded. Covered by browser ownership/policy, Librarian cleanup, query schema/indexer, researcher, file-picker, and specialist contract tests. | - -## Residual finding closure - -### Orchestrator and quality - -| Finding | Resolution | -| ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Reviewer crash text did not match state | Reviewer protocol failures now have an explicit bounded reviewer-only retry, a stopped/blocked state after the second failure, and an explicit challenge-bound bypass. Protocol failures are never sent to repair-editor. | -| Doc writer was hardwired and aux routing was over-broad | Aux routing uses affected source/workspace evidence and project-aware doc/test targets; generated aux files do not reopen the same snapshot indefinitely. | -| Mixed-package test writer used one command | Targets are grouped by nearest workspace and test command; build-only commands are rejected as test evidence. | -| Absolute gate paths and public context-pruner | Gate paths are normalized/contained. Context-pruner remains runtime-declared but is removed from model-visible tools, prompts, and stream metadata. | -| Git branch/staging/secret policy | Dirty-state inspection precedes branch mutation; staging is path-scoped; high-impact git operations are centrally classified and approval-gated; direct default-branch push is prohibited. | - -### Execution and runtime - -| Finding | Resolution | -| ---------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Basher labeled rather than summarized | Basher now performs deterministic semantic line extraction while preserving command/CWD/exit/job/log metadata and a structured result. | -| Timeout/spawn errors escaped deterministic results | Subagent timeout/cancellation is bounded and abort-aware; shell timeout/cancel paths return structured results and terminate owned work. | -| Basher logs and Librarian clones lacked ownership | Full-log retention is explicit; non-retained logs are deleted. Librarian cleanup derives its target only from trusted programmatic history and a validated URL-derived owned path. | -| Invalid agent configuration continued after `prompt-error` | Invalid configuration is terminal for the prompt; execution no longer emits a second start/finish sequence. | -| Home/project precedence and `maxSpawnDepth` drift | Executable project agents require explicit trust, project precedence is deterministic, and dynamic/template configuration carries `maxSpawnDepth`. | -| Code-search guards and web cancellation | Search outputs/flags are bounded; web retrieval is direct, abortable, timeout-aware, redirect-revalidated, and byte-capped. | - -### CLI and evaluations - -| Finding | Resolution | -| ---------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Registry refresh/listing and source-extension drift | Derived mode listings invalidate with registry refresh and CLI source discovery matches the SDK-supported agent extensions. | -| Expanded dependency rows were not keyboard reachable | Expanded dependency rows are flattened into focus/scroll order and covered by component tests. | -| External CLI permission profile was hidden | Structured output now includes `outputKind: "external-cli"` and `permissionProfile: "tmux-test"`; CLI rendering displays it. | -| Eval checks were serial and collapsed timeout/cancel | Object-form checks support stable IDs, dependencies, bounded parallelism, per-check timeout, skipped/configuration-error states, and duration evidence; legacy strings retain sequential behavior. | -| Best-of-N E2E was stale/vacuous | The obsolete E2E file is deleted, removed IDs are guarded by reachability tests, and the stale TypeScript exclusion has been removed. | - -### Discovery and research - -| Finding | Resolution | -| ------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| File-list retrieval ignored directory scope and oversized subtree budgets | `query_index.pathPrefixes` filters lexical, semantic, and related-file candidates before ranking; file-lister/picker results are scope-validated and capped to eight high-relevance files. | -| File-picker accepted prose/empty reads | Candidate parsing rejects prose-like, unsafe, non-file, and out-of-scope paths; empty safe results stop without issuing a read. | -| Researcher-docs lacked source/version/failure output | It now returns structured `status`, `answer`, `source`, `version`, and optional `failure`, using Context7 metadata without guessing. | -| Browser smoke always generated screenshot/PDF/recording | Browser evidence is proportional: screenshot for visual checks, PDF only for print/PDF behavior, and recording only for time-based or explicitly requested evidence. | - -## Architectural outcomes - -- Workspace mutation is serialized through the broker with revision/hash CAS, - authoritative commit receipts, journaling, rollback evidence, and - reconciliation. -- Orchestration state is replayable through typed ledger, plan, workflow, - discovery, lease, handoff, and receipt contracts. -- Agent selection accounts for capabilities, real writable path scope, leases, - dependencies, model context minima, and current workspace revision. -- Context is isolated per agent/model and shared through compact typed handoffs, - evidence, receipts, and task memory rather than copied parent histories. -- Gate snapshot attestation uses a single-line opaque hash. File details remain a - separate review input. Attestation failures retry the reviewer once and then - stop; they do not create source repair loops. - -## Validation evidence - -All validations below passed on the reconciled working tree: - -- TypeScript: agents, common, agent-runtime, indexer, SDK, CLI, and evals. -- Common: **723 passed**. -- Agent runtime: **1060 passed**. -- Indexer: **180 passed**. -- SDK: **923 passed, 1 skipped** (the skip is a pre-existing TODO integration case). -- Shipped agent unit suite: **591 passed**, plus **2 passed** specialist audit-contract tests. -- CLI: **2434 passed, 16 skipped** (environment-dependent tmux/clipboard/performance cases). -- Evals: **224 passed**. -- Focused residual cross-package suite: **303 passed**. -- Gate lifecycle/aux regression suite: **49 passed**. -- `git diff --check`: clean. - -There are no remaining source-level blockers from this audit. Reviewer -attestation failures are now protocol state, not repair findings; a fresh -matching reviewer receipt clears code findings, while repeated protocol failure -halts after the bounded retry instead of looping. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/MAP.md b/.agents/sessions/audit-agent-ecosystem-2026-07/MAP.md deleted file mode 100644 index f6c9f511aa..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/MAP.md +++ /dev/null @@ -1,350 +0,0 @@ -# Structural Map — openbuff - -- **Project root:** `/home/ben/Code/CLI/openbuff` -- **Built at:** 2026-07-11T07:21:35.959Z -- **Total files indexed:** 2138 -- **Graph:** 15036 nodes, 72053 edges - -> Pin this file in context. Every audit shard navigates from here instead of doing fuzzy round-trip discovery. - -## Entry points - -- `cli/src/index.tsx` -- `packages/code-map/src/index.ts` -- `packages/indexer/src/index.ts` -- `packages/internal/src/index.ts` - -## Directories (by size, biggest first) - -| dir | files | total size | top symbols | -| -------------------------- | ----- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `evals` | 597 | 12.4 MB | main, run, makeEvalRun, makeAgentResults, toolCall, log | -| `cli` | 426 | 3.0 MB | render, main, TestItem, tmux, createErrorMessage, formatTimestamp | -| `packages` | 312 | 2.3 MB | Greeting, start, greet, Greeter, flush, doGenerate | -| `sdk` | 161 | 1.4 MB | main, run, evaluate, createMockFs, log, resolveMcpEnv | -| `common` | 245 | 1.1 MB | createMockLogger, getStringProperty, getFileExtension, process, sleep, size | -| `agents` | 93 | 1.1 MB | extractInlineFunctionSource, parseGateStateBlock, feedJson, collectToolInputFiles, isFileChangingTool, hasEditArtifact | -| `scripts` | 67 | 515.7 KB | main, parseArgs, computeCost, ConversationMessage, TurnResult, makeConversationStreamRequest | -| `.omx` | 30 | 502.5 KB | — | -| `agents-graveyard` | 121 | 350.6 KB | createBase2WithTaskResearcher, getLatestEditToolResults, extractSpawnResults, getSpawnResults, createResearchImplementOrchestrator, createBase2Implementor | -| `bun.lock` | 1 | 270.1 KB | — | -| `docs` | 13 | 203.1 KB | — | -| `.agents` | 26 | 178.4 KB | publisher, getSpawnerPrompt, getSystemPrompt, getInstructionsPrompt, getDefaultReviewModeInstructions, getWorkModeInstructions | -| `.github` | 14 | 51.8 KB | — | -| `openbuff.d.example` | 4 | 22.5 KB | — | -| `LICENSE` | 1 | 11.1 KB | — | -| `.bin` | 1 | 8.5 KB | — | -| `README.zh-CN.md` | 1 | 8.1 KB | — | -| `README.md` | 1 | 8.0 KB | — | -| `WINDOWS.md` | 1 | 7.3 KB | — | -| `CONTRIBUTING.md` | 1 | 5.6 KB | — | -| `CODE_OF_CONDUCT.md` | 1 | 4.5 KB | — | -| `eslint.config.js` | 1 | 4.0 KB | — | -| `AGENTS.md` | 1 | 3.6 KB | — | -| `package.json` | 1 | 2.5 KB | — | -| `INFISICAL_SETUP_GUIDE.md` | 1 | 2.5 KB | — | -| `ROUTER.md` | 1 | 2.5 KB | — | -| `.env.example` | 1 | 1.7 KB | — | -| `tsconfig.json` | 1 | 839 B | — | -| `SECURITY.md` | 1 | 520 B | — | -| `.gitignore` | 1 | 487 B | — | -| `.vscode` | 1 | 438 B | — | -| `bunfig.toml` | 1 | 432 B | — | -| `.prettierrc` | 1 | 389 B | — | -| `tsconfig.base.json` | 1 | 386 B | — | -| `test` | 1 | 332 B | setup | -| `knowledge.md` | 1 | 287 B | — | -| `.e2e-scratch` | 2 | 279 B | add, greet, multiply | -| `NOTICE` | 1 | 156 B | — | -| `openbuff.json.example` | 1 | 118 B | — | -| `.envrc` | 1 | 30 B | — | -| `.bun-version` | 1 | 7 B | — | - -## Largest files per directory - -### `evals` - -- `evals/buffbench/logs/2026-07-04T17-30_base2/45-fork-read-files-base2-349a140.json` — 472.2 KB, 0 symbols -- `evals/buffbench/logs/2026-07-04T13-41_base2/2-add-deep-thinkers-base2-6c362c3.json` — 469.4 KB, 0 symbols -- `evals/buffbench/restrict-tool-types-base2-lite-error-ftj2.json` — 418.5 KB, 0 symbols -- `evals/buffbench/validate-custom-tools-base2-error-c6yk.json` — 418.3 KB, 0 symbols -- `evals/buffbench/logs/2026-07-04T17-30_base2/38-unify-agent-builder-base2-4852954.json` — 411.4 KB, 0 symbols - -### `cli` - -- `cli/bin/tree-sitter.wasm` — 200.7 KB, 0 symbols -- `cli/src/hooks/helpers/__tests__/send-message.test.ts` — 55.5 KB, 0 symbols -- `cli/src/utils/__tests__/message-block-helpers.test.ts` — 55.1 KB, 0 symbols -- `cli/src/chat.tsx` — 53.8 KB, 1 symbols -- `cli/src/utils/__tests__/send-message-helpers.test.ts` — 48.1 KB, 0 symbols - -### `packages` - -- `packages/agent-runtime/src/__tests__/read-files-edit-state.test.ts` — 126.5 KB, 2 symbols -- `packages/agent-runtime/src/__tests__/process-str-replace.test.ts` — 76.8 KB, 0 symbols -- `packages/agent-runtime/src/process-str-replace.ts` — 66.5 KB, 30 symbols -- `packages/agent-runtime/src/__tests__/run-programmatic-step.test.ts` — 65.2 KB, 0 symbols -- `packages/agent-runtime/src/run-agent-step.ts` — 53.3 KB, 4 symbols - -### `sdk` - -- `sdk/src/__tests__/model-provider.test.ts` — 80.0 KB, 3 symbols -- `sdk/src/provider-config.ts` — 72.5 KB, 30 symbols -- `sdk/src/tools/browser-logs.ts` — 62.8 KB, 30 symbols -- `sdk/src/impl/llm.ts` — 52.9 KB, 20 symbols -- `sdk/src/__tests__/run-cancellation.test.ts` — 44.5 KB, 1 symbols - -### `common` - -- `common/src/templates/initial-agents-dir/types/tools.ts` — 44.6 KB, 30 symbols -- `common/src/util/__tests__/messages.test.ts` — 41.4 KB, 0 symbols -- `common/src/tools/results/filesystem.ts` — 33.8 KB, 30 symbols -- `common/src/__tests__/agent-validation.test.ts` — 28.9 KB, 0 symbols -- `common/src/util/__tests__/saxy.test.ts` — 26.1 KB, 0 symbols - -### `agents` - -- `agents/base2/base2.ts` — 169.4 KB, 30 symbols -- `agents/__tests__/context-pruner.test.ts` — 120.8 KB, 2 symbols -- `agents/__tests__/base2.test.ts` — 111.3 KB, 4 symbols -- `agents/context-pruner.ts` — 74.0 KB, 30 symbols -- `agents/types/tools.ts` — 44.6 KB, 30 symbols - -### `scripts` - -- `scripts/test-fireworks-cache-intervals.ts` — 34.3 KB, 10 symbols -- `scripts/benchmark-providers.ts` — 33.1 KB, 16 symbols -- `scripts/test-fireworks-long.ts` — 30.7 KB, 6 symbols -- `scripts/test-canopywave-long.ts` — 29.0 KB, 5 symbols -- `scripts/test-siliconflow.ts` — 27.3 KB, 5 symbols - -### `.omx` - -- `.omx/ultragoal/worktree-baseline.json` — 247.5 KB, 0 symbols -- `.omx/ultragoal/brief.md` — 36.3 KB, 0 symbols -- `.omx/plans/prd-read-write-audit-remediation.md` — 36.3 KB, 0 symbols -- `.omx/state/todos-session.json` — 33.9 KB, 0 symbols -- `.omx/plans/traceability-read-write-audit-remediation.md` — 27.8 KB, 0 symbols - -### `agents-graveyard` - -- `agents-graveyard/base/base-prompts.ts` — 25.3 KB, 3 symbols -- `agents-graveyard/editor/best-of-n/editor-best-of-n.ts` — 18.5 KB, 4 symbols -- `agents-graveyard/base/ask.ts` — 12.2 KB, 0 symbols -- `agents-graveyard/registry/transform-agent.ts` — 12.0 KB, 0 symbols -- `agents-graveyard/base2/task-researcher/base2-with-task-researcher-planner-pro.ts` — 11.4 KB, 1 symbols - -### `bun.lock` - -- `bun.lock` — 270.1 KB, 0 symbols - -### `docs` - -- `docs/agents-and-tools.md` — 74.8 KB, 0 symbols -- `docs/codebuff-to-openbuff-migration.md` — 30.7 KB, 0 symbols -- `docs/configuration.md` — 20.6 KB, 0 symbols -- `docs/openbuff-provider-model-setup-ux.md` — 20.4 KB, 0 symbols -- `docs/architecture.md` — 12.3 KB, 0 symbols - -### `.agents` - -- `.agents/types/tools.ts` — 44.0 KB, 30 symbols -- `.agents/sessions/read-write-tooling-2026-07-10/findings/structured-results-plan.md` — 35.6 KB, 0 symbols -- `.agents/types/agent-definition.ts` — 13.7 KB, 3 symbols -- `.agents/lib/cli-agent-prompts.ts` — 13.7 KB, 5 symbols -- `.agents/sessions/read-write-tooling-2026-07-10/MAP.md` — 11.0 KB, 0 symbols - -### `.github` - -- `.github/workflows/cli-release-build.yml` — 11.7 KB, 0 symbols -- `.github/workflows/cli-release-staging.yml` — 8.8 KB, 0 symbols -- `.github/workflows/ci.yml` — 7.4 KB, 0 symbols -- `.github/knowledge.md` — 5.4 KB, 0 symbols -- `.github/workflows/cli-release-prod.yml` — 4.9 KB, 0 symbols - -### `openbuff.d.example` - -- `openbuff.d.example/providers.json` — 19.0 KB, 0 symbols -- `openbuff.d.example/routes.json` — 2.1 KB, 0 symbols -- `openbuff.d.example/hooks.json` — 1.2 KB, 0 symbols -- `openbuff.d.example/indexing.json` — 146 B, 0 symbols - -### `LICENSE` - -- `LICENSE` — 11.1 KB, 0 symbols - -### `.bin` - -- `.bin/bun` — 8.5 KB, 0 symbols - -### `README.zh-CN.md` - -- `README.zh-CN.md` — 8.1 KB, 0 symbols - -### `README.md` - -- `README.md` — 8.0 KB, 0 symbols - -### `WINDOWS.md` - -- `WINDOWS.md` — 7.3 KB, 0 symbols - -### `CONTRIBUTING.md` - -- `CONTRIBUTING.md` — 5.6 KB, 0 symbols - -### `CODE_OF_CONDUCT.md` - -- `CODE_OF_CONDUCT.md` — 4.5 KB, 0 symbols - -### `eslint.config.js` - -- `eslint.config.js` — 4.0 KB, 0 symbols - -### `AGENTS.md` - -- `AGENTS.md` — 3.6 KB, 0 symbols - -### `package.json` - -- `package.json` — 2.5 KB, 0 symbols - -### `INFISICAL_SETUP_GUIDE.md` - -- `INFISICAL_SETUP_GUIDE.md` — 2.5 KB, 0 symbols - -### `ROUTER.md` - -- `ROUTER.md` — 2.5 KB, 0 symbols - -### `.env.example` - -- `.env.example` — 1.7 KB, 0 symbols - -### `tsconfig.json` - -- `tsconfig.json` — 839 B, 0 symbols - -### `SECURITY.md` - -- `SECURITY.md` — 520 B, 0 symbols - -### `.gitignore` - -- `.gitignore` — 487 B, 0 symbols - -### `.vscode` - -- `.vscode/settings.json` — 438 B, 0 symbols - -### `bunfig.toml` - -- `bunfig.toml` — 432 B, 0 symbols - -### `.prettierrc` - -- `.prettierrc` — 389 B, 0 symbols - -### `tsconfig.base.json` - -- `tsconfig.base.json` — 386 B, 0 symbols - -### `test` - -- `test/setup-scm-loader.ts` — 332 B, 1 symbols - -### `knowledge.md` - -- `knowledge.md` — 287 B, 0 symbols - -### `.e2e-scratch` - -- `.e2e-scratch/widget.ts` — 275 B, 3 symbols -- `.e2e-scratch/browser-agent-note.txt` — 4 B, 0 symbols - -### `NOTICE` - -- `NOTICE` — 156 B, 0 symbols - -### `openbuff.json.example` - -- `openbuff.json.example` — 118 B, 0 symbols - -### `.envrc` - -- `.envrc` — 30 B, 0 symbols - -### `.bun-version` - -- `.bun-version` — 7 B, 0 symbols - -## Most-imported files (likely key modules) - -| in-degree | file | -| --------- | -------------------------------------------------------------- | -| 101 | `packages/agent-runtime/src/__tests__/rewrite-symbol.test.ts` | -| 82 | `common/src/util/messages.ts` | -| 75 | `cli/src/utils/arrays.ts` | -| 75 | `common/src/types/bun-test.d.ts` | -| 66 | `cli/src/__tests__/release/proxy-http-get.test.ts` | -| 62 | `common/src/util/error.ts` | -| 59 | `common/src/tools/params/utils.ts` | -| 54 | `sdk/e2e/utils/event-collector.ts` | -| 49 | `cli/src/utils/message-block-helpers.ts` | -| 44 | `cli/src/hooks/use-theme.tsx` | -| 43 | `packages/agent-runtime/src/__tests__/main-prompt.test.ts` | -| 40 | `sdk/src/provider-config.ts` | -| 40 | `common/src/util/plan-artifacts.ts` | -| 39 | `sdk/src/tools/filesystem-authority.ts` | -| 38 | `packages/agent-runtime/src/tools/handlers/tool/write-file.ts` | -| 36 | `cli/src/project-files.ts` | -| 34 | `common/src/testing/mocks/timers.ts` | -| 33 | `agents/base2/base2.ts` | -| 32 | `scripts/test-canopywave-long.ts` | -| 31 | `.e2e-scratch/widget.ts` | -| 31 | `common/src/util/content-hash.ts` | -| 30 | `common/src/types/session-state.ts` | -| 29 | `sdk/e2e/utils/get-api-key.ts` | -| 28 | `common/src/util/string.ts` | -| 28 | `sdk/src/run.ts` | - -## Cross-directory dependencies (architectural layering) - -| count | from → to | -| ----- | ----------------------------- | -| 251 | `packages` → `common` | -| 168 | `cli` → `common` | -| 132 | `sdk` → `common` | -| 125 | `cli` → `packages` | -| 62 | `cli` → `sdk` | -| 51 | `sdk` → `packages` | -| 25 | `cli` → `.e2e-scratch` | -| 23 | `cli` → `scripts` | -| 22 | `packages` → `cli` | -| 22 | `agents` → `sdk` | -| 21 | `sdk` → `cli` | -| 19 | `agents` → `agents-graveyard` | -| 19 | `agents-graveyard` → `common` | -| 18 | `packages` → `sdk` | -| 16 | `sdk` → `scripts` | -| 13 | `evals` → `common` | -| 12 | `scripts` → `packages` | -| 11 | `agents` → `common` | -| 9 | `evals` → `cli` | -| 8 | `common` → `packages` | -| 8 | `agents` → `cli` | -| 7 | `common` → `cli` | -| 7 | `evals` → `sdk` | -| 6 | `cli` → `agents` | -| 6 | `packages` → `agents` | -| 6 | `packages` → `.e2e-scratch` | -| 5 | `agents` → `packages` | -| 5 | `scripts` → `cli` | -| 5 | `evals` → `packages` | -| 5 | `common` → `sdk` | - -## Shard sizing hint - -Total indexed source: **23.4 MB** across **41** top-level directories. - -When sharding for an audit, aim for ~5–15 files per shard. Use the table above to group small dirs together and split huge dirs (e.g. split `src/` by subdirectory). diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/cli-ux-picker.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/cli-ux-picker.md deleted file mode 100644 index 0096a789aa..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/cli-ux-picker.md +++ /dev/null @@ -1,69 +0,0 @@ -# CLI/local-agent/eval UX audit — file-picker shard - -## Compact inventory for paired code-searcher - -- Selection/configuration: `agent-mode-toggle.tsx` switches top-level mode; `agent-checklist.tsx` searches/selects local agents and expands their spawnable dependency tree; `local-agent-registry.ts` merges bundled, project/parent/home agents and MCP config. -- Run display: `message-with-agents.tsx` renders message-tree children in responsive grids; `agent-branch-wrapper.tsx` / `agent-branch-item.tsx` render nested block-agent prompt, status, preview, body and selector summaries; `spawn-agents.tsx` renders the initiating tool call. -- Lifecycle: `message-block-helpers.ts` creates/moves nested agent blocks, attaches tool results, interruption notices, status/prompt metadata; `spawn-agent-matcher.ts` reconciles optimistic temporary blocks with runtime IDs. -- Local CLI agents: `.agents/lib/create-cli-agent.ts` shares schemas/prompts; individual Claude/Codex/Gemini/Openbuff definitions start tmux sessions and return structured results/captures/issues. -- Evaluation: `evals/buffbench/agent-runner.ts` runs agents under a 60-minute aggregate timeout and then sequential final checks; `plan-sharding-signals.ts` evaluates audit classification, parallelism, pair count and textual coverage. -- Tests found: strong registry integration and sharding-signal suites; component layout/message tests; only pure helper tests for mode toggle. No direct tests found for spawn reconciliation with duplicate agent types or the `spawn_agents` renderer semantics. - -## [HIGH] Correctness / state mutation — cli/src/utils/spawn-agent-matcher.ts:11 — Concurrent same-type agents can be reconciled to the wrong optimistic card - -- **Risk:** Matching uses only the base agent type and returns the first map entry, even though the block model already carries `spawnToolCallId` and `spawnIndex`; two concurrent `file-picker` (or other same-type) children can therefore exchange prompts, params, nesting, runtime IDs, and subsequent output depending on event arrival order. -- **Fix:** Key optimistic entries and start events by spawn tool-call ID plus index (or a runtime correlation token); use type only as a guarded legacy fallback and consume matched entries atomically. -- **Evidence:** `findMatchingSpawnAgent` normalizes `eventAgentType`, loops insertion order, and returns on `eventBaseName === storedBaseName` (lines 15-21), while `createAgentBlock` supports `spawnToolCallId`/`spawnIndex` at `message-block-helpers.ts:648-690` but this matcher ignores them. - -## [HIGH] Correctness / state mutation — cli/src/utils/local-agent-registry.ts:55 — Registry reinitialization leaves mode listing caches stale - -- **Risk:** Startup reloads (and any future watch/reload UX) replace `userAgentsCache`, paths, and MCP configuration but do not invalidate `cachedAgentsByMode`; after the first `loadLocalAgents()` call, added/removed/renamed agents and overrides can remain invisible until process restart or a test-only reset. -- **Fix:** Clear all derived listing/directory caches at the start or successful end of `initializeAgentRegistry`, and expose a supported reload API that returns validation diagnostics for the UI. -- **Evidence:** initialization assigns caches at lines 55-87; `loadLocalAgents` returns cached arrays at lines 242-249 and populates them at 299; only `__resetLocalAgentRegistryForTests` clears `cachedAgentsByMode` at lines 451-460. - -## [MEDIUM] UX / API contract — cli/src/components/tools/spawn-agents.tsx:8 — Every agent spawn is presented as “Review” - -- **Risk:** General research, file-picking, implementation, validation and orchestration spawns are mislabeled as a review stage, so users cannot understand plan execution or distinguish why concurrent agents are running; the collapsed history becomes a sequence of identical “Review” entries. -- **Fix:** Render a neutral “Spawned N agents” header by default and derive stage labels from explicit structured metadata (phase/purpose), with per-agent status counts and expandable prompt summaries. -- **Evidence:** the component hardcodes `const header = 'Review'` (line 23), shows only comma-joined types (24-28), and ignores prompts, status, results, failures and cancellation despite accepting prompts in its input type (18-20). - -## [MEDIUM] UX / correctness — cli/src/components/agent-checklist.tsx:194 — Keyboard scrolling assumes every row is one line even when dependency trees expand - -- **Risk:** Expanding an agent inserts arbitrarily many descendant rows, but focus scrolling still computes `focusedTop = focusedIndex * 1`; keyboard focus can move offscreen or scroll to the wrong visual agent, making large local-agent graphs difficult to configure. -- **Fix:** Measure row renderables or build a flattened visible-row model (agent and dependency rows) and scroll by actual offsets; include keyboard affordances for expand/collapse and announce dependency cycles/missing definitions. -- **Evidence:** the effect explicitly assumes `itemHeight = 1` at lines 194-215, while expanded recursive `DepTree` rows are inserted between checklist items at lines 365-377. - -## [MEDIUM] Error handling / UX — cli/src/utils/local-agent-registry.ts:55 — Invalid local-agent configuration degrades to an empty menu with log-only feedback - -- **Risk:** One SDK load failure clears every user agent and the CLI gives no actionable in-product file/schema diagnostic; unreadable files/directories during ID-to-path discovery are silently skipped, so “Open file” can disappear and users cannot tell whether an agent is invalid, shadowed, or inaccessible. -- **Fix:** Preserve successfully loaded agents, return per-file diagnostics (source, precedence, validation error), surface them in the agent picker/startup notice, and add a reload/fix action. -- **Evidence:** the catch replaces the complete cache with `{}` and only `logger.warn`s (61-69); path scanning swallows file and directory errors with empty catches (127-139); `LocalAgentInfo.filePath` falls back to an empty string (152-157). - -## [MEDIUM] Security / feature gap — .agents/codex-cli.ts:81 — External CLI agents require blanket unsafe permission modes - -- **Risk:** Codex, Claude, and Gemini automation is designed around disabling approval/sandbox protections, so a mistaken or prompt-injected task can perform unrestricted repository/system actions; the UX provides no per-run risk disclosure or safer selectable profile. -- **Fix:** Add explicit permission profiles (`read-only`, `workspace-write`, `full-access`), default smoke/review runs to the least privilege compatible with the task, require an acknowledged opt-in for full access, and record the chosen profile in structured results. -- **Evidence:** Codex starts with `-a never -s danger-full-access` and instructs always using it (lines 81-83); equivalent blanket flags are mandated in `.agents/claude-code-cli.ts:12` and `.agents/gemini-cli.ts:12`. - -## [LOW] Performance / evaluation feedback — evals/buffbench/agent-runner.ts:174 — Final checks are serial and share only the agent’s aggregate timeout - -- **Risk:** Independent checks run one-by-one, reducing eval throughput, and a late check inherits the nearly exhausted 60-minute outer signal; results do not distinguish timeout/abort from an ordinary exit code 1, weakening feedback about agent quality versus harness exhaustion. -- **Fix:** Define per-check timeout/concurrency metadata, run independent checks in a bounded pool, and emit typed outcomes (`passed`, `failed`, `timed_out`, `cancelled`) with elapsed time. -- **Evidence:** `runFinalCheckCommands` iterates `for (const command of commands)` and awaits each `execAsync` (lines 227-263); catch coerces nonnumeric abort/timeout codes to `1` (252-260); all work shares the outer 60-minute signal (89-97, 174-196). - -## [LOW] Test coverage gap — cli/src/utils/spawn-agent-matcher.ts:11 — Critical optimistic-to-runtime reconciliation has no focused regression suite - -- **Risk:** Ordering, duplicate-type concurrency, nested parents, namespaced types, cancellation-before-start, and unmatched events can regress while layout tests remain green. -- **Fix:** Add table-driven tests covering two identical agent types from one batch and different batches, out-of-order starts, correlation metadata, nested moves, unmatched starts, and cancellation cleanup; add renderer tests asserting phase-neutral labels and lifecycle summaries. -- **Evidence:** scoped test discovery found registry, mode-toggle, message/layout, and helper suites, but no test references to `findMatchingSpawnAgent` or `SpawnAgentsComponent`. - -## 8-domain disposition - -- Security: unsafe external-CLI permission defaults found; no credential rendering issue proven in this shard. -- Correctness: duplicate-type reconciliation and stale registry cache found. -- State mutation: derived registry cache invalidation and optimistic block identity covered. -- Error handling: local-agent failures are log-only; eval abort classification is lossy. -- Performance: sequential final checks and expanded-list calculation/scroll model reviewed. -- Dependency hygiene: no package-version/declaration issue established in the scoped files. -- Test coverage: reconciliation/renderer lifecycle gaps found; registry and sharding evaluator coverage are comparatively strong. -- API/ABI contracts: hardcoded “Review” is a misleading UI semantic contract; structured CLI output schema is shared, but lacks permission profile and typed cancellation/timeout fields. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/cli-ux-search.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/cli-ux-search.md deleted file mode 100644 index 4f94e37e3f..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/cli-ux-search.md +++ /dev/null @@ -1,72 +0,0 @@ -# CLI/local-agent/eval UX — code-search findings - -## [HIGH] correctness / state mutation / API contract — duplicate same-type starts attach to the wrong optimistic agent card - -- **Risk:** Two agents of the same type spawned concurrently can exchange prompts, params, nesting, IDs, streamed output, and final results when `subagent_start` events arrive out of request order. Users may see one file-picker/editor apparently performing another child’s task. -- **Fix:** Extend `subagent_start` with a spawn correlation field (`spawnToolCallId` + `spawnIndex`, or an opaque spawn token), match on it, and consume the exact optimistic entry. Keep base-type matching only as a guarded legacy fallback when exactly one candidate exists. -- **Evidence:** `findMatchingSpawnAgent` returns the first insertion-order entry whose base type matches (`cli/src/utils/spawn-agent-matcher.ts:11-24`). Optimistic blocks already store `spawnToolCallId` and `spawnIndex` (`cli/src/utils/sdk-event-handlers.ts:290-325`), but `PrintModeSubagentStart` exposes neither (`common/src/types/print-mode.ts:75-87`), and `handleSubagentStart` passes only `event.agentType` to the matcher (`cli/src/utils/sdk-event-handlers.ts:193-228`). Final spawn results do use index/ID correlation (`cli/src/utils/sdk-event-handlers.ts:424-445`), so start-time reconciliation is the inconsistent leg. - -## [HIGH] security / UX — shipped external-CLI agents silently launch with unrestricted permission bypasses - -- **Risk:** Selecting these agents starts third-party coding CLIs with approval and sandbox protections disabled, allowing prompt-injected or mistaken instructions to mutate anything available to the process. The structured result does not record a permission profile, and the user gets no per-run risk confirmation. -- **Fix:** Add explicit `read-only`, `workspace-write`, and `full-access` input profiles; default review/exploration to least privilege; require a visible confirmation before full access; include the selected profile in output and the agent card. -- **Evidence:** Codex is hardcoded to `-a never -s danger-full-access` in both definition and serialized handler (`.agents/codex-cli.ts:76-83,95-124`); Claude uses `--dangerously-skip-permissions` (`.agents/claude-code-cli.ts:5-21,40-45`); Gemini uses `--yolo` (`.agents/gemini-cli.ts:5-27,46-51`). `createCliAgent` exposes only work/review mode and no permission field (`.agents/lib/create-cli-agent.ts:20-62`). - -## [MEDIUM] correctness / UX — local-agent source links use a second, narrower loader that disagrees with the SDK - -- **Risk:** A valid `.tsx`, `.js`, `.mjs`, or `.cjs` agent can load and run but have no “Open file” link; regex scanning may also associate an ID with a comment/nested object or a different duplicate than the module the SDK actually selected. -- **Fix:** Use the SDK-provided `_sourceFilePath` on each loaded definition as the sole source of truth. Remove the regex directory rescan, or make the SDK return a typed public source-path field. -- **Evidence:** The SDK supports `.ts`, `.tsx`, `.js`, `.mjs`, and `.cjs` and records the actual module path as `_sourceFilePath` (`sdk/src/agents/load-agents.ts:142-157,243-246`). The CLI discards that field and rebuilds paths with `/id\s*:/` while accepting only `.ts` files (`cli/src/utils/local-agent-registry.ts:101-146`), then falls back to `filePath: ''` (`:152-157`). Existing CLI tests even describe non-TypeScript files as artifacts to ignore (`cli/src/__tests__/integration/local-agents.test.ts:292-316`), contradicting the SDK contract. - -## [MEDIUM] correctness / state mutation — registry reinitialization does not invalidate derived UI listings - -- **Risk:** If initialization is rerun after files or MCP config change, runtime definitions refresh while the `@` menu can continue returning the old cached array, so configuration and visible selection diverge. -- **Fix:** Clear `cachedAgentsByMode` and `cachedAgentsDir` at registry refresh boundaries; expose one supported reload operation that atomically returns agents, paths, MCP config, and diagnostics. -- **Evidence:** `initializeAgentRegistry` replaces `userAgentsCache`, `userAgentFilePaths`, and `mcpServersCache` (`cli/src/utils/local-agent-registry.ts:55-87`) but does not clear `cachedAgentsByMode`; `loadLocalAgents` returns the cached reference before reading refreshed state (`:232-249`) and populates it at `:299`. Only the test reset clears it (`:451-460`). The multi-initialize test checks deduplication but never caches a listing between changed initializations (`cli/src/__tests__/integration/local-agents.test.ts:1025-1051`). - -## [MEDIUM] UX / API contract — every `spawn_agents` tool call is mislabeled as “Review” - -- **Risk:** Discovery, research, implementation, testing, and debugging batches all appear as review stages. Collapsed histories become semantically misleading and users cannot tell why a batch exists or whether it is still queued/running/failed. -- **Fix:** Default to “Spawned N agents” and render structured phase/purpose when supplied; include queued/running/succeeded/failed counts and short per-agent prompt summaries. -- **Evidence:** `SpawnAgentsComponent` explicitly documents itself as “the reviewer stage,” hardcodes `const header = 'Review'`, and reduces all inputs to comma-separated types (`cli/src/components/tools/spawn-agents.tsx:8-46`). The runtime uses `spawn_agents` generically, and the CLI separately creates lifecycle-aware agent blocks for every requested child (`cli/src/utils/sdk-event-handlers.ts:290-330`). No focused renderer test was found. - -## [MEDIUM] UX / correctness — expanded dependency rows are absent from keyboard focus and scroll geometry - -- **Risk:** In a large local-agent graph, expanded descendants add many visual rows while focus remains indexed only over top-level agents. Arrow navigation can scroll the wrong location or leave the focused row offscreen, and dependency entries cannot be reached or expanded by keyboard. -- **Fix:** Flatten top-level agents and visible dependency nodes into one keyboard-navigation model with measured row offsets; add left/right or Enter affordances for expansion and expose missing/cyclic dependencies in-row. -- **Evidence:** Focus scrolling assumes one line per `filteredAgents` item and computes `focusedTop = focusedIndex` (`cli/src/components/agent-checklist.tsx:194-215`), while `DepTree` inserts recursive rows after that top-level item (`:257-381`). `needsScroll` also compares only `filteredAgents.length` to height (`:227`), excluding expanded rows. - -## [LOW] error handling / UX — catastrophic registry failures are log-only and erase all local-agent visibility - -- **Risk:** A loader-level failure (as opposed to an ordinary bad file) replaces every user agent/path with empty caches and surfaces only a log warning, leaving the picker indistinguishable from “no agents configured.” File-path scan failures are also swallowed, hiding why links are missing. -- **Fix:** Retain the last known-good registry on refresh failure and surface typed startup/picker diagnostics with source path and retry/open actions. -- **Evidence:** The outer loader catch empties both caches and only calls `logger.warn` (`cli/src/utils/local-agent-registry.ts:55-69`); path reads/directories use empty catches (`:110-139`). `getLoadedAgentsData` returns `null` when nothing remains (`:433-444`), providing no error state to the UI. - -## [LOW] performance / eval feedback — final checks are serial and timeout/cancellation collapse into exit code 1 - -- **Risk:** Independent eval checks consume wall time sequentially and a check aborted by the shared 60-minute signal is reported like an ordinary command failure, obscuring harness exhaustion versus product failure. -- **Fix:** Add typed outcomes and elapsed time; support per-check timeouts and bounded parallel groups while preserving explicit sequential groups for dependent commands. -- **Evidence:** Agent execution and final checks share one timeout signal (`evals/buffbench/agent-runner.ts:89-97,174-196`). `runFinalCheckCommands` awaits commands in a `for` loop and coerces nonnumeric abort codes to `1` (`:227-265`). - -## [LOW] test coverage gap — spawn reconciliation and spawn renderer lack focused regression coverage - -- **Risk:** Same-type batches, out-of-order starts, nested parents, cancellation-before-start, unmatched starts, and phase-neutral rendering can regress without a targeted failure. -- **Fix:** Add table-driven matcher/handler tests with two identical types across one and multiple batches, reversed starts, correlation metadata, nested parents, cancellation cleanup, and renderer snapshots/status assertions. -- **Evidence:** `cli/src/utils/__tests__/sdk-event-handlers.test.ts` covers failed finish state, but scoped search found no direct tests for `findMatchingSpawnAgent` or `SpawnAgentsComponent`; the event schema itself currently cannot express the required correlation. - -## Rejected or narrowed candidates - -- **Rejected:** “One invalid local-agent file clears the entire menu.” The SDK catches import/runtime/validation problems per file and continues (`sdk/src/agents/load-agents.ts:219-270`); CLI integration tests confirm a valid agent survives a syntax-error neighbor (`cli/src/__tests__/integration/local-agents.test.ts:920-954`). The remaining finding is only for catastrophic outer-loader failure and log-only diagnostics. -- **Narrowed:** Serial final checks are not inherently incorrect because commands may depend on earlier checks. The defect is lack of dependency/concurrency metadata and loss of timeout/cancel classification, so severity remains LOW. -- **Narrowed:** Cached listings are currently initialized mainly at startup (`cli/src/index.tsx:286-290`), so stale-cache impact is reload/development/future-watch UX rather than every ordinary run. - -## Coverage across all 8 domains - -- **Security:** unrestricted external CLI profiles verified; no credential text leak proven in the scoped renderers. -- **Correctness:** same-type spawn reconciliation, source-path contract mismatch, stale derived listings, and expanded-row geometry verified. -- **State mutation:** optimistic map consumption and registry cache invalidation traced through callers. -- **Error handling:** ordinary bad files are isolated; catastrophic registry failure and eval abort classification remain weak. -- **Performance:** expanded list geometry and serial final checks reviewed; no render-loop hotspot proven. -- **Dependency hygiene:** no undeclared or vulnerable package dependency established in this shard. -- **Test coverage:** focused matcher/renderer, JS-agent source-link, refresh-after-cache, keyboard-expanded-tree, and permission-profile tests are missing. -- **API/ABI contracts:** `subagent_start` lacks spawn correlation; SDK source-file support and CLI listing disagree; external CLI output lacks permission metadata; “Review” is a misleading generic tool semantic. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-picker.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-picker.md deleted file mode 100644 index a0d76e018a..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-picker.md +++ /dev/null @@ -1,127 +0,0 @@ -# Discovery / file-picker audit - -## Flow inventory - -- `file-picker` spawns one `file-lister`, parses newline-delimited paths, applies lexical containment, ranks by prompt tokens, caps at 12, reads the files, then gives the model one final synthesis step (`agents/file-explorer/file-picker.ts:298-378`). -- `file-lister` runs one graph search (`limit: 24`) and one tree read (`maxTokens: 500_000`), then asks the model to emit exactly paths (`agents/file-explorer/file-lister.ts:59-78`). Directory constraints affect only the tree read, not graph retrieval. -- `researcher-web` chooses the first URL in the prompt for direct fetch; otherwise heuristically decomposes numbered/questions/bulleted/comparison prompts, with three total search calls and one retry per failed subquery, then concatenates result prose and a global link list (`agents/researcher/researcher-web.ts:241-388`). -- `researcher-docs` is a prompt-only agent allowed one `read_docs` call and free-form final prose (`agents/researcher/researcher-docs.ts:17-26`). -- `browser-use` is a long-running structured-output browser operator with snapshot-first guidance and broad page/state/media tools (`agents/browser-use/browser-use.ts:50-158`, `200-266`). -- `librarian` validates a public GitHub URL, shallow-clones it into `/tmp`, then gives an LLM unrestricted shell exploration and returns answer/files/cloneDir for the parent to inspect and clean (`agents/librarian/librarian.ts:85-186`). - -## [HIGH] Security / UX — agents/browser-use/browser-use.ts:152-158 — Browser agent has powerful external side effects without an action boundary - -- **Risk:** The agent can submit forms, upload files, alter cookies/storage, and run page-context JavaScript, but its contract does not distinguish read-only inspection from consequential actions or require confirmation; a prompt-injected page can steer a browsing run into state-changing actions. -- **Fix:** Add an explicit `interactionPolicy` (`observe`, `safe-test`, `allow-side-effects`), default to observe/safe-test, prohibit credential/payment/account/message actions without parent authorization, and require the output to enumerate side effects. -- **Evidence:** `browser_logs`, `run_terminal_command`, upload, cookie, storage, and evaluate are all available while the workflow simply says “Execute the task step by step.” - -## [HIGH] Security — agents/librarian/librarian.ts:168-186 — Untrusted cloned repositories are explored with a general shell - -- **Risk:** Repository text is untrusted, yet the LLM receives `run_terminal_command`; prompt injection in README/source can induce commands beyond read-only inspection, including executing repository scripts or reading unrelated host paths. -- **Fix:** Replace general shell access with a jailed read-only `/tmp/` file/search tool, or enforce command/path allowlists and explicitly prohibit execution, network, symlink traversal, and reads outside `cloneDir`. -- **Evidence:** The added prompt recommends shell utilities but contains no enforceable command or filesystem boundary, and `yield 'STEP_ALL'` leaves subsequent commands model-directed. - -## [MEDIUM] Correctness / UX — agents/file-explorer/file-lister.ts:59-75 — Directory-scoped discovery is only half scoped - -- **Risk:** When callers specify directories, `query_index` still searches the entire repository, so its top 24 results can crowd out relevant in-scope files; the model then receives conflicting global graph results and scoped tree context. -- **Fix:** Pass directory/path filters into `query_index` (or filter its candidates) and expose whether discovery was repository-wide or scoped in the result. -- **Evidence:** `directories` is used only as `read_subtree.paths`; the index call accepts only `query` and `limit`. - -## [MEDIUM] Performance / UX — agents/file-explorer/file-lister.ts:63-75 — Every lookup can request a 500k-token subtree - -- **Risk:** A routine file pick can load an enormous repository tree after already querying the index, adding latency/cost and making relevance selection noisier; there is no progressive expansion or truncation feedback. -- **Fix:** Start with graph results and a bounded shallow tree, expand only relevant directories, and return truncation/coverage metadata to the picker. -- **Evidence:** The unconditional second tool call uses `maxTokens: 500_000`, including when no directory is specified. - -## [MEDIUM] Correctness — agents/file-explorer/file-picker.ts:273-320 — Path validation is lexical and may both reject valid files and accept non-file prose - -- **Risk:** Any string containing `..` is rejected (including legitimate names), while arbitrary newline text, directories, nonexistent paths, and decorated Markdown paths can pass and be sent to `read_files`; if every extracted path is dropped, the code still proceeds as a successful discovery. -- **Fix:** Require structured `{path, reason, score}` output from `file-lister`, normalize paths with a real path utility, validate existence/type through the file tool contract, and emit a distinct “no safe files” result. -- **Evidence:** `trimmed.includes('..')` is the traversal test; parsing is `fileListText.split('\n')`; success is based on extracted text before safety filtering. - -## [MEDIUM] Correctness / UX — agents/researcher/researcher-web.ts:275-350 — Research budget silently leaves decomposed questions unanswered - -- **Risk:** Up to five subquestions are produced, but only three total calls are allowed and retries consume the same global budget; later questions disappear entirely rather than being marked skipped, producing a report that looks comprehensive but is not. -- **Fix:** Make budget proportional/configurable, reserve at least one call per selected subquestion, prioritize explicitly, and include searched/skipped/failed coverage metadata in output. -- **Evidence:** `MAX_SUBQUERIES = 5` but `MAX_TOTAL_CALLS = 3`; the loop stops when total calls reaches three and only visited questions create sections. - -## [MEDIUM] Correctness / UX — agents/researcher/researcher-web.ts:317-360 — Citations are not attached to claims - -- **Risk:** Search-result prose is placed into sections while all links are deduplicated into one global list, so the parent/user cannot tell which source supports which claim or subquestion; link text is also trusted verbatim. -- **Fix:** Preserve per-result source IDs and URLs, render inline citations per section/claim, include publication date/domain, and distinguish quoted evidence from synthesis. -- **Evidence:** `sections` stores only `{question, result}` while links accumulate separately in `allLinks`, then render under one `Sources / Links` heading. - -## [MEDIUM] Feature gap / API contract — agents/researcher/researcher-web.ts:300-304,370-387 — Search has no source, date, locale, depth, or pagination controls - -- **Risk:** Users cannot request official-only evidence, recent sources, regional results, deeper research, or continuation; every query uses fixed `depth: 'standard'`, making quality and reproducibility opaque. -- **Fix:** Add structured params for recency/date range, allowed/blocked domains, locale, depth, max results/pages, and citation mode; report applied constraints and exhaustion/truncation. -- **Evidence:** The public schema exposes only `prompt`; all query calls hard-code standard depth. - -## [MEDIUM] Error handling / API contract — agents/researcher/researcher-docs.ts:17-26 — Docs lookup is entirely model-driven and unstructured - -- **Risk:** There is no deterministic handleSteps flow, library/version/topic input, empty/error contract, citations, or relevant-doc identifiers; “use once” prevents refinement after ambiguous or stale library resolution. -- **Fix:** Add explicit library/package/version/topic params, deterministic resolve-then-fetch behavior, one bounded fallback, and structured output containing answer, doc URLs/IDs, version, and unresolved ambiguities. -- **Evidence:** The definition only grants `read_docs` and relies on prompt instructions for a single call and prose response. - -## [MEDIUM] State mutation / UX — agents/librarian/librarian.ts:127-186 — Temporary clones have no lifecycle ownership - -- **Risk:** Cleanup is delegated to the parent via prose, so cancelled, failed, or forgotten runs leak clones in `/tmp`; there is no retention option, cleanup status, or automatic finalizer. -- **Fix:** Default to automatic cleanup after producing excerpts/artifacts, add `retainClone` explicitly, and have the runtime finalizer remove clones on errors/cancellation. -- **Evidence:** The spawner prompt tells the parent to `rm -rf`; the agent only returns `cloneDir` and never cleans it. - -## [LOW] Error handling — agents/librarian/librarian.ts:149-168 — Missing/malformed clone results are treated as success - -- **Risk:** If the tool result is absent, non-JSON, or lacks `exitCode`, execution continues to “Clone complete,” producing confusing downstream shell errors. -- **Fix:** Require a recognized result with `exitCode === 0` and verify the clone directory/repository before handing off. -- **Evidence:** Failure is handled only inside `if (result && result.type === 'json')`; all other shapes fall through. - -## [LOW] Error handling / Security — agents/librarian/librarian.ts:155-163 — Raw git stderr is returned to the parent - -- **Risk:** Clone failures may expose local paths, proxy details, credentials embedded by tooling, or noisy internals in user-visible output. -- **Fix:** Log redacted diagnostics and return a stable categorized error plus a short sanitized detail. -- **Evidence:** The output message directly concatenates `stderr`. - -## [LOW] UX — agents/browser-use/browser-use.ts:210-210,248-257 — Visual smoke guidance over-tests by default - -- **Risk:** Requiring screenshot, PDF, and recording for broadly worded visual/browser smoke tasks adds latency and artifacts even when a screenshot and console check would answer the question. -- **Fix:** Make evidence modes task-selectable and default to the minimum sufficient artifact; reserve PDF/recording for explicit requests or reproduction evidence. -- **Evidence:** Rule 5 says to explicitly exercise all three media actions when the task asks for visual/browser smoke coverage. - -## [LOW] Dependency hygiene — agents/librarian/librarian.ts:64-70 — Agent prompts assume host CLI utilities - -- **Risk:** `tree`, `grep`, `find`, and Git availability/behavior are environment-dependent, so runs can fail inconsistently across installations. -- **Fix:** Prefer runtime-owned filesystem/search primitives or probe capabilities and provide portable fallbacks. -- **Evidence:** The system prompt directs use of specific host commands without a capability contract. - -## [LOW] Test coverage gaps — agents/browser-use/browser-use.test.ts:1-11 — Browser coverage is an opt-in smoke script, not contract/error-path tests - -- **Risk:** Schema/prompt regressions, missing set_output, unsafe action-policy behavior, timeout/cancellation, and Chrome-unavailable handoff can ship without routine CI detection. -- **Fix:** Add cheap unit/contract tests for definition and structured output plus hermetic tool-sequence tests; retain opt-in live tests separately. -- **Evidence:** The file is environment-guarded and described as expensive trace generation; unlike picker/researcher tests it does not run as ordinary test coverage. - -## [LOW] Test coverage gaps — agents/**tests**/file-picker.test.ts:421-529 — Ranking tests validate filename token overlap, not retrieval quality - -- **Risk:** The picker can pass tests while missing semantically adjacent types/tests/configs, mishandling scoped queries, or reading zero safe paths. -- **Fix:** Add benchmark fixtures for graph adjacency, directory scope, nonexistent/decorated paths, all-paths-dropped, subtree truncation, and recall@12 against expected relevant sets. -- **Evidence:** Current tests focus on deterministic keyword ordering and the 12-path cap. - -## 8-domain disposition - -- Security: browser/librarian action boundaries and prompt-injection surfaces are material; URL and clone URL lexical guards are positive defense-in-depth. -- Correctness: scoped retrieval, path parsing, research-budget coverage, and malformed clone results need work. -- State mutation: temporary clone cleanup and browser side-effect ownership are unclear. -- Error handling: picker has an extraction error path, researcher records searched-subquery failures, but docs and malformed clone/browser environment paths are weak. -- Performance: 500k subtree reads, fixed research budgets, and mandatory media evidence are poorly adaptive. -- Dependency hygiene: librarian depends on host Git/shell utilities; no additional package-version issue was established in this shard. -- Test coverage: picker and web researcher have extensive generator tests; docs/browser/lifecycle and retrieval-quality evaluation remain thin. -- API/ABI contracts: free-form last-message outputs for picker/research agents lack coverage/citation/truncation metadata; browser is notably better with structured output. - -## Compact inventory for paired code-searcher - -- Inspect `query_index` contract for path filters, modes, truncation, `relatedFiles`, and pagination; compare with `file-lister.ts:63-67`. -- Inspect `read_subtree` behavior for empty paths and 500k token caps; compare with `file-lister.ts:70-75`. -- Inspect `read_files` behavior for `paths: []`, nonexistent paths, directories, Markdown-decorated paths, and absolute in-root paths; compare with `file-picker.ts:318-376`. -- Inspect `web_search` contract/backend for redirects, DNS/private-address checks, result/link provenance, domain/date filters, pagination, timeouts, and `max_links`; compare with `researcher-web.ts:241-387`. -- Inspect `read_docs`/Context7 contract for version selection, citations, ambiguity, token limits, and retries; compare with `researcher-docs.ts:19-25`. -- Inspect `browser_logs` authorization and URL/navigation policy, download/upload boundaries, evaluate sandbox, cancellation, and artifact cleanup; compare with `browser-use.ts:152-266`. -- Inspect terminal tool cwd/path enforcement and cancellation cleanup for `/tmp/librarian-*`; compare with `librarian.ts:127-186`. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-search.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-search.md deleted file mode 100644 index 6ccdbe2aeb..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-search.md +++ /dev/null @@ -1,97 +0,0 @@ -# Discovery / research — code-search findings - -## [CRITICAL] security / API contract — `web_search` direct fetch has no backend SSRF or redirect boundary - -- **Risk:** Any agent allowed `web_search` can fetch loopback, RFC1918, link-local/cloud-metadata, or redirect from a public URL to an internal address. `researcher-web` performs only a lexical first-URL check, so redirects and other direct callers bypass its defense. -- **Fix:** Enforce egress policy in the tool handler: resolve DNS, reject private/reserved addresses before every connection, revalidate every redirect hop, disable credential forwarding, cap response size while streaming, and add IPv4/IPv6/encoded-host/rebinding tests. -- **Evidence:** The public tool accepts any valid URL (`common/src/tools/params/tool/web-search.ts:17-23`). The handler calls `fetch(fetchUrl)` with default redirect behavior and no hostname/IP validation (`packages/agent-runtime/src/tools/handlers/tool/web-search.ts:67-105`). The agent-side comment explicitly delegates DNS rebinding to the backend (`agents/researcher/researcher-web.ts:34-39`), but the backend has no such guard. - -## [HIGH] correctness / state mutation / UX — all browser agents share one process-global Chrome session - -- **Risk:** Parallel browser-use runs share active tab, cookies, storage, logs, recording state, and navigation; one run’s `stop`, tab switch, or storage clear can corrupt another run and misattribute screenshots/results. -- **Fix:** Key browser sessions by client/run/agent ID and pass that identity through `browser_logs`; isolate user-data directories and cleanup per owner. Reject or serialize legacy unkeyed concurrent use. -- **Evidence:** `sdk/src/tools/browser-logs.ts:80` declares one module-global `browserSession`; `ensureBrowserSession` returns it to every caller (`:495-497`), and `stopBrowser` clears/kills that same singleton (`:611-629`). `BrowserSession` contains shared active target, logs, networks, and recording (`:48-59`). Existing browser tests cover schemas/helpers, not concurrent ownership (`sdk/src/__tests__/browser-logs.test.ts:17-180`). - -## [HIGH] security / UX — browser-use has state-changing actions without an interaction policy - -- **Risk:** A browsing task or prompt-injected page can submit forms, upload local files, change cookies/storage, execute page JavaScript, or interact with accounts without a read-only boundary or side-effect disclosure. -- **Fix:** Add `interactionPolicy: observe | safe-test | allow-side-effects`, default to `observe`; require parent/user authorization for authentication, messaging, purchase/payment, account changes, uploads, and destructive form submissions; report every side effect in structured output. -- **Evidence:** The agent exposes browser actions plus general terminal access (`agents/browser-use/browser-use.ts:152-158`) and explicitly permits click/type/upload/cookie/storage/evaluate (`:178-198,246-257`). Its input schema only has a URL (`:32-47`) and output has no side-effect ledger (`:50-148`). - -## [HIGH] security / state mutation — librarian explores untrusted clones with an unrestricted shell and no enforceable cwd - -- **Risk:** Prompt injection in README/source can induce execution of repository scripts or reads/writes outside the clone, including host/project data. Prompt rules are advisory, while `STEP_ALL` leaves commands model-directed. -- **Fix:** Provide read-only clone-scoped list/search/read tools; otherwise enforce an OS sandbox, fixed cwd, command allowlist, no network/exec, symlink containment, and read-only filesystem policy. -- **Evidence:** Librarian grants `run_terminal_command` (`agents/librarian/librarian.ts:58`) and instructs the model to use arbitrary shell utilities on untrusted content (`:60-83,170-186`). Only the initial clone URL/command is hardened (`:98-147`); subsequent exploration has no handler-enforced clone boundary. - -## [MEDIUM] correctness / UX — directory-scoped file listing cannot scope graph retrieval - -- **Risk:** For a narrow directory request, global top-24 index hits can crowd out in-scope files and conflict with the scoped tree, lowering recall and making results non-reproducible. -- **Fix:** Add `pathPrefixes`/exclude filters to `query_index`, or filter/overfetch results inside file-lister; return scope and truncation metadata. -- **Evidence:** `file-lister` applies `directories` only to `read_subtree` while `query_index` receives only query/limit (`agents/file-explorer/file-lister.ts:59-75`). The `query_index` schema supports file types and graph modes but no path/directory filter or pagination cursor (`common/src/tools/params/tool/query-index.ts:10-56`). - -## [MEDIUM] performance / UX — file-lister overrides the subtree tool’s normal budget by 50× - -- **Risk:** Every lookup can request a 500k-token whole-repository tree even after index retrieval, increasing scan/truncation work, latency, and irrelevant context. -- **Fix:** Start with index results and a <=10k shallow/scoped tree; expand selected directories progressively and expose truncation/recovery in the file-lister output. -- **Evidence:** `file-lister` unconditionally calls `read_subtree` with `maxTokens: 500_000` and empty paths mean project root (`agents/file-explorer/file-lister.ts:70-75`; handler `packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:322-345`). The tool contract says normal use should not exceed 10,000 (`common/src/tools/params/tool/read-subtree.ts:45-49`) and separately caps live scans at 1,000 nodes (`read-subtree.ts:24,171-199`). - -## [MEDIUM] correctness / error handling — file-picker treats unstructured prose as paths and can issue an empty read - -- **Risk:** Markdown bullets, explanations, directories, nonexistent names, and all-filtered results pass the initial “has results” check; after safety filtering the picker can call `read_files` with `paths: []`, yielding a tool validation failure rather than a clear discovery outcome. -- **Fix:** Make file-lister structured (`{path, reason, score}`), validate file existence/type before ranking, strip only documented legacy decoration, and return a distinct `no_safe_files`/partial-coverage result. -- **Evidence:** Spawn output is split on newlines with every nonempty line accepted (`agents/file-explorer/file-picker.ts:240-271`); safety is lexical and rejects any name containing `..` (`:273-320`); the filtered/capped array is passed directly to `read_files` (`:340-378`). `read_files` requires at least one selector (`common/src/tools/params/tool/read-files.ts:120`). - -## [MEDIUM] correctness / UX — researcher-web silently omits budget-exhausted questions - -- **Risk:** A decomposed five-part request may search only one or two questions when retries consume the three-call budget; later questions receive no section or “skipped” marker, so the report appears complete when it is not. -- **Fix:** Budget per selected question, prioritize explicitly, make depth/call budget configurable, and emit searched/failed/skipped coverage metadata. -- **Evidence:** Decomposition returns up to five questions but broad research permits only three total calls, including retries (`agents/researcher/researcher-web.ts:120-190,275-350`). Only loop-visited questions are added to `sections`; unvisited questions vanish (`:283-350`). - -## [MEDIUM] correctness / API contract — web reports lose claim-to-source provenance and research controls - -- **Risk:** The parent cannot tell which URL supports a section/claim, request official-only or recent sources, choose locale, paginate, or reproduce a search. Search results are embedded as JSON prose while fetched-page links are unrelated navigation links. -- **Fix:** Return structured results per subquestion with source IDs, title, URL, snippet/date/domain and inline citations; add allowed/blocked domains, recency/date, locale, result count/cursor, and depth params. -- **Evidence:** Researcher sections store only `{question, result}` and all links are merged globally (`agents/researcher/researcher-web.ts:278-280,317-360`). The public web tool offers query/url/depth and fetch-link controls only (`common/src/tools/params/tool/web-search.ts:9-48`); search handler serializes title/url/description records into one string (`packages/agent-runtime/src/tools/handlers/tool/web-search.ts:175-187`). - -## [MEDIUM] error handling / API contract — researcher-docs cannot disambiguate version, source, or failure - -- **Risk:** The one-call model-driven agent may choose the wrong library/version and cannot refine. Tool failures are returned in the same `documentation` string as successful docs, with no URL/library ID/version/citation contract. -- **Fix:** Add deterministic library resolution, explicit version/topic/token params, one ambiguity retry, and structured `{status, libraryId, version, documentation, sources}` output. -- **Evidence:** Agent input is prompt-only and mandates exactly one `read_docs` call plus prose (`agents/researcher/researcher-docs.ts:10-26`). The tool accepts library title/topic/tokens but no version and outputs only `documentation` (`common/src/tools/params/tool/read-docs.ts:9-32,75-84`). Handler encodes errors inside `documentation` (`packages/agent-runtime/src/tools/handlers/tool/read-docs.ts:76-84,107-131`). - -## [MEDIUM] state mutation / UX — librarian clone cleanup has no runtime owner - -- **Risk:** Success, failure after cloning, cancellation, or a forgotten parent cleanup leaks `/tmp/librarian-*` trees indefinitely; the output cannot state retained/cleaned status. -- **Fix:** Auto-clean in a runtime finalizer, default to excerpts rather than persistent clones, and add explicit `retainClone` with expiry/cleanup status. -- **Evidence:** Clone paths are timestamped under `/tmp` (`agents/librarian/librarian.ts:127-145`); the agent only returns `cloneDir`, while the spawner prompt tells the parent to run `rm -rf` later (`:13-14,34-55`). No cleanup occurs in `handleSteps` (`:85-186`). - -## [LOW] error handling — malformed librarian clone results fall through as success - -- **Risk:** Missing/non-JSON results or JSON without a numeric exit code lead to “Clone complete” and later confusing shell failures. -- **Fix:** Require a recognized result with `exitCode === 0`, verify `.git`/directory existence, and return a typed failure otherwise. -- **Evidence:** Clone failure handling exists only inside `if (result && result.type === 'json')`; all other shapes continue at `agents/librarian/librarian.ts:149-168`. - -## [LOW] performance / UX — browser smoke mandates screenshot, PDF, and recording by default - -- **Risk:** Routine visual checks create unnecessary latency and media artifacts, especially under the 30-minute browser-agent timeout. -- **Fix:** Add selectable evidence modes and default to the minimum sufficient artifact; use PDF/recording only when explicitly requested or needed for reproduction. -- **Evidence:** Browser rules require all three for visual/browser smoke (`agents/browser-use/browser-use.ts:200-210,248-253`). - -## Rejected or narrowed candidates - -- **Rejected:** `read_subtree` itself lacks path containment. The handler rejects absolute/outside-root paths, skips symlinks, and avoids leaking outside-path existence (`packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:142-159,176-196,286-310`). The problem is file-lister scope/budget, not subtree path security. -- **Rejected:** `read_files` blindly reads unsafe/nonexistent paths. Its handler validates every selector as a batch and returns structured per-item errors (`packages/agent-runtime/src/tools/handlers/tool/read-files.ts:47-94`). The picker issue is its unstructured handoff and empty-selector UX. -- **Narrowed:** Researcher-web has meaningful lexical SSRF protection for the first URL (`agents/researcher/researcher-web.ts:34-100,241-268`), so the critical SSRF finding is specifically the shared backend and redirects, not absence of all defense. -- **Narrowed:** Initial librarian clone command injection is already mitigated by a strict GitHub regex and shell quoting (`agents/librarian/librarian.ts:98-147`); risk begins after cloning, when untrusted repository content drives unrestricted shell exploration. - -## Coverage across all 8 domains - -- **Security:** backend web SSRF, browser action authority, and librarian shell isolation verified; positive path/URL guards noted. -- **Correctness:** directory-scoped retrieval, picker handoff, browser session ownership, research coverage/provenance, docs ambiguity, and clone result parsing covered. -- **State mutation:** singleton browser state and clone lifecycle are material gaps. -- **Error handling:** read tools are structured; docs/librarian and skipped research questions are weak. -- **Performance:** 500k subtree request and mandatory browser media verified; no index algorithm hotspot proven. -- **Dependency hygiene:** librarian assumes Git plus host shell utilities; no vulnerable package version established. -- **Test coverage:** missing concurrent browser-session, backend SSRF/redirect, scoped query-index, empty-safe-picker, docs version/error, and clone cancellation/cleanup tests. Browser/librarian live tests are opt-in rather than routine contracts. -- **API/ABI contracts:** query_index lacks path filters/cursor; research outputs lack coverage/citation/version metadata; browser_logs lacks session ownership; librarian output lacks cleanup state. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/execution-picker.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/execution-picker.md deleted file mode 100644 index 200b798b41..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/execution-picker.md +++ /dev/null @@ -1,76 +0,0 @@ -# Execution specialists — file-picker audit - -Scope: active editor/basher definitions, their assigned tests, the orphaned best-of-N E2E, and directly invoked runtime/SDK contracts. Evidence is from the current working tree (the four editor/basher source/unit-test files are already user-modified). Targeted unit tests pass: `bun test agents/__tests__/editor.test.ts agents/__tests__/basher.test.ts` (65 pass). - -## [HIGH] Security — agents/basher.ts:79 — arbitrary shell execution has no enforceable approval boundary - -- **Risk:** A parent can pass any shell string to basher and the runtime executes it directly; the warnings about permission are only model-facing prose, so destructive, network, credential-reading, or out-of-project commands are not technically gated. -- **Fix:** Add a runtime command-policy/approval decision before execution (risk classification, explicit user grant/capability, auditable denial shape), with bypass only for a narrow allowlist of non-effectful validation commands. -- **Evidence:** `agents/basher.ts:79-139` forwards `params.command` directly; `packages/agent-runtime/src/tools/handlers/tool/run-terminal-command.ts:20-32` forwards it as `mode: 'assistant'`; `sdk/src/tools/run-terminal-command.ts:232-236` calls `bash -c `. The only approval rules are advisory text in `common/src/tools/params/tool/run-terminal-command.ts:82-92`; cwd containment at `sdk/src/tools/run-terminal-command.ts:141-161` does not constrain absolute paths or side effects inside the shell program. - -## [HIGH] Test coverage gaps / API contract — agents/e2e/editor-best-of-n.e2e.test.ts:12 — the best-of-N E2E tests a feature that is no longer active, and can pass vacuously - -- **Risk:** The test name/documentation claims parallel implementors, selection, and applying the winner, but it registers a single leaf proposal agent and asserts words already present in the prompt/session; regressions or total absence of best-of-N are therefore invisible. -- **Fix:** Delete/rename this stale E2E if best-of-N is intentionally removed, or restore an active orchestrator and assert N child runs, selector input/verdict, canonical applied mutation receipt, and final file contents. -- **Evidence:** `agents/e2e/editor-best-of-n.e2e.test.ts:12-23` defines `editor-best-of-n-max` locally with `spawnableAgents: []` and no `handleSteps`; `:74-79` supplies `params.n` that nothing consumes; `:96-115` accepts terms such as `multiply`, `function`, and `number` that already occur in the request. `agents/tsconfig.json:14-19` excludes the test, while `agents/tool-reachability.test.ts:121-139` explicitly says best-of-N definitions were deleted; implementations remain only under `agents-graveyard/editor/best-of-n/`. - -## [MEDIUM] Correctness / UX — agents/basher.ts:141 — `what_to_summarize` does not summarize or extract the requested information - -- **Risk:** Callers pay the UX cost of specifying a focus but receive a fixed, truncated dump; a request such as “only failing tests” can still return unrelated stdout and omit the relevant tail, contradicting the spawner/system contract. -- **Fix:** Either rename it to `report_label`/`focus_label`, or implement deterministic focused extraction plus a clearly signaled fallback; if semantic summarization is desired, add an explicit provider step with failure-isolated fallback to the raw structured result. -- **Evidence:** The agent promises analysis at `agents/basher.ts:14,61-78`, but `:141-238` only appends fixed fields and bounded stdout/stderr. `agents/__tests__/basher.test.ts:257-285` explicitly verifies that no provider `STEP` occurs and only checks raw output inclusion. - -## [MEDIUM] Error handling — sdk/src/tools/run-terminal-command.ts:288 — timeouts/spawn failures reject instead of returning a deterministic command result - -- **Risk:** Basher advertises deterministic reporting, but a timeout or spawn error escapes the client tool promise; callers lose a normal `exitCode`/`timedOut`/partial-output shape and recovery becomes a generic tool/subagent failure. -- **Fix:** Convert timeout/spawn failures to the terminal output schema with `errorMessage`, `timedOut`, command, elapsed time, and bounded partial stdout/stderr; reserve promise rejection for cancellation/runtime corruption. -- **Evidence:** `sdk/src/tools/run-terminal-command.ts:288-301,372-389` rejects; `packages/agent-runtime/src/tools/handlers/tool/run-terminal-command.ts:31-32` and `agents/basher.ts:129-157` do not catch. `agents/__tests__/basher.test.ts:165-189` checks only timeout forwarding, with no timeout-result/recovery test. - -## [MEDIUM] State mutation / recovery — sdk/src/tools/background-jobs.ts:327 — timeout and `kill_job` target only the shell PID, not its process tree - -- **Risk:** Commands that spawn children (dev servers, package scripts, pipelines) may leave grandchildren running after timeout/kill while Openbuff reports the tracked shell as stopped, leaking ports/processes across turns. -- **Fix:** Launch and track a process group/job object, terminate the full tree (POSIX group signal; Windows tree/job-object equivalent), wait for confirmed exit, then report residual-process failure if cleanup is incomplete. -- **Evidence:** sync execution spawns a shell at `sdk/src/tools/run-terminal-command.ts:232-236` and timeout kills only `childProcess` at `:288-300`; background execution does the same at `sdk/src/tools/background-jobs.ts:327-343`, and `killBackgroundJob` calls only `job.child.kill(signal)`/`process.kill(pid)` at `:504-519`. `packages/agent-runtime/src/tools/handlers/tool/end-turn.ts:24-49` tells users `kill_job` manages leaked work, making complete cleanup part of the UX contract. - -## [MEDIUM] Security / state mutation — agents/basher.ts:112 — full-log capture creates persistent shell-owned `/tmp` logs with no cleanup contract - -- **Risk:** Test/build output can contain secrets or proprietary paths; `tee` creates a discoverable local file using ambient umask, the path is returned to the model, and neither basher nor its tests remove it. -- **Fix:** Move capture into the SDK background/log subsystem using exclusive restrictive file creation, add TTL/end-of-run cleanup and an explicit retain option, and redact known secret patterns before returning excerpts. -- **Evidence:** `agents/basher.ts:112-126` writes via `tee` to `/tmp/openbuff-basher-.log`; `:164-167` publishes the path. `agents/__tests__/basher.test.ts:288-353` validates creation/reporting but has no permissions, redaction, or cleanup assertion. By contrast, SDK background logs use exclusive/no-follow creation at `sdk/src/tools/background-jobs.ts:316-326`. - -## [MEDIUM] Correctness / UX flow — agents/editor/editor.ts:224 — incomplete target-file progress is informational only and can be derived from unrelated backticked paths - -- **Risk:** The editor can finish with `pendingTargetFiles`, yet no active consumer blocks completion or requests recovery; backticked examples/non-goals are also treated as targets, producing noisy progress and an unreliable validation handoff. -- **Fix:** Pass target files as structured spawn input, return an explicit `complete|incomplete|failed` status with reasons, and have the orchestrator gate incomplete required targets; keep prompt-regex parsing only as a compatibility fallback. -- **Evidence:** `agents/editor/editor.ts:224-239,368-386` emits progress; `:407-420` scans both a section and every backticked filename. Repository search finds `targetFileProgress` only in this file and `agents/__tests__/editor.test.ts:575-645`, not in base2/runtime/CLI consumers. The validation gate records only confirmed `changedFiles` (`agents/base2/base2.ts:1602-1637`). - -## [LOW] API/ABI contract — agents/basher.ts:45 — basher exposes a weaker subset of the terminal contract - -- **Risk:** `process_type` accepts any string in the agent schema and `cwd` is unavailable, causing late validation failures and forcing callers to encode directory changes inside shell strings. -- **Fix:** Mirror the terminal schema: enum `SYNC|BACKGROUND`, bounded/integer timeout rules, and project-contained `cwd`; share a schema/helper to prevent drift. -- **Evidence:** `agents/basher.ts:45-53` declares plain string `process_type` and no `cwd`; the canonical tool has an enum and cwd at `common/src/tools/params/tool/run-terminal-command.ts:53-70`. Existing basher schema tests cover timeout/full-log fields but not process type or cwd (`agents/__tests__/basher.test.ts:47-111`). - -## Eight-domain coverage - -| Domain | Result | -| ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Security | Approval enforcement and persistent log findings above. | -| Correctness | False best-of-N coverage, misleading summarization, and unenforced editor completeness. | -| State mutation | Process-tree/background cleanup and temp-log lifecycle gaps. | -| Error handling | Timeout/spawn failures lack deterministic structured recovery. | -| Performance | No standalone hotspot ranked; editor does return full `newMessages` at `agents/editor/editor.ts:222-235`, worth measuring for parent-context amplification before changing the contract. | -| Dependency hygiene | No direct dependency issue found in this shard. | -| Test coverage gaps | Best-of-N E2E is vacuous/excluded; timeout, process-tree cleanup, log lifecycle, and incomplete-target handoff are untested. | -| API/ABI contract breaks | Basher schema drift and dead best-of-N ID/test contract above. | - -## Compact inventory for paired code-searcher - -- Editor definition/UX: `agents/editor/editor.ts:22-199` (`createCodeEditor`, tools, read-before-write/recovery instructions); loop/output `:201-239`; canonical mutation receipt scan `:241-366`; target progress/parser `:368-455`. -- Editor tests: `agents/__tests__/editor.test.ts:42-274` prompt/tool contract; `:276-539` step/output/mutation evidence; `:575-645` target progress. Missing: abort/partial/rollback-incomplete output semantics, incomplete-target orchestration, parent-context size. -- Editor spawn contract: `packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:277-358` requires five prompt labels using substring matching. -- Read-before-write/application: `packages/agent-runtime/src/tools/handlers/tool/write-file.ts:54-137,179-190`; `str-replace.ts:87-165`; `replace-range.ts:41-122`; `rewrite-symbol.ts:53-169`; `edit-transaction.ts:76-177`; `edit-application-coordinator.ts:84-221` (notably ambiguous “Queued for approval” is rejected in its test at `__tests__/edit-application-coordinator.test.ts:173-201`). -- Validation handoff: `agents/base2/base2.ts:1602-1637` consumes changed paths, not `targetFileProgress`. -- Basher definition/flow: `agents/basher.ts:9-78` schema/prompts; `:79-139` execution; `:141-238` raw/deterministic reporting; tests `agents/__tests__/basher.test.ts:47-111,114-479`. -- Terminal/runtime: `common/src/tools/params/tool/run-terminal-command.ts:9-75,76-103`; runtime forwarding `packages/agent-runtime/src/tools/handlers/tool/run-terminal-command.ts:9-32`; SDK execution/timeout `sdk/src/tools/run-terminal-command.ts:120-203,206-389`. -- Background recovery: `common/src/tools/params/tool/check-job.ts:7-88`; `kill-job.ts:7-63`; SDK jobs `sdk/src/tools/background-jobs.ts:303-529`; end-turn visibility `packages/agent-runtime/src/tools/handlers/tool/end-turn.ts:9-52`. -- Best-of-N status: stale E2E `agents/e2e/editor-best-of-n.e2e.test.ts:8-119`; excluded at `agents/tsconfig.json:14-19`; removed-ID guard `agents/tool-reachability.test.ts:121-139`; historical implementation `agents-graveyard/editor/best-of-n/`. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/execution-search.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/execution-search.md deleted file mode 100644 index bd11e4b4fe..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/execution-search.md +++ /dev/null @@ -1,76 +0,0 @@ -# Execution specialists — code-searcher audit - -## Verified findings - -## [HIGH] Security / API contract — agents/basher.ts:79 — Shell safety is advisory, not enforceable - -- **Risk:** Any parent allowed to spawn basher can execute an arbitrary shell program without a runtime approval/capability decision; model prompt injection or mistaken routing can perform destructive, network, credential, production, or git actions despite documented “ask first” rules. -- **Fix:** Enforce command risk policy at the client/runtime boundary, require a scoped user grant for effectful categories, attach provenance/approval IDs to executions, and retain a narrow non-effectful validation allowlist. -- **Evidence:** basher forwards `params.command` at lines 79-139; runtime forwards it unchanged with `mode: 'assistant'` (`packages/agent-runtime/src/tools/handlers/tool/run-terminal-command.ts:20-32`); SDK invokes `bash -c` (`sdk/src/tools/run-terminal-command.ts:220-236`). `tool-executor.ts` enforces tool availability, not command semantics, while approval rules exist only as prose at `common/src/tools/params/tool/run-terminal-command.ts:82-92` and `docs/agents-and-tools.md:33`. - -## [HIGH] Correctness / test coverage — agents/e2e/editor-best-of-n.e2e.test.ts:12 — Best-of-N coverage is stale and vacuous - -- **Risk:** The suite implies parallel proposals, selection, and winner application remain protected, but it defines a single leaf agent and accepts words already present in the prompt, so removal or total failure of the workflow is invisible. -- **Fix:** Remove/rename the obsolete test if the feature is intentionally gone, or restore an active orchestrator test that asserts N starts, selector evidence, an applied mutation receipt, and exact final contents. -- **Evidence:** the local definition has `spawnableAgents: []` and no orchestration `handleSteps` (`:12-23`); `params.n` is unused (`:74-79`); assertions accept prompt terms (`:96-115`); `agents/tsconfig.json:14-19` excludes it and `agents/tool-reachability.test.ts:121-139` states the definitions were deleted. - -## [MEDIUM] Correctness / UX — agents/basher.ts:141 — `what_to_summarize` labels output but does not summarize it - -- **Risk:** Callers can request “only failures” yet receive a fixed head-truncated report containing unrelated output and possibly missing the relevant tail, creating misleading validation feedback. -- **Fix:** Rename the parameter to reflect deterministic labeling, or implement structured focus extraction with head/tail preservation and an explicit fallback; expose truncation metadata. -- **Evidence:** the system/spawner promises analysis and focus at lines 13-14 and 61-78, but lines 141-238 only append fixed fields and the first 8,000 stdout/4,000 stderr characters. `agents/__tests__/basher.test.ts:257-285` intentionally asserts no provider step. - -## [MEDIUM] Error handling / UX contract — sdk/src/tools/run-terminal-command.ts:288 — Timeout and spawn failures reject outside the normal command-result schema - -- **Risk:** Basher and other callers lose exit status, partial output, elapsed time, and a typed timeout indicator; users see a generic tool/agent crash rather than a recoverable validation result. -- **Fix:** Return a deterministic result for operational failures (`timedOut`, `spawnFailed`, partial stdout/stderr, command, elapsedMs); reserve rejection for caller cancellation or runtime corruption. -- **Evidence:** timeout rejects at lines 288-301 and spawn error rejects at 372-389; runtime and basher do not translate either (`run-terminal-command.ts:31-32`, `basher.ts:129-157`). Existing basher tests cover forwarding but not timeout recovery. - -## [MEDIUM] State mutation / cancellation — sdk/src/tools/run-terminal-command.ts:135 — Cancelling a request does not stop background commands - -- **Risk:** A user can cancel the owning agent while its dev server/watcher continues running; the cleanup burden is shifted to a later `kill_job` call that the cancelled orchestrator may never issue. -- **Fix:** Associate background jobs with request/agent ownership, support `detach: true` as an explicit choice, and cancel owned jobs by default on request cancellation/end while surfacing any retained jobs. -- **Evidence:** the API comment explicitly says background jobs are unaffected by `AbortSignal` at lines 135-139 and 248-252. End-turn only reports pending work (`packages/agent-runtime/src/tools/handlers/tool/end-turn.ts:24-49`); it does not guarantee cleanup on cancellation. - -## [MEDIUM] State mutation / process lifecycle — sdk/src/tools/background-jobs.ts:327 — Kill and timeout target only the shell process, not its tree - -- **Risk:** Package scripts, pipelines, and dev servers can leave grandchildren alive after Openbuff reports a job stopped, leaking ports and mutations across turns. -- **Fix:** spawn process groups/job objects, terminate the full tree on timeout/cancel/kill, await confirmed exit, and report residual cleanup failure. -- **Evidence:** background spawn has no detached process-group ownership at lines 327-343; kill calls only `job.child.kill(signal)` or one PID at 504-519. Sync timeout similarly kills only `childProcess` at `run-terminal-command.ts:288-300`. - -## [MEDIUM] Security / state lifecycle — agents/basher.ts:112 — Full-log mode persists unowned `/tmp` output - -- **Risk:** Test/build output containing secrets, internal paths, or source fragments remains in a shell-created temp file with no retention or cleanup contract; the model receives its path. -- **Fix:** reuse the SDK’s exclusive/no-follow log facility with owner-only permissions, redaction, TTL/end-of-run cleanup, and explicit retention consent. -- **Evidence:** lines 112-126 pipe output through `tee /tmp/openbuff-basher-.log`, and lines 164-167 publish it. Tests validate creation/reporting but not permissions or cleanup. SDK background logs use guarded creation at `sdk/src/tools/background-jobs.ts:316-326`. - -## [MEDIUM] Correctness / orchestrator UX — agents/editor/editor.ts:224 — Incomplete target progress is neither authoritative nor consumed - -- **Risk:** Editor can return pending required files without blocking completion, while unrelated backticked filenames from examples or non-goals can be misclassified as targets; the parent’s validation handoff therefore cannot reliably distinguish complete from partial implementation. -- **Fix:** pass target files structurally, emit `complete|partial|failed` plus per-target reasons, and require the orchestrator to resolve incomplete required targets before finalization. -- **Evidence:** `targetFileProgress` is emitted at lines 224-239; target extraction scans all backticked filenames at 403-420. Search finds consumers only in editor tests, while base2 tracks confirmed changed files rather than progress (`agents/base2/base2.ts:1602-1637`). - -## [LOW] API/ABI contract — agents/basher.ts:45 — Basher’s input schema drifts from the terminal tool - -- **Risk:** Invalid `process_type` values fail late, timeout constraints differ, and lack of `cwd` encourages fragile `cd ... &&` command strings. -- **Fix:** share the canonical schema subset: `SYNC|BACKGROUND` enum, bounded/integer timeout, and project-contained cwd. -- **Evidence:** basher declares `process_type` as unrestricted string and has no cwd at lines 45-53; canonical params define enum/cwd in `common/src/tools/params/tool/run-terminal-command.ts:53-70`. - -## Rejected / downgraded candidates - -- **Terminal cwd can escape project — rejected.** SDK now resolves lexical and realpath/symlink containment and returns an invalid-cwd result (`sdk/src/tools/run-terminal-command.ts:141-161`), with dedicated tests. -- **Terminal output grows without bound — rejected.** Sync output has streaming accumulation caps and final truncation (`run-terminal-command.ts:304-359`); background output is file-backed. The remaining issue is log retention, not memory growth. -- **Tool permissions are absent — rejected.** Runtime checks whether an agent may call a tool (`tool-executor.ts:505-515,677`); the verified gap is effect-level approval within an authorized shell tool. -- **Best-of-N implementation is silently active — rejected.** Active reachability tests explicitly treat the old IDs as removed; this is a stale test/documentation contract, not an active hidden workflow. -- **Dependency hygiene issue — not established.** No undeclared or vulnerable third-party dependency was evidenced in this scope. - -## Coverage across 8 domains - -- Security: enforceable shell approval and persistent log exposure. -- Correctness: misleading summary parameter, stale best-of-N test, incomplete editor handoff. -- State mutation: process-tree leaks, background ownership on cancellation, temp-log lifecycle. -- Error handling: timeout/spawn failures escape structured command results. -- Performance: bounded output is present; orphan processes and serial recovery create indirect resource/latency costs. -- Dependency hygiene: no concrete issue found. -- Test coverage: no meaningful best-of-N E2E; missing timeout-result, tree-kill, cancellation ownership, log cleanup, and editor-partial orchestration tests. -- API/ABI contract: basher schema drifts from canonical terminal params; `what_to_summarize` and editor progress overpromise their semantics. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/orchestrator-picker.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/orchestrator-picker.md deleted file mode 100644 index bb33aada6b..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/orchestrator-picker.md +++ /dev/null @@ -1,65 +0,0 @@ -# Orchestrator core — file-picker audit - -> **As-of audit snapshot (2026-07):** tables that mention "inline mirrors" in `base2.ts` describe the **audit-time** dual-copy layout. Today, `gate-paths` / `gate-reviewer` / `gate-repair` / `gate-concurrency` are generator-synced via `scripts/generate-gate-helpers.ts` into `` (not hand-maintained). Findings below remain historical evidence unless re-verified. - -## Compact inventory for paired code-searcher - -| Flow | Primary files / symbols | Tests to inspect | Search next | -| --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | -| Mode construction and tools | `agents/base2/base2.ts:22` `createBase2`; wrappers `base2-{plan,execute-plan,fast,fast-no-validation,evals}.ts`; `agents/base2/base-deep.ts:315` `createBaseDeep` | `agents/__tests__/base2.test.ts:80-350`, `:2582-2631`; slow-only `agents/e2e/base-deep.e2e.test.ts:15-600` | Mode/tool/spawnable parity, actual plan-only edit enforcement, execute-plan resume behavior | -| Discovery/spawn policy | `agents/base2/base2.ts:104-250`, `:3322-3330`; `base-deep.ts:39-68`; `quality-prompt-section.ts:60-74` | prompt assertions in `base2.test.ts:321-350`, `quality-prompt-snapshot.test.ts:53-74` | Runtime tests for scope classification, required joins, mentioned-agent routing, spawn failure/timeout | -| Gate lifecycle/completion | `base2.ts:351-1530` `handleSteps`; state `gate-state.ts:1-118` | `base2.test.ts:352-2679`; `gate-lifecycle.e2e.test.ts`; `reviewer-spawn-conditions.e2e.test.ts` | step-cap with dirty files; cancellation during gate/repair/reviewer; restart/resume from every phase | -| Changed-file/path tracking | `gate-files.ts:26-122`; `gate-paths.ts:12-59`; inline mirrors `base2.ts:1602-1691`, `:2089-2661` | `gate-changed-files`, `gate-files-parity`, `gate-paths` | absolute outside-cwd paths, symlinks, deletes/renames, dirty-tree overlap | -| Aux agents | `base2.ts:653-805`, selectors `:1693-1849` | `gate-aux-triggers.test.ts`; `gate-aux-ordering.e2e.test.ts` | Result/verdict consumption, multi-package sets, user opt-out, no-op agents, latency/cost | -| Validation repair | `gate-repair.ts:47-177`; inline `base2.ts:1026-1276`, `:3095-3319` | `gate-repair*`, telemetry assertions `base2.test.ts:3120-3210` | editor crash/no-op, command timeout, repeated same failure, cancellation, rollback/partial edit | -| Reviewer contract | `gate-reviewer.ts:23-349`; inline `base2.ts:1279-1398`, `:2663-3030` | `gate-reviewer.test.ts`; reviewer e2e tests | malformed/contradictory results, background job loss across turns, timeout UX | -| Ask/cancel/recovery | prompt guidance `base2.ts:142`, `:199-221`, `:3359`; cancellation docs `docs/request-flow.md:230-239` | **No scoped orchestrator test found** | `ask_user` resume/cancel, Escape during child/hook/editor/reviewer, state cleanup and user-visible message | - -## Findings - -### [HIGH] Correctness / UX — step cap marks an ungated dirty turn finalizable - -- **Evidence:** `agents/base2/base2.ts:557-569` handles `hitStepCap` by setting `currentPhase = 'final_response_allowed'`, enabling followups, and breaking before validation/review, even when `pendingGateFiles` already exist. The regression test explicitly expects this state with `src/a.ts` pending (`agents/__tests__/base2.test.ts:1010-1064`). -- **Risk:** reaching `maxAgentSteps` can present incomplete or unvalidated edits as a normal green completion, contradicting the documented rule that failed/timed-out validation blocks completion (`docs/agents-and-tools.md:36-37`). -- **Fix:** introduce an explicit `interrupted`/`step_cap_reached` phase; preserve pending gate state and return a non-green resume message instead of opening finalization. - -### [HIGH] Correctness / UX — security-reviewer output is awaited but never interpreted - -- **Evidence:** `agents/base2/base2.ts:772-799` yields `security-reviewer` and discards the result; no blocker/crash/verdict parsing follows. Yet docs say blocking security findings prevent completion (`docs/agents-and-tools.md:27`) and base-deep says `BLOCKING:` security findings block completion (`agents/base2/base-deep.ts:53`). -- **Risk:** a security reviewer can identify a critical auth/secrets issue and the orchestrator will continue to validation/code-review unchanged; crashes are also silent. -- **Fix:** define a structured security verdict contract and persist blockers/crashes in active state, or clearly make this advisory everywhere and surface its findings to the final reviewer. - -### [HIGH] Correctness / Performance / UX — automatic test/doc agents fire for nearly every source edit - -- **Evidence:** all recognized non-test source files trigger test-writer (`base2.ts:1693-1769`), and every file under `packages/*/src`, `agents/`, `common/src`, or `cli/src` is treated as public API (`:1772-1795`). The doc-writer always targets `docs/agents-and-tools.md` (`:756-764`). Docs confirm a CLI component edit typically spawns both agents (`docs/agents-and-tools.md:352-358`), although the routing policy says test/doc writers are for acceptance-criteria-required coverage (`docs/agents-and-tools.md:28`). -- **Risk:** tiny internal refactors incur two serial agents, may create unwanted tests/docs, and can pollute the agent-specific documentation with unrelated CLI/package changes. -- **Fix:** gate on behavior/public-contract evidence or explicit acceptance criteria; route docs by package/API ownership; expose an opt-out/preview in UX. - -### [MEDIUM] Correctness — mixed-package test-writer routing uses only the first package command - -- **Evidence:** `selectTestWriterTargets` returns every eligible target but derives one command from `targetFiles[0]` (`base2.ts:1759-1769`). The unit test codifies this first-target behavior (`gate-aux-triggers.test.ts:275`). Docs instead claim “For each package” the package’s own scripts run (`docs/agents-and-tools.md:354`). -- **Risk:** a cross-package edit gives test-writer misleading validation ownership; later packages can receive tests without the relevant command being reported/run. -- **Fix:** partition targets by package and spawn one writer per package, or pass a target-to-command map. - -### [MEDIUM] Security — gate path normalization does not reject absolute paths outside cwd - -- **Evidence:** `normalizeGateFilePath` rejects `..` but returns other absolute paths unchanged (`agents/base2/gate-paths.ts:12-40`; inline `base2.ts:1643-1672`). The fingerprint helper then calls `path.resolve(cwd, normalizedPath)` and synchronously reads/hashes it (`base2.ts:2577-2650`). Tests cover cwd absolute paths but not `/etc/...` or another workspace (`gate-paths.test.ts:96-150`). -- **Risk:** malformed/injected serialized gate state can make the orchestrator read/hash files outside the project boundary; at minimum this violates the stated project-relative invariant and can create confusing gate failures. -- **Fix:** resolve then require the result to be inside realpath(cwd), with explicit Windows-drive/UNC and symlink tests. - -### [MEDIUM] API/UX — `context-pruner` is user-spawnable despite an explicit prohibition - -- **Evidence:** `context-pruner` is in both spawn allowlists (`base2.ts:104-124`, `base-deep.ts:353-372`) while both prompts say never spawn it because `handleSteps` automatically does so (`base2.ts:222`, `base-deep.ts:67`; actual automatic spawn `base2.ts:523-531`). Docs also advertise it as orchestrator-spawnable (`docs/agents-and-tools.md:17`). -- **Risk:** model/user-mentioned routing can duplicate pruning, waste latency, and create confusing hidden child activity. -- **Fix:** remove it from public `spawnableAgents` and reserve `spawn_agent_inline` as a runtime-owned internal path; update docs. - -## Coverage gaps across the 8 domains - -- **Security:** absolute/symlink path containment and security-review verdict handling are uncovered. -- **Correctness:** no tests for aux-agent crash/no-op, mixed-package orchestration, or dirty step-cap recovery. -- **State mutation:** no scoped test cancels/restarts during `repair_loop`, `awaiting_review`, or a background reviewer job. -- **Error handling:** reviewer parsing is strong, but aux-agent failures and automatic context-pruner failure are not surfaced/tested. -- **Performance:** no latency/cost budget test for three serial aux spawns plus validation/review; proactive `query_index` also has prompt-only coverage. -- **Dependency hygiene:** no third-party dependency issue found in this shard; duplicated inline helper bodies remain a maintenance dependency enforced only partly by parity tests. -- **Test coverage:** cancellation (`docs/request-flow.md:230-239`), `ask_user` resume, plan-only behavioral enforcement, execute-plan artifact recovery, agent timeout/join, and spawn-policy behavior lack scoped end-to-end coverage. `base-deep.e2e.test.ts` is provider-gated/slow (`:20-21`) rather than deterministic CI coverage. -- **API/ABI contracts:** tool/spawnable arrays and `` are public behavioral contracts, but there are no full parity assertions across every mode/wrapper or compatibility tests for new active-work phases. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/orchestrator-search.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/orchestrator-search.md deleted file mode 100644 index c3196e345c..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/orchestrator-search.md +++ /dev/null @@ -1,72 +0,0 @@ -# Orchestrator core — code-searcher audit - -> **As-of audit snapshot (2026-07):** references to selective parity / inline helper duplication describe the **audit-time** workflow. Gate helpers in `SOURCE_MODULES` (including `gate-fingerprint`) are now generator-synced (`generate-gate-helpers` + freshness test); residual dual-copy for fingerprint is closed. Findings below remain historical evidence unless re-verified. - -## Verified findings - -## [HIGH] Correctness / UX — agents/base2/base2.ts:557 — Step-cap exhaustion bypasses required validation and review - -- **Risk:** A turn that exhausts `maxAgentSteps` after editing is moved to `final_response_allowed`, enables follow-up suggestions, and exits before its pending files are gated, making unfinished/unvalidated work look like a normal completion. -- **Fix:** Preserve `awaiting_validation`/pending files, emit a distinct `step_cap_reached` interrupted state, disable green completion affordances, and make the next turn resume the gate before unrelated work. -- **Evidence:** lines 557-569 explicitly skip the gate, set `currentPhase = 'final_response_allowed'`, and set `canSuggestFollowups = true`; this conflicts with the completion contract in `docs/agents-and-tools.md:36-37`. `agents/__tests__/base2.test.ts:1010-1064` codifies pending dirty files with this finalizable phase instead of testing safe recovery. - -## [HIGH] Security / correctness / error handling — agents/base2/base2.ts:772 — Security-reviewer results and failures are discarded - -- **Risk:** Critical findings, malformed output, crashes, and timeouts have no effect on orchestration; the final code reviewer may miss domain-specific security issues while the docs promise that blocking security findings prevent completion. -- **Fix:** Give `security-reviewer` a structured verdict contract, parse and persist blockers/crashes like `gate-reviewer.ts`, expose them in active state/UI, and require resolution or an explicit audited override. -- **Evidence:** lines 791-798 yield the reviewer without capturing `toolResult`; execution then only checks `auxGateFiredThisIteration` and continues. By contrast, the final reviewer captures/parses results at lines 1278-1400. The behavioral contract says security blockers prevent completion at `docs/agents-and-tools.md:27` and `agents/base2/base-deep.ts:53`. - -## [HIGH] Correctness / observability — agents/base2/base2.ts:711 — Aux-gate telemetry declares success before an agent runs - -- **Risk:** Operational telemetry and any UX built on it report both reviewer and validation as `passed` before test-writer, doc-writer, or security-reviewer has started; a crash/no-op can therefore leave a false-green audit trail. -- **Fix:** Add a first-class aux gate and lifecycle status (`started`, `passed`, `failed`, `timed_out`, `skipped`), emit success only after interpreting the child result, and never reuse final reviewer/validation fields for unrelated agents. -- **Evidence:** test-writer emits `reviewerStatus: 'passed'` and `validationStatus: 'passed'` at lines 711-718 before yielding at 720; doc-writer repeats this at 747-754; security-reviewer repeats it at 782-789. None captures or validates the subsequent result. - -## [MEDIUM] Correctness / performance / UX — agents/base2/base2.ts:680 — Broad predicates force serial test/doc agents for routine internal edits - -- **Risk:** Most source changes incur two blocking agent runs before validation, adding latency/cost and potentially generating unwanted tests or edits to a single unrelated documentation file; users receive no preview or opt-out. -- **Fix:** Trigger on changed behavior/public contract or acceptance criteria, route documentation to package ownership, batch independent advisory work where safe, and show the proposed gates before execution. -- **Evidence:** lines 680-805 run test-writer then doc-writer then security-reviewer sequentially; `selectDocWriterTargets` treats all `packages/*/src`, `agents/`, `common/src`, and `cli/src` as public API and always supplies `docs/agents-and-tools.md` at lines 743-763. `docs/agents-and-tools.md:358` confirms a normal CLI component edit fires both agents. - -## [MEDIUM] Correctness / API contract — agents/base2/base2.ts:1759 — Cross-package test writing is routed with only the first package’s command - -- **Risk:** A multi-package change sends all target files with one command derived from `targetFiles[0]`, so later packages get incorrect validation guidance and the implementation contradicts the documented “for each package” behavior. -- **Fix:** Partition by package and spawn per-package writers or pass a target-to-command mapping with package-specific ownership. -- **Evidence:** `selectTestWriterTargets` returns all eligible files but computes one `testCommand` from the first target at lines 1759-1769; `agents/__tests__/gate-aux-triggers.test.ts:275` locks in this behavior, while `docs/agents-and-tools.md:354` describes per-package commands. - -## [MEDIUM] Security — agents/base2/gate-paths.ts:12 — Absolute paths outside the project pass gate normalization - -- **Risk:** Injected or corrupted durable state can make gate fingerprinting read/hash files outside the repository, violating project containment and potentially exposing local file content to agent state/logging. -- **Fix:** Resolve and realpath each path, require it to remain under realpath(cwd), reject foreign Windows drives/UNC paths, and handle symlink escapes explicitly. -- **Evidence:** `normalizeGateFilePath` rejects `..` segments but returns `path.normalize(filePath)` for any absolute path at lines 12-40. The inline fingerprint code later resolves and reads gate paths (`agents/base2/base2.ts:2547-2650`). Tests cover cwd absolute paths but not foreign absolute paths or symlink escape (`agents/__tests__/gate-paths.test.ts:96-150`). - -## [MEDIUM] API/UX / performance — agents/base2/base2.ts:104 — Context-pruner is simultaneously internal-only and publicly spawnable - -- **Risk:** Mention-based routing or model choice can create a second visible context-pruner alongside the automatic hidden one, wasting a child run and confusing users about which pruning operation controls context. -- **Fix:** Remove it from public `spawnableAgents`, reserve `spawn_agent_inline` for runtime-owned pruning, and document it as an internal lifecycle service rather than a selectable specialist. -- **Evidence:** it appears in the base2 and base-deep spawn allowlists (`base2.ts:104-124`, `base-deep.ts:353-372`) while prompts prohibit spawning it (`base2.ts:222`, `base-deep.ts:67`); the runtime already invokes it every loop at `base2.ts:523-531`; docs advertise it as orchestrator-spawnable at `docs/agents-and-tools.md:17`. - -## [LOW] Test coverage / state mutation — docs/request-flow.md:230 — Cancellation and restart are not exercised across orchestrator phases - -- **Risk:** Escape/abort during an inline aux agent, validation repair, background reviewer wait, or `ask_user` can strand job IDs, done flags, blockers, or pending files and produce an incorrect resume phase. -- **Fix:** Add deterministic lifecycle tests that cancel and resume from each active phase, asserting pending files, background jobs, gate flags, user-visible status, and exactly-once child joins. -- **Evidence:** targeted test search found extensive gate predicate/parser tests but no scoped end-to-end cases for cancellation during `repair_loop`, aux spawns, `awaiting_review`, or background reviewer wait; cancellation behavior is only documented in `docs/request-flow.md:230-239`. - -## Rejected / downgraded candidates - -- **Reviewer result parsing is generally missing — rejected.** The final code-reviewer has substantial structured/text verdict parsing, blocker extraction, crash differentiation, durable fingerprints, and parity tests (`base2.ts:1278-1478`, `gate-reviewer.ts`, `agents/__tests__/gate-reviewer.test.ts`). The gap is specific to aux reviewers. -- **Aux agents race final validation — rejected.** `spawn_agent_inline` is intentionally blocking and the loop continues after each aux spawn (`base2.ts:680-805`); the issue is serial cost and discarded outcomes, not a concurrency race. -- **All absolute paths are invalid — rejected.** Paths inside cwd are intentionally normalized and tested. Only foreign absolute paths and symlink escapes lack containment. -- **Gate state has no CLI rendering — rejected.** CLI parses and renders dedicated gate blocks (`cli/src/utils/message-block-helpers.ts:30-90`, `cli/src/components/renderers/gate-state-box.tsx`, component tests). The stronger UX gap is false/underspecified aux lifecycle telemetry. -- **Dependency vulnerability — not established.** No third-party package/version issue was evidenced in this scope. - -## Coverage across 8 domains - -- Security: foreign absolute/symlink path containment; ignored security-review findings. -- Correctness: step-cap bypass, discarded aux outcomes, first-package routing, contradictory context-pruner contract. -- State mutation: pending gate state on step cap and cancellation/resume gaps. -- Error handling: aux crashes/timeouts/no-op outputs are not classified or surfaced. -- Performance: serial broad aux gates and duplicate-pruner possibility. -- Dependency hygiene: no concrete dependency issue found; inline/module helper duplication remains guarded by selective parity tests. -- Test coverage: dirty step-cap behavior is tested with the unsafe expectation; cancellation, aux failure, path escape, and multi-package orchestration need coverage. -- API/ABI contract: docs promise blocking security review and per-package commands, while implementation does neither; phase/telemetry vocabulary cannot represent aux lifecycle or step-cap interruption accurately. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/quality-picker.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/quality-picker.md deleted file mode 100644 index 448863f741..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/quality-picker.md +++ /dev/null @@ -1,92 +0,0 @@ -# Quality/recovery agent file-picker audit - -> **As-of audit snapshot (2026-07):** "duplicating serialized inline helpers, partly guarded by parity tests" was true at audit time. Paths/reviewer/repair/concurrency helpers are now generator-synced; do not treat hand-copy maintenance as current workflow. Findings below remain historical evidence unless re-verified. - -## Compact inventory for paired code-searcher - -| Flow | Primary files / symbols | Current contract | -| ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | -| Final code review | `agents/reviewer/code-reviewer.ts:9-105` (`createReviewer`); `agents/base2/base2.ts:1278-1398`; `agents/base2/gate-reviewer.ts:23-64,256-349` | Read-only reviewer, text/optional JSON verdict; `BLOCKING` and `coverage: missing` reopen the gate. | -| Validation repair | `agents/base2/base2.ts:983-1273`; `agents/base2/gate-repair.ts:47-177`; `agents/base2/gate-state.ts:49-90` | Parse hook errors, run up to 3 targeted editor repairs, then one broader editor escalation. | -| Auxiliary quality agents | `agents/base2/base2.ts:653-805,1693-1849`; `agents/base2/gate-state.ts:92-117` | For each aux-relevant pending set: test-writer -> doc-writer -> security-reviewer, then validation + code-reviewer. | -| Security review leaf | `agents/security-reviewer/security-reviewer.ts:5-48`; policy at `agents/base2/quality-prompt-section.ts:98-126` | Read-only adversarial review; unstructured last-message report. | -| Test mutation leaf | `agents/test-writer/test-writer.ts:6-65`; target selection at `agents/base2/base2.ts:1721-1770` | Reads source, writes tests, reports a parent-owned command; cannot validate. | -| Docs mutation leaf | `agents/doc-writer/doc-writer.ts:6-75`; target selection at `agents/base2/base2.ts:1772-1795` | Reads source/docs and mutates docs; automated gate always supplies `docs/agents-and-tools.md`. | -| Debugging leaf | `agents/debugger/debugger.ts:6-76` | Immediately runs supplied reproduce command, reads suspects, then model-driven diagnosis; report only, no fix. | -| Git workflow | `agents/git-committer/git-committer.ts:6-107`; policy at `agents/base2/quality-prompt-section.ts:128-149` | Optional branch, status/diff/log, optional `git add -A`, then model-driven stage/commit. | -| Tests/wiring | `agents/__tests__/code-reviewer.test.ts`; `gate-reviewer.test.ts`; `gate-repair.test.ts`; `gate-aux-triggers.test.ts`; `git-committer.test.ts`; `new-bundled-agents.test.ts`; `agents/e2e/gate-lifecycle.e2e.test.ts`; `reviewer-spawn-conditions.e2e.test.ts`; `gate-aux-ordering.e2e.test.ts` | Strong parser/order/happy-path coverage; leaf-agent tests are mostly prompt/schema checks. | - -Suggested searches: `spawn_agent_inline`, `testWriterGateDone|docWriterGateDone|preEditSecurityReviewDone`, `collectReviewerBlockers|getReviewerFinalizationVerdict|detectReviewerCrash`, `MAX_REPAIR_ROUNDS|repairEscalationDone`, `stage_all|branch_name|git add -A`, and all consumers of aux-agent results. - -## Findings - -### [HIGH] Security / correctness / error handling — auxiliary agent results are discarded - -- **Evidence:** each aux flag is marked done before spawning, and the yielded result is never assigned or parsed: test-writer `agents/base2/base2.ts:709-730`, doc-writer `:745-766`, security-reviewer `:780-798`; the loop simply continues at `:800-805`. By contrast, code-reviewer output is captured and gated at `:1282-1398`. The orchestrator promises that a `BLOCKING` security finding blocks completion at `agents/base2/base2.ts:203-206`. -- **Risk/UX:** a security reviewer can report a critical exploit, or any aux agent can crash/timeout, and the runtime still proceeds to final validation/review with no user-visible result. The done flag prevents retry for that pending set. -- **Fix:** give each aux agent a small structured result (`status`, `findings`, `files_changed`, `validation_command`), parse crash/blocking outcomes before setting done, persist them in active-work state, and emit a visible gate-state block. - -### [HIGH] State mutation / correctness / performance — tests and docs run for nearly every source edit, not when acceptance criteria require them - -- **Evidence:** every recognized non-test source file becomes a test-writer target (`agents/base2/base2.ts:1745-1770`), and nearly every `packages/*/src`, `agents/*`, `common/src`, or `cli/src` file is classified as public API for doc-writer (`:1772-1795`). The runtime automatically spawns both before validation (`:695-771`). This conflicts with the parent policy to spawn them when coverage/docs are required or implied (`:194,205-216`). Tests explicitly lock in this broad behavior at `agents/__tests__/gate-aux-triggers.test.ts:210-359`. -- **Risk/UX:** trivial refactors can incur two extra model calls and unsolicited test/doc mutations. Internal implementation changes can rewrite public docs, increasing latency, cost, review noise, and risk of inaccurate documentation. -- **Fix:** predicate on change intent/public contract or reviewer coverage verdict; first detect existing coverage/docs gaps, then spawn only when needed. Report why each mutation agent was triggered. - -### [HIGH] Correctness / UX — doc-writer is hard-wired to the wrong documentation destination - -- **Evidence:** all selected source files are handed to doc-writer with `target_doc_files: ['docs/agents-and-tools.md']` (`agents/base2/base2.ts:743-764`), although the leaf contract says infer/read the appropriate neighboring docs (`agents/doc-writer/doc-writer.ts:21-31,51-58`). The E2E test codifies the fixed destination (`agents/e2e/gate-aux-ordering.e2e.test.ts:35-42,113-124`). -- **Risk/UX:** SDK, CLI, common, or arbitrary package API changes may be documented in an agent-system guide, while the actual package README/API guide stays stale. -- **Fix:** add deterministic path-to-doc routing (package README/docs ownership), allow multiple candidate docs, and require the agent to return “no docs change needed” without mutation. - -### [HIGH] Security / API contract — “pre-edit” security policy is implemented post-edit and has no gate verdict contract - -- **Evidence:** policy requires advisory review before the editor (`agents/base2/quality-prompt-section.ts:110-125`), while the automated spawn occurs only after `editsHappened` and after test/doc writers (`agents/base2/base2.ts:680-799`). The state is nevertheless named `preEditSecurityReviewDone` (`agents/base2/gate-state.ts:92-98`). The security reviewer returns unstructured `last_message` (`agents/security-reviewer/security-reviewer.ts:30-48`) and does not share the code-reviewer labels/schema. -- **Risk/UX:** terminology and timing mislead maintainers; security guidance arrives too late to shape implementation and cannot participate reliably in blocking/finalization. -- **Fix:** separate `advisoryPreEditSecurityReview` from `blockingPostEditSecurityReview`; give the latter the same structured verdict/crash/coverage plumbing as code-reviewer. - -### [HIGH] Error handling / UX — reviewer crash recovery advertises a bypass that the state machine cannot perform - -- **Evidence:** on crash the runtime says retry once, switch reviewer, or “proceed without the reviewer gate” (`agents/base2/base2.ts:1350-1375`), but `runReviewerGate` is always identical to `runValidationGate` (`:388-390`), no retry counter/override/alternate reviewer is stored in `Base2ActiveWorkState`, and the branch ends with `continue` (`:1396`). The next completion attempt re-enters the same reviewer spawn. -- **Risk/UX:** persistent provider/model failures can trap a task in a reviewer loop despite instructions not to loop. -- **Fix:** implement bounded retry state, configurable fallback reviewer, and an explicit user-authorized bypass carrying a durable skipped gate-state reason. - -### [HIGH] Security / git safety — secret scanning and command restrictions are prose, not enforced controls - -- **Evidence:** policy claims git-committer scans secrets automatically (`agents/base2/quality-prompt-section.ts:142-147`), but the deterministic steps only run status, full diff, log, optional `git add -A`, then unrestricted model steps (`agents/git-committer/git-committer.ts:80-106`). `run_terminal_command` remains available (`:43-50`), so “no push/config/amend/rebase” is prompt-only (`:55-62`). -- **Risk/UX:** `stage_all` can stage unrelated or secret files, and a model error can execute prohibited git operations. The user receives no structured staged-file/secret-scan proof. -- **Fix:** replace raw shell git mutation with dedicated stage/commit tools, enforce a path allowlist and secret scan before commit, reject `.env`/credential patterns mechanically, and return structured commit/staged-file verification. - -### [MEDIUM] Correctness / UX — branch creation flow contradicts dirty-tree behavior - -- **Evidence:** schema warns `git_branch` refuses a dirty tree and suggests committing first (`agents/git-committer/git-committer.ts:26-35`), but `handleSteps` always calls `git_branch` first (`:64-78`); the prompt then suggests committing existing changes on the current branch before creating the requested branch (`:55-57`). The direct test asserts branch-first behavior only (`agents/__tests__/git-committer.test.ts:137-147`). -- **Risk/UX:** the common “dirty tree + create feature branch + commit” request fails on its first action; the suggested recovery may leave the commit on the old branch rather than the requested feature branch. -- **Fix:** inspect status first, then use an explicit user choice/`allow_dirty` branch switch or create the branch without switching and switch safely; add dirty-tree E2E coverage. - -### [MEDIUM] Security / error handling — debugger executes the supplied command before assessing safety or boundedness - -- **Evidence:** `reproduce_command` is yielded immediately (`agents/debugger/debugger.ts:60-67`) before the model can inspect it, with no explicit timeout/process type. The “max 3 attempts” rule is prompt-only (`:49-58`), and the generic leaf tests only inspect tools/prompt text (`agents/__tests__/new-bundled-agents.test.ts:117-133`). -- **Risk/UX:** a mistaken destructive, interactive, watcher, or production-affecting command can run automatically; hangs/timeouts and repeated-attempt behavior are not structurally controlled. -- **Fix:** add command classification/approval, explicit timeout and non-interactive defaults, mechanically count attempts, and return structured `reproduced | not_reproduced | timed_out | unsafe_command` status. - -### [MEDIUM] Correctness / recovery gap — repeated validation failures never route to debugger - -- **Evidence:** orchestrator policy says use debugger after repeated/unclear failures (`agents/base2/base2.ts:205,214-216`), but the automated repair path runs three editor rounds and one broader editor escalation (`:1042-1077,1157-1188`) and never spawns debugger. -- **Risk/UX:** diagnosis and mutation are conflated; repeated speculative edits can consume the repair budget without producing a durable root-cause report. -- **Fix:** after the first repeated identical failure (or unparseable output), invoke debugger read-only, feed its structured diagnosis into one final editor repair, then revalidate. - -### [LOW] Performance / review quality — code-reviewer inherits the entire conversation but has weak navigation - -- **Evidence:** reviewer includes message history (`agents/reviewer/code-reviewer.ts:37-38`) while its only tool is `read_files` (`:22-30`), despite being asked to inspect closely related context (`:53`). -- **Risk/UX:** long tasks pay high context cost, while the reviewer cannot search callers/symbols unless the parent already supplied paths. -- **Fix:** pass a compact review bundle (request, changed files, validation summary) with `includeMessageHistory: false`, and grant bounded outline/reference search or precompute related files. - -## Eight-domain coverage summary - -- **Security:** high-risk gaps in discarded security verdicts, post-edit timing, debugger auto-exec, and unenforced git controls. -- **Correctness:** coarse aux triggers/doc routing, branch-order contradiction, and absent debugger handoff. -- **State mutation:** unsolicited test/doc writes; done flags are committed before successful aux completion. -- **Error handling:** code-review parsing is strong, but aux crash/timeout recovery and reviewer bypass are missing. -- **Performance:** three serial aux model calls can fire for one ordinary edit; reviewer carries full history. -- **Dependency hygiene:** no new third-party dependency issue found in this shard; the main hygiene risk is duplicating serialized inline helpers, partly guarded by parity tests. -- **Test coverage gaps:** reviewer parser/repair/order paths are well covered; missing behavioral E2E cases include aux BLOCKING/crash/timeout, leaf mutation reports, debugger timeout/unsafe command, git secret detection/dirty-tree recovery, and multi-package test-writer routing. -- **API/ABI contracts:** code-reviewer has a parseable verdict contract; security/debug/test/doc/git remain heterogeneous `last_message` contracts, preventing reliable orchestration and user-visible status. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/quality-search.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/quality-search.md deleted file mode 100644 index 612ef9b4c8..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/quality-search.md +++ /dev/null @@ -1,82 +0,0 @@ -# Quality/recovery specialists — code-searcher audit - -## Verified findings - -## [HIGH] Security / correctness / error handling — agents/base2/base2.ts:695 — Auxiliary quality-agent outcomes are discarded - -- **Risk:** Security-reviewer can report a critical exploit, or test/doc/security agents can crash, time out, or no-op, yet orchestration proceeds and the pre-set done flag prevents retry for the same pending set. -- **Fix:** Standardize structured aux results (`status`, `findings`, `files_changed`, `validation_command`), interpret them before marking done, persist blockers/crashes in active state, and render a visible lifecycle result. -- **Evidence:** test/doc/security flags are set before yields and results are never assigned (`base2.ts:695-805`); final code-reviewer output is captured and gated at `:1278-1398`. Docs promise blocking security findings prevent completion (`docs/agents-and-tools.md:27`). - -## [HIGH] Correctness / UX — agents/base2/base2.ts:743 — Doc-writer routes every source domain to one agent-system guide - -- **Risk:** SDK, CLI, common, and arbitrary package changes can generate edits in `docs/agents-and-tools.md` while their actual README/API docs remain stale, producing misleading documentation churn. -- **Fix:** Maintain deterministic package/path-to-doc ownership, allow multiple candidate targets, and support a structured “no documentation change required” result. -- **Evidence:** every selected source file is passed with `target_doc_files: ['docs/agents-and-tools.md']` at lines 743-764; the leaf itself says to inspect appropriate neighboring docs (`agents/doc-writer/doc-writer.ts:21-31,51-58`); E2E locks in the fixed destination (`agents/e2e/gate-aux-ordering.e2e.test.ts:113-124`). - -## [HIGH] Security / API contract — agents/base2/quality-prompt-section.ts:110 — “Pre-edit” security review runs only after mutations - -- **Risk:** Maintainers and users may believe risky design is reviewed before implementation, but security feedback arrives after editor/test/doc mutations and cannot reliably block because it lacks a verdict contract. -- **Fix:** Split advisory pre-edit threat modeling from blocking post-edit security review; run the former before editor for high-risk scopes and give the latter code-reviewer-compatible verdict/crash plumbing. -- **Evidence:** policy requests review before editing (`quality-prompt-section.ts:110-125`), while the automatic gate requires `editsHappened` and runs security third after test/doc writers (`base2.ts:680-799`). State still calls it `preEditSecurityReviewDone` (`gate-state.ts:92-98`), and security-reviewer returns unstructured `last_message` (`security-reviewer.ts:30-48`). - -## [HIGH] Error handling / UX — agents/base2/base2.ts:1359 — Reviewer crash guidance describes recovery paths that do not exist - -- **Risk:** Persistent reviewer/provider failure can trap the workflow in the same blocked re-entry loop while telling the operator to switch reviewer or bypass the gate. -- **Fix:** Add durable retry count, configured fallback reviewer, and explicit user-authorized bypass with a skipped gate-state reason and audit trail. -- **Evidence:** crash text says retry, switch, or proceed without the gate at lines 1359-1375, but the branch only `continue`s at 1396; `runReviewerGate` mirrors `runValidationGate` (`base2.ts:388-390`), and active state has no override/fallback selection. - -## [HIGH] Security / state mutation — agents/git-committer/git-committer.ts:43 — Git safety and secret prevention are prompt-only - -- **Risk:** `stage_all` can include unrelated/secret files and unrestricted shell access can push, amend, rebase, alter config, or commit credentials despite the prose prohibition; output provides no machine-verifiable scan/staged-file proof. -- **Fix:** Replace raw mutation commands with dedicated stage/commit tools, enforce path scope and secret scanning, mechanically deny prohibited operations, and return a structured receipt containing staged files, scan result, commit hash, and subject. -- **Evidence:** tools include unrestricted `run_terminal_command` at lines 43-50; deterministic flow can execute `git add -A` at 97-103 then grants `STEP_ALL`; “do not push/commit secrets/amend/rebase” exists only in prompt lines 55-62. The shared policy claims automatic secret scanning (`quality-prompt-section.ts:142-147`) without an implementation step. - -## [MEDIUM] Correctness / UX — agents/git-committer/git-committer.ts:64 — Branch creation contradicts dirty-tree recovery guidance - -- **Risk:** The common “create feature branch and commit current dirty work” request fails immediately; suggested recovery commits on the old branch, contrary to user intent. -- **Fix:** Inspect status first, then explicitly choose safe dirty-work transfer semantics (create/switch with authorized dirty carry, stash, or user choice) before committing. -- **Evidence:** schema notes `git_branch` refuses dirty trees (`:26-35`), yet `handleSteps` invokes it before status/diff at `:64-85`; prompt simultaneously says commit existing changes first on the current branch (`:55-57`). Tests assert branch-first ordering but not dirty-tree behavior. - -## [MEDIUM] Security / error handling — agents/debugger/debugger.ts:60 — Debugger auto-executes an unclassified reproduce command - -- **Risk:** A destructive, interactive, watcher, network, or production-affecting command runs before the model can inspect safety; no explicit timeout/process type or structural attempt limit exists. -- **Fix:** Apply the runtime command approval classifier, default to bounded synchronous/non-interactive execution, validate command shape, and mechanically enforce attempt count with typed outcomes. -- **Evidence:** `reproduce_command` is yielded immediately at lines 60-67 with only `{command}`; “no more than 3 attempts” is prompt-only at 49-58. Leaf tests cover prompt/tool presence, not unsafe command/timeout behavior. - -## [MEDIUM] Correctness / recovery UX — agents/base2/base2.ts:1042 — Repeated validation failures never invoke the debugger specialist - -- **Risk:** Diagnosis and mutation remain conflated: three targeted editor attempts plus another broad editor can repeatedly guess at symptoms without producing an evidence-backed root cause. -- **Fix:** Detect repeated identical/unparseable failures, spawn debugger read-only, persist its diagnosis, then provide that evidence to one bounded final editor repair. -- **Evidence:** automated path performs editor repairs at lines 1042-1155 and a broader editor escalation at 1157-1200; no debugger spawn occurs. Orchestrator/docs prescribe debugger after repeated or unclear failures (`docs/agents-and-tools.md:26`, `base-deep.ts:55,63`). - -## [MEDIUM] Performance / state mutation — agents/base2/base2.ts:1745 — Broad aux predicates trigger unsolicited test/doc mutations - -- **Risk:** Routine internal refactors incur serial model calls and code/doc churn even when acceptance criteria do not require new tests or public documentation. -- **Fix:** Base triggers on behavior/public-contract changes, detected coverage/docs gaps, or explicit acceptance criteria; preview trigger reasons and allow scoped opt-out. -- **Evidence:** any recognized non-test source selects test-writer (`:1745-1769`), while nearly all source roots select doc-writer (`:1772-1795`). This conflicts with policy describing these agents for required/implied coverage (`docs/agents-and-tools.md:28`). - -## [LOW] Performance / review quality — agents/reviewer/code-reviewer.ts:22 — Reviewer receives full history but lacks search/navigation - -- **Risk:** Long conversations increase reviewer cost while the only tool cannot efficiently discover callers/references beyond parent-supplied paths, weakening cross-file review. -- **Fix:** Pass a compact review bundle with `includeMessageHistory: false`, plus bounded outline/reference search or precomputed related files. -- **Evidence:** reviewer sets `includeMessageHistory: true` at line 38 but permits only `read_files` at 22-29, while requiring closely related context at 53. - -## Rejected / downgraded candidates - -- **Final code-reviewer has no enforceable contract — rejected.** It has explicit text/JSON verdict labels, blocker/coverage parsing, crash detection, durable fingerprints, and strong parser/lifecycle tests. Heterogeneity is confined to the other quality agents. -- **Aux quality agents race each other — rejected.** Inline spawning blocks and order is tested; the problem is serial cost and ignored results, not a race. -- **Test-writer falsely claims it validates tests — rejected.** Current docs explicitly state it reports a parent-owned command and does not execute validation (`docs/agents-and-tools.md:140-155`). -- **Git branch tool itself cannot support dirty switching — downgraded.** SDK includes an `allow_dirty` capability, but git-committer’s schema/flow neither exposes nor uses it safely; the finding is agent UX/ordering. -- **Dependency hygiene issue — not established.** No undeclared/vulnerable package was evidenced; serialized helper duplication remains a maintenance concern covered partially by parity tests. - -## Coverage across 8 domains - -- Security: ignored security verdicts, misleading timing, prompt-only git controls, debugger command auto-execution. -- Correctness: doc routing, impossible reviewer recovery guidance, branch ordering, absent debugger handoff. -- State mutation: aux done-before-success, unsolicited test/doc writes, unsafe stage-all semantics. -- Error handling: aux and reviewer crash recovery gaps; debugger lacks typed timeout/unsafe outcomes. -- Performance: broad serial aux calls and full-history reviewer context. -- Dependency hygiene: no concrete third-party issue found. -- Test coverage: strong final-review parser/order coverage; missing aux blocker/crash/timeout, dirty branch, secret scan, debugger safety, and repeated-failure diagnosis tests. -- API/ABI contract: only code-reviewer has a reliable verdict; other quality agents use heterogeneous last-message outputs that orchestration cannot safely consume. diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-picker.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-picker.md deleted file mode 100644 index ff3e46f17b..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-picker.md +++ /dev/null @@ -1,128 +0,0 @@ -# Runtime/contracts/routing picker audit - -## Flow inventory for paired code-searcher - -- Registration: `sdk/src/agents/load-agents.ts:106-139,201-318` recursively imports project/parent/home `.agents`, resolves MCP env references, keys definitions by ID, and optionally validates; `packages/agent-runtime/src/templates/agent-registry.ts:21-87,93-110` resolves local templates first and has a disabled-in-local-mode database fallback. -- Routing/BYOK: `packages/agent-runtime/src/main-prompt.ts:93-151` chooses explicit agent ID or legacy cost-mode alias; `packages/agent-runtime/src/run-agent-step.ts:556-576` sends the stable agent type to SDK routing; `sdk/src/impl/agent-runtime.ts:64-107,129-137` wires local provider calls and disables remote registry/analytics/billing. -- Execution/output: `packages/agent-runtime/src/run-agent-step.ts:801-1398` builds prompts/tools, runs programmatic and LLM steps, prunes context, checkpoints, and returns `getAgentOutput`; `run-programmatic-step.ts:241-320,385-592` materializes generators, executes yields, and records errors/output. -- Permissions: spawn authorization is in `spawn-agent-utils.ts:176-249`; input validation is `:251-369`; filesystem path/commit policy is `sdk/src/tools/filesystem-authority.ts:165-243`; terminal containment is `sdk/src/tools/run-terminal-command.ts:141-161`. -- Spawn lifecycle: foreground/background dispatch is `spawn-agents.ts:95-393`; inline shared-history dispatch is `spawn-agent-inline.ts:97-200`; depth/timeout/cancellation is `spawn-agent-utils.ts:460-629`; background storage is `util/background-agent-jobs.ts:76-222`. -- Telemetry/cost: turn cost resets/finish event are `main-prompt.ts:182-190,245-265`; per-step cost/cache accumulation and budget checks are `run-agent-step.ts:367-383,440-454,705-755`; foreground subagent cost aggregation is `spawn-agents.ts:349-390`. -- Key tests present: main prompt, programmatic steps, spawn permission/nesting/history/images/depth/timeout, budgets/context pruning, agent loading/validation, terminal containment, code-search parsing/limits/abort, and filesystem authority. Important missing cases are called out below. - -## Findings - -## [HIGH] Security — `sdk/src/agents/load-agents.ts:73-79,243-255`; `packages/agent-runtime/src/run-agent-step.ts:456-474`; `packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:448-456` — resolved MCP secrets are logged inside whole agent templates - -- **Risk:** `$TOKEN` references are replaced with plaintext environment values, after which debug logs serialize the entire `agentTemplate`; local logs can therefore contain MCP/API credentials. -- **Fix:** retain secret references until process launch or redact all `mcpServers.*.env` values before logging templates; never log the full template object. -- **Evidence:** `resolveAgentMcpEnv(processedAgentDefinition)` mutates config, while both step and spawn logs include `agentTemplate` verbatim. - -## [HIGH] Security / API contract — `packages/agent-runtime/src/run-programmatic-step.ts:700-717,750-759` — programmatic agents bypass declared tool capabilities - -- **Risk:** a `handleSteps` agent may execute terminal, edit, spawn, or network tools even when absent from its `toolNames`, defeating the agent permission model and making a nominally read-only subagent write-capable. -- **Fix:** enforce membership in the effective allowed tool set, with an explicit narrowly-scoped privileged capability for trusted orchestrators instead of a blanket bypass. -- **Evidence:** the availability check is commented out with “You can run any tool from handleSteps now!”, then `executeToolCall` receives the arbitrary yielded name. - -## [HIGH] Security — `packages/agent-runtime/src/tools/handlers/tool/web-search.ts:67-105`; `common/src/tools/params/tool/web-search.ts:17-23` — direct URL fetch is an SSRF primitive - -- **Risk:** any syntactically valid URL is fetched with redirects and no protocol/private-address guard, allowing prompt-injected agents to probe localhost, LAN services, or cloud metadata endpoints. `response.text()` also buffers the complete body before the 50k-character truncation. -- **Fix:** allow only HTTP(S), resolve and block loopback/link-local/private/reserved addresses on every redirect, and stream with a hard byte cap. -- **Evidence:** schema uses only `z.string().url()` and handler calls `fetch(fetchUrl)` directly, then awaits the unbounded `response.text()`. - -## [HIGH] Correctness / state mutation — `packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts:95-203` — mixed background batches can create unreachable orphan work - -- **Risk:** background entries are validated and launched sequentially before foreground processing. If a later entry is invalid, the handler throws after earlier jobs started but before returning their job IDs, leaving detached work the parent cannot poll or cancel. -- **Fix:** pre-validate every batch entry before launching any, then use per-entry `allSettled` reporting and always return IDs for successfully-started jobs. -- **Evidence:** `backgroundReports` is returned only at the end, while each detached promise is attached during the initial `for` loop and validation errors are not caught per entry. - -## [HIGH] State mutation / performance / information exposure — `packages/agent-runtime/src/util/background-agent-jobs.ts:76-76,161-170,195-215`; adjacent `tools/handlers/tool/check-background-agent.ts:61-75,93-109` — completed background jobs retain and return full internal run state forever - -- **Risk:** the process-wide map has no deletion/TTL/cap; each settled result is the entire `executeSubagent` result, including `agentState` message history, system prompt, tool definitions, and proposal state. This grows memory across a session and injects oversized/internal prompt state back into the parent when polled. -- **Fix:** store a normalized result `{ output, cost, agentId, status }`, redact internal state, add TTL/LRU cleanup plus explicit consume/delete, and cap total jobs. -- **Evidence:** `job.result = result`, `jobs` is never pruned, and polling returns `result: job.result` verbatim. - -## [HIGH] Correctness / cancellation — `sdk/src/tools/run-terminal-command.ts:248-279,288-301` — timeout/abort can leave the command running after the tool has rejected - -- **Risk:** timeout sends only SIGTERM and marks `processFinished` before process exit, with no SIGKILL fallback. Abort's fallback checks `childProcess.killed`, which means a signal was sent, not that the process exited, so SIGKILL commonly never fires. Long-lived grandchildren may survive while the agent believes execution stopped. -- **Fix:** wait for `close`, escalate after a grace period based on observed exit, kill the process group where supported, and add stubborn-process timeout/abort tests. -- **Evidence:** timeout immediately rejects after `kill('SIGTERM')`; abort fallback is gated by `if (!childProcess.killed)`. - -## [HIGH] Correctness / BYOK compatibility — `packages/agent-runtime/src/run-agent-step.ts:835-838,1203-1236`; `util/context-pruning.ts:23-76` — runtime pruning/status is not connected to the resolved model context window - -- **Risk:** `maxContextLength` has no caller in the audited tree, so runtime semantic pruning and the CLI `context_window.max` default to 190k even for BYOK models with 8k/32k windows. The SDK emergency trim may save the request, but the agent prunes too late and the UX reports a false capacity. -- **Fix:** return resolved `contextWindowTokens` from routing before the loop or inject a model-capability resolver, then use `getModelContextMessageLimit` consistently for pruning and status. -- **Evidence:** the loop accepts an optional value and otherwise emits `DEFAULT_MAX_CONTEXT_TOKENS`; repository search found no value passed into this parameter. - -## [MEDIUM] Error handling / UX contract — `packages/agent-runtime/src/main-prompt.ts:200-241` — invalid agent configuration emits an error and then continues the run - -- **Risk:** the client can receive `prompt-error`, followed by `start`, streamed content, `finish`, and `prompt-response` for the same prompt. This creates ambiguous terminal state and may run a different surviving agent despite a broken configuration. -- **Fix:** either fail closed before `start`, or downgrade unrelated invalid definitions to a distinct non-terminal warning event with file paths and continue explicitly. -- **Evidence:** after sending `prompt-error` for non-empty `validationErrors`, execution unconditionally sends `start` and calls `mainPrompt`. - -## [MEDIUM] Correctness / UX — `sdk/src/agents/load-agents.ts:135-139,212-263`; `sdk/src/__tests__/load-agents.test.ts:694-731` — global agents silently override project agents and duplicates disappear before validation - -- **Risk:** directories are loaded project → parent → home but assignment is last-wins, so a home definition silently replaces the project-local definition. Duplicate IDs are collapsed before validation, and the existing test explicitly makes no duplicate assertion. -- **Fix:** define and test precedence (normally project > parent > home), retain provenance for every candidate, and report duplicate/override diagnostics. -- **Evidence:** `agents[id] = processedAgentDefinition` overwrites prior entries; the duplicate-ID test only asserts that a result exists. - -## [MEDIUM] API/ABI contract — `common/src/types/agent-template.ts:148-156`; `common/src/types/dynamic-agent-template.ts:175-217`; `packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts:527-533`; `common/src/constants/agents.ts:122-130` — advertised spawn-depth configuration is not loadable - -- **Risk:** runtime/type/docs tell users to configure `maxSpawnDepth` on an agent template or in `openbuff.json`, but the dynamic agent schema and provider config do not expose that field, so user configuration is stripped/ignored. -- **Fix:** add validated global and per-agent schema fields with precedence tests, or remove the unsupported configuration claims. -- **Evidence:** `AgentTemplate` reads the field and the error recommends it, while `DynamicAgentDefinitionSchema` omits it and no provider-config route exists. - -## [MEDIUM] Performance — `sdk/src/tools/code-search.ts:98-113,326-397`; `sdk/src/tools/find-files-matching-content.ts:587-665` — context flags can bypass code-search memory/output guards - -- **Risk:** `-A/-B/-C` accept arbitrary values; context events are always appended and do not trigger global/output-size stopping because the checks run only for match events. One match plus huge context can accumulate a file-sized in-memory array before final truncation. -- **Fix:** bound context counts, include context bytes/events in hard limits, and stop immediately when estimated output reaches the cap. -- **Evidence:** `shouldInclude = !isMatch || ...`, but limit checks are nested under `if (isMatch)`; flag parsing validates presence, not numeric range. - -## [MEDIUM] Security / UX permissions — `common/src/tools/params/tool/run-terminal-command.ts:76-103`; `packages/agent-runtime/src/tools/handlers/tool/run-terminal-command.ts:20-32`; `sdk/src/run.ts:1082-1091` — terminal approval is prompt guidance, not an enforceable capability - -- **Risk:** any agent with the tool (and every programmatic agent via the bypass above) can run arbitrary `bash -c` commands with the process environment; there is no typed approval token or side-effect policy between the model call and SDK execution. -- **Fix:** add command risk classification plus explicit user approval/capability receipts for destructive, network, credential, install, git-history, and out-of-project effects; apply it equally to subagents/programmatic calls. -- **Evidence:** the schema merely tells the model to ask, while the handler forwards directly and SDK executes directly after cwd containment. - -## [MEDIUM] Security / error handling — `packages/agent-runtime/src/run-agent-step.ts:1462-1497` — internal stack traces are included in user-visible agent output - -- **Risk:** non-HTTP failures expose local absolute paths and implementation details in CLI/error history, and may feed those internals into later model context. -- **Fix:** log stack traces only to the logger; return a stable sanitized code/message and an optional correlation ID. -- **Evidence:** `fallbackMessage` appends `error.stack` when no HTTP status, then returns it as `output.message`. - -## [MEDIUM] Telemetry / UX — `packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts:349-390`; `packages/agent-runtime/src/main-prompt.ts:245-253`; `packages/agent-runtime/src/tools/handlers/tool/end-turn.ts:1-19` — background-agent cost and lifecycle are absent from turn totals/end-turn warnings - -- **Risk:** only foreground child cost is added to the parent and `finish.totalCost`; shell jobs are warned about at `end_turn`, but the separate background-agent registry is not. Users can end a turn unaware of running agents and see understated spend. -- **Fix:** unify pending-job accounting, surface running background agents at end-turn/session exit, add kill/cancel, and aggregate settled cost into session telemetry without double counting. -- **Evidence:** foreground-only aggregation is explicit; background results remain separate and end-turn queries only `pending-background-jobs`. - -## [LOW] Correctness / state mutation — `packages/agent-runtime/src/run-programmatic-step.ts:244-263` — run-ID collision resumes the wrong generator after only a warning - -- **Risk:** if a dependency/test/custom run allocator returns a duplicate ID, one agent continues another agent's generator and mutable workflow state. -- **Fix:** fail closed, clear the conflicting registry entry, and mark both runs failed; do not continue after detecting mismatched owners. -- **Evidence:** owner mismatch logs `logger.warn` but leaves `generator` intact. - -## [LOW] Dependency hygiene / cancellation — `packages/agent-runtime/src/tools/handlers/tool/web-search-utils.ts:24-58`; `web-search.ts:75-81,144-145` — web-search timeout does not cancel underlying work and relies on an internal package subpath - -- **Risk:** `Promise.race` times out without aborting DuckDuckGo work, and `open-websearch/build/engines/...` is a private deep import likely to break on upstream layout changes. URL fetch also uses only an independent timeout signal, not the user/run abort signal. -- **Fix:** use a public package API with AbortSignal support and combine run cancellation with timeout. -- **Evidence:** the timer rejects separately while `searchDuckDuckGo` receives no signal; import targets `build/engines/duckduckgo/index.js`. - -## [LOW] Correctness / cache UX — `packages/agent-runtime/src/run-agent-step.ts:984-1019`; `packages/agent-runtime/src/main-prompt.ts:118-129` — system prompt cache invalidates only on agent-type changes - -- **Risk:** file tree, routed knowledge, patterns, and system/git context can change during a session but the cached system prompt remains byte-stable, so later turns may reason from stale project metadata. -- **Fix:** cache by a fingerprint of prompt-affecting file context/config, preserving provider cache hits while invalidating on relevant changes. -- **Evidence:** the only explicit invalidation clears `systemPrompt` on agent type change; otherwise the prior system prompt is reused. - -## Eight-domain coverage - -| Domain | Result | -| ------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Security | Secret logging, programmatic capability bypass, SSRF, prompt-only terminal permissions, stack leakage. | -| Correctness | Partial background launches, process-reaping bugs, model-window mismatch, config/event ambiguity, precedence/depth contracts. | -| State mutation | Unbounded background registry/full-state retention, orphan jobs/processes, collision continuation, stale cache. | -| Error handling | Error-then-success event sequence, internal stacks, partial batch failure, cancellation misclassification risks. | -| Performance | Retained job state, unbounded web body, context-line accumulation, uncancelled web work. | -| Dependency hygiene | Fragile `open-websearch/build/...` deep import; no other concrete dependency defect established in this shard. | -| Test coverage gaps | No web-search handler security/cancellation tests; no mixed background failure/cleanup/redaction tests; no stubborn terminal process tests; no programmatic tool-permission test; duplicate-agent test has no duplicate assertion; no model-specific runtime context test. | -| API/ABI contracts | Unsupported `maxSpawnDepth` config, silent agent precedence, background result shape exposes internals, invalid-config terminal event ambiguity. | diff --git a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-search.md b/.agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-search.md deleted file mode 100644 index d4a25a6a77..0000000000 --- a/.agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-search.md +++ /dev/null @@ -1,88 +0,0 @@ -# Runtime/contracts/routing — code-searcher audit - -## Verified findings - -## [HIGH] Security — sdk/src/agents/load-agents.ts:243 — Resolved MCP secrets can be logged inside agent templates - -- **Risk:** Environment references become plaintext credentials and full-template debug logging can persist them locally. -- **Fix:** Keep secret references opaque until process launch and centrally redact `mcpServers.*.env` and provider secrets from every log serializer. -- **Evidence:** local loading resolves MCP env values (`load-agents.ts:73-79,243-255`), while runtime step/spawn debug records include the whole `agentTemplate` (`run-agent-step.ts:456-474`, `spawn-agent-utils.ts:448-456`). - -## [HIGH] Security / permissions contract — packages/agent-runtime/src/run-programmatic-step.ts:700 — `handleSteps` bypasses declared tool capabilities - -- **Risk:** A template presented as read-only can yield terminal, edit, spawn, or network tools omitted from `toolNames`, undermining template permissions and docs claiming secure sandbox/tool declarations. -- **Fix:** Enforce the effective allowed-tool set for programmatic yields, with explicit narrowly scoped internal capabilities for trusted orchestrators. -- **Evidence:** the availability check is commented out with “You can run any tool from handleSteps now!” (`run-programmatic-step.ts:700-717`), then arbitrary calls reach execution (`:750-759`); docs state templates define tool permissions (`docs/agents-and-tools.md:11`). - -## [HIGH] Security / performance — packages/agent-runtime/src/tools/handlers/tool/web-search.ts:67 — Direct URL search enables SSRF and unbounded body buffering - -- **Risk:** Agents can fetch localhost, LAN, link-local/cloud-metadata, or redirected private endpoints; large responses are fully buffered before character truncation. -- **Fix:** Restrict to HTTP(S), resolve/block private/reserved addresses on every redirect, stream with a hard byte cap, and combine timeout with run cancellation. -- **Evidence:** schema only requires `z.string().url()` (`common/src/tools/params/tool/web-search.ts:17-23`); handler directly `fetch`es and awaits `response.text()` (`web-search.ts:67-105`). - -## [HIGH] Correctness / state mutation — packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts:95 — Mixed background batches can launch unreachable orphan agents - -- **Risk:** Earlier background entries start before a later invalid entry throws; because results/job IDs are returned only after batch processing, the parent cannot poll or cancel already-launched work. -- **Fix:** Pre-validate the full batch atomically, then launch; return per-entry settled reports and always expose IDs for successful starts. -- **Evidence:** background validation/launch happens sequentially in the initial loop (`spawn-agents.ts:95-203`), while `backgroundReports` is returned only at the end and validation failures are not isolated per entry. - -## [HIGH] State mutation / information exposure / performance — packages/agent-runtime/src/util/background-agent-jobs.ts:76 — Background agent registry retains full internal run state indefinitely - -- **Risk:** Settled jobs accumulate message history, prompts, tool definitions, and proposal state in a process-wide map, then polling injects that oversized internal state back into the parent. -- **Fix:** Store a normalized/redacted result, add consume/delete plus TTL/LRU limits, cap registry size, and expose explicit lifecycle cleanup. -- **Evidence:** `job.result = result` (`background-agent-jobs.ts:161-170`), no pruning exists (`:76,195-215`), and `check-background-agent.ts:61-109` returns the result verbatim. - -## [HIGH] Correctness / cancellation — sdk/src/tools/run-terminal-command.ts:248 — Timeout/abort does not reliably terminate the command tree - -- **Risk:** The tool rejects while a stubborn shell or descendants continue running; abort fallback tests `childProcess.killed`, which only means a signal was sent, not that exit occurred. -- **Fix:** own a process group, wait for observed close, escalate after grace based on actual exit, and test stubborn children/grandchildren. -- **Evidence:** timeout sends SIGTERM then immediately rejects without fallback (`:288-301`); abort escalation is gated by `if (!childProcess.killed)` (`:248-279`). - -## [HIGH] Correctness / BYOK UX — packages/agent-runtime/src/run-agent-step.ts:835 — Runtime context status/pruning is disconnected from resolved model capacity - -- **Risk:** BYOK 8k/32k models can be shown a 190k capacity and pruned too late; SDK emergency trimming may prevent failure but runtime behavior/status remains misleading. -- **Fix:** propagate resolved `contextWindowTokens` from model routing into the agent loop and use one capability source for semantic pruning, budgets, and CLI status. -- **Evidence:** runtime falls back to `DEFAULT_MAX_CONTEXT_TOKENS` when `maxContextLength` is absent (`run-agent-step.ts:835-838,1203-1236`; `util/context-pruning.ts:23-76`); routing already resolves optional context capacity (`sdk/src/impl/model-provider.ts:152-154`), but no audited caller connects it. - -## [MEDIUM] Error handling / UX contract — packages/agent-runtime/src/main-prompt.ts:200 — Invalid local-agent definitions emit an error and then continue - -- **Risk:** One prompt can produce `prompt-error` followed by start/content/finish events, making terminal state ambiguous and potentially running a surviving/fallback agent despite broken configuration. -- **Fix:** Fail closed before `start` for the selected agent, or emit non-terminal per-file warnings for unrelated invalid definitions with explicit continuation semantics. -- **Evidence:** validation errors trigger `prompt-error` at `main-prompt.ts:200-241`, after which execution unconditionally emits `start` and calls `mainPrompt`. - -## [MEDIUM] Correctness / discoverability — sdk/src/agents/load-agents.ts:135 — Home agents silently override project agents - -- **Risk:** A global definition can unexpectedly replace repository-local behavior; duplicate IDs disappear before validation, leaving users unable to diagnose provenance/shadowing. -- **Fix:** Define project-over-parent-over-home precedence, retain all candidate provenance, and surface override diagnostics in the agent picker/startup report. -- **Evidence:** directories are processed project → parent → home and assigned with last-wins `agents[id] = ...` (`load-agents.ts:135-139,212-263`); duplicate test only asserts a surviving result (`sdk/src/__tests__/load-agents.test.ts:694-731`). - -## [MEDIUM] API/ABI contract — common/src/types/agent-template.ts:148 — `maxSpawnDepth` is advertised but not loadable from dynamic config - -- **Risk:** Users are told to configure a depth value that the local-agent/provider schema omits, so the setting is stripped or ignored. -- **Fix:** Add validated global/per-agent fields with precedence tests, or remove the unsupported configuration guidance. -- **Evidence:** runtime/template types read and recommend `maxSpawnDepth` (`agent-template.ts:148-156`, `spawn-agent-utils.ts:527-533`), while `DynamicAgentDefinitionSchema` omits it (`dynamic-agent-template.ts:175-217`) and provider config has no corresponding route. - -## [MEDIUM] Telemetry / UX — packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts:349 — Background agent cost and lifecycle are missing from completion accounting - -- **Risk:** Users can end a turn unaware of running background agents and see understated spend; shell jobs receive end-turn warnings but agent jobs do not. -- **Fix:** Unify pending job accounting, surface background agents at end/session exit, add cancel actions, and aggregate settled cost without double-counting. -- **Evidence:** only foreground child costs are accumulated (`spawn-agents.ts:349-390`, `main-prompt.ts:245-253`); `end-turn.ts` queries shell pending jobs, not the background-agent registry. - -## Rejected / downgraded candidates - -- **Filesystem/terminal cwd containment is absent — rejected.** SDK applies lexical plus realpath/symlink containment and has focused tests; command side effects still require a separate approval policy. -- **Structured-output agents must expose `set_output` to the model — rejected.** The schema intentionally permits programmatic `handleSteps` to yield `set_output` without granting it to the LLM, with tests documenting the compatibility behavior. -- **BYOK model routing is hosted/fallback dependent — rejected.** Local provider resolution is implemented and explicit failures guide users to `openbuff.json`; the verified gap is propagation of model capabilities into runtime pruning/status. -- **All spawn cancellation is missing — rejected.** Foreground spawn timeout/depth/cancellation tests exist. Gaps are mixed partial launch, background ownership/cleanup, and terminal process reaping. -- **Dependency hygiene broadly broken — rejected.** One fragile deep import in web search remains (`open-websearch/build/...`), but no broader undeclared dependency issue was established. - -## Coverage across 8 domains - -- Security: secret logging, programmatic capability bypass, SSRF, command-process cleanup exposure. -- Correctness: partial background launch, context-window mismatch, invalid-config event ambiguity, precedence/depth contract drift. -- State mutation: retained background state, orphan agents/processes, missing lifecycle cleanup. -- Error handling: error-then-continue configuration flow and non-atomic batch failures. -- Performance: full background state retention, unbounded web body, incorrect pruning threshold. -- Dependency hygiene: fragile web-search deep import only; no broader issue proven. -- Test coverage: missing SSRF/redirect/body-cap, programmatic permission, mixed-batch orphan, registry cleanup/redaction, stubborn process, precedence, and model-capability integration cases. -- API/ABI contract: unsupported spawn-depth config, silent override precedence, oversized background result shape, ambiguous validation event semantics. diff --git a/.agents/sessions/audit-cli-2026-07-10/MAP.md b/.agents/sessions/audit-cli-2026-07-10/MAP.md deleted file mode 100644 index 94dec9098a..0000000000 --- a/.agents/sessions/audit-cli-2026-07-10/MAP.md +++ /dev/null @@ -1,141 +0,0 @@ -# Structural Map — cli - -- **Project root:** `/home/ben/Code/CLI/openbuff/cli` -- **Built at:** 2026-07-10T09:15:50.758Z -- **Total files indexed:** 421 -- **Graph:** 5521 nodes, 22480 edges - -> Pin this file in context. Every audit shard navigates from here instead of doing fuzzy round-trip discovery. - -## Entry points - -- `src/app.tsx` -- `src/index.tsx` - -## Directories (by size, biggest first) - -| dir | files | total size | top symbols | -| ------------------- | ----- | ---------- | ---------------------------------------------------------------------------------------------------------- | -| `src` | 398 | 2.6 MB | render, TestItem, tmux, createErrorMessage, updateBlocksRecursively, parseHistoryItem | -| `release` | 5 | 30.4 KB | resetTerminal, createConfig, getPostHogConfig, trackUpdateFailed, getLatestVersion, getLocalPackageVersion | -| `scripts` | 6 | 30.2 KB | main, log, AgentDefinition, getAllTsFiles, loadAgentDefinition, generateBundledAgentsFile | -| `knowledge.md` | 1 | 29.6 KB | — | -| `release-staging` | 5 | 29.4 KB | resetTerminal, createConfig, getPostHogConfig, trackUpdateFailed, getLatestVersion, getLocalPackageVersion | -| `tmux.knowledge.md` | 1 | 9.6 KB | — | -| `CHANGELOG.md` | 1 | 5.0 KB | — | -| `package.json` | 1 | 2.3 KB | — | -| `README.md` | 1 | 1.4 KB | — | -| `tsconfig.json` | 1 | 535 B | — | -| `.gitignore` | 1 | 103 B | — | - -## Largest files per directory - -### `src` - -- `src/hooks/helpers/__tests__/send-message.test.ts` — 55.5 KB, 0 symbols -- `src/utils/__tests__/message-block-helpers.test.ts` — 55.1 KB, 0 symbols -- `src/chat.tsx` — 53.8 KB, 1 symbols -- `src/utils/__tests__/send-message-helpers.test.ts` — 48.1 KB, 0 symbols -- `src/utils/__tests__/collapse-helpers.test.ts` — 43.1 KB, 0 symbols - -### `release` - -- `release/index.js` — 18.6 KB, 19 symbols -- `release/README.md` — 4.9 KB, 0 symbols -- `release/http.js` — 4.5 KB, 8 symbols -- `release/package.json` — 1.4 KB, 0 symbols -- `release/postinstall.js` — 929 B, 0 symbols - -### `scripts` - -- `scripts/build-binary.ts` — 11.9 KB, 8 symbols -- `scripts/smoke-binary.ts` — 6.8 KB, 2 symbols -- `scripts/prebuild-agents.ts` — 5.5 KB, 5 symbols -- `scripts/release.ts` — 3.0 KB, 6 symbols -- `scripts/test-sdk-file-hooks.sh` — 1.9 KB, 0 symbols - -### `knowledge.md` - -- `knowledge.md` — 29.6 KB, 0 symbols - -### `release-staging` - -- `release-staging/index.js` — 17.3 KB, 19 symbols -- `release-staging/README.md` — 5.3 KB, 0 symbols -- `release-staging/http.js` — 4.5 KB, 8 symbols -- `release-staging/package.json` — 1.4 KB, 0 symbols -- `release-staging/postinstall.js` — 901 B, 0 symbols - -### `tmux.knowledge.md` - -- `tmux.knowledge.md` — 9.6 KB, 0 symbols - -### `CHANGELOG.md` - -- `CHANGELOG.md` — 5.0 KB, 0 symbols - -### `package.json` - -- `package.json` — 2.3 KB, 0 symbols - -### `README.md` - -- `README.md` — 1.4 KB, 0 symbols - -### `tsconfig.json` - -- `tsconfig.json` — 535 B, 0 symbols - -### `.gitignore` - -- `.gitignore` — 103 B, 0 symbols - -## Most-imported files (likely key modules) - -| in-degree | file | -| --------- | ---------------------------------------------- | -| 86 | `src/__tests__/release/proxy-http-get.test.ts` | -| 50 | `src/utils/message-block-helpers.ts` | -| 42 | `src/hooks/use-theme.tsx` | -| 40 | `src/utils/arrays.ts` | -| 36 | `src/project-files.ts` | -| 28 | `scripts/release.ts` | -| 24 | `src/utils/implementor-helpers.ts` | -| 23 | `src/utils/env.ts` | -| 22 | `src/state/chat-store.ts` | -| 22 | `src/utils/openbuff-provider.ts` | -| 19 | `src/utils/terminal-enter-detection.ts` | -| 19 | `src/components/tools/types.ts` | -| 18 | `src/utils/theme-system.ts` | -| 18 | `src/__tests__/test-utils.ts` | -| 18 | `release/index.js` | -| 17 | `src/utils/text-layout.ts` | -| 16 | `src/utils/strings.ts` | -| 16 | `src/hooks/use-terminal-layout.ts` | -| 15 | `src/utils/block-operations.ts` | -| 15 | `src/utils/pending-attachments.ts` | -| 15 | `src/utils/analytics.ts` | -| 15 | `src/utils/local-agent-registry.ts` | -| 15 | `release-staging/http.js` | -| 14 | `scripts/build-binary.ts` | -| 14 | `src/hooks/helpers/send-message.ts` | - -## Cross-directory dependencies (architectural layering) - -| count | from → to | -| ----- | ----------------------------- | -| 31 | `src` → `scripts` | -| 18 | `release-staging` → `release` | -| 9 | `release-staging` → `src` | -| 9 | `release` → `src` | -| 9 | `release` → `release-staging` | -| 5 | `scripts` → `src` | -| 3 | `release-staging` → `scripts` | -| 3 | `src` → `release-staging` | -| 2 | `release` → `scripts` | - -## Shard sizing hint - -Total indexed source: **2.7 MB** across **11** top-level directories. - -When sharding for an audit, aim for ~5–15 files per shard. Use the table above to group small dirs together and split huge dirs (e.g. split `src/` by subdirectory). diff --git a/.agents/sessions/audit-cli-next-level-2026-07/AUDIT-REPORT.md b/.agents/sessions/audit-cli-next-level-2026-07/AUDIT-REPORT.md deleted file mode 100644 index 8ecbb87f4b..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/AUDIT-REPORT.md +++ /dev/null @@ -1,222 +0,0 @@ -# Openbuff CLI Independent Audit Report - -## Executive verdict - -Openbuff's current CLI is **feature-rich and technically ambitious, but not release-ready at the audited worktree state**. It already offers a genuinely strong terminal agent experience: broad provider support, sophisticated keyboard and attachment flows, rich tool/subagent progress, thoughtful terminal cleanup, multi-platform packaging, and a large focused test surface. The central weakness is not lack of capability; it is that several trust boundaries are less mature than the feature surface they protect. Configuration edits can overwrite unrelated settings, project histories can collide, cancellation can fork model history from visible history, project switching can retain the wrong project context, and update/release paths can interrupt sessions or expose publication authority. - -The CLI therefore feels closer to a strong advanced preview than a dependable general release. A user can accomplish a great deal, and the normal path is polished enough to impress, but several uncommon-looking paths are actually ordinary operations: changing a route, opening two same-named repositories, cancelling and immediately continuing, selecting a project, or receiving an update. Those paths need to become transactional and test-gated before the product can make a high-confidence local-first/BYOK reliability claim. - -### Scorecard — inference, not a measured benchmark - -| Area | Score | Inferred assessment | -| ----------------------------------- | ---------: | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Capability breadth | **8.5/10** | Commands, palettes, prompt history, attachments/images, providers, OAuth, local agents/skills, queueing, checkpoints, tool and subagent rendering, and broad platform packaging form an unusually complete CLI surface. | -| Interaction UX | **7.0/10** | Centralized keyboard classification and rich input behavior are strong, but file search truncation, uncancellable shell jobs, blocking attachment/history work, hidden shortcuts under overlays, and prompt-loss paths undermine predictability. | -| Onboarding/configuration | **5.5/10** | Provider readiness and OAuth security controls are good; destructive route writes, stale project bootstrap, basename project collisions, mouse-only core actions, and silent configuration/agent-load errors are serious onboarding trust defects. | -| Runtime reliability/state integrity | **5.0/10** | Cancellation propagation, retry logic, streaming batching, and checkpoints are strong foundations, but continuation races, non-atomic completed saves, dropped queued sends, ignored runtime errors, and async event-ordering contract breaks remain. | -| Presentation/accessibility | **6.5/10** | The baseline TUI is calm and information-rich, with good diff/tool rendering and terminal compatibility work; keyboard accessibility, overlay ownership, scroll state, color-independent semantics, and real render-failure isolation need work. | -| Distribution/operations | **4.5/10** | Platform coverage, smoke testing, proxy support, and terminal cleanup are substantial, but update ordering, staging workflow authority, unsigned downloads, destructive cache replacement, version drift, telemetry control, and ARM64 validation are below release-grade. | -| Test/release readiness | **5.0/10** | The suite is large and source help/typecheck pass, but the final isolated CLI suite is **2,327 pass / 30 fail / 15 skip**; release workflows do not require the relevant full validation, and several tests encode or bypass the failure modes found here. | -| **Overall** | **6.0/10** | **Strong product depth, medium user trust, low current release confidence.** The score weights state integrity, security, and release safety more heavily than cosmetic polish. | - -## Top 10 highest-leverage findings - -All ten are retained as **High** severity from the underlying evidence. Ordering within High prioritizes security, data/context isolation, irreversible mutation, and release/session integrity. - -1. **[HIGH] A PR-controlled staging path can reach write-and-publish authority.** A same-repository PR title can trigger PR-head code in a workflow with `contents: write`, inherited secrets, release credentials, and later npm publication authority. Remove the PR release trigger and require protected, reviewed, pinned-commit execution. Evidence: `.github/workflows/cli-release-staging.yml:3-28,118-127,232-237`. - -2. **[HIGH] Project history identity collides for repositories with the same basename.** `/work/client/app` and `/work/internal/app` share the same project data directory, so history, checkpoints, and `--continue` can expose or resume another project's state. Key storage by canonical absolute path plus a stable hash and migrate legacy directories. Evidence: `cli/src/project-files.ts:50-59`; `cli/src/utils/run-state-storage.ts:121-145`. - -3. **[HIGH] Editing one model route can overwrite unrelated configuration.** The editable draft omits supported configuration fields and the force-write can reset or delete indexing, vision, failover, hooks, run limits, and discovery/capability metadata. Persist source-aware fragments or complete lossless drafts atomically. Evidence: `cli/src/utils/openbuff-provider.ts:346-364`; `docs/configuration.md:33-48,135-186,231-260`. - -4. **[HIGH] Cancellation can permanently fork visible UI history from SDK/model history.** Escape admits a new send before the cancelled run returns its authoritative preserved state, and aborted completion is discarded rather than becoming the next continuation base. Introduce a `cancelling` state and serialize or explicitly merge continuation state. Evidence: `cli/src/hooks/helpers/send-message.ts:353-375`; `cli/src/hooks/use-send-message.ts:581-592`; `sdk/src/__tests__/run-cancellation.test.ts:945-1107`. - -5. **[HIGH] Auto-update can kill a healthy active session before a replacement exists.** The wrapper terminates the child, then downloads, validates inadequately, and swallows failure. Stage, integrity-check, smoke-test, and atomically activate before stopping or relaunching the current binary. Evidence: `cli/release/index.js:639-696` and the parallel staging wrapper. - -6. **[HIGH] Project selection changes `cwd` without reinitializing project-scoped state.** The first-run picker can leave direnv, agents/MCP, skills, caches, and index state bound to the launch directory. Make project switching one cancellable bootstrap transaction. Evidence: `cli/src/index.tsx:293-300,343-368`; `cli/src/init/init-app.ts:19-65`. - -7. **[HIGH] The SDK promises async event handlers but dispatches them fire-and-forget.** Events may reorder and `client.run()` may resolve before callback side effects finish. Serialize and drain an awaited event queue, or narrow the public contract. Evidence: `sdk/src/run.ts:152-168,530-570,719-725`; `common/src/types/contracts/client.ts:63`; `sdk/e2e/utils/event-collector.ts:27-38`. - -8. **[HIGH] A public local-agent loader contract changed silently in the dirty worktree.** Its new default excludes project agents despite JSDoc and published documentation promising project/parent discovery, breaking callers without a type error. Preserve the SDK default or make the trust change an explicit versioned API with migration guidance. Evidence: `sdk/src/agents/load-agents.ts:146-172,207-222`; `README.md:177`; `docs/agents-and-tools.md:10-12`. - -9. **[HIGH] Core onboarding actions are mouse-only.** Keyboard-only users can browse but cannot activate the project picker's `Open` action; the shared button primitive also affects OAuth configure/disconnect/retry/close actions. Add focus and keyboard activation semantics plus end-to-end keyboard tests. Evidence: `cli/src/components/button.tsx:43-69`; `cli/src/components/project-picker-screen.tsx:269-285,486-501`; `cli/src/components/chatgpt-connect-banner.tsx:174-205,226-239`. - -10. **[HIGH] The final isolated CLI test suite is not green.** The final result was **2,327 pass / 30 fail / 15 skip**: 29 failures in `cli/src/__tests__/integration/local-agents.test.ts` and one in `init-type-sources.test.ts`. Repair the trust/default/test-isolation contract, regenerate init type sources, and make this isolated suite a release gate. Evidence: `cli/src/__tests__/integration/local-agents.test.ts:1`; independent final validation summary. - -## Evidence - -- Five independent discovery/audit pairs covered onboarding/configuration, interaction/commands, runtime/state, presentation, and distribution/operations. Every pair evaluated security, correctness, state mutation, error handling, performance, dependency hygiene, test gaps, and API/ABI breaks. -- The findings include direct file-and-line inspection, bounded reproductions, targeted tests, built-binary smoke checks, and isolated terminal captures. Duplicate observations, notably the command-palette 50-file cap, are consolidated here. -- Final validation state: `bun run cli/src/index.tsx --help` passed; `bun run --cwd cli typecheck` passed; built `--help` and `--version` passed; isolated 120x36 and 80x24 startup captures passed; the final isolated CLI suite reported **2,327 pass / 30 fail / 15 skip**. -- Several positive controls were directly substantiated: provider endpoint validation, PKCE/state/loopback OAuth controls, cross-origin authorization protection, abort-aware bounded retries, cancellation cleanup, checkpoint temp-write/rename, sensitive-file filtering, terminal-mode restoration, binary boot/tree-sitter smoke tests, and proxy parity tests. - -## Inference - -- The scores and overall **6.0/10** are synthesis judgments, not empirical usability or reliability measurements. They intentionally weight security, user data/context integrity, release safety, and recovery more heavily than breadth or visual polish. -- The project-skill trust issue is a medium-confidence security inference from load order and prompt injection behavior, not a demonstrated exploit against a user. -- Synchronous filesystem findings are inferred responsiveness risks from render-thread I/O patterns; no large-repository latency benchmark was run. -- Plain-text terminal captures cannot establish color contrast or prove that styled text was absent, so presentation conclusions based on captures remain cautious. -- The public local-agent loader regression is evidence from the actively changing, uncommitted worktree and may not represent a released version. - -## Unknowns and limits - -- No live provider/model calls were made. Real network behavior, provider failover UX, rate-limit recovery, token refresh under adverse networks, and long-running model sessions were not exercised end to end. -- No production publication or updater activation was performed. Release conclusions come from workflow/wrapper behavior and deterministic tests. -- The worktree changed during the audit. Two transient parse failures were observed and later fixed by another actor; they are not final open findings. Time-sensitive version and architecture-gate observations are tied to the recorded audit state. -- Unrelated standalone SDK behavior, agent prompt quality, inactive historical code, editor settings, generated bundles, dependencies, and compiled outputs were outside scope except at explicit CLI boundaries. -- Accessibility was assessed from code and terminal captures, not with assistive-technology user testing. - -## Cross-cutting findings - -### Transactions are missing at state boundaries - -Several unrelated defects share one architectural cause: multi-step mutations are treated as ordinary sequential operations rather than transactions. Route configuration is reconstructed from an incomplete draft (`cli/src/utils/openbuff-provider.ts:346-364`); completed chats are written as two live files (`cli/src/utils/run-state-storage.ts:93-105`); project switching changes only part of project state (`cli/src/index.tsx:343-368`); history uses unlocked read-modify-rewrite (`cli/src/hooks/use-input-history.ts:65-75`); and updater/postinstall flows remove the working state before the replacement is committed (`cli/release/index.js:639-696`; `cli/release/postinstall.js:7-18`). A shared design rule—prepare, validate, atomically commit, preserve rollback—would eliminate multiple high- and medium-severity classes. - -### Trust policy is fragmented across agents, skills, configuration, telemetry, and release - -Project-agent trust was changed at one API boundary while project skills still load and can shadow global skills (`cli/src/utils/skill-registry.ts:22-31`; `sdk/src/skills/load-skills.ts:159-178,221-234`). Runtime telemetry lacks a user-facing disable control (`common/src/analytics.ts:59-89`). Release downloads have no independent integrity verification (`cli/release/index.js:448-546`), and staging authority is reachable from a PR trigger. These are all facets of one product promise: users should know which local, repository, network, and published inputs they are trusting. - -### The UI has rich states, but ownership and failure propagation are inconsistent - -Openbuff represents detailed tool and agent progress, yet provider retry/connection state is hard-coded or absent (`cli/src/hooks/use-connection-status.ts:6-10`), runtime error events are ignored (`cli/src/utils/sdk-event-handlers.ts:745-759`), shell jobs cannot be cancelled (`cli/src/commands/router.ts:76-82`), nested render fallbacks do not catch (`cli/src/components/error-boundary.tsx:10-30`), and global shortcuts remain active under some overlays (`cli/src/chat.tsx:1339-1347`). The next UX leap is less about adding panels and more about giving every asynchronous operation one visible owner, state machine, cancellation path, and recovery action. - -### Tests are abundant but miss cross-boundary failure contracts - -Focused unit coverage is a genuine strength, but the most consequential defects sit between components: cancel-A/send-B, queue-to-send rejection, two-file persistence failure, project switch rebootstrap, updater rollback, published version consistency, and production parser behavior. Some tests reconstruct a test-only parser (`cli/src/__tests__/cli-args.test.ts:15-51`), manually inject continuation state (`cli/src/hooks/helpers/__tests__/send-message.test.ts:1380-1394`), or assert helper caps without exercising real searchable scope (`cli/src/components/__tests__/command-palette-screen.test.ts:63-73`). Release confidence requires scenario tests around ownership and commit points. - -### Main-thread synchronous I/O appears in several interactive features - -First-run browsing, `@` mentions, directory/file attachments, and chat history all perform broad synchronous filesystem work (`cli/src/utils/directory-browser.ts:15-49`; `common/src/project-file-tree.ts:140-229`; `cli/src/utils/pending-attachments.ts:347-438`; `cli/src/utils/chat-history.ts:48-114`). These should converge on asynchronous, cancellable, indexed primitives rather than receive isolated micro-fixes. - -## What is already genuinely strong - -- **Breadth with coherent interaction primitives.** Keyboard classification is centralized and well tested; slash commands have one registry, aliases, typo suggestions, presets, palette discovery, and prompt-history search. Input supports ANSI-safe paste, long-paste attachment conversion, mentions, images, binary detection, compression, and provider-boundary normalization. -- **Provider and credential foundations.** Provider schemas constrain endpoints and key transport; discovery is cancellable and avoids leaking authorization cross-origin; OAuth uses PKCE, state validation, loopback callbacks, escaped HTML, sanitized exchange errors, and owner-only credential permissions. Readiness messages identify missing routes, keys, and OAuth before send. -- **Rich runtime observability.** Tool/subagent lifecycle, phases, context use/compaction, queue previews, failed agents, elapsed time, model, git stats, cost, and cache hit rate are represented. Streaming updates are batched and flushed at terminal states rather than naively rerendered for every chunk. -- **Cancellation and checkpoint foundations.** Runtime cancellation shares an abort signal, prevents new tools, waits for cooperative cleanup, stops browser sessions, and removes owned clone directories. Mid-turn checkpoints already use same-directory temporary files and rename, include turn identity, and reject stale or malformed snapshots. -- **Terminal craftsmanship.** Fatal handling restores raw mode, alternate screen, mouse/focus modes, bracketed paste, and cursor visibility. Theme/color compatibility is broad, diffs have parsed hunks, line gutters, truncation disclosure, collapsing, and unified `+`/`-` semantics. -- **Serious distribution effort.** The artifact matrix covers Linux x64/ARM64, macOS Intel/current and legacy, Apple Silicon, and Windows x64. Smoke testing includes long-lived startup and embedded tree-sitter initialization; proxy handling supports CONNECT, upper/lowercase variables, `NO_PROXY`, timeouts, and production/staging parity. -- **Large validation surface.** Even in the failing final run, 2,327 tests passed. Targeted onboarding tests passed 216/216, wrapper/proxy/analytics/error tests passed 60/60, and the source help/typecheck and built startup checks passed at final validation. - -## Prioritized next-level feature roadmap - -This roadmap derives from current gaps; it is not a generic competitor feature list. - -### Now — stabilize the trust core - -1. **Transactional project identity and switching.** Replace basename storage with canonical-path hashing and implement a single `switchProject()` transaction that reloads environment, config, agents/MCP, skills, indexes, and caches before chat becomes active. User impact: prevents cross-project history exposure and wrong-project execution. Gaps: `cli/src/project-files.ts:50-59`; `cli/src/index.tsx:293-300,343-368`. - -2. **Lossless, atomic configuration editing.** Preserve every source fragment and unrelated field; validate a complete draft; write temporary files and atomically rename. Add byte/semantic preservation tests. User impact: makes provider/model changes safe and reversible. Gap: `cli/src/utils/openbuff-provider.ts:346-364`. - -3. **Serialized run continuation and durable queue acceptance.** Add explicit run states (`running`, `cancelling`, `settling`), await or merge cancelled `RunState`, retain queued items until send acceptance, and drain async callbacks before run completion. User impact: no silent context or prompt loss after stop/retry. Gaps: `cli/src/hooks/helpers/send-message.ts:353-375`; `cli/src/hooks/use-send-message.ts:581-592`; `cli/src/hooks/use-message-queue.ts:294-302`; `sdk/src/run.ts:719-725`. - -4. **Safe release/update pipeline.** Remove PR publication authority, require protected reviewed commits, publish signed checksums/provenance, stage and smoke replacements before activation, retain rollback binaries, and gate publication on exact-version/source/typecheck/isolated-suite checks. User impact: updates cannot silently terminate work or replace a working binary with an unverified/unvalidated one. Gaps: `.github/workflows/cli-release-staging.yml:3-28`; `cli/release/index.js:448-546,639-696`; `cli/release/postinstall.js:7-18`; `.github/workflows/cli-release-prod.yml:81-90`. - -5. **Make the isolated CLI suite green and mandatory.** Resolve the 29 local-agent failures and generated init-type drift; test the production parser, project switch, cancellation continuation, queue rejection, persistence interruption, updater rollback, and async event ordering. User impact: converts the current broad test investment into release confidence. Gaps: `cli/src/__tests__/integration/local-agents.test.ts:1`; `init-type-sources.test.ts`; `cli/src/__tests__/cli-args.test.ts:15-51`; `cli/src/hooks/helpers/__tests__/send-message.test.ts:961-970,1380-1394`. - -### Next — make trust and recovery visible in the UX - -1. **Unified trust center / `openbuff doctor`.** Show project-agent and project-skill origins/shadowing, config parse diagnostics, telemetry state, credential/provider readiness, binary version/digest, and release provenance. User impact: users can understand what repository code and network services influence a run without reading logs. Gaps: `cli/src/utils/skill-registry.ts:22-31`; `sdk/src/provider-config.ts:1177-1194`; `common/src/analytics.ts:59-89`; `cli/release/index.js:448-546`. - -2. **Structured resilience timeline.** Add provider attempt, retry scheduled, failover, recovery, recoverable runtime error, and terminal error events; render them with retry/cancel/copy-diagnostic actions and redact internal stacks. User impact: users know whether the agent is thinking, offline, retrying, failed over, or needs action. Gaps: `cli/src/hooks/use-connection-status.ts:6-10`; `common/src/types/print-mode.ts:189-205`; `cli/src/utils/sdk-event-handlers.ts:745-759`; `packages/agent-runtime/src/run-agent-step.ts:1574-1609`. - -3. **Keyboard-complete modal and onboarding system.** Give the shared button focus/activation semantics, enforce one keyboard owner for every overlay, document shortcuts from the same registry, and add narrow/wide renderer-level keyboard journeys. User impact: the CLI becomes fully operable and predictable without a mouse. Gaps: `cli/src/components/button.tsx:43-69`; `cli/src/chat.tsx:1339-1347`; `cli/src/components/help-banner.tsx:59-105`. - -4. **Cancellable job and attachment manager.** Give bash, file attachment, and directory ingestion explicit job IDs, progress, timeout/cancel, and retry; move filesystem work off the render thread. User impact: large files or hung commands no longer freeze or strand the session. Gaps: `cli/src/commands/router.ts:76-82`; `cli/src/utils/pending-attachments.ts:347-438`. - -5. **Reliable session recovery.** Store completed chat state as one atomic versioned envelope (or committed multi-file generation), validate `--continue` containment, surface recovery choices, and preserve rejected queue/history writes. User impact: crashes and concurrent terminals stop turning into silent session loss or unexpected file ingestion. Gaps: `cli/src/utils/run-state-storage.ts:93-105,127-193`; `cli/src/hooks/use-input-history.ts:65-75`. - -### Later — differentiation built on existing strengths - -1. **Indexed, instant project navigation.** Reuse the code index or a persistent file metadata index for Ctrl+P, `@`, history, and directory attachments, with background invalidation and progressive results. User impact: monorepo-scale navigation stays responsive and complete. Gaps: `cli/src/components/command-palette-screen.tsx:170-191`; `cli/src/hooks/use-suggestion-engine.ts:641-671`; `cli/src/utils/chat-history.ts:48-114`. - -2. **First-class run replay and branchable recovery.** Build on preserved cancelled run state and atomic checkpoints to offer resume from checkpoint, retry from failure, branch from a turn, and compare continuation histories. User impact: advanced users can recover and experiment without losing provenance. Foundations/gaps: `sdk/src/__tests__/run-cancellation.test.ts:945-1107`; checkpoint behavior in `cli/src/utils/run-state-storage.ts:216-256`; current cancellation gap at `cli/src/hooks/helpers/send-message.ts:353-375`. - -3. **Verifiable local-first distribution.** Expose artifact digest, signature/provenance, telemetry policy, local data paths, cleanup/retention, and rollback directly in CLI diagnostics. User impact: BYOK/local-first becomes inspectable rather than merely asserted. Gaps: `cli/release/index.js:448-546`; `common/src/analytics.ts:59-89`; `cli/src/utils/clipboard-image.ts:18-23`. - -4. **Adaptive terminal presentation with accessibility contracts.** Unify responsive tokens, add color-independent diff semantics, preserve reading position during streams, and validate readable renderer output across width/color modes. User impact: trustworthy review and navigation on narrow, monochrome, tmux, and accessibility-constrained terminals. Gaps: `cli/src/hooks/use-terminal-layout.ts:7-23`; `cli/src/hooks/use-scroll-management.ts:92-130`; `cli/src/components/tools/diff-viewer.tsx:459-502`. - -## Remaining findings by audit domain and severity - -The Top 10 are not repeated here. The command-palette 50-file cap appeared independently in interaction and presentation shards and is listed once. - -### Security and privacy - -- **[MEDIUM] `--continue` permits path traversal outside the chat directory.** Require a basename and containment check before reads. `cli/src/utils/run-state-storage.ts:127-134`; caller path `cli/src/index.tsx:121-160`; safe deletion precedent `cli/src/utils/chat-history.ts:129-139`. -- **[MEDIUM] Convenience git commands interpolate raw shell arguments.** `/diff` and `/changes` can execute shell metacharacters despite presenting as constrained git helpers. `cli/src/commands/command-registry.ts:352-371`; `cli/src/data/slash-commands.ts:169-177`. -- **[MEDIUM, inference] Project skills bypass the project-agent trust boundary and can shadow globals.** `cli/src/utils/skill-registry.ts:22-31`; `sdk/src/skills/load-skills.ts:159-178,221-234`; `cli/src/commands/command-registry.ts:1013-1069`. -- **[MEDIUM] Downloaded executables lack independent integrity/provenance verification.** `cli/release/index.js:448-546`; `.github/workflows/cli-release-build.yml:278-298,431-443`. -- **[MEDIUM] Production runtime telemetry has no documented user-facing disable control.** `common/src/analytics.ts:59-89`; `.github/workflows/cli-release-build.yml:158-173`; `packages/agent-runtime/src/main-prompt.ts:67-88`; `docs/environment-variables.md:9-10`. -- **[LOW] Clipboard images persist as plaintext files in a predictable shared temp directory without cleanup.** `cli/src/utils/clipboard-image.ts:18-23,175-201`. - -### Correctness and state mutation - -- **[MEDIUM] Completed chat persistence is not crash-atomic.** Two live files can become mismatched or truncated. `cli/src/utils/run-state-storage.ts:93-105,152-193`; safer checkpoint precedent `:216-256`. -- **[MEDIUM] Rejected queued sends are irreversibly dropped.** `cli/src/hooks/use-message-queue.ts:121-127,294-302`; `cli/src/hooks/use-chat-streaming.ts:154-170`. -- **[MEDIUM] Cross-terminal prompt-history writes can overwrite each other.** `cli/src/hooks/use-input-history.ts:65-75`; `cli/src/utils/message-history.ts:106-123`. -- **[MEDIUM] Ctrl+B can move the input cursor to `-1`.** `cli/src/components/multiline-input.tsx:940-948`; safe arrow path `:962-966`. -- **[MEDIUM] Palette file search permanently excludes paths beyond the first 50.** `cli/src/components/command-palette-screen.tsx:170-191`; test gap `cli/src/components/__tests__/command-palette-screen.test.ts:63-73`. -- **[MEDIUM] Hidden chat shortcuts remain active under command/history overlays.** `cli/src/chat.tsx:1339-1347`; `cli/src/utils/keyboard-actions.ts:168-183,327-354`. -- **[MEDIUM] Page-up incorrectly re-enables follow mode and can snap readers back to streaming output.** `cli/src/hooks/use-scroll-management.ts:66,92-100,126-130`. -- **[MEDIUM] Scroll listeners can remain attached to a destroyed scrollbox after full-screen overlays.** `cli/src/hooks/use-chat-ui.ts:72-93`; `cli/src/hooks/use-scroll-management.ts:113-143`; `cli/src/chat.tsx:1543-1617`. -- **[MEDIUM] Release postinstall deletes the working offline fallback before replacement.** `cli/release/postinstall.js:7-18`; `cli/release/index.js:606-636`. -- **[MEDIUM] Version sources/tests mask wrapper, package, and binary drift.** `cli/release/index.js:274-288`; `cli/src/__tests__/release-wrapper.test.ts:15-18`; observed package versions at `cli/release/package.json:3`, `cli/release-staging/package.json:3`, and `cli/package.json:3`. -- **[MEDIUM] Published usage treats a positional directory as cwd, while production parses it as a prompt.** `cli/release/README.md:23-31`; `cli/src/index.tsx:125-163`. -- **[LOW] Per-turn cost displays an ambiguous raw cents value.** `cli/src/components/message-footer.tsx:221-235`; `cli/src/components/status-bar.tsx:192-201`; `packages/agent-runtime/src/run-agent-step.ts:1251`. -- **[LOW] Three responsive-layout contracts disagree.** `cli/src/hooks/use-terminal-layout.ts:7,23,121-130`; `cli/src/hooks/use-terminal-breakpoints.ts:20-29`; `cli/src/hooks/use-grid-layout.ts:13-22`; `cli/knowledge.md:306-320`. -- **[LOW] Side-by-side diffs rely primarily on red/green rather than explicit add/delete markers.** `cli/src/components/tools/diff-viewer.tsx:431-502`; `cli/src/components/tools/__tests__/diff-viewer.test.tsx:169-190`. - -### Error handling and resilience - -- **[MEDIUM] Provider connection/retry indicators are disconnected from real provider state.** `cli/src/hooks/use-connection-status.ts:6-10`; `cli/src/app.tsx:220`; `sdk/src/impl/llm.ts:1280-1318`; `common/src/types/print-mode.ts:189-205`. -- **[MEDIUM] Internal stack traces can reach the TUI and persisted history.** `packages/agent-runtime/src/run-agent-step.ts:1574-1609`; `cli/src/hooks/helpers/send-message.ts:450-453`; `sdk/src/error-utils.ts:107-124`. -- **[MEDIUM] Declared runtime `error` events are silently ignored by the CLI.** `common/src/types/print-mode.ts:12-16,189-205`; `sdk/src/run.ts:478-481`; `cli/src/utils/sdk-event-handlers.ts:745-759`. -- **[MEDIUM] Malformed ordinary configuration files disappear without diagnostics.** `sdk/src/provider-config.ts:1177-1194`; `cli/src/utils/openbuff-provider.ts:259-276`. -- **[MEDIUM] OAuth code exchange and refresh have no network timeout.** `cli/src/utils/chatgpt-oauth.ts:164-168,293-305`; `sdk/src/credentials.ts:221-292`. -- **[MEDIUM] Broken agent modules vanish from TUI validation diagnostics.** `sdk/src/agents/load-agents.ts:229-280`; `cli/src/utils/local-agent-registry.ts:63-76`. -- **[MEDIUM] `/init` can throw after partial scaffolding without coherent rollback/recovery.** `cli/src/commands/init.ts:61-90`; `cli/src/commands/router.ts:443-462`; `cli/src/commands/__tests__/init.test.ts:289-331`. -- **[MEDIUM] Interactive bash jobs have neither timeout nor cancellation.** `cli/src/commands/router.ts:76-82`; `cli/src/components/pending-bash-message.tsx`. -- **[MEDIUM] Nested-agent fallback is not an actual error boundary.** `cli/src/components/error-boundary.tsx:10-30`; `cli/src/components/message-with-agents.tsx:75-92`; `cli/src/index.tsx:386-423`. -- **[LOW] Release HTTP redirect following is unbounded and omits 307/308.** `cli/release/http.js:144-155`; `cli/src/__tests__/proxy-http-get.test.ts:161-239`. - -### Performance and responsiveness - -- **[MEDIUM] Every new `@` mention session rebuilds a project tree of up to 10,000 files.** `cli/src/hooks/use-suggestion-engine.ts:641-671`; `common/src/project-file-tree.ts:140-229`. -- **[MEDIUM] “Background” attachment work performs synchronous enumeration, sort, stat, and reads on the renderer thread.** `cli/src/utils/pending-attachments.ts:347-438`; send blocking at `cli/src/commands/router.ts:467-472`. -- **[MEDIUM] Chat-history search synchronously reads/parses up to 500 full message files.** `cli/src/utils/chat-history.ts:48-114`; `cli/src/components/chat-history-screen.tsx:51-60`. -- **[MEDIUM, inference] First-run directory browsing synchronously stats every entry and child `.git` directory.** `cli/src/utils/directory-browser.ts:15-49`; `cli/src/hooks/use-directory-browser.ts:41-47`. -- **[MEDIUM] Unknown terminals can incur two sequential 500 ms OSC probes before first paint.** `cli/src/utils/terminal-color-detection.ts:18-20,82-83,193-196,416-430`; `cli/src/index.tsx:246-258`. - -### Dependency, API, documentation, and release hygiene - -- **[MEDIUM] Linux ARM64 artifacts are published without executing the binary.** `.github/workflows/cli-release-build.yml:40-46,258-276`; checked-in AArch64 ripgrep makes native validation material. -- **[LOW] CLI imports undeclared `lodash`, relying on a root dev dependency/hoisting.** `cli/src/utils/send-message-helpers.ts:7`; `cli/src/utils/message-block-helpers.ts:1`; `cli/src/utils/sdk-event-handlers.ts:1`; `cli/package.json`; root `package.json:64`. -- **[LOW] CLI imports undeclared `picocolors`.** `cli/src/index.tsx:25`; `cli/package.json:34-71`. -- **[LOW] Unused `terminal-image` retains a duplicate/heavy image stack.** `cli/package.json:54`; corresponding `bun.lock` entries; custom path `cli/src/utils/terminal-images.ts` and direct `jimp` usage in `cli/src/utils/image-thumbnail.ts`. -- **[LOW] Clipboard storage retains the legacy `codebuff-clipboard-images` namespace.** `cli/src/utils/clipboard-image.ts:18-19`; identity contract `docs/architecture.md`. - -### Test and presentation coverage gaps - -- **[MEDIUM] Recovery tests simulate state production does not apply.** Add real deferred run cancellation/continuation, rejected queue restoration, async event drain, traversal, and injected mid-write tests. `cli/src/hooks/helpers/__tests__/send-message.test.ts:961-970,992,1380-1394`; production guard `cli/src/hooks/use-send-message.ts:581-592`. -- **[MEDIUM] CLI argument tests rebuild a smaller parser instead of exercising production flags.** `cli/src/__tests__/cli-args.test.ts:15-51`; production parser `cli/src/index.tsx:107-164`. -- **[LOW] Built-in help omits Ctrl+P, Ctrl+R, Ctrl+V, Shift+Enter, Tab behavior, and paging.** `cli/src/components/help-banner.tsx:59-105`; `cli/src/utils/keyboard-actions.ts`. -- **[LOW, inference] Presentation tests assert helpers/capture existence rather than readable rendered output.** `cli/src/components/__tests__/command-palette-screen.test.ts:32-297`; `debug/tmux-sessions/audit-cli-baseline/capture-003-command-palette.txt`. -- **[MEDIUM] Release publication lacks a required full validation gate and exact version assertion.** `.github/workflows/cli-release-prod.yml:81-90`; `.github/workflows/cli-release-build.yml:232-241,269-276,423-429`. - -## Coverage - -The accompanying `COVERAGE-MATRIX.md` records **five complete discovery/audit pairs**, all covered: - -1. Startup, onboarding, project selection, provider/model configuration, OAuth, and local-agent validation. -2. Input, keyboard, commands, bash, history, suggestions, clipboard, attachments, and images. -3. Send/stream lifecycle, queueing, cancellation, sessions, persistence, and SDK/runtime contracts. -4. Layout, themes, accessibility, scrolling, nested agents/tools, and rendering performance. -5. Packaging, release/update, platforms, CI, logging, analytics/privacy, diagnostics, and documentation. - -Across those pairs, all eight audit lenses were covered: security, correctness, state mutation, error handling, performance, dependency hygiene, test coverage, and API/ABI contracts. `cli/` was audited across all five pairs; CLI-facing boundaries in `sdk/`, `common/`, `packages/`, `.github/`, docs, root metadata, examples, and test setup were included as recorded in the matrix. - -Explicit exclusions were unrelated standalone SDK internals, agent prompt/quality review, inactive `agents-graveyard/`, editor settings, local shims, transient scratch fixtures, generated bundles, dependency directories, and compiled outputs except where a compiled binary was the subject of smoke/version evidence. Existing `.agents/sessions/**` audit findings and reports were excluded to preserve independence. No product code was edited by the audit. - -No live provider/model calls were executed. Runtime conclusions rely on source, deterministic/focused tests, isolated local TUI captures, and existing compiled-binary smoke behavior. - -The worktree was actively changing throughout the audit. Two transient parse failures were observed and subsequently fixed; final source help and typecheck pass. Time-sensitive observations—especially version drift, an environment-architecture check, and the public local-agent default—must be read as evidence of the audited dirty state. The final authoritative isolated CLI suite result is **2,327 pass / 30 fail / 15 skip**, which remains the release-readiness baseline for this report. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/COVERAGE-MATRIX.md b/.agents/sessions/audit-cli-next-level-2026-07/COVERAGE-MATRIX.md deleted file mode 100644 index 0e7f44d0e9..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/COVERAGE-MATRIX.md +++ /dev/null @@ -1,54 +0,0 @@ -# Coverage matrix - -| Functional domain | Discovery shard | Audit shard | Covered | -| -------------------------------------------------------------------------------------------- | --------------------------- | -------------------------- | ------- | -| Startup, onboarding, project selection, provider/model config, OAuth, local-agent validation | picker-onboarding-config | audit-onboarding-config | yes | -| Input, keyboard, commands, bash, history, suggestions, clipboard, attachments, images | picker-interaction-commands | audit-interaction-commands | yes | -| Send/stream lifecycle, queueing, cancellation, sessions, persistence, SDK/runtime contracts | picker-runtime-state | audit-runtime-state | yes | -| Layout, themes, accessibility, scrolling, nested agents/tools, rendering performance | picker-presentation-quality | audit-presentation-quality | yes | -| Packaging, release/update, platforms, CI, logging, analytics/privacy, diagnostics, docs | picker-distribution-quality | audit-distribution-quality | yes | -| Security | all five pairs | all five audit shards | yes | -| Correctness | all five pairs | all five audit shards | yes | -| State mutation | all five pairs | all five audit shards | yes | -| Error handling | all five pairs | all five audit shards | yes | -| Performance | all five pairs | all five audit shards | yes | -| Dependency hygiene | all five pairs | all five audit shards | yes | -| Test coverage gaps | all five pairs | all five audit shards | yes | -| API/ABI contract breaks | all five pairs | all five audit shards | yes | - -## Subsystem enumeration - -| Top-level subsystem | Disposition | -| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `cli/` | audited across five functional shard pairs; generated bundles, `node_modules`, and compiled outputs excluded | -| `sdk/` | audited at CLI-facing provider, run/event, persistence, filesystem/tool, packaging, and OAuth boundaries; unrelated standalone SDK surface out-of-scope | -| `common/` | audited at CLI contracts, messages/events, project tree, config/env, analytics, agent validation, and tool-result boundaries; unrelated utilities out-of-scope | -| `packages/` | audited at `agent-runtime`, `indexer`, `code-map`, `internal`, and build integration paths used by the CLI; package-internal behavior unrelated to CLI out-of-scope | -| `agents/` | out-of-scope except documented CLI rendering/command contracts; prompt and agent-quality audit is a separate product surface | -| `evals/` | out-of-scope except breadth-classification machinery used to select this audit workflow | -| `scripts/` | audited only for CLI smoke, tmux, structural-map, and release/test helpers; unrelated provider benchmarks/services out-of-scope | -| `agents-graveyard/` | out-of-scope: inactive historical code | -| `docs/` | audited for architecture, request flow, local mode, configuration, testing, environment, provider setup, and agent/tool CLI contracts | -| `.agents/` | structural map and audit instructions read; existing findings/reports explicitly excluded to preserve independence | -| `.github/` | audited CLI CI, nightly E2E, release build, staging, production, and adjacent SDK release workflows | -| `openbuff.d.example/` | audited as executable provider/routes/hooks/indexing configuration examples | -| `.bin/` | out-of-scope: local tool shim | -| `.vscode/` | out-of-scope: editor-only settings | -| `.e2e-scratch/` | out-of-scope: transient test fixtures | -| `test/` | audited only where root test setup affects CLI tests | - -## Root files and metadata - -| Root artifact | Disposition | -| ----------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | -| `package.json`, `bun.lock`, `.bun-version`, `bunfig.toml`, `tsconfig.json`, `tsconfig.base.json`, `eslint.config.js` | audited selectively for CLI dependency, runtime, workspace, build, and validation contracts | -| `README.md`, `README.zh-CN.md`, `WINDOWS.md`, `CONTRIBUTING.md`, `SECURITY.md`, `.env.example`, `openbuff.json.example` | audited where they make CLI install, platform, privacy, configuration, and operational claims | -| `AGENTS.md` | controlling audit instructions; not a product subsystem | -| `ROUTER.md`, `INFISICAL_SETUP_GUIDE.md`, `knowledge.md`, `.envrc`, `.gitignore`, `.prettierrc`, `CODE_OF_CONDUCT.md`, `LICENSE`, `NOTICE` | out-of-scope except incidental references; they do not materially implement the current CLI surface | - -## Exclusions - -- Existing audit reports and findings were not used as source evidence. -- No product code was edited by this audit. -- Live provider/model calls were not executed; runtime behavior was assessed from source, deterministic tests, and isolated local TUI captures. -- The worktree was actively changing during the audit. Time-sensitive validation is reported with its observed timestamp/state, and transient parse failures are separated from final-state findings. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/IMPLEMENTATION-REPORT.md b/.agents/sessions/audit-cli-next-level-2026-07/IMPLEMENTATION-REPORT.md deleted file mode 100644 index 440efec655..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/IMPLEMENTATION-REPORT.md +++ /dev/null @@ -1,48 +0,0 @@ -# CLI audit remediation report - -All findings in `AUDIT-REPORT.md` were reconciled against the current worktree and implemented from highest to lowest value. - -## Trust, release, and state integrity - -- Removed PR/push staging publication authority, narrowed secrets, added source/full-suite gates, exact version checks, ARM64 execution smoke, checksums, provenance attestations, and npm provenance. -- Made updates non-disruptive, staged pending activation, preserved offline binaries, bounded redirects, and corrected wrapper/version/documentation drift. -- Isolated project storage by canonical-path hash; validated continuation IDs; made chat/config/provider writes atomic and lossless. -- Made project switching transactional across cwd, direnv, index, agents/MCP, skills, SDK client, and chat state with rollback. -- Unified project agent/skill trust at the CLI boundary while retaining backward-compatible public SDK defaults. - -## Runtime and recovery - -- Serialized cancellation: a cancelling run retains continuation ownership until the SDK returns its authoritative preserved state. -- Serialized and drained async SDK event/stream callbacks before `run()` resolves. -- Restored rejected queued prompts at the queue head and paused safely instead of dropping or retry-looping. -- Added provider retry/failover/recovery events and a visible ordered resilience timeline. -- Surfaced runtime error events and removed internal stack frames from user-visible/persisted errors. -- Added OAuth exchange/refresh deadlines, atomic `/init` rollback/recovery, bounded cancellable bash commands, and append-safe cross-terminal prompt history. - -## Interaction, accessibility, and presentation - -- Added keyboard activation to the shared button primitive, explicit project-picker activation, and OAuth keyboard actions. -- Removed git-helper shell injection, fixed Ctrl+B underflow, isolated overlay keyboard ownership, corrected scroll-follow/listener ownership, and searched the complete project in the command palette. -- Added a real React render error boundary, complete shortcut help, unambiguous dollar cost formatting, canonical responsive breakpoints, and explicit side-by-side diff markers. -- Added `openbuff doctor` for provider, trust, agent, skill, MCP, and validation diagnostics. - -## Responsiveness, privacy, and dependencies - -- Moved project browsing, attachment reads, and chat-history loading off the renderer thread; bounded concurrent history reads and cached mention-tree refreshes. -- Reduced terminal theme probe latency. -- Added `OPENBUFF_TELEMETRY=0` and `DO_NOT_TRACK=1` runtime opt-outs. -- Secured and cleaned clipboard temp images under an Openbuff owner-only namespace. -- Removed CLI lodash usage, declared `picocolors`, and removed unused `terminal-image` plus its lockfile graph. -- Extracted the production CLI parser and made tests exercise it directly. - -## Final validation - -- Monorepo environment architecture and all workspace typechecks: pass. -- Isolated CLI suite: **2,397 pass / 0 fail / 16 environment-dependent skip**. -- SDK suite: **818 pass / 0 fail / 1 existing TODO skip**. -- Agent runtime suite: **922 pass / 0 fail**. -- Common suite: **647 pass / 0 fail**. -- Agents suite: **529 pass / 0 fail**. -- Keyboard/render overflow regression: passes separately under `NODE_ENV=test` (OpenTUI's test utility imports React `act`, which production React intentionally omits). -- Source `--help` smoke and `git diff --check`: pass. -- Fresh compiled binary build plus built `--version` and `--help` smoke: pass. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/MAP.md b/.agents/sessions/audit-cli-next-level-2026-07/MAP.md deleted file mode 100644 index fe4c902962..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/MAP.md +++ /dev/null @@ -1,341 +0,0 @@ -# Structural Map — openbuff - -- **Project root:** `/home/ben/Code/CLI/openbuff` -- **Built at:** 2026-07-11T19:56:45.763Z -- **Total files indexed:** 1594 -- **Graph:** 14030 nodes, 70654 edges - -> Pin this file in context. Every audit shard navigates from here instead of doing fuzzy round-trip discovery. - -## Entry points - -- `cli/src/index.tsx` -- `packages/code-map/src/index.ts` -- `packages/indexer/src/index.ts` -- `packages/internal/src/index.ts` - -## Directories (by size, biggest first) - -| dir | files | total size | top symbols | -| -------------------------- | ----- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `cli` | 434 | 2.9 MB | render, main, TestItem, tmux, createErrorMessage, formatTimestamp | -| `packages` | 318 | 2.4 MB | start, Greeting, greet, Greeter, flush, doGenerate | -| `sdk` | 166 | 1.4 MB | main, run, evaluate, createMockFs, log, resolveMcpEnv | -| `common` | 246 | 1.1 MB | createMockLogger, getStringProperty, getFileExtension, process, sleep, size | -| `agents` | 95 | 1.1 MB | extractInlineFunctionSource, parseGateStateBlock, feedJson, collectToolInputFiles, isFileChangingTool, hasEditArtifact | -| `evals` | 69 | 822.5 KB | main, run, makeEvalRun, makeAgentResults, toolCall, isRecord | -| `scripts` | 67 | 515.7 KB | main, parseArgs, computeCost, ConversationMessage, TurnResult, makeConversationStreamRequest | -| `agents-graveyard` | 121 | 350.6 KB | createBase2WithTaskResearcher, getLatestEditToolResults, extractSpawnResults, getSpawnResults, createResearchImplementOrchestrator, createBase2Implementor | -| `bun.lock` | 1 | 270.1 KB | — | -| `docs` | 13 | 207.8 KB | — | -| `.agents` | 18 | 119.1 KB | publisher, getSpawnerPrompt, getSystemPrompt, getInstructionsPrompt, getDefaultReviewModeInstructions, getWorkModeInstructions | -| `.github` | 14 | 55.6 KB | — | -| `openbuff.d.example` | 4 | 22.9 KB | — | -| `LICENSE` | 1 | 11.1 KB | — | -| `.bin` | 1 | 8.5 KB | — | -| `README.zh-CN.md` | 1 | 8.1 KB | — | -| `README.md` | 1 | 8.0 KB | — | -| `WINDOWS.md` | 1 | 7.3 KB | — | -| `CONTRIBUTING.md` | 1 | 5.6 KB | — | -| `CODE_OF_CONDUCT.md` | 1 | 4.5 KB | — | -| `eslint.config.js` | 1 | 4.0 KB | — | -| `AGENTS.md` | 1 | 3.6 KB | — | -| `ROUTER.md` | 1 | 3.3 KB | — | -| `package.json` | 1 | 2.6 KB | — | -| `INFISICAL_SETUP_GUIDE.md` | 1 | 2.5 KB | — | -| `.env.example` | 1 | 1.7 KB | — | -| `tsconfig.json` | 1 | 839 B | — | -| `SECURITY.md` | 1 | 520 B | — | -| `.gitignore` | 1 | 487 B | — | -| `.vscode` | 1 | 438 B | — | -| `bunfig.toml` | 1 | 432 B | — | -| `.prettierrc` | 1 | 389 B | — | -| `tsconfig.base.json` | 1 | 386 B | — | -| `test` | 1 | 332 B | setup | -| `knowledge.md` | 1 | 287 B | — | -| `.e2e-scratch` | 2 | 279 B | add, greet, multiply | -| `NOTICE` | 1 | 156 B | — | -| `openbuff.json.example` | 1 | 118 B | — | -| `.envrc` | 1 | 30 B | — | -| `.bun-version` | 1 | 7 B | — | - -## Largest files per directory - -### `cli` - -- `cli/src/data/initial-agent-type-sources.generated.ts` — 70.3 KB, 3 symbols -- `cli/src/hooks/helpers/__tests__/send-message.test.ts` — 56.4 KB, 0 symbols -- `cli/src/utils/__tests__/message-block-helpers.test.ts` — 55.1 KB, 0 symbols -- `cli/src/chat.tsx` — 55.1 KB, 1 symbols -- `cli/src/utils/__tests__/send-message-helpers.test.ts` — 48.1 KB, 0 symbols - -### `packages` - -- `packages/agent-runtime/src/__tests__/read-files-edit-state.test.ts` — 126.5 KB, 2 symbols -- `packages/agent-runtime/src/__tests__/process-str-replace.test.ts` — 76.8 KB, 0 symbols -- `packages/agent-runtime/src/__tests__/run-programmatic-step.test.ts` — 66.6 KB, 0 symbols -- `packages/agent-runtime/src/process-str-replace.ts` — 66.5 KB, 30 symbols -- `packages/agent-runtime/src/run-agent-step.ts` — 57.5 KB, 4 symbols - -### `sdk` - -- `sdk/src/__tests__/model-provider.test.ts` — 80.0 KB, 3 symbols -- `sdk/src/provider-config.ts` — 72.9 KB, 30 symbols -- `sdk/src/tools/browser-logs.ts` — 63.4 KB, 30 symbols -- `sdk/src/impl/llm.ts` — 52.9 KB, 20 symbols -- `sdk/src/run.ts` — 44.9 KB, 12 symbols - -### `common` - -- `common/src/templates/initial-agents-dir/types/tools.ts` — 47.8 KB, 30 symbols -- `common/src/util/__tests__/messages.test.ts` — 41.4 KB, 0 symbols -- `common/src/tools/results/filesystem.ts` — 33.9 KB, 30 symbols -- `common/src/__tests__/agent-validation.test.ts` — 28.9 KB, 0 symbols -- `common/src/util/__tests__/saxy.test.ts` — 26.1 KB, 0 symbols - -### `agents` - -- `agents/base2/base2.ts` — 179.3 KB, 30 symbols -- `agents/__tests__/context-pruner.test.ts` — 122.6 KB, 2 symbols -- `agents/__tests__/base2.test.ts` — 112.8 KB, 5 symbols -- `agents/context-pruner.ts` — 76.3 KB, 30 symbols -- `agents/types/tools.ts` — 47.8 KB, 30 symbols - -### `evals` - -- `evals/buffbench/eval-openbuff-v2.json` — 325.4 KB, 0 symbols -- `evals/buffbench/__tests__/plan-sharding-signals.test.ts` — 37.1 KB, 6 symbols -- `evals/buffbench/plan-sharding-signals.ts` — 30.9 KB, 26 symbols -- `evals/buffbench/run-buffbench.ts` — 23.7 KB, 10 symbols -- `evals/buffbench/pick-commits.ts` — 22.3 KB, 17 symbols - -### `scripts` - -- `scripts/test-fireworks-cache-intervals.ts` — 34.3 KB, 10 symbols -- `scripts/benchmark-providers.ts` — 33.1 KB, 16 symbols -- `scripts/test-fireworks-long.ts` — 30.7 KB, 6 symbols -- `scripts/test-canopywave-long.ts` — 29.0 KB, 5 symbols -- `scripts/test-siliconflow.ts` — 27.3 KB, 5 symbols - -### `agents-graveyard` - -- `agents-graveyard/base/base-prompts.ts` — 25.3 KB, 3 symbols -- `agents-graveyard/editor/best-of-n/editor-best-of-n.ts` — 18.5 KB, 4 symbols -- `agents-graveyard/base/ask.ts` — 12.2 KB, 0 symbols -- `agents-graveyard/registry/transform-agent.ts` — 12.0 KB, 0 symbols -- `agents-graveyard/base2/task-researcher/base2-with-task-researcher-planner-pro.ts` — 11.4 KB, 1 symbols - -### `bun.lock` - -- `bun.lock` — 270.1 KB, 0 symbols - -### `docs` - -- `docs/agents-and-tools.md` — 76.9 KB, 0 symbols -- `docs/codebuff-to-openbuff-migration.md` — 30.7 KB, 0 symbols -- `docs/configuration.md` — 23.1 KB, 0 symbols -- `docs/openbuff-provider-model-setup-ux.md` — 20.4 KB, 0 symbols -- `docs/architecture.md` — 12.3 KB, 0 symbols - -### `.agents` - -- `.agents/types/tools.ts` — 47.8 KB, 30 symbols -- `.agents/types/agent-definition.ts` — 16.3 KB, 3 symbols -- `.agents/lib/cli-agent-prompts.ts` — 13.7 KB, 5 symbols -- `.agents/codex-cli.ts` — 7.0 KB, 0 symbols -- `.agents/codebuff-local-cli.ts` — 5.9 KB, 0 symbols - -### `.github` - -- `.github/workflows/cli-release-build.yml` — 15.3 KB, 0 symbols -- `.github/workflows/cli-release-staging.yml` — 8.9 KB, 0 symbols -- `.github/workflows/ci.yml` — 7.3 KB, 0 symbols -- `.github/knowledge.md` — 5.4 KB, 0 symbols -- `.github/workflows/cli-release-prod.yml` — 4.9 KB, 0 symbols - -### `openbuff.d.example` - -- `openbuff.d.example/providers.json` — 19.0 KB, 0 symbols -- `openbuff.d.example/routes.json` — 2.1 KB, 0 symbols -- `openbuff.d.example/hooks.json` — 1.2 KB, 0 symbols -- `openbuff.d.example/indexing.json` — 554 B, 0 symbols - -### `LICENSE` - -- `LICENSE` — 11.1 KB, 0 symbols - -### `.bin` - -- `.bin/bun` — 8.5 KB, 0 symbols - -### `README.zh-CN.md` - -- `README.zh-CN.md` — 8.1 KB, 0 symbols - -### `README.md` - -- `README.md` — 8.0 KB, 0 symbols - -### `WINDOWS.md` - -- `WINDOWS.md` — 7.3 KB, 0 symbols - -### `CONTRIBUTING.md` - -- `CONTRIBUTING.md` — 5.6 KB, 0 symbols - -### `CODE_OF_CONDUCT.md` - -- `CODE_OF_CONDUCT.md` — 4.5 KB, 0 symbols - -### `eslint.config.js` - -- `eslint.config.js` — 4.0 KB, 0 symbols - -### `AGENTS.md` - -- `AGENTS.md` — 3.6 KB, 0 symbols - -### `ROUTER.md` - -- `ROUTER.md` — 3.3 KB, 0 symbols - -### `package.json` - -- `package.json` — 2.6 KB, 0 symbols - -### `INFISICAL_SETUP_GUIDE.md` - -- `INFISICAL_SETUP_GUIDE.md` — 2.5 KB, 0 symbols - -### `.env.example` - -- `.env.example` — 1.7 KB, 0 symbols - -### `tsconfig.json` - -- `tsconfig.json` — 839 B, 0 symbols - -### `SECURITY.md` - -- `SECURITY.md` — 520 B, 0 symbols - -### `.gitignore` - -- `.gitignore` — 487 B, 0 symbols - -### `.vscode` - -- `.vscode/settings.json` — 438 B, 0 symbols - -### `bunfig.toml` - -- `bunfig.toml` — 432 B, 0 symbols - -### `.prettierrc` - -- `.prettierrc` — 389 B, 0 symbols - -### `tsconfig.base.json` - -- `tsconfig.base.json` — 386 B, 0 symbols - -### `test` - -- `test/setup-scm-loader.ts` — 332 B, 1 symbols - -### `knowledge.md` - -- `knowledge.md` — 287 B, 0 symbols - -### `.e2e-scratch` - -- `.e2e-scratch/widget.ts` — 275 B, 3 symbols -- `.e2e-scratch/browser-agent-note.txt` — 4 B, 0 symbols - -### `NOTICE` - -- `NOTICE` — 156 B, 0 symbols - -### `openbuff.json.example` - -- `openbuff.json.example` — 118 B, 0 symbols - -### `.envrc` - -- `.envrc` — 30 B, 0 symbols - -### `.bun-version` - -- `.bun-version` — 7 B, 0 symbols - -## Most-imported files (likely key modules) - -| in-degree | file | -| --------- | -------------------------------------------------------------- | -| 101 | `packages/agent-runtime/src/__tests__/rewrite-symbol.test.ts` | -| 82 | `common/src/util/messages.ts` | -| 77 | `cli/src/utils/arrays.ts` | -| 75 | `common/src/types/bun-test.d.ts` | -| 66 | `cli/src/__tests__/release/proxy-http-get.test.ts` | -| 63 | `common/src/util/error.ts` | -| 59 | `common/src/tools/params/utils.ts` | -| 54 | `sdk/e2e/utils/event-collector.ts` | -| 49 | `cli/src/utils/message-block-helpers.ts` | -| 45 | `cli/src/hooks/use-theme.tsx` | -| 45 | `sdk/src/tools/filesystem-authority.ts` | -| 44 | `packages/agent-runtime/src/__tests__/main-prompt.test.ts` | -| 42 | `sdk/src/provider-config.ts` | -| 40 | `common/src/util/plan-artifacts.ts` | -| 38 | `packages/agent-runtime/src/tools/handlers/tool/write-file.ts` | -| 37 | `cli/src/project-files.ts` | -| 35 | `common/src/util/content-hash.ts` | -| 34 | `common/src/testing/mocks/timers.ts` | -| 33 | `agents/base2/base2.ts` | -| 33 | `packages/indexer/src/index-manager.ts` | -| 32 | `scripts/test-canopywave-long.ts` | -| 31 | `.e2e-scratch/widget.ts` | -| 30 | `common/src/types/session-state.ts` | -| 28 | `cli/scripts/build-binary.ts` | -| 28 | `sdk/src/run.ts` | - -## Cross-directory dependencies (architectural layering) - -| count | from → to | -| ----- | ----------------------------- | -| 275 | `packages` → `common` | -| 112 | `cli` → `packages` | -| 112 | `cli` → `common` | -| 101 | `sdk` → `common` | -| 68 | `agents` → `common` | -| 65 | `cli` → `sdk` | -| 43 | `common` → `packages` | -| 31 | `packages` → `cli` | -| 23 | `evals` → `sdk` | -| 23 | `cli` → `scripts` | -| 22 | `sdk` → `packages` | -| 20 | `packages` → `sdk` | -| 19 | `agents` → `agents-graveyard` | -| 19 | `agents-graveyard` → `common` | -| 19 | `agents` → `sdk` | -| 17 | `sdk` → `cli` | -| 17 | `evals` → `cli` | -| 15 | `agents` → `cli` | -| 15 | `cli` → `.e2e-scratch` | -| 15 | `evals` → `common` | -| 14 | `evals` → `scripts` | -| 13 | `agents` → `packages` | -| 10 | `agents` → `.e2e-scratch` | -| 10 | `scripts` → `packages` | -| 9 | `common` → `cli` | -| 8 | `.agents` → `packages` | -| 8 | `scripts` → `sdk` | -| 7 | `packages` → `agents` | -| 7 | `scripts` → `cli` | -| 6 | `cli` → `agents` | - -## Shard sizing hint - -Total indexed source: **11.4 MB** across **40** top-level directories. - -When sharding for an audit, aim for ~5–15 files per shard. Use the table above to group small dirs together and split huge dirs (e.g. split `src/` by subdirectory). diff --git a/.agents/sessions/audit-cli-next-level-2026-07/findings/distribution-quality.md b/.agents/sessions/audit-cli-next-level-2026-07/findings/distribution-quality.md deleted file mode 100644 index ff909de845..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/findings/distribution-quality.md +++ /dev/null @@ -1,81 +0,0 @@ -## [HIGH] Correctness / State mutation / Error handling — cli/release/index.js:639 — Auto-update can silently terminate a healthy in-progress session - -- **Risk:** The wrapper removes the child exit listener and kills the running CLI before the replacement has downloaded or validated, then swallows update failures, so a transient network, extraction, or filesystem error can end the user's active session with no actionable error while the old binary was still usable. -- **Fix:** Download, integrity-check, smoke-test, and atomically stage the replacement before asking the running CLI to exit, and preserve/relaunch the old binary with a visible warning if activation fails. -- **Evidence:** `checkForUpdates()` calls `runningProcess.removeListener('exit', exitListener)` at line 653 and `runningProcess.kill('SIGTERM')` at line 661 before `await downloadBinary(latestVersion)` at line 674; its catch at lines 694-696 is only `// Ignore update failures`. The staging wrapper duplicates this flow, while `release-wrapper.test.ts` has no update-failure/rollback scenario. -- **Confidence:** High — Evidence. - -## [HIGH] Security — .github/workflows/cli-release-staging.yml:3 — PR-controlled code can enter a write-and-publish workflow with inherited secrets - -- **Risk:** A same-repository pull request whose author adds `[codecane]` to its title can cause PR-head code and repository scripts to run in a workflow with `contents: write`, the release token, inherited client secrets, and later npm publication authority. -- **Fix:** Remove `pull_request` as a release trigger, require an environment-protected manual dispatch or trusted post-merge branch, pin the reviewed commit, and give each job only its minimum token/secrets. -- **Evidence:** The workflow listens to `pull_request` at lines 3-7, grants `contents: write` at lines 13-14, gates only on PR title text at lines 19-21, checks out `github.event.pull_request.head.sha` with a token at lines 25-28, and passes `secrets: inherit` into the reusable binary build at lines 118-127; the publish job then uses `NPM_TOKEN` at lines 232-237. -- **Confidence:** High — Evidence (fork PRs normally lose secrets, but same-repository PR heads remain in scope). - -## [MEDIUM] Security / Dependency hygiene — cli/release/index.js:448 — Downloaded executables have no independent integrity or provenance verification - -- **Risk:** The updater automatically executes a GitHub-release tarball after TLS download and extraction without checking a release manifest, digest, signature, or attestation, so a compromised release asset/CDN path or authorized release account becomes direct arbitrary-code execution on every updating client. -- **Fix:** Publish per-platform SHA-256 manifests plus signed provenance/Sigstore attestations, embed or fetch a trusted verification key/manifest, verify every archive before extraction, and expose the verified digest in diagnostics. -- **Evidence:** `downloadUrl` is constructed at lines 448-451 (and is runtime-overridable by `OPENBUFF_DOWNLOAD_BASE`), the response is streamed directly through gunzip and `tar.x()` at lines 498-507, and the resulting file is renamed into the executable cache at lines 527-546; `cli-release-build.yml` creates and uploads tarballs at lines 278-298 and 431-443 without checksum/signing/SBOM steps. -- **Confidence:** High — Evidence. - -## [MEDIUM] State mutation / Error handling — cli/release/postinstall.js:7 — Package installation destroys the offline fallback before replacement is available - -- **Risk:** Every npm install/upgrade deletes the cached working executable immediately, so the next invocation becomes network-dependent and cannot roll back or continue offline if the registry, GitHub, proxy, or new release is unavailable. -- **Fix:** Retain versioned binaries and metadata, download into a new slot on first launch, atomically switch only after verification/smoke, and provide `openbuff update --rollback` plus an explicit cache-prune command. -- **Evidence:** The postinstall script says `Clean up managed binaries so the wrapper downloads a fresh copy` and unconditionally calls `fs.unlinkSync(...)` at lines 7-18; `ensureBinaryExists()` then exits when it cannot resolve/download latest at `cli/release/index.js:606-636`. Staging repeats the destructive postinstall behavior. -- **Confidence:** High — Evidence. - -## [MEDIUM] Test coverage gaps / Correctness — .github/workflows/cli-release-prod.yml:81 — Release publication is not gated by repository validation or an exact version assertion - -- **Risk:** The production workflow can tag and publish a compiled binary from a repository state that fails the normal architecture/typecheck gate, and its smoke step accepts any successfully printed version rather than proving it equals the release version. -- **Fix:** Make release depend on the exact commit's required CI checks, run the env architecture check and CLI tests in the reusable workflow, and assert `"$($BIN --version)" = "${{ inputs.new-version }}"` before packaging. -- **Evidence:** Production goes directly from version bump to the reusable build at lines 81-90; the reusable workflow compiles at `cli-release-build.yml:232-241` and merely runs `"$BIN" --version` at lines 269-276/423-429. At 2026-07-11 23:22 EAT, `bun scripts/check-env-architecture.ts` failed on direct `process.env` use at `cli/src/native/ripgrep.ts:24`, while `bun run --cwd cli typecheck` passed, demonstrating a release-relevant repository gate that this workflow does not run. -- **Confidence:** High — Evidence, time-of-check stated because the dirty worktree was changing concurrently. - -## [MEDIUM] API/ABI contract breaks / Correctness — cli/release/index.js:274 — Version sources and tests mask release-wrapper/binary drift - -- **Risk:** Local wrapper tests and source invocations report the private workspace version instead of the published wrapper version, allowing release metadata, compiled binaries, and wrapper behavior to drift without a failing gate or trustworthy diagnostic. -- **Fix:** Define one release-version source, make source and packed-package tests exercise the actual published layout, assert wrapper/binary/npm/GitHub tag equality, and never keep a stale executable at the canonical development path. -- **Evidence:** `getLocalPackageVersion()` checks `../package.json` before its own package at lines 274-288; `release-wrapper.test.ts:15-18` deliberately reads `cli/package.json`, so both wrapper tests passed while source `node cli/release/index.js --version`, staging, and `cli/bin/openbuff --version` all reported `1.0.0`. At 2026-07-11 23:22 EAT, `cli/release/package.json:3` was `1.2.4`, `cli/release-staging/package.json:3` was `1.0.420`, `cli/package.json:3` was `1.0.0`, and the built binary was dated 2026-06-29 while current source files were dated 2026-07-11. -- **Confidence:** High — Evidence. - -## [MEDIUM] Security / API contract breaks — common/src/analytics.ts:59 — Production runtime telemetry has no user-facing disable control - -- **Risk:** Local/BYOK users can have runtime events sent to PostHog whenever production analytics configuration is embedded, but the code and documentation provide controls to increase telemetry detail rather than a documented opt-out, weakening the product's local-first privacy contract. -- **Fix:** Default telemetry off or add a clearly documented `OPENBUFF_TELEMETRY=off`/`DO_NOT_TRACK` gate honored by both runtime and wrapper paths, expose current telemetry state in `openbuff doctor`, and publish the exact event/property policy. -- **Evidence:** `common/src/analytics.ts` lazily creates a real PostHog client in production and captures events at lines 59-89; the production binary workflow exports client secrets/env at `cli-release-build.yml:158-173`; agent runtime calls it from `packages/agent-runtime/src/main-prompt.ts:67-88`. `docs/environment-variables.md:9-10` documents only `CODEBUFF_FULL_TELEMETRY*`, which disables sampling and may send full log payloads, not a disable switch. Separately, `cli/src/utils/analytics.ts:140-160` is now a no-op, making `common/src/analytics.knowledge.md:73-114` stale about CLI alias/identify behavior. -- **Confidence:** High — Evidence. - -## [MEDIUM] Test coverage gaps / Platform parity — .github/workflows/cli-release-build.yml:40 — Linux ARM64 is published without executing its binary - -- **Risk:** ABI, loader, native dependency, tree-sitter, and startup regressions specific to Linux ARM64 can pass release automation and reach users because that artifact is never booted. -- **Fix:** Add a native ARM64 runner or QEMU/container execution lane that runs the same version, tree-sitter, boot, and ripgrep smoke suite before upload. -- **Evidence:** The `linux-arm64` matrix entry explicitly sets `smoke_test: false` at lines 40-46, and the smoke job is skipped when false at lines 258-276; the checked-in `sdk/vendor/ripgrep/arm64-linux/rg` is an AArch64 dynamically linked ELF, making loader/platform validation material. -- **Confidence:** High — Evidence. - -## [LOW] Error handling / Performance — cli/release/http.js:144 — Redirect following is unbounded and incompletely tested - -- **Risk:** A bad proxy/CDN redirect loop can keep allocating requests and sockets until failure, while 307/308 responses are not followed even though artifact hosts may use them. -- **Fix:** Implement a small redirect budget, validate `Location`, support the intended redirect status set, preserve timeout/cancellation across the chain, and test loops, missing locations, downgrade attempts, proxy auth, custom CA, and `NO_PROXY` edge cases. -- **Evidence:** `httpGet()` recursively calls itself for 301/302 at lines 144-155 with no redirect counter; `proxy-http-get.test.ts:161-239` covers exactly one 302 redirect and has no loop/limit, TLS failure, CA, timeout, or `NO_PROXY` assertions. -- **Confidence:** High — Evidence. - -## [MEDIUM] API/ABI contract breaks / Documentation — cli/release/README.md:23 — Published usage documents a positional project directory that the CLI parses as a prompt - -- **Risk:** Users following `openbuff [project-directory]` can accidentally send a filesystem path as the initial agent prompt while Openbuff operates in the current directory, potentially editing the wrong project. -- **Fix:** Document `cd && openbuff` or `openbuff --cwd ` everywhere, add an end-to-end docs contract test, and consider rejecting a directory-looking positional argument with migration guidance. -- **Evidence:** The published README says `openbuff [project-directory]` and claims it changes the directory at lines 23-31 (staging says the same at lines 25-33), while `cli/src/index.tsx:125-140` defines only `--cwd ` plus positional `[prompt...]`, and lines 152-163 join all positional args into `initialPrompt`. -- **Confidence:** High — Evidence. - -## Strengths observed - -- The supported artifact matrix is explicit and now covers Linux x64/ARM64, macOS Intel/current, macOS Intel legacy, macOS Apple Silicon, and Windows x64; macOS legacy deployment targets are checked with `vtool`. -- The compiled-binary gate goes beyond `--help`/`--version`: `smoke-binary.ts` holds the process open, checks a visible boot signal, and separately validates tree-sitter initialization. The existing 2026-06-29 Linux binary passed this smoke at audit time. -- Proxy support is factored into a shared production/staging HTTP helper, uses CONNECT tunneling for HTTPS, honors upper/lowercase proxy variables and `NO_PROXY`, applies request timeouts, and has parity tests for the two wrappers. -- Runtime fatal handling deliberately restores raw mode, alternate-screen state, mouse/focus modes, bracketed paste, and cursor visibility; native crash diagnostics include platform/hardware/target details. -- Production/staging wrapper HTTP files were byte-identical, and their larger entrypoints were close enough in structure for direct parity comparison; targeted wrapper/proxy/analytics/error tests passed 60/60 at audit time. - -## Coverage / files actually read - -All eight domains were evaluated: Security, Correctness, State mutation, Error handling, Performance, Dependency hygiene, Test coverage gaps, and API/ABI contract breaks. Read or selectively inspected: `cli/release/{index.js,http.js,postinstall.js,package.json,README.md}`, all corresponding `cli/release-staging/*` files, `cli/scripts/{build-binary.ts,smoke-binary.ts,release.ts}`, `cli/package.json`, root `package.json`, `.bun-version`, `bunfig.toml`, relevant `bun.lock` entries, checked-in Linux ripgrep binary metadata, all four CLI/release workflows plus `ci.yml`, `nightly-e2e.yml`, and `sdk-release.yml`; `cli/src/index.tsx`, logger/analytics/error handling/error UI files; common analytics core/dispatcher/sampling/contracts/knowledge and their tests; wrapper/proxy/error tests; and the listed top-level/CLI/Windows/development/testing/environment/local-mode/configuration/security/contribution docs. Existing `.agents/sessions/**` audit reports/findings were not read. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/findings/interaction-commands.md b/.agents/sessions/audit-cli-next-level-2026-07/findings/interaction-commands.md deleted file mode 100644 index c899289ee5..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/findings/interaction-commands.md +++ /dev/null @@ -1,83 +0,0 @@ -# Interaction and commands audit findings - -## [MEDIUM] Correctness — cli/src/components/multiline-input.tsx:940 — Ctrl+B can move the cursor before the start of input - -- **Risk:** Pressing Ctrl+B at cursor position 0 emits `cursorPosition: -1`, which can desynchronize editing/rendering and make subsequent slice-based edits operate from the end of the string. -- **Fix:** Clamp the Ctrl+B target with `Math.max(0, cursorPosition - 1)` or route it through the existing clamping `moveCursor` helper, and add a component-level boundary test. -- **Evidence:** `handleNavigationKeys` directly calls `onChange({ cursorPosition: cursorPosition - 1 })` at lines 940-948 while the ordinary left-arrow path at lines 962-966 uses `moveCursor`; `cli/src/components/__tests__/multiline-input.test.tsx` contains no Ctrl+B/Ctrl+F boundary case. -- **Confidence:** High — Evidence. - -## [MEDIUM] Correctness — cli/src/components/command-palette-screen.tsx:170 — Palette search permanently excludes files after the first 50 - -- **Risk:** In repositories with more than 50 flattened entries, Ctrl+P cannot find or attach any file outside the tree's first 50 entries even after the user types an exact query. -- **Fix:** Keep an uncapped flattened search corpus and apply `MAX_EMPTY_FILE_ITEMS` only to the empty-query result, then cap matched render results after scoring. -- **Evidence:** `allEntries` always calls `buildEntries(..., LAYOUT.MAX_EMPTY_FILE_ITEMS)` at lines 170-175, so the later non-empty-query filter at lines 177-191 never sees later files; the comment says the cap is only for empty queries, while `command-palette-screen.test.ts` only verifies that `buildEntries` respects a supplied cap. -- **Confidence:** High — Evidence. - -## [MEDIUM] Performance — cli/src/hooks/use-suggestion-engine.ts:641 — Every new @ session rebuilds the project tree - -- **Risk:** Typing `@` after closing a prior mention session triggers a fresh traversal/stat of up to 10,000 files, causing avoidable input latency and filesystem churn on large repositories. -- **Fix:** Serve the existing `fileTree` immediately, refresh through a TTL/versioned background cache or index-backed path list, and invalidate from known filesystem/index events rather than every inactive-to-active transition. -- **Evidence:** the effect at lines 665-671 invokes `refreshFileSuggestions()` whenever `mentionContext.active` becomes true; that calls `getProjectFileTree` at lines 641-657, whose implementation in `common/src/project-file-tree.ts:140-229` breadth-first reads directories and stats entries up to `DEFAULT_MAX_FILES = 10_000`. -- **Confidence:** High — Evidence. - -## [MEDIUM] Error handling — cli/src/commands/router.ts:76 — Interactive bash jobs have no timeout or cancellation path - -- **Risk:** A command such as `!tail -f`, a hung subprocess, or a prompt-waiting program can remain running indefinitely, leave a permanent pending card, and cannot be stopped by the CLI's normal response-interrupt controls. -- **Fix:** Run interactive bash commands as cancellable jobs with an AbortSignal/job id, expose a Stop action and bounded default timeout, and reserve an explicit detach/background command for intentional long-running work. -- **Evidence:** `runBashCommand` calls `runTerminalCommand` with `process_type: 'SYNC'` and `timeout_seconds: -1` at lines 76-82 and retains only a UI message id, not a process/job handle; `PendingBashMessage` renders status text only, and `bash-command.test.ts` asserts state transitions but has no timeout/cancel test. -- **Confidence:** High — Evidence. - -## [MEDIUM] Security — cli/src/commands/command-registry.ts:352 — Convenience git commands interpolate raw shell arguments - -- **Risk:** `/diff` and `/changes` look like constrained git helpers but accept shell metacharacters and command substitution, so pasted or suggested arguments can execute unrelated local commands. -- **Fix:** Parse an allowlist of git flags/pathspecs and invoke git with an argv-based process API, or clearly route advanced users to explicit bash mode instead of concatenating strings. -- **Evidence:** `/diff` builds `git diff ${trimmedArgs}` at lines 352-359 and `/changes` builds `git status ${trimmedArgs}` at lines 363-371 before passing the string to the shell-backed `runBashCommand`; their registry descriptions promise only diff/status behavior in `cli/src/data/slash-commands.ts:169-177`. -- **Confidence:** High — Evidence. - -## [MEDIUM] State mutation — cli/src/hooks/use-input-history.ts:65 — Cross-terminal history writes can lose prompts - -- **Risk:** Two Openbuff processes saving near-simultaneously can both read the same history snapshot and overwrite each other, despite the code explicitly attempting cross-terminal history support. -- **Fix:** Use an append-only journal or an inter-process lock plus atomic temp-file rename, then reload/deduplicate after the committed write. -- **Evidence:** `saveToHistory` re-reads disk then constructs `[...diskHistory, message]` at lines 65-75, while `saveMessageHistory` rewrites the entire JSON file synchronously at `cli/src/utils/message-history.ts:106-123`; there is no lock, compare-and-swap, or append operation. -- **Confidence:** High — Evidence. - -## [MEDIUM] Performance — cli/src/utils/pending-attachments.ts:347 — “Background” attachment loading still blocks the renderer thread - -- **Risk:** Attaching a large directory or many files can freeze keyboard input because the deferred callback performs synchronous `readdirSync`, sorting, `statSync`, `readFileSync`, UTF-8 conversion, and store updates on the main event loop. -- **Fix:** Move attachment inspection to async filesystem APIs or a worker, cap work before sorting/reading, and report progressive/cancellable processing for batches. -- **Evidence:** line 347 describes asynchronous reading via `setTimeout`, but lines 354-438 execute synchronous directory enumeration, full sorting, stats, reads up to 1 MB, binary scanning, and conversion; the send route blocks while any file remains `processing` at `cli/src/commands/router.ts:467-472`. -- **Confidence:** High — Evidence. - -## [LOW] Security — cli/src/utils/clipboard-image.ts:18 — Clipboard images accumulate as persistent plaintext temp files - -- **Risk:** Screenshots and copied images can contain secrets and remain indefinitely in a predictable shared temp directory after removal, sending, or process exit. -- **Fix:** Use a per-process mode-0700 temp directory with restrictive file permissions and delete owned files on attachment removal, successful capture/send, startup TTL cleanup, and graceful exit. -- **Evidence:** `getClipboardTempDir` creates `${os.tmpdir()}/codebuff-clipboard-images` at lines 18-23 and platform readers write timestamped PNGs there (for example lines 175-201); repository search finds no production cleanup for that directory or its files. -- **Confidence:** High — Evidence. - -## [LOW] Test coverage gaps — cli/src/components/help-banner.tsx:59 — Help omits major implemented interaction shortcuts - -- **Risk:** Users cannot discover Ctrl+P palette, Ctrl+R prompt search, Ctrl+V attachments, Shift+Enter, Tab completion/mode switching, or PageUp/PageDown from the built-in help, weakening learnability and making shortcut behavior feel inconsistent. -- **Fix:** Generate help rows from a shared shortcut registry used by keyboard classification, include platform-specific labels, and add a contract test that every user-facing global action has help metadata. -- **Evidence:** `HelpBanner` lists only Ctrl+C/Esc, Ctrl+J/Opt+Enter, arrows, Ctrl+T, `/`, `@`, and `!` at lines 59-105, while `keyboard-actions.ts` implements Ctrl+P, Ctrl+R, Ctrl+V, Tab/Shift+Tab, and PageUp/PageDown and `MultilineInput` implements Shift+Enter and additional editing bindings. -- **Confidence:** High — Evidence. - -## [LOW] API/ABI contract breaks — cli/src/utils/clipboard-image.ts:18 — Clipboard storage retains the legacy Codebuff namespace - -- **Risk:** Openbuff writes new runtime artifacts under `codebuff-clipboard-images`, complicating support, cleanup, migration expectations, and privacy documentation for a renamed CLI. -- **Fix:** Move to an Openbuff-named temp namespace with one-time cleanup of the legacy directory and document the temporary-file lifecycle. -- **Evidence:** the current directory literal is `codebuff-clipboard-images` at line 19, while `docs/architecture.md` states `openbuff` is the primary CLI identity and retained Codebuff compatibility aliases are intentionally narrow. -- **Confidence:** High — Evidence. - -## Strengths observed - -- Keyboard classification is centralized and extensively unit-tested, with explicit priority ordering for modal input, streaming interruption, menus, history, queueing, and exit behavior. -- Attachment processing now records provenance/completeness, blocks mandatory sensitive paths for project file mentions, caps previews, detects binary content, and warns the model when supplied context is incomplete. -- Slash commands have a unified registry, aliases, typo suggestions, generated skill/game presets, and a command palette plus prompt-history search rather than relying only on memorized command names. -- Image handling has per-file and aggregate size controls, compression attempts, terminal fallbacks, and provider-boundary normalization tests. -- Dependency versions for the core terminal renderer are pinned consistently in `cli/package.json`, and the interaction surface has substantial focused test coverage. - -## Coverage / files actually read - -Evaluated all eight audit domains across the manifest's five subshards: Security, Correctness, State mutation, Error handling, Performance, Dependency hygiene, Test coverage gaps, and API/ABI contract breaks. Read or inspected bounded ranges/diffs/search evidence from: `cli/src/chat.tsx`; `cli/src/components/chat-input-bar.tsx`, `multiline-input.tsx`, `input-cursor.tsx`, `input-mode-banner.tsx`, `help-banner.tsx`, `suggestion-menu.tsx`, `command-palette-screen.tsx`, `prompt-history-search-screen.tsx`, `pending-bash-message.tsx`, `pending-attachments-banner.tsx`, attachment/image cards and blocks; `cli/src/hooks/use-chat-input.ts`, `use-chat-keyboard.ts`, `use-suggestion-engine.ts`, `use-input-history.ts`, `use-path-tab-completion.ts`, `use-clipboard.ts`, `use-send-message.ts`, and `hooks/helpers/send-message.ts`; `cli/src/utils/keyboard-actions.ts`, `path-completion.ts`, `chat-history.ts`, `message-history.ts`, `bash-context-processor.ts`, `bash-messages.ts`, `input-modes.ts`, `clipboard.ts`, `clipboard-image.ts`, `image-handler.ts`, `image-processor.ts`, `pending-attachments.ts`, image display/thumbnail/terminal helpers; `cli/src/data/slash-commands.ts`; `cli/src/commands/command-registry.ts`, `router.ts`, `router-utils.ts`, `help.ts`, `image.ts`, prompt/plan/index/init/info command integrations; `cli/src/state/chat-store.ts`; `cli/src/types/store.ts`, `chat.ts`, and send-message contract; `common/src/project-file-tree.ts`, engine profiles and game presets; `sdk/src/impl/chatgpt-backend-fetch.ts`; the manifest-listed keyboard, suggestions/history, command/bash, clipboard/image, attachment, and send-message tests; and `docs/architecture.md`, `docs/request-flow.md`, `docs/agents-and-tools.md`, `docs/testing.md`, `README.md`, `WINDOWS.md`, `cli/package.json`, and relevant lockfile entries. Existing `.agents/sessions/*` audit artifacts were not read. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/findings/onboarding-config.md b/.agents/sessions/audit-cli-next-level-2026-07/findings/onboarding-config.md deleted file mode 100644 index f3d55dbe14..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/findings/onboarding-config.md +++ /dev/null @@ -1,114 +0,0 @@ -# Onboarding and configuration audit findings - -## [HIGH] correctness — cli/src/utils/openbuff-provider.ts:346 — Route edits overwrite unrelated configuration - -- **Risk:** Changing one model route or removing one provider can silently delete vision routing, failover, hooks, run limits, discovery/capability metadata, and reset indexing, turning a harmless setup action into a broad configuration rewrite. -- **Fix:** Build edits from a complete, source-aware config draft and persist only the changed keys/fragments atomically, with regression tests that assert every unrelated schema field remains byte-for-byte or semantically unchanged. -- **Evidence:** `getEditableConfig()` at lines 354-364 clones only providers/default/modes/agents/reasoning fields, while `writeMergedConfig()` at lines 346-351 calls `writeProviderConfigFile(... force: true)`; a direct temporary-directory reproduction starting with `indexing.enabled: false` and writing only providers/routes rewrote `indexing.json` to the schema default `enabled: true`, and `docs/configuration.md:33-48,135-186,231-260` documents fragmented indexing, hooks, and failover as independent supported settings. -- **Confidence:** High — Evidence. - -## [HIGH] correctness — cli/src/project-files.ts:50 — Project history identity collides on directory basename - -- **Risk:** Two common repositories such as `/work/client/app` and `/work/internal/app` share `~/.config/openbuff/projects/app`, so history, checkpoints, and `--continue` can expose or resume the other project's conversation and run state. -- **Fix:** Derive the storage key from the canonical absolute path plus a stable hash (with a readable basename prefix), record the original root in metadata, and migrate legacy basename-only directories safely. -- **Evidence:** `getProjectDataDir()` uses only `path.basename(root)` at lines 50-59; `loadMostRecentChatState()` in `cli/src/utils/run-state-storage.ts:121-145` falls back to the most recent chat under that shared directory, and no collision test exists for `project-files.ts`. -- **Confidence:** High — Evidence. - -## [HIGH] state mutation — cli/src/index.tsx:343 — Project selection does not re-bootstrap project-scoped state - -- **Risk:** Selecting a project from the first-run picker changes `cwd` and the root, but leaves direnv variables, local agents/MCP, skills, and the index initialized for the launch directory, so the first usable session can run with missing or wrong project context. -- **Fix:** Replace the partial callback with one cancellable `switchProject()` transaction that validates the path, changes root, reloads environment/config/agents/skills, swaps the index instance, resets project caches, and only then reveals chat. -- **Evidence:** `handleProjectChange()` at lines 343-368 only calls `process.chdir`, `setProjectRoot`, `resetCodebuffClient`, recents, and file-tree state; the omitted initialization occurs once at `initializeApp()` (`cli/src/init/init-app.ts:19-65`) and `initializeAgentRegistry()` / `initializeSkillRegistry()` (`cli/src/index.tsx:293-300`), with no project-switch integration test. -- **Confidence:** High — Evidence. - -## [HIGH] API/ABI contract breaks — sdk/src/agents/load-agents.ts:207 — Public loader silently stopped loading project agents by default - -- **Risk:** Existing SDK callers of `loadLocalAgents({ verbose: true })` now receive only home agents, breaking project-local workflows without a type error and contradicting the published local-agent contract. -- **Fix:** Preserve the previous default for the exported SDK API or ship the trust change as an explicit major-version contract with a new clearly named safe loader/option and migration documentation; keep CLI trust policy at the CLI boundary. -- **Evidence:** The current worktree diff adds `includeProjectAgents = false` at lines 207-222 and changes default directories from cwd/parent/home to home-only, while the same file's JSDoc at lines 146-172 still says the default searches project and parent directories; `README.md:177` and `docs/agents-and-tools.md:10-12` promise project-local `.agents`, and tests predominantly pass `agentsPath` rather than asserting the default-location contract. -- **Confidence:** High — Evidence (current uncommitted regression). - -## [HIGH] correctness — cli/src/components/button.tsx:43 — Core onboarding actions are mouse-only - -- **Risk:** A keyboard-only user can browse directories but cannot activate the project picker's `Open` action, and the same primitive makes OAuth auto-configure, disconnect, retry, and close controls inaccessible. -- **Fix:** Give `Button` focus/activation semantics and a consistent keyboard model, then add explicit project-picker shortcuts (for example Enter to open the current path and a separate key to descend) plus end-to-end keyboard tests. -- **Evidence:** `Button` lines 43-69 only handles mouse down/up; `ProjectPickerScreen` maps Enter to directory navigation at `cli/src/components/project-picker-screen.tsx:269-285` and exposes `Open` solely through `onClick` at lines 486-501, while `ChatGptConnectBanner` actions at lines 174-205 and 226-239 are also `onClick`-only. -- **Confidence:** High — Evidence. - -## [MEDIUM] security — cli/src/utils/skill-registry.ts:22 — Project skills bypass the new project-trust boundary - -- **Risk:** Opening an untrusted repository directly can load a project skill that overrides a familiar global skill name and injects repository-controlled instructions into the next `/skill:*` invocation even when project agents and MCP are not trusted. -- **Fix:** Apply the same trust decision to project `.agents/skills` and `.claude/skills`, show each skill's origin in suggestions, and require confirmation when a project skill shadows a global skill. -- **Evidence:** `initializeSkillRegistry()` always calls `loadSkills({ cwd })` at lines 22-31; `sdk/src/skills/load-skills.ts:159-178,221-234` loads project skills last so they override globals, and `createSkillCommand()` injects the selected file's full content into the user prompt at `cli/src/commands/command-registry.ts:1013-1069`, whereas `--trust-project-agents` only gates agent/MCP loading in `cli/src/index.tsx:125-133,293-300`. -- **Confidence:** Medium — Inference grounded in the prompt-loading path. - -## [MEDIUM] error handling — sdk/src/provider-config.ts:1177 — Malformed normal configs disappear without a diagnostic - -- **Risk:** A syntax or schema error in global/project `openbuff.json` is silently ignored and later presented as missing setup or missing routes, preventing users from repairing the actual file and making first-run failures look unrelated. -- **Fix:** Preserve per-source parse diagnostics in `LoadedProviderConfig`, surface them in `/provider status` and readiness errors, and add a `openbuff doctor`/repair view that names the exact file and schema path without printing secrets. -- **Evidence:** `loadProviderConfigSync()` catches and discards every non-explicit config error at lines 1182-1194 and only records successful `sourceFilePaths`; `/provider status` renders only `describeLoadedProviderConfig()` plus missing env values (`cli/src/utils/openbuff-provider.ts:259-276`), while tests only require explicit malformed config to fail (`sdk/src/__tests__/model-provider.test.ts:1559` and 1682-1695). -- **Confidence:** High — Evidence. - -## [MEDIUM] error handling — cli/src/utils/chatgpt-oauth.ts:293 — OAuth token operations have no network timeout - -- **Risk:** A stalled token endpoint can leave manual authorization or background refresh pending indefinitely, with the module-global refresh promise blocking all later refresh attempts. -- **Fix:** Attach bounded abort signals to authorization-code exchange and refresh fetches, distinguish timeout from auth failure, and always clear/abort outstanding work when the banner or callback server closes. -- **Evidence:** `exchangeChatGptCodeForTokens()` calls `fetch()` without `signal` at lines 293-305 and `refreshChatGptOAuthToken()` does the same at `sdk/src/credentials.ts:235-247`; the five-minute callback timer only calls `stopChatGptOAuthServer()` (`chatgpt-oauth.ts:164-168`) and does not abort the fetch, while `chatGptRefreshPromise` remains occupied until the fetch settles (`credentials.ts:221-292`). -- **Confidence:** High — Evidence. - -## [MEDIUM] error handling — sdk/src/agents/load-agents.ts:229 — Broken agent modules vanish from validation diagnostics - -- **Risk:** Syntax errors, import-time exceptions, missing IDs, and missing environment variables make a custom agent disappear with no actionable TUI error because the CLI loads with `verbose: false` and only schema-valid imported agents reach validation. -- **Fix:** Collect file-aware load diagnostics for every skipped module and merge them with schema diagnostics, while still loading healthy agents. -- **Evidence:** Import and pre-validation failures at lines 229-280 only log when `verbose` is true, but `initializeAgentRegistry()` uses `{ verbose: false, validate: true }` at `cli/src/utils/local-agent-registry.ts:63-76`; `cli/src/__tests__/integration/local-agents.test.ts:187-217,915-993` merely verifies that healthy agents continue loading and never asserts that the broken file is reported. -- **Confidence:** High — Evidence. - -## [MEDIUM] error handling — cli/src/commands/init.ts:61 — `/init` can throw out of the command router after partial scaffolding - -- **Risk:** A read-only directory, disk-full condition, or failed directory creation can abort the slash command after some files were created, without a coherent repair message or rollback. -- **Fix:** Preflight writability, perform scaffold writes through a single error-returning operation using temporary files, and convert partial failures into explicit per-path TUI guidance with a safe retry. -- **Evidence:** Knowledge and directory creation at lines 61-90 are outside the per-type-file try/catch, `routeUserPrompt()` directly awaits command handlers without a catch at `cli/src/commands/router.ts:443-462`, and tests explicitly expect knowledge/mkdir failures to throw at `cli/src/commands/__tests__/init.test.ts:289-331` despite naming them graceful-handling cases. -- **Confidence:** High — Evidence. - -## [MEDIUM] performance — cli/src/utils/directory-browser.ts:15 — First-run browsing performs synchronous filesystem fan-out on the render thread - -- **Risk:** Opening a large home directory, network mount, or slow filesystem can freeze the TUI while every entry is synchronously statted and every child receives another `.git` stat. -- **Fix:** Enumerate asynchronously in cancellable batches, render partial results, limit concurrent git detection, and cache directory metadata while the picker is open. -- **Evidence:** `getDirectories()` synchronously calls `readdirSync`, `statSync(fullPath)`, and `hasGitDirectory(fullPath)` for every item at lines 15-49, and `useDirectoryBrowser()` invokes it during render memoization at `cli/src/hooks/use-directory-browser.ts:41-47`; tests cover small temporary directories but no high-cardinality or slow-I/O behavior. -- **Confidence:** Medium — Inference from synchronous hot-path I/O. - -## [MEDIUM] test coverage gaps — cli/src/**tests**/cli-args.test.ts:15 — CLI argument tests do not exercise the production parser - -- **Risk:** Flags central to onboarding (`--cwd`, `--continue`, `--plan`, `--local`, and the new `--trust-project-agents`) can regress while the argument suite remains green. -- **Fix:** Export a pure production `parseArgs(argv)` helper and test its exact Commander definition, help text, interactions, invalid paths, and trust behavior rather than rebuilding a smaller test-only command. -- **Evidence:** `parseTestArgs()` lines 15-51 constructs a separate command named `codecane` with only `--agent` and `--clear-logs`, while production `parseArgs()` defines the omitted flags at `cli/src/index.tsx:107-164`; the targeted suite passed all 12 tests without touching those production branches. -- **Confidence:** High — Evidence. - -## [LOW] dependency hygiene — cli/src/index.tsx:25 — Runtime import is not declared by the CLI package - -- **Risk:** `picocolors` currently resolves through workspace/transitive installation, so dependency graph changes or isolated packaging can break CLI startup unexpectedly. -- **Fix:** Add `picocolors` as a direct pinned CLI dependency and add an isolated-package/binary smoke check that rejects undeclared runtime imports. -- **Evidence:** `cli/src/index.tsx:25` imports `picocolors`, but `cli/package.json:34-71` does not declare it and repository package manifests contain no direct `picocolors` entry. -- **Confidence:** High — Evidence. - -## Strengths observed - -- Provider schemas reject non-HTTP(S) endpoints, require HTTPS when an API key is attached, and restrict plain HTTP to localhost (`sdk/src/provider-config.ts:177-221,253-293`). -- Model discovery has a 30-second cancellable timeout and avoids sending provider authorization to cross-origin custom endpoints by default (`sdk/src/model-discovery.ts:140-202,313-352`). -- OAuth uses PKCE, state validation, a loopback-only callback listener, escaped callback HTML, sanitized exchange errors, and owner-only credential permissions (`cli/src/utils/chatgpt-oauth.ts:72-106,156-237`; `sdk/src/credentials.ts:34-53,150-184`). -- Provider readiness produces actionable missing-route, missing-key, and missing-OAuth messages before a send (`cli/src/utils/openbuff-provider.ts:1176-1227`; `sdk/src/provider-config.ts:1577-1619`). -- Agent validation is local by default, filters invalid definitions before runtime, and the provider/config cache has meaningful invalidation coverage. -- The selected targeted suite completed with 216 passing tests and no failures across CLI args, project/path helpers, init, provider setup, OAuth sanitization, SDK credentials, local-agent loading, and validation. - -## Coverage / files actually read - -- **OC-1 startup:** `cli/src/index.tsx`, `cli/src/init/init-app.ts`, `cli/src/init/init-direnv.ts`, `cli/src/project-files.ts`, `cli/src/utils/env.ts`, `cli/src/utils/create-run-config.ts`, and the listed startup/argument/direnv tests by direct read or targeted symbol review. -- **OC-2 project selection:** `project-picker-screen.tsx`, `use-directory-browser.ts`, `directory-browser.ts`, `project-picker.ts`, `recent-projects.ts`, `use-path-tab-completion.ts`, `path-completion.ts`, `selectable-list.tsx`, `use-searchable-list.ts`, and all listed helper tests. -- **OC-3 init:** `commands/init.ts`, the init branch of `command-registry.ts`, the generator script and source-template surface, plus init and generated-source test references. -- **OC-4 settings/env/theme:** `settings.ts`, `auth.ts`, `theme-config.ts`, `use-theme.tsx`, relevant `chat-store.ts`, CLI/common/SDK env helpers and types, and their listed tests/docs. -- **OC-5 provider UX/readiness:** provider picker, model route picker, relevant `chat.tsx` and command/router branches, `openbuff-provider.ts`, send-readiness hooks/helpers, selectable/searchable primitives, and provider/router/readiness tests. -- **OC-6 provider SDK:** schema, load/merge/cache/write and model-resolution ranges of `provider-config.ts`; discovery implementation; model-provider integration ranges; example JSON files; and the relevant model-provider tests. -- **OC-7 OAuth:** CLI OAuth utility/banner/input routing, constants, credential storage, backend fetch/request transformation and OAuth policy ranges in `llm.ts`, plus listed OAuth/credential/backend tests. -- **OC-8 validation:** CLI registry and validation UX helpers, SDK agent loading/validation, common validator/schema ranges, and listed CLI/SDK/common tests (including `sdk/src/__tests__/load-agents.test.ts` as corroboration). -- **Docs/examples:** `README.md`, `.env.example`, `docs/configuration.md`, `docs/local-mode.md`, `docs/environment-variables.md`, `docs/request-flow.md`, `docs/development.md` references, `docs/agents-and-tools.md`, `docs/openbuff-provider-model-setup-ux.md`, and all four example configuration files. -- Existing `.agents/sessions/**` and prior audit finding/report files were not read. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/findings/presentation-quality.md b/.agents/sessions/audit-cli-next-level-2026-07/findings/presentation-quality.md deleted file mode 100644 index 40f88876f1..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/findings/presentation-quality.md +++ /dev/null @@ -1,89 +0,0 @@ -# Presentation-quality audit findings - -## [MEDIUM] Correctness — cli/src/components/command-palette-screen.tsx:172 — Palette search can never find files beyond the first 50 - -- **Risk:** In repositories with more than 50 flattened paths, Ctrl+P search silently excludes every later file even after the user types an exact filename, making the primary navigation feature unreliable at realistic monorepo scale. -- **Fix:** Keep the empty-query display capped, but build/search the complete flattened path index (or query the index lazily) and only cap the final rendered matches. -- **Evidence:** `allEntries` is permanently built with `buildEntries(slashCommands, fileTree, LAYOUT.MAX_EMPTY_FILE_ITEMS)` where `MAX_EMPTY_FILE_ITEMS = 50`, and `filteredEntries` only scores that already-truncated array; `command-palette-screen.test.ts:63-73` asserts the cap but has no test proving typed search reaches a path beyond it. -- **Confidence:** High (Evidence). - -## [MEDIUM] State mutation — cli/src/chat.tsx:1339 — Hidden chat shortcuts remain active beneath full-screen search overlays - -- **Risk:** While the command palette or prompt-history screen owns focus, PageUp/PageDown, Ctrl+T, Ctrl+R/Ctrl+P, and Tab can still mutate the hidden chat or open a second overlay because the global chat keyboard hook remains enabled. -- **Fix:** Include `commandPaletteOpen` and `promptHistoryOpen` in the `disabled` gate, and make each overlay the sole keyboard owner until it closes. -- **Evidence:** The `useChatKeyboard` disable expression at `chat.tsx:1342-1347` covers ask-user, review, model, provider, and plan pickers but omits both search overlays; `keyboard-actions.ts:168-183,327-354` continues resolving global overlay, collapse, mode, and scroll shortcuts, while the overlay interceptors intentionally return `false` for most of those keys. -- **Confidence:** High (Evidence). - -## [MEDIUM] Correctness — cli/src/hooks/use-scroll-management.ts:126 — Keyboard page-up incorrectly re-enables follow mode - -- **Risk:** A user paging upward to read earlier output is marked “at bottom,” so the next streamed message can snap the viewport back to the latest content and destroy their reading position. -- **Fix:** Distinguish intentional navigation (`scrollUp`/`scrollDown`) from follow-to-latest animation, disable auto-follow when the target is above the bottom, and derive `isAtBottom` from the actual resulting position. -- **Evidence:** Every animated scroll sets `programmaticScrollRef.current = true` (`:66`), and the change handler then unconditionally sets `autoScrollEnabledRef.current = true` and `setIsAtBottom(true)` (`:126-130`), including animations initiated by `scrollUp()` (`:92-100`); no dedicated scroll-management test exists in the manifest or test tree. -- **Confidence:** High (Evidence). - -## [MEDIUM] State mutation — cli/src/hooks/use-chat-ui.ts:72 — Scroll listeners stay attached to a destroyed scrollbox after overlays - -- **Risk:** Opening and closing a full-screen picker/search screen replaces the chat scrollbox, but overflow and at-bottom tracking can remain bound to the old instance, causing stale scrollbar visibility/status state and leaked listeners. -- **Fix:** Use a callback ref or renderer-instance state so listener effects depend on the current `ScrollBoxRenderable`, detach on ref replacement, and reattach after the chat viewport remounts. -- **Evidence:** `useChatUI` subscribes to `scrollbox.verticalScrollBar` in an effect with `[]` (`:72-93`), and `useChatScrollbox` similarly depends only on the stable ref object (`use-scroll-management.ts:113-143`); `chat.tsx:1543-1617` returns full-screen overlays that unmount the normal scrollbox without unmounting `Chat` or either hook. -- **Confidence:** High (Evidence). - -## [MEDIUM] Error handling — cli/src/components/error-boundary.tsx:26 — Nested-agent fallback is a passthrough, not an error boundary - -- **Risk:** A malformed or renderer-incompatible nested agent subtree can still take down the entire TUI instead of showing the promised local fallback, interrupting the session and hiding the output that caused the failure. -- **Fix:** Install a real OpenTUI-compatible error boundary at the root and nested-agent/tool boundaries, or pre-render risky subtrees through a supported isolation mechanism with a tested fallback. -- **Evidence:** `ErrorBoundaryPlaceholder` returns `children` unchanged and explicitly says it does not catch render errors (`:10-30`), yet `message-with-agents.tsx:75-92` imports the deprecated `ErrorBoundary` alias and wraps `AgentChildrenGrid` with an “Error rendering agent children” fallback that can never activate; fatal process handlers in `index.tsx:386-423` exit on uncaught render failures. -- **Confidence:** High (Evidence). - -## [MEDIUM] Performance — cli/src/utils/terminal-color-detection.ts:82 — Theme probing can add a full second before first paint on unknown terminals - -- **Risk:** Every interactive startup on an unrecognized TTY can wait through two sequential 500 ms OSC queries before OpenTUI renders, making the CLI feel slow precisely on terminals where detection is least likely to work. -- **Fix:** Treat only known/probed terminals as OSC-capable, prefer immediate environment fallbacks, cache capability, and move nonessential detection after first paint when safe. -- **Evidence:** `terminalSupportsOSC()` falls back to `process.stdin.isTTY === true` for every unknown terminal (`:82-83`), `detectTerminalThemeCore()` awaits OSC 11 and then OSC 10 sequentially (`:416-430`), each query has a 500 ms timeout (`:18-20,193-196`), and `index.tsx:246-258` awaits detection before argument parsing and renderer creation; tests cover known-positive terminals but no unknown-TTY fast path (`terminal-color-detection.test.ts:250-309`). -- **Confidence:** High (Evidence). - -## [LOW] API/ABI contract breaks — cli/src/hooks/use-terminal-layout.ts:7 — Three incompatible responsive contracts drift across the CLI - -- **Risk:** Components switch chrome and layouts at different widths, so the same terminal can be “narrow” to one screen and “standard” to another, producing hard-to-predict transitions and documentation that cannot be trusted. -- **Fix:** Define one exported responsive token set with named layout intents, migrate all screens to it, and add cross-component boundary tests plus synchronized documentation. -- **Evidence:** `use-terminal-layout.ts` documents `xs` as `<80` but implements `<50` (`:7,23,121-130`), `use-terminal-breakpoints.ts:20-29` uses 60/100 and 15/20/30, grid layout uses 100/150/200 (`use-grid-layout.ts:13-22`), and `cli/knowledge.md:306-320` still specifies a 70-column screen-mode contract; existing tests validate each local constant rather than a shared contract. -- **Confidence:** High (Evidence). - -## [LOW] Dependency hygiene — cli/package.json:54 — Unused `terminal-image` retains a heavy duplicate image stack - -- **Risk:** The workspace install and dependency graph carry an unused terminal rendering package plus its transitive legacy image decoder, increasing maintenance and supply-chain surface without providing CLI behavior. -- **Fix:** Remove `terminal-image` if the custom `terminal-images.ts` path is authoritative, or replace the custom protocol code with the dependency and delete the duplicate implementation. -- **Evidence:** `terminal-image` is declared at `cli/package.json:54` but has no imports anywhere under `cli/`; `bun.lock` shows it depends on `jimp@1.6.0` and `render-gif`, which additionally brings `jimp@0.14.0`, while the CLI already imports `jimp@1.6.0` directly in `utils/image-thumbnail.ts`. -- **Confidence:** High (Evidence). - -## [LOW] Test coverage gaps — cli/src/components/**tests**/command-palette-screen.test.ts:32 — Presentation tests do not assert what users can actually read - -- **Risk:** Layout/reconciliation regressions can pass helper-level tests while command rows, focus highlights, or text disappear in the real renderer, especially across terminal widths and color capabilities. -- **Fix:** Add ANSI-aware tmux golden checks at narrow/standard/wide sizes and light/dark/256-color modes, asserting visible command labels, selected-row markers, focus restoration, and overlay close behavior rather than only capture existence. -- **Evidence:** The command-palette test file exercises `buildEntries`, `scoreEntry`, and `entryToListItem` only (`:32-297`); the supplied built-binary capture `debug/tmux-sessions/audit-cli-baseline/capture-003-command-palette.txt` at 120x36 visibly contains `/` and `23 items` but no readable command labels (plain capture may omit styling, not expected text), and no integration test correlates the capture with the 23 rendered commands. -- **Confidence:** Medium (Inference corroborated by capture evidence). - -## [LOW] Correctness — cli/src/components/tools/diff-viewer.tsx:459 — Side-by-side diffs encode add/delete state primarily by red and green - -- **Risk:** In monochrome, low-color, or red-green color-impaired viewing, paired changed rows lose an explicit operation marker, reducing confidence about which side is removed versus added. -- **Fix:** Preserve `-` and `+` markers (or labeled OLD/NEW gutters) in side-by-side mode and add a no-color snapshot test. -- **Evidence:** Unified rows render `signChar` (`:431-454`), but side-by-side rows render only line numbers, colored text, and a separator (`:459-502`); the tests verify the separator and fallback width but not a color-independent semantic marker (`diff-viewer.test.tsx:169-190`). -- **Confidence:** High (Evidence). - -## Strengths observed - -- The render shell has explicit fatal terminal cleanup, alternate-screen reset handling, and listener removal during renderer handoff. -- Diff presentation now includes parsed hunks, line-number gutters, explicit truncation disclosure, automatic initial collapsing, and unified-mode `+`/`-` markers. -- Message and agent rendering has meaningful pagination, collapse controls, depth-limit disclosure, stable React keys, and substantial focused tests. -- Input handling is unusually comprehensive: paste routing strips ANSI, long paste becomes an attachment, Enter normalization is centralized, and ask-user forms are keyboard-operable. -- Theme and terminal compatibility work is broad: truecolor fallbacks, OSC timeouts/cleanup, light/dark palettes, 256-color awareness, tmux-specific knowledge, and native text wrapping are all present. -- The 120x36 startup capture shows a clean, calm baseline with clear directory context and a prominent input affordance; styling/contrast cannot be conclusively judged from plain text capture. - -## Coverage and files actually read - -- Evaluated all eight audit domains: Security, Correctness, State mutation, Error handling, Performance, Dependency hygiene, Test coverage gaps, and API/ABI contract breaks. No presentation-scoped critical/high security issue was substantiated; sensitive-file labeling, ANSI stripping on paste, base64 OSC52 payloads, and bounded diff rendering were positive controls. -- Read the audit rubric and the full presentation manifest; inspected current dirty-worktree status and scoped diffs/statistics without reading any existing `.agents/sessions/*` audit report. -- Read in full or targeted line-bounded sections: `cli/src/index.tsx`, `app.tsx`, `chat.tsx`, `use-chat-ui.ts`, `use-chat-state.ts`, `use-chat-messages.ts`, `use-scroll-management.ts`, terminal dimension/layout/breakpoint/grid hooks, `chat-scroll-accel.ts`, layout/text utilities, command/prompt search screens, suggestion/selectable list/input keyboard paths, message/agent/tool/collapse stores and renderers, diff/edit/read tool renderers, theme/color detection, terminal links/clipboard/images, ask-user and interaction primitives, animation/rerender instrumentation, and `cli/package.json`/`bun.lock` dependency metadata. -- Corroborated with the manifest's direct tests for command palette, layout/grid, keyboard, messages/agents, diffs/edit tools, theme detection, ask-user, and rerender performance; searched the complete manifest path set for listener, timeout, truncation, focus, keyboard, dependency, and error-handling patterns. -- Read relevant portions of `docs/architecture.md`, `docs/testing.md`, `docs/agents-and-tools.md`, `cli/knowledge.md`, and `cli/tmux.knowledge.md`. -- Inspected concrete built-binary captures `debug/tmux-sessions/audit-cli-baseline/capture-001-startup.txt` and `capture-003-command-palette.txt` from the isolated 120x36 run, treating absent styling cautiously. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/findings/runtime-state.md b/.agents/sessions/audit-cli-next-level-2026-07/findings/runtime-state.md deleted file mode 100644 index 1d8555bbd3..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/findings/runtime-state.md +++ /dev/null @@ -1,128 +0,0 @@ -# Runtime/state CLI audit findings - -## [HIGH] Correctness — cli/src/hooks/helpers/send-message.ts:353 — cancellation can fork UI history from SDK history - -- **Risk:** Escape immediately unlocks a new send while the cancelled SDK run is still producing its authoritative preserved session state, so the next prompt can start from stale `previousRunStateRef` and permanently omit the cancelled turn's prompt, partial work, tool results, and interruption marker from model context. -- **Fix:** Add an explicit `cancelling` state and either await the cancelled run's `RunState` before admitting the next send or merge that state into a per-run continuation chain before a newer run captures `previousRun`. -- **Evidence:** `setupStreamingContext` releases `updateChainInProgress(false)` and `setCanProcessQueue(...)` at lines 353-375 and explicitly accepts stale continuation state at lines 355-360; `use-send-message.ts:581-592` then discards every aborted completion instead of updating `previousRunStateRef`, while `sdk/src/__tests__/run-cancellation.test.ts:945-1107` proves the SDK's cancelled result preserves history specifically so it can be passed as the next run's `previousRun`. -- **Confidence:** High. -- **Basis:** Evidence. - -## [HIGH] API/ABI contract breaks — sdk/src/run.ts:719 — promised async event handlers are dispatched fire-and-forget - -- **Risk:** SDK consumers can receive reordered events or see `client.run()` resolve before their async `handleEvent`/`handleStreamChunk` side effects finish, despite the public types promising `Promise` support. -- **Fix:** Serialize response callbacks through an awaited event queue and drain it before returning the terminal `RunState`, or change `SendActionFn` to return a promise and await it throughout the runtime. -- **Evidence:** `OpenbuffClientOptions` declares both callbacks as `void | Promise` at `sdk/src/run.ts:152-168`, and `onResponseChunk` awaits them at lines 530-570, but `sendAction` calls `onResponseChunk(action)` and `onSubagentResponseChunk(action)` without awaiting at lines 719-725; `common/src/types/contracts/client.ts:63` fixes `SendActionFn` to a synchronous `void`, and the ordering tests use a synchronous collector (`sdk/e2e/utils/event-collector.ts:27-38`) so they do not exercise the advertised async ABI. -- **Confidence:** High. -- **Basis:** Evidence. - -## [MEDIUM] Security — cli/src/utils/run-state-storage.ts:127 — `--continue` accepts traversal outside the chat directory - -- **Risk:** A crafted conversation id such as `../../../some-dir` makes the CLI read `run-state.json` and `chat-messages.json` outside the project chat store, allowing unintended local-file ingestion into the UI and subsequent agent context. -- **Fix:** Require a single basename chat id and verify `path.relative(chatsDir, candidateDir)` remains contained before any stat/read, reusing the validation already applied by deletion. -- **Evidence:** `cli/src/index.tsx:121-160` accepts an arbitrary optional `--continue [conversation-id]`; `use-send-message.ts:153-161` passes it directly to `loadMostRecentChatState`; that function uses `path.join(baseDir, chatId.trim())` at lines 127-134 without containment validation, while `deleteChatSession` correctly rejects `.`, `..`, and non-basename ids at `cli/src/utils/chat-history.ts:129-139`. -- **Confidence:** High. -- **Basis:** Evidence. - -## [MEDIUM] State mutation — cli/src/utils/run-state-storage.ts:93 — completed chat persistence is not crash-atomic - -- **Risk:** A crash, disk-full error, or interruption between the two direct writes can leave a truncated file or a `run-state.json`/`chat-messages.json` pair from different turns, after which restore rejects the whole session and silently falls back. -- **Fix:** Persist one versioned envelope atomically, or write both versioned temp files plus a commit manifest and rename only after every write/fsync succeeds. -- **Evidence:** `saveChatState` writes the two live files sequentially with `writeFileSync` at lines 93-105; `loadMostRecentChatState` requires and parses both together at lines 152-193 and returns `null` on any failure; the same module already implements temp-file plus same-directory rename for checkpoints at lines 216-256, demonstrating the safer local pattern, while `run-state-storage.test.ts:266-338` only tests manual serialization shape rather than interrupted saves. -- **Confidence:** High. -- **Basis:** Evidence. - -## [MEDIUM] Correctness — cli/src/hooks/use-message-queue.ts:294 — rejected queued sends are irreversibly dropped - -- **Risk:** The queue removes its head before `sendMessage` is accepted, and any rejection is only logged, so initialization races or future preflight failures can erase a user's queued prompt and attachments without a retry affordance. -- **Fix:** Keep an explicit in-flight queue item until send acceptance, and on rejection either requeue it with retry metadata or restore it to the input while showing a visible actionable error. -- **Evidence:** `processNextMessage` slices the item out at lines 294-302 before calling `completeQueuedMessageProcessing`; the rejection handler at lines 121-127 only calls `logger.warn`; `use-chat-streaming.ts:154-170` explicitly rejects when the send ref is missing and logs that the message was dropped, but `use-queue-controls.test.ts:56-249` covers ownership/watchdogs and has no rejected-send preservation assertion. -- **Confidence:** High. -- **Basis:** Evidence. - -## [MEDIUM] Correctness — cli/src/hooks/use-connection-status.ts:6 — resilience indicators are disconnected from provider reality - -- **Risk:** Users cannot distinguish DNS loss, provider unreachability, rate-limit backoff, or failover from ordinary thinking because the CLI always reports connected and never sets retry state true, making queue/recovery behavior opaque during the moments trust matters most. -- **Fix:** Add structured `provider_attempt`, `retry_scheduled`, `failover`, and `provider_recovered` events carrying model/provider, attempt, delay, and status; derive connection and retry UI from those events rather than a constant hook. -- **Evidence:** `useConnectionStatus` returns `true` unconditionally at lines 6-10 and its tests require that behavior; `cli/src/app.tsx:220` hard-codes `authStatus = 'ok'`; retry logic in `sdk/src/impl/llm.ts:1280-1318` only logs attempts; the event union in `common/src/types/print-mode.ts:189-205` has no retry/failover event; CLI code only calls `setIsRetrying(false)` (`use-send-message.ts:301,597`, `sdk-event-handlers.ts:145-150`), while `formatRetryBannerMessage` is defined but unused. -- **Confidence:** High. -- **Basis:** Evidence. - -## [MEDIUM] Error handling — packages/agent-runtime/src/run-agent-step.ts:1574 — internal stacks can reach the user-facing error banner - -- **Risk:** Local paths, implementation frames, and provider/library internals can be exposed in the TUI and persisted chat history, and raw technical output makes errors less trustworthy and actionable. -- **Fix:** Preserve structured diagnostics only in redacted logs, and return a stable public error code, safe summary, provider/status metadata, and optional recovery action to the CLI. -- **Evidence:** The runtime appends `error.stack` when no structured server status is available at lines 1574-1583 and returns it in `output.message` at lines 1602-1609; the CLI deliberately passes that raw message into `UserErrorBanner` at `cli/src/hooks/helpers/send-message.ts:450-453`; `sdk/src/error-utils.ts:107-124` claims to sanitize but simply returns the original message, while `cli/src/utils/__tests__/error-handling.test.ts:38-47` asserts stacks must not reach user-facing messages on a different helper path. -- **Confidence:** High. -- **Basis:** Evidence. - -## [MEDIUM] Error handling — cli/src/utils/sdk-event-handlers.ts:745 — declared runtime error events are silently ignored - -- **Risk:** Parser/tool/runtime errors emitted during an otherwise continuing run disappear from the TUI, leaving users watching an apparently stuck agent without the warning the SDK explicitly delivered. -- **Fix:** Handle `PrintModeEvent.type === 'error'` as a non-destructive activity/error block with correlation metadata, and distinguish recoverable stream warnings from terminal run errors. -- **Evidence:** `printModeErrorSchema` is part of the public discriminated union at `common/src/types/print-mode.ts:12-16,189-205`; the SDK forwards action errors through `handleEvent({ type: 'error' ... })` at `sdk/src/run.ts:478-481`; `createEventHandler` matches tool/subagent/finish/phase/context events but falls through to `.otherwise(() => undefined)` at `cli/src/utils/sdk-event-handlers.ts:745-759`, and `sdk-event-handlers.test.ts` has no error-event assertion. -- **Confidence:** High. -- **Basis:** Evidence. - -## [MEDIUM] Performance — cli/src/utils/chat-history.ts:48 — chat history search blocks the TUI on synchronous full-file reads - -- **Risk:** Opening history with hundreds of long chats can freeze keyboard/render responsiveness while up to 500 JSON files are synchronously statted, read, and parsed on the main event loop. -- **Fix:** Maintain a small atomic per-chat metadata index and load pages asynchronously, moving legacy index reconstruction to a worker or bounded async concurrency. -- **Evidence:** `getAllChats` synchronously enumerates and stats every directory at lines 48-78, then synchronously reads/parses as many as 500 complete message files at lines 80-114; `chat-history-screen.tsx:51-60` calls it during initial state construction and again in `setTimeout(0)`, which defers work but does not move it off the UI thread. -- **Confidence:** High. -- **Basis:** Evidence. - -## [LOW] Dependency hygiene — cli/src/utils/send-message-helpers.ts:7 — CLI imports undeclared `lodash` - -- **Risk:** Isolated CLI installs/builds depend on the root workspace's hoisted development dependency, so a packaging or workspace-layout change can turn a working source tree into a missing-module startup failure. -- **Fix:** Declare `lodash` in `cli/package.json` or replace the small `has`/`isEqual` uses with local/native helpers and remove the transitive reliance. -- **Evidence:** `send-message-helpers.ts:7`, `message-block-helpers.ts:1`, and `sdk-event-handlers.ts:1` import `lodash`; `cli/package.json` does not declare it, while only root `package.json:64` declares `lodash` under `devDependencies` and the binary workflow installs with `bun install --frozen-lockfile --cwd cli` at `.github/workflows/cli-release-build.yml:84-85`. -- **Confidence:** High. -- **Basis:** Evidence. - -## [LOW] Correctness — cli/src/components/message-footer.tsx:221 — per-turn cost uses an ambiguous unit - -- **Risk:** A turn costing three cents is rendered as `cost 3` while the session status renders cents as dollars, so users can misread spend by roughly two orders of magnitude. -- **Fix:** Rename the legacy `credits` field to `costCents` at the UI boundary and format it with the same currency helper as the session total. -- **Evidence:** Runtime accounting explicitly stores provider cost in cents (`packages/agent-runtime/src/run-agent-step.ts:1251` and `cli/src/chat.tsx:389-391`); the status bar divides by 100 and prefixes `$` at `cli/src/components/status-bar.tsx:192-201`; the message footer prints the raw number as `cost ${cost}` at `message-footer.tsx:221-235`, and `message-block.completion.test.tsx:54-65` locks in `cost 3`. -- **Confidence:** High. -- **Basis:** Evidence. - -## [MEDIUM] Test coverage gaps — cli/src/hooks/helpers/**tests**/send-message.test.ts:961 — recovery tests simulate state that production never applies - -- **Risk:** The suite can pass while cancellation continuation loses history, async SDK callbacks escape ordering, queue rejection drops messages, and chat persistence tears, because current tests assert helper-local mechanics rather than the end-to-end ownership and failure boundaries. -- **Fix:** Add integration tests for abort-A/send-B with a deferred real `client.run`, rejected queued-send restoration, async callback ordering/drain-before-return, traversal rejection, and injected mid-write persistence failure/recovery. -- **Evidence:** The cancellation race description at lines 961-970 says B is blocked until A resolves, but the current test at lines 992 onward asserts B proceeds immediately; the later test manually assigns `previousRunState = runStateA` at lines 1380-1394 after run B is already set up, bypassing the production guard at `use-send-message.ts:581-592`; event-ordering uses synchronous handlers, queue tests omit rejection preservation, and run-state tests manually serialize files rather than exercising `saveChatState` failures. -- **Confidence:** High. -- **Basis:** Evidence. - -## Strengths observed - -- Streaming UI updates are deliberately batched at 100 ms and flushed before completion/error (`cli/src/utils/message-updater.ts:123-252`), balancing responsiveness with render pressure while preserving partial content. -- Tool and agent progress contracts are substantially richer than a basic spinner: queued/running/succeeded/failed/cancelled tool lifecycle, `tool_start`, subagent correlation, phase, context-window, and context-compaction events are represented and rendered (`common/src/types/print-mode.ts`, `cli/src/utils/sdk-event-handlers.ts`). -- The runtime's retry loop uses bounded exponential backoff with jitter, does not retry after content has been yielded, and is abort-aware; focused retry/abort tests passed. -- Runtime cancellation now propagates a shared signal, blocks new tools, waits for cooperative cleanup, stops browser sessions, and cleans owned temporary clone directories before returning (`sdk/src/run.ts:858-874`). -- Mid-turn checkpoints use same-directory temp writes and rename, include a turn id, reject malformed/stale snapshots, and provide a credible base for a first-class resume/retry UX. -- Persistence sanitization and mandatory sensitive-file filtering are present, and tool renderers have an exhaustive metadata disposition with a safe generic fallback. -- Status and completion surfaces already expose elapsed time, context use, model, git diff stats, cost, cache hit rate, queue preview, failed agents, and a visible Escape stop action. -- Release builds compile the source entry and run both synchronous and long-lived binary smoke checks, including embedded tree-sitter initialization. - -## Verification notes - -- The manifest's earlier source parse failure at `sdk/src/tools/find-files-matching-content.ts:473` is no longer present: the line now contains a valid conditional, `bun run cli/src/index.tsx --help` exits 0, and the diff shows the invalid control flow was repaired in the dirty worktree. -- The manifest's earlier CLI typecheck parse failure at `packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts:243` is also no longer present: that line now closes a report object, and `bun run --cwd cli typecheck` exits 0. -- The existing built binary still exits 0 for `cli/bin/openbuff --help`. -- Focused validation ran 90 tests across queue, persistence/checkpoint, connection, event translation, SDK cancellation/error history/retry/event handling, and runtime abort parsing: 82 passed; 8 checkpoint tests failed only because they attempt to write `/home/ben/.config/openbuff/...` in this restricted audit sandbox rather than an injected temp directory. This itself indicates test-environment coupling, but it was not treated as a product failure. - -## Coverage / files actually read - -- Docs: `docs/request-flow.md`, `docs/architecture.md`, `docs/testing.md`, `docs/development.md`, `docs/agents-and-tools.md`, `docs/local-mode.md`, `docs/environment-variables.md`. -- Send/run setup: `cli/src/hooks/use-send-message.ts`, `cli/src/hooks/helpers/send-message.ts`, `cli/src/utils/create-run-config.ts`, `cli/src/utils/codebuff-client.ts`, `cli/src/utils/send-message-helpers.ts`, `cli/src/utils/send-message-timer.ts`, `cli/src/types/contracts/send-message.ts`, `cli/src/utils/yield-to-event-loop.ts`, `cli/src/project-files.ts`. -- Streaming/events: `cli/src/hooks/use-chat-streaming.ts`, `cli/src/hooks/stream-state.ts`, `cli/src/utils/sdk-event-handlers.ts`, `cli/src/utils/stream-chunk-processor.ts`, `cli/src/utils/create-event-handler-state.ts`, `cli/src/utils/message-updater.ts`, `cli/src/utils/tool-result-normalizer.ts`, `common/src/types/print-mode.ts`. -- Queue/input/cancellation: `cli/src/chat.tsx`, `cli/src/hooks/use-message-queue.ts`, `cli/src/hooks/use-queue-controls.ts`, `cli/src/hooks/use-queue-ui.ts`, `cli/src/hooks/use-chat-input.ts`, `cli/src/hooks/use-chat-keyboard.ts`, `cli/src/hooks/use-exit-handler.ts`, `cli/src/utils/chat-input-key-intercept.ts`, plus the adjacent prompt router needed to verify submit-vs-enqueue behavior. -- History/state: `cli/src/hooks/use-chat-state.ts`, `cli/src/hooks/use-chat-messages.ts`, `cli/src/state/chat-store.ts`, `cli/src/state/chat-history-store.ts`, `cli/src/state/message-block-store.ts`, `cli/src/utils/message-history.ts`, `cli/src/utils/chat-history.ts`, `cli/src/utils/run-state-storage.ts`, `cli/src/types/chat.ts`, `cli/src/types/chat-state.ts`. -- Resilience/status: `cli/src/hooks/use-connection-status.ts`, `cli/src/hooks/use-timeout.ts`, `cli/src/utils/error-handling.ts`, `cli/src/utils/error-messages.ts`, `cli/src/utils/format-timeout.ts`, `cli/src/utils/openbuff-provider.ts`, `cli/src/utils/validation-error-helpers.ts`, `cli/src/utils/status-indicator-state.ts`. -- Rendering/progress: `cli/src/components/message-block.tsx`, `cli/src/components/message-with-agents.tsx`, `cli/src/components/blocks/tool-branch.tsx`, `cli/src/components/blocks/tool-block-group.tsx`, `cli/src/components/tools/registry.ts`, `cli/src/components/tools/types.ts`, `cli/src/components/tools/tool-call-item.tsx`, `cli/src/components/status-bar.tsx`, `cli/src/components/progress-bar.tsx`, `cli/src/components/message-footer.tsx`. -- SDK/runtime: `sdk/src/client.ts`, `sdk/src/run.ts`, `sdk/src/run-state.ts`, `sdk/src/impl/llm.ts`, `sdk/src/retry-config.ts`, `sdk/src/error-utils.ts`, `common/src/types/session-state.ts`, `common/src/types/messages/codebuff-message.ts`, `common/src/types/contracts/llm.ts`, `common/src/util/error.ts`, `packages/agent-runtime/src/run-agent-step.ts`, `packages/agent-runtime/src/main-prompt.ts`, `packages/agent-runtime/src/prompt-agent-stream.ts`, `packages/agent-runtime/src/tool-stream-parser.ts`, `packages/agent-runtime/src/tools/stream-parser.ts`, and the explicitly requested `packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts` verification target. -- Source/build contract: `cli/src/index.tsx`, `cli/tsconfig.json`, `cli/package.json`, `sdk/package.json`, `sdk/src/index.ts`, `sdk/src/tools/index.ts`, `sdk/src/tools/find-files-matching-content.ts`, root `package.json`, `bunfig.toml`, `cli/scripts/build-binary.ts`, `cli/scripts/smoke-binary.ts`, `cli/src/pre-init/tree-sitter-wasm.ts`, `cli/src/native/ripgrep.ts`, `sdk/scripts/build.ts`, `.github/workflows/cli-release-build.yml`. -- Tests read or executed include the manifest's queue, send lifecycle, event translation, connection/timeout, history/checkpoint, SDK cancellation/error/retry/event, runtime abort/stream parser, event collector/order, rendering completion, and source/binary contract groups. Existing `.agents/sessions/**` audit findings/reports were not read. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/findings/validation-summary.md b/.agents/sessions/audit-cli-next-level-2026-07/findings/validation-summary.md deleted file mode 100644 index db8ab57978..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/findings/validation-summary.md +++ /dev/null @@ -1,22 +0,0 @@ -# Independent validation summary - -## [HIGH] Test coverage gaps — cli/src/**tests**/integration/local-agents.test.ts:1 — Current CLI suite is not green - -- **Risk:** The current worktree cannot be treated as release-ready because a fresh isolated-home run still has a broad local-agent integration failure cluster plus generated init-type drift. -- **Fix:** Repair the underlying local-agent default/trust/test-isolation contract, regenerate init type sources, and require the full isolated CLI suite in release gates. -- **Evidence:** Final run `HOME=/tmp/openbuff-test-home-final-20260711-2350 bun run --cwd cli test` produced 2,327 pass, 30 fail, 15 skip across 2,372 tests; 29 failures are in `integration/local-agents.test.ts` and one is `init-type-sources.test.ts`. -- **Confidence:** High — Evidence. - -## Current validation state - -- `bun run cli/src/index.tsx --help`: pass at final check. -- `bun run --cwd cli typecheck`: pass at final check. -- `cli/bin/openbuff --help` and `--version`: pass; local binary reports `1.0.0` and predates current source. -- Isolated 120x36 and 80x24 built-binary startup captures: pass with a clean input surface. -- The worktree changed during the audit. Two transient parse errors were directly observed, then fixed by another concurrent actor before the final check; they are historical validation evidence, not final open findings. - -## Strengths observed - -- The CLI test surface is large and detailed: 2,327 tests pass in the final isolated run. -- Source and compiled entry paths both have non-effectful smoke coverage. -- The built TUI starts successfully in isolated wide and narrow terminal sessions. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/manifests/distribution-quality.md b/.agents/sessions/audit-cli-next-level-2026-07/manifests/distribution-quality.md deleted file mode 100644 index f3885953a6..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/manifests/distribution-quality.md +++ /dev/null @@ -1,97 +0,0 @@ -# Distribution and operational quality file manifest - -## Scope - -Independent file-picker manifest for packaging/release wrappers and versioning; binary build and smoke validation; Linux/macOS/Windows/WSL compatibility; native dependencies; analytics/privacy/logging; crash/error containment; dependency hygiene; CLI-facing documentation; CI/release testing; release freshness and operational diagnostics. This is navigation only, not an audit or recommendation set. - -## Subshard A — production npm wrapper, updater, download, and crash diagnostics (1,187 LOC) - -- `cli/release/index.js` — production Node wrapper entry point. Key flow: `createConfig` → platform detection (`getHardwareArch`, `getMacOSVersion`, `getPlatformKey`, `assertSupportedPlatform`) → registry/version resolution (`getLatestVersion`, `getCurrentVersion`, `compareVersions`) → `downloadBinary`/extraction/metadata → `ensureBinaryExists` → child execution/update checking → `printCrashDiagnostics`. -- `cli/release/http.js` — proxy-aware HTTP(S), TLS, redirect and timeout client used for registry and binary downloads; main symbol `createReleaseHttpClient`. -- `cli/release/postinstall.js` — installation messaging and cached-binary cleanup behavior. -- `cli/release/package.json` — published package version, aliases (`openbuff`, `cb`), supported OS/CPU matrix, Node floor, tar dependency, uninstall cleanup. -- `cli/release/README.md` — user installation, platform support, cache/update and corporate proxy troubleshooting. - -Related tests: - -- `cli/src/__tests__/release-wrapper.test.ts` — wrapper/platform/version/update/crash behavior. -- `cli/src/__tests__/release/proxy-http-get.test.ts` — HTTP client proxy, TLS and redirect behavior. - -## Subshard B — staging wrapper parity and release freshness (1,166 LOC) - -- `cli/release-staging/index.js` — staging (`codecane`) counterpart to the production updater; compare flow and drift against `cli/release/index.js`. -- `cli/release-staging/http.js` — staging copy of release HTTP client. -- `cli/release-staging/postinstall.js` — staging cache cleanup/install messaging. -- `cli/release-staging/package.json` — staging package identity/version/platform matrix. -- `cli/release-staging/README.md` — staging install and troubleshooting documentation. - -Key flow: staging package/workflow version → npm registry lookup → platform artifact URL → local `~/.config/openbuff` cache → child binary. Treat production/staging duplication and synchronization as an explicit audit boundary. - -## Subshard C — binary compilation, native assets, smoke gates, and release trigger (802 LOC) - -- `cli/scripts/build-binary.ts` — Bun compile target selection, injected version/env values, OpenTUI native bundle acquisition/patching, tree-sitter WASM and ripgrep bundling, legacy macOS handling. Key symbols: `getTargetInfo`, `main`, `assertLegacyMacOSBuildConfig`, `patchOpenTuiNativeEntryForLegacy`, `patchOpenTuiAssetPaths`, `ensureOpenTuiNativeBundle`. -- `cli/scripts/smoke-binary.ts` — boot-level binary validation beyond `--help`/`--version`; detects boot signals and fatal patterns and exercises tree-sitter startup. -- `cli/scripts/release.ts` — local GitHub workflow-dispatch client and token/version input handling. - -Related manifests/native inputs: - -- `cli/package.json` — binary/release scripts and runtime/native-adjacent dependencies (`@opentui/*`, `jimp`, `systeminformation`, `terminal-image`, `yoga-layout`, `node-machine-id`). -- `package.json`, `bun.lock`, `.bun-version`, `bunfig.toml` — workspace runtime pinning, overrides and dependency-resolution source of truth. `bun.lock` is large (~270 KB); inspect selectively by relevant package names rather than linearly. -- `sdk/vendor/ripgrep/x64-linux/rg`, `sdk/vendor/ripgrep/arm64-linux/rg` — checked-in Linux native executables consumed by distribution-related flows; binary inspection only, no generated-bundle review. - -## Subshard D — release and CI automation matrix (1,240 LOC) - -- `.github/workflows/cli-release-build.yml` — cross-platform artifact build matrix, native dependency preparation, smoke/upload orchestration. Large single workflow (443 LOC), but group remains below ~3k LOC. -- `.github/workflows/cli-release-prod.yml` — production versioning/publish/release orchestration. -- `.github/workflows/cli-release-staging.yml` — staging versioning/publish/release orchestration. -- `.github/workflows/ci.yml` — normal typecheck/test/build gates and platform coverage. -- `.github/workflows/nightly-e2e.yml` — scheduled CLI/runtime end-to-end signal. -- `.github/workflows/sdk-release.yml` — adjacent release conventions and package-version hygiene for comparison. - -Key flow: workflow dispatch/tag/version selection → matrix binary build → smoke gate → GitHub release artifacts → npm wrapper publication. Verify OS/architecture runners, artifact naming, secrets, caching, permissions, concurrency, failure diagnostics and production/staging parity. - -Related tests/scripts: - -- `cli/src/__tests__/e2e-cli.test.ts`, `cli/src/__tests__/integration-tmux.test.ts`, `cli/scripts/validate-cli-with-tmux.sh` — executable/TUI integration coverage. -- `scripts/openbuff-smoke.ts` — provider-level local smoke path (`OPENBUFF_SMOKE_OK` contract), distinct from compiled-binary boot smoke. -- `scripts/run-tests-summary`, `docs/testing.md` — repository test execution/reporting conventions. - -## Subshard E — runtime fatal paths, logging, analytics, and privacy boundary (1,509 LOC) - -- `cli/src/index.tsx` — CLI version loading/argument parsing, initialization, platform-specific TTY handling, early `uncaughtException`/`unhandledRejection` handlers and transition to normal runtime. -- `cli/src/utils/logger.ts` — Pino file logging, log location/lifecycle, context enrichment, analytics-dispatch coupling and analytics error reporting. -- `cli/src/utils/analytics.ts` — current CLI analytics API/stubs, dependency seams, identification/error APIs. -- `cli/src/utils/error-handling.ts` — safe user-facing error extraction vs internal diagnostic logging. -- `cli/src/components/error-boundary.tsx` — React/OpenTUI error fallback boundary behavior. -- `cli/src/components/user-error-banner.tsx`, `cli/src/utils/error-messages.ts`, `cli/src/utils/validation-error-helpers.ts` — adjacent user-visible failure presentation. -- `common/src/analytics.ts`, `common/src/analytics-core.ts` — shared PostHog client/config and tracking behavior. -- `common/src/util/analytics-log.ts`, `common/src/util/analytics-dispatcher.ts`, `common/src/util/analytics-sampling.ts` — event recognition, routing and sampling policy. -- `common/src/constants/analytics-events.ts`, `common/src/types/contracts/analytics.ts`, `common/src/types/contracts/logger.ts`, `common/src/analytics.knowledge.md` — event/schema/contracts and intended policy. - -Related tests: - -- `common/src/util/__tests__/analytics-log.test.ts`, `common/src/util/__tests__/analytics-dispatcher.test.ts`, `common/src/util/__tests__/analytics-sampling.test.ts`. -- `cli/src/components/__tests__/user-error-banner.test.tsx`, `cli/src/utils/__tests__/validation-error-formatting.test.ts`. - -## Subshard F — cross-platform and operational documentation (800 LOC) - -- `README.md`, `README.zh-CN.md` — top-level install/upgrade/feature and support claims. -- `WINDOWS.md` — Windows native/WSL setup, shell/toolchain and known-platform guidance. -- `cli/README.md` — CLI development and invocation notes. -- `docs/development.md` — supported development/runtime workflow and operational logs. -- `docs/testing.md` — CI/local/TMUX test strategy. -- `docs/environment-variables.md` — environment-variable and diagnostics configuration contract. -- `docs/local-mode.md`, `docs/configuration.md`, `SECURITY.md`, `CONTRIBUTING.md` — adjacent BYOK/privacy, config, vulnerability-reporting and contributor operational claims. - -Cross-check documentation claims against the published wrapper OS/CPU matrix, workflow matrix, Bun/Node requirements, binary cache path, update behavior, proxy support, diagnostics/log paths and analytics state. - -## Explicit exclusions - -- `node_modules/**`, compiled `dist/**`/`build/**`, release archives, generated agent/type-source bundles, and binary contents beyond identifying checked-in native artifacts. -- Existing audit manifests/findings under `.agents/sessions/**`; this manifest was independently discovered from `MAP.md`, repository paths, manifests and symbol searches. -- Provider/model UX, chat interaction design, agent orchestration quality, indexing/retrieval quality, and general SDK API design except where directly required by CLI packaging, native assets, logging, analytics, tests or release operations. -- Historical changelog entries under `scripts/changelog/**` except a later auditor may sample them solely to compare documented release cadence/freshness. - -## Size notes - -No proposed subshard exceeds ~3k source LOC. The combined surface is intentionally split because the full set is ~7.6k LOC plus the lockfile and binary assets. The largest individual files are the production/staging wrappers (~800 LOC each) and `cli/scripts/build-binary.ts` (~500 LOC). diff --git a/.agents/sessions/audit-cli-next-level-2026-07/manifests/interaction-commands.md b/.agents/sessions/audit-cli-next-level-2026-07/manifests/interaction-commands.md deleted file mode 100644 index 5464b862f9..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/manifests/interaction-commands.md +++ /dev/null @@ -1,196 +0,0 @@ -# Interaction and commands file-picker manifest - -## Scope - -Independent source map for the CLI's conversational input surface: chat/input composition, terminal keyboard behavior, slash-command discovery and routing, bash mode, prompt history/completion/suggestions, attachments/images/clipboard, and user-facing help/discoverability. This is a file-selection manifest only; it contains no quality judgments or feature recommendations. - -The primary surface is large enough to split before auditing. Groups below are intentionally bounded to roughly 5–15 primary files. Tests and documentation are listed separately so an auditor can verify contracts without silently expanding a source shard. - -## Subshard IC-1 — Input widget and terminal-key dispatch - -Primary files (8; approximately 4.5k LOC, **over the ~3k LOC target — split `chat.tsx` into its own review pass if context is tight**): - -- `cli/src/chat.tsx` -- `cli/src/components/chat-input-bar.tsx` -- `cli/src/components/multiline-input.tsx` -- `cli/src/components/input-cursor.tsx` -- `cli/src/components/input-mode-banner.tsx` -- `cli/src/hooks/use-chat-input.ts` -- `cli/src/hooks/use-chat-keyboard.ts` -- `cli/src/utils/keyboard-actions.ts` - -Key symbols / flows: - -- `Chat` composes input state, overlay state, suggestion state, command routing, clipboard status, and global keyboard handlers. -- `ChatInputBar` renders the active input mode, previews, suggestions, and `MultilineInput`. -- `MultilineInput` owns cursor movement, editing, paste/IME handling, printable-key normalization, selection, wrapping, and newline shortcuts. -- `useChatKeyboard` maps global terminal events to `ChatKeyboardHandlers`; `dispatchAction` bridges classified actions to UI behavior. -- `keyboard-actions.ts` classifies keys by mode and interaction state; `useChatInput` provides input sizing and build/send helpers. - -Related tests: - -- `cli/src/components/__tests__/multiline-input.test.tsx` -- `cli/src/hooks/__tests__/use-chat-input.test.ts` -- `cli/src/utils/__tests__/keyboard-actions.test.ts` -- `cli/src/utils/__tests__/chat-input-key-intercept.test.ts` -- `cli/src/utils/__tests__/terminal-enter-detection.test.ts` -- `cli/src/hooks/__tests__/use-terminal-layout.test.ts` - -## Subshard IC-2 — Suggestions, mentions, completion, and prompt history - -Primary files (10; approximately 2.6k LOC): - -- `cli/src/hooks/use-suggestion-engine.ts` -- `cli/src/hooks/use-input-history.ts` -- `cli/src/hooks/use-path-tab-completion.ts` -- `cli/src/components/suggestion-menu.tsx` -- `cli/src/components/command-palette-screen.tsx` -- `cli/src/components/prompt-history-search-screen.tsx` -- `cli/src/utils/path-completion.ts` -- `cli/src/utils/chat-history.ts` -- `cli/src/project-files.ts` -- `cli/src/hooks/use-searchable-list.ts` - -Key symbols / flows: - -- `useSuggestionEngine` parses `/` and `@` trigger contexts, ranks slash commands, agents, and project files, and returns selection state. -- `usePathTabCompletion` delegates filesystem/path expansion to `path-completion.ts`. -- `useInputHistory` navigates persisted prompt entries; `PromptHistorySearchScreen` provides fuzzy full-screen retrieval. -- `CommandPaletteScreen` builds searchable command and file entries from `SlashCommand[]` plus the project file tree. -- `project-files.ts` supplies the file tree used by mention matching, palette entries, and generated game-development commands. - -Related tests: - -- `cli/src/hooks/__tests__/use-suggestion-engine.test.ts` -- `cli/src/hooks/__tests__/use-suggestion-engine-mention.test.ts` -- `cli/src/hooks/__tests__/use-input-history.test.ts` -- `cli/src/hooks/__tests__/use-path-tab-completion.test.ts` -- `cli/src/__tests__/path-completion.test.ts` -- `cli/src/components/__tests__/command-palette-screen.test.ts` -- `cli/src/components/__tests__/prompt-history-search-screen.test.ts` -- `cli/src/data/__tests__/slash-commands.test.ts` - -## Subshard IC-3 — Slash-command registry, routing, help, and bash mode - -Primary files (11; approximately 2.8k LOC): - -- `cli/src/data/slash-commands.ts` -- `cli/src/commands/command-registry.ts` -- `cli/src/commands/router.ts` -- `cli/src/commands/router-utils.ts` -- `cli/src/commands/help.ts` -- `cli/src/commands/image.ts` -- `cli/src/commands/prompt-builders.ts` -- `cli/src/utils/bash-context-processor.ts` -- `cli/src/utils/bash-messages.ts` -- `cli/src/components/pending-bash-message.tsx` -- `cli/src/utils/input-modes.ts` - -Key symbols / flows: - -- `SLASH_COMMANDS`, `SLASHLESS_COMMAND_IDS`, and `getSlashCommandsWithSkills` define discoverable command metadata, aliases, implicit commands, skill commands, and engine presets. -- `COMMAND_REGISTRY`, `findCommand`, and `findCommandSuggestions` define executable handlers and argument-aware resolution. -- `routeUserPrompt` distinguishes normal prompts, slash/implicit commands, queueing, and bash execution; `runBashCommand` and `addBashMessageToHistory` feed terminal results back into conversation context. -- `processBashContext`, `buildBashHistoryMessages`, and `formatBashContextForPrompt` translate local shell activity into model-visible message content. -- `help.ts` and command descriptions are the main built-in discoverability contract; `input-modes.ts` defines mode labels/behavior including bash/image/feedback modes. - -Related integration files: - -- `common/src/util/engine-profiles.ts` -- `common/src/util/game-dev-presets.ts` -- `cli/src/commands/plan-artifacts.ts` -- `cli/src/commands/plan-timeline.ts` -- `cli/src/commands/index-command.ts` -- `cli/src/commands/init.ts` -- `cli/src/commands/info.ts` - -Related tests: - -- `cli/src/commands/__tests__/command-args.test.ts` -- `cli/src/commands/__tests__/command-suggestions.test.ts` -- `cli/src/commands/__tests__/router-input.test.ts` -- `cli/src/commands/__tests__/bash-command.test.ts` -- `cli/src/__tests__/bash-mode.test.ts` -- `cli/src/utils/__tests__/bash-context-processor.test.ts` -- `cli/src/commands/__tests__/image.test.ts` -- `cli/src/commands/__tests__/plan-timeline.test.ts` - -## Subshard IC-4 — Clipboard and attachment ingestion - -Primary files (12; approximately 3.2k LOC, **likely just over the ~3k LOC target — split platform clipboard code from attachment state if needed**): - -- `cli/src/hooks/use-clipboard.ts` -- `cli/src/utils/clipboard.ts` -- `cli/src/utils/clipboard-image.ts` -- `cli/src/utils/image-handler.ts` -- `cli/src/utils/image-processor.ts` -- `cli/src/utils/pending-attachments.ts` -- `cli/src/state/chat-store.ts` -- `cli/src/types/store.ts` -- `cli/src/components/pending-attachments-banner.tsx` -- `cli/src/components/attachment-card.tsx` -- `cli/src/components/file-attachment-card.tsx` -- `cli/src/components/text-attachment-card.tsx` - -Key symbols / flows: - -- `useClipboard` registers the renderer and exposes transient copy/paste status; `clipboard.ts` selects platform tools, renderer APIs, or OSC52 for copy operations. -- `clipboard-image.ts` detects and reads image/file/text clipboard payloads across macOS, Linux, and Windows. -- `processImageFile`, `extractImagePaths`, and `processImagesForMessage` validate/compress image inputs and translate them into SDK message content. -- `pending-attachments.ts` creates, validates, de-duplicates, updates, and captures image/text/file attachments; `ChatStoreState.pendingAttachments` is the shared lifecycle state. -- Attachment banners/cards render processing, ready, partial, and error states before send. - -Related tests: - -- `cli/src/utils/__tests__/clipboard.test.ts` -- `cli/src/utils/__tests__/image-processor.test.ts` -- `cli/src/utils/__tests__/image-dimensions.test.ts` -- `cli/src/utils/__tests__/pending-attachments.test.ts` - -## Subshard IC-5 — Image rendering and message/SDK handoff - -Primary files (11; approximately 2.5k LOC): - -- `cli/src/components/image-card.tsx` -- `cli/src/components/image-thumbnail.tsx` -- `cli/src/components/blocks/image-block.tsx` -- `cli/src/utils/image-display.ts` -- `cli/src/utils/image-thumbnail.ts` -- `cli/src/utils/terminal-images.ts` -- `cli/src/hooks/use-send-message.ts` -- `cli/src/hooks/helpers/send-message.ts` -- `cli/src/types/chat.ts` -- `cli/src/types/contracts/send-message.ts` -- `sdk/src/impl/chatgpt-backend-fetch.ts` - -Key symbols / flows: - -- Image cards/blocks and terminal helpers choose inline image rendering, thumbnails, or textual fallback according to terminal support and dimensions. -- `useSendMessage` coordinates queue/run state; `prepareUserMessage` and the send helper collect pending bash context and attachments before calling the SDK. -- `PendingAttachment` variants become SDK `MessageContent`; `ChatMessage`/`ContentBlock` types retain the user-visible transcript representation. -- `chatgpt-backend-fetch.ts` is a provider-specific boundary that normalizes `image_url` parts and should be checked alongside the generic SDK message contract. - -Related tests: - -- `cli/src/hooks/helpers/__tests__/send-message.test.ts` -- `cli/src/utils/__tests__/send-message-helpers.test.ts` -- `cli/src/components/__tests__/message-block.completion.test.tsx` -- `cli/src/components/__tests__/message-block.streaming.test.tsx` - -## Documentation and user-facing contract references - -- `docs/request-flow.md` — authoritative CLI-to-SDK message preparation flow, including bash context and attachments. -- `docs/agents-and-tools.md` (especially “Slash Commands”) — documented registry, aliases, implicit commands, skills, and generated presets. -- `docs/architecture.md` — TUI ownership and high-level command/input responsibilities. -- `docs/testing.md` — expected TUI/tmux validation approach. -- `README.md` and `WINDOWS.md` — installation/platform expectations relevant to clipboard, shell, and terminal interaction. -- `cli/package.json` — terminal/image/clipboard runtime dependencies and CLI test scripts. - -## Explicit exclusions - -- Provider/model setup, auth/connect screens, project picker, and route picker except where they reuse `MultilineInput`; those belong to onboarding/configuration shards. -- Chat transcript rendering, tool blocks, streaming lifecycle, scrolling, and collapse behavior except the image/message-handoff files explicitly listed above. -- Durable-plan semantics, indexer behavior, agent runtime/tool execution, and SDK provider correctness beyond their direct command or message-content boundary. -- Generated agent bundles such as `cli/src/data/initial-agent-type-sources.generated.ts`. -- `node_modules`, build outputs, release artifacts, snapshots, and existing audit findings/session reports. -- Specialized modal inputs (`ask-user`, feedback, publish) except shared input-mode definitions; they should be reviewed with their owning feature shards to avoid duplication. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/manifests/onboarding-config.md b/.agents/sessions/audit-cli-next-level-2026-07/manifests/onboarding-config.md deleted file mode 100644 index f064eca193..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/manifests/onboarding-config.md +++ /dev/null @@ -1,261 +0,0 @@ -# File manifest — onboarding, project selection, configuration, providers, OAuth, validation - -## Scope statement - -This manifest covers the CLI path from process start to a usable local/BYOK session: argument parsing and bootstrap, project-root selection, first-project initialization, user settings and environment loading, provider/model setup and routing configuration, ChatGPT/Codex OAuth, and local-agent validation surfaced in the TUI. It identifies implementation, integration, tests, docs, and example configuration only. It intentionally does **not** assess quality or recommend changes. - -The structural map used for discovery is `.agents/sessions/audit-cli-next-level-2026-07/MAP.md`. Generated bundles, existing audit findings, and `node_modules` were not searched as implementation sources. - -## Primary files grouped into audit subshards - -### OC-1 — Process startup and application bootstrap (13 files, ~2.6k LOC) - -- `cli/src/index.tsx` -- `cli/src/app.tsx` -- `cli/src/init/init-app.ts` -- `cli/src/init/init-direnv.ts` -- `cli/src/pre-init/tree-sitter-wasm.ts` -- `cli/src/project-files.ts` -- `cli/src/utils/env.ts` -- `cli/src/utils/create-run-config.ts` -- `cli/src/__tests__/cli-args.test.ts` -- `cli/src/__tests__/e2e-cli.test.ts` -- `cli/src/__tests__/home-directory-detection.test.ts` -- `cli/src/init/__tests__/init-direnv.test.ts` -- `cli/src/__tests__/unit/create-run-config.test.ts` - -Key symbols and flow: - -- `main()` / `parseArgs()` parse `--cwd`, `--agent`, `--continue`, `--plan`, and the compatibility `--local` flag. -- `initializeApp()` applies `cwd`, establishes the module-level project root, initializes analytics/direnv/theme/timestamps, starts the index, and refreshes stored ChatGPT credentials in the background. -- `setProjectRoot()` / `getProjectRoot()` are the root authority used by CLI storage, tools, and plan artifacts. -- `AppWithAsyncAuth` loads the initial file tree, initializes local agents and skills, and passes project-picker state into `App`. -- `createRunConfig()` imports provider-config values (`maxAgentSteps`, indexing) into each SDK run and installs the sensitive-file filter. - -### OC-2 — Project and directory selection (11 files, ~2.1k LOC) - -- `cli/src/components/project-picker-screen.tsx` -- `cli/src/hooks/use-directory-browser.ts` -- `cli/src/utils/directory-browser.ts` -- `cli/src/utils/project-picker.ts` -- `cli/src/utils/recent-projects.ts` -- `cli/src/hooks/use-path-tab-completion.ts` -- `cli/src/utils/path-completion.ts` -- `cli/src/__tests__/utils/project-picker.test.ts` -- `cli/src/hooks/__tests__/use-directory-browser.test.ts` -- `cli/src/hooks/__tests__/use-path-tab-completion.test.ts` -- `cli/src/__tests__/path-completion.test.ts` - -Key symbols and flow: - -- `shouldShowProjectPicker(startCwd, homeDir)` gates first-screen project selection. -- `ProjectPickerScreen` combines recent projects, searchable directories, direct path entry, tab completion, keyboard navigation, and the final Open action. -- `useDirectoryBrowser()` delegates filesystem enumeration to `getDirectories()` and path expansion/navigation helpers. -- Selection returns to `cli/src/index.tsx::handleProjectChange()`, which calls `process.chdir`, updates `setProjectRoot`, resets the SDK client, persists recents, and reloads the file tree. -- `App` also exposes switching from a nested directory to the discovered Git root. - -### OC-3 — Project initialization and scaffold generation (8 files, ~3.6k LOC) - -**Sizing flag:** likely over the ~3k LOC audit target. Split the command/dispatch pair from template-generation sources if the auditor cannot keep all files in context. - -- `cli/src/commands/init.ts` -- `cli/src/commands/command-registry.ts` -- `cli/scripts/generate-init-type-sources.ts` -- `common/src/templates/initial-agents-dir/types/agent-definition.ts` -- `common/src/templates/initial-agents-dir/types/tools.ts` -- `common/src/templates/initial-agents-dir/types/util-types.ts` -- `cli/src/commands/__tests__/init.test.ts` -- `cli/src/__tests__/init-type-sources.test.ts` - -Key symbols and flow: - -- `command-registry.ts` maps slashless/`/init` input to `handleInitializationFlowLocally()`. -- `handleInitializationFlowLocally()` creates `knowledge.md`, `.agents/`, `.agents/types/`, and the three public type files, then returns a `postUserMessage` callback for TUI feedback. -- `generate-init-type-sources.ts` is the source-to-generated bridge used at build time; audit the source templates above rather than the generated payload. - -### OC-4 — User settings, config directory, environment, and theme persistence (14 files, ~2.2k LOC) - -- `cli/src/utils/settings.ts` -- `cli/src/utils/auth.ts` -- `cli/src/utils/theme-config.ts` -- `cli/src/hooks/use-theme.tsx` -- `cli/src/state/chat-store.ts` -- `cli/src/utils/env.ts` -- `cli/src/types/env.ts` -- `common/src/env-process.ts` -- `common/src/env.ts` -- `sdk/src/env.ts` -- `sdk/src/types/env.ts` -- `cli/src/__tests__/utils/env.test.ts` -- `common/src/__tests__/env-process.test.ts` -- `sdk/src/__tests__/env.test.ts` - -Key symbols and flow: - -- CLI storage is rooted at `~/.config/openbuff` via `cli/src/utils/auth.ts::getConfigDir()`. -- `loadSettings()` / `saveSettings()` manage `settings.json`; `loadModePreference()` feeds the initial Zustand chat-store mode. -- `theme-config.ts` and `use-theme.tsx` combine saved theme preferences with terminal detection performed during startup. -- `getBaseEnv()`, `getCliEnv()`, and `getSdkEnv()` define the environment-DI boundary; SDK helpers resolve `OPENBUFF_API_KEY`, the retained `CODEBUFF_API_KEY` alias, provider keys, and OAuth token aliases. -- Direnv mutation itself belongs to OC-1, but its resulting environment feeds this subshard. - -### OC-5A — Provider/model picker overlays and command dispatch (7 files, ~5k+ LOC) - -**Sizing flag:** over ~3k LOC. Prefer symbol-range review of `chat.tsx` and `command-registry.ts`; their unrelated chat/rendering and command implementations are not part of this scope. - -- `cli/src/components/provider-picker-screen.tsx` -- `cli/src/components/model-route-picker.tsx` -- `cli/src/chat.tsx` -- `cli/src/commands/command-registry.ts` -- `cli/src/components/selectable-list.tsx` -- `cli/src/hooks/use-searchable-list.ts` -- `cli/src/commands/__tests__/router-input.test.ts` - -Key symbols and flow: - -- `/setup`, `/provider`, and `/models` dispatch from `command-registry.ts` and return overlay-open or input-mode state. -- `Chat` owns `providerPickerOpen` / `modelRoutePickerOpen`, renders the full-screen overlays, and applies `ProviderPickerSelection` through provider helpers. -- `ProviderPickerScreen` covers built-in presets, Codex OAuth selection, environment-key status, and a custom-provider draft. -- `ModelRoutePicker` edits default, mode, agent, vision, and reasoning-effort routes against the currently loaded provider/model set. -- Shared searchable/selectable-list primitives define keyboard, filtering, focus, and selection behavior. - -### OC-5B — CLI provider setup, config mutation, and pre-send readiness (5 files, ~4.4k LOC) - -**Sizing flag:** over ~3k LOC, largely because the focused readiness test file also covers adjacent send-message behavior. Review only provider-readiness symbols in that test if necessary. - -- `cli/src/utils/openbuff-provider.ts` -- `cli/src/hooks/use-send-message.ts` -- `cli/src/hooks/helpers/send-message.ts` -- `cli/src/utils/__tests__/openbuff-provider.test.ts` -- `cli/src/hooks/helpers/__tests__/send-message.test.ts` - -Key symbols and flow: - -- `setupOpenbuffProviderFromArgs()`, `addCustomOpenbuffProvider()`, and `handleOpenbuffProviderCommand()` implement presets, custom providers, provider removal, discovery commands, and Codex connect/disconnect dispatch. -- `configureOpenbuffModelFromArgs()`, `setRouteModel()`, and picker-facing helpers mutate route configuration. -- `writeMergedConfig()` / `getEditableConfig()` bridge CLI choices into SDK config parsing and persistence. -- `getOpenbuffProviderReadiness()` resolves the selected agent/mode before a message is sent and returns user-facing setup, missing-route, missing-key, or missing-OAuth status. -- `use-send-message.ts` invokes readiness; `cleanupProviderReadinessFailure()` restores UI/queue state when the gate blocks a send. - -### OC-6 — SDK provider configuration, discovery, and model resolution (9 files, ~6.8k+ LOC) - -**Sizing flag:** substantially over ~3k LOC. Split `provider-config.ts` schema/load/write logic from model discovery/resolution tests, or review the named symbols below by ranges. - -- `sdk/src/provider-config.ts` -- `sdk/src/model-discovery.ts` -- `sdk/src/impl/model-provider.ts` -- `sdk/src/__tests__/model-provider.test.ts` -- `sdk/src/impl/__tests__/provider-options-metadata.test.ts` -- `openbuff.d.example/providers.json` -- `openbuff.d.example/routes.json` -- `openbuff.d.example/indexing.json` -- `openbuff.json.example` - -Key symbols and flow: - -- `providerConfigFileSchema` defines providers, routes, reasoning effort, vision fallback, failover, indexing, hooks, and run limits. -- `loadProviderConfigSync()` loads the explicit `OPENBUFF_PROVIDER_CONFIG` path or merges global/project/ancestor config, `extends`/`include` fragments, and implicit `openbuff.d/*.json`, with source attribution and mtime caching. -- `writeProviderConfigFile()` validates and writes monolithic or fragmented configuration; `createProviderPresetConfig()` materializes built-in setup presets. -- `resolveConfiguredAgentModelConfig()` applies mode → agent → default → explicit fallback routing; `resolveConfiguredProviderModel()` maps the routable model to provider model and environment key. -- `discoverProviderModels()`, cache helpers, and `addDiscoveredModelToProviderConfig()` cover provider model discovery and persistence. -- `getModelForRequest()` is the SDK integration boundary that consumes the resolved configuration. Transport internals beyond that boundary belong to the provider/runtime audit. - -### OC-7A — ChatGPT/Codex OAuth CLI and TUI flow (10 files, ~2.4k LOC) - -- `cli/src/utils/chatgpt-oauth.ts` -- `cli/src/components/chatgpt-connect-banner.tsx` -- `cli/src/components/input-mode-banner.tsx` -- `cli/src/components/chat-input-bar.tsx` -- `cli/src/utils/input-modes.ts` -- `cli/src/commands/router.ts` -- `cli/src/data/slash-commands.ts` -- `common/src/constants/chatgpt-oauth.ts` -- `cli/src/utils/__tests__/chatgpt-oauth.test.ts` -- `cli/src/commands/__tests__/router-connect-chatgpt.test.ts` - -Key symbols and flow: - -- `/provider connect codex` or the provider picker sets the `connect:chatgpt` input mode. -- `ChatGptConnectBanner` starts the flow, displays the browser URL/status, supports retry/disconnect, and offers post-connect Codex preset creation. -- `connectChatGptOAuth()` creates PKCE verifier/state, opens the browser, and starts the localhost callback server; `exchangeChatGptCodeForTokens()` supports callback URL or manual-code input. -- `routeUserPrompt()` handles authorization-code input when the connect mode is active. -- `CHATGPT_OAUTH_*` constants define endpoints, redirect URI, token aliases, and the allowlisted model mapping. - -### OC-7B — OAuth credential storage and SDK request integration (9 files, ~6.4k LOC) - -**Sizing flag:** substantially over ~3k LOC because `llm.ts` and `model-provider.test.ts` are high-fanout files. Restrict those files to OAuth credential refresh, OAuth model resolution, backend request transformation, and stream fallback symbols. - -- `sdk/src/credentials.ts` -- `sdk/src/env.ts` -- `sdk/src/impl/model-provider.ts` -- `sdk/src/impl/chatgpt-backend-fetch.ts` -- `sdk/src/impl/llm.ts` -- `sdk/src/__tests__/credentials.test.ts` -- `sdk/src/__tests__/chatgpt-backend-fetch.test.ts` -- `sdk/src/impl/__tests__/llm-chatgpt-oauth-policy.test.ts` -- `sdk/src/__tests__/model-provider.test.ts` - -Key symbols and flow: - -- `getChatGptOAuthCredentials()` resolves environment override before `~/.config/openbuff/credentials.json`. -- `saveChatGptOAuthCredentials()`, `clearChatGptOAuthCredentials()`, `refreshChatGptOAuthToken()`, and `getValidChatGptOAuthCredentials()` own persistence, permissions, expiry, refresh, and refresh deduplication. -- `initializeApp()` in OC-1 triggers best-effort background refresh when credentials exist. -- `getModelForRequest()` selects a configured `chatgpt-oauth` provider or the allowlisted direct-OAuth path. -- `createChatGptBackendFetch()` / request-body transforms adapt AI SDK requests to the ChatGPT backend; `llm.ts` owns stream auth/rate-limit classification, refresh/retry, and fallback behavior. - -### OC-8 — Local-agent loading, validation contracts, and validation UX (14 files, ~5.2k LOC) - -**Sizing flag:** over ~3k LOC. Split local-agent load/common-schema validation from the small TUI formatting/popover group if necessary. - -- `cli/src/utils/local-agent-registry.ts` -- `sdk/src/agents/load-agents.ts` -- `sdk/src/validate-agents.ts` -- `common/src/templates/agent-validation.ts` -- `common/src/types/dynamic-agent-template.ts` -- `cli/src/hooks/use-agent-validation.ts` -- `cli/src/components/validation-error-popover.tsx` -- `cli/src/utils/validation-error-formatting.ts` -- `cli/src/utils/validation-error-helpers.ts` -- `cli/src/utils/format-validation-errors-for-message.ts` -- `cli/src/__tests__/integration/local-agents.test.ts` -- `sdk/src/__tests__/validate-agents.test.ts` -- `common/src/__tests__/agent-validation.test.ts` -- `cli/src/utils/__tests__/validation-error-formatting.test.ts` - -Key symbols and flow: - -- Startup calls `initializeAgentRegistry()`, which discovers project/user agent directories and delegates loading to the SDK. -- `loadLocalAgents()` imports agent modules, resolves MCP environment references, optionally validates, removes invalid definitions, and returns file-aware validation diagnostics. -- SDK `validateAgents()` adapts arrays to the common validator; common `validateAgents()` / `validateSingleAgent()` enforce the dynamic-agent schemas and duplicate-ID rules. -- `useAgentValidation()` runs local validation before send, filters network-only errors, and exposes state to the chat UI. -- `ValidationErrorPopover` and formatting helpers map validator IDs/messages back to local files and user-facing field labels. - -## Related documentation and example configuration - -- `docs/configuration.md` — authoritative config locations, fragment/merge semantics, routing resolution, discovery, capabilities, indexing, and hook configuration. -- `docs/local-mode.md` — active local/BYOK provider setup and routing behavior. -- `docs/openbuff-provider-model-setup-ux.md` — design/proposal context for the current provider/model setup surfaces; distinguish proposed phases from shipped behavior. -- `docs/environment-variables.md` — environment naming, DI helpers, and load order. -- `docs/request-flow.md` — CLI → SDK → configured provider request path and validation-gate context. -- `docs/development.md` — development-time direnv and local setup behavior. -- `docs/agents-and-tools.md` — local agent structure and validation/tool contracts. -- `README.md` — public CLI start, custom-agent init, and provider configuration promises. -- `.env.example` — documented key/env surface. -- `openbuff.d.example/providers.json`, `openbuff.d.example/routes.json`, `openbuff.d.example/indexing.json`, `openbuff.json.example` — executable examples that should be checked against the SDK schema and CLI-generated output. - -## Shared/high-fanout files - -- `cli/src/index.tsx`, `cli/src/app.tsx`, `cli/src/chat.tsx`, and `cli/src/commands/command-registry.ts` connect multiple subshards. Audit only the startup, project-picker, setup/model/provider, init, and OAuth branches listed above. -- `sdk/src/provider-config.ts`, `sdk/src/impl/model-provider.ts`, and `sdk/src/__tests__/model-provider.test.ts` are intentionally cross-referenced because configuration, model routing, provider selection, and OAuth converge there. -- `cli/src/utils/local-agent-registry.ts` participates in both startup and validation; its implementation is assigned to OC-8, while OC-1 should only verify the startup call order. - -## Explicit exclusions - -- `node_modules/`, lockfile contents, compiled binaries, release artifacts, and generated agent bundles. -- `cli/src/data/initial-agent-type-sources.generated.ts` as an audit source. Its generator and source templates are included in OC-3. -- Existing audit/session findings under `.agents/sessions/**/findings/` and `/tmp/openbuff-cli-audit-2026-07/findings/`; this manifest was independently derived from source, tests, docs, examples, and the structural map. -- `docs/authentication.md` and legacy hosted Codebuff login/device-code/cloud-auth behavior. That document explicitly describes the removed cloud flow; only the shared credentials-file boundary is relevant here. -- Deep provider transport behavior, token accounting, retry policy unrelated to ChatGPT OAuth, and agent-runtime execution after `getModelForRequest()`; those belong to request/runtime/provider audit shards. -- Index construction/ranking internals; only bootstrap, config, status, and run-config integration are in scope. -- General chat rendering, message history, attachments, plans, review UI, command implementations unrelated to `init`, `setup`, `provider`, `models`, or ChatGPT connect. -- MCP configuration, agent prompt quality, shipped agent behavior, evals, website/backend code, release/install/update UX, and `agents-graveyard/`. -- The future phases in `docs/openbuff-provider-model-setup-ux.md` are reference context, not assumed shipped functionality. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/manifests/presentation-quality.md b/.agents/sessions/audit-cli-next-level-2026-07/manifests/presentation-quality.md deleted file mode 100644 index 14aa99ed65..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/manifests/presentation-quality.md +++ /dev/null @@ -1,291 +0,0 @@ -# CLI presentation-quality file manifest - -## Scope - -Independent file selection for the CLI presentation half of audit shard pair 4. This manifest covers the rendered OpenTUI surface and the code that directly controls its terminal behavior: - -- root TUI composition and chat viewport -- terminal dimensions, responsive layout, resizing, scrolling, and collapse/navigation behavior -- keyboard reachability, input editing, suggestions, search, and full-screen picker screens -- message, nested-agent, thinking, plan/gate, and tool-call presentation -- theme selection, color capability detection, contrast inputs, markdown/text wrapping, images, and terminal links -- interactive primitives, banners, status/error UI, ask-user forms, and mode controls -- perceived performance: memo boundaries, store subscription boundaries, stable render trees, animations, and rerender instrumentation -- OpenTUI-specific integration constraints and terminal/platform compatibility - -This is a discovery manifest only. It intentionally contains no quality findings or feature proposals. - -## Primary source subshards - -All line counts are approximate `wc -l` counts from the current tree. Primary groups target 5-15 files and stay below the audit pattern's ~3k LOC risk threshold. - -### P1 — OpenTUI bootstrap and chat shell (~2.9k LOC, 6 files) - -- `cli/src/index.tsx` — renderer creation, renderer options/cleanup, root mounting, process lifecycle -- `cli/src/app.tsx` — top-level authenticated/project surfaces, logo/header composition, terminal-focus integration -- `cli/src/chat.tsx` — central chat layout, overlay routing, scrollbox, message list, status bar, input bar, keyboard callback wiring -- `cli/src/hooks/use-chat-ui.ts` — aggregation point for dimensions, responsive state, theme palette, scrolling, and overflow state -- `cli/src/hooks/use-chat-state.ts` — stabilized chat/store selections consumed by the render shell -- `cli/src/components/error-boundary.tsx` — React class boundary adapted to OpenTUI JSX constraints - -Sizing note: this group is close to the ~3k LOC threshold because `chat.tsx` is 1,779 LOC. Do not add adjacent streaming/business hooks to this subshard. - -### P2 — Responsive layout, resize, and scroll mechanics (~1.2k LOC, 11 files) - -- `cli/src/hooks/use-terminal-dimensions.ts` -- `cli/src/hooks/use-terminal-layout.ts` -- `cli/src/hooks/use-terminal-breakpoints.ts` -- `cli/src/hooks/use-grid-layout.ts` -- `cli/src/components/grid-layout.tsx` -- `cli/src/hooks/use-scroll-management.ts` -- `cli/src/utils/chat-scroll-accel.ts` -- `cli/src/utils/layout-helpers.ts` -- `cli/src/utils/text-layout.ts` -- `cli/src/utils/ui-constants.ts` -- `cli/src/utils/renderer-cleanup.ts` - -### P3 — Input editing and global keyboard behavior (~3.0k LOC, 10 files) - -- `cli/src/components/chat-input-bar.tsx` -- `cli/src/components/multiline-input.tsx` -- `cli/src/components/input-cursor.tsx` -- `cli/src/hooks/use-chat-keyboard.ts` -- `cli/src/hooks/use-chat-input.ts` -- `cli/src/hooks/use-input-history.ts` -- `cli/src/utils/keyboard-actions.ts` -- `cli/src/utils/chat-input-key-intercept.ts` -- `cli/src/utils/terminal-enter-detection.ts` -- `cli/src/utils/keypad-keys.ts` - -Sizing note: approximately 2,961 LOC; keep suggestion/search files in P4. - -### P4 — Suggestions, searchable lists, and command/history overlays (~2.0k LOC, 6 files) - -- `cli/src/hooks/use-suggestion-engine.ts` -- `cli/src/components/suggestion-menu.tsx` -- `cli/src/hooks/use-searchable-list.ts` -- `cli/src/components/selectable-list.tsx` -- `cli/src/components/command-palette-screen.tsx` -- `cli/src/components/prompt-history-search-screen.tsx` - -### P5 — Full-screen project/session/provider/model pickers (~2.85k LOC, 6 files) - -- `cli/src/components/project-picker-screen.tsx` -- `cli/src/components/chat-history-screen.tsx` -- `cli/src/components/model-route-picker.tsx` -- `cli/src/components/provider-picker-screen.tsx` -- `cli/src/components/plan-session-picker-screen.tsx` -- `cli/src/hooks/use-directory-browser.ts` - -Sizing note: close to the ~3k threshold because the model and provider pickers are 861 and 663 LOC. Keep picker data/config semantics out of this presentation pass. - -### P6 — Message composition, markdown handoff, and collapse state (~2.65k LOC, 11 files) - -- `cli/src/components/message-with-agents.tsx` -- `cli/src/components/message-block.tsx` -- `cli/src/components/message-footer.tsx` -- `cli/src/components/blocks/blocks-renderer.tsx` -- `cli/src/components/blocks/single-block.tsx` -- `cli/src/components/blocks/content-with-markdown.tsx` -- `cli/src/components/blocks/user-content-copy.tsx` -- `cli/src/components/collapse-button.tsx` -- `cli/src/utils/collapse-helpers.ts` -- `cli/src/utils/block-processor.ts` -- `cli/src/state/message-block-store.ts` - -### P7 — Nested agents, thinking, and grouped branches (~2.0k LOC, 10 files) - -- `cli/src/components/blocks/agent-block-grid.tsx` -- `cli/src/components/blocks/agent-list-branch.tsx` -- `cli/src/components/blocks/agent-branch-item.tsx` -- `cli/src/components/blocks/agent-branch-wrapper.tsx` -- `cli/src/components/blocks/implementor-row.tsx` -- `cli/src/components/blocks/thinking-block.tsx` -- `cli/src/components/thinking.tsx` -- `cli/src/components/blocks/tool-branch.tsx` -- `cli/src/components/blocks/tool-block-group.tsx` -- `cli/src/components/blocks/ask-user-branch.tsx` - -### P8 — Tool renderer framework and discovery/command tools (~1.56k LOC, 12 files) - -- `cli/src/components/tools/types.ts` -- `cli/src/components/tools/registry.ts` -- `cli/src/components/tools/tool-call-item.tsx` -- `cli/src/components/tools/discovery-output.tsx` -- `cli/src/components/tools/query-index.tsx` -- `cli/src/components/tools/code-search.tsx` -- `cli/src/components/tools/run-terminal-command.tsx` -- `cli/src/components/tools/spawn-agents.tsx` -- `cli/src/components/tools/list-directory.tsx` -- `cli/src/components/tools/read-subtree.tsx` -- `cli/src/components/tools/read-docs.tsx` -- `cli/src/components/tools/glob.tsx` - -### P9 — Edit/action tools, diffs, and interactive tool output (~2.6k LOC, 13 files) - -- `cli/src/components/tools/diff-viewer.tsx` -- `cli/src/components/tools/apply-patch.tsx` -- `cli/src/components/tools/str-replace.tsx` -- `cli/src/components/tools/edit-transaction.tsx` -- `cli/src/components/tools/read-files.tsx` -- `cli/src/components/tools/write-file.tsx` -- `cli/src/components/tools/run-file-change-hooks.tsx` -- `cli/src/components/tools/proposal-actions.tsx` -- `cli/src/components/tools/render-ui.tsx` -- `cli/src/components/tools/write-todos.tsx` -- `cli/src/components/tools/suggest-followups.tsx` -- `cli/src/components/tools/skill.tsx` -- `cli/src/components/tools/task-completed.tsx` - -### P10 — Theme resolution and terminal color capability (~2.1k LOC, 5 files) - -- `cli/src/hooks/use-theme.tsx` -- `cli/src/utils/theme-system.ts` -- `cli/src/utils/theme-config.ts` -- `cli/src/utils/terminal-color-detection.ts` -- `cli/src/types/theme-system.ts` - -### P11 — Markdown, text wrapping, links, images, and terminal transport (~2.4k LOC, 11 files) - -- `cli/src/utils/markdown-renderer.tsx` -- `cli/src/utils/syntax-highlighter.tsx` -- `cli/src/utils/word-wrap-utils.ts` -- `cli/src/components/terminal-link.tsx` -- `cli/src/components/highlighted-text.tsx` -- `cli/src/utils/clipboard.ts` -- `cli/src/utils/terminal-images.ts` -- `cli/src/utils/image-display.ts` -- `cli/src/components/image-card.tsx` -- `cli/src/components/image-thumbnail.tsx` -- `cli/src/utils/image-thumbnail.ts` - -### P12 — Interactive primitives, status/error UI, ask-user, and mode controls (~2.75k LOC, 15 files) - -- `cli/src/components/clickable.tsx` -- `cli/src/components/button.tsx` -- `cli/src/components/top-banner.tsx` -- `cli/src/components/bottom-banner.tsx` -- `cli/src/components/help-banner.tsx` -- `cli/src/components/status-bar.tsx` -- `cli/src/components/user-error-banner.tsx` -- `cli/src/components/validation-error-popover.tsx` -- `cli/src/components/ask-user/index.tsx` -- `cli/src/components/ask-user/components/accordion-question.tsx` -- `cli/src/components/ask-user/components/options-list.tsx` -- `cli/src/components/ask-user/components/question-option.tsx` -- `cli/src/components/ask-user/components/custom-answer-input.tsx` -- `cli/src/components/segmented-control.tsx` -- `cli/src/components/agent-mode-toggle.tsx` - -### P13 — Animation and rerender instrumentation (~0.53k LOC, 5 files) - -- `cli/src/hooks/use-why-did-you-update.ts` -- `cli/src/utils/yield-to-event-loop.ts` -- `cli/src/hooks/use-now.ts` -- `cli/src/hooks/use-sheen-animation.tsx` -- `cli/src/components/shimmer-text.tsx` - -## Related tests - -The tests below are the most direct presentation verification surfaces. They are intentionally not one giant subshard: the combined set is well over 3k LOC, and several individual files exceed 600-1,200 LOC. - -### Layout, resize, scrolling, and performance - -- `cli/src/__tests__/rerender-perf.integration.test.ts` -- `cli/src/components/__tests__/grid-layout.test.tsx` — 1,033 LOC; oversized single test file -- `cli/src/components/__tests__/grid-layout.integration.test.tsx` -- `cli/src/components/__tests__/subagent-card-layout.test.tsx` -- `cli/src/components/__tests__/agent-branch-overflow.test.tsx` -- `cli/src/hooks/__tests__/use-terminal-layout.test.ts` — 675 LOC -- `cli/src/hooks/__tests__/use-grid-layout.test.ts` -- `cli/src/utils/__tests__/layout-helpers.test.ts` -- `cli/src/utils/__tests__/text-layout.test.ts` - -### Input, keyboard reachability, search, and pickers - -- `cli/src/components/__tests__/multiline-input.test.tsx` -- `cli/src/hooks/__tests__/use-chat-input.test.ts` -- `cli/src/hooks/__tests__/use-input-history.test.ts` -- `cli/src/hooks/__tests__/use-suggestion-engine.test.ts` -- `cli/src/hooks/__tests__/use-suggestion-engine-mention.test.ts` -- `cli/src/utils/__tests__/keyboard-actions.test.ts` — 650 LOC -- `cli/src/utils/__tests__/chat-input-key-intercept.test.ts` -- `cli/src/utils/__tests__/terminal-enter-detection.test.ts` -- `cli/src/components/__tests__/selectable-list.test.ts` -- `cli/src/components/__tests__/command-palette-screen.test.ts` -- `cli/src/components/__tests__/prompt-history-search-screen.test.ts` -- `cli/src/components/__tests__/plan-session-picker-screen.test.ts` -- `cli/src/__tests__/utils/project-picker.test.ts` - -### Messages, agents, tools, collapse, and markdown - -- `cli/src/components/__tests__/message-with-agents.test.tsx` — 765 LOC -- `cli/src/components/__tests__/message-block.streaming.test.tsx` -- `cli/src/components/__tests__/message-block.completion.test.tsx` -- `cli/src/utils/__tests__/block-processor.test.ts` -- `cli/src/utils/__tests__/message-block-helpers.test.ts` -- `cli/src/utils/__tests__/collapse-helpers.test.ts` — 1,272 LOC; oversized single test file -- `cli/src/utils/__tests__/markdown-renderer.test.tsx` -- `cli/src/components/tools/__tests__/registry-metadata.test.ts` -- `cli/src/components/tools/__tests__/diff-viewer.test.tsx` -- `cli/src/components/tools/__tests__/apply-patch.test.tsx` -- `cli/src/components/tools/__tests__/str-replace.test.tsx` -- `cli/src/components/tools/__tests__/edit-transaction.test.tsx` -- `cli/src/components/tools/__tests__/read-files.test.tsx` -- `cli/src/components/tools/__tests__/discovery-tools.test.tsx` -- `cli/src/components/tools/__tests__/query-index.test.tsx` -- `cli/src/components/tools/__tests__/code-search.test.tsx` -- `cli/src/components/tools/__tests__/run-terminal-command.test.ts` -- `cli/src/components/tools/__tests__/render-ui.test.tsx` - -### Theme, interaction primitives, and end-to-end terminal rendering - -- `cli/src/utils/__tests__/terminal-color-detection.test.ts` -- `cli/src/__tests__/unit/copy-button.test.ts` -- `cli/src/__tests__/unit/segmented-control.test.ts` -- `cli/src/__tests__/unit/agent-mode-toggle.test.ts` -- `cli/src/components/__tests__/status-indicator.test.tsx` -- `cli/src/components/__tests__/user-error-banner.test.tsx` -- `cli/src/components/ask-user/__tests__/multiple-choice-form.test.ts` -- `cli/src/components/ask-user/__tests__/validation.test.ts` -- `cli/src/__tests__/integration-tmux.test.ts` -- `cli/src/__tests__/e2e-cli.test.ts` - -## Related docs and integration metadata - -- `docs/architecture.md` — canonical CLI entry flow and package responsibility -- `docs/testing.md` — tmux-based render verification and capture expectations -- `docs/agents-and-tools.md` — current nested agent/tool block rendering contract and narrow-width behavior -- `cli/knowledge.md` — OpenTUI flex/resize rules, text-node constraints, clickable primitives, markdown wrapping, reconciliation hazards, menu navigation, and streaming-markdown policy -- `cli/tmux.knowledge.md` — bracketed-paste requirement, terminal captures, ANSI capture, and OpenTUI keyboard behavior under tmux -- `cli/README.md` — user-facing CLI capability summary -- `WINDOWS.md` — Windows terminal/shell compatibility expectations -- `cli/package.json` — OpenTUI 0.2.2, React 19, React reconciler, terminal-image, string-width, Yoga, and Zustand integration versions -- `cli/src/types/react19-compat.d.ts` — local React/OpenTUI JSX compatibility declaration - -## Key symbols and render flows - -1. **Process to root surface:** `main()` / `createRoot(renderer).render(...)` in `index.tsx` → `App` → `Chat`. -2. **Chat layout:** `Chat` calls `useChatUI`, renders the OpenTUI ``, maps `visibleTopLevelMessages` into `MessageWithAgents`, then renders `StatusBar` and `ChatInputBar` in a non-scrolling footer. -3. **Dimensions and responsive state:** OpenTUI renderer dimensions → `useTerminalDimensions` → `useTerminalLayout` / `useTerminalBreakpoints` / `useGridLayout` → component width, compact-height, narrow-width, and column decisions. -4. **Scrolling:** `useChatUI` owns the `ScrollBoxRenderable` ref → `useChatScrollbox` controls sticky/latest/page scrolling → `createChatScrollAcceleration` supplies inertial acceleration → `StatusBar` exposes the scroll-to-bottom affordance. -5. **Keyboard and input:** OpenTUI `useKeyboard` is used by `MultilineInput`, `useChatKeyboard`, provider/ask-user screens, and related overlays. `Chat` builds `ChatKeyboardState`/handlers; `keyboard-actions.ts` and `chat-input-key-intercept.ts` decide routing among input editing, suggestions, agent focus, scrolling, collapse, and overlays. -6. **Suggestions and search:** `useSuggestionEngine` derives slash/agent/file contexts → `SuggestionMenu`; full-screen command/history/picker screens share `SelectableList`, `useSearchableList`, and focused `` renderables. -7. **Message rendering:** `MessageWithAgents` → `MessageBlock` → `BlocksRenderer` → `processBlocks` routes reasoning, tools, implementors, agent groups, images, and single blocks. -8. **Nested agents:** `AgentBlockGrid` computes responsive groups → `AgentBranchWrapper` owns agent body/preview state → `AgentBranchItem` renders the bordered collapsible surface; nested blocks recurse through the same flow. -9. **Tool rendering:** `ToolBlockGroup` → `ToolBranch` → `renderToolComponent` in the registry → custom `ToolComponent` render config or fallback `ToolCallItem`; collapse/streaming previews and `availableWidth` flow through this boundary. -10. **Markdown/text:** `ContentWithMarkdown` selects `renderMarkdown` or `renderStreamingMarkdown`; output must remain valid within OpenTUI text-node constraints. Width is propagated into code blocks, tables, wrapping, diffs, links, and nested cards. -11. **Theme/color:** `initializeThemeStore` / `detectSystemTheme` combine OSC detection, environment/IDE/platform detection, manual overrides, and theme config → `useTheme` → `createMarkdownPalette` and component color props. -12. **Terminal compatibility:** focus-reporting escape sequences, OSC color queries, OSC52 clipboard, truecolor checks, bracketed paste, keypad/enter normalization, terminal-image sizing, renderer cleanup, and Windows/IDE terminal inference span P1/P3/P10/P11. -13. **Rerender boundaries:** `memo` is pervasive in message/agent/tool/layout components; `BlocksRenderer` stores changing props in a ref behind stable processor handlers; `message-block-store` supplies context/callbacks; `rerender-perf.integration.test.ts` is the direct regression surface. - -## Explicit exclusions - -- Existing audit reports, findings, manifests, or session conclusions, including everything under `.agents/sessions/**/findings/`; this selection was made from source/docs/tests only. -- `node_modules`, built binaries, `dist/`, release bundles/wrappers, generated agent-source bundles, lockfiles, and generated declarations other than the small React compatibility declaration named above. -- SDK/provider/runtime correctness, model routing, authentication, analytics, billing/cost computation, indexing internals, and tool execution semantics except where their already-produced data is rendered by the CLI. -- Command implementation semantics in `cli/src/commands/`; command palette presentation is included, but command execution behavior belongs to another shard. -- Chat streaming/event processing, queue ownership, persistence, file mutation, and API/network behavior except for their direct render/store subscription boundaries. -- Project discovery, provider config mutation, and plan artifact semantics behind picker screens; only their visible layout, focus, keyboard, validation display, and navigation surfaces are in scope here. -- Image decoding/resizing/upload internals; only terminal sizing/transport and rendered cards/thumbnails are included. -- Init/release/build scripts and packaging, except `cli/package.json` as OpenTUI integration metadata. diff --git a/.agents/sessions/audit-cli-next-level-2026-07/manifests/runtime-state.md b/.agents/sessions/audit-cli-next-level-2026-07/manifests/runtime-state.md deleted file mode 100644 index c05968bed8..0000000000 --- a/.agents/sessions/audit-cli-next-level-2026-07/manifests/runtime-state.md +++ /dev/null @@ -1,253 +0,0 @@ -# Runtime/state file-picker manifest - -## Scope - -Independent source map for the CLI message lifecycle: user submission, SDK run setup, streaming/event translation, cancellation, queue ownership, retry/timeout/error display, chat/session/history persistence, tool-progress rendering, and the CLI-to-SDK/runtime contracts beneath those flows. This is a file-selection artifact only; it contains no audit findings or feature recommendations. - -The primary groups below are sized for later review at roughly 5–15 files and normally below ~3,000 source LOC. LOC figures are approximate `wc -l` totals and exclude the related tests/docs lists. - -## Primary files by subshard - -### RS-1 — Send lifecycle and run setup (~1,962 LOC; 9 files) - -- `cli/src/hooks/use-send-message.ts` -- `cli/src/hooks/helpers/send-message.ts` -- `cli/src/utils/create-run-config.ts` -- `cli/src/utils/codebuff-client.ts` -- `cli/src/utils/send-message-helpers.ts` -- `cli/src/utils/send-message-timer.ts` -- `cli/src/types/contracts/send-message.ts` -- `cli/src/utils/yield-to-event-loop.ts` -- `cli/src/project-files.ts` - -Key symbols/flow: `useSendMessage` -> provider/agent preparation -> `prepareUserMessage` -> `setupStreamingContext` -> `createEventHandlerState`/`createRunConfig` -> `OpenbuffClient.run`; run ownership, checkpoint resume, `AbortController`, completion/error cleanup, queue finalization, timer outcomes, current chat ID. - -### RS-2 — Stream state and SDK-event translation (~1,865 LOC; 8 files) - -- `cli/src/hooks/use-chat-streaming.ts` -- `cli/src/hooks/stream-state.ts` -- `cli/src/utils/sdk-event-handlers.ts` -- `cli/src/utils/stream-chunk-processor.ts` -- `cli/src/utils/create-event-handler-state.ts` -- `cli/src/utils/message-updater.ts` -- `cli/src/utils/tool-result-normalizer.ts` -- `common/src/types/print-mode.ts` - -Key symbols/flow: `useChatStreaming`, `createStreamController`, `createEventHandler`, `createStreamChunkHandler`, root/subagent chunk routing, batched message mutation, event-to-block correlation, `tool_call`/`tool_start`/`tool_result`, phases, context-window/compaction, finish/cost, streaming-agent membership, and the discriminated `PrintModeEvent` contract. - -### RS-3 — Queue, keyboard, submit, and cancellation controls (~2,872 LOC; 8 files) - -- `cli/src/chat.tsx` -- `cli/src/hooks/use-message-queue.ts` -- `cli/src/hooks/use-queue-controls.ts` -- `cli/src/hooks/use-queue-ui.ts` -- `cli/src/hooks/use-chat-input.ts` -- `cli/src/hooks/use-chat-keyboard.ts` -- `cli/src/hooks/use-exit-handler.ts` -- `cli/src/utils/chat-input-key-intercept.ts` - -Key symbols/flow: `Chat` composition root; `useMessageQueue`, processing-owner symbols and 60s watchdog; queue pause/resume/clear; stream status gates; submit-vs-enqueue decisions; Escape/Ctrl-C handling; exit confirmation; input interception while overlays or tool interactions are active. - -### RS-4 — Chat store, history, checkpoints, and persisted run state (~2,246 LOC; 10 files) - -- `cli/src/hooks/use-chat-state.ts` -- `cli/src/hooks/use-chat-messages.ts` -- `cli/src/state/chat-store.ts` -- `cli/src/state/chat-history-store.ts` -- `cli/src/state/message-block-store.ts` -- `cli/src/utils/message-history.ts` -- `cli/src/utils/chat-history.ts` -- `cli/src/utils/run-state-storage.ts` -- `cli/src/types/chat.ts` -- `cli/src/types/chat-state.ts` - -Key symbols/flow: Zustand/Immer `useChatStore`; message snapshots/undo/redo; UI refs versus store state; batched/visible message loading; prompt history; per-chat directory discovery/deletion; `run-state.json`, `chat-messages.json`, `turn-checkpoint.json`; reload/continue-chat restoration; `ChatMessage`/content-block shape. - -### RS-5 — CLI resilience and user-facing status/error state (~1,481 LOC; 7 files) - -- `cli/src/hooks/use-connection-status.ts` -- `cli/src/hooks/use-timeout.ts` -- `cli/src/utils/error-handling.ts` -- `cli/src/utils/error-messages.ts` -- `cli/src/utils/format-timeout.ts` -- `cli/src/utils/openbuff-provider.ts` -- `cli/src/utils/validation-error-helpers.ts` - -Key symbols/flow: online/offline connection observation, named timer registry/cleanup, SDK error sanitization/status extraction, retry-banner formatting, provider readiness/discovery errors before a run starts, network validation IDs, and timeout text presented by the CLI. - -### RS-6 — Tool/subagent progress rendering (~2,450 LOC; 11 files) - -- `cli/src/components/message-block.tsx` -- `cli/src/components/message-with-agents.tsx` -- `cli/src/components/blocks/tool-branch.tsx` -- `cli/src/components/blocks/tool-block-group.tsx` -- `cli/src/components/tools/registry.ts` -- `cli/src/components/tools/types.ts` -- `cli/src/components/tools/tool-call-item.tsx` -- `cli/src/components/status-bar.tsx` -- `cli/src/utils/status-indicator-state.ts` -- `cli/src/components/progress-bar.tsx` -- `cli/src/components/message-footer.tsx` - -Key symbols/flow: streaming/completed message presentation, agent child grids, queued/pending/active tool state, renderer registry and fallback disposition, elapsed/cost/context status, auth/retry/unreachable/queue status derivation, completion footer, and progress visualization. - -### RS-7 — SDK run/session contract (~2,690 LOC; 5 files) - -- `sdk/src/client.ts` -- `sdk/src/run.ts` -- `sdk/src/run-state.ts` -- `common/src/types/session-state.ts` -- `common/src/types/messages/codebuff-message.ts` - -Key symbols/flow: `OpenbuffClient.run`; `RunOptions`/`OpenbuffClientOptions`; `handleEvent`; `previousRun`; composed user and timeout abort signals; cancellation-state reconstruction; filesystem mutation callbacks; `RunState`; initial/restored `SessionState`; and `mainAgentState.messageHistory`. - -This group is close to the ~3k ceiling. Do not add more source files without splitting it. - -### RS-8 — Provider streaming, retry, timeout, and abort classification (~2,424 LOC; 4 files) - -- `sdk/src/impl/llm.ts` -- `sdk/src/retry-config.ts` -- `common/src/types/contracts/llm.ts` -- `common/src/util/error.ts` - -Key symbols/flow: `promptAiSdkStream`, transient-network and HTTP retry classification, no-retry-after-yield rule, exponential backoff/jitter, abort-aware delay, OAuth refresh/failover paths, post-stream metadata timeout, provider stream chunk contract, and shared abort/error utilities. - -This is intentionally a 4-file group because `llm.ts` and `error.ts` are large; adding a fifth substantial source would push it toward the pruning threshold. - -### RS-9 — Agent-runtime stream/event production (~2,993 LOC; 5 files) - -- `packages/agent-runtime/src/run-agent-step.ts` -- `packages/agent-runtime/src/main-prompt.ts` -- `packages/agent-runtime/src/prompt-agent-stream.ts` -- `packages/agent-runtime/src/tool-stream-parser.ts` -- `packages/agent-runtime/src/tools/stream-parser.ts` - -Key symbols/flow: `runAgentStep`/`loopAgentSteps`; prompt stream consumption; response chunk callbacks; text/reasoning/tool-call parsing; parallel tool ordering; abort propagation; subagent and phase events; runtime errors translated into `PrintModeEvent` payloads. - -This group is effectively at the ~3k ceiling. Review it as listed; if tests or `tools/tool-executor.ts` are promoted into primary scope, split the group first. - -### RS-10 — Source entry and workspace SDK resolution (~1,715 LOC; 9 files) - -- `cli/src/index.tsx` -- `cli/tsconfig.json` -- `cli/package.json` -- `sdk/package.json` -- `sdk/src/index.ts` -- `sdk/src/tools/index.ts` -- `sdk/src/tools/find-files-matching-content.ts` -- `package.json` -- `bunfig.toml` - -Key symbols/flow: CLI source entry/import graph, `@openbuff/sdk` workspace mapping to `sdk/src/index.ts`, SDK barrel exports, eager parsing of tool modules, package export conditions, and root workspace scripts. - -Baseline contract evidence supplied by the parent audit: `bun run cli/src/index.tsx --help` currently fails to parse `sdk/src/tools/find-files-matching-content.ts:473` (`continue` outside a loop), while the already-built `cli/bin/openbuff --help` succeeds. The source file and resolution chain above must be reviewed together; the packaged binary itself is excluded as generated output. - -### RS-11 — Binary/package build and smoke path (~1,955 LOC including workflow YAML; 6 files) - -- `cli/scripts/build-binary.ts` -- `cli/scripts/smoke-binary.ts` -- `cli/src/pre-init/tree-sitter-wasm.ts` -- `cli/src/native/ripgrep.ts` -- `sdk/scripts/build.ts` -- `.github/workflows/cli-release-build.yml` - -Key symbols/flow: SDK prebuild before CLI compilation, source entry chosen for `bun build --compile`, production conditions/defines, embedded tree-sitter and ripgrep assets, binary startup smoke checks, and release-job invocation of the same build script. - -## Related tests - -### Send/setup tests - -- `cli/src/hooks/helpers/__tests__/send-message.test.ts` (1,745 LOC) -- `cli/src/utils/__tests__/send-message-helpers.test.ts` (1,762 LOC) -- `cli/src/utils/__tests__/send-message-timer.test.ts` -- `cli/src/hooks/__tests__/use-send-message-timer.test.ts` -- `cli/src/__tests__/helpers/mock-api-client.ts` - -Do not place the two 1.7k test files in one audit group: together they exceed ~3.5k LOC before implementation files. - -### Stream/event and rendering tests - -- `cli/src/utils/__tests__/sdk-event-handlers.test.ts` -- `cli/src/utils/__tests__/message-updater.test.ts` -- `cli/src/components/__tests__/message-block.streaming.test.tsx` -- `cli/src/components/__tests__/message-block.completion.test.tsx` -- `cli/src/components/__tests__/message-with-agents.test.tsx` -- `cli/src/components/__tests__/status-indicator.test.tsx` -- `cli/src/components/tools/__tests__/registry-metadata.test.ts` - -### Queue/input/cancellation tests - -- `cli/src/hooks/__tests__/use-queue-controls.test.ts` -- `cli/src/hooks/__tests__/use-chat-input.test.ts` -- `cli/src/utils/__tests__/chat-input-key-intercept.test.ts` -- `cli/src/hooks/__tests__/use-connection-status.test.ts` -- `cli/src/hooks/__tests__/use-timeout.test.ts` -- `cli/src/utils/__tests__/osc-timeout-scenarios.test.ts` - -### History/persistence tests - -- `cli/src/utils/__tests__/chat-history.test.ts` -- `cli/src/utils/__tests__/run-state-storage.test.ts` -- `cli/src/utils/__tests__/turn-checkpoint.test.ts` -- `cli/src/hooks/__tests__/use-input-history.test.ts` -- `cli/src/components/__tests__/prompt-history-search-screen.test.ts` - -### SDK/runtime contract tests - -- `sdk/src/__tests__/run-cancellation.test.ts` (1,394 LOC) -- `sdk/src/__tests__/run-error-preserves-history.test.ts` -- `sdk/src/__tests__/run-handle-event.test.ts` -- `sdk/src/__tests__/retry-config.test.ts` -- `sdk/src/impl/__tests__/failover-integration.test.ts` -- `sdk/src/impl/__tests__/llm-chatgpt-oauth-policy.test.ts` -- `packages/agent-runtime/src/__tests__/loop-agent-steps-abort.test.ts` -- `packages/agent-runtime/src/__tests__/subagent-streaming.test.ts` -- `packages/agent-runtime/src/__tests__/subagent-timeout.test.ts` -- `packages/agent-runtime/src/__tests__/spawn-agents-message-history.test.ts` -- `packages/agent-runtime/src/__tests__/stream-parser-abort.test.ts` -- `packages/agent-runtime/src/__tests__/stream-parser-parallelism.test.ts` -- `packages/agent-runtime/src/__tests__/stream-parser-reasoning.test.ts` - -### SDK e2e event/continuation tests - -- `sdk/e2e/utils/event-collector.ts` -- `sdk/e2e/utils/__tests__/event-collector.test.ts` -- `sdk/e2e/integration/event-ordering.integration.test.ts` -- `sdk/e2e/integration/event-types.integration.test.ts` -- `sdk/e2e/integration/stream-chunks.integration.test.ts` -- `sdk/e2e/streaming/concurrent-streams.e2e.test.ts` -- `sdk/e2e/streaming/subagent-streaming.e2e.test.ts` -- `sdk/e2e/workflows/error-recovery.e2e.test.ts` -- `sdk/e2e/workflows/multi-turn-conversation.e2e.test.ts` - -### Source/binary contract tests - -- `cli/src/__tests__/cli-args.test.ts` -- `cli/src/__tests__/release-wrapper.test.ts` -- `cli/src/__tests__/release/proxy-http-get.test.ts` -- `sdk/src/__tests__/find-files-matching-content.test.ts` -- `sdk/smoke-test-dist.ts` -- `sdk/scripts/verify.ts` -- `sdk/test/esm-compatibility/test-imports.js` -- `sdk/test/cjs-compatibility/test-imports.js` -- `sdk/test/ripgrep-bundling/test-ripgrep.js` - -## Related docs - -- `docs/request-flow.md` — authoritative CLI -> SDK -> runtime -> provider lifecycle, streaming, state, and cancellation narrative. -- `docs/architecture.md` — package boundaries, `client.run()`, `handleSteps`, and stream routing. -- `docs/testing.md` — DI-over-mocking and tmux CLI validation expectations. -- `docs/development.md` — workspace/Bun development and binary workflow context. -- `docs/agents-and-tools.md` — tool/subagent lifecycle, step caps, and completion/gate behavior. -- `docs/local-mode.md` — BYOK/local provider routing assumptions used by run readiness and error handling. -- `docs/environment-variables.md` — process/env loading contract across CLI and SDK. - -## Explicit exclusions - -- `node_modules/**`, cache directories, coverage output, SDK `dist/**`, compiled assets, source maps, and generated bundles. -- `cli/bin/openbuff` and other compiled binaries: behavior may be smoke-tested, but the binary is generated and must not be source-audited. -- Generated agent/type-source files such as `cli/src/agents/bundled-agents.generated.*` and `cli/src/data/initial-agent-type-sources.generated.ts`. -- Existing audit findings, reports, and manifests under `.agents/sessions/**` other than the pinned structural `MAP.md`; this picker was independent of current audits. -- Individual tool-specific renderer implementations (`components/tools/read-files.tsx`, `apply-patch.tsx`, etc.) except the shared registry/types/fallback/progress surface listed in RS-6. -- Provider/model picker UX, OAuth setup screens, analytics/feedback/publishing, indexing internals, command palette/slash-command feature breadth, and terminal layout/theme concerns except where directly imported by the scoped lifecycle files. -- Agent prompt quality and tool semantics beyond the runtime event/state contract. `packages/agent-runtime/src/tools/tool-executor.ts` is adjacent but intentionally not primary in RS-9 because adding it would push that group well over ~4k LOC; assign it to a dedicated tool-execution shard if required. diff --git a/.agents/sessions/audit-read-write-flow-2026-07-11-independent/AUDIT-REPORT.md b/.agents/sessions/audit-read-write-flow-2026-07-11-independent/AUDIT-REPORT.md deleted file mode 100644 index bd93228483..0000000000 --- a/.agents/sessions/audit-read-write-flow-2026-07-11-independent/AUDIT-REPORT.md +++ /dev/null @@ -1,173 +0,0 @@ -# Independent audit: read/write tools and flow - -Date: 2026-07-11 - -## Conclusion - -Openbuff has a strong internal foundation for safe filesystem work: canonical path checks, sensitive-file policy, structured read results, read-before-edit state, mutation authority, receipts, bounded reads, and rollback modeling are all present. The largest gaps occur at boundaries between those mechanisms. - -The audit found six high-risk clusters: - -1. write paths advertise stronger concurrency/verification guarantees than they consistently enforce; -2. some read recovery instructions cannot succeed with the default filesystem adapter; -3. canonical result contracts conflict with accepted legacy schemas and are degraded across SDK/runtime/CLI handoffs; -4. structural and adjacent filesystem tools bypass the primary adapter/policy path; -5. cancellation, proposal, hook, partial-read, and unconfirmed states are not represented coherently to users; -6. the public SDK surface does not expose or negotiate the capabilities implemented internally. - -No critical vulnerability was established. Eight findings are high severity, two are medium-high, and six are medium. Confidence is high for all direct code-path findings. UX impact is high-confidence where render logic is deterministic, but frequency and visual prominence remain inference because no live provider-backed TUI session was run. - -## Ranked findings - -### 1. High — atomic commit capability is bypassed by common update paths - -**Evidence:** `CodebuffFileSystem` exposes `conditionalCommit`, but `write_file`, `str_replace`, and transaction update commits validate state and later call ordinary `writeFile` (`common/src/types/filesystem.ts:54-58`, `sdk/src/tools/change-file.ts:625-655`, `sdk/src/tools/change-file.ts:783-834`). In-process authority locks do not cover external processes or adapter users. - -**Inference:** another writer can change a file between final validation and write, after which Openbuff can overwrite unseen content and still issue a successful receipt. This is the highest-confidence correctness risk in the write path. - -**Improvement:** route updates through conditional commit using the prepared hash; expose a clearly weaker authority tier when an adapter cannot provide it. Add conditional delete/native move and, eventually, a native multi-path transaction capability. - -### 2. High — large-file read recovery is a dead end on the default filesystem - -**Evidence:** the default adapter is `fs.promises` (`sdk/src/run.ts:351`). Whole reads over 10 MB recommend retrying with a range, but oversized range reads require optional `readTextRange`; without it they return `unsupported` (`sdk/src/tools/read-files.ts:371-386`, `sdk/src/tools/read-files.ts:648-660`). Tests explicitly accept that failure. - -**Inference:** the default CLI/SDK path can tell an agent exactly how to recover and then reject the recovery attempt. Large source/generated files are therefore unreadable through the primary tool unless the host supplies a custom capability. - -**Improvement:** implement bounded line/byte range reading for the default Node adapter, or change the initial error to accurately describe the required capability and alternative route. - -### 3. High — symbol and structural reads lack per-selector failure isolation - -**Evidence:** `requestOptionalFile` throws for blocked, binary, unsupported-encoding, too-large, and I/O states (`sdk/src/run.ts:621-667`). `read_files.symbols`, `read_outline`, and `read_slices` await it without per-selector/state classification (`packages/agent-runtime/src/tools/handlers/tool/read-files.ts:176-200`, `packages/agent-runtime/src/tools/handlers/tool/read-outline.ts:55-93`, `packages/agent-runtime/src/tools/handlers/tool/read-slices.ts:44-53`). - -**Inference:** one bad symbol selector can reject a mixed batch after other reads succeeded, defeating the canonical partial-result design and losing actionable recovery information. - -**Improvement:** return typed optional-file results or catch/classify failures per selector, preserve earlier successes, and continue the batch. - -### 4. High — `mutation_v1` schemas and runtime normalization disagree - -**Evidence:** active mutation schemas accept canonical results and several legacy success/error shapes (`common/src/tools/params/tool/str-replace.ts:14-25`, `common/src/tools/params/tool/apply-patch.ts:9-23`, `common/src/tools/params/tool/edit-transaction.ts:170-194`). Metadata advertises `mutation_v1`, while runtime normalization rejects schema-valid legacy shapes as malformed/unconfirmed (`common/src/tools/metadata.ts:157-166`, `packages/agent-runtime/src/tools/tool-executor.ts:84-143`). A current test codifies this contradiction. - -**Inference:** schema consumers cannot rely on accepted output meaning accepted runtime evidence. Compatibility behavior is effectively split across two incompatible contracts. - -**Improvement:** make the canonical envelope the only active internal/runtime output; isolate legacy translation at an explicitly versioned boundary and test output exclusivity. - -### 5. High — mutation failures lose recovery detail and can leak proposed content - -**Evidence:** single-file mutation failures are collapsed toward generic `application_rejected` results (`sdk/src/tools/change-file.ts:108-138`). Failure logging prints the entire proposed content (`sdk/src/tools/change-file.ts:873-875`), and the runtime also logs full `write_file` content (`packages/agent-runtime/src/tools/handlers/tool/write-file.ts:278`). - -**Inference:** users/models lose the precise stale/policy/I/O recovery branch, while logs can capture secrets or proprietary source that was never successfully committed. - -**Improvement:** preserve typed sanitized error codes, retryability, and recovery; log only operation metadata, path fingerprint, and byte counts—never mutation payloads. - -### 6. High — transaction success is finalized before receipt verification completes - -**Evidence:** transaction code calls `finishCommit(...succeeded: true)` before `issueCommittedReceipt` re-reads and verifies final hashes (`sdk/src/tools/change-file.ts:342-418`). Authority state is terminal once committed, while observed-failure receipts can still be created for committed operations (`sdk/src/tools/filesystem-authority.ts:311-334`, `sdk/src/tools/filesystem-authority.ts:393-403`). - -**Inference:** a verification failure can trigger rollback logic after authority state already says committed, producing ambiguous lease/receipt/rollback semantics. - -**Improvement:** treat verification as part of the commit lease and finalize success only after verified receipt issuance; model verification failure as a distinct terminal state with explicit rollback evidence. - -### 7. High — canonical successful mutations lose their diff and create/update identity in the CLI - -**Evidence:** canonical results store action type/path/patch in `actions[]` (`sdk/src/tools/change-file.ts:74-103`). CLI diff and create detection still primarily inspect legacy top-level fields/messages (`cli/src/utils/implementor-helpers.ts:615-675`, `cli/src/utils/implementor-helpers.ts:824-840`). Existing component fixtures use legacy payloads. - -**Inference:** an applied canonical edit can display with no diff, and a created file can be labeled/counted as an edit. The primary user proof of what changed is lost even though the result contains it. - -**Improvement:** create one canonical mutation normalizer for cards, activity, and summaries using action-level `action`, paths, patch, receipt, and outcome. - -### 8. High — cancellation can misreport a mutation that later completes - -**Evidence:** interrupt handling marks unresolved writes terminally cancelled (`cli/src/utils/block-operations.ts:537-552`). A late authoritative result is attached without changing lifecycle (`cli/src/utils/sdk-event-handlers.ts:596-608`), and edit cards give cancellation precedence over canonical outcome (`cli/src/components/tools/str-replace.tsx:189-207`). The behavior is asserted by test. Native `apply_patch` also receives no abort signal (`sdk/src/run.ts:1084-1091`). - -**Inference:** the CLI can say “Cancelled before completion” even when disk changed. For signal-blind mutations, cancellation is partly a presentation state rather than a guarantee about side effects. - -**Improvement:** separate run interruption from authoritative mutation outcome; show applied-after-interrupt/not-applied/unconfirmed explicitly and reconcile unknown late states by re-reading. - -### 9. Medium-high — override range correlation can accept the wrong content/capability - -**Evidence:** same-path range results are correlated by index/kind/path rather than requested coordinates (`sdk/src/tools/read-files.ts:1035-1066`, `sdk/src/tools/read-files.ts:1115-1146`, `packages/agent-runtime/src/get-file-reading-updates.ts:217-236`). - -**Inference:** a reordered or incorrect override response can be associated with the wrong request and potentially mint/retain an edit capability for content that does not cover the requested range. - -**Improvement:** use explicit selector IDs and validate returned start/end/completeness against each request before accepting content or capability. - -### 10. Medium-high — `read_subtree` bypasses injected filesystem authority and mixes live/stale state - -**Evidence:** `read_subtree` directly uses host `node:fs`, not `CodebuffFileSystem`, `fileFilter`, canonical authorization, or the primary sensitive-file policy (`packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:1-25`, `packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:164-235`). It combines live filesystem paths with cached `fileTokenScores` symbols (`packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:130-139`, `packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:254-266`). - -**Inference:** virtual/remote/sandbox adapters can see a different tree than `read_files`, and recently changed files can show current paths with stale symbol data without provenance. - -**Improvement:** route subtree through the same adapter/policy path and expose current-vs-cached provenance or refresh symbols from current content. - -### 11. Medium — structured reads exist, but the SDK defaults to lossy legacy output - -**Evidence:** `filesystemResultFormat` defaults to `legacy-v0` (`sdk/src/run.ts:170-173`, `sdk/src/run.ts:346`). The mapper flattens ordered selectors, typed errors, truncation, recovery, and capabilities into path-keyed strings (`sdk/src/tools/read-files.ts:838-888`). - -**Inference:** new integrations are nudged toward the least observable and least evolvable ABI, and duplicate selectors for one path become ambiguous. - -**Improvement:** default new client usage to structured v1 and retain legacy behavior only behind explicit compatibility selection. - -### 12. Medium — `fsSource` is not a complete filesystem abstraction - -**Evidence:** read/write core uses the injected adapter, but `read_image`, file-change hooks, subtree, and search/process-backed paths use direct Node filesystem/process views or receive no adapter (`sdk/src/tools/read-image.ts:44-49`, `sdk/src/tools/file-change-hooks.ts:53-124`, `sdk/src/run.ts:1122-1160`). - -**Inference:** hosts using overlays, remote workspaces, browser adapters, or permission-mediating filesystems cannot assume all tools see or enforce the same authority boundary. - -**Improvement:** publish a capability/support matrix; route compatible tools through adapter capabilities and fail closed with typed unsupported results when host-backed operations are unavailable. - -### 13. Medium — proposal lifecycle is model-only and lacks a coherent review UX - -**Evidence:** proposal accept/reject/apply requires IDs, revisions, and base hashes (`common/src/tools/params/tool/proposal-actions.ts:12-65`). Preview tools have edit renderers, but lifecycle tools fall back to generic rendering and there are no direct user actions (`cli/src/components/tools/registry.ts:36-61`, `agents/base2/base2.ts:167-170`). Proposal `str_replace` also claims parity while omitting direct-tool options such as `atomic`, `occurrenceIndex`, and `skipIfMissing` (`common/src/tools/params/tool/propose-str-replace.ts:43-77`, `common/src/tools/params/tool/str-replace.ts:35-95`). - -**Inference:** users cannot directly inspect and control the CAS lifecycle, and some valid direct edits cannot be previewed with equivalent semantics. - -**Improvement:** add a proposal review/apply card with lifecycle state and actions; share replacement schemas/semantics or state the unsupported differences explicitly. - -### 14. Medium — unconfirmed mutations disappear from completion summaries - -**Evidence:** canonical `unconfirmed` outcomes are excluded from edited counts and are not added as failures/warnings; canonical `errors[]` is not used for summary classification (`cli/src/utils/completion-summary.ts:57-70`, `cli/src/utils/completion-summary.ts:143-151`, `cli/src/utils/completion-summary.ts:178-186`). A test intentionally expects zero edited and zero failed. - -**Inference:** the final status can look clean or be absent precisely when disk state is unknown and user reconciliation is required. - -**Improvement:** add explicit unconfirmed, rolled-back, and rollback-incomplete summary states with affected paths and a re-read/reconcile action. - -### 15. Medium — validation hooks are effectively invisible in the run UX - -**Evidence:** hook results support named commands, terminal outputs, errors, and validation statuses (`common/src/tools/params/tool/run-file-change-hooks.ts:43-61`). The CLI card mostly shows changed paths and inspects only the first result for special status; completion summaries only aggregate generic terminal commands (`cli/src/components/tools/run-file-change-hooks.tsx:32-65`, `cli/src/utils/completion-summary.ts:91-104`). - -**Inference:** configured validation can fail without an adequate per-hook card or final aggregate, and “no hooks configured” is not clearly distinguished from verified success. - -**Improvement:** render per-hook pass/fail/skipped rows with bounded output and include hook validation in completion summaries. - -### 16. Medium — tool reachability and mutation composition have product gaps - -**Evidence:** transactions cannot compose `replace_range`, `rewrite_symbol`, patch hunks, or whole-file overwrite primitives (`common/src/tools/params/tool/edit-transaction.ts:72-168`). Documentation says `apply_patch` is active and `apply_smart_patch` quarantined, while live quarantine is empty and primary agents expose smart patch (`common/src/tools/constants.ts:137-143`, `agents/base2/base2.ts:69-104`, `docs/request-flow.md:57-59`). `read_slices` is marked deprecated but remains prompt-visible and required by reachability tests (`common/src/tools/metadata.ts:173-175`, `agents/tool-reachability.test.ts:16-32`). - -**Inference:** agents must trade coordinated rollback for safer edit primitives, while contradictory tool availability/deprecation signals increase selection ambiguity and maintenance surface. - -**Improvement:** expand transaction action algebra, establish one generated capability/reachability manifest, and remove deprecated aliases from new prompts while retaining compatibility handling. - -## Cross-cutting feature improvements - -- Byte-safe, encoding-aware mutation APIs and metadata-preserving move/delete/rollback. Current transaction snapshots are UTF-8 text and cannot guarantee restoration of permissions, symlinks, ACLs, xattrs, or binary identity. -- A public filesystem capability preflight so tools can be filtered/annotated before model execution instead of failing late. -- A public, typed mutation event callback carrying actions, paths, hashes, receipt, operation ID, and awaitable delivery semantics; current `onFilesChanged(): void` is insufficient for precise host synchronization. -- First-class dry-run/preflight and receipt query/export APIs for destructive or multi-file changes. -- Output budgets/pagination for symbols and outlines, plus clustered reads for far-apart ranges in oversized files. -- A single canonical read envelope for outline/subtree/symbol tools and a single canonical mutation normalizer across runtime, SDK, CLI, and completion summaries. - -## Existing strengths - -- Canonical path containment, sensitive-file filtering, binary/encoding classification, bounded concurrency, truncation modeling, and ordered structured read results are well-developed in `read_files`. -- Filesystem authority provides policy phases, locks, leases, receipts, redaction, and optional stronger adapter capabilities. -- Runtime read-before-edit state and write barriers cover many stale-edit and scheduling cases. -- External mutation results are conservatively downgraded to unconfirmed instead of trusting self-certified receipts. -- Current tests cover a wide range of component behavior; the gaps are concentrated in cross-boundary and adversarial scenarios. - -## Validation and limits - -213 focused tests passed: 14 common, 47 agent-runtime, 101 SDK, 43 CLI, and 8 agents. This confirms the repository is currently internally consistent with the observed behavior; it does not negate the findings. Several tests explicitly lock in problematic compatibility or UX behavior. - -This was a source-and-test audit, not a live provider/TUI usability study. No hostile concurrent filesystem process was run, so race impact is derived from direct check-then-write code paths. Filesystem metadata behavior of external adapters is unknown. Existing audit reports and remediation plans were not read or used as evidence. - -See `COVERAGE-MATRIX.md` for subsystem, domain, test, and out-of-scope accounting. diff --git a/.agents/sessions/audit-read-write-flow-2026-07-11-independent/COVERAGE-MATRIX.md b/.agents/sessions/audit-read-write-flow-2026-07-11-independent/COVERAGE-MATRIX.md deleted file mode 100644 index d1c62a41fc..0000000000 --- a/.agents/sessions/audit-read-write-flow-2026-07-11-independent/COVERAGE-MATRIX.md +++ /dev/null @@ -1,59 +0,0 @@ -# Independent read/write audit — coverage matrix - -Date: 2026-07-11 - -This audit used live source, current tests, and current product documentation. Existing audit reports, remediation plans, `.agents/sessions/*` findings, `.omx/plans/*`, graveyard code, and evaluation logs were excluded as evidence. - -## Repository coverage - -| Area | Scope | Coverage | Result | -| ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | ---------------------: | ---------------------------------------------------------------------------------------------------------------------- | -| `common/` | Tool names, metadata, read/write schemas, canonical result envelopes, filesystem adapter types, generated tool definitions | Audited | Contract drift, missing transaction primitives, and reachability/deprecation gaps found | -| `packages/agent-runtime/` | Registration, scheduling, handlers, read-before-edit state, structural reads, output validation/normalization, proposal coordination | Audited | Symbol failure isolation, legacy/canonical normalization conflict, subtree authority bypass, and scheduler drift found | -| `sdk/` | Client options, dispatch, filesystem adapter, reads, mutations, authority/receipts, callbacks, hooks, public exports | Audited | Highest-risk concurrency, recovery, adapter, cancellation, and public API gaps found | -| `agents/` | Base/editor tool reachability, model-facing edit/read guidance, generated agent types | Audited | Patch reachability/docs drift, shell fallback guidance, deprecated tool exposure, and generated type drift found | -| `cli/` | Event lifecycle, result normalization, read/write cards, proposal UX, hooks, completion summaries | Audited | Canonical diff loss, misleading cancellation state, unconfirmed-result omission, and weak recovery UX found | -| `docs/` | Architecture, request flow, deterministic edit behavior, tools, testing, SDK-facing guidance | Relevant docs audited | Several live-behavior contradictions and SDK integration omissions found | -| Tests in the areas above | Focused unit/component/integration tests around read, write, authority, handlers, reachability, and rendering | Audited and executed | 213 focused tests passed; current tests permit the reported gaps | -| `packages/code-map/` | Structural-read dependency behavior | Sampled where called | Parser/cached-symbol behavior traced; package internals were not independently audited | -| `packages/indexer/` | `query_index`/cached context interaction with read flow | Sampled where adjacent | Index freshness implications noted; indexing architecture was not independently audited | -| `scripts/` | Tool-definition generation and registration checks | Sampled | Semantic parity checks are incomplete | - -## Explicitly out of scope - -| Area | Reason | -| -------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | -| `evals/` and evaluation logs | Product evaluation framework is not part of the read/write execution path; logs were excluded as historical evidence | -| `agents-graveyard/` | Inactive code | -| Existing `.agents/sessions/*` except this new audit directory | Prior audits/findings were expressly excluded as source material | -| `.omx/`, including plans and state | Prior plans and workflow state were expressly excluded as source material | -| Provider implementations and model quality | Audit concerns filesystem tools/flow, not provider correctness | -| Release pipelines, packaging infrastructure, and platform installers | Sampled only where needed to verify SDK public exports/docs | -| Live provider-backed TUI behavior | No real-provider interactive session was run; UI conclusions come from current event/render code and component tests | -| Third-party filesystem adapters outside this repository | Their capabilities and metadata preservation cannot be established from this repository | - -## Engineering-domain coverage - -| Domain | Evidence inspected | Outcome | -| ------------------ | -------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | -| Security | Canonical path policy, sensitive-file policy, adapter boundaries, shell-routing guidance, mutation logging | Full-content failure logging and authority-boundary inconsistencies found; no direct traversal exploit established | -| Correctness | Selector correlation, hashes/capabilities, conditional commits, transaction verification/rollback, canonical outputs | Multiple high-confidence defects found | -| State mutation | Read-before-edit, locks/leases, create/update/delete/move, rollback, cancellation, proposals | Lost-update and lifecycle/state-reporting gaps found | -| Error handling | Typed read errors, mutation normalization, recovery guidance, partial/unconfirmed/rollback outcomes | Errors are frequently collapsed or hidden at integration/UI boundaries | -| Performance | Read concurrency, output limits, oversized ranges, structural output, cancellation | Large-file and structural-output gaps found; no broad performance regression established | -| Dependency hygiene | Tool registries, generated types, public exports, Node filesystem bypasses | Multiple sources of truth and incomplete public surfaces found | -| Test coverage | 213 focused tests plus negative-edge inventory | Strong component baseline, but missing race, canonical-UI, adapter-parity, and failure-isolation tests | -| API/ABI | `read_v1`, `mutation_v1`, legacy compatibility, SDK options/overrides/callbacks | Significant advertised-vs-runtime contract drift found | - -## Validation totals - -| Package/surface | Passing tests | -| ------------------------ | ------------: | -| `common` | 14 | -| `packages/agent-runtime` | 47 | -| `sdk` | 101 | -| `cli` | 43 | -| `agents` | 8 | -| **Total** | **213** | - -Passing tests are evidence that the current suite accepts the observed behavior, not evidence that the gaps are harmless. The most important missing cases are external-writer races, receipt-verification failure after bytes are written, mixed symbol-selector failures, default-filesystem large-range recovery, canonical mutation rendering, late results after cancellation, and non-Node adapter parity. diff --git a/.agents/sessions/audit-read-write-flow-2026-07-11-independent/MAP.md b/.agents/sessions/audit-read-write-flow-2026-07-11-independent/MAP.md deleted file mode 100644 index 8bd6d0260f..0000000000 --- a/.agents/sessions/audit-read-write-flow-2026-07-11-independent/MAP.md +++ /dev/null @@ -1,350 +0,0 @@ -# Structural Map — openbuff - -- **Project root:** `/home/ben/Code/CLI/openbuff` -- **Built at:** 2026-07-11T13:02:02.679Z -- **Total files indexed:** 2160 -- **Graph:** 15492 nodes, 73423 edges - -> Pin this file in context. Every audit shard navigates from here instead of doing fuzzy round-trip discovery. - -## Entry points - -- `cli/src/index.tsx` -- `packages/code-map/src/index.ts` -- `packages/indexer/src/index.ts` -- `packages/internal/src/index.ts` - -## Directories (by size, biggest first) - -| dir | files | total size | top symbols | -| -------------------------- | ----- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `evals` | 597 | 12.4 MB | main, run, makeEvalRun, makeAgentResults, toolCall, log | -| `cli` | 427 | 3.0 MB | render, main, TestItem, tmux, createErrorMessage, formatTimestamp | -| `packages` | 314 | 2.3 MB | start, Greeting, greet, Greeter, flush, doGenerate | -| `sdk` | 163 | 1.4 MB | main, run, evaluate, createMockFs, log, resolveMcpEnv | -| `common` | 245 | 1.1 MB | createMockLogger, getStringProperty, getFileExtension, process, sleep, size | -| `agents` | 95 | 1.1 MB | extractInlineFunctionSource, parseGateStateBlock, feedJson, collectToolInputFiles, isFileChangingTool, hasEditArtifact | -| `scripts` | 67 | 515.7 KB | main, parseArgs, computeCost, ConversationMessage, TurnResult, makeConversationStreamRequest | -| `.omx` | 30 | 503.1 KB | — | -| `agents-graveyard` | 121 | 350.6 KB | createBase2WithTaskResearcher, getLatestEditToolResults, extractSpawnResults, getSpawnResults, createResearchImplementOrchestrator, createBase2Implementor | -| `.agents` | 41 | 344.4 KB | publisher, getSpawnerPrompt, getSystemPrompt, getInstructionsPrompt, getDefaultReviewModeInstructions, getWorkModeInstructions | -| `bun.lock` | 1 | 270.1 KB | — | -| `docs` | 13 | 204.8 KB | — | -| `.github` | 14 | 51.9 KB | — | -| `openbuff.d.example` | 4 | 22.5 KB | — | -| `LICENSE` | 1 | 11.1 KB | — | -| `.bin` | 1 | 8.5 KB | — | -| `README.zh-CN.md` | 1 | 8.1 KB | — | -| `README.md` | 1 | 8.0 KB | — | -| `WINDOWS.md` | 1 | 7.3 KB | — | -| `CONTRIBUTING.md` | 1 | 5.6 KB | — | -| `CODE_OF_CONDUCT.md` | 1 | 4.5 KB | — | -| `eslint.config.js` | 1 | 4.0 KB | — | -| `AGENTS.md` | 1 | 3.6 KB | — | -| `package.json` | 1 | 2.5 KB | — | -| `INFISICAL_SETUP_GUIDE.md` | 1 | 2.5 KB | — | -| `ROUTER.md` | 1 | 2.5 KB | — | -| `.env.example` | 1 | 1.7 KB | — | -| `tsconfig.json` | 1 | 839 B | — | -| `SECURITY.md` | 1 | 520 B | — | -| `.gitignore` | 1 | 487 B | — | -| `.vscode` | 1 | 438 B | — | -| `bunfig.toml` | 1 | 432 B | — | -| `.prettierrc` | 1 | 389 B | — | -| `tsconfig.base.json` | 1 | 386 B | — | -| `test` | 1 | 332 B | setup | -| `knowledge.md` | 1 | 287 B | — | -| `.e2e-scratch` | 2 | 279 B | add, greet, multiply | -| `NOTICE` | 1 | 156 B | — | -| `openbuff.json.example` | 1 | 118 B | — | -| `.envrc` | 1 | 30 B | — | -| `.bun-version` | 1 | 7 B | — | - -## Largest files per directory - -### `evals` - -- `evals/buffbench/logs/2026-07-04T17-30_base2/45-fork-read-files-base2-349a140.json` — 472.2 KB, 0 symbols -- `evals/buffbench/logs/2026-07-04T13-41_base2/2-add-deep-thinkers-base2-6c362c3.json` — 469.4 KB, 0 symbols -- `evals/buffbench/restrict-tool-types-base2-lite-error-ftj2.json` — 418.5 KB, 0 symbols -- `evals/buffbench/validate-custom-tools-base2-error-c6yk.json` — 418.3 KB, 0 symbols -- `evals/buffbench/logs/2026-07-04T17-30_base2/38-unify-agent-builder-base2-4852954.json` — 411.4 KB, 0 symbols - -### `cli` - -- `cli/bin/tree-sitter.wasm` — 200.7 KB, 0 symbols -- `cli/src/hooks/helpers/__tests__/send-message.test.ts` — 55.5 KB, 0 symbols -- `cli/src/utils/__tests__/message-block-helpers.test.ts` — 55.1 KB, 0 symbols -- `cli/src/chat.tsx` — 53.8 KB, 1 symbols -- `cli/src/utils/__tests__/send-message-helpers.test.ts` — 48.1 KB, 0 symbols - -### `packages` - -- `packages/agent-runtime/src/__tests__/read-files-edit-state.test.ts` — 126.5 KB, 2 symbols -- `packages/agent-runtime/src/__tests__/process-str-replace.test.ts` — 76.8 KB, 0 symbols -- `packages/agent-runtime/src/__tests__/run-programmatic-step.test.ts` — 66.6 KB, 0 symbols -- `packages/agent-runtime/src/process-str-replace.ts` — 66.5 KB, 30 symbols -- `packages/agent-runtime/src/run-agent-step.ts` — 54.3 KB, 4 symbols - -### `sdk` - -- `sdk/src/__tests__/model-provider.test.ts` — 80.0 KB, 3 symbols -- `sdk/src/provider-config.ts` — 72.5 KB, 30 symbols -- `sdk/src/tools/browser-logs.ts` — 63.4 KB, 30 symbols -- `sdk/src/impl/llm.ts` — 52.9 KB, 20 symbols -- `sdk/src/__tests__/run-cancellation.test.ts` — 44.5 KB, 1 symbols - -### `common` - -- `common/src/templates/initial-agents-dir/types/tools.ts` — 46.6 KB, 30 symbols -- `common/src/util/__tests__/messages.test.ts` — 41.4 KB, 0 symbols -- `common/src/tools/results/filesystem.ts` — 33.8 KB, 30 symbols -- `common/src/__tests__/agent-validation.test.ts` — 28.9 KB, 0 symbols -- `common/src/util/__tests__/saxy.test.ts` — 26.1 KB, 0 symbols - -### `agents` - -- `agents/base2/base2.ts` — 180.2 KB, 30 symbols -- `agents/__tests__/context-pruner.test.ts` — 120.8 KB, 2 symbols -- `agents/__tests__/base2.test.ts` — 111.9 KB, 5 symbols -- `agents/context-pruner.ts` — 74.0 KB, 30 symbols -- `agents/types/tools.ts` — 46.6 KB, 30 symbols - -### `scripts` - -- `scripts/test-fireworks-cache-intervals.ts` — 34.3 KB, 10 symbols -- `scripts/benchmark-providers.ts` — 33.1 KB, 16 symbols -- `scripts/test-fireworks-long.ts` — 30.7 KB, 6 symbols -- `scripts/test-canopywave-long.ts` — 29.0 KB, 5 symbols -- `scripts/test-siliconflow.ts` — 27.3 KB, 5 symbols - -### `.omx` - -- `.omx/ultragoal/worktree-baseline.json` — 247.5 KB, 0 symbols -- `.omx/ultragoal/brief.md` — 36.3 KB, 0 symbols -- `.omx/plans/prd-read-write-audit-remediation.md` — 36.3 KB, 0 symbols -- `.omx/state/todos-session.json` — 34.4 KB, 0 symbols -- `.omx/plans/traceability-read-write-audit-remediation.md` — 27.8 KB, 0 symbols - -### `agents-graveyard` - -- `agents-graveyard/base/base-prompts.ts` — 25.3 KB, 3 symbols -- `agents-graveyard/editor/best-of-n/editor-best-of-n.ts` — 18.5 KB, 4 symbols -- `agents-graveyard/base/ask.ts` — 12.2 KB, 0 symbols -- `agents-graveyard/registry/transform-agent.ts` — 12.0 KB, 0 symbols -- `agents-graveyard/base2/task-researcher/base2-with-task-researcher-planner-pro.ts` — 11.4 KB, 1 symbols - -### `.agents` - -- `.agents/types/tools.ts` — 46.6 KB, 30 symbols -- `.agents/sessions/read-write-tooling-2026-07-10/findings/structured-results-plan.md` — 35.6 KB, 0 symbols -- `.agents/sessions/audit-agent-ecosystem-2026-07/findings/runtime-picker.md` — 16.2 KB, 0 symbols -- `.agents/types/agent-definition.ts` — 16.1 KB, 3 symbols -- `.agents/sessions/audit-agent-ecosystem-2026-07/findings/discovery-picker.md` — 13.9 KB, 0 symbols - -### `bun.lock` - -- `bun.lock` — 270.1 KB, 0 symbols - -### `docs` - -- `docs/agents-and-tools.md` — 76.5 KB, 0 symbols -- `docs/codebuff-to-openbuff-migration.md` — 30.7 KB, 0 symbols -- `docs/configuration.md` — 20.6 KB, 0 symbols -- `docs/openbuff-provider-model-setup-ux.md` — 20.4 KB, 0 symbols -- `docs/architecture.md` — 12.3 KB, 0 symbols - -### `.github` - -- `.github/workflows/cli-release-build.yml` — 11.7 KB, 0 symbols -- `.github/workflows/cli-release-staging.yml` — 8.8 KB, 0 symbols -- `.github/workflows/ci.yml` — 7.4 KB, 0 symbols -- `.github/knowledge.md` — 5.4 KB, 0 symbols -- `.github/workflows/cli-release-prod.yml` — 4.9 KB, 0 symbols - -### `openbuff.d.example` - -- `openbuff.d.example/providers.json` — 19.0 KB, 0 symbols -- `openbuff.d.example/routes.json` — 2.1 KB, 0 symbols -- `openbuff.d.example/hooks.json` — 1.2 KB, 0 symbols -- `openbuff.d.example/indexing.json` — 146 B, 0 symbols - -### `LICENSE` - -- `LICENSE` — 11.1 KB, 0 symbols - -### `.bin` - -- `.bin/bun` — 8.5 KB, 0 symbols - -### `README.zh-CN.md` - -- `README.zh-CN.md` — 8.1 KB, 0 symbols - -### `README.md` - -- `README.md` — 8.0 KB, 0 symbols - -### `WINDOWS.md` - -- `WINDOWS.md` — 7.3 KB, 0 symbols - -### `CONTRIBUTING.md` - -- `CONTRIBUTING.md` — 5.6 KB, 0 symbols - -### `CODE_OF_CONDUCT.md` - -- `CODE_OF_CONDUCT.md` — 4.5 KB, 0 symbols - -### `eslint.config.js` - -- `eslint.config.js` — 4.0 KB, 0 symbols - -### `AGENTS.md` - -- `AGENTS.md` — 3.6 KB, 0 symbols - -### `package.json` - -- `package.json` — 2.5 KB, 0 symbols - -### `INFISICAL_SETUP_GUIDE.md` - -- `INFISICAL_SETUP_GUIDE.md` — 2.5 KB, 0 symbols - -### `ROUTER.md` - -- `ROUTER.md` — 2.5 KB, 0 symbols - -### `.env.example` - -- `.env.example` — 1.7 KB, 0 symbols - -### `tsconfig.json` - -- `tsconfig.json` — 839 B, 0 symbols - -### `SECURITY.md` - -- `SECURITY.md` — 520 B, 0 symbols - -### `.gitignore` - -- `.gitignore` — 487 B, 0 symbols - -### `.vscode` - -- `.vscode/settings.json` — 438 B, 0 symbols - -### `bunfig.toml` - -- `bunfig.toml` — 432 B, 0 symbols - -### `.prettierrc` - -- `.prettierrc` — 389 B, 0 symbols - -### `tsconfig.base.json` - -- `tsconfig.base.json` — 386 B, 0 symbols - -### `test` - -- `test/setup-scm-loader.ts` — 332 B, 1 symbols - -### `knowledge.md` - -- `knowledge.md` — 287 B, 0 symbols - -### `.e2e-scratch` - -- `.e2e-scratch/widget.ts` — 275 B, 3 symbols -- `.e2e-scratch/browser-agent-note.txt` — 4 B, 0 symbols - -### `NOTICE` - -- `NOTICE` — 156 B, 0 symbols - -### `openbuff.json.example` - -- `openbuff.json.example` — 118 B, 0 symbols - -### `.envrc` - -- `.envrc` — 30 B, 0 symbols - -### `.bun-version` - -- `.bun-version` — 7 B, 0 symbols - -## Most-imported files (likely key modules) - -| in-degree | file | -| --------- | -------------------------------------------------------------- | -| 101 | `packages/agent-runtime/src/__tests__/rewrite-symbol.test.ts` | -| 82 | `common/src/util/messages.ts` | -| 76 | `cli/src/utils/arrays.ts` | -| 75 | `common/src/types/bun-test.d.ts` | -| 66 | `cli/src/__tests__/release/proxy-http-get.test.ts` | -| 63 | `common/src/util/error.ts` | -| 59 | `common/src/tools/params/utils.ts` | -| 54 | `sdk/e2e/utils/event-collector.ts` | -| 49 | `cli/src/utils/message-block-helpers.ts` | -| 44 | `cli/src/hooks/use-theme.tsx` | -| 44 | `packages/agent-runtime/src/__tests__/main-prompt.test.ts` | -| 40 | `sdk/src/provider-config.ts` | -| 40 | `common/src/util/plan-artifacts.ts` | -| 40 | `sdk/src/tools/filesystem-authority.ts` | -| 38 | `packages/agent-runtime/src/tools/handlers/tool/write-file.ts` | -| 36 | `cli/src/project-files.ts` | -| 34 | `common/src/testing/mocks/timers.ts` | -| 34 | `agents/base2/base2.ts` | -| 32 | `scripts/test-canopywave-long.ts` | -| 31 | `.e2e-scratch/widget.ts` | -| 31 | `common/src/util/content-hash.ts` | -| 30 | `common/src/types/session-state.ts` | -| 28 | `common/src/util/string.ts` | -| 28 | `sdk/src/run.ts` | -| 28 | `cli/scripts/build-binary.ts` | - -## Cross-directory dependencies (architectural layering) - -| count | from → to | -| ----- | ----------------------------- | -| 251 | `packages` → `common` | -| 168 | `cli` → `common` | -| 133 | `sdk` → `common` | -| 125 | `cli` → `packages` | -| 61 | `cli` → `sdk` | -| 51 | `sdk` → `packages` | -| 25 | `cli` → `.e2e-scratch` | -| 23 | `packages` → `cli` | -| 23 | `cli` → `scripts` | -| 22 | `sdk` → `cli` | -| 20 | `packages` → `sdk` | -| 19 | `agents` → `sdk` | -| 19 | `agents` → `agents-graveyard` | -| 19 | `agents-graveyard` → `common` | -| 16 | `sdk` → `scripts` | -| 14 | `evals` → `common` | -| 12 | `scripts` → `packages` | -| 10 | `agents` → `common` | -| 9 | `evals` → `cli` | -| 8 | `common` → `packages` | -| 7 | `agents` → `cli` | -| 7 | `common` → `cli` | -| 7 | `evals` → `sdk` | -| 6 | `cli` → `agents` | -| 6 | `packages` → `agents` | -| 6 | `packages` → `.e2e-scratch` | -| 5 | `agents` → `packages` | -| 5 | `scripts` → `cli` | -| 5 | `evals` → `packages` | -| 5 | `common` → `sdk` | - -## Shard sizing hint - -Total indexed source: **23.6 MB** across **41** top-level directories. - -When sharding for an audit, aim for ~5–15 files per shard. Use the table above to group small dirs together and split huge dirs (e.g. split `src/` by subdirectory). diff --git a/.agents/sessions/background-job-push-model-2026-08/EVENTS.jsonl b/.agents/sessions/background-job-push-model-2026-08/EVENTS.jsonl deleted file mode 100644 index 97901f9069..0000000000 --- a/.agents/sessions/background-job-push-model-2026-08/EVENTS.jsonl +++ /dev/null @@ -1,3 +0,0 @@ -{"ts":"2026-08-02T10:29:36.463Z","kind":"append_lesson","summary":"Appended entry \"Session closed — M1–M5 delivered, R7 deferred\" to STATUS.md","payload":{"heading":"Session closed — M1–M5 delivered, R7 deferred","artifact":"STATUS.md"}} -{"ts":"2026-08-02T10:30:15.315Z","kind":"append_lesson","summary":"Appended entry \"Closure lessons\" to LESSONS.md","payload":{"heading":"Closure lessons","artifact":"LESSONS.md"}} -{"ts":"2026-08-02T10:30:15.315Z","kind":"session_status","summary":"Session status -> completed","payload":{"status":"completed"}} diff --git a/.agents/sessions/background-job-push-model-2026-08/LESSONS.md b/.agents/sessions/background-job-push-model-2026-08/LESSONS.md deleted file mode 100644 index 04c2aaf867..0000000000 --- a/.agents/sessions/background-job-push-model-2026-08/LESSONS.md +++ /dev/null @@ -1,120 +0,0 @@ -# Background Job Push Model & Drain Inversion — LESSONS - -## Decisions - -- **Push metadata, pull content.** The split is by KIND of data, not by timing - (running vs finished). Metadata is small, bucketable, change-gated → cache - safe. Content is unbounded and mutates every step → injecting it invalidates - the prompt-cache suffix with no ceiling. Fold the digest into the SAME per-step - observation block that already carries `git_status` so the cache-invalidation - boundary does not move versus today. -- **Fix ordering is a hard dependency chain, not a preference.** The checkJob - poll loop exists specifically to drive `readNewJobOutput`. It cannot be - removed (M3) before draining is inverted (M2). The digest's pending-lines - bucket is meaningless (M4) before output enters the registry on write (M2). - M1 is independent and cheap, so it goes first. -- **Tee for live, file-read for recovery.** Live jobs get a real in-process - pipe + line splitter (write-time timestamps, per-line events). The log-file - read path is retained ONLY as the cross-session recovery fallback, where no - in-process pipe can exist by definition. This preserves the durable disk - projection without letting it dictate live semantics. -- **The digest is declarative state, not a prompt.** Fixed "No action required - unless you need this output." line, asserted in a test. It never creates a - step, never wakes a finished turn, never asks a question — so a chatty dev - server cannot interrupt the agent mid-task. - -## Gotchas / traps - -- `truncatedAtCursor` returning `first.sequence <= cursor` is true for a HEALTHY - buffer — that is why every `check_job` reported `truncated:true, dropped:0`. - The fix must key off events actually EVICTED past the cursor, not the oldest - retained event. -- `CHECK_JOB_OUTPUT_LIMIT` (50_000) only bounds the `wait_for` match window via - `appendBoundedCollected`; the returned `events` array is separately unbounded. - Both need bounding. -- Event `timestamp` is currently POLL time, not write time (blob emission). - Per-line teeing is what makes timestamps and the pending bucket truthful. -- Switching `startBackgroundJob` stdio from `['ignore', outFd, outFd]` to a pipe - is the most safety-critical edit: it touches detach, log-quota monitor, and - kill/exit settlement. Change the read source only; leave quota + settle paths - intact. -- `sdk/src/__tests__/check-job.test.ts` asserts `readNewJobOutput` output - directly (e.g. `expect(readNewJobOutput(job)).toBe('hello\n')`) — these lock - in blob semantics and need REWORK, not extension, once draining is inverted. - -## Corrections to earlier framing - -- Earlier claim: "log file as source of truth contradicts PLAN.md's 'disk is a - projection'." Half right. That invariant holds for LIFECYCLE state (genuinely - registry-owned). It was never stated for OUTPUT. This is an UNSPECIFIED area - of the unified-background-jobs design, not a violated invariant. - -## Open risk to watch during M4 - -- If a job settles and its digest entry is dropped by `agentStep` TTL before the - agent acts on it, the agent would never learn it finished. Requires a - settlement tombstone that persists until acknowledged — not just a state read. - - - -## Closure lessons — 2026-08-02T10:30:15.314Z - -### Deviations accepted at closure (2026-08-02) - -**Behavioral goal vs. specced mechanism are separable — and the goal is what matters.** -All three deviations (D1 interval drainer vs piped stdio, D2 retained `while` loop -with `wait()` as the wake, D3 real `list_jobs` + SDK gate vs bespoke base2 block) -satisfy the underlying requirement while diverging from the DESIGN's prescribed -mechanism. Writing the plan's DESIGN section as an implementation transcript -("switch stdio to a pipe", "delete the loop") rather than as an observable -contract ("an unpolled job still accrues events and settles") made these look -like failures when they were tradeoffs. Future plans should state DESIGN as -requirements + acceptance criteria, and keep prescribed mechanisms clearly marked -as one candidate implementation. - -**D1 specifically: the risk calculus beat the elegance.** -Piped stdio was the cleanest way to get write-time timestamps, but it touches -detach, the log-quota monitor, and kill/exit settlement simultaneously — the three -most failure-prone paths in `startBackgroundJob`. A 250 ms interval drainer got -R3 (per-line events) and R4 (drain independent of `check_job`) with a fraction of -the blast radius. The cost is honest and bounded: timestamps are drain-time at -≤250 ms granularity, so sub-250ms bursts collapse into one tick. That was worth -it; a lost line or broken detach would not have been. - -**D3 turned out better than the spec.** -Routing the digest through the real `list_jobs` tool instead of a hand-assembled -base2 block gave one source of truth for the digest shape (the tool's own schema -and tests), and the SDK-side fingerprint gate applies to model-initiated -`list_jobs` calls too — not just the pushed ones. When a plan specifies a bespoke -formatter that duplicates an existing tool's output, prefer the tool. - -### The open risk flagged during planning came true - -This file's pre-existing "Open risk to watch during M4" entry predicted exactly -what happened: the settlement tombstone (R7) was the piece that did not get -built. Flagging a risk in LESSONS is not the same as tracking it as a task — R7 -was a numbered requirement in SPEC.md yet had no dedicated milestone task of its -own (it was folded into M4-T1's prose), so it fell through. **Requirements that -survive only inside another task's description are the ones that get dropped.** -Give each acceptance-criteria-bearing requirement its own checkbox. - -### Residual risk of deferring R7 (accepted, bounded) - -If a job settles and its digest entry is expired by `agentStep` TTL before the -agent acts, the agent can miss the completion. Four mitigations bound this: -settled jobs remain listable for the settled TTL; `end_turn` warns on running -process jobs; the live `job_update` rail surfaces settlement to the user -immediately; and a status/`exitCode`/`completedAt` change busts the digest -fingerprint, so the settlement digest IS emitted at least once. The gap is purely -TTL expiry before acknowledgement, never suppression. Three M4-T3 test cases -(force-after-compaction, unacked settlement, MEASURED token ceiling) remain -untestable until R7 lands. - -### Process note: plan artifacts drifted from the code - -At the time the user asked whether the plans were complete, `STATUS.md` claimed -M1–M4 done while `PLAN.md` still had every box unchecked and a `current-task` -pointer at M1 — and the session had no `STATE.json` at all. The code was well -ahead of both artifacts. Update the plan artifact at each milestone boundary, not -retroactively at closure; a stale artifact makes a finished session look -abandoned and forces a re-audit of the code to answer a simple status question. diff --git a/.agents/sessions/background-job-push-model-2026-08/PLAN.md b/.agents/sessions/background-job-push-model-2026-08/PLAN.md deleted file mode 100644 index 8728a18c25..0000000000 --- a/.agents/sessions/background-job-push-model-2026-08/PLAN.md +++ /dev/null @@ -1,133 +0,0 @@ -# Background Job Push Model & Drain Inversion — PLAN - - - -Status legend: `[ ]` pending, `[~]` in_progress, `[x]` done, `[/]` cancelled, `[!]` blocked. - -**Session closed 2026-08-02.** M1–M5 delivered with three recorded deviations -from the original DESIGN (see §Deviations). One requirement (R7 settlement -tombstone) is explicitly DEFERRED, not delivered — its residual risk is stated -below. Shipped in commits `e5797f4fe` (push model + list_jobs) and -`a7650cfdd` (reviewer nits); further list_jobs schema/doc nits are uncommitted -in the working tree at closure time. - -Ordering was strict: M2 depended on M1 landing; M3 depended on M2 (the poll loop -existed ONLY to drain, so it could not be removed until draining was inverted); -M4 depended on M3 (the digest's pending-lines bucket is only meaningful once -output enters the registry on write). - -## Milestones - -### M1 — Independent registry/bounding bugs (no design risk) — DONE - -- [x] M1-T1 Fix `truncatedAtCursor` in `common/src/util/job-registry.ts` to signal a real gap. (Now returns `first.sequence > cursor + 1` — a gap only when unread events were actually evicted; 3 regression tests in `job-registry.test.ts` cover healthy-buffer, evicted, and caught-up cursors.) -- [x] M1-T2 Bound `check_job` returned `events` by an explicit ceiling in `sdk/src/tools/check-job.ts`. (`CHECK_JOB_POLL_ACCUMULATION_CAP` = 2× `CHECK_JOB_OUTPUT_LIMIT` bounds both the `wait_for` match window and the returned payload; `truncated: true` when trimmed. Covered by the oversized-single-event and chatty-follow OOM tests.) -- [x] M1-T3 Validate: `common` + `sdk` typecheck; job-registry + check-job suites. (check-job 41/41; job-registry suites green; typechecks clean.) - -### M2 — Invert the drain (write-time, per-line events) — DONE WITH DEVIATION (D1) - -- [x] M2-T1 Live path drains independently of `check_job` and emits per-line `output` events. **Deviation D1:** implemented as a job-owned 250 ms interval drainer (`logQuotaTimer` → `readNewJobOutput` → `emitJobOutputLines`) plus a `hasLiveDrainer` single-drainer gate, NOT the specced piped stdio + in-process splitter. `stdio` remains `['ignore', outFd, outFd]`. -- [x] M2-T2 Keep `readNewJobOutput`/file-read strictly as the cross-session RECOVERY fallback. (Partially superseded by D1: `readNewJobOutput` is now ALSO the live drain source, but it is called from exactly one owner per job — the interval for live jobs, `check_job` for recovered jobs — so the shared `readOffset`/`decoder`/`lineCarry` cursor is never double-drained.) -- [x] M2-T3 Validate: `sdk` background-job suites. (`MAX_LINE_BYTES` force-flush, `hasLiveDrainer` pre-drain, `wait_for`-via-`lineCarry`, and settled-TTL prune/recovery tests added; blob-semantics assertions reworked.) - -### M3 — Replace poll loop with jobRegistry.wait() — DONE WITH DEVIATION (D2) - -- [x] M3-T1 Follow mode is event-driven via `jobRegistry.wait()`. **Deviation D2:** the `while (true)` loop in `check-job.ts` is RETAINED, but each iteration now awaits `jobRegistry.wait(registryJobId, { timeoutMs: Math.min(POLL_INTERVAL_MS, remaining) })` instead of `sleep(200)`. Matches wake immediately; the 200 ms cap only bounds how long a quiet iteration blocks. -- [x] M3-T2 Preserve external contract: poll vs follow, matched latch, `kill_on_timeout`, full-window events. (Verified: `timeout_seconds: 0`/absent = single non-blocking poll; follow without `wait_for` blocks to exit or deadline; `matched` only emitted when `wait_for` was supplied.) -- [x] M3-T3 Validate: check-job suite + agent-runtime check-job handler. (41/41 check-job; agent-runtime handler suites green.) - -### M4 — Pushed status digest + list_jobs pending signal — DONE WITH DEVIATION (D3) - -- [x] M4-T1 Change-gated background-job digest in the per-step observation. **Deviation D3:** base2 yields a plain programmatic `list_jobs` after the initial and mid-loop `git_status` (same rail, `agentStep` TTL), and the change-gate lives in the SDK dispatch (`applyListJobsDigestGate` in `sdk/src/run.ts`) rather than in base2. Per-turn row fingerprint (jobId|status|pending|gap|completedAt|exitCode, folded with `truncatedCount`) suppresses an unchanged repeat digest into `{ unchanged: true, note }`. **NOT delivered:** the settlement tombstone (R7), the coarse elapsed bucket, the `+N more (list_jobs)` line (`truncatedCount` is used instead), and the MEASURED token-ceiling test. -- [x] M4-T2 Bucketed pending-output + gap signal on `list_jobs`. (`common/src/util/list-jobs-view.ts` pure helpers; pending buckets relative to the `check_job` consumer cursor, `gap` from ring truncation, terminal tail ≤10, row cap 10 with `truncatedCount`, fixed no-action note. `list_jobs` never advances `lastCheckCursor`. Dual-id recovered jobs reverse-resolve to the user-facing jobId.) -- [x] M4-T3 Validate: regression tests for omit-when-unchanged, no-action-line contract, cursor immutability, gap, terminal tail, dual-id rediscovery, and the suppressed-variant schema. (34/34 across list-jobs-view, list-jobs, run-list-jobs-gate, list-jobs-params.) **NOT covered:** force-on-first-step-after-compaction, unacked-settlement forcing, and the MEASURED token ceiling — all downstream of the deferred R7. - -### M5 — Final validation & review — DONE - -- [x] M5-T1 Cross-package typecheck. (All 11 packages green via the `script:typecheck` hook.) -- [x] M5-T2 Live end-to-end background-job smoke. (Real `startBackgroundJob` → `check_job` follow matched `READY-TOKEN` → `list_jobs` digest with matching jobId and pending bucket → clean kill. **Not exercised:** "settlement surfaces exactly once" — unverifiable while R7 is deferred.) -- [x] M5-T3 Address automated reviewer gate blockers. (All reviewer gates reached LOOKS_GOOD or NON_BLOCKING; blocking findings on the dual-id remap, follow-without-`wait_for`, and fd leak were repaired and re-reviewed.) - -## Deviations from DESIGN (accepted at closure) - -**D1 — Interval drainer instead of piped stdio (M2-T1).** -Spec called for `spawn(..., { stdio: ['ignore', 'pipe', 'pipe'] })` with an -in-process line splitter teeing to both the log fd and `emitJobOutput`. Shipped -instead: the existing fd-based spawn plus a per-job 250 ms interval that calls -`readNewJobOutput` → `emitJobOutputLines`. - -- _Requirements met:_ R3 (per-line `output` events) and R4 (draining no longer - depends on `check_job`; an unpolled job still accrues events and settles). -- _Requirement partially met:_ R3's "write-time timestamps". Timestamps are - DRAIN-time with ≤250 ms granularity, not true write-time. AC3 ("lines written - seconds apart appear as distinct events with distinct timestamps") holds at - second scale; sub-250ms bursts collapse into one drain tick. -- _Why accepted:_ the piped-stdio edit is the single most safety-critical change - in the plan (detach, log-quota monitor, kill/exit settlement). The interval - drainer achieves the behavioral goal without touching detach semantics, and - the final drain + `flushJobLineCarry` on exit/error guarantees no trailing - line is lost. - -**D2 — `while (true)` retained around `jobRegistry.wait()` (M3-T1).** -Spec said delete the loop. Shipped: the loop remains, but the body awaits -`jobRegistry.wait()` with `timeoutMs: Math.min(POLL_INTERVAL_MS, remaining)`. - -- _Requirement met in substance:_ R5's real goal — no 200 ms quantization on a - match, no dead `wait()` on the shell path — is achieved. `wait()` is now the - wake mechanism. -- _Residual:_ the periodic re-entry still exists so the loop can re-drain a - recovered (non-live-drainer) job and re-evaluate `lineCarry`, which `wait()` - alone cannot observe. Removing the loop entirely would require moving carry - inspection into the registry. - -**D3 — Digest = plain `list_jobs` + SDK-side gate (M4-T1).** -Spec described a bespoke digest block assembled in base2 with per-job elapsed -buckets and a `+N more (list_jobs)` overflow line. Shipped: base2 yields the -real `list_jobs` tool, and suppression is a per-run fingerprint gate in -`sdk/src/run.ts`. - -- _Better than spec:_ one source of truth for the digest shape (the tool's own - schema/tests), and the gate applies to model-initiated `list_jobs` calls too. -- _Divergent details:_ `truncatedCount` replaces `+N more`; no elapsed bucket; - gate state is per-`run()` (per turn) rather than per-agent-context. - -## Deferred requirement - -**R7 — settlement tombstone: DEFERRED (not implemented).** -Requirement: "Settlement is surfaced at least once even if `agentStep` TTL -expires the digest before acknowledgement." No tombstone exists in the job code -(`grep tombstone` matches only `sdk/src/services/workspace-mutation-broker.ts`, -unrelated). - -_Residual risk:_ if a job settles and the digest entry carrying its terminal -state is expired by `agentStep` TTL before the agent acts on it, the agent can -miss the completion. Mitigations already in place that reduce (not eliminate) -the exposure: - -- Settled jobs stay listable for the whole settled TTL, so any later `list_jobs` - still reports the terminal state and `exitCode`. -- `end_turn` warns about still-running process jobs. -- The live `job_update` rail surfaces settlement to the USER immediately (M5 of - the unified-background-jobs session), so a human sees it even when the agent - does not. -- A status/`exitCode`/`completedAt` change busts the digest fingerprint, so the - settlement digest is emitted at least once — the gap is purely TTL expiry - before acknowledgement, not suppression. - -_To implement later:_ persist an unacknowledged-settlement flag per job in the -registry, force digest emission while any flag is set, and clear the flag when -the agent observes the terminal row. That also unlocks the three untested M4-T3 -cases (force-after-compaction, unacked settlement, MEASURED token ceiling). - -## DESIGN (historical — as originally specified) - -The original M1–M4 design specification is preserved verbatim in this file's -git history (see the pre-closure revision). It is intentionally not restated -here, because §Deviations above is now the authoritative record of what shipped -and how it differs. - -## Current state / resume - -Closed. Nothing to resume. The only carried-forward work item is R7 -(settlement tombstone) plus its three dependent M4-T3 test cases; open a new -session if that becomes a priority. diff --git a/.agents/sessions/background-job-push-model-2026-08/SPEC.md b/.agents/sessions/background-job-push-model-2026-08/SPEC.md deleted file mode 100644 index 0e556aa592..0000000000 --- a/.agents/sessions/background-job-push-model-2026-08/SPEC.md +++ /dev/null @@ -1,119 +0,0 @@ -# Background Job Push Model & Drain Inversion — SPEC - -## Problem - -The agent is the only background-job consumer still forced to poll. Hands-on -exercise of `run_terminal_command` (BACKGROUND) + `check_job` / `list_jobs` / -`read_logs` / `kill_job`, plus a background `spawn_agents` file-picker, surfaced -concrete defects rooted in one design flaw: **job output only enters the -registry as a side effect of the agent calling `check_job`.** - -Live push already exists for the human (`sdk/src/run.ts` → -`jobRegistry.subscribeAll` → `sdk/src/job-update-forwarder.ts` → -`cli/src/utils/sdk-event-handlers.ts:handleJobUpdate`), and the runtime already -pushes a per-step `git_status` observation into the agent's context. Background -jobs were never wired into that push rail for the agent. - -## Confirmed defects (evidence) - -1. `truncated: true, dropped: 0` on every `check_job`. `truncatedAtCursor` - (common/src/util/job-registry.ts) returns `first.sequence <= cursor`, true - for a HEALTHY buffer. It must compare against the highest EVICTED sequence. -2. Returned `events` array is unbounded. `CHECK_JOB_OUTPUT_LIMIT` (50_000) in - sdk/src/tools/check-job.ts only bounds the `wait_for` match window - (`appendBoundedCollected`), never the returned events. One poll of a chatty - job returned ~4,000 lines in a single event. -3. Output arrives as giant undifferentiated blobs. `readNewJobOutput` - (sdk/src/tools/background-jobs.ts) emits all bytes since last offset as ONE - `output` event, stamped at POLL time, not write time. Job-3's three lines - written 25s apart collapsed into one event. -4. Draining is a side effect of polling. `readNewJobOutput` is called from ONE - production site: the `while(true)` loop in check-job.ts. Unpolled jobs - produce no events; `read_logs` and `check_job` mutate shared cursor/offset - state. -5. `list_jobs` gives no progress signal (no cursor, no pending-output count). -6. `jobRegistry.wait()` (event-driven, no sleep-poll, self-cleaning) is DEAD - CODE on the shell path because checkJob must poll to drain. - -## Goals - -- Push bounded job STATE metadata into the agent's per-step observation; keep - job CONTENT pull-only. -- Invert draining so output enters the registry on write, per line, with - write-time timestamps — independent of polling. -- Replace checkJob's hand-rolled poll loop with `jobRegistry.wait()`. -- Fix the two independent registry/bounding bugs first (cheap, no design risk). - -## Non-goals - -- No change to the human-facing `job_update` → CLI live render pipeline - (already correct). -- No change to lifecycle state ownership (registry-owned; genuinely correct). -- No auto-kill of long-runners; end_turn still warns, never kills. -- No change to cross-session recovery's on-disk projection contract (it stays - the durable fallback where no in-process pipe can exist). - -## Requirements - -- R1: `truncated` is true IFF unread events were actually evicted for that - cursor; `dropped` and `truncated` never contradict. -- R2: `check_job` returned `events` are bounded by an explicit ceiling - regardless of how chatty the job is. -- R3: Live shell jobs emit per-line `output` events with write-time timestamps. -- R4: Draining no longer depends on `check_job` being called; an unpolled job - still accrues registry events and settles. -- R5: `checkJob` follow mode delegates to `jobRegistry.wait()` — no sleep loop. -- R6: A change-gated, bounded background-job digest is injected into the agent's - per-step observation block (same rail/TTL as `git_status`), declarative only, - with a fixed "no action required unless you need this output" contract line. -- R7: Settlement is surfaced at least once even if `agentStep` TTL expires the - digest before acknowledgement (settlement tombstone). -- R8: `list_jobs` exposes a per-job pending-output signal (bucketed) and a gap - flag. - -## Acceptance criteria - -- AC1: A live dev-server-style job, never polled, shows accruing registry - events and a terminal state after exit (R4). -- AC2: A chatty job's `check_job` response is bounded and reports - `truncated`/`dropped` consistently (R1, R2). -- AC3: Lines written seconds apart appear as distinct events with distinct - write-time timestamps (R3). -- AC4: `wait_for` resolves without 200ms quantization and with no sleep loop - (R5); existing check-job suite semantics preserved. -- AC5: The digest appears in the agent observation, is omitted when nothing - changed, force-emits on first step of a turn / after compaction / on - unacknowledged settlement, and carries the fixed no-action line asserted in a - test (R6, R7). -- AC6: Cross-package typecheck clean; job suites green; live end-to-end smoke. - -## Relevant systems / files - -- common/src/util/job-registry.ts — registry, `snapshot`, `wait`, - `truncatedAtCursor`, ring buffer, `subscribeAll`. -- sdk/src/tools/background-jobs.ts — `startBackgroundJob`, `readNewJobOutput`, - `emitJobOutput`, `settleBackgroundJob`, stdio wiring. -- sdk/src/tools/check-job.ts — poll loop to be replaced; `CHECK_JOB_OUTPUT_LIMIT`. -- sdk/src/tools/list-jobs.ts, read-logs.ts, kill-job.ts. -- sdk/src/run.ts (:541 subscribeAll; :1326 end_turn branch). -- sdk/src/job-update-forwarder.ts (human push; reference). -- agents/base2/base2.ts (git_status yield pattern → digest injection site). -- packages/agent-runtime/src/util/messages.ts (`expireMessages`, agentStep TTL). -- packages/agent-runtime/src/run-programmatic-step.ts - (`formatProgrammaticToolResultMessage`). -- Tests: common/src/util/**tests**/job-registry.test.ts, - sdk/src/**tests**/check-job.test.ts (asserts readNewJobOutput semantics — will - need rework), packages/agent-runtime/.../check-job.test.ts (asserts dropped:0). - -## Risks - -- Switching `startBackgroundJob` stdio from `['ignore', outFd, outFd]` to a pipe - touches the most safety-critical detach/quota/kill code. -- sdk/src/**tests**/check-job.test.ts asserts `readNewJobOutput` output directly - and will need rework, not just extension. -- The digest is a new context consumer landing on top of in-flight - `context-budget.ts` / `measure-context-baseline.ts` work; needs a MEASURED - token ceiling. -- PLAN.md of unified-background-jobs claims "disk is a projection, never - consulted for live state" — that holds for lifecycle, NOT output. This spec - formalizes the output source-of-truth question left unspecified there. diff --git a/.agents/sessions/background-job-push-model-2026-08/STATE.json b/.agents/sessions/background-job-push-model-2026-08/STATE.json deleted file mode 100644 index 86ee5be726..0000000000 --- a/.agents/sessions/background-job-push-model-2026-08/STATE.json +++ /dev/null @@ -1,10 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "background-job-push-model-2026-08", - "status": "completed", - "currentTask": null, - "revision": 1, - "checkpoint": null, - "createdAt": "2026-08-02T10:30:15.315Z", - "updatedAt": "2026-08-02T10:30:15.315Z" -} diff --git a/.agents/sessions/background-job-push-model-2026-08/STATUS.md b/.agents/sessions/background-job-push-model-2026-08/STATUS.md deleted file mode 100644 index 07396e3e96..0000000000 --- a/.agents/sessions/background-job-push-model-2026-08/STATUS.md +++ /dev/null @@ -1,53 +0,0 @@ -# Background Job Push Model & Drain Inversion — STATUS - -## Current state — M1–M4 implemented; awaiting gate on M4 files - -### Done - -- M1: truncatedAtCursor gap predicate; bound returned events; tests. -- M2: live 250ms drain, per-line events, hasLiveDrainer, peekJobLineCarry. -- M3: checkJob wait() wake; matched=false; follow without wait_for; OOM bounds tests. -- Adapter TTL prune + recovery settledAt / no post-TTL re-cache; lastCheckCursor for live observability. -- Reviewer nits: wait_for-via-lineCarry test; MAX_LINE_BYTES force-flush test (export MAX_LINE_BYTES). -- **M4:** - - Pure helpers `common/src/util/list-jobs-view.ts` + unit tests. - - `list_jobs` schema + SDK: pending bucket, gap, optional terminal tail, note line; does not advance lastCheckCursor. - - base2 yields programmatic `list_jobs` after initial + mid-loop `git_status` (pushed status digest). - -### Validation (local) - -- check-job 41/41; list-jobs-view + list-jobs 12/12; combined suites 53/53. -- typecheck: common, sdk, agents clean. - -### Pending - -- Runtime validation/reviewer gate on M4 dirty set. -- M5: cross-package smoke + live dev-server if desired. - -## Resume - -If gate returns blockers, fix finding-scoped issues only. M5 is optional next after LOOKS_GOOD. - - - -## Session closed — M1–M5 delivered, R7 deferred — 2026-08-02T10:29:36.462Z - -The session is closed with M1–M5 delivered and one requirement (R7) explicitly deferred. Rather than checking boxes silently, `PLAN.md` was rewritten to be honest about what shipped: every milestone line now carries the concrete evidence for its claim, and a new §Deviations section records three places where the implementation diverged from the original DESIGN. - -### What shipped - -- **M1 (registry/bounding bugs):** `truncatedAtCursor` now reports a gap only when unread events were actually evicted (`first.sequence > cursor + 1`) instead of firing on every healthy buffer; `CHECK_JOB_POLL_ACCUMULATION_CAP` bounds both the `wait_for` match window and the returned `events` payload. check-job 41/41. -- **M2 (drain inversion):** live jobs drain independently of `check_job` and emit per-line `output` events, with a `hasLiveDrainer` single-drainer gate so the shared `readOffset`/`decoder`/`lineCarry` cursor is never double-drained. Final drain + `flushJobLineCarry` on exit/error guarantees no trailing line is lost. -- **M3 (poll loop):** follow mode is event-driven through `jobRegistry.wait()`; a match wakes immediately with no 200 ms quantization. -- **M4 (digest + list_jobs):** base2 pushes a programmatic `list_jobs` on the same rail as `git_status`; the SDK gates unchanged repeat digests into `{ unchanged: true, note }`; `list_jobs` gained pending buckets relative to the `check_job` cursor, a `gap` flag, terminal tails, a 10-row cap with `truncatedCount`, and dual-id reverse-resolution so recovered jobs expose the user-facing jobId. 34/34 across the four related suites. -- **M5 (validation):** all 11 packages typecheck; live real-spawn smoke passed (start → follow matched `READY-TOKEN` → digest → clean kill); every reviewer gate reached LOOKS_GOOD or NON_BLOCKING after repairing the dual-id, follow-without-`wait_for`, and fd-leak blockers. - -### Three recorded deviations - -- **D1** — a 250 ms interval drainer replaced the specced piped stdio + line splitter. Achieves R3/R4, but timestamps are drain-time at ≤250 ms granularity rather than true write-time. Accepted because piped stdio is the single most safety-critical edit in the plan (detach, quota monitor, kill/exit settlement). -- **D2** — the `while (true)` loop in `check-job.ts` was retained, with `jobRegistry.wait()` as the wake mechanism instead of `sleep(200)`. R5's real goal is met; the periodic re-entry survives so the loop can still re-drain recovered jobs and re-inspect `lineCarry`, which `wait()` alone cannot observe. -- **D3** — the digest is the real `list_jobs` tool plus an SDK-side fingerprint gate, not a bespoke base2 block. Better in that the digest has one source of truth and the gate also covers model-initiated calls; divergent in that `truncatedCount` replaces `+N more`, there is no elapsed bucket, and gate state is per-run rather than per-agent-context. - -### Deferred: R7 settlement tombstone - -Not implemented — `grep tombstone` matches only the unrelated workspace-mutation broker. If a job settles and its digest entry is expired by `agentStep` TTL before the agent acts, the agent can miss the completion. Four mitigations bound the exposure: settled jobs stay listable for the settled TTL, `end_turn` warns on running process jobs, the live `job_update` rail shows the user immediately, and a status/`exitCode`/`completedAt` change busts the fingerprint so the settlement digest is emitted at least once. The gap is purely TTL expiry before acknowledgement, not suppression. This also leaves three M4-T3 cases untested (force-after-compaction, unacked settlement, MEASURED token ceiling). diff --git a/.agents/sessions/context-baseline-25k/EVENTS.jsonl b/.agents/sessions/context-baseline-25k/EVENTS.jsonl deleted file mode 100644 index 2020ea0b6f..0000000000 --- a/.agents/sessions/context-baseline-25k/EVENTS.jsonl +++ /dev/null @@ -1,11 +0,0 @@ -{"ts":"2026-08-04T21:12:43.937Z","kind":"append_lesson","summary":"Appended entry \"M1-T3/T4/T5 verified landed — 2026-08-04\" to STATUS.md","payload":{"heading":"M1-T3/T4/T5 verified landed — 2026-08-04","artifact":"STATUS.md"}} -{"ts":"2026-08-04T23:20:04.489Z","kind":"append_lesson","summary":"Appended entry \"Smoke set + M4 baselines — 2026-08-04\" to STATUS.md","payload":{"heading":"Smoke set + M4 baselines — 2026-08-04","artifact":"STATUS.md"}} -{"ts":"2026-08-05T00:38:00.947Z","kind":"append_lesson","summary":"Appended entry \"NF-1/NF-2 resolved — 2026-08-04\" to STATUS.md","payload":{"heading":"NF-1/NF-2 resolved — 2026-08-04","artifact":"STATUS.md"}} -{"ts":"2026-08-05T02:59:40.287Z","kind":"append_lesson","summary":"Appended entry \"Deliverables committed — 2026-08-05\" to STATUS.md","payload":{"heading":"Deliverables committed — 2026-08-05","artifact":"STATUS.md"}} -{"ts":"2026-08-05T06:02:55.271Z","kind":"append_lesson","summary":"Appended entry \"AC-A1 pre-flip smoke baselines — 2026-08-05\" to STATUS.md","payload":{"heading":"AC-A1 pre-flip smoke baselines — 2026-08-05","artifact":"STATUS.md"}} -{"ts":"2026-08-05T10:11:05.320Z","kind":"append_lesson","summary":"Appended entry \"M3 tree/knowledge reductions — 2026-08-05\" to STATUS.md","payload":{"heading":"M3 tree/knowledge reductions — 2026-08-05","artifact":"STATUS.md"}} -{"ts":"2026-08-05T17:37:42.620Z","kind":"append_lesson","summary":"Appended entry \"Resume 2026-08-05 — continue open plan\" to STATUS.md","payload":{"heading":"Resume 2026-08-05 — continue open plan","artifact":"STATUS.md"}} -{"ts":"2026-08-05T17:44:16.775Z","kind":"append_lesson","summary":"Appended entry \"M5-T1 ranking — 2026-08-05\" to STATUS.md","payload":{"heading":"M5-T1 ranking — 2026-08-05","artifact":"STATUS.md"}} -{"ts":"2026-08-05T17:46:34.461Z","kind":"append_lesson","summary":"Appended entry \"M5 schema diet complete — 2026-08-05\" to STATUS.md","payload":{"heading":"M5 schema diet complete — 2026-08-05","artifact":"STATUS.md"}} -{"ts":"2026-08-05T17:59:42.119Z","kind":"append_lesson","summary":"Appended entry \"M5 security review cleared\" to STATUS.md","payload":{"heading":"M5 security review cleared","artifact":"STATUS.md"}} -{"ts":"2026-08-05T17:59:42.119Z","kind":"session_status","summary":"Session status -> validating","payload":{"status":"validating"}} diff --git a/.agents/sessions/context-baseline-25k/LESSONS.md b/.agents/sessions/context-baseline-25k/LESSONS.md deleted file mode 100644 index cb5f6e1fbb..0000000000 --- a/.agents/sessions/context-baseline-25k/LESSONS.md +++ /dev/null @@ -1,17 +0,0 @@ -# Context Baseline 25–30k — LESSONS - -Session: context-baseline-25k -Running notes on gotchas, decisions, and reusable patterns. Append-only. - -## general-agent dual-site mirror trap — 2026-08-04 - -`agents/general-agent/general-agent.ts` has its own inline `shouldProactivelyQueryIndex` (length check + generic code-intent regex only). Mirroring base2's post-M4 strong-intent gate onto it is NOT behavior-neutral: `agents/__tests__/general-agent.test.ts` audit-loop tests drive prompt `'Audit service completeness'`, and once the classifier recognizes audit verbs the first yield flips from `spawn_agent_inline` (context-pruner) to `query_index`, breaking 3 tests (rejects audit completion without a structural receipt; breaks the audit loop once a matching structural receipt is present; breaks the audit loop after exhausting completion retries). Any M4-T2-style mirror must bundle those test updates (accept a query_index first yield) and ship as its own gate-scoped change — never fold it into an advisory-repair turn. The out-of-scope mirror was reverted via `git restore agents/general-agent/general-agent.ts`; suite back to 7/7. - -## Editor scope discipline — 2026-08-04 - -When delegating "advisory repair" fixes to an editor, name the finding IDs AND explicitly forbid expanding scope items (dual-site mirrors, adjacent cleanups). The NF-1/NF-2 repair silently included a descoped dual-site change because the prompt mentioned M4-T2 in discovery context rather than as a non-goal. - -## Gate timing notes — 2026-08-04 - -- `bun run typecheck` at repo root takes >30s (11 packages); run per-package or accept the timeout and let `run_file_change_hooks` do it. -- `git restore ` is safe for out-of-scope editor drift in this repo's hook configuration (typecheck-only hooks pass regardless; the gate scope follows dirty files). diff --git a/.agents/sessions/context-baseline-25k/PLAN.md b/.agents/sessions/context-baseline-25k/PLAN.md deleted file mode 100644 index cb75481f60..0000000000 --- a/.agents/sessions/context-baseline-25k/PLAN.md +++ /dev/null @@ -1,156 +0,0 @@ -# Context Baseline 25–30k — PLAN - -Session: context-baseline-25k -Spec: ./SPEC.md -Status: active - - - -Sequencing: **measure first (M0)**, then **tool tiers (M1)** as largest fixed win, parallel **prompt default-on (M2)** + **tree/knowledge (M3)**, then **proactive lean (M4)**, optional **schema diet (M5)**, **history polish (M6)**. Gate e2e after every default flip. - -Predecessor complete (do not re-implement): ledger, `/context`, prompt disclosure _flag_, proactive cache+epoch, git gate, semantic compaction, tool-result lifecycle — see `../context-budget-architecture-2026-08/STATUS.md`. - ---- - -## Milestone 0 — Measurement lock (no behavior change) - -Goal: production-faithful numbers; stop optimizing the wrong baseline. - -- [x] M0-T1 Extend `scripts/measure-context-baseline.ts`: - - Measure `FILE_TREE_PROMPT_SMALL` (2.5k) and full 10k separately - - Measure progressive prompt disclosure off vs on (authored surface via `createBase2`) - - Emit single **default base2 assembled fixed** line (production placeholders) - - Optionally stub tool-tier totals once tiers exist (core vs full) -- [x] M0-T2 Record corrected numbers in STATUS.md (fixed vs first-turn with proactive) -- [x] M0-T3 Soft assertion helpers or documented phase targets (≤32k / ≤28k / ≤30k) -- Validation: `bun run scripts/measure-context-baseline.ts` exit 0; STATUS updated — **done 2026-08-04** (fixed 49,345 this worktree; ~39.4k if git clean) - ---- - -## Milestone 1 — Progressive tool disclosure (R1, AC-F2, AC-G\*) - -Goal: cut tool definition tax from ~23.5k toward core ≤12k. - -- [x] M1-T1 Create `agents/base2/tool-tiers.ts` with CORE / IMPLEMENT / AUDIT / MEDIA_3D / JOB_EXTRA and `resolveModelToolNames` -- [x] M1-T2 Wire `createBase2` `toolNames` from tiers + existing mode gates (`planOnly`, `fast`, `executePlan`, `noAskUser`); keep `programmaticToolNames` unchanged -- [ ] M1-T3 Agent state: `unlockedToolTiers`; deterministic unlock in `handleSteps` on implement intent / active work / audit classifier / media paths - - Acceptance: `publishUnlockedToolTiers` runs before each STEP inside serialized handleSteps; inline copy guarded by a sync test against `resolveUnlockedTiersForPhase(deriveIntentSignals(...))` - - Validate: bun test agents/**tests**/base2-progressive-tool-disclosure.test.ts -- [ ] M1-T4 Runtime: only unlocked tools in `getToolSet` / token ledger; locked call → actionable message (tool-executor or spawn path) - - Acceptance: loopAgentSteps test shows per-step expand AND shrink; `tool-executor` rejects still-locked tier tools with the listed current surface - - Validate: bun test agents/**tests**/base2-progressive-tool-disclosure.test.ts -- [ ] M1-T5 Short always-on “Tool surface” prompt block - - Acceptance: `## Tool surface` block present in the base2 system prompt exactly when progressiveToolDisclosure is on - - Validate: bun test agents/**tests**/base2-progressive-tool-disclosure.test.ts -- [x] M1-T6 Flag: `progressiveToolDisclosure` option + `OPENBUFF_PROGRESSIVE_TOOL_DISCLOSURE` canary (default off until canary green) -- [x] M1-T7 Tests (partial — unit canary + core <12k + planOnly gates; gate e2e unlock + locked-tool path now landed and green): - - [x] Core-only token total ≤ ~12k; full ≤ 25k - - [x] Plan mode never exposes `edit_transaction` - - [x] Gate e2e with IMPLEMENT unlocked after first edit - - [x] Locked tool error path -- Validation: agents + agent-runtime typecheck; unit + gate e2e subset -- Sequencing note: M0 recommended for before/after numbers -- Risk: prompt-cache miss on unlock — accept per phase; keep CORE stable - -### M1 rollout - -| Phase | Behavior | -| ----------- | ---------------------------------------------------------------------------------------------------------------------- | -| Canary | env on: core + auto-unlock IMPLEMENT on implement intent (does **not** satisfy AC-A1) | -| Default-on | only after **AC-A1**: full gate e2e **and** buffbench subset or fixed smoke set; results no worse than STATUS baseline | -| Kill switch | explicit false → full 33-tool surface | - ---- - -## Milestone 2 — Default-on progressive prompt disclosure (R2, AC-P1, AC-G2) - -Goal: thinner always-on authored surface; gate text remains full. - -- [ ] M2-T1 Confirm guides exist and tests pass with flag on (`agents/__tests__/base2-progressive-disclosure.test.ts`) -- [ ] M2-T2 Default `progressivePromptDisclosure` **true** (explicit false still wins; update env docs: default on) -- [ ] M2-T3 Ensure `gateAwarenessSection` is **never** passed through `disclose()` — always full text -- [ ] M2-T4 Optional: shrink `knowledgeFilesPrompt` to short blurb + `agents/guides/knowledge-files.md` -- [ ] M2-T5 Optional second wave: compress long spawning guidelines / multi examples to index + one example -- [ ] M2-T6 Mode-thin instructions where cheap (conversation / plan vs implement) without dropping gate in implement modes -- [ ] M2-T7 Gate e2e + progressive-disclosure suite green; update `docs/configuration.md` / `docs/environment-variables.md` -- [ ] M2-T8 Before default-on: satisfy **AC-A1** (gate e2e + buffbench subset or fixed smoke tasks; no worse than STATUS baseline). Record evidence in STATUS. -- Validation: agents tests; gate e2e; **AC-A1** before default true -- Sequencing note: independent of M1; canary dogfood recommended before default flip -- Risk: hidden guidance — keep trigger pointers; canary first; **AC-A1** hard bar for default-on - ---- - -## Milestone 3 — Cheaper always-on project context (R3) - -Goal: lower SMALL tree + knowledge instruction cost. - -- [ ] M3-T1 Lower `FILE_TREE_PROMPT_SMALL` budget in `packages/agent-runtime/src/templates/strings.ts` from 2_500 → **1_500–2_000** -- [ ] M3-T2 Bias truncation toward path-only / drop symbols earlier for agent-mode small tree (`truncate-file-tree.ts` if needed) -- [ ] M3-T3 Shrink static `knowledgeFilesPrompt` (or pair with M2-T4 guide) -- [ ] M3-T4 Update baseline script expectations; document that LARGE remains for search agents -- Validation: measure script; prompts-ledger tests if any; no gate impact expected -- Sequencing note: uses M0 for proof of savings -- Risk: low — discovery tools remain - ---- - -## Milestone 4 — Lean proactive inject (R4, AC-R1) - -Goal: fewer firings + smaller first-hit payload. - -- [x] M4-T1 Tighten `classifyProactiveRetrieval` in `agents/base2/base2.ts` (stronger intent; strip bare generic solo triggers; skip pure Q&A) -- [x] M4-T2 Mirror intent policy in `agents/general-agent/general-agent.ts` if still dual-sited — N/A: no dual-site exists (verified absence of `classifyProactiveRetrieval` there) -- [x] M4-T3 Lower default limits (unknown 8, multi-file 12, cross-subsystem 16) -- [x] M4-T4 Compact proactive result envelope (<1.5k); full envelope on explicit `query_index` only — via inline `toCompactProactiveRetrievalResult` -- [x] M4-T5 Defer or summarize cross-subsystem structure/list_directory extras -- [x] M4-T6 Tests: no-fire weak prompts; compact size; cache hit still pointer; epoch invalidation unchanged -- Validation: base2 proactive tests; agents typecheck — gate reviewer verdict NON_BLOCKING 2026-08-04, coverage covered; two advisory nits (NF-1 EXPLORE_PROMPT mentions dropped matchedSnippets/relatedFiles; NF-2 compact cache result read-once) recorded in STATUS.md -- Sequencing note: uses M0 for size assertions -- Risk: audit breadth — keep full path on explicit tool / confirmed audit - ---- - -## Milestone 5 — Schema / description diet (R5, optional) - -Goal: extra 1–3k on CORE without removing tools. - -- [ ] M5-T1 Rank tools by token cost (script or unit helper) -- [ ] M5-T2 Shorten top CORE descriptions / redundant field describes -- [ ] M5-T3 Re-check core ≤12k and schema compile/generate guards -- Validation: `base2-context-budget`; tool definition generate if required -- Sequencing note: prefer after M1 (diet CORE after tiers exist) -- Risk: model misuse from terse schemas — keep critical constraints in schema - ---- - -## Milestone 6 — History hygiene polish (R6) - -Goal: remaining window lasts longer (not fixed baseline). - -- [ ] M6-T1 Ensure proactive injects are lifecycle-normal (not pinned/high) -- [ ] M6-T2 Verify spawn/read digest bounds; document any gaps -- [ ] M6-T3 Optional prompt nudge: prefer post-edit receipts over full re-reads -- Validation: tool-result-lifecycle + messages trim tests -- Sequencing note: no critical dependencies - ---- - -## Cross-cutting - -- [ ] X-T1 Flags documented (`OPENBUFF_PROGRESSIVE_TOOL_DISCLOSURE`, disclosure default change) -- [ ] X-T2 Architecture/config docs updated for tool tiers + 25–30k program -- [ ] X-T3 Full validation before default-on flips: typecheck agents/agent-runtime/cli/common; gate e2e; context-pruning; progressive-disclosure; baseline script -- [ ] X-T4 Do **not** remove automated reviewer/hooks to save tokens -- [ ] X-T5 Document the fixed smoke task set (or buffbench subset commands) in STATUS **before** the first M1/M2 default-on attempt; record pre-flip baseline scores for **AC-A1** - ---- - -## Suggested calendar (indicative) - -| Window | Work | -| -------- | --------------------------------------------------- | -| Week 1 | M0 measurement lock | -| Week 1–2 | M1 tool tiers (canary) | -| Week 2 | M2 prompt default-on + M3 tree/knowledge (parallel) | -| Week 3 | M4 proactive + M5 schema diet | -| Week 4 | M1/M2 default-on only after | diff --git a/.agents/sessions/context-baseline-25k/SPEC.md b/.agents/sessions/context-baseline-25k/SPEC.md deleted file mode 100644 index 85706518ae..0000000000 --- a/.agents/sessions/context-baseline-25k/SPEC.md +++ /dev/null @@ -1,262 +0,0 @@ -# Context Baseline 25–30k — SPEC - -Status: ready -Session: context-baseline-25k -Owner: orchestrator (Buffy) -Created: 2026-08-04 -Predecessor: `.agents/sessions/context-budget-architecture-2026-08/` (M0–M6 largely complete) - -## Problem statement - -The model context window fills quickly. Cost lives in two pools; the prior architecture session instrumented and partially reduced both, but **fixed per-turn overhead remains ~40–50k** on this repo before useful conversation work. - -### Measured baseline (openbuff repo, 2026-08-04) - -From `bun run scripts/measure-context-baseline.ts` (gpt-tokenizer + 1.35× Anthropic fudge): - -| Component | Tokens | Notes | -| ---------------------------------------- | ---------: | ------------------------------------------------------------------------------------------ | -| Tool definitions (33 base2 tools) | **23,506** | Dominant fixed cost; regression cap 25k in `agents/__tests__/base2-context-budget.test.ts` | -| base2 systemPrompt (raw template) | **11,612** | Placeholders not expanded | -| File tree @ 10k budget | **9,126** | **Overstates default base2** — production uses `FILE_TREE_PROMPT_SMALL` (2.5k) | -| Knowledge files instruction (static) | 1,036 | Always-on “how to write knowledge” | -| System info / git / patterns / language | ~1.5k | Small | -| Proactive `query_index` (representative) | ~4.9k | Variable; often every coding turn | -| git_status | ~52 | Cheap; already gated by SDK | - -**Script total (incl. 10k tree + injections):** ~51.9k (~27% of 190k). -**Realistic default-mode fixed cost (SMALL tree, no proactive):** ~**40–45k**. - -### What already shipped (predecessor session) - -- Context budget ledger + `/context` (alias `/ctx`) -- Progressive **prompt** disclosure (`progressivePromptDisclosure`, default **off**; env canary `OPENBUFF_PROGRESSIVE_PROMPT_DISCLOSURE`) -- Proactive retrieval cache + workspace revision + index mutation epoch invalidation -- Git observation gating (SDK `applyGitStatusGate`) -- Model-aware semantic compaction before mechanical 190k brake -- Tool-result lifecycle tagging / keep-N policy - -### What is still open (this program) - -1. **Tool schemas never tiered** — full ~23.5k every request -2. **Progressive prompt disclosure not default-on** -3. **Measurement overstates tree** — script measures 10k path; production SMALL unmeasured in “default fixed” line -4. **Proactive envelope still fat** on first miss (~5k + cross-subsystem extras) -5. **Knowledge essay** always-on (~1k) -6. Gate-critical text mixed with relocatable advisory bulk (M4 flag off keeps everything inline) - -## Goals - -- **G1.** Cut default-mode **fixed** baseline (system + tools + small tree + knowledge + profiles; **exclude** proactive) from ~40–45k to **≤30k**, stretch **≤25k**. -- **G2.** Progressive **tool** disclosure: core tools always registered; implement/audit/media tiers unlock deterministically without weakening the gate. -- **G3.** Default-on progressive **prompt** disclosure with gate text remaining fully inline. -- **G4.** Cheaper always-on project context (path-biased small tree + short knowledge blurb). -- **G5.** Lean proactive inject (tighter classifier + compact envelope); full envelope only on explicit `query_index`. -- **G6.** Optional schema/description diet on CORE tools. -- **G7.** Production-faithful measurement so claims are provable via the baseline script and `/context`. -- **G8.** No silent ability regression on default-on flips: gate e2e plus buffbench subset or fixed smoke tasks must be no worse than the pre-flip baseline recorded in STATUS. - -## Non-goals - -- Not changing provider/model routing or BYOK architecture. -- Not rewriting context-pruner summarization heuristics beyond wiring already done. -- Not altering the deterministic edit / read-capability system. -- **Not weakening the reviewer/validation gate contract** (hooks → automated reviewer; basher / `run_targeted_validation` remain optional evidence only). -- Not removing programmatic tools required by `handleSteps` (`spawn_agent_inline`, `git_status`, `run_file_change_hooks`, `inspect_codebase_structure`). -- No new third-party dependencies. -- Not counting proactive inject as part of the fixed 25–30k target (report fixed vs first-turn separately). -- Not hard-gating requests that exceed per-component budgets (ledger remains advisory unless a later program decides otherwise). - -## Gate invariants (non-negotiable) - -Any optimization MUST preserve: - -1. **Runtime owns the gate** — on turn end, hooks then automated code-reviewer; model does not “run” the gate. -2. **Pinned authority** — `GATE: PENDING | PASSED`, `pendingGateFiles`, fail-closed repair loops. -3. **What is not the gate** — basher typecheck/test/lint and `run_targeted_validation` are optional evidence only; they do not unlock `git-committer`. -4. **Withholds while PENDING** — `suggest_followups` rejected; `git-committer` withheld; no manual code-reviewer re-spawn for the same pending set. -5. **Re-arm on edit** — any new mutation returns to PENDING. -6. **Programmatic tools stay available** to `handleSteps` even if model-facing `toolNames` shrink. - -**Always-on inline (never guide-only):** - -- Full `gateAwarenessSection` (`agents/base2/quality-prompt-section.ts`) -- Minimal harness recovery: end turn when GATE PENDING; do not treat typecheck as gate - -**OK to relocate / lazy-load:** craftsmanship, git discipline detail, security pre-edit procedure, specialist catalog, broad-audit procedure, fat schemas for rare tools, full file tree, full knowledge essay. - -## Requirements - -### R0 — Measurement lock - -- Extend `scripts/measure-context-baseline.ts` to report: - - `FILE_TREE_PROMPT_SMALL` (2.5k) and full (10k) separately - - progressive prompt disclosure off vs on (authored surface) - - tool tiers once defined (core vs full) - - a single **“default base2 assembled fixed”** line matching production placeholders -- Soft CI targets: phase-1 fixed ≤ 32k; phase-2 ≤ 28k; program target ≤ 30k / stretch 25k -- `/context` categories remain accurate (`tools`, `fileTree`, `system`, etc.) - -### R1 — Progressive tool disclosure - -- Define tiers in a single source of truth (e.g. `agents/base2/tool-tiers.ts`): - - **CORE** — discovery/orchestration always model-visible - - **IMPLEMENT** — edit/plan/validation helpers when implementation starts - - **AUDIT** — structure/completeness/coverage tools for broad audits - - **MEDIA_3D** — image/3d tools when media paths appear - - **JOB_EXTRA** — e.g. `kill_job` as needed -- Wire `createBase2` `toolNames` from tiers + mode (`planOnly` / `fast` / `executePlan` gates preserved). -- Only unlocked tools enter provider ToolSet / token accounting (`getToolSet` / `run-agent-step`). -- Locked tool call → actionable error (or one-shot deterministic unlock) — no silent no-op. -- Prefer **deterministic unlock from `handleSteps`** (classifier + phase) over free-form model `enable_tools`. -- `programmaticToolNames` unchanged. -- Short always-on prompt block describing tool surface / unlock rules. -- Flags: canary `OPENBUFF_PROGRESSIVE_TOOL_DISCLOSURE`; explicit option on `createBase2`; safe default until proven. - -### R2 — Default-on progressive prompt disclosure - -- Ship M4 behavior default **on** (or env default on) after canary. -- Relocate: quality, git discipline, security review, specialist routing, broad-audit (existing guides under `agents/guides/*`). -- Optionally second wave: compress long spawning/examples; shrink `knowledgeFilesPrompt` to ~150 tok + guide. -- Mode-thin instructions: conversation-only / plan / default / fast as specified in PLAN. -- Keep ≥25% authored-surface reduction test (`agents/__tests__/base2-progressive-disclosure.test.ts`). - -### R3 — Cheaper project context - -- Default orchestrator tree budget **1,500–2,000** (from 2,500 SMALL); path-oriented / symbol-stripped earlier. -- On-demand full tree via tools or LARGE placeholder for search agents only. -- Shrink static knowledge instruction; keep root knowledge **contents**. - -### R4 — Lean proactive inject - -- Tighten `classifyProactiveRetrieval`: stronger intent; no fire on bare generic words alone; Q&A without edit intent skips proactive. -- Lower limits: unknown ~8, multi-file ~12, cross-subsystem ~16 (from 14/24/30). -- Compact proactive envelope (paths, scores, top symbols, snapshotId; omit fat `status.coverage` / long explanations). Target **<1.5k** tokens when fired. -- Defer or summarize cross-subsystem `inspect_codebase_structure` + root `list_directory` extras. -- Do not pin proactive results; lifecycle normal importance. - -### R5 — Schema / description diet (optional stack) - -- Rank tools by tokens; cap verbose descriptions; leaner Zod→JSON Schema where safe. -- CORE total still ≤ ~12k after diet. - -### R6 — History hygiene (non-fixed) - -- Tag proactive for aggressive simplify; keep spawn/read bounds; do not put fixed baseline inside history budget (already true). - -### R7 — Backward compatibility & docs - -- Feature flags with safe rollout (canary → default-on → kill switch). -- Document in `docs/configuration.md`, `docs/environment-variables.md`, `docs/architecture.md` context-budget section. -- Existing gate e2e, context-pruning, and progressive-disclosure tests must keep passing (updated only where behavior intentionally changes). - -### R8 — Ability regression bar (hard for default-on) - -- Before flipping **M1** (progressive tool disclosure) or **M2** (progressive prompt disclosure) to **default-on**, record evidence in STATUS that satisfies **AC-A1**. -- Gate e2e alone is **necessary but not sufficient** for ability: also run a buffbench subset **or** a fixed smoke task set documented in STATUS. -- Results must be **no worse** than the pre-flip baseline (task success / gate pass rates) recorded in STATUS. -- Canary-only shipping (env flag on for dogfood) does **not** satisfy AC-A1 for a production default flip. - -## Acceptance criteria - -| ID | Criterion | -| ----- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| AC-F1 | Default-mode **fixed** baseline ≤ **30k**; stretch ≤ **25k** (no proactive) | -| AC-F2 | Core-only tools ≤ **12k**; full tools ≤ **25k** | -| AC-G1 | Gate e2e suites pass (`gate-lifecycle`, `gate-aux-ordering`, reviewer spawn conditions) | -| AC-G2 | `gateAwarenessSection` still fully inline when progressive prompt is on | -| AC-G3 | Programmatic hooks/git/spawn_inline work with locked model tools | -| AC-P1 | Progressive prompt default-on ≥ **25%** authored reduction | -| AC-R1 | Weak-intent prompts do not proactive-inject; compact proactive < **1.5k** when fired | -| AC-T1 | Baseline script reports production-faithful default fixed line | -| AC-C1 | `/context` reflects tool tiers and tree budget after changes | -| AC-A1 | Before flipping M1 or M2 to **default-on**, run (a) full gate e2e suite **and** (b) either a buffbench subset or a fixed smoke task set documented in STATUS; results must be **no worse than the pre-flip baseline** recorded in STATUS (task success / gate pass rates). Canary-only shipping does not satisfy AC-A1. | - -## Relevant systems (exact files) - -- `agents/base2/base2.ts` — `toolNames`, `programmaticToolNames`, progressive disclosure, `classifyProactiveRetrieval`, proactive yield/cache -- `agents/base2/quality-prompt-section.ts` — gate + relocatable sections -- `agents/guides/*` — on-demand prompt guides -- `agents/__tests__/base2-progressive-disclosure.test.ts`, `base2-context-budget.test.ts` -- `agents/e2e/gate-lifecycle.e2e.test.ts`, `gate-aux-ordering.e2e.test.ts`, reviewer-spawn e2e -- `packages/agent-runtime/src/run-agent-step.ts` — ledger, tools, system prompt assembly -- `packages/agent-runtime/src/tools/prompts.ts`, `tool-executor.ts` — ToolSet / locked-tool behavior -- `packages/agent-runtime/src/templates/strings.ts` — `FILE_TREE_PROMPT_SMALL` (2.5k), LARGE, placeholders -- `packages/agent-runtime/src/system-prompt/prompts.ts`, `truncate-file-tree.ts` -- `packages/agent-runtime/src/util/context-budget.ts`, `context-pruning.ts`, `tool-result-lifecycle.ts`, `token-counter.ts` -- `scripts/measure-context-baseline.ts` -- `cli/src/commands/context.ts` -- `common/src/tools/list.ts` — tool schemas -- `docs/configuration.md`, `docs/environment-variables.md`, `docs/architecture.md` -- Predecessor: `.agents/sessions/context-budget-architecture-2026-08/{SPEC,PLAN,STATUS}.md` - -## Key interfaces (pseudo-code) - -```ts -// agents/base2/tool-tiers.ts (NEW) -export const CORE_TOOLS = [ - /* spawn_agents, query_index, read_*, list_directory, glob, ask_user, skill, jobs minimal, ... */ -] as const -export const IMPLEMENT_TOOLS = [ - /* edit_transaction, create_plan, update_plan_status, inspect_*, get_affected_tests, run_targeted_validation, ... */ -] as const -export const AUDIT_TOOLS = [ - /* inspect_codebase_structure, inspect_feature_completeness, evaluate_audit_coverage, get_change_review_bundle, get_task */ -] as const -export const MEDIA_3D_TOOLS = [ - /* read_image, inspect_3d_asset, render_3d_preview, edit_3d_asset */ -] as const - -export type ToolTier = 'core' | 'implement' | 'audit' | 'media_3d' | 'job_extra' -export function resolveModelToolNames(params: { - mode: 'default' | 'fast' - planOnly?: boolean - executePlan?: boolean - unlockedTiers: ToolTier[] -}): string[] - -// agent state (additive) -// unlockedToolTiers?: ToolTier[] -// progressiveToolDisclosure?: boolean - -// proactive compact envelope (proactive path only) -// { kind: 'query_index_result_compact', results: [{ path, score, topSymbols?, reason? }], totalIndexed, indexAge, snapshotId } -``` - -## Token budget math (how to hit 25–30k) - -| Workstream | Est. save | Mechanism | -| ------------------------------------- | --------------: | ---------------------------------------- | -| Progressive tool disclosure | **12–16k** | Core ~12–15 tools always; rest on demand | -| Default progressive prompt disclosure | **3–5k** | M4 flag on; gate text inline | -| Cheaper tree + knowledge | **1–3k** | 1.5–2k tree; short knowledge blurb | -| Lean proactive | **2–4k/firing** | Classifier + compact envelope (variable) | -| Schema diet | **1–3k** | Shorter CORE schemas | - -Stack: ~42k realistic fixed − 14k tools − 4k prompt − 2k tree/knowledge ≈ **22–26k** fixed. - -## Risks and mitigations - -| Risk | Mitigation | -| -------------------------------------- | ---------------------------------------------------------------------------- | -| Model cannot find locked tool | Deterministic phase unlock + clear error | -| Hidden craftsmanship/git rules | Trigger pointers; canary; **AC-A1** buffbench/smoke before default-on | -| Prompt-cache thrash on tool unlock | Unlock once per phase; stabilize CORE all session | -| Audit quality drop from lean proactive | Full envelope on explicit tool; structure when audit confirmed | -| Gate weakened by thinner prompts | Never relocate `gateAwarenessSection`; gate e2e is merge bar | -| Ability regression on default-on flips | **AC-A1** hard bar: gate e2e + buffbench/smoke no worse than STATUS baseline | -| Measuring wrong tree budget | R0 production-faithful script | -| e2e asserting full tool list | Update only intentional assertions; keep gate semantics | - -## Out of scope (future) - -- Hard per-component request refusal over budget -- Embedding-based retrieval dedup -- Cross-session budget persistence -- Removing the automated reviewer to save tokens - -## Related reading - -- Predecessor session STATUS (completed architecture work) -- `docs/configuration.md` — Context budget and proactive retrieval -- `docs/architecture.md` — Context budget, retrieval caching, git observation gating diff --git a/.agents/sessions/context-baseline-25k/STATE.json b/.agents/sessions/context-baseline-25k/STATE.json deleted file mode 100644 index 52ae327ff5..0000000000 --- a/.agents/sessions/context-baseline-25k/STATE.json +++ /dev/null @@ -1,10 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "context-baseline-25k", - "status": "validating", - "currentTask": null, - "revision": 1, - "checkpoint": null, - "createdAt": "2026-08-05T17:59:42.119Z", - "updatedAt": "2026-08-05T17:59:42.119Z" -} diff --git a/.agents/sessions/context-baseline-25k/STATUS.md b/.agents/sessions/context-baseline-25k/STATUS.md deleted file mode 100644 index aaaf7bbf01..0000000000 --- a/.agents/sessions/context-baseline-25k/STATUS.md +++ /dev/null @@ -1,334 +0,0 @@ -# Context Baseline 25–30k — STATUS - -Session: context-baseline-25k -Last updated: 2026-08-04 -Lifecycle: **active** (M1 canary slice landed; default-on blocked until AC-A1) - -## Current state - -M0 complete. M1 progressive **tool** disclosure canary slice landed (`tool-tiers.ts` + `createBase2` wiring, default **off**). - -**Predecessor** `.agents/sessions/context-budget-architecture-2026-08/` remains complete for ledger/`/context`/disclosure _flag_/proactive cache/git gate/compaction/lifecycle. This program hits **≤30k fixed** (stretch 25k) without weakening the gate. - -## Ability regression bar (AC-A1) - -Added to SPEC (G8, R8, AC-A1) and PLAN (M1 default-on, M2-T8, X-T5, validation gates). - -Before flipping **M1** or **M2** to **default-on**: - -1. Full gate e2e suite, **and** -2. Buffbench subset **or** fixed smoke task set documented below -3. Results **no worse** than pre-flip baseline recorded here - -Canary-only shipping does **not** satisfy AC-A1. - -### Fixed smoke task set (draft for X-T5; finalize before first default-on) - -Placeholder until first default-on attempt — record commands + scores here: - -| ID | Task | Command / procedure | Pre-flip baseline | -| --- | --------------------------- | ---------------------------------------------------------------- | ----------------- | -| S1 | Gate lifecycle e2e | `bun test agents/e2e/gate-lifecycle.e2e.test.ts` | TBD | -| S2 | Gate aux ordering e2e | `bun test agents/e2e/gate-aux-ordering.e2e.test.ts` | TBD | -| S3 | Progressive disclosure unit | `bun test agents/__tests__/base2-progressive-disclosure.test.ts` | TBD | -| S4 | Context budget tools unit | `bun test agents/__tests__/base2-context-budget.test.ts` | TBD | -| S5 | Optional buffbench subset | document eval id + runner when used | TBD | - -## Baseline snapshot - -### Legacy script numbers (pre-M0, 10k tree — overstated) - -| Component | Tokens | -| --------------------------------------- | ------: | -| Tool definitions (33 tools) | 23,506 | -| base2 systemPrompt raw template | 11,612 | -| File tree (10k budget) | 9,126 | -| Knowledge instruction static | 1,036 | -| System info / git / patterns / language | ~1.5k | -| Proactive query_index (rep. 24) | ~4,912 | -| git_status rep. | ~52 | -| Script “fixed excl. injections” | ~46,931 | -| Script total + injections | ~51,895 | - -### M0 production-faithful numbers (2026-08-04T05:49Z) - -Source: `bun run scripts/measure-context-baseline.ts` exit 0 after M0 script rewrite. - -| Metric | Tokens | -| -------------------------------------------------------------- | -----------------------------------: | -| Default fixed (prod, disclosure off, SMALL tree, no proactive) | **49,345** | -| Default fixed if disclosure on (est.) | **46,661** | -| Authored surface off → on | 15,288 → 11,287 (**−4,001 / 26.2%**) | -| Injections (rep. proactive + git) | 4,960 | -| First-turn (fixed + injections) | 54,305 | -| Soft targets phase1/2/program/stretch | all OVER advisory vs 49,345 | - -#### Per-component (prod-fixed unless noted) - -| Component | Tokens | Tag | -| --------------------------------------- | --------: | ------------------------------------ | -| base2 systemPrompt raw (disclosure off) | 11,612 | prod-fixed | -| Tool definitions (33 tools) | 23,506 | prod-fixed | -| File tree SMALL (2.5k budget) | **2,016** | prod-fixed | -| File tree FULL (10k) | 9,175 | comparison | -| Knowledge contents | 98 | prod-fixed | -| Knowledge instruction | 1,036 | prod-fixed | -| System info | 407 | prod-fixed | -| Git changes prompt | **9,975** | prod-fixed (dirty worktree inflated) | -| Patterns index | 307 | prod-fixed | -| Language + engine profile | 388 | prod-fixed | -| Proactive query_index rep. | 4,912 | injection | -| git_status rep. | 48 | injection | - -#### Interpretation for planning - -- **SMALL tree is ~2.0k**, not ~9k — prior overstatement confirmed. -- **Tools (23.5k) still dominate** fixed cost → M1 is still the largest lever. -- **Git changes ~10k** on this run is session-dependent (large dirty/diff in worktree). On a clean tree this line is near 0–few hundred; **clean-ish fixed ≈ 49,345 − 9,975 ≈ 39.4k**. -- Disclosure on saves ~2.7k on raw system alone (est. fixed 46.7k) and **26.2% authored surface** (meets AC-P1 canary metric). -- AC-F1 (≤30k) is **not** met yet; gap is mostly tools + template + dirty git. - -## Completed work (this session) - -- [x] SPEC.md / PLAN.md / STATUS.md initial plan artifacts -- [x] AC-A1 ability-regression bar added to SPEC + PLAN + STATUS -- [x] M0-T1/T2/T3: production-faithful baseline measurement -- [x] M1-T1: `agents/base2/tool-tiers.ts` (CORE/IMPLEMENT/AUDIT/MEDIA_3D/JOB_EXTRA + `resolveModelToolNames`) -- [x] M1-T2: `createBase2` wires `toolNames` via resolver (mode gates preserved) -- [x] M1-T6: `progressiveToolDisclosure` option + `OPENBUFF_PROGRESSIVE_TOOL_DISCLOSURE` canary (default off) -- [x] M1 unit tests: `agents/__tests__/base2-progressive-tool-disclosure.test.ts` (core <12k, env canary, planOnly gates) -- [x] Docs: `docs/environment-variables.md`, `docs/configuration.md` -- [ ] M1-T3: handleSteps `unlockedToolTiers` deterministic unlock -- [ ] M1-T4: locked-tool runtime error path -- [ ] M1-T5: always-on “Tool surface” prompt block -- [ ] M1 default-on: **blocked until AC-A1** - -## Completed work (predecessor — do not redo) - -- Context budget ledger + `/context` -- Progressive prompt disclosure implementation (flag/env; guides; ≥25% test) -- Proactive retrieval cache + indexMutationEpoch invalidation -- Git status observation gating (SDK) -- Semantic compaction at model-aware trigger -- Tool-result lifecycle policy - -## Pending work - -1. **M0** — complete -2. **M1** — canary surface done; remaining unlock/runtime/prompt + default-on after AC-A1 -3. **M2** — Progressive prompt disclosure default-on **only after AC-A1** -4. **M3** — Smaller SMALL tree + knowledge instruction -5. **M4** — Lean proactive classifier + compact envelope -6. **M5** — Optional CORE schema diet -7. **M6** — History hygiene polish -8. **X-\*** — Docs, full validation, default flips - -## Next checkpoint - -**M1-T3/T4/T5:** deterministic unlock in handleSteps + locked-tool error + tool-surface prompt (still canary-only; no default-on). - -## Resume instructions - -1. Read `SPEC.md` then `PLAN.md`. -2. Run baseline script; update this STATUS with M0 numbers. -3. Gate e2e + **AC-A1** before any M1/M2 default-on. -4. Never relocate `gateAwarenessSection`. - -## Target scoreboard - -| | M0 measured | After program | -| --------------------------- | -------------------: | --------------------------: | -| Tools | 23.5k | 8–12k core (full on unlock) | -| Authored surface | 15.3k off / 11.3k on | keep ≥25% reduction | -| Tree (production SMALL) | **2.0k** | ~1.5–2k | -| Knowledge instruction | 1.0k | ~0.2k | -| Git changes (this worktree) | **10.0k** (dirty) | ~0 on clean | -| **Fixed total (this run)** | **49.3k** | **~25–30k** | -| **Fixed if clean git** | **~39.4k est.** | **~25–30k** | -| Proactive first hit | ~5k | ~0–1.5k | -| Gate | full strength | **unchanged** | -| Ability (default-on) | n/a | **AC-A1** hard bar | - -## Blockers - -None for M0 (complete). Default-on blocked until AC-A1 evidence is recorded. - - - -## M1-T3/T4/T5 verified landed — 2026-08-04 — 2026-08-04T21:12:43.936Z - -## M1-T3/T4/T5 verified landed — 2026-08-04 - -Resume at M1-T3 found the unlock runtime already shipped; verified against live code, no new edits this turn. - -- **M1-T3** deterministic unlock: `publishUnlockedToolTiers` defined inside serialized `handleSteps` (inlines deriveIntentSignals/resolveUnlockedTiersForPhase; canary via programmaticConfig). Inline-vs-canonical sync guard test passes across the phase/prompt matrix; canary-off clears stale unlocks. -- **M1-T4** runtime surface: `run-agent-step` filters custom/MCP `additionalToolDefinitions` via `getEffectiveAgentToolNames(template, agentState)`; loopAgentSteps test proves per-step expand AND shrink of the offered ToolSet; `tool-executor` rejects still-locked tier tools with `buildUnavailableToolMessage` (lists current available surface, actionable, fail-closed). -- **M1-T5** always-on `## Tool surface` prompt block present when progressiveToolDisclosure is on. - -Validation: `bun test agents/__tests__/base2-progressive-tool-disclosure.test.ts` 50/50 pass; `agent-tool-names.test.ts` 3/3; `agents` + `packages/agent-runtime` typecheck clean. - -PLAN checkboxes M1-T3/T4/T5 and the two open M1-T7 test bullets marked done. Next checkpoint unchanged: M1 default-on still **blocked until AC-A1** (gate e2e + buffbench subset or fixed smoke set, no worse than baseline). M2/M3/M4 remain pending. - - - -## Smoke set + M4 baselines — 2026-08-04 — 2026-08-04T23:20:04.488Z - -## Smoke set + M4 baselines — 2026-08-04 - -Context: after M4 landed (lean proactive inject), the fixed smoke set that gates any M1/M2 default-on is now recorded. All commands were run inside this session against the working tree at snapshot v3:34e577b5d0f3f5. - -| ID | Task | Command | Result | -| --- | --------------------------- | ---------------------------------------------------------------- | ------------------------------------------- | -| S1 | Gate lifecycle e2e | `bun test agents/e2e/gate-lifecycle.e2e.test.ts` | 3 pass / 0 fail | -| S2 | Gate aux ordering e2e | `bun test agents/e2e/gate-aux-ordering.e2e.test.ts` | 14 pass / 0 fail | -| S3 | Progressive disclosure unit | `bun test agents/__tests__/base2-progressive-disclosure.test.ts` | part of 61/61 combined with S4 | -| S4 | Context budget tools unit | `bun test agents/__tests__/base2-context-budget.test.ts` | part of 61/61 combined with S3 | -| S5 | Optional buffbench subset | document eval id + runner when used | deferred until the first default-on attempt | - -## Follow-up fixes (post-M4) - -The M4 code-reviewer pass returned `NON_BLOCKING` with two advisory findings; both are deferred for follow-up without re-arming the gate: - -- **NF-1 (prompt/payload mismatch):** `EXPLORE_PROMPT` in `agents/base2/base2.ts` still tells consuming subagents to "deduplicate its candidates, matchedSnippets, and relatedFiles", but the M4 compact envelope deliberately drops those fields from the proactive route note. Reconcile the wording (or note the fields come from the live query_index tool, not the injected note). -- **NF-2 (persisted `result` never re-read):** `toCompactProactiveRetrievalResult` returns `[{ type:'json', value }]` and the hit path yields only the route-note string; the stored `proactiveRetrievalCache.result` is asserted only by tests. Either consume `.result` on a cache hit or drop the persisted body if it is genuinely unused. - -Both findings are advisory; the gate passed with full validation and six satisfied requirement-coverages (serialization self-containment, cache contract, cache-hit semantics, regex safety, over-tightening check for broad/audit prompts, and regression-test presence). - - - -## NF-1/NF-2 resolved — 2026-08-04 — 2026-08-05T00:38:00.947Z - -Both M4 reviewer advisories fixed and locally validated (agents suites: base2 198/198, general-agent 7/7): - -- **NF-1:** `EXPLORE_PROMPT` now says "deduplicate its candidates by path, score, reason, and kind" — matching the compact proactive envelope's surviving fields instead of dropped `relatedFiles`/`matchedSnippets`. -- **NF-2:** cache-HIT route note now embeds `cachedProactiveRetrieval.result` (compact envelope) alongside the route metadata, so the persisted envelope is genuinely consumed rather than write-only. Defensive fallback omits the suffix when the entry is somehow absent. -- **Scope correction:** an attempted M4 mirror of the classifier onto `general-agent`'s `shouldProactivelyQueryIndex` (M4-T2) broke 3 audit-loop tests (audit prompts began firing query_index first, shifting the expected yield sequence). The file was reverted to HEAD; the mirror is re-queued as its own scoped change with test updates if pursued. -- Gate: hooks green on `agents/base2/base2.ts` + `agents/__tests__/base2.test.ts` (+ session artifacts); reviewer pass returned LOOKS_GOOD on the NF-1/NF-2 scope; deliverables committed as `77f403b`. - -## AC-A1 evaluation + M2 default-on flip — 2026-08-05 - -Pre-flip smoke baselines (S1–S4) were recorded on the pre-flip tree; the M2 default-on flip was then applied (`DEFAULT_PROGRESSIVE_PROMPT_DISCLOSURE = true`, resolver = `option ?? (envCanary || default)`) and the same suites re-run. - -Post-flip vs. baseline: - -| Suite | Baseline | Post-flip | Verdict | -| ---------------------------------- | --------------------------------------- | -------------------------------------------------------------------------- | -------- | -| S1 gate-lifecycle e2e | 3/3 | 3/3 (17/17 combined with S2) | no worse | -| S2 gate-aux-ordering e2e | 14/14 | 14/14 (combined above) | no worse | -| S3 progressive-disclosure unit | 9/9 | 9/9 (3 assertions realigned to default-on contract, equal strength) | no worse | -| S4 tool-tier + context-budget unit | 52/52 (61/61 combined with S3 pre-flip) | 52/52 | no worse | -| base2.test.ts (regression net) | 137/137 | 137/137 (2 stale verbatim-prompt assertions moved to explicit-off surface) | no worse | -| agents typecheck | clean | clean | no worse | - -AC4 metric confirmed post-flip by `scripts/measure-context-baseline.ts`: authored surface 15,290 -> 11,288 tok (-4,002, 26.2% >= 25% target). Caveat: the script's `Default fixed (prod)` line (48,034 tok) is unchanged because it measures the raw system template with unreplaced placeholders; the disclosure saving flows only through production runtime assembly (known SPEC open item #3 measurement gap). - -AC-A1 **satisfied for the prompt-disclosure flip**: no suite is worse than its recorded baseline. The M2 flip is ready to gate and commit. - - - -## Deliverables committed — 2026-08-05 — 2026-08-05T02:59:40.287Z - -H1 landed on `77f403b0110da31acff3b3e6b18db509bf44d8fa` (`feat(base2): tier model-visible tools by phase + slim proactive retrieval`, +3004/−129 across 14 files). M1 (progressive tool disclosure) and M4 (lean proactive inject) shipped; NF-1/NF-2 follow-ups resolved; session artifacts committed. The branch is ahead 1 of origin; no push was requested or performed. - -**Still outstanding:** M2 default-on for progressive prompt disclosure is blocked on AC-A1 smoke evidence (S1–S4 commands + pre-flip baselines not yet recorded in STATUS.md). That is the only remaining commitment-gate item. - - - -## AC-A1 pre-flip smoke baselines — 2026-08-05 — 2026-08-05T06:02:55.271Z - -## AC-A1 pre-flip smoke baselines — 2026-08-05 - -Recorded against HEAD `77f403b` (post-M1/M4, disclosure default OFF — the M2 flip candidate). All ran on this tree. - -| ID | Task | Command | Pre-flip result | -| --- | --------------------------- | ------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------- | -| S1 | Gate lifecycle e2e | `bun test agents/e2e/gate-lifecycle.e2e.test.ts` | **3 pass / 0 fail** (3 gate events visible: blocked→awaiting_review→final_response_allowed per test) | -| S2 | Gate aux ordering e2e | `bun test agents/e2e/gate-aux-ordering.e2e.test.ts` | **14 pass / 0 fail** | -| S3 | Progressive disclosure unit | `bun test agents/__tests__/base2-progressive-disclosure.test.ts` | **9 pass / 0 fail** | -| S4 | Context budget tools unit | `bun test agents/__tests__/base2-context-budget.test.ts` + `base2-progressive-tool-disclosure.test.ts` | **52 pass / 0 fail** (combined) | -| S5 | buffbench subset | (none wired) | deferred — no default-on attempt yet | - -AC-A1 rule: M2 default-on is acceptable only if the post-flip runs match these results on the same tree. - - - -## M3 tree/knowledge reductions — 2026-08-05 — 2026-08-05T10:11:05.319Z - -M3-T1..T4 landed (AC-A1 for M2 already satisfied). - -**Code** - -- `FILE_TREE_PROMPT_SMALL` budget **2_500 → 1_750** (`packages/agent-runtime/src/templates/strings.ts`). -- Agent-mode truncation **preferPathOnly** (`truncate-file-tree.ts` + `getProjectFileTreePrompt` passes `preferPathOnly: mode === 'agent'`); search/LARGE unchanged (symbol-rich). -- Static `knowledgeFilesPrompt` shrunk to short blurb + pointer; full essay at `agents/guides/knowledge-files.md`. Root knowledge **contents** injection unchanged. -- Baseline script `FILE_TREE_SMALL_BUDGET = 1_750` + comments. - -**Validation** - -- `bun test` truncate-file-tree + prompts-ledger: **10/10 pass** -- `packages/agent-runtime` typecheck: clean -- `bun run scripts/measure-context-baseline.ts`: - - File tree SMALL (1750 budget): **1,773** tok (was ~2.0k @ 2500) - - Knowledge instruction static: **99** tok (was ~1,036) - - Default fixed (prod, SMALL, no proactive): **46,808** tok (session-dependent git still inflates) - -**Next:** gate this M3 diff; optional commit after GATE: PASSED. M5 schema diet / M1 tool default-on still separate. - - - -## Resume 2026-08-05 — continue open plan — 2026-08-05T17:37:42.619Z - -Resumed context-baseline-25k. Verified live tree: - -- M1 canary surface (tool tiers + unlock + locked-tool path + Tool surface prompt): landed earlier; default-on still blocked on AC-A1 for tools (prompt AC-A1 already satisfied for M2). -- M2 progressive prompt disclosure default-on: committed `71eb68b44`. -- M3 cheaper SMALL tree (1750) + knowledge blurb: committed `9ce22c079`. -- M4 lean proactive: committed in `77f403b01` + NF-1/NF-2 follow-ups. - -PLAN checkboxes were stale; syncing done markers. Next implementation: **M5 schema/description diet** (rank CORE tools, shorten top descriptions, keep core ≤12k), then M6 history hygiene polish and remaining X-T docs if needed. - - - -## M5-T1 ranking — 2026-08-05 — 2026-08-05T17:44:16.775Z - -Measured via `scripts/rank-core-tool-tokens.ts` (new): - -| Metric | Tokens | -| ----------------------- | ---------: | -| CORE total | **14,183** | -| progressive core-only | **14,183** | -| full surface (33 tools) | **23,598** | -| AC-F2 core target | ≤12,000 | - -Top CORE costs: read_files 3002, spawn_agents 2876, query_index 1219, ask_user 1078, check_job 954, check_background_agent 936, suggest_followups 754, write_todos 627, list_jobs 569, glob 549. - -Next: M5-T2 shorten top CORE tool descriptions (~2.5–4k savings needed). - - - -## M5 schema diet complete — 2026-08-05 — 2026-08-05T17:46:34.460Z - -M5-T1 ranking + M5-T2/T3 description diet landed. - -**Token budget (scripts/rank-core-tool-tokens.ts):** -| Metric | Before | After | -|---|---:|---:| -| CORE / progressive core-only | 14,183 | **9,680** (AC-F2 ≤12k met) | -| Full surface (33 tools) | 23,598 | **19,095** (≤25k still met) | - -**Edits:** shortened model-facing `description` prose on top CORE tools under `common/src/tools/params/tool/` (read_files, spawn_agents, query_index, ask_user, check_job, check_background_agent, suggest_followups, write_todos, list_jobs, glob, read_subtree, read_logs). Ranking helper: `scripts/rank-core-tool-tokens.ts`. - -**Validation (local):** - -- `bun test` coerce-to-array + base2-progressive-tool-disclosure + base2-context-budget: green (core-only <12k assertion pass) -- `cd common && bun run typecheck`: clean - -**M6 quick check:** `tool-result-lifecycle.ts` already tags `query_index` as VERBOSE + normal importance (not pinned/high); spawn tools stay high. M6-T1 largely already satisfied; residual M6-T2/T3 optional polish only. - -**Still open on this plan:** M1 tool default-on (AC-A1 for tools), M5 default-on N/A (diet is always-on), optional M6 polish, X-T docs for tool tiers + schema diet. - - - -## M5 security review cleared — 2026-08-05T17:59:42.118Z - -Snapshot-bound security-reviewer returned LOOKS_GOOD (receipt 0xRtSQH7XPo / snapshot v3:8b8a8c283b248aeb571eb42f715cc448ac9d77f39c473a974eb8ecdeb6eb1db5). Coverage covered; no findingIds. Pending files: rank script + 12 CORE tool description diets. Local checks earlier: coreTotal 9680 (≤12k), 138+8 tests pass, common typecheck clean. Ending turn for automated hooks+reviewer gate. diff --git a/.agents/sessions/context-budget-architecture-2026-08/EVENTS.jsonl b/.agents/sessions/context-budget-architecture-2026-08/EVENTS.jsonl deleted file mode 100644 index 40ece89684..0000000000 --- a/.agents/sessions/context-budget-architecture-2026-08/EVENTS.jsonl +++ /dev/null @@ -1,12 +0,0 @@ -{"ts":"2026-08-01T06:32:17.067Z","kind":"append_lesson","summary":"Appended entry \"M0 numbers CORRECTED after under-measurement fix\" to STATUS.md","payload":{"heading":"M0 numbers CORRECTED after under-measurement fix","artifact":"STATUS.md"}} -{"ts":"2026-08-02T23:25:00.121Z","kind":"append_lesson","summary":"Appended entry \"M1 complete — context budget ledger + /context telemetry (gate-verified)\" to STATUS.md","payload":{"heading":"M1 complete — context budget ledger + /context telemetry (gate-verified)","artifact":"STATUS.md"}} -{"ts":"2026-08-02T23:25:00.121Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-08-03T12:57:34.912Z","kind":"append_lesson","summary":"Appended entry \"M4 complete + M5 verified already-implemented\" to STATUS.md","payload":{"heading":"M4 complete + M5 verified already-implemented","artifact":"STATUS.md"}} -{"ts":"2026-08-03T12:57:34.912Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-08-03T13:08:03.830Z","kind":"append_lesson","summary":"Appended entry \"M6 implementation in progress — tool-result lifecycle\" to STATUS.md","payload":{"heading":"M6 implementation in progress — tool-result lifecycle","artifact":"STATUS.md"}} -{"ts":"2026-08-03T13:12:52.889Z","kind":"append_lesson","summary":"Appended entry \"M6 complete — tool-result lifecycle (gate-verified)\" to STATUS.md","payload":{"heading":"M6 complete — tool-result lifecycle (gate-verified)","artifact":"STATUS.md"}} -{"ts":"2026-08-03T13:12:52.889Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-08-03T13:51:26.914Z","kind":"append_lesson","summary":"Appended entry \"M2 residual + M3 + X-T1/X-T2 complete (gate-verified NON_BLOCKING)\" to STATUS.md","payload":{"heading":"M2 residual + M3 + X-T1/X-T2 complete (gate-verified NON_BLOCKING)","artifact":"STATUS.md"}} -{"ts":"2026-08-03T13:51:26.914Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-08-03T17:58:41.623Z","kind":"append_lesson","summary":"Appended entry \"Optional tail complete — canary + tool schema cost\" to STATUS.md","payload":{"heading":"Optional tail complete — canary + tool schema cost","artifact":"STATUS.md"}} -{"ts":"2026-08-03T17:58:41.624Z","kind":"session_status","summary":"Session status -> completed","payload":{"status":"completed"}} diff --git a/.agents/sessions/context-budget-architecture-2026-08/PLAN.md b/.agents/sessions/context-budget-architecture-2026-08/PLAN.md deleted file mode 100644 index e2e408107e..0000000000 --- a/.agents/sessions/context-budget-architecture-2026-08/PLAN.md +++ /dev/null @@ -1,107 +0,0 @@ -# Context Budget Architecture — PLAN - -Session: context-budget-architecture-2026-08 -Spec: ./SPEC.md -Status: draft - -Sequencing principle: measure first (M1) so every later change is validated against real numbers, then the three independent reductions (M2 retrieval, M3 git, M4 prompt), then earlier compaction (M5), then lifecycle polish (M6). M2–M4 are independent and can be parallelized after M1. - -## Milestone 0 — Baseline measurement (no behavior change) - -Goal: capture current per-component token costs so reductions are provable. - -- [ ] M0-T1 Add a throwaway instrumentation harness (script under scripts/) that builds the orchestrator system prompt + file tree + knowledge for this repo and prints token counts per block using countTokensJson (packages/agent-runtime/src/util/token-counter.ts). Record numbers in STATUS.md. -- [ ] M0-T2 Measure one representative proactive query_index result and one git_status injection; record tokens. -- Validation: script runs via bun; numbers recorded. No source behavior change. - -## Milestone 1 — Context budget ledger + telemetry (R1, R2, AC1) - -Goal: per-turn, per-component accounting surfaced to the CLI. - -- [ ] M1-T1 Create packages/agent-runtime/src/util/context-budget.ts with ContextCategory, BudgetLine, ContextBudgetLedger, recordBlock, finalizeLedger, formatLedgerForCli (see SPEC interfaces). Reuse TOKEN_COUNT_CACHE. -- [ ] M1-T2 Instrument system-prompt assembly (packages/agent-runtime/src/system-prompt/prompts.ts: getProjectFileTreePrompt, knowledgeFilesPrompt, additionalSystemPrompts, getSystemInfoPrompt) to record blocks into a ledger passed via params or returned alongside. -- [ ] M1-T3 Instrument run-agent-step.ts contextTokenCount computation to attach ledger to session state/telemetry. -- [ ] M1-T4 Ship /context command (follow agents/patterns/ship-a-cli-command.md): register in cli/src/commands/command-registry.ts and cli/src/data/slash-commands.ts; render formatLedgerForCli. Optionally fold into /usage. -- [ ] M1-T5 Unit tests: ledger sums match countTokensJson of assembled request within 5%; format output stable. -- Validation: bun typecheck (packages/agent-runtime, cli); new unit tests pass; manual /context run. -- Depends on: M0. - -## Milestone 2 — Retrieval dedup + tighter classifier (R3, AC2) - -Goal: stop re-injecting identical proactive query_index results. - -- [ ] M2-T1 Add retrieval cache keyed by stableHash(normalizedQuery) + workspace revision (common/src/types/workspace-state.ts). Store last entry in agentState (mutableAgentState) with tokens + timestamp. -- [ ] M2-T2 In agents/base2/base2.ts around the proactive query_index yield (~line 704): on cache hit with unchanged revision, yield an add_message pointer (<200 tokens) instead of the full tool call; on miss, run query_index and record the entry. -- [ ] M2-T3 Invalidate cache on workspace revision change (advanceWorkspaceState) and on index markPathsChanged (packages/indexer/src/index-manager.ts). -- [ ] M2-T4 Tighten classifyProactiveRetrieval (base2.ts:7893): remove/weight-down generic triggers (context, index, flow) so they alone do not fire; keep strong-intent words. Mirror the change in agents/general-agent/general-agent.ts shouldProactivelyQueryIndex. -- [ ] M2-T5 Tests: extend agents/**tests**/base2.test.ts "base2 proactive index lookup" — assert second equivalent turn injects pointer not full result; assert generic-word prompt no longer triggers; assert invalidation after revision bump. -- Validation: bun test agents/**tests**/base2.test.ts; typecheck agents. -- Depends on: M1 (uses ledger tokens for assertions). Independent of M3/M4. - -## Milestone 3 — Git delta helper (R4, AC3) - -Goal: one guarded git observation instead of 9 redundant yields. - -- [ ] M3-T1 Add maybeYieldGitObservation helper (new util consumed by base2 handleSteps). Computes a worktree fingerprint, compares to state.lastGitFingerprint; returns full git_status on first/changed, compact delta add_message on minor change, undefined when unchanged. -- [ ] M3-T2 Replace all 9 inline git_status yields in agents/base2/base2.ts (lines 787, 1097, 2480, 3133, 3354, 3474, 4064, 4301, 4449) with calls to the helper. Preserve the semantic reason each site existed (turn start vs post-gate) via the helper's note field. -- [ ] M3-T3 Store lastGitFingerprint + workspaceRevision in agentState; reset on revision change. -- [ ] M3-T4 Tests: base2 handleSteps test asserting zero git blocks across unchanged consecutive turns and exactly one compact delta after a simulated change. Update affected e2e expectations (agents/e2e/gate-lifecycle.e2e.test.ts, reviewer-spawn-conditions.e2e.test.ts) only where the redundant yields were asserted. -- Validation: bun test agents; typecheck agents. -- Depends on: M1. Independent of M2/M4. -- Risk: many e2e tests assert git_status yields — audit assertions before replacing; keep first-observation behavior identical. - -## Milestone 4 — Progressive prompt disclosure (R5, AC4) - -Goal: shrink the always-on orchestrator system prompt; keep guidance retrievable. - -- [ ] M4-T1 Inventory the orchestrator system prompt (agents/base2/base2.ts systemPrompt + quality-prompt-section.ts + patterns index) and measure each section (from M0/M1 ledger). Identify verbose, rarely-needed blocks (specialist routing detail, security pattern list, full recovery workflow prose). -- [ ] M4-T2 Relocate verbose detail into skills/knowledge (existing skill loader; see cli/src/data/slash-commands.ts getSlashCommandsWithSkills and agents/patterns/). Replace inline prose with a compact always-on index: one line per rule block with trigger keywords and the skill/pattern name to load. -- [ ] M4-T3 Ensure trigger keywords in the compact index cause the model to load the relevant skill/pattern on demand (verify against skill tool behavior). -- [ ] M4-T4 Gate behind a config flag (default off) until evals show no task-success regression. -- [ ] M4-T5 Tests: assert always-on prompt token count reduced >= 25% vs M0 baseline; assert each relocated block is reachable via its named skill/pattern; snapshot test of the compact index. -- Validation: bun typecheck agents + cli; prompt budget test; run a small buffbench subset (evals/) comparing before/after if available. -- Depends on: M1. Independent of M2/M3. -- Risk: hidden guidance regression — mitigate with flag + eval comparison; do not drop any mandate, only relocate detail. - -## Milestone 5 — Earlier semantic compaction (R6, AC5) - -Goal: compact before the 190k emergency brake. - -- [ ] M5-T1 In packages/agent-runtime/src/run-agent-step.ts, compute getSemanticCompactionBudget(getEffectiveContextLimits(...)) for the active model and trigger semantic compaction at triggerBudgetTokens (not DEFAULT_MAX_CONTEXT_TOKENS). Keep maybePruneContext as the hard fallback. -- [ ] M5-T2 Ensure pinned control-plane memory (extractPinnedContextBlocks) and fixed baseline stay outside the history budget. -- [ ] M5-T3 Wire the LLM context-pruner (agents/context-pruner.ts) spawn to the budget trigger where smarter summarization is wanted; keep deterministic trim as fallback. -- [ ] M5-T4 Tests: extend packages/agent-runtime/src/util/**tests**/context-pruning.test.ts and agents/e2e/context-pruning-threshold.e2e.test.ts — synthetic growing history triggers compaction at triggerBudgetTokens, before 190k; pinned blocks retained. -- Validation: bun test packages/agent-runtime; e2e context-pruning tests. -- Depends on: M1. Can run parallel to M2–M4. - -## Milestone 6 — Tool-result lifecycle (R7) - -Goal: deterministic TTL/importance compression for verbose results. - -- [ ] M6-T1 Tag verbose tool results (query_index, read_files, spawn_agents) with importance + turn-born metadata at creation. -- [ ] M6-T2 Replace the blunt numToolResultsToKeep cutoff in simplifyToolResultHelper (messages.ts:153) with a deterministic policy: compress to a receipt after N turns unless pinned/important. Keep behavior monotonic and testable. -- [ ] M6-T3 Tests: unit tests for the lifecycle policy; ensure no pinned/active-work result is ever dropped. -- Validation: bun test packages/agent-runtime; messages.test.ts. -- Depends on: M5 (shares compaction path). - -## Cross-cutting - -- [ ] X-T1 Keep all behavior changes behind config flags with safe defaults; document flags in docs/configuration.md and docs/environment-variables.md. -- [ ] X-T2 Update docs/architecture.md "Deterministic Edits, Reviewer Gates, and Plan Artifacts" area with a short context-budget section once M1 lands. -- [ ] X-T3 Run full validation suite before finalizing: bun typecheck across packages/agent-runtime, agents, cli, common; bun test for touched packages; relevant e2e (context-pruning, base2 proactive, gate-lifecycle). - -## Validation gates (per milestone) - -- M1: ledger sum within 5% of request tokens; /context renders. -- M2: second-turn pointer <200 tokens; generic-word no-trigger; invalidation on revision bump. -- M3: zero git blocks on unchanged turns; one delta after change. -- M4: >=25% always-on prompt reduction; all blocks retrievable; no eval regression (flag-gated). -- M5: compaction at triggerBudgetTokens before 190k; pinned retained. -- M6: lifecycle compression deterministic; pinned never dropped. - -## Risks tracker - -- e2e tests asserting redundant git_status yields (M3) — audit first. -- prompt disclosure regressions (M4) — flag + eval comparison. -- token-count drift vs provider — advisory ledger first, not a hard gate. -- stale retrieval cache after external edits (M2) — revision + index-staleness invalidation. diff --git a/.agents/sessions/context-budget-architecture-2026-08/SPEC.md b/.agents/sessions/context-budget-architecture-2026-08/SPEC.md deleted file mode 100644 index 49f0ec9cb8..0000000000 --- a/.agents/sessions/context-budget-architecture-2026-08/SPEC.md +++ /dev/null @@ -1,120 +0,0 @@ -# Context Budget Architecture — SPEC - -Status: draft -Session: context-budget-architecture-2026-08 -Owner: orchestrator (Buffy) -Related audits: docs/audits/read-write-indexing-2026-07-14 (workspace revision now exists in common/src/types/workspace-state.ts) - -## Problem statement - -The model context window fills quickly because cost lives in two pools and only one is managed: - -1. Fixed per-turn baseline (NEVER pruned). Re-sent on every request: - - Orchestrator system prompt (core mandates, edit mandates, harness recovery, spawning guidelines, gate semantics, git discipline, security patterns, specialist routing) plus inlined AGENTS.md, patterns index, language profile. - - Project file tree via getProjectFileTreePrompt (packages/agent-runtime/src/system-prompt/prompts.ts:121), token-budgeted but large. - - Knowledge files via knowledgeFilesPrompt (prompts.ts:14) and docs/architecture.md. -2. Conversation history (the ONLY pruned pool). Pruning is reactive: maybePruneContext (packages/agent-runtime/src/util/context-pruning.ts:262) does nothing until total tokens exceed DEFAULT_MAX_CONTEXT_TOKENS = 190_000, then trimMessagesToFitTokenLimitWithReport (packages/agent-runtime/src/util/messages.ts:385) drops oldest messages and simplifies old tool results (keep N most recent full via numToolResultsToKeep). - -Automatic injections compound the problem: - -- Proactive query_index: classifyProactiveRetrieval (agents/base2/base2.ts:7893) fires on a very broad keyword set (code, file, repo, project, module, package, function, class, component, hook, api, schema, config, test, implement, fix, debug, refactor, audit, review, investigate, architecture, flow, index, context). Result is a verbose programmatic_tool_result block (often ~10k tokens) with NO cross-turn dedup. A second site exists in agents/general-agent/general-agent.ts (shouldProactivelyQueryIndex). -- git_status: yielded at 9 sites in agents/base2/base2.ts handleSteps (lines 787, 1097, 2480, 3133, 3354, 3474, 4064, 4301, 4449). Cheap individually but redundant; re-fires even when the worktree is unchanged, and each injection rides along in history. - -Root structural gap: there is no per-component token accounting and no per-turn budget enforcement. The biggest, most repetitive costs are in the unmanaged fixed baseline. - -## Goals - -- G1. Measure: produce a per-turn, per-component token ledger (system sections, file tree, knowledge, proactive retrieval, git observations, conversation) with telemetry. -- G2. Retrieval dedup: make proactive query_index cached and deduped by (normalized query, workspace revision); inject a compact pointer on repeats instead of the full result. -- G3. Git deltas: collapse the 9 git_status yield sites behind one guarded helper that injects a compact delta only when the worktree changed since the last observation. -- G4. Progressive prompt disclosure: move large static rule blocks into retrievable skills/knowledge and keep only a compact index inline; fetch full text on demand. -- G5. Earlier, gentler compaction: wire the existing model-aware semantic budget into the main loop so compaction starts before the 190k emergency brake. - -## Non-goals - -- Not changing the provider/model routing or BYOK architecture. -- Not rewriting the LLM-based context-pruner agent's summarization heuristics (agents/context-pruner.ts) beyond wiring it to the budget; its knowledge-memory and pinned-active-work logic stays intact. -- Not altering the deterministic edit / read-capability system. -- Not changing the reviewer/validation gate contract. -- No new third-party dependencies. - -## Requirements - -- R1 (Ledger). Introduce a ContextBudgetLedger type and a collector that measures each injected block with countTokens/countTokensJson (packages/agent-runtime/src/util/token-counter.ts) and records category, tokens, and cacheability. Must be O(blocks) and reuse the existing TOKEN_COUNT_CACHE. -- R2 (Telemetry surface). Expose the ledger to the CLI via a new /context command (or extend /usage) showing per-component token cost and % of window. Follow agents/patterns/ship-a-cli-command.md and register in cli/src/commands/command-registry.ts + cli/src/data/slash-commands.ts. -- R3 (Retrieval cache). Add a retrieval cache keyed by stableHash(normalizedQuery) + workspace revision (common/src/types/workspace-state.ts stableHash/advanceWorkspaceState). On hit with unchanged revision, emit a one-line pointer message instead of the full query_index result. Tighten classifyProactiveRetrieval so generic words (context, index, flow) alone do not trigger; require stronger intent. -- R4 (Git delta helper). Add a single guarded helper (e.g. maybeYieldGitObservation) that compares current worktree fingerprint against the last-observed fingerprint stored in agentState; yields a compact delta ("+N modified since last check" or full status on first observation / after change) and nothing when unchanged. Replace all 9 inline git_status yields in base2.ts with calls to it. -- R5 (Progressive disclosure). Split the orchestrator system prompt into a compact always-on index plus on-demand detail blocks served through the existing skill/knowledge mechanism. Keep behavior-equivalent guidance reachable; do not drop any mandate, only relocate verbose detail. -- R6 (Early compaction). Call getSemanticCompactionBudget (context-pruning.ts:109) in the main loop (packages/agent-runtime/src/run-agent-step.ts) and trigger semantic compaction at triggerBudgetTokens rather than waiting for DEFAULT_MAX_CONTEXT_TOKENS. Preserve pinned control-plane memory and the fixed baseline (they sit outside the history budget). -- R7 (Tool-result lifecycle). Tag verbose tool results (query_index, read_files, spawn_agents) with a TTL/importance so they compress to receipts after a configurable number of turns, replacing the blunt numToolResultsToKeep cutoff. Must remain deterministic and testable. -- R8 (Backward compat). All changes behind feature flags / config with safe defaults; existing tests (agents/**tests**/base2.test.ts "base2 proactive index lookup", packages/agent-runtime/src/util/**tests**/context-pruning.test.ts, agents/e2e/context-pruning-threshold.e2e.test.ts) must keep passing. - -## Acceptance criteria - -- AC1. A /context (or /usage) invocation prints a per-component token breakdown summing to within 5% of the actual request token count. -- AC2. Repeating an equivalent prompt within an unchanged workspace revision injects a pointer (<200 tokens) instead of a full query_index result; verified by a unit test asserting the second turn's injected block size. -- AC3. With an unchanged worktree, consecutive turns inject zero git_status blocks; after a real change, exactly one compact delta is injected. Verified by a base2 handleSteps test. -- AC4. The always-on orchestrator system prompt token count drops by a measurable target (measure baseline first; target >= 25% reduction) with all guidance still retrievable on demand. -- AC5. Semantic compaction triggers at the model-aware triggerBudgetTokens on a synthetic growing history, before reaching DEFAULT_MAX_CONTEXT_TOKENS; verified by an e2e/property test. -- AC6. No regression in existing context-pruning, proactive-retrieval, and gate tests. - -## Relevant systems (exact files) - -- packages/agent-runtime/src/util/context-pruning.ts — budgets, maybePruneContext, getSemanticCompactionBudget, getEffectiveContextLimits, DEFAULT_MAX_CONTEXT_TOKENS. -- packages/agent-runtime/src/util/messages.ts — trimMessagesToFitTokenLimitWithReport, simplifyToolResultHelper, getContextCategoryTelemetry, extractPinnedContextBlocks, numToolResultsToKeep, shortenedMessageTokenFactor. -- packages/agent-runtime/src/util/token-counter.ts — countTokens, countTokensJson, countTokensForFiles, TOKEN_COUNT_CACHE. -- packages/agent-runtime/src/system-prompt/prompts.ts — getProjectFileTreePrompt, getSystemInfoPrompt, getGitChangesPrompt, knowledgeFilesPrompt, additionalSystemPrompts. -- packages/agent-runtime/src/system-prompt/truncate-file-tree.ts — truncateFileTreeBasedOnTokenBudget. -- packages/agent-runtime/src/system-prompt/search-system-prompt.ts — getSearchSystemPrompt. -- packages/agent-runtime/src/run-agent-step.ts — loopAgentSteps, contextTokenCount computation, where maybePruneContext is called. -- agents/base2/base2.ts — classifyProactiveRetrieval (7893), proactive query_index yield (~704), 9 git_status yields. -- agents/general-agent/general-agent.ts — shouldProactivelyQueryIndex (second proactive site). -- agents/context-pruner.ts — LLM summarization, knowledge memory, pinned active work. -- common/src/types/workspace-state.ts — WorkspaceStateV1, stableHash, advanceWorkspaceState, createInitialWorkspaceState. -- sdk/src/services/workspace-journal.ts, workspace-mutation-broker.ts — revision/journal machinery. -- cli/src/commands/command-registry.ts, cli/src/data/slash-commands.ts — command registration. -- packages/indexer/src/index-manager.ts, cli/src/utils/index-workspace-watcher.ts — index staleness/refresh. - -## Key interfaces (pseudo-code) - -// packages/agent-runtime/src/util/context-budget.ts (NEW) -export type ContextCategory = -| 'system-core' | 'system-rules' | 'file-tree' | 'knowledge' -| 'proactive-retrieval' | 'git-observation' | 'conversation' | 'tool-result' - -export interface BudgetLine { category: ContextCategory; label: string; tokens: number; cacheable: boolean } -export interface ContextBudgetLedger { -lines: BudgetLine[] -totalTokens: number -windowTokens: number -reservedTokens: number -byCategory: Record -} -export function recordBlock(ledger, category, label, content, opts?): BudgetLine -export function finalizeLedger(ledger, windowTokens): ContextBudgetLedger -export function formatLedgerForCli(ledger): string - -// retrieval cache (NEW, near base2 or a shared util) -export interface RetrievalCacheEntry { queryHash: string; workspaceRevision: number; tokens: number; at: number } -export function retrievalCacheKey(query: string, revision: number): string -export function shouldReuseRetrieval(entry, query, revision): boolean - -// git delta helper (NEW helper consumed by base2 handleSteps) -export function maybeYieldGitObservation(state: { lastGitFingerprint?: string; workspaceRevision?: number }): -| { toolName: 'git_status'; input: {}; note: 'first'|'changed' } -| { addMessage: string } // compact delta or nothing -| undefined // unchanged -> inject nothing - -## Risks and mitigations - -- Risk: progressive disclosure hides guidance the model needs. Mitigation: keep a compact always-on index with trigger keywords; measure task success on evals before/after; gate behind flag. -- Risk: retrieval dedup serves stale candidates after external edits. Mitigation: key on workspace revision from workspace-state.ts and index staleness; invalidate on markPathsChanged. -- Risk: git delta suppresses a needed observation. Mitigation: always emit full status on first observation and after any revision change; delta only within an unchanged revision. -- Risk: early compaction drops pinned control-plane memory. Mitigation: keep pinned blocks (extractPinnedContextBlocks) and fixed baseline outside the history budget, as today. -- Risk: token counting drift vs provider. Mitigation: reuse existing gpt-tokenizer + ANTHROPIC_TOKEN_FUDGE_FACTOR; treat ledger as advisory telemetry, not a hard request gate initially. - -## Out of scope (future) - -- Hard per-component request gating (refuse to inject over budget) — start advisory. -- Semantic (embedding) retrieval dedup — start with exact normalized-query + revision key. -- Cross-session budget persistence. diff --git a/.agents/sessions/context-budget-architecture-2026-08/STATE.json b/.agents/sessions/context-budget-architecture-2026-08/STATE.json deleted file mode 100644 index e1d933dd80..0000000000 --- a/.agents/sessions/context-budget-architecture-2026-08/STATE.json +++ /dev/null @@ -1,10 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "context-budget-architecture-2026-08", - "status": "completed", - "currentTask": null, - "revision": 5, - "checkpoint": null, - "createdAt": "2026-08-02T23:25:00.119Z", - "updatedAt": "2026-08-03T17:58:41.620Z" -} diff --git a/.agents/sessions/context-budget-architecture-2026-08/STATUS.md b/.agents/sessions/context-budget-architecture-2026-08/STATUS.md deleted file mode 100644 index b2aab98ef1..0000000000 --- a/.agents/sessions/context-budget-architecture-2026-08/STATUS.md +++ /dev/null @@ -1,226 +0,0 @@ -# Context Budget Architecture — STATUS - -Session: context-budget-architecture-2026-08 -Last updated: 2026-08-01 - -## Current state - -Milestone 0 (baseline measurement) complete. Script created and run. Numbers recorded below. - -## M0 Baseline Numbers (2026-08-01) - -Measured via `bun run scripts/measure-context-baseline.ts` on the openbuff repo itself. -Token counts use gpt-tokenizer with 1.35x Anthropic fudge factor. - -### Per-component breakdown - -| Component | Tokens | Notes | -| ------------------------------------- | ---------- | -------------------------------------------------------- | -| base2 systemPrompt (raw template) | 10,674 | 39,283 chars; includes unreplaced placeholders | -| Knowledge files instruction (static) | 1,036 | The "how to use knowledge files" prompt | -| Patterns index prompt | 307 | Compact catalog from agents/patterns/INDEX.md | -| File tree prompt (agent, 10k budget) | 190 | UNDERESTIMATE — session-state extraction missed fileTree | -| System info prompt | 109 | OS, shell, chrome, recently-read files | -| Language + engine profile | 0 | UNDERESTIMATE — depends on fileTree which was empty | -| Knowledge files (root-level contents) | 0 | UNDERESTIMATE — sessionState.knowledgeFiles was empty | -| Git changes prompt | 0 | Clean tree at measurement time | -| **Fixed baseline subtotal** | **12,316** | Excludes injections; underestimates true cost | -| Proactive query_index (24 results) | 4,912 | scope=multi-file, mode=explain, limit=24 | -| git_status injection | 52 | Compact: branch + dirty paths | -| **Automatic injections subtotal** | **4,964** | Per-turn when proactive retrieval fires | -| **Total measured** | **17,280** | 9.1% of 190k max | - -### Known underestimates - -The script extracts ProjectFileContext fields from `initialSessionState()`, but the -session-state shape does not directly expose `knowledgeFiles`, `fileTree`, and -`userKnowledgeFiles` at the top level the script expected. In production: - -- **Knowledge file contents** (knowledge.md, AGENTS.md, docs/architecture.md) add - roughly 3,000–5,000 tokens based on the files present in this repo. -- **File tree** at 10k budget adds roughly 2,000–4,000 tokens for a repo this size. -- **Language profile** adds roughly 500–1,000 tokens. -- **Tool definitions** (schemas for ~40+ tools) are sent alongside the system prompt - but NOT measured by this script. Estimated 15,000–25,000 tokens based on the - tool count and schema verbosity. - -### Realistic per-turn fixed cost estimate - -| Component | Est. tokens | -| ------------------------------------------------ | ------------------ | -| System prompt (assembled, placeholders replaced) | ~15,000–18,000 | -| Tool definitions (schemas) | ~15,000–25,000 | -| Per-turn injections (query_index + git_status) | ~5,000 | -| **Total fixed per-turn overhead** | **~35,000–48,000** | -| As % of 190k max | ~18–25% | - -### Key takeaways for prioritization - -1. **base2 systemPrompt template (10,674 tok)** is the single largest measured fixed - cost and the primary target for M4 (progressive disclosure). A 25% reduction - saves ~2,700 tokens/turn. -2. **Proactive query_index (~4,912 tok/injection)** is the largest per-turn variable - cost. M2 (dedup + tighter classifier) can eliminate most repeat injections. -3. **git_status (52 tok)** is cheap individually but fires at 9 sites; M3 (delta - helper) eliminates redundant fires. Cost savings are small but it removes - noise from history that compounds over turns. -4. **Tool definitions** are a large unmeasured cost. Not in scope for this plan - but worth a future investigation (lazy tool registration, schema compression). -5. **Pruning threshold (190k)** means the context grows to ~190k before any cleanup. - M5 (earlier semantic compaction) addresses this. - -## Completed work - -- [x] SPEC.md written -- [x] PLAN.md written -- [x] M0-T1: scripts/measure-context-baseline.ts created and runs clean (exit 0) -- [x] M0-T2: Representative query_index and git_status injections measured -- [x] Numbers recorded in this file - -## Next steps - -- M1: Context budget ledger + /context telemetry -- M2: Retrieval dedup (can parallelize with M3, M4 after M1) -- M3: Git delta helper -- M4: Progressive prompt disclosure -- M5: Earlier semantic compaction - -## Resume instructions - -1. Read SPEC.md and PLAN.md in this directory. -2. The baseline script is at scripts/measure-context-baseline.ts. -3. Fix the session-state → ProjectFileContext extraction to get accurate - knowledge/tree/profile numbers before starting M1. -4. Start M1-T1: create packages/agent-runtime/src/util/context-budget.ts. - - - -## M0 numbers CORRECTED after under-measurement fix — 2026-08-01T06:32:17.067Z - -The initial M0 run under-measured because the script stub-extracted ProjectFileContext fields. The hardened script now uses the real `sessionState.fileContext`, and the base2 template measurement reflects the fully-assembled default-mode systemPrompt (createBase2('default')). Re-run 2026-08-01T06:30Z: - -| Component | Tokens (corrected) | Was | -| ------------------------------------- | ------------------ | ---------- | -| base2 systemPrompt (raw template) | 11,631 | 10,674 | -| File tree prompt (agent, 10k) | 9,729 | 190 | -| Knowledge files (root-level contents) | 98 | 0 | -| Knowledge files instruction (static) | 1,036 | 1,036 | -| System info prompt | 326 | 109 | -| Git changes prompt | 363 | 0 | -| Patterns index prompt | 307 | 307 | -| Language + engine profile | 388 | 0 | -| **Fixed baseline** | **23,878** | **12,316** | -| Proactive query_index (24) | 4,912 | 4,912 | -| git_status | 52 | 52 | -| **Injections/turn** | **4,964** | **4,964** | -| **Total measured** | **28,842** | **17,280** | -| % of 190k max | **15.2%** | 9.1% | - -Key revision: the fixed baseline was ~2x larger than first measured, dominated by the file tree (9,729) and the assembled systemPrompt template (11,631). Priorities for M4 (progressive disclosure) and M1 telemetry are unchanged but better justified. Tool-definition schemas remain the largest UNMEASURED cost (est 15–25k). Script hardening complete: isInjection/isRawTemplate flags, KNOWLEDGE_FILE_NAMES_LOWERCASE reuse, and exit-1 on component error. - - - -## M1 complete — context budget ledger + /context telemetry (gate-verified) — 2026-08-02T23:25:00.119Z - -M1 is complete and gate-verified (validation hooks + code-reviewer NON_BLOCKING on snapshot v3:c9d9450c41c52). - -- M1-T1 ledger module (`packages/agent-runtime/src/util/context-budget.ts`) — already existed with createBudgetLedger/recordBlock/applyMeasure/finalizeLedger + full unit tests. -- M1-T2 instrument system-prompt assembly — already existed: `applyMeasure` wired into `prompts.ts` (fileTree, systemInfo, gitChanges) via `formatPrompt`'s optional `ledger` param. -- M1-T3 wire ledger into run-agent-step — DONE: `createBudgetLedger` before system-prompt assembly in `run-agent-step.ts`, threaded through the `systemPrompt` getAgentPrompt call, persisted to `initialAgentState.contextBudgetLedger` via `finalizeLedger` guarded by `builtSystemPromptThisTurn` (cached-prompt turns keep the prior ledger by identity). Added `AgentState.contextBudgetLedger?: ContextBudgetLedger` + `BudgetLine`/`BudgetCategory` types to `common/src/types/session-state.ts`. -- M1-T4 /context CLI command — DONE: `cli/src/commands/context.ts` `handleContextCommand` reads `runState.sessionState.mainAgentState.contextBudgetLedger`, renders via `formatLedgerForCli`, registered in `command-registry.ts` + `slash-commands.ts` as `/context` (alias `/ctx`). -- M1-T5 unit tests — DONE: existing ledger tests + new `cli/src/commands/__tests__/context.test.ts` (breakdown + no-data fallback) and a `loop-agent-steps.test.ts` case (ledger populated on build turn, kept by identity on cached-prompt turn). - -Repairs during gate: moved `formatLedgerForCli` into `common/src/util/context-budget.ts` (shared single implementation, re-exported from agent-runtime); tightened the `BudgetLine.category` mirror from `string` to the canonical `BudgetCategory` union; documented the stale-windowTokens edge on cached-prompt turns. - -Two NON_BLOCKING reviewer nits left for later: brittle substring assertions in context.test.ts (could pin exact formatted lines), and no zero-window/empty-ledger formatter branch coverage. - -Next: M2 (retrieval dedup), M3 (git delta helper), M4 (progressive prompt disclosure) — independent, can parallelize. M5 (earlier semantic compaction) after M1. - - - -## M4 complete + M5 verified already-implemented — 2026-08-03T12:57:34.911Z - -## Progress checkpoint (2026-08-03) - -### M4 — Progressive prompt disclosure — DONE (committed `a1882a65`) - -- `createBase2({ progressivePromptDisclosure })` default off; flag-on relocates five advisory sections to `agents/guides/*`. -- Tests: `agents/__tests__/base2-progressive-disclosure.test.ts` (7/7) including qualitySection/code-craftsmanship coverage after reviewer repair. -- Gate: LOOKS_GOOD; typecheck agents + agent-runtime green. Not pushed. - -### M5 — Earlier semantic compaction — ALREADY IMPLEMENTED (no new code needed) - -Verified against SPEC R6 / AC5 by reading source + running focused suites: - -- `getSemanticCompactionBudget` + model-aware trigger/target in `packages/agent-runtime/src/util/context-pruning.ts`. -- Runtime: `run-agent-step.ts` computes semantic budget, detects retained ``, emits `semantic_compaction` before `maybePruneContext` (mechanical emergency brake uses `providerSafeMessageLimit`, not bare 190k-only first response). -- Pruner injection: `run-programmatic-step.ts` injects authoritative `semanticBudget` into context-pruner generator params. -- Pinned/control-plane: extractPinnedContextBlocks + knowledge_memory retention; ledger annotate via `annotateLedgerAfterCompaction`. -- Tests green (33/33): `context-pruning.test.ts` + `loop-agent-steps.test.ts` including "runs semantic programmatic compaction before the mechanical brake" and small-model budget cases; context-pruner unit cases scale trigger under/over across windows (8k–1M). - -### Remaining plan work - -- **M3** (git delta helper): SPEC still open — no `maybeYieldGitObservation` symbol in tree; plan checkboxes unchecked. (Earlier session notes may have partially addressed git_status noise; re-audit before implementing.) -- **M2** residual: base2 has per-session proactive retrieval cache; TODO remains for index-manager `markPathsChanged` invalidation. -- **M6** (tool-result lifecycle): still uses blunt `numToolResultsToKeep = 1` in `messages.ts` simplifyToolResultHelper — next implementation milestone. -- Cross-cutting X-T1/X-T2 docs flags still open. - -### Next checkpoint - -Start **M6 tool-result lifecycle** (R7) unless user redirects to residual M3 git delta. - - - -## M6 implementation in progress — tool-result lifecycle — 2026-08-03T13:08:03.830Z - -M6 (R7) code landed locally, awaiting gate: - -- New pure policy: `packages/agent-runtime/src/util/tool-result-lifecycle.ts` (tags + shouldKeepFullToolResult; protected does not consume keep-N budget). -- Wired into `simplifyToolResultHelper` / trim first pass in `messages.ts`. -- Tagging at creation in both `tool-executor.ts` ToolMessage sites (`lifecycleTagsForToolResult` + `sentAt`). -- Tests: `tool-result-lifecycle.test.ts` + messages trim case for pinned full vs normal simplify; local suite 57/57 pass. - -Note: hard-budget second pass in trim may still content-simplify even `keepDuringTruncation` tool bodies (pre-existing emergency behavior; existing tests require it). M6 protects the first-pass keep-N policy + never _drops_ pinned messages. - - - -## M6 complete — tool-result lifecycle (gate-verified) — 2026-08-03T13:12:52.888Z - -M6 (SPEC R7) complete and gate-verified LOOKS_GOOD on snapshot v3:77a0076fba5b5. - -Shipped: - -- `packages/agent-runtime/src/util/tool-result-lifecycle.ts` — pure policy (tags, isProtectedToolResult, shouldKeepFullToolResult); protected does not consume keep-N budget; default N=1 for newest unprotected summarizable results. -- Creation tagging in both tool-executor ToolMessage sites (`lifecycleTagsForToolResult` + `sentAt`). -- `messages.ts` simplifyToolResultHelper wired to policy on the first trim pass. -- Tests: tool-result-lifecycle unit + messages trim integration (pinned full vs normal simplify). - -Caveat (documented): hard-budget second pass may still content-simplify keepDuringTruncation tool bodies under extreme pressure (pre-existing emergency behavior). - -Also verified earlier this turn: M5 already implemented (no new code); M4 committed as `a1882a65`. - -Remaining plan work: residual M3 git delta helper re-audit; M2 markPathsChanged invalidation TODO; X-T1/X-T2 docs. - - - -## M2 residual + M3 + X-T1/X-T2 complete (gate-verified NON_BLOCKING) — 2026-08-03T13:51:26.913Z - -All three suggested follow-ups are complete and gate-verified (NON_BLOCKING, 3 minor nits accepted without code changes): - -- **M2 residual** — `IndexManager.indexMutationEpoch` (in-process monotonic, additive getter; increments on markStale/markPathsChanged); `query_index` result includes the epoch; base2 proactive cache invalidates when the newest epoch in messageHistory differs from the cache entry at the same workspace revision. Tests: IndexManager epoch + base2 epoch-guard invalidation. -- **M3** — already satisfied by SDK `applyGitStatusGate` (`sdk/src/run.ts`): unchanged per-turn git_status repeats compact to an "unchanged" note; no new base2 helper needed. Verified git-status-gate suite 12/12 + base2 suite green. -- **X-T1/X-T2** — context-budget documentation added to `docs/architecture.md` (Context Budget, Retrieval Caching, and Git Observation Gating), `docs/configuration.md` (Context budget and proactive retrieval), and `docs/environment-variables.md` (no new OPENBUFF\_\* vars). - -Context-budget plan is now substantially complete across M0–M6 + cross-cutting docs. Remaining optional work: canary progressivePromptDisclosure, measure tool-definition schema cost, eval comparison. - - - -## Optional tail complete — canary + tool schema cost — 2026-08-03T17:58:41.620Z - -Follow-up optional work delivered: - -- **Canary progressivePromptDisclosure**: `OPENBUFF_PROGRESSIVE_PROMPT_DISCLOSURE` (1/true/yes/on) enables disclosure when `createBase2` option is omitted; explicit boolean still wins. Tests cover env on + explicit false override. Docs updated in environment-variables.md + configuration.md. -- **Tool-definition schema cost measured**: baseline script now reports ~23.5k tokens for base2 default tools (33 tools). Runtime `run-agent-step` records category `tools` / label `tool definitions` on ledger rebuild turns. -- **Eval comparison**: no buffbench run; AC4 (>=25% authored surface reduction) remains the unit-test canary acceptance metric in base2-progressive-disclosure.test.ts. - -Session can be treated as completed for plan purposes. diff --git a/.agents/sessions/context-budget-architecture-2026-08/baseline-output.txt b/.agents/sessions/context-budget-architecture-2026-08/baseline-output.txt deleted file mode 100644 index 30ba1e0ba2..0000000000 --- a/.agents/sessions/context-budget-architecture-2026-08/baseline-output.txt +++ /dev/null @@ -1,33 +0,0 @@ -Using environment: dev -=== Context Budget Baseline Measurement === -Project: /home/ben/Code/CLI/openbuff -Date: 2026-08-01T06:30:37.529Z - -Building real project context via SDK discovery... - OK — project context built. - -=== Per-Component Token Breakdown === - - base2 systemPrompt (raw template) 11,631 tok (42464 chars, includes placeholders) - File tree prompt (agent, 10k budget) 9,729 tok (25171 chars) - Knowledge files (root-level contents) 98 tok (8 total knowledge files, 306 chars rendered) - Knowledge files instruction (static) 1,036 tok (3699 chars) - System info prompt 326 tok (1008 chars) - Git changes prompt 363 tok (1210 chars) - Proactive query_index (representative, 24 results) 4,912 tok (scope=multi-file, mode=explain, limit=24) - git_status injection (representative) 52 tok (compact: branch + dirty paths only) - Patterns index prompt 307 tok (907 chars) - Language + engine profile prompt 388 tok (1562 chars) - ---- Summary --- - Fixed per-turn baseline (excl. injections): 23,878 tokens - Automatic injections (per-turn): 4,964 tokens - Total measured: 28,842 tokens - DEFAULT_MAX_CONTEXT_TOKENS: 190,000 tokens - Baseline as % of max: 15.2% - -Note: token counts use gpt-tokenizer with 1.35x Anthropic fudge factor. -Note: the base2 systemPrompt is measured as the raw template; after placeholder - replacement the assembled prompt is larger (file tree + knowledge + profiles - are inlined). The true per-turn system cost is approximately: - raw template (11,631) + inlined components (12,247) = 23,878 tokens diff --git a/.agents/sessions/dynamic-cross-session-memory/LESSONS.md b/.agents/sessions/dynamic-cross-session-memory/LESSONS.md deleted file mode 100644 index 7cce756995..0000000000 --- a/.agents/sessions/dynamic-cross-session-memory/LESSONS.md +++ /dev/null @@ -1,11 +0,0 @@ -# LESSONS: Dynamic Cross-Session Memory - -## 2026-08-22 - -- The runtime already deduped merges (`uniqueRecent`, `normalizeEvidence`) and excluded stale evidence (`evidenceIsFresh`) — verify existing helpers before writing new ones; the real gap was persistence + reconciliation only. -- `TaskMemoryV1.evidence[]` was schema-designed for cross-session staleness (`freshnessHash`, `workspaceRevision`, `supersedes`, `stale`) months before anyone wired persistence; read schemas for intent, not just current usage. -- `runOnce` has many terminal paths; the only reliable success discriminator is `terminalState.output.type !== 'error'` (cancelled/error states always emit `type:'error'`). -- Persisted evidence paths and journal move destinations are untrusted input: fail closed on path escape before any file read. -- Fixed `.tmp` filenames race concurrent saves in one cwd; unique suffix (pid+random) + best-effort unlink on failure. -- Delegation reliability: this session saw thinker/web-researcher/architect/editor spawn failures; when a specialist fails twice with the same class of error, implement directly with edit_transaction rather than burning retries. -- FNV-1a now lives in three places (sdk store, run.ts git-status fingerprint, agent-runtime task-memory); the @codebuff/common shared helper should land before a fourth copy appears. diff --git a/.agents/sessions/dynamic-cross-session-memory/PLAN.md b/.agents/sessions/dynamic-cross-session-memory/PLAN.md deleted file mode 100644 index 4a4446c182..0000000000 --- a/.agents/sessions/dynamic-cross-session-memory/PLAN.md +++ /dev/null @@ -1,26 +0,0 @@ -# PLAN: Dynamic Cross-Session Memory - -## M1 — Persistence substrate + hydration (Phase 0+1) - -- [ ] M1-T1 Implement `task-memory-store.ts`: load (schema+checksum validated), persist (atomic), reconcile (hash-based, DI fs), move-rebinding from workspaceMoves. IDs: stable strings `ev-` style preserved. -- [ ] M1-T2 Thread through `initialSessionState` (options: persistentMemory, workspaceMoves; hydrate reconciled memory) and `runOnce` (post-run persist merged memory; derive moves from WorkspaceJournalService when instantiated). -- [ ] M1-T3 Unit/integration tests `sdk/src/__tests__/task-memory-store.test.ts` covering AC1–AC4. - -## M2 — Runtime gap fixes - -- [ ] M2-T1 Dedupe exact-duplicate strings in mergeTaskMemoryDraft list fields (respecting caps/order). -- [ ] M2-T2 compileTaskMemoryContext drops stale:true evidence from serialized block. -- [ ] M2-T3 Extend existing task-memory runtime tests for M2 behaviors. - -## M3 — Eval harness - -- [ ] M3-T1 `evals/memory-retention/` scenario runner + README: cold vs warm comparison, stale/rebound assertions (AC5/AC7 metrics), runnable via `bun test`. - -## M4 — Validation & docs - -- [ ] M4-T1 Typecheck: sdk, evals (+ common if touched indirectly). -- [ ] M4-T2 Run new tests + neighboring suites (initial-session-state, task-memory runtime tests). -- [ ] M4-T3 Update STATUS.md/LESSONS.md; note follow-on phases (consolidation, procedural library) as explicitly deferred. - -Dependencies: M1→M2→M3→M4 (M3 depends on exported store API). -Risks: run.ts touch-points are large/fragile — minimal surgical edits only; sdk standalone constraint forbids indexer/runtime imports in sdk. diff --git a/.agents/sessions/dynamic-cross-session-memory/SPEC.md b/.agents/sessions/dynamic-cross-session-memory/SPEC.md deleted file mode 100644 index d17307b92a..0000000000 --- a/.agents/sessions/dynamic-cross-session-memory/SPEC.md +++ /dev/null @@ -1,39 +0,0 @@ -# SPEC: Dynamic Cross-Session Memory - -## Goal - -Eliminate forced codebase re-analysis at every conversation start by persisting structured operational memory (TaskMemoryV1) per project root, reconciling its evidence against live file state at session start, and injecting only trustworthy (fresh or honestly-stale-marked) memory. - -## Non-goals (this iteration) - -- No background consolidation passes (sleep-time compute), no embedding/vector stores, no server. -- No changes to human-curated knowledge.md handling or memory-drift-guard. -- No auto-derived markdown injection into knowledge files. - -## Requirements - -1. R1 Persist committed task memory per project root at `/.openbuff/memory/task-memory.json`; schema-validated (zod TaskMemoryV1), checksum-checked on load, atomic tmp+rename writes, corrupt/missing file degrades silently to no memory. -2. R2 Hydrate at session start: `initialSessionState` loads + reconciles persisted memory into `mainAgentState.taskMemory` when cwd is provided. Option `persistentMemory: boolean` (default true) disables entirely; failures never break session start. -3. R3 Reconciliation (DI-injectable fs): for each evidence item with a path — missing file ⇒ `{stale: true}`; sha256(content) ≠ freshnessHash ⇒ `{stale: true}`; match ⇒ refresh `verifiedAt`. Never delete entries (drift-detection-over-trust). -4. R4 Move rebinding: optional `workspaceMoves: {from,to}[]` input (caller derives from WorkspaceStateV1 journal change records with action 'move'); a stale-missing evidence path matching a move source is rebound to the destination (path updated, staleness recomputed against destination hash) instead of being orphaned. -5. R5 Cross-run merge on save: persisted memory merges with this run's final memory via existing append-only `mergeTaskMemoryDraft` semantics (dedupe exact-duplicate list strings); bounded by existing schema array caps. -6. R6 Stale-aware compilation: `compileTaskMemoryContext` excludes `stale: true` evidence from the injected `` block (counts still visible), so unverified facts are surfaced-not-trusted. -7. R7 Eval harness (deterministic, no LLM): scripted multi-session scenario proving warm-start advantage and correct stale/rebound behavior; metrics: reconcile outcome counts, compiled-context chars cold vs warm, wrong-memory count (must be 0). - -## Acceptance criteria - -- AC1 Second `initialSessionState` on unchanged project yields hydrated taskMemory whose evidence verifies fresh. -- AC2 Mutating an evidenced file marks exactly that evidence stale; unrelated evidence stays fresh. -- AC3 Renaming (move) an evidenced file with `workspaceMoves` supplied rebinds the entry; without it, entry is stale (not deleted). -- AC4 Corrupt/truncated memory file is ignored (fresh start) without throwing. -- AC5 Compiled context contains zero stale-evidence text; contains fresh decisions text. -- AC6 All new + existing affected tests pass; sdk + evals typechecks pass. - -## Systems touched - -- sdk/src/services/task-memory-store.ts (new): load/persist/reconcile. -- sdk/src/run-state.ts: hydration + options threading. -- sdk/src/run.ts: persist final memory post-run; supply workspace moves from journal when available. -- packages/agent-runtime/src/util/task-memory.ts: dedupe in merge/normalize + stale-aware compile. -- sdk/src/**tests**/task-memory-store.test.ts (new), evals/memory-retention/ (new harness). -- Plan artifacts: .agents/sessions/dynamic-cross-session-memory/{PLAN,STATUS}.md diff --git a/.agents/sessions/dynamic-cross-session-memory/STATUS.md b/.agents/sessions/dynamic-cross-session-memory/STATUS.md deleted file mode 100644 index f10eb4aa5f..0000000000 --- a/.agents/sessions/dynamic-cross-session-memory/STATUS.md +++ /dev/null @@ -1,25 +0,0 @@ -# STATUS: Dynamic Cross-Session Memory - -Updated: 2026-08-22 (implementation complete, awaiting final review gate) - -## Completed - -- M1: `sdk/src/services/task-memory-store.ts` — load (schema+checksum validated, corrupt-safe), atomic save with unique tmp suffix + 0o600 perms, hash reconciliation with move rebinding, path-traversal fail-closed guard, Promise.all batched hashing, exported `stableHash` with FNV-1a vector tests. -- M1: hydration wired in `initialSessionState` (options `persistentMemory`, `workspaceMoves`; never blocks session start) and post-run persistence in `runOnce` success path only (`collectWorkspaceMoves` at module scope derives moves from the workspace journal). -- M2: runtime gap fixes verified already-present (`uniqueRecent`/`normalizeEvidence` dedupe; `evidenceIsFresh` excludes stale from ``); covered by new runtime tests. -- M3: deterministic eval harness `evals/memory-retention/` (S1–S4, no LLM). -- M4: typechecks green (script:typecheck all packages, typecheck-sdk, typecheck-agent-runtime); tests green: sdk store 11/11, runtime task-memory 8/8, eval scenarios 4/4. -- Review hardening round applied: safeParse never-throw, traversal test with spying fs, missing-file case, FNV vectors, 0o600 assertion, module-scope helper, batched hashing, symlink-exposure doc note. - -## Deferred (explicitly out of scope this iteration) - -- Background consolidation (sleep-time distillation), procedural recipe library, trust-scored retrieval ranking — see SPEC non-goals. -- Shared FNV-1a helper in @codebuff/common (three copies now annotated; consolidation pending). - -## Compatibility notes - -- Pre-checksum `task-memory.json` records: none exist in any deployment (the store shipped alongside checksum enforcement), so `loadPersistedTaskMemory` discards them fail-closed rather than tolerate-and-upgrade. Revisit only if a deployed legacy format ever materializes. - -## Blockers - -- None. Fresh reviewer pass pending on final snapshot. diff --git a/.agents/sessions/external-read-roots-2026-08/SPEC.md b/.agents/sessions/external-read-roots-2026-08/SPEC.md deleted file mode 100644 index 8ed85103eb..0000000000 --- a/.agents/sessions/external-read-roots-2026-08/SPEC.md +++ /dev/null @@ -1,151 +0,0 @@ -# SPEC — Read-only external read roots - -## Overview - -Openbuff tools are contained to the project root, plus a narrow exception for the -openbuff-owned OS temp namespace (`owned-temp`). Users cannot read files outside -the project — including their own openbuff config directory (logs, harness -state) — even though those are the user's own files and reading them is a -legitimate, non-mutating action. - -This work adds a third containment scope, `external-read`: a **read-only**, -**default-closed**, **configure-once** allowlist of roots outside the project -that path-taking READ tools may reach. The write path is structurally excluded. - -Wave 1 (the `common` primitive) is COMPLETE, gate-approved, and committed. -Wave 2 (product wiring) is implemented in the working tree but NOT yet green. - -## Goals - -- A user can read files in their openbuff config dir (e.g. logs, harness state) - through `read_files` / `read_logs` / `read_image` / `list_directory`. -- A user can allowlist additional absolute roots via `openbuff.json` - → `readableRoots: string[]`. -- Mandatory-sensitive files stay blocked **inside** allowlisted roots - (`credentials.json`, `.env`, private keys, kubeconfig, tfstate, …). -- Writes can never reach an allowlisted root, enforced structurally rather than - by handler discipline. -- Default posture is closed: with nothing configured, behavior is byte-identical - to before this work. - -## Non-goals - -- `glob` stays project-only. It is a pattern-driven directory walk; widening it - would let a pattern enumerate an allowlisted root. Deliberately out of scope. -- No write, move, delete, or `cwd` access to external roots, ever. -- No relative entries in `readableRoots` (ambiguous in a global config file; - dropped rather than guessed). -- No mid-session reconfiguration. Changing `readableRoots` requires a restart. -- No UI/slash-command surface for managing roots in this iteration. - -## Requirements - -### R1 — Containment primitive (`common`) — DONE, committed - -- `ContainedProjectPath['scope']` is `'project' | 'owned-temp' | 'external-read'`. -- `configureExternalReadRoots(roots)` normalizes/dedupes/sorts; skips filesystem - roots and `..` entries; idempotent for an equivalent set; THROWS on a - differing set. -- `ensureExternalReadRootsConfigured(roots)` — non-throwing wrapper returning - `'configured' | 'unchanged' | 'refused-changed'`. -- `getExternalReadRoots()`, `resetExternalReadRootsForTesting()`, - `isExternalReadPath(input)`. -- `resolveProjectPathForRead` / `resolveProjectPathForFileSystemRead` — the ONLY - entry points that can produce `external-read`. They delegate to the existing - write resolvers first and only fall back to the external branch on `null`. -- `resolveProjectPath` / `resolveProjectPathForFileSystem` are UNCHANGED. -- The external resolver refuses mandatory-sensitive basenames on BOTH the - lexical and dereferenced path, so consumers inherit the refusal fail-closed. -- `credentials.json` / `.yaml` / `.yml` added to `SENSITIVE_BASENAMES`. - -### R2 — Config schema (`sdk/src/provider-config.ts`) — implemented, unverified - -- `readableRoots: z.array(z.string().min(1)).default([])`. -- `.default([])` is intentional: downstream code never handles `undefined`. - Consequence: `readableRoots: string[]` is REQUIRED in the schema OUTPUT type - and therefore in `LoadedProviderConfig['config']`. This is what breaks the - hand-built test fixtures (see PLAN task T1). - -### R3 — Read-only operation resolvers (`sdk/src/tools/path-utils.ts`) — implemented - -- `resolveFilePathForReadOperation`, - `resolveFilePathForFileSystemReadOperation`. -- Follow-symlink shape only; no `followFinalSymlink: false` (that option exists - for unlink-style mutations). -- The existing `resolveFilePathFor*Operation` write twins are unchanged. - -### R4 — Read handlers rewired — implemented - -`read-files.ts` (`authorizeReadTarget`), `read-logs.ts`, `read-image.ts`, -`list-directory.ts`. Plus: `read-files.ts` extends its owned-temp fileFilter -alias block to cover `external-read` (both scopes carry an ABSOLUTE -`relativePath`, so a host filter written against project-relative globs would -silently fail OPEN). Alias key: `external-read/`. - -### R5 — Run-start configuration (`sdk/src/run.ts`) — implemented - -In `runOnce`, before tool dispatch: `ensureExternalReadRootsConfigured([...])` -with the openbuff config dir (`getConfigDir(env)`) plus absolute-only entries -from `loadProviderConfigSync().config.readableRoots`. Wrapped in try/catch so a -malformed config cannot block a run. `'refused-changed'` logs a warn naming -counts, not paths (home-dir paths are mildly sensitive and logs get shared). - -### R6 — Agent-runtime read backstop (`tool-executor.ts`) — implemented - -The pre-dispatch scope check's `ownedTempRead` condition extended so an -allowlisted external READ is not hard-blocked. Writes there still hard-block. - -### R7 — Fixtures + validation — BLOCKED (this is the remaining work) - -Six typecheck failures, all "missing required `readableRoots`" in hand-built -`LoadedProviderConfig` fixtures. See PLAN task T1 for exact sites. - -### R8 — Security review — NOT STARTED - -`security-reviewer` on the permission-boundary widening. - -### R9 — Documentation — NOT STARTED - -`docs/configuration.md` + `openbuff.json.example`. - -## Acceptance criteria - -- AC1 — `bun run typecheck` exits 0 for all 11 workspace packages. -- AC2 — `bun test sdk/src/__tests__/` reports 0 fail AND 0 error. -- AC3 — `bun test common/src/util/__tests__/` and the agent-runtime tool tests - report 0 fail. -- AC4 — Test named for the write-path invariant still passes: `resolveProjectPath` - returns `null` for a path inside a configured allowlisted root. -- AC5 — `credentials.json` inside an allowlisted root is refused by BOTH - `isExternalReadPath` and the resolver. -- AC6 — With `readableRoots` unset and no config dir on the allowlist, external - paths are refused exactly as before (default-closed). -- AC7 — Full-directory SDK test run is order-independent: the read-logs external - suite passes both alone and after a suite that exercises the SDK run path. -- AC8 — `security-reviewer` returns no unresolved BLOCKING finding. -- AC9 — Docs state: reads only, absolute-only entries, sensitive files still - blocked, restart required to apply changes. - -## Relevant files - -Committed (wave 1): -- `common/src/util/project-path-containment.ts` -- `common/src/util/sensitive-paths.ts` -- `common/src/util/__tests__/{project-path-containment,sensitive-paths}.test.ts` -- `sdk/src/tools/filesystem-authority.ts` — `toAuthorizedPath` fails closed on - `external-read` with code `external_read_scope_unsupported` - -Dirty (wave 2, in working tree): -- `sdk/src/provider-config.ts`, `sdk/src/run.ts` -- `sdk/src/tools/{path-utils,read-files,read-logs,read-image,list-directory}.ts` -- `packages/agent-runtime/src/tools/tool-executor.ts` -- `sdk/src/__tests__/{path-utils,read-files,read-logs,model-provider}.test.ts` -- `packages/agent-runtime/src/__tests__/run-agent-step-tools.test.ts` -- `common/src/util/project-path-containment.ts` (+ its test) — `ensureExternalReadRootsConfigured` - -To touch in T1: -- `sdk/src/__tests__/model-provider.test.ts` -- `sdk/src/impl/__tests__/failover.test.ts` - -To touch in T4: -- `docs/configuration.md`, `openbuff.json.example` diff --git a/.agents/sessions/harness-cohesion-audit-2026-07/EVENTS.jsonl b/.agents/sessions/harness-cohesion-audit-2026-07/EVENTS.jsonl deleted file mode 100644 index d1c53e2c12..0000000000 --- a/.agents/sessions/harness-cohesion-audit-2026-07/EVENTS.jsonl +++ /dev/null @@ -1,53 +0,0 @@ -{"ts":"2026-07-21T20:48:18.666Z","kind":"append_lesson","summary":"Appended entry \"Confirmed Decisions (2026-07)\" to STATUS.md","payload":{"heading":"Confirmed Decisions (2026-07)","artifact":"STATUS.md"}} -{"ts":"2026-07-21T20:48:18.666Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-21T20:50:21.223Z","kind":"append_lesson","summary":"Appended entry \"M1.1 Roster Inventory Matrix\" to STATUS.md","payload":{"heading":"M1.1 Roster Inventory Matrix","artifact":"STATUS.md"}} -{"ts":"2026-07-21T20:53:10.549Z","kind":"task_update","summary":"Updated 1 task line(s): M1.1","payload":{"matched":["M1.1"]}} -{"ts":"2026-07-21T20:53:20.550Z","kind":"task_update","summary":"Updated 1 task line(s): M1.2","payload":{"matched":["M1.2"]}} -{"ts":"2026-07-21T20:53:20.550Z","kind":"current_task","summary":"Current task -> \"M1.2\"","payload":{"currentTask":"M1.2"}} -{"ts":"2026-07-21T20:59:42.088Z","kind":"task_update","summary":"Updated 2 task line(s): M1.2, M1.3","payload":{"matched":["M1.2","M1.3"]}} -{"ts":"2026-07-21T20:59:42.088Z","kind":"current_task","summary":"Current task -> \"M1.3\"","payload":{"currentTask":"M1.3"}} -{"ts":"2026-07-21T21:02:17.853Z","kind":"task_update","summary":"Updated 2 task line(s): M1.3, M1.4","payload":{"matched":["M1.3","M1.4"]}} -{"ts":"2026-07-21T21:02:17.853Z","kind":"current_task","summary":"Current task -> \"M1.4\"","payload":{"currentTask":"M1.4"}} -{"ts":"2026-07-21T21:05:19.252Z","kind":"task_update","summary":"Updated 2 task line(s): M1.4, M1.5","payload":{"matched":["M1.4","M1.5"]}} -{"ts":"2026-07-21T21:05:19.252Z","kind":"current_task","summary":"Current task -> \"M1.5\"","payload":{"currentTask":"M1.5"}} -{"ts":"2026-07-21T21:06:15.318Z","kind":"task_update","summary":"Updated 2 task line(s): M1.5, M2.1","payload":{"matched":["M1.5","M2.1"]}} -{"ts":"2026-07-21T21:06:15.318Z","kind":"current_task","summary":"Current task -> \"M2.1\"","payload":{"currentTask":"M2.1"}} -{"ts":"2026-07-21T21:06:44.764Z","kind":"append_lesson","summary":"Appended entry \"M1 Complete (2026-07)\" to STATUS.md","payload":{"heading":"M1 Complete (2026-07)","artifact":"STATUS.md"}} -{"ts":"2026-07-21T21:09:19.791Z","kind":"task_update","summary":"Updated 2 task line(s): M2.1, M2.2","payload":{"matched":["M2.1","M2.2"]}} -{"ts":"2026-07-21T21:09:19.791Z","kind":"current_task","summary":"Current task -> \"M2.2\"","payload":{"currentTask":"M2.2"}} -{"ts":"2026-07-21T21:10:12.840Z","kind":"task_update","summary":"Updated 2 task line(s): M2.2, M2.3","payload":{"matched":["M2.2","M2.3"]}} -{"ts":"2026-07-21T21:10:12.840Z","kind":"current_task","summary":"Current task -> \"M2.3\"","payload":{"currentTask":"M2.3"}} -{"ts":"2026-07-21T21:12:14.767Z","kind":"task_update","summary":"Updated 2 task line(s): M2.3, M3.1","payload":{"matched":["M2.3","M3.1"]}} -{"ts":"2026-07-21T21:12:14.767Z","kind":"current_task","summary":"Current task -> \"M3.1\"","payload":{"currentTask":"M3.1"}} -{"ts":"2026-07-21T21:14:12.501Z","kind":"task_update","summary":"Updated 2 task line(s): M3.1, M3.3","payload":{"matched":["M3.1","M3.3"]}} -{"ts":"2026-07-21T21:14:12.501Z","kind":"current_task","summary":"Current task -> \"M3.3\"","payload":{"currentTask":"M3.3"}} -{"ts":"2026-07-21T21:16:17.295Z","kind":"task_update","summary":"Updated 3 task line(s): M3.3, M3.2, M3.4","payload":{"matched":["M3.3","M3.2","M3.4"]}} -{"ts":"2026-07-21T21:16:17.295Z","kind":"current_task","summary":"Current task -> \"M3.4\"","payload":{"currentTask":"M3.4"}} -{"ts":"2026-07-21T21:17:10.256Z","kind":"append_lesson","summary":"Appended entry \"M3.4 Decision: gateAwarenessSection gating rule\" to LESSONS.md","payload":{"heading":"M3.4 Decision: gateAwarenessSection gating rule","artifact":"LESSONS.md"}} -{"ts":"2026-07-21T21:17:57.842Z","kind":"task_update","summary":"Updated 2 task line(s): M3.4, M4.1","payload":{"matched":["M3.4","M4.1"]}} -{"ts":"2026-07-21T21:17:57.842Z","kind":"current_task","summary":"Current task -> \"M4.1\"","payload":{"currentTask":"M4.1"}} -{"ts":"2026-07-21T21:20:06.175Z","kind":"task_update","summary":"Updated 2 task line(s): M4.1, M4.2","payload":{"matched":["M4.1","M4.2"]}} -{"ts":"2026-07-21T21:20:06.175Z","kind":"current_task","summary":"Current task -> \"M4.2\"","payload":{"currentTask":"M4.2"}} -{"ts":"2026-07-21T21:22:54.380Z","kind":"task_update","summary":"Updated 2 task line(s): M4.2, M5.1","payload":{"matched":["M4.2","M5.1"]}} -{"ts":"2026-07-21T21:22:54.380Z","kind":"current_task","summary":"Current task -> \"M5.1\"","payload":{"currentTask":"M5.1"}} -{"ts":"2026-07-21T21:24:57.924Z","kind":"task_update","summary":"Updated 2 task line(s): M5.1, M5.2","payload":{"matched":["M5.1","M5.2"]}} -{"ts":"2026-07-21T21:24:57.924Z","kind":"current_task","summary":"Current task -> \"M5.2\"","payload":{"currentTask":"M5.2"}} -{"ts":"2026-07-21T21:26:26.377Z","kind":"append_lesson","summary":"Appended entry \"M5.2 BLOCKED — conflict with compatibility invariants\" to STATUS.md","payload":{"heading":"M5.2 BLOCKED — conflict with compatibility invariants","artifact":"STATUS.md"}} -{"ts":"2026-07-21T21:26:44.622Z","kind":"task_update","summary":"Updated 2 task line(s): M5.2, M5.3","payload":{"matched":["M5.2","M5.3"]}} -{"ts":"2026-07-21T21:26:44.622Z","kind":"current_task","summary":"Current task -> \"M5.3\"","payload":{"currentTask":"M5.3"}} -{"ts":"2026-07-21T21:31:15.966Z","kind":"task_update","summary":"Updated 2 task line(s): M5.3, M6.1","payload":{"matched":["M5.3","M6.1"]}} -{"ts":"2026-07-21T21:31:15.966Z","kind":"current_task","summary":"Current task -> \"M6.1\"","payload":{"currentTask":"M6.1"}} -{"ts":"2026-07-21T21:34:04.802Z","kind":"task_update","summary":"Updated 2 task line(s): M6.1, M6.2","payload":{"matched":["M6.1","M6.2"]}} -{"ts":"2026-07-21T21:34:04.802Z","kind":"current_task","summary":"Current task -> \"M6.2\"","payload":{"currentTask":"M6.2"}} -{"ts":"2026-07-21T21:35:09.654Z","kind":"task_update","summary":"Updated 2 task line(s): M6.2, M7.1","payload":{"matched":["M6.2","M7.1"]}} -{"ts":"2026-07-21T21:35:09.654Z","kind":"current_task","summary":"Current task -> \"M7.1\"","payload":{"currentTask":"M7.1"}} -{"ts":"2026-07-21T21:37:38.329Z","kind":"task_update","summary":"Updated 2 task line(s): M7.1, M5.2","payload":{"matched":["M7.1","M5.2"]}} -{"ts":"2026-07-21T21:37:38.329Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-21T21:38:02.119Z","kind":"append_lesson","summary":"Appended entry \"Execution Complete (2026-07)\" to STATUS.md","payload":{"heading":"Execution Complete (2026-07)","artifact":"STATUS.md"}} -{"ts":"2026-07-21T21:38:02.119Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-21T21:49:12.918Z","kind":"task_update","summary":"Updated 1 task line(s): M3.2","payload":{"matched":["M3.2"]}} -{"ts":"2026-07-21T21:49:48.771Z","kind":"append_lesson","summary":"Appended entry \"M3.2 Blocker Resolved + Non-Blocking Cleanup (2026-07)\" to STATUS.md","payload":{"heading":"M3.2 Blocker Resolved + Non-Blocking Cleanup (2026-07)","artifact":"STATUS.md"}} -{"ts":"2026-07-22T03:57:28.377Z","kind":"append_lesson","summary":"Appended entry \"Reviewer read-budget fix (2026-07)\" to LESSONS.md","payload":{"heading":"Reviewer read-budget fix (2026-07)","artifact":"LESSONS.md"}} -{"ts":"2026-07-22T03:58:01.891Z","kind":"append_lesson","summary":"Appended entry \"Reviewer Read-Budget Fix Complete (2026-07)\" to STATUS.md","payload":{"heading":"Reviewer Read-Budget Fix Complete (2026-07)","artifact":"STATUS.md"}} -{"ts":"2026-08-02T21:32:41.648Z","kind":"append_lesson","summary":"Appended entry \"M5.2 RESOLVED — read_slices fully removed from live surface\" to STATUS.md","payload":{"heading":"M5.2 RESOLVED — read_slices fully removed from live surface","artifact":"STATUS.md"}} -{"ts":"2026-08-02T21:32:41.649Z","kind":"session_status","summary":"Session status -> completed","payload":{"status":"completed"}} diff --git a/.agents/sessions/harness-cohesion-audit-2026-07/LESSONS.md b/.agents/sessions/harness-cohesion-audit-2026-07/LESSONS.md deleted file mode 100644 index 04976cc41f..0000000000 --- a/.agents/sessions/harness-cohesion-audit-2026-07/LESSONS.md +++ /dev/null @@ -1,34 +0,0 @@ -# LESSONS — Harness Cohesion Audit - -## Gotchas discovered during the audit - -- **`handleSteps` is serialized (`.toString()` → `new Function`).** Reconstructed generators lose module closure, so gate helpers cannot be imported at runtime. **Do not hand-edit the inline region.** Edit the pure modules in `scripts/generate-gate-helpers.ts` `SOURCE_MODULES` (`gate-paths.ts`, `gate-reviewer.ts`, `gate-repair.ts`, `gate-concurrency.ts`, `gate-fingerprint.ts`), then regenerate into the `` block of `agents/base2/base2.ts` via `bun run scripts/generate-gate-helpers.ts --write agents/base2/base2.ts` (or `prebuild:agents`). Freshness is enforced by `agents/__tests__/gate-helpers-freshness.test.ts`; parity matrices still cover paths/reviewer/repair/concurrency/fingerprint. - -- **Prompt-section snapshot freeze.** `quality-prompt-snapshot.test.ts` byte-freezes `qualitySection` and shared prompt text. Prompt edits (M2/M3) require an intentional snapshot update; never blind-accept the new snapshot. - -- **Generated artifacts.** `cli/src/agents/bundled-agents.generated.ts` and `agents/types/tools.ts` are generated (`prebuild-agents.ts`, tool-def generator). Edit source + regenerate; never hand-edit. - -- **Five parallel rosters, no single source of truth.** agent `.ts` default exports (de-facto truth) → generated bundle → `routes.json` (drifted, 12 dead ids) → `AGENT_PERSONAS`/`AGENT_IDS` (heavily drifted) → `spawnableAgents` arrays. This is the root cause of the "lack of cohesion" the user felt. - -- **Two orchestration mechanisms coexist.** The authoritative flow is the base2 `handleSteps` gate + `spawn-agent-utils` capability clamping. `packages/agent-runtime/src/orchestration/*` (esp. `workflow-engine`) is invoked but advisory/telemetry-only — a dual-system smell that can mislead future readers. - -- **Prompt/capability mismatch is real cohesion debt.** `buildBroadAuditSection` instructs the coordinator to collect `structuralReceipt`s from file-picker/code-searcher shards, but only `general-agent` + `write_audit_findings` emit those receipts. The produce-path and the prompted spawn-path don't line up — a concrete example of features added without wiring them into coordination. - -- **Capability clamps are layered (good).** `deriveSpawnTemplateCapabilities` clamps at spawn time AND `executeToolCall` re-enforces filesystem/tool scope at execution — plan-only propagation forces child terminal profiles to read-only. Preserve both layers when editing spawn logic. - -## Decisions (record as confirmed) - -- (confirmed) persona-map strategy: reconcile + guard against the canonical `agents/**/*.ts` roster (not a full generate-from-source refactor this pass). See STATUS.md "Confirmed Decisions (2026-07)". -- (confirmed) orchestration/ fate: keep `orchestration/workflow-engine`, marked advisory/telemetry-only. See STATUS.md "Execution Complete (2026-07)". - - - -## M3.4 Decision: gateAwarenessSection gating rule — 2026-07-21T21:17:10.256Z - -Documented rule: gateAwarenessSection is present in an orchestrator's system prompt IFF that orchestrator runs the validation/reviewer gate. base2 encodes this as `isDefault ? gateAwarenessSection : ''` (fast modes are non-default and skip the gate, so they correctly omit it). base-deep composes createBase2('default') and always runs the gate, so its hand-written system prompt interpolates gateAwarenessSection unconditionally — which is equivalent to the isDefault rule because base-deep is always default+gated. No behavioral divergence exists; the two sites obey one rule. Kept as-is (no code change) rather than forcing base-deep through the isDefault ternary, since base-deep never runs a non-default/fast mode. - - - -## Reviewer read-budget fix (2026-07) — 2026-07-22T03:57:28.376Z - -Root cause of the reviewer-gate attestation failures: MAX_RENDER_CHARS = 100_000 in sdk/src/tools/read-files.ts truncated any whole-file read over 100k chars. base2.ts is 313 KB, so reviewers (code-reviewer, security-reviewer, all 14 specialists — all read exclusively via read_files) received a FILE_TOO_LARGE stub and physically could not attest to that pending file, producing 'reviewer did not attest to every pending file: agents/base2/base2.ts'. Fix (user-directed): set MAX_RENDER_CHARS = MAX_FILE_BYTES so the existing 10MB byte gate is the single effective read ceiling; sub-10MB source files now render fully for both whole-file and range reads. MAX_FILE_BYTES (10MB) and MAX_RANGE_READ_BYTES unchanged as the remaining safety valve. Coupled tests in sdk/src/**tests**/read-files.test.ts updated; SDK read-files suite 62 pass / 0 fail, SDK typecheck clean. Tradeoff accepted by user: a large (<10MB) generated file can now dump fully into a reader's context; the 10MB byte gate remains the hard cap. diff --git a/.agents/sessions/harness-cohesion-audit-2026-07/PLAN.md b/.agents/sessions/harness-cohesion-audit-2026-07/PLAN.md deleted file mode 100644 index 7ffd078b4c..0000000000 --- a/.agents/sessions/harness-cohesion-audit-2026-07/PLAN.md +++ /dev/null @@ -1,129 +0,0 @@ -# PLAN — Harness Cohesion Remediation - - - -Milestones are ordered by risk/leverage. Each executable task has a stable ID, dependencies, acceptance, and validation. Do the roster + prompt-mismatch milestones first (highest cohesion payoff, lowest blast radius), gate/tool hygiene next, orchestration-duality decision last. - -Validation routing (per AGENTS.md path→suite map): - -- `agents/*` → agents typecheck + relevant `agents/__tests__/*` and e2e -- `common/*` → common checks + dependent package typechecks -- `packages/agent-runtime/*` → runtime typecheck/tests -- `cli/*` → CLI typecheck (+ visual smoke if components) - ---- - -## M1 — Single source of truth for the agent roster - -- [x] M1.1 Inventory + classify every id referenced across the 5 rosters (Inventory matrix recorded in STATUS.md; analysis-only task.) - - Acceptance: a checked-in matrix (this session) listing each id × {shipped-bundled, spawnable, routed, persona, external-cli-allowlisted, dead} - - Validate: n/a (analysis) -- [x] M1.2 Reconcile `common/src/constants/agents.ts` (`AGENT_PERSONAS`/`AGENT_IDS`) (Claimed: reconcile AGENT_PERSONAS/AGENT_IDS against bundled roster.) - - Depends on: M1.1 - - Acceptance: remove non-shipped ids (`ask`, `planner`, `agent-builder`, `reviewer`, `file-explorer`, `researcher`); add missing shipped ids OR derive the map from bundled agents; no runtime consumer breaks - - Validate: `bun test` common + `cli` typecheck -- [x] M1.3 Prune dead ids from `openbuff.d.example/routes.json` or allowlist external CLI agents explicitly (routes.json pruned to shipped/bundled + external-CLI/eval allowlist; 0 dead ids verified.) - - Depends on: M1.1 - - Acceptance: every routes.json id is shipped/bundled or in a documented external-CLI allowlist - - Validate: new guard test (M1.4) -- [x] M1.4 Add a roster-drift guard test (Building roster-drift guard test.) (roster-drift guard test created + green (4 pass). Validates the whole M1 milestone.) - - Depends on: M1.2, M1.3 - - Acceptance: test fails if a persona/routes/spawnable id is neither bundled nor allowlisted, AND if a bundled non-root agent is unreachable by any orchestrator or pattern - - Validate: `bun test agents/__tests__/` (new test file) -- [x] M1.5 Resolve `directory-lister` / `glob-matcher` reachability (directory-lister/glob-matcher reachability.) (Decision recorded: directory-lister/glob-matcher stay bundled, intentionally excluded from orchestrator spawnability; guard encodes intentionallyExcluded. M1 milestone complete.) - - Depends on: M1.4 - - Acceptance: either added to an orchestrator/pattern spawnable path or removed from bundling; guard from M1.4 passes - - Validate: `bun test agents/__tests__/` - -## M2 — Prompt ↔ capability alignment (core cohesion fix) - -- [x] M2.1 Fix `buildBroadAuditSection` shard→receipt path (Claiming: fix buildBroadAuditSection shard->receipt path.) (buildBroadAuditSection produce/consume path fixed; snapshot + base2 tests green.) - - Acceptance: the section routes audit/reasoning shards that must emit `structuralReceipt` to `general-agent` + `write_audit_findings` (not file-picker/code-searcher); discovery-only shards are named as discovery-only; produce-path and consume-path (`evaluate_audit_coverage`) connect - - Validate: `bun test agents/__tests__/quality-prompt-snapshot.test.ts` (update snapshot intentionally) + `agents/__tests__/base2.test.ts` -- [x] M2.2 Surface the durable-findings / synthesizer flow in orchestrator guidance (Surface durable-findings/synthesizer flow in orchestrator guidance + docs.) (Doc and prompt agree on general-agent + write_audit_findings + synthesizer audit flow. No doc edit needed (docs already correct at lines 24, 84, 951-962).) - - Depends on: M2.1 - - Acceptance: `buildBroadAuditSection` (or a sibling) tells the coordinator to spawn `general-agent` audit shards with `sessionSlug`/`shardId`/`snapshotId` and to reduce via `synthesizer`; `docs/agents-and-tools.md` matches - - Validate: snapshot test + doc read-back -- [x] M2.3 Remove the production-dead `frontendSection` re-export (Claiming: remove production-dead frontendSection re-export or rewire snapshot test to canonical prompt-sections.ts.) (frontendSection dead re-export removed; test rewired to canonical source; snapshot + typecheck green.) - - Acceptance: `quality-prompt-section.ts:77` re-export removed OR snapshot test rewired to import from the canonical `prompt-sections.ts`; production path unchanged - - Validate: `bun test agents/__tests__/quality-prompt-snapshot.test.ts` - -## M3 — Orchestrator family consistency - -- [x] M3.1 Reconcile `base-deep` spawnable list with `base2` (Claiming: reconcile base-deep spawnable list with base2.) (base-deep override removed; inherits base2 computed list. Also resolves the base-deep half of M3.2.) (Verified: browser-use is unconditional across all modes incl. fast; per-mode deltas are only the coded planOnly/isDefault/isFast gates, now guarded by roster-drift + specialists tests. No code change needed.) (Documented intentional per-mode spawnable deltas in base2.ts + test-asserted delta block in roster-drift.test.ts. Reviewer blocker RF-1/RF-2 resolved. roster-drift 7 pass, typechecks green.) - - Depends on: M1.4 - - Acceptance: base-deep no longer silently drops `context-pruner`/`tmux-cli` (or the deltas are intentional + documented + test-asserted); prefer generating base-deep's list from the same computed source as base2 - - Validate: `bun test agents/__tests__/base2.test.ts` + new consistency assertion -- [x] M3.2 Align `base2-fast` spawnable set (browser-use) with the other modes (Documented intentional per-mode deltas in base2.ts and froze them with a roster-drift assertion.) (browser-use is unconditional across default/fast/plan/execute-plan; the only fast delta vs default is the default-only editor family + thinker, and the only plan delta is the implementation-only mutation agents. Documented inline in `agents/base2/base2.ts` above `spawnableAgents` and test-asserted by the new `intentional per-mode spawnable deltas (M3.2)` block in `agents/__tests__/roster-drift.test.ts`.) - - Depends on: M3.1 - - Acceptance: intentional per-mode deltas only; documented - - Validate: `bun test agents/__tests__/` -- [x] M3.3 Share editor-handoff guidance between DEFAULT and EXECUTE_PLAN step prompts (Share editor-handoff guidance between DEFAULT and EXECUTE_PLAN step prompts.) (buildExecutePlanStepPrompt composes buildImplementationStepPrompt; EXECUTE_PLAN now carries editor-handoff guidance. Gate green.) - - Acceptance: `buildExecutePlanStepPrompt` carries the same editor-handoff / "don't manually spawn code-reviewer" guidance as `buildImplementationStepPrompt`; PLAN builder composes from shared builders instead of reimplementing - - Validate: snapshot + `agents/__tests__/base2.test.ts` -- [x] M3.4 Make `gateAwarenessSection` gating consistent across base2/base-deep (Claiming: make gateAwarenessSection gating consistent across base2/base-deep.) (gateAwarenessSection gating rule documented; equivalent behavior confirmed, no code change needed.) - - Depends on: M3.1 - - Acceptance: one documented rule for when the section is included (both conditional or both unconditional with justification) - - Validate: snapshot test - -## M4 — Gate / reviewer drift guards - -- [x] M4.1 Add aux-path parity test for `gate-paths.ts` helpers (Claiming: add aux-path parity test for gate-paths.ts helpers.) (gate-paths parity test added and green (3 pass).) - - Acceptance: `gate-aux-triggers.test.ts` (or new file) asserts inline `normalizeGateFilePath`/`normalizeGateFileList`/`gateFileSetsEqual` equal the `gate-paths.ts` exports, matching the existing gate-repair/gate-reviewer parity pattern - - Validate: `bun test agents/__tests__/gate-aux-triggers.test.ts` -- [x] M4.2 Single frozen source for the security-sensitive glob list (Single frozen source + parity test for the security-sensitive glob list.) (security-glob-parity guard green (4 pass).) - - Acceptance: inline gate predicate and `securityReviewSection` derive from / are parity-tested against one list - - Validate: `bun test agents/__tests__/` - -## M5 — Tool registry hygiene - -- [x] M5.1 Resolve the 4 dead tools (`lookup_agent_info`, `render_ui`, `find_files`, `find_files_matching_content`) (Resolve 4 dead tools: lookup_agent_info, render_ui, find_files, find_files_matching_content.) (4 dead tools quarantined; gate green.) - - Acceptance: each is either granted to an agent that should have it, or set non-promptVisible/quarantined, with a rationale; no dangling promptVisible-but-ungranted tools - - Validate: `bun test agents/tool-reachability.test.ts` + `common/src/tools/__tests__/` -- [/] M5.2 Remove `read_slices` from published/generated type surface (Remove read_slices from publishedTools + regenerated agents/types/tools.ts.) (Conflicts with test-encoded compatibility invariants; read_slices already prompt-invisible via quarantine. Awaiting user decision — see STATUS.) (CANCELLED (resolved-by-quarantine). Removing read_slices from publishedTools + generated types would break two deliberate compatibility invariants (quarantined-tools-stay-published; generated-types-include-every-published-tool) and change the external custom-agent type contract. read_slices is already prompt-invisible via M5.1 quarantine metadata, which achieves the real goal. Reopen only if intentionally breaking the published/type compatibility contract.) - - Acceptance: `read_slices` out of `publishedTools` and regenerated `agents/types/tools.ts`; quarantine metadata unchanged - - Validate: `bun test common/src/tools/__tests__/` + regenerate tool defs -- [x] M5.3 Broaden `tool-reachability.test.ts` to enumerate all structured-output agents (Claiming: broaden tool-reachability set_output coverage to all structured-output agents.) (Broadened + fixed structured-output guard; 11 pass.) - - Depends on: M5.1 - - Acceptance: test covers every structured-output agent's auto-injected `set_output` - - Validate: `bun test agents/tool-reachability.test.ts` - -## M6 — Discovery agent boundaries - -- [x] M6.1 Give `basher` an explicit terminal permission profile (Give basher an explicit terminalPermissionProfile.) (basher terminalPermissionProfile made explicit (workspace-write); gate green.) - - Acceptance: basher declares a profile consistent with debugger/git-committer/etc.; no capability regression - - Validate: `bun test agents/__tests__/basher.test.ts` + spawn-permissions runtime tests -- [x] M6.2 Clarify `file-picker` / `file-lister` boundary (Clarify file-picker/file-lister boundary.) (file-lister documented as file-picker internal worker; tests green.) - - Acceptance: file-lister documented as file-picker's internal worker (or merged); no orphan spawnerPrompt confusion - - Validate: `bun test agents/__tests__/file-picker.test.ts` + `file-lister.test.ts` - -## M7 — Orchestration subsystem decision - -- [x] M7.1 Decide fate of `packages/agent-runtime/src/orchestration/*` (Deciding fate of packages/agent-runtime/src/orchestration/\*.) (orchestration/ kept advisory/telemetry with doc marker in workflow-engine.ts; runtime typecheck green.) - - Acceptance: a documented decision (keep-as-telemetry / remove `workflow-engine` / promote to authoritative); if kept, a doc note marks it advisory so it is not mistaken for the authoritative gate - - Validate: runtime typecheck/tests if code changes; else doc-only - ---- - -## Risks / Blockers / Open Questions - -- **Snapshot churn:** `quality-prompt-snapshot.test.ts` byte-freezes shared prompt text; M2/M3 edits require deliberate snapshot updates. Do not blind-update — confirm the diff is the intended change. -- **Generated artifacts:** `bundled-agents.generated.ts` and `agents/types/tools.ts` are generated; edit the source + regenerate, never hand-edit. -- **Open question (needs user):** Should `AGENT_PERSONAS`/`AGENT_IDS` be fully derived from bundled agents (bigger refactor) or reconciled + drift-guarded (smaller)? Default assumption: reconcile + guard. -- **Open question (needs user):** For `orchestration/`, is `workflow-engine` intended future work or removable? Default assumption: keep, mark advisory. -- **base-deep list generation** may reduce flexibility if some deltas are intentional; confirm intended deltas before collapsing to a generated list. - -## Validation Gates (per milestone) - -- M1: new roster-drift guard test green; common + cli typecheck green. -- M2/M3: intentional snapshot update + `agents/__tests__/base2.test.ts` green. -- M4: new parity tests green. -- M5: tool-reachability + tool metadata tests green; tool defs regenerated. -- M6: agent + spawn-permission tests green. -- M7: decision recorded; typecheck green if code touched. - -## Checkpoint / Update Rules - -- Update STATUS.md via `update_plan_status` at each task start/finish, blocker discovery/resolution, and validation result. -- Append to LESSONS.md via `update_plan_status` whenever a snapshot/generated-artifact gotcha or an intentional-delta decision is confirmed. -- Use `create_plan` only to rewrite SPEC.md/PLAN.md if scope materially changes. diff --git a/.agents/sessions/harness-cohesion-audit-2026-07/SPEC.md b/.agents/sessions/harness-cohesion-audit-2026-07/SPEC.md deleted file mode 100644 index 50a6f2058e..0000000000 --- a/.agents/sessions/harness-cohesion-audit-2026-07/SPEC.md +++ /dev/null @@ -1,92 +0,0 @@ -# SPEC — Agent / Tool / Reviewer Harness Cohesion - -## Overview - -The Openbuff harness has accumulated a large, capable set of agents (46 shipped), 61 registered tools, 14 specialists, 3 aux gates + a validation/reviewer gate, and a multi-mode orchestrator family (`base2` / `base2-plan` / `base2-execute-plan` / `base-deep` / fast variants). A recent burst of feature additions left the pieces individually strong but weakly coordinated: multiple hand-maintained rosters drift against each other, some prompt guidance tells the coordinator to do things the tools/gates don't actually support (and omits capabilities that do exist), and gate logic is duplicated inline without full parity guards. - -> **Historical note (post-remediation):** The five handleSteps gate helpers in `scripts/generate-gate-helpers.ts` `SOURCE_MODULES` (`gate-paths`, `gate-reviewer`, `gate-repair`, `gate-concurrency`, `gate-fingerprint`) are now **generator-synced** into `` (freshness + parity tests). The overview line about "duplicated inline without full parity guards" and the goal to "close drift risk in the inline-mirrored gate logic" describe the **audit-time** debt; fingerprint is no longer a residual dual-copy. - -This spec captures a source-backed audit and a remediation plan to restore cohesion so the orchestrator fully and correctly uses every part of the harness. - -## Goals - -- Establish a single (or generated) source of truth for the agent roster so spawnable/routed/bundled/persona lists cannot silently drift. -- Align orchestrator prompt guidance with actual harness capabilities: remove instructions the harness can't honor, and surface real capabilities the coordinator currently under-uses. -- Make the multi-mode orchestrator family (DEFAULT / PLAN / EXECUTE_PLAN / base-deep / fast) consistent in spawnable agents, gate wiring, and step-prompt guidance except where a difference is intentional and documented. -- Close drift risk in the inline-mirrored gate logic with parity guards and shared frozen lists. -- Clean the tool registry of dead/orphaned/deprecated-but-visible tools. -- Decide the fate of the parallel `orchestration/` subsystem (authoritative vs telemetry-only). - -## Non-Goals - -- No behavioral redesign of the reviewer verdict contract, deterministic-edit system, or context-compaction budgets (those are cohesive already). -- No model-routing/provider changes. -- No new agents or tools beyond what's needed to close a gap. -- Not touching `agents-graveyard/` (already dead, no shipped imports). - -## Key Findings (source-backed, from sharded audit at snapshot `68f0ffb8…`) - -### Roster drift (no single source of truth) - -- Agent roster is maintained in 5 independent places: `agents/**/*.ts` default exports (de-facto truth), `cli/src/agents/bundled-agents.generated.ts` (generated, in sync), `openbuff.d.example/routes.json` (hand, **12 dead unshipped ids**), `common/src/constants/agents.ts` `AGENT_PERSONAS`/`AGENT_IDS` (hand, **missing ~18 shipped, includes 6 non-shipped**: `ask`, `planner`, `agent-builder`, `reviewer`, `file-explorer`, `researcher`), and `spawnableAgents` arrays in `base2.ts:116`, `base-deep.ts:352`, `general-agent.ts`. -- `directory-lister` and `glob-matcher` are bundled + registered + routed but **NOT spawnable by any orchestrator** (dead-end agents). - -### Orchestrator family inconsistency - -- `base-deep.ts:352` hand-overrides `spawnableAgents`, dropping `context-pruner` and `tmux-cli` that `base2`'s computed list includes. -- `base2-fast` omits `browser-use` (present in DEFAULT/PLAN/EXECUTE_PLAN). -- EXECUTE_PLAN step prompt (`buildExecutePlanStepPrompt`, base2.ts:6424) **drops** the editor-handoff / "don't manually spawn code-reviewer" guidance that DEFAULT's `buildImplementationStepPrompt` carries. -- `gateAwarenessSection` is conditional (`isDefault`) in base2 but **unconditional** in base-deep — divergent gating. -- PLAN's instruction builder (`buildPlanOnlyInstructionsPrompt`) reimplements blocks rather than composing from the implementation builder — the main prompt drift surface. - -### Prompt ↔ capability mismatch (core of the complaint) - -- **HIGH:** `buildBroadAuditSection` (quality-prompt-section.ts step 4) tells the coordinator to call `evaluate_audit_coverage` with each shard's `structuralReceipt`, but the shards it names in steps 1–3 are `file-picker`/`code-searcher` (discovery-only) which **cannot emit `structuralReceipt`**. Only `general-agent` audit shards emit it via `write_audit_findings`. The prompt's produce-path and consume-path don't connect. -- `write_audit_findings` / `synthesizer` / `general-agent` durable-findings flow is essentially invisible in the orchestrator prompt, so the coordinator won't reliably use it. -- `frontendSection` re-export at `quality-prompt-section.ts:77` is production-dead (only the snapshot test imports it; production uses the `{CODEBUFF_FRONTEND_SECTION}` placeholder). - -### Gate / reviewer drift risk - -- Entire gate lifecycle is inline-mirrored inside `createBase2.handleSteps` (serialized via `.toString()`), with canonical copies in `gate-*.ts`. -- Parity guards exist for `gate-repair.ts` and `gate-reviewer.ts` (good), but the aux-path helpers `normalizeGateFilePath`, `normalizeGateFileList`, `gateFileSetsEqual` (`gate-paths.ts`) have **no cross-implementation parity test** — `gate-aux-triggers.test.ts` only tests the inline copies. -- Security-sensitive glob list is duplicated (inline gate predicate vs `securityReviewSection`), kept in sync by convention only. - -### Tool registry hygiene - -- schema/handler/metadata form a compile-enforced bijection (good). -- Dead tools (active + promptVisible, granted to no agent): `lookup_agent_info`, `render_ui`, `find_files`, `find_files_matching_content`. -- `read_slices` is correctly quarantined in metadata but still in `publishedTools` and the generated `agents/types/tools.ts` type surface. -- `tool-reachability.test.ts` does not enumerate all structured-output agents (coverage gap). - -### Orchestration subsystem duality - -- `packages/agent-runtime/src/orchestration/` (`select-agent-attempt`, `workflow-engine`, `discovery-coordinator`) is invoked but advisory/bookkeeping; the authoritative orchestration is the base2 inline gate. `workflow-engine` is telemetry-only — a dual-system smell. - -### Discovery agent boundaries - -- `basher` has no `terminalPermissionProfile` while debugger/git-committer/dependency-manager/librarian all do. -- `file-picker` ↔ `file-lister` overlap (file-lister is file-picker's internal worker yet carries its own spawnerPrompt). - -## Relevant Files / Systems - -- Orchestrator: `agents/base2/base2.ts`, `base2-plan.ts`, `base2-execute-plan.ts`, `base2-fast*.ts`, `base-deep.ts` -- Prompt sections: `agents/base2/quality-prompt-section.ts`, `common/src/constants/prompt-sections.ts`, `common/src/constants/git-discipline.ts`, `packages/agent-runtime/src/templates/{strings,types}.ts` -- Gates: `agents/base2/gate-{state,files,paths,repair,reviewer}.ts`, parity tests in `agents/__tests__/gate-*.test.ts` -- Reviewers/specialists: `agents/reviewer/code-reviewer.ts`, `agents/security-reviewer/security-reviewer.ts`, `agents/specialists/create-specialist.ts`, `common/src/agents/specialist-risk-router.ts` -- Roster/routing: `common/src/constants/agents.ts`, `openbuff.d.example/routes.json`, `cli/scripts/prebuild-agents.ts`, `cli/src/agents/bundled-agents.generated.ts` -- Tools: `common/src/tools/{constants,list,metadata}.ts`, `packages/agent-runtime/src/tools/handlers/list.ts`, `agents/tool-reachability.test.ts` -- Runtime coordination: `packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts`, `.../orchestration/*` -- Docs: `docs/agents-and-tools.md`, `agents/patterns/INDEX.md` - -## Acceptance Criteria - -- A guard test fails if any `routes.json` / persona / spawnable entry references an id that is neither a shipped/bundled agent nor an explicitly-allowlisted external CLI agent. -- `AGENT_PERSONAS`/`AGENT_IDS` either derived from bundled agents or reconciled + guarded. -- `directory-lister` / `glob-matcher` are either reachable or removed, with a test asserting no bundled agent is silently unreachable. -- `base-deep` and `base2` spawnable lists are consistent (test-guarded superset relationship) with intentional deltas documented. -- EXECUTE_PLAN and DEFAULT step prompts share editor-handoff guidance; PLAN composes from shared builders. -- `buildBroadAuditSection` routes audit shards to the agent(s) that actually emit receipts, and the audit doc + prompt agree. -- Aux-path gate helpers gain a parity test; security-glob list has one frozen source + parity test. -- Dead tools resolved; `read_slices` removed from published/generated surface. -- A documented decision on `orchestration/` (keep-as-telemetry vs remove vs promote), with a test or doc note reflecting it. -- All touched packages pass their typecheck + tests. diff --git a/.agents/sessions/harness-cohesion-audit-2026-07/STATE.json b/.agents/sessions/harness-cohesion-audit-2026-07/STATE.json deleted file mode 100644 index ed7289a860..0000000000 --- a/.agents/sessions/harness-cohesion-audit-2026-07/STATE.json +++ /dev/null @@ -1,17 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "harness-cohesion-audit-2026-07", - "status": "completed", - "currentTask": null, - "revision": 42, - "checkpoint": { - "taskId": "M3.2", - "phase": "validation", - "passed": true, - "summary": "M3.2 resolved reviewer blocker RF-1/RF-2-6e7bf5ad: documented intentional per-mode spawnable deltas inline in base2.ts + added test-asserted delta block to roster-drift.test.ts (browser-use unconditional across all modes; fast delta = default-only editor family + thinker; plan delta = implementation-only mutation agents). roster-drift 7 pass, all typechecks green.", - "receiptIds": ["hqcVDZOlRfU", "hqYwB06XYSs"], - "recordedAt": "2026-07-21T21:49:00.880Z" - }, - "createdAt": "2026-07-21T20:48:18.665Z", - "updatedAt": "2026-08-02T21:32:41.648Z" -} diff --git a/.agents/sessions/harness-cohesion-audit-2026-07/STATUS.md b/.agents/sessions/harness-cohesion-audit-2026-07/STATUS.md deleted file mode 100644 index 7cf71627d2..0000000000 --- a/.agents/sessions/harness-cohesion-audit-2026-07/STATUS.md +++ /dev/null @@ -1,131 +0,0 @@ -# STATUS — Harness Cohesion Audit & Remediation - -## Current state - -- **Phase:** Execution complete — all executable milestones (M1–M7) implemented and validated; awaiting the automated validation/reviewer gate on the changed source files. (Superseded the initial plan-only phase; see appended execution entries below.) -- **Mode:** EXECUTE_PLAN — the audit + durable packet were produced during planning, then source remediation was applied. - -## Audit coverage (complete) - -All six harness domains + prompt layer audited via parallel shards against snapshot `68f0ffb8…`: - -1. Orchestrator family (base2 / base2-plan / base2-execute-plan / base2-fast / base-deep) — ✅ -2. Gates & reviewers (aux gates, validation+reviewer gate, verdict parsing, parity tests) — ✅ -3. Specialists (14) + risk router — ✅ -4. Discovery/execution agents (19) — ✅ -5. Tool-grant / reachability layer — ✅ -6. Runtime coordination (spawn/handoff/receipt, orchestration/) — ✅ -7. Prompt-section ↔ consumer cohesion — ✅ - -## Completed - -- SPEC.md, PLAN.md, STATUS.md, LESSONS.md written. -- Findings synthesized into 7 milestones (M1–M7) with stable task IDs. -- All executable milestones implemented and validated: M1 (single source of truth), M2 (prompt↔capability), M3 (orchestrator family), M4 (gate drift guards), M5 (tool hygiene; M5.2 cancelled/resolved-by-quarantine), M6 (discovery boundaries), M7 (orchestration/ kept advisory). See the appended execution entries below. -- Follow-up reviewer read-budget fix landed in `sdk/src/tools/read-files.ts` (see the appended 2026-07-22 entry). - -## Pending - -- Automated validation/reviewer gate re-run against the fresh snapshot on the changed source files. - -## Blocked / needs user decision - -- M5.2 (genuinely remove `read_slices` from the published/generated tool type surface): left cancelled/resolved-by-quarantine because forcing it would break the external custom-agent published-tool type contract. Reopen only on explicit user direction. - -## Next checkpoint - -Await the automated validation/reviewer gate on the changed source files. If it clears, the remediation is complete; otherwise address any gate findings. - -## Resume instructions - -1. Read SPEC.md for the ranked findings + evidence. -2. Read PLAN.md for milestone/task IDs and validation gates. -3. Milestones M1–M7 are implemented; verify the appended execution entries against the live source before resuming. -4. Update this file via `update_plan_status` at each task boundary. - - - -## Confirmed Decisions (2026-07) — 2026-07-21T20:48:18.665Z - -- Canonical roster source of truth = `agents/**/*.ts` default exports (what the runtime actually loads/bundles). -- `openbuff.d.example/routes.json` and `common/src/constants/agents.ts` persona maps are DERIVED-OR-GUARDED against that canonical roster (reconcile current contents + add a drift-guard test), not a full gener-from-source refactor this pass. -- External CLI agents (claude-code-cli, codex-cli, gemini-cli, codebuff-local-cli, notion-\*) and judges are an explicit documented allowlist for routes.json, not dead ids to delete blindly. -- Persona-map open question RESOLVED: reconcile + guard (smaller change), per user. - - - -## M1.1 Roster Inventory Matrix — 2026-07-21T20:50:21.223Z - -Canonical roster source = default exports under agents/\*_/_.ts (what the runtime loads), mirrored by getBundledAgentIds() in cli/src/agents/bundled-agents.generated.ts (46 ids). Guard test uses that generated list as truth. - -routes.json DEAD ids (no shipped agent, no .agents/ local agent) -> REMOVE: base, file-explorer, researcher. - -routes.json EXTERNAL/LOCAL ALLOWLIST (exist in .agents/ or evals judges) -> KEEP + document: claude-code-cli, codebuff-local-cli, codex-cli, gemini-cli, notion-query-agent, notion-researcher, judge-gpt, judge-gemini, judge-claude. - -agents.ts AGENT_PERSONAS stale keys (not shipped) -> remove/remap: base (keep? used by graveyard only), ask, file-explorer, researcher, planner, agent-builder, reviewer. Note reviewer->code-reviewer rename; base retained as orchestrator persona alias but base2/base-deep are the real ids. - -Consumers of AGENT_PERSONAS/AGENT_NAMES/AGENT_IDS (blast radius): common/src/util/agent-name-resolver.ts (Object.entries over AGENT_PERSONAS), cli display via AGENT_NAMES; graveyard base-factory.ts (dead, ignore). No shipped hard dependency on the stale keys except agent-name-resolver which tolerates any key set. - -Not spawnable by any orchestrator (M1.5): directory-lister, glob-matcher. - -Decision (confirmed by user): agents/\*_/_.ts exports are canonical; routes.json + persona maps are reconciled + drift-guarded, not fully generated this pass. - - - -## M1 Complete (2026-07) — 2026-07-21T21:06:44.763Z - -M1 single-source-of-truth milestone done. AGENT_PERSONAS reconciled to shipped roster (stale keys removed, ~18 shipped added, satisfies constraint relaxed to string-keyed record). routes.json pruned of 3 dead ids (base, file-explorer, researcher); 55 agents remain (shipped + external-CLI/eval allowlist). New guard agents/**tests**/roster-drift.test.ts (4 pass) checks personas/routes/spawnable reference only shipped-or-allowlisted ids AND every shipped non-root agent is reachable or intentionallyExcluded. M1.5: directory-lister/glob-matcher kept bundled, intentionally excluded from spawnability (mechanical work = direct glob/list_directory tools). common + cli typecheck green. Now on M2.1. - - - -## M5.2 BLOCKED — conflict with compatibility invariants — 2026-07-21T21:26:26.377Z - -M5.2 (remove read_slices from publishedTools + generated agents/types/tools.ts) contradicts two deliberate, test-encoded invariants in common/src/tools/**tests**/tool-registration-consistency.test.ts: - -1. 'quarantined compatibility tools remain registered and published' asserts every quarantinedToolName (now read_slices + apply_smart_patch + the 4 newly-quarantined dead tools) MUST stay in publishedTools so persisted histories / external callers get a compatibility response, not an unknown-tool error. -2. 'generated agent tool types include every published-style tool name' requires the generated type surface to include every non-internal toolName; read_slices is not in the internal-only exclusion set. - -read_slices is ALREADY prompt-invisible via quarantine metadata (reachability=quarantined, promptVisible=false), which is the real 'don't show the model' fix. Removing it from the published/generated surface would require weakening both guards AND changes the external custom-agent type contract (a compatibility break). Recommend M5.2 be re-scoped to no-op/already-resolved-by-quarantine. Awaiting user decision. - - - -## Execution Complete (2026-07) — 2026-07-21T21:38:02.119Z - -All executable milestones done and validated green: - -- M1 (single source of truth): AGENT_PERSONAS reconciled; routes.json pruned of 3 dead ids; new roster-drift.test.ts (4 pass); directory-lister/glob-matcher intentionally-excluded decision encoded. common+cli typecheck green. -- M2 (prompt<->capability): buildBroadAuditSection now routes audit shards to general-agent + write_audit_findings and marks file-picker/code-searcher discovery-only; docs already matched; dead frontendSection re-export removed + test rewired. snapshot + base2 green. -- M3 (orchestrator family): base-deep now inherits base2's computed spawnable list (no more dropped context-pruner/tmux-cli/browser-use); EXECUTE_PLAN step prompt composes from buildImplementationStepPrompt (editor-handoff guidance restored); M3.2/M3.4 resolved as consistent-by-construction/documented. typecheck + 89 tests green. -- M4 (gate drift guards): gate-paths-parity.test.ts (3 pass) + security-glob-parity.test.ts (4 pass) added. -- M5 (tool hygiene): 4 dead tools (lookup_agent_info, render_ui, find_files, find_files_matching_content) quarantined -> prompt-invisible; tool-reachability broadened to all structured-output agents (caught a latent no-op guard + fixed effective-tool resolution). M5.2 CANCELLED (resolved-by-quarantine; forcing it would break the published/type compatibility contract). -- M6 (discovery boundaries): basher given explicit workspace-write terminal profile; file-lister documented as file-picker's internal worker. -- M7: orchestration/ kept advisory with doc marker. - -Only open item: M5.2 needs a user decision if you want to intentionally break the external custom-agent published-tool type contract; otherwise it stays resolved by quarantine. All touched packages typecheck clean; all new + existing targeted tests pass. Awaiting the automated validation/reviewer gate on the changed source files. - - - -## M3.2 Blocker Resolved + Non-Blocking Cleanup (2026-07) — 2026-07-21T21:49:48.771Z - -Reviewer blocker RF-1/RF-2-6e7bf5ad (M3.2 uncertain) resolved: intentional per-mode spawnable deltas documented in base2.ts and test-frozen in roster-drift.test.ts (browser-use unconditional; fast/plan deltas asserted). roster-drift now 7 pass, all typechecks green. M3.2 marked done. Also addressing 2 non-blocking reviewer doc nits (stale frontendSection docstring + mislabeled M3.3->M2.1 milestone comment) in quality-prompt-section.ts since they are the exact stale-reference drift this audit targets. - - - -## Reviewer Read-Budget Fix Complete (2026-07) — 2026-07-22T03:58:01.890Z - -Fixed the reviewer-gate attestation failure at its root: raised MAX_RENDER_CHARS to equal MAX_FILE_BYTES (10MB) in sdk/src/tools/read-files.ts so the byte gate is the single read ceiling. Reviewers (which read only via read_files) can now fully read large files like base2.ts (313 KB) and attest to them. SDK read-files suite 62 pass / 0 fail; SDK typecheck clean. Coupled test expectations updated. Awaiting runtime reviewer gate re-run against the fresh snapshot; the prior did-not-attest-to-base2.ts blocker should now clear because the file renders fully. Open decision still outstanding: M5.2 (whether to genuinely remove read_slices from the published/generated type surface, a compatibility break), left cancelled/quarantined unless user directs otherwise. - - - -## M5.2 RESOLVED — read_slices fully removed from live surface — 2026-08-02T21:32:41.647Z - -M5.2 is resolved. Verification against live source confirmed `read_slices` was already absent from `toolNames`, `publishedTools`, and `quarantinedToolNames` in `common/src/tools/constants.ts` (removed during the earlier read-tool-unification work). The two test-encoded invariants that originally blocked removal no longer apply because `read_slices` is no longer in `quarantinedToolNames`. - -Remaining dead-code references were cleaned up in this pass: - -- `packages/agent-runtime/src/tools/tool-executor.ts:1449` — removed unreachable `toolName === 'read_slices'` branch from the permission classifier. -- `packages/agent-runtime/src/structural-read.ts` — updated two stale comments that referenced `read_slices` / the deprecated alias. - -Negative assertions and historical docs were intentionally kept: `agents/tool-reachability.test.ts` (regression guards asserting `read_slices` does not appear in prompts/docs), `agents/__tests__/editor.test.ts` (negative assertion), and `docs/audits/read-write-indexing-2026-07-14/findings/read-path.md` (historical audit doc). - -Validation: gate NON_BLOCKING (3 minor nits unrelated to this change: window<=0 clamping in buildWindowBlock, missing direct unit tests for build\*Block builders, per-call z.toJSONSchema in coerceInputScalarsBySchema). Typecheck green across all packages. diff --git a/.agents/sessions/harness-ui-overhaul-2026-08/PLAN.md b/.agents/sessions/harness-ui-overhaul-2026-08/PLAN.md deleted file mode 100644 index 7232313eb6..0000000000 --- a/.agents/sessions/harness-ui-overhaul-2026-08/PLAN.md +++ /dev/null @@ -1,82 +0,0 @@ -# PLAN — Harness UI Overhaul (Full Sweep) - -Slug: `harness-ui-overhaul-2026-08` -Snapshot: `399c28986835a71e7ee7b45b6dcaf9bf2b9f8ef85181a443bdcfaeebbe6a137c` -Source: SPEC.md (this session), prior harness inventory (399c… snapshot, 22 subsystems) -Current workspace: `feat/task-memory-evidence-pipeline`, clean worktree. - -## Milestones - -### M0 — Session bootstrap (done) - -- SPEC.md landed. This PLAN.md + STATUS.md land this turn. Gate clears on session artifacts only (no src/ changes this turn). - -### M1 — Primitive + Completion summary (P1, vertical slice 1) - -Goal: one visual language — no more `✅ 3 files edited | ❌ Hooks: 1 failed` pipe-joined text line. - -- [ ] M1-T1 Extract `cli/src/components/renderers/harness-box.tsx` (`HarnessBox`, `HarnessSection`, `HarnessRow`): `borderStyle: single` + `BORDER_CHARS` + `paddingLeft/Right 1 + gap`, theme token prop (`secondary` default, `success/error/warning` for status). Adopt in `PlanBox` + `GateStateBox` (no visual regression). -- [ ] M1-T2 Extend `cli/src/types/chat.ts`: `CompletionSummaryContentBlock { type:'completion-summary', summary: CompletionSummary }` + guard `isCompletionSummaryBlock`. Extend `ContentBlock` union. -- [ ] M1-T3 New renderer `cli/src/components/renderers/completion-summary-box.tsx`: sections Files / Hooks / Review / Tests / Auxiliary / Errors, icon+color mapping (`STATUS_ICON` style: `✓ ✗ ⚠` or reuse emoji but themed), bordered via `HarnessBox`. Empty → null (honor `computeCompletionSummary` null). -- [ ] M1-T4 Wire `cli/src/utils/sdk-event-handlers.ts:handleFinish`: emit typed block instead of `type:'text'` with `formatCompletionSummary(summary)`. Keep `formatCompletionSummary` export for logs/tests. -- [ ] M1-T5 Route in `cli/src/components/blocks/single-block.tsx` (`case 'completion-summary'` → `CompletionSummaryBox`), thread `availableWidth`/`markdownPalette`/`onInsertCommand` as needed (no new props needed, but keep parity with plan/gate). -- [ ] M1-T6 Tests: extend `cli/src/utils/__tests__/completion-summary.test.ts` (data unchanged), add `cli/src/components/__tests__/completion-summary-box.test.tsx` (dark/light, null/empty, mixed verdict), `cli/src/utils/__tests__/sdk-event-handlers.test.ts` for handleFinish block type. -- Validate: `bun --cwd cli run typecheck`, `bun --cwd cli test` (targeted), `tmux-cli` smoke: stream → completion box renders. - -### M2 — Memory interactive box (P1, vertical slice 2) - -- [ ] M2-T1 Extend `chat.ts`: `MemoryContentBlock { type:'memory', status:'empty'|'status'|'prune-result', revision?, updatedAt?, goal?, counts?, evidence:{live,stale,total}, stalePaths?, hint?, pruneOutcome? }` or reuse raw `TaskMemoryV1` + reconciled evidence; decide in slice kickoff (prefer minimal view-model to keep renderer pure). -- [ ] M2-T2 New renderer `cli/src/components/renderers/memory-box.tsx`: header `revision · age` (`formatAge`), goal preview 120 chars with expand, counts grid, evidence badge (`fresh/stale` colored), collapsible `Stale paths (5)` via `CollapseButton`/`Button`, empty state copy ("written after your first successful run"), error banner for prune failure reasons. -- [ ] M2-T3 Refactor `cli/src/commands/memory-command.ts`: `handleMemoryCommand` returns `{ blocks: ContentBlock[] }` or typed message payload instead of `string`; keep `string` fallback export for `handleMemoryCommand` tests via wrapper. Preserve `WorkspaceJournalService.create → collectWorkspaceMoves` move-rebinding for both status & prune. -- [ ] M2-T4 Wire `cli/src/commands/command-registry.ts` memory handler: `getSystemMessage(string)` → typed memory block insertion (mirror `appendLocalMessage` but block-aware; add `appendLocalBlocks` helper if needed, keep `appendLocalMessage` for other commands until M3). -- [ ] M2-T5 Interactions: `Button` "Prune stale evidence" → `onInsertCommand('/memory prune')`, hover `borderColor theme.foreground`, `DASHED_BORDER` not used (harness = solid). -- [ ] M2-T6 Tests: update `cli/src/commands/__tests__/memory-command.test.ts` (block shape), add `memory-box.test.tsx`, add move-aware prune integration test (rename fixture → stale rebounds, not deleted). -- Validate: `bun --cwd cli test` memory + command-registry, tmux `/memory` → box, `/memory prune` flow. - -### M3 — Sweep remaining plain-text reports (P2) - -Each command gets minimal typed block + box, reusing `HarnessBox`. - -- [ ] M3-T1 `context` (`cli/src/commands/context.ts`): `context` block + `ContextBox` (ledger breakdown, trigger/target budgets). -- [ ] M3-T2 `info` (`cli/src/commands/info.ts`): `info` block + `InfoBox`. -- [ ] M3-T3 `doctor` (`formatOpenbuffProviderStatus` + diagnostics): `doctor` block + `DoctorBox` (split provider status vs diagnostics sections). -- [ ] M3-T4 `index` (`cli/src/commands/index-command.ts`): `index-status` block + `IndexStatusBox`. -- [ ] M3-T5 `plan-status` + `plans` (`formatPlanStatusReport`/`formatPlanListReport`): `plan-status-list` block + `PlanStatusBox` (preserve `STATUS_BADGE` `[active]/[paused]/…`, `progress done/total`, `currentTask`). -- [ ] M3-T6 `help` audit: if it bypasses box system, migrate; otherwise mark out-of-scope with reason in STATUS matrix. -- [ ] M3-T7 Registry wiring: replace `getSystemMessage(string)` calls for each with block helper; keep `appendLocalMessage(string)` deprecated path until all migrated, then remove or keep shim for skills. -- Validate: per-command tests + one combined `command-registry` sweep test; tmux checklist `/context /info /doctor /index /plan-status /plans`. - -### M4 — Hardening + docs - -- [ ] M4-T1 OpenTUI safety audit: every box wraps markdown/`span` in ``, no `{' '}` whitespace, `minWidth:0` on flex cols, resize 1→2 col not collapsing. Add to `cli/knowledge.md` "HarnessBox" note. -- [ ] M4-T2 Theme parity: `dark` + `light` via `useTheme()`; `messageTextAttributes` preserved. -- [ ] M4-T3 Update `docs/architecture.md` (CLI TUI section) pending-gate box inventory. -- [ ] M4-T4 `LESSONS.md` capture: string→typed-block migration pattern for future harness surfaces. - -## Dependencies - -M1 primitive extraction before M1-T3/M2-T2/M3 boxes. M1-T2 (chat types) before any box. M2-T3 (memory command contract) before registry wiring. M3 can parallelize T1-T5 after harness-box lands. - -## Risks - -- `CompletionSummaryBox` icon drift from `formatCompletionSummary` emoji → mitigate: keep string formatter for logs, box uses theme tokens (not emoji parsing). -- `handleFinish` text→block change breaks `sdk-event-handlers.test.ts` snapshots → update snapshots, keep null-path. -- `memory-command` string→blocks breaks existing tests expecting `string` → keep backward-compat wrapper or update test to assert block shape. -- Move-rebinding regression (rename → stale → prune deletes) → explicit test with `WorkspaceJournalService` mock moves. -- OpenTUI reconciler fragility (`` nesting) → reuse `PlanBox`/`GateStateBox` patterns verbatim. - -## Validation gates - -- `bun --cwd cli run typecheck` -- `bun --cwd cli test` (or `bun test cli/src/utils/__tests__/completion-summary.test.ts cli/src/utils/__tests__/sdk-event-handlers.test.ts cli/src/commands/__tests__/memory-command.test.ts`) -- `tmux-cli` (via `tmux-cli` agent): streaming → completion box, `/memory`, `/memory prune` (move-aware), `/context /info /doctor /index /plan-status /plans` sweep. -- No new `common/` contract beyond `chat.ts` ContentBlock extension; no `sdk/` provider change. - -## Out of scope - -- `read-edit` / `shell-policy` surfaces (separate sessions). -- `common/src/templates` agent examples. - -## Resume pointer - -Next concrete step after gate: `M1-T1 + M1-T2` (harness-box + chat types) — single editor slice, no validation bypass. diff --git a/.agents/sessions/harness-ui-overhaul-2026-08/SPEC.md b/.agents/sessions/harness-ui-overhaul-2026-08/SPEC.md deleted file mode 100644 index e885993492..0000000000 --- a/.agents/sessions/harness-ui-overhaul-2026-08/SPEC.md +++ /dev/null @@ -1,72 +0,0 @@ -# SPEC — Harness UI Overhaul (Full Sweep) - -## Goal - -Unify the Openbuff harness into one visual language. Every system surface that currently emits a plain `string → getSystemMessage → ` must land on the same bordered, themed, OpenTUI-safe renderer system that `PlanBox` / `GateStateBox` / `AgentBranchWrapper` already use. No more split between premium interactive chrome and `console.log`-style dumps for memory or session summaries. - -## Non-goals - -- No redesign of already-mature surfaces (`PlanBox`, `GateStateBox`, `AgentBranchWrapper`/`AgentBlockGrid`/`ImplementorRow`/`DiffViewer`, `StatusBar`) beyond extracting a shared primitive they adopt. -- No new backend or provider APIs; no agent-runtime prompt contract change (`task_completed` stays empty). -- No global theming overhaul; reuse existing `ChatTheme` + `BORDER_CHARS` tokens. -- No migration of log/debug artifacts outside `cli/src`. - -## Requirements - -### R1 — Completion summary becomes a first-class box (P1) - -- Data source stays `computeCompletionSummary(blocks)` (`cli/src/utils/completion-summary.ts`). -- New typed block `type: 'completion-summary'` + renderer `CompletionSummaryBox` replaces the `formatCompletionSummary()` → `type:'text'` injection in `cli/src/utils/sdk-event-handlers.ts:handleFinish`. -- Sections: Files (edited/failed/unconfirmed/rolled_back/rollback_incomplete), Hooks (passed/failed/skipped), Review verdict (`BLOCKING/NON_BLOCKING/LOOKS_GOOD/…` → `error/warning/success`), Tests, Auxiliary, Errors. Status → border color + icon mapping mirrors `GateStateBox` (success/warning/error). -- `formatCompletionSummary()` retained for logs/fallback; renderer is the TUI source of truth. - -### R2 — Memory becomes interactive and on-system (P1) - -- Data source stays `sdk/task-memory-store` + `reconcileTaskMemoryEvidence`/`pruneStale…` via `WorkspaceJournalService`. -- New typed block `type: 'memory'` + `MemoryBox` replaces `string` return from `cli/src/commands/memory-command.ts` (`handleMemoryCommand` → `runStatus`/`runPrune`). -- Layout: header `revision · age` (via `formatAge`), goal preview (120 chars, expand affordance if truncated), counts grid (Decisions · Requirements · Edits / Validations · Blockers · Next actions), evidence `fresh/stale` with colored badge, collapsible `Stale paths (5)` list, empty/no-record state that explains when memory is written. -- Prune affordance: clickable `Button` wired to `onInsertCommand('/memory prune')` (same pattern as `PlanBox` command pills), hover border `theme.secondary → theme.foreground`. Prune failure reasons surfaced verbatim (invalid-record / concurrent-write / write-failed) with no phrasing as "nothing to prune". -- Workspace-move rebinding contract preserved: both status and prune pass `WorkspaceMoveRecord[]` from `WorkspaceJournalService.create`. - -### R3 — Sweep remaining plain-text reports (P2) - -Convert every `getSystemMessage(string)` report in `cli/src/commands/command-registry.ts` and helpers to typed blocks/boxes sharing the same primitive: - -- `/context` (`cli/src/commands/context.ts` — ledger breakdown) -- `/info` (`cli/src/commands/info.ts`) -- `/doctor` (`formatOpenbuffProviderStatus` + agent diagnostics) -- `/index` (`cli/src/commands/index-command.ts`) -- `/plan-status` + `/plans` (`formatPlanStatusReport`/`formatPlanListReport` — retain `STATUS_BADGE`/`progress done/total` + `currentTask` semantics) -- `/help` if it still bypasses the box system; otherwise leave its existing structured screen. - Each gets a minimal typed block (e.g. `context`, `info`, `doctor`, `index-status`, `plan-status-list`) and a `*Box` renderer. No one-off inline styles. - -### R4 — Shared harness chrome primitive (P2, extracted alongside R1/R2) - -- Extract `HarnessBox` (and `HarnessSection`/`HarnessRow` helpers) that codifies the common pattern: `borderStyle:'single' + BORDER_CHARS + theme token + paddingLeft/right 1 + gap`. Adopted by `PlanBox`, `GateStateBox`, `CompletionSummaryBox`, `MemoryBox`, and the sweep boxes. One change propagates. -- Tokens: `DASHED_BORDER_CHARS` / `IMPLEMENTOR_BORDER_CHARS` remain reserved for ghost/implementor contexts; harness uses rounded `BORDER_CHARS`. - -### R5 — Wiring and single-point injection - -- `sdk-event-handlers.ts:handleFinish` and `command-registry.ts` (`appendLocalMessage` path) are the only UI injection seams. Change is additive (new block types) not string-format surgery. - -## Acceptance criteria - -- AC1: A completed run with edits+hooks+review renders a bordered `CompletionSummaryBox` (not a `| -joined` text line); empty runs produce no box (`computeCompletionSummary` null path unchanged). -- AC2: `/memory` (no record) shows the two-line "not yet written" empty state inside a box; `/memory` with record shows revision/age/goal/counts/evidence + stale-list affordance; `/memory prune` with stale entries shows the button and correctly rebinds renamed-file evidence (move-aware). -- AC3: `/context` `/info` `/doctor` `/index` `/plan-status` `/plans` each render as bordered boxes with themed headings (no raw `lines.join('\n')` text blocks). -- AC4: `SingleBlock` routes all new block types; `chat.ts`→`MessageBlock`→`BlocksRenderer` threading of `onInsertCommand`/`markdownPalette`/`availableWidth` matches `PlanBox` precedent. -- AC5: OpenTUI-safe: every markdown / `` / `` fragment is wrapped in ``; no `{' '}` JSX whitespace; no box-inside-text violations; resize (1→2 column) does not collapse `minWidth`. -- AC6: Theme-correct in `dark` and `light` (`useTheme()` tokens, border `success/error/warning/secondary` mapping, `TextAttributes.DIM/BOLD` where appropriate). -- AC7: Backward-compatible: `formatCompletionSummary` and `formatAge`/`pluralizeEntries` remain exported for logs/tests; block types extend the `ContentBlock` union, never retype existing `text`/`tool`/`agent` fields. -- AC8: Tests: `completion-summary.test.ts` (data), new `completion-summary-box.test.tsx`/`memory-box.test.tsx` (render), `command-registry` integration for move-aware prune, plus one `tmux-cli` smoke for streaming→completion and `/memory` flow. - -## Relevant systems - -- `cli/src/types/chat.ts` — `ContentBlock` union, `GateStateStatus`, `PlanArtifactMetadata`. -- `cli/src/types/theme-system.ts` + `cli/src/utils/ui-constants.ts` + `cli/src/hooks/use-theme.tsx` — `ChatTheme`, `BORDER_CHARS`. -- `cli/src/utils/completion-summary.ts` + `cli/src/utils/sdk-event-handlers.ts` (finish seam) + `cli/src/utils/message-block-helpers.ts`. -- `cli/src/commands/memory-command.ts` + `sdk/src/services/task-memory-store.ts` + `common/src/types/task-memory.ts`. -- `cli/src/commands/command-registry.ts` + `cli/src/commands/context.ts`/`info.ts`/`index-command.ts`/`plan-artifacts.ts`. -- `cli/src/components/renderers/{plan-box,gate-state-box}.tsx` + `cli/src/components/blocks/single-block.tsx` + `cli/src/components/message-block.tsx` + `cli/src/utils/markdown-renderer.tsx`. -- `cli/knowledge.md` — autoCollapse, toggle, suggestion-menu, streaming markdown constraints. -- Validation: `cli` Vitest/Bun + `tmux-cli` / `scripts/tmux/tmux-cli.sh`. diff --git a/.agents/sessions/harness-ui-overhaul-2026-08/STATUS.md b/.agents/sessions/harness-ui-overhaul-2026-08/STATUS.md deleted file mode 100644 index d32ff7aaae..0000000000 --- a/.agents/sessions/harness-ui-overhaul-2026-08/STATUS.md +++ /dev/null @@ -1,54 +0,0 @@ -# STATUS — Harness UI Overhaul (Full Sweep) - -Slug: `harness-ui-overhaul-2026-08` -Snapshot: `399c28986835a71e7ee7b45b6dcaf9bf2b9f8ef85181a443bdcfaeebbe6a137c` -Branch: `feat/task-memory-evidence-pipeline` (clean) - -## Current state - -- **Phase:** Plan complete — SPEC.md + PLAN.md landed and under gate review. Ready for `M1` execution. -- **Mode:** PLAN (no `cli/src` edits this turn; gate is artifacts-only) - -## Milestone tracker - -- M0 Session bootstrap — done ✅ (SPEC.md, PLAN.md) -- M1 Primitive + Completion summary (P1) — pending - - M1-T1 `harness-box.tsx` + adopt in PlanBox/GateStateBox - - M1-T2 `chat.ts` `completion-summary` block + guard - - M1-T3 `completion-summary-box.tsx` - - M1-T4 `sdk-event-handlers.ts:handleFinish` wired - - M1-T5 `single-block.tsx` routing - - M1-T6 tests + tmux smoke -- M2 Memory interactive box (P1) — pending -- M3 Sweep remaining reports (P2) — pending -- M4 Hardening + docs — pending - -## Coverage matrix (plan scope) - -| Domain | Shard / File | Covered | -| ------------------------------------ | ------------------------------------------------------------------------------------------------------ | ------------------------ | -| cli TUI renderers | `cli/src/components/renderers/*` | yes — SPEC R4, M1 | -| cli chat types | `cli/src/types/chat.ts` | yes — R1/R2, M1-T2/M2-T1 | -| cli completion summary | `cli/src/utils/completion-summary.ts` | yes — R1, M1 | -| cli sdk-event-handlers (finish seam) | `cli/src/utils/sdk-event-handlers.ts` | yes — R5, M1-T4 | -| cli memory command | `cli/src/commands/memory-command.ts` | yes — R2, M2 | -| cli command registry + helpers | `cli/src/commands/command-registry.ts` + `context.ts`/`info.ts`/`index-command.ts`/`plan-artifacts.ts` | yes — R3, M3 | -| cli blocks routing | `cli/src/components/blocks/single-block.tsx` | yes — M1-T5/M2 | -| sdk task memory store | `sdk/src/services/task-memory-store.ts` | yes — R2 (data source) | -| common task-memory types | `common/src/types/task-memory.ts` | yes — R2 | -| theme/tokens | `cli/src/types/theme-system.ts` + `ui-constants.ts` + `hooks/use-theme.tsx` | yes — R4 | - -Domains explicitly out-of-scope for this plan: `sdk/provider`, `agent-runtime` prompts, `packages/indexer`, `common/templates`. - -## Validation gates (next) - -- `bun --cwd cli run typecheck` -- `bun --cwd cli test` (completion-summary, sdk-event-handlers, memory-command, command-registry) -- `tmux-cli` smoke: streaming→completion box, `/memory` + `/memory prune` (move-aware), `/context /info /doctor /index /plan-status /plans` - -## Resume instructions - -1. Read `SPEC.md` + `PLAN.md` in this session dir. -2. Start at `M1-T1 + M1-T2` (harness-box + chat types) — single editor slice. -3. Keep `formatCompletionSummary`/`formatAge` exported for logs/tests (AC7). -4. Preserve move-rebinding: `WorkspaceJournalService.create → collectWorkspaceMoves` in memory status & prune. diff --git a/.agents/sessions/read-edit-auth-unification-2026-07/EVENTS.jsonl b/.agents/sessions/read-edit-auth-unification-2026-07/EVENTS.jsonl deleted file mode 100644 index 16b98dd937..0000000000 --- a/.agents/sessions/read-edit-auth-unification-2026-07/EVENTS.jsonl +++ /dev/null @@ -1,42 +0,0 @@ -{"ts":"2026-07-22T07:20:12.216Z","kind":"task_update","summary":"Updated 1 task line(s): M1.1","payload":{"matched":["M1.1"]}} -{"ts":"2026-07-22T07:20:12.216Z","kind":"current_task","summary":"Current task -> \"M1.1\"","payload":{"currentTask":"M1.1"}} -{"ts":"2026-07-22T07:36:28.623Z","kind":"append_lesson","summary":"Appended entry \"M1.1 Finding — str_replace path already unifies (2026-07)\" to STATUS.md","payload":{"heading":"M1.1 Finding — str_replace path already unifies (2026-07)","artifact":"STATUS.md"}} -{"ts":"2026-07-22T08:45:45.930Z","kind":"task_update","summary":"Updated 2 task line(s): M1.1, M1.2","payload":{"matched":["M1.1","M1.2"]}} -{"ts":"2026-07-22T08:45:45.930Z","kind":"append_lesson","summary":"Appended entry \"M1.1 validation\" to PLAN.md","payload":{"heading":"M1.1 validation","artifact":"PLAN.md"}} -{"ts":"2026-07-22T08:45:45.930Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-22T08:45:45.930Z","kind":"current_task","summary":"Current task -> \"M1.2\"","payload":{"currentTask":"M1.2"}} -{"ts":"2026-07-22T09:07:51.632Z","kind":"task_update","summary":"Updated 2 task line(s): M1.2, M2.1","payload":{"matched":["M1.2","M2.1"]}} -{"ts":"2026-07-22T09:07:51.633Z","kind":"append_lesson","summary":"Appended entry \"M1.2 validation\" to PLAN.md","payload":{"heading":"M1.2 validation","artifact":"PLAN.md"}} -{"ts":"2026-07-22T09:07:51.633Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-22T09:07:51.633Z","kind":"current_task","summary":"Current task -> \"M2.1\"","payload":{"currentTask":"M2.1"}} -{"ts":"2026-07-22T09:27:19.405Z","kind":"task_update","summary":"Updated 2 task line(s): M2.1, M3.1","payload":{"matched":["M2.1","M3.1"]}} -{"ts":"2026-07-22T09:27:19.405Z","kind":"append_lesson","summary":"Appended entry \"M2.1 validation\" to PLAN.md","payload":{"heading":"M2.1 validation","artifact":"PLAN.md"}} -{"ts":"2026-07-22T09:27:19.405Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-22T09:27:19.405Z","kind":"current_task","summary":"Current task -> \"M3.1\"","payload":{"currentTask":"M3.1"}} -{"ts":"2026-07-22T10:12:24.826Z","kind":"task_update","summary":"Updated 2 task line(s): M4.1, M3.1","payload":{"matched":["M4.1","M3.1"]}} -{"ts":"2026-07-22T10:12:24.826Z","kind":"append_lesson","summary":"Appended entry \"Repair scope and caller boundary lesson\" to PLAN.md","payload":{"heading":"Repair scope and caller boundary lesson","artifact":"PLAN.md"}} -{"ts":"2026-07-22T10:12:24.826Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-22T10:12:24.826Z","kind":"current_task","summary":"Current task -> \"M3.1\"","payload":{"currentTask":"M3.1"}} -{"ts":"2026-07-22T10:15:36.521Z","kind":"append_lesson","summary":"Appended entry \"AC2 reconciliation validation\" to STATUS.md","payload":{"heading":"AC2 reconciliation validation","artifact":"STATUS.md"}} -{"ts":"2026-07-22T10:15:36.521Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-22T10:22:14.677Z","kind":"task_update","summary":"Updated 2 task line(s): M3.1, M4.1","payload":{"matched":["M3.1","M4.1"]}} -{"ts":"2026-07-22T10:22:14.677Z","kind":"append_lesson","summary":"Appended entry \"M3.1 completion evidence\" to PLAN.md","payload":{"heading":"M3.1 completion evidence","artifact":"PLAN.md"}} -{"ts":"2026-07-22T10:22:14.677Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-22T10:22:14.677Z","kind":"current_task","summary":"Current task -> \"M4.1\"","payload":{"currentTask":"M4.1"}} -{"ts":"2026-07-22T13:43:35.102Z","kind":"task_update","summary":"Updated 2 task line(s): M4.1, M4.2","payload":{"matched":["M4.1","M4.2"]}} -{"ts":"2026-07-22T13:43:35.103Z","kind":"append_lesson","summary":"Appended entry \"M4.1 completion\" to PLAN.md","payload":{"heading":"M4.1 completion","artifact":"PLAN.md"}} -{"ts":"2026-07-22T13:43:35.103Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-07-22T13:43:35.103Z","kind":"current_task","summary":"Current task -> \"M4.2\"","payload":{"currentTask":"M4.2"}} -{"ts":"2026-07-22T14:30:21.276Z","kind":"current_task","summary":"Current task -> \"M4.2\"","payload":{"currentTask":"M4.2"}} -{"ts":"2026-07-22T14:31:08.397Z","kind":"task_update","summary":"Updated 2 task line(s): M4.2, M5.1","payload":{"matched":["M4.2","M5.1"]}} -{"ts":"2026-07-22T14:31:08.397Z","kind":"current_task","summary":"Current task -> \"M5.1\"","payload":{"currentTask":"M5.1"}} -{"ts":"2026-07-22T14:31:56.376Z","kind":"append_lesson","summary":"Appended entry \"M4.2 completion — dead-tool removal\" to STATUS.md","payload":{"heading":"M4.2 completion — dead-tool removal","artifact":"STATUS.md"}} -{"ts":"2026-07-22T14:34:15.160Z","kind":"current_task","summary":"Current task -> \"M5.1\"","payload":{"currentTask":"M5.1"}} -{"ts":"2026-07-22T14:34:27.452Z","kind":"task_update","summary":"Updated 2 task line(s): M5.1, M5.2","payload":{"matched":["M5.1","M5.2"]}} -{"ts":"2026-07-22T14:34:27.452Z","kind":"current_task","summary":"Current task -> \"M5.2\"","payload":{"currentTask":"M5.2"}} -{"ts":"2026-07-22T14:35:33.475Z","kind":"current_task","summary":"Current task -> \"M5.2\"","payload":{"currentTask":"M5.2"}} -{"ts":"2026-07-22T14:36:28.088Z","kind":"current_task","summary":"Current task -> \"M5.2\"","payload":{"currentTask":"M5.2"}} -{"ts":"2026-07-22T14:38:02.848Z","kind":"task_update","summary":"Updated 1 task line(s): M5.2","payload":{"matched":["M5.2"]}} -{"ts":"2026-07-22T14:38:02.848Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} -{"ts":"2026-07-22T14:38:19.432Z","kind":"append_lesson","summary":"Appended entry \"Session complete — 2026-07-22\" to STATUS.md","payload":{"heading":"Session complete — 2026-07-22","artifact":"STATUS.md"}} -{"ts":"2026-07-22T14:38:19.432Z","kind":"session_status","summary":"Session status -> completed","payload":{"status":"completed"}} diff --git a/.agents/sessions/read-edit-auth-unification-2026-07/LESSONS.md b/.agents/sessions/read-edit-auth-unification-2026-07/LESSONS.md deleted file mode 100644 index 59fae7e99d..0000000000 --- a/.agents/sessions/read-edit-auth-unification-2026-07/LESSONS.md +++ /dev/null @@ -1,15 +0,0 @@ -# LESSONS — Read/Edit Authorization Unification - -## M4.2 dead-tool removal lessons - -- Removing a tool name from `toolNames`/`publishedTools` in `common/src/tools/constants.ts` triggers `satisfies` errors across every parallel registry: `list.ts`, `metadata.ts` (READ/MUTATION/NAMED_PATH sets + PATH_INPUTS), `input-aliases.ts`, runtime `handlers/list.ts`, and both generated `types/tools.ts` files (`agents/types/tools.ts` and `common/src/templates/initial-agents-dir/types/tools.ts`). Remove the name everywhere plus delete the dedicated param schema and handler in one transaction. -- Residual references hide where a grep surfaces them but typecheck pins precisely: `sdk/src/tool-execution-deadline.ts` (`FILE_MUTATION_TOOLS` set) and `common/src/tools/__tests__/tool-metadata.test.ts` (mutation-schema loop) both hard-listed `apply_smart_patch`. -- The AC4 legacy-format removal dropped `filesystemResultFormat` from `OpenbuffClientOptions`, so `cli/src/utils/codebuff-client.ts` had to stop passing it — an unrelated-looking CLI typecheck error that was actually part of the same unified-model cleanup. -- Gate/pruner string literals for `apply_smart_patch` in `agents/base2/gate-files.ts`, `base2.ts`, `editor.ts`, and `context-pruner.ts` are intentional backward-compat for classifying persisted tool-call history and must stay; they are asserted by context-pruner/base2 tests. -- `mintSliceCapability` is scope-gated (mints a token only when a handler passes `scope`), so `extractSlices` core tests asserting a `readCapability` had to be updated: the shared core returns tokenless slices and the `read_files`/`rewrite_symbol` handlers re-mint scoped caps. -- `agents/tool-reachability.test.ts` asserts the docs no longer contain `### apply_smart_patch` or the `read_slices` deprecated-alias section, so `docs/agents-and-tools.md` sections had to be removed to match the source-of-truth removal. - -## Process lessons - -- Never run `git stash` for inspection during an active edit session; it silently shelved all working changes. Recovered with `git stash pop`. Inspect committed baselines with `git show HEAD:path` instead. -- `edit_transaction` preflight is atomic and a multi-edit doc removal where an earlier edit shifts later line numbers invalidates the later `replace_range` hash. Apply shifting edits sequentially, or re-read/retry with the runtime-provided fresh `readCapability` (no extra read round-trip needed). diff --git a/.agents/sessions/read-edit-auth-unification-2026-07/PLAN.md b/.agents/sessions/read-edit-auth-unification-2026-07/PLAN.md deleted file mode 100644 index 5262d5d62f..0000000000 --- a/.agents/sessions/read-edit-auth-unification-2026-07/PLAN.md +++ /dev/null @@ -1,82 +0,0 @@ -# PLAN — Read/edit auth unification + mutation results + editor fix + legacy removal - - - -Execute milestones in order. Each milestone has a validation gate; do not mark done until it passes. - -## M1 — Unify authorization on content-correctness (the core) - -- [x] M1.1 Make cap.v3 the single authority: validation re-hashes the current targeted content and compares to the capability hash; remove whole-file-vs-range authority branching in `process-str-replace.ts` / `process-edit-transaction.ts`. (Claiming: make cap.v3 the single authority; validation re-hashes current targeted content vs capability hash; remove whole-file-vs-range authority branching. Design fork resolved: keep observed-bytes floor (partial range read cannot mint whole-file authority).) (Validated: agent-runtime typecheck + 178 targeted tests passed; implementation is existing scope+content-hash behavior plus clarifying docs.) - - Acceptance: AC1 — a range cap authorizes a matching edit within its observed range; it cannot authorize a whole-file overwrite unless it covers the whole current file. A whole-file cap authorizes a matching sub-range or whole-file edit; authority is content-hash equality within the observed-bytes floor. - - Validate: `cd packages/agent-runtime && bun run typecheck` + `bun test src/__tests__/process-str-replace.test.ts src/__tests__/process-edit-transaction.test.ts src/__tests__/read-files-edit-state.test.ts` -- [x] M1.2 Keep the anti-footgun: a whole-file overwrite still requires a hash covering the whole current file (range caps continue to also mint a whole-file-scoped cap when the full file was observed, per the existing `wholeFileReadCapability` gates). Fold both into one uniform cap emission so the model carries one token. (Claimed after M1.1 validation; preserve observed-bytes floor while unifying model-facing capability emission.) (Validated: one capability per successful selector; common/SDK typechecks and 60 read-files tests passed.) - - Acceptance: whole-file overwrite from a partial-only observation is still refused; from a full observation it succeeds. - - Validate: same suite as M1.1. - -## M2 — Mutation results show new file state (D-B) - -- [x] M2.1 Include post-edit rendered content (bounded by the 10MB read ceiling) + a fresh whole-file cap.v3 per applied file in the model-visible `file_mutation_result`, via `edit-application-coordinator.ts` + `change-file.ts` + `filesystem.ts` result schema. (Claimed: add bounded post-edit file state and fresh capability to model-visible mutation results.) (Validated: action-local exact afterContent with afterHash correlation; 47 targeted tests and all affected typechecks passed.) - - Acceptance: AC2 — after an applied edit the tool result shows new content + usable cap; a follow-up edit needs no re-read. - - Validate: `cd sdk && bun run typecheck && bun test src/__tests__/change-file.test.ts src/__tests__/replace-range.test.ts` + agent-runtime edit-application-coordinator test. - -## M3 — Fix the editor subagent (D-C) - -- [x] M3.1 Ensure the editor returns `completed` with correct `changedFiles` whenever mutations applied; reconcile against actual mutation receipts instead of emitting `blocked`/null. (Common receipt reconciliation now preserves receipt-correlated handler state; editor caller implementation and validation remain pending outside this gate's writable files.) (Active repair task; dependent caller scope must include sdk/src/run.ts when validation names it.) (Validated receipt-backed completed status, changedFiles, non-null reconciled output, and forgery rejection.) - - Acceptance: AC3. - - Validate: `cd agents && bun run typecheck && bun test __tests__/editor.test.ts` - -## M4 — Remove legacy read/edit compatibility (D-D) - -- [x] M4.1 Remove cap.v2 (pathless) tokens + object-form basedOnRead + `expectedHash` legacy replace_range form + legacy path-keyed read-result map (`LegacyReadFilesMap`). Collapse `basedOnRead` schema to the single cap.v3 string. (`replace_range` and SDK path-keyed override compatibility are removed; remaining repository surfaces require the full M4 validation gate.) (Move M4.1 back to pending so M3.1 is the only active task; dependent caller scope must include sdk/src/run.ts when validation names it.) (Active: remove legacy read/edit forms and update canonical callers including sdk/src/run.ts.) (validated cap.v3/structured-only model and absence proof) - - Acceptance: AC4 (auth forms). - - Validate: common + agent-runtime + sdk typechecks; content-hash/based-on-read/edit-transaction schema tests. -- [x] M4.2 Remove quarantined dead tools (`read_slices`, `apply_smart_patch`) and their schemas/handlers/registrations/type surface if unreferenced by shipped agents. (claimed after M4.1 validation) (read_slices + apply_smart_patch removed across schemas/handlers/registrations/type surface + all residual source refs (tool-execution-deadline, tool-metadata test, cli codebuff-client, docs, structural-read test); typecheck clean, 33/33 targeted tests green.) - - Acceptance: AC4 (dead tools). - - Validate: common typecheck + tool-registration-consistency + tool-metadata tests + agents typecheck. - -## M5 — Full validation + docs - -- [x] M5.1 Update `docs/deterministic-edit-system.md` to describe the single unified model. - - Validate: configured hooks. -- [x] M5.2 Full monorepo typecheck + all touched test suites green. (Full typecheck clean; 330 read/edit/auth tests green across agent-runtime/SDK/common.) - - Validate: `bun run typecheck` (root) + the union of suites above. - -## Open decision (needs user before M1) - -- ODA: Resolved in favor of the security-preserving interpretation: whole-file overwrite requires a whole-file-covering hash; range edits require a matching hash for the observed range; both use one uniform cap.v3 validation path. A partial range capability never authorizes rewriting unobserved bytes. - - - -## M1.1 validation — 2026-07-22T08:45:45.926Z - -M1.1 validated: agent-runtime typecheck passed; 178/178 tests passed across process-str-replace, process-edit-transaction, and read-files-edit-state. Source verification showed cap.v3 scope + current-range hash are already the single in-range authority decision; only clarifying documentation was required. - - - -## M1.2 validation — 2026-07-22T09:07:51.630Z - -M1.2 validated: removed range-derived wholeFileReadCapability from SDK producer, common result schema, and replace_range guidance. Updated read-files tests so range reads expose only their own token and whole-file-to-subrange tests start from a genuine whole-file editAnchor. Common + SDK typechecks passed; 60/60 SDK read-files tests passed. Remaining string references are negative regression assertions only. - - - -## M2.1 validation — 2026-07-22T09:27:19.403Z - -M2.1 validated: model-visible file_mutation_result actions now include exact afterContent for applied text create/update/move actions, hash-correlated to afterHash; deletes/failures omit it. Common typecheck + 18 filesystem result tests, SDK typecheck + 21 change-file tests, and agent-runtime typecheck + 8 coordinator tests passed. Validation exposed and fixed a pre-existing move receipt inconsistency: committed move actions must use the independently verified destination final hash, not null. - - - -## Repair scope and caller boundary lesson — 2026-07-22T10:12:24.819Z - -The repair editor was blocked because validation named sdk/src/run.ts outside its writable handoff scope. The attempted fallback to widen read-files.ts would have violated AC4 by restoring legacy compatibility; dependent callers must be updated to the canonical structured result instead. - - - -## M3.1 completion evidence — 2026-07-22T10:22:14.675Z - -RF-2/7/11/16 implemented and validated. Runtime mutation attestations now override stale blocked/null editor output only when no receipt errors exist, producing a completed AgentReceipt with authoritative changedFiles and a non-null reconciled output. Batch, background, and inline spawn surfaces return receipt.output. Agent-runtime and agents typechecks passed; 24 spawn receipt tests and 43 editor tests passed, including blocked/null mutation cases and forgery rejection. - - - -## M4.1 completion — 2026-07-22T13:43:35.097Z - -M4.1 completed. Unified model now uses structured ReadFilesResultV1 overrides and scoped cap.v3 string inputs. replace_range freshness is derived from the capability; optional contained target bounds are checked in original snapshot coordinates and shifted only for application. Mutation afterContent/fresh-capability correlation uses byte-exact receipt hashes while cap.v3 tokens retain normalized read hashes, including CRLF coverage. Validation: common typecheck + 108 tests, SDK typecheck + 57 tests, agent-runtime typecheck + 231 tests passed. Production-only scan found no active cap.v2, legacy read-map, object-input basedOnRead, wholeFileCapabilityHash, or replace_range expectedHash compatibility implementation. diff --git a/.agents/sessions/read-edit-auth-unification-2026-07/SPEC.md b/.agents/sessions/read-edit-auth-unification-2026-07/SPEC.md deleted file mode 100644 index 9d3437fae8..0000000000 --- a/.agents/sessions/read-edit-auth-unification-2026-07/SPEC.md +++ /dev/null @@ -1,44 +0,0 @@ -# SPEC — Unify read/edit authorization, fix mutation results, fix editor subagent, remove legacy compat - -## Goal (user, verbatim intent) - -"Make a whole file and read range authorization the same so edits apply regardless of either as long as the content is correct. And fix the mutation results so it shows you the new file state. And fix the editor subagent. Also get rid of the legacy compatibility in this and any other tools that have legacy compatibility. We don't need legacy compat — one uniform best implementation, update anything relying on legacy to work with it. There are no third parties. This is a self-contained CLI." - -## Confirmed design decisions - -- **D-A (auth unification):** Content correctness is the authority for edits **within the content observed by the capability**, subject to the observed-bytes floor. Normative rule (identical in SPEC, PLAN, and STATUS): a fresh authenticated `cap.v3` capability bound to (project, path, run) authorizes an edit **only within the byte range it observed** when the targeted current content matches the capability's hash. Whole-file reads and range reads use the same content-hash validation path within their observed scope, so an edit applies regardless of which read produced the capability — but a partial range capability can NEVER authorize a whole-file overwrite, because it does not establish authority over unobserved bytes. A whole-file overwrite requires a capability covering the whole current file. A whole-file read mints a cap.v3 over lines 1..N; a range read mints a cap.v3 over its lines. -- **D-B (mutation results show new state):** `file_mutation_result` returned to the model includes, per applied file, the post-edit rendered content (bounded by the existing 10MB read ceiling) plus a fresh whole-file `cap.v3` readCapability, so the model both sees new state and can immediately chain edits without a re-read. -- **D-C (editor subagent):** The editor must return a completed structured receipt with `changedFiles` whenever its edits actually applied; it must not return `blocked`/null when the filesystem shows applied mutations. Reconcile the receipt against actual mutation receipts. -- **D-D (legacy removal):** Remove cap.v2 (pathless) tokens, object-form `{startLine,endLine,hash}` basedOnRead, legacy path-keyed read-result maps (`Record`), the `expectedHash` legacy replace_range form, and the now-redundant separate `wholeFileReadCapability` field (subsumed by D-A). Delete quarantined dead tools (`read_slices`, `apply_smart_patch`) and their schemas/handlers/registrations if unreferenced by shipped agents. Update every dependent call site + test to the single cap.v3 path. - -## Non-goals - -- No change to the reviewer/validation gate semantics (separate, already-fixed subsystem). -- No provider/model routing changes. - -## Key systems (source-backed) - -- `common/src/util/content-hash.ts` — cap.v2/v3 encode/decode, scope fingerprint. (auth token core) -- `common/src/tools/params/based-on-read.ts` — basedOnRead schema (string | object union). -- `packages/agent-runtime/src/process-str-replace.ts` — `normalizeBasedOnRead`, `validateReadCapabilityAuthority`, `getReadCapabilityKey` (legacy branches). -- `packages/agent-runtime/src/process-edit-transaction.ts` — replace_range capability resolution, whole-file-sub-range branch. -- `packages/agent-runtime/src/tools/handlers/tool/write-file.ts` — whole-file read-authorization state (`grant/has/isFresh/getUsableWholeFileAuthorizationHash`). -- `sdk/src/tools/read-files.ts` — capability minting (`encodeReadCapabilityToken`, `wholeFileReadCapability`), render. -- `sdk/src/tools/change-file.ts` + `common/src/tools/results/filesystem.ts` — mutation result shape (`file_mutation_result`, before/afterHash, freshCapabilities). -- `packages/agent-runtime/src/tools/handlers/tool/edit-application-coordinator.ts` — reconciles client mutation output back to the model. -- `agents/editor/editor.ts` — editor subagent receipt/changedFiles logic. -- Legacy/dead: `common/src/tools/params/tool/read-slices.ts`, `apply-smart-patch.ts`, `common/src/tools/constants.ts` (`quarantinedToolNames`, `publishedTools`), `common/src/types/contracts/client.ts` (`LegacyReadFilesMap`). - -## Acceptance criteria - -- AC1: A range-read cap.v3 authorizes an edit within its observed range when the targeted content hash matches current content; it cannot authorize a whole-file overwrite unless the capability covers the whole current file. A whole-file cap authorizes a matching sub-range or whole-file edit. One validation path applies within the capability's observed scope, with the observed-bytes floor preserved. -- AC2: After any applied edit, the model-visible tool result contains the new file content (bounded) + a fresh usable cap.v3 for the edited path. -- AC3: The editor subagent returns `completed` with correct `changedFiles` whenever mutations applied. -- AC4: No cap.v2, object-form basedOnRead, legacy path-keyed read map, or `expectedHash` legacy form remains in the active code path; `read_slices`/`apply_smart_patch` removed if unreferenced. -- AC5: All package typechecks pass and the read/edit test suites pass (updated to the unified model). - -## Risks - -- R1: Auth core rewrite has the widest blast radius in the repo; heavy test churn in `process-str-replace.test.ts`, `read-files-edit-state.test.ts`, `content-hash.test.ts`, `edit-transaction.schema.test.ts`, `process-edit-transaction.test.ts`. -- R2: Removing the legacy path-keyed read-result parser could break resuming persisted chat histories authored before cap.v3. Self-contained CLI + user directive says remove; flag if a persisted-format migration is needed. -- R3: Embedding post-edit content in mutation results increases context cost for large files — bounded by the existing 10MB read ceiling and only for edited files. diff --git a/.agents/sessions/read-edit-auth-unification-2026-07/STATE.json b/.agents/sessions/read-edit-auth-unification-2026-07/STATE.json deleted file mode 100644 index 75a9474e18..0000000000 --- a/.agents/sessions/read-edit-auth-unification-2026-07/STATE.json +++ /dev/null @@ -1,17 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "read-edit-auth-unification-2026-07", - "status": "completed", - "currentTask": null, - "revision": 17, - "checkpoint": { - "taskId": "M5.2", - "phase": "validation", - "passed": true, - "summary": "Full monorepo typecheck clean; 231 agent-runtime + 57 SDK + 42 common read/edit/auth tests green.", - "receiptIds": ["61f91818-45a7-4f47-b99a-791ca8753c55"], - "recordedAt": "2026-07-22T14:36:28.088Z" - }, - "createdAt": "2026-07-22T07:20:12.216Z", - "updatedAt": "2026-07-22T14:38:19.432Z" -} diff --git a/.agents/sessions/read-edit-auth-unification-2026-07/STATUS.md b/.agents/sessions/read-edit-auth-unification-2026-07/STATUS.md deleted file mode 100644 index 3067f04e28..0000000000 --- a/.agents/sessions/read-edit-auth-unification-2026-07/STATUS.md +++ /dev/null @@ -1,81 +0,0 @@ -# STATUS — Read/Edit Authorization Unification - -## Current state - -- **Phase:** Session claims M1-M5 complete, but completion is not treated as verified evidence in this reviewed snapshot; the implementation and direct test artifacts must be included in the review scope before those claims are accepted. No current task is active. -- **Confirmed design decision (user, 2026-07):** Unify whole-file and range read - authorization so an edit applies regardless of which produced the authority, - **as long as the supplied content matches current bytes** — BUT keep the - observed-bytes safety floor: a partial/range read may authorize edits within - what was observed and may not mint whole-file authority for content the model - never saw. Interchangeability is about "content-correct edits apply"; it is - NOT "a 3-line read authorizes rewriting the whole file." - -## Scope (4 sub-goals from the user request) - -1. **Unify authorization** — whole-file authorization and cap.v3 range - capability become one interchangeable content-correctness check on the edit - path, subject to the observed-bytes floor above. -2. **Mutation results show new file state** — `edit_transaction` / edit results - echo the resulting file state (or a fresh whole-file editAnchor/content - marker) so the model can see what it just wrote without a re-read. -3. **Fix the editor subagent** — the `status: blocked` / null-receipt pattern - when its first transaction applied but a follow-up hit strict-read; make the - editor return a coherent completed receipt reflecting applied changes. -4. **Remove legacy read/edit compatibility** — drop cap.v2/base64 legacy - tokens, the legacy `{startLine,endLine,hash}` object form, and legacy - path-keyed override normalization in favor of one uniform cap.v3 + structured - implementation; update every dependent call site + test. - -## Session-reported completion claims - -The following are session history claims, not independently verified evidence for the current reviewed snapshot: - -- M1 authorization unification. -- M2 post-edit content and fresh capability results. -- M3 editor receipt reconciliation and authoritative `changedFiles` behavior. -- M4 cap.v3-only authorization, structured read results, and dead-tool removal. -- M5 documentation and validation gates. - -Review acceptance requires the corresponding implementation and direct test files to be present in the reviewed snapshot. - -## Pending - -- Include or review the implementation and direct test artifacts for AC1-AC5 before treating the session completion claims as verified. -- Add direct formatter coverage for missing `str_replace` replacement fields and mixed validation issues. - -## Blocked / needs user decision - -- None; the authorization design fork is resolved with the observed-bytes floor kept. - -## Next checkpoint - -- None; the session is complete. - -## Resume instructions - -- None; there is no active current-task pointer. Preserve the observed-bytes floor and the completed cap.v3-only model in future work. - - - -## M1.1 Finding — str_replace path already unifies (2026-07) — 2026-07-22T07:36:28.622Z - -Ground truth after reading validateReadCapability + validateReadCapabilityAuthority in process-str-replace.ts: the str_replace authority path ALREADY treats whole-file and range capabilities uniformly. validateReadCapability re-hashes the current [startLine,endLine] slice and compares to the capability hash with NO whole-file-vs-range branch; validateReadCapabilityAuthority checks only cap.v3 scope. So M1.1's 'remove whole-file-vs-range authority branching in process-str-replace.ts' is a no-op there by design — there is no such branch to remove. The editor correctly landed a doc-comment-only change documenting the authenticity-vs-content-correctness split (typecheck green, exit 0). The remaining real M1.1 logic surface is narrower than the plan assumed: it lives only in the replace_range / process-edit-transaction.ts whole-file-sub-range path, and that path is the anti-footgun FLOOR the user chose to KEEP (option 2), so it must not be collapsed. Net: M1.1 in the str_replace path is satisfied-by-documentation; edit-transaction needs verification, not necessarily change. - - - -## AC2 reconciliation validation — 2026-07-22T10:15:36.520Z - -RF-1/5/6/10/14/15 verified in live source: reconcileFileMutationResultV1 preserves handler afterContent only when action identity and afterHash correlate with the matching receipt, and retains fresh capabilities only when their snapshot hash is a committed receipt afterHash. Common typecheck and 19 filesystem result tests passed, including `preserves hash-correlated handler content and fresh capabilities with a matching receipt`. - - - -## M4.2 completion — dead-tool removal — 2026-07-22T14:31:56.376Z - -M4.2 validated. `read_slices` and `apply_smart_patch` fully removed: common params/tool files deleted, handler files deleted, list/metadata/constants/input-aliases registrations removed, agents/types/tools.ts and template tools.ts updated, sdk/tool-execution-deadline.ts deadline map cleaned, tool-metadata test updated, structural-read test capability assertions updated to reflect scope-gated minting, docs/agents-and-tools.md sections removed, cli/codebuff-client.ts `filesystemResultFormat` removed. Full monorepo `bun run typecheck` clean; tool-metadata + structural-read + read-outline-slices + tool-reachability suites green (33/33). Next: M5.1 docs update. - - - -## Session complete — 2026-07-22 — 2026-07-22T14:38:19.432Z - -Session history reports all milestones M1–M5 validated, including M4.2 removal of the quarantined dead tools (`read_slices`, `apply_smart_patch`) and their schemas/handlers/registrations/type surface. These historical validation statements are not independent evidence for the current reviewed snapshot; the referenced implementation and test files must be included in the review scope. diff --git a/.agents/sessions/read-edit-pipeline-create-capability/LESSONS.md b/.agents/sessions/read-edit-pipeline-create-capability/LESSONS.md deleted file mode 100644 index 0eef74e5d8..0000000000 --- a/.agents/sessions/read-edit-pipeline-create-capability/LESSONS.md +++ /dev/null @@ -1,29 +0,0 @@ -# LESSONS / DECISIONS - -## Medium tier (M3) — implemented and validated - -- #4: `coordinateEditApplication` applied branch appends one additive `postEditCapabilities` json part (path + contentHash + cap.v3 readCapability) only when anchors are granted. This output reaches the model via message history but not the user-facing CLI rows (which render separately), so the token is model-visible only. Immutably appended; error/rejected/threw paths untouched. -- #5: `strictEditAuthorizationError` gained a created/edited-this-session cause branch (fires when a `confirmedPostEditAnchorsByPath[path]` anchor exists and no higher-precedence cause applies) plus an `effectiveFreshReadCapability` fallback echoing that anchor's token for recovery/basedOnRead. -- #6 was delivered by the High tier: `commitAppliedEditPaths` grants sticky auth + clears reread markers from runtime-known content even without a client-echoed anchor, while preserving the marker for blind allowMultiple replace-alls via `preserveRereadRequirementsForPaths`. -- Gotcha: improving recovery-message wording (#5) broke a pre-existing test that asserted brittle message text (`/context compaction|read_files/i`). The test's intent (write_file stays blocked + marker preserved) still held; updated the assertion to match the new, more accurate wording. Lesson: prefer asserting behavior/intent over exact message strings for user-facing recovery text. -- Tests: 117/117 pass across edit-application-coordinator + read-files-edit-state; monorepo typecheck clean. - -## M4 Lower tier (#7–#8) — DESIGN DECISION: DEFER BOTH (design-doc only, per user) - -Decision recorded 2026-07-28. No implementation authorized. Rationale below. - -### #7 — Narrow the whole-file-anchor requirement for the _grant_ (not the _edit_) - -- Idea: a confirmed-but-scoped anchor (post-edit bytes known but anchor covers a region) grants read authorization for that region only, recorded distinctly from a whole-file grant. -- Risks/open questions: a scoped grant must never be promoted to whole-file authorization by a later blind apply; must not clear `context_compacted` for whole-file-overwrite purposes; must not authorize delete/move of the whole file (which #3 keys off a whole-file anchor hash match). `confirmedPostEditAnchorsByPath` currently implies whole-file, so a scoped grant needs a separate/tagged representation. -- Recommendation: DEFER. Marginal benefit over current High/Medium behavior (which already grants sticky from runtime-known whole-file content on confirmed applies) is small; fail-closed risk surface is meaningful. Revisit only if a concrete workflow is blocked by scoped-anchor-only grants. - -### #8 — First-class "created this session" state - -- Idea: a `createdThisSessionByPath` (or `knownContentOrigin: 'agent-written' | 'read' | 'confirmed-apply'`) marker so the gate can relax delete/move/overwrite for agent-created files on principle rather than via the anchor-hash heuristic. -- Risks/open questions: external modification (another process can still change an agent-created file on disk) — the marker MUST be invalidated on any hash mismatch against the live snapshot, else it fails open; reviewer-gate working-tree markers (`sha256::`) must not desynchronize; must not become a blank check for whole-file overwrite after compaction (intentionally blocked today); needs a per-turn hydration lifecycle like readAuthorizationsByPath. -- Recommendation: DEFER. Largest design surface in the backlog (state shape, compaction, gate markers). The Medium tier's created/edited-this-session recovery message already captures most of the user-facing benefit without new state. Observe whether High/Medium removes the practical friction first. - -### If either is later approved - -Specify which item and the safety constraints: for #7, no whole-file-overwrite promotion and a distinct scoped-grant representation; for #8, external-modification invalidation and gate-marker synchronization. Then produce a full implementation plan with state-shape and gate review before any code change. diff --git a/.agents/sessions/read-edit-pipeline-create-capability/PLAN.md b/.agents/sessions/read-edit-pipeline-create-capability/PLAN.md deleted file mode 100644 index 8a0cbc7810..0000000000 --- a/.agents/sessions/read-edit-pipeline-create-capability/PLAN.md +++ /dev/null @@ -1,55 +0,0 @@ - - -# Plan: Create→Edit/Delete Capability & Read/Edit Pipeline Improvements - -## Problem - -`edit_transaction` `create` (and `write_file`) do not reliably mint or surface a reusable read capability, so a follow-up `delete` / `str_replace` / `write_file` on a session-created file is blocked by strict read-before-edit and forces a redundant `read_files` round-trip. This is the exact friction seen in the wild (agent had to `read_files` a throwaway file it had just created before it could `delete` it). - -## Root cause (verified against source) - -- `edit-transaction.ts` strict gate exempts `create` (`if (edit.type === 'create' && initialContentByPath.get(edit.path) === null) return`), so no `freshWholeFileAuthorizationPaths` entry is established at preflight. -- Post-edit sticky authorization is only granted by `commitAppliedEditPaths` (`edit-application-coordinator.ts`), and only when **all** of: `wholeFileContentByPath.get(path)` is a string, a `confirmedAnchor` exists, and `strictReadBeforeEdit` is on. -- `confirmedAnchor` is only produced by `getPositiveApplicationEvidence` when the client echoes a post-edit `editAnchor` that passes a 7-point check: whole-file covering (`startLine===1`, `endLine===totalLines`), `contentHash === getContentHash(content)`, valid `cap.v3` decode, scope match on `{projectId, path, runId}`, and decoded bounds/hash equal to the record. -- The granted authorization lives in invisible in-memory maps (`readAuthorizationsByPath` / `confirmedPostEditAnchorsByPath`); the tool output (`1. path • create • applied`) never surfaces the minted `readCapability` to the model. -- `delete` is **not** exempted the way `create` is; it falls through the strict gate and hits the generic `Edit blocked: strict read-before-edit is enabled and no fresh read authorization exists`. - -## Key files - -- `packages/agent-runtime/src/tools/handlers/tool/edit-transaction.ts` — strict gate, lifecycle preflight, `wholeFileContentByPath` population, `coordinateEditApplication` wiring. -- `packages/agent-runtime/src/tools/handlers/tool/edit-application-coordinator.ts` — `getPositiveApplicationEvidence`, `commitAppliedEditPaths`, `coordinateEditApplication`, `invalidatePreparedEditPaths`. -- `packages/agent-runtime/src/tools/handlers/tool/write-file.ts` — `FileProcessingState`, `grantWholeFileReadAuthorization`, `hasWholeFileReadAuthorization`, `isWholeFileReadAuthorizationFresh`, `revokeWholeFileReadAuthorization`, `normalizeToolPath`. -- `packages/agent-runtime/src/tools/handlers/tool/edit-read-state.ts` — `markEditRequiresFreshRead`, `clearEditRereadRequirement`, `strictEditAuthorizationError`. -- `docs/deterministic-edit-system.md` — policy documentation to keep in sync. - -## Improvement backlog (see SPEC.md for full detail) - -- High: #1 server-side mint + surface capability on create/write_file; #2 create grants sticky auth directly; #3 relax delete auth on fresh confirmed anchor. -- Medium: #4 echo capability in every mutating tool result; #5 differentiate recovery messages; #6 decouple marker-clearing from anchor check for known-content creates. -- Lower: #7 narrow whole-file-anchor requirement for the grant (not the edit); #8 first-class "created this session" state. - -## Milestones - -- [ ] M1 — Spec & design (SPEC.md) capturing all 8 improvements, security/compaction interplay, and acceptance criteria. -- [ ] M2 — (future, gated on user approval) Implement High tier (#1–#3) with tests. -- [ ] M3 — (future) Medium tier (#4–#6). -- [ ] M4 — (future) Lower tier (#7–#8) design decision + optional implementation. - -## Resolution (2026-08-03) - -**Session closed — substantially implemented by other work.** Verified against live source: - -- **#1** server-side mint + surface: `synthesizePostEditAnchor` (edit-application-coordinator.ts:287-310) mints cap.v3 from known content; `coordinateEditApplication` appends `postEditCapabilities` to model-facing output (lines 496-509). ✅ -- **#2** create grants sticky auth: `commitAppliedEditPaths` (lines 351-377) grants from runtime-known create bytes without client anchor. ✅ -- **#3** relax delete auth: strict gate delete/move branch (edit-transaction.ts:507-537) authorizes on fresh whole-file confirmed anchor hash-matched to snapshot. ✅ -- **#4** echo capability: done via `postEditCapabilities`. ✅ -- **#6** decouple marker-clearing: done in `commitAppliedEditPaths`. ✅ -- **#5** differentiated recovery messages: partial (stale vs never-read vs compacted, but no 'created this session' message). ⚠️ -- **#7/#8**: not implemented (lower tier, intentional). ❌ - -Milestones M2–M4 as written are moot. Remaining polish (#5 message wording, #7/#8 design) can be reopened as a fresh scoped task if desired. - -## Validation gates - -- Spec reviewed for correctness against the four source files above. -- Implementation milestones (future) must add/adjust tests under `packages/agent-runtime/src/**/__tests__` and pass `packages/agent-runtime` typecheck + targeted tests, and keep `docs/deterministic-edit-system.md` in sync. diff --git a/.agents/sessions/read-edit-pipeline-create-capability/SPEC.md b/.agents/sessions/read-edit-pipeline-create-capability/SPEC.md deleted file mode 100644 index 1306e2ecec..0000000000 --- a/.agents/sessions/read-edit-pipeline-create-capability/SPEC.md +++ /dev/null @@ -1,128 +0,0 @@ -# SPEC: Create→Edit/Delete Capability & Read/Edit Pipeline Improvements - -Status: design only. No code changes are authorized by this document. Implementation is split into future milestones (M2–M4) and requires explicit user sign-off per tier. - -## 1. Background & problem statement - -The deterministic edit system enforces staged read-before-edit: under strict mode, an edit to an existing path requires either a fresh whole-file read authorization or an explicit scoped capability (`basedOnRead` / `readCapability`). This is correct and fail-closed. - -The friction: an agent that **creates** a file knows its exact bytes (it supplied them), yet the runtime does not reliably convert that known content into reusable edit authorization or hand the model a capability token. A subsequent `delete` / `str_replace` / `write_file` on that same file is then blocked, and the model is forced into a redundant `read_files` round-trip. Observed in a real session: an agent created a throwaway smoke-test file, then had to `read_files` it before `edit_transaction` `delete` would apply. - -The system is choosing correctness over convenience. The goal of this work is to recover the convenience **without** weakening the fail-closed guarantee for genuinely-unknown files. - -## 2. Verified mechanism (source of truth) - -File: `packages/agent-runtime/src/tools/handlers/tool/edit-transaction.ts` - -- Strict gate exempts create: `if (edit.type === 'create' && initialContentByPath.get(edit.path) === null) return`. So no `freshWholeFileAuthorizationPaths` entry is established for a create at preflight. -- `delete`/`move` are validated only for existence in lifecycle preflight (`Delete source does not exist`), but are **not** exempted from the strict gate — they fall through to the generic authorization check. -- `wholeFileContentByPath` is populated for `create` (`set(edit.path, edit.content)`) and `move` (`set(destinationPath, sourceContent)`). - -File: `packages/agent-runtime/src/tools/handlers/tool/edit-application-coordinator.ts` - -- `commitAppliedEditPaths` grants sticky auth only when `typeof wholeFileContent === 'string' && confirmedAnchor && strictReadBeforeEdit`. -- `getPositiveApplicationEvidence` builds `confirmedAnchor` only when the applied action's `editAnchor` passes a 7-point check: content string present, `readCapability` string, `startLine===1`, `endLine===normalizeLineEndings(content).split('\n').length`, `contentHash === getContentHash(content)`, `cap.v3` decodes, `readCapabilityMatchesScope({projectId, path, runId})`, and decoded `startLine/endLine/hash` equal the record. It also cross-checks `matchingAction.afterHash === getExactContentHash(content)` for every `wholeFileContentByPath` entry. - -File: `packages/agent-runtime/src/tools/handlers/tool/write-file.ts` - -- `FileProcessingState` holds `readAuthorizationsByPath`, `readAuthorizationHashesByPath`, `confirmedPostEditAnchorsByPath`, `modelVisibleReadAuthorizationHashesByPath`, `editRereadRequirementsByPath`. -- `grantWholeFileReadAuthorization` writes the sticky hash from known content. - -### Why the gap happens (three distinct causes) - -1. **Anchor dependence.** The post-edit grant requires the client to echo a whole-file-covering `cap.v3` anchor. Post-edit anchors are optional ("may return"), and create/delete actions are the least likely to carry one. No anchor → no grant, even though the create content is known exactly. -2. **Invisibility.** Even when the grant succeeds, it is written to in-memory maps only. The tool output (`1. path • create • applied`) never surfaces the `readCapability`, so the model has no token to pass and no signal that a re-read is unnecessary. -3. **Asymmetric lifecycle handling.** `create` is exempt from the strict gate; `delete` is not. So the very next lifecycle op on a just-created file is the one that blocks. - -## 3. Goals / non-goals - -### Goals - -- Eliminate the redundant read-before-delete/edit round-trip for files whose content is already known to the runtime (creates, whole-file writes, confirmed edits). -- Make granted authorization explicit and compaction-resilient by surfacing capability tokens to the model. -- Preserve fail-closed behavior for any file whose current bytes are not positively known. - -### Non-goals - -- Weakening the 7-point `confirmedAnchor` evidence check for _confirming an apply happened_. -- Allowing `write_file` whole-file overwrite on scoped/partial capabilities (the current floor stays). -- Changing the `context_compacted` semantics for `write_file` (sticky hash alone still must not authorize a blind overwrite). -- Auto-reread behavior changes for `str_replace` (out of scope). - -## 4. Improvement backlog - -### HIGH tier - -#### #1 — Server-side mint + surface a capability on create / whole-file write - -- **What.** When a `create` (or a confirmed whole-file `write_file`) is applied, the runtime already has the exact post-edit bytes. Mint a `cap.v3` read capability from that known content server-side instead of depending on the client to echo an anchor, and return it in the tool output (structured `editAnchor.readCapability`, plus a short non-secret indicator in the human-readable line). -- **Where.** `edit-application-coordinator.ts` (`commitAppliedEditPaths` / `getPositiveApplicationEvidence` fallback) and the output shaping in `edit-transaction.ts` / `write-file.ts`. -- **Why high.** Directly removes the reported round-trip; the content needed is already in `wholeFileContentByPath` for creates. -- **Security note.** The minted token must be scope-bound to `{projectId, path, runId}` exactly like read-minted tokens, and must only be minted when the apply is positively confirmed. Server-side minting from known content is _stronger_ than trusting a client anchor, so this does not lower the security bar. -- **Compaction note.** A token visible in the transcript survives compaction; invisible in-memory maps accrue `context_compacted` markers. Surfacing the token therefore improves robustness under compaction. - -#### #2 — Create grants sticky authorization directly from known content - -- **What.** After a confirmed `create`, call `grantWholeFileReadAuthorization(fileProcessingState, path, edit.content)` unconditionally. The runtime supplied the bytes; it does not need client evidence to trust them. -- **Where.** `edit-transaction.ts` post-apply path (or `commitAppliedEditPaths` with a create-aware branch). -- **Why high.** Removes the dependence on the 7-point anchor check specifically for the create case, closing cause #1. -- **Interaction with #1.** #2 establishes the internal grant; #1 surfaces the token. They are independent but complementary; together they fully close the create gap. - -#### #3 — Relax delete authorization on a fresh confirmed post-edit anchor - -- **What.** A delete's safety bound is "the file is in the state I believe it is." If the path has a fresh confirmed post-edit anchor (e.g. from #2) whose `contentHash` matches current content, allow `delete` to proceed on that anchor without an additional whole-file read. -- **Where.** `edit-transaction.ts` strict gate: add `delete` (and evaluate `move`) to the set of edits that can be authorized by a matching `confirmedPostEditAnchorsByPath[path]` fresh against the snapshotted content. -- **Why high.** This is the exact operation that blocked in the observed session. -- **Risk / care.** Must still verify hash freshness against the transaction snapshot (`initialContentByPath`), and must still fail closed when there is no confirmed anchor or the hash is stale. `move` should be treated cautiously because it also touches a destination path. - -### MEDIUM tier - -#### #4 — Echo the granted capability in every successful mutating tool result - -- **What.** Generalize #1: whenever a mutation results in a granted whole-file authorization, include the capability (or a reference to it) in the structured tool output so the model can reuse it explicitly. -- **Why medium.** Turns implicit state into an explicit, transcript-resident token; reduces a whole class of "I edited it, why am I blocked" cases beyond create. -- **Care.** Keep user-facing CLI rows free of raw tokens (docs already require: show a short hash + whether a capability exists, never the token in CLI rows). The token belongs in the model-facing structured output only. - -#### #5 — Differentiate recovery messages by actual cause - -- **What.** The generic `Edit blocked: strict read-before-edit is enabled and no fresh read authorization exists` should distinguish: never-read vs. created-this-session vs. compacted vs. stale. After a create, the message should say "pass the capability from the create result" rather than "read the file first." -- **Why medium.** The current message actively misleads the model into an unnecessary full re-read. -- **Where.** `strictEditAuthorizationError` in `edit-read-state.ts` and the inline failure strings in `edit-transaction.ts`. - -#### #6 — Decouple reread-marker clearing from the strict anchor check for known-content creates - -- **What.** `commitAppliedEditPaths` already calls `clearEditRereadRequirement` per path, but the sticky grant is gated on the anchor check. For known-content creates, clear `failedEditRequiresReadByPath` / `editRereadRequirementsByPath` markers even when the client anchor is absent. -- **Why medium.** Removes a class of spurious blocks where the marker outlives a known-good create. - -### LOWER tier - -#### #7 — Narrow the whole-file-anchor requirement for the _grant_ (not the _edit_) - -- **What.** The 7-point check currently does two jobs: (a) prove the apply succeeded, (b) prove the anchor is whole-file. For granting _sticky read authorization_, arguably only (a) plus a trusted content hash is needed; a confirmed-but-scoped anchor could still authorize reads of that region. -- **Why lower.** Broader semantic change; needs a design decision and careful reasoning about what "scoped sticky" means for later whole-file overwrites. -- **Decision required.** Architect/owner sign-off before implementation. - -#### #8 — First-class "created this session" state - -- **What.** Distinguish "content known because I wrote it" from "content known because I read it" in `FileProcessingState`. The former can authorize more aggressively within the mutation broker's authority and gives the gate a principled reason to relax delete/overwrite for agent-created files. -- **Why lower.** Largest design surface; touches state shape, compaction semantics, and the reviewer gate's working-tree markers. Should follow, not precede, #1–#3. - -## 5. Cross-cutting concerns - -- **Fail-closed invariant.** Every relaxation must keep failing closed when current bytes are not positively known. Hash freshness against the live snapshot is always required. -- **Compaction.** Surfacing tokens (#1/#4) is the primary compaction-resilience lever. `context_compacted` blocking of blind `write_file` overwrites is intentionally preserved. -- **Security.** Server-side minting must reuse the existing `cap.v3` issuer and scope binding. No new trust in client-supplied anchors is introduced; #1 actually _reduces_ reliance on client anchors. -- **Reviewer/validation gate.** #8 interacts with the gate's `sha256::` working-tree markers; defer until the gate implications are designed. -- **Backwards compatibility.** Surfacing an extra structured field in tool output is additive and safe. Changing grant semantics is internal and must be covered by tests. - -## 6. Acceptance criteria (per tier, for future implementation) - -- **High (#1–#3).** An agent can `create` a file and then `delete`, `str_replace`, or `write_file` it in a later step **without** an intervening `read_files`, in strict mode, provided no external modification occurred. A new regression test reproduces the originally-observed create→delete flow and asserts no spurious block. External modification still blocks. -- **Medium (#4–#6).** Mutating tool results expose a reusable capability to the model; recovery messages name the actual cause; known-content creates clear stale reread markers. Covered by targeted tests. -- **Lower (#7–#8).** Documented design decision; implementation only after explicit approval, with full state-shape and gate review. - -## 7. Testing & docs (for future implementation) - -- Tests under `packages/agent-runtime/src/**/__tests__` (see existing `read-files-edit-state.test.ts`, `edit-application-coordinator.test.ts`, `write-file.test.ts`). -- Add a regression test that mirrors the observed session: create → delete with no read, assert success in strict mode; and create → external-change → delete, assert block. -- Update `docs/deterministic-edit-system.md` to document the new create/delete authorization behavior and the surfaced capability field. diff --git a/.agents/sessions/read-tool-unification-2026-07/EVENTS.jsonl b/.agents/sessions/read-tool-unification-2026-07/EVENTS.jsonl deleted file mode 100644 index 4e8b6eec74..0000000000 --- a/.agents/sessions/read-tool-unification-2026-07/EVENTS.jsonl +++ /dev/null @@ -1,5 +0,0 @@ -{"ts":"2026-08-02T17:34:06.608Z","kind":"append_lesson","summary":"Appended entry \"M1 + M2 complete (gate-verified)\" to STATUS.md","payload":{"heading":"M1 + M2 complete (gate-verified)","artifact":"STATUS.md"}} -{"ts":"2026-08-02T17:34:06.608Z","kind":"session_status","summary":"Session status -> executing","payload":{"status":"executing"}} -{"ts":"2026-08-02T20:18:50.156Z","kind":"append_lesson","summary":"Appended entry \"M3 complete: read_blocks fully removed (gate-verified)\" to STATUS.md","payload":{"heading":"M3 complete: read_blocks fully removed (gate-verified)","artifact":"STATUS.md"}} -{"ts":"2026-08-02T21:34:04.265Z","kind":"append_lesson","summary":"Appended entry \"Session closed — all milestones complete\" to STATUS.md","payload":{"heading":"Session closed — all milestones complete","artifact":"STATUS.md"}} -{"ts":"2026-08-02T21:34:04.265Z","kind":"session_status","summary":"Session status -> completed","payload":{"status":"completed"}} diff --git a/.agents/sessions/read-tool-unification-2026-07/PLAN.md b/.agents/sessions/read-tool-unification-2026-07/PLAN.md deleted file mode 100644 index 855d3fe9a8..0000000000 --- a/.agents/sessions/read-tool-unification-2026-07/PLAN.md +++ /dev/null @@ -1,310 +0,0 @@ -# PLAN — Unify the read tool surface onto `read_files` - - - -Three milestones. M1 is correctness-only with no public schema change. M2 is the contract -change. M3 is deprecation/cleanup. Each milestone is independently shippable and gated. - -Validation commands used throughout: - -- `cd packages/agent-runtime && bun run typecheck && bun test src/__tests__/read-blocks.test.ts src/__tests__/read-files-edit-state.test.ts` -- `cd common && bun run typecheck && bun test src/tools` -- `cd sdk && bun run typecheck && bun test src/__tests__/read-files.test.ts` -- `cd cli && bun run typecheck && bun test src/components/tools` -- `cd agents && bun run typecheck && bun test __tests__` -- repo-wide: `bun run typecheck` - ---- - -## M1 — Correctness: authority ladder + bounds + ecosystem wiring (no schema change) - -### M1-T1 — Extract the shared coverage→authority ladder - -New file `packages/agent-runtime/src/tools/handlers/tool/read-authority-ladder.ts`. - -```ts -export type ReadBlockAuthority = 'whole_file' | 'scoped' | 'none' - -export type ReadBlockCoverage = { - complete: boolean - startLine: number - endLine: number - totalLines: number - /** Exact undecorated normalized text. Numbered display content is NOT accepted. */ - sourceContent: string | undefined - /** False for heuristic regex symbol slices (no parser proof). */ - capabilityEligible?: boolean -} - -/** - * Single source of truth for "what does this observed block authorize". - * whole_file requires ALL of: complete, startLine === 1, endLine === totalLines, - * a real sourceContent string. Anything else is scoped at best. Fails closed. - */ -export function classifyReadBlockAuthority( - c: ReadBlockCoverage, -): ReadBlockAuthority { - if (!c.complete) return 'none' - if (c.capabilityEligible === false) return 'none' - if (c.sourceContent === undefined) return 'none' - if (c.startLine === 1 && c.endLine === c.totalLines && c.totalLines > 0) - return 'whole_file' - return 'scoped' -} -``` - -Do NOT move `grantWholeFileReadAuthorization` — it stays in `write-file.ts`; the ladder only -decides, callers act. - -### M1-T2 — Rewire `read_files` onto the ladder - -In `packages/agent-runtime/src/tools/handlers/tool/read-files.ts` (~lines 215–290) replace -the two hand-rolled branches (`selector === 'file'` and the `selector === 'range'` + -`startLine === 1 && endLine === totalLines` check) with one loop calling -`classifyReadBlockAuthority`. Behavior must be byte-identical to today: - -- `'whole_file'` → `wholeFileGrantPaths.add(path)`, delete - `confirmedPostEditAnchorsByPath[path]`, and when `strictReadBeforeEdit` call - `grantWholeFileReadAuthorization(fileProcessingState, path, sourceContent)`. -- `'scoped'` / `'none'` → no grant. -- The existing `context_compacted` post-loop rule (preserve unless in `wholeFileGrantPaths`) - is unchanged. - -For the `file` selector, pass `sourceContent: result.content` and -`{startLine: 1, endLine: , totalLines: }` so a complete -whole-file read classifies as `whole_file` exactly as before. - -### M1-T3 — `read_blocks` grants whole-file authority on full coverage (R1) - -In `read-blocks.ts`: - -- import `classifyReadBlockAuthority`, `grantWholeFileReadAuthorization`. -- Track `const wholeFileGrantPaths = new Set()` alongside `successfulReadPaths`. -- After building each `window` / `around` item, classify with the block's real - `sourceContent` and `totalLines`; on `'whole_file'` add to `wholeFileGrantPaths`, delete - `confirmedPostEditAnchorsByPath[path]`, and grant when `strictReadBeforeEdit`. -- Symbol slices pass `capabilityEligible: Boolean(slice.readCapability)` and never classify - as whole_file unless they genuinely span 1..totalLines. -- Replace the unconditional `context_compacted` `continue` with the same - `!wholeFileGrantPaths.has(path)` guard `read_files` uses. - -### M1-T4 — Bound `windowSize` / `contextLines` (R3) - -In `common/src/tools/params/tool/read-blocks.ts` add shared caps next to the existing -defaults and export them for reuse in M2: - -```ts -export const MAX_WINDOW_SIZE = 5_000 // lines -export const MAX_CONTEXT_LINES = 2_000 // lines per side -``` - -Apply `.max(MAX_WINDOW_SIZE)` / `.max(MAX_CONTEXT_LINES)` and mention the cap in each -`.describe()`. - -In `read-blocks.ts`, after slicing a block, check its byte length against -`MAX_RANGE_READ_BYTES` (re-export it from `common` or mirror the constant — do NOT import -`sdk` into `agent-runtime`; the value already lives in `sdk/src/tools/read-files.ts:41`, so -introduce `MAX_READ_BLOCK_BYTES = 4_194_304` in the `common` read-blocks params module and -have the SDK constant reference it in M2). Over budget → emit -`status:'error'`, `code:'too_large'`, `recovery:'read_smaller_range'`, no `editAnchor`. - -### M1-T5 — Context-pruner recognizes `read_blocks` (R7) - -In `agents/context-pruner.ts`: - -- add `'read_blocks'` to the tool list at ~line 161; -- add a `case 'read_blocks':` next to `case 'read_files':` (~line 229) summarizing - `windows: path:win`, `around: path@match#occ`, `symbols: path#name`, and ending with the - same "(re-fetch with … if needed)" pointer, naming `read_blocks`; -- extend the `kind === 'read_files_result'` failure-text collector (~line 962) to also match - `'read_blocks_result'`; -- extend the inspection-path tracking at ~1062/~1105 to accept `read_blocks` inputs and - `read_blocks_result` values. - -### M1-T6 — CLI renderer for `read_blocks` (R7) - -New `cli/src/components/tools/read-blocks.tsx`, modeled directly on `read-files.tsx`: -reuse the `ReadDiagnostics` shape (`findToolResultByKind(outputRaw, 'read_blocks_result')`), -`recoveryLabel`, and `getReadStatus` logic; selector labels are -`path:win N`, `path@"match"#occ`, `path#symbol`. Register it in -`cli/src/components/tools/registry.ts`. - -Note the load-time invariant in `registry.ts`: it throws when -`toolMetadata[tool].renderer === 'custom'` and no component is registered. If you also set -`read_blocks` to `renderer: 'custom'` in `common/src/tools/metadata.ts`, both edits must land -together. Prefer registering the component first and updating metadata in the same -transaction. - -### M1-T7 — Grant `read_blocks` to the remaining read-capable agents (R7) - -Add `'read_blocks'` to `toolNames` for the agents that already have `read_files` but not -`read_blocks`: `agents/test-writer/test-writer.ts`, `agents/doc-writer/doc-writer.ts`, -`agents/reviewer/code-reviewer.ts`, `agents/security-reviewer/security-reviewer.ts`, -`agents/general-agent/general-agent.ts`, `agents/thinker/thinker.ts`, -`agents/debugger/debugger.ts`, `agents/synthesizer/synthesizer.ts`, and the specialists via -`agents/specialists/create-specialist.ts`. Confirm each file's current list by reading it -first — do not add the tool to an agent that deliberately has no file-read access -(`file-picker`, `code-searcher`, `basher`, `git-committer` are read-discovery/exec agents; -check before touching). - -### M1 validation gate - -- `packages/agent-runtime` typecheck + `read-blocks.test.ts` + `read-files-edit-state.test.ts` -- `common` typecheck + `bun test src/tools` -- `agents` typecheck + `bun test __tests__` (context-pruner, tool-reachability, roster-drift) -- `cli` typecheck + `bun test src/components/tools` -- New tests: A1, A2, A3, A4, A5 from SPEC. - -### M1 tests to author - -In `packages/agent-runtime/src/__tests__/read-blocks.test.ts`: - -1. window covering `1..totalLines` on a strict-mode state → `readAuthorizationsByPath[path]` - is `true` and `readAuthorizationHashesByPath[path]` equals the content hash. -2. window covering a sub-range → neither map is set, but `editAnchor.readCapability` decodes - to the block bounds. -3. `around` block that happens to span the whole file → whole-file grant. -4. seeded `editRereadRequirementsByPath[path] = {reason:'context_compacted'}` → cleared by (1), - preserved by (2). -5. `windowSize` at the cap succeeds; a resolved block over `MAX_READ_BLOCK_BYTES` returns - `too_large` with no `editAnchor`. -6. heuristic (non-parser) symbol slice still exposes no `editAnchor` and no grant. - -In `common/src/tools/__tests__/read-files-schema.test.ts`: `windowSize`/`contextLines` above -the caps are rejected. - ---- - -## M2 — Contract: unified selector surface on `read_files` - -### M2-T1 — Extend the `read_files` input schema (R4) - -In `common/src/tools/params/tool/read-files.ts`, add `windows` and `around` arrays with the -exact shapes already defined in `read-blocks.ts` (reuse by importing the per-selector object -schemas — export them from the read-blocks params module rather than duplicating). Update: - -- the `superRefine` emptiness check to include the two new arrays; -- `inferSingleSelectorPath` so `windows` and `around` also inherit a single `paths[0]` - shorthand (this is the silent-break risk in SPEC); -- the tool `description` to document the five selectors and the authority ladder; -- the example call to include one `windows` and one `around` entry. - -### M2-T2 — Unify the result item union (R4, R5) - -In `common/src/tools/results/filesystem.ts`: - -- move `readBlocksWindowItemSchema` / `readBlocksAroundItemSchema` / - `readBlocksSymbolItemSchema` into `readFilesItemV1Schema`'s union (the error item's - `selector` enum already lists `window`/`around`/`symbol`); -- add optional `referencedBy: z.record(z.string(), z.string().array()).optional()` to the - range/window/around/symbol item schemas (R5), and keep the existing `.strict()` calls valid; -- keep `readBlocksResultV1Schema` / `buildReadBlocksResultV1` / `isReadBlocksResultV1` - exported for M3's forwarding surface, but define them in terms of the shared item union; -- preserve every existing `superRefine` invariant, notably "partial results cannot expose - exact source content or edit capabilities". - -### M2-T3 — One handler serves both tools - -Refactor `read-files.ts` to accept the five selector groups and produce one ordered result -array. Extract the per-selector block builders currently in `read-blocks.ts` -(window/around/symbol) into a shared module — natural home: -`packages/agent-runtime/src/structural-read.ts` already owns `findLiteralOccurrences` and -`selectSymbolSlice`, so add `buildWindowBlock` / `buildAroundBlock` there and have both -handlers call them. `requestIndex` ordering is -`paths → ranges → windows → around → symbols`. - -### M2-T4 — Manifest-first oversized reads (R6) - -In `read-files.ts`, when a `paths` read is rejected/truncated for size, instead of returning -only the failure, emit the window manifest for that path plus its first window (a `window` -selector item with `windowSize`/`windowCount`/`totalLines`) so the agent can page -immediately. Keep the failure information in the same result (`status:'partial'` + -`truncation`), and mint no whole-file capability for it. - -### M2-T5 — Regenerate the public type surface (R7, A8) - -Run `bun scripts/generate-tool-definitions.ts` — it writes -`common/src/templates/initial-agents-dir/types/tools.ts`, `agents/types/tools.ts`, -`.agents/types/tools.ts`, then chains `cli/scripts/generate-init-type-sources.ts` for -`cli/src/data/initial-agent-type-sources.generated.ts`. All four must be committed; -`cli/knowledge.md:868` documents that CI verifies they are current. - -### M2 validation gate - -Repo-wide `bun run typecheck`, plus `common`/`agent-runtime`/`sdk`/`cli`/`agents` test -suites, plus `cli/src/__tests__/init-type-sources.test.ts` and -`common/src/tools/__tests__/tool-registration-consistency.test.ts`. - -New tests: A6, A8 from SPEC. - ---- - -## M3 — Deprecate `read_blocks` to a forwarding surface - -### M3-T1 — Forward and mark deprecated - -`read-blocks.ts` becomes a thin adapter: map its input to the unified selector groups, call -the shared handler, wrap the result with `buildReadBlocksResultV1`. Prefix its -`description` with a deprecation note pointing at `read_files`. Keep the tool registered so -`agents/tool-reachability.test.ts` and `scripts/check-tool-registration.ts` stay green and no -agent's `toolNames` breaks. - -### M3-T2 — Update prompts and docs - -- `agents/editor/editor.ts` (~lines 96, 108), `agents/editor/repair-editor.ts` (~line 28), - `agents/base2/base2.ts` (~lines 214, 265): change "prefer read_blocks for large files" to - "prefer `read_files` windows/around for large files". -- `docs/agents-and-tools.md` and `packages/agent-runtime/docs/deterministic-edit-system.md`: - document the five selectors and the three-tier authority ladder. -- `AGENTS.md` retrieval-conventions bullet. - -### M3 validation gate - -`agents` typecheck + `bun test __tests__` (includes `quality-prompt-snapshot.test.ts` — the -snapshot will need updating), plus A7 and A9 from SPEC. - ---- - -## Dependencies - -- M1-T2 and M1-T3 both depend on M1-T1. -- M1-T6 depends on M1-T3 only for the result fields it renders; it may be done in parallel. -- M2-T3 depends on M2-T1 + M2-T2. -- M2-T5 depends on all other M2 tasks. -- M3 depends on M2 completing. - -## Task status - -- [ ] M1-T1 extract shared coverage→authority ladder -- [ ] M1-T2 rewire read_files onto the ladder -- [ ] M1-T3 read_blocks whole-file grant + context_compacted parity -- [ ] M1-T4 bound windowSize/contextLines + block byte budget -- [ ] M1-T5 context-pruner read_blocks recognition -- [ ] M1-T6 CLI read-blocks renderer + registry -- [ ] M1-T7 grant read_blocks to remaining read-capable agents -- [ ] M1-T8 author M1 regression tests (A1–A5) -- [ ] M1-GATE run M1 validation suites -- [ ] M2-T1 read_files accepts windows + around -- [ ] M2-T2 unified result item union + referencedBy everywhere -- [ ] M2-T3 one shared handler for both tools -- [ ] M2-T4 manifest-first oversized reads -- [ ] M2-T5 regenerate the four generated type mirrors -- [ ] M2-T6 author M2 tests (A6, A8) -- [ ] M2-GATE repo-wide typecheck + all package suites -- [ ] M3-T1 read_blocks forwards + deprecation note -- [ ] M3-T2 prompt and docs updates -- [ ] M3-GATE agents suites incl. prompt snapshot (A7, A9) - - - -## M1-T1 through M1-T4 + M1-T8 complete — 2026-07-30 — 2026-07-30T18:25:38.435Z - -All four implementation tasks done and verified: - -- M1-T1: read-authority-ladder.ts created (classifyReadBlockAuthority, 9 unit tests) -- M1-T2: read_files rewired onto the ladder (both file/range branches unified) -- M1-T3: read_blocks gains whole-file grant + context_compacted parity (6 regression tests) -- M1-T4: windowSize/contextLines capped, MAX_READ_BLOCK_BYTES enforced (8 schema tests) -- M1-T8: 37 tests pass total, typecheck clean for agent-runtime + common - -Next: M1-T5 (context-pruner), M1-T6 (CLI renderer), M1-T7 (agent grants). diff --git a/.agents/sessions/read-tool-unification-2026-07/SPEC.md b/.agents/sessions/read-tool-unification-2026-07/SPEC.md deleted file mode 100644 index fa920d625b..0000000000 --- a/.agents/sessions/read-tool-unification-2026-07/SPEC.md +++ /dev/null @@ -1,97 +0,0 @@ -# SPEC — Unify the read tool surface onto `read_files` - -## Goal - -Make the read surface as mature as `edit_transaction`: one tool, one result kind, one -authority ladder, byte-bounded, fully wired into the CLI/pruner/agent ecosystem — without -weakening strict read-before-edit. - -## Non-goals - -- Changing `edit_transaction`, `replace_range`, `rewrite_symbol`, `str_replace`, or - `write_file` authorization semantics. The read side must keep feeding them exactly the - authority classes they accept today. -- Changing `read_outline`, `read_subtree`, `read_image`, `read_logs`, `read_docs`. -- Deleting `read_blocks`. It is retained as a forwarding surface (M3), not removed. -- Loosening `sdk/src/tools/read-policy.ts` / `sensitive-paths` behavior. - -## Current behavior (verified against source) - -| Fact | Evidence | -| ----------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `read_blocks` DOES mint `editAnchor` (`startLine`/`endLine`/`contentHash`/cap.v3) | `packages/agent-runtime/src/tools/handlers/tool/read-blocks.ts` `mintBlockEditAnchor` | -| `read_blocks` NEVER grants sticky whole-file auth | no `grantWholeFileReadAuthorization` import in `read-blocks.ts`; callers are `read-files.ts`, `replace-range.ts`, `edit-transaction.ts`, `write-file.ts` | -| `read_files` grants whole-file auth for a complete `paths` read AND a complete `ranges` read covering `1..totalLines` (hashing `sourceContent`) | `read-files.ts` ~lines 215–290, `wholeFileGrantPaths` | -| Only a whole-file grant may clear `context_compacted` | `read-files.ts` post-loop; `read-blocks.ts` unconditionally `continue`s on that reason | -| `windowSize` / `contextLines` have NO upper bound | `common/src/tools/params/tool/read-blocks.ts` (`.min(1)` / `.min(0)` only) | -| `read_files` bounds range reads at 4 MiB | `MAX_RANGE_READ_BYTES = 4_194_304`, `sdk/src/tools/read-files.ts:41` | -| `read_blocks` reads the FULL untruncated file then slices in memory | `requestOptionalFile` → `getFileForEditResult` (`sdk/src/run.ts` ~857, comment says "MUST be the full, untruncated file") | -| Read policy is intact for both tools | `getFileForEditResult` applies `fileFilter`; `sdk/src/__tests__/run-file-filter.test.ts` `[SEC-H02]` | -| No CLI renderer for `read_blocks` | `cli/src/components/tools/registry.ts` has `ReadFilesComponent`/`ReadSubtreeComponent`, no read-blocks entry | -| Context-pruner does not know `read_blocks_result` | `agents/context-pruner.ts` keys on `read_files` / `read_files_result` at lines 161, 229, 962, 1062, 1105, 2302 | -| Only 4 agents may call `read_blocks` | `agents/editor/editor.ts`, `agents/editor/repair-editor.ts`, `agents/base2/base2.ts`, `agents/base2/base-deep.ts` | -| `referencedBy` exists only on the whole-file item | `readFilesFileItemSchema` in `common/src/tools/results/filesystem.ts` | -| The result schema already anticipates the merge | `readFilesErrorItemSchema.selector` enum is already `['file','range','symbols','window','around','symbol']` | - -## Requirements - -- **R1 — Whole-file-covering block grants whole-file authority.** A complete `window` or - `around` block whose `[startLine,endLine] === [1,totalLines]` must call - `grantWholeFileReadAuthorization` with its `sourceContent` and may clear - `context_compacted`. Sub-file blocks must not. -- **R2 — One authority ladder.** The coverage→authority decision exists in exactly one - helper, consumed by both handlers. No third copy. -- **R3 — Bounded blocks.** `windowSize` and `contextLines` are schema-capped, and an - oversized resolved block fails with `too_large` or returns `status:'partial'` + - `truncation`, never unbounded content. `partial` blocks still mint no capability. -- **R4 — Unified selector surface.** `read_files` accepts `paths | ranges | windows | -around | symbols` in one call, returning one `read_files_result` with contiguous - `requestIndex`. -- **R5 — Uniform metadata.** `referencedBy` is available on every non-error selector, not - just whole-file reads. -- **R6 — Manifest-first oversized reads.** A `paths` read that would exceed limits returns - the window manifest (`totalLines`/`windowSize`/`windowCount`) plus the first window - instead of a bare failure. -- **R7 — Ecosystem parity.** CLI renderer, context-pruner recognition, and agent grants - cover the new selectors; generated tool-definition mirrors stay current. -- **R8 — No authorization regression.** Legacy/absent state still fails closed; partial, - truncated, and heuristic-regex slices still mint nothing. - -## Acceptance criteria - -| ID | Behavior | Verification | -| --- | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | -| A1 | `read_blocks`/`read_files` window covering the whole file authorizes a following `write_file` under strict mode | new case in `packages/agent-runtime/src/__tests__/read-blocks.test.ts` + `read-files-edit-state.test.ts` | -| A2 | Sub-file window does NOT authorize `write_file`, still authorizes `replace_range` via its capability | same suites | -| A3 | Whole-file-covering block clears `context_compacted`; sub-file block does not | `read-blocks.test.ts` | -| A4 | `windowSize`/`contextLines` above the cap are rejected by the schema | `common/src/tools/__tests__/read-files-schema.test.ts` | -| A5 | A resolved block over the byte budget yields `too_large`/`partial` with no `editAnchor` | `read-blocks.test.ts` | -| A6 | One `read_files` call with all five selector kinds returns one result, contiguous indexes | `read-files-edit-state.test.ts` + `filesystem.test.ts` | -| A7 | `read_blocks` still works and returns the unified shape | `read-blocks.test.ts` | -| A8 | Generated type mirrors are current | `bun scripts/generate-tool-definitions.ts` produces no diff; `cli/src/__tests__/init-type-sources.test.ts` | -| A9 | Every read-capable agent that has `read_files` also has the new selectors reachable | `agents/tool-reachability.test.ts`, `scripts/check-tool-registration.ts` | - -## Relevant systems - -- `common/src/tools/params/tool/read-files.ts`, `.../read-blocks.ts` — input schemas + descriptions -- `common/src/tools/results/filesystem.ts` — result item unions, builders, guards -- `common/src/tools/metadata.ts` — READ_TOOLS set, renderer intent -- `packages/agent-runtime/src/tools/handlers/tool/read-files.ts`, `read-blocks.ts`, `write-file.ts` (grant fns), `edit-read-state.ts` -- `packages/agent-runtime/src/structural-read.ts`, `get-file-reading-updates.ts` -- `sdk/src/tools/read-files.ts` (`MAX_RANGE_READ_BYTES`, `authorizeReadTarget`), `sdk/src/run.ts` (`requestFiles`/`requestOptionalFile`) -- `cli/src/components/tools/registry.ts` + new renderer; `agents/context-pruner.ts` -- Generated: `agents/types/tools.ts`, `.agents/types/tools.ts`, `common/src/templates/initial-agents-dir/types/tools.ts`, `cli/src/data/initial-agent-type-sources.generated.ts` - -## Risks - -- **Silent authority widening.** A coverage predicate that is too loose grants whole-file - auth from a partial read. Mitigation: the shared helper takes `{complete, startLine, -endLine, totalLines, sourceContent}` and returns `'whole_file' | 'scoped' | 'none'`; - `'whole_file'` requires all of complete + 1 + totalLines + a real `sourceContent`. -- **Preprocessor gap.** `inferSingleSelectorPath` currently infers a missing `path` only - for `ranges`/`symbols`. Omitting `windows`/`around` breaks single-path shorthand silently. -- **CI-verified generated drift.** `cli/knowledge.md:868` states CI verifies the generated - tool sources; forgetting the regen chain fails CI, not local typecheck. -- **Registry invariant.** `cli/src/components/tools/registry.ts` throws at module load when - metadata declares `renderer: 'custom'` without a registered component. Metadata and - registration must land in the same change. diff --git a/.agents/sessions/read-tool-unification-2026-07/STATE.json b/.agents/sessions/read-tool-unification-2026-07/STATE.json deleted file mode 100644 index 9e3e544f39..0000000000 --- a/.agents/sessions/read-tool-unification-2026-07/STATE.json +++ /dev/null @@ -1,10 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "read-tool-unification-2026-07", - "status": "completed", - "currentTask": null, - "revision": 2, - "checkpoint": null, - "createdAt": "2026-08-02T17:34:06.608Z", - "updatedAt": "2026-08-02T21:34:04.262Z" -} diff --git a/.agents/sessions/read-tool-unification-2026-07/STATUS.md b/.agents/sessions/read-tool-unification-2026-07/STATUS.md deleted file mode 100644 index 3bc6edc74d..0000000000 --- a/.agents/sessions/read-tool-unification-2026-07/STATUS.md +++ /dev/null @@ -1,90 +0,0 @@ -# STATUS — read tool unification - -## Current state - -Planning complete. No implementation code has been touched. `SPEC.md` and `PLAN.md` are -written and awaiting user review before M1 starts. - -Branch at plan time: `fix/reviewer-gate-hardening` (clean worktree). - -## Completed - -- Source-verified audit of the read surface (`read_files`, `read_blocks`, `read_outline`, - `read_subtree`, SDK read policy, capability minting, edit-authorization consumers). -- Decision recorded: merge direction is **into `read_files`**, because whole-file authority - minting (`grantWholeFileReadAuthorization`), the strict-mode recovery selector in - `strictEditAuthorizationError`, the context-pruner keys, and the CLI renderer are all - already anchored on `read_files`. Making `read_blocks` the superset would create a second - tool that mints whole-file authority. -- `SPEC.md` — requirements R1–R8, acceptance criteria A1–A9, evidence table, risks. -- `PLAN.md` — M1 (correctness, no schema change), M2 (unified selector contract), - M3 (deprecate `read_blocks` to a forwarding surface). - -## Pending - -All of M1, M2, M3. Next checkpoint is **M1-T1**: extract -`classifyReadBlockAuthority` into -`packages/agent-runtime/src/tools/handlers/tool/read-authority-ladder.ts`. - -## Blocked - -Nothing. M1 needs no decisions beyond what SPEC records. M2 changes the public tool schema -and regenerates four committed type mirrors, so it should get explicit go-ahead before it -starts. - -## Resume instructions - -1. Read `SPEC.md` (evidence table + acceptance criteria) then `PLAN.md`. -2. Start at the `` pointer in `PLAN.md`. -3. M1 is safe to ship alone: it grants no new authority class, only makes an already-complete - whole-file observation grant what an identical `read_files` observation already grants. -4. Do not begin M2 until M1's gate is green — M2-T3 collapses both handlers onto the shared - block builders, which is only mechanical once the authority ladder is single-sourced. - -## Key invariants to preserve - -- `whole_file` authority requires ALL of: `complete`, `startLine === 1`, - `endLine === totalLines`, and a real undecorated `sourceContent`. Numbered display content - must never be used for hashing or granting. -- Partial/truncated blocks and heuristic (non-parser) symbol slices mint no capability. -- Only a whole-file grant may clear `context_compacted`. -- `cli/src/components/tools/registry.ts` throws at module load when metadata declares - `renderer: 'custom'` with no registered component — metadata and registration land together. - - - -## M1 + M2 complete (gate-verified) — 2026-08-02T17:34:06.607Z - -Milestone M1 (correctness: authority ladder + bounds + ecosystem wiring) and Milestone M2 (unified selector surface on read_files) are complete and gate-verified. - -**M1 (T1–T8 + GATE):** shared coverage→authority ladder (`classifyReadBlockAuthority`), read_files + read_blocks rewired onto it, whole-file-covering blocks grant sticky auth, windowSize/contextLines capped + MAX_READ_BLOCK_BYTES enforced, context-pruner recognizes read_blocks (FILE_INSPECTION_TOOLS, summarizeToolCall case, read_blocks_result collector, path tracking), CLI read-blocks.tsx renderer registered + metadata renderer:custom landed together, read_blocks granted to all read-capable agents. Per user decision, thinker also gained read-only access (read_files + read_blocks; no-history/no-spawn preserved). M1-GATE green: agents 767, cli 71, agent-runtime 124, common 220. - -**M2 (T1–T6 + GATE):** read_files input schema gained windows/around selectors (reusing exported read-blocks selector schemas); five-selector emptiness check; inferSingleSelectorPath infers windows/around single-path shorthand; description documents all five selectors + authority ladder. Result item union unified (readBlocksItemV1Schema = readFilesItemV1Schema; window/around/symbol items moved into the shared union; referencedBy added to range/window/around/symbol strict schemas). Shared block builders (buildWindowBlock/buildAroundBlock/buildSymbolBlock + ReadBlockBuilderContext) extracted to structural-read.ts; both handlers share them; read_files serves five selectors ordered paths→ranges→windows→around→symbols. Manifest-first oversized reads synthesize a window manifest + first window after a truncated whole-file read (no whole-file capability). One consumer fix: simplify-tool-results.ts restricted content-stripping to file/range (window/around/symbol pass through) to resolve a TS2322. Four generated type mirrors regenerated and stable; init-type-sources 3/3. - -**Validation:** all repo typechecks green (script:typecheck, typecheck-common, typecheck-agents, typecheck-agent-runtime); targeted suites green (read-files-edit-state 126, read-blocks, read-files-schema, filesystem, simplify-tool-results 43, init-type-sources 3). Reviewer gate: NON_BLOCKING with 4 non-blocking nits recorded for later cleanup: (1) loadFile/mintBlockEditAnchor/applyBlockAuthority/overBudgetError duplicated between read-files.ts and read-blocks.ts — hoist a small factory into structural-read.ts; (2) readBlocksResultV1Schema superRefine re-implements the readFilesResultV1Schema invariant block — extract a shared helper; (3) selectSymbolSlice occurrence>1 relies on slice-truncation semantics — prefer astMatches[occurrence-1]; (4) no direct test of the over-budget too_large branch in buildWindowBlock/buildAroundBlock. - -**Next:** M3 (deprecate read_blocks to a forwarding surface + prompt/docs updates). M3-T1 makes read_blocks a thin adapter over the shared handler; M3-T2 updates editor/base2 prompts + docs + AGENTS.md; M3-GATE runs agents suites incl. quality-prompt-snapshot. - - - -## M3 complete: read_blocks fully removed (gate-verified) — 2026-08-02T20:18:50.156Z - -M3 is complete. Per user decision (full removal, not a forwarding alias), the `read_blocks` tool was deleted across every layer after `read_files` became a strict functional superset. - -**Prerequisite (occurrence gap closed first):** added the occurrence-aware `symbol` selector to `read_files` (`symbol: [{ path, name, occurrence? }]`, mirroring rewrite_symbol occurrence semantics) so `read_files` is a strict superset of `read_blocks`: paths, ranges, windows, around, occurrence-aware `symbol`, batch `symbols`. Context-pruner + CLI renderer brought to parity for the new selectors; mirrors regenerated; gate NON_BLOCKING. - -**M3 removal (2 editor waves + follow-ups):** - -- Wave 1 (core): removed `read_blocks` from `constants.ts` (toolNames + publishedTools), `list.ts`, `metadata.ts` (READ*TOOLS/CUSTOM_RENDERERS/PATH_INPUTS); deleted `params/tool/read-blocks.ts` (selector schemas + MAX\*\* constants relocated into `read-files.ts`); removed `readBlocksResultV1Schema`/`buildReadBlocksResultV1`/`isReadBlocksResultV1` + types from `filesystem.ts` (kept the `readBlocks*ItemSchema`window/around/symbol item kinds inside the shared`readFilesItemV1Schema`union); deleted the runtime handler + its registration; deleted the CLI renderer + registry entry; deleted`read-blocks.test.ts`+`read-blocks-schema.test.ts`; removed all `read_blocks`/`read_blocks_result`handling from`agents/context-pruner.ts`. -- Wave 2 (agents + docs): removed `read_blocks` from every agent `toolNames` (editor, repair-editor, thinker, code-reviewer, security-reviewer, debugger, doc-writer, test-writer, synthesizer, general-agent, base2, base-deep, create-specialist); rewrote prompt prose to point at `read_files` windows/around/symbol selectors; updated agent tests (thinker/code-reviewer/editor/base2/gate-lifecycle e2e); folded `read_blocks` docs into `read_files` in `docs/agents-and-tools.md` + `docs/deterministic-edit-system.md`. -- Follow-up fixes surfaced by validation/review: input-aliases map — the occurrence-aware `symbol` is a real canonical selector, NOT aliased onto batch `symbols` (self-alias with coerce:'array'+coerceCanonical so a singular object coerces to a one-element array); added `window`/`around`/`symbol` alias entries; updated the input-aliases test. `metadata.ts` PATH_INPUTS for `read_files` extended with `windows[].path`/`around[].path`/`symbol[].path`. `structural-read.ts` user-facing error strings/comments re-pointed from `read_blocks` to `read_files`. Model-facing `read_files` description re-pointed the `symbol` selector's occurrence semantics from `read_blocks` to `rewrite_symbol`. Four generated type mirrors regenerated (init-type-sources 3/3; tool-registration-consistency green). - -**Validation:** all typechecks green (script:typecheck, typecheck-common, typecheck-cli, typecheck-agents, typecheck-agent-runtime); common tools 223/223; input-aliases 9/9; read-files-schema 15/15; agent-runtime read tests green (one environmental tree-sitter c_sharp .scm build error in a single occurrence test, unrelated to this change); cli tools 71/71. Reviewer gate: NON_BLOCKING (cosmetic nits: sequential per-selector block-builder awaits, duplicated DEFAULT_WINDOW_SIZE/DEFAULT_CONTEXT_LINES constants, a CLI legacy-heuristic comment, symbols 100k-char slice cap documentation, duplicated totalLines derivation helper). - -**Plan complete:** M1 (authority ladder + ecosystem), M2 (unified selector surface), M3 (full read_blocks removal) all done and gate-verified. The read surface is now a single `read_files` tool with six selectors, one authority ladder, cap.v3 minting, byte budgets, and full CLI/pruner/agent parity. - - - -## Session closed — all milestones complete — 2026-08-02T21:34:04.262Z - -All three milestones (M1 authority ladder, M2 unified selector surface, M3 read_blocks removal) are complete and gate-verified per the appended entries above. The read surface is now unified on `read_files` with six selectors (paths, ranges, windows, around, symbol, symbols). `read_blocks` is fully removed from every layer (registry, handler, schemas, CLI renderer, agent grants, prompts, docs). Flipping session state to completed. diff --git a/.agents/sessions/read-write-tooling-2026-07-10/AUDIT-REPORT.md b/.agents/sessions/read-write-tooling-2026-07-10/AUDIT-REPORT.md deleted file mode 100644 index 1ef4c3681d..0000000000 --- a/.agents/sessions/read-write-tooling-2026-07-10/AUDIT-REPORT.md +++ /dev/null @@ -1,67 +0,0 @@ -# Read/write tooling audit report - -## Implemented in this pass - -1. Preserve `str_replace` failure pressure across intervening exact successes so alternating retry cascades reach the circuit breaker. -2. Keep `rewrite_symbol` available as a trusted structural recovery path and clear the failure budget after confirmed structural recovery. -3. Identify the exact failed replacement index in atomic batch errors. -4. Prevent SDK read-failure marker strings from granting strict read-before-edit authorization while preserving the failure marker in model-visible output. -5. Align `read_files` visible line counts and hashes with `replace_range` for newline-terminated files. -6. Allow `apply_patch` updates whose valid result is an empty file. -7. Show bounded multiline edit diagnostics in the CLI. -8. Correct model guidance for sequential/overlapping replacements and atomic batch recovery. -9. Honor unified patch coordinates and reject ambiguous coordinate-less repeated context. -10. Make context-pruner file/edit facts depend on successful matching tool results and preserve bounded head-and-tail diagnostics for every edit tool. -11. Reject unsafe runtime paths consistently before file or client I/O. -12. Invalidate prepared `str_replace`/`write_file` state when client application fails or throws. -13. Validate strict-mode `basedOnRead` anchors against current content, including small files and transaction edits. -14. Prevent range capabilities and symbol/range reads from authorizing whole-file overwrites or whole-file edit access. -15. Render explicit queued/pending/applied/failed states for `apply_patch` and queued/pending/read/partial/failed states for all `read_files` selector forms. -16. Let fresh per-replacement capabilities and `rewrite_symbol` recover through a failed-edit gate, clearing the gate only after client-confirmed success. -17. Enforce whole-file `write_file` authorization after resolving prior same-path edits, and revoke sticky authorization on processing/client failures found in any output part. -18. Require a matching fresh capability on every strict-mode replacement and emit strict-specific invalid-anchor guidance. -19. Give missing symbol reads an explicit structured failure reason instead of an ambiguous empty slice list. -20. Require positive `apply_patch` success evidence in the CLI and reject empty, malformed, nested-error, `applied: false`, and plain-error envelopes without showing the requested diff. -21. Prevent negated success text such as “not applied successfully” from becoming a persisted context-pruner edit fact. -22. Fix zero-context unified diff insertions so their old-file coordinate determines the insertion point. -23. Replace cross-turn Boolean read authorization with content-hash authorization that is revoked when disk content changes and advanced only after confirmed writes. -24. Add injected-filesystem-aware realpath containment across SDK reads, writes, patches, ranges, listings, and image reads. -25. Route direct edit tools through one prepare/apply/commit coordinator that commits only after positive client application evidence. -26. Treat empty or ambiguous client edit output as unconfirmed, preserve explicit rejection diagnostics, and require a fresh read after indeterminate application. -27. Report rollback failures accurately when an atomic SDK change cannot fully restore every path. -28. Introduce canonical `read_files` result version 1 with typed selector identity, status, errors, omissions, strict summary invariants, and legacy compatibility. -29. Build native structured reads from typed read metadata rather than reparsing rendered legacy marker strings. -30. Reconcile structured read results against the original selector index, kind, and normalized path before granting authorization. - -## Resolved architectural findings - -1. **Versioned whole-file authorization:** whole-file permission is now tied to a content hash across turns. External changes and rejected/failed edits revoke it; confirmed writes advance it. -2. **Unified edit application coordination:** `write_file`, `str_replace`, `replace_range`, `edit_transaction`, and runtime `apply_patch` now share confirmation, invalidation, and commit behavior. -3. **Realpath containment:** SDK filesystem operations resolve containment through the same injected filesystem used for the operation, including virtual-filesystem symlink escape tests. -4. **Structured `read_files` compatibility slice:** canonical v1 results now preserve selector identity, typed failures, aggregate status, omission semantics, and legacy histories/overrides. -5. **Accurate rollback reporting:** partial rollback failure is surfaced with affected paths instead of claiming complete atomic recovery. -6. **Fail-closed client confirmation:** empty or ambiguous client output cannot synthesize edit success. Legacy confirmation is tied to the expected tool/path shape, and original client/preflight diagnostics are preserved. - -## Highest-priority remaining findings - -1. **MEDIUM — structured edit results:** the coordinator still decodes legacy edit envelopes heuristically. Migrate mutation tools to a canonical `file_edit_result` contract with explicit `changed`, `atomic`, per-file status, failed operation index, and recovery requirements. -2. **MEDIUM — coherent multi-selector reads:** native structured whole/range/symbol selectors are independently read. A single logical request can therefore observe different file versions; introduce shared snapshot/read deduplication where selectors overlap. -3. **MEDIUM — broader filesystem result migration:** extend canonical versioned results beyond `read_files` to `read_subtree`, outlines/slices, discovery/listing tools, and generated output/result types. -4. **LOW — migration lifecycle:** add v0/v1/malformed telemetry, document deprecation, switch defaults only after compatibility evidence, and retain legacy decoding for persisted histories through the deprecation window. - -## Validation - -- Five workspace typechecks passed: `common`, `sdk`, `packages/agent-runtime`, `cli`, and `agents`. -- Runtime: 879 passed. -- SDK: 763 passed, 1 pre-existing skipped integration test. -- Agents: 523 passed. -- Common: 616 passed. -- Focused CLI read-result tests: 7 passed. -- Focused edit coordinator/recovery tests: 69 passed. -- SDK build and packaged-consumer verification passed, including CJS, ESM/types/compile, bundled ripgrep, and tree-sitter query checks. -- Final independent reviewer gate: **APPROVE**. -- Final architect gate: **APPROVE**. - -## Coverage - -See `COVERAGE-MATRIX.md`. The audit covered runtime edit state/matching, SDK filesystem application, common tool contracts, model-facing recovery prompts/context pruning, and CLI tool rendering. diff --git a/.agents/sessions/read-write-tooling-2026-07-10/COVERAGE-MATRIX.md b/.agents/sessions/read-write-tooling-2026-07-10/COVERAGE-MATRIX.md deleted file mode 100644 index 4b501cb0f7..0000000000 --- a/.agents/sessions/read-write-tooling-2026-07-10/COVERAGE-MATRIX.md +++ /dev/null @@ -1,28 +0,0 @@ -# Coverage matrix - -| Domain | Shard IDs | Covered | -| --------------------------------------------- | ------------------ | ------- | -| Runtime edit state and matching | runtime-edit-audit | yes | -| SDK filesystem execution and common contracts | sdk-contract-audit | yes | -| Agent prompts, pruning, and CLI UX | ux-prompt-audit | yes | - -## Subsystem enumeration - -- `packages/agent-runtime`: audited — deterministic read/edit processing, handlers, and focused tests. -- `sdk`: audited — read, patch, range, and client file-application paths. -- `common`: audited — read-anchor and edit-tool schemas/contracts. -- `agents`: audited — editor/base recovery prompts and context-pruning behavior. -- `cli`: audited — read/write/edit tool rendering and focused tests. -- `.agents`: out-of-scope except for this audit's artifacts. -- `.github`: out-of-scope — CI configuration is not part of read/write tool behavior. -- `agents-graveyard`: out-of-scope — inactive implementations. -- `common-legacy`: out-of-scope — no active read/write tool paths selected by the map. -- `docs`: audited only where deterministic edit contracts are documented. -- `evals`: out-of-scope — evaluation runners do not implement file I/O semantics. -- `node_modules`: out-of-scope — dependencies. -- `openbuff.d.example`: out-of-scope — example configuration. -- `packages/code-map`: out-of-scope — structural parsing is consumed by rewrite_symbol but file mutation is outside this package. -- `packages/indexer`: out-of-scope — retrieval/indexing, not deterministic file mutation. -- `packages/internal`: out-of-scope — provider internals. -- `prototype`: out-of-scope — user project content, not harness implementation. -- `scripts`: out-of-scope except the structural-map builder used for audit setup. diff --git a/.agents/sessions/read-write-tooling-2026-07-10/MAP.md b/.agents/sessions/read-write-tooling-2026-07-10/MAP.md deleted file mode 100644 index 70d353a3d8..0000000000 --- a/.agents/sessions/read-write-tooling-2026-07-10/MAP.md +++ /dev/null @@ -1,346 +0,0 @@ -# Structural Map — openbuff - -- **Project root:** `/home/ben/Code/CLI/openbuff` -- **Built at:** 2026-07-10T12:08:35.932Z -- **Total files indexed:** 2093 -- **Graph:** 14405 nodes, 70651 edges - -> Pin this file in context. Every audit shard navigates from here instead of doing fuzzy round-trip discovery. - -## Entry points - -- `cli/src/index.tsx` -- `packages/code-map/src/index.ts` -- `packages/indexer/src/index.ts` -- `packages/internal/src/index.ts` - -## Directories (by size, biggest first) - -| dir | files | total size | top symbols | -| -------------------------- | ----- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `evals` | 597 | 12.4 MB | main, run, makeEvalRun, makeAgentResults, toolCall, log | -| `cli` | 423 | 2.9 MB | render, main, TestItem, tmux, createErrorMessage, formatTimestamp | -| `packages` | 307 | 2.2 MB | Greeting, start, greet, Greeter, flush, doGenerate | -| `sdk` | 162 | 1.3 MB | main, run, createMockFs, log, resolveMcpEnv, errorResult | -| `agents` | 93 | 1.1 MB | extractInlineFunctionSource, parseGateStateBlock, feedJson, collectToolInputFiles, isFileChangingTool, hasEditArtifact | -| `common` | 237 | 1.0 MB | createMockLogger, getStringProperty, getFileExtension, process, sleep, size | -| `scripts` | 67 | 515.7 KB | main, parseArgs, computeCost, ConversationMessage, TurnResult, makeConversationStreamRequest | -| `agents-graveyard` | 121 | 350.6 KB | createBase2WithTaskResearcher, getLatestEditToolResults, extractSpawnResults, getSpawnResults, createResearchImplementOrchestrator, createBase2Implementor | -| `bun.lock` | 1 | 270.1 KB | — | -| `docs` | 13 | 202.1 KB | — | -| `.agents` | 25 | 138.5 KB | publisher, getSpawnerPrompt, getSystemPrompt, getInstructionsPrompt, getDefaultReviewModeInstructions, getWorkModeInstructions | -| `.github` | 14 | 51.8 KB | — | -| `.omx` | 1 | 33.9 KB | — | -| `openbuff.d.example` | 4 | 22.5 KB | — | -| `LICENSE` | 1 | 11.1 KB | — | -| `.bin` | 1 | 8.5 KB | — | -| `README.zh-CN.md` | 1 | 8.1 KB | — | -| `README.md` | 1 | 8.0 KB | — | -| `WINDOWS.md` | 1 | 7.3 KB | — | -| `CONTRIBUTING.md` | 1 | 5.6 KB | — | -| `CODE_OF_CONDUCT.md` | 1 | 4.5 KB | — | -| `eslint.config.js` | 1 | 4.0 KB | — | -| `AGENTS.md` | 1 | 3.6 KB | — | -| `package.json` | 1 | 2.5 KB | — | -| `INFISICAL_SETUP_GUIDE.md` | 1 | 2.5 KB | — | -| `ROUTER.md` | 1 | 2.5 KB | — | -| `.env.example` | 1 | 1.7 KB | — | -| `tsconfig.json` | 1 | 839 B | — | -| `SECURITY.md` | 1 | 520 B | — | -| `.gitignore` | 1 | 487 B | — | -| `.vscode` | 1 | 438 B | — | -| `bunfig.toml` | 1 | 432 B | — | -| `.prettierrc` | 1 | 389 B | — | -| `tsconfig.base.json` | 1 | 386 B | — | -| `test` | 1 | 332 B | setup | -| `knowledge.md` | 1 | 287 B | — | -| `.e2e-scratch` | 2 | 279 B | add, greet, multiply | -| `NOTICE` | 1 | 156 B | — | -| `openbuff.json.example` | 1 | 118 B | — | -| `.envrc` | 1 | 30 B | — | -| `.bun-version` | 1 | 7 B | — | - -## Largest files per directory - -### `evals` - -- `evals/buffbench/logs/2026-07-04T17-30_base2/45-fork-read-files-base2-349a140.json` — 472.2 KB, 0 symbols -- `evals/buffbench/logs/2026-07-04T13-41_base2/2-add-deep-thinkers-base2-6c362c3.json` — 469.4 KB, 0 symbols -- `evals/buffbench/restrict-tool-types-base2-lite-error-ftj2.json` — 418.5 KB, 0 symbols -- `evals/buffbench/validate-custom-tools-base2-error-c6yk.json` — 418.3 KB, 0 symbols -- `evals/buffbench/logs/2026-07-04T17-30_base2/38-unify-agent-builder-base2-4852954.json` — 411.4 KB, 0 symbols - -### `cli` - -- `cli/bin/tree-sitter.wasm` — 200.7 KB, 0 symbols -- `cli/src/hooks/helpers/__tests__/send-message.test.ts` — 55.5 KB, 0 symbols -- `cli/src/utils/__tests__/message-block-helpers.test.ts` — 55.1 KB, 0 symbols -- `cli/src/chat.tsx` — 53.8 KB, 1 symbols -- `cli/src/utils/__tests__/send-message-helpers.test.ts` — 48.1 KB, 0 symbols - -### `packages` - -- `packages/agent-runtime/src/__tests__/read-files-edit-state.test.ts` — 113.5 KB, 1 symbols -- `packages/agent-runtime/src/__tests__/process-str-replace.test.ts` — 76.8 KB, 0 symbols -- `packages/agent-runtime/src/process-str-replace.ts` — 66.5 KB, 30 symbols -- `packages/agent-runtime/src/__tests__/run-programmatic-step.test.ts` — 65.2 KB, 0 symbols -- `packages/agent-runtime/src/run-agent-step.ts` — 53.3 KB, 4 symbols - -### `sdk` - -- `sdk/src/__tests__/model-provider.test.ts` — 80.0 KB, 3 symbols -- `sdk/src/provider-config.ts` — 72.5 KB, 30 symbols -- `sdk/src/tools/browser-logs.ts` — 62.8 KB, 30 symbols -- `sdk/src/impl/llm.ts` — 52.9 KB, 20 symbols -- `sdk/src/__tests__/run-cancellation.test.ts` — 41.6 KB, 1 symbols - -### `agents` - -- `agents/base2/base2.ts` — 166.8 KB, 30 symbols -- `agents/__tests__/context-pruner.test.ts` — 117.8 KB, 1 symbols -- `agents/__tests__/base2.test.ts` — 111.3 KB, 4 symbols -- `agents/context-pruner.ts` — 71.8 KB, 30 symbols -- `agents/types/tools.ts` — 44.0 KB, 30 symbols - -### `common` - -- `common/src/templates/initial-agents-dir/types/tools.ts` — 44.0 KB, 30 symbols -- `common/src/util/__tests__/messages.test.ts` — 41.4 KB, 0 symbols -- `common/src/__tests__/agent-validation.test.ts` — 28.9 KB, 0 symbols -- `common/src/util/__tests__/saxy.test.ts` — 26.1 KB, 0 symbols -- `common/src/browser-actions.ts` — 25.4 KB, 30 symbols - -### `scripts` - -- `scripts/test-fireworks-cache-intervals.ts` — 34.3 KB, 10 symbols -- `scripts/benchmark-providers.ts` — 33.1 KB, 16 symbols -- `scripts/test-fireworks-long.ts` — 30.7 KB, 6 symbols -- `scripts/test-canopywave-long.ts` — 29.0 KB, 5 symbols -- `scripts/test-siliconflow.ts` — 27.3 KB, 5 symbols - -### `agents-graveyard` - -- `agents-graveyard/base/base-prompts.ts` — 25.3 KB, 3 symbols -- `agents-graveyard/editor/best-of-n/editor-best-of-n.ts` — 18.5 KB, 4 symbols -- `agents-graveyard/base/ask.ts` — 12.2 KB, 0 symbols -- `agents-graveyard/registry/transform-agent.ts` — 12.0 KB, 0 symbols -- `agents-graveyard/base2/task-researcher/base2-with-task-researcher-planner-pro.ts` — 11.4 KB, 1 symbols - -### `bun.lock` - -- `bun.lock` — 270.1 KB, 0 symbols - -### `docs` - -- `docs/agents-and-tools.md` — 74.1 KB, 0 symbols -- `docs/codebuff-to-openbuff-migration.md` — 30.7 KB, 0 symbols -- `docs/configuration.md` — 20.6 KB, 0 symbols -- `docs/openbuff-provider-model-setup-ux.md` — 20.4 KB, 0 symbols -- `docs/architecture.md` — 12.3 KB, 0 symbols - -### `.agents` - -- `.agents/types/tools.ts` — 44.0 KB, 30 symbols -- `.agents/types/agent-definition.ts` — 13.7 KB, 3 symbols -- `.agents/lib/cli-agent-prompts.ts` — 13.7 KB, 5 symbols -- `.agents/sessions/read-write-tooling-2026-07-10/MAP.md` — 11.0 KB, 0 symbols -- `.agents/codex-cli.ts` — 7.0 KB, 0 symbols - -### `.github` - -- `.github/workflows/cli-release-build.yml` — 11.7 KB, 0 symbols -- `.github/workflows/cli-release-staging.yml` — 8.8 KB, 0 symbols -- `.github/workflows/ci.yml` — 7.4 KB, 0 symbols -- `.github/knowledge.md` — 5.4 KB, 0 symbols -- `.github/workflows/cli-release-prod.yml` — 4.9 KB, 0 symbols - -### `.omx` - -- `.omx/state/todos-session.json` — 33.9 KB, 0 symbols - -### `openbuff.d.example` - -- `openbuff.d.example/providers.json` — 19.0 KB, 0 symbols -- `openbuff.d.example/routes.json` — 2.1 KB, 0 symbols -- `openbuff.d.example/hooks.json` — 1.2 KB, 0 symbols -- `openbuff.d.example/indexing.json` — 146 B, 0 symbols - -### `LICENSE` - -- `LICENSE` — 11.1 KB, 0 symbols - -### `.bin` - -- `.bin/bun` — 8.5 KB, 0 symbols - -### `README.zh-CN.md` - -- `README.zh-CN.md` — 8.1 KB, 0 symbols - -### `README.md` - -- `README.md` — 8.0 KB, 0 symbols - -### `WINDOWS.md` - -- `WINDOWS.md` — 7.3 KB, 0 symbols - -### `CONTRIBUTING.md` - -- `CONTRIBUTING.md` — 5.6 KB, 0 symbols - -### `CODE_OF_CONDUCT.md` - -- `CODE_OF_CONDUCT.md` — 4.5 KB, 0 symbols - -### `eslint.config.js` - -- `eslint.config.js` — 4.0 KB, 0 symbols - -### `AGENTS.md` - -- `AGENTS.md` — 3.6 KB, 0 symbols - -### `package.json` - -- `package.json` — 2.5 KB, 0 symbols - -### `INFISICAL_SETUP_GUIDE.md` - -- `INFISICAL_SETUP_GUIDE.md` — 2.5 KB, 0 symbols - -### `ROUTER.md` - -- `ROUTER.md` — 2.5 KB, 0 symbols - -### `.env.example` - -- `.env.example` — 1.7 KB, 0 symbols - -### `tsconfig.json` - -- `tsconfig.json` — 839 B, 0 symbols - -### `SECURITY.md` - -- `SECURITY.md` — 520 B, 0 symbols - -### `.gitignore` - -- `.gitignore` — 487 B, 0 symbols - -### `.vscode` - -- `.vscode/settings.json` — 438 B, 0 symbols - -### `bunfig.toml` - -- `bunfig.toml` — 432 B, 0 symbols - -### `.prettierrc` - -- `.prettierrc` — 389 B, 0 symbols - -### `tsconfig.base.json` - -- `tsconfig.base.json` — 386 B, 0 symbols - -### `test` - -- `test/setup-scm-loader.ts` — 332 B, 1 symbols - -### `knowledge.md` - -- `knowledge.md` — 287 B, 0 symbols - -### `.e2e-scratch` - -- `.e2e-scratch/widget.ts` — 275 B, 3 symbols -- `.e2e-scratch/browser-agent-note.txt` — 4 B, 0 symbols - -### `NOTICE` - -- `NOTICE` — 156 B, 0 symbols - -### `openbuff.json.example` - -- `openbuff.json.example` — 118 B, 0 symbols - -### `.envrc` - -- `.envrc` — 30 B, 0 symbols - -### `.bun-version` - -- `.bun-version` — 7 B, 0 symbols - -## Most-imported files (likely key modules) - -| in-degree | file | -| --------- | ------------------------------------------------------------- | -| 101 | `packages/agent-runtime/src/__tests__/rewrite-symbol.test.ts` | -| 80 | `common/src/util/messages.ts` | -| 75 | `common/src/types/bun-test.d.ts` | -| 73 | `cli/src/utils/arrays.ts` | -| 66 | `cli/src/__tests__/release/proxy-http-get.test.ts` | -| 62 | `common/src/util/error.ts` | -| 59 | `common/src/tools/params/utils.ts` | -| 54 | `sdk/e2e/utils/event-collector.ts` | -| 49 | `cli/src/utils/message-block-helpers.ts` | -| 43 | `cli/src/hooks/use-theme.tsx` | -| 43 | `packages/agent-runtime/src/__tests__/main-prompt.test.ts` | -| 40 | `sdk/src/provider-config.ts` | -| 40 | `common/src/util/plan-artifacts.ts` | -| 37 | `agents/base2/base2.ts` | -| 36 | `cli/src/project-files.ts` | -| 34 | `common/src/testing/mocks/timers.ts` | -| 32 | `scripts/test-canopywave-long.ts` | -| 31 | `.e2e-scratch/widget.ts` | -| 30 | `common/src/types/session-state.ts` | -| 29 | `sdk/e2e/utils/get-api-key.ts` | -| 28 | `common/src/util/string.ts` | -| 28 | `sdk/src/run.ts` | -| 28 | `cli/scripts/build-binary.ts` | -| 27 | `cli/src/utils/env.ts` | -| 27 | `common/src/util/lru-cache.ts` | - -## Cross-directory dependencies (architectural layering) - -| count | from → to | -| ----- | ----------------------------- | -| 232 | `packages` → `common` | -| 167 | `cli` → `common` | -| 125 | `cli` → `packages` | -| 124 | `sdk` → `common` | -| 62 | `cli` → `sdk` | -| 50 | `sdk` → `packages` | -| 25 | `cli` → `.e2e-scratch` | -| 23 | `cli` → `scripts` | -| 22 | `agents` → `sdk` | -| 21 | `packages` → `cli` | -| 20 | `sdk` → `cli` | -| 19 | `agents` → `agents-graveyard` | -| 19 | `agents-graveyard` → `common` | -| 16 | `packages` → `sdk` | -| 16 | `sdk` → `scripts` | -| 13 | `evals` → `common` | -| 12 | `scripts` → `packages` | -| 10 | `agents` → `common` | -| 9 | `evals` → `cli` | -| 8 | `common` → `packages` | -| 8 | `agents` → `cli` | -| 7 | `common` → `cli` | -| 7 | `evals` → `sdk` | -| 6 | `packages` → `agents` | -| 6 | `packages` → `.e2e-scratch` | -| 5 | `agents` → `packages` | -| 5 | `cli` → `agents` | -| 5 | `scripts` → `cli` | -| 5 | `evals` → `packages` | -| 5 | `common` → `sdk` | - -## Shard sizing hint - -Total indexed source: **22.5 MB** across **41** top-level directories. - -When sharding for an audit, aim for ~5–15 files per shard. Use the table above to group small dirs together and split huge dirs (e.g. split `src/` by subdirectory). diff --git a/.agents/sessions/read-write-tooling-2026-07-10/findings/runtime-edit.md b/.agents/sessions/read-write-tooling-2026-07-10/findings/runtime-edit.md deleted file mode 100644 index d20e717332..0000000000 --- a/.agents/sessions/read-write-tooling-2026-07-10/findings/runtime-edit.md +++ /dev/null @@ -1,17 +0,0 @@ -# Runtime edit findings - -- Fixed alternating and partial-success `str_replace` retry loops by retaining a per-path failure budget. -- Fixed `rewrite_symbol` recovery so it bypasses the raw retry breaker and resets the budget only after client-confirmed success. -- Fixed atomic errors so they identify the failed replacement index. -- Fixed failed SDK read markers so they remain visible but cannot authorize edits. -- Added lexical path hardening across runtime read/write handlers before I/O. -- Invalidated prepared direct-edit state after client rejection/throws. -- Enforced fresh strict-mode anchors and scoped range/symbol authorization. -- Allowed capability-bearing and structural recovery through failed-edit gates, clearing state only after confirmed application. -- Enforced whole-file authorization after prior same-path edits and revoked write authorization on any non-syntax processing/client failure. -- Scanned every client output part for write errors and corrected strict per-replacement capability guidance. -- Added content-hash-backed whole-file authorization across turns, including external-change revocation and confirmed-write hash advancement. -- Unified direct edit confirmation through a prepare/apply/commit coordinator for `write_file`, `str_replace`, `replace_range`, `edit_transaction`, and runtime `apply_patch`. -- Empty or ambiguous client output now fails closed; explicit client/preflight errors retain their actionable diagnostics. -- Canonical structured reads are reconciled against the requested selector before authorization, preventing mismatched or truncated results from granting whole-file access. -- Remaining work: canonical structured edit results and a coherent shared snapshot for overlapping selectors in one read request. diff --git a/.agents/sessions/read-write-tooling-2026-07-10/findings/sdk-contracts.md b/.agents/sessions/read-write-tooling-2026-07-10/findings/sdk-contracts.md deleted file mode 100644 index 6c9493d2a9..0000000000 --- a/.agents/sessions/read-write-tooling-2026-07-10/findings/sdk-contracts.md +++ /dev/null @@ -1,13 +0,0 @@ -# SDK and contract findings - -- Fixed trailing-newline and CRLF range hashes so `read_files` output is immediately valid for range-aware edit tools. -- Fixed valid `apply_patch` updates that delete all file content. -- Clarified generated `basedOnRead` contracts and regenerated tool definitions. -- Unified patch coordinates now disambiguate repeated context; unresolved coordinate-less ambiguity fails without writing. -- Zero-context unified insertion hunks now honor their declared old-file coordinate. -- Missing symbol reads now carry an explicit error reason instead of only `{ slices: [] }`. -- Added injected-filesystem-aware realpath containment for SDK reads, images, writes, patches, ranges, and directory listings, including an escaping virtual-filesystem symlink regression test. -- Added canonical `read_files` v1 results built directly from typed read metadata, with per-selector identity, typed errors, strict summary invariants, omissions, and legacy adapters. -- Preserved public `getFiles()` behavior while adding `getFilesStructured()` and a structured format opt-in used by the CLI. -- Atomic rollback now reports incomplete recovery and affected paths when any rollback operation fails. -- Remaining work: extend structured results to the rest of the filesystem tools and replace legacy edit-result envelopes with the shared canonical mutation result. diff --git a/.agents/sessions/read-write-tooling-2026-07-10/findings/structured-results-plan.md b/.agents/sessions/read-write-tooling-2026-07-10/findings/structured-results-plan.md deleted file mode 100644 index 8419aca69f..0000000000 --- a/.agents/sessions/read-write-tooling-2026-07-10/findings/structured-results-plan.md +++ /dev/null @@ -1,557 +0,0 @@ -# Structured filesystem result contracts: inventory and migration plan - -Date: 2026-07-10 - -## Decision - -This is primarily a harness/contract issue, with a secondary model-recovery issue. - -The model can still choose stale replacements or retry an atomic batch poorly, but the harness currently makes recovery harder than it should be: read failures are embedded in content strings, write success/failure is inferred from prose fields, several consumers recursively guess status from arbitrary JSON, and declared output schemas are not validated before native results are streamed. The repeated `str_replace` pattern in the motivating trace is therefore best understood as a model retry weakness amplified by ambiguous harness feedback. - -The recommended fix is to introduce versioned, discriminated filesystem result payloads while retaining adapters for every current envelope. Do not change the outer `ToolResultOutput[]` transport in this migration; keep the existing single JSON part (`[{ type: 'json', value: ... }]`) and version the JSON value inside it. - -## Scope and non-goals - -In scope: - -- `read_files`, including whole-file, range, and symbol selectors. -- `read_subtree`. -- Closely related filesystem reads: `read_outline`, `read_slices`, `find_files`, `list_directory`, and `glob`. -- File-edit result status shared by `str_replace`, `write_file`, `replace_range`, `rewrite_symbol`, `apply_patch`, and `edit_transaction`. -- SDK override compatibility, runtime authorization, context pruning/simplification, print-mode events, and CLI rendering. -- Generated agent tool types. - -Non-goals for the first implementation slice: - -- Replacing the general `ToolResultOutput[]` media/JSON transport. -- Changing MCP/custom-tool result contracts. -- Removing legacy result decoding in the same release that structured results are introduced. -- Combining this work with the separate realpath-containment or unified edit-coordinator projects. - -## Current result-flow inventory - -### Shared transport and type layer - -- `common/src/types/messages/content-part.ts:48-59` defines only the outer JSON/media result parts. A JSON result can contain any `JSONValue`; it has no filesystem status semantics. -- `common/src/tools/params/utils.ts:123-133` wraps a value in the stable one-part JSON transport. -- `common/src/util/messages.ts:788-805` likewise preserves a bare array or object as the JSON `value`; there is no additional envelope. -- `common/src/tools/list.ts:117-132` infers `CodebuffToolOutput` from each tool's Zod `outputSchema`. -- `packages/agent-runtime/src/tools/tool-executor.ts:587-630` accepts a native handler's output, streams it, and appends it to history without parsing it through `toolParams[toolName].outputSchema`. -- `common/src/tools/compile-tool-definitions.ts:10-66` generates input parameter types only. Programmatic agent templates receive no generated `ToolResultMap`/`GetToolResult` contract. - -### SDK `read_files` - -`sdk/src/tools/read-files.ts:33-38,208-399` currently returns `Record`: - -- A successful whole-file read is the raw content string. -- A blocked, missing, outside-root, over-10-MB, or I/O failure is a marker-prefixed string from `FILE_READ_STATUS` (`common/src/constants/paths.ts:43-64`). -- An example/template file is a success string prefixed with `[TEMPLATE]`. -- A range read is a rendered string with an embedded `[RANGE_BLOCK ...]` header, hash, capability, and numbered body (`sdk/src/tools/read-files.ts:349-396`). -- Multiple ranges for one path are concatenated into one string; a range result replaces a whole-file result for the same path. -- `null` remains possible for SDK overrides and empty/missing map entries, but the native reader usually uses marker strings for failures. - -The public and runtime-facing contracts preserve this map: - -- `common/src/types/contracts/client.ts:26-35` (`FileLineRange`, `RequestFilesFn`). -- `sdk/src/run.ts:85-95` (`OpenbuffClientOptions.overrideTools.read_files`). -- `sdk/src/run.ts:494-524,689-715` (runtime dependency wiring and legacy override lookup). -- `sdk/src/index.ts:9-11` publicly exports `getFiles`. - -This means failures have at least three representations before the runtime builds its tool result: a marker string, `null`, or a missing key. - -### Runtime `read_files` - -- `packages/agent-runtime/src/get-file-reading-updates.ts:20-39` drops `null`, `undefined`, and missing map entries, preserving marker strings as if they were ordinary content. -- `packages/agent-runtime/src/util/render-read-files-result.ts:17-48` constructs the current legacy value: - - ```ts - [ - { summary: { ok, failed, requested } }, - { path, content, referencedBy? }, - ] - ``` - - It detects failures by a regular expression over `content`. - -- `packages/agent-runtime/src/tools/handlers/tool/read-files.ts:41-77` combines whole, range, and symbol paths for path validation. -- `packages/agent-runtime/src/tools/handlers/tool/read-files.ts:93-125` uses `toOptionalFile()` marker parsing to clear edit gates and grant whole-file authorization. -- `packages/agent-runtime/src/tools/handlers/tool/read-files.ts:143-151` computes `requestedReadCount` from whole and range paths only. Symbol requests are excluded. -- `packages/agent-runtime/src/tools/handlers/tool/read-files.ts:159-201` appends `{ path, slices, errorMessage? }` entries after the legacy file array. - -Two concrete inconsistencies follow: - -1. `common/src/tools/params/tool/read-files.ts:11-43` does not declare `errorMessage` on the slice entry even though the runtime emits it. -2. `common/src/tools/params/tool/read-files.ts:47-113` requires `paths`, while the handler and tests support range-only and symbol-only calls. A true symbol-only model call without `paths: []` can be rejected before reaching the handler. - -The motivating ambiguity is reproducible from the code: a symbol-only miss can append `{ path, slices: [], errorMessage }` after a summary computed as `{ ok: 0, failed: 0, requested: 0 }`. - -### Runtime authorization and pruning - -- `common/src/types/session-state.ts:101,197` stores sticky whole-file authorization as `Record` rather than tying it to a content generation/hash. -- `packages/agent-runtime/src/tools/handlers/tool/read-files.ts:99-123` grants it from successful whole-file paths after marker-string classification. -- Range and symbol reads intentionally use scoped capabilities rather than whole-file authorization (`packages/agent-runtime/src/tools/handlers/tool/read-files.ts:190-198`). -- `packages/agent-runtime/src/util/simplify-tool-results.ts:24-41` turns every non-summary `read_files` entry into `{ path, contentOmittedForLength: true }`. That includes symbol results and symbol errors, so the simplifier can erase both selector kind and failure detail. -- `packages/agent-runtime/src/util/simplify-tool-results.ts:127-165` uses separate ad hoc omission variants for `read_subtree`. -- `agents/context-pruner.ts:763-879` recursively searches strings and fields for failure evidence. -- `agents/context-pruner.ts:935-995` separately reconstructs successful `read_files` paths from summary/content/slice heuristics, with a legacy fallback that treats an unrecognized non-failure result as success. -- `agents/context-pruner.ts:1547-1684` correctly correlates calls/results by `toolCallId`, but it still has to interpret the unstructured payloads described above. - -### Runtime `read_subtree` - -There is no SDK-native `read_subtree` implementation. It is a runtime-native operation over the cached/live project tree: - -- `packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:26-105` builds directory/file entries. -- `packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:278-303` builds `{ path, errorMessage }` failures. -- `packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts:307-352` returns an array of directory, file, or error entries with no top-level summary/status. -- `common/src/tools/params/tool/read-subtree.ts:56-99` also declares pruned `printedTreeOmittedForLength` and `variablesOmittedForLength` variants that the handler itself never emits. - -The current legacy value is therefore: - -```ts -Array< - | { path, type: 'directory', printedTree, tokenCount, truncationLevel } - | { path, type: 'file', variables } - | { path, errorMessage } - | pruned omission variants -> -``` - -It cannot distinguish an empty successful result from a malformed/empty result, and request-to-result identity is only the normalized path string. - -### Related read/filesystem contracts - -- `read_outline` puts a failure string in the successful `outline` field (`packages/agent-runtime/src/tools/handlers/tool/read-outline.ts:46-62`; schema at `common/src/tools/params/tool/read-outline.ts:36-46`). -- `read_slices` uses `slices: []` for both a missing file and no matching symbols (`packages/agent-runtime/src/tools/handlers/tool/read-slices.ts:19-35`; schema at `common/src/tools/params/tool/read-slices.ts:46-68`). -- `find_files` reuses `fileContentsSchema.array()` or `{ message }` (`common/src/tools/params/tool/find-files.ts:47-59`) and shares `RequestFilesFn`/`renderReadFilesResult` (`packages/agent-runtime/src/tools/handlers/tool/find-files.ts:92-128`). -- `list_directory` and `glob` use success/error unions distinguished only by property presence (`common/src/tools/params/tool/list-directory.ts:41-58`; `common/src/tools/params/tool/glob.ts:55-74`). - -### Write/edit contracts relevant to the observed retry loop - -- `str_replace`, `write_file`, `replace_range`, `rewrite_symbol`, and several plan tools share `{ file, message } | { file, errorMessage, patch? }` (`common/src/tools/params/tool/str-replace.ts:13-23`). -- `apply_patch` uses `{ message, applied[] } | { errorMessage }` (`common/src/tools/params/tool/apply-patch.ts:8-21`). -- `edit_transaction` has another union with success files or failures (`common/src/tools/params/tool/edit-transaction.ts:139-162`). -- `cli/src/components/tools/str-replace.tsx:117-200` infers applied/failed from `message`, `errorMessage`, and formatted text. -- `cli/src/components/tools/apply-patch.tsx:121-210` recursively searches arbitrary nested output for failure and requires hand-coded positive success evidence. -- `agents/context-pruner.ts:769-875,997-1036` performs another independent prose/field heuristic for edit failure and success. - -The runtime now emits better atomic diagnostics and maintains retry pressure, but the result contract still does not expose stable fields such as `atomic`, `changed`, `failedOperationIndex`, `retryable`, or `requiresFreshRead`. Those are exactly the facts a model, CLI, and pruner need to stop an unproductive retry pattern deterministically. - -### CLI ingestion and rendering - -- `common/src/types/print-mode.ts:51-59` transports the generic `ToolResultOutput[]` on `tool_result` events. -- `cli/src/utils/sdk-event-handlers.ts:560-590` passes the raw event output to block storage. -- `cli/src/utils/message-block-helpers.ts:1065-1090` stores both a JSON string and `outputRaw: unknown` without normalization. -- `cli/src/components/tools/read-files.tsx:97-181` recursively counts guessed successes/failures across summaries, content markers, errors, and slice arrays. -- `cli/src/components/tools/read-subtree.tsx:12-43` ignores output entirely and always renders `List deeply` with no pending/success/partial/failure state. -- `cli/src/types/chat.ts:37-53` types the raw result as `unknown`, so renderer contracts cannot be checked statically. - -## Target contract - -### Stable outer transport - -Keep this unchanged for all phases: - -```ts -type ToolResultOutput = - | { type: 'json'; value: JSONValue } - | { type: 'media'; data: string; mediaType: string } -``` - -Structured filesystem values live under the existing JSON part. Legacy unversioned values are called `v0`; the first structured contract is `version: 1`. - -### Shared fields - -Create `common/src/tools/results/filesystem.ts` with strict Zod schemas and inferred types: - -```ts -type FilesystemResultStatus = 'ok' | 'partial' | 'error' - -type FilesystemError = { - code: - | 'not_found' - | 'blocked' - | 'outside_project' - | 'too_large' - | 'io_error' - | 'invalid_request' - | 'stale_read' - | 'no_match' - | 'ambiguous_match' - | 'application_rejected' - message: string - retryable: boolean - requiresFreshRead?: boolean - recovery?: - | 'discover_path' - | 'read_again' - | 'read_smaller_range' - | 'choose_symbol' - | 'change_edit_strategy' -} -``` - -The error `message` remains model-readable, but state machines must use `status`, `code`, and recovery fields rather than parse the prose. - -### `ReadFilesResultV1` - -```ts -type ReadFilesResultV1 = { - kind: 'read_files_result' - version: 1 - status: FilesystemResultStatus - summary: { - requested: number - ok: number - partial: number - failed: number - uniquePaths: number - } - results: ReadFilesItemV1[] -} -``` - -Every normalized selector produces exactly one result item, identified by its input order: - -```ts -type ReadFilesItemV1 = - | { - selector: 'file' - requestIndex: number - path: string - status: 'ok' | 'partial' - content: string - complete: boolean - template: boolean - truncation?: { - reason: 'character_limit' - omittedStartLine?: number - omittedEndLine?: number - } - referencedBy?: Record - } - | { - selector: 'range' - requestIndex: number - path: string - status: 'ok' | 'partial' - content: string - startLine: number - endLine: number - totalLines: number - complete: boolean - rangeHash?: string - readCapability?: string - truncation?: { reason: 'character_limit' } - } - | { - selector: 'symbols' - requestIndex: number - path: string - status: 'ok' | 'partial' - requestedSymbols: string[] - missingSymbols: string[] - slices: ExtractedSlice[] - } - | { - selector: 'file' | 'range' | 'symbols' - requestIndex: number - path: string - status: 'error' - error: FilesystemError - } -``` - -Counting rules: - -- `requested` is the number of selector entries, not the number of unique paths. Each `paths[]` element, each `ranges[]` element, and each `symbols[]` group counts once. -- `summary.ok + summary.partial + summary.failed === summary.requested`. -- Duplicate paths remain distinct through `requestIndex`; do not concatenate ranges into an opaque string in the structured path. -- A symbol group with some slices found is `partial` and lists `missingSymbols`; no slices found is `error/no_match`. -- A 100k-character truncated whole/range read is `partial` with `complete: false`. It must not grant whole-file authorization. Do not emit a capability covering unseen range content; require a smaller range. -- An over-10-MB refusal is `error/too_large`, not a content string. - -### `ReadSubtreeResultV1` - -```ts -type ReadSubtreeResultV1 = { - kind: 'read_subtree_result' - version: 1 - status: FilesystemResultStatus - summary: { requested: number; ok: number; partial: number; failed: number } - results: Array< - | { requestIndex: number; path: string; status: 'ok' | 'partial'; type: 'directory'; printedTree?: string; printedTreeOmittedForLength?: true; tokenCount: number; truncationLevel: ... } - | { requestIndex: number; path: string; status: 'ok' | 'partial'; type: 'file'; variables?: string[]; variablesOmittedForLength?: true } - | { requestIndex: number; path: string; status: 'error'; error: FilesystemError } - > -} -``` - -Pruning changes payload presence (`printedTree` to `printedTreeOmittedForLength`, for example), never the status or error identity. - -### Related reads - -Use the same `{ kind, version, status, error? }` vocabulary for `read_outline`, `read_slices`, `list_directory`, and `glob`. `find_files` should return either a structured discovery result or reuse `ReadFilesResultV1`; it must not reuse the legacy array union indefinitely. - -### `FileEditResultV1` - -Introduce one shared write result for filesystem mutation tools: - -```ts -type FileEditResultV1 = { - kind: 'file_edit_result' - version: 1 - tool: - | 'str_replace' - | 'write_file' - | 'replace_range' - | 'rewrite_symbol' - | 'apply_patch' - | 'edit_transaction' - status: FilesystemResultStatus - atomic: boolean - changed: boolean - files: Array<{ - path: string - status: 'applied' | 'unchanged' | 'error' - action: 'create' | 'update' | 'delete' - patch?: string - error?: FilesystemError - }> - operations?: Array<{ - index: number - status: 'applied' | 'skipped' | 'error' - error?: FilesystemError - }> - failedOperationIndex?: number - message: string -} -``` - -For an aborted atomic batch, require `status: 'error'`, `changed: false`, all non-failing operations marked `skipped`, the failing operation marked `error`, and `failedOperationIndex`. This lets the model and harness know that successful-looking earlier matches were not committed and that the whole batch must be rebuilt from a fresh read. - -## Minimal migration sequence - -### Phase 0: Freeze legacy behavior with characterization tests - -No emitted shape changes yet. - -1. Add fixtures covering every current `read_files` representation: content, template, marker failure, null/missing override, whole+range same path, multiple ranges, range truncation, symbol success, partial symbol match, symbol miss, and symbol-only input. -2. Add `read_subtree` fixtures for full/pruned directory, full/pruned file, mixed success/error, empty paths (root), and unsafe/missing paths. -3. Add edit fixtures for success, plain error, nested error, partial non-atomic result, and aborted atomic result. - -Exact tests: - -- Update `sdk/src/__tests__/read-files.test.ts`. -- Update `packages/agent-runtime/src/__tests__/read-files-edit-state.test.ts`. -- Create `packages/agent-runtime/src/util/__tests__/render-read-files-result.test.ts`. -- Update `packages/agent-runtime/src/tools/handlers/__tests__/read-subtree.test.ts`. -- Update `packages/agent-runtime/src/util/__tests__/simplify-tool-results.test.ts`. -- Update `agents/__tests__/context-pruner.test.ts`. -- Update `cli/src/components/tools/__tests__/read-files.test.tsx` and create `cli/src/components/tools/__tests__/read-subtree.test.tsx`. -- Update `cli/src/components/tools/__tests__/str-replace.test.tsx` and `cli/src/components/tools/__tests__/apply-patch.test.tsx`. - -### Phase 1: Add shared schemas, decoders, and generated result types - -Create: - -- `common/src/tools/results/filesystem.ts`: strict v1 schemas/types, legacy v0 schemas, `decodeReadFilesResult`, `decodeReadSubtreeResult`, `decodeFileEditResult`, and status aggregation helpers. -- `common/src/tools/results/__tests__/filesystem.test.ts`: canonical parsing, legacy normalization, malformed payload rejection, idempotence, and summary-invariant tests. - -Modify: - -- `common/src/tools/params/tool/read-files.ts`: make `paths` default to `[]`; replace the current output union with `z.union([readFilesResultV1Schema, legacyReadFilesValueSchema])` during migration. -- `common/src/tools/params/tool/read-subtree.ts`: accept canonical v1 plus legacy v0. -- `common/src/tools/params/tool/read-outline.ts`, `read-slices.ts`, `find-files.ts`, `list-directory.ts`, and `glob.ts`: adopt the shared status/error vocabulary or explicitly register their legacy schemas for later conversion. -- `common/src/tools/params/tool/str-replace.ts`, `apply-patch.ts`, and `edit-transaction.ts`: export canonical-plus-legacy schemas; dependent write tools continue importing the shared schema. -- `common/src/tools/list.ts`: export named `ReadFilesResult`, `ReadSubtreeResult`, and `FileEditResult` aliases in addition to `CodebuffToolOutput`. -- `common/src/tools/compile-tool-definitions.ts`: generate `ToolResultMap` and `GetToolResult` from the JSON value schema, not only `ToolParamsMap`/`GetToolParams`. -- `common/src/tools/__tests__/compile-tool-definitions.test.ts`: assert result-map generation and discriminants. -- Regenerate `agents/types/tools.ts` and `common/src/templates/initial-agents-dir/types/tools.ts`; re-export `GetToolResult` from both agent-definition templates. - -Compatibility rule: all decoders return canonical v1 in memory, but schemas accept legacy v0. No native handler emits v1 yet. - -### Phase 2: Structure SDK reads without breaking `getFiles` or overrides - -Modify `sdk/src/tools/read-files.ts`: - -- Refactor `ReadOneFileResult` into a structured internal result with explicit error code/status. -- Add `getFilesStructured(...) => Promise`. -- Keep `getFiles(...) => Promise>` as a deprecated legacy adapter implemented from the structured reader. Preserve byte-for-byte legacy rendering for existing public callers. -- Keep `getFileForEdit` as the full-content editing path, but use structured errors internally rather than converting through display strings. - -Modify public/runtime wiring: - -- `common/src/types/contracts/client.ts`: widen `RequestFilesFn` to return `ReadFilesResultV1 | LegacyReadFilesMap`; add named `LegacyReadFilesMap` and `RequestFilesResult` types. -- `sdk/src/run.ts`: widen `overrideTools.read_files` to accept existing legacy implementations and new structured implementations. Add `filesystemResultFormat?: 'legacy-v0' | 'structured-v1'` to `OpenbuffClientOptions`, initially defaulting to legacy for public SDK users. -- `sdk/src/impl/agent-runtime.ts` and `common/src/types/contracts/agent-runtime.ts`: thread the optional format capability into runtime scoped dependencies. -- `sdk/src/index.ts`: export `getFilesStructured` and the structured result types. -- `sdk/scripts/build.ts` and `sdk/test/esm-compatibility/test-types.ts`: include/check the new public exports. - -Update tests: - -- `sdk/src/__tests__/read-files.test.ts`: assert structured per-selector results and unchanged legacy `getFiles` snapshots. -- `sdk/src/__tests__/run-file-filter.test.ts`: assert blocked/template status in both formats. -- `sdk/src/__tests__/run-handle-event.test.ts`: assert legacy default and structured opt-in event payloads. -- Add a public type compatibility case to `sdk/test/esm-compatibility/test-types.ts` showing that the old override signature still typechecks. - -### Phase 3: Normalize runtime reads and authorization - -Modify: - -- `packages/agent-runtime/src/get-file-reading-updates.ts`: accept `RequestFilesResult`, normalize legacy maps immediately, and return typed per-selector results instead of `{ path, content }[]`. -- `packages/agent-runtime/src/util/render-read-files-result.ts`: convert it into the canonical builder/legacy adapter boundary. Rename only if desired; avoiding a file rename reduces churn. -- `packages/agent-runtime/src/tools/handlers/tool/read-files.ts`: count all selectors, preserve `requestIndex`, emit canonical v1 when opted in, and derive edit-state updates from typed status/coverage rather than marker strings. -- `packages/agent-runtime/src/tools/handlers/tool/find-files.ts`: consume normalized results and either emit `ReadFilesResultV1` or adapt to its declared legacy shape. -- `packages/agent-runtime/src/tools/handlers/tool/read-subtree.ts`: build `ReadSubtreeResultV1` directly and adapt to v0 only at the compatibility boundary. -- `packages/agent-runtime/src/tools/handlers/tool/read-outline.ts` and `read-slices.ts`: stop encoding failures in `outline`/empty slices. -- `packages/agent-runtime/src/util/simplify-tool-results.ts`: preserve `kind`, `version`, status, selector, request index, path, summary, and errors. Replace only large payload fields with omission metadata. -- `packages/agent-runtime/src/tools/tool-executor.ts`: for filesystem tools, decode/validate the handler output before streaming. During the compatibility window accept v0 and normalize it; malformed native output should become a logged harness error and a valid canonical error result, never an unvalidated payload. - -Authorization rules after this phase: - -- Only `selector: 'file'`, `status: 'ok'`, and `complete: true` grants sticky whole-file authorization. -- A truncated whole read is `partial` and does not grant whole-file authorization. -- Range/symbol results never grant whole-file authorization; their exact capabilities remain the proof for scoped edits. -- Error items never clear failed-edit gates. -- Replace `Record` with a versioned authorization record in a follow-up-compatible shape, for example `{ contentHash, readAt, source: 'whole-file' }`, so external changes can invalidate it. This can land after structured results if kept as a separate risk-controlled change. - -Update tests: - -- `packages/agent-runtime/src/__tests__/read-files-edit-state.test.ts`: symbol-only counts, partial symbol match, truncated read authorization denial, marker-v0 compatibility, and canonical-v1 authorization. -- `packages/agent-runtime/src/tools/handlers/__tests__/read-subtree.test.ts`: canonical summary invariants and per-request error identity. -- `packages/agent-runtime/src/util/__tests__/simplify-tool-results.test.ts`: errors survive pruning; symbol entries do not become fake content omissions; simplification is idempotent. -- `packages/agent-runtime/src/__tests__/tool-validation-error.test.ts`: malformed native filesystem output is contained and a valid paired tool result remains. -- `packages/agent-runtime/src/tools/handlers/tool/__tests__/runtime-path-hardening.test.ts`: unsafe paths yield canonical `outside_project`/`invalid_request` codes without I/O. - -### Phase 4: Move pruner and CLI consumers to one status decoder - -Context pruning: - -- Update the inline helpers in `agents/context-pruner.ts` to recognize canonical `kind/version/status/results` first and retain the existing v0 heuristics only as a fallback for old message history. This agent's `handleSteps` is serialized, so it cannot import the common decoder directly. -- Preserve the current `toolCallId` correlation. -- Record a read fact only for `ok` items and intentionally selected `partial` items; never infer success from an unknown canonical object. -- Preserve structured error code/recovery text within the existing bounded tool-error budget. -- Update `agents/__tests__/context-pruner.test.ts` and `agents/e2e/context-pruner.e2e.test.ts` with mixed v0/v1 histories and re-compaction. - -CLI normalization: - -- Create `cli/src/utils/filesystem-tool-results.ts` as the single UI decoder for canonical and legacy results. -- Change `cli/src/types/chat.ts` so `outputRaw` is typed as `ToolResultOutput[] | LegacyPersistedToolOutput`, rather than unconstrained `unknown`. -- Keep `cli/src/utils/message-block-helpers.ts:1065-1090` responsible for storage only; do not embed tool-specific status inference there. -- Update `cli/src/components/tools/read-files.tsx` to render the normalized aggregate status and optional `ok/partial/failed` counts without recursive JSON inspection. -- Update `cli/src/components/tools/read-subtree.tsx` to render queued/pending/listed/partial/failed states and the first bounded error. -- Update `cli/src/components/tools/str-replace.tsx`, `apply-patch.tsx`, and `edit-transaction.tsx` to use `FileEditResultV1`; retain v0 fallback in the decoder. -- Update `cli/src/utils/sdk-event-handlers.ts` only if normalization is chosen at ingestion time. Prefer decoding at render/use sites so persisted old blocks remain readable. - -Tests: - -- `cli/src/components/tools/__tests__/read-files.test.tsx`: canonical ok/partial/error, symbol miss, malformed canonical error, and v0 parity. -- New `cli/src/components/tools/__tests__/read-subtree.test.tsx`: pending/ok/partial/error and pruned results. -- `cli/src/components/tools/__tests__/str-replace.test.tsx`: aborted atomic batch shows failed index and no applied diff. -- `cli/src/components/tools/__tests__/apply-patch.test.tsx`: canonical positive success and canonical rejection. -- `cli/src/utils/__tests__/message-block-helpers.test.ts`: raw v1 payload preservation. -- `cli/src/utils/__tests__/sdk-event-handlers.test.ts`: v1 result pairing by `toolCallId`. - -### Phase 5: Emit structured writes and turn retry guidance into state - -Modify the common schemas and runtime handlers for the six edit tools listed above so native results emit `FileEditResultV1` in structured mode. The SDK application functions should return the same typed value before the JSON transport wrapper. - -Key exact files: - -- `sdk/src/tools/change-file.ts`, `sdk/src/tools/apply-patch.ts`, and `sdk/src/tools/replace-range.ts`. -- `packages/agent-runtime/src/tools/handlers/tool/str-replace.ts`, `write-file.ts`, `replace-range.ts`, `rewrite-symbol.ts`, `apply-patch.ts`, and `edit-transaction.ts`. -- `common/src/tools/params/tool/str-replace.ts`, `write-file.ts`, `replace-range.ts`, `rewrite-symbol.ts`, `apply-patch.ts`, and `edit-transaction.ts`. - -For atomic `str_replace`, populate per-operation statuses from the existing batch diagnostics. The circuit breaker and model prompt can then key off `error.code`, `failedOperationIndex`, and `requiresFreshRead` instead of matching prose. Preserve the human message for readability. - -Update the existing SDK/runtime edit tests rather than creating parallel suites: - -- `sdk/src/__tests__/change-file.test.ts`, `apply-patch.test.ts`, and `replace-range.test.ts`. -- `packages/agent-runtime/src/__tests__/read-files-edit-state.test.ts` for read/edit recovery state. -- `packages/agent-runtime/src/tools/handlers/tool/__tests__/str-replace-circuit-breaker.test.ts`. -- `packages/agent-runtime/src/__tests__/process-edit-transaction.test.ts` and `apply-smart-patch.test.ts` where they consume shared results. - -### Phase 6: Default switch and eventual legacy removal - -1. Ship at least one release with SDK default `legacy-v0`, CLI opting into `structured-v1`, and dual decoders everywhere. -2. Add telemetry/log counters for decoded v0, decoded v1, malformed, and adapter fallback. Do not include file content in telemetry. -3. In the next breaking SDK release, switch the public default to `structured-v1`; retain an explicit `legacy-v0` option for one deprecation cycle. -4. Remove v0 emission only after persisted-history, SDK override, CLI, and context-pruner compatibility tests demonstrate that old payloads still decode. -5. Keep v0 decoding longer than v0 emission because saved run states and copied message histories can outlive a release. - -## Backward-compatibility strategy - -- Preserve the outer `ToolResultOutput[]` transport and print-mode event fields. -- Preserve the public `getFiles()` return type and behavior; add `getFilesStructured()` rather than changing it in place. -- Widen `overrideTools.read_files` to accept legacy or v1 returns. Existing legacy override functions remain assignable. -- Treat missing `kind/version` as v0 and decode using a narrowly documented adapter. -- Never dual-emit v0 and v1 entries in the same result; that doubles model tokens and risks duplicate counting. -- Do not infer v1 success from absence of an error. Unknown/malformed v1 is a harness error. Only the v0 adapter may use legacy heuristics. -- Keep CLI and context-pruner v0 decoders for persisted histories after structured emission becomes the default. -- Use a format capability/option rather than user-agent/version guessing. - -## Validation commands - -Run in this order after each phase that touches the named package: - -```sh -bun test common/src/tools/results common/src/tools/__tests__/compile-tool-definitions.test.ts -bun run --cwd common typecheck -bun test sdk/src/__tests__/read-files.test.ts sdk/src/__tests__/run-file-filter.test.ts sdk/src/__tests__/run-handle-event.test.ts -bun run --cwd sdk typecheck -bun test packages/agent-runtime/src/util/__tests__/simplify-tool-results.test.ts packages/agent-runtime/src/tools/handlers/__tests__/read-subtree.test.ts packages/agent-runtime/src/__tests__/read-files-edit-state.test.ts packages/agent-runtime/src/__tests__/tool-validation-error.test.ts -bun run --cwd packages/agent-runtime typecheck -bun test agents/__tests__/context-pruner.test.ts agents/e2e/context-pruner.e2e.test.ts -bun test cli/src/components/tools/__tests__/read-files.test.tsx cli/src/components/tools/__tests__/read-subtree.test.tsx cli/src/components/tools/__tests__/str-replace.test.tsx cli/src/components/tools/__tests__/apply-patch.test.tsx cli/src/utils/__tests__/sdk-event-handlers.test.ts -bun run --cwd cli typecheck -bun run --cwd sdk build -bun run --cwd sdk verify:skip-build -``` - -Use the repository's configured hooks/reviewer gate after the focused suites pass. - -## Acceptance criteria - -1. A symbol-only `read_files` call is valid without a synthetic `paths: []` field. -2. Every whole/range/symbol selector has exactly one typed result item and stable `requestIndex`. -3. Symbol misses cannot produce a zero-request/zero-failure summary. -4. `summary.ok + summary.partial + summary.failed === summary.requested` for every v1 batch. -5. Native v1 failures are never encoded solely inside `content`, `outline`, `message`, or an empty array. -6. Pruning/simplification preserves status, selector, path, request identity, and bounded error/recovery fields. -7. Whole-file authorization is granted only from complete successful whole reads; range/symbol/truncated reads do not grant it. -8. CLI read/subtree/write status is derived from normalized enums, not recursive string matching, when v1 is present. -9. An aborted atomic replacement result explicitly states `changed: false`, the failed operation index, skipped operations, and whether a fresh read is required. -10. Native filesystem outputs are schema-validated before being streamed or persisted. -11. Existing `getFiles()` callers and legacy `overrideTools.read_files` implementations continue to work unchanged during the compatibility window. -12. Old persisted v0 tool results still render and prune correctly after v1 becomes the CLI default. - -## Risks and rollback boundaries - -- The highest-risk semantic change is denying whole-file authorization for truncated reads. Land it behind focused authorization tests and, if necessary, a separate flag from result-format emission. -- Generating output types can expose existing schema/handler mismatches. Start with filesystem tools and strict canonical schemas; do not make all native tool outputs strict in one patch. -- The context-pruner cannot import common helpers because its generator is serialized. Keep its canonical decoder small and fixture-driven rather than duplicating the full general adapter. -- Public SDK event consumers may inspect `event.output[0].value` directly. The legacy-default/structured-opt-in release is the rollback boundary for this risk. -- Each phase can be reverted independently because the outer transport and v0 decoders remain stable until the final removal phase. - -## Recommended first implementation slice - -Implement Phases 0-3 for `read_files` only, plus the common decoder and CLI/pruner dual-read support. That fixes the clearest ambiguity (symbol-only counts and status-as-content) without simultaneously rewriting every edit handler. Follow immediately with `read_subtree`, then structured writes. This order gives the edit state machine trustworthy read evidence before changing write result semantics. - -## Implementation status - -Completed in this audit: - -- The recommended `read_files` v1 compatibility slice is implemented across common schemas, SDK native/override adapters, runtime normalization and authorization, CLI rendering, and context pruning. -- Native structured reads are built from typed read metadata; marker-like ordinary source content is not interpreted as status. -- Summary counts and aggregate status are validated against actual result items. -- Structured results are reconciled with requested selector index, kind, and normalized path before authorization. -- Public SDK `getFiles()` remains legacy-compatible; `getFilesStructured()` and the format option are exported and packaged-consumer verification covers ESM/CJS/types. -- Content-hash authorization, injected-filesystem realpath containment, unified edit application coordination, and accurate rollback-failure reporting landed alongside the result-contract slice. - -Still planned: - -- `ReadSubtreeResultV1` and canonical contracts for outline/slices/discovery/listing tools. -- `FileEditResultV1` emission for every mutation tool, replacing heuristic success/error decoding in the coordinator, CLI, and pruner. -- Generated output/result maps for programmatic agent tool types. -- Shared snapshot/read deduplication for overlapping selectors in one request. -- Telemetry, deprecation milestones, default switching, and eventual legacy-emission removal. diff --git a/.agents/sessions/read-write-tooling-2026-07-10/findings/ux-prompts.md b/.agents/sessions/read-write-tooling-2026-07-10/findings/ux-prompts.md deleted file mode 100644 index dd1b952435..0000000000 --- a/.agents/sessions/read-write-tooling-2026-07-10/findings/ux-prompts.md +++ /dev/null @@ -1,13 +0,0 @@ -# UX and prompt findings - -- Fixed sequential-overlap and atomic-batch recovery guidance for the model. -- Fixed generic retry wording and distinguished syntax-only recovery from stale/no-match recovery. -- Fixed CLI edit failures to retain a bounded head-and-tail diagnostic with an actionable recovery line. -- Made context pruning correlate tool calls/results and persist only successful read/edit facts. -- Added accurate queued/pending/applied/failed patch rendering and visible queued/pending/read/partial/failed path/range/symbol read states. -- Required positive patch success evidence, rejected malformed/nested-error envelopes, and suppressed unapplied diffs. -- Prevented negated success wording from being persisted as an edit fact during context pruning. -- CLI and context pruning now prefer canonical `read_files` status/path/error fields while retaining legacy-history fallback. -- Simplification preserves canonical result identity and bounded error information instead of collapsing selector failures into fake content entries. -- Empty edit application output now renders as an unconfirmed rejection rather than synthesized success, while explicit rejection diagnostics remain visible. -- Remaining architectural work is tracked in the main report: structured edit results, coherent overlapping-selector snapshots, broader filesystem v1 coverage, and the compatibility/deprecation lifecycle. diff --git a/.agents/sessions/review-gate-correctness/LESSONS.md b/.agents/sessions/review-gate-correctness/LESSONS.md deleted file mode 100644 index 0bfba4f827..0000000000 --- a/.agents/sessions/review-gate-correctness/LESSONS.md +++ /dev/null @@ -1,100 +0,0 @@ -# Lessons — review-gate correctness & convergence - -Companion to `PLAN.md` (design) and `STATUS.md` (progress). This file holds decisions and gotchas that outlived the slice that produced them. - -## Decision record: T1.4d — embedder guide fallback - -**Status: implemented** as the hybrid below, with recovery keyed PER POINTER. `common/src/util/guides.ts` owns the guide→body tables and detection; one `ON_DEMAND_GUIDE_FALLBACK_` placeholder per relocated guide in `packages/agent-runtime/src/templates/{types,strings}.ts` is the provider surface; base2 appends, after its pointers, exactly the placeholders whose pointers that mode actually emitted. The record is kept in full because the two rejected framings are the ones a future reader will reach for first. - -Source: `architect` specialist, run against the working tree at `v3:4a19a075615be`. `PLAN.md` framed T1.4d as a choice between (A) gate the disclosure default to workspaces containing `agents/guides/`, (B) have `disclose()` emit the pointer plus an inline section copy when the guide is unreachable, or (C) keep the compact degrade clause only. **Both A and B are wrong as framed, and the defect is more severe than "open follow-up" implied.** - -### The defect is worse than described - -`PLAN.md` says the pointer read "fails" in an embedder workspace. The stronger finding: the advertised read contract is **unsatisfiable by construction**, and no published artifact ships the guides. - -- `cli/release/package.json` and `cli/release-staging/package.json` publish only `index.js`, `http.js`, `postinstall.js`, `README.md`. `sdk/package.json` publishes only `dist`, `README.md`, `CHANGELOG.md`. `cli/scripts/build-binary.ts` copies native libs and wasm only — no `agents/guides` copy step. -- Path rewriting cannot rescue it: `normalizeToolPath` (`packages/agent-runtime/src/tools/handlers/tool/write-file.ts`) rejects absolute and drive-relative paths and enforces project-relative, so a pointer cannot be redirected at a package-root guide directory. Content injection is the only delivery mechanism short of adding the guides to a published `files` list. -- The same defect exists independently of base2: `packages/agent-runtime/src/system-prompt/prompts.ts` emits an `agents/guides/knowledge-files.md` pointer from the runtime package. - -### Why option A cannot work at the layer PLAN.md implies - -Not merely because prompt assembly is synchronous. `cli/scripts/prebuild-agents.ts` imports every module under `agents/` and `JSON.stringify`s the **resolved** definitions into `cli/src/agents/bundled-agents.generated.ts` at CLI build time — inside the openbuff worktree. `cli/package.json` wires that into the build and `cli/src/utils/local-agent-registry.ts` consumes the bundle in the shipped CLI. So any `createBase2`-layer probe is frozen at build time and would resolve "guides present" for **every** embedder: a guaranteed false negative, independent of the sync/async question. - -### Why option B cannot ship as authored text - -`agents/__tests__/base2-progressive-disclosure.test.ts` measures `authoredSurface` (systemPrompt + instructionsPrompt + stepPrompt) **before** placeholder injection and asserts `(off - on) / off >= 0.25`. The disclosed/explicit-off delta _is_ the six section bodies, so inlining them as authored text destroys the acceptance metric it was created to protect. - -### Recommended: hybrid C + runtime placeholder - -Keep today's authored surface (compact pointer + "If that guide is unavailable" clause) and recover full bodies through an **additive** runtime placeholder whose provider does the detection. - -- **Detection belongs in** `packages/agent-runtime/src/templates/strings.ts` `toInject` — async-capable, filesystem-capable, keyed on the embedder's `fileContext.projectRoot`. `PATTERNS_INDEX` (with `common/src/util/patterns.ts`) and `FRONTEND_SECTION` are exact precedents, including the collapse-to-empty-string behavior. -- **Single-sourcing:** re-home the six section bodies into `common/` and re-export them unchanged from `agents/base2/quality-prompt-section.ts`. Re-export rather than copy: it keeps `qualitySection` byte-identical under `quality-prompt-snapshot.test.ts` and keeps `review-rubric-parity.test.ts` the single drift owner. `agent-runtime` must not import from `agents/` (`packages/agent-runtime/src/util/base2-tool-tiers.ts` documents this), which is why the bodies move rather than being imported. -- **Shape:** add `ON_DEMAND_GUIDE_FALLBACK` to `placeholderNames` (`packages/agent-runtime/src/templates/types.ts`); in `common/src/util/guides.ts` expose `GUIDE_FALLBACK_SECTIONS`, `findMissingGuides(projectRoot, logger)`, `formatGuideFallbackSections({ missing })` returning `''` when nothing is missing. -- **Additive, never a replacement.** If the placeholder replaces `guideSections` instead of following it, the pointer-presence assertions and the >=25% metric both become **vacuous rather than preserved**. The provider must return `''` in-repo so the resolved in-repo prompt stays byte-identical to today. -- **Keep every degrade clause verbatim.** It is unverified that all embedder entry points (notably SDK-direct consumers of the bundled definitions) run `injectPlaceholders`, so the compact inline clause remains the last line of defense. - -### Metric consequence - -The >=25% authored reduction survives byte-for-byte, because `authoredSurface` is measured pre-injection and a placeholder is a short marker. But that metric is then **structurally blind** to resolved-prompt regrowth in embedder workspaces. Add (do not replace) a resolved-surface budget: inject against both the repo root and a synthetic guide-less temp root, and assert the guide-less resolved surface is no larger than the resolved explicit-off surface. - -### Falsifying test - -Inject placeholders for `createBase2('default')` against a temp root with no `agents/guides/` and assert every relocated body appears in the resolved prompt; inject against the repo root and assert no body appears and the >=25% reduction still holds. Derive the section list from `GUIDE_POINTERS` so the loop cannot pass vacuously. - -### Unknowns the architect could not close - -- Exact token counts per option (requires running `countTokens`). -- Whether every embedder entry point runs `injectPlaceholders` — only the `strings.ts` provider table was inspected. **Still open**, which is why every pointer keeps its compact degrade clause as the last line of defense. -- Whether any publish pipeline outside `cli/release*`, `sdk/package.json`, and `build-binary.ts` copies `agents/guides/*.md`. Evidence is strongly negative but not exhaustive. - -### What implementation added to the record - -- The re-home is a **move plus re-export**, not a copy: `agents/base2/quality-prompt-section.ts` now re-exports the six bodies from `common/src/constants/prompt-sections.ts`, so `qualitySection` stays byte-identical under `quality-prompt-snapshot.test.ts` and every existing consumer import path is unchanged. `gateAwarenessSection` deliberately did NOT move — it is not relocatable to a guide, so it has no fallback body. -- `findMissingGuides` returns `[]` for a falsy or non-string `projectRoot`. "Unknown root" must not mean "everything is missing", or every prompt formatted without a real root regrows by six full sections. -- The resolved-surface budget the architect asked for landed as its own case: inject against a guide-less temp root and against the repo root, and compare against the resolved explicit-off surface. The pre-injection >=25% authored metric is structurally blind to embedder-workspace regrowth, so it was kept AND supplemented rather than replaced. -- `GUIDE_FALLBACK_SECTIONS` is keyed by plain `string`, not base2's `GuidePath` union — `common/` cannot import from `agents/`. The two drift-guard assertions comparing pointer paths to table keys therefore need `String(guide)` widening; comparing the narrower union to `string[]` has no matching `toEqual` overload and fails typecheck rather than at runtime. - -### Recovery must mirror the mode's exclusions (review repair) - -The first implementation emitted ONE `ON_DEMAND_GUIDE_FALLBACK` placeholder whose provider re-inlined all six bodies. That is wrong for two mode-specific reasons, both found in review: - -- **Plan mode omits git-discipline deliberately** (`!planOnly && disclose(GUIDE_PATHS.gitDiscipline)`, pinned by `base2-progressive-disclosure.test.ts`). An all-six recovery handed a guide-less embedder commit/push guidance back in a read-only mode. Fix: one placeholder per pointer (`GuidePointerRow.fallbackPlaceholder`), emitted from the same `buildArray` entry that emits the pointer, so an omitted pointer omits its recovery by construction. -- **The broad-audit body is clause-parameterized.** Plan mode's pointer tail says "do not implement", so recovering `buildBroadAuditSection('proceed to implementation or the answer')` there produced directly contradictory finalize instructions. Fix: `BROAD_AUDIT_FALLBACK_SECTIONS` keyed by `BroadAuditFinalizeClause`, plus a plan-clause placeholder base2 substitutes in plan mode. `GUIDE_FALLBACK_SECTIONS` keeps the implementation variant as the table default because that is what the guide file documents. - -Two further consequences of the same review: - -- **Recovered bodies are recorded in the shared `ContextBudgetLedger`** (`applyMeasure`, category `systemPrompt`, label `guide-fallback:`), the way `getProjectFileTreePrompt`/`getGitChangesPrompt` do. They are the largest block this path adds, so an unrecorded block silently under-counts an embedder's context budget. A collapsed (in-repo) block records nothing, which is what keeps the ledger honest. -- **`findMissingGuides` has no try/catch and no `logger`.** `fs.existsSync` reports a failed probe as `false` instead of throwing and `path.join` only ever sees the type-guarded root plus a literal table key, so the guard was unreachable dead code with an unreachable `logger?.warn` inside it. - -The filesystem probe is memoized once per formatted prompt: six providers now ask the same question, and each one runs only when its placeholder is present in the prompt. - -### Scope note - -This is cross-package (`common/`, `packages/agent-runtime/`, `agents/`, plus a new test), unlike every other Tier 1 item which stayed inside `agents/`. Schedule it as its own slice rather than bundling it with the reviewer-loop work. - -## Gotchas worth carrying forward - -**An advisory channel is only real once it has a display surface on every path that persists it.** The first advisory slice wrote `receipt.advisories` from all three reviewer families but rendered them only on the gate-pass `` block, so intermediate `NON_BLOCKING` receipts and every security/specialist receipt stored advisories invisibly. Review caught the prompt/behavior mismatch ("shown to the user" vs shown only on pass). Two valid fixes exist — narrow the claim or add the surfaces — and they are not equivalent: adding surfaces is the one that keeps the reviewer's mental model true. - -**Advisory display on the aux-pass paths must be conditional on a non-empty list.** `base2.test.ts` and `gate-lifecycle.e2e.test.ts` advance the generator yield by yield, so an unconditional `add_message` on a passing security or specialist gate shifts every subsequent expectation and fails as a confusing off-by-one yield mismatch rather than as "a new message appeared". - -**On the blocker/repair path there is no receipt to read yet.** `recordSuccessfulReviewReceipt` runs only once a finalization verdict exists, so that surface must read `collectReviewerAdvisories(reviewerToolResult)` directly. Using the shared collector (rather than a second inline `result.advisories` read) is what keeps the displayed semantics identical to the persisted ones. - -**Reconstructed inline helpers need their whole closure.** `base2.test.ts` rebuilds `formatGateStateBlock` with `extractInlineFunctionSource` + `new Function`. Extracting the advisory bounding into a shared `boundAdvisoryLines` helper broke that test at call time until the helper was added to the reconstruction list — the typecheck cannot see it, because the reconstruction is string-based. Any new inline helper called by an already-reconstructed one must be added to the same list. - -**Sanitize before the bounds check, and collapse whitespace before stripping controls.** In `formatGateStateBlock` and the CLI's `parseGateStateAdvisories`, collapsing `\s+` first turns tabs/newlines into spaces; stripping `/[\x00-\x1f\x7f]/` first would delete them and glue words together. Applying the 240-char cap after the strip is what makes the bound describe the text actually emitted. On the parse side, sanitizing after the emptiness check would let a controls-only entry pass as non-empty. - -**T1.5's id-keying only bites when the reviewer supplies an id.** Minted `RF--` ids embed the blocker's position in the round's list, so identical text at a different index yields a different id. They are deliberately excluded from `::id:` keying. Consequence for T1.2(c): the round ledger must instruct **verbatim** re-raise text regardless of ids, because bare-string findings still condone on `(class, text)`. - -**The condone/merge logic cannot import from `gate-reviewer.ts`.** It lives inside the serialized `handleSteps` generator (`.toString()` + `new Function(...)`), so module-scope closures are unavailable at reconstruction time. Any change is duplicated by hand into the `` region; `scripts/generate-gate-helpers.ts` is the source of truth and the parity tests enforce it. Always re-run `--check` after touching `base2.ts`. - -**`lastPinnedStateMessage` is an invalidation sentinel, not a history.** `markActiveWorkStateChanged` resets it to `''` on every gate-state write. Any "did this change since last time" comparison must use a separate emitted-value baseline (`lastEmittedPinnedStateMessage`) or the branch is dead. This exact mistake shipped once and was caught by review. - -**A reviewer "crash" is not always a code defect.** Four of this session's gate stalls were a provider switch, a user interrupt, a provider billing error (`预扣费额度失败`, insufficient prepaid credit), and a transient `Unable to connect` whose provider host answered HTTP 200 on a probe moments later. None warranted a bypass. Read the crash detail — and probe the host — before proposing `BYPASS REVIEWER`. - -**A stable review bundle is worth more than an extra slice of progress.** After a specialist crash, the correct move is to end the turn without editing: any edit moves the worktree and invalidates the bundle the specialist must attest against, re-triggering the same snapshot-mismatch refresh that preceded the crash. - -**Do not hand mutating git commands to a subagent.** A basher spawn ran `git stash push --include-untracked` despite an explicit read-only instruction, reverting the entire uncommitted working set; recovery was `git stash apply stash@{0}`. It also produced a misleading test failure (`Unable to find inline stripReviewerVerdictPrefix declaration`) because the test file reconstructs helpers from a `base2.ts` that had just been reverted. Use `git show HEAD:` or the read tools for historical comparison instead. - -**`.base2-test-scratch` has a pre-existing cleanup race.** Two `base2.test.ts` cases can `mkdtemp` into the shared scratch root after `afterAll` removes it, producing `ENOENT` as an unhandled-between-tests error with 0 failures. Predates this work; unrelated to any gate change. diff --git a/.agents/sessions/review-gate-correctness/PLAN.md b/.agents/sessions/review-gate-correctness/PLAN.md deleted file mode 100644 index 64039d445b..0000000000 --- a/.agents/sessions/review-gate-correctness/PLAN.md +++ /dev/null @@ -1,293 +0,0 @@ -# Review-gate correctness & convergence plan (rev 3) - -> **Progress is tracked outside this file.** This document is the design: conclusions, evidence, tiers, sequencing. It carries no status markers by design, so nothing here goes stale as work lands. -> -> - `STATUS.md` — what has landed (with verified code citations), what is open and why, and resume instructions. -> - `LESSONS.md` — decision records (incl. the T1.4d architect decision) and gotchas. -> -> Tier 0 is closed and Tier 1 is mostly closed; T1.2(a), T1.2(c), and T1.4d remain open. Tier 2/3 are still gated on the sequencing step-8 re-measurement below. - -Rev 3 supersedes rev 2 after adversarial review by `architect` and `thinker`. Eleven architect findings and three thinker findings are folded in. Material changes from rev 2: - -- **New Tier 0** — a live authority hole: the condoned-pass branch bypasses the gate's own coverage/requirement hard rules. Found by the architect while verifying rev 2's evidence; independently confirmed at `base2.ts:4650-4655`. -- **H2's identity half is restored** as T1.5. Rev 2's decisive drop reason ("adds another persisted structure") was **factually false** — `openReviewerFindings` already is an id-keyed ledger with rehydration wired. Only the changed-files admissibility rule stays dropped. -- **T1.2(c) is withheld** until id-keyed condoning lands. As written it defeats the only convergence mechanism rev 2 keeps. -- **T1.3 resequenced before T1.2(a)** — the "advisory channel" T1.2(a) writes into does not exist yet. -- **New T1.6** — fingerprint cycle detection. The existing no-progress guard compares only against the immediately preceding fingerprint, so an A→B→A oscillation never trips it. Neither rev 1 nor rev 2 saw this. -- Tier gating stated once; T1.1 durability resolved; T1.3 migration hazard named. - -## Governing conclusions - -1. **The current prompt makes termination logically impossible.** "Find ways to improve the code changes" is a search that succeeds on any non-trivial file; "Do not emit `LOOKS_GOOD` while any findings remain" then forbids stopping. The acceptance predicate is unsatisfiable by construction — a divergence proof readable off `code-reviewer.ts` alone, not a probabilistic argument. Fixing the generator (T1.2) is the highest-confidence item here. -2. **Cross-round memory is an optimization for expected-case termination and a necessity for any worst-case bound.** With a stateless reviewer there is no monotone decreasing quantity to induct on. Rev 2's "fix the generator and drop the ledger" was half right: fix the generator _and_ repair the memory that already exists. -3. **Locality does not bound the observed pathology.** New findings decompose into repair-induced (scales with edit size; contracts as repairs shrink) and **rediscovery** (findings about code the previous sample happened not to mention — not caused by the edit, so locality does nothing). Rediscovery dominates under an unbounded rubric. This is why "repair edits are small" never saved this loop, and why bounding the rubric precedes any output filtering. -4. **Enforcement belongs on the orchestrator side, never in new required reviewer output fields.** More required fields raise schema-non-compliance probability, and non-compliance routes to `currentPhase = 'blocked'` after one bounded retry — worse than an extra nit round. The condone credit is orchestrator-owned and costs zero protocol risk. - -## Prior attempts (read before proposing anything) - -| Commit | Subject | -| ----------- | -------------------------------------------------------------------- | -| `6b5b035db` | Harden gate TUI and reviewer loop convergence | -| `bf31b9f3d` | harden the base2 specialist reviewer gate and its repair loop | -| `4573e2753` | Harden reviewer and validation gate against stalls, loops, and churn | -| `2d6ad7c27` | fix structured reviewer retry loop | -| `933dd440e` | add explicit MAX_REVIEWER_REPAIR_ROUNDS cap to reviewer-repair loop | -| `ff2ff4e24` | Run reviewer repair loop until findings clear | - -The last two are opposite directions — a cap added, then removed for "run until findings clear." The fossil is still in `common/src/util/gate-repair-budgets.ts`: `DEFAULT_MAX_REVIEWER_REPAIR_ROUNDS` is `null`, commented _"@deprecated Omitted option/env means unlimited, not these values."_ - -Every one of those fixes was unfalsifiable when it shipped: no per-round finding telemetry exists, so each author decided on judgment. **Do not add a seventh judgment-based fix without instrumentation.** - -## Evidence base (verified reads) - -| Fact | Location | -| ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- | -| Only `LOOKS_GOOD` finalizes; `NON_BLOCKING` is repair fuel | `gate-reviewer.ts:464-499`, `:319-334` | -| **Condoned pass pre-sets the verdict** | `base2.ts:4121-4127` | -| **…and the derivation is guarded, so the hard rules never run** | `base2.ts:4650-4655` (`if (!reviewerFinalizationVerdict)`) | -| `getReviewerFinalizationVerdict` is the only enforcement of `coverage: missing` and in-scope `missing`/`uncertain` | `gate-reviewer.ts:471-491` | -| Those hard rules are emitted as plain strings the condone capture can absorb | `gate-reviewer.ts:288-315` vs `base2.ts:4529-4539` | -| `recordSuccessfulReviewReceipt` returns early on `BLOCKING` ⇒ all-condoned BLOCKING round passes with **no receipt** | `base2.ts:6927-6932` | -| Fabricated verdict is persisted | `base2.ts:4842-4843` | -| Reviewer prompt is an unbounded generator, and self-contradicts | `code-reviewer.ts` instructionsPrompt | -| Condoning is exact-string after stripping the verdict prefix | `base2.ts:4093-4103`, capture `:4529-4539`, cleared `:4828` | -| Prefix stripping is severity-blind ⇒ NON_BLOCKING→BLOCKING escalation of identical text is swallowed | `base2.ts:4097-4101`, `:4530-4532` | -| **`openReviewerFindings` already is an id-keyed ledger** (`id`, `gateId`, `text`, `status: open\|resolved\|condoned`, `files[]`, `snapshotFingerprint`, `reviewer`) | `gate-state.ts:111-122` | -| …with rehydration already wired | `base2.ts:904` | -| …and `mergeReviewerFindings` already flips records to `condoned` | `base2.ts:5192-5206` | -| …and repair reconciliation is already id-keyed | `base2.ts:4493-4511`, `:4525-4528` | -| Condone credit is unverified self-report; `reviewerRepairHasProgress` (any changed file) short-circuits completeness | `base2.ts:4498-4511` | -| Finding identity has three owners: reviewer-supplied id, FNV hash of text, raw condone text | `gate-reviewer.ts:627-634`; `base2.ts:7459-7466`; `:4090-4101` | -| code-reviewer path is the only one that does **not** correlate reviewer ids (security/specialist do) | `base2.ts:4218-4232` vs `:2047-2054`, `:2688-2696` | -| Object findings render as `[id] summary` ⇒ ids churn every text-keyed identity | `gate-reviewer.ts:548-560` | -| `retainedBlockers` matches by substring | `base2.ts:5214-5221` | -| No-progress guard compares **only** the immediately preceding fingerprint | `base2.ts:4551-4575`; specialist `:2962-2995` | -| Bare-string findings never become `findingRecords` (`if (!id \|\| !text) return []`) | `gate-reviewer.ts:623-634` | -| ⇒ a `LOOKS_GOOD` receipt with bare-string nits records `findings: []` | `base2.ts:7042`, `:7069` | -| Receipt state + parsers already accept optional `severity`/`dimension` | `gate-state.ts:32-43`; `gate-reviewer.ts:639-644` | -| Findings carry `evidence: string[]` prose; `files[]` exists but is populated with the **whole pending set** | `code-reviewer.ts` schema; `base2.ts:4183`, `:4224` | -| Rubric reaches the model only as a pointer, uniquely targeting a `.ts` module | `base2.ts:154-155`, `:516`; siblings `:146-153` | -| Dead-env-canary trap: with a default-ON flag, `envFlag \|\| DEFAULT` can never read the env | `base2.ts:87-94` | -| Repair budgets resolve missing→null→unlimited; surfaced in `/context` | `base2.ts:100-123`; `gate-repair-budgets.ts:28-46`; `cli/src/commands/context.ts` | - ---- - -# Tier 0 — live defect, fix before anything else - -## T0.1 — The condoned pass bypasses the gate's own hard rules - -**This is a defect in shipped code, not in a proposal.** It is also the only item here that can let a genuinely incomplete change through, so it precedes every improvement. - -Mechanism, end to end: - -1. `collectReviewerBlockers` emits coverage/requirement hard rules as plain strings, e.g. `"BLOCKING: test coverage missing for changed behavior (add a case to the relevant *.test.ts)"` (`gate-reviewer.ts:288-315`). -2. A prior repair round reported some finding addressed, so its text entered `condonedFindingTexts` (`base2.ts:4529-4539`). -3. On re-review the condone filter suppresses every collected blocker, so `blockers.length === 0` while `collectedBlockers.length > 0` (`:4093-4103`). -4. That branch assigns `reviewerFinalizationVerdict = 'LOOKS_GOOD'` directly (`:4121-4127`). -5. The derivation at `:4650-4655` is guarded by `if (!reviewerFinalizationVerdict)`, so **`getReviewerFinalizationVerdict` never runs** — and it is the only place `coverage: "missing"` and in-scope `requirementCoverage` `missing`/`uncertain` are enforced (`gate-reviewer.ts:471-491`). -6. `recordSuccessfulReviewReceipt` returns early for a `BLOCKING` verdict (`:6927-6932`), so the pass is credited with **no review receipt at all**, and the fabricated verdict is persisted as `gatePassedReviewerVerdict` (`:4842-4843`). - -Fix — make condoning **blocker suppression only, never verdict authority**: - -``` -// at the condoned-pass branch (base2.ts:4121) -// do NOT assign reviewerFinalizationVerdict here. -// Suppress blockers, then let the normal derivation decide: -const condonedVerdict = getReviewerFinalizationVerdict(reviewerToolResult) -if (condonedVerdict === 'LOOKS_GOOD') { ...credit as today... } -else { keep the round open — coverage/requirement rules still stand } -``` - -Remove the `if (!reviewerFinalizationVerdict)` guard's dependence on the condone path, or at minimum re-run the coverage and requirement checks before crediting. Never let the condone path fabricate a verdict. - -Note the interaction with T1.4b: rev 2 promised to teach implementers that "uncertain blocks exactly like missing," which is currently _not_ reliably true. Fixing the code and the guideline together avoids documenting an aspiration. - -Acceptance: an all-condoned round whose receipt carries `coverage: "missing"` or an in-scope `uncertain` requirement does **not** pass. Add e2e coverage in `agents/e2e/gate-lifecycle.e2e.test.ts`. - -## T0.2 — Condone credit must be backed by evidence - -Condoning is credited purely from the repair-editor's self-reported `findingsAddressed`, with no check that the claimed finding was touched. Worse, `reviewerRepairHasProgress` — true when _any_ changed file path exists — short-circuits the completeness check, so a receipt with `status: 'blocked'` and unaddressed ids still condones every id it lists (`base2.ts:4498-4511`, `:4525-4539`). - -This is the rubber-stamping risk, located in the orchestrator rather than the reviewer. Fix (zero protocol-failure cost, because no reviewer output changes): - -- Only condone a finding id the receipt both lists in `findingsAddressed` **and** backs with at least one `changedFiles` entry. -- Require `receipt.status === 'completed'` for condoning even when other progress exists. Keep `reviewerRepairHasProgress` for loop-continuation decisions; do not let it authorize condoning. -- Log rejected condone claims via `emitGateTelemetry` so T1.1 can count them. - -Acceptance: a repair receipt that lists ids without changed files condones nothing; the finding stays open. - ---- - -# Tier 1 — ship these - -Gate statement, stated once: **T1.1 data gates Tier 2. Tier 3 follows re-measurement.** No other tier gate exists. - -## T1.1 — Per-round telemetry + shadow mode ← first - -`emitGateTelemetry` already carries `repairRound` and `skipReason`. Add per round: - -- `findingCount`; severity histogram once T1.3 lands (before that, one `unlabeled` bucket) -- `newFindingCount` vs `carriedFindingCount` -- reviewer wall-clock; reviewed-file count -- terminal outcome: `passed` / `blocked` / which `lastReviewerGateSkipReason` -- **rejected condone claims** from T0.2 - -**Durability (architect finding 9).** Rev 2 defined round-over-round comparison against generator locals while forbidding new persisted state — so the counterfactual would silently degrade across a serialized turn, which is exactly when the metric matters (`base2.ts:4863` resets round count only on gate pass; rounds are expected to span serialization). Resolution: derive the comparison from **already-persisted** state — the prior round's `openReviewerFindings` plus `reviewReceipts`. If that proves insufficient, grant T1.1 one explicit persisted field with a `??=` default at `base2.ts:889-917` and say so; do not leave it on locals. - -**Shadow mode.** Compute what a severity threshold _would_ decide and log it without enforcing: - -``` -would-suppress: 4 of 6 findings (severity low / unlabeled-hygiene) -would-have-passed-at-round: 1 (actual: 4) -``` - -Acceptance: after N real turns you can state the nit-driven share of rounds and what a threshold would have let through. - -## T1.2 — Fix the generator - -The highest-confidence behavioral item: it removes a logical obstruction, not a heuristic one. - -**(a) Resolve the contradiction.** One rule: `LOOKS_GOOD` when nothing requires a code or contract change; cosmetic observations go to the advisory channel. **Sequenced after T1.3** (architect finding 4) — bare-string findings on a `LOOKS_GOOD` receipt are surfaced nowhere today (`gate-reviewer.ts:275-334` emits no blockers for `LOOKS_GOOD`; `:623-634` drops findings without ids; `base2.ts:7042`/`:7069` therefore record `findings: []`). Shipping (a) before the channel exists silently discards the nits. Acceptance gate: a `LOOKS_GOOD` receipt's advisory findings must appear in `reviewReceipts` and in the CLI. - -**(b) Bound the generator.** Replace "find ways to improve" with a finite completeness criterion: enumerate every violation requiring a change in one pass, then stop. Make "do not drip-feed" the framing rather than an aside. Per conclusion 3, this is what shrinks the rediscovery term; nothing else in this plan does. - -Target the three convergence conditions explicitly: a satisfiable empty set ("nothing requires a change", not "nothing could be improved"); monotonicity under repair (findings must be over properties a repair can clear — taste-based findings are not monotone); and low churn sensitivity (violation-driven, not proportional to code in view). - -**(c) Round ledger as a verification checklist — WITHHELD until T1.5.** The intended prompt: - -``` -Repair round: N. This is a re-review. -Findings raised earlier and reported addressed — verify each is genuinely -fixed and cite the line that fixes it. If a fix is wrong or incomplete, -re-raise it and say why: - - -Files the last repair changed: -``` - -**Why withheld (architect finding 2, decisive).** Condoning is exact-string equality (`base2.ts:4093-4103`). "Re-raise it and say why" changes the text, so the re-raise escapes the filter and re-enters the repair loop — while a _verbatim_ re-raise of a genuinely unfixed blocker gets swallowed. T1.2(c) is therefore simultaneously suppressed and loop-amplifying depending on wording: the exact "reword problem" rev 2 cited as grounds to drop H2, reintroduced via prompt. Ship (c) only after T1.5 re-keys condoning on finding id — or, as a fallback, instruct re-raises to repeat the original text verbatim with the reason on a separate line the matcher strips. Also note: shipping (c) before T0.1 would worsen the escaped-defect path. - -## T1.3 — Optional `id` / `severity` / `dimension` metadata - -No thresholding. Nothing gates on it. Extend `code-reviewer.ts` findings to accept the specialist-style object shape **alongside** bare strings, with new fields **optional, never required** (conclusion 4). The plumbing already exists (`gate-state.ts:32-43`, `gate-reviewer.ts:639-644`); only the emitting schema is missing. - -**Also route code-reviewer through id correlation (architect finding 7).** Finding identity currently has three owners, and the code-reviewer path always uses `buildReviewerFindingId(text, index)` (`base2.ts:4218-4232`) while security (`:2047-2054`) and specialist (`:2688-2696`) paths correlate reviewer-supplied ids via `record?.id ?? buildReviewerFindingId(...)`. Without this, T1.3's ids never reach `openReviewerFindings` and T1.5 has nothing to key on. Add the parity assertion to `gate-reviewer` tests. - -**Migration hazard (architect finding 8).** Once findings carry ids, the parser renders each as `[id] summary` (`gate-reviewer.ts:548-560`), changing blocker text — which churns `condonedFindingTexts`, the FNV `buildReviewerFindingId` hashes, and `mergeReviewerFindings`' substring retention (`base2.ts:5214-5221`). Handle sessions mid-loop across the change: either clear `condonedFindingTexts`/`openReviewerFindings` on shape mismatch, or make the matcher strip a leading `[id] ` token. Rollback must not leave `[id] `-prefixed condone keys that can never match. - -Acceptance: severity/dimension appear in telemetry and the advisory list; gate decisions byte-identical; legacy bare-string findings still parse and still **block** (never silently downgraded to advisory). - -## T1.4 — Fix the guidelines - -### T1.4a — the rubric barely reaches the model - -`progressivePromptDisclosure` defaults ON, so `base2.ts:516` emits only `preReviewSelfCheckPointer`, and that pointer (`:154-155`) uniquely targets a TypeScript module while every sibling targets `agents/guides/*.md` (`:146-153`). - -- Create `agents/guides/pre-review-self-check.md` with the full rubric (T1.4b content). -- Repoint to that guide. -- Extend the pointer/guide pairs at `agents/__tests__/base2-progressive-disclosure.test.ts:186-192`. -- Leave `agents/editor/editor.ts:220` interpolating the section in full. - -### T1.4b — mirror what actually blocks - -Add to `preReviewSelfCheckSection`: requirement coverage (`uncertain` blocks like `missing` — subject to T0.1 making that true); file attestation (every pending file read and accounted for; changed tests are first-class targets); coverage naming (name the exact test file and case; `coverage: "missing"` auto-blocks); advisory vs blocking (cosmetic observations do not hold the turn; do not pre-emptively refactor for style). - -Keep the existing 7 bullets. The section is explicitly not byte-frozen (`agents/__tests__/quality-prompt-snapshot.test.ts:59-72`); extend those topic assertions. **Do not edit `qualitySection`** — byte-frozen with a snapshot test and duplicated into `agents/guides/code-craftsmanship.md`. - -### T1.4c — stop rubric/reviewer drift - -Add `agents/__tests__/review-rubric-parity.test.ts` asserting every blocking rule in `code-reviewer.ts` has a matching topic in `preReviewSelfCheckSection` (keyword table: `requirementCoverage`, `uncertain`, `coverage: missing`, `reviewedFiles`, the five dimensions). Follows the `gate-helpers-freshness` / `gate-reviewer-parity` precedent. - -### T1.4d — inline fallback for guide pointers in external workspaces (open) - -Progressive prompt disclosure defaults ON, and every guide pointer names a path under `agents/guides/` that `read_files` resolves against the user's workspace root. Inside the openbuff repo that resolves; in any embedder workspace the read fails and the model silently loses all five relocated sections. Resolution options, not yet chosen: gate the default to workspaces that actually contain `agents/guides/`, or have `disclose()` emit the pointer plus an inline copy of the section when the guide is unreachable. Partially mitigated: every pointer now carries an explicit "if that guide is unavailable" clause with the compact inline rules, so an embedder degrades to summarized guidance instead of a failed read; emitting the full section bodies on an unreachable guide remains open. The guide-pointer comment block in `agents/base2/base2.ts` references this item. - -## T1.5 — Re-key condoning on finding id (H2's identity half, restored) - -Rev 2 dropped this claiming it "adds another persisted structure with its own migration and desync modes." **That was false.** `openReviewerFindings` (`gate-state.ts:111-122`) already carries `id`, `gateId`, `text`, `status: 'open' | 'resolved' | 'condoned'`, `files[]`, `snapshotFingerprint`, `reviewer`; rehydration is wired at `base2.ts:904`; `mergeReviewerFindings` already sets `condoned` (`:5192-5206`); repair reconciliation is already id-keyed (`:4493-4511`). This is a refactor of existing state, not new state. - -Changes: - -- Condone by `openReviewerFindings[].id` instead of raw text, in both the filter (`:4090-4101`) and `mergeReviewerFindings` (`:5197-5206`). -- **Key on (verdict class, id), not prefix-stripped text** (architect finding 6). Today both prefixes map to one key, so a NON_BLOCKING finding escalated to BLOCKING with identical text is silently suppressed and can trigger the all-condoned pass. -- Keep reading legacy `condonedFindingTexts` on resume so an in-flight session does not lose convergence progress and restart the loop. -- Unblocks T1.2(c). - -**Still dropped: the changed-files admissibility rule.** Rev 2's reason was imprecise (architect finding 10) — the repair side is structured `{ path: string }[]` (`base2.ts:4500-4503`) and `files[]` exists on every finding (`gate-state.ts:117`); the real defect is that it is populated with the entire pending set (`:4183`, `:4224`), so it carries no per-finding attribution. That makes admissibility a _field-semantics_ problem, not an impossibility. Revisit only if T1.1 shows persistent `newFindingCount > 0` on untouched files after T1.2 — and if so, populate `files[]` from the finding's cited path first. - -## T1.6 — Fingerprint cycle detection (new; missed by rev 1 and rev 2) - -The no-progress guard compares only against the **immediately preceding** fingerprint (`base2.ts:4551-4575`), so an A→B→A oscillation changes the fingerprint every round and never trips it. Keep a `Set` of snapshot fingerprints seen this turn in `handleSteps` **loop scope** — no persisted state — and fail closed on a repeat. - -This is strictly stronger than a round cap and is **not** a re-litigation of `933dd440e`/`ff2ff4e24`: it fires on demonstrated non-progress, not a guessed budget. Apply the same treatment to the specialist loop (`:2962-2995`). - -Acceptance: a synthetic A→B→A repair sequence terminates with a `reviewer-repair-cycle` skip reason instead of looping. - ---- - -# Tier 2 — evidence-gated (needs T1.1 data) - -## T2.1 — Severity thresholding - -Unresolved design problem to answer first: **severity is self-reported by the finding's author.** Dimension-binding does not fix it — the reviewer picks the dimension too, so a correctness bug labeled `hygiene` is capped automatically. And severity is a property of finding × context, not of the finding: "unnecessary try/catch" is cosmetic in a script and a swallowed auth error in a permission path, which the reviewer's own security checklist says to flag. - -Candidates, not yet chosen: derive severity from dimension plus the file's risk class (reuse `matchesSecuritySensitiveGlob`) rather than trusting the label; or have a second cheap pass classify severity independently of the finder. - -Needs a kill switch on the `createBase2` option + `OPENBUFF_*` env pattern (`base2.ts:100-123`), surfaced in `/context`. Trap: for a default-ON flag resolve with `??` on an explicit boolean — `envFlag || DEFAULT` is the documented dead canary at `:87-94`. - -**Thresholding may not touch the runaway loop at all** — that loop is driven by findings new each round; if those are medium-or-above, a threshold changes nothing. Rev 1 wrongly presented this as the top fix for both symptoms. - -## T2.2 — Scope re-review to what changed - -After a repair, pass the full pending set for _attestation_ but direct deep review only at the repair receipt's `changedFiles`, citing the prior verdict for the rest. Attacks the "runs for a while" cost. Keep `collectReviewerAttestationIssues` unchanged so coverage gaps still fail closed. Promote if T1.1 shows reviewer wall-clock dominates. - -## T2.3 — Requirement ledger through the editor handoff - -Carry verbatim acceptance criteria in the editor handoff `Requirements` field, have the editor self-score each in its receipt, and pass that to the reviewer as _claimed_ coverage. Reviewer contradicting a claim is a real finding; silence is not. - -## T2.4 — Nit-ratchet on the no-progress guard - -Requires T1.5's id ledger plus T1.1 data. Revisit only if T1.2 + T1.6 prove insufficient. - ---- - -# Tier 3 — after re-measurement - -- **T3.1** Normalize/dedupe finding text: `dedupeExactStringsPreserveOrder` is exact-match; reuse T1.5's key. (Rev 2 conditioned this on the evidence-gated T2.1, which was incoherent — architect finding 11. It depends on T1.5, which is Tier 1.) -- **T3.2** Docs: `docs/agents-and-tools.md:587` (`LOOKS_GOOD`-only paragraph, budget table) and the repair-loop table in `agents/guides/editor-writers-and-repair.md`. -- **T3.3 / T3.4** Advisory surface: render advisories distinctly in `cli/src/components/renderers/gate-state-box.tsx`; require them in the completion summary and offer a "fix the N nits" followup so the channel is not a silent dumping ground. **Partially pulled forward** — T1.2(a) cannot ship without the minimal version of these. -- **T3.5** Tests for whatever lands. - ---- - -# Sequencing - -1. **T0.1, T0.2** — live authority holes; independent of everything else -2. **T1.1** telemetry + shadow mode (durability resolved per above) -3. **T1.4** guidelines — fully independent, can run parallel with 1–2 -4. **T1.6** cycle detection — small, self-contained, no persisted state -5. **T1.3** optional id/severity + code-reviewer id correlation, with the `[id] ` migration handled -6. **T1.5** id-keyed condoning — unblocks T1.2(c) -7. **T1.2** generator fix: (b) any time after 1; (a) after T1.3 + minimal advisory surface; (c) after T1.5 -8. **Re-measure.** Real decision point: if `newFindingCount` collapses, Tier 2 may be unnecessary -9. Tier 2 individually gated on that data; Tier 3 last - -## Falsification criteria - -- **Escaped-defect rate** — a finding raised in a later turn on a file whose earlier-turn finding was suppressed or downgraded. Rising ⇒ suppression is too loose. Measurable only because T1.1 records suppression decisions. -- **Blocked-turn rate** — share of turns ending in `currentPhase = 'blocked'` (protocol failure, incomplete receipt, no-progress, cycle). Any change that raises this is a regression even if round counts fall. -- **Rejected condone claims** (T0.2) — a nonzero rate is direct evidence the repair-editor was over-claiming, and retroactively justifies T0.1/T0.2. - -## Implementation traps - -- The condone/merge logic (`base2.ts:4084-4148`, `:4521-4540`, `:5187-5229`) lives inside the serialized `handleSteps` generator and **cannot import** from `gate-reviewer.ts` — reconstructed functions lose their module closure. Any id-keying change must be duplicated by hand or moved into the `` region; `scripts/generate-gate-helpers.ts` is the source of truth and `gate-*-parity.test.ts` enforces it. -- New `gate-state.ts` fields need a `??=` default at `base2.ts:889-917`, must stay plain JSON (three explicit "never a Set" / "never a Map" comments), and must handle sessions updating mid-loop. -- `createReviewer` is referenced from `agents/__tests__/code-reviewer.test.ts` and `agents/__tests__/base2-writer-spawn-rules.test.ts`; `docs/agents-and-tools.md` documents a bare-string text-mode contract that persisted sessions and third-party reviewers still emit. Any schema change needs dual-shape fixtures. - -## Missing runtime evidence to collect before Tier 2 - -1. Per-round finding text/id sets from real turns — how often is a re-raise verbatim vs reworded? -2. How often the condoned-pass branch (`base2.ts:4121`) fires, and whether any such pass carried `coverage: "missing"` or an in-scope `uncertain` requirement. That single number decides whether T0.1 is an authority hole in practice or only in theory. - -## Validation per slice - -`agents` typecheck plus `agents/__tests__/gate-reviewer*.test.ts`, `agents/__tests__/base2*.test.ts`, `agents/__tests__/quality-prompt-snapshot.test.ts`. T1.4 additionally `agents/__tests__/base2-progressive-disclosure.test.ts`. T0.1, T0.2, T1.2, T1.3, T1.5, T1.6 additionally `agents/e2e/gate-lifecycle.e2e.test.ts`. diff --git a/.agents/sessions/review-gate-correctness/STATUS.md b/.agents/sessions/review-gate-correctness/STATUS.md deleted file mode 100644 index 141f79e1c2..0000000000 --- a/.agents/sessions/review-gate-correctness/STATUS.md +++ /dev/null @@ -1,66 +0,0 @@ -# Status — review-gate correctness & convergence - -Companion to `PLAN.md` (rev 3), which is a design document with no progress tracking. This file is the progress record. All citations were verified by reading the files at snapshot `v3:4a19a075615be` (uncommitted, branch `feat/task-memory-evidence-pipeline`); line numbers drift as the files change, so treat them as locators rather than addresses. - -## Current state - -Tier 0 closed. Tier 1 **closed**: T1.2(a) and T1.4d both shipped in this session, leaving only T1.2(c). Tier 2/3 not started **by design** — `PLAN.md` sequencing step 8 gates them on re-measuring `newFindingCount` with the T1.1 telemetry, which now has a durable sink. - -Nothing is committed. The last gate pass was reviewer `LOOKS_GOOD` with zero findings; validation was typecheck clean across 11 packages, agents unit suite 1047 pass / 0 fail, e2e 54/54, agent-runtime + common suites 1808 pass / 0 fail across 77 files, generated gate-helpers region fresh. - -## Landed (verified in source) - -| Item | Evidence | -| -------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **T0.1** condone path is no longer verdict authority | `base2.ts` `receiptHasHardRule` re-asserts the coverage/requirement hard rules before crediting a condoned pass | -| **T0.2** condone credit requires evidence | `base2.ts` `condoneEvidenceIsSufficient` requires `status === 'completed'` **and** non-empty `changedFiles`; rejections emit `condoneClaimsRejected` + `condoneRejectReason` | -| **T1.1** per-round telemetry + shadow mode | `round-findings` event carries `findingCount`, `rawFindingCount`, `newFindingCount`, `carriedFindingCount`, `suppressibleFindingCount`, `escalatedFindingCount`, `wouldPassAtThisRound`. Derived from already-persisted `openReviewerFindings`, so it survives serialization (the durability requirement in T1.1) | -| **T1.2(b)** generator bounded | `code-reviewer.ts` instructionsPrompt: "that REQUIRES A CHANGE, in a single pass, and then stop", plus the three named convergence conditions (satisfiable empty set / monotonicity under repair / low churn sensitivity) | -| **T1.3** optional finding metadata + id correlation | `code-reviewer.ts` `findings.items.anyOf` = [string, object with `required: ['text']`]; `base2.ts` `correlateReviewerFindingRecord` matches `[id]` → exact stripped text → single-unambiguous-substring | -| **T1.4a** rubric reaches the model as a guide | `agents/guides/pre-review-self-check.md` exists and the pointer targets it (previously pointed at a `.ts` module) | -| **T1.4b** rubric mirrors what blocks | `quality-prompt-section.ts` `preReviewSelfCheckSection`: `Test coverage (blocking)`, `Requirement coverage (blocking)` incl. "`uncertain` blocks exactly like `missing`", `Advisory vs blocking` | -| **T1.4c** drift guard | `agents/__tests__/review-rubric-parity.test.ts` — `REVIEWER_SCHEMA_RULES` schema↔rubric table with exact enum comparison, plus a guide-drift sweep | -| **T1.5** condoning re-keyed on (verdict class, identity) | `gate-state.ts` `condonedFindingKeys`; `base2.ts` `condonedKeyMatches` (same-class, plus one-directional BLOCKING→NON_BLOCKING de-escalation), `legacyCondonedTextMatches` (single owner of the pre-T1.5 fallback), `boundCondonedEntries` (200-entry cap) | -| **T1.6** fingerprint cycle detection | `base2.ts` `reviewer-repair-cycle` and `specialist-repair-cycle` skip reasons, turn-scoped fingerprint set, no persisted state | -| **Telemetry sink** (resume step 1) | `common/src/util/gate-telemetry.ts` + `packages/agent-runtime/src/orchestration/gate-telemetry-sink.ts` write the `base2.gate` events to a durable JSONL sink under `.openbuff/` (gitignored), so the Tier 2 gate is now mechanically answerable | -| **Advisory channel** (resume step 2) | `code-reviewer.ts` declares an OPTIONAL additive `advisories` output field; `gate-reviewer.ts` `collectReviewerAdvisories` reads the last `schemaVersion`-shaped entry; `recordSuccessfulReviewReceipt` persists `advisories` + `advisoryCount`; `cli/src/types/chat.ts` / `message-block-helpers.ts` / `gate-state-box.tsx` render them from a real `advisories` field rather than smuggling them into `details` | -| **T1.2(a)** LOOKS_GOOD contradiction resolved | `code-reviewer.ts`: `LOOKS_GOOD` when nothing REQUIRES A CHANGE even with cosmetic observations, which go to `advisories` with `findings` empty. Unblocked by the advisory channel above | -| **T1.4d** embedder guide fallback | `common/src/util/guides.ts` (`FALLBACK_GUIDES`, `GUIDE_FALLBACK_SECTIONS`, `BROAD_AUDIT_FALLBACK_SECTIONS`, `findMissingGuides`, `formatGuideFallbackSection`) + one `ON_DEMAND_GUIDE_FALLBACK_` placeholder per relocated guide in `packages/agent-runtime/src/templates/{types,strings}.ts`, appended additively after base2's pointers. Recovery is per pointer and mirrors each mode's exclusions: plan mode omits git-discipline's recovery exactly as it omits the pointer, and takes the plan-clause broad-audit body. Recovered blocks are recorded in the shared context-budget ledger. See LESSONS.md for why this shape and not the two `PLAN.md` framings | - -Also landed but **not tracked in PLAN.md**: delta-only pinned active-work state ("Win 4a"). `base2.ts` `lastEmittedPinnedStateMessage` is the emitted-block baseline, deliberately distinct from `lastPinnedStateMessage`, which `markActiveWorkStateChanged` resets to `''` on every gate-state write. - -## Open - -### T1.2(c) — round ledger — UNBLOCKED, the only remaining Tier 1 item - -T1.5 landed, so the withhold reason (exact-string condoning turning "re-raise it and say why" into a filter escape) is resolved. - -Carry-forward caveat, not in `PLAN.md`: T1.5's id-keying only bites when the reviewer supplies a stable id. Minted `RF--` ids are deliberately excluded from id-keying because they embed list position. For bare-string findings condoning still falls back to `(class, text)`, so **the ledger must instruct verbatim re-raise text regardless of ids** — the fallback the plan named, not only the id path. - -### Advisory display asymmetry — resolved, with one deliberate non-goal - -The first advisory-channel slice displayed advisories only on the gate-pass `` block, so an intermediate `NON_BLOCKING` receipt and every security/specialist receipt persisted them invisibly. A follow-up slice added an `Advisories (non-blocking; no change required):` block to the reviewer-blocker/repair path (read through `collectReviewerAdvisories(reviewerToolResult)`, because `recordSuccessfulReviewReceipt` has not run yet at that point) and to the security pass, specialist parent-owned-only pass, and specialist normal pass. All three aux surfaces emit ONLY when the bounded list is non-empty — an unconditional yield would shift the generator step sequence that `base2.test.ts` and `gate-lifecycle.e2e.test.ts` advance yield-by-yield. - -Still deliberately absent: no advisory surface on the crash / no-verdict / attestation-failure paths. Those have no trustworthy receipt to read. - -## Not started by design - -Tier 2 (T2.1 severity thresholding, T2.2 scoped re-review, T2.3 requirement ledger through the editor handoff, T2.4 nit-ratchet) is gated on T1.1 data from real turns. Tier 3 follows that. This is the plan's own decision point: if `newFindingCount` collapses after T1.2 + T1.6, Tier 2 may be unnecessary. - -The former blocker is cleared: the telemetry sink now persists the `base2.gate` events, so "after N real turns you can state the nit-driven share of rounds" is mechanically answerable. What remains is accumulating those real turns — the decision is data-gated, not implementation-gated. - -## Resume instructions - -One gate cycle per slice (each edit re-runs validation + reviewer). Steps 1, 2, 3, and 5 of the original sequence are done: - -1. ~~Telemetry sink~~ — done; persists to a gitignored `.openbuff/` JSONL sink. -2. ~~Advisory channel~~ — done; schema + `reviewReceipts` + CLI render, plus the blocker/repair and aux-pass surfaces. -3. ~~T1.2(a) prompt fix~~ — done, after 2. -4. **T1.2(c) round ledger** — the only remaining Tier 1 item, with the verbatim-re-raise instruction above. -5. ~~T1.4d~~ — done via the architect's hybrid recommendation (compact clause retained + additive runtime placeholder). - -After 4, the next decision is Tier 2, and it is data-gated: read the accumulated sink events and re-measure `newFindingCount` before implementing any of T2.1–T2.4. - -## Validation per slice - -`agents` typecheck plus `agents/__tests__/gate-reviewer*.test.ts`, `agents/__tests__/base2*.test.ts`, `agents/__tests__/quality-prompt-snapshot.test.ts`. T1.4-family slices add `agents/__tests__/base2-progressive-disclosure.test.ts`, `common/src/util/__tests__/guides.test.ts`, and `packages/agent-runtime/src/templates/__tests__/strings.test.ts`. Behavioral gate slices add `agents/e2e/gate-lifecycle.e2e.test.ts`. CLI-facing gate-state slices add `cli/src/utils/__tests__/message-block-helpers.test.ts` and `cli/src/components/__tests__/gate-state-box.test.tsx`. Any `base2.ts` change also needs `bun run scripts/generate-gate-helpers.ts --check agents/base2/base2.ts`. diff --git a/.agents/sessions/reviewer-coupling-followups/DESIGN-PROPOSAL.md b/.agents/sessions/reviewer-coupling-followups/DESIGN-PROPOSAL.md deleted file mode 100644 index e7a03d1691..0000000000 --- a/.agents/sessions/reviewer-coupling-followups/DESIGN-PROPOSAL.md +++ /dev/null @@ -1,294 +0,0 @@ -# Reviewer subsystem — design proposals - -Status: Proposal A IMPLEMENTED (via `scripts/generate-gate-helpers.ts` + -freshness test); Proposal B IMPLEMENTED (merge ledger via -`mergeReviewerFindings` on security + final code-reviewer blocking paths, -keeping the existing arrays rather than a second keyed field); Proposal C -already IMPLEMENTED (aux-ownership variant — see below). Three safe fixes were -shipped earlier (see "Already shipped" below). - -Source of truth at time of writing: `agents/base2/base2.ts` (the serialized -`createBase2` `handleSteps` generator), `agents/base2/gate-state.ts`, -`agents/base2/gate-reviewer.ts`, `agents/base2/gate-paths.ts`, -`agents/base2/gate-repair.ts`, and their parity tests under -`agents/__tests__/`. - -## Already shipped (context) - -- **Removed dead `staticReviewOnly` scaffolding.** The flag, `staticReviewerJobId`, - the guarded `check_background_agent { cancel: true }` blocks, the - `= undefined` resets, and the field-only round-trip tests were deleted. The - reviewer was never actually spawned with `background: true`, so no reviewer - ever ran concurrently with validation; enabling the flag did nothing. This - was pure cleanup. -- **Recorded specialist reviewer provenance.** `openReviewerFindings[].reviewer` - was widened to `'code-reviewer' | 'security-reviewer' | SpecialistReviewerAgent` - and the specialist blocking path now stamps `reviewer: agentType`. This - provenance is now what drives Proposal C's aux-ownership routing. -- **Implemented Proposal C (aux-ownership variant).** `requiredReviewerRevalidation` - was widened to hold a `SpecialistReviewerAgent`, and an inline - `revalidationFamily` classifier routes each family's revalidation back to the - aux block that already owns its correct attestation contract: security -> - security aux block (params `changed_files`+`snapshot_fingerprint`), specialist - -> specialist aux block (fresh `get_change_review_bundle` `snapshot_id` + - one-refresh retry), code -> final block. Fire-guards are family-scoped and - marker clears never clobber another family's marker. This also FIXED a latent - bug: security-reviewer revalidation previously flowed into the paramless final - spawn, which could never satisfy security-reviewer's required params. See the - aux-ownership note appended to Proposal C below. - ---- - -## Proposal A — Reduce `gate-reviewer.ts` / inline parity duplication - -### Problem - -The reviewer-gate helpers exist twice: - -1. As clean, testable module exports in `agents/base2/gate-reviewer.ts` - (`collectReviewerBlockers`, `getReviewerFinalizationVerdict`, - `detectReviewerCrash`, `collectStructuredReviewerOutputs`, - `collectReviewerAttestationIssues`, `stripReviewerPreamble`, - `isTestCoverageReviewerFinding`, ...). -2. As inline copies inside the `createBase2` `handleSteps` generator body - (`collectReviewerFindingRecordsInline`, `selectSpecialistReviewersInline`, - and the inline verdict/blocker/attestation helpers). - -Root cause (documented in `gate-reviewer.ts`'s header comment): `handleSteps` -is serialized with `handleSteps.toString()` and reconstructed via -`new Function(...)`. Reconstructed functions lose their module closure, so the -generator cannot reference imports from `gate-reviewer.ts` at runtime. -Everything the generator calls must be defined inline in the body. - -Current mitigation: parity tests (`gate-reviewer-parity.test.ts`, -`gate-paths-parity.test.ts`, `gate-repair-parity.test.ts`) use -`extractInlineFunctionSource` / `loadInline*Helpers` to pull each inline copy -out of the serialized generator and assert byte/behavior equivalence with the -exported version. Drift is _detected_, not _prevented_, and the extraction is -fragile (it string-parses the generator body). - -### Options - -- **A1 — Build-step inlining (single source of truth).** Keep only the - `gate-reviewer.ts` module. Add a prebuild step that mechanically injects the - helper sources into the generator body (or emits a generated file that the - generator references) before `handleSteps.toString()` serialization runs. - The parity tests are replaced by a "generated block is fresh" check - (same pattern as `cli/src/agents/bundled-agents.generated.ts` freshness). - - Pros: one source of truth; drift becomes impossible rather than detected. - - Cons: adds generator complexity; must run before the existing - `cli/scripts/prebuild-agents.ts` bundling; another generated artifact to - keep fresh in CI. -- **A2 — Serializable helper injection.** Pass the helpers into the generator - through a mechanism that survives `new Function` (e.g. stringify named helper - sources into a preamble the generator evals once). Fights the execution - model; high risk. Not recommended. -- **A3 — Keep parity tests, harden extraction.** Lowest effort: leave the two - copies, but make `extractInlineFunctionSource` more robust and add a lint that - fails when a `gate-reviewer.ts` export lacks a matching parity test. - - Pros: minimal, no execution-model risk. - - Cons: still two sources of truth; still a maintenance tax. - -### Recommendation - -A1 if we are willing to invest in the prebuild step (it is the only option that -actually eliminates the duplication the user flagged). A3 as the low-risk -fallback if we want to keep the change small. A2 is not recommended. - -### Risk / blast radius - -High. `base2.ts` is the gate. Any inlining bug changes review verdict parsing -for every edit turn. Must land behind the full gate + parity/`base2.test.ts` / -e2e suites, and the generated bundle must be regenerated. - ---- - -## Proposal B — Merged multi-reviewer finding ledger (single-slot -> keyed) - -### Problem - -`activeWorkState.openReviewerBlockers` / `openReviewerFindings` are a single -slot. Whichever tier blocks last wins: the security gate sets them, a specialist -gate can overwrite them, and the final code-reviewer overwrites them again. -Today this is safe ONLY because the tiers are strictly sequential and each -blocking tier `continue`s or `break`s the loop before the next tier runs — so -exactly one tier's findings are ever "live" at once. There is no place that -merges findings from multiple reviewers. - -### Why it is worth changing - -- If tiers ever run concurrently (e.g. a future real static-review path), the - single slot becomes a lost-update race. -- Operators only ever see one tier's findings at a time; a change that is both a - security risk and a reliability risk surfaces only the last-writing tier. -- Repair routing has to reconstruct provenance (partially addressed by the - shipped provenance metadata). - -### Proposed shape - -Replace the single slot with a keyed ledger on the gate state: - -```ts -openReviewerFindingsByGate?: Record -``` - -- Derive the flat `openReviewerBlockers` (still consumed by pinning/messaging - and `buildPinnedActiveWorkMessage`) as a computed projection over all `open` - ledger entries, so the user-visible contract is unchanged. -- Each tier writes/updates only its own keyed entry; clearing a finding marks - that entry `resolved` instead of blowing away the whole slot. -- The gate passes only when every ledger entry is `resolved`. - -### Migration / compatibility - -- Backward-compatible: older serialized state lacks the field -> treat as the - legacy single-slot behavior (fail closed). -- `context-pruner.ts` reads `openReviewerBlockers`; keep that field as the - derived projection so the pruner and gate-state block are unaffected. -- Parity: the ledger logic lives inline in the generator, so it needs the same - parity-test treatment as Proposal A (another argument for doing A first). - -### Risk - -Medium-high. Changes the core block/clear bookkeeping. Must preserve exact -finalization semantics (gate passes iff no open findings) and the pinned -`` contract. - ---- - -## Proposal C — Specialist / security revalidation routing (the Q2 behavior change) - -### Why this is NOT a small fix - -The Q2 request was "route revalidation back to that specialist." Verification -shows the final-reviewer block cannot do this as-is: - -- `base2.ts:2307`: `requiredReviewerAgentType = requiredReviewerRevalidation ?? 'code-reviewer'`. -- The final reviewer spawn (`base2.ts:3038`) passes **only a `prompt`** — no - `params`. -- But `security-reviewer` requires `params.changed_files` + `snapshot_fingerprint`, - and specialists require `params.snapshot_id`. -- Specialists attest against the `get_change_review_bundle` `snapshotId`, which - is a DIFFERENT fingerprint than the `reviewSnapshotFingerprint` the final - block builds/validates against. - -So routing a specialist through `requiredReviewerRevalidation` would fail spawn -params AND snapshot attestation. Today specialists effectively "self-revalidate" -by re-running their own gate block on loop re-entry (they `continue` on block -and never set `requiredReviewerRevalidation`). `requiredReviewerRevalidation` is -currently only ever `'security-reviewer'` or `'code-reviewer'`. - -### Options - -- **C1 — Keep current self-revalidation, do nothing to routing.** The shipped - provenance metadata already records which specialist found the issue. - Specialists re-run their own block after repair. Lowest risk; the flagged - `reviewerOriginFromGateId` reconstruction stays but is now backed by real - metadata. -- **C2 — Generalize the revalidation dispatcher.** Make the final block - parameterize the reviewer spawn by agent family: build the correct `params` - (`snapshot_id` for specialists via a fresh `get_change_review_bundle`, - `changed_files`+`snapshot_fingerprint` for security, none for code-reviewer) - and attest against the matching fingerprint per family. Then - `requiredReviewerRevalidation` can legitimately hold a specialist type. - - Pros: unified revalidation path; provenance drives routing. - - Cons: the final block must branch attestation by reviewer family; the - snapshot-fingerprint mismatch between the specialist bundle id and the - reviewable-scope fingerprint has to be reconciled. This is the single most - fragile part of the gate. - -### Recommendation - -C1 now (already effectively in place; the provenance ship makes it coherent). -Pursue C2 only alongside Proposal A/B, because it touches the same inline gate -logic and needs the same parity + e2e coverage. Do not attempt C2 as a -standalone quick edit. - -### Risk - -High (C2). Attestation is the anti-clash core of the whole subsystem; a wrong -fingerprint branch would let a review of stale bytes pass. - -### IMPLEMENTED (aux-ownership, chosen over C2 dispatch) - -Rather than turning the fragile final block into a family-aware dispatcher -(C2), the implemented design keeps each family's attestation in the aux block -that already encodes it correctly: - -- `requiredReviewerRevalidation` is now a persisted "revalidation-owed" family - marker that CAN hold a `SpecialistReviewerAgent`, but it does not drive the - final-block spawn. An inline `revalidationFamily(marker)` classifier maps the - marker to `'none' | 'code' | 'security' | 'specialist'`. -- Fire-guards: the security aux block fires when the family is `none` or - `security`; the specialist aux block fires when the family is `none` or - `specialist` (re-including the owed specialist and dropping it from - `specialistReviewGatesDone` so its spawn+attestation re-runs); the final - block spawns a reviewer only when the family is `none` or `code`. -- Each owner block clears the marker only when it owns that family, so one - block never clobbers another family's marker. The initial marker inference - now prefers `openReviewerFindings[0].reviewer` (the shipped provenance) with - a gateId-prefix fallback, which is how a specialist blocking finding from one - turn re-fires the specialist aux block on the next. -- This fixed the latent security-reviewer revalidation bug (paramless final - spawn) as a side effect, since security now always revalidates through its - params-bearing aux block. - -Note: on a blocking specialist finding the specialist aux block parks -`blocked` with the provenance-stamped finding and `continue`s (it does not -spawn repair-editor inline the way the security block does); the marker is -rehydrated to the specialist family at the next turn's setup from that -persisted finding, which is what drives the aux re-fire. - -Remaining open: C2's full unified dispatcher is intentionally NOT implemented; -aux-ownership was chosen for lower blast radius on the attestation core. - ---- - -## Suggested sequencing - -1. Proposal A (single source of truth) first — it unblocks safe iteration on the - inline gate logic that B and C both need. -2. Proposal B (keyed ledger) second — depends on A's parity approach. -3. Proposal C2 (revalidation dispatcher) last — highest risk, reuses A+B - infrastructure. C1 is the no-op-now default. - -Each must land behind: `cd agents && bun run typecheck`, `bun test base2.test.ts`, -the `gate-reviewer` / `gate-*-parity` suites, the `agents/e2e` gate lifecycle + -reviewer-spawn-conditions e2e tests, and a regenerated -`cli/src/agents/bundled-agents.generated.ts`. - ---- - -## Lessons learned during implementation (2026-08-21) - -- **TDZ hazard when hoisting consts inside the serialized generator.** Moving - `reliabilityCodeStems` / `reliabilityCodeExtension` "above" - `selectSpecialistReviewersInline` must mean above it AND before every runtime - call site: `const` bindings initialize only when execution reaches them, and - `handleSteps` executes top-to-bottom from its opening (~line 554) while the - first router call site sits at ~line 2199. Placing the hoisted block next to - the function's declaration site (~line 6667) crashed every gate e2e at - runtime with `Cannot access 'reliabilityCodeExtension' before -initialization` even though typecheck and the parity suite stayed green — - the parity harness rebuilds scope instead of executing the generator, and - the file-change hooks only typecheck. Rule: any hoisted binding inside - `handleSteps` goes at the very top of the generator body, keeping the consts - contiguous with the function so the parity slice keeps working. Only - executing the full test suite catches this class of bug. -- **Router vocabulary widening couples to test fixtures.** Fixtures chosen - under old routing rules silently change meaning when vocabulary widens: - `cli/src/auth/session.ts` began routing reliability-reviewer once exact - filename stems (`session`) were added, breaking three aux-ordering e2e tests. - When widening the router, grep e2e fixtures for newly-matching paths and - prefer fixtures whose basename cannot match any family (e.g. - `token-store.ts`). -- **Full-suite validation belongs before push, not after.** Both regressions - above were invisible to focused suites and typecheck; only - `cd agents && bun test` over all 51 files surfaced them. `check:ci-local` - now includes the full agents suite as Step E so the pre-push hook and local - CI mirror catch TDZ-class regressions before they reach origin/main. diff --git a/.agents/sessions/reviewer-gate-concurrency-fix3-2026-07/SPEC.md b/.agents/sessions/reviewer-gate-concurrency-fix3-2026-07/SPEC.md deleted file mode 100644 index 697ed0f794..0000000000 --- a/.agents/sessions/reviewer-gate-concurrency-fix3-2026-07/SPEC.md +++ /dev/null @@ -1,128 +0,0 @@ -# Fix 3 (deferred): Concurrent-instance isolation for the base2 validation/reviewer gate - -Status: IMPLEMENTED — absorption uses task-related (`changedFiles`) + -runtime-published `selfMutatedPaths` only. The agent runtime -(`packages/agent-runtime/src/run-agent-step.ts` `publishSelfMutatedPaths`) -records confirmed broker/tool mutation paths, agent-receipt changedFiles, and -terminal/basher `touchedPaths` (pre/post `git status --porcelain -uall` dirty -delta) onto `agentState.selfMutatedPaths` as a JSON-safe `string[]` after each -stream step so concurrent isolation can credit process-owned writes — including -formatter/codegen side effects outside the mutation broker. SYNC commands emit -`touchedPaths` on the command result; BACKGROUND jobs capture a pre-start dirty -snapshot and emit a one-shot settlement delta on the first settled `check_job` -observation (not on start, not on re-polls; soft-fail omits when not a git repo -or recovered without snapshot). -Pure helper: `agents/base2/gate-concurrency.ts` `shouldAbsorbGitStatusFile`. -The handleSteps inline is **generator-synced** (not hand-maintained): `scripts/generate-gate-helpers.ts` emits it into the `` region of `agents/base2/base2.ts` (same as gate-paths/reviewer/repair). Edit the pure module and regenerate (`--write` / `prebuild:agents`); freshness is enforced by `agents/__tests__/gate-helpers-freshness.test.ts` and the concurrency parity matrix. - -## Problem - -When multiple Openbuff instances (or the user + an instance) share one worktree, -instance B's base2 gate can absorb instance A's in-flight edits into its own -`pendingGateFiles` and try to validate/review files it never touched. Symptoms: -spurious `awaiting_validation`, reviewer spawns over unrelated files, and -attestation churn. - -### What already works (do NOT re-fix) - -- **Pre-existing dirty files (Issue 3a) are already isolated.** The turn-start - snapshot `initialGitStatusFiles` (`agents/base2/base2.ts` ~705-712) is - subtracted from the mid-turn git-status sweep (~938: - `!initialGitStatusFiles.includes(file)`), and the fresh-turn pending - population (~606-617) only seeds from `changedFiles` (empty at turn start). - A file already dirty when the turn begins never enters this turn's pending set. -- **The commit guard is already scoped** to `taskRelatedFiles` - (`touchedFiles`/`changedFiles`/`pendingGateFiles`/`gatePassedFiles`) via - `uncommittedUnvalidatedFiles` (~804-813), so unrelated dirty files do not - block commits. - -### The remaining gap (Issue 3b) - -The mid-turn git-status sweep (~930-947, `recordChangedFiles([file], { -fromStatusObservation: true })`) pulls in ANY newly-dirty path that appeared -_after_ turn start, regardless of whether THIS agent authored it. If instance A -writes `foo.ts` during instance B's turn, B's post-step `git_status` reports -`foo.ts` as newly dirty (not in B's `initialGitStatusFiles`), so B absorbs it. - -The sweep is load-bearing, not a backstop: in the test harness and for -`{ file }`-shaped step results, files enter `pendingGateFiles` ONLY via this -sweep (~15 tests in `agents/__tests__/base2.test.ts` depend on it). So it cannot -simply be scoped to `taskRelatedFiles` without dropping legitimately -self-authored changes that only `git_status` (not `extractChangedFiles` / -`extractChangedFilesFromMessages`) observed — e.g. a formatter/codegen write -triggered by a `basher`/`run_terminal_command` step. - -## Why the naive fixes fail - -- **Scope sweep to `taskRelatedFiles`:** drops self-authored files that only - git saw (terminal/codegen writes) and breaks the ~15 harness tests that seed - pending exclusively through the sweep. -- **Thinker's `stepResult.toolName` gate:** wrong for this architecture. A STEP - is a full multi-tool model step; there is no single `toolName` on the step - result. Tool calls are scanned from the message-history delta - (`extractChangedFilesFromMessages`, ~5257), not a scalar tool id. - -## Proposed design: per-process ownership token from the mutation broker - -The SDK already journals every mutation per originating process through the -worktree-scoped cooperative mutation broker -(`sdk/src/services/workspace-mutation-broker.ts`). That journal is the correct -authority for "did THIS process write this file," rather than inferring -authorship from repo-global `git status`. - -### Sketch - -1. **Expose an owned-path set from the broker to the runtime.** The broker - already records exact-byte conditional commits/creates/deletes per process. - Add a read-only accessor that returns the set of project-relative paths this - process's broker has mutated during the session (or since a passed cursor). - Surface it on the agent runtime state the base2 generator can read, e.g. - `mutableAgentState.selfMutatedPaths` (a `Set` / string[]). - -2. **Gate the sweep absorption on self-authorship OR self-mutation evidence.** - In the mid-turn git-status sweep, keep the existing - `!initialGitStatusFiles.includes(file) && !gatePassedFiles.has(file)` - predicate and add: absorb the file only if it is already task-related - (`taskRelatedFiles.has(file)`) OR present in `selfMutatedPaths`. A file - dirtied during the turn that is neither task-related nor self-mutated is - attributed to a concurrent instance and excluded from `pendingGateFiles` - (but STILL recorded into `gitStatusObservedFiles` / `gitStatusObservedDirty` - so the existing committed-file pruning telemetry keeps working — narrow only - the absorption branch, not the observation bookkeeping). - -3. **Test-harness compatibility.** The ~15 base2 tests feed `{ file: 'src/a.ts' }` - step results — those already route through `extractChangedFiles` → - `recordChangedFiles`, making the file task-related BEFORE the sweep, so they - remain unaffected. Add a test double for `selfMutatedPaths` so the - terminal/codegen self-authored case is covered explicitly. - -### Residual, bounded ambiguity (state honestly) - -If instance B runs a mutating terminal command in the SAME step that instance A -writes a file, and B's broker did not record A's file (it wouldn't — different -process), B correctly EXCLUDES A's file. The only unavoidable ambiguity is a -tool/codegen path the broker does not observe (a raw child process writing -outside the broker); those remain attributed via the `taskRelatedFiles` fallback -only, i.e. excluded unless independently self-authored. This errs toward NOT -absorbing another instance's file — the correct safety bias for isolation, and a -strict improvement over today. - -## Scope / risk - -- Cross-package: `sdk/src/services/workspace-mutation-broker.ts` (accessor) + - runtime state plumbing + `agents/base2/base2.ts` (sweep predicate). Medium - size; touches the fragile serialized `handleSteps` generator, so it needs its - own e2e coverage (a two-instance simulation feeding disjoint - `selfMutatedPaths` + overlapping `git_status`). -- Must preserve every existing gate e2e invariant - (`agents/e2e/gate-*.e2e.test.ts`, `agents/__tests__/base2.test.ts`). - -## Open questions for the user - -1. Is per-session self-mutation tracking sufficient, or do we want per-turn - cursoring (reset the owned-path set at turn boundaries)? -2. Should a file dirtied by an unobserved raw child process (outside the broker) - fail open (absorb + validate — safer for correctness, worse for isolation) - or fail closed (exclude — better isolation, risk of an unvalidated self-edit)? - The commit guard already fails closed independently, which argues for - fail-closed here too. diff --git a/.agents/sessions/shell-policy-audit-2026-07/LESSONS.md b/.agents/sessions/shell-policy-audit-2026-07/LESSONS.md deleted file mode 100644 index 6b47501fb5..0000000000 --- a/.agents/sessions/shell-policy-audit-2026-07/LESSONS.md +++ /dev/null @@ -1,36 +0,0 @@ -# LESSONS — Shell/Terminal Security Restriction Audit - -## Key insight - -The git-committer's unreliability is not a bug in git-committer — it's a -contradiction: `gitCommitGuidePrompt` tells the model to use -`git commit -m "$(cat <<'EOF' ... EOF)"`, but the `git-commit` terminal profile's -`hasActiveShellSyntaxAnywhere` rejects `$(`, backticks, `<`, `>` anywhere (and -raw newlines). Guidance and enforcement must be co-designed or they fight. - -## Enforcement layering - -Three independent layers touch a git command: - -1. `evaluateTerminalCommandPolicy` (profile allowlist), -2. `classifyTerminalHarnessAction` + `evaluateHarnessActionPolicy` (approval), -3. `validateStagedCommit` (staged-diff safety re-scan). - Redundant for an already-locked profile; a single authoritative owner per - concern would reduce friction. - -## Distinguish friction from security value - -- Real value (keep): traversal/outside-path containment, force/delete-push - block, privilege escalation, env dump, sensitive-file staged scan, tmux - write-through-shell block. -- Mostly friction (relax): quote-blind raw-syntax ban on git-commit, git - read-only composition limited to git-only segments, /git slash-command - bracket/brace ban. - -## Gotcha - -The raw-syntax guard is intentionally NOT quote-aware because `bash -c` expands -substitution/redirection even inside quotes. Any relaxation (Approach A) must -use a bounded, quoted-delimiter heredoc parse (like the existing -`stripBoundedDiagnosticHeredoc`) so body text stays inert — do not just make the -guard quote-aware. diff --git a/.agents/sessions/shell-policy-audit-2026-07/PLAN.md b/.agents/sessions/shell-policy-audit-2026-07/PLAN.md deleted file mode 100644 index 63afc1abcd..0000000000 --- a/.agents/sessions/shell-policy-audit-2026-07/PLAN.md +++ /dev/null @@ -1,79 +0,0 @@ -# PLAN — Shell/Terminal Security Restriction Remediation - -Tiered so you can approve only what you want. Tier 1 fixes the git-committer -pain directly; higher tiers are optional cleanups. Every change is -security-sensitive → advisory `security-reviewer` before edit + full -validation/reviewer gate after. - -## Tier 1 — Fix git-committer reliability (recommended) - -- [ ] T1.1 Allow multi-line commit messages under the git-commit profile - - Depends on: none - - Approach A (minimal): extend the existing bounded-heredoc allowance (already - used by `validation-diagnosis` via `stripBoundedDiagnosticHeredoc`) to accept - the `git commit -m "$(cat <<'EOF' ... EOF)"` shape for git-commit, treating - the quoted-delimiter body as inert data. Keep the raw-syntax guard for every - other git-commit command. - - Approach B (cleaner, larger): give git-committer a structured commit path so - the message travels as a tool parameter and never through the shell (no `$(`, - no heredoc). Removes the contradiction entirely. - - Acceptance: git-committer produces a 2+ line commit message end-to-end. - - Validate: `bun test sdk/src/__tests__/terminal-command-policy.test.ts` + - `agents/__tests__/git-committer.test.ts`. - -- [ ] T1.2 Reconcile `gitCommitGuidePrompt` with the policy - - Depends on: T1.1 - - If Approach A: keep the heredoc guidance (now permitted). If Approach B: - rewrite the guide to describe the structured path and drop the shell HEREDOC. - - Acceptance: guidance no longer instructs a form the active profile blocks. - - Validate: `agents/__tests__/quality-prompt-snapshot.test.ts`. - -- [ ] T1.3 Widen read-only git composition to allow safe pagers/filters - - Depends on: none - - In `isReadOnlyGitCommand` composition path, permit non-git segments drawn - from a tight allowlist (`head`, `tail`, `cat`, `wc`, `nl`, `grep`, `rg`, - `sort`, `uniq`) so `git log | head` works; keep the no-substitution / no- - redirection guards. - - Acceptance: `git log --oneline | head -20` allowed; `git log | sh` still denied. - - Validate: `bun test sdk/src/__tests__/terminal-command-policy.test.ts`. - -## Tier 2 — Reduce redundant enforcement layers (optional) - -- [ ] T2.1 De-duplicate git-action gating for the git-commit profile - - Depends on: T1.\* - - Decide the git-commit profile is authoritative; avoid double-gating `commit` - through the harness approval layer for that profile (push/force still gated). - Document which layer owns which concern in `docs/agents-and-tools.md`. - - Acceptance: a normal git-committer commit is evaluated by one authoritative - layer; no behavior change for push/force/default-branch. - - Validate: `bun test sdk/src/__tests__/harness-enforcement.test.ts` + - `sdk/src/__tests__/run-terminal-command.test.ts`. - -- [ ] T2.2 Soften the git add allowlist failure mode (F3) - - Depends on: T2.1 - - Keep the exact-subset rule but improve the denial reason and consider - accepting `./`-prefixed / normalized-equivalent paths without failing. - - Acceptance: equivalent path spellings no longer spuriously rejected. - -## Tier 3 — CLI convenience (optional, low risk) - -- [ ] T3.1 Narrow FORBIDDEN_SHELL_CHARACTERS for /git slash commands (F5) - - Depends on: none - - Stop blocking `[ ] { } ( )` (still block `; $ \` | & < > \\` and newlines) so - git pathspec/brace syntax works; args are already shell-quoted individually. - - Acceptance: `/git diff 'app/{a,b}'` parses; injection cases still throw. - - Validate: `bun test cli/src/commands/__tests__/git-command-args.test.ts`. - -## Risks / blockers - -- These are security controls. Any relaxation must be reviewed by - `security-reviewer` and must not weaken the (C) keep-list in SPEC. -- Approach B (T1.1) is a larger surface (new structured executor path); Approach - A is the minimal fix. -- Open question: which tiers to execute, and A vs B for T1.1. - -## Validation gates - -- Per-task `bun test` on the named suites. -- Full runtime validation + reviewer gate before finalizing any tier. -- security-reviewer advisory pass on `terminal-command-policy.ts` changes. diff --git a/.agents/sessions/shell-policy-audit-2026-07/SPEC.md b/.agents/sessions/shell-policy-audit-2026-07/SPEC.md deleted file mode 100644 index 2afd95f7d1..0000000000 --- a/.agents/sessions/shell-policy-audit-2026-07/SPEC.md +++ /dev/null @@ -1,115 +0,0 @@ -# SPEC — Shell/Terminal Security Restriction Audit - -## Overview - -Audit the terminal-command permission layer for restrictions that are more -punishing than protective — controls that block legitimate agent work (the -git-committer is the motivating example) without adding meaningful security -value. Produce ranked removal/relaxation candidates, clearly separating -low-value friction from genuine security controls that must stay. - -The trigger: `git-committer` (profile `git-commit`) cannot reliably run git. -Root cause confirmed below is a direct contradiction between the guidance the -model receives and what the policy permits. - -## Goals - -- Enumerate every enforcement point in the terminal/shell policy layer. -- Classify each restriction as: (A) safe to relax/remove, (B) consolidate, or - (C) keep (real security value). -- Give concrete, minimal change candidates for the (A)/(B) items. -- Fix the git-committer reliability problem specifically. - -## Non-Goals - -- Removing path-traversal, force-push, privilege-escalation, env-dump, or - sensitive-file protections (these are (C) — real value). -- Rewriting the harness approval system. -- Touching `full-access`/`user` mode behavior. - -## Relevant systems / files - -- `sdk/src/tools/terminal-command-policy.ts` — `evaluateTerminalCommandPolicy`, - per-profile guards (git-commit, dependency-mutation, read-only, - validation-diagnosis, tmux-test, workspace-write). ~1200 lines. Primary surface. -- `sdk/src/tools/run-terminal-command.ts` — calls the policy, then - `classifyTerminalHarnessAction`, then `validateStagedCommit` for git commits. -- `sdk/src/services/harness-enforcement.ts` — `classifyTerminalHarnessAction`, - `evaluateHarnessActionPolicy` (second enforcement layer, approval-based). -- `agents/git-committer/git-committer.ts` — profile `git-commit`; handleSteps - yields git commands; `owned_paths` → `allowed_paths`. -- `common/src/constants/git-discipline.ts` — `gitCommitGuidePrompt` (recommends - HEREDOC commit form that the git-commit profile blocks). -- `common/src/tools/params/tool/run-terminal-command.ts` — injects - `gitCommitGuidePrompt` into the tool description (line ~135). -- `cli/src/commands/git-command-args.ts` — `parseSafeGitArgs` / - `FORBIDDEN_SHELL_CHARACTERS` for the `/git diff` and `/git status` slash commands. -- Profile assignments: `agents/git-committer`, `agents/dependency-manager`, - `agents/debugger`, `agents/librarian`, `agents/browser-use`, `agents/tmux-cli`, - `agents/basher`, `.agents/lib/create-cli-agent.ts`. - -## Findings (evidence-backed) - -### F1 — [HIGH friction] git-commit policy contradicts its own guidance - -- `hasActiveShellSyntaxAnywhere` rejects `$(`, backtick, `<`, `>` ANYWHERE in a - git-commit command (deliberately not quote-aware because `bash -c` expands - inside quotes). Raw newlines are also rejected for non-full-access profiles. -- `gitCommitGuidePrompt` explicitly instructs: `git commit -m "$(cat <<'EOF' ... EOF)"`. -- Net effect: the documented multi-line commit workflow is impossible under - `git-commit`; the agent is limited to single-line `-m "..."`. This is the - "can't reliably manage git" symptom. - -### F2 — [MEDIUM] Triple, overlapping enforcement for git actions - -A single `git commit` from git-committer passes through: - -1. `evaluateTerminalCommandPolicy` (git-commit profile — already fully constrains it), -2. `classifyTerminalHarnessAction` → `commit` action (approval layer; gated in - `strict` mode), -3. `validateStagedCommit` re-scan in run-terminal-command.ts. - The profile is already authoritative for git-committer; layers 2–3 add friction - and cognitive load with little marginal safety for this already-locked profile. - -### F3 — [LOW] git add requires non-empty allowed_paths + exact subset - -`git add` under git-commit fails unless every staged path is an exact, -normalization-matched member of `owned_paths`. Legitimate but a real source of -"it refused my command" surprises when owned_paths is omitted or mismatched. - -### F4 — [MEDIUM friction] Read-only git composition allowlist is too narrow - -Under git-commit, shell composition (`|`/`&&`/`;`) requires EVERY segment to be -an allowlisted read-only _git_ command (`splitReadOnlyShellSegments().every(isReadOnlyGitCommand)`). -So ordinary inspection like `git log --oneline | head -20` is rejected because -`head` isn't a git command. Over-restrictive for an agent whose whole job is -inspecting git state. - -### F5 — [LOW] CLI /git slash-command arg parser blocks pathspec syntax - -`FORBIDDEN_SHELL_CHARACTERS = /[\n\r;$\`|&<>()[\]{}\\]/`blocks`[ ] { } ( )`, -which are legitimate in git pathspecs/brace globs (e.g. `:(exclude)`, `app/{a,b}`). -User-facing convenience command only. - -### F6 — [KEEP / note] tmux-test TMUX_UNSAFE_EXECUTABLES is very broad - -Blocks node/bun/make/find/awk/sed/python/etc. Niche profile; intentional. Note -but do not relax without a dedicated review. - -## Keep — real security value (do NOT remove) - -- Path traversal + outside-project absolute-path containment. -- Force/delete-push block; default-branch push approval gate. -- Privilege escalation (`sudo`/`su`), system package managers, env dumping. -- Sensitive-file / private-key staged-commit scan (`validateStagedCommit`). -- tmux write-through-shell fixture block. -- Interpreter one-liners that read env / spawn subprocesses in read-only mode. - -## Acceptance criteria - -- Each finding has a concrete file:line-backed cause and a proposed minimal change. -- git-committer can produce a multi-line commit and do normal read-only - inspection composition after Tier 1. -- No (C) control is weakened. -- Existing `terminal-command-policy.test.ts` security cases still pass; new - cases cover the relaxed paths. diff --git a/.agents/sessions/shell-policy-audit-2026-07/STATUS.md b/.agents/sessions/shell-policy-audit-2026-07/STATUS.md deleted file mode 100644 index 225965c731..0000000000 --- a/.agents/sessions/shell-policy-audit-2026-07/STATUS.md +++ /dev/null @@ -1,36 +0,0 @@ -# STATUS — Shell/Terminal Security Restriction Audit - -## Current state - -Audit complete (plan mode). No source changed. Enforcement layer fully read and -mapped; findings F1–F6 recorded with file/line evidence in SPEC.md. - -## Completed - -- Read full policy engine (`terminal-command-policy.ts`), executor - (`run-terminal-command.ts`), approval layer (`harness-enforcement.ts`), - git-committer, git-discipline prompt, git-branch, CLI git-command-args. -- Mapped all profile assignments and `gitCommitGuidePrompt` injection point. -- Confirmed root cause of git-committer unreliability (F1: policy vs. guidance - contradiction). - -## Pending (awaiting user decision) - -- Which tier(s) to execute (Tier 1 recommended). -- T1.1 Approach A (minimal heredoc allowance) vs B (structured commit path). - -## Blocked - -- All implementation blocked in plan mode + pending user go-ahead (security- - sensitive changes). - -## Next checkpoint - -User selects tiers + A/B. Then exit plan mode, spawn security-reviewer -(advisory) on `terminal-command-policy.ts`, implement selected tasks, run the -named per-task suites, then full gate. - -## Resume instructions - -Re-read SPEC.md findings + PLAN.md tiers. Start with T1.1 on -`sdk/src/tools/terminal-command-policy.ts`. diff --git a/.agents/sessions/terminal-policy-repair-2026-08/EVENTS.jsonl b/.agents/sessions/terminal-policy-repair-2026-08/EVENTS.jsonl deleted file mode 100644 index 3c15396974..0000000000 --- a/.agents/sessions/terminal-policy-repair-2026-08/EVENTS.jsonl +++ /dev/null @@ -1,4 +0,0 @@ -{"ts":"2026-08-04T15:19:32.662Z","kind":"append_lesson","summary":"Appended entry \"RF tee findings status\" to STATUS.md","payload":{"heading":"RF tee findings status","artifact":"STATUS.md"}} -{"ts":"2026-08-04T15:19:32.663Z","kind":"session_status","summary":"Session status -> validating","payload":{"status":"validating"}} -{"ts":"2026-08-04T21:04:52.636Z","kind":"append_lesson","summary":"Appended entry \"Session complete — 2026-08-04\" to STATUS.md","payload":{"heading":"Session complete — 2026-08-04","artifact":"STATUS.md"}} -{"ts":"2026-08-04T21:04:52.639Z","kind":"session_status","summary":"Session status -> completed","payload":{"status":"completed"}} diff --git a/.agents/sessions/terminal-policy-repair-2026-08/LESSONS.md b/.agents/sessions/terminal-policy-repair-2026-08/LESSONS.md deleted file mode 100644 index 4e2ef36861..0000000000 --- a/.agents/sessions/terminal-policy-repair-2026-08/LESSONS.md +++ /dev/null @@ -1,15 +0,0 @@ -# LESSONS — Terminal policy repair - -## Lessons captured during planning - -- Segment-parsed safety detectors must fail closed: `segments?.some(unsafe) ?? false` is a fail-open hole whenever the segment splitter returns undefined (background `&`, empty segments from trailing/leading `;`, `;;`). Read-only profile already denies on `!segments`; tmux-test detectors skipped that posture. -- Blanket-bans vs. composition-aware checks: removing the raw-newline ban was correct UX, but every downstream guard that parsed "commands" needed re-auditing for the new separator class. Policy changes that widen the input alphabet must be paired with a fail-open review of all segment consumers. -- Reviewer findings are snapshot-bound and RF-ID-keyed: they cannot be cleared conversationally; each repair edit must cite the finding IDs, and only a fresh matching reviewer pass clears them. -- repair-editor requires the structured `handoff` object, not a bare prompt — a prompt-only spawn failed handler validation. -- Consistency between allow guards and message helpers matters: the allow regex accepted only `-m` while placeholder/strip helpers already handled `--message`/`--message=`, producing a confusing generic deny for a documented form. - -## Gotchas for execution - -- `splitReadOnlyShellSegments` treats `\r\n` as one separator; any new test with CRLF should account for that. -- Existing positive tmux-test tests (`normalizes tmux executable quoting…`, `applies outside-absolute-path containment…`) are the regression canary for fail-closed changes. -- `\r|\n` multi-line composition under validation-diagnosis is intentionally still fail-closed unless it matches the bounded `cat > file <<'EOF'…EOF` heredoc — do not loosen this while fixing tmux-test. diff --git a/.agents/sessions/terminal-policy-repair-2026-08/PLAN.md b/.agents/sessions/terminal-policy-repair-2026-08/PLAN.md deleted file mode 100644 index 8c56edb76d..0000000000 --- a/.agents/sessions/terminal-policy-repair-2026-08/PLAN.md +++ /dev/null @@ -1,56 +0,0 @@ -# PLAN — Terminal policy repair (RF-1..RF-6) - - - -Single milestone: close all six reviewer findings, re-validate, pass a fresh reviewer pass. - -## Tasks - -- [ ] T1 — Inventory fail-open tmux detectors - - Role: editor (read phase) or direct read - - Read `sdk/src/tools/terminal-command-policy.ts` fresh and list every tmux-test detector that consumes `splitReadOnlyShellSegments(command)` with `segments?.some(...) ?? false` or equivalent: `hasUnsafeTmuxFileMutation`, `hasUnsafeTmuxSedInPlace`, `hasUnsafeTmuxExecutable`, `hasUnsafeTmuxGitCommand`, `hasUnsafeTmuxWriteRedirection`, `hasActiveTmuxCompoundShellSyntax`. - - Acceptance: complete list of `?? false`/fail-open sites confirmed against live file, not memory. - - Validate: code-search `segments\?\.` and `?? false` in terminal-command-policy.ts. - -- [ ] T2 — Make tmux-test detectors fail closed (RF-1, RF-5) - - Depends on: T1 - - For each detector identified in T1, change the unparseable path from `?? false` to `?? true` (undefined segments ⇒ treat composition as unsafe). Do not touch detectors that genuinely don't parse segments. Keep each function's name/signature. - - Rationale to preserve in a brief comment: `splitReadOnlyShellSegments` returns undefined on background `&`, substitution, or empty segments; bash still executes those forms, so tmux-test must deny rather than skip the guard. - - Acceptance: `touch workspace.txt;`, `tee workspace.txt;`, `rm -rf src &`, `touch x;;echo y`, `; touch x` all denied under `tmux-test`. - - Validate: run policy test file (T5 gate) — no new fail-closed false-positives on the existing allow cases in `blocks tmux agents from direct workspace mutation` / `normalizes tmux executable quoting`. - -- [ ] T3 — Align git-commit allow with `--message` (RF-3, RF-6) - - Depends on: T2 - - Extend the git-commit commit-allow regex from `(?=.*-m(?:\s|$))` to also accept `--message` forms: `(?=.*(?:-m|--message)(?:\s|=|$))`. Keep placeholder-message rejection and non-amend guard unchanged. - - Acceptance: `git commit --message "Fix the parser"` allowed (with real message); `git commit --message probe` still denied as placeholder; `--message="Fix"` allowed. - - Validate: policy tests. - -- [ ] T4 — Add failing-closed test cases (RF-2, RF-4, RF-6) - - Depends on: T2, T3 - - In `sdk/src/__tests__/terminal-command-policy.test.ts` add: - - a tmux-test test asserting `allowed === false` for trailing `;`, `;;`, leading `;`, and background `&` around `touch`/`tee`/`rm` mutators (e.g. `tmux run-shell 'touch /tmp/x;'` shape if fixtures are wrapped, per existing test idioms — mirror the style of `blocks tmux agents from direct workspace mutation`); - - git-commit allow cases: `git commit --message "Fix the parser"`, `git commit --message="Fix the parser"` → allowed true; placeholder via `--message` → false. - - Acceptance: new tests fail against the pre-T2/T3 code and pass after. - - Validate: run policy test file. - -- [ ] T5 — Validate - - Depends on: T4 - - Run `bun test sdk/src/__tests__/terminal-command-policy.test.ts` (must be all-pass) and end turn so hooks (`bun run typecheck`, `cd sdk && bun run typecheck`) run. - - Acceptance: 0 failures; hooks green. - - Validate: basher output + gate hooks summary. - -- [ ] T6 — Fresh reviewer pass (RF-1..RF-6) - - Depends on: T5 - - End turn; harness runs the automated reviewer against the new snapshot. If any finding re-opens, do exactly one targeted repair for that finding ID and re-validate (no broad rewrites). - - Acceptance: GATE: PASSED; all six RF records cleared. - -## Execution notes (execute mode) - -- Edit through `repair-editor` with the full handoff contract (schemaVersion, taskId, role='repair-editor', objective, requirements[] one per RF ID, acceptanceCriteria[] one per RF ID, context: [], nonGoals, findings[] with files + snapshotFingerprint, permissions{readablePaths, writablePaths, allowedTools}). A previous repair-editor spawn failed validation because only a prose prompt was sent — always include the structured `handoff` object and cite finding IDs (RF-1-4391b95f, RF-2-327c10c4, RF-3-fa741f2a, RF-4-7b925458, RF-5-e9fa653a, RF-6-7a8c09df). Use the full snapshot fingerprint from the harness state at execute time (prefix `v3:7fa30d019b80a…`). -- Sequential discipline: read fresh → one repair transaction → run policy tests → end turn. No parallel reviewer during repair. -- Preserve unrelated dirty work: `scripts/measure-context-baseline.ts`, `agents/base2/*`, `docs/*`, `.agents/sessions/context-baseline-25k/` are not ours — do not stage or edit them. - -## Risks / open questions - -- Fail-closed `?? true` could over-deny exotic-but-safe tmux commands whose segment parse returns undefined (e.g. `tmux new-session -d && tmux ls` — currently parsed, fine; substitution forms already denied by hasActiveCommandSubstitution). Existing tmux-test allow tests will surface any regression in T5. -- RF-3 is labeled "Optional consistency" by the reviewer, but it is open-BLOCKING in the gate, so it must be resolved (align or explicit intentional-deny test). diff --git a/.agents/sessions/terminal-policy-repair-2026-08/SPEC.md b/.agents/sessions/terminal-policy-repair-2026-08/SPEC.md deleted file mode 100644 index 9f94563084..0000000000 --- a/.agents/sessions/terminal-policy-repair-2026-08/SPEC.md +++ /dev/null @@ -1,45 +0,0 @@ -# SPEC — Terminal policy repair after reviewer gate (2026-08) - -## Overview - -The blanket "no raw newlines" terminal policy was removed from `evaluateTerminalCommandPolicy` (done, policy tests 33/33 green, typecheck hooks green). The reviewer gate returned 6 blocking findings (RF-1..RF-6) that must be repaired before the gate clears. All work is confined to two files: - -- `sdk/src/tools/terminal-command-policy.ts` -- `sdk/src/__tests__/terminal-command-policy.test.ts` - -Current snapshot fingerprint (use the harness-provided full value at execute time): `v3:7fa30d019b80a…` (files=sdk/src/tools/terminal-command-policy.ts, sdk/src/**tests**/terminal-command-policy.test.ts). - -## Open findings (each repair edit must cite at least one) - -- RF-1-4391b95f: tmux-test mutation/executable/git guards fail open when `splitReadOnlyShellSegments` returns undefined (`segments?.some(...) ?? false`). `touch workspace.txt;`, `tee workspace.txt;`, `rm -rf src &` never hit `hasUnsafeTmuxFileMutation`/`hasUnsafeTmuxExecutable`. Read-only correctly denies on `!segments`; tmux must fail closed the same way. -- RF-2-327c10c4: add tmux-test cases for trailing `;`, `;;`, leading `;`, and background `&` around mutators (touch/tee/rm) asserting `allowed===false`. -- RF-3-fa741f2a: git-commit allow regex only accepts `-m`; `hasPlaceholderCommitMessage`/`stripCommitMessageArgs` handle `--message`/`--message=`. Real `git commit --message "Fix"` gets a generic deny. Align allow with `--message` or encode the intentional deny in a test. -- RF-4-7b925458: test coverage missing for changed behavior. -- RF-5-e9fa653a: requirement: restricted profiles fail closed on unsafe/unparseable shell composition. -- RF-6-7a8c09df: requirement: behavior-changing policy paths have meaningful test coverage. - -## Requirements - -- R1 (RF-1, RF-5): Every tmux-test unsafe-detector built on `splitReadOnlyShellSegments` must treat an `undefined` parse as unsafe (fail closed), matching the existing read-only `!segments → deny` posture. -- R2 (RF-2, RF-4, RF-6): New adversarial test cases in `terminal-command-policy.test.ts` for unparseable/malformed composition around mutators under tmux-test, plus allow/deny coverage for any git-commit message-flag change. -- R3 (RF-3): `git commit --message "Fix"` and `--message=…` are either accepted by the allow clause or explicitly tested as intentionally denied. Preferred: extend the allow regex to accept `-m`/`--message` (both spaced and `=` forms), since other helpers already parse them. -- R4: Do not weaken other guards (git-commit substitution denial incl. double-quoted `$(`/backticks, path containment, workspace deny patterns, validation-diagnosis heredoc handling). - -## Non-goals - -- No refactor of the policy module structure; minimal diff. -- No changes to `run-terminal-command.ts`, git-discipline guidance, docs, or `.agents/sessions/*` history files. -- No re-introduction of the blanket raw-newline ban (already intentionally removed). - -## Acceptance criteria - -- A1: `bun test sdk/src/__tests__/terminal-command-policy.test.ts` passes including the new tmux-test fail-closed cases and git-commit `--message` cases. -- A2: File-change hooks (`bun run typecheck`, `cd sdk && bun run typecheck`) pass. -- A3: A fresh snapshot-bound code-reviewer clears RF-1..RF-6 with no new blockers. - -## Relevant code anchors (verify fresh before editing) - -- `splitReadOnlyShellSegments` (~line 458): returns `undefined` for substitution/backtick, background `&`, or any empty segment (trailing/leading `;`, `;;`). -- tmux-test block in `evaluateTerminalCommandPolicy` (~line 1103): `workspaceWriteSyntax` array of detectors; dependent detectors use `segments?.some(...) ?? false`. -- `hasUnsafeTmuxExecutable` (~line 427) and its siblings (`hasUnsafeTmuxFileMutation`, `hasUnsafeTmuxSedInPlace`, `hasUnsafeTmuxGitCommand`, `hasUnsafeTmuxWriteRedirection`, `hasActiveTmuxCompoundShellSyntax`) — the fail-open `?? false` sites. -- git-commit allow clause (~line 1210): `^git\s+commit\s+(?=.*-m(?:\s|$)).+` guards the commit allow. diff --git a/.agents/sessions/terminal-policy-repair-2026-08/STATE.json b/.agents/sessions/terminal-policy-repair-2026-08/STATE.json deleted file mode 100644 index 9053a0b81a..0000000000 --- a/.agents/sessions/terminal-policy-repair-2026-08/STATE.json +++ /dev/null @@ -1,10 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "terminal-policy-repair-2026-08", - "status": "completed", - "currentTask": null, - "revision": 2, - "checkpoint": null, - "createdAt": "2026-08-04T15:19:32.655Z", - "updatedAt": "2026-08-04T21:04:52.635Z" -} diff --git a/.agents/sessions/terminal-policy-repair-2026-08/STATUS.md b/.agents/sessions/terminal-policy-repair-2026-08/STATUS.md deleted file mode 100644 index 2533c018a0..0000000000 --- a/.agents/sessions/terminal-policy-repair-2026-08/STATUS.md +++ /dev/null @@ -1,39 +0,0 @@ -# STATUS — Terminal policy repair (2026-08-04) - -## Current state - -- Mode: plan. Gate: PENDING (blocked) with 6 open reviewer findings RF-1..RF-6 on snapshot `v3:7fa30d019b80a…`. -- Raw-newline ban removal: implemented; `bun test sdk/src/__tests__/terminal-command-policy.test.ts` was 33/33 green; typecheck hooks green. -- Reviewer pass: BLOCKING on tmux-test fail-open `?? false` guards, missing tmux fail-closed tests, git-commit `-m` vs `--message` inconsistency, missing coverage/requirements. -- Repair-editor spawn attempt failed earlier — next execution must include the full structured `handoff` object (see PLAN.md execution notes). - -## Completed - -- Blanket raw-newline ban removed from `evaluateTerminalCommandPolicy`. -- Newline-aware composition handling added (`normalizeCommand` preserves newlines; `hasUnquotedShellSyntax` treats unquoted newlines as syntax; `splitReadOnlyShellSegments` splits on newlines incl. `\r\n`; validation-diagnosis heredoc strip retained with narrow multi-line fail-closed guard). -- Test update: multi-line multi-command composition still denied under restricted profiles; reason no longer the blanket newline message. -- Local validation: policy tests 33/33 pass; `bun run typecheck` + sdk typecheck pass. - -## Blocked on - -- RF-1-4391b95f, RF-2-327c10c4, RF-3-fa741f2a, RF-4-7b925458, RF-5-e9fa653a, RF-6-7a8c09df — see PLAN.md T1–T6. - -## Next checkpoint - -T5 validation run after repair; then T6 fresh reviewer pass. GATE: PASSED is the completion signal. - -## Resume instructions - -In execute mode: work PLAN.md T1→T6 in order. Use repair-editor with full handoff citing the RF IDs and the harness-provided snapshot fingerprint. Do not touch the unrelated dirty paths listed in PLAN.md risks. - - - -## RF tee findings status — 2026-08-04T15:19:32.653Z - -RF-1-999e85ef / RF-2-356e9c97 claim tee is missing from TMUX_UNSAFE_EXECUTABLES. Live code at sdk/src/tools/terminal-command-policy.ts:267 already includes 'tee'. hasUnsafeTmuxExecutable + resolveTmuxCommand cover bare, command/env-wrapped, and /usr/bin/tee. Suite bun test sdk/src/**tests**/terminal-command-policy.test.ts: 35 pass / 0 fail including "blocks tmux agents from direct workspace mutation". No further source edit required for these RF IDs; needs fresh matching reviewer pass to clear open records. - - - -## Session complete — 2026-08-04 — 2026-08-04T21:04:52.635Z - -All six reviewer findings RF-1..RF-6 verified as already resolved in the live tree during resume; no new source edits were required. Validation green: `bun test sdk/src/__tests__/terminal-command-policy.test.ts` 35/35 pass; `cd sdk && bun run typecheck` clean. Runtime gate: GATE PASSED (no edited files; reviewer verdict LOOKS_GOOD). Session is complete — safe to archive. diff --git a/.agents/sessions/unified-background-jobs/EVENTS.jsonl b/.agents/sessions/unified-background-jobs/EVENTS.jsonl deleted file mode 100644 index b9767fcc01..0000000000 --- a/.agents/sessions/unified-background-jobs/EVENTS.jsonl +++ /dev/null @@ -1,8 +0,0 @@ -{"ts":"2026-07-27T09:28:59.884Z","kind":"task_update","summary":"Updated 1 task line(s): M0.1 Dispatch discovery shard pairs","payload":{"matched":["M0.1 Dispatch discovery shard pairs"]}} -{"ts":"2026-07-27T09:28:59.884Z","kind":"session_status","summary":"Session status -> active","payload":{"status":"active"}} -{"ts":"2026-07-27T09:28:59.884Z","kind":"current_task","summary":"Current task -> \"M0.1 Dispatch discovery shard pairs\"","payload":{"currentTask":"M0.1 Dispatch discovery shard pairs"}} -{"ts":"2026-08-01T14:43:39.938Z","kind":"append_lesson","summary":"Appended entry \"M5 Live UI Complete (verified implemented) — 2026-08-01\" to STATUS.md","payload":{"heading":"M5 Live UI Complete (verified implemented) — 2026-08-01","artifact":"STATUS.md"}} -{"ts":"2026-08-01T14:58:36.879Z","kind":"append_lesson","summary":"Appended entry \"M6 Prompts/Docs/Evals Complete — 2026-08-01\" to STATUS.md","payload":{"heading":"M6 Prompts/Docs/Evals Complete — 2026-08-01","artifact":"STATUS.md"}} -{"ts":"2026-08-02T08:08:21.265Z","kind":"task_update","summary":"Updated 1 task line(s): M7.1 typecheck all; M7.2 test all touched suites; M7.3 live dev-server smoke; M7.4 coverage gate","payload":{"matched":["M7.1 typecheck all; M7.2 test all touched suites; M7.3 live dev-server smoke; M7.4 coverage gate"]}} -{"ts":"2026-08-02T08:08:21.265Z","kind":"session_status","summary":"Session status -> completed","payload":{"status":"completed"}} -{"ts":"2026-08-02T08:08:21.265Z","kind":"current_task","summary":"Current task pointer cleared","payload":{"currentTask":null}} diff --git a/.agents/sessions/unified-background-jobs/LESSONS.md b/.agents/sessions/unified-background-jobs/LESSONS.md deleted file mode 100644 index e61a25e0a4..0000000000 --- a/.agents/sessions/unified-background-jobs/LESSONS.md +++ /dev/null @@ -1,22 +0,0 @@ -# Unified Background Jobs — Lessons & Security Findings - -## Security review (BLOCKING) — ownership must be enforced, not just present - -The security-reviewer flagged the M4 tool migration as BLOCKING. Root cause: when I deleted the authorize/foreign/recover gate + pending-background-jobs Map, I assumed `jobRegistry.assertOwned` enforced ownership — but nothing on the process-job path (check_job/kill_job/read_logs/list_jobs) actually calls it. `assertOwned` was dead code for shell jobs. Result: any reachable jobId (predictable `job-N-hex`, or leaked via end_turn/list_jobs) allowed cross-session log reads and process-group kills. - -**Lesson:** deleting a gate without wiring the replacement enforcer into every operation is a regression, not a simplification. Ownership must be _checked at the point of action_ (esp. the mutating kill path), with the owner derived from **trusted run/session state — never from model/tool input**. - -Fix plan (see SECURITY findings SEC-1..SEC-6): - -1. run.ts computes a trusted owner from sessionState (clientSessionId=run clientSessionId/promptId, rootRunId=sessionState.mainAgentState.runId ?? agentId) and injects it into checkJob/killJob/readLogs/listJobs/runTerminalCommand. Never read `owner` from model input (it was even feeding approval binding via terminalInput.owner.rootRunId — model-controlled). -2. sdk tools call jobRegistry.assertOwned(jobId, trustedOwner); foreign → generic not_found. kill_job gates terminateProcessTree on ownership. -3. getBackgroundJob restampOwner comes only from trusted owner. -4. Recovered-pid liveness binding before kill (SEC-3); orphaned-file sweep fail-closed (SEC-4); recovered-metadata authenticity (SEC-5) deferred if scope grows; list_jobs always scoped + check_job logFile redacted for non-owners (SEC-6). - -## Design lesson — dual-key ids are a trap - -The M3 agent adapter briefly had two ids per job (core `job-N-hex` vs adapter `bg-agent-`), causing consumer/adapter mismatches. Fixed by single-keying: the adapter creates the registry job with the final `bg-agent-` id directly (via an optional explicit jobId on core create). One job, one id, everywhere. - -## Test-authoring lesson - -Parallel editor (impl) + test-writer (tests) against a _sketched_ contract causes drift (job.id vs jobId, chunk vs chunkType/data, matched as boolean vs the matched event, snapshot/wait returning undefined). Write tests against the REAL module, or author impl+tests together. diff --git a/.agents/sessions/unified-background-jobs/PLAN.md b/.agents/sessions/unified-background-jobs/PLAN.md deleted file mode 100644 index 70a09ef3fb..0000000000 --- a/.agents/sessions/unified-background-jobs/PLAN.md +++ /dev/null @@ -1,184 +0,0 @@ -# Unified Background Job Architecture — PLAN - - - -Status legend: `[ ]` pending, `[~]` in_progress, `[x]` done, `[!]` blocked. - -## Milestones - -### M0 — Discovery & blast-radius (DONE) - -- [x] M0.1 Discovery shards machine-confirmed the exact blast radius (see DESIGN §Consumers). - -### M1 — Unified core in `common` (single source of truth) — DONE - -- [x] M1.1 New module `common/src/util/job-registry.ts` + wait/snapshot/stream primitives + 63 unit tests (63/63 pass, common typecheck clean). - -### M2 — Shell (process) adapter in `sdk` — DONE (49/49 sdk tests) - -- [x] M2.1 ProcessJobAdapter: spawn/kill/log capture/quota. -- [x] M2.2 Cross-session recovery as write-only disk projection; removed pendingJobs gate for shell. -- [x] M2.3 Re-pointed run-terminal-command BACKGROUND branch. - -### M3 — Agent adapter in `agent-runtime` — DONE (24/24 agent tests) - -- [x] M3.1 AgentJobAdapter: coroutine, AbortController cancel, chunk streaming, capacity limits (shares singleton, per-kind agent bounds). -- [x] M3.2 Re-pointed spawn-agents background path + background-agent-jobs. - -### M4 — Tool migration (thin wrappers) — DONE (31/31 agent-runtime; all typechecks clean) - -- [x] M4.1 check_job unified event result; M4.2 check_background_agent; M4.3 kill/list/read_logs + end_turn on unified core; M4.4 schemas (full JobState/JobEventPayload/kind); M4.5 DELETED authorize-background-job.ts + common/util/pending-background-jobs.ts. Single-key bg-agent ids. - -### M5 — Live UI — DONE (verified implemented + unit-tested; no new code needed) - -- [x] M5.1 Run loop consumes job event stream (sdk/src/job-update-forwarder.ts + subscribeAll/dispose in run.ts; printModeJobUpdateSchema; run-job-updates.test.ts 6/6); M5.2 CLI renders live job status/output (handleJobUpdate in cli/src/utils/sdk-event-handlers.ts + backgroundJobId correlation + terminal-command-display.tsx; sdk-event-handlers.test.ts 26/26 incl. 13 job_update cases). Live real-terminal dev-server smoke deferred to M7.3. - -### M6 — Prompts, docs, evals, generated bundles — DONE (verified current; no edits needed) - -- [x] M6.1 agent prompts (base2.ts:292 already documents the unified model; editor.ts + base-deep.ts have zero stale authorize/foreign/recover refs); M6.2 docs/deterministic-edit-system.md:28 + agents-and-tools.md:581-628 already document the unified JobRegistry / assertOwned / job*update; no doc still describes the deleted tri-state gate. M6.3 evals clean (evals/buffbench/* reference no job tooling, assert no ownership behavior); cli/release\_/index.js regeneration deferred to the CI release workflow (build-binary.ts is the release path — regenerate, do not hand-edit). - -### M7 — Validation & review - -- [x] M7.1 typecheck all; M7.2 test all touched suites; M7.3 live dev-server smoke; M7.4 coverage gate. (typecheck all packages green; 174/174 sdk+common + 26/26 CLI + 21/21 context-budget; live real-spawn smoke passed; reviewer gates LOOKS_GOOD/NON_BLOCKING with nits applied; commits e5797f4fe (push-model) + 4ce235b51 (context-budget)) - -## DESIGN (authoritative implementation spec) - -### Core types (`common/src/util/job-registry.ts`) - -```ts -export type JobKind = 'process' | 'agent' -export type JobState = - | 'queued' - | 'running' - | 'stopping' - | 'completed' - | 'error' - | 'stopped' - | 'lost' - | 'cancelled' -export const TERMINAL_STATES: ReadonlySet // completed|error|stopped|lost|cancelled - -export interface JobOwner { - clientSessionId: string - rootRunId: string - parentRunId: string - parentAgentId: string -} - -export type JobEventPayload = - | { type: 'output'; data: string } // shell stdout/stderr bytes (agent 'text' chunk) - | { type: 'agent_chunk'; chunkType: string; data: unknown } // agent structured chunk (tool_call/tool_result/subagent_*) - | { - type: 'lifecycle' - state: JobState - exitCode?: number | null - error?: string - } - | { type: 'status'; message?: string } - -export interface JobEvent { - sequence: number - jobId: string - timestamp: number - payload: JobEventPayload -} - -export interface Job { - jobId: string - kind: JobKind - state: JobState - owner: JobOwner - label: string // shell: command; agent: agentType - createdAt: number - startedAt?: number - completedAt?: number - exitCode?: number | null - error?: string - result?: unknown // agent result -} -``` - -### State machine (enforced, un-bypassable) - -- `queued -> running -> stopping -> {completed|error|stopped|lost|cancelled}` -- `queued -> running`; `running -> stopping`; `stopping -> `; `running -> `. -- No transitions out of a terminal state. Adapters NEVER mutate state; they call `registry.emit()` and the registry folds lifecycle events into state. - -### Registry API - -```ts -class JobRegistry { - create(params: { kind: JobKind; label: string; owner: JobOwner }): Job // state=queued, emits lifecycle(queued) - start(jobId): Job // queued->running - emit(jobId, payload: Omit): JobEvent // appends to ring buffer; lifecycle payloads fold into state - get(jobId): Job | undefined - list(owner?: Pick): Job[] // running + settled-within-TTL - listRunning(owner?): Job[] - assertOwned(jobId, owner): Job | { error: 'not_found' | 'foreign' } // ownership enforced INSIDE registry - snapshot( - jobId, - cursor = 0, - ): { - events: JobEvent[] - nextCursor: number - state: JobState - truncated: boolean - dropped: number - } - wait( - jobId, - opts: { - predicate?: (e: JobEvent) => boolean - timeoutMs?: number - cursor?: number - }, - ): Promise<{ - events: JobEvent[] - nextCursor: number - state: JobState - matched: boolean - timedOut: boolean - dropped: number - }> - stream(jobId, cursor?): AsyncIterable // in-process push for UI/run-loop - cancel(jobId): void // emits lifecycle(cancelled/stopping); adapter does the real kill/abort -} -export const jobRegistry: JobRegistry // process-wide singleton (replaces pending-background-jobs Map) -``` - -- Ring buffer: bounded (e.g. 500 events / 256KB per job), oldest evicted, `dropped` counter tracked. Per-consumer cursors; cursor = last-consumed sequence (0 = from start). NO job-global read offset. -- `wait` resolves on (a) predicate match over new events, (b) terminal state, or (c) timeout. Subsumes follow-mode + dev-server readiness. -- TTL sweep for settled jobs (24h shell-consistent, or kind-aware). -- Test hooks: `__clearJobRegistryForTest()`. - -### Adapter pattern (dependency inversion) - -```ts -interface JobAdapter { - readonly kind: JobKind - // adapter owns the real process/coroutine; it ONLY emits events via registry.emit(). -} -// sdk registers ProcessJobAdapter; agent-runtime registers AgentJobAdapter. -// The registry never imports either. Disk metadata (shell) is a write-only recovery projection. -``` - -### Consumers to migrate (M2/M3/M4) - -- common/src/util/pending-background-jobs.ts → REPLACED by job-registry.ts (keep as thin deprecated re-export during transition, or delete + fix all importers). -- sdk: tools/background-jobs.ts, run-terminal-command.ts, check-job.ts, kill-job.ts, list-jobs.ts, read-logs.ts, run.ts(1323). -- agent-runtime: util/background-agent-jobs.ts, tools/handlers/tool/{spawn-agents,check-background-agent,check-job,kill-job,list-jobs,read-logs,authorize-background-job,end-turn}.ts, run-agent-step.ts(32,1096), util/step-loop-guard.ts. -- common tool schemas: tools/params/tool/{check-job,check-background-agent,kill-job,list-jobs,read-logs}.ts. -- cli: components/tools/background-job-tools.tsx, run-terminal-command.tsx, terminal-command-display.tsx, utils/sdk-event-handlers.ts. -- prompts/docs/evals: agents/base2/base2.ts, agents/editor/editor.ts, agents/base2/base-deep.ts, docs/deterministic-edit-system.md, evals/buffbench/\*. -- generated: cli/release/index.js, cli/release-staging/index.js (regenerate, do not hand-edit). - -### Key invariants - -- Single source of truth: the JobRegistry. Disk metadata (shell) is a projection, never consulted for live state. -- Ownership is a job attribute checked in `assertOwned`, inside the registry. The authorize/foreign/recover tri-state gate is DELETED. -- Per-consumer cursors only; at-least-once semantics; explicit cursor is idempotent. -- Bounded memory: ring buffer caps + log-size quota for shell jobs. - -## Current state / resume - -Design + SPEC + M0 discovery complete. Implementing M1 (unified core). Resume at `` above. After M1 lands, do M2 and M3 in parallel, then M4 (coherent single-pass tool migration), then M5/M6, then M7 validation. diff --git a/.agents/sessions/unified-background-jobs/SPEC.md b/.agents/sessions/unified-background-jobs/SPEC.md deleted file mode 100644 index bd9c47216e..0000000000 --- a/.agents/sessions/unified-background-jobs/SPEC.md +++ /dev/null @@ -1,59 +0,0 @@ -# Unified Background Job Architecture — SPEC - -## Goal - -Replace the fragmented background-job system (two job models, three state stores, two read/cursor semantics, an authorize/re-stamp gate, manual sleep-polling) with a single, unified architecture that is excellent for dev servers and long-running agents, and renders live activity in the CLI. Breaking changes to tool contracts are allowed (user approved). - -## Current-state problems (verified by reading source) - -- Two job models: shell `ChildProcess` jobs (sdk/tools/background-jobs.ts, log file + metadata file) and in-process agent coroutines (agent-runtime/util/background-agent-jobs.ts, in-memory ring buffer). Different lifecycles, TTLs (24h vs 30min), recovery models, cancellation. -- Three state stores for shell jobs: SDK in-memory `jobs` Map, shared `pendingJobs` Map (common/util/pending-background-jobs.ts), and disk metadata files. They drift; "recover" exists only to reconcile them. -- Ownership via authorize-background-job.ts: owned/foreign/recover tri-state; "recover" re-stamps ownership from disk. This produced the user's "unavailable to this run" failure. -- Two read semantics: shell uses a single job-global `readOffset` mutated on read (read starvation across consumers); agents use readOffset + consumerCursors + cursor/nextCursor + droppedChunks. -- Polling UX is archaic: follow mode is a manual `await sleep(200)` loop in check-job.ts; no join/wait primitive; a separate step-loop-guard must special-case polling tools. - -## Non-goals - -- Do not change foreground (SYNC) run_terminal_command behavior. -- Do not change the foreground spawn_agents aggregation path. -- Do not add a real agent-facing push/subscribe channel (the model can only act via request/response tool calls; live UI consumes the stream in-process instead). -- No production/deploy/git-commit actions. - -## Requirements (acceptance criteria) - -1. ONE `JobRegistry` (common package) is the single source of truth for all background jobs, with `kind: 'process' | 'agent'`. No dependency on child_process/fs/agent-runtime in the registry core. -2. ONE lifecycle state machine: `queued -> running -> stopping -> {completed|error|stopped|lost|cancelled}`. Adapters emit events; the registry folds them into state (state machine is un-bypassable). -3. ONE event model: a bounded, sequenced ring buffer of `JobEvent` envelopes `{ sequence, timestamp, jobId, type: 'output'|'status'|'lifecycle', payload: { kind, ... } }`. Per-consumer sequence cursors; no job-global read offset. -4. A `wait(jobId, { predicate?, timeoutMs? })` join primitive that resolves on lifecycle settle or when output matches a predicate (regex/substring) — subsumes follow-mode and dev-server readiness. Plus a non-blocking `snapshot(jobId, cursor?)` read. -5. Ownership is a job attribute enforced inside the registry; remove the authorize/foreign/recover gate and the metadata re-stamp path. -6. Shell jobs: spawn (bash, windows bash resolution), process-group kill (SIGTERM w/ SIGKILL escalation), log capture to a per-job temp file (O_EXCL+O_NOFOLLOW, 0o600), log-size quota termination, and dev-server readiness via `wait`. Cross-session recovery persists as a write-only disk projection of the same event stream; a dead live job reconciles to `lost`. -7. Agent jobs: coroutine-backed, AbortController cancel, chunk streaming into the same EventLog, capacity limits, no disk recovery (process-scoped). -8. Tools become thin wrappers over the registry: `check_job`, `check_background_agent` (new `{events, nextCursor, state, truncated}` + settled result/error), `kill_job`, `list_jobs`, `read_logs`. `wait_for` becomes a predicate over the event stream. end_turn leak detection reads the unified registry. -9. Live UI: the run loop consumes the job event stream in-process and emits job activity to the CLI chat so users see live status/output without the agent polling. CLI renders live job status in the terminal-command and job tool components. -10. Prompts, docs (docs/deterministic-edit-system.md), and evals updated for the new contracts. Generated release bundles (cli/release, cli/release-staging) regenerated, not hand-edited. -11. step-loop-guard polling-loop detection reads the unified model (or is simplified if join replaces most polling). - -## Architecture (decided with thinker) - -- Registry in `common` (dependency inversion): `sdk` and `agent-runtime` each register a kind-adapter; the registry never imports them. Preserves `sdk→common`, `agent-runtime→common` DAG. -- Adapters: `ProcessJobAdapter` (sdk) and `AgentJobAdapter` (agent-runtime). Adapters emit events; registry folds to state. -- Internal push stream (async iterator) for UI/run-loop; agent-facing surface is `wait`/`snapshot` over per-consumer cursors. -- Event envelope single type; payload variant distinguishes shell output bytes vs agent structured chunk. -- Disk metadata = optional recovery projection, not a parallel truth. - -## Relevant systems / files - -- common: util/pending-background-jobs.ts (to be superseded by new registry module) -- sdk: tools/background-jobs.ts, tools/run-terminal-command.ts, tools/check-job.ts, tools/kill-job.ts, tools/list-jobs.ts, tools/read-logs.ts -- agent-runtime: util/background-agent-jobs.ts, tools/handlers/tool/{check-background-agent,check-job,kill-job,list-jobs,read-logs,authorize-background-job,end-turn,spawn-agents}.ts, util/step-loop-guard.ts, tools/stream-parser.ts (fileProcessingState / job wiring) -- common tool schemas: common/src/tools/params/tool/{check-job,check-background-agent,kill-job,list-jobs,read-logs}.ts -- cli: components/tools/background-job-tools.tsx, components/tools/run-terminal-command.tsx, components/terminal-command-display.tsx, utils/sdk-event-handlers.ts -- prompts/docs/evals: agents/base2/base2.ts, agents/editor/editor.ts, agents/base2/base-deep.ts, docs/deterministic-edit-system.md, evals/buffbench/\* -- generated: cli/release/index.js, cli/release-staging/index.js, cli/src/data/initial-agent-type-sources.generated.ts - -## Validation gates - -- bun typecheck across sdk, agent-runtime, common, cli. -- bun test for all touched packages' background-job and tool tests. -- Regenerate release bundles via their scripts; run cli release-wrapper tests. -- Live dev-server smoke: start a dev server as a background job, `wait` for readiness, read_logs, kill_job — verify no manual polling needed. diff --git a/.agents/sessions/unified-background-jobs/STATE.json b/.agents/sessions/unified-background-jobs/STATE.json deleted file mode 100644 index 58df2b16d0..0000000000 --- a/.agents/sessions/unified-background-jobs/STATE.json +++ /dev/null @@ -1,24 +0,0 @@ -{ - "schemaVersion": 2, - "slug": "unified-background-jobs", - "status": "completed", - "currentTask": null, - "revision": 2, - "checkpoint": { - "taskId": "M7.1", - "phase": "validation", - "passed": true, - "summary": "All M7 validation complete: typecheck across all packages green (gate hooks); job suites 174/174, CLI sdk-event-handlers 26/26, context-budget 21/21; live real-spawn smoke passed (startBackgroundJob → check_job follow matched READY-TOKEN → list_jobs digest with matching jobId → kill clean); reviewer gates LOOKS_GOOD then NON_BLOCKING with the 4 nits applied and re-validated 26/26.", - "receiptIds": [ - "wWpA16kCa58", - "wWpA4NvQGRk", - "wXa1kYgnpek", - "wWvc22T4Tsc", - "wYL0Vejndyo", - "wYL0XN5WtCE" - ], - "recordedAt": "2026-08-02T08:08:21.264Z" - }, - "createdAt": "2026-07-27T09:28:59.883Z", - "updatedAt": "2026-08-02T08:08:21.264Z" -} diff --git a/.agents/sessions/unified-background-jobs/STATUS.md b/.agents/sessions/unified-background-jobs/STATUS.md deleted file mode 100644 index 4a1f7b782a..0000000000 --- a/.agents/sessions/unified-background-jobs/STATUS.md +++ /dev/null @@ -1,63 +0,0 @@ -# Unified Background Job Architecture — STATUS - -## Current state (workspace.v1.827) — M0–M4 + security fix GREEN - -### Done - -- **M0** Discovery/blast-radius machine-confirmed. -- **M1** Unified `common/src/util/job-registry.ts` core (lifecycle state machine, sequenced ring buffer, per-consumer cursors, wait/snapshot/stream, in-registry ownership). 63 unit tests. -- **M2** Shell (process) adapter in `sdk` on the unified core (spawn/kill/log capture/quota, cross-session recovery as write-only disk projection). -- **M3** Agent adapter in `agent-runtime` on the shared singleton (per-kind bounds, single-key ids, coroutine cancel, chunk streaming). -- **M4** Tool surface migrated (check_job/check_background_agent/kill_job/list_jobs/read_logs/end_turn) onto the unified core; legacy `pending-background-jobs` Map and `authorize/foreign/recover` gate deleted. -- **Security fix (SEC-1..6)** landed and validated: - - Process-job ops stamp a TRUSTED owner from agentState/session (never model input), enforced via `jobRegistry.assertOwned`. - - Follow-timeout hang fixed: `deadline = Date.now() + timeoutMs` moved to function entry in `sdk/src/tools/check-job.ts`. - - TS2590 fixed: `ProcessJobClientToolCall` is a standalone structural type; handlers cast at the forward boundary. - - Recovery re-attach lockout fixed: added `JobRegistry.restampOwner()`; `getBackgroundJob` upgrades only a placeholder owner to the trusted owner (never overwrites a real owner). - - 3 agent-runtime handler tests retyped `forwardedToolCall` to `ProcessJobClientToolCall`. - -### Validation (workspace.v1.827) - -- Typechecks: common:0, sdk:0, agent-runtime:0 (clean). -- sdk background-job suites: 55 pass / 0 fail. -- agent-runtime + common job suites: 94 pass / 0 fail. - -### Remaining - -- **M5** Wire live job event stream to CLI; render live job activity. -- **M6** Update prompts, docs, evals; regenerate release bundles (controlled breaks). -- **M7** Final cross-package validation + dev-server smoke + coverage gate; final review. - -## Resume instructions - -Backend unification + security hardening is complete and green. Next checkpoint is M5 (live UI). The full design lives in `PLAN.md`; the security findings/decisions live in `LESSONS.md`. - - - -## M5 Live UI Complete (verified implemented) — 2026-08-01 — 2026-08-01T14:43:39.935Z - -M5 needed NO new code — the live-UI chain was already implemented and unit-tested; the plan's IN PROGRESS checkbox was stale. Verified by reading source + running the two relevant suites. - -Run-loop forwarding (M5.1): `sdk/src/job-update-forwarder.ts` (`createJobUpdateForwarder`) is subscribed once in `sdk/src/run.ts` with the trusted owner and disposed on every terminal path (abort + normal completion), so the process-wide `jobRegistry` singleton never leaks listeners. `common/src/types/print-mode.ts` registers `printModeJobUpdateSchema` (additive union member). Owner-scoped: rejects foreign + `UNKNOWN_JOB_OWNER`; forwards lifecycle+output only. - -CLI render (M5.2): `handleJobUpdate` in `cli/src/utils/sdk-event-handlers.ts` (registered in the match) updates correlated tool blocks (lifecycle + 50k tail-bounded output + flag-deduped error append) and agent blocks (status + truncated error). Correlation wired in production: `tool_call` carries `backgroundJobId`; `handleRegularToolCall` stores it. `cli/src/components/terminal-command-display.tsx` renders `job · status · detached · log`. - -Validation: `bun test sdk/src/__tests__/run-job-updates.test.ts` = 6 pass/0 fail; `bun test cli/src/utils/__tests__/sdk-event-handlers.test.ts` = 26 pass/0 fail (incl. all 13 job_update cases). - -Remaining for M5's spirit: the LIVE real-terminal dev-server smoke — that is M7.3, intentionally deferred. - -Next: M6 (prompts/docs/evals/generated bundles). M6.1 base2 prompt (line 292) already documents live-surface behavior; verify editor.ts/base-deep.ts prompts, docs/deterministic-edit-system.md, and regenerate cli/release bundles. - - - -## M6 Prompts/Docs/Evals Complete — 2026-08-01 — 2026-08-01T14:58:36.878Z - -M6 needed NO new edits — prompts, docs, and evals were already aligned with the unified job contracts. Verified by source search, not assumed. - -M6.1 (agent prompts): `agents/base2/base2.ts:292` already documents the unified model (BACKGROUND process_type → jobId; check_job/wait_for readiness; kill_job; list_jobs rediscovers shell+agent jobs; "live job status and output surfaced automatically"). `agents/editor/editor.ts` and `agents/base2/base-deep.ts` contain ZERO references to the deleted authorize/foreign/recover/pending-background-jobs model. - -M6.2 (docs): `docs/deterministic-edit-system.md:28` and `docs/agents-and-tools.md:581-628` already describe the unified JobRegistry (process+agent kinds, one lifecycle machine, assertOwned foreign→not-found, restampOwner, additive job_update). No doc still documents the deleted tri-state gate or pending-background-jobs Map. - -M6.3 (evals + bundles): evals/buffbench/_ are clean — no job-tool references, no ownership-behavior assertions. cli/release_/index.js regeneration DEFERRED to the CI release workflow (user decision): build-binary.ts is the CI release path (requires version, compiles platform binaries), and the bundles are regenerate-not-hand-edit artifacts. - -Next: M7 (final cross-package validation + live dev-server smoke + coverage gate + final review). diff --git a/.agents/types/agent-definition.ts b/.agents/types/agent-definition.ts index ec085ac848..7296e09816 100644 --- a/.agents/types/agent-definition.ts +++ b/.agents/types/agent-definition.ts @@ -39,16 +39,6 @@ export interface AgentDefinition { */ model?: ModelName - /** - * Optional wall-clock timeout in milliseconds for a single execution of this - * agent as a subagent. When set, executeSubagent uses this as the deadline - * (overridable per-spawn via spawn_agents' timeout_seconds). Undefined falls - * back to the shared DEFAULT_SUBAGENT_TIMEOUT_MS, which is -1 (disabled): by - * default there is no wall-clock timeout, so long-running agents run to - * completion. Set a positive value to opt this agent into a wall-clock bound. - */ - defaultTimeoutMs?: number - /** Maximum subagent nesting depth. Defaults to the runtime limit. */ maxSpawnDepth?: number diff --git a/.agents/types/tools.ts b/.agents/types/tools.ts index 172984f5bd..32cd35f189 100644 --- a/.agents/types/tools.ts +++ b/.agents/types/tools.ts @@ -794,7 +794,7 @@ export interface RunTerminalCommandParams { detach?: boolean /** The working directory to run the command in. Default is the project root. */ cwd?: string - /** Set to -1 for no timeout. Does not apply for BACKGROUND commands. Default 30 */ + /** Wall-clock bound in seconds for SYNC commands. Omit or use -1 for no timeout (the default). Does not apply to BACKGROUND commands. */ timeout_seconds?: number /** Runtime-managed background job owner; agents must omit. */ owner?: { @@ -945,15 +945,13 @@ export interface SpawnAgentsParams { constraints?: string[] } | Record - /** Optional wall-clock deadline seconds; omit or -1 for none. Agent defaultTimeoutMs still applies when set. */ - timeout_seconds?: number /** Parameters object for the agent */ params?: { /** Terminal command to run (basher, tmux-cli) */ command?: string /** What information from the command output is desired (basher) */ what_to_summarize?: string - /** Timeout for command. Set to -1 for no timeout. Default 30 (basher) */ + /** Timeout for command in seconds. Omit or -1 for no timeout (default). */ timeout_seconds?: number /** Save full command output to a /tmp log and extract failure lines for long SYNC command output (basher) */ save_full_log?: boolean diff --git a/agents/__tests__/base2.test.ts b/agents/__tests__/base2.test.ts index 97a0125db2..6cb5c6e093 100644 --- a/agents/__tests__/base2.test.ts +++ b/agents/__tests__/base2.test.ts @@ -722,12 +722,7 @@ describe('base2 validation/reviewer coordination prompts', () => { expect(base2.systemPrompt).toContain( 'validation failure/timeout blocks completion even if review looks good', ) - expect(base2.systemPrompt).toContain( - 'Omit top-level `timeout_seconds` for editor and other productive subagents', - ) - expect(base2.systemPrompt).toContain( - 'omitted and `-1` mean no wall-clock deadline', - ) + expect(base2.systemPrompt).not.toContain('timeout_seconds` for editor') // specialistRoutingSection is relocated to a guide under default-on // disclosure; assert the relocation pointer in systemPrompt and keep the // verbatim-line contract on the explicit-off surface instead. @@ -1010,6 +1005,21 @@ describe('base2 validation/reviewer coordination prompts', () => { 'After completing the user request, summarize your changes', ) }) + + test('tells the orchestrator to size delegated work to the child window', () => { + const base2 = createBase2('default') + + expect(base2.systemPrompt).toContain( + "Size work to the child's context window", + ) + }) + + test('names the catalog window suffix and the receipt contextUsage field', () => { + const base2 = createBase2('default') + + expect(base2.systemPrompt).toContain('[context ~200k]') + expect(base2.systemPrompt).toContain('contextUsage') + }) }) describe('base-deep prompt naming and tool guidance', () => { diff --git a/agents/__tests__/dependency-manager.test.ts b/agents/__tests__/dependency-manager.test.ts index 2bee202bbe..77b835bb74 100644 --- a/agents/__tests__/dependency-manager.test.ts +++ b/agents/__tests__/dependency-manager.test.ts @@ -135,7 +135,7 @@ describe('dependency-manager', () => { expect(advancePastSnapshotRead(generator, environment).value).toMatchObject( { toolName: 'run_terminal_command', - input: { command: expected, timeout_seconds: 600 }, + input: { command: expected, timeout_seconds: -1 }, }, ) }) diff --git a/agents/base2/base2.ts b/agents/base2/base2.ts index d90f9dad0a..60555b1234 100644 --- a/agents/base2/base2.ts +++ b/agents/base2/base2.ts @@ -616,7 +616,7 @@ ${ - **Plan artifact maintenance:** In PLAN mode create and maintain durable artifacts; in EXECUTE_PLAN keep STATUS.md and LESSONS.md current at phase boundaries, blocker discovery/resolution, validation/review results, and finalization. Use update_plan_status for incremental STATUS/LESSONS updates and create_plan for SPEC/PLAN rewrites or missing artifacts. Do not update plan artifacts for ordinary implementation mode unless the user requested plan/session work. - **Tool choice:** Prefer dedicated tools over shell fallbacks: repository status and configured file-change hooks are runtime-owned and injected automatically; use read_files/read_outline/read_subtree/glob/list_directory/query_index for source inspection — tiered policy: small files (≤~400 lines) use paths or full-file range 1..totalLines for Tier1 whole-file auth; large/targeted blocks use windows/around/symbol for Tier2 scoped caps (must be complete:true to mint). After successful edit_transaction, compress body to path/pointer but retain whole-file postEditCapabilities verbatim. Don't force windows for small files. Inspect_3d_asset/render_3d_preview for 3D assets, read_image for other screenshots/images, edit_3d_asset for guarded Blender changes, edit_transaction for text project mutations, browser_use/codebuff_local_cli for visual smoke tests, and basher only for commands without a dedicated tool. \`run_targeted_validation\` is scoped evidence only — it never unlocks the gate/commit path; hooks + automated reviewer remain runtime-owned. - **Sequence agents properly:** Keep in mind dependencies when spawning different agents. Don't spawn agents in parallel that depend on each other. -- **Subagent deadlines:** Omit top-level \`timeout_seconds\` for editor and other productive subagents; omitted and \`-1\` mean no wall-clock deadline. Set a positive deadline only when the user explicitly requests one or the child is intentionally bounded diagnostic work. +- **Size work to the child's context window:** The spawnable-agent catalog reports each agent's context window (e.g. \`[context ~200k]\`). Scope one child to work that fits well inside its window: split broad discovery into several narrower shards rather than giving one agent a whole subsystem, and prefer targeted \`read_files\` selectors over whole-file reads in the handoff. A returned agent receipt reports \`contextUsage\` (tokens, window, percent): when a child came back above roughly 70% of its window, or reported repeated compaction, split the next comparable task into more shards instead of retrying it whole. - **Parallel join discipline:** When spawning agents in parallel, wait for every required result before moving to the next dependent phase. A timeout, failed validation, or \`BLOCKING:\` reviewer/security finding blocks completion until repaired or explicitly scoped out. - **Validation selection:** Validate every non-trivial or risky edit with the narrowest relevant typecheck/test/lint/build command or configured file-change hooks. Map changed paths to suites deterministically when possible: agents/base2/* -> agents typecheck plus prompt/gate tests or e2e subset when behavior changes; agents/* -> agents typecheck and relevant agent tests; packages/sdk/* -> SDK typecheck/tests; packages/agent-runtime/* -> runtime typecheck/tests; common/* -> common checks plus dependent package typechecks; cli/src/components/* or cli/src/hooks/* -> CLI typecheck plus CLI visual smoke; docs/prompt-only changes -> configured hooks or explicit skip reason. Skip validation only for docs/prompt-only changes, tiny low-risk edits, explicit no-validation modes, or when the user forbids it; state the skip reason. Validation failures/timeouts are blocking and must be repaired or explicitly scoped out. Green basher typechecks or \`run_targeted_validation\` are optional evidence only — never a substitute for the runtime hooks+reviewer gate. - **Reviewer selection:** Use the automated reviewer gate for edited code in default mode. Spawn code-reviewer manually only for user-requested extra review, advisory/pre-edit review, significant diffs outside the automated gate, or changed code whose risk warrants another perspective; spawn security-reviewer for auth, crypto, secrets, permissions, injection, sandboxing, path/process/network handling, supply-chain, or production-risk changes;${planOnly ? '' : ' spawn test-writer when behavior changes lack coverage;'} spawn debugger after repeated validation failure, runtime failure, or unclear crash behavior. Do not duplicate the same post-edit review manually. @@ -2028,7 +2028,6 @@ ${guideSections} command: group.testCommand, what_to_summarize: 'Report whether the writer-requested validation command passed, including exact failure lines.', - timeout_seconds: 300, }, }, ], diff --git a/agents/basher.ts b/agents/basher.ts index 046ab224a6..8ec66bf49f 100644 --- a/agents/basher.ts +++ b/agents/basher.ts @@ -49,7 +49,8 @@ const basher: AgentDefinition = { }, timeout_seconds: { type: 'number', - description: 'Set to -1 for no timeout. Default 30', + description: + 'Optional wall-clock bound in seconds. Omit or -1 for no timeout (the default).', }, process_type: { type: 'string', diff --git a/agents/dependency-manager/dependency-manager.ts b/agents/dependency-manager/dependency-manager.ts index f08f029151..7eb4ebb5be 100644 --- a/agents/dependency-manager/dependency-manager.ts +++ b/agents/dependency-manager/dependency-manager.ts @@ -52,7 +52,7 @@ const definition: SecretAgentDefinition = { minimum: 1, maximum: 1800, description: - 'Bounded timeout for each package-manager command. Defaults to 600 seconds.', + 'Optional wall-clock bound for each package-manager command. Package-manager commands are unbounded by default.', }, }, required: ['manager', 'operation'], @@ -90,9 +90,10 @@ const definition: SecretAgentDefinition = { typeof params?.workspace === 'string' ? params.workspace.trim() : '' const timeoutSeconds = typeof params?.timeout_seconds === 'number' && - Number.isFinite(params.timeout_seconds) + Number.isFinite(params.timeout_seconds) && + params.timeout_seconds > 0 ? Math.max(1, Math.min(1800, Math.floor(params.timeout_seconds))) - : 600 + : -1 const quote = (value: string) => `'${value.replaceAll("'", `'\\''`)}'` const packageArgs = packages.map(quote).join(' ') const emit = ( diff --git a/agents/guides/editor-writers-and-repair.md b/agents/guides/editor-writers-and-repair.md index 3e13d1f461..e06a56c768 100644 --- a/agents/guides/editor-writers-and-repair.md +++ b/agents/guides/editor-writers-and-repair.md @@ -42,8 +42,6 @@ Use phase-triggered delegation, not random spawns: - **`repair-editor`** — validation/reviewer repairs with exact diagnostics or finding IDs. Prefer runtime-owned repair loops over free-form re-edits when the gate already owns the findings. - **`test-writer` / `doc-writer`** — when documentation or test coverage is required or directly implied by acceptance criteria. Pass `params.target_files` / `params.source_files` (and `test_command` / optional `target_doc_files`) plus a self-contained verified contract in the prompt; writers do not inherit parent history. -Subagent deadlines: omit top-level `timeout_seconds` for productive editors/writers (`-1` / omitted = no wall-clock deadline) unless the user requests a bound or the child is intentionally diagnostic. - ## Automated aux gates (pre-reviewer) When the automated gate is on and edits produced a non-empty pending file set, base2 runs **pre-reviewer aux work once per distinct aux-relevant pending set**, then the final hooks + `code-reviewer` gate. diff --git a/agents/librarian/librarian.ts b/agents/librarian/librarian.ts index a96a08ecc6..66197fe025 100644 --- a/agents/librarian/librarian.ts +++ b/agents/librarian/librarian.ts @@ -174,7 +174,6 @@ When you are done, call set_output with status: "answered", your answer, all rel shellQuote(repoUrl) + ' ' + shellQuote(cloneDir), - timeout_seconds: 180, }, } diff --git a/agents/tmux-cli.ts b/agents/tmux-cli.ts index 11ce048316..8ff68c9e9a 100644 --- a/agents/tmux-cli.ts +++ b/agents/tmux-cli.ts @@ -530,7 +530,6 @@ esac toolName: 'run_terminal_command', input: { command: setupScript, - timeout_seconds: 30, }, includeToolCall: false, } diff --git a/agents/types/agent-definition.ts b/agents/types/agent-definition.ts index 78cdf572ca..cd84b0c5c8 100644 --- a/agents/types/agent-definition.ts +++ b/agents/types/agent-definition.ts @@ -40,16 +40,6 @@ export interface AgentDefinition { */ model?: ModelName - /** - * Optional wall-clock timeout in milliseconds for a single execution of this - * agent as a subagent. When set, executeSubagent uses this as the deadline - * (overridable per-spawn via spawn_agents' timeout_seconds). Undefined falls - * back to the shared DEFAULT_SUBAGENT_TIMEOUT_MS, which is -1 (disabled): by - * default there is no wall-clock timeout, so long-running agents run to - * completion. Set a positive value to opt this agent into a wall-clock bound. - */ - defaultTimeoutMs?: number - /** Maximum subagent nesting depth. Defaults to the runtime limit. */ maxSpawnDepth?: number diff --git a/agents/types/tools.ts b/agents/types/tools.ts index 172984f5bd..32cd35f189 100644 --- a/agents/types/tools.ts +++ b/agents/types/tools.ts @@ -794,7 +794,7 @@ export interface RunTerminalCommandParams { detach?: boolean /** The working directory to run the command in. Default is the project root. */ cwd?: string - /** Set to -1 for no timeout. Does not apply for BACKGROUND commands. Default 30 */ + /** Wall-clock bound in seconds for SYNC commands. Omit or use -1 for no timeout (the default). Does not apply to BACKGROUND commands. */ timeout_seconds?: number /** Runtime-managed background job owner; agents must omit. */ owner?: { @@ -945,15 +945,13 @@ export interface SpawnAgentsParams { constraints?: string[] } | Record - /** Optional wall-clock deadline seconds; omit or -1 for none. Agent defaultTimeoutMs still applies when set. */ - timeout_seconds?: number /** Parameters object for the agent */ params?: { /** Terminal command to run (basher, tmux-cli) */ command?: string /** What information from the command output is desired (basher) */ what_to_summarize?: string - /** Timeout for command. Set to -1 for no timeout. Default 30 (basher) */ + /** Timeout for command in seconds. Omit or -1 for no timeout (default). */ timeout_seconds?: number /** Save full command output to a /tmp log and extract failure lines for long SYNC command output (basher) */ save_full_log?: boolean diff --git a/cli/src/components/__tests__/sweep-boxes.test.tsx b/cli/src/components/__tests__/sweep-boxes.test.tsx index 93b092b6c1..be4745dd5b 100644 --- a/cli/src/components/__tests__/sweep-boxes.test.tsx +++ b/cli/src/components/__tests__/sweep-boxes.test.tsx @@ -41,6 +41,15 @@ const compactionBlock = ( ...overrides, }) +/** + * `renderToStaticMarkup` can separate adjacent text children with `` + * markers, and the ProgressBar renders its percent as ' ', the number and '%'. + * Stripping the markers lets the percent be asserted as the substring it + * actually displays as. + */ +const withoutTextSeparators = (markup: string): string => + markup.replaceAll('', '') + describe('CompactionBox', () => { test('renders headline counts, per-category rows and the retained-memory line', () => { const markup = renderToStaticMarkup( @@ -105,6 +114,89 @@ describe('CompactionBox', () => { expect(markup).toContain('Window 200k · trigger 150k · target 70k') }) + test('renders a live progress bar for a pending pass', () => { + const markup = renderToStaticMarkup( + , + ) + + expect(markup).toContain('Compacting context…') + // 40% of the 24-column bar: 10 filled cells, 14 empty ones, and the percent + // so the pass is not a silent live card. + expect(markup).toContain('█'.repeat(10)) + expect(markup).not.toContain('█'.repeat(11)) + expect(markup).toContain('░'.repeat(14)) + expect(withoutTextSeparators(markup)).toContain(' 40%') + }) + + test('renders the completed bar of a settled transient pass before its hold elapses', () => { + const markup = renderToStaticMarkup( + , + ) + + // A self-dismissing card holds at 100% first, so the bar is full with no + // empty cells left and the settled result lines still render. + expect(markup).toContain('█'.repeat(24)) + expect(markup).not.toContain('░') + expect(withoutTextSeparators(markup)).toContain(' 100%') + expect(markup).toContain('Context compacted') + expect(markup).toContain('190k → 120k tokens (−37%)') + }) + + test('renders no progress bar for a degraded pass and keeps its warning lines', () => { + const markup = renderToStaticMarkup( + , + ) + + expect(markup).not.toContain('█') + expect(markup).not.toContain('░') + // The existing warning lines are untouched by the progress affordance. + expect(markup).toContain('Context trimmed (emergency)') + expect(markup).toContain('Still over budget by 12.4k tokens') + expect(markup).toContain('No knowledge memory retained') + + // Control: the same pass marked transient does draw the completed bar, so + // the absence above is the degradation gate rather than a renderer that + // never draws one. + expect( + renderToStaticMarkup( + , + ), + ).toContain('█'.repeat(24)) + }) + test('renders a replayed pending block as an interrupted pass, never as live', () => { // Persisted blocks are replayed on reload. A pass the user aborted // mid-compaction was written by an earlier process, so its liveSessionId diff --git a/cli/src/components/renderers/compaction-box.tsx b/cli/src/components/renderers/compaction-box.tsx index eed7815dba..4f9c2da71e 100644 --- a/cli/src/components/renderers/compaction-box.tsx +++ b/cli/src/components/renderers/compaction-box.tsx @@ -1,7 +1,8 @@ -import { memo } from 'react' +import { memo, useEffect, useState } from 'react' import { HarnessBox } from './harness-box' import { useTheme } from '../../hooks/use-theme' +import { ProgressBar } from '../progress-bar' import { CLI_LIVE_SESSION_ID } from '../../types/chat' import { formatStatusTokenCount } from '../../utils/status-bar-chips' @@ -56,6 +57,18 @@ const isLiveCompaction = (block: CompactionContentBlock): boolean => /** Shown for a pass that ended before it reported a result. */ const INTERRUPTED_TEXT = 'Interrupted before this pass reported a result.' +/** Rendered width of the compaction progress bar, in columns. */ +const PROGRESS_BAR_WIDTH = 24 + +/** + * How long a settled `transient` card stays visible after it reaches 100% + * before it hides itself: long enough to see the bar complete, short enough + * that a healthy pass does not linger in the transcript. Hiding is purely + * visual — the event handler is what removes the block from state (at turn end + * or on abort), so the two cannot fight over ownership. + */ +const TRANSIENT_COMPACTION_HOLD_MS = 1_200 + /** * The pending/interrupted/declined triple that both the tone and every rendered * line depend on. @@ -135,6 +148,24 @@ interface CompactionBoxProps { export const CompactionBox = memo(({ block }: CompactionBoxProps) => { const theme = useTheme() + // A healthy settled pass is a transient progress affordance: it holds at 100% + // briefly and then renders nothing. The timer lives here rather than in the + // event handler so state stays a pure function of the events, and its cleanup + // is what makes a card dropped mid-hold harmless. + const transient = block.transient === true + const [holdExpired, setHoldExpired] = useState(false) + useEffect(() => { + setHoldExpired(false) + if (!transient) return + const timeout = setTimeout( + () => setHoldExpired(true), + TRANSIENT_COMPACTION_HOLD_MS, + ) + return () => clearTimeout(timeout) + // Re-armed when the block identity changes, so a card replaced in place by a + // later pass gets its own hold instead of inheriting an expired one. + }, [transient, block]) + const presentation = derivePresentation(block) const { unsettled, pending, interrupted, declined } = presentation const tone = deriveTone(block, presentation) @@ -201,6 +232,14 @@ export const CompactionBox = memo(({ block }: CompactionBoxProps) => { ? 'Knowledge memory retained' : 'No knowledge memory retained' + // The bar reports live movement for a pending pass and the completed 100% for + // a transient one; `sanitizeCount` absorbs a missing or garbage percent from a + // replayed block, and `ProgressBar` clamps the upper bound itself. + const showProgressBar = pending || transient + const progressValue = sanitizeCount(block.progressPercent) + + if (transient && holdExpired) return null + return ( {unsettled ? ( @@ -220,6 +259,9 @@ export const CompactionBox = memo(({ block }: CompactionBoxProps) => { {messagesText} )} + {showProgressBar ? ( + + ) : null} {categoryDeltas.length > 0 ? categoryDeltas.map((delta, index) => ( 0 ? formatTimeout(timeoutSeconds) : null diff --git a/cli/src/data/initial-agent-type-sources.generated.ts b/cli/src/data/initial-agent-type-sources.generated.ts index 40c450e8f5..761860bf8f 100644 --- a/cli/src/data/initial-agent-type-sources.generated.ts +++ b/cli/src/data/initial-agent-type-sources.generated.ts @@ -4,8 +4,8 @@ * not depend on runtime text-import support. */ -export const agentDefinitionSource = "/**\n * Openbuff Agent Type Definitions\n *\n * This file provides TypeScript type definitions for creating custom Openbuff agents.\n * Import these types in your agent files to get full type safety and IntelliSense.\n *\n * Usage in .agents/your-agent.ts:\n * import { AgentDefinition, ToolName, ModelName } from './types/agent-definition'\n *\n * const definition: AgentDefinition = {\n * // ... your agent configuration with full type safety ...\n * }\n *\n * export default definition\n */\n\n// ============================================================================\n// Agent Definition and Utility Types\n// ============================================================================\n\nexport interface AgentDefinition {\n /** Unique identifier for this agent. Must contain only lowercase letters, numbers, and hyphens, e.g. 'code-reviewer' */\n id: string\n\n /** Version string (if not provided, will default to '0.0.1' and be bumped on each publish) */\n version?: string\n\n /** Publisher ID for the agent. Must be provided if you want to publish the agent. */\n publisher?: string\n\n /** Human-readable name for the agent */\n displayName: string\n\n /**\n * AI model to use for this agent. Can be any model in OpenRouter: https://openrouter.ai/models\n *\n * Optional: if omitted, the model is resolved entirely from the user's openbuff.json via\n * `agents[agentId]` or `defaultModel`. An error is thrown at runtime if neither is configured.\n */\n model?: ModelName\n\n /**\n * Optional wall-clock timeout in milliseconds for a single execution of this\n * agent as a subagent. When set, executeSubagent uses this as the deadline\n * (overridable per-spawn via spawn_agents' timeout_seconds). Undefined falls\n * back to the shared DEFAULT_SUBAGENT_TIMEOUT_MS, which is -1 (disabled): by\n * default there is no wall-clock timeout, so long-running agents run to\n * completion. Set a positive value to opt this agent into a wall-clock bound.\n */\n defaultTimeoutMs?: number\n\n /** Maximum subagent nesting depth. Defaults to the runtime limit. */\n maxSpawnDepth?: number\n\n /**\n * https://openrouter.ai/docs/use-cases/reasoning-tokens\n * One of `max_tokens` or `effort` is required.\n * If `exclude` is true, reasoning will be removed from the response. Default is false.\n */\n reasoningOptions?: {\n enabled?: boolean\n exclude?: boolean\n } & (\n | {\n max_tokens: number\n }\n | {\n effort: 'high' | 'medium' | 'low' | 'minimal' | 'none'\n }\n )\n\n /**\n * Provider routing options for OpenRouter.\n * Controls which providers to use and fallback behavior.\n * See https://openrouter.ai/docs/features/provider-routing\n */\n providerOptions?: {\n /**\n * List of provider slugs to try in order (e.g. [\"anthropic\", \"openai\"])\n */\n order?: string[]\n /**\n * Whether to allow backup providers when primary is unavailable (default: true)\n */\n allow_fallbacks?: boolean\n /**\n * Only use providers that support all parameters in your request (default: false)\n */\n require_parameters?: boolean\n /**\n * Control whether to use providers that may store data\n */\n data_collection?: 'allow' | 'deny'\n /**\n * List of provider slugs to allow for this request\n */\n only?: string[]\n /**\n * List of provider slugs to skip for this request\n */\n ignore?: string[]\n /**\n * List of quantization levels to filter by (e.g. [\"int4\", \"int8\"])\n */\n quantizations?: Array<\n | 'int4'\n | 'int8'\n | 'fp4'\n | 'fp6'\n | 'fp8'\n | 'fp16'\n | 'bf16'\n | 'fp32'\n | 'unknown'\n >\n /**\n * Sort providers by price, throughput, or latency\n */\n sort?: 'price' | 'throughput' | 'latency'\n /**\n * Maximum pricing you want to pay for this request\n */\n max_price?: {\n prompt?: number | string\n completion?: number | string\n image?: number | string\n audio?: number | string\n request?: number | string\n }\n }\n\n /**\n * Optional per-run cost cap in US cents. When set, the agent runtime\n * enforces this as a hard spend ceiling — the turn ends if cumulative\n * creditsUsed exceeds it. Useful for BYOK configurations to guard\n * against runaway spend. Undefined = no cap.\n */\n maxCostCents?: number\n\n /**\n * Optional per-step input token cap. When set, the agent runtime ends\n * the turn if a single step's total input tokens exceed this threshold.\n * Undefined = no cap.\n */\n maxTokensPerTurn?: number\n\n // ============================================================================\n // Tools and Subagents\n // ============================================================================\n\n /** MCP servers by name. Names cannot contain `/`. */\n mcpServers?: Record\n\n /**\n * Tools this agent can use.\n *\n * By default, all tools are available from any specified MCP server. In\n * order to limit the tools from a specific MCP server, add the tool name(s)\n * in the format `'mcpServerName/toolName1'`, `'mcpServerName/toolName2'`,\n * etc.\n */\n toolNames?: (ToolName | (string & {}))[]\n\n /** Tools callable only from `handleSteps`; these are hidden from the model. */\n programmaticToolNames?: (ToolName | (string & {}))[]\n /**\n * Controls whether every spawnable agent is exposed as a separate native\n * tool (`direct`) or only through the generic `spawn_agents` tool\n * (`generic`). Defaults to `direct` for compatibility.\n */\n spawnableAgentToolMode?: 'direct' | 'generic'\n\n /** Enforced shell capability for this agent. Defaults to workspace-write. */\n terminalPermissionProfile?:\n | 'read-only'\n | 'librarian-read-only'\n | 'git-commit'\n | 'dependency-mutation'\n | 'validation-diagnosis'\n | 'tmux-test'\n | 'workspace-write'\n | 'full-access'\n /** Runtime-enforced project-relative glob allowlists for filesystem tools. */\n filesystemScope?: {\n read?: string[]\n write?: string[]\n }\n programmaticConfig?: Record\n\n /** Other agents this agent can spawn, like 'openbuff/file-picker@0.0.1'.\n *\n * Use the fully qualified agent id from the agent store, including publisher and version, for example: 'openbuff/file-picker@0.0.1'\n * (publisher and version are required!)\n *\n * Or, use the agent id from a local agent file in your .agents directory: 'file-picker'.\n */\n spawnableAgents?: string[]\n\n // ============================================================================\n // Input and Output\n // ============================================================================\n\n /** The input schema required to spawn the agent. Provide a prompt string and/or a params object or none.\n * 80% of the time you want just a prompt string with a description:\n * inputSchema: {\n * prompt: { type: 'string', description: 'A description of what info would be helpful to the agent' }\n * }\n */\n inputSchema?: {\n prompt?: { type: 'string'; description?: string }\n params?: JsonObjectSchema\n }\n\n /** How the agent should output a response to its parent (defaults to 'last_message')\n *\n * last_message: The last message from the agent, typically after using tools.\n *\n * all_messages: All messages from the agent, including tool calls and results.\n *\n * structured_output: Make the agent output a JSON object. Can be used with outputSchema or without if you want freeform json output.\n */\n outputMode?: 'last_message' | 'all_messages' | 'structured_output'\n\n /** JSON schema for structured output (when outputMode is 'structured_output') */\n outputSchema?: JsonObjectSchema\n\n // ============================================================================\n // Prompts\n // ============================================================================\n\n /** Prompt for when and why to spawn this agent. Include the main purpose and use cases.\n *\n * This field is key if the agent is intended to be spawned by other agents. */\n spawnerPrompt?: string\n\n /** Whether to include conversation history from the parent agent in context.\n *\n * Defaults to false.\n * Use this when the agent needs to know all the previous messages in the conversation.\n */\n includeMessageHistory?: boolean\n /** Bounded parent-history transfer policy. Defaults from includeMessageHistory. */\n messageHistoryMode?: 'none' | 'pinned' | 'full'\n /** Explicit capability for inline history-editor agents. Defaults to false. */\n propagateMessageHistoryChanges?: boolean\n\n /** Whether to append model reasoning chunks to this agent's message history.\n *\n * Defaults to false for better prompt-cache stability. Enable only when an\n * agent explicitly needs its hidden reasoning replayed on later turns.\n */\n includeReasoningInMessageHistory?: boolean\n\n /** Whether to inherit the parent agent's system prompt instead of using this agent's own systemPrompt.\n *\n * Defaults to false.\n * Use this when you want to enable prompt caching by preserving the same system prompt prefix.\n * Cannot be used together with the systemPrompt field.\n */\n inheritParentSystemPrompt?: boolean\n\n /** Background information for the agent. Fairly optional. Prefer using instructionsPrompt for agent instructions. */\n systemPrompt?: string\n\n /** Instructions for the agent.\n *\n * IMPORTANT: Updating this prompt is the best way to shape the agent's behavior.\n * This prompt is inserted after each user input. */\n instructionsPrompt?: string\n\n /** Prompt inserted at each agent step.\n *\n * Powerful for changing the agent's behavior, but usually not necessary for smart models.\n * Prefer instructionsPrompt for most instructions. */\n stepPrompt?: string\n\n // ============================================================================\n // Handle Steps\n // ============================================================================\n\n /** Programmatically step the agent forward and run tools.\n *\n * You can either yield:\n * - A tool call object with toolName and input properties.\n * - 'STEP' to run agent's model and generate one assistant message.\n * - 'STEP_ALL' to run the agent's model until it uses the end_turn tool or stops includes no tool calls in a message.\n *\n * Or use 'return' to end the turn.\n *\n * Example 1:\n * function* handleSteps({ agentState, prompt, params, logger }) {\n * logger.info('Starting file read process')\n * const { toolResult } = yield {\n * toolName: 'read_files',\n * input: { paths: ['file1.txt', 'file2.txt'] }\n * }\n * yield 'STEP_ALL'\n *\n * // Optionally do a post-processing step here...\n * logger.info('Files read successfully, setting output')\n * yield {\n * toolName: 'set_output',\n * input: {\n * output: 'The files were read successfully.',\n * },\n * }\n * }\n *\n * Example 2:\n * handleSteps: function* ({ agentState, prompt, params, logger }) {\n * while (true) {\n * logger.debug('Spawning thinker agent')\n * yield {\n * toolName: 'spawn_agents',\n * input: {\n * agents: [\n * {\n * agent_type: 'thinker',\n * prompt: 'Think deeply about the user request',\n * },\n * ],\n * },\n * }\n * const { stepsComplete } = yield 'STEP'\n * if (stepsComplete) break\n * }\n * }\n */\n handleSteps?: (context: AgentStepContext) => Generator<\n ToolCall | 'STEP' | 'STEP_ALL' | StepText | GenerateN,\n void,\n {\n agentState: AgentState\n toolResult: ToolResultOutput[] | undefined\n stepsComplete: boolean\n nResponses?: string[]\n }\n >\n}\n\n// ============================================================================\n// Supporting Types\n// ============================================================================\n\nexport interface AgentState {\n agentId: string\n runId: string\n parentId: string | undefined\n\n /** The agent's conversation history: messages from the user and the assistant. */\n messageHistory: Message[]\n\n /** The last value set by the set_output tool. This is a plain object or undefined if not set. */\n output: Record | undefined\n\n /** The system prompt for this agent. */\n systemPrompt: string\n\n /** The tool definitions for this agent. */\n toolDefinitions: Record<\n string,\n { description: string | undefined; inputSchema: {} }\n >\n\n /**\n * The token count from the Anthropic API.\n * This is updated on every agent step via the /api/v1/token-count endpoint.\n */\n contextTokenCount: number\n\n /** Context window resolved from the active model/provider, when known. */\n contextWindowTokens?: number\n\n /** Runtime-owned orchestrator state preserved independently of messages. */\n base2ActiveWork?: Record\n}\n\n/**\n * Context provided to handleSteps generator function\n */\nexport interface AgentStepContext {\n agentState: AgentState\n prompt?: string\n params?: Record\n logger: Logger\n config?: Record\n}\n\nexport type StepText = { type: 'STEP_TEXT'; text: string }\nexport type GenerateN = { type: 'GENERATE_N'; n: number }\n\n/**\n * Tool call object for handleSteps generator\n */\nexport type ToolCall = {\n [K in T]: {\n toolName: K\n input: GetToolParams\n includeToolCall?: boolean\n }\n}[T]\n\n// ============================================================================\n// Available Tools\n// ============================================================================\n\n/**\n * File operation tools\n */\nexport type FileEditingTools = 'read_files' | 'write_file' | 'str_replace'\n\n/**\n * Code analysis tools\n */\nexport type CodeAnalysisTools = 'code_search' | 'find_files' | 'read_files'\n\n/**\n * Terminal and system tools\n */\nexport type TerminalTools = 'run_terminal_command' | 'code_search'\n\n/**\n * Web and browser tools\n */\nexport type WebTools = 'web_search' | 'read_docs'\n\n/**\n * Agent management tools\n */\nexport type AgentTools = 'spawn_agents'\n\n/**\n * Output and control tools\n */\nexport type OutputTools = 'set_output'\n\n// ============================================================================\n// Available Models (see: https://openrouter.ai/models)\n// ============================================================================\n\n/**\n * AI models available for agents. Pick from our selection of recommended models or choose any model in OpenRouter.\n *\n * See available models at https://openrouter.ai/models\n */\nexport type ModelName =\n // Recommended Models\n\n // OpenAI\n | 'openai/gpt-5.5'\n | 'openai/gpt-5.4'\n | 'openai/gpt-5.4-mini'\n | 'openai/gpt-5.4-nano'\n | 'openai/gpt-5.3'\n | 'openai/gpt-5.3-codex'\n | 'openai/gpt-5.2'\n | 'openai/gpt-5.2-chat-latest'\n | 'openai/gpt-5.1'\n | 'openai/gpt-5.1-chat'\n\n // Anthropic\n | 'anthropic/claude-sonnet-4.6'\n | 'anthropic/claude-opus-4.7'\n | 'anthropic/claude-opus-4.6'\n | 'anthropic/claude-opus-4.5'\n | 'anthropic/claude-haiku-4.5'\n | 'anthropic/claude-sonnet-4.5'\n | 'anthropic/claude-opus-4.1'\n\n // Gemini\n | 'google/gemini-3.1-pro-preview'\n | 'google/gemini-3-pro-preview'\n | 'google/gemini-3-flash-preview'\n | 'google/gemini-3.1-flash-lite-preview'\n | 'google/gemini-2.5-pro'\n | 'google/gemini-2.5-flash'\n | 'google/gemini-2.5-flash-lite'\n\n // X-AI\n | 'x-ai/grok-4-fast'\n | 'x-ai/grok-4.1-fast'\n | 'x-ai/grok-code-fast-1'\n\n // Qwen\n | 'qwen/qwen3-max'\n | 'qwen/qwen3-coder-plus'\n | 'qwen/qwen3-coder'\n | 'qwen/qwen3-coder:nitro'\n | 'qwen/qwen3-coder-flash'\n | 'qwen/qwen3-235b-a22b-2507'\n | 'qwen/qwen3-235b-a22b-2507:nitro'\n | 'qwen/qwen3-235b-a22b-thinking-2507'\n | 'qwen/qwen3-235b-a22b-thinking-2507:nitro'\n | 'qwen/qwen3-30b-a3b'\n | 'qwen/qwen3-30b-a3b:nitro'\n\n // DeepSeek\n | 'deepseek/deepseek-v4-pro'\n | 'deepseek-v4-pro'\n | 'deepseek/deepseek-v4-flash'\n | 'deepseek-v4-flash'\n | 'deepseek/deepseek-chat-v3-0324'\n | 'deepseek/deepseek-chat-v3-0324:nitro'\n | 'deepseek/deepseek-r1-0528'\n | 'deepseek/deepseek-r1-0528:nitro'\n\n // Other open source models\n | 'moonshotai/kimi-k2'\n | 'moonshotai/kimi-k2:nitro'\n | 'moonshotai/kimi-k2.6'\n | 'z-ai/glm-5'\n | 'z-ai/glm-5.1'\n | 'z-ai/glm-4.6'\n | 'z-ai/glm-4.6:nitro'\n | 'z-ai/glm-4.7'\n | 'z-ai/glm-4.7:nitro'\n | 'z-ai/glm-4.7-flash'\n | 'z-ai/glm-4.7-flash:nitro'\n | 'minimax/minimax-m2.5'\n | 'minimax/minimax-m2.7'\n | (string & {})\n\nimport type { ToolName, GetToolParams } from './tools'\nimport type {\n Message,\n ToolResultOutput,\n JsonObjectSchema,\n MCPConfig,\n Logger,\n} from './util-types'\n\nexport type { ToolName, GetToolParams }\n" +export const agentDefinitionSource = "/**\n * Openbuff Agent Type Definitions\n *\n * This file provides TypeScript type definitions for creating custom Openbuff agents.\n * Import these types in your agent files to get full type safety and IntelliSense.\n *\n * Usage in .agents/your-agent.ts:\n * import { AgentDefinition, ToolName, ModelName } from './types/agent-definition'\n *\n * const definition: AgentDefinition = {\n * // ... your agent configuration with full type safety ...\n * }\n *\n * export default definition\n */\n\n// ============================================================================\n// Agent Definition and Utility Types\n// ============================================================================\n\nexport interface AgentDefinition {\n /** Unique identifier for this agent. Must contain only lowercase letters, numbers, and hyphens, e.g. 'code-reviewer' */\n id: string\n\n /** Version string (if not provided, will default to '0.0.1' and be bumped on each publish) */\n version?: string\n\n /** Publisher ID for the agent. Must be provided if you want to publish the agent. */\n publisher?: string\n\n /** Human-readable name for the agent */\n displayName: string\n\n /**\n * AI model to use for this agent. Can be any model in OpenRouter: https://openrouter.ai/models\n *\n * Optional: if omitted, the model is resolved entirely from the user's openbuff.json via\n * `agents[agentId]` or `defaultModel`. An error is thrown at runtime if neither is configured.\n */\n model?: ModelName\n\n /** Maximum subagent nesting depth. Defaults to the runtime limit. */\n maxSpawnDepth?: number\n\n /**\n * https://openrouter.ai/docs/use-cases/reasoning-tokens\n * One of `max_tokens` or `effort` is required.\n * If `exclude` is true, reasoning will be removed from the response. Default is false.\n */\n reasoningOptions?: {\n enabled?: boolean\n exclude?: boolean\n } & (\n | {\n max_tokens: number\n }\n | {\n effort: 'high' | 'medium' | 'low' | 'minimal' | 'none'\n }\n )\n\n /**\n * Provider routing options for OpenRouter.\n * Controls which providers to use and fallback behavior.\n * See https://openrouter.ai/docs/features/provider-routing\n */\n providerOptions?: {\n /**\n * List of provider slugs to try in order (e.g. [\"anthropic\", \"openai\"])\n */\n order?: string[]\n /**\n * Whether to allow backup providers when primary is unavailable (default: true)\n */\n allow_fallbacks?: boolean\n /**\n * Only use providers that support all parameters in your request (default: false)\n */\n require_parameters?: boolean\n /**\n * Control whether to use providers that may store data\n */\n data_collection?: 'allow' | 'deny'\n /**\n * List of provider slugs to allow for this request\n */\n only?: string[]\n /**\n * List of provider slugs to skip for this request\n */\n ignore?: string[]\n /**\n * List of quantization levels to filter by (e.g. [\"int4\", \"int8\"])\n */\n quantizations?: Array<\n | 'int4'\n | 'int8'\n | 'fp4'\n | 'fp6'\n | 'fp8'\n | 'fp16'\n | 'bf16'\n | 'fp32'\n | 'unknown'\n >\n /**\n * Sort providers by price, throughput, or latency\n */\n sort?: 'price' | 'throughput' | 'latency'\n /**\n * Maximum pricing you want to pay for this request\n */\n max_price?: {\n prompt?: number | string\n completion?: number | string\n image?: number | string\n audio?: number | string\n request?: number | string\n }\n }\n\n /**\n * Optional per-run cost cap in US cents. When set, the agent runtime\n * enforces this as a hard spend ceiling — the turn ends if cumulative\n * creditsUsed exceeds it. Useful for BYOK configurations to guard\n * against runaway spend. Undefined = no cap.\n */\n maxCostCents?: number\n\n /**\n * Optional per-step input token cap. When set, the agent runtime ends\n * the turn if a single step's total input tokens exceed this threshold.\n * Undefined = no cap.\n */\n maxTokensPerTurn?: number\n\n // ============================================================================\n // Tools and Subagents\n // ============================================================================\n\n /** MCP servers by name. Names cannot contain `/`. */\n mcpServers?: Record\n\n /**\n * Tools this agent can use.\n *\n * By default, all tools are available from any specified MCP server. In\n * order to limit the tools from a specific MCP server, add the tool name(s)\n * in the format `'mcpServerName/toolName1'`, `'mcpServerName/toolName2'`,\n * etc.\n */\n toolNames?: (ToolName | (string & {}))[]\n\n /** Tools callable only from `handleSteps`; these are hidden from the model. */\n programmaticToolNames?: (ToolName | (string & {}))[]\n /**\n * Controls whether every spawnable agent is exposed as a separate native\n * tool (`direct`) or only through the generic `spawn_agents` tool\n * (`generic`). Defaults to `direct` for compatibility.\n */\n spawnableAgentToolMode?: 'direct' | 'generic'\n\n /** Enforced shell capability for this agent. Defaults to workspace-write. */\n terminalPermissionProfile?:\n | 'read-only'\n | 'librarian-read-only'\n | 'git-commit'\n | 'dependency-mutation'\n | 'validation-diagnosis'\n | 'tmux-test'\n | 'workspace-write'\n | 'full-access'\n /** Runtime-enforced project-relative glob allowlists for filesystem tools. */\n filesystemScope?: {\n read?: string[]\n write?: string[]\n }\n programmaticConfig?: Record\n\n /** Other agents this agent can spawn, like 'openbuff/file-picker@0.0.1'.\n *\n * Use the fully qualified agent id from the agent store, including publisher and version, for example: 'openbuff/file-picker@0.0.1'\n * (publisher and version are required!)\n *\n * Or, use the agent id from a local agent file in your .agents directory: 'file-picker'.\n */\n spawnableAgents?: string[]\n\n // ============================================================================\n // Input and Output\n // ============================================================================\n\n /** The input schema required to spawn the agent. Provide a prompt string and/or a params object or none.\n * 80% of the time you want just a prompt string with a description:\n * inputSchema: {\n * prompt: { type: 'string', description: 'A description of what info would be helpful to the agent' }\n * }\n */\n inputSchema?: {\n prompt?: { type: 'string'; description?: string }\n params?: JsonObjectSchema\n }\n\n /** How the agent should output a response to its parent (defaults to 'last_message')\n *\n * last_message: The last message from the agent, typically after using tools.\n *\n * all_messages: All messages from the agent, including tool calls and results.\n *\n * structured_output: Make the agent output a JSON object. Can be used with outputSchema or without if you want freeform json output.\n */\n outputMode?: 'last_message' | 'all_messages' | 'structured_output'\n\n /** JSON schema for structured output (when outputMode is 'structured_output') */\n outputSchema?: JsonObjectSchema\n\n // ============================================================================\n // Prompts\n // ============================================================================\n\n /** Prompt for when and why to spawn this agent. Include the main purpose and use cases.\n *\n * This field is key if the agent is intended to be spawned by other agents. */\n spawnerPrompt?: string\n\n /** Whether to include conversation history from the parent agent in context.\n *\n * Defaults to false.\n * Use this when the agent needs to know all the previous messages in the conversation.\n */\n includeMessageHistory?: boolean\n /** Bounded parent-history transfer policy. Defaults from includeMessageHistory. */\n messageHistoryMode?: 'none' | 'pinned' | 'full'\n /** Explicit capability for inline history-editor agents. Defaults to false. */\n propagateMessageHistoryChanges?: boolean\n\n /** Whether to append model reasoning chunks to this agent's message history.\n *\n * Defaults to false for better prompt-cache stability. Enable only when an\n * agent explicitly needs its hidden reasoning replayed on later turns.\n */\n includeReasoningInMessageHistory?: boolean\n\n /** Whether to inherit the parent agent's system prompt instead of using this agent's own systemPrompt.\n *\n * Defaults to false.\n * Use this when you want to enable prompt caching by preserving the same system prompt prefix.\n * Cannot be used together with the systemPrompt field.\n */\n inheritParentSystemPrompt?: boolean\n\n /** Background information for the agent. Fairly optional. Prefer using instructionsPrompt for agent instructions. */\n systemPrompt?: string\n\n /** Instructions for the agent.\n *\n * IMPORTANT: Updating this prompt is the best way to shape the agent's behavior.\n * This prompt is inserted after each user input. */\n instructionsPrompt?: string\n\n /** Prompt inserted at each agent step.\n *\n * Powerful for changing the agent's behavior, but usually not necessary for smart models.\n * Prefer instructionsPrompt for most instructions. */\n stepPrompt?: string\n\n // ============================================================================\n // Handle Steps\n // ============================================================================\n\n /** Programmatically step the agent forward and run tools.\n *\n * You can either yield:\n * - A tool call object with toolName and input properties.\n * - 'STEP' to run agent's model and generate one assistant message.\n * - 'STEP_ALL' to run the agent's model until it uses the end_turn tool or stops includes no tool calls in a message.\n *\n * Or use 'return' to end the turn.\n *\n * Example 1:\n * function* handleSteps({ agentState, prompt, params, logger }) {\n * logger.info('Starting file read process')\n * const { toolResult } = yield {\n * toolName: 'read_files',\n * input: { paths: ['file1.txt', 'file2.txt'] }\n * }\n * yield 'STEP_ALL'\n *\n * // Optionally do a post-processing step here...\n * logger.info('Files read successfully, setting output')\n * yield {\n * toolName: 'set_output',\n * input: {\n * output: 'The files were read successfully.',\n * },\n * }\n * }\n *\n * Example 2:\n * handleSteps: function* ({ agentState, prompt, params, logger }) {\n * while (true) {\n * logger.debug('Spawning thinker agent')\n * yield {\n * toolName: 'spawn_agents',\n * input: {\n * agents: [\n * {\n * agent_type: 'thinker',\n * prompt: 'Think deeply about the user request',\n * },\n * ],\n * },\n * }\n * const { stepsComplete } = yield 'STEP'\n * if (stepsComplete) break\n * }\n * }\n */\n handleSteps?: (context: AgentStepContext) => Generator<\n ToolCall | 'STEP' | 'STEP_ALL' | StepText | GenerateN,\n void,\n {\n agentState: AgentState\n toolResult: ToolResultOutput[] | undefined\n stepsComplete: boolean\n nResponses?: string[]\n }\n >\n}\n\n// ============================================================================\n// Supporting Types\n// ============================================================================\n\nexport interface AgentState {\n agentId: string\n runId: string\n parentId: string | undefined\n\n /** The agent's conversation history: messages from the user and the assistant. */\n messageHistory: Message[]\n\n /** The last value set by the set_output tool. This is a plain object or undefined if not set. */\n output: Record | undefined\n\n /** The system prompt for this agent. */\n systemPrompt: string\n\n /** The tool definitions for this agent. */\n toolDefinitions: Record<\n string,\n { description: string | undefined; inputSchema: {} }\n >\n\n /**\n * The token count from the Anthropic API.\n * This is updated on every agent step via the /api/v1/token-count endpoint.\n */\n contextTokenCount: number\n\n /** Context window resolved from the active model/provider, when known. */\n contextWindowTokens?: number\n\n /** Runtime-owned orchestrator state preserved independently of messages. */\n base2ActiveWork?: Record\n}\n\n/**\n * Context provided to handleSteps generator function\n */\nexport interface AgentStepContext {\n agentState: AgentState\n prompt?: string\n params?: Record\n logger: Logger\n config?: Record\n}\n\nexport type StepText = { type: 'STEP_TEXT'; text: string }\nexport type GenerateN = { type: 'GENERATE_N'; n: number }\n\n/**\n * Tool call object for handleSteps generator\n */\nexport type ToolCall = {\n [K in T]: {\n toolName: K\n input: GetToolParams\n includeToolCall?: boolean\n }\n}[T]\n\n// ============================================================================\n// Available Tools\n// ============================================================================\n\n/**\n * File operation tools\n */\nexport type FileEditingTools = 'read_files' | 'write_file' | 'str_replace'\n\n/**\n * Code analysis tools\n */\nexport type CodeAnalysisTools = 'code_search' | 'find_files' | 'read_files'\n\n/**\n * Terminal and system tools\n */\nexport type TerminalTools = 'run_terminal_command' | 'code_search'\n\n/**\n * Web and browser tools\n */\nexport type WebTools = 'web_search' | 'read_docs'\n\n/**\n * Agent management tools\n */\nexport type AgentTools = 'spawn_agents'\n\n/**\n * Output and control tools\n */\nexport type OutputTools = 'set_output'\n\n// ============================================================================\n// Available Models (see: https://openrouter.ai/models)\n// ============================================================================\n\n/**\n * AI models available for agents. Pick from our selection of recommended models or choose any model in OpenRouter.\n *\n * See available models at https://openrouter.ai/models\n */\nexport type ModelName =\n // Recommended Models\n\n // OpenAI\n | 'openai/gpt-5.5'\n | 'openai/gpt-5.4'\n | 'openai/gpt-5.4-mini'\n | 'openai/gpt-5.4-nano'\n | 'openai/gpt-5.3'\n | 'openai/gpt-5.3-codex'\n | 'openai/gpt-5.2'\n | 'openai/gpt-5.2-chat-latest'\n | 'openai/gpt-5.1'\n | 'openai/gpt-5.1-chat'\n\n // Anthropic\n | 'anthropic/claude-sonnet-4.6'\n | 'anthropic/claude-opus-4.7'\n | 'anthropic/claude-opus-4.6'\n | 'anthropic/claude-opus-4.5'\n | 'anthropic/claude-haiku-4.5'\n | 'anthropic/claude-sonnet-4.5'\n | 'anthropic/claude-opus-4.1'\n\n // Gemini\n | 'google/gemini-3.1-pro-preview'\n | 'google/gemini-3-pro-preview'\n | 'google/gemini-3-flash-preview'\n | 'google/gemini-3.1-flash-lite-preview'\n | 'google/gemini-2.5-pro'\n | 'google/gemini-2.5-flash'\n | 'google/gemini-2.5-flash-lite'\n\n // X-AI\n | 'x-ai/grok-4-fast'\n | 'x-ai/grok-4.1-fast'\n | 'x-ai/grok-code-fast-1'\n\n // Qwen\n | 'qwen/qwen3-max'\n | 'qwen/qwen3-coder-plus'\n | 'qwen/qwen3-coder'\n | 'qwen/qwen3-coder:nitro'\n | 'qwen/qwen3-coder-flash'\n | 'qwen/qwen3-235b-a22b-2507'\n | 'qwen/qwen3-235b-a22b-2507:nitro'\n | 'qwen/qwen3-235b-a22b-thinking-2507'\n | 'qwen/qwen3-235b-a22b-thinking-2507:nitro'\n | 'qwen/qwen3-30b-a3b'\n | 'qwen/qwen3-30b-a3b:nitro'\n\n // DeepSeek\n | 'deepseek/deepseek-v4-pro'\n | 'deepseek-v4-pro'\n | 'deepseek/deepseek-v4-flash'\n | 'deepseek-v4-flash'\n | 'deepseek/deepseek-chat-v3-0324'\n | 'deepseek/deepseek-chat-v3-0324:nitro'\n | 'deepseek/deepseek-r1-0528'\n | 'deepseek/deepseek-r1-0528:nitro'\n\n // Other open source models\n | 'moonshotai/kimi-k2'\n | 'moonshotai/kimi-k2:nitro'\n | 'moonshotai/kimi-k2.6'\n | 'z-ai/glm-5'\n | 'z-ai/glm-5.1'\n | 'z-ai/glm-4.6'\n | 'z-ai/glm-4.6:nitro'\n | 'z-ai/glm-4.7'\n | 'z-ai/glm-4.7:nitro'\n | 'z-ai/glm-4.7-flash'\n | 'z-ai/glm-4.7-flash:nitro'\n | 'minimax/minimax-m2.5'\n | 'minimax/minimax-m2.7'\n | (string & {})\n\nimport type { ToolName, GetToolParams } from './tools'\nimport type {\n Message,\n ToolResultOutput,\n JsonObjectSchema,\n MCPConfig,\n Logger,\n} from './util-types'\n\nexport type { ToolName, GetToolParams }\n" -export const toolsSource = "/**\n * Union type of all available tool names\n */\nexport type ToolName =\n | 'add_message'\n | 'ask_user'\n | 'check_background_agent'\n | 'check_job'\n | 'code_search'\n | 'end_turn'\n | 'edit_transaction'\n | 'edit_3d_asset'\n | 'find_files'\n | 'find_files_matching_content'\n | 'git_status'\n | 'git_branch'\n | 'get_task'\n | 'get_change_review_bundle'\n | 'inspect_workspace'\n | 'inspect_environment'\n | 'inspect_3d_asset'\n | 'get_affected_tests'\n | 'get_build_targets'\n | 'inspect_codebase_structure'\n | 'inspect_feature_completeness'\n | 'evaluate_audit_coverage'\n | 'glob'\n | 'kill_job'\n | 'list_directory'\n | 'list_jobs'\n | 'lookup_agent_info'\n | 'query_index'\n | 'read_docs'\n | 'read_files'\n | 'read_image'\n | 'render_3d_preview'\n | 'read_logs'\n | 'read_outline'\n | 'read_subtree'\n | 'replace_range'\n | 'rewrite_symbol'\n | 'render_ui'\n | 'run_file_change_hooks'\n | 'run_targeted_validation'\n | 'run_terminal_command'\n | 'set_messages'\n | 'set_output'\n | 'skill'\n | 'spawn_agents'\n | 'str_replace'\n | 'suggest_followups'\n | 'task_completed'\n | 'think_deeply'\n | 'update_plan_status'\n | 'web_search'\n | 'write_file'\n | 'write_audit_findings'\n | 'write_todos'\n\n/**\n * Map of tool names to their parameter types\n */\nexport interface ToolParamsMap {\n add_message: AddMessageParams\n ask_user: AskUserParams\n check_background_agent: CheckBackgroundAgentParams\n check_job: CheckJobParams\n code_search: CodeSearchParams\n end_turn: EndTurnParams\n edit_transaction: EditTransactionParams\n edit_3d_asset: Edit3dAssetParams\n find_files: FindFilesParams\n find_files_matching_content: FindFilesMatchingContentParams\n git_status: GitStatusParams\n git_branch: GitBranchParams\n get_task: GetTaskParams\n get_change_review_bundle: GetChangeReviewBundleParams\n inspect_workspace: InspectWorkspaceParams\n inspect_environment: InspectEnvironmentParams\n inspect_3d_asset: Inspect3dAssetParams\n get_affected_tests: GetAffectedTestsParams\n get_build_targets: GetBuildTargetsParams\n inspect_codebase_structure: InspectCodebaseStructureParams\n inspect_feature_completeness: InspectFeatureCompletenessParams\n evaluate_audit_coverage: EvaluateAuditCoverageParams\n glob: GlobParams\n kill_job: KillJobParams\n list_directory: ListDirectoryParams\n list_jobs: ListJobsParams\n lookup_agent_info: LookupAgentInfoParams\n query_index: QueryIndexParams\n read_docs: ReadDocsParams\n read_files: ReadFilesParams\n read_image: ReadImageParams\n render_3d_preview: Render3dPreviewParams\n read_logs: ReadLogsParams\n read_outline: ReadOutlineParams\n read_subtree: ReadSubtreeParams\n replace_range: ReplaceRangeParams\n rewrite_symbol: RewriteSymbolParams\n render_ui: RenderUiParams\n run_file_change_hooks: RunFileChangeHooksParams\n run_targeted_validation: RunTargetedValidationParams\n run_terminal_command: RunTerminalCommandParams\n set_messages: SetMessagesParams\n set_output: SetOutputParams\n skill: SkillParams\n spawn_agents: SpawnAgentsParams\n str_replace: StrReplaceParams\n suggest_followups: SuggestFollowupsParams\n task_completed: TaskCompletedParams\n think_deeply: ThinkDeeplyParams\n update_plan_status: UpdatePlanStatusParams\n web_search: WebSearchParams\n write_file: WriteFileParams\n write_audit_findings: WriteAuditFindingsParams\n write_todos: WriteTodosParams\n}\n\n/**\n * Add a new message to the conversation history. To be used for complex requests that can't be solved in a single step, as you may forget what happened!\n */\nexport interface AddMessageParams {\n role: 'user' | 'assistant'\n content: string\n}\n\n/**\n * Ask the user a list of multiple choice questions. Each question must have at least 2 options. The agent execution will pause until the user submits their answers.\n */\nexport interface AskUserParams {\n /** List of multiple choice questions to ask the user */\n questions: {\n /** The question to ask the user */\n question: string\n /** Optional short display label. Values longer than 18 Unicode code points are truncated instead of rejecting the question. */\n header?: string\n /** Array of answer options with label and optional description. */\n options: {\n /** The display text for this option */\n label: string\n /** Explanation shown when option is focused */\n description?: string\n }[]\n /** If true, allows selecting multiple options (checkbox). If false, single selection only (radio). */\n multiSelect?: boolean\n /** Validation rules for \"Other\" text input */\n validation?: {\n /** Maximum length for \"Other\" text input */\n maxLength?: number\n /** Minimum length for \"Other\" text input */\n minLength?: number\n /** Regex pattern for \"Other\" text input */\n pattern?: string\n /** Custom error message when pattern fails */\n patternError?: string\n }\n }[]\n}\n\n/**\n * Join/wait on a background agent turn started by spawn_agents({ background: true }): returns the sequenced agent_chunk events produced since the cursor plus the unified job state. Use it to observe a long-running background agent without blocking the turn.\n */\nexport interface CheckBackgroundAgentParams {\n /** The jobId returned by spawn_agents({ background: true }) for the background agent turn. */\n jobId: string\n /** Optional sequence cursor from a prior response. Polling is idempotent for an explicit cursor; nextCursor can be supplied on the next call. */\n cursor?: number\n /** Optional substring to wait for in the new streamed chunks before returning (follow mode). Returns early as soon as it appears in any chunk payload. Useful for waiting until a background agent emits a specific milestone (e.g. a tool_result or a text marker). */\n wait_for?: string\n /** Max seconds to wait for new chunks / the wait_for pattern. 0 (default) returns immediately with whatever new chunks exist (poll mode); >0 blocks up to this long (follow mode). */\n timeout_seconds?: number\n /** When true, explicitly cancel the running background agent before returning its final status. Defaults to false. */\n cancel?: boolean\n}\n\n/**\n * Join/wait on a background job started by run_terminal_command: returns the sequenced output events produced since the last check plus the unified job state and exit code. Use it to observe a long-running process without blocking the turn. To watch an arbitrary log file, start a `tail -f ` BACKGROUND job and check_job it with a wait_for pattern.\n */\nexport interface CheckJobParams {\n /** The jobId returned by run_terminal_command with process_type: BACKGROUND. */\n jobId: string\n /** Optional substring to wait for in the new output before returning (follow mode). Returns early as soon as it appears (e.g. \"Listening on\" / \"compiled successfully\"). */\n wait_for?: string\n /** Max seconds to wait for new output / the wait_for pattern. 0 (default) returns immediately with whatever new output exists (poll mode); >0 blocks up to this long (follow mode). */\n timeout_seconds?: number\n /** Follow mode only: SIGTERM the job on follow-timeout. Poll mode never kills. Default false. */\n kill_on_timeout?: boolean\n}\n\n/**\n * Search for string patterns in the project's files. This tool uses ripgrep (rg), a fast line-oriented search tool. Use this tool only when read_files is not sufficient to find the files you need.\n */\nexport interface CodeSearchParams {\n /** The pattern to search for. */\n pattern: string\n /** Optional safe ripgrep flags as one string or argv tokens (e.g., \"-i -g *.ts -A 2\" or [\"-i\", \"-g\", \"*.ts\", \"-A\", \"2\"]). Allowed: -i/--ignore-case, -S/--smart-case, -s/--case-sensitive, -w/--word-regexp, -F/--fixed-strings, -U/--multiline, --multiline-dotall, -g/--glob, -t/--type, -T/--type-not, plus context -A/-B/-C (and long forms). JSON quotes delimit the string; do not embed another quote pair around the entire expression. Line numbers are automatic; -n/--line-number are ignored. Output-shape flags such as -c/--count, --count-matches, -l, -v/--invert-match, -r/--replace, --exec, and -z/--null are rejected. */\n flags?: string | string[]\n /** Optional working directory or single file to search within, relative to the project root or absolute. Absolute paths may be outside the project. A directory becomes ripgrep's cwd and scopes the search under that path (plus existing blessed hidden dirs when no paths are given); a file scopes the search to that file only (process cwd = project root when the file is under the project, else the file's parent). Defaults to searching the entire project root. */\n cwd?: string\n /** Optional list of file and/or directory paths to search (relative to the project root, or absolute). When non-empty, ripgrep searches only these targets instead of the whole cwd tree (and does not auto-expand hidden dirs). Can be combined with a file cwd. */\n paths?: string[]\n /** Maximum number of results to return per file. Defaults to 15. There is also a global limit of 250 results across all files. */\n maxResults?: number\n}\n\n/**\n * End your turn, regardless of any new tool results that might be coming. This will allow the user to type another prompt.\n */\nexport interface EndTurnParams {}\n\n/**\n * Parameters for edit_transaction tool\n */\nexport interface EditTransactionParams {\n edits: (\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'str_replace'\n replacements: {\n oldString: string\n newString: string\n allowMultiple?: boolean\n occurrenceIndex?: number\n /** Optional authenticated cap.v3 readCapability copied verbatim from the matching fresh read_files editAnchor. */\n basedOnRead?: string\n /** For deletion replacements only (newString is empty): treat a missing oldString as an already-applied no-op. Use only for explicit idempotent cleanup retries, never for ordinary edits. When every requested change resolves to such a no-op - every replacement of a standalone str_replace call, or every edit of an edit_transaction - the call succeeds with zero file changes and the skip messages rather than failing. When combined with occurrenceIndex, a partially-applied cleanup also skips: fewer remaining exact occurrences than the requested index means that occurrence is treated as already applied. Only valid when newString is empty; both the input and provider schemas reject any other combination. */\n skipIfMissing?: boolean\n }[]\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n /** A structured edit dispatched by operation kind. */\n type: 'structured'\n /** Structured edit operation to apply to this file. */\n operation:\n | {\n /** Deterministic text insertion. */\n kind: 'insert_text'\n /** 1-indexed insertion position. */\n position: {\n /** 1-indexed target line. */\n line: number\n /** 1-indexed target column. */\n column: number\n }\n text: string\n }\n | {\n /** Language-aware import insertion. */\n kind: 'insert_import'\n /** Complete language-native import statement to add, e.g. \"import { foo } from 'bar'\", \"from app import value\", or \"use crate::value\". */\n importStatement: string\n }\n | {\n /** Language-aware import removal. */\n kind: 'remove_import'\n /** Complete language-native import statement to remove. Required unless moduleSpecifier is provided. */\n importStatement?: string\n /** Module specifier to remove imports from, e.g. \"react\" or \"./helper\". */\n moduleSpecifier?: string\n }\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'create'\n /** Exact bytes to write to the new file. */\n content: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'delete'\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'move'\n /** New project-relative path. The destination must be absent. */\n destinationPath: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'replace_range'\n readCapability: string\n startLine?: number\n endLine?: number\n occurrence?: {\n match: string\n occurrence?: number\n }\n newContent: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'rewrite_symbol'\n symbol: string\n content: string\n occurrence?: number\n /** Optional cap.v3 copied from the matching read_files symbol slice. It authorizes exactly the symbol and its contiguous preceding comment block. */\n readCapability?: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'patch'\n diff: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'write_file'\n content: string\n /** Optional whole-file-covering cap.v3 from a fresh complete whole-file read. Only a full-file capability with a hash matching current content may authorize overwrite; partial ranges never authorize write_file. */\n basedOnRead?: string\n }\n )[]\n}\n\n/**\n * Parameters for edit_3d_asset tool\n */\nexport interface Edit3dAssetParams {\n /** Project-relative .blend path. */\n path: string\n /** Exact source hash returned by inspect_3d_asset. */\n source_hash: string\n operations: (\n | {\n type: 'rename_object'\n object: string\n new_name: string\n }\n | {\n type: 'set_object_transform'\n object: string\n location?: any[]\n rotation_degrees?: any[]\n scale?: any[]\n }\n | {\n type: 'set_render_resolution'\n width: number\n height: number\n percentage?: number\n }\n | {\n type: 'set_frame_range'\n start: number\n end: number\n }\n )[]\n}\n\n/**\n * Find several files related to a brief natural language description of the files or the name of a function or class you are looking for.\n */\nexport interface FindFilesParams {\n /** A brief natural language description of the files or the name of a function or class you are looking for. It's also helpful to mention a directory or two to look within. */\n prompt: string\n}\n\n/**\n * List unique file paths whose content matches a pattern, with optional symbol grouping. Built on top of ripgrep (rg).\n */\nexport interface FindFilesMatchingContentParams {\n /** Regex pattern (ripgrep syntax) to match file content against. */\n pattern: string\n /** Optional safe ripgrep flags as one string or argv tokens. Allowed: -i/--ignore-case, -S/--smart-case, -s/--case-sensitive, -w/--word-regexp, -F/--fixed-strings, -U/--multiline, --multiline-dotall, -g/--glob, -t/--type, -T/--type-not. Examples: \"-g *.ts -g *.tsx\" or [\"-g\", \"*.ts\", \"-g\", \"*.tsx\"]. Do not quote the entire expression inside the JSON string. Output-shape flags such as -c/--count, --count-matches, -l, -v/--invert-match, context -A/-B/-C, -r/--replace, --exec, and -z/--null are rejected (this tool forces -l or --json itself). Redundant -n/--line-number inputs are ignored. */\n flags?: string | string[]\n /** Optional working directory or single file to search within, relative to the project root or absolute. Absolute paths may be outside the project. A directory becomes ripgrep's cwd and scopes the search under that path (plus existing blessed hidden dirs); a file scopes the search to that file only (process cwd = project root when the file is under the project, else the file's parent). Defaults to the project root. */\n cwd?: string\n /** Maximum number of unique files to return. Defaults to 100. */\n maxFiles?: number\n /** When true, also return the names of the top-level symbols (functions, classes, methods, exports, constants) that contain each match, plus the per-file match count. Symbol extraction is heuristic and works best for JS/TS/Python/Go/Rust source files; languages without a recognized declaration shape produce an empty symbols list. */\n groupBySymbol?: boolean\n /** Maximum seconds to let ripgrep run before returning partial results. Defaults to 15. */\n timeoutSeconds?: number\n}\n\n/**\n * Read-only git status and (optionally) diff for the current project.\n */\nexport interface GitStatusParams {\n /** When true, also return the unified diff of uncommitted changes. */\n include_diff?: boolean\n /** When true with include_diff, returns the staged diff instead of unstaged. */\n staged?: boolean\n /** Optional path to scope status/diff to (relative to project root). */\n path?: string\n /** Maximum characters of diff output to return. Defaults to 40,000. */\n max_chars?: number\n}\n\n/**\n * Create a new git branch, optionally switching to it. Refuses to branch when the working tree is dirty unless `allow_dirty` is true.\n */\nexport interface GitBranchParams {\n /** Name of the branch to create. Must start with an alphanumeric character and contain only [a-zA-Z0-9._/-]. */\n branch_name: string\n /** When true (default), create AND switch to the branch (`git checkout -b`). When false, only create the branch (`git branch`), leaving the current branch checked out. */\n switch?: boolean\n /** When true, skip the dirty-tree refusal check. Defaults to false — the tool refuses to branch when the working tree has uncommitted changes. */\n allow_dirty?: boolean\n}\n\n/**\n * Parameters for get_task tool\n */\nexport interface GetTaskParams {\n /** Optional plan session slug. Defaults to .agents/ACTIVE_SESSION. */\n session?: string\n}\n\n/**\n * Parameters for get_change_review_bundle tool\n */\nexport interface GetChangeReviewBundleParams {\n max_chars?: number\n}\n\n/**\n * Inspect the current repository/worktree identity and Git state without modifying it.\n */\nexport interface InspectWorkspaceParams {}\n\n/**\n * Parameters for inspect_environment tool\n */\nexport interface InspectEnvironmentParams {}\n\n/**\n * Parameters for inspect_3d_asset tool\n */\nexport interface Inspect3dAssetParams {\n /** Project-relative 3D asset path. */\n path: string\n}\n\n/**\n * Parameters for get_affected_tests tool\n */\nexport interface GetAffectedTestsParams {\n files: string[]\n}\n\n/**\n * Parameters for get_build_targets tool\n */\nexport interface GetBuildTargetsParams {\n files: string[]\n}\n\n/**\n * Parameters for inspect_codebase_structure tool\n */\nexport interface InspectCodebaseStructureParams {\n scope?: string[]\n}\n\n/**\n * Parameters for inspect_feature_completeness tool\n */\nexport interface InspectFeatureCompletenessParams {\n feature: string\n snapshot_id: string\n scope?: string[]\n}\n\n/**\n * Parameters for evaluate_audit_coverage tool\n */\nexport interface EvaluateAuditCoverageParams {\n snapshot_id: string\n structural_receipts: {\n schema_version: 1\n snapshot_id: string\n shard_id: string\n subsystem_ids: string[]\n files: string[]\n domains: (\n | 'security'\n | 'correctness'\n | 'state-mutation'\n | 'error-handling'\n | 'performance'\n | 'dependency-hygiene'\n | 'test-coverage'\n | 'api-contract'\n )[]\n }[]\n features: {\n schema_version: 1\n snapshot_id: string\n feature: string\n evidence_kind: 'heuristic' | 'verified'\n evidence: {\n entrypoints: string[]\n implementation: string[]\n consumers: string[]\n tests: string[]\n docs: string[]\n failure_states: string[]\n }\n }[]\n out_of_scope?: {\n id: string\n reason: string\n }[]\n scope?: string[]\n}\n\n/**\n * Search for files matching a glob pattern. Returns matching file paths sorted by modification time (newest first, then path for deterministic ties).\n */\nexport interface GlobParams {\n /** Glob pattern to match files against (e.g., *.js, src/glob/*.ts, glob/test/glob/*.go). */\n pattern: string\n /** Optional working directory or file path, relative to project root. If a directory, the glob pattern is matched against paths relative to this cwd, while returned files remain project-relative. If a file path, the pattern is matched against that file only (full path or basename). If not provided, searches from project root. */\n cwd?: string\n}\n\n/**\n * Cancel a background job started by run_terminal_command.\n */\nexport interface KillJobParams {\n /** The jobId returned by run_terminal_command with process_type: BACKGROUND. */\n jobId: string\n /** Signal to send. Defaults to SIGTERM; use SIGKILL only if graceful termination fails. */\n signal?: 'SIGTERM' | 'SIGKILL'\n}\n\n/**\n * List files and directories in the specified path. Returns separate arrays of file names and directory names.\n */\nexport interface ListDirectoryParams {\n /** Directory path to list, relative to the project root. */\n path: string\n}\n\n/**\n * List this run's background jobs (shell processes and background agents, running and settled) with statuses, bucketed pending process/log output relative to the last check_job consumer cursor (agents usually show pending: 'none'), and a gap flag.\n */\nexport interface ListJobsParams {}\n\n/**\n * Retrieve information about an agent by ID\n */\nexport interface LookupAgentInfoParams {\n /** Agent ID (short local or full published format) */\n agentId: string\n}\n\n/**\n * Query the local codebase graph index to find relevant files ranked by symbol names, imports, headings, paths, doc concepts, and graph relationships. The index is built automatically on startup.\n */\nexport interface QueryIndexParams {\n /** Natural language query or keyword terms describing the files you are looking for. Optional for graph modes when from/to paths are provided. For example: \"authentication\", \"database migrations\", \"editor mutation logic\", \"React components\". */\n query?: string\n /** Maximum number of results to return. Defaults to 20. */\n limit?: number\n /** Optional list of file extensions to filter results (without dot). E.g. [\"ts\", \"tsx\"] for TypeScript only. */\n fileTypes?: string[]\n /** Optional normalized project-relative directory prefixes. Results outside every prefix are excluded before ranking/limiting. */\n pathPrefixes?: string[]\n /** search|explain|neighbors|path|commands|references — see tool description. */\n mode?: 'search' | 'neighbors' | 'path' | 'explain' | 'commands' | 'references'\n /** Optional source file path for neighbors, path, and references modes. */\n from?: string\n /** Optional target file path for path mode. Also used as the seed file for references mode when from is omitted or not indexed. */\n to?: string\n}\n\n/**\n * Fetch up-to-date documentation for libraries and frameworks using Context7 API.\n */\nexport interface ReadDocsParams {\n /** The library or framework name (e.g., \"Next.js\", \"MongoDB\", \"React\"). Use the official name as it appears in documentation if possible. Only public libraries available in Context7's database are supported, so small or private libraries may not be available. */\n libraryTitle: string\n /** Specific topic to focus on (e.g., \"routing\", \"hooks\", \"authentication\") */\n topic: string\n /** Optional maximum number of tokens to return. Defaults to 10000. Values less than 10000 are automatically increased to 10000. */\n max_tokens?: number\n}\n\n/**\n * Read multiple files from disk and return their contents. Use this tool to read as many files as would be helpful to answer the user's request.\n */\nexport interface ReadFilesParams {\n /** Whole-file paths to read. Complete results include editAnchor.readCapability for follow-up edits. */\n paths?: string[]\n /** 1-indexed inclusive line ranges. Sole `paths` entry infers missing path. */\n ranges?: {\n /** Project-relative file path. */\n path: string\n /** 1-indexed inclusive start line. Defaults to 1. */\n startLine?: number\n /** 1-indexed inclusive end line. Defaults to the last line. */\n endLine?: number\n }[]\n /** Contiguous line windows; each complete window mints a scoped cap.v3 editAnchor. */\n windows?: {\n /** File path to read in contiguous line windows, relative to the project root. */\n path: string\n /** Lines per window. Defaults to 400, capped at 5000. */\n windowSize?: number\n /** 1-indexed window number to return. Omit to get the window manifest (totalLines, windowSize, windowCount) plus the first window. */\n window?: number\n }[]\n /** Literal-anchored context blocks with a scoped cap.v3 editAnchor per block. */\n around?: {\n /** File path to read a content-anchored block from, relative to the project root. */\n path: string\n /** Exact literal string to anchor on. Robust to line-number drift. */\n match: string\n /** 1-indexed occurrence of `match` to anchor on. Defaults to 1. */\n occurrence?: number\n /** Lines of context to include on each side of the match, clamped at file boundaries. Defaults to 40, capped at 2000. */\n contextLines?: number\n }[]\n /** Nth top-level symbol by name (rewrite_symbol occurrence semantics); prefer batch `symbols` when possible. */\n symbol?: {\n /** File path to extract a symbol slice from, relative to the project root. */\n path: string\n /** Top-level symbol name (function, class, interface, method) to pull, as shown by read_outline. */\n name: string\n /** When multiple top-level symbols share this name, the 1-indexed one to return. Defaults to 1. Matches rewrite_symbol occurrence semantics. */\n occurrence?: number\n }[]\n /** Named symbol slices with editAnchors; prefer over full reads when names are known. */\n symbols?: {\n /** Project-relative file path. */\n path: string\n /** Symbol names to slice. */\n names: string[]\n }[]\n}\n\n/**\n * Read image files from disk and return them as model-visible image media.\n */\nexport interface ReadImageParams {\n /** List of image file paths to read. */\n paths: string[]\n}\n\n/**\n * Parameters for render_3d_preview tool\n */\nexport interface Render3dPreviewParams {\n /** Project-relative 3D asset path. */\n path: string\n views?: ('camera' | 'perspective' | 'front' | 'side' | 'top')[]\n mode?: 'material' | 'clay' | 'wireframe'\n width?: number\n height?: number\n}\n\n/**\n * Read the last N lines from a log/text file or background job log without starting a background tail process.\n */\nexport interface ReadLogsParams {\n /** Path to the log file, relative to the project root unless absolute. Required unless jobId is provided. */\n path?: string\n /** Background job id returned by run_terminal_command(process_type: BACKGROUND). When provided, reads the job log file directly. */\n jobId?: string\n /** Number of trailing lines to read. Defaults to 200. */\n lines?: number\n /** Maximum characters to return. Defaults to 20,000. */\n max_chars?: number\n}\n\n/**\n * Generate an outline of imports, exports, classes, methods, and function signatures in a source file without reading the entire implementation.\n */\nexport interface ReadOutlineParams {\n /** File path to generate the AST-like outline for, relative to the project root. */\n path: string\n}\n\n/**\n * Read one or more directory subtrees (as a blob including subdirectories, file names, and parsed variables within each source file) or return parsed variable names for files. If no paths are provided, returns the entire project tree.\n */\nexport interface ReadSubtreeParams {\n /** List of paths to directories or files. Relative to the project root. If omitted, the entire project tree is used. */\n paths?: string[]\n /** Maximum token budget for the subtree blob; the tree will be truncated to fit within this budget by first dropping file variables and then removing the most-nested files and directories. */\n maxTokens?: number\n}\n\n/**\n * Replace all of, a contained sub-range of, or the Nth literal occurrence inside content observed through one fresh cap.v3 read capability.\n */\nexport interface ReplaceRangeParams {\n /** The path to the file to edit. */\n path: string\n /** Copy the cap.v3 readCapability verbatim from the matching fresh read_files editAnchor. The token supplies the observed line bounds and content hash. */\n readCapability: string\n /** Optional 1-indexed target start within the capability-covered range. Omit with endLine to replace the complete observed range. */\n startLine?: number\n /** Optional 1-indexed target end within the capability-covered range. Omit with startLine to replace the complete observed range. */\n endLine?: number\n /** Optional occurrence targeting: replace the 1-indexed occurrence (default 1) of the exact literal match found inside the capability-authorized range. Mutually exclusive with startLine/endLine. */\n occurrence?: {\n match: string\n occurrence?: number\n }\n /** Complete replacement content for the selected line range. */\n newContent: string\n}\n\n/**\n * Replace a whole symbol's definition by name using the file's syntax tree, without copying its current text. Resolves the exact AST range and applies it through the safe str_replace path (atomic, anchored).\n */\nexport interface RewriteSymbolParams {\n /** File path containing the symbol, relative to the project root. */\n path: string\n /** Name of the function/class/method/type/interface to replace (as shown by read_outline). */\n symbol: string\n /** The complete new source for the symbol, replacing its entire current definition (e.g. the whole function including its signature and body). Provide REAL newlines/tabs in the string — literal backslash-n (\\n) and backslash-t (\\t) sequences are not interpreted and will be written verbatim into the file. This matches str_replace. */\n content: string\n /** When multiple top-level symbols share this name, the 1-indexed one to replace. */\n occurrence?: number\n /** Optional cap.v3 copied from the matching read_files symbol slice. Under strict read-before-edit this authorizes exactly the symbol and its contiguous preceding comment block. */\n readCapability?: string\n}\n\n/**\n * Render a small interactive UI widget in the Openbuff CLI. Currently supports a button that opens a link.\n */\nexport interface RenderUiParams {\n /** The UI widget to render. */\n widget: {\n /** Widget type. Currently, the only supported widget is button. */\n type: 'button'\n /** Short button label shown to the user. */\n text: string\n /** The http:// or https:// URL to open when the user clicks the button. */\n link: string\n /** Theme-aware color treatment. Use primary for the main action and secondary for lower-emphasis actions. */\n variant?: 'primary' | 'secondary'\n }\n}\n\n/**\n * Parameters for run_file_change_hooks tool\n */\nexport interface RunFileChangeHooksParams {\n /** List of file paths that were changed and should trigger file change hooks */\n files: string[]\n}\n\n/**\n * Parameters for run_targeted_validation tool\n */\nexport interface RunTargetedValidationParams {\n snapshot_id: string\n files: string[]\n artifact_kinds?: string[]\n}\n\n/**\n * Execute a CLI command from the **project root** (different from the user's cwd).\n */\nexport interface RunTerminalCommandParams {\n /** CLI command valid for user's OS. */\n command: string\n /** SYNC (default) for finite commands that exit: waits and returns output. BACKGROUND only for long-running or never-exiting processes (dev servers, watchers, log tails): starts a detached job and returns a jobId immediately so the turn is not blocked. Live job_update already drives the user UI; use check_job for agent-side readiness/exitCode/join, not solely for user progress. */\n process_type?: 'SYNC' | 'BACKGROUND'\n /** For BACKGROUND commands only: keep the job running if the owning request is cancelled. Defaults to false. */\n detach?: boolean\n /** The working directory to run the command in. Default is the project root. */\n cwd?: string\n /** Set to -1 for no timeout. Does not apply for BACKGROUND commands. Default 30 */\n timeout_seconds?: number\n /** Runtime-managed background job owner; agents must omit. */\n owner?: {\n clientSessionId: string\n rootRunId: string\n parentRunId: string\n parentAgentId: string\n }\n}\n\n/**\n * Atomically replace conversation history and, when supplied, commit a validated structured task-memory revision.\n */\nexport interface SetMessagesParams {\n messages: any\n taskMemory?: {\n schemaVersion: 1\n goal?: string\n requirements?: string[]\n decisions?: string[]\n filesInspected?: string[]\n editsMade?: string[]\n validationResults?: string[]\n reviewReceipts?: string[]\n blockers?: string[]\n nextActions?: string[]\n historicalSummary?: string\n evidence?: {\n id: string\n kind:\n | 'requirement'\n | 'decision'\n | 'read'\n | 'edit'\n | 'validation'\n | 'review'\n | 'blocker'\n | 'handoff'\n | 'note'\n summary: string\n source?: string\n path?: string\n freshnessHash?: string\n workspaceRevision?: number\n verifiedAt?: number\n supersedes?: string[]\n stale?: boolean\n }[]\n workspaceRevision?: number\n workspaceSnapshotId?: string\n }\n expectedTaskMemoryRevision?: number\n}\n\n/**\n * JSON object to set as the agent output. The shape of the parameters are specified dynamically further down in the conversation. This completely replaces any previous output. If the agent was spawned, this value will be passed back to its parent. If the agent has an outputSchema defined, the output will be validated against it.\n */\nexport interface SetOutputParams {\n data?: Record\n [key: string]: any\n}\n\n/**\n * Load a skill by name to get its full instructions. Skills provide reusable behaviors and instructions.\n */\nexport interface SkillParams {\n /** The name of the skill to load */\n name: string\n}\n\n/**\n * Spawn up to 12 agents and send a prompt and/or parameters to each of them. These agents will run in parallel. Note that that means they will run independently. Split larger work into bounded waves. If you need to run agents sequentially, use spawn_agents with one agent at a time instead.\n */\nexport interface SpawnAgentsParams {\n agents: {\n /** Agent to spawn. Must be a name from the live \"You can spawn the following agents\" catalog (hyphenated ids; underscores accepted). */\n agent_type: string\n /** Prompt to send to the agent */\n prompt?: string\n /** If true, return jobId immediately and run as in-process coroutine; poll with check_background_agent. Defaults to false (blocking). Cannot outlive this CLI session. */\n background?: boolean\n /** Optional structured handoff; additive — non-consumers still get prompt/params. */\n handoff?:\n | {\n schemaVersion: 1\n taskId: string\n role:\n | 'orchestrator'\n | 'explorer'\n | 'thinker'\n | 'editor'\n | 'repair-editor'\n | 'test-writer'\n | 'doc-writer'\n | 'dependency-manager'\n | 'debugger'\n | 'validator'\n | 'reviewer'\n | 'security-reviewer'\n | 'committer'\n | 'synthesizer'\n | 'specialist'\n | 'general'\n objective: string\n requirements: {\n id: string\n text: string\n required: boolean\n }[]\n acceptanceCriteria: {\n id: string\n behavior: string\n verification: string\n }[]\n context:\n | {\n path: string\n symbols: string[]\n reason: string\n confidence: 'confirmed' | 'inferred' | 'unknown'\n freshnessHash?: string\n workspaceRevision?: number\n }[]\n | Record\n | string\n currentBehavior?: string\n desiredBehavior?: string\n invariants?: string[]\n nonGoals: string[]\n risks?: string[]\n unknowns?: string[]\n findings: {\n id: string\n text: string\n files: string[]\n snapshotFingerprint: string\n }[]\n permissions: {\n readablePaths: string[]\n writablePaths: string[]\n allowedTools: string[]\n }\n workspaceRevision?: number\n workspaceSnapshotId?: string\n summary?: string\n artifacts?: string[]\n successCriteria?: string[]\n constraints?: string[]\n }\n | Record\n /** Optional wall-clock deadline seconds; omit or -1 for none. Agent defaultTimeoutMs still applies when set. */\n timeout_seconds?: number\n /** Parameters object for the agent */\n params?: {\n /** Terminal command to run (basher, tmux-cli) */\n command?: string\n /** What information from the command output is desired (basher) */\n what_to_summarize?: string\n /** Timeout for command. Set to -1 for no timeout. Default 30 (basher) */\n timeout_seconds?: number\n /** Save full command output to a /tmp log and extract failure lines for long SYNC command output (basher) */\n save_full_log?: boolean\n /** grep -E failure extraction pattern used with save_full_log (basher) */\n failure_pattern?: string\n /** Maximum extracted failure lines to return with save_full_log (basher) */\n max_failure_lines?: number\n /** Relevant file paths to read (general-agent) */\n filePaths?: string[]\n /** Relevant directory paths to inventory (general-agent) */\n directoryPaths?: string[]\n /** Directories to search within (file-picker) */\n directories?: string[]\n /** Starting URL to navigate to (browser-use) */\n url?: string\n /** Exact task-owned paths eligible for staging (git-committer) */\n owned_paths?: string[]\n /** Optional branch to create or switch to (git-committer) */\n branch_name?: string\n /** Create and switch to branch_name when true (git-committer) */\n branch_switch?: boolean\n /** Allow branch create/switch on a dirty worktree (git-committer) */\n allow_dirty_branch?: boolean\n /** Push the resulting feature branch when authorized (git-committer) */\n push?: boolean\n /** Remote used for fetch/push (git-committer) */\n remote?: string\n /** Assigned gate snapshot fingerprint (reviewer specialists) */\n snapshot_id?: string\n /** Changed file paths to review (security-reviewer) */\n changed_files?: string[]\n /** Opaque snapshot token to echo (security-reviewer) */\n snapshot_fingerprint?: string\n /** Package manager selected from repository manifests (dependency-manager) */\n manager?: string\n /** Dependency operation: add, remove, sync, restore, or update (dependency-manager) */\n operation?: string\n /** Exact package specifications (dependency-manager) */\n packages?: string[]\n /** Optional workspace selector (dependency-manager) */\n workspace?: string\n /** GitHub repository URL to clone (librarian) */\n repoUrl?: string\n /** Retain the owned /tmp clone after completion (librarian) */\n retainClone?: boolean\n /** Optional search or path patterns */\n patterns?: string[]\n /** Exact files in scope (reviewer specialists) */\n files?: string[]\n /** Optional agent-specific prompts */\n prompts?: string[]\n [key: string]: any\n }\n }[]\n}\n\n/**\n * Parameters for str_replace tool\n */\nexport interface StrReplaceParams {\n /** The file to edit. */\n path: string\n atomic?: boolean\n replacements: {\n oldString: string\n newString: string\n allowMultiple?: boolean\n occurrenceIndex?: number\n /** Optional authenticated cap.v3 readCapability copied verbatim from the matching fresh read_files editAnchor. */\n basedOnRead?: string\n /** For deletion replacements only (newString is empty): treat a missing oldString as an already-applied no-op. Use only for explicit idempotent cleanup retries, never for ordinary edits. When every requested change resolves to such a no-op - every replacement of a standalone str_replace call, or every edit of an edit_transaction - the call succeeds with zero file changes and the skip messages rather than failing. When combined with occurrenceIndex, a partially-applied cleanup also skips: fewer remaining exact occurrences than the requested index means that occurrence is treated as already applied. Only valid when newString is empty; both the input and provider schemas reject any other combination. */\n skipIfMissing?: boolean\n }[]\n}\n\n/**\n * Suggest clickable followup prompts to the user. Each followup becomes a card the user can click to send that prompt.\n */\nexport interface SuggestFollowupsParams {\n /** List of suggested followup prompts the user can click to send */\n followups: {\n /** The full prompt text to send as a user message when clicked */\n prompt: string\n /** Short display label for the card (defaults to truncated prompt if not provided) */\n label?: string\n }[]\n}\n\n/**\n * Signal that the task is complete. Use this tool when:\n- The user's request is completely fulfilled\n- You need clarification from the user before continuing\n- You are stuck or need help from the user to continue\n\nThis tool explicitly marks the end of your work on the current task.\n */\nexport interface TaskCompletedParams {}\n\n/**\n * Deeply consider complex tasks by brainstorming approaches and tradeoffs step-by-step.\n */\nexport interface ThinkDeeplyParams {\n /** Detailed step-by-step analysis. Initially keep each step concise (max ~5-7 words per step). */\n thought: string\n}\n\n/**\n * Parameters for update_plan_status tool\n */\nexport interface UpdatePlanStatusParams {\n /** Artifact path. Must be `.agents/sessions//PLAN.md`, `.agents/sessions//STATUS.md`, or `.agents/sessions//LESSONS.md`. Absolute paths and `..` traversal are rejected. Editing PLAN.md is permitted only for tri-state task toggles (not full overwrites). */\n path: string\n /** Targeted updates applied in order. Each entry rewrites at most one matching checklist line; unmatched updates fall through to `append`. */\n updates?: {\n /** Stable task ID at the start of a checklist line (for example `P2-T3`). Preferred over substring matching. */\n taskId?: string\n /** Substring of the existing task/checklist line to match (case-insensitive). The first matching `- [ ]`/`-[x]`/`-[~]`/`-[/]`/`-[!]` line in the artifact will be updated in place. */\n task?: string\n /** When provided, sets the checkbox state of the matched line (true -> `[x]`, false -> `[ ]`). Ignored when `status` is also provided. */\n completed?: boolean\n /** Explicit tri-state task status. When provided, overrides `completed`. Transitions a task to `in_progress` (`[~]`), `done` (`[x]`), `cancelled` (`[/]`), `blocked` (`[!]`), or back to `pending` (`[ ]`). */\n status?: 'pending' | 'in_progress' | 'done' | 'cancelled' | 'blocked'\n /** Optional short note to append to the matched line in parentheses. Preserves any existing trailing text on the line. */\n note?: string\n }[]\n /** Optional delimited entry appended at the end of the artifact (used when there is no matching task line for the change being recorded). */\n append?: {\n /** Short heading for an appended entry. Used to form a clearly delimited block (`## `). */\n heading: string\n /** Markdown body for the appended entry. Written verbatim under the heading. */\n body: string\n }\n /** Optional session-level status transition. When provided, `.agents/sessions//STATE.json` is created or updated to reflect the new lifecycle status. */\n sessionStatus?:\n | 'draft'\n | 'ready'\n | 'active'\n | 'executing'\n | 'validating'\n | 'reviewing'\n | 'blocked'\n | 'paused'\n | 'completed'\n | 'archived'\n /** Optional current-task pointer written as a `` annotation in PLAN.md. Pass an empty string or omit to clear the pointer. Only takes effect when path targets PLAN.md. */\n currentTask?: string\n /** Optional STATE.json compare-and-swap revision. The update fails without writing when the current revision differs. */\n expectedRevision?: number\n /** Validation or review evidence associated with a stable task ID. Completing a PLAN task requires a passed validation checkpoint with receiptIds. */\n checkpoint?: {\n taskId: string\n phase: 'validation' | 'review'\n passed: boolean\n summary?: string\n receiptIds?: string[]\n }\n}\n\n/**\n * Search the web for current information, or fetch the content of a specific URL.\n */\nexport interface WebSearchParams {\n /** The search query to find relevant web content. Required unless url is provided. */\n query?: string\n /** A specific URL to fetch and read the full text content of. When provided, fetches this page directly instead of searching. Useful for reading documentation, GitHub READMEs, blog posts, or any public web page. */\n url?: string\n /** Search depth - 'standard' for quick results, 'deep' for more comprehensive search. Default is 'standard'. Ignored when url is provided. */\n depth?: 'standard' | 'deep'\n /** When fetching a URL, also extract and return links found on the page. Enables navigation by letting you see what pages are linked. Default: true. */\n include_links?: boolean\n /** Maximum number of links to extract when include_links is true. Default: 40. */\n max_links?: number\n}\n\n/**\n * Create or overwrite a file with the given content.\n */\nexport interface WriteFileParams {\n /** Path to the file relative to the **project root** */\n path: string\n /** What the change is intended to do in only one sentence. */\n instructions: string\n /** Complete file content to write to the file. */\n content: string\n /** Optional whole-file-covering cap.v3 from a fresh complete whole-file read (paths or full-file range). Only a capability that covers the entire current file (startLine=1 through the current line count) with a hash matching current content may authorize overwrite; partial range capabilities never authorize write_file. */\n basedOnRead?: string\n}\n\n/**\n * Parameters for write_audit_findings tool\n */\nexport interface WriteAuditFindingsParams {\n /** Existing durable audit session slug under .agents/sessions/. Accepts only a short identifier token: 1 to 100 characters of letters, digits, dot, underscore, or dash, and neither `.` nor `..` on its own. */\n sessionSlug: string\n /** Unique shard identifier used as the findings filename. Accepts only a short identifier token: 1 to 100 characters of letters, digits, dot, underscore, or dash, and neither `.` nor `..` on its own. */\n shardId: string\n /** Exact snapshotId returned by inspect_codebase_structure, such as its 64-character sha256 digest. Required for a directly composable structuralReceipt; omitted only for legacy callers. Accepts only a short identifier token: 1 to 100 characters of letters, digits, dot, underscore, or dash, and neither `.` nor `..` on its own. When snapshotId and coverage.domains are both present the call receives a structuralReceipt, so coverage.subsystemIds and coverage.files must each name at least one entry: evaluate_audit_coverage rejects a receipt whose subsystem_ids or files list is empty. */\n snapshotId?: string\n /** Each findings entry rejects control and Unicode format characters in title, risk, fix, and evidence — NUL, any other control character, and the U+2028/U+2029 line separators — while still accepting tabs and line breaks in that prose. findings[].path is a location rather than prose, so it must be a single-line value with none of those characters and no tabs or line breaks; it is trimmed, and the trimmed value is the one rendered into the finding heading. */\n findings: {\n severity: 'CRITICAL' | 'HIGH' | 'MEDIUM' | 'LOW'\n domain:\n | 'security'\n | 'correctness'\n | 'state-mutation'\n | 'error-handling'\n | 'performance'\n | 'dependency-hygiene'\n | 'test-coverage'\n | 'api-contract'\n | 'api-abi'\n path: string\n line?: number\n title: string\n risk: string\n fix: string\n evidence: string\n }[]\n /** Every coverage list must name each entry at most once: a repeated file, subsystemId, featureId, or domain is rejected rather than counted twice. Entries are compared after trimming surrounding whitespace, and the trimmed value is what reaches the artifact and the receipt, so two spellings that differ only in whitespace are the same entry. Every coverage files, subsystemIds, and featureIds entry must be a single-line value: tabs, carriage returns, newlines, NUL, any other control or Unicode format character, and the U+2028/U+2029 line separators are rejected. Entries are trimmed, and the trimmed value is the one uniqueness is judged on. */\n coverage: {\n /** See the coverage description for the single-line hygiene rule, which applies to this list too. See the coverage description for the uniqueness rule, which applies to this list too. */\n subsystemIds: string[]\n /** See the coverage description for the single-line hygiene rule, which applies to this list too. See the coverage description for the uniqueness rule, which applies to this list too. */\n featureIds: string[]\n /** See the coverage description for the single-line hygiene rule, which applies to this list too. See the coverage description for the uniqueness rule, which applies to this list too. */\n files: string[]\n /** coverage.domains accepts canonical domain ids only, so use api-contract there: the legacy api-abi alias is accepted only in findings[].domain. When coverage.domains is present it must name at least one domain: an empty list is rejected rather than treated as an omitted field. See the coverage description for the uniqueness rule, which applies to this list too. */\n domains?: (\n | 'security'\n | 'correctness'\n | 'state-mutation'\n | 'error-handling'\n | 'performance'\n | 'dependency-hygiene'\n | 'test-coverage'\n | 'api-contract'\n )[]\n }\n /** Set noIssuesFound=true exactly when findings is empty and false whenever findings is non-empty; any other combination is rejected. */\n noIssuesFound?: boolean\n}\n\n/**\n * Write a todo list to track tasks for multi-step implementations. Use this frequently to maintain an updated step-by-step plan.\n */\nexport interface WriteTodosParams {\n /** List of todos with their completion status. Add ALL of the applicable tasks to the list, so you don't forget to do anything. Try to order the todos the same way you will complete them. Do not mark todos as completed if you have not completed them yet! */\n todos: {\n /** Description of the task */\n task: string\n /** Whether the task is completed */\n completed: boolean\n }[]\n}\n\n/**\n * Get parameters type for a specific tool\n */\nexport type GetToolParams = ToolParamsMap[T]\n" +export const toolsSource = "/**\n * Union type of all available tool names\n */\nexport type ToolName =\n | 'add_message'\n | 'ask_user'\n | 'check_background_agent'\n | 'check_job'\n | 'code_search'\n | 'end_turn'\n | 'edit_transaction'\n | 'edit_3d_asset'\n | 'find_files'\n | 'find_files_matching_content'\n | 'git_status'\n | 'git_branch'\n | 'get_task'\n | 'get_change_review_bundle'\n | 'inspect_workspace'\n | 'inspect_environment'\n | 'inspect_3d_asset'\n | 'get_affected_tests'\n | 'get_build_targets'\n | 'inspect_codebase_structure'\n | 'inspect_feature_completeness'\n | 'evaluate_audit_coverage'\n | 'glob'\n | 'kill_job'\n | 'list_directory'\n | 'list_jobs'\n | 'lookup_agent_info'\n | 'query_index'\n | 'read_docs'\n | 'read_files'\n | 'read_image'\n | 'render_3d_preview'\n | 'read_logs'\n | 'read_outline'\n | 'read_subtree'\n | 'replace_range'\n | 'rewrite_symbol'\n | 'render_ui'\n | 'run_file_change_hooks'\n | 'run_targeted_validation'\n | 'run_terminal_command'\n | 'set_messages'\n | 'set_output'\n | 'skill'\n | 'spawn_agents'\n | 'str_replace'\n | 'suggest_followups'\n | 'task_completed'\n | 'think_deeply'\n | 'update_plan_status'\n | 'web_search'\n | 'write_file'\n | 'write_audit_findings'\n | 'write_todos'\n\n/**\n * Map of tool names to their parameter types\n */\nexport interface ToolParamsMap {\n add_message: AddMessageParams\n ask_user: AskUserParams\n check_background_agent: CheckBackgroundAgentParams\n check_job: CheckJobParams\n code_search: CodeSearchParams\n end_turn: EndTurnParams\n edit_transaction: EditTransactionParams\n edit_3d_asset: Edit3dAssetParams\n find_files: FindFilesParams\n find_files_matching_content: FindFilesMatchingContentParams\n git_status: GitStatusParams\n git_branch: GitBranchParams\n get_task: GetTaskParams\n get_change_review_bundle: GetChangeReviewBundleParams\n inspect_workspace: InspectWorkspaceParams\n inspect_environment: InspectEnvironmentParams\n inspect_3d_asset: Inspect3dAssetParams\n get_affected_tests: GetAffectedTestsParams\n get_build_targets: GetBuildTargetsParams\n inspect_codebase_structure: InspectCodebaseStructureParams\n inspect_feature_completeness: InspectFeatureCompletenessParams\n evaluate_audit_coverage: EvaluateAuditCoverageParams\n glob: GlobParams\n kill_job: KillJobParams\n list_directory: ListDirectoryParams\n list_jobs: ListJobsParams\n lookup_agent_info: LookupAgentInfoParams\n query_index: QueryIndexParams\n read_docs: ReadDocsParams\n read_files: ReadFilesParams\n read_image: ReadImageParams\n render_3d_preview: Render3dPreviewParams\n read_logs: ReadLogsParams\n read_outline: ReadOutlineParams\n read_subtree: ReadSubtreeParams\n replace_range: ReplaceRangeParams\n rewrite_symbol: RewriteSymbolParams\n render_ui: RenderUiParams\n run_file_change_hooks: RunFileChangeHooksParams\n run_targeted_validation: RunTargetedValidationParams\n run_terminal_command: RunTerminalCommandParams\n set_messages: SetMessagesParams\n set_output: SetOutputParams\n skill: SkillParams\n spawn_agents: SpawnAgentsParams\n str_replace: StrReplaceParams\n suggest_followups: SuggestFollowupsParams\n task_completed: TaskCompletedParams\n think_deeply: ThinkDeeplyParams\n update_plan_status: UpdatePlanStatusParams\n web_search: WebSearchParams\n write_file: WriteFileParams\n write_audit_findings: WriteAuditFindingsParams\n write_todos: WriteTodosParams\n}\n\n/**\n * Add a new message to the conversation history. To be used for complex requests that can't be solved in a single step, as you may forget what happened!\n */\nexport interface AddMessageParams {\n role: 'user' | 'assistant'\n content: string\n}\n\n/**\n * Ask the user a list of multiple choice questions. Each question must have at least 2 options. The agent execution will pause until the user submits their answers.\n */\nexport interface AskUserParams {\n /** List of multiple choice questions to ask the user */\n questions: {\n /** The question to ask the user */\n question: string\n /** Optional short display label. Values longer than 18 Unicode code points are truncated instead of rejecting the question. */\n header?: string\n /** Array of answer options with label and optional description. */\n options: {\n /** The display text for this option */\n label: string\n /** Explanation shown when option is focused */\n description?: string\n }[]\n /** If true, allows selecting multiple options (checkbox). If false, single selection only (radio). */\n multiSelect?: boolean\n /** Validation rules for \"Other\" text input */\n validation?: {\n /** Maximum length for \"Other\" text input */\n maxLength?: number\n /** Minimum length for \"Other\" text input */\n minLength?: number\n /** Regex pattern for \"Other\" text input */\n pattern?: string\n /** Custom error message when pattern fails */\n patternError?: string\n }\n }[]\n}\n\n/**\n * Join/wait on a background agent turn started by spawn_agents({ background: true }): returns the sequenced agent_chunk events produced since the cursor plus the unified job state. Use it to observe a long-running background agent without blocking the turn.\n */\nexport interface CheckBackgroundAgentParams {\n /** The jobId returned by spawn_agents({ background: true }) for the background agent turn. */\n jobId: string\n /** Optional sequence cursor from a prior response. Polling is idempotent for an explicit cursor; nextCursor can be supplied on the next call. */\n cursor?: number\n /** Optional substring to wait for in the new streamed chunks before returning (follow mode). Returns early as soon as it appears in any chunk payload. Useful for waiting until a background agent emits a specific milestone (e.g. a tool_result or a text marker). */\n wait_for?: string\n /** Max seconds to wait for new chunks / the wait_for pattern. 0 (default) returns immediately with whatever new chunks exist (poll mode); >0 blocks up to this long (follow mode). */\n timeout_seconds?: number\n /** When true, explicitly cancel the running background agent before returning its final status. Defaults to false. */\n cancel?: boolean\n}\n\n/**\n * Join/wait on a background job started by run_terminal_command: returns the sequenced output events produced since the last check plus the unified job state and exit code. Use it to observe a long-running process without blocking the turn. To watch an arbitrary log file, start a `tail -f ` BACKGROUND job and check_job it with a wait_for pattern.\n */\nexport interface CheckJobParams {\n /** The jobId returned by run_terminal_command with process_type: BACKGROUND. */\n jobId: string\n /** Optional substring to wait for in the new output before returning (follow mode). Returns early as soon as it appears (e.g. \"Listening on\" / \"compiled successfully\"). */\n wait_for?: string\n /** Max seconds to wait for new output / the wait_for pattern. 0 (default) returns immediately with whatever new output exists (poll mode); >0 blocks up to this long (follow mode). */\n timeout_seconds?: number\n /** Follow mode only: SIGTERM the job on follow-timeout. Poll mode never kills. Default false. */\n kill_on_timeout?: boolean\n}\n\n/**\n * Search for string patterns in the project's files. This tool uses ripgrep (rg), a fast line-oriented search tool. Use this tool only when read_files is not sufficient to find the files you need.\n */\nexport interface CodeSearchParams {\n /** The pattern to search for. */\n pattern: string\n /** Optional safe ripgrep flags as one string or argv tokens (e.g., \"-i -g *.ts -A 2\" or [\"-i\", \"-g\", \"*.ts\", \"-A\", \"2\"]). Allowed: -i/--ignore-case, -S/--smart-case, -s/--case-sensitive, -w/--word-regexp, -F/--fixed-strings, -U/--multiline, --multiline-dotall, -g/--glob, -t/--type, -T/--type-not, plus context -A/-B/-C (and long forms). JSON quotes delimit the string; do not embed another quote pair around the entire expression. Line numbers are automatic; -n/--line-number are ignored. Output-shape flags such as -c/--count, --count-matches, -l, -v/--invert-match, -r/--replace, --exec, and -z/--null are rejected. */\n flags?: string | string[]\n /** Optional working directory or single file to search within, relative to the project root or absolute. Absolute paths may be outside the project. A directory becomes ripgrep's cwd and scopes the search under that path (plus existing blessed hidden dirs when no paths are given); a file scopes the search to that file only (process cwd = project root when the file is under the project, else the file's parent). Defaults to searching the entire project root. */\n cwd?: string\n /** Optional list of file and/or directory paths to search (relative to the project root, or absolute). When non-empty, ripgrep searches only these targets instead of the whole cwd tree (and does not auto-expand hidden dirs). Can be combined with a file cwd. */\n paths?: string[]\n /** Maximum number of results to return per file. Defaults to 15. There is also a global limit of 250 results across all files. */\n maxResults?: number\n}\n\n/**\n * End your turn, regardless of any new tool results that might be coming. This will allow the user to type another prompt.\n */\nexport interface EndTurnParams {}\n\n/**\n * Parameters for edit_transaction tool\n */\nexport interface EditTransactionParams {\n edits: (\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'str_replace'\n replacements: {\n oldString: string\n newString: string\n allowMultiple?: boolean\n occurrenceIndex?: number\n /** Optional authenticated cap.v3 readCapability copied verbatim from the matching fresh read_files editAnchor. */\n basedOnRead?: string\n /** For deletion replacements only (newString is empty): treat a missing oldString as an already-applied no-op. Use only for explicit idempotent cleanup retries, never for ordinary edits. When every requested change resolves to such a no-op - every replacement of a standalone str_replace call, or every edit of an edit_transaction - the call succeeds with zero file changes and the skip messages rather than failing. When combined with occurrenceIndex, a partially-applied cleanup also skips: fewer remaining exact occurrences than the requested index means that occurrence is treated as already applied. Only valid when newString is empty; both the input and provider schemas reject any other combination. */\n skipIfMissing?: boolean\n }[]\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n /** A structured edit dispatched by operation kind. */\n type: 'structured'\n /** Structured edit operation to apply to this file. */\n operation:\n | {\n /** Deterministic text insertion. */\n kind: 'insert_text'\n /** 1-indexed insertion position. */\n position: {\n /** 1-indexed target line. */\n line: number\n /** 1-indexed target column. */\n column: number\n }\n text: string\n }\n | {\n /** Language-aware import insertion. */\n kind: 'insert_import'\n /** Complete language-native import statement to add, e.g. \"import { foo } from 'bar'\", \"from app import value\", or \"use crate::value\". */\n importStatement: string\n }\n | {\n /** Language-aware import removal. */\n kind: 'remove_import'\n /** Complete language-native import statement to remove. Required unless moduleSpecifier is provided. */\n importStatement?: string\n /** Module specifier to remove imports from, e.g. \"react\" or \"./helper\". */\n moduleSpecifier?: string\n }\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'create'\n /** Exact bytes to write to the new file. */\n content: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'delete'\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'move'\n /** New project-relative path. The destination must be absent. */\n destinationPath: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'replace_range'\n readCapability: string\n startLine?: number\n endLine?: number\n occurrence?: {\n match: string\n occurrence?: number\n }\n newContent: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'rewrite_symbol'\n symbol: string\n content: string\n occurrence?: number\n /** Optional cap.v3 copied from the matching read_files symbol slice. It authorizes exactly the symbol and its contiguous preceding comment block. */\n readCapability?: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'patch'\n diff: string\n }\n | {\n /** Optional stable edit identifier echoed in diagnostics. */\n id?: string\n /** The file to edit. */\n path: string\n type: 'write_file'\n content: string\n /** Optional whole-file-covering cap.v3 from a fresh complete whole-file read. Only a full-file capability with a hash matching current content may authorize overwrite; partial ranges never authorize write_file. */\n basedOnRead?: string\n }\n )[]\n}\n\n/**\n * Parameters for edit_3d_asset tool\n */\nexport interface Edit3dAssetParams {\n /** Project-relative .blend path. */\n path: string\n /** Exact source hash returned by inspect_3d_asset. */\n source_hash: string\n operations: (\n | {\n type: 'rename_object'\n object: string\n new_name: string\n }\n | {\n type: 'set_object_transform'\n object: string\n location?: any[]\n rotation_degrees?: any[]\n scale?: any[]\n }\n | {\n type: 'set_render_resolution'\n width: number\n height: number\n percentage?: number\n }\n | {\n type: 'set_frame_range'\n start: number\n end: number\n }\n )[]\n}\n\n/**\n * Find several files related to a brief natural language description of the files or the name of a function or class you are looking for.\n */\nexport interface FindFilesParams {\n /** A brief natural language description of the files or the name of a function or class you are looking for. It's also helpful to mention a directory or two to look within. */\n prompt: string\n}\n\n/**\n * List unique file paths whose content matches a pattern, with optional symbol grouping. Built on top of ripgrep (rg).\n */\nexport interface FindFilesMatchingContentParams {\n /** Regex pattern (ripgrep syntax) to match file content against. */\n pattern: string\n /** Optional safe ripgrep flags as one string or argv tokens. Allowed: -i/--ignore-case, -S/--smart-case, -s/--case-sensitive, -w/--word-regexp, -F/--fixed-strings, -U/--multiline, --multiline-dotall, -g/--glob, -t/--type, -T/--type-not. Examples: \"-g *.ts -g *.tsx\" or [\"-g\", \"*.ts\", \"-g\", \"*.tsx\"]. Do not quote the entire expression inside the JSON string. Output-shape flags such as -c/--count, --count-matches, -l, -v/--invert-match, context -A/-B/-C, -r/--replace, --exec, and -z/--null are rejected (this tool forces -l or --json itself). Redundant -n/--line-number inputs are ignored. */\n flags?: string | string[]\n /** Optional working directory or single file to search within, relative to the project root or absolute. Absolute paths may be outside the project. A directory becomes ripgrep's cwd and scopes the search under that path (plus existing blessed hidden dirs); a file scopes the search to that file only (process cwd = project root when the file is under the project, else the file's parent). Defaults to the project root. */\n cwd?: string\n /** Maximum number of unique files to return. Defaults to 100. */\n maxFiles?: number\n /** When true, also return the names of the top-level symbols (functions, classes, methods, exports, constants) that contain each match, plus the per-file match count. Symbol extraction is heuristic and works best for JS/TS/Python/Go/Rust source files; languages without a recognized declaration shape produce an empty symbols list. */\n groupBySymbol?: boolean\n /** Maximum seconds to let ripgrep run before returning partial results. Defaults to 15. */\n timeoutSeconds?: number\n}\n\n/**\n * Read-only git status and (optionally) diff for the current project.\n */\nexport interface GitStatusParams {\n /** When true, also return the unified diff of uncommitted changes. */\n include_diff?: boolean\n /** When true with include_diff, returns the staged diff instead of unstaged. */\n staged?: boolean\n /** Optional path to scope status/diff to (relative to project root). */\n path?: string\n /** Maximum characters of diff output to return. Defaults to 40,000. */\n max_chars?: number\n}\n\n/**\n * Create a new git branch, optionally switching to it. Refuses to branch when the working tree is dirty unless `allow_dirty` is true.\n */\nexport interface GitBranchParams {\n /** Name of the branch to create. Must start with an alphanumeric character and contain only [a-zA-Z0-9._/-]. */\n branch_name: string\n /** When true (default), create AND switch to the branch (`git checkout -b`). When false, only create the branch (`git branch`), leaving the current branch checked out. */\n switch?: boolean\n /** When true, skip the dirty-tree refusal check. Defaults to false — the tool refuses to branch when the working tree has uncommitted changes. */\n allow_dirty?: boolean\n}\n\n/**\n * Parameters for get_task tool\n */\nexport interface GetTaskParams {\n /** Optional plan session slug. Defaults to .agents/ACTIVE_SESSION. */\n session?: string\n}\n\n/**\n * Parameters for get_change_review_bundle tool\n */\nexport interface GetChangeReviewBundleParams {\n max_chars?: number\n}\n\n/**\n * Inspect the current repository/worktree identity and Git state without modifying it.\n */\nexport interface InspectWorkspaceParams {}\n\n/**\n * Parameters for inspect_environment tool\n */\nexport interface InspectEnvironmentParams {}\n\n/**\n * Parameters for inspect_3d_asset tool\n */\nexport interface Inspect3dAssetParams {\n /** Project-relative 3D asset path. */\n path: string\n}\n\n/**\n * Parameters for get_affected_tests tool\n */\nexport interface GetAffectedTestsParams {\n files: string[]\n}\n\n/**\n * Parameters for get_build_targets tool\n */\nexport interface GetBuildTargetsParams {\n files: string[]\n}\n\n/**\n * Parameters for inspect_codebase_structure tool\n */\nexport interface InspectCodebaseStructureParams {\n scope?: string[]\n}\n\n/**\n * Parameters for inspect_feature_completeness tool\n */\nexport interface InspectFeatureCompletenessParams {\n feature: string\n snapshot_id: string\n scope?: string[]\n}\n\n/**\n * Parameters for evaluate_audit_coverage tool\n */\nexport interface EvaluateAuditCoverageParams {\n snapshot_id: string\n structural_receipts: {\n schema_version: 1\n snapshot_id: string\n shard_id: string\n subsystem_ids: string[]\n files: string[]\n domains: (\n | 'security'\n | 'correctness'\n | 'state-mutation'\n | 'error-handling'\n | 'performance'\n | 'dependency-hygiene'\n | 'test-coverage'\n | 'api-contract'\n )[]\n }[]\n features: {\n schema_version: 1\n snapshot_id: string\n feature: string\n evidence_kind: 'heuristic' | 'verified'\n evidence: {\n entrypoints: string[]\n implementation: string[]\n consumers: string[]\n tests: string[]\n docs: string[]\n failure_states: string[]\n }\n }[]\n out_of_scope?: {\n id: string\n reason: string\n }[]\n scope?: string[]\n}\n\n/**\n * Search for files matching a glob pattern. Returns matching file paths sorted by modification time (newest first, then path for deterministic ties).\n */\nexport interface GlobParams {\n /** Glob pattern to match files against (e.g., *.js, src/glob/*.ts, glob/test/glob/*.go). */\n pattern: string\n /** Optional working directory or file path, relative to project root. If a directory, the glob pattern is matched against paths relative to this cwd, while returned files remain project-relative. If a file path, the pattern is matched against that file only (full path or basename). If not provided, searches from project root. */\n cwd?: string\n}\n\n/**\n * Cancel a background job started by run_terminal_command.\n */\nexport interface KillJobParams {\n /** The jobId returned by run_terminal_command with process_type: BACKGROUND. */\n jobId: string\n /** Signal to send. Defaults to SIGTERM; use SIGKILL only if graceful termination fails. */\n signal?: 'SIGTERM' | 'SIGKILL'\n}\n\n/**\n * List files and directories in the specified path. Returns separate arrays of file names and directory names.\n */\nexport interface ListDirectoryParams {\n /** Directory path to list, relative to the project root. */\n path: string\n}\n\n/**\n * List this run's background jobs (shell processes and background agents, running and settled) with statuses, bucketed pending process/log output relative to the last check_job consumer cursor (agents usually show pending: 'none'), and a gap flag.\n */\nexport interface ListJobsParams {}\n\n/**\n * Retrieve information about an agent by ID\n */\nexport interface LookupAgentInfoParams {\n /** Agent ID (short local or full published format) */\n agentId: string\n}\n\n/**\n * Query the local codebase graph index to find relevant files ranked by symbol names, imports, headings, paths, doc concepts, and graph relationships. The index is built automatically on startup.\n */\nexport interface QueryIndexParams {\n /** Natural language query or keyword terms describing the files you are looking for. Optional for graph modes when from/to paths are provided. For example: \"authentication\", \"database migrations\", \"editor mutation logic\", \"React components\". */\n query?: string\n /** Maximum number of results to return. Defaults to 20. */\n limit?: number\n /** Optional list of file extensions to filter results (without dot). E.g. [\"ts\", \"tsx\"] for TypeScript only. */\n fileTypes?: string[]\n /** Optional normalized project-relative directory prefixes. Results outside every prefix are excluded before ranking/limiting. */\n pathPrefixes?: string[]\n /** search|explain|neighbors|path|commands|references — see tool description. */\n mode?: 'search' | 'neighbors' | 'path' | 'explain' | 'commands' | 'references'\n /** Optional source file path for neighbors, path, and references modes. */\n from?: string\n /** Optional target file path for path mode. Also used as the seed file for references mode when from is omitted or not indexed. */\n to?: string\n}\n\n/**\n * Fetch up-to-date documentation for libraries and frameworks using Context7 API.\n */\nexport interface ReadDocsParams {\n /** The library or framework name (e.g., \"Next.js\", \"MongoDB\", \"React\"). Use the official name as it appears in documentation if possible. Only public libraries available in Context7's database are supported, so small or private libraries may not be available. */\n libraryTitle: string\n /** Specific topic to focus on (e.g., \"routing\", \"hooks\", \"authentication\") */\n topic: string\n /** Optional maximum number of tokens to return. Defaults to 10000. Values less than 10000 are automatically increased to 10000. */\n max_tokens?: number\n}\n\n/**\n * Read multiple files from disk and return their contents. Use this tool to read as many files as would be helpful to answer the user's request.\n */\nexport interface ReadFilesParams {\n /** Whole-file paths to read. Complete results include editAnchor.readCapability for follow-up edits. */\n paths?: string[]\n /** 1-indexed inclusive line ranges. Sole `paths` entry infers missing path. */\n ranges?: {\n /** Project-relative file path. */\n path: string\n /** 1-indexed inclusive start line. Defaults to 1. */\n startLine?: number\n /** 1-indexed inclusive end line. Defaults to the last line. */\n endLine?: number\n }[]\n /** Contiguous line windows; each complete window mints a scoped cap.v3 editAnchor. */\n windows?: {\n /** File path to read in contiguous line windows, relative to the project root. */\n path: string\n /** Lines per window. Defaults to 400, capped at 5000. */\n windowSize?: number\n /** 1-indexed window number to return. Omit to get the window manifest (totalLines, windowSize, windowCount) plus the first window. */\n window?: number\n }[]\n /** Literal-anchored context blocks with a scoped cap.v3 editAnchor per block. */\n around?: {\n /** File path to read a content-anchored block from, relative to the project root. */\n path: string\n /** Exact literal string to anchor on. Robust to line-number drift. */\n match: string\n /** 1-indexed occurrence of `match` to anchor on. Defaults to 1. */\n occurrence?: number\n /** Lines of context to include on each side of the match, clamped at file boundaries. Defaults to 40, capped at 2000. */\n contextLines?: number\n }[]\n /** Nth top-level symbol by name (rewrite_symbol occurrence semantics); prefer batch `symbols` when possible. */\n symbol?: {\n /** File path to extract a symbol slice from, relative to the project root. */\n path: string\n /** Top-level symbol name (function, class, interface, method) to pull, as shown by read_outline. */\n name: string\n /** When multiple top-level symbols share this name, the 1-indexed one to return. Defaults to 1. Matches rewrite_symbol occurrence semantics. */\n occurrence?: number\n }[]\n /** Named symbol slices with editAnchors; prefer over full reads when names are known. */\n symbols?: {\n /** Project-relative file path. */\n path: string\n /** Symbol names to slice. */\n names: string[]\n }[]\n}\n\n/**\n * Read image files from disk and return them as model-visible image media.\n */\nexport interface ReadImageParams {\n /** List of image file paths to read. */\n paths: string[]\n}\n\n/**\n * Parameters for render_3d_preview tool\n */\nexport interface Render3dPreviewParams {\n /** Project-relative 3D asset path. */\n path: string\n views?: ('camera' | 'perspective' | 'front' | 'side' | 'top')[]\n mode?: 'material' | 'clay' | 'wireframe'\n width?: number\n height?: number\n}\n\n/**\n * Read the last N lines from a log/text file or background job log without starting a background tail process.\n */\nexport interface ReadLogsParams {\n /** Path to the log file, relative to the project root unless absolute. Required unless jobId is provided. */\n path?: string\n /** Background job id returned by run_terminal_command(process_type: BACKGROUND). When provided, reads the job log file directly. */\n jobId?: string\n /** Number of trailing lines to read. Defaults to 200. */\n lines?: number\n /** Maximum characters to return. Defaults to 20,000. */\n max_chars?: number\n}\n\n/**\n * Generate an outline of imports, exports, classes, methods, and function signatures in a source file without reading the entire implementation.\n */\nexport interface ReadOutlineParams {\n /** File path to generate the AST-like outline for, relative to the project root. */\n path: string\n}\n\n/**\n * Read one or more directory subtrees (as a blob including subdirectories, file names, and parsed variables within each source file) or return parsed variable names for files. If no paths are provided, returns the entire project tree.\n */\nexport interface ReadSubtreeParams {\n /** List of paths to directories or files. Relative to the project root. If omitted, the entire project tree is used. */\n paths?: string[]\n /** Maximum token budget for the subtree blob; the tree will be truncated to fit within this budget by first dropping file variables and then removing the most-nested files and directories. */\n maxTokens?: number\n}\n\n/**\n * Replace all of, a contained sub-range of, or the Nth literal occurrence inside content observed through one fresh cap.v3 read capability.\n */\nexport interface ReplaceRangeParams {\n /** The path to the file to edit. */\n path: string\n /** Copy the cap.v3 readCapability verbatim from the matching fresh read_files editAnchor. The token supplies the observed line bounds and content hash. */\n readCapability: string\n /** Optional 1-indexed target start within the capability-covered range. Omit with endLine to replace the complete observed range. */\n startLine?: number\n /** Optional 1-indexed target end within the capability-covered range. Omit with startLine to replace the complete observed range. */\n endLine?: number\n /** Optional occurrence targeting: replace the 1-indexed occurrence (default 1) of the exact literal match found inside the capability-authorized range. Mutually exclusive with startLine/endLine. */\n occurrence?: {\n match: string\n occurrence?: number\n }\n /** Complete replacement content for the selected line range. */\n newContent: string\n}\n\n/**\n * Replace a whole symbol's definition by name using the file's syntax tree, without copying its current text. Resolves the exact AST range and applies it through the safe str_replace path (atomic, anchored).\n */\nexport interface RewriteSymbolParams {\n /** File path containing the symbol, relative to the project root. */\n path: string\n /** Name of the function/class/method/type/interface to replace (as shown by read_outline). */\n symbol: string\n /** The complete new source for the symbol, replacing its entire current definition (e.g. the whole function including its signature and body). Provide REAL newlines/tabs in the string — literal backslash-n (\\n) and backslash-t (\\t) sequences are not interpreted and will be written verbatim into the file. This matches str_replace. */\n content: string\n /** When multiple top-level symbols share this name, the 1-indexed one to replace. */\n occurrence?: number\n /** Optional cap.v3 copied from the matching read_files symbol slice. Under strict read-before-edit this authorizes exactly the symbol and its contiguous preceding comment block. */\n readCapability?: string\n}\n\n/**\n * Render a small interactive UI widget in the Openbuff CLI. Currently supports a button that opens a link.\n */\nexport interface RenderUiParams {\n /** The UI widget to render. */\n widget: {\n /** Widget type. Currently, the only supported widget is button. */\n type: 'button'\n /** Short button label shown to the user. */\n text: string\n /** The http:// or https:// URL to open when the user clicks the button. */\n link: string\n /** Theme-aware color treatment. Use primary for the main action and secondary for lower-emphasis actions. */\n variant?: 'primary' | 'secondary'\n }\n}\n\n/**\n * Parameters for run_file_change_hooks tool\n */\nexport interface RunFileChangeHooksParams {\n /** List of file paths that were changed and should trigger file change hooks */\n files: string[]\n}\n\n/**\n * Parameters for run_targeted_validation tool\n */\nexport interface RunTargetedValidationParams {\n snapshot_id: string\n files: string[]\n artifact_kinds?: string[]\n}\n\n/**\n * Execute a CLI command from the **project root** (different from the user's cwd).\n */\nexport interface RunTerminalCommandParams {\n /** CLI command valid for user's OS. */\n command: string\n /** SYNC (default) for finite commands that exit: waits and returns output. BACKGROUND only for long-running or never-exiting processes (dev servers, watchers, log tails): starts a detached job and returns a jobId immediately so the turn is not blocked. Live job_update already drives the user UI; use check_job for agent-side readiness/exitCode/join, not solely for user progress. */\n process_type?: 'SYNC' | 'BACKGROUND'\n /** For BACKGROUND commands only: keep the job running if the owning request is cancelled. Defaults to false. */\n detach?: boolean\n /** The working directory to run the command in. Default is the project root. */\n cwd?: string\n /** Wall-clock bound in seconds for SYNC commands. Omit or use -1 for no timeout (the default). Does not apply to BACKGROUND commands. */\n timeout_seconds?: number\n /** Runtime-managed background job owner; agents must omit. */\n owner?: {\n clientSessionId: string\n rootRunId: string\n parentRunId: string\n parentAgentId: string\n }\n}\n\n/**\n * Atomically replace conversation history and, when supplied, commit a validated structured task-memory revision.\n */\nexport interface SetMessagesParams {\n messages: any\n taskMemory?: {\n schemaVersion: 1\n goal?: string\n requirements?: string[]\n decisions?: string[]\n filesInspected?: string[]\n editsMade?: string[]\n validationResults?: string[]\n reviewReceipts?: string[]\n blockers?: string[]\n nextActions?: string[]\n historicalSummary?: string\n evidence?: {\n id: string\n kind:\n | 'requirement'\n | 'decision'\n | 'read'\n | 'edit'\n | 'validation'\n | 'review'\n | 'blocker'\n | 'handoff'\n | 'note'\n summary: string\n source?: string\n path?: string\n freshnessHash?: string\n workspaceRevision?: number\n verifiedAt?: number\n supersedes?: string[]\n stale?: boolean\n }[]\n workspaceRevision?: number\n workspaceSnapshotId?: string\n }\n expectedTaskMemoryRevision?: number\n}\n\n/**\n * JSON object to set as the agent output. The shape of the parameters are specified dynamically further down in the conversation. This completely replaces any previous output. If the agent was spawned, this value will be passed back to its parent. If the agent has an outputSchema defined, the output will be validated against it.\n */\nexport interface SetOutputParams {\n data?: Record\n [key: string]: any\n}\n\n/**\n * Load a skill by name to get its full instructions. Skills provide reusable behaviors and instructions.\n */\nexport interface SkillParams {\n /** The name of the skill to load */\n name: string\n}\n\n/**\n * Spawn up to 12 agents and send a prompt and/or parameters to each of them. These agents will run in parallel. Note that that means they will run independently. Split larger work into bounded waves. If you need to run agents sequentially, use spawn_agents with one agent at a time instead.\n */\nexport interface SpawnAgentsParams {\n agents: {\n /** Agent to spawn. Must be a name from the live \"You can spawn the following agents\" catalog (hyphenated ids; underscores accepted). */\n agent_type: string\n /** Prompt to send to the agent */\n prompt?: string\n /** If true, return jobId immediately and run as in-process coroutine; poll with check_background_agent. Defaults to false (blocking). Cannot outlive this CLI session. */\n background?: boolean\n /** Optional structured handoff; additive — non-consumers still get prompt/params. */\n handoff?:\n | {\n schemaVersion: 1\n taskId: string\n role:\n | 'orchestrator'\n | 'explorer'\n | 'thinker'\n | 'editor'\n | 'repair-editor'\n | 'test-writer'\n | 'doc-writer'\n | 'dependency-manager'\n | 'debugger'\n | 'validator'\n | 'reviewer'\n | 'security-reviewer'\n | 'committer'\n | 'synthesizer'\n | 'specialist'\n | 'general'\n objective: string\n requirements: {\n id: string\n text: string\n required: boolean\n }[]\n acceptanceCriteria: {\n id: string\n behavior: string\n verification: string\n }[]\n context:\n | {\n path: string\n symbols: string[]\n reason: string\n confidence: 'confirmed' | 'inferred' | 'unknown'\n freshnessHash?: string\n workspaceRevision?: number\n }[]\n | Record\n | string\n currentBehavior?: string\n desiredBehavior?: string\n invariants?: string[]\n nonGoals: string[]\n risks?: string[]\n unknowns?: string[]\n findings: {\n id: string\n text: string\n files: string[]\n snapshotFingerprint: string\n }[]\n permissions: {\n readablePaths: string[]\n writablePaths: string[]\n allowedTools: string[]\n }\n workspaceRevision?: number\n workspaceSnapshotId?: string\n summary?: string\n artifacts?: string[]\n successCriteria?: string[]\n constraints?: string[]\n }\n | Record\n /** Parameters object for the agent */\n params?: {\n /** Terminal command to run (basher, tmux-cli) */\n command?: string\n /** What information from the command output is desired (basher) */\n what_to_summarize?: string\n /** Timeout for command in seconds. Omit or -1 for no timeout (default). */\n timeout_seconds?: number\n /** Save full command output to a /tmp log and extract failure lines for long SYNC command output (basher) */\n save_full_log?: boolean\n /** grep -E failure extraction pattern used with save_full_log (basher) */\n failure_pattern?: string\n /** Maximum extracted failure lines to return with save_full_log (basher) */\n max_failure_lines?: number\n /** Relevant file paths to read (general-agent) */\n filePaths?: string[]\n /** Relevant directory paths to inventory (general-agent) */\n directoryPaths?: string[]\n /** Directories to search within (file-picker) */\n directories?: string[]\n /** Starting URL to navigate to (browser-use) */\n url?: string\n /** Exact task-owned paths eligible for staging (git-committer) */\n owned_paths?: string[]\n /** Optional branch to create or switch to (git-committer) */\n branch_name?: string\n /** Create and switch to branch_name when true (git-committer) */\n branch_switch?: boolean\n /** Allow branch create/switch on a dirty worktree (git-committer) */\n allow_dirty_branch?: boolean\n /** Push the resulting feature branch when authorized (git-committer) */\n push?: boolean\n /** Remote used for fetch/push (git-committer) */\n remote?: string\n /** Assigned gate snapshot fingerprint (reviewer specialists) */\n snapshot_id?: string\n /** Changed file paths to review (security-reviewer) */\n changed_files?: string[]\n /** Opaque snapshot token to echo (security-reviewer) */\n snapshot_fingerprint?: string\n /** Package manager selected from repository manifests (dependency-manager) */\n manager?: string\n /** Dependency operation: add, remove, sync, restore, or update (dependency-manager) */\n operation?: string\n /** Exact package specifications (dependency-manager) */\n packages?: string[]\n /** Optional workspace selector (dependency-manager) */\n workspace?: string\n /** GitHub repository URL to clone (librarian) */\n repoUrl?: string\n /** Retain the owned /tmp clone after completion (librarian) */\n retainClone?: boolean\n /** Optional search or path patterns */\n patterns?: string[]\n /** Exact files in scope (reviewer specialists) */\n files?: string[]\n /** Optional agent-specific prompts */\n prompts?: string[]\n [key: string]: any\n }\n }[]\n}\n\n/**\n * Parameters for str_replace tool\n */\nexport interface StrReplaceParams {\n /** The file to edit. */\n path: string\n atomic?: boolean\n replacements: {\n oldString: string\n newString: string\n allowMultiple?: boolean\n occurrenceIndex?: number\n /** Optional authenticated cap.v3 readCapability copied verbatim from the matching fresh read_files editAnchor. */\n basedOnRead?: string\n /** For deletion replacements only (newString is empty): treat a missing oldString as an already-applied no-op. Use only for explicit idempotent cleanup retries, never for ordinary edits. When every requested change resolves to such a no-op - every replacement of a standalone str_replace call, or every edit of an edit_transaction - the call succeeds with zero file changes and the skip messages rather than failing. When combined with occurrenceIndex, a partially-applied cleanup also skips: fewer remaining exact occurrences than the requested index means that occurrence is treated as already applied. Only valid when newString is empty; both the input and provider schemas reject any other combination. */\n skipIfMissing?: boolean\n }[]\n}\n\n/**\n * Suggest clickable followup prompts to the user. Each followup becomes a card the user can click to send that prompt.\n */\nexport interface SuggestFollowupsParams {\n /** List of suggested followup prompts the user can click to send */\n followups: {\n /** The full prompt text to send as a user message when clicked */\n prompt: string\n /** Short display label for the card (defaults to truncated prompt if not provided) */\n label?: string\n }[]\n}\n\n/**\n * Signal that the task is complete. Use this tool when:\n- The user's request is completely fulfilled\n- You need clarification from the user before continuing\n- You are stuck or need help from the user to continue\n\nThis tool explicitly marks the end of your work on the current task.\n */\nexport interface TaskCompletedParams {}\n\n/**\n * Deeply consider complex tasks by brainstorming approaches and tradeoffs step-by-step.\n */\nexport interface ThinkDeeplyParams {\n /** Detailed step-by-step analysis. Initially keep each step concise (max ~5-7 words per step). */\n thought: string\n}\n\n/**\n * Parameters for update_plan_status tool\n */\nexport interface UpdatePlanStatusParams {\n /** Artifact path. Must be `.agents/sessions//PLAN.md`, `.agents/sessions//STATUS.md`, or `.agents/sessions//LESSONS.md`. Absolute paths and `..` traversal are rejected. Editing PLAN.md is permitted only for tri-state task toggles (not full overwrites). */\n path: string\n /** Targeted updates applied in order. Each entry rewrites at most one matching checklist line; unmatched updates fall through to `append`. */\n updates?: {\n /** Stable task ID at the start of a checklist line (for example `P2-T3`). Preferred over substring matching. */\n taskId?: string\n /** Substring of the existing task/checklist line to match (case-insensitive). The first matching `- [ ]`/`-[x]`/`-[~]`/`-[/]`/`-[!]` line in the artifact will be updated in place. */\n task?: string\n /** When provided, sets the checkbox state of the matched line (true -> `[x]`, false -> `[ ]`). Ignored when `status` is also provided. */\n completed?: boolean\n /** Explicit tri-state task status. When provided, overrides `completed`. Transitions a task to `in_progress` (`[~]`), `done` (`[x]`), `cancelled` (`[/]`), `blocked` (`[!]`), or back to `pending` (`[ ]`). */\n status?: 'pending' | 'in_progress' | 'done' | 'cancelled' | 'blocked'\n /** Optional short note to append to the matched line in parentheses. Preserves any existing trailing text on the line. */\n note?: string\n }[]\n /** Optional delimited entry appended at the end of the artifact (used when there is no matching task line for the change being recorded). */\n append?: {\n /** Short heading for an appended entry. Used to form a clearly delimited block (`## `). */\n heading: string\n /** Markdown body for the appended entry. Written verbatim under the heading. */\n body: string\n }\n /** Optional session-level status transition. When provided, `.agents/sessions//STATE.json` is created or updated to reflect the new lifecycle status. */\n sessionStatus?:\n | 'draft'\n | 'ready'\n | 'active'\n | 'executing'\n | 'validating'\n | 'reviewing'\n | 'blocked'\n | 'paused'\n | 'completed'\n | 'archived'\n /** Optional current-task pointer written as a `` annotation in PLAN.md. Pass an empty string or omit to clear the pointer. Only takes effect when path targets PLAN.md. */\n currentTask?: string\n /** Optional STATE.json compare-and-swap revision. The update fails without writing when the current revision differs. */\n expectedRevision?: number\n /** Validation or review evidence associated with a stable task ID. Completing a PLAN task requires a passed validation checkpoint with receiptIds. */\n checkpoint?: {\n taskId: string\n phase: 'validation' | 'review'\n passed: boolean\n summary?: string\n receiptIds?: string[]\n }\n}\n\n/**\n * Search the web for current information, or fetch the content of a specific URL.\n */\nexport interface WebSearchParams {\n /** The search query to find relevant web content. Required unless url is provided. */\n query?: string\n /** A specific URL to fetch and read the full text content of. When provided, fetches this page directly instead of searching. Useful for reading documentation, GitHub READMEs, blog posts, or any public web page. */\n url?: string\n /** Search depth - 'standard' for quick results, 'deep' for more comprehensive search. Default is 'standard'. Ignored when url is provided. */\n depth?: 'standard' | 'deep'\n /** When fetching a URL, also extract and return links found on the page. Enables navigation by letting you see what pages are linked. Default: true. */\n include_links?: boolean\n /** Maximum number of links to extract when include_links is true. Default: 40. */\n max_links?: number\n}\n\n/**\n * Create or overwrite a file with the given content.\n */\nexport interface WriteFileParams {\n /** Path to the file relative to the **project root** */\n path: string\n /** What the change is intended to do in only one sentence. */\n instructions: string\n /** Complete file content to write to the file. */\n content: string\n /** Optional whole-file-covering cap.v3 from a fresh complete whole-file read (paths or full-file range). Only a capability that covers the entire current file (startLine=1 through the current line count) with a hash matching current content may authorize overwrite; partial range capabilities never authorize write_file. */\n basedOnRead?: string\n}\n\n/**\n * Parameters for write_audit_findings tool\n */\nexport interface WriteAuditFindingsParams {\n /** Existing durable audit session slug under .agents/sessions/. Accepts only a short identifier token: 1 to 100 characters of letters, digits, dot, underscore, or dash, and neither `.` nor `..` on its own. */\n sessionSlug: string\n /** Unique shard identifier used as the findings filename. Accepts only a short identifier token: 1 to 100 characters of letters, digits, dot, underscore, or dash, and neither `.` nor `..` on its own. */\n shardId: string\n /** Exact snapshotId returned by inspect_codebase_structure, such as its 64-character sha256 digest. Required for a directly composable structuralReceipt; omitted only for legacy callers. Accepts only a short identifier token: 1 to 100 characters of letters, digits, dot, underscore, or dash, and neither `.` nor `..` on its own. When snapshotId and coverage.domains are both present the call receives a structuralReceipt, so coverage.subsystemIds and coverage.files must each name at least one entry: evaluate_audit_coverage rejects a receipt whose subsystem_ids or files list is empty. */\n snapshotId?: string\n /** Each findings entry rejects control and Unicode format characters in title, risk, fix, and evidence — NUL, any other control character, and the U+2028/U+2029 line separators — while still accepting tabs and line breaks in that prose. findings[].path is a location rather than prose, so it must be a single-line value with none of those characters and no tabs or line breaks; it is trimmed, and the trimmed value is the one rendered into the finding heading. */\n findings: {\n severity: 'CRITICAL' | 'HIGH' | 'MEDIUM' | 'LOW'\n domain:\n | 'security'\n | 'correctness'\n | 'state-mutation'\n | 'error-handling'\n | 'performance'\n | 'dependency-hygiene'\n | 'test-coverage'\n | 'api-contract'\n | 'api-abi'\n path: string\n line?: number\n title: string\n risk: string\n fix: string\n evidence: string\n }[]\n /** Every coverage list must name each entry at most once: a repeated file, subsystemId, featureId, or domain is rejected rather than counted twice. Entries are compared after trimming surrounding whitespace, and the trimmed value is what reaches the artifact and the receipt, so two spellings that differ only in whitespace are the same entry. Every coverage files, subsystemIds, and featureIds entry must be a single-line value: tabs, carriage returns, newlines, NUL, any other control or Unicode format character, and the U+2028/U+2029 line separators are rejected. Entries are trimmed, and the trimmed value is the one uniqueness is judged on. */\n coverage: {\n /** See the coverage description for the single-line hygiene rule, which applies to this list too. See the coverage description for the uniqueness rule, which applies to this list too. */\n subsystemIds: string[]\n /** See the coverage description for the single-line hygiene rule, which applies to this list too. See the coverage description for the uniqueness rule, which applies to this list too. */\n featureIds: string[]\n /** See the coverage description for the single-line hygiene rule, which applies to this list too. See the coverage description for the uniqueness rule, which applies to this list too. */\n files: string[]\n /** coverage.domains accepts canonical domain ids only, so use api-contract there: the legacy api-abi alias is accepted only in findings[].domain. When coverage.domains is present it must name at least one domain: an empty list is rejected rather than treated as an omitted field. See the coverage description for the uniqueness rule, which applies to this list too. */\n domains?: (\n | 'security'\n | 'correctness'\n | 'state-mutation'\n | 'error-handling'\n | 'performance'\n | 'dependency-hygiene'\n | 'test-coverage'\n | 'api-contract'\n )[]\n }\n /** Set noIssuesFound=true exactly when findings is empty and false whenever findings is non-empty; any other combination is rejected. */\n noIssuesFound?: boolean\n}\n\n/**\n * Write a todo list to track tasks for multi-step implementations. Use this frequently to maintain an updated step-by-step plan.\n */\nexport interface WriteTodosParams {\n /** List of todos with their completion status. Add ALL of the applicable tasks to the list, so you don't forget to do anything. Try to order the todos the same way you will complete them. Do not mark todos as completed if you have not completed them yet! */\n todos: {\n /** Description of the task */\n task: string\n /** Whether the task is completed */\n completed: boolean\n }[]\n}\n\n/**\n * Get parameters type for a specific tool\n */\nexport type GetToolParams = ToolParamsMap[T]\n" export const utilTypesSource = "// ===== JSON Types =====\nexport type JSONValue =\n | null\n | string\n | number\n | boolean\n | JSONObject\n | JSONArray\n\nexport type JSONObject = { [key: string]: JSONValue }\n\nexport type JSONArray = JSONValue[]\n\n/**\n * JSON Schema definition (for prompt schema or output schema)\n */\nexport type JsonSchema = {\n type?:\n | 'object'\n | 'array'\n | 'string'\n | 'number'\n | 'boolean'\n | 'null'\n | 'integer'\n description?: string\n properties?: Record\n required?: string[]\n enum?: Array\n [k: string]: unknown\n}\nexport type JsonObjectSchema = JsonSchema & { type: 'object' }\n\n// ===== Data Content Types =====\nexport type DataContent = string | Uint8Array | ArrayBuffer | Buffer\n\n// ===== Provider Metadata Types =====\nexport type ProviderMetadata = Record>\n\n// ===== Content Part Types =====\nexport type TextPart = {\n type: 'text'\n text: string\n providerOptions?: ProviderMetadata\n}\n\nexport type ImagePart = {\n type: 'image'\n image: DataContent\n mediaType?: string\n providerOptions?: ProviderMetadata\n}\n\nexport type FilePart = {\n type: 'file'\n data: DataContent\n filename?: string\n mediaType: string\n providerOptions?: ProviderMetadata\n}\n\nexport type ReasoningPart = {\n type: 'reasoning'\n text: string\n providerOptions?: ProviderMetadata\n}\n\nexport type ToolCallPart = {\n type: 'tool-call'\n toolCallId: string\n toolName: string\n input: Record\n providerOptions?: ProviderMetadata\n providerExecuted?: boolean\n}\n\nexport type ToolResultOutput =\n | {\n type: 'json'\n value: JSONValue\n }\n | {\n type: 'media'\n data: string\n mediaType: string\n }\n\n// ===== Message Types =====\nexport type AuxiliaryMessageData = {\n providerOptions?: ProviderMetadata\n tags?: string[]\n\n /** @deprecated Use tags instead. */\n timeToLive?: 'agentStep' | 'userPrompt'\n /** @deprecated Use tags instead. */\n keepDuringTruncation?: boolean\n /** @deprecated Use tags instead. */\n keepLastTags?: string[]\n}\n\nexport type SystemMessage = {\n role: 'system'\n content: TextPart[]\n} & AuxiliaryMessageData\n\nexport type UserMessage = {\n role: 'user'\n content: (TextPart | ImagePart | FilePart)[]\n} & AuxiliaryMessageData\n\nexport type AssistantMessage = {\n role: 'assistant'\n content: (TextPart | ReasoningPart | ToolCallPart)[]\n} & AuxiliaryMessageData\n\nexport type ToolMessage = {\n role: 'tool'\n toolCallId: string\n toolName: string\n content: ToolResultOutput[]\n} & AuxiliaryMessageData\n\nexport type Message =\n | SystemMessage\n | UserMessage\n | AssistantMessage\n | ToolMessage\n\n// ===== MCP Server Types =====\n\n/**\n * MCP server configuration for stdio-based servers.\n *\n * Environment variables in `env` can be:\n * - A plain string value (hardcoded, e.g., `'production'`)\n * - A `$VAR_NAME` reference to read from local environment (e.g., `'$NOTION_TOKEN'`)\n *\n * The `$VAR_NAME` syntax reads from `process.env.VAR_NAME` at agent load time.\n * This keeps secrets out of your agent definitions - store them in `.env.local` instead.\n *\n * @example\n * ```typescript\n * env: {\n * // Read NOTION_TOKEN from local .env file\n * NOTION_TOKEN: '$NOTION_TOKEN',\n * // Read MY_API_KEY from local env, pass as API_KEY to MCP server\n * API_KEY: '$MY_API_KEY',\n * // Hardcoded value (non-secret)\n * NODE_ENV: 'production',\n * }\n * ```\n */\nexport type MCPConfig =\n | {\n type?: 'stdio'\n command: string\n args?: string[]\n env?: Record\n }\n | {\n type?: 'http' | 'sse'\n url: string\n params?: Record\n headers?: Record\n }\n\n// ============================================================================\n// Logger Interface\n// ============================================================================\nexport interface Logger {\n debug: (data: any, msg?: string) => void\n info: (data: any, msg?: string) => void\n warn: (data: any, msg?: string) => void\n error: (data: any, msg?: string) => void\n}\n" diff --git a/cli/src/hooks/helpers/send-message.ts b/cli/src/hooks/helpers/send-message.ts index e1a004919b..a7fd331eba 100644 --- a/cli/src/hooks/helpers/send-message.ts +++ b/cli/src/hooks/helpers/send-message.ts @@ -13,6 +13,7 @@ import { logger } from '../../utils/logger' import { getFileAttachmentContextMetadata } from '../../utils/pending-attachments' import { appendInterruptionNotice, + dropTransientCompactionBlocks, markPendingCompactionInterrupted, } from '../../utils/message-block-helpers' import { getUserMessage } from '../../utils/message-history' @@ -377,9 +378,13 @@ export const setupStreamingContext = (params: { // A compaction pass that was still running is terminated here, in the // same composed update: the SDK drops every post-abort event, so neither // `settled` nor `finish` arrives to end the pending state, and this is - // the last write before the turn's blocks are persisted. + // the last write before the turn's blocks are persisted. A transient + // (self-dismissing) card is dropped in the same pass, so an abort mid-hold + // cannot persist one either. return appendInterruptionNotice( - markPendingCompactionInterrupted(cancelledBlocks), + dropTransientCompactionBlocks( + markPendingCompactionInterrupted(cancelledBlocks), + ), ) }) updater.markComplete() diff --git a/cli/src/types/chat.ts b/cli/src/types/chat.ts index bf000afc93..4a4268fff0 100644 --- a/cli/src/types/chat.ts +++ b/cli/src/types/chat.ts @@ -318,6 +318,27 @@ export type CompactionContentBlock = { fitsBudget?: boolean shortfallTokens?: number escalated?: boolean + /** + * Monotonic 0..100 completion estimate of a PENDING pass, raised by each + * `context_compaction_progress` event of the producing run and never lowered, + * so an out-of-order or garbage percent cannot rewind the bar. A settled + * healthy pass is stamped 100 so the card reads as complete before it + * self-dismisses. Optional: an older CLI replaying a block that carries it + * simply ignores the field and renders exactly as it did before. + */ + progressPercent?: number + /** + * True only for a SETTLED pass with nothing worth keeping in scrollback: such + * a card self-dismisses shortly after reaching 100% and is dropped from the + * blocks at turn end (and on abort), so it never persists to + * chat-messages.json. Degraded outcomes — an emergency/mechanical trim, a + * request-time trim, a pass that did not fit the budget, an escalated pass, a + * low-yield streak — are NEVER transient, and neither is a `declined` or + * `interrupted` pass: those stay as a permanent warning card. Optional: an + * older CLI replaying a block that carries it ignores the field, which at + * worst leaves one extra completed-pass card in the transcript. + */ + transient?: boolean } /** @@ -384,6 +405,13 @@ export type CompactionNotice = { * pass has no run id to match against. */ pendingRunIds?: string[] + /** + * Monotonic 0..100 progress of the newest live pass, so the status chip can + * report real movement instead of a bare '⇲ compacting…'. Raised only while a + * pass is pending and absent once nothing is live, which is also what every + * notice produced before this field existed carries. + */ + progressPercent?: number } export type AskUserContentBlock = { diff --git a/cli/src/utils/__tests__/sdk-event-handlers.test.ts b/cli/src/utils/__tests__/sdk-event-handlers.test.ts index 13306df28a..7a0ec29f04 100644 --- a/cli/src/utils/__tests__/sdk-event-handlers.test.ts +++ b/cli/src/utils/__tests__/sdk-event-handlers.test.ts @@ -24,6 +24,7 @@ import { import type { Logger } from '@codebuff/common/types/contracts/logger' import type { PrintModeContextCompaction, + PrintModeContextCompactionProgress, PrintModeEvent, PrintModeJobUpdate, } from '@codebuff/common/types/print-mode' @@ -2775,4 +2776,390 @@ describe('sdk-event-handlers', () => { // persist alongside the terminal state. expect(compactionBlocks[0]).not.toHaveProperty('liveSessionId') }) + + // Payload fixtures for the live-progress cases below. They build the same + // shapes the compaction tests above dispatch inline; only the fields under + // test vary per case. + const compactionCategories = { + toolResults: { tokens: 10, percent: 10, messages: 1 }, + todos: { tokens: 10, percent: 10, messages: 1 }, + fileReads: { tokens: 20, percent: 20, messages: 2 }, + subagents: { tokens: 20, percent: 20, messages: 2 }, + userAssistantMessages: { tokens: 40, percent: 40, messages: 4 }, + } + + const startedEvent = (runId = 'root-run', contextTokens = 152_000) => ({ + type: 'context_compaction_status' as const, + state: 'started' as const, + runId, + ancestorRunIds: [], + contextTokens, + }) + + const progressEvent = ( + overrides: Partial = {}, + ): PrintModeContextCompactionProgress => ({ + type: 'context_compaction_progress', + // Root turn by default: empty lineage is what allows root-level card state. + runId: 'root-run', + ancestorRunIds: [], + percent: 40, + phase: 'summarizing', + ...overrides, + }) + + const compactionResultEvent = ( + overrides: Partial = {}, + ): PrintModeContextCompaction => ({ + type: 'context_compaction', + action: 'semantic_compaction', + runId: 'root-run', + ancestorRunIds: [], + before: { tokens: 152_000, messages: 20, categories: compactionCategories }, + after: { tokens: 60_000, messages: 8, categories: compactionCategories }, + removedCategories: [], + retainedKnowledgeMemory: true, + recovery: 'Resume from .', + ...overrides, + }) + + const compactionCards = ( + messages: ChatMessage[], + ): CompactionContentBlock[] => + (messages[0].blocks ?? []).filter( + (block): block is CompactionContentBlock => block.type === 'compaction', + ) + + /** + * The single card left behind by one announced root pass that reported + * `overrides` as its result. Each call gets a fresh handler context, so the + * transient gate below is asserted per outcome rather than across accumulated + * state. + */ + const settledCardFor = ( + overrides: Partial = {}, + ): CompactionContentBlock => { + const { ctx, getMessages } = createTestContext() + const handleEvent = createEventHandler(ctx) + dispatchValidEvent(handleEvent, startedEvent()) + dispatchValidEvent(handleEvent, progressEvent({ percent: 60 })) + dispatchValidEvent(handleEvent, compactionResultEvent(overrides)) + const cards = compactionCards(getMessages()) + expect(cards).toHaveLength(1) + return cards[0] + } + + test('context_compaction_progress raises the pending card progress for the root run', () => { + const { ctx, getMessages } = createTestContext() + const handleEvent = createEventHandler(ctx) + + dispatchValidEvent(handleEvent, startedEvent()) + // A freshly announced pass starts its bar at 0; only progress moves it. + expect(compactionCards(getMessages())[0]).toMatchObject({ + status: 'pending', + runId: 'root-run', + progressPercent: 0, + }) + + dispatchValidEvent(handleEvent, progressEvent({ percent: 45 })) + + const cards = compactionCards(getMessages()) + // The live card advances in place rather than gaining a sibling. + expect(cards).toHaveLength(1) + expect(cards[0]).toMatchObject({ + status: 'pending', + runId: 'root-run', + progressPercent: 45, + }) + }) + + test('a lower progress percent never rewinds the pending card', () => { + const { ctx, getMessages } = createTestContext() + const notices: Array = [] + let notice: CompactionNotice | null = null + ctx.streaming.setCompactionNotice = (update) => { + notice = update(notice) + notices.push(notice) + } + const handleEvent = createEventHandler(ctx) + + dispatchValidEvent(handleEvent, startedEvent()) + dispatchValidEvent(handleEvent, progressEvent({ percent: 60 })) + // Two producers report for one pass, so an out-of-order or duplicated + // percent is expected: every write takes the maximum already recorded. + dispatchValidEvent( + handleEvent, + progressEvent({ percent: 25, phase: 'applying' }), + ) + + expect(compactionCards(getMessages())[0]).toMatchObject({ + progressPercent: 60, + }) + expect(notices.at(-1)).toEqual({ + count: 0, + action: 'semantic_compaction', + degraded: false, + pending: true, + pendingRunIds: ['root-run'], + progressPercent: 60, + }) + }) + + test('a non-finite or out-of-range progress percent is clamped instead of thrown on', () => { + const { ctx, getMessages } = createTestContext() + const handleEvent = createEventHandler(ctx) + + dispatchValidEvent(handleEvent, startedEvent()) + + // Garbage percents are dispatched without the schema on purpose: a replayed + // or cross-version payload can carry them, and the percent is documented as + // best-effort telemetry, so they must degrade to a renderable number. + for (const percent of [ + Number.NaN, + Number.POSITIVE_INFINITY, + Number.NEGATIVE_INFINITY, + -20, + ]) { + expect(() => handleEvent(progressEvent({ percent }))).not.toThrow() + } + // Still the announced 0 rather than NaN or a negative bar width. + expect(compactionCards(getMessages())[0]).toMatchObject({ + progressPercent: 0, + }) + + dispatchValidEvent( + handleEvent, + progressEvent({ percent: 150, phase: 'applying' }), + ) + // An over-range estimate reads as a finished bar, never as 150%. + expect(compactionCards(getMessages())[0]).toMatchObject({ + progressPercent: 100, + }) + }) + + test('a nested progress event advances the shared chip without touching the root card', () => { + const { ctx, getMessages } = createTestContext() + const notices: Array = [] + let notice: CompactionNotice | null = null + ctx.streaming.setCompactionNotice = (update) => { + notice = update(notice) + notices.push(notice) + } + const handleEvent = createEventHandler(ctx) + + dispatchValidEvent(handleEvent, startedEvent()) + // A foreground subagent / inline agent loop renders no root-level card, so + // its progress has none to advance -- but the status chip is shared, so it + // still reports movement while a pass is live. + dispatchValidEvent( + handleEvent, + progressEvent({ + runId: 'child-run', + ancestorRunIds: ['root-run'], + agentId: 'child-agent', + percent: 55, + phase: 'analyzing', + }), + ) + + const cards = compactionCards(getMessages()) + expect(cards).toHaveLength(1) + expect(cards[0]).toMatchObject({ + status: 'pending', + runId: 'root-run', + progressPercent: 0, + }) + expect(notices.at(-1)).toEqual({ + count: 0, + action: 'semantic_compaction', + degraded: false, + pending: true, + pendingRunIds: ['root-run'], + progressPercent: 55, + }) + }) + + test('a progress event with no pass pending neither creates a notice nor revives a settled one', () => { + const { ctx, getMessages } = createTestContext() + const notices: Array = [] + let notice: CompactionNotice | null = null + ctx.streaming.setCompactionNotice = (update) => { + notice = update(notice) + notices.push(notice) + } + const handleEvent = createEventHandler(ctx) + + // Nothing was announced: progress is telemetry, so it must not invent a + // live chip or a card of its own. The notice is consulted and left null + // rather than being created as a pending one. + dispatchValidEvent( + handleEvent, + progressEvent({ percent: 30, phase: 'analyzing' }), + ) + expect(notices).toEqual([null]) + expect(compactionCards(getMessages())).toHaveLength(0) + + // A completed pass settles the notice, and its healthy card holds at 100%. + dispatchValidEvent(handleEvent, compactionResultEvent()) + const settledNotice: CompactionNotice = { + count: 1, + action: 'semantic_compaction', + degraded: false, + } + expect(notices.at(-1)).toEqual(settledNotice) + expect(compactionCards(getMessages())[0]).toMatchObject({ + status: 'complete', + progressPercent: 100, + transient: true, + }) + + // A late progress event must not reopen the notice as live, and has no + // pending card left to move. + dispatchValidEvent( + handleEvent, + progressEvent({ percent: 70, phase: 'applying' }), + ) + expect(notices.at(-1)).toEqual(settledNotice) + const cards = compactionCards(getMessages()) + expect(cards).toHaveLength(1) + expect(cards[0]).toMatchObject({ + status: 'complete', + progressPercent: 100, + transient: true, + }) + }) + + test('a healthy compaction result settles the card at 100% and marks it transient', () => { + // Nothing here needs the user's attention: the bar visibly finishes and the + // card is marked for self-dismissal so it never reaches the transcript. + expect( + settledCardFor({ compactionCount: 1, fitsBudget: true }), + ).toMatchObject({ + status: 'complete', + action: 'semantic_compaction', + progressPercent: 100, + transient: true, + }) + }) + + test('an emergency mechanical trim result is never marked transient', () => { + const trimmed = settledCardFor({ + action: 'mechanical_trim', + retainedKnowledgeMemory: false, + recovery: 'Re-gather exact constraints.', + }) + + expect(trimmed).toMatchObject({ + status: 'complete', + action: 'mechanical_trim', + }) + // A degraded outcome stays in scrollback as a permanent warning card, so it + // gets neither the self-dismissal flag nor a completed bar. + expect(trimmed).not.toHaveProperty('transient') + expect(trimmed).not.toHaveProperty('progressPercent') + // Control: the same path with a healthy result does mark the card, so the + // absence above is the degradation gate rather than a missing feature. + expect(settledCardFor()).toMatchObject({ + progressPercent: 100, + transient: true, + }) + }) + + test('a compaction result that does not fit the budget is never marked transient', () => { + const overBudget = settledCardFor({ + fitsBudget: false, + shortfallTokens: 12_400, + }) + + expect(overBudget).toMatchObject({ + status: 'complete', + action: 'semantic_compaction', + fitsBudget: false, + shortfallTokens: 12_400, + }) + // Still over budget is exactly what the user must act on. + expect(overBudget).not.toHaveProperty('transient') + expect(overBudget).not.toHaveProperty('progressPercent') + expect(settledCardFor({ fitsBudget: true })).toMatchObject({ + progressPercent: 100, + transient: true, + }) + }) + + test('a low-yield compaction streak is never marked transient', () => { + const thrashing = settledCardFor({ consecutiveNoProgressCompactions: 2 }) + + expect(thrashing).toMatchObject({ + status: 'complete', + action: 'semantic_compaction', + consecutiveNoProgressCompactions: 2, + }) + // Compaction that stopped reclaiming space keeps its warning card. + expect(thrashing).not.toHaveProperty('transient') + expect(thrashing).not.toHaveProperty('progressPercent') + // One low-yield pass is still below the streak threshold, so it settles as a + // healthy transient card. + expect( + settledCardFor({ consecutiveNoProgressCompactions: 1 }), + ).toMatchObject({ + progressPercent: 100, + transient: true, + }) + }) + + test('handleFinish drops a transient compaction card and terminates a still-pending one', () => { + const { ctx, getMessages } = createTestContext() + const handleEvent = createEventHandler(ctx) + + // Two root passes in one turn: the first is still live, the second + // completed healthily and is therefore transient. + dispatchValidEvent(handleEvent, startedEvent('run-live')) + dispatchValidEvent(handleEvent, startedEvent('run-done', 120_000)) + dispatchValidEvent( + handleEvent, + compactionResultEvent({ runId: 'run-done' }), + ) + expect( + compactionCards(getMessages()).map((card) => card.transient === true), + ).toEqual([false, true]) + + dispatchValidEvent(handleEvent, { type: 'finish', totalCost: 0 }) + + const cards = compactionCards(getMessages()) + // The renderer only HIDES a self-dismissing card; the turn boundary is what + // keeps it out of the persisted blocks, while the unfinished pass is + // terminated rather than deleted. + expect(cards).toHaveLength(1) + expect(cards[0]).toMatchObject({ + status: 'interrupted', + runId: 'run-live', + beforeTokens: 152_000, + }) + expect(cards[0]).not.toHaveProperty('liveSessionId') + }) + + test('a newly announced pass drops the previous transient card instead of stacking cards', () => { + const { ctx, getMessages } = createTestContext() + const handleEvent = createEventHandler(ctx) + + dispatchValidEvent(handleEvent, startedEvent('run-1')) + dispatchValidEvent(handleEvent, compactionResultEvent({ runId: 'run-1' })) + expect(compactionCards(getMessages())[0]).toMatchObject({ + status: 'complete', + transient: true, + }) + + // The previous pass's card may still be inside its render hold, so the next + // announced pass drops it rather than leaving it stacked underneath. + dispatchValidEvent(handleEvent, startedEvent('run-2', 120_000)) + + const cards = compactionCards(getMessages()) + expect(cards).toHaveLength(1) + expect(cards[0]).toMatchObject({ + status: 'pending', + runId: 'run-2', + beforeTokens: 120_000, + progressPercent: 0, + }) + }) }) diff --git a/cli/src/utils/__tests__/status-bar-chips.test.ts b/cli/src/utils/__tests__/status-bar-chips.test.ts index 4d140b2d5e..511d75ea5e 100644 --- a/cli/src/utils/__tests__/status-bar-chips.test.ts +++ b/cli/src/utils/__tests__/status-bar-chips.test.ts @@ -599,6 +599,126 @@ describe('selectStatusBarChips', () => { expect(degradedPending?.tone).toBe('warning') }) + // Same (widthSize, notice) shape as the `compactionAt` helper of the + // label/tone test above, hoisted so the progress cases below share it. + const compactionChipAt = ( + widthSize: 'xs' | 'sm' | 'md' | 'lg', + notice: NonNullable, + ) => + byId( + selectStatusBarChips({ + ...full, + widthSize, + terminalWidth: 400, + compactionNotice: notice, + }).chips, + ).compaction + + test('a live pass reports its progress percent at md and lg', () => { + const pendingWithProgress = { + count: 0, + action: 'semantic_compaction', + degraded: false, + pending: true, + progressPercent: 62, + } as const + + // The wide sizes have room to report real movement instead of an + // indefinite ellipsis. + expect(compactionChipAt('md', pendingWithProgress)?.label).toBe( + '⇲ compacting 62%', + ) + expect(compactionChipAt('lg', pendingWithProgress)?.label).toBe( + '⇲ compacting 62%', + ) + // A live pass still reads as in progress rather than as a failed one. + expect(compactionChipAt('lg', pendingWithProgress)?.tone).toBe('warning') + + // The narrow sizes have no room for the percent, so they keep their exact + // previous label. + expect(compactionChipAt('xs', pendingWithProgress)?.label).toBe('⇲ …') + expect(compactionChipAt('sm', pendingWithProgress)?.label).toBe('⇲ …') + }) + + test('a live percent is rounded and clamped to the renderable range', () => { + const labelFor = (progressPercent: number) => + compactionChipAt('lg', { + count: 0, + action: 'semantic_compaction', + degraded: false, + pending: true, + progressPercent, + })?.label + + expect(labelFor(62.6)).toBe('⇲ compacting 63%') + expect(labelFor(150)).toBe('⇲ compacting 100%') + }) + + test('a live pass with no usable percent falls back to the ellipsis label', () => { + const base = { + count: 0, + action: 'semantic_compaction', + degraded: false, + pending: true, + } as const + + // Absent, zero, negative and non-finite all mean "no usable estimate": the + // percent is best-effort telemetry, so the chip keeps its previous label. + expect(compactionChipAt('md', base)?.label).toBe('⇲ compacting…') + expect(compactionChipAt('lg', base)?.label).toBe('⇲ compacting…') + for (const progressPercent of [ + 0, + -5, + Number.NaN, + Number.POSITIVE_INFINITY, + ]) { + expect(compactionChipAt('md', { ...base, progressPercent })?.label).toBe( + '⇲ compacting…', + ) + expect(compactionChipAt('lg', { ...base, progressPercent })?.label).toBe( + '⇲ compacting…', + ) + } + + // Control: a usable percent does reach the label, so the fallbacks above + // are the no-estimate branch rather than a chip that never reports one. + const usable = compactionChipAt('lg', { ...base, progressPercent: 62 }) + expect(usable?.label).toBe('⇲ compacting 62%') + }) + + test('a settled notice ignores a carried progress percent', () => { + // Progress only ever describes a live pass, so a settled notice that still + // carries one keeps its exact count labels. + expect( + compactionChipAt('md', { + count: 2, + action: 'semantic_compaction', + degraded: false, + progressPercent: 62, + })?.label, + ).toBe('⇲ compacted ×2') + expect( + compactionChipAt('lg', { + count: 3, + action: 'mechanical_trim', + degraded: true, + progressPercent: 62, + })?.label, + ).toBe('⇲ trimmed ×3') + + // Control: the same percent on a PENDING notice is reported, so the + // settled labels above are unchanged by choice rather than by accident. + expect( + compactionChipAt('lg', { + count: 2, + action: 'semantic_compaction', + degraded: false, + pending: true, + progressPercent: 62, + })?.label, + ).toBe('⇲ compacting 62%') + }) + test('an idle run stops reporting a pending pass as live', () => { // The run aborted mid-compaction, so no settling event will ever arrive. // The chip must not keep claiming a compaction is running. diff --git a/cli/src/utils/message-block-helpers.ts b/cli/src/utils/message-block-helpers.ts index 1b80f53353..d4f1006266 100644 --- a/cli/src/utils/message-block-helpers.ts +++ b/cli/src/utils/message-block-helpers.ts @@ -937,6 +937,29 @@ export const markPendingCompactionInterrupted = ( return changed ? next : blocks } +/** + * Removes every `transient` compaction block. Such a card is a purely transient + * progress affordance — a healthy pass that reclaimed space and has nothing the + * user must act on — so it must never reach persistence: the renderer only + * HIDES it after a short hold, and this is what actually takes it out of state, + * composed into the same abort/turn-end block update as + * {@link markPendingCompactionInterrupted}. Degraded, declined and interrupted + * passes are never marked transient, so they are untouched here and keep their + * permanent warning card. + * + * Returns the ORIGINAL array reference when nothing was dropped so React skips + * a re-render. Root-level only: compaction blocks are never nested under an + * agent block. + */ +export const dropTransientCompactionBlocks = ( + blocks: ContentBlock[], +): ContentBlock[] => { + const next = blocks.filter( + (block) => !(block.type === 'compaction' && block.transient === true), + ) + return next.length === blocks.length ? blocks : next +} + /** * Recursively finds an agent block by ID and returns its agent type. * Returns undefined if not found. diff --git a/cli/src/utils/sdk-event-handlers.ts b/cli/src/utils/sdk-event-handlers.ts index 6ccde51f70..0969788c41 100644 --- a/cli/src/utils/sdk-event-handlers.ts +++ b/cli/src/utils/sdk-event-handlers.ts @@ -16,6 +16,7 @@ import { import { shouldHideAgent } from './constants' import { createAgentBlock, + dropTransientCompactionBlocks, extractPlanFromBuffer, extractSpawnAgentResultContent, findAgentTypeById, @@ -63,6 +64,7 @@ import type { Logger } from '@codebuff/common/types/contracts/logger' import type { PrintModeContextWindow, PrintModeContextCompaction, + PrintModeContextCompactionProgress, PrintModeContextCompactionStatus, PrintModeContextRequestTrim, PrintModeEvent as SDKEvent, @@ -1505,6 +1507,9 @@ const handleContextCompactionStatus = ( retainedKnowledgeMemory: false, recovery: '', categoryDeltas: [], + // A freshly announced pass starts its bar at 0; only + // `context_compaction_progress` moves it, and only upwards. + progressPercent: 0, ...(event.resolvedContextWindowTokens !== undefined && { resolvedContextWindowTokens: event.resolvedContextWindowTokens, }), @@ -1516,7 +1521,10 @@ const handleContextCompactionStatus = ( }), } state.message.updater.updateAiMessageBlocks((blocks) => [ - ...blocks, + // A previous pass's self-dismissing card is dropped rather than left to + // stack under the new live one: it is a transient progress affordance, + // and its renderer may still be inside its hold timer. + ...dropTransientCompactionBlocks(blocks), pendingBlock, ]) } @@ -1564,6 +1572,60 @@ const handleContextCompactionStatus = ( }) } +/** + * Bounds a producer-supplied percent to the renderable 0..100 whole range. The + * event contract calls `percent` a best-effort estimate, so a replayed or + * cross-version payload carrying a non-finite or out-of-range value degrades to + * a usable number here instead of reaching the renderer. + */ +const clampCompactionPercent = (percent: number): number => + Number.isFinite(percent) ? Math.min(100, Math.max(0, Math.round(percent))) : 0 + +/** + * Live progress inside an announced pass (`context_compaction_progress`). The + * pruner runs inline and is hidden, so without this the pending card would sit + * at a silent 0 for the whole pass. + * + * Two producers emit for one pass (the agent loop's milestones and the inline + * spawn path's activity ticks), so an out-of-order or duplicated percent is + * expected: every write takes the MAXIMUM of what is already recorded, which is + * what makes the bar monotonic no matter what order the events arrive in. The + * original array is returned when nothing changed so React skips a re-render. + * + * Card updates are root-scoped exactly like {@link handleContextCompactionStatus} + * (a nested run renders no root-level card, so it has none to advance), while + * the shared status-bar chip tracks progress for root and nested passes alike — + * but only while a pass is actually live, so a progress event can never create a + * notice or revive a settled one. + */ +const handleContextCompactionProgress = ( + state: EventHandlerState, + event: PrintModeContextCompactionProgress, +) => { + const percent = clampCompactionPercent(event.percent) + + if (isRootCompactionEvent(event)) { + state.message.updater.updateAiMessageBlocks((blocks) => { + const pendingIndex = findLastPendingCompactionIndex(blocks, event.runId) + if (pendingIndex === -1) return blocks + const block = blocks[pendingIndex] + if (block.type !== 'compaction') return blocks + const nextPercent = Math.max(block.progressPercent ?? 0, percent) + if (nextPercent === block.progressPercent) return blocks + const next = [...blocks] + next[pendingIndex] = { ...block, progressPercent: nextPercent } + return next + }) + } + + state.streaming.setCompactionNotice((previous) => { + if (!previous || previous.pending !== true) return previous + const nextPercent = Math.max(previous.progressPercent ?? 0, percent) + if (nextPercent === previous.progressPercent) return previous + return { ...previous, progressPercent: nextPercent } + }) +} + /** * Request-time emergency trim (`context_request_trim`): the SDK dropped * messages at dispatch time because the request still exceeded the @@ -1641,6 +1703,18 @@ const handleContextRequestTrim = ( })) } +/** + * A pass worth keeping in scrollback: anything the user may need to act on. The + * single decision site for {@link CompactionContentBlock.transient}, so a + * degraded outcome can never be dismissed as a transient progress affordance. + */ +const compactionResultIsDegraded = (block: CompactionContentBlock): boolean => + block.action === 'mechanical_trim' || + block.trimSource === 'request' || + block.fitsBudget === false || + block.escalated === true || + (block.consecutiveNoProgressCompactions ?? 0) >= 2 + const handleContextCompaction = ( state: EventHandlerState, event: PrintModeContextCompaction, @@ -1722,15 +1796,26 @@ const handleContextCompaction = ( ...(event.escalated !== undefined && { escalated: event.escalated }), } + // A healthy pass has nothing the user must act on, so its card is a purely + // transient progress affordance: stamped complete at 100% so the bar visibly + // finishes, then self-dismissed by the renderer and dropped from the blocks at + // turn end so it never persists. A degraded pass gets neither field and stays + // in scrollback as a permanent warning card. + const settledBlock: CompactionContentBlock = compactionResultIsDegraded( + resultBlock, + ) + ? resultBlock + : { ...resultBlock, progressPercent: 100, transient: true } + // The live pending card settles into the result in place; with no pending // card for this run (e.g. a mechanical trim with no preceding start, or a // subagent result while the root card is live) the result appends, which is // the pre-existing behavior. state.message.updater.updateAiMessageBlocks((blocks) => { const pendingIndex = findLastPendingCompactionIndex(blocks, event.runId) - if (pendingIndex === -1) return [...blocks, resultBlock] + if (pendingIndex === -1) return [...blocks, settledBlock] const next = [...blocks] - next[pendingIndex] = resultBlock + next[pendingIndex] = settledBlock return next }) @@ -1778,7 +1863,13 @@ const handleFinish = (state: EventHandlerState, event: PrintModeFinish) => { // result. Kept separate from the recursive agent/tool settling below, which // walks nested blocks. A user abort never reaches this handler: the abort // listener in hooks/helpers/send-message.ts applies the same rewrite. - const rootBlocks = markPendingCompactionInterrupted(blocks) + // + // Transient cards are dropped in the same composed update: the renderer only + // HIDES a self-dismissing card, so this is what keeps it out of the turn's + // persisted transcript. + const rootBlocks = dropTransientCompactionBlocks( + markPendingCompactionInterrupted(blocks), + ) const settledBlocks = settleOrphanedForegroundAgents(rootBlocks, settledIds) const summary = computeCompletionSummary(settledBlocks) if (!summary) return settledBlocks @@ -1921,6 +2012,9 @@ export const createEventHandler = .with({ type: 'context_compaction' }, (e) => handleContextCompaction(state, e), ) + .with({ type: 'context_compaction_progress' }, (e) => + handleContextCompactionProgress(state, e), + ) .with({ type: 'context_compaction_status' }, (e) => handleContextCompactionStatus(state, e), ) diff --git a/cli/src/utils/status-bar-chips.ts b/cli/src/utils/status-bar-chips.ts index 5b8476733b..6c42cb9fc4 100644 --- a/cli/src/utils/status-bar-chips.ts +++ b/cli/src/utils/status-bar-chips.ts @@ -341,13 +341,28 @@ export const buildContextLabel = ( * form at 'md'/'lg' that distinguishes a semantic compaction from an emergency * mechanical trim. A pass that is still running reports the live state instead * of a count, which may still be 0 when nothing has completed yet. + * + * A live pass with a usable `progressPercent` reports it at 'md'/'lg' ('⇲ + * compacting 62%') so the chip shows real movement rather than an indefinite + * ellipsis. The percent is best-effort telemetry, so a missing, non-finite or + * zero value falls back to the previous ellipsis label, and the narrow sizes — + * which have no room for it — keep their exact previous labels. */ const buildCompactionLabel = ( widthSize: StatusBarWidthSize, - notice: Pick, + notice: Pick< + CompactionNotice, + 'count' | 'action' | 'pending' | 'progressPercent' + >, ): string => { const narrow = widthSize === 'xs' || widthSize === 'sm' - if (notice.pending) return narrow ? '⇲ …' : '⇲ compacting…' + if (notice.pending) { + if (narrow) return '⇲ …' + const percent = notice.progressPercent + return typeof percent === 'number' && Number.isFinite(percent) && percent > 0 + ? `⇲ compacting ${Math.min(100, Math.round(percent))}%` + : '⇲ compacting…' + } if (narrow) return `⇲ ${notice.count}` const verb = notice.action === 'mechanical_trim' ? 'trimmed' : 'compacted' return `⇲ ${verb} ×${notice.count}` @@ -499,6 +514,9 @@ export function selectStatusBarChips(input: SelectStatusBarChipsInput): { count: compactionNotice.count, action: compactionNotice.action, pending: compactionPending, + ...(compactionNotice.progressPercent !== undefined && { + progressPercent: compactionNotice.progressPercent, + }), }), // A live pass reads as in-progress, not as a failed one: a degraded // earlier pass only tones the chip red once it has settled. diff --git a/common/src/__tests__/images.test.ts b/common/src/__tests__/images.test.ts new file mode 100644 index 0000000000..f817a73c7c --- /dev/null +++ b/common/src/__tests__/images.test.ts @@ -0,0 +1,114 @@ +import { describe, expect, test } from 'bun:test' + +import { + IMAGE_EXTENSION_TO_MIME, + detectImageMediaTypeFromBytes, +} from '../constants/images' + +/** Buffer holding `signature` at `offset`, zero-padded before it. */ +function withSignature(offset: number, signature: number[]): Buffer { + const buffer = Buffer.alloc(offset + signature.length) + Buffer.from(signature).copy(buffer, offset) + return buffer +} + +describe('detectImageMediaTypeFromBytes', () => { + test('recognizes every supported signature', () => { + expect( + detectImageMediaTypeFromBytes( + withSignature(0, [0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]), + ), + ).toBe('image/png') + expect( + detectImageMediaTypeFromBytes(withSignature(0, [0xff, 0xd8, 0xff])), + ).toBe('image/jpeg') + // "GIF8" covers GIF87a and GIF89a alike. + expect(detectImageMediaTypeFromBytes(Buffer.from('GIF89a', 'ascii'))).toBe( + 'image/gif', + ) + expect(detectImageMediaTypeFromBytes(Buffer.from('GIF87a', 'ascii'))).toBe( + 'image/gif', + ) + expect(detectImageMediaTypeFromBytes(Buffer.from('BMxx', 'ascii'))).toBe( + 'image/bmp', + ) + }) + + test('requires the WEBP form tag at offset 8, not just the RIFF prefix', () => { + const webp = Buffer.concat([ + Buffer.from('RIFF', 'ascii'), + Buffer.alloc(4), + Buffer.from('WEBP', 'ascii'), + ]) + expect(detectImageMediaTypeFromBytes(webp)).toBe('image/webp') + + // RIFF also fronts WAV and AVI, so a prefix-only check would misreport + // audio as an image. + const wav = Buffer.concat([ + Buffer.from('RIFF', 'ascii'), + Buffer.alloc(4), + Buffer.from('WAVE', 'ascii'), + ]) + expect(detectImageMediaTypeFromBytes(wav)).toBeNull() + }) + + test('accepts TIFF in both byte orders', () => { + expect( + detectImageMediaTypeFromBytes(withSignature(0, [0x49, 0x49, 0x2a, 0x00])), + ).toBe('image/tiff') + expect( + detectImageMediaTypeFromBytes(withSignature(0, [0x4d, 0x4d, 0x00, 0x2a])), + ).toBe('image/tiff') + }) + + test('returns null for short buffers instead of throwing', () => { + // A truncated PNG header shares a prefix with the real signature, so an + // unguarded read would compare against undefined bytes. + expect(detectImageMediaTypeFromBytes(Buffer.alloc(0))).toBeNull() + expect( + detectImageMediaTypeFromBytes(Buffer.from([0x89, 0x50, 0x4e])), + ).toBeNull() + expect(detectImageMediaTypeFromBytes(Buffer.from('RIFF', 'ascii'))).toBeNull() + expect(detectImageMediaTypeFromBytes(Buffer.from([0x49, 0x49]))).toBeNull() + }) + + test('returns null for plain text wearing no signature', () => { + expect( + detectImageMediaTypeFromBytes( + Buffer.from('just some plain text, definitely not an image'), + ), + ).toBeNull() + }) + + test('accepts a Uint8Array as well as a Buffer', () => { + expect( + detectImageMediaTypeFromBytes( + new Uint8Array([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]), + ), + ).toBe('image/png') + }) + + test('only ever returns MIME strings the extension map already publishes', () => { + // The two sources must not disagree: a sniffed type that is not in the map + // would be announced to a provider under a name the extension path can + // never produce. + const published = new Set(Object.values(IMAGE_EXTENSION_TO_MIME)) + const detected = [ + withSignature(0, [0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]), + withSignature(0, [0xff, 0xd8, 0xff]), + Buffer.from('GIF89a', 'ascii'), + Buffer.from('BMxx', 'ascii'), + Buffer.concat([ + Buffer.from('RIFF', 'ascii'), + Buffer.alloc(4), + Buffer.from('WEBP', 'ascii'), + ]), + withSignature(0, [0x49, 0x49, 0x2a, 0x00]), + ].map((bytes) => detectImageMediaTypeFromBytes(bytes)) + + expect(detected.every((mime) => mime !== null)).toBe(true) + for (const mime of detected) { + expect(published.has(mime!)).toBe(true) + } + }) +}) diff --git a/common/src/actions.ts b/common/src/actions.ts index bd3c272d4c..0d6e9e6524 100644 --- a/common/src/actions.ts +++ b/common/src/actions.ts @@ -203,7 +203,6 @@ type ServerActionToolCallRequest = { requestId: string toolName: string input: unknown - timeout?: number mcpConfig?: MCPConfig } diff --git a/common/src/constants/images.ts b/common/src/constants/images.ts index f430e64770..a2788a91e3 100644 --- a/common/src/constants/images.ts +++ b/common/src/constants/images.ts @@ -37,6 +37,61 @@ export function getImageMimeType(ext: string): string | null { return IMAGE_EXTENSION_TO_MIME[ext.toLowerCase()] ?? null } +/** + * Detect an image MIME type from raw bytes by matching magic-number signatures. + * Returns null when the bytes do not match a supported image format. + */ +export function detectImageMediaTypeFromBytes( + bytes: Uint8Array | Buffer, +): string | null { + const hasPrefix = (signature: number[], offset = 0): boolean => { + if (bytes.length < offset + signature.length) { + return false + } + for (let i = 0; i < signature.length; i++) { + if (bytes[offset + i] !== signature[i]) { + return false + } + } + return true + } + const hasAscii = (text: string, offset = 0): boolean => { + if (bytes.length < offset + text.length) { + return false + } + for (let i = 0; i < text.length; i++) { + if (bytes[offset + i] !== text.charCodeAt(i)) { + return false + } + } + return true + } + + if (hasPrefix([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a])) { + return 'image/png' + } + if (hasPrefix([0xff, 0xd8, 0xff])) { + return 'image/jpeg' + } + // "GIF8" covers both GIF87a and GIF89a. + if (hasAscii('GIF8')) { + return 'image/gif' + } + if (hasAscii('BM')) { + return 'image/bmp' + } + if (hasAscii('RIFF') && hasAscii('WEBP', 8)) { + return 'image/webp' + } + if ( + hasPrefix([0x49, 0x49, 0x2a, 0x00]) || + hasPrefix([0x4d, 0x4d, 0x00, 0x2a]) + ) { + return 'image/tiff' + } + return null +} + /** * Image extensions as a regex alternation pattern (without dots) * e.g., "jpg|jpeg|png|webp|gif|bmp|tiff|tif" diff --git a/common/src/templates/initial-agents-dir/types/agent-definition.ts b/common/src/templates/initial-agents-dir/types/agent-definition.ts index 4de8028a6d..69e3ecd332 100644 --- a/common/src/templates/initial-agents-dir/types/agent-definition.ts +++ b/common/src/templates/initial-agents-dir/types/agent-definition.ts @@ -39,16 +39,6 @@ export interface AgentDefinition { */ model?: ModelName - /** - * Optional wall-clock timeout in milliseconds for a single execution of this - * agent as a subagent. When set, executeSubagent uses this as the deadline - * (overridable per-spawn via spawn_agents' timeout_seconds). Undefined falls - * back to the shared DEFAULT_SUBAGENT_TIMEOUT_MS, which is -1 (disabled): by - * default there is no wall-clock timeout, so long-running agents run to - * completion. Set a positive value to opt this agent into a wall-clock bound. - */ - defaultTimeoutMs?: number - /** Maximum subagent nesting depth. Defaults to the runtime limit. */ maxSpawnDepth?: number diff --git a/common/src/templates/initial-agents-dir/types/tools.ts b/common/src/templates/initial-agents-dir/types/tools.ts index 172984f5bd..32cd35f189 100644 --- a/common/src/templates/initial-agents-dir/types/tools.ts +++ b/common/src/templates/initial-agents-dir/types/tools.ts @@ -794,7 +794,7 @@ export interface RunTerminalCommandParams { detach?: boolean /** The working directory to run the command in. Default is the project root. */ cwd?: string - /** Set to -1 for no timeout. Does not apply for BACKGROUND commands. Default 30 */ + /** Wall-clock bound in seconds for SYNC commands. Omit or use -1 for no timeout (the default). Does not apply to BACKGROUND commands. */ timeout_seconds?: number /** Runtime-managed background job owner; agents must omit. */ owner?: { @@ -945,15 +945,13 @@ export interface SpawnAgentsParams { constraints?: string[] } | Record - /** Optional wall-clock deadline seconds; omit or -1 for none. Agent defaultTimeoutMs still applies when set. */ - timeout_seconds?: number /** Parameters object for the agent */ params?: { /** Terminal command to run (basher, tmux-cli) */ command?: string /** What information from the command output is desired (basher) */ what_to_summarize?: string - /** Timeout for command. Set to -1 for no timeout. Default 30 (basher) */ + /** Timeout for command in seconds. Omit or -1 for no timeout (default). */ timeout_seconds?: number /** Save full command output to a /tmp log and extract failure lines for long SYNC command output (basher) */ save_full_log?: boolean diff --git a/common/src/tools/params/tool/run-terminal-command.ts b/common/src/tools/params/tool/run-terminal-command.ts index 7fc851c2d7..03a750c53a 100644 --- a/common/src/tools/params/tool/run-terminal-command.ts +++ b/common/src/tools/params/tool/run-terminal-command.ts @@ -102,10 +102,10 @@ const inputSchema = z ), timeout_seconds: z .number() - .default(30) + .default(-1) .optional() .describe( - `Set to -1 for no timeout. Does not apply for BACKGROUND commands. Default 30`, + `Wall-clock bound in seconds for SYNC commands. Omit or use -1 for no timeout (the default). Does not apply to BACKGROUND commands.`, ), owner: z .object({ diff --git a/common/src/tools/params/tool/spawn-agents.ts b/common/src/tools/params/tool/spawn-agents.ts index 9bee457686..1d805e15e7 100644 --- a/common/src/tools/params/tool/spawn-agents.ts +++ b/common/src/tools/params/tool/spawn-agents.ts @@ -73,12 +73,6 @@ const spawnAgentEntryFields = { .describe( 'Optional structured handoff; additive — non-consumers still get prompt/params.', ), - timeout_seconds: z - .number() - .optional() - .describe( - 'Optional wall-clock deadline seconds; omit or -1 for none. Agent defaultTimeoutMs still applies when set.', - ), params: z .preprocess( coerceToObject, @@ -99,7 +93,7 @@ const spawnAgentEntryFields = { .number() .optional() .describe( - 'Timeout for command. Set to -1 for no timeout. Default 30 (basher)', + 'Timeout for command in seconds. Omit or -1 for no timeout (default).', ), save_full_log: z .boolean() diff --git a/common/src/types/agent-handoff.ts b/common/src/types/agent-handoff.ts index 22bebadc8b..92020cc5e9 100644 --- a/common/src/types/agent-handoff.ts +++ b/common/src/types/agent-handoff.ts @@ -169,6 +169,23 @@ export const agentReceiptSchema = z }) .strict() .optional(), + /** + * Observational telemetry about how much context the finished child + * actually used: the tokens it ended on, the window it ran against, and + * that ratio as a percent (clamped, so a child that overran its declared + * window still produces a valid receipt). Purely informational — the + * parent may use it to size later delegations (more, narrower shards when + * a child came back near its window). Optional, so every existing receipt + * stays valid and no consumer migration is required. + */ + contextUsage: z + .object({ + tokens: z.number().int().min(0), + windowTokens: z.number().int().min(0).optional(), + percentOfWindow: z.number().int().min(0).max(100).optional(), + compactionCount: z.number().int().min(0).optional(), + }) + .optional(), }) .strict() diff --git a/common/src/types/agent-template.ts b/common/src/types/agent-template.ts index 0d2526e93e..545ee31eda 100644 --- a/common/src/types/agent-template.ts +++ b/common/src/types/agent-template.ts @@ -138,15 +138,6 @@ export type AgentTemplate< * Undefined = no cap. */ maxTokensPerTurn?: number - /** - * Optional wall-clock timeout in milliseconds for a single execution of this - * agent as a subagent. When set, executeSubagent uses this as the deadline - * (overridable per-spawn via spawn_agents' timeout_seconds). Undefined falls - * back to the shared DEFAULT_SUBAGENT_TIMEOUT_MS, which is -1 (disabled): by - * default there is no wall-clock timeout, so long-running agents run to - * completion. Set a positive value to opt this agent into a wall-clock bound. - */ - defaultTimeoutMs?: number /** * Optional maximum nesting depth for this agent when spawned as a subagent. diff --git a/common/src/types/dynamic-agent-template.ts b/common/src/types/dynamic-agent-template.ts index 8be7585aa6..97be9e995b 100644 --- a/common/src/types/dynamic-agent-template.ts +++ b/common/src/types/dynamic-agent-template.ts @@ -181,10 +181,6 @@ export const DynamicAgentDefinitionSchema = z.object({ // turn if a single step's total input tokens exceed this threshold. maxTokensPerTurn: z.number().int().positive().optional(), - // Optional wall-clock timeout (ms) for a single subagent execution. -1 - // disables the timeout. Undefined falls back to the shared - // DEFAULT_SUBAGENT_TIMEOUT_MS, which is -1 (disabled) by default. - defaultTimeoutMs: z.number().optional(), maxSpawnDepth: z.number().int().min(0).optional(), // Tools and spawnable agents diff --git a/common/src/types/print-mode.ts b/common/src/types/print-mode.ts index 80bad6f811..5047fe91df 100644 --- a/common/src/types/print-mode.ts +++ b/common/src/types/print-mode.ts @@ -132,9 +132,8 @@ export const printModeSubagentFinishSchema = z.object({ prompt: z.string().optional(), spawnToolCallId: z.string().optional(), spawnIndex: z.number().int().nonnegative().optional(), - // Present when the subagent finished due to an error (e.g. wall-clock - // timeout) rather than completing normally. Lets the UI distinguish a - // failed finish from a successful one. + // Present when the subagent finished due to an error rather than completing + // normally. Lets the UI distinguish a failed finish from a successful one. error: z.string().optional(), }) export type PrintModeSubagentFinish = z.infer< @@ -399,6 +398,65 @@ export type PrintModeContextCompactionStatus = z.infer< typeof printModeContextCompactionStatusSchema > +/** + * Live progress WITHIN an announced context-compaction pass. ADDITIVE, + * non-breaking public-contract change: this is a NEW member of the + * {@link printModeEventSchema} discriminated union — no existing event variant + * is removed, renamed, or retyped, so no consumer migration or deprecation is + * required. In particular the `state` enum of + * {@link printModeContextCompactionStatusSchema} is deliberately NOT widened to + * carry progress: an added enum member would break a consumer that switches + * exhaustively over it, while an unknown event `type` is already contractually + * a no-op (see the forward-compatibility clause below). + * + * PRODUCER CONTRACT. This event only ever appears BETWEEN a + * `context_compaction_status` `started` and its matching `settled` for the SAME + * `runId`. The runtime emits it from deterministic milestones inside the + * announced pass — `analyzing` as soon as the pass is announced, `summarizing` + * per unit of observed pruner activity, `applying` once the pass has returned — + * and emits nothing at all outside an announced pass, so a suppressed or + * never-triggered iteration produces no progress events just as it produces + * neither half of the status pair. Correlation mirrors + * {@link printModeContextCompactionStatusSchema}: `runId` identifies the + * emitting run and `ancestorRunIds` is its lineage (empty ONLY for the root + * run); both are forwarded verbatim by every hop, while `agentId` is a display + * hint only, because the `spawn_agents` forwarding path rewrites it and at + * nesting depth >= 2 the delivered value names the nearest forwarding child + * rather than the emitter. + * + * `percent` is a BEST-EFFORT monotonic 0..100 estimate of how far the announced + * pass has progressed, never a guarantee. The pruner reports no total, two + * producers emit for the same pass (the agent loop and the inline spawn path), + * and events can be dropped or replayed, so a consumer MUST clamp for itself — + * `Math.max` against the last value it rendered, bounded to 0..100 — rather + * than trusting the sequence to arrive ordered or in range. `percent: 100` is + * NOT a claim that any space was reclaimed: the terminal + * {@link printModeContextCompactionSchema} result remains the ONLY signal that + * context was actually reclaimed, exactly as it was before this variant + * existed. Progress is telemetry, and no part of the compaction contract is + * gated on it. + * + * Forward-compatibility contract for `handleEvent` consumers: an exhaustive + * `switch`/match over `event.type` should treat unknown variants as no-ops + * (the SDK's own default handler only branches on `error`, and the CLI handler + * uses a catch-all `.otherwise`). Consumers may ignore this variant entirely + * and keep their previous behavior. + */ +export const printModeContextCompactionProgressSchema = z.object({ + type: z.literal('context_compaction_progress'), + runId: z.string(), + ancestorRunIds: z.string().array(), + agentId: z.string().optional(), + /** Monotonic 0..100 completion estimate for the announced pass. */ + percent: z.number(), + phase: z.enum(['analyzing', 'summarizing', 'applying']), + contextTokens: z.number().optional(), + targetBudgetTokens: z.number().optional(), +}) +export type PrintModeContextCompactionProgress = z.infer< + typeof printModeContextCompactionProgressSchema +> + /** * Request-time emergency context trim. ADDITIVE, non-breaking public-contract * change: this is a NEW member of the {@link printModeEventSchema} @@ -529,6 +587,7 @@ export const printModeEventSchema = z.discriminatedUnion('type', [ printModeToolStartSchema, printModeContextCompactionSchema, + printModeContextCompactionProgressSchema, printModeContextCompactionStatusSchema, printModeContextRequestTrimSchema, printModeContextWindowSchema, diff --git a/docs/agents-and-tools.md b/docs/agents-and-tools.md index bb384f20c8..baf00dc99b 100644 --- a/docs/agents-and-tools.md +++ b/docs/agents-and-tools.md @@ -106,7 +106,7 @@ diagnostics (`file`, range, severity, code, message, command, source) while the original bounded stdout/stderr remains available for recovery. - Automated security/test/doc auxiliary agents have explicit lifecycle handling. Their done flags are written only after successful completion; crashes and blocking security verdicts persist as blockers. Test/doc writers run automatically only when the user request explicitly includes those deliverables, and mixed-package test targets are routed to package-specific commands. -- Productive agent steps and subagent wall-clock duration are unlimited by default. A repeated-step watchdog stops identical no-progress loops, while cancellation, explicitly configured subagent deadlines, cost/token budgets, spawn-depth limits, and context compaction remain independent safeguards. Users may still configure a positive `maxAgentSteps`, agent-template `defaultTimeoutMs`, or per-spawn `timeout_seconds` cap; `-1` explicitly selects unlimited mode. Reviewer crashes retry once; repeated crashes require the explicit user phrase `bypass reviewer gate` before finalization can continue. +- Productive agent steps, subagent duration, file mutations, configured file-change hooks, and terminal commands are all unbounded in wall-clock time by default. The remaining safeguards are user cancellation, the repeated-step no-progress watchdog, cost/token budgets, spawn-depth limits, and context compaction; observational poll bounds (`check_job`, `check_background_agent`) and network/lock timeouts are unchanged. Reviewer crashes retry once; repeated crashes require the explicit user phrase `bypass reviewer gate` before finalization can continue. - Root-orchestrator mutating/control gate operations such as Git-status observation, file-change hooks, and structural inventory are model-hidden programmatic tools. Their results are injected when needed after edits, so the harness remains active without paying for those schemas on every provider request. The orchestrator must **not** treat basher typechecks or `run_targeted_validation` as gate substitutes — only the runtime-owned hooks→reviewer cycle clears the gate. The read-only `get_change_review_bundle` tool remains model-visible so an orchestrator can refresh a stale reviewer snapshot after compaction. Fresh greetings and simple gratitude prompts take a narrow conversational fast path only when no pending work or reviewer blocker exists. **Pattern-specific agents** are intentionally **excluded** from `spawnableAgents` because they have a narrow contract that only makes sense within a specific workflow pattern. They are spawned by the pattern flow itself, not by the orchestrator: @@ -1286,9 +1286,6 @@ Input fields: - `handoff` (object, optional) — structured handoff payload forwarded to the child spawn entry. - `background` (boolean, optional) — launches the child as a background job. -- `timeout_seconds` (number, optional) — opt-in per-spawn wall-clock deadline. - Omit it or pass `-1` for no deadline. A positive agent-template - `defaultTimeoutMs` remains an explicit configuration override. `spawn_agents.agents` also performs bounded repair for one- or double-stringified arrays and stringified object entries. Malformed or diff --git a/packages/agent-runtime/src/__tests__/loop-agent-steps-abort.test.ts b/packages/agent-runtime/src/__tests__/loop-agent-steps-abort.test.ts index 53e384fa8f..8418822666 100644 --- a/packages/agent-runtime/src/__tests__/loop-agent-steps-abort.test.ts +++ b/packages/agent-runtime/src/__tests__/loop-agent-steps-abort.test.ts @@ -17,12 +17,11 @@ import type { AgentState } from '@codebuff/common/types/session-state' * loopAgentSteps. These guard the `signal.aborted` checkpoints at * run-agent-step.ts lines ~889, ~1104, and ~1369. * - * The wall-clock timeout fix in spawn-agent-utils.ts relies on these - * checkpoints: when executeSubagent's timeout controller aborts the - * combined signal, loopAgentSteps must actually exit (as 'cancelled') so - * the stuck LLM stream is cancelled rather than orphaned. If a future - * refactor removes these checks, the timeout will reject the outer - * promise but the inner stream keeps running — these tests catch that. + * These checkpoints are what make user/parent cancellation real: when the + * caller aborts the signal a subagent runs on, loopAgentSteps must actually + * exit (as 'cancelled') so the in-flight LLM stream is cancelled rather than + * orphaned. If a future refactor removes these checks, an abort would resolve + * the caller while the inner stream keeps running — these tests catch that. */ describe('loopAgentSteps abort signal handling', () => { let agentTemplate: AgentTemplate diff --git a/packages/agent-runtime/src/__tests__/loop-agent-steps.test.ts b/packages/agent-runtime/src/__tests__/loop-agent-steps.test.ts index 36cfb3deed..9d8d4ec04f 100644 --- a/packages/agent-runtime/src/__tests__/loop-agent-steps.test.ts +++ b/packages/agent-runtime/src/__tests__/loop-agent-steps.test.ts @@ -825,6 +825,104 @@ describe('loopAgentSteps', () => { }) }) + // `context_compaction_progress` is best-effort telemetry emitted from + // deterministic milestones INSIDE an announced pass. The next cases pin the + // two invariants a consumer relies on: percents never rewind, and no emission + // ever falls outside a started/settled pair of the same run. + it('emits monotonic compaction progress strictly inside an announced pass', async () => { + setup() + const events: any[] = [] + agentState.messageHistory = [ + userMessage('small-window evidence '.repeat(8_000)), + userMessage('Continue from the retained goal.'), + ] + agentTemplate.handleSteps = + contextPruner.handleSteps as AgentTemplate['handleSteps'] + + await loopAgentSteps({ + ...baseParams, + agentState, + resolveModelContextWindow: mock(() => 32_000), + localAgentTemplates: { 'test-agent': agentTemplate }, + onResponseChunk: (event) => events.push(event), + }) + + const statusEvents = events.filter( + (event) => event.type === 'context_compaction_status', + ) + const started = statusEvents.find((event) => event.state === 'started') + const settled = statusEvents.find((event) => event.state === 'settled') + expect(started).toBeDefined() + expect(settled).toBeDefined() + + const progress = events.filter( + (event) => event.type === 'context_compaction_progress', + ) + // Both deterministic milestones: analysis as soon as the pass is announced, + // application once the step that owns the pruner has returned. + expect(progress.length).toBeGreaterThanOrEqual(2) + expect(progress[0]).toMatchObject({ + phase: 'analyzing', + percent: 20, + runId: started.runId, + agentId: 'test-agent-id', + ancestorRunIds: [], + contextTokens: expect.any(Number), + targetBudgetTokens: 8_400, + }) + expect(progress.at(-1)).toMatchObject({ phase: 'applying', percent: 90 }) + + const startedIndex = events.indexOf(started) + const settledIndex = events.indexOf(settled) + let previousPercent = 0 + for (const event of progress) { + const index = events.indexOf(event) + expect(index).toBeGreaterThan(startedIndex) + expect(index).toBeLessThan(settledIndex) + expect(event.runId).toBe(started.runId) + expect(event.percent).toBeGreaterThanOrEqual(previousPercent) + expect(event.percent).toBeLessThanOrEqual(100) + previousPercent = event.percent + } + }) + + it('emits no compaction progress for an iteration that announces no pass', async () => { + setup() + const events: any[] = [] + // A huge window keeps the request below the semantic trigger, so nothing is + // announced and a consumer must see no progress for a pass that never ran. + agentState.messageHistory = [userMessage('old evidence '.repeat(4_000))] + agentTemplate.handleSteps = function* () { + yield { + toolName: 'set_messages', + input: { + messages: [ + userMessage( + '\nPinned structured knowledge memory.\n', + ), + ], + }, + includeToolCall: false, + } + yield 'STEP' + } as () => StepGenerator + + await loopAgentSteps({ + ...baseParams, + agentState, + resolveModelContextWindow: mock(() => 1_000_000), + localAgentTemplates: { 'test-agent': agentTemplate }, + onResponseChunk: (event) => events.push(event), + }) + + expect( + events.filter((event) => event.type === 'context_compaction_status'), + ).toHaveLength(0) + expect( + events.filter((event) => event.type === 'context_compaction_progress'), + ).toHaveLength(0) + }) + it('emits a recovery-rich event when emergency mechanical trim is required', async () => { setup() const events: any[] = [] @@ -986,6 +1084,45 @@ describe('loopAgentSteps', () => { expect(result.agentState.suppressSemanticCompaction).toBe(true) }) + it('emits compaction progress only while an announced pass is unsettled', async () => { + setup() + const events: any[] = [] + // The suppression fixture drives several over-trigger iterations, only the + // first two of which announce a pass. Every later (suppressed) iteration + // must contribute no progress at all. + seedZeroReclaimAnnouncedPasses() + + await loopAgentSteps({ + ...baseParams, + agentState, + resolveModelContextWindow: mock(() => 64_000), + localAgentTemplates: { 'test-agent': agentTemplate }, + onResponseChunk: (event) => events.push(event), + }) + + const started = events.filter( + (event) => + event.type === 'context_compaction_status' && event.state === 'started', + ) + expect(started.length).toBeGreaterThanOrEqual(2) + + // Walk the stream: a progress event may only appear while an announced pass + // of the SAME run is still unsettled. + const liveRunIds = new Set() + let progressCount = 0 + for (const event of events) { + if (event.type === 'context_compaction_status') { + if (event.state === 'started') liveRunIds.add(event.runId) + else liveRunIds.delete(event.runId) + continue + } + if (event.type !== 'context_compaction_progress') continue + progressCount++ + expect(liveRunIds.has(event.runId)).toBe(true) + } + expect(progressCount).toBeGreaterThanOrEqual(started.length) + }) + it('leaves semantic compaction unsuppressed when a pass genuinely shrinks history', async () => { setup() const events: any[] = [] diff --git a/packages/agent-runtime/src/__tests__/prompts-schema-handling.test.ts b/packages/agent-runtime/src/__tests__/prompts-schema-handling.test.ts index 6aea429656..bff9223eae 100644 --- a/packages/agent-runtime/src/__tests__/prompts-schema-handling.test.ts +++ b/packages/agent-runtime/src/__tests__/prompts-schema-handling.test.ts @@ -233,7 +233,7 @@ describe('Schema handling error recovery', () => { }) describe('direct agent control envelope', () => { - test('exposes background and timeout controls for every direct agent tool', () => { + test('exposes the background control for every direct agent tool', () => { const schema = buildAgentToolInputSchema({ id: 'editor', displayName: 'Editor', @@ -253,7 +253,6 @@ describe('Schema handling error recovery', () => { schema.safeParse({ prompt: 'Implement it', background: true, - timeout_seconds: 120, }).success, ).toBe(true) }) @@ -442,12 +441,14 @@ describe('Schema handling error recovery', () => { }) }) - test('preserves background and timeout controls on direct agent tool calls', () => { + test('preserves the background control on direct agent tool calls', () => { const transformed = tryTransformAgentToolCall({ toolName: 'editor', input: { prompt: 'Implement the change', background: true, + // Stray deadline field from an older model habit: it is no longer part + // of the spawn entry contract, so it must not be forwarded. timeout_seconds: 90, }, spawnableAgents: ['openbuff/editor@1.0.0'], @@ -461,7 +462,6 @@ describe('Schema handling error recovery', () => { agent_type: 'openbuff/editor@1.0.0', prompt: 'Implement the change', background: true, - timeout_seconds: 90, }, ], }, diff --git a/packages/agent-runtime/src/__tests__/spawn-agent-inline-nesting.test.ts b/packages/agent-runtime/src/__tests__/spawn-agent-inline-nesting.test.ts index c0e2db8344..f0d1fdd0d0 100644 --- a/packages/agent-runtime/src/__tests__/spawn-agent-inline-nesting.test.ts +++ b/packages/agent-runtime/src/__tests__/spawn-agent-inline-nesting.test.ts @@ -1,6 +1,10 @@ import { TEST_USER_ID } from '@codebuff/common/old-constants' import { TEST_AGENT_RUNTIME_IMPL } from '@codebuff/common/testing/impl/agent-runtime' -import { getInitialSessionState } from '@codebuff/common/types/session-state' +import { agentReceiptSchema } from '@codebuff/common/types/agent-handoff' +import { + getInitialAgentState, + getInitialSessionState, +} from '@codebuff/common/types/session-state' import { afterEach, beforeEach, @@ -23,6 +27,7 @@ import type { AgentTemplate } from '@codebuff/common/types/agent-template' import type { CodebuffToolCall } from '@codebuff/common/tools/list' import type { ParamsExcluding } from '@codebuff/common/types/function-params' import type { PrintModeEvent } from '@codebuff/common/types/print-mode' +import type { AgentState } from '@codebuff/common/types/session-state' /** * Filters the writeToClient mock's captured calls, returning only the @@ -1255,3 +1260,119 @@ describe('spawn_agent_inline onResponseChunk parentAgentId nesting', () => { ).toBe(true) }) }) + +/** + * Finished child state carrying only the context telemetry these cases + * exercise. `contextTokenCount` is required on AgentState, so a child whose + * runtime never reported one is modelled by dropping the key through a Partial + * view rather than loosening the fixture's type. + */ +function childContextState(context: { + contextTokenCount?: number + contextWindowTokens?: number +}): AgentState { + const agentState: AgentState = { + ...getInitialAgentState(), + contextTokenCount: context.contextTokenCount ?? 0, + contextWindowTokens: context.contextWindowTokens, + } + if (context.contextTokenCount === undefined) { + delete (agentState as Partial).contextTokenCount + } + return agentState +} + +function receiptForChildContext( + agentId: string, + context: { contextTokenCount?: number; contextWindowTokens?: number }, +) { + return buildRuntimeAgentReceipt({ + agentType: 'file-picker', + agentId, + output: { type: 'structuredOutput', value: { status: 'completed' } }, + agentState: childContextState(context), + }) +} + +describe('buildRuntimeAgentReceipt context usage telemetry', () => { + it('reports the child tokens, window, and percent of window', () => { + const receipt = receiptForChildContext('file-picker-75', { + contextTokenCount: 150_000, + contextWindowTokens: 200_000, + }) + + expect(receipt.contextUsage).toEqual({ + tokens: 150_000, + windowTokens: 200_000, + percentOfWindow: 75, + }) + }) + + // A child may overrun its declared window, so the percent is clamped rather + // than emitted out of range, which would make the whole receipt unparseable. + it('clamps percentOfWindow to 100 for a child that overran its window', () => { + const receipt = receiptForChildContext('file-picker-overrun', { + contextTokenCount: 260_000, + contextWindowTokens: 200_000, + }) + + expect(receipt.contextUsage?.percentOfWindow).toBe(100) + expect(agentReceiptSchema.safeParse(receipt).success).toBe(true) + }) + + it('omits contextUsage when the child never reported a token count', () => { + const withoutTokens = receiptForChildContext('file-picker-no-tokens', { + contextWindowTokens: 200_000, + }) + const withTokens = receiptForChildContext('file-picker-with-tokens', { + contextTokenCount: 10_000, + contextWindowTokens: 200_000, + }) + + expect(withoutTokens.contextUsage).toBeUndefined() + expect(JSON.stringify(withoutTokens)).not.toContain('contextUsage') + // Presence is driven by the token count, not by the window alone. + expect(withTokens.contextUsage).toBeDefined() + }) + + it('keeps the token count and omits window fields without a usable window', () => { + expect( + receiptForChildContext('file-picker-no-window', { + contextTokenCount: 120_000, + }).contextUsage, + ).toEqual({ tokens: 120_000 }) + expect( + receiptForChildContext('file-picker-zero-window', { + contextTokenCount: 120_000, + contextWindowTokens: 0, + }).contextUsage, + ).toEqual({ tokens: 120_000 }) + }) + + // AgentState carries no compaction counter, so the receipt must never + // fabricate one for the parent's shard-sizing decision. + it('never invents a compactionCount', () => { + const contexts: Array<{ + contextTokenCount?: number + contextWindowTokens?: number + }> = [ + { contextTokenCount: 150_000, contextWindowTokens: 200_000 }, + { contextTokenCount: 260_000, contextWindowTokens: 200_000 }, + { contextTokenCount: 120_000 }, + { contextTokenCount: 120_000, contextWindowTokens: 0 }, + { contextWindowTokens: 200_000 }, + ] + + for (const [index, context] of contexts.entries()) { + const receipt = receiptForChildContext(`file-picker-${index}`, context) + + // Control: the telemetry itself is emitted whenever tokens are known, + // so the missing counter is an omission and not an absent feature. + expect(receipt.contextUsage !== undefined).toBe( + context.contextTokenCount !== undefined, + ) + expect(receipt.contextUsage?.compactionCount).toBeUndefined() + expect(JSON.stringify(receipt)).not.toContain('compactionCount') + } + }) +}) diff --git a/packages/agent-runtime/src/__tests__/subagent-timeout.test.ts b/packages/agent-runtime/src/__tests__/subagent-timeout.test.ts index ba5f8fe5bd..f83219ce59 100644 --- a/packages/agent-runtime/src/__tests__/subagent-timeout.test.ts +++ b/packages/agent-runtime/src/__tests__/subagent-timeout.test.ts @@ -3,7 +3,6 @@ import { spyOn, beforeEach, afterEach, describe, expect, it } from 'bun:test' import * as spawnAgentUtils from '../tools/handlers/tool/spawn-agent-utils' import { - resolveSubagentTimeoutMs, executeSubagent, createCombinedAbortSignal, } from '../tools/handlers/tool/spawn-agent-utils' @@ -35,38 +34,6 @@ function makeTemplate(overrides: Partial = {}): AgentTemplate { } as AgentTemplate } -describe('resolveSubagentTimeoutMs', () => { - it('uses the explicit per-spawn override when provided', () => { - const tpl = makeTemplate({ defaultTimeoutMs: 5 * 60 * 1000 }) - expect(resolveSubagentTimeoutMs(tpl, 30 * 1000)).toBe(30 * 1000) - }) - - it('falls back to the agent template defaultTimeoutMs when no override', () => { - const tpl = makeTemplate({ defaultTimeoutMs: 5 * 60 * 1000 }) - expect(resolveSubagentTimeoutMs(tpl, undefined)).toBe(5 * 60 * 1000) - }) - - it('has no wall-clock timeout when neither override nor template default is set', () => { - const tpl = makeTemplate() - expect(resolveSubagentTimeoutMs(tpl, undefined)).toBe(-1) - }) - - it('override takes precedence over template default even when default is -1', () => { - const tpl = makeTemplate({ defaultTimeoutMs: -1 }) - expect(resolveSubagentTimeoutMs(tpl, 1000)).toBe(1000) - }) - - it('template default of -1 (no timeout) is respected when no override', () => { - const tpl = makeTemplate({ defaultTimeoutMs: -1 }) - expect(resolveSubagentTimeoutMs(tpl, undefined)).toBe(-1) - }) - - it('explicit override of -1 (no timeout) is respected', () => { - const tpl = makeTemplate({ defaultTimeoutMs: 5 * 60 * 1000 }) - expect(resolveSubagentTimeoutMs(tpl, -1)).toBe(-1) - }) -}) - describe('withTimeout abort support', () => { it('aborts the controller on deadline before rejecting', async () => { const controller = new AbortController() @@ -113,7 +80,7 @@ describe('withTimeout abort support', () => { }) }) -describe('spawn_agents timeout_seconds override wiring', () => { +describe('spawn_agents no longer applies a wall-clock deadline', () => { let mockAgentTemplate: AgentTemplate let baseParams: any @@ -150,7 +117,7 @@ describe('spawn_agents timeout_seconds override wiring', () => { spyOn(spawnAgentUtils, 'executeSubagent').mockRestore?.() }) - it('passes timeout_seconds (seconds → ms) as subagentTimeoutMs to executeSubagent', async () => { + it('never forwards a subagentTimeoutMs option, even when an entry still carries timeout_seconds', async () => { const spy = spyOn(spawnAgentUtils, 'executeSubagent').mockResolvedValue({ agentState: { ...getInitialAgentState(), agentId: 'sub' }, output: { type: 'lastMessage' as const, value: [] }, @@ -174,86 +141,7 @@ describe('spawn_agents timeout_seconds override wiring', () => { } as any) expect(spy).toHaveBeenCalledTimes(1) - const passedTimeout = spy.mock.calls[0][0].subagentTimeoutMs - expect(passedTimeout).toBe(120 * 1000) - }) - - it('passes undefined when timeout_seconds is omitted (template default applies)', async () => { - const spy = spyOn(spawnAgentUtils, 'executeSubagent').mockResolvedValue({ - agentState: { ...getInitialAgentState(), agentId: 'sub' }, - output: { type: 'lastMessage' as const, value: [] }, - } as any) - - await handleSpawnAgents({ - ...baseParams, - toolCall: { - toolName: 'spawn_agents', - toolCallId: 'c1', - input: { - agents: [{ agent_type: 'test-agent', prompt: 'do thing' }], - }, - } as any, - } as any) - - expect(spy).toHaveBeenCalledTimes(1) - expect(spy.mock.calls[0][0].subagentTimeoutMs).toBeUndefined() - }) - - it('passes -1 → -1000 ms when timeout_seconds is -1 (no timeout)', async () => { - const spy = spyOn(spawnAgentUtils, 'executeSubagent').mockResolvedValue({ - agentState: { ...getInitialAgentState(), agentId: 'sub' }, - output: { type: 'lastMessage' as const, value: [] }, - } as any) - - await handleSpawnAgents({ - ...baseParams, - toolCall: { - toolName: 'spawn_agents', - toolCallId: 'c1', - input: { - agents: [ - { - agent_type: 'test-agent', - prompt: 'long thing', - timeout_seconds: -1, - }, - ], - }, - } as any, - } as any) - - expect(spy).toHaveBeenCalledTimes(1) - // -1 seconds → -1000 ms, which resolveSubagentTimeoutMs treats as - // no-timeout (non-positive). - expect(spy.mock.calls[0][0].subagentTimeoutMs).toBe(-1000) - }) - - it('applies the same override to background (detached) spawns', async () => { - const spy = spyOn(spawnAgentUtils, 'executeSubagent').mockResolvedValue({ - agentState: { ...getInitialAgentState(), agentId: 'sub' }, - output: { type: 'lastMessage' as const, value: [] }, - } as any) - - await handleSpawnAgents({ - ...baseParams, - toolCall: { - toolName: 'spawn_agents', - toolCallId: 'c1', - input: { - agents: [ - { - agent_type: 'test-agent', - prompt: 'bg thing', - background: true, - timeout_seconds: 300, - }, - ], - }, - } as any, - } as any) - - expect(spy).toHaveBeenCalledTimes(1) - expect(spy.mock.calls[0][0].subagentTimeoutMs).toBe(300 * 1000) + expect('subagentTimeoutMs' in spy.mock.calls[0][0]).toBe(false) }) }) diff --git a/packages/agent-runtime/src/run-agent-step.ts b/packages/agent-runtime/src/run-agent-step.ts index 405a987e70..1b9e625231 100644 --- a/packages/agent-runtime/src/run-agent-step.ts +++ b/packages/agent-runtime/src/run-agent-step.ts @@ -1412,9 +1412,14 @@ export async function loopAgentSteps( // its session persistence would replay forever. The settle carries this run's // correlation, so it can only ever settle the pending state this run started. let unsettledCompactionStart = false + // Highest progress percent already reported for the CURRENT announced pass. + // Declared alongside the pending flag because the two are settled together. + let lastCompactionProgressPercent = 0 const settleCompactionStatus = () => { if (!unsettledCompactionStart) return unsettledCompactionStart = false + // A later pass in this same run announces its own progress from 0 again. + lastCompactionProgressPercent = 0 onResponseChunk({ type: 'context_compaction_status', state: 'settled', @@ -1422,6 +1427,36 @@ export async function loopAgentSteps( }) } + // Best-effort progress inside an announced pass, carrying this run's own + // correlation so it can only ever advance the card this run opened. The + // pruner reports no total, so the percent is a deterministic milestone + // estimate rather than a measurement: it is clamped to 0..100 and forced + // monotonic here, because the inline spawn path emits activity ticks for the + // same pass and an 'applying' milestone can otherwise be undercut by a late + // tick. Gated on `unsettledCompactionStart` so nothing is emitted outside an + // announced pass (a suppressed or below-trigger iteration stays silent, just + // as it emits neither half of the status pair). Purely telemetry: it never + // gates or aborts a pass. + const emitCompactionProgress = ( + phase: 'analyzing' | 'summarizing' | 'applying', + percent: number, + extra?: { contextTokens?: number; targetBudgetTokens?: number }, + ) => { + if (!unsettledCompactionStart) return + const next = Math.max( + lastCompactionProgressPercent, + Math.min(100, Math.max(0, Math.round(percent))), + ) + lastCompactionProgressPercent = next + onResponseChunk({ + type: 'context_compaction_progress', + ...compactionCorrelation, + phase, + percent: next, + ...extra, + }) + } + // Outer try/finally guarantees this run's in-memory programmatic-step state // is torn down on EVERY exit path after runId is assigned — including if the // prompt/tool setup below throws before the main loop's own try/catch is @@ -1975,6 +2010,12 @@ export async function loopAgentSteps( triggerBudgetTokens: semanticBudget.triggerBudgetTokens, targetBudgetTokens: semanticBudget.targetBudgetTokens, }) + // First milestone of the announced pass, so the UI shows real movement + // instead of an idle bar while the inline pruner starts up. + emitCompactionProgress('analyzing', 20, { + contextTokens: contextTokensBeforeProgrammatic, + targetBudgetTokens: semanticBudget.targetBudgetTokens, + }) } // 1. Run programmatic step first if it exists @@ -2100,9 +2141,16 @@ export async function loopAgentSteps( system, tools, userInputId, + onCompactionProgress: emitCompactionProgress, }) } + // The pruner has returned (or the generator step that owned it has), so + // the announced pass is down to applying whatever it produced. Emitted + // for a suppressed/unannounced iteration too, where the gate inside the + // emitter makes it a no-op. + emitCompactionProgress('applying', 90) + // Capture the request goal once per step for the root agent. The // compaction branch below scrapes only when a // session actually compacts, so a session that never compacts used to diff --git a/packages/agent-runtime/src/templates/__tests__/strings.test.ts b/packages/agent-runtime/src/templates/__tests__/strings.test.ts index 35c7d00b15..02461e7252 100644 --- a/packages/agent-runtime/src/templates/__tests__/strings.test.ts +++ b/packages/agent-runtime/src/templates/__tests__/strings.test.ts @@ -33,6 +33,7 @@ import compatibilityReviewer from '../../../../../agents/specialists/compatibili import { createBase2, GUIDE_POINTERS } from '../../../../../agents/base2/base2' import type { AgentTemplate } from '../types' +import type { ResolveModelContextWindow } from '../prompts' import type { ContextBudgetLedger } from '../../util/context-budget' import type { AgentState } from '@codebuff/common/types/session-state' import type { ProjectFileContext } from '@codebuff/common/util/file' @@ -887,6 +888,227 @@ describe('getAgentPrompt', () => { } } }) + + test('threads a resolved context window into the spawnable agent catalog', async () => { + const filePickerTemplate = createMockAgentTemplate({ + id: 'file-picker', + displayName: 'File Picker', + spawnerPrompt: 'Spawn to find relevant files in a codebase', + }) + + const mainAgentTemplate = createMockAgentTemplate({ + id: 'main-agent', + displayName: 'Main Agent', + spawnableAgents: ['file-picker'], + instructionsPrompt: 'Main agent instructions.', + }) + + const agentTemplates: Record = { + 'main-agent': mainAgentTemplate, + 'file-picker': filePickerTemplate, + } + + const result = await getAgentPrompt({ + agentTemplate: mainAgentTemplate, + promptType: { type: 'instructionsPrompt' }, + fileContext: createMockFileContext(), + agentState: createMockAgentState('main-agent'), + agentTemplates, + additionalToolDefinitions: async () => ({}), + logger: createMockLogger(), + apiKey: TEST_AGENT_RUNTIME_IMPL.apiKey, + databaseAgentCache: TEST_AGENT_RUNTIME_IMPL.databaseAgentCache, + fetchAgentFromDatabase: TEST_AGENT_RUNTIME_IMPL.fetchAgentFromDatabase, + resolveModelContextWindow: () => 200_000, + }) + + expect(result).toContain('You can spawn the following agents:') + expect(result).toContain( + '- file-picker: Spawn to find relevant files in a codebase [context ~200k]', + ) + }) + + test('keeps the catalog byte-identical when no window resolver is injected', async () => { + const filePickerTemplate = createMockAgentTemplate({ + id: 'file-picker', + displayName: 'File Picker', + spawnerPrompt: 'Spawn to find relevant files in a codebase', + }) + + const mainAgentTemplate = createMockAgentTemplate({ + id: 'main-agent', + displayName: 'Main Agent', + spawnableAgents: ['file-picker'], + instructionsPrompt: 'Main agent instructions.', + }) + + const agentTemplates: Record = { + 'main-agent': mainAgentTemplate, + 'file-picker': filePickerTemplate, + } + + const withoutResolver = await getAgentPrompt({ + agentTemplate: mainAgentTemplate, + promptType: { type: 'instructionsPrompt' }, + fileContext: createMockFileContext(), + agentState: createMockAgentState('main-agent'), + agentTemplates, + additionalToolDefinitions: async () => ({}), + logger: createMockLogger(), + apiKey: TEST_AGENT_RUNTIME_IMPL.apiKey, + databaseAgentCache: TEST_AGENT_RUNTIME_IMPL.databaseAgentCache, + fetchAgentFromDatabase: TEST_AGENT_RUNTIME_IMPL.fetchAgentFromDatabase, + }) + + const withResolver = await getAgentPrompt({ + agentTemplate: mainAgentTemplate, + promptType: { type: 'instructionsPrompt' }, + fileContext: createMockFileContext(), + agentState: createMockAgentState('main-agent'), + agentTemplates, + additionalToolDefinitions: async () => ({}), + logger: createMockLogger(), + apiKey: TEST_AGENT_RUNTIME_IMPL.apiKey, + databaseAgentCache: TEST_AGENT_RUNTIME_IMPL.databaseAgentCache, + fetchAgentFromDatabase: TEST_AGENT_RUNTIME_IMPL.fetchAgentFromDatabase, + resolveModelContextWindow: () => 200_000, + }) + + const expectedCatalog = [ + 'Main agent instructions.', + '', + 'You can spawn the following agents:', + '', + '- file-picker: Spawn to find relevant files in a codebase', + ].join('\n') + + expect(withoutResolver).toBe(expectedCatalog) + expect(withoutResolver).not.toContain('[context ~') + // Control: the only delta an injected resolver adds is the suffix. + expect(withResolver).toBe(`${expectedCatalog} [context ~200k]`) + }) + + test('suffixes only the children whose context window resolves', async () => { + const filePickerTemplate = createMockAgentTemplate({ + id: 'file-picker', + displayName: 'File Picker', + spawnerPrompt: 'Spawn to find relevant files in a codebase', + }) + + const globMatcherTemplate = createMockAgentTemplate({ + id: 'glob-matcher', + displayName: 'Glob Matcher', + spawnerPrompt: 'Mechanically runs multiple glob pattern matches', + }) + + const mainAgentTemplate = createMockAgentTemplate({ + id: 'main-agent', + displayName: 'Main Agent', + spawnableAgents: ['file-picker', 'glob-matcher'], + instructionsPrompt: 'Main agent instructions.', + }) + + const agentTemplates: Record = { + 'main-agent': mainAgentTemplate, + 'file-picker': filePickerTemplate, + 'glob-matcher': globMatcherTemplate, + } + + // Only one route has a declared window; the other resolves to undefined. + const resolveModelContextWindow: ResolveModelContextWindow = ({ + agentId, + }) => (agentId === 'file-picker' ? 200_000 : undefined) + + const result = await getAgentPrompt({ + agentTemplate: mainAgentTemplate, + promptType: { type: 'instructionsPrompt' }, + fileContext: createMockFileContext(), + agentState: createMockAgentState('main-agent'), + agentTemplates, + additionalToolDefinitions: async () => ({}), + logger: createMockLogger(), + apiKey: TEST_AGENT_RUNTIME_IMPL.apiKey, + databaseAgentCache: TEST_AGENT_RUNTIME_IMPL.databaseAgentCache, + fetchAgentFromDatabase: TEST_AGENT_RUNTIME_IMPL.fetchAgentFromDatabase, + resolveModelContextWindow, + }) + + expect(result).toContain( + '- file-picker: Spawn to find relevant files in a codebase [context ~200k]', + ) + expect(result).toContain( + '- glob-matcher: Mechanically runs multiple glob pattern matches', + ) + expect((result ?? '').split('[context ~').length - 1).toBe(1) + }) + + test('compact catalog appends the window after the required-params hint', () => { + const gitCommitterTemplate = createMockAgentTemplate({ + id: 'git-committer', + displayName: 'Git Committer', + spawnerPrompt: 'Safely delivers task-owned changes through git', + inputSchema: { + params: z.object({ + owned_paths: z.array(z.string()), + }), + }, + }) + + const withoutWindow = formatCompactAgentCatalogLine( + 'git-committer', + gitCommitterTemplate, + ) + const withWindow = formatCompactAgentCatalogLine( + 'git-committer', + gitCommitterTemplate, + 200_000, + ) + + // The two-argument call form is unchanged, and the window is appended + // after the hint rather than spliced into it. + expect(withoutWindow).toBe( + '- git-committer: Safely delivers task-owned changes through git Required params: `owned_paths`.', + ) + expect(withWindow).toBe(`${withoutWindow} [context ~200k]`) + }) + + test('compact catalog appends nothing for an unknown, zero, or non-finite window', () => { + const filePickerTemplate = createMockAgentTemplate({ + id: 'file-picker', + displayName: 'File Picker', + spawnerPrompt: 'Spawn to find relevant files in a codebase', + }) + const baseline = formatCompactAgentCatalogLine( + 'file-picker', + filePickerTemplate, + ) + + expect(baseline).toBe( + '- file-picker: Spawn to find relevant files in a codebase', + ) + for (const windowTokens of [ + undefined, + 0, + Number.POSITIVE_INFINITY, + Number.NaN, + ]) { + expect( + formatCompactAgentCatalogLine( + 'file-picker', + filePickerTemplate, + windowTokens, + ), + ).toBe(baseline) + } + // Control: a finite positive window is still honored for this fixture. + expect( + formatCompactAgentCatalogLine( + 'file-picker', + filePickerTemplate, + 200_000, + ), + ).toBe(`${baseline} [context ~200k]`) + }) }) test('uses harvested-text addendum when set_output is programmatic-only', async () => { diff --git a/packages/agent-runtime/src/templates/prompts.ts b/packages/agent-runtime/src/templates/prompts.ts index 692a76a864..4d6750cf04 100644 --- a/packages/agent-runtime/src/templates/prompts.ts +++ b/packages/agent-runtime/src/templates/prompts.ts @@ -7,11 +7,21 @@ import { z } from 'zod/v4' import { getAgentTemplate } from './agent-registry' import type { AgentTemplate } from '@codebuff/common/types/agent-template' +import type { AgentRuntimeDeps } from '@codebuff/common/types/contracts/agent-runtime' import type { Logger } from '@codebuff/common/types/contracts/logger' import type { ParamsExcluding } from '@codebuff/common/types/function-params' import type { AgentTemplateType } from '@codebuff/common/types/session-state' import type { ToolSet } from 'ai' +/** + * Injected resolver for a child's declared context window, threaded through + * `AgentRuntimeDeps`. Synchronous and may return `undefined`, so callers must + * never await it and must stay byte-identical when it is absent. + */ +export type ResolveModelContextWindow = NonNullable< + AgentRuntimeDeps['resolveModelContextWindow'] +> + function ensureJsonSchemaCompatible(schema: z.ZodType): z.ZodType { try { z.toJSONSchema(schema, { io: 'input' }) @@ -92,26 +102,43 @@ function formatRequiredSpawnContractHint( return `Required params: ${missing.map((key) => '`' + key + '`').join(', ')}.` } +/** Compact token rendering for catalog lines: `200_000` -> `200k`. */ +function formatContextWindowTokens(tokens: number): string { + return tokens >= 1_000 + ? `${Math.round(tokens / 1000)}k` + : `${Math.round(tokens)}` +} + /** * Compact catalog line for the "You can spawn the following agents" addendum. * Appends a one-line required-params/handoff hint when the child's contract - * is not already named in `spawnerPrompt`. + * is not already named in `spawnerPrompt`, then the child's context window as + * ` [context ~200k]` when the caller resolved one, so the parent can size + * delegated work. An unknown window appends nothing, keeping the catalog + * byte-identical for callers that inject no resolver. */ export function formatCompactAgentCatalogLine( agentType: AgentTemplateType, agentTemplate: AgentTemplate | null | undefined, + contextWindowTokens?: number, ): string { - if (!agentTemplate) return `- ${agentType}` + const windowSuffix = + typeof contextWindowTokens === 'number' && + Number.isFinite(contextWindowTokens) && + contextWindowTokens > 0 + ? ` [context ~${formatContextWindowTokens(contextWindowTokens)}]` + : '' + if (!agentTemplate) return `- ${agentType}${windowSuffix}` const prompt = agentTemplate.spawnerPrompt const hint = formatRequiredSpawnContractHint(agentType, agentTemplate) if (prompt) { return hint - ? `- ${agentType}: ${prompt} ${hint}` - : `- ${agentType}: ${prompt}` + ? `- ${agentType}: ${prompt} ${hint}${windowSuffix}` + : `- ${agentType}: ${prompt}${windowSuffix}` } - if (hint) return `- ${agentType}: ${hint}` - return `- ${agentType}` + if (hint) return `- ${agentType}: ${hint}${windowSuffix}` + return `- ${agentType}${windowSuffix}` } /** @@ -158,12 +185,6 @@ export function buildAgentToolInputSchema( .boolean() .optional() .describe('Launch the agent as a background job when true.') - schemaFields.timeout_seconds = z - .number() - .optional() - .describe( - 'Optional per-spawn wall-clock timeout in seconds; -1 disables it.', - ) return z .object(schemaFields) @@ -235,6 +256,7 @@ export async function buildAgentToolSet( function buildSingleAgentDescription( agentType: AgentTemplateType, agentTemplate: AgentTemplate | null, + contextWindowTokens?: number, ): string { if (!agentTemplate) { // Fallback for unknown agents @@ -252,7 +274,11 @@ params: None` : ['prompt: None', 'params: None'].join('\n') return buildArray( - formatCompactAgentCatalogLine(agentType, agentTemplate), + formatCompactAgentCatalogLine( + agentType, + agentTemplate, + contextWindowTokens, + ), agentTemplate.includeMessageHistory && 'This agent can see the current message history.', agentTemplate.inheritParentSystemPrompt && @@ -270,6 +296,8 @@ export async function buildFullSpawnableAgentsSpec( spawnableAgents: AgentTemplateType[] agentTemplates: Record logger: Logger + /** Optional; when absent the spec is byte-identical to the pre-window output. */ + resolveModelContextWindow?: ResolveModelContextWindow } & ParamsExcluding< typeof getAgentTemplate, 'agentId' | 'localAgentTemplates' @@ -285,20 +313,29 @@ export async function buildFullSpawnableAgentsSpec( const subAgentTypesAndTemplates = await Promise.all( spawnableAgents.map(async (agentType) => { + const agentTemplate = await getAgentTemplate({ + ...params, + agentId: agentType, + localAgentTemplates: agentTemplates, + }) return [ agentType, - await getAgentTemplate({ - ...params, - agentId: agentType, - localAgentTemplates: agentTemplates, + agentTemplate, + params.resolveModelContextWindow?.({ + agentId: agentTemplate?.id ?? agentType, + model: agentTemplate?.model, }), ] as const }), ) const agentsDescription = subAgentTypesAndTemplates - .map(([agentType, agentTemplate]) => - buildSingleAgentDescription(agentType, agentTemplate), + .map(([agentType, agentTemplate, contextWindowTokens]) => + buildSingleAgentDescription( + agentType, + agentTemplate, + contextWindowTokens, + ), ) .filter(Boolean) .join('\n\n') diff --git a/packages/agent-runtime/src/templates/strings.ts b/packages/agent-runtime/src/templates/strings.ts index 5c7059b747..f1e1d26701 100644 --- a/packages/agent-runtime/src/templates/strings.ts +++ b/packages/agent-runtime/src/templates/strings.ts @@ -39,6 +39,7 @@ import { applyMeasure } from '../util/context-budget' import { parseUserMessage } from '../util/messages' import type { AgentTemplate, PlaceholderValue } from './types' +import type { ResolveModelContextWindow } from './prompts' import type { Logger } from '@codebuff/common/types/contracts/logger' import type { ContextBudgetLedger } from '../util/context-budget' import type { ParamsExcluding } from '@codebuff/common/types/function-params' @@ -338,6 +339,9 @@ export async function getAgentPrompt( additionalToolDefinitions: () => Promise logger: Logger useParentTools?: boolean + /** Injected by the runtime (BYOK routing). Absent for tests and non-BYOK + * callers, in which case the catalog output stays byte-identical. */ + resolveModelContextWindow?: ResolveModelContextWindow } & ParamsExcluding< typeof formatPrompt, 'prompt' | 'tools' | 'spawnableAgents' @@ -406,7 +410,14 @@ export async function getAgentPrompt( agentId: agentType, localAgentTemplates: agentTemplates, }) - return formatCompactAgentCatalogLine(agentType, template) + return formatCompactAgentCatalogLine( + agentType, + template, + params.resolveModelContextWindow?.({ + agentId: template?.id ?? agentType, + model: template?.model, + }), + ) }), ) addendum += `\n\nYou can spawn the following agents:\n\n${agentDescriptions.join('\n')}` diff --git a/packages/agent-runtime/src/tools/handlers/tool/spawn-agent-inline.ts b/packages/agent-runtime/src/tools/handlers/tool/spawn-agent-inline.ts index 9a469cc565..b4ebf54674 100644 --- a/packages/agent-runtime/src/tools/handlers/tool/spawn-agent-inline.ts +++ b/packages/agent-runtime/src/tools/handlers/tool/spawn-agent-inline.ts @@ -261,6 +261,11 @@ export const handleSpawnAgentInline = (async ( // Extract common context params to avoid bugs from spreading all params const contextParams = extractSubagentContextParams(params) + // Observed pruner chunks, counted only so the parent's announced compaction + // pass can report live movement. The chunks themselves are still dropped: see + // the `else` branch of `onResponseChunk` below. + let prunerChunks = 0 + let result: Awaited> try { result = await executeSubagent({ @@ -334,6 +339,25 @@ export const handleSpawnAgentInline = (async ( } writeToClient(chunk) + } else { + // pruner output stays invisible; only a bounded activity tick is + // surfaced so the UI can show real progress instead of a silent stall. + // Stamped with the PARENT run's correlation, because the announced + // compaction pass this progress belongs to is the parent's, not the + // pruner child's. Nothing is emitted without a parent runId: there + // would be no announced pass to correlate the tick to. + prunerChunks += 1 + const parentRunId = parentAgentState.runId + if (parentRunId) { + writeToClient({ + type: 'context_compaction_progress', + runId: parentRunId, + ancestorRunIds: [...parentAgentState.ancestorRunIds], + agentId: parentAgentState.agentId, + phase: 'summarizing', + percent: Math.min(85, 45 + prunerChunks * 5), + }) + } } }, clearUserPromptMessagesAfterResponse: false, diff --git a/packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts b/packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts index ed940e76f3..9868bc8e5b 100644 --- a/packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts +++ b/packages/agent-runtime/src/tools/handlers/tool/spawn-agent-utils.ts @@ -7,7 +7,6 @@ import { normalizeAgentIdForLookup, parseAgentId, } from '@codebuff/common/util/agent-id-parsing' -import { withTimeout } from '@codebuff/common/util/promise' import { generateCompactId } from '@codebuff/common/util/string' import { containsStructuralAuditReceipt } from '@codebuff/common/util/audit-receipt' import { @@ -1295,6 +1294,46 @@ function extractReceiptEvidence(params: { return evidence.slice(-128) } +/** + * Observational context telemetry for the parent: how much of its window the + * finished child actually used. Built only from fields the child state already + * carries (`contextTokenCount` / `contextWindowTokens`) — there is no + * compaction counter on AgentState, so `compactionCount` is never invented. + * Omitted entirely when the token count is not a finite non-negative number, + * and `percentOfWindow` is clamped to 100 so a child that overran its declared + * window still yields a receipt that validates. + */ +function buildReceiptContextUsage( + agentState?: AgentState, +): AgentReceipt['contextUsage'] { + const rawTokens = agentState?.contextTokenCount + if ( + typeof rawTokens !== 'number' || + !Number.isFinite(rawTokens) || + rawTokens < 0 + ) { + return undefined + } + const tokens = Math.round(rawTokens) + const rawWindow = agentState?.contextWindowTokens + const windowTokens = + typeof rawWindow === 'number' && Number.isFinite(rawWindow) && rawWindow > 0 + ? Math.round(rawWindow) + : undefined + return { + tokens, + ...(windowTokens === undefined + ? {} + : { + windowTokens, + percentOfWindow: Math.min( + 100, + Math.round((tokens / windowTokens) * 100), + ), + }), + } +} + export function buildRuntimeAgentReceipt(params: { agentType: string agentId: string @@ -1480,6 +1519,7 @@ export function buildRuntimeAgentReceipt(params: { }, } : normalizedOutput + const contextUsage = buildReceiptContextUsage(params.agentState) const receipt = agentReceiptSchema.parse({ schemaVersion: 1, receiptId: generateCompactId(), @@ -1521,6 +1561,7 @@ export function buildRuntimeAgentReceipt(params: { artifacts: extractReceiptStringArray(receiptSources, 'artifacts'), errors, output: reconciledOutput as any, + ...(contextUsage ? { contextUsage } : {}), }) return receipt } @@ -1962,33 +2003,6 @@ export function logAgentSpawn(params: { ) } -/** - * Shared wall-clock default for a single subagent execution. - * - * Productive subagents are unlimited by default. Callers can opt into a - * deadline with timeout_seconds or a positive template-specific timeout. - */ -const DEFAULT_SUBAGENT_TIMEOUT_MS = -1 - -/** - * Resolves the wall-clock timeout (ms) for a subagent execution, in precedence - * order: explicit per-spawn override > agent template default > shared - * DEFAULT_SUBAGENT_TIMEOUT_MS. A non-positive explicit or template value - * (-1, 0) disables the timeout entirely. - */ -export function resolveSubagentTimeoutMs( - agentTemplate: AgentTemplate, - subagentTimeoutMs?: number, -): number { - if (subagentTimeoutMs !== undefined) { - return subagentTimeoutMs - } - if (agentTemplate.defaultTimeoutMs !== undefined) { - return agentTemplate.defaultTimeoutMs - } - return DEFAULT_SUBAGENT_TIMEOUT_MS -} - /** * Executes a subagent using loopAgentSteps */ @@ -2001,7 +2015,6 @@ export async function executeSubagent( onResponseChunk: (chunk: string | PrintModeEvent) => void isOnlyChild?: boolean ancestorRunIds: string[] - subagentTimeoutMs?: number spawnToolCallId?: string spawnIndex?: number } & ParamsExcluding, @@ -2021,7 +2034,6 @@ export async function executeSubagent( ancestorRunIds, prompt, spawnParams, - subagentTimeoutMs, spawnToolCallId, spawnIndex, } = withDefaults @@ -2055,59 +2067,29 @@ export async function executeSubagent( } onResponseChunk(startEvent) - // Thread an AbortController through withTimeout so the deadline actually - // cancels the underlying loopAgentSteps stream. The subagent's signal is the - // combination of the parent's signal (so a user/parent-level abort still - // propagates) and the timeout controller (so the deadline cancels this - // subagent without affecting its siblings). AbortSignal.any is available in - // Node 20+ and Bun. If unavailable at runtime, fall back to a manual - // EventTarget bridge so this stays safe on older runtimes. - const resolvedTimeoutMs = resolveSubagentTimeoutMs( - agentTemplate, - subagentTimeoutMs, - ) - const timeoutController = - resolvedTimeoutMs > 0 ? new AbortController() : undefined - const parentSignal = withDefaults.signal - let combinedSignal: (AbortSignal & { cleanup?: () => void }) | undefined - const subagentSignal = - timeoutController && parentSignal - ? (AbortSignal as any).any - ? (AbortSignal as any).any([parentSignal, timeoutController.signal]) - : (combinedSignal = createCombinedAbortSignal( - parentSignal, - timeoutController.signal, - )) - : timeoutController - ? timeoutController.signal - : parentSignal - + // The subagent runs on the parent's signal (carried by ...withDefaults), so + // user/parent cancellation still propagates. There is no wall-clock deadline: + // productive subagents are bounded only by cancellation, the repeated-step + // watchdog, spawn depth, and cost/token budgets. let result - let timedOut = false + let failed = false try { - result = await withTimeout( - loopAgentSteps({ - ...withDefaults, - signal: subagentSignal as AbortSignal, - onResponseChunk, - // Don't propagate parent's image content to subagents. - // If subagents need to see images, they get them through includeMessageHistory, - // not by creating new image-containing messages for their prompts. - content: undefined, - ancestorRunIds: [...ancestorRunIds, parentAgentState.runId ?? ''], - agentType: agentTemplate.id, - }), - resolvedTimeoutMs, - `Subagent ${agentTemplate.id} exceeded wall-clock timeout of ${resolvedTimeoutMs}ms`, - timeoutController ? { controller: timeoutController } : {}, - ) + result = await loopAgentSteps({ + ...withDefaults, + onResponseChunk, + // Don't propagate parent's image content to subagents. + // If subagents need to see images, they get them through includeMessageHistory, + // not by creating new image-containing messages for their prompts. + content: undefined, + ancestorRunIds: [...ancestorRunIds, parentAgentState.runId ?? ''], + agentType: agentTemplate.id, + }) } catch (error) { - // withTimeout rejects on deadline and has already aborted timeoutController, - // which cancels loopAgentSteps via the combined signal (loopAgentSteps checks - // signal.aborted at lines 889/1104/1369). Emit a finish event so the UI - // doesn't show a subagent that started but never finished, then re-throw so - // the parent sees the error via Promise.allSettled. - timedOut = true + // Any subagent failure (cancellation, budget exhaustion, thrown error) must + // still emit a finish event so the UI never shows a subagent that started + // but never finished. Re-throw so the parent sees the error via + // Promise.allSettled. + failed = true onResponseChunk({ type: 'subagent_finish', agentId: withDefaults.agentState.agentId, @@ -2122,13 +2104,9 @@ export async function executeSubagent( error: error instanceof Error ? error.message : String(error), }) throw error - } finally { - // RF-2: remove fallback listeners when neither signal fires to avoid - // retention on long-lived parentSignal. - combinedSignal?.cleanup?.() } - if (!timedOut) { + if (!failed) { onResponseChunk({ type: 'subagent_finish', agentId: result.agentState.agentId, diff --git a/packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts b/packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts index 35624958a5..8fd2573e75 100644 --- a/packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts +++ b/packages/agent-runtime/src/tools/handlers/tool/spawn-agents.ts @@ -387,7 +387,7 @@ export const handleSpawnAgents = (async ( subAgentState, spawnIndex, } = validated - const { prompt, timeout_seconds } = validated.input + const { prompt } = validated.input const contextParams = extractSubagentContextParams(params) @@ -397,8 +397,7 @@ export const handleSpawnAgents = (async ( // listener on BOTH inputs, and neither fires when the job settles // normally, so its cleanup() must run on settle or that listener (plus a // closure over this job's AbortController) stays attached to the - // long-lived parent signal for the rest of the run. Mirrors - // executeSubagent's `finally { combinedSignal?.cleanup?.() }`. + // long-lived parent signal for the rest of the run. const combinedSignal = contextParams.signal ? createCombinedAbortSignal( contextParams.signal, @@ -431,9 +430,6 @@ export const handleSpawnAgents = (async ( fingerprintId, spawnToolCallId: toolCall.toolCallId, spawnIndex, - // Per-spawn wall-clock override (seconds → ms; -1 → no timeout). - subagentTimeoutMs: - timeout_seconds === undefined ? undefined : timeout_seconds * 1000, // Background agents are detached; the parent never waits for them, so // the "only child" step-count semantics (tuned for blocking spawns the // parent blocks on) never apply. Force false regardless of how many @@ -620,7 +616,7 @@ export const handleSpawnAgents = (async ( subAgentState, spawnIndex, }) => { - const { prompt, timeout_seconds } = input + const { prompt } = input // Extract common context params to avoid bugs from spreading all params const contextParams = extractSubagentContextParams(params) @@ -640,9 +636,6 @@ export const handleSpawnAgents = (async ( spawnToolCallId: toolCall.toolCallId, spawnIndex, isOnlyChild: foregroundAgents.length === 1, - // Per-spawn wall-clock override (seconds → ms; -1 → no timeout). - subagentTimeoutMs: - timeout_seconds === undefined ? undefined : timeout_seconds * 1000, excludeToolFromMessageHistory: false, fromHandleSteps: false, parentSystemPrompt, diff --git a/packages/agent-runtime/src/tools/tool-executor.ts b/packages/agent-runtime/src/tools/tool-executor.ts index 05bceb97eb..06347f836d 100644 --- a/packages/agent-runtime/src/tools/tool-executor.ts +++ b/packages/agent-runtime/src/tools/tool-executor.ts @@ -3412,9 +3412,6 @@ export function tryTransformAgentToolCall(params: { if (Object.hasOwn(input, 'background')) { agentEntry.background = input.background } - if (Object.hasOwn(input, 'timeout_seconds')) { - agentEntry.timeout_seconds = input.timeout_seconds - } const spawnAgentsInput = { agents: [agentEntry], } diff --git a/packages/agent-runtime/src/util/runtime-semantic-compaction.ts b/packages/agent-runtime/src/util/runtime-semantic-compaction.ts index fdb6b87c5e..41b6974703 100644 --- a/packages/agent-runtime/src/util/runtime-semantic-compaction.ts +++ b/packages/agent-runtime/src/util/runtime-semantic-compaction.ts @@ -44,6 +44,16 @@ export async function runRuntimeSemanticCompaction( /** Parent tool surface, inherited by the pruner child. */ tools: ToolSet userInputId: string + /** + * Best-effort progress reporter owned by the caller's announced pass. Called + * once per observed pruner chunk; the caller clamps and enforces + * monotonicity. Optional, and never load-bearing: the pass runs identically + * without it. + */ + onCompactionProgress?: ( + phase: 'analyzing' | 'summarizing' | 'applying', + percent: number, + ) => void } & ParamsExcluding< typeof executeSubagent, | 'agentState' @@ -69,6 +79,10 @@ export async function runRuntimeSemanticCompaction( userInputId, } = params const runId = parentAgentState.runId ?? parentAgentState.agentId + // Counts observed pruner chunks so the announced pass can report live + // movement. The chunks themselves are still discarded (see the reporter passed + // to `executeSubagent` below): only the bounded tick is surfaced. + let prunerChunks = 0 // Recursion guard: the pruner's own run evaluates this same semantic trigger, // so a run that IS the pruner must never drive a nested runtime pass. Matched @@ -175,8 +189,16 @@ export async function runRuntimeSemanticCompaction( parentSystemPrompt: system, parentTools: tools, // The pruner is infrastructure, not conversation: its output stays - // invisible, matching the inline path's pruner-identity suppression. - onResponseChunk: () => {}, + // invisible, matching the inline path's pruner-identity suppression. The + // chunk is dropped and only a bounded progress tick is reported, so the UI + // can show real movement instead of a silent stall. + onResponseChunk: () => { + prunerChunks += 1 + params.onCompactionProgress?.( + 'summarizing', + Math.min(85, 45 + prunerChunks * 5), + ) + }, clearUserPromptMessagesAfterResponse: false, }) diff --git a/sdk/src/__tests__/file-change-hooks.test.ts b/sdk/src/__tests__/file-change-hooks.test.ts index e3d58724b7..24511a212c 100644 --- a/sdk/src/__tests__/file-change-hooks.test.ts +++ b/sdk/src/__tests__/file-change-hooks.test.ts @@ -226,7 +226,7 @@ describe('runFileChangeHooks', () => { expect(results).toBeDefined() expect(results).toHaveLength(1) expect(results![0]).toMatchObject({ hookName: 'slow', exitCode: 0 }) - // The configured timeout (30s) is forwarded, not the 180s default. + // The configured 30s bound is forwarded instead of the unbounded default. expect(paramsList).toHaveLength(1) expect(paramsList[0]).toMatchObject({ command: 'tsc', @@ -234,6 +234,28 @@ describe('runFileChangeHooks', () => { }) }) + test('forwards no timeout to runCommand when a hook omits timeoutSeconds', async () => { + const { run, paramsList } = fakeRunner({ + tsc: { exitCode: 0 }, + }) + const out = await runFileChangeHooks({ + files: ['src/a.ts'], + cwd: '/repo', + hooks: [{ name: 'typecheck', command: 'tsc' }], + runCommand: run, + }) + const results = jsonValue(out) + expect(results).toBeDefined() + expect(results).toHaveLength(1) + expect(results![0]).toMatchObject({ hookName: 'typecheck', exitCode: 0 }) + // No per-hook bound configured: -1 means the hook runs unbounded. + expect(paramsList).toHaveLength(1) + expect(paramsList[0]).toMatchObject({ + command: 'tsc', + timeout_seconds: -1, + }) + }) + test('runs per-file syntax hooks only for safe matching changed files', async () => { const { run, calls } = fakeRunner({ "php -l 'src/a.php' && php -l 'src/space file.php'": { exitCode: 0 }, diff --git a/sdk/src/__tests__/tool-execution-deadline.test.ts b/sdk/src/__tests__/tool-execution-deadline.test.ts deleted file mode 100644 index c83f4d3070..0000000000 --- a/sdk/src/__tests__/tool-execution-deadline.test.ts +++ /dev/null @@ -1,65 +0,0 @@ -import { describe, expect, it } from 'bun:test' - -import { - createToolExecutionDeadline, - FILE_MUTATION_TOOL_TIMEOUT_MS, - getDefaultToolExecutionTimeoutMs, -} from '../tool-execution-deadline' - -describe('tool execution deadlines', () => { - it('bounds file mutations while leaving interactive and self-bounded tools alone', () => { - expect(getDefaultToolExecutionTimeoutMs('edit_transaction')).toBe( - FILE_MUTATION_TOOL_TIMEOUT_MS, - ) - expect(getDefaultToolExecutionTimeoutMs('write_file')).toBe( - FILE_MUTATION_TOOL_TIMEOUT_MS, - ) - expect(getDefaultToolExecutionTimeoutMs('ask_user')).toBeUndefined() - expect( - getDefaultToolExecutionTimeoutMs('run_terminal_command'), - ).toBeUndefined() - }) - - it('aborts a hung mutation with a non-run-cancellation timeout error', async () => { - const parent = new AbortController() - const deadline = createToolExecutionDeadline({ - parentSignal: parent.signal, - timeoutMs: 5, - toolName: 'edit_transaction', - }) - - try { - await new Promise((resolve) => { - deadline.signal.addEventListener('abort', () => resolve(), { - once: true, - }) - }) - expect(deadline.signal.aborted).toBe(true) - expect(deadline.signal.reason).toBeInstanceOf(Error) - expect(deadline.signal.reason.name).toBe('ToolExecutionTimeoutError') - expect(deadline.signal.reason.message).toContain( - 'no successful result is confirmed', - ) - } finally { - deadline.dispose() - } - }) - - it('propagates parent cancellation without replacing its reason', () => { - const parent = new AbortController() - const reason = new Error('user cancelled') - const deadline = createToolExecutionDeadline({ - parentSignal: parent.signal, - timeoutMs: 1_000, - toolName: 'edit_transaction', - }) - - try { - parent.abort(reason) - expect(deadline.signal.aborted).toBe(true) - expect(deadline.signal.reason).toBe(reason) - } finally { - deadline.dispose() - } - }) -}) diff --git a/sdk/src/impl/__tests__/direct-agent-tool-repair.test.ts b/sdk/src/impl/__tests__/direct-agent-tool-repair.test.ts index 65a98f8276..c0ed3dced7 100644 --- a/sdk/src/impl/__tests__/direct-agent-tool-repair.test.ts +++ b/sdk/src/impl/__tests__/direct-agent-tool-repair.test.ts @@ -67,7 +67,6 @@ describe('direct agent tool repair', () => { input: { prompt: 'Run the command', background: true, - timeout_seconds: 90, params: JSON.stringify({ command: '["literal-shell-token"]', what_to_summarize: '{"keep":"as text"}', @@ -80,7 +79,6 @@ describe('direct agent tool repair', () => { agent_type: 'basher', prompt: 'Run the command', background: true, - timeout_seconds: 90, params: { command: '["literal-shell-token"]', what_to_summarize: '{"keep":"as text"}', diff --git a/sdk/src/impl/direct-agent-tool-repair.ts b/sdk/src/impl/direct-agent-tool-repair.ts index 0ecb29fef0..91f5763ac3 100644 --- a/sdk/src/impl/direct-agent-tool-repair.ts +++ b/sdk/src/impl/direct-agent-tool-repair.ts @@ -13,7 +13,7 @@ export function buildSpawnAgentsInputForDirectAgentCall(params: { const input = parsed as DirectAgentInput const entry: DirectAgentInput = { agent_type: params.agentType } - for (const key of ['prompt', 'background', 'timeout_seconds']) { + for (const key of ['prompt', 'background']) { if (Object.prototype.hasOwnProperty.call(input, key)) { entry[key] = input[key] } @@ -27,8 +27,7 @@ export function buildSpawnAgentsInputForDirectAgentCall(params: { } else { const legacyParams = Object.fromEntries( Object.entries(input).filter( - ([key]) => - !['prompt', 'handoff', 'background', 'timeout_seconds'].includes(key), + ([key]) => !['prompt', 'handoff', 'background'].includes(key), ), ) if (Object.keys(legacyParams).length > 0) entry.params = legacyParams diff --git a/sdk/src/provider-config.ts b/sdk/src/provider-config.ts index 633910725c..c69327553c 100644 --- a/sdk/src/provider-config.ts +++ b/sdk/src/provider-config.ts @@ -462,7 +462,7 @@ export const providerConfigFileSchema = z name: z.string().min(1).optional(), command: z.string().min(1), filePattern: z.string().min(1).optional(), - /** Per-hook override of the default 180s hook timeout, in seconds. */ + /** Optional per-hook wall-clock bound in seconds. Omitted means no timeout. */ timeoutSeconds: z.number().int().positive().max(3600).optional(), }), ) diff --git a/sdk/src/run.ts b/sdk/src/run.ts index 7e4ac95d79..cd05862c23 100644 --- a/sdk/src/run.ts +++ b/sdk/src/run.ts @@ -98,10 +98,6 @@ import { writeAuditFindings, } from './tools/write-audit-findings' import { createNodeFileSystem } from './tools/node-filesystem' -import { - createToolExecutionDeadline, - getDefaultToolExecutionTimeoutMs, -} from './tool-execution-deadline' import type { FilesystemAuthorityPolicy } from './tools/filesystem-authority' import type { CustomToolDefinition } from './custom-tool' @@ -999,119 +995,108 @@ async function runOnce({ if (cloneMatch?.[1]) ownedLibrarianCloneDirs.add(cloneMatch[1]) } } - const timeoutMs = getDefaultToolExecutionTimeoutMs(toolName) - const deadline = createToolExecutionDeadline({ - parentSignal: toolSignal ?? runSignal, - timeoutMs, - toolName, + const handled = await handleToolCall({ + action: { + type: 'tool-call-request', + requestId: callId ?? crypto.randomUUID(), + userInputId, + toolName, + input, + mcpConfig, + }, + overrides: overrideTools ?? {}, + onFilesChanged, + onFilesystemMutation, + verifyExternalMutation, + customToolDefinitions: customToolDefinitions + ? Object.fromEntries( + customToolDefinitions.map((def) => [def.toolName, def]), + ) + : {}, + cwd, + fs, + fileFilter, + filesystemPolicy, + trustedJobOwner, + logger, + capabilityIssuer: cwd + ? { + projectId: cwd, + runId: + sessionState.mainAgentState.runId ?? + sessionState.mainAgentState.agentId, + } + : undefined, + env, + harnessStateDir: resolvedHarnessStateDir, + approvalReceiptIds, + approvalMode, + requestApproval, + approvalService, + harnessWorkspaceIdentity: workspaceJournal + ? { + repositoryId: workspaceJournal.repositoryId, + workspaceId: workspaceJournal.workspaceId, + } + : undefined, + getWorkspaceState: () => sessionState.mainAgentState.workspaceState, + setWorkspaceState: (state) => { + sessionState.mainAgentState.workspaceState = state + }, + advanceWorkspaceJournal: workspaceJournal + ? (change) => + (() => { + if (!workspaceJournal) { + return advanceWorkspaceState( + sessionState.mainAgentState.workspaceState, + change, + ) + } + try { + return workspaceJournal.advance({ + runId: + sessionState.mainAgentState.runId ?? + sessionState.mainAgentState.agentId, + ...change, + }) + } catch (error) { + logger?.warn( + { error }, + 'Workspace journal write failed; continuing with in-memory workspace state', + ) + workspaceJournal = undefined + return advanceWorkspaceState( + sessionState.mainAgentState.workspaceState, + change, + ) + } + })() + : undefined, + signal: toolSignal ?? runSignal, }) - try { - const handled = await handleToolCall({ - action: { - type: 'tool-call-request', - requestId: callId ?? crypto.randomUUID(), - userInputId, - toolName, - input, - timeout: timeoutMs, - mcpConfig, - }, - overrides: overrideTools ?? {}, - onFilesChanged, - onFilesystemMutation, - verifyExternalMutation, - customToolDefinitions: customToolDefinitions - ? Object.fromEntries( - customToolDefinitions.map((def) => [def.toolName, def]), - ) - : {}, - cwd, - fs, - fileFilter, - filesystemPolicy, - trustedJobOwner, - logger, - capabilityIssuer: cwd - ? { - projectId: cwd, - runId: - sessionState.mainAgentState.runId ?? - sessionState.mainAgentState.agentId, - } - : undefined, - env, - harnessStateDir: resolvedHarnessStateDir, - approvalReceiptIds, - approvalMode, - requestApproval, - approvalService, - harnessWorkspaceIdentity: workspaceJournal - ? { - repositoryId: workspaceJournal.repositoryId, - workspaceId: workspaceJournal.workspaceId, - } - : undefined, - getWorkspaceState: () => sessionState.mainAgentState.workspaceState, - setWorkspaceState: (state) => { - sessionState.mainAgentState.workspaceState = state - }, - advanceWorkspaceJournal: workspaceJournal - ? (change) => - (() => { - if (!workspaceJournal) { - return advanceWorkspaceState( - sessionState.mainAgentState.workspaceState, - change, - ) - } - try { - return workspaceJournal.advance({ - runId: - sessionState.mainAgentState.runId ?? - sessionState.mainAgentState.agentId, - ...change, - }) - } catch (error) { - logger?.warn( - { error }, - 'Workspace journal write failed; continuing with in-memory workspace state', - ) - workspaceJournal = undefined - return advanceWorkspaceState( - sessionState.mainAgentState.workspaceState, - change, - ) - } - })() - : undefined, - signal: deadline.signal, - }) - // Intercept the single dispatch path (model- and agent-initiated calls - // alike) so an unchanged list_jobs digest doesn't re-inject the full - // table into the conversation every step. The gate owns the returned - // output; the per-turn fingerprint lives in this closure. - if (toolName === 'list_jobs') { - const gated = applyListJobsDigestGate( - lastListJobsFingerprint, - handled.output, - ) - lastListJobsFingerprint = gated.nextFingerprint - return { ...handled, output: gated.output } - } else if (toolName === 'git_status') { - // This interception runs AFTER the tool executed (`handled.output`), - // so every git_status observation still runs; only its context - // encoding is compacted when the worktree is byte-identical. - const gated = applyGitStatusGate( - lastGitStatusFingerprint, - handled.output, - ) - lastGitStatusFingerprint = gated.nextFingerprint - return { ...handled, output: gated.output } - } - return handled - } finally { - deadline.dispose() + // Intercept the single dispatch path (model- and agent-initiated calls + // alike) so an unchanged list_jobs digest doesn't re-inject the full + // table into the conversation every step. The gate owns the returned + // output; the per-turn fingerprint lives in this closure. + if (toolName === 'list_jobs') { + const gated = applyListJobsDigestGate( + lastListJobsFingerprint, + handled.output, + ) + lastListJobsFingerprint = gated.nextFingerprint + return { ...handled, output: gated.output } + } else if (toolName === 'git_status') { + // This interception runs AFTER the tool executed (`handled.output`), + // so every git_status observation still runs; only its context + // encoding is compacted when the worktree is byte-identical. + const gated = applyGitStatusGate( + lastGitStatusFingerprint, + handled.output, + ) + lastGitStatusFingerprint = gated.nextFingerprint + return { ...handled, output: gated.output } } + return handled }, requestMcpToolData: async ({ mcpConfig, toolNames }) => { const mcpClientId = await getMCPClient(mcpConfig) diff --git a/sdk/src/tool-execution-deadline.ts b/sdk/src/tool-execution-deadline.ts deleted file mode 100644 index e48fde1c90..0000000000 --- a/sdk/src/tool-execution-deadline.ts +++ /dev/null @@ -1,48 +0,0 @@ -import type { ToolName } from '@codebuff/common/tools/constants' - -export const FILE_MUTATION_TOOL_TIMEOUT_MS = 120_000 - -const FILE_MUTATION_TOOLS = new Set([ - 'create_plan', - 'edit_transaction', - 'replace_range', - 'rewrite_symbol', - 'str_replace', - 'update_plan_status', - 'write_file', - 'write_audit_findings', -]) - -export function getDefaultToolExecutionTimeoutMs( - toolName: string, -): number | undefined { - return FILE_MUTATION_TOOLS.has(toolName as ToolName) - ? FILE_MUTATION_TOOL_TIMEOUT_MS - : undefined -} - -export function createToolExecutionDeadline(params: { - parentSignal: AbortSignal - timeoutMs: number | undefined - toolName: string -}): { signal: AbortSignal; dispose: () => void } { - const { parentSignal, timeoutMs, toolName } = params - if (timeoutMs === undefined || timeoutMs <= 0) { - return { signal: parentSignal, dispose: () => {} } - } - - const timeoutController = new AbortController() - const timeout = setTimeout(() => { - const error = new Error( - `${toolName} timed out after ${Math.ceil(timeoutMs / 1000)} seconds. The operation was cancelled and no successful result is confirmed.`, - ) - error.name = 'ToolExecutionTimeoutError' - timeoutController.abort(error) - }, timeoutMs) - timeout.unref?.() - - return { - signal: AbortSignal.any([parentSignal, timeoutController.signal]), - dispose: () => clearTimeout(timeout), - } -} diff --git a/sdk/src/tools/concurrency.ts b/sdk/src/tools/concurrency.ts new file mode 100644 index 0000000000..c4a1947ce5 --- /dev/null +++ b/sdk/src/tools/concurrency.ts @@ -0,0 +1,76 @@ +// Shared bounded fan-out for SDK tools. A tool that maps over N paths must not +// open N concurrent filesystem operations, so every such tool routes through +// this one helper instead of carrying its own copy of the loop. + +/** + * Rethrows the signal's own `reason` when it carries one, so a caller-supplied + * cancellation error reaches the caller unchanged instead of being replaced by + * a generic abort. Lives next to the bounded loop below because that loop is + * what has to observe cancellation between items. + */ +export function throwIfAborted(signal?: AbortSignal): void { + if (!signal?.aborted) return + throw signal.reason instanceof Error + ? signal.reason + : new DOMException('Operation aborted', 'AbortError') +} + +/** + * Applies `map` over `values` with at most `concurrency` calls in flight. + * + * Results are index-aligned with `values`, so a caller may keep accounting that + * depends on input order. Abort is checked between items rather than mid-flight: + * an already-started operation settles, and the next one throws. + * + * Failure is bounded the same way, and is part of the contract: the first + * mapper rejection (or abort) wins, no further item is handed out, every + * already-started call is awaited, and only then is that first error rethrown. + * So a rejection never leaves workers issuing filesystem work in the background + * that the caller can no longer observe or await, and items past the failure + * point are never started at all. + * + * Precondition: `concurrency` must be a positive integer. A value below 1, a + * non-integer, or a non-finite one would size the worker pool at zero (or NaN) + * and resolve a full-length array of holes with no mapper having run, so it + * throws a `RangeError` naming the received value before anything is allocated + * or scheduled — including when `values` is empty, since an invalid limit is a + * caller bug whether or not there is work to do. + */ +export async function mapWithConcurrency( + values: readonly T[], + concurrency: number, + map: (value: T, index: number) => Promise, + signal?: AbortSignal, +): Promise { + // Checked before `results` is allocated and before anything is scheduled: a + // limit below 1 (or a non-integer/non-finite one) spawns zero workers, which + // would silently resolve an array of holes instead of failing. Validated + // regardless of `values.length`, because an invalid limit is a caller bug + // whether or not there is work to do. + if (!Number.isInteger(concurrency) || concurrency < 1) { + throw new RangeError( + `mapWithConcurrency requires a positive integer concurrency, received ${concurrency}: a limit below 1 spawns no workers, so no mapper would run and the call would resolve with holes.`, + ) + } + const results = new Array(values.length) + let nextIndex = 0 + // Boxed so a thrown `undefined` still registers as a failure. Workers settle + // normally and the error is rethrown below, which is what lets every started + // call finish before this function returns. + let failure: { error: unknown } | undefined + await Promise.all( + Array.from({ length: Math.min(concurrency, values.length) }, async () => { + while (!failure && nextIndex < values.length) { + try { + throwIfAborted(signal) + const index = nextIndex++ + results[index] = await map(values[index]!, index) + } catch (error) { + failure ??= { error } + } + } + }), + ) + if (failure) throw failure.error + return results +} diff --git a/sdk/src/tools/file-change-hooks.ts b/sdk/src/tools/file-change-hooks.ts index 46aca11010..17930b1f33 100644 --- a/sdk/src/tools/file-change-hooks.ts +++ b/sdk/src/tools/file-change-hooks.ts @@ -18,13 +18,15 @@ export type FileChangeHook = { command: string /** Optional glob; the hook runs only when a changed file matches it. */ filePattern?: string - /** Optional per-hook override of the default 180s hook timeout, in seconds. */ + /** Optional per-hook wall-clock bound in seconds. Omitted means no timeout. */ timeoutSeconds?: number /** Run the command once per matching changed file instead of project-wide. */ runPerFile?: boolean } -const HOOK_TIMEOUT_SECONDS = 180 +// Hooks are unbounded by default; -1 means no timeout. A project opts into a +// bound with the per-hook `timeoutSeconds`. +const HOOK_DEFAULT_TIMEOUT_SECONDS = -1 const MAX_HOOK_OUTPUT_CHARS = 6000 const MAX_MANIFEST_BYTES = 512_000 const MAX_PROJECT_SCAN_ENTRIES = 2_000 @@ -769,7 +771,7 @@ export async function runFileChangeHooks(params: { process_type: 'SYNC', cwd, projectRoot: cwd, - timeout_seconds: hook.timeoutSeconds ?? HOOK_TIMEOUT_SECONDS, + timeout_seconds: hook.timeoutSeconds ?? HOOK_DEFAULT_TIMEOUT_SECONDS, env, signal: params.signal, }) From 3aef2c06b1a09be7795d20dacf5ebff8044bad4b Mon Sep 17 00:00:00 2001 From: AnzoBenjamin Date: Sat, 5 Sep 2026 15:59:30 +0300 Subject: [PATCH 2/4] fix: resolving some review attestation issues. --- agents/reviewer/code-reviewer.ts | 2 +- common/src/types/session-state.ts | 2 + .../src/__tests__/loop-agent-steps.test.ts | 113 ++++++++++++++ packages/agent-runtime/src/run-agent-step.ts | 28 +++- .../tool/__tests__/set-output.test.ts | 145 ++++++++++++++++++ .../src/tools/handlers/tool/set-output.ts | 15 +- 6 files changed, 296 insertions(+), 9 deletions(-) diff --git a/agents/reviewer/code-reviewer.ts b/agents/reviewer/code-reviewer.ts index 4e10046364..0dc5c83f71 100644 --- a/agents/reviewer/code-reviewer.ts +++ b/agents/reviewer/code-reviewer.ts @@ -183,7 +183,7 @@ Before you emit a verdict, read every file in the pending list with read_files; Validation and other subagent work may be running in parallel with your review. You cannot observe results from parallel agents unless the prompt explicitly includes those completed results. If validation results are not included, treat your review as static code review only: do not say validation passed or failed, do not ask for a generic rerun just because results are absent, and only request validation when you see a concrete code-specific reason that a particular command or scenario must be checked. -Be brief: If you don't have much critical feedback, simply say it looks good in one sentence. No need to include a section on the good parts or "strengths" of the changes -- we just want the critical feedback for what could be improved. +Be brief: when nothing requires a change, leave \`findings\` empty and keep each \`dimensions\` value to one short clause — the verdict travels only through \`set_output\`, never as a prose reply. No need to include a section on the good parts or "strengths" of the changes -- we just want the critical feedback for what could be improved. Return the structured output required by your output schema with schemaVersion 1. The parent prompt supplies an opaque single-line snapshot fingerprint, a separate snapshot-details block, and a pending file list. Copy only the fingerprint token into snapshotFingerprint; do not copy the multiline details. List every file you actually read using the exact normalized project-relative path from the pending list (forward slashes, including directories such as __tests__). Evaluate correctness, security, tests, API compatibility, and performance separately. Enumerate each user requirement or plan acceptance criterion with satisfied/missing/uncertain evidence. Parent-owned process tasks (git rewrite/amend, commit/push, confirm CI green, operator validation already owned by the harness gate) are out of scope for requirementCoverage — omit them or do not let them alone force BLOCKING. Still require BLOCKING for incomplete implementation/source requirements and acceptance criteria the code change claims to satisfy: if ANY in-scope \`requirementCoverage[].status\` is \`missing\` or \`uncertain\`, the top-level \`verdict\` MUST be \`"BLOCKING"\` — never NON_BLOCKING or LOOKS_GOOD while those requirements are incomplete — and put each incomplete in-scope requirement into \`findings\` as a concrete next action. diff --git a/common/src/types/session-state.ts b/common/src/types/session-state.ts index 1f04c1d0e8..e7563dec03 100644 --- a/common/src/types/session-state.ts +++ b/common/src/types/session-state.ts @@ -137,6 +137,8 @@ export type AgentState = { repeatedStepProgressCount?: number /** Consecutive text-only turns without task_completed for explicit-completion agents (bounded fallback, resets on tool use). */ consecutiveTextOnlyWithoutCompletion?: number + /** Message from the most recent rejected set_output call, cleared once output is successfully set. Used to make the missing-structured-output retry name the real failure. */ + lastSetOutputError?: string creditsUsed: number directCreditsUsed: number /** diff --git a/packages/agent-runtime/src/__tests__/loop-agent-steps.test.ts b/packages/agent-runtime/src/__tests__/loop-agent-steps.test.ts index 9d8d4ec04f..febbc5f5c2 100644 --- a/packages/agent-runtime/src/__tests__/loop-agent-steps.test.ts +++ b/packages/agent-runtime/src/__tests__/loop-agent-steps.test.ts @@ -1852,4 +1852,117 @@ describe('loopAgentSteps', () => { ledger!.byCategory, ) }) + + // Regression: a structured agent that never populates output used to get only + // one retry. The fix introduces MAX_MISSING_OUTPUT_RETRIES (= 3) so the loop + // retries the missing-output nudge up to the cap before handing back null. + it('retries the missing-output nudge up to the cap and then returns null value', async () => { + setup() + let llmCallCount = 0 + agentTemplate.handleSteps = undefined + agentTemplate.toolNames = ['read_files'] + agentTemplate.outputSchema = z.object({ result: z.string() }) + runtimeParams.promptAiSdkStream = mock(async function* () { + llmCallCount++ + // The LLM only ever returns prose — it never calls set_output. + yield { type: 'text' as const, text: 'Still no structured output.' } + return promptSuccess(`mock-message-${llmCallCount}`) + }) + + const result = await loopAgentSteps({ + ...baseParams, + // More steps than the retry budget so the loop ends because the retry cap + // is hit, not because steps ran out. + agentState: { ...agentState, stepsRemaining: 20 }, + promptAiSdkStream: runtimeParams.promptAiSdkStream, + localAgentTemplates: { 'test-agent': agentTemplate }, + }) + + // 1 initial call + 3 retries = MAX_MISSING_OUTPUT_RETRIES (hard-coded 4; + // the constant is module-private in run-agent-step.ts). The old one-shot + // behavior yielded 2. + expect(llmCallCount).toBe(4) + // The structured-output envelope is returned with a null/absent value + // rather than throwing. + expect(result.output).toEqual({ + type: 'structuredOutput', + value: null, + }) + }) + + // Regression: when a set_output call is rejected (schema validation), the + // missing-output retry must name the actual rejection so the model can fix + // the reported fields, instead of a generic reminder. + it('injects a retry message that names the recorded set_output rejection', async () => { + setup() + let llmCallCount = 0 + agentTemplate.handleSteps = undefined + agentTemplate.toolNames = ['read_files', 'set_output'] + agentTemplate.outputSchema = z.object({ result: z.string() }) + runtimeParams.promptAiSdkStream = mock(async function* () { + llmCallCount++ + // Always call set_output with a payload that FAILS the outputSchema + // (number where a string is required), so output stays unset. + yield createToolCallChunk('set_output', { result: 42 }) + return promptSuccess(`mock-message-${llmCallCount}`) + }) + + const result = await loopAgentSteps({ + ...baseParams, + // The set_output handler resolves the outputSchema from + // agentState.agentType (not params.agentType), so the state must name + // the registered template id or validation is silently skipped. + agentState: { + ...agentState, + agentType: 'test-agent', + stepsRemaining: 20, + }, + promptAiSdkStream: runtimeParams.promptAiSdkStream, + localAgentTemplates: { 'test-agent': agentTemplate }, + }) + + // The injected nudge is appended to message history (not onResponseChunk), + // so observe it there. + const history = result.agentState.messageHistory + .map((message) => + typeof message.content === 'string' + ? message.content + : JSON.stringify(message.content), + ) + .join('\n') + expect(history).toContain('Your set_output call was rejected') + expect(history).toContain('Output validation error') + }) + + // Regression: when the agent simply never calls set_output (no recorded + // rejection), the retry must fall back to the generic wording. + it('injects the generic missing-output message when there is no recorded rejection', async () => { + setup() + let llmCallCount = 0 + agentTemplate.handleSteps = undefined + agentTemplate.toolNames = ['read_files'] + agentTemplate.outputSchema = z.object({ result: z.string() }) + runtimeParams.promptAiSdkStream = mock(async function* () { + llmCallCount++ + yield { type: 'text' as const, text: 'Just prose, no tool call.' } + return promptSuccess(`mock-message-${llmCallCount}`) + }) + + const result = await loopAgentSteps({ + ...baseParams, + agentState: { ...agentState, stepsRemaining: 20 }, + promptAiSdkStream: runtimeParams.promptAiSdkStream, + localAgentTemplates: { 'test-agent': agentTemplate }, + }) + + const history = result.agentState.messageHistory + .map((message) => + typeof message.content === 'string' + ? message.content + : JSON.stringify(message.content), + ) + .join('\n') + expect(history).toContain('does not populate structured output') + expect(history).not.toContain('Your set_output call was rejected') + }) }) diff --git a/packages/agent-runtime/src/run-agent-step.ts b/packages/agent-runtime/src/run-agent-step.ts index 1b9e625231..b1f5c5ab3e 100644 --- a/packages/agent-runtime/src/run-agent-step.ts +++ b/packages/agent-runtime/src/run-agent-step.ts @@ -1748,7 +1748,7 @@ export async function loopAgentSteps( } let shouldEndTurn = false - let hasRetriedOutputSchema = false + let outputSchemaRetryCount = 0 let currentPrompt = prompt let currentParams = spawnParams let totalSteps = 0 @@ -2432,21 +2432,32 @@ export async function loopAgentSteps( !agentTemplate.handleSteps && currentAgentState.output === undefined && shouldEndTurn && - !hasRetriedOutputSchema + outputSchemaRetryCount < MAX_MISSING_OUTPUT_RETRIES ) { - hasRetriedOutputSchema = true + outputSchemaRetryCount += 1 + // The set_output handler records its rejection on this same agent + // state object, so the retry names the real validation failure + // instead of a generic reminder the model cannot act on. + const rejection = currentAgentState.lastSetOutputError logger.warn( { agentType, agentId: currentAgentState.agentId, runId, + outputSchemaRetryCount, + // Flag only: the rejection text embeds the original output value + // and can be large. + hadSetOutputRejection: rejection !== undefined, }, 'Agent finished without setting required output, restarting loop', ) // Add system message instructing to use set_output + const attemptSuffix = `(attempt ${outputSchemaRetryCount} of ${MAX_MISSING_OUTPUT_RETRIES})` const outputSchemaMessage = withSystemTags( - `You must use the "set_output" tool to provide a result that matches the output schema before ending your turn. The output schema is required for this agent.`, + rejection + ? `Your set_output call was rejected and your output is still unset ${attemptSuffix}. Fix exactly the reported fields and call set_output again with native object/array values. Rejection: ${rejection}` + : `You must use the "set_output" tool to provide a result that matches the output schema before ending your turn ${attemptSuffix}. A prose answer or a Markdown JSON block does not populate structured output; the parent receives null.`, ) currentAgentState.messageHistory = [ @@ -2693,6 +2704,15 @@ const buildCompactionNoProgressClause = ( COMPACTION_NO_PROGRESS_FRACTION * 100, )}%.` +/** + * Bounded retries for a structured-output agent that ended its turn without + * setting required output. A failed `set_output` ends the turn (set_output is in + * TOOLS_WHICH_WONT_FORCE_NEXT_STEP), so a single retry left a reviewer that + * botched one field with no step to act on the validation error, and the parent + * received `value: null`. Bounded, not unlimited: each retry is one more LLM step. + */ +const MAX_MISSING_OUTPUT_RETRIES = 3 + /** * How many steps before the cap the one-time near-cap checkpoint nudge fires. * Compared with `===` against the per-step-decrementing stepsRemaining, so it diff --git a/packages/agent-runtime/src/tools/handlers/tool/__tests__/set-output.test.ts b/packages/agent-runtime/src/tools/handlers/tool/__tests__/set-output.test.ts index 3c813058c0..9b7cf216c0 100644 --- a/packages/agent-runtime/src/tools/handlers/tool/__tests__/set-output.test.ts +++ b/packages/agent-runtime/src/tools/handlers/tool/__tests__/set-output.test.ts @@ -556,4 +556,149 @@ describe('handleSetOutput', () => { expect(output).toEqual([{ type: 'json', value: { message: 'Output set' } }]) expect(agentState.output).toEqual(inner) }) + + test('records a schema-validation rejection on agentState.lastSetOutputError', async () => { + const template: AgentTemplate = { + id: 'reviewer-error-test', + displayName: 'Reviewer Error Test', + spawnerPrompt: 'Review code', + model: 'claude-3-5-sonnet-20241022', + inputSchema: {}, + outputMode: 'structured_output', + outputSchema: z.object({ reviewedFiles: z.array(z.string()) }), + includeMessageHistory: false, + inheritParentSystemPrompt: false, + mcpServers: {}, + toolNames: ['set_output'], + spawnableAgents: [], + systemPrompt: 'Test system prompt', + instructionsPrompt: 'Test instructions', + stepPrompt: 'Test step prompt', + } + const agentState = getInitialSessionState(mockFileContext).mainAgentState + agentState.agentType = template.id + const toolCall = { + toolName: 'set_output', + toolCallId: 'invalid-review-output', + input: { reviewedFiles: 'not-json' }, + } as unknown as CodebuffToolCall<'set_output'> + + const { output } = await handleSetOutput({ + ...TEST_AGENT_RUNTIME_IMPL, + previousToolCallFinished: Promise.resolve(), + toolCall, + agentState, + apiKey: 'test-api-key', + localAgentTemplates: { [template.id]: template }, + } as unknown as Parameters[0]) + + const message = output[0]?.type === 'json' ? output[0].value.message : '' + expect(typeof agentState.lastSetOutputError).toBe('string') + expect(agentState.lastSetOutputError).toContain('Output validation error') + // The recorded error must be exactly the message handed back to the model, + // so the loop's retry nudge and the tool result cannot drift apart. + expect(agentState.lastSetOutputError).toBe(message) + }) + + test('clears lastSetOutputError once a later set_output succeeds', async () => { + const template: AgentTemplate = { + id: 'reviewer-error-test', + displayName: 'Reviewer Error Test', + spawnerPrompt: 'Review code', + model: 'claude-3-5-sonnet-20241022', + inputSchema: {}, + outputMode: 'structured_output', + outputSchema: z.object({ reviewedFiles: z.array(z.string()) }), + includeMessageHistory: false, + inheritParentSystemPrompt: false, + mcpServers: {}, + toolNames: ['set_output'], + spawnableAgents: [], + systemPrompt: 'Test system prompt', + instructionsPrompt: 'Test instructions', + stepPrompt: 'Test step prompt', + } + const agentState = getInitialSessionState(mockFileContext).mainAgentState + agentState.agentType = template.id + + // Drive a rejecting call first so the error field is populated. + const rejectingCall = { + toolName: 'set_output', + toolCallId: 'invalid-review-output', + input: { reviewedFiles: 'not-json' }, + } as unknown as CodebuffToolCall<'set_output'> + + await handleSetOutput({ + ...TEST_AGENT_RUNTIME_IMPL, + previousToolCallFinished: Promise.resolve(), + toolCall: rejectingCall, + agentState, + apiKey: 'test-api-key', + localAgentTemplates: { [template.id]: template }, + } as unknown as Parameters[0]) + + expect(agentState.lastSetOutputError).toBeDefined() + + // A subsequent successful call on the SAME agentState must supersede the + // earlier rejection: a stale rejection must never leak into a later + // missing-output retry message. + const validCall = { + toolName: 'set_output', + toolCallId: 'valid-review-output', + input: { reviewedFiles: ['src/a.ts'] }, + } as unknown as CodebuffToolCall<'set_output'> + + await handleSetOutput({ + ...TEST_AGENT_RUNTIME_IMPL, + previousToolCallFinished: Promise.resolve(), + toolCall: validCall, + agentState, + apiKey: 'test-api-key', + localAgentTemplates: { [template.id]: template }, + } as unknown as Parameters[0]) + + expect(agentState.lastSetOutputError).toBeUndefined() + expect(agentState.output).toEqual({ reviewedFiles: ['src/a.ts'] }) + }) + + test('records a malformed-JSON rejection on agentState.lastSetOutputError', async () => { + const template: AgentTemplate = { + id: 'reviewer-test', + displayName: 'Reviewer Test', + spawnerPrompt: 'Review code', + model: 'claude-3-5-sonnet-20241022', + inputSchema: {}, + outputMode: 'structured_output', + outputSchema: z.object({ verdict: z.string() }), + includeMessageHistory: false, + inheritParentSystemPrompt: false, + mcpServers: {}, + toolNames: ['set_output'], + spawnableAgents: [], + systemPrompt: 'Test system prompt', + instructionsPrompt: 'Test instructions', + stepPrompt: 'Test step prompt', + } + const agentState = getInitialSessionState(mockFileContext).mainAgentState + agentState.agentType = template.id + const toolCall = { + toolName: 'set_output', + toolCallId: 'incomplete-review-output', + input: { data: '{"foo":' }, + } as unknown as CodebuffToolCall<'set_output'> + + await handleSetOutput({ + ...TEST_AGENT_RUNTIME_IMPL, + previousToolCallFinished: Promise.resolve(), + toolCall, + agentState, + apiKey: 'test-api-key', + localAgentTemplates: { [template.id]: template }, + } as unknown as Parameters[0]) + + expect(agentState.lastSetOutputError).toContain( + 'malformed or incomplete JSON text', + ) + expect(agentState.output).toBeUndefined() + }) }) diff --git a/packages/agent-runtime/src/tools/handlers/tool/set-output.ts b/packages/agent-runtime/src/tools/handlers/tool/set-output.ts index c38b2f137d..20c193ff71 100644 --- a/packages/agent-runtime/src/tools/handlers/tool/set-output.ts +++ b/packages/agent-runtime/src/tools/handlers/tool/set-output.ts @@ -37,11 +37,14 @@ export const handleSetOutput = (async (params: { const rawOutput = toolCall.input as Record const decodedData = decodeJsonObjectString(rawOutput?.data) if (typeof rawOutput?.data === 'string' && decodedData === rawOutput.data) { + const malformedJsonMessage = + 'Output was not set because data contained malformed or incomplete JSON text. Retry set_output with a real object value, not JSON.stringify(...). Keep findings and evidence concise enough to complete one tool call.' + // Recorded on agent state because this rejection also ends the turn + // (set_output is in TOOLS_WHICH_WONT_FORCE_NEXT_STEP), so the loop's + // missing-output retry is the only place that can report it to the model. + agentState.lastSetOutputError = malformedJsonMessage return { - output: jsonToolResult({ - message: - 'Output was not set because data contained malformed or incomplete JSON text. Retry set_output with a real object value, not JSON.stringify(...). Keep findings and evidence concise enough to complete one tool call.', - }), + output: jsonToolResult({ message: malformedJsonMessage }), } } const decodedOutput = @@ -167,6 +170,7 @@ export const handleSetOutput = (async (params: { }, 'set_output validation error', ) + agentState.lastSetOutputError = errorMessage return { output: jsonToolResult({ message: errorMessage }) } } } else { @@ -179,6 +183,9 @@ export const handleSetOutput = (async (params: { // Set the output (completely replaces previous output) agentState.output = finalOutput as Record + // A successful receipt supersedes any earlier rejection, so a later generic + // missing-output retry can never quote a stale one. + agentState.lastSetOutputError = undefined return { output: jsonToolResult({ message: 'Output set' }) } }) satisfies CodebuffToolHandlerFunction From fdc4670bad81bbbd230466d3a469cf10f944542e Mon Sep 17 00:00:00 2001 From: AnzoBenjamin Date: Sat, 5 Sep 2026 23:06:55 +0300 Subject: [PATCH 3/4] fix: resolving some review attestation issues. --- .../phase3-context-scale-up/EVENTS.jsonl | 2 + .../sessions/phase3-context-scale-up/PLAN.md | 63 ++ agents/__tests__/base2.test.ts | 320 +++++++++ agents/base2/base2.ts | 617 +++++++++++++++--- agents/base2/gate-state.ts | 53 ++ agents/e2e/gate-aux-ordering.e2e.test.ts | 401 ++++++++++++ .../components/renderers/compaction-box.tsx | 1 + cli/src/types/chat.ts | 1 + common/src/types/print-mode.ts | 5 + docs/agents-and-tools.md | 6 +- .../src/util/__tests__/messages.test.ts | 99 ++- packages/agent-runtime/src/util/messages.ts | 83 ++- 12 files changed, 1544 insertions(+), 107 deletions(-) create mode 100644 .agents/sessions/phase3-context-scale-up/EVENTS.jsonl create mode 100644 .agents/sessions/phase3-context-scale-up/PLAN.md diff --git a/.agents/sessions/phase3-context-scale-up/EVENTS.jsonl b/.agents/sessions/phase3-context-scale-up/EVENTS.jsonl new file mode 100644 index 0000000000..a3f67eb9c2 --- /dev/null +++ b/.agents/sessions/phase3-context-scale-up/EVENTS.jsonl @@ -0,0 +1,2 @@ +{"ts":"2026-09-05T18:53:14.100Z","kind":"task_update","summary":"Updated 1 task line(s): P3.1","payload":{"matched":["P3.1"]}} +{"ts":"2026-09-05T18:53:14.100Z","kind":"append_lesson","summary":"Appended entry \"Phase 3a landed (2026-09-05)\" to PLAN.md","payload":{"heading":"Phase 3a landed (2026-09-05)","artifact":"PLAN.md"}} diff --git a/.agents/sessions/phase3-context-scale-up/PLAN.md b/.agents/sessions/phase3-context-scale-up/PLAN.md new file mode 100644 index 0000000000..e92b6ebd48 --- /dev/null +++ b/.agents/sessions/phase3-context-scale-up/PLAN.md @@ -0,0 +1,63 @@ +# Phase 3: Scaled-up index + dynamic memory (search→read handoff) + +Decision record: user chose **Design + implement** with **Hybrid, measured** scope +(prompts/defaults first, measure context growth, add hard guardrails only where +measurements show drift). + +## Existing surfaces (verified 2026-09-05) + +- Guidance already present: reviewer instructions + spawn prompt already mandate + `read_files windows/around/symbol` for large files (agents/reviewer/code-reviewer.ts:180, + agents/base2/base2.ts:4740); orchestrator mandates carry the tiered read policy + (base2.ts:546, base-deep.ts:45-52); EXPLORE_PROMPT mandates query_index-first + discovery (base2.ts:11074). Editor large-file deterministic editing at editor.ts:122-128. +- Telemetry exists but is too coarse: `getContextCategory` lumps read_files / + find_files / read_subtree / read_outline / query_index into ONE `fileReads` + category (packages/agent-runtime/src/util/messages.ts:239-247). Whole-file + bodies and cheap structured slices are indistinguishable, so "is the + search→read handoff actually reducing read-token share?" is unanswerable. +- Compaction capture points already emit category before/after + (run-agent-step.ts:2024-2029, 2254-2256; ContextTrimReport.beforeCategories / + afterCategories / removedCategories). +- Dynamic memory: knowledge_memory pinned block survives compaction today + (extractPinnedContextBlocks, retainedKnowledgeMemory). A durable per-task read + ledger is DEFERRED until measurement shows the pinned block is insufficient. + +## Milestones + +- [x] P3.1 Design checkpoint (this document) — no code. (PLAN authored before implementation; no code changes in this task) +- [ ] P3.2 Phase 3a prompts: add the bounded search→read line to the specialist + scoped review prompt (`buildSpecialistScopedReviewPrompt`, base2.ts ~8599-8635) + so specialists get the same handoff reviewers already have. Do NOT alter the + existing reviewer prompt line at base2.ts:4740 (tests assert its exact text). +- [ ] P3.3 Phase 3a measurement: split the `fileReads` context category into + `boundedFileReads` (read_files with windows/around/symbol input, read_outline, + query_index, find_files) and whole-file `fileReads` (read_files paths selector, + read_subtree). Update ContextCategory union, summary initializer, + CONTEXT_EVICTION_PRIORITY (bounded reads cheapest-to-recover → evict first), + and additive tests in messages.test.ts. +- [ ] P3.4 Phase 3a validation: typecheck agent-runtime + agents; run messages, + context-pruner, and the base2 prompt-assertion tests. +- [ ] P3.5 Measurement window (manual, no code): after a few real tasks, compare + `boundedFileReads` vs `fileReads` token share from compaction telemetry. +- [ ] P3.6 DEFERRED hard guardrails (read_files auto-windowing past a size + threshold, handoff contract in spawn params): implement ONLY if M5 shows + whole-file share stays high or compaction churn recurs. Trigger, evidence, and + scope to be recorded here before any implementation. +- [ ] P3.7 DEFERRED dynamic-memory durable read ledger: only if M5 shows eviction + of needed reads (re-read churn after compaction). + +## Risks / dependencies + +- ContextCategory may be referenced in Record exhaustively; + every key site must be updated or typecheck fails (good — it fails loud). +- gate telemetry fixtures in common/src/testing may embed category summaries. +- Prompt additions must stay one line to avoid prompt-size growth. +- M2/M3 are independent (agents vs agent-runtime workspaces) and can land in + either order; both are small, additive, and locally validated like Phase 4. + + +## Phase 3a landed (2026-09-05) — 2026-09-05T18:53:14.099Z + +Phase 3a implemented and locally validated (gate receipts pending). Files: packages/agent-runtime/src/util/messages.ts (boundedFileReads category split + eviction-first ordering), messages.test.ts (additive telemetry + eviction tests, 55/55), common/src/types/print-mode.ts (optional boundedFileReads so replayed pre-split events keep validating), cli/src/types/chat.ts (CompactionCategoryDelta union + boundedFileReads), cli/src/components/renderers/compaction-box.tsx (exhaustive 'bounded reads' label), agents/base2/base2.ts (specialist scoped-prompt bounded-read line at :9051; final-reviewer line untouched). Validation evidence: common/agent-runtime/agents/cli typechecks exit 0; gate suites 273/273; sdk-event-handlers 70/70; sweep-boxes 27/27. P3.2-P3.4 task toggles deferred until the gate mints receipts; P3.6/P3.7 remain DEFERRED per hybrid-measured scope (trigger = P3.5 measurement window showing whole-file share stays high or re-read churn after compaction). + diff --git a/agents/__tests__/base2.test.ts b/agents/__tests__/base2.test.ts index 6cb5c6e093..9a381bf6bd 100644 --- a/agents/__tests__/base2.test.ts +++ b/agents/__tests__/base2.test.ts @@ -2294,6 +2294,317 @@ describe('base2 verification and reviewer gates', () => { }) }) + test('reuses the newest full-assurance validationEvidence receipt for unchanged bytes without re-running hooks', () => { + // Phase 4 per-file gate credit: the NEWEST validationEvidence entry with + // `assurance: 'full'`, files exactly equal to the gate scope, and content + // markers still matching the live bytes is reused IN PLACE — the hook run + // is skipped and the entry is kept verbatim (same summary/recordedAt). + const base2 = createBase2('default') + const tmpDir = makeProjectTempDir('base2-receipt-reuse-') + try { + const tmpFile = join(tmpDir, 'a.ts') + const gateFile = normalizeGateFilePath(tmpFile) + writeFileSync(tmpFile, 'export const value = 1\n') + const validationSummary = + 'Configured file-change hooks passed: typecheck.' + const recordedAt = '2025-01-01T00:00:00.000Z' + const agentState = { + agentId: 'base2-custom', + base2ActiveWork: { + changedFiles: [gateFile], + touchedFiles: [gateFile], + pendingGateFiles: [gateFile], + currentPhase: 'awaiting_validation', + latestWorkSummary: 'Pending gate previously validated.', + openReviewerBlockers: [], + lastValidationSummary: validationSummary, + nextRequiredAction: '', + lastPinnedStateMessage: '', + validationEvidence: [ + { + gateId: 'receipt-reuse-gate', + files: [gateFile], + snapshotFingerprint: 'seed-receipt-snapshot', + summary: validationSummary, + assurance: 'full', + recordedAt, + // Markers are built from the live bytes so the Phase-4 reuse + // predicate sees a byte-identical covered file set. + fileMarkers: { [gateFile]: buildContentMarker(tmpFile) }, + }, + ], + }, + } + const gen = base2.handleSteps!({ + agentState, + prompt: 'Finish the previous response.', + params: {}, + } as any) + + expect(gen.next().value).toMatchObject({ toolName: 'git_status' }) + expect( + gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any).value, + ).toMatchObject({ toolName: 'spawn_agent_inline' }) + const maybePinnedState = gen.next().value + if (maybePinnedState !== 'STEP') { + expect(maybePinnedState).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + } + expect( + gen.next({ stepsComplete: true, toolResult: [], agentState } as any) + .value, + ).toMatchObject({ toolName: 'git_status' }) + + // The receipt is reused: NO run_file_change_hooks yield. The very next + // yield is the post-validation dirty-scope re-check (git_status), which + // then advances to the code-reviewer gate as in a normal passing pass. + const postValidationStatus = gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any) + expect(postValidationStatus.value).not.toMatchObject({ + toolName: 'run_file_change_hooks', + }) + expect(postValidationStatus.value).toMatchObject({ + toolName: 'git_status', + }) + const reviewerSpawn = gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any) + expect(reviewerSpawn.value).toMatchObject({ + toolName: 'spawn_agents', + input: { agents: [{ agent_type: 'code-reviewer' }] }, + }) + + // The newest entry is kept verbatim: same summary, recordedAt, and + // markers (no hook rewrite replaced it). + const kept = (agentState as any).base2ActiveWork + .validationEvidence as Array> + expect(kept).toHaveLength(1) + expect(kept[0].summary).toBe(validationSummary) + expect(kept[0].recordedAt).toBe(recordedAt) + expect(kept[0].assurance).toBe('full') + expect(kept[0].fileMarkers).toEqual({ + [gateFile]: buildContentMarker(tmpFile), + }) + } finally { + rmSync(tmpDir, { recursive: true, force: true }) + } + }) + + test('does not reuse a validationEvidence receipt whose summary starts with REDUCED_ASSURANCE', () => { + const base2 = createBase2('default') + const tmpDir = makeProjectTempDir('base2-receipt-reduced-') + try { + const tmpFile = join(tmpDir, 'a.ts') + const gateFile = normalizeGateFilePath(tmpFile) + writeFileSync(tmpFile, 'export const value = 1\n') + const reducedSummary = + 'REDUCED_ASSURANCE: Validation hooks could not run for this snapshot.' + const agentState = { + agentId: 'base2-custom', + base2ActiveWork: { + changedFiles: [gateFile], + touchedFiles: [gateFile], + pendingGateFiles: [gateFile], + currentPhase: 'awaiting_validation', + latestWorkSummary: 'Pending gate previously validated.', + openReviewerBlockers: [], + lastValidationSummary: reducedSummary, + nextRequiredAction: '', + lastPinnedStateMessage: '', + // Matching markers and full assurance are not enough: a + // REDUCED_ASSURANCE summary must never be reused. + validationEvidence: [ + { + gateId: 'receipt-reduced-gate', + files: [gateFile], + snapshotFingerprint: 'seed-reduced-snapshot', + summary: reducedSummary, + assurance: 'full', + recordedAt: '2025-01-01T00:00:00.000Z', + fileMarkers: { [gateFile]: buildContentMarker(tmpFile) }, + }, + ], + }, + } + const gen = base2.handleSteps!({ + agentState, + prompt: 'Finish the previous response.', + params: {}, + } as any) + + expect(gen.next().value).toMatchObject({ toolName: 'git_status' }) + expect( + gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any).value, + ).toMatchObject({ toolName: 'spawn_agent_inline' }) + const maybePinnedState = gen.next().value + if (maybePinnedState !== 'STEP') { + expect(maybePinnedState).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + } + expect( + gen.next({ stepsComplete: true, toolResult: [], agentState } as any) + .value, + ).toMatchObject({ toolName: 'git_status' }) + const next = gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any) + + // Reduced-assurance summary -> no receipt reuse -> validation hooks rerun. + expect(next.value).toMatchObject({ + toolName: 'run_file_change_hooks', + input: { files: [gateFile] }, + }) + } finally { + rmSync(tmpDir, { recursive: true, force: true }) + } + }) + + test('does not reuse a validationEvidence receipt whose fileMarkers no longer match the live bytes', () => { + const base2 = createBase2('default') + const tmpDir = makeProjectTempDir('base2-receipt-stale-marker-') + try { + const tmpFile = join(tmpDir, 'a.ts') + const gateFile = normalizeGateFilePath(tmpFile) + writeFileSync(tmpFile, 'export const value = 1\n') + const staleMarker = buildContentMarker(tmpFile) + // The bytes drift after the receipt was recorded: the stored marker no + // longer matches, so the reuse must fail closed and re-run the hooks. + writeFileSync(tmpFile, 'export const value = 2\n') + const agentState = { + agentId: 'base2-custom', + base2ActiveWork: { + changedFiles: [gateFile], + touchedFiles: [gateFile], + pendingGateFiles: [gateFile], + currentPhase: 'awaiting_validation', + latestWorkSummary: 'Pending gate previously validated.', + openReviewerBlockers: [], + lastValidationSummary: + 'Configured file-change hooks passed: typecheck.', + nextRequiredAction: '', + lastPinnedStateMessage: '', + validationEvidence: [ + { + gateId: 'receipt-stale-marker-gate', + files: [gateFile], + snapshotFingerprint: 'seed-stale-marker-snapshot', + summary: 'Configured file-change hooks passed: typecheck.', + assurance: 'full', + recordedAt: '2025-01-01T00:00:00.000Z', + fileMarkers: { [gateFile]: staleMarker }, + }, + ], + }, + } + const gen = base2.handleSteps!({ + agentState, + prompt: 'Finish the previous response.', + params: {}, + } as any) + + expect(gen.next().value).toMatchObject({ toolName: 'git_status' }) + expect( + gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any).value, + ).toMatchObject({ toolName: 'spawn_agent_inline' }) + const maybePinnedState = gen.next().value + if (maybePinnedState !== 'STEP') { + expect(maybePinnedState).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + } + expect( + gen.next({ stepsComplete: true, toolResult: [], agentState } as any) + .value, + ).toMatchObject({ toolName: 'git_status' }) + const next = gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any) + + // Marker mismatch -> no receipt reuse -> validation hooks rerun. + expect(next.value).toMatchObject({ + toolName: 'run_file_change_hooks', + input: { files: [gateFile] }, + }) + } finally { + rmSync(tmpDir, { recursive: true, force: true }) + } + }) + + test('does not reuse a legacy validationEvidence receipt that lacks fileMarkers (fail closed)', () => { + const base2 = createBase2('default') + const tmpDir = makeProjectTempDir('base2-receipt-legacy-') + try { + const tmpFile = join(tmpDir, 'a.ts') + const gateFile = normalizeGateFilePath(tmpFile) + writeFileSync(tmpFile, 'export const value = 1\n') + const agentState = { + agentId: 'base2-custom', + base2ActiveWork: { + changedFiles: [gateFile], + touchedFiles: [gateFile], + pendingGateFiles: [gateFile], + currentPhase: 'awaiting_validation', + latestWorkSummary: 'Pending gate previously validated.', + openReviewerBlockers: [], + lastValidationSummary: + 'Configured file-change hooks passed: typecheck.', + nextRequiredAction: '', + lastPinnedStateMessage: '', + // Older serialized entries carry no per-file marker map: without + // byte evidence the receipt must never be reused. + validationEvidence: [ + { + gateId: 'legacy-receipt-gate', + files: [gateFile], + snapshotFingerprint: 'legacy-receipt-snapshot', + summary: 'Configured file-change hooks passed: typecheck.', + assurance: 'full', + recordedAt: '2025-01-01T00:00:00.000Z', + }, + ], + }, + } + const gen = base2.handleSteps!({ + agentState, + prompt: 'Finish the previous response.', + params: {}, + } as any) + + expect(gen.next().value).toMatchObject({ toolName: 'git_status' }) + expect( + gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any).value, + ).toMatchObject({ toolName: 'spawn_agent_inline' }) + const maybePinnedState = gen.next().value + if (maybePinnedState !== 'STEP') { + expect(maybePinnedState).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + } + expect( + gen.next({ stepsComplete: true, toolResult: [], agentState } as any) + .value, + ).toMatchObject({ toolName: 'git_status' }) + const next = gen.next({ + toolResult: [{ type: 'json', value: { status: ` M ${tmpFile}` } }], + } as any) + + // No fileMarkers -> no receipt reuse -> validation hooks rerun. + expect(next.value).toMatchObject({ + toolName: 'run_file_change_hooks', + input: { files: [gateFile] }, + }) + } finally { + rmSync(tmpDir, { recursive: true, force: true }) + } + }) + test('hitStepCap breaks out instead of falling through to the validation/reviewer gate', () => { // Regression: when an explicit fixed cap (stepsRemaining === 0) fires, the LLM // step returns shouldEndTurn=true. Before the hitStepCap flag was threaded @@ -9670,6 +9981,15 @@ describe('base2 specialist attestation tolerance', () => { ) expect(fingerprint).toMatch(/^v3:[a-f0-9]{64}$/) + // The specialist brief must steer large-file reads through bounded + // read_files block selectors so the reviewer's accumulated read context + // stays bounded (mirrors the final-reviewer prompt instruction). + expect( + String((spawn.value as any).input.agents[0].prompt ?? ''), + ).toContain( + 'Read large files via read_files windows/around/symbol selectors (bounded block reads)', + ) + // LOOKS_GOOD attesting ONLY the readable file, plus a stale-snapshot // finding record. Pre-fix the omitted deleted path was a coverage gap and // the stale record escalated to the bundle-refresh retry; now the diff --git a/agents/base2/base2.ts b/agents/base2/base2.ts index 60555b1234..4f29604e88 100644 --- a/agents/base2/base2.ts +++ b/agents/base2/base2.ts @@ -1141,6 +1141,13 @@ ${guideSections} // guarantee they are always plain objects so older serialized state // (which lacks them) is treated as `{}` and fails closed. activeWorkState.specialistReviewGateFingerprints ??= {} + // Phase 4 per-file credit records are deliberately NOT defaulted here. + // Their ABSENCE is the legacy signal: an absent securityReviewFileMarkers + // falls back to the scalar securityReviewGateFingerprint comparison, and + // an absent specialistReviewFileMarkers falls back to the scalar + // specialistReviewGateFingerprints comparison (seeded test fixtures + // depend on this). Both are defaulted lazily with ??= at the write + // sites once a fresh pass mints markers. Never seeded to `{}` here. activeWorkState.specialistRepairRoundCount ??= 0 activeWorkState.specialistNoVerdictCounts ??= {} activeWorkState.reviewReceipts ??= [] @@ -2231,10 +2238,28 @@ ${guideSections} // satisfies the gate; a stored fingerprint re-fires only on real byte // drift, so fail-closed drift detection and owed-security revalidation // (which never marks done on block) are both preserved. + // Phase 4: when securityReviewFileMarkers EXISTS it is authoritative + // over the scalar — the credit is fresh only while EVERY currently + // security-sensitive reviewable file carries a stored marker equal to + // its current content marker; a missing entry, a mismatch, or an + // `unreadable:*` marker is NOT fresh (fail closed) and the refire + // below is scoped to exactly the drifted subset. When the map is + // ABSENT the EXACT legacy scalar comparison applies, so seeded/legacy + // state storing only the scalar keeps its behavior. + const securityStoredFileMarkers = + activeWorkState.securityReviewFileMarkers const securityCreditIsFresh = - activeWorkState.securityReviewGateFingerprint === undefined || - activeWorkState.securityReviewGateFingerprint === - securitySnapshotFingerprint + securityStoredFileMarkers !== undefined + ? securitySensitiveReviewableFiles.every( + (file) => + securityStoredFileMarkers[file] !== undefined && + !securityStoredFileMarkers[file].startsWith('unreadable:') && + securityStoredFileMarkers[file] === + readGateFileContentMarker(file), + ) + : activeWorkState.securityReviewGateFingerprint === undefined || + activeWorkState.securityReviewGateFingerprint === + securitySnapshotFingerprint const owedReviewers = activeWorkState.owedReviewerRevalidations ?? [] const securityWouldRefire = runValidationGate && @@ -2261,6 +2286,12 @@ ${guideSections} activeWorkState.preEditSecurityReviewDone = true activeWorkState.securityReviewGateFingerprint = securitySnapshotFingerprint + // Phase 4: publish an (empty) per-file marker map so the + // authoritative freshness record exists; the current pending set + // is no longer security-sensitive here, and any later + // security-sensitive file without a stored marker fails closed + // and re-fires the gate. + activeWorkState.securityReviewFileMarkers ??= {} if ( activeWorkState.currentPhase === 'blocked' && nextRequired.includes( @@ -2284,6 +2315,10 @@ ${guideSections} activeWorkState.preEditSecurityReviewDone = true activeWorkState.securityReviewGateFingerprint = securitySnapshotFingerprint + // Phase 4: publish an (empty) per-file marker map — the sensitive + // reviewable set is empty here, and any later sensitive file + // without a stored marker fails closed and re-fires the gate. + activeWorkState.securityReviewFileMarkers ??= {} if ( activeWorkState.currentPhase === 'blocked' && (activeWorkState.nextRequiredAction ?? '').includes( @@ -2300,6 +2335,36 @@ ${guideSections} matchesSecuritySensitiveGlob(currentPendingGateFiles) && securityChangedFiles.length > 0 ) { + // Phase 4 scoped re-review: with the marker map present, spawn ONLY + // the drifted subset (sensitive reviewable files whose stored marker + // is missing or stale) and derive the attestation fingerprint and + // deleted-file set from EXACTLY that spawned list so the prompt + // file list and the echo contract stay coherent. With the map + // absent (legacy), reuse the full sensitive set and the unchanged + // full-set fingerprint — byte-identical to the pre-Phase-4 spawn. + const securityDriftedFiles = + securityStoredFileMarkers !== undefined + ? securitySensitiveReviewableFiles.filter( + (file) => + securityStoredFileMarkers[file] === undefined || + securityStoredFileMarkers[file] !== + readGateFileContentMarker(file), + ) + : securityChangedFiles + const securitySpawnFiles = + securityDriftedFiles.length > 0 + ? securityDriftedFiles + : securityChangedFiles + const securitySpawnScoped = securitySpawnFiles !== securityChangedFiles + const securitySpawnDetails = securitySpawnScoped + ? buildGateSnapshotDetails(securitySpawnFiles, '') + : securitySnapshotDetails + const securitySpawnFingerprint = securitySpawnScoped + ? hashGateSnapshotDetails(securitySpawnDetails) + : securitySnapshotFingerprint + const securitySpawnDeletedFiles = securitySpawnScoped + ? collectDeletedFilesFromSnapshotDetails(securitySpawnDetails) + : securityDeletedFiles auxGateFiredThisIteration = true const securityReviewResult = yield { toolName: 'spawn_agent_inline', @@ -2307,13 +2372,13 @@ ${guideSections} agent_type: 'security-reviewer', prompt: [ 'Perform the required snapshot-bound security review.', - `Pending changed files: ${securityChangedFiles.join(', ')}`, - `Snapshot fingerprint: ${securitySnapshotFingerprint}`, + `Pending changed files: ${securitySpawnFiles.join(', ')}`, + `Snapshot fingerprint: ${securitySpawnFingerprint}`, 'Return only the declared structured output.', ].join('\n'), params: { - changed_files: securityChangedFiles, - snapshot_fingerprint: securitySnapshotFingerprint, + changed_files: securitySpawnFiles, + snapshot_fingerprint: securitySpawnFingerprint, }, }, includeToolCall: false, @@ -2337,9 +2402,9 @@ ${guideSections} ) const securityAttestationIssues = collectReviewerAttestationIssues( securityToolResult, - securitySnapshotFingerprint, - securityChangedFiles, - securityDeletedFiles, + securitySpawnFingerprint, + securitySpawnFiles, + securitySpawnDeletedFiles, ) const securityVerdict = getReviewerFinalizationVerdict(securityToolResult) @@ -2354,7 +2419,7 @@ ${guideSections} const record = correlateReviewerFindingRecord(text, records) return { id: record?.id ?? buildReviewerFindingId(text, index), - gateId: `security-reviewer:${securitySnapshotFingerprint}`, + gateId: `security-reviewer:${securitySpawnFingerprint}`, // The PREFIXED blocker string, like the code-reviewer path: // `reviewerVerdictClass` derives the condone key's verdict // class from this text, so storing the record's unprefixed @@ -2363,8 +2428,8 @@ ${guideSections} // BLOCKING re-raise. Only the id is adopted from the record. text, status: 'open' as const, - files: securityChangedFiles, - snapshotFingerprint: securitySnapshotFingerprint, + files: securitySpawnFiles, + snapshotFingerprint: securitySpawnFingerprint, reviewer: 'security-reviewer' as const, createdAt: new Date().toISOString(), } @@ -2386,6 +2451,9 @@ ${guideSections} activeWorkState.securityReviewGateDone = false activeWorkState.preEditSecurityReviewDone = false activeWorkState.securityReviewGateFingerprint = undefined + // Phase 4: the per-file marker map is credit; drop it with the + // scalar so no stale marker can outlive the cleared credit. + delete activeWorkState.securityReviewFileMarkers markActiveWorkStateChanged() emitGateTelemetry({ currentPhase: 'repair_loop', @@ -2554,6 +2622,9 @@ ${guideSections} activeWorkState.securityReviewGateDone = false activeWorkState.preEditSecurityReviewDone = false activeWorkState.securityReviewGateFingerprint = undefined + // Phase 4: the per-file marker map is credit; drop it with the + // scalar so no stale marker can outlive the cleared credit. + delete activeWorkState.securityReviewFileMarkers markActiveWorkStateChanged() emitGateTelemetry({ currentPhase: 'blocked', @@ -2580,7 +2651,7 @@ ${guideSections} recordSuccessfulReviewReceipt( securityToolResult, 'security-reviewer', - securitySnapshotFingerprint, + securitySpawnFingerprint, ) markActiveWorkStateChanged() emitGateTelemetry({ @@ -2617,6 +2688,17 @@ ${guideSections} // the gate re-fires instead of reusing credit for unreviewed bytes. activeWorkState.securityReviewGateFingerprint = securitySnapshotFingerprint + // Phase 4: record per-file markers for exactly the files this pass + // attested, MERGING into any stored map so credit for files reviewed + // earlier survives a scoped re-review of the drifted subset. The + // marker map is authoritative for freshness; the scalar above stays + // written for legacy readers. + { + const markerMap = (activeWorkState.securityReviewFileMarkers ??= {}) + for (const file of securitySpawnFiles) { + markerMap[file] = readGateFileContentMarker(file) + } + } // The security aux block owns security-family revalidation; clear its // owed entry once it passes, but never clobber a code/specialist one. clearOwedReviewer('security-reviewer') @@ -2661,6 +2743,13 @@ ${guideSections} delete activeWorkState.specialistReviewGateFingerprints[ owedSpecialist ] + // Phase 4: the per-file marker map is credit too; drop it with + // the scalar so no stale marker survives the eviction. + if (activeWorkState.specialistReviewFileMarkers) { + delete activeWorkState.specialistReviewFileMarkers[ + owedSpecialist + ] + } markActiveWorkStateChanged() } } @@ -2676,9 +2765,24 @@ ${guideSections} // every sweep. The accepted consequence is that byte drift confined to // those test files alone does not force a specialist re-review; drift // in any aux-relevant source file still does. + // + // Phase 4: per-file credit via specialistReviewFileMarkers refines + // this further. When a marker map exists for a specialist the map is + // authoritative, a byte change re-opens only that file's credit, and + // the specialist is re-reviewed on exactly the drifted subset (see + // specialistCreditIsFresh / specialistDriftedFiles) instead of the + // whole aux-relevant set; state without a map keeps the scalar + // full-set semantics above unchanged. const specialistPendingFiles = selectReviewableGateFiles( currentPendingGateFiles, ) + // Computed once here and threaded into the serialized helpers below + // (specialistCreditIsFresh / specialistDriftedFiles live outside + // this block's closure, where currentPendingGateFiles does not + // exist). + const specialistAuxRelevantFiles = selectReviewableGateFiles( + selectAuxRelevantFiles(currentPendingGateFiles), + ) const specialistCreditFingerprint = hashGateSnapshotDetails( buildGateSnapshotDetails( selectReviewableGateFiles( @@ -2709,8 +2813,43 @@ ${guideSections} : baseRoutedSpecialists ).filter( (agentType) => - !specialistCreditIsFresh(agentType, specialistCreditFingerprint), + !specialistCreditIsFresh( + agentType, + specialistCreditFingerprint, + specialistAuxRelevantFiles, + ), ) + // Phase 4 scoped re-review: each routed specialist is spawned with + // only the subset of aux-relevant reviewable files lacking a fresh + // stored marker (the FULL aux-relevant set when its marker map is + // absent — legacy), and the expected attestation fingerprint is + // computed over EXACTLY that spawned list so the prompt file list + // and the echo-attestation contract stay coherent. + const specialistScopedFileSets = new Map() + const specialistScopedFingerprints = new Map() + const specialistScopedDeletedFiles = new Map() + for (const agentType of routedSpecialists) { + const driftedFiles = specialistDriftedFiles( + agentType, + specialistAuxRelevantFiles, + ) + const scopedFiles = + driftedFiles.length > 0 + ? driftedFiles + : selectReviewableGateFiles( + selectAuxRelevantFiles(currentPendingGateFiles), + ) + specialistScopedFileSets.set(agentType, scopedFiles) + const scopedDetails = buildGateSnapshotDetails(scopedFiles, '') + specialistScopedFingerprints.set( + agentType, + hashGateSnapshotDetails(scopedDetails), + ) + specialistScopedDeletedFiles.set( + agentType, + collectDeletedFilesFromSnapshotDetails(scopedDetails), + ) + } if ( routedSpecialists.length > 0 && specialistPendingFiles.length === 0 @@ -2746,7 +2885,18 @@ ${guideSections} // Gate-owned v3 fingerprint is the sole specialist attestation // token (same family as security/code-reviewer). Fail closed when // crypto is unavailable rather than spawning with a bare bundle id. - if (!isAttestableSnapshotFingerprint(specialistCreditFingerprint)) { + // Each scoped spawn fingerprint must attest independently. + const scopedFingerprintsAreAttestable = routedSpecialists.every( + (agentType) => + isAttestableSnapshotFingerprint( + specialistScopedFingerprints.get(agentType) ?? + specialistCreditFingerprint, + ), + ) + if ( + !isAttestableSnapshotFingerprint(specialistCreditFingerprint) || + !scopedFingerprintsAreAttestable + ) { activeWorkState.currentPhase = 'blocked' activeWorkState.openReviewerBlockers = [ 'Specialist review cannot attest: gate snapshot fingerprint is non-attestable (crypto unavailable).', @@ -2797,13 +2947,23 @@ ${guideSections} prompt: buildSpecialistScopedReviewPrompt({ title: 'Perform the routed post-edit specialist review.', agentType, - files: specialistPendingFiles, - snapshotFingerprint: specialistCreditFingerprint, + files: specialistScopedFileSets.get(agentType) ?? [], + snapshotFingerprint: + specialistScopedFingerprints.get(agentType) ?? + specialistCreditFingerprint, userPrompt: prompt ?? '', + extraLines: buildReviewerRoundLedgerLines(agentType, { + scopeFiles: specialistScopedFileSets.get(agentType) ?? [], + currentFingerprint: + specialistScopedFingerprints.get(agentType) ?? + specialistCreditFingerprint, + }), }), params: { - files: specialistPendingFiles, - snapshot_id: specialistCreditFingerprint, + files: specialistScopedFileSets.get(agentType) ?? [], + snapshot_id: + specialistScopedFingerprints.get(agentType) ?? + specialistCreditFingerprint, }, })), }, @@ -2829,9 +2989,10 @@ ${guideSections} // snapshot drift never triggers a pointless refresh+retry. const attestationIssues = collectReviewerAttestationIssues( result, - specialistCreditFingerprint, - specialistPendingFiles, - specialistDeletedFiles, + specialistScopedFingerprints.get(agentType) ?? + specialistCreditFingerprint, + specialistScopedFileSets.get(agentType) ?? [], + specialistScopedDeletedFiles.get(agentType) ?? [], ) return ( attestationIssues.length > 0 && @@ -2886,16 +3047,20 @@ ${guideSections} title: 'Retry the routed specialist review after snapshot/file attestation failure.', agentType, - files: specialistPendingFiles, - snapshotFingerprint: retryCreditFingerprint, + files: specialistScopedFileSets.get(agentType) ?? [], + snapshotFingerprint: + specialistScopedFingerprints.get(agentType) ?? + specialistCreditFingerprint, userPrompt: prompt ?? '', extraLines: [ 'Correct the structured output directly; do not request source edits for this protocol error.', ], }), params: { - files: specialistPendingFiles, - snapshot_id: retryCreditFingerprint, + files: specialistScopedFileSets.get(agentType) ?? [], + snapshot_id: + specialistScopedFingerprints.get(agentType) ?? + specialistCreditFingerprint, }, })), }, @@ -2916,14 +3081,15 @@ ${guideSections} for (const agentType of routedSpecialists) { const expectedSnapshotId = specialistSnapshots.get(agentType) ?? + specialistScopedFingerprints.get(agentType) ?? specialistCreditFingerprint const specialistToolResult = specialistResults.get(agentType) const specialistAttestationIssues = collectReviewerAttestationIssues( specialistToolResult, expectedSnapshotId, - specialistPendingFiles, - specialistDeletedFiles, + specialistScopedFileSets.get(agentType) ?? [], + specialistScopedDeletedFiles.get(agentType) ?? [], ) // Fingerprint-only drift on a fully-attesting review is NOT a // terminal protocol failure: only a FILE-COVERAGE gap or a @@ -3025,16 +3191,22 @@ ${guideSections} ;(activeWorkState.specialistReviewGateFingerprints ??= {})[ agentType ] = specialistCreditFingerprint - if ( - activeWorkState.lastReviewerGateSkipReason === - 'specialist-terminal-failure' || - activeWorkState.lastReviewerGateSkipReason === - 'specialist-rate-limited' - ) { - activeWorkState.lastReviewerGateSkipReason = '' + // Phase 4: merge markers for exactly the files this pass + // attested, PRESERVING previously stored markers for other + // files so their per-file credit survives the re-review. + { + const markerMap = ( + (activeWorkState.specialistReviewFileMarkers ??= {})[ + agentType + ] ??= {} + ) + for (const scopedFile of specialistScopedFileSets.get( + agentType, + ) ?? []) { + markerMap[scopedFile] = + readGateFileContentMarker(scopedFile) + } } - clearOwedReviewer(agentType) - markActiveWorkStateChanged() const parentOwnedPassAdvisories = boundAdvisoryLines( collectReviewerAdvisories(specialistToolResult), ) @@ -3104,6 +3276,13 @@ ${guideSections} agentType ] } + // Phase 4: the per-file marker entry is credit too; drop + // it with the scalar so no stale marker survives eviction. + if (activeWorkState.specialistReviewFileMarkers) { + delete activeWorkState.specialistReviewFileMarkers[ + agentType + ] + } const specialistRepairRound: number = Number(activeWorkState.specialistRepairRoundCount ?? 0) + 1 @@ -3546,6 +3725,22 @@ ${guideSections} ;(activeWorkState.specialistReviewGateFingerprints ??= {})[ agentType ] = specialistCreditFingerprint + // Phase 4: merge markers for exactly the files this pass + // attested, PRESERVING previously stored markers for other + // files so their per-file credit survives the re-review. + { + const markerMap = ( + (activeWorkState.specialistReviewFileMarkers ??= {})[ + agentType + ] ??= {} + ) + for (const scopedFile of specialistScopedFileSets.get( + agentType, + ) ?? []) { + markerMap[scopedFile] = + readGateFileContentMarker(scopedFile) + } + } if ( activeWorkState.lastReviewerGateSkipReason === 'specialist-terminal-failure' || @@ -3896,39 +4091,108 @@ ${guideSections} runValidationGate && !validationInfrastructureBypassed ) { - setGateProgress( - `gate: validation hooks running for ${gateScopeFiles.length} file(s)`, - ) - const verify = yield { - toolName: 'run_file_change_hooks', - input: { files: gateScopeFiles }, - } as any - let failures = collectHookFailures( - (verify as any) && (verify as any).toolResult, - ) - if (failures.length === 0) { - validationSummary = summarizeHookResults( - (verify as any) && (verify as any).toolResult, + // Phase 4 validation receipt reuse: the NEWEST evidence entry is + // reused IN PLACE — no hook run, no rewrite — only when it carries + // full assurance, covers EXACTLY the current gate-scope file set, + // and every captured content marker still matches the live bytes. + // Anything else (a legacy entry without markers, a missing or + // mismatched or `unreadable:*` marker, reduced/none assurance, or a + // REDUCED_ASSURANCE summary) falls through to the fresh hook run + // that rewrites the evidence, so the gate-lifecycle cadence only + // changes on an exact full-coverage, byte-identical reuse. + // Explicit element type via the named state type: without it TS + // infers this const through the later + // `activeWorkState.validationEvidence = ...` assignment (which + // references reusedEvidence) and reports TS7022 circularity. A + // `typeof activeWorkState...` annotation is itself circular + // (TS2502), so the named imported type is used instead. + const newestEvidenceEntry: + | NonNullable< + Base2ActiveWorkState['validationEvidence'] + >[number] + | undefined = + activeWorkState.validationEvidence?.[ + (activeWorkState.validationEvidence?.length ?? 0) - 1 + ] + const evidenceFileMarkers: Record | undefined = + newestEvidenceEntry?.fileMarkers + // Explicit annotations: these consts feed the later + // `activeWorkState.validationEvidence = reusedEvidence ? ...` + // assignment, and TS's inference walk through the state field + // reports TS7022 without them. + const evidenceCoversExactScope: boolean = + !!newestEvidenceEntry && + gateFileSetsEqual(newestEvidenceEntry.files ?? [], gateScopeFiles) + const reusedEvidence: + | NonNullable< + Base2ActiveWorkState['validationEvidence'] + >[number] + | null = + evidenceCoversExactScope && + newestEvidenceEntry.assurance === 'full' && + !newestEvidenceEntry.summary.startsWith('REDUCED_ASSURANCE:') && + evidenceFileMarkers !== undefined && + gateScopeFiles.every( + (file) => + evidenceFileMarkers[file] !== undefined && + !evidenceFileMarkers[file].startsWith('unreadable:') && + evidenceFileMarkers[file] === readGateFileContentMarker(file), ) - activeWorkState.lastValidationSummary = validationSummary - activeWorkState.validationAssurance = validationSummary.startsWith( - 'REDUCED_ASSURANCE:', + ? newestEvidenceEntry + : null + if (!reusedEvidence) { + setGateProgress( + `gate: validation hooks running for ${gateScopeFiles.length} file(s)`, ) - ? 'reduced' - : 'full' - activeWorkState.validationEvidence = [ - { - gateId: reviewSnapshotFingerprint, - files: gateScopeFiles, - snapshotFingerprint: buildGateFingerprint( - gateScopeFiles, - validationSummary, - ), - summary: validationSummary, - assurance: activeWorkState.validationAssurance, - recordedAt: new Date().toISOString(), - }, - ] + } + const verify = reusedEvidence + ? null + : (yield { + toolName: 'run_file_change_hooks', + input: { files: gateScopeFiles }, + } as any) + let failures = reusedEvidence + ? [] + : collectHookFailures( + (verify as any) && (verify as any).toolResult, + ) + if (failures.length === 0) { + if (!reusedEvidence) { + validationSummary = summarizeHookResults( + (verify as any) && (verify as any).toolResult, + ) + activeWorkState.lastValidationSummary = validationSummary + activeWorkState.validationAssurance = + validationSummary.startsWith('REDUCED_ASSURANCE:') + ? 'reduced' + : 'full' + } + activeWorkState.validationEvidence = reusedEvidence + ? // Reuse keeps the entry verbatim (summary, assurance, + // recordedAt, markers); only the mirrors above advance. + activeWorkState.validationEvidence + : [ + { + gateId: reviewSnapshotFingerprint, + files: gateScopeFiles, + snapshotFingerprint: buildGateFingerprint( + gateScopeFiles, + validationSummary, + ), + summary: validationSummary, + assurance: activeWorkState.validationAssurance, + recordedAt: new Date().toISOString(), + // Phase 4: per-file content markers captured at this + // pass so a later cycle with unchanged covered bytes can + // reuse this receipt instead of re-running hooks. + fileMarkers: Object.fromEntries( + gateScopeFiles.map((file) => [ + file, + readGateFileContentMarker(file), + ]), + ), + }, + ] activeWorkState.currentPhase = 'awaiting_review' markActiveWorkStateChanged() } else { @@ -4465,9 +4729,14 @@ ${guideSections} 'Snapshot details (read for file membership; do not echo):', reviewSnapshotDetails, `Validation gate summary: ${validationSummary}`, - // Re-review ledger; empty on round 0 so no stray heading or - // blank line appears in the first review's prompt. - ...buildReviewerRoundLedgerLines(requiredReviewerAgentType), + // Re-review ledger (repair round + already-attested + // block); empty on round 0 with no prior receipts so no + // stray heading or blank line appears in the first + // review's prompt. + ...buildReviewerRoundLedgerLines(requiredReviewerAgentType, { + scopeFiles: reviewableGateScopeFiles, + currentFingerprint: reviewSnapshotFingerprint, + }), 'Read large files via read_files windows (bounded block reads) instead of whole-file reads so your accumulated read context stays bounded; still attest to every pending file in reviewedFiles.', '', 'Return the required structured review object. Echo snapshotFingerprint exactly, list every pending changed file in reviewedFiles (including tests), evaluate all review dimensions, and map every user requirement to evidence. Changed tests are first-class review targets and may also be cited as coverage evidence. Use coverage: missing only when no covering test exists in the changed files or elsewhere in the repo.', @@ -4513,6 +4782,10 @@ ${guideSections} 'Snapshot details (read for file membership; do not echo):', reviewSnapshotDetails, `Validation gate summary: ${validationSummary}`, + ...buildReviewerRoundLedgerLines(requiredReviewerAgentType, { + scopeFiles: reviewableGateScopeFiles, + currentFingerprint: reviewSnapshotFingerprint, + }), '', 'Protocol errors from the prior response:', ...attestationIssues, @@ -5989,10 +6262,10 @@ ${guideSections} // T1.2(c) re-review ledger for the reviewer spawn packet. The reviewer is // stateless across repair rounds, so without this it re-derives every // finding from scratch instead of verifying the ones a repair round - // already reported as addressed. Returns [] on round 0 so the first - // review's prompt is byte-identical to the pre-ledger surface, and reads - // ONLY already-persisted state (openReviewerFindings / - // reviewerRepairRoundCount) — no new gate state. + // already reported as addressed. The repair-round header/findings block + // is emitted only when reviewerRepairRoundCount > 0, and the function + // reads ONLY already-persisted state (openReviewerFindings / + // reviewReceipts / reviewerRepairRoundCount) — no new gate state. // // Findings are filtered to the spawned reviewer's own family: // openReviewerFindings can hold security-reviewer and specialist records, @@ -6000,29 +6273,108 @@ ${guideSections} // scope. `finding.text` is rendered VERBATIM (keeping the // NON_BLOCKING:/BLOCKING: prefix and any `[id] ` segment) because that is // exactly the string the condone matcher compares a re-raise against. - // Inline because handleSteps is serialized via .toString() + - // new Function(...), so it must not reference module-scope imports. - function buildReviewerRoundLedgerLines(reviewer: string): string[] { + // + // Phase 4 already-attested ledger: options.scopeFiles bounds the block + // to the spawn's review scope (the reviewable gate-scope files for the + // final reviewer, the scoped spawn subset for specialists). LOOKS_GOOD + // receipts from the SAME reviewer family whose reviewedFiles intersect + // that scope are summarized in a handful of bounded lines: files whose + // receipt fingerprint still equals options.currentFingerprint are + // attested with unchanged bytes; the rest were attested in a prior round + // and their bytes may have changed. The reviewer is told to focus review + // depth on unattested/drifted files while still reporting any issue + // found anywhere. Scope and fingerprint arrive as explicit options (not + // closure reads) so the function stays callable from the specialist + // block before the reviewer-block `let` bindings initialize. Inline + // because handleSteps is serialized via .toString() + new Function(...), + // so it must not reference module-scope imports. + function buildReviewerRoundLedgerLines( + reviewer: string, + options?: { scopeFiles?: string[]; currentFingerprint?: string }, + ): string[] { + const lines: string[] = [] const repairRound = Number( activeWorkState.reviewerRepairRoundCount ?? 0, ) - if (!(repairRound > 0)) return [] - const lines = [`Repair round: ${repairRound}. This is a re-review.`] - const ownFindings = (activeWorkState.openReviewerFindings ?? []).filter( - (finding) => finding.reviewer === reviewer, - ) - if (ownFindings.length === 0) return lines - lines.push( - 'Findings raised earlier and reported addressed are listed below. Verify each is genuinely fixed and cite the line that fixes it. If a fix is wrong or incomplete, re-raise the finding with its ORIGINAL text repeated VERBATIM and put your reason on a separate line: the gate matches re-raises by exact text (and by stable finding id when you supplied one), so a reworded re-raise is treated as a brand-new finding and the repair loop cannot converge.', - ) - const shown = ownFindings.slice(0, 12) - for (const finding of shown) { - lines.push(` - ${finding.text}`) + if (repairRound > 0) { + lines.push(`Repair round: ${repairRound}. This is a re-review.`) + const ownFindings = ( + activeWorkState.openReviewerFindings ?? [] + ).filter((finding) => finding.reviewer === reviewer) + if (ownFindings.length > 0) { + lines.push( + 'Findings raised earlier and reported addressed are listed below. Verify each is genuinely fixed and cite the line that fixes it. If a fix is wrong or incomplete, re-raise the finding with its ORIGINAL text repeated VERBATIM and put your reason on a separate line: the gate matches re-raises by exact text (and by stable finding id when you supplied one), so a reworded re-raise is treated as a brand-new finding and the repair loop cannot converge.', + ) + const shown = ownFindings.slice(0, 12) + for (const finding of shown) { + lines.push(` - ${finding.text}`) + } + if (ownFindings.length > shown.length) { + lines.push( + ` - (+${ownFindings.length - shown.length} more earlier findings omitted)`, + ) + } + } } - if (ownFindings.length > shown.length) { - lines.push( - ` - (+${ownFindings.length - shown.length} more earlier findings omitted)`, + // Phase 4 already-attested block. Bounded: at most 12 files per list + // plus one count line, so the prompt cannot grow without bound across + // a long plan run. + const scopeSet = new Set( + selectReviewableGateFiles(options?.scopeFiles ?? []), + ) + if (scopeSet.size > 0) { + const familyReceipts = (activeWorkState.reviewReceipts ?? []).filter( + (receipt) => + receipt && + receipt.reviewer === reviewer && + receipt.verdict === 'LOOKS_GOOD' && + Array.isArray(receipt.reviewedFiles) && + receipt.reviewedFiles.some((file) => scopeSet.has(file)), ) + if (familyReceipts.length > 0) { + const unchangedFiles = new Set() + const priorFiles = new Set() + for (const receipt of familyReceipts) { + const matchesCurrentBytes = + typeof options?.currentFingerprint === 'string' && + options.currentFingerprint === receipt.snapshotFingerprint + for (const file of receipt.reviewedFiles) { + if (!scopeSet.has(file)) continue + if (matchesCurrentBytes) unchangedFiles.add(file) + else priorFiles.add(file) + } + } + const unchangedList = Array.from(unchangedFiles).slice(0, 12) + const priorList = Array.from(priorFiles) + .filter((file) => !unchangedFiles.has(file)) + .slice(0, 12) + if (unchangedList.length > 0 || priorList.length > 0) { + lines.push( + 'Already attested LOOKS_GOOD by an earlier round of this reviewer family:', + ) + if (unchangedList.length > 0) { + lines.push( + ` - attested with unchanged bytes (fingerprint still matches): ${unchangedList.join(', ')}`, + ) + } + if (priorList.length > 0) { + lines.push( + ` - attested in a prior round; bytes may have changed since: ${priorList.join(', ')}`, + ) + } + if ( + unchangedFiles.size > unchangedList.length || + priorFiles.size > priorList.length + ) { + lines.push( + ` - (+${unchangedFiles.size - unchangedList.length + priorFiles.size - priorList.length} more already-attested files omitted)`, + ) + } + lines.push( + 'These files were already attested LOOKS_GOOD; focus your review depth on files NOT listed above (unattested or drifted). Still report any issue you find anywhere, including in the listed files.', + ) + } + } } return lines } @@ -6370,12 +6722,30 @@ ${guideSections} function specialistCreditIsFresh( agentType: string, fingerprint: string, + auxRelevantReviewableFiles: string[], ): boolean { if ( !(activeWorkState.specialistReviewGatesDone ?? []).includes(agentType) ) { return false } + // Phase 4: when a per-file marker map exists for this specialist it + // is AUTHORITATIVE over the scalar fingerprint — the credit is fresh + // only while EVERY currently aux-relevant reviewable file carries a + // stored marker equal to its current content marker. A missing entry, + // a mismatch, or an `unreadable:*` marker is NOT fresh (fail closed); + // the scoped-drift helper below then narrows the re-review to exactly + // those files instead of the whole set. + const storedFileMarkers = + activeWorkState.specialistReviewFileMarkers?.[agentType] + if (storedFileMarkers !== undefined) { + return auxRelevantReviewableFiles.every( + (file) => + storedFileMarkers[file] !== undefined && + !storedFileMarkers[file].startsWith('unreadable:') && + storedFileMarkers[file] === readGateFileContentMarker(file), + ) + } const stored = (activeWorkState.specialistReviewGateFingerprints ?? {})[ agentType ] @@ -6386,6 +6756,26 @@ ${guideSections} return stored === fingerprint } + // Phase 4 scoped re-review: the subset of currently aux-relevant + // reviewable files whose stored marker for this specialist is missing + // or drifted (fail closed — an `unreadable:*` stored marker is drifted). + // Empty when the specialist's per-file map is ABSENT (legacy scalar + // state): the full aux-relevant set is then re-reviewed, byte-identical + // to the pre-Phase-4 spawn. Inline for the same serialization reason. + function specialistDriftedFiles( + agentType: string, + auxRelevantReviewableFiles: string[], + ): string[] { + const storedFileMarkers = + activeWorkState.specialistReviewFileMarkers?.[agentType] + if (storedFileMarkers === undefined) return auxRelevantReviewableFiles + return auxRelevantReviewableFiles.filter( + (file) => + storedFileMarkers[file] === undefined || + storedFileMarkers[file] !== readGateFileContentMarker(file), + ) + } + // T1.5: condoning is keyed on (verdict class, finding identity) rather // than finding text alone. Text-only keying strips the // NON_BLOCKING/BLOCKING prefix before comparing, so a nit condoned as @@ -7907,9 +8297,47 @@ function hashGateSnapshotDetails(details: string): string { gatePassedFiles.delete(file) // A re-edited file leaves the gate-passed ledger; drop its marker so // the eviction guard cannot see a stale marker and no orphan remains. + // The validation receipt markers covering the file are dropped so + // hook reuse fails closed into a fresh hook run for changed bytes. if (activeWorkState.gatePassedFileMarkers) { delete activeWorkState.gatePassedFileMarkers[file] } + // Phase 4: security per-file credit is CONTENT-keyed, so a mere + // re-absorption of the same still-dirty path (the scoped git-status + // sweep re-reports it every iteration) must NOT evict credit the + // aux block just wrote — an unconditional delete here re-fired + // security reviews in a loop on unchanged bytes. Evict only on + // confirmed byte drift; a genuinely re-edited file loses its marker + // immediately, and every freshness check also fails closed on a + // mismatch, so no stale credit can be reused. An `unreadable:*` + // stored marker is left in place: freshness already fails closed on + // it, and eviction would only mask the unreadable condition. + if (activeWorkState.securityReviewFileMarkers) { + const securityStoredMarker = + activeWorkState.securityReviewFileMarkers[file] + if ( + securityStoredMarker !== undefined && + !securityStoredMarker.startsWith('unreadable:') && + readGateFileContentMarker(file) !== securityStoredMarker + ) { + delete activeWorkState.securityReviewFileMarkers[file] + } + } + if ( + Array.isArray(activeWorkState.validationEvidence) && + activeWorkState.validationEvidence.length > 0 && + activeWorkState.validationEvidence[ + activeWorkState.validationEvidence.length - 1 + ]?.fileMarkers + ) { + const evidenceEntry = + activeWorkState.validationEvidence[ + activeWorkState.validationEvidence.length - 1 + ] + if (evidenceEntry.fileMarkers && file in evidenceEntry.fileMarkers) { + delete evidenceEntry.fileMarkers[file] + } + } activeWorkState.gatePassedFiles = activeWorkState.gatePassedFiles.filter( (passedFile) => passedFile !== file, @@ -8620,6 +9048,7 @@ function hashGateSnapshotDetails(details: string): string { '- Do NOT treat parent workflow as review requirements: rewriting git commits, running full validation, commit/push, confirming CI/CD green, or other operator/orchestrator duties. Omit those from requirementCoverage (or if mentioned only as context, never mark them missing/uncertain for the gate).', `Changed files: ${input.files.join(', ') || '(none)'}`, `Snapshot fingerprint (echo exactly): ${input.snapshotFingerprint}`, + 'Read large files via read_files windows/around/symbol selectors (bounded block reads) instead of whole-file reads so your accumulated read context stays bounded.', ] if (truncatedIntent) { lines.push( diff --git a/agents/base2/gate-state.ts b/agents/base2/gate-state.ts index c4957f8085..053427c6f1 100644 --- a/agents/base2/gate-state.ts +++ b/agents/base2/gate-state.ts @@ -323,6 +323,19 @@ export type Base2ActiveWorkState = Base2GateState & { summary: string assurance: 'full' | 'reduced' | 'none' recordedAt: string + /** + * Phase 4: per-file content markers (readGateFileContentMarker) captured + * at the hook pass, keyed by covered file path. Enables receipt reuse: + * when the NEWEST entry has full assurance, covers exactly the current + * gate-scope file set, and every marker still matches the live bytes, the + * hook run is skipped and this entry is kept verbatim (summary and + * assurance included). Legacy entries without this field, any marker + * mismatch or absence, an `unreadable:*` marker, and any + * REDUCED_ASSURANCE / reduced / none entry are NEVER reusable (fail + * closed). MUST stay a plain JSON-serializable record (never a Map/Set). + * Backward-compatible: older serialized entries lack this field. + */ + fileMarkers?: Record }> lastValidationSummary: string nextRequiredAction: string @@ -441,6 +454,46 @@ export type Base2ActiveWorkState = Base2GateState & { * gate rather than reusing unearned credit. */ securityReviewGateFingerprint?: string + /** + * Phase 4 per-file specialist credit: maps each credited specialist agent + * type to a record of normalized reviewable file path → content marker + * (readGateFileContentMarker) captured when that file's specialist credit + * was stored. When a map exists for a specialist it is AUTHORITATIVE over + * specialistReviewGateFingerprints: the specialist stays fresh only while + * every currently aux-relevant reviewable file carries a stored marker + * equal to its current content marker. A missing entry, a mismatch, or an + * unreadable file (`unreadable:*`) reopens that specialist's gate for a + * SCOPED re-review of exactly the drifted files instead of the whole set, + * and previously stored markers for the other files survive the pass. When + * the map is ABSENT for a specialist (legacy serialized state, including + * fixtures that seed only the scalar fingerprint), the gate falls back to + * the scalar specialistReviewGateFingerprints comparison and a not-fresh + * specialist re-reviews the full aux-relevant set (today's behavior). A + * state with neither record re-reviews (fail closed). Bounded: the + * per-specialist entry (and the scalar entry) are deleted wherever + * specialist credit is evicted. MUST stay a plain JSON-serializable + * record of plain records (never a Map/Set). Backward-compatible: older + * serialized state lacks this field (treated as no per-file credit). + */ + specialistReviewFileMarkers?: Record> + /** + * Phase 4 per-file security credit: maps each security-sensitive reviewable + * file to the content marker (readGateFileContentMarker) captured when its + * security credit was stored. When the map is PRESENT it is authoritative + * over securityReviewGateFingerprint: the security gate stays fresh only + * while every currently security-sensitive reviewable file carries a + * stored marker equal to its current content marker, so a scoped re-review + * of exactly the drifted files is spawned instead of the whole set (the + * stored scalar is still updated for legacy readers). When the map is + * ABSENT, the EXACT legacy scalar semantics apply + * (`undefined || === securitySnapshotFingerprint`), so seeded/legacy state + * that stores only the scalar keeps round-tripping unchanged. Fail closed + * on missing entries, mismatches, and `unreadable:*` markers. Deleted + * wherever the scalar security credit is cleared. MUST stay a plain + * JSON-serializable record (never a Map/Set). Backward-compatible: older + * serialized state lacks this field. + */ + securityReviewFileMarkers?: Record /** * Repair rounds for the specialist -> repair -> re-review loop. Telemetry * only by default (unlimited / progress-gated via no-progress and incomplete diff --git a/agents/e2e/gate-aux-ordering.e2e.test.ts b/agents/e2e/gate-aux-ordering.e2e.test.ts index 3d266fc071..a4d48e36e8 100644 --- a/agents/e2e/gate-aux-ordering.e2e.test.ts +++ b/agents/e2e/gate-aux-ordering.e2e.test.ts @@ -1121,6 +1121,407 @@ describe('base2 pre-reviewer aux gate ordering e2e', () => { }) }) + test('security per-file marker credit does not re-spawn security-reviewer when attested bytes are unchanged', () => { + // Phase 4 regression: after a passing security pass stores + // securityReviewFileMarkers, that map is the AUTHORITATIVE freshness + // record. A second iteration over the same bytes must NOT re-spawn the + // reviewer (the pre-Phase-4 eviction loop re-fired security on its own + // fresh credit); the next yield goes straight to the validation hooks. + const base2 = createBase2('default') + // Seed test/doc done so only security runs on the first aux pass; keep + // auxGatesLastPendingFiles aligned so resetAuxGateFlags cannot re-arm the + // writers mid-test (same seed shape as the STATUS-path freshness test). + const agentState = { + agentId: 'base2-custom', + base2ActiveWork: { + changedFiles: [AUX_TRIPLE_FILE], + touchedFiles: [AUX_TRIPLE_FILE], + pendingGateFiles: [AUX_TRIPLE_FILE], + currentPhase: 'awaiting_validation', + openReviewerBlockers: [], + openReviewerFindings: [], + lastValidationSummary: '', + nextRequiredAction: '', + lastPinnedStateMessage: '', + gatePassedFiles: [], + gatePassedPendingFiles: [], + gatePassedReviewerVerdict: '', + gatePassedValidationSummary: '', + gatePassedFingerprint: '', + lastReviewerGateSkipReason: '', + reviewReceipts: [], + testWriterGateDone: true, + docWriterGateDone: true, + securityReviewGateDone: false, + preEditSecurityReviewDone: false, + specialistReviewGatesDone: [], + auxGatesLastPendingFiles: [AUX_TRIPLE_FILE], + }, + } + const gen = base2.handleSteps!({ + agentState, + prompt: 'Please finish the pending auth session gate item.', + params: {}, + } as any) + + // Resumed-state prelude. + expect(gen.next().value).toMatchObject({ + toolName: 'git_status', + input: {}, + }) + expect( + gen.next(feedJson({ status: ` M ${AUX_TRIPLE_FILE}` })).value, + ).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'context-pruner' }, + }) + expect(gen.next().value).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + expect(gen.next(finishStepWithToolResult({})).value).toMatchObject({ + toolName: 'git_status', + input: {}, + }) + + // First iteration: security-reviewer spawns for the sensitive auth file. + const securityReviewerYield = gen.next( + feedJson({ status: ` M ${AUX_TRIPLE_FILE}` }), + ) + expect(securityReviewerYield.value).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { + agent_type: 'security-reviewer', + params: { changed_files: [AUX_TRIPLE_FILE] }, + }, + includeToolCall: false, + }) + const securityFingerprint = (securityReviewerYield.value as any).input + .params.snapshot_fingerprint as string + expect(securityFingerprint).toMatch(/^v3:[a-f0-9]{64}$/) + + // The LOOKS_GOOD pass publishes the per-file marker map. + expect( + gen.next(reviewerResult(securityFingerprint, [AUX_TRIPLE_FILE])).value, + ).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'context-pruner' }, + }) + const storedMarkers = (agentState as any).base2ActiveWork + .securityReviewFileMarkers as Record + expect(storedMarkers).toBeDefined() + expect(Object.keys(storedMarkers)).toContain(AUX_TRIPLE_FILE) + + // Second iteration with UNCHANGED bytes: the per-file map keeps the credit + // fresh, so security-reviewer must not re-spawn. The next yield is the + // final validation gate (run_file_change_hooks). + expect(gen.next().value).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + expect(gen.next(finishStepWithToolResult({})).value).toMatchObject({ + toolName: 'git_status', + }) + const secondIterationNext = gen.next( + feedJson({ status: ` M ${AUX_TRIPLE_FILE}` }), + ) + expect((secondIterationNext.value as any)?.input?.agent_type).not.toBe( + 'security-reviewer', + ) + expect(secondIterationNext.value).toMatchObject({ + toolName: 'run_file_change_hooks', + input: { files: [AUX_TRIPLE_FILE] }, + }) + expect((agentState as any).base2ActiveWork).toMatchObject({ + securityReviewGateDone: true, + preEditSecurityReviewDone: true, + }) + }) + + test('security per-file marker credit scopes a re-review to the newly drifted security-sensitive file', () => { + // A real scratch file backs the second sensitive path so its bytes (and + // content marker) can drift between passes; the `auth` path segment is in + // SECURITY_SENSITIVE_GLOBS and no specialist router stem matches policy.ts. + mkdirSync(`${SPECIALIST_SCRATCH_ROOT}/auth`, { recursive: true }) + const driftedFile = `${SPECIALIST_SCRATCH_ROOT}/auth/policy.ts` + writeFileSync(driftedFile, 'export const policy = "v1"\n') + const base2 = createBase2('default') + const agentState = { + agentId: 'base2-custom', + base2ActiveWork: { + changedFiles: [AUX_TRIPLE_FILE, driftedFile], + touchedFiles: [AUX_TRIPLE_FILE, driftedFile], + pendingGateFiles: [AUX_TRIPLE_FILE, driftedFile], + currentPhase: 'awaiting_validation', + openReviewerBlockers: [], + openReviewerFindings: [], + lastValidationSummary: '', + nextRequiredAction: '', + lastPinnedStateMessage: '', + gatePassedFiles: [], + gatePassedPendingFiles: [], + gatePassedReviewerVerdict: '', + gatePassedValidationSummary: '', + gatePassedFingerprint: '', + lastReviewerGateSkipReason: '', + reviewReceipts: [], + testWriterGateDone: true, + docWriterGateDone: true, + securityReviewGateDone: false, + preEditSecurityReviewDone: false, + specialistReviewGatesDone: [], + auxGatesLastPendingFiles: [AUX_TRIPLE_FILE, driftedFile], + }, + } + const gen = base2.handleSteps!({ + agentState, + prompt: 'Please finish the pending auth session gate item.', + params: {}, + } as any) + + // Resumed-state prelude. + expect(gen.next().value).toMatchObject({ + toolName: 'git_status', + input: {}, + }) + expect( + gen.next( + feedJson({ status: ` M ${AUX_TRIPLE_FILE}\n M ${driftedFile}` }), + ).value, + ).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'context-pruner' }, + }) + expect(gen.next().value).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + expect(gen.next(finishStepWithToolResult({})).value).toMatchObject({ + toolName: 'git_status', + input: {}, + }) + + // First pass: no marker map exists yet, so the spawn covers BOTH + // security-sensitive files (legacy full-set scope). + const securityReviewerYield = gen.next( + feedJson({ status: ` M ${AUX_TRIPLE_FILE}\n M ${driftedFile}` }), + ) + expect(securityReviewerYield.value).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'security-reviewer' }, + includeToolCall: false, + }) + const firstParams = (securityReviewerYield.value as any).input.params + expect((firstParams.changed_files as string[]).sort()).toEqual( + [AUX_TRIPLE_FILE, driftedFile].sort(), + ) + const firstFingerprint = firstParams.snapshot_fingerprint as string + expect(firstFingerprint).toMatch(/^v3:[a-f0-9]{64}$/) + + // The pass stores markers for exactly the attested file set. + expect( + gen.next( + reviewerResult(firstFingerprint, [AUX_TRIPLE_FILE, driftedFile]), + ).value, + ).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'context-pruner' }, + }) + const markersAfterFirstPass = (agentState as any).base2ActiveWork + .securityReviewFileMarkers as Record + expect(Object.keys(markersAfterFirstPass).sort()).toEqual( + [AUX_TRIPLE_FILE, driftedFile].sort(), + ) + const stableMarker = markersAfterFirstPass[AUX_TRIPLE_FILE] + // String snapshot BEFORE the drift: markersAfterFirstPass aliases the + // live map, so a property read at assertion time would see the refreshed + // value and the not.toBe below would compare the map against itself. + const driftedMarkerBefore = markersAfterFirstPass[driftedFile] + + // Second iteration: only the drifted file's BYTES change. + writeFileSync(driftedFile, 'export const policy = "v2"\n') + expect(gen.next().value).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + expect(gen.next(finishStepWithToolResult({})).value).toMatchObject({ + toolName: 'git_status', + }) + + // The re-review spawn is SCOPED to the drifted subset: changed_files + // contains ONLY the newly drifted file, and the attestation fingerprint is + // derived from exactly that list, so it differs from the first pass. + const reReviewSpawn = gen.next( + feedJson({ status: ` M ${AUX_TRIPLE_FILE}\n M ${driftedFile}` }), + ) + expect(reReviewSpawn.value).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'security-reviewer' }, + includeToolCall: false, + }) + const reReviewParams = (reReviewSpawn.value as any).input.params + expect(reReviewParams.changed_files).toEqual([driftedFile]) + const reReviewFingerprint = reReviewParams.snapshot_fingerprint as string + expect(reReviewFingerprint).toMatch(/^v3:[a-f0-9]{64}$/) + expect(reReviewFingerprint).not.toBe(firstFingerprint) + + // The scoped pass MERGES into the map: the unchanged file's stored marker + // survives and the drifted file's marker is refreshed. + expect( + gen.next(reviewerResult(reReviewFingerprint, [driftedFile])).value, + ).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'context-pruner' }, + }) + const markersAfterReReview = (agentState as any).base2ActiveWork + .securityReviewFileMarkers as Record + expect(markersAfterReReview[AUX_TRIPLE_FILE]).toBe(stableMarker) + expect(markersAfterReReview[driftedFile]).not.toBe(driftedMarkerBefore) + }) + + test('specialist per-file marker credit re-reviews only the drifted file and retains the unchanged marker', () => { + // Two real reliability-routed files (state/ dir + session/queue stems) so + // one specialist pass attests BOTH with real, drift-able bytes on disk. + mkdirSync(`${SPECIALIST_SCRATCH_ROOT}/state`, { recursive: true }) + const unchangedFile = `${SPECIALIST_SCRATCH_ROOT}/state/queue.ts` + writeFileSync(SPECIALIST_FILE, 'export const session = "v1"\n') + writeFileSync(unchangedFile, 'export const queue = "v1"\n') + const base2 = createBase2('default') + const agentState = { + agentId: 'base2-custom', + base2ActiveWork: specialistSeed({ + changedFiles: [SPECIALIST_FILE, unchangedFile], + touchedFiles: [SPECIALIST_FILE, unchangedFile], + pendingGateFiles: [SPECIALIST_FILE, unchangedFile], + auxGatesLastPendingFiles: [SPECIALIST_FILE, unchangedFile], + }), + } + const gen = base2.handleSteps!({ + agentState, + prompt: 'Please finish the pending reliability finding.', + params: {}, + } as any) + + // Resumed-state prelude. + expect(gen.next().value).toMatchObject({ + toolName: 'git_status', + input: {}, + }) + expect( + gen.next( + feedJson({ status: ` M ${SPECIALIST_FILE}\n M ${unchangedFile}` }), + ).value, + ).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'context-pruner' }, + }) + expect(gen.next().value).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + expect(gen.next(finishStepWithToolResult({})).value).toMatchObject({ + toolName: 'git_status', + input: {}, + }) + + // First specialist pass: no marker map yet, so the spawn covers BOTH + // routed files and the bundle freezes before the spawn. + expect( + gen.next( + feedJson({ status: ` M ${SPECIALIST_FILE}\n M ${unchangedFile}` }), + ).value, + ).toMatchObject({ + toolName: 'get_change_review_bundle', + includeToolCall: false, + }) + const firstSpawn = gen.next( + feedJson({ + snapshotId: 'specialist-credit-snapshot', + files: [SPECIALIST_FILE, unchangedFile], + }), + ) + expect(firstSpawn.value).toMatchObject({ + toolName: 'spawn_agents', + input: { agents: [{ agent_type: 'reliability-reviewer' }] }, + includeToolCall: false, + }) + const firstParams = (firstSpawn.value as any).input.agents[0].params + expect((firstParams.files as string[]).sort()).toEqual( + [SPECIALIST_FILE, unchangedFile].sort(), + ) + const firstFingerprint = specialistFingerprintFromSpawn(firstSpawn.value) + + // The passing pass stores per-specialist markers for exactly the files + // this pass attested. + expect( + gen.next( + spawnedReviewerResult('reliability-reviewer', firstFingerprint, [ + SPECIALIST_FILE, + unchangedFile, + ]), + ).value, + ).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'context-pruner' }, + }) + const markersAfterFirstPass = (agentState as any).base2ActiveWork + .specialistReviewFileMarkers?.[ + 'reliability-reviewer' + ] as Record + expect(Object.keys(markersAfterFirstPass).sort()).toEqual( + [SPECIALIST_FILE, unchangedFile].sort(), + ) + const unchangedMarker = markersAfterFirstPass[unchangedFile] + const driftedMarkerBefore = markersAfterFirstPass[SPECIALIST_FILE] + expect(unchangedMarker).toBeDefined() + expect(driftedMarkerBefore).toBeDefined() + + // Only the drifted file's bytes change before the second iteration. + writeFileSync(SPECIALIST_FILE, 'export const session = "v2"\n') + expect(gen.next().value).toMatchObject({ toolName: 'add_message' }) + expect(gen.next().value).toBe('STEP') + expect(gen.next(finishStepWithToolResult({})).value).toMatchObject({ + toolName: 'git_status', + }) + + // The re-review spawn is SCOPED: params.files contains ONLY the drifted + // file, and its scoped attestation fingerprint differs from the first pass. + expect( + gen.next( + feedJson({ status: ` M ${SPECIALIST_FILE}\n M ${unchangedFile}` }), + ).value, + ).toMatchObject({ + toolName: 'get_change_review_bundle', + includeToolCall: false, + }) + const secondSpawn = gen.next( + feedJson({ + snapshotId: 'specialist-credit-snapshot-refreshed', + files: [SPECIALIST_FILE], + }), + ) + expect(secondSpawn.value).toMatchObject({ + toolName: 'spawn_agents', + input: { agents: [{ agent_type: 'reliability-reviewer' }] }, + includeToolCall: false, + }) + const secondParams = (secondSpawn.value as any).input.agents[0].params + expect(secondParams.files).toEqual([SPECIALIST_FILE]) + const secondFingerprint = specialistFingerprintFromSpawn(secondSpawn.value) + expect(secondFingerprint).not.toBe(firstFingerprint) + + // The scoped pass MERGES markers: the unchanged file RETAINS its stored + // marker and the drifted file's marker is refreshed. + expect( + gen.next( + spawnedReviewerResult('reliability-reviewer', secondFingerprint, [ + SPECIALIST_FILE, + ]), + ).value, + ).toMatchObject({ + toolName: 'spawn_agent_inline', + input: { agent_type: 'context-pruner' }, + }) + const markersAfterSecondPass = (agentState as any).base2ActiveWork + .specialistReviewFileMarkers?.[ + 'reliability-reviewer' + ] as Record + expect(markersAfterSecondPass[unchangedFile]).toBe(unchangedMarker) + expect(markersAfterSecondPass[SPECIALIST_FILE]).not.toBe( + driftedMarkerBefore, + ) + }) + test('revalidates an owed specialist reviewer as aux-owned across turns before the final code-reviewer', () => { const base2 = createBase2('default') // Seed the turn so the marker is already owed to a specialist, simulating diff --git a/cli/src/components/renderers/compaction-box.tsx b/cli/src/components/renderers/compaction-box.tsx index 4f9c2da71e..04685608c7 100644 --- a/cli/src/components/renderers/compaction-box.tsx +++ b/cli/src/components/renderers/compaction-box.tsx @@ -33,6 +33,7 @@ const CATEGORY_LABEL: Record = { toolResults: 'tool results', todos: 'todos', fileReads: 'file reads', + boundedFileReads: 'bounded reads', subagents: 'subagents', userAssistantMessages: 'conversation', } diff --git a/cli/src/types/chat.ts b/cli/src/types/chat.ts index 4a4268fff0..08d270b0fa 100644 --- a/cli/src/types/chat.ts +++ b/cli/src/types/chat.ts @@ -208,6 +208,7 @@ export type CompactionCategoryDelta = { | 'toolResults' | 'todos' | 'fileReads' + | 'boundedFileReads' | 'subagents' | 'userAssistantMessages' beforeTokens: number diff --git a/common/src/types/print-mode.ts b/common/src/types/print-mode.ts index 5047fe91df..b9d8294f07 100644 --- a/common/src/types/print-mode.ts +++ b/common/src/types/print-mode.ts @@ -218,10 +218,14 @@ const contextCategoryStatsSchema = z.object({ messages: z.number(), }) +// `boundedFileReads` is optional so persisted/replayed compaction events +// emitted before the bounded-vs-whole-file telemetry split (which lack the +// key) keep validating; the runtime always emits it for new events. const contextCategorySummarySchema = z.object({ toolResults: contextCategoryStatsSchema, todos: contextCategoryStatsSchema, fileReads: contextCategoryStatsSchema, + boundedFileReads: contextCategoryStatsSchema.optional(), subagents: contextCategoryStatsSchema, userAssistantMessages: contextCategoryStatsSchema, }) @@ -292,6 +296,7 @@ export const printModeContextCompactionSchema = z.object({ 'toolResults', 'todos', 'fileReads', + 'boundedFileReads', 'subagents', 'userAssistantMessages', ]) diff --git a/docs/agents-and-tools.md b/docs/agents-and-tools.md index baf00dc99b..4ac72d7ac3 100644 --- a/docs/agents-and-tools.md +++ b/docs/agents-and-tools.md @@ -502,7 +502,7 @@ For backward compatibility, the `codebuff` command prefix may still work as a co > The canonical Gate vs Specialists matrix and Params Contract live in [agents/base2/quality-prompt-section.ts](../agents/base2/quality-prompt-section.ts) (`specialistRoutingSection`) and [agents/guides/specialist-routing.md](../agents/guides/specialist-routing.md) — this doc links there rather than duplicating. -The orchestrator (`base2` / `base-deep`, via the shared `createBase2` generator) runs three automated phase-gates around the existing validation + code-reviewer gate. Each gate is idempotent per pending gate-file set: it fires exactly once for a given set of edited files, and its done-flag resets only when the pending file set changes (order-insensitive). All three gates are guarded by the `runValidationGate` flag, so `base2-fast` / `base2-fast-no-validation` skip them. +The orchestrator (`base2` / `base-deep`, via the shared `createBase2` generator) runs three automated phase-gates around the existing validation + code-reviewer gate. Each gate is idempotent per pending gate-file set: it fires exactly once for a given set of edited files, and its done-flag resets only when the pending file set changes (order-insensitive). Security-reviewer and specialist credit additionally carry a PHASE-4 per-file, content-keyed ledger (`securityReviewFileMarkers` / `specialistReviewFileMarkers`): re-absorbing the same still-dirty path with UNCHANGED bytes does not re-fire the gate, while a confirmed byte change to an attested file re-fires it as a SCOPED re-review of exactly the drifted subset (legacy serialized state without the marker maps keeps the scalar whole-set fingerprints and their semantics). All three gates are guarded by the `runValidationGate` flag, so `base2-fast` / `base2-fast-no-validation` skip them. The gate predicates are self-contained string/regex matchers defined inline inside `createBase2.handleSteps`. They intentionally do NOT import `micromatch` or any module-scope binding, because `handleSteps` is serialized via `.toString()` and reconstructed with `new Function(...)`; module-scope imports would be `undefined` at reconstruction time. The glob list mirrors the advisory `securityReviewSection` in `agents/base2/quality-prompt-section.ts` so the automated gate and the advisory prompt agree on what counts as security-sensitive. @@ -518,6 +518,8 @@ Each aux gate is predicate-gated: if no pending file matches its relevance predi The three done-flags (`testWriterGateDone`, `docWriterGateDone`, `preEditSecurityReviewDone`) and the `auxGatesLastPendingFiles` snapshot live on `Base2ActiveWorkState` (`agents/base2/gate-state.ts`). `detectPendingGateFileSetChange` + `resetAuxGateFlags` reset the flags when the pending file set changes (compared via `gateFileSetsEqual`, order-insensitive). The reset predicate compares the AUX-RELEVANT subset of pending files — files that at least one aux predicate would act on — so newly-written aux outputs (test files created by `test-writer`, doc files updated by `doc-writer`) do not perturb the snapshot and do not re-trigger the aux gates for the same pending file set. +PHASE 4 also makes reviewer-family credit content-keyed per file. On a security or specialist pass the gate stores `readGateFileContentMarker` hashes for exactly the attested files (`securityReviewFileMarkers`, `specialistReviewFileMarkers[agentType]`); a later encounter re-reviews only the files whose markers no longer match the live bytes, merging fresh markers on pass so credit for unchanged files survives. Marker eviction in the changed-files ledger is content-aware: only confirmed byte drift drops a file's security credit, so the per-iteration git-status re-absorption of an unchanged dirty path cannot loop security-reviewer. The validation side mirrors this: the newest full-assurance `validationEvidence` entry carries `fileMarkers`, and when it covers exactly the current gate-scope set with every marker still matching, the hook run is skipped and the receipt is reused verbatim (any reduced-assurance summary, legacy entry without markers, or marker mismatch falls through to a fresh hook run). + ## Concurrent gate isolation (`selfMutatedPaths`) Mid-turn git-status absorption must not claim foreign worktree dirt from concurrent Openbuff instances or external editors. Terminal steps no longer auto-absorb every newly dirty path; basher/codegen writes re-enter the validation/reviewer gate only through published ownership. `touchedPaths` is best-effort ownership attribution for that absorb path, not authorization to mutate or finalize. @@ -584,7 +586,7 @@ include every pending file. Review guidance also covers meaningful test assertions, public and persisted compatibility, package boundaries, generated-artifact freshness, migration safety, and bounded resource use. -The `code-reviewer` gate decides whether a turn may finish green. **Only structured `verdict === 'LOOKS_GOOD'` permits gate pass / finalization** (after coverage and requirement adequacy checks). `NON_BLOCKING` does **not** finalize: its findings are collected as open repair targets and enter the same repair-editor / test-writer re-review loop used for `BLOCKING`. Both BLOCKING and NON_BLOCKING rounds increment the reviewer repair counter for telemetry. Repair loops default to **unlimited / progress-gated** (no-progress fingerprint and incomplete-receipt exits); optional hard caps remain via `maxReviewerRepairRounds` / `OPENBUFF_MAX_REVIEWER_REPAIR_ROUNDS` (max `20`). Validation-hook and specialist repair loops are likewise unlimited by default, with optional caps via `maxRepairRounds` / `maxSpecialistRepairRounds` and envs `OPENBUFF_MAX_REPAIR_ROUNDS` / `OPENBUFF_MAX_SPECIALIST_REPAIR_ROUNDS` (max `20`). Already-credited (`gatePassedFiles`) dirty task files stay out of gate scope so they do not re-arm validation/review while remaining dirty for commit UX. Coverage-missing and **in-scope** incomplete requirements still hard-block. The orchestrator parses the reviewer's tool result to extract a finalization verdict (`LOOKS_GOOD` or empty string `''`) and to surface any repair findings. The parser prefers structured (parsed-object) verdicts over text-mode fallbacks. The parsing helpers live in `agents/base2/gate-reviewer.ts` and are mirrored inline inside `createBase2.handleSteps` (the mirror is parity-tested by `agents/__tests__/gate-reviewer.test.ts`). +The `code-reviewer` gate decides whether a turn may finish green. **Only structured `verdict === 'LOOKS_GOOD'` permits gate pass / finalization** (after coverage and requirement adequacy checks). `NON_BLOCKING` does **not** finalize: its findings are collected as open repair targets and enter the same repair-editor / test-writer re-review loop used for `BLOCKING`. Both BLOCKING and NON_BLOCKING rounds increment the reviewer repair counter for telemetry. Repair loops default to **unlimited / progress-gated** (no-progress fingerprint and incomplete-receipt exits); optional hard caps remain via `maxReviewerRepairRounds` / `OPENBUFF_MAX_REVIEWER_REPAIR_ROUNDS` (max `20`). Validation-hook and specialist repair loops are likewise unlimited by default, with optional caps via `maxRepairRounds` / `maxSpecialistRepairRounds` and envs `OPENBUFF_MAX_REPAIR_ROUNDS` / `OPENBUFF_MAX_SPECIALIST_REPAIR_ROUNDS` (max `20`). Already-credited (`gatePassedFiles`) dirty task files stay out of gate scope so they do not re-arm validation/review while remaining dirty for commit UX, and reviewer prompts list files already attested `LOOKS_GOOD` in earlier rounds (with an unchanged-bytes marker where the recorded fingerprint still matches) so review depth concentrates on unattested or drifted files without suppressing fresh findings. Coverage-missing and **in-scope** incomplete requirements still hard-block. The orchestrator parses the reviewer's tool result to extract a finalization verdict (`LOOKS_GOOD` or empty string `''`) and to surface any repair findings. The parser prefers structured (parsed-object) verdicts over text-mode fallbacks. The parsing helpers live in `agents/base2/gate-reviewer.ts` and are mirrored inline inside `createBase2.handleSteps` (the mirror is parity-tested by `agents/__tests__/gate-reviewer.test.ts`). ### Parent-owned / process requirements diff --git a/packages/agent-runtime/src/util/__tests__/messages.test.ts b/packages/agent-runtime/src/util/__tests__/messages.test.ts index 9744519b04..95ba6800b7 100644 --- a/packages/agent-runtime/src/util/__tests__/messages.test.ts +++ b/packages/agent-runtime/src/util/__tests__/messages.test.ts @@ -322,6 +322,100 @@ describe('getContextCategoryTelemetry', () => { Object.values(telemetry).reduce((sum, entry) => sum + entry.messages, 0), ).toBe(messages.length) }) + + it('classifies read_files calls with window selectors as bounded reads', () => { + spyOn(tokenCounter, 'countTokensJson').mockImplementation( + (value) => JSON.stringify(value).length, + ) + + const messages: Message[] = [ + { + role: 'assistant', + content: [ + { + type: 'tool-call', + toolCallId: 'call-windowed', + toolName: 'read_files', + input: { windows: [{ path: 'src/big.ts', window: 2 }] }, + }, + ], + }, + { + role: 'tool', + toolName: 'read_files', + toolCallId: 'call-windowed', + content: jsonToolResult([{ path: 'src/big.ts', content: 'block' }]), + }, + ] + + const telemetry = getContextCategoryTelemetry(messages) + + expect(telemetry.boundedFileReads.messages).toBe(1) + expect(telemetry.fileReads.messages).toBe(0) + }) + + it('classifies read_files calls without selectors (paths-only) as whole-file reads', () => { + spyOn(tokenCounter, 'countTokensJson').mockImplementation( + (value) => JSON.stringify(value).length, + ) + + const messages: Message[] = [ + { + role: 'assistant', + content: [ + { + type: 'tool-call', + toolCallId: 'call-whole', + toolName: 'read_files', + input: { paths: ['src/file.ts'] }, + }, + ], + }, + { + role: 'tool', + toolName: 'read_files', + toolCallId: 'call-whole', + content: jsonToolResult([{ path: 'src/file.ts', content: 'whole' }]), + }, + ] + + const telemetry = getContextCategoryTelemetry(messages) + + expect(telemetry.fileReads.messages).toBe(1) + expect(telemetry.boundedFileReads.messages).toBe(0) + }) + + it('classifies read_outline and query_index as bounded reads and read_subtree as whole-file', () => { + spyOn(tokenCounter, 'countTokensJson').mockImplementation( + (value) => JSON.stringify(value).length, + ) + + const messages: Message[] = [ + { + role: 'tool', + toolName: 'read_outline', + toolCallId: 'outline-1', + content: jsonToolResult({ outline: 'outline body' }), + }, + { + role: 'tool', + toolName: 'query_index', + toolCallId: 'query-1', + content: jsonToolResult({ matches: [] }), + }, + { + role: 'tool', + toolName: 'read_subtree', + toolCallId: 'subtree-1', + content: jsonToolResult({ files: [] }), + }, + ] + + const telemetry = getContextCategoryTelemetry(messages) + + expect(telemetry.boundedFileReads.messages).toBe(2) + expect(telemetry.fileReads.messages).toBe(1) + }) }) describe('trimMessagesToFitTokenLimit', () => { @@ -1645,12 +1739,13 @@ describe('trimMessagesToFitTokenLimitWithReport eviction policy', () => { mock.restore() }) - it('evicts fileReads and toolResults before user+assistant turns', () => { + it('evicts bounded file reads and toolResults before user+assistant turns', () => { spyOn(tokenCounter, 'countTokensJson').mockImplementation( (value) => JSON.stringify(value).length, ) expect(CONTEXT_EVICTION_PRIORITY).toEqual([ + 'boundedFileReads', 'fileReads', 'toolResults', 'subagents', @@ -1706,7 +1801,7 @@ describe('trimMessagesToFitTokenLimitWithReport eviction policy', () => { ), ).toBe(false) expect(report.removedCategories).toEqual( - expect.arrayContaining(['fileReads', 'toolResults']), + expect.arrayContaining(['boundedFileReads', 'toolResults']), ) expect(report.removedCategories).not.toContain('userAssistantMessages') expect(report.fitsBudget).toBe(true) diff --git a/packages/agent-runtime/src/util/messages.ts b/packages/agent-runtime/src/util/messages.ts index dade0ed117..67604fc631 100644 --- a/packages/agent-runtime/src/util/messages.ts +++ b/packages/agent-runtime/src/util/messages.ts @@ -211,6 +211,7 @@ export type ContextCategory = | 'toolResults' | 'todos' | 'fileReads' + | 'boundedFileReads' | 'subagents' | 'userAssistantMessages' @@ -223,11 +224,52 @@ const emptyContextCategorySummary = (): ContextCategorySummary => ({ toolResults: { tokens: 0, percent: 0, messages: 0 }, todos: { tokens: 0, percent: 0, messages: 0 }, fileReads: { tokens: 0, percent: 0, messages: 0 }, + boundedFileReads: { tokens: 0, percent: 0, messages: 0 }, subagents: { tokens: 0, percent: 0, messages: 0 }, userAssistantMessages: { tokens: 0, percent: 0, messages: 0 }, }) -function getContextCategory(message: Message): ContextCategory { +/** + * Assistant tool-call inputs indexed by toolCallId. Tool-result messages carry + * only `toolName`/`toolCallId`, so classifying a read_files result as bounded + * or whole-file requires replaying the paired assistant call's input. + */ +function buildToolCallInputsByCallId( + messages: Message[], +): Map> { + const inputsByCallId = new Map>() + for (const message of messages) { + if (message.role !== 'assistant') continue + for (const part of message.content) { + if (part.type === 'tool-call') { + inputsByCallId.set(part.toolCallId, part.input) + } + } + } + return inputsByCallId +} + +/** Selector keys whose non-empty presence makes a read_files call bounded. */ +const BOUNDED_READ_SELECTOR_KEYS = [ + 'windows', + 'around', + 'symbol', + 'symbols', +] as const + +function isBoundedReadFilesCall(input: unknown): boolean { + if (!input || typeof input !== 'object') return false + const record = input as Record + return BOUNDED_READ_SELECTOR_KEYS.some((key) => { + const selectors = record[key] + return Array.isArray(selectors) && selectors.length > 0 + }) +} + +function getContextCategory( + message: Message, + toolCallInputsByCallId: ReadonlyMap>, +): ContextCategory { if (message.role !== 'tool') { return 'userAssistantMessages' } @@ -237,12 +279,22 @@ function getContextCategory(message: Message): ContextCategory { } if ( - message.toolName === 'read_files' || - message.toolName === 'find_files' || - message.toolName === 'read_subtree' || message.toolName === 'read_outline' || - message.toolName === 'query_index' + message.toolName === 'query_index' || + message.toolName === 'find_files' ) { + return 'boundedFileReads' + } + + if (message.toolName === 'read_files') { + return isBoundedReadFilesCall( + toolCallInputsByCallId.get(message.toolCallId), + ) + ? 'boundedFileReads' + : 'fileReads' + } + + if (message.toolName === 'read_subtree') { return 'fileReads' } @@ -253,8 +305,14 @@ function getContextCategory(message: Message): ContextCategory { return 'toolResults' } -/** Eviction order: cheapest-to-recover context first, conversation last. */ +/** + * Eviction order: cheapest-to-recover context first, conversation last. + * Bounded block reads (read_files windows/around/symbol selectors, plus + * read_outline, query_index, find_files) re-fetch cheapest, so they evict + * before whole-file reads; both precede generic tool results and conversation. + */ export const CONTEXT_EVICTION_PRIORITY: readonly ContextCategory[] = [ + 'boundedFileReads', 'fileReads', 'toolResults', 'subagents', @@ -266,10 +324,11 @@ export function getContextCategoryTelemetry( messages: Message[], ): ContextCategorySummary { const summary = emptyContextCategorySummary() + const toolCallInputsByCallId = buildToolCallInputsByCallId(messages) let totalTokens = 0 for (const message of messages) { - const category = getContextCategory(message) + const category = getContextCategory(message, toolCallInputsByCallId) const tokens = countTokensJson(message) summary[category].tokens += tokens summary[category].messages += 1 @@ -509,6 +568,10 @@ export function trimMessagesToFitTokenLimitWithReport(params: { } const initialContextCategoryTelemetry = getContextCategoryTelemetry(messages) + // Built from the ORIGINAL history once: read_files results classify as + // bounded vs whole-file by replaying the paired assistant call's input, and + // reconciliation later in the trim may drop assistant tool-call parts. + const toolCallInputsByCallId = buildToolCallInputsByCallId(messages) const shortenedMessages: Message[] = [] let numKept = 0 @@ -611,7 +674,8 @@ export function trimMessagesToFitTokenLimitWithReport(params: { const message = shortenedMessages[index] if (message.keepDuringTruncation || removedIndices.has(index)) continue if (answersPinnedToolResult(message)) continue - if (getContextCategory(message) !== category) continue + if (getContextCategory(message, toolCallInputsByCallId) !== category) + continue const mergesPrevious = removedIndices.has(index - 1) const mergesNext = removedIndices.has(index + 1) removedIndices.add(index) @@ -704,7 +768,8 @@ export function trimMessagesToFitTokenLimitWithReport(params: { if (runningTokens <= maxMessageTokens) break const message = trimmedMessages[index] if (escalationRemoved.has(index)) continue - if (getContextCategory(message) !== category) continue + if (getContextCategory(message, toolCallInputsByCallId) !== category) + continue if (!isEvictableDuringEscalation(message)) continue escalationRemoved.add(index) // Maintain the running total per drop instead of recounting the whole From 31b7576e1f4b5f2075c6d35762a049cc31c1ffa1 Mon Sep 17 00:00:00 2001 From: AnzoBenjamin Date: Sat, 5 Sep 2026 23:53:55 +0300 Subject: [PATCH 4/4] docs(knowledge): refresh cli + common knowledge for compaction progress changes Record the 2026-09-05 cli/src and common/src changes: the additive context_compaction_progress event and its monotonic-percent consumer, self-dismissing transient compaction cards, the boundedFileReads category, detectImageMediaTypeFromBytes, removal of the subagent wall-clock timeout surface, receipt contextUsage, and AgentState.lastSetOutputError. Clears the two guard:memory-drift staleness findings that were blocking the pre-push check:ci-local hook. --- cli/knowledge.md | 2 ++ common/knowledge.md | 2 ++ 2 files changed, 4 insertions(+) diff --git a/cli/knowledge.md b/cli/knowledge.md index bbc3a44568..467fac6dba 100644 --- a/cli/knowledge.md +++ b/cli/knowledge.md @@ -915,3 +915,5 @@ Streaming markdown renders as plain text until the message or agent finishes. Th - _Knowledge refresh 2026-08-31 (followups): `handleRuntimeError` in `cli/src/utils/sdk-event-handlers.ts` now splits runtime error events by `autoRecovering` — auto-recovering notices log at `debug` (`'SDK auto-recovering runtime notice'`) with no visible error banner, while genuine failures still log at `error` (`'SDK runtime error event'`) and render. Tool-ordering rejections (the `suggest_followups` gate/ordering rejections and the pre-gate `git-committer withheld` chunk) arrive with a concise `userMessage` plus `autoRecovering: true`, so the model still receives the full `message` via the `TOOL_CALL_ERROR` path in `packages/agent-runtime/src/tools/stream-parser.ts`; `git-committer blocked by unvalidated dirty file(s)` deliberately stays user-visible because it asks the user to reply `COMMIT ANYWAY`._ - _Knowledge refresh 2026-08-31 (UI polish): `cli/src/components/status-bar.tsx` is now three regions — status label left, chip cluster left-aligned in the growing middle (`flexGrow: 1` + `flexBasis: 0`), and every width-varying control (scroll-to-bottom, then the `■ Esc` stop hint) in a `flexShrink: 0` right region with no `minWidth: 0`, so a hover cannot reflow the label or the chips. `cli/src/components/scroll-to-bottom-button.tsx` exports `SCROLL_HINT_LABEL`/`SCROLL_GLYPH` plus `string-width`-derived `SCROLL_BUTTON_WIDTH` (10) and `SCROLL_BUTTON_COMPACT_WIDTH` (3) and renders at a fixed width in both hover states; its `isScrollButtonCompact(width)` predicate is the single source `StatusBar` also passes as `scrollButtonCompact`, so `statusBarChipBudget`'s duplicated `SCROLL_BUTTON_RESERVATION`/`SCROLL_BUTTON_COMPACT_RESERVATION` in `cli/src/utils/status-bar-chips.ts` always reserve the columns actually rendered (test-enforced agreement, since the util must not import a component module). `cli/src/components/renderers/completion-summary-box.tsx` renders a titled `Run summary` `HarnessBox` (`gap={0}`, `paddingBottom={0}`) of aligned `Label value` rows built from `ROW_LABELS` + derived `LABEL_COLUMN_WIDTH`, with no status emoji — meaning lives in the value words, not color. Reconciler-level coverage for the status bar lives in `cli/src/components/__tests__/status-bar.test.tsx`, which reuses the dev-only `renderTest`/`renderFrame` convention from `text-nesting.test.tsx` (`@opentui/react/test-utils` cannot be imported under `NODE_ENV=production`)._ + +- _Knowledge refresh 2026-09-05 (compaction progress + self-dismissing cards): `cli/src/utils/sdk-event-handlers.ts` consumes the additive `context_compaction_progress` event in `handleContextCompactionProgress`, clamping each reported percent to a whole 0..100 and writing only the MAXIMUM of what the card or notice already holds, so two producers for one pass (the agent loop's milestones and the inline spawn path's activity ticks) plus replayed or out-of-order events can never rewind the bar. Card updates stay root-scoped and paired by `runId`, while the status-bar notice tracks root and nested passes alike; a progress event never creates a notice or revives a settled one. `compactionResultIsDegraded` is the single site deciding `CompactionContentBlock.transient`: a healthy settled pass is stamped `progressPercent: 100` plus `transient: true`, while a mechanical or request-time trim, a pass that missed its budget, an escalated pass, and a low-yield streak all stay permanent warning cards, as do declined and interrupted passes. `cli/src/components/renderers/compaction-box.tsx` renders `cli/src/components/progress-bar.tsx` for pending and transient passes and hides a transient card after a short hold, but hiding is purely visual: `dropTransientCompactionBlocks` in `cli/src/utils/message-block-helpers.ts` is what actually removes it from state, composed into both the turn-end path in `handleFinish` and the abort path in `cli/src/hooks/helpers/send-message.ts`, so a self-dismissing card can never persist to the transcript. `cli/src/utils/status-bar-chips.ts` reports `⇲ compacting NN%` at md/lg when a live percent is finite and above zero and otherwise keeps the previous ellipsis label (xs/sm labels unchanged). `cli/src/types/chat.ts` carries the new optional `progressPercent` and `transient` block fields, the notice `progressPercent`, and the `boundedFileReads` category. `cli/src/components/terminal-command-display.tsx` now shows a timeout label only for a finite positive bound, because `timeout_seconds` defaults to no timeout and the removed 30s default would otherwise be implied. Coverage: `cli/src/utils/__tests__/sdk-event-handlers.test.ts`, `cli/src/utils/__tests__/status-bar-chips.test.ts`, and `cli/src/components/__tests__/sweep-boxes.test.tsx`._ diff --git a/common/knowledge.md b/common/knowledge.md index 6cb61b563d..952e2a86ad 100644 --- a/common/knowledge.md +++ b/common/knowledge.md @@ -69,6 +69,8 @@ This package contains code shared across the Openbuff monorepo, especially the l - _Knowledge refresh 2026-09-01 (request-time context trim): `common/src/types/print-mode.ts` gained the additive `context_request_trim` variant reporting the SDK's request-time emergency trim — the last-line-of-defense drop applied at dispatch when a request's messages still exceed the provider-safe budget after every runtime brake ran. It carries required `messageBudgetTokens`/`beforeTokens`/`afterTokens`/`beforeMessages`/`afterMessages` plus optional `runId`, `ancestorRunIds`, `agentId`, `resolvedContextWindowTokens`, and `model`. It is a DIFFERENT brake from `context_compaction` and its `context_compaction_status` pair, so the two must never be merged or counted as one pass. The existing `context_window` variant gained optional `compactionTriggerTokens`/`compactionTargetTokens` reporting the runtime's model-aware semantic-compaction budget; both are derived from the RAW resolved model window and are deliberately NOT clamped by `maxContextLength`, so `compactionTriggerTokens > max` is a legitimate payload and a consumer rendering trigger against `max` must clamp or suppress it itself. `common/src/types/contracts/llm.ts` re-exports the `RequestContextTrimInfo` payload type (declared in `print-mode`) for the optional `onRequestContextTrimmed` callback on the published `promptAiSdk`/`promptAiSdkStream`/`promptAiSdkStructured` signatures: purely observational, fires only when the trim actually dropped messages, can never affect the trim result, and a throwing consumer is caught and logged rather than aborting dispatch. All three additions are additive/optional, so callers that omit them and consumers that ignore unknown `event.type` values keep their previous behavior._ +- _Knowledge refresh 2026-09-05 (compaction progress, image sniffing, subagent timeout removal): `common/src/types/print-mode.ts` gained the additive `context_compaction_progress` variant (`runId`, `ancestorRunIds`, optional `agentId`, `percent`, a `phase` of `analyzing`/`summarizing`/`applying`, optional `contextTokens`/`targetBudgetTokens`). It is a NEW member of the discriminated union rather than a widened `context_compaction_status` `state` enum, precisely because an added enum member breaks a consumer switching exhaustively over that enum while an unknown event `type` is already contractually a no-op. It appears only BETWEEN a `started` and its matching `settled` for the same `runId`; `percent` is a best-effort monotonic estimate, so a consumer must clamp with its own maximum rather than trust arrival order or range, and `percent: 100` is never a claim that space was reclaimed — the terminal `context_compaction` result stays the only signal of that. The category schemas gained an optional `boundedFileReads` (optional so replayed events emitted before the bounded-vs-whole-file split still validate). `common/src/constants/images.ts` gained `detectImageMediaTypeFromBytes`, which matches PNG, JPEG, GIF, BMP, WEBP, and TIFF magic numbers, requires the WEBP form tag at offset 8 so RIFF-fronted audio is not misreported as an image, and returns null for short or unsigned buffers; `common/src/__tests__/images.test.ts` pins that it only ever returns MIME strings the extension map already publishes. The subagent wall-clock timeout surface was removed: `defaultTimeoutMs` is gone from `common/src/types/agent-template.ts` and `common/src/types/dynamic-agent-template.ts`, the per-spawn `timeout_seconds` entry field is gone from `common/src/tools/params/tool/spawn-agents.ts`, `timeout` is gone from the tool-call request in `common/src/actions.ts`, and `common/src/tools/params/tool/run-terminal-command.ts` now defaults `timeout_seconds` to -1 (no timeout). That is a public tool-schema change, so regenerated tool definition sources must land in the same commit. `common/src/types/agent-handoff.ts` gained an optional observational `contextUsage` on the agent receipt (`tokens` plus optional `windowTokens`, `percentOfWindow`, `compactionCount`) so a parent can size later delegations, and `common/src/types/session-state.ts` gained `AgentState.lastSetOutputError` so the missing-structured-output retry names the real rejection._ + ## Scope Notes Openbuff is CLI/SDK-focused and local/BYOK. Do not add new dependencies from `common/` to hosted web, billing, credit, subscription, or BigQuery product surfaces. Provider-owned billing, quota, token usage, and OAuth flows may still be documented when they refer to the user's configured provider rather than an Openbuff-hosted product.