Skip to content

Commit c1b5cf2

Browse files
review: refuse --scripted-model with --scenario/--prompt; document dump's default limit (README, site en/zh); describe Stop as prompt-path dependent
1 parent f52359b commit c1b5cf2

5 files changed

Lines changed: 26 additions & 9 deletions

File tree

‎docs/audits/2026-09-03-host-lineage-matrix.md‎

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -151,8 +151,10 @@ Findings specific to this run:
151151
- **Turn boundaries.** Four `UserPromptSubmit` payloads (rows 2, 58, 105,
152152
137), four distinct `prompt_id`s; the sequential and nested spawns ran under
153153
the `<task-notification>` prompt (row 58), and the registry's `generation`
154-
on their lineage is that `prompt_id`. Each `-p` process produced its own
155-
`Stop` + `SessionEnd` (rows 34/102–103, 129–130, 135, 140–141); resuming
154+
on their lineage is that `prompt_id`. Each `-p` process produced exactly one
155+
`SessionEnd` (rows 103, 130, 135, 141), while `Stop` follows the prompt
156+
path: twice in turn 1 (row 34 before the re-prompt, row 102 after), once in
157+
turns 2 and 4 (rows 129, 140), never in the `/compact` turn; resuming
156158
re-registers nothing new — the root node stays the same and every resumed
157159
root-side hook still resolves `depth 0, resolution: native`.
158160
- **MCP correlation held at every depth.** The five `probe` calls (rows 29,
@@ -382,7 +384,7 @@ header table. What changed and what did not:
382384
| `PostToolUseFailure` "not observed" | Emitted by 2.1.259 for a failed MCP tool call, at the root and inside subagents, with the same `session_id`/`agent_id` carrier as `PostToolUse` (orchestration rows 77, 94, 116) | Corrected: observed; the registry already closed windows on `tool/failure` |
383385
| `PreCompact`/`PostCompact` "not observed" | Fired by a manual `/compact` sent as a resumed `-p` prompt, bracketing a `SessionStart source: compact`; root-only, no `agent_id`, own `prompt_id` (orchestration rows 132–134) | Corrected: inducible on demand |
384386
| Two spawns in one message would leave the registry's `toolCallId` uncertain (`siblingsUncertain`) | The host serialised them: `SubagentStart` for the first fired before the second `Agent` `PreToolUse` opened, so both claims were certain and agree with `tool_response.agentId` and the stream's `task_started` | Held stronger than assumed |
385-
| Single-turn `-p` sessions only | Four `-p` turns resumed into one `session_id`; `SessionStart source: resume`, one `Stop`/`SessionEnd` per turn, root-side lineage unchanged across turns | New coverage; no framework impact |
387+
| Single-turn `-p` sessions only | Four `-p` turns resumed into one `session_id`; `SessionStart source: resume`, one `SessionEnd` per invocation while `Stop` follows the prompt path (two in turn 1 around the `<task-notification>` re-prompt, none in the `/compact` turn), root-side lineage unchanged across turns | New coverage; no framework impact |
386388

387389
No lineage or projection code change was needed: the registry replay test
388390
(`packages/rsc-runtime/tests/lineage-registry.test.ts`) now runs against all

‎examples/host-test/README.md‎

Lines changed: 12 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -36,6 +36,9 @@ Two MCP servers ship in the plugin:
3636
- `host-test` (generated routes, `src/mcp/host-test/tools/`): `dump` (filter by
3737
any conversation/session/subagent id, `full` for raw lines) and `reset`. Each
3838
`dump` call records the request context the generated server mounted for it.
39+
A bare `dump` returns the newest 50 matching records — a whole log of a few
40+
hundred records overflows the tool-result document — so pass `limit` (up to
41+
5000) for more; `matched` and `total` always count the whole log.
3942
- `host-test-raw` (hand-rolled stdio factory, `src/mcp/host-test-raw.ts`):
4043
`probe` records the raw SDK request context — session id, JSON-RPC id,
4144
`_meta`, lifted envelope, negotiated client info — so hook↔MCP correlation is
@@ -82,9 +85,15 @@ pnpm --filter @agent-bundle-example/host-test probe:uninstall claude
8285
`{ "prompt": "..." }`). For Claude every turn after the first runs
8386
`claude -p --resume <session_id>` against the session the first turn's
8487
`system/init` envelope reported, so one capture holds a multi-turn session
85-
with `SessionStart source: resume`, a fresh `prompt_id` per turn, and one
86-
`Stop`/`SessionEnd` per turn. Codex and Cursor drivers take the first turn
87-
only and refuse longer scenarios. `scenarios/claude-orchestration.json` is
88+
with `SessionStart source: resume` and one `SessionEnd` per invocation.
89+
`Stop` follows the prompt path rather than the invocation: a turn whose
90+
background subagents finish re-prompts itself with a `<task-notification>`
91+
and stops twice, while a `/compact` turn submits no prompt and never stops
92+
(`fixtures/host-lineage/claude-2.1.259-orchestration.ndjson`). If the first
93+
turn reports no `session_id`, the capture fails instead of running the
94+
remaining turns as fresh sessions; `--scripted-model` plays a fixed
95+
transcript and refuses `--scenario`/`--prompt`. Codex and Cursor drivers take
96+
the first turn only and refuse longer scenarios. `scenarios/claude-orchestration.json` is
8897
the checked-in orchestration scenario (two parallel `Agent` spawns, one
8998
sequential spawn that nests another, the `host-test:host-test` skill, a
9099
plugin-command probe, a manual `/compact`, a final stop).

‎examples/host-test/scripts/probe.mjs‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
// uninstall. Nothing here touches the real ~/.claude, ~/.codex, or ~/.cursor.
55
//
66
// node scripts/probe.mjs install <claude|codex|cursor> [--no-auth] [--root <dir>]
7-
// node scripts/probe.mjs capture <claude|codex|cursor> [--prompt <text> | --scenario <file.json>] [--model <m>] [--timeout <ms>] [--scripted-model]
7+
// node scripts/probe.mjs capture <claude|codex|cursor> [--prompt <text> | --scenario <file.json> | --scripted-model] [--model <m>] [--timeout <ms>]
88
// node scripts/probe.mjs uninstall <claude|codex|cursor> [--keep-home]
99
// node scripts/probe.mjs status <claude|codex|cursor>
1010
import { spawn, spawnSync } from 'node:child_process';
@@ -296,6 +296,12 @@ const describeStream = (envelopes) => {
296296
*/
297297
const captureClaude = () => {
298298
const scripted = flags.scriptedModel === true;
299+
// The scripted model answers one canned transcript; it cannot follow a
300+
// scenario, so refuse the combination instead of silently running the
301+
// canned turn under the scenario's name.
302+
if (scripted && (flags.scenario !== undefined || flags.prompt !== undefined)) {
303+
throw new Error('--scripted-model plays a fixed transcript and ignores prompts: drop --scenario/--prompt, or run the scenario against the real model');
304+
}
299305
const turns = scripted ? ['You are exercising the host-test probe plugin. Do the scripted steps.'] : scenarioTurns();
300306
const baseArgs = [
301307
'--output-format', 'stream-json',

‎website/docs/en/examples/index.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -23,7 +23,7 @@ walked through here:
2323
| --- | --- |
2424
| [Worktree Proximity](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/worktree-proximity) | One root task and two child agents in linked worktrees, coordinated through durable notices. |
2525
| [RSC Agent Runtime](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/rsc-agent-runtime) | An opt-in architecture experiment: one RSC runtime shared by hooks, MCP tools, and an MCP App timeline. Not a public API. |
26-
| [Host Test](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/host-test) | A probe plugin that records every hook payload and MCP call a host delivers, with the `request.lineage` each one resolved to. Its `probe:install` / `probe:capture` / `probe:uninstall` scripts drive a real host in an isolated home; `probe:capture <host> --scenario <file.json>` runs an ordered list of prompts as one multi-turn session (Claude: `claude -p --output-format stream-json`, later turns `--resume`d), saving the model's own tool-use stream beside the hook log. The Workbench walkthrough needs no signed-in host; the probe lifecycle does. |
26+
| [Host Test](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/host-test) | A probe plugin that records every hook payload and MCP call a host delivers, with the `request.lineage` each one resolved to. Its `probe:install` / `probe:capture` / `probe:uninstall` scripts drive a real host in an isolated home; `probe:capture <host> --scenario <file.json>` runs an ordered list of prompts as one multi-turn session (Claude: `claude -p --output-format stream-json`, later turns `--resume`d), saving the model's own tool-use stream beside the hook log. The plugin's `dump` MCP tool returns the newest 50 matching records unless `limit` is given (`matched` counts the whole log). The Workbench walkthrough needs no signed-in host; the probe lifecycle does. |
2727

2828
## Running them
2929

‎website/docs/zh/examples/index.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,7 @@ description: '可运行的 agent-bundle 示例:从 Skill 起步项目到完整
2121
| --- | --- |
2222
| [Worktree Proximity](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/worktree-proximity) | 一个根任务与两个位于关联 worktree 中的子智能体,通过持久通知协调。 |
2323
| [RSC Agent Runtime](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/rsc-agent-runtime) | 一个可选启用的架构实验:一个 RSC 运行时同时服务钩子、MCP 工具与 MCP App 时间线。不是公开 API。 |
24-
| [Host Test](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/host-test) | 一个探针插件,记录宿主投递的每个钩子负载与 MCP 调用,以及各自解析出的 `request.lineage`。它的 `probe:install` / `probe:capture` / `probe:uninstall` 脚本在隔离的 home 中驱动真实宿主;`probe:capture <host> --scenario <file.json>` 把一组有序提示作为一个多轮会话运行(Claude 使用 `claude -p --output-format stream-json`,后续轮次以 `--resume` 续接),并把模型自身的工具调用流保存在钩子日志旁。Workbench 导览不需要已登录的宿主;探针生命周期则需要。 |
24+
| [Host Test](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/host-test) | 一个探针插件,记录宿主投递的每个钩子负载与 MCP 调用,以及各自解析出的 `request.lineage`。它的 `probe:install` / `probe:capture` / `probe:uninstall` 脚本在隔离的 home 中驱动真实宿主;`probe:capture <host> --scenario <file.json>` 把一组有序提示作为一个多轮会话运行(Claude 使用 `claude -p --output-format stream-json`,后续轮次以 `--resume` 续接),并把模型自身的工具调用流保存在钩子日志旁。插件的 `dump` MCP 工具默认只返回最新的 50 条匹配记录,除非传入 `limit`(`matched` 统计整份日志)。Workbench 导览不需要已登录的宿主;探针生命周期则需要。 |
2525

2626
## 如何运行
2727

0 commit comments

Comments
 (0)