Skip to content

Commit 8e181ec

Browse files
review: fail probe:capture when a multi-turn scenario cannot resume; document Host Test on the examples pages; releasable changeset
1 parent 344d765 commit 8e181ec

4 files changed

Lines changed: 26 additions & 20 deletions

File tree

‎.changeset/claude-orchestration-lineage-evidence.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -2,4 +2,4 @@
22
"agent-bundle": patch
33
---
44

5-
Record the live Claude Code 2.1.259 orchestration session (2026-09-03) on the Claude capability table's `lineage.*` rows: two `Agent` spawns issued in one message arrive serialised so each child binds to its own spawn call, an `Explore` subagent carries the same lineage fields and reaches the plugin's MCP tools, `PostToolUseFailure` inside a subagent carries the subagent's `agent_id`, the root `session_id` survives `--resume` and a manual `/compact`, and `claudecode/toolUseId` matched at depth 0, 1, and 2 (fixtures `fixtures/host-lineage/claude-2.1.259-orchestration{,.stream}.ndjson`). Emitted host output and `adapterRevision` are unchanged.
5+
Extend the Claude host capability table (`claude-2.1.250.json`, rendered on the hosts reference page) with live Claude Code 2.1.259 evidence on the `lineage.subagent-events`, `lineage.root`, `lineage.parent`, `lineage.depth`, and `lineage.mcp-correlation` rows: two `Agent` spawns issued in one message bind to their own spawn calls, `Explore` subagents carry the same lineage fields as `general-purpose` ones, `PostToolUseFailure` carries the subagent's `agent_id`, the root `session_id` survives `--resume` and `/compact`, and `claudecode/toolUseId` correlates MCP calls at depth 0, 1, and 2. Emitted host output and `adapterRevision` are unchanged. (#455)

‎examples/host-test/scripts/probe.mjs‎

Lines changed: 13 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -330,12 +330,14 @@ const captureClaude = () => {
330330
const envelopes = parseStream(result.stdout ?? '');
331331
sessionId ??= envelopes.find((envelope) => typeof envelope.session_id === 'string')?.session_id;
332332
log(`turn ${String(index + 1)}/${String(turns.length)}: exit ${String(result.status)}, session ${sessionId ?? 'unknown'}; ${describeStream(envelopes)}`);
333-
results.push({ ...result, stdoutExtension: 'stream.ndjson' });
334-
if (result.status !== 0) break;
335-
if (sessionId === undefined && index + 1 < turns.length) {
336-
log('the first turn reported no session_id, so the remaining turns cannot resume it');
333+
if (result.status === 0 && sessionId === undefined && index + 1 < turns.length) {
334+
// A multi-turn scenario without a session to resume is not the
335+
// scenario: fail the turn so `capture()` refuses the partial run.
336+
results.push({ ...result, failure: 'the first turn reported no session_id, so the remaining turns cannot resume it', stdoutExtension: 'stream.ndjson' });
337337
break;
338338
}
339+
results.push({ ...result, stdoutExtension: 'stream.ndjson' });
340+
if (result.status !== 0) break;
339341
}
340342
} finally {
341343
mock?.kill();
@@ -397,8 +399,10 @@ const capture = () => {
397399
writeFileSync(join(paths.captures, `${name}.${turn.stdoutExtension}`), turn.stdout ?? '');
398400
writeFileSync(join(paths.captures, `${name}.stderr.txt`), turn.stderr ?? '');
399401
}
400-
// The session failed if any turn did.
401-
const result = turns.find((turn) => turn.status !== 0) ?? turns[turns.length - 1];
402+
// The session failed if any turn did, or if the driver could not run the
403+
// whole scenario (`failure`).
404+
const result = turns.find((turn) => turn.status !== 0 || turn.failure !== undefined) ?? turns[turns.length - 1];
405+
const failed = result.status !== 0 || result.failure !== undefined;
402406
log(`host exit ${String(result.status)} over ${String(turns.length)} turn(s); transcript at ${join(paths.captures, `session-${stamp}.*`)}`);
403407
const appended = existsSync(logFile) ? readFileSync(logFile).subarray(startOffset).toString('utf8') : '';
404408
const records = appended.split('\n').filter(Boolean).map((line) => JSON.parse(line));
@@ -417,9 +421,9 @@ const capture = () => {
417421
// session nor a session that produced no hook AND no MCP evidence is a
418422
// capture: automation must see the failure.
419423
const missing = ['event', 'mcp'].filter((kind) => !kinds.has(kind));
420-
if (result.status !== 0) {
421-
log(`host session failed with exit ${String(result.status)}; captures above are partial evidence at best`);
422-
process.exitCode = result.status ?? 1;
424+
if (failed) {
425+
log(`host session failed (${result.failure ?? `exit ${String(result.status)}`}); captures above are partial evidence at best`);
426+
process.exitCode = result.status === 0 || result.status === null ? 1 : result.status;
423427
} else if (missing.length > 0) {
424428
log(`host session exited 0 but produced no ${missing.join(' and no ')} record; the scenario requires both`);
425429
process.exitCode = 1;

‎website/docs/en/examples/index.mdx‎

Lines changed: 6 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,10 @@
11
---
2-
description: 'The runnable agent-bundle examples: four walkthroughs from a Skills starter to a complete media-management plugin, plus two advanced composition references.'
2+
description: 'The runnable agent-bundle examples: four walkthroughs from a Skills starter to a complete media-management plugin, plus three advanced composition references.'
33
---
44

55
# Examples
66

7-
The repository ships six examples, and they are products rather than fixtures. Each one uses
7+
The repository ships seven examples, and they are products rather than fixtures. Each one uses
88
only public `agent-bundle` exports and `workspace:*` dependencies, each one builds and validates
99
through the same public CLI you would use, and none of them needs an API key or a signed-in host
1010
to reach a useful state. Four have a walkthrough here:
@@ -16,13 +16,14 @@ to reach a useful state. Four have a walkthrough here:
1616
| [MCP App](./mcp-app.mdx) | A generated stdio MCP server plus an interactive MCP App resource. | `pnpm example:mcp-app` |
1717
| [Audiobook Curator](./audiobook-curator.mdx) | A complete route-tree application: MCP, projected CLI, providers, durable state. | `pnpm example:audiobook` |
1818

19-
Two more are advanced composition references, documented by their own READMEs rather than
19+
Three more are advanced composition references, documented by their own READMEs rather than
2020
walked through here:
2121

2222
| Example | What it proves |
2323
| --- | --- |
2424
| [Worktree Proximity](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/worktree-proximity) | One root task and two child agents in linked worktrees, coordinated through durable notices. |
2525
| [RSC Agent Runtime](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/rsc-agent-runtime) | An opt-in architecture experiment: one RSC runtime shared by hooks, MCP tools, and an MCP App timeline. Not a public API. |
26+
| [Host Test](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/host-test) | A probe plugin that records every hook payload and MCP call a host delivers, with the `request.lineage` each one resolved to. Its `probe:install` / `probe:capture` / `probe:uninstall` scripts drive a real host in an isolated home; `probe:capture <host> --scenario <file.json>` runs an ordered list of prompts as one multi-turn session (Claude: `claude -p --output-format stream-json`, later turns `--resume`d), saving the model's own tool-use stream beside the hook log. The Workbench walkthrough needs no signed-in host; the probe lifecycle does. |
2627

2728
## Running them
2829

@@ -46,8 +47,8 @@ pnpm examples:check
4647

4748
Each example package also exposes the same public commands directly, once the repository-level
4849
`pnpm build` has built the local `agent-bundle` workspace dependency. Every example has `validate`,
49-
`build`, and `check`; the four walkthrough examples and Worktree Proximity also have `dev`, while the
50-
RSC Agent Runtime demo is driven through its tests and `eval:hosts` instead:
50+
`build`, and `check`; the four walkthrough examples, Worktree Proximity, and Host Test also have `dev`,
51+
while the RSC Agent Runtime demo is driven through its tests and `eval:hosts` instead:
5152

5253
```sh
5354
cd examples/skills-starter

‎website/docs/zh/examples/index.mdx‎

Lines changed: 6 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,10 @@
11
---
2-
description: '可运行的 agent-bundle 示例:从 Skill 起步项目到完整媒体管理插件的四篇导览,外加两个进阶组合参考。'
2+
description: '可运行的 agent-bundle 示例:从 Skill 起步项目到完整媒体管理插件的四篇导览,外加三个进阶组合参考。'
33
---
44

55
# 示例
66

7-
仓库中带有六个示例,它们是产品,而不是测试夹具。每个示例只使用公开的 `agent-bundle` 导出与
7+
仓库中带有七个示例,它们是产品,而不是测试夹具。每个示例只使用公开的 `agent-bundle` 导出与
88
`workspace:*` 依赖,都通过你自己也会用的那套公开命令行来构建与校验,并且都不需要 API key 或已登录的
99
宿主就能进入有意义的状态。其中四个在这里有导览:
1010

@@ -15,12 +15,13 @@ description: '可运行的 agent-bundle 示例:从 Skill 起步项目到完整
1515
| [MCP App](./mcp-app.mdx) | 一个生成的 stdio MCP 服务器,加上一个交互式 MCP App 资源。 | `pnpm example:mcp-app` |
1616
| [有声书策展器](./audiobook-curator.mdx) | 一个完整的路由树应用:MCP、投影出的 CLI、提供者与持久状态。 | `pnpm example:audiobook` |
1717

18-
另外两个是进阶组合参考,由各自的 README 说明,这里不做导览:
18+
另外三个是进阶组合参考,由各自的 README 说明,这里不做导览:
1919

2020
| 示例 | 证明什么 |
2121
| --- | --- |
2222
| [Worktree Proximity](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/worktree-proximity) | 一个根任务与两个位于关联 worktree 中的子智能体,通过持久通知协调。 |
2323
| [RSC Agent Runtime](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/rsc-agent-runtime) | 一个可选启用的架构实验:一个 RSC 运行时同时服务钩子、MCP 工具与 MCP App 时间线。不是公开 API。 |
24+
| [Host Test](https://github.com/ScriptedAlchemy/agent-bundle/tree/main/examples/host-test) | 一个探针插件,记录宿主投递的每个钩子负载与 MCP 调用,以及各自解析出的 `request.lineage`。它的 `probe:install` / `probe:capture` / `probe:uninstall` 脚本在隔离的 home 中驱动真实宿主;`probe:capture <host> --scenario <file.json>` 把一组有序提示作为一个多轮会话运行(Claude 使用 `claude -p --output-format stream-json`,后续轮次以 `--resume` 续接),并把模型自身的工具调用流保存在钩子日志旁。Workbench 导览不需要已登录的宿主;探针生命周期则需要。 |
2425

2526
## 如何运行
2627

@@ -41,8 +42,8 @@ pnpm examples:check
4142
```
4243

4344
在仓库级 `pnpm build` 构建过本地 `agent-bundle` 工作区依赖之后,每个示例包也直接暴露同样的公开命令。每个示例都有
44-
`validate`、`build` 与 `check`;四个导览示例与 Worktree Proximity 还有 `dev`,而 RSC Agent Runtime 演示则通过其测试与
45-
`eval:hosts` 驱动:
45+
`validate`、`build` 与 `check`;四个导览示例、Worktree Proximity 与 Host Test 还有 `dev`,而 RSC Agent Runtime
46+
演示则通过其测试与 `eval:hosts` 驱动:
4647

4748
```sh
4849
cd examples/skills-starter

0 commit comments

Comments
 (0)