Skip to content

fix tool-result media loss, add count_tokens, and close the audit's remaining gaps - #44

Merged
tzzs merged 6 commits into
mainfrom
claude/tool-optimization-features-0da45f
Sep 20, 2026
Merged

tzzs merged 6 commits into
mainfrom
claude/tool-optimization-features-0da45f

Conversation

@tzzs

@tzzs tzzs commented Sep 19, 2026

Copy link
Copy Markdown
Owner

背景

对照当前 main 做了一轮审计,并查证了 OpenCode、LiteLLM、LangChain、Vercel AI SDK 和 claude-code-router 在同类问题上的实际实现。这个 PR 落地审计中结论明确的部分。

主要修复:tool_result 中的图片

工具可以在 tool_result.content 里返回图片(Claude Code 的 Read 工具读图片文件时就会)。此前所有转换路径都用 JSON.stringify 序列化整个 block 数组,导致:

  • 图片彻底失效 —— base64 变成字面文本,上游视觉模型看不到图;
  • token 成本爆炸 —— 一张 1MB 截图约 137 万字符直接进入上下文。

现在文本和图片在 toolResultContent 里拆分一次,各路径按自己协议真正接受的形态放置:

上游协议 图片去向
anthropic 原样透传(该路径本来就是直通)
responses 内联为 function_call_output 数组 output 中的 input_image
chat-completions 抽出,以紧随其后的一条 user 消息发送(role:"tool" 只接受字符串)

这个分流思路来自 OpenCode 的 message-v2.ts,它用 supportsMediaInToolResult 做同样的按能力分流。顺带说明:claude-code-router(musistudio/llms)当前的行为和本 PR 之前的 agentx 一字不差 —— 同样丢弃、同样 JSON.stringify。

目录元数据明确表示模型不接受图片输入时,图片被丢弃并在 tool 消息中说明;只含图片的结果仍会得到非空文本(若干上游的硬性要求)。元数据未知时按"支持"处理,与普通 user 消息中的图片一直以来不做能力检查的行为保持一致。

其余改动

POST /v1/messages/count_tokens —— 此前返回 404,Claude Code 只能退回本地盲猜,表现为上下文百分比不准、auto-compact 触发时机偏移。现在原生 Anthropic 上游被请求精确计数,其余协议按请求本身本地估算(图片按固定额度计费而非当作散文),上游失败则回落到估算。

GET /v1/models/{id},以及目录路由现在需要适配器 token —— 它们会暴露本机配置的 Provider 和模型。/health 保持开放供进程守护轮询。

--max-concurrency / AGENTX_MAX_CONCURRENCY —— 限制在途的上游请求数,超出的在本地排队。默认不限,因此不主动改变任何现有行为。

自定义 Provider 请求头(--header k=v,可重复) —— 此前请求头写死,需要归因头或非 Authorization: Bearer 形态的网关根本无法接入,而自定义 Provider 正是最近新增的能力。

agentx usage --json / --session <id> —— --session 让 UsageStore.sessionTotals 终于有了调用方;它此前已实现并有测试,但不可达。

agentx doctor 的上游探测 —— --offline 此前什么都没跳过,因为 doctor 根本不发网络请求。现在它会探测 Provider 的模型列表,并区分"不可达"与"可达但拒绝了 Key"。doctor 过去只能报告 Key 被找到,不能报告它被接受,而后者才是启动失败最常见的原因。

重试加入 408,理由与 504 相同。500 刻意保持不重试 —— 仓库里有一条测试明确编码了这个决定,本 PR 不推翻它。

未知 finish_reason 现在映射为 end_turn 而非 max_tokens。说明:这条路径在代理流程中够不到(chatResponseFailure 会先拦截),属于防御性修正,不是用户可见的 bug 修复。

刻意没做的事

  • is_error 透传:调研显示生态共识是"目标协议有该字段就设,没有就不编码"——AI SDK 和 LangChain 在 →Anthropic 方向认真设 is_error,在 →OpenAI 方向一律不编码;LiteLLM 的代码注释直接承认这个信息损失。agentx 里 is_error 只在"目标协议无此字段"的方向丢失(原生 Anthropic 上游是直通,不受影响),且 Claude Code 已经把错误写进 content 文本,is_error 对模型是冗余的。保持现状。
  • 工具输出截断:OpenCode 有这一层,但它是有状态的会话持有者;agentx 是无状态代理,主动改写用户内容有争议。
  • 文档中的 G(移除客户端自动安装)与 K(跨进程排队):都需要方向性拍板,K 的既有结论是"暂不实施"。

验证

make check 通过:375 项测试(原 347,新增 28 项覆盖上述每一项改动),加 npm 包校验。

🤖 Generated with Claude Code

tzzs and others added 6 commits September 20, 2026 05:19
A tool can return an image inside `tool_result.content` — Claude Code's
Read tool does for an image file. Every conversion path serialized that
block array with `JSON.stringify`, so the image reached the upstream as
literal base64 text: the model could not see it, and a 1MB screenshot
entered the context as roughly 1.37M characters.

Text and images are now split once, in `toolResultContent`, and each
caller places the image where its protocol actually accepts one:

- responses: inline `input_image` parts of the array `output`
- chat-completions: lifted into a following user message, because a
  `role:"tool"` message accepts only a string
- anthropic: untouched, that path already passes the request through

A model whose catalog metadata says it takes no image input has the
media dropped with a note instead, and a media-only result still gets
non-empty text, which several upstreams require. Unknown metadata counts
as capable, matching how images in ordinary user messages have always
been forwarded without a capability check.

Also maps an unrecognized chat `finish_reason` to `end_turn` rather than
`max_tokens`, which used to tell clients a complete reply was truncated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…currency

- `POST /v1/messages/count_tokens` replaces a 404. A native Anthropic
  upstream is asked for the exact count; the other protocols have no
  equivalent endpoint and get a local estimate computed from the request
  itself, with an image charged a flat rate rather than counted as prose.
  An upstream failure falls back to the estimate — a token count must
  never be why a session fails.
- `GET /v1/models/{id}` serves one model, and the catalog routes now
  require the adapter token: they reveal which provider and models this
  machine is configured for. `/health` stays open for supervisors.
- `--max-concurrency` / `AGENTX_MAX_CONCURRENCY` bounds in-flight
  upstream requests, queueing the rest locally. Unlimited stays the
  default, so nothing changes unless asked for.
- 408 joins the retryable statuses, for the same reason 504 is in it.
  500 deliberately stays out — see the existing test for that decision.
- `ProviderModel.headers` is merged over the default request headers, so
  a gateway can ask for its key under a name of its own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`--header k=v` is repeatable on `agentx config --base-url` and on the
launch commands, and persists with the provider in runtime.json. A
private gateway that wants an attribution header, or its key under a
name of its own, is now reachable at all — previously the request
headers were fixed and such an endpoint could not be used.

`agentx usage` gains `--json`, for scripts and status lines, and
`--session <id>`, which reports one client session's totals. The latter
finally gives `UsageStore.sessionTotals` a caller; it was implemented
and tested but unreachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`--offline` gated nothing: doctor made no network call at all, so the
flag only added a line to the report. It now skips a real check.

The check asks the provider's model list — which every protocol serves
beside its request path — and distinguishes unreachable from reachable
but rejecting the key. doctor could previously report a key as *found*
but never as *accepted*, and a rejected key is the most common reason a
launch fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Covers count_tokens and single-model lookup, the authentication boundary
around the catalog routes, --max-concurrency, --header, usage --json /
--session, the doctor upstream probe, and a table of where a tool
result's images land per upstream protocol. Both READMEs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…eing

Two problems found reviewing this branch.

A tool result carrying both text and an image, sent to a model whose
metadata says it takes no image input, dropped the image and said
nothing — the note only appeared when the result had no text at all,
which is the case where it mattered least. The note is now appended
whenever media is dropped, so a text-only model is never silently
missing content it was not told about.

The concurrency slot was also taken before authentication, letting an
unauthenticated request queue ahead of a legitimate one whenever
--max-concurrency is set. Authentication now runs first, and the
repeated 401 literal moved into one helper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tzzs
tzzs merged commit 5f5c422 into main Sep 20, 2026
12 checks passed
tzzs added a commit that referenced this pull request Sep 20, 2026
Keep PR #44's behaviour — tool-result images on every upstream protocol,
count_tokens, authenticated model catalog, per-model headers, concurrency
gate — while retaining this branch's typed wire payloads and single route
table, so the new endpoints arrive already expressed in that shape.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant