fix tool-result media loss, add count_tokens, and close the audit's remaining gaps - #44
Merged
Merged
Conversation
A tool can return an image inside `tool_result.content` — Claude Code's Read tool does for an image file. Every conversion path serialized that block array with `JSON.stringify`, so the image reached the upstream as literal base64 text: the model could not see it, and a 1MB screenshot entered the context as roughly 1.37M characters. Text and images are now split once, in `toolResultContent`, and each caller places the image where its protocol actually accepts one: - responses: inline `input_image` parts of the array `output` - chat-completions: lifted into a following user message, because a `role:"tool"` message accepts only a string - anthropic: untouched, that path already passes the request through A model whose catalog metadata says it takes no image input has the media dropped with a note instead, and a media-only result still gets non-empty text, which several upstreams require. Unknown metadata counts as capable, matching how images in ordinary user messages have always been forwarded without a capability check. Also maps an unrecognized chat `finish_reason` to `end_turn` rather than `max_tokens`, which used to tell clients a complete reply was truncated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…currency
- `POST /v1/messages/count_tokens` replaces a 404. A native Anthropic
upstream is asked for the exact count; the other protocols have no
equivalent endpoint and get a local estimate computed from the request
itself, with an image charged a flat rate rather than counted as prose.
An upstream failure falls back to the estimate — a token count must
never be why a session fails.
- `GET /v1/models/{id}` serves one model, and the catalog routes now
require the adapter token: they reveal which provider and models this
machine is configured for. `/health` stays open for supervisors.
- `--max-concurrency` / `AGENTX_MAX_CONCURRENCY` bounds in-flight
upstream requests, queueing the rest locally. Unlimited stays the
default, so nothing changes unless asked for.
- 408 joins the retryable statuses, for the same reason 504 is in it.
500 deliberately stays out — see the existing test for that decision.
- `ProviderModel.headers` is merged over the default request headers, so
a gateway can ask for its key under a name of its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`--header k=v` is repeatable on `agentx config --base-url` and on the launch commands, and persists with the provider in runtime.json. A private gateway that wants an attribution header, or its key under a name of its own, is now reachable at all — previously the request headers were fixed and such an endpoint could not be used. `agentx usage` gains `--json`, for scripts and status lines, and `--session <id>`, which reports one client session's totals. The latter finally gives `UsageStore.sessionTotals` a caller; it was implemented and tested but unreachable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`--offline` gated nothing: doctor made no network call at all, so the flag only added a line to the report. It now skips a real check. The check asks the provider's model list — which every protocol serves beside its request path — and distinguishes unreachable from reachable but rejecting the key. doctor could previously report a key as *found* but never as *accepted*, and a rejected key is the most common reason a launch fails. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Covers count_tokens and single-model lookup, the authentication boundary around the catalog routes, --max-concurrency, --header, usage --json / --session, the doctor upstream probe, and a table of where a tool result's images land per upstream protocol. Both READMEs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…eing Two problems found reviewing this branch. A tool result carrying both text and an image, sent to a model whose metadata says it takes no image input, dropped the image and said nothing — the note only appeared when the result had no text at all, which is the case where it mattered least. The note is now appended whenever media is dropped, so a text-only model is never silently missing content it was not told about. The concurrency slot was also taken before authentication, letting an unauthenticated request queue ahead of a legitimate one whenever --max-concurrency is set. Authentication now runs first, and the repeated 401 literal moved into one helper. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tzzs
added a commit
that referenced
this pull request
Sep 20, 2026
Keep PR #44's behaviour — tool-result images on every upstream protocol, count_tokens, authenticated model catalog, per-model headers, concurrency gate — while retaining this branch's typed wire payloads and single route table, so the new endpoints arrive already expressed in that shape.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
背景
对照当前 main 做了一轮审计,并查证了 OpenCode、LiteLLM、LangChain、Vercel AI SDK 和 claude-code-router 在同类问题上的实际实现。这个 PR 落地审计中结论明确的部分。
主要修复:tool_result 中的图片
工具可以在
tool_result.content里返回图片(Claude Code 的 Read 工具读图片文件时就会)。此前所有转换路径都用JSON.stringify序列化整个 block 数组,导致:现在文本和图片在
toolResultContent里拆分一次,各路径按自己协议真正接受的形态放置:anthropicresponsesfunction_call_output数组output中的input_imagechat-completionsrole:"tool"只接受字符串)这个分流思路来自 OpenCode 的
message-v2.ts,它用supportsMediaInToolResult做同样的按能力分流。顺带说明:claude-code-router(musistudio/llms)当前的行为和本 PR 之前的 agentx 一字不差 —— 同样丢弃、同样JSON.stringify。目录元数据明确表示模型不接受图片输入时,图片被丢弃并在 tool 消息中说明;只含图片的结果仍会得到非空文本(若干上游的硬性要求)。元数据未知时按"支持"处理,与普通 user 消息中的图片一直以来不做能力检查的行为保持一致。
其余改动
POST /v1/messages/count_tokens—— 此前返回 404,Claude Code 只能退回本地盲猜,表现为上下文百分比不准、auto-compact 触发时机偏移。现在原生 Anthropic 上游被请求精确计数,其余协议按请求本身本地估算(图片按固定额度计费而非当作散文),上游失败则回落到估算。GET /v1/models/{id},以及目录路由现在需要适配器 token —— 它们会暴露本机配置的 Provider 和模型。/health保持开放供进程守护轮询。--max-concurrency/AGENTX_MAX_CONCURRENCY—— 限制在途的上游请求数,超出的在本地排队。默认不限,因此不主动改变任何现有行为。自定义 Provider 请求头(
--header k=v,可重复) —— 此前请求头写死,需要归因头或非Authorization: Bearer形态的网关根本无法接入,而自定义 Provider 正是最近新增的能力。agentx usage --json/--session <id>——--session让UsageStore.sessionTotals终于有了调用方;它此前已实现并有测试,但不可达。agentx doctor的上游探测 ——--offline此前什么都没跳过,因为 doctor 根本不发网络请求。现在它会探测 Provider 的模型列表,并区分"不可达"与"可达但拒绝了 Key"。doctor 过去只能报告 Key 被找到,不能报告它被接受,而后者才是启动失败最常见的原因。重试加入 408,理由与 504 相同。500 刻意保持不重试 —— 仓库里有一条测试明确编码了这个决定,本 PR 不推翻它。
未知
finish_reason现在映射为end_turn而非max_tokens。说明:这条路径在代理流程中够不到(chatResponseFailure会先拦截),属于防御性修正,不是用户可见的 bug 修复。刻意没做的事
is_error透传:调研显示生态共识是"目标协议有该字段就设,没有就不编码"——AI SDK 和 LangChain 在 →Anthropic 方向认真设is_error,在 →OpenAI 方向一律不编码;LiteLLM 的代码注释直接承认这个信息损失。agentx 里is_error只在"目标协议无此字段"的方向丢失(原生 Anthropic 上游是直通,不受影响),且 Claude Code 已经把错误写进 content 文本,is_error对模型是冗余的。保持现状。验证
make check通过:375 项测试(原 347,新增 28 项覆盖上述每一项改动),加 npm 包校验。🤖 Generated with Claude Code