简单用法:改一个配置文件 → 双击重启 → ZCode 填本地地址。 Windows / macOS 中文说明
npm install -g @sevoniva/llm-coding-bridge@latest
llm-coding-bridge init-files
这会生成 config.json 和启动、重启、停止、自启动管理脚本。只需填写配置里的 baseUrl、model、apiKey,Windows 双击 restart.cmd(macOS 双击 restart.command)即可加载配置、启动服务并启用登录自启动。默认在 ZCode 填 http://127.0.0.1:37629/v1,模型填配置中的 model,API Key 填 local。
以后修改配置,保存后再次双击重启即可。向导 setup、多模型和系统密钥存储仍可选用,见进阶配置流程。
A local, production-oriented protocol bridge for coding clients that need one stable OpenAI-compatible endpoint in front of one or more upstream models.
The bridge keeps client configuration, model aliases, upstream model IDs, credentials, streaming behavior, and route health in one controlled local service. It does not modify client source code. The simple file workflow accepts upstream.apiKey in a private local config; the optional setup wizard stores keys in the system credential store instead.
Codex / Claude / OpenAI-compatible client
|
v
http://127.0.0.1:37629/v1
|
LLM Coding Bridge
- route by model alias
- normalize protocols
- enforce timeouts and limits
- emit streaming heartbeats
|
v
OpenAI-compatible upstream
Supported client surfaces:
| Client or protocol | Bridge endpoint | Primary use |
|---|---|---|
| Codex CLI / Codex Desktop | /v1/responses |
Responses API |
| Claude-compatible clients | /v1/messages |
Anthropic-compatible messages |
| OpenAI-compatible clients | /v1/chat/completions |
Chat Completions |
| Client discovery | /v1/models, /health |
Model and service checks |
The default listener is loopback-only. The service has no runtime dependencies and can be managed by macOS launchd or Windows Task Scheduler.
Requires Node.js 18 or newer.
npm install -g @sevoniva/llm-coding-bridge@latest
llm-coding-bridge --helpnpm install -g @sevoniva/llm-coding-bridge@latest
llm-coding-bridge init-files
Open the printed folder (default ~/.llm-coding-bridge, or %USERPROFILE%\.llm-coding-bridge on Windows). Edit config.json: replace upstream.baseUrl, upstream.model, and upstream.apiKey. Double-click restart.cmd on Windows or restart.command on macOS. The script reloads that exact file, enables autostart after login, waits for the new configuration to become healthy, and shows the client connection details.
For ZCode, use http://127.0.0.1:37629/v1, the configured upstream model ID, and local as the local API key. After changing the file, double-click restart again. Existing config.json is preserved when init-files is rerun. Use init-files --out <directory> to choose a visible folder. The file contains your upstream API key, so keep it local.
The generated folder includes start, restart, stop, status, and disable-autostart scripts. Stop is temporary; disable-autostart stops and removes the login task. On package or Node upgrades, rerun init-files in the same directory to refresh script paths, then double-click restart.
Optional guided setup with system-managed credentials and automatic client configuration: llm-coding-bridge setup. See the Chinese quick start, guided setup, or complete configuration reference.
DeepSeek Harness is supported as a custom provider. No Harness source change is required, and the built-in deepseek-official provider remains available alongside the bridge route.
Create a separate openai-completions provider with:
baseURL: http://127.0.0.1:37629/v1- the bridge model alias as the selected model
- a dedicated credential reference for this route
supportsReasoningEffort: truewhen the selected upstream acceptsreasoning_effort
Example provider shape:
agent-default-model:
provider: local-bridge
model: your-model
llm-pi-ai:
providers:
local-bridge:
displayName: Local Bridge
api: openai-completions
baseURL: http://127.0.0.1:37629/v1
apiKeyEnv: LOCAL_BRIDGE_API_KEY
compat:
thinkingFormat: openai
supportsReasoningEffort: true
models:
- id: your-modelThe bridge applies client-specific transport handling at the boundary:
- preserves empty assistant
contentin tool-call history; - forwards only the documented Harness identity and attribution headers;
- emits an SSE comment plus an empty delta heartbeat during idle periods;
- waits for successful upstream SSE headers so upstream HTTP errors remain HTTP errors;
- normalizes embedded error envelopes that report a real 4xx/5xx status inside an HTTP 200 response.
For upstream gateways that reject the top-level thinking extension or impose an output limit, configure the route:
{
"upstream": {
"translateThinkingToReasoningEffort": true,
"maxOutputTokens": 131072
}
}translateThinkingToReasoningEffort removes the unsupported top-level field, preserves an explicit client reasoning_effort, and maps disabled thinking to reasoning_effort: "none". maxOutputTokens caps numeric max_tokens and max_completion_tokens. In version 2 configuration, set either option on a provider or override it on an individual model.
A minimal configuration uses one upstream:
{
"server": {
"host": "127.0.0.1",
"port": 37629
},
"upstream": {
"name": "Custom Provider",
"baseUrl": "https://api.example.com/v1",
"model": "model-name",
"apiKeyEnv": "LLM_API_KEY",
"temperature": 0
}
}For multiple routes, use upstreams. Requests are matched by the client model field; an unknown model returns 404 model_not_found instead of silently selecting a different route.
{
"server": { "host": "127.0.0.1", "port": 37629 },
"upstreams": [
{
"name": "Fast",
"baseUrl": "https://fast.example.com/v1",
"model": "coding-fast",
"apiKeyEnv": "FAST_API_KEY"
},
{
"name": "Strong",
"baseUrl": "https://strong.example.com/v1",
"model": "coding-strong",
"apiKeyEnv": "STRONG_API_KEY"
}
]
}Version 2 separates the client-facing alias, upstream model ID, provider endpoint, and credential source:
{
"version": 2,
"providers": [
{
"id": "local-provider",
"name": "Local Provider",
"baseUrl": "https://api.example.com/v1",
"translateThinkingToReasoningEffort": true,
"maxOutputTokens": 131072,
"models": [
{
"alias": "coding-strong",
"upstreamModel": "provider-model-id",
"credentialRef": "coding-strong-key"
}
]
}
],
"credentials": {
"coding-strong-key": {
"source": "env",
"env": "LLM_API_KEY"
}
}
}Use llm-coding-bridge config migrate --dry-run before migrating an existing version 1 file. Use llm-coding-bridge config show --effective to inspect the resolved route without printing credential values.
The bridge supports:
apiKeyEnv: read a key from an environment variable;apiKeyCommand: resolve a key from a command or secret manager;apiKeySource: "client": forward the key supplied by a local provider switcher or client.
For client-provided keys, the bridge checks x-upstream-api-key, then Authorization: Bearer ..., then x-api-key. If local authentication is enabled, put the local token in Authorization and the upstream key in x-upstream-api-key.
Keys resolved from apiKeyCommand are cached in memory for 10 minutes by default. Configure upstream.apiKeyCacheTtlMs or set it to 0 to disable caching. A 401 response invalidates the cached key immediately.
The bridge is deliberately conservative about retries: it retries only before semantic output has been observed. Once text, reasoning, refusal, function calls, tool calls, or audio has been emitted, the request is not replayed.
It also provides:
- independent header, first-data, idle, total, and streaming deadlines;
- bounded
Retry-Afterhandling and per-route cooldown; - optional FIFO concurrency limits per provider, scoped to each upstream attempt so stalled attempts yield to queued work before retrying;
- optional per-provider start pacing so rolling request-rate limits are respected, including retries;
- protocol-specific Responses, Chat Completions, and Anthropic-compatible streaming;
- SSE heartbeats that do not count as upstream model output;
- conversion of valid non-SSE JSON responses to SSE for streaming clients;
- aggregation of complete SSE responses for non-streaming Chat Completions;
- request-body, response-body, and SSE-event size limits;
- backpressure-aware streaming and connection cleanup.
For Chat Completions, server.heartbeatIntervalMs defaults to 15 seconds. Set it to 0 to disable downstream heartbeats. Set server.maxConcurrentRequestsPerProvider to a positive integer to queue excess upstream work per provider; 0 keeps concurrency unlimited. Set server.minRequestIntervalMsPerProvider to pace every upstream attempt for a provider; 0 disables pacing.
Generate a dedicated profile:
llm-coding-bridge codex-profile --name bridge
codex --profile bridge exec --skip-git-repo-check "Reply exactly: OK"Or print the template:
llm-coding-bridge template codexThe generated profile uses:
base_url = http://127.0.0.1:37629/v1
wire_api = responses
The bridge exposes /v1/messages and /v1/messages/count_tokens:
export ANTHROPIC_BASE_URL="http://127.0.0.1:37629"
export ANTHROPIC_AUTH_TOKEN="local"
export ANTHROPIC_DEFAULT_SONNET_MODEL="your-model"
export ANTHROPIC_DEFAULT_OPUS_MODEL="your-model"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="your-model"Print the generated template with:
llm-coding-bridge template claudeUse:
Base URL: http://127.0.0.1:37629/v1
Endpoint: /chat/completions
Model: one of the aliases returned by /v1/models
llm-coding-bridge install-service installs a per-user launchd agent on macOS or a Task Scheduler logon task on Windows, and starts it immediately. Both start after the current user logs in. Windows runs without a persistent console window, appends logs, and retries failures every minute up to 999 times. This is user-session autostart, not a pre-login system service.
Use service-status to inspect the registered task even when the bridge is down. stop-service temporarily stops it; uninstall-service stops it and removes autostart while keeping configuration and credentials. restart-service reuses the installed absolute config path unless --config is supplied, and refreshes Node/npm paths after upgrades.
llm-coding-bridge install-service
launchctl list | grep llm-coding-bridge
llm-coding-bridge status
llm-coding-bridge logs --lines 80
llm-coding-bridge restart-service
llm-coding-bridge uninstall-serviceThe service cannot keep an open request alive while macOS is asleep. After wake, clients should retry according to their normal transport policy.
The default loopback listener does not require a token. For a shared or non-loopback listener, set server.localToken:
{
"server": {
"host": "127.0.0.1",
"port": 37629,
"localToken": "replace-with-a-random-secret"
}
}Clients can send Authorization: Bearer <token> or x-api-key: <token>. /health remains unauthenticated. Non-loopback listeners are rejected without a local token.
Security defaults:
- generated config, client profiles, and backups use private file permissions;
- request bodies are capped at 10 MB by default;
- upstream responses are subject to complete-response and cumulative-size limits;
- diagnostic logs contain validated request metadata only; fallback logging never serializes upstream error objects;
- API keys are never committed by the project and should not be placed in the bridge config;
- command-backed keys should use the object form to avoid shell interpretation;
- the package has no install script and does not execute network code during installation.
| Symptom | Check |
|---|---|
ECONNREFUSED 127.0.0.1:37629 |
Run llm-coding-bridge status, then llm-coding-bridge restart-service. |
404 model_not_found |
Compare the client model with curl http://127.0.0.1:37629/v1/models. |
401 from the upstream |
Re-check the selected credential source; command-backed keys are refreshed after a 401. |
| Harness receives a stream error as success | Confirm the custom provider uses the bridge baseURL and keep the built-in provider separate. |
| Requests stall during long thinking | Keep heartbeats enabled and set maxOutputTokens to the tested upstream limit. |
| Works in a shell but not after login | Use apiKeyCommand or a launchd-visible credential source instead of a shell-only environment variable. |
git clone https://github.com/sevoniva/llm-coding-bridge.git
cd llm-coding-bridge
npm ci
npm run verifyThe verification gate runs linting, the complete test suite, security and secret scans, the repository and release gates, dependency audit, and an npm pack dry run.
Further reference: Configuration Guide, release notes, and MIT License.
@sevoniva/llm-coding-bridge 是一个本地协议桥接服务,把 Codex、Claude 类客户端、DeepSeek Harness 和其他 OpenAI-compatible 客户端统一接到稳定的本地 /v1 端点,再按模型别名转发到一个或多个上游。它不需要修改客户端源码,默认只监听本机回环地址。简单配置支持在私有本地文件中填写 upstream.apiKey;可选向导则使用系统密钥存储。
安装并生成配置文件:
npm install -g @sevoniva/llm-coding-bridge@latest
llm-coding-bridge init-filesHarness 配置重点:
- 新增独立的
openai-completions自定义 Provider; baseURL使用http://127.0.0.1:37629/v1;- 选择 bridge 暴露的模型别名;
- 使用独立凭据引用;
- 保留内置的
deepseek-officialProvider,不需要修改 Harness 源码。
如果上游不接受顶层 thinking 或有输出上限,在 bridge 路由中配置 translateThinkingToReasoningEffort 和 maxOutputTokens。完整配置、协议行为和安全边界见 Configuration Guide。
Windows / macOS 自启动:
llm-coding-bridge install-service
llm-coding-bridge status
llm-coding-bridge logs --lines 80Windows 使用任务计划程序,macOS 使用 launchd:开机登录当前账户后启动,异常退出自动重试。详细步骤、配置路径和故障处理见中文安装配置与自启动指南。
MIT