Repository navigation
feat(server): count hybrid upstream generation with the model's tokenizer - #620
Conversation
…izer A model with local weights that an OpenAI-compatible upstream serves is now counted on the queued path as if it had run locally. The worker renders the model's chat template and counts the prompt with the model's tokenizer, and counts the completion from every text the upstream returned, private reasoning included. The terminal chunk reports that count without cached tokens and carries the upstream's own counts in usage.upstream_usage. The gateway decodes upstream_usage on the usage block for consumers of the stream outcome and never serializes it, so no response shows it. A request whose messages the local template cannot render, or whose counted prompt exceeds the context length, fails before dispatch. Without a loadable local tokenizer the upstream's counts are reported as before. Remote-backed models, sie upstreams and single-node serving are unchanged.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (16)
Included review availability: This review used your included allowance. 7 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 8 reviews per hour. 📝 WalkthroughWalkthroughQueued generations with locally backed models can use worker-tokenizer counts alongside upstream usage. The server can preserve upstream reasoning internally during counting. The gateway decodes optional upstream counts without adding them to serialized usage. ChangesQueued upstream token accounting
Sequence Diagram(s)sequenceDiagram
participant Client
participant StreamingProcessor
participant RemoteChatAdapter
participant LocalTokenizer
participant Gateway
Client->>StreamingProcessor: Submit queued generation
StreamingProcessor->>LocalTokenizer: Render and count prompt
StreamingProcessor->>RemoteChatAdapter: Request generation with reasoning preservation
RemoteChatAdapter->>StreamingProcessor: Return generated content and upstream usage
StreamingProcessor->>LocalTokenizer: Count generated content
StreamingProcessor->>Gateway: Send terminal chunk with local and upstream counts
Gateway->>Client: Return usage without upstream_usage in serialized output
Suggested reviewers: Priority: ⬇️ Low Merge Risk: ⚪ Minimal · up to This change adds worker-tokenizer token counting for locally backed models served through OpenAI-compatible upstreams. No merge-blocking risk was identified in the supplied context. 🚥 Pre-merge checks | ✅ 4 | ❓ 1❌ Failed checks (1 inconclusive)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 29.55% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 88 functions across 14 files. (2 skipped: 1 unsupported, 1 too large.) ✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
Summary
A model with local weights can also be served by an OpenAI-compatible upstream through a remote profile (a hybrid model). On the queued path its generation usage was the upstream's own count, so one request reported different counts depending on which profile served it. This change counts that generation as if it had run locally and keeps the upstream's counts beside the worker's count.
reasoning_content(orreasoning) only when the worker asks for it to count; the text is counted and dropped and never published.usage.upstream_usage(prompt_tokens,completion_tokens, optionalcached_tokens).upstream_usageintoUsageBlock::upstream_usagefor consumers of the stream outcome. The field is never serialized, so no response surface shows it.Behaviour
openaiupstream, queued chat or raw completionupstream_usagesieupstream, single-node servingA completion the upstream counted but that returned no text counts as one token, and no completion counts above the request's output bound.
Tests
tests/processors/test_hybrid_usage.py: the counting wrapper (reasoning-only chunks dropped, reasoning stripped from visible chunks, tool calls andn > 1candidates counted per choice, bound and floor, pass-through terminals, iterator close) and the queued processor end to end: streamed and buffered chat, tools rendered into the counted prompt, raw completions with and without the tool-call parser, refusals before dispatch, no tokenizer, remote-backed models, ansieupstream, and cancellation.tests/adapters/test_remote_openai_chat.py: reasoning is kept only on request and must be text.tests/processors/test_remote_chat.py: the strict-tools test now allows the OpenAI upstream to consult the tokenizer, to count; chat ownership is unchanged.queue::streaming:upstream_usagedecodes, clampscached_tokenstoprompt_tokens, rejects malformed values, and never serializes.Validation
pytestovertests/processors/, the remote adapter chat and generation suites andtests/api/test_remote_chat.py: 964 passed.ty check packages/sie_server,ruff format --checkandruff checkon the changed files: clean.cargo clippy --all-targets -- -D warnings,cargo fmt --checkandcargo test --libforsie-gateway: 1697 passed. The NATS-backed integration tests run in CI.tools/check_response_chunk_protocol.pyandtools/check_ipc_types_parity.py: OK.Review notes
Not an authentication or authorization change. The wire addition is additive: older gateways ignore
usage.upstream_usage, and a usage block without it decodes as before.Summary by CodeRabbit
New Features
Bug Fixes