Skip to content

feat(server): count hybrid upstream generation with the model's tokenizer - #620

Merged
krisztian-gajdar merged 1 commit into
mainfrom
feat/remote-hybrid-usage
Oct 8, 2026
Merged

krisztian-gajdar merged 1 commit into
mainfrom
feat/remote-hybrid-usage

Conversation

@krisztian-gajdar

@krisztian-gajdar krisztian-gajdar commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A model with local weights can also be served by an OpenAI-compatible upstream through a remote profile (a hybrid model). On the queued path its generation usage was the upstream's own count, so one request reported different counts depending on which profile served it. This change counts that generation as if it had run locally and keeps the upstream's counts beside the worker's count.

  • The worker renders the model's chat template and counts the prompt with the model's tokenizer, under the same special-token rule as the context-length guard.
  • It counts the completion from every text the upstream returned, private reasoning included, with the same tokenizer. The chat parser keeps an upstream's reasoning_content (or reasoning) only when the worker asks for it to count; the text is counted and dropped and never published.
  • The terminal chunk reports the worker's count, without cached tokens, and carries the upstream's counts in usage.upstream_usage (prompt_tokens, completion_tokens, optional cached_tokens).
  • The gateway decodes upstream_usage into UsageBlock::upstream_usage for consumers of the stream outcome. The field is never serialized, so no response surface shows it.

Behaviour

Case Reported usage
Hybrid model, openai upstream, queued chat or raw completion Worker count, plus upstream_usage
Messages the local template cannot render, or a counted prompt over the context length Refused before dispatch, as the local profile would refuse it
No loadable local tokenizer, or no chat template for a chat request The upstream's counts, as before
Remote-backed model, sie upstream, single-node serving Unchanged

A completion the upstream counted but that returned no text counts as one token, and no completion counts above the request's output bound.

Tests

  • tests/processors/test_hybrid_usage.py: the counting wrapper (reasoning-only chunks dropped, reasoning stripped from visible chunks, tool calls and n > 1 candidates counted per choice, bound and floor, pass-through terminals, iterator close) and the queued processor end to end: streamed and buffered chat, tools rendered into the counted prompt, raw completions with and without the tool-call parser, refusals before dispatch, no tokenizer, remote-backed models, an sie upstream, and cancellation.
  • tests/adapters/test_remote_openai_chat.py: reasoning is kept only on request and must be text.
  • tests/processors/test_remote_chat.py: the strict-tools test now allows the OpenAI upstream to consult the tokenizer, to count; chat ownership is unchanged.
  • queue::streaming: upstream_usage decodes, clamps cached_tokens to prompt_tokens, rejects malformed values, and never serializes.
  • Mutation check: 16 mutants over the wrapper, the processor plumbing, the tool-call parser propagation, the chat processor and the parser were each killed by the tests above.

Validation

  • pytest over tests/processors/, the remote adapter chat and generation suites and tests/api/test_remote_chat.py: 964 passed.
  • ty check packages/sie_server, ruff format --check and ruff check on the changed files: clean.
  • cargo clippy --all-targets -- -D warnings, cargo fmt --check and cargo test --lib for sie-gateway: 1697 passed. The NATS-backed integration tests run in CI.
  • tools/check_response_chunk_protocol.py and tools/check_ipc_types_parity.py: OK.

Review notes

Not an authentication or authorization change. The wire addition is additive: older gateways ignore usage.upstream_usage, and a usage block without it decodes as before.

Summary by CodeRabbit

  • New Features

    • For locally hosted models using an OpenAI-compatible upstream, token usage can now reflect counts from the local tokenizer, including tool-call content. Upstream counts remain available separately.
    • Private reasoning can be retained internally when enabled, while remaining excluded from caller-visible output.
  • Bug Fixes

    • Requests that cannot be counted locally are rejected before dispatch; when local counting is unavailable, usage falls back to upstream counts.
    • Token counts are capped to valid prompt and completion limits.

…izer

A model with local weights that an OpenAI-compatible upstream serves is now
counted on the queued path as if it had run locally. The worker renders the
model's chat template and counts the prompt with the model's tokenizer, and
counts the completion from every text the upstream returned, private reasoning
included. The terminal chunk reports that count without cached tokens and
carries the upstream's own counts in usage.upstream_usage.

The gateway decodes upstream_usage on the usage block for consumers of the
stream outcome and never serializes it, so no response shows it. A request
whose messages the local template cannot render, or whose counted prompt
exceeds the context length, fails before dispatch. Without a loadable local
tokenizer the upstream's counts are reported as before. Remote-backed models,
sie upstreams and single-node serving are unchanged.
@krisztian-gajdar
krisztian-gajdar requested a review from a team as a code owner October 7, 2026 21:26
@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Team
  • Run ID: 0d967ac9-0aa2-4d21-8e12-8b373348c8dd
📥 Commits

Reviewing files that changed from the base of the PR and between 6163120 and 7f7f0ae.

📒 Files selected for processing (16)
  • packages/sie_gateway/src/handlers/proxy.rs
  • packages/sie_gateway/src/handlers/sse.rs
  • packages/sie_gateway/src/queue/streaming.rs
  • packages/sie_server/REMOTE_BACKENDS.md
  • packages/sie_server/src/sie_server/adapters/_generation_base.py
  • packages/sie_server/src/sie_server/adapters/remote/_chat_transport.py
  • packages/sie_server/src/sie_server/adapters/remote/_openai_chat.py
  • packages/sie_server/src/sie_server/adapters/remote/openai.py
  • packages/sie_server/src/sie_server/adapters/remote/sie.py
  • packages/sie_server/src/sie_server/processors/hybrid_usage.py
  • packages/sie_server/src/sie_server/processors/remote_chat.py
  • packages/sie_server/src/sie_server/processors/streaming.py
  • packages/sie_server/src/sie_server/processors/tool_call_parser.py
  • packages/sie_server/tests/adapters/test_remote_openai_chat.py
  • packages/sie_server/tests/processors/test_hybrid_usage.py
  • packages/sie_server/tests/processors/test_remote_chat.py

Included review availability: This review used your included allowance. 7 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 8 reviews per hour.


📝 Walkthrough

Walkthrough

Queued generations with locally backed models can use worker-tokenizer counts alongside upstream usage. The server can preserve upstream reasoning internally during counting. The gateway decodes optional upstream counts without adding them to serialized usage.

Changes

Queued upstream token accounting

Layer / File(s) Summary
Upstream usage and reasoning data
packages/sie_server/src/sie_server/adapters/_generation_base.py, packages/sie_server/src/sie_server/adapters/remote/*, packages/sie_server/src/sie_server/processors/remote_chat.py, packages/sie_server/tests/adapters/test_remote_openai_chat.py
Generation chunks now carry optional upstream token usage and reasoning deltas. Remote chat adapters can preserve reasoning when requested, and remote chat processing emits reasoning deltas for buffered and streaming responses.
Local hybrid counting and generation wiring
packages/sie_server/src/sie_server/processors/hybrid_usage.py, packages/sie_server/src/sie_server/processors/streaming.py, packages/sie_server/src/sie_server/processors/tool_call_parser.py, packages/sie_server/tests/processors/*, packages/sie_server/REMOTE_BACKENDS.md
The processor locally renders and counts prompts for eligible upstream requests, then combines worker counts with upstream usage. Hybrid counting handles generated text, tool calls, and reasoning. Tests cover counting, fallback, validation, and cancellation paths.
Gateway usage decoding and fixtures
packages/sie_gateway/src/queue/streaming.rs, packages/sie_gateway/src/handlers/proxy.rs, packages/sie_gateway/src/handlers/sse.rs
The gateway decodes optional upstream counts, clamps cached tokens to upstream prompt counts, and omits upstream usage from serialized usage. Test fixtures initialize the optional field.

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant StreamingProcessor
  participant RemoteChatAdapter
  participant LocalTokenizer
  participant Gateway
  Client->>StreamingProcessor: Submit queued generation
  StreamingProcessor->>LocalTokenizer: Render and count prompt
  StreamingProcessor->>RemoteChatAdapter: Request generation with reasoning preservation
  RemoteChatAdapter->>StreamingProcessor: Return generated content and upstream usage
  StreamingProcessor->>LocalTokenizer: Count generated content
  StreamingProcessor->>Gateway: Send terminal chunk with local and upstream counts
  Gateway->>Client: Return usage without upstream_usage in serialized output
Loading

Suggested reviewers: huronat

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to 7f7f0

This change adds worker-tokenizer token counting for locally backed models served through OpenAI-compatible upstreams. No merge-blocking risk was identified in the supplied context.

🚥 Pre-merge checks | ✅ 4 | ❓ 1

❌ Failed checks (1 inconclusive)

Check name Status Explanation Resolution
Docstring Coverage ❓ Inconclusive Docstring coverage is 29.55% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 88 functions across 14 files. (2 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: counting hybrid upstream generation with the model tokenizer.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 29.55% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 88 functions across 14 files. (2 skipped: 1 unsupported, 1 too large.)

✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@krisztian-gajdar
krisztian-gajdar merged commit 2f9c107 into main Oct 8, 2026
21 checks passed
@krisztian-gajdar
krisztian-gajdar deleted the feat/remote-hybrid-usage branch October 8, 2026 09:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant