Repository navigation
Proposal: remote backends, serving a catalog model from an upstream #415
Description
Activity
- added 13 commits that reference this issue
on Sep 30, 2026 krisztian-gajdar commented
on Oct 3, 2026 ContributorAuthorMore actionsImplementation audit and continuation, 2026-10-04:
The proposal's “nothing is built yet” description is now historical. Public
mainhas upstream configuration and credential/egress controls, remote SIE encode/score/extract (including sparse/multivector and supported image inputs), OpenAI-compatible dense embeddings and reranking, remote profile/routing validation, request-level remote forbidding and disclosure, the remote worker bundle/Helm support, and per-worker upstream limits/circuit breaking. Single-node extraction and generation fallback are implemented. Exact fresh OpenAI and immutable SIE hybrid encode/score admission are available for one concrete worker; fleet admission and threshold policies remain guarded/refused.Merged in this continuation:
-
feat(server): stream generation through a remote SIE upstream #514: native remote SIE buffered/streaming generation, with bounded stream parsing, terminal usage validation and cancellation/client cleanup fixes.
-
feat(server): serve native SIE chat through remote profiles #523: single-node SIE-upstream chat serving, shared response validation, strict-output checks, disclosure and stream cleanup. Independent review found and fixed private reasoning leaking through associated logprobs.
-
docs(server): explain remote backends and migration to self-hosted models #522: remote-backend operator guide and hosted-to-self-hosted migration guide, documenting supported behavior and current limits.
-
feat(queue): support load-only work and reply to backend fallback refusals #525: worker-side load-only items, backend retry outcomes answered as fallback refusals, and retry hints carried through Python IPC/sidecar/gateway. Load-only items are NATS-only; local ingest rejects them because ACK emits no local completion. This does not activate gateway fallback routing or implement all refusal paths.
-
fix(server): suppress private chat reasoning logprobs independently of thinking policy #527: suppress private reasoning token logprobs in the local chat sibling, including when thinking configuration is absent. This closes the local sibling found by the required security sweep for feat(server): serve native SIE chat through remote profiles #523.
-
feat(server): serve OpenAI upstream chat and raw generation #528: OpenAI upstream single-node chat and native raw generation, shared bounded chat transport, pinned endpoints/token limits, exact usage and cancellation/privacy coverage.
-
feat(server): serve remote chat through queued generation #529: queued SIE/OpenAI chat without local weights/tokenizers; onboarded queued raw-template preference, tools, multiple choices, strict JSON Schema verification and exact usage. Security review found two fragmented-tool/terminal-ordering bugs, both fixed and covered by regressions.
-
feat(server): expose conservative immutable profile identity #530: conservative immutable local-profile identity metadata for native BGE-M3, with pinned tokenizer/source/runtime settings. Unidentified engines and checkpoint-selected custom code remain refused. Hybrid admission is still closed.
-
feat(server): measure remote equivalence against local noise #531: versioned equivalence records and SDK-based probe with two measured local runs, complete numerical/layout checks, and before/after server-bound endpoint/model contract verification. Records do not activate hybrid admission.
-
fix(server): refuse remote generation before streaming headers #532: remote native generation streams are primed before HTTP success, retaining typed pre-output refusals and closing primed iterators on cancellation/disconnect.
-
fix(server): honor explicitly selected default profiles #534: explicitly selecting the default profile stays local; native generation forwards the caller selector to the shared router.
-
feat(server): bridge cold single-node generation before output #533: single-node generation fallback across native generation, chat, completions and Responses, with early validation, bounded pre-output attempts, original refusal/retry restoration and worker/config-service admission parity.
-
fix(remote): bind profile identity to hardware and BLAS #536: local profile identity is bound to observed CPU/GPU hardware, kernel/driver, numerical library builds and actual BLAS kernel/thread selection. Unknown observations remain closed.
-
feat(sdk): accept origin-confined configured HTTP clients #539: Python SIEClient can adopt a configured synchronous HTTP client while preserving credential/auth/transport policy, with same-origin confinement and SDK ownership of cleanup.
-
feat(remote): admit exact fresh OpenAI equivalence evidence #535: exact fresh version-2 OpenAI equivalence admission for one concrete worker device, including process binding, per-bridge freshness/default checks and startup-owned Helm policy rendering.
-
feat(server): preserve onboarded templates for direct remote chat #537: single-node onboarded raw-template chat preference, pinned local tokenizers without loading model weights, tools/strict-output checks, and native SIE reasoning initialization across receiving surfaces.
-
feat(remote): admit fresh matching SIE profile identity #540: fresh bounded SIE profile/weights identity admission for one concrete worker, plus per-socket absolute deadlines and bounded asynchronous response headers. TLS proxy trust behavior and metadata confinement were independently reviewed.
-
feat(gateway): publish load-only model readiness work #541: NATS gateway load-only producer, broker durability and transport refusal, without inference inputs/result collectors. This is the producer prerequisite; cluster routing activation is next.
-
feat(gateway): preserve validated remote routing policies #542: validated gateway routing policies, atomic snapshot/delta handling and explicit-profile policy clearing. Threshold remains refused until its coordinator is implemented.
All twenty-seven merged PRs were approved by CodeRabbit on their submitted commits and merged after green CI. Independent adversarial review covered transport, queue settlement, reasoning privacy, hardware identity and evidence authority. Transient Rust runner/tool-download failures were retried; no red check was bypassed.
-
fix(worker): enforce live configuration authority through execution #543: live worker execution authority, including Python configuration leases, hashed model/profile membership, grammar routing, selected batch operation payloads, and every sidecar dispatch barrier. Merged after exact-head CodeRabbit approval and green CI.
-
feat(worker): fence verified inference with versioned IPC methods #544: updated-only backend IPC entrypoints and default-false capability, merged at 5f832a4 after exact-head CodeRabbit approval and all 46 checks green.
-
feat(worker): fence verified dispatch with versioned queues #545: versioned direct queue/stream, current authority checks, exact subject/payload model matching, all-child capability aggregation and conservative health advertisement. CodeRabbit's two availability findings were fixed and threads resolved; final exact-head approval and green CI preceded merge at 2ba3d6e. Full post-rebase sidecar suites passed 652 default/670 cloud library tests, 14 binary and 30 integration tests each.
-
feat(gateway): enforce remote-forbid with verified worker dispatch #546: gateway X-SIE-Remote: forbid across inference ingress, fresh positive worker capability, exact current hash, versioned direct dispatch and no legacy retries. CodeRabbit's balancing finding and independent review's pressure-overflow finding were fixed; all checks green, exact-head approval and the resolved thread preceded merge at 19be847 (2026-10-03 23:46 UTC).
-
feat(gateway): bridge buffered cluster generation refusals #547: buffered cluster generation fallback across native generation, chat, completions and supported Responses. Cold local demand is retained; unloaded local workers receive durably accepted load-only work before one exact-authority remote attempt. Explicit caller selectors remain authoritative. Full gateway suites, post-rebase conformance, both Clippy configurations and independent adversarial review passed. Exact-head CodeRabbit approval and all CI green preceded merge at 6ec8dec (2026-10-04 00:06 UTC).
-
feat(gateway): restore cluster stream refusals before output #548: streaming cluster generation fallback on native generation, chat and completions. HTTP success waits for the first valid event; pre-output errors and cancelled/failed terminals restore the original local refusal, while later errors stay in-stream without replay. Full gateway suites passed 1591 default/1602 cloud library and binary tests plus NATS integrations. Both Clippy configurations and independent review passed. After parent merge/rebase, fresh exact-head CodeRabbit approval and all CI green preceded merge at 0ebc55c (2026-10-04 00:23 UTC).
The extraction/audio layer was delivered in #549. It shares local warm-up and exact worker admission, validates cheap caller-controlled input boundaries before demand/dispatch, preserves display-model names and fallback/retry headers, and retains the local refusal until the requested audio format is ready. Independent review findings on audio formatting, metadata sizing and typed input limits were fixed, with source clearance and focused regressions.
Remaining deliveries, in dependency order:
- Complete cluster fallback: gateway profile choice, local warm-up/load-item production, bounded remote attempts across all pre-acceptance refusal paths, original-local-refusal restoration, opt-in triggers, fleet equivalence admission, request forbid control, telemetry and conformance tests.
- Flagged threshold policy with coordinated demand/wake behavior across gateway replicas.
- Recorded cold-start and equivalence evidence, with the required product review. No performance/equivalence claims should be made before those measurements.
The issue remains open. The new guides explain migration checks without claiming that model names or matching dimensions prove equivalence.
-
feat(gateway): bridge cluster extraction and audio refusals #549: cluster native extraction and audio transcription cold/loading bridges, bounded worker-compatible input validation, exact metadata size accounting, requested-model preservation, and original-refusal restoration through audio response formatting. Merged at 2026-10-04 00:46:11 UTC as
b009e9f40c62a20a6f3a36ec0e5c5a4db95d1467, with CodeRabbit approval on the exact submitted head and green CI. -
feat(gateway): opt in to pre-acceptance remote spill #550: opt-in saturation/unhealthy fallback across buffered and streaming surfaces, scoped refusal evidence and pure pre-dispatch pressure checks. The unhealthy retry-hint finding was fixed. Exact-head CodeRabbit approval and all CI green preceded merge at
4f3a62efbc7ea05023afa8adf37815447dc9f5d1(2026-10-04 01:31:12 UTC).
Current continuation: #551 adds bounded fallback telemetry, a persistent-remote alert/dashboard, idle-period reset and valid-local-output reset. Full default/cloud-storage gateway validation, both Clippy configurations and independent adversarial review passed; CodeRabbit/CI are running. Coordinated threshold implementation is next. Numerical fleet admission and actual acceptance measurements remain open; issue #415 is not complete.
-
- added 15 commits that reference this issue
on Oct 3, 2026
Summary
This proposes that a model in the SIE catalog can be served by a remote backend. The remote backend can be the only place the model runs, or it can stand in for local capacity that is not ready.
The contributing guide asks for direction to be confirmed before a cross-package change. This issue is that request. Nothing is built yet.
Motivation
Three situations come up for people who run SIE.
503withRetry-AfterandX-SIE-Error-Code: PROVISIONINGuntil a worker is ready, and the SDK waits. A model that is not loaded on a live worker answersMODEL_LOADINGin the same way.Comparable servers have added this in the last year. llama-swap has peers and selectors with
warmandspilloverstrategies. LocalAI has failover chains over local and remote targets. GPUStack added public model providers.Proposal
Terms
openaifor an OpenAI-compatible endpoint.siefor another SIE deploymentmodel:remoteWhere the call is made
The remote adapter runs in a worker. It is one more engine behind the adapter boundary. The gateway stays queue-only, holds no upstream credential and makes no upstream call. Its only new job is to choose which profile serves.
Three adapters are already OpenAI-protocol clients aimed at a local port:
mlx,tensorrt_llmand the SGLang embedding adapter. A remote adapter is that client with a configured base URL and credential.One implementation serves the single-node server and the cluster.
Configuration
An upstream is defined in deployment configuration: Helm values in a cluster, startup configuration on a single node. The model configuration API can name an upstream. It can never define one.
A model gains a remote profile and a routing block.
The public model name never encodes where the model runs. A model that is served remotely today keeps its name when it is later served locally. A request that names a profile is honoured as written. The routing policy applies only to the bare model name.
Policies
remote_onlyfallbackthresholdfallbacktriggers:provisioningmodel_loadingsaturatedunhealthyA bridged request always triggers the local warm-up. Serving remotely without waking local capacity would make the upstream the permanent server.
thresholdwould ship last and behind a flag.Equivalence rule for
encodeandscoreUnder
fallbackandthreshold, one model name is served by two backends. Forencodea difference between them is permanent, because the vectors are stored. Forscoreit changes the score scale. There is public evidence that the same weights can produce different vectors on different stacks, for example this report.The proposal is that hybrid serving of these two primitives is refused at configuration load unless one of two conditions holds.
sieupstream. The upstream reports the same weights revision and profile identity as the local profile.openaiupstream. A passing equivalence record exists for that exact upstream and model. A probe suite produces it. It covers short and long inputs, the truncation boundary, instruction prefixes and score scale.remote_onlyis not affected. It has one backend.Caller contract
Every request served by an upstream passes through the same validation and response shaping as a local request. A field SIE does not accept is rejected, also when the upstream would accept it. Under
fallbackthe caller cannot know which side will serve, so the accepted fields cannot depend on it.An operator can set or strip upstream parameters per upstream. A caller cannot pass extra fields upstream.
X-SIE-Remote: forbidX-SIE-Served-BylocalorremoteX-SIE-UpstreamX-SIE-Fallback-ReasonX-SIE-Fallback-Error/v1/modelswould report the routing policy and the upstream kind for each model.Generation
For a model SIE has onboarded, on an upstream that offers raw completions, SIE renders the prompt with its own chat template and parses tool calls itself. Otherwise the message list is forwarded at chat level. A request with a strict grammar uses the upstream's
response_format, and the worker verifies the finished output.The reason for preferring the first mode is that published comparisons of hosts serving the same weights point at templates and parsers as the main source of differences. The K2 Vendor Verifier is one example.
Egress and credentials
X-SIE-Remote: forbidFailure rules
Retry-After.Out of scope
Suggested order of work
sieupstream. This carries the upstream definition, credential handling and the global switchsieupstream, with the identity checkopenaiupstream for embeddings, rerank and generationfallbackon a single node, the request header and SDK supportremote_onlythrough the queuefallbackin a cluster, with the breaker, the rate cap and the opt-in triggersthreshold, behind a flagThe credential and egress handling in step 1 is a trust-boundary change and would be opened as its own small pull request for review.
Open questions
sieupstream compute and comparethresholddemand estimate is shared across gateway replicasmodel_loadingtriggerFeedback wanted