Summary
This proposes that a model in the SIE catalog can be served by a remote backend. The remote backend can be the only place the model runs, or it can stand in for local capacity that is not ready.
The contributing guide asks for direction to be confirmed before a cross-package change. This issue is that request. Nothing is built yet.
Motivation
Three situations come up for people who run SIE.
- A model has no adapter yet. The model is reachable somewhere else through an OpenAI-compatible endpoint. Today an application needs a second base URL, a second key and a second client for it.
- The first request after scale-from-zero waits. The gateway answers
503 with Retry-After and X-SIE-Error-Code: PROVISIONING until a worker is ready, and the SDK waits. A model that is not loaded on a live worker answers MODEL_LOADING in the same way.
- Hardware is fixed. When local capacity is full, a request is refused. Some operators have a second SIE deployment, or a hosted endpoint, that could take the excess.
Comparable servers have added this in the last year. llama-swap has peers and selectors with warm and spillover strategies. LocalAI has failover chains over local and remote targets. GPUStack added public model providers.
Proposal
Terms
| Term |
Meaning |
| Upstream |
A named endpoint outside this SIE deployment, with a kind, a base URL and a reference to a credential |
| Upstream kind |
openai for an OpenAI-compatible endpoint. sie for another SIE deployment |
| Remote profile |
A profile whose adapter is a remote adapter. Addressable as model:remote |
| Remote-backed model |
A model with no local weights. Every profile it has is a remote profile |
| Routing policy |
How a request for the bare model name chooses between the local profile and the remote profile |
Where the call is made
The remote adapter runs in a worker. It is one more engine behind the adapter boundary. The gateway stays queue-only, holds no upstream credential and makes no upstream call. Its only new job is to choose which profile serves.
Three adapters are already OpenAI-protocol clients aimed at a local port: mlx, tensorrt_llm and the SGLang embedding adapter. A remote adapter is that client with a configured base URL and credential.
One implementation serves the single-node server and the cluster.
Configuration
An upstream is defined in deployment configuration: Helm values in a cluster, startup configuration on a single node. The model configuration API can name an upstream. It can never define one.
upstreams:
team-sie:
kind: sie
base_url: https://sie.example.internal
api_key_secret: team-sie-key
rate_cap:
requests_per_minute: 600
max_concurrency: 32
A model gains a remote profile and a routing block.
sie_id: BAAI/bge-m3
hf_id: BAAI/bge-m3
routing:
policy: fallback
fallback_profile: remote
profiles:
default:
adapter_path: <the local adapter>
remote:
adapter_path: sie_server.adapters.remote.<adapter>
adapter_options:
loadtime:
upstream: team-sie
upstream_model: BAAI/bge-m3
The public model name never encodes where the model runs. A model that is served remotely today keeps its name when it is later served locally. A request that names a profile is honoured as written. The routing policy applies only to the bare model name.
Policies
| Policy |
Behaviour |
remote_only |
The bare model name resolves to the remote profile |
fallback |
Local is tried first. When it refuses before accepting the work, the remote profile serves |
threshold |
The remote profile serves while demand is low and the lane stays asleep. Sustained demand wakes the lane |
fallback triggers:
| Trigger |
Local signal |
Default |
provisioning |
No healthy worker for the lane |
On |
model_loading |
A healthy worker exists and the model is not loaded |
On |
saturated |
Backpressure or resource exhaustion |
Opt-in |
unhealthy |
Local capacity is down |
Opt-in |
A bridged request always triggers the local warm-up. Serving remotely without waking local capacity would make the upstream the permanent server.
threshold would ship last and behind a flag.
Equivalence rule for encode and score
Under fallback and threshold, one model name is served by two backends. For encode a difference between them is permanent, because the vectors are stored. For score it changes the score scale. There is public evidence that the same weights can produce different vectors on different stacks, for example this report.
The proposal is that hybrid serving of these two primitives is refused at configuration load unless one of two conditions holds.
- Identity, for an
sie upstream. The upstream reports the same weights revision and profile identity as the local profile.
- Proof, for an
openai upstream. A passing equivalence record exists for that exact upstream and model. A probe suite produces it. It covers short and long inputs, the truncation boundary, instruction prefixes and score scale.
remote_only is not affected. It has one backend.
Caller contract
Every request served by an upstream passes through the same validation and response shaping as a local request. A field SIE does not accept is rejected, also when the upstream would accept it. Under fallback the caller cannot know which side will serve, so the accepted fields cannot depend on it.
An operator can set or strip upstream parameters per upstream. A caller cannot pass extra fields upstream.
| Header |
Direction |
Meaning |
X-SIE-Remote: forbid |
Request |
The request is never served remotely |
X-SIE-Served-By |
Response |
local or remote |
X-SIE-Upstream |
Response |
The upstream name, when served remotely |
X-SIE-Fallback-Reason |
Response |
The trigger that caused remote serving |
X-SIE-Fallback-Error |
Response |
Present when the upstream attempt failed and the local refusal is returned |
/v1/models would report the routing policy and the upstream kind for each model.
Generation
For a model SIE has onboarded, on an upstream that offers raw completions, SIE renders the prompt with its own chat template and parses tool calls itself. Otherwise the message list is forwarded at chat level. A request with a strict grammar uses the upstream's response_format, and the worker verifies the finished output.
The reason for preferring the first mode is that published comparisons of hosts serving the same weights point at templates and parsers as the main source of differences. The K2 Vendor Verifier is one example.
Egress and credentials
| Layer |
Control |
| Deployment |
Nothing leaves until an operator defines an upstream and a model declares a policy. One setting disables all remote serving |
| Model |
The routing block |
| Request |
X-SIE-Remote: forbid |
| Network |
Only the remote worker pool holds credentials. A NetworkPolicy limits its egress to the declared hosts |
| Transport |
Redirects are refused. TLS is required outside loopback. A URL that carries credentials is rejected |
| Secrets |
Given by reference. Never returned by a read API. Never logged or traced |
Failure rules
- Fallback happens only before local capacity accepts the work. No request is executed twice.
- No fallback happens after the first byte has reached the caller.
- A client error is never retried.
- Each upstream has a circuit breaker and a required rate cap.
- When the upstream also fails, or the cap is reached, the caller receives the original local refusal with its
Retry-After.
Out of scope
- Weighted or cost-based routing across several providers
- Gateway-issued keys, budgets, semantic caching and guardrails
- Provider-native dialects. An operator who needs them can place a translating gateway behind SIE as the upstream
- Caller-supplied URLs or credentials
- Storing request or response bodies
Suggested order of work
- A remote-backed embedding model on a single node through an
sie upstream. This carries the upstream definition, credential handling and the global switch
- All four primitives through an
sie upstream, with the identity check
- An
openai upstream for embeddings, rerank and generation
fallback on a single node, the request header and SDK support
- A remote worker pool in the Helm chart, and
remote_only through the queue
fallback in a cluster, with the breaker, the rate cap and the opt-in triggers
- The equivalence probe and the configuration gate
threshold, behind a flag
- Documentation, including a guide for moving from a hosted API to self-hosted one model at a time
The credential and egress handling in step 1 is a trust-boundary change and would be opened as its own small pull request for review.
Open questions
- Which profile identity can both sides of an
sie upstream compute and compare
- Whether the rate cap applies per remote worker replica or across replicas
- How the
threshold demand estimate is shared across gateway replicas
- How the gateway requests a model load without a work item, for the
model_loading trigger
- Which thresholds the equivalence probe uses. They need a measured noise floor between two local runs
Feedback wanted
- Is a remote backend the right scope, as opposed to a general gateway
- Is the worker the right place for the outbound call
- Is the equivalence rule too strict or not strict enough
- Are the default triggers and the required rate cap the right defaults
Summary
This proposes that a model in the SIE catalog can be served by a remote backend. The remote backend can be the only place the model runs, or it can stand in for local capacity that is not ready.
The contributing guide asks for direction to be confirmed before a cross-package change. This issue is that request. Nothing is built yet.
Motivation
Three situations come up for people who run SIE.
503withRetry-AfterandX-SIE-Error-Code: PROVISIONINGuntil a worker is ready, and the SDK waits. A model that is not loaded on a live worker answersMODEL_LOADINGin the same way.Comparable servers have added this in the last year. llama-swap has peers and selectors with
warmandspilloverstrategies. LocalAI has failover chains over local and remote targets. GPUStack added public model providers.Proposal
Terms
openaifor an OpenAI-compatible endpoint.siefor another SIE deploymentmodel:remoteWhere the call is made
The remote adapter runs in a worker. It is one more engine behind the adapter boundary. The gateway stays queue-only, holds no upstream credential and makes no upstream call. Its only new job is to choose which profile serves.
Three adapters are already OpenAI-protocol clients aimed at a local port:
mlx,tensorrt_llmand the SGLang embedding adapter. A remote adapter is that client with a configured base URL and credential.One implementation serves the single-node server and the cluster.
Configuration
An upstream is defined in deployment configuration: Helm values in a cluster, startup configuration on a single node. The model configuration API can name an upstream. It can never define one.
A model gains a remote profile and a routing block.
The public model name never encodes where the model runs. A model that is served remotely today keeps its name when it is later served locally. A request that names a profile is honoured as written. The routing policy applies only to the bare model name.
Policies
remote_onlyfallbackthresholdfallbacktriggers:provisioningmodel_loadingsaturatedunhealthyA bridged request always triggers the local warm-up. Serving remotely without waking local capacity would make the upstream the permanent server.
thresholdwould ship last and behind a flag.Equivalence rule for
encodeandscoreUnder
fallbackandthreshold, one model name is served by two backends. Forencodea difference between them is permanent, because the vectors are stored. Forscoreit changes the score scale. There is public evidence that the same weights can produce different vectors on different stacks, for example this report.The proposal is that hybrid serving of these two primitives is refused at configuration load unless one of two conditions holds.
sieupstream. The upstream reports the same weights revision and profile identity as the local profile.openaiupstream. A passing equivalence record exists for that exact upstream and model. A probe suite produces it. It covers short and long inputs, the truncation boundary, instruction prefixes and score scale.remote_onlyis not affected. It has one backend.Caller contract
Every request served by an upstream passes through the same validation and response shaping as a local request. A field SIE does not accept is rejected, also when the upstream would accept it. Under
fallbackthe caller cannot know which side will serve, so the accepted fields cannot depend on it.An operator can set or strip upstream parameters per upstream. A caller cannot pass extra fields upstream.
X-SIE-Remote: forbidX-SIE-Served-BylocalorremoteX-SIE-UpstreamX-SIE-Fallback-ReasonX-SIE-Fallback-Error/v1/modelswould report the routing policy and the upstream kind for each model.Generation
For a model SIE has onboarded, on an upstream that offers raw completions, SIE renders the prompt with its own chat template and parses tool calls itself. Otherwise the message list is forwarded at chat level. A request with a strict grammar uses the upstream's
response_format, and the worker verifies the finished output.The reason for preferring the first mode is that published comparisons of hosts serving the same weights point at templates and parsers as the main source of differences. The K2 Vendor Verifier is one example.
Egress and credentials
X-SIE-Remote: forbidFailure rules
Retry-After.Out of scope
Suggested order of work
sieupstream. This carries the upstream definition, credential handling and the global switchsieupstream, with the identity checkopenaiupstream for embeddings, rerank and generationfallbackon a single node, the request header and SDK supportremote_onlythrough the queuefallbackin a cluster, with the breaker, the rate cap and the opt-in triggersthreshold, behind a flagThe credential and egress handling in step 1 is a trust-boundary change and would be opened as its own small pull request for review.
Open questions
sieupstream compute and comparethresholddemand estimate is shared across gateway replicasmodel_loadingtriggerFeedback wanted