Skip to content

Proposal: remote backends, serving a catalog model from an upstream #415

Description

@krisztian-gajdar

Summary

This proposes that a model in the SIE catalog can be served by a remote backend. The remote backend can be the only place the model runs, or it can stand in for local capacity that is not ready.

The contributing guide asks for direction to be confirmed before a cross-package change. This issue is that request. Nothing is built yet.

Motivation

Three situations come up for people who run SIE.

  • A model has no adapter yet. The model is reachable somewhere else through an OpenAI-compatible endpoint. Today an application needs a second base URL, a second key and a second client for it.
  • The first request after scale-from-zero waits. The gateway answers 503 with Retry-After and X-SIE-Error-Code: PROVISIONING until a worker is ready, and the SDK waits. A model that is not loaded on a live worker answers MODEL_LOADING in the same way.
  • Hardware is fixed. When local capacity is full, a request is refused. Some operators have a second SIE deployment, or a hosted endpoint, that could take the excess.

Comparable servers have added this in the last year. llama-swap has peers and selectors with warm and spillover strategies. LocalAI has failover chains over local and remote targets. GPUStack added public model providers.

Proposal

Terms

Term Meaning
Upstream A named endpoint outside this SIE deployment, with a kind, a base URL and a reference to a credential
Upstream kind openai for an OpenAI-compatible endpoint. sie for another SIE deployment
Remote profile A profile whose adapter is a remote adapter. Addressable as model:remote
Remote-backed model A model with no local weights. Every profile it has is a remote profile
Routing policy How a request for the bare model name chooses between the local profile and the remote profile

Where the call is made

The remote adapter runs in a worker. It is one more engine behind the adapter boundary. The gateway stays queue-only, holds no upstream credential and makes no upstream call. Its only new job is to choose which profile serves.

Three adapters are already OpenAI-protocol clients aimed at a local port: mlx, tensorrt_llm and the SGLang embedding adapter. A remote adapter is that client with a configured base URL and credential.

One implementation serves the single-node server and the cluster.

Configuration

An upstream is defined in deployment configuration: Helm values in a cluster, startup configuration on a single node. The model configuration API can name an upstream. It can never define one.

upstreams:
  team-sie:
    kind: sie
    base_url: https://sie.example.internal
    api_key_secret: team-sie-key
    rate_cap:
      requests_per_minute: 600
      max_concurrency: 32

A model gains a remote profile and a routing block.

sie_id: BAAI/bge-m3
hf_id: BAAI/bge-m3
routing:
  policy: fallback
  fallback_profile: remote
profiles:
  default:
    adapter_path: <the local adapter>
  remote:
    adapter_path: sie_server.adapters.remote.<adapter>
    adapter_options:
      loadtime:
        upstream: team-sie
        upstream_model: BAAI/bge-m3

The public model name never encodes where the model runs. A model that is served remotely today keeps its name when it is later served locally. A request that names a profile is honoured as written. The routing policy applies only to the bare model name.

Policies

Policy Behaviour
remote_only The bare model name resolves to the remote profile
fallback Local is tried first. When it refuses before accepting the work, the remote profile serves
threshold The remote profile serves while demand is low and the lane stays asleep. Sustained demand wakes the lane

fallback triggers:

Trigger Local signal Default
provisioning No healthy worker for the lane On
model_loading A healthy worker exists and the model is not loaded On
saturated Backpressure or resource exhaustion Opt-in
unhealthy Local capacity is down Opt-in

A bridged request always triggers the local warm-up. Serving remotely without waking local capacity would make the upstream the permanent server.

threshold would ship last and behind a flag.

Equivalence rule for encode and score

Under fallback and threshold, one model name is served by two backends. For encode a difference between them is permanent, because the vectors are stored. For score it changes the score scale. There is public evidence that the same weights can produce different vectors on different stacks, for example this report.

The proposal is that hybrid serving of these two primitives is refused at configuration load unless one of two conditions holds.

  • Identity, for an sie upstream. The upstream reports the same weights revision and profile identity as the local profile.
  • Proof, for an openai upstream. A passing equivalence record exists for that exact upstream and model. A probe suite produces it. It covers short and long inputs, the truncation boundary, instruction prefixes and score scale.

remote_only is not affected. It has one backend.

Caller contract

Every request served by an upstream passes through the same validation and response shaping as a local request. A field SIE does not accept is rejected, also when the upstream would accept it. Under fallback the caller cannot know which side will serve, so the accepted fields cannot depend on it.

An operator can set or strip upstream parameters per upstream. A caller cannot pass extra fields upstream.

Header Direction Meaning
X-SIE-Remote: forbid Request The request is never served remotely
X-SIE-Served-By Response local or remote
X-SIE-Upstream Response The upstream name, when served remotely
X-SIE-Fallback-Reason Response The trigger that caused remote serving
X-SIE-Fallback-Error Response Present when the upstream attempt failed and the local refusal is returned

/v1/models would report the routing policy and the upstream kind for each model.

Generation

For a model SIE has onboarded, on an upstream that offers raw completions, SIE renders the prompt with its own chat template and parses tool calls itself. Otherwise the message list is forwarded at chat level. A request with a strict grammar uses the upstream's response_format, and the worker verifies the finished output.

The reason for preferring the first mode is that published comparisons of hosts serving the same weights point at templates and parsers as the main source of differences. The K2 Vendor Verifier is one example.

Egress and credentials

Layer Control
Deployment Nothing leaves until an operator defines an upstream and a model declares a policy. One setting disables all remote serving
Model The routing block
Request X-SIE-Remote: forbid
Network Only the remote worker pool holds credentials. A NetworkPolicy limits its egress to the declared hosts
Transport Redirects are refused. TLS is required outside loopback. A URL that carries credentials is rejected
Secrets Given by reference. Never returned by a read API. Never logged or traced

Failure rules

  • Fallback happens only before local capacity accepts the work. No request is executed twice.
  • No fallback happens after the first byte has reached the caller.
  • A client error is never retried.
  • Each upstream has a circuit breaker and a required rate cap.
  • When the upstream also fails, or the cap is reached, the caller receives the original local refusal with its Retry-After.

Out of scope

  • Weighted or cost-based routing across several providers
  • Gateway-issued keys, budgets, semantic caching and guardrails
  • Provider-native dialects. An operator who needs them can place a translating gateway behind SIE as the upstream
  • Caller-supplied URLs or credentials
  • Storing request or response bodies

Suggested order of work

  1. A remote-backed embedding model on a single node through an sie upstream. This carries the upstream definition, credential handling and the global switch
  2. All four primitives through an sie upstream, with the identity check
  3. An openai upstream for embeddings, rerank and generation
  4. fallback on a single node, the request header and SDK support
  5. A remote worker pool in the Helm chart, and remote_only through the queue
  6. fallback in a cluster, with the breaker, the rate cap and the opt-in triggers
  7. The equivalence probe and the configuration gate
  8. threshold, behind a flag
  9. Documentation, including a guide for moving from a hosted API to self-hosted one model at a time

The credential and egress handling in step 1 is a trust-boundary change and would be opened as its own small pull request for review.

Open questions

  • Which profile identity can both sides of an sie upstream compute and compare
  • Whether the rate cap applies per remote worker replica or across replicas
  • How the threshold demand estimate is shared across gateway replicas
  • How the gateway requests a model load without a work item, for the model_loading trigger
  • Which thresholds the equivalence probe uses. They need a measured noise floor between two local runs

Feedback wanted

  • Is a remote backend the right scope, as opposed to a general gateway
  • Is the worker the right place for the outbound call
  • Is the equivalence rule too strict or not strict enough
  • Are the default triggers and the required rate cap the right defaults

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions