Skip to content

Proposal: remote backends, serving a catalog model from an upstream #415

Description

@krisztian-gajdar

Summary

This proposes that a model in the SIE catalog can be served by a remote backend. The remote backend can be the only place the model runs, or it can stand in for local capacity that is not ready.

The contributing guide asks for direction to be confirmed before a cross-package change. This issue is that request. Nothing is built yet.

Motivation

Three situations come up for people who run SIE.

  • A model has no adapter yet. The model is reachable somewhere else through an OpenAI-compatible endpoint. Today an application needs a second base URL, a second key and a second client for it.
  • The first request after scale-from-zero waits. The gateway answers 503 with Retry-After and X-SIE-Error-Code: PROVISIONING until a worker is ready, and the SDK waits. A model that is not loaded on a live worker answers MODEL_LOADING in the same way.
  • Hardware is fixed. When local capacity is full, a request is refused. Some operators have a second SIE deployment, or a hosted endpoint, that could take the excess.

Comparable servers have added this in the last year. llama-swap has peers and selectors with warm and spillover strategies. LocalAI has failover chains over local and remote targets. GPUStack added public model providers.

Proposal

Terms

Term Meaning
Upstream A named endpoint outside this SIE deployment, with a kind, a base URL and a reference to a credential
Upstream kind openai for an OpenAI-compatible endpoint. sie for another SIE deployment
Remote profile A profile whose adapter is a remote adapter. Addressable as model:remote
Remote-backed model A model with no local weights. Every profile it has is a remote profile
Routing policy How a request for the bare model name chooses between the local profile and the remote profile

Where the call is made

The remote adapter runs in a worker. It is one more engine behind the adapter boundary. The gateway stays queue-only, holds no upstream credential and makes no upstream call. Its only new job is to choose which profile serves.

Three adapters are already OpenAI-protocol clients aimed at a local port: mlx, tensorrt_llm and the SGLang embedding adapter. A remote adapter is that client with a configured base URL and credential.

One implementation serves the single-node server and the cluster.

Configuration

An upstream is defined in deployment configuration: Helm values in a cluster, startup configuration on a single node. The model configuration API can name an upstream. It can never define one.

upstreams:
  team-sie:
    kind: sie
    base_url: https://sie.example.internal
    api_key_secret: team-sie-key
    rate_cap:
      requests_per_minute: 600
      max_concurrency: 32

A model gains a remote profile and a routing block.

sie_id: BAAI/bge-m3
hf_id: BAAI/bge-m3
routing:
  policy: fallback
  fallback_profile: remote
profiles:
  default:
    adapter_path: <the local adapter>
  remote:
    adapter_path: sie_server.adapters.remote.<adapter>
    adapter_options:
      loadtime:
        upstream: team-sie
        upstream_model: BAAI/bge-m3

The public model name never encodes where the model runs. A model that is served remotely today keeps its name when it is later served locally. A request that names a profile is honoured as written. The routing policy applies only to the bare model name.

Policies

Policy Behaviour
remote_only The bare model name resolves to the remote profile
fallback Local is tried first. When it refuses before accepting the work, the remote profile serves
threshold The remote profile serves while demand is low and the lane stays asleep. Sustained demand wakes the lane

fallback triggers:

Trigger Local signal Default
provisioning No healthy worker for the lane On
model_loading A healthy worker exists and the model is not loaded On
saturated Backpressure or resource exhaustion Opt-in
unhealthy Local capacity is down Opt-in

A bridged request always triggers the local warm-up. Serving remotely without waking local capacity would make the upstream the permanent server.

threshold would ship last and behind a flag.

Equivalence rule for encode and score

Under fallback and threshold, one model name is served by two backends. For encode a difference between them is permanent, because the vectors are stored. For score it changes the score scale. There is public evidence that the same weights can produce different vectors on different stacks, for example this report.

The proposal is that hybrid serving of these two primitives is refused at configuration load unless one of two conditions holds.

  • Identity, for an sie upstream. The upstream reports the same weights revision and profile identity as the local profile.
  • Proof, for an openai upstream. A passing equivalence record exists for that exact upstream and model. A probe suite produces it. It covers short and long inputs, the truncation boundary, instruction prefixes and score scale.

remote_only is not affected. It has one backend.

Caller contract

Every request served by an upstream passes through the same validation and response shaping as a local request. A field SIE does not accept is rejected, also when the upstream would accept it. Under fallback the caller cannot know which side will serve, so the accepted fields cannot depend on it.

An operator can set or strip upstream parameters per upstream. A caller cannot pass extra fields upstream.

Header Direction Meaning
X-SIE-Remote: forbid Request The request is never served remotely
X-SIE-Served-By Response local or remote
X-SIE-Upstream Response The upstream name, when served remotely
X-SIE-Fallback-Reason Response The trigger that caused remote serving
X-SIE-Fallback-Error Response Present when the upstream attempt failed and the local refusal is returned

/v1/models would report the routing policy and the upstream kind for each model.

Generation

For a model SIE has onboarded, on an upstream that offers raw completions, SIE renders the prompt with its own chat template and parses tool calls itself. Otherwise the message list is forwarded at chat level. A request with a strict grammar uses the upstream's response_format, and the worker verifies the finished output.

The reason for preferring the first mode is that published comparisons of hosts serving the same weights point at templates and parsers as the main source of differences. The K2 Vendor Verifier is one example.

Egress and credentials

Layer Control
Deployment Nothing leaves until an operator defines an upstream and a model declares a policy. One setting disables all remote serving
Model The routing block
Request X-SIE-Remote: forbid
Network Only the remote worker pool holds credentials. A NetworkPolicy limits its egress to the declared hosts
Transport Redirects are refused. TLS is required outside loopback. A URL that carries credentials is rejected
Secrets Given by reference. Never returned by a read API. Never logged or traced

Failure rules

  • Fallback happens only before local capacity accepts the work. No request is executed twice.
  • No fallback happens after the first byte has reached the caller.
  • A client error is never retried.
  • Each upstream has a circuit breaker and a required rate cap.
  • When the upstream also fails, or the cap is reached, the caller receives the original local refusal with its Retry-After.

Out of scope

  • Weighted or cost-based routing across several providers
  • Gateway-issued keys, budgets, semantic caching and guardrails
  • Provider-native dialects. An operator who needs them can place a translating gateway behind SIE as the upstream
  • Caller-supplied URLs or credentials
  • Storing request or response bodies

Suggested order of work

  1. A remote-backed embedding model on a single node through an sie upstream. This carries the upstream definition, credential handling and the global switch
  2. All four primitives through an sie upstream, with the identity check
  3. An openai upstream for embeddings, rerank and generation
  4. fallback on a single node, the request header and SDK support
  5. A remote worker pool in the Helm chart, and remote_only through the queue
  6. fallback in a cluster, with the breaker, the rate cap and the opt-in triggers
  7. The equivalence probe and the configuration gate
  8. threshold, behind a flag
  9. Documentation, including a guide for moving from a hosted API to self-hosted one model at a time

The credential and egress handling in step 1 is a trust-boundary change and would be opened as its own small pull request for review.

Open questions

  • Which profile identity can both sides of an sie upstream compute and compare
  • Whether the rate cap applies per remote worker replica or across replicas
  • How the threshold demand estimate is shared across gateway replicas
  • How the gateway requests a model load without a work item, for the model_loading trigger
  • Which thresholds the equivalence probe uses. They need a measured noise floor between two local runs

Feedback wanted

  • Is a remote backend the right scope, as opposed to a general gateway
  • Is the worker the right place for the outbound call
  • Is the equivalence rule too strict or not strict enough
  • Are the default triggers and the required rate cap the right defaults

Activity

  1. krisztian-gajdar commented on Oct 3, 2026

    @krisztian-gajdar
    ContributorAuthor

    Implementation audit and continuation, 2026-10-04:

    The proposal's “nothing is built yet” description is now historical. Public main has upstream configuration and credential/egress controls, remote SIE encode/score/extract (including sparse/multivector and supported image inputs), OpenAI-compatible dense embeddings and reranking, remote profile/routing validation, request-level remote forbidding and disclosure, the remote worker bundle/Helm support, and per-worker upstream limits/circuit breaking. Single-node extraction and generation fallback are implemented. Exact fresh OpenAI and immutable SIE hybrid encode/score admission are available for one concrete worker; fleet admission and threshold policies remain guarded/refused.

    Merged in this continuation:

    All twenty-seven merged PRs were approved by CodeRabbit on their submitted commits and merged after green CI. Independent adversarial review covered transport, queue settlement, reasoning privacy, hardware identity and evidence authority. Transient Rust runner/tool-download failures were retried; no red check was bypassed.

    • fix(worker): enforce live configuration authority through execution #543: live worker execution authority, including Python configuration leases, hashed model/profile membership, grammar routing, selected batch operation payloads, and every sidecar dispatch barrier. Merged after exact-head CodeRabbit approval and green CI.

    • feat(worker): fence verified inference with versioned IPC methods #544: updated-only backend IPC entrypoints and default-false capability, merged at 5f832a4 after exact-head CodeRabbit approval and all 46 checks green.

    • feat(worker): fence verified dispatch with versioned queues #545: versioned direct queue/stream, current authority checks, exact subject/payload model matching, all-child capability aggregation and conservative health advertisement. CodeRabbit's two availability findings were fixed and threads resolved; final exact-head approval and green CI preceded merge at 2ba3d6e. Full post-rebase sidecar suites passed 652 default/670 cloud library tests, 14 binary and 30 integration tests each.

    • feat(gateway): enforce remote-forbid with verified worker dispatch #546: gateway X-SIE-Remote: forbid across inference ingress, fresh positive worker capability, exact current hash, versioned direct dispatch and no legacy retries. CodeRabbit's balancing finding and independent review's pressure-overflow finding were fixed; all checks green, exact-head approval and the resolved thread preceded merge at 19be847 (2026-10-03 23:46 UTC).

    • feat(gateway): bridge buffered cluster generation refusals #547: buffered cluster generation fallback across native generation, chat, completions and supported Responses. Cold local demand is retained; unloaded local workers receive durably accepted load-only work before one exact-authority remote attempt. Explicit caller selectors remain authoritative. Full gateway suites, post-rebase conformance, both Clippy configurations and independent adversarial review passed. Exact-head CodeRabbit approval and all CI green preceded merge at 6ec8dec (2026-10-04 00:06 UTC).

    • feat(gateway): restore cluster stream refusals before output #548: streaming cluster generation fallback on native generation, chat and completions. HTTP success waits for the first valid event; pre-output errors and cancelled/failed terminals restore the original local refusal, while later errors stay in-stream without replay. Full gateway suites passed 1591 default/1602 cloud library and binary tests plus NATS integrations. Both Clippy configurations and independent review passed. After parent merge/rebase, fresh exact-head CodeRabbit approval and all CI green preceded merge at 0ebc55c (2026-10-04 00:23 UTC).

    The extraction/audio layer was delivered in #549. It shares local warm-up and exact worker admission, validates cheap caller-controlled input boundaries before demand/dispatch, preserves display-model names and fallback/retry headers, and retains the local refusal until the requested audio format is ready. Independent review findings on audio formatting, metadata sizing and typed input limits were fixed, with source clearance and focused regressions.

    Remaining deliveries, in dependency order:

    1. Complete cluster fallback: gateway profile choice, local warm-up/load-item production, bounded remote attempts across all pre-acceptance refusal paths, original-local-refusal restoration, opt-in triggers, fleet equivalence admission, request forbid control, telemetry and conformance tests.
    2. Flagged threshold policy with coordinated demand/wake behavior across gateway replicas.
    3. Recorded cold-start and equivalence evidence, with the required product review. No performance/equivalence claims should be made before those measurements.

    The issue remains open. The new guides explain migration checks without claiming that model names or matching dimensions prove equivalence.

    • feat(gateway): bridge cluster extraction and audio refusals #549: cluster native extraction and audio transcription cold/loading bridges, bounded worker-compatible input validation, exact metadata size accounting, requested-model preservation, and original-refusal restoration through audio response formatting. Merged at 2026-10-04 00:46:11 UTC as b009e9f40c62a20a6f3a36ec0e5c5a4db95d1467, with CodeRabbit approval on the exact submitted head and green CI.

    • feat(gateway): opt in to pre-acceptance remote spill #550: opt-in saturation/unhealthy fallback across buffered and streaming surfaces, scoped refusal evidence and pure pre-dispatch pressure checks. The unhealthy retry-hint finding was fixed. Exact-head CodeRabbit approval and all CI green preceded merge at 4f3a62efbc7ea05023afa8adf37815447dc9f5d1 (2026-10-04 01:31:12 UTC).

    Current continuation: #551 adds bounded fallback telemetry, a persistent-remote alert/dashboard, idle-period reset and valid-local-output reset. Full default/cloud-storage gateway validation, both Clippy configurations and independent adversarial review passed; CodeRabbit/CI are running. Coordinated threshold implementation is next. Numerical fleet admission and actual acceptance measurements remain open; issue #415 is not complete.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions