Skip to content
@modelsphere

ModelSphere

Industrial-grade LLM inference with instant deployment, heterogeneous hardware support, latest-model readiness, and continuous optimization

ModelSphere

An open-source LLM inference platform for Kubernetes: deploy a model with a tuned configuration, route each request to the replica that already holds its context, scale on LLM signals and SLOs, and tune the serving configuration on your own workload.

Start here: modelsphere/modelsphere — the deployment repository, with the install guide and the architecture.

Area Repository What it does
Install modelsphere Installs the whole stack: Ansible, helmfile, offline bundle
helm-charts Helm charts for the sglang and vllm engines and the routing components
model-catalog The models swiss can deploy, and how to serve each one: engine, image, flags, GPUs, tuned variants
Routing llm-openresty Session-affinity router, request logging and metrics
cache_aware_router Routes to the replica holding the longest matching prompt prefix
autoconfig Keeps the routing layer's config in sync with the model backends
Scaling and health llm-operator Autoscales on KV-cache utilization, queue depth and TPM
slo-scaler-decision-gen Turns SLO targets and live signals into replica decisions
slo-api HTTP API to read and change a service's SLO
hang-watcher Restarts an engine that stopped making progress
Operate swiss Deploy control plane for the engine charts
console Web console: users and roles, model deployment, playground
Tune and measure llm-autotune Searches serving configurations on idle GPUs
llm-autotune-policies Search policies and the policy SDK for llm-autotune
llm-bench Benchmark platform for LLM serving endpoints

Everything here is licensed under Apache-2.0. Contributing · Code of Conduct · Security

Pinned Loading

  1. modelsphere modelsphere Public

    Deploy an LLM inference stack on Kubernetes: cache-aware routing, LLM-aware autoscaling, etc.

    Shell 4 2

Repositories

Showing 10 of 16 repositories
  • llm-openresty Public

    OpenResty-based session-affinity router for LLM inference backends

    modelsphere/llm-openresty's past year of commit activity
    Lua 0 Apache-2.0 0 0 1 Updated Oct 5, 2026
  • llm-operator Public

    Kubernetes operator that autoscales LLM inference workloads on LLM-specific signals — KV-cache utilization, queue depth, TPM — rather than CPU and memory

    modelsphere/llm-operator's past year of commit activity
    Go 1 Apache-2.0 1 0 7 Updated Oct 5, 2026
  • autoconfig Public

    Kubernetes operator that keeps routing-layer config (OpenResty / cache-aware-router / monitoring) in sync with discovered model backends

    modelsphere/autoconfig's past year of commit activity
    Go 1 Apache-2.0 0 0 3 Updated Oct 5, 2026
  • llm-bench Public

    An open benchmark platform for LLM serving endpoints: throughput sweeps, functional acceptance, real-traffic replay, scored leaderboards

    modelsphere/llm-bench's past year of commit activity
    Python 0 Apache-2.0 0 0 1 Updated Oct 5, 2026
  • helm-charts Public

    Helm charts for the ModelSphere LLM inference stack (sglang, vllm, cart, rdma-injector)

    modelsphere/helm-charts's past year of commit activity
    Go Template 1 Apache-2.0 1 0 4 Updated Oct 5, 2026
  • cache_aware_router Public

    Cache-aware router for LLM inference backends: routes each request to the replica that already holds the longest matching prompt prefix

    modelsphere/cache_aware_router's past year of commit activity
    Rust 0 Apache-2.0 0 0 1 Updated Oct 5, 2026
  • swiss Public

    Deploy control plane for the ModelSphere sglang and vllm charts, driven by the model catalog

    modelsphere/swiss's past year of commit activity
    Go 2 Apache-2.0 1 0 3 Updated Oct 5, 2026
  • model-catalog Public

    The models swiss can deploy, and how to serve each one: engine, image, flags, GPUs and tuned variants

    modelsphere/model-catalog's past year of commit activity
    HTML 1 Apache-2.0 1 0 0 Updated Oct 4, 2026
  • console Public

    ModelSphere community portal — identity (users/roles/login) + federation gateway to swiss and other backends

    modelsphere/console's past year of commit activity
    TypeScript 1 Apache-2.0 2 0 1 Updated Oct 3, 2026
  • llm-autotune Public

    Tunes LLM serving configurations overnight on your own GPUs: search spaces, pluggable search policies, benchmarked by LLMBench, winners promoted as merge requests

    modelsphere/llm-autotune's past year of commit activity
    Python 0 Apache-2.0 0 0 19 Updated Oct 2, 2026

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…