A distributed VRAM gateway for load-balancing LLM inference across GPU nodes running vLLM.
Plinth sits in front of multiple vLLM instances and provides:
- Load balancing — Round-robin request distribution across healthy GPU nodes
- Health monitoring — Continuous health checks with configurable failure thresholds
- GPU metrics collection — Scrapes NVIDIA GPU metrics (memory, utilization, temperature) via NVML
- Prometheus metrics — Exposes request latency, counts, and node health gauges
- OpenAI-compatible API — Proxies
/v1/chat/completionsand/v1/completionsrequests
┌─────────────────┐
│ Client (LLM) │
└────────┬────────┘
│
┌────────▼────────┐
│ Gateway │ ← Load balancer + health monitor
│ (plinth) │
└────────┬────────┘
│
┌────┴────┐
│ │
┌───▼───┐ ┌──▼────┐
│ Node1 │ │ Node2 │ ← vLLM + gpu-exporter
│ :8000 │ │ :8000 │
│ :9100 │ │ :9100 │
└───────┘ └───────┘
Each GPU node runs:
- vLLM — LLM inference server (default port 8000)
- gpu-exporter — Per-GPU Prometheus metrics (memory, utilization, temperature) for all NVIDIA GPUs, UUID-labeled (default port 9100)
- Go 1.22+
- NVIDIA GPU drivers + NVML (for gpu-exporter)
- vLLM instances running on GPU nodes
# Build all binaries (gateway, gpu-exporter, plinth-provision)
make all
# Or build individually
make build-gateway
make build-gpu-exporter
make build-provision
# Or use go directly
go build -o gateway ./cmd/gateway
go build -o gpu-exporter ./cmd/gpu-exporter
go build -o plinth-provision ./cmd/provisionThe Makefile also provides: test, vet, fmt, cover, clean, and versioned build targets (build-gateway-versioned, build-gpu-exporter-versioned, build-provision-versioned) that inject git commit hash and build time.
cp config/gateway.toml.example config/gateway.toml
# Edit config/gateway.toml with your cluster settings# Start gateway
./gateway -config config/gateway.toml
# Start gpu-exporter on each GPU node
./gpu-exporterSet [provision] ssh_host_key in config/gateway.toml first (the SHA256 fingerprint printed by ssh-keyscan -t ed25519 <node-ip>). Then:
# Provision a specific node
./plinth-provision -config config/gateway.toml provision gpu-node-1
# Provision all nodes
./plinth-provision -config config/gateway.toml provision --all
# Download model weights to controller
./plinth-provision -config config/gateway.toml weights pull Qwen/Qwen2.5-7B-Instruct
# Push weights to specific nodes
./plinth-provision -config config/gateway.toml weights push Qwen/Qwen2.5-7B-Instruct gpu-node-1
# Push weights to all nodes serving that model
./plinth-provision -config config/gateway.toml weights push Qwen/Qwen2.5-7B-Instruct --allNode provisioning installs drivers, Python, vLLM, the service user, and model weights. Run weights pull <model> before provision so the models step has local weights to push. weights push --all requires the model to be listed in [[models.available]]. Systemd service installation is handled separately by scripts/provision-node.sh.
See config/gateway.toml.example for a complete example.
| Section | Key | Description | Default |
|---|---|---|---|
[cluster] |
name |
Cluster identifier | — |
[gateway] |
listen |
API listen address | :8000 |
[gateway] |
metrics_listen |
Prometheus metrics address | :9090 |
[gateway] |
health_interval |
Health check interval | 3s |
[gateway] |
health_fail_threshold |
Failures before marking dead | 3 |
[gateway] |
request_overload_threshold |
Consecutive request-level 5xx before a node is relegated to the degraded fallback pool | 3 |
[gateway] |
max_buffered_response_bytes |
Max buffered non-streaming response bytes | 8388608 |
[gateway] |
max_request_body_bytes |
Max accepted request body bytes | 33554432 |
[[nodes]] |
ip |
Node IP address | — |
[[nodes]] |
name |
Node display name | — |
[[nodes]] |
vllm_port |
vLLM API port | 8000 |
[[nodes]] |
metrics_port |
gpu-exporter port | 9100 |
[[nodes]] |
ring |
Ring this node belongs to (empty for standalone) | — |
[models] |
default |
Default model name | — |
[[models.available]] |
name |
Model identifier | — |
[[models.available]] |
pipeline_stages |
Informational metadata only | 1 |
[[models.available]] |
ring |
Ring whose nodes serve this model (empty for standalone) | — |
[provision] |
weights_dir |
Local path for model weights | /var/lib/plinth/weights |
[provision] |
ssh_key |
Path to SSH private key | ~/.ssh/id_rsa |
[provision] |
ssh_user |
SSH login user | root |
[provision] |
ssh_port |
SSH port | 22 |
[provision] |
ssh_host_key |
Node host key fingerprint (required for provision/weights push) | — |
Node IPs must be unique — the gateway identifies nodes by IP.
| Method | Path | Description |
|---|---|---|
GET |
/health |
Cluster health status |
GET |
/v1/models |
List available models |
POST |
/v1/chat/completions |
Chat completion (proxied to vLLM) |
POST |
/v1/completions |
Text completion (proxied to vLLM) |
| Method | Path | Description |
|---|---|---|
GET |
/metrics |
Prometheus metrics |
| Method | Path | Description |
|---|---|---|
GET |
/metrics |
GPU metrics (memory, utilization, temperature) |
| State | Description |
|---|---|
Healthy |
Node responding to health checks |
Degraded |
Node failing health checks (below threshold) or relegated by a request-level 5xx streak |
Dead |
Node exceeded failure threshold |
The gateway prefers healthy nodes. If none are available, it falls back to degraded nodes. Dead nodes are excluded from load balancing.
Deployment scripts are provided in scripts/:
provision-node.sh— Set up a GPU node (vLLM + gpu-exporter)setup-gateway.sh— Set up the gateway nodesetup-keepalived.sh— Configure keepalived for HA gatewayvllm.service— systemd unit for vLLMgpu-exporter.service— systemd unit for gpu-exporterkeepalived-primary.conf/keepalived-secondary.conf— keepalived configs
GitHub Actions workflows in .github/workflows/:
- PR checks (
pr.yml) — Runsgo vet, tests with race detector, and builds gateway + gpu-exporter on linux/amd64 and linux/arm64 - Release (
release.yml) — On merge to main: runs tests, builds versioned gateway + gpu-exporter binaries, creates a GitHub release with commit hash version, and attaches the built binaries
plinth/
├── cmd/
│ ├── gateway/ # Gateway binary
│ ├── gpu-exporter/ # GPU metrics exporter
│ └── provision/ # Node provisioning + model weights CLI
├── config/ # Configuration files
├── docs/ # Detailed documentation
├── .github/workflows/ # CI/CD (PR checks + release)
├── internal/
│ ├── api/ # HTTP handlers + reverse proxy
│ ├── balancer/ # Round-robin load balancer
│ ├── config/ # TOML config loader
│ ├── gpumetrics/ # NVML collector + Prometheus exporter
│ ├── health/ # Node health monitoring
│ ├── log/ # Structured logging
│ ├── metrics/ # Prometheus gateway metrics
│ ├── provision/ # SSH client, provisioners, weight manager
│ └── version/ # Build version info
├── Makefile # Build, test, and utility targets
└── scripts/ # Deployment scripts
- Config reload —
SIGHUPreloads nodes, models, rings, and health settings at runtime. In-flight requests finish against the old config; health state carries over for unchanged nodes; an invalid config is rejected with the current config kept. Listen address changes require a restart. - Per-node model awareness (exact id match) — Routing still starts from ring membership (ring models round-robin among that ring's nodes, other models among ring-less nodes), but nodes are additionally filtered by each node's
GET /v1/modelslist. Matching is exact string equality againstdata[].id.models.defaultand every[[models.available]].namemust equal the node's served model id (no aliases). Unknown/ad-hoc request models use the ring-less pool and are then filtered the same way, so a model missing from every known list returns 503. Discovery failures never mark a node unhealthy; afterhealth_fail_thresholdconsecutive listing failures the node is excluded from model routing until listing recovers. A node that reports zero models on two consecutive checks is excluded from all model routing. - Overload sensing — vLLM's
/healthis a liveness probe and stays 200 while the engine is saturated, so node health alone cannot see overload. The gateway feeds each proxied attempt's status back to the monitor:request_overload_threshold(default 3) consecutive 5xx responses relegate a node to the degraded fallback pool (never dead — death stays probe-owned). The streak clears on the first non-5xx response from that node, or decays after 10s without another 5xx, whichever comes first, so recovery needs no operator action./healthexposes per-nodereq_streak. Transport failures and the connection-failure offline path are unchanged. - Streaming —
stream: truecompletion requests are streamed chunk-by-chunk to the client. Retry is allowed until the first chunk is flushed; once committed, an upstream failure terminates the stream (clean EOF, no fabricated[DONE], never a 502 after commit). Stalled streams without a first chunk are abandoned and retried after a fixed 10s deadline. Non-streaming requests are buffered up to max_buffered_response_bytes (default 8 MiB) so the attempt stays retryable; larger responses are committed early and passed through to the client. - Connection-level retry only — Requests whose proxy attempt fails to establish a connection (or times out before any response) are retried on other nodes in the pool. HTTP errors from vLLM are never retried (completions are not idempotent), and errors after response headers start return 502 unless the response has already been committed. Retrying on a response-header timeout carries a small duplicate-generation risk. For streaming requests, failures before the first chunk reach the client are retried; failures after commit terminate the stream instead of returning 502. Once a response is committed (buffered complete, or oversized and passed through), an upstream failure yields a truncated body and never a 502.
- Provisioning —
plinth provisioninstalls drivers, Python, vLLM, the service user, and model weights, but not systemd services; service setup remains inscripts/provision-node.sh.
plinth is Copyright (c) 2026 Blaine Motsinger under the MIT license.