A desktop GUI launcher for llama.cpp, written in
Rust with egui/eframe. It lets you configure, run and
monitor multiple llama-server instances at once, and exposes a built-in
OpenAI-compatible router that dispatches requests to the right instance by model alias.
Status: early development (
0.1.0). Windows is the primary target; the code is cross-platform but tested mainly on Windows.
- Multiple instances — run several
llama-serverprocesses in parallel, each with its own model, port, GPU device and parameters. - Live status & logs — per-instance status (starting / running / stopped / failed), console output and generation statistics (tokens/s).
- OpenAI-compatible router — a single endpoint in front of all instances that routes
/v1/chat/completions,/v1/completions,/v1/embeddings, … by the request'smodelfield, matched against each instance's alias. Supports streaming (SSE), API-key injection and CORS. - Dynamic routing — the routing table is rebuilt every frame from instances that are running and have routing enabled, so starting/stopping an instance updates routes instantly without restarting the router.
- Persistent configuration — global settings and every instance are stored as TOML and reloaded on startup.
- Inference device detection — enumerates inference devices via
llama-cli --list-devices.
- A working llama.cpp build containing
llama-serverandllama-clibinaries. - Rust (edition 2024 toolchain) to build from source.
cargo run --releaseThe app uses the glow OpenGL renderer (eframe's wgpu default is disabled in
Cargo.toml because it crashes on some Windows setups).
- Open ⚙ Settings and set the folder that contains your
llama-server/llama-clibinaries. Detected GPUs are listed there. - In the left panel, click ➕ Add to create an instance.
- Select the instance and configure it in the Settings tab:
- Model — pick a local
.gguffile or a Hugging Face reference (repo/model:file). - Inference — context size, threads, GPU layers, GPU device, offloading flags.
- Routing — enable routing, set an alias (the model name clients will request), and an optional API key.
- Model — pick a local
- Click ▶ Start (or 🔁 Restart / ⏹ Stop). Logs and stats appear in the Logs tab.
- Configure the router host/port in ⚙ Settings (default
127.0.0.1:11434). - Start it from the 🔀 Router menu (Start / Restart / Stop). Use Logs window to see proxied requests.
- Point any OpenAI-compatible client at the router and select an instance by its alias:
curl http://127.0.0.1:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<instance-alias>",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'GET /v1/modelslists the currently available aliases.- If exactly one route is active, the
modelfield may be omitted. - Unknown model →
404with the list of available aliases; unreachable instance →502.
Stored under the per-user config directory
(%APPDATA%\.llamacpp-launcher on Windows, $HOME/.llamacpp-launcher otherwise):
.llamacpp-launcher/
├── global.toml # llama.cpp dir, router host/port/enabled
└── instances/
└── <id>.toml # one file per instance