Skip to content

About

GUI interface for llama.cpp with the ability to route models

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

35 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llamacpp-launcher

A desktop GUI launcher for llama.cpp, written in Rust with egui/eframe. It lets you configure, run and monitor multiple llama-server instances at once, and exposes a built-in OpenAI-compatible router that dispatches requests to the right instance by model alias.

Status: early development (0.1.0). Windows is the primary target; the code is cross-platform but tested mainly on Windows.


Features

  • Multiple instances — run several llama-server processes in parallel, each with its own model, port, GPU device and parameters.
  • Live status & logs — per-instance status (starting / running / stopped / failed), console output and generation statistics (tokens/s).
  • OpenAI-compatible router — a single endpoint in front of all instances that routes /v1/chat/completions, /v1/completions, /v1/embeddings, … by the request's model field, matched against each instance's alias. Supports streaming (SSE), API-key injection and CORS.
  • Dynamic routing — the routing table is rebuilt every frame from instances that are running and have routing enabled, so starting/stopping an instance updates routes instantly without restarting the router.
  • Persistent configuration — global settings and every instance are stored as TOML and reloaded on startup.
  • Inference device detection — enumerates inference devices via llama-cli --list-devices.

Requirements

  • A working llama.cpp build containing llama-server and llama-cli binaries.
  • Rust (edition 2024 toolchain) to build from source.

Build & run

cargo run --release

The app uses the glow OpenGL renderer (eframe's wgpu default is disabled in Cargo.toml because it crashes on some Windows setups).

Getting started

  1. Open ⚙ Settings and set the folder that contains your llama-server / llama-cli binaries. Detected GPUs are listed there.
  2. In the left panel, click ➕ Add to create an instance.
  3. Select the instance and configure it in the Settings tab:
    • Model — pick a local .gguf file or a Hugging Face reference (repo/model:file).
    • Inference — context size, threads, GPU layers, GPU device, offloading flags.
    • Routing — enable routing, set an alias (the model name clients will request), and an optional API key.
  4. Click ▶ Start (or 🔁 Restart / ⏹ Stop). Logs and stats appear in the Logs tab.

Using the router

  1. Configure the router host/port in ⚙ Settings (default 127.0.0.1:11434).
  2. Start it from the 🔀 Router menu (Start / Restart / Stop). Use Logs window to see proxied requests.
  3. Point any OpenAI-compatible client at the router and select an instance by its alias:
curl http://127.0.0.1:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<instance-alias>",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'
  • GET /v1/models lists the currently available aliases.
  • If exactly one route is active, the model field may be omitted.
  • Unknown model → 404 with the list of available aliases; unreachable instance → 502.

Configuration files

Stored under the per-user config directory (%APPDATA%\.llamacpp-launcher on Windows, $HOME/.llamacpp-launcher otherwise):

.llamacpp-launcher/
├── global.toml            # llama.cpp dir, router host/port/enabled
└── instances/
    └── <id>.toml          # one file per instance

About

GUI interface for llama.cpp with the ability to route models

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages