The benchmark submitter for llamabench.ai — the crowd-sourced local-LLM speed leaderboard.
It's a single, self-contained CLI (llamabench) that bundles nothing: it shells out to
your existing llama.cpp build (llama-bench for standardized prefill/decode speed, and
llama-server for deterministic multi-turn output-correctness checks), assembles a result, and
submits it to the leaderboard. It's open source so you can see exactly what runs on your machine
before you curl … | sh it.
curl -fsSL https://llamabench.ai/install.sh | shThis downloads the prebuilt binary for your OS/arch from the latest release
and puts llamabench on your PATH. Prefer to do it by hand? Grab the archive for your platform
from the Releases page and drop the binary somewhere on your PATH.
Supported prebuilt targets: Linux x86_64, macOS (Intel + Apple Silicon), Windows x86_64.
Take the command you already run and swap the program name. Your exact configuration is benchmarked, verified, recorded verbatim as the reproduce command, and submitted:
# 1. Save your token once (get one at https://llamabench.ai/account).
llamabench auth <token>
# 2a. You run llama-bench? Drop the dash:
# llama-bench -m model.gguf -ngl 99 -fa on -ub 2048 -ot "ffn=CPU"
llamabench -m model.gguf -ngl 99 -fa on -ub 2048 -ot "ffn=CPU"
# 2b. You run llama-server? Prefix it:
# llama-server -m model.gguf -c 8192 -np 2 --jinja
llamabench llama-server -m model.gguf -c 8192 -np 2 --jinjaEvery flag is passed through to the real tool untouched (matrix runs like
-ngl 0,99 submit one result per configuration). llamabench adds its own flags
on top — they never collide with llama.cpp's: --dry-run (don't submit),
--no-verify (skip the output-correctness pass), --token <t>,
--handle <@you>, --family <fork>, --llama-dir <bin-dir>, --api <url>,
--download-llama. Bare llamabench <flags> picks the tool automatically
(server-only flags like --port/-c ⇒ llama-server); force it with
llamabench llama-bench … or llamabench llama-server ….
In llama-bench mode the speed table you know streams as usual and the numbers
are read from llama-bench's own per-test output (-oe jsonl is appended); TTFT
is probed on the verification server with the same standardized ~512-token
prompt the server mode uses, so the two modes' TTFTs are comparable. In
llama-server mode the server runs with your args verbatim and prefill/decode/TTFT
come from the server's own timings on standardized requests (temp 0, ~512-token
prompt, 128 generated tokens, median of 3).
Every submission of a local file records the GGUF's SHA-256 (hashed once per file, then cached by size/mtime). Linking that file to the Hugging Face repo it came from happens on llamabench.ai: open the result and name the repo — the server verifies the hash against the repo's published LFS hashes and, once one person has linked a file, every past and future submission of the same bytes is attributed automatically. Nothing to do in the CLI.
Prefer to pin provenance locally (offline / scripted runs)? The CLI link store still works and takes precedence:
llamabench link ./gemma-4-12b-it-UD-Q4_K_XL.gguf unsloth/gemma-4-12b-it-GGUF
llamabench link --list # show all links
llamabench link --forget <path> # remove oneThe original flag-based interface still works (and is what the submit page generates):
# Fetch the model from Hugging Face AND a prebuilt llama.cpp — no local setup:
llamabench run --hf-model bartowski/Llama-3.1-8B-Instruct-GGUF --quant Q4_K_M --download-llama
# Local model + your own llama.cpp build:
llamabench run --model /path/to/model.gguf --llama-dir /path/to/llama.cpp/build/bin
# Benchmarking a llama.cpp fork? Name it with --family so the result is recorded
# under that engine (ik_llama.cpp, beellama.cpp, or Xpress AI's ve_llama.cpp for the
# NEC Vector Engine). Forks have no prebuilt download — point --llama-dir at your build.
llamabench run --model /path/to/model.gguf \
--family ik_llama.cpp --llama-dir /path/to/ik_llama.cpp/build/bin
# One-off provenance without a persistent link (hash-verified this run only):
llamabench run --model /path/to/Llama-3.1-8B-Instruct-Q4_K_M.gguf \
--hf-model bartowski/Llama-3.1-8B-Instruct-GGUF --quant Q4_K_M
# Speed only / verification only / build-but-don't-submit:
llamabench bench --model /path/to/model.gguf
llamabench verify --model /path/to/model.gguf
llamabench run --model /path/to/model.gguf --dry-run--hf-model <repo> --quant <Q>downloads a GGUF straight from Hugging Face (streamed to a per-user cache, skipped if already present), picking the.ggufwhose name matches the quant.--quantalso sets the quant recorded in the result. Use--model <path>instead to point at a local file.- Model attribution: when you pass
--hf-model, the submission is attributed to the GGUF's base/finetune model (its Hugging Facebase_model, e.g.unsloth/gemma-4-12b-it-GGUF→google/gemma-4-12b-it) rather than the per-quant llama-bench label, so every GGUF repack of the same model groups together on the leaderboard. The repo is still recorded as provenance inhfModel. If nobase_modelis published (or no--hf-modelis given), the original per-quant label is kept. --model <path> --hf-model <repo> --quant <Q>(given together) benchmarks the local file but records its Hugging Face provenance and verifies it: the runner streams the local file through SHA-256 and compares it against the repo's published hash (thelfs.oidfrom HF's tree API) for the matching quant. The result carrieshfModelandhfVerified(✓match /⚠mismatch). A provenance check that can't be resolved recordshfVerified: falseand never fails the run.--download-llamagrabs the latest prebuilt llama.cpp release for your OS/arch. This is the standard CPU/Metal build only — GPU builds (CUDA / HIP / Vulkan) are NOT auto-selected. If you have a GPU, build llama.cpp yourself and point--llama-dirat it for full speed. With neither--llama-dirnor--download-llama, the runner usesllama-bench/llama-serverfrom yourPATH, and falls back to the prebuilt CPU/Metal build if they aren't found.--family <llama.cpp|ik_llama.cpp|beellama.cpp|ve_llama.cpp>records which llama.cpp variant the build is (defaultllama.cpp), so results from different engines stay comparable but distinct on the leaderboard. The forks share the samellama-bench/llama-serverCLI, so the runner drives them identically — but only upstream llama.cpp has prebuilt downloads, so build the fork and point--llama-dirat it (or put its binaries onPATH).ve_llama.cppis Xpress AI's fork adding NEC SX-Aurora Vector Engine support.
run resolves the submission token in this order: --token flag →
LLAMABENCH_TOKEN env var → the token saved by llamabench auth. If none is found
(and you're not using --dry-run), it errors and points you at llamabench auth.
Common flags (see --help for the full list): --ngl, --fa, --ctk/--ctv (KV cache type),
--n-prompt/--n-gen, --spec-decode, --seed, --turns, --reps.
Pass extra flags straight through to llama-server (handy for the many speculative-decoding
options) with either:
--server-args "<flags>"— one whitespace-delimited string, e.g.--server-args "--spec-type draft-mtp --spec-draft-n-max 2". Easiest for a bunch at once.--server-arg <value>— repeatable, one value each (--server-arg --foo --server-arg "two words"). Use it when a value contains spaces.
Both are appended (repeatable --server-arg first, then the split --server-args).
cargo build --release
# binary at target/release/llamabenchRequires a stable Rust toolchain. The only dependencies are crates.io packages — no submodules, no codegen.
Results are submitted under a token tied to your llamabench.ai account and land unverified; a
✓ verified badge is reserved for independently reproduced results. The runner records the exact
configuration and the llama.cpp revision so any result is reproducible. See the
Methodology page for details.
GPL-3.0-or-later. The llamabench.ai web app is a separate, proprietary project; the runner talks to it only over the documented result API.