Local AI in one static Go binary.
LLMs, vision, speech, images and decisions. No Python, no cgo. CPU, or GPU with -vulkan.
Quickstart · Speed · What golem does differently · How it is checked
A golem is inert matter given a voice. That is what these engines do to a file of weights.
Golem runs Gemma 4, Qwen3, Qwen3.8, Bonsai 2, Kyutai's speech models, Krea 2 and Laya on your own machine. It is one Go binary of about twelve megabytes, with nothing else to install, and it keeps pace with llama.cpp while doing it.
go install github.com/ThiraSoft/golem/cmd/golem-cli@latest
golem-cli -model gemma-4-E2B-it-QAT-Q4_0.gguf -p "Explain a mutex in one sentence."No Go toolchain? Take a static binary from the releases for Linux or macOS, amd64 or arm64. There is no configuration file: the engine reads the architecture out of the GGUF and opens the implementation that matches.
On one RX 9070 XT, every figure a benchmark in this repository:
- Gemma 4 26B A4B at 153 tokens a second, against 125 for llama.cpp on the same file.
- The same 26B in 5 GB of VRAM at 56 tokens a second, experts left in system memory.
- Bonsai 2 27B at 46 tokens a second, five times Prism's own llama.cpp fork, and 63 when it drafts.
- A Krea 2 picture in 8.9 seconds, the same picture ComfyUI draws for the same seed, at half its time a step.
- A Laya decision in 7.8 ms, two and a half times faster than the checkpoint's own PyTorch.

Drawn by golem's Krea 2 engine: 1280 × 720, eight steps, 10.4 s on an RX 9070 XT.
| Text | Gemma 4 (E2B, 12B, 26B A4B), Qwen3, Qwen3.8 27B, and Prism's ternary Bonsai 2 27B |
| Vision | Gemma 4 and Qwen3.8, from a projector file |
| Audio in | Gemma 4 hears WAV, MP3 and FLAC. Kyutai STT transcribes English and French, live from the microphone |
| Audio out | Kyutai Pocket TTS, 12 shipped models across 6 languages, and voice cloning from 20 seconds of audio |
| Images | Krea 2 from ComfyUI's fp8 files, LoRA included, the same picture as ComfyUI for the same seed |
| Decisions | Laya, the open Jev: a state and typed questions in, calibrated probabilities out, in one encoder pass |
| Embeddings | nomic-embed-text-v2-moe, behind OpenAI's and ollama's endpoints |
| Serving | OpenAI-compatible HTTP API, Jev's decision endpoint, tool calls, continuous batching, JSON schemas and GBNF grammars |
| Weights | GGUF with every K-quant llama.cpp writes, Prism's ternary PQ2_0 and PTQ1_0, and golem's own .golem format |
Every command is go install github.com/ThiraSoft/golem/cmd/<name>@latest, or go build ./cmd/<name> from a clone. Add -vulkan to any of them to run on the GPU.
golem-cli -model Qwen3-4B-Q4_0.gguf -p "Explain a mutex in one sentence." -statsServe an OpenAI-compatible API
golem-server -model Qwen3-4B-Q4_0.gguf -addr 127.0.0.1:8080 -parallel 4-parallel N cuts the context into N slots and holds N conversations at once. Whatever is waiting when a pass is built rides in that pass, so four clients wanting a token are one read of the weights instead of four. The same server answers transcriptions (-stt), embeddings (-embed), image generation (-krea2) and Jev's decision endpoint (-laya). Details in cmd/golem-server/README.md.
Look at a picture, listen to a recording
golem-cli -model gemma-4-E2B-it-QAT-Q4_0.gguf \
-mmproj mmproj-gemma-4-E2B-it-QAT-BF16.gguf \
-image photo.png -p "What is in this picture?"The same flags with -audio question.wav make Gemma listen. See gemma/README.md.
Speak, clone a voice, transcribe
pocket-tts -voice voice.safetensors -o hello.wav "Bonjour le monde."
pocket-tts -clone someone.wav -save-voice someone.safetensors # 20s of audio is enough
golem-cli -stt ~/models/stt-1b-en_fr -listen # microphone in, words outVoice cloning needs no training. Transcription streams at twice real time on eight CPU cores, and about three times on the card. See stt/README.md and pockettts/README.md.
Draw a picture with Krea 2
GOLEM_KREA2_COMFY=~/ComfyUI krea2 -prompt "a red fox in the snow, photo" -seed 5 -out fox.png
krea2 -prompt "..." -lora style.safetensors -lora-strength 0.8It reads ComfyUI's own three files where ComfyUI keeps them, and wants a Vulkan card with about 14 GB. See krea2/README.md.
Ask Laya a decision
laya -model convaiinnovations/laya -vulkan request.json{
"state": {"subject": "Refund not received", "body": "..."},
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this email?",
"criteria": {"billing": "payments, refunds", "support": "technical problems"}},
"spam": {"type": "noul", "instructions": "This email is spam."}
}
}Every option comes back with a calibrated probability, in Jev's shape. golem-server -laya serves the same on POST /v1/systemone, Jev's own endpoint, so a Jev client needs only its base URL changed. The English, typed-decisions and multilingual checkpoints open as Hugging Face ships them. See laya/README.md.
Ask for JSON and get JSON
golem-cli -model Qwen3-4B-Q4_0.gguf -json-schema city.json -p "Give me the city of Lyon."A model drawing inside a grammar cannot break the document, because every token that would is refused before the draw. -json, -json-schema and -grammar on the command line, response_format and grammar over the API. The engine is a port of llama.cpp's own, compared against it rule by rule. See grammar/README.md.
Use it as a library
m, err := engine.Open("Qwen3-4B-Q4_0.gguf", 4096) // reads the architecture, opens the engine
defer m.Close()
ids := m.Vocab.Encode("Explain a mutex in one sentence.", true, true)
hidden := m.Forward.ForwardBatch(ids, 0)
logits := make([]float32, m.Vocabulary)
m.Forward.Logits(hidden[len(hidden)-1], logits)engine.Open hands back one shape whichever family the file is: the forward pass, the vocabulary, the chat template the checkpoint carries, and what to sample with. The commands in cmd/ are thin wrappers over it, and pkg.go.dev has the rest.
Three ideas from this repository that go past running a model as it was shipped.
Gemma 4 26B A4B keeps 12.85 GB of experts and reads eight matrices out of a hundred and twenty-eight per block, so eleven of those twelve gigabytes sit untouched on any given token. GOLEM_MOE_EXPERTS_HOST=1 leaves the experts in system memory the card can address, takes the model's footprint on the card from 13.6 GiB to 1.3, and a cache of the experts a token keeps asking for buys the speed back:
| experts kept in VRAM | that much VRAM | tokens/s |
|---|---|---|
| 2 of 128 (the floor) | 0.2 GB | 7.8 |
| 32 of 128 | 3.2 GB | 30.2 |
| 51 of 128 | 5.1 GB | 56.2 |
| all 128 (fully resident) | 12.9 GB | 131.6 |
Two fifths of the pool buys four fifths of the tokens, and the answers are identical in every row. gemma/README.md has the full curve and the block-by-block residency that runs a 27 GB checkpoint on a 16 GB card.
The Qwen3.8 checkpoint ships a block whose job is to guess the token after the one just decided. Guess right and the next pass verifies two tokens for the price of one; guess wrong and it costs nothing beyond the pass it rode on. Every token is drawn from the model's own distribution, and a test asserts the answer is the same either way.
On an RX 9070 XT: 27.6 tokens a second one at a time, 40.1 drafting, 70 % of drafts accepted. Bonsai 2 ships without the block, and the original's grafts onto it unchanged: 46.0 becomes 62.8. See qwen35/README.md.
golem reads GGUF like everyone else. It also writes .golem, a trellis code after QTIP, which trades speed for size. On Qwen3-4B against the same bf16 reference, 4088 tokens of wikitext:
| size | perplexity | KL divergence | top-1 agreement | |
|---|---|---|---|---|
| bf16 | 7.5 GiB | 19.29 | ||
| Q4_K_M | 2.33 GiB | 20.04 | 0.0715 | 90.1 % |
.golem H4G |
1.96 GiB | 20.07 | 0.0495 | 90.2 % |
| Q3_K_M | 1.93 GiB | 24.03 | 0.2453 | 79.8 % |
.golem H3G |
1.54 GiB | 21.67 | 0.1715 | 84.6 % |
H3G is 20 % smaller than Q3_K_M and ahead of it on every column. On Qwen3.8-27B it is 10.45 GiB, closer to bf16 than Q3_K_M, and faster than Q4_0 on the same card: 38.8 tokens a second against 34.1.
golemquant -model Qwen3-4B-BF16.gguf -out Qwen3-4B.golem -bits 3 -calib wiki.txtcompress/README.md has the method and the measurements.
On an RX 9070 XT against llama.cpp's Vulkan build (ba1df050f, b9603), same Q4_0 files, both sides warmed, measured on 2026-09-03 in one sitting. Tokens a second, higher is better:
| golem gen | llama gen | golem pp512 | llama pp512 | |
|---|---|---|---|---|
| Gemma 4 26B A4B | 153.4 | 125.1 | 4434 | 4059 |
| Gemma 4 12B | 72.3 | 64.6 | 2901 | 2978 |
| Qwen3 4B | 180.5 | 172.0 | 5977 | 7117 |
| Qwen3 0.6B | 368.2 | 407.9 | 23445 | 23677 |
Bonsai 2 27B does not load in stock llama.cpp, so it is held against Prism's own fork (PrismML-Eng/llama.cpp, bdc23b5), same card, 2026-09-23:
| golem gen | fork gen | golem pp512 | fork pp512 | |
|---|---|---|---|---|
| Bonsai 2 27B PQ2_0 | 46.0 | 9.2 | 1253 | 924 |
| Bonsai 2 27B PTQ1_0 | 15.8 | 9.4 | 1128 | 423 |
Against the reference each model was written from, same card, 2026-09-28:
| golem | reference | |
|---|---|---|
| Krea 2, one DiT step at 768 × 1024 | 1.02 s | 1.89 s (ComfyUI) |
| Laya, one question | 7.8 ms | 19.9 ms (PyTorch) |
| Laya, 32 questions | 51 ms | 56 ms (PyTorch) |
On an i7-9700K with eight threads and Q4_0 weights, Gemma E2B draws 22.6 tokens a second and reads 204, the 26B A4B does 13.1 and 51, Qwen3 4B does 14.6 and 110: a tie with llama.cpp on generation, and between ×1.06 and ×1.33 reading a prompt.
Read these tables as a whole: golem reaches llama.cpp's level, and that is the claim. It is one card and one driver. A kernel that wins on RDNA 4 need not win elsewhere, and the column worth reading is the gap between the two engines rather than the rate. The one model golem loses is Qwen3 0.6B, where the weights are small enough that arithmetic rather than memory sets the pace, and qwen/README.md says why. ARCHITECTURE.md has the kernel work, the per-op breakdown against llama.cpp, and the next bottleneck.
If you ship Go and you need a model to run on the machine your software already runs on, this removes the part of that job that hurts. No Python runtime beside your binary, no four-gigabyte CUDA image, no cgo: GOOS=linux GOARCH=amd64 go build still produces something you can scp to a box you do not control.
That covers on-premise deployments, edge and embedded boxes, air-gapped machines, CI, and any product where "please install Python 3.11 first" is not a sentence to put in front of a customer. If you run models on your own workstation and a package manager is fine, Ollama is friendlier.
Why not bind to llama.cpp? Because a cgo binding is no longer a static binary. It brings back the C toolchain and a build that breaks differently on every distribution. The price of leaving it out was a full set of kernels: AVX2 and NEON in Go assembly, Vulkan compute bound through purego at runtime rather than linked. Four dependencies in all, and CGO_ENABLED=0 go build ./... passes.
No layer is deemed correct until its intermediate activations match the reference implementation.
Scripts load the real weights, inject a deterministic input, and write every intermediate quantity into testdata/. The Go tests read those files back, so they need neither Python nor llama.cpp at test time. For gemma/ the reference is llama.cpp itself, instrumented, because a bf16 reference would bury a mistake under its own quantization error. For pockettts/, stt/, krea2/ and laya/ it is PyTorch, layer by layer.
The code in this repository was written by an AI agent, Claude, directed and reviewed by a human. Some of the Bonsai 2 work was written by Gemini under Claude's review. The git history says so on the commits themselves.
That is the reason the section above exists. An agent will happily produce a layer that runs, returns plausible tokens and is quietly wrong in the fourth decimal place, and reading the diff does not catch that. Recorded activations do. Every claim in these files is a test in the repository or a benchmark in it. Judge it the way you would judge any dependency you did not write: run the tests, check the numbers, read the parts you are about to trust.
cmd/golem-cli,cmd/golem-server,cmd/pocket-tts,cmd/krea2,cmd/laya,cmd/golemquantare the commands.engine/reads the architecture out of a GGUF and opens the engine that implements it.gemma/,qwen/,qwen35/,pockettts/,stt/,nomic/,krea2/,laya/are standalone engines. They do not import one another.nn/andvk/are the shared kernels: quantized AVX2 and NEON, and Vulkan compute.compress/is the.golemformat: calibration, the pair-trellis codec, and the conversion pipeline.grammar/is GBNF and the JSON Schema converter that feeds it.internal/kyutai/is the Mimi codec, shared by both directions of speech.tensors/,token/,chat/,sample/,audio/,imageio/are the rest of the shared layer.ref/is what recorded each test fixture, and how to record it again.
go test ./... # correctness on every model and the card, about 17 minutesWeights are not in this repository, and every test that needs one skips cleanly when it cannot find it, which is why the same command is safe on a machine with no models and no card, and why CI runs it.
Use -p 1 if you have a GPU. Several packages put whole models on the card, and two at once want more than 16 GB. Anything slower than thirty seconds waits behind GOLEM_FULL_TEST=1; CONTRIBUTING.md says which tests those are and why.
The thing we need most is ARM benchmarks. The arm64 kernels are correct and tuned by nobody: written and verified under emulation, never once timed on real hardware. If you have Apple Silicon or a Graviton instance, run ./benchmark-arm.sh and share what it prints.
After that, bug reports (especially a parity test failing on your setup) and NEON tuning. CONTRIBUTING.md has the rest.
Golem is MIT Licensed.
Standing on the shoulders of giants: llama.cpp and ggml, Kyutai Pocket TTS, Google Gemma, Krea and ComfyUI, Convai Innovations' Laya, QTIP for the bitshift trellis under .golem, and FreeToken, which asked the expert-caching question first and answered it differently.