Typed questions in, calibrated answers out, in milliseconds.
One binary, models pulled by name and a TypeSafe-compatible API, the way Ollama runs LLMs.
Website · Models · Docs · Releases · Hugging Face
curl -fsSL https://ollaya.dev/install.sh | sh
ollaya run laya --preset triage "I was charged twice this month and want a refund."| Model | Author | |
|---|---|---|
laya |
Convai Innovations | The fastest: 8–10 ms for five questions on an RTX 4090. English and 100+ languages. |
decider |
Mapika | The most accurate: Qwen3.5 decoders, 2B and 0.8B. |
nli |
Moritz Laurer | Zero-shot NLI classifiers, the most accurate encoder. |
gliclass |
Knowledgator | Instruction-following zero-shot classifier. |
Weights always come from their authors' own Hugging Face repositories, pinned to a commit and verified by sha256. Ollaya never re-hosts them.
Apache-2.0. Ollaya is an independent project, not affiliated with Ollama or TypeSafe.