a hypervisor is to a VM what a runtime is to a local LLM (or so i think). a runtime allows your LLM to utilize your hardware for max potency.
llama.cpp: run your models on pretty much anything.
ollama: lets you run local and cloud models from the CLI π
harnesses are what give your LLM agentic capabilities like executing code and interacting with the screen.
goose: desktop, cli, api. "use it for research, writing, automation, data analysis..."
hermes harness that prides itself on its ability to improve over time.
crush: agentic coding from charmbracelet. they have their own provider named hyper. π
opencode
pi: barebones, highly extensible coding agent
firecrawl: gives an LLM the ability to scrape the web
π: zero-data retention policy; privacy-focused