Local inference platform for K/IQ-quant GGUF models on Apple Silicon
-
Updated
Sep 16, 2026 - Python
Local inference platform for K/IQ-quant GGUF models on Apple Silicon
Run GGUF through llama.cpp and SafeTensors through vLLM behind one OpenAI-compatible endpoint. Your coding tools select a model; the switchboard manages the local runtime, process, and resident-model change.
a GGUF inference runner in Rust
GGUF-Runner - Want to run LLMs locally, use this guide, and run with LLAMA.cpp
Render is a lightweight, easy-to-use CLI and local server for generating AI images. Run Stable Diffusion (SD 1.5, SDXL) and FLUX models locally with simple commands like pull and run. Powered by Vulkan acceleration and GGUF support for ultra-fast performance. Think Ollama, but for local image generation.
Autonomous, Sovereign & Zero-Cloud Local Generative AI Workstation
A complete inference runtime for open-weight large language models, enabling efficient execution through streaming weights, quantization, and memory-aware scheduling
To associate your repository with the gguf-runner topic, visit your repo's landing page and select "manage topics."