Skip to content
#

gguf-runner

Here are 7 public repositories matching this topic...

Run GGUF through llama.cpp and SafeTensors through vLLM behind one OpenAI-compatible endpoint. Your coding tools select a model; the switchboard manages the local runtime, process, and resident-model change.

  • Updated Sep 15, 2026
  • Rust

Render is a lightweight, easy-to-use CLI and local server for generating AI images. Run Stable Diffusion (SD 1.5, SDXL) and FLUX models locally with simple commands like pull and run. Powered by Vulkan acceleration and GGUF support for ultra-fast performance. Think Ollama, but for local image generation.

  • Updated Aug 31, 2026
  • Go

A complete inference runtime for open-weight large language models, enabling efficient execution through streaming weights, quantization, and memory-aware scheduling

  • Updated Aug 14, 2026
  • Python

Add this topic to your repo

To associate your repository with the gguf-runner topic, visit your repo's landing page and select "manage topics."

Learn more