You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A local app that rents a GPU in your own Colab account and serves Qwen3.8 on it: 27B on an A100-40G, or Flash-Next (125B-A6B MoE) on an A100-80G. The cost is shown before it starts; chat with images; Codex or any OpenAI client uses a fixed local /v1. Auto-stop, CU ledger, live prefill/decode. Windows, macOS, Linux.
Run a 177B model (Qwen3.8-Flash-Next) on an 8 GB laptop GPU. Measured: 6.6 GiB VRAM, 47.8 GiB RAM, 34-35 tok/s. The n-gram table stays on disk; the experts run on CPU.
Run Qwen3.8-Flash-Next EXL3 (turboderp pack) on one AMD Strix Halo / Framework Desktop (gfx1151) via the exllamav3-amd runtime fork. 41 tok/s mean, 47.5 peak, MTP self-speculation. Setup encodes the five ROCm/gfx1151 traps.
A llama-server-compatible HTTP front for ExLlamaV3: /props, /health, /slots and /v1/chat/completions with llama-server's timings object, so llama.cpp tooling drives an EXL3 model unchanged.
Blackwell (sm120) build of Mia'a AI Lab's Qwen3.8-27B EXL3 deployment kit. Measured on one RTX 5090: 193 tok/s decode, full native 262k context in 25.4 GiB.
Ollama-style CLI for ExLlamaV3 + TabbyAPI: pull, list, fit and rank local models with VRAM-aware auto-fit context for consumer NVIDIA GPUs. Serve and chat commands coming.