GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
-
Updated
Sep 20, 2026 - Python
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
A tiered-memory system design for workloads that don't fit in RAM: measure the working set, pin the hot tier, stream the cold tier from flash. Ships the residency calculator, measurement harnesses, and the build recipes behind it. Predictions validated against public benchmarks.
Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights
Qwen3.8-27B EXL3 (4.00 bpw) + DFlash2 speculative decoding for ExLlamaV3, validated at 262k context on a 24 GB RTX 3090
High-performance runtime extensions for vLLM.
Production-tuned deployment recipe: GLM-5.3-Flash 320B (EXL3 4-bit) on 2x NVIDIA DGX Spark - 850K context, DFlash2 + adaptive-k speculative decoding, full .env tuning and pitfalls
DeepSeek-V4-Flash-Vision-Exp (EXL3 MixedK, 256 experts, uncensored) on one NVIDIA DGX Spark with vLLM + sparkinfer: 245,760 context, vision + DSpark speculative decoding, CUDA graphs. Recipe, overlay patches, benchmarks, receipts.
LLM inference server for ExLlamaV3 / EXL3, with OpenAI- and Anthropic-compatible APIs optimized for Agent workloads.
LLM / AI inference server for Windows + NVIDIA (EXL3/ExLlamaV3). OpenAI-compatible API, multi-GPU, Blazor admin. Ollama-like, no Docker.
Qwen3.8-27B EXL3 on AMD gfx1201: RX 9070 XT (16 GB, tested) and Radeon AI PRO R9700 (32 GB). vLLM plugin with custom kernels: ~80 tok/s with MTP speculative decoding, int8 KV cache, 32k-64k context on 16 GB (up to 262k on 32 GB), vision and tool calls.
Run Qwen3.8-Flash-Next EXL3 (turboderp pack) on one AMD Strix Halo / Framework Desktop (gfx1151) via the exllamav3-amd runtime fork. 41 tok/s mean, 47.5 peak, MTP self-speculation. Setup encodes the five ROCm/gfx1151 traps.
Containerized private AI lab: TabbyAPI EXL3 + SillyTavern + Open WebUI + Ollama + SearXNG
GLM-5.3-Flash EXL3 4bpw on 2x NVIDIA DGX Spark (GB10): production recipe, boot ladder, quality gate, benchmarks and lessons learned — reproducible from CLAUDE.md/AGENTS.md
GLM-5.3 Flash on three DGX Sparks: eight concurrent 500k-token sessions with vision, a measured scheduler A/B against upstream, video enablement, reconstructed build, evidence and explicit limits. Fork and continuation of NNNtrance's recipe.
DeepSeek-V4.1-Flash at TP3 on 3x RTX PRO 6000 (SM120): EXL3 vs official checkpoint + UVA offload, same-box benchmarks, failed attempts, reproducible configs.
To associate your repository with the exl3 topic, visit your repo's landing page and select "manage topics."