Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
-
Updated
Sep 30, 2026 - Python
Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt; prefill and agent 128K profiles), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode (64 GB profile).
Qwen3.8-Flash-Next on 4x RTX 3090 with vLLM: 806,792 kv pool, 160-175tok/s single decode, 3 x 262K sessions resident, MTP, host-mapped PLE, pinned build and container recipes.
Qwen3.8-27B at 262K context on dual or quad RTX 3090s: stock vLLM 0.30.0 + 5 patches, FP8 KV pool 794K (dual) / 1.95M (quad), MTP K=3
Swift 1.5 post-train of Qwen3.8-Flash-Next on 4x RTX 3090 with vLLM: W4A16 checkpoint on the unchanged Flash-Next v2.2.0 engine, TP2 x PP2 + EP, 806,792-token FP8 KV pool, 262K context, MTP K=3, recalibrated KV scales.
Run Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on the CPU, INT8 KV cache, Docker, OpenAI-compatible API.
MTP (Model-Translation-Postprocess) is a ground-breaking technology designed to enhance the capabilities of Qwen, a sophisticated language model
What is Qwen3
A tower defense made in Three js to test Qwen3.8-27B performance ( hf download hf://unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q3_K_XL.gguf , 128K context )
Run Qwen3.8-27B locally in Claude Code Desktop with vision, tools, native reasoning and Claude Code CLI support. Validated on an NVIDIA RTX 5090 32 GB.
Qwen3
Run Qwen 3.8 27B locally with uncensored chat via Ollama, LM Studio, or llama.cpp — no API key needed.
In the realm of artificial intelligence, the Qwen3
To associate your repository with the qwen38 topic, visit your repo's landing page and select "manage topics."