Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
blackwell long-context fp8 local-llm llm-inference qwen speculative-decoding sglang qwen3 nixl sm120 nvfp4 rtx-pro-6000 hicache qwen3-8 qwen3-8-27b qwen38 dflash2 qwen3-8-flash-next fr-spec
-
Updated
Oct 1, 2026 - Python