SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
-
Updated
Sep 30, 2026 - Python
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Pure Rust Inference Engine
Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.
Fastest measured Qwen3.8-27B config for DGX Spark (GB10): SGLang + NVFP4 + DFlash2, 72 tok/s greedy median, 1M context, deterministic boots. Flash-Next 176B on one box. One command, everything pinned, and a local cockpit to drive it.
Hand-written NVFP4 W4A16 CUDA kernels for Volta
GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
AdaLLM is an NVFP4-first inference runtime for Ada Lovelace (RTX 4090) with FP8 KV cache and custom decode kernels. This repo targets NVFP4 weights and keeps the entire decode path in FP8
Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs + Flux 2 Dev / LTX 2.3 22B / ACE-Step v1.5 XL Turbo pre-bundled with abliterated text-encoder paths.
Serving 4-bit Qwen3.8-27B on a single DGX Spark (GB10): 75 tok/s single-stream, 246 tok/s aggregate at 8-way concurrency. NVFP4 vs MixedInt4-AutoRound vs the FP8 baseline, measured on one harness — including why the quantization advantage collapses to +0.2% by c16.
Native Windows port of NInfer engine for RTX 5090. Features Qwen3.8-27B with QUASAR and NInfer models, MTP/DFlash2 with vision and 262,144 context
LLM server for one RTX 5090. Runs Qwen3.8-Flash-Next (512-expert MoE) on a single 32 GB card at 75-80 tok/s. Decodes 17-191 % faster than llama.cpp on the same models. Built for agents: tool calls, long context, many streams at once.
Crazy Fast Qwen3.8–27b With a 420k Context on an RTX 5090
A production-ready Docker setup for ComfyUI that unlocks the full potential of NVIDIA Blackwell GPUs (RTX 50 series) through 4-bit quantization with NVFP4.
To associate your repository with the nvfp4 topic, visit your repo's landing page and select "manage topics."