Skip to content
#

nvfp4

Here are 250 public repositories matching this topic...

Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.

  • Updated Sep 15, 2026
  • Python

Serving 4-bit Qwen3.8-27B on a single DGX Spark (GB10): 75 tok/s single-stream, 246 tok/s aggregate at 8-way concurrency. NVFP4 vs MixedInt4-AutoRound vs the FP8 baseline, measured on one harness — including why the quantization advantage collapses to +0.2% by c16.

  • Updated Sep 29, 2026
  • Python

LLM server for one RTX 5090. Runs Qwen3.8-Flash-Next (512-expert MoE) on a single 32 GB card at 75-80 tok/s. Decodes 17-191 % faster than llama.cpp on the same models. Built for agents: tool calls, long context, many streams at once.

  • Updated Oct 1, 2026
  • C++

Add this topic to your repo

To associate your repository with the nvfp4 topic, visit your repo's landing page and select "manage topics."

Learn more