g
Pure Rust & CUDA LLM Inference
Universal Inference At Unimaginable Speeds
This fork also develops GLM-5.3-Flash NVFP4 on two DGX Sparks. Start with the fork release-candidate handoff for exact tested image/source pins, bounded profiles, current failures and rollback. The qualification bundles keep short-context concurrent MTP separate from ordinary long-context serving. The upstream quick starts and performance claims below do not qualify this GLM fork or its experimental context limits.
Atlas is a high-performance, pure Rust & CUDA LLM inference engine purpose-built for prosumer workstations (NVIDIA DGX Spark / GB10 SM121 and AMD Strix Halo). No Python, no PyTorch, no bloated dependency trees—just one compact binary with hand-tuned micro-kernels.
- Sub-90s First Token: Boots in seconds with cached weights; zero JIT compile or Python startup lag.
- Default Flagship Qwen 3.8 27B: Dense hybrid GDN + Attention running at 23.59 tok/s single-stream with MTP speculative decoding on a single GB10.
- Nemotron 3.5 Lightning + DSpark: Full bring-up of hybrid Mamba-2 SSM + MoE paired with DSpark speculative decoding drafters for sub-10ms token decode latencies.
- Qwen 3.8 Flash-Next Support: Stream massive ~180B hybrid MoE models inside ~90 GB resident VRAM using direct parallel
preadNVMe offloading. - Turnkey Sparkrun Integration: Launch any verified model recipe instantly with
sparkrun run @atlas/<recipe>. - OpenAI & Anthropic Compatible: Drop-in API endpoint supporting streaming, tool calling, and reasoning traces.
- MLPerf Proven: Official contributor to the MLPerf Inference v6.1 Edge Agentic benchmark.
The default flagship recipe deploys Qwen 3.8 27B in NVFP4 on a single GB10 with native MTP speculative decoding:
# Step 1: Install sparkrun
pip install sparkrun # or: uvx sparkrun setup install
# Step 2: Download weights (or let sparkrun fetch automatically)
huggingface-cli download unsloth/Qwen3.8-27B-NVFP4 \
--local-dir ~/.cache/huggingface/hub/models--unsloth--Qwen3.8-27B-NVFP4
# Step 3: Launch service (port 8888)
sparkrun run @atlas/qwen3.8-27b-nvfp4 --hosts localhostPrefer the one-line quickstart script?
curl -fsSL https://atlasinference.dev/quickstart.sh | shPairs the hybrid Mamba-2 + Attention + MoE backbone with NVIDIA's 6-layer DSpark drafter (gamma=4 / K=3 verify) for sub-10ms token generation:
# Export recommended performance environment
export ATLAS_DFLASH_OPTION_B=1
export ATLAS_NO_TOOL_INJECT=1 # +15.58 BFCL accuracy boost
# Launch Nemotron 3.5 Lightning with DSpark
sparkrun run @atlas/nemotron-3.5-lightning-30b-a3b-nvfp4-dspark --hosts localhostAtlas dynamically streams the 47.7 GB PLE n-gram table off NVMe using parallel pread workers, keeping peak resident memory under 90 GB on a single 128 GB GB10:
huggingface-cli download RadixArk/Qwen3.8-Flash-Next-NVFP4 \
--local-dir ~/.cache/huggingface/hub/models--RadixArk--Qwen3.8-Flash-Next-NVFP4
sparkrun run @atlas/qwen3.8-flash-next-nvfp4 --hosts localhostAtlas converts CUDA sources natively to AMD RDNA 3.5. Bring-up is active across two dedicated branches:
- Linux Leg (Ubuntu 24.04 / ROCm 6.2+): Branch
port/qwen3.8-strix-linux(PR #8) - Windows Leg (DirectX 12 / native MSVC): Branch
port/qwen3.8-windows(PR #9)
# Clone the AMD Strix Halo port branch
git clone -b port/qwen3.8-strix-linux https://github.com/Atlas-Inf/atlas.git
cd atlas
# Build native binary targeting AMD gfx1151 silicon
cargo build --release --features scale,rocm
# Launch Qwen 3.8 27B on Strix Halo
./target/release/spark serve unsloth/Qwen3.8-27B-NVFP4 --bind 127.0.0.1 --port 8888Atlas serves an OpenAI-compatible API on the designated port:
curl http://localhost:8888/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "atlas",
"messages": [{"role": "user", "content": "Explain quantum computing in three sentences."}],
"max_tokens": 256
}'Every recipe is maintained in the sparkrun-recipes SSOT repository and verified against committed gate baselines:
| Vendor | Family | Model Recipe | Quant | Topology | Highlights |
|---|---|---|---|---|---|
| Qwen | Qwen3.8 | @atlas/qwen3.8-27b-nvfp4 |
NVFP4 | Single GB10 | Default Flagship. Dense hybrid GDN + Attn, MTP spec decode, FP8 KV, 23.59 tok/s |
| Qwen | Qwen3.8 | @atlas/qwen3.8-27b-nvfp4-latency |
NVFP4 | Single GB10 | Low-concurrency / interactive profile tuned for minimal single-stream latency |
| Qwen | Qwen3.8 | @atlas/qwen3.8-27b-nvfp4-throughput |
NVFP4 | Single GB10 | Concurrency profile beating vLLM from 1 to 128 streams on GB10 |
| Nemotron | Nemotron-3.5 | @atlas/nemotron-3.5-lightning-30b-a3b-nvfp4-dspark |
NVFP4 | Single GB10 | New. Hybrid Mamba-2 SSM + MoE with 1.3 GB DSpark drafter (K=3 verify) |
| Nemotron | Nemotron-3 | @atlas/nemotron-3-nano-30b-a3b-nvfp4 |
NVFP4 | Single GB10 | 30B / 3B active Mamba-2 + MoE |
| Nemotron | Nemotron-3 | @atlas/nemotron-3-super-120b-a12b-nvfp4 |
NVFP4 | Single GB10 | 120B / 12B active hybrid architecture |
| Qwen | Qwen3.8 | @atlas/qwen3.8-flash-next-nvfp4 |
NVFP4 | Single GB10 | ~180B hybrid MoE, 8K context, parallel pread NVMe offload (750–800 tok/s prefill, 36.7 tok/s decode, ~90 GB resident) |
| Qwen | Qwen3.8 | @atlas/qwen3.8-flash-next-nvfp4-throughput |
NVFP4 | Single GB10 | Throughput-tuned 8K context profile |
| Qwen | Qwen3.6 | @atlas/qwen3.6-35b-a3b-fp8-mtp |
FP8 | Single GB10 | 35B / 3B active GDN + MoE + vision, MTP speculative |
| Gemma | Gemma-4 | @atlas/gemma-4-26b-a4b-nvfp4 |
NVFP4 | Single GB10 | 26B / 4B active MoE with GeGLU |
| DeepSeek | DeepSeek-V4 | @atlas/deepseek-v4-flash-nvfp4-ep2 |
NVFP4 | EP=2 (2 Sparks) | Dual-node Expert Parallelism |
Browse the interactive recipe browser at atlasinference.dev/#models.
- Double-Buffered Mamba-2 Chunked Scans: Hand-tuned SM121 PTX kernels delivering an 8.4x prefill latency reduction over generic vLLM implementations.
- Native FP4 Tensor Core Prefill GEMMs: Direct execution in Blackwell NVFP4 precision without dequantization overhead.
- PLE N-Gram NVMe Streaming (Direct Parallel
preadvs.mmap): Traditional engines (like baseline llama.cpp) suffer from thousands of scattered 4KBmmappage faults for tiny ~90-byte rows, stalling prefill at ~300 tok/s. Atlas implements an asynchronousO_DIRECTworker pool (ATLAS_PLE_FAULT_THREADS=32) using parallelpreaddirectly off NVMe storage (similar to the optimization in llama.cpp PR #28136). This delivers 750–800 tok/s cold prefill on DGX Spark (2.5x faster) and +20–32% on Strix Halo 128GB, completely bypassing OS page-cache faults while keeping the entire 47.7 GB n-gram table off RAM/VRAM. - Recurrent State Checkpoint & Rollback: Enables multi-token speculative decoding with DSpark on recurrent state models (Mamba-2 / GDN) without state divergence.
- TurboQuant+ KV Cache: Symmetric and asymmetric KV quantization (
bf16,fp8,nvfp4,turbo4) with Randomized Hadamard Rotation and Lloyd-Max codebooks.
| Flag | Bits/elem | Storage | Description |
|---|---|---|---|
--kv-cache-dtype bf16 |
16 | BF16 | Uncompressed baseline. Recommended for short-context or high-precision needs. |
--kv-cache-dtype fp8 |
8 | FP8 E4M3 | Default. Halves memory with minimal quality degradation across all benchmarks. |
--kv-cache-dtype nvfp4 |
4 | E2M1 | 4x compression vs BF16. Excellent for long context windows. |
--kv-cache-dtype turbo4 |
4 | E2M1 + WHT | ~2x lower reconstruction MSE than standard NVFP4 via Lloyd-Max codebooks. |
- Website: atlasinference.dev
- Discord: Join our Discord — Active daily development, live kernel tuning, and model requests.
- Recipes Repository: Atlas-Inf/sparkrun-recipes
- Deployment Guide: GB10 Deployment Guide
- Community Edition: Licensed under AGPLv3. Free and open for personal use, research, and non-commercial local deployments.
- Enterprise Edition: Commercial licensing for proprietary applications, SaaS hosting without AGPLv3 copyleft obligations, dedicated support, and custom hardware/kernel porting. Contact
debaterishaqui@gmail.com.
Continuity notice. Atlas is continuing. This repository, the Atlas-Inf GitHub organization, and atlasinference.dev are the replacement official Atlas channels. The existing website and GitHub repository remain disputed Atlas assets that have not been relinquished.
