Starter platform for testing and evaluating open source LLMs on NVIDIA DGX Spark (GB10 Grace Blackwell)
Optimized vLLM configurations for multiple models with automated performance testing. Designed to run directly on the Spark device (/opt/inference recommended) with AI-assisted management via Claude Code.
# 1. Clone to Spark device
cd /opt
git clone <repo-url> inference
cd inference
# 2. Setup environment
echo "HF_TOKEN=hf_your_token_here" > .env
sudo mkdir -p /opt/hf /opt/ollama
sudo chown -R $(id -u):$(id -g) /opt/hf /opt/ollama
# 3. Choose and start a model (only one vLLM service at a time)
# Speed Priority (8B-12B models):
docker compose up -d vllm-qwen3-8b-fp8 # ⚡ Fastest: ~10 tok/s
docker compose up -d vllm-llama31-8b-fp8 # NVIDIA-optimized: ~10 tok/s
docker compose up -d vllm-mistral-nemo-12b-fp8 # Long-context (65K): ~9 tok/s
# Quality Priority (27B-70B models):
docker compose up -d vllm-qwen38-27b-nvfp4 # Dense 27B reasoning, NVFP4 + MTP: ~28 tok/s code
docker compose up -d vllm-qwen35-35b-a3b-fp8 # ⚡ Fastest single-req: ~48 tok/s
docker compose up -d vllm-qwen3-30b-a3b-fp8 # MoE efficient: ~42 tok/s
docker compose up -d vllm-qwen3-32b-fp8 # Dense baseline: ~7 tok/s
docker compose up -d vllm-llama33-70b-fp8 # Max quality: ~6 tok/s
# Development/Testing (Ollama):
docker compose up -d ollama-qwen3-32b-fp8 # Port 11434: ~5-8 tok/s
# 4. Test inference (vLLM models on port 8000)
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "Qwen/Qwen3-8B-FP8", "messages": [{"role": "user", "content": "Hello!"}]}'Run one vLLM model at a time - all share port 8000:
| Service | Model | Type | Context | Best For |
|---|---|---|---|---|
vllm-qwen3-8b-fp8 |
Qwen3-8B-FP8 | Dense 8B | 32K | Max Speed ⚡ (~10 tok/s) |
vllm-llama31-8b-fp8 |
Llama-3.1-8B-FP8 | Dense 8B | 32K | NVIDIA-optimized (~10 tok/s) |
vllm-mistral-nemo-12b-fp8 |
Mistral-NeMo-12B-FP8 | Dense 12B | 65K | Long-context (128K native) |
vllm-qwen3-32b-fp8 |
Qwen3-32B-FP8 | Dense 32B | 32K | Balanced (~7 tok/s) |
vllm-qwen3-30b-a3b-fp8 |
Qwen3-30B-A3B-FP8 | MoE (3B active) | 32K | Efficient MoE (~42 tok/s) |
vllm-qwen35-35b-a3b-fp8 |
Qwen3.5-35B-A3B-FP8 | DeltaNet MoE (3B active) | 32K | Fastest + Reasoning (~48 tok/s) |
vllm-qwen38-27b-nvfp4 |
Qwen3.8-27B-NVFP4 | Dense 27B reasoning | 128K | Reasoning + tool-calls + MTP (~28 tok/s code, ~20 prose) |
vllm-llama33-70b-fp8 |
Llama 3.3 70B-FP8 | Dense 70B | 65K | Max Quality (~6 tok/s) |
Can run alongside vLLM (different port):
| Service | Model | Type | Context | Best For |
|---|---|---|---|---|
ollama-qwen3-32b-fp8 |
Qwen3:32b-q8_0 | Dense 32B | 32K | Development/Testing (~5-8 tok/s) |
Performance estimates for single-request throughput on DGX Spark GB10
# Start model
docker compose up -d <service-name>
# View logs
docker compose logs -f <service-name>
# Stop model
docker compose stop <service-name>
# Stop all vLLM services
docker compose stop vllm-qwen3-8b-fp8 vllm-llama31-8b-fp8 vllm-mistral-nemo-12b-fp8 vllm-qwen3-32b-fp8 vllm-qwen3-30b-a3b-fp8 vllm-qwen35-35b-a3b-fp8 vllm-qwen38-27b-nvfp4 vllm-llama33-70b-fp8
# GPU status
nvidia-smi
docker exec <service-name> nvidia-smi
# vLLM metrics (for running service)
curl http://localhost:8000/metricsMaximum Speed & Throughput:
- Qwen3-8B-FP8 - Fastest overall, best batched throughput (450-500 tok/s)
- Llama-3.1-8B-FP8 - NVIDIA-optimized, equally fast with better instruction following
Long-Context Processing:
- Mistral-NeMo-12B-FP8 - 65K configured (128K native), ideal for documents
- Llama 3.3 70B-FP8 - 65K configured with maximum quality
Fastest Single-Request + Reasoning:
- Qwen3.5-35B-A3B-FP8 - Hybrid DeltaNet MoE, ~48 tok/s, built-in thinking/reasoning, tool calling
Dense Reasoning with Tool Calling + Long Context:
- Qwen3.8-27B-NVFP4 - Dense 27B, NVFP4 pre-quantized, MTP speculative decoding, native XML tool-calls, 128K context. Sits in the vLLM Recipes DGX Spark reference band.
Balanced Performance:
- Qwen3-30B-A3B-FP8 - MoE architecture, efficient memory usage, ~42 tok/s
- Qwen3-32B-FP8 - Dense baseline, proven performance
Maximum Quality:
- Llama 3.3 70B-FP8 - Highest quality, complex reasoning, long-context analysis
Development/Testing:
- Ollama Qwen3-32B-FP8 - Simple API, easy model management, runs on port 11434
| Model | Model Size | KV Cache | Total Memory |
|---|---|---|---|
| Qwen3-8B / Llama-3.1-8B | ~8 GB | ~80-85 GB | ~88-93 GB |
| Mistral-NeMo-12B | ~12 GB | ~75-80 GB | ~87-92 GB |
| Qwen3-30B-A3B | ~30 GB | ~55-70 GB | ~85-100 GB |
| Qwen3.5-35B-A3B | ~37.5 GB | ~55-70 GB | ~90-108 GB |
| Qwen3.8-27B-NVFP4 | ~21 GB (NVFP4) | ~22 GB (bf16 KV) | ~45 GB @ util 0.60 |
| Qwen3-32B | ~32 GB | ~66 GB | ~98 GB |
| Llama 3.3 70B | ~35 GB | ~40-60 GB | ~75-95 GB |
Automated CLI testing tool included:
# Test all vLLM models
./tools/perftest-models.js
# Skip specific models
./tools/perftest-models.js --skip qwen3-32b-fp8
./tools/perftest-models.js --skip qwen3-32b-fp8,llama33-70b-fp8
# Custom iterations
./tools/perftest-models.js --iterations 5Tests: Single-request latency, long-context handling (1K/8K/16K/32K tokens)
Test Date: 2025-11-11 | Hardware: DGX Spark GB10 (128GB unified memory)
Single-Request Latency (500 tokens output)
| Model | Architecture | TTFT | Tokens/sec | Winner |
|---|---|---|---|---|
| Llama-3.1-8B | Dense 8B | 52ms | 23.88 | ⚡ FASTEST |
| Qwen3-8B | Dense 8B | 71ms | 21.96 | Close second |
| Mistral-NeMo-12B | Dense 12B | 85ms | 15.66 | Long-context specialist |
Long-Context Performance (32K tokens input)
| Model | TTFT | Total Time |
|---|---|---|
| Llama-3.1-8B | 1,232ms | 6.4s ⚡ |
| Qwen3-8B | 1,418ms | 7.1s |
| Mistral-NeMo-12B | 2,229ms | 7.2s |
Winner: Llama-3.1-8B-FP8 (NVIDIA-optimized) delivers fastest performance across all tests with 23.88 tok/s
Full report: docs/reports/model-comparison-2025-11-11T10-13-00-374Z.txt
Hardware: DGX Spark GB10 (128GB unified memory) | Image: vllm/vllm-openai:v0.24.0-ubuntu2404 | Config: MTP num_speculative_tokens=3, --gpu-memory-utilization 0.60, --max-model-len 131072, KV cache bf16
Single-Request Decode (256-token completions, temperature=0.1)
| Content type | Tokens/sec | TTFT (ms) |
|---|---|---|
| math | 26.07 | 273 |
| code | 28.38 | 273 |
| prose | 20.36 | 273 |
Concurrent Aggregate (prose prompt)
| Concurrency | Aggregate tokens/sec |
|---|---|
| c=4 | 62.29 |
| c=8 | 129.16 |
Uplift vs same service with MTP off + kv-fp8 baseline: +75% prose, +144% code, +124% math on single-stream; +56% at c=8. TTFT roughly doubles because MTP runs the target model plus one draft step before the first token.
Sits at or above the vLLM Recipes DGX Spark NVFP4 published headline (~24.5 tok/s prose-weighted) on code and math; prose is slightly below because prose has lower spec-decoding acceptance than structured content. Full rationale, tuning table, and lever verdicts in docs/vllm/qwen38-27b-nvfp4.md.
Test Date: 2025-11-09 | Hardware: DGX Spark GB10 (128GB unified memory)
Single-Request Latency (500 tokens output)
| Model | Architecture | TTFT | Tokens/sec | Winner |
|---|---|---|---|---|
| Qwen3-30B-A3B | MoE (3B active) | 78ms | 42.02 | ⚡ FASTEST |
| Qwen3-32B | Dense 32B | 205ms | 6.25 | Baseline |
| Llama 3.3 70B | Dense 70B | 396ms | 2.74 | Highest quality |
Long-Context Performance (32K tokens input)
| Model | TTFT | Total Time |
|---|---|---|
| Qwen3-30B-A3B | 881ms | 4.1s ⚡ |
| Qwen3-32B | 5,798ms | 23.8s |
| Llama 3.3 70B | 10,333ms | 49.0s |
Winner: Qwen3-30B-A3B-FP8 (MoE) delivers 6.7x faster single-request performance than Dense 32B baseline
Full report: docs/reports/model-comparison-2025-11-09T17-11-44-162Z.txt
- CLAUDE.md - AI assistant guide (for Claude Code)
- Hardware Specs - DGX Spark GB10 specifications
- Qwen3-8B-FP8 - Fastest 8B model (recommended)
- Llama-3.1-8B-FP8 - NVIDIA-optimized 8B
- Mistral-NeMo-12B-FP8 - Long-context specialist (65K/128K)
- Qwen3-32B-FP8 - Dense baseline
- Qwen3-30B-A3B-FP8 - MoE model
- Qwen3.5-35B-A3B-FP8 - Hybrid DeltaNet MoE (fastest + reasoning)
- Qwen3.8-27B-NVFP4 - Dense 27B reasoning, NVFP4 + MTP, native XML tool-calls, 128K
- Llama 3.3 70B-FP8 - Maximum quality
- Ollama Qwen3-32B - Provider comparison
- vLLM Docs: https://docs.vllm.ai
- Ollama Docs: https://docs.ollama.com
- DGX Spark: https://docs.nvidia.com/dgx/dgx-spark/
- Qwen Models: https://huggingface.co/Qwen
- Llama Models: https://huggingface.co/meta-llama
This repository is optimized for Claude Code with:
- CLAUDE.md - Comprehensive guidance for AI assistant
- docs/ - Technical documentation for hardware, vLLM, and per-model configs
- Enables Claude to configure, manage, and troubleshoot the Spark device
Simply open this repository in Claude Code to get intelligent assistance with model deployment, performance tuning, and troubleshooting.
NVIDIA DGX Spark - GB10 Grace Blackwell Superchip
- CPU: 20-core Arm (10x Cortex-X925 + 10x Cortex-A725)
- GPU: NVIDIA Blackwell (6,144 CUDA cores, 5th Gen Tensor Cores)
- Memory: 128 GB LPDDR5x unified (273 GB/s bandwidth)
- Key Feature: Coherent unified memory (no separate VRAM)
Performance Characteristics:
- Memory bandwidth (273 GB/s) is primary bottleneck
- FP8 quantization essential for optimal performance
- Single-request latency limited by hardware
- Batching and MoE architectures maximize throughput
Full specs: docs/nvidia-spark.md
.
├── README.md # This file
├── CLAUDE.md # AI assistant guidance (Claude Code)
├── docker-compose.yml # vLLM & Ollama service definitions
├── .env # Environment variables (create this)
├── tools/
│ └── perftest-models.js # Model-vs-model comparison
├── docs/
│ ├── nvidia-spark.md # Hardware specs & setup
│ ├── vllm/ # vLLM model configurations
│ │ ├── qwen3-8b-fp8.md # 8B model (fastest)
│ │ ├── llama31-8b-fp8.md # 8B NVIDIA-optimized
│ │ ├── mistral-nemo-12b-fp8.md # 12B long-context
│ │ ├── qwen3-32b-fp8.md # 32B baseline
│ │ ├── qwen3-30b-a3b-fp8.md # 30B MoE
│ │ ├── qwen35-35b-a3b-fp8.md # 35B DeltaNet MoE (fastest)
│ │ ├── qwen38-27b-nvfp4.md # 27B dense reasoning (NVFP4 + MTP)
│ │ └── llama33-70b-fp8.md # 70B max quality
│ ├── ollama/ # Ollama configurations
│ │ └── qwen3-32b-fp8.md # Provider comparison
│ └── reports/ # Performance test results
└── models/
└── ollama/
└── Modelfile-qwen3-32b-fp8
- One vLLM model at a time - All use port 8000
- First load is slow - Model downloads: 8-15 minutes (cached locally afterward)
- Memory bandwidth limited - 273 GB/s vs 900+ GB/s on datacenter GPUs
- MoE/DeltaNet architecture wins - Qwen3.5-35B-A3B (~48 tok/s) and Qwen3-30B-A3B (~42 tok/s) significantly outperform dense models
- FP8 quantization required - Essential for Spark's unified memory architecture
- Batching increases throughput - Single-request performance is hardware-limited
Provided as-is for NVIDIA DGX Spark systems. Model licenses apply:
- Qwen Models: Tongyi Qianwen License
- Llama Models: Llama 3 Community License
- Mistral Models: Apache 2.0 License
Built for NVIDIA DGX Spark GB10 | Optimized for unified memory architecture | Powered by vLLM