Skip to content

Latest commit

 

History

38 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

NVIDIA Spark Model Testing Platform

Starter platform for testing and evaluating open source LLMs on NVIDIA DGX Spark (GB10 Grace Blackwell)

Optimized vLLM configurations for multiple models with automated performance testing. Designed to run directly on the Spark device (/opt/inference recommended) with AI-assisted management via Claude Code.


Quick Start

# 1. Clone to Spark device
cd /opt
git clone <repo-url> inference
cd inference

# 2. Setup environment
echo "HF_TOKEN=hf_your_token_here" > .env
sudo mkdir -p /opt/hf /opt/ollama
sudo chown -R $(id -u):$(id -g) /opt/hf /opt/ollama

# 3. Choose and start a model (only one vLLM service at a time)
# Speed Priority (8B-12B models):
docker compose up -d vllm-qwen3-8b-fp8        # ⚡ Fastest: ~10 tok/s
docker compose up -d vllm-llama31-8b-fp8      # NVIDIA-optimized: ~10 tok/s
docker compose up -d vllm-mistral-nemo-12b-fp8 # Long-context (65K): ~9 tok/s

# Quality Priority (27B-70B models):
docker compose up -d vllm-qwen38-27b-nvfp4     # Dense 27B reasoning, NVFP4 + MTP: ~28 tok/s code
docker compose up -d vllm-qwen35-35b-a3b-fp8   # ⚡ Fastest single-req: ~48 tok/s
docker compose up -d vllm-qwen3-30b-a3b-fp8   # MoE efficient: ~42 tok/s
docker compose up -d vllm-qwen3-32b-fp8        # Dense baseline: ~7 tok/s
docker compose up -d vllm-llama33-70b-fp8      # Max quality: ~6 tok/s

# Development/Testing (Ollama):
docker compose up -d ollama-qwen3-32b-fp8     # Port 11434: ~5-8 tok/s

# 4. Test inference (vLLM models on port 8000)
curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "Qwen/Qwen3-8B-FP8", "messages": [{"role": "user", "content": "Hello!"}]}'

Available Models

vLLM Models (Port 8000)

Run one vLLM model at a time - all share port 8000:

Service Model Type Context Best For
vllm-qwen3-8b-fp8 Qwen3-8B-FP8 Dense 8B 32K Max Speed ⚡ (~10 tok/s)
vllm-llama31-8b-fp8 Llama-3.1-8B-FP8 Dense 8B 32K NVIDIA-optimized (~10 tok/s)
vllm-mistral-nemo-12b-fp8 Mistral-NeMo-12B-FP8 Dense 12B 65K Long-context (128K native)
vllm-qwen3-32b-fp8 Qwen3-32B-FP8 Dense 32B 32K Balanced (~7 tok/s)
vllm-qwen3-30b-a3b-fp8 Qwen3-30B-A3B-FP8 MoE (3B active) 32K Efficient MoE (~42 tok/s)
vllm-qwen35-35b-a3b-fp8 Qwen3.5-35B-A3B-FP8 DeltaNet MoE (3B active) 32K Fastest + Reasoning (~48 tok/s)
vllm-qwen38-27b-nvfp4 Qwen3.8-27B-NVFP4 Dense 27B reasoning 128K Reasoning + tool-calls + MTP (~28 tok/s code, ~20 prose)
vllm-llama33-70b-fp8 Llama 3.3 70B-FP8 Dense 70B 65K Max Quality (~6 tok/s)

Ollama Model (Port 11434)

Can run alongside vLLM (different port):

Service Model Type Context Best For
ollama-qwen3-32b-fp8 Qwen3:32b-q8_0 Dense 32B 32K Development/Testing (~5-8 tok/s)

Performance estimates for single-request throughput on DGX Spark GB10

Commands

# Start model
docker compose up -d <service-name>

# View logs
docker compose logs -f <service-name>

# Stop model
docker compose stop <service-name>

# Stop all vLLM services
docker compose stop vllm-qwen3-8b-fp8 vllm-llama31-8b-fp8 vllm-mistral-nemo-12b-fp8 vllm-qwen3-32b-fp8 vllm-qwen3-30b-a3b-fp8 vllm-qwen35-35b-a3b-fp8 vllm-qwen38-27b-nvfp4 vllm-llama33-70b-fp8

# GPU status
nvidia-smi
docker exec <service-name> nvidia-smi

# vLLM metrics (for running service)
curl http://localhost:8000/metrics

Model Selection Guide

By Use Case

Maximum Speed & Throughput:

  • Qwen3-8B-FP8 - Fastest overall, best batched throughput (450-500 tok/s)
  • Llama-3.1-8B-FP8 - NVIDIA-optimized, equally fast with better instruction following

Long-Context Processing:

  • Mistral-NeMo-12B-FP8 - 65K configured (128K native), ideal for documents
  • Llama 3.3 70B-FP8 - 65K configured with maximum quality

Fastest Single-Request + Reasoning:

  • Qwen3.5-35B-A3B-FP8 - Hybrid DeltaNet MoE, ~48 tok/s, built-in thinking/reasoning, tool calling

Dense Reasoning with Tool Calling + Long Context:

  • Qwen3.8-27B-NVFP4 - Dense 27B, NVFP4 pre-quantized, MTP speculative decoding, native XML tool-calls, 128K context. Sits in the vLLM Recipes DGX Spark reference band.

Balanced Performance:

  • Qwen3-30B-A3B-FP8 - MoE architecture, efficient memory usage, ~42 tok/s
  • Qwen3-32B-FP8 - Dense baseline, proven performance

Maximum Quality:

  • Llama 3.3 70B-FP8 - Highest quality, complex reasoning, long-context analysis

Development/Testing:

  • Ollama Qwen3-32B-FP8 - Simple API, easy model management, runs on port 11434

Memory Footprint

Model Model Size KV Cache Total Memory
Qwen3-8B / Llama-3.1-8B ~8 GB ~80-85 GB ~88-93 GB
Mistral-NeMo-12B ~12 GB ~75-80 GB ~87-92 GB
Qwen3-30B-A3B ~30 GB ~55-70 GB ~85-100 GB
Qwen3.5-35B-A3B ~37.5 GB ~55-70 GB ~90-108 GB
Qwen3.8-27B-NVFP4 ~21 GB (NVFP4) ~22 GB (bf16 KV) ~45 GB @ util 0.60
Qwen3-32B ~32 GB ~66 GB ~98 GB
Llama 3.3 70B ~35 GB ~40-60 GB ~75-95 GB

Performance Testing

Automated CLI testing tool included:

Model-vs-Model Comparison (vLLM)

# Test all vLLM models
./tools/perftest-models.js

# Skip specific models
./tools/perftest-models.js --skip qwen3-32b-fp8
./tools/perftest-models.js --skip qwen3-32b-fp8,llama33-70b-fp8

# Custom iterations
./tools/perftest-models.js --iterations 5

Tests: Single-request latency, long-context handling (1K/8K/16K/32K tokens)


Latest Performance Results

Fast Chat Models (8B-12B) - NEW! ⚡

Test Date: 2025-11-11 | Hardware: DGX Spark GB10 (128GB unified memory)

Single-Request Latency (500 tokens output)

Model Architecture TTFT Tokens/sec Winner
Llama-3.1-8B Dense 8B 52ms 23.88 ⚡ FASTEST
Qwen3-8B Dense 8B 71ms 21.96 Close second
Mistral-NeMo-12B Dense 12B 85ms 15.66 Long-context specialist

Long-Context Performance (32K tokens input)

Model TTFT Total Time
Llama-3.1-8B 1,232ms 6.4s ⚡
Qwen3-8B 1,418ms 7.1s
Mistral-NeMo-12B 2,229ms 7.2s

Winner: Llama-3.1-8B-FP8 (NVIDIA-optimized) delivers fastest performance across all tests with 23.88 tok/s

Full report: docs/reports/model-comparison-2025-11-11T10-13-00-374Z.txt


Qwen3.8-27B-NVFP4 with MTP speculative decoding

Hardware: DGX Spark GB10 (128GB unified memory) | Image: vllm/vllm-openai:v0.24.0-ubuntu2404 | Config: MTP num_speculative_tokens=3, --gpu-memory-utilization 0.60, --max-model-len 131072, KV cache bf16

Single-Request Decode (256-token completions, temperature=0.1)

Content type Tokens/sec TTFT (ms)
math 26.07 273
code 28.38 273
prose 20.36 273

Concurrent Aggregate (prose prompt)

Concurrency Aggregate tokens/sec
c=4 62.29
c=8 129.16

Uplift vs same service with MTP off + kv-fp8 baseline: +75% prose, +144% code, +124% math on single-stream; +56% at c=8. TTFT roughly doubles because MTP runs the target model plus one draft step before the first token.

Sits at or above the vLLM Recipes DGX Spark NVFP4 published headline (~24.5 tok/s prose-weighted) on code and math; prose is slightly below because prose has lower spec-decoding acceptance than structured content. Full rationale, tuning table, and lever verdicts in docs/vllm/qwen38-27b-nvfp4.md.


Dense & MoE Models (30B-70B)

Test Date: 2025-11-09 | Hardware: DGX Spark GB10 (128GB unified memory)

Single-Request Latency (500 tokens output)

Model Architecture TTFT Tokens/sec Winner
Qwen3-30B-A3B MoE (3B active) 78ms 42.02 ⚡ FASTEST
Qwen3-32B Dense 32B 205ms 6.25 Baseline
Llama 3.3 70B Dense 70B 396ms 2.74 Highest quality

Long-Context Performance (32K tokens input)

Model TTFT Total Time
Qwen3-30B-A3B 881ms 4.1s ⚡
Qwen3-32B 5,798ms 23.8s
Llama 3.3 70B 10,333ms 49.0s

Winner: Qwen3-30B-A3B-FP8 (MoE) delivers 6.7x faster single-request performance than Dense 32B baseline

Full report: docs/reports/model-comparison-2025-11-09T17-11-44-162Z.txt


Documentation

Quick Reference

Model Configuration Guides

External Resources


AI-Assisted Management

This repository is optimized for Claude Code with:

  • CLAUDE.md - Comprehensive guidance for AI assistant
  • docs/ - Technical documentation for hardware, vLLM, and per-model configs
  • Enables Claude to configure, manage, and troubleshoot the Spark device

Simply open this repository in Claude Code to get intelligent assistance with model deployment, performance tuning, and troubleshooting.


Hardware

NVIDIA DGX Spark - GB10 Grace Blackwell Superchip

  • CPU: 20-core Arm (10x Cortex-X925 + 10x Cortex-A725)
  • GPU: NVIDIA Blackwell (6,144 CUDA cores, 5th Gen Tensor Cores)
  • Memory: 128 GB LPDDR5x unified (273 GB/s bandwidth)
  • Key Feature: Coherent unified memory (no separate VRAM)

Performance Characteristics:

  • Memory bandwidth (273 GB/s) is primary bottleneck
  • FP8 quantization essential for optimal performance
  • Single-request latency limited by hardware
  • Batching and MoE architectures maximize throughput

Full specs: docs/nvidia-spark.md


Repository Structure

.
├── README.md                      # This file
├── CLAUDE.md                      # AI assistant guidance (Claude Code)
├── docker-compose.yml             # vLLM & Ollama service definitions
├── .env                           # Environment variables (create this)
├── tools/
│   └── perftest-models.js         # Model-vs-model comparison
├── docs/
│   ├── nvidia-spark.md               # Hardware specs & setup
│   ├── vllm/                         # vLLM model configurations
│   │   ├── qwen3-8b-fp8.md           # 8B model (fastest)
│   │   ├── llama31-8b-fp8.md         # 8B NVIDIA-optimized
│   │   ├── mistral-nemo-12b-fp8.md   # 12B long-context
│   │   ├── qwen3-32b-fp8.md          # 32B baseline
│   │   ├── qwen3-30b-a3b-fp8.md      # 30B MoE
│   │   ├── qwen35-35b-a3b-fp8.md     # 35B DeltaNet MoE (fastest)
│   │   ├── qwen38-27b-nvfp4.md       # 27B dense reasoning (NVFP4 + MTP)
│   │   └── llama33-70b-fp8.md        # 70B max quality
│   ├── ollama/                       # Ollama configurations
│   │   └── qwen3-32b-fp8.md          # Provider comparison
│   └── reports/                      # Performance test results
└── models/
    └── ollama/
        └── Modelfile-qwen3-32b-fp8

Important Notes

  • One vLLM model at a time - All use port 8000
  • First load is slow - Model downloads: 8-15 minutes (cached locally afterward)
  • Memory bandwidth limited - 273 GB/s vs 900+ GB/s on datacenter GPUs
  • MoE/DeltaNet architecture wins - Qwen3.5-35B-A3B (~48 tok/s) and Qwen3-30B-A3B (~42 tok/s) significantly outperform dense models
  • FP8 quantization required - Essential for Spark's unified memory architecture
  • Batching increases throughput - Single-request performance is hardware-limited

License

Provided as-is for NVIDIA DGX Spark systems. Model licenses apply:


Built for NVIDIA DGX Spark GB10 | Optimized for unified memory architecture | Powered by vLLM

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages