You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Reproducible LLM inference on four NVIDIA DGX Spark (GB10) nodes wired as a direct switchless ConnectX ring — serving profiles, ring launchers, pinned ARM64 runtime images and measured results, newest tested model first.
Measured serving recipes for DeepSeek-V4.1-Flash on 4x NVIDIA DGX Spark (GB10): 1M context on vLLM (CUDA graphs, vision, tools, DSpark) and a switchless-ring SGLang TP4 lane, plus a cross-project reference table. EN + 中文.
NCCL over a switchless 4-node DGX Spark / GB10 ring using both PCIe halves of every QSFP cable: ~193 Gb/s per cable instead of ~112, up to +34% vLLM prefill. One patch on top of switchless-nccl.