Deploy DeepSeek-V4-Flash-0731 on dual RTX PRO 6000 GPUs with vLLM and DSpark speculative decoding, achieving ~200-227 tok/s offline.