The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
macos metal mtp mlx inference-engine apple-silicon local-llm llm-inference local-ai qwen speculative-decoding openai-compatible claude-code qwen3-next anthropic-compatible native-mtp mtplx qwen3-8 qwen3-8-flash-next flash-next
-
Updated
Sep 19, 2026 - Python