Apple Neural Engine (ANE/NPU) vs GPU for local LLM inference on Apple Silicon / macOS. Core ML (CoreML), Core AI (CoreAI), MLX and Metal benchmarks: prefill, throughput, latency, memory, thermals, INT4/INT8, W4A16/A8W4 quantization, grouped scales and FP16 arithmetic. Reproducible component tests, compatibility findings, English/Chinese articles.
-
Updated
Sep 17, 2026 - Python