|
When running local language models like Qwen 2.5 1.5B or Phi-4 mini on Android devices, what are the primary throughput and power efficiency trade-offs between utilizing the Hexagon NPU via LiteRT delegate versus CPU int4 multithreaded GEMM execution? |
Answered by
PrinceBad
Sep 7, 2026
Replies: 1 comment
Technical Analysis: NPU (Hexagon via LiteRT) vs. CPU int4 Multithreading
|
0 replies
Answer selected by
PrinceBad
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Technical Analysis: NPU (Hexagon via LiteRT) vs. CPU int4 Multithreading
Tokens Per Second (Prefill & Decode):
Thermal & Energy Footprint: