Skip to content
Discussion options

You must be logged in to vote

Technical Analysis: NPU (Hexagon via LiteRT) vs. CPU int4 Multithreading

  1. Tokens Per Second (Prefill & Decode):

    • Hexagon NPU: Provides massive parallel compute for matrix-vector multiplications during decode phase, achieving up to 2.5x - 4x lower latency per token (18-28 tok/sec on Snapdragon 8 Gen 2/3) compared to CPU.
    • Kryo/Cortex CPU: Bound by memory bandwidth and thermal throttling; decode hovers around 6-11 tok/sec.
  2. Thermal & Energy Footprint:

    • LiteRT with NPU offload operates within a strict DSP power budget (~2.5W - 3.8W), avoiding thermal runaway during sustained reasoning runs.
    • Multithreaded CPU saturation causes aggressive thermal throttling after 40-60 seconds of continuous …

Replies: 1 comment

Comment options

PrinceBad
Sep 7, 2026
Maintainer Author

You must be logged in to vote
0 replies
Answer selected by PrinceBad
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
1 participant