Skip to content

cuda: Q4_K mmvq at vdr 4 on Pascal (+ Q4_K geometry) -- 121 -> 81 us at n=1 - #1

Open
mewsian wants to merge 3 commits into
Kmic-68:p100-optimizationsfrom
mewsian:q4k-pascal-pr
Open

mewsian wants to merge 3 commits into
Kmic-68:p100-optimizationsfrom
mewsian:q4k-pascal-pr

Conversation

@mewsian

@mewsian mewsian commented Sep 25, 2026

Copy link
Copy Markdown

Q4_K decode matvec was the slowest common K-quant on sm_60 (121 us at m=4096 k=14336, vs 91 for q6_K and 81 for q4_0) despite reading fewer bytes than q6_K. This applies the same approach as your Q6_K work: the kernel is bound by memory instructions, so each lane now takes 4 quant ints (vdr 4) and pays the scale/min unpack, dm load and ds loads once per 32 values instead of once per 16. The min-term q8 sum is an exact integer byte sum instead of two emulated dp4a.

Second commit gives Q4_K its own Pascal launch geometry (the shared one was tuned on a Q6_K model): 2 warps x 4 rows at n=1, 2x8 at n=5 (the full MTP verify batch at --spec-draft-n-max 4), other n unchanged. Other types are untouched.

Kernel, us/run at m4096 k14336

n=1 n=2 n=5 n=8
before 121.0 148.4 317.0 429.6
vdr 4 87.6 132.5 270.4 385.9
+ geometry 80.9 248.1

llama-bench tg256, -sm tensor, q8_0 KV, 2x P100

  • gemma-4-31B i1-Q4_K_M: 21.67 -> 24.72 t/s
  • Qwen3.8-27B i1-Q4_K_M: 25.18 -> 28.93 t/s

Accuracy (llama-perplexity --kl-divergence -ub 8, so it goes through mmvq)

  • Qwen3.8 Q4_K_M: PPL 3.4920 vs 3.4917, KLD 0.0007, 98.9% same top token
  • gemma-4-31B Q4_K_M: KLD 0.0078, but mainline vs the same reference is 0.0077; a control build (vdr 2, exact sum) reproduces the reference at KLD 0. Integer sums are exact; only float summation order moves.

Third commit keeps vdr 4 to sm_60-only builds (same __CUDA_ARCH_LIST__ == 600 test as the Pascal geometry). It was only measured on sm_60, where dp4a is emulated; sm_61+ has native dp4a, so other builds keep upstream's vdr 2.

test-backend-ops MUL_MAT (64) and MUL_MAT_ID (84) q4_K cases pass on the P100.

Compile-time switches are left in for A/B (P100_Q4K_VDR, P100_Q4K_DSMIN, P100_Q4K_{NWARPS,ROWS}_*). DSMIN=1 (min term from q8_1 ds.y) was 9% faster but is not exact (KLD 0.019 on gemma), so it is off. Happy to strip those if you'd rather not carry them.

Details are in the commit messages. Thanks for the fork, the Q6_K work made this one straightforward.

🤖 Generated with Claude Code

@rankaiyx

Copy link
Copy Markdown

Let me take this opportunity to share a tip:
If the vision component is not frequently used, the weights can be stored in system memory instead of video memory.

@Kmic-68

Kmic-68 commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Hey! Will check this out at some point. May even download every quant and tune. May be some time until I get around to it

mewsian and others added 3 commits September 29, 2026 22:27
The Q4_K decode matvec was the slowest common K-quant on sm_60 (121 us at m=4096 k=14336
against 91 for q6_K and 81 for q4_0) despite reading a third fewer bytes than q6_K. Same fix
as the fork's Q6_K work: this kernel is bound by memory instructions, so each lane now takes
4 of a sub-block pair's 8 quant ints (two 8-byte loads) and pays the 6-bit scale/min unpack,
the dm load and the ds loads once per 32 values instead of once per 16. The min-term sum of
q8 values is an exact integer SIMD byte sum instead of two emulated dp4a.

  kernel us/run   n=1     n=2     n=5     n=8
  before          121.0   148.4   317.0   429.6
  after            87.6   132.5   270.4   385.9

  llama-bench tg256, -sm tensor, q8_0 KV (fork build, interleaved):
    gemma-4-31B heretic i1-Q4_K_M   21.67 -> 24.72 t/s
    Qwen3.8-27B heretic i1-Q4_K_M   25.18 -> 28.93 t/s

Accuracy (llama-perplexity --kl-divergence at -ub 8, which routes through mmvq; the default
large batch goes to MMQ and would not exercise this):
  Qwen3.8 Q4_K_M: PPL 3.4920 vs 3.4917, KLD 0.0007, 98.9% same top token.
  gemma-4-31B Q4_K_M: KLD 0.0078 -- but mainline vs the same reference is 0.0077, and a
  control build (vdr 2, exact sum) reproduces the reference at KLD 0.000000. All integer sums
  are exact; only float summation order moved. Gemma on raw text (PPL ~65) amplifies that.

Kept as compile-time switches for A/B: P100_Q4K_VDR (4, or 2 for upstream's layout) and, at
vdr 2, P100_Q4K_DSMIN. DSMIN=1 (min term from q8_1 ds.y) was 9% faster but is rejected:
ds.y is the sum of the unquantized activations, which does not match the quantized dot
(KLD 0.019 on gemma). DSMIN=2 (exact byte sum at vdr 2) measured no gain.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Pascal launch geometry was tuned against a Q6_K model and shared by every type. Q4_K
now has its own (calc_rows_per_block takes the type for this), swept at m4096 k14336 on
top of the vdr-4 kernel:

  n=1: 2 warps x 4 rows instead of 2x2          89.2 -> 80.9 us
       (1x2 92.2, 2x1 116.6, 4x1 121.0, 4x2 92.2, 1x4 86.8, 4x4 83.4, 2x8 96.6, 1x8 104.0)
  n=5: 2x8 instead of the shared 4x16           270.4 -> 248.1 us
       The shared geometry steps up between n=4 (193) and n=5 (270). 2x8 is worse at
       n=3 (171 vs 162), n=4 (209 vs 193) and n=8 (419 vs 386), so it is used at n=5 only --
       the full verify batch at --spec-draft-n-max 4.
  n=2..4, 6..8: unchanged (4x16).

Macros P100_Q4K_{NWARPS,ROWS}_{1,N,N_LO} for re-sweeping; other types are untouched.
test-backend-ops MUL_MAT (64) and MUL_MAT_ID (84) q4_K cases pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
P100_Q4K_VDR defaulted to 4 for every arch. It was tuned on sm_60, where dp4a is emulated and the
exact byte sum beats it; sm_61+ has native dp4a and was never measured. Default to upstream's
vdr 2 unless the build targets sm_60 only, the same __CUDA_ARCH_LIST__ test the Pascal geometry uses.

Checked: nvcc sm_60 -> 4, sm_61 -> 2, sm_60+sm_86 -> 2. P100 build unchanged: q4_K m4096 k14336
n=1 80.8 us; test-backend-ops MUL_MAT (64) and MUL_MAT_ID (84) q4_K pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants