Conversation
|
Let me take this opportunity to share a tip: |
Owner
|
Hey! Will check this out at some point. May even download every quant and tune. May be some time until I get around to it |
The Q4_K decode matvec was the slowest common K-quant on sm_60 (121 us at m=4096 k=14336
against 91 for q6_K and 81 for q4_0) despite reading a third fewer bytes than q6_K. Same fix
as the fork's Q6_K work: this kernel is bound by memory instructions, so each lane now takes
4 of a sub-block pair's 8 quant ints (two 8-byte loads) and pays the 6-bit scale/min unpack,
the dm load and the ds loads once per 32 values instead of once per 16. The min-term sum of
q8 values is an exact integer SIMD byte sum instead of two emulated dp4a.
kernel us/run n=1 n=2 n=5 n=8
before 121.0 148.4 317.0 429.6
after 87.6 132.5 270.4 385.9
llama-bench tg256, -sm tensor, q8_0 KV (fork build, interleaved):
gemma-4-31B heretic i1-Q4_K_M 21.67 -> 24.72 t/s
Qwen3.8-27B heretic i1-Q4_K_M 25.18 -> 28.93 t/s
Accuracy (llama-perplexity --kl-divergence at -ub 8, which routes through mmvq; the default
large batch goes to MMQ and would not exercise this):
Qwen3.8 Q4_K_M: PPL 3.4920 vs 3.4917, KLD 0.0007, 98.9% same top token.
gemma-4-31B Q4_K_M: KLD 0.0078 -- but mainline vs the same reference is 0.0077, and a
control build (vdr 2, exact sum) reproduces the reference at KLD 0.000000. All integer sums
are exact; only float summation order moved. Gemma on raw text (PPL ~65) amplifies that.
Kept as compile-time switches for A/B: P100_Q4K_VDR (4, or 2 for upstream's layout) and, at
vdr 2, P100_Q4K_DSMIN. DSMIN=1 (min term from q8_1 ds.y) was 9% faster but is rejected:
ds.y is the sum of the unquantized activations, which does not match the quantized dot
(KLD 0.019 on gemma). DSMIN=2 (exact byte sum at vdr 2) measured no gain.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Pascal launch geometry was tuned against a Q6_K model and shared by every type. Q4_K
now has its own (calc_rows_per_block takes the type for this), swept at m4096 k14336 on
top of the vdr-4 kernel:
n=1: 2 warps x 4 rows instead of 2x2 89.2 -> 80.9 us
(1x2 92.2, 2x1 116.6, 4x1 121.0, 4x2 92.2, 1x4 86.8, 4x4 83.4, 2x8 96.6, 1x8 104.0)
n=5: 2x8 instead of the shared 4x16 270.4 -> 248.1 us
The shared geometry steps up between n=4 (193) and n=5 (270). 2x8 is worse at
n=3 (171 vs 162), n=4 (209 vs 193) and n=8 (419 vs 386), so it is used at n=5 only --
the full verify batch at --spec-draft-n-max 4.
n=2..4, 6..8: unchanged (4x16).
Macros P100_Q4K_{NWARPS,ROWS}_{1,N,N_LO} for re-sweeping; other types are untouched.
test-backend-ops MUL_MAT (64) and MUL_MAT_ID (84) q4_K cases pass.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
P100_Q4K_VDR defaulted to 4 for every arch. It was tuned on sm_60, where dp4a is emulated and the exact byte sum beats it; sm_61+ has native dp4a and was never measured. Default to upstream's vdr 2 unless the build targets sm_60 only, the same __CUDA_ARCH_LIST__ test the Pascal geometry uses. Checked: nvcc sm_60 -> 4, sm_61 -> 2, sm_60+sm_86 -> 2. P100 build unchanged: q4_K m4096 k14336 n=1 80.8 us; test-backend-ops MUL_MAT (64) and MUL_MAT_ID (84) q4_K pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mewsian
force-pushed
the
q4k-pascal-pr
branch
from
September 30, 2026 03:46
4f4f237 to
12aff2a
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Q4_K decode matvec was the slowest common K-quant on sm_60 (121 us at m=4096 k=14336, vs 91 for q6_K and 81 for q4_0) despite reading fewer bytes than q6_K. This applies the same approach as your Q6_K work: the kernel is bound by memory instructions, so each lane now takes 4 quant ints (vdr 4) and pays the scale/min unpack, dm load and ds loads once per 32 values instead of once per 16. The min-term q8 sum is an exact integer byte sum instead of two emulated dp4a.
Second commit gives Q4_K its own Pascal launch geometry (the shared one was tuned on a Q6_K model): 2 warps x 4 rows at n=1, 2x8 at n=5 (the full MTP verify batch at
--spec-draft-n-max 4), other n unchanged. Other types are untouched.Kernel, us/run at m4096 k14336
llama-bench tg256,
-sm tensor, q8_0 KV, 2x P100Accuracy (
llama-perplexity --kl-divergence -ub 8, so it goes through mmvq)Third commit keeps vdr 4 to sm_60-only builds (same
__CUDA_ARCH_LIST__ == 600test as the Pascal geometry). It was only measured on sm_60, where dp4a is emulated; sm_61+ has native dp4a, so other builds keep upstream's vdr 2.test-backend-ops MUL_MAT (64) and MUL_MAT_ID (84) q4_K cases pass on the P100.
Compile-time switches are left in for A/B (
P100_Q4K_VDR,P100_Q4K_DSMIN,P100_Q4K_{NWARPS,ROWS}_*). DSMIN=1 (min term from q8_1 ds.y) was 9% faster but is not exact (KLD 0.019 on gemma), so it is off. Happy to strip those if you'd rather not carry them.Details are in the commit messages. Thanks for the fork, the Q6_K work made this one straightforward.
🤖 Generated with Claude Code