Repository navigation
Conversation
hf_ptq ran the IQ search twice for every weight: once on the first forward after quantization, where fake quant packs the weight and caches the payload, and again in export, which called the format's encoder on the same unchanged weight. Instrumenting the TinyLlama example showed 154 encodes from fake quant and 154 more from export for 154 weights, in all five formats. On a 27B model at IQ1_M's 318 M elem/s the second pass is about 85 seconds of repeated work. IQFormat.pack now returns the quantizer's cached payload when it was packed from the weight being exported, and runs the search only otherwise. Fake quant is handed the weight reshaped into 256-value blocks, so the cache is keyed on that view rather than on the weight export holds. pack therefore also accepts another contiguous view of the same storage, version and length -- the same values in the same order, hence the same GGML blocks -- and reshapes the payload to the weight's layout. The version counter still makes a weight edited after its forward pack afresh. Reusing the payload also means the checkpoint holds exactly the bytes the evaluated model decoded. The HF exporter calls it; the Megatron exporter still packs itself, since it can remap or slice a weight before packing it. The cache key is now a NamedTuple so the view comparison can name its fields. The reuse test uses a 512-wide weight so the blocked view really differs from the weight, as it does on real models; with an exact-key match only, it fails for every format. The IQ payload export test now covers all five formats rather than two. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Every format had a CUDA encoder but decoded with PyTorch ops, which was fine while decoding looked like a one-off. Fake quant decodes every weight on every forward, though, so once packing was fast and cached, decoding became the cost: in the TinyLlama hf_ptq example it was 32 to 46 seconds of a 55 to 70 second run, 3 ms per weight per forward, for the 100 preview tokens. On a 27B model it would be about 10 seconds per forward. Each format now has a CUDA unpacker beside its packer. One thread decodes one 8-value vector: it reads that vector's index, sign or delta bits and local scale from the block, and writes the eight values in the requested dtype. common.cuh holds the launcher, the bit readers and the two value forms, sign-flipped for IQ2 and delta-shifted for IQ1. Every float operation is explicitly rounded and follows the PyTorch decoder's order, so the compiler cannot fuse a multiply into an add, and the output is bit-identical: across random payloads, including block scales that decode to inf or NaN, and real encodings, in float32, bfloat16, float16 and float64. The CUDA path also reproduces llama.cpp's own values on the conformance blocks. dequantize_<format> uses it for CUDA payloads and keeps the PyTorch path otherwise. Decoding a 5632x2048 weight drops from 3.3 to 5.2 ms to 0.06 to 0.14 ms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
The five IQ encoders were five copies of the same search. IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and launcher, and IQ2_S 113 of them; IQ1_S and IQ1_M had the same structure with a different choice space. Every scaled packer also validated its scales twice, in the pybind wrapper and again in the CUDA entry point. iq2_family.cuh now holds the one IQ2 encoder: the grid in shared memory, the 16 local scales scored per group, then the best entry per vector under the chosen scale. A format supplies its group shape, whether it stores seven sign bits and recovers the eighth from parity, and a store() that writes the chosen entries, sign masks and local scales into its layout. iq1_family.cuh holds the one IQ1 encoder over the shared ternary grid. Each group picks one of kChoices options; with kSharedShift the option also fixes the +/-1/8 delta (IQ1_S: shift * 8 + local), otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now writes FP16 scales, so both IQ1 formats take the same input. Each format file is now one Format struct -- layout constants, store() and the existing decode() -- plus its two entry points. Validation lives once in common.cuh, as check_pack_inputs and check_scaled_pack_inputs, and ggml.cpp binds the CUDA entry points directly, as it already did the unpackers. Kernel sources drop from 1,747 to 1,445 lines. Nothing changes in the output. Every encoder and decoder hashes the same as before on float32, bfloat16, float16 and float64 inputs with zero, tiny, oversized and non-finite blocks, and encode speed is unchanged within 1% on an idle RTX PRO 6000. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (19)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review. 📝 WalkthroughWalkthroughThe change adds shared CUDA encoders and decoders for five IQ formats, exposes CUDA unpackers through the extension, and uses them for CUDA-resident dequantization. IQ weight export now uses cache-aware packing, with tests for payload reuse and updates to weights. ChangesIQ CUDA packing, decoding, and export
Priority: ⬇️ Low Estimated code review effort: 4 (Complex) | ~45 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant Dequantize as dequantize_iq1_s
participant Extension as GGML extension
participant Unpacker as iq1_s_unpack_cuda
participant Decoder as decode_blocks
Dequantize->>Extension: Call iq1_s_unpack with payload, grid, and dtype
Extension->>Unpacker: Invoke CUDA unpack binding
Unpacker->>Decoder: Decode blocks with the IQ1_S format
Decoder-->>Dequantize: Return decoded tensor for reshaping
Possibly related PRs
Suggested reviewers: Merge Risk: ⚪ Minimal · up to This change consolidates the IQ CUDA encoders, adds CUDA decoding, and reuses cached payloads during export. No concrete defect was identified. The reported bit-identical output hashes and passing tests are consistent with this, so the risk to merging appears minimal. 🚥 Pre-merge checks | ✅ 5 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (5 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 69.81% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 53 functions across 16 files. (3 skipped: 3 unsupported.)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
|
|
Folded into #2604, which now carries this refactor on top of its two commits (same head, |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #2614 +/- ##
==========================================
+ Coverage 69.45% 78.87% +9.41%
==========================================
Files 611 611
Lines 68219 68272 +53
==========================================
+ Hits 47383 53850 +6467
+ Misses 20836 14422 -6414
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
What does this PR do?
Type of change: refactor (no behaviour change)
The five GGML IQ CUDA encoders were five copies of the same search. IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and launcher code, and IQ2_S 113 of them. IQ1_S and IQ1_M had the same structure with a different choice space. Every scaled packer also validated its scales twice, in the
ggml.cpppybind wrapper and again in the CUDA entry point.This PR keeps one encoder per family, as two templates:
iq2_family.cuhfor IQ2_XS, IQ2_XXS and IQ2_S. The grid sits in shared memory, the 16 local scales are scored per group, and each vector then takes its best entry under the chosen scale. A format supplies its group shape, whether it stores seven sign bits and recovers the eighth from parity, and astore()that writes the chosen entries, sign masks and local scales into its layout.iq1_family.cuhfor IQ1_S and IQ1_M, over the shared ternary grid. Each group picks one ofkChoicesoptions. WithkSharedShiftthe option also fixes the ±1/8 delta (IQ1_S:shift * 8 + local); otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now writes FP16 scales, so both IQ1 formats take the same input.Each format file is now one
Formatstruct, holding its layout constants andstore(), plus its entry point: 58–100 lines each. Validation lives once incommon.cuh, ascheck_pack_inputsandcheck_scaled_pack_inputs.ggml.cppbinds the CUDA entry points directly instead of through five wrappers. The kernel sources shrink from 1,536 to 1,241 lines (+665 / −960).This is the first of two PRs. #2604 builds on it: it adds CUDA decoders as a
decode()next to each format'sstore(), and makes export reuse fake quant's packed payloads.Testing
Nothing changes in the output. Before the refactor I hashed 40 outputs: 5 formats × float32/bfloat16/float16/float64 inputs × encode and decode, on a weight with zero, tiny, oversized and non-finite blocks. All 40 hash the same afterwards.
Encode speed is unchanged. Old and new were timed alternately for four rounds, in both orders, on an idle RTX PRO 6000 with a 5632×2048 weight. They were within 1% for every format: IQ1_S 37.6 / 37.6 ms, IQ1_M 37.0 / 37.0, IQ2_XXS 10.9 / 10.9, IQ2_XS 11.9 / 12.0, IQ2_S 15.5 / 15.4.
tests/gpu/torch/quantization/test_iq_formats_cuda.py,test_iq1_s_cuda.py,test_iq2_xs_cuda.py: 49 passedtests/gpu/_extensions/test_torch_extensions.py: the validation-message tests pass. [OMNIML-5899] Add Q8_0 CUDA packing kernel #2515's two Q8_0 tests fail identically on a cleanmainon this GPU.test_ggml_backend.py,test_iq_formats.py,test_convert_hf_config.py,test_presets.py,test_export_weight.py): 173 passedBefore your PR is "Ready for review"
CONTRIBUTING.md: ✅ No new code sources or dependencies.Additional Information
Merge order: this → #2604 (pack each IQ weight once and decode on CUDA).
🤖 Generated with Claude Code