Skip to content

Share one CUDA encoder per IQ family - #2614

Closed
cjluo-nv wants to merge 3 commits into
mainfrom
chenjiel/iq-kernel-dedup
Closed

cjluo-nv wants to merge 3 commits into
mainfrom
chenjiel/iq-kernel-dedup

Conversation

@cjluo-nv

@cjluo-nv cjluo-nv commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Type of change: refactor (no behaviour change)

The five GGML IQ CUDA encoders were five copies of the same search. IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and launcher code, and IQ2_S 113 of them. IQ1_S and IQ1_M had the same structure with a different choice space. Every scaled packer also validated its scales twice, in the ggml.cpp pybind wrapper and again in the CUDA entry point.

This PR keeps one encoder per family, as two templates:

  • iq2_family.cuh for IQ2_XS, IQ2_XXS and IQ2_S. The grid sits in shared memory, the 16 local scales are scored per group, and each vector then takes its best entry under the chosen scale. A format supplies its group shape, whether it stores seven sign bits and recovers the eighth from parity, and a store() that writes the chosen entries, sign masks and local scales into its layout.
  • iq1_family.cuh for IQ1_S and IQ1_M, over the shared ternary grid. Each group picks one of kChoices options. With kSharedShift the option also fixes the ±1/8 delta (IQ1_S: shift * 8 + local); otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now writes FP16 scales, so both IQ1 formats take the same input.

Each format file is now one Format struct, holding its layout constants and store(), plus its entry point: 58–100 lines each. Validation lives once in common.cuh, as check_pack_inputs and check_scaled_pack_inputs. ggml.cpp binds the CUDA entry points directly instead of through five wrappers. The kernel sources shrink from 1,536 to 1,241 lines (+665 / −960).

This is the first of two PRs. #2604 builds on it: it adds CUDA decoders as a decode() next to each format's store(), and makes export reuse fake quant's packed payloads.

Testing

Nothing changes in the output. Before the refactor I hashed 40 outputs: 5 formats × float32/bfloat16/float16/float64 inputs × encode and decode, on a weight with zero, tiny, oversized and non-finite blocks. All 40 hash the same afterwards.

Encode speed is unchanged. Old and new were timed alternately for four rounds, in both orders, on an idle RTX PRO 6000 with a 5632×2048 weight. They were within 1% for every format: IQ1_S 37.6 / 37.6 ms, IQ1_M 37.0 / 37.0, IQ2_XXS 10.9 / 10.9, IQ2_XS 11.9 / 12.0, IQ2_S 15.5 / 15.4.

  • tests/gpu/torch/quantization/test_iq_formats_cuda.py, test_iq1_s_cuda.py, test_iq2_xs_cuda.py: 49 passed
  • tests/gpu/_extensions/test_torch_extensions.py: the validation-message tests pass. [OMNIML-5899] Add Q8_0 CUDA packing kernel #2515's two Q8_0 tests fail identically on a clean main on this GPU.
  • IQ unit tests (test_ggml_backend.py, test_iq_formats.py, test_convert_hf_config.py, test_presets.py, test_export_weight.py): 173 passed

Before your PR is "Ready for review"

  • Is this change backward compatible?: ✅ Same bindings, messages and bytes.
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: ✅ No new code sources or dependencies.
  • Did you write any new necessary tests?: N/A. A refactor with no behaviour change, verified by the hashes above and the existing GPU tests.
  • Did you update Changelog?: N/A
  • Did you get Claude approval on this PR?: ❌ Not yet run.

Additional Information

Merge order: this → #2604 (pack each IQ weight once and decode on CUDA).

🤖 Generated with Claude Code

cjluo-nv and others added 3 commits October 1, 2026 05:24
hf_ptq ran the IQ search twice for every weight: once on the first forward
after quantization, where fake quant packs the weight and caches the
payload, and again in export, which called the format's encoder on the
same unchanged weight. Instrumenting the TinyLlama example showed 154
encodes from fake quant and 154 more from export for 154 weights, in all
five formats. On a 27B model at IQ1_M's 318 M elem/s the second pass is
about 85 seconds of repeated work.

IQFormat.pack now returns the quantizer's cached payload when it was packed
from the weight being exported, and runs the search only otherwise. Fake
quant is handed the weight reshaped into 256-value blocks, so the cache is
keyed on that view rather than on the weight export holds. pack therefore
also accepts another contiguous view of the same storage, version and
length -- the same values in the same order, hence the same GGML blocks --
and reshapes the payload to the weight's layout. The version counter still
makes a weight edited after its forward pack afresh. Reusing the payload
also means the checkpoint holds exactly the bytes the evaluated model
decoded. The HF exporter calls it; the Megatron exporter still packs
itself, since it can remap or slice a weight before packing it. The cache
key is now a NamedTuple so the view comparison can name its fields.

The reuse test uses a 512-wide weight so the blocked view really differs
from the weight, as it does on real models; with an exact-key match only,
it fails for every format. The IQ payload export test now covers all five
formats rather than two.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Every format had a CUDA encoder but decoded with PyTorch ops, which was
fine while decoding looked like a one-off. Fake quant decodes every weight
on every forward, though, so once packing was fast and cached, decoding
became the cost: in the TinyLlama hf_ptq example it was 32 to 46 seconds
of a 55 to 70 second run, 3 ms per weight per forward, for the 100 preview
tokens. On a 27B model it would be about 10 seconds per forward.

Each format now has a CUDA unpacker beside its packer. One thread decodes
one 8-value vector: it reads that vector's index, sign or delta bits and
local scale from the block, and writes the eight values in the requested
dtype. common.cuh holds the launcher, the bit readers and the two value
forms, sign-flipped for IQ2 and delta-shifted for IQ1. Every float
operation is explicitly rounded and follows the PyTorch decoder's order,
so the compiler cannot fuse a multiply into an add, and the output is
bit-identical: across random payloads, including block scales that decode
to inf or NaN, and real encodings, in float32, bfloat16, float16 and
float64. The CUDA path also reproduces llama.cpp's own values on the
conformance blocks. dequantize_<format> uses it for CUDA payloads and
keeps the PyTorch path otherwise.

Decoding a 5632x2048 weight drops from 3.3 to 5.2 ms to 0.06 to 0.14 ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
The five IQ encoders were five copies of the same search. IQ2_XS and
IQ2_XXS shared 139 of their roughly 150 lines of encoder and launcher, and
IQ2_S 113 of them; IQ1_S and IQ1_M had the same structure with a different
choice space. Every scaled packer also validated its scales twice, in the
pybind wrapper and again in the CUDA entry point.

iq2_family.cuh now holds the one IQ2 encoder: the grid in shared memory, the
16 local scales scored per group, then the best entry per vector under the
chosen scale. A format supplies its group shape, whether it stores seven
sign bits and recovers the eighth from parity, and a store() that writes the
chosen entries, sign masks and local scales into its layout.
iq1_family.cuh holds the one IQ1 encoder over the shared ternary grid. Each
group picks one of kChoices options; with kSharedShift the option also fixes
the +/-1/8 delta (IQ1_S: shift * 8 + local), otherwise each vector picks
its own (IQ1_M). IQ1_S's scale kernel now writes FP16 scales, so both IQ1
formats take the same input.

Each format file is now one Format struct -- layout constants, store() and
the existing decode() -- plus its two entry points. Validation lives once
in common.cuh, as check_pack_inputs and check_scaled_pack_inputs, and
ggml.cpp binds the CUDA entry points directly, as it already did the
unpackers. Kernel sources drop from 1,747 to 1,445 lines.

Nothing changes in the output. Every encoder and decoder hashes the same as
before on float32, bfloat16, float16 and float64 inputs with zero, tiny,
oversized and non-finite blocks, and encode speed is unchanged within 1% on
an idle RTX PRO 6000.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
@cjluo-nv
cjluo-nv requested review from a team as code owners October 1, 2026 05:56
@cjluo-nv
cjluo-nv requested review from jingyu-ml and mxinO October 1, 2026 05:56
@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 28490cb0-1f7e-46fc-9eee-70e99f358cd7

📥 Commits

Reviewing files that changed from the base of the PR and between aa89722 and 7a32325.

📒 Files selected for processing (19)
  • modelopt/torch/export/unified_export_hf.py
  • modelopt/torch/kernels/quantization/ggml/common.cuh
  • modelopt/torch/kernels/quantization/ggml/ggml.cpp
  • modelopt/torch/kernels/quantization/ggml/iq1_family.cuh
  • modelopt/torch/kernels/quantization/ggml/iq1_m.cu
  • modelopt/torch/kernels/quantization/ggml/iq1_s.cu
  • modelopt/torch/kernels/quantization/ggml/iq2_family.cuh
  • modelopt/torch/kernels/quantization/ggml/iq2_s.cu
  • modelopt/torch/kernels/quantization/ggml/iq2_xs.cu
  • modelopt/torch/kernels/quantization/ggml/iq2_xxs.cu
  • modelopt/torch/quantization/extensions.py
  • modelopt/torch/quantization/ggml/common.py
  • modelopt/torch/quantization/ggml/iq1_m.py
  • modelopt/torch/quantization/ggml/iq1_s.py
  • modelopt/torch/quantization/ggml/iq2_s.py
  • modelopt/torch/quantization/ggml/iq2_xs.py
  • modelopt/torch/quantization/ggml/iq2_xxs.py
  • tests/gpu/torch/quantization/test_iq_formats_cuda.py
  • tests/unit/torch/export/test_export_weight.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.


📝 Walkthrough

Walkthrough

The change adds shared CUDA encoders and decoders for five IQ formats, exposes CUDA unpackers through the extension, and uses them for CUDA-resident dequantization. IQ weight export now uses cache-aware packing, with tests for payload reuse and updates to weights.

Changes

IQ CUDA packing, decoding, and export

Layer / File(s) Summary
Shared CUDA validation, encoding, and decoding
modelopt/torch/kernels/quantization/ggml/common.cuh, modelopt/torch/kernels/quantization/ggml/iq1_family.cuh, modelopt/torch/kernels/quantization/ggml/iq2_family.cuh
Common helpers validate pack inputs and decode blocks. Shared IQ1 and IQ2 templates implement encoding for use by format-specific adapters.
IQ format layouts and CUDA entry points
modelopt/torch/kernels/quantization/ggml/iq1_*.cu, modelopt/torch/kernels/quantization/ggml/iq2_*.cu
The five format files define payload layouts and delegate packing and unpacking to the shared CUDA implementations.
CUDA unpacker bindings and Python dispatch
modelopt/torch/kernels/quantization/ggml/ggml.cpp, modelopt/torch/quantization/ggml/iq1_m.py, modelopt/torch/quantization/ggml/iq1_s.py, modelopt/torch/quantization/ggml/iq2_*.py, modelopt/torch/quantization/extensions.py, tests/gpu/torch/quantization/test_iq_formats_cuda.py
The extension exposes five CUDA unpackers. Python dequantizers use them for CUDA payloads when available. GPU tests compare CUDA output with PyTorch output and captured llama.cpp values.
Cache-aware IQ packing and export
modelopt/torch/quantization/ggml/common.py, modelopt/torch/export/unified_export_hf.py, tests/unit/torch/export/test_export_weight.py
IQFormat.pack reuses a matching cached payload across contiguous views or quantizes the weight when no match exists. Export passes the weight quantizer to the registered packer. Tests cover reuse and repacking after an in-place weight update.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant Dequantize as dequantize_iq1_s
  participant Extension as GGML extension
  participant Unpacker as iq1_s_unpack_cuda
  participant Decoder as decode_blocks
  Dequantize->>Extension: Call iq1_s_unpack with payload, grid, and dtype
  Extension->>Unpacker: Invoke CUDA unpack binding
  Unpacker->>Decoder: Decode blocks with the IQ1_S format
  Decoder-->>Dequantize: Return decoded tensor for reshaping
Loading

Possibly related PRs

  • NVIDIA/Model-Optimizer#2505: Adds the IQ formats and CUDA entry points that this change refactors into shared encoders and direct bindings.
  • NVIDIA/Model-Optimizer#2511: Adds the IQ2_XXS CUDA packer and payload layout that this change refactors to use the shared IQ2 encoder.

Suggested reviewers: hychiang-git

Merge Risk: ⚪ Minimal · up to 7a323

This change consolidates the IQ CUDA encoders, adds CUDA decoding, and reuses cached payloads during export. No concrete defect was identified. The reported bit-identical output hashes and passing tests are consistent with this, so the risk to merging appears minimal.

🚥 Pre-merge checks | ✅ 5 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 69.81% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 53 functions across 16 files. (3 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Security Anti-Patterns ✅ Passed No listed security anti-pattern was introduced. The review-scoped diff changes eight modelopt Python files and no examples Python files; added lines contain no `torch.load(..., weights_only=False)…
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the primary change: sharing CUDA encoders across IQ format families. It matches the refactoring described in the changeset.
Full details: Docstring Coverage

Explanation

Docstring coverage is 69.81% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 53 functions across 16 files. (3 skipped: 3 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-01 06:08 UTC

@cjluo-nv

cjluo-nv commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Folded into #2604, which now carries this refactor on top of its two commits (same head, 7a32325f9), so the change is reviewed as one PR.

@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.38710% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 78.87%. Comparing base (aa89722) to head (7a32325).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
modelopt/torch/quantization/ggml/common.py 97.22% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2614      +/-   ##
==========================================
+ Coverage   69.45%   78.87%   +9.41%     
==========================================
  Files         611      611              
  Lines       68219    68272      +53     
==========================================
+ Hits        47383    53850    +6467     
+ Misses      20836    14422    -6414     
Flag Coverage Δ
examples-diffusers 21.38% <22.58%> (+0.04%) ⬆️
examples-gpt-oss 13.62% <22.58%> (+0.05%) ⬆️
examples-hf_ptq 23.19% <95.16%> (-0.04%) ⬇️
examples-llm_distill 13.68% <22.58%> (+0.05%) ⬆️
examples-llm_eval 17.62% <22.58%> (+0.04%) ⬆️
examples-llm_qat 17.79% <22.58%> (+0.01%) ⬆️
examples-llm_sparsity 16.07% <22.58%> (+0.05%) ⬆️
examples-megatron_bridge 26.61% <22.58%> (-0.08%) ⬇️
examples-specdec_bench 13.38% <22.58%> (+0.05%) ⬆️
examples-speculative_decoding 18.00% <22.58%> (-0.02%) ⬇️
examples-torch_onnx 21.94% <22.58%> (+0.04%) ⬆️
examples-torch_trt 15.42% <22.58%> (+0.05%) ⬆️
examples-vllm_serve 13.86% <22.58%> (+0.05%) ⬆️
gpu 58.52% <62.90%> (+36.99%) ⬆️
regression 15.31% <22.58%> (+0.04%) ⬆️
unit 59.16% <66.12%> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@cjluo-nv

cjluo-nv commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

GitHub would not reopen this after its branch was force-pushed, so the encoder refactor continues in #2615. It is now split out to land before #2604 rather than on top of it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant