Repository navigation
[OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [2/4] - #2577
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (3)
🚧 Files skipped from review as they are similar to previous changes (2)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review. 📝 WalkthroughWalkthroughThe PR adds model-specific PTQ support for Falcon, Llama 4, and Nemotron-H. The Hugging Face quantization plugin imports the model-specific modules. Package documentation describes the module and registration conventions. ChangesModel-specific PTQ support
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Refactor Suggested reviewers: Merge Risk: ⚪ Minimal · up to The inspected PTQ paths retain their previous behavior, with no identified issue blocking merge after normal checks. 🚥 Pre-merge checks | ✅ 5 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (5 passed)
✨ Finishing Touches 💡 2📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
🛠️ Fix failing CI checks 💡
Comment |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #2577 +/- ##
==========================================
+ Coverage 71.54% 78.07% +6.53%
==========================================
Files 640 641 +1
Lines 71316 71273 -43
==========================================
+ Hits 51020 55644 +4624
+ Misses 20296 15629 -4667
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
19a86e7 to
926a490
Compare
|
926a490 to
eb3f74b
Compare
eb3f74b to
f3d094c
Compare
87b07c5 to
8d064a3
Compare
8d064a3 to
204455b
Compare
470dc61 to
cc5b7bc
Compare
) ### What does this PR do? Type of change: Refactor ### Summary of the series Model-specific PTQ modeling moves out of `modelopt/torch/quantization/plugins/huggingface.py` into the per-model packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`, next to each model's `specs.py`. By the end of the series `huggingface.py` shrinks from ~2025 to ~1525 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no `ModelSpec` get one. - Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and `is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily. - `huggingface.py` imports every `modeling_ptq` from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first. - Each package's `__init__.py` keeps importing only `specs`, so `import modelopt.torch.models` does not pull in quantization or transformers. - No public API or quantization behavior changes; only private (`_`-prefixed) names change module. Merge in order. #2576 targets `main`; each later PR targets the one before it. 1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7 ← **this PR** 2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon and Llama4 3. #2578 — Move GPT-OSS and Qwen3-VL-MoE 4. #2580 — Move the Step family (step3p5, shared by step3p7) The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. ### This PR [1/4] Registers a `ModelSpec` for each model type that gains `modeling_ptq.py` later in the series but has no spec yet. Each spec records only facts checked against transformers 4.57 and 5.14: - `llama4` and `qwen3_vl_moe`: fused MoE layout (`gate_up_proj`/`down_proj`), no gate/up pair, grouped export off. This is data only: both blocks were already detected as MoE (structurally and by name), and their expert containers are exported without name lookups. - `falcon`: dense, no sections. - `step3p5` / `step3p7`: `modeling_source="remote_code"` and intentionally no `MoESpec`. Step's expert projections sit directly on the MoE MLP with no `experts` container. Declaring the block would make `is_moe` claim it and send AWQ export into `get_experts_list`, which does not support that layout. `step3p7` gets its own package, following the `gemma4_text` / `gemma4` precedent. The exhaustive spec tables in `tests/unit/torch/models/test_model_specs.py` (`EXPECTED_MOE_LAYOUTS`, `root_class_names`) gain the two MoE rows. ### Usage No API change. The new specs are read through the existing registry, e.g. `get_spec("llama4").moe_spec`. ### Testing Each branch of the stack checked out and tested with `pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe` on CPU (torch 2.11, transformers 5.14), on `main` at 9f902ae (#2649): 2248 passed, 7 skipped. The new block names resolve in `test_specs_vs_transformers.py`, and the step3p5/step3p7 remote-code absence checks pass. Class names were also checked against the transformers 4.57 wheel. GPU tests and the transformers 4.57 matrix are left to CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Only private (`_`-prefixed) names change module; public API and quantization behavior are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ Spec tables extended (`EXPECTED_MOE_LAYOUTS`, `root_class_names`). - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal refactor) - Did you get Claude approval on this PR?: ❌ ### Additional Information Part of a 5-PR stack; see the series list above. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added model support for Falcon, Llama 4, Qwen3-VL-MoE, Step-3.5, and Step-3.7. * Llama 4 and Qwen3-VL-MoE support includes recognition of their fused expert layers. * Falcon, Llama 4, and Qwen3-VL-MoE require Transformers 4.57 or later. Step-3.5 and Step-3.7 use remote-code modeling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
c373800 to
91c945a
Compare
cjluo-nv
left a comment
There was a problem hiding this comment.
Bot review (gpt-6.1-sol) — DM the bot to share feedback.
Nudge: the moves look behavior-preserving, but the new import-time registration plumbing needs an automated regression test.
Needs action:
- Add subprocess tests in
tests/unit/torch/modelsfor importingmodeling_ptq, quantization, and export first; assert Falcon/Llama4 registrations, the Falcon callback, and Nemotron-H discoverer precedence without duplicate entries.
No action needed:
- The design retains existing quantization and calibration registries rather than introducing replacements. The PR body explains model-package colocation, keeping specs separate, and delayed explicit imports for dependency ordering.
- Wrapper and discoverer bodies move unchanged; no existing tests are weakened. Tests were inspected, not executed.
|
@cjluo-nv re: the registration-test item in your review — added in adbe38e: |
adbe38e to
69f0f03
Compare
cjluo-nv
left a comment
There was a problem hiding this comment.
Bot review (gpt-6.1-sol) — DM the bot to share feedback.
Approve: the prior import-registration coverage concern is resolved, and the refactor preserves existing wrappers and registration behavior.
No action needed:
- ✔️ Resolved since the last review: fresh-interpreter tests cover all five entry imports, Falcon/Llama4 registrations, the Falcon callback, and unique Nemotron-H discoverer precedence.
- The documented design reuses existing registries and keeps model specs separate from PTQ dependencies. New headers match the standard NVIDIA license text.
- Tests were inspected, not executed.
Large PR: spans 5 directories (≥ 5). The review came back clean, so this is an LGTM — a human should take the final look and approve.
69f0f03 to
8c8a494
Compare
…odels Introduce <model_type>/modeling_ptq.py for model-specific quantized-module wrappers and registrations, and move the first three there verbatim: Nemotron-H's layerwise decoder discovery, the Falcon linear registration and _QuantLlama4TextExperts. The HF plugin imports each modeling_ptq from an explicit list, placed after the generic wrappers they build on and before the homogeneous decoder discoverer, so Nemotron-H's more specific discoverer still matches first. Package __init__.py files keep importing only specs. Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Each case imports one module first in a fresh interpreter (a modeling_ptq, quantization, or export) and checks that Falcon and Llama4 are registered, the Falcon callback is installed, and Nemotron-H's decoder discoverer is registered once and ahead of the homogeneous HF one. Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> #2670 raised the minimum supported transformers to 5.5 and moved every spec's min_transformers_version to the floor, but llama4, qwen3_vl_moe and falcon landed with 4.57 in #2576. It also removed the pre-5.12 Qwen3-VL-MoE PTQ wrapper, so drop the qwen3_vl_moe spec's pointer to it.
Each first import needs its own interpreter; starting them one after another added over a minute to the unit-test job, which already runs close to its 20-minute limit. Launch them all at once and collect the results, which reports every failing first import together. Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
36e3222 to
d6ce3d3
Compare
…dels [3/4] (#2578) ### What does this PR do? Type of change: Refactor ### Summary of the series Model-specific PTQ modeling moves out of `modelopt/torch/quantization/plugins/huggingface.py` into the per-model packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`, next to each model's `specs.py`. By the end of the series `huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no `ModelSpec` get one. - Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and `is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily. - `huggingface.py` imports every `modeling_ptq` from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first. - Each package's `__init__.py` keeps importing only `specs`, so `import modelopt.torch.models` does not pull in quantization or transformers. - No public API or quantization behavior changes; only private (`_`-prefixed) names change module. Merge in order. #2576 and #2577 have merged; #2578 now targets `main`, and #2580 targets #2578's branch. 1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7 (merged) 2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon and Llama4 (merged) 3. #2578 — Move GPT-OSS ← **this PR** 4. #2580 — Move the Step family (step3p5, shared by step3p7) The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was dropped: #2670 removed its pre-5.12 wrapper. ### This PR [3/4] Moves `_QuantGptOssExperts` and its registration into `gpt_oss/modeling_ptq.py`, adds `gpt_oss` to the import loop, and extends the registration test to it. This slice originally also moved the pre-5.12 Qwen3-VL-MoE wrapper; #2670 removed that wrapper, so that half was dropped. ### Usage No API change. To add PTQ support for a new model, put its wrapper in `modelopt/torch/models/<model_type>/modeling_ptq.py` and add `<model_type>` to the import loop at the end of `quantization/plugins/huggingface.py`. ### Testing Each branch of the stack checked out and tested on `main` at 90ba9fb, CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe`, 2371 passed, 1 skipped. With transformers 5.5 (the new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and `test_huggingface.py`, 206 passed, 1 skipped. GPU tests are left to CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Only private (`_`-prefixed) names change module; public API and quantization behavior are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (verbatim move; existing tests re-pointed) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal refactor) - Did you get Claude approval on this PR?: ❌ ### Additional Information Part of a 5-PR stack; see the series list above. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added post-training quantization support for GPT-OSS models. * Added support for quantizing legacy Qwen3-VL-MoE layouts, while retaining support for newer fused layouts. * Quantization support now accommodates both legacy and newer Qwen3-VL-MoE model layouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
…dels [4/4] (#2580) ### What does this PR do? Type of change: Refactor ### Summary of the series Model-specific PTQ modeling moves out of `modelopt/torch/quantization/plugins/huggingface.py` into the per-model packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`, next to each model's `specs.py`. By the end of the series `huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no `ModelSpec` get one. - Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and `is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily. - `huggingface.py` imports every `modeling_ptq` from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first. - Each package's `__init__.py` keeps importing only `specs`, so `import modelopt.torch.models` does not pull in quantization or transformers. - No public API or quantization behavior changes; only private (`_`-prefixed) names change module. Merge in order. #2576, #2577 and #2578 have merged; this last slice, #2580, now targets `main`. 1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7 (merged) 2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon and Llama4 (merged) 3. #2578 — Move GPT-OSS (merged) 4. #2580 — Move the Step family (step3p5, shared by step3p7) ← **this PR** The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was dropped: #2670 removed its pre-5.12 wrapper. ### This PR [4/4] Moves `_QuantMoELinear`, `_is_expert_indexed_moe_linear`, `_is_step_family_model`, `register_moe_linear_on_the_fly` and `_reconstruct_fused_moe_linear` into `step3p5/modeling_ptq.py`, which also serves `step3p7`. This carries over #2569's switch of `_is_step_family_model` to `hf_model_type`; `huggingface.py` drops its now-unused `hf_model_type` and `re` imports. Export (`hf_export_handlers.py`, `layerwise_export.py`, `unified_export_hf.py`, `unified_export_hf_streaming.py`) imports these names from the new module. The imports stay deferred, because `modelopt.torch.quantization` imports `modelopt.torch.export` and importing at module scope would risk an import cycle; the comments now give that reason. The PTQ skill references (`unsupported-models.md`, `checkpoint-validation.md`) now point agents at `modelopt/torch/models/<model_type>/modeling_ptq.py` for model-specific patches. Rebasing onto #2649 also replaces the README's DBRX example of a model-specific wrapper with Llama4's fused BMM experts. `step3p7` also gets its own `modeling_ptq.py`, which imports the Step-3.5 module: Step-3.7 remote-code checkpoints use the same `MoELinear`, and today they are covered only because the HF plugin imports every `modeling_ptq`. With its own module, a future loader that imports PTQ modeling by model type finds it too. ### Usage No API change. To add PTQ support for a new model, put its wrapper in `modelopt/torch/models/<model_type>/modeling_ptq.py` and add `<model_type>` to the import loop at the end of `quantization/plugins/huggingface.py`. ### Testing Each branch of the stack checked out and tested on `main` at 90ba9fb, CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe`, 2372 passed, 1 skipped. With transformers 5.5 (the new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and `test_huggingface.py`, 207 passed, 1 skipped. GPU tests are left to CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Only private (`_`-prefixed) names change module; public API and quantization behavior are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (verbatim move; existing tests re-pointed) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal refactor) - Did you get Claude approval on this PR?: ❌ ### Additional Information Part of a 5-PR stack; see the series list above. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added post-training quantization support for Step-family models with expert-indexed MoE layers. * Checkpoint exports preserve the original layout of quantized expert weights and scales. * **Documentation** * Updated MoE architecture examples and guidance for configuring model-specific quantization support, including Step model revisions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
What does this PR do?
Type of change: Refactor
Summary of the series
Model-specific PTQ modeling moves out of
modelopt/torch/quantization/plugins/huggingface.pyinto the per-model packages undermodelopt/torch/models/<model_type>/modeling_ptq.py, next to each model'sspecs.py. By the end of the serieshuggingface.pyshrinks from ~1880 to ~1500 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention,FP8Linear,CompressedLinear, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had noModelSpecget one.CUSTOM_MODEL_PLUGINSby their own module, andis_homogeneous_hf_modelimportsis_nemotron_h_modellazily.huggingface.pyimports everymodeling_ptqfrom an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first.__init__.pykeeps importing onlyspecs, soimport modelopt.torch.modelsdoes not pull in quantization or transformers._-prefixed) names change module.Merge in order. #2576 targets
main; each later PR targets the one before it.modeling_ptq.pyconvention; move Nemotron-H, Falcon and Llama4 ← this PRThe DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was dropped: #2670 removed its pre-5.12 wrapper.
This PR [2/4]
Introduces
<model_type>/modeling_ptq.pyand moves the first three models there:nemotron_h:is_nemotron_h_model/get_nemotron_h_decoder_layersand their layerwise-calibration registration.falcon: theFalconLinearregistration andregister_falcon_linears_on_the_fly.llama4:_QuantLlama4TextExpertsand its registration.The registrations move in the form #2670 gave them (a top-level transformers import plus
if ... not in QuantModuleRegistry), andhuggingface.pydrops the imports it no longer uses. It also adds theimportlibloop inhuggingface.pythat imports eachmodeling_ptq, documents the convention inmodelopt/torch/models/README.mdand the package docstring, and addstests/unit/torch/models/test_modeling_ptq_registration.py, which checks in a fresh interpreter per case that registration does not depend on which module is imported first.It also raises the
llama4,qwen3_vl_moeandfalconspecs from #2576 to the transformers 5.5 floor that #2670 set for every other spec, and drops theqwen3_vl_moespec's pointer to the pre-5.12 wrapper #2670 removed.Usage
No API change. To add PTQ support for a new model, put its wrapper in
modelopt/torch/models/<model_type>/modeling_ptq.pyand add<model_type>to the import loop at the end ofquantization/plugins/huggingface.py.Testing
Each branch of the stack checked out and tested on
mainat 90ba9fb, CPU, torch 2.11. With transformers 5.18:pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe, 2370 passed, 1 skipped. With transformers 5.5 (the new floor):tests/unit/torch/modelsplustest_moe_linear.pyandtest_huggingface.py, 205 passed, 1 skipped. GPU tests are left to CI.Before your PR is "Ready for review"
_-prefixed) names change module; public API and quantization behavior are unchanged.CONTRIBUTING.md: N/AAdditional Information
Part of a 5-PR stack; see the series list above.
🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Bug Fixes
Documentation