Skip to content

[OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [2/4] - #2577

Merged
shengliangxu merged 4 commits into
mainfrom
shengliangx/move-modeling_2
Oct 8, 2026
Merged

shengliangxu merged 4 commits into
mainfrom
shengliangx/move-modeling_2

Conversation

@shengliangxu

@shengliangxu shengliangxu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Type of change: Refactor

Summary of the series

Model-specific PTQ modeling moves out of modelopt/torch/quantization/plugins/huggingface.py into the per-model packages under modelopt/torch/models/<model_type>/modeling_ptq.py, next to each model's specs.py. By the end of the series huggingface.py shrinks from ~1880 to ~1500 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, FP8Linear, CompressedLinear, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no ModelSpec get one.

  • Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to CUSTOM_MODEL_PLUGINS by their own module, and is_homogeneous_hf_model imports is_nemotron_h_model lazily.
  • huggingface.py imports every modeling_ptq from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first.
  • Each package's __init__.py keeps importing only specs, so import modelopt.torch.models does not pull in quantization or transformers.
  • No public API or quantization behavior changes; only private (_-prefixed) names change module.

Merge in order. #2576 targets main; each later PR targets the one before it.

  1. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [1/4] #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7
  2. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [2/4] #2577 — Add the modeling_ptq.py convention; move Nemotron-H, Falcon and Llama4 ← this PR
  3. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [3/4] #2578 — Move GPT-OSS
  4. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [4/4] #2580 — Move the Step family (step3p5, shared by step3p7)

The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was dropped: #2670 removed its pre-5.12 wrapper.

This PR [2/4]

Introduces <model_type>/modeling_ptq.py and moves the first three models there:

  • nemotron_h: is_nemotron_h_model / get_nemotron_h_decoder_layers and their layerwise-calibration registration.
  • falcon: the FalconLinear registration and register_falcon_linears_on_the_fly.
  • llama4: _QuantLlama4TextExperts and its registration.

The registrations move in the form #2670 gave them (a top-level transformers import plus if ... not in QuantModuleRegistry), and huggingface.py drops the imports it no longer uses. It also adds the importlib loop in huggingface.py that imports each modeling_ptq, documents the convention in modelopt/torch/models/README.md and the package docstring, and adds tests/unit/torch/models/test_modeling_ptq_registration.py, which checks in a fresh interpreter per case that registration does not depend on which module is imported first.

It also raises the llama4, qwen3_vl_moe and falcon specs from #2576 to the transformers 5.5 floor that #2670 set for every other spec, and drops the qwen3_vl_moe spec's pointer to the pre-5.12 wrapper #2670 removed.

Usage

No API change. To add PTQ support for a new model, put its wrapper in modelopt/torch/models/<model_type>/modeling_ptq.py and add <model_type> to the import loop at the end of quantization/plugins/huggingface.py.

Testing

Each branch of the stack checked out and tested on main at 90ba9fb, CPU, torch 2.11. With transformers 5.18: pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe, 2370 passed, 1 skipped. With transformers 5.5 (the new floor): tests/unit/torch/models plus test_moe_linear.py and test_huggingface.py, 205 passed, 1 skipped. GPU tests are left to CI.

Before your PR is "Ready for review"

  • Is this change backward compatible?: ✅ Only private (_-prefixed) names change module; public API and quantization behavior are unchanged.
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A
  • Did you write any new necessary tests?: N/A (verbatim move; covered by existing tests)
  • Did you update Changelog?: N/A (internal refactor)
  • Did you get Claude approval on this PR?: ❌

Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added post-training quantization support for Llama 4 expert layers.
    • Expanded Falcon quantization coverage to include linear layers discovered at runtime and older remote-code checkpoints.
    • Added Nemotron-H model detection and decoder-layer discovery for quantization.
  • Bug Fixes

    • Improved Qwen3-VL mixture-of-experts compatibility by applying the legacy quantization wrapper only when needed.
  • Documentation

    • Clarified how model-specific quantization support is loaded and where model support information belongs.

@copy-pr-bot

copy-pr-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 81c42328-d704-4c4c-9fad-1dc72e0e7bb0
📥 Commits

Reviewing files that changed from the base of the PR and between 8d064a3 and 204455b.

📒 Files selected for processing (3)
  • modelopt/torch/models/__init__.py
  • modelopt/torch/models/falcon/specs.py
  • modelopt/torch/quantization/plugins/huggingface.py
🚧 Files skipped from review as they are similar to previous changes (2)
  • modelopt/torch/models/init.py
  • modelopt/torch/models/falcon/specs.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

The PR adds model-specific PTQ support for Falcon, Llama 4, and Nemotron-H. The Hugging Face quantization plugin imports the model-specific modules. Package documentation describes the module and registration conventions.

Changes

Model-specific PTQ support

Layer / File(s) Summary
Document and load model-specific modules
modelopt/torch/models/README.md, modelopt/torch/models/__init__.py, modelopt/torch/quantization/plugins/huggingface.py
The documentation describes where model-specific PTQ implementations belong. The plugin imports the model-specific modules.
Falcon linear registration
modelopt/torch/models/falcon/modeling_ptq.py, modelopt/torch/models/falcon/specs.py, modelopt/torch/quantization/plugins/huggingface.py
The Falcon module registers Transformers’ FalconLinear and runtime attention dense types. The plugin’s corresponding registrations are removed. The specs comment identifies modeling_ptq.py as the handler for older remote-code checkpoints.
Llama 4 expert quantization
modelopt/torch/models/llama4/modeling_ptq.py, modelopt/torch/quantization/plugins/huggingface.py
The Llama 4 module adds a quantized text-experts implementation and registers it when the Transformers class is available. The plugin’s implementation and registry entry are removed.
Nemotron-H layer discovery
modelopt/torch/models/nemotron_h/modeling_ptq.py, modelopt/torch/quantization/plugins/huggingface.py
The Nemotron-H module detects supported models and registers decoder-layer discovery with LayerActivationCollector. The plugin imports the model detector and removes its local discovery code.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Refactor

Suggested reviewers: edwardf0t1

Merge Risk: ⚪ Minimal · up to 20445

The inspected PTQ paths retain their previous behavior, with no identified issue blocking merge after normal checks.

🚥 Pre-merge checks | ✅ 5 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 30.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 10 functions across 6 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Security Anti-Patterns ✅ Passed No listed security anti-pattern was introduced. The reviewed Python diff adds no torch.load(..., weights_only=False), numpy.load(..., allow_pickle=True), hardcoded trust_remote_code=True, extern…
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: moving model-specific PTQ modeling into the model-specific models package.
✨ Finishing Touches 💡 2
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 74.54545% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 78.07%. Comparing base (90ba9fb) to head (d6ce3d3).
⚠️ Report is 4 commits behind head on main.

Files with missing lines Patch % Lines
modelopt/torch/models/llama4/modeling_ptq.py 50.00% 10 Missing ⚠️
modelopt/torch/models/falcon/modeling_ptq.py 76.92% 3 Missing ⚠️
modelopt/torch/models/nemotron_h/modeling_ptq.py 94.11% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2577      +/-   ##
==========================================
+ Coverage   71.54%   78.07%   +6.53%     
==========================================
  Files         640      641       +1     
  Lines       71316    71273      -43     
==========================================
+ Hits        51020    55644    +4624     
+ Misses      20296    15629    -4667     
Flag Coverage Δ
examples-diffusers 20.25% <56.36%> (-1.12%) ⬇️
examples-gpt-oss 13.47% <54.54%> (+0.02%) ⬆️
examples-hf_ptq 23.41% <72.72%> (-0.02%) ⬇️
examples-llm_distill 13.53% <54.54%> (+0.01%) ⬆️
examples-llm_eval 17.31% <56.36%> (+0.01%) ⬆️
examples-llm_qat 17.53% <56.36%> (+0.01%) ⬆️
examples-llm_sparsity 15.85% <54.54%> (+0.01%) ⬆️
examples-megatron_bridge 26.47% <56.36%> (-0.12%) ⬇️
examples-specdec_bench 13.25% <54.54%> (+0.02%) ⬆️
examples-speculative_decoding 17.60% <56.36%> (-0.16%) ⬇️
examples-torch_onnx 21.60% <56.36%> (+0.01%) ⬆️
examples-torch_trt 15.24% <56.36%> (+0.01%) ⬆️
examples-vllm_serve 13.91% <54.54%> (+0.02%) ⬆️
gpu 58.84% <74.54%> (+25.32%) ⬆️
regression 15.10% <54.54%> (+0.01%) ⬆️
unit 59.73% <74.54%> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch from 19a86e7 to 926a490 Compare October 2, 2026 01:12
@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-08 20:10 UTC

@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch from 926a490 to eb3f74b Compare October 2, 2026 17:13
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch from eb3f74b to f3d094c Compare October 2, 2026 17:17
@shengliangxu
shengliangxu marked this pull request as ready for review October 2, 2026 17:19
@shengliangxu
shengliangxu requested review from a team as code owners October 2, 2026 17:19
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch 3 times, most recently from 87b07c5 to 8d064a3 Compare October 6, 2026 06:17
@shengliangxu
shengliangxu requested a review from a team as a code owner October 6, 2026 08:44
@shengliangxu
shengliangxu removed this pull request from stack #2631 October 6, 2026 09:15
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch from 8d064a3 to 204455b Compare October 6, 2026 09:15
@shengliangxu shengliangxu changed the title Move model-specific PTQ modeling into modelopt/torch/models [2/5] Move model-specific PTQ modeling into modelopt/torch/models [2/4] Oct 6, 2026
@shengliangxu
shengliangxu added this pull request to stack #2672 October 6, 2026 09:17
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch 3 times, most recently from 470dc61 to cc5b7bc Compare October 6, 2026 21:43
Base automatically changed from shengliangx/move-modeling_1 to main October 6, 2026 23:12
shengliangxu added a commit that referenced this pull request Oct 6, 2026
)

### What does this PR do?

Type of change: Refactor

### Summary of the series

Model-specific PTQ modeling moves out of
`modelopt/torch/quantization/plugins/huggingface.py` into the per-model
packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`,
next to each model's `specs.py`. By the end of the series
`huggingface.py` shrinks from ~2025 to ~1525 lines and keeps only the
code shared across models: sequential/fused MoE auto-detection,
attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the
transposed-quantization helpers that gpt_oss and llama4 share. Model
types that gain PTQ modeling but had no `ModelSpec` get one.

- Class and function bodies move verbatim. The only rewrites: on-the-fly
callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and
`is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily.
- `huggingface.py` imports every `modeling_ptq` from an explicit list.
The list sits after the generic wrappers these modules build on and
before the homogeneous decoder discoverer is registered, so Nemotron-H's
more specific discoverer still matches first.
- Each package's `__init__.py` keeps importing only `specs`, so `import
modelopt.torch.models` does not pull in quantization or transformers.
- No public API or quantization behavior changes; only private
(`_`-prefixed) names change module.

Merge in order. #2576 targets `main`; each later PR targets the one
before it.

1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and
step3p7 ← **this PR**
2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon
and Llama4
3. #2578 — Move GPT-OSS and Qwen3-VL-MoE
4. #2580 — Move the Step family (step3p5, shared by step3p7)

The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there
was nothing left to move.

### This PR [1/4]

Registers a `ModelSpec` for each model type that gains `modeling_ptq.py`
later in the series but has no spec yet. Each spec records only facts
checked against transformers 4.57 and 5.14:

- `llama4` and `qwen3_vl_moe`: fused MoE layout
(`gate_up_proj`/`down_proj`), no gate/up pair, grouped export off. This
is data only: both blocks were already detected as MoE (structurally and
by name), and their expert containers are exported without name lookups.
- `falcon`: dense, no sections.
- `step3p5` / `step3p7`: `modeling_source="remote_code"` and
intentionally no `MoESpec`. Step's expert projections sit directly on
the MoE MLP with no `experts` container. Declaring the block would make
`is_moe` claim it and send AWQ export into `get_experts_list`, which
does not support that layout. `step3p7` gets its own package, following
the `gemma4_text` / `gemma4` precedent.

The exhaustive spec tables in
`tests/unit/torch/models/test_model_specs.py` (`EXPECTED_MOE_LAYOUTS`,
`root_class_names`) gain the two MoE rows.

### Usage

No API change. The new specs are read through the existing registry,
e.g. `get_spec("llama4").moe_spec`.

### Testing

Each branch of the stack checked out and tested with `pytest
tests/unit/torch/models tests/unit/torch/quantization
tests/unit/torch/export tests/unit/recipe` on CPU (torch 2.11,
transformers 5.14), on `main` at 9f902ae (#2649): 2248 passed, 7
skipped. The new block names resolve in `test_specs_vs_transformers.py`,
and the step3p5/step3p7 remote-code absence checks pass. Class names
were also checked against the transformers 4.57 wheel. GPU tests and the
transformers 4.57 matrix are left to CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Only private (`_`-prefixed)
names change module; public API and quantization behavior are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ Spec tables extended
(`EXPECTED_MOE_LAYOUTS`, `root_class_names`).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal refactor)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added model support for Falcon, Llama 4, Qwen3-VL-MoE, Step-3.5, and
Step-3.7.
* Llama 4 and Qwen3-VL-MoE support includes recognition of their fused
expert layers.
* Falcon, Llama 4, and Qwen3-VL-MoE require Transformers 4.57 or later.
Step-3.5 and Step-3.7 use remote-code modeling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch 2 times, most recently from c373800 to 91c945a Compare October 6, 2026 23:45

@cjluo-nv cjluo-nv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bot review (gpt-6.1-sol) — DM the bot to share feedback.

Nudge: the moves look behavior-preserving, but the new import-time registration plumbing needs an automated regression test.

Needs action:

  • Add subprocess tests in tests/unit/torch/models for importing modeling_ptq, quantization, and export first; assert Falcon/Llama4 registrations, the Falcon callback, and Nemotron-H discoverer precedence without duplicate entries.

No action needed:

  • The design retains existing quantization and calibration registries rather than introducing replacements. The PR body explains model-package colocation, keeping specs separate, and delayed explicit imports for dependency ordering.
  • Wrapper and discoverer bodies move unchanged; no existing tests are weakened. Tests were inspected, not executed.

@shengliangxu

Copy link
Copy Markdown
Collaborator Author

@cjluo-nv re: the registration-test item in your review — added in adbe38e: tests/unit/torch/models/test_modeling_ptq_registration.py imports one module first in a fresh interpreter (each modeling_ptq in this PR, modelopt.torch.quantization, modelopt.torch.export) and asserts the Falcon and Llama4 registrations, the Falcon on-the-fly callback, and that Nemotron-H's decoder discoverer is registered exactly once and ahead of the homogeneous HF one. Checked that it fails if the modeling_ptq imports move after the homogeneous registration, or if the Falcon callback is dropped. #2578 and #2580 extend the same test to GPT-OSS / Qwen3-VL-MoE and the Step callback.

@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch from adbe38e to 69f0f03 Compare October 7, 2026 17:23

@cjluo-nv cjluo-nv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bot review (gpt-6.1-sol) — DM the bot to share feedback.

Approve: the prior import-registration coverage concern is resolved, and the refactor preserves existing wrappers and registration behavior.

No action needed:

  • ✔️ Resolved since the last review: fresh-interpreter tests cover all five entry imports, Falcon/Llama4 registrations, the Falcon callback, and unique Nemotron-H discoverer precedence.
  • The documented design reuses existing registries and keeps model specs separate from PTQ dependencies. New headers match the standard NVIDIA license text.
  • Tests were inspected, not executed.

Large PR: spans 5 directories (≥ 5). The review came back clean, so this is an LGTM — a human should take the final look and approve.

@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch from 69f0f03 to 8c8a494 Compare October 8, 2026 01:03
…odels

Introduce <model_type>/modeling_ptq.py for model-specific quantized-module
wrappers and registrations, and move the first three there verbatim:
Nemotron-H's layerwise decoder discovery, the Falcon linear registration and
_QuantLlama4TextExperts.

The HF plugin imports each modeling_ptq from an explicit list, placed after
the generic wrappers they build on and before the homogeneous decoder
discoverer, so Nemotron-H's more specific discoverer still matches first.
Package __init__.py files keep importing only specs.

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Each case imports one module first in a fresh interpreter (a modeling_ptq, quantization, or export) and checks that Falcon and Llama4 are registered, the Falcon callback is installed, and Nemotron-H's decoder discoverer is registered once and ahead of the homogeneous HF one.

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>

#2670 raised the minimum supported transformers to 5.5 and moved every spec's min_transformers_version to the floor, but llama4, qwen3_vl_moe and falcon landed with 4.57 in #2576. It also removed the pre-5.12 Qwen3-VL-MoE PTQ wrapper, so drop the qwen3_vl_moe spec's pointer to it.
Each first import needs its own interpreter; starting them one after another added over a minute to the unit-test job, which already runs close to its 20-minute limit. Launch them all at once and collect the results, which reports every failing first import together.

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_2 branch from 36e3222 to d6ce3d3 Compare October 8, 2026 16:47
@shengliangxu
shengliangxu merged commit e862bdf into main Oct 8, 2026
61 of 63 checks passed
@shengliangxu
shengliangxu deleted the shengliangx/move-modeling_2 branch October 8, 2026 20:10
@shengliangxu shengliangxu changed the title Move model-specific PTQ modeling into modelopt/torch/models [2/4] [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [2/4] Oct 8, 2026
shengliangxu added a commit that referenced this pull request Oct 9, 2026
…dels [3/4] (#2578)

### What does this PR do?

Type of change: Refactor

### Summary of the series

Model-specific PTQ modeling moves out of
`modelopt/torch/quantization/plugins/huggingface.py` into the per-model
packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`,
next to each model's `specs.py`. By the end of the series
`huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the
code shared across models: sequential/fused MoE auto-detection,
attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the
transposed-quantization helpers that gpt_oss and llama4 share. Model
types that gain PTQ modeling but had no `ModelSpec` get one.

- Class and function bodies move verbatim. The only rewrites: on-the-fly
callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and
`is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily.
- `huggingface.py` imports every `modeling_ptq` from an explicit list.
The list sits after the generic wrappers these modules build on and
before the homogeneous decoder discoverer is registered, so Nemotron-H's
more specific discoverer still matches first.
- Each package's `__init__.py` keeps importing only `specs`, so `import
modelopt.torch.models` does not pull in quantization or transformers.
- No public API or quantization behavior changes; only private
(`_`-prefixed) names change module.

Merge in order. #2576 and #2577 have merged; #2578 now targets `main`,
and #2580 targets #2578's branch.

1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and
step3p7 (merged)
2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon
and Llama4 (merged)
3. #2578 — Move GPT-OSS  ← **this PR**
4. #2580 — Move the Step family (step3p5, shared by step3p7)

The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there
was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was
dropped: #2670 removed its pre-5.12 wrapper.

### This PR [3/4]

Moves `_QuantGptOssExperts` and its registration into
`gpt_oss/modeling_ptq.py`, adds `gpt_oss` to the import loop, and
extends the registration test to it. This slice originally also moved
the pre-5.12 Qwen3-VL-MoE wrapper; #2670 removed that wrapper, so that
half was dropped.

### Usage

No API change. To add PTQ support for a new model, put its wrapper in
`modelopt/torch/models/<model_type>/modeling_ptq.py` and add
`<model_type>` to the import loop at the end of
`quantization/plugins/huggingface.py`.

### Testing

Each branch of the stack checked out and tested on `main` at 90ba9fb,
CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models
tests/unit/torch/quantization tests/unit/torch/export
tests/unit/recipe`, 2371 passed, 1 skipped. With transformers 5.5 (the
new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and
`test_huggingface.py`, 206 passed, 1 skipped. GPU tests are left to CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Only private (`_`-prefixed)
names change module; public API and quantization behavior are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (verbatim move; existing
tests re-pointed)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal refactor)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added post-training quantization support for GPT-OSS models.
* Added support for quantizing legacy Qwen3-VL-MoE layouts, while
retaining support for newer fused layouts.
* Quantization support now accommodates both legacy and newer
Qwen3-VL-MoE model layouts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
shengliangxu added a commit that referenced this pull request Oct 9, 2026
…dels [4/4] (#2580)

### What does this PR do?

Type of change: Refactor

### Summary of the series

Model-specific PTQ modeling moves out of
`modelopt/torch/quantization/plugins/huggingface.py` into the per-model
packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`,
next to each model's `specs.py`. By the end of the series
`huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the
code shared across models: sequential/fused MoE auto-detection,
attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the
transposed-quantization helpers that gpt_oss and llama4 share. Model
types that gain PTQ modeling but had no `ModelSpec` get one.

- Class and function bodies move verbatim. The only rewrites: on-the-fly
callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and
`is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily.
- `huggingface.py` imports every `modeling_ptq` from an explicit list.
The list sits after the generic wrappers these modules build on and
before the homogeneous decoder discoverer is registered, so Nemotron-H's
more specific discoverer still matches first.
- Each package's `__init__.py` keeps importing only `specs`, so `import
modelopt.torch.models` does not pull in quantization or transformers.
- No public API or quantization behavior changes; only private
(`_`-prefixed) names change module.

Merge in order. #2576, #2577 and #2578 have merged; this last slice,
#2580, now targets `main`.

1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and
step3p7 (merged)
2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon
and Llama4 (merged)
3. #2578 — Move GPT-OSS (merged)
4. #2580 — Move the Step family (step3p5, shared by step3p7) ← **this
PR**

The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there
was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was
dropped: #2670 removed its pre-5.12 wrapper.

### This PR [4/4]

Moves `_QuantMoELinear`, `_is_expert_indexed_moe_linear`,
`_is_step_family_model`, `register_moe_linear_on_the_fly` and
`_reconstruct_fused_moe_linear` into `step3p5/modeling_ptq.py`, which
also serves `step3p7`. This carries over #2569's switch of
`_is_step_family_model` to `hf_model_type`; `huggingface.py` drops its
now-unused `hf_model_type` and `re` imports.

Export (`hf_export_handlers.py`, `layerwise_export.py`,
`unified_export_hf.py`, `unified_export_hf_streaming.py`) imports these
names from the new module. The imports stay deferred, because
`modelopt.torch.quantization` imports `modelopt.torch.export` and
importing at module scope would risk an import cycle; the comments now
give that reason. The PTQ skill references (`unsupported-models.md`,
`checkpoint-validation.md`) now point agents at
`modelopt/torch/models/<model_type>/modeling_ptq.py` for model-specific
patches.

Rebasing onto #2649 also replaces the README's DBRX example of a
model-specific wrapper with Llama4's fused BMM experts.

`step3p7` also gets its own `modeling_ptq.py`, which imports the
Step-3.5 module: Step-3.7 remote-code checkpoints use the same
`MoELinear`, and today they are covered only because the HF plugin
imports every `modeling_ptq`. With its own module, a future loader that
imports PTQ modeling by model type finds it too.

### Usage

No API change. To add PTQ support for a new model, put its wrapper in
`modelopt/torch/models/<model_type>/modeling_ptq.py` and add
`<model_type>` to the import loop at the end of
`quantization/plugins/huggingface.py`.

### Testing

Each branch of the stack checked out and tested on `main` at 90ba9fb,
CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models
tests/unit/torch/quantization tests/unit/torch/export
tests/unit/recipe`, 2372 passed, 1 skipped. With transformers 5.5 (the
new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and
`test_huggingface.py`, 207 passed, 1 skipped. GPU tests are left to CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Only private (`_`-prefixed)
names change module; public API and quantization behavior are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (verbatim move; existing
tests re-pointed)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal refactor)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added post-training quantization support for Step-family models with
expert-indexed MoE layers.
* Checkpoint exports preserve the original layout of quantized expert
weights and scales.
* **Documentation**
* Updated MoE architecture examples and guidance for configuring
model-specific quantization support, including Step model revisions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants