diff --git a/CHANGELOG.rst b/CHANGELOG.rst index e32358de351..bc9d618d645 100755 --- a/CHANGELOG.rst +++ b/CHANGELOG.rst @@ -13,6 +13,7 @@ Changelog *Quantization* +- Backfill checkpoint aliases for three more published NVFP4 releases: ``moonshotai/Kimi-K2.7-Code``, ``google/gemma-4-26B-A4B-it`` and ``google/diffusiongemma-26B-A4B-it``. Each imports an existing general recipe wholesale rather than copying its body. - Backfill checkpoint aliases for nine more published NVFP4 releases, so each is reachable from its source model's hub path: ``zai-org/GLM-5.1`` and ``GLM-5.2``, ``MiniMaxAI/MiniMax-M2.5`` and ``MiniMax-M3``, ``deepseek-ai/DeepSeek-V3.1`` and ``DeepSeek-V3.2``, ``Qwen/Qwen3-235B-A22B-Instruct-2507`` and ``-Thinking-2507``, and ``Qwen/Qwen3.6-27B``. Each imports an existing general or architecture recipe wholesale rather than copying its body. - Add composed Hugging Face AutoQuantize recipes that run fixed PTQ or weight AutoQuantize before a separate KV-cache AutoQuantize stage, with independent resumable checkpoints for the weight and diff --git a/modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml b/modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml index 90decca0902..08894fec91d 100644 --- a/modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml +++ b/modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml @@ -20,7 +20,7 @@ imports: metadata: description: >- Qwen3.8-27B local-Hessian PTQ using the NVFP4 W4A4 and FP8 W8A8 assignment exported by - the 5.5-bit NVFP4-max AutoQuantize sweep. + the 5.5-bit NVFP4-max AutoQuantize sweep, as published in nvidia/Qwen3.8-27B-NVFP4. quantize: $import: base diff --git a/modelopt_recipes/models/google/diffusiongemma-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml b/modelopt_recipes/models/google/diffusiongemma-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml new file mode 100644 index 00000000000..d41599700b3 --- /dev/null +++ b/modelopt_recipes/models/google/diffusiongemma-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml @@ -0,0 +1,30 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +# Alias recipe for nvidia/diffusiongemma-26B-A4B-it-NVFP4 (source google/diffusiongemma-26B-A4B-it). Same recipe as `general/ptq/nvfp4_experts_only-kv_fp8_cast`: NVFP4 on the routed experts +# only, max calibration, with an FP8 KV cache in cast mode -- the same scheme as the +# gemma-4-26B-A4B-it release. This model nests its decoder under `model.decoder.layers` +# rather than `model.language_model.layers`, which the recipe's `*.experts.*` patterns +# match either way. + +imports: + base: general/ptq/nvfp4_experts_only-kv_fp8_cast + +$import: base +metadata: + description: >- + google/diffusiongemma-26B-A4B-it quantized with the general expert-only NVFP4 scheme + (max calibration) and an FP8 KV cache in cast mode, as published in + nvidia/diffusiongemma-26B-A4B-it-NVFP4. diff --git a/modelopt_recipes/models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml b/modelopt_recipes/models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml new file mode 100644 index 00000000000..4f38d435a43 --- /dev/null +++ b/modelopt_recipes/models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml @@ -0,0 +1,28 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +# Alias recipe for nvidia/Gemma-4-26B-A4B-NVFP4 (source google/gemma-4-26B-A4B-it). Same recipe as `general/ptq/nvfp4_experts_only-kv_fp8_cast`: NVFP4 on the routed experts +# only, max calibration, with an FP8 KV cache in cast mode. The dense MLPs, router, +# attention, vision tower and `lm_head` stay BF16. + +imports: + base: general/ptq/nvfp4_experts_only-kv_fp8_cast + +$import: base +metadata: + description: >- + google/gemma-4-26B-A4B-it quantized with the general expert-only NVFP4 scheme (max + calibration) and an FP8 KV cache in cast mode, as published in + nvidia/Gemma-4-26B-A4B-NVFP4. diff --git a/modelopt_recipes/models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast.yaml b/modelopt_recipes/models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast.yaml new file mode 100644 index 00000000000..119cf553713 --- /dev/null +++ b/modelopt_recipes/models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast.yaml @@ -0,0 +1,30 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +# Alias recipe for nvidia/Kimi-K2.7-Code-NVFP4 (source moonshotai/Kimi-K2.7-Code). Same recipe as +# `general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast`: NVFP4 on the routed experts +# with the expert `input_scale` pinned to 1.0 -- no activation calibration -- and an FP8 KV +# cache in cast mode. The release's exported expert `input_scale` tensors are exactly 1.0, +# and its shared experts, attention, `lm_head` and vision tower are all left unquantized. + +imports: + base: general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast + +$import: base +metadata: + description: >- + moonshotai/Kimi-K2.7-Code quantized with the general expert-only NVFP4 scheme, expert + `input_scale` pinned to 1.0 (calibration-free), and an FP8 KV cache in cast mode, as + published in nvidia/Kimi-K2.7-Code-NVFP4. diff --git a/modelopt_recipes/ptq.md b/modelopt_recipes/ptq.md index a8c5a887416..42c84e73d38 100644 --- a/modelopt_recipes/ptq.md +++ b/modelopt_recipes/ptq.md @@ -553,6 +553,15 @@ entry is a thin **alias** that imports that recipe wholesale and overrides only `model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4 (MSE static weights) on the routed experts, ModelOpt-default FP8 elsewhere, and an FP8 KV cache — as published in `nvidia/Qwen3.5-397B-A17B-NVFP4-V2`. +- **`models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast`** + aliases `general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast` — expert-only NVFP4 with + the expert `input_scale` pinned to 1.0 (no activation calibration) and an FP8 KV cache in + cast mode — as published in `nvidia/Kimi-K2.7-Code-NVFP4`. +- **`models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast`** and + **`models/google/diffusiongemma-26B-A4B-it/ptq/...`** alias + `general/ptq/nvfp4_experts_only-kv_fp8_cast` — expert-only NVFP4 with max calibration and an + FP8 KV cache in cast mode — as published in `nvidia/Gemma-4-26B-A4B-NVFP4` and + `nvidia/diffusiongemma-26B-A4B-it-NVFP4`. - **`models/zai-org/GLM-5.1/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast`** and **`models/zai-org/GLM-5.2/ptq/...`** alias `general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast` — expert-only NVFP4 with the