From 22ad7e8a6378a9b01a7648a917f2776e646a8174 Mon Sep 17 00:00:00 2001 From: Shengliang Xu Date: Fri, 2 Oct 2026 21:44:34 +0000 Subject: [PATCH 1/2] Backfill checkpoint aliases for three more published NVFP4 releases Continues the 2026 published-checkpoint backfill. Each entry imports an existing general recipe wholesale; nothing is duplicated. moonshotai/Kimi-K2.7-Code general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast google/gemma-4-26B-A4B-it general/ptq/nvfp4_experts_only-kv_fp8_cast google/diffusiongemma-26B-A4B-it general/ptq/nvfp4_experts_only-kv_fp8_cast Each engineer described the scheme in prose rather than naming a recipe file, so the mapping was checked against the published checkpoints before being recorded: - All three are NVFP4 with an FP8 KV cache, and every model card documents post-training quantization with no quantization-aware training or distillation. - Kimi-K2.7-Code's exported expert input_scale tensors are exactly 1.0, which is what the input_scale1 recipe produces by pinning amax to a constant 2688.0 (E2M1_MAX * E4M3_MAX); its shared experts and attention are excluded. That distinguishes it from the plain experts-only recipe, whose expert input scales are calibrated. - Both Gemma releases quantize only the routed experts -- dense MLPs, router, attention, vision tower and lm_head all appear in exclude_modules -- with calibrated expert input scales and no exported k_scale, i.e. KV cast mode. diffusiongemma nests its decoder under model.decoder.layers rather than model.language_model.layers, which the recipe's `*.experts.*` patterns match either way. Calibration is the one property a checkpoint cannot reveal, so it rests on what each engineer said: the Gemma pair used the default calibration, which is max. Co-Authored-By: Claude Opus 5 (1M context) Signed-off-by: Shengliang Xu --- CHANGELOG.rst | 1 + .../ptq/nvfp4_experts_only-kv_fp8_cast.yaml | 30 +++++++++++++++++++ .../ptq/nvfp4_experts_only-kv_fp8_cast.yaml | 28 +++++++++++++++++ ...experts_only_input_scale1-kv_fp8_cast.yaml | 30 +++++++++++++++++++ modelopt_recipes/ptq.md | 9 ++++++ 5 files changed, 98 insertions(+) create mode 100644 modelopt_recipes/models/google/diffusiongemma-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml create mode 100644 modelopt_recipes/models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml create mode 100644 modelopt_recipes/models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast.yaml diff --git a/CHANGELOG.rst b/CHANGELOG.rst index e32358de351..bc9d618d645 100755 --- a/CHANGELOG.rst +++ b/CHANGELOG.rst @@ -13,6 +13,7 @@ Changelog *Quantization* +- Backfill checkpoint aliases for three more published NVFP4 releases: ``moonshotai/Kimi-K2.7-Code``, ``google/gemma-4-26B-A4B-it`` and ``google/diffusiongemma-26B-A4B-it``. Each imports an existing general recipe wholesale rather than copying its body. - Backfill checkpoint aliases for nine more published NVFP4 releases, so each is reachable from its source model's hub path: ``zai-org/GLM-5.1`` and ``GLM-5.2``, ``MiniMaxAI/MiniMax-M2.5`` and ``MiniMax-M3``, ``deepseek-ai/DeepSeek-V3.1`` and ``DeepSeek-V3.2``, ``Qwen/Qwen3-235B-A22B-Instruct-2507`` and ``-Thinking-2507``, and ``Qwen/Qwen3.6-27B``. Each imports an existing general or architecture recipe wholesale rather than copying its body. - Add composed Hugging Face AutoQuantize recipes that run fixed PTQ or weight AutoQuantize before a separate KV-cache AutoQuantize stage, with independent resumable checkpoints for the weight and diff --git a/modelopt_recipes/models/google/diffusiongemma-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml b/modelopt_recipes/models/google/diffusiongemma-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml new file mode 100644 index 00000000000..d41599700b3 --- /dev/null +++ b/modelopt_recipes/models/google/diffusiongemma-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml @@ -0,0 +1,30 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +# Alias recipe for nvidia/diffusiongemma-26B-A4B-it-NVFP4 (source google/diffusiongemma-26B-A4B-it). Same recipe as `general/ptq/nvfp4_experts_only-kv_fp8_cast`: NVFP4 on the routed experts +# only, max calibration, with an FP8 KV cache in cast mode -- the same scheme as the +# gemma-4-26B-A4B-it release. This model nests its decoder under `model.decoder.layers` +# rather than `model.language_model.layers`, which the recipe's `*.experts.*` patterns +# match either way. + +imports: + base: general/ptq/nvfp4_experts_only-kv_fp8_cast + +$import: base +metadata: + description: >- + google/diffusiongemma-26B-A4B-it quantized with the general expert-only NVFP4 scheme + (max calibration) and an FP8 KV cache in cast mode, as published in + nvidia/diffusiongemma-26B-A4B-it-NVFP4. diff --git a/modelopt_recipes/models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml b/modelopt_recipes/models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml new file mode 100644 index 00000000000..4f38d435a43 --- /dev/null +++ b/modelopt_recipes/models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast.yaml @@ -0,0 +1,28 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +# Alias recipe for nvidia/Gemma-4-26B-A4B-NVFP4 (source google/gemma-4-26B-A4B-it). Same recipe as `general/ptq/nvfp4_experts_only-kv_fp8_cast`: NVFP4 on the routed experts +# only, max calibration, with an FP8 KV cache in cast mode. The dense MLPs, router, +# attention, vision tower and `lm_head` stay BF16. + +imports: + base: general/ptq/nvfp4_experts_only-kv_fp8_cast + +$import: base +metadata: + description: >- + google/gemma-4-26B-A4B-it quantized with the general expert-only NVFP4 scheme (max + calibration) and an FP8 KV cache in cast mode, as published in + nvidia/Gemma-4-26B-A4B-NVFP4. diff --git a/modelopt_recipes/models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast.yaml b/modelopt_recipes/models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast.yaml new file mode 100644 index 00000000000..119cf553713 --- /dev/null +++ b/modelopt_recipes/models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast.yaml @@ -0,0 +1,30 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +# Alias recipe for nvidia/Kimi-K2.7-Code-NVFP4 (source moonshotai/Kimi-K2.7-Code). Same recipe as +# `general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast`: NVFP4 on the routed experts +# with the expert `input_scale` pinned to 1.0 -- no activation calibration -- and an FP8 KV +# cache in cast mode. The release's exported expert `input_scale` tensors are exactly 1.0, +# and its shared experts, attention, `lm_head` and vision tower are all left unquantized. + +imports: + base: general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast + +$import: base +metadata: + description: >- + moonshotai/Kimi-K2.7-Code quantized with the general expert-only NVFP4 scheme, expert + `input_scale` pinned to 1.0 (calibration-free), and an FP8 KV cache in cast mode, as + published in nvidia/Kimi-K2.7-Code-NVFP4. diff --git a/modelopt_recipes/ptq.md b/modelopt_recipes/ptq.md index a8c5a887416..42c84e73d38 100644 --- a/modelopt_recipes/ptq.md +++ b/modelopt_recipes/ptq.md @@ -553,6 +553,15 @@ entry is a thin **alias** that imports that recipe wholesale and overrides only `model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4 (MSE static weights) on the routed experts, ModelOpt-default FP8 elsewhere, and an FP8 KV cache — as published in `nvidia/Qwen3.5-397B-A17B-NVFP4-V2`. +- **`models/moonshotai/Kimi-K2.7-Code/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast`** + aliases `general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast` — expert-only NVFP4 with + the expert `input_scale` pinned to 1.0 (no activation calibration) and an FP8 KV cache in + cast mode — as published in `nvidia/Kimi-K2.7-Code-NVFP4`. +- **`models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast`** and + **`models/google/diffusiongemma-26B-A4B-it/ptq/...`** alias + `general/ptq/nvfp4_experts_only-kv_fp8_cast` — expert-only NVFP4 with max calibration and an + FP8 KV cache in cast mode — as published in `nvidia/Gemma-4-26B-A4B-NVFP4` and + `nvidia/diffusiongemma-26B-A4B-it-NVFP4`. - **`models/zai-org/GLM-5.1/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast`** and **`models/zai-org/GLM-5.2/ptq/...`** alias `general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast` — expert-only NVFP4 with the From 553e45da3ea60ce35bc3b5b449cedc0a17f91755 Mon Sep 17 00:00:00 2001 From: Shengliang Xu Date: Fri, 2 Oct 2026 22:13:30 +0000 Subject: [PATCH 2/2] Name the published release in the Qwen3.8-27B recipe description The recipe reproduces nvidia/Qwen3.8-27B-NVFP4 -- ptq.md already says so -- but the recipe's own metadata.description did not, which makes it read like an unrelated AutoQuantize artifact when you encounter the file on its own. Every other checkpoint entry names the release it corresponds to, so an audit of backfill coverage that reads the recipes can miss this one. Verified the claim against the published checkpoint before asserting it: MLP gate/up/down_proj NVFP4, self_attn and linear_attn projections FP8, lm_head NVFP4, mtp excluded, no KV quantization, and the model card names local-Hessian calibration on 2,048 samples -- all matching this recipe, including the unusual NVFP4 lm_head. Co-Authored-By: Claude Opus 5 (1M context) Signed-off-by: Shengliang Xu --- .../Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml b/modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml index 90decca0902..08894fec91d 100644 --- a/modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml +++ b/modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian.yaml @@ -20,7 +20,7 @@ imports: metadata: description: >- Qwen3.8-27B local-Hessian PTQ using the NVFP4 W4A4 and FP8 W8A8 assignment exported by - the 5.5-bit NVFP4-max AutoQuantize sweep. + the 5.5-bit NVFP4-max AutoQuantize sweep, as published in nvidia/Qwen3.8-27B-NVFP4. quantize: $import: base