Backfill checkpoint aliases for Kimi-K2.7-Code and the Gemma-4 MoE releases - #2647
Draft
shengliangxu wants to merge 2 commits into
Draft
shengliangxu wants to merge 2 commits into
shengliangxu wants to merge 2 commits into
Conversation
Continues the 2026 published-checkpoint backfill. Each entry imports an existing general recipe wholesale; nothing is duplicated. moonshotai/Kimi-K2.7-Code general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast google/gemma-4-26B-A4B-it general/ptq/nvfp4_experts_only-kv_fp8_cast google/diffusiongemma-26B-A4B-it general/ptq/nvfp4_experts_only-kv_fp8_cast Each engineer described the scheme in prose rather than naming a recipe file, so the mapping was checked against the published checkpoints before being recorded: - All three are NVFP4 with an FP8 KV cache, and every model card documents post-training quantization with no quantization-aware training or distillation. - Kimi-K2.7-Code's exported expert input_scale tensors are exactly 1.0, which is what the input_scale1 recipe produces by pinning amax to a constant 2688.0 (E2M1_MAX * E4M3_MAX); its shared experts and attention are excluded. That distinguishes it from the plain experts-only recipe, whose expert input scales are calibrated. - Both Gemma releases quantize only the routed experts -- dense MLPs, router, attention, vision tower and lm_head all appear in exclude_modules -- with calibrated expert input scales and no exported k_scale, i.e. KV cast mode. diffusiongemma nests its decoder under model.decoder.layers rather than model.language_model.layers, which the recipe's `*.experts.*` patterns match either way. Calibration is the one property a checkpoint cannot reveal, so it rests on what each engineer said: the Gemma pair used the default calibration, which is max. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
The recipe reproduces nvidia/Qwen3.8-27B-NVFP4 -- ptq.md already says so -- but the recipe's own metadata.description did not, which makes it read like an unrelated AutoQuantize artifact when you encounter the file on its own. Every other checkpoint entry names the release it corresponds to, so an audit of backfill coverage that reads the recipes can miss this one. Verified the claim against the published checkpoint before asserting it: MLP gate/up/down_proj NVFP4, self_attn and linear_attn projections FP8, lm_head NVFP4, mtp excluded, no KV quantization, and the model card names local-Hessian calibration on 2,048 samples -- all matching this recipe, including the unusual NVFP4 lm_head. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #2647 +/- ##
==========================================
- Coverage 68.77% 68.68% -0.10%
==========================================
Files 619 619
Lines 69515 69890 +375
==========================================
+ Hits 47807 48001 +194
- Misses 21708 21889 +181
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Type of change: New recipes (checkpoint aliases)
Continues the 2026 published-checkpoint backfill started in #2643, with three more NVFP4 releases. Each entry is a thin alias importing an existing general recipe wholesale — no duplicated
quant_cfg, and editing a base recipe flows through to every alias pointing at it.moonshotai/Kimi-K2.7-Codegeneral/ptq/nvfp4_experts_only_input_scale1-kv_fp8_castnvidia/Kimi-K2.7-Code-NVFP4google/gemma-4-26B-A4B-itgeneral/ptq/nvfp4_experts_only-kv_fp8_castnvidia/Gemma-4-26B-A4B-NVFP4google/diffusiongemma-26B-A4B-itgeneral/ptq/nvfp4_experts_only-kv_fp8_castnvidia/diffusiongemma-26B-A4B-it-NVFP4How these were verified
Unlike #2643, where each engineer named a recipe file outright, these three were described in prose ("NVFP4 experts only, input activation scale 1.0, calibration free"; "nvfp4 experts only, default dataset, 1024 samples"). So the mapping was checked against the published checkpoints before being recorded:
lm_headinexclude_modules, leaving only the routed experts quantized. Kimi-K2.7-Code additionally excludes its shared experts.input_scale1vs plain experts-only. This is what separates the Kimi recipe from the Gemma one. Theinput_scale1recipe pins expert activation amax to a constant 2688.0 (E2M1_MAX * E4M3_MAX), so the exportedinput_scaleis exactly 1.0. Kimi-K2.7-Code's sampled expertinput_scaletensors are exactly 1.0; the Gemma releases' are calibrated values, so they take the plain recipe.k_scaletensors, which is what cast mode produces (use_constant_amax, no KV calibration) — consistent with the already-mergedGLM-5.1/Kimi-K2.6entries, which use-kv_fp8_castrecipes and likewise export none. A calibrated KV cache exports a non-1.0k_scale.diffusiongemmanests its decoder undermodel.decoder.layersrather thanmodel.language_model.layers. The recipe's*.experts.*patterns match either way, so no model-specific body is needed.What this cannot establish: calibration. A checkpoint does not record whether max, MSE, or something else was used, so that rests on what each engineer said — the Gemma pair used the default, which is max. Kimi-K2.7-Code is calibration-free on the activation side by construction.
Second commit: name the release in the Qwen3.8-27B description
models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessianalready reproducesnvidia/Qwen3.8-27B-NVFP4, andptq.mdsays so — but the recipe's ownmetadata.descriptiondid not, so on its own the file reads like an unrelated AutoQuantize artifact. Every other checkpoint entry names its release, so a coverage audit that reads the recipes can miss this one. (It missed mine.)Checked against the published checkpoint before asserting it: MLP
gate/up/down_projNVFP4,self_attnandlinear_attnprojections FP8,lm_headNVFP4,mtp*excluded, no KV quantization — and the model card names local-Hessian calibration on 2,048 samples. All matching, including the unusual NVFP4lm_head.Usage
python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path google/gemma-4-26B-A4B-it \ --recipe models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast \ --export_path <output>Testing
tests/unit/recipe/— 460 passed. All three aliases load throughload_recipe, inheritalgorithm: maxfrom their base, and resolve to 32quant_cfgentries.tests/unit/recipe/test_loader.py, which finds every PTQ recipe from disk, so they are covered without a bespoke test.pre-commitpasses on all changed files.transformer_engine.soin the dev venv; they reproduce identically on unmodifiedmain.Not covered: numerics. Nothing here asserts accuracy or that running a recipe reproduces a released checkpoint's weights bit-for-bit.
Before your PR is "Ready for review"
CONTRIBUTING.md: ✅ — no new dependencies, no source changes.Additional Information
Three more releases from the same tracking list are verified for scope and KV mode but still need their calibration method confirmed by the engineer who produced them (a checkpoint cannot reveal it):
moonshotai/Kimi-K2.5,Qwen/Qwen3.5-397B-A17B(the non-V2 release), andzai-org/GLM-5.1. They will follow separately.🤖 Generated with Claude Code