Skip to content

Backfill checkpoint aliases for Kimi-K2.7-Code and the Gemma-4 MoE releases - #2647

Draft
shengliangxu wants to merge 2 commits into
mainfrom
shengliangx/backfill-2026-group-a
Draft

shengliangxu wants to merge 2 commits into
mainfrom
shengliangx/backfill-2026-group-a

Conversation

@shengliangxu

@shengliangxu shengliangxu commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Type of change: New recipes (checkpoint aliases)

Continues the 2026 published-checkpoint backfill started in #2643, with three more NVFP4 releases. Each entry is a thin alias importing an existing general recipe wholesale — no duplicated quant_cfg, and editing a base recipe flows through to every alias pointing at it.

Source model Recipe Published as
moonshotai/Kimi-K2.7-Code general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast nvidia/Kimi-K2.7-Code-NVFP4
google/gemma-4-26B-A4B-it general/ptq/nvfp4_experts_only-kv_fp8_cast nvidia/Gemma-4-26B-A4B-NVFP4
google/diffusiongemma-26B-A4B-it general/ptq/nvfp4_experts_only-kv_fp8_cast nvidia/diffusiongemma-26B-A4B-it-NVFP4

How these were verified

Unlike #2643, where each engineer named a recipe file outright, these three were described in prose ("NVFP4 experts only, input activation scale 1.0, calibration free"; "nvfp4 experts only, default dataset, 1024 samples"). So the mapping was checked against the published checkpoints before being recorded:

  • Format and method. All three are NVFP4 with an FP8 KV cache, and every model card documents post-training quantization with no mention of quantization-aware training or distillation. That check matters because QAD retrains the weights while leaving the quantization layout identical — a PTQ recipe could match every module and format and still not reproduce such a release.
  • Scope. Both Gemma releases list dense MLPs, router, attention, vision tower and lm_head in exclude_modules, leaving only the routed experts quantized. Kimi-K2.7-Code additionally excludes its shared experts.
  • input_scale1 vs plain experts-only. This is what separates the Kimi recipe from the Gemma one. The input_scale1 recipe pins expert activation amax to a constant 2688.0 (E2M1_MAX * E4M3_MAX), so the exported input_scale is exactly 1.0. Kimi-K2.7-Code's sampled expert input_scale tensors are exactly 1.0; the Gemma releases' are calibrated values, so they take the plain recipe.
  • KV mode. Neither Gemma release exports k_scale tensors, which is what cast mode produces (use_constant_amax, no KV calibration) — consistent with the already-merged GLM-5.1/Kimi-K2.6 entries, which use -kv_fp8_cast recipes and likewise export none. A calibrated KV cache exports a non-1.0 k_scale.
  • A naming quirk worth knowing: diffusiongemma nests its decoder under model.decoder.layers rather than model.language_model.layers. The recipe's *.experts.* patterns match either way, so no model-specific body is needed.

What this cannot establish: calibration. A checkpoint does not record whether max, MSE, or something else was used, so that rests on what each engineer said — the Gemma pair used the default, which is max. Kimi-K2.7-Code is calibration-free on the activation side by construction.

Second commit: name the release in the Qwen3.8-27B description

models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_local_hessian already reproduces nvidia/Qwen3.8-27B-NVFP4, and ptq.md says so — but the recipe's own metadata.description did not, so on its own the file reads like an unrelated AutoQuantize artifact. Every other checkpoint entry names its release, so a coverage audit that reads the recipes can miss this one. (It missed mine.)

Checked against the published checkpoint before asserting it: MLP gate/up/down_proj NVFP4, self_attn and linear_attn projections FP8, lm_head NVFP4, mtp* excluded, no KV quantization — and the model card names local-Hessian calibration on 2,048 samples. All matching, including the unusual NVFP4 lm_head.

Usage

python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path google/gemma-4-26B-A4B-it \
    --recipe models/google/gemma-4-26B-A4B-it/ptq/nvfp4_experts_only-kv_fp8_cast \
    --export_path <output>

Testing

  • tests/unit/recipe/ — 460 passed. All three aliases load through load_recipe, inherit algorithm: max from their base, and resolve to 32 quant_cfg entries.
  • The aliases are swept by the shipped-recipe discovery in tests/unit/recipe/test_loader.py, which finds every PTQ recipe from disk, so they are covered without a bespoke test.
  • pre-commit passes on all changed files.
  • The 14 failures in the local run are pre-existing and environment-only — a broken transformer_engine .so in the dev venv; they reproduce identically on unmodified main.

Not covered: numerics. Nothing here asserts accuracy or that running a recipe reproduces a released checkpoint's weights bit-for-bit.

Before your PR is "Ready for review"

  • Is this change backward compatible?: ✅ — purely additive. Three new alias files plus a docs and changelog entry; no existing recipe path, body or API changes.
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: ✅ — no new dependencies, no source changes.
  • Did you write any new necessary tests?: ✅ — covered by the shipped-recipe sweep.
  • Did you update Changelog?: ✅
  • Did you get Claude approval on this PR?: ❌ — not yet run.

Additional Information

Three more releases from the same tracking list are verified for scope and KV mode but still need their calibration method confirmed by the engineer who produced them (a checkpoint cannot reveal it): moonshotai/Kimi-K2.5, Qwen/Qwen3.5-397B-A17B (the non-V2 release), and zai-org/GLM-5.1. They will follow separately.

🤖 Generated with Claude Code

Continues the 2026 published-checkpoint backfill. Each entry imports an existing
general recipe wholesale; nothing is duplicated.

  moonshotai/Kimi-K2.7-Code         general/ptq/nvfp4_experts_only_input_scale1-kv_fp8_cast
  google/gemma-4-26B-A4B-it         general/ptq/nvfp4_experts_only-kv_fp8_cast
  google/diffusiongemma-26B-A4B-it  general/ptq/nvfp4_experts_only-kv_fp8_cast

Each engineer described the scheme in prose rather than naming a recipe file, so
the mapping was checked against the published checkpoints before being recorded:

- All three are NVFP4 with an FP8 KV cache, and every model card documents
  post-training quantization with no quantization-aware training or distillation.
- Kimi-K2.7-Code's exported expert input_scale tensors are exactly 1.0, which is
  what the input_scale1 recipe produces by pinning amax to a constant 2688.0
  (E2M1_MAX * E4M3_MAX); its shared experts and attention are excluded. That
  distinguishes it from the plain experts-only recipe, whose expert input scales
  are calibrated.
- Both Gemma releases quantize only the routed experts -- dense MLPs, router,
  attention, vision tower and lm_head all appear in exclude_modules -- with
  calibrated expert input scales and no exported k_scale, i.e. KV cast mode.
  diffusiongemma nests its decoder under model.decoder.layers rather than
  model.language_model.layers, which the recipe's `*.experts.*` patterns match
  either way.

Calibration is the one property a checkpoint cannot reveal, so it rests on what
each engineer said: the Gemma pair used the default calibration, which is max.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 2, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

The recipe reproduces nvidia/Qwen3.8-27B-NVFP4 -- ptq.md already says so -- but
the recipe's own metadata.description did not, which makes it read like an
unrelated AutoQuantize artifact when you encounter the file on its own. Every
other checkpoint entry names the release it corresponds to, so an audit of
backfill coverage that reads the recipes can miss this one.

Verified the claim against the published checkpoint before asserting it: MLP
gate/up/down_proj NVFP4, self_attn and linear_attn projections FP8, lm_head
NVFP4, mtp excluded, no KV quantization, and the model card names local-Hessian
calibration on 2,048 samples -- all matching this recipe, including the unusual
NVFP4 lm_head.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
@codecov

codecov Bot commented Oct 2, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 68.68%. Comparing base (4e19e0b) to head (553e45d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2647      +/-   ##
==========================================
- Coverage   68.77%   68.68%   -0.10%     
==========================================
  Files         619      619              
  Lines       69515    69890     +375     
==========================================
+ Hits        47807    48001     +194     
- Misses      21708    21889     +181     
Flag Coverage Δ
unit 58.72% <ø> (-0.02%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant