Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.rst
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ Changelog

*Quantization*

- Add Q8_0 weight-only quantization with 32-value GGML blocks, packed unified HF and Megatron export, and a built-in ``q8_0`` PTQ recipe.
- Backfill checkpoint aliases for nine more published NVFP4 releases, so each is reachable from its source model's hub path: ``zai-org/GLM-5.1`` and ``GLM-5.2``, ``MiniMaxAI/MiniMax-M2.5`` and ``MiniMax-M3``, ``deepseek-ai/DeepSeek-V3.1`` and ``DeepSeek-V3.2``, ``Qwen/Qwen3-235B-A22B-Instruct-2507`` and ``-Thinking-2507``, and ``Qwen/Qwen3.6-27B``. Each imports an existing general or architecture recipe wholesale rather than copying its body.
- Add composed Hugging Face AutoQuantize recipes that run fixed PTQ or weight AutoQuantize before
a separate KV-cache AutoQuantize stage, with independent resumable checkpoints for the weight and
Expand Down
40 changes: 23 additions & 17 deletions docs/source/deployment/3_unified_hf.rst
Original file line number Diff line number Diff line change
Expand Up @@ -52,39 +52,41 @@ The unified HF export API supports the following quantization formats:
6. W4A8_AWQ - 4-bit weights and 8-bit activations with AWQ optimization
7. IQ1_S - 1-bit codebook quantization using the GGML block layout
8. IQ2_XS - 2-bit codebook quantization using the GGML block layout
9. Q8_0 - 8-bit symmetric integer quantization using the GGML block layout

.. note::
GGML has no equivalent for ModelOpt's per-tensor FP8 weight-and-activation format. In particular,
GGML does not define a first-class FP8 tensor type with the corresponding per-tensor weight and
activation scale semantics. Converting a ModelOpt FP8 checkpoint to GGUF therefore requires
conversion to another GGML-supported tensor type rather than a lossless FP8 encoding.

IQ weight representation
~~~~~~~~~~~~~~~~~~~~~~~~
GGML weight representation
~~~~~~~~~~~~~~~~~~~~~~~~~~

For IQ1_S and IQ2_XS, unified export replaces each floating-point ``<module>.weight`` with a
``uint8`` tensor containing byte-exact GGML blocks. Its shape is
``[*logical_shape[:-1], logical_shape[-1] // 256, payload_bytes]``, where ``payload_bytes`` is 50
for IQ1_S and 74 for IQ2_XS. No separate shape tensor is stored: a loader recovers the logical
shape as ``[*weight.shape[:-2], weight.shape[-2] * 256]``. This is unambiguous because IQ export
requires the logical last dimension to be divisible by 256.
For IQ1_S, IQ2_XS, and Q8_0, unified export replaces each floating-point
``<module>.weight`` with a ``uint8`` tensor containing byte-exact GGML blocks. Its shape is
``[*logical_shape[:-1], logical_shape[-1] // block_size, payload_bytes]``. The block size and
payload size are 256 and 50 for IQ1_S, 256 and 74 for IQ2_XS, and 32 and 34 for Q8_0. No separate
shape tensor is stored: a loader recovers the logical shape as
``[*weight.shape[:-2], weight.shape[-2] * block_size]``. This is unambiguous because export
requires each logical row to contain complete blocks.

.. note::
Megatron IQ export currently requires tensor and pipeline model parallel sizes of 1. Packing
Megatron GGML export currently requires tensor and pipeline model parallel sizes of 1. Packing
happens during export, so a tensor-parallel shard would be packed as if it were a whole
weight, and a pipeline stage holding no IQ layer would not reach the same rejection as its
weight, and a pipeline stage holding no GGML layer would not reach the same rejection as its
peers. Expert parallelism is supported, assuming every expert uses the same format.

.. warning::
Megatron fused-MoE IQ export is not currently supported. Its packed tensor would require the
Megatron fused-MoE GGML export is not currently supported. Its packed tensor would require the
deployment consumer to understand
``[num_experts, out_features, in_features // 256, payload_bytes]`` rather than the ordinary HF
fused-expert order. The exporter raises ``NotImplementedError`` until a deployment loader owns
this layout and is covered by an integration test. Dense and individually named expert weights
continue to use the representation above.
``[num_experts, out_features, in_features // block_size, payload_bytes]`` rather than the
ordinary HF fused-expert order. The exporter raises ``NotImplementedError`` until a
deployment loader owns this layout and is covered by an integration test. Dense and
individually named expert weights continue to use the representation above.

The generated configuration records ``quant_method: modelopt``, ``packing: ggml``, the 256-value
block size, and the payload byte count. IQ payloads are not represented as compressed-tensors
The generated configuration records ``quant_method: modelopt``, ``packing: ggml``, the format's
block size, and the payload byte count. GGML payloads are not represented as compressed-tensors
integer ``weights`` groups because all scales and indices are embedded in each packed block.

Each 74-byte IQ2_XS block represents 256 logical weights:
Expand All @@ -99,6 +101,10 @@ Each 74-byte IQ2_XS block represents 256 logical weights:
The canonical 512-by-8 IQ2_XS codebook is part of the implementation rather than the checkpoint.
The complete block therefore costs ``74 * 8 / 256 = 2.3125`` bits per logical weight.

Each 34-byte Q8_0 block represents 32 logical weights: bytes 0--1 hold the little-endian FP16
scale, and bytes 2--33 hold 32 signed int8 quants. The block costs
``34 * 8 / 32 = 8.5`` bits per logical weight.

Minimum Framework Versions
--------------------------

Expand Down
22 changes: 11 additions & 11 deletions modelopt/torch/export/convert_hf_config.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,9 +19,9 @@
from collections import defaultdict
from typing import Any

from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY
from modelopt.torch.quantization.ggml import GGML_FORMAT_REGISTRY

from .quant_format import IQ_FORMATS
from .quant_format import GGML_FORMATS


def _quant_algo_to_group_config(quant_algo: str, group_size: int | None = None) -> dict[str, Any]:
Expand All @@ -34,7 +34,7 @@ def _quant_algo_to_group_config(quant_algo: str, group_size: int | None = None)
Returns:
Dictionary with ``input_activations`` and ``weights`` entries suitable for
a compressed-tensors ``config_groups`` entry, or ModelOpt-owned metadata for
self-contained IQ payloads.
self-contained GGML payloads.
"""
if quant_algo == "FP8":
return {
Expand Down Expand Up @@ -122,13 +122,13 @@ def _quant_algo_to_group_config(quant_algo: str, group_size: int | None = None)
},
"weights": {"dynamic": False, "num_bits": 8, "type": "float", "group_size": gs},
}
elif quant_algo.lower() in IQ_FORMATS:
iq_format = IQ_FORMAT_REGISTRY[quant_algo.lower()]
block_size, payload_bytes = iq_format.block_size, iq_format.block_bytes
effective_bits = iq_format.effective_bits
elif quant_algo.lower() in GGML_FORMATS:
ggml_format = GGML_FORMAT_REGISTRY[quant_algo.lower()]
block_size, payload_bytes = ggml_format.block_size, ggml_format.block_bytes
effective_bits = ggml_format.effective_bits
if group_size not in (None, block_size):
raise ValueError(f"{quant_algo} requires group size {block_size}, got {group_size}")
# IQ payloads are self-contained blocks, not compressed-tensors integer groups.
# GGML payloads are self-contained blocks, not compressed-tensors integer groups.
# Keep their format marker outside a ``weights`` quantization scheme.
return {
"quant_algo": quant_algo,
Expand Down Expand Up @@ -229,13 +229,13 @@ def convert_hf_quant_config_format(input_config: dict[str, Any]) -> dict[str, An
"targets": ["Linear"],
}
new_config["config_groups"] = {"group_0": config_group_details}
elif str(quant_algo_value).lower() in IQ_FORMATS:
elif str(quant_algo_value).lower() in GGML_FORMATS:
# Forward the caller's group size so a mismatched one is rejected rather than rewritten
# to the format's block size.
iq_metadata = _quant_algo_to_group_config(
ggml_metadata = _quant_algo_to_group_config(
quant_algo_value, original_quantization_details.get("group_size")
)
new_config.update(iq_metadata)
new_config.update(ggml_metadata)
elif quant_algo_value == "NVFP4_SVD":
# NVFP4 + SVDQuant: NVFP4 weights/activations plus an AWQ-style
# pre_quant_scale and a low-rank residual (svdquant_lora_a/b) stored as
Expand Down
14 changes: 8 additions & 6 deletions modelopt/torch/export/quant_format.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@
constants, for example, are in :mod:`modelopt.torch.export.trtllm.model_config`.
"""

from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY
from modelopt.torch.quantization.ggml import GGML_FORMAT_REGISTRY, IQ_FORMAT_REGISTRY

QUANTIZATION_NONE = None
QUANTIZATION_FP8 = "fp8"
Expand All @@ -43,22 +43,24 @@
QUANTIZATION_IQ2_XXS = "iq2_xxs"
QUANTIZATION_IQ2_XS = "iq2_xs"
QUANTIZATION_IQ2_S = "iq2_s"
QUANTIZATION_Q8_0 = "q8_0"

# Every GGML IQ format, derived from the registry the quantization backend dispatches through, so
# export and dispatch cannot disagree about which formats exist. They share the weight-only,
# 256-value-block, per-module-scale shape, so export treats them as one family. A format's block
# geometry and packer are read from IQ_FORMAT_REGISTRY directly.
# Every GGML format is derived from the registry the quantization backend dispatches through, so
# export and dispatch cannot disagree about which formats exist. A format's block geometry and
# packer are read from GGML_FORMAT_REGISTRY directly. IQ_FORMATS remains as the compatibility
# subset used by IQ-specific conformance tests.
#
# Registering a format therefore declares it exportable, and that is intended rather than a side
# effect: fake quant is dequantize(quantize(w)), so a format cannot be dispatched without the
# packer and block geometry that are all export reads.
IQ_FORMATS = frozenset(IQ_FORMAT_REGISTRY)
Comment thread
hychiang-git marked this conversation as resolved.
GGML_FORMATS = frozenset(GGML_FORMAT_REGISTRY)
Comment thread
coderabbitai[bot] marked this conversation as resolved.


# Formats whose scales are purely per-module, so export never merges them across the q/k/v
# and gate/up groups that share an input. Every other format unifies input_amax (and, for
# NVFP4, weight_scale_2) across such a group, which only a whole-model forward can discover.
FUSION_FREE_FORMATS = IQ_FORMATS | frozenset(
FUSION_FREE_FORMATS = GGML_FORMATS | frozenset(
{
QUANTIZATION_FP8,
QUANTIZATION_NONE,
Expand Down
32 changes: 16 additions & 16 deletions modelopt/torch/export/quant_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@

from modelopt import __version__
from modelopt.torch.models import get_spec, list_all_possible
from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY
from modelopt.torch.quantization.ggml import GGML_FORMAT_REGISTRY
from modelopt.torch.quantization.model_calib import (
enable_stats_collection,
finish_stats_collection,
Expand All @@ -52,7 +52,7 @@
from ..quantization.nn import NVFP4StaticQuantizer, SequentialQuantizer, TensorQuantizer
from .model_utils import TiedWeightMap, get_language_model_from_vl
from .quant_format import (
IQ_FORMATS,
GGML_FORMATS,
KV_CACHE_FP8,
KV_CACHE_FP8_K_NVFP4_V,
KV_CACHE_INT8,
Expand Down Expand Up @@ -444,23 +444,23 @@ def get_weight_block_size(module: nn.Module, weight_name: str = "weight") -> int


def uses_iq_quantization(module) -> bool:
"""Whether any weight quantizer in ``module`` or its children targets an IQ format.
"""Whether any weight quantizer in ``module`` or its children targets a GGML format.

``get_quantization_format`` returns the *first* non-``NONE`` format it finds, so in a
mixed-format model IQ layers sitting behind, say, an FP8 layer are invisible to it. Callers
that must reject IQ specifically need to see every layer.
mixed-format model GGML layers sitting behind, say, an FP8 layer are invisible to it. Callers
that must reject GGML specifically need to see every layer.

This reads ``num_bits`` directly rather than resolving each layer's full format, so an
unrelated unsupported quantizer elsewhere in the model cannot turn the check into an error.
"""
for weight_name in weight_attr_names(module):
weight_quantizer = representative_weight_quantizer(module, weight_name)
# getattr: a SequentialQuantizer has is_enabled but no num_bits, and is never IQ --
# IQ is a single quantizer with backend="ggml".
# getattr: a SequentialQuantizer has is_enabled but no num_bits, and is never GGML --
# GGML is a single quantizer with backend="ggml".
if (
weight_quantizer is not None
and weight_quantizer.is_enabled
and getattr(weight_quantizer, "num_bits", None) in IQ_FORMATS
and getattr(weight_quantizer, "num_bits", None) in GGML_FORMATS
):
return True
return any(uses_iq_quantization(child) for _, child in module.named_children())
Expand Down Expand Up @@ -500,21 +500,21 @@ def _get_quantization_from_layer(layer, quantizer_attr_names: QuantizerAttrNames
return QUANTIZATION_W4A8_AWQ

# Handle individual num_bits cases
if weight_quantizer.num_bits in IQ_FORMATS:
if weight_quantizer.num_bits in GGML_FORMATS:
if weight_quantizer.backend != "ggml":
raise ValueError("IQ formats require the built-in 'ggml' quantization backend")
raise ValueError("GGML formats require the built-in 'ggml' quantization backend")
# Both exporters return before collecting input_scale and before the pre_quant_scale
# handling below, so an enabled activation quantizer would be dropped without a trace
# and the checkpoint would load as weight-only. Refuse instead.
if input_quantizer is not None and input_quantizer.is_enabled:
raise NotImplementedError(
"IQ1_S/IQ2_XS export is weight-only, but this layer has an enabled input "
"GGML export is weight-only, but this layer has an enabled input "
"quantizer. The GGML block payload carries no activation scale, so the "
"activation quantization would be silently lost."
)
if input_quantizer is not None and hasattr(input_quantizer, "_pre_quant_scale"):
raise NotImplementedError(
"IQ1_S/IQ2_XS export does not support an AWQ-style pre_quant_scale."
"GGML export does not support an AWQ-style pre_quant_scale."
)
return weight_quantizer.num_bits

Expand Down Expand Up @@ -766,10 +766,10 @@ def process_layer_quant_config(layer_config_dict):
"quant_algo": "MXFP8",
"group_size": block_size_value,
}
elif v in IQ_FORMATS:
iq_format = IQ_FORMAT_REGISTRY[v]
block_size, payload_bytes = iq_format.block_size, iq_format.block_bytes
effective_bits = iq_format.effective_bits
elif v in GGML_FORMATS:
ggml_format = GGML_FORMAT_REGISTRY[v]
block_size, payload_bytes = ggml_format.block_size, ggml_format.block_bytes
effective_bits = ggml_format.effective_bits
if block_size_value != block_size:
raise ValueError(
f"{v.upper()} requires block size {block_size}, got {block_size_value}"
Expand Down
10 changes: 5 additions & 5 deletions modelopt/torch/export/unified_export_hf.py
Original file line number Diff line number Diff line change
Expand Up @@ -67,7 +67,7 @@
from modelopt.torch.opt.conversion import ModeloptStateManager, modelopt_state
from modelopt.torch.opt.plugins.huggingface import _MODELOPT_STATE_SAVE_NAME
from modelopt.torch.quantization import set_quantizer_by_cfg_context
from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY
from modelopt.torch.quantization.ggml import GGML_FORMAT_REGISTRY
from modelopt.torch.quantization.nn import SequentialQuantizer, TensorQuantizer
from modelopt.torch.quantization.qtensor import MXFP8QTensor, NVFP4QTensor
from modelopt.torch.quantization.qtensor.base_qtensor import QTensorWrapper
Expand Down Expand Up @@ -101,7 +101,7 @@
)
from .quant_format import (
FUSION_FREE_FORMATS,
IQ_FORMATS,
GGML_FORMATS,
QUANTIZATION_FP8,
QUANTIZATION_FP8_PB_REAL,
QUANTIZATION_FP8_PC_PT,
Expand Down Expand Up @@ -633,13 +633,13 @@ def _export_quantized_weight(
"which dispatches to the streaming writer that materialises weights layer-by-layer."
)

if quantization_format in IQ_FORMATS:
if quantization_format in GGML_FORMATS:
if weight_name != "weight":
raise NotImplementedError(
"IQ unified export currently supports modules with a standard 'weight' "
"GGML unified export currently supports modules with a standard 'weight' "
f"attribute, got {weight_name!r} on {type(sub_module).__name__}"
)
packed_weight = IQ_FORMAT_REGISTRY[quantization_format].pack(
packed_weight = GGML_FORMAT_REGISTRY[quantization_format].pack(
weight.to(dtype), getattr(sub_module, quantizer_attrs.weight_quantizer, None)
)
setattr(sub_module, weight_name, nn.Parameter(packed_weight, requires_grad=False))
Expand Down
Loading
Loading