Describe the bug
GGUFParameter.__new__ accepts quant_type=None as its default and then immediately performs an
unguarded dict lookup on it:
# src/diffusers/quantizers/gguf/utils.py
def __new__(cls, data, requires_grad=False, quant_type=None):
data = data if data is not None else torch.empty(0)
self = torch.Tensor._make_subclass(cls, data, requires_grad)
self.quant_type = quant_type
block_size, type_size = GGML_QUANT_SIZES[quant_type] # <-- KeyError when quant_type is None
self.quant_shape = _quant_shape_from_byte_shape(self.shape, type_size, block_size)
return self
accelerate's offload path (set_module_tensor_to_device) re-wraps parameters as
param_cls(new_value, requires_grad=old_value.requires_grad) — without forwarding
quant_type. Every such re-wrap therefore raises KeyError: None.
Impact. This makes enable_sequential_cpu_offload() unusable with any GGUF-quantised
transformer. Because enable_model_cpu_offload() moves the whole transformer to the GPU at once
(and fails with a driver-level OOM when it does not fit), the practical effect is that no GGUF
model larger than available VRAM can be run at all — precisely the case GGUF quantisation exists
to serve. On an 8 GB card this ruled out FLUX.1-schnell Q8_0 (12.7 GB) and Chroma1-HD Q8_0 (9.7 GB).
The error is also misleading: KeyError: None surfacing from inside accelerate reads as a corrupt
or unsupported model file. We wrote off three different models as "broken quants" before running a
known-good file under seq offload and finding it failed identically.
Reproduction
Minimal — no model download required:
import torch
from diffusers.quantizers.gguf.utils import GGUFParameter
p = GGUFParameter(torch.zeros(32, 32), quant_type=8) # Q8_0
GGUFParameter(p, requires_grad=False) # KeyError: None
End-to-end:
import torch
from diffusers import FluxPipeline, FluxTransformer2DModel, GGUFQuantizationConfig
tr = FluxTransformer2DModel.from_single_file(
"https://huggingface.co/city96/FLUX.1-schnell-gguf/blob/main/flux1-schnell-Q4_K_S.gguf",
quantization_config=GGUFQuantizationConfig(compute_dtype=torch.bfloat16),
torch_dtype=torch.bfloat16,
config="black-forest-labs/FLUX.1-schnell", subfolder="transformer",
)
pipe = FluxPipeline.from_pretrained("black-forest-labs/FLUX.1-schnell",
transformer=tr, torch_dtype=torch.bfloat16)
pipe.enable_sequential_cpu_offload() # works with enable_model_cpu_offload()
pipe("a cat", num_inference_steps=1) # KeyError: None
Suggested fix
Inherit quant_type from the incoming tensor when it is not passed explicitly, and tolerate a
genuinely-absent type rather than raising. quant_shape is assigned but never read elsewhere in
diffusers, so leaving it None in that case is safe.
if quant_type is None:
quant_type = getattr(data, "quant_type", None)
self = torch.Tensor._make_subclass(cls, data, requires_grad)
self.quant_type = quant_type
if quant_type is None:
self.quant_shape = None
else:
block_size, type_size = GGML_QUANT_SIZES[quant_type]
self.quant_shape = _quant_shape_from_byte_shape(self.shape, type_size, block_size)
_extract_quant_type should probably also inspect kwargs, not only positional args.
Result after patching
enable_sequential_cpu_offload() works, and peak VRAM for FLUX.1-schnell Q4_K_S drops from
7.4 GB to 1.81 GB (26.9s vs 18s per 1024x1024 image, 4 steps, RTX 5070 Laptop 8 GB).
Models that previously could not be loaded at all now run.
System info
- diffusers 0.40.0, torch 2.11.0+cu128, accelerate (bundled), gguf 0.19.0
- Python 3.12, Linux (WSL2), RTX 5070 Laptop 8 GB (sm_120)
Describe the bug
GGUFParameter.__new__acceptsquant_type=Noneas its default and then immediately performs anunguarded dict lookup on it:
accelerate's offload path (set_module_tensor_to_device) re-wraps parameters asparam_cls(new_value, requires_grad=old_value.requires_grad)— without forwardingquant_type. Every such re-wrap therefore raisesKeyError: None.Impact. This makes
enable_sequential_cpu_offload()unusable with any GGUF-quantisedtransformer. Because
enable_model_cpu_offload()moves the whole transformer to the GPU at once(and fails with a driver-level OOM when it does not fit), the practical effect is that no GGUF
model larger than available VRAM can be run at all — precisely the case GGUF quantisation exists
to serve. On an 8 GB card this ruled out FLUX.1-schnell Q8_0 (12.7 GB) and Chroma1-HD Q8_0 (9.7 GB).
The error is also misleading:
KeyError: Nonesurfacing from insideacceleratereads as a corruptor unsupported model file. We wrote off three different models as "broken quants" before running a
known-good file under
seqoffload and finding it failed identically.Reproduction
Minimal — no model download required:
End-to-end:
Suggested fix
Inherit
quant_typefrom the incoming tensor when it is not passed explicitly, and tolerate agenuinely-absent type rather than raising.
quant_shapeis assigned but never read elsewhere indiffusers, so leaving it
Nonein that case is safe._extract_quant_typeshould probably also inspectkwargs, not only positionalargs.Result after patching
enable_sequential_cpu_offload()works, and peak VRAM for FLUX.1-schnell Q4_K_S drops from7.4 GB to 1.81 GB (26.9s vs 18s per 1024x1024 image, 4 steps, RTX 5070 Laptop 8 GB).
Models that previously could not be loaded at all now run.
System info