Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
27 commits
Select commit Hold shift + click to select a range
47a78c6
Hailuo MiniMax H3 training support
Aug 5, 2026
19071f0
fix: constrain duplicated conditioning paths
Aug 5, 2026
3a8bb78
fix: configure SDNQ compile mode before import
Aug 5, 2026
cd81836
feat: add MiniMax H3 drift distillation
Aug 5, 2026
1396552
feat: expand MiniMax H3 training and validation support
Aug 5, 2026
ad79d46
test: align regional Dynamo compile mode expectation
Aug 5, 2026
57b7a2a
examples: add MiniMax H3 ConvRot INT8 VRAM presets
Aug 5, 2026
7cca0d7
fix: default missing SDNQ compile mode to auto
Aug 6, 2026
24020ff
minimax h3 docs update, distillation tutorial
Aug 6, 2026
fa58172
Support mixed-rank LoRA adapter metadata
Aug 6, 2026
606cef3
Respect model flow target direction in distillers
Aug 6, 2026
25c23c9
Avoid double-quantizing SDNQ base models
Aug 6, 2026
e535eeb
Keep fused LoRA adapters tracked after partial unfuse
Aug 6, 2026
fc5fd9b
Fix templated attention backward layout
Aug 6, 2026
6798298
Fix LTX2 VAE up-block channel widths
Aug 6, 2026
0caa74f
Cover LTX2 dynamic shift sequence length
Aug 6, 2026
4fe2b93
Harden validation media export paths
Aug 6, 2026
2ae8194
Preserve validation context in text embed cache
Aug 6, 2026
9002dd9
Allow H3 drift to compose another distiller
Aug 6, 2026
2fcb807
Document H3 drift distiller composition
Aug 6, 2026
adbddcc
Harden MiniMax H3 ConvRot loading and validation
Aug 6, 2026
02c3c4e
Fix Webshart video sample and cache handling
Aug 7, 2026
20664aa
Update MiniMax H3 distillation and validation support
Aug 7, 2026
bbdbbb7
Stabilize Webshart metadata and cache preparation
Aug 7, 2026
ac2830c
Respect active model dropout cache policy in collate
Aug 7, 2026
b108848
Skip H3 drift reference pass when disabled
Aug 7, 2026
a469a52
webshart and validation logging fixes for video training
Aug 7, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.es.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ SimpleTuner es compatible con las siguientes familias de modelos. El soporte det
| **LTX Video 2** | 19B | Apache-2.0 | Sí |
| **Lumina2** | 2B | Apache-2.0 | Sí |
| **Mage-Flow** | 4B | MIT | Sí |
| **MiniMax H3** | 33B | MiniMax H3 Community License | Aplican condiciones (exclusiones territoriales; autorización requerida en EE. UU./UE/Reino Unido/Corea del Sur) |
| **OmniGen** | 3.8B | MIT | Sí |
| **PixArt Sigma** | 0.6B-0.9B | OpenRAIL++ | Sí (restringido) |
| **Qwen Image** | 20B | Apache-2.0 | Sí |
Expand Down
1 change: 1 addition & 0 deletions README.hi.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ SimpleTuner निम्नलिखित मॉडल families का सम
| **LTX Video 2** | 19B | Apache-2.0 | हाँ |
| **Lumina2** | 2B | Apache-2.0 | हाँ |
| **Mage-Flow** | 4B | MIT | हाँ |
| **MiniMax H3** | 33B | MiniMax H3 Community License | शर्तें लागू (territory exclusions; US/EU/UK/KR में authorization आवश्यक) |
| **OmniGen** | 3.8B | MIT | हाँ |
| **PixArt Sigma** | 0.6B-0.9B | OpenRAIL++ | हाँ (restricted) |
| **Qwen Image** | 20B | Apache-2.0 | हाँ |
Expand Down
1 change: 1 addition & 0 deletions README.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ SimpleTunerは以下のモデルファミリーをサポートしています。
| **LTX Video 2** | 19B | Apache-2.0 | 可 |
| **Lumina2** | 2B | Apache-2.0 | 可 |
| **Mage-Flow** | 4B | MIT | 可 |
| **MiniMax H3** | 33B | MiniMax H3 Community License | 条件付き(地域除外あり;米国/EU/英国/韓国は認可が必要) |
| **OmniGen** | 3.8B | MIT | 可 |
| **PixArt Sigma** | 0.6B-0.9B | OpenRAIL++ | 可(制限あり) |
| **Qwen Image** | 20B | Apache-2.0 | 可 |
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ SimpleTuner supports the following model families. Detailed training feature sup
| **LTX Video 2** | 19B | Apache-2.0 | Yes |
| **Lumina2** | 2B | Apache-2.0 | Yes |
| **Mage-Flow** | 4B | MIT | Yes |
| **MiniMax H3** | 33B | MiniMax H3 Community License | Conditions apply (territory exclusions; authorization required in US/EU/UK/KR) |
| **OmniGen** | 3.8B | MIT | Yes |
| **PixArt Sigma** | 0.6B-0.9B | OpenRAIL++ | Yes (restricted) |
| **Qwen Image** | 20B | Apache-2.0 | Yes |
Expand Down
1 change: 1 addition & 0 deletions README.pt-BR.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ SimpleTuner suporta as seguintes familias de modelos. O suporte detalhado a recu
| **LTX Video 2** | 19B | Apache-2.0 | Sim |
| **Lumina2** | 2B | Apache-2.0 | Sim |
| **Mage-Flow** | 4B | MIT | Sim |
| **MiniMax H3** | 33B | MiniMax H3 Community License | Condicoes aplicaveis (exclusoes territoriais; autorizacao exigida nos EUA/UE/Reino Unido/Coreia do Sul) |
| **OmniGen** | 3.8B | MIT | Sim |
| **PixArt Sigma** | 0.6B-0.9B | OpenRAIL++ | Sim (restrito) |
| **Qwen Image** | 20B | Apache-2.0 | Sim |
Expand Down
1 change: 1 addition & 0 deletions README.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ SimpleTuner 支持以下模型系列。详细的训练功能支持请参阅[快
| **LTX Video 2** | 19B | Apache-2.0 | 是 |
| **Lumina2** | 2B | Apache-2.0 | 是 |
| **Mage-Flow** | 4B | MIT | 是 |
| **MiniMax H3** | 33B | MiniMax H3 Community License | 有条件(含地区排除;美国/EU/英国/韩国需要授权) |
| **OmniGen** | 3.8B | MIT | 是 |
| **PixArt Sigma** | 0.6B-0.9B | OpenRAIL++ | 是(受限) |
| **Qwen Image** | 20B | Apache-2.0 | 是 |
Expand Down
6 changes: 3 additions & 3 deletions documentation/DATALOADER.md
Original file line number Diff line number Diff line change
Expand Up @@ -682,8 +682,8 @@ For example, with 4 GPUs, `train_batch_size=4`, and `gradient_accumulation_steps
To automatically adjust `repeats` when your dataset is smaller than the effective batch size, use the `--allow_dataset_oversubscription` flag (documented in [OPTIONS.md](OPTIONS.md#--allow_dataset_oversubscription)).

When enabled, SimpleTuner will:
- Calculate the minimum repeats needed for training
- Automatically increase `repeats` to meet the requirement
- Calculate the minimum repeats needed for each undersized aspect bucket
- Pad only those buckets to meet the effective batch size
- Log a warning showing the adjustment
- **Respect manually-set repeats values** - if you explicitly configure `repeats` in your dataset config, the automatic adjustment will be skipped

Expand Down Expand Up @@ -1370,7 +1370,7 @@ Webshart datasets load WebDataset-style tar shards through the `webshart` packag
- `metadata` is optional and points to a separate metadata location when captions or shard metadata are stored outside the shard source. For Hugging Face metadata repos such as `webshart/conceptual-captions-12m-webdataset-metadata`, pass the repo id; Webshart follows the source shard subfolder layout such as `data/`.
- `metadata_backend` must be `webshart`; it reads dimensions and captions from Webshart metadata.
- `caption_strategy` should be `webshart` to train from metadata captions, or `instanceprompt` to ignore stored captions.
- `webshart.cache_dir` stores SimpleTuner metadata plus Webshart metadata and shard caches. `shard_cache_gb` and `parallel_downloads` are passed to Webshart's shard cache.
- `webshart.cache_dir` stores SimpleTuner metadata plus Webshart metadata and shard caches. `shard_cache_gb` and `parallel_downloads` are passed to Webshart's shard cache; set `shard_cache_gb` to `0` to disable whole-shard caching and retain indexed range reads.

This backend requires a Webshart build with `TarDataLoader.list_shard_sample_aspect_buckets()`.

Expand Down
17 changes: 15 additions & 2 deletions documentation/OPTIONS.es.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,15 @@ Donde `foo` es tu entorno de configuración; o simplemente usa `config/config.js
- `diffusers` es el esquema estándar de PEFT/Diffusers.
- `comfyui` convierte hacia/desde claves estilo ComfyUI (`diffusion_model.*` con tensores `lora_A/lora_B` y `.alpha`). Flux, Flux2, Lumina2 y Z-Image detectarán automáticamente entradas ComfyUI incluso si esto se deja en `diffusers`, pero cámbialo a `comfyui` para forzar salida ComfyUI al guardar.

### `--minimax_h3_target_mode`

- **Qué**: Controla si MiniMax-H3 incluye filas de audio objetivo.
- **Opciones**: `auto`, `video`, `av`
- **Predeterminado**: `auto`
- **Notas**:
- `auto` se resuelve como solo video, omitiendo caché VAE de audio, colación y filas de audio objetivo para H3.
- Define `minimax_h3_target_mode` o `h3_target_mode` como `av` en una entrada de data backend para activar entrenamiento conjunto audio-video en un backend de audio auto-split o explícito.

### `--fuse_qkv_projections`

- **Qué**: Fusiona las proyecciones QKV en los bloques de atención del modelo para un uso más eficiente del hardware.
Expand Down Expand Up @@ -1880,6 +1889,7 @@ usage: train.py [-h] --model_family
[--flow_beta_schedule_beta FLOW_BETA_SCHEDULE_BETA]
[--flow_schedule_shift FLOW_SCHEDULE_SHIFT]
[--flow_schedule_auto_shift [FLOW_SCHEDULE_AUTO_SHIFT]]
[--audio_flow_schedule_shift AUDIO_FLOW_SCHEDULE_SHIFT]
[--flow_custom_timesteps FLOW_CUSTOM_TIMESTEPS]
[--flow_timesteps_mode {fixed-list,round-robin}]
[--flux_guidance_mode {constant,random-range}]
Expand Down Expand Up @@ -2000,7 +2010,7 @@ usage: train.py [-h] --model_family
[--rescale_betas_zero_snr [RESCALE_BETAS_ZERO_SNR]]
[--webhook_config WEBHOOK_CONFIG]
[--webhook_reporting_interval WEBHOOK_REPORTING_INTERVAL]
[--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow}]
[--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow,h3_drift}]
[--distillation_config DISTILLATION_CONFIG]
[--ema_validation {none,ema_only,comparison}]
[--local_rank LOCAL_RANK] [--ltx_train_mode {t2v,i2v}]
Expand Down Expand Up @@ -2351,6 +2361,9 @@ options:
Shift the noise schedule for flow-matching models
--flow_schedule_auto_shift [FLOW_SCHEDULE_AUTO_SHIFT]
Auto-adjust schedule shift based on image resolution
--audio_flow_schedule_shift AUDIO_FLOW_SCHEDULE_SHIFT
Shift the audio noise schedule for flow-matching
models with audio latents
--flow_custom_timesteps FLOW_CUSTOM_TIMESTEPS
Override flow-matching timestep sampling with a fixed
comma-separated list. The list is interpreted as
Expand Down Expand Up @@ -2733,7 +2746,7 @@ options:
Path to webhook configuration file
--webhook_reporting_interval WEBHOOK_REPORTING_INTERVAL
Interval for webhook reports (seconds)
--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow}
--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow,h3_drift}
Method for model distillation
Distillation methods cannot be combined with
--train_text_encoder.
Expand Down
17 changes: 15 additions & 2 deletions documentation/OPTIONS.hi.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,15 @@ simpletuner configure config/foo/config.json
- `diffusers` standard PEFT/Diffusers layout है।
- `comfyui` keys को ComfyUI‑style में convert करता है (`diffusion_model.*` के साथ `lora_A/lora_B` और `.alpha` tensors)। Flux, Flux2, Lumina2, और Z‑Image ComfyUI inputs को auto‑detect करेंगे भले ही यह `diffusers` पर हो, लेकिन saving के लिए ComfyUI output force करने के लिए `comfyui` सेट करें।

### `--minimax_h3_target_mode`

- **What**: MiniMax-H3 target audio rows शामिल करे या नहीं, इसे नियंत्रित करता है।
- **Choices**: `auto`, `video`, `av`
- **Default**: `auto`
- **Notes**:
- `auto` video-only में resolve होता है, जिससे H3 के लिए audio VAE cache, collate, और target audio rows skip होते हैं।
- auto-split या explicit audio backend को joint audio-video training में opt in करने के लिए data backend entry में `minimax_h3_target_mode` या `h3_target_mode` को `av` सेट करें।

### `--fuse_qkv_projections`

- **What**: मॉडल के attention blocks में QKV projections को fuse करता है ताकि hardware का अधिक कुशल उपयोग हो।
Expand Down Expand Up @@ -1878,6 +1887,7 @@ usage: train.py [-h] --model_family
[--flow_beta_schedule_beta FLOW_BETA_SCHEDULE_BETA]
[--flow_schedule_shift FLOW_SCHEDULE_SHIFT]
[--flow_schedule_auto_shift [FLOW_SCHEDULE_AUTO_SHIFT]]
[--audio_flow_schedule_shift AUDIO_FLOW_SCHEDULE_SHIFT]
[--flow_custom_timesteps FLOW_CUSTOM_TIMESTEPS]
[--flow_timesteps_mode {fixed-list,round-robin}]
[--flux_guidance_mode {constant,random-range}]
Expand Down Expand Up @@ -1998,7 +2008,7 @@ usage: train.py [-h] --model_family
[--rescale_betas_zero_snr [RESCALE_BETAS_ZERO_SNR]]
[--webhook_config WEBHOOK_CONFIG]
[--webhook_reporting_interval WEBHOOK_REPORTING_INTERVAL]
[--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow}]
[--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow,h3_drift}]
[--distillation_config DISTILLATION_CONFIG]
[--ema_validation {none,ema_only,comparison}]
[--local_rank LOCAL_RANK] [--ltx_train_mode {t2v,i2v}]
Expand Down Expand Up @@ -2349,6 +2359,9 @@ options:
Shift the noise schedule for flow-matching models
--flow_schedule_auto_shift [FLOW_SCHEDULE_AUTO_SHIFT]
Auto-adjust schedule shift based on image resolution
--audio_flow_schedule_shift AUDIO_FLOW_SCHEDULE_SHIFT
Shift the audio noise schedule for flow-matching
models with audio latents
--flow_custom_timesteps FLOW_CUSTOM_TIMESTEPS
Override flow-matching timestep sampling with a fixed
comma-separated list. The list is interpreted as
Expand Down Expand Up @@ -2731,7 +2744,7 @@ options:
Path to webhook configuration file
--webhook_reporting_interval WEBHOOK_REPORTING_INTERVAL
Interval for webhook reports (seconds)
--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow}
--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow,h3_drift}
Method for model distillation
Distillation methods cannot be combined with
--train_text_encoder.
Expand Down
17 changes: 15 additions & 2 deletions documentation/OPTIONS.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,15 @@ simpletuner configure config/foo/config.json
- `diffusers` は標準の PEFT/Diffusers 形式です。
- `comfyui` は ComfyUI 形式(`diffusion_model.*` と `lora_A/lora_B` + `.alpha`)に変換します。Flux、Flux2、Lumina2、Z-Image は `diffusers` のままでも ComfyUI 入力を自動検出しますが、保存時に ComfyUI 出力を強制したい場合は `comfyui` を指定してください。

### `--minimax_h3_target_mode`

- **内容**: MiniMax-H3 がターゲット音声行を含めるかを制御します。
- **選択肢**: `auto`, `video`, `av`
- **既定**: `auto`
- **注記**:
- `auto` は video-only として扱われ、H3 の audio VAE cache、collate、ターゲット音声行を省略します。
- auto-split または明示的な audio backend で joint audio-video training を使う場合は、data backend entry に `minimax_h3_target_mode` または `h3_target_mode` を `av` として設定します。

### `--fuse_qkv_projections`

- **内容**: モデルのアテンションブロック内 QKV 投影を融合し、ハードウェア効率を高めます。
Expand Down Expand Up @@ -1881,6 +1890,7 @@ usage: train.py [-h] --model_family
[--flow_beta_schedule_beta FLOW_BETA_SCHEDULE_BETA]
[--flow_schedule_shift FLOW_SCHEDULE_SHIFT]
[--flow_schedule_auto_shift [FLOW_SCHEDULE_AUTO_SHIFT]]
[--audio_flow_schedule_shift AUDIO_FLOW_SCHEDULE_SHIFT]
[--flow_custom_timesteps FLOW_CUSTOM_TIMESTEPS]
[--flow_timesteps_mode {fixed-list,round-robin}]
[--flux_guidance_mode {constant,random-range}]
Expand Down Expand Up @@ -2001,7 +2011,7 @@ usage: train.py [-h] --model_family
[--rescale_betas_zero_snr [RESCALE_BETAS_ZERO_SNR]]
[--webhook_config WEBHOOK_CONFIG]
[--webhook_reporting_interval WEBHOOK_REPORTING_INTERVAL]
[--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow}]
[--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow,h3_drift}]
[--distillation_config DISTILLATION_CONFIG]
[--ema_validation {none,ema_only,comparison}]
[--local_rank LOCAL_RANK] [--ltx_train_mode {t2v,i2v}]
Expand Down Expand Up @@ -2351,6 +2361,9 @@ options:
Shift the noise schedule for flow-matching models
--flow_schedule_auto_shift [FLOW_SCHEDULE_AUTO_SHIFT]
Auto-adjust schedule shift based on image resolution
--audio_flow_schedule_shift AUDIO_FLOW_SCHEDULE_SHIFT
Shift the audio noise schedule for flow-matching
models with audio latents
--flow_custom_timesteps FLOW_CUSTOM_TIMESTEPS
Override flow-matching timestep sampling with a fixed
comma-separated list. The list is interpreted as
Expand Down Expand Up @@ -2733,7 +2746,7 @@ options:
Path to webhook configuration file
--webhook_reporting_interval WEBHOOK_REPORTING_INTERVAL
Interval for webhook reports (seconds)
--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow}
--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow,h3_drift}
Method for model distillation
Distillation methods cannot be combined with
--train_text_encoder.
Expand Down
17 changes: 15 additions & 2 deletions documentation/OPTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,15 @@ Where `foo` is your config environment - or just use `config/config.json` if you
- `diffusers` is the standard PEFT/Diffusers layout.
- `comfyui` converts to/from ComfyUI-style keys (`diffusion_model.*` with `lora_A/lora_B` and `.alpha` tensors). Flux, Flux2, Lumina2, and Z-Image will auto-detect ComfyUI inputs even if this is left at `diffusers`, but set it to `comfyui` to force ComfyUI output when saving.

### `--minimax_h3_target_mode`

- **What**: Controls whether MiniMax-H3 includes target audio rows.
- **Choices**: `auto`, `video`, `av`
- **Default**: `auto`
- **Notes**:
- `auto` resolves to video-only, skipping audio VAE caching, collation, and target audio rows for H3.
- Set `minimax_h3_target_mode` or `h3_target_mode` to `av` in a data backend entry to opt an auto-split or explicit audio backend into joint audio-video training.

### `--fuse_qkv_projections`

- **What**: Fuses the QKV projections in the model's attention blocks to make more efficient use of hardware.
Expand Down Expand Up @@ -1884,6 +1893,7 @@ usage: train.py [-h] --model_family
[--flow_beta_schedule_beta FLOW_BETA_SCHEDULE_BETA]
[--flow_schedule_shift FLOW_SCHEDULE_SHIFT]
[--flow_schedule_auto_shift [FLOW_SCHEDULE_AUTO_SHIFT]]
[--audio_flow_schedule_shift AUDIO_FLOW_SCHEDULE_SHIFT]
[--flow_custom_timesteps FLOW_CUSTOM_TIMESTEPS]
[--flow_timesteps_mode {fixed-list,round-robin}]
[--flux_guidance_mode {constant,random-range}]
Expand Down Expand Up @@ -2004,7 +2014,7 @@ usage: train.py [-h] --model_family
[--rescale_betas_zero_snr [RESCALE_BETAS_ZERO_SNR]]
[--webhook_config WEBHOOK_CONFIG]
[--webhook_reporting_interval WEBHOOK_REPORTING_INTERVAL]
[--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow}]
[--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow,h3_drift}]
[--distillation_config DISTILLATION_CONFIG]
[--ema_validation {none,ema_only,comparison}]
[--local_rank LOCAL_RANK] [--ltx_train_mode {t2v,i2v}]
Expand Down Expand Up @@ -2355,6 +2365,9 @@ options:
Shift the noise schedule for flow-matching models
--flow_schedule_auto_shift [FLOW_SCHEDULE_AUTO_SHIFT]
Auto-adjust schedule shift based on image resolution
--audio_flow_schedule_shift AUDIO_FLOW_SCHEDULE_SHIFT
Shift the audio noise schedule for flow-matching
models with audio latents
--flow_custom_timesteps FLOW_CUSTOM_TIMESTEPS
Override flow-matching timestep sampling with a fixed
comma-separated list. The list is interpreted as
Expand Down Expand Up @@ -2739,7 +2752,7 @@ options:
Path to webhook configuration file
--webhook_reporting_interval WEBHOOK_REPORTING_INTERVAL
Interval for webhook reports (seconds)
--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow}
--distillation_method {lcm,dcm,dmd,perflow,flow_dpo,anyflow,h3_drift}
Method for model distillation
Distillation methods cannot be combined with
--train_text_encoder.
Expand Down
Loading