Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions documentation/OPTIONS.es.md
Original file line number Diff line number Diff line change
Expand Up @@ -253,6 +253,13 @@ Donde `foo` es tu entorno de configuración; o simplemente usa `config/config.js
- **Qué**: Ruta al modelo Gemma preentrenado o su identificador en <https://huggingface.co/models>.
- **Por qué**: Al entrenar modelos basados en Gemma (por ejemplo LTX-2, Sana o Lumina2), puedes apuntar a un checkpoint Gemma compartido sin cambiar la ruta del modelo base de difusión.

### `--qwen_text_encoder_model_name_or_path`

- **Qué**: Ruta a un codificador de texto Qwen preentrenado o su identificador en <https://huggingface.co/models>.
- **Predeterminado**: `None` (usa la fuente del codificador de texto Qwen definida por el modelo seleccionado).
- **Por qué**: Úsalo para compartir o reemplazar el codificador de texto Qwen en familias de modelos basadas en Qwen sin editar la caché de Hugging Face.
- **Notas**: Se aplica a familias de modelos con un solo codificador de texto Qwen. Si una familia define varios codificadores Qwen, la opción se ignora y SimpleTuner registra una advertencia.

### `--max_grounding_entities`
- Numero maximo de entidades de grounding por imagen para anotaciones espaciales estilo GLIGEN. Por defecto: 0 (deshabilitado). Valores tipicos: 4-16.

Expand Down Expand Up @@ -1748,6 +1755,7 @@ usage: train.py [-h] --model_family
[--pretrained_unet_subfolder PRETRAINED_UNET_SUBFOLDER]
[--pretrained_t5_model_name_or_path PRETRAINED_T5_MODEL_NAME_OR_PATH]
[--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH]
[--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH]
[--revision REVISION] [--variant VARIANT]
[--base_model_default_dtype {bf16,fp32}]
[--unet_attention_slice [UNET_ATTENTION_SLICE]]
Expand Down Expand Up @@ -2079,6 +2087,8 @@ options:
Path to pretrained T5 model
--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH
Path to pretrained Gemma model
--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH
Path to pretrained Qwen text encoder model
--revision REVISION Git branch/tag/commit for model version
--variant VARIANT Model variant (e.g., fp16, bf16)
--base_model_default_dtype {bf16,fp32}
Expand Down
10 changes: 10 additions & 0 deletions documentation/OPTIONS.hi.md
Original file line number Diff line number Diff line change
Expand Up @@ -253,6 +253,13 @@ simpletuner configure config/foo/config.json
- **What**: pretrained Gemma model का path या <https://huggingface.co/models> से उसका identifier.
- **Why**: Gemma‑based models (जैसे LTX-2, Sana, Lumina2) ट्रेन करते समय आप base diffusion model path बदले बिना Gemma weights का source specify कर सकते हैं।

### `--qwen_text_encoder_model_name_or_path`

- **What**: pretrained Qwen text encoder model का path या <https://huggingface.co/models> से उसका identifier.
- **Default**: `None` (selected model में defined Qwen text encoder source उपयोग होता है).
- **Why**: Qwen-based model families में Qwen text encoder को share या replace करने के लिए इसका उपयोग करें, बिना Hugging Face cache edit किए।
- **Notes**: यह उन model families पर लागू होता है जिनमें एक Qwen text encoder है। अगर कोई model family multiple Qwen encoders define करती है, तो option ignore होता है और SimpleTuner warning log करता है।

### `--max_grounding_entities`
- GLIGEN-style spatial annotations के लिए प्रति image grounding entities की अधिकतम संख्या। Default: 0 (disabled)। सामान्य मान: 4-16।

Expand Down Expand Up @@ -1746,6 +1753,7 @@ usage: train.py [-h] --model_family
[--pretrained_unet_subfolder PRETRAINED_UNET_SUBFOLDER]
[--pretrained_t5_model_name_or_path PRETRAINED_T5_MODEL_NAME_OR_PATH]
[--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH]
[--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH]
[--revision REVISION] [--variant VARIANT]
[--base_model_default_dtype {bf16,fp32}]
[--unet_attention_slice [UNET_ATTENTION_SLICE]]
Expand Down Expand Up @@ -2077,6 +2085,8 @@ options:
Path to pretrained T5 model
--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH
Path to pretrained Gemma model
--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH
Path to pretrained Qwen text encoder model
--revision REVISION Git branch/tag/commit for model version
--variant VARIANT Model variant (e.g., fp16, bf16)
--base_model_default_dtype {bf16,fp32}
Expand Down
10 changes: 10 additions & 0 deletions documentation/OPTIONS.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -254,6 +254,13 @@ simpletuner configure config/foo/config.json
- **内容**: 事前学習済み Gemma モデルのパス、または <https://huggingface.co/models> の識別子。
- **理由**: Gemma 系モデル(例: LTX-2、Sana、Lumina2)を学習する際、ベース拡散モデルのパスを変えずに Gemma 重みの参照先を指定できます。

### `--qwen_text_encoder_model_name_or_path`

- **内容**: 事前学習済み Qwen テキストエンコーダーモデルのパス、または <https://huggingface.co/models> の識別子。
- **既定**: `None`(選択したモデルが定義する Qwen テキストエンコーダーの参照元を使用します)
- **理由**: Hugging Face キャッシュを編集せずに、Qwen 系モデルファミリーの Qwen テキストエンコーダーを共有または置き換えるために使用します。
- **注記**: Qwen テキストエンコーダーが 1 つのモデルファミリーに適用されます。複数の Qwen エンコーダーを定義するモデルファミリーでは、このオプションは無視され、SimpleTuner が警告を記録します。

### `--max_grounding_entities`
- GLIGEN スタイルの空間アノテーション用に、画像あたりのグラウンディングエンティティの最大数を指定します。デフォルト: 0(無効)。一般的な値: 4-16。

Expand Down Expand Up @@ -1749,6 +1756,7 @@ usage: train.py [-h] --model_family
[--pretrained_unet_subfolder PRETRAINED_UNET_SUBFOLDER]
[--pretrained_t5_model_name_or_path PRETRAINED_T5_MODEL_NAME_OR_PATH]
[--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH]
[--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH]
[--revision REVISION] [--variant VARIANT]
[--base_model_default_dtype {bf16,fp32}]
[--unet_attention_slice [UNET_ATTENTION_SLICE]]
Expand Down Expand Up @@ -2079,6 +2087,8 @@ options:
Path to pretrained T5 model
--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH
Path to pretrained Gemma model
--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH
Path to pretrained Qwen text encoder model
--revision REVISION Git branch/tag/commit for model version
--variant VARIANT Model variant (e.g., fp16, bf16)
--base_model_default_dtype {bf16,fp32}
Expand Down
10 changes: 10 additions & 0 deletions documentation/OPTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -253,6 +253,13 @@ Where `foo` is your config environment - or just use `config/config.json` if you
- **What**: Path to the pretrained Gemma model or its identifier from <https://huggingface.co/models>.
- **Why**: When training Gemma-based models (for example LTX-2, Sana, or Lumina2), you can point at a shared Gemma checkpoint without changing the base diffusion model path.

### `--qwen_text_encoder_model_name_or_path`

- **What**: Path to a pretrained Qwen text encoder model or its identifier from <https://huggingface.co/models>.
- **Default**: `None` (use the Qwen text encoder source defined by the selected model).
- **Why**: Use this to share or replace the Qwen text encoder used by Qwen-based model families without editing the Hugging Face cache.
- **Notes**: This applies to model families with one Qwen text encoder. If a model family defines multiple Qwen text encoders, the option is ignored and SimpleTuner logs a warning.

### `--max_grounding_entities`

- **What**: Maximum number of grounding entities per image for GLIGEN-style spatial annotations.
Expand Down Expand Up @@ -1752,6 +1759,7 @@ usage: train.py [-h] --model_family
[--pretrained_unet_subfolder PRETRAINED_UNET_SUBFOLDER]
[--pretrained_t5_model_name_or_path PRETRAINED_T5_MODEL_NAME_OR_PATH]
[--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH]
[--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH]
[--revision REVISION] [--variant VARIANT]
[--base_model_default_dtype {bf16,fp32}]
[--unet_attention_slice [UNET_ATTENTION_SLICE]]
Expand Down Expand Up @@ -2083,6 +2091,8 @@ options:
Path to pretrained T5 model
--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH
Path to pretrained Gemma model
--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH
Path to pretrained Qwen text encoder model
--revision REVISION Git branch/tag/commit for model version
--variant VARIANT Model variant (e.g., fp16, bf16)
--base_model_default_dtype {bf16,fp32}
Expand Down
10 changes: 10 additions & 0 deletions documentation/OPTIONS.pt-BR.md
Original file line number Diff line number Diff line change
Expand Up @@ -253,6 +253,13 @@ Onde `foo` e seu ambiente de config — ou use `config/config.json` se nao estiv
- **O que**: Caminho para o modelo Gemma pre-treinado ou seu identificador em <https://huggingface.co/models>.
- **Por que**: Ao treinar modelos baseados em Gemma (por exemplo LTX-2, Sana ou Lumina2), voce pode apontar para um checkpoint Gemma compartilhado sem mudar o caminho do modelo base de difusao.

### `--qwen_text_encoder_model_name_or_path`

- **O que**: Caminho para um encoder de texto Qwen pre-treinado ou seu identificador em <https://huggingface.co/models>.
- **Padrao**: `None` (usa a origem do encoder de texto Qwen definida pelo modelo selecionado).
- **Por que**: Use para compartilhar ou substituir o encoder de texto Qwen em familias de modelos baseadas em Qwen sem editar o cache do Hugging Face.
- **Notas**: Aplica-se a familias de modelos com um unico encoder de texto Qwen. Se uma familia definir varios encoders Qwen, a opcao e ignorada e o SimpleTuner registra um aviso.

### `--max_grounding_entities`
- Numero maximo de entidades de grounding por imagem para anotacoes espaciais no estilo GLIGEN. Padrao: 0 (desabilitado). Valores tipicos: 4-16.

Expand Down Expand Up @@ -1744,6 +1751,7 @@ usage: train.py [-h] --model_family
[--pretrained_unet_subfolder PRETRAINED_UNET_SUBFOLDER]
[--pretrained_t5_model_name_or_path PRETRAINED_T5_MODEL_NAME_OR_PATH]
[--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH]
[--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH]
[--revision REVISION] [--variant VARIANT]
[--base_model_default_dtype {bf16,fp32}]
[--unet_attention_slice [UNET_ATTENTION_SLICE]]
Expand Down Expand Up @@ -2074,6 +2082,8 @@ options:
Path to pretrained T5 model
--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH
Path to pretrained Gemma model
--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH
Path to pretrained Qwen text encoder model
--revision REVISION Git branch/tag/commit for model version
--variant VARIANT Model variant (e.g., fp16, bf16)
--base_model_default_dtype {bf16,fp32}
Expand Down
10 changes: 10 additions & 0 deletions documentation/OPTIONS.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -254,6 +254,13 @@ simpletuner configure config/foo/config.json
- **内容**:预训练 Gemma 模型路径或 <https://huggingface.co/models> 上的标识符。
- **原因**:训练 Gemma 系模型(例如 LTX-2、Sana、Lumina2)时,可单独指定 Gemma 权重来源,而无需更换基础扩散模型路径。

### `--qwen_text_encoder_model_name_or_path`

- **内容**:预训练 Qwen 文本编码器模型路径,或 <https://huggingface.co/models> 上的标识符。
- **默认**:`None`(使用所选模型定义的 Qwen 文本编码器来源)。
- **原因**:用于在 Qwen 系模型家族中共享或替换 Qwen 文本编码器,而无需编辑 Hugging Face 缓存。
- **说明**:此选项适用于只有一个 Qwen 文本编码器的模型家族。如果某个模型家族定义了多个 Qwen 编码器,该选项会被忽略,SimpleTuner 会记录警告。

### `--max_grounding_entities`
- 每张图像用于 GLIGEN 风格空间标注的最大 grounding 实体数。默认值:0(禁用)。典型值:4-16。

Expand Down Expand Up @@ -1751,6 +1758,7 @@ usage: train.py [-h] --model_family
[--pretrained_unet_subfolder PRETRAINED_UNET_SUBFOLDER]
[--pretrained_t5_model_name_or_path PRETRAINED_T5_MODEL_NAME_OR_PATH]
[--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH]
[--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH]
[--revision REVISION] [--variant VARIANT]
[--base_model_default_dtype {bf16,fp32}]
[--unet_attention_slice [UNET_ATTENTION_SLICE]]
Expand Down Expand Up @@ -2081,6 +2089,8 @@ options:
Path to pretrained T5 model
--pretrained_gemma_model_name_or_path PRETRAINED_GEMMA_MODEL_NAME_OR_PATH
Path to pretrained Gemma model
--qwen_text_encoder_model_name_or_path QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH
Path to pretrained Qwen text encoder model
--revision REVISION Git branch/tag/commit for model version
--variant VARIANT Model variant (e.g., fp16, bf16)
--base_model_default_dtype {bf16,fp32}
Expand Down
2 changes: 1 addition & 1 deletion documentation/quickstart/FLUX2.es.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Para seleccionar una variante, configura `model_flavour` en tu configuración:
}
```

> **Importante**: Para `klein-4b` y `klein-9b`, deja `pretrained_text_encoder_model_name_or_path` sin definir a menos que realmente quieras reemplazar el codificador Qwen3 incluido. Si configuras ese campo, anulas el valor predeterminado de Klein y puedes provocar la descarga de otro codificador de texto.
> **Importante**: Para `klein-4b` y `klein-9b`, deja `qwen_text_encoder_model_name_or_path` sin definir para usar el codificador Qwen3 incluido. Configuralo solo cuando reemplaces el codificador Qwen3 incluido por otra fuente compatible con Qwen.

## Resumen del modelo

Expand Down
2 changes: 1 addition & 1 deletion documentation/quickstart/FLUX2.hi.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ FLUX.2 तीन वेरिएंट में आता है:
}
```

> **महत्वपूर्ण**: `klein-4b` और `klein-9b` के लिए `pretrained_text_encoder_model_name_or_path` को unset छोड़ें, जब तक कि आप bundled Qwen3 text encoder को जानबूझकर बदलना न चाहते हों। इस field को सेट करने पर Klein का default override हो जाता है और किसी दूसरे text encoder का download शुरू हो सकता है
> **महत्वपूर्ण**: `klein-4b` और `klein-9b` के लिए bundled Qwen3 text encoder उपयोग करने के लिए `qwen_text_encoder_model_name_or_path` को unset छोड़ें। इसे केवल तब सेट करें जब bundled Qwen3 text encoder को किसी दूसरे Qwen-compatible source से बदलना हो

## मॉडल ओवरव्यू

Expand Down
2 changes: 1 addition & 1 deletion documentation/quickstart/FLUX2.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ FLUX.2は3つのバリアントがあります:
}
```

> **重要**: `klein-4b` と `klein-9b` では、同梱のQwen3テキストエンコーダーを意図的に置き換えたい場合を除き、`pretrained_text_encoder_model_name_or_path` は設定しないでください。この項目を設定するとKleinのデフォルトを上書きし、別のテキストエンコーダーのダウンロードが発生することがあります
> **重要**: `klein-4b` と `klein-9b` では、同梱の Qwen3 テキストエンコーダーを使用するため `qwen_text_encoder_model_name_or_path` は未設定のままにします。同梱の Qwen3 テキストエンコーダーを別の Qwen 互換ソースに置き換える場合のみ設定してください

## モデル概要

Expand Down
2 changes: 1 addition & 1 deletion documentation/quickstart/FLUX2.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ To select a variant, set `model_flavour` in your config:
}
```

> **Important**: For `klein-4b` and `klein-9b`, leave `pretrained_text_encoder_model_name_or_path` unset unless you intentionally want to replace the bundled Qwen3 text encoder. Setting that field overrides the Klein default and can trigger downloads of a different text encoder.
> **Important**: For `klein-4b` and `klein-9b`, leave `qwen_text_encoder_model_name_or_path` unset to use the bundled Qwen3 text encoder. Set it only when replacing the bundled Qwen3 text encoder with another Qwen-compatible source.

## Model Overview

Expand Down
2 changes: 1 addition & 1 deletion documentation/quickstart/FLUX2.pt-BR.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Para selecionar uma variante, defina `model_flavour` na sua configuração:
}
```

> **Importante**: Para `klein-4b` e `klein-9b`, deixe `pretrained_text_encoder_model_name_or_path` sem definir, a menos que você realmente queira substituir o encoder Qwen3 incluído. Ao definir esse campo, você sobrescreve o padrão do Klein e pode disparar o download de outro encoder de texto.
> **Importante**: Para `klein-4b` e `klein-9b`, deixe `qwen_text_encoder_model_name_or_path` sem definir para usar o encoder Qwen3 incluído. Defina-o apenas ao substituir o encoder Qwen3 incluído por outra fonte compatível com Qwen.

## Visão geral do modelo

Expand Down
2 changes: 1 addition & 1 deletion documentation/quickstart/FLUX2.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ FLUX.2 有三个变体:
}
```

> **重要**:对于 `klein-4b` 和 `klein-9b`,除非你明确想替换内置的 Qwen3 文本编码器,否则不要设置 `pretrained_text_encoder_model_name_or_path`。设置这个字段会覆盖 Klein 的默认行为,并可能触发下载其他文本编码器
> **重要**:对于 `klein-4b` 和 `klein-9b`,请将 `qwen_text_encoder_model_name_or_path` 留空以使用内置的 Qwen3 文本编码器。仅在需要用另一个兼容 Qwen 的来源替换内置 Qwen3 文本编码器时设置它

## 模型概述

Expand Down
1 change: 1 addition & 0 deletions simpletuner/helpers/configuration/env_file.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@
"MODEL_TYPE": "--model_type",
"MODEL_NAME": "--pretrained_model_name_or_path",
"MODEL_FAMILY": "--model_family",
"QWEN_TEXT_ENCODER_MODEL_NAME_OR_PATH": "--qwen_text_encoder_model_name_or_path",
"TRAIN_BATCH_SIZE": "--train_batch_size",
"USE_GRADIENT_CHECKPOINTING": "--gradient_checkpointing",
"ENABLE_CHUNKED_FEED_FORWARD": "--enable_chunked_feed_forward",
Expand Down
8 changes: 6 additions & 2 deletions simpletuner/helpers/models/ace_step/model.py
Original file line number Diff line number Diff line change
Expand Up @@ -378,9 +378,13 @@ def _resolve_v15_layout(self, base_path: Optional[str] = None) -> Optional[Dict[
if available_variants:
variant_dir = available_variants[0]

qwen_text_encoder_path = self._get_optional_config_model_path("qwen_text_encoder_model_name_or_path")
tokenizer_dir = shared_root / self.V15_SHARED_TEXT_ENCODER_SUBFOLDER
vae_dir = shared_root / self.V15_SHARED_VAE_SUBFOLDER
if variant_dir is None or not tokenizer_dir.is_dir() or not vae_dir.is_dir():
if variant_dir is None or not vae_dir.is_dir():
self._v15_layout = None
return None
if not qwen_text_encoder_path and not tokenizer_dir.is_dir():
self._v15_layout = None
return None

Expand All @@ -404,7 +408,7 @@ def _resolve_v15_layout(self, base_path: Optional[str] = None) -> Optional[Dict[
self._v15_layout = {
"root_path": str(shared_root),
"variant_path": str(variant_dir),
"tokenizer_path": str(tokenizer_dir),
"tokenizer_path": qwen_text_encoder_path or str(tokenizer_dir),
"vae_path": str(vae_dir),
"silence_latent_path": str(silence_path),
}
Expand Down
Loading
Loading