ERNIE-Image
Baidu’s 8B single-stream DiT for text-to-image. It reads a Ministral 3 text encoder’s second-to-last hidden states and works in the FLUX.2 VAE latent space, 2×2 patchified to 128 channels. The repo also ships a prompt-enhancer LLM that ai-toolkit does not load.
- weights
- baidu/ERNIE-Image ↗
- org
- Baidu
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- total params
- 11.97B
- model.arch
- ernie_image
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | ERNIE-Image DiT (8B) diffusers.ErnieImageTransformer2DModel ai-toolkit uses its own patched copy of the class so batch sizes above 1 work. | 8.03B 8,033,490,048 | 16.07 GB | bf16 | yes |
| Text encoder | Mistral3 (Ministral 3 text model + Pixtral vision tower) transformers.Mistral3Model Text model: 26 layers, 3072-d, 131k vocab. The folder also holds a 24-layer Pixtral vision tower that text-only prompting never runs. | 3.85B 3,849,090,048 | 7.70 GB | bf16 | no |
| Tokenizer | Mistral tokenizer (131k vocab) | — | — | — | — |
| Prompt enhancer | Ministral 3 causal LMnot loaded transformers.Ministral3ForCausalLM Rewrites prompts before encoding in the reference pipeline. ai-toolkit does not load it, nor pe_tokenizer/. | 3.83B 3,831,659,520 | 7.66 GB | bf16 | no |
| VAE | FLUX.2 VAE diffusers.AutoencoderKLFlux2 The config names FLUX.2-dev’s VAE as its source. The int64 tensor is the batch-norm step counter. | 84.0M 84,046,372 | 168 MB | bf16+int64 | no |
| total | 11.97B | 23.93 GB | |||
Latent space
- autoencoder
- FLUX.2 VAE
- pixels per token
- 16×16
- notes
- The VAE downsamples 8× to 32 channels. The pipeline then packs 2×2 patches into channels (128 at 16×) and normalizes them with the VAE’s batch-norm running mean and variance. The transformer config therefore reads in_channels 128, patch_size 1.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 32×1×128×128 | 4,096 | ai-toolkit sample default |
| 768×768 | 32×1×96×96 | 2,304 |
Architecture
- Blocks
- 36
- Hidden size
- 4096 (32 heads × 128)
- FFN size
- 12288
- Text conditioning
- Single stream. Text states (3072-d) are projected to 4096 and appended to the image tokens
- Modulation
- One AdaLN modulation shared by all blocks
- Objective
- Rectified flow. The Hub scheduler config uses shift 4.0
- Norm / position
- QK RMSNorm, 3-axis RoPE (32/48/48, theta 256)
In AI Toolkit
- model.arch
- ernie_image
- UI label
- ERNIE-Image (image)
- model.name_or_path
- baidu/ERNIE-Image
- source
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- unload_text_encoder
- false
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- network.conv
- disabled (linear LoRA only)
Specifics
- Resolution
- Buckets and sample sizes snap to multiples of 32. One token covers 16×16 pixels, so 16 would be enough.
- Prompt encoding
- Each prompt is encoded on its own, unpadded, and the second-to-last hidden state is kept. Prompts are padded to the batch maximum only at the transformer call.
- Component sources
- The transformer comes from name_or_path; the text encoder, tokenizer and VAE from extras_name_or_path (defaults to name_or_path). A local folder with a text_encoder/ subfolder is used for everything.
- Scheduler
- Training and sampling use flowmatch with shift 3.0, not the 4.0 in the Hub scheduler config.
- LoRA target
- ErnieImageTransformer2DModel
- Saving
- LoRAs use the ComfyUI key prefix. Full fine-tunes save transformer/ in Diffusers format.
- Metadata base version
- ernie_image
Example config
not verified
job: extensionconfig: name: "my_ernie_image_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "baidu/ERNIE-Image" arch: "ernie_image" quantize: true qtype: "qfloat8" quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 30 prompts: - "woman with red hair, playing chess at the park, bomb going off in the background"Links
FLUX.2 [klein] 4B Base
The smallest FLUX.2 model: a 4B transformer conditioned on Qwen3-4B, using the same 32-channel FLUX.2 VAE and reference-image editing as FLUX.2 [dev]. The Base release is undistilled (no step or guidance distillation), Apache 2.0, and meant for fine-tuning.
FLUX.2 [klein] 9B Base
The larger FLUX.2 [klein]: a 9B transformer conditioned on Qwen3-8B, using the same 32-channel FLUX.2 VAE and reference-image editing as FLUX.2 [dev]. The Base release is undistilled (no step or guidance distillation) and meant for fine-tuning. Unlike the 4B, it is under the non-commercial FLUX license.