Docs
AI ToolkitModels

ERNIE-Image

Baidu’s 8B single-stream DiT for text-to-image. It reads a Ministral 3 text encoder’s second-to-last hidden states and works in the FLUX.2 VAE latent space, 2×2 patchified to 128 channels. The repo also ships a prompt-enhancer LLM that ai-toolkit does not load.

org
Baidu
modality
image
tasks
text-to-image
license
Apache 2.0
total params
11.97B
model.arch
ernie_image

Components

rolemodelparamssizedtypetrained
Transformer
ERNIE-Image DiT (8B)
diffusers.ErnieImageTransformer2DModel

ai-toolkit uses its own patched copy of the class so batch sizes above 1 work.

8.03B
8,033,490,048
16.07 GBbf16yes
Text encoder
Mistral3 (Ministral 3 text model + Pixtral vision tower)
transformers.Mistral3Model

Text model: 26 layers, 3072-d, 131k vocab. The folder also holds a 24-layer Pixtral vision tower that text-only prompting never runs.

3.85B
3,849,090,048
7.70 GBbf16no
Tokenizer
Mistral tokenizer (131k vocab)
————
Prompt enhancer
Ministral 3 causal LMnot loaded
transformers.Ministral3ForCausalLM

Rewrites prompts before encoding in the reference pipeline. ai-toolkit does not load it, nor pe_tokenizer/.

3.83B
3,831,659,520
7.66 GBbf16no
VAE
FLUX.2 VAE
diffusers.AutoencoderKLFlux2

The config names FLUX.2-dev’s VAE as its source. The int64 tensor is the batch-norm step counter.

84.0M
84,046,372
168 MBbf16+int64no
total11.97B23.93 GB

Latent space

spatial
8×
channels
32
patch
2×2
autoencoder
FLUX.2 VAE
pixels per token
16×16
notes
The VAE downsamples 8× to 32 channels. The pipeline then packs 2×2 patches into channels (128 at 16×) and normalizes them with the VAE’s batch-norm running mean and variance. The transformer config therefore reads in_channels 128, patch_size 1.
inputlatent (c×t×h×w)tokens
1024×102432×1×128×1284,096ai-toolkit sample default
768×76832×1×96×962,304

Architecture

Blocks
36
Hidden size
4096 (32 heads × 128)
FFN size
12288
Text conditioning
Single stream. Text states (3072-d) are projected to 4096 and appended to the image tokens
Modulation
One AdaLN modulation shared by all blocks
Objective
Rectified flow. The Hub scheduler config uses shift 4.0
Norm / position
QK RMSNorm, 3-axis RoPE (32/48/48, theta 256)

In AI Toolkit

model.arch
ernie_image
UI label
ERNIE-Image (image)
model.name_or_path
baidu/ERNIE-Image
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true
qtype
qfloat8
low_vram
true
unload_text_encoder
false
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
network.conv
disabled (linear LoRA only)

Specifics

Resolution
Buckets and sample sizes snap to multiples of 32. One token covers 16×16 pixels, so 16 would be enough.
Prompt encoding
Each prompt is encoded on its own, unpadded, and the second-to-last hidden state is kept. Prompts are padded to the batch maximum only at the transformer call.
Component sources
The transformer comes from name_or_path; the text encoder, tokenizer and VAE from extras_name_or_path (defaults to name_or_path). A local folder with a text_encoder/ subfolder is used for everything.
Scheduler
Training and sampling use flowmatch with shift 3.0, not the 4.0 in the Hub scheduler config.
LoRA target
ErnieImageTransformer2DModel
Saving
LoRAs use the ComfyUI key prefix. Full fine-tunes save transformer/ in Diffusers format.
Metadata base version
ernie_image

Example config

not verified

job: extensionconfig:  name: "my_ernie_image_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "baidu/ERNIE-Image"        arch: "ernie_image"        quantize: true        qtype: "qfloat8"        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 30        prompts:          - "woman with red hair, playing chess at the park, bomb going off in the background"

On this page