Docs
AI ToolkitModels

Z-Image

'The undistilled 6B foundation model of the Z-Image family: a single-stream DiT on Qwen3-4B text features and the FLUX.1 VAE, with full CFG and negative prompts. It is the Z-Image checkpoint meant for fine-tuning and trains directly, with no adapter.'

org
Tongyi-MAI (Alibaba)
modality
image
tasks
text-to-image
license
Apache 2.0
released
2026-01-23not verified
native output
512×512 to 2048×2048 (any aspect), 28–50 steps, CFG 3–5
total params
10.26B
model.arch
zimage

Components

rolemodelparamssizedtypetrained
Transformer
Z-Image DiT
diffusers.ZImageTransformer2DModel

Same shape as Z-Image Turbo. ai-toolkit loads this Diffusers folder directly (no ComfyUI file is registered for this repo).

6.15B
6,154,908,736
12.31 GBbf16yes
Text encoder
Qwen3-4B
transformers.Qwen3ForCausalLM

36 layers, 2560 hidden, tied embeddings (151,936 vocab). Only the second-to-last hidden state is used.

4.02B
4,022,468,096
8.04 GBbf16no
Tokenizer
Qwen3 BPE tokenizer
transformers.Qwen2Tokenizer
————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

Same config as the FLUX.1 VAE (scaling 0.3611, shift 0.1159).

83.8M
83,819,683
168 MBbf16no
total10.26B20.52 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
The FLUX.1 latent space. The transformer patchifies 2×2 internally (all_patch_size [2]), so one token covers 16×16 pixels.
inputlatent (c×t×h×w)tokens
512×51216×1×64×641,024
1024×102416×1×128×1284,096ai-toolkit sample default
2048×204816×1×256×25616,384

Architecture

Blocks
30 single-stream blocks, plus 2 noise-refiner blocks (image only) and 2 context-refiner blocks (text only)
Hidden size
3840 (30 heads × 128)
FFN size
10240 (SwiGLU)
Text conditioning
Qwen3-4B second-to-last hidden states (2560-d) through the chat template, up to 512 tokens, padding dropped. Text and image tokens share one sequence
Timestep
adaLN modulation from a 256-d timestep embedding (context refiner is unmodulated)
Distillation
None. Full CFG and negative prompts
Objective
Rectified flow, shift 6.0 in the repo scheduler config
Norm / position
RMSNorm, QK RMSNorm, 3-axis RoPE [32, 48, 48], theta 256

In AI Toolkit

model.arch
zimage
UI label
Z-Image (image)
model.name_or_path
Tongyi-MAI/Z-Image
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true
qtype
qfloat8
low_vram
true
train.unload_text_encoder
false
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
sample_steps
30
network.conv
disabled (linear LoRA only)

Specifics

Transformer source
transformer/ from name_or_path. A single .safetensors file also works (config from Tongyi-MAI/Z-Image-Turbo, which has the same shape).
Training adapter
Not needed: the model is undistilled. assistant_lora_path is still accepted and handled like Turbo’s.
Text encoder and VAE source
text_encoder/, tokenizer/ and vae/ from extras_name_or_path (defaults to name_or_path). Single .safetensors checkpoints fall back to Tongyi-MAI/Z-Image-Turbo.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Timesteps
The model takes t = (1000 − timestep) / 1000 (1 = clean) and its output is negated to match the noise − latents flow target.
Quantization
t_embedder, cap_embedder, all_x_embedder and all_final_layer stay in full precision.
Sampling
guidance_scale is shifted down by 1 before it reaches the pipeline (the UI’s 4 becomes 3). Flowmatch Euler with ai-toolkit’s fixed shift 3.0, not the repo’s 6.0; training uses the same 3.0 scheduler.
LoRA target
ZImageTransformer2DModel
Saving
Full fine-tunes save one ComfyUI-layout file (fused qkv). LoRA keys use the ComfyUI diffusion_model prefix.
Metadata base version
zimage

Example config

not verified

job: extensionconfig:  name: "my_zimage_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "Tongyi-MAI/Z-Image"        arch: "zimage"        quantize: true        qtype: "qfloat8"        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 30        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page