Docs
AI ToolkitModels

Z-Image Turbo

'The step-distilled 6B member of Z-Image: a single-stream DiT that makes images in about 8 steps without CFG. Training it directly breaks the distillation, so ai-toolkit trains it through a de-distilling training adapter that is merged in for training and removed for sampling.'

org
Tongyi-MAI (Alibaba)
modality
image
tasks
text-to-image
license
Apache 2.0
released
2025-11-25not verified
native output
8 steps, no CFG
total params
10.26B
model.arch
zimage:turbo

Components

rolemodelparamssizedtypetrained
Transformer
Z-Image Turbo DiT (Diffusers)
diffusers.ZImageTransformer2DModel

Stored as fp32 in the Diffusers repo. ai-toolkit supplies the config from here but takes the weights from the ComfyUI file below.

6.15B
6,154,908,736
24.62 GBfp32yes
Transformer (loaded)
Z-Image Turbo, ComfyUI bf16 filealternate file
diffusers.ZImageTransformer2DModel

What ai-toolkit downloads when name_or_path is Tongyi-MAI/Z-Image-Turbo and qtype is qfloat8 or quantization is off. With a convrot8 qtype it prefers z_image_turbo_int8_convrot.safetensors from the same folder. The fused ComfyUI qkv is split into to_q/to_k/to_v on load.

6.15B
6,154,908,736
12.31 GBbf16yes
Text encoder
Qwen3-4B
transformers.Qwen3ForCausalLM

36 layers, 2560 hidden, tied embeddings (151,936 vocab). Only the second-to-last hidden state is used.

4.02B
4,022,468,096
8.04 GBbf16no
Tokenizer
Qwen3 BPE tokenizer
transformers.Qwen2Tokenizer
————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

Same config as the FLUX.1 VAE (scaling 0.3611, shift 0.1159).

83.8M
83,819,683
168 MBbf16no
Training adapter
Z-Image Turbo training adapter v2 (LoRA, rank 64)adapter

A de-distillation LoRA trained on Turbo outputs. ai-toolkit merges it into the transformer before training and subtracts it while sampling, so your LoRA learns only the new concept and still runs at 8 steps without it. v1 (half the size) is in the same repo.

170.1M
170,065,920
340 MBbf16no
total10.26B32.83 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
The FLUX.1 latent space. The transformer patchifies 2×2 internally (all_patch_size [2]), so one token covers 16×16 pixels.
inputlatent (c×t×h×w)tokens
512×51216×1×64×641,024
1024×102416×1×128×1284,096ai-toolkit sample default
2048×204816×1×256×25616,384

Architecture

Blocks
30 single-stream blocks, plus 2 noise-refiner blocks (image only) and 2 context-refiner blocks (text only)
Hidden size
3840 (30 heads × 128)
FFN size
10240 (SwiGLU)
Text conditioning
Qwen3-4B second-to-last hidden states (2560-d) through the chat template, up to 512 tokens, padding dropped. Text and image tokens share one sequence
Timestep
adaLN modulation from a 256-d timestep embedding (context refiner is unmodulated)
Distillation
Step-distilled for ~8 steps and RL-tuned; runs without CFG
Objective
Rectified flow, shift 3.0
Norm / position
RMSNorm, QK RMSNorm, 3-axis RoPE [32, 48, 48], theta 256

In AI Toolkit

model.arch
zimage:turbo
UI label
Z-Image Turbo (w/ Training Adapter) (image)
model.name_or_path
Tongyi-MAI/Z-Image-Turbo
extra UI sections
model.low_vram, model.layer_offloading, model.assistant_lora_path

UI defaults

quantize / quantize_te
true / true
qtype
qfloat8
low_vram
true
assistant_lora_path
ostris/zimage_turbo_training_adapter/zimage_turbo_training_adapter_v2.safetensors
train.unload_text_encoder
false
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
sample guidance_scale / steps
1 / 9
network.conv
disabled (linear LoRA only)

Specifics

Arch name
zimage:turbo is a UI preset. The config loader strips everything after the colon, so the job runs as arch zimage with these defaults.
Training adapter
The adapter is merged into the full-precision transformer at weight 1.0 before quantization, then inverted (multiplier −1) during sampling so samples show the distilled model plus your LoRA. qfloat8 is switched to float8 when an adapter is used.
Transformer source
For the Tongyi-MAI/Z-Image-Turbo repo id, weights come from Comfy-Org/z_image_turbo (bf16, or int8_convrot for a convrot8 qtype); a local copy under MODELS_PATH wins over a download. Set model_kwargs.use_comfy_weights: false to load the Diffusers folder.
Text encoder and VAE source
text_encoder/, tokenizer/ and vae/ from extras_name_or_path (defaults to name_or_path). Single .safetensors checkpoints fall back to Tongyi-MAI/Z-Image-Turbo.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Timesteps
The model takes t = (1000 − timestep) / 1000 (1 = clean) and its output is negated to match the noise − latents flow target.
Quantization
t_embedder, cap_embedder, all_x_embedder and all_final_layer stay in full precision.
Sampling
guidance_scale is shifted down by 1 before it reaches the pipeline, so 1 means no CFG. Flowmatch Euler, shift 3.0.
LoRA target
ZImageTransformer2DModel
Saving
Full fine-tunes save one ComfyUI-layout file (fused qkv). LoRA keys use the ComfyUI diffusion_model prefix.
Metadata base version
zimage

Example config

not verified

job: extensionconfig:  name: "my_zimage_turbo_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "Tongyi-MAI/Z-Image-Turbo"        arch: "zimage:turbo"        assistant_lora_path: "ostris/zimage_turbo_training_adapter/zimage_turbo_training_adapter_v2.safetensors"        quantize: true        qtype: "qfloat8"        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 1        sample_steps: 9        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page