Z-Image Turbo
'The step-distilled 6B member of Z-Image: a single-stream DiT that makes images in about 8 steps without CFG. Training it directly breaks the distillation, so ai-toolkit trains it through a de-distilling training adapter that is merged in for training and removed for sampling.'
- weights
- Tongyi-MAI/Z-Image-Turbo ↗
- org
- Tongyi-MAI (Alibaba)
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2025-11-25not verified
- native output
- 8 steps, no CFG
- total params
- 10.26B
- model.arch
- zimage:turbo
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Z-Image Turbo DiT (Diffusers) diffusers.ZImageTransformer2DModel Stored as fp32 in the Diffusers repo. ai-toolkit supplies the config from here but takes the weights from the ComfyUI file below. | 6.15B 6,154,908,736 | 24.62 GB | fp32 | yes |
| Transformer (loaded) | Z-Image Turbo, ComfyUI bf16 filealternate file diffusers.ZImageTransformer2DModel What ai-toolkit downloads when name_or_path is Tongyi-MAI/Z-Image-Turbo and qtype is qfloat8 or quantization is off. With a convrot8 qtype it prefers z_image_turbo_int8_convrot.safetensors from the same folder. The fused ComfyUI qkv is split into to_q/to_k/to_v on load. | 6.15B 6,154,908,736 | 12.31 GB | bf16 | yes |
| Text encoder | Qwen3-4B transformers.Qwen3ForCausalLM 36 layers, 2560 hidden, tied embeddings (151,936 vocab). Only the second-to-last hidden state is used. | 4.02B 4,022,468,096 | 8.04 GB | bf16 | no |
| Tokenizer | Qwen3 BPE tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL Same config as the FLUX.1 VAE (scaling 0.3611, shift 0.1159). | 83.8M 83,819,683 | 168 MB | bf16 | no |
| Training adapter | Z-Image Turbo training adapter v2 (LoRA, rank 64)adapter A de-distillation LoRA trained on Turbo outputs. ai-toolkit merges it into the transformer before training and subtracts it while sampling, so your LoRA learns only the new concept and still runs at 8 steps without it. v1 (half the size) is in the same repo. | 170.1M 170,065,920 | 340 MB | bf16 | no |
| total | 10.26B | 32.83 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- The FLUX.1 latent space. The transformer patchifies 2×2 internally (all_patch_size [2]), so one token covers 16×16 pixels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 16×1×64×64 | 1,024 | |
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 2048×2048 | 16×1×256×256 | 16,384 |
Architecture
- Blocks
- 30 single-stream blocks, plus 2 noise-refiner blocks (image only) and 2 context-refiner blocks (text only)
- Hidden size
- 3840 (30 heads × 128)
- FFN size
- 10240 (SwiGLU)
- Text conditioning
- Qwen3-4B second-to-last hidden states (2560-d) through the chat template, up to 512 tokens, padding dropped. Text and image tokens share one sequence
- Timestep
- adaLN modulation from a 256-d timestep embedding (context refiner is unmodulated)
- Distillation
- Step-distilled for ~8 steps and RL-tuned; runs without CFG
- Objective
- Rectified flow, shift 3.0
- Norm / position
- RMSNorm, QK RMSNorm, 3-axis RoPE [32, 48, 48], theta 256
In AI Toolkit
- model.arch
- zimage:turbo
- UI label
- Z-Image Turbo (w/ Training Adapter) (image)
- model.name_or_path
- Tongyi-MAI/Z-Image-Turbo
- source
- extensions_built_in/diffusion_models/z_image/z_image.py
- toolkit/models/v2/diffusion_models/z_image.py
- toolkit/models/v2/text_encoders/qwen3.py
- toolkit/models/v2/vae/autoencoder_kl.py
- toolkit/models/v2/resolver.py
- toolkit/config_modules.py
- toolkit/models/registry.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading, model.assistant_lora_path
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- assistant_lora_path
- ostris/zimage_turbo_training_adapter/zimage_turbo_training_adapter_v2.safetensors
- train.unload_text_encoder
- false
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- sample guidance_scale / steps
- 1 / 9
- network.conv
- disabled (linear LoRA only)
Specifics
- Arch name
- zimage:turbo is a UI preset. The config loader strips everything after the colon, so the job runs as arch zimage with these defaults.
- Training adapter
- The adapter is merged into the full-precision transformer at weight 1.0 before quantization, then inverted (multiplier −1) during sampling so samples show the distilled model plus your LoRA. qfloat8 is switched to float8 when an adapter is used.
- Transformer source
- For the Tongyi-MAI/Z-Image-Turbo repo id, weights come from Comfy-Org/z_image_turbo (bf16, or int8_convrot for a convrot8 qtype); a local copy under MODELS_PATH wins over a download. Set model_kwargs.use_comfy_weights: false to load the Diffusers folder.
- Text encoder and VAE source
- text_encoder/, tokenizer/ and vae/ from extras_name_or_path (defaults to name_or_path). Single .safetensors checkpoints fall back to Tongyi-MAI/Z-Image-Turbo.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Timesteps
- The model takes t = (1000 − timestep) / 1000 (1 = clean) and its output is negated to match the noise − latents flow target.
- Quantization
- t_embedder, cap_embedder, all_x_embedder and all_final_layer stay in full precision.
- Sampling
- guidance_scale is shifted down by 1 before it reaches the pipeline, so 1 means no CFG. Flowmatch Euler, shift 3.0.
- LoRA target
- ZImageTransformer2DModel
- Saving
- Full fine-tunes save one ComfyUI-layout file (fused qkv). LoRA keys use the ComfyUI diffusion_model prefix.
- Metadata base version
- zimage
Example config
not verified
job: extensionconfig: name: "my_zimage_turbo_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "Tongyi-MAI/Z-Image-Turbo" arch: "zimage:turbo" assistant_lora_path: "ostris/zimage_turbo_training_adapter/zimage_turbo_training_adapter_v2.safetensors" quantize: true qtype: "qfloat8" quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 1 sample_steps: 9 prompts: - "a bear building a log cabin in the snow covered mountains"Links
FLUX.2 [dev]
The 32B, guidance-distilled FLUX.2 model. A new double/single-stream transformer conditioned on Mistral Small 3.1 (24B) hidden states, working in the 32-channel latent space of the new FLUX.2 VAE. Reference images go in as extra tokens, so one model generates and edits.
Z-Image
'The undistilled 6B foundation model of the Z-Image family: a single-stream DiT on Qwen3-4B text features and the FLUX.1 VAE, with full CFG and negative prompts. It is the Z-Image checkpoint meant for fine-tuning and trains directly, with no adapter.'