Z-Image De-Turbo
'Z-Image Turbo with the step distillation trained back out, made by fine-tuning Turbo on its own outputs. It runs with normal CFG at 20–30 steps and can be trained directly, with no adapter, for LoRAs or long fine-tunes that stay compatible with the Turbo weights.'
- weights
- ostris/Z-Image-De-Turbo ↗
- org
- Ostris
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2025-12-04not verified
- native output
- 20–30 steps, CFG 2–3
- total params
- 10.26B
- model.arch
- zimage:deturbo
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Z-Image De-Turbo DiT diffusers.ZImageTransformer2DModel Same shape as Z-Image Turbo. ai-toolkit loads this Diffusers folder. The repo also has a ComfyUI single file, z_image_de_turbo_v1_bf16.safetensors (same parameter count). | 6.15B 6,154,908,736 | 12.31 GB | bf16 | yes |
| Text encoder | Qwen3-4B transformers.Qwen3ForCausalLM Not in the De-Turbo repo. ai-toolkit loads it from extras_name_or_path, Tongyi-MAI/Z-Image-Turbo. 36 layers, 2560 hidden; only the second-to-last hidden state is used. | 4.02B 4,022,468,096 | 8.04 GB | bf16 | no |
| Tokenizer | Qwen3 BPE tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL Loaded from Tongyi-MAI/Z-Image-Turbo. Same config as the FLUX.1 VAE (scaling 0.3611, shift 0.1159). | 83.8M 83,819,683 | 168 MB | bf16 | no |
| total | 10.26B | 20.52 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- The FLUX.1 latent space. The transformer patchifies 2×2 internally (all_patch_size [2]), so one token covers 16×16 pixels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 16×1×64×64 | 1,024 | |
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 2048×2048 | 16×1×256×256 | 16,384 |
Architecture
- Blocks
- 30 single-stream blocks, plus 2 noise-refiner blocks (image only) and 2 context-refiner blocks (text only)
- Hidden size
- 3840 (30 heads × 128)
- FFN size
- 10240 (SwiGLU)
- Text conditioning
- Qwen3-4B second-to-last hidden states (2560-d) through the chat template, up to 512 tokens, padding dropped. Text and image tokens share one sequence
- Timestep
- adaLN modulation from a 256-d timestep embedding (context refiner is unmodulated)
- Distillation
- Removed by fine-tuning Turbo on Turbo-generated images; runs with CFG
- Objective
- Rectified flow (ai-toolkit shift 3.0; the repo ships no scheduler config)
- Norm / position
- RMSNorm, QK RMSNorm, 3-axis RoPE [32, 48, 48], theta 256
In AI Toolkit
- model.arch
- zimage:deturbo
- UI label
- Z-Image De-Turbo (De-Distilled) (image)
- model.name_or_path
- ostris/Z-Image-De-Turbo
- source
- extensions_built_in/diffusion_models/z_image/z_image.py
- toolkit/models/v2/diffusion_models/z_image.py
- toolkit/models/v2/text_encoders/qwen3.py
- toolkit/models/v2/vae/autoencoder_kl.py
- toolkit/models/v2/resolver.py
- toolkit/config_modules.py
- toolkit/models/registry.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- extras_name_or_path
- Tongyi-MAI/Z-Image-Turbo
- train.unload_text_encoder
- false
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- sample guidance_scale / steps
- 3 / 25
- network.conv
- disabled (linear LoRA only)
Specifics
- Arch name
- zimage:deturbo is a UI preset. The config loader strips everything after the colon, so the job runs as arch zimage with these defaults.
- Training adapter
- None needed. The distillation is already gone, so LoRAs train directly on this transformer.
- Transformer source
- transformer/ from ostris/Z-Image-De-Turbo. No ComfyUI file is registered for this repo, so no swap happens.
- Text encoder and VAE source
- text_encoder/, tokenizer/ and vae/ from extras_name_or_path, which the UI sets to Tongyi-MAI/Z-Image-Turbo because the De-Turbo repo ships only the transformer.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Timesteps
- The model takes t = (1000 − timestep) / 1000 (1 = clean) and its output is negated to match the noise − latents flow target.
- Quantization
- t_embedder, cap_embedder, all_x_embedder and all_final_layer stay in full precision.
- Sampling
- guidance_scale is shifted down by 1 before it reaches the pipeline (the UI’s 3 becomes 2). Flowmatch Euler, shift 3.0.
- LoRA target
- ZImageTransformer2DModel
- Saving
- Full fine-tunes save one ComfyUI-layout file (fused qkv). LoRA keys use the ComfyUI diffusion_model prefix. LoRAs are meant to stay usable on Turbo.
- Metadata base version
- zimage
Example config
not verified
job: extensionconfig: name: "my_zimage_deturbo_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "ostris/Z-Image-De-Turbo" extras_name_or_path: "Tongyi-MAI/Z-Image-Turbo" arch: "zimage:deturbo" quantize: true qtype: "qfloat8" quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 3 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Z-Image
'The undistilled 6B foundation model of the Z-Image family: a single-stream DiT on Qwen3-4B text features and the FLUX.1 VAE, with full CFG and negative prompts. It is the Z-Image checkpoint meant for fine-tuning and trains directly, with no adapter.'
FLUX.2 [klein] 4B Base
The smallest FLUX.2 model: a 4B transformer conditioned on Qwen3-4B, using the same 32-channel FLUX.2 VAE and reference-image editing as FLUX.2 [dev]. The Base release is undistilled (no step or guidance distillation), Apache 2.0, and meant for fine-tuning.