Qwen-Image
'The original 20B Qwen-Image text-to-image model: a 60-block double-stream MMDiT conditioned on Qwen2.5-VL-7B, working in the 8× / 16-channel latent space of a Wan-style VAE. Strong at text rendering, especially Chinese.'
- weights
- Qwen/Qwen-Image ↗
- org
- Qwen (Alibaba)
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2025-08-04
- native output
- 1328×1328 (1:1), 1664×928 (16:9) and other ~1.76 MP aspect ratios
- total params
- 28.85B
- model.arch
- qwen_image
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Qwen-Image MMDiT (20B) diffusers.QwenImageTransformer2DModel With the default qtype (qfloat8) and name_or_path Qwen/Qwen-Image, ai-toolkit loads the ComfyUI file Comfy-Org/Qwen-Image_ComfyUI split_files/diffusion_models/qwen_image_fp8mixed.safetensors (20,533,672,381 bytes) instead, using this repo only for the config. Unquantized or other qtypes prefer qwen_image_bf16.safetensors from the same repo. | 20.43B 20,430,401,088 | 40.86 GB | bf16 | yes |
| Text encoder | Qwen2.5-VL-7B-Instruct transformers.Qwen2_5_VLForConditionalGeneration Count includes the vision tower. ai-toolkit drops the vision tower right after loading (before quantization), since text-to-image never uses it. | 8.29B 8,292,166,656 | 16.58 GB | bf16 | no |
| Tokenizer | Qwen2 tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage A Wan 2.1-style causal video VAE. Images go through it as a single frame. Shared by every Qwen-Image 1.x model. | 126.9M 126,892,531 | 254 MB | bf16 | no |
| Training adapter | Accuracy recovery adapter, 3-bitadapter Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_torchao_uint3.safetensors" (the UI option "3 bit with ARA"). The transformer is quantized to 3-bit torchao and this LoRA-shaped adapter runs alongside it to recover accuracy. The example config says 3-bit is required for 24 GB. | 148.0M 147,961,856 | 296 MB | fp16 | no |
| total | 28.85B | 57.70 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- The VAE downsamples 3 times (dim_mult [1, 2, 4, 4]) for 8× spatial. The transformer packs 2×2 latent patches into 64-channel tokens, so one token covers 16×16 pixels. Latents are normalized with the per-channel latents_mean / latents_std from the VAE config.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 1328×1328 | 16×1×166×166 | 6,889 | native 1:1 |
| 1664×928 | 16×1×116×208 | 6,032 | native 16:9 |
Architecture
- Blocks
- 60 double-stream (MMDiT) blocks
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288 (GELU-tanh), separate image and text FFNs
- Text conditioning
- Joint attention with Qwen2.5-VL last hidden states (3584-d). System-prompt template tokens (34) are dropped; up to 1024 tokens
- Guidance
- No guidance embedding. True CFG with a negative prompt
- Objective
- Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
- Norm / position
- QK RMSNorm, 3D RoPE (16/56/56 axes)
In AI Toolkit
- model.arch
- qwen_image
- UI label
- Qwen-Image (image)
- model.name_or_path
- Qwen/Qwen-Image
- source
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- network.conv
- disabled (linear LoRA only)
- ARA option
- 3 bit with ARA (uint3)
Specifics
- Resolution
- Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
- Transformer source
- For Qwen/Qwen-Image the transformer weights come from Comfy-Org/Qwen-Image_ComfyUI: fp8mixed when qtype is qfloat8, bf16 when unquantized or for other qtypes. A local copy in the ComfyUI models folder wins over a download. Set model_kwargs.use_comfy_weights: false to load the Diffusers repo instead.
- Text encoder
- Qwen2.5-VL from text_encoder/ in name_or_path (or extras_name_or_path), vision tower removed. Never trained. With low_vram it stays on the CPU and moves to the GPU only to encode.
- Single-file checkpoints
- A .safetensors name_or_path loads with the Qwen/Qwen-Image transformer config, and the text encoder, tokenizer and VAE come from Qwen/Qwen-Image.
- Accuracy recovery adapter
- qtype "uint3|<adapter>" quantizes the adapter-covered linears to that qtype and everything else to uint8, then keeps the adapter live on the model. Cannot be combined with assistant_lora_path.
- Loss
- Flow-matching velocity target: noise − latents.
- LoRA target
- QwenImageTransformer2DModel
- Saving
- LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file (the Diffusers key layout is the ComfyUI layout for this model).
- Sampling
- Flowmatch Euler with true CFG (guidance_scale is true_cfg_scale). low_vram also turns on VAE tiling for the decode. Control images are not supported.
- Metadata base version
- qwen_image
Example config
not verified
job: extensionconfig: name: "my_qwen_image_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Qwen/Qwen-Image" arch: "qwen_image" quantize: true qtype: "qfloat8" # 24 GB: qtype: "uint3|ostris/accuracy_recovery_adapters/qwen_image_torchao_uint3.safetensors" quantize_te: true qtype_te: "qfloat8" low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 3 sample_steps: 25 prompts: - "a man holding a sign that says, 'this is a sign'"Links
Lumina-Image 2.0
A 2.6B flow-matching diffusion transformer that reads Gemma 2 2B hidden states and works in the FLUX.1 latent space. Text and image tokens pass through separate refiner layers, then share one stack of single-stream blocks.
Qwen-Image-2512
'The December 2025 update of the Qwen-Image text-to-image model. Same 20B MMDiT, Qwen2.5-VL-7B text encoder and VAE as the original, retrained for more realistic people, finer natural detail and better text rendering.'