Qwen-Image-2512
'The December 2025 update of the Qwen-Image text-to-image model. Same 20B MMDiT, Qwen2.5-VL-7B text encoder and VAE as the original, retrained for more realistic people, finer natural detail and better text rendering.'
- weights
- Qwen/Qwen-Image-2512 ↗
- org
- Qwen (Alibaba)
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2025-12-31not verified
- native output
- 1328×1328 (1:1), 1664×928 (16:9) and other ~1.76 MP aspect ratios
- total params
- 28.85B
- model.arch
- qwen_image:2512
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Qwen-Image-2512 MMDiT (20B) diffusers.QwenImageTransformer2DModel Same shape and parameter count as the original Qwen-Image transformer. ai-toolkit loads it from this repo: no ComfyUI file is registered for Qwen-Image-2512, although Comfy-Org/Qwen-Image_ComfyUI ships qwen_image_2512_bf16 and fp8_e4m3fn files for ComfyUI. | 20.43B 20,430,401,088 | 40.86 GB | bf16 | yes |
| Text encoder | Qwen2.5-VL-7B-Instruct transformers.Qwen2_5_VLForConditionalGeneration Count includes the vision tower. ai-toolkit drops the vision tower right after loading (before quantization), since text-to-image never uses it. | 8.29B 8,292,166,656 | 16.58 GB | bf16 | no |
| Tokenizer | Qwen2 tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage A Wan 2.1-style causal video VAE. Images go through it as a single frame. Shared by every Qwen-Image 1.x model. | 126.9M 126,892,531 | 254 MB | bf16 | no |
| Training adapter | Accuracy recovery adapter, 3-bitadapter Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_2512_torchao_uint3.safetensors" (the UI option "3 bit with ARA"). The transformer is quantized to 3-bit torchao and this LoRA-shaped adapter runs alongside it to recover accuracy. Trained for 2512; the original Qwen-Image adapter is a different file. | 147.5M 147,456,000 | 295 MB | bf16 | no |
| Training adapter | Accuracy recovery adapter, 4-bitadapternot verified The UI option "4 bit with ARA". | 147.5M 147,456,000 | 295 MB | bf16 | no |
| total | 28.85B | 57.70 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- The VAE downsamples 3 times (dim_mult [1, 2, 4, 4]) for 8× spatial. The transformer packs 2×2 latent patches into 64-channel tokens, so one token covers 16×16 pixels. Latents are normalized with the per-channel latents_mean / latents_std from the VAE config.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 1328×1328 | 16×1×166×166 | 6,889 | native 1:1 |
| 1664×928 | 16×1×116×208 | 6,032 | native 16:9 |
Architecture
- Blocks
- 60 double-stream (MMDiT) blocks
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288 (GELU-tanh), separate image and text FFNs
- Text conditioning
- Joint attention with Qwen2.5-VL last hidden states (3584-d). System-prompt template tokens (34) are dropped; up to 1024 tokens
- Guidance
- No guidance embedding. True CFG with a negative prompt
- Objective
- Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
- Norm / position
- QK RMSNorm, 3D RoPE (16/56/56 axes)
In AI Toolkit
- model.arch
- qwen_image:2512
- UI label
- Qwen-Image-2512 (image)
- model.name_or_path
- Qwen/Qwen-Image-2512
- source
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- network.conv
- disabled (linear LoRA only)
- ARA options
- 3 bit with ARA (uint3), 4 bit with ARA (uint4)
Specifics
- Arch variant
- qwen_image:2512 is a UI tag. ModelConfig strips everything after the colon, so it trains with the same QwenImageModel code as qwen_image; only the default name_or_path and adapter list differ.
- Resolution
- Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
- Transformer source
- transformer/ in Qwen/Qwen-Image-2512. The ComfyUI-file substitution ai-toolkit does for Qwen/Qwen-Image is only registered for that repo id, not for 2512.
- Text encoder
- Qwen2.5-VL from text_encoder/ in name_or_path (or extras_name_or_path), vision tower removed. Never trained. With low_vram it stays on the CPU and moves to the GPU only to encode.
- Single-file checkpoints
- A .safetensors name_or_path loads with the Qwen/Qwen-Image transformer config, and the text encoder, tokenizer and VAE come from Qwen/Qwen-Image (identical in size to the 2512 ones).
- Accuracy recovery adapter
- qtype "uint3|<adapter>" quantizes the adapter-covered linears to that qtype and everything else to uint8, then keeps the adapter live on the model. Cannot be combined with assistant_lora_path.
- Loss
- Flow-matching velocity target: noise − latents.
- LoRA target
- QwenImageTransformer2DModel
- Saving
- LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file (the Diffusers key layout is the ComfyUI layout for this model).
- Sampling
- Flowmatch Euler with true CFG (guidance_scale is true_cfg_scale). low_vram also turns on VAE tiling for the decode. Control images are not supported.
- Metadata base version
- qwen_image
Example config
not verified
job: extensionconfig: name: "my_qwen_image_2512_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Qwen/Qwen-Image-2512" arch: "qwen_image:2512" quantize: true qtype: "qfloat8" # 3-bit: qtype: "uint3|ostris/accuracy_recovery_adapters/qwen_image_2512_torchao_uint3.safetensors" quantize_te: true qtype_te: "qfloat8" low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 3 sample_steps: 25 prompts: - "a man holding a sign that says, 'this is a sign'"Links
Qwen-Image
'The original 20B Qwen-Image text-to-image model: a 60-block double-stream MMDiT conditioned on Qwen2.5-VL-7B, working in the 8× / 16-channel latent space of a Wan-style VAE. Strong at text rendering, especially Chinese.'
Qwen-Image-2.1
'One 7B single-stream DiT that does both text-to-image and editing with up to 10 reference images, conditioned on Qwen3-VL-8B. Its new 16×, 64-channel VAE is natively RGBA, so it can generate and edit transparent images.'