Docs
AI ToolkitModels

Qwen-Image

'The original 20B Qwen-Image text-to-image model: a 60-block double-stream MMDiT conditioned on Qwen2.5-VL-7B, working in the 8× / 16-channel latent space of a Wan-style VAE. Strong at text rendering, especially Chinese.'

org
Qwen (Alibaba)
modality
image
tasks
text-to-image
license
Apache 2.0
released
2025-08-04
native output
1328×1328 (1:1), 1664×928 (16:9) and other ~1.76 MP aspect ratios
total params
28.85B
model.arch
qwen_image

Components

rolemodelparamssizedtypetrained
Transformer
Qwen-Image MMDiT (20B)
diffusers.QwenImageTransformer2DModel

With the default qtype (qfloat8) and name_or_path Qwen/Qwen-Image, ai-toolkit loads the ComfyUI file Comfy-Org/Qwen-Image_ComfyUI split_files/diffusion_models/qwen_image_fp8mixed.safetensors (20,533,672,381 bytes) instead, using this repo only for the config. Unquantized or other qtypes prefer qwen_image_bf16.safetensors from the same repo.

20.43B
20,430,401,088
40.86 GBbf16yes
Text encoder
Qwen2.5-VL-7B-Instruct
transformers.Qwen2_5_VLForConditionalGeneration

Count includes the vision tower. ai-toolkit drops the vision tower right after loading (before quantization), since text-to-image never uses it.

8.29B
8,292,166,656
16.58 GBbf16no
Tokenizer
Qwen2 tokenizer
transformers.Qwen2Tokenizer
————
VAE
Qwen-Image VAE
diffusers.AutoencoderKLQwenImage

A Wan 2.1-style causal video VAE. Images go through it as a single frame. Shared by every Qwen-Image 1.x model.

126.9M
126,892,531
254 MBbf16no
Training adapter
Accuracy recovery adapter, 3-bitadapter

Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_torchao_uint3.safetensors" (the UI option "3 bit with ARA"). The transformer is quantized to 3-bit torchao and this LoRA-shaped adapter runs alongside it to recover accuracy. The example config says 3-bit is required for 24 GB.

148.0M
147,961,856
296 MBfp16no
total28.85B57.70 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Qwen-Image VAE
pixels per token
16×16
notes
The VAE downsamples 3 times (dim_mult [1, 2, 4, 4]) for 8× spatial. The transformer packs 2×2 latent patches into 64-channel tokens, so one token covers 16×16 pixels. Latents are normalized with the per-channel latents_mean / latents_std from the VAE config.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
1328×132816×1×166×1666,889native 1:1
1664×92816×1×116×2086,032native 16:9

Architecture

Blocks
60 double-stream (MMDiT) blocks
Hidden size
3072 (24 heads × 128)
FFN size
12288 (GELU-tanh), separate image and text FFNs
Text conditioning
Joint attention with Qwen2.5-VL last hidden states (3584-d). System-prompt template tokens (34) are dropped; up to 1024 tokens
Guidance
No guidance embedding. True CFG with a negative prompt
Objective
Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
Norm / position
QK RMSNorm, 3D RoPE (16/56/56 axes)

In AI Toolkit

model.arch
qwen_image
UI label
Qwen-Image (image)
model.name_or_path
Qwen/Qwen-Image
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true
qtype
qfloat8
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
network.conv
disabled (linear LoRA only)
ARA option
3 bit with ARA (uint3)

Specifics

Resolution
Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
Transformer source
For Qwen/Qwen-Image the transformer weights come from Comfy-Org/Qwen-Image_ComfyUI: fp8mixed when qtype is qfloat8, bf16 when unquantized or for other qtypes. A local copy in the ComfyUI models folder wins over a download. Set model_kwargs.use_comfy_weights: false to load the Diffusers repo instead.
Text encoder
Qwen2.5-VL from text_encoder/ in name_or_path (or extras_name_or_path), vision tower removed. Never trained. With low_vram it stays on the CPU and moves to the GPU only to encode.
Single-file checkpoints
A .safetensors name_or_path loads with the Qwen/Qwen-Image transformer config, and the text encoder, tokenizer and VAE come from Qwen/Qwen-Image.
Accuracy recovery adapter
qtype "uint3|<adapter>" quantizes the adapter-covered linears to that qtype and everything else to uint8, then keeps the adapter live on the model. Cannot be combined with assistant_lora_path.
Loss
Flow-matching velocity target: noise − latents.
LoRA target
QwenImageTransformer2DModel
Saving
LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file (the Diffusers key layout is the ComfyUI layout for this model).
Sampling
Flowmatch Euler with true CFG (guidance_scale is true_cfg_scale). low_vram also turns on VAE tiling for the decode. Control images are not supported.
Metadata base version
qwen_image

Example config

not verified

job: extensionconfig:  name: "my_qwen_image_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Qwen/Qwen-Image"        arch: "qwen_image"        quantize: true        qtype: "qfloat8"        # 24 GB: qtype: "uint3|ostris/accuracy_recovery_adapters/qwen_image_torchao_uint3.safetensors"        quantize_te: true        qtype_te: "qfloat8"        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 3        sample_steps: 25        prompts:          - "a man holding a sign that says, 'this is a sign'"

On this page