Qwen-Image-2.1
'One 7B single-stream DiT that does both text-to-image and editing with up to 10 reference images, conditioned on Qwen3-VL-8B. Its new 16×, 64-channel VAE is natively RGBA, so it can generate and edit transparent images.'
- weights
- Qwen/Qwen-Image-2.1 ↗
- org
- Qwen (Alibaba)
- modality
- image
- tasks
- text-to-image · image-editing · multi-image-editing · transparent (RGBA) generation
- license
- Qwen Research License
- released
- 2026-09-14not verified
- native output
- 2048×2048 (1:1), 2752×1536 (16:9) and other ~4 MP aspect ratios
- total params
- 16.22B
- model.arch
- qwen_image_2
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Qwen-Image-2.1 DiT (7B) QwenImage21Transformer2DModel (vendored in ai-toolkit) ai-toolkit loads the Comfy-Org repack instead: diffusion_models/qwen_image_2.1_int8_convrot.safetensors (7,256,783,064 bytes, pre-quantized int8 convrot) with the default convrot8 qtype, or qwen_image_2.1_bf16.safetensors (14,230,280,616 bytes) unquantized. The Qwen repo supplies only the config. | 7.12B 7,115,124,736 | 14.23 GB | bf16 | yes |
| Text encoder | Qwen3-VL-8B transformers.Qwen3VLForConditionalGeneration Count includes the vision tower, which stays loaded because any prompt may carry reference images. ai-toolkit loads a Comfy-Org repack file (text_encoders/qwen3vl_8b_int8_convrot.safetensors, 9,350,798,360 bytes, or qwen3vl_8b_bf16.safetensors, 17,534,334,616 bytes) and quantizes to qtype_te after loading. | 8.77B 8,767,123,696 | 17.53 GB | bf16 | no |
| Processor | Qwen3-VL processor (tokenizer + image processor) transformers.Qwen3VLProcessor Always loaded from the Qwen repo: the Comfy-Org repack does not carry it. | — | — | — | — |
| VAE | Qwen-Image-2.1 VAE (RGBA) AutoencoderKLQwenImage21 (vendored in ai-toolkit) 4 input and output channels (RGBA). The decoder is wider than the encoder (base dim 144 vs 96). ai-toolkit loads the Comfy-Org bf16 copy, vae/qwen_image_2.1_vae_bf16.safetensors (675,509,688 bytes), with the same parameter count. | 337.7M 337,740,404 | 1.35 GB | fp32 | no |
| total | 16.22B | 33.12 GB | |||
Latent space
- autoencoder
- Qwen-Image-2.1 VAE
- pixels per token
- 16×16
- notes
- The VAE downsamples 4 times for 16× spatial (scale_factor_spatial 16) into 64 channels, and the transformer takes latents unpatched, so one token still covers 16×16 pixels. Qwen3-VL vision tokens cover 32×32 pixels, so each reference image slot in the prompt maps to a 2×2 group of latent tokens. The VAE is a video VAE (8× temporal); images use one frame.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 64×1×64×64 | 4,096 | ai-toolkit sample default |
| 2048×2048 | 64×1×128×128 | 16,384 | native 1:1 |
| 2752×1536 | 64×1×96×172 | 16,512 | native 16:9 |
Architecture
- Blocks
- 32 single-stream blocks, one shared modulation projection
- Hidden size
- 4096 (32 heads × 128)
- FFN size
- 12288 (SwiGLU, mlp_ratio 3)
- Text conditioning
- Qwen3-VL last decoder layer before its final RMSNorm (4096-d), in the same sequence as the image tokens
- Image conditioning
- Each reference fills the <|image_pad|> slots Qwen3-VL reserved for it in the text stream, 4 latent tokens per slot
- Attention
- Block-causal: causal across the sequence, bidirectional inside each image block
- Timesteps
- causal_condition: text and reference tokens are modulated at t = 0, only the target at the sampled t
- Objective
- Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
- Position
- 3D RoPE (16/56/56 axes)
In AI Toolkit
- model.arch
- qwen_image_2
- UI label
- Qwen-Image-2.1 (image)
- model.name_or_path
- Comfy-Org/Qwen-Image-2.1
- source
- extensions_built_in/diffusion_models/qwen_image_2/qwen_image_2.py
- extensions_built_in/diffusion_models/qwen_image_2/src/pipeline.py
- extensions_built_in/diffusion_models/qwen_image_2/src/text_encoder.py
- extensions_built_in/diffusion_models/qwen_image_2/src/transformer.py
- extensions_built_in/diffusion_models/qwen_image_2/src/vae.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true
- qtype / qtype_te
- convrot8 / convrot8
- low_vram
- true
- train.unload_text_encoder
- false (section hidden)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- shift
- sample.guidance_scale
- 3.0
- model_kwargs.rgba
- false (Transparency (RGBA) checkbox)
- network.conv
- disabled (linear LoRA only)
Specifics
- One arch, two modes
- A dataset without control paths trains plain text-to-image. With control paths it trains editing. The prompt decides: references are only fed to the transformer when the prompt embedding reserved slots for them, so a dropped caption encoded as plain text trains as T2I.
- Training data
- Give reference folders as control_path_1 to control_path_3 (or a control_path list). Each target image is paired with the same-named file in every folder, in folder order. References keep their own size in the dataloader.
- Reference sizing
- By default (model_kwargs.match_target_res, default true) each reference is scaled to the target bucket area, keeping its own aspect ratio, on the 32 px grid. With match_target_res: false they are only shrunk to fit model_kwargs.control_image_max_pixels (default 1024×1024). The same size is used for the Qwen3-VL pass and the VAE pass so slot and token counts agree.
- Text embedding cache
- The cache key includes the control paths, the target bucket size when match_target_res is on, and the sizing rule, so changing either re-encodes. Batches larger than 1 need references with the same token count.
- Transparency
- model_kwargs.rgba: true loads dataset and reference images with alpha, encodes all four channels and saves samples as RGBA PNGs. RGB images get an opaque alpha. Qwen3-VL sees references composited over white. Changing it re-caches latents; it cannot be combined with alpha_mask.
- Weight sources
- Transformer, text encoder and VAE come from the Comfy-Org/Qwen-Image-2.1 files (a local copy in the ComfyUI models folder wins over a download). Configs and the processor come from Qwen/Qwen-Image-2.1. Other qtypes re-quantize layer by layer from the loaded file.
- Text encoder file choice
- The bf16 text encoder file is downloaded and quantized to qtype_te. A local int8 convrot copy is used when the bf16 file is not present.not verified
- Quantization
- img_in, txt_in*, time_text_embed*, modulation*, norm_out*, proj_out stay in full precision.
- Resolution
- Buckets and sample sizes snap to multiples of 32 (one Qwen3-VL vision token).
- Loss
- Flow-matching velocity target: noise − latents, on the target tokens only.
- LoRA target
- QwenImage21Transformer2DModel
- Saving
- The vendored transformer keeps ComfyUI’s fused img_mlp.gate_up projection, so LoRA module names match ComfyUI; keys use the diffusion_model. prefix. Full fine-tunes save a single ComfyUI-format file.
- Sampling
- ai-toolkit’s own flowmatch Euler loop. References from sample.ctrl_img and ctrl_img_1 to ctrl_img_3 (repeats dropped). CFG runs only when guidance_scale > 1. Qwen-Image 2.1 is meant to be sampled without guidance (guidance_scale 1); the UI default is 3.0. Decodes above 1 MP, or with low_vram, are tiled.
- Metadata base version
- qwen_image_2
Example config
not verified
job: extensionconfig: name: "my_qwen_image_2_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/target/images" # leave out control paths to train text-to-image control_path_1: "/path/to/reference/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "shift" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Comfy-Org/Qwen-Image-2.1" arch: "qwen_image_2" quantize: true qtype: "convrot8" quantize_te: true qtype_te: "convrot8" low_vram: true model_kwargs: rgba: false sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 3 sample_steps: 30 samples: - prompt: "change the background to a sunset beach" ctrl_img_1: "/path/to/reference.png" - prompt: "a neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night"Links
Qwen-Image-2512
'The December 2025 update of the Qwen-Image text-to-image model. Same 20B MMDiT, Qwen2.5-VL-7B text encoder and VAE as the original, retrained for more realistic people, finer natural detail and better text rendering.'
Ming-Image 0.1 Design
A 6B Z-Image-style DiT for UI, infographics, posters and other text-heavy design, with RGBA output. It is conditioned by a Ling-mini-2.0 MoE multimodal LLM through two caption streams. ai-toolkit trains it from the int8 ComfyUI repack with a training adapter.