Docs
AI ToolkitModels

Qwen-Image-2.1

'One 7B single-stream DiT that does both text-to-image and editing with up to 10 reference images, conditioned on Qwen3-VL-8B. Its new 16×, 64-channel VAE is natively RGBA, so it can generate and edit transparent images.'

org
Qwen (Alibaba)
modality
image
tasks
text-to-image · image-editing · multi-image-editing · transparent (RGBA) generation
license
Qwen Research License
released
2026-09-14not verified
native output
2048×2048 (1:1), 2752×1536 (16:9) and other ~4 MP aspect ratios
total params
16.22B
model.arch
qwen_image_2

Components

rolemodelparamssizedtypetrained
Transformer
Qwen-Image-2.1 DiT (7B)
QwenImage21Transformer2DModel (vendored in ai-toolkit)

ai-toolkit loads the Comfy-Org repack instead: diffusion_models/qwen_image_2.1_int8_convrot.safetensors (7,256,783,064 bytes, pre-quantized int8 convrot) with the default convrot8 qtype, or qwen_image_2.1_bf16.safetensors (14,230,280,616 bytes) unquantized. The Qwen repo supplies only the config.

7.12B
7,115,124,736
14.23 GBbf16yes
Text encoder
Qwen3-VL-8B
transformers.Qwen3VLForConditionalGeneration

Count includes the vision tower, which stays loaded because any prompt may carry reference images. ai-toolkit loads a Comfy-Org repack file (text_encoders/qwen3vl_8b_int8_convrot.safetensors, 9,350,798,360 bytes, or qwen3vl_8b_bf16.safetensors, 17,534,334,616 bytes) and quantizes to qtype_te after loading.

8.77B
8,767,123,696
17.53 GBbf16no
Processor
Qwen3-VL processor (tokenizer + image processor)
transformers.Qwen3VLProcessor

Always loaded from the Qwen repo: the Comfy-Org repack does not carry it.

————
VAE
Qwen-Image-2.1 VAE (RGBA)
AutoencoderKLQwenImage21 (vendored in ai-toolkit)

4 input and output channels (RGBA). The decoder is wider than the encoder (base dim 144 vs 96). ai-toolkit loads the Comfy-Org bf16 copy, vae/qwen_image_2.1_vae_bf16.safetensors (675,509,688 bytes), with the same parameter count.

337.7M
337,740,404
1.35 GBfp32no
total16.22B33.12 GB

Latent space

spatial
16×
channels
64
patch
1×1
autoencoder
Qwen-Image-2.1 VAE
pixels per token
16×16
notes
The VAE downsamples 4 times for 16× spatial (scale_factor_spatial 16) into 64 channels, and the transformer takes latents unpatched, so one token still covers 16×16 pixels. Qwen3-VL vision tokens cover 32×32 pixels, so each reference image slot in the prompt maps to a 2×2 group of latent tokens. The VAE is a video VAE (8× temporal); images use one frame.
inputlatent (c×t×h×w)tokens
1024×102464×1×64×644,096ai-toolkit sample default
2048×204864×1×128×12816,384native 1:1
2752×153664×1×96×17216,512native 16:9

Architecture

Blocks
32 single-stream blocks, one shared modulation projection
Hidden size
4096 (32 heads × 128)
FFN size
12288 (SwiGLU, mlp_ratio 3)
Text conditioning
Qwen3-VL last decoder layer before its final RMSNorm (4096-d), in the same sequence as the image tokens
Image conditioning
Each reference fills the <|image_pad|> slots Qwen3-VL reserved for it in the text stream, 4 latent tokens per slot
Attention
Block-causal: causal across the sequence, bidirectional inside each image block
Timesteps
causal_condition: text and reference tokens are modulated at t = 0, only the target at the sampled t
Objective
Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
Position
3D RoPE (16/56/56 axes)

In AI Toolkit

model.arch
qwen_image_2
UI label
Qwen-Image-2.1 (image)
model.name_or_path
Comfy-Org/Qwen-Image-2.1
extra UI sections
datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true
qtype / qtype_te
convrot8 / convrot8
low_vram
true
train.unload_text_encoder
false (section hidden)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
shift
sample.guidance_scale
3.0
model_kwargs.rgba
false (Transparency (RGBA) checkbox)
network.conv
disabled (linear LoRA only)

Specifics

One arch, two modes
A dataset without control paths trains plain text-to-image. With control paths it trains editing. The prompt decides: references are only fed to the transformer when the prompt embedding reserved slots for them, so a dropped caption encoded as plain text trains as T2I.
Training data
Give reference folders as control_path_1 to control_path_3 (or a control_path list). Each target image is paired with the same-named file in every folder, in folder order. References keep their own size in the dataloader.
Reference sizing
By default (model_kwargs.match_target_res, default true) each reference is scaled to the target bucket area, keeping its own aspect ratio, on the 32 px grid. With match_target_res: false they are only shrunk to fit model_kwargs.control_image_max_pixels (default 1024×1024). The same size is used for the Qwen3-VL pass and the VAE pass so slot and token counts agree.
Text embedding cache
The cache key includes the control paths, the target bucket size when match_target_res is on, and the sizing rule, so changing either re-encodes. Batches larger than 1 need references with the same token count.
Transparency
model_kwargs.rgba: true loads dataset and reference images with alpha, encodes all four channels and saves samples as RGBA PNGs. RGB images get an opaque alpha. Qwen3-VL sees references composited over white. Changing it re-caches latents; it cannot be combined with alpha_mask.
Weight sources
Transformer, text encoder and VAE come from the Comfy-Org/Qwen-Image-2.1 files (a local copy in the ComfyUI models folder wins over a download). Configs and the processor come from Qwen/Qwen-Image-2.1. Other qtypes re-quantize layer by layer from the loaded file.
Text encoder file choice
The bf16 text encoder file is downloaded and quantized to qtype_te. A local int8 convrot copy is used when the bf16 file is not present.not verified
Quantization
img_in, txt_in*, time_text_embed*, modulation*, norm_out*, proj_out stay in full precision.
Resolution
Buckets and sample sizes snap to multiples of 32 (one Qwen3-VL vision token).
Loss
Flow-matching velocity target: noise − latents, on the target tokens only.
LoRA target
QwenImage21Transformer2DModel
Saving
The vendored transformer keeps ComfyUI’s fused img_mlp.gate_up projection, so LoRA module names match ComfyUI; keys use the diffusion_model. prefix. Full fine-tunes save a single ComfyUI-format file.
Sampling
ai-toolkit’s own flowmatch Euler loop. References from sample.ctrl_img and ctrl_img_1 to ctrl_img_3 (repeats dropped). CFG runs only when guidance_scale > 1. Qwen-Image 2.1 is meant to be sampled without guidance (guidance_scale 1); the UI default is 3.0. Decodes above 1 MP, or with low_vram, are tiled.
Metadata base version
qwen_image_2

Example config

not verified

job: extensionconfig:  name: "my_qwen_image_2_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/target/images"          # leave out control paths to train text-to-image          control_path_1: "/path/to/reference/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "shift"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Comfy-Org/Qwen-Image-2.1"        arch: "qwen_image_2"        quantize: true        qtype: "convrot8"        quantize_te: true        qtype_te: "convrot8"        low_vram: true        model_kwargs:          rgba: false      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 3        sample_steps: 30        samples:          - prompt: "change the background to a sunset beach"            ctrl_img_1: "/path/to/reference.png"          - prompt: "a neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night"

On this page