Docs
AI ToolkitModels

Qwen-Image-Edit-2511

'The December 2025 update of the multi-image Qwen-Image-Edit line. Same 20B MMDiT and Qwen2.5-VL-7B encoder as 2509, plus t = 0 modulation of the reference tokens. Qwen reports less image drift, better character and multi-person consistency, and popular community LoRA effects built in.'

org
Qwen (Alibaba)
modality
image
tasks
image-editing · multi-image-editing
license
Apache 2.0
released
2025-12-23not verified
native output
About 1 MP at the last reference image aspect ratio
total params
28.85B
model.arch
qwen_image_edit_plus:2511

Components

rolemodelparamssizedtypetrained
Transformer
Qwen-Image-Edit-2511 MMDiT (20B)
diffusers.QwenImageTransformer2DModel

Same shape and parameter count as 2509. Its config adds zero_cond_t: true, which the Diffusers transformer reads. ai-toolkit loads it from this repo.

20.43B
20,430,401,088
40.86 GBbf16yes
Text encoder
Qwen2.5-VL-7B-Instruct
transformers.Qwen2_5_VLForConditionalGeneration

Includes the vision tower, which the edit models keep: every reference image is shown to it.

8.29B
8,292,166,656
16.58 GBbf16no
Tokenizer
Qwen2 tokenizer
transformers.Qwen2Tokenizer
————
Processor
Qwen2-VL image processor
transformers.Qwen2VLProcessor
————
VAE
Qwen-Image VAE
diffusers.AutoencoderKLQwenImage

The same Wan 2.1-style VAE as Qwen-Image. Images go through it as a single frame.

126.9M
126,892,531
254 MBbf16no
Training adapter
Accuracy recovery adapter, 3-bitadapter

Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_2511_torchao_uint3.safetensors" (the UI option "3 bit with ARA"). Trained for 2511; the 2509 adapter is a different file.

148.0M
147,961,856
296 MBbf16no
total28.85B57.70 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Qwen-Image VAE
pixels per token
16×16
notes
8× spatial, 2×2 patches into 64-channel tokens, so one token covers 16×16 pixels. Each reference is VAE-encoded at about 1 MP by default and appended after the target tokens, so every reference adds roughly another 1 MP worth of tokens.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default (target only)
1328×132816×1×166×1666,889Qwen-Image native 1:1 (target only)

Architecture

Blocks
60 double-stream (MMDiT) blocks
Hidden size
3072 (24 heads × 128)
FFN size
12288 (GELU-tanh), separate image and text FFNs
Text conditioning
Joint attention with Qwen2.5-VL last hidden states (3584-d). Each reference is inserted as "Picture N:" vision tokens before the prompt; template tokens (64) are dropped
Image conditioning
Each reference latent appended to the image tokens as its own RoPE frame (1, 2, 3, ...)
Timesteps
zero_cond_t: the target tokens are modulated with the sampled timestep, reference tokens with t = 0, and the text stream with the sampled timestep
Guidance
No guidance embedding. True CFG with a negative prompt
Objective
Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
Norm / position
QK RMSNorm, 3D RoPE (16/56/56 axes)

In AI Toolkit

model.arch
qwen_image_edit_plus:2511
UI label
Qwen-Image-Edit-2511 (instruction)
model.name_or_path
Qwen/Qwen-Image-Edit-2511
extra UI sections
datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res

UI defaults

quantize / quantize_te
true / true
qtype
qfloat8
low_vram
true
train.unload_text_encoder
false (section hidden)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
model_kwargs.match_target_res
false
network.conv
disabled (linear LoRA only)
ARA option
3 bit with ARA (uint3)

Specifics

Arch variant
qwen_image_edit_plus:2511 is a UI tag. ModelConfig strips everything after the colon, so it trains with the same QwenImageEditPlusModel code as 2509. The t = 0 reference modulation comes from the transformer config (zero_cond_t), applied by Diffusers from the per-image shapes ai-toolkit passes in.
Training data
Give up to three reference folders as control_path_1, control_path_2, control_path_3 (or a control_path list). Each target image is paired with the same-named file in every folder, in folder order; that order is Picture 1, 2, 3 in the prompt. Captions are the edit instruction.
References into the text encoder
References keep their own size and aspect ratio in the dataloader. For Qwen2.5-VL each one is resized to a 384×384 pixel area (own aspect, snapped to 32 px). Each sample is encoded separately and the embeddings are right-padded to batch them. With cache_text_embeddings the control paths are part of the cache key.
References into the transformer
Each reference is resized to a 1024×1024 pixel area (own aspect, snapped to 32 px), VAE-encoded, packed and appended after the noisy target tokens. With model_kwargs.match_target_res: true they are sized to the target bucket area instead. References stay clean; only the target tokens go into the loss.
Sampling
Set sample.ctrl_img_1 to ctrl_img_3 (or ctrl_img) per prompt. The pipeline is a custom QwenImageEditPlus pipeline with do_cfg_norm off by default: ai-toolkit found the official CFG renormalization hurts more often than it helps.
Text encoder
Qwen2.5-VL with its vision tower, plus the processor. Never trained. The UI forces unload_text_encoder off. The model refuses a prompt without references, so static prompts (blank, trigger, unconditional) are encoded with a random-noise 224×224 reference instead.
Resolution
Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
Single-file checkpoints
A .safetensors name_or_path loads with the Qwen/Qwen-Image transformer config, which has no zero_cond_t, so a 2511 single file would train without the t = 0 reference modulation. The text encoder and VAE also fall back to Qwen/Qwen-Image, which has no processor; set extras_name_or_path to the Edit repo.not verified
Loss
Flow-matching velocity target: noise − latents, on the target tokens only.
LoRA target
QwenImageTransformer2DModel
Saving
LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file.
Metadata base version
qwen_image (the edit class does not override it)

Example config

not verified

job: extensionconfig:  name: "my_qwen_image_edit_2511_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/target/images"          control_path_1: "/path/to/reference/images_1"          control_path_2: "/path/to/reference/images_2"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Qwen/Qwen-Image-Edit-2511"        arch: "qwen_image_edit_plus:2511"        quantize: true        qtype: "qfloat8"        # 3-bit: qtype: "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_2511_torchao_uint3.safetensors"        quantize_te: true        qtype_te: "qfloat8"        low_vram: true        model_kwargs:          match_target_res: false      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 3        sample_steps: 25        samples:          - prompt: "the person in Picture 1 wearing the jacket from Picture 2"            ctrl_img_1: "/path/to/image1.png"            ctrl_img_2: "/path/to/image2.png"

On this page