Qwen-Image-Edit-2511
'The December 2025 update of the multi-image Qwen-Image-Edit line. Same 20B MMDiT and Qwen2.5-VL-7B encoder as 2509, plus t = 0 modulation of the reference tokens. Qwen reports less image drift, better character and multi-person consistency, and popular community LoRA effects built in.'
- weights
- Qwen/Qwen-Image-Edit-2511 ↗
- org
- Qwen (Alibaba)
- modality
- image
- tasks
- image-editing · multi-image-editing
- license
- Apache 2.0
- released
- 2025-12-23not verified
- native output
- About 1 MP at the last reference image aspect ratio
- total params
- 28.85B
- model.arch
- qwen_image_edit_plus:2511
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Qwen-Image-Edit-2511 MMDiT (20B) diffusers.QwenImageTransformer2DModel Same shape and parameter count as 2509. Its config adds zero_cond_t: true, which the Diffusers transformer reads. ai-toolkit loads it from this repo. | 20.43B 20,430,401,088 | 40.86 GB | bf16 | yes |
| Text encoder | Qwen2.5-VL-7B-Instruct transformers.Qwen2_5_VLForConditionalGeneration Includes the vision tower, which the edit models keep: every reference image is shown to it. | 8.29B 8,292,166,656 | 16.58 GB | bf16 | no |
| Tokenizer | Qwen2 tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| Processor | Qwen2-VL image processor transformers.Qwen2VLProcessor | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage The same Wan 2.1-style VAE as Qwen-Image. Images go through it as a single frame. | 126.9M 126,892,531 | 254 MB | bf16 | no |
| Training adapter | Accuracy recovery adapter, 3-bitadapter Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_2511_torchao_uint3.safetensors" (the UI option "3 bit with ARA"). Trained for 2511; the 2509 adapter is a different file. | 148.0M 147,961,856 | 296 MB | bf16 | no |
| total | 28.85B | 57.70 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- 8× spatial, 2×2 patches into 64-channel tokens, so one token covers 16×16 pixels. Each reference is VAE-encoded at about 1 MP by default and appended after the target tokens, so every reference adds roughly another 1 MP worth of tokens.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default (target only) |
| 1328×1328 | 16×1×166×166 | 6,889 | Qwen-Image native 1:1 (target only) |
Architecture
- Blocks
- 60 double-stream (MMDiT) blocks
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288 (GELU-tanh), separate image and text FFNs
- Text conditioning
- Joint attention with Qwen2.5-VL last hidden states (3584-d). Each reference is inserted as "Picture N:" vision tokens before the prompt; template tokens (64) are dropped
- Image conditioning
- Each reference latent appended to the image tokens as its own RoPE frame (1, 2, 3, ...)
- Timesteps
- zero_cond_t: the target tokens are modulated with the sampled timestep, reference tokens with t = 0, and the text stream with the sampled timestep
- Guidance
- No guidance embedding. True CFG with a negative prompt
- Objective
- Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
- Norm / position
- QK RMSNorm, 3D RoPE (16/56/56 axes)
In AI Toolkit
- model.arch
- qwen_image_edit_plus:2511
- UI label
- Qwen-Image-Edit-2511 (instruction)
- model.name_or_path
- Qwen/Qwen-Image-Edit-2511
- source
- extensions_built_in/diffusion_models/qwen_image/qwen_image_edit_plus.py
- extensions_built_in/diffusion_models/qwen_image/qwen_image_pipelines.py
- extensions_built_in/diffusion_models/qwen_image/qwen_image.py
- toolkit/models/v2/diffusion_models/qwen_image.py
- toolkit/models/v2/text_encoders/qwen25_vl.py
- toolkit/models/v2/vae/qwen_image.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- train.unload_text_encoder
- false (section hidden)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- model_kwargs.match_target_res
- false
- network.conv
- disabled (linear LoRA only)
- ARA option
- 3 bit with ARA (uint3)
Specifics
- Arch variant
- qwen_image_edit_plus:2511 is a UI tag. ModelConfig strips everything after the colon, so it trains with the same QwenImageEditPlusModel code as 2509. The t = 0 reference modulation comes from the transformer config (zero_cond_t), applied by Diffusers from the per-image shapes ai-toolkit passes in.
- Training data
- Give up to three reference folders as control_path_1, control_path_2, control_path_3 (or a control_path list). Each target image is paired with the same-named file in every folder, in folder order; that order is Picture 1, 2, 3 in the prompt. Captions are the edit instruction.
- References into the text encoder
- References keep their own size and aspect ratio in the dataloader. For Qwen2.5-VL each one is resized to a 384×384 pixel area (own aspect, snapped to 32 px). Each sample is encoded separately and the embeddings are right-padded to batch them. With cache_text_embeddings the control paths are part of the cache key.
- References into the transformer
- Each reference is resized to a 1024×1024 pixel area (own aspect, snapped to 32 px), VAE-encoded, packed and appended after the noisy target tokens. With model_kwargs.match_target_res: true they are sized to the target bucket area instead. References stay clean; only the target tokens go into the loss.
- Sampling
- Set sample.ctrl_img_1 to ctrl_img_3 (or ctrl_img) per prompt. The pipeline is a custom QwenImageEditPlus pipeline with do_cfg_norm off by default: ai-toolkit found the official CFG renormalization hurts more often than it helps.
- Text encoder
- Qwen2.5-VL with its vision tower, plus the processor. Never trained. The UI forces unload_text_encoder off. The model refuses a prompt without references, so static prompts (blank, trigger, unconditional) are encoded with a random-noise 224×224 reference instead.
- Resolution
- Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
- Single-file checkpoints
- A .safetensors name_or_path loads with the Qwen/Qwen-Image transformer config, which has no zero_cond_t, so a 2511 single file would train without the t = 0 reference modulation. The text encoder and VAE also fall back to Qwen/Qwen-Image, which has no processor; set extras_name_or_path to the Edit repo.not verified
- Loss
- Flow-matching velocity target: noise − latents, on the target tokens only.
- LoRA target
- QwenImageTransformer2DModel
- Saving
- LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file.
- Metadata base version
- qwen_image (the edit class does not override it)
Example config
not verified
job: extensionconfig: name: "my_qwen_image_edit_2511_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/target/images" control_path_1: "/path/to/reference/images_1" control_path_2: "/path/to/reference/images_2" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Qwen/Qwen-Image-Edit-2511" arch: "qwen_image_edit_plus:2511" quantize: true qtype: "qfloat8" # 3-bit: qtype: "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_2511_torchao_uint3.safetensors" quantize_te: true qtype_te: "qfloat8" low_vram: true model_kwargs: match_target_res: false sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 3 sample_steps: 25 samples: - prompt: "the person in Picture 1 wearing the jacket from Picture 2" ctrl_img_1: "/path/to/image1.png" ctrl_img_2: "/path/to/image2.png"Links
Qwen-Image-Edit-2509
'The September 2025 update of Qwen-Image-Edit, and the first to take several reference images at once (best with 1 to 3). Same 20B MMDiT and Qwen2.5-VL-7B encoder, with better identity, product and text consistency than the original edit model.'
HiDream-E1.1
'The instruction-based image editing model built on HiDream-I1: the same 17B mixture-of-experts transformer and four text encoders, fine-tuned to take a source image placed beside the target in the latent and edit it from a text instruction.'