Qwen-Image-Edit-2509
'The September 2025 update of Qwen-Image-Edit, and the first to take several reference images at once (best with 1 to 3). Same 20B MMDiT and Qwen2.5-VL-7B encoder, with better identity, product and text consistency than the original edit model.'
- weights
- Qwen/Qwen-Image-Edit-2509 ↗
- org
- Qwen (Alibaba)
- modality
- image
- tasks
- image-editing · multi-image-editing
- license
- Apache 2.0
- released
- 2025-09-22not verified
- native output
- About 1 MP at the last reference image aspect ratio
- total params
- 28.85B
- model.arch
- qwen_image_edit_plus
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Qwen-Image-Edit-2509 MMDiT (20B) diffusers.QwenImageTransformer2DModel Same shape as the Qwen-Image transformer. ai-toolkit loads it from this repo. | 20.43B 20,430,401,088 | 40.86 GB | bf16 | yes |
| Text encoder | Qwen2.5-VL-7B-Instruct transformers.Qwen2_5_VLForConditionalGeneration Includes the vision tower, which the edit models keep: every reference image is shown to it. | 8.29B 8,292,166,656 | 16.58 GB | bf16 | no |
| Tokenizer | Qwen2 tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| Processor | Qwen2-VL image processor transformers.Qwen2VLProcessor | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage The same Wan 2.1-style VAE as Qwen-Image. Images go through it as a single frame. | 126.9M 126,892,531 | 254 MB | bf16 | no |
| Training adapter | Accuracy recovery adapter, 3-bitadapter Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_2509_torchao_uint3.safetensors" (the UI option "3 bit with ARA", and the ai-toolkit 32 GB example config). | 148.0M 147,961,856 | 296 MB | bf16 | no |
| total | 28.85B | 57.70 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- 8× spatial, 2×2 patches into 64-channel tokens, so one token covers 16×16 pixels. Each reference is VAE-encoded at about 1 MP by default and appended after the target tokens, so every reference adds roughly another 1 MP worth of tokens.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default (target only) |
| 1328×1328 | 16×1×166×166 | 6,889 | Qwen-Image native 1:1 (target only) |
Architecture
- Blocks
- 60 double-stream (MMDiT) blocks
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288 (GELU-tanh), separate image and text FFNs
- Text conditioning
- Joint attention with Qwen2.5-VL last hidden states (3584-d). Each reference is inserted as "Picture N:" vision tokens before the prompt; template tokens (64) are dropped
- Image conditioning
- Each reference latent appended to the image tokens as its own RoPE frame (1, 2, 3, ...)
- Guidance
- No guidance embedding. True CFG with a negative prompt
- Objective
- Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
- Norm / position
- QK RMSNorm, 3D RoPE (16/56/56 axes)
In AI Toolkit
- model.arch
- qwen_image_edit_plus
- UI label
- Qwen-Image-Edit-2509 (instruction)
- model.name_or_path
- Qwen/Qwen-Image-Edit-2509
- source
- extensions_built_in/diffusion_models/qwen_image/qwen_image_edit_plus.py
- extensions_built_in/diffusion_models/qwen_image/qwen_image_pipelines.py
- extensions_built_in/diffusion_models/qwen_image/qwen_image.py
- toolkit/models/v2/diffusion_models/qwen_image.py
- toolkit/models/v2/text_encoders/qwen25_vl.py
- toolkit/models/v2/vae/qwen_image.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- train.unload_text_encoder
- false (section hidden)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- model_kwargs.match_target_res
- false
- network.conv
- disabled (linear LoRA only)
- ARA option
- 3 bit with ARA (uint3)
Specifics
- Training data
- Give up to three reference folders as control_path_1, control_path_2, control_path_3 (or a control_path list). Each target image is paired with the same-named file in every folder, in folder order; that order is Picture 1, 2, 3 in the prompt. Captions are the edit instruction.
- References into the text encoder
- References keep their own size and aspect ratio in the dataloader. For Qwen2.5-VL each one is resized to a 384×384 pixel area (own aspect, snapped to 32 px). Each sample is encoded separately and the embeddings are right-padded to batch them. With cache_text_embeddings the control paths are part of the cache key.
- References into the transformer
- Each reference is resized to a 1024×1024 pixel area (own aspect, snapped to 32 px), VAE-encoded, packed and appended after the noisy target tokens. With model_kwargs.match_target_res: true they are sized to the target bucket area instead. References stay clean; only the target tokens go into the loss.
- Sampling
- Set sample.ctrl_img_1 to ctrl_img_3 (or ctrl_img) per prompt. The pipeline is a custom QwenImageEditPlus pipeline with do_cfg_norm off by default: ai-toolkit found the official CFG renormalization hurts more often than it helps.
- Text encoder
- Qwen2.5-VL with its vision tower, plus the processor. Never trained. The UI forces unload_text_encoder off. The model refuses a prompt without references, so static prompts (blank, trigger, unconditional) are encoded with a random-noise 224×224 reference instead.
- Resolution
- Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
- Single-file checkpoints
- A .safetensors name_or_path falls back to Qwen/Qwen-Image for the transformer config, text encoder and VAE, and that repo has no processor. Set extras_name_or_path to the Edit repo.not verified
- Loss
- Flow-matching velocity target: noise − latents, on the target tokens only.
- LoRA target
- QwenImageTransformer2DModel
- Saving
- LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file.
- Metadata base version
- qwen_image (the edit class does not override it)
Example config
not verified
job: extensionconfig: name: "my_qwen_image_edit_2509_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/target/images" control_path_1: "/path/to/reference/images_1" control_path_2: "/path/to/reference/images_2" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Qwen/Qwen-Image-Edit-2509" arch: "qwen_image_edit_plus" quantize: true qtype: "qfloat8" # 32 GB: qtype: "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_2509_torchao_uint3.safetensors" quantize_te: true qtype_te: "qfloat8" low_vram: true model_kwargs: match_target_res: false sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 3 sample_steps: 25 samples: - prompt: "the person in Picture 1 wearing the jacket from Picture 2" ctrl_img_1: "/path/to/image1.png" ctrl_img_2: "/path/to/image2.png"Links
Qwen-Image-Edit
'The first image-editing version of the 20B Qwen-Image model. One reference image goes to both Qwen2.5-VL (for meaning) and the VAE (for appearance), so it handles both semantic edits like restyling and precise local edits, including editing text in the image.'
Qwen-Image-Edit-2511
'The December 2025 update of the multi-image Qwen-Image-Edit line. Same 20B MMDiT and Qwen2.5-VL-7B encoder as 2509, plus t = 0 modulation of the reference tokens. Qwen reports less image drift, better character and multi-person consistency, and popular community LoRA effects built in.'