Docs
AI ToolkitModels

Qwen-Image-Edit-2509

'The September 2025 update of Qwen-Image-Edit, and the first to take several reference images at once (best with 1 to 3). Same 20B MMDiT and Qwen2.5-VL-7B encoder, with better identity, product and text consistency than the original edit model.'

org
Qwen (Alibaba)
modality
image
tasks
image-editing · multi-image-editing
license
Apache 2.0
released
2025-09-22not verified
native output
About 1 MP at the last reference image aspect ratio
total params
28.85B
model.arch
qwen_image_edit_plus

Components

rolemodelparamssizedtypetrained
Transformer
Qwen-Image-Edit-2509 MMDiT (20B)
diffusers.QwenImageTransformer2DModel

Same shape as the Qwen-Image transformer. ai-toolkit loads it from this repo.

20.43B
20,430,401,088
40.86 GBbf16yes
Text encoder
Qwen2.5-VL-7B-Instruct
transformers.Qwen2_5_VLForConditionalGeneration

Includes the vision tower, which the edit models keep: every reference image is shown to it.

8.29B
8,292,166,656
16.58 GBbf16no
Tokenizer
Qwen2 tokenizer
transformers.Qwen2Tokenizer
————
Processor
Qwen2-VL image processor
transformers.Qwen2VLProcessor
————
VAE
Qwen-Image VAE
diffusers.AutoencoderKLQwenImage

The same Wan 2.1-style VAE as Qwen-Image. Images go through it as a single frame.

126.9M
126,892,531
254 MBbf16no
Training adapter
Accuracy recovery adapter, 3-bitadapter

Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_2509_torchao_uint3.safetensors" (the UI option "3 bit with ARA", and the ai-toolkit 32 GB example config).

148.0M
147,961,856
296 MBbf16no
total28.85B57.70 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Qwen-Image VAE
pixels per token
16×16
notes
8× spatial, 2×2 patches into 64-channel tokens, so one token covers 16×16 pixels. Each reference is VAE-encoded at about 1 MP by default and appended after the target tokens, so every reference adds roughly another 1 MP worth of tokens.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default (target only)
1328×132816×1×166×1666,889Qwen-Image native 1:1 (target only)

Architecture

Blocks
60 double-stream (MMDiT) blocks
Hidden size
3072 (24 heads × 128)
FFN size
12288 (GELU-tanh), separate image and text FFNs
Text conditioning
Joint attention with Qwen2.5-VL last hidden states (3584-d). Each reference is inserted as "Picture N:" vision tokens before the prompt; template tokens (64) are dropped
Image conditioning
Each reference latent appended to the image tokens as its own RoPE frame (1, 2, 3, ...)
Guidance
No guidance embedding. True CFG with a negative prompt
Objective
Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
Norm / position
QK RMSNorm, 3D RoPE (16/56/56 axes)

In AI Toolkit

model.arch
qwen_image_edit_plus
UI label
Qwen-Image-Edit-2509 (instruction)
model.name_or_path
Qwen/Qwen-Image-Edit-2509
extra UI sections
datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res

UI defaults

quantize / quantize_te
true / true
qtype
qfloat8
low_vram
true
train.unload_text_encoder
false (section hidden)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
model_kwargs.match_target_res
false
network.conv
disabled (linear LoRA only)
ARA option
3 bit with ARA (uint3)

Specifics

Training data
Give up to three reference folders as control_path_1, control_path_2, control_path_3 (or a control_path list). Each target image is paired with the same-named file in every folder, in folder order; that order is Picture 1, 2, 3 in the prompt. Captions are the edit instruction.
References into the text encoder
References keep their own size and aspect ratio in the dataloader. For Qwen2.5-VL each one is resized to a 384×384 pixel area (own aspect, snapped to 32 px). Each sample is encoded separately and the embeddings are right-padded to batch them. With cache_text_embeddings the control paths are part of the cache key.
References into the transformer
Each reference is resized to a 1024×1024 pixel area (own aspect, snapped to 32 px), VAE-encoded, packed and appended after the noisy target tokens. With model_kwargs.match_target_res: true they are sized to the target bucket area instead. References stay clean; only the target tokens go into the loss.
Sampling
Set sample.ctrl_img_1 to ctrl_img_3 (or ctrl_img) per prompt. The pipeline is a custom QwenImageEditPlus pipeline with do_cfg_norm off by default: ai-toolkit found the official CFG renormalization hurts more often than it helps.
Text encoder
Qwen2.5-VL with its vision tower, plus the processor. Never trained. The UI forces unload_text_encoder off. The model refuses a prompt without references, so static prompts (blank, trigger, unconditional) are encoded with a random-noise 224×224 reference instead.
Resolution
Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
Single-file checkpoints
A .safetensors name_or_path falls back to Qwen/Qwen-Image for the transformer config, text encoder and VAE, and that repo has no processor. Set extras_name_or_path to the Edit repo.not verified
Loss
Flow-matching velocity target: noise − latents, on the target tokens only.
LoRA target
QwenImageTransformer2DModel
Saving
LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file.
Metadata base version
qwen_image (the edit class does not override it)

Example config

not verified

job: extensionconfig:  name: "my_qwen_image_edit_2509_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/target/images"          control_path_1: "/path/to/reference/images_1"          control_path_2: "/path/to/reference/images_2"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Qwen/Qwen-Image-Edit-2509"        arch: "qwen_image_edit_plus"        quantize: true        qtype: "qfloat8"        # 32 GB: qtype: "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_2509_torchao_uint3.safetensors"        quantize_te: true        qtype_te: "qfloat8"        low_vram: true        model_kwargs:          match_target_res: false      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 3        sample_steps: 25        samples:          - prompt: "the person in Picture 1 wearing the jacket from Picture 2"            ctrl_img_1: "/path/to/image1.png"            ctrl_img_2: "/path/to/image2.png"

On this page