Qwen-Image-Edit
'The first image-editing version of the 20B Qwen-Image model. One reference image goes to both Qwen2.5-VL (for meaning) and the VAE (for appearance), so it handles both semantic edits like restyling and precise local edits, including editing text in the image.'
- weights
- Qwen/Qwen-Image-Edit ↗
- org
- Qwen (Alibaba)
- modality
- image
- tasks
- image-editing
- license
- Apache 2.0
- released
- 2025-08-18not verified
- native output
- About 1 MP at the input image aspect ratio
- total params
- 28.85B
- model.arch
- qwen_image_edit
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Qwen-Image-Edit MMDiT (20B) diffusers.QwenImageTransformer2DModel Same shape as the Qwen-Image transformer, fine-tuned for editing. ai-toolkit loads it from this repo (no ComfyUI file substitution for the edit models). | 20.43B 20,430,401,088 | 40.86 GB | bf16 | yes |
| Text encoder | Qwen2.5-VL-7B-Instruct transformers.Qwen2_5_VLForConditionalGeneration Includes the vision tower, which the edit models keep: the reference image is shown to it. | 8.29B 8,292,166,656 | 16.58 GB | bf16 | no |
| Tokenizer | Qwen2 tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| Processor | Qwen2-VL image processor transformers.Qwen2VLProcessor | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage The same Wan 2.1-style VAE as Qwen-Image. Images go through it as a single frame. | 126.9M 126,892,531 | 254 MB | bf16 | no |
| Training adapter | Accuracy recovery adapter, 3-bitadapter Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_torchao_uint3.safetensors" (the UI option "3 bit with ARA"). The transformer is quantized to 3-bit torchao and this LoRA-shaped adapter runs alongside it to recover accuracy. | 148.0M 147,961,856 | 296 MB | bf16 | no |
| total | 28.85B | 57.70 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- 8× spatial, 2×2 patches into 64-channel tokens, so one token covers 16×16 pixels. The reference image is encoded the same way and its tokens are appended after the target tokens, so the sequence is about twice the target token count.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default (target only) |
| 1328×1328 | 16×1×166×166 | 6,889 | Qwen-Image native 1:1 (target only) |
Architecture
- Blocks
- 60 double-stream (MMDiT) blocks
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288 (GELU-tanh), separate image and text FFNs
- Text conditioning
- Joint attention with Qwen2.5-VL last hidden states (3584-d) of the prompt plus the reference image. Template tokens (64) are dropped
- Image conditioning
- Reference latents appended to the image tokens as a second frame in RoPE; the same image is also seen by Qwen2.5-VL
- Guidance
- No guidance embedding. True CFG with a negative prompt
- Objective
- Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
- Norm / position
- QK RMSNorm, 3D RoPE (16/56/56 axes)
In AI Toolkit
- model.arch
- qwen_image_edit
- UI label
- Qwen-Image-Edit (instruction)
- model.name_or_path
- Qwen/Qwen-Image-Edit
- source
- extra UI sections
- datasets.control_path, sample.ctrl_img, model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- network.conv
- disabled (linear LoRA only)
- ARA option
- 3 bit with ARA (uint3)
Specifics
- Training data
- Set datasets.control_path to a folder of reference images. Each target image is paired with the file of the same name (any of .jpg, .jpeg, .png, .webp) in that folder. Captions are the edit instruction.
- Reference into the text encoder
- The reference is resized to about 1 MP (1024×1024 area, its own aspect ratio, snapped to 32 px) and encoded with the prompt by Qwen2.5-VL. With cache_text_embeddings the control path is part of the cache key.
- Reference into the transformer
- The reference is resized to the target bucket size, VAE-encoded, packed and appended after the noisy target tokens. It stays clean, and only the target tokens are kept from the prediction for the loss.
- Sampling
- Set sample.ctrl_img per prompt; it is resized to the sample width and height. True CFG (guidance_scale is true_cfg_scale); low_vram turns on VAE tiling.
- Resolution
- Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
- Text encoder
- Qwen2.5-VL with its vision tower kept, plus the processor from processor/. Never trained. With low_vram it stays on the CPU and moves to the GPU only to encode.
- Single-file checkpoints
- A .safetensors name_or_path falls back to Qwen/Qwen-Image for the transformer config, text encoder and VAE, and that repo has no processor. Set extras_name_or_path to Qwen/Qwen-Image-Edit.not verified
- Accuracy recovery adapter
- qtype "uint3|<adapter>" quantizes the adapter-covered linears to that qtype and everything else to uint8, then keeps the adapter live on the model.
- Loss
- Flow-matching velocity target: noise − latents, on the target tokens only.
- LoRA target
- QwenImageTransformer2DModel
- Saving
- LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file.
- Metadata base version
- qwen_image (the edit class does not override it)
Example config
not verified
job: extensionconfig: name: "my_qwen_image_edit_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/target/images" control_path: "/path/to/reference/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Qwen/Qwen-Image-Edit" arch: "qwen_image_edit" quantize: true qtype: "qfloat8" quantize_te: true qtype_te: "qfloat8" low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 3 sample_steps: 25 samples: - prompt: "make the sky a sunset" ctrl_img: "/path/to/reference.png"Links
FLUX.1 Kontext [dev]
The instruction-editing member of FLUX.1. Same 12B transformer shape, text encoders and VAE as FLUX.1 [dev]; the input image goes in as a second set of tokens next to the image being generated. Trained in ai-toolkit on before/after pairs.
Qwen-Image-Edit-2509
'The September 2025 update of Qwen-Image-Edit, and the first to take several reference images at once (best with 1 to 3). Same 20B MMDiT and Qwen2.5-VL-7B encoder, with better identity, product and text consistency than the original edit model.'