Docs
AI ToolkitModels

Qwen-Image-Edit

'The first image-editing version of the 20B Qwen-Image model. One reference image goes to both Qwen2.5-VL (for meaning) and the VAE (for appearance), so it handles both semantic edits like restyling and precise local edits, including editing text in the image.'

org
Qwen (Alibaba)
modality
image
tasks
image-editing
license
Apache 2.0
released
2025-08-18not verified
native output
About 1 MP at the input image aspect ratio
total params
28.85B
model.arch
qwen_image_edit

Components

rolemodelparamssizedtypetrained
Transformer
Qwen-Image-Edit MMDiT (20B)
diffusers.QwenImageTransformer2DModel

Same shape as the Qwen-Image transformer, fine-tuned for editing. ai-toolkit loads it from this repo (no ComfyUI file substitution for the edit models).

20.43B
20,430,401,088
40.86 GBbf16yes
Text encoder
Qwen2.5-VL-7B-Instruct
transformers.Qwen2_5_VLForConditionalGeneration

Includes the vision tower, which the edit models keep: the reference image is shown to it.

8.29B
8,292,166,656
16.58 GBbf16no
Tokenizer
Qwen2 tokenizer
transformers.Qwen2Tokenizer
————
Processor
Qwen2-VL image processor
transformers.Qwen2VLProcessor
————
VAE
Qwen-Image VAE
diffusers.AutoencoderKLQwenImage

The same Wan 2.1-style VAE as Qwen-Image. Images go through it as a single frame.

126.9M
126,892,531
254 MBbf16no
Training adapter
Accuracy recovery adapter, 3-bitadapter

Used with qtype "uint3|ostris/accuracy_recovery_adapters/qwen_image_edit_torchao_uint3.safetensors" (the UI option "3 bit with ARA"). The transformer is quantized to 3-bit torchao and this LoRA-shaped adapter runs alongside it to recover accuracy.

148.0M
147,961,856
296 MBbf16no
total28.85B57.70 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Qwen-Image VAE
pixels per token
16×16
notes
8× spatial, 2×2 patches into 64-channel tokens, so one token covers 16×16 pixels. The reference image is encoded the same way and its tokens are appended after the target tokens, so the sequence is about twice the target token count.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default (target only)
1328×132816×1×166×1666,889Qwen-Image native 1:1 (target only)

Architecture

Blocks
60 double-stream (MMDiT) blocks
Hidden size
3072 (24 heads × 128)
FFN size
12288 (GELU-tanh), separate image and text FFNs
Text conditioning
Joint attention with Qwen2.5-VL last hidden states (3584-d) of the prompt plus the reference image. Template tokens (64) are dropped
Image conditioning
Reference latents appended to the image tokens as a second frame in RoPE; the same image is also seen by Qwen2.5-VL
Guidance
No guidance embedding. True CFG with a negative prompt
Objective
Rectified flow, resolution-dependent exponential shift (0.5 to 0.9)
Norm / position
QK RMSNorm, 3D RoPE (16/56/56 axes)

In AI Toolkit

model.arch
qwen_image_edit
UI label
Qwen-Image-Edit (instruction)
model.name_or_path
Qwen/Qwen-Image-Edit
extra UI sections
datasets.control_path, sample.ctrl_img, model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true
qtype
qfloat8
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
network.conv
disabled (linear LoRA only)
ARA option
3 bit with ARA (uint3)

Specifics

Training data
Set datasets.control_path to a folder of reference images. Each target image is paired with the file of the same name (any of .jpg, .jpeg, .png, .webp) in that folder. Captions are the edit instruction.
Reference into the text encoder
The reference is resized to about 1 MP (1024×1024 area, its own aspect ratio, snapped to 32 px) and encoded with the prompt by Qwen2.5-VL. With cache_text_embeddings the control path is part of the cache key.
Reference into the transformer
The reference is resized to the target bucket size, VAE-encoded, packed and appended after the noisy target tokens. It stays clean, and only the target tokens are kept from the prediction for the loss.
Sampling
Set sample.ctrl_img per prompt; it is resized to the sample width and height. True CFG (guidance_scale is true_cfg_scale); low_vram turns on VAE tiling.
Resolution
Buckets and sample sizes snap to multiples of 32 (8× VAE × 2×2 patch).
Text encoder
Qwen2.5-VL with its vision tower kept, plus the processor from processor/. Never trained. With low_vram it stays on the CPU and moves to the GPU only to encode.
Single-file checkpoints
A .safetensors name_or_path falls back to Qwen/Qwen-Image for the transformer config, text encoder and VAE, and that repo has no processor. Set extras_name_or_path to Qwen/Qwen-Image-Edit.not verified
Accuracy recovery adapter
qtype "uint3|<adapter>" quantizes the adapter-covered linears to that qtype and everything else to uint8, then keeps the adapter live on the model.
Loss
Flow-matching velocity target: noise − latents, on the target tokens only.
LoRA target
QwenImageTransformer2DModel
Saving
LoRA keys use the diffusion_model. prefix, which ComfyUI loads. Full fine-tunes save a single ComfyUI-format file.
Metadata base version
qwen_image (the edit class does not override it)

Example config

not verified

job: extensionconfig:  name: "my_qwen_image_edit_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/target/images"          control_path: "/path/to/reference/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Qwen/Qwen-Image-Edit"        arch: "qwen_image_edit"        quantize: true        qtype: "qfloat8"        quantize_te: true        qtype_te: "qfloat8"        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 3        sample_steps: 25        samples:          - prompt: "make the sky a sunset"            ctrl_img: "/path/to/reference.png"

On this page