HiDream-E1.1
'The instruction-based image editing model built on HiDream-I1: the same 17B mixture-of-experts transformer and four text encoders, fine-tuned to take a source image placed beside the target in the latent and edit it from a text instruction.'
- weights
- HiDream-ai/HiDream-E1-1 ↗
- org
- HiDream.ai
- modality
- image
- tasks
- image editing
- license
- MIT
- released
- 2025-07-16not verified
- native output
- 768×768 output beside a 768×768 source
- total params
- 30.80B
- model.arch
- hidream_e1
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | HiDream-E1.1 sparse DiT (MoE) diffusers.HiDreamImageTransformer2DModel Same shape and parameter count as HiDream-I1; only max_resolution differs (96×192 latents instead of 128×128). Includes all 4 routed experts per block; only 2 run per token. | 17.11B 17,105,733,184 | 34.21 GB | bf16 | yes |
| Text encoder 1 | CLIP ViT-L/14 text (with projection) transformers.CLIPTextModelWithProjection Pooled output only (768-d). | 123.8M 123,781,632 | 495 MB | fp32 | no |
| Text encoder 2 | OpenCLIP ViT-bigG/14 text (with projection) transformers.CLIPTextModelWithProjection Pooled output only (1280-d). Concatenated with CLIP-L into the 2048-d pooled vector. | 694.8M 694,840,320 | 2.78 GB | fp32 | no |
| Text encoder 3 | T5-XXL v1.1 (encoder only) transformers.T5EncoderModel | 4.76B 4,762,310,656 | 9.52 GB | bf16 | no |
| Text encoder 4 | Llama 3.1 8B Instruct transformers.LlamaForCausalLM Not in the HiDream repo. Upstream code pulls the gated meta-llama/Meta-Llama-3.1-8B-Instruct; ai-toolkit defaults to the ungated unsloth mirror (override with model_kwargs.llama_model_path). Hidden states from all 32 layers are used, not just the last. | 8.03B 8,030,261,248 | 16.06 GB | bf16 | no |
| Tokenizers | CLIP BPE ×2, T5 SentencePiece, Llama 3 BPE tokenizer/, tokenizer_2/, tokenizer_3/ in the HiDream repo; the Llama tokenizer comes from the Llama repo. | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL The FLUX.1 [schnell] autoencoder (scaling 0.3611, shift 0.1159). | 83.8M 83,819,683 | 168 MB | bf16 | no |
| total | 30.80B | 63.24 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- The transformer config sets max_resolution to 96×192 latents, so 48×96 = 4608 tokens: one 768×768 target and one 768×768 source side by side. The token counts below are for the target alone; the source doubles them.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 768×768 | 16×1×96×96 | 2,304 | native, per image |
| 512×512 | 16×1×64×64 | 1,024 |
Architecture
- Blocks
- 16 dual-stream + 32 single-stream
- Hidden size
- 2560 (20 heads × 128)
- Feed-forward
- Image tokens: MoE SwiGLU, 4 routed experts (top-2) of 6912 + 1 shared expert of 3584. Text tokens: dense SwiGLU 6912
- Text conditioning
- Same as HiDream-I1: T5 and last-layer Llama tokens in the sequence, plus a different Llama layer per block. 128 tokens max each
- Image conditioning
- Source latent concatenated along width with the noisy latent; no extra weights. Only the target half of the output is used
- Pooled conditioning
- CLIP-L + CLIP-G pooled (2048-d) added to the timestep embedding (adaLN)
- Objective
- Rectified flow, shift 3.0
- Norm / position
- QK RMSNorm, 3-axis RoPE (64/32/32)
In AI Toolkit
- model.arch
- hidream_e1
- UI label
- HiDream E1 (instruction)
- model.name_or_path
- HiDream-ai/HiDream-E1-1
- source
- extensions_built_in/diffusion_models/hidream/hidream_e1_model.py
- extensions_built_in/diffusion_models/hidream/hidream_model.py
- extensions_built_in/diffusion_models/hidream/src/pipelines/hidream_image/pipeline_hidream_image_editing.py
- toolkit/models/v2/diffusion_models/hidream.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.control_path, sample.ctrl_img, model.low_vram
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- lr
- 0.0001
- network_kwargs.ignore_if_contains
- ff_i.experts, ff_i.gate
- network.conv
- disabled (linear LoRA only)
Specifics
- Control images
- datasets.control_path supplies the source image. It is resized to the target size, VAE-encoded and concatenated along width, and the loss uses only the target half of the prediction. Samples need sample.ctrl_img.
- Resolution
- Buckets snap to multiples of 16. Keep each image at or under 768×768 of area: the target and source together must fit the 4608-token sequence.
- Model code
- Uses the Diffusers HiDreamImageTransformer2DModel with force_inference_output on, not the vendored class the hidream arch uses. Loading, text encoders, loss and saving are inherited from hidream.
- MoE
- The UI keeps LoRA off the routed experts and the router (ff_i.experts, ff_i.gate).
- Llama source
- unsloth/Meta-Llama-3.1-8B-Instruct unless model_kwargs.llama_model_path is set.
- Other text encoders and VAE
- Loaded from extras_name_or_path (defaults to name_or_path). None of the four are trained.
- Quantization
- quantize_te applies to T5 and Llama; the two CLIPs stay in the training dtype.
- Loss
- Target is noise − latents; the transformer output is negated to match.
- LoRA target
- HiDreamImageTransformer2DModel (double_stream_blocks, single_stream_blocks)
- Saving
- LoRA keys use the diffusion_model. prefix (ComfyUI style). Full fine-tunes save transformer/ in Diffusers format.
- Sampling
- HiDreamImageEditingPipeline with FlowUniPC multistep, shift 3.0.
- Metadata base version
- hidream_i1
Example config
not verified
job: extensionconfig: name: "my_hidream_e1_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 network_kwargs: ignore_if_contains: - "ff_i.experts" - "ff_i.gate" save: dtype: bfloat16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/edited/images" control_path: "/path/to/source/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768] train: batch_size: 1 steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "HiDream-ai/HiDream-E1-1" arch: "hidream_e1" quantize: true quantize_te: true sample: sampler: "flowmatch" sample_every: 250 width: 768 height: 768 guidance_scale: 5 sample_steps: 28 samples: - prompt: "make it snow" ctrl_img: "/path/to/source.jpg"Links
Qwen-Image-Edit-2511
'The December 2025 update of the multi-image Qwen-Image-Edit line. Same 20B MMDiT and Qwen2.5-VL-7B encoder as 2509, plus t = 0 modulation of the reference tokens. Qwen reports less image drift, better character and multi-person consistency, and popular community LoRA effects built in.'
Mage-Flow Edit Base
'The undistilled instruction-editing base of Microsoft’s Mage-Flow: the same 4.1B dual-stream NR-MMDiT and Mage-VAE as Mage-Flow Base, trained to edit. Reference images reach the model twice, through Qwen3-VL with the instruction and as clean latents packed after the target.'