Docs
AI ToolkitModels

HiDream-E1.1

'The instruction-based image editing model built on HiDream-I1: the same 17B mixture-of-experts transformer and four text encoders, fine-tuned to take a source image placed beside the target in the latent and edit it from a text instruction.'

org
HiDream.ai
modality
image
tasks
image editing
license
MIT
released
2025-07-16not verified
native output
768×768 output beside a 768×768 source
total params
30.80B
model.arch
hidream_e1

Components

rolemodelparamssizedtypetrained
Transformer
HiDream-E1.1 sparse DiT (MoE)
diffusers.HiDreamImageTransformer2DModel

Same shape and parameter count as HiDream-I1; only max_resolution differs (96×192 latents instead of 128×128). Includes all 4 routed experts per block; only 2 run per token.

17.11B
17,105,733,184
34.21 GBbf16yes
Text encoder 1
CLIP ViT-L/14 text (with projection)
transformers.CLIPTextModelWithProjection

Pooled output only (768-d).

123.8M
123,781,632
495 MBfp32no
Text encoder 2
OpenCLIP ViT-bigG/14 text (with projection)
transformers.CLIPTextModelWithProjection

Pooled output only (1280-d). Concatenated with CLIP-L into the 2048-d pooled vector.

694.8M
694,840,320
2.78 GBfp32no
Text encoder 3
T5-XXL v1.1 (encoder only)
transformers.T5EncoderModel
4.76B
4,762,310,656
9.52 GBbf16no
Text encoder 4
Llama 3.1 8B Instruct
transformers.LlamaForCausalLM

Not in the HiDream repo. Upstream code pulls the gated meta-llama/Meta-Llama-3.1-8B-Instruct; ai-toolkit defaults to the ungated unsloth mirror (override with model_kwargs.llama_model_path). Hidden states from all 32 layers are used, not just the last.

8.03B
8,030,261,248
16.06 GBbf16no
Tokenizers
CLIP BPE ×2, T5 SentencePiece, Llama 3 BPE

tokenizer/, tokenizer_2/, tokenizer_3/ in the HiDream repo; the Llama tokenizer comes from the Llama repo.

————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

The FLUX.1 [schnell] autoencoder (scaling 0.3611, shift 0.1159).

83.8M
83,819,683
168 MBbf16no
total30.80B63.24 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
The transformer config sets max_resolution to 96×192 latents, so 48×96 = 4608 tokens: one 768×768 target and one 768×768 source side by side. The token counts below are for the target alone; the source doubles them.
inputlatent (c×t×h×w)tokens
768×76816×1×96×962,304native, per image
512×51216×1×64×641,024

Architecture

Blocks
16 dual-stream + 32 single-stream
Hidden size
2560 (20 heads × 128)
Feed-forward
Image tokens: MoE SwiGLU, 4 routed experts (top-2) of 6912 + 1 shared expert of 3584. Text tokens: dense SwiGLU 6912
Text conditioning
Same as HiDream-I1: T5 and last-layer Llama tokens in the sequence, plus a different Llama layer per block. 128 tokens max each
Image conditioning
Source latent concatenated along width with the noisy latent; no extra weights. Only the target half of the output is used
Pooled conditioning
CLIP-L + CLIP-G pooled (2048-d) added to the timestep embedding (adaLN)
Objective
Rectified flow, shift 3.0
Norm / position
QK RMSNorm, 3-axis RoPE (64/32/32)

In AI Toolkit

model.arch
hidream_e1
UI label
HiDream E1 (instruction)
model.name_or_path
HiDream-ai/HiDream-E1-1
extra UI sections
datasets.control_path, sample.ctrl_img, model.low_vram

UI defaults

quantize / quantize_te
true / true (qfloat8)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
lr
0.0001
network_kwargs.ignore_if_contains
ff_i.experts, ff_i.gate
network.conv
disabled (linear LoRA only)

Specifics

Control images
datasets.control_path supplies the source image. It is resized to the target size, VAE-encoded and concatenated along width, and the loss uses only the target half of the prediction. Samples need sample.ctrl_img.
Resolution
Buckets snap to multiples of 16. Keep each image at or under 768×768 of area: the target and source together must fit the 4608-token sequence.
Model code
Uses the Diffusers HiDreamImageTransformer2DModel with force_inference_output on, not the vendored class the hidream arch uses. Loading, text encoders, loss and saving are inherited from hidream.
MoE
The UI keeps LoRA off the routed experts and the router (ff_i.experts, ff_i.gate).
Llama source
unsloth/Meta-Llama-3.1-8B-Instruct unless model_kwargs.llama_model_path is set.
Other text encoders and VAE
Loaded from extras_name_or_path (defaults to name_or_path). None of the four are trained.
Quantization
quantize_te applies to T5 and Llama; the two CLIPs stay in the training dtype.
Loss
Target is noise − latents; the transformer output is negated to match.
LoRA target
HiDreamImageTransformer2DModel (double_stream_blocks, single_stream_blocks)
Saving
LoRA keys use the diffusion_model. prefix (ComfyUI style). Full fine-tunes save transformer/ in Diffusers format.
Sampling
HiDreamImageEditingPipeline with FlowUniPC multistep, shift 3.0.
Metadata base version
hidream_i1

Example config

not verified

job: extensionconfig:  name: "my_hidream_e1_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32        network_kwargs:          ignore_if_contains:            - "ff_i.experts"            - "ff_i.gate"      save:        dtype: bfloat16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/edited/images"          control_path: "/path/to/source/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768]      train:        batch_size: 1        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "HiDream-ai/HiDream-E1-1"        arch: "hidream_e1"        quantize: true        quantize_te: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 768        height: 768        guidance_scale: 5        sample_steps: 28        samples:          - prompt: "make it snow"            ctrl_img: "/path/to/source.jpg"

On this page