Docs
AI ToolkitModels

HiDream-I1 Full

'A 17B sparse diffusion transformer with a mixture-of-experts feed-forward in every block, conditioned on four text encoders: CLIP-L, CLIP-G, T5-XXL and Llama 3.1 8B. Full is the undistilled base, the only HiDream-I1 variant meant for training.'

org
HiDream.ai
modality
image
tasks
text-to-image
license
MIT
released
2025-04-06not verified
native output
1024×1024 (4096 image tokens max)
total params
30.80B
model.arch
hidream

Components

rolemodelparamssizedtypetrained
Transformer
HiDream-I1 sparse DiT (MoE)
HiDreamImageTransformer2DModel (vendored in ai-toolkit)

Includes all 4 routed experts per block; only 2 run per token. ai-toolkit loads it in the training dtype (bf16 by default).

17.11B
17,105,733,184
34.21 GBfp16yes
Text encoder 1
CLIP ViT-L/14 text (with projection)
transformers.CLIPTextModelWithProjection

Pooled output only (768-d).

123.8M
123,781,632
495 MBfp32no
Text encoder 2
OpenCLIP ViT-bigG/14 text (with projection)
transformers.CLIPTextModelWithProjection

Pooled output only (1280-d). Concatenated with CLIP-L into the 2048-d pooled vector.

694.8M
694,840,320
2.78 GBfp32no
Text encoder 3
T5-XXL v1.1 (encoder only)
transformers.T5EncoderModel
4.76B
4,762,310,656
9.52 GBbf16no
Text encoder 4
Llama 3.1 8B Instruct
transformers.LlamaForCausalLM

Not in the HiDream repo. Upstream code pulls the gated meta-llama/Meta-Llama-3.1-8B-Instruct; ai-toolkit defaults to the ungated unsloth mirror (override with model_kwargs.llama_model_path). Hidden states from all 32 layers are used, not just the last.

8.03B
8,030,261,248
16.06 GBbf16no
Tokenizers
CLIP BPE ×2, T5 SentencePiece, Llama 3 BPE

tokenizer/, tokenizer_2/, tokenizer_3/ in the HiDream repo; the Llama tokenizer comes from the Llama repo.

————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

The FLUX.1 [schnell] autoencoder (scaling 0.3611, shift 0.1159).

83.8M
83,819,683
168 MBbf16no
total30.80B63.24 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
The transformer config caps the latent at 128×128 (max_resolution), so 64×64 = 4096 image tokens, or 1024×1024 pixels of area. Non-square inputs are padded up to that 4096-token sequence and masked.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
768×134416×1×168×964,032portrait, same area
512×51216×1×64×641,024

Architecture

Blocks
16 dual-stream + 32 single-stream
Hidden size
2560 (20 heads × 128)
Feed-forward
Image tokens: MoE SwiGLU, 4 routed experts (top-2) of 6912 + 1 shared expert of 3584. Text tokens: dense SwiGLU 6912
Text conditioning
T5 tokens and last-layer Llama tokens are joined into the sequence; each block also gets a different Llama layer (layers 0–31, then 31 repeated), all projected from 4096-d. 128 tokens max each
Pooled conditioning
CLIP-L + CLIP-G pooled (2048-d) added to the timestep embedding (adaLN)
Objective
Rectified flow, shift 3.0
Norm / position
QK RMSNorm, 3-axis RoPE (64/32/32)

In AI Toolkit

model.arch
hidream
UI label
HiDream (image)
model.name_or_path
HiDream-ai/HiDream-I1-Full
extra UI sections
model.low_vram

UI defaults

quantize / quantize_te
true / true (qfloat8)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
shift
lr
0.0002
network_kwargs.ignore_if_contains
ff_i.experts, ff_i.gate
network.conv
disabled (linear LoRA only)
quantization option
3 bit with ARA (uint3 + accuracy recovery adapter)

Specifics

Train Full only
The Dev and Fast variants are distilled; the example config warns that training them breaks.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch). Above 1024×1024 of area the image no longer fits the 4096-token sequence.
MoE
The UI keeps LoRA off the routed experts and the router (ff_i.experts, ff_i.gate); the shared expert and attention still get LoRA. No auxiliary load-balancing loss is applied, so routing is not regularized during training.
Llama source
unsloth/Meta-Llama-3.1-8B-Instruct unless model_kwargs.llama_model_path is set.
Other text encoders and VAE
Loaded from extras_name_or_path (defaults to name_or_path). None of the four are trained.
Quantization
quantize_te applies to T5 and Llama; the two CLIPs stay in the training dtype. The UI also offers a 3-bit transformer with ostris/accuracy_recovery_adapters/hidream_i1_full_torchao_uint3.safetensors.
Loss
Target is noise − latents; the transformer output is negated to match.
LoRA target
HiDreamImageTransformer2DModel (double_stream_blocks, single_stream_blocks)
Saving
LoRA keys use the diffusion_model. prefix (ComfyUI style). Full fine-tunes save transformer/ in Diffusers format.
Sampling
FlowUniPC multistep, shift 3.0. low_vram unloads encoders aggressively during sampling.
Metadata base version
hidream_i1

Example config

not verified

job: extensionconfig:  name: "my_hidream_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32        network_kwargs:          ignore_if_contains:            - "ff_i.experts"            - "ff_i.gate"      save:        dtype: bfloat16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "shift"        optimizer: "adamw8bit"        lr: 2e-4        dtype: bf16      model:        name_or_path: "HiDream-ai/HiDream-I1-Full"        arch: "hidream"        quantize: true        quantize_te: true        model_kwargs:          llama_model_path: "unsloth/Meta-Llama-3.1-8B-Instruct"      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page