HiDream-I1 Full
'A 17B sparse diffusion transformer with a mixture-of-experts feed-forward in every block, conditioned on four text encoders: CLIP-L, CLIP-G, T5-XXL and Llama 3.1 8B. Full is the undistilled base, the only HiDream-I1 variant meant for training.'
- weights
- HiDream-ai/HiDream-I1-Full ↗
- org
- HiDream.ai
- modality
- image
- tasks
- text-to-image
- license
- MIT
- released
- 2025-04-06not verified
- native output
- 1024×1024 (4096 image tokens max)
- total params
- 30.80B
- model.arch
- hidream
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | HiDream-I1 sparse DiT (MoE) HiDreamImageTransformer2DModel (vendored in ai-toolkit) Includes all 4 routed experts per block; only 2 run per token. ai-toolkit loads it in the training dtype (bf16 by default). | 17.11B 17,105,733,184 | 34.21 GB | fp16 | yes |
| Text encoder 1 | CLIP ViT-L/14 text (with projection) transformers.CLIPTextModelWithProjection Pooled output only (768-d). | 123.8M 123,781,632 | 495 MB | fp32 | no |
| Text encoder 2 | OpenCLIP ViT-bigG/14 text (with projection) transformers.CLIPTextModelWithProjection Pooled output only (1280-d). Concatenated with CLIP-L into the 2048-d pooled vector. | 694.8M 694,840,320 | 2.78 GB | fp32 | no |
| Text encoder 3 | T5-XXL v1.1 (encoder only) transformers.T5EncoderModel | 4.76B 4,762,310,656 | 9.52 GB | bf16 | no |
| Text encoder 4 | Llama 3.1 8B Instruct transformers.LlamaForCausalLM Not in the HiDream repo. Upstream code pulls the gated meta-llama/Meta-Llama-3.1-8B-Instruct; ai-toolkit defaults to the ungated unsloth mirror (override with model_kwargs.llama_model_path). Hidden states from all 32 layers are used, not just the last. | 8.03B 8,030,261,248 | 16.06 GB | bf16 | no |
| Tokenizers | CLIP BPE ×2, T5 SentencePiece, Llama 3 BPE tokenizer/, tokenizer_2/, tokenizer_3/ in the HiDream repo; the Llama tokenizer comes from the Llama repo. | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL The FLUX.1 [schnell] autoencoder (scaling 0.3611, shift 0.1159). | 83.8M 83,819,683 | 168 MB | bf16 | no |
| total | 30.80B | 63.24 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- The transformer config caps the latent at 128×128 (max_resolution), so 64×64 = 4096 image tokens, or 1024×1024 pixels of area. Non-square inputs are padded up to that 4096-token sequence and masked.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 768×1344 | 16×1×168×96 | 4,032 | portrait, same area |
| 512×512 | 16×1×64×64 | 1,024 |
Architecture
- Blocks
- 16 dual-stream + 32 single-stream
- Hidden size
- 2560 (20 heads × 128)
- Feed-forward
- Image tokens: MoE SwiGLU, 4 routed experts (top-2) of 6912 + 1 shared expert of 3584. Text tokens: dense SwiGLU 6912
- Text conditioning
- T5 tokens and last-layer Llama tokens are joined into the sequence; each block also gets a different Llama layer (layers 0–31, then 31 repeated), all projected from 4096-d. 128 tokens max each
- Pooled conditioning
- CLIP-L + CLIP-G pooled (2048-d) added to the timestep embedding (adaLN)
- Objective
- Rectified flow, shift 3.0
- Norm / position
- QK RMSNorm, 3-axis RoPE (64/32/32)
In AI Toolkit
- model.arch
- hidream
- UI label
- HiDream (image)
- model.name_or_path
- HiDream-ai/HiDream-I1-Full
- source
- extensions_built_in/diffusion_models/hidream/hidream_model.py
- extensions_built_in/diffusion_models/hidream/src/models/transformers/transformer_hidream_image.py
- extensions_built_in/diffusion_models/hidream/src/models/moe.py
- extensions_built_in/diffusion_models/hidream/src/pipelines/hidream_image/pipeline_hidream_image.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- shift
- lr
- 0.0002
- network_kwargs.ignore_if_contains
- ff_i.experts, ff_i.gate
- network.conv
- disabled (linear LoRA only)
- quantization option
- 3 bit with ARA (uint3 + accuracy recovery adapter)
Specifics
- Train Full only
- The Dev and Fast variants are distilled; the example config warns that training them breaks.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch). Above 1024×1024 of area the image no longer fits the 4096-token sequence.
- MoE
- The UI keeps LoRA off the routed experts and the router (ff_i.experts, ff_i.gate); the shared expert and attention still get LoRA. No auxiliary load-balancing loss is applied, so routing is not regularized during training.
- Llama source
- unsloth/Meta-Llama-3.1-8B-Instruct unless model_kwargs.llama_model_path is set.
- Other text encoders and VAE
- Loaded from extras_name_or_path (defaults to name_or_path). None of the four are trained.
- Quantization
- quantize_te applies to T5 and Llama; the two CLIPs stay in the training dtype. The UI also offers a 3-bit transformer with ostris/accuracy_recovery_adapters/hidream_i1_full_torchao_uint3.safetensors.
- Loss
- Target is noise − latents; the transformer output is negated to match.
- LoRA target
- HiDreamImageTransformer2DModel (double_stream_blocks, single_stream_blocks)
- Saving
- LoRA keys use the diffusion_model. prefix (ComfyUI style). Full fine-tunes save transformer/ in Diffusers format.
- Sampling
- FlowUniPC multistep, shift 3.0. low_vram unloads encoders aggressively during sampling.
- Metadata base version
- hidream_i1
Example config
not verified
job: extensionconfig: name: "my_hidream_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 network_kwargs: ignore_if_contains: - "ff_i.experts" - "ff_i.gate" save: dtype: bfloat16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "shift" optimizer: "adamw8bit" lr: 2e-4 dtype: bf16 model: name_or_path: "HiDream-ai/HiDream-I1-Full" arch: "hidream" quantize: true quantize_te: true model_kwargs: llama_model_path: "unsloth/Meta-Llama-3.1-8B-Instruct" sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Ming-Image 0.1 Design
A 6B Z-Image-style DiT for UI, infographics, posters and other text-heavy design, with RGBA output. It is conditioned by a Ling-mini-2.0 MoE multimodal LLM through two caption streams. ai-toolkit trains it from the int8 ComfyUI repack with a training adapter.
Stable Diffusion XL 1.0 Base
The 2.6B UNet latent diffusion model with two CLIP text encoders (CLIP-L and OpenCLIP bigG) and size/crop micro-conditioning. Still the base of a large ecosystem of fine-tunes, LoRAs and ControlNets.