HiDream-O1-Image
An 8B pixel-space image model built inside a Qwen3-VL language model. Text, timestep and 32×32 pixel patches share one token sequence in the same transformer, with no VAE and no separate text encoder. Generates up to 2048×2048.
- org
- HiDream.ai
- modality
- image
- tasks
- text-to-image · image editing · subject-driven generation
- license
- MIT
- released
- 2026-05-08not verified
- native output
- up to 2048×2048
- total params
- 8.80B
- model.arch
- hidream_o1
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Unified transformer | Qwen3-VL 8B with diffusion heads Qwen3VLForConditionalGeneration (vendored in ai-toolkit) One checkpoint holds everything. From the safetensors headers: 36 decoder layers 6.95B, token embeddings 622M, lm_head 622M (unused for images), Qwen3-VL vision tower 576M (not used by ai-toolkit, which trains text-to-image only), and the image heads: x_embedder 7.3M, t_embedder1 17.8M, final_layer2 12.6M. About 17.6 GB in bf16. | 8.80B 8,804,887,792 | 35.22 GB | fp32 | yes |
| Processor | Qwen3-VL processor (151,936-token vocab) transformers.AutoProcessor ai-toolkit adds the boi, bor, eor, bot and tms special-token shortcuts the pipeline uses. | — | — | — | — |
| total | 8.80B | 35.22 GB | |||
Latent space
- pixels per token
- 32×32
- notes
- No autoencoder: the model works on RGB pixels. Each 32×32×3 patch (3072 values) goes through a bottleneck embed (3072 → 1024 → 4096) to become one token, so a 2048×2048 image is 4096 tokens.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 2048×2048 | 3×1×2048×2048 | 4,096 | ai-toolkit sample default |
| 1024×1024 | 3×1×1024×1024 | 1,024 | |
| 2560×1440 | 3×1×1440×2560 | 3,600 | pipeline default |
Architecture
- Layers
- 36 Qwen3 decoder layers
- Hidden size
- 4096 (32 query heads × 128, 8 KV heads)
- FFN size
- 12288
- Text conditioning
- Chat-templated prompt tokens in the same sequence. Text attends causally; image tokens attend to everything
- Timestep
- Embedded by t_embedder1 and written into a <|tms_token|> slot after the prompt
- Output
- final_layer2 predicts clean pixels (x0) per patch. The noise is scaled by 8.0 (noise_scale)
- Position
- Interleaved 3D M-RoPE (sections 24/20/20), theta 5M
- Vision tower
- 27-layer ViT (1152-d, patch 16), for image inputs in editing and personalizationnot verified
In AI Toolkit
- model.arch
- hidream_o1
- UI label
- HiDream-O1 (image)
- model.name_or_path
- HiDream-ai/HiDream-O1-Image
- source
- extensions_built_in/diffusion_models/hidream/hidream_o1_model.py
- extensions_built_in/diffusion_models/hidream/src/hidream_o1/pipeline.py
- extensions_built_in/diffusion_models/hidream/src/hidream_o1/qwen3_vl_transformers.py
- extensions_built_in/diffusion_models/hidream/src/hidream_o1/model_config.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize
- true (qfloat8)
- timestep_type
- weighted
- max_loss
- 1.0
- network_kwargs.ignore_if_contains
- lm_head, patch_embed, visual
- network.transformer_only
- false
- model_kwargs
- noise_scale: 8.0, noise_scale_inference: 8.0
- sample size
- 2048×2048
- network.conv
- disabled (linear LoRA only)
- quantize_te / unload_text_encoder
- hidden (there is no separate text encoder)
Specifics
- Pixel space
- A stand-in VAE passes pixels through unchanged, so latent caching stores images. Buckets snap to multiples of 32 (the patch size).
- Noise
- Training noise is multiplied by noise_scale (8.0 by default) before interpolation: x_t = (1 − t)·x0 + t·8·ε. Sampling uses noise_scale_inference.
- Loss
- The model predicts x0 and the target is the clean image. The UI clamps the loss at 1.0 (max_loss) to stop outliers.
- Prompts
- Prompts are tokenized, not encoded: the chat template plus <|boi_token|><|tms_token|> is stored as token ids and run through the same transformer on every step.
- LoRA scope
- transformer_only is off so the image heads (x_embedder, t_embedder1, final_layer2) get LoRA too; lm_head and the vision tower are excluded.
- ComfyUI weights
- name_or_path may be a single .safetensors (ComfyUI layout). The missing lm_head is filled with zeros and the processor comes from HiDream-ai/HiDream-O1-Image.
- LoRA target
- HidreamO1Transformer (model.language_model.layers)
- Saving
- LoRA keys become diffusion_model.* (ComfyUI style). Full fine-tunes save a Transformers folder plus the processor, or a single file without lm_head when loaded from ComfyUI weights.
- Sampling
- Flowmatch Euler, shift 3.0, with the scaled noise.
- Metadata base version
- hidream_o1
Example config
not verified
job: extensionconfig: name: "my_hidream_o1_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 transformer_only: false network_kwargs: ignore_if_contains: - "lm_head" - "patch_embed" - "visual" save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [1024, 1536, 2048] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" max_loss: 1.0 optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "HiDream-ai/HiDream-O1-Image" arch: "hidream_o1" quantize: true low_vram: true model_kwargs: noise_scale: 8.0 noise_scale_inference: 8.0 sample: sampler: "flowmatch" sample_every: 250 width: 2048 height: 2048 guidance_scale: 5 sample_steps: 28 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Nucleus-Image
A 17B sparse mixture-of-experts DiT that activates about 2B parameters per step. Text from Qwen3-VL-8B enters only as keys and values, and images use the Qwen-Image VAE. Released as a pre-trained base model with no preference tuning.
Z-Image L2P (pixel space)
'A pixel-space version of Z-Image made with the Latent-to-Pixel (L2P) transfer method: the VAE is gone, 16×16 RGB patches go straight into the Z-Image trunk, and a small U-Net decoder turns the transformer features back into pixels.'