Docs
AI ToolkitModels

HiDream-O1-Image

An 8B pixel-space image model built inside a Qwen3-VL language model. Text, timestep and 32×32 pixel patches share one token sequence in the same transformer, with no VAE and no separate text encoder. Generates up to 2048×2048.

org
HiDream.ai
modality
image
tasks
text-to-image · image editing · subject-driven generation
license
MIT
released
2026-05-08not verified
native output
up to 2048×2048
total params
8.80B
model.arch
hidream_o1

Components

rolemodelparamssizedtypetrained
Unified transformer
Qwen3-VL 8B with diffusion heads
Qwen3VLForConditionalGeneration (vendored in ai-toolkit)

One checkpoint holds everything. From the safetensors headers: 36 decoder layers 6.95B, token embeddings 622M, lm_head 622M (unused for images), Qwen3-VL vision tower 576M (not used by ai-toolkit, which trains text-to-image only), and the image heads: x_embedder 7.3M, t_embedder1 17.8M, final_layer2 12.6M. About 17.6 GB in bf16.

8.80B
8,804,887,792
35.22 GBfp32yes
Processor
Qwen3-VL processor (151,936-token vocab)
transformers.AutoProcessor

ai-toolkit adds the boi, bor, eor, bot and tms special-token shortcuts the pipeline uses.

————
total8.80B35.22 GB

Latent space

spatial
1×
channels
3
patch
32×32
pixels per token
32×32
notes
No autoencoder: the model works on RGB pixels. Each 32×32×3 patch (3072 values) goes through a bottleneck embed (3072 → 1024 → 4096) to become one token, so a 2048×2048 image is 4096 tokens.
inputlatent (c×t×h×w)tokens
2048×20483×1×2048×20484,096ai-toolkit sample default
1024×10243×1×1024×10241,024
2560×14403×1×1440×25603,600pipeline default

Architecture

Layers
36 Qwen3 decoder layers
Hidden size
4096 (32 query heads × 128, 8 KV heads)
FFN size
12288
Text conditioning
Chat-templated prompt tokens in the same sequence. Text attends causally; image tokens attend to everything
Timestep
Embedded by t_embedder1 and written into a <|tms_token|> slot after the prompt
Output
final_layer2 predicts clean pixels (x0) per patch. The noise is scaled by 8.0 (noise_scale)
Position
Interleaved 3D M-RoPE (sections 24/20/20), theta 5M
Vision tower
27-layer ViT (1152-d, patch 16), for image inputs in editing and personalizationnot verified

In AI Toolkit

model.arch
hidream_o1
UI label
HiDream-O1 (image)
model.name_or_path
HiDream-ai/HiDream-O1-Image
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

quantize
true (qfloat8)
timestep_type
weighted
max_loss
1.0
network_kwargs.ignore_if_contains
lm_head, patch_embed, visual
network.transformer_only
false
model_kwargs
noise_scale: 8.0, noise_scale_inference: 8.0
sample size
2048×2048
network.conv
disabled (linear LoRA only)
quantize_te / unload_text_encoder
hidden (there is no separate text encoder)

Specifics

Pixel space
A stand-in VAE passes pixels through unchanged, so latent caching stores images. Buckets snap to multiples of 32 (the patch size).
Noise
Training noise is multiplied by noise_scale (8.0 by default) before interpolation: x_t = (1 − t)·x0 + t·8·ε. Sampling uses noise_scale_inference.
Loss
The model predicts x0 and the target is the clean image. The UI clamps the loss at 1.0 (max_loss) to stop outliers.
Prompts
Prompts are tokenized, not encoded: the chat template plus <|boi_token|><|tms_token|> is stored as token ids and run through the same transformer on every step.
LoRA scope
transformer_only is off so the image heads (x_embedder, t_embedder1, final_layer2) get LoRA too; lm_head and the vision tower are excluded.
ComfyUI weights
name_or_path may be a single .safetensors (ComfyUI layout). The missing lm_head is filled with zeros and the processor comes from HiDream-ai/HiDream-O1-Image.
LoRA target
HidreamO1Transformer (model.language_model.layers)
Saving
LoRA keys become diffusion_model.* (ComfyUI style). Full fine-tunes save a Transformers folder plus the processor, or a single file without lm_head when loaded from ComfyUI weights.
Sampling
Flowmatch Euler, shift 3.0, with the scaled noise.
Metadata base version
hidream_o1

Example config

not verified

job: extensionconfig:  name: "my_hidream_o1_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32        transformer_only: false        network_kwargs:          ignore_if_contains:            - "lm_head"            - "patch_embed"            - "visual"      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [1024, 1536, 2048]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        max_loss: 1.0        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "HiDream-ai/HiDream-O1-Image"        arch: "hidream_o1"        quantize: true        low_vram: true        model_kwargs:          noise_scale: 8.0          noise_scale_inference: 8.0      sample:        sampler: "flowmatch"        sample_every: 250        width: 2048        height: 2048        guidance_scale: 5        sample_steps: 28        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page