Docs
AI ToolkitModels

Ideogram 4

Ideogram’s first open-weight model: a 9.3B single-stream DiT trained from scratch on structured JSON captions, with strong text rendering and bounding-box layout control. It reads 13 layers of Qwen3-VL-8B hidden states and works in the FLUX.2 VAE latent space. Upstream CFG uses a separate unconditional transformer; ai-toolkit swaps it for a LoRA.

org
Ideogram
modality
image
tasks
text-to-image
license
Ideogram 4 Non-Commercial
released
2026-06-03
native output
256 to 2048 px per side (multiples of 16), aspect ratios up to 6:1
total params
17.51B
model.arch
ideogram4

Components

rolemodelparamssizedtypetrained
Transformer
Ideogram 4 DiT (conditional)
Ideogram4Transformer2DModel

Weight-only fp8 with a per-output-channel fp32 scale beside each linear; the count includes the scales. ai-toolkit folds the scales back into bf16 on load, then quantizes again if quantize is on.

9.28B
9,281,557,760
9.29 GBfp8_e4m3+fp32+bf16yes
Transformer (unconditional)
Ideogram 4 DiT (unconditional branch)not loaded
Ideogram4Transformer2DModel

Same architecture, separate weights, used for the negative branch of dual-branch CFG. ai-toolkit does not load it; the unconditional LoRA below stands in for it.

9.28B
9,281,557,760
9.29 GBfp8_e4m3+fp32+bf16no
Text encoder
Qwen3-VL-8B-Instruct (fp8 copy)
transformers.Qwen3VLModel

Not loaded by ai-toolkit, which uses the public bf16 model below instead.

8.15B
8,146,501,216
8.78 GBfp8_e4m3+bf16+fp32no
Text encoder (loaded)
Qwen3-VL-8B-Instructalternate file
transformers.Qwen3VLModel

Stock, frozen. Loaded from this repo (with its tokenizer) because it is faster and more precise than dequantizing the fp8 copy. Override with model_kwargs.text_encoder_path.

8.77B
8,767,123,696
17.53 GBbf16no
Tokenizer
Qwen2 tokenizer (Qwen3-VL)
transformers.Qwen2Tokenizer

ai-toolkit uses the tokenizer from Qwen/Qwen3-VL-8B-Instruct.

————
VAE
FLUX.2 VAE
diffusers.AutoencoderKLFlux2

Converted on load to the toolkit’s original-layout FLUX.2 autoencoder. The int64 tensor is the batch-norm step counter, which ai-toolkit does not use.

84.0M
84,046,372
168 MBbf16+int64no
Sampling adapter
Ideogram 4 unconditional LoRA (rank 16)adapter

Initialized from the difference between the conditional and unconditional weights, then distilled against the unconditional model. Wraps every linear in the conditional DiT and is switched on only for the unconditional CFG pass. Never active during training.

54.5M
54,456,320
109 MBbf16no
total17.51B18.24 GB

Latent space

spatial
8×
channels
32
patch
2×2
autoencoder
FLUX.2 VAE
pixels per token
16×16
notes
The VAE downsamples 8× to 32 channels. ai-toolkit keeps the encoder mean (no sampling), packs 2×2 patches into 128 channels at 16×, and normalizes with Ideogram’s fixed per-channel shift and scale (not the VAE’s batch-norm stats). The transformer takes the 128-channel tokens directly.
inputlatent (c×t×h×w)tokens
1024×102432×1×128×1284,096ai-toolkit sample default
2048×204832×1×256×25616,384maximum native size

Architecture

Blocks
34
Hidden size
4608 (18 heads × 256)
FFN size
12288
Modulation
AdaLN, 512-d conditioning
Text conditioning
Single stream. Hidden states from 13 Qwen3-VL layers (0, 3, … 33, 35) are concatenated (53248-d) and packed ahead of the image tokens
Guidance
Dual-branch CFG: the unconditional branch is image-only, with no text tokens
Objective
Rectified flow. Internally t = 1 is clean and the model predicts clean − noise
Position
Multimodal RoPE (sections 24/20/20, theta 5,000,000)

In AI Toolkit

model.arch
ideogram4
UI label
Ideogram4 (experimental)
model.name_or_path
ideogram-ai/ideogram-4-fp8
extra UI sections
model.low_vram, model.layer_offloading, ideogram_4_prompt, model.unconditional_lora_path

UI defaults

quantize / quantize_te
true / true
low_vram
true
timestep_type
linear
unconditional_lora_path
ostris/ideogram_4_unconditional_lora/ideogram_4_unconditional_lora_r16.safetensors
sample
Ideogram JSON prompt preset, 1024×1024, guidance 4, 30 steps
network.conv
disabled (linear LoRA only)

Specifics

Gated
The weights are gated on Hugging Face. Accept the license and set HF_TOKEN.
Resolution
Buckets and sample sizes snap to multiples of 16 (8× VAE × 2×2 patch).
Captions
Trained on structured JSON captions. JSON prompts are normalized to the official schema (older formats are migrated) and serialized compactly before encoding; plain-text prompts pass through. The UI allows multi-line prompts for this.
Prompt encoding
Qwen3-VL chat template with a generation prompt. Each caption is encoded at its natural length, capped at 3072 tokens (model_kwargs.max_text_length), and padded to the batch maximum only at the model call.
Unconditional pass
The negative prompt is ignored: the unconditional branch has no text tokens. With unconditional_lora_path set, that LoRA is switched on for the unconditional pass only.
Sampling schedule
A resolution-aware logit-normal sigma schedule (std 1.75, mean shifted by image area), not the flowmatch shift. Tunable with model_kwargs.ideogram_schedule_mu / ideogram_schedule_std.
Cache invalidation
Cached text embeddings are keyed to ideogram4_te_v2, so embeddings from older ai-toolkit versions are re-encoded.
LoRA target
Ideogram4Transformer2DModel
Saving
LoRAs use the ComfyUI key prefix. Full fine-tunes save one dequantized .safetensors of the transformer in save.dtype.
Metadata base version
ideogram4

Example config

not verified

job: extensionconfig:  name: "my_ideogram4_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "ideogram-ai/ideogram-4-fp8"        arch: "ideogram4"        quantize: true        quantize_te: true        low_vram: true        unconditional_lora_path: "ostris/ideogram_4_unconditional_lora/ideogram_4_unconditional_lora_r16.safetensors"      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 30        prompts:          - '{"high_level_description":"A red-haired woman playing chess in a park.","compositional_deconstruction":{"background":"An overcast city park.","elements":[{"type":"obj","desc":"Woman at a chess table, mid-move."}]}}'

On this page