Ideogram 4
Ideogram’s first open-weight model: a 9.3B single-stream DiT trained from scratch on structured JSON captions, with strong text rendering and bounding-box layout control. It reads 13 layers of Qwen3-VL-8B hidden states and works in the FLUX.2 VAE latent space. Upstream CFG uses a separate unconditional transformer; ai-toolkit swaps it for a LoRA.
- weights
- ideogram-ai/ideogram-4-fp8 ↗
- org
- Ideogram
- modality
- image
- tasks
- text-to-image
- license
- Ideogram 4 Non-Commercial
- released
- 2026-06-03
- native output
- 256 to 2048 px per side (multiples of 16), aspect ratios up to 6:1
- total params
- 17.51B
- model.arch
- ideogram4
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Ideogram 4 DiT (conditional) Ideogram4Transformer2DModel Weight-only fp8 with a per-output-channel fp32 scale beside each linear; the count includes the scales. ai-toolkit folds the scales back into bf16 on load, then quantizes again if quantize is on. | 9.28B 9,281,557,760 | 9.29 GB | fp8_e4m3+fp32+bf16 | yes |
| Transformer (unconditional) | Ideogram 4 DiT (unconditional branch)not loaded Ideogram4Transformer2DModel Same architecture, separate weights, used for the negative branch of dual-branch CFG. ai-toolkit does not load it; the unconditional LoRA below stands in for it. | 9.28B 9,281,557,760 | 9.29 GB | fp8_e4m3+fp32+bf16 | no |
| Text encoder | Qwen3-VL-8B-Instruct (fp8 copy) transformers.Qwen3VLModel Not loaded by ai-toolkit, which uses the public bf16 model below instead. | 8.15B 8,146,501,216 | 8.78 GB | fp8_e4m3+bf16+fp32 | no |
| Text encoder (loaded) | Qwen3-VL-8B-Instructalternate file transformers.Qwen3VLModel Stock, frozen. Loaded from this repo (with its tokenizer) because it is faster and more precise than dequantizing the fp8 copy. Override with model_kwargs.text_encoder_path. | 8.77B 8,767,123,696 | 17.53 GB | bf16 | no |
| Tokenizer | Qwen2 tokenizer (Qwen3-VL) transformers.Qwen2Tokenizer ai-toolkit uses the tokenizer from Qwen/Qwen3-VL-8B-Instruct. | — | — | — | — |
| VAE | FLUX.2 VAE diffusers.AutoencoderKLFlux2 Converted on load to the toolkit’s original-layout FLUX.2 autoencoder. The int64 tensor is the batch-norm step counter, which ai-toolkit does not use. | 84.0M 84,046,372 | 168 MB | bf16+int64 | no |
| Sampling adapter | Ideogram 4 unconditional LoRA (rank 16)adapter Initialized from the difference between the conditional and unconditional weights, then distilled against the unconditional model. Wraps every linear in the conditional DiT and is switched on only for the unconditional CFG pass. Never active during training. | 54.5M 54,456,320 | 109 MB | bf16 | no |
| total | 17.51B | 18.24 GB | |||
Latent space
- autoencoder
- FLUX.2 VAE
- pixels per token
- 16×16
- notes
- The VAE downsamples 8× to 32 channels. ai-toolkit keeps the encoder mean (no sampling), packs 2×2 patches into 128 channels at 16×, and normalizes with Ideogram’s fixed per-channel shift and scale (not the VAE’s batch-norm stats). The transformer takes the 128-channel tokens directly.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 32×1×128×128 | 4,096 | ai-toolkit sample default |
| 2048×2048 | 32×1×256×256 | 16,384 | maximum native size |
Architecture
- Blocks
- 34
- Hidden size
- 4608 (18 heads × 256)
- FFN size
- 12288
- Modulation
- AdaLN, 512-d conditioning
- Text conditioning
- Single stream. Hidden states from 13 Qwen3-VL layers (0, 3, … 33, 35) are concatenated (53248-d) and packed ahead of the image tokens
- Guidance
- Dual-branch CFG: the unconditional branch is image-only, with no text tokens
- Objective
- Rectified flow. Internally t = 1 is clean and the model predicts clean − noise
- Position
- Multimodal RoPE (sections 24/20/20, theta 5,000,000)
In AI Toolkit
- model.arch
- ideogram4
- UI label
- Ideogram4 (experimental)
- model.name_or_path
- ideogram-ai/ideogram-4-fp8
- source
- extensions_built_in/diffusion_models/ideogram4/ideogram4.py
- extensions_built_in/diffusion_models/ideogram4/src/transformer.py
- extensions_built_in/diffusion_models/ideogram4/src/pipeline.py
- extensions_built_in/diffusion_models/ideogram4/src/latent_norm.py
- toolkit/ideogram_caption.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- toolkit/models/v2/vae/flux2_kl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading, ideogram_4_prompt, model.unconditional_lora_path
UI defaults
- quantize / quantize_te
- true / true
- low_vram
- true
- timestep_type
- linear
- unconditional_lora_path
- ostris/ideogram_4_unconditional_lora/ideogram_4_unconditional_lora_r16.safetensors
- sample
- Ideogram JSON prompt preset, 1024×1024, guidance 4, 30 steps
- network.conv
- disabled (linear LoRA only)
Specifics
- Gated
- The weights are gated on Hugging Face. Accept the license and set HF_TOKEN.
- Resolution
- Buckets and sample sizes snap to multiples of 16 (8× VAE × 2×2 patch).
- Captions
- Trained on structured JSON captions. JSON prompts are normalized to the official schema (older formats are migrated) and serialized compactly before encoding; plain-text prompts pass through. The UI allows multi-line prompts for this.
- Prompt encoding
- Qwen3-VL chat template with a generation prompt. Each caption is encoded at its natural length, capped at 3072 tokens (model_kwargs.max_text_length), and padded to the batch maximum only at the model call.
- Unconditional pass
- The negative prompt is ignored: the unconditional branch has no text tokens. With unconditional_lora_path set, that LoRA is switched on for the unconditional pass only.
- Sampling schedule
- A resolution-aware logit-normal sigma schedule (std 1.75, mean shifted by image area), not the flowmatch shift. Tunable with model_kwargs.ideogram_schedule_mu / ideogram_schedule_std.
- Cache invalidation
- Cached text embeddings are keyed to ideogram4_te_v2, so embeddings from older ai-toolkit versions are re-encoded.
- LoRA target
- Ideogram4Transformer2DModel
- Saving
- LoRAs use the ComfyUI key prefix. Full fine-tunes save one dequantized .safetensors of the transformer in save.dtype.
- Metadata base version
- ideogram4
Example config
not verified
job: extensionconfig: name: "my_ideogram4_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "ideogram-ai/ideogram-4-fp8" arch: "ideogram4" quantize: true quantize_te: true low_vram: true unconditional_lora_path: "ostris/ideogram_4_unconditional_lora/ideogram_4_unconditional_lora_r16.safetensors" sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 30 prompts: - '{"high_level_description":"A red-haired woman playing chess in a park.","compositional_deconstruction":{"background":"An overcast city park.","elements":[{"type":"obj","desc":"Woman at a chess table, mid-move."}]}}'Links
Zeta-Chroma
'A work-in-progress pixel-space model from lodestones (the Chroma team), built on the Z-Image transformer. It has no VAE: 32×32 RGB patches go straight into the trunk, and a per-token MLP decoder predicts the clean pixels.'
Krea 2 Raw (edit training)
'Krea 2 Raw loaded for reference-image (edit) training. The weights are plain Krea 2 Raw text-to-image; ai-toolkit feeds control images to Qwen3-VL and appends them as clean latents so a LoRA learns to edit, restyle or follow a reference.'