OmniGen2
A 4B Lumina-style diffusion transformer conditioned on Qwen2.5-VL 3B. Reference images enter as extra latent tokens through their own refiner, so one model does text-to-image, instruction editing and subject-driven generation.
- weights
- OmniGen2/OmniGen2 ↗
- org
- VectorSpaceLab
- modality
- image
- tasks
- text-to-image · image editing · in-context generation
- license
- Apache 2.0
- released
- 2025-06-06not verified
- native output
- 1024×1024not verified
- total params
- 7.81B
- model.arch
- omnigen2
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | OmniGen2 DiT OmniGen2Transformer2DModel (vendored in ai-toolkit) Stored as fp32, about 7.9 GB in bf16. | 3.97B 3,967,161,400 | 15.87 GB | fp32 | yes |
| Text encoder (MLLM) | Qwen2.5-VL 3B Instruct transformers.Qwen2_5_VLForConditionalGeneration From the safetensors headers: 36 LM layers 2.77B, token embeddings 311M (tied, no separate LM head), vision tower 669M. ai-toolkit loads the vision tower but only encodes text. | 3.75B 3,754,622,976 | 15.02 GB | fp32 | no |
| Processor | Qwen2 tokenizer + image processor CLIPProcessor (loaded from processor/) The repo has a second copy with the same file sizes in mllm_processor/; ai-toolkit uses processor/. | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL Same config as the FLUX.1 autoencoder (scaling 0.3611, shift 0.1159). | 83.8M 83,819,683 | 335 MB | fp32 | no |
| total | 7.81B | 31.22 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- Reference images use the same VAE and patch size, through a separate patch embedder.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 768×1344 | 16×1×168×96 | 4,032 | portrait, same area |
| 512×512 | 16×1×64×64 | 1,024 |
Architecture
- Blocks
- 32 single-stream, plus 2 each of noise-refiner, reference-image-refiner and context-refiner layers
- Hidden size
- 2520 (21 query heads × 120, 7 KV heads)
- FFN size
- 10240 (SwiGLU)
- Text conditioning
- Last-layer Qwen2.5-VL hidden states (2048-d) of a chat-templated prompt, 256 tokens max
- Image conditioning
- Reference latents get their own patch embedder, a learned index embedding (up to 5 images) and 2 refiner layers
- Objective
- Rectified flow. Time runs from 0 = noise to 1 = image
- Position
- 3-axis RoPE (40/40/40), axis lengths 1024 / 1664 / 1664
In AI Toolkit
- model.arch
- omnigen2
- UI label
- OmniGen2 (image)
- model.name_or_path
- OmniGen2/OmniGen2
- source
- extensions_built_in/diffusion_models/omnigen2/__init__.py
- extensions_built_in/diffusion_models/omnigen2/src/models/transformers/transformer_omnigen2.py
- extensions_built_in/diffusion_models/omnigen2/src/pipelines/omnigen2/pipeline_omnigen2.py
- toolkit/models/v2/text_encoders/qwen25_vl.py
- config/examples/train_lora_omnigen2_24gb.yaml
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.control_path, sample.ctrl_img
UI defaults
- quantize / quantize_te
- false / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- network.conv
- disabled (linear LoRA only)
Specifics
- Control images
- Optional. With datasets.control_path set, the control image is resized to the target, VAE-encoded and passed as one reference image. Without it, training is plain text-to-image.
- Reference refiner
- LoRA skips ref_image_refiner unless model_kwargs.use_image_refiner is true; noise_refiner, context_refiner and layers are trained.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Prompts
- Wrapped in the Qwen chat template with the system prompt "You are a helpful assistant that generates high-quality images based on user instructions."
- Timesteps and loss
- Training uses a flowmatch scheduler with no shift. The transformer gets 1 − t, and the target is latents − noise (the output is not negated).
- Text encoder and VAE source
- mllm/, processor/ and vae/ from extras_name_or_path (defaults to name_or_path). Never trained.
- LoRA target
- OmniGen2Transformer2DModel
- Saving
- LoRA keys use the diffusion_model. prefix (ComfyUI style). Full fine-tunes save transformer/ in Diffusers format.
- Sampling
- OmniGen2's own flow-match Euler with dynamic time shift. Reference-image guidance is fixed at 1.0.
- Metadata base version
- omnigen2
Example config
not verified
job: extensionconfig: name: "my_omnigen2_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "sigmoid" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "OmniGen2/OmniGen2" arch: "omnigen2" quantize_te: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Stable Diffusion 1.5
The original 860M UNet latent diffusion model with a single CLIP ViT-L text encoder, trained at 512×512. Small and fast to train, with the largest library of community fine-tunes of any model.
FLUX.2 [dev]
The 32B, guidance-distilled FLUX.2 model. A new double/single-stream transformer conditioned on Mistral Small 3.1 (24B) hidden states, working in the 32-channel latent space of the new FLUX.2 VAE. Reference images go in as extra tokens, so one model generates and edits.