Docs
AI ToolkitModels

OmniGen2

A 4B Lumina-style diffusion transformer conditioned on Qwen2.5-VL 3B. Reference images enter as extra latent tokens through their own refiner, so one model does text-to-image, instruction editing and subject-driven generation.

org
VectorSpaceLab
modality
image
tasks
text-to-image · image editing · in-context generation
license
Apache 2.0
released
2025-06-06not verified
native output
1024×1024not verified
total params
7.81B
model.arch
omnigen2

Components

rolemodelparamssizedtypetrained
Transformer
OmniGen2 DiT
OmniGen2Transformer2DModel (vendored in ai-toolkit)

Stored as fp32, about 7.9 GB in bf16.

3.97B
3,967,161,400
15.87 GBfp32yes
Text encoder (MLLM)
Qwen2.5-VL 3B Instruct
transformers.Qwen2_5_VLForConditionalGeneration

From the safetensors headers: 36 LM layers 2.77B, token embeddings 311M (tied, no separate LM head), vision tower 669M. ai-toolkit loads the vision tower but only encodes text.

3.75B
3,754,622,976
15.02 GBfp32no
Processor
Qwen2 tokenizer + image processor
CLIPProcessor (loaded from processor/)

The repo has a second copy with the same file sizes in mllm_processor/; ai-toolkit uses processor/.

————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

Same config as the FLUX.1 autoencoder (scaling 0.3611, shift 0.1159).

83.8M
83,819,683
335 MBfp32no
total7.81B31.22 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
Reference images use the same VAE and patch size, through a separate patch embedder.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
768×134416×1×168×964,032portrait, same area
512×51216×1×64×641,024

Architecture

Blocks
32 single-stream, plus 2 each of noise-refiner, reference-image-refiner and context-refiner layers
Hidden size
2520 (21 query heads × 120, 7 KV heads)
FFN size
10240 (SwiGLU)
Text conditioning
Last-layer Qwen2.5-VL hidden states (2048-d) of a chat-templated prompt, 256 tokens max
Image conditioning
Reference latents get their own patch embedder, a learned index embedding (up to 5 images) and 2 refiner layers
Objective
Rectified flow. Time runs from 0 = noise to 1 = image
Position
3-axis RoPE (40/40/40), axis lengths 1024 / 1664 / 1664

In AI Toolkit

model.arch
omnigen2
UI label
OmniGen2 (image)
model.name_or_path
OmniGen2/OmniGen2
extra UI sections
datasets.control_path, sample.ctrl_img

UI defaults

quantize / quantize_te
false / true (qfloat8)
noise_scheduler / sampler
flowmatch / flowmatch
network.conv
disabled (linear LoRA only)

Specifics

Control images
Optional. With datasets.control_path set, the control image is resized to the target, VAE-encoded and passed as one reference image. Without it, training is plain text-to-image.
Reference refiner
LoRA skips ref_image_refiner unless model_kwargs.use_image_refiner is true; noise_refiner, context_refiner and layers are trained.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Prompts
Wrapped in the Qwen chat template with the system prompt "You are a helpful assistant that generates high-quality images based on user instructions."
Timesteps and loss
Training uses a flowmatch scheduler with no shift. The transformer gets 1 − t, and the target is latents − noise (the output is not negated).
Text encoder and VAE source
mllm/, processor/ and vae/ from extras_name_or_path (defaults to name_or_path). Never trained.
LoRA target
OmniGen2Transformer2DModel
Saving
LoRA keys use the diffusion_model. prefix (ComfyUI style). Full fine-tunes save transformer/ in Diffusers format.
Sampling
OmniGen2's own flow-match Euler with dynamic time shift. Reference-image guidance is fixed at 1.0.
Metadata base version
omnigen2

Example config

not verified

job: extensionconfig:  name: "my_omnigen2_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "sigmoid"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "OmniGen2/OmniGen2"        arch: "omnigen2"        quantize_te: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page