Docs
AI ToolkitModels

Lumina-Image 2.0

A 2.6B flow-matching diffusion transformer that reads Gemma 2 2B hidden states and works in the FLUX.1 latent space. Text and image tokens pass through separate refiner layers, then share one stack of single-stream blocks.

org
Alpha-VLLM (Shanghai AI Lab)
modality
image
tasks
text-to-image
license
Apache 2.0
released
2025-01-22not verified
native output
1024×1024not verified
total params
5.31B
model.arch
lumina2

Components

rolemodelparamssizedtypetrained
Transformer
Lumina-Image 2.0 Unified Next-DiT
diffusers.Lumina2Transformer2DModel

Stored as fp32 (about 5.2 GB in bf16). The repo root also holds the original-format weights (consolidated.00-of-01.pth, 10.4 GB), which ai-toolkit does not use.

2.61B
2,609,769,152
10.44 GBfp32yes
Text encoder
Gemma 2 2B
transformers.Gemma2Model

Decoder-only LM used as an encoder (no LM head). Includes the 256k-token embedding table.

2.61B
2,614,341,888
10.46 GBfp32no
Tokenizer
Gemma SentencePiece (256k vocab)
transformers.GemmaTokenizer
————
VAE
FLUX.1 VAE
diffusers.AutoencoderKL

Same config as the FLUX.1 autoencoder (scaling 0.3611, shift 0.1159).

83.8M
83,819,683
335 MBfp32no
total5.31B21.23 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
768×134416×1×168×964,032portrait, same area
512×51216×1×64×641,024

Architecture

Blocks
26 single-stream, plus 2 noise-refiner and 2 context-refiner layers
Hidden size
2304 (24 query heads × 96, 8 KV heads)
FFN size
9216 (SwiGLU)
Text conditioning
Penultimate Gemma 2 hidden states (2304-d) joined into the sequence after the context refiner. 256 tokens max in ai-toolkit
Prompt format
The pipeline prepends a fixed system prompt and "<Prompt Start>" to every caption
Objective
Rectified flow, shift 6.0. Time runs from 0 = noise to 1 = image
Position
3-axis RoPE (32/32/32), axis lengths 300 / 512 / 512

In AI Toolkit

model.arch
lumina2
UI label
Lumina2 (image)
model.name_or_path
Alpha-VLLM/Lumina-Image-2.0
extra UI sections

UI defaults

quantize / quantize_te
false / true (qfloat8)
noise_scheduler / sampler
flowmatch / flowmatch
network.conv
disabled (linear LoRA only)

Specifics

Model code
Handled by the legacy StableDiffusion class (arch lumina2 sets is_lumina2), not a model plugin.
Loading
Transformer from transformer/ in name_or_path; VAE, scheduler, tokenizer and Gemma from the original name_or_path. A local folder that has text_encoder/ is used as the base for all of them. te_name_or_path swaps the text encoder.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Timesteps
Training uses the flowmatch scheduler with Lumina's config (shift 6.0, no dynamic shifting). The example config uses timestep_type lumina2_shift, which is the same code path as shift.
Time direction
ai-toolkit passes 1 − t to the transformer and negates its output to match its noise − latents target.
Quantization
quantize_te quantizes Gemma 2. The transformer stays unquantized by default; its state dict is dequantized on save when it is.
LoRA target
Lumina2Transformer2DModel (layers, noise_refiner, context_refiner)
Saving
LoRAs save in PEFT format with transformer. keys. Full fine-tunes save transformer/ in Diffusers format.
Not supported
split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
Metadata base version
lumina2

Example config

not verified

job: extensionconfig:  name: "my_lumina2_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "lumina2_shift"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "Alpha-VLLM/Lumina-Image-2.0"        arch: "lumina2"        quantize_te: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4.0        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page