Lumina-Image 2.0
A 2.6B flow-matching diffusion transformer that reads Gemma 2 2B hidden states and works in the FLUX.1 latent space. Text and image tokens pass through separate refiner layers, then share one stack of single-stream blocks.
- org
- Alpha-VLLM (Shanghai AI Lab)
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2025-01-22not verified
- native output
- 1024×1024not verified
- total params
- 5.31B
- model.arch
- lumina2
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Lumina-Image 2.0 Unified Next-DiT diffusers.Lumina2Transformer2DModel Stored as fp32 (about 5.2 GB in bf16). The repo root also holds the original-format weights (consolidated.00-of-01.pth, 10.4 GB), which ai-toolkit does not use. | 2.61B 2,609,769,152 | 10.44 GB | fp32 | yes |
| Text encoder | Gemma 2 2B transformers.Gemma2Model Decoder-only LM used as an encoder (no LM head). Includes the 256k-token embedding table. | 2.61B 2,614,341,888 | 10.46 GB | fp32 | no |
| Tokenizer | Gemma SentencePiece (256k vocab) transformers.GemmaTokenizer | — | — | — | — |
| VAE | FLUX.1 VAE diffusers.AutoencoderKL Same config as the FLUX.1 autoencoder (scaling 0.3611, shift 0.1159). | 83.8M 83,819,683 | 335 MB | fp32 | no |
| total | 5.31B | 21.23 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 768×1344 | 16×1×168×96 | 4,032 | portrait, same area |
| 512×512 | 16×1×64×64 | 1,024 |
Architecture
- Blocks
- 26 single-stream, plus 2 noise-refiner and 2 context-refiner layers
- Hidden size
- 2304 (24 query heads × 96, 8 KV heads)
- FFN size
- 9216 (SwiGLU)
- Text conditioning
- Penultimate Gemma 2 hidden states (2304-d) joined into the sequence after the context refiner. 256 tokens max in ai-toolkit
- Prompt format
- The pipeline prepends a fixed system prompt and "<Prompt Start>" to every caption
- Objective
- Rectified flow, shift 6.0. Time runs from 0 = noise to 1 = image
- Position
- 3-axis RoPE (32/32/32), axis lengths 300 / 512 / 512
In AI Toolkit
- model.arch
- lumina2
- UI label
- Lumina2 (image)
- model.name_or_path
- Alpha-VLLM/Lumina-Image-2.0
- source
- extra UI sections
UI defaults
- quantize / quantize_te
- false / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- network.conv
- disabled (linear LoRA only)
Specifics
- Model code
- Handled by the legacy StableDiffusion class (arch lumina2 sets is_lumina2), not a model plugin.
- Loading
- Transformer from transformer/ in name_or_path; VAE, scheduler, tokenizer and Gemma from the original name_or_path. A local folder that has text_encoder/ is used as the base for all of them. te_name_or_path swaps the text encoder.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Timesteps
- Training uses the flowmatch scheduler with Lumina's config (shift 6.0, no dynamic shifting). The example config uses timestep_type lumina2_shift, which is the same code path as shift.
- Time direction
- ai-toolkit passes 1 − t to the transformer and negates its output to match its noise − latents target.
- Quantization
- quantize_te quantizes Gemma 2. The transformer stays unquantized by default; its state dict is dequantized on save when it is.
- LoRA target
- Lumina2Transformer2DModel (layers, noise_refiner, context_refiner)
- Saving
- LoRAs save in PEFT format with transformer. keys. Full fine-tunes save transformer/ in Diffusers format.
- Not supported
- split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
- Metadata base version
- lumina2
Example config
not verified
job: extensionconfig: name: "my_lumina2_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "lumina2_shift" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "Alpha-VLLM/Lumina-Image-2.0" arch: "lumina2" quantize_te: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4.0 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Chroma1-Base
'An 8.9B text-to-image model built from FLUX.1-schnell, with CLIP removed and the per-block modulation replaced by one small approximator network. Chroma1-Base is the neutral checkpoint meant for fine-tuning; Chroma1-HD shares its architecture and size.'
Qwen-Image
'The original 20B Qwen-Image text-to-image model: a 60-block double-stream MMDiT conditioned on Qwen2.5-VL-7B, working in the 8× / 16-channel latent space of a Wan-style VAE. Strong at text rendering, especially Chinese.'