FLUX.2 [dev]
The 32B, guidance-distilled FLUX.2 model. A new double/single-stream transformer conditioned on Mistral Small 3.1 (24B) hidden states, working in the 32-channel latent space of the new FLUX.2 VAE. Reference images go in as extra tokens, so one model generates and edits.
- org
- Black Forest Labs
- modality
- image
- tasks
- text-to-image · image-editing · multi-reference editing
- license
- FLUX Non-Commercial License
- released
- 2025-11-25not verified
- native output
- 1024×1024, up to 4 MPnot verified
- total params
- 56.32B
- model.arch
- flux2
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | FLUX.2 [dev] DiT ai-toolkit loads this single file in the original BFL layout into its own port of the reference model (Flux2 in flux2/src/model.py). The repo also has a Diffusers transformer/ folder (Flux2Transformer2DModel) with the same parameter count. | 32.22B 32,223,281,152 | 64.45 GB | bf16 | yes |
| Text encoder | Mistral Small 3.1 24B Instructnot verified transformers.Mistral3ForConditionalGeneration The full vision-language model, vision tower included. ai-toolkit loads mistralai/Mistral-Small-3.1-24B-Instruct-2503 instead, whose sharded files have the same parameter count and byte size. | 24.01B 24,011,361,280 | 48.02 GB | bf16 | no |
| Tokenizer | Mistral / Pixtral processor (131k vocab) transformers.PixtralProcessor ai-toolkit uses the processor from the Mistral repo, with fix_mistral_regex off. | — | — | — | — |
| VAE | FLUX.2 VAE (32 channels) Not the FLUX.1 VAE. Original-layout file; the int64 part is the BatchNorm step counter. Also shipped as a Diffusers vae/ folder (AutoencoderKLFlux2). ai-toolkit loads ae.safetensors from name_or_path. | 84.0M 84,046,372 | 336 MB | fp32+int64 | no |
| total | 56.32B | 112.81 GB | |||
Latent space
- autoencoder
- FLUX.2 VAE
- pixels per token
- 16×16
- notes
- The 2×2 patchify is part of the VAE (patch_size [2, 2] in its config), followed by BatchNorm with stored running statistics instead of a fixed scale and shift. ai-toolkit's VAE wrapper does both inside encode, so cached latents are 128 channels at 1/16 resolution and the transformer runs with patch size 1 (in_channels 128). Each token covers 16×16 pixels, as in FLUX.1.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 32×1×64×64 | 1,024 | |
| 1024×1024 | 32×1×128×128 | 4,096 | ai-toolkit sample default |
| 2048×2048 | 32×1×256×256 | 16,384 | 4 MP |
Architecture
- Blocks
- 8 double-stream + 48 single-stream
- Hidden size
- 6144 (48 heads × 128)
- FFN size
- 18432 (mlp_ratio 3)
- Text conditioning
- Mistral hidden states from layers 10, 20 and 30, concatenated to 15360-d, 512 tokens. No pooled vector
- Image conditioning
- Reference images VAE-encoded and appended as tokens, each offset on the RoPE time axis
- Guidance
- Guidance-distilled, guidance embedding
- Objective
- Rectified flow, resolution-dependent shift
- Position
- 4-axis RoPE (32, 32, 32, 32), theta 2000
In AI Toolkit
- model.arch
- flux2
- UI label
- FLUX.2 (image)
- model.name_or_path
- black-forest-labs/FLUX.2-dev
- source
- extensions_built_in/diffusion_models/flux2/flux2_model.py
- extensions_built_in/diffusion_models/flux2/src/model.py
- extensions_built_in/diffusion_models/flux2/src/pipeline.py
- extensions_built_in/diffusion_models/flux2/src/sampling.py
- toolkit/models/v2/vae/flux2_kl.py
- toolkit/models/v2/text_encoders/mistral3.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- train.unload_text_encoder
- false
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- model_kwargs.match_target_res
- false
- network.conv
- disabled (linear LoRA only)
Specifics
- Gated
- FLUX.2-dev is gated: accept the license on the Hub and set HF_TOKEN. The Mistral repo is not gated.
- Transformer source
- flux2-dev.safetensors from name_or_path (a Hub repo or a local folder containing that file), loaded into ai-toolkit's own Flux2 module, not Diffusers.
- Text encoder source
- Always mistralai/Mistral-Small-3.1-24B-Instruct-2503, whatever name_or_path is. Prompts are wrapped in a chat template with a fixed system message and padded to 512 tokens. Never trained.
- VAE source
- ae.safetensors from name_or_path, or model.vae_path (a local file or a repo/file.safetensors Hub path).
- Resolution
- Buckets snap to multiples of 16.
- Reference images
- Up to three control images per sample (ctrl_img_1–3); datasets.multi_control_paths for training. Each is capped at 1024×1024 pixels (1 MP) with its aspect ratio kept, or resized to the target's pixel count when match_target_res is on. Only the target tokens are kept from the prediction.
- Guidance during training
- The guidance embedding gets train.cfg_scale (default 1.0).
- low_vram
- Transformer and text encoder are loaded to the CPU and moved to the GPU when needed.
- Quantization
- Transformer uses qtype, the text encoder uses qtype_te. layer_offloading is supported.
- LoRA target
- Flux2 (double_blocks and single_blocks)
- Saving
- LoRA keys use the diffusion_model. prefix and the original BFL module names, which ComfyUI loads. Full fine-tunes save one safetensors file in the original layout.
- Metadata base version
- flux2
Example config
not verified
job: extensionconfig: name: "my_flux2_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "black-forest-labs/FLUX.2-dev" arch: "flux2" quantize: true quantize_te: true qtype: "qfloat8" low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
OmniGen2
A 4B Lumina-style diffusion transformer conditioned on Qwen2.5-VL 3B. Reference images enter as extra latent tokens through their own refiner, so one model does text-to-image, instruction editing and subject-driven generation.
Z-Image Turbo
'The step-distilled 6B member of Z-Image: a single-stream DiT that makes images in about 8 steps without CFG. Training it directly breaks the distillation, so ai-toolkit trains it through a de-distilling training adapter that is merged in for training and removed for sampling.'