Docs
AI ToolkitModels

FLUX.2 [dev]

The 32B, guidance-distilled FLUX.2 model. A new double/single-stream transformer conditioned on Mistral Small 3.1 (24B) hidden states, working in the 32-channel latent space of the new FLUX.2 VAE. Reference images go in as extra tokens, so one model generates and edits.

org
Black Forest Labs
modality
image
tasks
text-to-image · image-editing · multi-reference editing
license
FLUX Non-Commercial License
released
2025-11-25not verified
native output
1024×1024, up to 4 MPnot verified
total params
56.32B
model.arch
flux2

Components

rolemodelparamssizedtypetrained
Transformer
FLUX.2 [dev] DiT

ai-toolkit loads this single file in the original BFL layout into its own port of the reference model (Flux2 in flux2/src/model.py). The repo also has a Diffusers transformer/ folder (Flux2Transformer2DModel) with the same parameter count.

32.22B
32,223,281,152
64.45 GBbf16yes
Text encoder
Mistral Small 3.1 24B Instructnot verified
transformers.Mistral3ForConditionalGeneration

The full vision-language model, vision tower included. ai-toolkit loads mistralai/Mistral-Small-3.1-24B-Instruct-2503 instead, whose sharded files have the same parameter count and byte size.

24.01B
24,011,361,280
48.02 GBbf16no
Tokenizer
Mistral / Pixtral processor (131k vocab)
transformers.PixtralProcessor

ai-toolkit uses the processor from the Mistral repo, with fix_mistral_regex off.

————
VAE
FLUX.2 VAE (32 channels)

Not the FLUX.1 VAE. Original-layout file; the int64 part is the BatchNorm step counter. Also shipped as a Diffusers vae/ folder (AutoencoderKLFlux2). ai-toolkit loads ae.safetensors from name_or_path.

84.0M
84,046,372
336 MBfp32+int64no
total56.32B112.81 GB

Latent space

spatial
8×
channels
32
patch
2×2
autoencoder
FLUX.2 VAE
pixels per token
16×16
notes
The 2×2 patchify is part of the VAE (patch_size [2, 2] in its config), followed by BatchNorm with stored running statistics instead of a fixed scale and shift. ai-toolkit's VAE wrapper does both inside encode, so cached latents are 128 channels at 1/16 resolution and the transformer runs with patch size 1 (in_channels 128). Each token covers 16×16 pixels, as in FLUX.1.
inputlatent (c×t×h×w)tokens
512×51232×1×64×641,024
1024×102432×1×128×1284,096ai-toolkit sample default
2048×204832×1×256×25616,3844 MP

Architecture

Blocks
8 double-stream + 48 single-stream
Hidden size
6144 (48 heads × 128)
FFN size
18432 (mlp_ratio 3)
Text conditioning
Mistral hidden states from layers 10, 20 and 30, concatenated to 15360-d, 512 tokens. No pooled vector
Image conditioning
Reference images VAE-encoded and appended as tokens, each offset on the RoPE time axis
Guidance
Guidance-distilled, guidance embedding
Objective
Rectified flow, resolution-dependent shift
Position
4-axis RoPE (32, 32, 32, 32), theta 2000

In AI Toolkit

model.arch
flux2
UI label
FLUX.2 (image)
model.name_or_path
black-forest-labs/FLUX.2-dev
extra UI sections
datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res

UI defaults

quantize / quantize_te
true / true
qtype
qfloat8
low_vram
true
train.unload_text_encoder
false
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
model_kwargs.match_target_res
false
network.conv
disabled (linear LoRA only)

Specifics

Gated
FLUX.2-dev is gated: accept the license on the Hub and set HF_TOKEN. The Mistral repo is not gated.
Transformer source
flux2-dev.safetensors from name_or_path (a Hub repo or a local folder containing that file), loaded into ai-toolkit's own Flux2 module, not Diffusers.
Text encoder source
Always mistralai/Mistral-Small-3.1-24B-Instruct-2503, whatever name_or_path is. Prompts are wrapped in a chat template with a fixed system message and padded to 512 tokens. Never trained.
VAE source
ae.safetensors from name_or_path, or model.vae_path (a local file or a repo/file.safetensors Hub path).
Resolution
Buckets snap to multiples of 16.
Reference images
Up to three control images per sample (ctrl_img_1–3); datasets.multi_control_paths for training. Each is capped at 1024×1024 pixels (1 MP) with its aspect ratio kept, or resized to the target's pixel count when match_target_res is on. Only the target tokens are kept from the prediction.
Guidance during training
The guidance embedding gets train.cfg_scale (default 1.0).
low_vram
Transformer and text encoder are loaded to the CPU and moved to the GPU when needed.
Quantization
Transformer uses qtype, the text encoder uses qtype_te. layer_offloading is supported.
LoRA target
Flux2 (double_blocks and single_blocks)
Saving
LoRA keys use the diffusion_model. prefix and the original BFL module names, which ComfyUI loads. Full fine-tunes save one safetensors file in the original layout.
Metadata base version
flux2

Example config

not verified

job: extensionconfig:  name: "my_flux2_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "black-forest-labs/FLUX.2-dev"        arch: "flux2"        quantize: true        quantize_te: true        qtype: "qfloat8"        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page