FLUX.2 [klein] 4B Base
The smallest FLUX.2 model: a 4B transformer conditioned on Qwen3-4B, using the same 32-channel FLUX.2 VAE and reference-image editing as FLUX.2 [dev]. The Base release is undistilled (no step or guidance distillation), Apache 2.0, and meant for fine-tuning.
- org
- Black Forest Labs
- modality
- image
- tasks
- text-to-image · image-editing · multi-reference editing
- license
- Apache 2.0
- released
- 2026-01-15not verified
- native output
- 1024×1024not verified
- total params
- 7.98B
- model.arch
- flux2_klein_4b
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | FLUX.2 [klein] 4B Base DiT ai-toolkit loads this single file in the original BFL layout into its own Flux2 module. The repo also has a Diffusers transformer/ folder with the same parameter count. | 3.88B 3,875,544,576 | 7.75 GB | bf16 | yes |
| Text encoder | Qwen3-4Bnot verified transformers.Qwen3ForCausalLM ai-toolkit loads Qwen/Qwen3-4B instead, which has the same parameter count (8,044,982,000 bytes across its shards). | 4.02B 4,022,468,096 | 8.04 GB | bf16 | no |
| Tokenizer | Qwen2 BPE (152k vocab) transformers.Qwen2TokenizerFast ai-toolkit loads the tokenizer from Qwen/Qwen3-4B. | — | — | — | — |
| VAE | FLUX.2 VAE (32 channels) diffusers.AutoencoderKLFlux2 Same VAE as FLUX.2 [dev]. The klein repo only has it as a Diffusers folder, so ai-toolkit loads ai-toolkit/flux2_vae (ae.safetensors, fp32, 336,211,292 bytes, same parameter count) unless vae_path is set. | 84.0M 84,046,372 | 168 MB | bf16+int64 | no |
| total | 7.98B | 15.96 GB | |||
Latent space
- autoencoder
- FLUX.2 VAE
- pixels per token
- 16×16
- notes
- Same latent space as FLUX.2 [dev]: the VAE patchifies 2×2 and normalizes with BatchNorm, so the transformer sees 128-channel tokens at patch size 1, each covering 16×16 pixels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 32×1×64×64 | 1,024 | |
| 1024×1024 | 32×1×128×128 | 4,096 | ai-toolkit sample default |
| 1344×768 | 32×1×96×168 | 4,032 | landscape |
Architecture
- Blocks
- 5 double-stream + 20 single-stream
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 9216 (mlp_ratio 3)
- Text conditioning
- Qwen3 hidden states from layers 9, 18 and 27, concatenated to 7680-d, 512 tokens. No pooled vector
- Image conditioning
- Reference images VAE-encoded and appended as tokens, each offset on the RoPE time axis
- Guidance
- No guidance embedding. Undistilled, sampled with real CFG
- Objective
- Rectified flow, resolution-dependent shift
- Position
- 4-axis RoPE (32, 32, 32, 32), theta 2000
In AI Toolkit
- model.arch
- flux2_klein_4b
- UI label
- FLUX.2-klein-base-4B (image)
- model.name_or_path
- black-forest-labs/FLUX.2-klein-base-4B
- source
- extensions_built_in/diffusion_models/flux2/flux2_klein_model.py
- extensions_built_in/diffusion_models/flux2/flux2_model.py
- extensions_built_in/diffusion_models/flux2/src/model.py
- extensions_built_in/diffusion_models/flux2/src/pipeline.py
- extensions_built_in/diffusion_models/flux2/src/sampling.py
- toolkit/models/v2/vae/flux2_kl.py
- toolkit/models/v2/text_encoders/qwen3.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.qie.match_target_res
UI defaults
- quantize / quantize_te
- true / true
- qtype
- qfloat8
- low_vram
- true
- train.unload_text_encoder
- false
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- model_kwargs.match_target_res
- false
- network.conv
- disabled (linear LoRA only)
Specifics
- Code path
- Flux2KleinModel subclasses the FLUX.2 [dev] model class; only the text encoder, VAE source and guidance handling differ.
- Transformer source
- flux-2-klein-base-4b.safetensors from name_or_path (a Hub repo or a local folder containing that file).
- Text encoder source
- Always Qwen/Qwen3-4B, whatever name_or_path is. Prompts go through the Qwen chat template with thinking disabled and are padded to 512 tokens. Never trained.
- VAE source
- ai-toolkit/flux2_vae (ae.safetensors) unless name_or_path is a folder with ae.safetensors or model.vae_path is set.
- Resolution
- Buckets snap to multiples of 16.
- Reference images
- Same as FLUX.2 [dev]: up to three control images per sample, datasets.multi_control_paths for training, each capped at 1 MP unless match_target_res is on.
- Sampling
- Not guidance-distilled, so samples run real CFG with the negative prompt at guidance_scale.
- low_vram
- Transformer and text encoder are loaded to the CPU and moved to the GPU when needed.
- LoRA target
- Flux2 (double_blocks and single_blocks)
- Saving
- LoRA keys use the diffusion_model. prefix and BFL module names (ComfyUI layout). LoKr saves in the new format by default. Full fine-tunes save one safetensors file.
- Metadata base version
- flux2_klein_4b
Example config
not verified
job: extensionconfig: name: "my_flux2_klein_4b_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "black-forest-labs/FLUX.2-klein-base-4B" arch: "flux2_klein_4b" quantize: true quantize_te: true qtype: "qfloat8" low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 30 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Z-Image De-Turbo
'Z-Image Turbo with the step distillation trained back out, made by fine-tuning Turbo on its own outputs. It runs with normal CFG at 20–30 steps and can be trained directly, with no adapter, for LoRAs or long fine-tunes that stay compatible with the Turbo weights.'
ERNIE-Image
Baidu’s 8B single-stream DiT for text-to-image. It reads a Ministral 3 text encoder’s second-to-last hidden states and works in the FLUX.2 VAE latent space, 2×2 patchified to 128 channels. The repo also ships a prompt-enhancer LLM that ai-toolkit does not load.