FLUX.1 [dev]
The 12B guidance-distilled FLUX.1 model. A double- and single-stream rectified flow transformer conditioned on T5-XXL and CLIP-L, working in a 16-channel, 8× VAE latent space. The base most FLUX LoRAs are trained on.
- org
- Black Forest Labs
- modality
- image
- tasks
- text-to-image
- license
- FLUX.1 [dev] Non-Commercial License
- released
- 2024-08-01not verified
- native output
- 1024×1024, other aspect ratios from about 0.1 to 2 MPnot verified
- total params
- 16.87B
- model.arch
- flux
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | FLUX.1 [dev] DiT diffusers.FluxTransformer2DModel The repo also ships the same weights as a single file in the original layout (flux1-dev.safetensors), which ComfyUI uses. ai-toolkit loads the Diffusers transformer/ folder. | 11.90B 11,901,408,320 | 23.80 GB | bf16 | yes |
| Text encoder | CLIP ViT-L/14 text model transformers.CLIPTextModel Only the pooled output (768-d) is used. 77 tokens max. | 123.1M 123,060,480 | 246 MB | bf16 | no |
| Tokenizer | CLIP BPE (49k vocab) transformers.CLIPTokenizer | — | — | — | — |
| Text encoder 2 | T5 v1.1 XXL (encoder only) transformers.T5EncoderModel Per-token hidden states (4096-d) go into the transformer sequence. ai-toolkit pads to 512 tokens. | 4.76B 4,762,310,656 | 9.52 GB | bf16 | no |
| Tokenizer 2 | T5 SentencePiece (32k vocab) transformers.T5TokenizerFast | — | — | — | — |
| VAE | FLUX.1 VAE (16 channels) diffusers.AutoencoderKL Also shipped as ae.safetensors in fp32 in the original layout. Shared by every FLUX.1 model, Flex.1 and Flex.2. | 83.8M 83,819,683 | 168 MB | bf16 | no |
| total | 16.87B | 33.74 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- Latents are scaled with scaling_factor 0.3611 and shift_factor 0.1159. The 2×2 patching happens outside the transformer: latents are packed into 64-value tokens (in_channels 64, patch_size 1 in the transformer config), so each token covers 16×16 pixels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 16×1×64×64 | 1,024 | |
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 1344×768 | 16×1×96×168 | 4,032 | landscape |
Architecture
- Blocks
- 19 double-stream + 38 single-stream
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288
- Text conditioning
- T5 hidden states (4096-d) joined into the sequence; CLIP pooled vector (768-d) added to the timestep embedding
- Guidance
- Guidance-distilled: guidance scale goes in through an embedding, no CFG pass
- Objective
- Rectified flow, resolution-dependent shift (0.5 at 256 tokens to 1.15 at 4096)
- Norm / position
- QK RMSNorm, 3-axis RoPE (16, 56, 56)
In AI Toolkit
- model.arch
- flux
- UI label
- FLUX.1 (image)
- model.name_or_path
- black-forest-labs/FLUX.1-dev
- source
- extra UI sections
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- sigmoid (not set by the UI, config default)
- network.conv
- disabled (linear LoRA only)
Specifics
- Loader
- A legacy arch: loaded by the StableDiffusion class in toolkit/stable_diffusion_model.py, not a model extension. Old configs with is_flux: true still work.
- Gated
- Accept the license on the Hub and set HF_TOKEN before the first run.
- Resolution
- Buckets snap to multiples of 32 (the legacy loader doubles the 16 px token size).
- Text encoders
- Both come from name_or_path and are never trained. T5 is quantized with the transformer qtype when quantize_te is on; CLIP is not quantized. T5 prompts are padded to 512 tokens.
- Guidance during training
- The guidance embedding gets train.cfg_scale, which defaults to 1.0. bypass_guidance_embedding skips it entirely (used for Flex.1).
- Quantization
- The transformer is quantized with qtype plus any model.quantize_kwargs (for example exclude patterns).
- LoRA target
- FluxTransformer2DModel (transformer_blocks and single_transformer_blocks)
- Saving
- save_format is forced to diffusers. LoRAs are saved with Diffusers/PEFT keys (transformer.…lora_A / lora_B, no alpha keys), which ComfyUI can load. In this format linear_alpha is ignored: alpha is set to the rank. Full fine-tunes save only the transformer/ folder.
- Assistant LoRA
- assistant_lora_path is merged into the transformer before quantizing and removed (weight −1) when sampling. This is how FLUX.1 [schnell] is trained with a training adapter. inference_lora_path merges a LoRA that stays on.
- Other options
- split_model_over_gpus spreads the blocks across GPUs. use_flux_cfg samples with true CFG and bypasses the guidance embedding.
- Sampling
- Flowmatch Euler with dynamic shift. Negative prompts are ignored unless use_flux_cfg is on.
- Metadata base version
- flux.1
Example config
not verified
job: extensionconfig: name: "my_first_flux_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "black-forest-labs/FLUX.1-dev" arch: "flux" quantize: true quantize_te: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 20 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Anima Base v1.0
A 2B anime and illustration text-to-image model built on the Cosmos-Predict2 2B DiT. A small Qwen3 0.6B encoder feeds a learned 6-layer text conditioner, and images live in the 8× latent space of the Qwen-Image VAE. The Base version is the one meant for LoRA training.
Flex.1-alpha
An 8B, Apache 2.0 descendant of FLUX.1 [schnell] with 8 double-stream blocks instead of 19. It uses the FLUX.1 text encoders and VAE unchanged, and has a guidance embedder that can be bypassed, which is how it is meant to be fine-tuned.