Flex.2-preview
The follow-up to Flex.1-alpha: the same 8B FLUX-style transformer with inpainting and a universal control input (line, pose, depth) trained into the base model. The extra inputs are concatenated on the latent channels, so the transformer takes 196 input channels per token instead of 64.
- weights
- ostris/Flex.2-preview ↗
- org
- Ostris
- modality
- image
- tasks
- text-to-image · inpainting · controlled generation
- license
- Apache 2.0
- released
- 2025-04-24not verified
- native output
- 1024×1024not verified
- total params
- 13.13B
- model.arch
- flex2
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Flex.2-preview DiT diffusers.FluxTransformer2DModel Flex.1 layout (8 double + 38 single blocks) with a wider input projection: in_channels 196, out_channels 64. Also shipped as Flex.2-preview.safetensors for ComfyUI, which needs the Flex2 Conditioner node from ComfyUI-FlexTools. | 8.16B 8,163,669,568 | 16.33 GB | bf16 | yes |
| Text encoder | CLIP ViT-L/14 text model transformers.CLIPTextModel Only the pooled output (768-d) is used. | 123.1M 123,060,480 | 246 MB | bf16 | no |
| Tokenizer | CLIP BPE (49k vocab) transformers.CLIPTokenizer | — | — | — | — |
| Text encoder 2 | T5 v1.1 XXL (encoder only) transformers.T5EncoderModel 512 tokens. | 4.76B 4,762,310,656 | 9.52 GB | bf16 | no |
| Tokenizer 2 | T5 SentencePiece (32k vocab) transformers.T5TokenizerFast | — | — | — | — |
| VAE | FLUX.1 VAE (16 channels) diffusers.AutoencoderKL Also encodes the inpaint source and the control image. | 83.8M 83,819,683 | 168 MB | bf16 | no |
| total | 13.13B | 26.27 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- The transformer input is 49 latent channels before packing: 16 noisy latent + 16 masked inpaint latent + 1 inpaint mask + 16 control latent. Packed 2×2 that is 196 values per token. The output is the usual 16 channels (64 packed). Token counts are the same as FLUX.1.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 16×1×64×64 | 1,024 | |
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 1344×768 | 16×1×96×168 | 4,032 | landscape |
Architecture
- Blocks
- 8 double-stream + 38 single-stream
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288
- Text conditioning
- T5 hidden states (4096-d, 512 tokens) joined into the sequence; CLIP pooled vector added to the timestep embedding
- Image conditioning
- Inpaint latent, mask and control latent concatenated on the channels (in_channels 196)
- Guidance
- Guidance embedder, bypassed for training
- Objective
- Rectified flow, resolution-dependent shift (0.5 at 256 tokens to 1.15 at 4096)
- Norm / position
- QK RMSNorm, 3-axis RoPE (16, 56, 56)
- Lineage
- FLUX.1 [schnell] → OpenFLUX.1 → Flex.1-alpha → Flex.2-preview
In AI Toolkit
- model.arch
- flex2
- UI label
- Flex.2 (image)
- model.name_or_path
- ostris/Flex.2-preview
- source
- extra UI sections
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- train.bypass_guidance_embedding
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- datasets.controls
- depth, line, pose, inpaint (added to new datasets)
- model_kwargs
- invert_inpaint_mask_chance 0.2, inpaint_dropout 0.5, control_dropout 0.5, inpaint_random_chance 0.2, do_random_inpainting, random_blur_mask, random_dialate_mask
- network.conv
- disabled (linear LoRA only)
Specifics
- Control images
- datasets.controls generates control images for every training image and saves them to a _controls folder next to it: depth with Depth-Anything-V2-Large, line with TEED, pose with DWPose (needs easy_dwpose), and inpaint images whose alpha erases the subject found by BiRefNet_HR. One of the control images is picked at random per step.
- Inpaint conditioning
- The target latent is masked in latent space, then mask and masked latent are concatenated. Without a mask, do_random_inpainting draws random blobs. With dropout, the inpaint latent is zeros and the mask is all ones.
- model_kwargs
- Chances for dropping the control (control_dropout), dropping inpainting (inpaint_dropout), forcing a random mask (inpaint_random_chance), inverting the mask, and blurring or dilating it. All default to off in code; the UI turns them on.
- Guidance embedding
- Bypassed during training by default, as for Flex.1.
- Loading
- Diffusers from_pretrained for every part. T5 is quantized with the transformer qtype when quantize_te is on; CLIP is not quantized. Text encoders are never trained.
- Resolution
- Buckets snap to multiples of 16.
- LoRA target
- FluxTransformer2DModel
- Saving
- LoRAs use Diffusers/PEFT keys (transformer.…lora_A / lora_B), alpha set to the rank. Full fine-tunes save the transformer/ folder plus aitk_meta.yaml.
- Sampling
- ctrl_img is optional. A file with .inpaint. in its name must be RGBA and is used for inpainting (alpha is the mask); anything else is a control image.
- Metadata base version
- flex2
Example config
not verified
job: extensionconfig: name: "my_first_flex2_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" controls: - "depth" - "line" - "pose" - "inpaint" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 bypass_guidance_embedding: true steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "shift" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "ostris/Flex.2-preview" arch: "flex2" quantize: true quantize_te: true model_kwargs: invert_inpaint_mask_chance: 0.5 inpaint_dropout: 0.5 control_dropout: 0.5 inpaint_random_chance: 0.5 do_random_inpainting: false random_blur_mask: true random_dialate_mask: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Boogu-Image 0.1 Base
'The undistilled text-to-image base of Boogu-Image 0.1: a 10.3B Lumina2-style DiT with 8 double-stream and 32 single-stream layers, conditioned on Qwen3-VL-8B and decoding through the FLUX.1 VAE. Its authors pitch it for fine-tuning and dense Chinese and English text rendering.'
FLUX.1 Kontext [dev]
The instruction-editing member of FLUX.1. Same 12B transformer shape, text encoders and VAE as FLUX.1 [dev]; the input image goes in as a second set of tokens next to the image being generated. Trained in ai-toolkit on before/after pairs.