Docs
AI ToolkitModels

Flex.2-preview

The follow-up to Flex.1-alpha: the same 8B FLUX-style transformer with inpainting and a universal control input (line, pose, depth) trained into the base model. The extra inputs are concatenated on the latent channels, so the transformer takes 196 input channels per token instead of 64.

org
Ostris
modality
image
tasks
text-to-image · inpainting · controlled generation
license
Apache 2.0
released
2025-04-24not verified
native output
1024×1024not verified
total params
13.13B
model.arch
flex2

Components

rolemodelparamssizedtypetrained
Transformer
Flex.2-preview DiT
diffusers.FluxTransformer2DModel

Flex.1 layout (8 double + 38 single blocks) with a wider input projection: in_channels 196, out_channels 64. Also shipped as Flex.2-preview.safetensors for ComfyUI, which needs the Flex2 Conditioner node from ComfyUI-FlexTools.

8.16B
8,163,669,568
16.33 GBbf16yes
Text encoder
CLIP ViT-L/14 text model
transformers.CLIPTextModel

Only the pooled output (768-d) is used.

123.1M
123,060,480
246 MBbf16no
Tokenizer
CLIP BPE (49k vocab)
transformers.CLIPTokenizer
————
Text encoder 2
T5 v1.1 XXL (encoder only)
transformers.T5EncoderModel

512 tokens.

4.76B
4,762,310,656
9.52 GBbf16no
Tokenizer 2
T5 SentencePiece (32k vocab)
transformers.T5TokenizerFast
————
VAE
FLUX.1 VAE (16 channels)
diffusers.AutoencoderKL

Also encodes the inpaint source and the control image.

83.8M
83,819,683
168 MBbf16no
total13.13B26.27 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
The transformer input is 49 latent channels before packing: 16 noisy latent + 16 masked inpaint latent + 1 inpaint mask + 16 control latent. Packed 2×2 that is 196 values per token. The output is the usual 16 channels (64 packed). Token counts are the same as FLUX.1.
inputlatent (c×t×h×w)tokens
512×51216×1×64×641,024
1024×102416×1×128×1284,096ai-toolkit sample default
1344×76816×1×96×1684,032landscape

Architecture

Blocks
8 double-stream + 38 single-stream
Hidden size
3072 (24 heads × 128)
FFN size
12288
Text conditioning
T5 hidden states (4096-d, 512 tokens) joined into the sequence; CLIP pooled vector added to the timestep embedding
Image conditioning
Inpaint latent, mask and control latent concatenated on the channels (in_channels 196)
Guidance
Guidance embedder, bypassed for training
Objective
Rectified flow, resolution-dependent shift (0.5 at 256 tokens to 1.15 at 4096)
Norm / position
QK RMSNorm, 3-axis RoPE (16, 56, 56)
Lineage
FLUX.1 [schnell] → OpenFLUX.1 → Flex.1-alpha → Flex.2-preview

In AI Toolkit

model.arch
flex2
UI label
Flex.2 (image)
model.name_or_path
ostris/Flex.2-preview
extra UI sections

UI defaults

quantize / quantize_te
true / true (qfloat8)
train.bypass_guidance_embedding
true
noise_scheduler / sampler
flowmatch / flowmatch
datasets.controls
depth, line, pose, inpaint (added to new datasets)
model_kwargs
invert_inpaint_mask_chance 0.2, inpaint_dropout 0.5, control_dropout 0.5, inpaint_random_chance 0.2, do_random_inpainting, random_blur_mask, random_dialate_mask
network.conv
disabled (linear LoRA only)

Specifics

Control images
datasets.controls generates control images for every training image and saves them to a _controls folder next to it: depth with Depth-Anything-V2-Large, line with TEED, pose with DWPose (needs easy_dwpose), and inpaint images whose alpha erases the subject found by BiRefNet_HR. One of the control images is picked at random per step.
Inpaint conditioning
The target latent is masked in latent space, then mask and masked latent are concatenated. Without a mask, do_random_inpainting draws random blobs. With dropout, the inpaint latent is zeros and the mask is all ones.
model_kwargs
Chances for dropping the control (control_dropout), dropping inpainting (inpaint_dropout), forcing a random mask (inpaint_random_chance), inverting the mask, and blurring or dilating it. All default to off in code; the UI turns them on.
Guidance embedding
Bypassed during training by default, as for Flex.1.
Loading
Diffusers from_pretrained for every part. T5 is quantized with the transformer qtype when quantize_te is on; CLIP is not quantized. Text encoders are never trained.
Resolution
Buckets snap to multiples of 16.
LoRA target
FluxTransformer2DModel
Saving
LoRAs use Diffusers/PEFT keys (transformer.…lora_A / lora_B), alpha set to the rank. Full fine-tunes save the transformer/ folder plus aitk_meta.yaml.
Sampling
ctrl_img is optional. A file with .inpaint. in its name must be RGBA and is used for inpainting (alpha is the mask); anything else is a control image.
Metadata base version
flex2

Example config

not verified

job: extensionconfig:  name: "my_first_flex2_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          controls:            - "depth"            - "line"            - "pose"            - "inpaint"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        bypass_guidance_embedding: true        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "shift"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "ostris/Flex.2-preview"        arch: "flex2"        quantize: true        quantize_te: true        model_kwargs:          invert_inpaint_mask_chance: 0.5          inpaint_dropout: 0.5          control_dropout: 0.5          inpaint_random_chance: 0.5          do_random_inpainting: false          random_blur_mask: true          random_dialate_mask: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page