Docs
AI ToolkitModels

Flex.1-alpha

An 8B, Apache 2.0 descendant of FLUX.1 [schnell] with 8 double-stream blocks instead of 19. It uses the FLUX.1 text encoders and VAE unchanged, and has a guidance embedder that can be bypassed, which is how it is meant to be fine-tuned.

org
Ostris
modality
image
tasks
text-to-image
license
Apache 2.0
released
2025-01-18not verified
native output
1024×1024not verified
total params
13.13B
model.arch
flex1

Components

rolemodelparamssizedtypetrained
Transformer
Flex.1-alpha DiT
diffusers.FluxTransformer2DModel

Same block design as FLUX.1 with 8 double-stream blocks (FLUX.1 has 19). The guidance embedder was trained separately from the rest of the weights so it can be switched off. The repo also has Flex.1-alpha.safetensors, an all-in-one ComfyUI checkpoint with the text encoders and VAE (mixed bf16/fp8/fp16/fp32).

8.16B
8,163,264,064
16.33 GBbf16yes
Text encoder
CLIP ViT-L/14 text model
transformers.CLIPTextModel

Same size as the FLUX.1 CLIP-L. Only the pooled output (768-d) is used.

123.1M
123,060,480
246 MBbf16no
Tokenizer
CLIP BPE (49k vocab)
transformers.CLIPTokenizer
————
Text encoder 2
T5 v1.1 XXL (encoder only)
transformers.T5EncoderModel

Same size as the FLUX.1 T5. 512 tokens.

4.76B
4,762,310,656
9.52 GBbf16no
Tokenizer 2
T5 SentencePiece (32k vocab)
transformers.T5TokenizerFast
————
VAE
FLUX.1 VAE (16 channels)
diffusers.AutoencoderKL
83.8M
83,819,683
168 MBbf16no
total13.13B26.27 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
Identical to FLUX.1: latents are packed 2×2 into 64-value tokens outside the transformer, so each token covers 16×16 pixels.
inputlatent (c×t×h×w)tokens
512×51216×1×64×641,024
1024×102416×1×128×1284,096ai-toolkit sample default
1344×76816×1×96×1684,032landscape

Architecture

Blocks
8 double-stream + 38 single-stream
Hidden size
3072 (24 heads × 128)
FFN size
12288
Text conditioning
T5 hidden states (4096-d, 512 tokens) joined into the sequence; CLIP pooled vector added to the timestep embedding
Guidance
Guidance embedder that can be bypassed. Without it the model runs with true CFG
Objective
Rectified flow, resolution-dependent shift (0.5 at 256 tokens to 1.15 at 4096)
Norm / position
QK RMSNorm, 3-axis RoPE (16, 56, 56)
Lineage
FLUX.1 [schnell] → OpenFLUX.1 → Flex.1-alpha

In AI Toolkit

model.arch
flex1
UI label
Flex.1 (image)
model.name_or_path
ostris/Flex.1-alpha
extra UI sections

UI defaults

quantize / quantize_te
true / true (qfloat8)
train.bypass_guidance_embedding
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
sigmoid (not set by the UI, config default)
network.conv
disabled (linear LoRA only)

Specifics

Loader
ModelConfig rewrites arch flex1 to flux, so Flex.1 goes through exactly the same legacy FLUX.1 code path. Everything on the FLUX.1 card applies.
Guidance embedding
bypass_guidance_embedding replaces the time/text embedding forward so the guidance embedder is skipped during training steps, then restores it. The model card recommends training this way.
Quantization
The example config keeps *time_text_embed* out of quantization (model.quantize_kwargs.exclude). The UI does not add this.
Text encoders
CLIP-L and T5 from name_or_path, never trained. T5 is quantized when quantize_te is on; CLIP is not.
Resolution
Buckets snap to multiples of 32, as for FLUX.1.
LoRA target
FluxTransformer2DModel (transformer_blocks and single_transformer_blocks)
Saving
Diffusers/PEFT LoRA keys (transformer.…lora_A / lora_B), alpha set to the rank. Full fine-tunes save the transformer/ folder only.
Metadata base version
flux.1 (the arch is flux after the rewrite)

Example config

not verified

job: extensionconfig:  name: "my_first_flex_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512, 768, 1024]      train:        batch_size: 1        bypass_guidance_embedding: true        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "ostris/Flex.1-alpha"        arch: "flex1"        quantize: true        quantize_te: true        quantize_kwargs:          exclude:            - "*time_text_embed*"      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page