Flex.1-alpha
An 8B, Apache 2.0 descendant of FLUX.1 [schnell] with 8 double-stream blocks instead of 19. It uses the FLUX.1 text encoders and VAE unchanged, and has a guidance embedder that can be bypassed, which is how it is meant to be fine-tuned.
- weights
- ostris/Flex.1-alpha ↗
- org
- Ostris
- modality
- image
- tasks
- text-to-image
- license
- Apache 2.0
- released
- 2025-01-18not verified
- native output
- 1024×1024not verified
- total params
- 13.13B
- model.arch
- flex1
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Flex.1-alpha DiT diffusers.FluxTransformer2DModel Same block design as FLUX.1 with 8 double-stream blocks (FLUX.1 has 19). The guidance embedder was trained separately from the rest of the weights so it can be switched off. The repo also has Flex.1-alpha.safetensors, an all-in-one ComfyUI checkpoint with the text encoders and VAE (mixed bf16/fp8/fp16/fp32). | 8.16B 8,163,264,064 | 16.33 GB | bf16 | yes |
| Text encoder | CLIP ViT-L/14 text model transformers.CLIPTextModel Same size as the FLUX.1 CLIP-L. Only the pooled output (768-d) is used. | 123.1M 123,060,480 | 246 MB | bf16 | no |
| Tokenizer | CLIP BPE (49k vocab) transformers.CLIPTokenizer | — | — | — | — |
| Text encoder 2 | T5 v1.1 XXL (encoder only) transformers.T5EncoderModel Same size as the FLUX.1 T5. 512 tokens. | 4.76B 4,762,310,656 | 9.52 GB | bf16 | no |
| Tokenizer 2 | T5 SentencePiece (32k vocab) transformers.T5TokenizerFast | — | — | — | — |
| VAE | FLUX.1 VAE (16 channels) diffusers.AutoencoderKL | 83.8M 83,819,683 | 168 MB | bf16 | no |
| total | 13.13B | 26.27 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- Identical to FLUX.1: latents are packed 2×2 into 64-value tokens outside the transformer, so each token covers 16×16 pixels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 16×1×64×64 | 1,024 | |
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 1344×768 | 16×1×96×168 | 4,032 | landscape |
Architecture
- Blocks
- 8 double-stream + 38 single-stream
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288
- Text conditioning
- T5 hidden states (4096-d, 512 tokens) joined into the sequence; CLIP pooled vector added to the timestep embedding
- Guidance
- Guidance embedder that can be bypassed. Without it the model runs with true CFG
- Objective
- Rectified flow, resolution-dependent shift (0.5 at 256 tokens to 1.15 at 4096)
- Norm / position
- QK RMSNorm, 3-axis RoPE (16, 56, 56)
- Lineage
- FLUX.1 [schnell] → OpenFLUX.1 → Flex.1-alpha
In AI Toolkit
- model.arch
- flex1
- UI label
- Flex.1 (image)
- model.name_or_path
- ostris/Flex.1-alpha
- source
- extra UI sections
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- train.bypass_guidance_embedding
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- sigmoid (not set by the UI, config default)
- network.conv
- disabled (linear LoRA only)
Specifics
- Loader
- ModelConfig rewrites arch flex1 to flux, so Flex.1 goes through exactly the same legacy FLUX.1 code path. Everything on the FLUX.1 card applies.
- Guidance embedding
- bypass_guidance_embedding replaces the time/text embedding forward so the guidance embedder is skipped during training steps, then restores it. The model card recommends training this way.
- Quantization
- The example config keeps *time_text_embed* out of quantization (model.quantize_kwargs.exclude). The UI does not add this.
- Text encoders
- CLIP-L and T5 from name_or_path, never trained. T5 is quantized when quantize_te is on; CLIP is not.
- Resolution
- Buckets snap to multiples of 32, as for FLUX.1.
- LoRA target
- FluxTransformer2DModel (transformer_blocks and single_transformer_blocks)
- Saving
- Diffusers/PEFT LoRA keys (transformer.…lora_A / lora_B), alpha set to the rank. Full fine-tunes save the transformer/ folder only.
- Metadata base version
- flux.1 (the arch is flux after the rewrite)
Example config
not verified
job: extensionconfig: name: "my_first_flex_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768, 1024] train: batch_size: 1 bypass_guidance_embedding: true steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "ostris/Flex.1-alpha" arch: "flex1" quantize: true quantize_te: true quantize_kwargs: exclude: - "*time_text_embed*" sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
FLUX.1 [dev]
The 12B guidance-distilled FLUX.1 model. A double- and single-stream rectified flow transformer conditioned on T5-XXL and CLIP-L, working in a 16-channel, 8× VAE latent space. The base most FLUX LoRAs are trained on.
Chroma1-Base
'An 8.9B text-to-image model built from FLUX.1-schnell, with CLIP removed and the per-block modulation replaced by one small approximator network. Chroma1-Base is the neutral checkpoint meant for fine-tuning; Chroma1-HD shares its architecture and size.'