FLUX.1 Kontext [dev]
The instruction-editing member of FLUX.1. Same 12B transformer shape, text encoders and VAE as FLUX.1 [dev]; the input image goes in as a second set of tokens next to the image being generated. Trained in ai-toolkit on before/after pairs.
- org
- Black Forest Labs
- modality
- image
- tasks
- image-editing · text-to-image
- license
- FLUX.1 [dev] Non-Commercial License
- released
- 2025-06-26not verified
- native output
- 1024×1024, output size follows the input imagenot verified
- total params
- 16.87B
- model.arch
- flux_kontext
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | FLUX.1 Kontext [dev] DiT diffusers.FluxTransformer2DModel Same parameter count and config as FLUX.1 [dev] (in_channels 64, guidance embeds). Also shipped as flux1-kontext-dev.safetensors in the original layout. | 11.90B 11,901,408,320 | 23.80 GB | bf16 | yes |
| Text encoder | CLIP ViT-L/14 text model transformers.CLIPTextModel Only the pooled output (768-d) is used. | 123.1M 123,060,480 | 246 MB | bf16 | no |
| Tokenizer | CLIP BPE (49k vocab) transformers.CLIPTokenizer | — | — | — | — |
| Text encoder 2 | T5 v1.1 XXL (encoder only) transformers.T5EncoderModel 512 tokens in ai-toolkit. | 4.76B 4,762,310,656 | 9.52 GB | bf16 | no |
| Tokenizer 2 | T5 SentencePiece (32k vocab) transformers.T5TokenizerFast | — | — | — | — |
| VAE | FLUX.1 VAE (16 channels) diffusers.AutoencoderKL Encodes both the target and the input (control) image. | 83.8M 83,819,683 | 168 MB | bf16 | no |
| total | 16.87B | 33.74 GB | |||
Latent space
- autoencoder
- FLUX.1 VAE
- pixels per token
- 16×16
- notes
- Same packing as FLUX.1. The control image is encoded the same way and appended as extra tokens, so with one control image at the target size the sequence is twice the token counts below.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 512×512 | 16×1×64×64 | 1,024 | target only; ×2 with control |
| 1024×1024 | 16×1×128×128 | 4,096 | target only; ×2 with control |
| 1344×768 | 16×1×96×168 | 4,032 | landscape, target only |
Architecture
- Blocks
- 19 double-stream + 38 single-stream
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 12288
- Text conditioning
- T5 hidden states (4096-d) joined into the sequence; CLIP pooled vector (768-d) added to the timestep embedding
- Image conditioning
- Input image latents appended as tokens; their RoPE ids are marked with 1 on the first axis
- Guidance
- Guidance-distilled, guidance embedding
- Objective
- Rectified flow, resolution-dependent shift (0.5 at 256 tokens to 1.15 at 4096)
- Norm / position
- QK RMSNorm, 3-axis RoPE (16, 56, 56)
In AI Toolkit
- model.arch
- flux_kontext
- UI label
- FLUX.1-Kontext-dev (instruction)
- model.name_or_path
- black-forest-labs/FLUX.1-Kontext-dev
- source
- extensions_built_in/diffusion_models/flux_kontext/flux_kontext.py
- toolkit/models/v2/diffusion_models/flux.py
- toolkit/models/v2/text_encoders/t5.py
- toolkit/models/v2/text_encoders/clip.py
- toolkit/models/v2/vae/autoencoder_kl.py
- toolkit/models/flux.py
- toolkit/train_tools.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.control_path, sample.ctrl_img
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- network.conv
- disabled (linear LoRA only)
Specifics
- Gated
- Accept the license on the Hub and set HF_TOKEN before the first run.
- Training data
- Pairs: datasets.control_path holds the input images, with the same file names as the edited targets in folder_path. Captions are the edit instructions.not verified
- Control image
- Resized (bilinear) to the target crop size, VAE-encoded, then packed and appended after the target tokens. Only the target tokens are kept from the prediction, so the loss covers only them.
- Resolution
- Buckets snap to multiples of 16.
- Loading
- Transformer from name_or_path. T5, CLIP and VAE come from extras_name_or_path, which defaults to name_or_path, or from name_or_path if it is a local folder with text_encoder/.
- Quantization
- Transformer uses qtype, T5 uses qtype_te. CLIP is never quantized. layer_offloading is supported.
- low_vram
- Keeps the text encoders on the CPU and moves them to the GPU only to encode prompts.
- Guidance during training
- The guidance embedding gets train.cfg_scale (default 1.0).
- LoRA target
- FluxTransformer2DModel
- Saving
- LoRAs use Diffusers/PEFT keys (transformer.…lora_A / lora_B), alpha set to the rank. Full fine-tunes save the transformer/ folder plus aitk_meta.yaml.
- Sampling
- Every sample prompt needs --ctrl_img. The control image is resized to the sample width × height, which are floored to multiples of 16.
- Metadata base version
- flux.1_kontext
Example config
not verified
job: extensionconfig: name: "my_first_flux_kontext_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/edited/images" control_path: "/path/to/input/images" caption_ext: "txt" caption_dropout_rate: 0.05 cache_latents_to_disk: true resolution: [512, 768] train: batch_size: 1 steps: 3000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "black-forest-labs/FLUX.1-Kontext-dev" arch: "flux_kontext" quantize: true quantize_te: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 20 prompts: - "make the person smile --ctrl_img /path/to/input/images/person1.jpg"Links
Flex.2-preview
The follow-up to Flex.1-alpha: the same 8B FLUX-style transformer with inpainting and a universal control input (line, pose, depth) trained into the base model. The extra inputs are concatenated on the latent channels, so the transformer takes 196 input channels per token instead of 64.
Qwen-Image-Edit
'The first image-editing version of the 20B Qwen-Image model. One reference image goes to both Qwen2.5-VL (for meaning) and the VAE (for appearance), so it handles both semantic edits like restyling and precise local edits, including editing text in the image.'