Docs
AI ToolkitModels

FLUX.1 Kontext [dev]

The instruction-editing member of FLUX.1. Same 12B transformer shape, text encoders and VAE as FLUX.1 [dev]; the input image goes in as a second set of tokens next to the image being generated. Trained in ai-toolkit on before/after pairs.

org
Black Forest Labs
modality
image
tasks
image-editing · text-to-image
license
FLUX.1 [dev] Non-Commercial License
released
2025-06-26not verified
native output
1024×1024, output size follows the input imagenot verified
total params
16.87B
model.arch
flux_kontext

Components

rolemodelparamssizedtypetrained
Transformer
FLUX.1 Kontext [dev] DiT
diffusers.FluxTransformer2DModel

Same parameter count and config as FLUX.1 [dev] (in_channels 64, guidance embeds). Also shipped as flux1-kontext-dev.safetensors in the original layout.

11.90B
11,901,408,320
23.80 GBbf16yes
Text encoder
CLIP ViT-L/14 text model
transformers.CLIPTextModel

Only the pooled output (768-d) is used.

123.1M
123,060,480
246 MBbf16no
Tokenizer
CLIP BPE (49k vocab)
transformers.CLIPTokenizer
————
Text encoder 2
T5 v1.1 XXL (encoder only)
transformers.T5EncoderModel

512 tokens in ai-toolkit.

4.76B
4,762,310,656
9.52 GBbf16no
Tokenizer 2
T5 SentencePiece (32k vocab)
transformers.T5TokenizerFast
————
VAE
FLUX.1 VAE (16 channels)
diffusers.AutoencoderKL

Encodes both the target and the input (control) image.

83.8M
83,819,683
168 MBbf16no
total16.87B33.74 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
FLUX.1 VAE
pixels per token
16×16
notes
Same packing as FLUX.1. The control image is encoded the same way and appended as extra tokens, so with one control image at the target size the sequence is twice the token counts below.
inputlatent (c×t×h×w)tokens
512×51216×1×64×641,024target only; ×2 with control
1024×102416×1×128×1284,096target only; ×2 with control
1344×76816×1×96×1684,032landscape, target only

Architecture

Blocks
19 double-stream + 38 single-stream
Hidden size
3072 (24 heads × 128)
FFN size
12288
Text conditioning
T5 hidden states (4096-d) joined into the sequence; CLIP pooled vector (768-d) added to the timestep embedding
Image conditioning
Input image latents appended as tokens; their RoPE ids are marked with 1 on the first axis
Guidance
Guidance-distilled, guidance embedding
Objective
Rectified flow, resolution-dependent shift (0.5 at 256 tokens to 1.15 at 4096)
Norm / position
QK RMSNorm, 3-axis RoPE (16, 56, 56)

In AI Toolkit

model.arch
flux_kontext
UI label
FLUX.1-Kontext-dev (instruction)
model.name_or_path
black-forest-labs/FLUX.1-Kontext-dev
extra UI sections
datasets.control_path, sample.ctrl_img

UI defaults

quantize / quantize_te
true / true (qfloat8)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
network.conv
disabled (linear LoRA only)

Specifics

Gated
Accept the license on the Hub and set HF_TOKEN before the first run.
Training data
Pairs: datasets.control_path holds the input images, with the same file names as the edited targets in folder_path. Captions are the edit instructions.not verified
Control image
Resized (bilinear) to the target crop size, VAE-encoded, then packed and appended after the target tokens. Only the target tokens are kept from the prediction, so the loss covers only them.
Resolution
Buckets snap to multiples of 16.
Loading
Transformer from name_or_path. T5, CLIP and VAE come from extras_name_or_path, which defaults to name_or_path, or from name_or_path if it is a local folder with text_encoder/.
Quantization
Transformer uses qtype, T5 uses qtype_te. CLIP is never quantized. layer_offloading is supported.
low_vram
Keeps the text encoders on the CPU and moves them to the GPU only to encode prompts.
Guidance during training
The guidance embedding gets train.cfg_scale (default 1.0).
LoRA target
FluxTransformer2DModel
Saving
LoRAs use Diffusers/PEFT keys (transformer.…lora_A / lora_B), alpha set to the rank. Full fine-tunes save the transformer/ folder plus aitk_meta.yaml.
Sampling
Every sample prompt needs --ctrl_img. The control image is resized to the sample width × height, which are floored to multiples of 16.
Metadata base version
flux.1_kontext

Example config

not verified

job: extensionconfig:  name: "my_first_flux_kontext_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/edited/images"          control_path: "/path/to/input/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          cache_latents_to_disk: true          resolution: [512, 768]      train:        batch_size: 1        steps: 3000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "black-forest-labs/FLUX.1-Kontext-dev"        arch: "flux_kontext"        quantize: true        quantize_te: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 20        prompts:          - "make the person smile --ctrl_img /path/to/input/images/person1.jpg"

On this page