Docs
AI ToolkitModels

Krea 2 Turbo (edit training)

'Krea 2 Turbo loaded for reference-image (edit) training. The weights are plain Krea 2 Turbo text-to-image; ai-toolkit feeds control images to Qwen3-VL and appends them as clean latents so a LoRA learns to edit, restyle or follow a reference, trained through the de-distill adapter.'

org
Krea
modality
image
tasks
text-to-image
license
Krea 2 Community License
released
2026-06-22
total params
17.38B
model.arch
krea2:o_edit_turbo

Components

rolemodelparamssizedtypetrained
Transformer
Krea 2 SingleStreamDiT (Turbo)
SingleStreamDiT (ai-toolkit krea2/src/mmdit.py)

The original-layout single file that ai-toolkit loads. The 321.6M fp32 params are the norm scales and the per-block modulation offsets. The repo also ships a Diffusers copy in transformer/ (Krea2Transformer2DModel, same parameter count) that ai-toolkit does not use.

12.82B
12,820,073,036
26.28 GBbf16+fp32yes
Text encoder
Qwen3-VL-4B-Instruct
transformers.Qwen3VLModel

Includes the vision tower. ai-toolkit loads Qwen/Qwen3-VL-4B-Instruct instead; every tensor is identical to this copy (compared tensor by tensor). Edit training keeps the vision tower, since reference images are encoded with the prompt.

4.44B
4,437,815,808
8.88 GBbf16no
Training adapter
Krea 2 Turbo training adapter v1 (LoRA, rank 32)adapter

A de-distill LoRA by Ostris, trained at lr 1e-5 on images generated by Turbo. It is merged into the transformer while your LoRA trains and inverted while sampling, so the new LoRA runs on the distilled model at Turbo speed.

114.3M
114,262,016
229 MBbf16no
Tokenizer
Qwen2 tokenizer
transformers.Qwen2Tokenizer
————
VAE
Qwen-Image VAE
diffusers.AutoencoderKLQwenImage

The Qwen-Image VAE upcast to fp32 (every tensor equals Qwen/Qwen-Image vae/ once cast back to bf16). ai-toolkit loads the bf16 original from Qwen/Qwen-Image.

126.9M
126,892,531
508 MBfp32no
total17.38B35.67 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Qwen-Image VAE
pixels per token
16×16
notes
The Qwen-Image VAE is a video VAE (Wan 2.1 lineage). Images are encoded as a single frame and latents are normalized per channel with the latents_mean / latents_std from its config. In edit training each reference image is encoded to its own latent and appended after the target tokens.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
1280×128016×1×160×1606,400top of the time-shift range
2048×204816×1×256×25616,384Turbo README example

Architecture

Type
Single-stream MMDiT: text and image tokens run through the same blocks
Blocks
28
Hidden size
6144 (48 query heads × 128, 12 KV heads)
FFN size
16384 (SwiGLU)
Text conditioning
Hidden states from 12 Qwen3-VL layers (2, 5, … 35), 2560-d each, fused by a 4-block text transformer and a learned layer mix, then projected to 6144
Timestep modulation
One shared projection for all blocks, plus a learned offset per block
Norm / position
QK RMSNorm, sigmoid-gated attention output, 3-axis RoPE (32/48/48 dims, θ 1000)
Objective
Rectified flow (target noise − clean), exponential time shift μ 0.5 → 1.15 over 256–6400 image tokens
Reference images
None in the weights. Edit training adds them as clean latent tokens at t = 0, each on its own RoPE frame index
Distillation
Step-distilled, post-trained from Raw (model_index is_distilled: true)
Recommended sampling
8 steps, CFG off, fixed μ 1.15 (Krea README)

In AI Toolkit

model.arch
krea2:o_edit_turbo
UI label
Krea 2 Turbo (w/ Training Adapter) [Edit Training] (experimental)
model.name_or_path
krea/Krea-2-Turbo
extra UI sections
datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.assistant_lora_path, model.qie.match_target_res, model.model_kwargs.kv_cache

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
timestep_type
linear
assistant_lora_path
ostris/krea2_turbo_training_adapter/krea2_turbo_training_adapter_v1.safetensors
sample guidance_scale / sample_steps
1 / 8
model_kwargs
{ edit: true, match_target_res: true, kv_cache: true }
train.unload_text_encoder
false (section hidden)
network.conv
disabled (linear LoRA only)

Specifics

Gated repo
Accept the Krea 2 Community License on the Hub and set a Hugging Face token before training.
Training adapter
assistant_lora_path is downloaded (repo/file on the Hub, or a local path), its rank read from the weights, and merged into the transformer at weight 1.0 before quantization. While sampling it is applied at −1.0, cancelling the merge, so previews show Turbo plus your LoRA. The adapter README warns that long runs still wear the distillation down.
Quantization with the adapter
When an adapter is loaded, qtype qfloat8 is switched to float8.
Transformer file
Downloads turbo.safetensors from name_or_path (the file name comes from the repo name’s last segment). Override with model_kwargs.checkpoint_filename, or point name_or_path at a local .safetensors file or folder.
Edit training
Not a separate model. The :o_edit variants load the same Krea 2 weights with model_kwargs.edit: true, and the backend strips the suffix and runs arch krea2. Each training pair is a target image, its caption or instruction, and up to three control images (control_path_1 to _3). The LoRA learns to produce the target given the references, for edits, style transfer or subject reference.
Reference path 1: Qwen3-VL
Control images are downscaled (never up) to fit 384² pixels (model_kwargs.vlm_max_pixels) and put in the user message as “Picture N:” vision placeholders ahead of the prompt, so the text embeddings see them. Cached text embeddings are keyed on the control paths.
Reference path 2: latents
Each control image keeps its aspect ratio, is resized to the target’s pixel area with match_target_res (otherwise capped at 1 MP, model_kwargs.control_image_max_pixels), snapped to 16 px, VAE-encoded, and appended after the noisy target tokens. They get t = 0 modulation, RoPE frame index 1, 2, 3, and are left out of the loss.
kv_cache
Trains with an asymmetric mask: reference tokens attend only to each other. Their K/V then do not change across steps, so inference can compute them once. The base model was trained with full attention, so a LoRA must be trained with kv_cache on to use kv-cached inference. Sampling previews use the cache when it is on.
Running the LoRA
Needs the ComfyUI-Krea2-Ostris-Edit custom nodes or the ostris/Krea2OstrisEdit diffusers community pipeline. Stock Krea 2 pipelines ignore reference images.
Text encoder source
Loads Qwen/Qwen3-VL-4B-Instruct (model_kwargs.text_encoder_path), not text_encoder/ from the Krea repo. The vision tower stays, with its Conv3d patch embed swapped for an equivalent linear layer. The UI keeps the text encoder loaded (unload_text_encoder off). Prompts are wrapped in Krea’s fixed system template, and the 34 template tokens are sliced off the output.
VAE source
Loads vae/ from Qwen/Qwen-Image (model_kwargs.vae_path), not the fp32 copy in the Krea repo.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Timesteps
Training uses flowmatch with Krea’s resolution-dependent shift (use_dynamic_shifting, μ 0.5 at 256 px to 1.15 at 1280 px).
Turbo schedule
The sampler uses the resolution-dependent μ unless model_kwargs.schedule_mu is set. The Turbo README samples with a fixed μ 1.15; the UI does not set it.
Prompt length
Capped at 512 tokens (model_kwargs.max_text_length).
Quantization
first, tmlp*, tproj*, txtmlp*, txtfusion.projector and last* stay in full precision.
LoRA target
SingleStreamDiT
Saving
LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a single safetensors of the SingleStreamDiT state dict.
Sampling
Built-in Euler sampler with the same time shift. CFG is zero-normalized: the sampler uses guidance_scale − 1, so guidance_scale 1 turns CFG off. low_vram tiles the VAE decode.
Metadata base version
krea2

Example config

not verified

job: extensionconfig:  name: "my_krea2_turbo_edit_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/targets"          control_path_1: "/path/to/references"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "krea/Krea-2-Turbo"        arch: "krea2:o_edit_turbo"        quantize: true        quantize_te: true        low_vram: true        model_kwargs:          edit: true          match_target_res: true          kv_cache: true        assistant_lora_path: "ostris/krea2_turbo_training_adapter/krea2_turbo_training_adapter_v1.safetensors"      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 1        sample_steps: 8        samples:          - prompt: "make it a watercolor painting"            ctrl_img_1: "/path/to/reference.jpg"

On this page