Krea 2 Turbo (edit training)
'Krea 2 Turbo loaded for reference-image (edit) training. The weights are plain Krea 2 Turbo text-to-image; ai-toolkit feeds control images to Qwen3-VL and appends them as clean latents so a LoRA learns to edit, restyle or follow a reference, trained through the de-distill adapter.'
- weights
- krea/Krea-2-Turbo ↗
- org
- Krea
- modality
- image
- tasks
- text-to-image
- license
- Krea 2 Community License
- released
- 2026-06-22
- total params
- 17.38B
- model.arch
- krea2:o_edit_turbo
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Krea 2 SingleStreamDiT (Turbo) SingleStreamDiT (ai-toolkit krea2/src/mmdit.py) The original-layout single file that ai-toolkit loads. The 321.6M fp32 params are the norm scales and the per-block modulation offsets. The repo also ships a Diffusers copy in transformer/ (Krea2Transformer2DModel, same parameter count) that ai-toolkit does not use. | 12.82B 12,820,073,036 | 26.28 GB | bf16+fp32 | yes |
| Text encoder | Qwen3-VL-4B-Instruct transformers.Qwen3VLModel Includes the vision tower. ai-toolkit loads Qwen/Qwen3-VL-4B-Instruct instead; every tensor is identical to this copy (compared tensor by tensor). Edit training keeps the vision tower, since reference images are encoded with the prompt. | 4.44B 4,437,815,808 | 8.88 GB | bf16 | no |
| Training adapter | Krea 2 Turbo training adapter v1 (LoRA, rank 32)adapter A de-distill LoRA by Ostris, trained at lr 1e-5 on images generated by Turbo. It is merged into the transformer while your LoRA trains and inverted while sampling, so the new LoRA runs on the distilled model at Turbo speed. | 114.3M 114,262,016 | 229 MB | bf16 | no |
| Tokenizer | Qwen2 tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage The Qwen-Image VAE upcast to fp32 (every tensor equals Qwen/Qwen-Image vae/ once cast back to bf16). ai-toolkit loads the bf16 original from Qwen/Qwen-Image. | 126.9M 126,892,531 | 508 MB | fp32 | no |
| total | 17.38B | 35.67 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- The Qwen-Image VAE is a video VAE (Wan 2.1 lineage). Images are encoded as a single frame and latents are normalized per channel with the latents_mean / latents_std from its config. In edit training each reference image is encoded to its own latent and appended after the target tokens.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 1280×1280 | 16×1×160×160 | 6,400 | top of the time-shift range |
| 2048×2048 | 16×1×256×256 | 16,384 | Turbo README example |
Architecture
- Type
- Single-stream MMDiT: text and image tokens run through the same blocks
- Blocks
- 28
- Hidden size
- 6144 (48 query heads × 128, 12 KV heads)
- FFN size
- 16384 (SwiGLU)
- Text conditioning
- Hidden states from 12 Qwen3-VL layers (2, 5, … 35), 2560-d each, fused by a 4-block text transformer and a learned layer mix, then projected to 6144
- Timestep modulation
- One shared projection for all blocks, plus a learned offset per block
- Norm / position
- QK RMSNorm, sigmoid-gated attention output, 3-axis RoPE (32/48/48 dims, θ 1000)
- Objective
- Rectified flow (target noise − clean), exponential time shift μ 0.5 → 1.15 over 256–6400 image tokens
- Reference images
- None in the weights. Edit training adds them as clean latent tokens at t = 0, each on its own RoPE frame index
- Distillation
- Step-distilled, post-trained from Raw (model_index is_distilled: true)
- Recommended sampling
- 8 steps, CFG off, fixed μ 1.15 (Krea README)
In AI Toolkit
- model.arch
- krea2:o_edit_turbo
- UI label
- Krea 2 Turbo (w/ Training Adapter) [Edit Training] (experimental)
- model.name_or_path
- krea/Krea-2-Turbo
- source
- extensions_built_in/diffusion_models/krea2/krea2.py
- extensions_built_in/diffusion_models/krea2/src/mmdit.py
- extensions_built_in/diffusion_models/krea2/src/pipeline.py
- extensions_built_in/diffusion_models/krea2/src/text_encoder.py
- toolkit/models/v2/vae/qwen_image.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.multi_control_paths, sample.multi_ctrl_imgs, model.low_vram, model.layer_offloading, model.assistant_lora_path, model.qie.match_target_res, model.model_kwargs.kv_cache
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- timestep_type
- linear
- assistant_lora_path
- ostris/krea2_turbo_training_adapter/krea2_turbo_training_adapter_v1.safetensors
- sample guidance_scale / sample_steps
- 1 / 8
- model_kwargs
- { edit: true, match_target_res: true, kv_cache: true }
- train.unload_text_encoder
- false (section hidden)
- network.conv
- disabled (linear LoRA only)
Specifics
- Gated repo
- Accept the Krea 2 Community License on the Hub and set a Hugging Face token before training.
- Training adapter
- assistant_lora_path is downloaded (repo/file on the Hub, or a local path), its rank read from the weights, and merged into the transformer at weight 1.0 before quantization. While sampling it is applied at −1.0, cancelling the merge, so previews show Turbo plus your LoRA. The adapter README warns that long runs still wear the distillation down.
- Quantization with the adapter
- When an adapter is loaded, qtype qfloat8 is switched to float8.
- Transformer file
- Downloads turbo.safetensors from name_or_path (the file name comes from the repo name’s last segment). Override with model_kwargs.checkpoint_filename, or point name_or_path at a local .safetensors file or folder.
- Edit training
- Not a separate model. The :o_edit variants load the same Krea 2 weights with model_kwargs.edit: true, and the backend strips the suffix and runs arch krea2. Each training pair is a target image, its caption or instruction, and up to three control images (control_path_1 to _3). The LoRA learns to produce the target given the references, for edits, style transfer or subject reference.
- Reference path 1: Qwen3-VL
- Control images are downscaled (never up) to fit 384² pixels (model_kwargs.vlm_max_pixels) and put in the user message as “Picture N:” vision placeholders ahead of the prompt, so the text embeddings see them. Cached text embeddings are keyed on the control paths.
- Reference path 2: latents
- Each control image keeps its aspect ratio, is resized to the target’s pixel area with match_target_res (otherwise capped at 1 MP, model_kwargs.control_image_max_pixels), snapped to 16 px, VAE-encoded, and appended after the noisy target tokens. They get t = 0 modulation, RoPE frame index 1, 2, 3, and are left out of the loss.
- kv_cache
- Trains with an asymmetric mask: reference tokens attend only to each other. Their K/V then do not change across steps, so inference can compute them once. The base model was trained with full attention, so a LoRA must be trained with kv_cache on to use kv-cached inference. Sampling previews use the cache when it is on.
- Running the LoRA
- Needs the ComfyUI-Krea2-Ostris-Edit custom nodes or the ostris/Krea2OstrisEdit diffusers community pipeline. Stock Krea 2 pipelines ignore reference images.
- Text encoder source
- Loads Qwen/Qwen3-VL-4B-Instruct (model_kwargs.text_encoder_path), not text_encoder/ from the Krea repo. The vision tower stays, with its Conv3d patch embed swapped for an equivalent linear layer. The UI keeps the text encoder loaded (unload_text_encoder off). Prompts are wrapped in Krea’s fixed system template, and the 34 template tokens are sliced off the output.
- VAE source
- Loads vae/ from Qwen/Qwen-Image (model_kwargs.vae_path), not the fp32 copy in the Krea repo.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Timesteps
- Training uses flowmatch with Krea’s resolution-dependent shift (use_dynamic_shifting, μ 0.5 at 256 px to 1.15 at 1280 px).
- Turbo schedule
- The sampler uses the resolution-dependent μ unless model_kwargs.schedule_mu is set. The Turbo README samples with a fixed μ 1.15; the UI does not set it.
- Prompt length
- Capped at 512 tokens (model_kwargs.max_text_length).
- Quantization
- first, tmlp*, tproj*, txtmlp*, txtfusion.projector and last* stay in full precision.
- LoRA target
- SingleStreamDiT
- Saving
- LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a single safetensors of the SingleStreamDiT state dict.
- Sampling
- Built-in Euler sampler with the same time shift. CFG is zero-normalized: the sampler uses guidance_scale − 1, so guidance_scale 1 turns CFG off. low_vram tiles the VAE decode.
- Metadata base version
- krea2
Example config
not verified
job: extensionconfig: name: "my_krea2_turbo_edit_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/targets" control_path_1: "/path/to/references" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "krea/Krea-2-Turbo" arch: "krea2:o_edit_turbo" quantize: true quantize_te: true low_vram: true model_kwargs: edit: true match_target_res: true kv_cache: true assistant_lora_path: "ostris/krea2_turbo_training_adapter/krea2_turbo_training_adapter_v1.safetensors" sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 1 sample_steps: 8 samples: - prompt: "make it a watercolor painting" ctrl_img_1: "/path/to/reference.jpg"