Docs
AI ToolkitModels

Krea 2 Raw

'The undistilled base checkpoint of Krea 2, a 12.8B single-stream MMDiT conditioned on stacked Qwen3-VL-4B hidden states and working in the Qwen-Image VAE latent space. Krea ships it as the checkpoint to fine-tune, not to sample from.'

org
Krea
modality
image
tasks
text-to-image
license
Krea 2 Community License
released
2026-06-22
total params
17.38B
model.arch
krea2

Components

rolemodelparamssizedtypetrained
Transformer
Krea 2 SingleStreamDiT (Raw)
SingleStreamDiT (ai-toolkit krea2/src/mmdit.py)

The original-layout single file that ai-toolkit loads. The 321.6M fp32 params are the norm scales and the per-block modulation offsets. The repo also ships a Diffusers copy in transformer/ (Krea2Transformer2DModel, same parameter count) that ai-toolkit does not use.

12.82B
12,820,073,036
26.28 GBbf16+fp32yes
Text encoder
Qwen3-VL-4B-Instruct
transformers.Qwen3VLModel

Includes the vision tower. ai-toolkit loads Qwen/Qwen3-VL-4B-Instruct instead; every tensor is identical to this copy (compared tensor by tensor). The vision tower is dropped after loading for text-to-image training.

4.44B
4,437,815,808
8.88 GBbf16no
Tokenizer
Qwen2 tokenizer
transformers.Qwen2Tokenizer
————
VAE
Qwen-Image VAE
diffusers.AutoencoderKLQwenImage

The Qwen-Image VAE upcast to fp32 (every tensor equals Qwen/Qwen-Image vae/ once cast back to bf16). ai-toolkit loads the bf16 original from Qwen/Qwen-Image.

126.9M
126,892,531
508 MBfp32no
total17.38B35.67 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Qwen-Image VAE
pixels per token
16×16
notes
The Qwen-Image VAE is a video VAE (Wan 2.1 lineage). Images are encoded as a single frame and latents are normalized per channel with the latents_mean / latents_std from its config.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
1280×128016×1×160×1606,400top of the time-shift range
2048×204816×1×256×25616,384Turbo README example

Architecture

Type
Single-stream MMDiT: text and image tokens run through the same blocks
Blocks
28
Hidden size
6144 (48 query heads × 128, 12 KV heads)
FFN size
16384 (SwiGLU)
Text conditioning
Hidden states from 12 Qwen3-VL layers (2, 5, … 35), 2560-d each, fused by a 4-block text transformer and a learned layer mix, then projected to 6144
Timestep modulation
One shared projection for all blocks, plus a learned offset per block
Norm / position
QK RMSNorm, sigmoid-gated attention output, 3-axis RoPE (32/48/48 dims, θ 1000)
Objective
Rectified flow (target noise − clean), exponential time shift μ 0.5 → 1.15 over 256–6400 image tokens
Distillation
None (model_index is_distilled: false)
Recommended sampling
52 steps, CFG 3.5 (Krea README)

In AI Toolkit

model.arch
krea2
UI label
Krea 2 (raw) (image)
model.name_or_path
krea/Krea-2-Raw
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
timestep_type
linear
network.conv
disabled (linear LoRA only)

Specifics

Gated repo
Accept the Krea 2 Community License on the Hub and set a Hugging Face token before training.
Transformer file
Downloads raw.safetensors from name_or_path (the file name comes from the repo name’s last segment). Override with model_kwargs.checkpoint_filename, or point name_or_path at a local .safetensors file or folder.
Text encoder source
Loads Qwen/Qwen3-VL-4B-Instruct (model_kwargs.text_encoder_path), not text_encoder/ from the Krea repo. The vision tower is dropped. Prompts are wrapped in Krea’s fixed system template, and the 34 template tokens are sliced off the output.
VAE source
Loads vae/ from Qwen/Qwen-Image (model_kwargs.vae_path), not the fp32 copy in the Krea repo.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Timesteps
Training uses flowmatch with Krea’s resolution-dependent shift (use_dynamic_shifting, μ 0.5 at 256 px to 1.15 at 1280 px).
Prompt length
Capped at 512 tokens (model_kwargs.max_text_length).
Quantization
first, tmlp*, tproj*, txtmlp*, txtfusion.projector and last* stay in full precision.
LoRA target
SingleStreamDiT
Saving
LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a single safetensors of the SingleStreamDiT state dict.
Sampling
Built-in Euler sampler with the same time shift. CFG is zero-normalized: the sampler uses guidance_scale − 1, so guidance_scale 1 turns CFG off. low_vram tiles the VAE decode.
Metadata base version
krea2

Example config

not verified

job: extensionconfig:  name: "my_krea2_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: bf16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "krea/Krea-2-Raw"        arch: "krea2"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        guidance_scale: 4        sample_steps: 25        prompts:          - "a fox walking in the snow"

On this page