Krea 2 Raw
'The undistilled base checkpoint of Krea 2, a 12.8B single-stream MMDiT conditioned on stacked Qwen3-VL-4B hidden states and working in the Qwen-Image VAE latent space. Krea ships it as the checkpoint to fine-tune, not to sample from.'
- weights
- krea/Krea-2-Raw ↗
- org
- Krea
- modality
- image
- tasks
- text-to-image
- license
- Krea 2 Community License
- released
- 2026-06-22
- total params
- 17.38B
- model.arch
- krea2
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Krea 2 SingleStreamDiT (Raw) SingleStreamDiT (ai-toolkit krea2/src/mmdit.py) The original-layout single file that ai-toolkit loads. The 321.6M fp32 params are the norm scales and the per-block modulation offsets. The repo also ships a Diffusers copy in transformer/ (Krea2Transformer2DModel, same parameter count) that ai-toolkit does not use. | 12.82B 12,820,073,036 | 26.28 GB | bf16+fp32 | yes |
| Text encoder | Qwen3-VL-4B-Instruct transformers.Qwen3VLModel Includes the vision tower. ai-toolkit loads Qwen/Qwen3-VL-4B-Instruct instead; every tensor is identical to this copy (compared tensor by tensor). The vision tower is dropped after loading for text-to-image training. | 4.44B 4,437,815,808 | 8.88 GB | bf16 | no |
| Tokenizer | Qwen2 tokenizer transformers.Qwen2Tokenizer | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage The Qwen-Image VAE upcast to fp32 (every tensor equals Qwen/Qwen-Image vae/ once cast back to bf16). ai-toolkit loads the bf16 original from Qwen/Qwen-Image. | 126.9M 126,892,531 | 508 MB | fp32 | no |
| total | 17.38B | 35.67 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- The Qwen-Image VAE is a video VAE (Wan 2.1 lineage). Images are encoded as a single frame and latents are normalized per channel with the latents_mean / latents_std from its config.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 1280×1280 | 16×1×160×160 | 6,400 | top of the time-shift range |
| 2048×2048 | 16×1×256×256 | 16,384 | Turbo README example |
Architecture
- Type
- Single-stream MMDiT: text and image tokens run through the same blocks
- Blocks
- 28
- Hidden size
- 6144 (48 query heads × 128, 12 KV heads)
- FFN size
- 16384 (SwiGLU)
- Text conditioning
- Hidden states from 12 Qwen3-VL layers (2, 5, … 35), 2560-d each, fused by a 4-block text transformer and a learned layer mix, then projected to 6144
- Timestep modulation
- One shared projection for all blocks, plus a learned offset per block
- Norm / position
- QK RMSNorm, sigmoid-gated attention output, 3-axis RoPE (32/48/48 dims, θ 1000)
- Objective
- Rectified flow (target noise − clean), exponential time shift μ 0.5 → 1.15 over 256–6400 image tokens
- Distillation
- None (model_index is_distilled: false)
- Recommended sampling
- 52 steps, CFG 3.5 (Krea README)
In AI Toolkit
- model.arch
- krea2
- UI label
- Krea 2 (raw) (image)
- model.name_or_path
- krea/Krea-2-Raw
- source
- extensions_built_in/diffusion_models/krea2/krea2.py
- extensions_built_in/diffusion_models/krea2/src/mmdit.py
- extensions_built_in/diffusion_models/krea2/src/pipeline.py
- extensions_built_in/diffusion_models/krea2/src/text_encoder.py
- toolkit/models/v2/vae/qwen_image.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- timestep_type
- linear
- network.conv
- disabled (linear LoRA only)
Specifics
- Gated repo
- Accept the Krea 2 Community License on the Hub and set a Hugging Face token before training.
- Transformer file
- Downloads raw.safetensors from name_or_path (the file name comes from the repo name’s last segment). Override with model_kwargs.checkpoint_filename, or point name_or_path at a local .safetensors file or folder.
- Text encoder source
- Loads Qwen/Qwen3-VL-4B-Instruct (model_kwargs.text_encoder_path), not text_encoder/ from the Krea repo. The vision tower is dropped. Prompts are wrapped in Krea’s fixed system template, and the 34 template tokens are sliced off the output.
- VAE source
- Loads vae/ from Qwen/Qwen-Image (model_kwargs.vae_path), not the fp32 copy in the Krea repo.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Timesteps
- Training uses flowmatch with Krea’s resolution-dependent shift (use_dynamic_shifting, μ 0.5 at 256 px to 1.15 at 1280 px).
- Prompt length
- Capped at 512 tokens (model_kwargs.max_text_length).
- Quantization
- first, tmlp*, tproj*, txtmlp*, txtfusion.projector and last* stay in full precision.
- LoRA target
- SingleStreamDiT
- Saving
- LoRA keys are saved with the diffusion_model. prefix (ComfyUI layout). Full fine-tunes save a single safetensors of the SingleStreamDiT state dict.
- Sampling
- Built-in Euler sampler with the same time shift. CFG is zero-normalized: the sampler uses guidance_scale − 1, so guidance_scale 1 turns CFG off. low_vram tiles the VAE decode.
- Metadata base version
- krea2
Example config
not verified
job: extensionconfig: name: "my_krea2_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: bf16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "krea/Krea-2-Raw" arch: "krea2" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 guidance_scale: 4 sample_steps: 25 prompts: - "a fox walking in the snow"Links
PRXPixel
Photoroom’s 7B pixel-space variant of PRX. There is no VAE: the transformer denoises raw RGB in 16×16 pixel patches, predicts the clean image rather than a velocity, and starts from noise with std 2.0. Text comes from a Qwen3-VL text tower.
Krea 2 Turbo
'The post-trained, step-distilled release of Krea 2: the same 12.8B single-stream MMDiT as Raw, sampled in about 8 steps without CFG. ai-toolkit trains it through a de-distill training adapter so LoRAs keep the fast sampling.'