AI Toolkit
Models
Every model AI Toolkit can train, with its parts, sizes, latent space and training specifics.
image
| model | model.arch | total params | latent | license |
|---|---|---|---|---|
| Anima Base v1.0 | anima | 2.81B | 8× · 16ch | CircleStone Labs Non-Commercial License v1.0 |
| FLUX.1 [dev] | flux | 16.87B | 8× · 16ch | FLUX.1 [dev] Non-Commercial License |
| Flex.1-alpha | flex1 | 13.13B | 8× · 16ch | Apache 2.0 |
| Chroma1-Base | chroma | 13.75B | 8× · 16ch | Apache 2.0 |
| Lumina-Image 2.0 | lumina2 | 5.31B | 8× · 16ch | Apache 2.0 |
| Qwen-Image | qwen_image | 28.85B | 8× · 16ch | Apache 2.0 |
| Qwen-Image-2512 | qwen_image:2512 | 28.85B | 8× · 16ch | Apache 2.0 |
| Qwen-Image-2.1 | qwen_image_2 | 16.22B | 16× · 64ch | Qwen Research License |
| Ming-Image 0.1 Design | ming_image | 24.86B | 8× · 16ch | MIT |
| HiDream-I1 Full | hidream | 30.80B | 8× · 16ch | MIT |
| Stable Diffusion XL 1.0 Base | sdxl | 3.47B | 8× · 4ch | CreativeML Open RAIL++-M |
| Stable Diffusion 1.5 | sd15 | 1.07B | 8× · 4ch | CreativeML Open RAIL-M |
| OmniGen2 | omnigen2 | 7.81B | 8× · 16ch | Apache 2.0 |
| FLUX.2 [dev] | flux2 | 56.32B | 8× · 32ch | FLUX Non-Commercial License |
| Z-Image Turbo | zimage:turbo | 10.26B | 8× · 16ch | Apache 2.0 |
| Z-Image | zimage | 10.26B | 8× · 16ch | Apache 2.0 |
| Z-Image De-Turbo | zimage:deturbo | 10.26B | 8× · 16ch | Apache 2.0 |
| FLUX.2 [klein] 4B Base | flux2_klein_4b | 7.98B | 8× · 32ch | Apache 2.0 |
| ERNIE-Image | ernie_image | 11.97B | 8× · 32ch | Apache 2.0 |
| FLUX.2 [klein] 9B Base | flux2_klein_9b | 17.35B | 8× · 32ch | FLUX Non-Commercial License |
| Nucleus-Image | nucleus_image | 25.82B | 8× · 16ch | Apache 2.0 |
| HiDream-O1-Image | hidream_o1 | 8.80B | pixel space | MIT |
| Z-Image L2P (pixel space) | zimage_l2p | 10.19B | pixel space | Apache 2.0 |
| PRXPixel | prx_pixel | 8.72B | pixel space | Apache 2.0 |
| Krea 2 Raw | krea2 | 17.38B | 8× · 16ch | Krea 2 Community License |
| Krea 2 Turbo | krea2:turbo | 17.38B | 8× · 16ch | Krea 2 Community License |
| Mage-Flow Base | mageflow | 8.73B | 16× · 128ch | MIT |
| Boogu-Image 0.1 Base | boogu_image | 19.14B | 8× · 16ch | Apache 2.0 |
| Flex.2-preview | flex2 | 13.13B | 8× · 16ch | Apache 2.0 |
instruction
| model | model.arch | total params | latent | license |
|---|---|---|---|---|
| FLUX.1 Kontext [dev] | flux_kontext | 16.87B | 8× · 16ch | FLUX.1 [dev] Non-Commercial License |
| Qwen-Image-Edit | qwen_image_edit | 28.85B | 8× · 16ch | Apache 2.0 |
| Qwen-Image-Edit-2509 | qwen_image_edit_plus | 28.85B | 8× · 16ch | Apache 2.0 |
| Qwen-Image-Edit-2511 | qwen_image_edit_plus:2511 | 28.85B | 8× · 16ch | Apache 2.0 |
| HiDream-E1.1 | hidream_e1 | 30.80B | 8× · 16ch | MIT |
| Mage-Flow Edit Base | mageflow_edit | 8.73B | 16× · 128ch | MIT |
| Boogu-Image 0.1 Edit | boogu_image_edit | 19.14B | 8× · 16ch | Apache 2.0 |
video
| model | model.arch | total params | latent | license |
|---|---|---|---|---|
| Wan 2.1 T2V 1.3B | wan21:1b | 7.23B | 8× · 4×t · 16ch | Apache 2.0 |
| Wan 2.1 I2V 14B 480P | wan21_i2v:14b480p | 22.83B | 8× · 4×t · 16ch | Apache 2.0 |
| Wan 2.1 I2V 14B 720P | wan21_i2v:14b | 22.83B | 8× · 4×t · 16ch | Apache 2.0 |
| Wan 2.1 T2V 14B | wan21:14b | 20.10B | 8× · 4×t · 16ch | Apache 2.0 |
| Wan 2.2 T2V A14B | wan22_14b:t2v | 34.38B | 8× · 4×t · 16ch | Apache 2.0 |
| Wan 2.2 I2V A14B | wan22_14b_i2v | 34.39B | 8× · 4×t · 16ch | Apache 2.0 |
| Wan 2.2 TI2V 5B | wan22_5b | 11.39B | 16× · 4×t · 48ch | Apache 2.0 |
| MiniMax-H3 (FL2VA) | minimax_h3 | 48.62B | 16× · 4×t · 24ch + audio 800× · 32ch | MiniMax-H3 Community License |
| MiniMax-H3 Ref2VA | minimax_h3_ref2va | 48.62B | 16× · 4×t · 24ch + audio 800× · 32ch | MiniMax-H3 Community License |
| LTX-2 19B | ltx2 | 33.83B | 32× · 8×t · 128ch + audio 640× · 8ch | LTX-2 Community License |
| LTX-2.3 22B | ltx2.3 | 35.26B | 32× · 8×t · 128ch + audio 640× · 8ch | LTX-2 Community License |
| LTX-2.5 22B | ltx2.5 | 34.98B | 32× · 8×t · 128ch + audio 640× · 8ch | LTX-2.x Community License |
audio
| model | model.arch | total params | latent | license |
|---|---|---|---|---|
| ACE-Step 1.5 XL | ace_step_15_xl | 7.61B | audio 1,920× · 64ch | MIT |
| YuE2 3B | yue2 | 5.10B | audio 1,920× · 64ch | CC BY-NC 4.0 |
| ACE-Step 1.5 | ace_step_15 | 5.01B | audio 1,920× · 64ch | MIT |
llm
| model | model.arch | total params | latent | license |
|---|---|---|---|---|
| Qwen2.5-Omni 7B (thinker) | qwen25_omni | 8.93B | — | Apache 2.0 |
experimental
| model | model.arch | total params | latent | license |
|---|---|---|---|---|
| Zeta-Chroma | zeta_chroma | 10.52B | pixel space | Apache 2.0 |
| Ideogram 4 | ideogram4 | 17.51B | 8× · 32ch | Ideogram 4 Non-Commercial |
| Krea 2 Raw (edit training) | krea2:o_edit | 17.38B | 8× · 16ch | Krea 2 Community License |
| Krea 2 Turbo (edit training) | krea2:o_edit_turbo | 17.38B | 8× · 16ch | Krea 2 Community License |
Docs
Documentation for Ostris tools and projects.
Anima Base v1.0
A 2B anime and illustration text-to-image model built on the Cosmos-Predict2 2B DiT. A small Qwen3 0.6B encoder feeds a learned 6-layer text conditioner, and images live in the 8× latent space of the Qwen-Image VAE. The Base version is the one meant for LoRA training.