Docs
AI ToolkitModels

Anima Base v1.0

A 2B anime and illustration text-to-image model built on the Cosmos-Predict2 2B DiT. A small Qwen3 0.6B encoder feeds a learned 6-layer text conditioner, and images live in the 8× latent space of the Qwen-Image VAE. The Base version is the one meant for LoRA training.

org
CircleStone Labs / Comfy Org
modality
image
tasks
text-to-image
license
CircleStone Labs Non-Commercial License v1.0
native output
512² to 1536² pixels
total params
2.81B
model.arch
anima

Components

rolemodelparamssizedtypetrained
Transformer
Anima DiT (Cosmos-Predict2 2B)
diffusers.CosmosTransformer3DModel

A video DiT used for single frames: images get a frame dimension of 1. The original single file is split_files/diffusion_models/anima-base-v1.0.safetensors in circlestone-labs/Anima.

1.96B
1,956,405,248
3.91 GBbf16yes
Text conditioner
Anima LLM adapter (6 layers)
diffusers.AnimaTextConditioner

Learned embeddings for T5 token ids (32128 vocab) cross-attend to the Qwen3 hidden states. Its output is what the DiT cross-attends to. Not trained unless model_kwargs.train_text_conditioner is set. In ComfyUI it lives inside the diffusion model file as llm_adapter.

134.7M
134,663,680
269 MBbf16no
Text encoder
Qwen3 0.6B (base)
transformers.Qwen3Model

The base model without the LM head. Last hidden state, 1024-d.

596.0M
596,049,920
1.19 GBbf16no
Tokenizer
Qwen3 tokenizer
transformers.AutoTokenizer
————
Tokenizer
T5 tokenizer (32128 vocab)
transformers.AutoTokenizer

Only the token ids are used, as queries for the text conditioner. There is no T5 model.

————
VAE
Qwen-Image VAE
diffusers.AutoencoderKLQwenImage

The same VAE as Qwen-Image. A Wan-style video VAE; images are encoded as one frame.

126.9M
126,892,531
254 MBbf16no
total2.81B5.63 GB

Latent space

spatial
8×
channels
16
patch
2×2
autoencoder
Qwen-Image VAE
pixels per token
16×16
notes
Three 2× downsampling stages (dim_mult [1, 2, 4, 4]). Latents are normalized with the per-channel latents_mean / latents_std from the VAE config. The transformer config patch size is [1, 2, 2]: one frame, 2×2 spatial. The DiT also takes a padding-mask channel (concat_padding_mask), which ai-toolkit fills with zeros.
inputlatent (c×t×h×w)tokens
1024×102416×1×128×1284,096ai-toolkit sample default
512×51216×1×64×641,024lower end of the supported range
1536×153616×1×192×1929,216upper end of the supported range

Architecture

Blocks
28
Hidden size
2048 (16 heads × 128)
FFN size
8192 (mlp_ratio 4.0)
Modulation
AdaLN with a 256-d low-rank adaLN-LoRA
Text conditioning
Cross-attention on the text conditioner output (1024-d), padded to at least 512 tokens
Objective
Rectified flow, shift 3.0
Position
3D RoPE (rope_scale 1, 4, 4)
Base model
nvidia/Cosmos-Predict2-2B-Text2Image

In AI Toolkit

model.arch
anima
UI label
Anima (image)
model.name_or_path
circlestone-labs/Anima-Base-v1.0-Diffusers
extra UI sections
model.low_vram, model.layer_offloading

UI defaults

quantize / quantize_te
false / false
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
sample.neg
worst quality, low quality, score_1, score_2, score_3, blurry, jpeg artifacts, sepia, signature, artist name
network.conv
disabled (linear LoRA only)

Specifics

Resolution
Buckets and sample sizes snap to multiples of 32. The 8× VAE and 2×2 patch only need 16.
Loading
Every component comes from name_or_path (transformer/, text_conditioner/, text_encoder/, vae/, tokenizer/, t5_tokenizer/) through the diffusers modular Anima pipeline.
Prompt encoding
Each prompt is tokenized twice: Qwen3 for hidden states and T5 for conditioner query ids, both capped at 512 tokens (model_kwargs.max_sequence_length). Empty prompts keep one unmasked position so the conditioner always has something to attend to.
Text conditioner
Runs inside the training step, so it can be trained. Set model_kwargs.train_text_conditioner: true to add AnimaTextConditioner to the LoRA targets. Quantized with the transformer flag, at qtype_te.
LoRA target
CosmosTransformer3DModel (+ AnimaTextConditioner)
Saving
LoRAs are converted to the ComfyUI key layout (diffusion_model.blocks.*, conditioner keys under diffusion_model.llm_adapter.*). ComfyUI-format LoRAs load back in. Full fine-tunes save transformer/ and text_conditioner/ in Diffusers format.
Sampling
Flowmatch Euler, shift 3.0, real CFG with the negative prompt. The registry default guidance is 4.5.
Metadata base version
anima

Example config

not verified

job: extensionconfig:  name: "my_anima_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images"          caption_ext: "txt"          caption_dropout_rate: 0.05          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16      model:        name_or_path: "circlestone-labs/Anima-Base-v1.0-Diffusers"        arch: "anima"        quantize: false        quantize_te: false      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        neg: "worst quality, low quality, score_1, score_2, score_3, blurry, jpeg artifacts, sepia, signature, artist name"        guidance_scale: 4.5        sample_steps: 30        prompts:          - "masterpiece, best quality, score_7, safe, 1girl, red hair, playing chess in a park"

On this page