Anima Base v1.0
A 2B anime and illustration text-to-image model built on the Cosmos-Predict2 2B DiT. A small Qwen3 0.6B encoder feeds a learned 6-layer text conditioner, and images live in the 8× latent space of the Qwen-Image VAE. The Base version is the one meant for LoRA training.
- org
- CircleStone Labs / Comfy Org
- modality
- image
- tasks
- text-to-image
- license
- CircleStone Labs Non-Commercial License v1.0
- native output
- 512² to 1536² pixels
- total params
- 2.81B
- model.arch
- anima
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Anima DiT (Cosmos-Predict2 2B) diffusers.CosmosTransformer3DModel A video DiT used for single frames: images get a frame dimension of 1. The original single file is split_files/diffusion_models/anima-base-v1.0.safetensors in circlestone-labs/Anima. | 1.96B 1,956,405,248 | 3.91 GB | bf16 | yes |
| Text conditioner | Anima LLM adapter (6 layers) diffusers.AnimaTextConditioner Learned embeddings for T5 token ids (32128 vocab) cross-attend to the Qwen3 hidden states. Its output is what the DiT cross-attends to. Not trained unless model_kwargs.train_text_conditioner is set. In ComfyUI it lives inside the diffusion model file as llm_adapter. | 134.7M 134,663,680 | 269 MB | bf16 | no |
| Text encoder | Qwen3 0.6B (base) transformers.Qwen3Model The base model without the LM head. Last hidden state, 1024-d. | 596.0M 596,049,920 | 1.19 GB | bf16 | no |
| Tokenizer | Qwen3 tokenizer transformers.AutoTokenizer | — | — | — | — |
| Tokenizer | T5 tokenizer (32128 vocab) transformers.AutoTokenizer Only the token ids are used, as queries for the text conditioner. There is no T5 model. | — | — | — | — |
| VAE | Qwen-Image VAE diffusers.AutoencoderKLQwenImage The same VAE as Qwen-Image. A Wan-style video VAE; images are encoded as one frame. | 126.9M 126,892,531 | 254 MB | bf16 | no |
| total | 2.81B | 5.63 GB | |||
Latent space
- autoencoder
- Qwen-Image VAE
- pixels per token
- 16×16
- notes
- Three 2× downsampling stages (dim_mult [1, 2, 4, 4]). Latents are normalized with the per-channel latents_mean / latents_std from the VAE config. The transformer config patch size is [1, 2, 2]: one frame, 2×2 spatial. The DiT also takes a padding-mask channel (concat_padding_mask), which ai-toolkit fills with zeros.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 | 16×1×128×128 | 4,096 | ai-toolkit sample default |
| 512×512 | 16×1×64×64 | 1,024 | lower end of the supported range |
| 1536×1536 | 16×1×192×192 | 9,216 | upper end of the supported range |
Architecture
- Blocks
- 28
- Hidden size
- 2048 (16 heads × 128)
- FFN size
- 8192 (mlp_ratio 4.0)
- Modulation
- AdaLN with a 256-d low-rank adaLN-LoRA
- Text conditioning
- Cross-attention on the text conditioner output (1024-d), padded to at least 512 tokens
- Objective
- Rectified flow, shift 3.0
- Position
- 3D RoPE (rope_scale 1, 4, 4)
- Base model
- nvidia/Cosmos-Predict2-2B-Text2Image
In AI Toolkit
- model.arch
- anima
- UI label
- Anima (image)
- model.name_or_path
- circlestone-labs/Anima-Base-v1.0-Diffusers
- source
- extra UI sections
- model.low_vram, model.layer_offloading
UI defaults
- quantize / quantize_te
- false / false
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- sample.neg
- worst quality, low quality, score_1, score_2, score_3, blurry, jpeg artifacts, sepia, signature, artist name
- network.conv
- disabled (linear LoRA only)
Specifics
- Resolution
- Buckets and sample sizes snap to multiples of 32. The 8× VAE and 2×2 patch only need 16.
- Loading
- Every component comes from name_or_path (transformer/, text_conditioner/, text_encoder/, vae/, tokenizer/, t5_tokenizer/) through the diffusers modular Anima pipeline.
- Prompt encoding
- Each prompt is tokenized twice: Qwen3 for hidden states and T5 for conditioner query ids, both capped at 512 tokens (model_kwargs.max_sequence_length). Empty prompts keep one unmasked position so the conditioner always has something to attend to.
- Text conditioner
- Runs inside the training step, so it can be trained. Set model_kwargs.train_text_conditioner: true to add AnimaTextConditioner to the LoRA targets. Quantized with the transformer flag, at qtype_te.
- LoRA target
- CosmosTransformer3DModel (+ AnimaTextConditioner)
- Saving
- LoRAs are converted to the ComfyUI key layout (diffusion_model.blocks.*, conditioner keys under diffusion_model.llm_adapter.*). ComfyUI-format LoRAs load back in. Full fine-tunes save transformer/ and text_conditioner/ in Diffusers format.
- Sampling
- Flowmatch Euler, shift 3.0, real CFG with the negative prompt. The registry default guidance is 4.5.
- Metadata base version
- anima
Example config
not verified
job: extensionconfig: name: "my_anima_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images" caption_ext: "txt" caption_dropout_rate: 0.05 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 model: name_or_path: "circlestone-labs/Anima-Base-v1.0-Diffusers" arch: "anima" quantize: false quantize_te: false sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 neg: "worst quality, low quality, score_1, score_2, score_3, blurry, jpeg artifacts, sepia, signature, artist name" guidance_scale: 4.5 sample_steps: 30 prompts: - "masterpiece, best quality, score_7, safe, 1girl, red hair, playing chess in a park"Links
Models
Every model AI Toolkit can train, with its parts, sizes, latent space and training specifics.
FLUX.1 [dev]
The 12B guidance-distilled FLUX.1 model. A double- and single-stream rectified flow transformer conditioned on T5-XXL and CLIP-L, working in a 16-channel, 8× VAE latent space. The base most FLUX LoRAs are trained on.