Wan 2.1 T2V 1.3B
'The small text-to-video member of Wan 2.1. A 1.4B-parameter DiT in the 8× spatial / 4× temporal latent space of the Wan 2.1 VAE, with UMT5-XXL text conditioning. The cheapest Wan to train: it fits on a consumer GPU without quantizing the transformer.'
- org
- Wan-AI (Alibaba)
- modality
- video
- tasks
- text-to-video · text-to-image
- license
- Apache 2.0
- released
- 2025-02-25not verified
- native output
- 832×480, 81 frames, 16 fpsnot verified
- total params
- 7.23B
- model.arch
- wan21:1b
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Wan 2.1 T2V 1.3B DiT diffusers.WanTransformer3DModel Stored as fp32 in the Diffusers repo. By default ai-toolkit loads the ComfyUI repack instead (Comfy-Org/Wan_2.1_ComfyUI_repackaged, wan2.1_t2v_1.3B_bf16), with the config from this folder. | 1.42B 1,418,996,800 | 5.68 GB | fp32 | yes |
| Text encoder | UMT5-XXL (encoder only) transformers.UMT5EncoderModel Stored as fp32 here (the Wan 2.2 repos ship it in bf16). Includes the 256k-token embedding table. ai-toolkit does not load this folder: it uses ai-toolkit/umt5_xxl_encoder (bf16, same parameter count), whose weights are swapped for the ComfyUI umt5_xxl file by default. | 5.68B 5,680,910,336 | 22.72 GB | fp32 | no |
| Tokenizer | UMT5 SentencePiece (256k vocab) transformers.T5TokenizerFast ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead. | — | — | — | — |
| VAE | Wan 2.1 VAE diffusers.AutoencoderKLWan The same VAE is used by every Wan 2.1 model and by the Wan 2.2 14B models. | 126.9M 126,892,531 | 508 MB | fp32 | no |
| total | 7.23B | 28.91 GB | |||
Latent space
- autoencoder
- Wan 2.1 VAE
- pixels per token
- 16×16 × 4 frames
- frame count
- 4n + 1
- notes
- Causal 3D VAE: 8× spatial downsampling and two temporal downsamples (temperal_downsample [false, true, true]). The first frame gets its own latent frame, which is why frame counts are 4n + 1. Latents are normalized with per-channel latents_mean / latents_std from the VAE config.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 × 41f | 16×11×128×128 | 45,056 | ai-toolkit sample default |
| 832×480 × 81f | 16×21×60×104 | 32,760 | native 480p |
| 1024×1024 | 16×1×128×128 | 4,096 | still image |
Architecture
- Blocks
- 30
- Hidden size
- 1536 (12 heads × 128)
- FFN size
- 8960
- Text conditioning
- Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
- Image conditioning
- None (text-to-video only)
- Timesteps
- One timestep per sample
- Objective
- Rectified flow; the upstream scheduler uses shift 3.0
- Norm / position
- QK RMSNorm across heads, 3D RoPE
In AI Toolkit
- model.arch
- wan21:1b
- UI label
- Wan 2.1 (1.3B) (video)
- model.name_or_path
- Wan-AI/Wan2.1-T2V-1.3B-Diffusers
- source
- extra UI sections
- datasets.num_frames, model.low_vram, datasets.auto_frame_count
UI defaults
- quantize / quantize_te
- false / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- sigmoid (UI default, not set by the model)
- sample size
- 1024×1024, 41 frames, 16 fps
- datasets.fps
- 16
- network.conv
- disabled (linear LoRA only)
Specifics
- Arch string
- The UI sends wan21:1b. ModelConfig drops everything after the colon, so it loads the same Wan21 class as wan21:14b; only name_or_path differs.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Transformer source
- For this repo id the weights come from Comfy-Org/Wan_2.1_ComfyUI_repackaged (bf16 file preferred, fp16 fallback), with the config from name_or_path. A local folder loads as-is. model_kwargs.use_comfy_weights: false opts out.
- Text encoder source
- Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
- VAE source
- vae/ inside extras_name_or_path, which defaults to name_or_path.
- Loss
- Flow-matching target (noise − latents) over every latent frame. Training scheduler shift is 3.0.
- LoRA target
- WanTransformer3DModel
- Saving
- LoRAs are converted to the original Wan key layout, which ComfyUI can load. Full fine-tunes save a single file in the original layout.
- Sampling
- Uses flowmatch Euler with the training scheduler (shift 3.0). UniPC is disabled until a diffusers regression is fixed. low_vram also turns on VAE tiling and a pipeline that moves each model to the GPU only while it runs.
- Not supported
- split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
- Metadata base version
- wan_2.1
Example config
not verified
job: extensionconfig: name: "my_wan21_1b_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 41 fps: 16 resolution: [480, 632] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "sigmoid" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Wan-AI/Wan2.1-T2V-1.3B-Diffusers" arch: "wan21:1b" quantize: false quantize_te: true sample: sampler: "flowmatch" sample_every: 250 width: 832 height: 480 num_frames: 41 fps: 16 guidance_scale: 5 sample_steps: 30 prompts: - "woman playing the guitar, on stage, singing a song, laser lights, punk rocker"Links
Boogu-Image 0.1 Edit
'The instruction-editing model of Boogu-Image 0.1: the same 10.3B Lumina2-style DiT, Qwen3-VL-8B encoder and FLUX.1 VAE as the Base, trained to edit. The reference image is read twice, by Qwen3-VL with the instruction and as VAE latents through a dedicated refiner.'
Wan 2.1 I2V 14B 480P
'The 480p image-to-video model of Wan 2.1. A 14B DiT that takes the start frame twice: as a VAE latent concatenated onto the noisy input, and as CLIP ViT-H/14 features read through extra cross-attention. Uses the Wan 2.1 VAE and UMT5-XXL.'