Docs
AI ToolkitModels

Wan 2.1 T2V 1.3B

'The small text-to-video member of Wan 2.1. A 1.4B-parameter DiT in the 8× spatial / 4× temporal latent space of the Wan 2.1 VAE, with UMT5-XXL text conditioning. The cheapest Wan to train: it fits on a consumer GPU without quantizing the transformer.'

org
Wan-AI (Alibaba)
modality
video
tasks
text-to-video · text-to-image
license
Apache 2.0
released
2025-02-25not verified
native output
832×480, 81 frames, 16 fpsnot verified
total params
7.23B
model.arch
wan21:1b

Components

rolemodelparamssizedtypetrained
Transformer
Wan 2.1 T2V 1.3B DiT
diffusers.WanTransformer3DModel

Stored as fp32 in the Diffusers repo. By default ai-toolkit loads the ComfyUI repack instead (Comfy-Org/Wan_2.1_ComfyUI_repackaged, wan2.1_t2v_1.3B_bf16), with the config from this folder.

1.42B
1,418,996,800
5.68 GBfp32yes
Text encoder
UMT5-XXL (encoder only)
transformers.UMT5EncoderModel

Stored as fp32 here (the Wan 2.2 repos ship it in bf16). Includes the 256k-token embedding table. ai-toolkit does not load this folder: it uses ai-toolkit/umt5_xxl_encoder (bf16, same parameter count), whose weights are swapped for the ComfyUI umt5_xxl file by default.

5.68B
5,680,910,336
22.72 GBfp32no
Tokenizer
UMT5 SentencePiece (256k vocab)
transformers.T5TokenizerFast

ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead.

————
VAE
Wan 2.1 VAE
diffusers.AutoencoderKLWan

The same VAE is used by every Wan 2.1 model and by the Wan 2.2 14B models.

126.9M
126,892,531
508 MBfp32no
total7.23B28.91 GB

Latent space

spatial
8×
temporal
4×
channels
16
patch
1×2×2
autoencoder
Wan 2.1 VAE
pixels per token
16×16 × 4 frames
frame count
4n + 1
notes
Causal 3D VAE: 8× spatial downsampling and two temporal downsamples (temperal_downsample [false, true, true]). The first frame gets its own latent frame, which is why frame counts are 4n + 1. Latents are normalized with per-channel latents_mean / latents_std from the VAE config.
inputlatent (c×t×h×w)tokens
1024×1024 × 41f16×11×128×12845,056ai-toolkit sample default
832×480 × 81f16×21×60×10432,760native 480p
1024×102416×1×128×1284,096still image

Architecture

Blocks
30
Hidden size
1536 (12 heads × 128)
FFN size
8960
Text conditioning
Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
Image conditioning
None (text-to-video only)
Timesteps
One timestep per sample
Objective
Rectified flow; the upstream scheduler uses shift 3.0
Norm / position
QK RMSNorm across heads, 3D RoPE

In AI Toolkit

model.arch
wan21:1b
UI label
Wan 2.1 (1.3B) (video)
model.name_or_path
Wan-AI/Wan2.1-T2V-1.3B-Diffusers
extra UI sections
datasets.num_frames, model.low_vram, datasets.auto_frame_count

UI defaults

quantize / quantize_te
false / true (qfloat8)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
sigmoid (UI default, not set by the model)
sample size
1024×1024, 41 frames, 16 fps
datasets.fps
16
network.conv
disabled (linear LoRA only)

Specifics

Arch string
The UI sends wan21:1b. ModelConfig drops everything after the colon, so it loads the same Wan21 class as wan21:14b; only name_or_path differs.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Transformer source
For this repo id the weights come from Comfy-Org/Wan_2.1_ComfyUI_repackaged (bf16 file preferred, fp16 fallback), with the config from name_or_path. A local folder loads as-is. model_kwargs.use_comfy_weights: false opts out.
Text encoder source
Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
VAE source
vae/ inside extras_name_or_path, which defaults to name_or_path.
Loss
Flow-matching target (noise − latents) over every latent frame. Training scheduler shift is 3.0.
LoRA target
WanTransformer3DModel
Saving
LoRAs are converted to the original Wan key layout, which ComfyUI can load. Full fine-tunes save a single file in the original layout.
Sampling
Uses flowmatch Euler with the training scheduler (shift 3.0). UniPC is disabled until a diffusers regression is fixed. low_vram also turns on VAE tiling and a pipeline that moves each model to the GPU only while it runs.
Not supported
split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
Metadata base version
wan_2.1

Example config

not verified

job: extensionconfig:  name: "my_wan21_1b_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 41          fps: 16          resolution: [480, 632]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "sigmoid"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Wan-AI/Wan2.1-T2V-1.3B-Diffusers"        arch: "wan21:1b"        quantize: false        quantize_te: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 832        height: 480        num_frames: 41        fps: 16        guidance_scale: 5        sample_steps: 30        prompts:          - "woman playing the guitar, on stage, singing a song, laser lights, punk rocker"

On this page