Docs
AI ToolkitModels

Wan 2.1 T2V 14B

'The full-size text-to-video model of Wan 2.1. A 14B-parameter DiT in the 8× spatial / 4× temporal latent space of the Wan 2.1 VAE, with UMT5-XXL text conditioning. Same architecture as each Wan 2.2 14B expert, as a single model.'

org
Wan-AI (Alibaba)
modality
video
tasks
text-to-video · text-to-image
license
Apache 2.0
released
2025-02-25not verified
native output
1280×720 or 832×480, 81 frames, 16 fpsnot verified
total params
20.10B
model.arch
wan21:14b

Components

rolemodelparamssizedtypetrained
Transformer
Wan 2.1 T2V 14B DiT
diffusers.WanTransformer3DModel

Stored as fp32 in the Diffusers repo. By default ai-toolkit loads a ComfyUI repack instead (Comfy-Org/Wan_2.1_ComfyUI_repackaged: wan2.1_t2v_14B_fp8_scaled when quantizing with qfloat8, wan2.1_t2v_14B_bf16 otherwise), with the config from this folder.

14.29B
14,288,491,584
57.15 GBfp32yes
Text encoder
UMT5-XXL (encoder only)
transformers.UMT5EncoderModel

Stored as fp32 here (the Wan 2.2 repos ship it in bf16). Includes the 256k-token embedding table. ai-toolkit does not load this folder: it uses ai-toolkit/umt5_xxl_encoder (bf16, same parameter count), whose weights are swapped for the ComfyUI umt5_xxl file by default.

5.68B
5,680,910,336
22.72 GBfp32no
Tokenizer
UMT5 SentencePiece (256k vocab)
transformers.T5TokenizerFast

ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead.

————
VAE
Wan 2.1 VAE
diffusers.AutoencoderKLWan

The same VAE is used by every Wan 2.1 model and by the Wan 2.2 14B models.

126.9M
126,892,531
508 MBfp32no
total20.10B80.39 GB

Latent space

spatial
8×
temporal
4×
channels
16
patch
1×2×2
autoencoder
Wan 2.1 VAE
pixels per token
16×16 × 4 frames
frame count
4n + 1
notes
Causal 3D VAE: 8× spatial downsampling and two temporal downsamples (temperal_downsample [false, true, true]). The first frame gets its own latent frame, which is why frame counts are 4n + 1. Latents are normalized with per-channel latents_mean / latents_std from the VAE config.
inputlatent (c×t×h×w)tokens
1024×1024 × 41f16×11×128×12845,056ai-toolkit sample default
832×480 × 81f16×21×60×10432,760native 480p
1280×720 × 81f16×21×90×16075,600native 720p
1024×102416×1×128×1284,096still image

Architecture

Blocks
40
Hidden size
5120 (40 heads × 128)
FFN size
13824
Text conditioning
Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
Image conditioning
None (text-to-video only)
Timesteps
One timestep per sample
Objective
Rectified flow; the upstream scheduler uses shift 3.0
Norm / position
QK RMSNorm across heads, 3D RoPE

In AI Toolkit

model.arch
wan21:14b
UI label
Wan 2.1 (14B) (video)
model.name_or_path
Wan-AI/Wan2.1-T2V-14B-Diffusers
extra UI sections
datasets.num_frames, model.low_vram, datasets.auto_frame_count

UI defaults

quantize / quantize_te
true / true (qfloat8)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
sigmoid (UI default, not set by the model)
sample size
1024×1024, 41 frames, 16 fps
datasets.fps
16
network.conv
disabled (linear LoRA only)

Specifics

Arch string
The UI sends wan21:14b. ModelConfig drops everything after the colon, so it loads the same Wan21 class as wan21:1b; only name_or_path differs.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Transformer source
For this repo id the weights come from Comfy-Org/Wan_2.1_ComfyUI_repackaged, with the config from name_or_path. quantize with qfloat8 picks wan2.1_t2v_14B_fp8_scaled; no quantization picks the bf16 file. A file already on disk wins over a download. A local folder loads as-is. model_kwargs.use_comfy_weights: false opts out.
Text encoder source
Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
VAE source
vae/ inside extras_name_or_path, which defaults to name_or_path.
Loss
Flow-matching target (noise − latents) over every latent frame. Training scheduler shift is 3.0.
LoRA target
WanTransformer3DModel
Saving
LoRAs are converted to the original Wan key layout, which ComfyUI can load. Full fine-tunes save a single file in the original layout.
Sampling
Uses flowmatch Euler with the training scheduler (shift 3.0). UniPC is disabled until a diffusers regression is fixed. low_vram also turns on VAE tiling and a pipeline that moves each model to the GPU only while it runs.
low_vram
Off in the UI defaults. The repo's 24 GB example config (train_lora_wan21_14b_24gb.yaml) turns it on with quantize: the transformer and text encoder load on the CPU and move to the GPU when needed.
Not supported
split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
Metadata base version
wan_2.1

Example config

not verified

job: extensionconfig:  name: "my_wan21_14b_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 41          fps: 16          resolution: [480, 632]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "sigmoid"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Wan-AI/Wan2.1-T2V-14B-Diffusers"        arch: "wan21:14b"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 832        height: 480        num_frames: 41        fps: 16        guidance_scale: 5        sample_steps: 30        prompts:          - "woman playing the guitar, on stage, singing a song, laser lights, punk rocker"

On this page