Wan 2.1 T2V 14B
'The full-size text-to-video model of Wan 2.1. A 14B-parameter DiT in the 8× spatial / 4× temporal latent space of the Wan 2.1 VAE, with UMT5-XXL text conditioning. Same architecture as each Wan 2.2 14B expert, as a single model.'
- org
- Wan-AI (Alibaba)
- modality
- video
- tasks
- text-to-video · text-to-image
- license
- Apache 2.0
- released
- 2025-02-25not verified
- native output
- 1280×720 or 832×480, 81 frames, 16 fpsnot verified
- total params
- 20.10B
- model.arch
- wan21:14b
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Wan 2.1 T2V 14B DiT diffusers.WanTransformer3DModel Stored as fp32 in the Diffusers repo. By default ai-toolkit loads a ComfyUI repack instead (Comfy-Org/Wan_2.1_ComfyUI_repackaged: wan2.1_t2v_14B_fp8_scaled when quantizing with qfloat8, wan2.1_t2v_14B_bf16 otherwise), with the config from this folder. | 14.29B 14,288,491,584 | 57.15 GB | fp32 | yes |
| Text encoder | UMT5-XXL (encoder only) transformers.UMT5EncoderModel Stored as fp32 here (the Wan 2.2 repos ship it in bf16). Includes the 256k-token embedding table. ai-toolkit does not load this folder: it uses ai-toolkit/umt5_xxl_encoder (bf16, same parameter count), whose weights are swapped for the ComfyUI umt5_xxl file by default. | 5.68B 5,680,910,336 | 22.72 GB | fp32 | no |
| Tokenizer | UMT5 SentencePiece (256k vocab) transformers.T5TokenizerFast ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead. | — | — | — | — |
| VAE | Wan 2.1 VAE diffusers.AutoencoderKLWan The same VAE is used by every Wan 2.1 model and by the Wan 2.2 14B models. | 126.9M 126,892,531 | 508 MB | fp32 | no |
| total | 20.10B | 80.39 GB | |||
Latent space
- autoencoder
- Wan 2.1 VAE
- pixels per token
- 16×16 × 4 frames
- frame count
- 4n + 1
- notes
- Causal 3D VAE: 8× spatial downsampling and two temporal downsamples (temperal_downsample [false, true, true]). The first frame gets its own latent frame, which is why frame counts are 4n + 1. Latents are normalized with per-channel latents_mean / latents_std from the VAE config.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 × 41f | 16×11×128×128 | 45,056 | ai-toolkit sample default |
| 832×480 × 81f | 16×21×60×104 | 32,760 | native 480p |
| 1280×720 × 81f | 16×21×90×160 | 75,600 | native 720p |
| 1024×1024 | 16×1×128×128 | 4,096 | still image |
Architecture
- Blocks
- 40
- Hidden size
- 5120 (40 heads × 128)
- FFN size
- 13824
- Text conditioning
- Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
- Image conditioning
- None (text-to-video only)
- Timesteps
- One timestep per sample
- Objective
- Rectified flow; the upstream scheduler uses shift 3.0
- Norm / position
- QK RMSNorm across heads, 3D RoPE
In AI Toolkit
- model.arch
- wan21:14b
- UI label
- Wan 2.1 (14B) (video)
- model.name_or_path
- Wan-AI/Wan2.1-T2V-14B-Diffusers
- source
- extra UI sections
- datasets.num_frames, model.low_vram, datasets.auto_frame_count
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- sigmoid (UI default, not set by the model)
- sample size
- 1024×1024, 41 frames, 16 fps
- datasets.fps
- 16
- network.conv
- disabled (linear LoRA only)
Specifics
- Arch string
- The UI sends wan21:14b. ModelConfig drops everything after the colon, so it loads the same Wan21 class as wan21:1b; only name_or_path differs.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Transformer source
- For this repo id the weights come from Comfy-Org/Wan_2.1_ComfyUI_repackaged, with the config from name_or_path. quantize with qfloat8 picks wan2.1_t2v_14B_fp8_scaled; no quantization picks the bf16 file. A file already on disk wins over a download. A local folder loads as-is. model_kwargs.use_comfy_weights: false opts out.
- Text encoder source
- Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
- VAE source
- vae/ inside extras_name_or_path, which defaults to name_or_path.
- Loss
- Flow-matching target (noise − latents) over every latent frame. Training scheduler shift is 3.0.
- LoRA target
- WanTransformer3DModel
- Saving
- LoRAs are converted to the original Wan key layout, which ComfyUI can load. Full fine-tunes save a single file in the original layout.
- Sampling
- Uses flowmatch Euler with the training scheduler (shift 3.0). UniPC is disabled until a diffusers regression is fixed. low_vram also turns on VAE tiling and a pipeline that moves each model to the GPU only while it runs.
- low_vram
- Off in the UI defaults. The repo's 24 GB example config (train_lora_wan21_14b_24gb.yaml) turns it on with quantize: the transformer and text encoder load on the CPU and move to the GPU when needed.
- Not supported
- split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
- Metadata base version
- wan_2.1
Example config
not verified
job: extensionconfig: name: "my_wan21_14b_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 41 fps: 16 resolution: [480, 632] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "sigmoid" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Wan-AI/Wan2.1-T2V-14B-Diffusers" arch: "wan21:14b" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 832 height: 480 num_frames: 41 fps: 16 guidance_scale: 5 sample_steps: 30 prompts: - "woman playing the guitar, on stage, singing a song, laser lights, punk rocker"Links
Wan 2.1 I2V 14B 720P
'The 720p image-to-video model of Wan 2.1. A 14B DiT that takes the start frame twice: as a VAE latent concatenated onto the noisy input, and as CLIP ViT-H/14 features read through extra cross-attention. Uses the Wan 2.1 VAE and UMT5-XXL.'
Wan 2.2 T2V A14B
'The text-to-video flagship of Wan 2.2: two 14B experts, one for the high-noise steps and one for the low-noise steps, so 14B run per step out of 28.6B total. Each expert has the Wan 2.1 14B architecture and uses the Wan 2.1 VAE. ai-toolkit trains both experts together, or either one alone.'