Docs
AI ToolkitModels

Wan 2.1 I2V 14B 480P

'The 480p image-to-video model of Wan 2.1. A 14B DiT that takes the start frame twice: as a VAE latent concatenated onto the noisy input, and as CLIP ViT-H/14 features read through extra cross-attention. Uses the Wan 2.1 VAE and UMT5-XXL.'

org
Wan-AI (Alibaba)
modality
video
tasks
image-to-video
license
Apache 2.0
released
2025-02-25not verified
native output
832×480, 81 frames, 16 fpsnot verified
total params
22.83B
model.arch
wan21_i2v:14b480p

Components

rolemodelparamssizedtypetrained
Transformer
Wan 2.1 I2V 14B DiT (480P)
diffusers.WanTransformer3DModel

2,106,592,000 more parameters than the T2V 14B, mostly the image K/V projections added to every block. Stored as fp32 in the Diffusers repo. By default ai-toolkit loads a ComfyUI repack instead (Comfy-Org/Wan_2.1_ComfyUI_repackaged: wan2.1_i2v_480p_14B_fp8_scaled when quantizing with qfloat8, the bf16 file otherwise), with the config from this folder.

16.40B
16,395,083,584
65.58 GBfp32yes
Image encoder
CLIP ViT-H/14 vision tower (LAION-2B)
transformers.CLIPVisionModelWithProjection

Config names laion/CLIP-ViT-H-14-laion2B-s32B-b79K. ai-toolkit loads it as CLIPVisionModel (no projection) and feeds the penultimate hidden states, 257 tokens × 1280, to the transformer.

632.1M
632,076,800
1.26 GBbf16no
Image processor
CLIP image processor (224×224)
transformers.CLIPImageProcessor
————
Text encoder
UMT5-XXL (encoder only)
transformers.UMT5EncoderModel

Stored as fp32 here. Includes the 256k-token embedding table. ai-toolkit does not load this folder: it uses ai-toolkit/umt5_xxl_encoder (bf16, same parameter count), whose weights are swapped for the ComfyUI umt5_xxl file by default.

5.68B
5,680,910,336
22.72 GBfp32no
Tokenizer
UMT5 SentencePiece (256k vocab)
transformers.T5TokenizerFast

ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead.

————
VAE
Wan 2.1 VAE
diffusers.AutoencoderKLWan

The same VAE as every other Wan 2.1 model and the Wan 2.2 14B models.

126.9M
126,892,531
508 MBfp32no
total22.83B90.08 GB

Latent space

spatial
8×
temporal
4×
channels
16
patch
1×2×2
autoencoder
Wan 2.1 VAE
pixels per token
16×16 × 4 frames
frame count
4n + 1
notes
The transformer takes 36 input channels: 16 noisy latent + 4 mask + 16 conditioning latent. The conditioning latent is the VAE encoding of the start frame followed by black frames for the rest of the clip. The mask is 1 for the start frame and 0 elsewhere; its 4 channels are the 4 pixel frames folded into each latent frame. Output is 16 channels.
inputlatent (c×t×h×w)tokens
1024×1024 × 41f16×11×128×12845,056ai-toolkit sample default
832×480 × 81f16×21×60×10432,760native 480p
1024×102416×1×128×1284,096still image

Architecture

Blocks
40
Hidden size
5120 (40 heads × 128)
FFN size
13824
Text conditioning
Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
Image conditioning
Start-frame latent + mask concatenated on channels (in_channels 36), plus CLIP tokens (1280-d) through separate image K/V projections in each cross-attention
Timesteps
One timestep per sample
Objective
Rectified flow; the upstream scheduler uses shift 3.0
Norm / position
QK RMSNorm across heads, 3D RoPE

In AI Toolkit

model.arch
wan21_i2v:14b480p
UI label
Wan 2.1 I2V (14B-480P) (video)
model.name_or_path
Wan-AI/Wan2.1-I2V-14B-480P-Diffusers
extra UI sections
sample.ctrl_img, datasets.num_frames, model.low_vram, datasets.auto_frame_count

UI defaults

quantize / quantize_te
true / true (qfloat8)
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
sample size
1024×1024, 41 frames, 16 fps
datasets.fps
16
network.conv
disabled (linear LoRA only)

Specifics

Arch string
The UI sends wan21_i2v:14b480p. ModelConfig drops everything after the colon, so it loads the same Wan21I2V class as the 720P entry; only name_or_path differs.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Image-to-video
Always on. Each step takes the first frame of the clip (or the still image), builds the 20-channel mask + conditioning latent by VAE-encoding that frame followed by black frames for the full clip length, and runs the frame through CLIP (bilinear resize to 224×224, no crop). Nothing is cached, so every step pays a full-length VAE encode.
Loss
Flow-matching target (noise − latents) over every latent frame, including the start frame. Training scheduler shift is 3.0.
Image encoder source
image_encoder/ and image_processor/ from extras_name_or_path (defaults to name_or_path), falling back to name_or_path. Never trained.
Transformer source
For this repo id the weights come from Comfy-Org/Wan_2.1_ComfyUI_repackaged, with the config from name_or_path. quantize with qfloat8 picks wan2.1_i2v_480p_14B_fp8_scaled; no quantization picks the bf16 file. A file already on disk wins over a download. A local folder loads as-is. model_kwargs.use_comfy_weights: false opts out.
Text encoder source
Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
VAE source
vae/ inside extras_name_or_path, which defaults to name_or_path.
LoRA target
WanTransformer3DModel
Saving
LoRAs are converted to the original Wan key layout, which ComfyUI can load. Full fine-tunes save a single file in the original layout.
Sampling
Every sample needs ctrl_img; sampling raises an error without one. Size is floored to a multiple of 16 and the control image is resized to it. Uses flowmatch Euler with the training scheduler (shift 3.0); UniPC is disabled until a diffusers regression is fixed.
low_vram
Keeps the text encoder and CLIP on the CPU between uses and swaps the VAE, CLIP, text encoder and transformer on and off the GPU while sampling.
Not supported
split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
Metadata base version
wan_2.1

Example config

not verified

job: extensionconfig:  name: "my_wan21_i2v_480p_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 41          fps: 16          resolution: [480, 632]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Wan-AI/Wan2.1-I2V-14B-480P-Diffusers"        arch: "wan21_i2v:14b480p"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 832        height: 480        num_frames: 41        fps: 16        guidance_scale: 5        sample_steps: 30        samples:          - prompt: "a woman turns and smiles at the camera"            ctrl_img: "/path/to/start_frame.jpg"

On this page