Wan 2.1 I2V 14B 480P
'The 480p image-to-video model of Wan 2.1. A 14B DiT that takes the start frame twice: as a VAE latent concatenated onto the noisy input, and as CLIP ViT-H/14 features read through extra cross-attention. Uses the Wan 2.1 VAE and UMT5-XXL.'
- org
- Wan-AI (Alibaba)
- modality
- video
- tasks
- image-to-video
- license
- Apache 2.0
- released
- 2025-02-25not verified
- native output
- 832×480, 81 frames, 16 fpsnot verified
- total params
- 22.83B
- model.arch
- wan21_i2v:14b480p
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Wan 2.1 I2V 14B DiT (480P) diffusers.WanTransformer3DModel 2,106,592,000 more parameters than the T2V 14B, mostly the image K/V projections added to every block. Stored as fp32 in the Diffusers repo. By default ai-toolkit loads a ComfyUI repack instead (Comfy-Org/Wan_2.1_ComfyUI_repackaged: wan2.1_i2v_480p_14B_fp8_scaled when quantizing with qfloat8, the bf16 file otherwise), with the config from this folder. | 16.40B 16,395,083,584 | 65.58 GB | fp32 | yes |
| Image encoder | CLIP ViT-H/14 vision tower (LAION-2B) transformers.CLIPVisionModelWithProjection Config names laion/CLIP-ViT-H-14-laion2B-s32B-b79K. ai-toolkit loads it as CLIPVisionModel (no projection) and feeds the penultimate hidden states, 257 tokens × 1280, to the transformer. | 632.1M 632,076,800 | 1.26 GB | bf16 | no |
| Image processor | CLIP image processor (224×224) transformers.CLIPImageProcessor | — | — | — | — |
| Text encoder | UMT5-XXL (encoder only) transformers.UMT5EncoderModel Stored as fp32 here. Includes the 256k-token embedding table. ai-toolkit does not load this folder: it uses ai-toolkit/umt5_xxl_encoder (bf16, same parameter count), whose weights are swapped for the ComfyUI umt5_xxl file by default. | 5.68B 5,680,910,336 | 22.72 GB | fp32 | no |
| Tokenizer | UMT5 SentencePiece (256k vocab) transformers.T5TokenizerFast ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead. | — | — | — | — |
| VAE | Wan 2.1 VAE diffusers.AutoencoderKLWan The same VAE as every other Wan 2.1 model and the Wan 2.2 14B models. | 126.9M 126,892,531 | 508 MB | fp32 | no |
| total | 22.83B | 90.08 GB | |||
Latent space
- autoencoder
- Wan 2.1 VAE
- pixels per token
- 16×16 × 4 frames
- frame count
- 4n + 1
- notes
- The transformer takes 36 input channels: 16 noisy latent + 4 mask + 16 conditioning latent. The conditioning latent is the VAE encoding of the start frame followed by black frames for the rest of the clip. The mask is 1 for the start frame and 0 elsewhere; its 4 channels are the 4 pixel frames folded into each latent frame. Output is 16 channels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 × 41f | 16×11×128×128 | 45,056 | ai-toolkit sample default |
| 832×480 × 81f | 16×21×60×104 | 32,760 | native 480p |
| 1024×1024 | 16×1×128×128 | 4,096 | still image |
Architecture
- Blocks
- 40
- Hidden size
- 5120 (40 heads × 128)
- FFN size
- 13824
- Text conditioning
- Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
- Image conditioning
- Start-frame latent + mask concatenated on channels (in_channels 36), plus CLIP tokens (1280-d) through separate image K/V projections in each cross-attention
- Timesteps
- One timestep per sample
- Objective
- Rectified flow; the upstream scheduler uses shift 3.0
- Norm / position
- QK RMSNorm across heads, 3D RoPE
In AI Toolkit
- model.arch
- wan21_i2v:14b480p
- UI label
- Wan 2.1 I2V (14B-480P) (video)
- model.name_or_path
- Wan-AI/Wan2.1-I2V-14B-480P-Diffusers
- source
- toolkit/models/wan21/wan21_i2v.py
- toolkit/models/wan21/wan21.py
- toolkit/models/wan21/wan_utils.py
- toolkit/models/wan21/wan_lora_convert.py
- toolkit/models/v2/diffusion_models/wan.py
- toolkit/models/v2/text_encoders/umt5.py
- toolkit/models/v2/vae/wan.py
- toolkit/models/v2/vision_encoders/clip_vision.py
- toolkit/models/v2/resolver.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- sample.ctrl_img, datasets.num_frames, model.low_vram, datasets.auto_frame_count
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- sample size
- 1024×1024, 41 frames, 16 fps
- datasets.fps
- 16
- network.conv
- disabled (linear LoRA only)
Specifics
- Arch string
- The UI sends wan21_i2v:14b480p. ModelConfig drops everything after the colon, so it loads the same Wan21I2V class as the 720P entry; only name_or_path differs.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Image-to-video
- Always on. Each step takes the first frame of the clip (or the still image), builds the 20-channel mask + conditioning latent by VAE-encoding that frame followed by black frames for the full clip length, and runs the frame through CLIP (bilinear resize to 224×224, no crop). Nothing is cached, so every step pays a full-length VAE encode.
- Loss
- Flow-matching target (noise − latents) over every latent frame, including the start frame. Training scheduler shift is 3.0.
- Image encoder source
- image_encoder/ and image_processor/ from extras_name_or_path (defaults to name_or_path), falling back to name_or_path. Never trained.
- Transformer source
- For this repo id the weights come from Comfy-Org/Wan_2.1_ComfyUI_repackaged, with the config from name_or_path. quantize with qfloat8 picks wan2.1_i2v_480p_14B_fp8_scaled; no quantization picks the bf16 file. A file already on disk wins over a download. A local folder loads as-is. model_kwargs.use_comfy_weights: false opts out.
- Text encoder source
- Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
- VAE source
- vae/ inside extras_name_or_path, which defaults to name_or_path.
- LoRA target
- WanTransformer3DModel
- Saving
- LoRAs are converted to the original Wan key layout, which ComfyUI can load. Full fine-tunes save a single file in the original layout.
- Sampling
- Every sample needs ctrl_img; sampling raises an error without one. Size is floored to a multiple of 16 and the control image is resized to it. Uses flowmatch Euler with the training scheduler (shift 3.0); UniPC is disabled until a diffusers regression is fixed.
- low_vram
- Keeps the text encoder and CLIP on the CPU between uses and swaps the VAE, CLIP, text encoder and transformer on and off the GPU while sampling.
- Not supported
- split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
- Metadata base version
- wan_2.1
Example config
not verified
job: extensionconfig: name: "my_wan21_i2v_480p_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 41 fps: 16 resolution: [480, 632] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Wan-AI/Wan2.1-I2V-14B-480P-Diffusers" arch: "wan21_i2v:14b480p" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 832 height: 480 num_frames: 41 fps: 16 guidance_scale: 5 sample_steps: 30 samples: - prompt: "a woman turns and smiles at the camera" ctrl_img: "/path/to/start_frame.jpg"Links
Wan 2.1 T2V 1.3B
'The small text-to-video member of Wan 2.1. A 1.4B-parameter DiT in the 8× spatial / 4× temporal latent space of the Wan 2.1 VAE, with UMT5-XXL text conditioning. The cheapest Wan to train: it fits on a consumer GPU without quantizing the transformer.'
Wan 2.1 I2V 14B 720P
'The 720p image-to-video model of Wan 2.1. A 14B DiT that takes the start frame twice: as a VAE latent concatenated onto the noisy input, and as CLIP ViT-H/14 features read through extra cross-attention. Uses the Wan 2.1 VAE and UMT5-XXL.'