Wan 2.2 I2V A14B
'The image-to-video model of Wan 2.2: two 14B experts, one for the high-noise steps and one for the low-noise steps, so 14B run per step out of 28.6B total. The start frame enters as a VAE latent concatenated on channels; unlike Wan 2.1 I2V there is no CLIP image encoder. Uses the Wan 2.1 VAE.'
- org
- Wan-AI (Alibaba)
- modality
- video
- tasks
- image-to-video
- license
- Apache 2.0
- released
- 2025-07-28not verified
- native output
- 1280×720 or 832×480, 81 frames, 16 fpsnot verified
- total params
- 34.39B
- model.arch
- wan22_14b_i2v
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer (high noise) | Wan 2.2 I2V A14B high-noise expert diffusers.WanTransformer3DModel Upstream boundary_ratio is 0.9 (runs while t ≥ 900), but ai-toolkit switches at 875; see Specifics. in_channels 36, so 409,600 more parameters than a T2V expert. Stored as fp32 upstream. ai-toolkit defaults to ai-toolkit/Wan2.2-I2V-A14B-Diffusers-bf16 (bf16, 28,577,914,880 bytes, same parameter count) for the config, and by default loads the weights from the ComfyUI file wan2.2_i2v_high_noise_14B_fp8_scaled. | 14.29B 14,288,901,184 | 57.16 GB | fp32 | yes |
| Transformer (low noise) | Wan 2.2 I2V A14B low-noise expert diffusers.WanTransformer3DModel Runs for the rest of the schedule. Same config as the high-noise expert, different weights. ai-toolkit loads it the same way, from transformer_2 and wan2.2_i2v_low_noise_14B_fp8_scaled. | 14.29B 14,288,901,184 | 57.16 GB | fp32 | yes |
| Text encoder | UMT5-XXL (encoder only) transformers.UMT5EncoderModel Includes the 256k-token embedding table. The ai-toolkit bf16 repo has no text encoder: ai-toolkit uses ai-toolkit/umt5_xxl_encoder (same parameter count and byte size), whose weights are swapped for the ComfyUI umt5_xxl file by default. | 5.68B 5,680,910,336 | 11.36 GB | bf16 | no |
| Tokenizer | UMT5 SentencePiece (256k vocab) transformers.T5TokenizerFast ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead. | — | — | — | — |
| VAE | Wan 2.1 VAE diffusers.AutoencoderKLWan The Wan 2.1 VAE, not the 16× Wan 2.2 VAE the 5B model uses. ai-toolkit always loads ai-toolkit/wan2.1-vae (bf16, 253,806,966 bytes, same parameter count) for this arch. | 126.9M 126,892,531 | 508 MB | fp32 | no |
| Training adapter | Accuracy recovery adapter, uint4 (both experts)adapter Rank-16 LoRA over both experts (812 tensors each) that recovers the accuracy lost to 4-bit quantization. Used only when qtype is "uint4|ostris/accuracy_recovery_adapters/wan22_14b_i2v_torchao_uint4.safetensors" (the "4 bit with ARA" option in the UI). | 155.8M 155,789,312 | 312 MB | bf16 | no |
| total | 34.39B | 126.18 GB | |||
Latent space
- autoencoder
- Wan 2.1 VAE
- pixels per token
- 16×16 × 4 frames
- frame count
- 4n + 1
- notes
- Same latent space as Wan 2.1. The transformer takes 36 input channels: 16 noisy latent + 4 mask + 16 conditioning latent (the start frame followed by black frames, VAE-encoded). The mask is 1 for the start frame and 0 elsewhere; its 4 channels are the 4 pixel frames folded into each latent frame. Output is 16 channels.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 × 41f | 16×11×128×128 | 45,056 | ai-toolkit sample default |
| 832×480 × 81f | 16×21×60×104 | 32,760 | native 480p |
| 1280×720 × 81f | 16×21×90×160 | 75,600 | native 720p |
| 1024×1024 | 16×1×128×128 | 4,096 | still image |
Architecture
- Experts
- 2 × 14,288,901,184 params; one runs per step
- Expert switch
- High-noise expert for t ≥ 900, low-noise expert below (upstream boundary_ratio 0.9)
- Blocks (per expert)
- 40
- Hidden size
- 5120 (40 heads × 128)
- FFN size
- 13824
- Text conditioning
- Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
- Image conditioning
- Start-frame latent + mask concatenated on channels (in_channels 36). No CLIP encoder, no image cross-attention
- Objective
- Rectified flow; the upstream scheduler uses shift 3.0
- Norm / position
- QK RMSNorm across heads, 3D RoPE
In AI Toolkit
- model.arch
- wan22_14b_i2v
- UI label
- Wan 2.2 I2V (14B) (video)
- model.name_or_path
- ai-toolkit/Wan2.2-I2V-A14B-Diffusers-bf16
- source
- extensions_built_in/diffusion_models/wan22/wan22_14b_i2v_model.py
- extensions_built_in/diffusion_models/wan22/wan22_14b_model.py
- extensions_built_in/diffusion_models/wan22/wan22_pipeline.py
- extensions_built_in/diffusion_models/wan22/wan22_5b_model.py
- toolkit/models/wan21/wan21.py
- toolkit/models/wan21/wan_utils.py
- toolkit/models/wan21/wan_lora_convert.py
- toolkit/models/v2/diffusion_models/wan.py
- toolkit/models/v2/text_encoders/umt5.py
- toolkit/models/v2/vae/wan.py
- toolkit/models/v2/resolver.py
- toolkit/util/quantize.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- sample.ctrl_img, datasets.num_frames, model.low_vram, model.multistage, model.layer_offloading, datasets.auto_frame_count
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- linear
- sample size
- 1024×1024, 41 frames, 16 fps
- datasets.fps
- 16
- model_kwargs
- train_high_noise: true, train_low_noise: true
- qtype options
- adds "4 bit with ARA" (uint4 + accuracy recovery adapter)
- network.conv
- disabled (linear LoRA only)
Specifics
- Arch string
- wan22_14b_i2v is its own class (Wan2214bI2VModel), a subclass of the wan22_14b T2V class.
- Expert boundary
- ai-toolkit switches experts at 0.875 for both training and sampling, the same boundary as T2V. The upstream model_index.json uses 0.9.
- Image-to-video
- Always on. Each step takes the first frame of the clip (or the still image) and builds the 20-channel mask + conditioning latent by VAE-encoding that frame followed by black frames for the full clip length. Nothing is cached, so every step pays a full-length VAE encode.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Two experts
- Both transformers are wrapped in one DualWanTransformer3DModel that routes each forward by mean timestep: above 875 to the high-noise expert, otherwise the low-noise one. With low_vram, the inactive expert is moved to the CPU when the route changes.
- Stage training
- Training alternates between the stages: timesteps are drawn from 1000–875 for the high-noise expert and 875–0 for the low-noise one, switching every train.switch_boundary_every steps (UI default 1). model_kwargs.train_high_noise / train_low_noise pick which stages train; at least one must be on.
- Transformer source
- The ai-toolkit bf16 repo (and the upstream Wan-AI repo) supply the configs; for either repo id the weights come from Comfy-Org/Wan_2.2_ComfyUI_Repackaged wan2.2_i2v_{high,low}_noise_14B_fp8_scaled, the only registered files, even when not quantizing. A local folder loads as-is. model_kwargs.use_comfy_weights: false loads the repo's own weights.
- Text encoder source
- Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
- VAE source
- Always ai-toolkit/wan2.1-vae, whatever name_or_path is.
- Quantization
- Each expert is quantized separately. condition_embedder* and proj_out* stay in full precision. With the ARA, the adapter-covered linears go to uint4 and the rest of the transformer to uint8.
- Loss
- Flow-matching target (noise − latents) over every latent frame, including the start frame. Training scheduler shift is 5.0, not the 3.0 in the upstream scheduler config.
- LoRA target
- DualWanTransformer3DModel when training both stages; WanTransformer3DModel (just the chosen expert) when training one.
- Saving
- LoRAs are split into <name>_high_noise.safetensors and <name>_low_noise.safetensors in the original Wan key layout (set network.split_multistage_loras: false for one combined file). Full fine-tunes save one file per expert with the same suffixes.
- Sampling
- Uses a Wan 2.2 pipeline with both experts and flowmatch Euler (shift 5.0). num_frames is rounded down to 4n + 1. Give every sample a ctrl_img: the code only builds the 36-channel input when one is set. Size is floored to a multiple of 16 and the image is resized to it. UniPC is disabled until a diffusers regression is fixed.
- Not supported
- split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
- Metadata base version
- wan_2.2_14b
Example config
not verified
job: extensionconfig: name: "my_wan22_14b_i2v_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images/or/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 41 fps: 16 resolution: [512, 768] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 switch_boundary_every: 10 cache_text_embeddings: true model: name_or_path: "ai-toolkit/Wan2.2-I2V-A14B-Diffusers-bf16" arch: "wan22_14b_i2v" quantize: true qtype: "uint4|ostris/accuracy_recovery_adapters/wan22_14b_i2v_torchao_uint4.safetensors" quantize_te: true qtype_te: "qfloat8" low_vram: true model_kwargs: train_high_noise: true train_low_noise: true sample: sampler: "flowmatch" sample_every: 250 width: 832 height: 480 num_frames: 41 fps: 16 guidance_scale: 3.5 sample_steps: 25 samples: - prompt: "a woman turns and smiles at the camera" ctrl_img: "/path/to/start_frame.jpg"Links
Wan 2.2 T2V A14B
'The text-to-video flagship of Wan 2.2: two 14B experts, one for the high-noise steps and one for the low-noise steps, so 14B run per step out of 28.6B total. Each expert has the Wan 2.1 14B architecture and uses the Wan 2.1 VAE. ai-toolkit trains both experts together, or either one alone.'
Wan 2.2 TI2V 5B
The dense 5B member of Wan 2.2. One transformer does both text-to-video and image-to-video, working in the 16× spatial / 4× temporal latent space of the new Wan 2.2 VAE. Small enough to train on a single consumer GPU.