Docs
AI ToolkitModels

Wan 2.2 I2V A14B

'The image-to-video model of Wan 2.2: two 14B experts, one for the high-noise steps and one for the low-noise steps, so 14B run per step out of 28.6B total. The start frame enters as a VAE latent concatenated on channels; unlike Wan 2.1 I2V there is no CLIP image encoder. Uses the Wan 2.1 VAE.'

org
Wan-AI (Alibaba)
modality
video
tasks
image-to-video
license
Apache 2.0
released
2025-07-28not verified
native output
1280×720 or 832×480, 81 frames, 16 fpsnot verified
total params
34.39B
model.arch
wan22_14b_i2v

Components

rolemodelparamssizedtypetrained
Transformer (high noise)
Wan 2.2 I2V A14B high-noise expert
diffusers.WanTransformer3DModel

Upstream boundary_ratio is 0.9 (runs while t ≥ 900), but ai-toolkit switches at 875; see Specifics. in_channels 36, so 409,600 more parameters than a T2V expert. Stored as fp32 upstream. ai-toolkit defaults to ai-toolkit/Wan2.2-I2V-A14B-Diffusers-bf16 (bf16, 28,577,914,880 bytes, same parameter count) for the config, and by default loads the weights from the ComfyUI file wan2.2_i2v_high_noise_14B_fp8_scaled.

14.29B
14,288,901,184
57.16 GBfp32yes
Transformer (low noise)
Wan 2.2 I2V A14B low-noise expert
diffusers.WanTransformer3DModel

Runs for the rest of the schedule. Same config as the high-noise expert, different weights. ai-toolkit loads it the same way, from transformer_2 and wan2.2_i2v_low_noise_14B_fp8_scaled.

14.29B
14,288,901,184
57.16 GBfp32yes
Text encoder
UMT5-XXL (encoder only)
transformers.UMT5EncoderModel

Includes the 256k-token embedding table. The ai-toolkit bf16 repo has no text encoder: ai-toolkit uses ai-toolkit/umt5_xxl_encoder (same parameter count and byte size), whose weights are swapped for the ComfyUI umt5_xxl file by default.

5.68B
5,680,910,336
11.36 GBbf16no
Tokenizer
UMT5 SentencePiece (256k vocab)
transformers.T5TokenizerFast

ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead.

————
VAE
Wan 2.1 VAE
diffusers.AutoencoderKLWan

The Wan 2.1 VAE, not the 16× Wan 2.2 VAE the 5B model uses. ai-toolkit always loads ai-toolkit/wan2.1-vae (bf16, 253,806,966 bytes, same parameter count) for this arch.

126.9M
126,892,531
508 MBfp32no
Training adapter
Accuracy recovery adapter, uint4 (both experts)adapter

Rank-16 LoRA over both experts (812 tensors each) that recovers the accuracy lost to 4-bit quantization. Used only when qtype is "uint4|ostris/accuracy_recovery_adapters/wan22_14b_i2v_torchao_uint4.safetensors" (the "4 bit with ARA" option in the UI).

155.8M
155,789,312
312 MBbf16no
total34.39B126.18 GB

Latent space

spatial
8×
temporal
4×
channels
16
patch
1×2×2
autoencoder
Wan 2.1 VAE
pixels per token
16×16 × 4 frames
frame count
4n + 1
notes
Same latent space as Wan 2.1. The transformer takes 36 input channels: 16 noisy latent + 4 mask + 16 conditioning latent (the start frame followed by black frames, VAE-encoded). The mask is 1 for the start frame and 0 elsewhere; its 4 channels are the 4 pixel frames folded into each latent frame. Output is 16 channels.
inputlatent (c×t×h×w)tokens
1024×1024 × 41f16×11×128×12845,056ai-toolkit sample default
832×480 × 81f16×21×60×10432,760native 480p
1280×720 × 81f16×21×90×16075,600native 720p
1024×102416×1×128×1284,096still image

Architecture

Experts
2 × 14,288,901,184 params; one runs per step
Expert switch
High-noise expert for t ≥ 900, low-noise expert below (upstream boundary_ratio 0.9)
Blocks (per expert)
40
Hidden size
5120 (40 heads × 128)
FFN size
13824
Text conditioning
Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
Image conditioning
Start-frame latent + mask concatenated on channels (in_channels 36). No CLIP encoder, no image cross-attention
Objective
Rectified flow; the upstream scheduler uses shift 3.0
Norm / position
QK RMSNorm across heads, 3D RoPE

In AI Toolkit

model.arch
wan22_14b_i2v
UI label
Wan 2.2 I2V (14B) (video)
model.name_or_path
ai-toolkit/Wan2.2-I2V-A14B-Diffusers-bf16
extra UI sections
sample.ctrl_img, datasets.num_frames, model.low_vram, model.multistage, model.layer_offloading, datasets.auto_frame_count

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
linear
sample size
1024×1024, 41 frames, 16 fps
datasets.fps
16
model_kwargs
train_high_noise: true, train_low_noise: true
qtype options
adds "4 bit with ARA" (uint4 + accuracy recovery adapter)
network.conv
disabled (linear LoRA only)

Specifics

Arch string
wan22_14b_i2v is its own class (Wan2214bI2VModel), a subclass of the wan22_14b T2V class.
Expert boundary
ai-toolkit switches experts at 0.875 for both training and sampling, the same boundary as T2V. The upstream model_index.json uses 0.9.
Image-to-video
Always on. Each step takes the first frame of the clip (or the still image) and builds the 20-channel mask + conditioning latent by VAE-encoding that frame followed by black frames for the full clip length. Nothing is cached, so every step pays a full-length VAE encode.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Two experts
Both transformers are wrapped in one DualWanTransformer3DModel that routes each forward by mean timestep: above 875 to the high-noise expert, otherwise the low-noise one. With low_vram, the inactive expert is moved to the CPU when the route changes.
Stage training
Training alternates between the stages: timesteps are drawn from 1000–875 for the high-noise expert and 875–0 for the low-noise one, switching every train.switch_boundary_every steps (UI default 1). model_kwargs.train_high_noise / train_low_noise pick which stages train; at least one must be on.
Transformer source
The ai-toolkit bf16 repo (and the upstream Wan-AI repo) supply the configs; for either repo id the weights come from Comfy-Org/Wan_2.2_ComfyUI_Repackaged wan2.2_i2v_{high,low}_noise_14B_fp8_scaled, the only registered files, even when not quantizing. A local folder loads as-is. model_kwargs.use_comfy_weights: false loads the repo's own weights.
Text encoder source
Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
VAE source
Always ai-toolkit/wan2.1-vae, whatever name_or_path is.
Quantization
Each expert is quantized separately. condition_embedder* and proj_out* stay in full precision. With the ARA, the adapter-covered linears go to uint4 and the rest of the transformer to uint8.
Loss
Flow-matching target (noise − latents) over every latent frame, including the start frame. Training scheduler shift is 5.0, not the 3.0 in the upstream scheduler config.
LoRA target
DualWanTransformer3DModel when training both stages; WanTransformer3DModel (just the chosen expert) when training one.
Saving
LoRAs are split into <name>_high_noise.safetensors and <name>_low_noise.safetensors in the original Wan key layout (set network.split_multistage_loras: false for one combined file). Full fine-tunes save one file per expert with the same suffixes.
Sampling
Uses a Wan 2.2 pipeline with both experts and flowmatch Euler (shift 5.0). num_frames is rounded down to 4n + 1. Give every sample a ctrl_img: the code only builds the 36-channel input when one is set. Size is floored to a multiple of 16 and the image is resized to it. UniPC is disabled until a diffusers regression is fixed.
Not supported
split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
Metadata base version
wan_2.2_14b

Example config

not verified

job: extensionconfig:  name: "my_wan22_14b_i2v_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images/or/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 41          fps: 16          resolution: [512, 768]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        switch_boundary_every: 10        cache_text_embeddings: true      model:        name_or_path: "ai-toolkit/Wan2.2-I2V-A14B-Diffusers-bf16"        arch: "wan22_14b_i2v"        quantize: true        qtype: "uint4|ostris/accuracy_recovery_adapters/wan22_14b_i2v_torchao_uint4.safetensors"        quantize_te: true        qtype_te: "qfloat8"        low_vram: true        model_kwargs:          train_high_noise: true          train_low_noise: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 832        height: 480        num_frames: 41        fps: 16        guidance_scale: 3.5        sample_steps: 25        samples:          - prompt: "a woman turns and smiles at the camera"            ctrl_img: "/path/to/start_frame.jpg"

On this page