Docs
AI ToolkitModels

Wan 2.2 TI2V 5B

The dense 5B member of Wan 2.2. One transformer does both text-to-video and image-to-video, working in the 16× spatial / 4× temporal latent space of the new Wan 2.2 VAE. Small enough to train on a single consumer GPU.

org
Wan-AI (Alibaba)
modality
video
tasks
text-to-video · image-to-video · text-to-image
license
Apache 2.0
released
2025-07-28not verified
native output
1280×704, up to 121 frames, 24 fpsnot verified
total params
11.39B
model.arch
wan22_5b

Components

rolemodelparamssizedtypetrained
Transformer
Wan 2.2 TI2V 5B DiT
diffusers.WanTransformer3DModel

Stored as fp32 in the Diffusers repo, which ai-toolkit reads only for the config. The weights it loads are Comfy-Org/Wan_2.2_ComfyUI_Repackaged split_files/diffusion_models/wan2.2_ti2v_5B_fp16.safetensors (9,999,658,848 bytes).

5.00B
4,999,787,712
20.00 GBfp32yes
Text encoder
UMT5-XXL (encoder only)
transformers.UMT5EncoderModel

Includes the 256k-token embedding table. ai-toolkit does not load this copy: it uses the config and tokenizer from ai-toolkit/umt5_xxl_encoder and the weights from Comfy-Org/Wan_2.1_ComfyUI_repackaged (umt5_xxl_fp8_e4m3fn_scaled or umt5_xxl_fp16).

5.68B
5,680,910,336
11.36 GBbf16no
Tokenizer
UMT5 SentencePiece (256k vocab)
transformers.T5TokenizerFast
————
VAE
Wan 2.2 VAE
diffusers.AutoencoderKLWan

Not the Wan 2.1 VAE. The 14B Wan 2.2 models still use the 2.1 VAE; only this one uses the 2.2 VAE.

704.7M
704,688,668
2.82 GBfp32no
total11.39B34.18 GB

Latent space

spatial
16×
temporal
4×
channels
48
patch
1×2×2
autoencoder
Wan 2.2 VAE
pixels per token
32×32 × 4 frames
frame count
4n + 1
notes
The VAE does a 2×2 pixel unshuffle before encoding (12 input channels = 3 × 2 × 2), then 8× downsampling, for 16× total. The first frame gets its own latent frame, which is why frame counts are 4n + 1.
inputlatent (c×t×h×w)tokens
768×768 × 121f48×31×48×4817,856ai-toolkit sample default
1280×704 × 121f48×31×44×8027,280native 720p
1024×102448×1×64×641,024still image

Architecture

Blocks
30
Hidden size
3072 (24 heads × 128)
FFN size
14336
Text conditioning
Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
Image conditioning
None in the weights. I2V puts the clean first-frame latent in latent frame 0
Timesteps
Per token (expand_timesteps), so conditioned tokens can sit at t = 0
Objective
Rectified flow, shift 5.0
Norm / position
QK RMSNorm across heads, 3D RoPE

In AI Toolkit

model.arch
wan22_5b
UI label
Wan 2.2 TI2V (5B) (video)
model.name_or_path
Wan-AI/Wan2.2-TI2V-5B-Diffusers
extra UI sections
sample.ctrl_img, datasets.num_frames, datasets.do_i2v, datasets.auto_frame_count, model.low_vram

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
sample size
768×768, 121 frames, 24 fps
datasets.do_i2v
true
datasets.fps
24
network.conv
disabled (linear LoRA only)

Specifics

Resolution
Buckets snap to multiples of 32 (16× VAE × 2×2 patch). Wan 2.1 uses 16.
Frame count
Sampling rounds num_frames down to 4n + 1.
Image-to-video
With do_i2v on, the first frame is encoded and placed in latent frame 0 at t = 0 and masked out of the loss (the mask is renormalized). Uses cached first-frame latents when available. Still images train as plain text-to-image.
Weight source
For a Hub repo id, ai-toolkit loads ComfyUI repackaged weights and uses the named repo only for configs: the transformer from wan2.2_ti2v_5B_fp16.safetensors, and UMT5 from umt5_xxl_fp8_e4m3fn_scaled or umt5_xxl_fp16 (ranked by the requested qtype; a local copy of either wins). The UMT5 config and tokenizer come from ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with text_encoder/. model_kwargs.use_comfy_weights: false opts out. The text encoder is never trained.
VAE source
vae/ inside name_or_path (extras_name_or_path defaults to it).
Quantization
condition_embedder* and proj_out* stay in full precision.
LoRA target
WanTransformer3DModel
Saving
LoRAs are converted to the original Wan key layout, which ComfyUI can load. Full fine-tunes save a single file in the original layout.
Sampling
Uses flowmatch Euler (shift 5.0). UniPC is disabled until a diffusers regression is fixed. low_vram also turns on VAE tiling.
Not supported
split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
Metadata base version
wan_2.2_5b

Example config

not verified

job: extensionconfig:  name: "my_wan22_5b_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 121          fps: 24          do_i2v: true          resolution: [512, 768]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Wan-AI/Wan2.2-TI2V-5B-Diffusers"        arch: "wan22_5b"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 768        height: 768        num_frames: 121        fps: 24        guidance_scale: 4        sample_steps: 30        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page