Wan 2.2 TI2V 5B
The dense 5B member of Wan 2.2. One transformer does both text-to-video and image-to-video, working in the 16× spatial / 4× temporal latent space of the new Wan 2.2 VAE. Small enough to train on a single consumer GPU.
- org
- Wan-AI (Alibaba)
- modality
- video
- tasks
- text-to-video · image-to-video · text-to-image
- license
- Apache 2.0
- released
- 2025-07-28not verified
- native output
- 1280×704, up to 121 frames, 24 fpsnot verified
- total params
- 11.39B
- model.arch
- wan22_5b
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | Wan 2.2 TI2V 5B DiT diffusers.WanTransformer3DModel Stored as fp32 in the Diffusers repo, which ai-toolkit reads only for the config. The weights it loads are Comfy-Org/Wan_2.2_ComfyUI_Repackaged split_files/diffusion_models/wan2.2_ti2v_5B_fp16.safetensors (9,999,658,848 bytes). | 5.00B 4,999,787,712 | 20.00 GB | fp32 | yes |
| Text encoder | UMT5-XXL (encoder only) transformers.UMT5EncoderModel Includes the 256k-token embedding table. ai-toolkit does not load this copy: it uses the config and tokenizer from ai-toolkit/umt5_xxl_encoder and the weights from Comfy-Org/Wan_2.1_ComfyUI_repackaged (umt5_xxl_fp8_e4m3fn_scaled or umt5_xxl_fp16). | 5.68B 5,680,910,336 | 11.36 GB | bf16 | no |
| Tokenizer | UMT5 SentencePiece (256k vocab) transformers.T5TokenizerFast | — | — | — | — |
| VAE | Wan 2.2 VAE diffusers.AutoencoderKLWan Not the Wan 2.1 VAE. The 14B Wan 2.2 models still use the 2.1 VAE; only this one uses the 2.2 VAE. | 704.7M 704,688,668 | 2.82 GB | fp32 | no |
| total | 11.39B | 34.18 GB | |||
Latent space
- autoencoder
- Wan 2.2 VAE
- pixels per token
- 32×32 × 4 frames
- frame count
- 4n + 1
- notes
- The VAE does a 2×2 pixel unshuffle before encoding (12 input channels = 3 × 2 × 2), then 8× downsampling, for 16× total. The first frame gets its own latent frame, which is why frame counts are 4n + 1.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 768×768 × 121f | 48×31×48×48 | 17,856 | ai-toolkit sample default |
| 1280×704 × 121f | 48×31×44×80 | 27,280 | native 720p |
| 1024×1024 | 48×1×64×64 | 1,024 | still image |
Architecture
- Blocks
- 30
- Hidden size
- 3072 (24 heads × 128)
- FFN size
- 14336
- Text conditioning
- Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
- Image conditioning
- None in the weights. I2V puts the clean first-frame latent in latent frame 0
- Timesteps
- Per token (expand_timesteps), so conditioned tokens can sit at t = 0
- Objective
- Rectified flow, shift 5.0
- Norm / position
- QK RMSNorm across heads, 3D RoPE
In AI Toolkit
- model.arch
- wan22_5b
- UI label
- Wan 2.2 TI2V (5B) (video)
- model.name_or_path
- Wan-AI/Wan2.2-TI2V-5B-Diffusers
- source
- extra UI sections
- sample.ctrl_img, datasets.num_frames, datasets.do_i2v, datasets.auto_frame_count, model.low_vram
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- sample size
- 768×768, 121 frames, 24 fps
- datasets.do_i2v
- true
- datasets.fps
- 24
- network.conv
- disabled (linear LoRA only)
Specifics
- Resolution
- Buckets snap to multiples of 32 (16× VAE × 2×2 patch). Wan 2.1 uses 16.
- Frame count
- Sampling rounds num_frames down to 4n + 1.
- Image-to-video
- With do_i2v on, the first frame is encoded and placed in latent frame 0 at t = 0 and masked out of the loss (the mask is renormalized). Uses cached first-frame latents when available. Still images train as plain text-to-image.
- Weight source
- For a Hub repo id, ai-toolkit loads ComfyUI repackaged weights and uses the named repo only for configs: the transformer from wan2.2_ti2v_5B_fp16.safetensors, and UMT5 from umt5_xxl_fp8_e4m3fn_scaled or umt5_xxl_fp16 (ranked by the requested qtype; a local copy of either wins). The UMT5 config and tokenizer come from ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with text_encoder/. model_kwargs.use_comfy_weights: false opts out. The text encoder is never trained.
- VAE source
- vae/ inside name_or_path (extras_name_or_path defaults to it).
- Quantization
- condition_embedder* and proj_out* stay in full precision.
- LoRA target
- WanTransformer3DModel
- Saving
- LoRAs are converted to the original Wan key layout, which ComfyUI can load. Full fine-tunes save a single file in the original layout.
- Sampling
- Uses flowmatch Euler (shift 5.0). UniPC is disabled until a diffusers regression is fixed. low_vram also turns on VAE tiling.
- Not supported
- split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
- Metadata base version
- wan_2.2_5b
Example config
not verified
job: extensionconfig: name: "my_wan22_5b_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 121 fps: 24 do_i2v: true resolution: [512, 768] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Wan-AI/Wan2.2-TI2V-5B-Diffusers" arch: "wan22_5b" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 768 height: 768 num_frames: 121 fps: 24 guidance_scale: 4 sample_steps: 30 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Wan 2.2 I2V A14B
'The image-to-video model of Wan 2.2: two 14B experts, one for the high-noise steps and one for the low-noise steps, so 14B run per step out of 28.6B total. The start frame enters as a VAE latent concatenated on channels; unlike Wan 2.1 I2V there is no CLIP image encoder. Uses the Wan 2.1 VAE.'
MiniMax-H3 (FL2VA)
'MiniMax’s open 33B single-stream DiT that generates video with native stereo audio in one packed sequence of text, video and audio tokens. This is the first/last-frame (FL2VA) partition, guidance-distilled. ai-toolkit trains the Comfy-Org pruned repack, which drops most of the 13B AdaLN parameters.'