Wan 2.2 T2V A14B
'The text-to-video flagship of Wan 2.2: two 14B experts, one for the high-noise steps and one for the low-noise steps, so 14B run per step out of 28.6B total. Each expert has the Wan 2.1 14B architecture and uses the Wan 2.1 VAE. ai-toolkit trains both experts together, or either one alone.'
- org
- Wan-AI (Alibaba)
- modality
- video
- tasks
- text-to-video · text-to-image
- license
- Apache 2.0
- released
- 2025-07-28not verified
- native output
- 1280×720 or 832×480, 81 frames, 16 fpsnot verified
- total params
- 34.38B
- model.arch
- wan22_14b:t2v
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer (high noise) | Wan 2.2 T2V A14B high-noise expert diffusers.WanTransformer3DModel Runs while t > 875 (boundary_ratio 0.875). Stored as fp32 upstream. ai-toolkit defaults to ai-toolkit/Wan2.2-T2V-A14B-Diffusers-bf16 (bf16, 28,577,095,680 bytes, same parameter count) for the config, and by default loads the weights from the ComfyUI file wan2.2_t2v_high_noise_14B_fp8_scaled. | 14.29B 14,288,491,584 | 57.15 GB | fp32 | yes |
| Transformer (low noise) | Wan 2.2 T2V A14B low-noise expert diffusers.WanTransformer3DModel Runs while t ≤ 875. Same config as the high-noise expert, different weights. ai-toolkit loads it the same way, from transformer_2 and wan2.2_t2v_low_noise_14B_fp8_scaled. | 14.29B 14,288,491,584 | 57.15 GB | fp32 | yes |
| Text encoder | UMT5-XXL (encoder only) transformers.UMT5EncoderModel Includes the 256k-token embedding table. The ai-toolkit bf16 repo has no text encoder: ai-toolkit uses ai-toolkit/umt5_xxl_encoder (same parameter count and byte size), whose weights are swapped for the ComfyUI umt5_xxl file by default. | 5.68B 5,680,910,336 | 11.36 GB | bf16 | no |
| Tokenizer | UMT5 SentencePiece (256k vocab) transformers.T5TokenizerFast ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead. | — | — | — | — |
| VAE | Wan 2.1 VAE diffusers.AutoencoderKLWan The Wan 2.1 VAE, not the 16× Wan 2.2 VAE the 5B model uses. ai-toolkit always loads ai-toolkit/wan2.1-vae (bf16, 253,806,966 bytes, same parameter count) for this arch. | 126.9M 126,892,531 | 508 MB | fp32 | no |
| Training adapter | Accuracy recovery adapter, uint4 (both experts)adapter Rank-16 LoRA over both experts (812 tensors each) that recovers the accuracy lost to 4-bit quantization. Used only when qtype is "uint4|ostris/accuracy_recovery_adapters/wan22_14b_t2i_torchao_uint4.safetensors" (the "4 bit with ARA" option in the UI). | 155.8M 155,789,312 | 312 MB | bf16 | no |
| total | 34.38B | 126.18 GB | |||
Latent space
- autoencoder
- Wan 2.1 VAE
- pixels per token
- 16×16 × 4 frames
- frame count
- 4n + 1
- notes
- Same latent space as Wan 2.1. Causal 3D VAE: 8× spatial downsampling and two temporal downsamples. The first frame gets its own latent frame, which is why frame counts are 4n + 1.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 1024×1024 × 41f | 16×11×128×128 | 45,056 | ai-toolkit sample default |
| 832×480 × 81f | 16×21×60×104 | 32,760 | native 480p |
| 1280×720 × 81f | 16×21×90×160 | 75,600 | native 720p |
| 1024×1024 | 16×1×128×128 | 4,096 | still image |
Architecture
- Experts
- 2 × 14,288,491,584 params; one runs per step
- Expert switch
- High-noise expert for t > 875, low-noise expert below (boundary_ratio 0.875)
- Blocks (per expert)
- 40
- Hidden size
- 5120 (40 heads × 128)
- FFN size
- 13824
- Text conditioning
- Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
- Image conditioning
- None (text-to-video only)
- Objective
- Rectified flow; the upstream scheduler uses shift 3.0
- Norm / position
- QK RMSNorm across heads, 3D RoPE
In AI Toolkit
- model.arch
- wan22_14b:t2v
- UI label
- Wan 2.2 (14B) (video)
- model.name_or_path
- ai-toolkit/Wan2.2-T2V-A14B-Diffusers-bf16
- source
- extensions_built_in/diffusion_models/wan22/wan22_14b_model.py
- extensions_built_in/diffusion_models/wan22/wan22_pipeline.py
- extensions_built_in/diffusion_models/wan22/wan22_5b_model.py
- toolkit/models/wan21/wan21.py
- toolkit/models/wan21/wan_lora_convert.py
- toolkit/models/v2/diffusion_models/wan.py
- toolkit/models/v2/text_encoders/umt5.py
- toolkit/models/v2/vae/wan.py
- toolkit/models/v2/resolver.py
- toolkit/util/quantize.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- datasets.num_frames, model.low_vram, model.multistage, model.layer_offloading, datasets.auto_frame_count
UI defaults
- quantize / quantize_te
- true / true (qfloat8)
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- linear
- sample size
- 1024×1024, 41 frames, 16 fps
- datasets.fps
- 16
- model_kwargs
- train_high_noise: true, train_low_noise: true
- qtype options
- adds "4 bit with ARA" (uint4 + accuracy recovery adapter)
- network.conv
- disabled (linear LoRA only)
Specifics
- Arch string
- The UI sends wan22_14b:t2v. ModelConfig drops everything after the colon and loads the wan22_14b class.
- Resolution
- Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
- Two experts
- Both transformers are wrapped in one DualWanTransformer3DModel that routes each forward by mean timestep: above 875 to the high-noise expert, otherwise the low-noise one. With low_vram, the inactive expert is moved to the CPU when the route changes.
- Stage training
- Training alternates between the stages: timesteps are drawn from 1000–875 for the high-noise expert and 875–0 for the low-noise one, switching every train.switch_boundary_every steps (UI default 1). model_kwargs.train_high_noise / train_low_noise pick which stages train; at least one must be on.
- Transformer source
- The ai-toolkit bf16 repo (and the upstream Wan-AI repo) supply the configs; for either repo id the weights come from Comfy-Org/Wan_2.2_ComfyUI_Repackaged wan2.2_t2v_{high,low}_noise_14B_fp8_scaled, the only registered files, even when not quantizing. A local folder loads as-is. model_kwargs.use_comfy_weights: false loads the repo's own weights.
- Text encoder source
- Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
- VAE source
- Always ai-toolkit/wan2.1-vae, whatever name_or_path is.
- Quantization
- Each expert is quantized separately. condition_embedder* and proj_out* stay in full precision. With the ARA, the adapter-covered linears go to uint4 and the rest of the transformer to uint8.
- Loss
- Flow-matching target (noise − latents). Training scheduler shift is 5.0, not the 3.0 in the upstream scheduler config.
- LoRA target
- DualWanTransformer3DModel when training both stages; WanTransformer3DModel (just the chosen expert) when training one.
- Saving
- LoRAs are split into <name>_high_noise.safetensors and <name>_low_noise.safetensors in the original Wan key layout (set network.split_multistage_loras: false for one combined file). Full fine-tunes save one file per expert with the same suffixes.
- Sampling
- Uses a Wan 2.2 pipeline with both experts and flowmatch Euler (shift 5.0). UniPC is disabled until a diffusers regression is fixed. low_vram also turns on VAE tiling.
- Not supported
- split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
- Metadata base version
- wan_2.2_14b
Example config
not verified
job: extensionconfig: name: "my_wan22_14b_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/images/or/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 1 resolution: [512, 768, 1024] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "linear" optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 switch_boundary_every: 10 cache_text_embeddings: true model: name_or_path: "ai-toolkit/Wan2.2-T2V-A14B-Diffusers-bf16" arch: "wan22_14b:t2v" quantize: true qtype: "uint4|ostris/accuracy_recovery_adapters/wan22_14b_t2i_torchao_uint4.safetensors" quantize_te: true qtype_te: "qfloat8" low_vram: true model_kwargs: train_high_noise: true train_low_noise: true sample: sampler: "flowmatch" sample_every: 250 width: 1024 height: 1024 num_frames: 1 fps: 16 guidance_scale: 3.5 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains"Links
Wan 2.1 T2V 14B
'The full-size text-to-video model of Wan 2.1. A 14B-parameter DiT in the 8× spatial / 4× temporal latent space of the Wan 2.1 VAE, with UMT5-XXL text conditioning. Same architecture as each Wan 2.2 14B expert, as a single model.'
Wan 2.2 I2V A14B
'The image-to-video model of Wan 2.2: two 14B experts, one for the high-noise steps and one for the low-noise steps, so 14B run per step out of 28.6B total. The start frame enters as a VAE latent concatenated on channels; unlike Wan 2.1 I2V there is no CLIP image encoder. Uses the Wan 2.1 VAE.'