Docs
AI ToolkitModels

Wan 2.2 T2V A14B

'The text-to-video flagship of Wan 2.2: two 14B experts, one for the high-noise steps and one for the low-noise steps, so 14B run per step out of 28.6B total. Each expert has the Wan 2.1 14B architecture and uses the Wan 2.1 VAE. ai-toolkit trains both experts together, or either one alone.'

org
Wan-AI (Alibaba)
modality
video
tasks
text-to-video · text-to-image
license
Apache 2.0
released
2025-07-28not verified
native output
1280×720 or 832×480, 81 frames, 16 fpsnot verified
total params
34.38B
model.arch
wan22_14b:t2v

Components

rolemodelparamssizedtypetrained
Transformer (high noise)
Wan 2.2 T2V A14B high-noise expert
diffusers.WanTransformer3DModel

Runs while t > 875 (boundary_ratio 0.875). Stored as fp32 upstream. ai-toolkit defaults to ai-toolkit/Wan2.2-T2V-A14B-Diffusers-bf16 (bf16, 28,577,095,680 bytes, same parameter count) for the config, and by default loads the weights from the ComfyUI file wan2.2_t2v_high_noise_14B_fp8_scaled.

14.29B
14,288,491,584
57.15 GBfp32yes
Transformer (low noise)
Wan 2.2 T2V A14B low-noise expert
diffusers.WanTransformer3DModel

Runs while t ≤ 875. Same config as the high-noise expert, different weights. ai-toolkit loads it the same way, from transformer_2 and wan2.2_t2v_low_noise_14B_fp8_scaled.

14.29B
14,288,491,584
57.15 GBfp32yes
Text encoder
UMT5-XXL (encoder only)
transformers.UMT5EncoderModel

Includes the 256k-token embedding table. The ai-toolkit bf16 repo has no text encoder: ai-toolkit uses ai-toolkit/umt5_xxl_encoder (same parameter count and byte size), whose weights are swapped for the ComfyUI umt5_xxl file by default.

5.68B
5,680,910,336
11.36 GBbf16no
Tokenizer
UMT5 SentencePiece (256k vocab)
transformers.T5TokenizerFast

ai-toolkit loads the tokenizer from ai-toolkit/umt5_xxl_encoder instead.

————
VAE
Wan 2.1 VAE
diffusers.AutoencoderKLWan

The Wan 2.1 VAE, not the 16× Wan 2.2 VAE the 5B model uses. ai-toolkit always loads ai-toolkit/wan2.1-vae (bf16, 253,806,966 bytes, same parameter count) for this arch.

126.9M
126,892,531
508 MBfp32no
Training adapter
Accuracy recovery adapter, uint4 (both experts)adapter

Rank-16 LoRA over both experts (812 tensors each) that recovers the accuracy lost to 4-bit quantization. Used only when qtype is "uint4|ostris/accuracy_recovery_adapters/wan22_14b_t2i_torchao_uint4.safetensors" (the "4 bit with ARA" option in the UI).

155.8M
155,789,312
312 MBbf16no
total34.38B126.18 GB

Latent space

spatial
8×
temporal
4×
channels
16
patch
1×2×2
autoencoder
Wan 2.1 VAE
pixels per token
16×16 × 4 frames
frame count
4n + 1
notes
Same latent space as Wan 2.1. Causal 3D VAE: 8× spatial downsampling and two temporal downsamples. The first frame gets its own latent frame, which is why frame counts are 4n + 1.
inputlatent (c×t×h×w)tokens
1024×1024 × 41f16×11×128×12845,056ai-toolkit sample default
832×480 × 81f16×21×60×10432,760native 480p
1280×720 × 81f16×21×90×16075,600native 720p
1024×102416×1×128×1284,096still image

Architecture

Experts
2 × 14,288,491,584 params; one runs per step
Expert switch
High-noise expert for t > 875, low-noise expert below (boundary_ratio 0.875)
Blocks (per expert)
40
Hidden size
5120 (40 heads × 128)
FFN size
13824
Text conditioning
Cross-attention on UMT5 hidden states (4096-d), 512 tokens max
Image conditioning
None (text-to-video only)
Objective
Rectified flow; the upstream scheduler uses shift 3.0
Norm / position
QK RMSNorm across heads, 3D RoPE

In AI Toolkit

model.arch
wan22_14b:t2v
UI label
Wan 2.2 (14B) (video)
model.name_or_path
ai-toolkit/Wan2.2-T2V-A14B-Diffusers-bf16
extra UI sections
datasets.num_frames, model.low_vram, model.multistage, model.layer_offloading, datasets.auto_frame_count

UI defaults

quantize / quantize_te
true / true (qfloat8)
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
linear
sample size
1024×1024, 41 frames, 16 fps
datasets.fps
16
model_kwargs
train_high_noise: true, train_low_noise: true
qtype options
adds "4 bit with ARA" (uint4 + accuracy recovery adapter)
network.conv
disabled (linear LoRA only)

Specifics

Arch string
The UI sends wan22_14b:t2v. ModelConfig drops everything after the colon and loads the wan22_14b class.
Resolution
Buckets snap to multiples of 16 (8× VAE × 2×2 patch).
Two experts
Both transformers are wrapped in one DualWanTransformer3DModel that routes each forward by mean timestep: above 875 to the high-noise expert, otherwise the low-noise one. With low_vram, the inactive expert is moved to the CPU when the route changes.
Stage training
Training alternates between the stages: timesteps are drawn from 1000–875 for the high-noise expert and 875–0 for the low-noise one, switching every train.switch_boundary_every steps (UI default 1). model_kwargs.train_high_noise / train_low_noise pick which stages train; at least one must be on.
Transformer source
The ai-toolkit bf16 repo (and the upstream Wan-AI repo) supply the configs; for either repo id the weights come from Comfy-Org/Wan_2.2_ComfyUI_Repackaged wan2.2_t2v_{high,low}_noise_14B_fp8_scaled, the only registered files, even when not quantizing. A local folder loads as-is. model_kwargs.use_comfy_weights: false loads the repo's own weights.
Text encoder source
Loads ai-toolkit/umt5_xxl_encoder unless name_or_path is a local folder with a text_encoder/ subfolder. Its weights resolve to the ComfyUI umt5_xxl_fp8_e4m3fn_scaled file when quantize_te is on with qfloat8, otherwise umt5_xxl_fp16. Never trained.
VAE source
Always ai-toolkit/wan2.1-vae, whatever name_or_path is.
Quantization
Each expert is quantized separately. condition_embedder* and proj_out* stay in full precision. With the ARA, the adapter-covered linears go to uint4 and the rest of the transformer to uint8.
Loss
Flow-matching target (noise − latents). Training scheduler shift is 5.0, not the 3.0 in the upstream scheduler config.
LoRA target
DualWanTransformer3DModel when training both stages; WanTransformer3DModel (just the chosen expert) when training one.
Saving
LoRAs are split into <name>_high_noise.safetensors and <name>_low_noise.safetensors in the original Wan key layout (set network.split_multistage_loras: false for one combined file). Full fine-tunes save one file per expert with the same suffixes.
Sampling
Uses a Wan 2.2 pipeline with both experts and flowmatch Euler (shift 5.0). UniPC is disabled until a diffusers regression is fixed. low_vram also turns on VAE tiling.
Not supported
split_model_over_gpus, assistant_lora_path, inference_lora_path, lora_path
Metadata base version
wan_2.2_14b

Example config

not verified

job: extensionconfig:  name: "my_wan22_14b_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/images/or/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 1          resolution: [512, 768, 1024]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "linear"        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        switch_boundary_every: 10        cache_text_embeddings: true      model:        name_or_path: "ai-toolkit/Wan2.2-T2V-A14B-Diffusers-bf16"        arch: "wan22_14b:t2v"        quantize: true        qtype: "uint4|ostris/accuracy_recovery_adapters/wan22_14b_t2i_torchao_uint4.safetensors"        quantize_te: true        qtype_te: "qfloat8"        low_vram: true        model_kwargs:          train_high_noise: true          train_low_noise: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 1024        height: 1024        num_frames: 1        fps: 16        guidance_scale: 3.5        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains"

On this page