Docs
AI ToolkitModels

LTX-2.5 22B

'The next LTX-2.x release: the LTX-2.3 joint audio-video DiT with bias-free video feed-forwards and a Gemma 4 12B text encoder. Ships as ComfyUI-style split files, including pre-quantized int8 ConvRot transformer and text encoder files that ai-toolkit loads by default.'

org
Lightricks
modality
video
tasks
text-to-video · image-to-video · text-to-audio-video · image-to-audio-video
license
LTX-2.x Community License
released
2026-07-23not verified
total params
34.98B
model.arch
ltx2.5

Components

rolemodelparamssizedtypetrained
Transformer
LTX-2.5 22B audio-video DiT
diffusers.LTX2VideoTransformer3DModel

Counted from the bf16 file, which also holds the connector blocks (21.00B params, 42.02 GB in total); size is tensor bytes. ai-toolkit loads ltx-2.5-22b-dev-transformer-comfy-int8-convrot.safetensors by default (21.50 GB), pre-quantized int8 ConvRot, as is.

18.99B
18,987,859,200
37.99 GBbf16+fp32yes
Text connectors
LTX-2.5 text connectors
diffusers LTX2TextConnectors

Split across two files: the connector blocks (2.02B) are in the transformer file and text_embedding_projection (1.16B) is in the text encoder file. ai-toolkit joins them into one module. Same shape as LTX-2.3’s.

3.17B
3,172,227,584
6.34 GBbf16no
Text encoder
Gemma 4 12B (text decoder)
transformers Gemma4TextModel

The model.* tensors of the bf16 file (13.15B tensor elements, 26.26 GB in total). The file also holds the connectors’ text projection, a vision tower, multimodal and audio projectors and the tokenizer; ai-toolkit uses only the text decoder and the projection. Loads gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors by default (15.37 GB).

11.91B
11,907,350,320
23.81 GBbf16no
Tokenizer
Gemma 4 tokenizer (262k vocab)
transformers.AutoTokenizer

Stored inside the text encoder file as uint8 tensors. ai-toolkit writes it out to a folder next to the file on first load.

————
Video VAE
LTX-2.5 conv video VAE
diffusers.AutoencoderKLLTX2Video

The classic convolutional VAE. The repo’s default ltx-2.5-video-vae-bf16.safetensors has a new diffusion decoder (736M params) that Diffusers has no class for, so ai-toolkit uses this one. Both share the same encoder.

726.1M
726,106,225
1.45 GBbf16no
Audio VAE
LTX-2 audio VAE
diffusers.AutoencoderKLLTX2Audio

The audio_vae.* tensors of ltx-2.5-audio-vae-bf16.safetensors, which also holds the vocoder.

53.2M
53,248,402
106 MBbf16no
Vocoder
LTX-2.x vocoder with BWE
diffusers LTX2VocoderWithBWE

The vocoder.* tensors of the audio VAE file. Bandwidth extension to 48 kHz stereo. Only used for sampling.

129.1M
129,085,032
258 MBbf16no
total34.98B69.96 GB

Latent space

spatial
32×
temporal
8×
channels
128
patch
1×1×1
autoencoder
LTX-2.5 conv video VAE
pixels per token
32×32 × 8 frames
frame count
8n + 1
notes
Same latent space as LTX-2 and 2.3: a 4×4 pixel patchify, then 8× spatial and 8× temporal downsampling. The transformer patch is 1×1×1, so each latent voxel is one token.
inputlatent (c×t×h×w)tokens
768×768 × 121f128×16×24×249,216ai-toolkit sample default
768×512 × 121f128×16×16×246,144
1024×1024128×1×32×321,024still image

Audio latent space

sample rate
16 kHz
temporal
640×
channels
8
latent rate
25/s
autoencoder
LTX-2 audio VAE
audio channels
stereo
patch
1 latent steps per token
tokens per second
25.00
notes
Audio is resampled to 16 kHz and turned into a 64-bin log-mel spectrogram (hop 160, 100 frames/s). The VAE compresses time 4× and mel bins 4×, so each latent step is 8 channels × 16 mel bins, packed into one 128-d token. The BWE vocoder decodes to 48 kHz stereo.

Architecture

Blocks
48, each with a video and an audio stream
Video stream
4096 (32 heads × 128), FFN 16384 without biases
Audio stream
2048 (32 heads × 64), FFN 8192
Cross-modal
Audio-to-video and video-to-audio cross-attention in every block, modulated by the other stream’s timestep
Text conditioning
Gemma 4 12B (48 layers, 3840-d), all hidden states stacked, projected per modality, through 8-layer video and audio connectors, then cross-attention. 1024 tokens, left-padded
Image conditioning
None in the weights. I2V puts the clean first-frame latent in latent frame 0
Timesteps
Per token for video; one timestep for audio
Objective
Rectified flow, resolution-dependent exponential shift (0.95 to 2.05, terminal 0.1)
Changes from LTX-2.3
Gemma 4 text encoder, no video feed-forward biases, a learned keyframe position embedding (not used by the regular forward pass)
Norm / position
QK RMSNorm across heads, split RoPE (3D video with fps, 1D audio)

In AI Toolkit

model.arch
ltx2.5
UI label
LTX-2.5 (video)
model.name_or_path
Lightricks/LTX-2.5
extra UI sections
sample.ctrl_img, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, datasets.do_i2v, train.audio_loss_multiplier, datasets.auto_frame_count

UI defaults

quantize / qtype
true / convrot8
quantize_te / qtype_te
true / convrot8
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
sample size
768×768, 121 frames, 24 fps
train.audio_loss_multiplier
1.0
datasets.do_audio
true
datasets.do_i2v
false
datasets.cache_latents_to_disk
true
datasets.fps
24
datasets.auto_frame_count
false
network.conv
disabled (linear LoRA only)

Specifics

Gated repo
Lightricks/LTX-2.5 needs its terms accepted on the Hub and a token before the first download.
Loading
Four files resolve under the models folder at their ComfyUI paths and download there only when missing: the int8 ConvRot dev transformer, the int8 ConvRot Gemma 4 file, the conv video VAE and the audio VAE + vocoder. Each can be overridden with model_kwargs dit_path, text_encoder_path, video_vae_path and audio_vae_path.
Other files
A .safetensors name_or_path is used as the transformer file and te_name_or_path as the text encoder file, so the bf16 files work too.
Quantization
The int8 ConvRot layers attach as is (convrot8 matches the files, so nothing is re-quantized). Picking another qtype re-quantizes layer by layer. Fp32 scale-shift tables stay pinned in fp32.
Resolution
Buckets and sample sizes snap to multiples of 32.
Frame count
Sampling rounds num_frames down to 8n + 1.
Audio
With do_audio on, clip audio is made stereo, resampled to 16 kHz, turned into log-mel and encoded by the audio VAE. The audio stream adds its own loss, scaled by audio_loss_multiplier. Image batches and datasets without audio carry an empty audio stream and no audio loss.
Image-to-video
With do_i2v on, the first frame is encoded and placed in latent frame 0 at t = 0 and masked out of the loss (the mask is renormalized).
LoRA target
LTX2VideoTransformer3DModel
Saving
LoRAs are converted to the original LTX key layout with the diffusion_model prefix, which ComfyUI can load. Full fine-tunes save the transformer as a Diffusers folder.
Sampling
Samples are mp4 with audio. Same guidance extras as LTX-2.3: STG (scale 1, block 28), modality guidance 3, guidance rescale 0.7, audio guidance 7.
Metadata base version
ltx2

Example config

not verified

job: extensionconfig:  name: "my_ltx25_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 121          fps: 24          do_audio: true          do_i2v: false          cache_latents_to_disk: true          resolution: [512, 768]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        audio_loss_multiplier: 1.0        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Lightricks/LTX-2.5"        arch: "ltx2.5"        quantize: true        qtype: "convrot8"        quantize_te: true        qtype_te: "convrot8"        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 768        height: 768        num_frames: 121        fps: 24        guidance_scale: 3        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains, the sound of an axe chopping wood"

On this page