LTX-2.5 22B
'The next LTX-2.x release: the LTX-2.3 joint audio-video DiT with bias-free video feed-forwards and a Gemma 4 12B text encoder. Ships as ComfyUI-style split files, including pre-quantized int8 ConvRot transformer and text encoder files that ai-toolkit loads by default.'
- weights
- Lightricks/LTX-2.5 ↗
- org
- Lightricks
- modality
- video
- tasks
- text-to-video · image-to-video · text-to-audio-video · image-to-audio-video
- license
- LTX-2.x Community License
- released
- 2026-07-23not verified
- total params
- 34.98B
- model.arch
- ltx2.5
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | LTX-2.5 22B audio-video DiT diffusers.LTX2VideoTransformer3DModel Counted from the bf16 file, which also holds the connector blocks (21.00B params, 42.02 GB in total); size is tensor bytes. ai-toolkit loads ltx-2.5-22b-dev-transformer-comfy-int8-convrot.safetensors by default (21.50 GB), pre-quantized int8 ConvRot, as is. | 18.99B 18,987,859,200 | 37.99 GB | bf16+fp32 | yes |
| Text connectors | LTX-2.5 text connectors diffusers LTX2TextConnectors Split across two files: the connector blocks (2.02B) are in the transformer file and text_embedding_projection (1.16B) is in the text encoder file. ai-toolkit joins them into one module. Same shape as LTX-2.3’s. | 3.17B 3,172,227,584 | 6.34 GB | bf16 | no |
| Text encoder | Gemma 4 12B (text decoder) transformers Gemma4TextModel The model.* tensors of the bf16 file (13.15B tensor elements, 26.26 GB in total). The file also holds the connectors’ text projection, a vision tower, multimodal and audio projectors and the tokenizer; ai-toolkit uses only the text decoder and the projection. Loads gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors by default (15.37 GB). | 11.91B 11,907,350,320 | 23.81 GB | bf16 | no |
| Tokenizer | Gemma 4 tokenizer (262k vocab) transformers.AutoTokenizer Stored inside the text encoder file as uint8 tensors. ai-toolkit writes it out to a folder next to the file on first load. | — | — | — | — |
| Video VAE | LTX-2.5 conv video VAE diffusers.AutoencoderKLLTX2Video The classic convolutional VAE. The repo’s default ltx-2.5-video-vae-bf16.safetensors has a new diffusion decoder (736M params) that Diffusers has no class for, so ai-toolkit uses this one. Both share the same encoder. | 726.1M 726,106,225 | 1.45 GB | bf16 | no |
| Audio VAE | LTX-2 audio VAE diffusers.AutoencoderKLLTX2Audio The audio_vae.* tensors of ltx-2.5-audio-vae-bf16.safetensors, which also holds the vocoder. | 53.2M 53,248,402 | 106 MB | bf16 | no |
| Vocoder | LTX-2.x vocoder with BWE diffusers LTX2VocoderWithBWE The vocoder.* tensors of the audio VAE file. Bandwidth extension to 48 kHz stereo. Only used for sampling. | 129.1M 129,085,032 | 258 MB | bf16 | no |
| total | 34.98B | 69.96 GB | |||
Latent space
- autoencoder
- LTX-2.5 conv video VAE
- pixels per token
- 32×32 × 8 frames
- frame count
- 8n + 1
- notes
- Same latent space as LTX-2 and 2.3: a 4×4 pixel patchify, then 8× spatial and 8× temporal downsampling. The transformer patch is 1×1×1, so each latent voxel is one token.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 768×768 × 121f | 128×16×24×24 | 9,216 | ai-toolkit sample default |
| 768×512 × 121f | 128×16×16×24 | 6,144 | |
| 1024×1024 | 128×1×32×32 | 1,024 | still image |
Audio latent space
- autoencoder
- LTX-2 audio VAE
- audio channels
- stereo
- patch
- 1 latent steps per token
- tokens per second
- 25.00
- notes
- Audio is resampled to 16 kHz and turned into a 64-bin log-mel spectrogram (hop 160, 100 frames/s). The VAE compresses time 4× and mel bins 4×, so each latent step is 8 channels × 16 mel bins, packed into one 128-d token. The BWE vocoder decodes to 48 kHz stereo.
Architecture
- Blocks
- 48, each with a video and an audio stream
- Video stream
- 4096 (32 heads × 128), FFN 16384 without biases
- Audio stream
- 2048 (32 heads × 64), FFN 8192
- Cross-modal
- Audio-to-video and video-to-audio cross-attention in every block, modulated by the other stream’s timestep
- Text conditioning
- Gemma 4 12B (48 layers, 3840-d), all hidden states stacked, projected per modality, through 8-layer video and audio connectors, then cross-attention. 1024 tokens, left-padded
- Image conditioning
- None in the weights. I2V puts the clean first-frame latent in latent frame 0
- Timesteps
- Per token for video; one timestep for audio
- Objective
- Rectified flow, resolution-dependent exponential shift (0.95 to 2.05, terminal 0.1)
- Changes from LTX-2.3
- Gemma 4 text encoder, no video feed-forward biases, a learned keyframe position embedding (not used by the regular forward pass)
- Norm / position
- QK RMSNorm across heads, split RoPE (3D video with fps, 1D audio)
In AI Toolkit
- model.arch
- ltx2.5
- UI label
- LTX-2.5 (video)
- model.name_or_path
- Lightricks/LTX-2.5
- source
- extensions_built_in/diffusion_models/ltx2/ltx2.py
- extensions_built_in/diffusion_models/ltx2/convert_ltx2_to_diffusers.py
- toolkit/models/v2/diffusion_models/ltx2.py
- toolkit/models/v2/vae/ltx2.py
- toolkit/models/v2/text_encoders/gemma3.py
- toolkit/models/v2/resolver.py
- toolkit/util/comfy_quant_import.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- sample.ctrl_img, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, datasets.do_i2v, train.audio_loss_multiplier, datasets.auto_frame_count
UI defaults
- quantize / qtype
- true / convrot8
- quantize_te / qtype_te
- true / convrot8
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- sample size
- 768×768, 121 frames, 24 fps
- train.audio_loss_multiplier
- 1.0
- datasets.do_audio
- true
- datasets.do_i2v
- false
- datasets.cache_latents_to_disk
- true
- datasets.fps
- 24
- datasets.auto_frame_count
- false
- network.conv
- disabled (linear LoRA only)
Specifics
- Gated repo
- Lightricks/LTX-2.5 needs its terms accepted on the Hub and a token before the first download.
- Loading
- Four files resolve under the models folder at their ComfyUI paths and download there only when missing: the int8 ConvRot dev transformer, the int8 ConvRot Gemma 4 file, the conv video VAE and the audio VAE + vocoder. Each can be overridden with model_kwargs dit_path, text_encoder_path, video_vae_path and audio_vae_path.
- Other files
- A .safetensors name_or_path is used as the transformer file and te_name_or_path as the text encoder file, so the bf16 files work too.
- Quantization
- The int8 ConvRot layers attach as is (convrot8 matches the files, so nothing is re-quantized). Picking another qtype re-quantizes layer by layer. Fp32 scale-shift tables stay pinned in fp32.
- Resolution
- Buckets and sample sizes snap to multiples of 32.
- Frame count
- Sampling rounds num_frames down to 8n + 1.
- Audio
- With do_audio on, clip audio is made stereo, resampled to 16 kHz, turned into log-mel and encoded by the audio VAE. The audio stream adds its own loss, scaled by audio_loss_multiplier. Image batches and datasets without audio carry an empty audio stream and no audio loss.
- Image-to-video
- With do_i2v on, the first frame is encoded and placed in latent frame 0 at t = 0 and masked out of the loss (the mask is renormalized).
- LoRA target
- LTX2VideoTransformer3DModel
- Saving
- LoRAs are converted to the original LTX key layout with the diffusion_model prefix, which ComfyUI can load. Full fine-tunes save the transformer as a Diffusers folder.
- Sampling
- Samples are mp4 with audio. Same guidance extras as LTX-2.3: STG (scale 1, block 28), modality guidance 3, guidance rescale 0.7, audio guidance 7.
- Metadata base version
- ltx2
Example config
not verified
job: extensionconfig: name: "my_ltx25_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 121 fps: 24 do_audio: true do_i2v: false cache_latents_to_disk: true resolution: [512, 768] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" audio_loss_multiplier: 1.0 optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Lightricks/LTX-2.5" arch: "ltx2.5" quantize: true qtype: "convrot8" quantize_te: true qtype_te: "convrot8" low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 768 height: 768 num_frames: 121 fps: 24 guidance_scale: 3 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains, the sound of an axe chopping wood"Links
LTX-2.3 22B
'The 22B update to LTX-2: the same 48-block joint audio-video DiT with gated attention, much larger 8-layer text connectors and a bandwidth-extension vocoder for 48 kHz audio. Ships as one bf16 file; ai-toolkit pairs it with a separate Gemma 3 12B text encoder.'
ACE-Step 1.5 XL
The XL (4B) base DiT of ACE-Step 1.5. It writes full songs from a style caption, lyrics and metadata, using the same condition encoder, text encoder and 48 kHz stereo Oobleck VAE (1920×, 25 latent frames per second) as the 2B model.