LTX-2.3 22B
'The 22B update to LTX-2: the same 48-block joint audio-video DiT with gated attention, much larger 8-layer text connectors and a bandwidth-extension vocoder for 48 kHz audio. Ships as one bf16 file; ai-toolkit pairs it with a separate Gemma 3 12B text encoder.'
- weights
- Lightricks/LTX-2.3 ↗
- org
- Lightricks
- modality
- video
- tasks
- text-to-video · image-to-video · text-to-audio-video · image-to-audio-video
- license
- LTX-2 Community License
- released
- 2026-03-04not verified
- total params
- 35.26B
- model.arch
- ltx2.3
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | LTX-2.3 audio-video DiT diffusers.LTX2VideoTransformer3DModel Bundled in ltx-2.3-22b-dev.safetensors (23.07B params, 46.15 GB in total) with every part below except the text encoder. Counted from the model.diffusion_model.* tensors that are not connector weights; size is tensor bytes. | 18.99B 18,988,838,144 | 37.99 GB | bf16+fp32 | yes |
| Text connectors | LTX-2.3 text connectors diffusers LTX2TextConnectors The video and audio connector blocks (2.02B) plus the top-level text_embedding_projection (1.16B), which ai-toolkit folds into the connectors. Bundled in the single file. | 3.17B 3,172,227,584 | 6.34 GB | bf16 | no |
| Text encoder | Gemma 3 12B IT (QAT, unquantized) transformers.Gemma3ForConditionalGeneration Not in the LTX-2.3 repo. ai-toolkit loads this repo unless te_name_or_path is set. Includes the SigLIP vision tower, which ai-toolkit drops. | 12.19B 12,187,325,040 | 24.37 GB | bf16 | no |
| Tokenizer | Gemma 3 SentencePiece (262k vocab) transformers.GemmaTokenizerFast | — | — | — | — |
| Video VAE | LTX-2.3 video VAE diffusers.AutoencoderKLLTX2Video The vae.* tensors of the single file. Smaller decoder than the LTX-2 VAE, same latent space. | 726.1M 726,106,225 | 1.45 GB | bf16 | no |
| Audio VAE | LTX-2 audio VAE diffusers.AutoencoderKLLTX2Audio The audio_vae.* tensors of the single file. Same config and parameter count as LTX-2’s. | 53.2M 53,248,402 | 106 MB | bf16 | no |
| Vocoder | LTX-2.3 vocoder with BWE diffusers LTX2VocoderWithBWE The vocoder.* tensors of the single file. A 16 kHz vocoder followed by a bandwidth-extension stage that outputs 48 kHz stereo. Only used for sampling. | 129.1M 129,085,032 | 258 MB | bf16 | no |
| total | 35.26B | 70.52 GB | |||
Latent space
- autoencoder
- LTX-2.3 video VAE
- pixels per token
- 32×32 × 8 frames
- frame count
- 8n + 1
- notes
- Same latent space as LTX-2: a 4×4 pixel patchify, then 8× spatial and 8× temporal downsampling. The transformer patch is 1×1×1, so each latent voxel is one token. Width and height must be multiples of 32.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 768×768 × 121f | 128×16×24×24 | 9,216 | ai-toolkit sample default |
| 768×512 × 121f | 128×16×16×24 | 6,144 | |
| 1024×1024 | 128×1×32×32 | 1,024 | still image |
Audio latent space
- autoencoder
- LTX-2 audio VAE
- audio channels
- stereo
- patch
- 1 latent steps per token
- tokens per second
- 25.00
- notes
- Audio is resampled to 16 kHz and turned into a 64-bin log-mel spectrogram (hop 160, 100 frames/s). The VAE compresses time 4× and mel bins 4×, so each latent step is 8 channels × 16 mel bins, packed into one 128-d token. The BWE vocoder decodes to 48 kHz stereo.
Architecture
- Blocks
- 48, each with a video and an audio stream
- Video stream
- 4096 (32 heads × 128), FFN 16384
- Audio stream
- 2048 (32 heads × 64), FFN 8192
- Cross-modal
- Audio-to-video and video-to-audio cross-attention in every block, modulated by the other stream’s timestep
- Text conditioning
- Gemma 3 12B, all 49 hidden states (3840-d) stacked, projected per modality, through 8-layer video and audio connectors, then cross-attention. 1024 tokens, left-padded
- Image conditioning
- None in the weights. I2V puts the clean first-frame latent in latent frame 0
- Timesteps
- Per token for video; one timestep for audio
- Objective
- Rectified flow, resolution-dependent exponential shift (0.95 to 2.05, terminal 0.1)
- Changes from LTX-2
- Gated attention, cross-attention AdaLN, perturbed-attention support for STG guidance, bigger connectors with per-modality projections, BWE vocoder
- Norm / position
- QK RMSNorm across heads, split RoPE (3D video with fps, 1D audio)
In AI Toolkit
- model.arch
- ltx2.3
- UI label
- LTX-2.3 (video)
- model.name_or_path
- Lightricks/LTX-2.3/ltx-2.3-22b-dev.safetensors
- source
- extra UI sections
- sample.ctrl_img, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, datasets.do_i2v, train.audio_loss_multiplier, datasets.auto_frame_count
UI defaults
- quantize / quantize_te
- true / true
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- sample size
- 768×768, 121 frames, 24 fps
- train.audio_loss_multiplier
- 1.0
- datasets.do_audio
- true
- datasets.do_i2v
- false
- datasets.cache_latents_to_disk
- true
- datasets.fps
- 24
- datasets.auto_frame_count
- false
- network.conv
- disabled (linear LoRA only)
Specifics
- Resolution
- Buckets and sample sizes snap to multiples of 32.
- Frame count
- Sampling rounds num_frames down to 8n + 1.
- Loading
- name_or_path is a single .safetensors file. A repo/file path is looked for in the models folder first and downloaded there if missing. Transformer, connectors, both VAEs and the vocoder are converted from it to Diffusers modules; fp8 files are dequantized first.
- Text encoder source
- Lightricks/gemma-3-12b-it-qat-q4_0-unquantized, unless te_name_or_path is set (a Gemma 3 folder, repo, or ComfyUI Gemma 3 .safetensors). The vision tower is dropped. Never trained.
- Audio
- With do_audio on, clip audio is made stereo, resampled to 16 kHz, turned into log-mel and encoded by the audio VAE. The audio stream is noised at the same timestep as video and adds its own loss, scaled by audio_loss_multiplier. Image batches and datasets without audio carry an empty audio stream and no audio loss.
- Image-to-video
- With do_i2v on, the first frame is encoded and placed in latent frame 0 at t = 0 and masked out of the loss (the mask is renormalized).
- Cross timestep
- Training and sampling pass use_cross_timestep, the 2.3 behaviour, so each stream’s cross-attention is modulated by the other stream’s sigma.
- LoRA target
- LTX2VideoTransformer3DModel
- Saving
- LoRAs are converted to the original LTX-2.3 key layout with the diffusion_model prefix, which ComfyUI can load. Full fine-tunes save the transformer as a Diffusers folder.
- Sampling
- Samples are mp4 with audio. Adds STG (scale 1, block 28), modality guidance 3, guidance rescale 0.7 and audio guidance 7 on top of guidance_scale. low_vram turns on frame-wise VAE encode and decode.
- Metadata base version
- ltx2
Example config
not verified
job: extensionconfig: name: "my_ltx23_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 121 fps: 24 do_audio: true do_i2v: false cache_latents_to_disk: true resolution: [512, 768] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" audio_loss_multiplier: 1.0 optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Lightricks/LTX-2.3/ltx-2.3-22b-dev.safetensors" arch: "ltx2.3" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 768 height: 768 num_frames: 121 fps: 24 guidance_scale: 3 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains, the sound of an axe chopping wood"Links
LTX-2 19B
'Lightricks’ first joint audio-video DiT. One 19B model generates video and a synchronized stereo soundtrack together, with a 4096-wide video stream and a 2048-wide audio stream that cross-attend in every block. Video lives in a 32× spatial / 8× temporal latent space with 128 channels.'
LTX-2.5 22B
'The next LTX-2.x release: the LTX-2.3 joint audio-video DiT with bias-free video feed-forwards and a Gemma 4 12B text encoder. Ships as ComfyUI-style split files, including pre-quantized int8 ConvRot transformer and text encoder files that ai-toolkit loads by default.'