LTX-2 19B
'Lightricks’ first joint audio-video DiT. One 19B model generates video and a synchronized stereo soundtrack together, with a 4096-wide video stream and a 2048-wide audio stream that cross-attend in every block. Video lives in a 32× spatial / 8× temporal latent space with 128 channels.'
- weights
- Lightricks/LTX-2 ↗
- org
- Lightricks
- modality
- video
- tasks
- text-to-video · image-to-video · text-to-audio-video · image-to-audio-video
- license
- LTX-2 Community License
- released
- 2026-01-03not verified
- total params
- 33.83B
- model.arch
- ltx2
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | LTX-2 19B audio-video DiT diffusers.LTX2VideoTransformer3DModel Video and audio streams in one model. The single-file ltx-2-19b-dev.safetensors in the repo root bundles this, the connectors, both VAEs and the vocoder (21.64B params). | 18.88B 18,876,174,592 | 37.76 GB | bf16+fp32 | yes |
| Text connectors | LTX-2 text connectors diffusers LTX2TextConnectors Projects the stacked Gemma hidden states (49 × 3840) down, then runs separate 2-layer video and audio connector transformers with 128 learnable registers. Frozen in ai-toolkit. | 1.43B 1,431,475,200 | 2.86 GB | bf16 | no |
| Text encoder | Gemma 3 12B transformers.Gemma3ForConditionalGeneration Counted from the 11 model-*.safetensors shards, which transformers loads. The folder also holds a second 12-shard copy (diffusion_pytorch_model-*) that bundles the connectors, so the folder total on the Hub is about double. Includes the 417M SigLIP vision tower, which ai-toolkit drops after loading. | 12.19B 12,187,325,040 | 48.75 GB | fp32 | no |
| Tokenizer | Gemma 3 SentencePiece (262k vocab) transformers.GemmaTokenizerFast | — | — | — | — |
| Video VAE | LTX-2 video VAE diffusers.AutoencoderKLLTX2Video Causal encoder, non-causal decoder. | 1.22B 1,222,479,857 | 2.44 GB | bf16 | no |
| Audio VAE | LTX-2 audio VAE diffusers.AutoencoderKLLTX2Audio Encodes a stereo 64-bin log-mel spectrogram (16 kHz, hop 160), not the raw waveform. | 53.2M 53,248,402 | 107 MB | bf16 | no |
| Vocoder | LTX-2 vocoder diffusers LTX2Vocoder Mel spectrogram to 24 kHz stereo waveform. Only used for sampling. | 55.6M 55,592,802 | 111 MB | bf16 | no |
| total | 33.83B | 92.03 GB | |||
Latent space
- autoencoder
- LTX-2 video VAE
- pixels per token
- 32×32 × 8 frames
- frame count
- 8n + 1
- notes
- The VAE does a 4×4 pixel patchify, then 8× spatial and 8× temporal downsampling, for 32× spatial total. The transformer patch is 1×1×1, so each latent voxel is one token. The first frame gets its own latent frame, which is why frame counts are 8n + 1.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 768×768 × 121f | 128×16×24×24 | 9,216 | ai-toolkit sample default |
| 768×512 × 121f | 128×16×16×24 | 6,144 | diffusers example size |
| 1024×1024 | 128×1×32×32 | 1,024 | still image |
Audio latent space
- autoencoder
- LTX-2 audio VAE
- audio channels
- stereo
- patch
- 1 latent steps per token
- tokens per second
- 25.00
- notes
- Audio is resampled to 16 kHz and turned into a 64-bin log-mel spectrogram (hop 160, 100 frames/s). The VAE compresses time 4× and mel bins 4×, so each latent step is 8 channels × 16 mel bins, packed into one 128-d token. The vocoder turns decoded mel back into 24 kHz stereo.
Architecture
- Blocks
- 48, each with a video and an audio stream
- Video stream
- 4096 (32 heads × 128), FFN 16384
- Audio stream
- 2048 (32 heads × 64), FFN 8192
- Cross-modal
- Audio-to-video and video-to-audio cross-attention in every block
- Text conditioning
- Gemma 3 12B, all 49 hidden states (3840-d) stacked, through the connectors, then cross-attention in both streams. 1024 tokens, left-padded
- Image conditioning
- None in the weights. I2V puts the clean first-frame latent in latent frame 0
- Timesteps
- Per token for video, so conditioned tokens can sit at t = 0; one timestep for audio
- Objective
- Rectified flow, resolution-dependent exponential shift (0.95 to 2.05, terminal 0.1)
- Norm / position
- QK RMSNorm across heads, split RoPE (3D video with fps, 1D audio)
In AI Toolkit
- model.arch
- ltx2
- UI label
- LTX-2 (video)
- model.name_or_path
- Lightricks/LTX-2
- source
- extra UI sections
- sample.ctrl_img, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, datasets.do_i2v, train.audio_loss_multiplier, datasets.auto_frame_count
UI defaults
- quantize / quantize_te
- true / true
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- weighted
- sample size
- 768×768, 121 frames, 24 fps
- train.audio_loss_multiplier
- 1.0
- datasets.do_audio
- true
- datasets.do_i2v
- false
- datasets.fps
- 24
- datasets.auto_frame_count
- false
- network.conv
- disabled (linear LoRA only)
Specifics
- Resolution
- Buckets and sample sizes snap to multiples of 32.
- Frame count
- Sampling rounds num_frames down to 8n + 1.
- Loading
- A Diffusers folder or repo loads each component from its subfolder. A .safetensors name_or_path (local, or repo/file on the Hub, downloaded into the models folder) is read as a single-file checkpoint: transformer, connectors, both VAEs and vocoder come from it and fp8 weights are dequantized. Single files need te_name_or_path for the text encoder.
- Text encoder source
- text_encoder/ in name_or_path, or te_name_or_path (a Gemma 3 folder, repo, or ComfyUI Gemma 3 .safetensors). The vision tower is dropped. Never trained.
- Audio
- With do_audio on, clip audio is made stereo, resampled to 16 kHz, turned into log-mel and encoded by the audio VAE. The audio stream is noised at the same timestep as video and adds its own loss, scaled by audio_loss_multiplier. Image batches and datasets without audio carry an empty audio stream and no audio loss.
- Image-to-video
- With do_i2v on, the first frame is encoded and placed in latent frame 0 at t = 0 and masked out of the loss (the mask is renormalized). Uses cached first-frame latents when available.
- Prompt length
- Embeddings are cached at their real length and left-padded to 1024 tokens before the connectors.
- LoRA target
- LTX2VideoTransformer3DModel
- Saving
- LoRAs are converted to the original LTX-2 key layout with the diffusion_model prefix, which ComfyUI can load. Full fine-tunes save the transformer as a Diffusers folder.
- Sampling
- Samples are mp4 with the generated audio track. A ctrl_img switches to the image-to-video pipeline. low_vram turns on frame-wise VAE encode and decode.
- Metadata base version
- ltx2
Example config
not verified
job: extensionconfig: name: "my_ltx2_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 32 linear_alpha: 32 save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 121 fps: 24 do_audio: true do_i2v: false resolution: [512, 768] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "weighted" audio_loss_multiplier: 1.0 optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Lightricks/LTX-2" arch: "ltx2" quantize: true quantize_te: true low_vram: true sample: sampler: "flowmatch" sample_every: 250 width: 768 height: 768 num_frames: 121 fps: 24 guidance_scale: 4 sample_steps: 25 prompts: - "a bear building a log cabin in the snow covered mountains, the sound of an axe chopping wood"Links
MiniMax-H3 Ref2VA
'The omni-reference partition of MiniMax-H3: the same 33B single-stream audio-video DiT, trained to take images and video clips as subject and style references instead of first frames. ai-toolkit trains the Comfy-Org pruned repack and shares the text encoder and VAEs with the FL2VA arch.'
LTX-2.3 22B
'The 22B update to LTX-2: the same 48-block joint audio-video DiT with gated attention, much larger 8-layer text connectors and a bandwidth-extension vocoder for 48 kHz audio. Ships as one bf16 file; ai-toolkit pairs it with a separate Gemma 3 12B text encoder.'