Docs
AI ToolkitModels

LTX-2 19B

'Lightricks’ first joint audio-video DiT. One 19B model generates video and a synchronized stereo soundtrack together, with a 4096-wide video stream and a 2048-wide audio stream that cross-attend in every block. Video lives in a 32× spatial / 8× temporal latent space with 128 channels.'

org
Lightricks
modality
video
tasks
text-to-video · image-to-video · text-to-audio-video · image-to-audio-video
license
LTX-2 Community License
released
2026-01-03not verified
total params
33.83B
model.arch
ltx2

Components

rolemodelparamssizedtypetrained
Transformer
LTX-2 19B audio-video DiT
diffusers.LTX2VideoTransformer3DModel

Video and audio streams in one model. The single-file ltx-2-19b-dev.safetensors in the repo root bundles this, the connectors, both VAEs and the vocoder (21.64B params).

18.88B
18,876,174,592
37.76 GBbf16+fp32yes
Text connectors
LTX-2 text connectors
diffusers LTX2TextConnectors

Projects the stacked Gemma hidden states (49 × 3840) down, then runs separate 2-layer video and audio connector transformers with 128 learnable registers. Frozen in ai-toolkit.

1.43B
1,431,475,200
2.86 GBbf16no
Text encoder
Gemma 3 12B
transformers.Gemma3ForConditionalGeneration

Counted from the 11 model-*.safetensors shards, which transformers loads. The folder also holds a second 12-shard copy (diffusion_pytorch_model-*) that bundles the connectors, so the folder total on the Hub is about double. Includes the 417M SigLIP vision tower, which ai-toolkit drops after loading.

12.19B
12,187,325,040
48.75 GBfp32no
Tokenizer
Gemma 3 SentencePiece (262k vocab)
transformers.GemmaTokenizerFast
————
Video VAE
LTX-2 video VAE
diffusers.AutoencoderKLLTX2Video

Causal encoder, non-causal decoder.

1.22B
1,222,479,857
2.44 GBbf16no
Audio VAE
LTX-2 audio VAE
diffusers.AutoencoderKLLTX2Audio

Encodes a stereo 64-bin log-mel spectrogram (16 kHz, hop 160), not the raw waveform.

53.2M
53,248,402
107 MBbf16no
Vocoder
LTX-2 vocoder
diffusers LTX2Vocoder

Mel spectrogram to 24 kHz stereo waveform. Only used for sampling.

55.6M
55,592,802
111 MBbf16no
total33.83B92.03 GB

Latent space

spatial
32×
temporal
8×
channels
128
patch
1×1×1
autoencoder
LTX-2 video VAE
pixels per token
32×32 × 8 frames
frame count
8n + 1
notes
The VAE does a 4×4 pixel patchify, then 8× spatial and 8× temporal downsampling, for 32× spatial total. The transformer patch is 1×1×1, so each latent voxel is one token. The first frame gets its own latent frame, which is why frame counts are 8n + 1.
inputlatent (c×t×h×w)tokens
768×768 × 121f128×16×24×249,216ai-toolkit sample default
768×512 × 121f128×16×16×246,144diffusers example size
1024×1024128×1×32×321,024still image

Audio latent space

sample rate
16 kHz
temporal
640×
channels
8
latent rate
25/s
autoencoder
LTX-2 audio VAE
audio channels
stereo
patch
1 latent steps per token
tokens per second
25.00
notes
Audio is resampled to 16 kHz and turned into a 64-bin log-mel spectrogram (hop 160, 100 frames/s). The VAE compresses time 4× and mel bins 4×, so each latent step is 8 channels × 16 mel bins, packed into one 128-d token. The vocoder turns decoded mel back into 24 kHz stereo.

Architecture

Blocks
48, each with a video and an audio stream
Video stream
4096 (32 heads × 128), FFN 16384
Audio stream
2048 (32 heads × 64), FFN 8192
Cross-modal
Audio-to-video and video-to-audio cross-attention in every block
Text conditioning
Gemma 3 12B, all 49 hidden states (3840-d) stacked, through the connectors, then cross-attention in both streams. 1024 tokens, left-padded
Image conditioning
None in the weights. I2V puts the clean first-frame latent in latent frame 0
Timesteps
Per token for video, so conditioned tokens can sit at t = 0; one timestep for audio
Objective
Rectified flow, resolution-dependent exponential shift (0.95 to 2.05, terminal 0.1)
Norm / position
QK RMSNorm across heads, split RoPE (3D video with fps, 1D audio)

In AI Toolkit

model.arch
ltx2
UI label
LTX-2 (video)
model.name_or_path
Lightricks/LTX-2
extra UI sections
sample.ctrl_img, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, datasets.do_i2v, train.audio_loss_multiplier, datasets.auto_frame_count

UI defaults

quantize / quantize_te
true / true
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
sample size
768×768, 121 frames, 24 fps
train.audio_loss_multiplier
1.0
datasets.do_audio
true
datasets.do_i2v
false
datasets.fps
24
datasets.auto_frame_count
false
network.conv
disabled (linear LoRA only)

Specifics

Resolution
Buckets and sample sizes snap to multiples of 32.
Frame count
Sampling rounds num_frames down to 8n + 1.
Loading
A Diffusers folder or repo loads each component from its subfolder. A .safetensors name_or_path (local, or repo/file on the Hub, downloaded into the models folder) is read as a single-file checkpoint: transformer, connectors, both VAEs and vocoder come from it and fp8 weights are dequantized. Single files need te_name_or_path for the text encoder.
Text encoder source
text_encoder/ in name_or_path, or te_name_or_path (a Gemma 3 folder, repo, or ComfyUI Gemma 3 .safetensors). The vision tower is dropped. Never trained.
Audio
With do_audio on, clip audio is made stereo, resampled to 16 kHz, turned into log-mel and encoded by the audio VAE. The audio stream is noised at the same timestep as video and adds its own loss, scaled by audio_loss_multiplier. Image batches and datasets without audio carry an empty audio stream and no audio loss.
Image-to-video
With do_i2v on, the first frame is encoded and placed in latent frame 0 at t = 0 and masked out of the loss (the mask is renormalized). Uses cached first-frame latents when available.
Prompt length
Embeddings are cached at their real length and left-padded to 1024 tokens before the connectors.
LoRA target
LTX2VideoTransformer3DModel
Saving
LoRAs are converted to the original LTX-2 key layout with the diffusion_model prefix, which ComfyUI can load. Full fine-tunes save the transformer as a Diffusers folder.
Sampling
Samples are mp4 with the generated audio track. A ctrl_img switches to the image-to-video pipeline. low_vram turns on frame-wise VAE encode and decode.
Metadata base version
ltx2

Example config

not verified

job: extensionconfig:  name: "my_ltx2_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 121          fps: 24          do_audio: true          do_i2v: false          resolution: [512, 768]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        audio_loss_multiplier: 1.0        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Lightricks/LTX-2"        arch: "ltx2"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 768        height: 768        num_frames: 121        fps: 24        guidance_scale: 4        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains, the sound of an axe chopping wood"

On this page