Docs
AI ToolkitModels

LTX-2.3 22B

'The 22B update to LTX-2: the same 48-block joint audio-video DiT with gated attention, much larger 8-layer text connectors and a bandwidth-extension vocoder for 48 kHz audio. Ships as one bf16 file; ai-toolkit pairs it with a separate Gemma 3 12B text encoder.'

org
Lightricks
modality
video
tasks
text-to-video · image-to-video · text-to-audio-video · image-to-audio-video
license
LTX-2 Community License
released
2026-03-04not verified
total params
35.26B
model.arch
ltx2.3

Components

rolemodelparamssizedtypetrained
Transformer
LTX-2.3 audio-video DiT
diffusers.LTX2VideoTransformer3DModel

Bundled in ltx-2.3-22b-dev.safetensors (23.07B params, 46.15 GB in total) with every part below except the text encoder. Counted from the model.diffusion_model.* tensors that are not connector weights; size is tensor bytes.

18.99B
18,988,838,144
37.99 GBbf16+fp32yes
Text connectors
LTX-2.3 text connectors
diffusers LTX2TextConnectors

The video and audio connector blocks (2.02B) plus the top-level text_embedding_projection (1.16B), which ai-toolkit folds into the connectors. Bundled in the single file.

3.17B
3,172,227,584
6.34 GBbf16no
Text encoder
Gemma 3 12B IT (QAT, unquantized)
transformers.Gemma3ForConditionalGeneration

Not in the LTX-2.3 repo. ai-toolkit loads this repo unless te_name_or_path is set. Includes the SigLIP vision tower, which ai-toolkit drops.

12.19B
12,187,325,040
24.37 GBbf16no
Tokenizer
Gemma 3 SentencePiece (262k vocab)
transformers.GemmaTokenizerFast
————
Video VAE
LTX-2.3 video VAE
diffusers.AutoencoderKLLTX2Video

The vae.* tensors of the single file. Smaller decoder than the LTX-2 VAE, same latent space.

726.1M
726,106,225
1.45 GBbf16no
Audio VAE
LTX-2 audio VAE
diffusers.AutoencoderKLLTX2Audio

The audio_vae.* tensors of the single file. Same config and parameter count as LTX-2’s.

53.2M
53,248,402
106 MBbf16no
Vocoder
LTX-2.3 vocoder with BWE
diffusers LTX2VocoderWithBWE

The vocoder.* tensors of the single file. A 16 kHz vocoder followed by a bandwidth-extension stage that outputs 48 kHz stereo. Only used for sampling.

129.1M
129,085,032
258 MBbf16no
total35.26B70.52 GB

Latent space

spatial
32×
temporal
8×
channels
128
patch
1×1×1
autoencoder
LTX-2.3 video VAE
pixels per token
32×32 × 8 frames
frame count
8n + 1
notes
Same latent space as LTX-2: a 4×4 pixel patchify, then 8× spatial and 8× temporal downsampling. The transformer patch is 1×1×1, so each latent voxel is one token. Width and height must be multiples of 32.
inputlatent (c×t×h×w)tokens
768×768 × 121f128×16×24×249,216ai-toolkit sample default
768×512 × 121f128×16×16×246,144
1024×1024128×1×32×321,024still image

Audio latent space

sample rate
16 kHz
temporal
640×
channels
8
latent rate
25/s
autoencoder
LTX-2 audio VAE
audio channels
stereo
patch
1 latent steps per token
tokens per second
25.00
notes
Audio is resampled to 16 kHz and turned into a 64-bin log-mel spectrogram (hop 160, 100 frames/s). The VAE compresses time 4× and mel bins 4×, so each latent step is 8 channels × 16 mel bins, packed into one 128-d token. The BWE vocoder decodes to 48 kHz stereo.

Architecture

Blocks
48, each with a video and an audio stream
Video stream
4096 (32 heads × 128), FFN 16384
Audio stream
2048 (32 heads × 64), FFN 8192
Cross-modal
Audio-to-video and video-to-audio cross-attention in every block, modulated by the other stream’s timestep
Text conditioning
Gemma 3 12B, all 49 hidden states (3840-d) stacked, projected per modality, through 8-layer video and audio connectors, then cross-attention. 1024 tokens, left-padded
Image conditioning
None in the weights. I2V puts the clean first-frame latent in latent frame 0
Timesteps
Per token for video; one timestep for audio
Objective
Rectified flow, resolution-dependent exponential shift (0.95 to 2.05, terminal 0.1)
Changes from LTX-2
Gated attention, cross-attention AdaLN, perturbed-attention support for STG guidance, bigger connectors with per-modality projections, BWE vocoder
Norm / position
QK RMSNorm across heads, split RoPE (3D video with fps, 1D audio)

In AI Toolkit

model.arch
ltx2.3
UI label
LTX-2.3 (video)
model.name_or_path
Lightricks/LTX-2.3/ltx-2.3-22b-dev.safetensors
extra UI sections
sample.ctrl_img, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, datasets.do_i2v, train.audio_loss_multiplier, datasets.auto_frame_count

UI defaults

quantize / quantize_te
true / true
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
weighted
sample size
768×768, 121 frames, 24 fps
train.audio_loss_multiplier
1.0
datasets.do_audio
true
datasets.do_i2v
false
datasets.cache_latents_to_disk
true
datasets.fps
24
datasets.auto_frame_count
false
network.conv
disabled (linear LoRA only)

Specifics

Resolution
Buckets and sample sizes snap to multiples of 32.
Frame count
Sampling rounds num_frames down to 8n + 1.
Loading
name_or_path is a single .safetensors file. A repo/file path is looked for in the models folder first and downloaded there if missing. Transformer, connectors, both VAEs and the vocoder are converted from it to Diffusers modules; fp8 files are dequantized first.
Text encoder source
Lightricks/gemma-3-12b-it-qat-q4_0-unquantized, unless te_name_or_path is set (a Gemma 3 folder, repo, or ComfyUI Gemma 3 .safetensors). The vision tower is dropped. Never trained.
Audio
With do_audio on, clip audio is made stereo, resampled to 16 kHz, turned into log-mel and encoded by the audio VAE. The audio stream is noised at the same timestep as video and adds its own loss, scaled by audio_loss_multiplier. Image batches and datasets without audio carry an empty audio stream and no audio loss.
Image-to-video
With do_i2v on, the first frame is encoded and placed in latent frame 0 at t = 0 and masked out of the loss (the mask is renormalized).
Cross timestep
Training and sampling pass use_cross_timestep, the 2.3 behaviour, so each stream’s cross-attention is modulated by the other stream’s sigma.
LoRA target
LTX2VideoTransformer3DModel
Saving
LoRAs are converted to the original LTX-2.3 key layout with the diffusion_model prefix, which ComfyUI can load. Full fine-tunes save the transformer as a Diffusers folder.
Sampling
Samples are mp4 with audio. Adds STG (scale 1, block 28), modality guidance 3, guidance rescale 0.7 and audio guidance 7 on top of guidance_scale. low_vram turns on frame-wise VAE encode and decode.
Metadata base version
ltx2

Example config

not verified

job: extensionconfig:  name: "my_ltx23_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 32        linear_alpha: 32      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 121          fps: 24          do_audio: true          do_i2v: false          cache_latents_to_disk: true          resolution: [512, 768]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "weighted"        audio_loss_multiplier: 1.0        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Lightricks/LTX-2.3/ltx-2.3-22b-dev.safetensors"        arch: "ltx2.3"        quantize: true        quantize_te: true        low_vram: true      sample:        sampler: "flowmatch"        sample_every: 250        width: 768        height: 768        num_frames: 121        fps: 24        guidance_scale: 3        sample_steps: 25        prompts:          - "a bear building a log cabin in the snow covered mountains, the sound of an axe chopping wood"

On this page