Docs
AI ToolkitModels

MiniMax-H3 (FL2VA)

'MiniMax’s open 33B single-stream DiT that generates video with native stereo audio in one packed sequence of text, video and audio tokens. This is the first/last-frame (FL2VA) partition, guidance-distilled. ai-toolkit trains the Comfy-Org pruned repack, which drops most of the 13B AdaLN parameters.'

org
MiniMax
modality
video
tasks
text-to-video · image-to-video · text-to-audio-video · image-to-audio-video
license
MiniMax-H3 Community License
released
2026-07-28not verified
native output
Short side 768, 24 fps, 4–15 s, 32 kHz stereo audio
total params
48.62B
model.arch
minimax_h3

Components

rolemodelparamssizedtypetrained
Transformer
MiniMax-H3 FL2VA, pruned
MiniMaxH3Transformer (ai-toolkit)

Pruned: the timestep MLP is replaced by a small lookup table (8-d time embedding), which shrinks the per-block AdaLN projections. The unpruned file is 33.12B params. ai-toolkit loads minimax_h3_fl2va_pruned_int8_convrot.safetensors by default (20.97 GB), pre-quantized int8 ConvRot, as is.

20.11B
20,111,438,744
40.23 GBbf16+fp16+fp32yes
Text encoder
Qwen3-VL-32B (first 50 layers)
transformers.Qwen3VLForConditionalGeneration

The Comfy repack keeps only the 50 decoder layers the model reads, plus the vision tower; no final norm, no LM head. ai-toolkit loads qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors by default (15.69 GB): nvfp4 AWQ language linears, int8 embeddings, bf16 vision.

25.75B
25,753,095,920
51.51 GBbf16no
Tokenizer / processor
Qwen3-VL tokenizer with MiniMax-H3 special tokens
transformers.AutoTokenizer / AutoProcessor

From the original repo (FL2VA/tokenizer and FL2VA/processor), along with the text encoder config.

————
Video VAE
H3-VisualVAE (f16t4d24)
MiniMaxH3VideoVAE (ai-toolkit)

Causal 3D-CNN encoder and a 36-layer ViT decoder.

2.60B
2,603,871,080
5.21 GBfp16no
Audio VAE
H3-AudioVAE (DAC / BigVGAN)
MiniMaxH3AudioVAE (ai-toolkit)

Waveform autoencoder, no mel front end and no separate vocoder. Runs in fp32.

151.3M
151,287,320
605 MBfp32no
Training adapter
MiniMax-H3 training adapter v1 (LoRA rank 16)adapter

Made with ai-toolkit to hold off the breakdown of guidance distillation. Loaded as a live, frozen LoRA while training and switched off for samples. Never merged into the quantized weights.

77.5M
77,529,088
155 MBbf16no
total48.62B97.55 GB

Latent space

spatial
16×
temporal
4×
channels
24
patch
1×2×2
autoencoder
H3-VisualVAE
pixels per token
32×32 × 4 frames
frame count
17n + 5
notes
The encoder works in 17-frame chunks of 5 latent frames and drops 3 trailing latents, so 17n + 5 pixel frames map to 5n + 2 latent frames. A single frame maps to one latent frame. Each 1×2×2 patch is a 96-d token.
inputlatent (c×t×h×w)tokens
768×768 × 107f24×32×48×4818,432ai-toolkit sample default
1344×768 × 124f24×37×48×8437,296native 16:9 canvas, about 5 s
1024×102424×1×64×641,024still image

Audio latent space

sample rate
32 kHz
temporal
800×
channels
32
latent rate
40/s
autoencoder
H3-AudioVAE
audio channels
mono
patch
1 latent steps per token
streams
2 latent sequences per clip
tokens per second
80.00
notes
The VAE is mono. Stereo runs through it as two separate items, and each channel’s latents become their own rows in the packed sequence (channel-major), so a stereo clip is two latent sequences.

Architecture

Blocks
50, plus a 2-block text token refiner
Hidden size
5376; attention 56 heads × 128 (7168)
FFN
SwiGLU, 14336
Sequence
One packed sequence of text, condition video, audio and target video rows. Full self-attention, no cross-attention, no per-modality weights in attention or FFN
Modalities
Separate input projections and output heads for video and audio; AdaLN modulation per modality tag (video, text, audio)
Text conditioning
Unnormalized hidden state 50 of Qwen3-VL-32B (5120-d), 512 caption tokens max in ai-toolkit
Image conditioning
Keyframes enter twice: as <Picture i> vision blocks in the Qwen3-VL prompt and as noise-augmented latent rows ahead of the target
Timesteps
Per row; audio rows on their own schedule
Objective
Rectified flow predicting clean − noise, shift 12 for video and 3 for audio
Norm / position
QK RMSNorm per head, 3-axis MM-RoPE on 96 of 128 head channels
Guidance
Guidance-distilled: no CFG, one forward per step

In AI Toolkit

model.arch
minimax_h3
UI label
MiniMax-H3 (video)
model.name_or_path
Comfy-Org/MiniMax-H3
extra UI sections
sample.ctrl_img, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, datasets.do_i2v, train.audio_loss_multiplier, datasets.auto_frame_count, model.assistant_lora_path

UI defaults

quantize / qtype
true / convrot8
quantize_te / qtype_te
true / nvfp4
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
shift
cache_text_embeddings
true
do_guidance_loss / guidance_loss_target
true / 3.5
assistant_lora_path
ostris/minimax_h3_training_adapter/minimax_h3_training_adapter_v1.safetensors
network.linear / linear_alpha
16 / 16
network_kwargs.ignore_if_contains
["adaln_proj"]
sample
768×768, 107 frames, 24 fps, guidance 1, 28 steps
train.audio_loss_multiplier
1.0
datasets.do_audio / do_i2v
true / false
datasets.num_frames / fps
39 / 24
datasets.auto_frame_count
true
datasets.cache_latents_to_disk
true
network.conv
disabled (linear LoRA only)

Specifics

Loading
Files resolve under the models folder at their ComfyUI paths (diffusion_models/, text_encoders/, vae/) and download from Comfy-Org/MiniMax-H3 only when missing, about 43 GB. Override single files with model_kwargs dit_<partition>_path, text_encoder_path, video_vae_path and audio_vae_path.
Partition
model_kwargs.partition picks the transformer: fl2va_pruned (default), fl2va, ref2va or ref2va_pruned.
Quantization
convrot8 and nvfp4 match the shipped files, so nothing is re-quantized. Another qtype re-quantizes layer by layer. Patch projections, time embedder, final layer, condition projection, token refiner and AdaLN projections stay unquantized.
Text encoder
Truncated to 50 layers with the final norm removed; hidden state 50 is the conditioning. Captions are capped at 512 tokens (model_kwargs.max_text_length, 0 for no cap). Never trained.
Distillation handling
The model is guidance-distilled. The UI picks Contrastive Guidance (do_guidance_loss, target 3.5), the training adapter (assistant_lora_path), both (default) or none. The adapter is faster but still breaks down over a long run.
Resolution
Buckets and sample sizes snap to multiples of 32 (16× VAE × 2×2 patch).
Frame count
Clips snap down to 17n + 5 (5, 22, 39, 56, …, 107, 124). Video is fixed at 24 fps. num_frames 1 trains and samples single images.
Timesteps
The model uses t = 1 − sigma and predicts clean − noise; ai-toolkit flips the timestep and negates the prediction. The audio sigma is remapped from the video sigma (shift 12 to shift 3) every step.
Audio
With do_audio on, clip audio is made stereo, resampled to 32 kHz and encoded per channel. Its loss is scaled by audio_loss_multiplier. Without audio, noised silence rides along with no loss.
Image-to-video
With do_i2v on, the first frame is encoded with the released recipe (seed 42, fp16 rounding), noise-augmented to t = 0.999 and added as condition rows ahead of the target. Their predictions are dropped, so the target frames keep their full loss.
LoRA target
MiniMaxH3Transformer
Saving
LoRA keys use the diffusion_model prefix of the original checkpoint, which ComfyUI loads. Full fine-tunes dequantize and save transformer/model.safetensors, keeping the fp32 islands.
Sampling
The released sampler: flowmatch Euler, one forward per step, guidance_scale and negative prompts ignored. A ctrl_img is used as the first frame. Samples are mp4 with audio (model_kwargs.sample_audio).
Metadata base version
minimax_h3

Example config

not verified

job: extensionconfig:  name: "my_minimax_h3_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16        network_kwargs:          ignore_if_contains: ["adaln_proj"]      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 39          fps: 24          auto_frame_count: true          do_audio: true          do_i2v: false          cache_latents_to_disk: true          resolution: [512, 768]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "shift"        do_guidance_loss: true        guidance_loss_target: 3.5        audio_loss_multiplier: 1.0        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Comfy-Org/MiniMax-H3"        arch: "minimax_h3"        quantize: true        qtype: "convrot8"        quantize_te: true        qtype_te: "nvfp4"        low_vram: true        assistant_lora_path: "ostris/minimax_h3_training_adapter/minimax_h3_training_adapter_v1.safetensors"      sample:        sampler: "flowmatch"        sample_every: 250        width: 768        height: 768        num_frames: 107        fps: 24        guidance_scale: 1        sample_steps: 28        prompts:          - "a bear building a log cabin in the snow covered mountains, the sound of an axe chopping wood"

On this page