MiniMax-H3 (FL2VA)
'MiniMax’s open 33B single-stream DiT that generates video with native stereo audio in one packed sequence of text, video and audio tokens. This is the first/last-frame (FL2VA) partition, guidance-distilled. ai-toolkit trains the Comfy-Org pruned repack, which drops most of the 13B AdaLN parameters.'
- weights
- Comfy-Org/MiniMax-H3 ↗
- org
- MiniMax
- modality
- video
- tasks
- text-to-video · image-to-video · text-to-audio-video · image-to-audio-video
- license
- MiniMax-H3 Community License
- released
- 2026-07-28not verified
- native output
- Short side 768, 24 fps, 4–15 s, 32 kHz stereo audio
- total params
- 48.62B
- model.arch
- minimax_h3
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | MiniMax-H3 FL2VA, pruned MiniMaxH3Transformer (ai-toolkit) Pruned: the timestep MLP is replaced by a small lookup table (8-d time embedding), which shrinks the per-block AdaLN projections. The unpruned file is 33.12B params. ai-toolkit loads minimax_h3_fl2va_pruned_int8_convrot.safetensors by default (20.97 GB), pre-quantized int8 ConvRot, as is. | 20.11B 20,111,438,744 | 40.23 GB | bf16+fp16+fp32 | yes |
| Text encoder | Qwen3-VL-32B (first 50 layers) transformers.Qwen3VLForConditionalGeneration The Comfy repack keeps only the 50 decoder layers the model reads, plus the vision tower; no final norm, no LM head. ai-toolkit loads qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors by default (15.69 GB): nvfp4 AWQ language linears, int8 embeddings, bf16 vision. | 25.75B 25,753,095,920 | 51.51 GB | bf16 | no |
| Tokenizer / processor | Qwen3-VL tokenizer with MiniMax-H3 special tokens transformers.AutoTokenizer / AutoProcessor From the original repo (FL2VA/tokenizer and FL2VA/processor), along with the text encoder config. | — | — | — | — |
| Video VAE | H3-VisualVAE (f16t4d24) MiniMaxH3VideoVAE (ai-toolkit) Causal 3D-CNN encoder and a 36-layer ViT decoder. | 2.60B 2,603,871,080 | 5.21 GB | fp16 | no |
| Audio VAE | H3-AudioVAE (DAC / BigVGAN) MiniMaxH3AudioVAE (ai-toolkit) Waveform autoencoder, no mel front end and no separate vocoder. Runs in fp32. | 151.3M 151,287,320 | 605 MB | fp32 | no |
| Training adapter | MiniMax-H3 training adapter v1 (LoRA rank 16)adapter Made with ai-toolkit to hold off the breakdown of guidance distillation. Loaded as a live, frozen LoRA while training and switched off for samples. Never merged into the quantized weights. | 77.5M 77,529,088 | 155 MB | bf16 | no |
| total | 48.62B | 97.55 GB | |||
Latent space
- autoencoder
- H3-VisualVAE
- pixels per token
- 32×32 × 4 frames
- frame count
- 17n + 5
- notes
- The encoder works in 17-frame chunks of 5 latent frames and drops 3 trailing latents, so 17n + 5 pixel frames map to 5n + 2 latent frames. A single frame maps to one latent frame. Each 1×2×2 patch is a 96-d token.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 768×768 × 107f | 24×32×48×48 | 18,432 | ai-toolkit sample default |
| 1344×768 × 124f | 24×37×48×84 | 37,296 | native 16:9 canvas, about 5 s |
| 1024×1024 | 24×1×64×64 | 1,024 | still image |
Audio latent space
- autoencoder
- H3-AudioVAE
- audio channels
- mono
- patch
- 1 latent steps per token
- streams
- 2 latent sequences per clip
- tokens per second
- 80.00
- notes
- The VAE is mono. Stereo runs through it as two separate items, and each channel’s latents become their own rows in the packed sequence (channel-major), so a stereo clip is two latent sequences.
Architecture
- Blocks
- 50, plus a 2-block text token refiner
- Hidden size
- 5376; attention 56 heads × 128 (7168)
- FFN
- SwiGLU, 14336
- Sequence
- One packed sequence of text, condition video, audio and target video rows. Full self-attention, no cross-attention, no per-modality weights in attention or FFN
- Modalities
- Separate input projections and output heads for video and audio; AdaLN modulation per modality tag (video, text, audio)
- Text conditioning
- Unnormalized hidden state 50 of Qwen3-VL-32B (5120-d), 512 caption tokens max in ai-toolkit
- Image conditioning
- Keyframes enter twice: as <Picture i> vision blocks in the Qwen3-VL prompt and as noise-augmented latent rows ahead of the target
- Timesteps
- Per row; audio rows on their own schedule
- Objective
- Rectified flow predicting clean − noise, shift 12 for video and 3 for audio
- Norm / position
- QK RMSNorm per head, 3-axis MM-RoPE on 96 of 128 head channels
- Guidance
- Guidance-distilled: no CFG, one forward per step
In AI Toolkit
- model.arch
- minimax_h3
- UI label
- MiniMax-H3 (video)
- model.name_or_path
- Comfy-Org/MiniMax-H3
- source
- extensions_built_in/diffusion_models/minimax_h3/minimax_h3.py
- extensions_built_in/diffusion_models/minimax_h3/src/transformer.py
- extensions_built_in/diffusion_models/minimax_h3/src/packing.py
- extensions_built_in/diffusion_models/minimax_h3/src/pipeline.py
- extensions_built_in/diffusion_models/minimax_h3/src/text_encoder.py
- extensions_built_in/diffusion_models/minimax_h3/src/vae.py
- extensions_built_in/diffusion_models/minimax_h3/src/audio_vae.py
- toolkit/models/v2/resolver.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- sample.ctrl_img, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, datasets.do_i2v, train.audio_loss_multiplier, datasets.auto_frame_count, model.assistant_lora_path
UI defaults
- quantize / qtype
- true / convrot8
- quantize_te / qtype_te
- true / nvfp4
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- shift
- cache_text_embeddings
- true
- do_guidance_loss / guidance_loss_target
- true / 3.5
- assistant_lora_path
- ostris/minimax_h3_training_adapter/minimax_h3_training_adapter_v1.safetensors
- network.linear / linear_alpha
- 16 / 16
- network_kwargs.ignore_if_contains
- ["adaln_proj"]
- sample
- 768×768, 107 frames, 24 fps, guidance 1, 28 steps
- train.audio_loss_multiplier
- 1.0
- datasets.do_audio / do_i2v
- true / false
- datasets.num_frames / fps
- 39 / 24
- datasets.auto_frame_count
- true
- datasets.cache_latents_to_disk
- true
- network.conv
- disabled (linear LoRA only)
Specifics
- Loading
- Files resolve under the models folder at their ComfyUI paths (diffusion_models/, text_encoders/, vae/) and download from Comfy-Org/MiniMax-H3 only when missing, about 43 GB. Override single files with model_kwargs dit_<partition>_path, text_encoder_path, video_vae_path and audio_vae_path.
- Partition
- model_kwargs.partition picks the transformer: fl2va_pruned (default), fl2va, ref2va or ref2va_pruned.
- Quantization
- convrot8 and nvfp4 match the shipped files, so nothing is re-quantized. Another qtype re-quantizes layer by layer. Patch projections, time embedder, final layer, condition projection, token refiner and AdaLN projections stay unquantized.
- Text encoder
- Truncated to 50 layers with the final norm removed; hidden state 50 is the conditioning. Captions are capped at 512 tokens (model_kwargs.max_text_length, 0 for no cap). Never trained.
- Distillation handling
- The model is guidance-distilled. The UI picks Contrastive Guidance (do_guidance_loss, target 3.5), the training adapter (assistant_lora_path), both (default) or none. The adapter is faster but still breaks down over a long run.
- Resolution
- Buckets and sample sizes snap to multiples of 32 (16× VAE × 2×2 patch).
- Frame count
- Clips snap down to 17n + 5 (5, 22, 39, 56, …, 107, 124). Video is fixed at 24 fps. num_frames 1 trains and samples single images.
- Timesteps
- The model uses t = 1 − sigma and predicts clean − noise; ai-toolkit flips the timestep and negates the prediction. The audio sigma is remapped from the video sigma (shift 12 to shift 3) every step.
- Audio
- With do_audio on, clip audio is made stereo, resampled to 32 kHz and encoded per channel. Its loss is scaled by audio_loss_multiplier. Without audio, noised silence rides along with no loss.
- Image-to-video
- With do_i2v on, the first frame is encoded with the released recipe (seed 42, fp16 rounding), noise-augmented to t = 0.999 and added as condition rows ahead of the target. Their predictions are dropped, so the target frames keep their full loss.
- LoRA target
- MiniMaxH3Transformer
- Saving
- LoRA keys use the diffusion_model prefix of the original checkpoint, which ComfyUI loads. Full fine-tunes dequantize and save transformer/model.safetensors, keeping the fp32 islands.
- Sampling
- The released sampler: flowmatch Euler, one forward per step, guidance_scale and negative prompts ignored. A ctrl_img is used as the first frame. Samples are mp4 with audio (model_kwargs.sample_audio).
- Metadata base version
- minimax_h3
Example config
not verified
job: extensionconfig: name: "my_minimax_h3_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 network_kwargs: ignore_if_contains: ["adaln_proj"] save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 39 fps: 24 auto_frame_count: true do_audio: true do_i2v: false cache_latents_to_disk: true resolution: [512, 768] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "shift" do_guidance_loss: true guidance_loss_target: 3.5 audio_loss_multiplier: 1.0 optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Comfy-Org/MiniMax-H3" arch: "minimax_h3" quantize: true qtype: "convrot8" quantize_te: true qtype_te: "nvfp4" low_vram: true assistant_lora_path: "ostris/minimax_h3_training_adapter/minimax_h3_training_adapter_v1.safetensors" sample: sampler: "flowmatch" sample_every: 250 width: 768 height: 768 num_frames: 107 fps: 24 guidance_scale: 1 sample_steps: 28 prompts: - "a bear building a log cabin in the snow covered mountains, the sound of an axe chopping wood"Links
Wan 2.2 TI2V 5B
The dense 5B member of Wan 2.2. One transformer does both text-to-video and image-to-video, working in the 16× spatial / 4× temporal latent space of the new Wan 2.2 VAE. Small enough to train on a single consumer GPU.
MiniMax-H3 Ref2VA
'The omni-reference partition of MiniMax-H3: the same 33B single-stream audio-video DiT, trained to take images and video clips as subject and style references instead of first frames. ai-toolkit trains the Comfy-Org pruned repack and shares the text encoder and VAEs with the FL2VA arch.'