MiniMax-H3 Ref2VA
'The omni-reference partition of MiniMax-H3: the same 33B single-stream audio-video DiT, trained to take images and video clips as subject and style references instead of first frames. ai-toolkit trains the Comfy-Org pruned repack and shares the text encoder and VAEs with the FL2VA arch.'
- weights
- Comfy-Org/MiniMax-H3 ↗
- org
- MiniMax
- modality
- video
- tasks
- reference-to-video · reference-to-audio-video · text-to-video · text-to-audio-video
- license
- MiniMax-H3 Community License
- released
- 2026-07-28not verified
- native output
- Short side 768, 24 fps, 4–15 s, 32 kHz stereo audio
- total params
- 48.62B
- model.arch
- minimax_h3_ref2va
Components
| role | model | params | size | dtype | trained |
|---|---|---|---|---|---|
| Transformer | MiniMax-H3 Ref2VA, pruned MiniMaxH3Transformer (ai-toolkit) Same shape as the FL2VA transformer, different weights. Pruned: the timestep MLP is replaced by a small lookup table (8-d time embedding), which shrinks the per-block AdaLN projections. The unpruned file is 33.12B params. ai-toolkit loads minimax_h3_ref2va_pruned_int8_convrot.safetensors by default (20.97 GB), pre-quantized int8 ConvRot, as is. | 20.11B 20,111,438,744 | 40.23 GB | bf16+fp16+fp32 | yes |
| Text encoder | Qwen3-VL-32B (first 50 layers) transformers.Qwen3VLForConditionalGeneration Shared with the FL2VA arch. The Comfy repack keeps only the 50 decoder layers the model reads, plus the vision tower, which sees every reference. ai-toolkit loads qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors by default (15.69 GB): nvfp4 AWQ language linears, int8 embeddings, bf16 vision. | 25.75B 25,753,095,920 | 51.51 GB | bf16 | no |
| Tokenizer / processor | Qwen3-VL tokenizer with MiniMax-H3 special tokens transformers.AutoTokenizer / AutoProcessor Loaded from the original repo’s FL2VA folders for both arches, along with the text encoder config. | — | — | — | — |
| Video VAE | H3-VisualVAE (f16t4d24) MiniMaxH3VideoVAE (ai-toolkit) Causal 3D-CNN encoder and a 36-layer ViT decoder. Shared with the FL2VA arch. | 2.60B 2,603,871,080 | 5.21 GB | fp16 | no |
| Audio VAE | H3-AudioVAE (DAC / BigVGAN) MiniMaxH3AudioVAE (ai-toolkit) Waveform autoencoder, no mel front end and no separate vocoder. Runs in fp32. | 151.3M 151,287,320 | 605 MB | fp32 | no |
| Training adapter | MiniMax-H3 Ref2VA training adapter v1 (LoRA rank 16)adapter Made with ai-toolkit for the ref2va weights, to hold off the breakdown of guidance distillation. Loaded as a live, frozen LoRA while training and switched off for samples. Never merged into the quantized weights. | 77.5M 77,529,088 | 155 MB | bf16 | no |
| total | 48.62B | 97.55 GB | |||
Latent space
- autoencoder
- H3-VisualVAE
- pixels per token
- 32×32 × 4 frames
- frame count
- 17n + 5
- notes
- The encoder works in 17-frame chunks of 5 latent frames and drops 3 trailing latents, so 17n + 5 pixel frames map to 5n + 2 latent frames. A single frame maps to one latent frame. Each 1×2×2 patch is a 96-d token. References add their own latent rows on top of the target’s.
| input | latent (c×t×h×w) | tokens | |
|---|---|---|---|
| 768×768 × 107f | 24×32×48×48 | 18,432 | ai-toolkit sample default |
| 1344×768 × 124f | 24×37×48×84 | 37,296 | native 16:9 canvas, about 5 s |
| 1024×1024 | 24×1×64×64 | 1,024 | still image |
Audio latent space
- autoencoder
- H3-AudioVAE
- audio channels
- mono
- patch
- 1 latent steps per token
- streams
- 2 latent sequences per clip
- tokens per second
- 80.00
- notes
- The VAE is mono. Stereo runs through it as two separate items, and each channel’s latents become their own rows in the packed sequence (channel-major), so a stereo clip is two latent sequences. A reference video’s soundtrack can ride along as clean audio rows.
Architecture
- Blocks
- 50, plus a 2-block text token refiner
- Hidden size
- 5376; attention 56 heads × 128 (7168)
- FFN
- SwiGLU, 14336
- Sequence
- One packed sequence of text, reference blocks, audio and target video rows. Full self-attention, no cross-attention, no per-modality weights in attention or FFN
- Text conditioning
- Unnormalized hidden state 50 of Qwen3-VL-32B (5120-d), 512 caption tokens max in ai-toolkit
- Reference conditioning
- Each reference enters twice: as a <Picture i> or timestamped <Video k> vision block in the Qwen3-VL prompt, and as a noise-augmented latent block on its own rotary grid
- Timesteps
- Per row; audio rows on their own schedule, reference rows pinned near clean
- Objective
- Rectified flow predicting clean − noise, shift 12 for video and 3 for audio
- Norm / position
- QK RMSNorm per head, 3-axis MM-RoPE on 96 of 128 head channels
- Guidance
- Guidance-distilled: no CFG, one forward per step
In AI Toolkit
- model.arch
- minimax_h3_ref2va
- UI label
- MiniMax-H3 Ref2V (video)
- model.name_or_path
- Comfy-Org/MiniMax-H3
- source
- extensions_built_in/diffusion_models/minimax_h3/minimax_h3.py
- extensions_built_in/diffusion_models/minimax_h3/src/transformer.py
- extensions_built_in/diffusion_models/minimax_h3/src/packing.py
- extensions_built_in/diffusion_models/minimax_h3/src/pipeline.py
- extensions_built_in/diffusion_models/minimax_h3/src/text_encoder.py
- extensions_built_in/diffusion_models/minimax_h3/src/ref_video_cache.py
- extensions_built_in/diffusion_models/minimax_h3/src/vae.py
- extensions_built_in/diffusion_models/minimax_h3/src/audio_vae.py
- toolkit/models/v2/resolver.py
- toolkit/models/v2/text_encoders/qwen3_vl.py
- extensions_built_in/diffusion_models/ui.tsx
- extra UI sections
- sample.multi_ctrl_imgs, datasets.multi_control_paths, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, train.audio_loss_multiplier, datasets.auto_frame_count, model.assistant_lora_path
UI defaults
- quantize / qtype
- true / convrot8
- quantize_te / qtype_te
- true / nvfp4
- low_vram
- true
- noise_scheduler / sampler
- flowmatch / flowmatch
- timestep_type
- shift
- cache_text_embeddings
- true
- do_guidance_loss / guidance_loss_target
- true / 3.5
- assistant_lora_path
- ostris/minimax_h3_training_adapter/minimax_h3_ref2va_training_adapter_v1.safetensors
- network.linear / linear_alpha
- 16 / 16
- network_kwargs.ignore_if_contains
- ["adaln_proj"]
- sample
- 768×768, 107 frames, 24 fps, guidance 1, 28 steps
- train.audio_loss_multiplier
- 1.0
- datasets.do_audio
- true
- datasets.num_frames / fps
- 39 / 24
- datasets.auto_frame_count
- true
- datasets.cache_latents_to_disk
- true
- network.conv
- disabled (linear LoRA only)
Specifics
- Loading
- Same as MiniMax-H3: files resolve under the models folder at their ComfyUI paths and download from Comfy-Org/MiniMax-H3 only when missing. model_kwargs.partition is ref2va_pruned (default) or ref2va.
- Quantization
- convrot8 and nvfp4 match the shipped files, so nothing is re-quantized. Another qtype re-quantizes layer by layer. Patch projections, time embedder, final layer, condition projection, token refiner and AdaLN projections stay unquantized.
- References
- Training references come from the dataset control paths (several per item); sampling uses the sample ctrl images, always as references, never as first frames. Every item in a batch needs the same number of references with matching aspect ratios.
- Reference sizing
- Each reference keeps its own aspect and is matched to the target’s pixel area on a /32 grid. Images only scale down. A same-aspect video reference is exactly the target size.
- Video references
- Control videos get the dataset’s frame count and fps, snap to 17n + 5, and are VAE-encoded once and cached next to the video. Their soundtrack rides as clean reference audio rows when every item has one.
- Image Reference Presentation
- model_kwargs.image_refs_as_video holds a still image for image_ref_video_frames (default 5) frames and sends it through the video-reference path, for LoRAs trained on image refs but used with video refs. Changes the text-embedding cache key.
- Distillation handling
- Contrastive Guidance, the ref2va training adapter, both (default), none, or D-OPSD (model_kwargs.dopsd): a no-grad teacher pass that sees the target as its own reference produces the target for a reference-free pass, baking the reference into the trigger word. D-OPSD needs cached pixel tensors.
- Text encoder
- Truncated to 50 layers with the final norm removed; hidden state 50 is the conditioning. Captions are capped at 512 tokens; vision blocks are never trimmed. Never trained.
- Frame count
- Clips snap down to 17n + 5 (5, 22, 39, 56, …, 107, 124). Video is fixed at 24 fps. num_frames 1 trains and samples single images.
- Audio
- With do_audio on, clip audio is made stereo, resampled to 32 kHz and encoded per channel. Its loss is scaled by audio_loss_multiplier. Without audio, noised silence rides along with no loss.
- LoRA target
- MiniMaxH3Transformer
- Saving
- LoRA keys use the diffusion_model prefix of the original checkpoint, which ComfyUI loads. Full fine-tunes dequantize and save transformer/model.safetensors, keeping the fp32 islands.
- Metadata base version
- minimax_h3_ref2va
Example config
not verified
job: extensionconfig: name: "my_minimax_h3_ref2va_lora_v1" process: - type: "sd_trainer" training_folder: "output" device: cuda:0 network: type: "lora" linear: 16 linear_alpha: 16 network_kwargs: ignore_if_contains: ["adaln_proj"] save: dtype: float16 save_every: 250 max_step_saves_to_keep: 4 datasets: - folder_path: "/path/to/videos" control_path: "/path/to/references" caption_ext: "txt" caption_dropout_rate: 0.05 num_frames: 39 fps: 24 auto_frame_count: true do_audio: true cache_latents_to_disk: true resolution: [512, 768] train: batch_size: 1 steps: 2000 gradient_checkpointing: true noise_scheduler: "flowmatch" timestep_type: "shift" do_guidance_loss: true guidance_loss_target: 3.5 audio_loss_multiplier: 1.0 optimizer: "adamw8bit" lr: 1e-4 dtype: bf16 cache_text_embeddings: true model: name_or_path: "Comfy-Org/MiniMax-H3" arch: "minimax_h3_ref2va" quantize: true qtype: "convrot8" quantize_te: true qtype_te: "nvfp4" low_vram: true assistant_lora_path: "ostris/minimax_h3_training_adapter/minimax_h3_ref2va_training_adapter_v1.safetensors" sample: sampler: "flowmatch" sample_every: 250 width: 768 height: 768 num_frames: 107 fps: 24 guidance_scale: 1 sample_steps: 28 samples: - prompt: "the character from the reference image walks through a snowy forest, footsteps crunching" ctrl_img_1: "/path/to/reference.png"Links
MiniMax-H3 (FL2VA)
'MiniMax’s open 33B single-stream DiT that generates video with native stereo audio in one packed sequence of text, video and audio tokens. This is the first/last-frame (FL2VA) partition, guidance-distilled. ai-toolkit trains the Comfy-Org pruned repack, which drops most of the 13B AdaLN parameters.'
LTX-2 19B
'Lightricks’ first joint audio-video DiT. One 19B model generates video and a synchronized stereo soundtrack together, with a 4096-wide video stream and a 2048-wide audio stream that cross-attend in every block. Video lives in a 32× spatial / 8× temporal latent space with 128 channels.'