Docs
AI ToolkitModels

MiniMax-H3 Ref2VA

'The omni-reference partition of MiniMax-H3: the same 33B single-stream audio-video DiT, trained to take images and video clips as subject and style references instead of first frames. ai-toolkit trains the Comfy-Org pruned repack and shares the text encoder and VAEs with the FL2VA arch.'

org
MiniMax
modality
video
tasks
reference-to-video · reference-to-audio-video · text-to-video · text-to-audio-video
license
MiniMax-H3 Community License
released
2026-07-28not verified
native output
Short side 768, 24 fps, 4–15 s, 32 kHz stereo audio
total params
48.62B
model.arch
minimax_h3_ref2va

Components

rolemodelparamssizedtypetrained
Transformer
MiniMax-H3 Ref2VA, pruned
MiniMaxH3Transformer (ai-toolkit)

Same shape as the FL2VA transformer, different weights. Pruned: the timestep MLP is replaced by a small lookup table (8-d time embedding), which shrinks the per-block AdaLN projections. The unpruned file is 33.12B params. ai-toolkit loads minimax_h3_ref2va_pruned_int8_convrot.safetensors by default (20.97 GB), pre-quantized int8 ConvRot, as is.

20.11B
20,111,438,744
40.23 GBbf16+fp16+fp32yes
Text encoder
Qwen3-VL-32B (first 50 layers)
transformers.Qwen3VLForConditionalGeneration

Shared with the FL2VA arch. The Comfy repack keeps only the 50 decoder layers the model reads, plus the vision tower, which sees every reference. ai-toolkit loads qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors by default (15.69 GB): nvfp4 AWQ language linears, int8 embeddings, bf16 vision.

25.75B
25,753,095,920
51.51 GBbf16no
Tokenizer / processor
Qwen3-VL tokenizer with MiniMax-H3 special tokens
transformers.AutoTokenizer / AutoProcessor

Loaded from the original repo’s FL2VA folders for both arches, along with the text encoder config.

————
Video VAE
H3-VisualVAE (f16t4d24)
MiniMaxH3VideoVAE (ai-toolkit)

Causal 3D-CNN encoder and a 36-layer ViT decoder. Shared with the FL2VA arch.

2.60B
2,603,871,080
5.21 GBfp16no
Audio VAE
H3-AudioVAE (DAC / BigVGAN)
MiniMaxH3AudioVAE (ai-toolkit)

Waveform autoencoder, no mel front end and no separate vocoder. Runs in fp32.

151.3M
151,287,320
605 MBfp32no
Training adapter
MiniMax-H3 Ref2VA training adapter v1 (LoRA rank 16)adapter

Made with ai-toolkit for the ref2va weights, to hold off the breakdown of guidance distillation. Loaded as a live, frozen LoRA while training and switched off for samples. Never merged into the quantized weights.

77.5M
77,529,088
155 MBbf16no
total48.62B97.55 GB

Latent space

spatial
16×
temporal
4×
channels
24
patch
1×2×2
autoencoder
H3-VisualVAE
pixels per token
32×32 × 4 frames
frame count
17n + 5
notes
The encoder works in 17-frame chunks of 5 latent frames and drops 3 trailing latents, so 17n + 5 pixel frames map to 5n + 2 latent frames. A single frame maps to one latent frame. Each 1×2×2 patch is a 96-d token. References add their own latent rows on top of the target’s.
inputlatent (c×t×h×w)tokens
768×768 × 107f24×32×48×4818,432ai-toolkit sample default
1344×768 × 124f24×37×48×8437,296native 16:9 canvas, about 5 s
1024×102424×1×64×641,024still image

Audio latent space

sample rate
32 kHz
temporal
800×
channels
32
latent rate
40/s
autoencoder
H3-AudioVAE
audio channels
mono
patch
1 latent steps per token
streams
2 latent sequences per clip
tokens per second
80.00
notes
The VAE is mono. Stereo runs through it as two separate items, and each channel’s latents become their own rows in the packed sequence (channel-major), so a stereo clip is two latent sequences. A reference video’s soundtrack can ride along as clean audio rows.

Architecture

Blocks
50, plus a 2-block text token refiner
Hidden size
5376; attention 56 heads × 128 (7168)
FFN
SwiGLU, 14336
Sequence
One packed sequence of text, reference blocks, audio and target video rows. Full self-attention, no cross-attention, no per-modality weights in attention or FFN
Text conditioning
Unnormalized hidden state 50 of Qwen3-VL-32B (5120-d), 512 caption tokens max in ai-toolkit
Reference conditioning
Each reference enters twice: as a <Picture i> or timestamped <Video k> vision block in the Qwen3-VL prompt, and as a noise-augmented latent block on its own rotary grid
Timesteps
Per row; audio rows on their own schedule, reference rows pinned near clean
Objective
Rectified flow predicting clean − noise, shift 12 for video and 3 for audio
Norm / position
QK RMSNorm per head, 3-axis MM-RoPE on 96 of 128 head channels
Guidance
Guidance-distilled: no CFG, one forward per step

In AI Toolkit

model.arch
minimax_h3_ref2va
UI label
MiniMax-H3 Ref2V (video)
model.name_or_path
Comfy-Org/MiniMax-H3
extra UI sections
sample.multi_ctrl_imgs, datasets.multi_control_paths, datasets.num_frames, model.layer_offloading, model.low_vram, datasets.do_audio, datasets.audio_normalize, datasets.audio_preserve_pitch, train.audio_loss_multiplier, datasets.auto_frame_count, model.assistant_lora_path

UI defaults

quantize / qtype
true / convrot8
quantize_te / qtype_te
true / nvfp4
low_vram
true
noise_scheduler / sampler
flowmatch / flowmatch
timestep_type
shift
cache_text_embeddings
true
do_guidance_loss / guidance_loss_target
true / 3.5
assistant_lora_path
ostris/minimax_h3_training_adapter/minimax_h3_ref2va_training_adapter_v1.safetensors
network.linear / linear_alpha
16 / 16
network_kwargs.ignore_if_contains
["adaln_proj"]
sample
768×768, 107 frames, 24 fps, guidance 1, 28 steps
train.audio_loss_multiplier
1.0
datasets.do_audio
true
datasets.num_frames / fps
39 / 24
datasets.auto_frame_count
true
datasets.cache_latents_to_disk
true
network.conv
disabled (linear LoRA only)

Specifics

Loading
Same as MiniMax-H3: files resolve under the models folder at their ComfyUI paths and download from Comfy-Org/MiniMax-H3 only when missing. model_kwargs.partition is ref2va_pruned (default) or ref2va.
Quantization
convrot8 and nvfp4 match the shipped files, so nothing is re-quantized. Another qtype re-quantizes layer by layer. Patch projections, time embedder, final layer, condition projection, token refiner and AdaLN projections stay unquantized.
References
Training references come from the dataset control paths (several per item); sampling uses the sample ctrl images, always as references, never as first frames. Every item in a batch needs the same number of references with matching aspect ratios.
Reference sizing
Each reference keeps its own aspect and is matched to the target’s pixel area on a /32 grid. Images only scale down. A same-aspect video reference is exactly the target size.
Video references
Control videos get the dataset’s frame count and fps, snap to 17n + 5, and are VAE-encoded once and cached next to the video. Their soundtrack rides as clean reference audio rows when every item has one.
Image Reference Presentation
model_kwargs.image_refs_as_video holds a still image for image_ref_video_frames (default 5) frames and sends it through the video-reference path, for LoRAs trained on image refs but used with video refs. Changes the text-embedding cache key.
Distillation handling
Contrastive Guidance, the ref2va training adapter, both (default), none, or D-OPSD (model_kwargs.dopsd): a no-grad teacher pass that sees the target as its own reference produces the target for a reference-free pass, baking the reference into the trigger word. D-OPSD needs cached pixel tensors.
Text encoder
Truncated to 50 layers with the final norm removed; hidden state 50 is the conditioning. Captions are capped at 512 tokens; vision blocks are never trimmed. Never trained.
Frame count
Clips snap down to 17n + 5 (5, 22, 39, 56, …, 107, 124). Video is fixed at 24 fps. num_frames 1 trains and samples single images.
Audio
With do_audio on, clip audio is made stereo, resampled to 32 kHz and encoded per channel. Its loss is scaled by audio_loss_multiplier. Without audio, noised silence rides along with no loss.
LoRA target
MiniMaxH3Transformer
Saving
LoRA keys use the diffusion_model prefix of the original checkpoint, which ComfyUI loads. Full fine-tunes dequantize and save transformer/model.safetensors, keeping the fp32 islands.
Metadata base version
minimax_h3_ref2va

Example config

not verified

job: extensionconfig:  name: "my_minimax_h3_ref2va_lora_v1"  process:    - type: "sd_trainer"      training_folder: "output"      device: cuda:0      network:        type: "lora"        linear: 16        linear_alpha: 16        network_kwargs:          ignore_if_contains: ["adaln_proj"]      save:        dtype: float16        save_every: 250        max_step_saves_to_keep: 4      datasets:        - folder_path: "/path/to/videos"          control_path: "/path/to/references"          caption_ext: "txt"          caption_dropout_rate: 0.05          num_frames: 39          fps: 24          auto_frame_count: true          do_audio: true          cache_latents_to_disk: true          resolution: [512, 768]      train:        batch_size: 1        steps: 2000        gradient_checkpointing: true        noise_scheduler: "flowmatch"        timestep_type: "shift"        do_guidance_loss: true        guidance_loss_target: 3.5        audio_loss_multiplier: 1.0        optimizer: "adamw8bit"        lr: 1e-4        dtype: bf16        cache_text_embeddings: true      model:        name_or_path: "Comfy-Org/MiniMax-H3"        arch: "minimax_h3_ref2va"        quantize: true        qtype: "convrot8"        quantize_te: true        qtype_te: "nvfp4"        low_vram: true        assistant_lora_path: "ostris/minimax_h3_training_adapter/minimax_h3_ref2va_training_adapter_v1.safetensors"      sample:        sampler: "flowmatch"        sample_every: 250        width: 768        height: 768        num_frames: 107        fps: 24        guidance_scale: 1        sample_steps: 28        samples:          - prompt: "the character from the reference image walks through a snowy forest, footsteps crunching"            ctrl_img_1: "/path/to/reference.png"

On this page