MiniMax-H3-Ref2VA — Video Generation AI Model
MiniMax-H3 Ref2VA (omni-reference) generates a 5-15s 24fps video with native stereo audio from a prompt plus a reference image and a reference clip, that clip's own soundtrack included. Video and its soundtrack come out of one denoising loop, so lip movement and sound land in sync. ~144GB bf16 weights — offloaded component-by-component, minutes-scale per clip.
Details
- CategoryVideo Generation
- Year2026
- LicenseMiniMax-H3-Community-License
Compliance & Provenance
- ProviderMiniMax (open weights) · Specialized
- EU AI Act RiskLimited Risk
- Art. 50 TransparencyRequired — AI-generated outputs are marked
Inputs & Outputs
- TextInput · string
Describe the scene, the motion and the sound you want, and refer to the references by their tags — '<Picture 1>' for the image, '<Video 1>' for the clip, '<Audio 1>' for that clip's own soundtrack.
- ImageInput · image · optional
Optional reference image (subject, style or scene). It does not bind the generated geometry — the canvas is the one the resolution parameter selects.
- VideoInput · video · optional
Optional reference clip for motion and camera, conditioned on together with its own soundtrack, which has to be mono or stereo. Anything longer than the clip being generated is truncated to it, so match its length to num_frames — about 5.2s at the default 124 frames. 2s is the shortest clip MiniMax documents and 15s the longest.
- VideoOutput · video
Generated video with its synchronized stereo soundtrack muxed in
Tags
Alternatives in Video Generation
- Cosmos3 Nano
NVIDIA world model for text- and image-to-video, with optional synced audio.
- Helios Base
On-device text-to-video, up to 60s @ 24fps. 3-10 min.
- Helios Distilled
Fast distilled text-to-video, fixed 640x384. 1-5 min.
- Wan2.2 TI2V 5B
Wan2.2 dense 5B unified text+image-to-video. Text only → 5s video from prompt; text + image → image as the first frame of the video.