MiniMax-H3-Ref2VA — Video Generation AI Model

MiniMax-H3 Ref2VA (omni-reference) generates a 5-15s 24fps video with native stereo audio from a prompt plus a reference image and a reference clip, that clip's own soundtrack included. Video and its soundtrack come out of one denoising loop, so lip movement and sound land in sync. ~144GB bf16 weights — offloaded component-by-component, minutes-scale per clip.

Details

  • CategoryVideo Generation
  • Year2026
  • LicenseMiniMax-H3-Community-License

Compliance & Provenance

  • ProviderMiniMax (open weights) · Specialized
  • EU AI Act RiskLimited Risk
  • Art. 50 TransparencyRequired — AI-generated outputs are marked

Inputs & Outputs

  • TextInput · string

    Describe the scene, the motion and the sound you want, and refer to the references by their tags — '<Picture 1>' for the image, '<Video 1>' for the clip, '<Audio 1>' for that clip's own soundtrack.

  • ImageInput · image · optional

    Optional reference image (subject, style or scene). It does not bind the generated geometry — the canvas is the one the resolution parameter selects.

  • VideoInput · video · optional

    Optional reference clip for motion and camera, conditioned on together with its own soundtrack, which has to be mono or stereo. Anything longer than the clip being generated is truncated to it, so match its length to num_frames — about 5.2s at the default 124 frames. 2s is the shortest clip MiniMax documents and 15s the longest.

  • VideoOutput · video

    Generated video with its synchronized stereo soundtrack muxed in

Tags

  • video-generation
  • reference-to-video
  • audio-video-generation
  • synchronized-audio-video
  • lip-sync
  • omni-reference
  • diffusion
  • minimax
  • minimax-h3

Alternatives in Video Generation

  • Cosmos3 Nano

    NVIDIA world model for text- and image-to-video, with optional synced audio.

  • Helios Base

    On-device text-to-video, up to 60s @ 24fps. 3-10 min.

  • Helios Distilled

    Fast distilled text-to-video, fixed 640x384. 1-5 min.

  • Wan2.2 TI2V 5B

    Wan2.2 dense 5B unified text+image-to-video. Text only → 5s video from prompt; text + image → image as the first frame of the video.

Resources