For the complete documentation index, see llms.txt. This page is also available as Markdown.

Reference to Video

Reference to Video is MiniMax H3's most powerful task — and the one that makes it a true omni-modal model. Instead of animating an image directly (that's Image to Video), you hand H3 references — images, a video clip, an audio track, or any mix of them — and it builds a brand-new shot around them. Your character's face, your product, a dance from another video, a recorded voice line: H3 reads them all as one context and fuses them into a single coherent video.

Reference to Video vs. Image to Video

The distinction matters, because they solve different problems:

  • Image to Video uses your image as the literal first frame. The video starts from exactly that picture.

  • Reference to Video uses your images as ingredients. The character from your photo can appear mid-scene, from a new angle, in a new outfit, doing something entirely new — the composition is invented fresh from your prompt, which is also why this task gives you an aspect ratio control while Image to Video doesn't.

Rule of thumb: if you want this exact picture to move, use Image to Video. If you want this character / thing / style in a new shot, use Reference to Video.

The Three Reference Types

Image references (up to 5)

Lock the identity of anything: a character's face, an outfit, a product, a pet, a location, or an overall art style. H3 keeps them consistent throughout the clip. You can combine subjects — for example, two character photos plus one background photo — and direct them in the prompt: "the woman from <Picture 1> and the man from the <Picture 2> sit across from each other in the café from the third image."

Use <Picture N> in prompt to reference the image

Video reference (one clip)

Borrow motion from existing footage: choreography, a fight sequence, camera movement, a gesture. H3 extracts how things move and re-performs it with your subject and scene. The reference clip is trimmed to your target clip length, so pick the exact seconds of motion you want. Note that attaching a video reference raises the Standard-mode price from 6 to 7 credits per second.

Use <Video 1> in prompt to reference the video

Audio reference (one track, 2–15 seconds)

Drive the clip with sound: a recorded voice line, a sound effect, a snippet of music. H3 generates the video to the audio — a character speaking your recorded dialogue with matching lip movement and delivery is the classic use.

On H3, an audio reference is a complete job by itself: you can generate with only a voice recording and a prompt, no images required. Describe who is speaking and where, and H3 invents the performer around your audio.

Use <Audio 1> in prompt to reference the video

Mixing them

The real magic is combination — one generation can use all three at once. Character photos define who, a video reference defines how they move, an audio track defines what they say or dance to, and the prompt ties it together.

Last updated