Reference to Video
Reference to Video is MiniMax H3's most powerful task — and the one that makes it a true omni-modal model. Instead of animating an image directly (that's Image to Video), you hand H3 references — images, a video clip, an audio track, or any mix of them — and it builds a brand-new shot around them. Your character's face, your product, a dance from another video, a recorded voice line: H3 reads them all as one context and fuses them into a single coherent video.
Reference to Video vs. Image to Video
The distinction matters, because they solve different problems:
Image to Video uses your image as the literal first frame. The video starts from exactly that picture.
Reference to Video uses your images as ingredients. The character from your photo can appear mid-scene, from a new angle, in a new outfit, doing something entirely new — the composition is invented fresh from your prompt, which is also why this task gives you an aspect ratio control while Image to Video doesn't.
Rule of thumb: if you want this exact picture to move, use Image to Video. If you want this character / thing / style in a new shot, use Reference to Video.
The Three Reference Types
Image references (up to 5)
Lock the identity of anything: a character's face, an outfit, a product, a pet, a location, or an overall art style. H3 keeps them consistent throughout the clip. You can combine subjects — for example, two character photos plus one background photo — and direct them in the prompt: "the woman from <Picture 1> and the man from the <Picture 2> sit across from each other in the café from the third image."
Use <Picture N> in prompt to reference the image
Video reference (one clip)
Borrow motion from existing footage: choreography, a fight sequence, camera movement, a gesture. H3 extracts how things move and re-performs it with your subject and scene. The reference clip is trimmed to your target clip length, so pick the exact seconds of motion you want. Note that attaching a video reference raises the Standard-mode price from 6 to 7 credits per second.
Use <Video 1> in prompt to reference the video
Audio reference (one track, 2–15 seconds)
Drive the clip with sound: a recorded voice line, a sound effect, a snippet of music. H3 generates the video to the audio — a character speaking your recorded dialogue with matching lip movement and delivery is the classic use.
On H3, an audio reference is a complete job by itself: you can generate with only a voice recording and a prompt, no images required. Describe who is speaking and where, and H3 invents the performer around your audio.
Use <Audio 1> in prompt to reference the video
Mixing them
The real magic is combination — one generation can use all three at once. Character photos define who, a video reference defines how they move, an audio track defines what they say or dance to, and the prompt ties it together.
Last updated