> For the complete documentation index, see [llms.txt](https://docs.dreamerland.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.dreamerland.ai/video/minimax-h3/reference-to-video.md).

# Reference to Video

Reference to Video is MiniMax H3's most powerful task — and the one that makes it a true omni-modal model. Instead of animating an image directly (that's Image to Video), you hand H3 **references** — images, a video clip, an audio track, or any mix of them — and it builds a brand-new shot *around* them. Your character's face, your product, a dance from another video, a recorded voice line: H3 reads them all as one context and fuses them into a single coherent video.

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Ftbrc6W2n35oVxftMovDQ%2Fuploads%2FH1EVCP91DDsIb9d2uaw8%2Fc3vVxXbRZGkAGGLoL9TJ_0.mp4?alt=media&token=5e457d35-d811-4ac4-9b0c-7b38fd4868b0>" %}

\
**Reference to Video vs. Image to Video**

The distinction matters, because they solve different problems:

* **Image to Video** uses your image as the *literal first frame*. The video starts from exactly that picture.
* **Reference to Video** uses your images as *ingredients*. The character from your photo can appear mid-scene, from a new angle, in a new outfit, doing something entirely new — the composition is invented fresh from your prompt, which is also why this task gives you an **aspect ratio** control while Image to Video doesn't.

Rule of thumb: if you want *this exact picture to move*, use Image to Video. If you want *this character / thing / style in a new shot*, use Reference to Video.

### The Three Reference Types

#### Image references (up to 5)

Lock the identity of anything: a character's face, an outfit, a product, a pet, a location, or an overall art style. H3 keeps them consistent throughout the clip. You can combine subjects — for example, two character photos plus one background photo — and direct them in the prompt: *"the woman from \<Picture 1> and the man from the \<Picture 2> sit across from each other in the café from the third image."*

*Use \<Picture N> in prompt to reference the image*

#### Video reference (one clip)&#x20;

Borrow **motion** from existing footage: choreography, a fight sequence, camera movement, a gesture. H3 extracts how things move and re-performs it with your subject and scene. The reference clip is trimmed to your target clip length, so pick the exact seconds of motion you want. Note that attaching a video reference raises the Standard-mode price from 6 to 7 credits per second.

*Use \<Video 1> in prompt to reference the video*

#### Audio reference (one track, 2–15 seconds)

Drive the clip with **sound**: a recorded voice line, a sound effect, a snippet of music. H3 generates the video *to* the audio — a character speaking your recorded dialogue with matching lip movement and delivery is the classic use.

On H3, an audio reference is a complete job by itself: you can generate with **only** a voice recording and a prompt, no images required. Describe who is speaking and where, and H3 invents the performer around your audio.

*Use \<Audio 1> in prompt to reference the video*

#### Mixing them

The real magic is combination — one generation can use all three at once. Character photos define *who*, a video reference defines *how they move*, an audio track defines *what they say or dance to*, and the prompt ties it together.
