Intro to MiniMax H3

July 30, 2026

MiniMax H3: Create video with visuals, motion, and native audio

MiniMax H3 is a multimodal video model built to work with more than a visual prompt.

It can understand text, images, video, and audio together, then use those inputs to generate a unified audiovisual result. That means a single generation can combine the look of a character, the movement of a reference video, the sound of a voice, camera direction, dialogue, ambience, and music.

Instead of thinking only about what the video should look like, H3 works best when you describe what the complete scene should be like to watch and hear.

What makes H3 different

Many video workflows separate creation into different tasks: text-to-video, image-to-video, motion transfer, character reference, voice generation, sound effects, and editing.

H3 is designed to bring these signals into a shared context. MiniMax describes this as general-purpose multimodal generation: references and instructions can work together instead of being treated as isolated modes.

For example, you can define:

  • what a character looks like
  • how they move
  • how the camera behaves
  • what they say
  • what their voice should sound like
  • what happens in the environment
  • what sound effects are present
  • what music accompanies the scene

H3 then reasons across those elements as parts of the same video.

Start with the whole scene

A useful H3 prompt describes the creative result rather than listing disconnected keywords.

Instead of:

Woman, train, cinematic, rain, camera movement.

Try:

A cinematic close shot of a woman sitting beside the window of a moving train at night. Rain runs across the glass as city lights streak through her reflection. She looks down at a folded letter, then slowly raises her eyes toward the window. The camera gently tracks to the right. Train wheels rumble underneath quiet rainfall, with sparse piano playing in the background.

The second version gives H3 relationships to work with: subject, environment, action, camera, timing, ambience, and music.

Choose the right starting point

H3 can work from several kinds of creative context.

With text only, describe the complete scene from scratch.

With a starting image, treat the image as the opening state. Describe what begins moving and how the scene develops from there.

With starting and ending images, focus on the path connecting them. Describe the physical changes, camera movement, subject motion, and transitions that should naturally lead from the first composition to the last.

With references, different assets can contribute different parts of the result. An image might define the character, a video might define camera movement, and audio might define a voice or musical direction. MiniMax calls this its Omni Reference workflow.

Think beyond the picture

H3 generates audio together with video, so sound can be part of the original creative direction rather than an afterthought.

You can describe:

Dialogue
Without looking away from the window, she quietly says, ‘I think this is my stop.’

Environmental audio
Rain taps against the glass while the train produces a steady metallic rhythm.

Physical sounds
The paper softly unfolds in her hands.

Music
Sparse piano notes play at a slow tempo, joined by a sustained low cello.

The official H3 prompting guidance separates these ideas into the visual/action timeline, the overall soundscape, and non-diegetic music.

You do not need to make every prompt complicated. The useful principle is simply to think about the complete audiovisual moment.

Where H3 is especially useful

H3 is a strong fit when the video depends on several kinds of creative direction working together.

That includes character-driven scenes, commercials, product films, stylized content, music-driven work, narrative clips, motion references, dialogue, and scenes where sound is an important part of the result. MiniMax's own examples span brand films, social creative, short-form narrative, product advertising, UI motion, games, and stylized animation.

A helpful reminder

More inputs do not automatically make a generation better.

Give each instruction or reference a clear job. If an image defines the character, a video defines motion, and audio defines the voice, say so. Clear relationships give the model a stronger creative hierarchy than a collection of unexplained references.

Think of H3 as an audiovisual scene generator rather than only a video generator.

Describe what should happen, how it should look, how it should move, and what the audience should hear. When you use references, tell H3 what each reference contributes to the final result.