Get the Most from MiniMax H3 Prompts
Last Updated
Aug 5, 2026Fresh
Models Tested
MiniMax H3
Introduction to MiniMax H3 Video Generation
MiniMax H3 generates 15-second 2K video with sound built in: voices, ambient noise, and music come out of the same generation as the picture, with lips matched to every spoken line. It powers the highest-resolution tier of our AI video generator, and it has one ability the rest of the lineup lacks: it animates photos of real people.
That combination changes what a prompt needs to do. You are directing picture and soundtrack at once, and often a real face. This guide covers the prompting rules that make those clips land: how to write dialogue the model can perform, how to work with photos of real people responsibly, and how to use reference files to keep characters consistent.
Every technique here comes from hands-on testing with the model on Ambience AI, so you can skip the trial and error and start from prompts that work.
Understanding MiniMax H3's Capabilities
MiniMax H3 is a native-audio video model: it generates the picture and the soundtrack together, so voices land in sync with lips and footsteps land in sync with feet. Every clip is 15 seconds at 2K resolution, which gives you room for a full spoken beat rather than a silent loop.
Core Capabilities
What Makes H3 Different
- Native sound and dialogue - Voices, ambient noise, and music generated with the picture and lip-synced automatically
- 2K resolution - The highest-resolution video output on Ambience AI
- 15-second clips - Enough time for a setup, a spoken line, and a closing beat
- Real people from photos - The only native-audio model that accepts photoreal human images as input, where Seedance 2.0 screens them out
Generation Modes
- Text-to-Video - Describe the scene, the action, and the sound entirely in words
- Image-to-Video - Animate a photo; the clip follows your image's framing and aspect ratio
- Reference-to-Video - Provide up to 12 reference files and direct them by name in the prompt (Team and Business plans)
Good to Know Before You Start
- Every clip is 15 seconds and 800 credits, on every mode
- Generations take around 10 minutes; the extra time pays for 2K detail and a full soundtrack
- There is no seed setting, so two runs of the same prompt will come out differently
- Prefer stylized or illustrated characters? Our Seedance 2.0 guide covers the other native-audio model in the lineup
Animating Real People from Photos
This is the ability creators reach for H3 to get: hand it a photo of a real person and it produces a clip of that person moving and speaking, as long as your description matches the photo. A founder photo becomes a product announcement. A team headshot becomes a welcome message. Seedance 2.0 screens photoreal people out of its inputs, so among the models that generate sound, H3 is the one to pick for this work.
Use It with Consent
Animate your own photos and photos of people who agreed to appear in your video: yourself, your team, your clients, and collaborators who said yes. That covers every legitimate use we know of.
Leave everyone else out of it. Putting words in the mouth of someone who never agreed to appear can cause real harm, and content like that is against our terms.
Choosing a Source Photo
1. Sharp and Well Lit
A crisp, evenly lit photo gives the model the facial detail it needs to work from at 2K.
2. Face Clearly Visible
Front-facing or three-quarter angles animate best. Sunglasses, heavy shadows, and extreme angles fight the lip sync.
3. One Main Subject
A single clear subject keeps the motion and the voice attached to the right person.
4. Room to Move
Leave some space around the subject. Tight crops limit gestures and camera movement.
Prompt the Performance, Skip the Appearance
The photo already tells the model what the person looks like. Spend your prompt on what they do, how they feel, and what they say:
integrated_multimodal_description: [Shot 1] Live-action, a medium close-up holds on the woman from the photo at her desk. She looks up from the laptop, breaks into a warm smile, and with a confident, easy voice (S1) says: <d>[English] We just shipped the feature you have all been asking for.</d> She closes the laptop and leans back, holding the frame to the end. overall_soundscape: Quiet office ambience with a laptop lid closing and a chair creaking softly. non_diegetic_music: None.
Notice there is no physical description at all. Delivery cues like "warm smile" and "confident voice" steer the performance; the photo handles the rest.
Make the Prompt Agree with the Photo
H3 follows your words over your image. When the two disagree, the words win, and the clip can cut to a different person partway through to satisfy the description. We tested this both ways on the same photo of a bearded man: a prompt that said "she" held him for about four seconds, then cut to a woman for the rest of the clip, while a prompt that described him held his face, his clothes, and the room for all 15 seconds. Match the description to the photo and the likeness holds.
- Match the pronouns to the person in the photo, in every sentence
- Match the role and the setting too. A "barista behind the bar" over an office headshot invites the same swap
- Do not describe a second person unless you want one to appear
Writing Dialogue That Sounds Right
Native audio means the model performs your script. The difference between clean, natural speech and rushed, garbled speech almost always comes down to how much dialogue you ask for and how you format it.
The Dialogue Budget
- About 20 words of dialogue per 15-second clip. Speech needs room to breathe alongside the action.
- About 10 words per spoken line. Short lines get clean, well-paced delivery.
- One sentence per quoted line. Two sentences inside one quote is where garbled speech starts. Split them into separate quotes instead.
- Tag the spoken words. H3 performs whatever sits inside
<d>[English] ... </d>and treats everything outside it as action, so it never reads your stage directions aloud. Give each speaker an ID the first time they appear:(S1), then(S2).
Time Every Shot After the First
H3 cuts on a clock. Shot 1 carries no timestamp, and every shot after it opens with one: [Shot 2] At 00:05.000, the camera cuts to .... In our own testing the cuts landed within 0.12 seconds of the mark, while the same brief written without timestamps held one static framing for the entire clip.
- Give every shot a beat: a line or a named action, including the last one. Dialogue bunches at the start of a shot, so a closing shot with nothing to do is where dead air comes from
- Steady the mouth: add "realistic lip articulation, no exaggerated mouth opening" to keep the performance from going rubbery
- Hold the frame while speaking: "locked camera, no head turns while speaking" protects lip sync
- Say what the score is: H3 has a field for it, so write
non_diegetic_music: None.instead of stacking "no music" negatives at the end of the prompt
The Anatomy of a Spoken Clip
1. Setup
A beat of visual action before anyone speaks: who is here, where they are, what they are doing.
2. Voice Descriptor
Say how the line is delivered: "in a low, weathered voice" or "bright and energetic". It shapes tone, pace, and emotion.
3. The Line
One short quoted sentence. If there is more to say, give it a second, separate quote.
4. Closing Beat
End on a visual action after the last line, so the speech finishes cleanly instead of being cut off.
A Complete Example
H3 reads three named fields rather than one paragraph. Describe your clip in the Ambience chat and it writes this structure for you; write it yourself when you want the cuts exact.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium close-up frames a florist trimming stems at a sunlit market stall. The camera pushes in with small amplitude at slow speed as the florist with a warm, unhurried voice (S1) looks up and says: <d>[English] These came in this morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of her hands folding paper around the bouquet as she says: <d>[English] They will last all week.</d> [Shot 3] At 00:10.000, the camera cuts to a medium shot as she sets the finished bouquet on the counter and rests both hands beside it, holding the frame to the end. overall_soundscape: Market ambience with distant chatter, snipping shears, paper folding crisply around stems, and the soft knock of the bouquet set down on wood. non_diegetic_music: None.
Style up front: named once in Shot 1, then left alone
Timed cuts: Shot 2 and Shot 3 open on the clock, so the edit lands where you asked
Camera: a verb with amplitude and speed, kept separate from what the subject does
Speech: voice described once, speaker ID (S1), spoken words inside the tag
Sound: what a microphone in the room would pick up, with the score declared separately
Direct the Whole Soundtrack
H3 splits audio in two, which is the difference between sound that belongs to the room and sound the audience hears over it. Anything you leave out is left to the model, and it reaches for background music.
- overall_soundscape is what a microphone in the scene would pick up: rain on the window, a distant train, a cup set down. Name each one, and leave the dialogue out of it
- non_diegetic_music is the score. Describe it by instrument, rhythm, and dynamics: "solo piano, slow, fading under the last line"
- Want silence? Write
non_diegetic_music: None.That one line does what a pile of "no music" negatives is trying to do
Reference Files and How to Cite Them
Reference-to-video lets you hand H3 a set of files and direct them like a cast: this person, wearing this jacket, in this room, speaking over this soundtrack. It is available on Team and Business plans.
What You Can Attach
- Up to 9 images - people, products, wardrobe, and locations
- Up to 3 videos - motion or framing you want echoed
- Up to 3 audio files - a voice or sound bed to match
- 12 files maximum in total across all three types per generation
Label Each File and Say What Must Survive
H3 refers to attachments by label: <Picture 1> for a still, <Video 1> for a clip, <Audio 1> for a voice or sound, and <Subject 1> for the person or object those files describe. Numbering follows the order you attach them. What makes the difference is telling the model which parts are not allowed to change. A reference run also renames the description field to detailed_description and adds a one-line summary:
subject_definitions: <Subject 1> is the presenter shown in <Picture 1>, including her face and hair. <Subject 2> is the product shown in <Picture 2>. summary: [reference generation] <Subject 1> presents <Subject 2> to camera at a studio desk. retention_analysis: <Picture 1> (face, hairline): fully_preserved. <Picture 2> (product shape and label): fully_preserved. detailed_description: [Shot 1] Live-action, a medium shot frames <Subject 1> at a studio desk. She picks up <Subject 2>, turns it toward the camera, and with a friendly, even voice (S1) says: <d>[English] This is the part everyone asks about.</d> overall_soundscape: Quiet studio room tone with the soft handling of the product. non_diegetic_music: None.
The markers are fully_preserved, partially_preserved, attribute_transfer, and weak_reference. Use fully_preserved for a face you need to stay itself, and attribute_transfer when you only want a quality from the reference, like its colour or material. Ambience writes this structure for you when you attach files in chat.
Reference Tips That Pay Off
- Give every file a job. A reference with no role in the prompt is noise the model has to guess about.
- Fewer, sharper references win. Three files with clear roles beat nine loosely related ones.
- Consent applies here too. Reference photos of people should be your own or from people who agreed to appear.
- Count before you generate. Stay within 9 images, 3 videos, 3 audio files, and 12 total.
Technical Settings & Planning
H3 keeps its settings simple: one duration, one resolution, one price. What is left to plan is how you spend each generation.
Fixed Parameters
Duration & Resolution
• 15 seconds per clip, every time
• 2K resolution, the sharpest video output on Ambience AI
• 800 credits per generation on every mode
Aspect Ratio
• Text-to-video and reference-to-video: choose your dimensions as usual
• Image-to-video: the clip follows your source photo's aspect ratio
Generation Time
• Around 10 minutes per clip
• Keep creating elsewhere while it runs; your clip lands in your Library when it is done
Working Without a Seed
Expect Variation
H3 has no seed setting, so the same prompt produces a different take each run. Treat each generation as a take, the way a director would.
Save What Works
Keep a note of prompts that landed. The wording is your only lever for consistency across runs, so reuse it precisely.
Plan Important Shots
For a clip that matters, budget two or three takes and pick the best. Anchoring with a source photo or reference files also narrows the variation.
Troubleshooting & Solutions
Most H3 problems trace back to a handful of causes, and each has a reliable fix.
Garbled or Rushed Speech
Problem: Dialogue comes out slurred, hurried, or half-swallowed
Solution: Your lines are too long. Keep each quote to one sentence of about 10 words, and stay near 20 words of dialogue total
Unwanted Background Music
Problem: A soundtrack appears where you wanted quiet
Solution: Name the room sounds you do want in overall_soundscape, and write non_diegetic_music: None. so the score is settled rather than guessed
Flat or Robotic Delivery
Problem: The voice is intelligible but lifeless
Solution: Add a voice descriptor and an emotion cue: "says warmly, with a small laugh" gives the model a performance to aim for
A Different Result Every Run
Problem: Re-running a prompt changes the scene, the voice, or the pacing
Solution: That is expected without a seed. Reuse exact wording, anchor with a photo or references, and generate multiple takes of key shots
The Person Changes Mid-Clip
Problem: Your photo animates, then the clip cuts to someone else
Solution: Your prompt describes a different person than your photo shows. Match pronouns, role, and setting to the photo, and drop any second character
Reference Files Rejected
Problem: A reference generation is turned away before it starts
Solution: Check your counts: at most 9 images, 3 videos, 3 audio files, and 12 files in total. Drop the extras and retry
Still Processing After a While
Problem: The clip seems to be taking forever
Solution: Around 10 minutes is normal for 15 seconds of 2K video with sound. The progress bar keeps updating until the clip lands in your Library
A Debugging Habit That Works
- Change one thing at a time: dialogue length, then audio line, then references, so you know what fixed it
- Trim before you add: most speech problems are solved by cutting words, and most audio problems by naming sounds
- Keep a prompt log: with no seed to pin a result, your notes are the reproducibility system
Start Creating with MiniMax H3
MiniMax H3 turns a photo and a paragraph into 15 seconds of 2K video with a finished soundtrack. The craft is in the writing: budget your dialogue, describe the voice, name every sound, and end on a visual beat. With those habits, the clips come out sounding like you directed them, because you did.
Start with a single spoken line and a simple scene, then work up to reference casts and multi-line exchanges. Generate with H3 today in our AI video generator.
Want to go deeper on video prompting? Read the Seedance 2.0 Prompting Guide for the other native-audio model, the Kling Prompting Guide for cinematic camera work, or the WAN Prompting Guide for fast, atmospheric clips. Or explore the complete suite of AI creative tools on Ambience AI.
Sources & Citations
This guide is based on hands-on testing with MiniMax H3 on Ambience AI, alongside official documentation and provider sources.
- MiniMax Official Website - model announcements and research notes
Ready to Make Videos That Talk?
Put these prompting techniques to work with MiniMax H3. Animate your photos, write your lines, and get 2K clips with sound built in.