MiniMax H3 (Hailuo 3) is live on EvoLinkTry it with 10 free credits

MiniMax H3 Prompts and Video Examples

Browse 12 verified prompts with real MiniMax H3 outputs, input requirements, and reusable variables. Copy a template or send it directly to the EvoLink playground.

MiniMax H3 is also known as Hailuo 3, Hailuo 3.0, or Hailuo 03.

Generation mode

Use case

12 of 12 prompts

Sci-fi trailer clip generated with MiniMax H3: a lone figure before a vast circular cosmic gateway as a title resolves out of darknessMULTIMODAL REFERENCE
15s16:92 reference assets

Cinematic & VFX

Sci-Fi Mystery Trailer

A two-reference cinematic trailer beat: one image sets the atmosphere, the other locks the protagonist, and the prompt drives a single push-in with an on-screen title and matching audio.

View details

Full prompt

Realistic cinematic look, high-contrast lighting, and a tight pace. Use Figure 1 as the overall atmosphere and style reference, and Figure 2 as the protagonist reference. Shot 1 — Ultra-wide establishing shot. A huge circular cosmic gateway nearly fills the frame. The person is only a tiny figure seen from behind before the gateway, positioned toward the lower right. The ground is wet and reflective, and the center of the gateway is pitch black. The camera slowly pushes forward. A large title fades in from the edge of the darkness, blurred at first and then sharp: "THE STARS WERE LISTENING". Use an extremely condensed, heavy, all-caps typeface in dark red mixed with rust red, with subtle grain and misted edges. Audio: a deep low-frequency pulse, faint metallic vibrations in the distance, and a soft hit as the text becomes sharp. → Hard cut.

What you need to supply

  • Image 1: an atmosphere and style plate — the environment, palette and grade you want the shot to inherit.
  • Image 2: the protagonist, shot large enough that the face and silhouette stay readable when the figure is small in frame.
  • A title string you actually want burned in; H3 renders the text you write, so keep it short and spell it exactly.

Why it works

  • Each reference is given one job — Image 1 for atmosphere, Image 2 for the character — so the model never has to guess which plate owns the look.
  • The shot is described as one continuous push-in with a single event (the title resolving), which is what a 15-second budget can actually hold.
  • Audio is written as three specific layers (low pulse, distant metallic vibration, a hit on the text) instead of a vague 'cinematic soundtrack'.

Swap these out

  • title text and typeface treatment
  • gateway or landmark in the establishing shot
  • color of the title
  • audio bed
  • final cut behavior

Constraints

  • Reference the assets by array position — 'Figure 1', 'Figure 2' — matching image_urls order. The @image1 syntax is not part of this API contract.
  • The published prompt is a starter, deliberately incomplete: it ends on a hard cut so you can extend it with your own second shot.

Settings

Multimodal Reference · 15s · 16:9 · 2 reference assets

Text-to-video clip generated with MiniMax H3: a dusk kitchen filmed handheld while a hand-drawn glowing creature moves among the propsTEXT TO VIDEO
15s16:9No reference assets

Cinematic & VFX

Glowing Kitchen Creature

A pure text-to-video prompt that mixes live-action kitchen footage with hand-drawn glowing animation, and spends most of its words on camera imperfection and what must never appear.

View details

Full prompt

15-second, 16:9 landscape video. Blend live-action footage of a small kitchen at dusk with hand-drawn glowing animation. The last light of sunset lingers by the window. The lived-in kitchen contains an old wooden table, a half-washed mug, a slightly fogged glass bottle, and a hanging dishcloth. Give the footage subtle one-handed smartphone shake, hesitant close-range focusing, exposure fluctuations caused by backlight, and slightly coarse noise in the shadows. It should not look carefully arranged like an advertisement; instead, it should feel like someone hurriedly captured an unbelievable event at home. Do not show huge eyes, gaping mouths, fangs, threatening or lunging movements, sudden black frames, or jump scares. Use only kitchen room tone, cloth rubbing, the soft clink of a mug, water dripping from the faucet, the camera operator's footsteps and quiet breathing, plus gentle electronic sounds and tiny calls from the hand-drawn creature.

What you need to supply

  • Nothing to upload — this route takes the prompt only.
  • Decide duration and aspect ratio in the request, not in the prompt text.

Why it works

  • It specifies the camera's flaws — one-handed shake, hesitant focus, backlight exposure swings, noise in the shadows — which is what sells 'someone filmed this at home' over 'this is an advert'.
  • It carries an explicit do-not list (no huge eyes, no fangs, no lunging, no jump scares) that keeps a cute creature from drifting into horror.
  • The audio direction names each source separately: room tone, cloth, mug, tap, footsteps, breathing, plus the creature's own sounds.

Swap these out

  • room and time of day
  • creature design and behavior
  • props on the table
  • which imperfections the camera shows
  • audio layers

Constraints

  • The prompt states its own length and framing in words; the API still needs duration and aspect_ratio as request fields.
  • Negative instructions work best as a short explicit list. Stacking dozens of prohibitions costs prompt budget you need for the action.

Settings

Text to Video · 15s · 16:9 · No reference assets

Vertical short-drama trailer clip generated with MiniMax H3: a vampire lead and a human heroine in a candlelit castle interiorMULTIMODAL REFERENCE
15s9:162 reference assets

Short Drama & Narrative

Vampire Romance Short Drama

A vertical short-drama hook that pins both leads to one reference image and the location to another, then spends the rest of the prompt on relationship beats and shot sizes rather than plot.

View details

Full prompt

Generate a 15-second, 9:16 vertical trailer segment for an international live-action vampire romance short drama. Use Figure 1 as the appearance reference for the male and female leads, and Figure 2 as the scene reference. Keep both leads' identities consistent, with a realistic live-action look and premium short-drama production quality. Story: an innocent human heroine accidentally enters a forbidden area of an old castle and awakens a sleeping aristocratic vampire. He discovers that she carries an aura connected to an ancient war, which sparks a powerful urge to control her and a dangerous fascination with her. She fears him but does not completely submit, resisting his pressure. Overall style: an international ReelShort / DramaBox vampire-romance trailer. Dark romance, dangerous attraction, fate, intense control, brooding oppression, and a striking reversal. Keep the visuals premium, restrained, and tightly paced, like the opening 15-second hook of a hit short drama. No gore, cheap horror, Halloween aesthetic, or modern street feel. Format: 9:16 vertical composition for TikTok / ReelShort / DramaBox. Use primarily medium close-ups, close-ups, and extreme close-ups, emphasizing faces, eye contact, pressure, and relationship tension within the vertical frame.

What you need to supply

  • Image 1: both leads together, so the model reads them as one consistent casting.
  • Image 2: the location plate — the castle interior, its light and its materials.
  • A one-line premise and a stated emotional reversal; a 15-second hook can carry one turn, not a full episode.

Why it works

  • Naming the genre reference (ReelShort / DramaBox trailer) transfers pacing, grade and framing conventions in a few words.
  • It fixes the shot vocabulary — medium close-up, close-up, extreme close-up — which is what makes a vertical frame read as premium instead of cramped.
  • The exclusion list (no gore, no cheap horror, no Halloween look, no modern street feel) removes the four ways this genre usually degrades.

Swap these out

  • lead appearance references
  • location
  • premise and reversal
  • target platform look
  • shot-size mix

Constraints

  • Vertical output is requested through aspect_ratio on the reference route, not by describing '9:16' in the prompt alone.
  • Identity holds far better when the leads arrive as one image than as two separately cropped portraits.

Settings

Multimodal Reference · 15s · 9:16 · 2 reference assets

Vertical eyewear commercial clip generated with MiniMax H3: two models in a seamless white studio wearing futuristic wraparound glassesMULTIMODAL REFERENCE
15s9:163 reference assets

Brand & Product Ads

Futuristic Eyewear Campaign

A three-image fashion commercial where each reference owns a different layer: the key visual, the models' faces, and the product design itself.

View details

Full prompt

Generate a vertical screen 9:16 high-end fashion glasses commercial, taking overall reference to the storyboard rhythm, editing speed, white studio texture and cool fashion atmosphere of the given video. The picture is a minimalist white booth, a seamless white background, a strong sense of high-end advertising, and a clean, simple, handsome, avant-garde, international first-line fashion blockbuster texture. Key visual character reference picture 1, two full-body female models, one black female model and one European and American model, maintain their high-end clothing, body posture, white studio light and shadow, fashion show temperament and overall cool attitude. Both of them wear futuristic high-end glasses. The design of the glasses refers to Figure 3, emphasizing the covered curved surface, sharp geometric cat-eye/goggle hybrid outline, mirror reflection, streamlined temples, and the texture of high-end fashion accessories. Please refer to Figure 2 for the appearance details of the two characters.

What you need to supply

  • Image 1: the key visual — full-body models, wardrobe, studio light and attitude.
  • Image 2: appearance detail for the two characters, so faces stay consistent through the cuts.
  • Image 3: the product, shot clearly enough that its silhouette and materials survive at speed.
  • A seamless studio background you are willing to keep for the whole clip.

Why it works

  • Three references with three explicit jobs is the pattern the reference route is built for; ambiguity is what makes multi-image prompts collapse.
  • The product is described by its geometry — wrap curvature, cat-eye/goggle hybrid outline, mirrored surface, streamlined temples — not just by its name.
  • A single-material white studio removes background variance, so the model spends its budget on the product and the performers.

Swap these out

  • product category
  • model casting
  • studio color
  • editing speed
  • wardrobe

Constraints

  • The published prompt also asks for the pacing of a given video. The released asset set for this case is three images, so treat that line as a style instruction, or attach your own clip as Video 1.
  • Up to 9 images, 3 videos and 3 audio clips per reference request, and no more than 12 files in total; audio can never be the only reference type.

Settings

Multimodal Reference · 15s · 9:16 · 3 reference assets

Claymation clip generated with MiniMax H3: a clay fox leaping across a lava canyon as the camera sweeps beneath itFIRST / LAST FRAME
10s16:91 reference asset

Character & Motion

Clay Fox Canyon Leap

A single start frame animated into one decisive action, with the camera move described as precisely as the action itself.

View details

Full prompt

Claymation style. A sprinting fox reaches the edge of a cliff and launches without hesitation, making a dramatically tense, heroic slow-motion leap across a vast lava canyon. While the fox is airborne, the camera rushes at high speed beneath its belly in a sweeping dynamic move, fully revealing the terrifying depth of the chasm and the fox's clay body at maximum extension in midair.

What you need to supply

  • A start image that already carries the style — here the claymation fox and its material.
  • One action you want to happen. Not three.

Why it works

  • The prompt describes change, not the picture. The start frame already holds the appearance, so every word buys motion.
  • The camera has its own instruction — a high-speed sweep beneath the fox's belly — which turns a jump into a shot.
  • Naming the peak moment ('maximum extension in midair') gives the model a target pose to build the timing around.

Swap these out

  • character and material style
  • environment and hazard
  • camera path
  • slow-motion emphasis

Constraints

  • The image-to-video route derives the output ratio from the input image and rejects aspect_ratio outright.
  • Do not re-describe what the start frame already shows; repeating static appearance is the most common way to waste an image-to-video prompt.

Settings

First / Last Frame · 10s · 16:9 · 1 reference asset

Video edit generated with MiniMax H3: a source clip with the newspaper, chair, sunglasses and burning car all replaced or removedMULTIMODAL REFERENCE
10s16:91 reference asset

Editing & Transformation

Multi-Element Scene Edit

Six separate edits to one source clip, written as a flat list of change instructions with no scene description at all.

View details

Full prompt

Replace the newspaper in the reference video with a green-covered book; change the chair the character is sitting on to a red sofa; remove the sunglasses worn by the character to retain a clear face; remove the car burning effect to keep the vehicle in a normal state; change the photo the character takes out of his arms to a small black book; and add a tree on the left side of the screen

What you need to supply

  • Video 1: the clip to edit, 2–15 seconds, MP4 or MOV, H.264 or H.265, up to 50 MB.
  • A list of edits, each naming what to change and what it becomes.

Why it works

  • Every instruction is a pair — target plus result — so nothing is left to interpretation ('the newspaper' becomes 'a green-covered book').
  • Removals state the intended end state ('remove the sunglasses to retain a clear face'), which stops the model from leaving a hole where the object was.
  • It carries no scene description whatsoever, so the model treats the source clip as ground truth and only applies the deltas.

Swap these out

  • number of edits
  • objects replaced
  • effects removed
  • elements added

Constraints

  • EvoLink exposes three H3 routes; edit-style work goes through reference-to-video with the source clip as Video 1.
  • Reference video duration is billable, so trim the source to the section you actually need before uploading.
  • State what must stay unchanged when an edit sits next to something you care about — unstated regions are fair game for the model.

Settings

Multimodal Reference · 10s · 16:9 · 1 reference asset

Visual-novel interface transition generated with MiniMax H3 between a fixed opening frame and a fixed final frameFIRST / LAST FRAME
15s16:92 reference assets

Short Drama & Narrative

Otome Visual Novel Transition

A true first-and-last-frame prompt: two images fix both ends of the shot, and the text only has to describe the journey between them.

View details

Full prompt

Use the first image as the opening frame and the second image as the exact final frame to generate an otome visual-novel interface transition. Overall feel: a premium Chinese otome romance-interaction interface capturing an intimate moment before and after a performance. Transition naturally from "choose to watch his performance" to "Han Xu is drawn in by the heroine's words and reacts with intrigued interest." UI text, choices, and dialogue boxes should appear with refined otome-game presentation. Keep the transition silky smooth and the emotion suggestive yet restrained.

What you need to supply

  • Image 1: the opening frame, complete with its UI state.
  • Image 2: the exact closing frame you want to land on.
  • A one-line description of the emotional change between the two states.

Why it works

  • Both endpoints are locked, so the model solves interpolation instead of composition — the most reliable way to get a predictable shot.
  • The prompt names the emotional transition ('choose to watch his performance' → 'drawn in and intrigued') rather than listing frames, so the performance carries the cut.
  • It states the interface elements should animate in the genre's own idiom, which keeps the UI from being redrawn.

Swap these out

  • opening and closing frames
  • emotional arc between them
  • UI presentation style
  • transition speed

Constraints

  • Send both frames as image_start and image_end on the image-to-video route; the reference route does not accept these fields at all.
  • The closer the two frames are in framing and lighting, the smoother the interpolation. Two unrelated compositions produce a cut, not a transition.

Settings

First / Last Frame · 15s · 16:9 · 2 reference assets

Motion-transfer clip generated with MiniMax H3: two referenced characters performing street dance copied from a reference videoMULTIMODAL REFERENCE
10s16:93 reference assets

Character & Motion

Street Dance Motion Transfer

Twenty-odd words that move choreography from a reference clip onto two characters supplied as images — the clearest demonstration of what ordered references buy you.

View details

Full prompt

Have the characters perform street dance following the movements in Video 1. Use Figure 1 and Figure 2 as the character references.

What you need to supply

  • Image 1 and Image 2: the two characters, one clean full-body reference each.
  • Video 1: the movement to copy, 2–15 seconds, with the performer fully in frame.

Why it works

  • The prompt is short because the references carry the information — the video owns the motion, the images own the identities.
  • Each asset is addressed by its array position, so there is no ambiguity about which reference supplies what.
  • It asks for nothing else. No lighting, no camera, no style — every extra instruction would compete with the motion it is trying to copy.

Swap these out

  • characters
  • source choreography
  • environment
  • number of performers

Constraints

  • Total reference video duration must stay within 15 seconds, and each clip must be 2–15 seconds at 23.976–60 FPS.
  • Reference videos must be MP4 or MOV with H.264 or H.265, up to 50 MB each, with the whole JSON body under 64 MB.

Settings

Multimodal Reference · 10s · 16:9 · 3 reference assets

Dialogue replacement generated with MiniMax H3: a character's spoken line swapped for a new line from a reference audio clipMULTIMODAL REFERENCE
10s16:92 reference assets

Native Audio & Dialogue

Dialogue and Performance Replacement

Swapping a spoken line inside an existing clip: the old line is quoted, the new line is quoted, and the performance is allowed to shift just enough to match.

View details

Full prompt

Replace the girl's line in Video 1, "We can't be together. It's not that we don't love each other; we truly can't make it to the end," with the line from Audio 1: "Don't go, okay? This time, let's not let go of each other." Slightly adjust the corresponding performance.

What you need to supply

  • Video 1: the clip containing the line to replace.
  • Audio 1: the replacement line, WAV or MP3, up to 15 MB and 15 seconds.
  • Both lines written out verbatim in the prompt.

Why it works

  • Quoting the outgoing line tells the model exactly which span of the clip to operate on, instead of 'the dialogue near the middle'.
  • Quoting the incoming line means the lip sync has a target rather than being inferred from the audio alone.
  • 'Slightly adjust the corresponding performance' grants a bounded licence to change the acting — bounded, so the rest of the take survives.

Swap these out

  • source clip
  • outgoing line
  • replacement line and voice
  • how much performance may change

Constraints

  • Audio can never be the only reference type; it must arrive with an image or a video.
  • Reference audio and video each cap at 15 seconds of total duration per request.
  • Say what stays fixed. Framing, wardrobe and background will drift if the prompt only talks about the line.

Settings

Multimodal Reference · 10s · 16:9 · 2 reference assets

Voice-reference clip generated with MiniMax H3: a character speaking a written line in a timbre taken from a reference audio clipMULTIMODAL REFERENCE
10s16:92 reference assets

Native Audio & Dialogue

Wind Voice Clone

The minimum viable audio-reference prompt: the line to speak, and one clip that defines whose voice speaks it.

View details

Full prompt

Character dialogue: "Follow the wind, live free. Leave worries behind, enjoy the moment." Use Audio 1 as the voice-timbre reference.

What you need to supply

  • Video 1: the character who will deliver the line.
  • Audio 1: a clean sample of the target voice, 2–15 seconds, ideally without music underneath.
  • The exact line of dialogue, written out.

Why it works

  • The dialogue is quoted rather than paraphrased, so timing and lip sync have something concrete to lock onto.
  • Audio 1 is assigned one narrow job — voice timbre — instead of being handed over as a general 'soundtrack'.
  • Nothing else is specified, so the reference clip keeps full control of framing and performance.

Swap these out

  • line of dialogue
  • voice reference
  • character clip
  • delivery pace

Constraints

  • A voice reference transfers timbre, not accent, emotion or pacing; write those into the prompt if they matter.
  • Music or overlapping speakers in the reference audio degrade the result — supply an isolated voice.

Settings

Multimodal Reference · 10s · 16:9 · 2 reference assets

Stage clip generated with MiniMax H3: two magicians swapping suit colors in a puff of smoke as the curtain shifts from red to blueFIRST / LAST FRAME
7s16:91 reference asset

Editing & Transformation

Magician Costume Swap

A staged trick used as an instruction-following test: two costumes swap, one detail must not change, and the background completes a color transition on cue.

View details

Full prompt

Two magicians stand onstage facing the audience and perform a "swap" trick. They wave their wands at the same time, and a cloud of smoke rises. When it clears, their suit colors have switched: the person on the left wears a white suit, and the person on the right now wears a black suit, while both magicians' glove colors remain unchanged. They bow to thank the audience. The red curtain behind them closes, transitioning from deep red to deep blue.

What you need to supply

  • A start image holding both performers, their costumes and the stage.
  • A clear before/after state for whatever swaps.

Why it works

  • The swap is masked by an event — the smoke cloud — giving the model a legitimate moment to make the change instead of morphing on camera.
  • It names what must not change (the glove colors) next to what must, which is exactly how you keep an edit from spreading.
  • The shot ends on a defined final state: the bow, the curtain closing, the color landing on deep blue.

Swap these out

  • performers and costumes
  • the detail that stays fixed
  • cover event for the swap
  • closing color transition

Constraints

  • The image-to-video route derives its aspect ratio from the input image and rejects the aspect_ratio field.
  • Every unstated attribute is fair game for the model. If a detail must survive the change, write it down.

Settings

First / Last Frame · 7s · 16:9 · 1 reference asset

Product film generated with MiniMax H3: premium over-ear headphones rotating above a reflective pedestal in a black studioTEXT TO VIDEO
15s16:9No reference assets

Brand & Product Ads

Luxury Headphones Showcase

A 15-second product film written as four timed blocks, each with its own camera move and job — the most directly reusable template here for anyone with no reference assets.

View details

Full prompt

Create a 15-second luxury cinematic product showcase for premium wireless over-ear headphones. 0–4s: Begin with an extreme macro tracking shot moving across the soft memory-foam ear cushion, fine fabric texture, brushed-metal hinge and precision-machined controls. A narrow light band travels across the surface, revealing realistic materials against a deep black studio background. 4–8s: Pull back into a three-quarter hero view. The headphones rotate slowly above a glossy reflective pedestal. The ear cups pivot naturally while the adjustable headband extends slightly, demonstrating flexible construction and comfort. Maintain exact symmetry, stable geometry and consistent proportions. 8–12s: Transition into an elegant exploded-view reveal. The ear cushion, acoustic driver, internal sound chamber, control ring and outer shell separate smoothly in perfect alignment. Subtle luminous sound waves pulse outward from the driver while the camera performs a restrained side orbit. 12–15s: Every component reconnects seamlessly. The headphones settle into a centered front-facing hero composition as soft rim lighting defines the silhouette. Complete a gentle dolly-in toward the ear cups. Premium technology-commercial finish, controlled reflections, realistic shadows, shallow depth of field, crisp surface detail, stable product shape, no hands, no distortion, no onscreen text.

What you need to supply

  • Nothing to upload — text-to-video takes the prompt alone.
  • A product you can describe by material and mechanism, not just by name.

Why it works

  • Time is split into 0–4s, 4–8s, 8–12s and 12–15s, so the model gets a shot list instead of a wish list.
  • Each block moves the camera differently — macro track, pull back, side orbit, dolly in — which is what makes the clip read as edited rather than drifting.
  • The closing line is a list of prohibitions (no hands, no distortion, no onscreen text) that cover the three ways product renders usually fail.

Swap these out

  • product
  • materials and finish
  • timing of each block
  • background and lighting
  • whether the exploded view appears

Constraints

  • Duration is an integer from 4 to 15 seconds and is set in the request; the timed blocks in the prompt must add up to it.
  • Timed blocks are direction, not a hard timeline. Keep them to four or fewer for a 15-second clip.
  • This clip was published by its author at 720p; the H3 routes on EvoLink output 2K.

Settings

Text to Video · 15s · 16:9 · No reference assets

The MiniMax H3 prompt framework

H3 takes ordered references and generates its own audio. Write to the workflow instead of relying on a generic subject + action + camera + style formula.

The H3 prompt skeleton

Goal → ordered references → subject and identity anchors → chronological action beats → camera path → audio or dialogue direction → what must stay unchanged → final state. Work down that list and you will not forget the two lines most prompts miss: the audio and the ending.

Text to Video

You own everything, so spend the words on a timeline the clip can actually finish. Name the subject, order the beats, give the camera one continuous path, then describe the sound as separate layers — dialogue, ambience, music — rather than as a mood.

First / Last Frame

The frames already carry the appearance. Describe only the change between them: the motion added, the camera move, what must be preserved, and the state you land on. Re-describing the picture is the most common way to waste this route.

Multimodal Reference

Give every asset one job and address it by array position — Image 1 for the character, Image 2 for the location, Video 1 for the camera path, Audio 1 for the voice. Up to 9 images, 3 videos and 3 audio clips per request, and no more than 12 files in total; audio can never travel alone.

Editing an existing clip

Write two lists, not one paragraph: Change (target → result, one line each) and Preserve (anything adjacent that must survive). Send the source clip as Video 1 on the reference route, trimmed to the section you actually need.

Audio and dialogue

Quote the lines you want spoken, verbatim. Assign a voice reference one narrow job — timbre — and write accent, emotion and pacing into the prompt separately. Keep dialogue, ambience and music from competing for the same seconds.

Failure clinic

The seven ways an H3 prompt usually goes wrong, and what to change.

The clip ignores half the prompt
Likely causePrompt overload — too many subjects, styles and events competing inside one budget.
FixCut to one subject, one location and the number of beats the duration can hold. Move the rest into a second generation.
The camera drifts or fights itself
Likely causeTwo conflicting moves in one shot — an orbit and a push-in, or a handheld feel plus a locked-off frame.
FixGive the shot a single continuous camera path. If you need a second move, describe it as a cut.
The wrong reference drives the wrong thing
Likely causeAssets attached without stated roles, so the model guesses which plate owns the look.
FixAddress every asset by array position and give it exactly one job: Image 1 for identity, Video 1 for motion, Audio 1 for timbre.
The story is rushed and unreadable
Likely causeA full narrative squeezed into a short duration.
FixBudget roughly one beat per three to four seconds. A 10-second clip carries two or three beats, not seven.
Nothing really moves
Likely causeThe prompt describes appearance instead of change — common when a start image is attached.
FixWrite verbs. What moves, in what order, and what state does the shot end on?
An edit spreads beyond the thing you asked for
Likely causeNo preserved region declared, so unstated attributes are fair game.
FixList what must stay identical next to every change: faces, wardrobe details, background, framing.
The audio is a muddle
Likely causeDialogue, ambience and music all requested at once with no hierarchy.
FixSay which layer leads, keep the others quiet, and leave silence where the line lands.

MiniMax H3 prompt FAQ

Is MiniMax H3 the same thing as Hailuo 3?

Yes. MiniMax H3 is also known as Hailuo 3, Hailuo 3.0 or Hailuo 03. EvoLink uses MiniMax H3 as the product name and exposes it through three routes: text-to-video, image-to-video and reference-to-video.

Can I run these prompts without writing any code?

Yes. Use this prompt sends the text straight into the MiniMax H3 playground on the model page with the matching workflow already selected. You supply any reference assets there and generate in the browser.

How long can a MiniMax H3 clip be?

Duration is an integer from 4 to 15 seconds on all three routes, and the current routes output 2K. The prompts on this page were measured against their published clips, so the duration shown on each card is what that example actually runs.

How many reference assets can one prompt use?

The reference-to-video route accepts up to 9 images, 3 videos and 3 audio clips in a single request, capped at 12 files in total — so a full 9 + 3 + 3 set is rejected. At least one image or video is required; audio can never be the only reference type. Reference video and audio clips must each run 2 to 15 seconds.

Why do the prompts say Figure 1 or Video 1?

The reference route resolves assets by their position in the image_urls, video_urls and audio_urls arrays. Referring to Image 1 or Video 1 in the prompt is how you assign a job to a specific asset. The @image1 syntax is not part of this API contract.

Where do these prompts and clips come from?

Every card names its source. Official MiniMax samples are published with permission and their prompt text is the English edition of the original Chinese prompt. Community samples link back to the creator’s original post on X.

Can I use these prompts through the API?

Yes. The prompt text is identical whether you paste it into the playground or send it as the prompt field on the API. The MiniMax H3 API guide covers request structure, async tasks, callbacks and error handling.

Pick a prompt and generate it

Every prompt maps to a live MiniMax H3 route on EvoLink. Send one to the playground, or build the same request against the unified API.