GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

MiniMax H3 Prompt Guide & Video Examples

Browse 40 verified prompts with real MiniMax H3 outputs, input requirements, and reusable variables. Copy a template or send it directly to the EvoLink playground.

MiniMax H3 is also known as Hailuo 3, Hailuo 3.0, or Hailuo 03.

Generation mode

Use case

40 of 40 prompts

MiniMax H3 Brand & Product Ads prompts

Vertical eyewear commercial clip generated with MiniMax H3: two models in a seamless white studio wearing futuristic wraparound glassesMULTIMODAL REFERENCE
15s9:163 reference assets

Brand & Product Ads

Futuristic Eyewear Campaign

A three-image fashion commercial where each reference owns a different layer: the key visual, the models' faces, and the product design itself.

View details

Full prompt

Generate a vertical screen 9:16 high-end fashion glasses commercial, taking overall reference to the storyboard rhythm, editing speed, white studio texture and cool fashion atmosphere of the given video. The picture is a minimalist white booth, a seamless white background, a strong sense of high-end advertising, and a clean, simple, handsome, avant-garde, international first-line fashion blockbuster texture. Key visual character reference picture 1, two full-body female models, one black female model and one European and American model, maintain their high-end clothing, body posture, white studio light and shadow, fashion show temperament and overall cool attitude. Both of them wear futuristic high-end glasses. The design of the glasses refers to Figure 3, emphasizing the covered curved surface, sharp geometric cat-eye/goggle hybrid outline, mirror reflection, streamlined temples, and the texture of high-end fashion accessories. Please refer to Figure 2 for the appearance details of the two characters.

What you need to supply

  • Image 1: the key visual — full-body models, wardrobe, studio light and attitude.
  • Image 2: appearance detail for the two characters, so faces stay consistent through the cuts.
  • Image 3: the product, shot clearly enough that its silhouette and materials survive at speed.
  • A seamless studio background you are willing to keep for the whole clip.

Why it works

  • Three references with three explicit jobs is the pattern the reference route is built for; ambiguity is what makes multi-image prompts collapse.
  • The product is described by its geometry — wrap curvature, cat-eye/goggle hybrid outline, mirrored surface, streamlined temples — not just by its name.
  • A single-material white studio removes background variance, so the model spends its budget on the product and the performers.

Swap these out

  • product category
  • model casting
  • studio color
  • editing speed
  • wardrobe

Constraints

  • The published prompt also asks for the pacing of a given video. The released asset set for this case is three images, so treat that line as a style instruction, or attach your own clip as Video 1.
  • Up to 9 images, 3 videos and 3 audio clips per reference request, and no more than 12 files in total; audio can never be the only reference type.

Settings

Multimodal Reference · 15s · 9:16 · 3 reference assets

Product film generated with MiniMax H3: premium over-ear headphones rotating above a reflective pedestal in a black studioTEXT TO VIDEO
15s16:9No reference assets

Brand & Product Ads

Luxury Headphones Showcase

A 15-second product film written as four timed blocks, each with its own camera move and job — the most directly reusable template here for anyone with no reference assets.

View details

Full prompt

Create a 15-second luxury cinematic product showcase for premium wireless over-ear headphones. 0–4s: Begin with an extreme macro tracking shot moving across the soft memory-foam ear cushion, fine fabric texture, brushed-metal hinge and precision-machined controls. A narrow light band travels across the surface, revealing realistic materials against a deep black studio background. 4–8s: Pull back into a three-quarter hero view. The headphones rotate slowly above a glossy reflective pedestal. The ear cups pivot naturally while the adjustable headband extends slightly, demonstrating flexible construction and comfort. Maintain exact symmetry, stable geometry and consistent proportions. 8–12s: Transition into an elegant exploded-view reveal. The ear cushion, acoustic driver, internal sound chamber, control ring and outer shell separate smoothly in perfect alignment. Subtle luminous sound waves pulse outward from the driver while the camera performs a restrained side orbit. 12–15s: Every component reconnects seamlessly. The headphones settle into a centered front-facing hero composition as soft rim lighting defines the silhouette. Complete a gentle dolly-in toward the ear cups. Premium technology-commercial finish, controlled reflections, realistic shadows, shallow depth of field, crisp surface detail, stable product shape, no hands, no distortion, no onscreen text.

What you need to supply

  • Nothing to upload — text-to-video takes the prompt alone.
  • A product you can describe by material and mechanism, not just by name.

Why it works

  • Time is split into 0–4s, 4–8s, 8–12s and 12–15s, so the model gets a shot list instead of a wish list.
  • Each block moves the camera differently — macro track, pull back, side orbit, dolly in — which is what makes the clip read as edited rather than drifting.
  • The closing line is a list of prohibitions (no hands, no distortion, no onscreen text) that cover the three ways product renders usually fail.

Swap these out

  • product
  • materials and finish
  • timing of each block
  • background and lighting
  • whether the exploded view appears

Constraints

  • Duration is an integer from 4 to 15 seconds and is set in the request; the timed blocks in the prompt must add up to it.
  • Timed blocks are direction, not a hard timeline. Keep them to four or fewer for a 15-second clip.
  • This clip was published by its author at 720p; the H3 routes on EvoLink output 2K.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 beverage ad: a heat-exhausted walker sips a chilled juice and the street blooms into lush greenery around the bottleTEXT TO VIDEO
15s16:9No reference assets

Brand & Product Ads

Summer Heat Beverage Ad

A classic problem-relief beverage spot: heat-exhausted subject, one sip, and the whole environment transforms — ending on a droplet-covered product hero shot.

View details

Full prompt

A young person walks under the blazing summer sun, looking exhausted and sweating heavily. The road shimmers with heat waves, and everything appears dry and dull. Suddenly, they grab a chilled bottle of premium fruit juice from a cooler and take a refreshing sip. Instantly, the environment transforms—lush green trees bloom, vibrant flowers appear, a cool breeze flows, water splashes through the air, and glowing particles surround the scene. Ice cubes and fresh fruit slices (orange, mango, or according to the flavour) swirl around the bottle in cinematic slow motion. End with a stunning close-up of the juice bottle covered in cold water droplets against a bright, refreshing background. Ultra-realistic, premium commercial, 4K, cinematic lighting, high-detail, smooth camera movements, vibrant colours, luxury beverage advertisement. Tagline ideas: Beat the Heat. Taste the Freshness. Every Sip Brings Life. Refresh Your Day, Naturally. Stay Cool. Stay Fresh.

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • The ad is built as a before/after state change — dry heat versus lush freshness — which gives the model one clear transformation to execute instead of a list of moods.
  • The product enters late and ends the clip as a close-up hero shot, the standard beverage-ad beat order the model can pattern-match.
  • Flavor elements (ice, fruit slices) are staged as physical objects swirling in slow motion, not abstract adjectives.

Swap these out

  • beverage type and flavor cues
  • the exhausted-subject setting
  • transformation environment
  • tagline text

Constraints

  • The trailing tagline list is copy inspiration, not burned-in text — move one tagline into the visual direction if you want it rendered on screen.
  • One transformation is the budget; adding a second product or scene change will crowd the 15-second window.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 UGC & Creator Ads prompts

MiniMax H3 stream opening from five references: an anime-styled streamer reads her chat rail on a Twitch-like layout with a LIVE badge and follower bannerMULTIMODAL REFERENCE
15s16:95 reference assets

UGC & Creator Ads

VTuber Stream Opening

A five-asset production: four images split identity, platform chrome, room and opening card into separate jobs, while an audio reference drives real lip-synced streamer performance — including its silent lead-in.

View details

Full prompt

Use @Image4as the opening card only: circular Luna avatar, black background, cream “LUNALIVE”, rose “STREAM STARTING”, warm circles, tiny mint accent. Static, chime. Use @Image1 for Luna’s identity only: same face, long center-parted black hair, blue-gray eyes, pink anime hoodie, white headphones around neck, delicate necklace, pale nails. Do not copy the drink, pose, or background from @Image1 . Use @Image2 only for Twitch-like platform chrome: dark top bar with “LUNALIVE”, red “LIVE” badge, “2.4K viewers”, right “STREAM CHAT” rail, bottom title area, rounded “FOLLOW” and pink “SUBSCRIBE” buttons. Do not copy Luna, pose, drink, or room from @Image2 Use @Image3 only for the cozy room behind Luna inside the video area: desk, monitor, white PC, plush shelves, curtain fairy lights, soft pink/purple light. No empty-room showcase shot. Use @Audio1as Luna’s actual vocal performance and behavior reference. @Audio1has a silent lead-in: 0.0–2.4s must be treated as no speech. Preserve her voice identity, cadence, tone, breaths, pauses, emphasis, warmth, and streamer mannerisms. Do not replace the voice, do not generate a different influencer voice, and do not add extra spoken lines beyond @Audio1 Lip sync, mouth shapes, jaw movement, smiles, glances, nods, and small hand movements must follow the audio waveform after 2.4s. Create a 15-second 16:9 Twitch-like stream opening. Important performance direction: Luna is reading chat, not delivering a camera monologue. Place Luna slightly left of center in the video area with the right “STREAM CHAT” rail clearly visible. Whenever she speaks, her eyes angle screen-right toward the chat rail as if she is reading the messages out loud. She returns to camera only for brief reactions. Add constant small movement: eye darts to chat, eyebrow lifts, tiny nods, head tilts, shoulders shifting, one subtle hand gesture near the desk. No stiff talking-head pose. One cursor only, no cursor trail, no duplicate panels, no duplicate buttons. Render only large UI text cleanly. Chat feels alive with typing dots and soft short blurred lines, but only these chat lines are readable: “hi Luna”, “welcome back”, “so cozy”, “gugugaga?”, “RAID INCOMING!”. Do not invent usernames. Chat pops stay quieter than Luna’s voice. [0–2 seconds] Open on @Image4 “LUNALIVE / STREAM STARTING”. Absolute no-speech zone: @Audio1is silent here, Luna is not shown, no mouth movement, no voice on the card. One soft chime only. No movement. [2–7.4 seconds] Hard cut to Luna live at 2.0s, slightly left of center in the cozy room from @Image3 with @Image2 chrome active: “LUNALIVE”, red “LIVE”, “2.4K viewers”, right “STREAM CHAT” rail, bottom title “COZY NEON” and “Just Chatting”. She settles for a beat, eyes already moving toward the chat rail. At about 2.4s when speech begins in @Audio1 , match lip sync exactly while she reads toward the chat rail, not into camera. Chat shows typing dots and soft blurred lines. [7.4–7.9 seconds] First audio pause = chat beat. Typing dots, then readable messages pop in: “hi Luna”, “welcome back”. Luna’s eyes track the new messages on the right rail; small nod and smile follow the audio pause. [7.9–12.5 seconds] Continue matching @Audio1Keep her gaze mostly on the chat rail while speaking, like she is reading and reacting. Add one more readable message: “so cozy”. During any softer phrase, she leans slightly forward as if reading; during brighter phrases, eyebrows lift and shoulders react. No frozen face. [12.5–13.9 seconds] Bigger audio pause = bigger chat beat. “gugugaga?” appears, then “RAID INCOMING!”, and a clean “NEW FOLLOWER” banner slides in with a gentle pop. Luna reads the raid message from the chat rail, then reacts brighter as the audio resumes. [13.9–14.8 seconds] Finish @Audio1 with accurate lip sync. If the audio winds down, Luna stops talking, gives a small wave toward chat, and settles into a warm listening pose. End on the stable live frame: Luna slightly left of center, eyes toward the right chat rail, red “LIVE”, “2.4K viewers”, no end card. Audio mix: 0–2s card is silent except one soft chime. @Audio1voice begins only after the cut to Luna and remains primary. Tiny chat pops under the voice. Warm low room tone. No crowd noise, no music lyrics, no rain, no traffic.

What you need to supply

  • Image 1: the streamer’s identity only — face, hair, hoodie; pose and background are explicitly not copied.
  • Image 2: the stream-platform chrome — chat rail, LIVE badge, buttons.
  • Image 3: the cozy room behind her.
  • Image 4: the static "stream starting" opening card.
  • Audio 1: her actual vocal take; note the 0–2.4s silence is scripted as a no-speech zone.

Why it works

  • Each reference gets one job and two explicit "do not copy" exclusions, the cleanest per-asset role assignment in the library.
  • The performance note — "she is reading chat, not delivering a monologue" — redirects eye-line and micro-gestures, which is what makes the clip feel live.
  • Readable chat messages are whitelisted to five exact strings and everything else stays blurred, so UI text renders clean.

Swap these out

  • streamer identity plate
  • platform chrome styling
  • the whitelisted chat lines
  • audio take and its pauses

Constraints

  • The original post writes @Image1…@Audio1; on EvoLink address assets by array position and keep at least one image alongside the audio — audio can never travel alone.
  • The route accepts up to 9 images, 3 videos and 3 audio clips, capped at 12 files in total.

Settings

Multimodal Reference · 15s · 16:9 · 5 reference assets

MiniMax H3 UGC ad: a woman films herself applying hair serum from a dropper bottle in a bright bedroom, smartphone-style framingMULTIMODAL REFERENCE
15s3:41 reference asset

UGC & Creator Ads

UGC Hair Serum Ad

A four-scene TikTok-style serum testimonial with one product reference: selfie intro, application close-up, mirror result, counter-top product shot — imperfection is the styling.

View details

Full prompt

Create a 15-second authentic UGC-style hair growth serum ad using the provided product image as the exact product reference. SCENE 1 — 0–3s A young woman films herself in a bright bedroom using a smartphone front camera. Natural lighting, handheld movement, casual appearance. She looks at the camera and says: “I’ve been trying this hair growth serum lately…” SCENE 2 — 3–7s Cut to a close handheld shot of the woman holding the serum bottle. She removes the dropper, applies a few drops directly to her scalp, and gently massages it in. Keep the movement natural and slightly imperfect like real UGC content. SCENE 3 — 7–11s Mirror selfie shot. She runs her fingers through her hair, showing healthy-looking, fuller hair while casually talking to the camera: “And honestly, I love how easy it is to add to my routine.” SCENE 4 — 11–15s Close-up product shot on her bathroom counter. She picks up the bottle and smiles toward the camera. End with natural on-screen text: “Simple hair care. Every day.” Style: authentic TikTok/Reels UGC, smartphone camera, realistic skin texture, natural expressions, subtle handheld motion, imperfect framing, casual home environment, soft daylight, realistic audio, no cinematic commercial look, no excessive beauty filters. Preserve the exact product packaging, label, bottle shape, and branding from the reference image.

What you need to supply

  • Image 1: your product shot — packaging, label and bottle shape are locked from this plate across all four scenes.

Why it works

  • The scene order mirrors organic creator content (hook → demo → result → product), so the ad reads native in a feed.
  • "Slightly imperfect like real UGC" plus the no-cinematic negative list is what defeats the polished-commercial look buyers scroll past.
  • Spoken lines are short, quoted and conversational — native audio delivers them as natural testimonial speech.

Swap these out

  • product plate
  • the two spoken lines
  • bedroom/bathroom settings
  • closing on-screen text

Constraints

  • Product identity lives entirely in the reference image; describing the label in text as well invites conflict.
  • The published clip measures 2:3, which the API does not accept — request 3:4 or 9:16 for a feed-native vertical.

Settings

Multimodal Reference · 15s · 3:4 · 1 reference asset

MiniMax H3 food vlog: a young woman lifts a gourmet cheeseburger toward her phone camera mid-laugh between quick jump cutsMULTIMODAL REFERENCE
15s21:91 reference asset

UGC & Creator Ads

Burger Night UGC Vlog

An eight-shot food-vlog template with a locked burger reference: selfie hook, box open, zoom punch, cheese pull, bite reaction — hard jump cuts doing the pacing work.

View details

Full prompt

Duration: 15 seconds | Aspect Ratio: 16:9 | Style: Authentic UGC / iPhone selfie-vlog, handheld, natural light, TikTok/Reels aesthetic. Product Reference: Use the uploaded gourmet burger image as the only product reference. Preserve the bun shape, patty thickness, cheese melt, lettuce, tomato, sauces, and proportions exactly in every shot. Character Description Name: Hana A young Japanese woman in her early 20s with natural beauty, long dark hair in a loose ponytail, oversized cream sweatshirt, minimal makeup, bright smile, friendly lifestyle-vlogger personality. Shot Breakdown SHOT 1 (0–2s) — Selfie showing the burger box. Dialogue: "Burger night!" SHOT 2 (2–4s) — Opens the box. SHOT 3 (4–6s) — Quick zoom on the burger. SHOT 4 (6–8s) — Hands lifting the burger with cheese stretching naturally. SHOT 5 (8–10s) — Bite reaction. Dialogue: "Okay... that's incredible." SHOT 6 (10–12s) — Casual close-up b-roll while reaching for fries. SHOT 7 (12–14s) — Toasting the burger toward the camera. Dialogue: "You need this." SHOT 8 (14–15s) — Freeze frame with overlay: "burger cravings = solved 🍔" Look & Feel Warm apartment lighting, genuine phone footage, slight grain, natural autofocus breathing, handheld imperfections, fast jump cuts. Negative Prompt cinematic grading, commercial production, CGI burger, fake cheese, distorted hands, warped food, perfect stabilization, studio lighting, text glitches, logo distortion.

What you need to supply

  • Image 1: the food hero shot — bun, patty, cheese melt and proportions stay identical in every cut.

Why it works

  • Eight micro-shots of ~2 seconds each match how real food creators actually cut, so the energy is native to the format.
  • The named character ("Hana") with two short quoted lines gives the model a stable face and voice without overspecifying.
  • Food physics gets its own instruction (cheese stretching naturally) — the money shot is scripted, not hoped for.

Swap these out

  • the food item and its plate
  • character styling
  • the two dialogue lines
  • freeze-frame end text

Constraints

  • Keep the food description in the reference image only; the negative list (no CGI burger, no fake cheese) guards the realism.

Settings

Multimodal Reference · 15s · 21:9 · 1 reference asset

MiniMax H3 Typography & Text Motion prompts

MiniMax H3 fashion film: watercolor ribbons pull the camera around a woman in a white high-neck dress as the phrase LET SILENCE BLOOM forms in physical lettersTEXT TO VIDEO
15s16:9No reference assets

Typography & Text Motion

Watercolor Couture One-Take

An avant-garde fashion film where watercolor ribbons pull a single continuous camera through space, typography exists as physical objects, and every move lands on the music.

View details

Full prompt

## Concept AQUARELLE No.7 — an avant-garde haute couture film where watercolor becomes a living medium. 15-second cinematic sequence. The rhythm controls the visual world: 0–4s: restrained silence and negative space 4–8s: gradual density buildup 8s: drop moment, expanding into fluid long-form motion 12–15s: transition into a final fashion poster composition ## Character Identity Lock Maintain the exact same female character throughout the entire video. Identity: - Young woman - Long black hair - Calm and refined facial features - White high-neck pigment dress - Black wide belt - Pigment heels Strict consistency: - Same face - Same hairstyle - Same age - Same body proportions - Same garment structure Any new colors must only appear through watercolor gradually absorbing into the fabric. Core concept: Her fingertips can extract transparent watercolor ribbons from the air. These watercolor ribbons can: - Pull the camera through space - Transform the environment - Shape physical typography - Interact with depth and materials The world follows real cinematic physics: water, paper, fabric, light, shadows, and depth of field must feel physically believable. ## Camera Direction Strict one-take shot. No cuts. No teleportation. No hidden transitions. Camera journey: Macro shot of a floating water droplet → watercolor ribbon emerges → camera pulls back to reveal the woman → camera circles around her → enters a paper art gallery → passes through dimensional typography → rises into a final overhead fashion poster. ## Typography Only allow these words: LET SILENCE BLOOM AQUARELLE No.7 WEAR THE UNSEEN Typography is not a flat overlay. Letters must have: - Physical depth - Wet watercolor reflections - Paper fiber edges - Shadows - Occlusion - Material interaction ## Visual Style Avant-garde fashion editorial. Inspired by: - Museum catalog composition - Handmade ivory paper - Sculptural negative space - Elegant Didone serif typography - Translucent watercolor calligraphy Color palette: - Pale cyan - Crimson lake - Smoky violet - Ink black ## Motion Language Every movement follows the music: Kick: Camera movement and paper folding. Wooden snare: Paper structures physically fold and transform. Sub-bass: Changes the perception of spatial scale. ## Storyboard 30 beats, 0.5 seconds each. 01 0.0–0.5 Macro shot: A transparent water droplet floats in the air, reflecting a blurred silhouette of a black-haired woman. 02 0.5–1.0 The droplet stretches with the breath-like vocal, becoming a pale cyan watercolor thread. 03 1.0–1.5 Camera travels backward along the thread as paper fibers slowly emerge into focus. 04 1.5–2.0 The thread wraps around the lens. Focus shifts to her raised fingertip. 05 2.0–3.0 Camera continues pulling back, revealing her face and white high-neck dress. 06 3.0–4.0 She moves her wrist. The watercolor thread guides a smooth camera arc. 07 4.0–8.0 Additional watercolor colors emerge from her movement. Paper folds, typography begins forming, and the phrase "LET SILENCE BLOOM" appears as a physical object. 08 8.0–12.0 The camera passes through a transparent paper flower structure. The environment expands into an endless ivory paper gallery. Her dress absorbs watercolor naturally. The ribbons create sculptural forms around her body. 09 12.0–15.0 Camera cranes upward. The composition transforms into a luxury fashion advertisement poster. Typography appears: AQUARELLE No.7 WEAR THE UNSEEN Final frame: A museum-level fashion editorial poster. The woman remains centered, calm, and elegant. A final watercolor droplet remains suspended in the air. ## Negative Prompt No: - Cuts - Scene changes - Identity change - Face swap - Extra limbs - Deformed hands - Random costume changes - Explosive paint effects without physical cause - Incorrect typography - Chinese characters - Extra subtitles - Extra logos - Watermarks

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • The camera journey is written as one continuous chain (droplet → ribbon → reveal → gallery → overhead poster) with "no cuts" stated as a hard rule.
  • Typography is whitelisted — only three phrases may appear — and given material properties (paper fiber edges, wet reflections, occlusion), which is why the words render cleanly.
  • A 30-beat storyboard at 0.5s per beat maps sound to space: kick folds paper, snare transforms structures, sub-bass changes scale.

Swap these out

  • the three allowed phrases
  • pigment palette
  • garment structure
  • gallery environment

Constraints

  • The phrase whitelist is the typography quality mechanism — adding more text reintroduces the spelling drift the prompt is built to avoid.
  • One-take prompts fail loudly: if any beat implies a cut, the whole spatial chain breaks.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 educational animation: a rounded letter A inflates into a smiling red apple beside the word APPLE in a soft pastel sceneTEXT TO VIDEO
15s16:9No reference assets

Typography & Text Motion

A-B-C-D Learning Animation

A children’s phonics animation with a fixed teaching loop — letter, sound, object, action, word — where each letter physically morphs into its object and the narration is scripted per second.

View details

Full prompt

Create a 15-second animated educational video that teaches young children the letters A, B, C, and D. The learning pattern for every letter must be: LETTER → SOUND → OBJECT → PLAYFUL ACTION → OBJECT NAME Target audience: children ages 3 to 6. Visual style: Use adorable rounded 3D characters, soft pastel colors, gentle facial expressions, and simple recognizable objects. Combine this with a premium minimalist technology aesthetic featuring clean white space, elegant composition, soft studio lighting, subtle reflections, smooth gradients, rounded geometry, crisp typography, and extremely polished transitions. The animation should feel playful and child-friendly while remaining calm, uncluttered, and beautifully designed. Use a clean off-white background with a different soft color glow behind each letter. 0:00–0:01 | Introduction A small smiling star mascot bounces into the center of the screen. Colorful letters briefly float around it. Display the text: “Let’s learn!” The mascot taps the screen, creating a soft ripple that reveals the first letter. 0:01–0:04 | A is for Apple Show a large uppercase “A” and smaller lowercase “a” beside it. Use thick, rounded, highly readable typography. The narrator says: “A. A says ah. A is for Apple.” The uppercase A gently inflates and transforms into a shiny red apple. Its top point becomes the apple stem, and a small green leaf unfolds from the side. The apple gains a cute smiling face and performs one soft bounce. Display the word: “APPLE” Highlight the first letter A in red. Add a soft pop and a tiny crunchy sound. 0:04–0:07 | B is for Ball The apple rolls across the screen and leaves behind a curved red trail. The trail loops twice and forms a large uppercase “B,” with a lowercase “b” appearing beside it. The narrator says: “B. B says buh. B is for Ball.” The two rounded sections of the B expand and merge into a colorful striped ball. The ball bounces twice with playful squash-and-stretch animation. Display the word: “BALL” Highlight the first letter B in blue. Synchronize each bounce with a soft musical note. 0:07–0:10 | C is for Cat On its final bounce, the ball stretches into a curved shape and becomes a large uppercase “C.” A lowercase “c” slides gently into place beside it. The narrator says: “C. C says kuh. C is for Cat.” The C rotates and becomes the curled tail of a cute orange cat. The rest of the cat forms from soft rounded shapes. The cat stretches, blinks, and gives one gentle wave with its paw. Display the word: “CAT” Highlight the first letter C in orange. Add a quiet and friendly “meow.” 0:10–0:13 | D is for Duck The cat’s tail uncurls and transforms into the curved side of a large uppercase “D.” A lowercase “d” pops up beside it. The narrator says: “D. D says duh. D is for Duck.” The straight line of the D becomes the duck’s neck. The curved section becomes its round yellow body. A small orange beak and two tiny wings pop into place. The duck waddles forward, flaps its wings, and gives one cheerful quack. Display the word: “DUCK” Highlight the first letter D in yellow. Add tiny water ripples beneath its feet. 0:13–0:15 | Recap The apple, ball, cat, and duck slide into four clean rounded tiles. Place their letters above them: “A B C D” The mascot returns and points to each object as they bounce once in sequence. Narrator: “A, B, C, D. Great job!” Finish with the text: “Great job!” Use a small sparkle animation and a warm musical chime. Animation requirements: Keep each letter fully visible for a moment before it transforms. Show uppercase and lowercase versions clearly. Make every object instantly recognizable. Use smooth shape morphing so children can visually understand how the letter becomes the object. Maintain stable spelling, clean letterforms, accurate object shapes, and consistent character design. Use gentle squash-and-stretch, soft motion blur, subtle shadows, polished lighting, and precisely synchronized sound effects. Avoid fast camera movement, cluttered backgrounds, harsh colors, tiny text, warped letters, random symbols, duplicated objects, scary expressions, or overly complex transformations. The final video should feel cute, educational, memorable, calming, and exceptionally polished.

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • The LETTER → SOUND → OBJECT → ACTION → NAME loop repeats four times with identical structure, and repetition is the strongest stabilizer H3 has.
  • Every morph is a geometric explanation (the A’s point becomes the apple stem, the B’s bowls become the ball) so the letterforms survive the transformation.
  • Narration lines are quoted exactly with per-segment timestamps, driving native audio and keeping caption text in sync.

Swap these out

  • the four letters and objects
  • mascot design
  • color-glow per letter
  • narration voice

Constraints

  • Spelling stability depends on the "keep each letter fully visible before it transforms" rule — cutting it reintroduces warped glyphs.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 kinetic typography: the words Every great change emerge from darkness in thin serif letters with drifting golden particlesTEXT TO VIDEO
15s16:9No reference assets

Typography & Text Motion

Kinetic Quote Typography

Pure kinetic typography: one quote revealed phrase by phrase, with each phrase changing the scene’s atmosphere, motion language and palette until the full line locks on an off-white end card.

View details

Full prompt

Create a 15-second cinematic text-animation video built around the quote: “Every great change begins quietly, grows through courage, and becomes impossible to ignore.” The quote should appear gradually as a visual story. Each new phrase must transform the design, atmosphere, movement, and emotional intensity of the scene. Use elegant typography, accurate spelling, cinematic lighting, smooth transitions, and perfectly readable text. 0:00–0:03 | “Every great change” Begin with a completely black screen. A tiny point of warm light slowly appears in the center, like the first spark of an idea. The words “Every great change” emerge softly from the darkness, one word at a time. Use thin, elegant serif typography with wide letter spacing. “Every” fades in gently. “Great” grows slightly larger. “Change” forms from small drifting particles that gather into solid letters. Keep the scene quiet, minimal, and mysterious. 0:03–0:06 | “begins quietly,” The camera slowly moves closer to the text. The previous words shrink and reposition toward the upper-left corner as the phrase “begins quietly,” appears in delicate lowercase letters. Animate the phrase as though it is being written by an invisible hand. Each letter should create a subtle ripple in the darkness. Introduce faint textures, soft shadows, floating dust, and gentle light rays. The comma should appear last and create a small circular pulse. 0:06–0:09 | “grows through courage,” The pulse expands and transforms the scene from darkness into a rich sunrise gradient with deep orange, red, and golden tones. The words “grows through courage” rise upward from the bottom of the frame. Animate “grows” by gradually increasing its size and weight. Animate “through” along a curved path. Animate “courage” in bold uppercase letters that push through a translucent barrier, causing it to crack into geometric fragments. The movement should feel powerful but controlled. 0:09–0:12 | “and becomes” The fragments rotate in slow motion and reorganize into a clean editorial grid. The phrase “AND BECOMES” appears across the frame in condensed sans-serif typography. Animate the letters with fast tracking changes, vertical stretching, masking, and perspective movement. The camera accelerates forward through the center of the word “BECOMES.” The sound and visual energy should steadily build. 0:12–0:14 | “impossible to ignore.” Reveal a vast bright space filled with light, moving shapes, and large-scale typography. The words “IMPOSSIBLE TO IGNORE” appear one after another. “IMPOSSIBLE” expands beyond the edges of the screen. “TO” remains small and perfectly centered. “IGNORE” slams into place with strong visual impact, briefly shaking the surrounding grid and shapes. Use bold contrast, dramatic scale, sharp shadows, and synchronized motion. 0:14–0:15 | Final quote All movement stops instantly. The complete quote appears centered on a clean off-white background: “Every great change begins quietly, grows through courage, and becomes impossible to ignore.” Use refined black typography with “change,” “courage,” and “impossible” highlighted in deep red. Hold the final composition clearly for the last second. Maintain one continuous visual journey from darkness to light, silence to impact, and simplicity to complexity. Keep every phrase connected through visual transformations rather than hard cuts. Use realistic motion blur, precise kerning, clean masks, stable letterforms, smooth camera movement, subtle film grain, cinematic sound design, rising ambient music, soft particles, controlled color transitions, and a final deep impact sound. Avoid misspelled words, warped letters, duplicated characters, unreadable text, random symbols, excessive flickering, chaotic layouts, inconsistent fonts.

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • Each phrase owns a 3-second block with its own animation verb (emerge, handwrite, rise, slam), so the energy build is structural, not adjectival.
  • The emotional arc is mapped to design language — darkness to sunrise gradient to editorial grid — giving the model a palette script, not just words.
  • A hard full stop ("all movement stops instantly") plus a held end card guarantees a readable final frame for thumbnails.

Swap these out

  • the quote and its phrase splits
  • highlight words and accent color
  • per-phrase animation verbs
  • end-card styling

Constraints

  • Keep phrases under ~6 words each; H3 renders short display text far more reliably than sentence-length lines.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 Short Drama & Narrative prompts

Vertical short-drama trailer clip generated with MiniMax H3: a vampire lead and a human heroine in a candlelit castle interiorMULTIMODAL REFERENCE
15s9:162 reference assets

Short Drama & Narrative

Vampire Romance Short Drama

A vertical short-drama hook that pins both leads to one reference image and the location to another, then spends the rest of the prompt on relationship beats and shot sizes rather than plot.

View details

Full prompt

Generate a 15-second, 9:16 vertical trailer segment for an international live-action vampire romance short drama. Use Figure 1 as the appearance reference for the male and female leads, and Figure 2 as the scene reference. Keep both leads' identities consistent, with a realistic live-action look and premium short-drama production quality. Story: an innocent human heroine accidentally enters a forbidden area of an old castle and awakens a sleeping aristocratic vampire. He discovers that she carries an aura connected to an ancient war, which sparks a powerful urge to control her and a dangerous fascination with her. She fears him but does not completely submit, resisting his pressure. Overall style: an international ReelShort / DramaBox vampire-romance trailer. Dark romance, dangerous attraction, fate, intense control, brooding oppression, and a striking reversal. Keep the visuals premium, restrained, and tightly paced, like the opening 15-second hook of a hit short drama. No gore, cheap horror, Halloween aesthetic, or modern street feel. Format: 9:16 vertical composition for TikTok / ReelShort / DramaBox. Use primarily medium close-ups, close-ups, and extreme close-ups, emphasizing faces, eye contact, pressure, and relationship tension within the vertical frame.

What you need to supply

  • Image 1: both leads together, so the model reads them as one consistent casting.
  • Image 2: the location plate — the castle interior, its light and its materials.
  • A one-line premise and a stated emotional reversal; a 15-second hook can carry one turn, not a full episode.

Why it works

  • Naming the genre reference (ReelShort / DramaBox trailer) transfers pacing, grade and framing conventions in a few words.
  • It fixes the shot vocabulary — medium close-up, close-up, extreme close-up — which is what makes a vertical frame read as premium instead of cramped.
  • The exclusion list (no gore, no cheap horror, no Halloween look, no modern street feel) removes the four ways this genre usually degrades.

Swap these out

  • lead appearance references
  • location
  • premise and reversal
  • target platform look
  • shot-size mix

Constraints

  • Vertical output is requested through aspect_ratio on the reference route, not by describing '9:16' in the prompt alone.
  • Identity holds far better when the leads arrive as one image than as two separately cropped portraits.

Settings

Multimodal Reference · 15s · 9:16 · 2 reference assets

Visual-novel interface transition generated with MiniMax H3 between a fixed opening frame and a fixed final frameFIRST / LAST FRAME
15s16:92 reference assets

Short Drama & Narrative

Otome Visual Novel Transition

A true first-and-last-frame prompt: two images fix both ends of the shot, and the text only has to describe the journey between them.

View details

Full prompt

Use the first image as the opening frame and the second image as the exact final frame to generate an otome visual-novel interface transition. Overall feel: a premium Chinese otome romance-interaction interface capturing an intimate moment before and after a performance. Transition naturally from "choose to watch his performance" to "Han Xu is drawn in by the heroine's words and reacts with intrigued interest." UI text, choices, and dialogue boxes should appear with refined otome-game presentation. Keep the transition silky smooth and the emotion suggestive yet restrained.

What you need to supply

  • Image 1: the opening frame, complete with its UI state.
  • Image 2: the exact closing frame you want to land on.
  • A one-line description of the emotional change between the two states.

Why it works

  • Both endpoints are locked, so the model solves interpolation instead of composition — the most reliable way to get a predictable shot.
  • The prompt names the emotional transition ('choose to watch his performance' → 'drawn in and intrigued') rather than listing frames, so the performance carries the cut.
  • It states the interface elements should animate in the genre's own idiom, which keeps the UI from being redrawn.

Swap these out

  • opening and closing frames
  • emotional arc between them
  • UI presentation style
  • transition speed

Constraints

  • Send both frames as image_start and image_end on the image-to-video route; the reference route does not accept these fields at all.
  • The closer the two frames are in framing and lighting, the smoother the interpolation. Two unrelated compositions produce a cut, not a transition.

Settings

First / Last Frame · 15s · 16:9 · 2 reference assets

MiniMax H3 period sequence: soldiers in 1940s uniforms take cover behind vintage cars on a smoke-filled rural American street, shot like archival filmTEXT TO VIDEO
15s16:9No reference assets

Short Drama & Narrative

1940s War Newsreel Realism

A period-accurate 1947 newsreel simulation whose realism comes from three stacked systems: era-consistent set dressing, a physics contract, and a negative list that bans every modern object.

View details

Full prompt

Create a 15-second ultra-photorealistic live-action war sequence set in the United States in 1947, designed to look like authentic historical footage captured on a 1940s film camera. The entire scene must feel grounded, documentary-like, raw, and physically realistic. Environment: A rural American town in 1947 with wooden houses, old brick buildings, telephone poles, dirt roads, vintage American cars from the 1940s, wooden fences, farmland, and period-accurate street details. Overcast afternoon light, light fog, drifting smoke, dust in the air, damaged buildings, scattered debris, and a tense wartime atmosphere. Characters: American soldiers wearing historically accurate late-1940s military uniforms, helmets, boots, and equipment. Civilians wear authentic 1940s American clothing. Natural faces, realistic skin texture, sweat, dirt, fatigue, and believable body movements. 0–3s — Establishing Shot: Wide handheld shot of a quiet rural American street suddenly filled with smoke and confusion. Vintage 1940s vehicles are parked along the road while soldiers move quickly between wooden buildings. Civilians rush toward safer areas. 3–6s — Tension: Camera moves through the street at shoulder height, following several soldiers as distant gunfire is heard. They immediately react and take cover behind a vintage vehicle and a brick wall. Their movements are cautious and realistic. 6–10s — Combat: Fast handheld tracking shot as the soldiers move between cover while distant gunfire impacts the environment. Small pieces of wood, dust, and debris fall naturally from nearby impacts. Weapon recoil, movement, and body weight must be physically accurate. Keep the violence realistic and restrained. 10–13s — Human Moment: Camera briefly focuses on a soldier helping an injured civilian move behind cover. Their breathing, facial expressions, body language, and movement should feel natural and unscripted. 13–15s — Final Shot: Camera pulls back into a wide shot of the American town as smoke slowly moves through the street. Soldiers remain behind cover while vintage vehicles and damaged buildings fill the background. The scene ends with an authentic, tense 1940s documentary feeling. Visual Style: Ultra-photorealistic live-action, authentic 1940s American environment, vintage 35mm film texture, subtle film grain, natural imperfections, realistic exposure, handheld documentary cinematography, muted historical color palette, realistic smoke and dust, natural shadows, accurate depth of field. Physics: Strictly obey real-world gravity, momentum, inertia, friction, recoil, weight, collision physics, and human biomechanics. No exaggerated explosions, impossible movements, superhero behavior, or choreographed-looking combat. Negative Prompt: modern buildings, modern cars, smartphones, modern clothing, modern weapons, futuristic technology, CGI appearance, video-game graphics, fantasy, superhero action, excessive explosions, excessive blood, gore, impossible physics, unrealistic recoil, slow-motion physics, distorted faces, extra limbs, floating objects, plastic skin, artificial-looking environments.

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • Period accuracy is enforced twice — positively (1940s cars, uniforms, telephone poles) and negatively (no smartphones, no modern buildings) — closing both failure directions.
  • The physics paragraph ("strictly obey gravity, momentum, recoil…") reads like a render spec and visibly suppresses superhero-style motion.
  • A quiet human beat (a soldier helping a civilian) is scheduled at 10–13s, giving the sequence documentary credibility rather than nonstop action.

Swap these out

  • era and location
  • the human moment
  • film-stock look
  • intensity of the combat beats

Constraints

  • The restraint clauses ("violence realistic and restrained", no gore) are part of why this output is usable — keep them when adapting.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 Character & Motion prompts

Claymation clip generated with MiniMax H3: a clay fox leaping across a lava canyon as the camera sweeps beneath itFIRST / LAST FRAME
10s16:91 reference asset

Character & Motion

Clay Fox Canyon Leap

A single start frame animated into one decisive action, with the camera move described as precisely as the action itself.

View details

Full prompt

Claymation style. A sprinting fox reaches the edge of a cliff and launches without hesitation, making a dramatically tense, heroic slow-motion leap across a vast lava canyon. While the fox is airborne, the camera rushes at high speed beneath its belly in a sweeping dynamic move, fully revealing the terrifying depth of the chasm and the fox's clay body at maximum extension in midair.

What you need to supply

  • A start image that already carries the style — here the claymation fox and its material.
  • One action you want to happen. Not three.

Why it works

  • The prompt describes change, not the picture. The start frame already holds the appearance, so every word buys motion.
  • The camera has its own instruction — a high-speed sweep beneath the fox's belly — which turns a jump into a shot.
  • Naming the peak moment ('maximum extension in midair') gives the model a target pose to build the timing around.

Swap these out

  • character and material style
  • environment and hazard
  • camera path
  • slow-motion emphasis

Constraints

  • The image-to-video route derives the output ratio from the input image and rejects aspect_ratio outright.
  • Do not re-describe what the start frame already shows; repeating static appearance is the most common way to waste an image-to-video prompt.

Settings

First / Last Frame · 10s · 16:9 · 1 reference asset

Motion-transfer clip generated with MiniMax H3: two referenced characters performing street dance copied from a reference videoMULTIMODAL REFERENCE
10s16:93 reference assets

Character & Motion

Street Dance Motion Transfer

Twenty-odd words that move choreography from a reference clip onto two characters supplied as images — the clearest demonstration of what ordered references buy you.

View details

Full prompt

Have the characters perform street dance following the movements in Video 1. Use Figure 1 and Figure 2 as the character references.

What you need to supply

  • Image 1 and Image 2: the two characters, one clean full-body reference each.
  • Video 1: the movement to copy, 2–15 seconds, with the performer fully in frame.

Why it works

  • The prompt is short because the references carry the information — the video owns the motion, the images own the identities.
  • Each asset is addressed by its array position, so there is no ambiguity about which reference supplies what.
  • It asks for nothing else. No lighting, no camera, no style — every extra instruction would compete with the motion it is trying to copy.

Swap these out

  • characters
  • source choreography
  • environment
  • number of performers

Constraints

  • Total reference video duration must stay within 15 seconds, and each clip must be 2–15 seconds at 23.976–60 FPS.
  • Reference videos must be MP4 or MOV with H.264 or H.265, up to 50 MB each, with the whole JSON body under 64 MB.

Settings

Multimodal Reference · 10s · 16:9 · 3 reference assets

MiniMax H3 character introduction generated from a reference sheet: a black-haired warrior with a braided ponytail revealed from boots to full-body hero pose in stone ruinsMULTIMODAL REFERENCE
15s1:11 reference asset

Character & Motion

Character-Sheet Hero Intro

The library’s most-liked community template: one character reference sheet drives a boots-to-face-to-full-body cinematic introduction that works for any original character design.

View details

Full prompt

Use @[char ref] as the sole character reference. Preserve the exact identity, face, body proportions, hairstyle, outfit, colors, materials and overall silhouette of the character throughout the entire video. Do not redesign, simplify or replace any defining visual features. Create a cinematic character introduction focused on presence, silhouette, attitude and controlled motion. 0–4s Begin with a close shot of a defining lower-body or detail element such as boots, shoes, feet, hands, clothing hem or an important accessory. The character enters frame or settles into position. The camera slowly tracks upward while hair, clothing and secondary elements move naturally in the wind or environment. 4–8s Reveal more of the body with a medium or medium-wide shot from the back, side or three-quarter angle. The character stands in a calm, composed way inside the environment. The camera makes a smooth orbit, arc or lateral move to gradually reveal the character’s face and silhouette. 8–12s Move into a tight cinematic portrait or upper-body shot. The character performs one subtle signature action that fits their personality, such as lifting the chin, turning the head, adjusting clothing, brushing hair aside, opening a hand, looking toward camera, or shifting posture. Keep the motion minimal and intentional. The expression should match the character’s vibe. 12–15s End with a strong full-body hero shot that clearly presents the entire design and silhouette. Use a low-angle, eye-level or slightly dramatic framing depending on the character’s personality. The character settles into a natural final pose and holds it confidently for a clean final reveal. VISUAL DIRECTION Premium cinematic presentation. Match the visual medium and rendering style of @[char ref]. Emphasize clean silhouette, elegant staging, subtle secondary motion, believable hair and cloth movement, strong composition, atmospheric depth and polished lighting. The scene should feel like a high-end anime, game or film character introduction. CAMERA Use a clear progression from detail reveal to partial reveal to face reveal to full-body hero reveal. Camera movement should be smooth, controlled and intentional. Avoid chaotic motion. ENVIRONMENT Place the character in a fitting environment that supports their identity and mood. The background should enhance the character without distracting from them.

What you need to supply

  • Image 1: your character reference — a design sheet or clean full-body render; the video inherits identity, outfit, materials and rendering style from it.

Why it works

  • The reveal ladder (detail → partial → face → full-body hero) is a fixed camera grammar, so the model spends its variance on the character, not the shot plan.
  • "Match the visual medium and rendering style of the reference" makes one prompt work for anime, game-render and film-real characters alike.
  • The single "signature action" slot in the third block is where personality lives — one gesture, deliberately small.

Swap these out

  • character reference sheet
  • the signature action
  • environment mood
  • final pose framing

Constraints

  • The original post writes references as @[char ref]; on EvoLink’s route, assets are addressed by array position — say "Image 1" and it maps to image_urls[0].
  • One character, one environment: this template intentionally never cuts location, which is why the silhouette stays stable.

Settings

Multimodal Reference · 15s · 1:1 · 1 reference asset

MiniMax H3 Cinematic & VFX prompts

Sci-fi trailer clip generated with MiniMax H3: a lone figure before a vast circular cosmic gateway as a title resolves out of darknessMULTIMODAL REFERENCE
15s16:92 reference assets

Cinematic & VFX

Sci-Fi Mystery Trailer

A two-reference cinematic trailer beat: one image sets the atmosphere, the other locks the protagonist, and the prompt drives a single push-in with an on-screen title and matching audio.

View details

Full prompt

Realistic cinematic look, high-contrast lighting, and a tight pace. Use Figure 1 as the overall atmosphere and style reference, and Figure 2 as the protagonist reference. Shot 1 — Ultra-wide establishing shot. A huge circular cosmic gateway nearly fills the frame. The person is only a tiny figure seen from behind before the gateway, positioned toward the lower right. The ground is wet and reflective, and the center of the gateway is pitch black. The camera slowly pushes forward. A large title fades in from the edge of the darkness, blurred at first and then sharp: "THE STARS WERE LISTENING". Use an extremely condensed, heavy, all-caps typeface in dark red mixed with rust red, with subtle grain and misted edges. Audio: a deep low-frequency pulse, faint metallic vibrations in the distance, and a soft hit as the text becomes sharp. → Hard cut.

What you need to supply

  • Image 1: an atmosphere and style plate — the environment, palette and grade you want the shot to inherit.
  • Image 2: the protagonist, shot large enough that the face and silhouette stay readable when the figure is small in frame.
  • A title string you actually want burned in; H3 renders the text you write, so keep it short and spell it exactly.

Why it works

  • Each reference is given one job — Image 1 for atmosphere, Image 2 for the character — so the model never has to guess which plate owns the look.
  • The shot is described as one continuous push-in with a single event (the title resolving), which is what a 15-second budget can actually hold.
  • Audio is written as three specific layers (low pulse, distant metallic vibration, a hit on the text) instead of a vague 'cinematic soundtrack'.

Swap these out

  • title text and typeface treatment
  • gateway or landmark in the establishing shot
  • color of the title
  • audio bed
  • final cut behavior

Constraints

  • Reference the assets by array position — 'Figure 1', 'Figure 2' — matching image_urls order. The @image1 syntax is not part of this API contract.
  • The published prompt is a starter, deliberately incomplete: it ends on a hard cut so you can extend it with your own second shot.

Settings

Multimodal Reference · 15s · 16:9 · 2 reference assets

Text-to-video clip generated with MiniMax H3: a dusk kitchen filmed handheld while a hand-drawn glowing creature moves among the propsTEXT TO VIDEO
15s16:9No reference assets

Cinematic & VFX

Glowing Kitchen Creature

A pure text-to-video prompt that mixes live-action kitchen footage with hand-drawn glowing animation, and spends most of its words on camera imperfection and what must never appear.

View details

Full prompt

15-second, 16:9 landscape video. Blend live-action footage of a small kitchen at dusk with hand-drawn glowing animation. The last light of sunset lingers by the window. The lived-in kitchen contains an old wooden table, a half-washed mug, a slightly fogged glass bottle, and a hanging dishcloth. Give the footage subtle one-handed smartphone shake, hesitant close-range focusing, exposure fluctuations caused by backlight, and slightly coarse noise in the shadows. It should not look carefully arranged like an advertisement; instead, it should feel like someone hurriedly captured an unbelievable event at home. Do not show huge eyes, gaping mouths, fangs, threatening or lunging movements, sudden black frames, or jump scares. Use only kitchen room tone, cloth rubbing, the soft clink of a mug, water dripping from the faucet, the camera operator's footsteps and quiet breathing, plus gentle electronic sounds and tiny calls from the hand-drawn creature.

What you need to supply

  • Nothing to upload — this route takes the prompt only.
  • Decide duration and aspect ratio in the request, not in the prompt text.

Why it works

  • It specifies the camera's flaws — one-handed shake, hesitant focus, backlight exposure swings, noise in the shadows — which is what sells 'someone filmed this at home' over 'this is an advert'.
  • It carries an explicit do-not list (no huge eyes, no fangs, no lunging, no jump scares) that keeps a cute creature from drifting into horror.
  • The audio direction names each source separately: room tone, cloth, mug, tap, footsteps, breathing, plus the creature's own sounds.

Swap these out

  • room and time of day
  • creature design and behavior
  • props on the table
  • which imperfections the camera shows
  • audio layers

Constraints

  • The prompt states its own length and framing in words; the API still needs duration and aspect_ratio as request fields.
  • Negative instructions work best as a short explicit list. Stacking dozens of prohibitions costs prompt budget you need for the action.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 anime race generated from a first frame: two hover-bikes trade positions through a wet mountain hairpin trailing cyan and crimson lightFIRST / LAST FRAME
15s16:91 reference asset

Cinematic & VFX

Hover-Bike Anime Race

A first-frame-driven sports-anime race with a CRITICAL ENTITY LOCK: exactly two named riders on two fully-specified bikes, tracked through timed hairpins, drafts and a photo-finish.

View details

Full prompt

Cinematic Anime Video Scene Generate a 15-second horizontal 16:9 original high-speed hover-bike racing anime video from the provided first frame. CRITICAL ENTITY LOCK: There must be exactly 2 racers and 2 bikes in the entire video: RENJI on VALKYRIE-01 (cyan/black drift bike) and ELENA on AERO-X (crimson/white draft bike). Do not add extra racers, drone support vehicles, spectators, or traffic. Maintain total visual consistency for both bikes, helmet visors, suit patterns, repulsor spark colors, and bike liveries throughout the sequence. Entity identity: VALKYRIE-01: Matte-black and cyan angular hover-bike, exposed repulsor pads, lateral drift brakes, blue plasma exhaust trails, ridden by Renji (cyan trim suit). AERO-X: Pearl-white and neon-crimson aerodynamic hover-bike, enclosed canopy, crimson energy draft aura, white-hot central booster, ridden by Elena (crimson/gold visor suit). Video style: High-budget modern sports anime, sakuga-level velocity animation, crisp line art, vibrant neon lighting contrast, high-speed camera tracking, hyper-realistic friction and energy particle effects. Set on a wet downhill mountain pass at dawn. Camera and pacing: Continuous forward velocity, zero slow-motion interruptions: 0.0s - 3.0s: High-speed rear-tracking shot diving into the first downhill hairpin curve; instant drift initiation. 3.0s - 7.5s: Tight side-parallel tracking shot as bikes navigate rock debris and trade positions through S-curves. 7.5s - 11.5s: Close camera lock on the draft-slingshot maneuver; high-energy particle displacement as booster ignition occurs. 11.5s - 15.0s: Low-angle front-facing camera lock on the final sprint to the finish line bridge, ending on a hyper-speed photo-finish freeze. Action timing: 0.0s - 1.5s: Sequence begins at speed. VALKYRIE-01 leads downhill; AERO-X locks onto its rear bumper. Anti-gravity repulsors spray road water and blue sparks into the frame. 1.5s - 4.0s: First sharp hairpin. VALKYRIE-01 deploys lateral drift airbrakes with a burst of blue thruster fire, sliding sideways at 300 km/h. AERO-X stays glued inside its slipstream aura. 4.0s - 7.0s: Mountain debris hazard. VALKYRIE-01 hops over a boulder using a repulsor burst. AERO-X ducks under it, scraping the neon magenta guardrail in a cloud of friction sparks. 7.0s - 10.0s: S-Curve exchange. Bikes lean side-by-side; their repulsor fields collide, creating a bright electrical shockwave. ELENA pulls the overdrive lever; AERO-X's rear fins extend. 10.0s - 13.0s: Slingshot maneuver. AERO-X bursts out of VALKYRIE-01's draft, igniting its central white plasma booster. Both bikes roar down the final straightaway side-by-side. 13.0s - 15.0s: Final sprint toward the finish light gate. Water sprays violently behind them. Both nose cones cross the finish line simultaneously in a flash of light. Final freeze frame. Motion quality: Fluid 2D animation, extreme speed-line integration, stable bike geometry, flawless vehicle reflection rendering, zero limb or body clipping, high-frame-rate kinetic realism. Environment: Wet mountain pass asphalt, sheer cliff walls, neon cyan and magenta guardrail lights, early dawn sky with pink/purple clouds, water spray, floating spark particles. Final output: 15 seconds, horizontal 16:9, original high-budget sports racing anime, exactly 2 racers, relentless kinetic pacing, dynamic cinematography, no subtitles, no watermarks, no logos.

What you need to supply

  • Start frame: a still of the two racers and bikes in your art style — the whole sequence extends from this image.

Why it works

  • Entity lock by census ("exactly 2 racers and 2 bikes… no spectators, no traffic") kills the crowd-inflation failure that racing prompts usually hit.
  • Both machines get liveries, physics quirks and rider suits as named identities, so the model can keep them apart at speed.
  • Camera and action live on separate timelines that reference each other, keeping relentless pace without ever asking for two camera moves at once.

Swap these out

  • start frame art style
  • bike identities and liveries
  • track hazards
  • finish-line staging

Constraints

  • This is the first-frame route: the start image carries the art style, and image-to-video accepts no aspect_ratio parameter — the frame defines it.

Settings

First / Last Frame · 15s · 16:9 · 1 reference asset

MiniMax H3 Native Audio & Dialogue prompts

Dialogue replacement generated with MiniMax H3: a character's spoken line swapped for a new line from a reference audio clipMULTIMODAL REFERENCE
10s16:92 reference assets

Native Audio & Dialogue

Dialogue and Performance Replacement

Swapping a spoken line inside an existing clip: the old line is quoted, the new line is quoted, and the performance is allowed to shift just enough to match.

View details

Full prompt

Replace the girl's line in Video 1, "We can't be together. It's not that we don't love each other; we truly can't make it to the end," with the line from Audio 1: "Don't go, okay? This time, let's not let go of each other." Slightly adjust the corresponding performance.

What you need to supply

  • Video 1: the clip containing the line to replace.
  • Audio 1: the replacement line, WAV or MP3, up to 15 MB and 15 seconds.
  • Both lines written out verbatim in the prompt.

Why it works

  • Quoting the outgoing line tells the model exactly which span of the clip to operate on, instead of 'the dialogue near the middle'.
  • Quoting the incoming line means the lip sync has a target rather than being inferred from the audio alone.
  • 'Slightly adjust the corresponding performance' grants a bounded licence to change the acting — bounded, so the rest of the take survives.

Swap these out

  • source clip
  • outgoing line
  • replacement line and voice
  • how much performance may change

Constraints

  • Audio can never be the only reference type; it must arrive with an image or a video.
  • Reference audio and video each cap at 15 seconds of total duration per request.
  • Say what stays fixed. Framing, wardrobe and background will drift if the prompt only talks about the line.

Settings

Multimodal Reference · 10s · 16:9 · 2 reference assets

Voice-reference clip generated with MiniMax H3: a character speaking a written line in a timbre taken from a reference audio clipMULTIMODAL REFERENCE
10s16:92 reference assets

Native Audio & Dialogue

Wind Voice Clone

The minimum viable audio-reference prompt: the line to speak, and one clip that defines whose voice speaks it.

View details

Full prompt

Character dialogue: "Follow the wind, live free. Leave worries behind, enjoy the moment." Use Audio 1 as the voice-timbre reference.

What you need to supply

  • Video 1: the character who will deliver the line.
  • Audio 1: a clean sample of the target voice, 2–15 seconds, ideally without music underneath.
  • The exact line of dialogue, written out.

Why it works

  • The dialogue is quoted rather than paraphrased, so timing and lip sync have something concrete to lock onto.
  • Audio 1 is assigned one narrow job — voice timbre — instead of being handed over as a general 'soundtrack'.
  • Nothing else is specified, so the reference clip keeps full control of framing and performance.

Swap these out

  • line of dialogue
  • voice reference
  • character clip
  • delivery pace

Constraints

  • A voice reference transfers timbre, not accent, emotion or pacing; write those into the prompt if they matter.
  • Music or overlapping speakers in the reference audio degrade the result — supply an isolated voice.

Settings

Multimodal Reference · 10s · 16:9 · 2 reference assets

MiniMax H3 montage from character and audio references: a horned white-haired character cut across five environments in rapid beat-synced shotsMULTIMODAL REFERENCE
15s1:12 reference assets

Native Audio & Dialogue

Audio-Synced Environment Montage

A dual-reference montage where a character image locks identity and an audio clip dictates the edit: five environments, six burst-cut shots each, every cut landing on the track’s accents.

View details

Full prompt

Use @[char ref] as the strict character reference and @[audio ref] as the timing, rhythm and editing reference. Keep the character’s exact identity, proportions, hairstyle, outfit, colors and overall style consistent throughout. Create a 15-second cinematic burst-cut video showcasing the character across 5 different environments that naturally fit their design, vibe and world. AUDIO SYNC Synchronize the entire edit to @[audio ref]. Cuts, camera accents, transitions and environment changes should land precisely on strong beats, half-beats and musical accents. Let audio1 control the pacing and intensity of the montage. STRUCTURE - 5 environments total - 3 seconds per environment - 6 burst-cut shots per environment - 30 shots total Each environment must be clearly different in atmosphere, lighting, scale and visual language. Show each environment through rapid cinematic angles: wide establishing shots, aerials, low angles, side views, tracking shots, close environmental details, medium shots and hero frames. Every cut must reveal a new angle, distance, composition or spatial relationship. Avoid repeated framing. Mix static shots, push-ins, pull-backs, tracking, orbit and crane-like movement. Keep character movement subtle and natural. The focus is environmental variety, cinematic framing and tight synchronization with audio1. Hard constraints: - exactly 5 environments - exactly 6 shots per environment - exactly 30 shots total - environment changes must follow audio1’s musical phrasing - cuts and motion accents synchronized to audio1 - no outfit changes - no character duplication - no morphing - no text or UI - no blurry unreadable frames - maintain strict character consistency

What you need to supply

  • Image 1: the character whose identity every shot must keep.
  • Audio 1: the track that owns pacing — cuts, transitions and environment changes follow its beats.

Why it works

  • It divides labor cleanly between modalities: the image answers "who", the audio answers "when" — neither fights the text.
  • Exact arithmetic (5 environments × 6 shots = 30 cuts) is stated as a hard constraint, turning a vague montage into a countable structure.
  • "Character movement stays subtle" pushes all the energy into camera variety, which montages survive far better than action variety.

Swap these out

  • character reference
  • audio track and its phrasing
  • the five environments
  • shot-type mix

Constraints

  • The original post writes @[char ref] and @[audio ref]; on EvoLink address them by array position — Image 1 and Audio 1 — and remember audio can never be the only reference type.
  • Reference audio clips must run 2 to 15 seconds on this route.
  • The published clip measures 8:9, which the API does not accept — request 1:1 for the same near-square framing.

Settings

Multimodal Reference · 15s · 1:1 · 2 reference assets

MiniMax H3 Game & Interface prompts

MiniMax H3 UI animation from nine references: a minimal creature encyclopedia where a cursor selects MEADOW CROWN, a fluffy horned creature, in a fieldMULTIMODAL REFERENCE
15s16:99 reference assets

Game & Interface

Creature Encyclopedia UI Demo

Nine references mapped to explicit roles — one UI layout plate plus eight creature cards — animated as a locked-camera encyclopedia demo where a cursor clicks through entries and the last creature eats it.

View details

Full prompt

Use Image 1 as the exact UI/layout/style reference for the creature encyclopedia screen. Use Images 2–9 as the exact creature references. Map them like this: Image 2 = card A = LUMI HARE Image 3 = card B = CLOUD WISP Image 4 = card C = EMBER FENNEC Image 5 = card D = TIDE BEHEMOTH Image 6 = card E = PETAL VULPIN Image 7 = card F = ORCHARD EYE Image 8 = card G = MEADOW CROWN Image 9 = card H = FROST GLIDER Create a 15-second 16:9 video. Keep the camera locked. Keep the interface, layout, typography, panels, icons and overall composition stable, elegant and readable. The UI should feel like a modern minimal digital creature encyclopedia, similar to a sleek pokedex. No scene cuts, no extra text, no extra buttons, no UI distortion. Sequence: 0–2.5s: Cursor clicks card A. Main creature becomes LUMI HARE. Title changes to “LUMI HARE”. Creature blinks and rotates slightly. 2.5–5s: Cursor clicks card C. Main creature becomes EMBER FENNEC. Title changes to “EMBER FENNEC”. Cursor drags to rotate it left and right. 5–7.5s: Cursor clicks card E. Main creature becomes PETAL VULPIN. Title changes to “PETAL VULPIN”. Cursor pokes it a few times. It reacts, annoyed. 7.5–10s: Cursor clicks card G. Main creature becomes MEADOW CROWN. Title changes to “MEADOW CROWN”. Cursor taps near the face/horns. It recoils slightly. 10–12s: Cursor clicks card D. Main creature becomes TIDE BEHEMOTH. Title changes to “TIDE BEHEMOTH”. Cursor keeps poking it. 12–15s: TIDE BEHEMOTH gets angry, opens its mouth very wide, lunges forward, and swallows the cursor. Then it returns to idle. Title stays “TIDE BEHEMOTH”. Rules: - When a card is selected, both the main creature and the main title must update. - Only animate cursor, selection state, title change, and the selected creature. - Only one cursor. - Keep motion subtle and clean until the final swallow. - No cuts, no camera move, no UI distortion, no extra text. Audio: soft UI click sounds, subtle hover sounds, tiny creature reaction sounds, then a sharper aggressive creature sound and one comedic swallow gulp at the end.

What you need to supply

  • Image 1: the encyclopedia UI layout that owns typography, panels and composition.
  • Images 2–9: one creature per card, each named in the prompt’s card map.

Why it works

  • The image-to-card mapping table (Image 2 = card A = LUMI HARE…) is the most literal role assignment possible — the model never guesses which asset is which.
  • The camera is locked and only four things may animate (cursor, selection, title, active creature), shrinking the failure surface of UI motion to near zero.
  • The comedic beat at the end (creature swallows cursor) is scheduled last so the demo stays clean until the payoff.

Swap these out

  • UI style plate
  • the eight creatures and names
  • interaction script
  • sound-effect set

Constraints

  • Nine images is the per-type maximum on this route — and the 12-file total cap means you cannot add three videos and three audio clips on top.
  • UI text stability relies on "no camera move, no UI distortion"; freeing the camera reintroduces warped panels.

Settings

Multimodal Reference · 15s · 16:9 · 9 reference assets

MiniMax H3 gameplay-style clip: a first-person view down a rifle scope advancing through a smoky military base with a generic FPS HUDTEXT TO VIDEO
15s16:9No reference assets

Game & Interface

FPS Gameplay Simulation

A first-person shooter sequence that reads as captured gameplay: player-controlled camera grammar, a fully-specified generic HUD, and pacing written as tactics instead of choreography.

View details

Full prompt

Camera: First-person perspective at eye level with authentic handheld player movement, as if recorded directly from a modern AAA military shooter. The player carries a highly detailed assault rifle with realistic animations, visible hands, tactical gloves, dynamic reload mechanics, and weapon sway. **Opening Action:** The video immediately begins with the player already aiming down a roadway inside a modern military base. Multiple enemy soldiers are visible in the distance near sandbags, barricades, and military vehicles. The player carefully tracks one target, making small aim corrections while maintaining ADS (aim down sights). Fire several controlled bursts immediately at the visible enemies, producing realistic muzzle flashes, shell casings ejecting, smoke, recoil, hit reactions, and dust impacts around the targets. Continue firing in multiple short bursts while adjusting aim between enemies, simulating authentic FPS gameplay rather than scripted animation. **Movement:** After the opening firefight, lower slightly from ADS and begin advancing cautiously along the road beside concrete barriers, Hesco walls, and parked military vehicles. Frequently check left and right corners, briefly stop to reacquire targets, then raise the weapon and fire additional controlled bursts whenever enemies appear ahead. Continue pushing forward with deliberate player-controlled movement, using cover naturally and maintaining believable tactical pacing. **Environment:** Large modern military base with guard towers, armored vehicles, shipping containers, blast barriers, damaged buildings, smoke plumes, burning debris, scattered shell casings, dust clouds, and atmospheric battlefield haze. Cool natural daylight mixed with smoke and orange firelight creates a cinematic battlefield atmosphere. **Camera Motion:** Authentic player-controlled movement with subtle head bob, weapon sway, natural mouse-look adjustments, small left-right corrections while aiming, realistic recoil, smooth tracking of moving targets, brief pauses before shooting, and fluid forward progression. Avoid cinematic camera moves—everything should feel like genuine live gameplay captured by a skilled player. **Visual Quality:** Ultra-photorealistic, AAA game graphics with realistic PBR materials, detailed weapon models, physically accurate lighting, volumetric smoke, dynamic particle effects, crisp textures, realistic bullet impacts, muzzle flash illumination, motion blur only during rapid movement, and high-end military shooter presentation. **Gameplay UI:** Display a realistic modern FPS HUD inspired by games like PUBG, Battlefield, or Call of Duty (without copying exact copyrighted assets). Include: * Central dynamic crosshair or reticle * Ammo counter with magazine and reserve ammunition * Fire mode indicator * Compass at the top * Squad/team status panel * Mini-map in the upper corner * Health bar * Tactical equipment icons (grenades, medkit) * Hit markers when bullets connect * Directional damage indicators * Kill notification feed * Objective marker in the distance * Subtle interaction prompts and realistic HUD animations The HUD should feel polished, modern, and fully integrated into the gameplay, enhancing the illusion of authentic recorded footage from a contemporary military FPS. #MiniMaxH3

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • The realism target is "recorded from a skilled player", so head bob, aim corrections and pause-then-fire pacing are specified as camera behavior.
  • The HUD is itemized down to hit markers and kill feed while explicitly staying generic — dense enough to read as a real game, safe enough to publish.
  • "Avoid cinematic camera moves" is the key inversion: what most prompts want is exactly what would break this one.

Swap these out

  • environment and faction styling
  • HUD element set
  • engagement rhythm
  • weather and light

Constraints

  • The prompt keeps the HUD "inspired by, without copying" — preserve that clause; cloning a specific game’s HUD is a different (and riskier) task.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 Vlog & Selfie Camera prompts

MiniMax H3 selfie-style clip: a woman filming herself in a forest turns her phone to reveal a crashed smoking UFO between the treesTEXT TO VIDEO
15s1:1No reference assets

Vlog & Selfie Camera

Selfie-Cam UFO Discovery

A phone-realism showcase: handheld selfie footage with autofocus breathing and rolling shutter, scripted Japanese dialogue, and a camera flip from face to crashed UFO.

View details

Full prompt

Ultra photorealistic live-action captured on an iPhone 17. Authentic handheld selfie footage with premium cinematic documentary color grading, realistic HDR, deep green foliage, warm sunlight, subtle teal shadows, natural skin tones, gentle filmic contrast, rolling shutter, autofocus breathing, slight motion blur, and natural handheld shake. A lush forest in daytime with dense trees, wild plants, an uneven dirt trail, scattered leaves, soft sunlight through the canopy, and a gentle breeze. The atmosphere is quiet and slightly unsettling. A cute Japanese woman in her early twenties wearing a stylish bikini walks through the forest while recording herself in selfie mode. She suddenly notices something ahead, looks shocked, turns the camera, and points into the distance. A large crashed UFO is partially embedded in the forest floor. Its metallic hull is badly damaged with broken panels, scorch marks, exposed internal structures, thick gray smoke, and occasional sparks. She says in Japanese: 「ちょっと待って! あそこ見て! UFOじゃない!? 完全に墜落してるんだけど! 煙まで出てる! やばい、本物かもしれない! ちょっと近づいてみる!」 She alternates between filming herself and the UFO while continuing to point at it. Continuous single take. Natural walking movement, realistic hand tremors, slight framing imperfections, quick pans, and autofocus shifts between her face and the UFO. Natural sunlight creates cinematic highlights, soft shadows, realistic reflections on the UFO, subtle volumetric light, and realistic smoke. Audio: footsteps on leaves, gentle wind, birds becoming quieter near the crash site, creaking branches, faint electrical crackling from the UFO, and distant eerie unidentified animal calls echoing through the forest. Negative: no blood, no visible aliens, no monsters, no horror creature reveal, no excessive explosions, no CGI, no cartoon style, no text, no subtitles, no watermark, no logo.

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • Phone artifacts are enumerated (autofocus hunting, rolling shutter, hand tremor, framing imperfections) — realism comes from named flaws, not the word "realistic".
  • The dialogue is quoted verbatim in Japanese, so H3’s native audio generates actual speech with matching lip sync instead of gibberish.
  • Audio is layered diegetically — footsteps, wind, birds going quiet, electrical crackle — with no soundtrack to break the found-phone illusion.

Swap these out

  • spoken lines and language
  • the discovered object
  • forest or urban setting
  • outfit and character styling

Constraints

  • The negative list (no aliens, no horror reveal) is load-bearing: it keeps the clip in teaser territory and stops the model escalating the scene.
  • The published clip ran 1:1 — set the frame via the aspect parameter, and keep the alternating self/object framing in the text.

Settings

Text to Video · 15s · 1:1 · No reference assets

MiniMax H3 documentary clip: a street photographer frames an elderly man and his terrier outside a café, then shows the captured photo to the viewerTEXT TO VIDEO
15s16:9No reference assets

Vlog & Selfie Camera

Street Photographer Moment

A compact documentary beat with a camera-inside-the-camera: a photographer frames a candid street scene, takes the shot, then turns her camera to show the viewer the photo she just captured.

View details

Full prompt

A young Western female street photographer walks through a lively downtown street and notices an elderly man sitting outside a café with his small dog. She carefully composes the candid moment through her camera, captures the photo, then turns the camera toward the viewer to proudly show the shot she just took. She smiles, says “Look at that,” then continues walking through the city. Ultra-photorealistic visuals, natural handheld documentary movement, realistic camera interaction, authentic facial expressions, accurate hand movements, realistic dog behavior, natural daylight, cinematic depth of field, continuous character consistency, immersive city ambience, premium documentary realism.

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • The show-the-photo beat forces H3 to render a coherent still image inside the video — a quiet capability demo that reads as one natural gesture.
  • Interaction chain is fully specified (compose → capture → flip → "Look at that" → walk on), so the clip has a complete arc in one take.
  • The subjects are described by role, not identity — elderly man, small dog, café — keeping the street scene generic and safe.

Swap these out

  • the candid subject
  • city and light
  • the spoken line
  • camera prop type

Constraints

  • The in-camera photo must match the scene it was taken from; if you change the subjects, change both descriptions together.

Settings

Text to Video · 15s · 16:9 · No reference assets

MiniMax H3 travel vlog: a woman opens curtains onto a seaside terrace, then flips to selfie mode saying good morning with the ocean behind herTEXT TO VIDEO
15s9:16No reference assets

Vlog & Selfie Camera

Seaside Morning Vlog Arc

A second-by-second morning vlog with a deliberate mode switch: the first five seconds are cinematic third-person, and only then does the character start filming herself — the moment the "vlog" begins.

View details

Full prompt

Create a 15-second ultra-realistic cinematic lifestyle vlog video, vertical 9:16, featuring the same young woman throughout the entire video. Preserve her facial identity, facial proportions, hairstyle, skin tone and overall appearance consistently in every shot. She wears the same outfit throughout: fitted white V-neck T-shirt with a small subtle logo, blue denim jeans, natural makeup, long softly wavy brown hair. 0:00–0:01 — Wake-up: Close-up inside a beautiful bright bedroom. The woman is lying comfortably on the bed, slowly wakes up, stretches naturally and opens her eyes. She is NOT filming a vlog yet and does not hold a phone or camera. Soft morning sunlight enters through the curtains. 0:01–0:02 — Gets up: Medium shot. She sits up on the bed, smiles softly, fixes her hair and gets ready to start her morning. Natural, effortless movement. 0:02–0:03 — Walks to window: She walks toward the large glass balcony door/window. Camera follows her naturally from behind/side. 0:03–0:04 — Seaside reveal: She opens the curtains/door and looks outside. Reveal a breathtaking blue ocean, coastal hills, flowers, balcony and beautiful morning sunlight. She smiles happily while taking in the view. 0:04–0:05 — Steps outside: She walks out onto the seaside terrace. Gentle ocean breeze moves her hair naturally. Wide cinematic shot showing the beautiful surroundings. 0:05–0:06 — VLOG START: Only now she starts filming herself in handheld selfie-vlog style. She looks into the camera with a bright natural smile and says: “Good morning!” 0:06–0:07 — Show the view: She turns the camera away from herself and slowly pans across the stunning ocean, coastal mountains, flowers and terrace. Smooth handheld vlog movement. 0:07–0:08 — Back to selfie: Selfie shot. She looks into the camera and happily says: “This place is just perfect!” 0:08–0:09 — Location reveal: Wide cinematic shot of the cozy seaside terrace with wooden table, chairs, plants and flowers overlooking the ocean. 0:09–0:10 — Walk to table: Medium tracking shot as she walks toward the table, enjoying the view. Her hair and T-shirt move gently in the sea breeze. 0:10–0:11 — Sit and relax: She sits at the seaside table, smiling peacefully and enjoying the ocean view. A refreshing orange-colored juice is placed on the table. 0:11–0:12 — Juice close-up: Cinematic close-up of her hand picking up the glass of fresh orange juice. Beautiful ocean bokeh in the background, natural sunlight reflecting through the glass. 0:12–0:13 — Vlog toast: Selfie shot. She raises the juice toward the camera with a cheerful smile and says: “Cheers to good days!” 0:13–0:14 — Happy close-up: Beautiful close-up of her smiling naturally at the camera, ocean and warm sunlight softly blurred behind her. 0:14–0:15 — Ending: Camera moves from her toward the sparkling ocean and peaceful coastal landscape. Warm sunlight, gentle waves and a relaxing cinematic ending. Overall Style Ultra-realistic, cinematic travel vlog, natural handheld camera movement, realistic human motion, smooth transitions, soft morning sunlight, realistic ocean waves, gentle wind in hair and clothes, beautiful coastal atmosphere, premium lifestyle aesthetic, natural expressions, authentic vlog feeling, shallow depth of field, cinematic composition, realistic skin texture, high detail, 4K quality.

What you need to supply

  • Nothing to upload — this route takes the prompt only.

Why it works

  • The "she is NOT filming yet" note in the wake-up beats prevents the most common vlog artifact — a phone appearing before the vlog starts.
  • Fifteen one-second beats alternate selfie cam and scenic cutaways exactly like real travel-vlog grammar.
  • Three short quoted lines ("Good morning!") give native audio natural checkpoints without a monologue to sustain.

Swap these out

  • location reveal
  • the three spoken lines
  • outfit and identity lock
  • the drink prop

Constraints

  • The prompt asks for 9:16 and that is the ratio to request; the published clip was re-encoded to roughly 3:2, another reason to set orientation in the request rather than in prose.

Settings

Text to Video · 15s · 9:16 · No reference assets

MiniMax H3 Editing & Transformation prompts

Video edit generated with MiniMax H3: a source clip with the newspaper, chair, sunglasses and burning car all replaced or removedMULTIMODAL REFERENCE
10s16:91 reference asset

Editing & Transformation

Multi-Element Scene Edit

Six separate edits to one source clip, written as a flat list of change instructions with no scene description at all.

View details

Full prompt

Replace the newspaper in the reference video with a green-covered book; change the chair the character is sitting on to a red sofa; remove the sunglasses worn by the character to retain a clear face; remove the car burning effect to keep the vehicle in a normal state; change the photo the character takes out of his arms to a small black book; and add a tree on the left side of the screen

What you need to supply

  • Video 1: the clip to edit, 2–15 seconds, MP4 or MOV, H.264 or H.265, up to 50 MB.
  • A list of edits, each naming what to change and what it becomes.

Why it works

  • Every instruction is a pair — target plus result — so nothing is left to interpretation ('the newspaper' becomes 'a green-covered book').
  • Removals state the intended end state ('remove the sunglasses to retain a clear face'), which stops the model from leaving a hole where the object was.
  • It carries no scene description whatsoever, so the model treats the source clip as ground truth and only applies the deltas.

Swap these out

  • number of edits
  • objects replaced
  • effects removed
  • elements added

Constraints

  • EvoLink exposes three H3 routes; edit-style work goes through reference-to-video with the source clip as Video 1.
  • Reference video duration is billable, so trim the source to the section you actually need before uploading.
  • State what must stay unchanged when an edit sits next to something you care about — unstated regions are fair game for the model.

Settings

Multimodal Reference · 10s · 16:9 · 1 reference asset

Stage clip generated with MiniMax H3: two magicians swapping suit colors in a puff of smoke as the curtain shifts from red to blueFIRST / LAST FRAME
7s16:91 reference asset

Editing & Transformation

Magician Costume Swap

A staged trick used as an instruction-following test: two costumes swap, one detail must not change, and the background completes a color transition on cue.

View details

Full prompt

Two magicians stand onstage facing the audience and perform a "swap" trick. They wave their wands at the same time, and a cloud of smoke rises. When it clears, their suit colors have switched: the person on the left wears a white suit, and the person on the right now wears a black suit, while both magicians' glove colors remain unchanged. They bow to thank the audience. The red curtain behind them closes, transitioning from deep red to deep blue.

What you need to supply

  • A start image holding both performers, their costumes and the stage.
  • A clear before/after state for whatever swaps.

Why it works

  • The swap is masked by an event — the smoke cloud — giving the model a legitimate moment to make the change instead of morphing on camera.
  • It names what must not change (the glove colors) next to what must, which is exactly how you keep an edit from spreading.
  • The shot ends on a defined final state: the bow, the curtain closing, the color landing on deep blue.

Swap these out

  • performers and costumes
  • the detail that stays fixed
  • cover event for the swap
  • closing color transition

Constraints

  • The image-to-video route derives its aspect ratio from the input image and rejects the aspect_ratio field.
  • Every unstated attribute is fair game for the model. If a detail must survive the change, write it down.

Settings

First / Last Frame · 7s · 16:9 · 1 reference asset

Looking for the MiniMax H3 API?

Model page with per-second pricing, three generation routes, and integration docs.

Get API access

MiniMax H3 prompt guide: the framework

H3 takes ordered references and generates its own audio. Write to the workflow instead of relying on a generic subject + action + camera + style formula.

The H3 prompt skeleton

Goal → ordered references → subject and identity anchors → chronological action beats → camera path → audio or dialogue direction → what must stay unchanged → final state. Work down that list and you will not forget the two lines most prompts miss: the audio and the ending.

Text to Video

You own everything, so spend the words on a timeline the clip can actually finish. Name the subject, order the beats, give the camera one continuous path, then describe the sound as separate layers — dialogue, ambience, music — rather than as a mood.

First / Last Frame

The frames already carry the appearance. Describe only the change between them: the motion added, the camera move, what must be preserved, and the state you land on. Re-describing the picture is the most common way to waste this route.

Multimodal Reference

Give every asset one job and address it by array position — Image 1 for the character, Image 2 for the location, Video 1 for the camera path, Audio 1 for the voice. Up to 9 images, 3 videos and 3 audio clips per request, and no more than 12 files in total; audio can never travel alone.

Editing an existing clip

Write two lists, not one paragraph: Change (target → result, one line each) and Preserve (anything adjacent that must survive). Send the source clip as Video 1 on the reference route, trimmed to the section you actually need.

Audio and dialogue

Quote the lines you want spoken, verbatim. Assign a voice reference one narrow job — timbre — and write accent, emotion and pacing into the prompt separately. Keep dialogue, ambience and music from competing for the same seconds.

MiniMax H3 prompt FAQ

Is MiniMax H3 the same thing as Hailuo 3?

Yes. MiniMax H3 is also known as Hailuo 3, Hailuo 3.0 or Hailuo 03. EvoLink uses MiniMax H3 as the product name and exposes it through three routes: text-to-video, image-to-video and reference-to-video.

Can I run these prompts without writing any code?

Yes. Use this prompt sends the text straight into the MiniMax H3 playground on the model page with the matching workflow already selected. You supply any reference assets there and generate in the browser.

How long can a MiniMax H3 clip be?

Duration is an integer from 4 to 15 seconds on all three routes, and the current routes output 2K. The prompts on this page were measured against their published clips, so the duration shown on each card is what that example actually runs.

How many reference assets can one prompt use?

The reference-to-video route accepts up to 9 images, 3 videos and 3 audio clips in a single request, capped at 12 files in total — so a full 9 + 3 + 3 set is rejected. At least one image or video is required; audio can never be the only reference type. Reference video and audio clips must each run 2 to 15 seconds.

Why do the prompts say Figure 1 or Video 1?

The reference route resolves assets by their position in the image_urls, video_urls and audio_urls arrays. Referring to Image 1 or Video 1 in the prompt is how you assign a job to a specific asset. The @image1 syntax is not part of this API contract.

Where do these prompts and clips come from?

Every card names its source. Official MiniMax samples are published with permission and their prompt text is the English edition of the original Chinese prompt. Community samples link back to the creator’s original post on X.

Can I use these prompts through the API?

Yes. The prompt text is identical whether you paste it into the playground or send it as the prompt field on the API. The MiniMax H3 API guide covers request structure, async tasks, callbacks and error handling.

Pick a prompt and generate it

Every prompt maps to a live MiniMax H3 route on EvoLink. Send one to the playground, or build the same request against the unified API.