GPT Image 2.5 Flare 与 Sunburst 已上线 EvoLink立即体验

MiniMax H3 提示词与视频案例

浏览 40 个附带真实 MiniMax H3 生成结果、输入要求和可替换变量的已核验提示词。可直接复制模板,或将它发送到 EvoLink Playground。

MiniMax H3 也常被称为海螺 3、Hailuo 3.0 或 Hailuo 03。

生成模式

使用场景

显示 40 / 40 条提示词

MiniMax H3 品牌与产品广告提示词

MiniMax H3 生成的竖屏眼镜广告:两位模特在无缝白色影棚中佩戴未来感包覆式眼镜多模态参考
15s9:163 个参考素材

品牌与产品广告

未来感眼镜广告

三张图驱动的时尚广告:主视觉、模特面孔和产品设计分别由不同参考图负责。

查看详情

完整提示词(已验证英文原文)

Generate a vertical screen 9:16 high-end fashion glasses commercial, taking overall reference to the storyboard rhythm, editing speed, white studio texture and cool fashion atmosphere of the given video. The picture is a minimalist white booth, a seamless white background, a strong sense of high-end advertising, and a clean, simple, handsome, avant-garde, international first-line fashion blockbuster texture. Key visual character reference picture 1, two full-body female models, one black female model and one European and American model, maintain their high-end clothing, body posture, white studio light and shadow, fashion show temperament and overall cool attitude. Both of them wear futuristic high-end glasses. The design of the glasses refers to Figure 3, emphasizing the covered curved surface, sharp geometric cat-eye/goggle hybrid outline, mirror reflection, streamlined temples, and the texture of high-end fashion accessories. Please refer to Figure 2 for the appearance details of the two characters.

需要提供的素材

  • 图片 1:主视觉,包含模特全身、服装、影棚光线和态度。
  • 图片 2:两位人物的外观细节,用于保持剪辑中的面孔一致性。
  • 图片 3:产品参考,需足够清晰,以便快速镜头中仍能保留轮廓和材质。
  • 一个愿意贯穿整条视频的无缝影棚背景。

为什么有效

  • 三个参考素材各有一个明确职责,正是多模态参考路由的最佳用法;歧义才是多图提示词失败的主因。
  • 产品通过包覆式曲率、猫眼与护目镜混合轮廓、镜面材质和流线镜腿等几何特征来描述,而不只是写产品名称。
  • 单一材质的白色影棚降低背景变化,让模型把资源留给产品和演员。

可替换变量

  • 产品类别
  • 模特选角
  • 影棚颜色
  • 剪辑速度
  • 服装

限制与注意事项

  • 原提示词还要求参考给定视频的节奏,但当前公开素材只包含三张图。可将该句视为风格指令,或自行添加视频 1。
  • 每次参考请求最多支持 9 张图片、3 段视频和 3 段音频,且文件总数不超过 12 个;音频不能作为唯一参考类型。

生成设置

多模态参考 · 15s · 9:16 · 3 个参考素材

MiniMax H3 生成的产品片:高端头戴式耳机在黑色影棚的反光展台上旋转文生视频
15s16:9无参考素材

品牌与产品广告

高端耳机产品展示

把 15 秒产品片拆成四个时间段,每段都有独立运镜和任务;这是本页最适合无参考素材用户直接复用的模板。

查看详情

完整提示词(已验证英文原文)

Create a 15-second luxury cinematic product showcase for premium wireless over-ear headphones. 0–4s: Begin with an extreme macro tracking shot moving across the soft memory-foam ear cushion, fine fabric texture, brushed-metal hinge and precision-machined controls. A narrow light band travels across the surface, revealing realistic materials against a deep black studio background. 4–8s: Pull back into a three-quarter hero view. The headphones rotate slowly above a glossy reflective pedestal. The ear cups pivot naturally while the adjustable headband extends slightly, demonstrating flexible construction and comfort. Maintain exact symmetry, stable geometry and consistent proportions. 8–12s: Transition into an elegant exploded-view reveal. The ear cushion, acoustic driver, internal sound chamber, control ring and outer shell separate smoothly in perfect alignment. Subtle luminous sound waves pulse outward from the driver while the camera performs a restrained side orbit. 12–15s: Every component reconnects seamlessly. The headphones settle into a centered front-facing hero composition as soft rim lighting defines the silhouette. Complete a gentle dolly-in toward the ear cups. Premium technology-commercial finish, controlled reflections, realistic shadows, shallow depth of field, crisp surface detail, stable product shape, no hands, no distortion, no onscreen text.

需要提供的素材

  • 无需上传素材,文生视频只需提示词。
  • 选择一个能用材质和机械结构描述的产品,不要只写名称。

为什么有效

  • 时间被拆成 0–4 秒、4–8 秒、8–12 秒和 12–15 秒,让模型获得镜头清单,而不是一串愿望。
  • 每个时间段采用不同运镜:微距跟移、拉远、侧向环绕和向前推进,让片子像经过剪辑,而不是漫无目的地漂移。
  • 结尾的禁止列表覆盖产品渲染最常见的三种问题:手部、变形和屏幕文字。

可替换变量

  • 产品
  • 材质与表面处理
  • 每个时间段的分配
  • 背景与光线
  • 是否使用爆炸分解视图

限制与注意事项

  • 时长是 4–15 的整数,并在请求中设置;提示词里各时间段的总和必须与之一致。
  • 分时指令是引导,不是硬时间轴;15 秒视频建议不超过四个时间段。
  • 该案例由作者以 720p 发布;EvoLink 当前 H3 路由输出 2K。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 饮料广告:中暑的行人喝下一口冰镇果汁,街道随即在瓶身周围绽放出葱郁绿意文生视频
15s16:9无参考素材

品牌与产品广告

夏日高温饮料广告

经典的"问题—缓解"式饮料广告:被酷暑折磨的主角喝下一口,整个环境随之焕然一新,最后落在挂满水珠的产品特写主镜头上。

查看详情

完整提示词(已验证英文原文)

A young person walks under the blazing summer sun, looking exhausted and sweating heavily. The road shimmers with heat waves, and everything appears dry and dull. Suddenly, they grab a chilled bottle of premium fruit juice from a cooler and take a refreshing sip. Instantly, the environment transforms—lush green trees bloom, vibrant flowers appear, a cool breeze flows, water splashes through the air, and glowing particles surround the scene. Ice cubes and fresh fruit slices (orange, mango, or according to the flavour) swirl around the bottle in cinematic slow motion. End with a stunning close-up of the juice bottle covered in cold water droplets against a bright, refreshing background. Ultra-realistic, premium commercial, 4K, cinematic lighting, high-detail, smooth camera movements, vibrant colours, luxury beverage advertisement. Tagline ideas: Beat the Heat. Taste the Freshness. Every Sip Brings Life. Refresh Your Day, Naturally. Stay Cool. Stay Fresh.

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • 广告被构建成前后状态对比——干燥酷热对比葱郁清新——给模型一个明确的转变任务去执行,而不是一堆情绪词的罗列。
  • 产品在后段才登场,并以特写主镜头收尾,这正是模型可以直接模式匹配的标准饮料广告节拍顺序。
  • 风味元素(冰块、水果切片)被写成在电影感慢动作中环绕瓶身旋转的实体物件,而不是抽象形容词。

可替换变量

  • 饮料类型与风味元素
  • 主角中暑场景的设定
  • 转变后的环境
  • 标语文字

限制与注意事项

  • 结尾的标语列表只是文案灵感,不会烧录进画面——若希望某句标语渲染上屏,要把它写进视觉指令。
  • 15 秒的预算只够一次转变;再加第二个产品或场景变化会挤爆时间窗口。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 UGC 与达人广告提示词

MiniMax H3 由五个参考生成的开播画面:动漫风格主播在类 Twitch 布局中看着聊天栏,带 LIVE 徽章和关注横幅多模态参考
15s16:95 个参考素材

UGC 与达人广告

VTuber 开播场景

五素材制作:四张图分别负责身份、平台界面、房间和开播卡片,音频参考驱动真实的口型同步主播表演——连它的静音前奏都写进了脚本。

查看详情

完整提示词(已验证英文原文)

Use @Image4as the opening card only: circular Luna avatar, black background, cream “LUNALIVE”, rose “STREAM STARTING”, warm circles, tiny mint accent. Static, chime. Use @Image1 for Luna’s identity only: same face, long center-parted black hair, blue-gray eyes, pink anime hoodie, white headphones around neck, delicate necklace, pale nails. Do not copy the drink, pose, or background from @Image1 . Use @Image2 only for Twitch-like platform chrome: dark top bar with “LUNALIVE”, red “LIVE” badge, “2.4K viewers”, right “STREAM CHAT” rail, bottom title area, rounded “FOLLOW” and pink “SUBSCRIBE” buttons. Do not copy Luna, pose, drink, or room from @Image2 Use @Image3 only for the cozy room behind Luna inside the video area: desk, monitor, white PC, plush shelves, curtain fairy lights, soft pink/purple light. No empty-room showcase shot. Use @Audio1as Luna’s actual vocal performance and behavior reference. @Audio1has a silent lead-in: 0.0–2.4s must be treated as no speech. Preserve her voice identity, cadence, tone, breaths, pauses, emphasis, warmth, and streamer mannerisms. Do not replace the voice, do not generate a different influencer voice, and do not add extra spoken lines beyond @Audio1 Lip sync, mouth shapes, jaw movement, smiles, glances, nods, and small hand movements must follow the audio waveform after 2.4s. Create a 15-second 16:9 Twitch-like stream opening. Important performance direction: Luna is reading chat, not delivering a camera monologue. Place Luna slightly left of center in the video area with the right “STREAM CHAT” rail clearly visible. Whenever she speaks, her eyes angle screen-right toward the chat rail as if she is reading the messages out loud. She returns to camera only for brief reactions. Add constant small movement: eye darts to chat, eyebrow lifts, tiny nods, head tilts, shoulders shifting, one subtle hand gesture near the desk. No stiff talking-head pose. One cursor only, no cursor trail, no duplicate panels, no duplicate buttons. Render only large UI text cleanly. Chat feels alive with typing dots and soft short blurred lines, but only these chat lines are readable: “hi Luna”, “welcome back”, “so cozy”, “gugugaga?”, “RAID INCOMING!”. Do not invent usernames. Chat pops stay quieter than Luna’s voice. [0–2 seconds] Open on @Image4 “LUNALIVE / STREAM STARTING”. Absolute no-speech zone: @Audio1is silent here, Luna is not shown, no mouth movement, no voice on the card. One soft chime only. No movement. [2–7.4 seconds] Hard cut to Luna live at 2.0s, slightly left of center in the cozy room from @Image3 with @Image2 chrome active: “LUNALIVE”, red “LIVE”, “2.4K viewers”, right “STREAM CHAT” rail, bottom title “COZY NEON” and “Just Chatting”. She settles for a beat, eyes already moving toward the chat rail. At about 2.4s when speech begins in @Audio1 , match lip sync exactly while she reads toward the chat rail, not into camera. Chat shows typing dots and soft blurred lines. [7.4–7.9 seconds] First audio pause = chat beat. Typing dots, then readable messages pop in: “hi Luna”, “welcome back”. Luna’s eyes track the new messages on the right rail; small nod and smile follow the audio pause. [7.9–12.5 seconds] Continue matching @Audio1Keep her gaze mostly on the chat rail while speaking, like she is reading and reacting. Add one more readable message: “so cozy”. During any softer phrase, she leans slightly forward as if reading; during brighter phrases, eyebrows lift and shoulders react. No frozen face. [12.5–13.9 seconds] Bigger audio pause = bigger chat beat. “gugugaga?” appears, then “RAID INCOMING!”, and a clean “NEW FOLLOWER” banner slides in with a gentle pop. Luna reads the raid message from the chat rail, then reacts brighter as the audio resumes. [13.9–14.8 seconds] Finish @Audio1 with accurate lip sync. If the audio winds down, Luna stops talking, gives a small wave toward chat, and settles into a warm listening pose. End on the stable live frame: Luna slightly left of center, eyes toward the right chat rail, red “LIVE”, “2.4K viewers”, no end card. Audio mix: 0–2s card is silent except one soft chime. @Audio1voice begins only after the cut to Luna and remains primary. Tiny chat pops under the voice. Warm low room tone. No crowd noise, no music lyrics, no rain, no traffic.

需要提供的素材

  • 图片 1:只提供主播身份——面孔、发型、卫衣;姿势和背景被明确排除在复制范围外。
  • 图片 2:直播平台界面——聊天栏、LIVE 徽章、按钮。
  • 图片 3:她身后的温馨房间。
  • 图片 4:静态的"stream starting"开播卡片。
  • 音频 1:她的真实人声素材;注意 0–2.4 秒的静音被脚本标记为禁言区。

为什么有效

  • 每个参考素材领到一项职责和两条明确的 "do not copy" 排除项,是本库最干净的按素材分工。
  • 表演指令——"she is reading chat, not delivering a monologue"——重新导向了视线和微动作,正是这点让片段有直播感。
  • 可读的聊天消息被白名单限定为五条精确字符串,其余全部保持模糊,UI 文字因此渲染干净。

可替换变量

  • 主播身份底板
  • 平台界面样式
  • 白名单聊天内容
  • 音频素材及其停顿

限制与注意事项

  • 原帖写作 @Image1…@Audio1;在 EvoLink 上按数组位置寻址,且必须至少有一张图片与音频同行——音频永远不能单独作为参考。
  • 该路由最多接受 9 张图片、3 段视频和 3 段音频,文件总数上限 12 个。

生成设置

多模态参考 · 15s · 16:9 · 5 个参考素材

MiniMax H3 UGC 广告:女子在明亮卧室中自拍,用滴管瓶为头皮涂抹精华,手机拍摄式构图多模态参考
15s3:41 个参考素材

UGC 与达人广告

UGC 护发精华广告

四场景 TikTok 风格精华液证言广告,只用一张产品参考图:自拍开场、涂抹特写、镜前效果、台面产品镜头——不完美本身就是造型。

查看详情

完整提示词(已验证英文原文)

Create a 15-second authentic UGC-style hair growth serum ad using the provided product image as the exact product reference. SCENE 1 — 0–3s A young woman films herself in a bright bedroom using a smartphone front camera. Natural lighting, handheld movement, casual appearance. She looks at the camera and says: “I’ve been trying this hair growth serum lately…” SCENE 2 — 3–7s Cut to a close handheld shot of the woman holding the serum bottle. She removes the dropper, applies a few drops directly to her scalp, and gently massages it in. Keep the movement natural and slightly imperfect like real UGC content. SCENE 3 — 7–11s Mirror selfie shot. She runs her fingers through her hair, showing healthy-looking, fuller hair while casually talking to the camera: “And honestly, I love how easy it is to add to my routine.” SCENE 4 — 11–15s Close-up product shot on her bathroom counter. She picks up the bottle and smiles toward the camera. End with natural on-screen text: “Simple hair care. Every day.” Style: authentic TikTok/Reels UGC, smartphone camera, realistic skin texture, natural expressions, subtle handheld motion, imperfect framing, casual home environment, soft daylight, realistic audio, no cinematic commercial look, no excessive beauty filters. Preserve the exact product packaging, label, bottle shape, and branding from the reference image.

需要提供的素材

  • 图片 1:你的产品图——包装、标签和瓶身造型在全部四个场景中都从这张底板锁定。

为什么有效

  • 场景顺序复刻了原生创作者内容(钩子 → 演示 → 效果 → 产品),广告在信息流里读起来像原生内容。
  • "Slightly imperfect like real UGC" 加上反商业质感的负面清单,正是击败用户会直接划走的精致广告感的关键。
  • 口播台词简短、带引号、口语化——原生音频会把它们演绎成自然的证言语气。

可替换变量

  • 产品底板
  • 两句口播台词
  • 卧室/浴室场景
  • 收尾屏幕文字

限制与注意事项

  • 产品身份完全由参考图承载;再在文字里描述标签会引入冲突。
  • 已发布片段实测为 2:3,而 API 并不接受这一比例——想要贴合信息流的竖幅就请求 3:4 或 9:16。

生成设置

多模态参考 · 15s · 3:4 · 1 个参考素材

MiniMax H3 美食 vlog:年轻女子在快速跳切之间笑着把芝士汉堡举向手机镜头多模态参考
15s21:91 个参考素材

UGC 与达人广告

汉堡之夜 UGC vlog

八镜头美食 vlog 模板,汉堡参考图全程锁定:自拍钩子、开盒、急推特写、芝士拉丝、咬一口的反应——节奏全靠硬跳切完成。

查看详情

完整提示词(已验证英文原文)

Duration: 15 seconds | Aspect Ratio: 16:9 | Style: Authentic UGC / iPhone selfie-vlog, handheld, natural light, TikTok/Reels aesthetic. Product Reference: Use the uploaded gourmet burger image as the only product reference. Preserve the bun shape, patty thickness, cheese melt, lettuce, tomato, sauces, and proportions exactly in every shot. Character Description Name: Hana A young Japanese woman in her early 20s with natural beauty, long dark hair in a loose ponytail, oversized cream sweatshirt, minimal makeup, bright smile, friendly lifestyle-vlogger personality. Shot Breakdown SHOT 1 (0–2s) — Selfie showing the burger box. Dialogue: "Burger night!" SHOT 2 (2–4s) — Opens the box. SHOT 3 (4–6s) — Quick zoom on the burger. SHOT 4 (6–8s) — Hands lifting the burger with cheese stretching naturally. SHOT 5 (8–10s) — Bite reaction. Dialogue: "Okay... that's incredible." SHOT 6 (10–12s) — Casual close-up b-roll while reaching for fries. SHOT 7 (12–14s) — Toasting the burger toward the camera. Dialogue: "You need this." SHOT 8 (14–15s) — Freeze frame with overlay: "burger cravings = solved 🍔" Look & Feel Warm apartment lighting, genuine phone footage, slight grain, natural autofocus breathing, handheld imperfections, fast jump cuts. Negative Prompt cinematic grading, commercial production, CGI burger, fake cheese, distorted hands, warped food, perfect stabilization, studio lighting, text glitches, logo distortion.

需要提供的素材

  • 图片 1:食物主图——面包胚、肉饼、芝士融化状态和比例在每次剪切中保持一致。

为什么有效

  • 八个约 2 秒的微镜头与真实美食创作者的剪法一致,能量因此是这个格式原生的。
  • 具名角色("Hana")加两句简短的引用台词,给了模型稳定的面孔和声音,又不至于过度设定。
  • 食物物理有专属指令(芝士自然拉丝)——名场面是写出来的,不是碰运气。

可替换变量

  • 食物及其呈现
  • 角色造型
  • 两句对白
  • 定格收尾文字

限制与注意事项

  • 食物描述只放在参考图里;负面列表(no CGI burger、no fake cheese)守护真实感。

生成设置

多模态参考 · 15s · 21:9 · 1 个参考素材

MiniMax H3 文字与动态排版提示词

MiniMax H3 时装片:水彩缎带牵引镜头环绕身穿白色高领长裙的女子,短语 LET SILENCE BLOOM 以实体字母成形文生视频
15s16:9无参考素材

文字与动态排版

水彩高定时装一镜到底

一部先锋时装片:水彩缎带牵引单一连续镜头穿越空间,文字以实体物件的形式存在,每次运动都落在音乐节拍上。

查看详情

完整提示词(已验证英文原文)

## Concept AQUARELLE No.7 — an avant-garde haute couture film where watercolor becomes a living medium. 15-second cinematic sequence. The rhythm controls the visual world: 0–4s: restrained silence and negative space 4–8s: gradual density buildup 8s: drop moment, expanding into fluid long-form motion 12–15s: transition into a final fashion poster composition ## Character Identity Lock Maintain the exact same female character throughout the entire video. Identity: - Young woman - Long black hair - Calm and refined facial features - White high-neck pigment dress - Black wide belt - Pigment heels Strict consistency: - Same face - Same hairstyle - Same age - Same body proportions - Same garment structure Any new colors must only appear through watercolor gradually absorbing into the fabric. Core concept: Her fingertips can extract transparent watercolor ribbons from the air. These watercolor ribbons can: - Pull the camera through space - Transform the environment - Shape physical typography - Interact with depth and materials The world follows real cinematic physics: water, paper, fabric, light, shadows, and depth of field must feel physically believable. ## Camera Direction Strict one-take shot. No cuts. No teleportation. No hidden transitions. Camera journey: Macro shot of a floating water droplet → watercolor ribbon emerges → camera pulls back to reveal the woman → camera circles around her → enters a paper art gallery → passes through dimensional typography → rises into a final overhead fashion poster. ## Typography Only allow these words: LET SILENCE BLOOM AQUARELLE No.7 WEAR THE UNSEEN Typography is not a flat overlay. Letters must have: - Physical depth - Wet watercolor reflections - Paper fiber edges - Shadows - Occlusion - Material interaction ## Visual Style Avant-garde fashion editorial. Inspired by: - Museum catalog composition - Handmade ivory paper - Sculptural negative space - Elegant Didone serif typography - Translucent watercolor calligraphy Color palette: - Pale cyan - Crimson lake - Smoky violet - Ink black ## Motion Language Every movement follows the music: Kick: Camera movement and paper folding. Wooden snare: Paper structures physically fold and transform. Sub-bass: Changes the perception of spatial scale. ## Storyboard 30 beats, 0.5 seconds each. 01 0.0–0.5 Macro shot: A transparent water droplet floats in the air, reflecting a blurred silhouette of a black-haired woman. 02 0.5–1.0 The droplet stretches with the breath-like vocal, becoming a pale cyan watercolor thread. 03 1.0–1.5 Camera travels backward along the thread as paper fibers slowly emerge into focus. 04 1.5–2.0 The thread wraps around the lens. Focus shifts to her raised fingertip. 05 2.0–3.0 Camera continues pulling back, revealing her face and white high-neck dress. 06 3.0–4.0 She moves her wrist. The watercolor thread guides a smooth camera arc. 07 4.0–8.0 Additional watercolor colors emerge from her movement. Paper folds, typography begins forming, and the phrase "LET SILENCE BLOOM" appears as a physical object. 08 8.0–12.0 The camera passes through a transparent paper flower structure. The environment expands into an endless ivory paper gallery. Her dress absorbs watercolor naturally. The ribbons create sculptural forms around her body. 09 12.0–15.0 Camera cranes upward. The composition transforms into a luxury fashion advertisement poster. Typography appears: AQUARELLE No.7 WEAR THE UNSEEN Final frame: A museum-level fashion editorial poster. The woman remains centered, calm, and elegant. A final watercolor droplet remains suspended in the air. ## Negative Prompt No: - Cuts - Scene changes - Identity change - Face swap - Extra limbs - Deformed hands - Random costume changes - Explosive paint effects without physical cause - Incorrect typography - Chinese characters - Extra subtitles - Extra logos - Watermarks

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • 摄影机旅程被写成一条连续链路(水滴 → 缎带 → 人物揭示 → 纸艺画廊 → 顶拍海报),并把 "no cuts" 定为硬规则。
  • 排版走白名单制——只允许出现三句短语——并被赋予材质属性(纸纤维边缘、湿润水彩反光、遮挡关系),这正是文字能干净渲染的原因。
  • 30 拍、每拍 0.5 秒的分镜把声音映射到空间:底鼓折叠纸张、军鼓让纸结构变形、次低频改变空间尺度感。

可替换变量

  • 三句允许出现的短语
  • 颜料色板
  • 服装结构
  • 画廊环境

限制与注意事项

  • 短语白名单就是排版质量机制——加更多文字会重新引入这份提示词专门规避的拼写漂移。
  • 一镜到底的提示词失败起来非常明显:任何一拍暗示了剪切,整条空间链路就会断裂。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 教育动画:圆润的字母 A 在柔和粉彩场景中膨胀成一颗带笑脸的红苹果,旁边写着单词 APPLE文生视频
15s16:9无参考素材

文字与动态排版

A-B-C-D 识字动画

一部儿童自然拼读动画,采用固定教学循环——字母、发音、物体、动作、单词——每个字母都物理变形为对应物体,旁白按秒写好脚本。

查看详情

完整提示词(已验证英文原文)

Create a 15-second animated educational video that teaches young children the letters A, B, C, and D. The learning pattern for every letter must be: LETTER → SOUND → OBJECT → PLAYFUL ACTION → OBJECT NAME Target audience: children ages 3 to 6. Visual style: Use adorable rounded 3D characters, soft pastel colors, gentle facial expressions, and simple recognizable objects. Combine this with a premium minimalist technology aesthetic featuring clean white space, elegant composition, soft studio lighting, subtle reflections, smooth gradients, rounded geometry, crisp typography, and extremely polished transitions. The animation should feel playful and child-friendly while remaining calm, uncluttered, and beautifully designed. Use a clean off-white background with a different soft color glow behind each letter. 0:00–0:01 | Introduction A small smiling star mascot bounces into the center of the screen. Colorful letters briefly float around it. Display the text: “Let’s learn!” The mascot taps the screen, creating a soft ripple that reveals the first letter. 0:01–0:04 | A is for Apple Show a large uppercase “A” and smaller lowercase “a” beside it. Use thick, rounded, highly readable typography. The narrator says: “A. A says ah. A is for Apple.” The uppercase A gently inflates and transforms into a shiny red apple. Its top point becomes the apple stem, and a small green leaf unfolds from the side. The apple gains a cute smiling face and performs one soft bounce. Display the word: “APPLE” Highlight the first letter A in red. Add a soft pop and a tiny crunchy sound. 0:04–0:07 | B is for Ball The apple rolls across the screen and leaves behind a curved red trail. The trail loops twice and forms a large uppercase “B,” with a lowercase “b” appearing beside it. The narrator says: “B. B says buh. B is for Ball.” The two rounded sections of the B expand and merge into a colorful striped ball. The ball bounces twice with playful squash-and-stretch animation. Display the word: “BALL” Highlight the first letter B in blue. Synchronize each bounce with a soft musical note. 0:07–0:10 | C is for Cat On its final bounce, the ball stretches into a curved shape and becomes a large uppercase “C.” A lowercase “c” slides gently into place beside it. The narrator says: “C. C says kuh. C is for Cat.” The C rotates and becomes the curled tail of a cute orange cat. The rest of the cat forms from soft rounded shapes. The cat stretches, blinks, and gives one gentle wave with its paw. Display the word: “CAT” Highlight the first letter C in orange. Add a quiet and friendly “meow.” 0:10–0:13 | D is for Duck The cat’s tail uncurls and transforms into the curved side of a large uppercase “D.” A lowercase “d” pops up beside it. The narrator says: “D. D says duh. D is for Duck.” The straight line of the D becomes the duck’s neck. The curved section becomes its round yellow body. A small orange beak and two tiny wings pop into place. The duck waddles forward, flaps its wings, and gives one cheerful quack. Display the word: “DUCK” Highlight the first letter D in yellow. Add tiny water ripples beneath its feet. 0:13–0:15 | Recap The apple, ball, cat, and duck slide into four clean rounded tiles. Place their letters above them: “A B C D” The mascot returns and points to each object as they bounce once in sequence. Narrator: “A, B, C, D. Great job!” Finish with the text: “Great job!” Use a small sparkle animation and a warm musical chime. Animation requirements: Keep each letter fully visible for a moment before it transforms. Show uppercase and lowercase versions clearly. Make every object instantly recognizable. Use smooth shape morphing so children can visually understand how the letter becomes the object. Maintain stable spelling, clean letterforms, accurate object shapes, and consistent character design. Use gentle squash-and-stretch, soft motion blur, subtle shadows, polished lighting, and precisely synchronized sound effects. Avoid fast camera movement, cluttered backgrounds, harsh colors, tiny text, warped letters, random symbols, duplicated objects, scary expressions, or overly complex transformations. The final video should feel cute, educational, memorable, calming, and exceptionally polished.

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • LETTER → SOUND → OBJECT → ACTION → NAME 的循环以完全相同的结构重复四次,而重复正是 H3 手里最强的稳定器。
  • 每次变形都是几何层面的解释(A 的尖角变成苹果梗、B 的两个圆弧合成皮球),字形因此能在变形中幸存。
  • 旁白台词逐句引用并配有分段时间戳,驱动原生音频的同时保证字幕文字同步。

可替换变量

  • 四个字母及对应物体
  • 吉祥物设计
  • 每个字母的背景光晕色
  • 旁白声音

限制与注意事项

  • 拼写稳定性依赖"每个字母先完整展示再变形"这条规则——删掉它就会重新出现扭曲字形。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 动态排版:细衬线字母组成的 Every great change 从黑暗中浮现,金色微粒缓缓漂浮文生视频
15s16:9无参考素材

文字与动态排版

动态引语排版

纯动态排版:一句引语按短语逐段揭示,每个短语都改变场景的氛围、运动语言和色板,最终完整句子定格在米白色收尾卡上。

查看详情

完整提示词(已验证英文原文)

Create a 15-second cinematic text-animation video built around the quote: “Every great change begins quietly, grows through courage, and becomes impossible to ignore.” The quote should appear gradually as a visual story. Each new phrase must transform the design, atmosphere, movement, and emotional intensity of the scene. Use elegant typography, accurate spelling, cinematic lighting, smooth transitions, and perfectly readable text. 0:00–0:03 | “Every great change” Begin with a completely black screen. A tiny point of warm light slowly appears in the center, like the first spark of an idea. The words “Every great change” emerge softly from the darkness, one word at a time. Use thin, elegant serif typography with wide letter spacing. “Every” fades in gently. “Great” grows slightly larger. “Change” forms from small drifting particles that gather into solid letters. Keep the scene quiet, minimal, and mysterious. 0:03–0:06 | “begins quietly,” The camera slowly moves closer to the text. The previous words shrink and reposition toward the upper-left corner as the phrase “begins quietly,” appears in delicate lowercase letters. Animate the phrase as though it is being written by an invisible hand. Each letter should create a subtle ripple in the darkness. Introduce faint textures, soft shadows, floating dust, and gentle light rays. The comma should appear last and create a small circular pulse. 0:06–0:09 | “grows through courage,” The pulse expands and transforms the scene from darkness into a rich sunrise gradient with deep orange, red, and golden tones. The words “grows through courage” rise upward from the bottom of the frame. Animate “grows” by gradually increasing its size and weight. Animate “through” along a curved path. Animate “courage” in bold uppercase letters that push through a translucent barrier, causing it to crack into geometric fragments. The movement should feel powerful but controlled. 0:09–0:12 | “and becomes” The fragments rotate in slow motion and reorganize into a clean editorial grid. The phrase “AND BECOMES” appears across the frame in condensed sans-serif typography. Animate the letters with fast tracking changes, vertical stretching, masking, and perspective movement. The camera accelerates forward through the center of the word “BECOMES.” The sound and visual energy should steadily build. 0:12–0:14 | “impossible to ignore.” Reveal a vast bright space filled with light, moving shapes, and large-scale typography. The words “IMPOSSIBLE TO IGNORE” appear one after another. “IMPOSSIBLE” expands beyond the edges of the screen. “TO” remains small and perfectly centered. “IGNORE” slams into place with strong visual impact, briefly shaking the surrounding grid and shapes. Use bold contrast, dramatic scale, sharp shadows, and synchronized motion. 0:14–0:15 | Final quote All movement stops instantly. The complete quote appears centered on a clean off-white background: “Every great change begins quietly, grows through courage, and becomes impossible to ignore.” Use refined black typography with “change,” “courage,” and “impossible” highlighted in deep red. Hold the final composition clearly for the last second. Maintain one continuous visual journey from darkness to light, silence to impact, and simplicity to complexity. Keep every phrase connected through visual transformations rather than hard cuts. Use realistic motion blur, precise kerning, clean masks, stable letterforms, smooth camera movement, subtle film grain, cinematic sound design, rising ambient music, soft particles, controlled color transitions, and a final deep impact sound. Avoid misspelled words, warped letters, duplicated characters, unreadable text, random symbols, excessive flickering, chaotic layouts, inconsistent fonts.

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • 每个短语独占一个 3 秒区块并配有专属的动画动词(浮现、手写、上升、砸落),能量的递进因此是结构性的,而不是靠形容词。
  • 情绪弧线被映射到设计语言上——从黑暗到日出渐变再到编辑排版网格——给模型的是一份色板脚本,而不只是文字。
  • 一个硬停顿("all movement stops instantly")加上停留展示的收尾卡,保证了可读的最终帧,方便做封面缩略图。

可替换变量

  • 引语及其短语切分
  • 高亮词与强调色
  • 每个短语的动画动词
  • 收尾卡样式

限制与注意事项

  • 每个短语控制在约 6 个词以内;H3 渲染简短的展示性文字远比整句长行可靠。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 短剧与叙事提示词

MiniMax H3 生成的竖屏短剧预告:烛光城堡中的吸血鬼男主与人类女主多模态参考
15s9:162 个参考素材

短剧与叙事

吸血鬼爱情短剧

竖屏短剧开场模板:一张图锁定男女主外观,另一张图锁定场景,提示词把篇幅留给关系节拍和镜头景别,而不是塞入完整剧情。

查看详情

完整提示词(已验证英文原文)

Generate a 15-second, 9:16 vertical trailer segment for an international live-action vampire romance short drama. Use Figure 1 as the appearance reference for the male and female leads, and Figure 2 as the scene reference. Keep both leads' identities consistent, with a realistic live-action look and premium short-drama production quality. Story: an innocent human heroine accidentally enters a forbidden area of an old castle and awakens a sleeping aristocratic vampire. He discovers that she carries an aura connected to an ancient war, which sparks a powerful urge to control her and a dangerous fascination with her. She fears him but does not completely submit, resisting his pressure. Overall style: an international ReelShort / DramaBox vampire-romance trailer. Dark romance, dangerous attraction, fate, intense control, brooding oppression, and a striking reversal. Keep the visuals premium, restrained, and tightly paced, like the opening 15-second hook of a hit short drama. No gore, cheap horror, Halloween aesthetic, or modern street feel. Format: 9:16 vertical composition for TikTok / ReelShort / DramaBox. Use primarily medium close-ups, close-ups, and extreme close-ups, emphasizing faces, eye contact, pressure, and relationship tension within the vertical frame.

需要提供的素材

  • 图片 1:男女主同框参考,让模型把他们视为一组稳定选角。
  • 图片 2:场景参考,例如城堡内部、光线和材质。
  • 一句话故事前提和一个明确情绪反转;15 秒钩子只适合承载一次反转,不是整集故事。

为什么有效

  • 直接指定 ReelShort / DramaBox 预告片风格,可以用少量文字传递节奏、调色和构图惯例。
  • 限定中近景、近景和特写的镜头语汇,让竖屏画面更精致,而不是显得拥挤。
  • 排除血腥、廉价恐怖、万圣节风和现代街头感,避免这一类型最常见的四种偏移。

可替换变量

  • 主角外观参考
  • 场景
  • 前提与反转
  • 目标平台风格
  • 景别组合

限制与注意事项

  • 竖屏输出需通过参考路由的 aspect_ratio 请求,不能只在提示词中写“9:16”。
  • 男女主同时出现在一张图中,通常比两张独立裁切肖像更容易保持身份一致。

生成设置

多模态参考 · 15s · 9:16 · 2 个参考素材

MiniMax H3 生成的视觉小说界面转场,画面在锁定的开始帧和结束帧之间过渡首尾帧
15s16:92 个参考素材

短剧与叙事

乙女视觉小说转场

标准的首尾帧提示词:两张图锁定镜头的起点和终点,文字只需描述两个状态之间如何过渡。

查看详情

完整提示词(已验证英文原文)

Use the first image as the opening frame and the second image as the exact final frame to generate an otome visual-novel interface transition. Overall feel: a premium Chinese otome romance-interaction interface capturing an intimate moment before and after a performance. Transition naturally from "choose to watch his performance" to "Han Xu is drawn in by the heroine's words and reacts with intrigued interest." UI text, choices, and dialogue boxes should appear with refined otome-game presentation. Keep the transition silky smooth and the emotion suggestive yet restrained.

需要提供的素材

  • 图片 1:包含完整 UI 状态的开场帧。
  • 图片 2:希望准确落入的最终帧。
  • 一句话描述两个状态之间的情绪变化。

为什么有效

  • 起点和终点都被锁定,模型要解决的是插值而不是重新构图,这是获得可预测镜头的稳定方法。
  • 提示词描述情绪转变(“选择观看他的表演” → “被女主的话吸引并产生兴趣”),而不是逐帧罗列画面。
  • 明确要求界面元素按乙女游戏的表达方式动画,可减少 UI 被重绘。

可替换变量

  • 开始帧与结束帧
  • 两帧之间的情绪弧线
  • UI 表现风格
  • 转场速度

限制与注意事项

  • 在图生视频路由中分别以 image_start 和 image_end 传入两帧;参考路由不接受这两个字段。
  • 两帧的构图和光照越接近,过渡越平滑;完全无关的构图更像硬切,而不是转场。

生成设置

首尾帧 · 15s · 16:9 · 2 个参考素材

MiniMax H3 年代片段:身穿 1940 年代军服的士兵在烟雾弥漫的美国乡村街道上依托老式汽车掩护,画面如档案胶片文生视频
15s16:9无参考素材

短剧与叙事

1940 年代战地新闻片

一段考据严谨的 1947 年新闻片模拟:真实感来自三层叠加的系统——年代一致的场景陈设、一份物理契约,以及封禁一切现代物件的负面列表。

查看详情

完整提示词(已验证英文原文)

Create a 15-second ultra-photorealistic live-action war sequence set in the United States in 1947, designed to look like authentic historical footage captured on a 1940s film camera. The entire scene must feel grounded, documentary-like, raw, and physically realistic. Environment: A rural American town in 1947 with wooden houses, old brick buildings, telephone poles, dirt roads, vintage American cars from the 1940s, wooden fences, farmland, and period-accurate street details. Overcast afternoon light, light fog, drifting smoke, dust in the air, damaged buildings, scattered debris, and a tense wartime atmosphere. Characters: American soldiers wearing historically accurate late-1940s military uniforms, helmets, boots, and equipment. Civilians wear authentic 1940s American clothing. Natural faces, realistic skin texture, sweat, dirt, fatigue, and believable body movements. 0–3s — Establishing Shot: Wide handheld shot of a quiet rural American street suddenly filled with smoke and confusion. Vintage 1940s vehicles are parked along the road while soldiers move quickly between wooden buildings. Civilians rush toward safer areas. 3–6s — Tension: Camera moves through the street at shoulder height, following several soldiers as distant gunfire is heard. They immediately react and take cover behind a vintage vehicle and a brick wall. Their movements are cautious and realistic. 6–10s — Combat: Fast handheld tracking shot as the soldiers move between cover while distant gunfire impacts the environment. Small pieces of wood, dust, and debris fall naturally from nearby impacts. Weapon recoil, movement, and body weight must be physically accurate. Keep the violence realistic and restrained. 10–13s — Human Moment: Camera briefly focuses on a soldier helping an injured civilian move behind cover. Their breathing, facial expressions, body language, and movement should feel natural and unscripted. 13–15s — Final Shot: Camera pulls back into a wide shot of the American town as smoke slowly moves through the street. Soldiers remain behind cover while vintage vehicles and damaged buildings fill the background. The scene ends with an authentic, tense 1940s documentary feeling. Visual Style: Ultra-photorealistic live-action, authentic 1940s American environment, vintage 35mm film texture, subtle film grain, natural imperfections, realistic exposure, handheld documentary cinematography, muted historical color palette, realistic smoke and dust, natural shadows, accurate depth of field. Physics: Strictly obey real-world gravity, momentum, inertia, friction, recoil, weight, collision physics, and human biomechanics. No exaggerated explosions, impossible movements, superhero behavior, or choreographed-looking combat. Negative Prompt: modern buildings, modern cars, smartphones, modern clothing, modern weapons, futuristic technology, CGI appearance, video-game graphics, fantasy, superhero action, excessive explosions, excessive blood, gore, impossible physics, unrealistic recoil, slow-motion physics, distorted faces, extra limbs, floating objects, plastic skin, artificial-looking environments.

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • 年代准确性被双向强制——正向(1940 年代汽车、军服、电线杆)与反向(禁止智能手机、现代建筑)——同时堵住两个失败方向。
  • 物理段落("strictly obey gravity, momentum, recoil…")读起来像渲染规格书,明显抑制了超级英雄式动作。
  • 一个安静的人性节拍(士兵帮助受伤平民)被排在 10–13 秒,让整段获得纪录片的可信度,而不是无休止的动作戏。

可替换变量

  • 年代与地点
  • 人性时刻
  • 胶片质感
  • 战斗节拍的强度

限制与注意事项

  • 克制条款("violence realistic and restrained"、无血腥)正是这个输出可用的原因之一——改编时要保留。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 角色与动作提示词

MiniMax H3 生成的黏土动画:黏土狐狸飞跃熔岩峡谷,摄影机从它下方掠过首尾帧
10s16:91 个参考素材

角色与动作

黏土狐狸飞跃峡谷

把一张起始图像变成一个明确动作,并像描述角色动作一样精确地描述摄影机路径。

查看详情

完整提示词(已验证英文原文)

Claymation style. A sprinting fox reaches the edge of a cliff and launches without hesitation, making a dramatically tense, heroic slow-motion leap across a vast lava canyon. While the fox is airborne, the camera rushes at high speed beneath its belly in a sweeping dynamic move, fully revealing the terrifying depth of the chasm and the fox's clay body at maximum extension in midair.

需要提供的素材

  • 一张已经包含风格的起始图,例如黏土狐狸和它的材质。
  • 只选一个要发生的主动作,不要同时写三个。

为什么有效

  • 提示词描述的是“变化”而不是图片外观;起始帧已经承载静态信息,因此每个字都用来购买动作。
  • 摄影机有独立指令:从狐狸腹部下方高速掠过,让一次跳跃真正变成一个镜头。
  • “在空中完全舒展”为模型提供了一个目标姿势,时序可围绕该峰值时刻展开。

可替换变量

  • 角色与材质风格
  • 环境和危险元素
  • 摄影机路径
  • 慢动作重点

限制与注意事项

  • 图生视频路由会根据输入图片确定输出比例,不接受 aspect_ratio。
  • 不要重复描述起始帧已经呈现的内容;这是浪费图生视频提示词的最常见方式。

生成设置

首尾帧 · 10s · 16:9 · 1 个参考素材

MiniMax H3 动作迁移案例:两个参考角色复制参考视频中的街舞动作多模态参考
10s16:93 个参考素材

角色与动作

街舞动作迁移

只用二十多个英文单词,就能把参考视频中的舞蹈动作迁移给两个图片角色,直观展示了“按顺序指定参考素材”的价值。

查看详情

完整提示词(已验证英文原文)

Have the characters perform street dance following the movements in Video 1. Use Figure 1 and Figure 2 as the character references.

需要提供的素材

  • 图片 1 和图片 2:两个角色各一张干净的全身参考图。
  • 视频 1:要复制的动作,2–15 秒,表演者需完整出现在画面中。

为什么有效

  • 提示词很短,因为信息由参考素材承载:视频提供动作,图片提供身份。
  • 每个素材都按数组位置被明确称呼,不会混淆哪个参考负责哪项信息。
  • 不再追加光照、运镜或风格指令,避免额外要求与动作复制竞争。

可替换变量

  • 角色
  • 源舞蹈动作
  • 环境
  • 表演者数量

限制与注意事项

  • 参考视频总时长需保持在 15 秒以内,每段视频必须为 2–15 秒、23.976–60 FPS。
  • 参考视频必须为 MP4 或 MOV,使用 H.264 或 H.265,单个最大 50 MB,整个 JSON 请求体小于 64 MB。

生成设置

多模态参考 · 10s · 16:9 · 3 个参考素材

MiniMax H3 根据设定图生成的角色出场:黑发编辫的战士在石头废墟中从靴子逐步揭示到全身英雄姿势多模态参考
15s1:11 个参考素材

角色与动作

角色设定图英雄出场

本库社区点赞最高的模板:一张角色设定参考图驱动"靴子—面孔—全身"的电影式出场,适用于任何原创角色设计。

查看详情

完整提示词(已验证英文原文)

Use @[char ref] as the sole character reference. Preserve the exact identity, face, body proportions, hairstyle, outfit, colors, materials and overall silhouette of the character throughout the entire video. Do not redesign, simplify or replace any defining visual features. Create a cinematic character introduction focused on presence, silhouette, attitude and controlled motion. 0–4s Begin with a close shot of a defining lower-body or detail element such as boots, shoes, feet, hands, clothing hem or an important accessory. The character enters frame or settles into position. The camera slowly tracks upward while hair, clothing and secondary elements move naturally in the wind or environment. 4–8s Reveal more of the body with a medium or medium-wide shot from the back, side or three-quarter angle. The character stands in a calm, composed way inside the environment. The camera makes a smooth orbit, arc or lateral move to gradually reveal the character’s face and silhouette. 8–12s Move into a tight cinematic portrait or upper-body shot. The character performs one subtle signature action that fits their personality, such as lifting the chin, turning the head, adjusting clothing, brushing hair aside, opening a hand, looking toward camera, or shifting posture. Keep the motion minimal and intentional. The expression should match the character’s vibe. 12–15s End with a strong full-body hero shot that clearly presents the entire design and silhouette. Use a low-angle, eye-level or slightly dramatic framing depending on the character’s personality. The character settles into a natural final pose and holds it confidently for a clean final reveal. VISUAL DIRECTION Premium cinematic presentation. Match the visual medium and rendering style of @[char ref]. Emphasize clean silhouette, elegant staging, subtle secondary motion, believable hair and cloth movement, strong composition, atmospheric depth and polished lighting. The scene should feel like a high-end anime, game or film character introduction. CAMERA Use a clear progression from detail reveal to partial reveal to face reveal to full-body hero reveal. Camera movement should be smooth, controlled and intentional. Avoid chaotic motion. ENVIRONMENT Place the character in a fitting environment that supports their identity and mood. The background should enhance the character without distracting from them.

需要提供的素材

  • 图片 1:你的角色参考——设定图或干净的全身渲染;视频从中继承身份、服装、材质和渲染风格。

为什么有效

  • 揭示阶梯(细节 → 局部 → 面孔 → 全身英雄镜头)是固定的镜头语法,模型可以把变化预算全部花在角色上,而不是镜头设计上。
  • "Match the visual medium and rendering style of the reference" 让同一份提示词对动漫、游戏渲染和影视真人角色都成立。
  • 第三段里唯一的"标志性动作"槽位是个性所在——一个刻意保持克制的小动作。

可替换变量

  • 角色设定参考图
  • 标志性动作
  • 环境氛围
  • 最终姿势构图

限制与注意事项

  • 原帖用 @[char ref] 称呼参考素材;在 EvoLink 的路由上素材按数组位置寻址——写"图片 1"即对应 image_urls[0]。
  • 一个角色、一个环境:这个模板刻意从不切换地点,这正是轮廓保持稳定的原因。

生成设置

多模态参考 · 15s · 1:1 · 1 个参考素材

MiniMax H3 电影感与视觉特效提示词

MiniMax H3 生成的科幻预告片:孤独人影站在巨大环形宇宙传送门前,标题从黑暗中逐渐显现多模态参考
15s16:92 个参考素材

电影感与视觉特效

科幻悬疑预告片

用两张参考图构建电影感预告片:一张定氛围,一张锁定主角,再用提示词控制一次连续推镜、屏幕标题和同步音效。

查看详情

完整提示词(已验证英文原文)

Realistic cinematic look, high-contrast lighting, and a tight pace. Use Figure 1 as the overall atmosphere and style reference, and Figure 2 as the protagonist reference. Shot 1 — Ultra-wide establishing shot. A huge circular cosmic gateway nearly fills the frame. The person is only a tiny figure seen from behind before the gateway, positioned toward the lower right. The ground is wet and reflective, and the center of the gateway is pitch black. The camera slowly pushes forward. A large title fades in from the edge of the darkness, blurred at first and then sharp: "THE STARS WERE LISTENING". Use an extremely condensed, heavy, all-caps typeface in dark red mixed with rust red, with subtle grain and misted edges. Audio: a deep low-frequency pulse, faint metallic vibrations in the distance, and a soft hit as the text becomes sharp. → Hard cut.

需要提供的素材

  • 图片 1:氛围与风格参考,包含希望镜头继承的环境、色板和调色。
  • 图片 2:主角参考,人脸和轮廓需足够清晰,以便人物在远景中仍可识别。
  • 最终要显示在画面中的标题文字;H3 会尝试渲染你写下的文字,因此要简短并确保拼写准确。

为什么有效

  • 每个参考素材只负责一件事:图片 1 控制氛围,图片 2 控制人物,避免模型猜测风格应继承自哪一张图。
  • 整个镜头只有一次连续推进和一个主事件(标题逐渐清晰),符合 15 秒视频可承载的信息量。
  • 音频被拆成低频脉冲、远处金属振动和文字变清晰时的轻击三层,比“电影感配乐”更可执行。

可替换变量

  • 标题文字和字体处理
  • 建立镜头中的传送门或地标
  • 标题颜色
  • 背景音
  • 结尾剪辑方式

限制与注意事项

  • 按数组顺序称呼素材,例如“图片 1”、“图片 2”,并与 image_urls 顺序一致。`@image1` 不属于当前 API 契约。
  • 这是一个故意停在硬切处的起步模板,方便你继续扩展第二个镜头。

生成设置

多模态参考 · 15s · 16:9 · 2 个参考素材

MiniMax H3 文生视频案例:黄昏厨房的手持镜头中,一只手绘发光生物在道具间移动文生视频
15s16:9无参考素材

电影感与视觉特效

厨房里的发光生物

纯文生视频模板,把黄昏厨房实拍与手绘发光动画结合,重点描述手持镜头的不完美感和必须排除的恐怖元素。

查看详情

完整提示词(已验证英文原文)

15-second, 16:9 landscape video. Blend live-action footage of a small kitchen at dusk with hand-drawn glowing animation. The last light of sunset lingers by the window. The lived-in kitchen contains an old wooden table, a half-washed mug, a slightly fogged glass bottle, and a hanging dishcloth. Give the footage subtle one-handed smartphone shake, hesitant close-range focusing, exposure fluctuations caused by backlight, and slightly coarse noise in the shadows. It should not look carefully arranged like an advertisement; instead, it should feel like someone hurriedly captured an unbelievable event at home. Do not show huge eyes, gaping mouths, fangs, threatening or lunging movements, sudden black frames, or jump scares. Use only kitchen room tone, cloth rubbing, the soft clink of a mug, water dripping from the faucet, the camera operator's footsteps and quiet breathing, plus gentle electronic sounds and tiny calls from the hand-drawn creature.

需要提供的素材

  • 无需上传素材,该路由只需提示词。
  • 时长和画面比例应在请求参数中设置,不要只写在提示词中。

为什么有效

  • 明确写出单手抖动、犹豫对焦、逆光曝光波动和暗部噪点,才能呈现“在家中仓促拍到”的质感,而不是广告片。
  • 简短的禁止列表排除巨大眼睛、尖牙、扑击和突然惊吓,防止可爱生物滑向恐怖风。
  • 音频指令分别指定房间底噪、布料摩擦、杯子碰撞、水滴、脚步、呼吸和生物发声。

可替换变量

  • 房间和时间
  • 生物设计与行为
  • 桌面道具
  • 镜头不完美特征
  • 音频层次

限制与注意事项

  • 提示词虽然写了时长和构图,API 请求仍需提供 duration 和 aspect_ratio。
  • 负面指令适合使用简短明确的列表;过多禁止项会挤占动作描述的空间。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 由首帧生成的动漫竞速:两台悬浮摩托拖着青色与深红光轨,在湿滑山道发夹弯中交换位置首尾帧
15s16:91 个参考素材

电影感与视觉特效

悬浮摩托动漫竞速

首帧驱动的运动番竞速,带 CRITICAL ENTITY LOCK:恰好两位具名车手骑乘两台完整设定的赛车,经历计时的发夹弯、尾流对拉和终点摄影判定。

查看详情

完整提示词(已验证英文原文)

Cinematic Anime Video Scene Generate a 15-second horizontal 16:9 original high-speed hover-bike racing anime video from the provided first frame. CRITICAL ENTITY LOCK: There must be exactly 2 racers and 2 bikes in the entire video: RENJI on VALKYRIE-01 (cyan/black drift bike) and ELENA on AERO-X (crimson/white draft bike). Do not add extra racers, drone support vehicles, spectators, or traffic. Maintain total visual consistency for both bikes, helmet visors, suit patterns, repulsor spark colors, and bike liveries throughout the sequence. Entity identity: VALKYRIE-01: Matte-black and cyan angular hover-bike, exposed repulsor pads, lateral drift brakes, blue plasma exhaust trails, ridden by Renji (cyan trim suit). AERO-X: Pearl-white and neon-crimson aerodynamic hover-bike, enclosed canopy, crimson energy draft aura, white-hot central booster, ridden by Elena (crimson/gold visor suit). Video style: High-budget modern sports anime, sakuga-level velocity animation, crisp line art, vibrant neon lighting contrast, high-speed camera tracking, hyper-realistic friction and energy particle effects. Set on a wet downhill mountain pass at dawn. Camera and pacing: Continuous forward velocity, zero slow-motion interruptions: 0.0s - 3.0s: High-speed rear-tracking shot diving into the first downhill hairpin curve; instant drift initiation. 3.0s - 7.5s: Tight side-parallel tracking shot as bikes navigate rock debris and trade positions through S-curves. 7.5s - 11.5s: Close camera lock on the draft-slingshot maneuver; high-energy particle displacement as booster ignition occurs. 11.5s - 15.0s: Low-angle front-facing camera lock on the final sprint to the finish line bridge, ending on a hyper-speed photo-finish freeze. Action timing: 0.0s - 1.5s: Sequence begins at speed. VALKYRIE-01 leads downhill; AERO-X locks onto its rear bumper. Anti-gravity repulsors spray road water and blue sparks into the frame. 1.5s - 4.0s: First sharp hairpin. VALKYRIE-01 deploys lateral drift airbrakes with a burst of blue thruster fire, sliding sideways at 300 km/h. AERO-X stays glued inside its slipstream aura. 4.0s - 7.0s: Mountain debris hazard. VALKYRIE-01 hops over a boulder using a repulsor burst. AERO-X ducks under it, scraping the neon magenta guardrail in a cloud of friction sparks. 7.0s - 10.0s: S-Curve exchange. Bikes lean side-by-side; their repulsor fields collide, creating a bright electrical shockwave. ELENA pulls the overdrive lever; AERO-X's rear fins extend. 10.0s - 13.0s: Slingshot maneuver. AERO-X bursts out of VALKYRIE-01's draft, igniting its central white plasma booster. Both bikes roar down the final straightaway side-by-side. 13.0s - 15.0s: Final sprint toward the finish light gate. Water sprays violently behind them. Both nose cones cross the finish line simultaneously in a flash of light. Final freeze frame. Motion quality: Fluid 2D animation, extreme speed-line integration, stable bike geometry, flawless vehicle reflection rendering, zero limb or body clipping, high-frame-rate kinetic realism. Environment: Wet mountain pass asphalt, sheer cliff walls, neon cyan and magenta guardrail lights, early dawn sky with pink/purple clouds, water spray, floating spark particles. Final output: 15 seconds, horizontal 16:9, original high-budget sports racing anime, exactly 2 racers, relentless kinetic pacing, dynamic cinematography, no subtitles, no watermarks, no logos.

需要提供的素材

  • 起始帧:用你的画风绘制的两位车手与赛车静帧——整段序列从这张图延展而来。

为什么有效

  • 用"清点数量"实现实体锁定("exactly 2 racers and 2 bikes… no spectators, no traffic"),杜绝竞速提示词最常见的人群膨胀问题。
  • 两台机体各自拥有涂装、物理特性和车手战衣等具名身份,模型在高速中也能把它们区分开。
  • 镜头与动作各自拥有相互引用的独立时间轴,既维持不间断的节奏,又从不同时要求两种运镜。

可替换变量

  • 起始帧画风
  • 赛车身份与涂装
  • 赛道障碍
  • 终点画面编排

限制与注意事项

  • 这是首帧路由:起始图承载画风,且图生视频不接受 aspect_ratio 参数——画幅由这张图决定。

生成设置

首尾帧 · 15s · 16:9 · 1 个参考素材

MiniMax H3 原生音频与对白提示词

MiniMax H3 对白替换案例:角色的台词被替换为参考音频中的新台词多模态参考
10s16:92 个参考素材

原生音频与对白

对白与表演替换

替换现有视频中的一句台词:准确引用旧台词和新台词,并允许表演只在匹配新对白所需的范围内调整。

查看详情

完整提示词(已验证英文原文)

Replace the girl's line in Video 1, "We can't be together. It's not that we don't love each other; we truly can't make it to the end," with the line from Audio 1: "Don't go, okay? This time, let's not let go of each other." Slightly adjust the corresponding performance.

需要提供的素材

  • 视频 1:包含待替换台词的原片。
  • 音频 1:新台词,WAV 或 MP3,最大 15 MB、最长 15 秒。
  • 在提示词中逐字写出新旧两句台词。

为什么有效

  • 引用待替换台词,可准确指定视频中要处理的片段,比“中间那句对白”更可靠。
  • 引用新台词后,口型同步有了明确目标,而不需只从音频猜测。
  • “轻微调整对应表演”给出了有边界的修改授权,让其他镜头内容尽可能保留。

可替换变量

  • 源视频
  • 待替换台词
  • 新台词与音色
  • 允许表演变化的程度

限制与注意事项

  • 音频不能是唯一参考类型,必须与图片或视频一起提交。
  • 每次请求的参考音频和视频总时长均不超过 15 秒。
  • 明确写出要保持不变的内容;如果只谈台词,构图、服装和背景都可能偏移。

生成设置

多模态参考 · 10s · 16:9 · 2 个参考素材

MiniMax H3 音色参考案例:角色使用参考音频中的音色说出指定台词多模态参考
10s16:92 个参考素材

原生音频与对白

“随风而行”音色参考

最小可用的音频参考提示词:写明要说的台词,再用一段音频规定由谁的音色说出它。

查看详情

完整提示词(已验证英文原文)

Character dialogue: "Follow the wind, live free. Leave worries behind, enjoy the moment." Use Audio 1 as the voice-timbre reference.

需要提供的素材

  • 视频 1:要说出台词的角色。
  • 音频 1:2–15 秒的干净目标音色样本,尽量不带背景音乐。
  • 逐字写出要说的台词。

为什么有效

  • 台词使用引号准确写出,时序和口型同步就有具体目标。
  • 音频 1 只被赋予“音色”这一项职责,而不是被模糊当作整体配乐。
  • 没有额外指定构图和表演,因此这些信息继续由参考视频控制。

可替换变量

  • 台词
  • 音色参考
  • 角色视频
  • 语速

限制与注意事项

  • 音色参考只传递音色,不自动传递口音、情绪和节奏;如果这些很重要,需另写在提示词中。
  • 音乐或多人重叠说话会降低效果,应提供隔离的单人语音。

生成设置

多模态参考 · 10s · 16:9 · 2 个参考素材

MiniMax H3 由角色与音频参考生成的蒙太奇:白发有角的角色在五个环境间随节拍快速切换多模态参考
15s1:12 个参考素材

原生音频与对白

音频同步环境蒙太奇

双参考蒙太奇:角色图锁定身份,音频片段决定剪辑——五个环境、每个环境六个快切镜头,每次剪切都落在音轨重音上。

查看详情

完整提示词(已验证英文原文)

Use @[char ref] as the strict character reference and @[audio ref] as the timing, rhythm and editing reference. Keep the character’s exact identity, proportions, hairstyle, outfit, colors and overall style consistent throughout. Create a 15-second cinematic burst-cut video showcasing the character across 5 different environments that naturally fit their design, vibe and world. AUDIO SYNC Synchronize the entire edit to @[audio ref]. Cuts, camera accents, transitions and environment changes should land precisely on strong beats, half-beats and musical accents. Let audio1 control the pacing and intensity of the montage. STRUCTURE - 5 environments total - 3 seconds per environment - 6 burst-cut shots per environment - 30 shots total Each environment must be clearly different in atmosphere, lighting, scale and visual language. Show each environment through rapid cinematic angles: wide establishing shots, aerials, low angles, side views, tracking shots, close environmental details, medium shots and hero frames. Every cut must reveal a new angle, distance, composition or spatial relationship. Avoid repeated framing. Mix static shots, push-ins, pull-backs, tracking, orbit and crane-like movement. Keep character movement subtle and natural. The focus is environmental variety, cinematic framing and tight synchronization with audio1. Hard constraints: - exactly 5 environments - exactly 6 shots per environment - exactly 30 shots total - environment changes must follow audio1’s musical phrasing - cuts and motion accents synchronized to audio1 - no outfit changes - no character duplication - no morphing - no text or UI - no blurry unreadable frames - maintain strict character consistency

需要提供的素材

  • 图片 1:每个镜头都必须保持身份一致的角色。
  • 音频 1:掌控节奏的音轨——剪切、转场和环境切换都跟随它的节拍。

为什么有效

  • 它在模态之间干净分工:图片回答"是谁",音频回答"何时"——两者都不与文字争抢。
  • 精确算术(5 个环境 × 6 个镜头 = 30 次剪切)被写成硬约束,把模糊的蒙太奇变成可计数的结构。
  • "角色动作保持克制"把全部能量推向镜头变化,而蒙太奇消化镜头变化的能力远强于动作变化。

可替换变量

  • 角色参考
  • 音轨及其乐句结构
  • 五个环境
  • 镜头类型组合

限制与注意事项

  • 原帖写作 @[char ref] 和 @[audio ref];在 EvoLink 上按数组位置称呼——图片 1、音频 1——并记住音频永远不能是唯一的参考类型。
  • 该路由的参考音频片段必须为 2–15 秒。
  • 已发布片段实测为 8:9,而 API 并不接受这一比例——想要同样的近方形画面就请求 1:1。

生成设置

多模态参考 · 15s · 1:1 · 2 个参考素材

MiniMax H3 游戏与界面提示词

MiniMax H3 由九张参考生成的 UI 动画:极简生物图鉴中,光标选中草地上毛茸茸的有角生物 MEADOW CROWN多模态参考
15s16:99 个参考素材

游戏与界面

生物图鉴 UI 演示

九张参考图映射到明确角色——一张 UI 布局底板加八张生物卡片——以锁定机位动画演绎图鉴:光标逐条点击词条,最后一只生物把光标吞掉。

查看详情

完整提示词(已验证英文原文)

Use Image 1 as the exact UI/layout/style reference for the creature encyclopedia screen. Use Images 2–9 as the exact creature references. Map them like this: Image 2 = card A = LUMI HARE Image 3 = card B = CLOUD WISP Image 4 = card C = EMBER FENNEC Image 5 = card D = TIDE BEHEMOTH Image 6 = card E = PETAL VULPIN Image 7 = card F = ORCHARD EYE Image 8 = card G = MEADOW CROWN Image 9 = card H = FROST GLIDER Create a 15-second 16:9 video. Keep the camera locked. Keep the interface, layout, typography, panels, icons and overall composition stable, elegant and readable. The UI should feel like a modern minimal digital creature encyclopedia, similar to a sleek pokedex. No scene cuts, no extra text, no extra buttons, no UI distortion. Sequence: 0–2.5s: Cursor clicks card A. Main creature becomes LUMI HARE. Title changes to “LUMI HARE”. Creature blinks and rotates slightly. 2.5–5s: Cursor clicks card C. Main creature becomes EMBER FENNEC. Title changes to “EMBER FENNEC”. Cursor drags to rotate it left and right. 5–7.5s: Cursor clicks card E. Main creature becomes PETAL VULPIN. Title changes to “PETAL VULPIN”. Cursor pokes it a few times. It reacts, annoyed. 7.5–10s: Cursor clicks card G. Main creature becomes MEADOW CROWN. Title changes to “MEADOW CROWN”. Cursor taps near the face/horns. It recoils slightly. 10–12s: Cursor clicks card D. Main creature becomes TIDE BEHEMOTH. Title changes to “TIDE BEHEMOTH”. Cursor keeps poking it. 12–15s: TIDE BEHEMOTH gets angry, opens its mouth very wide, lunges forward, and swallows the cursor. Then it returns to idle. Title stays “TIDE BEHEMOTH”. Rules: - When a card is selected, both the main creature and the main title must update. - Only animate cursor, selection state, title change, and the selected creature. - Only one cursor. - Keep motion subtle and clean until the final swallow. - No cuts, no camera move, no UI distortion, no extra text. Audio: soft UI click sounds, subtle hover sounds, tiny creature reaction sounds, then a sharper aggressive creature sound and one comedic swallow gulp at the end.

需要提供的素材

  • 图片 1:决定字体、面板与构图的图鉴 UI 布局。
  • 图片 2–9:每张卡片一只生物,均按提示词中的卡片映射表命名。

为什么有效

  • 图片到卡片的映射表(Image 2 = card A = LUMI HARE…)是可能写出的最直白的角色分配——模型永远不用猜哪个素材对应什么。
  • 机位锁定且只允许四样东西动(光标、选中态、标题、当前生物),把 UI 动画的失败面收缩到近乎为零。
  • 喜剧节拍(生物吞掉光标)被排在最后,演示在包袱抖出之前始终保持干净。

可替换变量

  • UI 风格底板
  • 八只生物及其命名
  • 交互脚本
  • 音效组合

限制与注意事项

  • 九张图片是该路由单一类型的上限——而 12 个文件的总数上限意味着不能再叠加三段视频和三段音频。
  • UI 文字的稳定依赖 "no camera move, no UI distortion";解放机位就会重新出现扭曲面板。

生成设置

多模态参考 · 15s · 16:9 · 9 个参考素材

MiniMax H3 游戏实况风片段:第一人称视角沿步枪瞄具推进,穿过烟雾弥漫的军事基地,带通用 FPS HUD文生视频
15s16:9无参考素材

游戏与界面

FPS 游戏实况模拟

一段读起来像真实录屏的第一人称射击序列:玩家操控式的镜头语法、一套完整设定的通用 HUD,以及写成战术而非编舞的节奏。

查看详情

完整提示词(已验证英文原文)

Camera: First-person perspective at eye level with authentic handheld player movement, as if recorded directly from a modern AAA military shooter. The player carries a highly detailed assault rifle with realistic animations, visible hands, tactical gloves, dynamic reload mechanics, and weapon sway. **Opening Action:** The video immediately begins with the player already aiming down a roadway inside a modern military base. Multiple enemy soldiers are visible in the distance near sandbags, barricades, and military vehicles. The player carefully tracks one target, making small aim corrections while maintaining ADS (aim down sights). Fire several controlled bursts immediately at the visible enemies, producing realistic muzzle flashes, shell casings ejecting, smoke, recoil, hit reactions, and dust impacts around the targets. Continue firing in multiple short bursts while adjusting aim between enemies, simulating authentic FPS gameplay rather than scripted animation. **Movement:** After the opening firefight, lower slightly from ADS and begin advancing cautiously along the road beside concrete barriers, Hesco walls, and parked military vehicles. Frequently check left and right corners, briefly stop to reacquire targets, then raise the weapon and fire additional controlled bursts whenever enemies appear ahead. Continue pushing forward with deliberate player-controlled movement, using cover naturally and maintaining believable tactical pacing. **Environment:** Large modern military base with guard towers, armored vehicles, shipping containers, blast barriers, damaged buildings, smoke plumes, burning debris, scattered shell casings, dust clouds, and atmospheric battlefield haze. Cool natural daylight mixed with smoke and orange firelight creates a cinematic battlefield atmosphere. **Camera Motion:** Authentic player-controlled movement with subtle head bob, weapon sway, natural mouse-look adjustments, small left-right corrections while aiming, realistic recoil, smooth tracking of moving targets, brief pauses before shooting, and fluid forward progression. Avoid cinematic camera moves—everything should feel like genuine live gameplay captured by a skilled player. **Visual Quality:** Ultra-photorealistic, AAA game graphics with realistic PBR materials, detailed weapon models, physically accurate lighting, volumetric smoke, dynamic particle effects, crisp textures, realistic bullet impacts, muzzle flash illumination, motion blur only during rapid movement, and high-end military shooter presentation. **Gameplay UI:** Display a realistic modern FPS HUD inspired by games like PUBG, Battlefield, or Call of Duty (without copying exact copyrighted assets). Include: * Central dynamic crosshair or reticle * Ammo counter with magazine and reserve ammunition * Fire mode indicator * Compass at the top * Squad/team status panel * Mini-map in the upper corner * Health bar * Tactical equipment icons (grenades, medkit) * Hit markers when bullets connect * Directional damage indicators * Kill notification feed * Objective marker in the distance * Subtle interaction prompts and realistic HUD animations The HUD should feel polished, modern, and fully integrated into the gameplay, enhancing the illusion of authentic recorded footage from a contemporary military FPS. #MiniMaxH3

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • 真实感目标是"由高手玩家录制",因此头部晃动、瞄准修正和停顿后开火的节奏都被写成镜头行为。
  • HUD 被逐项列到命中标记和击杀信息流的级别,同时明确保持通用——密度足以像真游戏,安全到可以发布。
  • "Avoid cinematic camera moves" 是关键的反转:多数提示词追求的东西恰恰会毁掉这一条。

可替换变量

  • 环境与阵营风格
  • HUD 元素组合
  • 交火节奏
  • 天气与光线

限制与注意事项

  • 提示词把 HUD 保持在"借鉴而不复制"的边界内——保留这个条款;克隆某款具体游戏的 HUD 是另一项(且风险更高的)任务。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 Vlog 与自拍镜头提示词

MiniMax H3 自拍风格片段:在森林里自拍的女子把手机转向树木之间一架冒烟的坠毁 UFO文生视频
15s1:1无参考素材

Vlog 与自拍镜头

自拍镜头 UFO 发现

手机实拍质感的展示样本:带自动对焦呼吸和果冻效应的手持自拍画面、成稿的日语台词,以及从面孔翻转到坠毁 UFO 的镜头切换。

查看详情

完整提示词(已验证英文原文)

Ultra photorealistic live-action captured on an iPhone 17. Authentic handheld selfie footage with premium cinematic documentary color grading, realistic HDR, deep green foliage, warm sunlight, subtle teal shadows, natural skin tones, gentle filmic contrast, rolling shutter, autofocus breathing, slight motion blur, and natural handheld shake. A lush forest in daytime with dense trees, wild plants, an uneven dirt trail, scattered leaves, soft sunlight through the canopy, and a gentle breeze. The atmosphere is quiet and slightly unsettling. A cute Japanese woman in her early twenties wearing a stylish bikini walks through the forest while recording herself in selfie mode. She suddenly notices something ahead, looks shocked, turns the camera, and points into the distance. A large crashed UFO is partially embedded in the forest floor. Its metallic hull is badly damaged with broken panels, scorch marks, exposed internal structures, thick gray smoke, and occasional sparks. She says in Japanese: 「ちょっと待って! あそこ見て! UFOじゃない!? 完全に墜落してるんだけど! 煙まで出てる! やばい、本物かもしれない! ちょっと近づいてみる!」 She alternates between filming herself and the UFO while continuing to point at it. Continuous single take. Natural walking movement, realistic hand tremors, slight framing imperfections, quick pans, and autofocus shifts between her face and the UFO. Natural sunlight creates cinematic highlights, soft shadows, realistic reflections on the UFO, subtle volumetric light, and realistic smoke. Audio: footsteps on leaves, gentle wind, birds becoming quieter near the crash site, creaking branches, faint electrical crackling from the UFO, and distant eerie unidentified animal calls echoing through the forest. Negative: no blood, no visible aliens, no monsters, no horror creature reveal, no excessive explosions, no CGI, no cartoon style, no text, no subtitles, no watermark, no logo.

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • 手机瑕疵被逐一列举(对焦拉风箱、果冻效应、手部颤抖、构图不完美)——真实感来自被点名的缺陷,而不是 "realistic" 这个词本身。
  • 台词以日语逐字引用,H3 的原生音频因此能生成口型匹配的真实语音,而非乱语。
  • 音频按画内声源分层——脚步声、风声、靠近坠机点逐渐安静的鸟叫、电流噼啪——不加配乐,守住"捡到的手机录像"的错觉。

可替换变量

  • 台词内容与语言
  • 被发现的物体
  • 森林或城市场景
  • 服装与角色造型

限制与注意事项

  • 负面列表(不出现外星人、不揭示恐怖生物)承担实际功能:它把片子留在预告氛围里,阻止模型自行升级场面。
  • 已发布片段为 1:1——画幅要通过比例参数设置,同时在文字里保留自拍与目标物交替的构图。

生成设置

文生视频 · 15s · 1:1 · 无参考素材

MiniMax H3 纪录片片段:街头摄影师取景拍下咖啡馆外的老人与梗犬,随后向观众展示拍到的照片文生视频
15s16:9无参考素材

Vlog 与自拍镜头

街头摄影师瞬间

一个紧凑的纪录片节拍,镜头里还有一台相机:摄影师取景抓拍街头一幕,按下快门,再把相机转向观众展示她刚拍到的照片。

查看详情

完整提示词(已验证英文原文)

A young Western female street photographer walks through a lively downtown street and notices an elderly man sitting outside a café with his small dog. She carefully composes the candid moment through her camera, captures the photo, then turns the camera toward the viewer to proudly show the shot she just took. She smiles, says “Look at that,” then continues walking through the city. Ultra-photorealistic visuals, natural handheld documentary movement, realistic camera interaction, authentic facial expressions, accurate hand movements, realistic dog behavior, natural daylight, cinematic depth of field, continuous character consistency, immersive city ambience, premium documentary realism.

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • 展示照片这个节拍迫使 H3 在视频内部渲染一张连贯的静态照片——一次低调的能力演示,读起来却只是一个自然的动作。
  • 交互链条被完整写明(取景 → 拍摄 → 翻转 → "Look at that" → 继续行走),片段在一个镜头里拥有完整弧线。
  • 拍摄对象按角色而非身份描述——老人、小狗、咖啡馆——街头场景保持通用与安全。

可替换变量

  • 抓拍对象
  • 城市与光线
  • 口播台词
  • 相机道具类型

限制与注意事项

  • 相机里的照片必须与它拍摄的场景一致;如果更换拍摄对象,两处描述要一起改。

生成设置

文生视频 · 15s · 16:9 · 无参考素材

MiniMax H3 旅行 vlog:女子拉开窗帘走上海滨露台,随后切到自拍模式说早安,身后是大海文生视频
15s9:16无参考素材

Vlog 与自拍镜头

海滨清晨 vlog 弧线

逐秒编排的清晨 vlog,带一次刻意的模式切换:前五秒是电影感第三人称,之后角色才开始自拍——"vlog"正是从那一刻开始的。

查看详情

完整提示词(已验证英文原文)

Create a 15-second ultra-realistic cinematic lifestyle vlog video, vertical 9:16, featuring the same young woman throughout the entire video. Preserve her facial identity, facial proportions, hairstyle, skin tone and overall appearance consistently in every shot. She wears the same outfit throughout: fitted white V-neck T-shirt with a small subtle logo, blue denim jeans, natural makeup, long softly wavy brown hair. 0:00–0:01 — Wake-up: Close-up inside a beautiful bright bedroom. The woman is lying comfortably on the bed, slowly wakes up, stretches naturally and opens her eyes. She is NOT filming a vlog yet and does not hold a phone or camera. Soft morning sunlight enters through the curtains. 0:01–0:02 — Gets up: Medium shot. She sits up on the bed, smiles softly, fixes her hair and gets ready to start her morning. Natural, effortless movement. 0:02–0:03 — Walks to window: She walks toward the large glass balcony door/window. Camera follows her naturally from behind/side. 0:03–0:04 — Seaside reveal: She opens the curtains/door and looks outside. Reveal a breathtaking blue ocean, coastal hills, flowers, balcony and beautiful morning sunlight. She smiles happily while taking in the view. 0:04–0:05 — Steps outside: She walks out onto the seaside terrace. Gentle ocean breeze moves her hair naturally. Wide cinematic shot showing the beautiful surroundings. 0:05–0:06 — VLOG START: Only now she starts filming herself in handheld selfie-vlog style. She looks into the camera with a bright natural smile and says: “Good morning!” 0:06–0:07 — Show the view: She turns the camera away from herself and slowly pans across the stunning ocean, coastal mountains, flowers and terrace. Smooth handheld vlog movement. 0:07–0:08 — Back to selfie: Selfie shot. She looks into the camera and happily says: “This place is just perfect!” 0:08–0:09 — Location reveal: Wide cinematic shot of the cozy seaside terrace with wooden table, chairs, plants and flowers overlooking the ocean. 0:09–0:10 — Walk to table: Medium tracking shot as she walks toward the table, enjoying the view. Her hair and T-shirt move gently in the sea breeze. 0:10–0:11 — Sit and relax: She sits at the seaside table, smiling peacefully and enjoying the ocean view. A refreshing orange-colored juice is placed on the table. 0:11–0:12 — Juice close-up: Cinematic close-up of her hand picking up the glass of fresh orange juice. Beautiful ocean bokeh in the background, natural sunlight reflecting through the glass. 0:12–0:13 — Vlog toast: Selfie shot. She raises the juice toward the camera with a cheerful smile and says: “Cheers to good days!” 0:13–0:14 — Happy close-up: Beautiful close-up of her smiling naturally at the camera, ocean and warm sunlight softly blurred behind her. 0:14–0:15 — Ending: Camera moves from her toward the sparkling ocean and peaceful coastal landscape. Warm sunlight, gentle waves and a relaxing cinematic ending. Overall Style Ultra-realistic, cinematic travel vlog, natural handheld camera movement, realistic human motion, smooth transitions, soft morning sunlight, realistic ocean waves, gentle wind in hair and clothes, beautiful coastal atmosphere, premium lifestyle aesthetic, natural expressions, authentic vlog feeling, shallow depth of field, cinematic composition, realistic skin texture, high detail, 4K quality.

需要提供的素材

  • 无需上传素材,文生视频路由只需提示词。

为什么有效

  • 起床段落里"她还没有开始拍摄"的注记,防住了 vlog 最常见的穿帮——手机在 vlog 开始前就出现在画面里。
  • 十五个一秒节拍在自拍镜头与风景空镜之间交替,与真实旅行 vlog 的语法完全一致。
  • 三句简短的引用台词("Good morning!")为原生音频提供自然的锚点,而无需撑起一段独白。

可替换变量

  • 地点揭示
  • 三句口播台词
  • 服装与身份锁定
  • 饮品道具

限制与注意事项

  • 提示词要求 9:16,请求里也应当填这个比例;已发布片段被重新编码为约 3:2——又一个说明画面方向该写进请求参数、而非散文描述的理由。

生成设置

文生视频 · 15s · 9:16 · 无参考素材

MiniMax H3 编辑与转换提示词

MiniMax H3 视频编辑案例:源视频中的报纸、椅子、墨镜和燃烧车辆被替换或移除多模态参考
10s16:91 个参考素材

编辑与转换

多元素场景编辑

对一段源视频执行六个独立修改,全部写成平铺的“目标 → 结果”指令,完全不重述场景。

查看详情

完整提示词(已验证英文原文)

Replace the newspaper in the reference video with a green-covered book; change the chair the character is sitting on to a red sofa; remove the sunglasses worn by the character to retain a clear face; remove the car burning effect to keep the vehicle in a normal state; change the photo the character takes out of his arms to a small black book; and add a tree on the left side of the screen

需要提供的素材

  • 视频 1:待编辑片段,2–15 秒,MP4 或 MOV,H.264 或 H.265,最大 50 MB。
  • 一份编辑列表,每条都指明要改什么、改成什么。

为什么有效

  • 每条指令都包含“目标 + 结果”,例如报纸变成绿色封面的书,不留猜测空间。
  • 删除操作同时写明期望终态,例如“去掉墨镜以保留清晰面孔”,防止原位留下空洞或畸变。
  • 不描述场景,让模型把源视频当作事实,只应用明确差异。

可替换变量

  • 编辑项数量
  • 替换的物体
  • 移除的效果
  • 新增元素

限制与注意事项

  • EvoLink 对外提供三条 H3 路由;编辑类任务通过参考生视频路由执行,源视频作为视频 1。
  • 参考视频时长会计费,上传前应先剪掉不需要的片段。
  • 修改项附近如果有必须保留的内容,要明确写出;未声明区域都可能被模型改动。

生成设置

多模态参考 · 10s · 16:9 · 1 个参考素材

MiniMax H3 舞台案例:两位魔术师在烟雾中交换西装颜色,幕布从红色转为蓝色首尾帧
7s16:91 个参考素材

编辑与转换

魔术师服装互换

用一场舞台魔术测试指令遵循:两套服装互换,一处细节必须不变,背景则按时完成颜色过渡。

查看详情

完整提示词(已验证英文原文)

Two magicians stand onstage facing the audience and perform a "swap" trick. They wave their wands at the same time, and a cloud of smoke rises. When it clears, their suit colors have switched: the person on the left wears a white suit, and the person on the right now wears a black suit, while both magicians' glove colors remain unchanged. They bow to thank the audience. The red curtain behind them closes, transitioning from deep red to deep blue.

需要提供的素材

  • 一张包含两位表演者、服装和舞台的起始图。
  • 需要互换的内容要有清晰的前后状态。

为什么有效

  • 使用烟雾事件遮挡交换,让模型有一个合理时机完成变化,而不是在画面中直接融化。
  • 在写明需要改变的服装时,紧邻写出手套颜色必须保持不变,防止编辑范围扩散。
  • 镜头有明确的最终状态:鞠躬、幕布闭合、颜色落在深蓝色。

可替换变量

  • 表演者与服装
  • 必须保持不变的细节
  • 用于遮挡交换的事件
  • 结尾颜色转换

限制与注意事项

  • 图生视频路由根据输入图片确定画面比例,不接受 aspect_ratio 字段。
  • 模型可能改动任何未声明属性;如果某个细节必须在变化中保留,就要写明。

生成设置

首尾帧 · 7s · 16:9 · 1 个参考素材

在找 MiniMax H3 API?

模型页含按秒计价、三条生成路由与接入文档。

获取 API 接入

MiniMax H3 提示词框架

H3 可以理解按顺序指定的参考素材,并原生生成音频。应围绕实际工作流来写,而不是只用通用公式。

H3 提示词骨架

目标 → 按顺序指定参考素材 → 主体与身份锚点 → 按时间排列的动作节拍 → 摄影机路径 → 音频或对白指令 → 必须保持不变的内容 → 最终状态。按这个顺序检查,就不会漏掉最常被忽略的音频和结尾。

文生视频

所有画面信息都由你定义,因此应把篇幅用在视频能够完成的时间线上。命名主体、排列动作节拍、只给摄影机一条连续路径,再把声音拆成对白、环境音和音乐,不要只写抽象氛围。

首尾帧

首帧和尾帧已经承载外观,文字只需描述两者之间的变化:新增动作、运镜、必须保留的内容以及最终状态。重复描述图片是浪费这条路由的最常见方式。

多模态参考

每个素材只负责一件事,并按数组顺序称呼:图片 1 控制角色,图片 2 控制地点,视频 1 控制运镜,音频 1 控制音色。单次最多 9 张图片、3 段视频和 3 段音频,且文件总数不超过 12 个;音频不能单独使用。

编辑现有视频

不要写成一大段,而是拆成两份列表:“修改”逐行写目标 → 结果,“保留”写出附近必须原样保留的内容。将源视频裁剪到必要片段,作为视频 1 传入参考路由。

音频与对白

逐字引用希望角色说出的台词。只给声音参考一项职责——音色,口音、情绪和节奏需另外写入提示词。对白、环境音和音乐不要抢占同一时段。

MiniMax H3 提示词常见问题

MiniMax H3 和海螺 3 是同一个模型吗?

是。MiniMax H3 也常被称为 Hailuo 3、Hailuo 3.0 或 Hailuo 03。EvoLink 对外统一使用 MiniMax H3 作为产品名称,并提供文生视频、图生视频和参考生视频三条路由。

不写代码也能使用这些提示词吗?

可以。点击“使用此提示词”后,页面会将英文提示词写入 MiniMax H3 Playground,并选中匹配的工作流。你只需在模型页补充参考素材并在浏览器中生成。

MiniMax H3 可以生成多长的视频?

当前三条路由的 duration 为 4–15 的整数,输出为 2K。本页每张卡片显示的时长来自已发布样片的实际测量。

一条提示词最多可以使用多少参考素材?

参考生视频路由单次最多接收 9 张图片、3 段视频和 3 段音频。至少需要一张图片或一段视频,音频不能是唯一参考类型。每段参考视频和音频需为 2–15 秒。

为什么提示词中会写“图片 1”或“视频 1”?

参考路由按 image_urls、video_urls 和 audio_urls 中的位置识别素材。在提示词中称呼“图片 1”或“视频 1”,是为了给具体素材分配职责。`@image1` 语法不属于当前 API 契约。

这些提示词和视频来自哪里?

每张卡片都标明来源。MiniMax 官方案例在获得授权后发布,卡片中的英文提示词是官方中文原稿的已验证英文版。社区案例会链接到创作者在 X 上的原始帖子。

可以通过 API 使用这些提示词吗?

可以。无论粘贴到 Playground,还是作为 API 的 prompt 字段发送,提示词文本都完全一致。MiniMax H3 API 指南详细介绍了请求结构、异步任务、回调和错误处理。

选择一条提示词,立即生成

本页每条提示词都对应 EvoLink 上可用的 MiniMax H3 路由,可发送到 Playground,也可通过统一 API 构建请求。