GPT Image 2.5 Flare & Sunburst が EvoLink で利用可能にGPT Image 2.5 を試す

MiniMax H3プロンプトと動画例

実際のMiniMax H3出力、必要な入力、置き換え可能な変数を備えた検証済みプロンプト40件を掲載。テンプレートをコピーするか、EvoLink Playgroundに直接送信できます。

MiniMax H3はHailuo 3、Hailuo 3.0、Hailuo 03とも呼ばれます。

生成モード

用途

40件中40件のプロンプト

MiniMax H3 ブランド・商品広告のプロンプト

MiniMax H3 で生成した縦型アイウェア広告:シームレスな白いスタジオで未来的なラップアラウンド型メガネを着用した2人のモデルマルチモーダル参照
15s9:16参照素材 3件

ブランド・商品広告

未来的なアイウェア広告

キービジュアル、モデルの顔、製品デザインを3枚の画像で分担するファッション広告です。

詳細を表示

完全なプロンプト(検証済み英語原文)

Generate a vertical screen 9:16 high-end fashion glasses commercial, taking overall reference to the storyboard rhythm, editing speed, white studio texture and cool fashion atmosphere of the given video. The picture is a minimalist white booth, a seamless white background, a strong sense of high-end advertising, and a clean, simple, handsome, avant-garde, international first-line fashion blockbuster texture. Key visual character reference picture 1, two full-body female models, one black female model and one European and American model, maintain their high-end clothing, body posture, white studio light and shadow, fashion show temperament and overall cool attitude. Both of them wear futuristic high-end glasses. The design of the glasses refers to Figure 3, emphasizing the covered curved surface, sharp geometric cat-eye/goggle hybrid outline, mirror reflection, streamlined temples, and the texture of high-end fashion accessories. Please refer to Figure 2 for the appearance details of the two characters.

必要な入力

  • 画像1:キービジュアル。モデルの全身、衣装、スタジオ照明、佇まいを含みます。
  • 画像2:2人の顔のディテール。カットをまたいでも同一人物に保つために使います。
  • 画像3:製品。速いカットでもシルエットと素材感が残るよう、十分に鮮明に撮影します。
  • 動画全体を通して維持できるシームレスなスタジオ背景。

機能する理由

  • 3つの参照に3つの明確な役割を割り当てるのは、参照ルートが本来想定している使い方です。多画像プロンプトが崩れる原因は曖昧さにあります。
  • 製品は名前だけでなく形状で記述されています:包み込む曲面、キャットアイとゴーグルを混ぜた鋭い輪郭、ミラー反射、流線型のテンプル。
  • 単一素材の白いスタジオは背景のばらつきを排除し、モデルの処理能力を製品と人物に集中させます。

置き換え可能な変数

  • 製品カテゴリ
  • モデルのキャスティング
  • スタジオの色
  • 編集の速さ
  • 衣装

制約

  • 公開プロンプトは参照動画のテンポにも言及していますが、このケースで公開されている素材は画像3枚です。この一文はスタイル指示として扱うか、自分の映像を Video 1 として追加してください。
  • 参照リクエストは画像9枚、動画3本、音声3本まで、かつ合計12ファイルが上限です。音声だけを参照素材にすることはできません。

設定

マルチモーダル参照 · 15s · 9:16 · 参照素材 3件

MiniMax H3 で生成した製品映像:黒いスタジオの反射する台の上で回転する高級オーバーイヤーヘッドホンテキストから動画
15s16:9参照素材なし

ブランド・商品広告

高級ヘッドホンの製品ショーケース

15秒を4つの時間ブロックに分け、それぞれにカメラ動作と役割を与えた製品動画です。

詳細を表示

完全なプロンプト(検証済み英語原文)

Create a 15-second luxury cinematic product showcase for premium wireless over-ear headphones. 0–4s: Begin with an extreme macro tracking shot moving across the soft memory-foam ear cushion, fine fabric texture, brushed-metal hinge and precision-machined controls. A narrow light band travels across the surface, revealing realistic materials against a deep black studio background. 4–8s: Pull back into a three-quarter hero view. The headphones rotate slowly above a glossy reflective pedestal. The ear cups pivot naturally while the adjustable headband extends slightly, demonstrating flexible construction and comfort. Maintain exact symmetry, stable geometry and consistent proportions. 8–12s: Transition into an elegant exploded-view reveal. The ear cushion, acoustic driver, internal sound chamber, control ring and outer shell separate smoothly in perfect alignment. Subtle luminous sound waves pulse outward from the driver while the camera performs a restrained side orbit. 12–15s: Every component reconnects seamlessly. The headphones settle into a centered front-facing hero composition as soft rim lighting defines the silhouette. Complete a gentle dolly-in toward the ear cups. Premium technology-commercial finish, controlled reflections, realistic shadows, shallow depth of field, crisp surface detail, stable product shape, no hands, no distortion, no onscreen text.

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。
  • 名前だけでなく、素材と構造で説明できる製品。

機能する理由

  • 時間が 0–4秒、4–8秒、8–12秒、12–15秒 に分割されているため、モデルは願望リストではなくショットリストを受け取ります。
  • 各ブロックでカメラの動きが異なります(マクロ追従、引き、サイドオービット、ドリーイン)。これが「漂う映像」ではなく「編集された映像」に見せます。
  • 最後の一文が禁止リスト(手を写さない、歪ませない、画面文字を出さない)で、製品CGが失敗する典型的な3パターンを潰しています。

置き換え可能な変数

  • 製品
  • 素材と仕上げ
  • 各ブロックの時間配分
  • 背景と照明
  • 分解ビューを入れるかどうか

制約

  • 尺は4〜15秒の整数でリクエストに指定します。プロンプト内の時間ブロックの合計はこれと一致させてください。
  • 時間ブロックは演出指示であり、厳密なタイムラインではありません。15秒なら4つ以下に抑えます。
  • このクリップは作者が720pで公開したものです。EvoLink の H3 ルートは 2K を出力します。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 の飲料広告:炎天下でばてた歩行者が冷えたジュースを飲むと、ボトルの周囲で街並みが瑞々しい緑に変わる様子テキストから動画
15s16:9参照素材なし

ブランド・商品広告

真夏の清涼飲料広告

「問題→解決」の王道構成による飲料CMです。炎天下でばてた人物がひと口飲むと環境全体が一変し、水滴をまとった製品のヒーローショットで締めくくります。

詳細を表示

完全なプロンプト(検証済み英語原文)

A young person walks under the blazing summer sun, looking exhausted and sweating heavily. The road shimmers with heat waves, and everything appears dry and dull. Suddenly, they grab a chilled bottle of premium fruit juice from a cooler and take a refreshing sip. Instantly, the environment transforms—lush green trees bloom, vibrant flowers appear, a cool breeze flows, water splashes through the air, and glowing particles surround the scene. Ice cubes and fresh fruit slices (orange, mango, or according to the flavour) swirl around the bottle in cinematic slow motion. End with a stunning close-up of the juice bottle covered in cold water droplets against a bright, refreshing background. Ultra-realistic, premium commercial, 4K, cinematic lighting, high-detail, smooth camera movements, vibrant colours, luxury beverage advertisement. Tagline ideas: Beat the Heat. Taste the Freshness. Every Sip Brings Life. Refresh Your Day, Naturally. Stay Cool. Stay Fresh.

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • この広告は乾いた熱気と瑞々しい緑という before/after の状態変化として設計されています。ムードの羅列ではなく、実行すべき変身が1つだけモデルに与えられます。
  • 製品は終盤に登場し、クローズアップのヒーローショットで動画を締めます。飲料CMの定番のビート順であり、モデルがパターンとして再現しやすい構成です。
  • フレーバー要素(氷、フルーツの薄切り)は抽象的な形容詞ではなく、スローモーションで渦を巻く物理的なオブジェクトとして配置されています。

置き換え可能な変数

  • 飲料の種類とフレーバーの手がかり
  • ばてた人物のいる場面設定
  • 変身後の環境
  • タグラインの文言

制約

  • 末尾のタグライン一覧はコピー案であり、画面に焼き込まれる文字ではありません。画面に表示したい場合は、いずれか1つを映像指示の中へ移してください。
  • 変身は1回が予算の上限です。2つ目の製品やシーン転換を足すと15秒の枠が破綻します。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 UGC・クリエイター広告のプロンプト

MiniMax H3 が5つの参照から生成した配信オープニング:アニメ調の配信者が、LIVEバッジとフォロワーバナーの付いた Twitch 風レイアウトのチャット欄を読む様子マルチモーダル参照
15s16:9参照素材 5件

UGC・クリエイター広告

VTuber配信オープニング

5素材のプロダクションです。4枚の画像が同一性、配信プラットフォームUI、部屋、オープニングカードを別々の仕事に分割し、音声参照が本物のリップシンク付き配信者演技を駆動します。冒頭の無音部分まで台本化されています。

詳細を表示

完全なプロンプト(検証済み英語原文)

Use @Image4as the opening card only: circular Luna avatar, black background, cream “LUNALIVE”, rose “STREAM STARTING”, warm circles, tiny mint accent. Static, chime. Use @Image1 for Luna’s identity only: same face, long center-parted black hair, blue-gray eyes, pink anime hoodie, white headphones around neck, delicate necklace, pale nails. Do not copy the drink, pose, or background from @Image1 . Use @Image2 only for Twitch-like platform chrome: dark top bar with “LUNALIVE”, red “LIVE” badge, “2.4K viewers”, right “STREAM CHAT” rail, bottom title area, rounded “FOLLOW” and pink “SUBSCRIBE” buttons. Do not copy Luna, pose, drink, or room from @Image2 Use @Image3 only for the cozy room behind Luna inside the video area: desk, monitor, white PC, plush shelves, curtain fairy lights, soft pink/purple light. No empty-room showcase shot. Use @Audio1as Luna’s actual vocal performance and behavior reference. @Audio1has a silent lead-in: 0.0–2.4s must be treated as no speech. Preserve her voice identity, cadence, tone, breaths, pauses, emphasis, warmth, and streamer mannerisms. Do not replace the voice, do not generate a different influencer voice, and do not add extra spoken lines beyond @Audio1 Lip sync, mouth shapes, jaw movement, smiles, glances, nods, and small hand movements must follow the audio waveform after 2.4s. Create a 15-second 16:9 Twitch-like stream opening. Important performance direction: Luna is reading chat, not delivering a camera monologue. Place Luna slightly left of center in the video area with the right “STREAM CHAT” rail clearly visible. Whenever she speaks, her eyes angle screen-right toward the chat rail as if she is reading the messages out loud. She returns to camera only for brief reactions. Add constant small movement: eye darts to chat, eyebrow lifts, tiny nods, head tilts, shoulders shifting, one subtle hand gesture near the desk. No stiff talking-head pose. One cursor only, no cursor trail, no duplicate panels, no duplicate buttons. Render only large UI text cleanly. Chat feels alive with typing dots and soft short blurred lines, but only these chat lines are readable: “hi Luna”, “welcome back”, “so cozy”, “gugugaga?”, “RAID INCOMING!”. Do not invent usernames. Chat pops stay quieter than Luna’s voice. [0–2 seconds] Open on @Image4 “LUNALIVE / STREAM STARTING”. Absolute no-speech zone: @Audio1is silent here, Luna is not shown, no mouth movement, no voice on the card. One soft chime only. No movement. [2–7.4 seconds] Hard cut to Luna live at 2.0s, slightly left of center in the cozy room from @Image3 with @Image2 chrome active: “LUNALIVE”, red “LIVE”, “2.4K viewers”, right “STREAM CHAT” rail, bottom title “COZY NEON” and “Just Chatting”. She settles for a beat, eyes already moving toward the chat rail. At about 2.4s when speech begins in @Audio1 , match lip sync exactly while she reads toward the chat rail, not into camera. Chat shows typing dots and soft blurred lines. [7.4–7.9 seconds] First audio pause = chat beat. Typing dots, then readable messages pop in: “hi Luna”, “welcome back”. Luna’s eyes track the new messages on the right rail; small nod and smile follow the audio pause. [7.9–12.5 seconds] Continue matching @Audio1Keep her gaze mostly on the chat rail while speaking, like she is reading and reacting. Add one more readable message: “so cozy”. During any softer phrase, she leans slightly forward as if reading; during brighter phrases, eyebrows lift and shoulders react. No frozen face. [12.5–13.9 seconds] Bigger audio pause = bigger chat beat. “gugugaga?” appears, then “RAID INCOMING!”, and a clean “NEW FOLLOWER” banner slides in with a gentle pop. Luna reads the raid message from the chat rail, then reacts brighter as the audio resumes. [13.9–14.8 seconds] Finish @Audio1 with accurate lip sync. If the audio winds down, Luna stops talking, gives a small wave toward chat, and settles into a warm listening pose. End on the stable live frame: Luna slightly left of center, eyes toward the right chat rail, red “LIVE”, “2.4K viewers”, no end card. Audio mix: 0–2s card is silent except one soft chime. @Audio1voice begins only after the cut to Luna and remains primary. Tiny chat pops under the voice. Warm low room tone. No crowd noise, no music lyrics, no rain, no traffic.

必要な入力

  • 画像1:配信者の同一性のみ。顔、髪、パーカー。ポーズと背景は複製しないと明記されています。
  • 画像2:配信プラットフォームのUI。チャット欄、LIVEバッジ、ボタン類。
  • 画像3:彼女の背後にある居心地のよい部屋。
  • 画像4:静止画の「stream starting」オープニングカード。
  • Audio 1:実際のボーカルテイク。0〜2.4秒の無音は「話さない区間」として台本化されています。

機能する理由

  • 各参照に仕事が1つずつ与えられ、さらに「複製しないもの」が2つずつ明示されています。ライブラリ中で最も整った素材別の役割分担です。
  • 演技の指示——「彼女はチャットを読んでいるのであって、独白しているのではない」——が視線とマイクロジェスチャーを方向付けます。クリップが生配信に見えるのはこの働きです。
  • 読めるチャットメッセージは5つの文字列に限定され、それ以外はぼかしたままにするため、UIテキストが崩れずに描画されます。

置き換え可能な変数

  • 配信者の同一性の参照
  • プラットフォームUIのスタイル
  • ホワイトリスト化するチャット行
  • 音声テイクとその間の取り方

制約

  • 原投稿は @Image1…@Audio1 という表記です。EvoLink では素材を配列上の位置で指定し、音声には必ず画像を1枚以上添えてください。音声は単独では送れません。
  • このルートは画像9枚、動画3本、音声3本まで、合計12ファイルが上限です。

設定

マルチモーダル参照 · 15s · 16:9 · 参照素材 5件

MiniMax H3 のUGC広告:明るい寝室でスポイトボトルのヘアセラムを塗る自分を、スマホ風のフレーミングで撮影する女性マルチモーダル参照
15s3:4参照素材 1件

UGC・クリエイター広告

UGC風ヘアセラム広告

製品参照1枚で作る TikTok 風セラム体験談の4シーン構成です。自撮りの導入、塗布のクローズアップ、鏡の前の結果、カウンターの製品ショット。不完全さこそがスタイリングです。

詳細を表示

完全なプロンプト(検証済み英語原文)

Create a 15-second authentic UGC-style hair growth serum ad using the provided product image as the exact product reference. SCENE 1 — 0–3s A young woman films herself in a bright bedroom using a smartphone front camera. Natural lighting, handheld movement, casual appearance. She looks at the camera and says: “I’ve been trying this hair growth serum lately…” SCENE 2 — 3–7s Cut to a close handheld shot of the woman holding the serum bottle. She removes the dropper, applies a few drops directly to her scalp, and gently massages it in. Keep the movement natural and slightly imperfect like real UGC content. SCENE 3 — 7–11s Mirror selfie shot. She runs her fingers through her hair, showing healthy-looking, fuller hair while casually talking to the camera: “And honestly, I love how easy it is to add to my routine.” SCENE 4 — 11–15s Close-up product shot on her bathroom counter. She picks up the bottle and smiles toward the camera. End with natural on-screen text: “Simple hair care. Every day.” Style: authentic TikTok/Reels UGC, smartphone camera, realistic skin texture, natural expressions, subtle handheld motion, imperfect framing, casual home environment, soft daylight, realistic audio, no cinematic commercial look, no excessive beauty filters. Preserve the exact product packaging, label, bottle shape, and branding from the reference image.

必要な入力

  • 画像1:自社の製品写真。パッケージ、ラベル、ボトル形状は全4シーンでこの1枚から固定されます。

機能する理由

  • シーンの順序がオーガニックなクリエイター投稿(フック → デモ → 結果 → 製品)をなぞるため、フィードの中で広告がネイティブに見えます。
  • 「本物のUGCのように少し不完全に」と、シネマティック調を禁じる否定リストの組み合わせが、視聴者がスクロールで飛ばす作り込まれたCM感を打ち消します。
  • 話す台詞は短く、引用され、会話調です。ネイティブ音声がそれを自然な体験談として発話します。

置き換え可能な変数

  • 製品写真
  • 2つの台詞
  • 寝室・バスルームの舞台
  • 締めの画面内テキスト

制約

  • 製品の同一性は参照画像だけに宿らせます。ラベルをテキストでも記述すると衝突を招きます。
  • 公開クリップの実測は2:3で、APIはこの比率を受け付けません。フィードに馴染む縦型なら3:4か9:16をリクエストしてください。

設定

マルチモーダル参照 · 15s · 3:4 · 参照素材 1件

MiniMax H3 のフードVlog:素早いジャンプカットの合間に、若い女性が笑いながらグルメチーズバーガーをスマホのカメラへ持ち上げる様子マルチモーダル参照
15s21:9参照素材 1件

UGC・クリエイター広告

バーガーナイトのUGC Vlog

バーガーの参照を固定した8ショット構成のフードVlogテンプレートです。自撮りのフック、箱を開ける、ズームパンチ、チーズプル、ひと口目のリアクション。テンポはハードなジャンプカットが担います。

詳細を表示

完全なプロンプト(検証済み英語原文)

Duration: 15 seconds | Aspect Ratio: 16:9 | Style: Authentic UGC / iPhone selfie-vlog, handheld, natural light, TikTok/Reels aesthetic. Product Reference: Use the uploaded gourmet burger image as the only product reference. Preserve the bun shape, patty thickness, cheese melt, lettuce, tomato, sauces, and proportions exactly in every shot. Character Description Name: Hana A young Japanese woman in her early 20s with natural beauty, long dark hair in a loose ponytail, oversized cream sweatshirt, minimal makeup, bright smile, friendly lifestyle-vlogger personality. Shot Breakdown SHOT 1 (0–2s) — Selfie showing the burger box. Dialogue: "Burger night!" SHOT 2 (2–4s) — Opens the box. SHOT 3 (4–6s) — Quick zoom on the burger. SHOT 4 (6–8s) — Hands lifting the burger with cheese stretching naturally. SHOT 5 (8–10s) — Bite reaction. Dialogue: "Okay... that's incredible." SHOT 6 (10–12s) — Casual close-up b-roll while reaching for fries. SHOT 7 (12–14s) — Toasting the burger toward the camera. Dialogue: "You need this." SHOT 8 (14–15s) — Freeze frame with overlay: "burger cravings = solved 🍔" Look & Feel Warm apartment lighting, genuine phone footage, slight grain, natural autofocus breathing, handheld imperfections, fast jump cuts. Negative Prompt cinematic grading, commercial production, CGI burger, fake cheese, distorted hands, warped food, perfect stabilization, studio lighting, text glitches, logo distortion.

必要な入力

  • 画像1:フードのヒーローショット。バンズ、パティ、チーズの溶け具合、比率がすべてのカットで同一に保たれます。

機能する理由

  • 約2秒×8つのマイクロショットは、実際のフード系クリエイターの編集そのものです。エネルギーがフォーマットにネイティブになります。
  • 名前付きのキャラクター(「Hana」)と短い引用台詞2つが、指定過剰にならずに安定した顔と声をモデルに与えます。
  • フードの物理表現に専用の指示があります(チーズが自然に伸びる)。マネーショットは祈るものではなく台本にするものです。

置き換え可能な変数

  • フードとその盛り付け
  • キャラクターのスタイリング
  • 2つの台詞
  • フリーズフレームの締めテキスト

制約

  • フードの記述は参照画像の中だけに留めてください。否定リスト(CGバーガーなし、作り物のチーズなし)がリアリズムを守ります。

設定

マルチモーダル参照 · 15s · 21:9 · 参照素材 1件

MiniMax H3 タイポグラフィ・文字アニメーションのプロンプト

MiniMax H3 のファッションフィルム:白いハイネックドレスの女性の周囲を水彩のリボンがカメラごと引き回し、LET SILENCE BLOOM の文字が物理的な文字として現れる様子テキストから動画
15s16:9参照素材なし

タイポグラフィ・文字アニメーション

水彩オートクチュールのワンテイク

水彩のリボンが1本の連続したカメラを空間の中へ引き込むアバンギャルドなファッションフィルムです。タイポグラフィは物理的なオブジェクトとして存在し、すべての動きが音楽に着地します。

詳細を表示

完全なプロンプト(検証済み英語原文)

## Concept AQUARELLE No.7 — an avant-garde haute couture film where watercolor becomes a living medium. 15-second cinematic sequence. The rhythm controls the visual world: 0–4s: restrained silence and negative space 4–8s: gradual density buildup 8s: drop moment, expanding into fluid long-form motion 12–15s: transition into a final fashion poster composition ## Character Identity Lock Maintain the exact same female character throughout the entire video. Identity: - Young woman - Long black hair - Calm and refined facial features - White high-neck pigment dress - Black wide belt - Pigment heels Strict consistency: - Same face - Same hairstyle - Same age - Same body proportions - Same garment structure Any new colors must only appear through watercolor gradually absorbing into the fabric. Core concept: Her fingertips can extract transparent watercolor ribbons from the air. These watercolor ribbons can: - Pull the camera through space - Transform the environment - Shape physical typography - Interact with depth and materials The world follows real cinematic physics: water, paper, fabric, light, shadows, and depth of field must feel physically believable. ## Camera Direction Strict one-take shot. No cuts. No teleportation. No hidden transitions. Camera journey: Macro shot of a floating water droplet → watercolor ribbon emerges → camera pulls back to reveal the woman → camera circles around her → enters a paper art gallery → passes through dimensional typography → rises into a final overhead fashion poster. ## Typography Only allow these words: LET SILENCE BLOOM AQUARELLE No.7 WEAR THE UNSEEN Typography is not a flat overlay. Letters must have: - Physical depth - Wet watercolor reflections - Paper fiber edges - Shadows - Occlusion - Material interaction ## Visual Style Avant-garde fashion editorial. Inspired by: - Museum catalog composition - Handmade ivory paper - Sculptural negative space - Elegant Didone serif typography - Translucent watercolor calligraphy Color palette: - Pale cyan - Crimson lake - Smoky violet - Ink black ## Motion Language Every movement follows the music: Kick: Camera movement and paper folding. Wooden snare: Paper structures physically fold and transform. Sub-bass: Changes the perception of spatial scale. ## Storyboard 30 beats, 0.5 seconds each. 01 0.0–0.5 Macro shot: A transparent water droplet floats in the air, reflecting a blurred silhouette of a black-haired woman. 02 0.5–1.0 The droplet stretches with the breath-like vocal, becoming a pale cyan watercolor thread. 03 1.0–1.5 Camera travels backward along the thread as paper fibers slowly emerge into focus. 04 1.5–2.0 The thread wraps around the lens. Focus shifts to her raised fingertip. 05 2.0–3.0 Camera continues pulling back, revealing her face and white high-neck dress. 06 3.0–4.0 She moves her wrist. The watercolor thread guides a smooth camera arc. 07 4.0–8.0 Additional watercolor colors emerge from her movement. Paper folds, typography begins forming, and the phrase "LET SILENCE BLOOM" appears as a physical object. 08 8.0–12.0 The camera passes through a transparent paper flower structure. The environment expands into an endless ivory paper gallery. Her dress absorbs watercolor naturally. The ribbons create sculptural forms around her body. 09 12.0–15.0 Camera cranes upward. The composition transforms into a luxury fashion advertisement poster. Typography appears: AQUARELLE No.7 WEAR THE UNSEEN Final frame: A museum-level fashion editorial poster. The woman remains centered, calm, and elegant. A final watercolor droplet remains suspended in the air. ## Negative Prompt No: - Cuts - Scene changes - Identity change - Face swap - Extra limbs - Deformed hands - Random costume changes - Explosive paint effects without physical cause - Incorrect typography - Chinese characters - Extra subtitles - Extra logos - Watermarks

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • カメラの旅路が1本の連続した鎖(水滴 → リボン → 人物の出現 → ギャラリー → 俯瞰のポスター)として書かれ、「カットなし」がハードルールとして明示されています。
  • タイポグラフィはホワイトリスト制で、登場できるのは3つのフレーズだけです。さらに紙の繊維の縁、濡れた反射、オクルージョンといった物質特性が与えられており、文字が崩れず描画されるのはそのためです。
  • 0.5秒刻み30ビートの絵コンテが音と空間を対応付けます。キックが紙を折り、スネアが構造を変形させ、サブベースがスケール感を変えます。

置き換え可能な変数

  • 許可される3つのフレーズ
  • 顔料のパレット
  • 衣装の構造
  • ギャラリーの環境

制約

  • フレーズのホワイトリストこそがタイポグラフィ品質の仕組みです。テキストを追加すると、このプロンプトが避けるよう設計された綴りの揺れが再発します。
  • ワンテイクのプロンプトは派手に失敗します。どこか1つのビートでもカットを示唆すると、空間の連鎖全体が壊れます。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 の教育アニメ:丸みを帯びた文字 A がふくらんで笑顔の赤いリンゴになり、パステル調の画面に APPLE の単語が表示される様子テキストから動画
15s16:9参照素材なし

タイポグラフィ・文字アニメーション

A・B・C・D 学習アニメーション

「文字→音→モノ→アクション→単語」の固定された学習ループを持つ子ども向けフォニックスアニメです。各文字が対応するモノへ物理的にモーフィングし、ナレーションは秒単位で台本化されています。

詳細を表示

完全なプロンプト(検証済み英語原文)

Create a 15-second animated educational video that teaches young children the letters A, B, C, and D. The learning pattern for every letter must be: LETTER → SOUND → OBJECT → PLAYFUL ACTION → OBJECT NAME Target audience: children ages 3 to 6. Visual style: Use adorable rounded 3D characters, soft pastel colors, gentle facial expressions, and simple recognizable objects. Combine this with a premium minimalist technology aesthetic featuring clean white space, elegant composition, soft studio lighting, subtle reflections, smooth gradients, rounded geometry, crisp typography, and extremely polished transitions. The animation should feel playful and child-friendly while remaining calm, uncluttered, and beautifully designed. Use a clean off-white background with a different soft color glow behind each letter. 0:00–0:01 | Introduction A small smiling star mascot bounces into the center of the screen. Colorful letters briefly float around it. Display the text: “Let’s learn!” The mascot taps the screen, creating a soft ripple that reveals the first letter. 0:01–0:04 | A is for Apple Show a large uppercase “A” and smaller lowercase “a” beside it. Use thick, rounded, highly readable typography. The narrator says: “A. A says ah. A is for Apple.” The uppercase A gently inflates and transforms into a shiny red apple. Its top point becomes the apple stem, and a small green leaf unfolds from the side. The apple gains a cute smiling face and performs one soft bounce. Display the word: “APPLE” Highlight the first letter A in red. Add a soft pop and a tiny crunchy sound. 0:04–0:07 | B is for Ball The apple rolls across the screen and leaves behind a curved red trail. The trail loops twice and forms a large uppercase “B,” with a lowercase “b” appearing beside it. The narrator says: “B. B says buh. B is for Ball.” The two rounded sections of the B expand and merge into a colorful striped ball. The ball bounces twice with playful squash-and-stretch animation. Display the word: “BALL” Highlight the first letter B in blue. Synchronize each bounce with a soft musical note. 0:07–0:10 | C is for Cat On its final bounce, the ball stretches into a curved shape and becomes a large uppercase “C.” A lowercase “c” slides gently into place beside it. The narrator says: “C. C says kuh. C is for Cat.” The C rotates and becomes the curled tail of a cute orange cat. The rest of the cat forms from soft rounded shapes. The cat stretches, blinks, and gives one gentle wave with its paw. Display the word: “CAT” Highlight the first letter C in orange. Add a quiet and friendly “meow.” 0:10–0:13 | D is for Duck The cat’s tail uncurls and transforms into the curved side of a large uppercase “D.” A lowercase “d” pops up beside it. The narrator says: “D. D says duh. D is for Duck.” The straight line of the D becomes the duck’s neck. The curved section becomes its round yellow body. A small orange beak and two tiny wings pop into place. The duck waddles forward, flaps its wings, and gives one cheerful quack. Display the word: “DUCK” Highlight the first letter D in yellow. Add tiny water ripples beneath its feet. 0:13–0:15 | Recap The apple, ball, cat, and duck slide into four clean rounded tiles. Place their letters above them: “A B C D” The mascot returns and points to each object as they bounce once in sequence. Narrator: “A, B, C, D. Great job!” Finish with the text: “Great job!” Use a small sparkle animation and a warm musical chime. Animation requirements: Keep each letter fully visible for a moment before it transforms. Show uppercase and lowercase versions clearly. Make every object instantly recognizable. Use smooth shape morphing so children can visually understand how the letter becomes the object. Maintain stable spelling, clean letterforms, accurate object shapes, and consistent character design. Use gentle squash-and-stretch, soft motion blur, subtle shadows, polished lighting, and precisely synchronized sound effects. Avoid fast camera movement, cluttered backgrounds, harsh colors, tiny text, warped letters, random symbols, duplicated objects, scary expressions, or overly complex transformations. The final video should feel cute, educational, memorable, calming, and exceptionally polished.

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • LETTER → SOUND → OBJECT → ACTION → NAME のループが同一構造で4回繰り返されます。反復は H3 が持つ最強の安定化装置です。
  • すべてのモーフィングが幾何学的な説明になっています(A の頂点がリンゴの軸になり、B のふくらみがボールになる)。だから変形を経ても文字の形が保たれます。
  • ナレーションはセグメントごとのタイムスタンプ付きで正確に引用され、ネイティブ音声を駆動しながら字幕テキストとの同期を保ちます。

置き換え可能な変数

  • 4つの文字とモノ
  • マスコットのデザイン
  • 文字ごとのグローの色
  • ナレーションの声

制約

  • 綴りの安定性は「各文字を変形前に完全に見せる」ルールに依存しています。これを削ると文字の歪みが再発します。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 のキネティックタイポグラフィ:Every great change の文字が細身のセリフ体で闇から浮かび上がり、金色の粒子が漂う様子テキストから動画
15s16:9参照素材なし

タイポグラフィ・文字アニメーション

キネティック・タイポグラフィの名言

純粋なキネティックタイポグラフィです。1つの名言をフレーズごとに明かし、フレーズが変わるたびにシーンの空気、動きの言語、パレットが変化して、最後はオフホワイトのエンドカードに全文が固定されます。

詳細を表示

完全なプロンプト(検証済み英語原文)

Create a 15-second cinematic text-animation video built around the quote: “Every great change begins quietly, grows through courage, and becomes impossible to ignore.” The quote should appear gradually as a visual story. Each new phrase must transform the design, atmosphere, movement, and emotional intensity of the scene. Use elegant typography, accurate spelling, cinematic lighting, smooth transitions, and perfectly readable text. 0:00–0:03 | “Every great change” Begin with a completely black screen. A tiny point of warm light slowly appears in the center, like the first spark of an idea. The words “Every great change” emerge softly from the darkness, one word at a time. Use thin, elegant serif typography with wide letter spacing. “Every” fades in gently. “Great” grows slightly larger. “Change” forms from small drifting particles that gather into solid letters. Keep the scene quiet, minimal, and mysterious. 0:03–0:06 | “begins quietly,” The camera slowly moves closer to the text. The previous words shrink and reposition toward the upper-left corner as the phrase “begins quietly,” appears in delicate lowercase letters. Animate the phrase as though it is being written by an invisible hand. Each letter should create a subtle ripple in the darkness. Introduce faint textures, soft shadows, floating dust, and gentle light rays. The comma should appear last and create a small circular pulse. 0:06–0:09 | “grows through courage,” The pulse expands and transforms the scene from darkness into a rich sunrise gradient with deep orange, red, and golden tones. The words “grows through courage” rise upward from the bottom of the frame. Animate “grows” by gradually increasing its size and weight. Animate “through” along a curved path. Animate “courage” in bold uppercase letters that push through a translucent barrier, causing it to crack into geometric fragments. The movement should feel powerful but controlled. 0:09–0:12 | “and becomes” The fragments rotate in slow motion and reorganize into a clean editorial grid. The phrase “AND BECOMES” appears across the frame in condensed sans-serif typography. Animate the letters with fast tracking changes, vertical stretching, masking, and perspective movement. The camera accelerates forward through the center of the word “BECOMES.” The sound and visual energy should steadily build. 0:12–0:14 | “impossible to ignore.” Reveal a vast bright space filled with light, moving shapes, and large-scale typography. The words “IMPOSSIBLE TO IGNORE” appear one after another. “IMPOSSIBLE” expands beyond the edges of the screen. “TO” remains small and perfectly centered. “IGNORE” slams into place with strong visual impact, briefly shaking the surrounding grid and shapes. Use bold contrast, dramatic scale, sharp shadows, and synchronized motion. 0:14–0:15 | Final quote All movement stops instantly. The complete quote appears centered on a clean off-white background: “Every great change begins quietly, grows through courage, and becomes impossible to ignore.” Use refined black typography with “change,” “courage,” and “impossible” highlighted in deep red. Hold the final composition clearly for the last second. Maintain one continuous visual journey from darkness to light, silence to impact, and simplicity to complexity. Keep every phrase connected through visual transformations rather than hard cuts. Use realistic motion blur, precise kerning, clean masks, stable letterforms, smooth camera movement, subtle film grain, cinematic sound design, rising ambient music, soft particles, controlled color transitions, and a final deep impact sound. Avoid misspelled words, warped letters, duplicated characters, unreadable text, random symbols, excessive flickering, chaotic layouts, inconsistent fonts.

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • 各フレーズが3秒のブロックを所有し、固有のアニメーション動詞(浮かび上がる、手書きされる、立ち上がる、叩き込まれる)を持ちます。エネルギーの積み上げが形容詞ではなく構造で作られています。
  • 感情のアークがデザイン言語に対応付けられています。闇から朝焼けのグラデーション、そしてエディトリアルのグリッドへ。モデルは言葉だけでなくパレットの台本を受け取ります。
  • 「すべての動きが瞬時に止まる」というハードストップと静止したエンドカードが、サムネイルに使える読みやすい最終フレームを保証します。

置き換え可能な変数

  • 名言とフレーズの区切り方
  • 強調する単語とアクセントカラー
  • フレーズごとのアニメーション動詞
  • エンドカードのスタイリング

制約

  • 各フレーズは6語程度までに抑えてください。H3 は文の長さの行より、短いディスプレイテキストのほうをはるかに安定して描画します。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 ショートドラマ・物語のプロンプト

MiniMax H3 で生成した縦型ショートドラマ予告:ろうそくの灯る城の内部にいる吸血鬼の主役と人間のヒロインマルチモーダル参照
15s9:16参照素材 2件

ショートドラマ・物語

吸血鬼ロマンス・ショートドラマ

1枚で主役2人、別の1枚で場所を固定し、関係性、テンポ、画角に集中する縦型フックです。

詳細を表示

完全なプロンプト(検証済み英語原文)

Generate a 15-second, 9:16 vertical trailer segment for an international live-action vampire romance short drama. Use Figure 1 as the appearance reference for the male and female leads, and Figure 2 as the scene reference. Keep both leads' identities consistent, with a realistic live-action look and premium short-drama production quality. Story: an innocent human heroine accidentally enters a forbidden area of an old castle and awakens a sleeping aristocratic vampire. He discovers that she carries an aura connected to an ancient war, which sparks a powerful urge to control her and a dangerous fascination with her. She fears him but does not completely submit, resisting his pressure. Overall style: an international ReelShort / DramaBox vampire-romance trailer. Dark romance, dangerous attraction, fate, intense control, brooding oppression, and a striking reversal. Keep the visuals premium, restrained, and tightly paced, like the opening 15-second hook of a hit short drama. No gore, cheap horror, Halloween aesthetic, or modern street feel. Format: 9:16 vertical composition for TikTok / ReelShort / DramaBox. Use primarily medium close-ups, close-ups, and extreme close-ups, emphasizing faces, eye contact, pressure, and relationship tension within the vertical frame.

必要な入力

  • 画像1:主役2人が一緒に写った参照。モデルが2人を1つのキャスティングとして読み取れます。
  • 画像2:ロケーションの基準。城内部の光と素材感を含めます。
  • 一文の設定と、明示された感情の反転。15秒のフックが支えられるのは1回の反転であり、1話分の物語ではありません。

機能する理由

  • ジャンルの参照(ReelShort / DramaBox 風の予告編)を指定するだけで、テンポ、カラー、構図の慣習を少ない語数で伝えられます。
  • ミディアムクローズアップ、クローズアップ、エクストリームクローズアップと画角の語彙を固定しています。縦型が窮屈ではなく高品質に見えるのはこの指定によります。
  • 除外リスト(流血、安いホラー、ハロウィン風、現代の街の雰囲気)が、このジャンルが崩れる典型的な4パターンを潰しています。

置き換え可能な変数

  • 主役の外見参照
  • ロケーション
  • 設定と反転
  • 想定プラットフォームの見え方
  • 画角の組み合わせ

制約

  • 縦型出力は参照ルートの aspect_ratio で指定します。プロンプトに「9:16」と書くだけでは指定になりません。
  • 主役2人が1枚に収まった画像のほうが、個別に切り出した2枚のポートレートより同一性が安定します。

設定

マルチモーダル参照 · 15s · 9:16 · 参照素材 2件

MiniMax H3 で生成したビジュアルノベルのインターフェース遷移:固定した開始フレームと最終フレームの間の変化先頭/末尾フレーム
15s16:9参照素材 2件

ショートドラマ・物語

乙女ビジュアルノベルのトランジション

先頭と末尾フレームを固定し、その間の過程と感情変化だけをテキストで指定します。

詳細を表示

完全なプロンプト(検証済み英語原文)

Use the first image as the opening frame and the second image as the exact final frame to generate an otome visual-novel interface transition. Overall feel: a premium Chinese otome romance-interaction interface capturing an intimate moment before and after a performance. Transition naturally from "choose to watch his performance" to "Han Xu is drawn in by the heroine's words and reacts with intrigued interest." UI text, choices, and dialogue boxes should appear with refined otome-game presentation. Keep the transition silky smooth and the emotion suggestive yet restrained.

必要な入力

  • 画像1:UI の状態を含む開始フレーム。
  • 画像2:着地させたい正確な最終フレーム。
  • 2つの状態の間で起きる感情変化を一文で。

機能する理由

  • 両端が固定されているため、モデルは構図ではなく補間を解きます。予測可能なショットを得る最も確実な方法です。
  • フレームを列挙するのではなく感情の変化(「彼のパフォーマンスを見ることを選ぶ」→「引き込まれて興味を示す」)を名指ししているため、演技がカットを支えます。
  • インターフェース要素をジャンルの流儀でアニメーションさせるよう指定することで、UI が描き直されるのを防いでいます。

置き換え可能な変数

  • 開始フレームと終了フレーム
  • 2つの間の感情の起伏
  • UI の見せ方
  • トランジションの速度

制約

  • 2枚は画像から動画のルートで image_start と image_end として送ります。参照ルートはこれらの項目を受け付けません。
  • 2枚の構図と照明が近いほど補間は滑らかになります。無関係な構図同士では、トランジションではなくカットになります。

設定

先頭/末尾フレーム · 15s · 16:9 · 参照素材 2件

MiniMax H3 の時代物シーケンス:1940年代の軍服の兵士たちが、煙の立ちこめる米国の田舎町でビンテージカーの陰に身を隠す、アーカイブフィルム風の映像テキストから動画
15s16:9参照素材なし

ショートドラマ・物語

1940年代の戦時ニュース映画リアリズム

時代考証に忠実な1947年のニュース映画シミュレーションです。リアリズムは積み重ねた3つのシステムから生まれます。時代の揃った美術、物理の契約、そして現代の物体を全面的に禁じる否定リストです。

詳細を表示

完全なプロンプト(検証済み英語原文)

Create a 15-second ultra-photorealistic live-action war sequence set in the United States in 1947, designed to look like authentic historical footage captured on a 1940s film camera. The entire scene must feel grounded, documentary-like, raw, and physically realistic. Environment: A rural American town in 1947 with wooden houses, old brick buildings, telephone poles, dirt roads, vintage American cars from the 1940s, wooden fences, farmland, and period-accurate street details. Overcast afternoon light, light fog, drifting smoke, dust in the air, damaged buildings, scattered debris, and a tense wartime atmosphere. Characters: American soldiers wearing historically accurate late-1940s military uniforms, helmets, boots, and equipment. Civilians wear authentic 1940s American clothing. Natural faces, realistic skin texture, sweat, dirt, fatigue, and believable body movements. 0–3s — Establishing Shot: Wide handheld shot of a quiet rural American street suddenly filled with smoke and confusion. Vintage 1940s vehicles are parked along the road while soldiers move quickly between wooden buildings. Civilians rush toward safer areas. 3–6s — Tension: Camera moves through the street at shoulder height, following several soldiers as distant gunfire is heard. They immediately react and take cover behind a vintage vehicle and a brick wall. Their movements are cautious and realistic. 6–10s — Combat: Fast handheld tracking shot as the soldiers move between cover while distant gunfire impacts the environment. Small pieces of wood, dust, and debris fall naturally from nearby impacts. Weapon recoil, movement, and body weight must be physically accurate. Keep the violence realistic and restrained. 10–13s — Human Moment: Camera briefly focuses on a soldier helping an injured civilian move behind cover. Their breathing, facial expressions, body language, and movement should feel natural and unscripted. 13–15s — Final Shot: Camera pulls back into a wide shot of the American town as smoke slowly moves through the street. Soldiers remain behind cover while vintage vehicles and damaged buildings fill the background. The scene ends with an authentic, tense 1940s documentary feeling. Visual Style: Ultra-photorealistic live-action, authentic 1940s American environment, vintage 35mm film texture, subtle film grain, natural imperfections, realistic exposure, handheld documentary cinematography, muted historical color palette, realistic smoke and dust, natural shadows, accurate depth of field. Physics: Strictly obey real-world gravity, momentum, inertia, friction, recoil, weight, collision physics, and human biomechanics. No exaggerated explosions, impossible movements, superhero behavior, or choreographed-looking combat. Negative Prompt: modern buildings, modern cars, smartphones, modern clothing, modern weapons, futuristic technology, CGI appearance, video-game graphics, fantasy, superhero action, excessive explosions, excessive blood, gore, impossible physics, unrealistic recoil, slow-motion physics, distorted faces, extra limbs, floating objects, plastic skin, artificial-looking environments.

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • 時代考証が二重に強制されています。肯定側(1940年代の車、軍服、電柱)と否定側(スマートフォンなし、現代建築なし)の両方から失敗の方向を塞ぎます。
  • 物理の段落(「重力、運動量、反動に厳密に従う…」)はレンダリング仕様のように読め、スーパーヒーロー的な動きを目に見えて抑え込みます。
  • 静かな人間のビート(兵士が民間人を助ける)が10〜13秒に予定されており、ノンストップのアクションではなくドキュメンタリーとしての説得力を与えます。

置き換え可能な変数

  • 時代と場所
  • 人間味のある瞬間
  • フィルムストックの質感
  • 戦闘ビートの強度

制約

  • 抑制の条項(「暴力は現実的かつ抑制的に」、流血なし)は、この出力が実用に耐える理由の一部です。アレンジする際も残してください。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 キャラクター・動作のプロンプト

MiniMax H3 で生成したクレイアニメ:溶岩の渓谷を飛び越えるクレイの狐と、その下を通り抜けるカメラ先頭/末尾フレーム
10s16:9参照素材 1件

キャラクター・動作

クレイ・フォックスの渓谷ジャンプ

1枚の開始画像を1つの明確な動作にし、カメラ経路も同じ精度で記述します。

詳細を表示

完全なプロンプト(検証済み英語原文)

Claymation style. A sprinting fox reaches the edge of a cliff and launches without hesitation, making a dramatically tense, heroic slow-motion leap across a vast lava canyon. While the fox is airborne, the camera rushes at high speed beneath its belly in a sweeping dynamic move, fully revealing the terrifying depth of the chasm and the fox's clay body at maximum extension in midair.

必要な入力

  • スタイルをすでに含んだ開始画像。ここではクレイの狐とその質感です。
  • 起こしたい動作を1つ。3つではありません。

機能する理由

  • プロンプトは絵ではなく変化を書いています。開始フレームが見た目をすでに担っているので、すべての語が動きに使われます。
  • カメラに独自の指示(狐の腹の下を高速で通り抜ける動き)が与えられており、これが単なるジャンプを「ショット」に変えています。
  • ピークの瞬間(「空中で最大限に伸びた状態」)を名指しすることで、モデルにタイミングを組み立てる目標ポーズを与えています。

置き換え可能な変数

  • キャラクターと素材のスタイル
  • 環境と危険要素
  • カメラ経路
  • スローモーションの強調点

制約

  • 画像から動画のルートは入力画像から出力比率を決定し、aspect_ratio は受け付けません。
  • 開始フレームがすでに示している内容を書き直さないでください。静的な見た目の繰り返しは、このルートのプロンプトを最も無駄にする書き方です。

設定

先頭/末尾フレーム · 10s · 16:9 · 参照素材 1件

MiniMax H3 で生成したモーション転写:参照動画から複製したストリートダンスを踊る2人の参照キャラクターマルチモーダル参照
10s16:9参照素材 3件

キャラクター・動作

ストリートダンスのモーション転写

参照動画の振付を、画像で2人のキャラクターとして与える短いプロンプトです。

詳細を表示

完全なプロンプト(検証済み英語原文)

Have the characters perform street dance following the movements in Video 1. Use Figure 1 and Figure 2 as the character references.

必要な入力

  • 画像1と画像2:2人のキャラクター。それぞれ全身が写ったきれいな参照を1枚ずつ。
  • Video 1:複製したい動き。2〜15秒で、演者が全身フレーム内に収まっているもの。

機能する理由

  • 情報は参照素材が担っているため、プロンプトは短くて済みます。動画が動きを、画像が同一性を持っています。
  • 各素材が配列上の位置で指定されているので、どの参照が何を供給するかが曖昧になりません。
  • 他には何も要求していません。照明もカメラもスタイルも指定しないのは、余計な指示が転写したい動きと競合するからです。

置き換え可能な変数

  • キャラクター
  • 元となる振付
  • 環境
  • 演者の人数

制約

  • 参照動画の合計時間は15秒以内、各クリップは2〜15秒、23.976〜60 FPS である必要があります。
  • 参照動画は MP4 または MOV で H.264 または H.265、1本あたり50 MB まで、JSON 全体は 64 MB 未満に収めます。

設定

マルチモーダル参照 · 10s · 16:9 · 参照素材 3件

MiniMax H3 が参照シートから生成したキャラクター登場シーン:黒髪で三つ編みポニーテールの戦士が、石造りの遺跡でブーツから全身のヒーローポーズまで順に明かされる様子マルチモーダル参照
15s1:1参照素材 1件

キャラクター・動作

キャラクターシートのヒーロー登場

ライブラリで最も「いいね」を集めたコミュニティテンプレートです。1枚のキャラクター参照シートから、ブーツ → 顔 → 全身と進むシネマティックな登場シーンを生成でき、どんなオリジナルキャラクターのデザインにも機能します。

詳細を表示

完全なプロンプト(検証済み英語原文)

Use @[char ref] as the sole character reference. Preserve the exact identity, face, body proportions, hairstyle, outfit, colors, materials and overall silhouette of the character throughout the entire video. Do not redesign, simplify or replace any defining visual features. Create a cinematic character introduction focused on presence, silhouette, attitude and controlled motion. 0–4s Begin with a close shot of a defining lower-body or detail element such as boots, shoes, feet, hands, clothing hem or an important accessory. The character enters frame or settles into position. The camera slowly tracks upward while hair, clothing and secondary elements move naturally in the wind or environment. 4–8s Reveal more of the body with a medium or medium-wide shot from the back, side or three-quarter angle. The character stands in a calm, composed way inside the environment. The camera makes a smooth orbit, arc or lateral move to gradually reveal the character’s face and silhouette. 8–12s Move into a tight cinematic portrait or upper-body shot. The character performs one subtle signature action that fits their personality, such as lifting the chin, turning the head, adjusting clothing, brushing hair aside, opening a hand, looking toward camera, or shifting posture. Keep the motion minimal and intentional. The expression should match the character’s vibe. 12–15s End with a strong full-body hero shot that clearly presents the entire design and silhouette. Use a low-angle, eye-level or slightly dramatic framing depending on the character’s personality. The character settles into a natural final pose and holds it confidently for a clean final reveal. VISUAL DIRECTION Premium cinematic presentation. Match the visual medium and rendering style of @[char ref]. Emphasize clean silhouette, elegant staging, subtle secondary motion, believable hair and cloth movement, strong composition, atmospheric depth and polished lighting. The scene should feel like a high-end anime, game or film character introduction. CAMERA Use a clear progression from detail reveal to partial reveal to face reveal to full-body hero reveal. Camera movement should be smooth, controlled and intentional. Avoid chaotic motion. ENVIRONMENT Place the character in a fitting environment that supports their identity and mood. The background should enhance the character without distracting from them.

必要な入力

  • 画像1:キャラクターの参照。デザインシートまたはクリーンな全身レンダリング。動画は同一性、衣装、素材、レンダリングスタイルをこの1枚から引き継ぎます。

機能する理由

  • リビールの階段(ディテール → 部分 → 顔 → 全身のヒーローショット)が固定されたカメラ文法なので、モデルは変動幅をショットプランではなくキャラクターに使えます。
  • 「参照の視覚媒体とレンダリングスタイルに合わせる」の一文により、アニメ、ゲームレンダリング、実写風のどのキャラクターでも同じプロンプトが機能します。
  • 3ブロック目の「シグネチャーアクション」の枠に個性が宿ります。仕草は1つだけ、意図的に小さく。

置き換え可能な変数

  • キャラクターの参照シート
  • シグネチャーアクション
  • 環境のムード
  • 最終ポーズのフレーミング

制約

  • 原投稿は参照を @[char ref] と表記しています。EvoLink のルートでは素材は配列上の位置で指定するため、「Image 1」と書けば image_urls[0] に対応します。
  • キャラクター1人、環境1つ。このテンプレートは意図的にロケーションを切り替えないため、シルエットが安定し続けます。

設定

マルチモーダル参照 · 15s · 1:1 · 参照素材 1件

MiniMax H3 シネマティック・VFXのプロンプト

MiniMax H3 で生成したSF予告編:巨大な円形の宇宙ゲートの前に立つ小さな人影と、暗闇から浮かび上がるタイトルマルチモーダル参照
15s16:9参照素材 2件

シネマティック・VFX

SFミステリー予告編

2枚の参照画像で雰囲気と主人公を固定し、プロンプトでプッシュイン、タイトル、音響を制御する予告編です。

詳細を表示

完全なプロンプト(検証済み英語原文)

Realistic cinematic look, high-contrast lighting, and a tight pace. Use Figure 1 as the overall atmosphere and style reference, and Figure 2 as the protagonist reference. Shot 1 — Ultra-wide establishing shot. A huge circular cosmic gateway nearly fills the frame. The person is only a tiny figure seen from behind before the gateway, positioned toward the lower right. The ground is wet and reflective, and the center of the gateway is pitch black. The camera slowly pushes forward. A large title fades in from the edge of the darkness, blurred at first and then sharp: "THE STARS WERE LISTENING". Use an extremely condensed, heavy, all-caps typeface in dark red mixed with rust red, with subtle grain and misted edges. Audio: a deep low-frequency pulse, faint metallic vibrations in the distance, and a soft hit as the text becomes sharp. → Hard cut.

必要な入力

  • 画像1:雰囲気とスタイルの基準。カットに引き継ぎたい環境、配色、カラーグレーディングを含めます。
  • 画像2:主人公の参照。人物が画面内で小さくなっても顔とシルエットが判別できる大きさで撮影します。
  • 実際に焼き込みたいタイトル文字。H3 は書いたとおりに文字を描画しようとするため、短くし綴りを正確にします。

機能する理由

  • 参照素材ごとに役割が1つだけ与えられています(画像1が雰囲気、画像2が人物)。どちらがルックを決めるかをモデルが推測せずに済みます。
  • カットは1回の連続プッシュインと1つの出来事(タイトルが鮮明になる)だけで構成されており、15秒が実際に支えられる情報量に収まっています。
  • 音は「映画的な劇伴」といった曖昧な指定ではなく、低域のパルス、遠くの金属的な振動、文字が鮮明になる瞬間の打撃音という3層で書かれています。

置き換え可能な変数

  • タイトル文字と書体処理
  • 設定ショットのゲートやランドマーク
  • タイトルの色
  • 環境音のベッド
  • 終わりのカットの扱い

制約

  • 素材は配列上の位置で指定します(「Figure 1」「Figure 2」)。image_urls の並び順と一致させてください。@image1 という記法はこの API 契約には含まれません。
  • 公開されているプロンプトは意図的に途中で止めた出発点です。ハードカットで終わるため、2つ目のカットを自分で続けられます。

設定

マルチモーダル参照 · 15s · 16:9 · 参照素材 2件

MiniMax H3 で生成したテキスト動画:夕暮れの台所を手持ちで撮影し、手描きの光る生きものが小道具の間を動く様子テキストから動画
15s16:9参照素材なし

シネマティック・VFX

光る台所の生きもの

実写の台所と手描き発光アニメを組み合わせ、カメラの不完全さと禁止要素を明記したテキスト動画テンプレートです。

詳細を表示

完全なプロンプト(検証済み英語原文)

15-second, 16:9 landscape video. Blend live-action footage of a small kitchen at dusk with hand-drawn glowing animation. The last light of sunset lingers by the window. The lived-in kitchen contains an old wooden table, a half-washed mug, a slightly fogged glass bottle, and a hanging dishcloth. Give the footage subtle one-handed smartphone shake, hesitant close-range focusing, exposure fluctuations caused by backlight, and slightly coarse noise in the shadows. It should not look carefully arranged like an advertisement; instead, it should feel like someone hurriedly captured an unbelievable event at home. Do not show huge eyes, gaping mouths, fangs, threatening or lunging movements, sudden black frames, or jump scares. Use only kitchen room tone, cloth rubbing, the soft clink of a mug, water dripping from the faucet, the camera operator's footsteps and quiet breathing, plus gentle electronic sounds and tiny calls from the hand-drawn creature.

必要な入力

  • アップロードは不要です。このルートはプロンプトのみで動作します。
  • 尺とアスペクト比はプロンプト文ではなくリクエストのパラメータで指定します。

機能する理由

  • 片手持ちの手ぶれ、迷いのあるピント合わせ、逆光による露出の揺れ、暗部の粗いノイズと、カメラの欠点が具体的に指定されています。これが「広告」ではなく「家で慌てて撮った」質感を生みます。
  • 巨大な目、牙、飛びかかる動き、ジャンプスケアを禁じる短い明示リストが、かわいい生きものがホラーへ流れるのを防ぎます。
  • 音の指示が発生源ごとに分かれています:部屋の空気音、布、マグカップ、蛇口、足音、呼吸、そして生きもの自身の声。

置き換え可能な変数

  • 部屋と時間帯
  • 生きもののデザインと動き
  • テーブル上の小道具
  • カメラのどの不完全さを見せるか
  • 音のレイヤー

制約

  • プロンプト内に尺と画面比が書かれていますが、API には duration と aspect_ratio をリクエスト項目として別途渡す必要があります。
  • 否定指示は短く明示的なリストが最も有効です。禁止事項を積み上げるほど、動作の記述に使える予算が削られます。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 が先頭フレームから生成したアニメレース:濡れた山道のヘアピンで、シアンと深紅の光の尾を引きながら2台のホバーバイクが順位を入れ替える様子先頭/末尾フレーム
15s16:9参照素材 1件

シネマティック・VFX

ホバーバイクのアニメレース

先頭フレーム駆動のスポーツアニメレースです。CRITICAL ENTITY LOCK により、名前付きのライダー2人と仕様を完全に指定したバイク2台だけを、タイム指定のヘアピン、スリップストリーム、写真判定のゴールまで追跡します。

詳細を表示

完全なプロンプト(検証済み英語原文)

Cinematic Anime Video Scene Generate a 15-second horizontal 16:9 original high-speed hover-bike racing anime video from the provided first frame. CRITICAL ENTITY LOCK: There must be exactly 2 racers and 2 bikes in the entire video: RENJI on VALKYRIE-01 (cyan/black drift bike) and ELENA on AERO-X (crimson/white draft bike). Do not add extra racers, drone support vehicles, spectators, or traffic. Maintain total visual consistency for both bikes, helmet visors, suit patterns, repulsor spark colors, and bike liveries throughout the sequence. Entity identity: VALKYRIE-01: Matte-black and cyan angular hover-bike, exposed repulsor pads, lateral drift brakes, blue plasma exhaust trails, ridden by Renji (cyan trim suit). AERO-X: Pearl-white and neon-crimson aerodynamic hover-bike, enclosed canopy, crimson energy draft aura, white-hot central booster, ridden by Elena (crimson/gold visor suit). Video style: High-budget modern sports anime, sakuga-level velocity animation, crisp line art, vibrant neon lighting contrast, high-speed camera tracking, hyper-realistic friction and energy particle effects. Set on a wet downhill mountain pass at dawn. Camera and pacing: Continuous forward velocity, zero slow-motion interruptions: 0.0s - 3.0s: High-speed rear-tracking shot diving into the first downhill hairpin curve; instant drift initiation. 3.0s - 7.5s: Tight side-parallel tracking shot as bikes navigate rock debris and trade positions through S-curves. 7.5s - 11.5s: Close camera lock on the draft-slingshot maneuver; high-energy particle displacement as booster ignition occurs. 11.5s - 15.0s: Low-angle front-facing camera lock on the final sprint to the finish line bridge, ending on a hyper-speed photo-finish freeze. Action timing: 0.0s - 1.5s: Sequence begins at speed. VALKYRIE-01 leads downhill; AERO-X locks onto its rear bumper. Anti-gravity repulsors spray road water and blue sparks into the frame. 1.5s - 4.0s: First sharp hairpin. VALKYRIE-01 deploys lateral drift airbrakes with a burst of blue thruster fire, sliding sideways at 300 km/h. AERO-X stays glued inside its slipstream aura. 4.0s - 7.0s: Mountain debris hazard. VALKYRIE-01 hops over a boulder using a repulsor burst. AERO-X ducks under it, scraping the neon magenta guardrail in a cloud of friction sparks. 7.0s - 10.0s: S-Curve exchange. Bikes lean side-by-side; their repulsor fields collide, creating a bright electrical shockwave. ELENA pulls the overdrive lever; AERO-X's rear fins extend. 10.0s - 13.0s: Slingshot maneuver. AERO-X bursts out of VALKYRIE-01's draft, igniting its central white plasma booster. Both bikes roar down the final straightaway side-by-side. 13.0s - 15.0s: Final sprint toward the finish light gate. Water sprays violently behind them. Both nose cones cross the finish line simultaneously in a flash of light. Final freeze frame. Motion quality: Fluid 2D animation, extreme speed-line integration, stable bike geometry, flawless vehicle reflection rendering, zero limb or body clipping, high-frame-rate kinetic realism. Environment: Wet mountain pass asphalt, sheer cliff walls, neon cyan and magenta guardrail lights, early dawn sky with pink/purple clouds, water spray, floating spark particles. Final output: 15 seconds, horizontal 16:9, original high-budget sports racing anime, exactly 2 racers, relentless kinetic pacing, dynamic cinematography, no subtitles, no watermarks, no logos.

必要な入力

  • 開始フレーム:自分の画風で描いたレーサー2人とバイク2台の静止画。シーケンス全体がこの画像から延長されます。

機能する理由

  • 頭数によるエンティティロック(「レーサーはちょうど2人、バイクは2台…観客なし、通行車両なし」)が、レース系プロンプトが陥りがちな群衆の増殖を断ちます。
  • 2台のマシンそれぞれにカラーリング、挙動の癖、ライダースーツが名前付きの同一性として与えられ、高速でもモデルが両者を区別し続けられます。
  • カメラとアクションが互いを参照し合う別々のタイムラインに分かれており、2つのカメラ動作を同時に要求することなく容赦ないペースを保ちます。

置き換え可能な変数

  • 開始フレームの画風
  • バイクの同一性とカラーリング
  • コースの危険要素
  • ゴールラインの演出

制約

  • これは先頭フレームのルートです。開始画像が画風を運び、画像から動画は aspect_ratio パラメータを受け付けません。フレームが比率を決めます。

設定

先頭/末尾フレーム · 15s · 16:9 · 参照素材 1件

MiniMax H3 ネイティブ音声・台詞のプロンプト

MiniMax H3 で生成した台詞の置き換え:登場人物の台詞が参照音声の新しい台詞に差し替えられた状態マルチモーダル参照
10s16:9参照素材 2件

ネイティブ音声・台詞

台詞と演技の置き換え

元の台詞と新しい台詞を逐語的に引用し、演技の変化幅を限定します。

詳細を表示

完全なプロンプト(検証済み英語原文)

Replace the girl's line in Video 1, "We can't be together. It's not that we don't love each other; we truly can't make it to the end," with the line from Audio 1: "Don't go, okay? This time, let's not let go of each other." Slightly adjust the corresponding performance.

必要な入力

  • Video 1:置き換える台詞が含まれるクリップ。
  • Audio 1:新しい台詞。WAV または MP3 で、15 MB・15秒まで。
  • 新旧どちらの台詞もプロンプト内に一字一句書き出します。

機能する理由

  • 元の台詞を引用することで、「中盤あたりの台詞」ではなく、クリップのどの区間を処理するかを正確に指定できます。
  • 新しい台詞を引用すると、リップシンクが音声からの推測ではなく明確な目標を持てます。
  • 「対応する演技を少しだけ調整する」という書き方が、演技変更の許可に上限を設け、テイクの残りを保護します。

置き換え可能な変数

  • 元クリップ
  • 置き換える台詞
  • 新しい台詞と声
  • 演技をどこまで変えてよいか

制約

  • 音声だけを参照素材にすることはできません。必ず画像または動画と一緒に渡します。
  • 参照音声と参照動画は、1リクエストあたりそれぞれ合計15秒までです。
  • 固定したい要素を明記してください。台詞のことだけを書くと、構図、衣装、背景が動いてしまいます。

設定

マルチモーダル参照 · 10s · 16:9 · 参照素材 2件

MiniMax H3 で生成したボイス参照クリップ:参照音声の声色で書かれた台詞を話す登場人物マルチモーダル参照
10s16:9参照素材 2件

ネイティブ音声・台詞

「風のままに」ボイス参照

話す台詞と、その声色を定義する1本のクリップだけで構成する最小プロンプトです。

詳細を表示

完全なプロンプト(検証済み英語原文)

Character dialogue: "Follow the wind, live free. Leave worries behind, enjoy the moment." Use Audio 1 as the voice-timbre reference.

必要な入力

  • Video 1:台詞を話すキャラクター。
  • Audio 1:目標とする声のクリーンなサンプル。2〜15秒で、できれば音楽が重なっていないもの。
  • 話させたい台詞を一字一句書き出したもの。

機能する理由

  • 台詞が要約ではなく引用で書かれているため、タイミングとリップシンクに具体的な手がかりがあります。
  • Audio 1 には「声色」という限定的な役割だけが与えられており、一般的な「BGM」として渡されていません。
  • 他に何も指定していないため、構図と演技は参照クリップが完全に制御し続けます。

置き換え可能な変数

  • 台詞
  • ボイス参照
  • キャラクターのクリップ
  • 話す速さ

制約

  • ボイス参照が転写するのは声色であり、アクセント、感情、テンポは自動的には移りません。必要ならプロンプトに書いてください。
  • 参照音声に音楽や複数話者が混ざると品質が落ちます。単独の声だけを抽出して渡してください。

設定

マルチモーダル参照 · 10s · 16:9 · 参照素材 2件

MiniMax H3 がキャラクターと音声の参照から生成したモンタージュ:角のある白髪のキャラクターが、ビート同期の高速ショットで5つの環境を横断する様子マルチモーダル参照
15s1:1参照素材 2件

ネイティブ音声・台詞

音声同期のエンバイロメント・モンタージュ

キャラクター画像が同一性を固定し、音声クリップが編集を支配する二重参照のモンタージュです。5つの環境×各6つのバーストカットで、すべてのカットがトラックのアクセントに着地します。

詳細を表示

完全なプロンプト(検証済み英語原文)

Use @[char ref] as the strict character reference and @[audio ref] as the timing, rhythm and editing reference. Keep the character’s exact identity, proportions, hairstyle, outfit, colors and overall style consistent throughout. Create a 15-second cinematic burst-cut video showcasing the character across 5 different environments that naturally fit their design, vibe and world. AUDIO SYNC Synchronize the entire edit to @[audio ref]. Cuts, camera accents, transitions and environment changes should land precisely on strong beats, half-beats and musical accents. Let audio1 control the pacing and intensity of the montage. STRUCTURE - 5 environments total - 3 seconds per environment - 6 burst-cut shots per environment - 30 shots total Each environment must be clearly different in atmosphere, lighting, scale and visual language. Show each environment through rapid cinematic angles: wide establishing shots, aerials, low angles, side views, tracking shots, close environmental details, medium shots and hero frames. Every cut must reveal a new angle, distance, composition or spatial relationship. Avoid repeated framing. Mix static shots, push-ins, pull-backs, tracking, orbit and crane-like movement. Keep character movement subtle and natural. The focus is environmental variety, cinematic framing and tight synchronization with audio1. Hard constraints: - exactly 5 environments - exactly 6 shots per environment - exactly 30 shots total - environment changes must follow audio1’s musical phrasing - cuts and motion accents synchronized to audio1 - no outfit changes - no character duplication - no morphing - no text or UI - no blurry unreadable frames - maintain strict character consistency

必要な入力

  • 画像1:全ショットが維持すべき同一性を持つキャラクター。
  • Audio 1:テンポを支配するトラック。カット、トランジション、環境の切り替えはそのビートに従います。

機能する理由

  • モダリティ間の分業が明快です。画像が「誰」を、音声が「いつ」を答え、どちらもテキストと争いません。
  • 正確な算数(5環境 × 6ショット = 30カット)がハード制約として明示され、漠然としたモンタージュを数えられる構造に変えています。
  • 「キャラクターの動きは控えめに保つ」がエネルギーをカメラの多様性へ押し込みます。モンタージュはアクションの多様性よりカメラの多様性のほうがはるかに耐えます。

置き換え可能な変数

  • キャラクター参照
  • 音声トラックとそのフレージング
  • 5つの環境
  • ショットタイプの組み合わせ

制約

  • 原投稿は @[char ref] と @[audio ref] という表記です。EvoLink では配列上の位置(画像1、Audio 1)で指定し、音声だけを参照素材にはできない点も忘れないでください。
  • このルートの参照音声クリップは2〜15秒である必要があります。
  • 公開クリップの実測は8:9で、APIはこの比率を受け付けません。同じくほぼ正方形の画面にしたい場合は1:1をリクエストしてください。

設定

マルチモーダル参照 · 15s · 1:1 · 参照素材 2件

MiniMax H3 ゲーム・UIのプロンプト

MiniMax H3 が9枚の参照から生成したUIアニメーション:ミニマルなクリーチャー図鑑で、カーソルが草原にいる角のあるふわふわのクリーチャー MEADOW CROWN を選択する様子マルチモーダル参照
15s16:9参照素材 9件

ゲーム・UI

クリーチャー図鑑UIデモ

9枚の参照を明示的な役割に対応付けます。UIレイアウトの1枚とクリーチャーカードの8枚を固定カメラの図鑑デモとしてアニメーションさせ、カーソルがエントリーをクリックしていき、最後のクリーチャーがカーソルを食べてしまいます。

詳細を表示

完全なプロンプト(検証済み英語原文)

Use Image 1 as the exact UI/layout/style reference for the creature encyclopedia screen. Use Images 2–9 as the exact creature references. Map them like this: Image 2 = card A = LUMI HARE Image 3 = card B = CLOUD WISP Image 4 = card C = EMBER FENNEC Image 5 = card D = TIDE BEHEMOTH Image 6 = card E = PETAL VULPIN Image 7 = card F = ORCHARD EYE Image 8 = card G = MEADOW CROWN Image 9 = card H = FROST GLIDER Create a 15-second 16:9 video. Keep the camera locked. Keep the interface, layout, typography, panels, icons and overall composition stable, elegant and readable. The UI should feel like a modern minimal digital creature encyclopedia, similar to a sleek pokedex. No scene cuts, no extra text, no extra buttons, no UI distortion. Sequence: 0–2.5s: Cursor clicks card A. Main creature becomes LUMI HARE. Title changes to “LUMI HARE”. Creature blinks and rotates slightly. 2.5–5s: Cursor clicks card C. Main creature becomes EMBER FENNEC. Title changes to “EMBER FENNEC”. Cursor drags to rotate it left and right. 5–7.5s: Cursor clicks card E. Main creature becomes PETAL VULPIN. Title changes to “PETAL VULPIN”. Cursor pokes it a few times. It reacts, annoyed. 7.5–10s: Cursor clicks card G. Main creature becomes MEADOW CROWN. Title changes to “MEADOW CROWN”. Cursor taps near the face/horns. It recoils slightly. 10–12s: Cursor clicks card D. Main creature becomes TIDE BEHEMOTH. Title changes to “TIDE BEHEMOTH”. Cursor keeps poking it. 12–15s: TIDE BEHEMOTH gets angry, opens its mouth very wide, lunges forward, and swallows the cursor. Then it returns to idle. Title stays “TIDE BEHEMOTH”. Rules: - When a card is selected, both the main creature and the main title must update. - Only animate cursor, selection state, title change, and the selected creature. - Only one cursor. - Keep motion subtle and clean until the final swallow. - No cuts, no camera move, no UI distortion, no extra text. Audio: soft UI click sounds, subtle hover sounds, tiny creature reaction sounds, then a sharper aggressive creature sound and one comedic swallow gulp at the end.

必要な入力

  • 画像1:タイポグラフィ、パネル、構図を支配する図鑑UIのレイアウト。
  • 画像2〜9:カードごとにクリーチャー1体。それぞれプロンプトのカード対応表で名前が与えられています。

機能する理由

  • 画像とカードの対応表(Image 2 = カードA = LUMI HARE…)は、考えうる限り最も文字どおりの役割割り当てです。どの素材が何かをモデルが推測することはありません。
  • カメラは固定され、動いてよいのは4つだけ(カーソル、選択状態、タイトル、アクティブなクリーチャー)。UIモーションの失敗面積をほぼゼロまで縮めています。
  • 締めのコメディビート(クリーチャーがカーソルを飲み込む)は最後に予定されており、オチまでデモがクリーンに保たれます。

置き換え可能な変数

  • UIスタイルの参照
  • 8体のクリーチャーと名前
  • インタラクションの台本
  • 効果音の集合

制約

  • 画像9枚はこのルートの種類別の上限です。さらに合計12ファイルの上限があるため、その上に動画3本と音声3本をすべて積むことはできません。
  • UIテキストの安定性は「カメラを動かさない、UIを歪ませない」に依存しています。カメラを解放するとパネルの歪みが再発します。

設定

マルチモーダル参照 · 15s · 16:9 · 参照素材 9件

MiniMax H3 のゲームプレイ風クリップ:ジェネリックなFPSのHUDを表示しながら、煙る軍事基地をライフルのスコープ越しに進む一人称視点テキストから動画
15s16:9参照素材なし

ゲーム・UI

FPSゲームプレイ・シミュレーション

キャプチャされたゲームプレイに見える一人称シューティングのシーケンスです。プレイヤー操作のカメラ文法、完全指定のジェネリックHUD、振り付けではなく戦術として書かれたペース配分で構成します。

詳細を表示

完全なプロンプト(検証済み英語原文)

Camera: First-person perspective at eye level with authentic handheld player movement, as if recorded directly from a modern AAA military shooter. The player carries a highly detailed assault rifle with realistic animations, visible hands, tactical gloves, dynamic reload mechanics, and weapon sway. **Opening Action:** The video immediately begins with the player already aiming down a roadway inside a modern military base. Multiple enemy soldiers are visible in the distance near sandbags, barricades, and military vehicles. The player carefully tracks one target, making small aim corrections while maintaining ADS (aim down sights). Fire several controlled bursts immediately at the visible enemies, producing realistic muzzle flashes, shell casings ejecting, smoke, recoil, hit reactions, and dust impacts around the targets. Continue firing in multiple short bursts while adjusting aim between enemies, simulating authentic FPS gameplay rather than scripted animation. **Movement:** After the opening firefight, lower slightly from ADS and begin advancing cautiously along the road beside concrete barriers, Hesco walls, and parked military vehicles. Frequently check left and right corners, briefly stop to reacquire targets, then raise the weapon and fire additional controlled bursts whenever enemies appear ahead. Continue pushing forward with deliberate player-controlled movement, using cover naturally and maintaining believable tactical pacing. **Environment:** Large modern military base with guard towers, armored vehicles, shipping containers, blast barriers, damaged buildings, smoke plumes, burning debris, scattered shell casings, dust clouds, and atmospheric battlefield haze. Cool natural daylight mixed with smoke and orange firelight creates a cinematic battlefield atmosphere. **Camera Motion:** Authentic player-controlled movement with subtle head bob, weapon sway, natural mouse-look adjustments, small left-right corrections while aiming, realistic recoil, smooth tracking of moving targets, brief pauses before shooting, and fluid forward progression. Avoid cinematic camera moves—everything should feel like genuine live gameplay captured by a skilled player. **Visual Quality:** Ultra-photorealistic, AAA game graphics with realistic PBR materials, detailed weapon models, physically accurate lighting, volumetric smoke, dynamic particle effects, crisp textures, realistic bullet impacts, muzzle flash illumination, motion blur only during rapid movement, and high-end military shooter presentation. **Gameplay UI:** Display a realistic modern FPS HUD inspired by games like PUBG, Battlefield, or Call of Duty (without copying exact copyrighted assets). Include: * Central dynamic crosshair or reticle * Ammo counter with magazine and reserve ammunition * Fire mode indicator * Compass at the top * Squad/team status panel * Mini-map in the upper corner * Health bar * Tactical equipment icons (grenades, medkit) * Hit markers when bullets connect * Directional damage indicators * Kill notification feed * Objective marker in the distance * Subtle interaction prompts and realistic HUD animations The HUD should feel polished, modern, and fully integrated into the gameplay, enhancing the illusion of authentic recorded footage from a contemporary military FPS. #MiniMaxH3

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • リアリズムの目標が「上手いプレイヤーの録画」に設定され、ヘッドボブ、エイム補正、静止してから撃つ間合いがカメラ挙動として指定されています。
  • HUDはヒットマーカーやキルフィードまで項目化されつつ、明示的にジェネリックに保たれています。実在のゲームに見える密度と、公開しても安全な抽象度の両立です。
  • 「シネマティックなカメラ移動を避ける」が鍵となる逆転です。多くのプロンプトが欲しがるものこそ、この一本を壊すものだからです。

置き換え可能な変数

  • 環境と勢力のスタイリング
  • HUD要素の集合
  • 交戦のリズム
  • 天候と光

制約

  • プロンプトはHUDを「複製せず着想を得る」に留めています。この条項は残してください。特定のゲームのHUDの複製は、別の(そしてよりリスクの高い)タスクです。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 Vlog・自撮りカメラのプロンプト

MiniMax H3 の自撮り風クリップ:森で自撮りしていた女性がスマホを反転させ、木々の間で煙を上げる墜落UFOを映す様子テキストから動画
15s1:1参照素材なし

Vlog・自撮りカメラ

自撮りカメラのUFO発見

スマホ実写感のショーケースです。オートフォーカスの揺れやローリングシャッターを備えた手持ちの自撮り映像、台本化された日本語の台詞、顔から墜落UFOへのカメラ反転で構成します。

詳細を表示

完全なプロンプト(検証済み英語原文)

Ultra photorealistic live-action captured on an iPhone 17. Authentic handheld selfie footage with premium cinematic documentary color grading, realistic HDR, deep green foliage, warm sunlight, subtle teal shadows, natural skin tones, gentle filmic contrast, rolling shutter, autofocus breathing, slight motion blur, and natural handheld shake. A lush forest in daytime with dense trees, wild plants, an uneven dirt trail, scattered leaves, soft sunlight through the canopy, and a gentle breeze. The atmosphere is quiet and slightly unsettling. A cute Japanese woman in her early twenties wearing a stylish bikini walks through the forest while recording herself in selfie mode. She suddenly notices something ahead, looks shocked, turns the camera, and points into the distance. A large crashed UFO is partially embedded in the forest floor. Its metallic hull is badly damaged with broken panels, scorch marks, exposed internal structures, thick gray smoke, and occasional sparks. She says in Japanese: 「ちょっと待って! あそこ見て! UFOじゃない!? 完全に墜落してるんだけど! 煙まで出てる! やばい、本物かもしれない! ちょっと近づいてみる!」 She alternates between filming herself and the UFO while continuing to point at it. Continuous single take. Natural walking movement, realistic hand tremors, slight framing imperfections, quick pans, and autofocus shifts between her face and the UFO. Natural sunlight creates cinematic highlights, soft shadows, realistic reflections on the UFO, subtle volumetric light, and realistic smoke. Audio: footsteps on leaves, gentle wind, birds becoming quieter near the crash site, creaking branches, faint electrical crackling from the UFO, and distant eerie unidentified animal calls echoing through the forest. Negative: no blood, no visible aliens, no monsters, no horror creature reveal, no excessive explosions, no CGI, no cartoon style, no text, no subtitles, no watermark, no logo.

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • スマホ特有のアラ(オートフォーカスの迷い、ローリングシャッター、手の震え、構図の乱れ)が列挙されています。リアルさは「realistic」という単語ではなく、名指しされた欠点から生まれます。
  • 台詞が日本語で一字一句引用されているため、H3 のネイティブ音声が意味を成さない音ではなく、リップシンクの合った実際の発話を生成します。
  • 音声はダイエジェティックに層構造で書かれています。足音、風、静かになっていく鳥、電気的なパチパチ音。BGMを排して「拾われたスマホ映像」の錯覚を守ります。

置き換え可能な変数

  • 話す台詞と言語
  • 発見される物体
  • 森または都市の舞台
  • 衣装とキャラクターのスタイリング

制約

  • 否定リスト(エイリアンなし、ホラーの種明かしなし)が構造を支えています。クリップをティーザーの域にとどめ、モデルがシーンをエスカレートさせるのを防ぎます。
  • 公開クリップは 1:1 でした。フレームは aspect パラメータで指定し、自分と対象を交互に映すフレーミングはテキスト側に残してください。

設定

テキストから動画 · 15s · 1:1 · 参照素材なし

MiniMax H3 のドキュメンタリークリップ:ストリートフォトグラファーがカフェの前で年配の男性とテリアを構図に収め、撮った写真を視聴者へ見せる様子テキストから動画
15s16:9参照素材なし

Vlog・自撮りカメラ

ストリートフォトグラファーの瞬間

カメラの中にカメラがあるコンパクトなドキュメンタリービートです。写真家が街角の何気ない光景を構図に収めてシャッターを切り、続いてカメラを裏返して、いま撮った写真を視聴者に見せます。

詳細を表示

完全なプロンプト(検証済み英語原文)

A young Western female street photographer walks through a lively downtown street and notices an elderly man sitting outside a café with his small dog. She carefully composes the candid moment through her camera, captures the photo, then turns the camera toward the viewer to proudly show the shot she just took. She smiles, says “Look at that,” then continues walking through the city. Ultra-photorealistic visuals, natural handheld documentary movement, realistic camera interaction, authentic facial expressions, accurate hand movements, realistic dog behavior, natural daylight, cinematic depth of field, continuous character consistency, immersive city ambience, premium documentary realism.

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • 写真を見せるビートは、動画の中に整合した1枚の静止画を描画することを H3 に強制します。1つの自然な仕草として成立する、静かな能力デモです。
  • インタラクションの連鎖が完全に指定されています(構図 → 撮影 → 裏返す → 「Look at that」 → 歩き去る)。ワンテイクで完結したアークを持てるのはこのためです。
  • 被写体は同一性ではなく役割で記述されています。年配の男性、小型犬、カフェ。街の光景をジェネリックで安全なものに保ちます。

置き換え可能な変数

  • スナップの被写体
  • 都市と光
  • 話す一言
  • カメラの小道具の種類

制約

  • カメラ内の写真は、それを撮ったシーンと一致していなければなりません。被写体を変えるなら、両方の記述を一緒に変えてください。

設定

テキストから動画 · 15s · 16:9 · 参照素材なし

MiniMax H3 の旅行Vlog:女性がカーテンを開けて海辺のテラスに出たあと、自撮りモードに切り替えて海を背に good morning と言う様子テキストから動画
15s9:16参照素材なし

Vlog・自撮りカメラ

海辺の朝Vlogアーク

意図的なモード切り替えを持つ秒刻みの朝Vlogです。最初の5秒はシネマティックな三人称で、そのあと初めてキャラクターが自分を撮り始めます。そこが「Vlog」の始まる瞬間です。

詳細を表示

完全なプロンプト(検証済み英語原文)

Create a 15-second ultra-realistic cinematic lifestyle vlog video, vertical 9:16, featuring the same young woman throughout the entire video. Preserve her facial identity, facial proportions, hairstyle, skin tone and overall appearance consistently in every shot. She wears the same outfit throughout: fitted white V-neck T-shirt with a small subtle logo, blue denim jeans, natural makeup, long softly wavy brown hair. 0:00–0:01 — Wake-up: Close-up inside a beautiful bright bedroom. The woman is lying comfortably on the bed, slowly wakes up, stretches naturally and opens her eyes. She is NOT filming a vlog yet and does not hold a phone or camera. Soft morning sunlight enters through the curtains. 0:01–0:02 — Gets up: Medium shot. She sits up on the bed, smiles softly, fixes her hair and gets ready to start her morning. Natural, effortless movement. 0:02–0:03 — Walks to window: She walks toward the large glass balcony door/window. Camera follows her naturally from behind/side. 0:03–0:04 — Seaside reveal: She opens the curtains/door and looks outside. Reveal a breathtaking blue ocean, coastal hills, flowers, balcony and beautiful morning sunlight. She smiles happily while taking in the view. 0:04–0:05 — Steps outside: She walks out onto the seaside terrace. Gentle ocean breeze moves her hair naturally. Wide cinematic shot showing the beautiful surroundings. 0:05–0:06 — VLOG START: Only now she starts filming herself in handheld selfie-vlog style. She looks into the camera with a bright natural smile and says: “Good morning!” 0:06–0:07 — Show the view: She turns the camera away from herself and slowly pans across the stunning ocean, coastal mountains, flowers and terrace. Smooth handheld vlog movement. 0:07–0:08 — Back to selfie: Selfie shot. She looks into the camera and happily says: “This place is just perfect!” 0:08–0:09 — Location reveal: Wide cinematic shot of the cozy seaside terrace with wooden table, chairs, plants and flowers overlooking the ocean. 0:09–0:10 — Walk to table: Medium tracking shot as she walks toward the table, enjoying the view. Her hair and T-shirt move gently in the sea breeze. 0:10–0:11 — Sit and relax: She sits at the seaside table, smiling peacefully and enjoying the ocean view. A refreshing orange-colored juice is placed on the table. 0:11–0:12 — Juice close-up: Cinematic close-up of her hand picking up the glass of fresh orange juice. Beautiful ocean bokeh in the background, natural sunlight reflecting through the glass. 0:12–0:13 — Vlog toast: Selfie shot. She raises the juice toward the camera with a cheerful smile and says: “Cheers to good days!” 0:13–0:14 — Happy close-up: Beautiful close-up of her smiling naturally at the camera, ocean and warm sunlight softly blurred behind her. 0:14–0:15 — Ending: Camera moves from her toward the sparkling ocean and peaceful coastal landscape. Warm sunlight, gentle waves and a relaxing cinematic ending. Overall Style Ultra-realistic, cinematic travel vlog, natural handheld camera movement, realistic human motion, smooth transitions, soft morning sunlight, realistic ocean waves, gentle wind in hair and clothes, beautiful coastal atmosphere, premium lifestyle aesthetic, natural expressions, authentic vlog feeling, shallow depth of field, cinematic composition, realistic skin texture, high detail, 4K quality.

必要な入力

  • アップロードは不要です。テキストから動画はプロンプトのみで動作します。

機能する理由

  • 起床ビートの「まだ撮影していない」という注記が、Vlog系で最もよくあるアーティファクト——Vlog開始前にスマホが現れてしまう問題——を防ぎます。
  • 1秒×15のビートが自撮りカメラと風景のカットアウェイを交互に並べます。実際の旅行Vlogの文法そのものです。
  • 短い引用台詞3つ(「Good morning!」)が、持続すべき独白なしに、ネイティブ音声へ自然なチェックポイントを与えます。

置き換え可能な変数

  • ロケーションのリビール
  • 3つの台詞
  • 衣装と同一性の固定
  • ドリンクの小道具

制約

  • プロンプトは9:16を求めており、リクエストでもこの比率を指定します。公開クリップはおよそ3:2に再エンコードされており、向きは文章ではなくリクエストで指定すべきことのもう1つの理由です。

設定

テキストから動画 · 15s · 9:16 · 参照素材なし

MiniMax H3 編集・変換のプロンプト

MiniMax H3 で生成した動画編集:元クリップの新聞、椅子、サングラス、炎上する車が置換または削除された状態マルチモーダル参照
10s16:9参照素材 1件

編集・変換

複数要素のシーン編集

元動画に6つの変更を「対象 → 結果」で列挙し、シーン自体は説明し直しません。

詳細を表示

完全なプロンプト(検証済み英語原文)

Replace the newspaper in the reference video with a green-covered book; change the chair the character is sitting on to a red sofa; remove the sunglasses worn by the character to retain a clear face; remove the car burning effect to keep the vehicle in a normal state; change the photo the character takes out of his arms to a small black book; and add a tree on the left side of the screen

必要な入力

  • Video 1:編集対象のクリップ。2〜15秒、MP4 または MOV、H.264 または H.265、50 MB まで。
  • 編集リスト。各項目に「何を」「どう変えるか」を明記します。

機能する理由

  • すべての指示が「対象+結果」の組になっており、解釈の余地がありません(「新聞」が「緑の表紙の本」になる)。
  • 削除指示は望ましい最終状態も併記しています(「サングラスを外して顔をはっきりさせる」)。これにより、物があった場所に穴が残るのを防ぎます。
  • シーンの描写が一切ないため、モデルは元クリップを事実として扱い、差分だけを適用します。

置き換え可能な変数

  • 編集項目の数
  • 置き換える物
  • 取り除く効果
  • 追加する要素

制約

  • EvoLink が提供する H3 のルートは3つで、編集系の作業は元クリップを Video 1 とする reference-to-video で行います。
  • 参照動画の長さは課金対象です。アップロード前に必要な範囲だけに切り詰めてください。
  • 変更対象の隣に残したいものがある場合は、変わらない部分を明示してください。書かれていない領域はモデルが自由に扱えます。

設定

マルチモーダル参照 · 10s · 16:9 · 参照素材 1件

MiniMax H3 で生成したステージ映像:煙の中でスーツの色を交換する2人のマジシャンと、赤から青へ変わる幕先頭/末尾フレーム
7s16:9参照素材 1件

編集・変換

マジシャンの衣装入れ替え

2着の衣装を入れ替え、1つの細部を保持し、背景色も指定タイミングで変える指示追従テストです。

詳細を表示

完全なプロンプト(検証済み英語原文)

Two magicians stand onstage facing the audience and perform a "swap" trick. They wave their wands at the same time, and a cloud of smoke rises. When it clears, their suit colors have switched: the person on the left wears a white suit, and the person on the right now wears a black suit, while both magicians' glove colors remain unchanged. They bow to thank the audience. The red curtain behind them closes, transitioning from deep red to deep blue.

必要な入力

  • 2人の演者、衣装、ステージが写った開始画像。
  • 入れ替わるものについて、変化の前後が明確であること。

機能する理由

  • 入れ替えを煙という出来事で隠しているため、モデルは画面上でモーフィングさせずに変化させる正当なタイミングを得られます。
  • 変えてはいけないもの(手袋の色)が、変えるもののすぐ隣に書かれています。編集範囲の広がりを防ぐのはこの書き方です。
  • ショットは明確な最終状態で終わります:お辞儀、閉じる幕、深い青に着地する色。

置き換え可能な変数

  • 演者と衣装
  • 固定したままにする細部
  • 入れ替えを隠す出来事
  • 最後の色変化

制約

  • 画像から動画のルートは入力画像からアスペクト比を決定し、aspect_ratio 項目を受け付けません。
  • 書かれていない属性はモデルが自由に扱えます。変化の中で残したい細部があるなら、必ず書き出してください。

設定

先頭/末尾フレーム · 7s · 16:9 · 参照素材 1件

MiniMax H3 の API をお探しですか?

秒単位の料金、3つの生成ルート、連携ドキュメントを掲載したモデルページへ。

API ページを開く

MiniMax H3プロンプトのフレームワーク

H3は順序付き参照を理解し、音声も生成します。汎用公式だけでなく、実際のワークフローに合わせて書きます。

H3プロンプトの基本構造

目的 → 順序付き参照 → 主体とアイデンティティ → 時系列の動作 → カメラ経路 → 音声または台詞 → 保持する要素 → 最終状態の順で書きます。音と終わりの指定を忘れにくくなります。

テキストから動画

クリップ内で完了できる時系列を書きます。主体、少数の動作、1本の連続したカメラ経路、台詞・環境音・音楽の層を分けて指定します。

先頭 / 末尾フレーム

外観は画像で決まっています。間の変化、カメラ動作、保持する内容、最後の状態だけを記述します。

マルチモーダル参照

各アセットに1つだけ役割を与え、配列順で呼びます。Image 1はキャラクター、Image 2は場所、Video 1はモーション、Audio 1は声色に使います。1リクエストの合計は12ファイルまでです。

既存クリップの編集

「変更」は対象 → 結果、「保持」は変えない内容として2つのリストに分けます。必要な範囲だけをVideo 1として送ります。

音声と台詞

話してほしい台詞をそのまま引用します。声の参照は声色のみに使い、アクセント、感情、ペースは別に指定します。

MiniMax H3プロンプト FAQ

MiniMax H3とHailuo 3は同じモデルですか?

はい。MiniMax H3はHailuo 3、Hailuo 3.0、Hailuo 03とも呼ばれます。EvoLinkではMiniMax H3の名称を使い、テキスト動画、画像動画、参照動画の3ルートを提供します。

コードを書かずに使えますか?

はい。「このプロンプトを使う」を押すと、検証済みの英語プロンプトと対応ワークフローがPlaygroundに送られます。参照ファイルはそこで追加します。

MiniMax H3の動画は何秒ですか?

3ルートとも4〜15秒の整数を指定でき、現在の出力は2Kです。カードの時間は公開されたサンプルに対応します。

1つのプロンプトで何個の参照を使えますか?

参照ルートは最大9枚の画像、3本の動画、3本の音声を受け付けますが、合計12ファイルが上限です(9+3+3のフルセットは拒否されます)。画像または動画が少なくとも1つ必要で、音声だけは使えません。動画と音声は2〜15秒です。

なぜプロンプトに「Image 1」や「Video 1」と書くのですか?

APIはimage_urls、video_urls、audio_urls内の位置で参照を識別します。検証済みプロンプト内の英語表記は、特定のアセットに役割を割り当てます。

プロンプトと動画の出典はどこですか?

各カードに出典を表示します。MiniMax公式例は中国語原文の検証済み英語版を使い、コミュニティ例はXの元投稿へリンクします。

これらのプロンプトをAPIで使えますか?

はい。同じ英語プロンプトをPlaygroundとAPIのpromptフィールドで使えます。APIガイドでリクエスト、非同期タスク、コールバック、エラーを解説しています。

プロンプトを選んで生成

すべてのプロンプトがEvoLinkの稼働中MiniMax H3ルートに対応します。Playgroundへ送るか、統合APIで同じリクエストを構築できます。