GPT Image 2.5 Flare & Sunburst ya están disponibles en EvoLinkProbar GPT Image 2.5

Prompts de MiniMax H3 y ejemplos de vídeo

Explora 40 prompts verificados con resultados reales de MiniMax H3, requisitos de entrada y variables reutilizables. Copia una plantilla o envíala al Playground de EvoLink.

MiniMax H3 también se conoce como Hailuo 3, Hailuo 3.0 o Hailuo 03.

Modo de generación

Caso de uso

40 de 40 prompts

Prompts de MiniMax H3: Marca y anuncios de producto

Anuncio vertical de gafas generado con MiniMax H3: dos modelos en un estudio blanco sin costuras con gafas envolventes futuristasREFERENCIA MULTIMODAL
15s9:163 recursos de referencia

Marca y anuncios de producto

Campaña de gafas futuristas

Anuncio de moda con tres imágenes que controlan por separado el visual principal, los rostros y el producto.

Ver detalles

Prompt completo (original inglés verificado)

Generate a vertical screen 9:16 high-end fashion glasses commercial, taking overall reference to the storyboard rhythm, editing speed, white studio texture and cool fashion atmosphere of the given video. The picture is a minimalist white booth, a seamless white background, a strong sense of high-end advertising, and a clean, simple, handsome, avant-garde, international first-line fashion blockbuster texture. Key visual character reference picture 1, two full-body female models, one black female model and one European and American model, maintain their high-end clothing, body posture, white studio light and shadow, fashion show temperament and overall cool attitude. Both of them wear futuristic high-end glasses. The design of the glasses refers to Figure 3, emphasizing the covered curved surface, sharp geometric cat-eye/goggle hybrid outline, mirror reflection, streamlined temples, and the texture of high-end fashion accessories. Please refer to Figure 2 for the appearance details of the two characters.

Qué debes aportar

  • Imagen 1: el visual clave — modelos de cuerpo entero, vestuario, luz de estudio y actitud.
  • Imagen 2: el detalle de los rostros, para que los dos personajes se mantengan iguales entre cortes.
  • Imagen 3: el producto, fotografiado con nitidez suficiente para que su silueta y sus materiales sobrevivan a un montaje rápido.
  • Un fondo de estudio sin costuras que estés dispuesto a mantener durante todo el clip.

Por qué funciona

  • Tres referencias con tres funciones explícitas es justo el patrón para el que está hecha la ruta de referencia; la ambigüedad es lo que hunde los prompts multiimagen.
  • El producto se describe por su geometría — curvatura envolvente, contorno híbrido cat-eye/gafa de protección, superficie espejada, patillas aerodinámicas — y no solo por su nombre.
  • Un estudio blanco de un solo material elimina varianza del fondo, para que el presupuesto se vaya al producto y a las modelos.

Variables sustituibles

  • categoría de producto
  • casting de modelos
  • color del estudio
  • velocidad de montaje
  • vestuario

Restricciones

  • El prompt publicado pide además el ritmo de un vídeo dado. El material publicado para este caso son tres imágenes: trata esa línea como indicación de estilo o adjunta tu propio clip como Video 1.
  • Hasta 9 imágenes, 3 vídeos y 3 clips de audio por petición de referencia, con un máximo de 12 archivos en total; el audio nunca puede ser el único tipo de referencia.

Ajustes

Referencia multimodal · 15s · 9:16 · 3 recursos de referencia

Vídeo de producto generado con MiniMax H3: unos auriculares de diadema premium giran sobre un pedestal reflectante en un estudio negroTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Marca y anuncios de producto

Presentación de auriculares de lujo

Vídeo de producto de 15 segundos dividido en cuatro bloques, cada uno con movimiento de cámara y función propios.

Ver detalles

Prompt completo (original inglés verificado)

Create a 15-second luxury cinematic product showcase for premium wireless over-ear headphones. 0–4s: Begin with an extreme macro tracking shot moving across the soft memory-foam ear cushion, fine fabric texture, brushed-metal hinge and precision-machined controls. A narrow light band travels across the surface, revealing realistic materials against a deep black studio background. 4–8s: Pull back into a three-quarter hero view. The headphones rotate slowly above a glossy reflective pedestal. The ear cups pivot naturally while the adjustable headband extends slightly, demonstrating flexible construction and comfort. Maintain exact symmetry, stable geometry and consistent proportions. 8–12s: Transition into an elegant exploded-view reveal. The ear cushion, acoustic driver, internal sound chamber, control ring and outer shell separate smoothly in perfect alignment. Subtle luminous sound waves pulse outward from the driver while the camera performs a restrained side orbit. 12–15s: Every component reconnects seamlessly. The headphones settle into a centered front-facing hero composition as soft rim lighting defines the silhouette. Complete a gentle dolly-in toward the ear cups. Premium technology-commercial finish, controlled reflections, realistic shadows, shallow depth of field, crisp surface detail, stable product shape, no hands, no distortion, no onscreen text.

Qué debes aportar

  • Nada que subir: el texto a vídeo funciona solo con el prompt.
  • Un producto que sepas describir por sus materiales y su mecánica, no solo por su nombre.

Por qué funciona

  • El tiempo se reparte en 0–4 s, 4–8 s, 8–12 s y 12–15 s, así el modelo recibe un plan de planos en lugar de una lista de deseos.
  • Cada bloque mueve la cámara de forma distinta — travelling macro, retroceso, órbita lateral, avance — y eso es lo que hace que el clip se lea montado y no a la deriva.
  • La línea final es una lista de prohibiciones (sin manos, sin deformaciones, sin texto en pantalla) que cubre los tres fallos habituales de los renders de producto.

Variables sustituibles

  • producto
  • materiales y acabado
  • duración de cada bloque
  • fondo e iluminación
  • si aparece la vista despiezada

Restricciones

  • La duración es un entero de 4 a 15 segundos y se fija en la petición; los bloques del prompt deben sumar esa cifra.
  • Los bloques cronometrados son dirección, no una línea de tiempo estricta. No más de cuatro para un clip de 15 segundos.
  • Su autor publicó este clip en 720p; las rutas H3 de EvoLink entregan 2K.

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Anuncio de bebida con MiniMax H3: un caminante agotado por el calor bebe un sorbo de zumo helado y la calle florece en vegetación exuberante alrededor de la botellaTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Marca y anuncios de producto

Anuncio de bebida contra el calor veraniego

Spot clásico de bebida de problema y alivio: un sujeto agotado por el calor, un sorbo, y todo el entorno se transforma — cierre con un plano héroe del producto cubierto de gotas.

Ver detalles

Prompt completo (original inglés verificado)

A young person walks under the blazing summer sun, looking exhausted and sweating heavily. The road shimmers with heat waves, and everything appears dry and dull. Suddenly, they grab a chilled bottle of premium fruit juice from a cooler and take a refreshing sip. Instantly, the environment transforms—lush green trees bloom, vibrant flowers appear, a cool breeze flows, water splashes through the air, and glowing particles surround the scene. Ice cubes and fresh fruit slices (orange, mango, or according to the flavour) swirl around the bottle in cinematic slow motion. End with a stunning close-up of the juice bottle covered in cold water droplets against a bright, refreshing background. Ultra-realistic, premium commercial, 4K, cinematic lighting, high-detail, smooth camera movements, vibrant colours, luxury beverage advertisement. Tagline ideas: Beat the Heat. Taste the Freshness. Every Sip Brings Life. Refresh Your Day, Naturally. Stay Cool. Stay Fresh.

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • El anuncio se construye como un cambio de estado antes/después — calor seco frente a frescura exuberante — lo que da al modelo una única transformación clara que ejecutar en lugar de una lista de estados de ánimo.
  • El producto entra tarde y cierra el clip como primer plano héroe, el orden de beats estándar del anuncio de bebidas que el modelo puede reconocer por patrón.
  • Los elementos de sabor (hielo, rodajas de fruta) se escenifican como objetos físicos girando en cámara lenta, no como adjetivos abstractos.

Variables sustituibles

  • tipo de bebida y señales de sabor
  • el escenario del sujeto agotado
  • entorno de la transformación
  • texto del eslogan

Restricciones

  • La lista de eslóganes del final es inspiración de copy, no texto incrustado — pasa uno de ellos a la dirección visual si quieres que se renderice en pantalla.
  • El presupuesto es una única transformación; añadir un segundo producto o cambio de escena saturará la ventana de 15 segundos.

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Prompts de MiniMax H3: UGC y anuncios de creadores

Apertura de directo con MiniMax H3 desde cinco referencias: una streamer de estilo anime lee su carril de chat en un layout tipo Twitch con insignia LIVE y banner de seguidoresREFERENCIA MULTIMODAL
15s16:95 recursos de referencia

UGC y anuncios de creadores

Apertura de directo de una VTuber

Una producción de cinco recursos: cuatro imágenes reparten identidad, interfaz de la plataforma, habitación y tarjeta de apertura en trabajos separados, mientras una referencia de audio impulsa una actuación real de streamer con sincronía labial — incluida su entrada en silencio.

Ver detalles

Prompt completo (original inglés verificado)

Use @Image4as the opening card only: circular Luna avatar, black background, cream “LUNALIVE”, rose “STREAM STARTING”, warm circles, tiny mint accent. Static, chime. Use @Image1 for Luna’s identity only: same face, long center-parted black hair, blue-gray eyes, pink anime hoodie, white headphones around neck, delicate necklace, pale nails. Do not copy the drink, pose, or background from @Image1 . Use @Image2 only for Twitch-like platform chrome: dark top bar with “LUNALIVE”, red “LIVE” badge, “2.4K viewers”, right “STREAM CHAT” rail, bottom title area, rounded “FOLLOW” and pink “SUBSCRIBE” buttons. Do not copy Luna, pose, drink, or room from @Image2 Use @Image3 only for the cozy room behind Luna inside the video area: desk, monitor, white PC, plush shelves, curtain fairy lights, soft pink/purple light. No empty-room showcase shot. Use @Audio1as Luna’s actual vocal performance and behavior reference. @Audio1has a silent lead-in: 0.0–2.4s must be treated as no speech. Preserve her voice identity, cadence, tone, breaths, pauses, emphasis, warmth, and streamer mannerisms. Do not replace the voice, do not generate a different influencer voice, and do not add extra spoken lines beyond @Audio1 Lip sync, mouth shapes, jaw movement, smiles, glances, nods, and small hand movements must follow the audio waveform after 2.4s. Create a 15-second 16:9 Twitch-like stream opening. Important performance direction: Luna is reading chat, not delivering a camera monologue. Place Luna slightly left of center in the video area with the right “STREAM CHAT” rail clearly visible. Whenever she speaks, her eyes angle screen-right toward the chat rail as if she is reading the messages out loud. She returns to camera only for brief reactions. Add constant small movement: eye darts to chat, eyebrow lifts, tiny nods, head tilts, shoulders shifting, one subtle hand gesture near the desk. No stiff talking-head pose. One cursor only, no cursor trail, no duplicate panels, no duplicate buttons. Render only large UI text cleanly. Chat feels alive with typing dots and soft short blurred lines, but only these chat lines are readable: “hi Luna”, “welcome back”, “so cozy”, “gugugaga?”, “RAID INCOMING!”. Do not invent usernames. Chat pops stay quieter than Luna’s voice. [0–2 seconds] Open on @Image4 “LUNALIVE / STREAM STARTING”. Absolute no-speech zone: @Audio1is silent here, Luna is not shown, no mouth movement, no voice on the card. One soft chime only. No movement. [2–7.4 seconds] Hard cut to Luna live at 2.0s, slightly left of center in the cozy room from @Image3 with @Image2 chrome active: “LUNALIVE”, red “LIVE”, “2.4K viewers”, right “STREAM CHAT” rail, bottom title “COZY NEON” and “Just Chatting”. She settles for a beat, eyes already moving toward the chat rail. At about 2.4s when speech begins in @Audio1 , match lip sync exactly while she reads toward the chat rail, not into camera. Chat shows typing dots and soft blurred lines. [7.4–7.9 seconds] First audio pause = chat beat. Typing dots, then readable messages pop in: “hi Luna”, “welcome back”. Luna’s eyes track the new messages on the right rail; small nod and smile follow the audio pause. [7.9–12.5 seconds] Continue matching @Audio1Keep her gaze mostly on the chat rail while speaking, like she is reading and reacting. Add one more readable message: “so cozy”. During any softer phrase, she leans slightly forward as if reading; during brighter phrases, eyebrows lift and shoulders react. No frozen face. [12.5–13.9 seconds] Bigger audio pause = bigger chat beat. “gugugaga?” appears, then “RAID INCOMING!”, and a clean “NEW FOLLOWER” banner slides in with a gentle pop. Luna reads the raid message from the chat rail, then reacts brighter as the audio resumes. [13.9–14.8 seconds] Finish @Audio1 with accurate lip sync. If the audio winds down, Luna stops talking, gives a small wave toward chat, and settles into a warm listening pose. End on the stable live frame: Luna slightly left of center, eyes toward the right chat rail, red “LIVE”, “2.4K viewers”, no end card. Audio mix: 0–2s card is silent except one soft chime. @Audio1voice begins only after the cut to Luna and remains primary. Tiny chat pops under the voice. Warm low room tone. No crowd noise, no music lyrics, no rain, no traffic.

Qué debes aportar

  • Imagen 1: solo la identidad de la streamer — rostro, pelo, sudadera; la pose y el fondo explícitamente no se copian.
  • Imagen 2: la interfaz de la plataforma de streaming — carril de chat, insignia LIVE, botones.
  • Imagen 3: la habitación acogedora tras ella.
  • Imagen 4: la tarjeta estática de apertura «stream starting».
  • Audio 1: su toma vocal real; ojo, el silencio de 0–2,4 s está guionizado como zona sin habla.

Por qué funciona

  • Cada referencia recibe un trabajo y dos exclusiones explícitas de «no copiar», la asignación de roles por recurso más limpia de la biblioteca.
  • La nota de actuación — «está leyendo el chat, no soltando un monólogo» — redirige la mirada y los microgestos, que es lo que hace que el clip se sienta en directo.
  • Los mensajes de chat legibles se limitan a cinco cadenas exactas y todo lo demás queda desenfocado, así el texto de la interfaz se renderiza limpio.

Variables sustituibles

  • lámina de identidad de la streamer
  • estilo de la interfaz de la plataforma
  • las líneas de chat en lista blanca
  • la toma de audio y sus pausas

Restricciones

  • El post original escribe @Image1…@Audio1; en EvoLink direcciona los recursos por posición del array y mantén al menos una imagen junto al audio — el audio nunca puede viajar solo.
  • La ruta acepta hasta 9 imágenes, 3 vídeos y 3 clips de audio, con un tope de 12 archivos en total.

Ajustes

Referencia multimodal · 15s · 16:9 · 5 recursos de referencia

Anuncio UGC con MiniMax H3: una mujer se graba aplicándose sérum capilar de un frasco con cuentagotas en un dormitorio luminoso, con encuadre de smartphoneREFERENCIA MULTIMODAL
15s3:41 recurso de referencia

UGC y anuncios de creadores

Anuncio UGC de sérum capilar

Un testimonio de sérum estilo TikTok en cuatro escenas con una referencia de producto: intro selfie, primer plano de aplicación, resultado en el espejo, plano del producto en la encimera — la imperfección es el estilismo.

Ver detalles

Prompt completo (original inglés verificado)

Create a 15-second authentic UGC-style hair growth serum ad using the provided product image as the exact product reference. SCENE 1 — 0–3s A young woman films herself in a bright bedroom using a smartphone front camera. Natural lighting, handheld movement, casual appearance. She looks at the camera and says: “I’ve been trying this hair growth serum lately…” SCENE 2 — 3–7s Cut to a close handheld shot of the woman holding the serum bottle. She removes the dropper, applies a few drops directly to her scalp, and gently massages it in. Keep the movement natural and slightly imperfect like real UGC content. SCENE 3 — 7–11s Mirror selfie shot. She runs her fingers through her hair, showing healthy-looking, fuller hair while casually talking to the camera: “And honestly, I love how easy it is to add to my routine.” SCENE 4 — 11–15s Close-up product shot on her bathroom counter. She picks up the bottle and smiles toward the camera. End with natural on-screen text: “Simple hair care. Every day.” Style: authentic TikTok/Reels UGC, smartphone camera, realistic skin texture, natural expressions, subtle handheld motion, imperfect framing, casual home environment, soft daylight, realistic audio, no cinematic commercial look, no excessive beauty filters. Preserve the exact product packaging, label, bottle shape, and branding from the reference image.

Qué debes aportar

  • Imagen 1: tu foto de producto — el packaging, la etiqueta y la forma del frasco quedan fijados desde esta lámina en las cuatro escenas.

Por qué funciona

  • El orden de escenas replica el contenido orgánico de creadores (gancho → demo → resultado → producto), así el anuncio se lee nativo dentro del feed.
  • «Slightly imperfect like real UGC» más la lista negativa anticinematográfica es lo que vence al look de anuncio pulido que los compradores pasan de largo.
  • Las frases habladas son cortas, citadas y conversacionales — el audio nativo las entrega como habla testimonial natural.

Variables sustituibles

  • lámina del producto
  • las dos frases habladas
  • escenarios de dormitorio y baño
  • texto en pantalla de cierre

Restricciones

  • La identidad del producto vive por completo en la imagen de referencia; describir la etiqueta también en el texto invita al conflicto.
  • El clip publicado mide 2:3, una relación que la API no acepta — pide 3:4 o 9:16 para un vertical nativo del feed.

Ajustes

Referencia multimodal · 15s · 3:4 · 1 recurso de referencia

Vlog gastronómico con MiniMax H3: una joven levanta una hamburguesa con queso gourmet hacia la cámara de su móvil entre risas y jump cuts rápidosREFERENCIA MULTIMODAL
15s21:91 recurso de referencia

UGC y anuncios de creadores

Vlog UGC de noche de hamburguesas

Una plantilla de vlog gastronómico en ocho planos con la hamburguesa fijada por referencia: gancho selfie, apertura de la caja, zoom de impacto, queso que se estira, reacción al bocado — los jump cuts duros hacen el trabajo de ritmo.

Ver detalles

Prompt completo (original inglés verificado)

Duration: 15 seconds | Aspect Ratio: 16:9 | Style: Authentic UGC / iPhone selfie-vlog, handheld, natural light, TikTok/Reels aesthetic. Product Reference: Use the uploaded gourmet burger image as the only product reference. Preserve the bun shape, patty thickness, cheese melt, lettuce, tomato, sauces, and proportions exactly in every shot. Character Description Name: Hana A young Japanese woman in her early 20s with natural beauty, long dark hair in a loose ponytail, oversized cream sweatshirt, minimal makeup, bright smile, friendly lifestyle-vlogger personality. Shot Breakdown SHOT 1 (0–2s) — Selfie showing the burger box. Dialogue: "Burger night!" SHOT 2 (2–4s) — Opens the box. SHOT 3 (4–6s) — Quick zoom on the burger. SHOT 4 (6–8s) — Hands lifting the burger with cheese stretching naturally. SHOT 5 (8–10s) — Bite reaction. Dialogue: "Okay... that's incredible." SHOT 6 (10–12s) — Casual close-up b-roll while reaching for fries. SHOT 7 (12–14s) — Toasting the burger toward the camera. Dialogue: "You need this." SHOT 8 (14–15s) — Freeze frame with overlay: "burger cravings = solved 🍔" Look & Feel Warm apartment lighting, genuine phone footage, slight grain, natural autofocus breathing, handheld imperfections, fast jump cuts. Negative Prompt cinematic grading, commercial production, CGI burger, fake cheese, distorted hands, warped food, perfect stabilization, studio lighting, text glitches, logo distortion.

Qué debes aportar

  • Imagen 1: el plano héroe de la comida — el pan, la carne, el queso fundido y las proporciones se mantienen idénticos en cada corte.

Por qué funciona

  • Ocho microplanos de ~2 segundos cada uno replican cómo cortan de verdad los creadores gastronómicos, así la energía es nativa del formato.
  • El personaje con nombre («Hana») y dos frases cortas citadas dan al modelo un rostro y una voz estables sin sobreespecificar.
  • La física de la comida recibe su propia instrucción (el queso estirándose de forma natural) — el plano estrella está guionizado, no confiado a la suerte.

Variables sustituibles

  • el plato y su presentación
  • estilismo del personaje
  • las dos frases de diálogo
  • texto final sobre fotograma congelado

Restricciones

  • Mantén la descripción de la comida solo en la imagen de referencia; la lista negativa (sin hamburguesa CGI, sin queso falso) protege el realismo.

Ajustes

Referencia multimodal · 15s · 21:9 · 1 recurso de referencia

Prompts de MiniMax H3: Tipografía y texto animado

Película de moda con MiniMax H3: cintas de acuarela arrastran la cámara alrededor de una mujer con vestido blanco de cuello alto mientras la frase LET SILENCE BLOOM se forma en letras físicasTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Tipografía y texto animado

Plano secuencia de alta costura en acuarela

Una película de moda vanguardista donde cintas de acuarela arrastran una única cámara continua por el espacio, la tipografía existe como objeto físico y cada movimiento cae sobre la música.

Ver detalles

Prompt completo (original inglés verificado)

## Concept AQUARELLE No.7 — an avant-garde haute couture film where watercolor becomes a living medium. 15-second cinematic sequence. The rhythm controls the visual world: 0–4s: restrained silence and negative space 4–8s: gradual density buildup 8s: drop moment, expanding into fluid long-form motion 12–15s: transition into a final fashion poster composition ## Character Identity Lock Maintain the exact same female character throughout the entire video. Identity: - Young woman - Long black hair - Calm and refined facial features - White high-neck pigment dress - Black wide belt - Pigment heels Strict consistency: - Same face - Same hairstyle - Same age - Same body proportions - Same garment structure Any new colors must only appear through watercolor gradually absorbing into the fabric. Core concept: Her fingertips can extract transparent watercolor ribbons from the air. These watercolor ribbons can: - Pull the camera through space - Transform the environment - Shape physical typography - Interact with depth and materials The world follows real cinematic physics: water, paper, fabric, light, shadows, and depth of field must feel physically believable. ## Camera Direction Strict one-take shot. No cuts. No teleportation. No hidden transitions. Camera journey: Macro shot of a floating water droplet → watercolor ribbon emerges → camera pulls back to reveal the woman → camera circles around her → enters a paper art gallery → passes through dimensional typography → rises into a final overhead fashion poster. ## Typography Only allow these words: LET SILENCE BLOOM AQUARELLE No.7 WEAR THE UNSEEN Typography is not a flat overlay. Letters must have: - Physical depth - Wet watercolor reflections - Paper fiber edges - Shadows - Occlusion - Material interaction ## Visual Style Avant-garde fashion editorial. Inspired by: - Museum catalog composition - Handmade ivory paper - Sculptural negative space - Elegant Didone serif typography - Translucent watercolor calligraphy Color palette: - Pale cyan - Crimson lake - Smoky violet - Ink black ## Motion Language Every movement follows the music: Kick: Camera movement and paper folding. Wooden snare: Paper structures physically fold and transform. Sub-bass: Changes the perception of spatial scale. ## Storyboard 30 beats, 0.5 seconds each. 01 0.0–0.5 Macro shot: A transparent water droplet floats in the air, reflecting a blurred silhouette of a black-haired woman. 02 0.5–1.0 The droplet stretches with the breath-like vocal, becoming a pale cyan watercolor thread. 03 1.0–1.5 Camera travels backward along the thread as paper fibers slowly emerge into focus. 04 1.5–2.0 The thread wraps around the lens. Focus shifts to her raised fingertip. 05 2.0–3.0 Camera continues pulling back, revealing her face and white high-neck dress. 06 3.0–4.0 She moves her wrist. The watercolor thread guides a smooth camera arc. 07 4.0–8.0 Additional watercolor colors emerge from her movement. Paper folds, typography begins forming, and the phrase "LET SILENCE BLOOM" appears as a physical object. 08 8.0–12.0 The camera passes through a transparent paper flower structure. The environment expands into an endless ivory paper gallery. Her dress absorbs watercolor naturally. The ribbons create sculptural forms around her body. 09 12.0–15.0 Camera cranes upward. The composition transforms into a luxury fashion advertisement poster. Typography appears: AQUARELLE No.7 WEAR THE UNSEEN Final frame: A museum-level fashion editorial poster. The woman remains centered, calm, and elegant. A final watercolor droplet remains suspended in the air. ## Negative Prompt No: - Cuts - Scene changes - Identity change - Face swap - Extra limbs - Deformed hands - Random costume changes - Explosive paint effects without physical cause - Incorrect typography - Chinese characters - Extra subtitles - Extra logos - Watermarks

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • El viaje de cámara se escribe como una cadena continua (gota → cinta → revelación → galería → póster cenital) con «no cuts» declarado como regla dura.
  • La tipografía está en lista blanca — solo pueden aparecer tres frases — y recibe propiedades materiales (bordes de fibra de papel, reflejos húmedos, oclusión), y por eso las palabras se renderizan limpias.
  • Un storyboard de 30 beats a 0,5 s por beat mapea el sonido al espacio: el kick pliega papel, la caja transforma estructuras, el subgrave cambia la escala.

Variables sustituibles

  • las tres frases permitidas
  • paleta de pigmentos
  • estructura del vestido
  • entorno de la galería

Restricciones

  • La lista blanca de frases es el mecanismo de calidad tipográfica — añadir más texto reintroduce la deriva ortográfica que el prompt está construido para evitar.
  • Los prompts de plano secuencia fallan de forma estrepitosa: si algún beat implica un corte, toda la cadena espacial se rompe.

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Animación educativa con MiniMax H3: una letra A redondeada se infla hasta convertirse en una manzana roja sonriente junto a la palabra APPLE en una suave escena pastelTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Tipografía y texto animado

Animación educativa A-B-C-D

Animación infantil de fonética con un bucle didáctico fijo — letra, sonido, objeto, acción, palabra — donde cada letra se transforma físicamente en su objeto y la narración está guionizada al segundo.

Ver detalles

Prompt completo (original inglés verificado)

Create a 15-second animated educational video that teaches young children the letters A, B, C, and D. The learning pattern for every letter must be: LETTER → SOUND → OBJECT → PLAYFUL ACTION → OBJECT NAME Target audience: children ages 3 to 6. Visual style: Use adorable rounded 3D characters, soft pastel colors, gentle facial expressions, and simple recognizable objects. Combine this with a premium minimalist technology aesthetic featuring clean white space, elegant composition, soft studio lighting, subtle reflections, smooth gradients, rounded geometry, crisp typography, and extremely polished transitions. The animation should feel playful and child-friendly while remaining calm, uncluttered, and beautifully designed. Use a clean off-white background with a different soft color glow behind each letter. 0:00–0:01 | Introduction A small smiling star mascot bounces into the center of the screen. Colorful letters briefly float around it. Display the text: “Let’s learn!” The mascot taps the screen, creating a soft ripple that reveals the first letter. 0:01–0:04 | A is for Apple Show a large uppercase “A” and smaller lowercase “a” beside it. Use thick, rounded, highly readable typography. The narrator says: “A. A says ah. A is for Apple.” The uppercase A gently inflates and transforms into a shiny red apple. Its top point becomes the apple stem, and a small green leaf unfolds from the side. The apple gains a cute smiling face and performs one soft bounce. Display the word: “APPLE” Highlight the first letter A in red. Add a soft pop and a tiny crunchy sound. 0:04–0:07 | B is for Ball The apple rolls across the screen and leaves behind a curved red trail. The trail loops twice and forms a large uppercase “B,” with a lowercase “b” appearing beside it. The narrator says: “B. B says buh. B is for Ball.” The two rounded sections of the B expand and merge into a colorful striped ball. The ball bounces twice with playful squash-and-stretch animation. Display the word: “BALL” Highlight the first letter B in blue. Synchronize each bounce with a soft musical note. 0:07–0:10 | C is for Cat On its final bounce, the ball stretches into a curved shape and becomes a large uppercase “C.” A lowercase “c” slides gently into place beside it. The narrator says: “C. C says kuh. C is for Cat.” The C rotates and becomes the curled tail of a cute orange cat. The rest of the cat forms from soft rounded shapes. The cat stretches, blinks, and gives one gentle wave with its paw. Display the word: “CAT” Highlight the first letter C in orange. Add a quiet and friendly “meow.” 0:10–0:13 | D is for Duck The cat’s tail uncurls and transforms into the curved side of a large uppercase “D.” A lowercase “d” pops up beside it. The narrator says: “D. D says duh. D is for Duck.” The straight line of the D becomes the duck’s neck. The curved section becomes its round yellow body. A small orange beak and two tiny wings pop into place. The duck waddles forward, flaps its wings, and gives one cheerful quack. Display the word: “DUCK” Highlight the first letter D in yellow. Add tiny water ripples beneath its feet. 0:13–0:15 | Recap The apple, ball, cat, and duck slide into four clean rounded tiles. Place their letters above them: “A B C D” The mascot returns and points to each object as they bounce once in sequence. Narrator: “A, B, C, D. Great job!” Finish with the text: “Great job!” Use a small sparkle animation and a warm musical chime. Animation requirements: Keep each letter fully visible for a moment before it transforms. Show uppercase and lowercase versions clearly. Make every object instantly recognizable. Use smooth shape morphing so children can visually understand how the letter becomes the object. Maintain stable spelling, clean letterforms, accurate object shapes, and consistent character design. Use gentle squash-and-stretch, soft motion blur, subtle shadows, polished lighting, and precisely synchronized sound effects. Avoid fast camera movement, cluttered backgrounds, harsh colors, tiny text, warped letters, random symbols, duplicated objects, scary expressions, or overly complex transformations. The final video should feel cute, educational, memorable, calming, and exceptionally polished.

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • El bucle LETTER → SOUND → OBJECT → ACTION → NAME se repite cuatro veces con estructura idéntica, y la repetición es el estabilizador más potente que tiene H3.
  • Cada metamorfosis es una explicación geométrica (la punta de la A se convierte en el rabito de la manzana, los lóbulos de la B en la pelota), así las formas de las letras sobreviven a la transformación.
  • Las líneas de narración se citan exactas con marcas de tiempo por segmento, lo que guía el audio nativo y mantiene el texto de los rótulos en sincronía.

Variables sustituibles

  • las cuatro letras y sus objetos
  • diseño de la mascota
  • halo de color por letra
  • voz de la narración

Restricciones

  • La estabilidad ortográfica depende de la regla de «mantener cada letra totalmente visible antes de transformarse» — quitarla reintroduce glifos deformados.

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Tipografía cinética con MiniMax H3: las palabras Every great change emergen de la oscuridad en finas letras serif entre partículas doradas a la derivaTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Tipografía y texto animado

Tipografía cinética de una cita

Tipografía cinética pura: una cita revelada frase a frase, donde cada frase cambia la atmósfera, el lenguaje de movimiento y la paleta de la escena hasta que la línea completa se fija en una tarjeta final en blanco roto.

Ver detalles

Prompt completo (original inglés verificado)

Create a 15-second cinematic text-animation video built around the quote: “Every great change begins quietly, grows through courage, and becomes impossible to ignore.” The quote should appear gradually as a visual story. Each new phrase must transform the design, atmosphere, movement, and emotional intensity of the scene. Use elegant typography, accurate spelling, cinematic lighting, smooth transitions, and perfectly readable text. 0:00–0:03 | “Every great change” Begin with a completely black screen. A tiny point of warm light slowly appears in the center, like the first spark of an idea. The words “Every great change” emerge softly from the darkness, one word at a time. Use thin, elegant serif typography with wide letter spacing. “Every” fades in gently. “Great” grows slightly larger. “Change” forms from small drifting particles that gather into solid letters. Keep the scene quiet, minimal, and mysterious. 0:03–0:06 | “begins quietly,” The camera slowly moves closer to the text. The previous words shrink and reposition toward the upper-left corner as the phrase “begins quietly,” appears in delicate lowercase letters. Animate the phrase as though it is being written by an invisible hand. Each letter should create a subtle ripple in the darkness. Introduce faint textures, soft shadows, floating dust, and gentle light rays. The comma should appear last and create a small circular pulse. 0:06–0:09 | “grows through courage,” The pulse expands and transforms the scene from darkness into a rich sunrise gradient with deep orange, red, and golden tones. The words “grows through courage” rise upward from the bottom of the frame. Animate “grows” by gradually increasing its size and weight. Animate “through” along a curved path. Animate “courage” in bold uppercase letters that push through a translucent barrier, causing it to crack into geometric fragments. The movement should feel powerful but controlled. 0:09–0:12 | “and becomes” The fragments rotate in slow motion and reorganize into a clean editorial grid. The phrase “AND BECOMES” appears across the frame in condensed sans-serif typography. Animate the letters with fast tracking changes, vertical stretching, masking, and perspective movement. The camera accelerates forward through the center of the word “BECOMES.” The sound and visual energy should steadily build. 0:12–0:14 | “impossible to ignore.” Reveal a vast bright space filled with light, moving shapes, and large-scale typography. The words “IMPOSSIBLE TO IGNORE” appear one after another. “IMPOSSIBLE” expands beyond the edges of the screen. “TO” remains small and perfectly centered. “IGNORE” slams into place with strong visual impact, briefly shaking the surrounding grid and shapes. Use bold contrast, dramatic scale, sharp shadows, and synchronized motion. 0:14–0:15 | Final quote All movement stops instantly. The complete quote appears centered on a clean off-white background: “Every great change begins quietly, grows through courage, and becomes impossible to ignore.” Use refined black typography with “change,” “courage,” and “impossible” highlighted in deep red. Hold the final composition clearly for the last second. Maintain one continuous visual journey from darkness to light, silence to impact, and simplicity to complexity. Keep every phrase connected through visual transformations rather than hard cuts. Use realistic motion blur, precise kerning, clean masks, stable letterforms, smooth camera movement, subtle film grain, cinematic sound design, rising ambient music, soft particles, controlled color transitions, and a final deep impact sound. Avoid misspelled words, warped letters, duplicated characters, unreadable text, random symbols, excessive flickering, chaotic layouts, inconsistent fonts.

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • Cada frase posee un bloque de 3 segundos con su propio verbo de animación (emerger, escribirse a mano, ascender, estamparse), así la subida de energía es estructural, no adjetival.
  • El arco emocional se mapea a lenguaje de diseño — de la oscuridad al degradado de amanecer y a la retícula editorial — dando al modelo un guion de paleta, no solo palabras.
  • Un frenazo total («all movement stops instantly») más una tarjeta final sostenida garantiza un último fotograma legible para miniaturas.

Variables sustituibles

  • la cita y sus cortes de frase
  • palabras destacadas y color de acento
  • verbos de animación por frase
  • estilo de la tarjeta final

Restricciones

  • Mantén las frases por debajo de ~6 palabras cada una; H3 renderiza texto corto en pantalla con mucha más fiabilidad que líneas de longitud de oración.

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Prompts de MiniMax H3: Minidrama y narrativa

Tráiler vertical de minidrama generado con MiniMax H3: un vampiro y una heroína humana en el interior de un castillo iluminado con velasREFERENCIA MULTIMODAL
15s9:162 recursos de referencia

Minidrama y narrativa

Minidrama romántico de vampiros

Gancho vertical: una imagen fija a los protagonistas, otra el lugar, y el prompt se centra en relación, ritmo y tamaños de plano.

Ver detalles

Prompt completo (original inglés verificado)

Generate a 15-second, 9:16 vertical trailer segment for an international live-action vampire romance short drama. Use Figure 1 as the appearance reference for the male and female leads, and Figure 2 as the scene reference. Keep both leads' identities consistent, with a realistic live-action look and premium short-drama production quality. Story: an innocent human heroine accidentally enters a forbidden area of an old castle and awakens a sleeping aristocratic vampire. He discovers that she carries an aura connected to an ancient war, which sparks a powerful urge to control her and a dangerous fascination with her. She fears him but does not completely submit, resisting his pressure. Overall style: an international ReelShort / DramaBox vampire-romance trailer. Dark romance, dangerous attraction, fate, intense control, brooding oppression, and a striking reversal. Keep the visuals premium, restrained, and tightly paced, like the opening 15-second hook of a hit short drama. No gore, cheap horror, Halloween aesthetic, or modern street feel. Format: 9:16 vertical composition for TikTok / ReelShort / DramaBox. Use primarily medium close-ups, close-ups, and extreme close-ups, emphasizing faces, eye contact, pressure, and relationship tension within the vertical frame.

Qué debes aportar

  • Imagen 1: los dos protagonistas juntos, para que el modelo los lea como un único reparto coherente.
  • Imagen 2: la localización — el interior del castillo, su luz y sus materiales.
  • Una premisa en una frase y un giro emocional explícito; un gancho de 15 segundos aguanta un giro, no un episodio entero.

Por qué funciona

  • Nombrar la referencia de género (tráiler tipo ReelShort / DramaBox) transmite ritmo, etalonaje y convenciones de encuadre en pocas palabras.
  • El vocabulario de planos está fijado — plano medio, primer plano, primerísimo primer plano — y eso es lo que hace que un formato vertical se lea premium en vez de agobiado.
  • La lista de exclusiones (sin gore, sin terror barato, sin estética de Halloween, sin aire de calle contemporánea) elimina las cuatro degradaciones habituales del género.

Variables sustituibles

  • referencias de aspecto de los protagonistas
  • localización
  • premisa y giro
  • estética de la plataforma objetivo
  • mezcla de tamaños de plano

Restricciones

  • El formato vertical se pide con aspect_ratio en la ruta de referencia, no escribiendo «9:16» en el prompt.
  • La identidad se mantiene mucho mejor cuando ambos protagonistas llegan en una sola imagen que en dos retratos recortados por separado.

Ajustes

Referencia multimodal · 15s · 9:16 · 2 recursos de referencia

Transición de interfaz de novela visual generada con MiniMax H3 entre un fotograma de apertura y uno de cierre fijadosPRIMER / ÚLTIMO FOTOGRAMA
15s16:92 recursos de referencia

Minidrama y narrativa

Transición de novela visual otome

El primer y el último fotograma fijan ambos extremos; el texto solo describe el recorrido y el cambio emocional.

Ver detalles

Prompt completo (original inglés verificado)

Use the first image as the opening frame and the second image as the exact final frame to generate an otome visual-novel interface transition. Overall feel: a premium Chinese otome romance-interaction interface capturing an intimate moment before and after a performance. Transition naturally from "choose to watch his performance" to "Han Xu is drawn in by the heroine's words and reacts with intrigued interest." UI text, choices, and dialogue boxes should appear with refined otome-game presentation. Keep the transition silky smooth and the emotion suggestive yet restrained.

Qué debes aportar

  • Imagen 1: el fotograma de apertura, con el estado de su interfaz.
  • Imagen 2: exactamente el fotograma de cierre en el que quieres aterrizar.
  • Una frase que describa el cambio emocional entre ambos estados.

Por qué funciona

  • Ambos extremos están fijados, así que el modelo resuelve una interpolación en lugar de una composición: la vía más fiable hacia un plano predecible.
  • El prompt nombra la transición emocional («elegir ver su actuación» → «atraída e intrigada») en vez de enumerar fotogramas, de modo que la interpretación sostiene el corte.
  • Pedir que los elementos de interfaz se animen en el lenguaje propio del género evita que la UI se redibuje.

Variables sustituibles

  • fotogramas de apertura y cierre
  • arco emocional entre ambos
  • estilo de presentación de la interfaz
  • velocidad de la transición

Restricciones

  • Envía ambos fotogramas como image_start e image_end en la ruta imagen a vídeo; la ruta de referencia no acepta estos campos.
  • Cuanto más se parezcan los dos fotogramas en encuadre e iluminación, más suave será la interpolación. Dos composiciones sin relación producen un corte, no una transición.

Ajustes

Primer / último fotograma · 15s · 16:9 · 2 recursos de referencia

Secuencia de época con MiniMax H3: soldados con uniformes de los años 40 se cubren tras coches de época en una calle rural americana llena de humo, rodada como película de archivoTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Minidrama y narrativa

Realismo de noticiario bélico de los años 40

Una simulación de noticiario de 1947 fiel a la época cuyo realismo procede de tres sistemas apilados: ambientación coherente con la era, un contrato de física y una lista negativa que veta todo objeto moderno.

Ver detalles

Prompt completo (original inglés verificado)

Create a 15-second ultra-photorealistic live-action war sequence set in the United States in 1947, designed to look like authentic historical footage captured on a 1940s film camera. The entire scene must feel grounded, documentary-like, raw, and physically realistic. Environment: A rural American town in 1947 with wooden houses, old brick buildings, telephone poles, dirt roads, vintage American cars from the 1940s, wooden fences, farmland, and period-accurate street details. Overcast afternoon light, light fog, drifting smoke, dust in the air, damaged buildings, scattered debris, and a tense wartime atmosphere. Characters: American soldiers wearing historically accurate late-1940s military uniforms, helmets, boots, and equipment. Civilians wear authentic 1940s American clothing. Natural faces, realistic skin texture, sweat, dirt, fatigue, and believable body movements. 0–3s — Establishing Shot: Wide handheld shot of a quiet rural American street suddenly filled with smoke and confusion. Vintage 1940s vehicles are parked along the road while soldiers move quickly between wooden buildings. Civilians rush toward safer areas. 3–6s — Tension: Camera moves through the street at shoulder height, following several soldiers as distant gunfire is heard. They immediately react and take cover behind a vintage vehicle and a brick wall. Their movements are cautious and realistic. 6–10s — Combat: Fast handheld tracking shot as the soldiers move between cover while distant gunfire impacts the environment. Small pieces of wood, dust, and debris fall naturally from nearby impacts. Weapon recoil, movement, and body weight must be physically accurate. Keep the violence realistic and restrained. 10–13s — Human Moment: Camera briefly focuses on a soldier helping an injured civilian move behind cover. Their breathing, facial expressions, body language, and movement should feel natural and unscripted. 13–15s — Final Shot: Camera pulls back into a wide shot of the American town as smoke slowly moves through the street. Soldiers remain behind cover while vintage vehicles and damaged buildings fill the background. The scene ends with an authentic, tense 1940s documentary feeling. Visual Style: Ultra-photorealistic live-action, authentic 1940s American environment, vintage 35mm film texture, subtle film grain, natural imperfections, realistic exposure, handheld documentary cinematography, muted historical color palette, realistic smoke and dust, natural shadows, accurate depth of field. Physics: Strictly obey real-world gravity, momentum, inertia, friction, recoil, weight, collision physics, and human biomechanics. No exaggerated explosions, impossible movements, superhero behavior, or choreographed-looking combat. Negative Prompt: modern buildings, modern cars, smartphones, modern clothing, modern weapons, futuristic technology, CGI appearance, video-game graphics, fantasy, superhero action, excessive explosions, excessive blood, gore, impossible physics, unrealistic recoil, slow-motion physics, distorted faces, extra limbs, floating objects, plastic skin, artificial-looking environments.

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • La precisión de época se impone dos veces — en positivo (coches de los 40, uniformes, postes telefónicos) y en negativo (sin smartphones, sin edificios modernos) — cerrando ambas direcciones de fallo.
  • El párrafo de física («strictly obey gravity, momentum, recoil…») se lee como una especificación de render y suprime de forma visible el movimiento estilo superhéroe.
  • Un beat humano silencioso (un soldado ayudando a un civil) se programa en 10–13 s, dando a la secuencia credibilidad documental en lugar de acción sin pausa.

Variables sustituibles

  • época y localización
  • el momento humano
  • look del material fílmico
  • intensidad de los beats de combate

Restricciones

  • Las cláusulas de contención («violence realistic and restrained», sin sangre explícita) son parte de por qué este resultado es utilizable — consérvalas al adaptar.

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Prompts de MiniMax H3: Personaje y movimiento

Animación con plastilina generada con MiniMax H3: un zorro de arcilla salta un cañón de lava mientras la cámara pasa por debajoPRIMER / ÚLTIMO FOTOGRAMA
10s16:91 recurso de referencia

Personaje y movimiento

Zorro de plastilina salta un cañón

Una imagen inicial se anima en una sola acción, con la trayectoria de cámara descrita con la misma precisión que el salto.

Ver detalles

Prompt completo (original inglés verificado)

Claymation style. A sprinting fox reaches the edge of a cliff and launches without hesitation, making a dramatically tense, heroic slow-motion leap across a vast lava canyon. While the fox is airborne, the camera rushes at high speed beneath its belly in a sweeping dynamic move, fully revealing the terrifying depth of the chasm and the fox's clay body at maximum extension in midair.

Qué debes aportar

  • Una imagen inicial que ya lleve el estilo — aquí, el zorro de plastilina y su material.
  • Una acción que quieras que ocurra. No tres.

Por qué funciona

  • El prompt describe el cambio, no la imagen. El fotograma inicial ya contiene la apariencia, así que cada palabra compra movimiento.
  • La cámara recibe su propia instrucción — un barrido a gran velocidad por debajo del vientre del zorro — y eso convierte un salto en un plano.
  • Nombrar el instante culminante («máxima extensión en el aire») le da al modelo una pose objetivo sobre la que construir el tempo.

Variables sustituibles

  • personaje y estilo de material
  • entorno y peligro
  • trayectoria de cámara
  • énfasis de la cámara lenta

Restricciones

  • La ruta imagen a vídeo deduce la relación de aspecto de la imagen de entrada y rechaza aspect_ratio.
  • No vuelvas a describir lo que la imagen inicial ya muestra; repetir la apariencia estática es la forma más común de desperdiciar un prompt de imagen a vídeo.

Ajustes

Primer / último fotograma · 10s · 16:9 · 1 recurso de referencia

Transferencia de movimiento generada con MiniMax H3: dos personajes referenciados ejecutan una coreografía de street dance copiada de un vídeo de referenciaREFERENCIA MULTIMODAL
10s16:93 recursos de referencia

Personaje y movimiento

Transferencia de movimiento de street dance

Una instrucción breve transfiere la coreografía de un vídeo a dos personajes definidos con imágenes.

Ver detalles

Prompt completo (original inglés verificado)

Have the characters perform street dance following the movements in Video 1. Use Figure 1 and Figure 2 as the character references.

Qué debes aportar

  • Imagen 1 e Imagen 2: los dos personajes, con una referencia de cuerpo entero limpia para cada uno.
  • Video 1: el movimiento a copiar, de 2 a 15 segundos, con el intérprete completamente en cuadro.

Por qué funciona

  • El prompt es corto porque la información la llevan las referencias: el vídeo posee el movimiento y las imágenes poseen las identidades.
  • Cada recurso se nombra por su posición en el array, así que no hay ambigüedad sobre qué referencia aporta qué.
  • No se pide nada más. Ni luz, ni cámara, ni estilo: cualquier instrucción extra competiría con el movimiento que se intenta copiar.

Variables sustituibles

  • personajes
  • coreografía de origen
  • entorno
  • número de intérpretes

Restricciones

  • La duración total de los vídeos de referencia debe mantenerse por debajo de 15 segundos, y cada clip entre 2 y 15 segundos a 23,976–60 FPS.
  • Los vídeos de referencia deben ser MP4 o MOV con H.264 o H.265, hasta 50 MB cada uno, y el cuerpo JSON completo por debajo de 64 MB.

Ajustes

Referencia multimodal · 10s · 16:9 · 3 recursos de referencia

Introducción de personaje con MiniMax H3 generada desde una ficha de referencia: una guerrera de pelo negro con coleta trenzada revelada de las botas a la pose heroica de cuerpo entero entre ruinas de piedraREFERENCIA MULTIMODAL
15s1:11 recurso de referencia

Personaje y movimiento

Presentación heroica desde ficha de personaje

La plantilla comunitaria con más «me gusta» de la biblioteca: una ficha de referencia de personaje dirige una introducción cinematográfica de botas a rostro a cuerpo entero que funciona con cualquier diseño de personaje original.

Ver detalles

Prompt completo (original inglés verificado)

Use @[char ref] as the sole character reference. Preserve the exact identity, face, body proportions, hairstyle, outfit, colors, materials and overall silhouette of the character throughout the entire video. Do not redesign, simplify or replace any defining visual features. Create a cinematic character introduction focused on presence, silhouette, attitude and controlled motion. 0–4s Begin with a close shot of a defining lower-body or detail element such as boots, shoes, feet, hands, clothing hem or an important accessory. The character enters frame or settles into position. The camera slowly tracks upward while hair, clothing and secondary elements move naturally in the wind or environment. 4–8s Reveal more of the body with a medium or medium-wide shot from the back, side or three-quarter angle. The character stands in a calm, composed way inside the environment. The camera makes a smooth orbit, arc or lateral move to gradually reveal the character’s face and silhouette. 8–12s Move into a tight cinematic portrait or upper-body shot. The character performs one subtle signature action that fits their personality, such as lifting the chin, turning the head, adjusting clothing, brushing hair aside, opening a hand, looking toward camera, or shifting posture. Keep the motion minimal and intentional. The expression should match the character’s vibe. 12–15s End with a strong full-body hero shot that clearly presents the entire design and silhouette. Use a low-angle, eye-level or slightly dramatic framing depending on the character’s personality. The character settles into a natural final pose and holds it confidently for a clean final reveal. VISUAL DIRECTION Premium cinematic presentation. Match the visual medium and rendering style of @[char ref]. Emphasize clean silhouette, elegant staging, subtle secondary motion, believable hair and cloth movement, strong composition, atmospheric depth and polished lighting. The scene should feel like a high-end anime, game or film character introduction. CAMERA Use a clear progression from detail reveal to partial reveal to face reveal to full-body hero reveal. Camera movement should be smooth, controlled and intentional. Avoid chaotic motion. ENVIRONMENT Place the character in a fitting environment that supports their identity and mood. The background should enhance the character without distracting from them.

Qué debes aportar

  • Imagen 1: tu referencia de personaje — una ficha de diseño o un render limpio de cuerpo entero; el vídeo hereda de ella la identidad, el vestuario, los materiales y el estilo de render.

Por qué funciona

  • La escalera de revelación (detalle → parcial → rostro → héroe de cuerpo entero) es una gramática de cámara fija, así el modelo gasta su varianza en el personaje, no en el plan de planos.
  • «Match the visual medium and rendering style of the reference» hace que un mismo prompt funcione igual para personajes anime, de render de videojuego o de imagen real.
  • La única ranura de «signature action» del tercer bloque es donde vive la personalidad: un gesto, deliberadamente pequeño.

Variables sustituibles

  • ficha de referencia del personaje
  • la acción característica
  • atmósfera del entorno
  • encuadre de la pose final

Restricciones

  • El post original escribe las referencias como @[char ref]; en la ruta de EvoLink los recursos se direccionan por posición del array — di «Image 1» y se mapea a image_urls[0].
  • Un personaje, un entorno: esta plantilla nunca corta de localización a propósito, y por eso la silueta se mantiene estable.

Ajustes

Referencia multimodal · 15s · 1:1 · 1 recurso de referencia

Prompts de MiniMax H3: Cine y VFX

Tráiler de ciencia ficción generado con MiniMax H3: una figura solitaria ante un enorme portal cósmico circular mientras un título emerge de la oscuridadREFERENCIA MULTIMODAL
15s16:92 recursos de referencia

Cine y VFX

Tráiler de misterio y ciencia ficción

Tráiler con dos referencias: una imagen fija la atmósfera, otra al protagonista, y el prompt controla avance de cámara, título y sonido.

Ver detalles

Prompt completo (original inglés verificado)

Realistic cinematic look, high-contrast lighting, and a tight pace. Use Figure 1 as the overall atmosphere and style reference, and Figure 2 as the protagonist reference. Shot 1 — Ultra-wide establishing shot. A huge circular cosmic gateway nearly fills the frame. The person is only a tiny figure seen from behind before the gateway, positioned toward the lower right. The ground is wet and reflective, and the center of the gateway is pitch black. The camera slowly pushes forward. A large title fades in from the edge of the darkness, blurred at first and then sharp: "THE STARS WERE LISTENING". Use an extremely condensed, heavy, all-caps typeface in dark red mixed with rust red, with subtle grain and misted edges. Audio: a deep low-frequency pulse, faint metallic vibrations in the distance, and a soft hit as the text becomes sharp. → Hard cut.

Qué debes aportar

  • Imagen 1: la lámina de atmósfera y estilo — entorno, paleta y etalonaje que debe heredar el plano.
  • Imagen 2: el protagonista, con un encuadre lo bastante amplio para que rostro y silueta sigan siendo legibles cuando la figura ocupe poco espacio.
  • Un texto de título que de verdad quieras incrustar: H3 renderiza exactamente lo que escribas, así que sé breve y comprueba la ortografía.

Por qué funciona

  • Cada referencia tiene una sola función — la Imagen 1 la atmósfera, la Imagen 2 el personaje — y así el modelo nunca tiene que adivinar cuál define el look.
  • El plano se describe como un único avance continuo con un solo acontecimiento (el título que se vuelve nítido): justo lo que caben en 15 segundos.
  • El sonido se escribe en tres capas concretas (pulso grave, vibraciones metálicas lejanas y un golpe sobre el título) en lugar de una vaga «música de cine».

Variables sustituibles

  • texto del título y tratamiento tipográfico
  • portal o punto de referencia del plano general
  • color del título
  • colchón sonoro
  • comportamiento del corte final

Restricciones

  • Refiérete a los recursos por su posición en el array — «Figure 1», «Figure 2» — siguiendo el orden de image_urls. La sintaxis @image1 no forma parte de este contrato de API.
  • El prompt publicado es deliberadamente un punto de partida: termina en un corte seco para que puedas encadenar tu propio segundo plano.

Ajustes

Referencia multimodal · 15s · 16:9 · 2 recursos de referencia

Clip de texto a vídeo generado con MiniMax H3: una cocina al anochecer filmada a mano mientras una criatura luminosa dibujada a mano se mueve entre los objetosTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Cine y VFX

Criatura luminosa en la cocina

Plantilla de texto a vídeo que mezcla imagen real y animación dibujada, con defectos de cámara y exclusiones precisas.

Ver detalles

Prompt completo (original inglés verificado)

15-second, 16:9 landscape video. Blend live-action footage of a small kitchen at dusk with hand-drawn glowing animation. The last light of sunset lingers by the window. The lived-in kitchen contains an old wooden table, a half-washed mug, a slightly fogged glass bottle, and a hanging dishcloth. Give the footage subtle one-handed smartphone shake, hesitant close-range focusing, exposure fluctuations caused by backlight, and slightly coarse noise in the shadows. It should not look carefully arranged like an advertisement; instead, it should feel like someone hurriedly captured an unbelievable event at home. Do not show huge eyes, gaping mouths, fangs, threatening or lunging movements, sudden black frames, or jump scares. Use only kitchen room tone, cloth rubbing, the soft clink of a mug, water dripping from the faucet, the camera operator's footsteps and quiet breathing, plus gentle electronic sounds and tiny calls from the hand-drawn creature.

Qué debes aportar

  • Nada que subir: esta ruta funciona solo con el prompt.
  • La duración y la relación de aspecto se fijan en los parámetros de la petición, no solo en el texto del prompt.

Por qué funciona

  • Los defectos de la cámara están nombrados — temblor a una mano, enfoque dubitativo, oscilaciones de exposición a contraluz, ruido grueso en las sombras. Eso es lo que vende «alguien lo grabó deprisa en casa» y no «esto es un anuncio».
  • Una lista de prohibiciones corta y explícita (nada de ojos enormes, colmillos, embestidas ni sustos) evita que una criatura simpática derive hacia el terror.
  • La dirección de sonido nombra cada fuente por separado: tono de sala, tela, taza, grifo, pasos, respiración y los sonidos de la criatura.

Variables sustituibles

  • estancia y momento del día
  • diseño y comportamiento de la criatura
  • objetos sobre la mesa
  • qué imperfecciones muestra la cámara
  • capas de sonido

Restricciones

  • El prompt declara su duración y encuadre en el texto; la API sigue necesitando duration y aspect_ratio como campos de la petición.
  • Las instrucciones negativas funcionan mejor como lista corta y explícita. Acumular decenas de prohibiciones consume el presupuesto que necesitas para la acción.

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Carrera anime con MiniMax H3 generada desde un primer fotograma: dos motos deslizadoras se intercambian posiciones en una horquilla de montaña mojada dejando estelas de luz cian y carmesíPRIMER / ÚLTIMO FOTOGRAMA
15s16:91 recurso de referencia

Cine y VFX

Carrera anime de motos deslizadoras

Una carrera de anime deportivo dirigida por el primer fotograma con un CRITICAL ENTITY LOCK: exactamente dos pilotos con nombre sobre dos motos totalmente especificadas, seguidos por horquillas cronometradas, rebufos y un final al photo-finish.

Ver detalles

Prompt completo (original inglés verificado)

Cinematic Anime Video Scene Generate a 15-second horizontal 16:9 original high-speed hover-bike racing anime video from the provided first frame. CRITICAL ENTITY LOCK: There must be exactly 2 racers and 2 bikes in the entire video: RENJI on VALKYRIE-01 (cyan/black drift bike) and ELENA on AERO-X (crimson/white draft bike). Do not add extra racers, drone support vehicles, spectators, or traffic. Maintain total visual consistency for both bikes, helmet visors, suit patterns, repulsor spark colors, and bike liveries throughout the sequence. Entity identity: VALKYRIE-01: Matte-black and cyan angular hover-bike, exposed repulsor pads, lateral drift brakes, blue plasma exhaust trails, ridden by Renji (cyan trim suit). AERO-X: Pearl-white and neon-crimson aerodynamic hover-bike, enclosed canopy, crimson energy draft aura, white-hot central booster, ridden by Elena (crimson/gold visor suit). Video style: High-budget modern sports anime, sakuga-level velocity animation, crisp line art, vibrant neon lighting contrast, high-speed camera tracking, hyper-realistic friction and energy particle effects. Set on a wet downhill mountain pass at dawn. Camera and pacing: Continuous forward velocity, zero slow-motion interruptions: 0.0s - 3.0s: High-speed rear-tracking shot diving into the first downhill hairpin curve; instant drift initiation. 3.0s - 7.5s: Tight side-parallel tracking shot as bikes navigate rock debris and trade positions through S-curves. 7.5s - 11.5s: Close camera lock on the draft-slingshot maneuver; high-energy particle displacement as booster ignition occurs. 11.5s - 15.0s: Low-angle front-facing camera lock on the final sprint to the finish line bridge, ending on a hyper-speed photo-finish freeze. Action timing: 0.0s - 1.5s: Sequence begins at speed. VALKYRIE-01 leads downhill; AERO-X locks onto its rear bumper. Anti-gravity repulsors spray road water and blue sparks into the frame. 1.5s - 4.0s: First sharp hairpin. VALKYRIE-01 deploys lateral drift airbrakes with a burst of blue thruster fire, sliding sideways at 300 km/h. AERO-X stays glued inside its slipstream aura. 4.0s - 7.0s: Mountain debris hazard. VALKYRIE-01 hops over a boulder using a repulsor burst. AERO-X ducks under it, scraping the neon magenta guardrail in a cloud of friction sparks. 7.0s - 10.0s: S-Curve exchange. Bikes lean side-by-side; their repulsor fields collide, creating a bright electrical shockwave. ELENA pulls the overdrive lever; AERO-X's rear fins extend. 10.0s - 13.0s: Slingshot maneuver. AERO-X bursts out of VALKYRIE-01's draft, igniting its central white plasma booster. Both bikes roar down the final straightaway side-by-side. 13.0s - 15.0s: Final sprint toward the finish light gate. Water sprays violently behind them. Both nose cones cross the finish line simultaneously in a flash of light. Final freeze frame. Motion quality: Fluid 2D animation, extreme speed-line integration, stable bike geometry, flawless vehicle reflection rendering, zero limb or body clipping, high-frame-rate kinetic realism. Environment: Wet mountain pass asphalt, sheer cliff walls, neon cyan and magenta guardrail lights, early dawn sky with pink/purple clouds, water spray, floating spark particles. Final output: 15 seconds, horizontal 16:9, original high-budget sports racing anime, exactly 2 racers, relentless kinetic pacing, dynamic cinematography, no subtitles, no watermarks, no logos.

Qué debes aportar

  • Fotograma inicial: una imagen fija de los dos corredores y sus motos en tu estilo artístico — toda la secuencia se extiende desde esta imagen.

Por qué funciona

  • El bloqueo de entidades por censo («exactly 2 racers and 2 bikes… no spectators, no traffic») elimina la inflación de multitudes que suele golpear a los prompts de carreras.
  • Ambas máquinas reciben libreas, peculiaridades físicas y trajes de piloto como identidades con nombre, así el modelo puede distinguirlas a toda velocidad.
  • Cámara y acción viven en líneas de tiempo separadas que se referencian mutuamente, manteniendo un ritmo implacable sin pedir nunca dos movimientos de cámara a la vez.

Variables sustituibles

  • estilo artístico del fotograma inicial
  • identidades y libreas de las motos
  • peligros del trazado
  • puesta en escena de la línea de meta

Restricciones

  • Esta es la ruta de primer fotograma: la imagen inicial lleva el estilo artístico, e imagen a vídeo no acepta el parámetro aspect_ratio — el fotograma lo define.

Ajustes

Primer / último fotograma · 15s · 16:9 · 1 recurso de referencia

Prompts de MiniMax H3: Audio nativo y diálogo

Sustitución de diálogo generada con MiniMax H3: la frase de un personaje se ha reemplazado por otra procedente de un clip de audio de referenciaREFERENCIA MULTIMODAL
10s16:92 recursos de referencia

Audio nativo y diálogo

Sustitución de diálogo e interpretación

Sustituye una frase citando literalmente la original y la nueva, y limita cuánto puede cambiar la actuación.

Ver detalles

Prompt completo (original inglés verificado)

Replace the girl's line in Video 1, "We can't be together. It's not that we don't love each other; we truly can't make it to the end," with the line from Audio 1: "Don't go, okay? This time, let's not let go of each other." Slightly adjust the corresponding performance.

Qué debes aportar

  • Video 1: el clip que contiene la frase a sustituir.
  • Audio 1: la frase nueva, WAV o MP3, hasta 15 MB y 15 segundos.
  • Ambas frases escritas literalmente en el prompt.

Por qué funciona

  • Citar la frase saliente le dice al modelo exactamente qué tramo del clip debe intervenir, en lugar de «el diálogo que hay hacia la mitad».
  • Citar la frase entrante le da un objetivo a la sincronía labial en vez de deducirla solo del audio.
  • «Ajustar ligeramente la interpretación» concede un permiso acotado para cambiar la actuación: acotado, para que el resto de la toma sobreviva.

Variables sustituibles

  • clip de origen
  • frase sustituida
  • frase nueva y voz
  • cuánto puede cambiar la interpretación

Restricciones

  • El audio nunca puede ser el único tipo de referencia; debe llegar acompañado de una imagen o un vídeo.
  • El audio y el vídeo de referencia tienen un tope de 15 segundos de duración total por petición.
  • Di qué se mantiene fijo. El encuadre, el vestuario y el fondo se desvían si el prompt solo habla de la frase.

Ajustes

Referencia multimodal · 10s · 16:9 · 2 recursos de referencia

Clip de referencia de voz generado con MiniMax H3: un personaje pronuncia una frase escrita con el timbre tomado de un clip de audio de referenciaREFERENCIA MULTIMODAL
10s16:92 recursos de referencia

Audio nativo y diálogo

Referencia de voz «Sigue el viento»

El prompt mínimo con referencia de audio: la frase exacta y un clip que define el timbre.

Ver detalles

Prompt completo (original inglés verificado)

Character dialogue: "Follow the wind, live free. Leave worries behind, enjoy the moment." Use Audio 1 as the voice-timbre reference.

Qué debes aportar

  • Video 1: el personaje que pronunciará la frase.
  • Audio 1: una muestra limpia de la voz objetivo, de 2 a 15 segundos, idealmente sin música de fondo.
  • La frase exacta, escrita por completo.

Por qué funciona

  • El diálogo se cita en lugar de parafrasearse, así el tempo y la sincronía labial tienen algo concreto a lo que agarrarse.
  • A Audio 1 se le asigna una única función acotada — el timbre — en vez de entregarlo como «banda sonora» genérica.
  • No se especifica nada más, así que el clip de referencia conserva el control total del encuadre y la interpretación.

Variables sustituibles

  • frase de diálogo
  • referencia de voz
  • clip del personaje
  • velocidad de habla

Restricciones

  • Una referencia de voz transfiere timbre, no acento, emoción ni ritmo; escríbelos en el prompt si importan.
  • La música o varios hablantes solapados en el audio de referencia degradan el resultado: aporta una voz aislada.

Ajustes

Referencia multimodal · 10s · 16:9 · 2 recursos de referencia

Montaje con MiniMax H3 desde referencias de personaje y audio: un personaje cornudo de pelo blanco cortado entre cinco entornos en planos rápidos sincronizados con el beatREFERENCIA MULTIMODAL
15s1:12 recursos de referencia

Audio nativo y diálogo

Montaje de entornos sincronizado con audio

Un montaje de doble referencia donde una imagen de personaje fija la identidad y un clip de audio dicta el montaje: cinco entornos, seis planos en ráfaga cada uno, y cada corte cayendo sobre los acentos de la pista.

Ver detalles

Prompt completo (original inglés verificado)

Use @[char ref] as the strict character reference and @[audio ref] as the timing, rhythm and editing reference. Keep the character’s exact identity, proportions, hairstyle, outfit, colors and overall style consistent throughout. Create a 15-second cinematic burst-cut video showcasing the character across 5 different environments that naturally fit their design, vibe and world. AUDIO SYNC Synchronize the entire edit to @[audio ref]. Cuts, camera accents, transitions and environment changes should land precisely on strong beats, half-beats and musical accents. Let audio1 control the pacing and intensity of the montage. STRUCTURE - 5 environments total - 3 seconds per environment - 6 burst-cut shots per environment - 30 shots total Each environment must be clearly different in atmosphere, lighting, scale and visual language. Show each environment through rapid cinematic angles: wide establishing shots, aerials, low angles, side views, tracking shots, close environmental details, medium shots and hero frames. Every cut must reveal a new angle, distance, composition or spatial relationship. Avoid repeated framing. Mix static shots, push-ins, pull-backs, tracking, orbit and crane-like movement. Keep character movement subtle and natural. The focus is environmental variety, cinematic framing and tight synchronization with audio1. Hard constraints: - exactly 5 environments - exactly 6 shots per environment - exactly 30 shots total - environment changes must follow audio1’s musical phrasing - cuts and motion accents synchronized to audio1 - no outfit changes - no character duplication - no morphing - no text or UI - no blurry unreadable frames - maintain strict character consistency

Qué debes aportar

  • Imagen 1: el personaje cuya identidad debe conservar cada plano.
  • Audio 1: la pista que manda en el ritmo — cortes, transiciones y cambios de entorno siguen sus beats.

Por qué funciona

  • Reparte el trabajo limpiamente entre modalidades: la imagen responde al «quién», el audio al «cuándo» — ninguno pelea con el texto.
  • La aritmética exacta (5 entornos × 6 planos = 30 cortes) se enuncia como restricción dura, convirtiendo un montaje vago en una estructura contable.
  • «El movimiento del personaje se mantiene sutil» empuja toda la energía hacia la variedad de cámara, que los montajes soportan mucho mejor que la variedad de acción.

Variables sustituibles

  • referencia del personaje
  • pista de audio y su fraseo
  • los cinco entornos
  • mezcla de tipos de plano

Restricciones

  • El post original escribe @[char ref] y @[audio ref]; en EvoLink direcciónalos por posición del array — Image 1 y Audio 1 — y recuerda que el audio nunca puede ser el único tipo de referencia.
  • Los clips de audio de referencia deben durar de 2 a 15 segundos en esta ruta.
  • El clip publicado mide 8:9, una relación que la API no acepta — pide 1:1 para el mismo encuadre casi cuadrado.

Ajustes

Referencia multimodal · 15s · 1:1 · 2 recursos de referencia

Prompts de MiniMax H3: Juego e interfaz

Animación de interfaz con MiniMax H3 desde nueve referencias: una enciclopedia minimalista de criaturas donde un cursor selecciona a MEADOW CROWN, una criatura esponjosa con cuernos, en un pradoREFERENCIA MULTIMODAL
15s16:99 recursos de referencia

Juego e interfaz

Demo de interfaz de enciclopedia de criaturas

Nueve referencias mapeadas a roles explícitos — una lámina de layout de interfaz más ocho tarjetas de criatura — animadas como demo de enciclopedia con cámara fija donde un cursor recorre las entradas y la última criatura se lo come.

Ver detalles

Prompt completo (original inglés verificado)

Use Image 1 as the exact UI/layout/style reference for the creature encyclopedia screen. Use Images 2–9 as the exact creature references. Map them like this: Image 2 = card A = LUMI HARE Image 3 = card B = CLOUD WISP Image 4 = card C = EMBER FENNEC Image 5 = card D = TIDE BEHEMOTH Image 6 = card E = PETAL VULPIN Image 7 = card F = ORCHARD EYE Image 8 = card G = MEADOW CROWN Image 9 = card H = FROST GLIDER Create a 15-second 16:9 video. Keep the camera locked. Keep the interface, layout, typography, panels, icons and overall composition stable, elegant and readable. The UI should feel like a modern minimal digital creature encyclopedia, similar to a sleek pokedex. No scene cuts, no extra text, no extra buttons, no UI distortion. Sequence: 0–2.5s: Cursor clicks card A. Main creature becomes LUMI HARE. Title changes to “LUMI HARE”. Creature blinks and rotates slightly. 2.5–5s: Cursor clicks card C. Main creature becomes EMBER FENNEC. Title changes to “EMBER FENNEC”. Cursor drags to rotate it left and right. 5–7.5s: Cursor clicks card E. Main creature becomes PETAL VULPIN. Title changes to “PETAL VULPIN”. Cursor pokes it a few times. It reacts, annoyed. 7.5–10s: Cursor clicks card G. Main creature becomes MEADOW CROWN. Title changes to “MEADOW CROWN”. Cursor taps near the face/horns. It recoils slightly. 10–12s: Cursor clicks card D. Main creature becomes TIDE BEHEMOTH. Title changes to “TIDE BEHEMOTH”. Cursor keeps poking it. 12–15s: TIDE BEHEMOTH gets angry, opens its mouth very wide, lunges forward, and swallows the cursor. Then it returns to idle. Title stays “TIDE BEHEMOTH”. Rules: - When a card is selected, both the main creature and the main title must update. - Only animate cursor, selection state, title change, and the selected creature. - Only one cursor. - Keep motion subtle and clean until the final swallow. - No cuts, no camera move, no UI distortion, no extra text. Audio: soft UI click sounds, subtle hover sounds, tiny creature reaction sounds, then a sharper aggressive creature sound and one comedic swallow gulp at the end.

Qué debes aportar

  • Imagen 1: el layout de la enciclopedia que define tipografía, paneles y composición.
  • Imágenes 2–9: una criatura por tarjeta, cada una nombrada en el mapa de tarjetas del prompt.

Por qué funciona

  • La tabla de mapeo imagen-tarjeta (Image 2 = tarjeta A = LUMI HARE…) es la asignación de roles más literal posible — el modelo nunca adivina qué recurso es cuál.
  • La cámara está fija y solo cuatro cosas pueden animarse (cursor, selección, título, criatura activa), reduciendo casi a cero la superficie de fallo del movimiento de interfaz.
  • El beat cómico del final (la criatura se traga el cursor) se programa en último lugar para que la demo se mantenga limpia hasta el remate.

Variables sustituibles

  • lámina de estilo de la interfaz
  • las ocho criaturas y sus nombres
  • guion de interacción
  • conjunto de efectos de sonido

Restricciones

  • Nueve imágenes es el máximo por tipo en esta ruta — y el tope total de 12 archivos implica que no puedes sumarle además tres vídeos y tres clips de audio.
  • La estabilidad del texto de la interfaz depende de «sin movimiento de cámara, sin distorsión de la UI»; liberar la cámara reintroduce paneles deformados.

Ajustes

Referencia multimodal · 15s · 16:9 · 9 recursos de referencia

Clip estilo gameplay con MiniMax H3: una vista en primera persona a través de la mira de un rifle avanza por una base militar humeante con un HUD de FPS genéricoTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Juego e interfaz

Simulación de gameplay FPS

Una secuencia de shooter en primera persona que se lee como gameplay capturado: gramática de cámara controlada por el jugador, un HUD genérico totalmente especificado y un ritmo escrito como táctica y no como coreografía.

Ver detalles

Prompt completo (original inglés verificado)

Camera: First-person perspective at eye level with authentic handheld player movement, as if recorded directly from a modern AAA military shooter. The player carries a highly detailed assault rifle with realistic animations, visible hands, tactical gloves, dynamic reload mechanics, and weapon sway. **Opening Action:** The video immediately begins with the player already aiming down a roadway inside a modern military base. Multiple enemy soldiers are visible in the distance near sandbags, barricades, and military vehicles. The player carefully tracks one target, making small aim corrections while maintaining ADS (aim down sights). Fire several controlled bursts immediately at the visible enemies, producing realistic muzzle flashes, shell casings ejecting, smoke, recoil, hit reactions, and dust impacts around the targets. Continue firing in multiple short bursts while adjusting aim between enemies, simulating authentic FPS gameplay rather than scripted animation. **Movement:** After the opening firefight, lower slightly from ADS and begin advancing cautiously along the road beside concrete barriers, Hesco walls, and parked military vehicles. Frequently check left and right corners, briefly stop to reacquire targets, then raise the weapon and fire additional controlled bursts whenever enemies appear ahead. Continue pushing forward with deliberate player-controlled movement, using cover naturally and maintaining believable tactical pacing. **Environment:** Large modern military base with guard towers, armored vehicles, shipping containers, blast barriers, damaged buildings, smoke plumes, burning debris, scattered shell casings, dust clouds, and atmospheric battlefield haze. Cool natural daylight mixed with smoke and orange firelight creates a cinematic battlefield atmosphere. **Camera Motion:** Authentic player-controlled movement with subtle head bob, weapon sway, natural mouse-look adjustments, small left-right corrections while aiming, realistic recoil, smooth tracking of moving targets, brief pauses before shooting, and fluid forward progression. Avoid cinematic camera moves—everything should feel like genuine live gameplay captured by a skilled player. **Visual Quality:** Ultra-photorealistic, AAA game graphics with realistic PBR materials, detailed weapon models, physically accurate lighting, volumetric smoke, dynamic particle effects, crisp textures, realistic bullet impacts, muzzle flash illumination, motion blur only during rapid movement, and high-end military shooter presentation. **Gameplay UI:** Display a realistic modern FPS HUD inspired by games like PUBG, Battlefield, or Call of Duty (without copying exact copyrighted assets). Include: * Central dynamic crosshair or reticle * Ammo counter with magazine and reserve ammunition * Fire mode indicator * Compass at the top * Squad/team status panel * Mini-map in the upper corner * Health bar * Tactical equipment icons (grenades, medkit) * Hit markers when bullets connect * Directional damage indicators * Kill notification feed * Objective marker in the distance * Subtle interaction prompts and realistic HUD animations The HUD should feel polished, modern, and fully integrated into the gameplay, enhancing the illusion of authentic recorded footage from a contemporary military FPS. #MiniMaxH3

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • El objetivo de realismo es «grabado por un jugador hábil», así el balanceo de cabeza, las correcciones de puntería y el ritmo de pausa y disparo se especifican como comportamiento de cámara.
  • El HUD se detalla hasta los marcadores de impacto y el kill feed sin dejar de ser explícitamente genérico — lo bastante denso para leerse como un juego real, lo bastante seguro para publicarse.
  • «Avoid cinematic camera moves» es la inversión clave: lo que la mayoría de los prompts busca es exactamente lo que rompería este.

Variables sustituibles

  • entorno y estilismo de las facciones
  • conjunto de elementos del HUD
  • ritmo de los enfrentamientos
  • clima y luz

Restricciones

  • El prompt mantiene el HUD «inspirado en, sin copiar» — conserva esa cláusula; clonar el HUD de un juego concreto es una tarea distinta (y más arriesgada).

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Prompts de MiniMax H3: Vlog y cámara selfie

Clip estilo selfie con MiniMax H3: una mujer que se graba en un bosque gira su móvil para revelar un ovni estrellado y humeante entre los árbolesTEXTO A VÍDEO
15s1:1Sin recursos de referencia

Vlog y cámara selfie

Descubrimiento de un ovni a cámara selfie

Una demostración de realismo de móvil: metraje selfie en mano con respiración de autofoco y rolling shutter, diálogo guionizado en japonés y un giro de cámara del rostro al ovni estrellado.

Ver detalles

Prompt completo (original inglés verificado)

Ultra photorealistic live-action captured on an iPhone 17. Authentic handheld selfie footage with premium cinematic documentary color grading, realistic HDR, deep green foliage, warm sunlight, subtle teal shadows, natural skin tones, gentle filmic contrast, rolling shutter, autofocus breathing, slight motion blur, and natural handheld shake. A lush forest in daytime with dense trees, wild plants, an uneven dirt trail, scattered leaves, soft sunlight through the canopy, and a gentle breeze. The atmosphere is quiet and slightly unsettling. A cute Japanese woman in her early twenties wearing a stylish bikini walks through the forest while recording herself in selfie mode. She suddenly notices something ahead, looks shocked, turns the camera, and points into the distance. A large crashed UFO is partially embedded in the forest floor. Its metallic hull is badly damaged with broken panels, scorch marks, exposed internal structures, thick gray smoke, and occasional sparks. She says in Japanese: 「ちょっと待って! あそこ見て! UFOじゃない!? 完全に墜落してるんだけど! 煙まで出てる! やばい、本物かもしれない! ちょっと近づいてみる!」 She alternates between filming herself and the UFO while continuing to point at it. Continuous single take. Natural walking movement, realistic hand tremors, slight framing imperfections, quick pans, and autofocus shifts between her face and the UFO. Natural sunlight creates cinematic highlights, soft shadows, realistic reflections on the UFO, subtle volumetric light, and realistic smoke. Audio: footsteps on leaves, gentle wind, birds becoming quieter near the crash site, creaking branches, faint electrical crackling from the UFO, and distant eerie unidentified animal calls echoing through the forest. Negative: no blood, no visible aliens, no monsters, no horror creature reveal, no excessive explosions, no CGI, no cartoon style, no text, no subtitles, no watermark, no logo.

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • Los artefactos del móvil se enumeran (búsqueda de autofoco, rolling shutter, temblor de mano, imperfecciones de encuadre) — el realismo nace de defectos con nombre, no de la palabra «realistic».
  • El diálogo se cita literalmente en japonés, así el audio nativo de H3 genera habla real con sincronía labial correspondiente en lugar de galimatías.
  • El audio se estratifica de forma diegética — pasos, viento, pájaros que enmudecen, chisporroteo eléctrico — sin banda sonora que rompa la ilusión de móvil encontrado.

Variables sustituibles

  • frases habladas e idioma
  • el objeto descubierto
  • escenario de bosque o urbano
  • vestuario y estilismo del personaje

Restricciones

  • La lista negativa (sin alienígenas, sin revelación de terror) es estructural: mantiene el clip en territorio de teaser e impide que el modelo escale la escena.
  • El clip publicado salió en 1:1 — fija el encuadre con el parámetro de aspecto y conserva en el texto la alternancia de encuadre entre ella y el objeto.

Ajustes

Texto a vídeo · 15s · 1:1 · Sin recursos de referencia

Clip documental con MiniMax H3: una fotógrafa callejera encuadra a un hombre mayor y su terrier frente a una cafetería y luego muestra al espectador la foto capturadaTEXTO A VÍDEO
15s16:9Sin recursos de referencia

Vlog y cámara selfie

El instante de la fotógrafa callejera

Un beat documental compacto con una cámara dentro de la cámara: una fotógrafa encuadra una escena callejera espontánea, dispara y luego gira su cámara para enseñar al espectador la foto que acaba de capturar.

Ver detalles

Prompt completo (original inglés verificado)

A young Western female street photographer walks through a lively downtown street and notices an elderly man sitting outside a café with his small dog. She carefully composes the candid moment through her camera, captures the photo, then turns the camera toward the viewer to proudly show the shot she just took. She smiles, says “Look at that,” then continues walking through the city. Ultra-photorealistic visuals, natural handheld documentary movement, realistic camera interaction, authentic facial expressions, accurate hand movements, realistic dog behavior, natural daylight, cinematic depth of field, continuous character consistency, immersive city ambience, premium documentary realism.

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • El beat de enseñar la foto obliga a H3 a renderizar una imagen fija coherente dentro del vídeo — una demostración silenciosa de capacidad que se lee como un gesto natural.
  • La cadena de interacción está totalmente especificada (componer → capturar → girar → «Look at that» → seguir andando), así el clip tiene un arco completo en una sola toma.
  • Los sujetos se describen por rol, no por identidad — un hombre mayor, un perro pequeño, una cafetería — manteniendo la escena callejera genérica y segura.

Variables sustituibles

  • el sujeto espontáneo
  • ciudad y luz
  • la frase hablada
  • tipo de cámara como atrezo

Restricciones

  • La foto dentro de la cámara debe coincidir con la escena de la que se tomó; si cambias los sujetos, cambia ambas descripciones a la vez.

Ajustes

Texto a vídeo · 15s · 16:9 · Sin recursos de referencia

Vlog de viajes con MiniMax H3: una mujer abre las cortinas hacia una terraza junto al mar y luego pasa a modo selfie para dar los buenos días con el océano detrásTEXTO A VÍDEO
15s9:16Sin recursos de referencia

Vlog y cámara selfie

Arco de vlog matinal junto al mar

Un vlog matinal segundo a segundo con un cambio de modo deliberado: los primeros cinco segundos son tercera persona cinematográfica, y solo entonces el personaje empieza a grabarse — el momento en que el «vlog» comienza.

Ver detalles

Prompt completo (original inglés verificado)

Create a 15-second ultra-realistic cinematic lifestyle vlog video, vertical 9:16, featuring the same young woman throughout the entire video. Preserve her facial identity, facial proportions, hairstyle, skin tone and overall appearance consistently in every shot. She wears the same outfit throughout: fitted white V-neck T-shirt with a small subtle logo, blue denim jeans, natural makeup, long softly wavy brown hair. 0:00–0:01 — Wake-up: Close-up inside a beautiful bright bedroom. The woman is lying comfortably on the bed, slowly wakes up, stretches naturally and opens her eyes. She is NOT filming a vlog yet and does not hold a phone or camera. Soft morning sunlight enters through the curtains. 0:01–0:02 — Gets up: Medium shot. She sits up on the bed, smiles softly, fixes her hair and gets ready to start her morning. Natural, effortless movement. 0:02–0:03 — Walks to window: She walks toward the large glass balcony door/window. Camera follows her naturally from behind/side. 0:03–0:04 — Seaside reveal: She opens the curtains/door and looks outside. Reveal a breathtaking blue ocean, coastal hills, flowers, balcony and beautiful morning sunlight. She smiles happily while taking in the view. 0:04–0:05 — Steps outside: She walks out onto the seaside terrace. Gentle ocean breeze moves her hair naturally. Wide cinematic shot showing the beautiful surroundings. 0:05–0:06 — VLOG START: Only now she starts filming herself in handheld selfie-vlog style. She looks into the camera with a bright natural smile and says: “Good morning!” 0:06–0:07 — Show the view: She turns the camera away from herself and slowly pans across the stunning ocean, coastal mountains, flowers and terrace. Smooth handheld vlog movement. 0:07–0:08 — Back to selfie: Selfie shot. She looks into the camera and happily says: “This place is just perfect!” 0:08–0:09 — Location reveal: Wide cinematic shot of the cozy seaside terrace with wooden table, chairs, plants and flowers overlooking the ocean. 0:09–0:10 — Walk to table: Medium tracking shot as she walks toward the table, enjoying the view. Her hair and T-shirt move gently in the sea breeze. 0:10–0:11 — Sit and relax: She sits at the seaside table, smiling peacefully and enjoying the ocean view. A refreshing orange-colored juice is placed on the table. 0:11–0:12 — Juice close-up: Cinematic close-up of her hand picking up the glass of fresh orange juice. Beautiful ocean bokeh in the background, natural sunlight reflecting through the glass. 0:12–0:13 — Vlog toast: Selfie shot. She raises the juice toward the camera with a cheerful smile and says: “Cheers to good days!” 0:13–0:14 — Happy close-up: Beautiful close-up of her smiling naturally at the camera, ocean and warm sunlight softly blurred behind her. 0:14–0:15 — Ending: Camera moves from her toward the sparkling ocean and peaceful coastal landscape. Warm sunlight, gentle waves and a relaxing cinematic ending. Overall Style Ultra-realistic, cinematic travel vlog, natural handheld camera movement, realistic human motion, smooth transitions, soft morning sunlight, realistic ocean waves, gentle wind in hair and clothes, beautiful coastal atmosphere, premium lifestyle aesthetic, natural expressions, authentic vlog feeling, shallow depth of field, cinematic composition, realistic skin texture, high detail, 4K quality.

Qué debes aportar

  • Nada que subir: esta ruta de texto a vídeo funciona solo con el prompt.

Por qué funciona

  • La nota «she is NOT filming yet» en los beats del despertar evita el artefacto más común del vlog: un móvil que aparece antes de que el vlog empiece.
  • Quince beats de un segundo alternan cámara selfie y cortes escénicos exactamente como la gramática real del vlog de viajes.
  • Tres frases cortas citadas («Good morning!») dan al audio nativo puntos de control naturales sin un monólogo que sostener.

Variables sustituibles

  • revelación de la localización
  • las tres frases habladas
  • vestuario y fijación de identidad
  • la bebida como atrezo

Restricciones

  • El prompt pide 9:16 y esa es la relación que hay que solicitar; el clip publicado se recodificó aproximadamente a 3:2 — otra razón para fijar la orientación en la petición y no en la prosa.

Ajustes

Texto a vídeo · 15s · 9:16 · Sin recursos de referencia

Prompts de MiniMax H3: Edición y transformación

Edición de vídeo generada con MiniMax H3: en un clip fuente se han sustituido o eliminado el periódico, la silla, las gafas de sol y el coche en llamasREFERENCIA MULTIMODAL
10s16:91 recurso de referencia

Edición y transformación

Edición de escena con varios elementos

Seis cambios independientes sobre un vídeo fuente, escritos como lista objetivo → resultado sin volver a describir la escena.

Ver detalles

Prompt completo (original inglés verificado)

Replace the newspaper in the reference video with a green-covered book; change the chair the character is sitting on to a red sofa; remove the sunglasses worn by the character to retain a clear face; remove the car burning effect to keep the vehicle in a normal state; change the photo the character takes out of his arms to a small black book; and add a tree on the left side of the screen

Qué debes aportar

  • Video 1: el clip a editar, de 2 a 15 segundos, MP4 o MOV, H.264 o H.265, hasta 50 MB.
  • Una lista de ediciones, cada una indicando qué cambiar y en qué se convierte.

Por qué funciona

  • Cada instrucción es un par — objetivo más resultado — y no deja nada a la interpretación («el periódico» pasa a ser «un libro de tapa verde»).
  • Las eliminaciones declaran el estado final buscado («quitar las gafas de sol para dejar la cara despejada»), lo que evita que quede un hueco donde estaba el objeto.
  • No hay ninguna descripción de escena, así que el modelo trata el clip fuente como verdad y solo aplica las diferencias.

Variables sustituibles

  • número de ediciones
  • objetos sustituidos
  • efectos eliminados
  • elementos añadidos

Restricciones

  • EvoLink expone tres rutas H3; el trabajo de edición pasa por reference-to-video con el clip fuente como Video 1.
  • La duración del vídeo de referencia se factura, así que recorta la fuente al fragmento que realmente necesitas antes de subirla.
  • Indica qué debe permanecer intacto cuando una edición esté junto a algo que te importa: las zonas no mencionadas quedan a merced del modelo.

Ajustes

Referencia multimodal · 10s · 16:9 · 1 recurso de referencia

Plano de escenario generado con MiniMax H3: dos magos intercambian el color de sus trajes entre una nube de humo mientras el telón pasa de rojo a azulPRIMER / ÚLTIMO FOTOGRAMA
7s16:91 recurso de referencia

Edición y transformación

Intercambio de vestuario entre magos

Prueba de seguimiento: dos trajes cambian, un detalle permanece intacto y el fondo completa una transición de color.

Ver detalles

Prompt completo (original inglés verificado)

Two magicians stand onstage facing the audience and perform a "swap" trick. They wave their wands at the same time, and a cloud of smoke rises. When it clears, their suit colors have switched: the person on the left wears a white suit, and the person on the right now wears a black suit, while both magicians' glove colors remain unchanged. They bow to thank the audience. The red curtain behind them closes, transitioning from deep red to deep blue.

Qué debes aportar

  • Una imagen inicial con ambos artistas, su vestuario y el escenario.
  • Un estado antes/después claro para todo lo que se intercambie.

Por qué funciona

  • El intercambio queda tapado por un suceso — la nube de humo — lo que da al modelo un momento legítimo para hacer el cambio en vez de transformarlo a la vista.
  • Lo que no debe cambiar (el color de los guantes) está escrito justo al lado de lo que sí cambia: así es como se evita que una edición se extienda.
  • El plano termina en un estado definido: la reverencia, el telón que se cierra, el color que aterriza en azul profundo.

Variables sustituibles

  • artistas y vestuario
  • el detalle que permanece fijo
  • suceso que tapa el intercambio
  • transición de color final

Restricciones

  • La ruta imagen a vídeo deduce la relación de aspecto de la imagen de entrada y rechaza el campo aspect_ratio.
  • Cualquier atributo no mencionado queda a merced del modelo. Si un detalle debe sobrevivir al cambio, escríbelo.

Ajustes

Primer / último fotograma · 7s · 16:9 · 1 recurso de referencia

¿Buscas la API de MiniMax H3?

Página del modelo con precios por segundo, tres rutas de generación y docs de integración.

Ir a la API

El marco de prompts de MiniMax H3

H3 entiende referencias ordenadas y genera su propio audio. Escribe para el flujo concreto en lugar de depender solo de una fórmula genérica.

La estructura de un prompt de H3

Objetivo → referencias ordenadas → identidad del sujeto → acciones cronológicas → recorrido de cámara → audio o diálogo → elementos que no cambian → estado final. Así no se olvidan el sonido ni el cierre.

Texto a vídeo

Describe una secuencia que el clip pueda completar: sujeto, pocas acciones, un recorrido continuo de cámara y capas separadas de diálogo, ambiente y música.

Primer / último fotograma

Las imágenes ya definen el aspecto. Describe solo el cambio entre ellas, el movimiento de cámara, lo que debe conservarse y el estado final.

Referencias multimodales

Asigna una sola función a cada recurso y cita su posición: Image 1 para el personaje, Image 2 para el lugar, Video 1 para el movimiento y Audio 1 para la voz.

Editar un clip existente

Escribe dos listas: «Cambiar» con objetivo → resultado y «Conservar» para todo lo que debe quedar igual. Envía solo el fragmento necesario como Video 1.

Audio y diálogo

Cita literalmente las frases. Usa la referencia de voz solo para el timbre y define por separado el acento, la emoción y el ritmo.

Preguntas frecuentes sobre prompts de MiniMax H3

¿MiniMax H3 y Hailuo 3 son el mismo modelo?

Sí. MiniMax H3 también se conoce como Hailuo 3, Hailuo 3.0 o Hailuo 03. EvoLink usa MiniMax H3 y ofrece rutas de texto a vídeo, imagen a vídeo y referencia a vídeo.

¿Puedo usar estos prompts sin programar?

Sí. «Usar este prompt» envía el prompt original verificado en inglés al Playground de MiniMax H3 y selecciona el flujo adecuado. Allí puedes añadir las referencias.

¿Cuánto puede durar un vídeo de MiniMax H3?

Las tres rutas aceptan duraciones enteras de 4 a 15 segundos y actualmente generan en 2K. La duración de cada tarjeta corresponde al ejemplo publicado.

¿Cuántas referencias puede usar un prompt?

La ruta de referencia admite hasta 9 imágenes, 3 vídeos y 3 audios, con un tope de 12 archivos en total: un conjunto completo de 9 + 3 + 3 se rechaza. Se requiere al menos una imagen o un vídeo; el audio no puede usarse solo. Cada vídeo o audio debe durar entre 2 y 15 segundos.

¿Por qué los prompts dicen «Image 1» o «Video 1»?

La API resuelve los recursos por su posición en image_urls, video_urls y audio_urls. Esas referencias en inglés se conservan en el prompt verificado para asignar una función concreta.

¿De dónde proceden los prompts y los clips?

Cada tarjeta indica su fuente. Los ejemplos oficiales usan una versión inglesa verificada del prompt chino; los ejemplos de la comunidad enlazan la publicación original en X.

¿Puedo usar estos prompts mediante la API?

Sí. El mismo prompt en inglés funciona en el Playground y en el campo prompt de la API. La guía de MiniMax H3 explica solicitudes, tareas asíncronas, callbacks y errores.

Elige un prompt y genera

Cada prompt se asigna a una ruta activa de MiniMax H3 en EvoLink. Envíalo al Playground o crea la misma solicitud con la API unificada.