
Text & Image to Video
Generate high-quality cinematic videos from text prompts, images, or combined references. H3 understands characters, environments, visual styles, and storytelling intent to create consistent scenes.
Create cinematic videos with native multimodal understanding. Generate, edit, and transform videos using text, images, motion references, camera guidance, and audio inputs.
| Prompt | Reference Image | Generated Clip |
|---|---|---|
| Image 1 is the reference for the overall visual texture and atmosphere, while Image 2 is the reference for the character’s appearance. Generate a 15-second 16:9 widescreen trendy fashion film. Maintain character consistency: platinum blonde long hair, narrow black vintage sunglasses, a glossy black patent leather trench coat, a cool and confident fashion expression, with orange reflections from the flames illuminating the surface of the leather. The overall style should be a fast-cut cinematic fashion advertisement, blending a nighttime fire scene, black smoke, orange-red flames, VHS glitches, CCTV broadcast interruptions, 1990s analog film grain, scanlines, chromatic aberration, light leaks, white flash transitions, and subtle camera shake. | ![]() |
| Prompt | Reference Image | Reference Video | Generated Clip |
|---|---|---|---|
| Image 1 is the reference for the overall visual texture and atmosphere, while Image 2 is the reference for the character’s appearance. Generate a 15-second 16:9 widescreen trendy fashion film. Maintain character consistency: platinum blonde long hair, narrow black vintage sunglasses, a glossy black patent leather trench coat, a cool and confident fashion expression, with orange reflections from the flames illuminating the surface of the leather. The overall style should be a fast-cut cinematic fashion advertisement, blending a nighttime fire scene, black smoke, orange-red flames, VHS glitches, CCTV broadcast interruptions, 1990s analog film grain, scanlines, chromatic aberration, light leaks, white flash transitions, and subtle camera shake. | ![]() |
| Prompt | Reference Image | Reference Video | Generated Clip |
|---|---|---|---|
| Replace the child at the very back of Video 1 with the golden retriever from Image 1. Change the khaki jacket worn by the child on the far left side of Video 1 to the denim jacket from Image 2. | ![]() |

Generate high-quality cinematic videos from text prompts, images, or combined references. H3 understands characters, environments, visual styles, and storytelling intent to create consistent scenes.
Generate cinematic videos, edit existing footage, transfer motion, and create synchronized audio-visual experiences with one powerful multimodal AI model.
Explore MiniMax H3