Step 01
Write a MiniMax H3 prompt
Describe subject, action, camera movement, lighting, and mood. MiniMax H3 follows complex multimodal instructions with strong instruction-following accuracy.
MiniMax H3 is MiniMax’s omni-modal video model — generate 2K clips with native stereo audio from text, first frames, images, audio, or reference video in one MiniMax H3 workspace.
Direct with MiniMax H3. Render at 2K with native sound.
monster-bakery.pngUpload first frameMiniMax H3 is a general-purpose omni-modal generative model from MiniMax. It understands unified context across text, images, video, and audio, then generates video with native 32 kHz stereo sound at up to 2K resolution.
“A cinematic close-up of a woman turning toward soft window light, with natural movement and subtle facial expression.”
Create MiniMax H3 AI video from text or reference media in four steps — from prompt to downloadable 2K clip.
Step 01
Describe subject, action, camera movement, lighting, and mood. MiniMax H3 follows complex multimodal instructions with strong instruction-following accuracy.
Step 02
MiniMax H3 accepts first and last frames, up to 9 images, 3 video clips, and 3 audio clips — guide identity, motion, sound, and composition in one MiniMax H3 generation.
Step 03
Choose MiniMax H3 resolution up to 2K, duration from 5–15 seconds, aspect ratio, and native stereo audio before you render.
Step 04
MiniMax H3 renders at 24 FPS, then saves your clip with prompt and settings to your private library for preview and download.
Every clip here was created with MiniMax H3 from original prompts and reference frames. Reuse a MiniMax H3 prompt or video to start your next generation.
The enormous cake suddenly rises another foot and releases a soft burst of flour. The little orange monster stumbles backward, catches the wooden spoon, blinks twice in comic disbelief, then breaks into a delighted grin. Flour curls through the warm oven light and copper pans wobble gently. Preserve the original monster design, exactly two teal horns, bakery layout and theatrical family-animation style. Smooth expressive character motion, playful timing, no new characters, no text, no logos.
Epic IMAX-style aerial chase at bright golden hour. An entirely original crimson-masked web acrobat launches from the glass crown of a skyscraper, fires a silver cable-web and swings in one powerful arc through a sunlit canyon of towers above busy traffic. The camera dives behind the hero, passes close to reflective windows, then pulls wide to reveal the glowing city skyline. Athletic motion, believable gravity, dramatic parallax, premium superhero feature-film cinematography, original suit design with no recognizable emblem, no text, no logos.
Bright cinematic stop-motion macro shot of a colorful pop-up book opening by itself on a sunlit children’s art-studio table. A tiny scarlet paper dragon unfolds from the center crease, stretches its accordion wings, then leaps into the air and circles once above the book as cheerful paper trees bend in its wake. Colorful paper scraps spiral through warm midday sunbeams. Handmade paper fibers, practical stop-motion imperfections, playful monster energy, clean pastel background, precise readable motion, no text, no logos, no human hands.
Epic IMAX-style wide shot at a brilliant alpine sunrise. A colossal ancient stone titan slowly rises from beneath a blue glacier as sheets of ice break away and a controlled avalanche rolls into the valley. Two tiny rescue helicopters sweep past the camera to reveal the immense scale while the titan turns toward the first sunlight. Photorealistic feature-film visual effects, majestic rather than frightening, crisp bright atmosphere, natural rock and snow physics, no text, no logos.
GoPro footage of someone airgliding through the mountains, shaky camera footage.
Millions of red poppy flowers slowly grow from a grassy desert landscape under clear skies.
The subway car rocks gently as it moves. The giant sea-green monster carefully shifts its tiny briefcase to one claw, ducks under a hanging rail and gives a small sleepy yawn. Its long fur responds naturally to the train motion while nearby commuters glance up with restrained smiles. The camera makes a slow handheld push forward. Preserve the original friendly creature, passengers and photorealistic practical-effects look. No panic, no extra monster, no text, no logos.
A dragon flying over New York City, cinematic drone shot.
The dancer completes one slow, precise turn in the train aisle while the three porcelain koi orbit her in a smooth counter-clockwise path. Rain continues sliding down the windows, the carriage sways almost imperceptibly, fabric and hair respond naturally, and the camera makes a subtle push forward. Preserve the original face, costume, carriage composition and cool blue lighting. No new people, no extra fish, no text.
The botanist takes two careful steps through the shallow water and raises one hand. The tiny bioluminescent seeds gather into a loose spiral above their palm, casting faint cyan reflections that move across the flooded floor. Ferns stir in a soft draft and the camera drifts slowly to the right. Preserve the greenhouse architecture, yellow raincoat, moonlit atmosphere and realistic scale. No text, no logos.
Macro cinematic shot looking into a rain puddle on a quiet city street. Inside the reflection, a tiny night market comes alive with miniature canvas stalls, steam, warm lanterns and people moving naturally. A raindrop lands and sends a slow ripple across the entire miniature world while the camera gently lowers toward the water. Photorealistic, tactile scale, believable reflections, no text, no logos.
A silver translucent orb floats through the streets of London, handheld camera footage.
A cinematic wide shot of an old coastal library at night as a shallow tide slowly flows between the bookshelves. Loose pages lift gently in the sea breeze, small silver fish pass through reflected moonlight on the floor, and the camera makes one calm forward dolly. Photorealistic practical set, restrained motion, natural water physics, subtle 35mm film grain, no text, no logos.
A cinematic midnight laundromat in winter. Long white sheets slowly rise from open washing machines and billow through the room like sails in a quiet ocean wind. A lone attendant stands still at the back while fluorescent ceiling light mixes with deep blue street light. The camera moves laterally at a steady human pace, revealing realistic cloth physics, subtle reflections on the tile floor and drifting steam. Photorealistic practical set, restrained surrealism, no readable text, no logos.
A ship sailing through rough waters, grey skies, cinematic ocean spray.
A field of glowing asteroids drifts past the camera in deep space, with cinematic light and enormous scale.
A cinematic close-up of a woman turning toward soft window light, with natural movement and subtle facial expression.
MiniMax H3 reads unified text context and generates video with strong instruction following — ideal for product demos, brand films, and narrative beats at up to 2K.
Upload a first frame or start/end frames. MiniMax H3 animates stills while preserving composition, identity, and lighting from your reference image.
MiniMax H3 accepts up to 9 images, 3 videos, and 3 audio clips — guide subject, motion, sound, and style in one omni-modal MiniMax H3 generation.
MiniMax H3 is a general-purpose omni-modal model that understands text, images, video, and audio in shared context for coherent cinematic output.
Set MiniMax H3 resolution up to 2K, duration from 5–15 seconds, aspect ratio, and native 32 kHz stereo audio before you generate.
Every MiniMax H3 render saves with prompt and settings to your private library — preview, download, or reuse clips as MiniMax H3 reference media.
Describe subject, action, camera, lighting, and mood. MiniMax H3 turns your prompt into cinematic 2K footage with instruction-following accuracy for ads, branding, and short-form content.
Upload reference video and let MiniMax H3 transfer motion, camera language, and pacing. MiniMax H3 supports video-to-video workflows for consistent character and product shots.
Animate a first frame or lock start and end compositions with MiniMax H3. Reference images, audio, and video guide identity, sound, and facial expression in one MiniMax H3 render.
Explore prompts and clips generated with MiniMax H3 in this studio. Load any MiniMax H3 example, then edit subject, camera, motion, or style for your next render.
Subscriptions renew automatically and can be canceled anytime. One-time credit packs stay valid for one year.
Lite
For casual creators
Pro
For professionals
Ultra
For power users
MiniMax H3 is MiniMax’s omni-modal generative model for video. It understands text, images, video, and audio in unified context and generates up to 2K clips with native stereo sound.
Yes. MiniMax H3 supports image-to-video from a first frame, plus start and end frames. Add reference images to guide subject identity, composition, and visual style in MiniMax H3.
Most MiniMax H3 renders take several minutes. Jobs run in the background while MiniMax H3 processes motion, lighting, and native audio until your clip is ready.
The MiniMax H3 task is marked as failed, and credits used for that generation are automatically returned to your balance.
MiniMax H3 is built for commercial content — ads, branding, e-commerce, and short-form campaigns. Review your plan terms and provider policies before publishing paid work.
Yes. MiniMax H3 outputs native 32 kHz stereo audio synced to on-screen motion — dialogue, ambience, and effects in one MiniMax H3 pass without a silent placeholder.
Write a prompt or add reference media, then generate a downloadable MiniMax H3 clip at up to 2K with native stereo audio in one workspace.