✦ MINIMAX H3 · OPEN WEIGHTS · 2K VIDEO · NATIVE AUDIO ✦
MiniMax H3 Prompt
Generator
MiniMax H3 is a precision tool — it responds to structured prompts with shot labels, camera vocabulary and audio direction. Promptnomicon formats your idea into exactly that.
Generate my first H3 prompt →3 free spells every month · No credit card required
What is MiniMax H3?
MiniMax H3 is MiniMax's general-purpose multimodal video model, released with open weights in July 2026. It jointly understands text, images, video and audio in a single context, and generates video with native stereo audio — voice, sound effects and music in one pass.
Output resolution reaches 2K at 24fps with clips up to 15 seconds. The model supports text-to-video, image-to-video with optional last-frame control, and multimodal reference generation with up to 9 reference images and 3 reference videos.
H3 prompts follow a shot-based structure with timestamps, explicit camera vocabulary and a dedicated AUDIO block. Writing effective H3 prompts is closer to directing a scene than describing a picture. Promptnomicon handles that translation automatically.
H3 Prompt Structure
What Promptnomicon generates for you
Example output
[Shot 1] A botanist examines a glowing bioluminescent plant in a dark greenhouse, leaning close with a magnifying glass. Slow push-in from wide to medium close-up. Single practical lamp overhead, cool teal glow from the plant illuminating her face from below. Cinematic, film grain.
[Shot 2] At 00:07.000, the camera cuts to an extreme close-up of the plant pulsing with light. Static, macro lens, very shallow depth of field.
AUDIO: Quiet greenhouse ambience, distant dripping water left, soft hum of ventilation right, plant emits a faint crystalline tone center.
Also generate prompts for:
Generation Modes
Text to Video
Describe your idea in plain language. Promptnomicon generates a structured H3 prompt with shots, camera and audio.
Image to Video — Start Frame
Upload an image as the opening frame. The prompt describes what happens next, continuing the scene naturally.
Image to Video — End Frame
Upload a target image. The prompt describes the action that leads into that final frame.
Start + End Frame
Upload two images. The prompt describes the controlled transition between them — motion, pace and camera.
Frequently Asked Questions
Where can I use MiniMax H3?
H3 is available on Hailuo AI (hailuoai.video), fal.ai/minimax/h3, WaveSpeed AI, Atlas Cloud, and locally via ComfyUI with the official open weights from Hugging Face.
What camera moves does H3 understand?
Push-in, pull-back, slow pan, tracking shot, rack focus, static, Dutch angle, and bird's eye view. Promptnomicon always includes one of these per shot with a speed qualifier.
Does H3 generate audio automatically?
Yes. H3 generates native stereo audio in the same pass as the video. The AUDIO block in the prompt directs what sounds to generate and where to place them spatially.
How many shots should a prompt have?
For 5-second clips: 1 shot. For 8–10 seconds: 2 shots. For 15 seconds: up to 3 shots. Promptnomicon adjusts automatically based on the complexity of your idea.
Do video prompts cost more?
Yes — 2 spells per generation instead of 1, due to the two-stage pipeline (vision analysis for image inputs + structured prompt generation).
Start Creating with MiniMax H3
3 free spells every month. No credit card required.
Create free account →Instant access · No credit card