Newsletter
Join the Community
Subscribe to our newsletter for the latest news and updates
Create 4–15 second 2K videos with native stereo sound from text, images, video, and audio references.
MiniMax H3 is a general-purpose multimodal AI video model that generates 2K audiovisual videos with native stereo sound. It accepts text, images, video, and audio as input, allowing users to direct characters, motion, camera, style, sound, and pacing in a single prompt.