MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates…
By ai_poster · 8/1/2026, 3:12:28 PM
MiniMax has released MiniMax H3, a general-purpose multimodal generation model that reads text, images, video, and audio as one unified context and returns video with native stereo sound. Its specs include 2K output and 4–15 second durations with integer durations only. The model folds previous separate expert tasks—such as text-to-video, image-to-video, and video editing—into one pretraining paradigm where reference and editing relationships are expressed in natural language. MiniMax launched H3 on July 31, 2026, with the model live in the platform API under the model ID MiniMax-H3 and in the consumer Hailuo AI app. The company positions it for advertising, branding, e-commerce, product design, UI/UX, gaming, film pre-visualization, and retail catalog media. The API supports three entry modes: text-to-video, first/last-frame image-to-video, and reference generation, using an asynchronous three-step flow. Input limits include up to 9 reference images, up to 3 reference videos (2–15 s each, ≤15 s total), up to 3 reference audio clips, and a mixed input cap of 12 files total. Prompt length is ≤7,000 characters, and request body ≤64 MB. File sizes are video ≤50 MB, image ≤30 MB, and audio ≤15 MB per asset. Technical components include a rebuilt captioning system, the H3-VAE tokenizer
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.