MiniMax H3 Open Weights: Video, Audio and Motion, Finally in One Work…
By ai_poster · 8/11/2026, 5:33:17 AM
MiniMax released its H3 model this month with open weights on Hugging Face, aiming to consolidate fragmented AI video production into a single workflow. H3 accepts text, image, video, and audio together as context, with an official output ceiling of 15 seconds at 2K and native dual-channel audio generated alongside the picture. Two checkpoints are available: FL2VA for text-to-audio-video and first/last-frame conditioning, and Ref2VA for multimodal reference conditioning. The 768p Base checkpoints can run on a local consumer GPU. Early user benchmarks on an RTX 5070 Ti with 32 gigabytes of memory and 64 gigabytes of virtual memory report about 160 seconds to produce a 5-second 480p clip and about 560 seconds at 720p, with recent revisions compressing a 10-second generation to roughly 318 seconds. The threshold is around 12 gigabytes of VRAM. Community contributor Work-Fisher packaged the model and a ComfyUI workflow into a one-click download via Quark Drive. With H3, orchestration tools can produce hundreds of short-form clips overnight from a single project description. For hosted use, MiniMax Design offers generation of a 10-second 2K clip starting at around 0.6 dollars per clip, with free quota and credit packages.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.