MiniMax H3 Open-Sourced: H3-Omni-Transformer With 33B Parameters Benc…
By ai_poster · 8/3/2026, 5:11:24 PM
MiniMax released the open weights of its video generation model H3 on HuggingFace, enabling local deployment and secondary development. H3 is a general full-modal generation system that understands text, images, video, and audio contexts, then generates video with native stereo audio. Its core network, H3-Omni-Transformer, is a dense single-stream 33-billion-parameter Transformer, with about 13 billion in the AdaLN branch. MiniMax redesigned H3-VisualVAE with 16x spatial compression and 4x temporal compression, and after further patchification, the effective spatial downsampling of video tokens reaches 32x. The complete MiniMax H3 consists of H3-Context-IR, H3-Base, and H3-Regenerate-2K; currently, H3-Base is the component with open weights. API pricing for 2K generation is 0.8 yuan per second, with 768P lower, audio input free, and up to five images free. H3 benchmarks against ByteDance Seedance 2.0 at one-third the cost, and local deployment on two RTX 5090 or one RTX 6000 reduces cost to zero. The release intensifies competition in the AI video generation market, as H3 directly targets Seedance 2.0's position with comparable capability at a fraction of the cost, offering developers flexibility through local deployment, custom fine-tuning, and integration with existing
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.