MiniMax Launches H3, an Omni-Modal Model Generating 2K Video With Native Stereo Audio
MiniMax
MiniMax officially launched H3 on July 31, an omni-modal generation model that jointly understands text, image, video, and audio context to produce up to 15-second 2K video clips with native stereo sound, accepting up to 9 reference images, 3 video clips, and 3 audio clips as controllable inputs. The model is live via MiniMax's API and the Hailuo AI consumer app, with open weights promised in the coming days.
Why it matters
Native audio-video co-generation at commercial 2K quality, with open weights planned, pushes open Chinese labs (competing with ByteDance and Kuaishou) further into a segment previously led by closed Western video models, while undercutting rival pricing by roughly two-thirds.
Importance: 4/5
Major omni-modal video release with open weights planned and 4 independent confirmations (official + 3 media outlets).