At the WAIC 2026 conference, MiniMax announced H3 — a new generation of multimodal video models. The new model is designed for creating high-quality content in 2K resolution with video lengths of up to 15 seconds, offering native audio integration and advanced tools for controlling visual consistency.

What happened
MiniMax presented the multimodal H3 model, which supports video generation in 2K resolution with a duration of up to 15 seconds. Key technical features include the Omni Reference system for maintaining character consistency, improved handling of 2D animation, and built-in audio generation, including dialogue, sound effects, and lip-sync.
Context
The development of H3 marks a transition from the evolutionary development of the Hailuo 2.x architecture to a fundamentally new approach, where video, sound, and physical accuracy of movements are considered as a single product. The model is positioned as a competitor to leading solutions in the industry, such as Kling and Veo.
Why this matters for the industry
The emergence of H3 intensifies competition in the segment of high-quality video models, pushing the industry to move from specialized tools to comprehensive multimodal solutions (video + audio + consistency). This may accelerate the development of new UX patterns for character management and video production automation.
Why this matters for users
For content creators, the model allows generating longer and more stable videos with the same characters without the need for complex fragment stitching and separate voiceover. Built-in sound handling and lip-sync significantly reduce post-production workload.
What is currently unknown / limitations
At present, there is no information on API availability, usage costs, and latency, which makes it difficult to assess the possibility of integrating the model into real production processes.
Sources
- MiniMax Unveils M3 and H3 Models at WAIC 2026
- MiniMax H3 Review: Features, Quality, Pricing and Early Tests
Author
Look at AI, editorial team
