At the WAIC 2026 conference, MiniMax announced H3 — a new generation of multimodal video models. The new model is designed for creating high-quality content in 2K resolution with video lengths of up to 15 seconds, offering native audio integration and advanced tools for controlling visual consistency.

image

What happened

MiniMax presented the multimodal H3 model, which supports video generation in 2K resolution with a duration of up to 15 seconds. Key technical features include the Omni Reference system for maintaining character consistency, improved handling of 2D animation, and built-in audio generation, including dialogue, sound effects, and lip-sync.

Context

The development of H3 marks a transition from the evolutionary development of the Hailuo 2.x architecture to a fundamentally new approach, where video, sound, and physical accuracy of movements are considered as a single product. The model is positioned as a competitor to leading solutions in the industry, such as Kling and Veo.

Why this matters for the industry

The emergence of H3 intensifies competition in the segment of high-quality video models, pushing the industry to move from specialized tools to comprehensive multimodal solutions (video + audio + consistency). This may accelerate the development of new UX patterns for character management and video production automation.

Why this matters for users

For content creators, the model allows generating longer and more stable videos with the same characters without the need for complex fragment stitching and separate voiceover. Built-in sound handling and lip-sync significantly reduce post-production workload.

What is currently unknown / limitations

At present, there is no information on API availability, usage costs, and latency, which makes it difficult to assess the possibility of integrating the model into real production processes.

Sources

Author

Look at AI, editorial team