Article URL: https://minimaxh3.art/blog/what-is-minimax-h3 Comments URL: https://news.ycombinator.com/item?id=49130723 Points: 3 # Comments: 0

On July 31, 2026, Chinese AI company MiniMax officially launched MiniMax H3 — the third-generation model in its Hailuo video family, also known as Hailuo… On July 31, 2026, Chinese AI company MiniMax officially launched MiniMax H3 — the third-generation model in its Hailuo video family, also known as Hailuo 3.0. First previewed at WAIC 2026 just two weeks earlier, H3 arrives with a clear ambition: not just to generate moving images, but to produce complete short-form audiovisual scenes — with native 2K resolution, synchronized sound, and up to 15 seconds of continuous footage in a single generation. If you have been following the AI video space, you know the pace of progress has been relentless. New models arrive every few months, each claiming to push the boundary further. So what exactly does H3 bring to the table, and should you care? Here is everything you need to know. MiniMax H3 is a multimodal AI video-generation model that produces native 2K video at 24fps with built-in synchronized audio — dialogue, sound effects, and ambient atmosphere — from text prompts, images, or a combination of reference materials including video and audio clips. It is a video model, not a text or coding model. MiniMax also released its M3 language model (for text, agents, and reasoning) at the same WAIC 2026 event, and the two are frequently confused online. H3 is squarely focused on video creation. Those numbers matter, but they only tell part of the story. The real question is what MiniMax H3 actually does with them. One of the most persistent pain points in AI video has been character consistency. Generate a woman walking through a café in one shot, and she may look like a completely different person in the next. Faces drift. Clothing changes. Voices shift. H3 addresses this with what MiniMax calls Omni-Reference — a control system that lets you feed up to 9 reference images, 3 video clips, and 3 audio clips into a single generation request. The model reads all of these inputs as one unified context, extracting the character's face from a photo, borrowing camera motion from a video clip, and absorbing vocal tone from an audio sample.