MiniMax H3
Open-weights omni-modal video generator with native stereo audio, driven by text, image, video, or audio input.
About
Hailuo 3.0 reached open weights as MiniMax H3, an omni-modal video generator whose 33B dense H3-Base transformer conditions on any mix of text, images, video clips, and audio, and emits video with native 32 kHz stereo sound rather than adding audio in a second pass. Clips run 4 to 15 seconds at 24 fps across six aspect ratios with speech in 11 languages, and generation modes cover text-to-video, first and last frame control, and reference-to-video from identity images. The open model outputs 768p; MiniMax's H3-Regenerate-2K stage lifts that to 2K through in-context regeneration but remains API-only. Day-one support landed in SGLang, vLLM, diffusers, and ComfyUI, with BF16 checkpoints for the FL2VA and Ref2VA variants on Hugging Face; the reference SGLang deployment spans four GPUs, putting full local use in multi-GPU territory. Weights ship under the MiniMax H3 Community License, and the same model powers the paid Hailuo apps and platform API, making this one of the strongest open video stacks of 2026.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Advanced (4/5)
- License
- MiniMax H3 Community License
- Added
- Aug 24, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.