MiniMax-Music3
Open-weights model that generates complete five-minute stereo songs from lyrics and a style caption.
About
Full songs, not loops, are the target of MiniMax-Music3: given lyrics tagged with markers like [Verse] and [Chorus] plus a caption describing genre, instrumentation, and vocal style, it renders up to five minutes of 32 kHz, 16-bit stereo audio with coherent arrangements. The architecture is hierarchical: an 8B global language model plans semantic tokens and long-range song structure, a 0.6B local model fills in frame-level acoustic detail, and a 2.4B flow-matching stage with a Flow-VAE decoder synthesizes the waveform. MiniMax published the weights on Hugging Face in August 2026 under the MiniMax-Music3 Community License, which permits commercial use with attribution until yearly revenue passes 20 million dollars. Inference runs as a modular diffusers pipeline on a CUDA GPU, needing about 24 GB of VRAM at full precision, around 22 GB with CPU offloading, and 8 GB cards via a slower streaming mode; SGLang-Omni serves it at scale, and ComfyUI added official support at launch. It arrived as the strongest open-weights alternative to hosted services like Suno.
Should you use MiniMax-Music3?
Pick it when
Pick MiniMax-Music3 when you want a self-hosted alternative to hosted song services, five-minute stereo tracks from [Verse] and [Chorus] tagged lyrics, run through diffusers or SGLang-Omni on a CUDA GPU with about 24 GB.
Look elsewhere when
Skip it on small GPUs, since 8 GB cards only work through a slower streaming mode, or if revenue may pass 20 million dollars a year, where the community license stops covering you. DiffRhythm and HeartMuLa are Apache-2.0.
Alternatives to MiniMax-Music3
- MiniMax Music 3
Duplicate listing of the same weights that emphasizes ComfyUI workflows and the two-GPU reference split, more useful if you work in ComfyUI.
- DiffRhythm
Apache-2.0 with no revenue cap and 8 GB with chunked decoding, but no section-tagged lyric format and a shorter maximum near 4 minutes 45 seconds.
- YuE
Apache-2.0 multilingual 7B lyric-to-song with no revenue cap, needing 16 GB rather than 24 GB at full precision here, though setup is harder.
- HeartMuLa
Apache-2.0 3B model near real time on one GPU with lazy loading for smaller cards, trading model scale for clean commercial terms.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Music & Audio Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- MiniMax-Music3 Community License
- Minimum VRAM
- 8 GB
- Added
- Aug 24, 2026