MiniMax-Music3

Open-weights model that generates complete five-minute stereo songs from lyrics and a style caption.

Open SourceSelf HostedOffline CapableGPU Required (8GB+ VRAM)
0.0 (0)

About

Full songs, not loops, are the target of MiniMax-Music3: given lyrics tagged with markers like [Verse] and [Chorus] plus a caption describing genre, instrumentation, and vocal style, it renders up to five minutes of 32 kHz, 16-bit stereo audio with coherent arrangements. The architecture is hierarchical: an 8B global language model plans semantic tokens and long-range song structure, a 0.6B local model fills in frame-level acoustic detail, and a 2.4B flow-matching stage with a Flow-VAE decoder synthesizes the waveform. MiniMax published the weights on Hugging Face in August 2026 under the MiniMax-Music3 Community License, which permits commercial use with attribution until yearly revenue passes 20 million dollars. Inference runs as a modular diffusers pipeline on a CUDA GPU, needing about 24 GB of VRAM at full precision, around 22 GB with CPU offloading, and 8 GB cards via a slower streaming mode; SGLang-Omni serves it at scale, and ComfyUI added official support at launch. It arrived as the strongest open-weights alternative to hosted services like Suno.

Should you use MiniMax-Music3?

Pick it when

Pick MiniMax-Music3 when you want a self-hosted alternative to hosted song services, five-minute stereo tracks from [Verse] and [Chorus] tagged lyrics, run through diffusers or SGLang-Omni on a CUDA GPU with about 24 GB.

Look elsewhere when

Skip it on small GPUs, since 8 GB cards only work through a slower streaming mode, or if revenue may pass 20 million dollars a year, where the community license stops covering you. DiffRhythm and HeartMuLa are Apache-2.0.

Alternatives to MiniMax-Music3

  • MiniMax Music 3

    Duplicate listing of the same weights that emphasizes ComfyUI workflows and the two-GPU reference split, more useful if you work in ComfyUI.

  • DiffRhythm

    Apache-2.0 with no revenue cap and 8 GB with chunked decoding, but no section-tagged lyric format and a shorter maximum near 4 minutes 45 seconds.

  • YuE

    Apache-2.0 multilingual 7B lyric-to-song with no revenue cap, needing 16 GB rather than 24 GB at full precision here, though setup is harder.

  • HeartMuLa

    Apache-2.0 3B model near real time on one GPU with lazy loading for smaller cards, trading model scale for clean commercial terms.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
MiniMax-Music3 Community License
Minimum VRAM
8 GB
Added
Aug 24, 2026

Tags