MiniMax Music 3
Open-weights model that turns tagged lyrics and captions into full stereo songs up to five minutes long.
About
Complete songs rather than loops are the output of MiniMax Music 3: given lyrics annotated with section tags and a structured caption describing genre, vocal character, and arrangement, it renders up to five minutes of 32 kHz 16-bit stereo audio with coherent verse and chorus structure. The open-weights release exposes the full stack: an 8B global language model plans long-range musical structure, a 0.6B local model fills in acoustic detail over an 8-layer residual vector quantization tokenizer, and a 2.4B Flow Matching network with a Flow-VAE decoder synthesizes the waveform. The reference pipeline splits those stages across two CUDA GPUs, but CPU offloading brings it down to a single 8 GB card with 24 GB or more recommended, and ComfyUI shipped official workflow support at launch. Weights live on Hugging Face under the MiniMax-Music3 Community License, which permits commercial use below 20 million dollars of annual revenue with attribution in the product UI. Released in August 2026, it is the strongest open lyric-to-song system yet from a major lab.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Music & Audio Generation
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Intermediate (3/5)
- License
- MiniMax-Music3 Community License
- Minimum VRAM
- 8 GB
- Added
- Aug 24, 2026
Related Tools
Latent diffusion model for text-to-audio, music, and speech generation.
Audio super-resolution model for upsampling audio to higher sample rates.
State-of-the-art music source separation model by Meta for splitting tracks.
Fast music generation model producing full songs with lyrics in seconds.
Audio diffusion model by Harmonai for generating music samples.
PyTorch library for deep learning research on audio generation including MusicGen and AudioGen.