HeartMuLa
Open music foundation model family covering song generation, a 12.5 Hz codec, transcription, and alignment.
About
Four models make up the HeartMuLa family, an open music foundation stack: the HeartMuLa language model generates full songs conditioned on multilingual lyrics and style tags, HeartCodec compresses audio into 12.5 Hz tokens so long tracks stay tractable for the language model, HeartTranscriptor performs Whisper-based lyrics transcription, and HeartCLAP aligns music and text for retrieval and scoring. The released 3B generation model runs at roughly real-time speed on a single GPU, with lazy loading to fit smaller cards and multi-GPU placement for headroom, and a 7B variant is planned. Install is git clone plus pip on Python 3.10, with checkpoints pulled from Hugging Face or ModelScope; a January 2026 license update moved the whole project and its model weights to Apache 2.0, clearing commercial use. As one of the few serious open rivals to ACE-Step for lyric-conditioned song generation, with the codec and transcription pieces usable standalone, it has become a common base for open music research and tooling, sitting near 3.7k GitHub stars.
Should you use HeartMuLa?
Pick it when
Pick HeartMuLa when you want an Apache-2.0 lyric-to-song model with multilingual lyrics and style tags that you can build products on, plus reusable codec, lyrics transcription, and music-text alignment models.
Look elsewhere when
Skip it if you need songs in seconds, since the 3B model runs at roughly real time and ACE-Step is faster. The 7B variant is still only planned and no VRAM floor is listed, so test fit on small cards despite lazy loading.
Alternatives to HeartMuLa
- ACE-Step
Full songs with lyrics in seconds on an 8 GB GPU, much faster turnaround, but no standalone codec, transcription, or retrieval models.
- MiniMax-Music3
Larger 8B planner aimed at five-minute section-tagged songs, but a revenue-capped community license and about 24 GB at full precision.
- DiffRhythm
Same Apache-2.0 terms with diffusion instead of a language model and up to 4 minutes 45 seconds, but no codec, transcription, or retrieval models.
- YuE
Multilingual 7B models available now with in-context learning modes, also Apache-2.0, but they need 16 GB or more of VRAM.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Music & Audio Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026