HeartMuLa

Open music foundation model family covering song generation, a 12.5 Hz codec, transcription, and alignment.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Four models make up the HeartMuLa family, an open music foundation stack: the HeartMuLa language model generates full songs conditioned on multilingual lyrics and style tags, HeartCodec compresses audio into 12.5 Hz tokens so long tracks stay tractable for the language model, HeartTranscriptor performs Whisper-based lyrics transcription, and HeartCLAP aligns music and text for retrieval and scoring. The released 3B generation model runs at roughly real-time speed on a single GPU, with lazy loading to fit smaller cards and multi-GPU placement for headroom, and a 7B variant is planned. Install is git clone plus pip on Python 3.10, with checkpoints pulled from Hugging Face or ModelScope; a January 2026 license update moved the whole project and its model weights to Apache 2.0, clearing commercial use. As one of the few serious open rivals to ACE-Step for lyric-conditioned song generation, with the codec and transcription pieces usable standalone, it has become a common base for open music research and tooling, sitting near 3.7k GitHub stars.

Should you use HeartMuLa?

Pick it when

Pick HeartMuLa when you want an Apache-2.0 lyric-to-song model with multilingual lyrics and style tags that you can build products on, plus reusable codec, lyrics transcription, and music-text alignment models.

Look elsewhere when

Skip it if you need songs in seconds, since the 3B model runs at roughly real time and ACE-Step is faster. The 7B variant is still only planned and no VRAM floor is listed, so test fit on small cards despite lazy loading.

Alternatives to HeartMuLa

  • ACE-Step

    Full songs with lyrics in seconds on an 8 GB GPU, much faster turnaround, but no standalone codec, transcription, or retrieval models.

  • MiniMax-Music3

    Larger 8B planner aimed at five-minute section-tagged songs, but a revenue-capped community license and about 24 GB at full precision.

  • DiffRhythm

    Same Apache-2.0 terms with diffusion instead of a language model and up to 4 minutes 45 seconds, but no codec, transcription, or retrieval models.

  • YuE

    Multilingual 7B models available now with in-context learning modes, also Apache-2.0, but they need 16 GB or more of VRAM.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
Aug 24, 2026

Tags