Higgs Audio
Llama-based audio foundation model generating expressive speech, dialogue, and cloned voices.
About
Higgs Audio v2 extends Llama-3.2-3B into a text-audio foundation model, adding a 2.2B parameter DualFFN audio adapter for 5.8B parameters total and a unified tokenizer running at 25 frames per second, pretrained on 10 million hours of audio from Boson AI's AudioVerse pipeline. The result generates expressive speech with automatic prosody adaptation, zero-shot voice cloning, multi-speaker dialogue, melodic humming in a cloned voice, and speech over background music. On EmergentTTS-Eval it wins 75.71 percent of comparisons against gpt-4o-mini-tts in the Emotions category and 55.71 percent on Questions, with a 2.44 word error rate on Seed-TTS Eval. The repository installs through pip, venv, conda, or uv, includes a vLLM-backed OpenAI-compatible server for higher throughput, and recommends a GPU with at least 24 GB of memory. Code is Apache-2.0 while the weights carry the Boson Higgs Audio 2 Community License, a Llama-style grant that requires an expanded license beyond 100,000 annual active users. Boson has since shipped a standalone v3 model served through SGLang-Omni or its hosted API, so the v2 guide now lives in README_V2.md.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Intermediate (3/5)
- License
- Boson Higgs Audio 2 Community License
- Minimum VRAM
- 24 GB
- Added
- Jul 29, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.