Dia2
Nari Labs' streaming dialogue TTS models that begin speaking before the full text arrives.
About
Where most TTS systems wait for a full sentence, Dia2 begins speaking after the first few words arrive. Nari Labs' successor to its viral Dia model is a streaming dialogue engine released as open 1B and 2B checkpoints on Hugging Face, generating conversational audio progressively so a voice agent can respond with near-zero perceived latency. The model conditions on prior audio context, which keeps voices consistent across turns and lets developers steer delivery with an audio prefix, and it sustains up to about two minutes of English speech per generation. Built on Kyutai's Mimi codec, it runs in bfloat16 on CUDA 12.8 or newer GPUs with a CPU fallback, and installation goes through the uv package manager from the GitHub repository. Code and weights are both Apache 2.0, so commercial voice products can embed it freely. Nari Labs is candid that this is research-grade software: output quality varies between generations unless you fine-tune or anchor it with audio conditioning, but for real-time dialogue it occupies ground few open models touch.
Should you use Dia2?
Pick it when
Pick Dia2 when you are building an English voice agent that should start speaking while the LLM reply is still arriving, need voices to stay consistent across turns, and want Apache-2.0 code and weights you can ship commercially.
Look elsewhere when
Skip it if you need consistent, polished output without fine-tuning or audio anchoring, or non-English speech; GPU inference also needs CUDA 12.8 or newer. Kyutai TTS is the more production-ready streaming option.
Alternatives to Dia2
- Kyutai TTS
Streams text in so playback starts mid-reply, with a production Rust server and MLX builds, but its CC-BY-4.0 weights require attribution.
- Dia TTS
Original 1.6B Dia renders full multi-speaker scripts offline with nonverbal sounds, simpler for recorded content but without Dia2's streaming.
- Orpheus TTS
About 200 ms streaming, lower with input streaming, via vLLM with emotion tags and eight voices, aimed at single-voice agents rather than dialogue.
- RealtimeTTS
Glue library that streams LLM tokens into 20-plus engines with fallback, useful if you want to keep the model swappable instead of building around one.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026