Dia2
Nari Labs' streaming dialogue TTS models that begin speaking before the full text arrives.
About
Where most TTS systems wait for a full sentence, Dia2 begins speaking after the first few words arrive. Nari Labs' successor to its viral Dia model is a streaming dialogue engine released as open 1B and 2B checkpoints on Hugging Face, generating conversational audio progressively so a voice agent can respond with near-zero perceived latency. The model conditions on prior audio context, which keeps voices consistent across turns and lets developers steer delivery with an audio prefix, and it sustains up to about two minutes of English speech per generation. Built on Kyutai's Mimi codec, it runs in bfloat16 on CUDA 12.8 or newer GPUs with a CPU fallback, and installation goes through the uv package manager from the GitHub repository. Code and weights are both Apache 2.0, so commercial voice products can embed it freely. Nari Labs is candid that this is research-grade software: output quality varies between generations unless you fine-tune or anchor it with audio conditioning, but for real-time dialogue it occupies ground few open models touch.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.