Dia2

Nari Labs' streaming dialogue TTS models that begin speaking before the full text arrives.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Where most TTS systems wait for a full sentence, Dia2 begins speaking after the first few words arrive. Nari Labs' successor to its viral Dia model is a streaming dialogue engine released as open 1B and 2B checkpoints on Hugging Face, generating conversational audio progressively so a voice agent can respond with near-zero perceived latency. The model conditions on prior audio context, which keeps voices consistent across turns and lets developers steer delivery with an audio prefix, and it sustains up to about two minutes of English speech per generation. Built on Kyutai's Mimi codec, it runs in bfloat16 on CUDA 12.8 or newer GPUs with a CPU fallback, and installation goes through the uv package manager from the GitHub repository. Code and weights are both Apache 2.0, so commercial voice products can embed it freely. Nari Labs is candid that this is research-grade software: output quality varies between generations unless you fine-tune or anchor it with audio conditioning, but for real-time dialogue it occupies ground few open models touch.

Should you use Dia2?

Pick it when

Pick Dia2 when you are building an English voice agent that should start speaking while the LLM reply is still arriving, need voices to stay consistent across turns, and want Apache-2.0 code and weights you can ship commercially.

Look elsewhere when

Skip it if you need consistent, polished output without fine-tuning or audio anchoring, or non-English speech; GPU inference also needs CUDA 12.8 or newer. Kyutai TTS is the more production-ready streaming option.

Alternatives to Dia2

  • Kyutai TTS

    Streams text in so playback starts mid-reply, with a production Rust server and MLX builds, but its CC-BY-4.0 weights require attribution.

  • Dia TTS

    Original 1.6B Dia renders full multi-speaker scripts offline with nonverbal sounds, simpler for recorded content but without Dia2's streaming.

  • Orpheus TTS

    About 200 ms streaming, lower with input streaming, via vLLM with emotion tags and eight voices, aimed at single-voice agents rather than dialogue.

  • RealtimeTTS

    Glue library that streams LLM tokens into 20-plus engines with fallback, useful if you want to keep the model swappable instead of building around one.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
Aug 24, 2026

Tags