FireRedTTS 2

Streaming TTS for long-form multi-speaker dialogue with zero-shot voice cloning in seven languages.

Open SourceSelf HostedOffline CapableGPU Required (9GB+ VRAM)
0.0 (0)

About

FireRedTTS 2 tackles long-form conversational synthesis, generating podcast-style dialogues of up to 3 minutes with 4 distinct speakers while keeping speaker switches reliable and prosody consistent with conversational context. The system runs a dual-transformer over interleaved text and speech sequences on top of a 12.5 Hz streaming speech tokenizer, generating sentence by sentence, and reaches a 140 ms first-packet latency on an NVIDIA L20 in the team's tests, which suits live voice agents as well as offline production. It covers English, Chinese, Japanese, Korean, French, German, and Russian, and supports zero-shot voice cloning, including cross-lingual and code-switching scenarios, plus random timbre generation for building synthetic training data. Deployment means a Python 3.11 environment or Docker image with weights pulled from Hugging Face via Git LFS; bf16 inference brings VRAM needs from 14 GB down to about 9 GB. Code and weights from the Xiaohongshu FireRed team are Apache-2.0, though the README restricts the voice cloning feature to academic research, and a paper and demo page accompany the release.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
9 GB
Added
Jul 29, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools