Coqui TTS
Deep learning toolkit for text-to-speech synthesis
About
Coqui TTS is a deep learning library for text-to-speech that ships pretrained models, training and fine-tuning tools, and implementations of several spectrogram and vocoder architectures including Tacotron, Glow-TTS, VITS, and FastSpeech. It supports multi-speaker and multi-lingual training and a 16+ language pretrained set covering English, Spanish, French, German, and others. Released under MPL-2.0.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Audio & Speech
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- MPL-2.0
- Added
- Jan 29, 2026
Related Tools
Free text-to-speech generator with multiple voices, accents, and languages. No signup required.
Universal neural vocoder from NVIDIA that converts mel spectrograms into waveforms up to 44 kHz.
End-to-end Chinese and English spoken dialogue model from Zhipu AI with streaming speech output.
Audio foundation model unifying speech recognition, understanding, and conversation in one 7B model.
Python library for audio and music analysis, from spectrograms to beat tracking and features.
CTranslate2-based Whisper with 4x faster transcription