Zonos
Open-weight multilingual TTS from Zyphra with voice cloning and a mixture-of-experts second generation.
About
Zyphra open-weights the Zonos family of text to speech models. Zonos-v0.1 shipped in February 2025 in two variants, a transformer and an SSM hybrid, trained on more than 200,000 hours of multilingual speech and covering English, Japanese, Chinese, French, and German with zero-shot voice cloning from 10 to 30 second samples, audio prefix conditioning, and fine control over speaking rate, pitch, emotion, and audio quality at 44.1 kHz. It predicts Descript Audio Codec tokens after eSpeak phonemization, needs a GPU with 6 GB or more of VRAM, installs via uv, pip, or Docker with espeak-ng as a system dependency, and ships a Gradio interface; the hybrid additionally wants a 3000-series or newer NVIDIA card. ZONOS2, announced in June 2026 and living in its own repository, moved to a sparse mixture-of-experts design with 900 million active of 8 billion total parameters, trained on over six million hours, and is described as the first open-source MoE text to speech model with roughly 4x higher throughput. Zyphra also serves the models through its hosted playground and API. The v0.1 repo is Apache-2.0 and holds about 7.2k GitHub stars.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 6 GB
- Added
- Jul 29, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.