ViiTorVoice-NAR
Non-autoregressive TTS with word-level audio editing, zero-shot cloning, and 60 ms first-frame latency.
About
Editing one word inside finished narration normally means regenerating the whole take; the non-autoregressive design of ViiTorVoice-NAR makes the edit surgical instead, resynthesizing only the changed span from source audio, original text, and edited text while the surrounding speech is left untouched. The same model performs zero-shot voice cloning from a short prompt recording or a precomputed codebook, plus paralinguistic control through inline tags such as emotion markers, steered by classifier-free guidance scales. Because generation is parallel rather than token by token, first-frame latency lands around 60 ms end to end, which suits live agents and streaming pipelines. Deployment is a bash-scripted Python environment exposing gRPC services with an HTTP gateway, with weights fetched from Hugging Face, and demonstrations cover Chinese and English. Released in June 2026 under Apache 2.0 with no usage restrictions, it occupies a niche little open TTS touches: post-production word-level audio editing combined with low-latency synthesis.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.