ViiTorVoice-NAR
Non-autoregressive TTS with word-level audio editing, zero-shot cloning, and 60 ms first-frame latency.
About
Editing one word inside finished narration normally means regenerating the whole take; the non-autoregressive design of ViiTorVoice-NAR makes the edit surgical instead, resynthesizing only the changed span from source audio, original text, and edited text while the surrounding speech is left untouched. The same model performs zero-shot voice cloning from a short prompt recording or a precomputed codebook, plus paralinguistic control through inline tags such as emotion markers, steered by classifier-free guidance scales. Because generation is parallel rather than token by token, first-frame latency lands around 60 ms end to end, which suits live agents and streaming pipelines. Deployment is a bash-scripted Python environment exposing gRPC services with an HTTP gateway, with weights fetched from Hugging Face, and demonstrations cover Chinese and English. Released in June 2026 under Apache 2.0 with no usage restrictions, it occupies a niche little open TTS touches: post-production word-level audio editing combined with low-latency synthesis.
Should you use ViiTorVoice-NAR?
Pick it when
Pick ViiTorVoice-NAR when you edit finished narration or dubbing and want to fix single words without regenerating the take, or need about 60 ms first-frame latency for a live agent, in Chinese or English, under Apache-2.0.
Look elsewhere when
Skip it if you want a pip install or a simple Python call: deployment runs gRPC services behind an HTTP gateway, the June 2026 release is new, and demos cover only Chinese and English. CosyVoice 2 is the more established streamer.
Alternatives to ViiTorVoice-NAR
- VoiceCraft
Also does word-level speech editing and zero-shot TTS, but its code is CC BY-NC-SA and weights non-commercial, where ViiTorVoice is Apache-2.0.
- Step-Audio-EditX
Edits emotion, style, and paralinguistic events in existing recordings rather than words, with Mandarin, English, and dialects, on 12 GB.
- CosyVoice 2
Mature bidirectional streaming near 150 ms first packet with a web UI and Docker, but no word-level editing of existing audio.
- FireRedTTS 2
Streaming long-form dialogue with 4 speakers and 140 ms first packet, better for podcasts, though without targeted word edits.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026