Tools/Text-to-Speech (TTS)/ViiTorVoice-NAR

ViiTorVoice-NAR

Non-autoregressive TTS with word-level audio editing, zero-shot cloning, and 60 ms first-frame latency.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Editing one word inside finished narration normally means regenerating the whole take; the non-autoregressive design of ViiTorVoice-NAR makes the edit surgical instead, resynthesizing only the changed span from source audio, original text, and edited text while the surrounding speech is left untouched. The same model performs zero-shot voice cloning from a short prompt recording or a precomputed codebook, plus paralinguistic control through inline tags such as emotion markers, steered by classifier-free guidance scales. Because generation is parallel rather than token by token, first-frame latency lands around 60 ms end to end, which suits live agents and streaming pipelines. Deployment is a bash-scripted Python environment exposing gRPC services with an HTTP gateway, with weights fetched from Hugging Face, and demonstrations cover Chinese and English. Released in June 2026 under Apache 2.0 with no usage restrictions, it occupies a niche little open TTS touches: post-production word-level audio editing combined with low-latency synthesis.

Should you use ViiTorVoice-NAR?

Pick it when

Pick ViiTorVoice-NAR when you edit finished narration or dubbing and want to fix single words without regenerating the take, or need about 60 ms first-frame latency for a live agent, in Chinese or English, under Apache-2.0.

Look elsewhere when

Skip it if you want a pip install or a simple Python call: deployment runs gRPC services behind an HTTP gateway, the June 2026 release is new, and demos cover only Chinese and English. CosyVoice 2 is the more established streamer.

Alternatives to ViiTorVoice-NAR

  • VoiceCraft

    Also does word-level speech editing and zero-shot TTS, but its code is CC BY-NC-SA and weights non-commercial, where ViiTorVoice is Apache-2.0.

  • Step-Audio-EditX

    Edits emotion, style, and paralinguistic events in existing recordings rather than words, with Mandarin, English, and dialects, on 12 GB.

  • CosyVoice 2

    Mature bidirectional streaming near 150 ms first packet with a web UI and Docker, but no word-level editing of existing audio.

  • FireRedTTS 2

    Streaming long-form dialogue with 4 speakers and 140 ms first packet, better for podcasts, though without targeted word edits.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
Aug 24, 2026

Tags