Tools/Text-to-Speech (TTS)/ViiTorVoice-NAR

ViiTorVoice-NAR

Non-autoregressive TTS with word-level audio editing, zero-shot cloning, and 60 ms first-frame latency.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Editing one word inside finished narration normally means regenerating the whole take; the non-autoregressive design of ViiTorVoice-NAR makes the edit surgical instead, resynthesizing only the changed span from source audio, original text, and edited text while the surrounding speech is left untouched. The same model performs zero-shot voice cloning from a short prompt recording or a precomputed codebook, plus paralinguistic control through inline tags such as emotion markers, steered by classifier-free guidance scales. Because generation is parallel rather than token by token, first-frame latency lands around 60 ms end to end, which suits live agents and streaming pipelines. Deployment is a bash-scripted Python environment exposing gRPC services with an HTTP gateway, with weights fetched from Hugging Face, and demonstrations cover Chinese and English. Released in June 2026 under Apache 2.0 with no usage restrictions, it occupies a niche little open TTS touches: post-production word-level audio editing combined with low-latency synthesis.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
Aug 24, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
Browse all Text-to-Speech (TTS) tools