Supertonic

On-device ONNX text-to-speech around 99M parameters covering 31 languages with no cloud calls.

Open SourceSelf HostedOffline Capable
0.0 (0)

About

Supertonic goes where big TTS models cannot: about 99M parameters running through ONNX Runtime entirely on-device, synthesizing 44.1kHz audio faster than real time on plain CPUs, phones, browsers via WebGPU, and even Raspberry Pi boards and e-readers. Version 3, released May 2026, expanded coverage from 5 to 31 languages and ships ten preset voice styles plus inline expression tags for things like laughter and breaths. Integration is the selling point: ready-made SDKs cover Python, Node.js, browser JavaScript, Java, C++, C#, Go, Rust, Swift, and Flutter, a pip package installs the engine in one line, and a local HTTP server exposes OpenAI-compatible endpoints. Sample code is MIT while the model weights carry an OpenRAIL-M license, with everything downloadable from Hugging Face. One caveat: developer Supertone ended official support in July 2026 and its hosted Voice Builder is being retired, so the repo is frozen in a working state rather than evolving. For private, offline speech on weak hardware it remains hard to beat.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Easy (2/5)
License
MIT / OpenRAIL-M
Added
Aug 24, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
Browse all Text-to-Speech (TTS) tools