Supertonic
On-device ONNX text-to-speech around 99M parameters covering 31 languages with no cloud calls.
About
Supertonic goes where big TTS models cannot: about 99M parameters running through ONNX Runtime entirely on-device, synthesizing 44.1kHz audio faster than real time on plain CPUs, phones, browsers via WebGPU, and even Raspberry Pi boards and e-readers. Version 3, released May 2026, expanded coverage from 5 to 31 languages and ships ten preset voice styles plus inline expression tags for things like laughter and breaths. Integration is the selling point: ready-made SDKs cover Python, Node.js, browser JavaScript, Java, C++, C#, Go, Rust, Swift, and Flutter, a pip package installs the engine in one line, and a local HTTP server exposes OpenAI-compatible endpoints. Sample code is MIT while the model weights carry an OpenRAIL-M license, with everything downloadable from Hugging Face. One caveat: developer Supertone ended official support in July 2026 and its hosted Voice Builder is being retired, so the repo is frozen in a working state rather than evolving. For private, offline speech on weak hardware it remains hard to beat.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- MIT / OpenRAIL-M
- Added
- Aug 24, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.