Tag
Release of C++/GGML based implementations of Supertonic 3, MOSS-TTS, IndexTTS2, and Irodori-TTS in audio.cpp, capable of generating 10 hours of audio in 3 minutes on an RTX 5090.
Compares two locally running mobile TTS models, Kokoro and Supertonic, questioning their production quality beyond initial demos.
Supertonic 3 is a 99M parameter open-source TTS model that runs entirely on-device, beating ElevenLabs on a Raspberry Pi with 167x faster than real-time performance on a laptop CPU.
Supertonic is a lightning-fast, on-device TTS model with 99M parameters, supporting 31 languages. It runs locally with no API costs, outperforms cloud TTS on accuracy for numbers, phone numbers, and technical terms, and can be installed via Python, Node.js, Rust, Go, and more.
The article highlights Supertonic, an open-source text-to-speech model that runs entirely on-device, claiming superior speed and formatting accuracy compared to cloud-based services like ElevenLabs and OpenAI.