Tag
A user shares hands-on impressions of a thin, light Android-based device featuring microphone, speaker, OCR, text-to-speech, and a ChatGPT 5.2 chat, noting the only clear drawback is its lack of color.
NVIDIA's entire speech stack—ASR, TTS, and codec—is now quantized to GGUF and runs locally on-device via NeMo-Speech.cpp, with new model releases for Magpie-TTS, Nemotron Speech Streaming, and Parakeet.
Scenema Audio, an expressive text-to-speech model with zero-shot voice cloning, is now available as a native ComfyUI custom node, quantized to run on 8GB VRAM. The release adds inline stage direction cues, 12 preset voices, and simplifies the prompt format for ComfyUI.
Speechfony is a new fully local desktop app that reads PDFs and EPUBs aloud with sentence highlighting, semantic search, and MP3 audiobook export, using Kokoro TTS and on-device embeddings. It prioritizes privacy and offline use, with an open-source MIT license.
Qwen3-TTS voice cloning has been merged into mainline llama.cpp, enabling local text-to-speech with voice cloning from short reference audio via the llama-tts binary, supporting multiple languages. Limitations remain, including only the Base model and no server endpoint yet.
Presents DLLM-TTS, a block discrete diffusion language model for text-to-speech synthesis that processes X-Codec2 tokens in blocks, enabling parallel generation with RTF 0.15 while achieving competitive quality with only 20K hours of training data.
OpenVoice is an open-source instant voice cloning model with style and language control, now available on GitHub.
OpenMMLab has open-sourced Amphion, an audio generation toolbox supporting TTS, singing voice conversion, sound effect generation, and more. It is completely free for commercial use and supports local deployment.
Simba 3.2 from SpeechifyAI claims the #1 spot on the blind-test voice leaderboard, surpassing ElevenLabs, OpenAI, and Google DeepMind, at a significantly lower cost with a new API and free tier.
Gradium has released a new TTS model in public beta that accurately reads phone numbers, emails, IBANs, and time expressions natively, offering API access with 1M credits for testing.
Fish Audio launches S2.1 Pro, a production voice model with 90ms latency, support for 83 languages, voice cloning from short samples, and multi-speaker dialogue, available via API with a free tier for development.
Audio8 releases a 0.6B-parameter multilingual text-to-speech model with zero-shot voice cloning capabilities, available on Hugging Face under Apache 2.0 license.
Liso lets you highlight text on any webpage and converts it into audio for listening.
NeuTTS-2E is an open-source on-device TTS model that supports seven controllable emotions.
Dia is a 1.6B-parameter open-source text-to-speech model that generates English dialogue from transcripts, supporting two-speaker generation, audio conditioning, and nonverbal cues.
A curated list of text-to-speech resources for voice agent builders, organized around the decision between real-time streaming synthesis and offline high-fidelity synthesis, with emphasis on streaming latency and mid-stream cancellation.
A curated reference of open-source text-to-speech models organized by license type, highlighting which models can be used commercially.
Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a new text-to-speech model with Flash (real-time) and Plus (high-quality) versions, supporting 16 languages, natural language style control, and robust voice cloning.
A developer created a MacOS utility for local text-to-speech and dictation using advanced AI models, emphasizing speed and privacy.
Scylla's Band is a new TTS model and inference framework, including an Android sample for deployment.