Tag
TTS Arena launches as a blind benchmark for text-to-speech models, where users compare anonymous TTS outputs and vote for the more human-sounding one, updating a live leaderboard.
Proposes HybridCodec, a novel framework combining temporally compressed discrete tokens with continuous residuals to improve speaker characteristic retention in speech language models, reducing autoregressive steps while maintaining quality.
This paper investigates LoRA fine-tuning of the VoxCPM2 TTS model to improve quality for low-resource languages like Khmer, while showing no gain for Korean which the base model already handles well. The adapter yields significant MOS improvement for Khmer with minimal parameter training.
audio.cpp is a C++/ggml runtime that integrates 12 audio models including Qwen3-TTS, PocketTTS, and VeVo2, achieving TTS up to 5x faster than Python on CUDA.
Inflect-Micro-v2 is a compact text-to-speech model with under 10M parameters, supporting CPU/CUDA inference and long-text handling, released on Hugging Face.
Release of Inflect-Nano-v2, a fixed-voice English TTS model with under 4M parameters for local text-to-waveform synthesis, supporting CPU or CUDA inference and long-text handling.
This paper presents a modular end-to-end speech-to-speech conversational system for the low-resource Algerian Dialect, integrating ASR, NLU, RAG, and TTS with dedicated datasets and fine-tuned models.
A personal experiment building an AI commentator for World Cup matches reveals realistic results until fast-paced gameplay causes issues.
A guide to building a fully local voice assistant using Platypush on a Raspberry Pi, covering hotword detection, speech-to-text, text-to-speech, and home automation integration.
MOSS-TTS is an open-source voice cloning model introduced by Mosi Company. Users can clone a voice by reading a small amount of text, and then use the cloned voice to generate any speech with realistic results.
The article discusses the underutilized potential of voice as an output layer for AI agents, highlighting practical use cases and workflow challenges beyond simple text-to-speech.
NetEase Youdao open-sourced the 1.3B parameter Confucius4-TTS model, supporting zero-shot voice cloning and cross-lingual speech synthesis in 14 languages, fast and with excellent results.
MOSS-TTS-Local Transformer v1.5 is an open-source 48 kHz stereo TTS model with zero-shot voice cloning, native streaming, and support for 31 languages, built on a Qwen3-4B backbone and served via SGLang-Omni.
MosiAI has released MOSS-TTS Local Transformer v1.5, a text-to-speech model that supports voice cloning, over 30 languages, and high-quality 48 kHz output.
Inflect-Nano, an ultra-extreme tiny 4.63 million parameter text-to-speech model, has been released.
Google's Gemini TTS now supports streaming audio generation, allowing developers to build voice applications that start speaking instantly without waiting for full audio output.
VoxCPM2 is an open-source speech synthesis model from OpenBMB, using a tokenizer-free diffusion autoregressive architecture, supporting 30 languages, voice design, and controllable voice cloning. It can clone a voice with just one sentence, or create a brand new voice using text, outputting 48kHz high-quality audio, and is commercially usable.
Inflect-Nano-v1 is a tiny English text-to-speech model with 4.63M total inference parameters, including its vocoder, designed for local, efficient speech synthesis experiments.
Kokoro-82M is a highly natural text-to-speech model with 82 million parameters and over 11 million downloads, representing a significant advancement in AI voice generation.
Cartesia released Sonic-3.5 (text-to-speech) and Ink-2 (speech-to-text), claiming they are the #1 streaming models for voice agents, with potential to disrupt call centers.