Tag
Pipecat is an open-source Python framework for building real-time voice AI agents, handling speech recognition, text-to-speech, conversation logic, and supporting multiple AI service providers.
A detailed, unsponsored review of ElevenLabs rating it 8.1/10, highlighting its emotional range and low latency as strengths, but cautioning about high costs and the 'regeneration tax' for casual users and high-volume publishers.
Real World VoiceEQ is a new benchmark for evaluating the human quality of voice AI, based on over a million human ratings, assessing models across speech recognition, synthesis, and understanding in real-world conditions.
VocalVia is a tool that converts documents and articles into editable multi-voice audio.
FreyaTTS is a compact, tokenizer-free Turkish-first text-to-speech model based on a non-autoregressive conditional flow-matching Diffusion Transformer, achieving state-of-the-art performance with a fraction of the parameters of larger systems and released under Apache-2.0.
This paper identifies a confound in best-of-N TTS evaluation where the apparent quality of ASR verifiers depends strongly on which ASR family is used as evaluator. The authors propose cross-family rank ensembles that achieve lower word error rates across multiple evaluators.
A report on testing an AI pipeline that automatically generates a full podcast episode from a given topic without any manual editing.
This article introduces Kokoro, a lightweight 82M-parameter text-to-speech model that runs locally on CPU, providing high-quality speech synthesis across multiple languages while preserving privacy. It explains how to set up Kokoro via a Docker container with an OpenAI-compatible API for easy integration.
Gepard is a new streaming TTS model capable of real-time dialogue with ~50ms time-to-first-audio, supporting voice cloning and high parallelism, released under Apache 2.0.
GRAFT is a per-word pronunciation conditioning mechanism for zero-shot text-to-speech that uses a spoken sample of a target word to control its pronunciation, achieving significant improvements in target-word phoneme error rates across multiple languages while preserving speaker similarity.
This paper investigates using text-to-speech (TTS) to generate synthetic training data for spoken question answering in Luxembourgish, a low-resource language, and evaluates multi-source TTS configurations with a parameter-efficient SLAM-style architecture.
The author shares their experience swapping the TTS in their voice agent to a custom model (Banter 1) designed for bilingual Arabic-English conversations, which significantly reduced perceived lag.
Kyutai released Pocket TTS, a text-to-speech model capable of cloning a voice from just 5 seconds of audio, running on CPU and released under the MIT license. It was benchmarked against Kokoro, Supertonic, and Inflect-Nano for English TTS.
This paper introduces Audex, a unified audio-text LLM from NVIDIA that achieves state-of-the-art performance across multiple audio and speech tasks while preserving strong text reasoning capabilities without regression.
LuxTTS is a lightweight voice cloning TTS model, supporting 48kHz high-fidelity output, achieving 150x real-time speed on a single GPU, requiring only 1GB VRAM for local operation, with performance comparable to models ten times its size.
Alexandria is an open-source tool that transforms books into fully-voiced audiobooks using AI-powered script annotation and text-to-speech, with local/cloud LLM support, voice cloning, and a built-in Qwen3-TTS engine.
SPARCLE is a speaker-aware grapheme representation model that uses contrastive learning to align grapheme embeddings with acoustic representations, improving text-to-speech quality especially in low-resource settings.
An open-source, free local voice AI studio that supports voice cloning, voice generation, and global dictation. No API key required, runs entirely locally, and serves as a free alternative to ElevenLabs and WisprFlow.
A developer documents building a self-hosted text-to-speech app on an NVIDIA Jetson Orin Nano using Kokoro-82M and durable streams, enabling reliable local AI inference with shareable audio outputs.
TTS Arena launches as a blind benchmark for text-to-speech models, where users compare anonymous TTS outputs and vote for the more human-sounding one, updating a live leaderboard.