text-to-speech

Tag

Cards List
#text-to-speech

@Sumanth_077: Open-source framework for building real-time voice AI agents! Pipecat is a Python framework for orchestrating audio, vi…

X AI KOLs Timeline · 2026-07-16 Cached

Pipecat is an open-source Python framework for building real-time voice AI agents, handling speech recognition, text-to-speech, conversation logic, and supporting multiple AI service providers.

0 favorites 0 likes
#text-to-speech

​ElevenLabs Review (2026) — Why I rate it an 8.1/10 (and who should actually avoid paying for it)

Reddit r/ArtificialInteligence · 2026-07-15

A detailed, unsponsored review of ElevenLabs rating it 8.1/10, highlighting its emotional range and low latency as strengths, but cautioning about high costs and the 'regeneration tax' for casual users and high-volume publishers.

0 favorites 0 likes
#text-to-speech

Introducing Real World VoiceEQ: Measuring the human quality of voice AI

Hugging Face Blog · 2026-07-15 Cached

Real World VoiceEQ is a new benchmark for evaluating the human quality of voice AI, based on over a million human ratings, assessing models across speech recognition, synthesis, and understanding in real-world conditions.

0 favorites 0 likes
#text-to-speech

VocalVia

Product Hunt · 2026-07-13

VocalVia is a tool that converts documents and articles into editable multi-voice audio.

0 favorites 0 likes
#text-to-speech

FreyaTTS Technical Report

arXiv cs.CL · 2026-07-13 Cached

FreyaTTS is a compact, tokenizer-free Turkish-first text-to-speech model based on a non-autoregressive conditional flow-matching Diffusion Transformer, achieving state-of-the-art performance with a fraction of the parameters of larger systems and released under Apache-2.0.

0 favorites 0 likes
#text-to-speech

Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

arXiv cs.CL · 2026-07-10 Cached

This paper identifies a confound in best-of-N TTS evaluation where the apparent quality of ASR verifiers depends strongly on which ASR family is used as evaluator. The authors propose cross-family rank ensembles that achieve lower word error rates across multiple evaluators.

0 favorites 0 likes
#text-to-speech

I tested an AI pipeline that turns a topic into a full podcast episode without manual editing

Reddit r/ArtificialInteligence · 2026-07-07

A report on testing an AI pipeline that automatically generates a full podcast episode from a given topic without any manual editing.

0 favorites 0 likes
#text-to-speech

Local, CPU-Friendly, High-Quality TTS (Text-to-Speech) with Kokoro

Hacker News Top · 2026-07-07 Cached

This article introduces Kokoro, a lightweight 82M-parameter text-to-speech model that runs locally on CPU, providing high-quality speech synthesis across multiple languages while preserving privacy. It explains how to set up Kokoro via a Docker container with an OpenAI-compatible API for easy integration.

0 favorites 0 likes
#text-to-speech

Gepard : 0.6B streaming TTS built for real-time dialogue - 20× realtime factor, ~50ms time-to-first-audio, vLLM-native, Apache 2.0

Reddit r/LocalLLaMA · 2026-07-07 Cached

Gepard is a new streaming TTS model capable of real-time dialogue with ~50ms time-to-first-audio, supporting voice cloning and high parallelism, released under Apache 2.0.

0 favorites 0 likes
#text-to-speech

GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

arXiv cs.LG · 2026-07-07 Cached

GRAFT is a per-word pronunciation conditioning mechanism for zero-shot text-to-speech that uses a spoken sample of a target word to control its pronunciation, achieving significant improvements in target-word phoneme error rates across multiple languages while preserving speaker similarity.

0 favorites 0 likes
#text-to-speech

LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

arXiv cs.CL · 2026-07-07 Cached

This paper investigates using text-to-speech (TTS) to generate synthetic training data for spoken question answering in Luxembourgish, a low-resource language, and evaluates multi-source TTS configurations with a parameter-efficient SLAM-style architecture.

0 favorites 0 likes
#text-to-speech

I swapped the TTS in my voice agent and it cut the lag people actually feel more than anything else

Reddit r/AI_Agents · 2026-07-06

The author shares their experience swapping the TTS in their voice agent to a custom model (Banter 1) designed for bilingual Arabic-English conversations, which significantly reduced perceived lag.

0 favorites 0 likes
#text-to-speech

Kyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTS

Reddit r/LocalLLaMA · 2026-07-06

Kyutai released Pocket TTS, a text-to-speech model capable of cloning a voice from just 5 seconds of audio, running on CPU and released under the MIT license. It was benchmarked against Kokoro, Supertonic, and Inflect-Nano for English TTS.

0 favorites 0 likes
#text-to-speech

Unified Audio Intelligence Without Regressing on Text Intelligence

Hugging Face Daily Papers · 2026-07-06 Cached

This paper introduces Audex, a unified audio-text LLM from NVIDIA that achieves state-of-the-art performance across multiple audio and speech tasks while preserving strong text reasoning capabilities without regression.

0 favorites 0 likes
#text-to-speech

@NFTCPS: Fraud call centers have a new weapon — voice cloning has been pushed to new heights again. LuxTTS, a lightweight TTS model, after seeing it I can only say: truly insane. Fast: 150x real-time on a single GPU, even runs faster than real speech on CPU. Clear: 48kHz directly, most models are still stuck at 24kHz…

X AI KOLs Timeline · 2026-07-05 Cached

LuxTTS is a lightweight voice cloning TTS model, supporting 48kHz high-fidelity output, achieving 150x real-time speed on a single GPU, requiring only 1GB VRAM for local operation, with performance comparable to models ten times its size.

0 favorites 0 likes
#text-to-speech

@tom_doerr: Alexandria generates audiobooks from books using AI https://github.com/Finrandojin/alexandria-audiobook…

X AI KOLs Timeline · 2026-07-04 Cached

Alexandria is an open-source tool that transforms books into fully-voiced audiobooks using AI-powered script annotation and text-to-speech, with local/cloud LLM support, voice cloning, and a built-in Qwen3-TTS engine.

0 favorites 0 likes
#text-to-speech

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

arXiv cs.CL · 2026-07-03 Cached

SPARCLE is a speaker-aware grapheme representation model that uses contrastive learning to align grapheme embeddings with acoustic representations, improving text-to-speech quality especially in low-resource settings.

0 favorites 0 likes
#text-to-speech

@cevenif: Bro, it's time to say goodbye to those paid voice tools! The open-source and free Voicebox has arrived, completely crushing paid giants like ElevenLabs and WisprFlow. Features: Voice cloning - instantly become anyone, Global voice input - accessible anytime...

X AI KOLs Timeline · 2026-07-02 Cached

An open-source, free local voice AI studio that supports voice cloning, voice generation, and global dictation. No API key required, runs entirely locally, and serves as a free alternative to ElevenLabs and WisprFlow.

0 favorites 0 likes
#text-to-speech

Serving Local AI on my Jetson through Durable Streams

Lobsters Hottest · 2026-06-30 Cached

A developer documents building a self-hosted text-to-speech app on an NVIDIA Jetson Orin Nano using Kokoro-82M and durable streams, enabling reliable local AI inference with shareable audio outputs.

0 favorites 0 likes
#text-to-speech

@realmrfakename: TTS Arena is live. A blind benchmark for text-to-speech, rebuilt from the ground up - listen to two anonymous models, p…

X AI KOLs Following · 2026-06-29 Cached

TTS Arena launches as a blind benchmark for text-to-speech models, where users compare anonymous TTS outputs and vote for the more human-sounding one, updating a live leaderboard.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback