speech-synthesis

Tag

Cards List
#speech-synthesis

Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models

Reddit r/LocalLLaMA ↗ · 16h ago

The article analyzes over 100 audio models and finds that Qwen-family LLMs, especially Qwen3, are widely used as language backbones across various audio tasks like TTS, ASR, and music generation.

0 favorites 0 likes
#speech-synthesis

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Hugging Face Daily Papers ↗ · yesterday Cached

The paper introduces EmoRES-TTS, a training-free method for emotional speech generation that enhances controllability by decomposing emotion vectors into shared and residual components, achieving superior performance over existing methods on benchmarks.

0 favorites 0 likes
#speech-synthesis

BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech

arXiv cs.CL ↗ · 5d ago Cached

BanglaKontho presents a 20-hour single-speaker Bangla TTS corpus from professional audiobooks, along with an open-source text normalizer and preprocessing pipeline, to address the gap in long-form prosody for low-resource Bangla speech synthesis.

0 favorites 0 likes
#speech-synthesis

StillTalk

Product Hunt ↗ · 2026-09-17 Cached

StillTalk is a product that uses AI to turn a generated photo into a talking character, with live rendering in the browser and support for multiple languages and custom sentences.

0 favorites 0 likes
#speech-synthesis

@FinanceYF5: Voice AI is starting to make its way into phones. Open-source Audio8, which packages speech recognition and speech synt…

X AI KOLs Timeline ↗ · 2026-09-16 Cached

The post introduces the open-source Audio8 models, which enable on-device speech recognition and synthesis on phones and PCs, including an offline transcription version for iPhone.

0 favorites 0 likes
#speech-synthesis

@neilzegh: Most voice providers come with a "max" and a "flash" version, with the "max" one being reliable and the "flash" one bei…

X AI KOLs Timeline ↗ · 2026-08-31 Cached

Gradium has released a new Text-to-Speech model that combines high accuracy with low latency, achieving under 250ms time-to-first-audio through intelligent contextual handling and silence trimming.

0 favorites 0 likes
#speech-synthesis

@vvolhejn: Our open-source TTS just got even open-sourcer

X AI KOLs Timeline ↗ · 2026-08-25 Cached

Kyutai Labs has open-sourced their Pocket TTS training stack, including data pipeline, recipes, and evaluations, allowing developers to train text-to-speech models on GPUs and run them on CPUs.

0 favorites 0 likes
#speech-synthesis

Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

arXiv cs.CL ↗ · 2026-08-24 Cached

Poly-InstructTTS is a text-to-speech system that learns expressive speech from open-ended natural language instructions using a large-scale multi-modal dataset, improving instruction adherence and expressiveness in TTS models.

0 favorites 0 likes
#speech-synthesis

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

arXiv cs.CL ↗ · 2026-08-24 Cached

This paper introduces a synthetic Bengali speech dataset of 10,000 audio-text pairs for telecom customer care scenarios, generated using OmniVoice voice-cloning, and evaluates it with an ASR model, achieving low word error rates.

0 favorites 0 likes
#speech-synthesis

Ultrafast Qwen3-TTS at 34 ms Time-to-First-Audio, Handling 10 Requests Per Second [OSS]

Reddit r/LocalLLaMA ↗ · 2026-08-21 Cached

Nari Qwen3-TTS is a high-performance serving implementation for the Qwen3-TTS model, achieving sub-50ms time-to-first-audio and handling 10 requests per second on a single H100 GPU.

0 favorites 0 likes
#speech-synthesis

X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance

arXiv cs.CL ↗ · 2026-08-20 Cached

X2Streaming-TTS presents a causal token-level text-to-speech framework for true streaming synthesis, using causal commitment and speech-state inheritance to handle uncertain text prefixes and maintain acoustic continuity in low-latency spoken dialogue systems.

0 favorites 0 likes
#speech-synthesis

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

arXiv cs.CL ↗ · 2026-08-17 Cached

VoiceChat-TTS is a low-latency, continuous text-to-speech model designed for interactive agents, enabling real-time streaming and interruption handling without compromising speech quality.

0 favorites 0 likes
#speech-synthesis

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

arXiv cs.CL ↗ · 2026-08-04 Cached

Presents DLLM-TTS, a block discrete diffusion language model for text-to-speech synthesis that processes X-Codec2 tokens in blocks, enabling parallel generation with RTF 0.15 while achieving competitive quality with only 20K hours of training data.

0 favorites 0 likes
#speech-synthesis

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

arXiv cs.CL ↗ · 2026-07-03 Cached

SPARCLE is a speaker-aware grapheme representation model that uses contrastive learning to align grapheme embeddings with acoustic representations, improving text-to-speech quality especially in low-resource settings.

0 favorites 0 likes
#speech-synthesis

@cevenif: Bro, it's time to say goodbye to those paid voice tools! The open-source and free Voicebox has arrived, completely crushing paid giants like ElevenLabs and WisprFlow. Features: Voice cloning - instantly become anyone, Global voice input - accessible anytime...

X AI KOLs Timeline ↗ · 2026-07-02 Cached

An open-source, free local voice AI studio that supports voice cloning, voice generation, and global dictation. No API key required, runs entirely locally, and serves as a free alternative to ElevenLabs and WisprFlow.

0 favorites 0 likes
#speech-synthesis

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

arXiv cs.LG ↗ · 2026-06-29 Cached

Proposes HybridCodec, a novel framework combining temporally compressed discrete tokens with continuous residuals to improve speaker characteristic retention in speech language models, reducing autoregressive steps while maintaining quality.

0 favorites 0 likes
#speech-synthesis

owensong/Inflect-Nano-v2

Hugging Face Models Trending ↗ · 2026-06-25 Cached

Release of Inflect-Nano-v2, a fixed-voice English TTS model with under 4M parameters for local text-to-waveform synthesis, supporting CPU or CUDA inference and long-text handling.

0 favorites 0 likes
#speech-synthesis

@LinearUncle: Recommending an open-source voice cloning repository from a Chinese company called Mosi: MOSS-TTS. You read a passage, it clones your voice, then you can use your voice to read any text. Check the post details to see how I used it in practice—it works great and can be indistinguishable from the real thing. https://github.com/OpenMOS…

X AI KOLs Timeline ↗ · 2026-06-19 Cached

MOSS-TTS is an open-source voice cloning model introduced by Mosi Company. Users can clone a voice by reading a small amount of text, and then use the cloned voice to generate any speech with realistic results.

0 favorites 0 likes
#speech-synthesis

Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation

arXiv cs.CL ↗ · 2026-06-12 Cached

This paper presents a low-latency real-time audio game commentary system that uses LLM-based parallel text generation to reduce inter-utterance silence from 9.6 to 0.3 seconds, significantly improving perceived speaking rhythm compared to sequential baselines.

0 favorites 0 likes
#speech-synthesis

Latency matters more than model selection when building AI tutoring systems

Reddit r/AI_Agents ↗ · 2026-06-04

A practitioner argues that speech start latency—not model selection—is the critical factor in AI tutoring systems, recommending targets under 1 second for speech start and highlighting streaming TTS as the highest-leverage optimization. The post outlines a full pipeline from ASR through TTS and avatar sync, identifying where latency compounds most.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback