Tag
The article analyzes over 100 audio models and finds that Qwen-family LLMs, especially Qwen3, are widely used as language backbones across various audio tasks like TTS, ASR, and music generation.
The paper introduces EmoRES-TTS, a training-free method for emotional speech generation that enhances controllability by decomposing emotion vectors into shared and residual components, achieving superior performance over existing methods on benchmarks.
BanglaKontho presents a 20-hour single-speaker Bangla TTS corpus from professional audiobooks, along with an open-source text normalizer and preprocessing pipeline, to address the gap in long-form prosody for low-resource Bangla speech synthesis.
StillTalk is a product that uses AI to turn a generated photo into a talking character, with live rendering in the browser and support for multiple languages and custom sentences.
The post introduces the open-source Audio8 models, which enable on-device speech recognition and synthesis on phones and PCs, including an offline transcription version for iPhone.
Gradium has released a new Text-to-Speech model that combines high accuracy with low latency, achieving under 250ms time-to-first-audio through intelligent contextual handling and silence trimming.
Kyutai Labs has open-sourced their Pocket TTS training stack, including data pipeline, recipes, and evaluations, allowing developers to train text-to-speech models on GPUs and run them on CPUs.
Poly-InstructTTS is a text-to-speech system that learns expressive speech from open-ended natural language instructions using a large-scale multi-modal dataset, improving instruction adherence and expressiveness in TTS models.
This paper introduces a synthetic Bengali speech dataset of 10,000 audio-text pairs for telecom customer care scenarios, generated using OmniVoice voice-cloning, and evaluates it with an ASR model, achieving low word error rates.
Nari Qwen3-TTS is a high-performance serving implementation for the Qwen3-TTS model, achieving sub-50ms time-to-first-audio and handling 10 requests per second on a single H100 GPU.
X2Streaming-TTS presents a causal token-level text-to-speech framework for true streaming synthesis, using causal commitment and speech-state inheritance to handle uncertain text prefixes and maintain acoustic continuity in low-latency spoken dialogue systems.
VoiceChat-TTS is a low-latency, continuous text-to-speech model designed for interactive agents, enabling real-time streaming and interruption handling without compromising speech quality.
Presents DLLM-TTS, a block discrete diffusion language model for text-to-speech synthesis that processes X-Codec2 tokens in blocks, enabling parallel generation with RTF 0.15 while achieving competitive quality with only 20K hours of training data.
SPARCLE is a speaker-aware grapheme representation model that uses contrastive learning to align grapheme embeddings with acoustic representations, improving text-to-speech quality especially in low-resource settings.
An open-source, free local voice AI studio that supports voice cloning, voice generation, and global dictation. No API key required, runs entirely locally, and serves as a free alternative to ElevenLabs and WisprFlow.
Proposes HybridCodec, a novel framework combining temporally compressed discrete tokens with continuous residuals to improve speaker characteristic retention in speech language models, reducing autoregressive steps while maintaining quality.
Release of Inflect-Nano-v2, a fixed-voice English TTS model with under 4M parameters for local text-to-waveform synthesis, supporting CPU or CUDA inference and long-text handling.
MOSS-TTS is an open-source voice cloning model introduced by Mosi Company. Users can clone a voice by reading a small amount of text, and then use the cloned voice to generate any speech with realistic results.
This paper presents a low-latency real-time audio game commentary system that uses LLM-based parallel text generation to reduce inter-utterance silence from 9.6 to 0.3 seconds, significantly improving perceived speaking rhythm compared to sequential baselines.
A practitioner argues that speech start latency—not model selection—is the critical factor in AI tutoring systems, recommending targets under 1 second for speech start and highlighting streaming TTS as the highest-leverage optimization. The post outlines a full pipeline from ASR through TTS and avatar sync, identifying where latency compounds most.