text-to-speech

Tag

Cards List
#text-to-speech

@realmrfakename: TTS Arena is live. A blind benchmark for text-to-speech, rebuilt from the ground up - listen to two anonymous models, p…

X AI KOLs Following · 2026-06-29 Cached

TTS Arena launches as a blind benchmark for text-to-speech models, where users compare anonymous TTS outputs and vote for the more human-sounding one, updating a live leaderboard.

0 favorites 0 likes
#text-to-speech

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

arXiv cs.LG · 2026-06-29 Cached

Proposes HybridCodec, a novel framework combining temporally compressed discrete tokens with continuous residuals to improve speaker characteristic retention in speech language models, reducing autoregressive steps while maintaining quality.

0 favorites 0 likes
#text-to-speech

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

arXiv cs.CL · 2026-06-26 Cached

This paper investigates LoRA fine-tuning of the VoxCPM2 TTS model to improve quality for low-resource languages like Khmer, while showing no gain for Korean which the base model already handles well. The adapter yields significant MOS improvement for Khmer with minimal parameter training.

0 favorites 0 likes
#text-to-speech

audio.cpp: 12 audio models (Qwen3-TTS, PocketTTS, VeVo2 etc) in 1 C++/ggml runtime — TTS up to 5x faster than Python on CUDA

Reddit r/LocalLLaMA · 2026-06-25

audio.cpp is a C++/ggml runtime that integrates 12 audio models including Qwen3-TTS, PocketTTS, and VeVo2, achieving TTS up to 5x faster than Python on CUDA.

0 favorites 0 likes
#text-to-speech

owensong/Inflect-Micro-v2

Hugging Face Models Trending · 2026-06-25 Cached

Inflect-Micro-v2 is a compact text-to-speech model with under 10M parameters, supporting CPU/CUDA inference and long-text handling, released on Hugging Face.

0 favorites 0 likes
#text-to-speech

owensong/Inflect-Nano-v2

Hugging Face Models Trending · 2026-06-25 Cached

Release of Inflect-Nano-v2, a fixed-voice English TTS model with under 4M parameters for local text-to-waveform synthesis, supporting CPU or CUDA inference and long-text handling.

0 favorites 0 likes
#text-to-speech

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

arXiv cs.CL · 2026-06-25 Cached

This paper presents a modular end-to-end speech-to-speech conversational system for the low-resource Algerian Dialect, integrating ASR, NLU, RAG, and TTS with dedicated datasets and fine-tuned models.

0 favorites 0 likes
#text-to-speech

I tried making an AI World Cup commentator. It sounds real until the game gets fast

Reddit r/ArtificialInteligence · 2026-06-23

A personal experiment building an AI commentator for World Cup matches reveals realistic results until fast-paced gameplay causes issues.

0 favorites 0 likes
#text-to-speech

A fully local voice assistant setup

Lobsters Hottest · 2026-06-22 Cached

A guide to building a fully local voice assistant using Platypush on a Raspberry Pi, covering hotword detection, speech-to-text, text-to-speech, and home automation integration.

0 favorites 0 likes
#text-to-speech

@LinearUncle: Recommending an open-source voice cloning repository from a Chinese company called Mosi: MOSS-TTS. You read a passage, it clones your voice, then you can use your voice to read any text. Check the post details to see how I used it in practice—it works great and can be indistinguishable from the real thing. https://github.com/OpenMOS…

X AI KOLs Timeline · 2026-06-19 Cached

MOSS-TTS is an open-source voice cloning model introduced by Mosi Company. Users can clone a voice by reading a small amount of text, and then use the cloned voice to generate any speech with realistic results.

0 favorites 0 likes
#text-to-speech

Voice feels like the underrated output layer for AI agents

Reddit r/AI_Agents · 2026-06-18

The article discusses the underutilized potential of voice as an output layer for AI agents, highlighting practical use cases and workflow challenges beyond simple text-to-speech.

0 favorites 0 likes
#text-to-speech

@Gorden_Sun: NetEase Youdao open-sources Confucius4-TTS, a 1.3B TTS model, supports multilingual, supports voice cloning, good results, very fast. Github: https://github.com/netease-youdao/Confucius4-TTS… Online demo: …

X AI KOLs Timeline · 2026-06-18 Cached

NetEase Youdao open-sourced the 1.3B parameter Confucius4-TTS model, supporting zero-shot voice cloning and cross-lingual speech synthesis in 14 languages, fast and with excellent results.

0 favorites 0 likes
#text-to-speech

@lmsysorg: SGLang-Omni now serves MOSS-TTS-Local Transformer v1.5 from @Open_MOSS on day 0! This is an open 48 kHz stereo TTS mode…

X AI KOLs Timeline · 2026-06-18 Cached

MOSS-TTS-Local Transformer v1.5 is an open-source 48 kHz stereo TTS model with zero-shot voice cloning, native streaming, and support for 31 languages, built on a Qwen3-4B backbone and served via SGLang-Omni.

0 favorites 0 likes
#text-to-speech

@MosiAI_Official: MOSS-TTS Local Transformer v1.5 is here. Clone any voice. Speak any language. Hear every detail. 30+ languages, 48 kHz …

X AI KOLs Following · 2026-06-18 Cached

MosiAI has released MOSS-TTS Local Transformer v1.5, a text-to-speech model that supports voice cloning, over 30 languages, and high-quality 48 kHz output.

0 favorites 0 likes
#text-to-speech

I released Inflect-Nano, an ultra-extreme tiny 4.63m parameter TTS model.

Reddit r/LocalLLaMA · 2026-06-17

Inflect-Nano, an ultra-extreme tiny 4.63 million parameter text-to-speech model, has been released.

0 favorites 0 likes
#text-to-speech

@_philschmid: QoL for Speech Generation! You can now stream audio from Gemini TTS as it's generated. No more waiting. Build voice ass…

X AI KOLs Following · 2026-06-17 Cached

Google's Gemini TTS now supports streaming audio generation, allowing developers to build voice applications that start speaking instantly without waiting for full audio output.

0 favorites 0 likes
#text-to-speech

@FakeMaidenMaker: Explosive! This open-source project converts text to human-like voice for free, can clone anyone's voice, and adjust timbre with text! GitHub has garnered 30K stars, from Mianbao Intelligent OpenBMB, VoxCPM previously topped both GitHub and HuggingFace charts. Do...

X AI KOLs Timeline · 2026-06-17 Cached

VoxCPM2 is an open-source speech synthesis model from OpenBMB, using a tokenizer-free diffusion autoregressive architecture, supporting 30 languages, voice design, and controllable voice cloning. It can clone a voice with just one sentence, or create a brand new voice using text, outputting 48kHz high-quality audio, and is commercially usable.

0 favorites 0 likes
#text-to-speech

owensong/Inflect-Nano-v1

Hugging Face Models Trending · 2026-06-16 Cached

Inflect-Nano-v1 is a tiny English text-to-speech model with 4.63M total inference parameters, including its vocoder, designed for local, efficient speech synthesis experiments.

0 favorites 0 likes
#text-to-speech

@HuggingModels: Imagine a text-to-speech model that sounds this natural, with 82M parameters and 11M+ downloads. Kokoro-82M is here, an…

X AI KOLs Timeline · 2026-06-16 Cached

Kokoro-82M is a highly natural text-to-speech model with 82 million parameters and over 11 million downloads, representing a significant advancement in AI voice generation.

0 favorites 0 likes
#text-to-speech

@svpino: There's no way call centers stay in business after this. Listen to this conversation. You cannot tell I'm speaking to a…

X AI KOLs Following · 2026-06-15 Cached

Cartesia released Sonic-3.5 (text-to-speech) and Ink-2 (speech-to-text), claiming they are the #1 streaming models for voice agents, with potential to disrupt call centers.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback