text-to-speech

Tag

Cards List
#text-to-speech

@MaximeRivest: This device is very impressive. its got microphone, speaker support, text to speech not, ocr, a chat gpt 5.2 chat and i…

X AI KOLs Following · yesterday Cached

A user shares hands-on impressions of a thin, light Android-based device featuring microphone, speaker, OCR, text-to-speech, and a ChatGPT 5.2 chat, noting the only clear drawback is its lack of color.

0 favorites 0 likes
#text-to-speech

🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp

Reddit r/LocalLLaMA · 2d ago

NVIDIA's entire speech stack—ASR, TTS, and codec—is now quantized to GGUF and runs locally on-device via NeMo-Speech.cpp, with new model releases for Magpie-TTS, Nemotron Speech Streaming, and Parakeet.

0 favorites 0 likes
#text-to-speech

Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM

Reddit r/LocalLLaMA · 3d ago

Scenema Audio, an expressive text-to-speech model with zero-shot voice cloning, is now available as a native ComfyUI custom node, quantized to run on 8GB VRAM. The release adds inline stage direction cues, 12 preset voices, and simplifies the prompt format for ComfyUI.

0 favorites 0 likes
#text-to-speech

Building a Fully Local PDF Read-Aloud & PDF-to-Audiobook Desktop App with Kokoro 82M, Qwen, and llama.cpp

Reddit r/LocalLLaMA · 4d ago

Speechfony is a new fully local desktop app that reads PDFs and EPUBs aloud with sentence highlighting, semantic search, and MP3 audiobook export, using Kokoro TTS and on-device embeddings. It prioritizes privacy and offline use, with an open-source MIT license.

0 favorites 0 likes
#text-to-speech

Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support

Reddit r/LocalLLaMA · 4d ago

Qwen3-TTS voice cloning has been merged into mainline llama.cpp, enabling local text-to-speech with voice cloning from short reference audio via the llama-tts binary, supporting multiple languages. Limitations remain, including only the Base model and no server endpoint yet.

0 favorites 0 likes
#text-to-speech

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

arXiv cs.CL · 5d ago Cached

Presents DLLM-TTS, a block discrete diffusion language model for text-to-speech synthesis that processes X-Codec2 tokens in blocks, enabling parallel generation with RTF 0.15 while achieving competitive quality with only 20K hours of training data.

0 favorites 0 likes
#text-to-speech

@DanKornas: myshell-ai/OpenVoice instant voice cloning with style and language control GitHub: Archive:

X AI KOLs Timeline · 2026-08-02 Cached

OpenVoice is an open-source instant voice cloning model with style and language control, now available on GitHub.

0 favorites 0 likes
#text-to-speech

@CycleDecoded: Shanghai AI Laboratory (OpenMMLab) has completely demolished "sound creation" this time. Their open-source Amphion is simply an "all-purpose arsenal" for the audio-visual creation world. From speech synthesis to AI singing voice conversion, this thing can beat the vast majority of paid voice software. Basic Info Project Name:…

X AI KOLs Timeline · 2026-08-01 Cached

OpenMMLab has open-sourced Amphion, an audio generation toolbox supporting TTS, singing voice conversion, sound effect generation, and more. It is completely free for commercial use and supports local deployment.

0 favorites 0 likes
#text-to-speech

@PrajwalTomar_: You are massively underestimating what just dropped. The #1 model on the blind-test voice leaderboard is not ElevenLabs…

X AI KOLs Following · 2026-08-01 Cached

Simba 3.2 from SpeechifyAI claims the #1 spot on the blind-test voice leaderboard, surpassing ElevenLabs, OpenAI, and Google DeepMind, at a significantly lower cost with a new API and free tier.

0 favorites 0 likes
#text-to-speech

@RemiCadene: Gradium better than ElevenLabs and Cartesia?

X AI KOLs Following · 2026-07-30 Cached

Gradium has released a new TTS model in public beta that accurately reads phone numbers, emails, IBANs, and time expressions natively, offering API access with 1M credits for testing.

0 favorites 0 likes
#text-to-speech

Fish Audio launches S2.1 Pro with support for 83 languages (2 minute read)

TLDR AI · 2026-07-29 Cached

Fish Audio launches S2.1 Pro, a production voice model with 90ms latency, support for 83 languages, voice cloning from short samples, and multi-speaker dialogue, available via API with a free tier for development.

0 favorites 0 likes
#text-to-speech

Audio8/Audio8-TTS-Preview-0.6b

Hugging Face Models Trending · 2026-07-28 Cached

Audio8 releases a 0.6B-parameter multilingual text-to-speech model with zero-shot voice cloning capabilities, available on Hugging Face under Apache 2.0 license.

0 favorites 0 likes
#text-to-speech

Liso

Product Hunt · 2026-07-23

Liso lets you highlight text on any webpage and converts it into audio for listening.

0 favorites 0 likes
#text-to-speech

We built NeuTTS-2E, an open-source on-device TTS model with 7 controllable emotions

Reddit r/LocalLLaMA · 2026-07-22

NeuTTS-2E is an open-source on-device TTS model that supports seven controllable emotions.

0 favorites 0 likes
#text-to-speech

@DanKornas: Writing a two-person audio demo shouldn’t mean stitching together separate speech clips. Dia is a 1.6B-parameter text-t…

X AI KOLs Timeline · 2026-07-21 Cached

Dia is a 1.6B-parameter open-source text-to-speech model that generates English dialogue from transcripts, supporting two-speaker generation, audio conditioning, and nonverbal cues.

0 favorites 0 likes
#text-to-speech

TTS curated list for voice agent builders — focused on streaming latency and mid-stream cancellation

Reddit r/ArtificialInteligence · 2026-07-21 Cached

A curated list of text-to-speech resources for voice agent builders, organized around the decision between real-time streaming synthesis and offline high-fidelity synthesis, with emphasis on streaming latency and mid-stream cancellation.

0 favorites 0 likes
#text-to-speech

Curated open-source TTS reference — organized by license (because half the top open-weight models can't be shipped commercially)

Reddit r/AI_Agents · 2026-07-21

A curated reference of open-source text-to-speech models organized by license type, highlighting which models can be used commercially.

0 favorites 0 likes
#text-to-speech

@Ali_TongyiLab: Qwen-Audio-3.0-TTS is here. Our latest text-to-speech model, in two flavors: • Flash: real-time interaction • Plus: hig…

X AI KOLs Timeline · 2026-07-20 Cached

Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a new text-to-speech model with Flash (real-time) and Plus (high-quality) versions, supporting 16 languages, natural language style control, and robust voice cloning.

0 favorites 0 likes
#text-to-speech

I built a fully-local and speedy MacOS utility for text to speech and dictation, running top of the range AI models

Reddit r/artificial · 2026-07-20

A developer created a MacOS utility for local text-to-speech and dictation using advanced AI models, emphasizing speed and privacy.

0 favorites 0 likes
#text-to-speech

Introducing Scylla's Band, a new TTS model + inference framework with Android sample!

Reddit r/LocalLLaMA · 2026-07-20

Scylla's Band is a new TTS model and inference framework, including an Android sample for deployment.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback