tts

Tag

Cards List
#tts

@rohanpaul_ai: I wasn't expecting Soniox TTS v2 (a text-to-speech model) to sound this natural. They just released this TTS v2 > Reall…

X AI KOLs Following · 3d ago Cached

Soniox TTS v2 is a new text-to-speech model offering premium voice quality, expressive control via audio tags, high-fidelity voice cloning, support for 60+ languages, and low-latency streaming, priced at $0.70 per generated hour.

0 favorites 0 likes
#tts

How to achieve sub-800ms latency and 50% lower telephony costs in Voice AI pipelines (Architecture Breakdown)

Reddit r/AI_Agents · 5d ago

A technical breakdown of an enterprise Voice AI architecture that cuts telephony costs by 40-60% via wholesale carriers and achieves sub-500ms latency using Deepgram, Claude/GPT-4o-mini, and ElevenLabs/Cartesia, orchestrated through n8n and Supabase.

0 favorites 0 likes
#tts

@hank_aibtc: Whoa, this thing really blew my mind. Dug up a local tool called KrillinAI, free, video translation, precise subtitles, ultra-natural voiceover, voice cloning, the whole pipeline in one go. Chinese-English translation is especially stable, Whisper recognition accuracy is ridiculously high, LLM translates segment by segment without losing context, CosyV…

X AI KOLs Timeline · 5d ago Cached

Introducing KrillinAI, a free locally-run video translation tool that supports precise subtitles, natural voiceover, and voice cloning. It integrates Whisper, LLM, and CosyVoice, and supports Win/Mac and yt-dlp.

0 favorites 0 likes
#tts

@Huahuazo: The most annoying part of reposting overseas videos? Manually cutting subtitles, eye-straining translation alignment, and guessing audio sync—this workflow is so inefficient it makes you want to smash your keyboard. VideoLingo connects the entire process into an automated pipeline. It has 7k+ stars on GitHub and an MIT license, so it's reliable to use. …

X AI KOLs Timeline · 6d ago Cached

VideoLingo is an open-source video translation, localization, and dubbing tool that uses WhisperX and AI to deliver Netflix-level subtitles and multilingual dubbing. It supports downloading via yt-dlp and multiple TTS options, helping video reposters automate the entire workflow.

0 favorites 0 likes
#tts

I built a local realtime voice stack for Ollama: Parakeet STT → Qwen 2.5 7B → Qwen3-TTS

Reddit r/LocalLLaMA · 2026-08-08

The author built a local realtime voice stack using Parakeet STT, Qwen 2.5 7B, and Qwen3-TTS, integrated with Ollama.

0 favorites 0 likes
#tts

Shipped a Hindi-English voice agent for a fintech. Here's everything that broke and what actually fixed it

Reddit r/AI_Agents · 2026-08-07

A developer shares a postmortem of building a Hindi-English voice agent for fintech, highlighting challenges with number readback, code-mixed TTS, latency under load, and compliance. Key fix was choosing TTS with first-class support for Indian code-mixing and testing at real concurrency.

0 favorites 0 likes
#tts

@yiyirats: When relying on third-party voice services, data privacy and stability are always factors to consider. Voicebox is a locally run AI voice studio; all processing is done locally, data is not uploaded to the cloud, and no account registration is required. The features are quite comprehensive: supports Qwen3-TTS, LuxTTS, Chatterbo…

X AI KOLs Timeline · 2026-08-06 Cached

Voicebox is a locally run open-source AI voice studio that supports 7 TTS engines, 23 languages, and voice cloning. All processing is done locally to protect privacy. The project has received 33.8k stars on GitHub.

0 favorites 0 likes
#tts

@SamuelZengML: We’re genuinely surprised and grateful to see our latest open-source model, Audio8/Audio8-TTS-Preview-0.6b, reach #1 on…

X AI KOLs Following · 2026-08-03 Cached

Audio8's open-source TTS model reached #1 on Hugging Face's TTS trending list and #10 overall, with the team expressing gratitude and plans to keep improving.

0 favorites 0 likes
#tts

@RemiCadene: Gradium better than ElevenLabs and Cartesia?

X AI KOLs Following · 2026-07-30 Cached

Gradium has released a new TTS model in public beta that accurately reads phone numbers, emails, IBANs, and time expressions natively, offering API access with 1M credits for testing.

0 favorites 0 likes
#tts

You can now fine-tune my 3.96M-parameter TTS on your own voice or language

Reddit r/LocalLLaMA · 2026-07-27

A new toolkit enables fine-tuning the tiny Inflect-Nano/Micro TTS models on custom voice and language, supporting warm-start, resumption, and export to PyTorch/ONNX.

0 favorites 0 likes
#tts

[audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains

Reddit r/LocalLLaMA · 2026-07-24

Release 0.4 of audio.cpp adds C++/GGML inference for Higgs Audio v3 TTS 4B (10x real-time) and Fish Audio S2 Pro, with full GGUF loading and Q8 speed/VRAM gains.

0 favorites 0 likes
#tts

We built NeuTTS-2E, an open-source on-device TTS model with 7 controllable emotions

Reddit r/LocalLLaMA · 2026-07-22

NeuTTS-2E is an open-source on-device TTS model that supports seven controllable emotions.

0 favorites 0 likes
#tts

@LangChain: Full audio on your traces STT/TTS latency interruptions + VAD Only a few lines of code to set up

X AI KOLs Timeline · 2026-07-21 Cached

LangChain launches LangSmith tracing for voice frameworks (Pipecat, LiveKit, OpenAI Realtime, Gemini Live), enabling full audio monitoring, STT/TTS latency tracking, interruption detection, and VAD analysis with minimal code.

0 favorites 0 likes
#tts

TTS curated list for voice agent builders — focused on streaming latency and mid-stream cancellation

Reddit r/ArtificialInteligence · 2026-07-21 Cached

A curated list of text-to-speech resources for voice agent builders, organized around the decision between real-time streaming synthesis and offline high-fidelity synthesis, with emphasis on streaming latency and mid-stream cancellation.

0 favorites 0 likes
#tts

Curated open-source TTS reference — organized by license (because half the top open-weight models can't be shipped commercially)

Reddit r/AI_Agents · 2026-07-21

A curated reference of open-source text-to-speech models organized by license type, highlighting which models can be used commercially.

0 favorites 0 likes
#tts

@HuggingApps: NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words transcription, translation, sou…

X AI KOLs Following · 2026-07-20 Cached

NVIDIA released Nemotron, an audio-native model capable of transcription, translation, sound recognition, audio Q&A, TTS, and full speech-to-speech, with open weights in 2B and 30B sizes.

0 favorites 0 likes
#tts

Introducing Scylla's Band, a new TTS model + inference framework with Android sample!

Reddit r/LocalLLaMA · 2026-07-20

Scylla's Band is a new TTS model and inference framework, including an Android sample for deployment.

0 favorites 0 likes
#tts

[audio.cpp] 10 hours of audio generated in 3 minutes on RTX 5090 (demo included)! C++/GGML based Supertonic 3, MOSS-TTS, IndexTTS2, and Irodori-TTS released

Reddit r/LocalLLaMA · 2026-07-15

Release of C++/GGML based implementations of Supertonic 3, MOSS-TTS, IndexTTS2, and Irodori-TTS in audio.cpp, capable of generating 10 hours of audio in 3 minutes on an RTX 5090.

0 favorites 0 likes
#tts

Self-hosted voice for any agent/harness of your choice (open-source)

Reddit r/LocalLLaMA · 2026-07-13

The Cicero project enables self-hosted, bidirectional voice conversation with AI agents, supporting local TTS/STT and integration with various agent protocols and tools like Claude Code and Hermes Agent.

0 favorites 0 likes
#tts

@mattpocockuk: Thinking about a workflow like this for helping prevent comprehension debt on a fast-moving repo: Fast-moving repo w/lo…

X AI KOLs Following · 2026-07-12 Cached

Matt Pocock shares an idea for a workflow using LLMs to generate podcast summaries of code diffs, aiming to prevent comprehension debt in fast-moving repos; he has implemented it for his personal wiki with good results.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback