tts

Tag

Cards List
#tts

@OpenBMB: Build with VoxCPM: Whispera — Your Local AI Voice Assistant What if your AI assistant could listen, think, remember, an…

X AI KOLs Timeline · 2026-07-11 Cached

Whispera 是一个基于 VoxCPM 的 Windows 本地实时语音助手,集成了 SenseVoice ASR、llama-server 本地 LLM 推理、VoxCPM 流式 TTS 和 Mem0 长期记忆,完全离线运行。

0 favorites 0 likes
#tts

Maya-2-Native reaches #2 on Voice Arena's Hindi TTS leaderboard, trailing only Gemini 3.1 Flash.

Reddit r/artificial · 2026-07-07

Maya-2-Native has achieved the #2 position on Voice Arena's Hindi TTS leaderboard, surpassed only by Gemini 3.1 Flash.

0 favorites 0 likes
#tts

CPU TTS benchmark with UTMOS MOS scoring: Kokoro, Supertonic, Inflect-Nano, and Kyutai's new Pocket TTS [P]

Reddit r/MachineLearning · 2026-07-06

A CPU TTS benchmark compares Kokoro, Supertonic, Inflect-Nano, and Kyutai's Pocket TTS using UTMOS MOS scores, revealing interesting findings about RTF scaling, UTMOS limitations with small vocoders, and undocumented output caps. Pocket TTS offers unique zero-shot voice cloning on CPU.

0 favorites 0 likes
#tts

Voice agents, demystified: STT+TTS and 4 demo agents you can talk to in the browser + build yours with RAG and Tools

Reddit r/AI_Agents · 2026-07-03

A guide demystifying voice agents using speech-to-text and text-to-speech, with four browser-based demo agents and instructions to build your own using RAG and tools.

0 favorites 0 likes
#tts

@0x0SojalSec: Stop paying for ElevenLabs or cloud TTS. Free Clone voice in just 3 seconds fully Locally on Laptop Turns your docs int…

X AI KOLs Timeline · 2026-07-01 Cached

A free, fully local voice cloning and TTS tool powered by Qwen3-TTS and Kokoro runs on Apple Silicon via MLX, enabling studio-quality audiobook generation from PDFs without cloud services.

0 favorites 0 likes
#tts

[audio.cpp] VibeVoice 1.5B released — 90-min podcast in 22.95 min, 4.08x real-time, 2.86x faster than Python without quantization. Native C++/ggml

Reddit r/LocalLLaMA · 2026-07-01

VibeVoice 1.5B, a long-form multi-speaker TTS model, is now supported in audio.cpp, a native C++/ggml runtime, achieving 4.08x real-time speed on RTX 5090, 2.86x faster than Python baseline without quantization.

0 favorites 0 likes
#tts

Qwen3-tts.cpp + Compose Desktop GUI

Reddit r/LocalLLaMA · 2026-06-29

The developer improved qwen3-tts.cpp to run 5x realtime on RTX 5080 and created a cross-platform desktop GUI with Kotlin Compose Multiplatform, featuring voice cloning, streaming, and speaker embedding management.

0 favorites 0 likes
#tts

owensong/Inflect-Micro-v2

Hugging Face Models Trending · 2026-06-25 Cached

Inflect-Micro-v2 is a compact text-to-speech model with under 10M parameters, supporting CPU/CUDA inference and long-text handling, released on Hugging Face.

0 favorites 0 likes
#tts

@codersoar: Created an "open-source, completely free" HTTP API service that provides image OCR recognition, multi-language translation, web content retrieval, face and facial landmark detection, QR/barcode recognition, and text-to-speech (TTS). Feel free to try it if you need these features! Principle: macOS comes with...

X AI KOLs Timeline · 2026-06-21 Cached

A fully free and open-source HTTP API service based on macOS native capabilities, offering image OCR, multi-language translation, web content retrieval, face recognition, QR/barcode recognition, and text-to-speech.

0 favorites 0 likes
#tts

@lmsysorg: SGLang-Omni now serves MOSS-TTS-Local Transformer v1.5 from @Open_MOSS on day 0! This is an open 48 kHz stereo TTS mode…

X AI KOLs Timeline · 2026-06-18 Cached

MOSS-TTS-Local Transformer v1.5 is an open-source 48 kHz stereo TTS model with zero-shot voice cloning, native streaming, and support for 31 languages, built on a Qwen3-4B backbone and served via SGLang-Omni.

0 favorites 0 likes
#tts

@MosiAI_Official: MOSS-TTS Local Transformer v1.5 is here. Clone any voice. Speak any language. Hear every detail. 30+ languages, 48 kHz …

X AI KOLs Following · 2026-06-18 Cached

MosiAI has released MOSS-TTS Local Transformer v1.5, a text-to-speech model that supports voice cloning, over 30 languages, and high-quality 48 kHz output.

0 favorites 0 likes
#tts

I released Inflect-Nano, an ultra-extreme tiny 4.63m parameter TTS model.

Reddit r/LocalLLaMA · 2026-06-17

Inflect-Nano, an ultra-extreme tiny 4.63 million parameter text-to-speech model, has been released.

0 favorites 0 likes
#tts

Your voice agent probably isn't slow because of the LLM.

Reddit r/AI_Agents · 2026-06-17

A developer debunks the common belief that LLM latency is the primary cause of slow voice agents, explaining that delays often stem from earlier stages like audio capture, VAD, and STT. They recommend logging specific latency metrics and testing various STT/TTS providers and orchestration frameworks to diagnose issues.

0 favorites 0 likes
#tts

A structured path for learning to build voice agents, from your first STT call to production

Reddit r/AI_Agents · 2026-06-17

A curated, open-source learning path for building voice agents, covering from STT to production, with 190+ resources and a 5-week plan.

0 favorites 0 likes
#tts

@HuggingModels: Imagine a text-to-speech model that sounds this natural, with 82M parameters and 11M+ downloads. Kokoro-82M is here, an…

X AI KOLs Timeline · 2026-06-16 Cached

Kokoro-82M is a highly natural text-to-speech model with 82 million parameters and over 11 million downloads, representing a significant advancement in AI voice generation.

0 favorites 0 likes
#tts

Which is the better local mobile TTS: Kokoro or Supertonic?

Reddit r/LocalLLaMA · 2026-06-14

Compares two locally running mobile TTS models, Kokoro and Supertonic, questioning their production quality beyond initial demos.

0 favorites 0 likes
#tts

@Teknium: Had Hermes Agent with the Manim Video skill plus it's TTS tool create a video explaining Hermes' Agent.

X AI KOLs Following · 2026-06-14 Cached

Demonstrates Hermes Agent using its Manim Video skill and TTS tool to create a video explaining itself.

0 favorites 0 likes
#tts

@laowangbabababa: Shocked! Dr. Qi on Douyin sells a 500k digital human agent per day, and I built it in 2 minutes. Using the Pixelle-Video project, which already has 22k stars. It supports digital human lip-syncing, motion transfer, and image-to-video. Supports ComfyUI, input a topic, from script writing to adding...

X AI KOLs Timeline · 2026-06-13 Cached

Introducing the open-source project Pixelle-Video: a fully automated AI short video engine. Input a topic and it automatically generates a video with script, images, voiceover, and background music. Supports local and cloud models, modular design allows flexible replacement of each component model.

0 favorites 0 likes
#tts

@Gorden_Sun: ZONOS2: Open-source MoE TTS model. 8B total parameters, 0.9B activated parameters. Supports multilingual, voice cloning, Chinese, and Chinese results are good. Model:

X AI KOLs Timeline · 2026-06-13 Cached

Zyphra released ZONOS2, an open-source MoE text-to-speech model trained on over 6 million hours of multilingual speech, supporting voice cloning and high-quality synthesis across many languages.

0 favorites 0 likes
#tts

@seclink: Xiaomi is on a strong growth trajectory! MiMo v2.5-ASR Released on 2026-06-02

X AI KOLs Following · 2026-06-12 Cached

Xiaomi has released updates to its MiMo model series, including mimo-v2.5-asr (supporting multiple dialects and lyric transcription), mimo-v2.5-pro (trillion parameters, 1M context), mimo-v2.5 (full-modal perception), and a TTS series, significantly improving agent performance and recognition capability in complex acoustic scenarios.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback