Tag
Whispera 是一个基于 VoxCPM 的 Windows 本地实时语音助手,集成了 SenseVoice ASR、llama-server 本地 LLM 推理、VoxCPM 流式 TTS 和 Mem0 长期记忆,完全离线运行。
Maya-2-Native has achieved the #2 position on Voice Arena's Hindi TTS leaderboard, surpassed only by Gemini 3.1 Flash.
A CPU TTS benchmark compares Kokoro, Supertonic, Inflect-Nano, and Kyutai's Pocket TTS using UTMOS MOS scores, revealing interesting findings about RTF scaling, UTMOS limitations with small vocoders, and undocumented output caps. Pocket TTS offers unique zero-shot voice cloning on CPU.
A guide demystifying voice agents using speech-to-text and text-to-speech, with four browser-based demo agents and instructions to build your own using RAG and tools.
A free, fully local voice cloning and TTS tool powered by Qwen3-TTS and Kokoro runs on Apple Silicon via MLX, enabling studio-quality audiobook generation from PDFs without cloud services.
VibeVoice 1.5B, a long-form multi-speaker TTS model, is now supported in audio.cpp, a native C++/ggml runtime, achieving 4.08x real-time speed on RTX 5090, 2.86x faster than Python baseline without quantization.
The developer improved qwen3-tts.cpp to run 5x realtime on RTX 5080 and created a cross-platform desktop GUI with Kotlin Compose Multiplatform, featuring voice cloning, streaming, and speaker embedding management.
Inflect-Micro-v2 is a compact text-to-speech model with under 10M parameters, supporting CPU/CUDA inference and long-text handling, released on Hugging Face.
A fully free and open-source HTTP API service based on macOS native capabilities, offering image OCR, multi-language translation, web content retrieval, face recognition, QR/barcode recognition, and text-to-speech.
MOSS-TTS-Local Transformer v1.5 is an open-source 48 kHz stereo TTS model with zero-shot voice cloning, native streaming, and support for 31 languages, built on a Qwen3-4B backbone and served via SGLang-Omni.
MosiAI has released MOSS-TTS Local Transformer v1.5, a text-to-speech model that supports voice cloning, over 30 languages, and high-quality 48 kHz output.
Inflect-Nano, an ultra-extreme tiny 4.63 million parameter text-to-speech model, has been released.
A developer debunks the common belief that LLM latency is the primary cause of slow voice agents, explaining that delays often stem from earlier stages like audio capture, VAD, and STT. They recommend logging specific latency metrics and testing various STT/TTS providers and orchestration frameworks to diagnose issues.
A curated, open-source learning path for building voice agents, covering from STT to production, with 190+ resources and a 5-week plan.
Kokoro-82M is a highly natural text-to-speech model with 82 million parameters and over 11 million downloads, representing a significant advancement in AI voice generation.
Compares two locally running mobile TTS models, Kokoro and Supertonic, questioning their production quality beyond initial demos.
Demonstrates Hermes Agent using its Manim Video skill and TTS tool to create a video explaining itself.
Introducing the open-source project Pixelle-Video: a fully automated AI short video engine. Input a topic and it automatically generates a video with script, images, voiceover, and background music. Supports local and cloud models, modular design allows flexible replacement of each component model.
Zyphra released ZONOS2, an open-source MoE text-to-speech model trained on over 6 million hours of multilingual speech, supporting voice cloning and high-quality synthesis across many languages.
Xiaomi has released updates to its MiMo model series, including mimo-v2.5-asr (supporting multiple dialects and lyric transcription), mimo-v2.5-pro (trillion parameters, 1M context), mimo-v2.5 (full-modal perception), and a TTS series, significantly improving agent performance and recognition capability in complex acoustic scenarios.