Tag
Outcome School 公开了一篇详细教程,讲解如何设计实时语音 AI 智能体,涵盖级联管道(STT→LLM→TTS)、端到端 Speech-to-Speech 模型与混合方案、延迟预算、打断处理(barge-in)、工具调用、记忆与上下文、电话系统集成、规模化部署、边缘情况、可观测性、安全与成本估算。
This paper introduces randomized guidance as an efficient training method for tandem speech-to-speech models, enabling them to learn natural conversational behavior directly from real conversations without simulating backend LLM behavior.
This paper proposes a causality-aware framework for LLM-based simultaneous speech-to-speech translation, introducing a novel data pipeline and adaptive policy to improve quality-latency trade-off, achieving state-of-the-art results with reduced latency.
Gemini 3.8 Live with Live Avatar brings real-time visual presence to conversational AI, allowing enterprises to create more natural and interactive digital experiences with dynamic avatars and multilingual support.
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, new speech-to-speech AI models, and the author built a web UI tool for voice conversations using these models.
Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, advanced AI models for real-time reasoning and voice agents, achieving top scores in various benchmarks.
The article reports on real-world testing of OpenAI's GPT-Live-1 model for phone agents, highlighting issues with instruction adherence, alphanumeric errors, and language handling in production scenarios.
xAI has launched Grok Voice Think Fast 2.0, an improved speech-to-speech model with higher transcription accuracy and lower latency, outperforming competitors on benchmarks.
Introduces a fully local AI girlfriend project made by a Bilibili developer, integrating four models: Silero VAD, Whisper, llama.cpp, and Qwen3-TTS, all packed into 15G VRAM with hot-swapping capability.
Google Gemma announces that developers can now use the Gemma 4 31B model as the brain for voice AI, enabled by Hugging Face and Cerebras for ultra-fast inference, as part of an open-source cascaded speech-to-speech stack.
NVIDIA released Nemotron, an audio-native model capable of transcription, translation, sound recognition, audio Q&A, TTS, and full speech-to-speech, with open weights in 2B and 30B sizes.
Hugging Face has open-sourced a modular speech-to-speech tool, which separates the four layers (VAD, speech recognition, LLM, and speech synthesis) and supports replacement, allowing users to flexibly swap components or run locally without rewriting the entire pipeline.
Thom Wolf and Cerebras released a fully open-source realtime voice demo with models and code, showcasing state-of-the-art speech-to-speech capabilities.
This paper proposes a reference-based evaluation protocol for assessing prosody and rhythm in speech-to-speech AI systems, using matched human conversation data to provide interpretable behavioral plausibility checks.
Hugging Face and Cerebras demonstrate a real-time speech-to-speech pipeline combining open-source models (Nvidia's Parakeet, Gemma 4, Qwen3TTS) with Cerebras' fast inference, enabling natural conversational AI and powering robots like Reachy Mini.
This paper presents a modular end-to-end speech-to-speech conversational system for the low-resource Algerian Dialect, integrating ASR, NLU, RAG, and TTS with dedicated datasets and fine-tuned models.
Discusses leveraging Gemma 4 12B's encoder-free architecture for native voice input, seeking out-of-the-box solutions for low-latency streaming audio ingestion.
Gemini 3.5 Live Translate is a new audio model for real-time speech-to-speech translation.
Google DeepMind announces Live Translate, a feature that converts speech into over 70 languages in real-time while preserving tone, pace, and pitch for more natural conversations.
Google releases Gemini 3.5 Live Translate, an audio model for near real-time speech-to-speech translation in over 70 languages, preserving speaker intonation and pacing. It is rolling out across Google products including the Gemini Live API, Google Meet, and Google Translate.