Tag
ElevenLabs MCP in Claude allows users to create and manage voice agents within chat systems, integrating ElevenLabs' voice technology with Claude's AI capabilities.
LangChain is hosting a webinar on how to evaluate voice agents across execution, outcomes, and experience, covering tool use, task completion, latency, interruptions, and conversational friction.
NVIDIA announces Magpie Multilingual TTS, an open-weights text-to-speech model supporting 12 languages with low-latency deployment via NVIDIA NIM for building production voice agents.
Discusses how voice agents lose paralinguistic signals like tone, hesitation, and speaker identity when transcribing to text, and questions whether and how these features are captured and used downstream.
A developer asks what STT APIs people use in production voice agents, comparing Deepgram, AssemblyAI, and Smallest AI Pulse, and highlighting common failure points like endpointing, latency, and barge-in.
LiveKit Agents is an open-source real-time voice agent framework supporting WebRTC, telephony integration, semantic turn detection, MCP tool calling, and multi-agent handoff, helping developers quickly build real-time voice applications such as AI customer service and phone bots.
Google Devs shares how to use Google ADK, Gemini Live, and LangSmith to build voice agents with full traceability of interactions, including tool calls and token costs.
The article discusses the challenge of separating probabilistic language understanding from deterministic execution in voice agents, recommending a confidence check for conditional phrases before triggering structured actions, as implemented with Vomo AI.
Encore AI has raised $30M to develop AI voice agents that learn from customer interactions by analyzing call recordings and transcripts.
An open source profiler designed for voice agents to provide insights into internal operations and performance.
A deep dive into red-teaming voice agents, highlighting audio as an attack surface, the need for multi-turn testing, and practical baseline methodologies (1,200 calls) for pre-launch safety.
This article discusses the challenges of scaling voice agents, noting that failures occur at different layers, and identifies the most common bottleneck that limits performance first.
LangChain launches LangSmith tracing for voice frameworks (Pipecat, LiveKit, OpenAI Realtime, Gemini Live), enabling full audio monitoring, STT/TTS latency tracking, interruption detection, and VAD analysis with minimal code.
A curated list of text-to-speech resources for voice agent builders, organized around the decision between real-time streaming synthesis and offline high-fidelity synthesis, with emphasis on streaming latency and mid-stream cancellation.
Cars24 uses OpenAI's technology to build AI-powered voice and chat agents that handle over a million conversation minutes per month, automating the full customer journey for buying, selling, and financing cars.
Cekura introduces a self-improvement loop for voice agents, enabling continuous enhancement of conversational AI.
The article identifies a structural flaw in voice agents where they cannot detect when they over-promise across multiple conversation turns, and describes building a deterministic checker that flags contradictions without relying on LLM evaluation.
This paper evaluates the reliability of Gemini models as audio judges for scoring full-duplex voice agent conversations, finding that Gemini 2.5 Flash shows strong agreement with human raters on most dimensions, though model swaps require re-validation.
Eric Schmidt advises starting an agentic AI company to capitalize on high demand and low supply of builders, suggesting building and selling AI voice agents for $10k/month.
Simba 3.2, claimed as the world's #1 voice model, now powers new voice agents.