Tag
An open-source gateway, sarvam-bridge, lets existing ElevenLabs/OpenAI/Deepgram apps switch to Sarvam AI by changing only the base URL, handling Indic language quirks like chunking, audio reassembly, and language codes. The author details technical decisions, stress testing, and a crash bug fix.
Athens-based Omilia raises a $67M Series B led by Expedition Growth Capital to scale its AI-powered customer support platform, which combines generative AI with more traditional automation. The company reports $60M ARR and counts clients like Capital One and Taco Bell.
An analysis of 10,000 voice AI calls reveals key failure points: STT error rates, first-8-second chaos, interruption handling, extended silence, tool call latency, LLM failures, and broken escalation — offering practical insights for agent builders.
Bland launches Speech v3, claiming it's the world's first Human Speech Engine and top model in Design Arena's Audio Realism benchmark, surpassing ElevenLabs, Grok, Cartesia, and OpenAI.
Eight months after receiving a flawed AI cold call, the author's own AI voice agent naturally handled a call from another AI agent — a real-world buyer's agent asking about specific NVMe drives. The incident highlights rapid progress in AI agent communication and raises legal questions about the EU AI Act's transparency rules for machine-to-machine calls.
NVIDIA open-sourced PersonaPlex 7B, a real-time conversational model that listens and speaks simultaneously, handling natural interruptions and overlaps unlike most voice models.
OpenAI announces GPT-Live, a new architecture and stack for realtime audio that enables listening while speaking, with continuous audio flow for deeper reasoning and tool use without interrupting conversation.
OpenAI describes how they built GPT-Live, a full-duplex realtime voice AI system that eliminates the turn detector, enabling natural continuous conversation. The article details architecture improvements in inference, context management, and media transport over six months.
Voice AI is quietly rolling out across US fast-food drive-thrus, with chains like Taco Bell, Dairy Queen, and White Castle scaling up automated ordering after years of testing.
After months of experimentation, the author successfully ordered a pizza using OpenClaw integrated with vapi.ai and Twilio, and is now hosting an AMA about the setup.
Simba 3.2 from SpeechifyAI claims the #1 spot on the blind-test voice leaderboard, surpassing ElevenLabs, OpenAI, and Google DeepMind, at a significantly lower cost with a new API and free tier.
Smallest.ai raises $13M in Series A funding to develop small, specialized voice models enabling real-time, human-like conversation for AI agents, aiming to make voice interactions indistinguishable from human speech.
TEN is a framework for building real-time multimodal conversational AI agents, offering configurable STT, LLM, and TTS components, a visual designer, and deployment options including self-hosting and split deployment.
The article discusses the challenge of separating probabilistic language understanding from deterministic execution in voice agents, recommending a confidence check for conditional phrases before triggering structured actions, as implemented with Vomo AI.
Telli raises $15M seed round led by redalpine to build AI agents for B2C customer operations, handling calls, lead qualification, appointments, and follow-ups. Already used by Sky, Enpal, and Vaillant.
The authors open-sourced their AI voice agent stack, unexpectedly receiving significant attention.
Explores the challenges bootstrapped Voice AI tools face when dealing with expensive partner API paywalls in the hospitality and restaurant sectors.
David Turturean solved a 40-year-old open problem in p-adic Galois theory using voice input, in collaboration with problem proposer David Roe, under EpochAI Research's FrontierMath initiative.
A practitioner recounts experiences with voice AI in call centers, detailing hidden costs when solutions underperform in production, and asks for honest feedback from others with similar real-world experience.
An analysis of the real costs associated with Voice AI calls, covering factors like API usage, latency, and provider pricing.