Red-teaming voice agents: audio as the attack surface, multi-turn pressure, and closing the loop
Summary
A deep dive into red-teaming voice agents, highlighting audio as an attack surface, the need for multi-turn testing, and practical baseline methodologies (1,200 calls) for pre-launch safety.
Similar Articles
Building voice AI agents that take turns like humans — the gotchas nobody warns you about
This article shares hard-won lessons from building real-time voice AI agents, highlighting the importance of proper turn-taking, VAD handling, billing awareness, and avoiding echo loops.
Building an AI voice agent from scratch: the parts that actually took our time
The author shares a postmortem on building a production phone-based AI voice agent, revealing that most engineering time was consumed by telephony infrastructure, turn detection, observability, and failure handling rather than core LLM behavior. They suggest using managed platforms like Vapi, Retell, or Dasha from the start to focus engineering effort on business logic.
Why voice and messaging channels beat desktop dashboards for running autonomous agents
The article argues that voice and messaging channels are more effective than desktop dashboards for operating autonomous agents, describing practical solutions like duplex audio and shared chat dynamics implemented for Mentat on prompt2bot.
8 months running a voice agent in production: what broke, what fixed it, and the system prompt I use
A practitioner shares 8 months of experience running a voice agent for a law firm, detailing challenges like latency, turn-taking, and post-call workflows, and provides a working system prompt.
Expanding on how Voice Engine works and our safety research
OpenAI details the development history and safety approach for Voice Engine, from internal testing in 2022 through various limited deployments including ChatGPT Voice Mode and TTS API, emphasizing careful rollout with professional voice actors and ongoing collaboration with policymakers to address synthetic voice risks.