My voice agent sounded smart until one phone number was transcribed wrong.
Summary
The article argues that voice agent STT should be evaluated on entity accuracy (e.g., phone numbers, dates) rather than general word error rate, because missing critical fields can break workflows. It mentions testing with HubSpot fields and notes Smallest AI Pulse as an interesting tool for capturing workflow-critical entities in real time.
Similar Articles
Best STT API for voice agents? I’d test latency before accuracy
The author argues that for live voice agents, STT latency and real-time behavior are more critical than raw transcription accuracy, and proposes a different evaluation scorecard.
I analysed 10000 voice ai call and 40% of them had similar problems
An analysis of 10,000 voice AI calls reveals key failure points: STT error rates, first-8-second chaos, interruption handling, extended silence, tool call latency, LLM failures, and broken escalation — offering practical insights for agent builders.
What STT API are you using for production voice agents, and what broke first?
A developer asks what STT APIs people use in production voice agents, comparing Deepgram, AssemblyAI, and Smallest AI Pulse, and highlighting common failure points like endpointing, latency, and barge-in.
Before choosing an STT API, rank which transcript mistakes would actually hurt users.
A guide on evaluating speech-to-text APIs by ranking transcript mistakes based on their actual impact on users, rather than raw accuracy metrics.
@svpino: Why do so many AI-powered phone agents sound smart until you interrupt them? Even when these agents give you reasonable…
Deepgram released Flux TTS, a streaming conversation-native text-to-speech model that retains tone, pacing, and context across turns, handles interruptions, and runs with latency as low as 80ms to make voice AI feel more natural.