@garrytan: Everyone's bottleneck in voice AI is the same: retrieval. The agent thinks, network round-trips to a vector DB, and the…
Summary
Garry Tan highlights that retrieval is the key bottleneck in voice AI and introduces Moss, an open-source tool achieving sub-10ms vector search, alongside a hackathon at YC office on June 6-7.
View Cached Full Text
Cached at: 05/31/26, 04:53 PM
Everyone’s bottleneck in voice AI is the same: retrieval. The agent thinks, network round-trips to a vector DB, and the magic dies.
Moss runs search at sub-10ms (no hop). Open source. This is the layer voice agents were missing. Build on it June 6-7 at the YC office.
Pete Koomen (@koomen): Come build agents that can finally hold a fluid conversation at the 24-Hour Conversational AI Hackathon, hosted by @usemoss at the YC Office, June 6-7. First place wins an interview with a YC partner:
Similar Articles
@svpino: Why do so many AI-powered phone agents sound smart until you interrupt them? Even when these agents give you reasonable…
Deepgram released Flux TTS, a streaming conversation-native text-to-speech model that retains tone, pacing, and context across turns, handles interruptions, and runs with latency as low as 80ms to make voice AI feel more natural.
@garrytan: I just made some new GBrain evals that help prove that my retrieval-for-AI-agent open source layer is SOTA for reading …
Garry Tan has released new evaluations for gbrain, an open-source retrieval layer for AI agents, showcasing state-of-the-art performance in memory reading and writing without LLM-in-loop retrieval.
Building an AI voice agent from scratch: the parts that actually took our time
The author shares a postmortem on building a production phone-based AI voice agent, revealing that most engineering time was consumed by telephony infrastructure, turn detection, observability, and failure handling rather than core LLM behavior. They suggest using managed platforms like Vapi, Retell, or Dasha from the start to focus engineering effort on business logic.
@MaxForAI: If you are working on voice agents, you should try this project. A team from NTU, NUS, and Shanghai AI Lab released: Mega-ASR. This fully open-source ASR is built on Qwen3-ASR, aiming to break the long-standing bottleneck of ASR performance in noisy, reverberant, or other impaired real-world environments...
NTU, NUS, and Shanghai AI Lab jointly released Mega-ASR, a fully open-source ASR model built on Qwen3-ASR. Using the Voices-in-the-Wild-2M dataset and progressive acoustic-to-semantic optimization, it achieves up to 30% relative Word Error Rate (WER) reduction in real-world noisy environments. With only 1.7B parameters, it enables efficient inference on consumer-grade hardware.
OPEN AI: How we built a realtime system for responsive voice AI in six months
OpenAI describes how they built GPT-Live, a full-duplex realtime voice AI system that eliminates the turn detector, enabling natural continuous conversation. The article details architecture improvements in inference, context management, and media transport over six months.