Tag
A comparison of AI voice assistants ChatGPT-Live, Pi, Lucy OS1, and Gemini-Live focusing on which feels most natural to talk with, concluding that conversational quality is becoming the key differentiator as intelligence improves.
An analysis of half-duplex vs full-duplex architecture in AI voice models, discussing key features like overlap, backchannels, and barge-in that make voice agents sound robotic.
Google's Gemini Spark and Apple's Gemini-powered Siri 2.0 are launching in the next two weeks, representing major attempts to bring AI agents to mainstream consumers with billions of potential users.
Elba showcases its unified AI architecture that bridges voice and text channels in real-time, contrasting its single-agent system with competitors' fragmented 'three bots in a trenchcoat' approaches.
MoshiRAG combines a compact full-duplex speech language model with asynchronous retrieval-augmented generation to improve factuality while maintaining real-time interactivity. The approach leverages natural temporal gaps in conversation to retrieve external knowledge without disrupting the natural flow of dialogue.