Tag
Microsoft is testing a new native real-time voice model, MAI Realtime, in early access on its MAI Playground. The full-duplex system supports multiple languages, low latency, and configurable turn-taking, positioning it as a competitor to OpenAI's GPT Live and Sesame.
Thom Wolf and Cerebras released a fully open-source realtime voice demo with models and code, showcasing state-of-the-art speech-to-speech capabilities.
A claim that the Gemma-4-31B model running on Cerebras hardware outperforms ChatGPT's voice mode, demonstrated via a Hugging Face Space for real-time voice interaction.
We gave a Reachy Mini robot a real-time voice brain using GPT Realtime, allowing it to hear, see, talk, and physically react via motion tools. The project is open-source on GitHub.
Reachy Mini has a new fully open-source backend for real-time voice interaction, running audio models locally and leveraging LLM subscriptions to avoid per-second API costs.
LiveKit Agents is an open-source framework for building realtime, multimodal voice agents that can see, hear, and understand, with flexible STT/LLM/TTS integrations, job scheduling, telephony support, MCP compatibility, and a built-in test framework.