We gave a Reachy Mini a real-time voice brain

Reddit r/LocalLLaMA Tools

Summary

We gave a Reachy Mini robot a real-time voice brain using GPT Realtime, allowing it to hear, see, talk, and physically react via motion tools. The project is open-source on GitHub.

We attended an event the other day and found this little guy lying on our desk, a Reachy Mini from Hugging Face. It belongs to the daughter of the event organizer. We got curious about how it worked, and an hour later we'd given it a brain. The model basically becomes Reachy. It hears through its mic, sees through its camera, talks through its speaker, and calls motion tools to physically react while it talks. Repo: [https://github.com/opper-ai/reachy-voice-realtime](https://github.com/opper-ai/reachy-voice-realtime) Key things: * Web UI to watch the camera feed, transcript, and tool calls live. * 19 motion and perception tools the model calls mid-conversation (emotes, head/antenna/body movement, camera, sound direction). * Mimics you, wave and it waves back, nod and it nods, tilt your head and it tilts. * Runs on GPT Realtime 2, routed through Opper so the model is a one-line swap. * The realtime client and tool layer are separate, so you can also wire it straight to a provider or a local/OS realtime model. Setup's in the README (Python 3.12+), MIT licensed. We handed it back to his daugther so now she can finally talk to her robot.
Original Article

Similar Articles

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

Hugging Face Blog

Hugging Face and Cerebras demonstrate a real-time speech-to-speech pipeline combining open-source models (Nvidia's Parakeet, Gemma 4, Qwen3TTS) with Cerebras' fast inference, enabling natural conversational AI and powering robots like Reachy Mini.

Introducing gpt-realtime and Realtime API updates

OpenAI Blog

OpenAI is making the Realtime API generally available with a new advanced speech-to-speech model called gpt-realtime, featuring improved instruction following, tool calling, and natural speech quality. New capabilities include MCP server support, image inputs, SIP phone calling, and two new voices (Cedar and Marin).