@simplifyinAI: Most voice assistants make you wait your turn, you talk, then it talks, back and forth. NVIDIA just built one that does…
Summary
NVIDIA has launched Nemotron 3 VoiceChat, an AI model enabling real-time, full-duplex voice interactions that allow natural interruptions, available via a live demo as part of its open-source NeMo Speech framework.
View Cached Full Text
Cached at: 08/22/26, 07:21 AM
Most voice assistants make you wait your turn, you talk, then it talks, back and forth. NVIDIA just built one that doesn’t work that way, and you can try it right now.
It’s called Nemotron 3 VoiceChat, one of the models out of NVIDIA’s open-source NeMo Speech framework. It listens and responds in real time, full-duplex, so you can actually interrupt it mid-sentence like a real conversation.
the framework behind it needs GPU and CUDA to run yourself. But the model has a live demo anyone can just click and try.
Similar Articles
nvidia/NVIDIA-NemotronLabs-VoiceChat-11B · Hugging Face (full duplex)
NVIDIA released NemotronLabs VoiceChat 11B, an open end-to-end full-duplex speech model enabling real-time conversational AI with ~450ms turn-taking latency, barge-in, and live tool calling, the first open full-duplex model to support tool calling.
@oliviscusAI: NVIDIA just removed the biggest friction point in voice AI. They open-sourced PersonaPlex 7B, a real-time conversationa…
NVIDIA open-sourced PersonaPlex 7B, a real-time conversational model that listens and speaks simultaneously, handling natural interruptions and overlaps unlike most voice models.
OpenAI's New Voice Models Want to Do More Than Talk Back
OpenAI has launched three new real-time audio models to enable continuous, multitasking voice interactions that prioritize long-context reasoning, live translation, and seamless tool use.
@kwindla: https://x.com/kwindla/status/2062544580105359686
NVIDIA released Nemotron 3.5 ASR, an open-source multilingual speech-to-text model with the lowest latency tested, available in multilingual and English-only variants, ideal for voice agents and self-hosted deployments.
NVIDIA Launches Nemotron 3 Nano Omni Model, Unifying Vision, Audio and Language for up to 9x More Efficient AI Agents
NVIDIA announces Nemotron 3 Nano Omni, an open multimodal model that unifies vision, audio, and language processing to enable faster and more efficient AI agents, achieving up to 9x higher throughput compared to other open omni models.