@Thom_Wolf: Most people should probably update their priors on the state of open-source speech-to-speech. It's honestly kind of min…
Summary
Thom Wolf and Cerebras released a fully open-source realtime voice demo with models and code, showcasing state-of-the-art speech-to-speech capabilities.
View Cached Full Text
Cached at: 07/03/26, 08:33 AM
Most people should probably update their priors on the state of open-source speech-to-speech.
It’s honestly kind of mind-blowing.
We teamed up with @cerebras to build a fully open-source realtime voice demo (models + code) to show what’s possible today.
Demo : https://huggingface.co/spaces/smolagents/hf-realtime-voice…
Blog: https://huggingface.co/blog/cerebras-gemma4-voice-ai…
Go test it, fork it, tweak it, and impress your friends.
video is raw, no cut, no speed-up, first take
HF Realtime Voice - a Hugging Face Space by smolagents
Source: https://huggingface.co/spaces/smolagents/hf-realtime-voice Fetching metadata from the HF Docker repository...
Similar Articles
@LinusEkenstam: This is life-changing tech. We need less hype, and more stuff like this. Imagine how many people this can help
Bland launches Speech v3, claiming it's the world's first Human Speech Engine and top model in Design Arena's Audio Realism benchmark, surpassing ElevenLabs, Grok, Cartesia, and OpenAI.
@kwindla: OpenAI shipped a new speech-to-speech model today: gpt-realtime-2 This is the first speech-to-speech model good enough …
OpenAI has released gpt-realtime-2, a new speech-to-speech model optimized for real-time voice agent interactions with low-latency tool calling.
@vvolhejn: Our open-source TTS just got even open-sourcer
Kyutai Labs has open-sourced their Pocket TTS training stack, including data pipeline, recipes, and evaluations, allowing developers to train text-to-speech models on GPUs and run them on CPUs.
@oliviscusAI: NVIDIA just removed the biggest friction point in voice AI. They open-sourced PersonaPlex 7B, a real-time conversationa…
NVIDIA open-sourced PersonaPlex 7B, a real-time conversational model that listens and speaks simultaneously, handling natural interruptions and overlaps unlike most voice models.
@kwindla: https://x.com/kwindla/status/2062544580105359686
NVIDIA released Nemotron 3.5 ASR, an open-source multilingual speech-to-text model with the lowest latency tested, available in multilingual and English-only variants, ideal for voice agents and self-hosted deployments.