@rohanpaul_ai: Love this, another huge release from Meta. Lunched Muse Voice Transcribe for real-time voice dictation, with the lowest…

X AI KOLs Timeline Models

Summary

Meta has released Muse Voice Transcribe, a real-time speech-to-text model with a 3.1% word error rate and adaptive streaming capabilities, making it suitable for voice agents.

Love this, another huge release from Meta. Lunched Muse Voice Transcribe for real-time voice dictation, with the lowest 3.1% final-transcription word error rate with adaptive delay. which is significantly ahead of competing models. - The big deal is that Muse does something beyond just transcribing speech; it learns when to wait, when to commit a word, when a speaker changes, and when a turn is actually over, all inside the same streaming model. That makes it much closer to a real-time perception layer for voice agents than a conventional speech-to-text API. - Muse processes audio in 80ms chunks and chooses after each chunk whether to emit text or keep listening. Meta made it available through Meta Model API, Meta AI for Mac, and Muse Code
Original Article
View Cached Full Text

Cached at: 09/01/26, 05:49 PM

Love this, another huge release from Meta.

Lunched Muse Voice Transcribe for real-time voice dictation, with the lowest 3.1% final-transcription word error rate with adaptive delay. which is significantly ahead of competing models.

  • The big deal is that Muse does something beyond just transcribing speech; it learns when to wait, when to commit a word, when a speaker changes, and when a turn is actually over, all inside the same streaming model. That makes it much closer to a real-time perception layer for voice agents than a conventional speech-to-text API.

  • Muse processes audio in 80ms chunks and chooses after each chunk whether to emit text or keep listening.

Meta made it available through Meta Model API, Meta AI for Mac, and Muse Code

Mark Zuckerberg (@finkd): Muse Voice Transcribe is MSL’s first real-time audio perception model – rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.

Similar Articles

Meta launches Muse Image across its apps (3 minute read)

TLDR AI

Meta has launched Muse Image, its first image-generation model from Meta Superintelligence Labs, now available in Meta AI across its apps. The model supports advanced reasoning, multi-reference composition, and is ranked No.2 on Arena, with plans for Muse Video and Content Seal watermarking.