@FinanceYF5: Voice AI is starting to make its way into phones. Open-source Audio8, which packages speech recognition and speech synt…
Summary
The post introduces the open-source Audio8 models, which enable on-device speech recognition and synthesis on phones and PCs, including an offline transcription version for iPhone.
View Cached Full Text
Cached at: 09/16/26, 08:11 PM
Voice AI is starting to make its way into phones. Open-source Audio8, which packages speech recognition and speech synthesis into a set of models that can run locally on phones and PCs. It also includes an offline transcription version for iPhone. Beyond the cloud, on-device capabilities are becoming increasingly usable. @LeonaYangAGI Well done
Leona Yang (@LeonaYangAGI): Over the past two months we open-sourced a set of on-device audio models. The name is Audio8. The series can run on phones, PCs, and other hardware with limited compute and memory. You can use them for local inference right away.
What’s included: ASR: 0.1B, 0.3B, 0.6B, 3B TTS:
Similar Articles
@FinanceYF5: 1/ Voice Agent can finally listen and speak simultaneously just like a real person. GPT-Live-1's API has recently been …
GPT-Live-1's API allows developers to integrate voice agents that can listen and speak simultaneously, providing a natural and fluent conversational experience for applications.
Labs AI
Labs AI is an iPhone app that turns text into natural AI voiceovers.
@thekuchh: voice mode on your phone is the real unlock leave the mac open go get coffee keep working the whole walk in essence: ta…
The article highlights the benefits of voice mode on phones for seamless AI assistant access and announces Vellum's free, open-source availability on iOS and Android with features like conversational voice mode.
Introducing next-generation audio models in the API
OpenAI introduced next-generation audio models for the API, including improved speech-to-text (gpt-4o-transcribe, gpt-4o-mini-transcribe) and customizable text-to-speech models that enable developers to build more intelligent and expressive voice agents with enhanced accuracy across challenging scenarios.
OpenAI's New Voice Models Want to Do More Than Talk Back
OpenAI has launched three new real-time audio models to enable continuous, multitasking voice interactions that prioritize long-context reasoning, live translation, and seamless tool use.