Tag
Fish Audio S2 is a new open-weight audio model available on Hugging Face, offering two models for timing and acoustic details, with fast inference and a hosted version S2.1 Pro supporting 83 languages at lower cost than ElevenLabs.
Dia is a 1.6B-parameter open-source text-to-speech model that generates English dialogue from transcripts, supporting two-speaker generation, audio conditioning, and nonverbal cues.
Released a model (Stereo2Spatial) that converts stereo music tracks to spatialized binaural mixes, using flow-matching diffusion and amplitude lifting for stable training. The model and a Windows app are open-sourced under Apache 2.0.
OpenAI is rolling out a new bidirectional voice model (Bidi 1) for ChatGPT that allows simultaneous speaking, hearing, and listening, real-time translation, and improved conversation context handling. The upgrade is appearing in the web interface and app for some users, with a broader release expected soon.
OpenAI is preparing to release GPT-Bidi-1, a next-generation voice model for ChatGPT that supports bidirectional communication, interruptions, and mid-sentence adjustments, aiming to close the gap between voice and text capabilities.
Gemini 3.5 Live Translate is a new audio model for real-time speech-to-speech translation.
Google releases Gemini 3.5 Live Translate, an audio model for near real-time speech-to-speech translation in over 70 languages, preserving speaker intonation and pacing. It is rolling out across Google products including the Gemini Live API, Google Meet, and Google Translate.
Google has released Gemini 3.1 Flash Live, a new high-quality audio model designed for more natural and reliable real-time voice interactions with improved latency and reasoning capabilities.