Microsoft MAI-Voice-2
Summary
Microsoft has released MAI-Voice-2, an expressive text-to-speech system supporting voice cloning in 15 languages.
Similar Articles
Microsoft tests new MAI Realtime voice model (2 minute read)
Microsoft is testing a new native real-time voice model, MAI Realtime, in early access on its MAI Playground. The full-duplex system supports multiple languages, low latency, and configurable turn-taking, positioning it as a competitor to OpenAI's GPT Live and Sesame.
Microsoft's New MAI-Image and MAI-Voice (2 minute read)
Microsoft announces public preview of MAI-Image-2.5-Pro and MAI-Voice-2-Flash, their latest purpose-built generative AI models for image and voice, now available on Azure AI and powering Microsoft products like Bing Image Creator.
k2-fsa/OmniVoice
OmniVoice is a massively multilingual zero-shot text-to-speech model supporting over 600 languages, built on a diffusion language model architecture with fast inference and voice cloning capabilities.
@MosiAI_Official: MOSS-TTS Local Transformer v1.5 is here. Clone any voice. Speak any language. Hear every detail. 30+ languages, 48 kHz …
MosiAI has released MOSS-TTS Local Transformer v1.5, a text-to-speech model that supports voice cloning, over 30 languages, and high-quality 48 kHz output.
Voiser AI
Voiser AI offers human-like AI voiceovers in over 140 languages.