Microsoft tests new MAI Realtime voice model (2 minute read)

TLDR AI Models

Summary

Microsoft is testing a new native real-time voice model, MAI Realtime, in early access on its MAI Playground. The full-duplex system supports multiple languages, low latency, and configurable turn-taking, positioning it as a competitor to OpenAI's GPT Live and Sesame.

Microsoft's first native real-time voice model, MAI Realtime, has surfaced as a hidden early-access entry in the company's MAI Playground. The listing points to a bidirectional, full-duplex system that can listen and speak at the same time rather than trading turns. There are two voices available, both noticeably more natural than what Copilot's voice mode currently delivers. The model will likely be made available on Microsoft Foundry and Copilot voice, but no timeline is available.
Original Article
View Cached Full Text

Cached at: 08/03/26, 01:28 PM

# Exclusive: Microsoft tests new MAI Realtime voice model Source: [https://www.testingcatalog.com/exclusive-microsoft-tests-new-mai-realtime-voice-model/](https://www.testingcatalog.com/exclusive-microsoft-tests-new-mai-realtime-voice-model/) Microsoft appears to be preparing its first native real\-time voice model, referred to as MAI Realtime, which has surfaced as a hidden early\-access entry in the company’s MAI Playground\. The listing suggests a small group of partners already has hands\-on access, and what is visible points to a bidirectional, full\-duplex system, one that listens and speaks at the same time rather than trading turns, placing it in direct comparison with[OpenAI’s GPT Live 1](https://www.testingcatalog.com/openai-rolls-out-gpt-live-voice-for-chatgpt-on-web-and-mobile/)or[Sesame](https://www.testingcatalog.com/sesame-debutes-ios-app-in-preview-with-personal-voice-agents/)\. > \> It supports English, German, Spanish, French, Italian, Portuguese, Japanese, Korean, Chinese, Dutch, Hindi, Indonesian, Arabic, Russian, Turkish, Vietnamese, and Thai\.[pic\.twitter\.com/7BmcW2MBa0](https://t.co/7BmcW2MBa0?ref=testingcatalog.com) — 🚨 AI News \| TestingCatalog \(@testingcatalog\)[August 2, 2026](https://x.com/testingcatalog/status/2083928447450067196?ref_src=twsrc%5Etfw&ref=testingcatalog.com) Two voices are present so far, Victoria and Grant, both noticeably more natural than what Copilot’s voice mode currently delivers\. Language can be pinned explicitly or left on automatic detection, and the model switches languages mid\-conversation without losing its footing\. Turn\-taking is configurable through two listener options: a**Switchboard mode**built around an MAI\-Ears endpointer driven by inline control tokens, and a**deterministic setup**that pairs silence\-based endpointing with a Whisper semantic endpointer\. ![MAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/MAI-Playground-Microsoft-AI-08-02-2026_04_34_PM.jpg)![MAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/MAI-Playground-Microsoft-AI-08-02-2026_04_33_PM--1-.jpg)The practical difference between them is subtle in use, though interruptions are handled cleanly and response latency is low\. The model does not sing or produce non\-speech sounds, which keeps it squarely a conversational system rather than a general audio generator\. A debug panel exposes live latency figures, model thoughts and processing steps, and sample sharing looks set to arrive for playground users once access widens\. ![MAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/MAI-Playground-Microsoft-AI-08-02-2026_04_33_PM.jpg)That would fill a conspicuous gap\. Every[MAI speech model](https://www.testingcatalog.com/microsoft-previews-mai-image-2-5-pro-and-mai-voice-2-flash/)shipped so far runs in one direction: MAI\-Voice\-2 and its Flash variant for synthesis, and MAI\-Transcribe\-1\.5 for recognition, while the speech\-to\-speech layer in Azure Speech’s Voice Live API still relies on the GPT\-Realtime model\. A first\-party full\-duplex model would close that dependency for Mustafa Suleyman’s superintelligence team, which shipped seven in\-house models at Build 2026 and has been steadily swapping OpenAI components out of Copilot, Teams and Bing\.[Microsoft](https://www.testingcatalog.com/tag/microsoft/)Foundry is the likely developer destination, with Copilot voice the obvious consumer surface, though no timeline has been attached to either\.

Similar Articles

Microsoft's New MAI-Image and MAI-Voice (2 minute read)

TLDR AI

Microsoft announces public preview of MAI-Image-2.5-Pro and MAI-Voice-2-Flash, their latest purpose-built generative AI models for image and voice, now available on Azure AI and powering Microsoft products like Bing Image Creator.

Microsoft MAI-Voice-2

Product Hunt

Microsoft has released MAI-Voice-2, an expressive text-to-speech system supporting voice cloning in 15 languages.

Advancing voice intelligence with new models in the API

OpenAI Blog

OpenAI has announced three new voice models in its API: GPT-Realtime-2 with advanced reasoning, GPT-Realtime-Translate for live multilingual translation, and GPT-Realtime-Whisper for streaming transcription, aiming to enable more natural and action-oriented voice applications.

Introducing gpt-realtime and Realtime API updates

OpenAI Blog

OpenAI is making the Realtime API generally available with a new advanced speech-to-speech model called gpt-realtime, featuring improved instruction following, tool calling, and natural speech quality. New capabilities include MCP server support, image inputs, SIP phone calling, and two new voices (Cedar and Marin).