MAI (Microsoft AI) is very far behind on coding

Reddit r/ArtificialInteligence News

Summary

The article criticizes Microsoft AI's lackluster coding model performance compared to rivals like Kimi K3 and Deepseek V4, suggesting MAI is far behind despite vast resources.

Kimi K3 basically matches US frontier labs, Deepseek V4 is ~90% of frontier intelligence at ~5% of the cost, yet MAI team (one of the most well-resourced AI teams in the world) won't submit MAI-Thinking-1 or MAI-Code-Flash to Artificial Analysis for benchmarking, which is a telling sign of how far behind they are. I understand that MAI was first focused on lowering COGS for MS teams transcripts / image generation for Copilot (their audio and image models are at the frontier and super cost-effective, see them on Artificial Analysis), but being this far behind on coding and general intelligence is quite pathetic given their resources. Not sure what Satya is thinking. MSFT stock is likely stuck until they can put out a model that benchmarks well
Original Article

Similar Articles

MAI-Thinking-1

Hacker News Top

Microsoft AI introduces MAI-Thinking-1, a 35B-active parameter reasoning model trained from scratch without distillation, achieving strong performance on software engineering and math benchmarks while emphasizing clean data and self-sufficiency.

Microsoft's new MAI models

Simon Willison's Blog

Microsoft announced two new LLMs: MAI-Thinking-1 (35B reasoning model) and MAI-Code-1-Flash (5B code model), both trained on enterprise-grade, clean data without third-party distillation, with MAI-Thinking-1 claimed to be preferred over Sonnet 4.6 in blind evaluations.

Are AI coding agents hitting a wall, or are we just measuring them wrong?

Reddit r/AI_Agents

This article examines the gap between hype and reality for AI coding agents, arguing that they are effective for accelerating workflow parts but still require human oversight for architecture, debugging, and review, and questioning whether current benchmarks measure the right things.