@AdinaYakup: Xiaomi @XiaomiMiMo just released 2 SoTA models One might be the new BEST open model yet Both are: - Sparse MoE + 1M con…

X AI KOLs Timeline Models

Summary

Xiaomi has released two state-of-the-art AI models, MiMo-V2.6 Pro RL for maximum capability and MiMo-V2.6 Flash RL for maximum efficiency, both featuring sparse mixture-of-experts architecture, 1M context length, MIT licensing, and native omni-modal support for text, image, video, and audio.

Xiaomi @XiaomiMiMo just released 2 SoTA models One might be the new BEST open model yet🔥 Both are: - Sparse MoE + 1M context + MIT licensed - Native omni: text/image/video/audio ✨ MiMo- V2.6 Pro RL: Maximum capability - 1.02T / 42B - Strong coding + agent performance: 71.9 DeepSWE / 53.1 AutomationBench / 89.9 Terminal Bench 2.1 ✨ MiMo-V2.6 Flash RL: Maximum efficiency - 309B / 15B - 15B active params while staying close to Pro on many agent benchmarks
Original Article
View Cached Full Text

Cached at: 09/21/26, 09:39 PM

Xiaomi @XiaomiMiMo just released 2 SoTA models One might be the new BEST open model yet🔥

Both are:

  • Sparse MoE + 1M context + MIT licensed
  • Native omni: text/image/video/audio

✨ MiMo- V2.6 Pro RL: Maximum capability

  • 1.02T / 42B
  • Strong coding + agent performance: 71.9 DeepSWE / 53.1 AutomationBench / 89.9 Terminal Bench 2.1

✨ MiMo-V2.6 Flash RL: Maximum efficiency

  • 309B / 15B
  • 15B active params while staying close to Pro on many agent benchmarks

Similar Articles

Xiaomi MiMo v2.6

Hacker News Top

Xiaomi has released version 2.6 of its MiMo model, suggesting an incremental update to its AI capabilities.

XiaomiMiMo/MiMo-V2.5-Pro

Hugging Face Models Trending

Xiaomi releases MiMo-V2.5-Pro, an open-source MoE language model with 1.02T total parameters and 1M token context, optimized for complex agentic and software engineering tasks.

XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

Reddit r/LocalLLaMA

MiMo-V2.6-Flash-RL is a multimodal AI model that scales reinforcement learning for self-improvement, featuring a sparse mixture-of-experts architecture with 309B total parameters and 1M token context length.