Xiaomi is now serving MiMo V2.5 at 1000-3000tps using DFlash & Persistent kernel. DFLash model is out, open-source release promised coming soon
Summary
Xiaomi has released MiMo V2.5 with DFlash and Persistent kernel, achieving 1000-3000 tps. The DFlash model is now available and open-source release is promised soon.
Similar Articles
Xiaomi quietly uploaded MiMo-V2.5-DFlash — official DFlash weights are now on Hugging Face
Xiaomi has quietly released MiMo-V2.5-DFlash, a 300B-parameter model on Hugging Face, with DFlash potentially doubling inference speed.
XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash
XiaomiMiMo releases MiMo-V2.5-Pro-FP4-DFlash, an FP4-quantized MoE model with block-diffusion speculative decoding to reduce memory and bandwidth for trillion-parameter inference.
Xiaomi just claimed 1,000+ tps on a 1T model using a standard 8-GPU server
Xiaomi released MiMo-V2.5-Pro-UltraSpeed in collaboration with TileRT, achieving over 1000 tokens/s decode speed on a 1-trillion-parameter model, enabling real-time AI interaction and accelerating coding agents and reasoning tasks.
Xiaomi released their SOTA model, MiMo-V2.5-Pro.
Xiaomi launched MiMo-V2.5-Pro, claiming state-of-the-art performance.
Tested Xiaomi's MiMo V2.5 Pro for autonomous coding: 301 commits, 60+ pages, $70 in API costs. Now it's open-source.
Xiaomi has open-sourced its MiMo V2.5 Pro model, a 1.02T parameter MoE model designed for autonomous coding tasks. The article details a real-world test showing high efficiency with low API costs due to high cache hit rates.