Xiaomi quietly uploaded MiMo-V2.5-DFlash — official DFlash weights are now on Hugging Face

Reddit r/LocalLLaMA Models

Summary

Xiaomi has quietly released MiMo-V2.5-DFlash, a 300B-parameter model on Hugging Face, with DFlash potentially doubling inference speed.

https://huggingface.co/XiaomiMiMo/MiMo-V2.5-DFlash Xiaomi appears to have quietly uploaded MiMo-V2.5-DFlash to Hugging Face: there is dedicated dflash directory containing the Dflash model, anyone willing to GGUF it and try? I'd do it but I can't today. This model is pretty good IMO (300B + params) and runs at about 8-10 tk/s on 2x24gb cards + vram offload (96/128gb drr5), dflash could double that speed and make it very interesting. EDIT: the main reason it's interesting, is because the MTP head was shared already, but doesn't work yet il llama cpp. I speculate (pun intended) the Dflash does work instead. EDIT2: very cool! they shared also the SEPARATE MTP model. the reason Llama doesn't work already is because it has trouble identifying the MTP layers. a separate MTP model might work too.
Original Article

Similar Articles

XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash

Hugging Face Models Trending

XiaomiMiMo releases MiMo-V2.5-Pro-FP4-DFlash, an FP4-quantized MoE model with block-diffusion speculative decoding to reduce memory and bandwidth for trillion-parameter inference.

XiaomiMiMo/MiMo-V2.5-Pro

Hugging Face Models Trending

Xiaomi releases MiMo-V2.5-Pro, an open-source MoE language model with 1.02T total parameters and 1M token context, optimized for complex agentic and software engineering tasks.