Xiaomi quietly uploaded MiMo-V2.5-DFlash — official DFlash weights are now on Hugging Face

Reddit r/LocalLLaMA Models

Summary

Xiaomi has quietly released MiMo-V2.5-DFlash, a 300B-parameter model on Hugging Face, with DFlash potentially doubling inference speed.

https://huggingface.co/XiaomiMiMo/MiMo-V2.5-DFlash Xiaomi appears to have quietly uploaded MiMo-V2.5-DFlash to Hugging Face: there is dedicated dflash directory containing the Dflash model, anyone willing to GGUF it and try? I'd do it but I can't today. This model is pretty good IMO (300B + params) and runs at about 8-10 tk/s on 2x24gb cards + vram offload (96/128gb drr5), dflash could double that speed and make it very interesting. EDIT: the main reason it's interesting, is because the MTP head was shared already, but doesn't work yet il llama cpp. I speculate (pun intended) the Dflash does work instead. EDIT2: very cool! they shared also the SEPARATE MTP model. the reason Llama doesn't work already is because it has trouble identifying the MTP layers. a separate MTP model might work too.
Original Article

Similar Articles

XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash

Hugging Face Models Trending

XiaomiMiMo releases MiMo-V2.5-Pro-FP4-DFlash, an FP4-quantized MoE model with block-diffusion speculative decoding to reduce memory and bandwidth for trillion-parameter inference.

XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

Reddit r/LocalLLaMA

MiMo-V2.6-Flash-RL is a multimodal AI model that scales reinforcement learning for self-improvement, featuring a sparse mixture-of-experts architecture with 309B total parameters and 1M token context length.

Xiaomi MiMo v2.6

Hacker News Top

Xiaomi has released version 2.6 of its MiMo model, suggesting an incremental update to its AI capabilities.