any one else finds Mimo v2.5 better than deepseek v4 flash!?

Reddit r/LocalLLaMA Models

Summary

A user reports that Mimo v2.5 outperforms DeepSeek v4 Flash in coding tasks based on benchmarks like Codex, Oh My Pi, Hermes, and Terminal Bench v2.0, though both models are similar overall.

I noticed while using both, mimo was often better, after benchmarking mimo v2.5 via open code endpoint in diff harness like codex, oh my pi, hermes. i found that mimo is indeed better in coding tasks. and over all, hermes scored 55% with mimo v2.5 via terminal bench v2.0 others did under 50% too with any harness from my list or deepseek v4 flash Not that i dont like deep seek v4 flash its GOAT, i have used it more. but as per benchmark both models are same at most places but when u run real life complex problems solving mimo v2.5 seemed to me helping me more i tested hy3 preview too. idk to me it felt like benchmark trained. needs to try more , but scores were pretty low for me in terminal bench EDIT: also while comparing oh my pi vs hermes vs codex cli. found hermes better for some reason. (offc for low lvl models only in my casestudy)
Original Article

Similar Articles

deepseek-ai/DeepSeek-V4-Flash

Hugging Face Models Trending

DeepSeek releases DeepSeek-V4-Flash and DeepSeek-V4-Pro, new MoE language models supporting 1 million token contexts with improved efficiency and performance.