KAT Coder 2.5 dev: Do yourself a favor and try it!

Reddit r/LocalLLaMA Models

Summary

A developer enthusiastically recommends KAT Coder 2.5 dev, claiming it is faster, more accurate, and uses fewer tokens than Qwen 3.6 35b a3b, and outperforms Gemma 4 models on their setup, with a GitHub repo containing detailed benchmarks.

It is so good! I don't know why there aren't more people talking about it. Fewer tokens, faster and more accurate than Qwen 3.6 35b a3b. On my setup it's nearly as good as 27b, but 5x faster. And it completely trashes the Gemma 4 models. At least for my use case, it feels amazing. I'd love to hear other people's experience with it. If you want an actual measure of performance, I have a GitHub repo explaining how I tested it for my type of use case with a detailed performance comparison with other models . It has the quants I used, OpenCode and llama.cpp version along with all the flags for temp, top-p, top-k etc.; and if there's some detail missing please let me know. But really I think you should just download the model and try it out yourself, because we all have different use cases and those will always be more informative than benchmarks or one person's idiosyncratic experience.
Original Article

Similar Articles

Kwaipilot/KAT-Coder-V2.5-Dev

Hugging Face Models Trending

KAT-Coder-V2.5-Dev is an open-weight MoE coding model with 35B total parameters (3B active), achieving state-of-the-art results on agentic coding benchmarks through SFT and RL training.

Gemma 4 31B's competence surprised me

Reddit r/LocalLLaMA

A user shares anecdotal findings that Gemma 4 31B outperforms Qwen 3.6 models and matches Opus 4.7 in understanding and refactoring messy academic code, highlighting a benchmark (SciCode) where Gemma excels.