KAT Coder 2.5 dev: Do yourself a favor and try it!
Summary
A developer enthusiastically recommends KAT Coder 2.5 dev, claiming it is faster, more accurate, and uses fewer tokens than Qwen 3.6 35b a3b, and outperforms Gemma 4 models on their setup, with a GitHub repo containing detailed benchmarks.
Similar Articles
Kat Coder 2.5 is insane. Especially considering I ran it at Q4_K_M
Kat Coder 2.5 is a coding assistant model that delivers impressive performance even when run at Q4_K_M quantization.
35B-A3B tool calling benchmark: Original Qwen vs. KAT Coder, Ornith and Tiel-Coder
The article benchmarks fine-tuned Qwen3.6-35B-A3B models for tool calling capabilities, showing Ornith 1.5 and Tiel-Coder perform best, approaching scores of larger Qwen models. The study uses extensive GPU time and the tool-eval-bench utility for evaluation.
Kwaipilot/KAT-Coder-V2.5-Dev
KAT-Coder-V2.5-Dev is an open-weight MoE coding model with 35B total parameters (3B active), achieving state-of-the-art results on agentic coding benchmarks through SFT and RL training.
I ran the 35B agentic comparison someone asked for (stock vs Ornith vs KAT-Coder, 120 runs)
A detailed bakeoff of 35B coding models (KAT-Coder-V2.5-Dev, Qwen3.5, Ornith, etc.) with 120 runs shows KAT-Coder matching the best stock pass rate with cleaner tool behavior, while Ornith fails due to mechanical issues. Full methodology and results are linked.
Gemma 4 31B's competence surprised me
A user shares anecdotal findings that Gemma 4 31B outperforms Qwen 3.6 models and matches Opus 4.7 in understanding and refactoring messy academic code, highlighting a benchmark (SciCode) where Gemma excels.