We tested Deepseek v4 flash, GLM 5.2, and Kimi K3 on hard agentic tasks, and DeepSeek just crushed

Reddit r/LocalLLaMA News

Summary

Composio tested DeepSeek v4 flash, GLM 5.2, and Kimi K3 on hard agentic tasks, finding DeepSeek the fastest and cheapest with roughly the same success rate as the others, while frontier models still lead slightly.

https://preview.redd.it/uybjxyypj7hh1.png?width=1200&format=png&auto=webp&s=8293e8da332a14920b335caee52473762c09530d As part of our internal eval report, we (@Composio) regularly test the models on some of the hardest agentic tasks. And we just tested the big 3 of open-weight models on our new internal evals. The tasks include long-running workflows involving multiple applications (Pagerduty, Gmail, HubSpot, Airtable, Slack, etc). Each task is scored by a fixed set of deterministic checks, and it passes only if every check passes. No points for partial success. To keep the test impartial, we used Pi Agent as the harness paired with Composio MCP to access the apps it needs. Here's what we found out, Time taken to accomplish a task DeepSeek v4 flash, well, as the name suggests, was the fastest of all. It took 164s per task. Almost 2.5x faster than GLM and 1.4x faster than Kimi K3. Kimi K3, despite being bigger, was faster than GLM 5.2. We also calculated the cost for each task by each model. With input tokens accounting for roughly 95% of spend, the average cost per task came out to: DeepSeek: ~$0.08 GLM 5.2: ~$0.57 Kimi K3: ~$1.39 Despite the large differences in speed and price, success rates were almost identical: DeepSeek: 20/30 Kimi: 21/30 GLM: 21/30 None of the three models was far behind frontier models on the same suite: Fable 5 and GPT-5.6 Sol each scored 24/30, and Opus 5 scored 23/30. Each model had a distinct working style: GLM was slower and token-lean, DeepSeek was fast and token-hungry, and Kimi was somewhere in between. Avg tokens/task: GLM 5.2: 409k Kimi K3: 462k DeepSeek: 569k. Even so, DeepSeek was the fastest and cheapest. State-of-the-art models like Fable and GPT 5.6 Sol are still better when that extra raw IQ force is needed, but at a far higher cost. We absolutely expected Kimi and GLM to do well. But DeepSeek V4 flash gave us a pleasant surprise. The Whale bros are still in contention. Would love to know your experience so far with open-weight models on agentic workflows.
Original Article

Similar Articles