DeepSeek v4 Flash has a nice bump in Capability

Reddit r/LocalLLaMA Models

Summary

DeepSeek V4 Flash shows significant benchmark gains in preview updates, trading blows with GPT-5.6 Terra on agentic coding tasks.

DeepSeek V4 Flash: Preview → 2026-07-31 Benchmark Preview 0731 Δ Terminal Bench* 56.9 82.7 +25.8 Toolathlon 51.8 70.3 +18.5 NL2Repo — 54.2 new Cybergym — 76.7 new DeepSWE — 54.4 new Agent Last Exam — 25.2 new Automation Bench — 25.1 new DSBench-FullStack — 68.7 new DSBench-Hard — 59.6 new * Terminal Bench changed from v2.0 → v2.1, so the improvement isn't a strict apples-to-apples comparison. Compared to GPT-5.6 Terra Benchmark GPT-5.6 Terra DeepSeek V4 Flash Advantage Terminal Bench 78.4 82.7 Flash (+4.3) Toolathlon 53.1 70.3 Flash (+17.2) DeepSWE 69.6 54.4 Terra (+15.2) Agents' Last Exam 50.4 25.2 Terra (+25.2) Trading blows, but no clear winner... very interesting! Source: https://api-docs.deepseek.com/updates/
Original Article

Similar Articles