DeepSWE benchmarks indicate that DeepSeek v4 Pro only passes 8% of tasks
Summary
A discussion about DeepSWE benchmarks showing that DeepSeek v4 Pro passes only 8% of tasks, which is surprisingly low compared to its performance on similar tasks.
Similar Articles
DeepSWE Opus 4.8 results have been released.
The results of DeepSWE Opus 4.8 have been released, showcasing its performance on benchmarks.
How good is DeepSeek-V4 Flash, actually?
An evaluation of the performance and capabilities of DeepSeek-V4 Flash, assessing its real-world effectiveness.
Someone did an audit on the new DeepSWE, the results aren't pretty
DeepSWE is a new benchmark for evaluating AI coding agents on real-world software engineering tasks from active open-source repositories, comprising 113 tasks across TypeScript, Go, Python, JavaScript, and Rust with isolated environments and program-based verifiers.
I have (even faster) DeepSeek V4 Pro at home
A user reports successfully running the DeepSeek V4 Pro model locally using ktransformers and sharing detailed benchmark results across various context depths, demonstrating improved inference speeds.
How can Deepseek v4 top the coding leaderboards and still sit 8 months behind the frontier?
Analysis of DeepSeek V4's top coding scores versus its reported 8-month gap behind the frontier, highlighting differences between narrow benchmark optimization and broader reasoning tests, plus the practical performance hit when running quantized local versions.