@SanthProject: Now this is a bench i can get behind not the rigged as fuck deepswe benchmark
Summary
SanthProject praises Cognition's new FrontierCode coding evaluation benchmark, calling it a fair alternative to the DeepSwe benchmark.
View Cached Full Text
Cached at: 06/08/26, 11:29 PM
Now this is a bench i can get behind not the rigged as fuck deepswe benchmark
Cognition (@cognition): Introducing FrontierCode: a coding eval that raises the bar for difficulty & quality. Each task took 40+ hrs of work by leading open-source maintainers.
Models write sloppy code that works but isn’t maintainable. Our eval is first to measure: would you actually merge this code?
Similar Articles
@Suhail: Excited to see benchmarks headed in this direction and getting better!
Proximal has released FrontierSWE v2, an updated ultra-long horizon coding benchmark with expanded tasks and improved methodology, highlighting large performance gaps where Claude Fable 5.1 leads.
@denizbirlikci: To understand why we built FrontierCode, read @METR_Evals's blog post on why "many SWE-bench-passing PRs would not be m…
Cognition announces FrontierCode, a new coding evaluation benchmark that goes beyond unit tests to measure code quality, scope, test correctness, and human reviewer approval, addressing the issue of agents writing sloppy code that passes tests but is not maintainable.
@garrytan: This is the new standard for engineering evals
Announcing DeepSWE, a new benchmark for agentic coding that reveals true differences between models, reflecting real-world developer experiences.
@ArizePhoenix: The result: a 50% improvement over 5.2 on their in-house Code Bench, Terminal-Bench 3.0 from 4.6 → 28.3, DeepSWE v1.1 f…
DeepSWE v1.1 shows a 50% improvement on in-house Code Bench and sets new state-of-the-art scores on Terminal-Bench 3.0 and Agents' Last Exam, with emerging cyber capabilities advancing faster than expected.
@cognition: We’ve made improvements to the FrontierCode methodology and are releasing FrontierCode 1.1 with clearer guidelines for …
Cognition releases FrontierCode 1.1, an updated benchmark for evaluating code quality with refined guidelines for fair internet use and grading criteria, along with new model scores for Sonnet 5 and Fable 5.