@SanthProject: Now this is a bench i can get behind not the rigged as fuck deepswe benchmark

X AI KOLs Following Tools

Summary

SanthProject praises Cognition's new FrontierCode coding evaluation benchmark, calling it a fair alternative to the DeepSwe benchmark.

Now this is a bench i can get behind not the rigged as fuck deepswe benchmark
Original Article
View Cached Full Text

Cached at: 06/08/26, 11:29 PM

Now this is a bench i can get behind not the rigged as fuck deepswe benchmark

Cognition (@cognition): Introducing FrontierCode: a coding eval that raises the bar for difficulty & quality. Each task took 40+ hrs of work by leading open-source maintainers.

Models write sloppy code that works but isn’t maintainable. Our eval is first to measure: would you actually merge this code?

Similar Articles