CursorBench 3.1

Hacker News Top Tools

Summary

CursorBench 3.1 introduces new benchmark tasks focused on codebase understanding, bugfinding, planning, and code review, and presents updated scores and cost comparisons for various AI models.

No content available
Original Article
View Cached Full Text

Cached at: 07/02/26, 08:04 AM

# Cursor · CursorBench Source: [https://cursor.com/evals](https://cursor.com/evals) We evaluate agents on ambiguous, multi\-file tasks from real Cursor sessions\. Higher scores are better\. [More about CursorBench](https://cursor.com/blog/cursorbench)A scatter and line chart comparing Fable 5, Opus 4\.8, Opus 4\.7, GPT\-5\.5, Sonnet 5, Sonnet 4\.6, GLM 5\.2, Composer 2\.5, and Composer 2 scores against average cost per task\.75% CursorBench 3\.1 score70%65%60%55%50%45%$20$16$12$8$4$0Average cost per taskFable 5 highComposer 2\.5GPT\-5\.5 mediumGemini 3\.5 FlashOpus 4\.8 highSonnet 5 highKimi K2\.7 CodeGLM 5\.2 highModel1Fable 5 Max72\.9%$18\.0263,842762Fable 5 Extra High72\.0%$13\.7448,754633Fable 5 High70\.6%$10\.8137,173544Fable 5 Medium69\.8%$8\.2728,507475Opus 4\.7 Max64\.8%$11\.0262,989966GPT\-5\.5 Extra High64\.3%$4\.3717,905467Fable 5 Low64\.2%$5\.7018,882368Opus 4\.8 Max63\.8%$7\.5977,370609Composer 2\.563\.2%$0\.5515,1523710GPT\-5\.5 High62\.6%$3\.5913,3294011Opus 4\.8 Extra High62\.1%$6\.1455,6225412Opus 4\.7 Extra High61\.6%$7\.1143,9427213Sonnet 5 Max61\.2%$6\.8793,4859314Opus 4\.7 High59\.4%$5\.0132,2275915GPT\-5\.5 Medium59\.2%$2\.229,0653516Opus 4\.8 High58\.4%$4\.4136,7884517Sonnet 5 Extra High58\.4%$5\.2358,2288618Sonnet 5 High57\.0%$3\.7441,7356619Opus 4\.8 Medium56\.6%$3\.8331,6844120Sonnet 5 Medium54\.9%$2\.5727,4695321GLM 5\.2 Max54\.6%$3\.1151,3128322Opus 4\.8 Low54\.3%$2\.9322,7263623Opus 4\.7 Medium52\.7%$2\.9319,1934124Kimi K2\.7 Code52\.7%$1\.9232,9027025Composer 252\.2%$0\.5614,1634026GLM 5\.2 High50\.7%$2\.4630,6217627Gemini 3\.5 Flash49\.8%$1\.9435,1057928Sonnet 4\.6 Max49\.0%$3\.0940,2805529GPT\-5\.5 Low48\.8%$1\.194,9232430Sonnet 4\.6 High48\.8%$3\.0637,3525731Opus 4\.7 Low48\.3%$1\.8713,1642932Sonnet 5 Low47\.7%$1\.4617,0283733Kimi 2\.647\.6%$1\.2724,7835634Sonnet 4\.6 Medium46\.0%$2\.6431,3605035Sonnet 4\.6 Low41\.5%$1\.8921,2115036Kimi 2\.531\.9%$0\.879,44630 ## Changelog ### CursorBench 3\.1 - Introduced problems focused on codebase understanding, bugfinding, planning, and code review\. - Improved grading criteria for some edit tasks\. ### CursorBench 3\.0 - Initial set of tasks focused on edit, refactor, and bugfix problems\. Avg cost / task is computed by applying each model's published[per\-million\-token pricing](https://cursor.com/docs/models-and-pricing)\(input, cache read, cache write, and output\) to the tokens it used on each CursorBench 3\.1 task, then averaging across tasks\. Results are subject to variance; small differences in scores may not be statistically meaningful\.

Similar Articles

ProgramBench (5 minute read)

TLDR AI

ProgramBench is a new benchmark that evaluates AI agents' ability to reconstruct complete software projects from compiled binaries and documentation without access to source code or decompilation tools.

New benchmark dropped

Reddit r/singularity

A new benchmark has been released, likely for evaluating AI or software performance.