Hy3 Benchmark Roundup: from SWE-Bench Pro to 312 real-world workflow tasks
Summary
A roundup of benchmarks including SWE-Bench Pro and 312 real-world workflow tasks, likely evaluating AI performance on software engineering challenges.
Similar Articles
Ramp SWE-Bench (3 minute read)
Ramp built a private benchmark from 80 production backend tasks to evaluate AI coding agents on review-ready patches, exposing trade-offs in accuracy, latency, and cost without public-benchmark contamination.
New benchmark dropped
A new benchmark has been released, likely for evaluating AI or software performance.
SWE-rebench leaderboard update: GLM-5.2, Qwen3.6-27B, Qwen3.6-35B-A3B, Gemma 4 31B and more + improved UI
SWE-rebench leaderboard updated with new models (GLM-5.2, Qwen3.6, Gemma 4 31B, etc.) and an improved UI, showing performance rankings on software engineering tasks.
Senior SWE Bench: a new benchmark focussed on realistically underspecified feature tasks
Senior SWE-Bench is a new open-source benchmark designed to evaluate AI agents on realistic, underspecified software engineering tasks, emphasizing skills like intent alignment and code quality rather than overly detailed specifications.
CursorBench 3.1
CursorBench 3.1 introduces new benchmark tasks focused on codebase understanding, bugfinding, planning, and code review, and presents updated scores and cost comparisons for various AI models.