Terminal Bench 3 has been released. It’s a new benchmark that hasn’t been included in model training sets yet. (I’m not showing the results from third-party harnesses to keep things fair.)
Summary
Terminal Bench 3, a new benchmark for terminal-based AI models that excludes third-party harness results and has not been used in training sets, has been released.
Similar Articles
TerminalBench 2.1 from GPT‑5.6 Sol, Terra, and Luna
TerminalBench 2.1 is a benchmark suite derived from GPT‑5.6 Sol, Terra, and Luna models, likely used for evaluating AI performance on terminal-based tasks.
CursorBench 3.1
CursorBench 3.1 introduces new benchmark tasks focused on codebase understanding, bugfinding, planning, and code review, and presents updated scores and cost comparisons for various AI models.
Introducing BenchBench (5 minute read)
Introduces BenchBench, a benchmark that tests AI models' ability to create effective benchmarks for other models, with GPT 5.2 being the only successful winner so far while frontier models like GPT 5.5 and Opus 4.6 struggled.
New benchmark dropped
A new benchmark has been released, likely for evaluating AI or software performance.
New bench designed for smaller models: ObviousBench.com
ObviousBench is a new benchmark designed specifically for evaluating smaller AI models.