@gdb: benchmarks get saturated very quickly these days

X AI KOLs Following News

Summary

A tweet notes that benchmarks quickly become saturated, citing the example of a model called GPT-5.6 Sol Pro scoring 91/99 on prinzbench, with two questions remaining unsolved.

benchmarks get saturated very quickly these days
Original Article
View Cached Full Text

Cached at: 07/17/26, 12:19 AM

benchmarks get saturated very quickly these days

prinz (@deredleritt3r): Added to prinzbench: GPT-5.6 Sol Pro.

As previewed a few days ago, this model has saturated my benchmark, with a total score of 91/99.

For context, prinzbench contains two questions that no model tested to date has ever been able to solve (one requires extremely thorough

Similar Articles

Introducing BenchBench (5 minute read)

TLDR AI

Introduces BenchBench, a benchmark that tests AI models' ability to create effective benchmarks for other models, with GPT 5.2 being the only successful winner so far while frontier models like GPT 5.5 and Opus 4.6 struggled.