Tag
A tweet notes that benchmarks quickly become saturated, citing the example of a model called GPT-5.6 Sol Pro scoring 91/99 on prinzbench, with two questions remaining unsolved.