can someone explain why we think a 90%+ bench is considered saturated?

Reddit r/ArtificialInteligence News

Summary

The article questions why AI benchmarks are considered saturated at 90%+ accuracy and advocates for aiming at 100% or developing new benchmarks.

help me wrap my mind around this because I can't find an honest reason. I understand we get to 95%, but we want models that we can trust, why not aim for 100%, why do we add new 'benches' and not try to get 100%?
Original Article

Similar Articles

Life After Benchmark Saturation: A Case Study of CORE-Bench

arXiv cs.AI

This paper argues against the 'retire-and-replace' approach to saturated benchmarks, using CORE-Bench as a case study to demonstrate that measuring agent performance along dimensions such as construct validity, efficiency, reliability, and human-agent collaboration yields meaningful insights even after accuracy plateaus.