Tag
A Yale-led paper finds that frontier AI models' poor performance on physics benchmarks is largely due to benchmark flaws rather than model limitations, suggesting benchmark quality is a critical bottleneck for accurate evaluation.
The article discusses a research paper revealing that many science-based LLM benchmarks have incorrect answers, and when corrected, LLM performance scores rise significantly, suggesting current evaluations may underestimate AI capabilities in physics.
Expert re-grading of physics benchmarks reveals that frontier AI models perform better than previously evaluated, indicating broken assessments and near-saturation on closed-ended tasks, which underscores the need for more rigorous evaluations.