Tag
A systematic study defining and analyzing benchmark saturation across 60 language model benchmarks, finding nearly half exhibit saturation and that expert-curated benchmarks are more resilient, suggesting design choices for durable evaluation.
This paper introduces CurveShift, an analysis method that separates overall ability gains from difficulty-specific improvements in LLM agents. Using METR time-horizon data and LiveCodeBench, it finds that most apparent shifts toward harder tasks are ceiling effects, though a genuine hard-task effect exists for reasoning models in competitive programming.
A large-scale human evaluation of 10 frontier LLMs across 62 questions finds that all models exhibit response drift, with most converging to a 78-81% deviation ceiling, while two achieve lower deviation. Drift varies by domain and question, and automated metrics explain little of human judgments, highlighting the need for human evaluation.
The article predicts that the AI benchmark 'Agents Last Exam' will reach saturation (models maxing out performance) by February next year.
This paper argues against the 'retire-and-replace' approach to saturated benchmarks, using CORE-Bench as a case study to demonstrate that measuring agent performance along dimensions such as construct validity, efficiency, reliability, and human-agent collaboration yields meaningful insights even after accuracy plateaus.
This paper proposes offline preference-based trajectory evaluation for agentic systems, which compares trajectories via temporal preferences rather than binary success metrics. It shows that this approach reduces ties from roughly 75% to 35%, improving discriminative power and data efficiency across diverse benchmarks.
BenchEvolver is an evolutionary framework that automatically generates harder coding problems from existing ones, creating challenging benchmarks that maintain validity and diversity while enabling model self-improvement and enhanced training performance.
This paper introduces a population coupling trend and h-field diagnostic to analyze the relationship between coding and reasoning capabilities across frontier AI models, finding that capabilities cooperate but with varying emphasis per lab. It provides a playbook for measurement and predicts benchmark saturation trends.