benchmark-saturation

Tag

Cards List
#benchmark-saturation

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Hacker News Top · 6d ago Cached

A systematic study defining and analyzing benchmark saturation across 60 language model benchmarks, finding nearly half exhibit saturation and that expert-curated benchmarks are more resilient, suggesting design choices for durable evaluation.

0 favorites 0 likes
#benchmark-saturation

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

arXiv cs.CL · 2026-08-04 Cached

This paper introduces CurveShift, an analysis method that separates overall ability gains from difficulty-specific improvements in LLM agents. Using METR time-horizon data and LiveCodeBench, it finds that most apparent shifts toward harder tasks are ceiling effects, though a genuine hard-task effect exists for reasoning models in competitive programming.

0 favorites 0 likes
#benchmark-saturation

Response drift across frontier large language models

arXiv cs.CL · 2026-07-24 Cached

A large-scale human evaluation of 10 frontier LLMs across 62 questions finds that all models exhibit response drift, with most converging to a 78-81% deviation ceiling, while two achieve lower deviation. Drift varies by domain and question, and automated metrics explain little of human judgments, highlighting the need for human evaluation.

0 favorites 0 likes
#benchmark-saturation

Agents Last Exam will be saturated by next February at the latest.

Reddit r/singularity · 2026-07-20

The article predicts that the AI benchmark 'Agents Last Exam' will reach saturation (models maxing out performance) by February next year.

0 favorites 0 likes
#benchmark-saturation

Life After Benchmark Saturation: A Case Study of CORE-Bench

arXiv cs.AI · 2026-06-26 Cached

This paper argues against the 'retire-and-replace' approach to saturated benchmarks, using CORE-Bench as a case study to demonstrate that measuring agent performance along dimensions such as construct validity, efficiency, reliability, and human-agent collaboration yields meaningful insights even after accuracy plateaus.

0 favorites 0 likes
#benchmark-saturation

Offline Preference-Based Trajectory Evaluation

arXiv cs.LG · 2026-06-17 Cached

This paper proposes offline preference-based trajectory evaluation for agentic systems, which compares trajectories via temporal preferences rather than binary success metrics. It shows that this approach reduces ties from roughly 75% to 35%, improving discriminative power and data efficiency across diverse benchmarks.

0 favorites 0 likes
#benchmark-saturation

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

Hugging Face Daily Papers · 2026-05-31 Cached

BenchEvolver is an evolutionary framework that automatically generates harder coding problems from existing ones, creating challenging benchmarks that maintain validity and diversity while enabling model self-improvement and enhanced training performance.

0 favorites 0 likes
#benchmark-saturation

The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next

arXiv cs.LG · 2026-05-20

This paper introduces a population coupling trend and h-field diagnostic to analyze the relationship between coding and reasoning capabilities across frontier AI models, finding that capabilities cooperate but with varying emphasis per lab. It provides a playbook for measurement and predicts benchmark saturation trends.

0 favorites 0 likes
← Back to home

Submit Feedback