Tag
Researchers show agent leaderboards rank task specialization rather than capability, with agent main effects accounting for under 3% of variance across three benchmarks, and propose a Deployment Decision Reliability framework.
This paper applies Generalizability Theory to agent benchmarks, showing leaderboards rank specialization rather than capability, and proposes a framework (DDR) for sizing reliable deployment evaluations.