Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
Summary
Stanford researchers develop a Bayesian variance-decomposition framework using Generalizability theory to diagnose reliability in agent benchmarks, showing that adding more tasks cannot fix uncertainty caused by scaffold coverage and that pooling diverse benchmarks is a cheaper path to reliable rankings.
View Cached Full Text
Cached at: 10/02/26, 09:46 AM
# Agent Evaluation Reliability: More Tasks Won’t (Always) Fix An Agent Leaderboard
Source: [https://arxiv.org/html/2610.00651](https://arxiv.org/html/2610.00651)
Michael Hardy♢,∗\\diamondsuit,\*Ruhana Azam♣\\clubsuitAnka Reuel♢,†\\diamondsuit,\\daggerMykel Kochenderfer♢,†\\diamondsuit,\\daggerSanmi Koyejo♢,†\\diamondsuit,\\dagger♢\\diamondsuitStanford University,∗\*hardym\[α\]stanford⋅\\cdotedu♣\\clubsuitUniversity of Illinois Urbana\-Champaign†\\daggerSenior Author
###### Abstract
Agent evaluations are increasingly used to compare models, assess capabilities, and inform deployment decisions, yet observed scores can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks\. This makes reliability claim\-dependent: an evaluation that reliably ranks deployed systems may not reliably rank underlying models or produce reliable absolute scores\. We ask which conclusions current agent evaluations reliably support and what additional evaluation would actually improve them\. Using Generalizability theory, we develop a Bayesian variance\-decomposition framework for sparse, imbalanced agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index\. The framework separates*signal*, performance differences relevant to the intended claim, from*noise*, irrelevant variation that can still change scores or rankings\. We find four practical results: \(1\) Reliability depends on the measurement goal\. Fixed model–scaffold systems are ranked reliably \(Eρ2=0\.935E\\rho^\{2\}=0\.935–0\.9940\.994\), while underlying\-model reliability is substantially lower \(0\.1480\.148–0\.8410\.841\)\. \(2\) Scaffold choice can change evaluation conclusions\. We introduce*inter\-scaffold reliability*, measuring whether scaffolds preserve model rankings, and show that scaffold effects vary substantially across evaluations\. \(3\) More tasks cannot resolve all uncertainty\. Even infinitely many similarly constructed tasks improve model\-ranking reliability of the dataset by at most0\.0970\.097when uncertainty is dominated by limited scaffold coverage\. \(4\) Pooling diverse benchmarks can improve cross\-task rankings at lower cost\. For rankings across diverse agentic tasks, pooling benchmarks raises projected reliability from0\.440\.44to0\.750\.75at the same task budget and can reduce projected cost by up to83%83\\%\. Evaluation design should therefore follow the intended claim: practitioners should identify what a score or ranking should mean, diagnose what limits its reliability, and spend evaluation budget on the sources of uncertainty that matter\.111Code and data:[https://github\.com/hardy\-education/scaffold\_eval](https://github.com/hardy-education/scaffold_eval)
## 1Introduction
Agentic AI systems interact with software, tools, users, and environments to accomplish multi\-step goals\. Agent evaluations such as SWE\-bench\([Jimenez et al\., 2023](https://arxiv.org/html/2610.00651#bib.bib17)\)andτ\\tau\-bench\([Yao et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib22)\)are increasingly used to test and rank such systems in technical reports, system cards, and policy discussions\([Anthropic, 2026](https://arxiv.org/html/2610.00651#bib.bib27);[OpenAI, 2026](https://arxiv.org/html/2610.00651#bib.bib26)\)\. Yet, a score on such evaluations is only useful if we know what conclusions it can reliably support: Does an observed ranking reflect model performance differences that persist beyond the particular evaluation setup? Would the same score or ranking hold under a different scaffold, a different set of tasks, or a broader range of agentic settings?
These questions are particularly important for agent evaluations because performance depends on more than the underlying model\. A scaffold222In the current literature, the software infrastructure and post\-training methods surrounding the LLM allowing it to act is synonymous with “scaffold”, “harness”, and even “agent”\.manages the state, exposes tools, and translates model outputs into actions; together, the model333Throughout, “model” denotes a base large language model paired with a reasoning\-effort configuration; for example,gpt\-5\-highandgpt\-5\-minimalare treated as distinct, following our primary dataset\.and scaffold form the system\. This creates an immediate measurement choice: are we trying to evaluate the complete system as deployed, or isolate performance differences attributable to the underlying model? The answer changes what should count as signal and what should count as evaluation error\.
Reliability also depends on the conclusion we want to draw\. A practitioner may care whether a model ranking remains stable, whether an absolute score would remain similar, or whether performance generalizes across different agentic tasks\. These are different measurement goals and need not have the same reliability\. Existing agent reliability work primarily studies whether a fixed agent behaves consistently across repeated runs or prompt perturbations of the same task\([Rabanser et al\., 2026](https://arxiv.org/html/2610.00651#bib.bib6);[Razavi et al\., 2025](https://arxiv.org/html/2610.00651#bib.bib30);[Yao et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib22)\)\. We instead ask:what claims do current agent evaluations reliably support, and what additional evaluation would make those claims more reliable?
We use Generalizability Theory \(G\-theory\)\([Cronbach et al\., 1972](https://arxiv.org/html/2610.00651#bib.bib2);[Brennan, 2001](https://arxiv.org/html/2610.00651#bib.bib15)\)to separate variation associated with models, tasks, scaffolds, benchmarks, and their interactions\. To word it differently, our framework separates*signal*–performance differences that persist across the conditions an evaluation is meant to generalize over–from*noise*–variation that is irrelevant to the intended claim but can still change the resulting score or ranking; reliability increases when signal dominates this noise\. G\-theory lets us identify not only how reliable an evaluation is for a particular claim, but also what limits that reliability and whether adding more tasks, scaffolds, or benchmark coverage would help\. We develop a Bayesian variance\-decomposition approach for the sparse and imbalanced designs common in agent leaderboards and apply it to30,85930\{,\}859rollouts from 9 benchmarks in the Holistic Agent Leaderboard \(HAL\)\([Kapoor et al\., 2025](https://arxiv.org/html/2610.00651#bib.bib4)\)and 13 the Harbor Index\([Shi et al\., 2026](https://arxiv.org/html/2610.00651#bib.bib49)\)\.
Our results yield four practical findings:
- •Reliability depends on the measurement goal\. Fixed model–scaffold systems are ranked reliablyEρMA2∈\[0\.94E\\rho\_\{MA\}^\{2\}\\in\[0\.94,0\.99\]0\.99\], while rank reliability for the models is substantially lowerEρM2∈\[0\.148,0\.841\]E\\rho\_\{M\}^\{2\}\\in\[0\.148,0\.841\]\.
- •Changing the scaffold can change evaluation conclusions\. We introduce*inter\-scaffold reliability*, which measures whether different scaffolds produce similar model rankings when evaluating the same models on the same tasks\. Its posterior medians range from0\.1510\.151to0\.8520\.852across benchmarks\. We further show that scaffold choice can change which tasks a model solves even without improving its overall benchmark score\.
- •Adding more tasks does not always make an evaluation more reliable\. Under the observed scaffold coverage, even infinitely many similarly constructed tasks would improve model\-ranking reliability by at most approximately0\.100\.10\.
- •When the goal is to rank models across diverse agentic tasks, pooling diverse benchmarks can improve reliability at lower cost\. At the same task budget, projected ranking reliability rises from approximately0\.440\.44within one benchmark to0\.750\.75across the nine\-benchmark battery, while reliability\-aware allocation can reduce projected evaluation cost by up to83%83\\%\.
The broader lesson is that more evaluation is not automatically better evaluation\. Practitioners should first specify what they want an evaluation result to mean, then identify which sources of uncertainty limit that claim, and allocate evaluation effort accordingly\.
## 2Related Work and Background
##### Agent evaluation\.
Agent benchmarks span software engineering\([Jimenez et al\., 2023](https://arxiv.org/html/2610.00651#bib.bib17)\), web navigation\([Zhou et al\., 2023](https://arxiv.org/html/2610.00651#bib.bib25);[He et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib24)\), general assistance\([Mialon et al\., 2023](https://arxiv.org/html/2610.00651#bib.bib18)\), and customer service\([Yao et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib22)\)\. Existing reliability studies primarily test whether a fixed agent behaves consistently across repeated runs, prompt perturbations, tool configurations, or environmental failures\([Rabanser et al\., 2026](https://arxiv.org/html/2610.00651#bib.bib6);[Razavi et al\., 2025](https://arxiv.org/html/2610.00651#bib.bib30);[Wang et al\., 2026](https://arxiv.org/html/2610.00651#bib.bib29);[Kumar and Mishra, 2025](https://arxiv.org/html/2610.00651#bib.bib28)\)\. These analyses assess robustness of a particular system\. We instead study the reliability of the*comparative inference*: whether the reported ordering of models or systems would persist under new tasks, scaffolds, or benchmarks\. This distinction is consequential because a system can be repeatable within one harness while its rank is highly contingent on that harness\.
##### Generalizability theory separates signal from conditional advantage\.
Generalizability theory \(G\-theory\) treats evaluation conditions as measurement*facets*and decomposes score variation into their main effects and interactions\([Cronbach et al\., 1972](https://arxiv.org/html/2610.00651#bib.bib2);[Shavelson et al\., 1989](https://arxiv.org/html/2610.00651#bib.bib3);[Brennan, 2001](https://arxiv.org/html/2610.00651#bib.bib15);[Cronbach and Shavelson, 2004](https://arxiv.org/html/2610.00651#bib.bib8)\)\. Related work shows that AI benchmark conclusions depend on task sampling, evaluation conditions, and the model population\([Madaan et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib13);[Hardy and Kim, 2026](https://arxiv.org/html/2610.00651#bib.bib7);[Hardy et al\., 2026](https://arxiv.org/html/2610.00651#bib.bib31)\)\. The crucial first choice is the*object of measurement*\. Model–scaffold compatibility is signal when selecting a deployable system, but error when ranking models independently of scaffolding\. Reliability therefore belongs to an object, a generalization universe, and an evaluation design, not to a benchmark alone\.
For an objectooand design𝒟\\mathcal\{D\}, the relative generalizability coefficient and its signal\-to\-noise ratio are
Eρo2\(𝒟\)=σo2σo2\+σδ2\(𝒟\),SNRo\(𝒟\)=σo2σδ2\(𝒟\)=Eρo21−Eρo2\.E\\rho\_\{o\}^\{2\}\(\\mathcal\{D\}\)=\\frac\{\\sigma\_\{o\}^\{2\}\}\{\\sigma\_\{o\}^\{2\}\+\\sigma\_\{\\delta\}^\{2\}\(\\mathcal\{D\}\)\},\\qquad\\operatorname\{SNR\}\_\{o\}\(\\mathcal\{D\}\)=\\frac\{\\sigma\_\{o\}^\{2\}\}\{\\sigma\_\{\\delta\}^\{2\}\(\\mathcal\{D\}\)\}=\\frac\{E\\rho\_\{o\}^\{2\}\}\{1\-E\\rho\_\{o\}^\{2\}\}\.\(1\)Hereσo2\\sigma\_\{o\}^\{2\}is the variance of scores averaged over the specified universe, andσδ2\\sigma\_\{\\delta\}^\{2\}contains variation that can change relative standing\. UnderXo=Zo\+δoX\_\{o\}=Z\_\{o\}\+\\delta\_\{o\}, with uncorrelated universe score and error,Eρo2=Corr2\(Xo,Zo\)E\\rho\_\{o\}^\{2\}=\\operatorname\{Corr\}^\{2\}\(X\_\{o\},Z\_\{o\}\)\.
Facet main effects cancel from comparisons only when objects share the same conditions and weights\. A uniformly difficult task, for example, does not change relative latent scores under a common allocation\. Object\-by\-facet interactions do: they represent advantages that depend on the chosen task, scaffold, or benchmark\. G\-theory uses these components in a*Decision study*\(D\-study\) to project reliability before collecting more data\.
## 3Reliability and Generalizability for Agent Leaderboards
Agent leaderboards are often sparse, imbalanced, and partially crossed: models are evaluated through different scaffolds, task counts are unequal, and many combinations are absent\. In this section, we discuss several metrics to measure reliability in agent leaderboards given these common constraints\. We fit a Bayesian variance decomposition to binary task outcomes, then propagate its uncertainty into reliability, design projections, and model ranks\.
### 3\.1Observations and Generalization Universe
Letb∈ℬb\\in\\mathcal\{B\}index benchmarks,i∈ℐbi\\in\\mathcal\{I\}\_\{b\}tasks nested within benchmarkbb,m∈ℳm\\in\\mathcal\{M\}models, anda∈𝒜a\\in\\mathcal\{A\}agent scaffolds\. The responseybima∈\{0,1\}y\_\{bima\}\\in\\\{0,1\\\}indicates whether modelmm, operated through scaffoldaa, solves taskiifrom benchmarkbb\. Our primary object is the modelmm\. Its universe score is the component expected to persist across new tasks, benchmarks, and scaffolds resembling those represented in the leaderboard\. We also consider the model–scaffold pair\(m,a\)\(m,a\), the relevant object when selecting a deployable system\. We treat tasks as draws from benchmark\-specific construction and grading processes; benchmarks as instruments drawn from a battery of contemporary agent evaluations; scaffolds as contemporary evaluation and orchestration systems; and models as the frontier\-model population represented in the data\. These universes delimit the inference\. In particular, generalization to new benchmark\-like tasks does not establish that a benchmark represents all real\-world uses\.
### 3\.2A Latent Variance Decomposition for Binary Outcomes
We use Bayesian Bernoulli–logit mixed models:ybima∼Bernoulli\(pbima\)y\_\{bima\}\\sim\\operatorname\{Bernoulli\}\(p\_\{bima\}\),logit\(pbima\)=ηbima\\operatorname\{logit\}\(p\_\{bima\}\)=\\eta\_\{bima\}\. Equivalently,ybima=𝟙\{zbima\>0\}y\_\{bima\}=\\mathbb\{1\}\\\{z\_\{bima\}\>0\\\}, wherezbima=ηbima\+εbimaz\_\{bima\}=\\eta\_\{bima\}\+\\varepsilon\_\{bima\}andεbima∼Logistic\(0,1\)\\varepsilon\_\{bima\}\\sim\\operatorname\{Logistic\}\(0,1\), giving the standard latent logistic residual varianceπ2/3\\pi^\{2\}/3\. Modeling the binary responses directly avoids Gaussian approximations that are particularly misleading under floor or ceiling effects\. All random effects are mean\-zero Gaussian and mutually independent; for example,um\(M\)∼𝒩\(0,σM2\)u\_\{m\}^\{\(M\)\}\\sim\\mathcal\{N\}\(0,\\sigma\_\{M\}^\{2\}\)\. Variance components and reliability coefficients are defined on the common latent log\-odds scale \(see Appendix[G\.1](https://arxiv.org/html/2610.00651#A7.SS1)\)\. From an Item Response Theory \(IRT\) perspective, these are random\-item, many\-facet Rasch models\([Wang and Wilson, 2005](https://arxiv.org/html/2610.00651#bib.bib40);[Fox and Glas, 2001](https://arxiv.org/html/2610.00651#bib.bib41);[Linacre and Wright, 2002](https://arxiv.org/html/2610.00651#bib.bib42)\)\.
#### 3\.2\.1Benchmark\-level model and system reliability
For each benchmarkbb, we fit
ηima\(b\)=β0\(b\)\+ui\(I\)\+um\(M\)\+ua\(A\)\+uim\(IM\)\+uia\(IA\)\+uma\(MA\)\.\\eta\_\{ima\}^\{\(b\)\}=\\beta\_\{0\}^\{\(b\)\}\+u\_\{i\}^\{\(I\)\}\+u\_\{m\}^\{\(M\)\}\+u\_\{a\}^\{\(A\)\}\+u\_\{im\}^\{\(IM\)\}\+u\_\{ia\}^\{\(IA\)\}\+u\_\{ma\}^\{\(MA\)\}\.\(2\)Variance components are benchmark\-specific, with the indexbbsuppressed for readability\. The decomposition separates persistent model differences from task sensitivity, scaffold differences, and model–scaffold compatibility\. Because repeated observations of the same\(i,m,a\)\(i,m,a\)cell are rare, the task–model–scaffold interaction is not separately identifiable from latent response variation; we denote this terminal residual variance byσIMA,e2=π2/3\\sigma\_\{IMA,e\}^\{2\}=\\pi^\{2\}/3\. For an equal\-allocation design withnin\_\{i\}tasks andnan\_\{a\}scaffolds, model\-ranking reliability is
EρM\(b\)2\(ni,na\)=σM2σM2\+σIM2/ni\+σMA2/na\+σIMA,e2/\(nina\)\.E\\rho\_\{M\(b\)\}^\{2\}\(n\_\{i\},n\_\{a\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{IM\}^\{2\}/n\_\{i\}\+\\sigma\_\{MA\}^\{2\}/n\_\{a\}\+\\sigma\_\{IMA,e\}^\{2\}/\(n\_\{i\}n\_\{a\}\)\}\.\(3\)Scaffold main effects do not enter the denominator because shifting every model equally does not change their ordering\. In contrast, model–scaffold interactions are rank\-relevant: they encode which models benefit from which scaffolds\. When the object is the completemodel–scaffold system,
EρMA\(b\)2\(ni\)=σM2\+σA2\+σMA2σM2\+σA2\+σMA2\+\(σIM2\+σIA2\+σIMA,e2\)/ni\.E\\rho\_\{MA\(b\)\}^\{2\}\(n\_\{i\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{A\}^\{2\}\+\\sigma\_\{MA\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{A\}^\{2\}\+\\sigma\_\{MA\}^\{2\}\+\(\\sigma\_\{IM\}^\{2\}\+\\sigma\_\{IA\}^\{2\}\+\\sigma\_\{IMA,e\}^\{2\}\)/n\_\{i\}\}\.\(4\)
To measure whether scaffold choice preserves model ordering, we additionally define the inter\-scaffold reliability for two independently sampled scaffolds evaluated on the samenin\_\{i\}tasks:
ρAA′\(b\)\(ni\)=σM2\+σIM2/niσM2\+σIM2/ni\+σMA2\+σIMA,e2/ni\.\\rho\_\{AA^\{\\prime\}\}^\{\(b\)\}\(n\_\{i\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{IM\}^\{2\}/n\_\{i\}\}\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{IM\}^\{2\}/n\_\{i\}\+\\sigma\_\{MA\}^\{2\}\+\\sigma\_\{IMA,e\}^\{2\}/n\_\{i\}\}\.\(5\)This is analogous to inter\-rater reliability, with scaffolds acting as alternative measurement procedures\. A low value means that changing the scaffold can change which model appears strongest, even when each scaffold yields internally stable scores\.
#### 3\.2\.2Persistent model differences across pooled benchmarks
Each benchmark contains too few scaffolds to estimate all scaffold\-related components precisely in isolation\. We therefore also fit a joint model across the nine\-benchmark leaderboard:
ηbima=\\displaystyle\\eta\_\{bima\}=\{\}β0\+ub\(B\)\+ubi\(I\[B\]\)\+um\(M\)\+ua\(A\)\+ubm\(BM\)\+uba\(BA\)\+uma\(MA\)\+ubim\(IM\[B\]\)\+ubia\(IA\[B\]\)\+ubma\(BMA\)\.\\displaystyle\\beta\_\{0\}\+u\_\{b\}^\{\(B\)\}\+u\_\{bi\}^\{\(I\[B\]\)\}\+u\_\{m\}^\{\(M\)\}\+u\_\{a\}^\{\(A\)\}\+u\_\{bm\}^\{\(BM\)\}\+u\_\{ba\}^\{\(BA\)\}\+u\_\{ma\}^\{\(MA\)\}\+u\_\{bim\}^\{\(IM\[B\]\)\}\+u\_\{bia\}^\{\(IA\[B\]\)\}\+u\_\{bma\}^\{\(BMA\)\}\.\(6\)Items are nested within benchmarks, while models and scaffolds are crossed with benchmarks wherever supported by the observed incidence graph\. Shared models and scaffolds connect benchmarks and permit partial separation of their effects\. Hierarchical priors can produce estimates for weakly supported contrasts, but cannot supply empirical identification between disconnected components; connectivity and estimability diagnostics are reported in Appendix[E](https://arxiv.org/html/2610.00651#A5)\.
The leaderboard universe score for modelmmisum\(M\)u\_\{m\}^\{\(M\)\}, the component expected to persist across sampled benchmarks, tasks, and scaffolds\. Benchmark\-conditioned capabilityubm\(BM\)u\_\{bm\}^\{\(BM\)\}contributes to performance on benchmarkbb, but is not assumed to transfer to a new benchmark\. For a balanced design withnbn\_\{b\}benchmarks,nin\_\{i\}tasks per benchmark, andnan\_\{a\}scaffolds,
EρM2\(nb,ni,na\)=σM2σM2\+σBM2/nb\+σMA2/na\+σBMA2/\(nbna\)\+σIM\[B\]2/\(nbni\)\+σBIMA,e2/\(nbnina\)\.E\\rho\_\{M\}^\{2\}\(n\_\{b\},n\_\{i\},n\_\{a\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{BM\}^\{2\}/n\_\{b\}\+\\sigma\_\{MA\}^\{2\}/n\_\{a\}\+\\sigma\_\{BMA\}^\{2\}/\(n\_\{b\}n\_\{a\}\)\+\\sigma\_\{IM\[B\]\}^\{2\}/\(n\_\{b\}n\_\{i\}\)\+\\sigma\_\{BIMA,e\}^\{2\}/\(n\_\{b\}n\_\{i\}n\_\{a\}\)\}\.\(7\)HereσBIMA,e2=π2/3\\sigma\_\{BIMA,e\}^\{2\}=\\pi^\{2\}/3is terminal cell\-specific variation and the logistic residual \(see §[D\.3\.1](https://arxiv.org/html/2610.00651#A4.SS3.SSS1)for separated variance\)\. Equation[7](https://arxiv.org/html/2610.00651#S3.E7)maps each design intervention to the uncertainty it can reduce: tasks average task\-indexed error, benchmarks average benchmark\-conditioned differences, and scaffolds average scaffold\-conditioned differences\. Table[1](https://arxiv.org/html/2610.00651#S3.T1)consolidates the estimands used in the main body\.
Table 1:Reliability estimands\. Each row appliesEρo2=σo2/\(σo2\+σδ2\)E\\rho\_\{o\}^\{2\}=\\sigma\_\{o\}^\{2\}/\(\\sigma\_\{o\}^\{2\}\+\\sigma\_\{\\delta\}^\{2\}\),SNRo=σo2/σδ2\\text\{SNR\}\_\{o\}=\\sigma\_\{o\}^\{2\}/\\sigma\_\{\\delta\}^\{2\}, but changes the object whose ordering should generalize\. Thenfn\_\{f\}terms denote equal allocation over facetff\.EstimandObject of inference𝝈𝒐𝟐\\bm\{\\sigma\_\{o\}^\{2\}\}𝝈𝜹𝟐\(𝓓\)\\bm\{\\sigma\_\{\\delta\}^\{2\}\(\\mathcal\{D\}\)\}Eq\.EρM\(b\)2E\\rho^\{2\}\_\{M\(b\)\},SNRM\(b\)\\text\{SNR\}\_\{M\(b\)\}Does benchmarkbbpreserve model order?σM2\\sigma\_\{M\}^\{2\}σIM2ni\+σMA2na\+σIMA,e2nina\\dfrac\{\\sigma\_\{IM\}^\{2\}\}\{n\_\{i\}\}\+\\dfrac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\+\\dfrac\{\\sigma\_\{IMA,e\}^\{2\}\}\{n\_\{i\}n\_\{a\}\}\([3](https://arxiv.org/html/2610.00651#S3.E3)\)EρMA\(b\)2E\\rho^\{2\}\_\{MA\(b\)\}Does benchmarkbbpreserve system order?σM2\+σA2\+σMA2\\sigma\_\{M\}^\{2\}\+\\sigma\_\{A\}^\{2\}\+\\sigma\_\{MA\}^\{2\}σIM2\+σIA2\+σIMA,e2ni\\dfrac\{\\sigma\_\{IM\}^\{2\}\+\\sigma\_\{IA\}^\{2\}\+\\sigma\_\{IMA,e\}^\{2\}\}\{n\_\{i\}\}\([4](https://arxiv.org/html/2610.00651#S3.E4)\)ρAA′\(b\)\\rho^\{\(b\)\}\_\{AA^\{\\prime\}\}Do scaffolds inducethe same model order?σM2\+σIM2ni\\sigma\_\{M\}^\{2\}\+\\dfrac\{\\sigma\_\{IM\}^\{2\}\}\{n\_\{i\}\}σMA2\+σIMA,e2ni\\sigma\_\{MA\}^\{2\}\+\\dfrac\{\\sigma\_\{IMA,e\}^\{2\}\}\{n\_\{i\}\}\([5](https://arxiv.org/html/2610.00651#S3.E5)\)EρM2E\\rho^\{2\}\_\{M\},SNRM\\text\{SNR\}\_\{M\}Does the leaderboardpreserve model order?σM2\\sigma\_\{M\}^\{2\}σBM2nb\+σMA2na\+σBMA2nbna\+σIM\[B\]2nbni\+σBIMA,e2nbnina\\dfrac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\dfrac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\+\\dfrac\{\\sigma\_\{BMA\}^\{2\}\}\{n\_\{b\}n\_\{a\}\}\+\\dfrac\{\\sigma\_\{IM\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\dfrac\{\\sigma\_\{BIMA,e\}^\{2\}\}\{n\_\{b\}n\_\{i\}n\_\{a\}\}\([7](https://arxiv.org/html/2610.00651#S3.E7)\)
### 3\.3What Additional Evaluations Can Resolve
###### Proposition 3\.1\(Facet\-specific replication and reliability ceilings\)\.
For variance components, the reliability in Equations[3](https://arxiv.org/html/2610.00651#S3.E3)and[7](https://arxiv.org/html/2610.00651#S3.E7)is non\-decreasing in each sample size\. Holdingnbn\_\{b\}andnan\_\{a\}fixed,
limni→∞EρM\(b\)2\(ni,na\)=σM2σM2\+σMA2na,limni→∞EρM2\(nb,ni,na\)=σM2σM2\+σBM2nb\+σMA2na\+σBMA2nbna\.\\lim\_\{n\_\{i\}\\rightarrow\\infty\}E\\rho\_\{M\(b\)\}^\{2\}\(n\_\{i\},n\_\{a\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\frac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\},\\;\\lim\_\{n\_\{i\}\\rightarrow\\infty\}E\\rho\_\{M\}^\{2\}\(n\_\{b\},n\_\{i\},n\_\{a\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\frac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\+\\frac\{\\sigma\_\{BMA\}^\{2\}\}\{n\_\{b\}n\_\{a\}\}\}\.\(8\)
###### Proof\.
Increasingnin\_\{i\}monotonically decreases only denominator terms indexed by tasks\. Those terms converge to zero, while benchmark\- and scaffold\-indexed terms remain\. ∎
Thus, arbitrarily many tasks cannot overcome model–benchmark heterogeneity or model–scaffold coupling\. This result motivates evaluating diversity, rather than task count alone, as a design resource\.
###### Corollary 3\.2\(Benchmark breadth at a fixed task budget\)\.
Fixnan\_\{a\}and the total numberN=nbniN=n\_\{b\}n\_\{i\}of tasks per scaffold\. Under the balanced, exchangeable\-facet model,
σδ2=\\displaystyle\\sigma\_\{\\delta\}^\{2\}=\{\}σMA2na\+σBM2\+σBMA2/nanb\+σIM\[B\]2\+σBIMA,e2/naN\.\\displaystyle\\frac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\+\\frac\{\\sigma\_\{BM\}^\{2\}\+\\sigma\_\{BMA\}^\{2\}/n\_\{a\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{IM\[B\]\}^\{2\}\+\\sigma\_\{BIMA,e\}^\{2\}/n\_\{a\}\}\{N\}\.\(9\)Thus, distributing the same task budget across more benchmarks increases reliability wheneverσBM2\+σBMA2/na\>0\\sigma\_\{BM\}^\{2\}\+\\sigma\_\{BMA\}^\{2\}/n\_\{a\}\>0\.
###### Proof\.
Useni=N/nbn\_\{i\}=N/n\_\{b\}in Eq\.[7](https://arxiv.org/html/2610.00651#S3.E7)\. Only the benchmark\-conditioned term changes withnbn\_\{b\}\. ∎
Breadth does not create additional persistent model signal; it reduces contamination by condition\-specific advantages\. The corollary assumes similarly informative benchmark draws, adequate crossing, and no additional benchmark setup cost\. It does not imply that arbitrary new benchmarks outperform more tasks, or that benchmarks always offer greater value than scaffolds\.
### 3\.4Bayesian Estimation and Design Studies
We use Bayesian estimation with regularizing priors because sparse crossed designs can yield unstable variance estimates, especially near floor and ceiling performance\. For every posterior draw, we compute reliability, SNR, and the task\-only ceilings\. D\-study projections therefore retain uncertainty in the variance decomposition rather than substituting point estimates\. Partial pooling allows connected observations to inform common variance components while retaining uncertainty where scaffold or cross\-benchmark replication is limited\.
Primary D\-studies use common equal allocations; analyses reproducing the observed imbalance appear in Appendix[C](https://arxiv.org/html/2610.00651#A3)\. We measure evaluation volume in model–scaffold–task trials and use dashboard prices for dollar\-cost projections\. Repeated task subsampling checks whether reduced designs retain the ranking information predicted by the D\-study \(see §[F\.1](https://arxiv.org/html/2610.00651#A6.SS1)\)\.
### 3\.5Posterior Capability, Ranks, and Transportability
For benchmarkbb, the scaffold\-marginalized latent capability of modelmmin posterior drawssisθmb\(s\)=um\(M,s\)\+ubm\(BM,s\)\\theta\_\{mb\}^\{\(s\)\}=u\_\{m\}^\{\(M,s\)\}\+u\_\{bm\}^\{\(BM,s\)\}\. To test whether the estimated shared model component is transportable rather than an artifact of the nine\-benchmark panel, we evaluateθ^m=um\(M\)\\widehat\{\\theta\}\_\{m\}=u\_\{m\}^\{\(M\)\}on four contemporaneous benchmarks excluded from model estimation\. Among overlapping models, we compare its association with external performance against that of the conventional in\-panel mean\-accuracy aggregate\. This is an out\-of\-panel test of convergent predictive validity, not evidence of universal deployment validity nor internal construct validity\. Finally, we assess sensitivity to the link function, estimation method, variance\-component estimator, prior specification, and leave\-one\-benchmark\-out refits\. The latter analysis also identifies benchmarks that disproportionately contribute signal or connectivity to the leaderboard\. Appendices have full computational estimation details \(§[C](https://arxiv.org/html/2610.00651#A3)\), robustness checks and sensitivity analyses \(§[D](https://arxiv.org/html/2610.00651#A4)\), and methodological comparisons \(§[G](https://arxiv.org/html/2610.00651#A7)\)\.
## 4Data
We analyze agent rollouts from all nine benchmarks distributed via HAL[Kapoor et al\. \(2024\)](https://arxiv.org/html/2610.00651#bib.bib5): AssistantBench\([Yoran et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib23)\), CoreBench Hard\([Siegel et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib20)\), GAIA\([Mialon et al\., 2023](https://arxiv.org/html/2610.00651#bib.bib18)\), Online\-Mind2Web\([Xue et al\., 2025](https://arxiv.org/html/2610.00651#bib.bib45)\), SciCode\([Tian et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib21)\), ScienceAgentBench\([Chen et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib16)\), SWE\-bench Verified Mini\([Jimenez et al\., 2023](https://arxiv.org/html/2610.00651#bib.bib17)\),τ\\tau\-bench Airline\([Yao et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib22)\), and USACO\([Shi et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib19)\), covering web navigation, scientific programming tasks, multi\-step and user assistance, software engineering, customer\-service interaction, and competitive programming \(details in Appendix[A\.1](https://arxiv.org/html/2610.00651#A1.SS1)\)\. From each of AssistantBench, GAIA, SciCode and SWE\-bench Verified Mini, the HAL leaderboard uses a subset of tasks from the full benchmark\([Kapoor et al\., 2025](https://arxiv.org/html/2610.00651#bib.bib4)\)\.
HAL dataset covers29,92329\{,\}923total agent rollouts across 54 models and 13 scaffolds; 9 models appear on every benchmark each of which span at least 69% of the available scaffolds\. Each rollout is over a single task, model, reasoning\-effort, and scaffold\. Summary statistics are found in Table[2](https://arxiv.org/html/2610.00651#S4.T2)\. All outcomes are task\-level binary scores\. The incidence structure is incomplete at every level, a common problem in leaderboards\([Singh et al\., 2026](https://arxiv.org/html/2610.00651#bib.bib44)\)\. One scaffold is shared across eight benchmarks; AssistantBench connects the remaining benchmark through an additional shared scaffold; all other benchmarks contain at least one additional benchmark\-specific scaffold\. Models and scaffolds are neither fully crossed nor evenly replicated\. Nevertheless, shared models and scaffolds connect the observation graph, permitting partial separation of model, benchmark, and scaffold effects\.
Additionally, to corroborate our claims and provide external convergent validity, we use two supplemental datasets\. The Harbor Index dataset is a meta\-benchmark containing task level scores from 29 agent benchmarks for 9 LLMs and 4 scaffolds\. The second dataset consists of LLM\-level scores for four benchmarks that are contemporaneous with and include LLMs specified in the HAL data\. Appendices more detail for the datasets \(§[A](https://arxiv.org/html/2610.00651#A1)\) and data connectivity analysis \(§[E](https://arxiv.org/html/2610.00651#A5)\)\. These data represent current practice: because agent evaluations are expensive,many model–scaffold–benchmark combinations absent entirely\.
Table 2:HAL dataset design summary\. Counts denote unique levels of each facet\.BenchmarkModelsTasksAgent ScaffoldsModel–Scaffold PairsAssistantBench1833230CORE\-Bench Hard3445356GAIA21165234Online\-Mind2Web13300223SciCode1765337ScienceAgentBench19102225SWE\-bench Verified Mini2450226τ\\tau\-bench Airline2350342USACO13307214
## 5Results and Practical Recommendations
The results identify a structural limitation of task\-only scaling: current evaluations can distinguish fixed systems precisely while leaving persistent model differences unresolved\. The useful response is to change the measurement design, not merely enlarge it\.
### 5\.1Reliability depends on the measurement goal
Reliability depends on the measurement claim a practitioner wants to make\. In agent evaluations, two choices are relevant here: whether the object of measurement is the underlying model or the model–scaffold system, and whether the goal is relative \(rank\) or absolute interpretation of scores\.
##### Model and system reliability can differ substantially\.
Across the nine benchmarks, estimated system reliability \(i\.e\., ranking model–scaffold pairs\) isEρMA\(b\)2∈\[0\.935,0\.994\]E\\rho\_\{MA\(b\)\}^\{2\}\\in\[0\.935,0\.994\], whereas model reliability isEρM\(b\)2∈\[0\.148,0\.841\]E\\rho\_\{M\(b\)\}^\{2\}\\in\[0\.148,0\.841\]\(Figure[1](https://arxiv.org/html/2610.00651#S5.F1)\)\. These are not conflicting assessments of the same score\. System reliability counts scaffold differences and compatibility as signal; model reliability requires differences that persist after averaging over scaffolds\. A leaderboard can therefore be reliable for ranking model\-scaffold pairs but unreliable for solely ranking the underlying model\.
Figure 1:Reliability depends on the object being ranked & task replication cannot remove scaffold\-dependent model advantages\.D\-study reliability as tasks are added for base models and fixed model–scaffold systems\. Dashed lines are task\-only asymptotes for model rankings\. Estimates are posterior medians with68%68\\%standard error HDIs\. See also Table[16](https://arxiv.org/html/2610.00651#A6.T16)\.
##### Recommendation\.
Specify whether the intended object of measure is the model or a model\-\-scaffold system, especially in leaderboards using multiple agentic benchmarks, and whether the intended interpretation is rank comparison \(or whether absolute scores are intended to carry meaning, such as capability thresholds\.444There may be situations where reliability of actual numeric scores \(not just ranking\) is needed\. We illustrate this separate estimation on the observed scale in §[G\.5](https://arxiv.org/html/2610.00651#A7.SS5), which shows that score reliability is substantially lower for all benchmarks\. Reliablerankingsdo not imply reliablescores\.Provide rankings, calculate scores, and estimate reliability for that specific claim rather than reporting a single generic reliability statistic\.
### 5\.2Changing the scaffold can change evaluation conclusions
##### How much the scaffold matters depends on the benchmark\.
Inter\-scaffold reliability \(Eq\.[5](https://arxiv.org/html/2610.00651#S3.E5)\) measures whether different scaffolds produce similar rankings when evaluating the same models on the same tasks ranging \(with 95% HDI\) from 0\.151 \[0\.004,0\.577\] on OnlineMind2Web to 0\.852 \[0\.631,0\.947\] on CORE\-Bench Hard \(Figure[2](https://arxiv.org/html/2610.00651#S5.F2)\)\. At the low end, changing only the scaffold can substantially change which models appear to perform best on a benchmark, even when the models and tasks remain fixed\.
##### Scaffolds can change which tasks get solved without improving the overall benchmark score\.
The pooled decomposition \(Eq\.[6](https://arxiv.org/html/2610.00651#S3.E6)\) distinguishes persistent differences from task\-specific sensitivity\. In the pooled analysis, the contrastσM2−σA2\\sigma\_\{M\}^\{2\}\-\\sigma\_\{A\}^\{2\}favors neither direction \(posterior directional probability approximately50%50\\%\)\([Makowski et al\., 2019](https://arxiv.org/html/2610.00651#bib.bib43)\)\. However, task\-specific variation across scaffolds is larger than task\-specific variation across models with posterior probabilityPr\(σIA\[B\]2\>σIM\[B\]2∣y\)=84%\\Pr\\\!\(\\sigma\_\{IA\[B\]\}^\{2\}\>\\sigma\_\{IM\[B\]\}^\{2\}\\mid y\)=84\\%\(see Figure[10](https://arxiv.org/html/2610.00651#A6.F10)\)\. Benchmark specific decompositions \(Eq\.[2](https://arxiv.org/html/2610.00651#S3.E2)\) show the same pattern \(Figure[2](https://arxiv.org/html/2610.00651#S5.F2), middle and right\) This suggests that scaffold choice can strongly affect success on individual tasks, even without making one scaffold producing uniformly better on a benchmark as a whole\. Similarly, A scaffold can change which tasks a system solves without producing a uniformly stronger system\.
Figure 2:Scaffolds are rank\-relevant and result in distinct task\-level consequences\.Left: posterior inter\-scaffold reliabilityρAA′\(b\)\\rho\_\{AA^\{\\prime\}\}^\{\(b\)\}\.MiddleandRight: model\-minus\-scaffold system rank\-relevant variance contrasts \(over denominator from Eq\.[4](https://arxiv.org/html/2610.00651#S3.E4)\) at the benchmark and task levels, respectively; positive values favor the model contribution\. Task\-interaction contrasts measure task\-specific sensitivity\. Points and bars denote posterior means and medians, respectively, with68%68\\%standard error HDIs\.
##### Recommendation\.
If the goal is to evaluate the underlying model, test the same models across multiple scaffolds rather than relying on a single implementation\. Report how much conclusions change across scaffolds, and avoid attributing scaffold\-specific advantages to the model itself\. If the deployed model–scaffold system is the intended object of evaluation, scaffold variation can instead be treated as part of the system being measured\.
### 5\.3More tasks cannot always resolve evaluation uncertainty
##### More tasks only help when the main uncertainty comes from differences across tasks\.
Adding tasks helps when a model’s measured performance changes substantially depending on which tasks it is tested on \(Proposition[3\.1](https://arxiv.org/html/2610.00651#S3.Thmtheorem1)\)\. It does not fix uncertainty caused by other facets, such as the choice of scaffold\. In our data, even infinitely many similarly constructed tasks would improve model\-ranking reliability by at most0\.100\.10\. Only CORE\-Bench Hard and SciCode can exceedEρ2=0\.75E\\rho^\{2\}=0\.75through task scaling alone; for OnlineMind2Web, reliability increases only from0\.1480\.148to0\.1530\.153\.
##### A benchmark can run out of useful information before it runs out of tasks\.
Adding tasks repeatedly measures the same model–scaffold and benchmark\-specific effects\. Once these sources of uncertainty dominate, more tasks make the existing evaluation setup more precise without making the broader model claim substantially more reliable\. Low reliability does not mean that a benchmark measures an unimportant capability; it means that, for the models being compared, its scores do not reliably distinguish the quantity of interest\. Thus, reliability should be re\-estimated as the competitor population changes\.
##### Recommendation\.
Estimate how much reliability can improve from adding tasks before expanding a benchmark\. Add tasks when task sampling is the main source of uncertainty; otherwise, spend evaluation budget on relevant facets that drive uncertainty, such as broader scaffold coverage\.
\(a\)Model reliability\.\(b\)Rank\-relevant signal\-to\-noise ratio\.
Figure 3:Broader measurement conditions reduce error that an increase in the number of test items alone cannot\.Pooled D\-studies compare concentration within one benchmark against allocation across multiple benchmarks\. Estimates are posterior medians; panel \(a\) includes68%68\\%standard error HDIs\. HAL and Harbor are fitted separately\. The reference bands\[2,3\]\[2,3\]represent minimal detection limits used in laboratory measurement\([CDER, 2024](https://arxiv.org/html/2610.00651#bib.bib38);[Sheehan and Yost, 2026](https://arxiv.org/html/2610.00651#bib.bib39);[Taleuzzaman, 2018](https://arxiv.org/html/2610.00651#bib.bib37)\), and are descriptive rather than benchmark calibrated thresholds\.
### 5\.4Pooling diverse benchmarks can make model rankings more reliable
Agentic leaderboards often consist of multiple benchmarks from which model agentic capability is to be inferred\. If the goal is to rank models on their ability to perform on diverse agentic tasks, pooling benchmarks that test different kinds of tasks provides more information about which performance differences persist across settings\. At a fixed task budget, benchmark breadth averages condition\-specific model advantages that within\-benchmark replication leaves untouched \(Corollary[3\.2](https://arxiv.org/html/2610.00651#S3.Thmtheorem2)\)\.
Table 3:Agreement with external agent benchmarks\.Kendall’sτ\\taufor the unweighted HAL mean and reliability\-adjusted latent \(θ^\\widehat\{\\theta\}\) model effect \(median posterior\) with\[95%\]\[95\\%\]CIs\.External BenchmarkMean scoreθ^\\widehat\{\\theta\}DifferenceLLM ObservationsBFCL v40\.429\[−0\.282,0\.836\]0\.429\\,\[\-0\.282,\\,0\.836\]0\.714\[0\.147,0\.928\]\\mathbf\{0\.714\}\\,\[0\.147,\\,0\.928\]\+0\.285\+0\.2857Terminal\-Bench 2\.00\.524\[−0\.165,0\.869\]0\.524\\,\[\-0\.165,\\,0\.869\]0\.810\[0\.361,0\.954\]\\mathbf\{0\.810\}\\,\[0\.361,\\,0\.954\]\+0\.286\+0\.2867SWE\-bench Verified0\.565\[0\.306,0\.746\]0\.565\\,\[0\.306,\\,0\.746\]0\.765\[0\.595,0\.870\]\\mathbf\{0\.765\}\\,\[0\.595,\\,0\.870\]\+0\.200\+0\.20020τ2\\tau^\{2\}\-bench Core0\.333\[−0\.515,0\.852\]0\.333\\,\[\-0\.515,\\,0\.852\]0\.867\[0\.383,0\.977\]\\mathbf\{0\.867\}\\,\[0\.383,\\,0\.977\]\+0\.534\+0\.5346Meanτ\\tau0\.4630\.4630\.789\\mathbf\{0\.789\}\+0\.326\+0\.326—
##### Pooling benchmarks can achieve more reliable model rankings at lower cost\.
At the same task budget, distributing evaluations across the nine\-benchmark battery raises projected model\-ranking reliability from approximately 0\.44 for a single benchmark to 0\.75 \(Figure[3](https://arxiv.org/html/2610.00651#S5.F3)\)\. A complementary pooled analysis of the Harbor Index\([Shi et al\., 2026](https://arxiv.org/html/2610.00651#bib.bib49)\)\(Appendix[A\.2](https://arxiv.org/html/2610.00651#A1.SS2)\) supports the same qualitative pattern\. Moreover, the full HAL battery costs more than \($47,000\), while a balanced allocation achieves comparable reliability for approximately \($19,000\)\. For an illustrative target of \(SNR=2\.5,Eρ2≈0\.71SNR=2\.5,E\\rho^\{2\}\\approx 0\.71\), approximately \(14\) tasks per benchmark cost \($8,144\), an estimated \(83%\) reduction\. In Appendix Table[19](https://arxiv.org/html/2610.00651#A6.T19), we corroborate the diminishing returns using task subsampling\.
##### Task diversity leads to better model rank generalization\.
Model effect estimatesθ^\\hat\{\\theta\}also agree more closely with all four held\-out agent benchmarks than a simple average of benchmark scores \(mean Kendall’sτ\\tau:0\.7890\.789vs\.0\.4630\.463; Table[3](https://arxiv.org/html/2610.00651#S5.T3); see §[A\.3](https://arxiv.org/html/2610.00651#A1.SS3)for data and contamination prevention\)\. Together, these results suggest that pooling diverse benchmarks can help separate model differences that recur across settings from advantages specific to a particular benchmark or evaluation condition\. These results provide convergent predictive evidence, not proof of a universal one\-dimensional agent capability\.555Internally, the pooled estimation strongly outperforms per\-benchmark estimations in correlations of posterior rankings with benchmark\-wise observed scores, \[0\.62,0\.96\] and \[0\.08,0\.25\] respectively \(Table[10](https://arxiv.org/html/2610.00651#A4.T10), §[D\.2](https://arxiv.org/html/2610.00651#A4.SS2)\)\. Yet posterior rank intervals remain wide; many neighboring models have ordering probabilities near1/21/2\(Figure[13](https://arxiv.org/html/2610.00651#A6.F13)\), with some models moving across rank quartiles between reported and adjusted rankings \(Figures[11](https://arxiv.org/html/2610.00651#A6.F11)and[12](https://arxiv.org/html/2610.00651#A6.F12)\)\.
##### Breadth is most useful when the design is connected with informative benchmarks\.
For these data, CORE\-Bench Hard contains the greatest SNR \(see Figure[7](https://arxiv.org/html/2610.00651#A6.F7), §[D\.4](https://arxiv.org/html/2610.00651#A4.SS4)\), particularly with its own scaffold, CORE Agent \(§[D\.5](https://arxiv.org/html/2610.00651#A4.SS5)\)\. Removing this strongest discriminating signal increases pooled uncertainty more than any other LOO removals\. Thus, having a benchmark with high model signal can support anchoring a common scale through sufficient connectivity to other benchmarks\.
##### Recommendation\.
When the intended claim concerns performance across different kinds of agentic tasks, pool benchmarks that represent that range rather than evaluating each benchmark exhaustively\. Use the reliability analysis to determine how much evaluation is needed within each benchmark while preserving enough task diversity to support the broader cross\-task claim\. For imbalanced designs, prioritize identifying informative anchors to bridge conditions\.
## 6Conclusion
Agent evaluations should not be treated as having a single, intrinsic level of reliability\. Reliability depends on what practitioners want to learn from them: the underlying model or a complete model–scaffold system, a ranking or an absolute score, and performance on one benchmark or across a broader range of agentic tasks\. Our results show that these distinctions matter in practice and hence that evaluation design should follow the intended claim: Practitioners should first specify what they want a score or ranking to mean, then identify which sources of uncertainty prevent the evaluation from supporting that claim\. More evaluation is useful only when it addresses those sources\. Reliability analysis can therefore guide not only how evaluation results are interpreted, but also where additional evaluation effort is most valuable for differentiating a target model population\.
### Reproducibility statement
To facilitate reproducibility, the code and data are available online,666[https://github\.com/hardy\-education/scaffold\_eval](https://github.com/hardy-education/scaffold_eval)original data sources are linked in Appendix[A](https://arxiv.org/html/2610.00651#A1), and computation and estimation details are found in Appendix[C](https://arxiv.org/html/2610.00651#A3)\.
### AI use statement
AI was used during the writing phase of this study\. GPT 5\.6 Sol was used after initial drafts to reduce the text length of several sections in the main body, all of which required further editing after use\. Paragraphs describing tabular results in Appendices[D\.4](https://arxiv.org/html/2610.00651#A4.SS4)–[D\.6](https://arxiv.org/html/2610.00651#A4.SS6)were revised using GPT 5\.5\. Drafts of several sections of Appendix[E](https://arxiv.org/html/2610.00651#A5)were created from research notes using GPT 5\.5, which were revised and then rewritten using human\-edited combinations of text from GPT 5\.5, GPT 5\.6 Sol, and Claude Opus 4\.8\. Finally, the 100% human\-written code used in this study was refactored with expanded comments and validated against paper findings \(under subsampling\) to support reproducibility using Claude Opus 4\.8 with Claude Code\.
## References
- AnthropicSystem Card: Claude Opus 4\.7\.System CardAnthropic\.Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p1.1)\.
- Barreset al\.\(2025\)V\. Barres, H\. Dong, S\. Ray, X\. Si, and K\. Narasimhan$\\tau^2$\-Bench: Evaluating Conversational Agents in a Dual\-Control Environment\.\(en\)\.External Links:[Link](https://arxiv.org/abs/2506.07982v1)Cited by:[4th item](https://arxiv.org/html/2610.00651#A1.I1.i4.p1.1)\.
- Bateset al\.\(2015\)D\. Bates, M\. Mächler, B\. Bolker, and S\. WalkerFitting Linear Mixed\-Effects Models Usinglme4\.Journal of Statistical Software67\(1\) \(en\)\.External Links:ISSN 1548\-7660,[Link](http://www.jstatsoft.org/v67/i01/),[Document](https://dx.doi.org/10.18637/jss.v067.i01)Cited by:[§C\.4](https://arxiv.org/html/2610.00651#A3.SS4.p1.1)\.
- Bindoff \(2026\)A\. D\. BindoffPartial pooling predicts cross\-validation reliability: a closed\-form triage and Rao\-Blackwellised cure for hierarchical LOO\.arXiv\.Note:arXiv:2607\.18836 \[stat\.ME\]External Links:[Link](http://arxiv.org/abs/2607.18836),[Document](https://dx.doi.org/10.48550/arXiv.2607.18836)Cited by:[§D\.3](https://arxiv.org/html/2610.00651#A4.SS3.p3.1)\.
- Brennan \(2001\)R\. L\. BrennanGeneralizability Theory\.Springer,New York, NY\(en\)\.External Links:ISBN 978\-1\-4419\-2938\-9 978\-1\-4757\-3456\-0,[Link](http://link.springer.com/10.1007/978-1-4757-3456-0),[Document](https://dx.doi.org/10.1007/978-1-4757-3456-0)Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p4.1),[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px2.p1.1)\.
- Bürkner \(2021\)P\. BürknerBayesian Item Response Modeling in R with brms and Stan\.Journal of Statistical Software100,pp\. 1–54\(en\)\.External Links:ISSN 1548\-7660,[Link](https://doi.org/10.18637/jss.v100.i05),[Document](https://dx.doi.org/10.18637/jss.v100.i05)Cited by:[§C\.1\.3](https://arxiv.org/html/2610.00651#A3.SS1.SSS3.p2.1)\.
- CDER \(2024\)CDERQ2\(R2\) Validation of Analytical Procedures\.U\.S\. Department of Health and Human Services Food and Drug Administration\(en\)\.External Links:[Link](https://www.regulations.gov/document/FDA-2022-D-1503-0002)Cited by:[Figure 3](https://arxiv.org/html/2610.00651#S5.F3)\.
- Chenet al\.\(2024\)Z\. Chen, S\. Chen, Y\. Ning, Q\. Zhang, B\. Wang, B\. Yu, Y\. Li, Z\. Liao, C\. Wei, Z\. Lu,et al\.Scienceagentbench: toward rigorous assessment of language agents for data\-driven scientific discovery\.arXiv preprint arXiv:2410\.05080\.Cited by:[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Cronbachet al\.\(1972\)L\. J\. Cronbach, G\. Gleser, H\. Nanda, and N\. RajaratnamThe Dependability of behavioral measurements: theory of generalizability for scores and profiles\.Wiley,New York\(en\)\.External Links:ISBN 978\-0\-471\-18850\-6Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p4.1),[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px2.p1.1)\.
- Cronbach and Meehl \(1955\)L\. J\. Cronbach and P\. E\. MeehlConstruct validity in psychological tests\.Psychological Bulletin52\(4\),pp\. 281–302\.External Links:ISSN 1939\-1455,[Document](https://dx.doi.org/10.1037/h0040957)Cited by:[Reliability is necessary but not sufficient for valid evaluation\.](https://arxiv.org/html/2610.00651#Ax2.SSx2.SSS0.Px1.p1.1)\.
- Cronbach and Shavelson \(2004\)L\. J\. Cronbach and R\. J\. ShavelsonMy Current Thoughts on Coefficient Alpha and Successor Procedures\.Educational and Psychological Measurement64\(3\),pp\. 391–418\(EN\)\.External Links:ISSN 0013\-1644,[Link](https://doi.org/10.1177/0013164404266386),[Document](https://dx.doi.org/10.1177/0013164404266386)Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px2.p1.1)\.
- Fox and Glas \(2001\)J\. Fox and C\. A\. W\. GlasBayesian estimation of a multilevel IRT model using gibbs sampling\.Psychometrika66\(2\),pp\. 271–288\(en\)\.External Links:ISSN 1860\-0980,[Link](https://doi.org/10.1007/BF02294839),[Document](https://dx.doi.org/10.1007/BF02294839)Cited by:[§3\.2](https://arxiv.org/html/2610.00651#S3.SS2.p1.1)\.
- Hardy and Kim \(2026\)M\. Hardy and Y\. KimKnowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact\.arXiv\.Note:arXiv:2603\.00883 \[cs\]External Links:[Link](http://arxiv.org/abs/2603.00883),[Document](https://dx.doi.org/10.48550/arXiv.2603.00883)Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px2.p1.1)\.
- Hardyet al\.\(2026\)M\. Hardy, A\. Reuel, L\. Zhang, J\. M\. Casabianca, S\. Truong, Y\. Dave, H\. Lee, B\. Domingue, and S\. KoyejoAI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems\.arXiv\.Note:arXiv:2605\.25272 \[cs\.AI\]External Links:[Link](http://arxiv.org/abs/2605.25272),[Document](https://dx.doi.org/10.48550/arXiv.2605.25272)Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px2.p1.1)\.
- Harrison \(2015\)X\. A\. HarrisonA comparison of observation\-level random effect and Beta\-Binomial models for modelling overdispersion in Binomial data in ecology & evolution\.PeerJ3,pp\. e1114\.External Links:ISSN 2167\-8359,[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC4517959/),[Document](https://dx.doi.org/10.7717/peerj.1114)Cited by:[§D\.3\.1](https://arxiv.org/html/2610.00651#A4.SS3.SSS1.p1.2)\.
- Heet al\.\(2024\)H\. He, W\. Yao, K\. Ma, W\. Yu, Y\. Dai, H\. Zhang, Z\. Lan, and D\. YuWebvoyager: building an end\-to\-end web agent with large multimodal models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6864–6890\.Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1)\.
- Jimenezet al\.\(2023\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. R\. NarasimhanSwe\-bench: can language models resolve real\-world github issues?\.InThe twelfth international conference on learning representations,Cited by:[3rd item](https://arxiv.org/html/2610.00651#A1.I1.i3.p1.1),[§1](https://arxiv.org/html/2610.00651#S1.p1.1),[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Kapooret al\.\(2025\)S\. Kapoor, B\. Stroebl, P\. Kirgis, N\. Nadgir, Z\. S\. Siegel, B\. Wei, T\. Xue, Z\. Chen, F\. Chen, S\. Utpala, F\. Ndzomga, D\. Oruganty, S\. Luskin, K\. Liu, B\. Yu, A\. Arora, D\. Hahm, H\. Trivedi, H\. Sun, J\. Lee, T\. Jin, Y\. Mai, Y\. Zhou, Y\. Zhu, R\. Bommasani, D\. Kang, D\. Song, P\. Henderson, Y\. Su, P\. Liang, and A\. NarayananHolistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation\.arXiv\.Note:arXiv:2510\.11977 \[cs\]External Links:[Link](http://arxiv.org/abs/2510.11977),[Document](https://dx.doi.org/10.48550/arXiv.2510.11977)Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p4.1),[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Kapooret al\.\(2024\)S\. Kapoor, B\. Stroebl, Z\. S\. Siegel, N\. Nadgir, and A\. NarayananAI Agents That Matter\.\(en\)\.External Links:[Link](https://arxiv.org/abs/2407.01502v1)Cited by:[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Kumar and Mishra \(2025\)P\. Kumar and S\. MishraRobustness in large language models: a survey of mitigation strategies and evaluation metrics\.arXiv preprint arXiv:2505\.18658\.Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1)\.
- Linacre and Wright \(2002\)J\. Linacre and B\. WrightUnderstanding Rasch measurement: Construction of measures from many\-facet data\.Journal of applied measurement3,pp\. 486–512\.Cited by:[§3\.2](https://arxiv.org/html/2610.00651#S3.SS2.p1.1)\.
- Madaanet al\.\(2024\)L\. Madaan, A\. K\. Singh, R\. Schaeffer, A\. Poulton, S\. Koyejo, P\. Stenetorp, S\. Narang, and D\. HupkesQuantifying Variance in Evaluation Benchmarks\.arXiv\.Note:arXiv:2406\.10229External Links:[Link](http://arxiv.org/abs/2406.10229),[Document](https://dx.doi.org/10.48550/arXiv.2406.10229)Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px2.p1.1)\.
- Makowskiet al\.\(2019\)D\. Makowski, M\. S\. Ben\-Shachar, S\. H\. A\. Chen, and D\. LüdeckeIndices of Effect Existence and Significance in the Bayesian Framework\.Frontiers in Psychology10\(English\)\.External Links:ISSN 1664\-1078,[Link](https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.02767/full),[Document](https://dx.doi.org/10.3389/fpsyg.2019.02767)Cited by:[§5\.2](https://arxiv.org/html/2610.00651#S5.SS2.SSS0.Px2.p1.1)\.
- Merrillet al\.\(2026\)M\. A\. Merrill, A\. G\. Shaw, N\. Carlini, B\. Li, H\. Raj, I\. Bercovich, L\. Shi, J\. Y\. Shin, T\. Walshe, E\. K\. Buchanan, J\. Shen, G\. Ye, H\. Lin, J\. Poulos, M\. Wang, M\. Nezhurina, J\. Jitsev, D\. Lu, O\. M\. Mastromichalakis, Z\. Xu, Z\. Chen, Y\. Liu, R\. Zhang, L\. L\. Chen, A\. Kashyap, J\. Uslu, J\. Li, J\. Wu, M\. Yan, S\. Bian, V\. Sharma, K\. Sun, S\. Dillmann, A\. Anand, A\. Lanpouthakoun, B\. Koopah, C\. Hu, E\. Guha, G\. H\. S\. Dreiman, J\. Zhu, K\. Krauth, L\. Zhong, N\. Muennighoff, R\. Amanfu, S\. Tan, S\. Pimpalgaonkar, T\. Aggarwal, X\. Lin, X\. Lan, X\. Zhao, Y\. Liang, Y\. Wang, Z\. Wang, C\. Zhou, D\. Heineman, H\. Liu, H\. Trivedi, J\. Yang, J\. Lin, M\. Shetty, M\. Yang, N\. Omi, N\. Raoof, S\. Li, T\. Y\. Zhuo, W\. Lin, Y\. Dai, Y\. Wang, W\. Chai, S\. Zhou, D\. Wahdany, Z\. She, J\. Hu, Z\. Dong, Y\. Zhu, S\. Cui, A\. Saiyed, A\. Kolbeinsson, J\. Hu, C\. M\. Rytting, R\. Marten, Y\. Wang, A\. Dimakis, A\. Konwinski, and L\. SchmidtTerminal\-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces\.arXiv\(en\)\.Note:Version Number: 1External Links:[Link](https://arxiv.org/abs/2601.11868),[Document](https://dx.doi.org/10.48550/ARXIV.2601.11868)Cited by:[2nd item](https://arxiv.org/html/2610.00651#A1.I1.i2.p1.1)\.
- Mialonet al\.\(2023\)G\. Mialon, C\. Fourrier, T\. Wolf, Y\. LeCun, and T\. ScialomGaia: a benchmark for general ai assistants\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- OpenAI \(2026\)OpenAIGPT\-5\.5 System Card\.System CardOpenAI\.Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p1.1)\.
- Paananenet al\.\(2021\)T\. Paananen, J\. Piironen, P\. Bürkner, and A\. VehtariImplicitly adaptive importance sampling\.Statistics and Computing31\(2\),pp\. 16\(en\)\.External Links:ISSN 1573\-1375,[Link](https://doi.org/10.1007/s11222-020-09982-2),[Document](https://dx.doi.org/10.1007/s11222-020-09982-2)Cited by:[§D\.3](https://arxiv.org/html/2610.00651#A4.SS3.p1.1)\.
- Patilet al\.\(2025\)S\. G\. Patil, H\. Mao, F\. Yan, C\. C\. Ji, V\. Suresh, I\. Stoica, and J\. E\. GonzalezThe Berkeley Function Calling Leaderboard \(BFCL\): From Tool Use to Agentic Evaluation of Large Language Models\.InProceedings of the 42nd International Conference on Machine Learning,pp\. 48371–48392\(en\)\.External Links:ISSN 2640\-3498,[Link](https://proceedings.mlr.press/v267/patil25a.html)Cited by:[1st item](https://arxiv.org/html/2610.00651#A1.I1.i1.p1.1)\.
- Rabanseret al\.\(2026\)S\. Rabanser, S\. Kapoor, P\. Kirgis, K\. Liu, S\. Utpala, and A\. NarayananTowards a Science of AI Agent Reliability\.arXiv\.Note:arXiv:2602\.16666 \[cs\]External Links:[Link](http://arxiv.org/abs/2602.16666),[Document](https://dx.doi.org/10.48550/arXiv.2602.16666)Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p3.1),[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1)\.
- Razaviet al\.\(2025\)A\. Razavi, M\. Soltangheis, N\. Arabzadeh, S\. Salamat, M\. Zihayat, and E\. BagheriBenchmarking prompt sensitivity in large language models\.ArXivabs/2502\.06065\.External Links:[Link](https://api.semanticscholar.org/CorpusID:276249536)Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p3.1),[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1)\.
- Rizzo and Székely \(2010\)M\. L\. Rizzo and G\. J\. SzékelyDISCO analysis: A nonparametric extension of analysis of variance\.The Annals of Applied Statistics4\(2\)\.Note:arXiv:1011\.2288 \[stat\]External Links:ISSN 1932\-6157,[Link](http://arxiv.org/abs/1011.2288),[Document](https://dx.doi.org/10.1214/09-AOAS245)Cited by:[§C\.5](https://arxiv.org/html/2610.00651#A3.SS5.p1.1),[§G\.2](https://arxiv.org/html/2610.00651#A7.SS2.p1.1)\.
- Salaudeenet al\.\(2025\)O\. Salaudeen, A\. Reuel, A\. Ahmed, S\. Bedi, Z\. Robertson, S\. Sundar, B\. Domingue, A\. Wang, and S\. KoyejoMeasurement to Meaning: A Validity\-Centered Framework for AI Evaluation\.arXiv\.Note:arXiv:2505\.10573 \[cs\]External Links:[Link](http://arxiv.org/abs/2505.10573),[Document](https://dx.doi.org/10.48550/arXiv.2505.10573)Cited by:[Reliability is necessary but not sufficient for valid evaluation\.](https://arxiv.org/html/2610.00651#Ax2.SSx2.SSS0.Px1.p1.1)\.
- Shavelsonet al\.\(1989\)R\. J\. Shavelson, N\. M\. Webb, and G\. L\. RowleyGeneralizability theory\.American Psychologist44\(6\),pp\. 922–932\.External Links:ISSN 1935\-990X,[Document](https://dx.doi.org/10.1037/0003-066X.44.6.922)Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px2.p1.1)\.
- Sheehan and Yost \(2026\)T\. L\. Sheehan and R\. A\. YostWhat’s the Most Meaningful Standard for Mass Spectrometry: Instrument Detection Limit or Signal\-to\-Noise Ratio? \| Spectroscopy Online\.\(en\)\.External Links:[Link](https://www.spectroscopyonline.com/view/what-s-most-meaningful-standard-mass-spectrometry-instrument-detection-limit-or-signal-noise-ratio-0)Cited by:[Figure 3](https://arxiv.org/html/2610.00651#S5.F3)\.
- Shiet al\.\(2026\)L\. Shi, H\. Lin, Z\. Zhu, X\. Zhou, X\. Li, X\. Lin, Y\. Deng, H\. Xu, Y\. Li, S\. Li, Z\. Chen, H\. Xing, H\. Raj, B\. Chen, Q\. Shi, S\. Dillmann, Y\. Gao, P\. Khanna, R\. Lu, C\. B\. Zhou, M\. Yang, R\. Zhang, S\. Chai, J\. Chang, Y\. Chen, X\. Chen, Y\. Dai, W\. Yang, H\. Liu, M\. Liu, Z\. Wang, A\. E\. Assadi, B\. Stroebl, E\. K\. Buchanan, H\. Meng, J\. He, L\. Yu, R\. Shayanfar, Y\. Lee, Z\. Dong, A\. G\. Hart, A\. Wei, A\. Kashyap, A\. Khatua, A\. J\. Zheng, C\. Ma, D\. Heineman, D\. Chen, H\. Trinh, H\. Fang, H\. Zhang, H\. Shen, I\. Sugiura, J\. Sun, J\. Gao, J\. Lin, J\. Li, K\. Yang, L\. Hsiung, M\. Wang, M\. Tang, N\. Omi, N\. Raoof, N\. Edwards, O\. Guo, O\. M\. Mastromichalakis, P\. Ji, P\. Hejman, Q\. Qi, Q\. Lin, R\. Zhuang, R\. Yang, R\. Zheng, R\. Marten, S\. Fazliani, S\. Hou, S\. Jiang, S\. Li, B\. Yuan, M\. Glass, S\. Bian, T\. Y\. Zhuo, T\. Wu, T\. Tang, W\. Zhao, W\. Xuan, W\. Liang, X\. Liu, X\. Lan, X\. Zhang, X\. Zhao, Y\. Tang, Y\. Jiang, Y\. Li, Y\. Guan, Y\. Li, Y\. Liu, Y\. Tang, Yujun, Mao, Y\. Zhao, Y\. Wang, Y\. Tang, Z\. Tang, Z\. Li, Z\. Wang, Z\. She, K\. Liu, I\. Chaabane, Y\. Tang, X\. Li, S\. S\. S\. N\. GNVV, X\. Zheng, A\. Konwinski, B\. Li, L\. L\. Chen, A\. Dimakis, N\. Carlini, S\. Vosoughi, S\. Koyejo, D\. He, E\. Guha, B\. Feuer, M\. Merrill, L\. Schmidt, and A\. ShawHarbor Adapters and Harbor\-Index: Infrastructure and a Curated Meta\-Dataset for Large\-Scale Agentic Evaluation\.arXiv\(en\)\.Note:Version Number: 3External Links:[Link](https://arxiv.org/abs/2609.04298),[Document](https://dx.doi.org/10.48550/ARXIV.2609.04298)Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p4.1),[§5\.4](https://arxiv.org/html/2610.00651#S5.SS4.SSS0.Px1.p1.1)\.
- Shiet al\.\(2024\)Q\. Shi, M\. Tang, K\. Narasimhan, and S\. YaoCan language models solve olympiad programming?\.arXiv preprint arXiv:2404\.10952\.Cited by:[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Siegelet al\.\(2024\)Z\. S\. Siegel, S\. Kapoor, N\. Nagdir, B\. Stroebl, and A\. NarayananCore\-bench: fostering the credibility of published research through a computational reproducibility agent benchmark\.arXiv preprint arXiv:2409\.11363\.Cited by:[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Singhet al\.\(2026\)S\. Singh, Y\. Nan, A\. Wang, D\. Dsouza, S\. Kapoor, A\. Üstün, S\. Koyejo, Y\. Deng, S\. Longpre, N\. Smith, B\. Ermis, M\. Fadaee, and S\. HookerThe Leaderboard Illusion\.Advances in Neural Information Processing Systems38\(en\)\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/70a93f260a51123b3c0e33ecd1b4de97-Abstract-Datasets_and_Benchmarks_Track.html),[Document](https://dx.doi.org/10.52202/085713-2620)Cited by:[Appendix E](https://arxiv.org/html/2610.00651#A5.p1.1),[§4](https://arxiv.org/html/2610.00651#S4.p2.1)\.
- Székely and Rizzo \(2017\)G\. J\. Székely and M\. L\. RizzoThe Energy of Data\.Annual Review of Statistics and Its Application4\(1\),pp\. 447–479\(en\)\.External Links:ISSN 2326\-8298, 2326\-831X,[Link](https://www.annualreviews.org/doi/10.1146/annurev-statistics-060116-054026),[Document](https://dx.doi.org/10.1146/annurev-statistics-060116-054026)Cited by:[§C\.5](https://arxiv.org/html/2610.00651#A3.SS5.p1.1)\.
- Taleuzzaman \(2018\)M\. TaleuzzamanLimit of Blank \(LOB\), Limit of Detection \(LOD\), and Limit of Quantification \(LOQ\)\.Organic & Medicinal Chemistry International Journal7\(5\) \(en\)\.External Links:ISSN 24747610,[Link](https://juniperpublishers.com/omcij/OMCIJ.MS.ID.555722.php),[Document](https://dx.doi.org/10.19080/OMCIJ.2018.07.555722)Cited by:[Figure 3](https://arxiv.org/html/2610.00651#S5.F3)\.
- Tianet al\.\(2024\)M\. Tian, L\. Gao, S\. D\. Zhang, X\. Chen, C\. Fan, X\. Guo, R\. Haas, P\. Ji, K\. Krongchon, Y\. Li,et al\.Scicode: a research coding benchmark curated by scientists\.Advances in Neural Information Processing Systems37,pp\. 30624–30650\.Cited by:[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Vehtariet al\.\(2017\)A\. Vehtari, A\. Gelman, and J\. GabryPractical Bayesian model evaluation using leave\-one\-out cross\-validation and WAIC\.Statistics and Computing27\(5\),pp\. 1413–1432\(en\)\.External Links:ISSN 1573\-1375,[Link](https://doi.org/10.1007/s11222-016-9696-4),[Document](https://dx.doi.org/10.1007/s11222-016-9696-4)Cited by:[§D\.3](https://arxiv.org/html/2610.00651#A4.SS3.p1.1)\.
- Vehtariet al\.\(2024\)A\. Vehtari, D\. Simpson, A\. Gelman, Y\. Yao, and J\. GabryPareto Smoothed Importance Sampling\.Journal of Machine Learning Research25\(72\),pp\. 1–58\.External Links:ISSN 1533\-7928,[Link](http://jmlr.org/papers/v25/19-556.html)Cited by:[§D\.3](https://arxiv.org/html/2610.00651#A4.SS3.p1.1)\.
- Wanget al\.\(2026\)R\. Wang, Y\. Chen, Y\. Wang, C\. Wu, J\. Fang, X\. Cai, Q\. Gu, H\. Su, A\. Zhang, X\. Wang,et al\.AgentNoiseBench: benchmarking robustness of tool\-using llm agents under noisy condition\.arXiv preprint arXiv:2602\.11348\.Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1)\.
- Wang and Wilson \(2005\)W\. Wang and M\. WilsonExploring Local Item Dependence Using a Random\-Effects Facet Model\.Applied Psychological Measurement29\(4\),pp\. 296–318\(EN\)\.External Links:ISSN 0146\-6216,[Link](https://doi.org/10.1177/0146621605276281),[Document](https://dx.doi.org/10.1177/0146621605276281)Cited by:[§3\.2](https://arxiv.org/html/2610.00651#S3.SS2.p1.1)\.
- Xueet al\.\(2025\)T\. Xue, W\. Qi, T\. Shi, C\. H\. Song, B\. Gou, D\. Song, H\. Sun, and Y\. SuAn Illusion of Progress? Assessing the Current State of Web Agents\.arXiv\.Note:arXiv:2504\.01382 \[cs\.AI\]External Links:[Link](http://arxiv.org/abs/2504.01382),[Document](https://dx.doi.org/10.48550/arXiv.2504.01382)Cited by:[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Yaoet al\.\(2024\)S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhanτ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.arXiv preprint arXiv:2406\.12045\.Cited by:[§1](https://arxiv.org/html/2610.00651#S1.p1.1),[§1](https://arxiv.org/html/2610.00651#S1.p3.1),[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Yoranet al\.\(2024\)O\. Yoran, S\. J\. Amouyal, C\. Malaviya, B\. Bogin, O\. Press, and J\. BerantAssistantbench: can web agents solve realistic and time\-consuming tasks?\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 8938–8968\.Cited by:[§4](https://arxiv.org/html/2610.00651#S4.p1.1)\.
- Zhouet al\.\(2023\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.Webarena: a realistic web environment for building autonomous agents\.arXiv preprint arXiv:2307\.13854\.Cited by:[§2](https://arxiv.org/html/2610.00651#S2.SS0.SSS0.Px1.p1.1)\.
## Appendix
## Limitations and Future Work
##### Reliability is necessary but not sufficient for valid evaluation\.
The bounds we estimate describe how consistently a benchmark differentiates the objects it ranks, not whether the rankings correspond to underlying capability or to any external ground truth: a benchmark can be reliable but invalid\. The object\-of\-measurement problem we document is in this sense a construct validity problem in disguise, since reliability against an undefined construct cannot be assessed in principle\([Cronbach and Meehl, 1955](https://arxiv.org/html/2610.00651#bib.bib1);[Salaudeen et al\., 2025](https://arxiv.org/html/2610.00651#bib.bib10)\)\. Future work should complement the reliability framework with validity studies, including external\-criterion validation and cross\-benchmark transfer analyses\.
##### Scope conditions on the headline ceiling\.
The reliability bounds we report are joint properties of the benchmark, the scaffold sample, the task sample, and the competitor set being ranked\. A clustered frontier\-model pool depressesσM2\\sigma^\{2\}\_\{M\}regardless of benchmark quality, lowering reliability not because the benchmark is poorly designed but because it is being used to distinguish systems too close for its resolution; hence, the ceiling we find should be read as a statement about current leaderboards for the current frontier\-model competitor set\. Scaffolds are selected artifacts, often co\-developed with the benchmarks they evaluate, and per\-benchmark counts \(na∈\{2,3\}n\_\{a\}\\in\\\{2,3\\\}\) reflect standard practice across the field rather than a feature of HAL; pooling across benchmarks in Equation[6](https://arxiv.org/html/2610.00651#S3.E6)partially addresses but per\-benchmark scaffold\-variance claims remain weaker, and pooling itself depends onHAL\-Generalistas the cross\-benchmark harness anchor \(Appendix[E](https://arxiv.org/html/2610.00651#A5)\)\. Future work should construct evaluation datasets that systematically span larger scaffold libraries, and should track how the ceiling shifts as the model pool evolves\.
##### Scope of Agent Scaffolds and Systems\.
These results are conditional on the sampled model pool and scaffold set\. A compressed frontier depressesσM2\\sigma\_\{M\}^\{2\}and thus reliability independent of benchmark quality; a richer scaffold sample would sharpen the model–scaffold decomposition\. Our claims are therefore about*these leaderboards for this competitor set*, and generalize to new models and scaffolds only insofar as the sampled conditions represent the intended universe\. We view stating this scope explicitly as part of reliable practice rather than a caveat to it\. Inference remains conditional on the represented evaluation population; partial pooling does not by itself correct selective evaluation or reporting\.
##### Latent\-scale reliability\.
We estimate reliability on the latent logit scale, while leaderboards report observed proportions\. The two scales do not translate one\-to\-one, so our numbers should be read as approximate bounds on observed\-scale rank stability\. Future work could entail a sensitivity analysis comparing the scales\.
##### Raising reliability bounds\.
The empirical scope of this work is the characterization of reliability bounds rather than their displacement\. Validating interventions that raise the ceiling is the subject of subsequent work, since each candidate intervention introduces its own measurement\-design tradeoffs\. Candidate interventions include task selection using item discrimination methods, scaffold sampling under a defined scaffold family, and task\-quality auditing; empirical evaluation of these, alongside methods that reduce the measurement cost of the recommendations, are the most direct extensions\.
## Appendix ADataset Descriptions
### A\.1HAL Dataset
The Holistic Agent Leaderboard \(HAL\) dataset777[https://hal\.cs\.princeton\.edu/](https://hal.cs.princeton.edu/)is a standardized, cost\-aware, and third\-party evaluation platform and dataset initiative developed by the SAgE \(Science of Agent Evaluation\) research group at Princeton University\. The formatted and cleaned tasks\-level of this dataset is publicly released888[https://huggingface\.co/datasets/razam2/hal\-response\-matrix](https://huggingface.co/datasets/razam2/hal-response-matrix)\.
#### A\.1\.1Benchmarks
##### AssistantBench
Tasks consist of time\-consuming, busy\-work tasks that an average person may face, seeking information from the web\.
Which gyms near Tompkins Square Park \(<<200m\) have fitness classes before 7am?
##### CORE\-Bench Hard
Tasks ask an agent to computationally reproduce specific quantitative results from a published scientific paper, given a code repository, dataset, and research paper\.
Run the main\.py file three times\. First, withconfig/uci\.json, the preprocessing task, and the CTGCN\-C method\. Second, withconfig/uci\.json, the embedding task, and the CTGCN\-C method\. Third, using Python3 withconfig/uci\.jsonand the link\-pred task\.
##### GAIA
GAIA consists of tasks requiring multi\-step tool use — including web browsing, code execution, file reading \(PDFs, spreadsheets, audio\), and multimodal understanding\.
\(Given an Excel file\) The attached Excel file contains the sales of menu items for a local fast\-food chain\. What were the total sales that the chain made from food \(not including drinks\)? Express your answer in USD with two decimal places\.
##### OnlineMind2Web
Online Mind2Web is the live, online version of Mind2Web\. It does not rely on cached pages and allows for real\-time testing against dynamic, evolving web interfaces\. All tasks are sourced from 136 popular websites to reflect authentic user workflows\.
Browse used Audi cars made before 2015 and sort by lowest price on KBB\. Website:https://www\.kbb\.com/
##### SciCode
Tasks consist of research\-level coding problems decomposed into subproblems drawn from actual scientific work across physics, chemistry, biology, math, and materials science\.
Main problem: Reproduce the Chern number phase diagram of the Haldane model on a hexagonal lattice\. Subproblem 1\.1: Write a Haldane model Hamiltonian on a hexagonal lattice, given: wavevector componentskxk\_\{x\}andkyk\_\{y\}, lattice spacingaa, nearest\-neighbor coupling constantt1t\_\{1\}, next\-nearest\-neighbor coupling constantt2t\_\{2\}, phaseφ\\varphifor next\-nearest\-neighbor hopping, and on\-site energymm\. Output: a2×22\\times 2matrix\. Subproblem 1\.2: Calculate the Chern number using the Haldane Hamiltonian, given the grid sizeδ\\deltafor discretizing the Brillouin zone…
##### ScienceAgentBench
ScienceAgentBench is a benchmark for evaluating the ability of language agents to conduct data\-driven scientific discovery\.
\(Given separate training and test datasets\) Train a multitask model on the Clintox dataset to predict a drug’s toxicity and FDA approval status\. Save the test set predictions, including the SMILES representation of drugs and the probability of positive labels, to “pred\_results/clintox\_test\_pred\.csv”\.
##### SWE\-bench Verified Mini
A subset of SWE\-bench Verified where each task gives the agent a real GitHub issue description and a full Python repository, and requires the agent to produce a code patch that makes failing unit tests pass\.
\(Given a repo and issue description\) Navigate the repository, identify the relevant code path indjango/db/models/query\.py, and produce a\.patchfile that makes the FAIL\_TO\_PASS unit tests pass without breaking existing PASS\_TO\_PASS tests\.
##### τ\\tau\-bench Airline
Tasks simulate a realistic airline customer service scenario in which a human user \(played by another LLM\) contacts an agent with requests like rebooking a flight, adding a passenger, or canceling a reservation\. The agent must follow a detailed airline policy document and a limited set of callable functions\.
\(Given the airline’s policy and pre\-defined functions\) Your user id isdaiki\_muller\_1116\. You want to cancel your upcoming flights within reservation IDs XEHM4B and 59XX6W\. If the agent says either reservation has basic economy flights, ask to upgrade them to economy first and then cancel them\. You are very persistent and terse but clear\. After the third agent message, also ask whether you have any other upcoming flights and what their total cost is\.
##### USACO
Tasks are problems from the USA Computing Olympiad, spanning four difficulty tiers \(Bronze through Platinum\)\.
Farmer John hasNNcows \(2≤N≤1052\\leq N\\leq 10^\{5\}\), each liking exactly one type of hayhih\_\{i\}\. He can host focus groups over contiguous ranges of cows — if more than half the cows in a group prefer the same type, all cows switch to that type\. He wants to know which types of hay can become universally liked\. GivenTTtest cases, each withNNand a list ofhih\_\{i\}values, output all achievable universal hay types in increasing order, or−1\-1\.
#### A\.1\.2Scaffolds
The scaffolds included in the HAL benchmark dataset are in Table[7](https://arxiv.org/html/2610.00651#A2.T7)\. For each benchmark, the scaffold consist of a contrast between a generalist scaffold \(e\.g\., HAL generalist, Claude Code\) and and a specialist scaffold tailored to the specific benchmark \(e\.g\., SWE agent,τ\\tau\-bench tool\-calling\)\. For most benchmarks, the majority of LLMs have scores across multiple scaffolds, with the exceptions of USACO, SWE\-bench Verified Mini, and ScienceAgentBench\. USACO in particular has very little model–scaffold diversity, with uncertainty visible in its very large HDI intervals when trying to generalize scaffolds\.
### A\.2Harbor Index Dataset for Decision Study Corroboration
Harbor\-Index999[https://harbor\-index\.org/](https://harbor-index.org/)is a curated meta\-dataset and benchmark containing 82 difficult and diverse tasks designed for evaluating AI language model agents\. Distilled from a pool of over 6,000 candidate tasks across 54 benchmarks, the final dataset features 82 high\-quality tasks spanning 29 benchmarks and seven domains \(including software engineering, scientific research, tool use, mathematics, data analytics, and security\)\. It was built by passing candidate tasks through difficulty filtering, automated AI audits, human reviews, and iterative audit\-and\-fix loops to weed out structurally broken or flawed tasks\.
The final dataset lacks sufficient per\-benchmark task representation that prevent inclusion in all the analyses of this paper\. Descriptive statistics are in Table[4](https://arxiv.org/html/2610.00651#A1.T4)\. Thus, to estimate both task and benchmark effects, we take the subset of benchmarks with at least three items\. The cross\-benchmark model of Equation[6](https://arxiv.org/html/2610.00651#S3.E6)is fit and produces the sample the posterior draws used in Figure[3\(b\)](https://arxiv.org/html/2610.00651#S5.F3.sf2)\. For our analyses, we use their verified and judged outcome classifications as the measure of task success \(i\.e\., if the verifier marked a reward as a false negative and the judge determined it was a “true solve” it was marked as successful\)\. Because the dataset suffers from a small quantity of total tasks, both the proportion of variance due to persistent differences in LLM and overall reliability begins to drop more sharply if we take the subset of benchmarks with at least four items, decreasing further with the subset that has five items\. We conjecture this pattern may be in part due to the nonrandom nature of the item selection process used to represent each benchmark\.
Table 4:Descriptive statistics from the complementary Harbor Index dataset\.Boldeddatasets have at least three tasks and were used in the analysis of benchmark and item effects on agent system signal\.BenchmarkModelsTasksScaffoldModel\-Scaffold PairMean Scorealgotune954180\.078arcagi2954180\.067bigcodebench914180\.222bixbench954180\.122codepde914180\.000cybergym924180\.1 39dacode914180\.111featurebench944180\.042gaia934180\.185gaia2954180\.051gpqadiamond914180\.056gso974180\.024hle984180\.079labbench944180\.182omnimath924180\.278qcircuitbench914180\.000replicationbench914180\.444scicode934180\.019skillsbench924180\.167sldbench914180\.056spider2924180\.000swebenchpro944180\.139swebenchverified954180\.067swelancer924180\.139swesmith914180\.111swtbenchverified914180\.333tb934180\.241usaco914180\.000widesearch914180\.000
### A\.3External Validation Datasets for Model Rankings
We compare the posterior\-median model effectθ^m\\widehat\{\\theta\}\_\{m\}with four contemporaneous external benchmarks containing at least six overlapping LLMs and which report individual task\-level performance by model\. The main body reports concordance of the observed rankings withθ^m\\widehat\{\\theta\}\_\{m\}and mean score with these external benchmarks in Table[3](https://arxiv.org/html/2610.00651#S5.T3); §[A\.3\.1](https://arxiv.org/html/2610.00651#A1.SS3.SSS1)analyzes the differences between the two HAL ranking mechanisms\. To prevent data leakage for the external SWE\-bench andτ2\\tau^\{2\}\-bench, the semantically corresponding HAL benchmark is excluded before estimating bothθ^m\\widehat\{\\theta\}\_\{m\}and the raw\-score baseline \(see §[D\.4](https://arxiv.org/html/2610.00651#A4.SS4)\), as described in the details of the data and collection below\.
- •The Berkeley Function Calling Leaderboard\(BFCL\) V4\([Patil et al\., 2025](https://arxiv.org/html/2610.00651#bib.bib46)\)is a benchmark designed to evaluate the tool use, API execution, and agentic capabilities of large language models\.101010[https://gorilla\.cs\.berkeley\.edu/leaderboard\.html](https://gorilla.cs.berkeley.edu/leaderboard.html)\. We use BFCL’s holistic score for function\-calling models\.
- •
- •SWE\-bench Verified\([Jimenez et al\., 2023](https://arxiv.org/html/2610.00651#bib.bib17)\)is the superset of tasks that are associated with the HAL dataset SWE\-bench Verified Mini\. The data of several third\-party external leaderboard providers121212[https://www\.swebench\.com/](https://www.swebench.com/),[https://www\.vals\.ai/benchmarks/swebench](https://www.vals.ai/benchmarks/swebench),[https://llm\-stats\.com/benchmarks/swe\-bench\-verified](https://llm-stats.com/benchmarks/swe-bench-verified)were collected\. Accuracy scores were averaged across leaderboards for overlapping models for the standard Kendall’sτ\\tau, and where left separate for the multilevel partial correlation \(Appendix[A\.3\.2](https://arxiv.org/html/2610.00651#A1.SS3.SSS2)\)\. When performing analyses with this data, such as in Table[3](https://arxiv.org/html/2610.00651#S5.T3), the external correlations were calculated with estimates from the HAL dataset after removing SWE\-bench Verified Mini from the analysis\. The description of this process based on the pooled model \(Eq\.[6](https://arxiv.org/html/2610.00651#S3.E6)\) can be found in Appendix[D\.4](https://arxiv.org/html/2610.00651#A4.SS4)\.
- •τ𝟐\\mathbf\{\\tau^\{2\}\}\-bench Core\([Barres et al\., 2025](https://arxiv.org/html/2610.00651#bib.bib48)\)is an evaluation framework and benchmark for conversational AI and customer service agents that tests performance in a “dual\-control” environment where both the AI agent and the user take active actions\. It contains tasks that evaluate LLM agents on retail, airline \(of which HAL’sτ\\tau\-bench Airline is a subset\), and telecom customer\-service tasks, where the agent and the user both act on the world\.131313[https://taubench\.com/leaderboard?benchmark=core](https://taubench.com/leaderboard?benchmark=core)We use the overall score across all of these domains and, as with SWE\-bench Verified, we only calculate correlations with HAL data after removingτ\\tau\-bench Airline\. The description of this process based on the pooled model \(Eq\.[6](https://arxiv.org/html/2610.00651#S3.E6)\) can be found in Appendix[D\.4](https://arxiv.org/html/2610.00651#A4.SS4)\.
#### A\.3\.1Differences in concordance with external performance
In addition to Table[3](https://arxiv.org/html/2610.00651#S5.T3), we provide additional analyses to explore the differences between transportable ranking methods\. This section supplements the differences in correlations\. We compared the rank agreement of the more traditional aggregated mean scorex¯\\bar\{x\}and the latent estimated valueθ^\\hat\{\\theta\}with external performance separately for each benchmark\. For each model, the external reference was its mean reported benchmark score across the available sources and evaluation conditions\. Both scoring mechanisms were evaluated on the same models within each benchmark\. We used Kendall’sτb\\tau\_\{b\}, defining the paired difference as
Δb=τb\(x¯,Zb\)−τb\(θ^,Zb\),\\Delta\_\{b\}=\\tau\_\{b\}\(\\bar\{x\},Z\_\{b\}\)\-\\tau\_\{b\}\(\\hat\{\\theta\},Z\_\{b\}\),whereZbZ\_\{b\}denotes the aggregated external reference for benchmarkbb\. Thus,Δb<0\\Delta\_\{b\}<0favorsθ^\\hat\{\\theta\}\.
To characterize uncertainty, we constructed nominal 95% percentile intervals using a paired bootstrap within each benchmark \(with 20,000 bootstraps\)\. Each resampled unit was a complete model\-level observation containing mean score,θ^\\hat\{\\theta\}, and the external reference, preserving the dependence between the two estimated correlations\. Given the small numbers of models, particularly for benchmarks with six or seven observations, these intervals were interpreted as exploratory rather than as reliably calibrated confidence intervals\. We also recomputed each difference after removing each model in turn to assess sensitivity to individual observations\. The resulting leave\-one\-out ranges are influence diagnostics, not confidence intervals\.
Table 5:Within\-benchmark agreement with the aggregated external reference\. Negative differences favorθ^\\hat\{\\theta\}\. Bootstrap intervals are exploratory, marginal 95% percentile intervals and are not adjusted for multiple comparisons\.Benchmarknnτb\(x¯,Z\)\\tau\_\{b\}\(\\bar\{x\},Z\)τb\(θ^,Z\)\\tau\_\{b\}\(\\hat\{\\theta\},Z\)Δb\\Delta\_\{b\}Bootstrap intervalLOO rangeBFCL70\.4290\.714−0\.286\-0\.286\[−1\.111,0\.556\]\[\-1\.111,\\phantom\{\-\}0\.556\]\[−0\.533,0\.000\]\[\-0\.533,\\phantom\{\-\}0\.000\]SWE\-bench200\.5330\.755−0\.222\-0\.222\[−0\.503,0\.045\]\[\-0\.503,\\phantom\{\-\}0\.045\]\[−0\.258,−0\.129\]\[\-0\.258,\-0\.129\]τ2\\tau^\{2\}\-core60\.3330\.867−0\.533\-0\.533\[−1\.231,0\.000\]\[\-1\.231,\\phantom\{\-\}0\.000\]\[−0\.800,−0\.400\]\[\-0\.800,\-0\.400\]Terminal\-Bench 2\.070\.5240\.810−0\.286\-0\.286\[−0\.824,0\.000\]\[\-0\.824,\\phantom\{\-\}0\.000\]\[−0\.400,−0\.133\]\[\-0\.400,\-0\.133\]The observed correlations consistently favoredθ^\\hat\{\\theta\}\(Table[5](https://arxiv.org/html/2610.00651#A1.T5)\)\. Its advantage inτb\\tau\_\{b\}ranged from 0\.222 on SWE\-bench to 0\.533 onτ2\\tau^\{2\}\-core\. Both mechanisms were positively associated with the external reference in every benchmark, butθ^\\hat\{\\theta\}exhibited stronger agreement throughout\.
The direction of the comparison was also stable under single\-model deletion\. Every leave\-one\-out difference remained strictly negative for SWE\-bench,τ2\\tau^\{2\}\-core, and Terminal\-Bench 2\.0\. For BFCL, deleting an individual model could reduce the difference to zero, but no deletion reversed its sign\. Consequently, the observed direction was not dependent on retaining any single model, although this diagnostic does not establish stability to broader changes in the model sample\.
As expected, uncertainty remained substantial due to the constraint of having few candidate contemporaneous benchmarks that correspond with HAL benchmark models\. Thus predictably for small sample sizes, none of the bootstrap intervals excluded zero: those for BFCL and SWE\-bench extended above zero, whereas those forτ2\\tau^\{2\}\-core and Terminal\-Bench 2\.0 ended at zero\. The latter endpoints should not be interpreted as precise significance thresholds, given the discrete rank statistic and small samples\. Interval bounds below−1\-1are permissible because a difference between two Kendall correlations lies in\[−2,2\]\[\-2,2\]\. Undefined bootstrap differences were absent except forτ2\\tau^\{2\}\-core, where they occurred in 0\.01% of resamples; its interval was calculated from the defined replicates\.
These difference analyses provide consistent descriptive evidence thatθ^\\hat\{\\theta\}ranks the observed models more closely to their aggregated external performance than the mean score does\. They do not, however, establish a statistically conclusive advantage within individual benchmarks\. Moreover, the analysis treats each model’s aggregated external score as its reference value: it does not separately propagate uncertainty in that score or adjust for unequal coverage of sources and evaluation conditions\. The results therefore characterize agreement with the observed mean references, rather than with context\-adjusted latent performance\.
#### A\.3\.2Combined External Data Multilevel Partial Correlation
In addition to traditional Kendall’sτb\\tau\_\{b\}correlations, a multilevel partial correlation was calculated to clarify the relationships with greater statistical power\. For LLMmmevaluated on benchmarkbband external leaderboard providersℓ\\ell, letRmbℓR\_\{mb\\ell\}denote its observed performance rank andPmbℓP\_\{mb\\ell\}the rank induced by the calculated proxy\. Benchmark\- and leaderboard\-provider\-level heterogeneity can be modeled through the rank\-based mixed\-effects specifications
Rmbℓ=𝐱mbℓ⊤𝜷R\+ubR\+vℓR\+εmbℓR,Pmbℓ=𝐱mbℓ⊤𝜷P\+ubP\+vℓP\+εmbℓP,R\_\{mb\\ell\}=\\mathbf\{x\}\_\{mb\\ell\}^\{\\top\}\\bm\{\\beta\}\_\{R\}\+u^\{R\}\_\{b\}\+v^\{R\}\_\{\\ell\}\+\\varepsilon^\{R\}\_\{mb\\ell\},\\qquad P\_\{mb\\ell\}=\\mathbf\{x\}\_\{mb\\ell\}^\{\\top\}\\bm\{\\beta\}\_\{P\}\+u^\{P\}\_\{b\}\+v^\{P\}\_\{\\ell\}\+\\varepsilon^\{P\}\_\{mb\\ell\},whereub\(⋅\)u\_\{b\}^\{\(\\cdot\)\}andvℓ\(⋅\)v\_\{\\ell\}^\{\(\\cdot\)\}are random effects for benchmarks and leaderboards, respectively, and𝐱mbℓ\\mathbf\{x\}\_\{mb\\ell\}contains any control variables\. The inclusion of benchmark and leaderboard random effects removes the need to average across the SWE\-bench Verified leaderboards, increasing the overall number of observations to 52 for this analysis\. The multilevel partial Kendall correlation is then defined as the concordance between the adjusted residual ranks,
τpartial=𝔼\[sgn\(R~a−R~a′\)sgn\(P~a−P~a′\)\],\\tau\_\{\\mathrm\{partial\}\}=\\mathbb\{E\}\\\!\\left\[\\operatorname\{sgn\}\\\!\\left\(\\tilde\{R\}\_\{a\}\-\\tilde\{R\}\_\{a^\{\\prime\}\}\\right\)\\operatorname\{sgn\}\\\!\\left\(\\tilde\{P\}\_\{a\}\-\\tilde\{P\}\_\{a^\{\\prime\}\}\\right\)\\right\],whereR~mbℓ=Rmbℓ−𝐱mbℓ⊤𝜷^R−u^bR−v^ℓR\\tilde\{R\}\_\{mb\\ell\}=R\_\{mb\\ell\}\-\\mathbf\{x\}\_\{mb\\ell\}^\{\\top\}\\hat\{\\bm\{\\beta\}\}\_\{R\}\-\\hat\{u\}\_\{b\}^\{R\}\-\\hat\{v\}\_\{\\ell\}^\{R\}andP~mbℓ=Pmbℓ−𝐱mbℓ⊤𝜷^P−u^bP−v^ℓP\\tilde\{P\}\_\{mb\\ell\}=P\_\{mb\\ell\}\-\\mathbf\{x\}\_\{mb\\ell\}^\{\\top\}\\hat\{\\bm\{\\beta\}\}\_\{P\}\-\\hat\{u\}\_\{b\}^\{P\}\-\\hat\{v\}\_\{\\ell\}^\{P\}\. Thus,τpartial\\tau\_\{\\mathrm\{partial\}\}measures agreement between the LLM rankings and the proxy rankings after accounting for observed covariates and clustering attributable to benchmarks and leaderboards\. The results are in Table[6](https://arxiv.org/html/2610.00651#A1.T6)\.
Table 6:Hierarchical concordance with external agent benchmarks\.Kendall’sτ\\taufor the unweighted HAL mean and reliability\-adjusted latent \(θ^\\widehat\{\\theta\}\) model effect \(median posterior\) with\[95%\]\[95\\%\]CIs\.External BenchmarkMean scoreθ^\\widehat\{\\theta\}DifferenceLLM ObservationsMultilevelτpartial\\tau\_\{\\text\{partial\}\}0\.430\[0\.266,0\.570\]0\.430\\,\[0\.266,\\,0\.570\]0\.632\[0\.507,0\.732\]\\mathbf\{0\.632\}\\,\[0\.507,\\,0\.732\]\+0\.202\+0\.20252
## Appendix BHAL Coverage Matrix
Below is the list of agents used to test each of the benchmarks\. Each model was not uniformly tested with each benchmark and agent\. As shown in Table[7](https://arxiv.org/html/2610.00651#A2.T7), about nearly all \(11/13\) agents have only been tested on a single benchmark\. Similarly, many models \(12/54\) are tested with a single agent\. The full breakdown of which models, agents and benchmarks have been tested can be seen in Figure[5](https://arxiv.org/html/2610.00651#A2.F5)\.
Table 7:Agent harness/scaffold by benchmark\.BenchmarkAgent NamesAssistantBenchhal\_generalist, browser\-useCORE\-Bench Hardhal\_generalist, coreagent, claude\_codeGAIAhal\_generalist, hf\_open\_deep\_researchOnlineMind2Webbrowser\-use, seeactSciCodehal\_generalist, tool\_calling\_agent, scicode\_zeroScienceAgentBenchhal\_generalist, sab\_selfdebugSWE\-bench Verified Minihal\_generalist, sweagentτ\\tau\-bench Airlinehal\_generalist, taubench\_tool\_calling, taubench\_fewshotusacohal\_generalist, usaco\_episodic\_semanticFigure 4:Number of benchmarks tested \(left\) and agents tested \(right\) per model\.Figure 5:Matrix indicating all the variations of benchmarks, scaffold and models \(with reasoning\-effort\) that were included in our dataset\.
## Appendix CMethods, Estimation, and Computational Details
This appendix reports prior distributions, parameterization, sampler settings, convergence diagnostics, posterior predictive checks, and the treatment of repeated observations\. It also gives reliability expressions for the observed unbalanced allocation\. For each posterior draw, design\-specific error is computed using the realized cell weights rather than the equal\-allocation approximations used for the main D\-studies\.
We provide the full random\-effects structure for each model using standard mixed\-model notation\. For the main body studies, we estimate all variance components jointly using Bayesian generalized linear mixed models, which provide partial pooling for sparsely observed cells and propagate uncertainty into the resulting generalizability coefficients\.
### C\.1Estimation Details
#### C\.1\.1Leaderboard\-level pooled model
Letybima∈\{0,1\}y\_\{bima\}\\in\\\{0,1\\\}denote success on benchmarkbb, taskii, modelmm, and scaffoldaa\. We assume
ybima∣pbima∼Bernoulli\(pbima\),logit\(pbima\)=ηbima,y\_\{bima\}\\mid p\_\{bima\}\\sim\\operatorname\{Bernoulli\}\(p\_\{bima\}\),\\qquad\\operatorname\{logit\}\(p\_\{bima\}\)=\\eta\_\{bima\},with latent linear predictor \(from Equation[6](https://arxiv.org/html/2610.00651#S3.E6)\)
ηbima=\\displaystyle\\eta\_\{bima\}=\{\}μ\+ub\+ubi\+um\+ua\+ubim\+ubia\+uma\\displaystyle\\mu\+u\_\{b\}\+u\_\{bi\}\+u\_\{m\}\+u\_\{a\}\+u\_\{bim\}\+u\_\{bia\}\+u\_\{ma\}\+ubm\+uba\+ubma\.\\displaystyle\+u\_\{bm\}\+u\_\{ba\}\+u\_\{bma\}\.Herebibiidentifies a task nested within its benchmark\. Each term is a mean\-zero random intercept,
uF,j∣σF∼𝒩\(0,σF2\),F∈\{b,bi,m,a,bim,bia,ma,bm,ba,bma\},u\_\{F,j\}\\mid\\sigma\_\{F\}\\sim\\mathcal\{N\}\(0,\\sigma\_\{F\}^\{2\}\),\\qquad F\\in\\\{b,bi,m,a,bim,bia,ma,bm,ba,bma\\\},and random\-effect families are conditionally independent\. This decomposition separates persistent model and scaffold differences from variation attributable to benchmark choice, task composition, and model–scaffold compatibility\. In particular,ubmu\_\{bm\}captures benchmark\-dependent model performance, whileubmau\_\{bma\}captures benchmark\-specific compatibility between models and scaffolds; both can alter rankings even when average model effects are unchanged\.
The model was fitted to29,92329\{,\}923observations spanning99benchmarks,1,1171\{,\}117benchmark–task units,5454models, and1313scaffolds\. Because the likelihood is Bernoulli\-logit, all variance components are defined on the latent log\-odds scale\. When an observation\-level residual is required for a generalizability coefficient, the conventional logistic varianceπ2/3\\pi^\{2\}/3is used\. Reliability quantities are computed separately for every posterior draw, rather than from ratios of posterior mean variances, thereby preserving uncertainty in nonlinear variance decompositions\.
brmsspecification:
brm\(score~1\+\(1\|benchmark\)\+\(1\|benchmark:task\_id\)
\+\(1\|model\_name\)\+\(1\|agent\_name\)
\+\(1\|benchmark:task\_id:model\_name\)
\+\(1\|benchmark:task\_id:agent\_name\)
\+\(1\|model\_name:agent\_name\)
\+\(1\|benchmark:model\_name\)
\+\(1\|benchmark:agent\_name\)
\+\(1\|benchmark:model\_name:agent\_name\),
family=bernoulli\("logit"\),\.\.\.\)
#### C\.1\.2Benchmark\-level decomposition models
To quantify the signal supplied by each benchmark, we also fit a separate model within every benchmark using Equation[2](https://arxiv.org/html/2610.00651#S3.E2):
yima\\displaystyle y\_\{ima\}∼Bernoulli\(pima\),\\displaystyle\\sim\\operatorname\{Bernoulli\}\(p\_\{ima\}\),logit\(pima\)\\displaystyle\\operatorname\{logit\}\(p\_\{ima\}\)=μ\+ui\+um\+ua\+uim\+uia\+uma\.\\displaystyle=\\mu\+u\_\{i\}\+u\_\{m\}\+u\_\{a\}\+u\_\{im\}\+u\_\{ia\}\+u\_\{ma\}\.These fits distinguish stable model variation from task\-, scaffold\-, and interaction\-driven variation within an instrument\. Thus, a benchmark with many observations need not be highly informative for ranking models: its effective signal depends on the posterior magnitude of model\-related variance relative to the variance induced by tasks, scaffolds, and their interactions\.
brmsspecification:
brm\(score~1\+\(1\|task\_id\)\+\(1\|model\_name\)\+\(1\|agent\_name\)
\+\(1\|task\_id:model\_name\)\+\(1\|task\_id:agent\_name\)
\+\(1\|model\_name:agent\_name\),
family=bernoulli\("logit"\),\.\.\.\)
#### C\.1\.3Priors and posterior computation
We used the default weakly informativebrmspriors\. For every group\-level standard deviation,
σF∼Student\-t\+\(3,0,2\.5\),\\sigma\_\{F\}\\sim\\operatorname\{Student\}\\text\{\-\}t^\{\+\}\(3,0,2\.5\),wheret\+t^\{\+\}denotes truncation toσF≥0\\sigma\_\{F\}\\geq 0\. Equivalently, the estimated variance component isσF2\\sigma\_\{F\}^\{2\}\. The prior is concentrated near modest latent\-scale heterogeneity while retaining sufficiently heavy tails for large benchmark or task effects\. The intercept used the corresponding weakly informativeStudent\-t\(3,0,2\.5\)\\operatorname\{Student\}\\text\{\-\}t\(3,0,2\.5\)prior\. These priors regularize components supported by few levels—notably benchmark and scaffold effects—without forcing them toward equality\.
Posterior sampling used Stan’s Hamiltonian Monte Carlo implementation throughbrms\([Bürkner, 2021](https://arxiv.org/html/2610.00651#bib.bib12)\), with the conservative target average proposal acceptance probability during the adaptation period for the sampleradapt\_delta=0\.95\\texttt\{adapt\\\_delta\}=0\.95without any resultant divergent transitions\. The full model used six chains of9,0009\{,\}000iterations, including4,0004\{,\}000warm\-up iterations, with thinning by eight, yielding3,7503\{,\}750retained draws\. Each benchmark\-specific model used four chains of2,0002\{,\}000iterations, including1,0001\{,\}000warm\-up iterations, with thinning by three\. On an Apple M1 Max, the full fit required approximately5\.55\.5hours, while a benchmark\-specific fit required approximately6\.56\.5minutes on average\.
For the full model, all reported split\-R^\\widehat\{R\}values rounded to1\.001\.00; bulk effective sample sizes for variance components ranged from1,5591\{,\}559to3,7433\{,\}743, and tail effective sample sizes ranged from2,0622\{,\}062to3,7833\{,\}783\. Benchmark\-specific fits were also inspected individually\. For example, the largest split\-R^\\widehat\{R\}in the USACO fit was1\.031\.03, with uncertainty retained in all downstream posterior summaries rather than suppressed through plug\-in estimates\.
##### Convergence\.
All parameters across all models fit and all estimates derived from posterior draws throughout this study achieveR^<1\.05\\hat\{R\}<1\.05, the standard threshold for adequate mixing\. For the 95% of estimatesR^\\widehat\{R\}values rounded to1\.001\.00\. For the leaderboard\-level model, complete convergence information is in Table[8](https://arxiv.org/html/2610.00651#A3.T8)
Table 8:Full posterior summary\.Bayesian fit estimates for the full leaderboard\-level model \(Eq\.[6](https://arxiv.org/html/2610.00651#S3.E6)\)\. Group\-level entries are standard deviations of random intercepts on the latent log\-odds scale\. Intervals are equal\-tailed 95% posterior credible intervals\.EffectLevelsEstimateSD95% CrIR^\\widehat\{R\}Bulk ESSTail ESSGroup\-level standard deviationsScaffold,σa\\sigma\_\{a\}130\.790\.44\[0\.06,1\.74\]\[0\.06,\\ 1\.74\]1\.0034543360Benchmark,σb\\sigma\_\{b\}92\.450\.73\[1\.39,4\.21\]\[1\.39,\\ 4\.21\]1\.0037433783Benchmark×\\timesscaffold,σba\\sigma\_\{ba\}210\.690\.36\[0\.06,1\.43\]\[0\.06,\\ 1\.43\]1\.0035803615Benchmark×\\timesmodel,σbm\\sigma\_\{bm\}1820\.520\.17\[0\.12,0\.80\]\[0\.12,\\ 0\.80\]1\.0017712110Benchmark×\\timesmodel×\\timesscaffold,σbma\\sigma\_\{bma\}2870\.550\.14\[0\.24,0\.79\]\[0\.24,\\ 0\.79\]1\.0015592062Task within benchmark,σbt\\sigma\_\{bt\}11172\.410\.09\[2\.25,2\.58\]\[2\.25,\\ 2\.58\]1\.0033093350Task×\\timesscaffold in benchmark,σbta\\sigma\_\{bta\}23941\.030\.05\[0\.93,1\.14\]\[0\.93,\\ 1\.14\]1\.0034453540Task×\\timesmodel in benchmark,σbtm\\sigma\_\{btm\}188560\.880\.07\[0\.75,1\.01\]\[0\.75,\\ 1\.01\]1\.0025133241Model,σm\\sigma\_\{m\}540\.780\.14\[0\.52,1\.07\]\[0\.52,\\ 1\.07\]1\.0031523550Model×\\timesscaffold,σma\\sigma\_\{ma\}2090\.360\.14\[0\.05,0\.61\]\[0\.05,\\ 0\.61\]1\.0018522376Population\-level coefficientIntercept,μ\\mu—−1\.95\-1\.950\.88\[−3\.62,−0\.17\]\[\-3\.62,\-0\.17\]1\.0036933574
Note\.“Levels” denotes the number of observed levels of each grouping factor\. Estimate and SD are the posterior mean and posterior standard deviation, respectively\. For group\-level effects, the estimate summarizes the random\-intercept standard deviationσ\\sigma; for the intercept, it summarizes the coefficient itself\. ESS denotes effective sample size\. “Scaffold” corresponds toagent\_namein the estimation data\.
#### C\.1\.4Leave\-one\-benchmark\-out stability
To assess whether the leaderboard\-level decomposition was dominated by any single instrument, we refit Equation[6](https://arxiv.org/html/2610.00651#S3.E6)after removing each benchmark in turn\. These are full Bayesian refits, not importance\-sampling approximations\. Each refit used four chains of4,0004\{,\}000iterations, including2,0002\{,\}000warm\-up iterations, with thinning by five, yielding1,6001\{,\}600retained draws\. Stability was evaluated by comparing the posterior distributions of variance components and derived reliability quantities with those from the complete\-data fit\. This analysis directly tests whether conclusions about model, scaffold, and benchmark contributions persist under changes to the benchmark ecosystem\.
### C\.2Bayesian Estimation for Linear Mixed Effect Models
To establish methodological sensitivity, Equation[6](https://arxiv.org/html/2610.00651#S3.E6)was estimated using a Bayesian linear mixed effect model on the observation scale\. Additional detail about the differences in methods and results can be found in Appendix[G](https://arxiv.org/html/2610.00651#A7)\. All linear estimations of this use the same hyperparameters as above, except that the family is Gaussian with an identity link\.
### C\.3Frequentist Linear Mixed Effects \(LME\)
We fit a Gaussian identity\-link model via restricted maximum likelihood \(REML\)\. Although the binary outcome violates normality, LME provides a familiar baseline and is the most commonly used variance\-component estimator in G\-theory applications\. Variance components are extracted directly from the REML fit\.
### C\.4Frequentist Generalized Linear Mixed Effects \(GLME\)
A logistic mixed\-effects model with a Bernoulli likelihood and logit link is fit via Laplace approximation to the marginal likelihood, with parameter estimation by penalized iteratively reweighted least squares \(PIRLS\)\. This respects the binary nature of the data but estimates are on the logit scale; we convert variance proportions by computing the share of total variance \(including the logistic residual varianceπ2/3\\pi^\{2\}/3\) attributable to each component\. Both the Linear and Generalized Linear models were estimated usinglme4\([Bates et al\., 2015](https://arxiv.org/html/2610.00651#bib.bib14)\)\.
### C\.5Nonparametric Distance Components \(DISCO\)
We estimate a nonparametric variance decomposition \(see Appendix[G\.2](https://arxiv.org/html/2610.00651#A7.SS2)\) using distance components\([Rizzo and Székely, 2010](https://arxiv.org/html/2610.00651#bib.bib9);[Székely and Rizzo, 2017](https://arxiv.org/html/2610.00651#bib.bib11)\)from the energy\-statistics literature\. DISCO decomposes total dispersion—measured by pairwise Euclidean distances—into between\- and within\-group components without distributional assumptions\. For a single facet withKKgroups,
𝒮total=𝒮between\+𝒮within,\\mathcal\{S\}\_\{\\text\{total\}\}\\;=\\;\\mathcal\{S\}\_\{\\text\{between\}\}\+\\mathcal\{S\}\_\{\\text\{within\}\},\(10\)where𝒮\\mathcal\{S\}denotes the energy\-based dispersion statistic\. The G\-coefficient analog isEρdisco2=𝒮p/\(𝒮p\+𝒮within\)E\\rho^\{2\}\_\{\\textsc\{disco\}\}=\\mathcal\{S\}\_\{p\}/\(\\mathcal\{S\}\_\{p\}\+\\mathcal\{S\}\_\{\\text\{within\}\}\)\. DISCO captures nonlinear relationships and is robust to the heavy skewness observed in the Bayesian posteriors, particularly for facets with few levels \(e\.g\., agents\)\. However, it does not guarantee a positiveσϵ2/𝒮global within\\sigma^\{2\}\_\{\\epsilon\}/\\mathcal\{S\}\_\{\\text\{global within\}\}term, so its reported “variance”/dispersion shares sum to one across the total explained dispersion\.
### C\.6Details On Compute\-Usage
Estimations of the main variance decomposition models used in the body of the paper took 8 total hours on an Apple M1 Max\. LOO ablation studies in Appendix[D](https://arxiv.org/html/2610.00651#A4)took 57 hours\. Methodological contrasts reported in Appendix[G](https://arxiv.org/html/2610.00651#A7)took 2 hours for frequentist estimations \(both linear mixed effect models and generalized mixed effect models\), Bayesian linear mixed effect models took 8 hours, and nonparametric distance components estimates took 18 hours\.
### C\.7Posterior Estimands
Table[9](https://arxiv.org/html/2610.00651#A3.T9)contains the set of estimands found in the paper\.
Table 9:Reliability estimands\. Each row appliesEρo2=σo2/\(σo2\+σδ2\)E\\rho\_\{o\}^\{2\}=\\sigma\_\{o\}^\{2\}/\(\\sigma\_\{o\}^\{2\}\+\\sigma\_\{\\delta\}^\{2\}\),SNRo=σo2/σδ2\\text\{SNR\}\_\{o\}=\\sigma\_\{o\}^\{2\}/\\sigma\_\{\\delta\}^\{2\}, but changes the object whose ordering should generalize\. Thenfn\_\{f\}terms denote equal allocation over facetff\.EstimandObject of inference𝝈𝒐𝟐\\bm\{\\sigma\_\{o\}^\{2\}\}𝝈𝜹𝟐\(𝓓\)\\bm\{\\sigma\_\{\\delta\}^\{2\}\(\\mathcal\{D\}\)\}Eq\.EρM\(b\)2E\\rho^\{2\}\_\{M\(b\)\},SNRM\(b\)\\text\{SNR\}\_\{M\(b\)\}Does benchmarkbbpreserve model order?σM2\\sigma\_\{M\}^\{2\}σIM2ni\+σMA2na\+σIMA,e2nina\\dfrac\{\\sigma\_\{IM\}^\{2\}\}\{n\_\{i\}\}\+\\dfrac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\+\\dfrac\{\\sigma\_\{IMA,e\}^\{2\}\}\{n\_\{i\}n\_\{a\}\}\([3](https://arxiv.org/html/2610.00651#S3.E3)\)EρMA\(b\)2E\\rho^\{2\}\_\{MA\(b\)\}Does benchmarkbbpreserve system order?σM2\+σA2\+σMA2\\sigma\_\{M\}^\{2\}\+\\sigma\_\{A\}^\{2\}\+\\sigma\_\{MA\}^\{2\}σIM2\+σIA2\+σIMA,e2ni\\dfrac\{\\sigma\_\{IM\}^\{2\}\+\\sigma\_\{IA\}^\{2\}\+\\sigma\_\{IMA,e\}^\{2\}\}\{n\_\{i\}\}\([4](https://arxiv.org/html/2610.00651#S3.E4)\)ρAA′\(b\)\\rho^\{\(b\)\}\_\{AA^\{\\prime\}\}Do scaffolds inducethe same model order?σM2\+σIM2ni\\sigma\_\{M\}^\{2\}\+\\dfrac\{\\sigma\_\{IM\}^\{2\}\}\{n\_\{i\}\}σMA2\+σIMA,e2ni\\sigma\_\{MA\}^\{2\}\+\\dfrac\{\\sigma\_\{IMA,e\}^\{2\}\}\{n\_\{i\}\}\([5](https://arxiv.org/html/2610.00651#S3.E5)\)EρM2E\\rho^\{2\}\_\{M\},SNRM\\text\{SNR\}\_\{M\}Does the leaderboardpreserve model order?σM2\\sigma\_\{M\}^\{2\}σBM2nb\+σMA2na\+σBMA2nbna\+σIM\[B\]2nbni\+σBIMA,e2nbnina\\dfrac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\dfrac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\+\\dfrac\{\\sigma\_\{BMA\}^\{2\}\}\{n\_\{b\}n\_\{a\}\}\+\\dfrac\{\\sigma\_\{IM\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\dfrac\{\\sigma\_\{BIMA,e\}^\{2\}\}\{n\_\{b\}n\_\{i\}n\_\{a\}\}\([7](https://arxiv.org/html/2610.00651#S3.E7)\)limni→∞EρM\(⋅\)2\\displaystyle\\lim\_\{n\_\{i\}\\rightarrow\\infty\}E\\rho\_\{M\(\\cdot\)\}^\{2\}Does having unlimited taskspreserve model order?σM2\\sigma\_\{M\}^\{2\}σMA2na\\dfrac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}orσBM2nb\+σMA2na\+σBMA2nbna\\frac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\+\\frac\{\\sigma\_\{BMA\}^\{2\}\}\{n\_\{b\}n\_\{a\}\}\([8](https://arxiv.org/html/2610.00651#S3.E8)\)Δρ\(M−A\)\(b\)2\{\\Delta\\rho^\{2\}\_\{\(M\-A\)\(b\)\}\}Do models have more benchmark\-relevant signal than scaffolds?σM2−σA2\\sigma\_\{M\}^\{2\}\-\\sigma\_\{A\}^\{2\}σMA2\+σIM2\+σIA2\+σIMA,e2ni\\sigma\_\{MA\}^\{2\}\+\\frac\{\\sigma\_\{IM\}^\{2\}\+\\sigma\_\{IA\}^\{2\}\+\\sigma\_\{IMA,e\}^\{2\}\}\{n\_\{i\}\}\(11\)ΔρI\(M−A\)\(b\)2\\Delta\\rho^\{2\}\_\{I\(M\-A\)\(b\)\}Do models have more task\-level signal than scaffolds?σM2\+σIM2ni−σA2−σIA2ni\\sigma\_\{M\}^\{2\}\+\\frac\{\\sigma\_\{IM\}^\{2\}\}\{n\_\{i\}\}\-\\sigma\_\{A\}^\{2\}\-\\frac\{\\sigma\_\{IA\}^\{2\}\}\{n\_\{i\}\}σMA2\+σIM2\+σIA2\+σIMA,e2ni\\sigma\_\{MA\}^\{2\}\+\\frac\{\\sigma\_\{IM\}^\{2\}\+\\sigma\_\{IA\}^\{2\}\+\\sigma\_\{IMA,e\}^\{2\}\}\{n\_\{i\}\}\(12\)ΔρM−A2\\Delta\\rho^\{2\}\_\{M\-A\}Do models have more leaderboard\-relevant signal than scaffolds?σM2−σA2\\sigma\_\{M\}^\{2\}\-\\sigma\_\{A\}^\{2\}σMA2\+σBM2nb\+σBA2nb\+σBMA2nb\+σIM\[B\]2nbni\+σIA\[B\]2nbni\+σBIMA,e2nbni\\sigma\_\{MA\}^\{2\}\+\\frac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{BA\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{BMA\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{IM\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\frac\{\\sigma\_\{IA\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\frac\{\\sigma\_\{BIMA,e\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\(13\)ΔρB\(M−A\)2\\Delta\\rho^\{2\}\_\{B\(M\-A\)\}Do models have more benchmark\-level signal than scaffolds?σM2\+σBM2nb−σA2−σBA2nb\\sigma\_\{M\}^\{2\}\+\\frac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\-\\sigma\_\{A\}^\{2\}\-\\frac\{\\sigma\_\{BA\}^\{2\}\}\{n\_\{b\}\}σMA2\+σBM2nb\+σBA2nb\+σBMA2nb\+σIM\[B\]2nbni\+σIA\[B\]2nbni\+σBIMA,e2nbni\\sigma\_\{MA\}^\{2\}\+\\frac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{BA\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{BMA\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{IM\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\frac\{\\sigma\_\{IA\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\frac\{\\sigma\_\{BIMA,e\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\(14\)ΔρI\[B\]\(M−A\)2\\Delta\\rho^\{2\}\_\{I\[B\]\(M\-A\)\}Do models have more task\-level signal than scaffolds?σM2\+σBM2nb\+σIM2nbni−σA2−σBA2nb−σIA2nbni\\sigma\_\{M\}^\{2\}\+\\frac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{IM\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\-\\sigma\_\{A\}^\{2\}\-\\frac\{\\sigma\_\{BA\}^\{2\}\}\{n\_\{b\}\}\-\\frac\{\\sigma\_\{IA\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}σMA2\+σBM2nb\+σBA2nb\+σBMA2nb\+σIM\[B\]2nbni\+σIA\[B\]2nbni\+σBIMA,e2nbni\\sigma\_\{MA\}^\{2\}\+\\frac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{BA\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{BMA\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{IM\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\frac\{\\sigma\_\{IA\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\frac\{\\sigma\_\{BIMA,e\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\(15\)
## Appendix DAblations and Sensitivity Analyses
The benchmark ecosystem is sparse, unbalanced, and only partially connected\. Tasks are nested within benchmarks, most model–scaffold combinations are absent, and many interaction levels are observed only once\. Consequently, a fully crossed variance decomposition could in principle be driven by a small number of influential observations or by a particularly informative benchmark\. We therefore examine robustness at four complementary levels: pointwise predictive stability, stability of latent rankings, leave\-one\-benchmark\-out sensitivity, and leave\-one\-benchmark–scaffold\-out sensitivity\.
### D\.1Reference indexed variance decomposition
For observationnn, letb\[n\]b\[n\],i\[n\]i\[n\],m\[n\]m\[n\], anda\[n\]a\[n\]denote its benchmark, task, model, and scaffold\. The reference model with additional indexing from equation[6](https://arxiv.org/html/2610.00651#S3.E6)is
Yn\\displaystyle Y\_\{n\}∼Bernoulli\(pn\),\\displaystyle\\sim\\operatorname\{Bernoulli\}\(p\_\{n\}\),logit\(pn\)\\displaystyle\\operatorname\{logit\}\(p\_\{n\}\)=μ\+ub\[n\]B\+ub\[n\],i\[n\]I\+um\[n\]M\+ua\[n\]A\\displaystyle=\\mu\+u^\{B\}\_\{b\[n\]\}\+u^\{I\}\_\{b\[n\],i\[n\]\}\+u^\{M\}\_\{m\[n\]\}\+u^\{A\}\_\{a\[n\]\}\+ub\[n\],m\[n\]BM\+ub\[n\],a\[n\]BA\+um\[n\],a\[n\]MA\\displaystyle\\quad\+u^\{BM\}\_\{b\[n\],m\[n\]\}\+u^\{BA\}\_\{b\[n\],a\[n\]\}\+u^\{MA\}\_\{m\[n\],a\[n\]\}\+ub\[n\],i\[n\],m\[n\]IM\+ub\[n\],i\[n\],a\[n\]IA\+ub\[n\],m\[n\],a\[n\]BMA,\\displaystyle\\quad\+u^\{IM\}\_\{b\[n\],i\[n\],m\[n\]\}\+u^\{IA\}\_\{b\[n\],i\[n\],a\[n\]\}\+u^\{BMA\}\_\{b\[n\],m\[n\],a\[n\]\},where every random effect is independently distributed asuℓF∼𝒩\(0,σF2\)u^\{F\}\_\{\\ell\}\\sim\\mathcal\{N\}\(0,\\sigma\_\{F\}^\{2\}\)for facet or interactionFF\. On the latent logistic scale, the observation\-level residual variance isπ2/3\\pi^\{2\}/3\. The reference fit containsN=29,923N=29\{,\}923binary observations, 54 models, 13 scaffolds, nine benchmarks, and 1,117 benchmark\-specific tasks\.
The central object of measurement is the model\. For a fixed evaluation designDD, the relative generalizability coefficient has the schematic form from Equation[1](https://arxiv.org/html/2610.00651#S2.E1)
ρD2=σM2σM2\+σδ,D2,\\rho\_\{D\}^\{2\}=\\frac\{\\sigma\_\{M\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{\\delta,D\}^\{2\}\},\(16\)whereσδ,D2\\sigma\_\{\\delta,D\}^\{2\}contains only those components that change model contrasts under repeated realizations ofDD, divided by their effective replication counts\. Components that shift all models equally under a shared condition do not enter the relative error variance\.
##### Cancellation of common condition effects\.
For two modelsmmandm′m^\{\\prime\}evaluated under the same benchmark and scaffold,
ηbima−ηbim′a\\eta\_\{bima\}\-\\eta\_\{bim^\{\\prime\}a\}eliminates the benchmark main effectubBu^\{B\}\_\{b\}, scaffold main effectuaAu^\{A\}\_\{a\}, benchmark–scaffold effectubaBAu^\{BA\}\_\{ba\}, task–scaffold effectubiaIA\[B\]u^\{IA\[B\]\}\_\{bia\}, and task main effectubiIu^\{I\}\_\{bi\}\. Thus, these components can affect absolute scores but not the ordering of models evaluated under identical conditions\. This distinction is important below: instability inσA2\\sigma\_\{A\}^\{2\}orσBA2\\sigma\_\{BA\}^\{2\}need not imply instability in relative model reliability\.
All posterior chains for the reference model mixed adequately: the reported variance parameters hadR^=1\.00\\widehat\{R\}=1\.00, with bulk effective sample sizes between 1,559 and 3,743\. The following analyses address robustness to the data and specification rather than only Monte Carlo convergence\.
### D\.2Sensitivity of model rankings
Standard leaderboards rank systems by empirical mean accuracy,141414e\.g\., for HAL see “Accuracy” at[https://hal\.cs\.princeton\.edu/reliability/](https://hal.cs.princeton.edu/reliability/)for HELM see “mean score” at[https://crfm\.stanford\.edu/helm/](https://crfm.stanford.edu/helm/), etc\.
y¯bm=1Nbm∑n:b\[n\]=b,m\[n\]=myn,\\bar\{y\}\_\{bm\}=\\frac\{1\}\{N\_\{bm\}\}\\sum\_\{n:\\,b\[n\]=b,m\[n\]=m\}y\_\{n\},possibly aggregating over different scaffolds and nonidentical task sets\. Such rankings treat all observed variation as evidence about model capability\. The mixed models instead estimate latent model effects after separating task difficulty, scaffold effects, and their interactions\.
For each benchmark, we formed rankings from posterior summaries of the relevant latent model effects\. We considered two estimators:
1. 1\.Full\-model latent ranking, obtained from the joint model in Equation[6](https://arxiv.org/html/2610.00651#S3.E6), which partially pools information through models and scaffolds appearing elsewhere in the benchmark ecosystem\.
2. 2\.Per\-benchmark latent ranking, obtained by fitting, within each benchmark, logit\(pima\)=μ\+uiI\+umM\+uaA\+uimIM\+uiaIA\+umaMA\.\\operatorname\{logit\}\(p\_\{ima\}\)=\\mu\+u^\{I\}\_\{i\}\+u^\{M\}\_\{m\}\+u^\{A\}\_\{a\}\+u^\{IM\}\_\{im\}\+u^\{IA\}\_\{ia\}\+u^\{MA\}\_\{ma\}\.This estimator uses only the information available within one benchmark\.
These posterior conditional means are the Bayesian analogue of best linear unbiased predictors \(BLUPs\)\. We compared their induced rankings with the rankings from observed mean scores using Kendall’sτ\\tau, Spearman’sρ\\rho, and bias\-corrected squared distance correlationdCorn2\\text\{dCor\}^\{2\}\_\{n\}computed on rank vectors\. Kendall’sτ\\taumeasures pairwise ordering agreement; Spearman’sρ\\rhomeasures monotone rank association; and corrected squared distance correlation can detect more general dependence between the rank assignments\.
Table 10:Association between observed rankings and latent\-effect rankings\. “Full” uses the ecosystem model \(Eq\.[6](https://arxiv.org/html/2610.00651#S3.E6)\); “Per\-benchmark” uses an independently fitted model for each benchmark \(Eq\.[2](https://arxiv.org/html/2610.00651#S3.E2)\)EstimatorMetricAssistantBenchCORE\-HardGAIAMind2WebSciCodeScienceAgentBenchSWE\-miniτ\\tau\-benchUSACOFulldCorn2\{\}^\{2\}\_\{n\}\.404\.899\.780\.872\.703\.651\.833\.828\.430Per\-benchmarkdCorn2\{\}^\{2\}\_\{n\}\-\.005\.107\.046\.039\.013\.007\.046\.051\.032FullKendallτ\\tau\.478\.859\.771\.744\.731\.582\.764\.787\.590Per\-benchmarkKendallτ\\tau\.081\.248\.196\.151\.128\.101\.167\.180\.142FullSpearmanρ\\rho\.619\.961\.900\.896\.867\.795\.916\.930\.731Per\-benchmarkSpearmanρ\\rho\.111\.340\.271\.211\.171\.137\.236\.255\.199
Two findings are salient\. First, the jointly estimated rankings retain substantial agreement with observed leaderboards for CORE\-Bench Hard, GAIA, Online\-Mind2Web, SciCode, SWE\-bench Verified Mini, andτ\\tau\-Bench Airline\. For example, full\-model Spearman correlations are at least0\.850\.85on these benchmarks\. Their empirical rankings therefore contain recoverable cross\-system signal even after nuisance variation is separated\.
Second, the independently fitted per\-benchmark models exhibit weaker correspondence with raw rankings on every benchmark\. This result is not evidence that the per\-benchmark models are necessarily “less accurate\.” Instead, it exposes an information and estimand mismatch\. Within one sparse benchmark, model effects must be separated from model–scaffold and task–model interactions using few connected observations\. Strong posterior shrinkage is therefore appropriate, but it compresses distinctions that empirical average accuracy treats as model signal\. Moreover, the raw ranking may combine deployable model–scaffold performance, whereas the latent model ranking intentionally removes scaffold\-specific contributions\. The joint model can recover more stable model effects because shared models and scaffolds connect otherwise isolated benchmark\-specific designs\. In other words, the combined model, after accounting for sources of variation, provides more stable and reliable estimates by taking advantage of shared facet variation\.
Accordingly, rank correlation with the observed leaderboard is a sensitivity diagnostic, not a ground\-truth accuracy measure\. High agreement indicates that adjustment for known nuisance facets preserves the reported ordering\. Low agreement indicates that the ordering depends on whether task and scaffold variation are treated as capability or as measurement error\. It does not establish which ordering is externally valid without an independent criterion\.
### D\.3Pointwise predictive stability with PSIS\-LOO
We first assessed whether individual observations exert disproportionate influence on the posterior predictive distribution\. Exact leave\-one\-out cross\-validation would requireNNrefits\. We instead used Pareto\-smoothed importance\-sampling leave\-one\-out cross\-validation \(PSIS\-LOO\), with moment matching for observations having unstable importance ratios\([Vehtari et al\., 2017](https://arxiv.org/html/2610.00651#bib.bib32);[Paananen et al\., 2021](https://arxiv.org/html/2610.00651#bib.bib33);[Vehtari et al\., 2024](https://arxiv.org/html/2610.00651#bib.bib34)\)\. For observationnn, PSIS approximates
p\(yn∣y−n\)=∫p\(yn∣θ\)p\(θ∣y−n\)𝑑θp\(y\_\{n\}\\mid y\_\{\-n\}\)=\\int p\(y\_\{n\}\\mid\\theta\)\\,p\(\\theta\\mid y\_\{\-n\}\)\\,d\\thetausing draws from the full posteriorp\(θ∣y\)p\(\\theta\\mid y\)\. The generalized Pareto shape diagnostick^n\\widehat\{k\}\_\{n\}measures the tail behavior of the resulting importance ratios\. Valuesk^n≤0\.7\\widehat\{k\}\_\{n\}\\leq 0\.7generally indicate a reliable approximation, whereas larger values identify observations for which deleting the point substantially changes its posterior predictive distribution\.
Table 11:PSIS\-LOO diagnostics for the full Bayesian logistic mixed model\.QuantityEstimateStandard errorelpdloo\\operatorname\{elpd\}\_\{\\mathrm\{loo\}\}−11,217\.1\-11\{,\}217\.194\.294\.2ploop\_\{\\mathrm\{loo\}\}3,032\.03\{,\}032\.032\.032\.0LOOIC\\operatorname\{LOOIC\}22,434\.222\{,\}434\.2188\.5188\.5Pareto diagnosticCountPercentagek^≤0\.7\\widehat\{k\}\\leq 0\.729,90229\{,\}902\>99\.9%\>99\.9\\%0\.7<k^≤10\.7<\\widehat\{k\}\\leq 12121<0\.1%<0\.1\\%k^\>1\\widehat\{k\}\>1000\.0%0\.0\\%The diagnostics in Table[11](https://arxiv.org/html/2610.00651#A4.T11)are favorable for a model of this complexity\. The effective number of parameters is substantially smaller than bothNNand the nominal number of coefficients induced by the random effects, demonstrating strong regularization through partial pooling\. No observation hask^\>1\\widehat\{k\}\>1, and only 21 of 29,923 observations havek^\>0\.7\\widehat\{k\}\>0\.7\.
The few warnings are explained by the connectivity of the design rather than by broad model failure\. Approximately32%32\\%of observed benchmark–task–model groups are singletons, a known point of sensitivity for PSIS\-LOO\([Bindoff, 2026](https://arxiv.org/html/2610.00651#bib.bib35)\)\. Of the 21 observations withk^\>0\.7\\widehat\{k\}\>0\.7, 19 \(90\.5%90\.5\\%\) belong to such singleton groups\. The remaining two form the only observations in their group: SciCode task 14 evaluated with GPT\-4\.1 under two scaffolds\.
This behavior follows directly from the pointwise LOO target\. Ifyny\_\{n\}is the only observation informing a random\-effect leveluℓu\_\{\\ell\}, deletion ofyny\_\{n\}makes the leave\-one\-out distribution ofuℓu\_\{\\ell\}approximately prior\-predictive\. The full\-data posterior, by contrast, has adapteduℓu\_\{\\ell\}toyny\_\{n\}\. Importance sampling must therefore bridge two meaningfully different distributions, producing a heavy\-tailed importance ratio\. A largek^n\\widehat\{k\}\_\{n\}in this setting primarily says that the outcome of an isolated, data\-defined group cannot be predicted after removing its only observation\. It does not, by itself, imply that population\-level variance components are determined by that point\.
#### D\.3\.1Observation\-level random\-effect ablation\.
As an additional check, we augmented Equation[6](https://arxiv.org/html/2610.00651#S3.E6)with
ub,i,m,aIMA∼𝒩\(0,σIMA2\)\.u^\{IMA\}\_\{b,i,m,a\}\\sim\\mathcal\{N\}\(0,\\sigma\_\{IMA\}^\{2\}\)\.Because benchmark–task–model–scaffold cells generally lack within\-cell replication in the HAL dataset, this term acts as an observation\-level random effect \(OLRE\)\([Harrison, 2015](https://arxiv.org/html/2610.00651#bib.bib36)\)\. In a Bernoulli\-logit model, it competes with the fixed latent logistic residual rather than identifying a conventionally replicated three\-way interaction\. Including it did not materially change the reliability estimates: it absorbed variation previously treated as observation\-level noise, and its contribution is attenuated by task replication in the D\-study\. The principal generalizability conclusions are therefore not an artifact of omitting the highest\-order cell term\.
Importantly, this additional term should be used for studies where output stability is of interest where benchmark–task–model–scaffold levels have multiple runs\. This variation can support measuring the stochasticity of task scores under the same conditions\.
### D\.4Leave\-one\-benchmark\-out sensitivity
Pointwise LOO evaluates interpolation within the observed ecosystem\. A stronger test removes an entire benchmark and therefore deletes all of its tasks, score distribution, and benchmark\-specific interactions\. For eachb⋆∈ℬb^\{\\star\}\\in\\mathcal\{B\}, we refitted the complete variance decomposition to
𝒟−b⋆=\{yn:b\[n\]≠b⋆\}\\mathcal\{D\}\_\{\-b^\{\\star\}\}=\\\{y\_\{n\}:b\[n\]\\neq b^\{\\star\}\\\}and recomputed the posterior variance decomposition and D\-study reliability curves\. This procedure asks whether conclusions about the benchmark battery are broadly distributed across benchmarks or are driven by a single test\.
Figure 6:Posterior distributions for benchmark\-level LOO ablations for the paper’s quantities of interestThe reliability estimations, proportions of variance attributable to models, and the principal interaction facets were generally stable across these refits as shown in Fig\.[6](https://arxiv.org/html/2610.00651#A4.F6)\. Sensitivity nevertheless depended on which benchmark was withheld\. Removing CORE\-Bench Hard caused the largest reduction in estimated reliability\. This is substantively expected: CORE\-Bench Hard supplies the most runs, spans the widest range of observed accuracies, and has the highest estimated signal\-to\-noise ratio among the included benchmarks\. It therefore contributes unusually strong information both for distinguishing models and for connecting model performance to the rest of the ecosystem\.
Without CORE\-Bench Hard, the posterior signal\-to\-noise ratio from the D\-study no longer reaches the prespecified limit\-of\-detection band\. This finding should not be read merely as “more observations are better\.” A large but noisy benchmark can contribute little to the reliability of a battery\. CORE\-Bench Hard is influential because it combines replication with discrimination\. Thus, increasing the number of benchmarks is not equivalent to increasing effective measurement information\.
More generally, for a battery aggregating conditionally independent benchmark\-level measurements with signalSbS\_\{b\}and errorEbE\_\{b\}, the information supplied by benchmarkbbis governed by its discrimination relative to error, not simply its task count\. Although the fully crossed design includes interactions and is more complicated than this schematic case, the same principle applies: benchmarks with substantial model variation and controlled task\- and harness\-dependent error contribute disproportionately to stable ecosystem\-level rankings\.
This ablation also clarifies the intended scope of generalization\. The leave\-one\-benchmark\-out results do not claim that the fitted model can predict an arbitrary future benchmark with no shared structure\. Rather, they test whether the estimated variance decomposition and design recommendations survive removal of one observed measurement instrument\. The sensitivity to CORE\-Bench Hard indicates that the present nine\-benchmark ecosystem contains limited redundancy at the high\-signal end\.
### D\.5Leave\-one\-benchmark–scaffold\-out sensitivity
Because scaffold coverage is also unbalanced, we conducted a complementary sensitivity analysis that removed one observed benchmark–scaffold combination at a time and re\-estimated the full decomposition\. Table[12](https://arxiv.org/html/2610.00651#A4.T12)summarizes the resulting frequentist proportions of total variance\. This ablation is especially stringent for benchmark\-specific scaffolds, whose removal can eliminate nearly all direct information about a scaffold level\.
Table 12:Leave\-one\-benchmark–scaffold\-out proportions of variance\. SE is the standard error across refits, and CV is the coefficient of variation\.FacetMeanMedianMin\.Max\.SDSECVResidual\.237\.235\.226\.267\.008\.002\.035Scaffold \(AA\)\.020\.019\.008\.037\.007\.002\.338Benchmark \(BB\)\.258\.259\.229\.298\.015\.003\.060Benchmark–scaffold \(BABA\)\.026\.028\.000\.035\.009\.002\.345Benchmark–model \(BMBM\)\.020\.020\.009\.031\.004\.001\.208Benchmark–model–scaffold \(BMABMA\)\.014\.013\.010\.019\.003\.001\.180Task within benchmark \(II\)\.328\.329\.282\.365\.015\.003\.047Task–scaffold \(IAIA\)\.052\.053\.028\.066\.009\.002\.168Task–model \(IMIM\)\.003\.004\.000\.006\.001<\.001<\.001\.426Model \(MM\)\.033\.033\.026\.044\.004\.001\.110Model–scaffold \(MAMA\)\.008\.009\.003\.012\.002<\.001<\.001\.249
Table 13:Leave\-one\-LLM\-out proportions of variance\. SD is the standard deviation across refits, SE is the standard error across refits, and CV is the coefficient of variation\.FacetMeanMedianMin\.Max\.SDSECVResidual0\.2360\.2360\.2310\.2410\.00200\.008Scaffold \(AA\)0\.0200\.0200\.0160\.0290\.00200\.098Benchmark \(BB\)0\.2590\.2590\.2520\.2680\.00200\.009Benchmark–scaffold \(BABA\)0\.0260\.0260\.0220\.0300\.00100\.056Benchmark–model \(BMBM\)0\.0200\.0200\.0150\.0220\.00100\.065Benchmark–model–scaffold \(BMABMA\)0\.0140\.0140\.0080\.0160\.00100\.092Task within benchmark \(II\)0\.3280\.3280\.3200\.3340\.00200\.006Task–scaffold \(IAIA\)0\.0520\.0520\.0490\.0530\.00100\.018Task–model \(IMIM\)0\.0040\.0040\.0010\.0040\.00100\.166Model \(MM\)0\.0330\.0330\.0280\.0360\.00200\.048Model–scaffold \(MAMA\)0\.0090\.0090\.0060\.0120\.00100\.107
The dominant components are stable\. Task difficulty accounts for approximately32\.8%32\.8\\%of variance across refits, benchmark differences for25\.8%25\.8\\%, and residual variation for23\.7%23\.7\\%\. Their coefficients of variation are only0\.0470\.047,0\.0600\.060, and0\.0350\.035, respectively\. The model component is also stable in absolute terms, ranging from0\.0260\.026to0\.0440\.044\.
The largest relative variation occurs for the scaffold main effect, the benchmark–scaffold interaction, and the task–model interaction\. These cases require different interpretations\. The first two are weakly identified because most scaffolds occur in only one benchmark; removing a benchmark–scaffold cell can therefore remove much of the relevant connectivity\. However, common scaffold and benchmark–scaffold shifts cancel from same\-condition model contrasts and consequently do not enter the relative model reliability estimates used in the primary D\-study\.
The task–model component has the largest coefficient of variation \(0\.4260\.426\), but its estimated proportion is always between00and0\.0060\.006\. Its high relative variability is therefore primarily a small\-denominator effect: even its maximum is smaller than the typical contribution of every other reported component\. Coefficients of variation should not be interpreted without the corresponding absolute scale\. Additionally, both this component and the second largest CV value, the benchmark–scaffold component, are the only components that contain minimum values equal to zero; these sensitivities, paired with extreme ablation minima, are likely elevated due to the nature of frequentist estimations, which can result in singular values in highly unbalanced designs \(see Appendix[G\.3](https://arxiv.org/html/2610.00651#A7.SS3)\)\. These estimates may be more stable under a full Bayesian ablation where CV estimates would be less likely to decrease the mean in the denominator\.
The model\-related components that directly govern ranking stability remain comparatively well behaved\. The model proportion has CV0\.1100\.110, while the benchmark–model and model–scaffold components remain small across all deletions\. Hence, no individual benchmark–scaffold cell appears to create the principal conclusion that bare\-model rankings are less reliable than rankings would appear under a decomposition that treats all observed model–scaffold performance as signal\. However, it is worth noting that removing AssistantBench during the LOO ablation disconnects the OnlineMind2Web scaffold from the broader network, which means that, for that connection, identifiability conditions found in Proposition[E\.3](https://arxiv.org/html/2610.00651#A5.Thmtheorem3)would not be complete\. Nevertheless, the estimates remain stable, even with this disconnection\.
Importantly, the contribution to overall reliability of CORE\-Bench Hard discussed in §[D\.4](https://arxiv.org/html/2610.00651#A4.SS4)can be narrowed further to the contribution of the CORE Agent, which shows the most between\-LLM discrimination on its benchmark\. Removing this particular benchmark–scaffold accounts for the large increase in uncertainty across the pooled model\. As this scaffold is less connected than the HAL Generalist scaffold, this reaffirms that having a strong discriminating signal is critical to pooled reliability, potentially more than large connectivity\.
### D\.6Leave\-one\-LLM\-out sensitivity
We next test whether the estimated measurement structure is driven by a single LLM and whether the variance decomposition generalizes to new LLMs\. For each LLM modelm⋆∈ℳm^\{\\star\}\\in\\mathcal\{M\}, we removed all of its observations,
𝒟−m⋆=\{yn:m\[n\]≠m⋆\},\\mathcal\{D\}\_\{\-m^\{\\star\}\}=\\\{y\_\{n\}:m\[n\]\\neq m^\{\\star\}\\\},and refit the variance\-decomposition model\. This is a stronger perturbation than deleting one response: it simultaneously removes a model main\-effect level and every observed benchmark–model, model–scaffold, task–model, and benchmark–model–scaffold cell involving that LLM\. Because evaluation coverage differs substantially across models—only nine of the 54 models appear in every benchmark—the ablation also tests sensitivity to the most highly connected models in the design\. The LOO ablations for this and those of Appendix[D\.5](https://arxiv.org/html/2610.00651#A4.SS5)were conducted in a frequentist framework to reduce practical computation costs of hundreds of Bayesian refittings \(see Appendix[C\.6](https://arxiv.org/html/2610.00651#A3.SS6)\)\.
Table[13](https://arxiv.org/html/2610.00651#A4.T13)shows that the decomposition is remarkably insensitive to the removal of any LLM\. The three largest components are almost invariant: task\-within\-benchmark variance has mean proportion0\.3280\.328, range0\.3200\.320–0\.3340\.334, and CV0\.0060\.006; residual variance has mean0\.2360\.236, range0\.2310\.231–0\.2410\.241, and CV0\.0080\.008; and benchmark variance has mean0\.2590\.259, range0\.2520\.252–0\.2680\.268, and CV0\.0090\.009\. Thus, the conclusion that variation is dominated by differences among tasks and benchmarks, together with substantial unexplained response variation, is not attributable to the performance profile of a particular model\.
More importantly for relative reliability, the model variance is also stable\. Its proportion remains between0\.0280\.028and0\.0360\.036, with mean0\.0330\.033and CV0\.0480\.048\. The benchmark–model component ranges from0\.0150\.015to0\.0220\.022, while the benchmark–model–scaffold component ranges from0\.0080\.008to0\.0160\.016\. Removing even a highly connected or unusually capable LLM therefore does not qualitatively change the estimated amount of systematic between\-model variation or the extent to which model performance depends on the benchmark and scaffold\.
As in the benchmark–scaffold ablation, the largest relative variability occurs in small components\. The task–model interaction has CV0\.1660\.166, but its variance share is only0\.0010\.001–0\.0040\.004\. The model–scaffold interaction has CV0\.1070\.107and remains between0\.0060\.006and0\.0120\.012\. Their relative sensitivity should consequently not be confused with a large contribution to total variance\. In absolute terms, deletion\-induced changes in both components are small\.
The leave\-one\-model results are also more stable than the leave\-one\-benchmark results\. This asymmetry reflects the structure of the available evidence\. Each LLM contributes another sample from the population of systems, and its outcomes are partially pooled with those of the remaining 53 models\. By contrast, removing a benchmark deletes an entire measurement instrument, all of its unique tasks, and its characteristic signal\-to\-noise ratio\. The ecosystem therefore has greater redundancy across models than across high\-quality benchmarks\. Adding another LLM primarily improves estimation of the distribution of model capability, whereas adding a discriminating benchmark can alter the quality of the measurement battery itself\.
##### Implication for reliability\.
For a fixed D\-study design, relative reliability depends on the model variance and on model\-dependent error components such asBMBM,MAMA,IMIM, andBMABMA, after scaling by their effective replication counts\. The stability of these components under model deletion implies that the reported reliability conclusions are not generated by one extreme or unusually well\-connected LLM\. Nevertheless, this ablation evaluates*influence on the population variance decomposition*, not the ability to predict the performance or rank of a previously unseen model\. Generalization to a new LLM additionally requires that it be exchangeable with the sampled model population and evaluated on conditions that connect it to the existing design\.
No single LLM acts as a leverage point for the principal variance decomposition\. The greater sensitivity to removing an informative benchmark than to removing an LLM reinforces a central design recommendation: once a reasonably diverse model sample has been obtained, additional evaluation resources may yield greater reliability gains by improving benchmark quality, task replication, and cross\-scaffold connectivity than by adding sparsely evaluated models\.
### D\.7Methodological Robustness
We also estimate each model using five different approaches\. A separate Appendix[G](https://arxiv.org/html/2610.00651#A7)explains and discusses each method and reports the full results\.
### D\.8Rank stability across posterior
Ranking models within each draw yields posterior rank distributions and pairwise ordering probabilitiesPr\(θmb\>θm′b∣y\)\.\\Pr\(\\theta\_\{mb\}\>\\theta\_\{m^\{\\prime\}b\}\\mid y\)\.These quantities distinguish an estimated ordering from evidence that two models are meaningfully distinguishable\. We compare posterior and published rankings using Spearman correlation and ranked unbiased squared distance correlation,dCorn2\\operatorname\{dCor\}\_\{n\}^\{2\}; the latter remains informative in the presence of extensive ties\.
We compare posterior ranks with published ranks using Spearman correlation and ranked unbiased squared distance correlation, dCorn2\{\}^\{2\}\_\{n\}\(Table[14](https://arxiv.org/html/2610.00651#A4.T14)\)\. The latter is useful for benchmarks with extensive ties and provides a direct diagnostic of whether estimated capability ranking is statistically associated with the reported ordering\.
Posterior capability estimates differ from ranks obtained by sorting raw percent\-correct scores\. Raw scores credit the model for every condition with which it happens to be paired\. The pooled decomposition instead estimates the portion of performance that persists after averaging over the specified benchmark and scaffold universes\.
Table 14:Median correlation of reported and estimated ranks across posterior drawsCorr\.AssistantBenchCORE\-BenchGAIAMind2webSciCodeSci\.AgentBenchSWEbenchτ\\tau\-benchUSACOSpearman0\.4420\.8240\.790\.6210\.6620\.6110\.8080\.6830\.692dCorn2\{\}^\{2\}\_\{n\}0\.1410\.640\.5760\.330\.360\.3230\.6330\.4080\.395
### D\.9Implications for benchmark design
These checks support five practical conclusions\.
1. 1\.Sparse cells principally limit local prediction\.PSIS warnings are almost entirely confined to singleton or doubleton interaction levels\. The model cannot predict an isolated cell after its sole observation is removed, but the population\-level variance decomposition is stable to these observations\.
2. 2\.Partial pooling is necessary for ecosystem\-level ranking\.Per\-benchmark data are generally insufficient to cleanly distinguish model capability from scaffold and task interactions\. Cross\-benchmark connectivity substantially stabilizes latent model rankings, although the resulting rankings remain uncertain on low\-signal benchmarks\.
3. 3\.Benchmark quality is not interchangeable with benchmark quantity\.The leave\-one\-benchmark\-out analysis identifies CORE\-Bench Hard as a high\-information anchor\. A useful test battery should include multiple independently constructed benchmarks with high discrimination and controlled error, rather than merely adding more noisy tasks or near\-duplicate benchmarks\.
4. 4\.Absolute\-score instability and relative\-rank instability are distinct\.Scaffold and benchmark–scaffold effects can materially shift reported accuracies while canceling from same\-condition model comparisons\. Benchmark reports should therefore state whether reliability concerns absolute deployment performance, model ranking, or model–scaffold system ranking; these are different objects of measurement and induce different error terms\.
5. 5\.Representative LLM panel improves benchmark diagnostics\.once a reasonably diverse model sample has been obtained, additional evaluation resources may yield greater reliability gains by improving benchmark quality, task replication, and cross\-scaffold connectivity than by adding sparsely evaluated models\.
The robustness analyses do not imply that the sparse design is harmless\. Rather, they localize its consequences\. The main variance and reliability conclusions are not driven by a handful of observations or benchmark–scaffold cells, but the precision of individual rankings remains strongly dependent on cross\-benchmark connectivity and on the inclusion of at least one high\-signal measurement instrument\. This distinction is essential for designing future agentic benchmark batteries: additional evaluation should be allocated to conditions that improve connectivity and reduce model\-dependent error, not only to increasing the nominal number of tasks or sparsely connecting models\.
## Appendix EIdentifiability and connectivity of the leaderboard decomposition
The leaderboard data are sparse by construction: models are submitted selectively, scaffolds are not used uniformly, and tasks are unique to benchmarks\. Such missingness is common in public leaderboards\([Singh et al\., 2026](https://arxiv.org/html/2610.00651#bib.bib44)\)\. A fully crossed experiment would simplify estimation, but it is not necessary for the questions studied here\. What is required is sufficient overlap to distinguish persistent model variation from variation associated with benchmarks, scaffolds, and their interactions\.
This appendix establishes that the observed design provides the connectivity needed for those contrasts\. We distinguish three concepts that are often conflated:
1. 1\.Connectivity:whether observed model, benchmark, and scaffold levels belong to a common comparison network\.
2. 2\.Structural identifiability:whether distinct variance components imply distinct distributions over the observed responses\.
3. 3\.Practical estimability:whether the finite data determine those components precisely\.
Connectivity and structural identifiability are design properties; practical estimability also depends on sample size, outcome variation, and prior regularization\. Our Bayesian estimator can produce a proper posterior for a weakly informed component, but a proper posterior alone is not evidence that the component is strongly data\-identified\. We therefore use posterior uncertainty and sensitivity analyses, rather than existence of an estimate, to characterize practical estimability\.
### E\.1Incidence\-Graph Connectivity and Identifiability
Represent the observed design as a multipartite incidence graph whose vertices are benchmarks, models, and scaffolds, with edges induced by observed evaluations\. Shared models and scaffolds connect benchmark\-specific observations and support estimation of common variance components\. Contrasts within a connected component are informed by observed paths; contrasts across disconnected components are not identified without additional assumptions\. We report component membership, articulation vertices, benchmark degrees, and the change in connectivity produced by removing each benchmark\. Posterior regularization stabilizes weakly supported components but does not create evidence for disconnected contrasts\.
### E\.2Observed incidence structure
The full dataset contains nine benchmarks, 54 LLMs, and 13 scaffolds\. Nine models appear on every benchmark and span at least69%69\\%of the scaffold set\. One scaffold is shared by eight benchmarks\. AssistantBench connects the remaining benchmark through an additional shared scaffold, and every other benchmark contains at least one additional benchmark\-specific scaffold\. Models and scaffolds are consequently neither fully crossed nor evenly replicated\.
Let
Ω=\{\(b,i,m,a\):ybimais observed\}\\Omega=\\left\\\{\(b,i,m,a\):y\_\{bima\}\\ \\text\{is observed\}\\right\\\}denote the observed response cells\. The pooled decomposition is equation[6](https://arxiv.org/html/2610.00651#S3.E6):
ηbima=\\displaystyle\\eta\_\{bima\}=\{\}β0\+ub\(B\)\+ubi\(I\[B\]\)\+um\(M\)\+ua\(A\)\+ubm\(BM\)\+uba\(BA\)\+uma\(MA\)\\displaystyle\\beta\_\{0\}\+u\_\{b\}^\{\(B\)\}\+u\_\{bi\}^\{\(I\[B\]\)\}\+u\_\{m\}^\{\(M\)\}\+u\_\{a\}^\{\(A\)\}\+u\_\{bm\}^\{\(BM\)\}\+u\_\{ba\}^\{\(BA\)\}\+u\_\{ma\}^\{\(MA\)\}\+ubim\(IM\[B\]\)\+ubia\(IA\[B\]\)\+ubma\(BMA\),\(b,i,m,a\)∈Ω\.\\displaystyle\+u\_\{bim\}^\{\(IM\[B\]\)\}\+u\_\{bia\}^\{\(IA\[B\]\)\}\+u\_\{bma\}^\{\(BMA\)\},\\qquad\(b,i,m,a\)\\in\\Omega\.Each random effect has mean zero and a component\-specific variance\. The Bernoulli–logit likelihood fixes the latent scale through the standard logistic residual varianceπ2/3\\pi^\{2\}/3\.
It is useful to representΩ\\Omegaas a multipartite incidence graph\. Let
G=\(V,E\),V=ℬ∪ℳ∪𝒜,G=\(V,E\),\\qquad V=\\mathcal\{B\}\\cup\\mathcal\{M\}\\cup\\mathcal\{A\},where an edge joins two levels when they co\-occur in at least one observed response\. For example,mm–bbis an edge if modelmmis evaluated on benchmarkbb, andaa–bbis an edge if scaffoldaais used on benchmarkbb\. Items need not connect across benchmarks because they are intentionally nested within benchmark\.
The nine models observed on every benchmark form model\-side anchors\. The scaffold used on eight benchmarks forms a scaffold\-side anchor, while the second shared scaffold connects AssistantBench to the rest of the graph\. Benchmark\-specific scaffolds are leaves or local branches attached to this connected core\. Thus, all benchmarks and their associated observations belong to one comparison network\.
##### Lemma 1 \(connected additive contrasts\)\.
###### Lemma E\.1\(connected additive contrasts\)\.
Consider an additive model on an observed incidence graph,
ηma=μ\+αm\+γa,\\eta\_\{ma\}=\\mu\+\\alpha\_\{m\}\+\\gamma\_\{a\},with one centering constraint per facet\. If the model–scaffold incidence graph is connected, all estimable contrastsαm−αm′\\alpha\_\{m\}\-\\alpha\_\{m^\{\\prime\}\}andγa−γa′\\gamma\_\{a\}\-\\gamma\_\{a^\{\\prime\}\}are identified\.
###### Proof\.
Suppose two parameterizations produce the same linear predictor on every observed edge\. Their differences satisfyΔαm\+Δγa=0\\Delta\\alpha\_\{m\}\+\\Delta\\gamma\_\{a\}=0on each edge\. Along any path in a connected bipartite graph, these equalities imply that all model differences equal a common constant and all scaffold differences equal its negative\. Centering removes this remaining additive degree of freedom\. ∎
Lemma[E\.1](https://arxiv.org/html/2610.00651#A5.Thmtheorem1)concerns fixed additive effects, whereas Equation[6](https://arxiv.org/html/2610.00651#S3.E6)uses random effects and interactions\. It nevertheless provides the relevant intuition: an observation from a model or scaffold that is disconnected from the remainder of the leaderboard cannot support a common ranking\. Shared models and scaffolds create paths along which relative effects can be compared\.
### E\.3Identifiability of random\-effect variances
For observations ordered as a vector, write the latent linear predictor as
𝜼=β0𝟏\+∑k=1KZk𝐮k,𝐮k∼𝒩\(𝟎,σk2I\),\\bm\{\\eta\}=\\beta\_\{0\}\\mathbf\{1\}\+\\sum\_\{k=1\}^\{K\}Z\_\{k\}\\mathbf\{u\}\_\{k\},\\qquad\\mathbf\{u\}\_\{k\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\sigma\_\{k\}^\{2\}I\),whereZkZ\_\{k\}is the incidence matrix for componentkk\. On the latent scale, the random effects induce covariance
Cov\(𝜼\)=∑k=1Kσk2Kk,Kk=ZkZk⊤\.\\operatorname\{Cov\}\(\\bm\{\\eta\}\)=\\sum\_\{k=1\}^\{K\}\\sigma\_\{k\}^\{2\}K\_\{k\},\\qquad K\_\{k\}=Z\_\{k\}Z\_\{k\}^\{\\top\}\.\(17\)The entries ofKkK\_\{k\}indicate which pairs of observations share a level of facet or interactionkk\. For example, two observations share the model kernelKMK\_\{M\}when they use the same model, and shareKBMK\_\{BM\}only when they use both the same benchmark and model\.
###### Proposition E\.2\(variance\-component criterion\)\.
A collection of latent variance components\{σk2:k∈𝒦\}\\\{\\sigma\_\{k\}^\{2\}:k\\in\\mathcal\{K\}\\\}is structurally identifiable from the observed design if its covariance kernels\{Kk:k∈𝒦\}\\\{K\_\{k\}:k\\in\\mathcal\{K\}\\\}, restricted toΩ\\Omega, are linearly independent after removing components that are deterministically confounded with the terminal residual\.
###### Proof\.
If the restricted kernels are linearly independent, equality of two induced covariance matrices implies
∑k∈𝒦\(σk2−σ~k2\)Kk=0,\\sum\_\{k\\in\\mathcal\{K\}\}\(\\sigma\_\{k\}^\{2\}\-\\widetilde\{\\sigma\}\_\{k\}^\{2\}\)K\_\{k\}=0,which has only the trivial solutionσk2=σ~k2\\sigma\_\{k\}^\{2\}=\\widetilde\{\\sigma\}\_\{k\}^\{2\}for everykk\. If the kernels are linearly dependent, a nonzero perturbation of their coefficients leaves the induced covariance unchanged, so the corresponding components cannot be separated from the observed design\. ∎
For a Bernoulli GLMM, Proposition[E\.2](https://arxiv.org/html/2610.00651#A5.Thmtheorem2)is most directly interpreted as a design criterion on the latent scale\. The nonlinear likelihood can affect the amount of information, especially under floor effects, but it cannot create distinctions absent from the incidence matrices\.
A practical interpretation is the following: to distinguish two variance components, the design must contain pairs of observations that share the grouping represented by one component without always sharing the grouping represented by the other\. Complete crossing is sufficient for this condition but is not necessary\.
### E\.4How the observed design separates the required components
Table[15](https://arxiv.org/html/2610.00651#A5.T15)summarizes the principal replication patterns\. These conditions concern variance components, not estimation of every individual random\-effect realization\.
Table 15:Observed comparisons supporting the pooled variance decomposition\. “Separation” identifies the principal pair of components distinguished by each overlap pattern\.SeparationRequired comparisonSupport in the observed designMMversusBMBMThe same model appears on multiple benchmarks\.Nine models appear on all nine benchmarks; additional models provide partial cross\-benchmark replication\.AAversusBABAThe same scaffold appears on multiple benchmarks\.One scaffold spans eight benchmarks; an additional shared scaffold connects AssistantBench\.MAMAversusBMABMAThe same model–scaffold pair recurs across benchmarks\.Cross\-benchmark models evaluated through shared scaffolds create repeated model–scaffold pairs\.MMversusMAMAModels are observed under multiple scaffolds, with overlap across models\.The cross\-benchmark models span at least nine of the 13 scaffolds, and shared scaffolds provide common comparison conditions\.BMBMversusBMABMAWithin a benchmark, models are observed under more than one scaffold\.Each benchmark contains shared or locally replicated scaffold conditions in addition to benchmark\-specific scaffolds\.I\[B\]I\[B\]versusIM\[B\]IM\[B\]Each item is attempted by multiple models\.Benchmark items are repeatedly scored across leaderboard models\.I\[B\]I\[B\]versusIA\[B\]IA\[B\]Each item is attempted under multiple scaffold conditions\.Scaffold replication within benchmarks provides item–scaffold contrasts where observed\.#### E\.4\.1Model and benchmark–model variation
The distinction between persistent model capability and benchmark\-specific performance is central to the leaderboard reliability coefficient\. If every model appeared on only one benchmark, thenum\(M\)u\_\{m\}^\{\(M\)\}andubm\(BM\)u\_\{bm\}^\{\(BM\)\}would be inseparable: “model” would be nested within benchmark\. That failure does not occur here\. Nine models appear on every benchmark, producing observations that shareMMwhile differing inBMBM\.
Consequently, the model kernelKMK\_\{M\}and benchmark–model kernelKBMK\_\{BM\}have different support\. The former links observations from the same model across benchmarks; the latter links them only within a benchmark\. This overlap identifies the contrast between globally persistent model variationσM2\\sigma\_\{M\}^\{2\}and benchmark\-conditioned model variationσBM2\\sigma\_\{BM\}^\{2\}\.
#### E\.4\.2Scaffold and benchmark–scaffold variation
Benchmark\-specific scaffolds alone would not distinguish a scaffold main effect from a benchmark–scaffold interaction\. For a scaffold used on exactly one benchmark, itsAAandBABAcolumns coincide\. The design avoids complete confounding because one scaffold spans eight benchmarks and a second shared scaffold connects AssistantBench\. Observations using a shared scaffold have the sameAAlevel but differentBABAlevels, providing the comparisons required to distinguishσA2\\sigma\_\{A\}^\{2\}fromσBA2\\sigma\_\{BA\}^\{2\}\.
Benchmark\-specific scaffolds remain less individually informed than shared scaffolds\. Their effects are estimated through the exchangeability assumptions of the hierarchical model, with uncertainty propagated into the posterior\. The shared scaffolds identify the population\-level separation; partial pooling regularizes levels with little direct replication\.
#### E\.4\.3Model–scaffold and benchmark–model–scaffold variation
SeparatingMAMAfromBMABMArequires recurrence of a model–scaffold pair across benchmarks\. The cross\-benchmark models and shared scaffolds provide such recurrence: observations can share a model and scaffold while differing in benchmark\. Within\-benchmark scaffold variation then supplies the complementary comparisons needed to estimate how model–scaffold compatibility changes by benchmark\.
This component is more weakly informed than the model and benchmark–model components because repeated model–scaffold pairs are less frequent than repeated models\. Our estimands therefore integrate over its posterior uncertainty\. The leave\-one\-benchmark\-out and alternative\-estimator analyses in Appendix[D](https://arxiv.org/html/2610.00651#A4)assess whether substantive conclusions depend on a small number of these bridges\.
#### E\.4\.4Item\-indexed variation
Items are nested within benchmarks and are not expected to recur across benchmarks\. Their main effects are identified because the same item is attempted by multiple models and scaffold conditions\. Item–model variation is informed when a benchmark item is attempted by multiple models, while item–scaffold variation is informed by scaffold replication within benchmark\.
No assumption is made that an item from one benchmark is exchangeable with an identically labeled item from another benchmark\. Pooling occurs at the variance\-component level: the model estimates the typical magnitude of item\-conditioned interactions across the observed benchmark panel\.
### E\.5What is not separately identifiable
The design does not contain independent repeated executions of every\(b,i,m,a\)\(b,i,m,a\)cell\. A four\-way benchmark–item–model–scaffold interaction is therefore observationally confounded with cell\-level execution variation and the Bernoulli residual for most of such interactions\. We combine these sources into the terminal component
σBIMA,e2\.\\sigma\_\{BIMA,e\}^\{2\}\.This is not a limitation for the reported relative reliability coefficients: all unresolved cell\-specific variation belongs in the error term because it can change model ordering across repeated evaluation conditions\. Separating execution stochasticity, grading instability, and four\-way interaction would require repeated runs under the same recorded cell and, ideally, explicit grading replicates\.
More generally, the present design does not support:
- •an unregularized estimate for every unobserved model–scaffold cell;
- •precise scaffold\-specific effects for scaffolds used on only one benchmark;
- •causal claims about replacing one scaffold with another; or
- •generalization to scaffold or benchmark populations unrelated to the connected evaluation panel\.
These quantities are unnecessary for our research questions\. We require population\-level variance components and scaffold\-marginalized model contrasts, not predictions for every missing factorial cell\.
### E\.6Identifiability of the reliability estimands
The principal leaderboard coefficient is Equation[7](https://arxiv.org/html/2610.00651#S3.E7):
EρM2\(nb,ni,na\)=σM2σM2\+σBM2nb\+σMA2na\+σBMA2nbna\+σIM\[B\]2nbni\+σBIMA,e2nbnina\.E\\rho\_\{M\}^\{2\}\(n\_\{b\},n\_\{i\},n\_\{a\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\frac\{\\sigma\_\{BM\}^\{2\}\}\{n\_\{b\}\}\+\\frac\{\\sigma\_\{MA\}^\{2\}\}\{n\_\{a\}\}\+\\frac\{\\sigma\_\{BMA\}^\{2\}\}\{n\_\{b\}n\_\{a\}\}\+\\frac\{\\sigma\_\{IM\[B\]\}^\{2\}\}\{n\_\{b\}n\_\{i\}\}\+\\frac\{\\sigma\_\{BIMA,e\}^\{2\}\}\{n\_\{b\}n\_\{i\}n\_\{a\}\}\}\.This coefficient depends on a small set of variance components and their sums, not on every random effect in Equation[6](https://arxiv.org/html/2610.00651#S3.E6)\. The design requirements are correspondingly weaker than those needed to reconstruct the complete model–benchmark–scaffold response tensor\.
###### Proposition E\.3\(sufficiency for the leaderboard estimand\)\.
Suppose that:
1. \(i\)the benchmark–model incidence graph is connected and contains models observed on multiple benchmarks;
2. \(ii\)the benchmark–scaffold incidence graph is connected and contains scaffolds observed on multiple benchmarks;
3. \(iii\)at least some model–scaffold pairs recur across benchmarks; and
4. \(iv\)benchmark items are attempted by multiple models and scaffolds\.
Then the covariance kernels corresponding toMM,BMBM,MAMA,BMABMA, andIM\[B\]IM\[B\]have distinct observed replication patterns\. The variance combinations entering Equation[7](https://arxiv.org/html/2610.00651#S3.E7)are therefore structurally distinguishable, up to the explicitly combined terminal componentσBIMA,e2\\sigma\_\{BIMA,e\}^\{2\}\.
JustificationConditions \(i\)–\(iv\) provide observation pairs that share, respectively:MMbut notBMBM;AAbut notBABA;MAMAbut notBMABMA; andIM\[B\]IM\[B\]and not onlyI\[B\]I\[B\]\. These yield distinct covariance kernels under Proposition[E\.2](https://arxiv.org/html/2610.00651#A5.Thmtheorem2)\. The cell\-specific remainder is intentionally aggregated because no additional replication distinguishes its constituents\.□\\square
The observed leaderboard satisfies these conditions\. In particular, the nine cross\-benchmark models identify persistent versus benchmark\-conditioned model variation, while shared scaffolds and repeated model–scaffold pairs identify the scaffold\-related components that determine the task\-only reliability ceiling\.
###### Corollary E\.4\(identifiability of D\-study ceilings\)\.
Under Proposition[E\.3](https://arxiv.org/html/2610.00651#A5.Thmtheorem3), the task\-only limit
limni→∞EρM2\(nb,ni,na\)=σM2σM2\+σBM2/nb\+σMA2/na\+σBMA2/\(nbna\)\\lim\_\{n\_\{i\}\\rightarrow\\infty\}E\\rho\_\{M\}^\{2\}\(n\_\{b\},n\_\{i\},n\_\{a\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{BM\}^\{2\}/n\_\{b\}\+\\sigma\_\{MA\}^\{2\}/n\_\{a\}\+\\sigma\_\{BMA\}^\{2\}/\(n\_\{b\}n\_\{a\}\)\}is identified without separately decomposing terminal item\-level residual variation\.
Corollary[E\.4](https://arxiv.org/html/2610.00651#A5.Thmtheorem4)is important for the paper’s central design conclusion\. The claim that adding tasks cannot eliminate benchmark\- and scaffold\-conditioned error does not depend on a precise decomposition of every item\-level noise source\. It depends on the persistent components whose identification is supported by cross\-benchmark models and scaffolds\.
### E\.7Bayesian regularization and evidential scope
All variance components are estimated with proper, weakly informative priors\. The posterior is therefore proper even when a component is weakly informed near the boundary\. We use priors to stabilize finite\-sample estimation, not to substitute for disconnected comparisons\. Three features constrain interpretation:
1. 1\.Data\-supported contrasts\.Model rankings are inferred through the connected observation graph\. We do not report comparisons between disconnected components because none exist in the observed design\.
2. 2\.Posterior rather than plug\-in reliability\.Equation[7](https://arxiv.org/html/2610.00651#S3.E7)is evaluated within each posterior draw\. Weak separation among related components therefore appears as uncertainty in reliability and its D\-study projection\.
3. 3\.Sensitivity to bridges\.Leave\-one\-benchmark\-out refits test whether conclusions rely on a particular benchmark or scaffold connection\. Prior and estimator comparisons test whether weak components are driving the reported variance ordering\. These analyses are reported in Appendix[D](https://arxiv.org/html/2610.00651#A4)\.
The distinction between identifiability and precision is especially important for agent leaderboards\. Sparse designs can be connected enough to estimate whether scaffold variation is non\-negligible while still being too sparse to estimate its exact magnitude narrowly\. Large credible intervals in this setting are not model failure; they reveal that the leaderboard contains too few common evaluation conditions to support precise claims\.
### E\.8Connectivity Summary
The full design is incomplete but connected\. Nine models evaluated across every benchmark anchor the model scale; shared scaffolds connect the benchmark panel; and recurring model–scaffold pairs distinguish general scaffold compatibility from benchmark\-specific compatibility\. These overlaps are sufficient for the variance combinations underlying model\-ranking reliability, signal\-to\-noise ratios, scaffold\-versus\-model comparisons, and D\-study ceilings\.
The design does not identify every possible cell effect, nor is that required\. Unreplicated terminal variation is assigned to relative error, and uncertainty from sparsely observed scaffold interactions is propagated through the Bayesian posterior\. The resulting claims are therefore appropriately scoped: they characterize the reliability supported by this connected leaderboard ecosystem, rather than asserting a complete factorial decomposition of all possible models, scaffolds, tasks, and benchmarks\.
## Appendix FAdditional Tables and Figures
The coefficients of Table[1](https://arxiv.org/html/2610.00651#S3.T1)isolate*where*a benchmark’s signal lives and*which*facet dissipates it\. Because each coefficient shares a common numerator–denominator structure \(signal variance over signal plus generalizing error\), differences across coefficients within a single benchmark are directly interpretable: they reveal how much of the apparent capability signal survives when we generalize over tasks, over scaffolds, or over both\. Table[17](https://arxiv.org/html/2610.00651#A6.T17)reports posterior means from the per\-benchmark fit of Eq\.[2](https://arxiv.org/html/2610.00651#S3.E2)\.
Table 16:Reliability depends on the object being ranked\.Reliability estimates for each benchmark under model and model–scaffold systems as measurement systems, including the task limit for models as the object of measurement\.EstimandAssistantCORE\-HardGAIAMind2WebSciCodeScienceAgentSWE\-miniτ\\tau\-AirlineUSACOlimni→∞EρM2\\lim\_\{n\_\{i\}\\rightarrow\\infty\}E\\rho\_\{M\}^\{2\}\([8](https://arxiv.org/html/2610.00651#S3.E8)\)0\.2100\.8920\.3250\.1530\.8400\.6830\.5520\.6690\.377EρM2E\\rho\_\{M\}^\{2\}\([3](https://arxiv.org/html/2610.00651#S3.E3)\)0\.1820\.8410\.3180\.1480\.8110\.6280\.5430\.5720\.375EρMA2E\\rho\_\{MA\}^\{2\}\([4](https://arxiv.org/html/2610.00651#S3.E4)\)0\.9350\.9720\.9900\.9740\.9750\.9870\.9930\.9440\.994
Figure 7:Benchmarks differ sharply in ranking signal\.Posterior model\-ranking signal\-to\-noise ratio as tasks are added\. Shaded bands give external detection\-limit reference values\. Points and intervals are posterior medians and68%68\\%standard error HDIs\.Table 17:Posterior\-mean reliability and generalizability coefficients per benchmark\.EρM\(i\)2E\\rho^\{2\}\_\{M\(i\)\}andEρA\(i\)2E\\rho^\{2\}\_\{A\(i\)\}are task reliabilities \(σo2/\(σo2\+σio2\)\\sigma^\{2\}\_\{o\}/\(\\sigma^\{2\}\_\{o\}\+\\sigma^\{2\}\_\{io\}\)\) for model and scaffold rankings;ρAA′\(b\)\\rho^\{\(b\)\}\_\{AA^\{\\prime\}\}is inter\-scaffold reliability \(Eq\.[5](https://arxiv.org/html/2610.00651#S3.E5)\);EρM\(b\)2E\\rho^\{2\}\_\{M\(b\)\},EρA\(b\)2E\\rho^\{2\}\_\{A\(b\)\}, andEρMA\(b\)2E\\rho^\{2\}\_\{MA\(b\)\}are G\-coefficients by object of measurement \(Eqs\.[3](https://arxiv.org/html/2610.00651#S3.E3)–[4](https://arxiv.org/html/2610.00651#S3.E4)\); the last column is the task\-saturated model ceilinglimni→∞EρM\(b\)2\\lim\_\{n\_\{i\}\\to\\infty\}E\\rho^\{2\}\_\{M\(b\)\}\(Eq\.[8](https://arxiv.org/html/2610.00651#S3.E8)\)\.BenchmarkEρM\(i\)2E\\rho^\{2\}\_\{M\(i\)\}EρA\(i\)2E\\rho^\{2\}\_\{A\(i\)\}ρAA′\(b\)\\rho^\{\(b\)\}\_\{AA^\{\\prime\}\}EρM\(b\)2E\\rho^\{2\}\_\{M\(b\)\}EρA\(b\)2E\\rho^\{2\}\_\{A\(b\)\}EρMA\(b\)2E\\rho^\{2\}\_\{MA\(b\)\}EρM\(b\)2\(∞\)E\\rho^\{2\}\_\{M\(b\)\}\(\\infty\)scicode0\.6020\.5340\.7460\.7360\.5160\.9710\.765gaia0\.2630\.7050\.3330\.3280\.5790\.9900\.334taubench\_airline0\.5490\.5390\.5450\.5330\.7590\.9380\.619swebench\_verified\_mini0\.5990\.8570\.5040\.4990\.6390\.9920\.506usaco0\.2380\.2530\.3920\.4290\.4780\.9930\.432assistantbench0\.3620\.4450\.2870\.2600\.5210\.9110\.305corebench\_hard0\.6360\.9410\.8280\.8170\.8050\.9720\.866onlinemind2web0\.2960\.2280\.2050\.2030\.4300\.9720\.210scienceagentbench0\.5580\.7530\.5670\.5610\.8810\.9820\.610Figure 8:Reliability by Object of Measurement across all tasks in a benchmark\. Estimated using variance components of Equation[2](https://arxiv.org/html/2610.00651#S3.E2), and posteriors of Equations[3](https://arxiv.org/html/2610.00651#S3.E3)and[4](https://arxiv.org/html/2610.00651#S3.E4)\. Intervals are pseudo\-standard error at 0\.68 HDI, bar estimates are medians and points are means of posterior draw distributions\.Table 18:Internal Score Transportability\.Kendall’s tau correlation matrix of benchmarks and summary statistics against each benchmark’s mean accuracy, the unweighted HAL mean score, and reliability\-adjusted latent \(θ^\\widehat\{\\theta\}\) model effect \(median posterior\)AssistantCORE\-HardGAIAMind2WebSciCodeScienceAgentSWE\-miniτ\\tau\-AirlineUSACOMean Scoreθ^\\widehat\{\\theta\}assistantbench1\.00\.3000\.485\-0\.0610\.1030\.215\-0\.4040\.1800\.2150\.2050\.261corebench\_hard0\.3001\.00\.4870\.4000\.4920\.4580\.1480\.5810\.1670\.6870\.835gaia0\.4850\.4871\.00\.0280\.2600\.4790\.2150\.3630\.1820\.4530\.606onlinemind2web\-0\.0610\.4000\.0281\.00\.1710\.0560\.4240\.315\-0\.1410\.5810\.348scicode0\.1030\.4920\.2600\.1711\.00\.1840\.1170\.4770\.2250\.2020\.472scienceagentbench0\.2150\.4580\.4790\.0560\.1841\.00\.1290\.587\-0\.0850\.3920\.499swebench\_verified\_mini\-0\.4040\.1480\.2150\.4240\.1170\.1291\.00\.0230\.0320\.5090\.554taubench\_airline0\.1800\.5810\.3630\.3150\.4770\.5870\.0231\.00\.0730\.4040\.588usaco0\.2150\.1670\.182\-0\.1410\.225\-0\.0850\.0320\.0731\.00\.3900\.338
### F\.1Repeated Task Subsampling
We confirm the cost reduction findings by repeatedly samplingkktasks per benchmark without replacement, using the same sampled tasks for all models to ensure a paired comparison\. Within each subsample, scores are first averaged at the benchmark–model level and then averaged across benchmarks, giving each benchmark equal weight\. The resulting model ranking is compared with the ranking obtained from the complete dataset using Spearman’s rank correlation coefficient \(ρ\\rho\) and Kendall’s rank correlation coefficient \(τ\\tau\)\. Repeating this procedure 500 times across many random subsamples yields a distribution of rank correlations for eachkk, quantifying how stable the model rankings are with respect to the number and selection of evaluated tasks\. Task subsampling corroborates the diminishing returns\. Across500500paired subsamples,1515tasks per benchmark retain mean Spearman agreement of0\.9320\.932with the complete\-data ranking at an estimated82%82\\%cost reduction;3030tasks increase agreement to0\.9760\.976Agreement with the full ranking establishes information retention, but makes not claim to ranking validity\.
Table 19:Task subsampling shows rank information preservation\.Bootstrapped correlations of subsampled tasks and leaderboard Rankskmeanρmedianρρminρmaxmeanτmedianττminτmaxtop@1 agreementmean abs rank changecost savings50\.7900\.8010\.5840\.9170\.6340\.6400\.4580\.7660\.6407\.2210\.94100\.8910\.8960\.8040\.9520\.7420\.7450\.6460\.8290\.7565\.1100\.88150\.9320\.9360\.8700\.9660\.7990\.8020\.7200\.8580\.8344\.0520\.82200\.9540\.9570\.9150\.9780\.8370\.8390\.7780\.8870\.9123\.3230\.76250\.9670\.9690\.9400\.9820\.8630\.8650\.8110\.9010\.9602\.8220\.69300\.9760\.9770\.9570\.9880\.8860\.8890\.8420\.9230\.9902\.3850\.63
Figure 9:Proportion of Variance Explained in the posterior distributions from Eq\.[2](https://arxiv.org/html/2610.00651#S3.E2)\. “Item” from benchmarking literature represents “task” commonly found in agentic evaluation literature\. “Agent” represents “scaffold”\.Figure 10:Posterior density of Model\-Agent difference of rank\-relevant variance across the leaderboard \(left\), across a sampled benchmark \(middle\), and across a sampled task \(right\)\. Positive numbers suggest LLM variance contribution is greater than that of agent scaffold\.Figure 11:Published score versus posterior model percentile ranks using model\-wise mean published score \(x\-axis\) and draw\-wise posterior median rankings with 95% CIs\. Vertical intervals are95%95\\%posterior rank intervals; jitter separates horizontal ties\.Figure 12:Rank order changes by benchmark based on comparison with median posterior rank from full variance decomposition model of Eq\.[6](https://arxiv.org/html/2610.00651#S3.E6)\.Figure 13:Cumulative posterior dominance rankings across individuals\. The point estimate for each individualmmrepresents their global dominance score, defined as the expected probability that their latent parameterθm=um\(M\)\+umb\(MB\)\\theta\_\{m\}=u\_\{m\}^\{\(M\)\}\+u\_\{mb\}^\{\(\}MB\)is strictly greater than that of a randomly selected peerm′m^\{\\prime\}from the population, given the observed datayy:Global Dominance\(m\)=1M−1∑m′≠mPr\(θm\>θm′∣y\)\\text\{Global Dominance\}\(m\)=\\frac\{1\}\{M\-1\}\\sum\_\{m^\{\\prime\}\\neq m\}\\Pr\(\\theta\_\{m\}\>\\theta\_\{m^\{\\prime\}\}\\mid y\)whereMMis the total number of individuals, and the pairwise probabilities are empirically derived via Bayesian draws as:Pr\(θm\>θm′∣y\)≈1D∑d=1D𝕀\(θm\(d\)\>θm′\(d\)\)\\Pr\(\\theta\_\{m\}\>\\theta\_\{m^\{\\prime\}\}\\mid y\)\\approx\\frac\{1\}\{D\}\\sum\_\{d=1\}^\{D\}\\mathbb\{I\}\\left\(\\theta\_\{m\}^\{\(d\)\}\>\\theta\_\{m^\{\\prime\}\}^\{\(d\)\}\\right\)with𝕀\(⋅\)\\mathbb\{I\}\(\\cdot\)denoting the indicator function andDDrepresenting the total number of posterior draws\. Individuals are arranged along the vertical axis in ascending order of their global dominance scores\. The vertical dashed line atPr=0\.5\\Pr=0\.5denotes the theoretical baseline of a perfectly average individual who exhibits structural parity relative to the rest of the cohort\. Deviations toward1\.01\.0signify deterministic stochastic dominance over the population, whereas values approaching0\.00\.0indicate systematic underperformance relative to the group\.Figure 14:Pairwise Posterior Dominance, where color is Pr\(Row \> Column∣y\\mid y\) across posterior draws\. The point estimate for each individualmmrepresentsPr\(θm\>θm′∣y\)\\Pr\(\\theta\_\{m\}\>\\theta\_\{m^\{\\prime\}\}\\mid y\), defined by the expected probability of latent parameterθm=um\(M\)\+umb\(MB\)\\theta\_\{m\}=u\_\{m\}^\{\(M\)\}\+u\_\{mb\}^\{\(\}MB\), for each benchmarkbb\.
## Appendix GMethod and Scale Comparison
### G\.1Latent Space Estimations
Variance components in our generalized linear mixed model \(GLMM\) decompositions live on the*latent*\(logit\) scale, soρ2\\rho^\{2\}should be interpreted as reliability of the linear predictorη\\eta\(i\.e\., of differences in*log\-odds*\) rather than as reliability of raw percent\-correct\. This is standard in binary\-response measurement: the latent scale is the scale on which additive random effects and G\-theory D\-study algebra apply, and it is the scale on which “true score \+ error” is well\-defined without probability\-dependent heteroskedasticity\. Importantly, a given amount of latent variability implies different variability in percent\-correct depending on where a model sits on the sigmoid: nearp≈0\.5p\\approx 0\.5, small changes inη\\etatranslate to large changes inpp, while nearp≈0p\\approx 0or11they compress\. For that reason, a latent\-scale reliability such asρ2≈0\.45\\rho^\{2\}\\approx 0\.45*does*correspond to meaningful rank instability on the probability scale\. The latent\-scaleρ2\\rho^\{2\}is a mathematically coherent reliability target for binary data, and the associated posterior predictive mapping quantifies how it manifests as practically relevant leaderboard instability in percent\-correct\.
In addition to the interpretation benefits and alignment to measurement theoretic latent modeling, the logistic mixed effect approach also handles unbalanced data and values near zero, both common in leaderboard testing, better than the observation\-level linear approach\. This can be seen in Figure[15](https://arxiv.org/html/2610.00651#A7.F15), which shows the latent and observed modeling estimated reliabilities\. Rank order reliabilities are higher for the latent estimations proportionate to the quantity of low scoring models \(mean score per modely¯mb⪅0\.1\\bar\{y\}\_\{mb\}\\lessapprox 0\.1\)\. This is particularly true of SciCode where 100% of models have an accuracy score less than 0\.1\.
Figure 15:Rank order reliability, G\-coefficient, per benchmark, for both latent and observed estimations with LLM as object of measurement\. Bars are medians, points are means, and intervals are pseudo\-standard error 0\.68 HDI across the posterior using Eq\.[3](https://arxiv.org/html/2610.00651#S3.E3)#### G\.1\.1Variance shares and reliability on the latent scale
For GLMM/Bayesian GLMM, we report variance shares as proportions of total latent variance:
πk=σk2∑jσj2,\\pi\_\{k\}\\;=\\;\\frac\{\\sigma\_\{k\}^\{2\}\}\{\\sum\_\{j\}\\sigma\_\{j\}^\{2\}\},where the sum runs over included random effects and \(optionally\) a latent residual term\. For Bayesian fits,πk\\pi\_\{k\}is computed per posterior draw, yielding credible intervals\. These can be found in Tables[20](https://arxiv.org/html/2610.00651#A7.T20)and[21](https://arxiv.org/html/2610.00651#A7.T21)\.
##### D\-study scaling\.
When an interaction term involves a sampled facet, its contribution to the variance of an averaged score shrinks with the number of sampled levels\. For example, with tasks averaged \(nin\_\{i\}tasks\), anI×MI\\times Mcomponent contributesσIM2/ni\\sigma^\{2\}\_\{IM\}/n\_\{i\}\.
### G\.2Distance components \(DISCO\) quasi\-reliability
The DISCO decomposition\([Rizzo and Székely, 2010](https://arxiv.org/html/2610.00651#bib.bib9)\)generalizes classical ANOVA to arbitrary metric spaces\. ForKKgroups withnkn\_\{k\}observations each, define the within\-group dispersion as
𝒮W=∑k=1K1nk∑i<j‖Xki−Xkj‖,\\mathcal\{S\}\_\{W\}=\\sum\_\{k=1\}^\{K\}\\frac\{1\}\{n\_\{k\}\}\\sum\_\{i<j\}\\\|X\_\{ki\}\-X\_\{kj\}\\\|,\(18\)and the total dispersion as
𝒮T=1N∑i<j‖Xi−Xj‖,\\mathcal\{S\}\_\{T\}=\\frac\{1\}\{N\}\\sum\_\{i<j\}\\\|X\_\{i\}\-X\_\{j\}\\\|,\(19\)whereN=∑knkN=\\sum\_\{k\}n\_\{k\}\. The between\-group component is𝒮B=𝒮T−𝒮W\\mathcal\{S\}\_\{B\}=\\mathcal\{S\}\_\{T\}\-\\mathcal\{S\}\_\{W\}, and the proportion attributable to the grouping factor is𝒮B/𝒮T\\mathcal\{S\}\_\{B\}/\\mathcal\{S\}\_\{T\}\.
For multi\-facet designs, we compute DISCO sequentially for each facet, using the dispersion ratio as the analog ofη2\\eta^\{2\}\. Because DISCO does not produce a residual term, the facet proportions do not generally sum to one when computed independently; we normalize to aid comparison with the parametric estimates\.
DISCO decomposes dispersion using pairwise distances rather than squared deviations, improving robustness under non\-normality and heavy\-tailed random effects\. For an object groupinggg\(e\.g\., model or model–scaffold pair\), DISCO yields between\-group and within\-group dispersion terms\(Tg,Twithin\)\(T\_\{g\},T\_\{\\text\{within\}\}\), from which we define
Eρ2=TgTg\+Twithin\.E\\rho^\{2\}\\;=\\;\\frac\{T\_\{g\}\}\{T\_\{g\}\+T\_\{\\text\{within\}\}\}\.We use DISCO primarily as a methodological sensitivity analysis to validate that conclusions \(dominant facets; low model\-ranking reliability\) are not artifacts of Gaussian random\-effect assumptions\.
### G\.3Why estimation method changes conclusions—and what to do about it
The five estimation methods are not interchangeable, and their disagreements are informative\.
A notable outcome is that variance shares differ across linear mixed model \(LMM\), Bayesian LMM, generalized linera mixed model \(GLMM\), Bayesian GLMM, and DISCO, especially under sparse observations:
- •LMMtends to allocate a large portion to residualσ2\\sigma^\{2\}, which can understate structured interactions when the link is misspecified for Bernoulli data\.
- •GLMMcan collapse some components toward zero \(boundary estimates\) in unbalanced designs, yielding deceptively “clean” decompositions\.
- •Bayesian LMMallows for estimates of reliability and distributional uncertainty on observed scale, but less equipped for floor or ceiling effects or highly imbalanced data\.
- •Bayesian GLMMexposes skewness and large uncertainty, which is appropriate when the design under\-identifies components\.
- •DISCOcan attribute more to interaction\-like dispersion without parametric assumptions, often aligning with the intuition that agentic pipelines create nonlinear, heteroskedastic effects\.
Practical guidance\.
1. 1\.UseBayesianornonparametricmethods to communicate uncertainty when data are sparse; prefer point\-estimate GLMM/LMM only when coverage is dense\.
2. 2\.Triangulate: when all four methods agree on the*ranking*of dominant facets \(e\.g\., task vs interactions vs residual\), recommendations are robust; when they disagree, the correct conclusion is that the benchmark is under\-instrumented for that inference\.
##### Linear vs\. generalized models\.
The LME estimates generally attribute the largest share of variance to the residual \(with a borderline exception of USACO\), with correspondingly compressed named components\. The GLME and Bayesian estimates, by contrast, allocate substantially more variance to specific tasks and models\. This divergence is expected: the Gaussian identity\-link model treats binary\{0,1\}\\\{0,1\\\}responses as continuous, and its residual absorbs the Bernoulli variance floor \(p\(1−p\)p\(1\-p\)\) that the logit\-link models separate out via the distributional assumption\. The practical implication is that*linear G\-theory estimates applied to pass/fail benchmarks systematically underestimate the proportion of variance attributable to named facets and overestimate residual noise*, leading to artificially compressed—but not necessarily more conservative—reliability estimates\.
##### Bayesian vs\. frequentist GLME\.
The Bayesian and GLME estimates agree in broad strokes but diverge in two instructive ways\. First, the Bayesian posteriors for agent variance are markedly right\-skewed, with posterior means 2–5×\\timeslarger than posterior medians \(e\.g\., SWE\-bench agent mean=0\.27=0\.27, median=0\.19=0\.19\)\. The GLME point estimate, which approximates the posterior mode, misses this tail mass and underestimates the expected contribution of scaffold variance\. Second, the credible intervals from the Bayesian analysis expose the*decision\-relevant uncertainty*: for several benchmarks, the 95% HDI for model variance includes zero, meaning the data are consistent with no true model differentiation at all\. This uncertainty is invisible in frequentist analyses\.
##### Nonparametric DISCO\.
The DISCO estimates depart most dramatically from the parametric methods on the interaction terms, difference likely due to heterogeneous cluster imbalances \(and hence our preference for Bayesian estimates\)\. DISCO attributes 40–57% of total dispersion to task×\\timesmodel interactions across benchmarks, compared to 1–11% from the Bayesian estimates\. This divergence likely reflects two mechanisms: \(i\) DISCO’s sensitivity to nonlinear dependencies that the additive random\-effects models cannot capture, and \(ii\) the absence of a separate residual term in DISCO, which forces unexplained variation into the named interactions\. The practical upshot is that DISCO serves as a useful*upper bound*on interaction effects and a reminder that the additive decomposition assumed by mixed\-effects models may understate the complexity of the task×\\timesmodel relationship\.
### G\.4Variance decomposition explains the reliability ceiling and generalizability
Figure[16](https://arxiv.org/html/2610.00651#A7.F16)reports posterior variance shares from the leaderboard model\. Persistent model variation is smaller than the combined variation associated with scaffolds, benchmark\-conditioned model performance, and item\-conditioned interactions\. Most observed item outcomes therefore contain more condition\-specific variation than globally transferable model signal\.
Figure 16:Full variance decomposition of Eq\.[6](https://arxiv.org/html/2610.00651#S3.E6)as percentages of total variance\. Variance decompositions for individual benchmarks are in Appendix[F](https://arxiv.org/html/2610.00651#A6)\. Black represents residual variance\.This does not imply that model capability is unimportant\. Reliability is population\-dependent: contemporary frontier models are often similar, whereas tasks and scaffolds are heterogeneous\. A benchmark could have ranked a historically broader model population reliably and yet fail to distinguish the current frontier\. Likewise, a future benchmark may become uninformative as models saturate it\. Reliability must consequently be re\-estimated as the evaluated population and agent ecosystem change\.
For strong generalization to new tasks, we would want to see larger main effects for models and scaffolds, making decisions clearer\. However, the reality is that the second\-order task\-level interactions, model–task and scaffold–task, represent larger variation components than their respective main effects\. This quantifies that, instead of providing clear architecture choices, some models or scaffolds are better at certain tasks than others, implying that a developer should, for a given task, choose a specific model and scaffold for each new task, rather than for the group of tasks represented by a benchmark or leaderboard \(because of a smaller benchmark–model component\)\. This is not ideal for developers, as it means they would need to know about the demands of each task and the appropriateness of architecture selections for it\. This partly explains why there is such low measurement reliability for fewer numbers of items in Figures[1](https://arxiv.org/html/2610.00651#S5.F1)and[3](https://arxiv.org/html/2610.00651#S5.F3): a different model scaffold might be optimal for each task item\.
### G\.5Absolute Reliability on Linear Scale
Generalizability theory also permits estimation of reliability of numeric scores\. Absolute reliability, as it is known, is always less than or equal to the corresponding estimate of relative reliability\. But because scores are on the observed scale, all of the analyses in this section are based on the Bayesian linear mixed model \(§[C\.2](https://arxiv.org/html/2610.00651#A3.SS2)\)\. They are illustrative and serve as a framework for practitioners looking to measure reliability of assigned numeric scores\.
##### Reliablerankingsdo not imply reliablescores\.
Relative reliability \(Eρ2E\\rho^\{2\}\) asks whether the ranking of models stays similar across evaluation conditions, while absolute reliability \(Φ\\Phi\) asks whether their reported scores stay similar\. Figure[18](https://arxiv.org/html/2610.00651#A7.F18)shows that this distinction matters substantially in practice: absolute reliability is consistently lower across the nine benchmarks, and for several benchmarks remains near zero even as ranking reliability increases with additional tasks\. An evaluation can therefore support a stable ranking without supporting equally strong claims about a model’s absolute score or whether it exceeds a fixed threshold\.
While not recommended for the HAL dataset and benchmarks, there may be some agentic benchmarks where absolute scores are needed for decision\-making\. Absolute reliability is most interpretably performed on the observed scale, rather than the latent scale used throughout the main body of this study\. For the HAL dataset, there is simply not enough signal relative to noise in the benchmarks for any reliable absolute scoring\. Nevertheless, the concept of absolute reliability may be important for some agentic measures\.
#### G\.5\.1Absolute reliability: numeric values of scores carry meaning
Rank\-based reliability measures whether an evaluation preserves the ordering of models or systems, but it does not establish whether their score levels are reproducible\. Absolute reliability is important when scores are interpreted directly—for example, when assessing whether a model exceeds a deployment threshold, meets a minimum capability requirement, or achieves a specified performance target\. In these settings, a task set or scaffold that shifts every model’s score equally can change the decision even without changing the ranking\. The absolute reliability, or dependability coefficient,Φ\\Phi, therefore includes these common shifts in its error variance\.
Treating tasks and scaffolds as random facets, the absolute reliability of a model score under an equal\-allocation design is
ΦM\(b\)\(ni,na\)=σM2σM2\+\(σI2\+σIM2\)/ni\+\(σA2\+σMA2\)/na\+\(σIA2\+σIMA,e2\)/\(nina\)\.\\Phi\_\{M\(b\)\}\(n\_\{i\},n\_\{a\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\(\\sigma\_\{I\}^\{2\}\+\\sigma\_\{IM\}^\{2\}\)/n\_\{i\}\+\(\\sigma\_\{A\}^\{2\}\+\\sigma\_\{MA\}^\{2\}\)/n\_\{a\}\+\(\\sigma\_\{IA\}^\{2\}\+\\sigma\_\{IMA,e\}^\{2\}\)/\(n\_\{i\}n\_\{a\}\)\}\.\(20\)Relative to model\-ranking reliability, the denominator additionally includes task main effects, scaffold main effects, and task–scaffold interactions\. These components affect the absolute score of a model averaged over sampled tasks and scaffolds, even when they do not affect its position relative to other models\. Increasing the number of tasks reduces task\-related error, whereas increasing the number of scaffolds reduces scaffold\-related error\.
When the object of measurement is the completemodel–scaffold system, scaffold differences are part of the system’s universe score rather than measurement error\. Absolute reliability is then
ΦMA\(b\)\(ni\)=σM2\+σA2\+σMA2σM2\+σA2\+σMA2\+\(σI2\+σIM2\+σIA2\+σIMA,e2\)/ni\.\\Phi\_\{MA\(b\)\}\(n\_\{i\}\)=\\frac\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{A\}^\{2\}\+\\sigma\_\{MA\}^\{2\}\}\{\\sigma\_\{M\}^\{2\}\+\\sigma\_\{A\}^\{2\}\+\\sigma\_\{MA\}^\{2\}\+\(\\sigma\_\{I\}^\{2\}\+\\sigma\_\{IM\}^\{2\}\+\\sigma\_\{IA\}^\{2\}\+\\sigma\_\{IMA,e\}^\{2\}\)/n\_\{i\}\}\.\(21\)Here, task main effects enter the denominator because sampling an easier or harder task set can shift a system’s score across an absolute decision threshold\. In contrast, scaffold main effects and model–scaffold interactions remain signal because they distinguish the systems being evaluated\.
For either object of measurement, absolute reliability is no greater than the corresponding ranking reliability:
ΦM\(b\)\(ni,na\)≤EρM\(b\)2\(ni,na\),ΦMA\(b\)\(ni\)≤EρMA\(b\)2\(ni\)\.\\Phi\_\{M\(b\)\}\(n\_\{i\},n\_\{a\}\)\\leq E\\rho\_\{M\(b\)\}^\{2\}\(n\_\{i\},n\_\{a\}\),\\qquad\\Phi\_\{MA\(b\)\}\(n\_\{i\}\)\\leq E\\rho\_\{MA\(b\)\}^\{2\}\(n\_\{i\}\)\.\(22\)A large gap indicates that rankings may be reproducible even though reported score levels are sensitive to the sampled evaluation conditions\. Likewise, high inter\-scaffold ranking reliability does not guarantee absolute agreement: two scaffolds can preserve model ordering while producing systematically different scores\.
Figure 17:Standardized Absolute Inter\-scaffold Reliability\. Standardized inter\-scaffold reliability looks at the reliability for a single fixed task, rather than the composite reliability across all tasks in a benchmark\.Figure 18:Absolute vs Relative Reliability on Observed Scale
### G\.6Full Estimation Tables for Proportion of Latent Variation Explained
Full variance\-proportion tables for all models and all five estimation methods are provided in the following tables including Bayesian posterior summaries and also LME, GLME, and DISCO estimates to enable cross\-method comparison\.
Table 20:Proportion of variation explained per benchmark\. For comparability across methods, Bayesian estimates represent posterior means \(rather than medians found in most of the study\) to preserve proportional relationships\.BenchmarkFacetLMMBayes LMMGLMMBayes GLMMDISCOscicodeitem0\.2350\.2320\.7660\.7180\.209scicodeagent0\.0010\.0410\.0040\.0540\.001scicodemodel0\.0100\.0110\.0530\.0570\.012scicodeitem\_agent0\.0370\.0360\.0120\.0160\.252scicodeitem\_model0\.1300\.1230\.0120\.0350\.509scicodemodel\_agent0\.0020\.0040\.0030\.0150\.017scicodesigma0\.5840\.5530\.1500\.105NAgaiaitem0\.2100\.1130\.3420\.2800\.169gaiaagent0\.0410\.4710\.0290\.1940\.017gaiamodel0\.0300\.0200\.0400\.0420\.046gaiaitem\_agent0\.0360\.0200\.0350\.0360\.209gaiaitem\_model0\.0970\.0510\.0410\.1020\.481gaiamodel\_agent0\.0590\.0390\.0830\.0790\.077gaiasigma0\.5270\.2860\.4310\.267NAtaubench\_airlineitem0\.1660\.1400\.2770\.2440\.161taubench\_airlineagent0\.0430\.2370\.0440\.1800\.021taubench\_airlinemodel0\.0270\.0240\.0390\.0350\.025taubench\_airlineitem\_agent0\.0860\.0700\.1070\.1010\.250taubench\_airlineitem\_model0\.0350\.0240\.0000\.0330\.484taubench\_airlinemodel\_agent0\.0120\.0120\.0160\.0200\.059taubench\_airlinesigma0\.6310\.4930\.5180\.386NAswebench\_verified\_miniitem0\.1530\.0700\.3490\.2770\.092swebench\_verified\_miniagent0\.2460\.6450\.1030\.2600\.073swebench\_verified\_minimodel0\.0560\.0270\.2720\.1420\.119swebench\_verified\_miniitem\_agent0\.0680\.0340\.0240\.0260\.190swebench\_verified\_miniitem\_model0\.0890\.0390\.0090\.0590\.399swebench\_verified\_minimodel\_agent0\.0570\.0360\.0420\.1300\.126swebench\_verified\_minisigma0\.3310\.1490\.2020\.106NAusacoitem0\.2570\.1410\.5510\.4500\.239usacoagent0\.0720\.4690\.0000\.0900\.003usacomodel0\.0620\.0220\.0000\.0490\.027usacoitem\_agent0\.2000\.1140\.1300\.1720\.262usacoitem\_model0\.1840\.1030\.0000\.1160\.440usacomodel\_agent0\.0000\.0270\.1030\.0660\.030usacosigma0\.2250\.1240\.2160\.057NAassistantbenchitem0\.1160\.0740\.4370\.3310\.156assistantbenchagent0\.0000\.3770\.0000\.1810\.002assistantbenchmodel0\.0000\.0030\.0000\.0240\.017assistantbenchitem\_agent0\.0830\.0570\.1070\.1370\.219assistantbenchitem\_model0\.0470\.0220\.0000\.0570\.574assistantbenchmodel\_agent0\.0130\.0090\.0690\.0580\.032assistantbenchsigma0\.7410\.4580\.3860\.213NAcorebench\_harditem0\.2280\.1780\.4150\.3620\.153corebench\_hardagent0\.0630\.2860\.0470\.1800\.027corebench\_hardmodel0\.0710\.0560\.1340\.1130\.074corebench\_harditem\_agent0\.0100\.0090\.0000\.0050\.200corebench\_harditem\_model0\.0990\.0740\.0290\.0630\.462corebench\_hardmodel\_agent0\.0120\.0120\.0110\.0160\.083corebench\_hardsigma0\.5170\.3850\.3650\.262NAonlinemind2webitem0\.2830\.2050\.4140\.3720\.239onlinemind2webagent0\.0000\.2720\.0000\.0850\.000onlinemind2webmodel0\.0020\.0040\.0030\.0090\.007onlinemind2webitem\_agent0\.1330\.0970\.1640\.1630\.296onlinemind2webitem\_model0\.0220\.0140\.0000\.0230\.446onlinemind2webmodel\_agent0\.0160\.0130\.0270\.0290\.012onlinemind2websigma0\.5450\.3950\.3930\.319NAscienceagentbenchitem0\.2460\.1290\.5320\.4130\.208scienceagentbenchagent0\.0420\.5050\.0480\.2580\.011scienceagentbenchmodel0\.0160\.0080\.0260\.0200\.016scienceagentbenchitem\_agent0\.0860\.0450\.0490\.0460\.249scienceagentbenchitem\_model0\.0340\.0170\.0000\.0210\.493scienceagentbenchmodel\_agent0\.0000\.0030\.0080\.0140\.023scienceagentbenchsigma0\.5770\.2940\.3380\.227NA
Table 21:Proportion of variation explained full leaderboard\. For comparability across methods, Bayesian estimates represent posterior means \(rather than medians found in most of the study\) to preserve proportional relationships\.FacetLMMBayes LMMGLMMBayes GLMMDISCOagent0\.0160\.0390\.0200\.0400\.014benchmark0\.0750\.1080\.2590\.2990\.023benchmark\_agent0\.0280\.0360\.0260\.0310\.030benchmark\_item0\.2520\.2340\.3280\.2970\.117benchmark\_item\_agent0\.0700\.0650\.0520\.0550\.140benchmark\_item\_model0\.0470\.0430\.0040\.0400\.240benchmark\_model0\.0100\.0080\.0200\.0150\.040benchmark\_model\_agent0\.0130\.0140\.0140\.0160\.046model0\.0260\.0250\.0330\.0320\.015model\_agent0\.0070\.0060\.0090\.0080\.035benchmark\_task\_model\_agent \+ sigma0\.4560\.4220\.2360\.1680\.302Similar Articles
@dair_ai: There is much less signal in agent leaderboards than the rankings imply. A four-facet Generalizability Theory decomposi…
Researchers show agent leaderboards rank task specialization rather than capability, with agent main effects accounting for under 3% of variance across three benchmarks, and propose a Deployment Decision Reliability framework.
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
This paper applies Generalizability Theory to agent benchmarks, showing leaderboards rank specialization rather than capability, and proposes a framework (DDR) for sizing reliable deployment evaluations.
Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions
This empirical survey extends prior work on the bias-reliability tradeoff in LLM evaluation by measuring evaluator coupling, strategy diversity, and small-sample reliability across 11 conditions, confirming that low evaluator influence leads to high measurement noise while strong coupling reduces diversity and noise.
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
This paper argues that aggregate-score leaderboards for LLM agent benchmarks fail to capture deployment-relevant dimensions and show rank instability. It proposes ranking configurations by predictive validity—the correlation between in-sample and out-of-sample rank—and introduces a twelve-tier measurement apparatus along with falsifiable out-of-distribution criteria.
How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
This paper analyzes how many tasks are needed in partial evaluations of LLM agent benchmarks to reach the same pairwise conclusions as full benchmarks. It finds that required task fractions vary sharply across benchmarks and suggests reporting standards for partial evaluations.