Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

arXiv cs.AI Papers

Summary

This paper investigates the measurement limitations of LLM-judged benchmarks, showing that generalizability is constrained by judge variability and that protocol design significantly affects performance ceilings, with implications for AI benchmarking methodologies.

arXiv:2609.27787v1 Announce Type: new Abstract: Benchmarks scored by an LLM judge routinely adjudicate differences of a tenth of a point, but the resolution of those benchmarks has never been measured. Existing sample-complexity work covers accuracy benchmarks and leaves the judged case open. Treating the system as the object of measurement, we decompose 373,019 judgments into system, item, judge and interaction components using generalizability theory. The central result is structural: under a single judge, generalizability asymptotes to sigma2_s/(sigma2_s+sigma2_sj) regardless of item count, because the system-by-judge term carries no n_i. Items saturate; judges do not. The item cost of a target diverges as the target nears that ceiling. The ceiling is a property of pointwise rubric scoring, not of LLM judging. Run as a pairwise preference in both presentation orders, sigma2_sj falls two orders of magnitude below sigma2_s and the ceiling rises to 0.986 (bootstrap [0.934, 1.000] on 11 systems), so one judge suffices. Pairwise buys a different problem: a system presented first wins 8.6 percentage points more often than the same system presented second, a bias 1.23x the median improvement claimed in the 53 published win-rate comparisons we recovered. Protocol design dominates panel size. Measured floors are 0.41-1.24 points on a 0-5 scale at native item counts, against a median reported improvement of 0.28 points; on the one benchmark recurring often enough for an exactly matched comparison, all 17 recovered MT-Bench improvements fall below MT-Bench's own floor, and 70% of the win-rate claims fall below the pairwise floor. An audit of 628 arXiv papers, double-coded by two independent models and validated against blind human coding (kappa=0.73), finds fewer than one paper in four states whether its evaluation was run more than once, and only 46-67% report uncertainty of any kind.
Original Article
View Cached Full Text

Cached at: 09/24/26, 09:30 AM

# Ask Which, Not How Good:Sizing Benchmarks Scored by an LLM
Source: [https://arxiv.org/html/2609.27787](https://arxiv.org/html/2609.27787)
###### Abstract

Benchmarks scored by an LLM judge are used to adjudicate differences of a tenth of a point, but the resolution of those benchmarks has never been measured\. Existing sample\-complexity work covers*accuracy*benchmarks and leaves the judged case open\. We close it\. Treating the*system*as the object of measurement, we decompose 350,000 judgments into system, item, judge and interaction components, holding a 9\-judge panel, one pointwise 0–5 rubric and one scale fixed while varying the items, drawn from MT\-Bench, Arena\-Hard, AlpacaEval 2 and a discrimination\-screened set\. System panels are close but not identical across item sets, in two ways we record\.

Three results follow\. First, a structural one: under a single judge, generalizability asymptotes toσs2/\(σs2\+σs​j2\)\\sigma^\{2\}\_\{s\}/\(\\sigma^\{2\}\_\{s\}\+\\sigma^\{2\}\_\{sj\}\)*regardless of item count*, because the system\-by\-judge term carries nonin\_\{i\}\. Items saturate; judges do not, and the item cost of a target*diverges*as the target nears the ceiling: on MT\-Bench under its own protocol,two judges at thirty items match 77 items at one judgeatE​ρ2=0\.75E\\rho^\{2\}=0\.75, while at0\.800\.80—within0\.0060\.006of that benchmark’s ceiling—no item count suffices at all\. Second,*protocol design moves the ceiling itself*: on identical prompts, MT\-Bench’s native template, 1–10 scale and reference answers double the share of variance attributable to the system \(4\.8% to 8\.4% on matched material\) and cut the judges required from 18 to 3, at no cost per call\. Third, minimum detectable differences are 0\.41–1\.24 points on a 0–5 scale at native item counts, while the median of 60 improvements recovered from published papers is0\.28 points; on the one benchmark recurring often enough for an exactly matched comparison,all 17 recovered MT\-Bench improvements fall below MT\-Bench’s own floor\.

The*existence*of the ceiling is invariant; its*height*is not\. We take the six current\-generation judges as the primary universe, since the other three were included to probe the judge facet rather than because anyone would hire them, and report the nine\-judge panel as sensitivity\. That concession costs the alarm and not the argument: under the primary universe three of four item sets clearE​ρ2=0\.8E\\rho^\{2\}=0\.8and MT\-Bench’s ceiling rises from0\.6000\.600to0\.7980\.798, yet a single judge still cannot reach0\.80\.8on it at any item count, and still needs five judges to do so\.

*Protocol matters more than panel size, and we ran the protocol most of the field uses\.*Under a pairwise preference—both presentation orders, position as an estimable facet—σs​j2\\sigma^\{2\}\_\{sj\}falls two orders of magnitude belowσs2\\sigma^\{2\}\_\{s\}and the ceiling rises to0\.986\\mathbf\{0\.986\}\(bootstrap\[0\.934,1\.000\]\[0\.934,1\.000\]on 11 systems;0\.9540\.954–0\.9740\.974restricted to current models\), so one judge suffices\. The ceiling result is therefore a property of*pointwise rubric scoring*, not of LLM judging as such\. Pairwise buys its own problem instead: a system presented first wins8\.6 percentage pointsmore often than the same system presented second, which is1\.23×1\.23\\timesthe median improvement claimed in the 53 published win\-rate comparisons we recovered, and70%70\\%of those fall below the pairwise floor at native item counts\.

An audit of 628 arXiv papers—double\-coded by two independent models and validated against blind human coding \(κ=0\.73\\kappa=0\.73\)—finds that fewer than one paper in four states whether its evaluation was run more than once \(16\.7% human, 20\.0% automated,κ=0\.89\\kappa=0\.89\), and that only 46–67% report uncertainty of any kind\. The crossed judgment dataset, a D\-study calculator and the audit codebook are available from the author on request\.

## 1Introduction

A large fraction of reported progress in language modelling is adjudicated by an LLM judge\. A method is declared better because it scores 7\.4 rather than 7\.2 on MT\-Bench, or wins 34\.5% rather than 31\.8% on AlpacaEval\. These are measurements, and like all measurements they have a resolution, but that resolution is almost never reported and, for judged benchmarks, has not been characterised\.

For*accuracy*benchmarks the question is settled in outline\.[13](https://arxiv.org/html/2609.27787#bib.bib16)sets out the statistical case for reporting uncertainty on benchmark results and supplies the item\-sampling machinery; a subsequent industry analysis\([18](https://arxiv.org/html/2609.27787#bib.bib1)\)applies it to compute minimum detectable effects on MMLU, HumanEval and GSM8K and reports that many published gaps fall inside their own benchmark’s noise\. Both treat evaluations whose only noise sources are item sampling and generation, and both identify the judged case—where a judge facet with its own behaviour is added—as open\. This paper is that formula, and the measurement that goes with it\. The framing follows[15](https://arxiv.org/html/2609.27787#bib.bib15)in treating a believed empirical claim as a measurement artefact\.

The judge facet is not merely one more variance source\. Writing the measurement model out \(§[3](https://arxiv.org/html/2609.27787#S3)\) shows that the system\-by\-judge interaction enters the error variance asσs​j2/nj\\sigma^\{2\}\_\{sj\}/n\_\{j\}, a term containing nonin\_\{i\}\. Items therefore cannot reduce it\. This has an immediate and, to our knowledge, unremarked consequence: a single\-judge protocol has a reliability ceiling that no amount of data collection can raise\. Whether that ceiling sits above or below a usable threshold is an empirical question, and we answer it for the benchmarks the field actually publishes on\.

#### Contributions\.

1. 1\.A sample\-complexity result for judged benchmarks\.Under one judge,E​ρ2→σs2/\(σs2\+σs​j2\)E\\rho^\{2\}\\to\\sigma^\{2\}\_\{s\}/\(\\sigma^\{2\}\_\{s\}\+\\sigma^\{2\}\_\{sj\}\)asni→∞n\_\{i\}\\to\\infty, so items saturate and judges do not\. The exchange rate steepens as the target nears the ceiling: two judges at thirty items match 77 items at one judge atE​ρ2=0\.75E\\rho^\{2\}=0\.75, and no item count suffices at0\.800\.80\. The saturation is universe\-invariant; whether a given ceiling clears a usable threshold is not; we take the six current\-generation judges as the primary universe and report the wider panel as sensitivity \(§[5\.2](https://arxiv.org/html/2609.27787#S5.SS2), §[5\.5](https://arxiv.org/html/2609.27787#S5.SS5)\)\.
2. 2\.The ceiling is a property of pointwise rubric scoring, not of LLM judging\.Run as a pairwise preference in both presentation orders,σs​j2\\sigma^\{2\}\_\{sj\}falls two orders of magnitude belowσs2\\sigma^\{2\}\_\{s\}, the ceiling rises to 0\.986 \(\[0\.934,1\.000\]\[0\.934,1\.000\], 11 systems\), and one judge suffices, but a system presented first wins 8\.6 percentage points more often than the same system presented second, a bias1\.23×1\.23\\timesthe median published win\-rate claim\. With that floor measured,70%70\\%of the 53 win\-rate comparisons we recovered fall below it \(§[5\.4](https://arxiv.org/html/2609.27787#S5.SS4)\)\.
3. 3\.Protocol design moves the ceiling\.On identical prompts, MT\-Bench’s native protocol raisesσs2\\sigma^\{2\}\_\{s\}share from 4\.8% to 8\.4% against a generic holistic rubric and cuts the judges needed forE​ρ2=0\.8E\\rho^\{2\}=0\.8from 18 to 3, on matched material\. Rubric and scale are a cheaper lever than either items or judges, and are almost never treated as a design variable \(§[5\.3](https://arxiv.org/html/2609.27787#S5.SS3)\)\.
4. 4\.Measured noise floorsof 0\.41–1\.24 points on a 0–5 scale at native item counts, with a common judge panel, rubric and scale—system panels differ only as §[8](https://arxiv.org/html/2609.27787#S8)records—so cross\-set differences are attributable to the items; and a measured distribution of reported improvements whose median \(0\.28 points\) sits below every floor we measured—all 17 recovered MT\-Bench improvements fall below MT\-Bench’s own \(§[5\.1](https://arxiv.org/html/2609.27787#S5.SS1), §[5\.6](https://arxiv.org/html/2609.27787#S5.SS6)\)\.
5. 5\.Reliability is a property of the \(benchmark, judge\-universe\) pair\.MT\-Bench’s ceiling moves between 0\.600 and 0\.798 depending on which judges are admitted; no published work states its judge universe \(§[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px1)\)\.
6. 6\.The limit is a property of the design, not of LLM judges\.On a matched design LLM judges are*more*self\-consistent than human experts and far more so than crowd workers \(§[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px10)\); and discriminating power collapses when the systems compared are close, system variance falling to 1\.9–4\.2% on three of four item sets once restricted to ten current frontier models \(§[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px4)\)\.
7. 7\.Item\-discrimination screening is the one lever that raisesσs2\\sigma^\{2\}\_\{s\}\.Screening candidate items by between\-system variance movedσs2\\sigma^\{2\}\_\{s\}by roughly two orders of magnitude in a pilot, and yields a system\-variance share comparable to the best public set \(16\.6% against Arena\-Hard’s 19\.2%, on items chosen for that property rather than found to have it\)\. Since items saturate, this is the only axis that raises the ceiling rather than approaching it \(§[4](https://arxiv.org/html/2609.27787#S4)\)\.
8. 8\.Judge provenance is a further facet\.Replicating the design with five open\-weight judges on identical material, judges disagree2\.17×2\.17\\timesmore across model families than within them once scale usage is standardised away, so panel size overstates panel independence, though mixing families does not reduce the judges required \(§[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px3)\)\.
9. 9\.An audit of 628 papers, double\-coded by two independent models, showing that the disclosure needed to assess any of this is largely absent, and quantifying how reliably each disclosure can be coded at all \(§[6](https://arxiv.org/html/2609.27787#S6)\)\.

## 2Related work

#### Sizing and uncertainty for benchmarks\.

[13](https://arxiv.org/html/2609.27787#bib.bib16)makes the case for reporting uncertainty on evaluation results and supplies the item\-sampling machinery;[18](https://arxiv.org/html/2609.27787#bib.bib1)applies it to compute minimum detectable effects on accuracy benchmarks and finds many published gaps inside their own noise\. Both treat evaluations whose only stochastic facets are item sampling and generation, and both identify the judged case as open\. Within NLP,[3](https://arxiv.org/html/2609.27787#bib.bib17)showed that common experimental designs are underpowered for the effects they claim,[7](https://arxiv.org/html/2609.27787#bib.bib18)set out significance\-testing practice for the field, and[6](https://arxiv.org/html/2609.27787#bib.bib19)gave confidence intervals for summarisation metrics that account for both systems and items\. Our contribution is to add the judge facet, which is what changes the asymptotics: item sampling can be beaten by collecting more items and the judge facet cannot\.[15](https://arxiv.org/html/2609.27787#bib.bib15)is the closest in spirit, treating a believed empirical claim as an artifact of the measurement rather than a property of the model\.

#### Generalizability theory\.

The variance\-decomposition machinery is standard in educational measurement\([5](https://arxiv.org/html/2609.27787#bib.bib10);[2](https://arxiv.org/html/2609.27787#bib.bib9);[16](https://arxiv.org/html/2609.27787#bib.bib20)\), where the object of measurement is normally a person and raters are the nuisance facet\. Recent applications to LLM evaluation\([12](https://arxiv.org/html/2609.27787#bib.bib4);[17](https://arxiv.org/html/2609.27787#bib.bib6)\)keep the item as the object and therefore answer an annotation\-quality question\. Taking the*system*as the object is what makesσs​j2/nj\\sigma^\{2\}\_\{sj\}/n\_\{j\}the binding term, since it is the interaction between the thing being measured and the instrument measuring it\.

#### Judge reliability and bias\.

A substantial literature documents that LLM judges are noisy and biased evaluators:[19](https://arxiv.org/html/2609.27787#bib.bib3)and[14](https://arxiv.org/html/2609.27787#bib.bib5)characterise agreement, consistency and bias across judge models;[20](https://arxiv.org/html/2609.27787#bib.bib2)shows that conclusions move when the judge changes;[10](https://arxiv.org/html/2609.27787#bib.bib8)attributes much of the problem to design failures in judge benchmarks; and[4](https://arxiv.org/html/2609.27787#bib.bib7)applies item response theory to diagnose judge quality\. This work establishes*that*judges disagree\. What it does not derive is the consequence for sample size, which is the gap this paper fills: disagreement between judges enters the error variance in a term that no amount of item collection reduces\.

#### Benchmarks and leaderboards\.

We measure items drawn from MT\-Bench\([21](https://arxiv.org/html/2609.27787#bib.bib11)\), Arena\-Hard\([11](https://arxiv.org/html/2609.27787#bib.bib13)\)and AlpacaEval 2\([8](https://arxiv.org/html/2609.27787#bib.bib12)\)\. Work on leaderboard stability under resampling and on the statistical properties of Elo\-style aggregation is complementary to ours: it asks how a ranking moves under perturbation, where we ask how finely the underlying scores can be resolved at all\.[1](https://arxiv.org/html/2609.27787#bib.bib21)argues that benchmark progress claims outrun the evidence supporting them, which is the concern our instruments are intended to make checkable\.

## 3The measurement model

We treat the*system*as the object of measurement, in the generalizability\-theory sense of[5](https://arxiv.org/html/2609.27787#bib.bib10)and[2](https://arxiv.org/html/2609.27787#bib.bib9)\. This differs from prior applications of G\-theory to LLM evaluation\([12](https://arxiv.org/html/2609.27787#bib.bib4);[17](https://arxiv.org/html/2609.27787#bib.bib6)\), which take the item as the object and therefore answer an annotation\-quality question rather than a benchmark\-validity one\. Related measurement work studies judges through item response theory\([4](https://arxiv.org/html/2609.27787#bib.bib7)\), replicate\-aggregation curves\([19](https://arxiv.org/html/2609.27787#bib.bib3)\), judge substitution\([20](https://arxiv.org/html/2609.27787#bib.bib2)\), and design failures in judge benchmarks\([10](https://arxiv.org/html/2609.27787#bib.bib8);[14](https://arxiv.org/html/2609.27787#bib.bib5)\); none derives or measures the item\-count ceiling below\. A score is modelled as

Xs​i​j​r=μ\+νs\+νi\+νj\+νs​i\+νs​j\+νi​j\+νs​i​j\+νr⁡\(s​i​j\)\.X\_\{sijr\}=\\mu\+\\nu\_\{s\}\+\\nu\_\{i\}\+\\nu\_\{j\}\+\\nu\_\{si\}\+\\nu\_\{sj\}\+\\nu\_\{ij\}\+\\nu\_\{sij\}\+\\nu\_\{r\(sij\)\}\.\(1\)Withr=1r=1,σs​i​j2\\sigma^\{2\}\_\{sij\}and the residual are not separately identified:*a single\-run design cannot compute its own noise floor*\. We therefore user=3r=3\.

Relative error variance, generalizability, and the minimum detectable difference between two systems measured on the same items and judges are

σδ2\\displaystyle\\sigma^\{2\}\_\{\\delta\}=σs​i2ni\+σs​j2nj\+σs​i​j2ni​nj\+σe2ni​nj​nr,E​ρ2=σs2σs2\+σδ2,\\displaystyle=\\frac\{\\sigma^\{2\}\_\{si\}\}\{n\_\{i\}\}\+\\frac\{\\sigma^\{2\}\_\{sj\}\}\{n\_\{j\}\}\+\\frac\{\\sigma^\{2\}\_\{sij\}\}\{n\_\{i\}n\_\{j\}\}\+\\frac\{\\sigma^\{2\}\_\{e\}\}\{n\_\{i\}n\_\{j\}n\_\{r\}\},\\qquad E\\rho^\{2\}=\\frac\{\\sigma^\{2\}\_\{s\}\}\{\\sigma^\{2\}\_\{s\}\+\\sigma^\{2\}\_\{\\delta\}\},\(2\)MDD95\\displaystyle\\mathrm\{MDD\}\_\{95\}≈1\.96​2​σδ2\.\\displaystyle\\approx 1\.96\\sqrt\{2\\,\\sigma^\{2\}\_\{\\delta\}\}\.\(3\)
#### The ceiling\.

Settingnj=1n\_\{j\}=1and lettingni→∞n\_\{i\}\\to\\infty, every term ofσδ2\\sigma^\{2\}\_\{\\delta\}vanishes exceptσs​j2\\sigma^\{2\}\_\{sj\}, giving

limni→∞E​ρ2\|nj=1=σs2σs2\+σs​j2,limni→∞MDD95\|nj=1=1\.96​2​σs​j2\.\\lim\_\{n\_\{i\}\\to\\infty\}E\\rho^\{2\}\\big\|\_\{n\_\{j\}=1\}=\\frac\{\\sigma^\{2\}\_\{s\}\}\{\\sigma^\{2\}\_\{s\}\+\\sigma^\{2\}\_\{sj\}\},\\qquad\\lim\_\{n\_\{i\}\\to\\infty\}\\mathrm\{MDD\}\_\{95\}\\big\|\_\{n\_\{j\}=1\}=1\.96\\sqrt\{2\\,\\sigma^\{2\}\_\{sj\}\}\.\(4\)Items buy nothing againstσs​j2\\sigma^\{2\}\_\{sj\}; only judges do\. Equation \([4](https://arxiv.org/html/2609.27787#S3.E4)\) follows directly from standard D\-study algebra\([2](https://arxiv.org/html/2609.27787#bib.bib9)\); our claim is not the derivation but that its consequence has never been measured for judged benchmarks, nor drawn\. “Run more items” is the field’s reflexive response to evaluation noise, and against this component it is futile\.

#### Why MDD rather thanE​ρ2E\\rho^\{2\}\.

σδ2\\sigma^\{2\}\_\{\\delta\}is built from components estimated on tens of thousands of degrees of freedom, whereasσs2\\sigma^\{2\}\_\{s\}has onlyns−1n\_\{s\}\-1\. MDD contains noσs2\\sigma^\{2\}\_\{s\};E​ρ2E\\rho^\{2\}divides by it\. In our data the bootstrap interval on MDD is roughly±7%\\pm 7\\%of its point estimate while that onσs2\\sigma^\{2\}\_\{s\}exceeds100%100\\%\. We therefore report MDD as the primary quantity andE​ρ2E\\rho^\{2\}as an interval\.

## 4Experimental setup

#### Design\.

42 systems \(14 models×\\times3 prompt configurations, spanning 2024\-era to current frontier\),nin\_\{i\}items per benchmark, 9 judges, 3 replicates, pointwise scoring on a 0–5 scale\. The same judges, rubric and scale are used across all four item sets\. The system panels are close but not identical, and §[8](https://arxiv.org/html/2609.27787#S8)records the two ways they differ, so cross\-set differences are attributable to the items up to that caveat rather than unconditionally\.

Two system counts appear in the tables and the difference is deliberate\. The three public item sets are scored on the 42 systems above\. The screened set additionally carries a*null arm*: for twelve of the systems we drew two further independent generation samples under an identical configuration, giving pairs whose true difference is exactly zero, which is what makes the Type I error check in §[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px11)possible\. Those replicas are systems for the purposes of the decomposition, so the screened set has36\+23=5936\+23=59after one replica is dropped for incomplete coverage\. Its base panel is 36 rather than 42 becauseclaude\-opus\-5andclaude\-sonnet\-5were not available as systems when the screened set was collected, in any of the three prompt configurations; both remain in the judge panel\. Refitting the screened row on those 36 base systems alone changes it little \(MDD 1\.25 against 1\.24, ceiling 0\.735 against 0\.746\), so the cross\-set comparison in Table[1](https://arxiv.org/html/2609.27787#S5.T1)does not turn on the null arm\. The near\-frontier subset \(Appendix[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px4)\) is the ten current frontier models at three configurations, plus their null\-arm replicas where present, giving 38 on the screened set and 30 elsewhere\.

#### Judges\.

A crossed 2 families×\\times3 tiers panel of current models, plus three further judges: a deliberately weak judge, a legacy judge representing what published work typically used, and a*route replica*, the same weights served through a different provider route\.

All three are*included*in the nine\-judge universe that Table[1](https://arxiv.org/html/2609.27787#S5.T1)reports\. We report that universe as the headline because published work states no universe at all and a reader assembling “some LLM judges” could plausibly land on it\. The choice is not neutral—a universe containing a deliberately weak judge carries more system\-by\-judge disagreement, hence a lower ceiling and a larger judge requirement, which is the direction that flatters this paper’s thesis—and it turns out to be load\-bearing rather than cosmetic\. We therefore take the*six current\-generation judges*as the primary universe for interpretation, since it is what a practitioner assembling a panel today would build, and report the nine\-judge fit in the tables because it is what a reader reconstructing our full design would compute\. Under the primary universe three of the four item sets clearE​ρ2=0\.8E\\rho^\{2\}=0\.8; under the nine\-judge universe none does\. The structural result is identical either way\. §[5\.2](https://arxiv.org/html/2609.27787#S5.SS2)quantifies the gap and §[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px1)gives every headline quantity under four universes\.

#### Benchmarks\.

MT\-Bench\([21](https://arxiv.org/html/2609.27787#bib.bib11)\), Arena\-Hard\([11](https://arxiv.org/html/2609.27787#bib.bib13)\)and AlpacaEval 2\([8](https://arxiv.org/html/2609.27787#bib.bib12)\)items are random samples under a fixed seed\. A fourth set is constructed by*item discrimination screening*: candidates are scored by a held\-out judge and retained by between\-system variance, which is each item’s contribution toσs2\\sigma^\{2\}\_\{s\}\. In a 15\-system pilot this movedσs2\\sigma^\{2\}\_\{s\}from0\.00270\.0027to0\.69580\.6958, roughly two orders of magnitude; the pilot is small and its two arms differ in item count as well as in screening, so we take the direction and the order of magnitude rather than a point estimate\. What the screened set buys in the main study, and what it costs, is reported with the other item sets in §[5\.1](https://arxiv.org/html/2609.27787#S5.SS1)\.

The motivation is that classical item analysis has screened items for discrimination for a century while benchmark construction generally does not, selecting instead for difficulty, diversity or realism\. A benchmark can therefore carry hundreds of items and almost no system signal, and because items saturate \(§[5\.2](https://arxiv.org/html/2609.27787#S5.SS2)\), adding more cannot repair it\. Screening is the only lever we tested that raisesσs2\\sigma^\{2\}\_\{s\}itself, which is the numerator every other quantity in this paper depends on\.

#### Infrastructure\.

Three failure modes materially affected measurement and are reported because they generalise\. \(i\) The inference gateway performed*response caching*: 99\.8% of replicate cells initially returned a byte\-identical cached response, collapsing the replicate facet and makingσe2\\sigma^\{2\}\_\{e\}read as zero\. Any judge\-reliability study run through a caching gateway will conclude judges are perfectly self\-consistent\. \(ii\) Reasoning\-model judges silently returned empty completions when the token budget was sized for a score, removing a judge from the panel without removing it from the reported design\. \(iii\) Prompt caching did not function, so replicate cost is not discounted\. All calls log model version, timestamp, prompt hash and response identifier; the response identifier is what made \(i\) detectable\.

## 5Results

### 5\.1Noise floors

Table[1](https://arxiv.org/html/2609.27787#S5.T1)reports the decomposition\. Minimum detectable differences at a common thirty items, one judge and one run range from 0\.58 to 1\.24 points on a 0–5 scale, and system variance is 4\.1–4\.5% of total score variance on MT\-Bench and AlpacaEval 2: the overwhelming majority of what a judge score varies with is not the system\.

Table[1](https://arxiv.org/html/2609.27787#S5.T1)is fitted on all nine judges, which is the*sensitivity*universe, not the primary one: three of the nine were included to probe the judge facet rather than because anyone would hire them\. Section[5\.5](https://arxiv.org/html/2609.27787#S5.SS5)sets out the primary universe of six current\-generation judges and reports every ceiling under both, and readers who want the headline numbers for a panel they would actually assemble should read Table[4](https://arxiv.org/html/2609.27787#S5.T4)first\. We present the nine\-judge fit here because the rest of this section varies the item facet against a fixed judge set, and the wider set makes that comparison harder rather than easier on us\.

Table 1:Detectability and the single\-judge ceiling\. MDD andE​ρ2E\\rho^\{2\}at a commonni=30n\_\{i\}=30, one judge, one run, so benchmarks are compared at an equal item budget\. Intervals are parametric bootstrap \(600–800 resamples\);P\(≥0\.8\)P\(\\geq\\\!0\.8\)is the bootstrap probability that the ceiling reaches the conventional threshold\. The final column is the number of judges needed to reachE​ρ2=0\.8E\\rho^\{2\}=0\.8at 30 items and 3 replicates\.
### 5\.2Items cannot lift the ceiling

Figure[1](https://arxiv.org/html/2609.27787#S5.F1)plotsE​ρ2E\\rho^\{2\}against item count under a single judge\. Every curve flattens below the 0\.8 threshold\. The asymptotes are 0\.600 \(MT\-Bench\), 0\.706 \(AlpacaEval 2\), 0\.746 \(our screened items\) and 0\.792 \(Arena\-Hard\)\. Parametric bootstrap intervals \(800 resamples\) put the posterior probability that the ceiling reaches 0\.8 at 0\.0% for MT\-Bench, 6\.2% for our screened set and 2\.0% for AlpacaEval 2\.On these three, a single\-judge protocol does not reach conventional reliability at any item count under this judge universe\.Arena\-Hard is the honest exception: at 0\.792 with a 95% interval of\[0\.685,0\.858\]\[0\.685,\\,0\.858\]it has a 41% probability of clearing the threshold, and we do not claim it fails\. The corresponding floors on MDD at infinite items are 0\.40–1\.09 points\.

#### The exchange rate between the two axes\.

ReachingE​ρ2=0\.8E\\rho^\{2\}=0\.8at 30 items requires 25 judges for MT\-Bench, 7 for AlpacaEval 2, and 2 each for our screened set and Arena\-Hard \(Table[1](https://arxiv.org/html/2609.27787#S5.T1)\)\. Adding a*second*judge at 30 items exceeds what a single judge achieves at*any*item count on three of the four benchmarks, because the target sits above the single\-judge ceiling and no amount of item collection reaches it\. On the fourth, AlpacaEval 2, a single judge needs 153 items to match two judges at 30, a target far enough below that benchmark’s ceiling \(0\.6660\.666against0\.7060\.706\) for the count to be stable, unlike the near\-ceiling case in §[5\.3](https://arxiv.org/html/2609.27787#S5.SS3)\. A second judge is not a marginal improvement over more items; on most benchmarks it is the only thing that works\.

These asymptotes are computed over the nine\-judge universe; over the six current\-generation judges the same benchmarks give0\.9290\.929,0\.9180\.918,0\.8290\.829and0\.7980\.798, so three of four clear the threshold and only MT\-Bench does not \(§[5\.5](https://arxiv.org/html/2609.27787#S5.SS5)\)\. The saturation itself is universe\-invariant—no ceiling is reachable by adding items, for any universe, becauseσs​j2/nj\\sigma^\{2\}\_\{sj\}/n\_\{j\}carries nonin\_\{i\}—but whether a given ceiling is high enough to be usable is a joint property of the benchmark and the judges admitted\. We regard that sensitivity as a finding rather than a caveat: a reliability number reported without a stated judge universe is not conservative or liberal, it is undetermined\.

Figure 1:Generalizability against item count under a single judge\. Dashed lines are the asymptotes of Eq\. \([4](https://arxiv.org/html/2609.27787#S3.E4)\); the solid horizontal line is the conventionalE​ρ2=0\.8E\\rho^\{2\}=0\.8threshold\. All four flatten below it*under the nine\-judge universe*; under the six current\-generation judges three of the four asymptote above it \(§[5\.5](https://arxiv.org/html/2609.27787#S5.SS5)\)\.

### 5\.3Protocol design moves the ceiling

Everything above holds the judging protocol fixed—our generic holistic 0–5 rubric—so that differences are attributable to items\. That isolates the item effect but leaves a question a reader will reasonably ask: are these properties of the benchmark, or of our rubric? We answer it for MT\-Bench, the one pointwise benchmark whose protocol we can reproduce exactly, by re\-running the same prompts under FastChat’s ownsingle\-v1template, its 1–10 scale, and the GPT\-4 reference answers it supplies for the math, reasoning and coding categories\.

The two arms did not survive collection equally, and the difference runs against us\. They drew from different item pools—43 candidates for the generic arm, 50 for the native one—and under the native protocol’s longer prompts theoai\-cheapjudge returned empty completions often enough to be dropped entirely, taking 12 of its 50 items with it\. As collected that leaves38×38×838\\times 38\\times 8against42×42×942\\times 42\\times 9\. The arm favouring our conclusion is therefore also the arm with the smaller panel, which is exactly the shape of confound a reader should distrust\. We remove it by refitting both arms on the 38 systems, 33 items and 8 judges common to the two \(Table[2](https://arxiv.org/html/2609.27787#S5.T2), lower block\) and report that as the result\.

Matched, MT\-Bench under its native protocol recovers8\.4% of variance as system difference against 4\.8% under our rubric—close to double—and its single\-judge ceiling rises from 0\.623 to 0\.798\. The judges required forE​ρ2=0\.8E\\rho^\{2\}=0\.8at thirty items fall from 18 to 3\. Matching moves the generic arm slightly in its own favour \(4\.5% to 4\.8%, ceiling 0\.600 to 0\.623\) and the native arm slightly against it \(9\.1% to 8\.4%, 0\.804 to 0\.798\), so the lever is not an artefact of the dropped judge\.

#### The item exchange rate diverges near the ceiling\.

That last shift is small but it crosses a threshold, and the crossing is instructive\. As collected, the native arm’s ceiling is 0\.804 and a single judge reachesE​ρ2=0\.8E\\rho^\{2\}=0\.8at 860 items; matched, the ceiling is 0\.798 and*no*item count reaches 0\.8\. Both are true of the same protocol\. The reason is thatE​ρ2E\\rho^\{2\}approaches its asymptote hyperbolically, so the items required to hit a target diverge as the target approaches the ceiling, and 0\.8 sits within 0\.006 of this one\. Away from the asymptote the exchange rate is stable and far less dramatic: atE​ρ2=0\.75E\\rho^\{2\}=0\.75the matched native arm needs 2 judges at thirty items or 77 items at one judge, and at 0\.78, 3 judges or 209 items\. We therefore report the exchange rate as a property that*steepens*rather than as a single ratio, and treat any specific item count quoted near a ceiling—including our own 860—as an artefact of proximity to the asymptote rather than a stable quantity\.

Three consequences\. First, the cross\-benchmark comparison in §[5\.1](https://arxiv.org/html/2609.27787#S5.SS1)is*conservative*: holding a generic rubric fixed understates what each benchmark achieves under its own design, so the true floors are likely lower than Table[1](https://arxiv.org/html/2609.27787#S5.T1)reports\. Second, rubric and scale are a design lever that costs nothing per call—reference answers and a wider response scale bought more discriminating power here than tripling the judge panel would have—and protocol design is almost never reported as a choice, let alone justified\. Third, it revises the strongest form of our headline\. Under a generic rubric no item count reachesE​ρ2=0\.8E\\rho^\{2\}=0\.8on any item set studied; under the native protocol the answer depends on which fit you read, which is precisely why we state the surviving claim in its weaker form: a protocol change buys more than any feasible item budget, and the budget required explodes near whichever ceiling the protocol produces\.

Table 2:Protocol design on identical MT\-Bench prompts: our generic holistic 0–5 rubric versus MT\-Bench’s own FastChat template, 1–10 scale, and GPT\-4 reference answers for math, reasoning and coding\. The two arms drew from different item pools \(43 and 50 candidates\) and did not survive collection equally—under the native protocol’s longer prompts theoai\-cheapjudge returned empty completions often enough to be dropped, along with 12 of its 50 items—so the upper block reports each arm as collected and the lower block refits both on the 38 systems, 33 items and 8 judges common to the two, which is the contrast the text relies on\. MDD is on a common 0–5 basis\. Judges is the count needed forE​ρ2=0\.8E\\rho^\{2\}=0\.8at 30 items and 3 replicates; items the count at one judge and a single run, which diverges as the target approaches the ceiling\.Protocoldesignσs2\\sigma^\{2\}\_\{s\}E​ρ2E\\rho^\{2\}@30ceilingMDD0​\-​5\{\}\_\{0\\text\{\-\}5\}judgesitems*as collected*Generic 0–5 rubric42×\\times42×\\times94\.5%0\.4630\.6000\.6325neverMT\-Bench native38×\\times38×\\times89\.1%0\.7000\.8040\.563860*matched: common 38 systems, 33 items, 8 judges*Generic 0–5 rubric38×\\times33×\\times84\.8%0\.4830\.6230\.6118neverMT\-Bench native38×\\times33×\\times88\.4%0\.6850\.7980\.563never

### 5\.4The pairwise protocol, and where the ceiling goes

Everything to this point scores one response at a time against a rubric\. Most of the field does not: Arena\-Hard and AlpacaEval 2 are pairwise, Chatbot Arena is pairwise, and the 53 win\-rate comparisons we recovered from published papers are pairwise\. A result about pointwise rubrics that is silent on pairwise judging describes a minority of practice\. So we ran it\.

#### Design\.

Each of 12 systems is compared against a fixed baseline \(gpt\-4o\-mini\) on the same 24 items by the same 4 judges, twice, with the judge returning a preference rather than a score\. The whole grid is run in*both*presentation orders—system first and baseline first—so position is an estimable facet rather than a nuisance averaged away\. Responses are reused from the main harvest, so this arm costs judging only\. The crossed decomposition is unchanged and the ceiling derivation carries over verbatim, becauseσs​j2/nj\\sigma^\{2\}\_\{sj\}/n\_\{j\}contains nonin\_\{i\}whatever the judge returns\.

#### The ceiling largely disappears\.

Order\-balanced, the system term is35\.7%35\.7\\%of total variance and the single\-judge ceiling is0\.986\\mathbf\{0\.986\}; one judge reachesE​ρ2=0\.93E\\rho^\{2\}=0\.93at 24 items\. This is not a saturation artefact—no system’s win rate is pinned at 0 or 1, the observed range being0\.100\.10to0\.870\.87—but the 12 systems do span gpt\-3\.5\-turbo to gpt\-5\.4\-mini, and Section[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px4)shows that range inflatesσs2\\sigma^\{2\}\_\{s\}\. Dropping the 2023 models halves the system share, to17\.517\.5–19\.8%19\.8\\%, and the ceiling still comes back at0\.9540\.954–0\.9740\.974with one judge reachingE​ρ2≈0\.86E\\rho^\{2\}\\approx 0\.86\. The effect survives the restriction that kills it in the pointwise design\.

This arm is small, and the precision follows\. After reduction to a complete block it carries 11 systems against the main study’s 42, soσs2\\sigma^\{2\}\_\{s\}rests on 10 degrees of freedom rather than 41, and it is the numerator of the ceiling\. The parametric bootstrap gives\[0\.934,1\.000\]\[0\.934,1\.000\]for the balanced ceiling and\[0\.824,0\.983\]\[0\.824,0\.983\]for the tighter frontier\-only fit \(truncated at 1, which the estimator does not enforce\)\. Those are wide intervals and we do not want them read as precise\. What carries the claim is that the*lower*bounds,0\.9340\.934and0\.8240\.824, still sit above every pointwise ceiling we measured \(0\.6000\.600–0\.7980\.798\): the gap between protocols is larger than the uncertainty in the pairwise estimate, even though that uncertainty is substantial\.

We state the consequence plainly, because it cuts against our own headline\. Under a binary preference the judges barely disagree:σs​j2\\sigma^\{2\}\_\{sj\}is two orders of magnitude belowσs2\\sigma^\{2\}\_\{s\}, where under our generic 0–5 rubric it is comparable to it\. Asking “which of these two is better” is a far better conditioned question than asking “what score does this deserve,” and the ceiling result—no item count suffices—is a property of*pointwise rubric scoring*, not of LLM judging as such\. Together with Section[5\.3](https://arxiv.org/html/2609.27787#S5.SS3), where MT\-Bench’s native template moves the ceiling from0\.6230\.623to0\.7980\.798, the honest summary of this paper is thatprotocol design dominates panel size, and that the generic pointwise rubric is the worst case rather than the representative one\.

#### But pairwise buys a new facet, and it is not small\.

Position bias is a systematic shift, not noise\. A system placed in the first slot wins8\.6\\mathbf\{8\.6\}percentage points more often than the same system, judged by the same judges on the same items, placed second; the shift varies across systems with a standard deviation of7\.37\.3points and reaches2525points at its worst, so it is not a constant that cancels\. The median improvement claimed in the 53 published win\-rate comparisons is7\.07\.0points\.An evaluation that does not balance presentation order therefore carries a systematic bias1\.23×1\.23\\timesthe median effect it is trying to detect, which is worth more than the entire reliability gain from adding judges\.

#### The 53 deltas, no longer dropped\.

With a measured pairwise floor we can price the win\-rate comparisons we previously had to set aside\. Order\-balanced, the MDD is19\.119\.1points at our common 30 items and one judge,10\.110\.1at Arena\-Hard’s native 500, and9\.89\.8at AlpacaEval 2’s native 805, falling to7\.57\.5and7\.17\.1at two judges\. Against the floor at native item counts,37 of 53 \(70%70\\%\) published win\-rate improvements fall below itat one judge, and 27 of 53 \(51%51\\%\) at two\. The pairwise protocol is better conditioned than the pointwise one and the claims made on it are still, in the majority, smaller than the noise, not because the ceiling binds, but because a 7\-point win\-rate difference is simply a small effect on a binary outcome measured over a few hundred items\. The caveats of Section[6](https://arxiv.org/html/2609.27787#S6)apply here too: this is our baseline, our items and our judges, and it is a floor transferred to claims made on other instruments\.

Table 3:The pairwise arm: 12 systems against a fixedgpt\-4o\-minibaseline on 24 shared items, 4 judges, 2 replicates, run in both presentation orders\. The response is a win indicator, soσs2\\sigma^\{2\}\_\{s\}and MDD are in win\-rate units and are not comparable to the 0–5 rows elsewhere; ceiling andE​ρ2E\\rho^\{2\}are unitless and are\. “Balanced” is the mean of the two orders, which is the design a careful pairwise evaluation actually runs\. “Frontier\-only” drops the 2023 models, matching the restriction of Appendix[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px4)\. Under a binary preferenceσs​j2\\sigma^\{2\}\_\{sj\}falls two orders of magnitude belowσs2\\sigma^\{2\}\_\{s\}and the ceiling all but vanishes—at the cost of a position facet that a pointwise design does not have\.Armdesignσs2\\sigma^\{2\}\_\{s\}σs​j2\\sigma^\{2\}\_\{sj\}shareceilingE​ρ2E\\rho^\{2\}@24judges@0\.8System first \(AB\)11×\\times20×\\times40\.08740\.0003433\.7%0\.9960\.9371Baseline first \(BA\)12×\\times19×\\times40\.07050\.0007527\.6%0\.9890\.9141Balanced11×\\times19×\\times40\.08010\.0011035\.7%0\.9860\.9341AB, frontier\-only9×\\times20×\\times40\.04320\.0020719\.8%0\.9540\.8581BA, frontier\-only9×\\times19×\\times40\.04420\.0011917\.5%0\.9740\.8641*position bias:*slot\-A advantage8\.68\.6points, sd across systems7\.37\.3, max25\.025\.0*floor \(balanced, 1 judge\):*19\.119\.1pts @30 items,10\.110\.1@500,9\.89\.8@805*53 published win\-rate claims:*median7\.07\.0pts;3737\(70%70\\%\) below the native\-count floor

### 5\.5Which judges count: the primary universe

Every reliability quantity in this paper is conditional on a universe of judges, and we have to say which universe is the headline one\. The nine\-judge panel includes a deliberately weak judge, a legacy model, and a route replica of a judge already in the panel\. Those three were put there to probe the judge facet, not because anyone would hire them, and leaving them in the primary panel would inflateσs​j2\\sigma^\{2\}\_\{sj\}for reasons of our own construction\. A reviewer is entitled to read that as building the conclusion into the design, so we do not ask for the benefit of the doubt:the primary universe is the six current\-generation judges, and the nine\-judge panel is reported throughout as sensitivity \(Table[4](https://arxiv.org/html/2609.27787#S5.T4)\)\.

The concession is real and it costs the alarm, not the argument\. Restricting to the six raises every ceiling—MT\-Bench from0\.6000\.600to0\.7980\.798, our screened set from0\.7460\.746to0\.9290\.929—and cuts every judge requirement, most sharply on MT\-Bench, from 25 judges to 5\. Three of the four item sets now clear the conventional0\.80\.8threshold\. Anyone reading this paper as “LLM judging is hopeless” should read the primary column instead: with a panel of current models and a well\-chosen item set, single\-judge reliability above0\.90\.9is achievable, and we measure it\.

What does not change is the shape of the problem\. The ceiling is still finite and still set byσs​j2\\sigma^\{2\}\_\{sj\}, so the item cost of a target still diverges as the target approaches it: on MT\-Bench, the strongest of the three established item sets under our rubric, a single current\-generation judge still cannot reach0\.80\.8at any item count and still needs five judges to get there at thirty items\. The ordering of item sets is identical under both universes, and the screened set remains the cell where item construction has bought the most\. Every structural claim in Section[5\.2](https://arxiv.org/html/2609.27787#S5.SS2)is an algebraic consequence of the crossed design and holds under any universe whatever; only the numbers move\.

Two further points on the conditioning\. First, the six are not independent draws from some population of judges, they are two vendor families at three capability tiers, and Section[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px3)shows the family term is the larger one, so a panel of six judges from a*single*vendor is a different and worse universe than the one tabulated here\. Second, restricting the universe restricts the claim: results in the primary column describe what a panel of current frontier judges can resolve today, and say nothing about what the same benchmark will resolve when those judges are replaced, which on the evidence of the legacy judge is a live concern rather than a hypothetical one\.

Table 4:Primary judge universe: the six current\-generation judges, with the nine\-judge panel beside it as sensitivity\. The six drop the deliberately weak judge, the legacy model and the route replica, and are the universe a practitioner assembling a panel today would actually draw from\. Quantities are as in Table[1](https://arxiv.org/html/2609.27787#S5.T1): MDD andE​ρ2E\\rho^\{2\}at a commonni=30n\_\{i\}=30, one judge, one run; ceiling isσs2/\(σs2\+σs​j2\)\\sigma^\{2\}\_\{s\}/\(\\sigma^\{2\}\_\{s\}\+\\sigma^\{2\}\_\{sj\}\)with a parametric bootstrap interval; judges is the count needed forE​ρ2=0\.8E\\rho^\{2\}=0\.8at 30 items and 3 replicates\. Restricting to current\-generation judges raises every ceiling and lowers every judge requirement, and three of the four item sets then clear0\.80\.8—but the ordering of item sets, the finite ceiling, and the divergence of the item cost near that ceiling are unchanged\.
### 5\.6What improvements do papers actually claim?

The floor only matters relative to the effects being reported, so we measured those too\. From the audit frame \(§[6](https://arxiv.org/html/2609.27787#S6)\) we extracted each paper’s headline improvement on a judge\-scored benchmark, together with the scale it was reported on, recovering 157 usable comparisons from 227 eligible papers\. Those 157 divide into 60 reported on a pointwise scale we can convert to a common 0–5 basis, 53 reported as win rates, and 44 on scales we could not convert, accuracy on a bespoke subset, composite indices, or rubric totals whose range was not stated\. The 44 are excluded from every quantity below rather than converted on a guess, which is the same discipline we apply to win rates\.

Scale conversion matters and is usually left implicit\. MT\-Bench reports on 1–10, several benchmarks on 0–100, ours on 0–5; a 0\.2\-point MT\-Bench delta is 0\.11 points on a 0–5 basis\. We convert pointwise scales by their range,δ0​\-​5=δ⋅5/\(hi−lo\)\\delta\_\{0\\text\{\-\}5\}=\\delta\\cdot 5/\(\\text\{hi\}\-\\text\{lo\}\), and*do not*convert win rates: a pairwise win\-rate difference is a different estimand, and inventing an exchange rate would be precisely the unstated conversion this paper argues against\. They are reported separately\.

Across the 60 pointwise comparisons the median reported improvement is0\.28 pointson a 0–5 basis \(p25=0\.13=0\.13, p75=0\.63=0\.63\), smaller than every floor in Table[1](https://arxiv.org/html/2609.27787#S5.T1)\. Two comparisons against those floors are available, and they differ in how much they are worth\.

#### The matched comparison\.

Only one benchmark appears often enough in the recovered deltas to be compared against its own measured floor: 17 of the 60 are MT\-Bench\. Their median improvement is 0\.17 points on a 0–5 basis, andall seventeen fall below MT\-Bench’s floor, 0\.63 at 30 items, and still 0\.54 at MT\-Bench’s native 80 items, since items saturate and the floor barely moves\. Nine of the seventeen state an item count of their own, and pricing each of those against the floor at*its own*stated count—rather than at any common one—changes nothing: the count still below is seventeen\. The saturation result is what makes this robust, since a paper would have to have run orders of magnitude more items, not a few more, to buy itself a materially lower floor\. This is the only strictly like\-for\-like statement we can make, andn=17n=17is small, but it is exact\.

#### The pooled comparison, and why it is an upper bound\.

Applying a single measured floor to all 60 deltas puts 60–73% below it, depending on which floor \(AlpacaEval 2’s 0\.41 at its native 805 items to Arena\-Hard’s 0\.61 at its native 500\), and 85% below our screened set’s 1\.24\. We report the range rather than a point, and as an upper bound, for three reasons\. Most of the 60 come from bespoke one\-off instruments we never measured, so no floor of ours is strictly theirs\. Published work evaluates at native item counts higher than our commonni=30n\_\{i\}=30, which lowers the true floor\. And §[5\.3](https://arxiv.org/html/2609.27787#S5.SS3)shows our generic rubric is the*conservative*case, so native protocols would lower the floor further still\. Each correction pushes the fraction down\. The defensible claim is the direction and the order of magnitude—reported improvements sit at the scale of the measurement error, not comfortably above it—not a specific percentage\.

## 6What published work discloses

We audited a frame of 628 arXiv papers \(2024–2026\) retrieved by five preregistered queries, fixed before any paper was read; 320 were sampled under a fixed seed, 312 coded successfully and 227 report a judged system comparison\. Table[5](https://arxiv.org/html/2609.27787#S6.T5)reports rates over those 227\. As a check on draw size, the same rates computed on an earlier 120\-paper draw and on the full 320\-paper draw agree closely \(replication 20\.7% vs 19\.8%, uncertainty 45\.7% vs 45\.8%\); those two percentages are over the eligible subsets of each draw, not over the 227\. Coding was performed by a pinned model at temperature zero, constrained to quote verbatim evidence or answernot\_stated, with inference forbidden; hand validation of spot\-checked fields agreed at≈\\approx88%\. Two earlier regex\-based coders were discarded after validation showed they could not distinguish training runs from evaluation runs, or a data\-source model from a judge\.

Table 5:Disclosure among 92 papers that report a judged system comparison\.#### Do the checkable claims survive?

Disclosure being absent is not the same as conclusions being wrong, so we close the loop where the data allows\. Of the 60 pointwise improvements recovered, 17 are on MT\-Bench, the one benchmark whose floor we measure directly\. Setting each against MT\-Bench’s floor at its own native 80 items,none of the seventeen exceeds it\. Thirty\-nine of the 60 papers also state their item count, so the same test could be extended once floors exist for their benchmarks; we report the matched subset rather than transferring a floor across instruments\.

Combined with §[5\.1](https://arxiv.org/html/2609.27787#S5.SS1)and §[5\.6](https://arxiv.org/html/2609.27787#S5.SS6), the field is reporting differences well below its instruments’ resolution while omitting the information needed to notice\. The coding schema, both automated codings, the human coding and the codebook are available from the author on request\.

## 7Recommendations

1. 1\.Report the judge universe\.A reliability or significance claim is undefined without it\.
2. 2\.Use the benchmark’s native protocol, and say which you used\.This is the cheapest lever we found: it costs nothing per call, and on matched material it moved MT\-Bench’s ceiling from 0\.623 to 0\.798 and the judges needed forE​ρ2=0\.8E\\rho^\{2\}=0\.8from 18 to 3\. Rubric, scale and reference answers are design variables, not incidental formatting\.
3. 3\.Spend on judges before items\.Past a modest item count the marginal return on items is near zero; judges are the binding axis\.
4. 4\.Report MDD alongside the delta\.It is precisely estimable where reliability coefficients are not\.
5. 5\.Screen items for discriminationwhen constructing benchmarks\.
6. 6\.Disclose replication, and whether the inference path caches responses\.

## 8Limitations

The system panels behind Table[1](https://arxiv.org/html/2609.27787#S5.T1)differ between rows in two ways\. The screened set omitsclaude\-opus\-5andclaude\-sonnet\-5, unavailable as systems at collection time, so its base panel is 36 rather than 42; and it carries 23 null\-arm replicas the public sets do not, giving 59\. Refitting without the replicas moves nothing material \(ceiling 0\.735 against 0\.746, MDD 1\.25 against 1\.24\), but the rows are matched on judges, rubric and scale rather than on systems, and cross\-set statements should be read with that in mind\.

#### Judge panels\.

The closed panel spans two model families; a third was unavailable throughout, so panel\-composition generality is untested beyond the splits in §[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px1)and §[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px3)\. The open\-weight panel is five models at a single quantisation \(q4\_K\_M for the larger three\) served locally at a 4096\-token context\. Quantisation and serving configuration are plausible facets we did not vary, so the provenance contrast is between*these*panels rather than between open and closed weights in general\. Qwen3 8B was excluded after 13\.0% of its calls returned empty completions in thinking mode, leaving one family represented by a single model\.

#### Protocol\.

Items are scored under a common rubric rather than each benchmark’s native protocol\. That isolates the item effect but states results for a standard pointwise protocol applied to those items\. We replicate MT\-Bench natively \(§[5\.3](https://arxiv.org/html/2609.27787#S5.SS3)\) and find the generic rubric the conservative case, and we now measure a pairwise protocol directly \(§[5\.4](https://arxiv.org/html/2609.27787#S5.SS4)\), which is the form Arena\-Hard and AlpacaEval 2 actually take\. That arm is not those benchmarks as deployed: it uses our items, our judges and a single fixed baseline, rather than Arena\-Hard’s own prompt set, template and reference model\. So the gap for those two rows is now bounded rather than closed, we can say what a pairwise protocol does to the ceiling in general, but not what each benchmark’s specific pipeline does to it\. Given that the pairwise arm moves the ceiling more than any other manipulation in this paper, this remains the largest source of uncertainty about how our numbers transfer, and the most useful thing a replication could pin down\.

#### Data quality\.

The Arena\-Hard cell rests on 17 items against 32–43 for the other sets\. This is a collection shortfall rather than an attrition one: the harvest recorded no errors and no dropped cells, and the source pool holds 500 prompts, so the constraint was budget rather than the benchmark\. Itsσs2\\sigma^\{2\}\_\{s\}is accordingly the least precisely estimated of the four, and its ceiling carries the widest interval in Table[1](https://arxiv.org/html/2609.27787#S5.T1)\. Approximately 15% of generations were truncated at the token cap, which penalises verbose systems consistently across judges\.

#### Scope of the human and audit comparisons\.

The SummEval comparison is one dimension of one dated summarisation task and plausibly flatters judges\. The audit’s replication field conflates repetition for variance with repeated calls for order control, so 19\.8% is a conservative upper bound\.

## 9Conclusion

Judged benchmarks have a resolution and it is measurable\. Because the system\-by\-judge interaction carries no dependence on item count, a single\-judge protocol has a ceiling that more data cannot raise\. That much is algebra and holds for any judge panel, any rubric, any benchmark\.

Where that ceiling sits is not algebra, and we found it to be strongly conditional\. Under a nine\-judge universe that includes a legacy and a deliberately weak judge, none of our four item sets reaches conventional reliability at any item count; under six current\-generation judges, three of the four do\. Switching MT\-Bench from a generic rubric to its own native protocol moves its ceiling from 0\.623 to 0\.798 on matched material\. A reliability number for a judged benchmark is therefore not a property of the benchmark\. It is a property of the benchmark, the judge universe and the protocol together, and reporting it without the latter two leaves it undetermined rather than merely imprecise, which is what nearly all published work currently does\.

Two things follow that do not depend on the conditions\. The exchange rate between judges and items is steep enough to reverse the usual instinct, and it steepens as the target rises: two judges at thirty items are worth 77 items at one judge atE​ρ2=0\.75E\\rho^\{2\}=0\.75, and past a modest count items are nearly worthless, the count needed diverges as the target nears whatever ceiling the protocol produces\. And the improvements the field reports sit at the scale of the measurement error rather than comfortably above it, every one of the seventeen MT\-Bench improvements we recovered falls below MT\-Bench’s own floor\. The limit is not a failure of LLM judges, which are more self\-consistent than human experts on a matched design, but a property of the measurement design, which binds human raters equally\.

## Appendix ASupporting analyses

Each result below is stated with the numbers that carry it\. Full per\-cell tables are in the released artifacts \(Appendix[F](https://arxiv.org/html/2609.27787#A6)\)\.

#### The judge universe\.

E​ρ2E\\rho^\{2\}is defined relative to a universe of admissible judges, and published work never states one\. Recomputing the ceiling under four choices moves MT\-Bench from0\.6000\.600\(all nine\) to0\.6490\.649\(drop legacy\),0\.6910\.691\(drop weak\) and0\.7980\.798\(current\-generation six\); our screened set moves0\.746→0\.804→0\.840→0\.9290\.746\\to 0\.804\\to 0\.840\\to 0\.929, Arena\-Hard0\.792→0\.9180\.792\\to 0\.918, AlpacaEval 20\.706→0\.8290\.706\\to 0\.829\. The ordering is stable, the level is not, so a reliability figure reported without its judge universe is uninterpretable\. Section[5\.5](https://arxiv.org/html/2609.27787#S5.SS5)explains which one we take as primary\.

#### Route versus model identity\.

The route replica serves the same weights asanth\-cheapthrough a different provider endpoint, so the variance of their interaction difference isolates route\. It is0\.00100\.0010–0\.00740\.0074across the four item sets, or1\.4%1\.4\\%–5\.0%5\.0\\%of the disagreement distinct current\-generation models produce on the same systems\.σs​j2\\sigma^\{2\}\_\{sj\}is about*which model judges*, not where it is served; a panel of one model behind several endpoints buys almost none of the independence a panel of distinct models buys\.

#### Judge provenance is a facet\.

Five open\-weight judges on identical material disagree2\.17×2\.17\\timesmore across model families than within them once scale usage is standardised away \(2\.92×2\.92\\timeson raw scores, before standardisation, the difference is a scale\-compression artefact and we report the standardised figure\)\. Mixing families buys independence but not a smaller panel\.

#### Discriminating power at the frontier\.

Published comparisons are between closely matched frontier systems, not across a 2024–2026 capability range\. Restricted to ten current frontier models, system variance falls to1\.9%1\.9\\%on our screened items and2\.8%2\.8\\%on MT\-Bench\. Arena\-Hard is the exception at15\.7%15\.7\\%, consistent with its design intent and the one benchmark still separating frontier systems\.

#### Is the screened set circular?

Our items were selected for between\-system variance, so finding it afterwards would be guaranteed\. Two checks\. The screener \(gpt\-4\.1\) is deliberately outside the evaluation panel, and splitting that panel by vendor family shows the screened set lifts the system share3\.16×3\.16\\timesover MT\-Bench for OpenAI\-family judges and3\.82×3\.82\\timesfor Anthropic\-family, the gain is*larger*on the family least like the screener, the opposite of selection on one model’s taste\. And ranking items by discrimination within random halves of the system panel and correlating the two rankings gives a split\-half Spearmanρ=0\.757\\rho=0\.757\(95% range0\.4130\.413–0\.8930\.893\), the highest of the four item sets, against0\.5970\.597for MT\-Bench and0\.5510\.551for Arena\-Hard\. Screening captured a transferable property, not the screener’s preferences\.

#### What counts as a system\.

The 42 systems are 14 models×\\times3 prompt configurations, so configurations are nested in models rather than crossed\. Fitting the nested model, configuration\-given\-model carries55\.7%55\.7\\%of system variance on MT\-Bench,46\.7%46\.7\\%on AlpacaEval 2,26\.1%26\.1\\%on Arena\-Hard and0%0\\%on our screened items\. Treating the model as the object of measurement lowers MT\-Bench’s ceiling from0\.6000\.600to0\.4040\.404and AlpacaEval 2’s from0\.7060\.706to0\.5670\.567\. Whether a benchmark can separate*models*is a harder question than whether it can separate*system configurations*, and the results in the body answer the easier one\.

#### A panel and an ensemble are the same object\.

Almost nobody reportsnjn\_\{j\}separate scores; they average the panel into one\. G\-theory says the mean ofkkjudges carriesσs​j2/k\\sigma^\{2\}\_\{sj\}/k, which is falsifiable\. A single ensemble judge cannot be fitted—atnj=1n\_\{j\}=1the judge facet is unidentified—so we form*disjoint*ensembles of sizekk, refit on those⌊nj/k⌋\\lfloor n\_\{j\}/k\\rfloorensemble judges, and compare the resultingσs​j2\\sigma^\{2\}\_\{sj\}against the prediction\. Over 60 random partitions per cell atk=2k=2andk=3k=3on all four item sets, observed over predicted ranges0\.930\.93–1\.111\.11with median1\.011\.01\. Every judge count in this paper may be read as an ensemble size\.

#### Temperature is not separated\.

All judges run at a fixed decoding temperature, so what the replicate facet measures is run\-to\-run variation*at*that temperature, which isσe2\\sigma^\{2\}\_\{e\}and not a temperature effect\. Separating it needs an arm crossing temperature with judge, which we have not run\. This is the one place where “judge” remains a bundle\.

#### Reliability is not validity\.

E​ρ2E\\rho^\{2\}asks whether a panel reproduces itself, not whether it is right\. On SummEval\([9](https://arxiv.org/html/2609.27787#bib.bib14)\), moving from\(ni=10,nj=1\)\(n\_\{i\}\{=\}10,n\_\{j\}\{=\}1\)to\(ni=50,nj=9\)\(n\_\{i\}\{=\}50,n\_\{j\}\{=\}9\)raisesE​ρ2E\\rho^\{2\}from0\.7110\.711to0\.9590\.959\(\+0\.248\+0\.248\) while Kendallτ\\tauagainst the three\-expert consensus rises only from0\.5960\.596to0\.7100\.710\(\+0\.114\+0\.114\) and is flat over the last four configurations\. Recovery of the human top three sits at23%23\\%–32%32\\%across every configuration against a18\.8%18\.8\\%chance baseline, and does not improve on either axis\. Reliability is a precondition for a believable comparison, not a substitute for one\.

#### The limit is structural, not a judge deficiency\.

On SummEval’s fully crossed human design \(16 systems×\\times100 articles, 3 experts and 5 crowd workers\), rater\-by\-system variance per unit of system signal is0\.070\.07–0\.130\.13for LLM judges against0\.340\.34for experts and0\.670\.67for crowd workers\. Two judges match three experts; crowd workers produce essentially no system\-level signal \(σs2≈0\.002\\sigma^\{2\}\_\{s\}\\approx 0\.002at five raters\)\. Substituting humans makes the resolution worse at far greater cost\.

#### What does not explain the floor\.

Type I error is correctly calibrated: on pairs with a true difference of exactly zero—two independent generations of the same system—the modal protocol’s false\-positive rate is2\.5%2\.5\\%, below nominal\. The problem is power, not calibration\. Self\-preference is real but small, explaining0\.8%0\.8\\%–4\.8%4\.8\\%ofσs​j2\\sigma^\{2\}\_\{sj\}\. Length explains more: within\-item correlation between response length and score is0\.210\.21–0\.250\.25on the public benchmarks and length explains13%13\\%–17%17\\%of between\-system variance, so part of what those benchmarks measure as quality is verbosity\. On our screened items, which carry explicit format and length constraints, the correlation is−0\.004\-0\.004, item design can remove this\.

#### What the floor costs in practice\.

Three consequences a practitioner can act on\.*The naive bootstrap understates the interval\.*Resampling items only—what a careful author does today—gives a half\-width of0\.340\.34–0\.490\.49points against a true single\-judge MDD of0\.580\.58–1\.241\.24, understating by1\.3×1\.3\\timesto2\.8×2\.8\\timesbecause it ignores the judge facet entirely\.*Rank inversions are near chance for adjacent systems\.*For two systems adjacent in the true top ten, whose mean separation is0\.0100\.010–0\.0130\.013points, the probability of returning them in the wrong order is45%45\\%–49%49\\%at one judge and barely moves at three\.*Under a call budget the optimum is interior\.*Minimisingσδ2\\sigma^\{2\}\_\{\\delta\}subject to cost∝ni​nj​nr\\propto n\_\{i\}n\_\{j\}n\_\{r\}gives, on MT\-Bench,30×4×130\\times 4\\times 1at 120 calls,60×6×160\\times 6\\times 1at 360 and108×10×1108\\times 10\\times 1at 1080; on our screened items20×6×120\\times 6\\times 1,36×10×136\\times 10\\times 1and60×18×160\\times 18\\times 1\. The optimal judge count grows with budget rather than saturating, and a single run is optimal at every budget examined\.

## Appendix BReading the main tables

#### Raw points do not rank item sets\.

The four sets use different amounts of the scale, so a raw MDD is measured against a different ruler in each\. System means span2\.632\.63points on our screened set against1\.091\.09on MT\-Bench and1\.181\.18on AlpacaEval 2\. In units of each set’s own between\-system spread the ordering changes:1\.571\.57SD for Arena\-Hard,1\.771\.77for the screened set,2\.312\.31for AlpacaEval 2,2\.672\.67for MT\-Bench, so the screened set has the largest raw MDD and the second smallest standardised one\. This is the same scale\-compression artefact we document for judge panels in §[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px3): any absolute\-scale comparison of noise rewards an instrument for compressing its output\. Raw points are the right unit for comparing a benchmark’s floor against improvements reported*on that benchmark*\(§[5\.6](https://arxiv.org/html/2609.27787#S5.SS6)\), and the wrong unit for ranking benchmarks against each other\.

#### What screening bought, and what it cost\.

Screening raisedσs2\\sigma^\{2\}\_\{s\}by an order of magnitude over MT\-Bench \(0\.4560\.456against0\.0440\.044\) and roughly doubled the usable scale range, but raised the interaction terms alongside it \(σs​i2\\sigma^\{2\}\_\{si\}0\.4740\.474against0\.2860\.286;σs​i​j2\\sigma^\{2\}\_\{sij\}0\.6830\.683against0\.2670\.267\)\. That is unsurprising on reflection: items selected for separating systems will also separate them in*different orders*, which is system\-by\-item interaction by definition\. The net is a single\-judge ceiling of0\.7460\.746and a standardised MDD of1\.771\.77SD, in both cases second to Arena\-Hard’s0\.7920\.792and1\.571\.57\. Screening buys a better instrument than the two weakest public sets and does not beat the strongest: a lever worth pulling, not a solved problem\.

#### Three MDD figures, three questions\.

MDD@30 \(0\.580\.58–1\.241\.24\) evaluates every set at a common thirty items, which is what makes them comparable\. MDD at native counts \(0\.410\.41–1\.241\.24\) evaluates each at the count its users deploy, the right figure against a*published*result\. MDD at infinite items \(0\.400\.40–1\.091\.09\) is the floor no item budget can beat under one judge\. All three are single\-judge, single\-run, and none moves much, because items saturate\.

#### The object of measurement\.

σs2\\sigma^\{2\}\_\{s\}is defined over a*population of systems*, exactly asE​ρ2E\\rho^\{2\}is defined over a universe of judges\. “MT\-Bench has 4\.5% system variance” is shorthand for “across the 42 systems studied here”\. We report the near\-frontier subset separately \(§[A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px4)\) precisely because the choice of system population changes the answer: discriminating power is a property of the pair, not of the benchmark alone\. Neither facet should be reported without its universe\.

## Appendix CAudit and infrastructure detail

#### Validation\.

Coding reliability was established in two stages\. The sample was first coded a second time by a different model family \(GPT\-5\.6\) under an independently worded prompt\. A 30\-paper subsample was then coded by a human working*blind*—no automated answers, no evidence quotes—so that agreement measures coding difficulty rather than suggestibility\.

Against the human, overall agreement is 88\.0% \(κ=0\.73\\kappa=0\.73\)\. Crucially, the field the headline depends on validates almost perfectly: whether the evaluation was repeated reachesκ=0\.89\\kappa=0\.89with a single disagreement in thirty papers, the human coding 16\.7% and the automated coder 20\.0%\. Whether the scorer was named is perfect \(κ=1\.00\\kappa=1\.00\)\. Uncertainty reporting is substantial \(κ=0\.65\\kappa=0\.65\)\.

The human coding also resolves an ambiguity the two automated coders could not\. Between themselves they agreed only moderately on replication \(κ=0\.47\\kappa=0\.47, estimating 15% and 29%\), which we had provisionally reported as a range\. The human agrees closely with the first coder and not the second, indicating that the disagreement was one coder in error rather than genuine ambiguity in the papers\. We therefore report the lower, validated figure\.

One field remains genuinely ambiguous and is reported as a range: version pinning \(κ=0\.45\\kappa=0\.45; 53% automated, 73% human\), where the human applied a more generous standard for what counts as a pinned version\. Item counts show 86\.7% agreement, thoughκ\\kappais undefined there because the human coded every paper affirmatively\.

#### Closing the loop: does disclosure predict survival?

The audit so far shows that disclosure is absent, not that any conclusion was wrong, and those are different complaints\. The link we can actually test is whether the papers that*do*report enough to be checked are also the ones whose claims clear a floor\. Of the 227 papers reporting a judged comparison,146146\(64\.3%64\.3\\%\) state both a judge model and an item count, which is the minimum needed to compute an implied floor at all; the remaining third cannot be checked by anyone, including their own authors\.

Among the 60 recovered deltas, 35 come from papers stating both and 25 do not\. Applying the same MT\-Bench floor to both groups at each paper’s own stated item count,40%40\\%of the fully\-disclosing claims exceed it against24%24\\%of the rest\. The absolute level here is not trustworthy—it transfers one instrument’s floor to claims made on others, which is the error this paper argues against—but the*contrast*is, because both groups are priced against the identical floor, so whatever bias the transfer introduces is common to them\. The direction is worth stating plainly: better\-disclosed claims are not the fragile ones\. Disclosure is not what makes a result survive, but it does travel with results that do, most likely because papers that report an item count tend to report a larger one\.

This is the causal link between Section[6](https://arxiv.org/html/2609.27787#S6)and Section[5\.1](https://arxiv.org/html/2609.27787#S5.SS1), and it cuts against the most cynical reading of the audit\. The problem is not that the field is making claims it knows to be unsupported; it is that two\-thirds of a corpus can be checked and one\-third cannot, and no reader can tell from the outside which third a given paper is in\. That is a reporting failure with a cheap fix, and it is the one recommendation in this paper that costs nothing to adopt\.

## Appendix DEstimation notes

#### Where the two\-system difference comes from\.

Two systems measured on the*same*items and judges share those facets, so the item and judge main effects cancel in the difference and only the interactions with system survive\. That is whyσδ2\\sigma^\{2\}\_\{\\delta\}carriesσs​i2\\sigma^\{2\}\_\{si\},σs​j2\\sigma^\{2\}\_\{sj\},σs​i​j2\\sigma^\{2\}\_\{sij\}and the residual but notσi2\\sigma^\{2\}\_\{i\},σj2\\sigma^\{2\}\_\{j\}orσi​j2\\sigma^\{2\}\_\{ij\}: this is relative error, appropriate for ranking or comparing systems within a study, and not absolute error, which is what one would need to compare a score against a fixed external standard\. The variance of the difference is2​σδ22\\sigma^\{2\}\_\{\\delta\}because the two systems are independent draws from the system population given the shared design\.

#### The 1\.96 factor\.

MDD95≈1\.96​2​σδ2\\mathrm\{MDD\}\_\{95\}\\approx 1\.96\\sqrt\{2\\sigma^\{2\}\_\{\\delta\}\}assumes the difference of two system means is approximately normal\. Withni≥20n\_\{i\}\\geq 20items the central limit theorem is doing the work and the approximation is unremarkable; at very smallnin\_\{i\}it will be optimistic, and attquantile on the appropriate error degrees of freedom should be substituted\.

#### ANOVA rather than REML, and negative estimates\.

We estimate components by expected mean squares on the complete crossed block\. The estimator is unbiased and closed form, which matters here because we refit under many judge universes, system populations and item subsets, and a closed form makes those refits cheap and deterministic\. Its known cost is that individual component estimates can go negative when the true value is near zero; we truncate at zero when reporting variance*shares*and leave the raw value in place when computingσδ2\\sigma^\{2\}\_\{\\delta\}, so that truncation never flatters a floor\. REML would avoid negative estimates and is the better choice for a single definitive fit, but it is iterative and would make the bootstrap over 600–800 resamples materially more expensive\. Where we report intervals we use a parametric bootstrap rather than Wald intervals, precisely because the sampling distribution of a variance component near the boundary is not symmetric\.

#### Degrees of freedom onσs2\\sigma^\{2\}\_\{s\}\.

The system facet hasns−1n\_\{s\}\-1degrees of freedom, so with 42 systemsσs2\\sigma^\{2\}\_\{s\}is the least precisely estimated component in the design, and it is the numerator of every reliability quantity\. This is the main reason we report ceilings with bootstrap intervals rather than as point estimates, and why Arena\-Hard—which carries the fewest items and therefore the widest interval—is the cell we are least willing to make claims about\.

## Appendix EAdditional figures

Figure 2:Iso\-cost allocation\. At a fixed budget ofni​nj​nrn\_\{i\}n\_\{j\}n\_\{r\}calls, MDD against how the budget is split; the marked point is the optimum\. It is interior at every budget, and it moves toward more judges as the budget grows rather than saturating\.Figure 3:Probability that two adjacent systems in the true top ten are returned in the wrong order, against item count, at one and three judges\. The curves flatten well short of the 0\.5 chance line: items buy rank stability slowly, and a third judge buys more than doubling the items\.Figure 4:Sizing nomogram\. Contours ofE​ρ2E\\rho^\{2\}over judges and items; read off the panel a target requires\. Each contour flattens to a horizontal asymptote, and below that judge count no item budget reaches the target, the ceiling, drawn\.
## Appendix FArtifacts

#### Sieve: the discrimination\-screened item set\.

The screened row throughout this paper is a releasable benchmark, named here so it can be cited and criticised separately from our results\.Sieveis 38 open\-ended instruction items selected for between\-system discriminating power, constructed as follows\.*Pool:*71 items across four strata designed to force systems apart rather than to sample a task distribution,*pushback*\(24 items, a false premise or unsafe\-but\-plausible ask\),*compose*\(18, interacting constraints\),*constraint*\(16, a checkable format or length rule\),*verifiable*\(13, a checkable factual or arithmetic core\)\.*Screening design:*55systems×\\times71 items×\\times1 judge×\\times2 replicates, 720 calls, on the same holistic 0–5 rubric used in the main study; a checklist variant discriminated roughly4×4\\timesworse and was dropped\.*Screening judge:*openai/gpt\-4\.1\-2025\-04\-14, deliberately*not*a member of the nine\-judge evaluation panel, so no item is selected using a judge that later scores it\.*Selection rule:*retain items with between\-system variance≥1\.0\\geq 1\.0on the 0–5 scale; pool mean was3\.993\.99\(sd1\.361\.36\) and 38 of 71 cleared it\. The main study harvests 30 of the 38\. Sieve is a measurement instrument, not a capability benchmark: items were chosen because systems answer them differently, not because they represent any task distribution\.

#### Artifacts\.

The crossed judgment dataset \(373,019 calls across four item sets, nine judges, three replicates, with per\-call model, prompt configuration, replicate index, raw completion and parsed score\); Sieve, its stratum labels and per\-item discrimination scores, plus the 33 screened\-out items; the pairwise arm \(4,608 preference judgments in both orders\); a D\-study calculator that takes variance components and returns item and judge requirements; and the audit, 628\-paper frame, both automated codings, the blind human coding, and the codebook\.

#### Licensing and reproduction\.

Data and item sets under CC BY 4\.0, code under MIT\. Judgment records carry the dated model identifier requested rather than the alias, since the proxy echoes the alias back; readers reproducing against another endpoint should expect different snapshots behind the same names\. The endpoint used here enforces a client allowlist at the edge, so an external reader cannot replay the harvest against it and must supply their own provider credentials\.

#### Availability\.

The datasets and code described above are available from the author on request\.

## References

- Bowman and Dahl \(2021\)S\. R\. Bowman and G\. E\. DahlWhat will it take to fix benchmarking in natural language understanding?\.Proceedings of NAACL\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px4.p1.1)\.
- Brennan \(2001\)R\. L\. BrennanGeneralizability theory\.Springer,New York\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.SS0.SSS0.Px1.p1.2),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- Cardet al\.\(2020\)D\. Card, P\. Henderson, U\. Khandelwal, R\. Jia, K\. Mahowald, and D\. JurafskyWith little power comes great responsibility\.InProceedings of EMNLP,Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px1.p1.1)\.
- Choiet al\.\(2026\)J\. Choi, S\. Park, C\. Cho, H\. Park, and B\. KimDiagnosing the reliability of LLM\-as\-a\-judge via item response theory\.arXiv preprint arXiv:2602\.00521\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- Cronbachet al\.\(1972\)L\. J\. Cronbach, G\. C\. Gleser, H\. Nanda, and N\. RajaratnamThe dependability of behavioral measurements: theory of generalizability for scores and profiles\.Wiley,New York\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- Deutschet al\.\(2021\)D\. Deutsch, R\. Dror, and D\. RothA statistical analysis of summarization evaluation metrics using resampling methods\.Transactions of the Association for Computational Linguistics9\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px1.p1.1)\.
- Droret al\.\(2018\)R\. Dror, G\. Baumer, S\. Shlomov, and R\. ReichartThe hitchhiker’s guide to testing statistical significance in natural language processing\.Proceedings of ACL\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px1.p1.1)\.
- Duboiset al\.\(2024\)Y\. Dubois, B\. Galambosi, P\. Liang, and T\. B\. HashimotoLength\-controlled AlpacaEval: a simple way to debias automatic evaluators\.External Links:2404\.04475Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px4.p1.1),[§4](https://arxiv.org/html/2609.27787#S4.SS0.SSS0.Px3.p1.1)\.
- Fabbriet al\.\(2021\)A\. R\. Fabbri, W\. Kryściński, B\. McCann, C\. Xiong, R\. Socher, and D\. RadevSummEval: re\-evaluating summarization evaluation\.Transactions of the Association for Computational Linguistics9,pp\. 391–409\.Cited by:[Appendix A](https://arxiv.org/html/2609.27787#A1.SS0.SSS0.Px9.p1.1)\.
- Feueret al\.\(2025\)B\. Feuer, C\. Tseng, A\. S\. Lathe, O\. Elachqar, and J\. P\. DickersonWhen judgment becomes noise: how design failures in LLM judge benchmarks silently undermine validity\.arXiv preprint arXiv:2509\.20293\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- Liet al\.\(2024\)T\. Li, W\. Chiang, E\. Frick, L\. Dunlap, B\. Zhu, J\. E\. Gonzalez, and I\. StoicaFrom crowdsourced data to high\-quality benchmarks: Arena\-Hard and BenchBuilder pipeline\.External Links:2406\.11939Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px4.p1.1),[§4](https://arxiv.org/html/2609.27787#S4.SS0.SSS0.Px3.p1.1)\.
- Messing \(2026\)S\. MessingHidden measurement error in LLM pipelines distorts annotation, evaluation, and benchmarking\.arXiv preprint arXiv:2604\.11581\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- Miller \(2024\)E\. MillerAdding error bars to evals: a statistical approach to language model evaluations\.External Links:2411\.00640Cited by:[§1](https://arxiv.org/html/2609.27787#S1.p2.1),[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px1.p1.1)\.
- Normanet al\.\(2026\)J\. D\. Norman, M\. U\. Rivera, and D\. A\. HughesReliability without validity: a systematic, large\-scale evaluation of LLM\-as\-a\-judge models across agreement, consistency, and bias\.arXiv preprint arXiv:2606\.19544\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- Schaefferet al\.\(2023\)R\. Schaeffer, B\. Miranda, and S\. KoyejoAre emergent abilities of large language models a mirage?\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2609.27787#S1.p2.1),[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px1.p1.1)\.
- Shavelson and Webb \(1991\)R\. J\. Shavelson and N\. M\. WebbGeneralizability theory: a primer\.Sage\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px2.p1.1)\.
- Songet al\.\(2025\)D\. Song, W\. Lee, and H\. JiaoExploring LLM autoscoring reliability in large\-scale writing assessments using generalizability theory\.arXiv preprint arXiv:2507\.19980\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- The Machine Learning Society \(2026\)The Machine Learning SocietyThe sample complexity of LLM evaluation\.Note:Computes minimum detectable effects for accuracy benchmarks and identifies judge\-based sizing as future work[https://www\.tmls\.nyc/research/eval\-sample\-complexity](https://www.tmls.nyc/research/eval-sample-complexity)Cited by:[§1](https://arxiv.org/html/2609.27787#S1.p2.1),[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px1.p1.1)\.
- Yagubyan \(2026\)A\. YagubyanThe coin flip judge? reliability and bias in LLM\-as\-a\-judge evaluation\.arXiv preprint arXiv:2606\.13685\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- Yanget al\.\(2026\)Z\. Yang, Y\. Hou, and X\. YangWhen the judge changes, so does the measurement: auditing LLM\-as\-judge reliability\.arXiv preprint arXiv:2607\.08535\.Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.27787#S3.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,Cited by:[§2](https://arxiv.org/html/2609.27787#S2.SS0.SSS0.Px4.p1.1),[§4](https://arxiv.org/html/2609.27787#S4.SS0.SSS0.Px3.p1.1)\.

Similar Articles

Benchmarking LLMs

Reddit r/AI_Agents

A study or report on benchmarking large language models, likely comparing performance across various tasks.

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

arXiv cs.CL

This paper investigates whether standard benchmarks underestimate LLM performance by re-evaluating hallucination detection datasets using an LLM-first, human-adjudicated assessment method. The study finds that incorporating LLM reasoning into the adjudication process improves agreement and suggests that model-assisted re-evaluation yields more reliable benchmarks for ambiguity-prone tasks.

Quantifying Ranking Uncertainty in LLM Benchmarks

arXiv cs.LG

This paper analyzes sources of ranking uncertainty in LLM benchmarks like MMLU, proposing modifications to hypothesis tests for constructing rank confidence intervals, and shows that variability across subjects is substantial.