一个评委小组相当于多少人类?
摘要
本文介绍了量化语言模型小组模拟人类判断有效程度的方法,使用谱多样性和分布恢复指标来衡量有效表示。
arXiv:2609.21277v1 Announce Type: new
Abstract: How many human judgments does a panel of language models represent? The answer depends on what is matched. We audit categorical judge panels against empirical human label distributions, retaining disagreement that binary errors relative to one gold label collapse. We measure spectral residual diversity by matching the participation ratio of a normalized residual Gram matrix to conditionally independent human-reference draws, giving nu_H. We separately match distributional squared error, giving nu_MSE. Across three ChaosNLI tasks, the same 32-judge panels have nu_H=4.24--6.50 but nu_MSE=2.30--3.75. A spectral identity separates the eigenvalues, member energies, and averaging-direction weights that determine error. Realizable hard-label panels show that greater spectral diversity can accompany worse distribution recovery even with equal member energies and nonnegative correlations. In the observed panels, within-size ranking agreement varies sharply by task; some member additions produce conflicting changes that persist across two item halves. The consensus-direction share of centered residual variance is gamma_co=43.8% on MNLI-m and 33.7% on SNLI, quantifying shared variation retained by averaging. We provide aligned votes and analysis protocols for auditing these distinctions. Effective size is therefore a target-specific measurement: spectral diversity and distribution recovery should not be treated as interchangeable measures of panel quality or as general human-replacement rates.
查看缓存全文
缓存时间: 2026/09/21 09:05
# How Many Humans Is a Judge Panel Worth?
Source: [https://arxiv.org/html/2609.21277](https://arxiv.org/html/2609.21277)
\[BoldFont=texgyretermes\-bold\.otf,ItalicFont=texgyretermes\-italic\.otf,BoldItalicFont=texgyretermes\-bolditalic\.otf\] \[BoldFont=texgyreheros\-bold\.otf\]
Chao LiYingying YuYunfeng LiTsinghua UniversityUniversity College LondonAutonavilichaocasia@gmail\.comyuyingying311@gmail\.comliyunfengxiaolong@gmail\.com††thanks:Corresponding author\.
###### Abstract
How many human judgments does a panel of language models represent? The answer depends on what is matched\. We audit categorical judge panels against empirical human label distributions, retaining disagreement that binary errors relative to one gold label collapse\. We measure spectral residual diversity by matching the participation ratio of a normalized residual Gram matrix to conditionally independent human\-reference draws, givingνH\\nu\_\{H\}\. We separately match distributional squared error, givingνMSE\\nu\_\{\\mathrm\{MSE\}\}\. Across three ChaosNLI tasks, the same 32\-judge panels haveνH=4\.24\\nu\_\{H\}=4\.24–6\.506\.50butνMSE=2\.30\\nu\_\{\\mathrm\{MSE\}\}=2\.30–3\.753\.75\. A spectral identity separates the eigenvalues, member energies, and averaging\-direction weights that determine error\. Realizable hard\-label panels show that greater spectral diversity can accompany worse distribution recovery even with equal member energies and nonnegative correlations\. In the observed panels, within\-size ranking agreement varies sharply by task; some member additions produce conflicting changes that persist across two item halves\. The consensus\-direction share of centered residual variance isγco=43\.8%\\gamma\_\{\\mathrm\{co\}\}=43\.8\\%on MNLI\-m and33\.7%33\.7\\%on SNLI, quantifying shared variation retained by averaging\. We provide aligned votes and analysis protocols for auditing these distinctions\. Effective size is therefore a target\-specific measurement: spectral diversity and distribution recovery should not be treated as interchangeable measures of panel quality or as general human\-replacement rates\.
## 1Introduction
Language models increasingly supply judgments for evaluation and data annotation\. Their agreement with human judgments motivates their use in costly manual assessments\([Zheng et al\., 2023](https://arxiv.org/html/2609.21277#bib.bib34)\)\. Combining models can improve human agreement and reduce evaluation cost\([Verga et al\., 2024](https://arxiv.org/html/2609.21277#bib.bib31)\), but adding judges also adds correlated responses\. An effective\-vote analysis found that 9 judges could supply roughly two independent votes\([Kohli, 2026b](https://arxiv.org/html/2609.21277#bib.bib16)\)\. A panel count alone therefore says little about what its judgments contribute\.
An effective count also needs an explicit target\. Binary errors relative to a single gold label retain whether a judge agrees with that label, but collapse the identities of all alternatives\. If human label frequencies are\(0\.6,0\.3,0\.1\)\(0\.6,0\.3,0\.1\), choosing the second or third label incurs the same binary error while deviating from human judgments in different directions\. Human disagreement in language inference has reproducible structure\([Pavlick and Kwiatkowski, 2019](https://arxiv.org/html/2609.21277#bib.bib24)\), and label variation matters for evaluation\([Plank, 2022](https://arxiv.org/html/2609.21277#bib.bib25)\)\. Dense annotations in ChaosNLI\([Nie et al\., 2020](https://arxiv.org/html/2609.21277#bib.bib23)\)allow us to measure model residuals relative to that distribution\.
Retaining this reference leaves a second distinction: residual diversity versus distribution recovery\. A panel can span several residual directions while its average label frequencies remain far from the human distribution\. Participation ratio \(PR\) summarizes spectral dimension\([Mazzucato et al\., 2016](https://arxiv.org/html/2609.21277#bib.bib22)\); the squared error of an average also depends on member error energies and the alignment of residual directions with averaging\. Prior analyses study correlated model errors\([Kim et al\., 2025](https://arxiv.org/html/2609.21277#bib.bib13)\), diversity versus majority\-vote gains\([Kim, 2026](https://arxiv.org/html/2609.21277#bib.bib12)\), and common errors retained by aggregation\([Afrin and Shihab, 2026](https://arxiv.org/html/2609.21277#bib.bib1)\)\. These findings motivate a joint audit of what different panel measurements say about the same answers\.
We study this question in a fixed pool of 32 models from 10 provider families on three categorical language\-inference tasks\. Our target is each item’s empirical human label distribution\. We compare spectral sizeνH\\nu\_\{H\}, distribution\-error matched sizeνMSE\\nu\_\{\\mathrm\{MSE\}\}, and binary\-error effective votesneffn\_\{\\mathrm\{eff\}\}, then examine when the spectral and distributional rankings agree\. The contribution is the human\-distribution audit and its empirical distinctions; PR, spectral identities, and fixed\-direction variance projections are established tools\. We do not propose a new selection or aggregation algorithm\.
Our findings have three parts:
1. 1\.The dependence visible in fixed model answers depends on its representation and summary\. The 32\-judge panels haveνH=4\.24\\nu\_\{H\}=4\.24,6\.466\.46, and6\.506\.50, compared withνMSE=2\.30\\nu\_\{\\mathrm\{MSE\}\}=2\.30,3\.753\.75, and3\.443\.44\. Crossed comparisons distinguish the complete representation, including encoding, reference, and centering, from the dependence summary; a separate comparison varies only the residual anchor and its matching curve\.
2. 2\.PR omits information needed to determine error\. An exact decomposition identifies the role of member energy and spectral alignment with the averaging direction\. A realizable panel with equal member energies and nonnegative correlations raises PR by 14\.3% while raising squared error by 25%\. The consensus\-direction variance shareγco\\gamma\_\{\\mathrm\{co\}\}is 43\.8% on MNLI\-m and 33\.7% on SNLI, quantifying shared fluctuations retained by averaging\.
3. 3\.The practical extent of the mismatch is task\-dependent\. At selected fixed sizes, PR–error ranking agreement is 0\.960–0\.977 on MNLI\-m but only 0\.223–0\.419 onα\\alphaNLI\. Some member additions conflict across both item halves, while the overallα\\alphaNLI rankings are unstable\. These observations delimit how the measurements can be interpreted in the present pool\.
We supply 96,000 baseline votes and 122,000 presentation\-order records, with failure indicators, requested model identifiers, and analysis protocols\. The results first compare matched targets and their rankings, then visualize residual geometry and variation retained by averaging, and finally examine panel size, presentation changes, and reference sensitivity\. The appendices provide detailed definitions, protocols, and supplementary diagnostics\.
## 2Related Work
##### Model judges and panels\.
Strong agreement with human preferences motivates model\-based evaluation\([Zheng et al\., 2023](https://arxiv.org/html/2609.21277#bib.bib34)\), but judge behavior can be unstable or exploitable\([Thakur et al\., 2025](https://arxiv.org/html/2609.21277#bib.bib29)\)\. Multi\-model panels can improve human agreement and reduce cost\([Verga et al\., 2024](https://arxiv.org/html/2609.21277#bib.bib31)\)\. Cross\-model consensus has also shown benefits over resampling one model for selecting reasoning answers\([Liu, 2026](https://arxiv.org/html/2609.21277#bib.bib19)\)\. Our judges instead assign categorical inference labels; we study their residuals relative to human label distributions, not a general replacement rate for human evaluation\.
##### Dependence and ensemble gains\.
Independence underlies classical voting and variance\-reduction arguments\([Condorcet, 1785](https://arxiv.org/html/2609.21277#bib.bib6);[Dietterich, 2000](https://arxiv.org/html/2609.21277#bib.bib8)\)\. Correlated errors limit the gains from adding models\([Kim et al\., 2025](https://arxiv.org/html/2609.21277#bib.bib13);[Kohli, 2026b](https://arxiv.org/html/2609.21277#bib.bib16)\), and cross\-family similarity also appears in open\-ended generation\([Jiang et al\., 2025](https://arxiv.org/html/2609.21277#bib.bib11)\)\. Conditional\-information criteria address member selection\([Turkmen et al\., 2026](https://arxiv.org/html/2609.21277#bib.bib30)\); diversity audits distinguish member capability from diversity in majority\-vote gains\([Kim, 2026](https://arxiv.org/html/2609.21277#bib.bib12)\)\. Unlike binary\-error audits, we retain the full human label distribution as a residual reference\. Unlike audits of majority\-vote gains, we target the squared error of the panel’s label frequencies\. Applying standard PR and quadratic\-form identities to these residuals separates spectral diversity from distribution recovery; the main empirical distinction is their task\-dependent relationship with member energy\.
##### Human disagreement and annotation budgets\.
Variation among human labels need not be removable noise\([Pavlick and Kwiatkowski, 2019](https://arxiv.org/html/2609.21277#bib.bib24);[Plank, 2022](https://arxiv.org/html/2609.21277#bib.bib25)\)\. ChaosNLI provides dense item\-level labels for studying such variation\([Nie et al\., 2020](https://arxiv.org/html/2609.21277#bib.bib23)\)\. This differs from recovering a latent true label under annotator\-error models\([Dawid and Skene, 1979](https://arxiv.org/html/2609.21277#bib.bib7);[Raykar et al\., 2010](https://arxiv.org/html/2609.21277#bib.bib26)\): our target is a specified empirical distribution\. Annotation allocation also depends on the objective\. Stable ranking probabilities assess repeatability under labeling budgets\([Riley et al\., 2024](https://arxiv.org/html/2609.21277#bib.bib27)\); sufficiently large budgets can favor labeling more items over relabeling the same items under particular binary\-comparison assumptions\([Dorner and Hardt, 2026](https://arxiv.org/html/2609.21277#bib.bib9)\)\. Soft\-label learning can have different annotation saturation points for distribution matching and uncertainty ranking\([Kohli, 2026a](https://arxiv.org/html/2609.21277#bib.bib15)\)\. We hold model answers and items fixed while varying the labels used to estimate the reference\.
##### Effective size and common variation\.
Design effects express variance inflation as an effective sample size\([Kish, 1965](https://arxiv.org/html/2609.21277#bib.bib14)\), whereas participation ratio measures spectral dimension\([Mazzucato et al\., 2016](https://arxiv.org/html/2609.21277#bib.bib22)\)\. Human\-label savings in assisted evaluation\([Dorner et al\., 2025](https://arxiv.org/html/2609.21277#bib.bib10)\)and cost\-optimal allocation between judges\([Angelopoulos et al\., 2025](https://arxiv.org/html/2609.21277#bib.bib2)\)concern estimation precision and cost;νH\\nu\_\{H\}concerns residual diversity\. Common\-direction analyses include market modes in correlation matrices\([Laloux et al\., 1999](https://arxiv.org/html/2609.21277#bib.bib17)\), independent components under identification assumptions\([Comon, 1994](https://arxiv.org/html/2609.21277#bib.bib5)\), and constrained components informed by prior directions\([Lu and Rajapakse, 2005](https://arxiv.org/html/2609.21277#bib.bib20)\)\. CARE separates quality and shared confounding under a latent\-variable model\([Zhao et al\., 2026](https://arxiv.org/html/2609.21277#bib.bib33)\);[Afrin and Shihab \(2026\)](https://arxiv.org/html/2609.21277#bib.bib1)analyze common errors retained by aggregation and contamination from noisy references\. We use a fixed equal\-weight direction, without independent\-component identification, to relate shared residual variation to distribution recovery\.
## 3Human\-Referenced Panel Measurements
### 3\.1Judgments and measurement targets
Letnnbe the number of items,kkthe number of judges, andhi∈ℝLh\_\{i\}\\in\\mathbb\{R\}^\{L\}the empirical human distribution overLLlabel categories on itemii\. Judgeaareturns a categorical labell\(a,i\)l\(a,i\), represented by a one\-hot vectorel\(a,i\)e\_\{l\(a,i\)\}\. The referencehih\_\{i\}is an empirical target, not an error\-free population distribution\. We condition the measurements on this target and the fixed judge responses; the data and collection protocol are described in Section[4\.1](https://arxiv.org/html/2609.21277#S4.SS1)\.
Table[1](https://arxiv.org/html/2609.21277#S3.T1)distinguishes the quantities used throughout the audit\. The two human\-reference sizes match different panel statistics\. The variance share describes a component of residual averaging, while binary\-error effective votes supply a comparison based on agreement with one gold label\.
Table 1:Measurement targets in the panel audit\.
### 3\.2Residual representation
For judgeaa’s labell\(a,i\)l\(a,i\), letel\(a,i\)e\_\{l\(a,i\)\}be its one\-hot vector\. Its residual relative to the empirical human label distribution is
ra,i=el\(a,i\)−hi∈ℝL\.r\_\{a,i\}=e\_\{l\(a,i\)\}\-h\_\{i\}\\in\\mathbb\{R\}^\{L\}\.\(1\)Residual coordinates sum to zero\. We use an orthonormal basis of this\(L−1\)\(L\-1\)\-dimensional subspace: for the ordered labels entailment, neutral, contradiction,v1=\(1,0,−1\)⊤/2v\_\{1\}=\(1,0,\-1\)^\{\\top\}/\\sqrt\{2\}andv2=\(−1,2,−1\)⊤/6v\_\{2\}=\(\-1,2,\-1\)^\{\\top\}/\\sqrt\{6\}; for binary labels,v1=\(1,−1\)⊤/2v\_\{1\}=\(1,\-1\)^\{\\top\}/\\sqrt\{2\}\. The two three\-label coordinates represent polarity and neutrality versus the extremes\.
Letca,ic\_\{a,i\}contain the coordinatesvt⊤ra,iv\_\{t\}^\{\\top\}r\_\{a,i\}\. Averagingca,icb,i⊤c\_\{a,i\}c\_\{b,i\}^\{\\top\}over items gives a second\-moment blockBa,bB\_\{a,b\}\. We form a normalized residual Gram matrix
Ca,b=trBa,btrBa,atrBb,b\.C\_\{a,b\}=\\frac\{\\operatorname\{tr\}B\_\{a,b\}\}\{\\sqrt\{\\operatorname\{tr\}B\_\{a,a\}\\operatorname\{tr\}B\_\{b,b\}\}\}\.\(2\)This is the normalized inner product between flattened residual arrays\. We retain their across\-item means, soCCincludes average displacement from the reference\. With positive residual energy for each judge,CCis positive semidefinite, has unit diagonal, and has tracekk\. Taking block traces discards the full cross\-label block structure\.
### 3\.3Spectral effective sizeνH\\nu\_\{H\}
The participation ratio \(PR\) measures effective spectral dimension\([Mazzucato et al\., 2016](https://arxiv.org/html/2609.21277#bib.bib22)\)\. For eigenvaluesλj\\lambda\_\{j\}ofCCand the mean squared off\-diagonal inner productq¯\\bar\{q\}, its unit diagonal gives
PR=\(∑j=1kλj\)2∑j=1kλj2=k1\+\(k−1\)q¯\.\\mathrm\{PR\}=\\frac\{\(\\sum\_\{j=1\}^\{k\}\\lambda\_\{j\}\)^\{2\}\}\{\\sum\_\{j=1\}^\{k\}\\lambda\_\{j\}^\{2\}\}=\\frac\{k\}\{1\+\(k\-1\)\\bar\{q\}\}\.\(3\)PR lies in\[1,k\]\[1,k\]; it measures spectral concentration and does not retain inner\-product signs\. Fork=1k=1, we setPR=1\\mathrm\{PR\}=1\.
We express PR in human\-reference units using direct Monte Carlo sampling\([Robert and Casella, 2004](https://arxiv.org/html/2609.21277#bib.bib28)\)\. Conditional onhh, draw each ofmmsimulated labels per item independently fromCat\(hi\)\\operatorname\{Cat\}\(h\_\{i\}\)and compute the resulting panel PR\. At each grid size, the mean ofN=12N=12replicates isPR0\(m\)\\mathrm\{PR\}\_\{0\}\(m\)\. This fixed computational budget and its replicate spread are documented in Appendix[A](https://arxiv.org/html/2609.21277#A1)\. For the first adjacent grid points that bracketPRobs\\mathrm\{PR\}\_\{\\mathrm\{obs\}\}, define
νH=\\displaystyle\\nu\_\{H\}=\{\}mj\+\(mj\+1−mj\)\\displaystyle m\_\{j\}\+\(m\_\{j\+1\}\-m\_\{j\}\)\(4\)⋅PRobs−PR0\(mj\)PR0\(mj\+1\)−PR0\(mj\)\.\\displaystyle\\cdot\\frac\{\\mathrm\{PR\}\_\{\\mathrm\{obs\}\}\-\\mathrm\{PR\}\_\{0\}\(m\_\{j\}\)\}\{\\mathrm\{PR\}\_\{0\}\(m\_\{j\+1\}\)\-\\mathrm\{PR\}\_\{0\}\(m\_\{j\}\)\}\.The grid is2,…,12,16,24,32,48,64,96,1282,\\ldots,12,16,24,32,48,64,96,128\. All mean reference curves used here are strictly increasing\. Values at or below the first mean are reported asνH≤2\\nu\_\{H\}\\leq 2; those above the final mean are out of range\.
We callνH\\nu\_\{H\}the*human\-referenced effective panel size*\. It may be noninteger and differs from the nominal panel sizekk\. Within one strictly increasing reference curve, it preserves PR’s ranking: the matching provides units, not additional ranking information\. Its scale depends onhh, the item set, and the residual and simulation protocols\. At finitenn,PR0\(m\)\\mathrm\{PR\}\_\{0\}\(m\)is generally belowmm\. For fixedmm, independent item draws and a positive lower bound on average human label variance suffice forPR0\(m\)→m\\mathrm\{PR\}\_\{0\}\(m\)\\to masnngrows\. Matching one spectral statistic does not imply equal accuracy, distributional error, or annotation cost for real humans\.
### 3\.4What spectral size implies for distribution recovery
To compare residual diversity with distribution recovery, we measure the squared error of the panel’s label frequencies\. Letp¯i=k−1∑ael\(a,i\)\\bar\{p\}\_\{i\}=k^\{\-1\}\\sum\_\{a\}e\_\{l\(a,i\)\}be its label frequencies, with no majority\-vote threshold\. Define
ℰ=1n∑i‖p¯i−hi‖2\.\\mathcal\{E\}=\\frac\{1\}\{n\}\\sum\_\{i\}\\\|\\bar\{p\}\_\{i\}\-h\_\{i\}\\\|^\{2\}\.\(5\)Formmconditionally independent human\-reference labels per item, the expected error isJ/mJ/m, whereJ=n−1∑i\(1−‖hi‖2\)J=n^\{\-1\}\\sum\_\{i\}\(1\-\\\|h\_\{i\}\\\|^\{2\}\)\. Matching this target yields
νMSE=Jℰ\.\\nu\_\{\\mathrm\{MSE\}\}=\\frac\{J\}\{\\mathcal\{E\}\}\.\(6\)It is finite and positive whenJ,ℰ\>0J,\\mathcal\{E\}\>0; no finite match exists forℰ=0<J\\mathcal\{E\}=0<J, and the reference degenerates whenJ=0J=0\. Both effective sizes condition on empiricalhh, but match different objectives\.
Let⟨⋅⟩i\\langle\\cdot\\rangle\_\{i\}denote an item average,Ka,b=⟨ra,i⊤rb,i⟩iK\_\{a,b\}=\\langle r\_\{a,i\}^\{\\top\}r\_\{b,i\}\\rangle\_\{i\}the unnormalized Gram matrix,ωa=Ka,a\\omega\_\{a\}=K\_\{a,a\}the member energy, andD=diag\(ω1,…,ωk\)D=\\operatorname\{diag\}\(\\sqrt\{\\omega\_\{1\}\},\\ldots,\\sqrt\{\\omega\_\{k\}\}\)\. Then
C=D−1KD−1,ℰ=1k2𝟏⊤K𝟏\.C=D^\{\-1\}KD^\{\-1\},\\qquad\\mathcal\{E\}=\\frac\{1\}\{k^\{2\}\}\\mathbf\{1\}^\{\\top\}K\\mathbf\{1\}\.\(7\)PR uses the eigenvalues of normalizedCC, whereas error uses a quadratic form ofKK\. The following standard spectral identity states exactly which information the two summaries retain\.
##### Proposition 1 \(Error and spectral orientation\)\.
Assumeωa\>0\\omega\_\{a\}\>0for all members\. WriteC=∑j=1kλjwjwj⊤C=\\sum\_\{j=1\}^\{k\}\\lambda\_\{j\}w\_\{j\}w\_\{j\}^\{\\top\}with orthonormalwjw\_\{j\}, and letd=\(ωa\)a=1kd=\(\\sqrt\{\\omega\_\{a\}\}\)\_\{a=1\}^\{k\},Ω=d⊤d\\Omega=d^\{\\top\}d, andbj=\(wj⊤d\)2/Ωb\_\{j\}=\(w\_\{j\}^\{\\top\}d\)^\{2\}/\\Omega\. Then
ℰ=Ωk2∑j=1kλjbj,∑j=1kbj=1\.\\mathcal\{E\}=\\frac\{\\Omega\}\{k^\{2\}\}\\sum\_\{j=1\}^\{k\}\\lambda\_\{j\}b\_\{j\},\\qquad\\sum\_\{j=1\}^\{k\}b\_\{j\}=1\.\(8\)
Herebjb\_\{j\}is the squared alignment of the energy\-weighted averaging direction with eigenvectorwjw\_\{j\}\. Error depends on these weights and the energy scaleΩ\\Omega, while PR depends only on∑jλj2\\sum\_\{j\}\\lambda\_\{j\}^\{2\}\. The identity follows by substituting the eigendecomposition intod⊤Cd/k2d^\{\\top\}Cd/k^\{2\}; Appendix[C](https://arxiv.org/html/2609.21277#A3)discusses repeated eigenvalues and gives realizable equal\-spectrum panels with different errors\. Normalizing members removes their energy scales, and retaining eigenvalues alone discards orientation\.
Even equal member energies do not make PR an error ranking\. Ifωa=ω\\omega\_\{a\}=\\omega, andc¯C,vC\\bar\{c\}\_\{C\},v\_\{C\}are the mean and variance of off\-diagonalCa,bC\_\{a,b\}, then
q¯=c¯C2\+vC,ℰω=1\+\(k−1\)c¯Ck\.\\bar\{q\}=\\bar\{c\}\_\{C\}^\{\\,2\}\+v\_\{C\},\\qquad\\frac\{\\mathcal\{E\}\}\{\\omega\}=\\frac\{1\+\(k\-1\)\\bar\{c\}\_\{C\}\}\{k\}\.\(9\)The second momentq¯\\bar\{q\}controls PR, while the signed first moment controls error\. Their different dependence onvCv\_\{C\}permits conflicting changes\. A realizable construction withk=4k=4andhi=\(1/2,1/2\)h\_\{i\}=\(1/2,1/2\)has equal energies, zero mean residuals, and nonnegative correlations, yet increasing PR from22to16/716/7increasesℰ\\mathcal\{E\}from1/41/4to5/165/16\. Appendix[C](https://arxiv.org/html/2609.21277#A3)gives the labels and derivations\. Without additional assumptions, larger PR orνH\\nu\_\{H\}therefore does not guarantee smallerℰ\\mathcal\{E\}or largerνMSE\\nu\_\{\\mathrm\{MSE\}\}\. Section[4\.3](https://arxiv.org/html/2609.21277#S4.SS3)examines the empirical extent of this mismatch\.
### 3\.5Consensus\-direction variance share
To measure variation retained by equal\-weight averaging, we project onto the fixed judge\-space directionu1=𝟏/ku\_\{1\}=\\mathbf\{1\}/\\sqrt\{k\}\([Afrin and Shihab, 2026](https://arxiv.org/html/2609.21277#bib.bib1)\)\. For label coordinatett, formXi,a=\(ca,i\)tX\_\{i,a\}=\(c\_\{a,i\}\)\_\{t\}, center each column across items to obtainXcX\_\{c\}, and letS=Xc⊤Xc/nS=X\_\{c\}^\{\\top\}X\_\{c\}/n\.
Fork≥2k\\geq 2andtrS\>0\\operatorname\{tr\}S\>0, writev¯=trS/k\\bar\{v\}=\\operatorname\{tr\}S/kfor mean judge variance andc¯\\bar\{c\}for mean off\-diagonal covariance\. Their ratio is
ρ¯=c¯v¯=∑a≠bSa,b\(k−1\)trS,v¯=trSk\.\\bar\{\\rho\}=\\frac\{\\bar\{c\}\}\{\\bar\{v\}\}=\\frac\{\\sum\_\{a\\neq b\}S\_\{a,b\}\}\{\(k\-1\)\\operatorname\{tr\}S\},\\qquad\\bar\{v\}=\\frac\{\\operatorname\{tr\}S\}\{k\}\.\(10\)This normalized mean covariance generally differs from the mean of pairwise Pearson correlations\. The consensus\-direction variance share is
γco=u1⊤Su1trS=1\+\(k−1\)ρ¯k\.\\gamma\_\{\\mathrm\{co\}\}=\\frac\{u\_\{1\}^\{\\top\}Su\_\{1\}\}\{\\operatorname\{tr\}S\}=\\frac\{1\+\(k\-1\)\\bar\{\\rho\}\}\{k\}\.\(11\)It lies in\[0,1\]\[0,1\]and equals1/k1/kwhen mean covariance is zero, linking it to the design\-effect form\([Kish, 1965](https://arxiv.org/html/2609.21277#bib.bib14)\)\. It measures centered shared fluctuations, not statistical bias or the across\-item mean residual\. Across label coordinates, we use total\-variance weighting:γco,all=∑tu1⊤Stu1/∑ttrSt\\gamma\_\{\\mathrm\{co,all\}\}=\\sum\_\{t\}u\_\{1\}^\{\\top\}S\_\{t\}u\_\{1\}/\\sum\_\{t\}\\operatorname\{tr\}S\_\{t\}, with positive total trace\.
Only the consensus component survives the full equal\-weight average; components orthogonal to𝟏\\mathbf\{1\}cancel\. For a uniformly selected subsetIIofmmjudges from this fixed pool, letx¯I=m−1∑a∈I\(Xc\):,a\\bar\{x\}\_\{I\}=m^\{\-1\}\\sum\_\{a\\in I\}\(X\_\{c\}\)\_\{:,a\}\. Using empirical variance with denominatornn,
Vm\\displaystyle V\_\{m\}=𝔼I:\|I\|=m\[Varn\(x¯I\)\]\\displaystyle=\\mathbb\{E\}\_\{I:\|I\|=m\}\[\\operatorname\{Var\}\_\{n\}\(\\bar\{x\}\_\{I\}\)\]\(12\)=c¯\+v¯−c¯m,1≤m≤k\.\\displaystyle=\\bar\{c\}\+\\frac\{\\bar\{v\}\-\\bar\{c\}\}\{m\},\\qquad 1\\leq m\\leq k\.ThusVm/v¯=1/m\+\(1−1/m\)ρ¯V\_\{m\}/\\bar\{v\}=1/m\+\(1\-1/m\)\\bar\{\\rho\}\. Substituting the full panel’s observedρ¯\\bar\{\\rho\}gives an analytic curve for averaging within the fixed pool\.
The consensus share connects centered variation to the distributional error defined in Section[3\.4](https://arxiv.org/html/2609.21277#S3.SS4)\. Forμa=⟨ra,i⟩i\\mu\_\{a\}=\\langle r\_\{a,i\}\\rangle\_\{i\}andμ¯=k−1∑aμa\\bar\{\\mu\}=k^\{\-1\}\\sum\_\{a\}\\mu\_\{a\}, the raw\-scale channel covariances give
ℰ=‖μ¯‖2\+∑ttrStkγco,all\.\\mathcal\{E\}=\\\|\\bar\{\\mu\}\\\|^\{2\}\+\\frac\{\\sum\_\{t\}\\operatorname\{tr\}S\_\{t\}\}\{k\}\\gamma\_\{\\mathrm\{co,all\}\}\.\(13\)The two terms are the squared mean residual and the variation of the aggregated residual across items\. The shareγco,all\\gamma\_\{\\mathrm\{co,all\}\}must therefore be interpreted together with total variance and mean displacement\([Afrin and Shihab, 2026](https://arxiv.org/html/2609.21277#bib.bib1)\)\.
## 4Experiments
### 4\.1Data and audit protocol
We use MNLI\-m, SNLI, andα\\alphaNLI from ChaosNLI\([Nie et al\., 2020](https://arxiv.org/html/2609.21277#bib.bib23)\)\. MNLI\-m and SNLI distinguish entailment, neutral, and contradiction\([Williams et al\., 2018](https://arxiv.org/html/2609.21277#bib.bib32);[Bowman et al\., 2015](https://arxiv.org/html/2609.21277#bib.bib4)\);α\\alphaNLI chooses between two abductive hypotheses\([Bhagavatula et al\., 2020](https://arxiv.org/html/2609.21277#bib.bib3)\)\. Each selected item has 100 human labels\. We sample 1,000 items per dataset approximately equally across human\-entropy terciles, with seed 42\. Task agreement uses the dataset\-provided 100\-label modegoldi\\mathrm\{gold\}\_\{i\}, preserving its original tie choices\.
The panel contains 32 models from 10 provider families \(Appendix[A](https://arxiv.org/html/2609.21277#A1), Table[A1](https://arxiv.org/html/2609.21277#A1.T1)\)\. Each model responds separately at temperature 0, with one user message, no system message, and a label\-only instruction\. Requested model identifiers are recorded in the supplement; they do not independently authenticate the service’s underlying provider version\. Table[2](https://arxiv.org/html/2609.21277#S4.T2)distinguishes the response sets and the analyses performed on them\.
Table 2:Audit protocol\. Task\-specific counts follow MNLI\-m, SNLI, andα\\alphaNLI order\.The baseline results use the retained matrices; failure indicators distinguish placeholders from model judgments\. Appendix[D\.5](https://arxiv.org/html/2609.21277#A4.SS5.SSS0.Px1)reports the saved sensitivity results after excluding affected items\. The label, tie, filtering, and reference\-draw rules are in Appendix[A](https://arxiv.org/html/2609.21277#A1); the supplementary reproduction map links each result to its available records and saved outputs\.
### 4\.2Effective sizes and reference choices
For the fixed item sets and empirical human distributions, the 32\-judge panels haveνH=4\.24\\nu\_\{H\}=4\.24,6\.466\.46, and6\.506\.50\(Table[3](https://arxiv.org/html/2609.21277#S4.T3)\), only 13\.3%–20\.3% of their nominal size\. Retaining human disagreement therefore does not remove the large gap between the number of judgments and their effective residual dimension\.
Table 3:Different measurement targets for the same 32\-judge panels\.The distribution\-error matchesνMSE=2\.30\\nu\_\{\\mathrm\{MSE\}\}=2\.30,3\.753\.75, and3\.443\.44are lower thanνH\\nu\_\{H\}\. The spectral match is 1\.84, 1\.72, and 1\.89 times the error match, respectively\. These ratios describe the difference between targets in the full panels; they are not calibration factors transferable to other panels\. Table[3](https://arxiv.org/html/2609.21277#S4.T3)also reports the errors and consensus shares\.
Binary\-error effective votesneffn\_\{\\mathrm\{eff\}\}are approximately 2\([Kish, 1965](https://arxiv.org/html/2609.21277#bib.bib14);[Kohli, 2026b](https://arxiv.org/html/2609.21277#bib.bib16)\)\. The difference fromνH\\nu\_\{H\}involves both representation, including encoding, reference, and centering, and summarization\. Keeping the binary\-error Pearson matrix but replacing the signed\-correlation summary by PR gives 3\.54, 4\.38, and 3\.79\. Using human\-distribution residuals then gives PR of 4\.23, 6\.42, and 6\.43\. The conversion from this PR toνH\\nu\_\{H\}changes the values by only 0\.3%–1\.1%; the complete representation and summarization account for most of the difference from binary effective votes\. This comparison does not isolate human disagreement alone; Table[4](https://arxiv.org/html/2609.21277#S4.T4)varies the anchor within the same uncentered residual construction\. Appendix[C](https://arxiv.org/html/2609.21277#A3)gives the crossed comparison\.
We next keep the answers and spectral summary fixed while changing the residual anchorAA: uniform labels, a human\-mode vectorAmodeA\_\{\\mathrm\{mode\}\}, or the full distributionhh\. Ties inAmodeA\_\{\\mathrm\{mode\}\}take the first label in the order of Section[3\.2](https://arxiv.org/html/2609.21277#S3.SS2); this can differ from the suppliedgoldi\\mathrm\{gold\}\_\{i\}\. Simulated labels always come fromhh, while both model and simulated residuals subtract the sameAA, with a separate reference curve for each anchor\.
Table 4:Reference\-anchor comparison for the same 32 model judges\.The full distribution yields lowerq¯\\bar\{q\}and higherνH\\nu\_\{H\}on every task \(Table[4](https://arxiv.org/html/2609.21277#S4.T4)\)\. The mode and uniform anchors give similar effective sizes on MNLI\-m; using the mode already raises the value on SNLI andα\\alphaNLI\. Retaining the full distribution raises it further on all three\. Human disagreement changes the residual dependence visible in the same model answers\.
### 4\.3Spectral diversity versus distribution recovery
Section[3\.4](https://arxiv.org/html/2609.21277#S3.SS4)rules out a general monotone relationship betweenνH\\nu\_\{H\}and distribution recovery\. We now ask how closely the objectives agree in the observed model pool, and whether a single member addition can move them in opposite directions\.
#### 4\.3\.1Within\-size rankings and member energy
For eachk=2,…,31k=2,\\ldots,31, we deduplicate the 2,000 sampled panels by member set and compute the Spearman correlation of PR with−ℰ\-\\mathcal\{E\}\. The minus sign makes larger values favorable for both quantities\. Within a strictly increasing reference curve’s matched range, these rankings also correspond to those ofνH\\nu\_\{H\}andνMSE\\nu\_\{\\mathrm\{MSE\}\}\. Working with PR retains small panels below the human\-matching lower limit\.
Because normalization removes member energy, we also compare PR with−ℰ/ω¯\-\\mathcal\{E\}/\\bar\{\\omega\}, whereω¯=k−1∑aωa\\bar\{\\omega\}=k^\{\-1\}\\sum\_\{a\}\\omega\_\{a\}is mean member residual energy\. This measures error relative to the members’ average squared deviation\.
Figure 1:Agreement between spectral and distribution\-recovery rankings\. Each point is a within\-size Spearman correlation, denotedcorrS\\operatorname\{corr\}\_\{S\}, over distinct member sets after deduplicating 2,000 draws per size from the fixed pool\. Appendix[D](https://arxiv.org/html/2609.21277#A4)reports item\-half stability\.At sizesk=4,8,16,24,31k=4,8,16,24,31, the PR–−ℰ\-\\mathcal\{E\}correlation ranges from 0\.960 to 0\.977 on MNLI\-m and from 0\.711 to 0\.867 on SNLI, but only from 0\.223 to 0\.419 onα\\alphaNLI \(Figure[1](https://arxiv.org/html/2609.21277#S4.F1)\)\. Within this fixed pool, spectral and distributional rankings are closely aligned on MNLI\-m, less closely on SNLI, and substantially less closely onα\\alphaNLI\.
Scaling error by member energy raises theα\\alphaNLI correlations to 0\.937–0\.960; atk=16k=16, the increase is from 0\.338 to 0\.937\. The ranking mismatch is closely associated with member error energy: a panel can offer diverse residual directions while its members deviate farther from the human distribution\. This motivates reporting member error energy together with diversity\([Kim, 2026](https://arxiv.org/html/2609.21277#bib.bib12)\)\. The scaled error is a different target, however, not a causal control for member capability\.
#### 4\.3\.2Member additions and item stability
For each unique panel atk=4,8,16,24,31k=4,8,16,24,31, we add every model outside the panel\. Define spectral gainΔPR=PRnew−PRold\\Delta\\mathrm\{PR\}=\\mathrm\{PR\}\_\{\\mathrm\{new\}\}\-\\mathrm\{PR\}\_\{\\mathrm\{old\}\}and error improvementΔℰ=ℰold−ℰnew\\Delta\\mathcal\{E\}=\\mathcal\{E\}\_\{\\mathrm\{old\}\}\-\\mathcal\{E\}\_\{\\mathrm\{new\}\}\. Opposite signs mean the objectives conflict\.
We repeat the calculation on two fixed 500\-item halves\. Table[5](https://arxiv.org/html/2609.21277#S4.T5)counts additions with the same direction of conflict on the full set and both halves, and with both relative changes at least 1% on each set\. The relative changes use the original panel’s values:\|ΔPR\|/PRold\|\\Delta\\mathrm\{PR\}\|/\\mathrm\{PR\}\_\{\\mathrm\{old\}\}and\|Δℰ\|/ℰold\|\\Delta\\mathcal\{E\}\|/\\mathcal\{E\}\_\{\\mathrm\{old\}\}\. We use a fixed descriptive 1% filter; the split assesses stability within the observed items\.
Table 5:Conflicting additions stable across the full item set and both halves, with both relative changes at least 1%\. Parentheses give within\-size percentages\.For4→54\\to 5, 3\.89%, 0\.57%, and 2\.07% of additions meet this criterion\. A member can improve spectral diversity while worsening distribution recovery, or the reverse\. The frequencies generally decline with size; none of the31→3231\\to 32additions passes the filter\. Comparisons share members and items and are not independent experimental replicates\.
Panel rankings themselves depend on item composition\. The correlation of PR rankings between the two halves is only 0\.370–0\.554 onα\\alphaNLI, compared with 0\.915–0\.942 on MNLI\-m\. Appendix[D](https://arxiv.org/html/2609.21277#A4)gives the full curves and protocol\. For this audit, the within\-size associations and stable conflicts answer complementary questions: a high overall correlation can coexist with individual reversals, while a stable reversal does not establish that panel rankings transfer to new items\. Distributional error directly evaluates the recovery target; spectral diversity describes a different property of the same responses\.
### 4\.4Geometry of model and constructed residuals
To visualize how model residuals differ from human sampling, Figure[2](https://arxiv.org/html/2609.21277#S4.F2)compares their direction and shape on the two three\-label tasks\. For each judge, anisotropyAaA\_\{a\}is the difference between its two residual second\-moment eigenvalues divided by their sum; larger values mean greater concentration along one axis\. The angleΔθa\\Delta\\theta\_\{a\}measures rotation of that principal axis from the analytic human reference\. Simulated labels drawn itemwise fromhhgive the comparison cloud\. Appendix[B](https://arxiv.org/html/2609.21277#A2)defines these coordinates and cloud distances\.
Figure 2:Residual geometry relative to human sampling\. Black points are 32 model judges; crosses mark simulated cloud centers\. Ellipses markd=1,2d=1,2, not confidence regions\. Dashed lines show analytic human\-reference anisotropyA∗A\_\{\*\}; dotted lines show the simulated 99th percentile\. Both tasks use the same axis limits\. Stars and diamonds denote the human\-mode judgeJgoldJ\_\{\\mathrm\{gold\}\}and panel\-majority judgeJmajJ\_\{\\mathrm\{maj\}\}\.Most model points lie above and to the left of the simulated cloud center: 28 of 32 on MNLI\-m and 23 of 32 on SNLI\. Upward displacement means greater anisotropy; leftward displacement means a negative principal\-axis rotation\. The constructed judgeJgoldJ\_\{\\mathrm\{gold\}\}, which returnsgoldi\\mathrm\{gold\}\_\{i\}without model inference, also lies above the cloud\. Its anisotropy excess is 76\.9% and 99\.1% of the mean model excess\. Both this judge and the simulations output one\-hot labels, but one deterministically chooses the human mode while the other samples fromhh\. Thus mode selection can produce similar upward displacement; one\-hot encoding alone does not explain the contrast\.
Aggregating all 32 judges intoJmajJ\_\{\\mathrm\{maj\}\}shifts the point left relative toJgoldJ\_\{\\mathrm\{gold\}\}by approximately2\.10∘2\.10^\{\\circ\}and3\.09∘3\.09^\{\\circ\}\. HereJmajJ\_\{\\mathrm\{maj\}\}returns the panel’s modal label, resolving ties in canonical label order\. Majority vote retains a directional deviation shared by the individual model points\. This descriptive geometry does not identify the mechanism generating the deviation\. We next examine the distinct judge\-space direction whose variation survives equal\-weight averaging\.
### 4\.5Shared variation retained by averaging
Figure[3](https://arxiv.org/html/2609.21277#S4.F3)illustrates the alignment in a separate 10\-model panel containing one member per family\. Its elongated item clouds lie close to the projected consensus direction\. The plot shows only two principal components; the full consensus vector need not lie in that plane\. This illustrative projection is not an estimate of the full panel’s variance share\.
Figure 3:Illustrative 10\-model projection, distinct from the 32\-model estimates in Table[3](https://arxiv.org/html/2609.21277#S4.T3)\. Item residuals from one judge per family are shown in the first two principal components\. Red is the projected consensus direction; orange is orthogonal within the plane\. PC1 parentheses give the red line’s angle to PC1\. This projection does not estimate the full panel’sγco\\gamma\_\{\\mathrm\{co\}\}\.Using all 32 judges, the two channel\-specific consensus\-direction variance sharesγco\\gamma\_\{\\mathrm\{co\}\}are 44\.8% and 43\.4% on MNLI\-m, and 36\.0% and 32\.8% on SNLI\. Total\-variance weighting gives 43\.8% and 33\.7%; the singleα\\alphaNLI channel gives 35\.9% \(Table[3](https://arxiv.org/html/2609.21277#S4.T3)\)\. These shares quantify centered residual variation retained along the fixed equal\-weight direction\.
Figure 4:Expected variance of a random\-subpanel mean relative to mean individual variance\. Gray shows the zero\-covariance reference\.Equations \([11](https://arxiv.org/html/2609.21277#S3.E11)\)–\([12](https://arxiv.org/html/2609.21277#S3.E12)\) show how this common variation survives averaging and produces diminishing variance reduction \(Figure[4](https://arxiv.org/html/2609.21277#S4.F4)\)\. Decomposing full\-panel distributional error with Equation \([13](https://arxiv.org/html/2609.21277#S3.E13)\), the squared mean residual‖μ¯‖2\\\|\\bar\{\\mu\}\\\|^\{2\}accounts for 11\.3%, 4\.8%, and 0\.1% ofℰ\\mathcal\{E\}\. Most error thus comes from aggregated residuals that vary across items, rather than a fixed across\-item displacement\. This shared variation is not by itself evidence of a particular preference mechanism\.
### 4\.6Panel size and scope of expansion
The medianνH\\nu\_\{H\}of random subpanels increases with nominal size, with diminishing gains \(Figure[5](https://arxiv.org/html/2609.21277#S4.F5)\)\. These medians summarize sampled panels; they do not assert monotonicity for every member addition\.
Figure 5:Random\-subpanelνH\\nu\_\{H\}: medians and task\-colored 5th–95th percentile bands\. Open markers denote family\-diverse or full panels from the same model pool, not independent test results; the gray line isνH=k\\nu\_\{H\}=k\.The searched candidates and provider\-family comparisons yield no general expansion rule for this pool \(Appendix[D](https://arxiv.org/html/2609.21277#A4)\)\. In particular, the subpanel search selects the size on the second item half and so does not measure held\-out selection gains\. These results delimit the audit rather than validate a panel\-selection algorithm\. We also examine whether changing the presentation of the same options produces the association reduction expected from an independent\-transition reference\.
### 4\.7Presentation changes and residual association
We replace one judge at a time by its option\-reordered responses, keeping all other responses fixed\. DefineΔq¯=q¯alt−q¯base\\Delta\\bar\{q\}=\\bar\{q\}\_\{\\mathrm\{alt\}\}\-\\bar\{q\}\_\{\\mathrm\{base\}\}, so negative values indicate reduced squared residual association\. An independent\-transition reference samples each item’s replacement from that judge’s empirical transition probabilities conditional on its original label\. This reference includes unchanged labels and matches conditional flip rates and destination frequencies in expectation\.
Missing variants leave31\+31\+29=9131\+31\+29=91model–dataset combinations, using 489, 492, and 492 common complete\-case items\. The comparisons use the presentation experiment’s own stored original\-response panel, with cache reuse in the response records\. They cannot isolate presentation\-order effects from repeat\-call variation\. Appendix[A](https://arxiv.org/html/2609.21277#A1)gives the full response and simulation protocols\.
Figure 6:Observed presentation changes versus an independent\-transition reference\. \(a\) Each point is one model–dataset combination; horizontal bars span the 2\.5th–97\.5th percentiles of 200 reference simulations; both axes use10−310^\{\-3\}units\. \(b\) Absolute mean reference change divided by absolute observed change, on a logarithmic axis\. Cached responses and repeat\-call variation prevent interpreting this as an isolated order effect\.Under the independent\-transition reference, points would fluctuate aroundy=xy=xin Figure[6](https://arxiv.org/html/2609.21277#S4.F6)a\. Instead, 86 observed changes exceed their corresponding reference’s 97\.5th percentile; the other 5 fall within its interval\. These are descriptive simulation comparisons without multiplicity correction\. The reordered responses retain more residual association than the matched reference\. Its mean change is negative for every combination, whereas the observed change is positive in 37 cases\. The median absolute\-change ratios are 6\.57, 4\.91, and 2\.22 \(Figure[6](https://arxiv.org/html/2609.21277#S4.F6)b\)\. The observed response changes therefore differ from independent transitions with the same conditional flip rates\. Additional flip and subpanel diagnostics appear in Appendix[D](https://arxiv.org/html/2609.21277#A4)\.
### 4\.8Sensitivity to the human reference
Keeping all 32 model judges fixed, we estimate an anchorAMA\_\{M\}by drawingMMlabels without replacement from each item’s 100\-label count pool\. Model and simulated residuals subtractAMA\_\{M\}, while simulated labels still come from the fullhh\. This changes reference estimation while holding model answers fixed\.
MedianνH\\nu\_\{H\}is 3\.35, 4\.36, and 4\.32 atM=5M=5, and 3\.70, 5\.10, and 5\.12 atM=10M=10\. It approaches the full\-reference values 4\.24, 6\.46, and 6\.50 asMMincreases; distributions at adjacent largerMMoverlap substantially\. Appendix[A](https://arxiv.org/html/2609.21277#A1), Figure[A1](https://arxiv.org/html/2609.21277#A1.F1), gives the complete sensitivity curves\.
Anchor error enters every model residual as the same vector,ra,iA=ra,ih\+\(hi−AM,i\)r\_\{a,i\}^\{A\}=r\_\{a,i\}^\{h\}\+\(h\_\{i\}\-A\_\{M,i\}\)\. This shared displacement arises from estimating the reference, not from changing the model\. Appendix[A](https://arxiv.org/html/2609.21277#A1)gives its expected squared magnitude\.
### 4\.9Task agreement as a separate target
Panel majority vote agrees with human gold on 68\.6%, 86\.5%, and 93\.0% of items, exceeding the best individual model only on SNLI\. Agreement and binary\-error clustering are reported in Appendix[D](https://arxiv.org/html/2609.21277#A4)\. They use a thresholded label decision, whereasℰ\\mathcal\{E\}uses the unthresholded panel frequencies; the measurements in Table[3](https://arxiv.org/html/2609.21277#S4.T3)therefore do not predict majority\-vote accuracy\.
## 5Conclusion
A judge panel has no single human\-equivalent count across measurement targets\. In the same three 32\-model panels, matching residual spectral diversity givesνH=4\.24\\nu\_\{H\}=4\.24,6\.466\.46, and6\.506\.50, whereas matching distributional error givesνMSE=2\.30\\nu\_\{\\mathrm\{MSE\}\}=2\.30,3\.753\.75, and3\.443\.44\. Human disagreement changes the dependence visible in model answers, and the choice of summary changes what effective size means\.
Distribution recovery depends on member energy and the orientation of residual variation relative to averaging\. The exact spectral identity and realizable counterexamples explain why PR alone cannot order that error\. Empirically, spectral and recovery rankings agree strongly on MNLI\-m but much less onα\\alphaNLI; some additions conflict on both item halves, even though the overallα\\alphaNLI rankings are unstable\. The audit therefore pairs each reported size with its target and separates the descriptive properties of fixed responses from claims about better aggregation or human labor savings\.
### Limitations
Our evidence covers three categorical language\-inference datasets and one fixed model pool\. The empirical 100\-label distributions have sampling error\. In the anchor\-size experiment, simulations still use the full distributions, so the experiment does not validate deployment with only a few human labels\.
NeitherνH\\nu\_\{H\}norνMSE\\nu\_\{\\mathrm\{MSE\}\}is a general human\-replacement rate\. The study audits existing panels; it does not benchmark a new selection or aggregation algorithm\. The spectral identity explains a measurement distinction and does not supply a new estimator or an improvement guarantee\.
The 12 simulation replicates and finite grid quantify neither item\-sampling uncertainty nor reference\-estimation and member\-selection uncertainty\. Panels and additions overlap\. Fixed\-half diagnostics measure stability in already observed data; the 1% filter is not a validated utility threshold\. The weak cross\-half ranking onα\\alphaNLI further limits selection claims\. Finally, consensus projections describe equal\-weight residual averaging, not majority\-vote decision error, and the present model pool provides no bound on the value of future models\.
### Data Availability
The 32\-judge vote archive is hosted at[https://github\.com/Chao1208/chaosnli\-judge\-votes](https://github.com/Chao1208/chaosnli-judge-votes)\(snapshot5da92bb\)\. The accompanying[vote supplement](https://supplement/votes/README.md)contains the 96,000 baseline and 122,000 presentation\-order records used here, including cache reuse, with requested model identifiers and failure indicators\. The repository also retains presentation records excluded under its documented rules\. We distribute model judgments and item identifiers, not task text or human counts\. Users obtain ChaosNLI under its original terms\([Nie et al\., 2020](https://arxiv.org/html/2609.21277#bib.bib23)\)and join by dataset and item identifier\.
The accompanying[analysis protocol](https://supplement/analysis_protocol.json),[presentation item lists](https://supplement/order_item_ids.json), and[distribution\-error protocol](https://supplement/loss_diagnostics/protocol.json)specify member sets, item selection, splits, and tie rules\. Per\-call timestamps were not retained; available request metadata and its limits are documented in the supplement\. The accompanying[analysis code](https://supplement/analysis_code/ANALYSIS.md)computes the full\-panel measurements from the archived votes and external human counts, using saved calibration curves forνH\\nu\_\{H\}\. The code distinguishes the recorded matrices used in the paper from an analysis excluding items with failed responses\. This code supplement is separate from repository snapshot5da92bb\. A supplementary[reproduction map](https://supplement/reproduction_map.md)identifies available records, result fields, and missing inputs, with ordered calculation steps for the core tables\. The package does not provide executable reproduction of every reported result or the complete random\-number streams\.
## References
- Afrin and Shihab \(2026\)F\. Afrin and I\. F\. Shihab\. 2026\.[Evaluator ensembles under reward hacking: Covariance geometry and finite\-search guarantees](https://arxiv.org/abs/2608.08002v1)\.ArXiv:2608\.08002v1\.
- Angelopoulos et al\. \(2025\)A\. N\. Angelopoulos, J\. Eisenstein, J\. Berant, A\. Agarwal, and A\. Fisch\. 2025\.[Cost\-optimal active AI model evaluation](https://arxiv.org/abs/2506.07949v1)\.ArXiv:2506\.07949v1\.
- Bhagavatula et al\. \(2020\)C\. Bhagavatula, R\. Le Bras, C\. Malaviya, K\. Sakaguchi, A\. Holtzman, H\. Rashkin, D\. Downey, S\. W\.\-t\. Yih, and Y\. Choi\. 2020\.Abductive commonsense reasoning\.In International Conference on Learning Representations \(ICLR\)\.
- Bowman et al\. \(2015\)S\. R\. Bowman, G\. Angeli, C\. Potts, and C\. D\. Manning\. 2015\.A large annotated corpus for learning natural language inference\.In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing \(EMNLP\), 632–642\.
- Comon \(1994\)P\. Comon\. 1994\.Independent component analysis, A new concept?Signal Processing, 36\(3\), 287–314\.
- Condorcet \(1785\)M\. J\. A\. N\. de Caritat, marquis de Condorcet\. 1785\.Essai sur l’application de l’analyse à la probabilité des décisions rendues à la pluralité des voix\.Paris: Imprimerie Royale\.
- Dawid and Skene \(1979\)A\. P\. Dawid and A\. M\. Skene\. 1979\.Maximum likelihood estimation of observer error\-rates using the EM algorithm\.Journal of the Royal Statistical Society, Series C \(Applied Statistics\), 28\(1\), 20–28\.
- Dietterich \(2000\)T\. G\. Dietterich\. 2000\.Ensemble methods in machine learning\.In Multiple Classifier Systems \(MCS 2000\), Lecture Notes in Computer Science, 1857, 1–15\. Springer\.
- Dorner and Hardt \(2026\)F\. E\. Dorner and M\. Hardt\. 2026\.[Don’t label twice: Quantity beats quality when comparing binary classifiers on a budget](https://arxiv.org/abs/2402.02249v3)\.ArXiv:2402\.02249v3\.
- Dorner et al\. \(2025\)F\. E\. Dorner, V\. Y\. Nastl, and M\. Hardt\. 2025\.Limits to scalable evaluation at the frontier: LLM as judge won’t beat twice the data\.In International Conference on Learning Representations \(ICLR\)\.
- Jiang et al\. \(2025\)L\. Jiang, Y\. Chai, M\. Li, M\. Liu, R\. Fok, N\. Dziri, Y\. Tsvetkov, M\. Sap, A\. Albalak, and Y\. Choi\. 2025\.[Artificial hivemind: The open\-ended homogeneity of language models \(and beyond\)](https://arxiv.org/abs/2510.22954v1)\.In Advances in Neural Information Processing Systems \(NeurIPS 2025\)\. arXiv:2510\.22954v1\.
- Kim \(2026\)Donghwan Kim\. 2026\.[Are Diversity Metrics Measuring Diversity? A Capability\-Controlled Audit of Majority\-Vote Gain in LLM Ensembles](https://arxiv.org/abs/2607.20768v1)\.ArXiv:2607\.20768v1\.
- Kim et al\. \(2025\)E\. M\. Kim, A\. Garg, K\. Peng, and N\. Garg\. 2025\.Correlated errors in large language models\.In Proceedings of the 42nd International Conference on Machine Learning, PMLR, 267, 30038–30066\.
- Kish \(1965\)L\. Kish\. 1965\.Survey Sampling\.New York: John Wiley & Sons\.
- Kohli \(2026a\)G\. Kohli\. 2026a\.[Metric\-dependent annotation saturation for learning from label distributions](https://arxiv.org/abs/2605.29797v1)\.ArXiv:2605\.29797v1\.
- Kohli \(2026b\)G\. Kohli\. 2026b\.[Nine judges, two effective votes: Correlated errors undermine LLM evaluation panels](https://arxiv.org/abs/2605.29800v1)\.ArXiv:2605\.29800v1\.
- Laloux et al\. \(1999\)L\. Laloux, P\. Cizeau, J\.\-P\. Bouchaud, and M\. Potters\. 1999\.Noise dressing of financial correlation matrices\.Physical Review Letters, 83\(7\), 1467–1470\.
- Lipsitch et al\. \(2010\)M\. Lipsitch, E\. Tchetgen Tchetgen, and T\. Cohen\. 2010\.Negative controls: a tool for detecting confounding and bias in observational studies\.Epidemiology, 21\(3\), 383–388\.
- Liu \(2026\)N\. Liu\. 2026\.[LLMs as a jury: Cross\-model consensus can outperform process reward models for LLM reasoning](https://arxiv.org/abs/2607.10139v3)\.ArXiv:2607\.10139v3\.
- Lu and Rajapakse \(2005\)W\. Lu and J\. C\. Rajapakse\. 2005\.Approach and applications of constrained ICA\.IEEE Transactions on Neural Networks, 16\(1\), 203–212\.
- Mardia and Jupp \(2000\)K\. V\. Mardia and P\. E\. Jupp\. 2000\.Directional Statistics\.Chichester: John Wiley & Sons\.
- Mazzucato et al\. \(2016\)L\. Mazzucato, A\. Fontanini, and G\. La Camera\. 2016\.Stimuli reduce the dimensionality of cortical activity\.Frontiers in Systems Neuroscience, 10, 11\.
- Nie et al\. \(2020\)Y\. Nie, X\. Zhou, and M\. Bansal\. 2020\.What can we learn from collective human opinions on natural language inference data?In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\), 9131–9143\.
- Pavlick and Kwiatkowski \(2019\)E\. Pavlick and T\. Kwiatkowski\. 2019\.Inherent disagreements in human textual inferences\.Transactions of the Association for Computational Linguistics, 7, 677–694\.
- Plank \(2022\)B\. Plank\. 2022\.The “problem” of human label variation: On ground truth in data, modeling and evaluation\.In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing \(EMNLP\), 10671–10682\.
- Raykar et al\. \(2010\)V\. C\. Raykar, S\. Yu, L\. H\. Zhao, G\. Hermosillo Valadez, C\. Florin, L\. Bogoni, and L\. Moy\. 2010\.Learning from crowds\.Journal of Machine Learning Research, 11, 1297–1322\.
- Riley et al\. \(2024\)P\. Riley, D\. Deutsch, G\. Foster, V\. Ratnakar, A\. Dabirmoghaddam, and M\. Freitag\. 2024\.[Finding replicable human evaluations via stable ranking probability](https://arxiv.org/abs/2404.01474v1)\.ArXiv:2404\.01474v1\.
- Robert and Casella \(2004\)Christian P\. Robert and George Casella\. 2004\.[*Monte Carlo Statistical Methods*](https://doi.org/10.1007/978-1-4757-4145-2), 2 edition\.Springer, New York\.
- Thakur et al\. \(2025\)A\. S\. Thakur, K\. Choudhary, V\. S\. Ramayapally, S\. Vaidyanathan, and D\. Hupkes\. 2025\.Judging the judges: Evaluating alignment and vulnerabilities in LLMs\-as\-judges\.In Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics \(GEM²\), 404–430\. Association for Computational Linguistics\.
- Turkmen et al\. \(2026\)Y\. Turkmen, B\. Buyukates, and M\. Bastopcu\. 2026\.[Don’t always pick the highest\-performing model: An information theoretic view of LLM ensemble selection](https://arxiv.org/abs/2602.08003v1)\.ArXiv:2602\.08003v1\.
- Verga et al\. \(2024\)P\. Verga, S\. Hofstätter, S\. Althammer, Y\. Su, A\. Piktus, A\. Arkhangorodsky, M\. Xu, N\. White, and P\. Lewis\. 2024\.[Replacing judges with juries: Evaluating LLM generations with a panel of diverse models](https://arxiv.org/abs/2404.18796v2)\.ArXiv:2404\.18796v2\.
- Williams et al\. \(2018\)A\. Williams, N\. Nangia, and S\. R\. Bowman\. 2018\.A broad\-coverage challenge corpus for sentence understanding through inference\.In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL\-HLT\), Volume 1 \(Long Papers\), 1112–1122\.
- Zhao et al\. \(2026\)J\. Zhao, C\. Shin, T\.\-H\. Huang, S\. S\. S\. Namburi GNVV, and F\. Sala\. 2026\.[CARE: Confounder\-aware aggregation for reliable LLM evaluation](https://arxiv.org/abs/2603.00039v1)\.ArXiv:2603\.00039v1\.
- Zheng et al\. \(2023\)L\. Zheng, W\.\-L\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica\. 2023\.Judging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena\.In Advances in Neural Information Processing Systems 36 \(NeurIPS 2023\), Datasets and Benchmarks Track, 46595–46623\.
## Appendix ACollection and Analysis Protocols
### A\.1Model panel
Table[A1](https://arxiv.org/html/2609.21277#A1.T1)lists the fixed model pool\. Requested service identifiers and collection limits are documented in the vote supplement\.
Table A1:The fixed panel of 32 model judges\.
### A\.2Prompts and response sets
All tasks use temperature 0, one user message, and no system message\. MNLI\-m and SNLI share the following prompt:
Given the following premise and hypothesis, determine the relationship between them\.Premise: \{premise\}Hypothesis: \{hypothesis\}What is the relationship? Reply with ONLY one word: "entailment", "neutral", or "contradiction"\.Theα\\alphaNLI prompt is:
Given two observations and two hypotheses, determine which hypothesis better explains what happened between the two observations\.Observation 1: \{obs1\}Observation 2: \{obs2\}Hypothesis 1: \{hyp1\}Hypothesis 2: \{hyp2\}Which hypothesis is more plausible? Reply with ONLY one word: "1" or "2"\.Presentation variants reorder the three label names or exchange the two abductive hypotheses\. Responses are mapped back to canonical labels\. Parsing accepts the legal labels and specified aliases; unparseable responses are flagged\. Baseline placeholders are retained as stated in Section[4\.1](https://arxiv.org/html/2609.21277#S4.SS1); presentation comparisons require parseable paired responses\.
The initial 500 presentation items are the first 500 rows of the stored entropy\-stratified sample, ordered by original dataset row; there is no second random item sample\. MNLI\-m and SNLI require all 31 retained models to be parseable in v0, v1, and v2;α\\alphaNLI requires all 29 in v0 and v1\. This leaves 489, 492, and 492 items\. Baseline and presentation analyses use separately stored response snapshots\. Presentation analyses use their own v0 for every member of the original panel, not the full baseline matrix\. Figure[6](https://arxiv.org/html/2609.21277#S4.F6)compares v0 with v1 on the common all\-variant item set\. Model and item lists accompany the[protocol](https://supplement/analysis_protocol.json)and[item manifest](https://supplement/order_item_ids.json)\.
### A\.3Human simulation and anchor size
Each grid point uses 12 direct Monte Carlo replicates\. Inverting each replicate’s curve separately gives sample standard deviations ofνH\\nu\_\{H\}of 0\.0052, 0\.0134, and 0\.0203 for the three full panels\. These describe simulation spread only\. Replicate random numbers are reused across grid sizes and anchor configurations\. All nine[mean reference curves](https://supplement/calibration_curves.csv)for three tasks and three anchors are strictly increasing\.
The mode anchor in Table[4](https://arxiv.org/html/2609.21277#S4.T4)selects the first maximal count in the order entailment, neutral, contradiction \(or option 1, option 2\)\. Its tie rule differs from the suppliedgoldi\\mathrm\{gold\}\_\{i\}on 6, 2, and 2 items\.
ForM=5,10,20,30,50,75,100M=5,10,20,30,50,75,100, drawMMlabels without replacement from each item’s 100\-label pool and normalize to obtainAM,iA\_\{M,i\}\. Each partial\-count setting uses 60 anchor replicates; each replicate recomputes the observed PR and reference curve\. Simulated labels always come from fullhih\_\{i\}, and both model and simulated residuals subtractAM,iA\_\{M,i\}\. Figure[A1](https://arxiv.org/html/2609.21277#A1.F1)shows medians and 5th–95th percentiles;M=100M=100is deterministic\.
Figure A1:Human\-reference label countMM: median and 5th–95th percentiles ofνH\\nu\_\{H\}across 60 sampled anchors;M=100M=100uses the full reference\.Conditional on the finite label pool, the shared anchor\-estimation displacement has expected squared magnitude
𝔼A\|h\[1n∑i‖AM,i−hi‖2\]\\displaystyle\\mathbb\{E\}\_\{A\\mid h\}\\left\[\\frac\{1\}\{n\}\\sum\_\{i\}\\\|A\_\{M,i\}\-h\_\{i\}\\\|^\{2\}\\right\]\(A1\)=100−M99M1n∑i\(1−‖hi‖2\)\.\\displaystyle=\\frac\{100\-M\}\{99M\}\\frac\{1\}\{n\}\\sum\_\{i\}\(1\-\\\|h\_\{i\}\\\|^\{2\}\)\.It decreases to zero atM=100M=100\. This is a finite\-pool statement, not an error estimate relative to an unknown human population\.
### A\.4Subpanel sampling and search
For 9\-judge diverse panels, omit one family with probability proportional to1/sg1/s\_\{g\}, wheresgs\_\{g\}is its size, then select one member uniformly from each remaining family\. This samples uniformly over valid panels\. For size 16, require at least one member from every family and minimize∑g\(bg2\)\\sum\_\{g\}\\binom\{b\_\{g\}\}\{2\}over family allocationsbgb\_\{g\}\. Weight optimal allocations by∏g\(sgbg\)\\prod\_\{g\}\\binom\{s\_\{g\}\}\{b\_\{g\}\}and sample uniformly without replacement within families\. This again samples uniformly over valid panels\. Each design uses 2,000 draws\.
The search in Figure[D2](https://arxiv.org/html/2609.21277#A4.F2)uses sizes2,4,6,8,12,16,24,322,4,6,8,12,16,24,32\. At fixed size, minimizingQ\(I\)=∑a,b∈I,a≠bCa,b2Q\(I\)=\\sum\_\{a,b\\in I,\\ a\\neq b\}C\_\{a,b\}^\{2\}maximizesPRI=\|I\|2/\(\|I\|\+Q\(I\)\)\\mathrm\{PR\}\_\{I\}=\|I\|^\{2\}/\(\|I\|\+Q\(I\)\)\. Start from the model with smallest squared\-inner\-product row sum and from 20 random starts\. Add members greedily by the smallest increase inQQ, then apply improving one\-in/one\-out swaps and keep the best candidate\.
Panel search and evaluation use all items in Figure[D2](https://arxiv.org/html/2609.21277#A4.F2)a\. For Figure[D2](https://arxiv.org/html/2609.21277#A4.F2)b, 50 random half splits generate candidates on the first half and rebuild the human reference for evaluation on the second half\. The full panel is a candidate\. Selecting across sizes on the second half makes the gains exploratory, not independent test performance\.
### A\.5Independent\-transition reference
For each model, estimate a transition matrixTaT\_\{a\}from paired responses, conditional on the original label\. Independently sample each item’s replacement from the corresponding row, leaving other models unchanged\. An unobserved original\-label row is defined as no change\. This matches conditional flip rates and destination frequencies in expectation\.
Each model uses 200 replicates\. The bands in Figure[6](https://arxiv.org/html/2609.21277#S4.F6)a are their 2\.5th–97\.5th percentiles, without multiplicity correction\. Figure[6](https://arxiv.org/html/2609.21277#S4.F6)b plots\|𝔼simΔq¯sim\|/\|Δq¯swap\|\|\\mathbb\{E\}\_\{\\mathrm\{sim\}\}\\Delta\\bar\{q\}\_\{\\mathrm\{sim\}\}\|/\|\\Delta\\bar\{q\}\_\{\\mathrm\{swap\}\}\|, which is undefined for a zero denominator\.
### A\.6Binary\-error size and majority vote
Defineεa,i=𝟏\[l\(a,i\)≠goldi\]\\varepsilon\_\{a,i\}=\\mathbf\{1\}\[l\(a,i\)\\neq\\mathrm\{gold\}\_\{i\}\]and letϕ¯\\bar\{\\phi\}be the mean pairwise Pearson correlation of these errors\. Table[3](https://arxiv.org/html/2609.21277#S4.T3)usesneff=k/\[1\+\(k−1\)ϕ¯\]n\_\{\\mathrm\{eff\}\}=k/\[1\+\(k\-1\)\\bar\{\\phi\}\]\([Kish, 1965](https://arxiv.org/html/2609.21277#bib.bib14);[Kohli, 2026b](https://arxiv.org/html/2609.21277#bib.bib16)\), requiring nonzero error variances and positive denominator\. This design effect describes averaging standardized errors; it directly gives a variance ratio for raw mean errors only under equal error variances\.
Task majority vote selects among tied labels using a deterministic hash of the item and complete response vector\. The geometricJmajJ\_\{\\mathrm\{maj\}\}instead selects the first modal label in canonical order\. Results use their respective definitions\. Exact requested model identifiers, seeds, and tie algorithms are recorded in the[protocol](https://supplement/analysis_protocol.json)\.
## Appendix BGeometric References and Projections
### B\.1Individual residual geometry
For three\-label tasks, the2×22\\times 2second momentGa=Ba,aG\_\{a\}=B\_\{a,a\}describes an individual judge’s residual geometry\. With eigenvaluesσa,12≥σa,22\\sigma\_\{a,1\}^\{2\}\\geq\\sigma\_\{a,2\}^\{2\}and positive trace, define anisotropy
Aa=σa,12−σa,22σa,12\+σa,22∈\[0,1\]\.A\_\{a\}=\\frac\{\\sigma\_\{a,1\}^\{2\}\-\\sigma\_\{a,2\}^\{2\}\}\{\\sigma\_\{a,1\}^\{2\}\+\\sigma\_\{a,2\}^\{2\}\}\\in\[0,1\]\.\(B1\)LargerAaA\_\{a\}means greater concentration along one axis\. The principal\-axis angleθa\\theta\_\{a\}is defined moduloπ\\pirelative tov1v\_\{1\}\([Mardia and Jupp, 2000](https://arxiv.org/html/2609.21277#bib.bib21)\); it is unidentified when the eigenvalues coincide\. Binary residuals are one\-dimensional, so we do not apply this geometry toα\\alphaNLI\.
We compare each judge with simulated label sequences drawn itemwise fromhh\. The analytic human second moment sets the angular origin\. A Mahalanobis distance measures displacement in angle–anisotropy coordinates relative to the simulated cloud, accounting for their scale and covariance\. The definitions of these references follow below\.
Two constructed judges clarify these displacements\.JgoldJ\_\{\\mathrm\{gold\}\}returnsgoldi\\mathrm\{gold\}\_\{i\}on each item without model inference, providing a comparison inspired by negative controls\([Lipsitch et al\., 2010](https://arxiv.org/html/2609.21277#bib.bib18)\)\.JmajJ\_\{\\mathrm\{maj\}\}returns the panel’s modal label, resolving ties in the canonical order of Section[3\.2](https://arxiv.org/html/2609.21277#S3.SS2)\. It is used for geometry; task\-agreement results use the separate majority\-vote tie rule in Appendix[A](https://arxiv.org/html/2609.21277#A1)\.
### B\.2Human geometry and cloud distance
LetV=\[v1,v2\]V=\[v\_\{1\},v\_\{2\}\]\. For itemwise human\-reference labelsYi∼hiY\_\{i\}\\sim h\_\{i\}, the expected uncentered residual second moment is
G∗=1n∑iV⊤\[diag\(hi\)−hihi⊤\]V\.G\_\{\*\}=\\frac\{1\}\{n\}\\sum\_\{i\}V^\{\\top\}\[\\operatorname\{diag\}\(h\_\{i\}\)\-h\_\{i\}h\_\{i\}^\{\\top\}\]V\.\(B2\)Ifwaw\_\{a\}is the leading unit eigenvector ofGaG\_\{a\}, its axial angle isθa=atan2\(wa,2,wa,1\)modπ\\theta\_\{a\}=\\operatorname\{atan2\}\(w\_\{a,2\},w\_\{a,1\}\)\\bmod\\pi\. Applying the same definition toG∗G\_\{\*\}givesθ∗\\theta\_\{\*\}and reference anisotropyA∗A\_\{\*\}\. The plotted difference is
Δθa=\(\(θa−θ∗\+π/2\)modπ\)−π/2\.\\Delta\\theta\_\{a\}=\(\(\\theta\_\{a\}\-\\theta\_\{\*\}\+\\pi/2\)\\bmod\\pi\)\-\\pi/2\.\(B3\)It is converted to degrees and is invariant to eigenvector sign\. Repeated leading eigenvalues leave the axis unidentified\.
For a simulated sequencebb, setzb=\(\(180/π\)Δθb,Ab\)⊤z\_\{b\}=\(\(180/\\pi\)\\Delta\\theta\_\{b\},A\_\{b\}\)^\{\\top\}\. The mean of 6,400 simulated points isμ0\\mu\_\{0\}, and their sample covariance, with denominator 6,399, isΣ0\\Sigma\_\{0\}\. Assuming positive definiteness,
d\(z\)=\(z−μ0\)⊤Σ0−1\(z−μ0\)\.d\(z\)=\\sqrt\{\(z\-\\mu\_\{0\}\)^\{\\top\}\\Sigma\_\{0\}^\{\-1\}\(z\-\\mu\_\{0\}\)\}\.\(B4\)Ellipses atd=1,2d=1,2are equal\-distance contours, not confidence regions of specified coverage\. The angle origin uses the analytic reference; distances use the empirical cloud center\.
WithμA,0\\mu\_\{A,0\}the center’s anisotropy, the gold\-to\-model excess ratio is\(AJgold−μA,0\)/\(k−1∑aAa−μA,0\)\(A\_\{J\_\{\\mathrm\{gold\}\}\}\-\\mu\_\{A,0\}\)/\(k^\{\-1\}\\sum\_\{a\}A\_\{a\}\-\\mu\_\{A,0\}\), requiring nonzero denominator\. It compares displacement magnitudes and is not an additive fraction of explained error\.
### B\.3Projecting the consensus direction
Figure[3](https://arxiv.org/html/2609.21277#S4.F3)selects the first model key lexicographically within each family: Doubao\-seed1\.8, DeepSeekV3\.2, Gemini2\.5\-pro, GLM5, GPT4\.1, Grok4\.5, Claude\-haiku4\.5, KimiK2\.5, MiniMaxM3, and Qwen3\.5\-plus\.
LetP2P\_\{2\}contain the first two orthonormal eigenvectors ofSS, and setw=P2⊤u1=\(w1,w2\)⊤w=P\_\{2\}^\{\\top\}u\_\{1\}=\(w\_\{1\},w\_\{2\}\)^\{\\top\}\. For‖w‖\>0\\\|w\\\|\>0, the displayed directions are
u∥=P2w‖w‖,u⟂=P2\(−w2,w1\)⊤‖w‖\.u\_\{\\parallel\}=\\frac\{P\_\{2\}w\}\{\\\|w\\\|\},\\qquad u\_\{\\perp\}=\\frac\{P\_\{2\}\(\-w\_\{2\},w\_\{1\}\)^\{\\top\}\}\{\\\|w\\\|\}\.\(B5\)They form an orthonormal basis of the displayed plane, butu∥u\_\{\\parallel\}need not equal the fullu1u\_\{1\}\. The projection is undefined atw=0w=0; a tie between the second and third eigenvalues makes the plane nonunique\. Plot axes have equal scale and limits at the 99th percentile ofmax\(\|PC1\|,\|PC2\|\)\\max\(\|\\mathrm\{PC1\}\|,\|\\mathrm\{PC2\}\|\); a small number of points lie outside the frame\. The displayed PC scores in Figure[3](https://arxiv.org/html/2609.21277#S4.F3)use residual coordinates uniformly rescaled by2\\sqrt\{2\}before projection; this does not change directions or normalized variance shares\.
## Appendix CSpectral Size and Distributional Error
### C\.1Representation and summary
For a symmetric positive semidefinite unit\-diagonal matrixRRwithk≥2k\\geq 2, letr¯\\bar\{r\}andr2¯\\overline\{r^\{2\}\}be the mean and mean square over unordered off\-diagonal pairs\. Signed aggregationk/\[1\+\(k−1\)r¯\]k/\[1\+\(k\-1\)\\bar\{r\}\], when its denominator is positive, and PRk/\[1\+\(k−1\)r2¯\]k/\[1\+\(k\-1\)\\overline\{r^\{2\}\}\]summarize different moments\. Table[C1](https://arxiv.org/html/2609.21277#A3.T1)crosses these summaries with the binary\-error Pearson matrixΦ\\Phiand human\-residual Gram matrixCC\. The signed summary ofΦ\\Phiisneffn\_\{\\mathrm\{eff\}\}\. Applying the same expression toCCis descriptive and is not a separate human\-count equivalence\.
Moving fromΦ\\PhitoCCchanges the representation, including encoding, reference, and centering: Pearson correlations center binary errors, whereasCCretains the means of distributional residuals\. The crossed comparison does not isolate the effect of human disagreement alone\. Table[4](https://arxiv.org/html/2609.21277#S4.T4)separately varies the anchor within the same uncentered residual construction and recalibrates each anchor\.
Table C1:Crossed representations and dependence summaries for the same 32 judges\.
### C\.2Error identities
For a one\-hot labelY∼hiY\\sim h\_\{i\},𝔼Y=hi\\mathbb\{E\}Y=h\_\{i\}andtrCov\(Y\)=1−‖hi‖2\\operatorname\{tr\}\\operatorname\{Cov\}\(Y\)=1\-\\\|h\_\{i\}\\\|^\{2\}\. Averagingmmconditionally independent labels divides this variance bymm\. Averaging over items yieldsℰH\(m\)=J/m\\mathcal\{E\}\_\{H\}\(m\)=J/mand hence Equation \([6](https://arxiv.org/html/2609.21277#S3.E6)\)\. This expectation concerns within\-item draws and does not require independence across items\.
Decompose the raw Gram matrix asKa,b=∑t\(St\)a,b\+μa⊤μbK\_\{a,b\}=\\sum\_\{t\}\(S\_\{t\}\)\_\{a,b\}\+\\mu\_\{a\}^\{\\top\}\\mu\_\{b\}\. Substitution intok−2𝟏⊤K𝟏k^\{\-2\}\\mathbf\{1\}^\{\\top\}K\\mathbf\{1\}yields Equation \([13](https://arxiv.org/html/2609.21277#S3.E13)\)\. The covariances use denominatornnand the original residual scale; member\-normalized covariances cannot replace them\.
### C\.3Spectral orientation identity
Under the positive\-energy assumption,K=DCDK=DCDandD𝟏=dD\\mathbf\{1\}=d\. Substituting an orthonormal eigendecomposition gives
ℰ\\displaystyle\\mathcal\{E\}=d⊤Cdk2=1k2∑jλj\(wj⊤d\)2\\displaystyle=\\frac\{d^\{\\top\}Cd\}\{k^\{2\}\}=\\frac\{1\}\{k^\{2\}\}\\sum\_\{j\}\\lambda\_\{j\}\(w\_\{j\}^\{\\top\}d\)^\{2\}\(C1\)=Ωk2∑jλjbj\.\\displaystyle=\\frac\{\\Omega\}\{k^\{2\}\}\\sum\_\{j\}\\lambda\_\{j\}b\_\{j\}\.Parseval’s identity gives∑jbj=1\\sum\_\{j\}b\_\{j\}=1\. For a repeated eigenvalue, individualbjb\_\{j\}depend on the basis within its eigenspace, but their sum over that eigenspace and the weighted sum in Equation \([8](https://arxiv.org/html/2609.21277#S3.E8)\) do not\. The identity is a decomposition of the observed error, not a prediction from PR alone\. It uses the uncentered Gram matrix; replacing it by centered covariance would drop the mean\-residual term\.
### C\.4Realizable hard\-label counterexamples
First, equal spectra need not imply equal error\. With four items andhi=\(1/2,1/2\)h\_\{i\}=\(1/2,1/2\), encode labels by signs and seta=\(1,1,−1,−1\)⊤a=\(1,1,\-1,\-1\)^\{\\top\},b=\(1,−1,1,−1\)⊤b=\(1,\-1,1,\-1\)^\{\\top\}\. Panels\[a,a,b,b\]\[a,a,b,b\]and\[a,−a,b,−b\]\[a,\-a,b,\-b\]both have normalized Gram eigenvalues\(2,2,0,0\)\(2,2,0,0\), PR of 2, and member energies1/21/2\. Their distributional errors are1/41/4and 0\. What differs is the eigenvector orientation relative to equal\-weight averaging\.
The nonnegative\-correlation construction in Section[3\.4](https://arxiv.org/html/2609.21277#S3.SS4)uses rows as items and columns as judges:
Z\\displaystyle Z=\(111111−1−1−1−111−1−1−1−1\),\\displaystyle=\\begin\{pmatrix\}1&1&1&1\\\\ 1&1&\-1&\-1\\\\ \-1&\-1&1&1\\\\ \-1&\-1&\-1&\-1\\end\{pmatrix\},\(C2\)H\\displaystyle H=\(11111−11−111−1−11−1−11\)\.\\displaystyle=\\begin\{pmatrix\}1&1&1&1\\\\ 1&\-1&1&\-1\\\\ 1&1&\-1&\-1\\\\ 1&\-1&\-1&1\\end\{pmatrix\}\.FormZA=\[Z;Z\]Z\_\{A\}=\[Z;Z\]andZB=\[H;𝟏4𝟏4⊤\]Z\_\{B\}=\[H;\\mathbf\{1\}\_\{4\}\\mathbf\{1\}\_\{4\}^\{\\top\}\], then append each matrix’s negation:\[ZA;−ZA\]\[Z\_\{A\};\-Z\_\{A\}\]and\[ZB;−ZB\]\[Z\_\{B\};\-Z\_\{B\}\]\. Both panels now have 16 items with the samehi=\(1/2,1/2\)h\_\{i\}=\(1/2,1/2\)\. Mapping signs to one\-hot labels gives energy1/21/2for every member and zero across\-item mean residual\. Their normalized Gram matrices are
C\(A\)\\displaystyle C^\{\(A\)\}=\(1100110000110011\),\\displaystyle=\\begin\{pmatrix\}1&1&0&0\\\\ 1&1&0&0\\\\ 0&0&1&1\\\\ 0&0&1&1\\end\{pmatrix\},\(C3\)C\(B\)\\displaystyle C^\{\(B\)\}=12I4\+12𝟏4𝟏4⊤\.\\displaystyle=\\tfrac\{1\}\{2\}I\_\{4\}\+\\tfrac\{1\}\{2\}\\mathbf\{1\}\_\{4\}\\mathbf\{1\}\_\{4\}^\{\\top\}\.PanelAAhas\(c¯C,vC,q¯\)=\(1/3,2/9,1/3\)\(\\bar\{c\}\_\{C\},v\_\{C\},\\bar\{q\}\)=\(1/3,2/9,1/3\); panelBBhas\(1/2,0,1/4\)\(1/2,0,1/4\)\. Equations \([3](https://arxiv.org/html/2609.21277#S3.E3)\) and \([9](https://arxiv.org/html/2609.21277#S3.E9)\) give PR2→16/72\\to 16/7and error1/4→5/161/4\\to 5/16\. The conflict requires neither negative correlations, unequal energies, nor an across\-item mean displacement\.
## Appendix DPresentation and Ranking Diagnostics
### D\.1Individual agreement and error clustering
Agreement withgoldi\\mathrm\{gold\}\_\{i\}ranges from 59\.7% to 72\.0% on MNLI\-m, 71\.2% to 85\.3% on SNLI, and 84\.4% to 93\.5% onα\\alphaNLI; the corresponding model means are 66\.0%, 81\.2%, and 90\.6%\. Rankings differ by task: Grok4\.6 has the lowest agreement on MNLI\-m and the highest onα\\alphaNLI\. Panel majority vote achieves 68\.6%, 86\.5%, and 93\.0%, exceeding the best individual model only on SNLI\.
Errors also cluster across judges \(Figure[D1](https://arxiv.org/html/2609.21277#A4.F1)\)\. All 32 judges disagree with gold on 38, 8, and 4 items, respectively\. A reference that multiplies the judges’ marginal error probabilities predicts far fewer than one such item\. This comparison rejects the unconditional independence description of these data; item\-difficulty heterogeneity can also produce clustering, so it does not identify a specific source of shared model error\.
Figure D1:\(a\) Individual agreement with human gold; dashed lines show panel majority vote\. \(b\) Observed error counts versus an unconditional independence reference\.
### D\.2Exploratory search in the fixed pool
Larger panels do not always have largerνH\\nu\_\{H\}\. Among the searched candidates in Figure[D2](https://arxiv.org/html/2609.21277#A4.F2)a, MNLI\-m andα\\alphaNLI peak atk=16k=16, reaching 5\.067 and 6\.958\. SNLI peaks atk=24k=24with 7\.268, close to 7\.254 atk=16k=16\. All full 32\-model panels lie below these peaks\. This is a property of the searched candidates in this pool, distinct from the trend in random\-panel medians\.
Figure[D2](https://arxiv.org/html/2609.21277#A4.F2)b uses 50 half\-item splits\. Candidates are generated at each size on one half, then evaluated and selected across sizes on the other half\. Median gains over the full panel are 0\.690, 0\.656, and 0\.035\. Because the full panel is a candidate and the second half also selects the size, these gains are exploratory search diagnostics, not held\-out selection benefits\.
Figure D2:Subpanel search in the fixed pool\. \(a\) Best candidates by size; circles mark peaks and dashed lines the full panel\. \(b\) Gains after selecting the size on the second half of items; vertical lines mark medians\.
### D\.3Provider families
To test family membership as a diversity proxy, we compare 2,000 random 9\-judge panels with 2,000 panels containing at most one judge per family\. Family diversity yields small medianνH\\nu\_\{H\}gains on SNLI andα\\alphaNLI and almost no change on MNLI\-m \(Table[D1](https://arxiv.org/html/2609.21277#A4.T1)\)\. Its value depends on the task even within the same model pool\.
Table D1:MedianνH\\nu\_\{H\}for 9\-judge panels; differences are computed before rounding\.
### D\.4Additional presentation diagnostics
Section[4\.7](https://arxiv.org/html/2609.21277#S4.SS7)and Figure[6](https://arxiv.org/html/2609.21277#S4.F6)compare observed presentation changes with independent transitions\. Here we examine the changes in spectral size and the destinations of changed labels within the same saved responses\.
For 200 random 10\-judge subpanels, replacing each member once gives 2,000 comparisons per dataset\. Mean changes inνH\\nu\_\{H\}are\+0\.006\+0\.006,−0\.014\-0\.014, and\+0\.096\+0\.096\. The observed replacements provide no consistent cross\-task increase inνH\\nu\_\{H\}\.
##### Flip destinations\.
The fraction of originally incorrect responses among flips is 1\.62, 2\.81, and 5\.36 times the overall error rate among valid model–item pairs\. Pooling 1,524 MNLI\-m flips and 1,127 SNLI flips, 87\.4% and 91\.7% go to the highest\-human\-support label after excluding the original label; ties take the first label in canonical order\. For binaryα\\alphaNLI, removing the original leaves no destination choice\.
The originally correct subsets have 701 and 549 flips, with corresponding destination rates 90\.3% and 92\.5%\. The originally incorrect subsets have 823 and 578 flips, with rates 84\.9% and 90\.8%, equal to their rates of switching to gold\. These data therefore do not distinguish choosing a well\-supported alternative from choosing gold among originally incorrect responses\.
Onα\\alphaNLI, flips favor the first\-presented option for 22 judges and the other option for 7\. The median judge\-level fraction selecting the first\-presented option after a flip is 61\.5%, whereas pooling all flips gives 53\.9%\.
### D\.5Within\-size comparisons and item stability
For eachk=2,…,31k=2,\\ldots,31, we draw 2,000 panels without replacement within a panel but with possible repeats across draws\. Deduplicating member sets and including the full 32\-judge panel leaves 54,139 panels per dataset\. There is no within\-size correlation atk=32k=32, where only one panel exists\. The[protocol](https://supplement/loss_diagnostics/protocol.json)supplies seeds and member lists\.
Fork=4,8,16,24,31k=4,8,16,24,31, adding each absent model gives 150,372 original\-panel/added\-member records per dataset\. Different additions can yield the same enlarged panel and share both judges and items\. Item\-identifier hashes define two fixed 500\-item groups; raw Gram matrices and PR are recomputed within each group\. The split lists accompany the supplement\.
Table[5](https://arxiv.org/html/2609.21277#S4.T5)requires the same conflict direction on the full set and both halves, with both absolute relative changes at least 1% on every set\. Numerical ties are excluded at\|ΔPR\|≤10−10\|\\Delta\\mathrm\{PR\}\|\\leq 10^\{\-10\}or\|Δℰ\|≤10−12\|\\Delta\\mathcal\{E\}\|\\leq 10^\{\-12\}\. These overlapping records describe the fixed pool and are not independent replications\.
##### Failed\-item sensitivity\.
Removing the 1/0/5 affected items on MNLI\-m, SNLI, andα\\alphaNLI changes full\-panelℰ\\mathcal\{E\}from 0\.196911 to 0\.197088 on MNLI\-m and from 0\.048254 to 0\.048357 onα\\alphaNLI; SNLI is unchanged\. For4→54\\to 5, stable conflicts passing the same 1% filter change from 2,116/309/1,127 to 2,079/309/1,055 out of 54,348 additions per task\. The clean views retain the original item\-half memberships without refilling removed items\. These saved sensitivity results preserve the presence of conflicts and their task\-dependent frequencies\.
Figure[D3](https://arxiv.org/html/2609.21277#A4.F3)reports rank agreement between item halves across all panel sizes\. Onα\\alphaNLI, both PR and error rankings are less stable than on MNLI\-m, demonstrating sensitivity to item composition even within a fixed model pool\.
Figure D3:Within\-size rank correlations between fixed item halves for PR and distributional errorℰ\\mathcal\{E\}\.相似文章
用LLM评审员增强人工评估:你需要多少人工审核?
本文提出了一种两阶段抽样设计,其中LLM评估用于增强而非替代人工评分,并利用缺失数据文献中的双重稳健估计量,提供了确定人工和LLM评审样本量的指导。
十六个模型,不足两种声音:在无唯一正确答案时衡量集成离散度
这篇 arXiv 预印本研究形成集成的十六个语言模型的语义离散度,表明平均而言集成多样性较小,且模型身份仅部分解释了哪个模型的差异最大。作者提出了一种逐模型异议贡献度量,并发现离散度由临床内容而非解释开放性所组织。
MM-JudgeBias:评测 MLLM-as-a-Judge 组合偏差的基准
研究者发布 MM-JudgeBias 基准,揭示多模态大模型在充当自动评判器时的系统性组合偏差,对 26 个 SOTA MLLM 在 1,800 条样本上进行测试。
在现实对话环境中评估语言模型
本文介绍了UPHELD,这是一个用于评估LLMs人类规模对话能力的大型基准,并提出了一个Mixture-of-Judges框架,将与人类评估的相关性提高了约30%。
一百万人,一百万个人工智能代理,三个基础模型。这能构成多元化的审议吗?——又如何衡量?
针对仅用三个基础模型支撑数百万个人工智能代理是否能产生真正多元化的审议进行批判性反思,认为模型间的相关误差可能造成虚假的一致意见,并寻求从集成学习中得出的可操作指标来衡量真正的人类代表性多样性。