Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs
Summary
This paper investigates whether stochastic sampling (self-consistency) in LLMs can capture cross-question structure similar to diverse ensembles. Using a Marchenko–Pastur test, the authors find that within a single model, stochastic variation yields at most one significant dimension, while an ensemble of 24 models yields four, revealing a dimensionality gap that limits self-consistency as an ensemble substitute.
View Cached Full Text
Cached at: 07/24/26, 05:00 AM
# 1 Introduction
Source: [https://arxiv.org/html/2607.20464](https://arxiv.org/html/2607.20464)
marginparsep has been altered\. topmargin has been altered\. marginparpush has been altered\. The page layout violates the ICML style\.Please do not change the page layout, or include packages like geometry, savetrees, or fullpage, which change it for you\. We’re not able to reliably undo arbitrary changes to the style\. Please remove the offending package\(s\), or layout\-changing commands and try again\.
Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs
Izhar Ali1
###### Abstract
When a language model gives different answers on repeated runs, does that variation reveal what it does not know? Self\-consistency turns the variation into a per\-question uncertainty estimate via majority voting\. But does the same variation reveal cross\-question structure—related questions flipping together, the way a diverse ensemble does? We compare two regimes on the same questions: one model run100100times atτ=1\\tau=1versus an ensemble of2424LLMs run once each atτ=0\\tau=0\. A Marchenko–Pastur random\-matrix test separates signal from sampling noise on both sides\. Within any single model, at most one dimension rises above noise across five families and three benchmarks \(MMLU, HellaSwag, GSM8K\)\. Across the ensemble, four eigenvalues clear the noise edge, while a matched\-difficulty Bernoulli null produces at most one in500500Monte Carlo draws\. Self\-consistency gives accurate per\-question uncertainty but no detectable cross\-question structure; only a diverse ensemble surfaces what a model does not know\.
††footnotetext:1Rowan University\. Correspondence to: Izhar Ali <aliizh94@rowan\.edu\>\.
2nd Workshop on Epistemic Intelligence in Machine Learning \(EIML@ICML 2026\), Seoul, South Korea\. Copyright 2026 by the author\(s\)\.Self\-consistency\(Wanget al\.,[2023](https://arxiv.org/html/2607.20464#bib.bib8)\)has become the default low\-cost uncertainty estimator for LLMs, alongside descendants that read stochastic variation as a signal of what a model knows: answer\-recurrence confidence\(Kadavathet al\.,[2022](https://arxiv.org/html/2607.20464#bib.bib21)\), hallucination flagging by agreement\(Manakulet al\.,[2023](https://arxiv.org/html/2607.20464#bib.bib22)\), semantic entropy\(Farquharet al\.,[2024](https://arxiv.org/html/2607.20464#bib.bib7)\)\. Whether that signal can stand in for the more expensive alternative of deep ensembles\(Lakshminarayananet al\.,[2017](https://arxiv.org/html/2607.20464#bib.bib14)\)depends on what kind of information it carries\.
First, stochastic samples can estimate a*per\-question*success probabilitypip\_\{i\}by averaging independent noise\. Second, they can reveal*cross\-question*structure: correlated errors across related questions, the way a diverse ensemble does\(Kendall and Gal,[2017](https://arxiv.org/html/2607.20464#bib.bib23); Hüllermeier and Waegeman,[2021](https://arxiv.org/html/2607.20464#bib.bib24)\)\. To substitute for an ensemble, self\-consistency needs the second kind too\.
Cross\-model dimensionality is well characterized—a single factor captures 79% of variance on 5,000\+ models\(Kipniset al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib1)\), factor analysis yields 8 dimensions on 60 models\(Maimonet al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib2)\), low\-dimensional joint embeddings\(Yaoet al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib3)\)—but these all analyze a model×\\timesbenchmark score matrix\. Within\-model dimensionality, under a sampling\-noise null, has not been measured\. We close this gap with a Marchenko–Pastur test\(Marčenko and Pastur,[1967](https://arxiv.org/html/2607.20464#bib.bib10)\)on the run×\\timesquestion correctness matrix\.
One dimension within, four across: this is the*dimensionality gap*\.
#### Contributions\.
1. 1\.We define within\-model epistemic dimensionality as the number of independent directions of structured error in one model’s stochastic samples\.
2. 2\.Empirically, we find at most one above\-noise dimension within each of five LLMs across three benchmarks \(MMLU, HellaSwag, GSM8K chain\-of\-thought\), but four above\-noise dimensions across2424diverse models on MMLU—unmatched by any of500500matched\-difficulty Bernoulli null draws\.
## 2Method
### 2\.1Within\-Model Stochastic Probing
Letℳ\\mathcal\{M\}be a language model and\{q1,…,qN\}\\\{q\_\{1\},\\ldots,q\_\{N\}\\\}beNNquestions\. We generateKKstochastic completions per question at temperatureτ\>0\\tau\>0\. For runkkand questionii, define
cki=\{1ifℳanswersqicorrectly in runk,0otherwise\.c\_\{ki\}=\\begin\{cases\}1&\\text\{if \}\\mathcal\{M\}\\text\{ answers \}q\_\{i\}\\text\{ correctly in run \}k,\\\\ 0&\\text\{otherwise\.\}\\end\{cases\}This gives aK×NK\\times Nbinary matrix𝐂\\mathbf\{C\}\.
A question is*borderline*if its pass ratefi=1K∑kckif\_\{i\}=\\frac\{1\}\{K\}\\sum\_\{k\}c\_\{ki\}lies in\(ε,1−ε\)\(\\varepsilon,1\-\\varepsilon\)withε=0\.05\\varepsilon=0\.05\. Let𝐂B\\mathbf\{C\}\_\{B\}denote𝐂\\mathbf\{C\}restricted to theNbN\_\{b\}borderline questions; only these contribute non\-trivial variance to the correlation test below\.
### 2\.2Marchenko–Pastur \(MP\) Null Test
We test column independence in𝐂B\\mathbf\{C\}\_\{B\}using a random\-matrix null\. Compute theNb×NbN\_\{b\}\\times N\_\{b\}correlation matrix𝐑\\mathbf\{R\}of𝐂B\\mathbf\{C\}\_\{B\}\(each borderline question a variable, each of theKKruns an observation\) and its eigenvaluesλ1≥⋯≥λNb\\lambda\_\{1\}\\geq\\cdots\\geq\\lambda\_\{N\_\{b\}\}\. Under the null of independent Bernoulli columns,𝐑\\mathbf\{R\}’s spectrum follows the Marchenko–Pastur law\(Marčenko and Pastur,[1967](https://arxiv.org/html/2607.20464#bib.bib10)\)with upper edgeλ\+=\(1\+γ\)2\\lambda\_\{\+\}=\(1\+\\sqrt\{\\gamma\}\)^\{2\},γ:=Nb/K\\gamma:=N\_\{b\}/K, and the standardized top eigenvaluez:=\(λ1−λ\+\)/σTWz:=\(\\lambda\_\{1\}\-\\lambda\_\{\+\}\)/\\sigma\_\{\\mathrm\{TW\}\}converges to Tracy–WidomF1F\_\{1\}\(Prop\.[A\.2](https://arxiv.org/html/2607.20464#A1.Thmtheorem2)\)\. We usezzas the primary null test and report the count\|\{i:λi\>λ\+\}\|\|\\\{i:\\lambda\_\{i\}\>\\lambda\_\{\+\}\\\}\|as a coarse summary; cross\-checked against Horn’s parallel analysis\(Horn,[1965](https://arxiv.org/html/2607.20464#bib.bib9)\)\(App\.[B\.1](https://arxiv.org/html/2607.20464#A2.SS1)\)\.
Rejecting this null is direct evidence of structured internal uncertainty: temperature sampling injects independent noise per prompt, so cross\-question covariation requires a coupling mechanism it does not supply\. If each question flips independently, the correlation matrix is asymptotically pure noise—no eigenvalue exceedsλ\+\\lambda\_\{\+\}beyond Tracy–Widom \(TW\) fluctuations \(Proposition[A\.2](https://arxiv.org/html/2607.20464#A1.Thmtheorem2), App\.[A](https://arxiv.org/html/2607.20464#A1)\)\.
### 2\.3Across\-Model Comparison
We apply the same MP test to a diverse ensemble ofMMmodels, one run each atτ=0\\tau\\\!=\\\!0, with models \(rather than runs\) as observations\. We restrict to questions where models split \(neither all\-correct nor all\-incorrect\)—the across\-model analogue of the borderline filter from §[2\.1](https://arxiv.org/html/2607.20464#S2.SS1)\. The primary statistic is the MP signal count, calibrated against an independent\-Bernoulli null with per\-column rates matching observed per\-question pass rates \(500500Monte Carlo draws; §[3\.3](https://arxiv.org/html/2607.20464#S3.SS3)\)\. For comparability with prior factor\-analytic work\(Kipniset al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib1); Maimonet al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib2); Yaoet al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib3); Wenet al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib12)\)we also report the Shannon effective rankexp\(H\)\\exp\(H\); it saturates near theM−1M\\\!\-\\\!1ceiling under any heterogeneous\-rate null and is a comparability statistic, not an independence test \(App\.[B\.3](https://arxiv.org/html/2607.20464#A2.SS3)\)\.
### 2\.4Experimental Setup
#### Dataset\.
N=500N=500MMLU\(Hendryckset al\.,[2021](https://arxiv.org/html/2607.20464#bib.bib15)\)questions \(test split\), stratified across 57 subjects\.
#### Within\-model\.
Five models,K=100K=100stochastic completions per question atτ=1\.0\\tau=1\.0: Qwen2\.5\-7B\-Instruct, Mistral\-7B\-Instruct\-v0\.3, SmolLM2\-1\.7B\-Instruct, Phi\-3\-mini\-4k\-Instruct, Meta\-Llama\-3\-8B\-Instruct\. Temperature robustness:τ∈\{0\.5,1\.5\}\\tau\\in\\\{0\.5,1\.5\\\},K=30K=30,N=100N=100on Qwen2\.5\-7B\.
#### Across\-model\.
N=500N=500MMLU questions,K=1K=1run per model atτ=0\\tau=0, across 24 instruction\-tuned models spanning eleven families and 0\.5B–22B: Qwen \(5\), Mistral \(4\), Zephyr\-7B, Phi \(3\), SmolLM2\-1\.7B, OLMo\-2\-7B, Yi\-1\.5\-6B, DeepSeek\-7B, Granite\-3\.1\-8B, Llama \(4\), Gemma\-2 \(2\)\.
## 3Results
### 3\.1Within\-Model: No Structure Above Noise
Within a single model, the run×\\timesquestion correctness matrix is statistically indistinguishable from independent Bernoulli draws: at most one eigenvalue reaches the MP edge, and only within Tracy–Widom sampling noise\. Qwen2\.5\-7B\-Instruct \(71\.2%71\.2\\%atτ=1\.0\\tau\\\!=\\\!1\.0;77\.4%77\.4\\%greedy\) has121/500121/500borderline questions; the remaining379379are near\-deterministic\.
#### MP test\.
The top eigenvalue sits below the MP noise edge \(λ1=4\.17\\lambda\_\{1\}\\\!=\\\!4\.17vs\.λ\+=4\.41\\lambda\_\{\+\}\\\!=\\\!4\.41atγ=1\.21\\gamma\\\!=\\\!1\.21\): zero above\-noise eigenvalues \(Fig\.[1](https://arxiv.org/html/2607.20464#S3.F1)a\)\. Getting a history question right on a given pass predicts nothing about a chemistry question on the same pass nor about any same\-subject question \(within\- and between\-subject borderline correlations both≈0\\approx\\\!0\)\. Horn’s parallel analysis agrees\. Thedwithin≤1d\_\{\\text\{within\}\}\\\!\\leq\\\!1ceiling holds acrossε∈\[0\.01,0\.20\]\\varepsilon\\\!\\in\\\!\[0\.01,0\.20\]\(Fig\.[A1](https://arxiv.org/html/2607.20464#A2.F1)\)\.
#### Replication across five families\.
The same pattern holds on Mistral\-7B, SmolLM2\-1\.7B, Phi\-3\-mini and Meta\-Llama\-3\-8B \(K=100K\\\!=\\\!100,τ=1\.0\\tau\\\!=\\\!1\.0; borderline counts121121–403403\); SmolLM2’sλmax=9\.05\\lambda\_\{\\max\}\\\!=\\\!9\.05touchesλ\+=9\.04\\lambda\_\{\+\}\\\!=\\\!9\.04but stays within MP sampling noise \(TWz=\+0\.05z\\\!=\\\!\+0\.05; Table[1](https://arxiv.org/html/2607.20464#S3.T1)\)\. Five families,1\.71\.7B–88B parameters, none with MP\-significant signal\.
### 3\.2Generalization: HellaSwag and GSM8K
The null is not an MMLU artifact\. Three models on HellaSwag\(Zellerset al\.,[2019](https://arxiv.org/html/2607.20464#bib.bib18)\)\(N=200N\\\!=\\\!200,K=50K\\\!=\\\!50\) and Qwen2\.5\-7B on GSM8K chain\-of\-thought\(Cobbeet al\.,[2021](https://arxiv.org/html/2607.20464#bib.bib17)\)\(N=100N\\\!=\\\!100,K=30K\\\!=\\\!30\) also give zero signal eigenvalues \(Table[1](https://arxiv.org/html/2607.20464#S3.T1)\)\. Within\- and between\-category correlations on HellaSwag are near zero \(\|r\|≤0\.05\|r\|\\\!\\leq\\\!0\.05\): no semantic coupling even under the dataset’s activity grouping\.
Figure 1:The dimensionality gap\.\(a\)Eigenvalue spectrum of the correctness correlation matrix on borderline MMLU questions \(Qwen2\.5\-7B,K=100K\\\!=\\\!100,Nb=121N\_\{b\}\\\!=\\\!121\)\. No eigenvalue exceeds the MP noise edgeλ\+=4\.41\\lambda\_\{\+\}=4\.41\(red dashed\)\.\(b\)Same\-metric comparison: MP signal count is≤1\\leq\\\!1for every one of five families on three tasks but44for the2424\-model ensemble \(τ=0\\tau\\\!=\\\!0\), unmatched by any of500500matched\-difficulty independent\-Bernoulli null draws\.\(c\)Split\-halfpip\_\{i\}on Qwen2\.5\-7B \(r=0\.994r=0\.994\): successive samples are independent draws from a fixed per\-question Bernoulli\.Table 1:Marchenko–Pastur test across three benchmarks\. Marginλ\+−λmax\\lambda\_\{\+\}\\\!\-\\\!\\lambda\_\{\\max\}is non\-negative or within MP noise for every model \(SmolLM2≤1\\leq\\\!1signal dimension\)\. TWzz: standard deviations by whichλmax\\lambda\_\{\\max\}exceeds the MP edge under the null \(App\.[A](https://arxiv.org/html/2607.20464#A1)\); Monte Carlo mean≈−1\.60\\approx\\\!\-1\.60atK=100K\\\!=\\\!100\.ModelNbN\_\{b\}λ\+\\lambda\_\{\+\}λmax\\lambda\_\{\\max\}marginTWzz*MMLU*\(K=100K\\\!=\\\!100,N=500N\\\!=\\\!500\)Qwen2\.5\-7B1214\.414\.17\+0\.24\+0\.24−2\.01\-2\.01Mistral\-7B\-v0\.32837\.196\.85\+0\.34\+0\.34−2\.37\-2\.37SmolLM2\-1\.7B4039\.049\.05−0\.01\-0\.01\+0\.05\+0\.05Phi\-3\-mini2696\.976\.78\+0\.19\+0\.19−1\.36\-1\.36Meta\-Llama\-3\-8B1735\.365\.05\+0\.31\+0\.31−2\.37\-2\.37*HellaSwag*\(K=50K\\\!=\\\!50,N=200N\\\!=\\\!200\)Qwen2\.5\-7B152\.402\.21\+0\.18\+0\.18−1\.15\-1\.15Mistral\-7B\-v0\.3584\.314\.04\+0\.28\+0\.28−1\.45\-1\.45Meta\-Llama\-3\-8B825\.204\.80\+0\.40\+0\.40−1\.96\-1\.96*GSM8K*\(K=30K\\\!=\\\!30,N=100N\\\!=\\\!100\)Qwen2\.5\-7B656\.115\.72\+0\.39\+0\.39−1\.27\-1\.27
### 3\.3Across\-Model: Structured Multi\-Dimensional Disagreement
The 24\-model correctness correlation matrix has four eigenvalues above the MP edge, unmatched by any of500500matched\-difficulty independent\-Bernoulli draws \(p≤1/500p\\\!\\leq\\\!1/500\)\. Model accuracies span 40\.2–77\.4%; on 459/500 questions \(91\.8%\) at least one model disagrees\. Mean pairwise Pearsonr=0\.35r\\\!=\\\!0\.35\(range0\.010\.01–0\.630\.63\), higher on same\-family pairs \(Table[2](https://arxiv.org/html/2607.20464#S3.T2); full distribution in Fig\.[A2](https://arxiv.org/html/2607.20464#A5.F2)\)\.
#### MP count \(primary\)\.
λ1\.\.4=83\.0,43\.1,36\.5,34\.6\\lambda\_\{1\.\.4\}\\\!=\\\!83\.0,\\,43\.1,\\,36\.5,\\,34\.6againstλ\+=28\.9\\lambda\_\{\+\}\\\!=\\\!28\.9\. The matched\-difficulty null \(§[2\.3](https://arxiv.org/html/2607.20464#S2.SS3);500500draws\) produced at most one signal eigenvalue on every draw \(mean0\.0280\.028, max11\)\. Even matched for per\-question difficulty, independent Bernoullis cannot produce the four above\-noise directions we observe\.
#### Shannon effective rank \(continuity with prior work\)\.
The same matrix hasexp\(H\)≈18\.7\\exp\(H\)\\\!\\approx\\\!18\.7, below every null draw \(𝔼\[exp\(H\)\]≈22\.3\\mathbb\{E\}\[\\exp\(H\)\]\\\!\\approx\\\!22\.3;p≤1/500p\\\!\\leq\\\!1/500\)\. Butexp\(H\)\\exp\(H\)does not discriminate regimes: within\- and across\-model values both sit in a ceiling\-adjacent band \(Table[2](https://arxiv.org/html/2607.20464#S3.T2); App\.[B\.3](https://arxiv.org/html/2607.20464#A2.SS3)\)\.
Table 2:The dimensionality gap on 500 MMLU questions\. Within\-model columns report ranges across the five MMLU models \(Table[1](https://arxiv.org/html/2607.20464#S3.T1)\)\.λmax\\lambda\_\{\\max\}: top eigenvalue of the correctness correlation matrix \(both sides\)\. MP signal count is primary and separates the regimes \(≤1\\leq\\\!1vs\.44\); Shannonexp\(H\)\\exp\(H\)saturates near its ceiling on both sides \(6464–87%87\\%vs\.81%81\\%\)\.
### 3\.4Robustness and Downstream Impact
#### Temperature robustness\.
Atτ=1\.5\\tau\\\!=\\\!1\.5\(K=30K\\\!=\\\!30,N=100N\\\!=\\\!100, Qwen2\.5\-7B\),dwithin=0d\_\{\\text\{within\}\}\\\!=\\\!0\. Atτ=0\.5\\tau\\\!=\\\!0\.5only33questions are borderline—below the MP test’s data floor, so no conclusion is drawn at this temperature\.
#### Per\-question estimation\.
Split\-halfpip\_\{i\}on Qwen2\.5\-7B correlates atr=0\.994r\\\!=\\\!0\.994\(Fig\.[1](https://arxiv.org/html/2607.20464#S3.F1)c\), calibration gap below0\.020\.02: stochastic sampling captures each scalarpip\_\{i\}precisely but carries no cross\-question information\.
#### Selective prediction\.
Two peer models beat100100\-sample self\-consistency at∼1/40\{\\sim\}\\,1/40th the cost\. The task is predicting whether Qwen2\.5\-7B’sτ=0\\tau\\\!=\\\!0answer is correct, with no ground\-truth labels at inference\. Operational SC100\(fraction of the100100samples agreeing with the modal answer\) reaches AUROC0\.7120\.712\.
Llama\-3\.1\-8B \+ Gemma\-2\-9B, each run once atτ=0\\tau\\\!=\\\!0, reach0\.8070\.807at∼1/40\{\\sim\}\\,1/40th the cost \(Qwen\-7B\-equivalent forward passes; App\.[C](https://arxiv.org/html/2607.20464#A3)\)\. A single external model \(Llama\-3\.1\-8B\) edges past SC100at∼1/88\{\\sim\}\\,1/88th the cost \(AUROC0\.7490\.749vs\.0\.7120\.712\), though this single\-peer gap is marginal atN=500N\\\!=\\\!500\(DeLongp≈0\.2p\\\!\\approx\\\!0\.2\); the two\-peer gap is highly significant \(p<0\.005p\\\!<\\\!0\.005\)\.
Figure 2:Cost\-AUROC Pareto frontier for selective prediction on Qwen2\.5\-7B \(τ=0\\tau\\\!=\\\!0\), 500 MMLU questions\.Ensemble2\(AUROC0\.810\.81\) beats SC100\(0\.710\.71\) at∼1/40\{\\sim\}\\,1/40th the compute cost\. Grey dotted: SC oracle upper bound \(requires ground\-truth labels at inference, not deployable\)\. Ensemble Pareto\-dominates operational SC across the entire cost range\. Cost: Qwen\-7B\-equivalent forward passes\.
## 4Discussion
#### Why temperature variation is shallow\.
Temperature is a uniform dilation of the logits: it does not selectively modulate any knowledge domain\. Consistent with this, within\- and between\-subject correlations are both≈0\{\\approx\}\\,0\(§[3\.1](https://arxiv.org/html/2607.20464#S3.SS1)\), and TWzz\-scores across five models and three benchmarks are individually consistent with the null \(allz≤\+0\.05z\\\!\\leq\\\!\+0\.05; Prop\.[A\.2](https://arxiv.org/html/2607.20464#A1.Thmtheorem2), Table[1](https://arxiv.org/html/2607.20464#S3.T1)\)\.
#### Locally committed, globally incoherent\.
When Qwen2\.5\-7B errs on a borderline question, it picks the same wrong answer87\.4%87\.4\\%of the time across samples—yet these wrong\-answer preferences do not associate across questions \(Appendix[D](https://arxiv.org/html/2607.20464#A4)\)\. The model has strong local commitments; it just doesn’t chain them into global structure\.
#### Exception: SmolLM2\.
SmolLM2 reaches TWz≈\+1\.77z\\\!\\approx\\\!\+1\.77atε=0\.15\\varepsilon\\\!=\\\!0\.15\(∼\\sim99th percentile\), consistent with slight residual cross\-question coupling that Assumption[A\.1](https://arxiv.org/html/2607.20464#A1.Thmtheorem1)idealizes away; a candidate mechanism is chat\-template randomness\.
#### Implications for epistemic intelligence\.
Selective\-prediction research should budget for diverse models, not deeper sampling\. Stacking more samples from a single model cannot recover a dimensionality that is not there\.
## 5Limitations and Future Work
- •*Binary correctness\.*We test for structure in correctness; richer signals \(log\-probabilities, semantic clusters\) could carry structure that correctness misses\.
- •*MP power\.*The test operates atγ∈\[0\.3,4\.0\]\\gamma\\\!\\in\\\!\[0\.3,4\.0\];K≫NbK\\\!\\gg\\\!N\_\{b\}would sharpen detection of structure near the MP edge\.
- •*Across\-model ceiling\.*dacrossd\_\{\\text\{across\}\}has not asymptoted byM=24M\\\!=\\\!24, so “four” is a lower bound on the true rank\.
- •*Post\-training\.*Instruction\-tuned models may differ from pretrained\-only bases\.
#### Future work\.
- •*Scope\.*Open\-generation benchmarks \(TruthfulQA, HaluEval\) and a 70B\-tier within\-model test\.
- •*Richer probes\.*Log\-probability or semantic\-cluster within\-model null; prompt perturbation as a second stochastic axis\.
- •*Mechanism\.*Decompose the same\-family SC100win \(§[C\.3](https://arxiv.org/html/2607.20464#A3.SS3)\) into disagreement vs\. calibration; compare against a multi\-seed Pythia ensemble\(Lakshminarayananet al\.,[2017](https://arxiv.org/html/2607.20464#bib.bib14)\)\.
- •*Matched sweep\.*Temperature sweep at matchedKKto separate sample\-count from temperature effects\.
## 6Conclusion
Per\-question uncertainty lives within a model \(r=0\.994r\\\!=\\\!0\.994\); multi\-dimensional, cross\-question uncertainty lives between them\. Temperature sampling cannot couple errors across questions; model diversity can—and two peer models outperform100100\-sample operational self\-consistency at1/401/40th the cost \(§[3\.4](https://arxiv.org/html/2607.20464#S3.SS4)\)\. To go beyond a single scalar per question, ask a different model\.
## References
- Tracy–Widom law for the extreme eigenvalues of sample correlation matrices\.Electronic Journal of Probability17,pp\. 1–32\.Cited by:[Appendix A](https://arxiv.org/html/2607.20464#A1.1.p1.11)\.
- N\. Cecere, A\. Bacciu, I\. Fernández\-Tobías, and A\. Mantrach \(2025\)Monte carlo temperature: a robust sampling strategy for LLM’s uncertainty quantification methods\.InProceedings of the 5th Workshop on Trustworthy NLP \(TrustNLP\) at NAACL,Cited by:[Appendix F](https://arxiv.org/html/2607.20464#A6.SS0.SSS0.Px2.p1.1)\.
- K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, D\. Prafulla, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. Schulman \(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§3\.2](https://arxiv.org/html/2607.20464#S3.SS2.p1.5)\.
- S\. Farquhar, J\. Kuhn, Y\. Gal,et al\.\(2024\)Detecting hallucinations in large language models using semantic entropy\.Nature\.Cited by:[§1](https://arxiv.org/html/2607.20464#S1.p1.1)\.
- K\. Hamidieh, V\. Thost, W\. Gerych, M\. Yurochkin, and M\. Ghassemi \(2025\)Complementing self\-consistency with cross\-model disagreement for uncertainty quantification\.InNeurIPS Workshop on Reliable and Responsible Foundation Models,Cited by:[Appendix F](https://arxiv.org/html/2607.20464#A6.SS0.SSS0.Px2.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart,et al\.\(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.4](https://arxiv.org/html/2607.20464#S2.SS4.SSS0.Px1.p1.1)\.
- J\. L\. Horn \(1965\)A rationale and test for the number of factors in factor analysis\.Psychometrika30\(2\),pp\. 179–185\.Cited by:[§B\.1](https://arxiv.org/html/2607.20464#A2.SS1.p1.8),[§2\.2](https://arxiv.org/html/2607.20464#S2.SS2.p1.13)\.
- E\. Hüllermeier and W\. Waegeman \(2021\)Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods\.Machine Learning110,pp\. 457–506\.Cited by:[§1](https://arxiv.org/html/2607.20464#S1.p2.1)\.
- S\. Kadavath, T\. Conerly, A\. Askell,et al\.\(2022\)Language models \(mostly\) know what they know\.arXiv preprint arXiv:2207\.05221\.Cited by:[§1](https://arxiv.org/html/2607.20464#S1.p1.1)\.
- A\. Kendall and Y\. Gal \(2017\)What uncertainties do we need in Bayesian deep learning for computer vision?\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2607.20464#S1.p2.1)\.
- E\. Kim, A\. Garg, K\. Peng, and N\. Garg \(2025\)Correlated errors in large language models\.InInternational Conference on Machine Learning \(ICML\),Note:arXiv:2506\.07962Cited by:[Appendix F](https://arxiv.org/html/2607.20464#A6.SS0.SSS0.Px2.p1.1)\.
- A\. Kipnis, K\. Voudouris, L\. M\. Schulze Buschoff, and E\. Schulz \(2025\)Metabench – a sparse benchmark of reasoning and knowledge in large language models\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2407\.12844Cited by:[Appendix F](https://arxiv.org/html/2607.20464#A6.SS0.SSS0.Px1.p1.2),[§1](https://arxiv.org/html/2607.20464#S1.p3.2),[§2\.3](https://arxiv.org/html/2607.20464#S2.SS3.p1.5)\.
- B\. Lakshminarayanan, A\. Pritzel, and C\. Blundell \(2017\)Simple and scalable predictive uncertainty estimation using deep ensembles\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2607.20464#S1.p1.1),[3rd item](https://arxiv.org/html/2607.20464#S5.I2.i3.p1.1)\.
- A\. Maimon, A\. D\. N\. Cohen, G\. Vishne, S\. Ravfogel, and R\. Tsarfaty \(2025\)IQ test for LLMs: an evaluation framework for uncovering core skills in LLMs\.arXiv preprint arXiv:2507\.20208\.Cited by:[Appendix F](https://arxiv.org/html/2607.20464#A6.SS0.SSS0.Px1.p1.2),[§1](https://arxiv.org/html/2607.20464#S1.p3.2),[§2\.3](https://arxiv.org/html/2607.20464#S2.SS3.p1.5)\.
- P\. Manakul, A\. Liusie, and M\. J\. Gales \(2023\)SelfCheckGPT: zero\-resource black\-box hallucination detection for generative large language models\.InEmpirical Methods in Natural Language Processing \(EMNLP\),Cited by:[§1](https://arxiv.org/html/2607.20464#S1.p1.1)\.
- V\. A\. Marčenko and L\. A\. Pastur \(1967\)Distribution of eigenvalues for some sets of random matrices\.Mathematics of the USSR\-Sbornik1\(4\),pp\. 457\.Cited by:[Appendix A](https://arxiv.org/html/2607.20464#A1.1.p1.11),[§1](https://arxiv.org/html/2607.20464#S1.p3.2),[§2\.2](https://arxiv.org/html/2607.20464#S2.SS2.p1.13)\.
- N\. S\. Pillai and J\. Yin \(2012\)Edge universality of correlation matrices\.Annals of Statistics40\(3\),pp\. 1737–1763\.Note:arXiv:1112\.2381Cited by:[Appendix A](https://arxiv.org/html/2607.20464#A1.1.p1.11)\.
- X\. Wang, J\. Wei, D\. Schuurmans,et al\.\(2023\)Self\-consistency improves chain of thought reasoning in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2607.20464#S1.p1.1)\.
- Z\. Wen, Z\. Liu, Z\. Tian, S\. Pan, Z\. Huang, D\. Li, and M\. Huang \(2025\)Scenario\-independent uncertainty estimation for LLM\-based question answering via factor analysis\.InProceedings of the ACM Web Conference,Cited by:[Appendix F](https://arxiv.org/html/2607.20464#A6.SS0.SSS0.Px2.p1.1),[§2\.3](https://arxiv.org/html/2607.20464#S2.SS3.p1.5)\.
- L\. H\. Yao, N\. Jarvis, T\. Zhan, S\. Ghosh, L\. Liu, and T\. Jiang \(2025\)JE\-IRT: a geometric lens on LLM abilities through joint embedding item response theory\.arXiv preprint arXiv:2509\.22888\.Cited by:[Appendix F](https://arxiv.org/html/2607.20464#A6.SS0.SSS0.Px1.p1.2),[§1](https://arxiv.org/html/2607.20464#S1.p3.2),[§2\.3](https://arxiv.org/html/2607.20464#S2.SS3.p1.5)\.
- R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. Choi \(2019\)HellaSwag: can a machine really finish your sentence?\.InAssociation for Computational Linguistics \(ACL\),Cited by:[§3\.2](https://arxiv.org/html/2607.20464#S3.SS2.p1.5)\.
- L\. Zhou, L\. Pacchiardi, F\. Martínez\-Plumed, K\. M\. Collins, others, and J\. Hernández\-Orallo \(2025\)General scales unlock AI evaluation with explanatory and predictive power\.arXiv preprint arXiv:2503\.06378\.Cited by:[Appendix F](https://arxiv.org/html/2607.20464#A6.SS0.SSS0.Px1.p1.2)\.
## Appendix ANull Prediction and Test Statistic
This appendix formalizes the asymptotic null behavior of the MP test from §[2\.2](https://arxiv.org/html/2607.20464#S2.SS2)\.
###### Assumption A\.1\(Per\-question independence\)\.
Assumecki∼Ber\(pi\)c\_\{ki\}\\sim\\mathrm\{Ber\}\(p\_\{i\}\)are mutually independent across questionsiiwithin a run and across runskk\. Operationally: independent forward passes with fresh sampling RNG per question and no shared KV cache, as in the standard eval pipeline used here\.
###### Proposition A\.2\(Null spectrum and test statistic\)\.
Under Assumption[A\.1](https://arxiv.org/html/2607.20464#A1.Thmtheorem1)withpi∈\[ε,1−ε\]p\_\{i\}\\in\[\\varepsilon,1\-\\varepsilon\]for all borderline questions: \(i\) the population Pearson correlation matrix of𝐂B\\mathbf\{C\}\_\{B\}’s columns is𝐈Nb\\mathbf\{I\}\_\{N\_\{b\}\}; \(ii\) asK→∞K\\to\\inftywithγ\\gammafixed,𝐑\\mathbf\{R\}’s empirical spectrum converges to Marchenko–Pastur on\[λ−,λ\+\]\[\\lambda\_\{\-\},\\lambda\_\{\+\}\]and the standardized top eigenvaluez=\(λ1−λ\+\)/σTWz=\(\\lambda\_\{1\}\-\\lambda\_\{\+\}\)/\\sigma\_\{\\mathrm\{TW\}\}converges in distribution to Tracy–WidomF1F\_\{1\}, with edge scaleσTW=\(1\+γ\)\(1\+1/γ\)1/3K−2/3\\sigma\_\{\\mathrm\{TW\}\}=\(1\+\\sqrt\{\\gamma\}\)\(1\+1/\\sqrt\{\\gamma\}\)^\{1/3\}K^\{\-2/3\}\.zzis therefore the appropriate null test; the unbuffered count\|\{i:λi\>λ\+\}\|\|\\\{i:\\lambda\_\{i\}\>\\lambda\_\{\+\}\\\}\|has a non\-vanishingO\(1\)O\(1\)limit \(sinceℙ\(TW1\>0\)\\mathbb\{P\}\(\\mathrm\{TW\}\_\{1\}\>0\)is bounded away from zero\) and is reported as a coarse summary only\.
###### Sketch\.
Independence givesCov\(cki,ckj\)=0\\mathrm\{Cov\}\(c\_\{ki\},c\_\{kj\}\)=0fori≠ji\\neq jandpi\(1−pi\)\>0p\_\{i\}\(1\-p\_\{i\}\)\>0on borderline questions, so the population Pearson correlation is𝐈Nb\\mathbf\{I\}\_\{N\_\{b\}\}\. By Hoeffding with a union bound over theNNcandidate questions,supi\|fi−pi\|→0\\sup\_\{i\}\|f\_\{i\}\-p\_\{i\}\|\\to 0a\.s\. asK→∞K\\to\\infty, so the data\-driven borderline set\{i:fi∈\(ε,1−ε\)\}\\\{i:f\_\{i\}\\in\(\\varepsilon,1\-\\varepsilon\)\\\}coincides with its population counterpart with probability→1\\to 1\. For \(ii\), center bypip\_\{i\}\(harmless because sample Pearson correlation is shift\-invariant\) to obtain mean\-zero bounded entries with column\-heterogeneous variancepi\(1−pi\)p\_\{i\}\(1\-p\_\{i\}\); this fitsPillai and Yin \([2012](https://arxiv.org/html/2607.20464#bib.bib19)\)\(MP bulk; Theorem 1\.1 allows column\-dependent variances and sub\-exponential entries, extendingMarčenko and Pastur,[1967](https://arxiv.org/html/2607.20464#bib.bib10)\) andBaoet al\.\([2012](https://arxiv.org/html/2607.20464#bib.bib20)\)\(Tracy–Widom edge for sample correlation matrices\)\. ∎
## Appendix BSupplementary Analyses
### B\.1Parallel Analysis
As a complementary test to MP, we apply Horn’s parallel analysis\(Horn,[1965](https://arxiv.org/html/2607.20464#bib.bib9)\)\. We generate 100 surrogateK×NbK\\\!\\times\\\!N\_\{b\}binary matrices with independent columns matched to𝐂B\\mathbf\{C\}\_\{B\}in shape and per\-question pass rate, and compute each surrogate’s eigenvalues\. For each rankjj, we compareλj\(𝐑\)\\lambda\_\{j\}\(\\mathbf\{R\}\)to the 95th percentile ofλj\\lambda\_\{j\}across surrogates\. The reported PA count is the largestrrsuch thatλj\\lambda\_\{j\}exceeds its surrogate threshold for everyj≤rj\\\!\\leq\\\!r\(the classical contiguous rule\)\. PA agrees with MP on every model and benchmark reported in the main text\.
### B\.2Robustness to borderline thresholdε\\varepsilon
The main text usesε=0\.05\\varepsilon\\\!=\\\!0\.05to define borderline questions\. Figure[A1](https://arxiv.org/html/2607.20464#A2.F1)sweepsε\\varepsilonacross\{0\.01,0\.02,0\.05,0\.10,0\.15,0\.20\}\\\{0\.01,0\.02,0\.05,0\.10,0\.15,0\.20\\\}and confirms thedwithin≤1d\_\{\\text\{within\}\}\\\!\\leq\\\!1conclusion is not an artifact of this choice\.
Figure A1:Robustness to borderline threshold\.MP noise marginλ\+−λmax\\lambda\_\{\+\}\\\!\-\\\!\\lambda\_\{\\max\}plotted against borderline thresholdε∈\{0\.01,0\.02,0\.05,0\.10,0\.15,0\.20\}\\varepsilon\\in\\\{0\.01,0\.02,0\.05,0\.10,0\.15,0\.20\\\}for the five MMLU models \(K=100K\\\!=\\\!100\)\. Positive margin: no signal above the MP noise edge \(dwithin=0d\_\{\\text\{within\}\}\\\!=\\\!0\); negative margin: at least one above\-noise eigenvalue\. Grey band: Tracy–Widom±1σ\\pm 1\\sigma\(averaged across models perε\\varepsilon\), the typical MP\-null fluctuation scale of the margin\. Four of five models \(Qwen, Mistral, Phi\-3, Llama\-3\) stay strictly in the noise regime across the fullε\\varepsilon\-range\. SmolLM2 sits in the signal band throughout \(margin−0\.26\-0\.26to−0\.01\-0\.01\) but remains within≈1σTW\{\\approx\}\\,1\\,\\sigma\_\{\\mathrm\{TW\}\}of zero atε∈\{0\.05,0\.20\}\\varepsilon\\\!\\in\\\!\\\{0\.05,0\.20\\\}: at most one above\-noise dimension\. The claimdwithin≤1d\_\{\\text\{within\}\}\\\!\\leq\\\!1is not sensitive to the choice ofε\\varepsilon\.
### B\.3Shannon Effective Rank: Definition and Extended Analysis
#### Definition\.
The Shannon effective rank of the correctness correlation matrix isdacross=exp\(H\)d\_\{\\text\{across\}\}\\\!=\\\!\\exp\(H\), whereH=−∑jλ^jlogλ^jH\\\!=\\\!\-\\sum\_\{j\}\\hat\{\\lambda\}\_\{j\}\\log\\hat\{\\lambda\}\_\{j\}andλ^j=λj/∑ℓλℓ\\hat\{\\lambda\}\_\{j\}\\\!=\\\!\\lambda\_\{j\}/\\sum\_\{\\ell\}\\lambda\_\{\\ell\}are the normalized eigenvalues\.
#### Ceiling saturation\.
Under any heterogeneous\-rate Bernoulli null,exp\(H\)\\exp\(H\)saturates near theM−1M\\\!\-\\\!1ceiling \(the−1\-1comes from column\-demeaning: centered columns lie in the\(M−1\)\(M\{\-\}1\)\-dimensional hyperplane orthogonal to𝟏M\\mathbf\{1\}\_\{M\}\)\. Empirically, the matched\-difficulty null forM=24M\\\!=\\\!24gives𝔼\[exp\(H\)\]≈22\.3\\mathbb\{E\}\[\\exp\(H\)\]\\\!\\approx\\\!22\.3\(min22\.122\.1\)\. Within\-model values reach6464–87%87\\%of theK−1K\\\!\-\\\!1ceiling \(Table[2](https://arxiv.org/html/2607.20464#S3.T2)\), indistinguishable from the across\-model81%81\\%\.exp\(H\)\\exp\(H\)therefore does not discriminate the regimes that MP counting separates cleanly\.
## Appendix CSelective Prediction: Full Results
### C\.1Method
#### Target\.
For each ofN=500N\\\!=\\\!500MMLU questions we predict whether Qwen2\.5\-7B\-Instruct answers correctly atτ=0\\tau\\\!=\\\!0\(greedy decoding\)\. The base accuracy is0\.7740\.774\(387/500387/500correct\), so the trivial constant predictor achieves AUPR≈0\.774\\approx 0\.774and AUROC0\.50\.5\.
#### Operational SCK\(deployable\)\.
For eachK∈\{2,5,10,20,50,100\}K\\in\\\{2,5,10,20,50,100\\\}we drawKKstochastic samples from Qwen\-7B \(τ=1\.0\\tau\\\!=\\\!1\.0, the same passes used in §[3\.1](https://arxiv.org/html/2607.20464#S3.SS1)\) and define the confidence as the fraction of samples agreeing with the modal answer:p^i=maxa1K∑k=1K𝟏\[aki=a\]\\hat\{p\}\_\{i\}=\\max\_\{a\}\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\mathbf\{1\}\[a\_\{ki\}=a\]\. This is the standard self\-consistency confidence; it does not require ground\-truth labels and is therefore deployable\.
#### Oracle SCK\(not deployable, reported for completeness\)\.
The oracle confidence is the fraction of theKKsamples that match the gold answer:p^ioracle=1K∑kcki\\hat\{p\}\_\{i\}^\{\\text\{oracle\}\}=\\frac\{1\}\{K\}\\sum\_\{k\}c\_\{ki\}\. This is an upper bound on any sample\-based confidence and requires test\-time label access; we report it atK=2K\\\!=\\\!2andK=100K\\\!=\\\!100only \(Table[A1](https://arxiv.org/html/2607.20464#A3.T1)\), to bracket what self\-consistency*could*achieve in principle\.
#### Ensemblek\(deployable\)\.
Forkkother models\{ℳ1,…,ℳk\}\\\{\\mathcal\{M\}\_\{1\},\\ldots,\\mathcal\{M\}\_\{k\}\\\}run once atτ=0\\tau\\\!=\\\!0, the confidence is the fraction whose answer independently matches Qwen\-7B’sτ=0\\tau\\\!=\\\!0answer:p^iens=1k∑j=1k𝟏\[aj\(qi\)=aQwen\(qi\)\]\\hat\{p\}\_\{i\}^\{\\text\{ens\}\}=\\frac\{1\}\{k\}\\sum\_\{j=1\}^\{k\}\\mathbf\{1\}\[a\_\{j\}\(q\_\{i\}\)=a\_\{\\text\{Qwen\}\}\(q\_\{i\}\)\]\. This needs no ground truth—only Qwen\-7B’s own greedy answer and thekkexternal models’ greedy answers—so it is deployable today\.
#### Cost\.
We report inference cost in*Qwen\-7B\-equivalent forward passes*: a single greedy pass of aBB\-parameter model costsB/7B/7units\. SCKon Qwen\-7B costsKKunits; Ensemblekcosts∑j=1kBj/7\\sum\_\{j=1\}^\{k\}B\_\{j\}/7\. This isolates the parameter\-FLOPs comparison from sequence\-length and serving\-stack effects that vary across deployments\.
#### Selection rules\.
*Family\-diverse*: descend byτ=0\\tau\\\!=\\\!0MMLU accuracy, accepting one model per distinct family \(the 11 families enumerated in §[2\.4](https://arxiv.org/html/2607.20464#S2.SS4)\)\.*Top\-kkby accuracy*: descend byτ=0\\tau\\\!=\\\!0MMLU accuracy without family deduplication\. We use family\-diverse selection in the main figure to avoid trivial wins from stacking near\-identical models; §[C\.3](https://arxiv.org/html/2607.20464#A3.SS3)shows the family\-diversity bonus is small but real\.
#### Metrics\.
AUROC \(rank\-correctness of confidence\) and AUPR \(area under precision\-recall, sensitive to the positive base rate of0\.7740\.774\)\. We report both but emphasize AUROC because it is invariant to the positive prevalence and thus directly comparable across setups with different base accuracies\.
### C\.2Full AUROC/AUPR Table
Table[A1](https://arxiv.org/html/2607.20464#A3.T1)reports every predictor at every budget we swept, with cost in Qwen\-7B\-equivalent forward passes\. Two patterns stand out: \(i\) operational SCKimproves slowly withKK\(\+0\.20\+0\.20AUROC over a50×50\\\!\\timescost range\) and plateaus well below the cheapest ensemble—the marginal gain fromK=50K\\\!=\\\!50toK=100K\\\!=\\\!100is only\+0\.011\+0\.011AUROC; \(ii\) EnsemblekPareto\-dominates operational SCKat every cost\. Top\-kkselection slightly beats family\-diverse selection at matchedkk\(because it concentrates parameter mass on the best models\); family\-diverse selection costs fewer parameters perkkand remains above the SC frontier\.
Table A1:Full selective\-prediction results: AUROC and AUPR for predicting Qwen2\.5\-7B’sτ=0\\tau\\\!=\\\!0correctness on 500 MMLU questions\. Cost is in Qwen\-7B\-equivalent forward passes \(parameters/77B\)\. The oracle SCKrows are upper bounds that require label access at inference time and are shown for context only\. Operational SCKuses agreement\-with\-modal\-answer confidence; Ensemblekuses agreement with Qwen\-7B’sτ=0\\tau\\\!=\\\!0answer overkkother models\. Ensemble rows are cumulative: risingkkadds the next\-highest\-τ=0\\tau\\\!=\\\!0\-MMLU\-accuracy model under each selection rule \(family\-diverse: next distinct family; top\-kk: next model overall\)\. Thek=1k\\\!=\\\!1family\-diverse member is Llama\-3\.1\-8B \(the top non\-Qwen model\)\.
### C\.3Family\-Diversity Ablation
Does ensemble gain require diverse architectures, or is any collection of three independent runs enough? We compare twok=3k\\\!=\\\!3ensembles selected from the 24\-model pool:
- •*Same\-family*\(Qwen2\.5\-0\.5B, 1\.5B, 3B\): AUROC0\.8110\.811, AUPR0\.9090\.909, cost0\.710\.71Qwen\-7B units\.
- •*Different\-family*\(Llama\-3\.1\-8B, Gemma\-2\-9B, Phi\-3\-mini\): AUROC0\.8390\.839, AUPR0\.9240\.924, cost3\.003\.00Qwen\-7B units\.
Family diversity contributes\+0\.028\+0\.028AUROC—real but modest, on top of a much larger gap between either ensemble and operational SC100\(0\.7120\.712\)\. Even three same\-family small Qwens, at1/1401/140of SC100’s cost, beatK=100K\\\!=\\\!100operational SC by\+0\.099\+0\.099AUROC\. This is the strongest available reading of the dimensionality gap as a deployment claim: the cross\-model signal is so much richer than the within\-model signal that even within\-*family*\-but\-across\-*scale*disagreement clears the SC100bar by a wide margin\.
## Appendix DStrong Local Preferences, Zero Global Coherence
A question is*borderline*here as in §[2\.1](https://arxiv.org/html/2607.20464#S2.SS1): its pass rate over theK=100K\\\!=\\\!100samples lies in\(0\.05,0\.95\)\(0\.05,0\.95\)\. When Qwen2\.5\-7B errs on such a question, it picks the same wrong choice87\.4%87\.4\\%of the time \(mean across borderline questions,K=100K\\\!=\\\!100\)\. Yet these preferences are independent across questions: Cramér’s V and mutual\-information permutation tests find no pairwise associations above the 99th percentile of the null, across all five models \(1,2421\{,\}242borderline questions,200200permutations per model\)\. Locally committed, globally incoherent\.
## Appendix EPairwise Model Correlations
Fig\.[A2](https://arxiv.org/html/2607.20464#A5.F2)expands the across\-model pairwise\-correlation summary in §[3\.3](https://arxiv.org/html/2607.20464#S3.SS3)\(meanr=0\.35r\\\!=\\\!0\.35, range0\.010\.01–0\.630\.63\) to the full distribution over all\(242\)=276\\binom\{24\}\{2\}\\\!=\\\!276model pairs, split by same\-family vs\. cross\-family\.
Figure A2:Pairwise model correlations\.All\(242\)=276\\binom\{24\}\{2\}\\\!=\\\!276model\-pairs on MMLU \(τ=0\\tau\\\!=\\\!0\), sorted by Pearsonrrand split by same\-family \(red,n=30n\\\!=\\\!30\) vs\. cross\-family \(grey,n=246n\\\!=\\\!246\)\. Side panel: kernel density with mean barsr¯=0\.41\\bar\{r\}\\\!=\\\!0\.41\(same\) and0\.340\.34\(cross\)\. Same\-family pairs trend higher \(Cohen’sd=0\.58d\\\!=\\\!0\.58, KSp<10−3p\\\!<\\\!10^\{\-3\}\), but the same\-family and cross\-family densities overlap substantially\. The top\-3 pairs byrrare annotated with their model names\.
## Appendix FExtended Related Work
#### Across\-model dimensionality\.
Cross\-model competence dimensionality has been studied via 1D item response theory \(IRT\), with a single factor capturing79%79\\%of variance on5,0005\{,\}000\+ models\(Kipniset al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib1)\); 8\-factor analysis on 60 models\(Maimonet al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib2)\); low\-dimensional joint embeddings\(Yaoet al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib3)\); and 18 cognitive rubrics\(Zhouet al\.,[2025](https://arxiv.org/html/2607.20464#bib.bib4)\)\. We complement these with the first structural measurement of*within\-model*dimensionality across stochastic samples\.
#### Self\-consistency structure\.
Hamidiehet al\.\([2025](https://arxiv.org/html/2607.20464#bib.bib5)\)showed empirically that cross\-model disagreement captures uncertainty self\-consistency misses; we find the analogous gap structurally in the correctness covariance\.Kimet al\.\([2025](https://arxiv.org/html/2607.20464#bib.bib6)\)studied pairwise error correlations across 350\+ models; we measure the full factor structure\.Wenet al\.\([2025](https://arxiv.org/html/2607.20464#bib.bib12)\)applied factor analysis to within\-question answer\-choice variation, not the cross\-question covariance we analyze\.Cecereet al\.\([2025](https://arxiv.org/html/2607.20464#bib.bib16)\)proposed Monte Carlo temperature; none tested here produces above\-noise structure\. MP is standard in portfolio theory; its application to within\-model LLM uncertainty is, to our knowledge, novel in this context\.Similar Articles
Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs
This paper introduces a validity-diversity framework attributing diversity collapse in LLMs to order and shape miscalibration during decoding, validated across 14 language models.
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
This paper proposes a framework to test whether LLM estimates obey statistical self-consistency (law of total probability) across subpopulations, finding widespread violations and the 'macro fallacy' where fine-grained estimates align better with human data.
Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles
This paper investigates whether aggregating probability estimates from multiple LLMs exhibits a wisdom-of-crowds effect, finding that learned aggregators outperform individual models and that training cutoff contamination is a pervasive confound in such evaluations.
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
This paper audits five diversity measures for LLM ensembles, finding that their associations with majority-vote gain are heavily entangled with model capability and are unstable after controlling for capability. The only robust signal is a modest residual pairwise co-failure association.
More Is Not More: What Matters for Diversity in LLM Opinions?
A factorial experiment reveals that persona detail does not monotonically increase LLM opinion diversity; interaction architectures explore non-overlapping opinion regions; low-cost interventions like temperature scaling have negligible effects.