Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees
摘要
This paper theoretically characterizes reward hacking in evaluator ensembles using covariance geometry, proving common-mode error is not identifiable from judge scores alone and bounding overstatement in best-of-K selection.
查看缓存全文
缓存时间: 2026/08/11 08:08
# Covariance Geometry and Finite-Search Guarantees
Source: [https://arxiv.org/html/2608.08002](https://arxiv.org/html/2608.08002)
## Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite\-Search Guarantees
Ibne Farabi Shihab211footnotemark:1 1Department of Computer Science, Kalinga Institute of Industrial Technology 2Department of Computer Science, Iowa State University ishihab@iastate\.eduCorresponding author:ishihab@iastate\.edu\.
###### Abstract
Language\-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality\. We characterize this failure through the covariance geometry of evaluator ensembles\. For calibrated judges, the ensemble mean retains common\-mode error along the all\-ones direction, whereas cross\-judge disagreement captures only orthogonal error\. Consequently, disagreement can be high despite robust aggregation, or low while shared response\-dependent errors persist\. We prove that common\-mode error is not identifiable from internal judge scores alone\. Under a joint sub\-Gaussian model, we bound best\-of\-KKselection overstatement and target\-quality regret, extending the guarantees to predictably adaptive search under conditional calibration\. The resulting search terms scale aslogK\\sqrt\{\\log K\}and are asymptotically tight for Gaussian projected errors\. We further show that noisy quality proxies introduce artificial rank\-one covariance without changing disagreement, and propose a bounded two\-anchor Bernstein certificate for finite\-search error and regret\. Fixed\-seed Gaussian stress tests over 120\(J,ρ,K\)\(J,\\rho,K\)configurations and real\-model audits validate the theory while revealing the limits of disagreement\-based diagnostics under increasing search pressure\.
Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite\-Search Guarantees
Fariya Afrin1††thanks:Equal contribution\.and Ibne Farabi Shihab211footnotemark:1††thanks:Corresponding author:ishihab@iastate\.edu\.1Department of Computer Science, Kalinga Institute of Industrial Technology2Department of Computer Science, Iowa State Universityishihab@iastate\.edu
## 1Introduction
Learned reward models and language\-model judges have become central to reinforcement learning from human feedback, direct preference optimization, reranking, and automatic evaluation\(Christianoet al\.,[2017](https://arxiv.org/html/2608.08002#bib.bib2); Stiennonet al\.,[2020](https://arxiv.org/html/2608.08002#bib.bib13); Ouyanget al\.,[2022](https://arxiv.org/html/2608.08002#bib.bib10); Rafailovet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib11)\)\. Their scalability comes with a structural vulnerability\. Once a generator is optimized against a fixed evaluator, it can increase the evaluator’s score by exploiting imperfections in the proxy rather than by improving the quality that the proxy was intended to represent\. This behavior is commonly described as reward hacking, evaluator gaming, or reward overoptimization\(Amodeiet al\.,[2016](https://arxiv.org/html/2608.08002#bib.bib1); Gaoet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib6); Skalseet al\.,[2022](https://arxiv.org/html/2608.08002#bib.bib24)\)\.
An ensemble is a natural defense\. If different judges make different mistakes, averaging their scores can prevent a response from succeeding by exploiting only one member\. Empirical work confirms that ensembles often mitigate overoptimization, especially when their members differ before reward\-model fine\-tuning\(Costeet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib3); Eisensteinet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib4)\)\. Yet the same studies also expose a limit: judges trained from related data, architectures, and preference signals can share the errors that optimization discovers\. Adding more correlated judges then creates additional votes without adding commensurate information\(Kimet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib20); Goelet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib21); Kohli,[2026](https://arxiv.org/html/2608.08002#bib.bib22)\)\.
This paper makes the connection between error dependence and optimization pressure explicit\. The key object is not the raw number of judges but the projection of their joint error onto the aggregation direction\. For a uniform ensemble, this direction is the all\-ones vector\. The error component parallel to it changes the ensemble mean and can therefore be exploited by search\. The orthogonal component makes judges disagree but cancels under averaging\. This geometric view recovers the familiar equal\-variance formula
Var\(e¯\)=σ2\(1J\+J−1Jρ\),\\mathrm\{Var\}\(\\bar\{e\}\)=\\sigma^\{2\}\\left\(\\frac\{1\}\{J\}\+\\frac\{J\-1\}\{J\}\\rho\\right\),\(1\)but it also shows why that identity alone is not a reward\-hacking result: an operational guarantee must concern the candidate selected by the proxy and the target\-quality loss induced by that selection\.
We therefore analyze a generator that drawsKKcandidates and selects the one with the highest ensemble score\. Under a joint sub\-Gaussian error model, we bound the selected response’s proxy overstatement and its regret relative to the candidate with the highest target quality\. The bound separates prompt\-level offsets, which change absolute scores but not within\-prompt selection, from response\-dependent common\-mode errors, which remain exploitable after averaging\. It also makes the role of the search budget explicit\. The guarantee is an upper bound of orderlogK\\sqrt\{\\log K\}in general\. For independent Gaussian projected errors, classical extreme\-value theory makes this order tight for𝔼maxk≤Ke¯k\\mathbb\{E\}\\max\_\{k\\leq K\}\\bar\{e\}\_\{k\}; it does not by itself establish tightness for selected\-response overstatement or target\-quality regret under arbitrary target qualities\. We do not infer an exact empirical scaling law from the upper bound alone\. A conditional version covers finite adaptive proposal sequences, making explicit the calibration obligation created when search reacts to earlier evaluator outputs\.
The same geometry clarifies what disagreement measures\. Disagreement is often used for active learning and uncertainty sampling\(Seunget al\.,[1992](https://arxiv.org/html/2608.08002#bib.bib12); Houlsbyet al\.,[2011](https://arxiv.org/html/2608.08002#bib.bib8); Galet al\.,[2017](https://arxiv.org/html/2608.08002#bib.bib5)\), including for language reward models\(Gleave and Irving,[2022](https://arxiv.org/html/2608.08002#bib.bib7)\)\. It is informative about the orthogonal, judge\-specific component of error, but it is blind to an error shared by all judges\. This limitation is compounded by a measurement problem: true response quality is rarely observed, so empirical studies often subtract one stronger judge as a proxy\. We prove that independent noise in this quality proxy appears as a rank\-one common mode in every estimated judge error\. A two\-anchor cross\-covariance construction removes this particular contamination and separates the theoretical object from an artifact of the audit instrument\.
Our contribution is consequently narrower than introducing ensembles, correlated\-error analysis, or disagreement sampling\. Gleave and Irving report that ensemble\-based active learning for language reward models does not outperform random selection\(Gleave and Irving,[2022](https://arxiv.org/html/2608.08002#bib.bib7)\); Eisenstein et al\. show that shared reward\-model errors survive ensembling\(Eisensteinet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib4)\); Kohli quantifies effective independence in discrete judge panels\(Kohli,[2026](https://arxiv.org/html/2608.08002#bib.bib22)\); CARE models latent confounders for static judge aggregation\(Zhaoet al\.,[2026](https://arxiv.org/html/2608.08002#bib.bib25)\); and concurrent work shows that pairwise correlation cannot identify the all\-model co\-failure tail in discrete orchestration\(Chen,[2026](https://arxiv.org/html/2608.08002#bib.bib28)\)\. Against this background, we provide an optimization\-aware analysis for continuous evaluator scores\. Specifically, we characterize the decomposition of ensemble error into common\-mode and disagreement components and prove that the former is not identifiable from internal judge scores alone; derive finite and predictably adaptive search guarantees for selected\-response overstatement and target\-quality regret; and show how a noisy quality anchor induces a spurious rank\-one common mode together with a finite\-sample two\-anchor correction\. We evaluate these results using controlled Gaussian stress tests, a best\-of\-KKreal\-model audit, a multi\-family finite\-search audit over five eligible ensemble models from three families with a separately held\-out Llama\-family proxy excluded from every ensemble, two generators \(§[5\.2](https://arxiv.org/html/2608.08002#S5.SS2)\), an anchor\-sensitivity and verifier\-backed evaluation \(§[5\.3](https://arxiv.org/html/2608.08002#S5.SS3)\), and an exploratory DPO pilot reported in Appendix[G](https://arxiv.org/html/2608.08002#A7)\. These experiments measure the quantities predicted by the theory and assess its practical implications, without claiming universality beyond the evaluated settings\.
## 2Related Work
Reward overoptimization has been measured under best\-of\-nn, policy optimization, and direct alignment algorithms\(Gaoet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib6); Rafailovet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib15)\)\.Complementary work detects proxy gaming through invariance\-based evaluator stress tests that separate sensitivity to exploitable features from content\-driven improvements under controlled perturbations\(Akteret al\.,[2026](https://arxiv.org/html/2608.08002#bib.bib33)\)\. Whereas that work diagnoses whether proxy\-score gains arise from exploitable sensitivities, we characterize which components of correlated evaluator error survive ensemble aggregation and how best\-of\-KKsearch amplifies them\. Ensemble defenses include mean aggregation, worst\-case aggregation, uncertainty penalties, and weight averaging\(Costeet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib3); Raméet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib16)\)\. Eisenstein et al\. show that diversity introduced before reward\-model fine\-tuning is more useful than diversity introduced only during fine\-tuning, while common qualitative failures remain exploitable by every ensemble member\(Eisensteinet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib4)\)\. Our search analysis is complementary: it does not propose that the mean is the optimal aggregator, but characterizes how error in a chosen aggregation direction is amplified by a finite candidate budget\. The directly adjacent decision\-level audit ofLandesberg \([2026](https://arxiv.org/html/2608.08002#bib.bib31)\)shows on a5,0005\{,\}000\-prompt Chatbot Arena best\-of\-44benchmark that global judge–label correlation badly overstates within\-prompt selection value, because global agreement is dominated by prompt\-level baseline effects; that work diagnoses the score\-versus\-decision gap for a*single*judge, whereas our contribution is the covariance geometry of a judge*ensemble*under search—which error component averaging removes, what disagreement can and cannot certify, and a bounded variance certificate for the selected candidate—so the two are complementary rather than overlapping\. Likewise, inference\-time analyses of best\-of\-nnreward hacking and its guarantees\(Gaoet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib6); Huanget al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib32)\)treat the proxy–true\-reward gap for a given scalar proxy; we locate that gap in the common\-mode component of an ensemble’s error covariance and make it auditable\.
Disagreement is a long\-standing acquisition signal in query\-by\-committee and Bayesian active learning\(Seunget al\.,[1992](https://arxiv.org/html/2608.08002#bib.bib12); Houlsbyet al\.,[2011](https://arxiv.org/html/2608.08002#bib.bib8); Galet al\.,[2017](https://arxiv.org/html/2608.08002#bib.bib5)\)\. In the closest reward\-model study,Gleave and Irving \([2022](https://arxiv.org/html/2608.08002#bib.bib7)\)find that ensemble active learning does not beat random sampling and that estimated epistemic uncertainty is only weakly related to error\. More recent work documents correlated failures across large model families\(Kimet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib20); Goelet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib21)\), measures the effective number of votes in judge panels\(Kohli,[2026](https://arxiv.org/html/2608.08002#bib.bib22)\), and learns confounder\-aware static aggregators\(Zhaoet al\.,[2026](https://arxiv.org/html/2608.08002#bib.bib25)\)\. Concurrent work on routing and voting further shows that pairwise error correlation does not identify the probability that all models fail on the same item\(Chen,[2026](https://arxiv.org/html/2608.08002#bib.bib28)\)\. That result concerns a discrete co\-failure tail; our covariance is likewise not presented as a tail certificate\. The distinction here is between static aggregation accuracy and optimization\-induced selection under an explicit joint tail condition: even a small residual error can matter when a generator searches specifically for high proxy scores\.
Because no learned judge is ground truth, evaluation validity is also central\. RewardBench and JudgeBench provide preference and objective\-correctness tests for reward models and language\-model judges\(Lambertet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib9); Tanet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib14)\); self\-preference and model\-family effects further show why a held\-out judge need not be an independent quality anchor\(Panicksseryet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib18)\)\. Appendix[B](https://arxiv.org/html/2608.08002#A2)develops these connections and states the novelty boundary in greater detail\.
## 3Setup
Letx∈𝒳x\\in\\mathcal\{X\}be a prompt anda∈𝒜xa\\in\\mathcal\{A\}\_\{x\}a candidate response\. Judgej∈\{1,…,J\}j\\in\\\{1,\\ldots,J\\\}assigns a scalar scoresj\(x,a\)s\_\{j\}\(x,a\)\. Since independently trained judges can differ in scale and offset, we map their outputs to a common audited scale,
rj\(x,a\)=sj\(x,a\)−μjτj,τj\>0,r\_\{j\}\(x,a\)=\\frac\{s\_\{j\}\(x,a\)\-\\mu\_\{j\}\}\{\\tau\_\{j\}\},\\qquad\\tau\_\{j\}\>0,\(2\)where the calibration split is disjoint from search, training, and final evaluation\. A rank transform alone does not justify additive comparisons across judges; when raw outputs are not interval\-scaled, the mapping must be learned against a common calibration target, for example through held\-out isotonic calibration\.
For a fixed prompt, write the vector of calibrated scores as
𝐫\(x,a\)=q\(x,a\)𝟏\+𝐛\(x\)\+𝐞\(x,a\)\.\\mathbf\{r\}\(x,a\)=q\(x,a\)\\mathbf\{1\}\+\\mathbf\{b\}\(x\)\+\\mathbf\{e\}\(x,a\)\.\(3\)Hereq\(x,a\)q\(x,a\)is the target quality,𝐛\(x\)\\mathbf\{b\}\(x\)contains judge offsets for the prompt, and𝐞\(x,a\)\\mathbf\{e\}\(x,a\)is the response\-dependent residual error\. Expectations below are over a declared candidate distributiona∼νxa\\sim\\nu\_\{x\}, conditional onxx\. The offsets absorb the conditional error means, so𝔼νx\[𝐞\]=𝟎\\mathbb\{E\}\_\{\\nu\_\{x\}\}\[\\mathbf\{e\}\]=\\mathbf\{0\}\. A shared response\-dependent failure is part of𝐞\\mathbf\{e\}, not𝐛\\mathbf\{b\}: unlike a prompt\-constant offset, it can change which candidate wins the search\.
For weights𝐰\\mathbf\{w\}satisfying𝟏⊤𝐰=1\\mathbf\{1\}^\{\\top\}\\mathbf\{w\}=1, the aggregate score and error are
R𝐰\(x,a\)=𝐰⊤𝐫\(x,a\),z𝐰\(x,a\)=𝐰⊤𝐞\(x,a\)\.R\_\{\\mathbf\{w\}\}\(x,a\)=\\mathbf\{w\}^\{\\top\}\\mathbf\{r\}\(x,a\),\\qquad z\_\{\\mathbf\{w\}\}\(x,a\)=\\mathbf\{w\}^\{\\top\}\\mathbf\{e\}\(x,a\)\.\(4\)The main text uses the uniform mean𝐰=𝟏/J\\mathbf\{w\}=\\mathbf\{1\}/J, denoted byRR,b¯\\bar\{b\}, ande¯\\bar\{e\}\. LetΣx=Covνx\(𝐞\)\\Sigma\_\{x\}=\\mathrm\{Cov\}\_\{\\nu\_\{x\}\}\(\\mathbf\{e\}\)and let
P=I−1J𝟏𝟏⊤P=I\-\\frac\{1\}\{J\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\(5\)be the orthogonal projector away from the aggregation direction\. These two projections,𝟏𝟏⊤/J\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}/JandPP, organize the remainder of the analysis\. We construct preference data using the calibrated disagreement\-selection DPO arm described in Algorithm[1](https://arxiv.org/html/2608.08002#alg1)\.
### 4\.1The aggregation direction
We first retain the equal\-variance formulation because it gives a transparent scalar summary\.
###### Assumption 1\(Calibrated equal\-variance judge errors\)\.
For a fixed promptxx,𝔼\[ej\]=0\\mathbb\{E\}\[e\_\{j\}\]=0,0<Var\(ej\)=σ2<∞0<\\mathrm\{Var\}\(e\_\{j\}\)=\\sigma^\{2\}<\\infty, andρjk=Cov\(ej,ek\)/σ2\\rho\_\{jk\}=\\mathrm\{Cov\}\(e\_\{j\},e\_\{k\}\)/\\sigma^\{2\}\. The average pairwise correlation is
ρ=1J\(J−1\)∑j≠kρjk,−1J−1≤ρ≤1\.\\rho=\\frac\{1\}\{J\(J\-1\)\}\\sum\_\{j\\neq k\}\\rho\_\{jk\},\\qquad\-\\frac\{1\}\{J\-1\}\\leq\\rho\\leq 1\.
###### Lemma 2\(Variance of the ensemble\-mean error\)\.
Under Assumption[1](https://arxiv.org/html/2608.08002#Thmtheorem1),
Var\(e¯\)=σ2\(1J\+J−1Jρ\)\.\\mathrm\{Var\}\(\\bar\{e\}\)=\\sigma^\{2\}\\left\(\\frac\{1\}\{J\}\+\\frac\{J\-1\}\{J\}\\rho\\right\)\.\(6\)
Equation \([6](https://arxiv.org/html/2608.08002#S4.E6)\) is the continuous\-score analogue of the classical design effect\(Kish,[1965](https://arxiv.org/html/2608.08002#bib.bib26)\)\. Defining
Jeff=J1\+\(J−1\)ρJ\_\{\\mathrm\{eff\}\}=\\frac\{J\}\{1\+\(J\-1\)\\rho\}\(7\)givesVar\(e¯\)=σ2/Jeff\\mathrm\{Var\}\(\\bar\{e\}\)=\\sigma^\{2\}/J\_\{\\mathrm\{eff\}\}, directly connecting the analysis to recent effective\-vote measurements for discrete judge panels\(Kohli,[2026](https://arxiv.org/html/2608.08002#bib.bib22)\)\. The identity is exact but is not itself our novelty claim\. In the heteroskedastic case, the corresponding quantity is simply
Var\(e¯\)=1J2𝟏⊤Σx𝟏\.\\mathrm\{Var\}\(\\bar\{e\}\)=\\frac\{1\}\{J^\{2\}\}\\mathbf\{1\}^\{\\top\}\\Sigma\_\{x\}\\mathbf\{1\}\.\(8\)For a heterogeneous panel, we use the variance\-equivalent effective size
σ¯x2=1Jtr\(Σx\),Jeffhet=σ¯x2Var\(e¯\)=Jtr\(Σx\)𝟏⊤Σx𝟏\.\\bar\{\\sigma\}\_\{x\}^\{2\}=\\frac\{1\}\{J\}\\operatorname\{tr\}\(\\Sigma\_\{x\}\),\\qquad J\_\{\\mathrm\{eff\}\}^\{\\mathrm\{het\}\}=\\frac\{\\bar\{\\sigma\}\_\{x\}^\{2\}\}\{\\mathrm\{Var\}\(\\bar\{e\}\)\}=\\frac\{J\\operatorname\{tr\}\(\\Sigma\_\{x\}\)\}\{\\mathbf\{1\}^\{\\top\}\\Sigma\_\{x\}\\mathbf\{1\}\}\.\(9\)This quantity equals Equation \([7](https://arxiv.org/html/2608.08002#S4.E7)\) under equal marginal variance\. In the experiments,J^effhet\\widehat\{J\}\_\{\\mathrm\{eff\}\}^\{\\mathrm\{het\}\}is computed from the calibration\-split residual covariance exactly via Equation \([9](https://arxiv.org/html/2608.08002#S4.E9)\)—the implementation’sJ/\(1\+\(J−1\)c¯off/c¯diag\)J/\(1\+\(J\-1\)\\bar\{c\}\_\{\\rm off\}/\\bar\{c\}\_\{\\rm diag\}\), with mean off\-diagonal and diagonal entries ofΣ^cal\\widehat\{\\Sigma\}\_\{\\rm cal\}, is algebraically identical—as a single point estimate; no bootstrap interval is reported for it\.
### 4\.2Disagreement is the orthogonal component
Define calibrated cross\-judge disagreement by
D\(x,a\)=1J∑j=1J\(rj\(x,a\)−R\(x,a\)\)2=1J‖P𝐫\(x,a\)‖2\.\\begin\{split\}D\(x,a\)&=\\frac\{1\}\{J\}\\sum\_\{j=1\}^\{J\}\\left\(r\_\{j\}\(x,a\)\-R\(x,a\)\\right\)^\{2\}\\\\ &=\\frac\{1\}\{J\}\\left\\lVert P\\mathbf\{r\}\(x,a\)\\right\\rVert^\{2\}\.\\end\{split\}\(10\)
###### Proposition 3\(Exact projector decomposition\)\.
Under the setup in Section[3](https://arxiv.org/html/2608.08002#S3),
𝔼\[D\(x,a\)\]=1J‖P𝐛\(x\)‖2\+1Jtr\(PΣx\)\.\\mathbb\{E\}\[D\(x,a\)\]=\\frac\{1\}\{J\}\\left\\lVert P\\mathbf\{b\}\(x\)\\right\\rVert^\{2\}\+\\frac\{1\}\{J\}\\operatorname\{tr\}\(P\\Sigma\_\{x\}\)\.\(11\)If the calibrated prompt offsets are common across judges and Assumption[1](https://arxiv.org/html/2608.08002#Thmtheorem1)holds, this reduces to
𝔼\[D\(x,a\)\]=σ2J−1J\(1−ρ\)\.\\mathbb\{E\}\[D\(x,a\)\]=\\sigma^\{2\}\\frac\{J\-1\}\{J\}\(1\-\\rho\)\.\(12\)
The first term in Equation \([11](https://arxiv.org/html/2608.08002#S4.E11)\) is necessary: global score calibration does not automatically remove judge\-specific prompt offsets\. Once that term is controlled, disagreement measures energy orthogonal to the ensemble mean\. By contrast, Equation \([8](https://arxiv.org/html/2608.08002#S4.E8)\) measures energy parallel to the mean\. In the exchangeable positive\-correlation modelej=c\+uje\_\{j\}=c\+u\_\{j\}, suppose the centered shared componentccand centered judge\-specific componentsuju\_\{j\}are mutually independent\. The two quantities become
Var\(e¯\)\\displaystyle\\mathrm\{Var\}\(\\bar\{e\}\)=Var\(c\)\+Var\(uj\)J,\\displaystyle=\\mathrm\{Var\}\(c\)\+\\frac\{\\mathrm\{Var\}\(u\_\{j\}\)\}\{J\},\(13\)𝔼\[D\]\\displaystyle\\mathbb\{E\}\[D\]=J−1JVar\(uj\)\.\\displaystyle=\\frac\{J\-1\}\{J\}\\mathrm\{Var\}\(u\_\{j\}\)\.Thus a shared error can produce low disagreement and high ensemble risk, while decorrelated judge\-specific errors can produce high disagreement and low ensemble risk\.
The separation creates an identification boundary before any tail assumption or search analysis enters\.
###### Proposition 4\(Identification limit\)\.
For any scalar functionh\(x,a\)h\(x,a\)with𝔼νx\[h\]=0\\mathbb\{E\}\_\{\\nu\_\{x\}\}\[h\]=0, define
q\(h\)=q\+h,𝐞\(h\)=𝐞−h𝟏\.q^\{\(h\)\}=q\+h,\\qquad\\mathbf\{e\}^\{\(h\)\}=\\mathbf\{e\}\-h\\mathbf\{1\}\.\(14\)The scores and the centering convention are unchanged, while
P𝐞\(h\)=P𝐞,e¯\(h\)=e¯−h\.P\\mathbf\{e\}^\{\(h\)\}=P\\mathbf\{e\},\\qquad\\bar\{e\}^\{\(h\)\}=\\bar\{e\}\-h\.\(15\)Consequently, internal judge scores and their disagreement cannot identify response\-dependent common\-mode error without an external anchor or additional structural assumptions\.
Figure 1:Conceptual distinction between idiosyncratic and common\-mode judge errors\. Idiosyncratic errors generate disagreement and can be reduced through averaging when sufficiently decorrelated\. Common\-mode errors are shared across judges, remain after averaging, and can induce reward hacking despite apparent consensus\. The projectorPPmeasures the former, while the aggregation direction𝟏\\mathbf\{1\}measures the latter\.Figure[1](https://arxiv.org/html/2608.08002#S4.F1)illustrates the distinction between idiosyncratic disagreement and common\-mode judge error\. The former is captured by the projectorPP, whereas the latter lies along the aggregation direction𝟏\\mathbf\{1\}and therefore survives averaging\.
### 4\.3What finite search can exploit
The covariance identity controls a second moment, but reward hacking is a selection problem\. For candidatesa1,…,aKa\_\{1\},\\ldots,a\_\{K\}, define
kR\\displaystyle k\_\{R\}∈argmaxk≤KR\(x,ak\),\\displaystyle\\in\\arg\\max\_\{k\\leq K\}R\(x,a\_\{k\}\),\(16\)kq\\displaystyle k\_\{q\}∈argmaxk≤Kq\(x,ak\)\.\\displaystyle\\in\\arg\\max\_\{k\\leq K\}q\(x,a\_\{k\}\)\.
###### Assumption 5\(Joint sub\-Gaussian search errors\)\.
For a fixed promptxx, candidates are sampled from the declared search distribution\. Their centered error vectors satisfy, for every𝐭∈ℝJ\\mathbf\{t\}\\in\\mathbb\{R\}^\{J\},
𝔼exp\(𝐭⊤𝐞\)≤exp\(12𝐭⊤Γx𝐭\)\\mathbb\{E\}\\exp\(\\mathbf\{t\}^\{\\top\}\\mathbf\{e\}\)\\leq\\exp\\\!\\left\(\\frac\{1\}\{2\}\\mathbf\{t\}^\{\\top\}\\Gamma\_\{x\}\\mathbf\{t\}\\right\)\(17\)for a positive semidefinite proxy matrixΓx\\Gamma\_\{x\}\. For the uniform ensemble, letvx=J−2𝟏⊤Γx𝟏v\_\{x\}=J^\{\-2\}\\mathbf\{1\}^\{\\top\}\\Gamma\_\{x\}\\mathbf\{1\}\. Candidate errors share this conditional marginal bound; independence across candidates is sufficient but not required by the union\-bound argument\.
###### Theorem 6\(Selected overstatement and target\-quality regret\)\.
Under Assumption[5](https://arxiv.org/html/2608.08002#Thmtheorem5), the response selected by the ensemble satisfies
𝔼\[R\(x,akR\)−q\(x,akR\)\]≤b¯\(x\)\+2vxlogK\.\\mathbb\{E\}\\left\[R\(x,a\_\{k\_\{R\}\}\)\-q\(x,a\_\{k\_\{R\}\}\)\\right\]\\leq\\bar\{b\}\(x\)\+\\sqrt\{2v\_\{x\}\\log K\}\.\(18\)With probability at least1−δ1\-\\delta,
R\(x,akR\)−q\(x,akR\)≤b¯\(x\)\+2vxlog\(K/δ\)\.R\(x,a\_\{k\_\{R\}\}\)\-q\(x,a\_\{k\_\{R\}\}\)\\leq\\bar\{b\}\(x\)\+\\sqrt\{2v\_\{x\}\\log\(K/\\delta\)\}\.\(19\)The loss in target quality relative to the best searched candidate obeys
𝔼\[q\(x,akq\)−q\(x,akR\)\]\\displaystyle\\mathbb\{E\}\\left\[q\(x,a\_\{k\_\{q\}\}\)\-q\(x,a\_\{k\_\{R\}\}\)\\right\]≤22vxlogK,\\displaystyle\\leq 2\\sqrt\{2v\_\{x\}\\log K\},\(20\)q\(x,akq\)−q\(x,akR\)\\displaystyle q\(x,a\_\{k\_\{q\}\}\)\-q\(x,a\_\{k\_\{R\}\}\)≤22vxlog\(2K/δ\)\\displaystyle\\leq 2\\sqrt\{2v\_\{x\}\\log\(2K/\\delta\)\}\(21\)with probability at least1−δ1\-\\delta\.
The constant prompt offsetb¯\(x\)\\bar\{b\}\(x\)affects absolute overstatement but cancels from the within\-prompt regret\. The exploitable term is the response\-dependent error projected onto the ensemble direction\. For general sub\-Gaussian errors, Theorem[6](https://arxiv.org/html/2608.08002#Thmtheorem6)is an upper bound rather than an equality\. If the projected candidate errors are independent Gaussian variables with variancevxv\_\{x\}, classical extreme\-value theory gives𝔼maxke¯k=2vxlogK\(1\+o\(1\)\)\\mathbb\{E\}\\max\_\{k\}\\bar\{e\}\_\{k\}=\\sqrt\{2v\_\{x\}\\log K\}\(1\+o\(1\)\), so the order is asymptotically tight\(Leadbetteret al\.,[2012](https://arxiv.org/html/2608.08002#bib.bib27)\)\.
###### Corollary 7\(Error diversity\)\.
Under Assumption[1](https://arxiv.org/html/2608.08002#Thmtheorem1)and a Gaussian proxy, or a sub\-Gaussian proxy with matching covariance, the variance\-driven terms in Theorem[6](https://arxiv.org/html/2608.08002#Thmtheorem6)decrease asd=1−ρd=1\-\\rhoincreases, holdingJJ,KK, andσ\\sigmafixed\. Diversity does not remove a shared response\-dependent component or the prompt offset\.
The same theorem applies to any fixed weights𝐰\\mathbf\{w\}after replacingvxv\_\{x\}by𝐰⊤Γx𝐰\\mathbf\{w\}^\{\\top\}\\Gamma\_\{x\}\\mathbf\{w\}\. WhenΓx\\Gamma\_\{x\}is positive definite and unconstrained signed weights are allowed, the minimum\-variance weights are proportional toΓx−1𝟏\\Gamma\_\{x\}^\{\-1\}\\mathbf\{1\}\. Nonnegative weights require the corresponding simplex\-constrained quadratic program\. This observation motivates covariance\-aware ensemble construction but does not address unobserved bias by itself\.
The fixed\-distribution statement also has a finite adaptive counterpart\. It is deliberately conditional: adaptive search is safe only if calibration continues to hold after conditioning on the search history\.
###### Corollary 8\(Predictable finite adaptation\)\.
Letℱk−1\\mathcal\{F\}\_\{k\-1\}contain the prompts, candidates, scores, and random choices observed before proposalkk\. The proposal distribution foraka\_\{k\}may beℱk−1\\mathcal\{F\}\_\{k\-1\}\-measurable\. If, for everyk≤Kk\\leq Kandλ∈ℝ\\lambda\\in\\mathbb\{R\},
𝔼\[exp\(λzk\)∣ℱk−1\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\exp\(\\lambda z\_\{k\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]≤exp\(λ2vx/2\),\\displaystyle\\leq\\exp\(\\lambda^\{2\}v\_\{x\}/2\),\(22\)zk\\displaystyle z\_\{k\}=e¯\(x,ak\),\\displaystyle=\\bar\{e\}\(x,a\_\{k\}\),then all four bounds in Theorem[6](https://arxiv.org/html/2608.08002#Thmtheorem6)hold for selection from theKKadaptively proposed candidates\.
### 4\.4Noise in the quality anchor
Proposition[4](https://arxiv.org/html/2608.08002#Thmtheorem4)makes an external quality signal necessary\. Yet the covariance in Equation \([8](https://arxiv.org/html/2608.08002#S4.E8)\) is defined relative to the target quality, which is rarely observed\. Let a held\-out quality proxy beq~=q\+η\\widetilde\{q\}=q\+\\eta, and define proxy\-relative errors𝐞~=𝐫−q~𝟏−𝐛\\widetilde\{\\mathbf\{e\}\}=\\mathbf\{r\}\-\\widetilde\{q\}\\mathbf\{1\}\-\\mathbf\{b\}\.
###### Proposition 9\(Quality\-anchor contamination\)\.
Ifη\\etais centered, independent of𝐞\\mathbf\{e\}, and has varianceτ2\\tau^\{2\}, then
Cov\(𝐞~\)=Σx\+τ2𝟏𝟏⊤\.\\mathrm\{Cov\}\(\\widetilde\{\\mathbf\{e\}\}\)=\\Sigma\_\{x\}\+\\tau^\{2\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\.\(23\)Consequently, the uniform\-ensemble variance is inflated byτ2\\tau^\{2\}, whereasP𝐞~=P𝐞P\\widetilde\{\\mathbf\{e\}\}=P\\mathbf\{e\}and disagreement is unchanged\.
The contamination therefore looks exactly like an error shared by all judges\. A single stronger judge can still define a useful proxy\-relative audit, but it cannot identify true common\-mode risk without assumptions about its own error\.
###### Corollary 10\(Two\-anchor decontamination\)\.
Supposeq~\(1\)=q\+η1\\widetilde\{q\}^\{\(1\)\}=q\+\\eta\_\{1\}andq~\(2\)=q\+η2\\widetilde\{q\}^\{\(2\)\}=q\+\\eta\_\{2\}, where the anchor errors are centered and conditionally independent of each other and of𝐞\\mathbf\{e\}\. If𝐞~\(m\)=𝐫−q~\(m\)𝟏−𝐛\\widetilde\{\\mathbf\{e\}\}^\{\(m\)\}=\\mathbf\{r\}\-\\widetilde\{q\}^\{\(m\)\}\\mathbf\{1\}\-\\mathbf\{b\}, then
Cov\(𝐞~\(1\),𝐞~\(2\)\)=Σx\.\\mathrm\{Cov\}\\\!\\left\(\\widetilde\{\\mathbf\{e\}\}^\{\(1\)\},\\widetilde\{\\mathbf\{e\}\}^\{\(2\)\}\\right\)=\\Sigma\_\{x\}\.\(24\)
This cross\-covariance construction is most credible when the anchors arise from different mechanisms, such as an exact verifier and a held\-out judge, rather than two closely related language models\.
### 4\.5From estimated variance to a valid bounded tail
The covariance analysis identifies the direction that search can exploit, but covariance alone does not certify a sub\-Gaussian proxy\. We close this gap on a declared bounded score scale: one split fits the aggregator, a disjoint certification split estimates its residual variance; pairing two candidates from the same prompt removes prompt\-level offsets, two quality anchors remove the rank\-one contamination above under an explicit cross\-orthogonality condition, and Bernstein’s inequality converts the variance upper bound and the declared range into a valid search tail\.
Fix a predeclared task–generator stratumssand an aggregation ruleAAthat is fixed independently of the certification split\. For a random promptXXin this stratum and candidatea∼νXa\\sim\\nu\_\{X\}, write
Y\(X,a\)\\displaystyle Y\(X,a\)=A\(𝐫\(X,a\)\)−q\(X,a\),\\displaystyle=A\(\\mathbf\{r\}\(X,a\)\)\-q\(X,a\),\(25\)Z\(X,a\)\\displaystyle Z\(X,a\)=Y\(X,a\)−𝔼\[Y\(X,a\)∣X\]\.\\displaystyle=Y\(X,a\)\-\\mathbb\{E\}\[Y\(X,a\)\\mid X\]\.Thus𝔼\[Z∣X\]=0\\mathbb\{E\}\[Z\\mid X\]=0, and the prompt\-constant offset cancels from within\-prompt selection\.
###### Theorem 11\(Bounded two\-anchor Bernstein certificate\)\.
Let\(Xi,ai,ai′\)i=1m\(X\_\{i\},a\_\{i\},a\_\{i\}^\{\\prime\}\)\_\{i=1\}^\{m\}be independent certification triples \(XiX\_\{i\}from stratumss;ai,ai′a\_\{i\},a\_\{i\}^\{\\prime\}conditionally independent draws fromνXi\\nu\_\{X\_\{i\}\}\), independent of the subsequent search sample\. For anchorsq~\(ℓ\)=q\+η\(ℓ\)\\widetilde\{q\}^\{\(\\ell\)\}=q\+\\eta^\{\(\\ell\)\},ℓ∈\{1,2\}\\ell\\in\\\{1,2\\\}, defineU\(ℓ\)\(X,a\)=A\(𝐫\(X,a\)\)−q~\(ℓ\)\(X,a\)U^\{\(\\ell\)\}\(X,a\)=A\(\\mathbf\{r\}\(X,a\)\)\-\\widetilde\{q\}^\{\(\\ell\)\}\(X,a\)andΔi\(ℓ\)=U\(ℓ\)\(Xi,ai\)−U\(ℓ\)\(Xi,ai′\)\\Delta\_\{i\}^\{\(\\ell\)\}=U^\{\(\\ell\)\}\(X\_\{i\},a\_\{i\}\)\-U^\{\(\\ell\)\}\(X\_\{i\},a\_\{i\}^\{\\prime\}\)\. Assume, conditional onXX, that each anchor error is centered, each is uncorrelated withZZ, and the two anchor errors are mutually uncorrelated:
𝔼\[η\(ℓ\)∣X\]\\displaystyle\\mathbb\{E\}\[\\eta^\{\(\\ell\)\}\\mid X\]=0,\\displaystyle=0,𝔼\[Zη\(ℓ\)∣X\]\\displaystyle\\mathbb\{E\}\[Z\\eta^\{\(\\ell\)\}\\mid X\]=0,\\displaystyle=0,\(26\)𝔼\[η\(1\)η\(2\)∣X\]\\displaystyle\\mathbb\{E\}\[\\eta^\{\(1\)\}\\eta^\{\(2\)\}\\mid X\]=0\.\\displaystyle=0\.Suppose the declared score and anchor ranges imply\|Δi\(ℓ\)\|≤Cℓ,s\|\\Delta\_\{i\}^\{\(\\ell\)\}\|\\leq C\_\{\\ell,s\}and\|Z\|≤cs\|Z\|\\leq c\_\{s\}almost surely\. Set
v^s\\displaystyle\\widehat\{v\}\_\{s\}=12m∑i=1mΔi\(1\)Δi\(2\),\\displaystyle=\\frac\{1\}\{2m\}\\sum\_\{i=1\}^\{m\}\\Delta\_\{i\}^\{\(1\)\}\\Delta\_\{i\}^\{\(2\)\},\(27\)vsU\\displaystyle v\_\{s\}^\{\\mathrm\{U\}\}=max\{0,v^s\+C1,sC2,slog\(1/δest\)2m\}\.\\displaystyle=\\max\\\!\\left\\\{0,\\widehat\{v\}\_\{s\}\+C\_\{1,s\}C\_\{2,s\}\\sqrt\{\\frac\{\\log\(1/\\delta\_\{\\rm est\}\)\}\{2m\}\}\\right\\\}\.Then, with probability at least1−δest1\-\\delta\_\{\\rm est\}over the certification sample,
Var\(Z\)≤vsU\.\\mathrm\{Var\}\(Z\)\\leq v\_\{s\}^\{\\mathrm\{U\}\}\.\(28\)For a new random prompt fromssand anyKKsearched candidates having the same marginal candidate law, define
Bs\(K,δ;v\)=2vlog\(K/δ\)\+cs3log\(K/δ\)B\_\{s\}\(K,\\delta;v\)=\\sqrt\{2v\\log\(K/\\delta\)\}\+\\frac\{c\_\{s\}\}\{3\}\\log\(K/\\delta\)\(29\)andBs±\(K,δ;v\)=Bs\(2K,δ;v\)B\_\{s\}^\{\\pm\}\(K,\\delta;v\)=B\_\{s\}\(2K,\\delta;v\)\. IfkAk\_\{A\}maximizes the aggregate score andkqk\_\{q\}maximizes target quality, then, jointly over certification and search, with probability at least1−δest−δsearch1\-\\delta\_\{\\rm est\}\-\\delta\_\{\\rm search\},
Z\(X,akA\)\\displaystyle Z\(X,a\_\{k\_\{A\}\}\)≤Bs\(K,δsearch;vsU\),\\displaystyle\\leq B\_\{s\}\(K,\\delta\_\{\\rm search\};v\_\{s\}^\{\\mathrm\{U\}\}\),\(30\)q\(X,akq\)−q\(X,akA\)\\displaystyle q\(X,a\_\{k\_\{q\}\}\)\-q\(X,a\_\{k\_\{A\}\}\)≤2Bs±\(K,δsearch;vsU\)\.\\displaystyle\\leq 2B\_\{s\}^\{\\pm\}\(K,\\delta\_\{\\rm search\};v\_\{s\}^\{\\mathrm\{U\}\}\)\.\(31\)The probability statement is marginal over prompts within the predeclared stratum\. It becomes prompt\-conditional only when the assumptions and the variance upper bound are established separately for that prompt\.
The theorem does not require Gaussian errors and does not equate covariance with a sub\-Gaussian proxy; its linear term records the price of a bounded but otherwise unknown tail\. A tighter Bennett inversion, the declared\-constant conventions, and a Bonferroni\-simultaneous version appear with the proof in Appendix[A](https://arxiv.org/html/2608.08002#A1); other proofs are in Appendix[A](https://arxiv.org/html/2608.08002#A1)\. The assumptions, failure modes, and claim boundaries underlying these conclusions are detailed in Appendix[K](https://arxiv.org/html/2608.08002#A11)\.
## 5Empirical Audits
### 5\.1Controlled Gaussian stress test
Before using learned evaluators, we test the implementation in the model where the order of the maximum\-error envelope is known to be tight\. The parameter grid is
J\\displaystyle J∈\{1,2,4,8\},\\displaystyle\\in\\\{1,2,4,8\\\},\(32\)K\\displaystyle K∈\{2,4,8,16,32\},\\displaystyle\\in\\\{2,4,8,6,2\\\},ρ\\displaystyle\\rho∈\{0,\.25,\.50,\.65,\.75,\.90\}\.\\displaystyle\\in\\\{0,25,50,65,75,90\\\}\.For each setting, we generate50,00050\{,\}000prompts withqk∼𝒩\(0,1\)q\_\{k\}\\sim\\mathcal\{N\}\(0,1\)and
ejk=ρck\+1−ρujk,ck,ujk∼iid𝒩\(0,1\)\.e\_\{jk\}=\\sqrt\{\\rho\}\\,c\_\{k\}\+\\sqrt\{1\-\\rho\}\\,u\_\{jk\},\\qquad c\_\{k\},u\_\{jk\}\\stackrel\{\{\\scriptstyle\\mathrm\{iid\}\}\}\{\{\\sim\}\}\\mathcal\{N\}\(0,1\)\.\(33\)The fixed seed is 20260802, and nested candidate prefixes are reused acrossKK\. Across the 120 cells no Monte Carlo mean exceeds its population upper bound \(largest observed\-to\-bound ratio0\.7870\.787for the envelope,0\.5570\.557for overstatement,0\.1160\.116for regret; every MC standard error below0\.00420\.0042\)\. Figure[2](https://arxiv.org/html/2608.08002#A6.F2)also checks the quality\-anchor algebra: one noisy anchor inflates the estimated projected variance by≈τ2\\approx\\tau^\{2\}while the two\-anchor cross\-covariance stays within0\.0020\.002of the true0\.6250\.625at all noise levels\. These verify the code path and the rank\-one contamination identity under their stated model—not that real judge errors are Gaussian or real anchors independent \(full outputs: Appendix[F](https://arxiv.org/html/2608.08002#A6)\)\.
### 5\.2Multi\-family finite\-search audit
The archived multi\-family audit \(five models, three families, two generators, one panel per row; Appendix[C](https://arxiv.org/html/2608.08002#A3), Table[2](https://arxiv.org/html/2608.08002#A2.T2)\) retains three observations: the projected scale falls with panel size butJ^effhet≈1\.5\\widehat\{J\}\_\{\\rm eff\}^\{\\rm het\}\\approx 1\.5atJ=4J\{=\}4; overstatement falls withJJfaster than regret improves \(descriptive—the records audit below supplies the paired version\); the0\.810\.81–0\.990\.99plug\-in pass fractions are not coverage\. BothJ=2J\{=\}2panels are within\-DeBERTa, so no same\- versus cross\-family conclusion is supported\. The complementary two\-judge error\-envelope audit is reported in Appendix[E](https://arxiv.org/html/2608.08002#A5)\.
### 5\.3Anchor sensitivity and verifier\-backed evaluation \(summary\)
Because the audits above measure error relative to a held\-out learned judge, we additionally examine the sensitivity of the results to the choice of quality anchor by comparing two held\-out anchors from different model families \(Appendix[D](https://arxiv.org/html/2608.08002#A4)\)\. The anchors exhibit Spearman correlation0\.420\.42on GSM8K but−0\.21\-0\.21on open\-ended prompts, with corresponding anchor\-dependent regret, whereas HumanEval provides a degenerate verifiable anchor\. The empirical validation of the two\-anchor identity and its associated UCB is therefore confined to the controlled synthetic experiment\.
Table 1:Aggregation comparison on the largest eligible panel \(J=5J\{=\}5\) and search budget \(K=32K\{=\}32\): mean selected target quality and target regret with prompt\-paired bootstrap95%95\\%intervals over the120120locked test prompts\. Best\-single, penalty, covariance weights, and CARE hyperparameters are chosen without test access\.
### 5\.4All\-subset and aggregation audit
We freeze five eligible judges and enumerate every nonempty subset \(3131panels; held\-out anchors excluded\)\. Every panel receives the same tensors \(320320GSM8K prompts,Kmax=32K\_\{\\max\}\{=\}32nested candidates from a frozen SmolLM2\-360M sampler; input SHA\-256 digests released\);K∈\{2,4,8,16,32\}K\\in\\\{2,4,8,16,32\\\}uses nested prefixes, so all contrasts are prompt\-paired\. Disjoint fit/certification/test splits \(120/80/120120/80/120\) separate learned components, the variance UCB, and final metrics\. Six predeclared rules are compared \(best\-single, uniform mean, minimum, uncertainty\-penalized, Ledoit–Wolf simplex covariance weighting, CARE\-SVD at the pinned official revision; CARE marked not applicable onJ<3J\{<\}3panels rather than silently replaced\)\. Intervals resample prompts; method contrasts subtract the within\-prompt uniform mean; panel\-size contrasts average*all*\(5j\)\\binom\{5\}\{j\}panels of a size within prompt before subtracting the singleton average, so ensemble size is no longer confounded with judge identity—the confound that made the archivedJ=1J\{=\}1vs\.J=4J\{=\}4comparison descriptive\. Additional experimental\-design and reporting requirements, including cross\-family evaluation and acquisition diagnostics, are provided in Appendix[I](https://arxiv.org/html/2608.08002#A9)\. Scope: the generator solves only3%3\\%of items \(41%41\\%of prompts have a correct candidate\)—low\-prevalence contrasts\.
The paired results are informative in both directions \(Table[1](https://arxiv.org/html/2608.08002#S5.T1)\)\.*No aggregation rule separates from the uniform mean on selected target quality*atJ=5J\{=\}5,K=32K\{=\}32\(best\-single−0\.025\-0\.025\[−0\.067,0\.008\]\[\-0\.067,0\.008\]; CARE−0\.008\-0\.008\[−0\.050,0\.033\]\[\-0\.050,0\.033\]; the weighted rules\+0\.008\+0\.008\[−0\.017,0\.042\]\[\-0\.017,0\.042\]\)\. The rules separate on*proxy inflation*: best\-single overstates the selected centered score by\+0\.104\+0\.104\[0\.067,0\.144\]\[0\.067,0\.144\]versus the mean; covariance weighting is the only rule not worse \(−0\.018\-0\.018\[−0\.049,0\.009\]\[\-0\.049,0\.009\]\)\. Ensembling itself helps:J=5J\{=\}5beats singletons on selected quality \(\+0\.028\+0\.028\[0\.003,0\.058\]\[0\.003,0\.058\]\) and cuts centered overstatement \(−0\.107\-0\.107\[−0\.139,−0\.080\]\[\-0\.139,\-0\.080\]; covariance−0\.126\-0\.126\[−0\.162,−0\.094\]\[\-0\.162,\-0\.094\]\)—the common\-mode account on real records with prompt\-paired uncertainty\.
### 5\.5Real\-task bounded certificate \(summary\)
On the one predeclared non\-degenerate task \(GSM8K: exact match is the target and anchor 1, soη1≡0\\eta\_\{1\}\\equiv 0; GRM\-Llama3 is anchor 2\), Theorem[11](https://arxiv.org/html/2608.08002#Thmtheorem11)is*valid but uninformative across all855855cells*: atm=80m\{=\}80, the estimation correction \(≈0\.61\\approx 0\.61\) dominatesv^s\\widehat\{v\}\_\{s\}\(median0\.0170\.017\), every radius reaches its range cap, and the1\.01\.0pass fractions are trivial and therefore do not constitute coverage\. Achievingv^s\\widehat\{v\}\_\{s\}\-scale precision would require approximately10510^\{5\}pairs \(Appendix[M](https://arxiv.org/html/2608.08002#A13.SS0.SSS0.Px4)\)\. A three\-seed DPO pilot \(Appendix[G](https://arxiv.org/html/2608.08002#A7)\), while underpowered, provides no statistically reliable evidence for either superiority or equivalence\. The full calibrated disagreement\-selection protocol, including the matching controls and acquisition procedure, is provided in Appendix[H](https://arxiv.org/html/2608.08002#A8)\. A detailed analysis of the three\-seed DPO pilot \(Appendix[J](https://arxiv.org/html/2608.08002#A10)\) further delineates the interpretation and limitations of this result; in particular, the pilot neither establishes a benefit nor rules out a practically meaningful effect\. The corresponding seed\-level confidence interval and exact randomization\-test calculations are reported in Appendix[L](https://arxiv.org/html/2608.08002#A12)\. For visualization, see Appendix[N](https://arxiv.org/html/2608.08002#A14)\.
## 6Conclusion and Implications
Evaluator ensembles mitigate reward hacking only along directions in which their errors cancel: covariance projected onto the aggregation direction controls the error exposed to finite search; disagreement measures an orthogonal component\. This yields overstatement and regret bounds \(incl\. conditionally calibrated adaptive proposals\); only the maximum\-error envelope has a general Gaussian tightness result; a noisy proxy adds a rank\-one common mode, removable by two anchors only under explicit conditional independence\. Empirically, ensembling suppresses proxy inflation and modestly improves selected quality, no weighting rule beats the plain mean, and the real\-task certificate is valid but uninformative at audit scale\. Panels, budgets, calibration, and external quality evidence must be audited jointly\.
## Limitations
The Gaussian\-model search guarantees are conditional on a joint sub\-Gaussian proxy, or on history\-conditional control for predictable adaptation, and covariance alone does not imply either tail condition\. The bounded Bernstein certificate \(Theorem[11](https://arxiv.org/html/2608.08002#Thmtheorem11)\) removes the earlier Gaussian bridge, but it does not make the anchor assumptions automatic: its validity requires a predeclared bounded scale, a fit split independent of certification, and centered cross\-orthogonal anchor errors; the main statement is marginal over prompts in a declared stratum, not prompt\-conditional; the linear Bernstein term can make the certificate conservative, especially after simultaneous correction; and learned anchors from different families remain a sensitivity analysis unless their error relationship has independent design\-based support\. The archived real\-model experiments use covariance plug\-ins rather than certified proxy matrices\. The multi\-family audit spans five eligible ensemble judges from three families plus a held\-out Llama\-family proxy, two small generators, and200200locked test prompts, with one executed panel per row rather than all subsets; scaling to frontier\-size generators, larger panels, and non\-English tasks is unestablished\. The expectation\-scale benchmark pass fraction is as low as0\.810\.81, but this descriptive quantity is not nominal coverage because the real\-model audit does not certify a prompt\-conditional sub\-Gaussian proxy; real\-data nominal coverage of Equation \([19](https://arxiv.org/html/2608.08002#S4.E19)\) is unestablished\. Both executedJ=2J\{=\}2panels are within\-DeBERTa pairs, so no same\- versus cross\-family conclusion is available from this run, and no pairedJ=1J\{=\}1versusJ=4J\{=\}4contrast, frozen common\-mode thresholds, exact checkpoint revision hashes, or per\-prompt records are contained in the archived summaries\. Direct aggregation baselines \(best single judge, minimum, uncertainty\-penalized, simplex\-constrained covariance weighting, and CARE at the pinned official revision\) are now executed on the all\-subset records audit \(§[5\.4](https://arxiv.org/html/2608.08002#S5.SS4)\); that audit is single\-task \(GSM8K\), single\-generator, and low\-prevalence, so its contrasts do not generalize beyond that scope, and the real\-task certificate it feeds is uninformative atm=80m\{=\}80\(§[5\.5](https://arxiv.org/html/2608.08002#S5.SS5)\)\. The archivedzz\-scored runs were unbounded without a predeclared clipping rule, so no certificate value is derived from them\. The measured covariances remain proxy\-relative \(Proposition[9](https://arxiv.org/html/2608.08002#Thmtheorem9)\); the two\-anchor validation mitigates but does not eliminate this, and it shows the anchors themselves disagree beyond sign outside verifiable domains—GSM8K supplies the only non\-degenerate verifiable anchor here, since the1\.51\.5B generator passes zero HumanEval suites\. The DPO study contains only three paired training seeds and one reported external judge; it is underpowered for equivalence, and no human evaluation was available\. A power\-determined multi\-seed study with a predeclared smallest effect of interest remains necessary\. Corollary[8](https://arxiv.org/html/2608.08002#Thmtheorem8)covers finite predictable proposals only when conditional calibration is justified; it does not cover fully co\-evolving policy training in which that condition can fail\.
## 7Ethics Statement
This work studies failures of automated evaluation and does not introduce human\-subject data\. Nevertheless, reward\-hacking diagnostics can be dual use: the same analyses that identify evaluator weaknesses could help an adversary target them\.
## References
- S\. Akter, I\. F\. Shihab, and A\. Sharma \(2026\)Detecting proxy gaming in rl and llm alignment via evaluator stress tests\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 10554–10583\.Cited by:[§2](https://arxiv.org/html/2608.08002#S2.p1.5)\.
- D\. Amodei, C\. Olah, J\. Steinhardt, P\. Christiano, J\. Schulman, and D\. Mané \(2016\)Concrete problems in ai safety\.arXiv preprint arXiv:1606\.06565\.Cited by:[§1](https://arxiv.org/html/2608.08002#S1.p1.1)\.
- G\. Bennett \(1962\)Probability inequalities for the sum of independent random variables\.Journal of the American Statistical Association57\(297\),pp\. 33–45\.Cited by:[§A\.1](https://arxiv.org/html/2608.08002#A1.SS1.SSS0.Px6.p2.3)\.
- J\. Chen \(2026\)When does combining language models help? a co\-failure ceiling on routing, voting, and mixture\-of\-agents across 67 frontier models\.arXiv preprint arXiv:2606\.27288\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p5.1),[§1](https://arxiv.org/html/2608.08002#S1.p6.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1)\.
- P\. F\. Christiano, J\. Leike, T\. Brown, M\. Martic, S\. Legg, and D\. Amodei \(2017\)Deep reinforcement learning from human preferences\.Advances in neural information processing systems30\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p1.1)\.
- T\. Coste, U\. Anwar, R\. Kirk, and D\. Krueger \(2024\)Reward model ensembles help mitigate overoptimization\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 50905–50931\.Cited by:[Appendix K](https://arxiv.org/html/2608.08002#A11.p7.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px2.p1.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p2.1),[§2](https://arxiv.org/html/2608.08002#S2.p1.5)\.
- F\. E\. Dorner, V\. Nastl, and M\. Hardt \(2025\)Limits to scalable evaluation at the frontier: llm as judge won’t beat twice the data\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 26467–26491\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p2.1)\.
- J\. Eisenstein, C\. Nagpal, A\. Agarwal, A\. Beirami, A\. D’Amour, D\. Dvijotham, A\. Fisch, K\. Heller, S\. Pfohl, D\. Ramachandran,et al\.\(2023\)Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking\.arXiv preprint arXiv:2312\.09244\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px2.p2.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p2.1),[§1](https://arxiv.org/html/2608.08002#S1.p6.1),[§2](https://arxiv.org/html/2608.08002#S2.p1.5)\.
- Y\. Gal, R\. Islam, and Z\. Ghahramani \(2017\)Deep bayesian active learning with image data\.InInternational conference on machine learning,pp\. 1183–1192\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px4.p1.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p5.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1)\.
- L\. Gao, J\. Schulman, and J\. Hilton \(2023\)Scaling laws for reward model overoptimization\.InInternational conference on machine learning,pp\. 10835–10866\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px1.p2.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p1.1),[§2](https://arxiv.org/html/2608.08002#S2.p1.5)\.
- Y\. Geifman and R\. El\-Yaniv \(2017\)Selective classification for deep neural networks\.Advances in neural information processing systems30\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px5.p1.1)\.
- A\. Gleave and G\. Irving \(2022\)Uncertainty estimation for language reward models\.arXiv preprint arXiv:2203\.07472\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px2.p3.1),[§1](https://arxiv.org/html/2608.08002#S1.p5.1),[§1](https://arxiv.org/html/2608.08002#S1.p6.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1)\.
- S\. Goel, J\. Struber, I\. A\. Auzina, K\. K\. Chandra, P\. Kumaraguru, D\. Kiela, A\. Prabhu, M\. Bethge, and J\. Geiping \(2025\)Great models think alike and this undermines ai oversight\.arXiv preprint arXiv:2502\.04313\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p3.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p2.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1)\.
- W\. Hoeffding \(1963\)Probability inequalities for sums of bounded random variables\.Journal of the American statistical association58\(301\),pp\. 13–30\.Cited by:[§A\.1](https://arxiv.org/html/2608.08002#A1.SS1.SSS0.Px6.p1.10)\.
- N\. Houlsby, F\. Huszár, Z\. Ghahramani, and M\. Lengyel \(2011\)Bayesian active learning for classification and preference learning\.arXiv preprint arXiv:1112\.5745\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px4.p1.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p5.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1)\.
- A\. Huang, A\. Block, Q\. Liu, N\. Jiang, A\. Krishnamurthy, and D\. J\. Foster \(2025\)Is best\-of\-n the best of them? coverage, scaling, and optimality in inference\-time alignment\.arXiv preprint arXiv:2503\.21878\.Cited by:[§2](https://arxiv.org/html/2608.08002#S2.p1.5)\.
- E\. Kim, A\. Garg, K\. Peng, and N\. Garg \(2025\)Correlated errors in large language models\.arXiv preprint arXiv:2506\.07962\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p3.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p2.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1)\.
- L\. Kish \(1965\)Survey sampling\.John Wiley & Sons\.Cited by:[§4\.1](https://arxiv.org/html/2608.08002#S4.SS1.p2.5)\.
- G\. Kohli \(2026\)Nine judges, two effective votes: correlated errors undermine llm evaluation panels\.arXiv preprint arXiv:2605\.29800\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p4.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p2.1),[§1](https://arxiv.org/html/2608.08002#S1.p6.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.08002#S4.SS1.p2.1)\.
- N\. Lambert, V\. Pyatkin, J\. Morrison, L\. Miranda, B\. Y\. Lin, K\. Chandu, N\. Dziri, S\. Kumar, T\. Zick, Y\. Choi,et al\.\(2025\)Rewardbench: evaluating reward models for language modeling\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 1755–1797\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p1.1),[Appendix I](https://arxiv.org/html/2608.08002#A9.p4.1),[§2](https://arxiv.org/html/2608.08002#S2.p3.1)\.
- E\. Landesberg \(2026\)When llm judge scores look good but best\-of\-n decisions fail\.arXiv preprint arXiv:2603\.12520\.Cited by:[§2](https://arxiv.org/html/2608.08002#S2.p1.5)\.
- M\. R\. Leadbetter, G\. Lindgren, and H\. Rootzén \(2012\)Extremes and related properties of random sequences and processes\.Springer Science & Business Media\.Cited by:[§4\.3](https://arxiv.org/html/2608.08002#S4.SS3.p2.3)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray,et al\.\(2022\)Training language models to follow instructions with human feedback\.Advances in neural information processing systems35,pp\. 27730–27744\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p1.1)\.
- A\. Panickssery, S\. R\. Bowman, and S\. Feng \(2024\)Llm evaluators recognize and favor their own generations\.Advances in Neural Information Processing Systems37,pp\. 68772–68802\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p2.1),[§2](https://arxiv.org/html/2608.08002#S2.p3.1)\.
- R\. Rafailov, Y\. Chittepu, R\. Park, H\. Sikchi, J\. Hejna, W\. B\. Knox, C\. Finn, and S\. Niekum \(2024\)Scaling laws for reward model overoptimization in direct alignment algorithms\.Advances in Neural Information Processing Systems37,pp\. 126207–126242\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px1.p2.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px6.p1.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§2](https://arxiv.org/html/2608.08002#S2.p1.5)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.Advances in neural information processing systems36,pp\. 53728–53741\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p1.1)\.
- A\. Ramé, N\. Vieillard, L\. Hussenot, R\. Dadashi, G\. Cideron, O\. Bachem, and J\. Ferret \(2024\)Warm: on the benefits of weight averaged reward models\.arXiv preprint arXiv:2401\.12187\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px2.p2.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§2](https://arxiv.org/html/2608.08002#S2.p1.5)\.
- H\. S\. Seung, M\. Opper, and H\. Sompolinsky \(1992\)Query by committee\.InProceedings of the fifth annual workshop on Computational learning theory,pp\. 287–294\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px4.p1.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p5.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1)\.
- J\. Skalse, N\. Howe, D\. Krasheninnikov, and D\. Krueger \(2022\)Defining and characterizing reward gaming\.Advances in neural information processing systems35,pp\. 9460–9471\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px1.p2.1),[§1](https://arxiv.org/html/2608.08002#S1.p1.1)\.
- N\. Stiennon, L\. Ouyang, J\. Wu, D\. Ziegler, R\. Lowe, C\. Voss, A\. Radford, D\. Amodei, and P\. F\. Christiano \(2020\)Learning to summarize with human feedback\.Advances in neural information processing systems33,pp\. 3008–3021\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p1.1)\.
- S\. Tan, S\. Zhuang, K\. Montgomery, W\. Tang, A\. Cuadron, C\. Wang, R\. Popa, and I\. Stoica \(2025\)Judgebench: a benchmark for evaluating llm\-based judges\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 63277–63303\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p1.1),[Appendix I](https://arxiv.org/html/2608.08002#A9.p4.1),[§2](https://arxiv.org/html/2608.08002#S2.p3.1)\.
- P\. Verga, S\. Hofstatter, S\. Althammer, Y\. Su, A\. Piktus, A\. Arkhangorodsky, M\. Xu, N\. White, and P\. Lewis \(2024\)Replacing judges with juries: evaluating llm generations with a panel of diverse models\.arXiv preprint arXiv:2404\.18796\.Cited by:[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px4.p1.1)\.
- J\. Zhao, C\. Shin, T\. Huang, S\. S\. S\. N\. GNVV, and F\. Sala \(2026\)CARE: confounder\-aware aggregation for reliable llm evaluation\.arXiv preprint arXiv:2603\.00039\.Cited by:[Appendix K](https://arxiv.org/html/2608.08002#A11.p7.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px3.p6.1),[Appendix B](https://arxiv.org/html/2608.08002#A2.SS0.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2608.08002#S1.p6.1),[§2](https://arxiv.org/html/2608.08002#S2.p2.1)\.
## Appendix AProofs
#### Proof of Lemma[2](https://arxiv.org/html/2608.08002#Thmtheorem2)
Usinge¯=J−1∑jej\\bar\{e\}=J^\{\-1\}\\sum\_\{j\}e\_\{j\},
Var\(e¯\)\\displaystyle\\mathrm\{Var\}\(\\bar\{e\}\)=1J2\(∑jVar\(ej\)\+∑j≠kCov\(ej,ek\)\)\\displaystyle=\\frac\{1\}\{J^\{2\}\}\\left\(\\sum\_\{j\}\\mathrm\{Var\}\(e\_\{j\}\)\+\\sum\_\{j\\neq k\}\\mathrm\{Cov\}\(e\_\{j\},e\_\{k\}\)\\right\)=1J2\(Jσ2\+J\(J−1\)ρσ2\),\\displaystyle=\\frac\{1\}\{J^\{2\}\}\\left\(J\\sigma^\{2\}\+J\(J\-1\)\\rho\\sigma^\{2\}\\right\),which is Equation \([6](https://arxiv.org/html/2608.08002#S4.E6)\)\. Positive semidefiniteness impliesρ≥−1/\(J−1\)\\rho\\geq\-1/\(J\-1\)\. Atρ=1\\rho=1, averaging provides no variance reduction; atρ=0\\rho=0, the standard deviation falls as1/J1/\\sqrt\{J\}; and at the negative endpoint, the equal\-variance errors cancel in the mean\.
#### Proof of Proposition[3](https://arxiv.org/html/2608.08002#Thmtheorem3)
BecauseP𝟏=0P\\mathbf\{1\}=0, Equation \([3](https://arxiv.org/html/2608.08002#S3.E3)\) givesP𝐫=P𝐛\+P𝐞P\\mathbf\{r\}=P\\mathbf\{b\}\+P\\mathbf\{e\}\. Since𝔼\[𝐞\]=0\\mathbb\{E\}\[\\mathbf\{e\}\]=0,
𝔼\[D\]\\displaystyle\\mathbb\{E\}\[D\]=1J𝔼‖P𝐛\+P𝐞‖2\\displaystyle=\\frac\{1\}\{J\}\\mathbb\{E\}\\left\\lVert P\\mathbf\{b\}\+P\\mathbf\{e\}\\right\\rVert^\{2\}=1J‖P𝐛‖2\+1J𝔼\[𝐞⊤P𝐞\]\\displaystyle=\\frac\{1\}\{J\}\\left\\lVert P\\mathbf\{b\}\\right\\rVert^\{2\}\+\\frac\{1\}\{J\}\\mathbb\{E\}\[\\mathbf\{e\}^\{\\top\}P\\mathbf\{e\}\]=1J‖P𝐛‖2\+1Jtr\(PΣx\)\.\\displaystyle=\\frac\{1\}\{J\}\\left\\lVert P\\mathbf\{b\}\\right\\rVert^\{2\}\+\\frac\{1\}\{J\}\\operatorname\{tr\}\(P\\Sigma\_\{x\}\)\.IfP𝐛=0P\\mathbf\{b\}=0and all marginal variances equalσ2\\sigma^\{2\}, then
1Jtr\(PΣx\)\\displaystyle\\frac\{1\}\{J\}\\operatorname\{tr\}\(P\\Sigma\_\{x\}\)=1Jtr\(Σx\)−1J2𝟏⊤Σx𝟏\\displaystyle=\\frac\{1\}\{J\}\\operatorname\{tr\}\(\\Sigma\_\{x\}\)\-\\frac\{1\}\{J^\{2\}\}\\mathbf\{1\}^\{\\top\}\\Sigma\_\{x\}\\mathbf\{1\}=σ2−Var\(e¯\)\\displaystyle=\\sigma^\{2\}\-\\mathrm\{Var\}\(\\bar\{e\}\)=σ2J−1J\(1−ρ\)\.\\displaystyle=\\sigma^\{2\}\\frac\{J\-1\}\{J\}\(1\-\\rho\)\.
### A\.1Proof of Proposition[4](https://arxiv.org/html/2608.08002#Thmtheorem4)
Substitution gives
q\(h\)𝟏\+𝐛\+𝐞\(h\)=\(q\+h\)𝟏\+𝐛\+𝐞−h𝟏=𝐫\.q^\{\(h\)\}\\mathbf\{1\}\+\\mathbf\{b\}\+\\mathbf\{e\}^\{\(h\)\}=\(q\+h\)\\mathbf\{1\}\+\\mathbf\{b\}\+\\mathbf\{e\}\-h\\mathbf\{1\}=\\mathbf\{r\}\.Moreover,𝔼\[𝐞\(h\)\]=𝟎\\mathbb\{E\}\[\\mathbf\{e\}^\{\(h\)\}\]=\\mathbf\{0\}because both𝐞\\mathbf\{e\}andhhare centered underνx\\nu\_\{x\}\. Finally,P𝟏=0P\\mathbf\{1\}=0givesP𝐞\(h\)=P𝐞P\\mathbf\{e\}^\{\(h\)\}=P\\mathbf\{e\}, whereas uniform averaging givese¯\(h\)=e¯−h\\bar\{e\}^\{\(h\)\}=\\bar\{e\}\-h\. The observations and every statistic computed solely from them are therefore unchanged even though the common\-mode error differs\.
#### Proof of Theorem[6](https://arxiv.org/html/2608.08002#Thmtheorem6)
Letzk=e¯\(x,ak\)z\_\{k\}=\\bar\{e\}\(x,a\_\{k\}\)\. Assumption[5](https://arxiv.org/html/2608.08002#Thmtheorem5)with𝐭=λ𝟏/J\\mathbf\{t\}=\\lambda\\mathbf\{1\}/Jgives𝔼exp\(λzk\)≤exp\(λ2vx/2\)\\mathbb\{E\}\\exp\(\\lambda z\_\{k\}\)\\leq\\exp\(\\lambda^\{2\}v\_\{x\}/2\)\. For anyλ\>0\\lambda\>0,
𝔼maxkzk\\displaystyle\\mathbb\{E\}\\max\_\{k\}z\_\{k\}≤1λlog𝔼exp\(λmaxkzk\)\\displaystyle\\leq\\frac\{1\}\{\\lambda\}\\log\\mathbb\{E\}\\exp\\\!\\left\(\\lambda\\max\_\{k\}z\_\{k\}\\right\)≤1λlog∑k=1K𝔼exp\(λzk\)\\displaystyle\\leq\\frac\{1\}\{\\lambda\}\\log\\sum\_\{k=1\}^\{K\}\\mathbb\{E\}\\exp\(\\lambda z\_\{k\}\)≤logKλ\+λvx2\.\\displaystyle\\leq\\frac\{\\log K\}\{\\lambda\}\+\\frac\{\\lambda v\_\{x\}\}\{2\}\.Optimizing atλ=2logK/vx\\lambda=\\sqrt\{2\\log K/v\_\{x\}\}yields𝔼maxkzk≤2vxlogK\\mathbb\{E\}\\max\_\{k\}z\_\{k\}\\leq\\sqrt\{2v\_\{x\}\\log K\}\. Since
R\(x,akR\)−q\(x,akR\)=b¯\(x\)\+zkR≤b¯\(x\)\+maxkzk,R\(x,a\_\{k\_\{R\}\}\)\-q\(x,a\_\{k\_\{R\}\}\)=\\bar\{b\}\(x\)\+z\_\{k\_\{R\}\}\\leq\\bar\{b\}\(x\)\+\\max\_\{k\}z\_\{k\},Equation \([18](https://arxiv.org/html/2608.08002#S4.E18)\) follows\. The sub\-Gaussian tail and a union bound give
ℙ\(maxkzk\>t\)≤Kexp\(−t22vx\),\\mathbb\{P\}\\\!\\left\(\\max\_\{k\}z\_\{k\}\>t\\right\)\\leq K\\exp\\\!\\left\(\-\\frac\{t^\{2\}\}\{2v\_\{x\}\}\\right\),which proves Equation \([19](https://arxiv.org/html/2608.08002#S4.E19)\)\.
For regret, the definition ofkRk\_\{R\}gives
qkR\+zkR≥qkq\+zkq,q\_\{k\_\{R\}\}\+z\_\{k\_\{R\}\}\\geq q\_\{k\_\{q\}\}\+z\_\{k\_\{q\}\},because the prompt offset is common to all candidates\. Therefore
qkq−qkR≤zkR−zkq≤maxkzk−minkzk\.q\_\{k\_\{q\}\}\-q\_\{k\_\{R\}\}\\leq z\_\{k\_\{R\}\}\-z\_\{k\_\{q\}\}\\leq\\max\_\{k\}z\_\{k\}\-\\min\_\{k\}z\_\{k\}\.Applying the expected maximum bound to bothzkz\_\{k\}and−zk\-z\_\{k\}proves Equation \([20](https://arxiv.org/html/2608.08002#S4.E20)\)\. A union bound over both tails of allKKvariables givesmaxk\|zk\|≤2vxlog\(2K/δ\)\\max\_\{k\}\|z\_\{k\}\|\\leq\\sqrt\{2v\_\{x\}\\log\(2K/\\delta\)\}with probability at least1−δ1\-\\delta, and the range is at most twice this value, proving Equation \([21](https://arxiv.org/html/2608.08002#S4.E21)\)\.
#### Proof of Corollary[8](https://arxiv.org/html/2608.08002#Thmtheorem8)
Taking expectations in Equation \([22](https://arxiv.org/html/2608.08002#S4.E22)\) gives
𝔼exp\(λzk\)\\displaystyle\\mathbb\{E\}\\exp\(\\lambda z\_\{k\}\)=𝔼\[𝔼\{exp\(λzk\)∣ℱk−1\}\]\\displaystyle=\\mathbb\{E\}\\\!\\left\[\\mathbb\{E\}\\\{\\exp\(\\lambda z\_\{k\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\\}\\right\]≤exp\(λ2vx/2\)\.\\displaystyle\\leq\\exp\(\\lambda^\{2\}v\_\{x\}/2\)\.for every adaptively generated candidate\. The log\-sum\-exp argument in the proof of Theorem[6](https://arxiv.org/html/2608.08002#Thmtheorem6)therefore applies without independence\. Likewise,
ℙ\(zk\>t\)\\displaystyle\\mathbb\{P\}\(z\_\{k\}\>t\)=𝔼\[ℙ\(zk\>t∣ℱk−1\)\]\\displaystyle=\\mathbb\{E\}\\\!\\left\[\\mathbb\{P\}\(z\_\{k\}\>t\\mid\\mathcal\{F\}\_\{k\-1\}\)\\right\]≤exp\(−t2/2vx\)\.\\displaystyle\\leq\\exp\(\-t^\{2\}/2v\_\{x\}\)\.and the same statement holds for−zk\-z\_\{k\}\. Union bounds over the predictable sequence establish the two high\-probability conclusions\. The regret comparison remains pathwise after theKKcandidates have been generated, so its proof is unchanged\.
#### Proof of Corollary[7](https://arxiv.org/html/2608.08002#Thmtheorem7)
Letd=1−ρd=1\-\\rho\. The variance\-driven expected overstatement term is
B\(d\)=σ2logK\(1J\+J−1J\(1−d\)\)1/2\.B\(d\)=\\sigma\\sqrt\{2\\log K\}\\left\(\\frac\{1\}\{J\}\+\\frac\{J\-1\}\{J\}\(1\-d\)\\right\)^\{1/2\}\.At every interior point where the variance is positive,
∂B∂d\\displaystyle\\frac\{\\partial B\}\{\\partial d\}=−σ2logK2J−1J\\displaystyle=\-\\frac\{\\sigma\\sqrt\{2\\log K\}\}\{2\}\\frac\{J\-1\}\{J\}×\(1J\+J−1Jρ\)−1/2<0\.\\displaystyle\\qquad\\times\\left\(\\frac\{1\}\{J\}\+\\frac\{J\-1\}\{J\}\\rho\\right\)^\{\-1/2\}<0\.The same monotonicity applies to the regret bound\. Neither derivative contains the constant prompt offset, and the equal\-variance parameterization holdsσ\\sigmafixed\.
#### Proofs of Proposition[9](https://arxiv.org/html/2608.08002#Thmtheorem9)and Corollary[10](https://arxiv.org/html/2608.08002#Thmtheorem10)
The proxy\-relative error is𝐞~=𝐞−η𝟏\\widetilde\{\\mathbf\{e\}\}=\\mathbf\{e\}\-\\eta\\mathbf\{1\}\. Independence and centering give
Cov\(𝐞~\)\\displaystyle\\mathrm\{Cov\}\(\\widetilde\{\\mathbf\{e\}\}\)=Cov\(𝐞\)\+Var\(η\)𝟏𝟏⊤\\displaystyle=\\mathrm\{Cov\}\(\\mathbf\{e\}\)\+\\mathrm\{Var\}\(\\eta\)\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}=Σx\+τ2𝟏𝟏⊤\.\\displaystyle=\\Sigma\_\{x\}\+\\tau^\{2\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\.For uniform weights,J−2𝟏⊤\(τ2𝟏𝟏⊤\)𝟏=τ2J^\{\-2\}\\mathbf\{1\}^\{\\top\}\(\\tau^\{2\}\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\)\\mathbf\{1\}=\\tau^\{2\}\. In contrast,P𝐞~=P𝐞−ηP𝟏=P𝐞P\\widetilde\{\\mathbf\{e\}\}=P\\mathbf\{e\}\-\\eta P\\mathbf\{1\}=P\\mathbf\{e\}pointwise\.
With two anchors,
𝐞~\(1\)=𝐞−η1𝟏,𝐞~\(2\)=𝐞−η2𝟏\.\\widetilde\{\\mathbf\{e\}\}^\{\(1\)\}=\\mathbf\{e\}\-\\eta\_\{1\}\\mathbf\{1\},\\qquad\\widetilde\{\\mathbf\{e\}\}^\{\(2\)\}=\\mathbf\{e\}\-\\eta\_\{2\}\\mathbf\{1\}\.All mixed covariance terms vanish under the conditional independence assumptions, leaving
Cov\(𝐞~\(1\),𝐞~\(2\)\)=Cov\(𝐞\)=Σx\.\\mathrm\{Cov\}\\\!\\left\(\\widetilde\{\\mathbf\{e\}\}^\{\(1\)\},\\widetilde\{\\mathbf\{e\}\}^\{\(2\)\}\\right\)=\\mathrm\{Cov\}\(\\mathbf\{e\}\)=\\Sigma\_\{x\}\.
#### Bennett refinement, declared constants, and simultaneity
The closed\-form Bernstein radius is convenient but need not be the tightest valid use of the same variance estimate\. Leth\(u\)=\(1\+u\)log\(1\+u\)−uh\(u\)=\(1\+u\)\\log\(1\+u\)\-uandv¯=min\{vsU,cs2\}\\bar\{v\}=\\min\\\{v\_\{s\}^\{\\mathrm\{U\}\},c\_\{s\}^\{2\}\\\}\. The deterministic cap follows directly from the theorem’s assumption\|Z\|≤cs\|Z\|\\leq c\_\{s\}, which givesVar\(Z∣X\)≤𝔼\[Z2∣X\]≤cs2\\operatorname\{Var\}\(Z\\mid X\)\\leq\\mathbb\{E\}\[Z^\{2\}\\mid X\]\\leq c\_\{s\}^\{2\}; the sharper Popoviciu capcs2/4c\_\{s\}^\{2\}/4would additionally require the conditional support*width*to be at mostcsc\_\{s\}, which the theorem does not assume, so we do not use it\. Define
Ts\(K,δ;vsU\)=min\{cs,v¯csh−1\(cs2v¯logKδ\)\},T\_\{s\}\(K,\\delta;v\_\{s\}^\{\\mathrm\{U\}\}\)=\\min\\\!\\left\\\{c\_\{s\},\\frac\{\\bar\{v\}\}\{c\_\{s\}\}h^\{\-1\}\\\!\\left\(\\frac\{c\_\{s\}^\{2\}\}\{\\bar\{v\}\}\\log\\frac\{K\}\{\\delta\}\\right\)\\right\\\},\(34\)withTs=0T\_\{s\}=0whenv¯=0\\bar\{v\}=0\. Bennett’s bounded\-variance moment\-generating function bound implies that Equation \([30](https://arxiv.org/html/2608.08002#S4.E30)\) remains valid after replacingBsB\_\{s\}byTsT\_\{s\}\. If target quality has declared range widthWq,sW\_\{q,s\}, Equation \([31](https://arxiv.org/html/2608.08002#S4.E31)\) remains valid with the tighter radius
min\{Wq,s,2Ts\(2K,δsearch;vsU\)\}\.\\min\\\!\\left\\\{W\_\{q,s\},2T\_\{s\}\(2K,\\delta\_\{\\rm search\};v\_\{s\}^\{\\mathrm\{U\}\}\)\\right\\\}\.\(35\)We report the Bennett radius together with the closed\-form Bernstein radius and the deterministic cap\. The standard inequalityh\(u\)≥u2/\[2\(1\+u/3\)\]h\(u\)\\geq u^\{2\}/\[2\(1\+u/3\)\]shows pointwise that the Bennett inversion is no larger than the Bernstein relaxation, while the deterministic cap holds surely\. Thus reporting the tightest of these nested bounds introduces no post\-selection or additional failure probability\.
If the aggregate and each anchor are clipped to\[0,1\]\[0,1\]before either tuning or certification, thenC1,s=C2,s=2C\_\{1,s\}=C\_\{2,s\}=2\. If both aggregate and target quality lie in\[0,1\]\[0,1\], thencs=2c\_\{s\}=2is valid because centering a variable with range two produces an absolute deviation of at most two\. These constants must be declared rather than estimated from the observed extrema\.
ForMMpredeclared cells, choosingδest=δsearch=α/\(2M\)\\delta\_\{\\rm est\}=\\delta\_\{\\rm search\}=\\alpha/\(2M\)and applying a union bound makes Equations \([30](https://arxiv.org/html/2608.08002#S4.E30)\)–\([31](https://arxiv.org/html/2608.08002#S4.E31)\) simultaneous across allMMcells with family\-wise probability at least1−α1\-\\alpha\. Pointwise and simultaneous certificates should be reported in separate columns\.
#### Proof of Theorem[11](https://arxiv.org/html/2608.08002#Thmtheorem11)
For one certification pair, writeU\(ℓ\)=μX\+Z−η\(ℓ\)U^\{\(\\ell\)\}=\\mu\_\{X\}\+Z\-\\eta^\{\(\\ell\)\}, whereμX=𝔼\[Y∣X\]\\mu\_\{X\}=\\mathbb\{E\}\[Y\\mid X\]\. Candidate differencing removesμX\\mu\_\{X\}\. Conditional onXX, independence ofaaanda′a^\{\\prime\}together with Equation \([26](https://arxiv.org/html/2608.08002#S4.E26)\) gives
𝔼\[12Δ\(1\)Δ\(2\)∣X\]=Var\(Z∣X\)\.\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{2\}\\Delta^\{\(1\)\}\\Delta^\{\(2\)\}\\mid X\\right\]=\\mathrm\{Var\}\(Z\\mid X\)\.\(36\)Since𝔼\[Z∣X\]=0\\mathbb\{E\}\[Z\\mid X\]=0, averaging overXXyields
𝔼\[12Δ\(1\)Δ\(2\)\]=𝔼\[Var\(Z∣X\)\]=Var\(Z\)\.\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{2\}\\Delta^\{\(1\)\}\\Delta^\{\(2\)\}\\right\]=\\mathbb\{E\}\[\\mathrm\{Var\}\(Z\\mid X\)\]=\\mathrm\{Var\}\(Z\)\.\(37\)Moreover,12Δ\(1\)Δ\(2\)\\frac\{1\}\{2\}\\Delta^\{\(1\)\}\\Delta^\{\(2\)\}lies in an interval of widthC1,sC2,sC\_\{1,s\}C\_\{2,s\}\. The one\-sided Hoeffding inequality\(Hoeffding,[1963](https://arxiv.org/html/2608.08002#bib.bib30)\)therefore gives
ℙ\{Var\(Z\)\>v^s\+C1,sC2,slog\(1/δest\)2m\}≤δest,\\mathbb\{P\}\\\!\\left\\\{\\mathrm\{Var\}\(Z\)\>\\widehat\{v\}\_\{s\}\+C\_\{1,s\}C\_\{2,s\}\\sqrt\{\\frac\{\\log\(1/\\delta\_\{\\rm est\}\)\}\{2m\}\}\\right\\\}\\leq\\delta\_\{\\rm est\},\(38\)which proves Equation \([28](https://arxiv.org/html/2608.08002#S4.E28)\)\.
On this event, the classical bounded\-variable moment\-generating\-function bound implies, for0≤λ<3/cs0\\leq\\lambda<3/c\_\{s\},
log𝔼\[eλZ\]≤λ2Var\(Z\)2\(1−λcs/3\)≤λ2vsU2\(1−λcs/3\)\.\\log\\mathbb\{E\}\[e^\{\\lambda Z\}\]\\leq\\frac\{\\lambda^\{2\}\\mathrm\{Var\}\(Z\)\}\{2\(1\-\\lambda c\_\{s\}/3\)\}\\leq\\frac\{\\lambda^\{2\}v\_\{s\}^\{\\mathrm\{U\}\}\}\{2\(1\-\\lambda c\_\{s\}/3\)\}\.\(39\)The Chernoff argument consequently yields
ℙ\{Z\>2vsUt\+cst/3\}≤e−t\.\\mathbb\{P\}\\\!\\left\\\{Z\>\\sqrt\{2v\_\{s\}^\{\\mathrm\{U\}\}t\}\+c\_\{s\}t/3\\right\\\}\\leq e^\{\-t\}\.\(40\)A direct Bennett argument\(Bennett,[1962](https://arxiv.org/html/2608.08002#bib.bib29)\)instead gives
ℙ\{Z\>t\}≤exp\[−v¯cs2h\(cstv¯\)\],0≤t≤cs,\\mathbb\{P\}\\\{Z\>t\\\}\\leq\\exp\\\!\\left\[\-\\frac\{\\bar\{v\}\}\{c\_\{s\}^\{2\}\}h\\\!\\left\(\\frac\{c\_\{s\}t\}\{\\bar\{v\}\}\\right\)\\right\],\\qquad 0\\leq t\\leq c\_\{s\},\(41\)where replacing the unknown variance by its upper bound weakens the tail bound\. Inverting this expression after a union bound gives the radius in Equation \([34](https://arxiv.org/html/2608.08002#A1.E34)\)\. Applying it to both tails gives Equation \([35](https://arxiv.org/html/2608.08002#A1.E35)\)\.
For the closed\-form relaxation, a union bound overKKcandidate marginals witht=log\(K/δsearch\)t=\\log\(K/\\delta\_\{\\rm search\}\)proves Equation \([30](https://arxiv.org/html/2608.08002#S4.E30)\); independence among the searched candidates is unnecessary for this step\. Applying Equation \([40](https://arxiv.org/html/2608.08002#A1.E40)\) to bothZZand−Z\-Zwitht=log\(2K/δsearch\)t=\\log\(2K/\\delta\_\{\\rm search\}\)bounds every searched error in absolute value byBs±B\_\{s\}^\{\\pm\}\. Finally, aggregate\-score optimality gives
q\(X,akq\)−q\(X,akA\)≤Z\(X,akA\)−Z\(X,akq\),q\(X,a\_\{k\_\{q\}\}\)\-q\(X,a\_\{k\_\{A\}\}\)\\leq Z\(X,a\_\{k\_\{A\}\}\)\-Z\(X,a\_\{k\_\{q\}\}\),\(42\)so the two\-sided event proves Equation \([31](https://arxiv.org/html/2608.08002#S4.E31)\)\. Combining the estimation and search events by a final union bound completes the proof\.
## Appendix BAdditional Related Work
#### Reward overoptimization and reward hacking
RLHF fits a learned reward model to pairwise or scalar preference judgments and then optimizes a policy against that model as a proxy for the intended objective\(Christianoet al\.,[2017](https://arxiv.org/html/2608.08002#bib.bib2); Stiennonet al\.,[2020](https://arxiv.org/html/2608.08002#bib.bib13); Ouyanget al\.,[2022](https://arxiv.org/html/2608.08002#bib.bib10)\)\. The reward model is estimated from a finite, imperfectly labeled sample and must evaluate a policy distribution that changes as optimization proceeds\. Its reliability is therefore not a peripheral implementation detail: a region in which the reward model is an unreliable proxy can become the region toward which the optimized policy moves\.
Gaoet al\.\([2023](https://arxiv.org/html/2608.08002#bib.bib6)\)operationalize this failure using a large “gold” reward model as a stand\-in for human judgment\. Optimizing a smaller proxy through best\-of\-nnsampling or PPO initially improves gold reward, but continued optimization eventually decreases it\.Rafailovet al\.\([2024](https://arxiv.org/html/2608.08002#bib.bib15)\)observe an analogous pattern for direct alignment algorithms such as DPO, showing that the phenomenon is not specific to explicit reinforcement learning against a scalar reward\. We use*reward overoptimization*for this measurable degradation under a stronger evaluation signal\. We use*reward hacking*more broadly for behavior that obtains proxy reward in a way the principal would not endorse\(Skalseet al\.,[2022](https://arxiv.org/html/2608.08002#bib.bib24)\)\. Overoptimization is one empirically tractable manifestation of that broader problem\.
Our contribution is not to establish that proxy optimization can fail\. Instead, we ask which component of a learned evaluator’s error remains after an ensemble is aggregated and how a finite search budget can exploit that component\. This scope matters because it separates the existence of reward hacking, which prior work already establishes, from the narrower question of when averaging evaluators can mitigate it\.
#### Reward\-model ensembles and uncertainty
The most direct ensemble mitigation replaces a single reward model with several members and aggregates their predictions\.Costeet al\.\([2024](https://arxiv.org/html/2608.08002#bib.bib3)\)compare mean and worst\-case objectives with uncertainty weighting\. Their experiments show that conservative aggregation can substantially reduce overoptimization, particularly under label noise\. This result also shows why the mean should not be treated as the uniquely correct ensemble objective: different aggregators expose different directions of the joint error and make different bias–variance tradeoffs\.
Eisensteinet al\.\([2023](https://arxiv.org/html/2608.08002#bib.bib4)\)reach a complementary conclusion\. Reward models that perform similarly on their training distribution can be underspecified by the available preference data and assign sharply different rewards after the policy moves off distribution\. Ensembles that vary in pretraining seed transfer better under optimization than ensembles that vary only in fine\-tuning seed on top of a shared pretrained model\. Even pretraining\-diverse ensembles do not eliminate reward hacking, however, because some qualitative failures are shared by every member\.Raméet al\.\([2024](https://arxiv.org/html/2608.08002#bib.bib16)\)study weight\-space averaging as a less expensive alternative to prediction\-space ensembling\. Together, these studies establish that ensemble diversity matters empirically while also showing that nominal diversity in seeds or models does not guarantee diversity in the errors reached by optimization\.
The negative disagreement\-selection result is also not unprecedented\.Gleave and Irving \([2022](https://arxiv.org/html/2608.08002#bib.bib7)\)train bootstrap ensembles of language reward models and find that active learning driven by ensemble uncertainty does not outperform random sampling\. Their estimated epistemic uncertainty is only weakly related to model error, which they attribute in part to the similarity of members fine\-tuned from one underlying language model\. Our DPO pilot differs in both intervention and outcome: it selects model\-generated preference data and measures transfer after policy optimization\. Nevertheless, its three\-seed result should be interpreted as a scoped instance of the same broader warning, not as the first negative result for disagreement\-based acquisition\.
These works motivate rather than invalidate our analysis\. We do not propose that an ensemble mean is optimal, and we do not claim that covariance is a complete model of adversarial optimization\. We identify the projected response\-dependent error governing a fixed aggregation rule under finite search, derive selected\-response and regret guarantees, and state the information that must come from outside the ensemble before its common\-mode risk can be estimated\.
#### Language\-model judges and correlated errors
RewardBench evaluates reward models on challenging chosen–rejected triples spanning chat, reasoning, and safety, whereas JudgeBench emphasizes difficult response pairs for which preference and objective correctness can diverge\(Lambertet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib9); Tanet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib14)\)\. These benchmarks reveal distinct evaluator failure modes, but static accuracy on either benchmark is not a substitute for evaluating responses generated by a policy that has been optimized against the evaluator\.
Several studies explain why a held\-out judge should not automatically be treated as ground truth\.Panicksseryet al\.\([2024](https://arxiv.org/html/2608.08002#bib.bib18)\)document self\-preference: evaluators can recognize and favor outputs from related models even when human raters judge the outputs to be equally good\.Dorneret al\.\([2025](https://arxiv.org/html/2608.08002#bib.bib19)\)establish a complementary limitation on scalable evaluation\. When the judge is no more capable than the evaluated model, a small amount of ground\-truth data cannot in general produce an arbitrarily large reduction in evaluation sample complexity through debiasing alone\.
Correlated wrong answers make the same concern visible at the panel level\. Across more than 350 language models,Kimet al\.\([2025](https://arxiv.org/html/2608.08002#bib.bib20)\)find substantial dependence in model errors; architectural or provider diversity does not guarantee independent failures\.Goelet al\.\([2025](https://arxiv.org/html/2608.08002#bib.bib21)\)likewise connect similarity between models to correlated oversight failures and judge preferences\. These findings motivate evaluating error covariance on responses reachable by the optimized generator, rather than inferring independence from model names, providers, or checkpoint counts\.
Most directly,Kohli \([2026](https://arxiv.org/html/2608.08002#bib.bib22)\)apply the Kish design effect to a panel of nine frontier judges from seven model families and find only about two effective independent votes\. Their analysis concerns discrete panel reliability and the gap from an independent Condorcet model\. Equation \([7](https://arxiv.org/html/2608.08002#S4.E7)\) uses the same design\-effect algebra for continuous residual scores\. Our incremental step is to place the residual retained by aggregation inside an optimization\-induced selection event and bound both the selected response’s overstatement and its quality regret\.
Concurrent work byChen \([2026](https://arxiv.org/html/2608.08002#bib.bib28)\)studies a different but adjacent failure object: the probability that every member of a model pool gives the wrong discrete answer\. It proves that pairwise correlation cannot identify this all\-wrong tail and validates the distinction across a large model pool\. That result reinforces a boundary already explicit here: covariance alone is not a tail guarantee\. Our object is the continuous residual projected through a reward aggregator, and our finite\-search bounds require a joint moment\-generating\-function condition; the two analyses should not be collapsed into one correlation\-only claim\.
CARE takes a different next step\.Zhaoet al\.\([2026](https://arxiv.org/html/2608.08002#bib.bib25)\)model static judge scores through latent quality and shared confounders and prove identifiability and finite\-sample recovery without gold labels under their structural assumptions\. We do not offer a competing latent\-variable estimator\. Proposition[4](https://arxiv.org/html/2608.08002#Thmtheorem4)instead states the no\-assumption boundary: an unrestricted response\-dependent common mode cannot be identified from internal judge scores alone\. Proposition[9](https://arxiv.org/html/2608.08002#Thmtheorem9)then shows how one common audit practice biases the relevant covariance, while Corollary[10](https://arxiv.org/html/2608.08002#Thmtheorem10)gives a correction when two anchor errors satisfy explicit independence conditions\. Structured latent models and external\-anchor constructions are therefore complementary ways to cross the same identification boundary\.
#### Disagreement and active learning
Query\-by\-committee selects examples on which members of a committee disagree\(Seunget al\.,[1992](https://arxiv.org/html/2608.08002#bib.bib12)\)\. BALD gives a Bayesian interpretation in which acquisition reflects information about model parameters\(Houlsbyet al\.,[2011](https://arxiv.org/html/2608.08002#bib.bib8); Galet al\.,[2017](https://arxiv.org/html/2608.08002#bib.bib5)\)\. Panel\-based language\-model evaluation applies a related intuition:Vergaet al\.\([2024](https://arxiv.org/html/2608.08002#bib.bib17)\)show that a jury of smaller, diverse\-family evaluators can outperform a single large judge on several benchmarks\. In each case, disagreement is useful when it is coupled to the error or uncertainty that the intervention is meant to reduce\.
The reward\-hacking setting adds a specific caution\. A committee can disagree because its members make different errors, and that idiosyncratic variation is exactly what averaging can suppress\. Yet all members may also respond to the same spurious feature\. Such an error changes the ensemble mean without generating internal disagreement\. In the projector notation of Proposition[3](https://arxiv.org/html/2608.08002#Thmtheorem3), the shared component lies in the aggregation direction and is annihilated byPP\. Low disagreement is consequently evidence of internal consistency, not a certificate of external correctness\.
This distinction does not make disagreement useless\. It can identify ambiguous prompts, poorly calibrated score regions, examples on which additional labels are likely to change the committee, or cases dominated by member\-specific noise\. What it cannot do by itself is rank common\-mode errors that all judges endorse\. The value of disagreement\-based selection therefore depends on the downstream target: it may improve calibration or label efficiency without reducing the error exploited by an ensemble mean\.
#### Selective prediction and disagreement filtering
Selective prediction allows a model to abstain on low\-confidence cases, trading coverage for risk\(Geifman and El\-Yaniv,[2017](https://arxiv.org/html/2608.08002#bib.bib23)\)\. Disagreement filtering can be viewed as a form of selective evaluation in which the system refuses to trust examples on which its judges are inconsistent\. Standard risk–coverage reasoning requires the confidence score to rank the error of interest at least approximately: abstaining must remove errors faster than it removes correct predictions\.
A common\-mode evaluator failure can violate this premise\. Every judge may be confident and mutually consistent on the same wrong response\. Filtering high\-disagreement examples can then remove genuinely difficult, idiosyncratic cases while retaining the confidently wrong cases that matter most under optimization\. The issue is not merely a different point on a risk–coverage curve; the internal abstention score may be insensitive to the relevant risk component\. A valid selective\-evaluation system therefore needs an external error signal or a structural model linking disagreement to common\-mode risk\.
#### Preference optimization without explicit reward models
DPO reparameterizes preference optimization so that a policy can be trained directly from a preference dataset without an explicit online reward\-model stage\(Rafailovet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib11)\)\. Removing that stage does not remove dependence on the quality of the preference signal\. If model\-generated labels contain a correlated bias, the bias can be propagated into the policy even when no scalar reward model is optimized online\. The overoptimization results ofRafailovet al\.\([2024](https://arxiv.org/html/2608.08002#bib.bib15)\)make this connection empirical\.
We use DPO as a controlled setting in which the acquisition rule can change while the pair budget and optimization procedure remain fixed\. The pilot asks whether selecting prompts with greater calibrated disagreement changes transfer to an external judge\. It does not claim that the DPO objective itself follows from the covariance theorem, nor that a negative pilot result generalizes to other preference optimizers, larger models, or acquisition budgets\.
#### Positioning and research gap
The literature establishes a progression that constrains our novelty claim\. Proxy rewards can be overoptimized\(Gaoet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib6); Rafailovet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib15)\); reward\-model ensembles can mitigate but not eliminate this behavior\(Costeet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib3); Eisensteinet al\.,[2023](https://arxiv.org/html/2608.08002#bib.bib4); Raméet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib16)\); disagreement is a classical uncertainty and acquisition signal\(Seunget al\.,[1992](https://arxiv.org/html/2608.08002#bib.bib12); Houlsbyet al\.,[2011](https://arxiv.org/html/2608.08002#bib.bib8); Galet al\.,[2017](https://arxiv.org/html/2608.08002#bib.bib5)\); and recent work directly measures or models correlated judge errors\(Kimet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib20); Goelet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib21); Kohli,[2026](https://arxiv.org/html/2608.08002#bib.bib22); Zhaoet al\.,[2026](https://arxiv.org/html/2608.08002#bib.bib25)\)\. We therefore do not claim to introduce any of those ideas\.
The contribution is the connection between their remaining gaps in a finite\-search setting\. The projector decomposition separates the error changed by aggregation from the error observed as disagreement while retaining judge\-specific prompt offsets\. The identification result shows why internal consensus alone cannot recover the common mode\. The search theorem turns the retained direction into guarantees for the response actually selected and for its quality regret, rather than only for the maximum error among searched candidates\. Finally, the anchor analysis distinguishes true common covariance from rank\-one contamination induced by measuring all judges against the same noisy proxy\. These results provide an optimization\-aware audit framework; they do not replace conservative aggregation, latent\-confounder modeling, or external evaluation\.
Table 2:Multi\-family audit on the locked test split \(K=32K\{=\}32cells shown; brackets are marginal prompt\-bootstrap95%95\\%CIs for selected\-response overstatement\)\. Herev^\\widehat\{v\}is a pooled projected residual variance andJ^effhet\\widehat\{J\}\_\{\\rm eff\}^\{\\rm het\}is the heterogeneous variance\-equivalent panel size defined in Equation \([9](https://arxiv.org/html/2608.08002#S4.E9)\)\. “bench\. frac\.” is the descriptive fraction of prompts whose realized error envelope is below2v^logK\\sqrt\{2\\widehat\{v\}\\log K\}\. It has no nominal coverage interpretation and is not a test of Equation \([19](https://arxiv.org/html/2608.08002#S4.E19)\)\. Panels name the judges actually evaluated:D1D\_\{1\}–D3D\_\{3\}are DeBERTa checkpoints,GGis Gemma\-2B,MMis Mistral\-7B; bothJ=2J\{=\}2rows are within\-DeBERTa pairs\.
## Appendix CArchived Multi\-Family Audit in Full
A preliminary two\-judge envelope audit \(now superseded; Appendix[E](https://arxiv.org/html/2608.08002#A5)\) found the realized proxy\-relative envelope below the Gaussian plug\-in curve in every\(J,K\)\(J,K\)cell\. We subsequently scaled the evaluation to five eligible models from three families \(D1D\_\{1\}–D3D\_\{3\}OpenAssistant DeBERTa, Gemma\-2B RMGG, and Mistral\-7B RMMM\), together with a held\-out Llama\-family proxy excluded from every ensemble\. The evaluation used two generators,J∈\{1,2,4\}J\\in\\\{1,2,4\\\}, andK∈\{2,…,32\}K\\in\\\{2,\\dots,32\\\}; covariance was estimated on a calibration split, whereas all reported quantities were computed on a locked 200\-prompt test split \(Table[2](https://arxiv.org/html/2608.08002#A2.T2); artifact manifest in Appendix[M](https://arxiv.org/html/2608.08002#A13)\)\. The archived run evaluated one panel per row—\{D1\}\\\{D\_\{1\}\\\}, the within\-DeBERTa pairs, and\{D1,D3,G,M\}\\\{D\_\{1\},D\_\{3\},G,M\\\}—and therefore does not support same\- versus cross\-family comparisons; the earlier “mixed” finding is accordingly withdrawn\. The all\-subset audit of §[5\.4](https://arxiv.org/html/2608.08002#S5.SS4)addresses this limitation using prompt\-level records\.
Three observations nevertheless remain\. First, the projected scale decreases with panel size \(vv:1\.80→1\.091\.80\\\!\\to\\\!1\.09and2\.07→1\.282\.07\\\!\\to\\\!1\.28fromJ=1J\{=\}1toJ=4J\{=\}4\), whileJ^effhet≈1\.5\\widehat\{J\}\_\{\\rm eff\}^\{\\rm het\}\\approx 1\.5atJ=4J\{=\}4\(Equation \([9](https://arxiv.org/html/2608.08002#S4.E9)\)\), indicating that correlation forfeits most of the nominal averaging benefit\. Second, selected\-response overstatement atK=32K\{=\}32decreases withJJ\(1\.19→0\.801\.19\\\!\\to\\\!0\.80,1\.27→0\.691\.27\\\!\\to\\\!0\.69\), whereas target regret improves more modestly \(1\.80→1\.621\.80\\\!\\to\\\!1\.62,1\.72→1\.451\.72\\\!\\to\\\!1\.45\)\. This pattern is consistent with the theoretical separation between suppressing proxy over\-scoring and improving selection\. Because the archived summary lacks paired prompt\-level contrasts, these decreases are descriptive rather than inferential; the records audit above provides the paired analysis\. Third, the fraction of realized envelopes below the expectation\-scale benchmark2v^logK\\sqrt\{2\\widehat\{v\}\\log K\}ranges from0\.810\.81to0\.990\.99\. This quantity is a pass fraction rather than a coverage probability: it has no confidence parameter and uses an uncertifiedv^\\widehat\{v\}\. Consequently, it neither demonstrates undercoverage nor validates Theorem[6](https://arxiv.org/html/2608.08002#Thmtheorem6), and we make no nominal coverage claim for the real\-model experiments\.
Finally, we do not report a common\-mode comparison acrossJJ\. The archived rates were computed using within\-test\-cell thresholds rather than calibration\-frozen thresholds, and disagreement is identically zero atJ=1J\{=\}1\(yielding a rate of0\.250\.25by construction\)\. A frozen\-threshold recomputation requires per\-prompt records that are unavailable in the archived summary\. Artifact layout, contents, and remaining gaps are documented in Appendix[M](https://arxiv.org/html/2608.08002#A13)\.
## Appendix DAnchor Sensitivity and Verifier\-Backed Evaluation in Full
The audits above measure error against a held\-out learned judge, exactly the practice Proposition[9](https://arxiv.org/html/2608.08002#Thmtheorem9)warns about\. We therefore compare two held\-out learned anchors from different families \(Llama\-3\-8B RM, Mistral\-7B RM; fixed Qwen2\.5\-1\.5B generator,K=8K\{=\}8\) with objective verification where the task admits it—a sensitivity analysis, not an instantiation of Corollary[10](https://arxiv.org/html/2608.08002#Thmtheorem10), since family difference does not imply conditional independence\. On GSM8K \(exact\-match anchor, prevalence0\.370\.37\) the anchors correlate at Spearman0\.420\.42and rank a verified\-correct candidate first on0\.730\.73/0\.660\.66of solvable prompts \(DeBERTa ensemble:0\.620\.62\)\. On HumanEval the verifiable anchor is*degenerate*\(zero passing suites\), reported as such\. On open\-ended prompts the anchors correlate*negatively*\(Spearman−0\.21\-0\.21\) and mean selected\-response regret depends strongly on the anchor \(3\.83\.8vs6\.16\.1on GSM8K;4\.04\.0vs7\.37\.3open\-ended\)\. Only GSM8K supplies a non\-degenerate verifier\-backed target; the negative correlations reinforce the anchor\-validity concern, so the two\-anchor identity and its UCB are verified empirically only in the controlled synthetic experiment\.
## Appendix ETwo\-Judge Error\-Envelope Audit Table
Setup \(moved from §[5\.2](https://arxiv.org/html/2608.08002#S5.SS2)\): candidates from SmolLM2\-360M\-Instruct, twozz\-calibrated reward\-model judges, a stronger held\-out proxyq~\\widetilde\{q\}excluded from the ensemble, and the mean over120120prompts of the proxy\-relative envelopemaxk≤Ke¯\(x,ak\)\\max\_\{k\\leq K\}\\bar\{e\}\(x,a\_\{k\}\)—an upper bound on the selected response’s error, not the selected\-response statistic itself\. With plug\-in variancesv^J=1=0\.708\\widehat\{v\}\_\{J=1\}=0\.708andv^J=2=0\.696\\widehat\{v\}\_\{J=2\}=0\.696, the observed envelope stays below the Gaussian plug\-in curve in every\(J,K\)\(J,K\)cell\. The magnitudes are not precisely predicted—the plug\-in curves are loose and the run kept no prompt\-level intervals—and an equal\-variance fit \(σ2≈0\.84\\sigma^\{2\}\\approx 0\.84,ρ≈0\.65\\rho\\approx 0\.65\) predicts disagreement0\.1470\.147vs\. the measured0\.1760\.176: qualitatively consistent, not exact, with Proposition[9](https://arxiv.org/html/2608.08002#Thmtheorem9)explaining why the shared proxy prevents reading the fitted correlation as ground\-truth common\-mode error\.
Table 3:Finite\-search error\-envelope audit across all reported\(J,K\)\(J,K\)cells\. Observed proxy\-relative envelopes remain below the Gaussian plug\-in curves2v^logK\\sqrt\{2\\widehat\{v\}\\log K\}computed fromv^J=1=0\.708\\widehat\{v\}\_\{J=1\}=0\.708andv^J=2=0\.696\\widehat\{v\}\_\{J=2\}=0\.696\. The table supports numerical consistency with the upper envelope; it does not show equality, identify alogK\\sqrt\{\\log K\}scaling law, or isolate the causal effect of correlation from judge identity\.
## Appendix FSynthetic Stress\-Test Details
Figure 2:Fixed\-seed synthetic stress test\. Left: forρ=0\.65\\rho=0\.65, solid curves are Monte Carlo means ofmaxk≤Ke¯k\\max\_\{k\\leq K\}\\bar\{e\}\_\{k\}and dashed curves are2vlogK\\sqrt\{2v\\log K\}\. Right: one noisy anchor inflates projected variance by approximatelyτ2\\tau^\{2\}, while the cross\-covariance from independent anchors remains near the true value\. Because the data are generated from the assumed Gaussian model, this is an implementation and assumption stress test rather than real\-model validation\.The synthetic audit is fully specified by Equation \([33](https://arxiv.org/html/2608.08002#S5.E33)\)\. For each of the2424\(J,ρ\)\(J,\\rho\)settings, one draw contains50,00050\{,\}000prompts and3232candidates per prompt\. TheK∈\{2,4,8,16,32\}K\\in\\\{2,4,8,16,32\\\}conditions use nested prefixes of the same candidate tensor\. The target qualities and all common and judge\-specific error variables are independent standard Gaussians\. Ensemble selection maximizesqk\+e¯kq\_\{k\}\+\\bar\{e\}\_\{k\}; the reported quantities aremaxke¯k\\max\_\{k\}\\bar\{e\}\_\{k\},e¯kR\\bar\{e\}\_\{k\_\{R\}\}, andmaxkqk−qkR\\max\_\{k\}q\_\{k\}\-q\_\{k\_\{R\}\}\. Their reference curves are the corresponding terms in Theorem[6](https://arxiv.org/html/2608.08002#Thmtheorem6)withv=ρ\+\(1−ρ\)/Jv=\\rho\+\(1\-\\rho\)/J\.
The anchor experiment uses200,000200\{,\}000independent scalar projected errors withJ=4J=4andρ=0\.5\\rho=0\.5\. Two independent Gaussian anchor errors are added at each value ofτ2∈\{0,\.25,\.50,1\.00\}\\tau^\{2\}\\in\\\{0,\.25,\.50,1\.00\\\}\. The single\-anchor estimator is the sample variance of one proxy\-relative projected residual; the corrected estimator is the sample cross\-covariance between the two proxy\-relative residuals\. With true projected variance0\.6250\.625, the single\-anchor estimate rises from0\.6210\.621to0\.8710\.871,1\.1211\.121, and1\.6261\.626asτ2\\tau^\{2\}increases, while the two\-anchor estimates remain0\.6210\.621,0\.6220\.622,0\.6230\.623, and0\.6230\.623\. The experiment uses no fitted parameters and no discarded runs\. The accompanying simulation script and two CSV outputs reproduce Figure[2](https://arxiv.org/html/2608.08002#A6.F2)from seed 20260802\.
This stress test is narrower than a model evaluation\. Gaussian draws make the sub\-Gaussian proxy exact, while the construction makes the anchor errors independent\. The stress test can reveal implementation mistakes, arithmetic inconsistencies, or a mismatch between a theorem and its claimed statistic\. It cannot establish that learned judges satisfy the assumptions, which is why the real\-model audit and the missing confirmatory experiments remain separately identified\.
## Appendix GExploratory DPO Pilot
This appendix retains the original controlled DPO experiment, while narrowing its conclusion to the available evidence\. Two training arms start from the same Qwen2\.5\-1\.5B\-Instruct checkpoint and use the same DPO budget over three rounds\. One arm selects prompts with high calibrated cross\-judge disagreement; the other selects a size\-matched random sample\. Preference labels in both arms come from the calibrated ensemble mean\. Evaluation uses identical held\-out prompts and a held\-out judge\. The full acquisition protocol and the additional matching controls originally specified for a confirmatory study appear in Appendix[H](https://arxiv.org/html/2608.08002#A8)\.
Across three paired training seeds and 120 evaluation prompts, the held\-out\-score differences are0\.0350\.035,0\.0330\.033, and0\.0020\.002\. Their mean is0\.0230\.023with standard error0\.0110\.011\. The exact sign\-flip test gives one\-sidedp=0\.125p=0\.125and two\-sidedp=0\.250p=0\.250; the previously reportedp=0\.13p=0\.13was the rounded one\-sided value\. Attinterval with two degrees of freedom is\[−0\.023,0\.069\]\[\-0\.023,0\.069\]\. The exploratory three\-seed DPO pilot is summarized in Table[4](https://arxiv.org/html/2608.08002#A7.T4)\.
Table 4:Seed\-level DPO pilot\. All three estimates are positive, but the interval is too wide to establish either a reliable improvement or practical equivalence\. Prompts are repeated measurements within a trained model and are not treated as independent experimental units\.The pilot yields no statistically reliable evidence that high\-disagreement selection improves transfer at this scale\. It also does not show that the intervention has no practically meaningful effect\. The pooled\-prompt analysis reportedp=0\.41p=0\.41, but it is secondary because prompts sharing a trained model are not independent experimental units\. A confirmatory negative result would require a predeclared smallest effect of interest, an equivalence test, and a seed count chosen by power analysis\. Appendix[J](https://arxiv.org/html/2608.08002#A10)develops the possible interpretations without treating planned controls as completed results\. The accompanying seed\-summary artifact records the three values and the exact analysis; candidate\-level and training logs were not present in the supplied artifact and remain necessary for full reproduction\.
## Appendix HDisagreement\-Selection Protocol
The intervention is calibrated disagreement selection\. At each round, the current generator samples at least two responses for every prompt under a fixed decoding configuration\. Each training judge scores the same responses, and all transformations that place those scores on a common scale are fitted on a calibration split disjoint from prompt selection, policy training, and final evaluation\. Reusing the acquisition or evaluation prompts for calibration would allow the score mapping itself to absorb part of the intervention\.
For responseaka\_\{k\}to promptxx, the algorithm computes the calibrated disagreement in Equation \([10](https://arxiv.org/html/2608.08002#S4.E10)\)\. The prompt\-level acquisition score is the average of this quantity across itsKrK\_\{r\}sampled responses\. Averaging avoids making the acquisition decision depend on one anomalous sample, while keeping the response\-level scores available for diagnostic analysis\. The disagreement arm selects prompts from a prespecified top quantile\. The random arm draws a size\-matched subset from the same eligible pool, and the matched controls draw from that pool subject to the constraints in Table[6](https://arxiv.org/html/2608.08002#A9.T6)\.
Preference construction is held fixed after acquisition\. The calibrated ensemble mean labels each response pair, a prespecified margin removes ties and near\-ties, and resampling restores the same number of retained pairs in every arm\. Thus the intended intervention is the distribution of prompts entering preference construction, not the number of labels, the confidence of those labels, or the amount of optimization\. In later rounds, candidates are generated by the current policy rather than the initial checkpoint, so the acquisition rule is evaluated under the distribution shift it helps create\.
Algorithm 1Calibrated disagreement\-selection DPO arm0:initial policy
πθ0\\pi\_\{\\theta\_\{0\}\}; prompt pool
𝒳train\\mathcal\{X\}\_\{\\mathrm\{train\}\}
0:calibrated judges
1,…,J1,\\ldots,J; rounds
TT; candidates per prompt
KrK\_\{r\}
0:pairs per round
MM; selection rule
SS
1:set
θ←θ0\\theta\\leftarrow\\theta\_\{0\}
2:for
t=1,…,Tt=1,\\ldots,Tdo
3:sample
KrK\_\{r\}responses per prompt using the fixed decoding configuration
4:compute calibrated judge scores and per\-prompt disagreement
D^t\(x\)\\widehat\{D\}\_\{t\}\(x\)
5:select exactly
MMpreference pairs according to
SSand its matching constraints
6:label each pair by the calibrated ensemble mean
7:discard pairs below the tie margin and resample to retain
MMpairs
8:update
θ\\thetawith the fixed DPO objective, reference model, and optimizer schedule
9:endfor
10:returnfinal policy
πθ\\pi\_\{\\theta\}
The random comparison is the minimum causal control\. It measures the effect of performing the same amount of preference optimization without targeting disagreement\. Pair counts, generated tokens, judge calls, DPO steps, decoding settings, and tie\-handling rules must be identical\. Otherwise, an apparent acquisition gain could be explained by a larger data or compute budget\.
Length matching addresses a more specific alternative\. High\-disagreement prompts or responses may be longer, more likely to reach truncation limits, or more likely to elicit verbose outputs\. Matching prompt and response\-length distributions tests whether the result is caused by those properties rather than evaluator uncertainty\. Initial\-reward matching distinguishes disagreement from ordinary hard\-example mining: the control should reproduce the initial ensemble\-score distribution without conditioning on cross\-judge spread\.
Update\-norm matching addresses the possibility that selected examples merely induce larger optimization steps\. It matches either gradient norms during training or the resulting parameter displacement under a declared convention\. Finally, comparing mean aggregation with a minimum or another conservative rule asks whether any benefit is specific to the mean rather than to the selected data\. These controls isolate different mechanisms and are not interchangeable\.
Only the random comparison is present in the supplied pilot outputs\. The length\-, reward\-, and update\-norm\-matched arms and the conservative\-aggregation comparison therefore remain requirements for a confirmatory study\. They are retained here as part of the full experimental design, not presented as completed evidence\.
## Appendix IExperimental Design and Reporting Requirements
The intended unit of randomization is the training seed\. Each seed initializes all arms from the same checkpoint and uses identical train, calibration, and evaluation splits\. Randomness in candidate sampling, pair construction, data ordering, and optimization should be paired where doing so does not leak the acquisition decision\. Across arms, the number of preference pairs, DPO steps, generated tokens, and judge queries must be matched unless a separate cost\-normalized comparison is explicitly reported\.
This distinction determines the statistical analysis\. Evaluation prompts are repeated measurements within a trained policy, not independent replications of training\. Seed\-level paired differences are therefore the primary experimental units\. Prompt bootstrap or permutation intervals can describe uncertainty conditional on the trained policies, but they cannot replace inference across independently trained seeds\. Hyperparameters and stopping rules must be selected without examining the final paired outcomes\. For reproducibility, the minimum information that must be reported for the generator, decoding procedure, judges, DPO configuration, selection procedure, evaluation protocol, and statistical analysis is summarized in Table[5](https://arxiv.org/html/2608.08002#A9.T5)\. The checklist specifies the required fields, while the archival configuration should provide the actual values for each field\.
Table 5:Minimum information required for a reproducible implementation\. A checklist is not a substitute for reporting the actual values; every entry must be instantiated in the archival configuration\.The primary outcome should be selected\-response target quality, accompanied by the training ensemble score and the proxy–anchor gap\. Reporting all three distinguishes an intervention that genuinely improves external quality from one that merely increases the score of the evaluators used to generate its labels\. When candidates are retained, the analysis should also report regret relative to the best candidate under the external anchor, matching the quantity in Theorem[6](https://arxiv.org/html/2608.08002#Thmtheorem6)\.
A held\-out judge is itself imperfect and should not be described as ground truth\. Verifiable mathematical answers, unit\-tested executable code, exact structured\-output checks, or other task\-specific validators provide scalable anchors when human evaluation is unavailable\. RewardBench\-style preference triples and JudgeBench\-style correctness pairs can complement these outcomes because they probe different evaluator failures\(Lambertet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib9); Tanet al\.,[2025](https://arxiv.org/html/2608.08002#bib.bib14)\)\. Static benchmark accuracy nevertheless cannot replace evaluation of the optimized policy’s own generations, where the distribution shift induced by search or DPO is present\.
Table 6:Prespecified controls for attributing an improvement to calibrated disagreement selection\. The current pilot executes only the first row; the table defines the complete confirmatory design and does not report additional results\.### I\.1Evaluation across judge and generator families
The two\-judge pilot cannot separate the effect of adding a member from the identity of the member added\. A confirmatory matrix should therefore vary ensemble size and the source of diversity\. At minimum, it should distinguish members that share a pretrained checkpoint but use different reward\-model fine\-tuning seeds, members with different pretraining seeds, and members from different architectures or providers\. The relevant comparison is measured error covariance on the policy’s candidate distribution, not the nominal category alone\.
The generator side should vary both model family and search pressure\. For best\-of\-KKaudits, the same base candidates should be reused across nested budgets whenever possible, so changes withKKare not confounded by independent sampling noise\. For trained policies, evaluation should include the initial generator and every final arm under identical decoding\. ReportingJ∈\{1,2,4,…\}J\\in\\\{1,2,4,\\ldots\\\}and severalKKvalues permits direct comparison of nominal ensemble size, effective size, and the search budget that acts on the residual projected error\.
Every judge subset must be calibrated without access to its test responses\. Covariance, disagreement, external\-anchor error, and judge\-call cost should be reported for each subset\. This design can reveal whether cross\-family diversity genuinely reduces the aggregation direction or merely increases orthogonal disagreement\. It also prevents one favorable pair of judges from being treated as evidence for a general ensemble\-size law\.
### I\.2Diagnostics connecting acquisition to error
The first diagnostic plots calibrated disagreement against error under an independent anchor\. In addition to a correlation coefficient, it should display the full joint distribution and the low\-disagreement, high\-error quadrant that exposes common\-mode failures\. Stratification by prompt length, response length, task, and initial ensemble score tests whether the acquisition statistic is acting as a proxy for a simpler observable property\.
The second diagnostic compares the projected covariance in Equation \([8](https://arxiv.org/html/2608.08002#S4.E8)\) with selected\-response overstatement across judge subsets and search budgets\. The unit of analysis and uncertainty interval must be declared: prompt\-level intervals should be clustered by prompt, while comparisons between trained policies must retain seed\-level replication\. A third calibration diagnostic should compare raw and calibrated score distributions, because a disagreement trend that disappears after placing judges on a common scale is a units artifact rather than evidence of epistemic diversity\.
All diagnostics should report judge calls, generated tokens, and wall\-clock or accelerator cost\. A method that reduces error only by multiplying evaluator cost may still be useful, but its advantage is different from a covariance\-efficient ensemble\. Diagnostic plots are mechanism checks; they do not replace the primary external\-quality comparison\.
## Appendix JInterpreting the DPO Pilot
The three seed\-level differences are positive, but their interpretation is constrained by the size of the experiment\. The mean difference of0\.0230\.023is a point estimate, not evidence that the acquisition rule reliably improves transfer\. Thettinterval crosses zero and includes effects that would be practically meaningful in either direction, while the exact two\-sided randomization test has only four attainable levels below one\. Conversely, failure to reject zero is not evidence that the methods are equivalent\. An equivalence claim would require a prespecified smallest effect of interest and an interval lying entirely inside the corresponding equivalence region\.
The result is compatible with the covariance analysis but is not predicted by it\. The theory shows that disagreement observes the component orthogonal to mean aggregation\. It does not say that examples with large orthogonal error contain no useful training signal\. Such examples could still improve calibration, expose ambiguous preferences, diversify response styles, or produce harder preference pairs\. Whether those benefits transfer to an external evaluator is an empirical question determined by the data distribution, label mechanism, model capacity, and optimization budget\.
Several alternative mechanisms remain unresolved in the current pilot\. High\-disagreement prompts may be longer or harder, may receive lower initial ensemble scores, or may generate larger DPO updates\. A held\-out reward model can also share errors with the training judges, which would make transfer under that judge an incomplete measure of common\-mode failure\. The matched controls and independent anchors specified above are needed to distinguish these possibilities\.
The defensible conclusion is therefore deliberately local: under the reported checkpoint, three\-round DPO budget, acquisition procedure, and three paired seeds, the experiment does not provide statistically reliable evidence that disagreement selection improves held\-out judge score\. It neither establishes a benefit nor rules out a practically important one, and it does not support a universal claim about disagreement\-based data selection\.
## Appendix KClaim Boundaries and Failures
The formal claim is deliberately conditional\. Under calibrated scores, a declared candidate distribution, and a joint sub\-Gaussian error proxy, the error projected onto a fixed aggregation direction controls selected\-response overstatement and target\-quality regret\. Covariance alone does not imply the required tail condition\. Heavy\-tailed or adversarially dependent errors may make a covariance plug\-in severely optimistic even when the second moment is estimated accurately\. The empirical curves in Table[3](https://arxiv.org/html/2608.08002#A5.T3)must therefore remain working\-model diagnostics unless the proxy matrix or tail condition is independently justified\.
Finite search is another substantive boundary\. The union\-bound argument allows dependence among theKKcandidate errors\. Corollary[8](https://arxiv.org/html/2608.08002#Thmtheorem8)also allows proposal distributions to react predictably to earlier evaluator outputs, but only when the projected error remains conditionally centered and sub\-Gaussian after conditioning on that history\. Repeated policy training can invalidate precisely this condition by moving toward regions in which calibration fails\. The corollary is therefore a finite adaptive\-search result, not a certificate for unrestricted policy–evaluator co\-adaptation\.
The primary structural failure is a response\-dependent common error\. If every judge rewards the same spurious feature, the ensemble can agree confidently while assigning the wrong ordering to candidates\. Proposition[4](https://arxiv.org/html/2608.08002#Thmtheorem4)shows that the resulting common mode cannot be recovered from judge scores alone without additional structure\. A prompt\-only calibration offset is different: it changes absolute overstatement but cancels from within\-prompt selection\. Keeping these notions separate avoids attributing an optimization failure to a constant that search cannot exploit\.
Score calibration introduces a second boundary\. Reward models can differ in offset, scale, nonlinearity, and saturation\. Without a common interval interpretation, one judge can dominate the mean and disagreement can reflect units rather than uncertainty\. Global calibration also does not guarantee that judge\-specific prompt offsets vanish, which is why Proposition[3](https://arxiv.org/html/2608.08002#Thmtheorem3)retains the‖P𝐛‖2\\left\\lVert P\\mathbf\{b\}\\right\\rVert^\{2\}term\. Calibration must be audited on held\-out data and, when possible, by task and score region rather than only in aggregate\.
The external quality anchor is not exempt from evaluator error\. A single held\-out reward model can share training data, architecture, style preferences, or blind spots with the training ensemble\. Even independent anchor noise contributes the rank\-one term in Proposition[9](https://arxiv.org/html/2608.08002#Thmtheorem9); correlated anchor noise can be more difficult to separate\. The two\-anchor correction is valid only under its conditional independence assumptions\. Closely related language models should not be called independent anchors merely because they use different checkpoint names\.
Selection itself induces distribution shift\. Prompts selected this way may differ from random prompts in length, difficulty, topic, adversarial structure, initial reward, or the size of the resulting gradient\. If disagreement beats random selection but not a length\- or reward\-matched control, the appropriate conclusion is that the observed gain is explained by that matched property\. Table[6](https://arxiv.org/html/2608.08002#A9.T6)is designed to separate these mechanisms, but only its random comparison was executed in the supplied pilot\.
Finally, the analysis does not establish that the ensemble mean is optimal or that disagreement selection should improve preference optimization\. Worst\-case, uncertainty\-weighted, covariance\-aware, and latent\-confounder\-aware objectives remain competing approaches\(Costeet al\.,[2024](https://arxiv.org/html/2608.08002#bib.bib3); Zhaoet al\.,[2026](https://arxiv.org/html/2608.08002#bib.bib25)\)\. The theoretical contribution characterizes a chosen aggregation rule and an identification problem; the empirical pilot tests one acquisition intervention\. Neither should be generalized into a universal ranking of ensemble designs\.
## Appendix LStatistical Details for the DPO Pilot
For paired seed differencesd=\(0\.035,0\.033,0\.002\)d=\(0\.035,0\.033,0\.002\), the sample mean isd¯=0\.02333\\bar\{d\}=0\.02333, the sample standard deviation is0\.018500\.01850, and the standard error is0\.010680\.01068\. Usingt0\.975,2=4\.303t\_\{0\.975,2\}=4\.303gives
d¯±t0\.975,2SE\(d\)=\[−0\.0226,0\.0693\]\.\\bar\{d\}\\pm t\_\{0\.975,2\}\\operatorname\{SE\}\(d\)=\[\-0\.0226,0\.0693\]\.An exact sign\-flip randomization test enumerates the eight possible sign assignments\. One assignment is at least as large as the observed positive mean, giving one\-sidedp=1/8=0\.125p=1/8=0\.125; two assignments are at least as extreme in absolute value, giving two\-sidedp=2/8=0\.250p=2/8=0\.250\. With three seeds, exactpp\-values are necessarily coarse\. The prompt\-pooledp=0\.41p=0\.41is retained only as a secondary descriptive analysis because it treats repeated prompts within a trained model as if they supplied independent training replicates\.
A confirmatory study should specify a smallest effect of interest before looking at final outcomes\. Demonstrating no practically important benefit then requires the confidence interval to fall inside the equivalence region, not merely a failure to reject zero\. The pilot interval is too wide for such a conclusion\.
## Appendix MArtifact Manifest and Model Inventory
#### Model inventory
Table[7](https://arxiv.org/html/2608.08002#A13.T7)lists every model with its role and aggregation eligibility\. The executed runs loaded checkpoints by repository name without pinning immutable revision hashes; the hashes therefore cannot be reconstructed honestly after the fact and are marked as not recorded\. A rerun that pins revisions is required before hash\-level reproducibility can be claimed\.
Table 7:Model inventory\. Revision hashes were not pinned by the executed runs and are not recorded; the DeBERTa checkpoints are one family \(three checkpoints\), so the eligible pool spans three families \(DeBERTa, Gemma, Mistral\)\. Panels actually evaluated:\{D1\}\\\{D\_\{1\}\\\},\{D1,D3\}\\\{D\_\{1\},D\_\{3\}\\\},\{D1,D2\}\\\{D\_\{1\},D\_\{2\}\\\},\{D1,D3,G,M\}\\\{D\_\{1\},D\_\{3\},G,M\\\}; an all\-subsets design \(J=2J\{=\}2: three DeBERTa\-within\-family pairs and all seven cross\-family pairs;J=4J\{=\}4: the three three\-family panels\) was not executed\.
#### Aggregation baselines \(executed\)
For a comparison of aggregation rules on identical candidate tensors, calibration split, test prompts, and search budgets, the predeclared baseline set is: each individual judge with any “best single” designation chosen on calibration data only; the uniform mean; minimum aggregationminjrj\\min\_\{j\}r\_\{j\}; uncertainty\-penalized aggregationr¯−λD\\bar\{r\}\-\\lambda\\sqrt\{D\}withλ\\lambdachosen on calibration data only; simplex\-constrained covariance weighting𝐰^∈argmin𝐰≥0,𝟏⊤𝐰=1𝐰⊤Σ^cal𝐰\\widehat\{\\mathbf\{w\}\}\\in\\arg\\min\_\{\\mathbf\{w\}\\geq 0,\\mathbf\{1\}^\{\\top\}\\mathbf\{w\}=1\}\\mathbf\{w\}^\{\\top\}\\widehat\{\\Sigma\}\_\{\\rm cal\}\\mathbf\{w\}; and the closest executable confounder\-aware aggregation baseline \(CARE\-SVD\) at its public implementation revision\. All six are now executed in the all\-subset records audit \(§[5\.4](https://arxiv.org/html/2608.08002#S5.SS4)\) with selected target quality, regret, selected and centered overstatement, disagreement, judge calls, and prompt\-paired intervals against the uniform mean, released per \(panel,KK, method\) cell in the artifact CSVs\.
#### Certification protocol \(deferred from §[5\.5](https://arxiv.org/html/2608.08002#S5.SS5)\)\.
For the one predeclared non\-degenerate task \(GSM8K; exact match is target and anchor 1, soη1≡0\\eta\_\{1\}\\equiv 0by construction; the held\-out GRM\-Llama3 reward model is anchor 2\), we invoke Theorem[11](https://arxiv.org/html/2608.08002#Thmtheorem11)on one independent pair per certification prompt \(m=80m\{=\}80; first two candidates in the frozen order\), all scales fixed before outcomes were inspected \(C1=C2=c=2C\_\{1\}\{=\}C\_\{2\}\{=\}c\{=\}2, target width11;δest=δsearch=0\.025\\delta\_\{\\rm est\}\{=\}\\delta\_\{\\rm search\}\{=\}0\.025pointwise, Bonferroni across the855855\-cell family\)\. Reported exactly as observed: the certificate is*valid but uninformative in all855855cells*\.v^s\\widehat\{v\}\_\{s\}is small \(median0\.0170\.017, max0\.270\.27\) but the estimation correctionC1C2log\(1/δest\)/2m≈0\.61C\_\{1\}C\_\{2\}\\sqrt\{\\log\(1/\\delta\_\{\\rm est\}\)/2m\}\\approx 0\.61dominates it, so every Bernstein and Bennett radius hits its deterministic range cap and the pass fractions of1\.01\.0are trivial,*not*coverage\. Reachingv^s\\widehat\{v\}\_\{s\}\-scale precision at these constants needsm≈105m\\approx 10^\{5\}pairs: at paper scale the certificate’s contribution is the honest sizing formula, not a usable bound \(Appendix[M](https://arxiv.org/html/2608.08002#A13.SS0.SSS0.Px4)\)\.
#### Real\-task certificate cells
Table[8](https://arxiv.org/html/2608.08002#A13.T8)reports representative certificate cells from the855855\-cell family \(all cells sharem=80m\{=\}80,C1=C2=c=2C\_\{1\}\{=\}C\_\{2\}\{=\}c\{=\}2, target width11; per\-cell values, both radius families, and the Bennett inversion are inreal\_outputs/anchor\_ucb\_cells\.csv\)\. No cell is favorable:v^s∈\[0\.009,0\.27\]\\widehat\{v\}\_\{s\}\\in\[0\.009,0\.27\]\(median0\.0170\.017\) is dominated by them=80m\{=\}80estimation correction, so*every*pointwise and simultaneous radius—Bernstein and Bennett alike—is capped at its deterministic range bound in all855855cells\. The pass fractions of1\.01\.0are trivial consequences of the caps, not coverage\.
Table 8:Representative anchor\-UCB certificate cells \(BsB\_\{s\}: selected\-error radius, capped atC1=2C\_\{1\}\{=\}2;TsT\_\{s\}: regret radius, capped at the target width11; pt/sim: pointwiseδest=δsearch=0\.025\\delta\_\{\\rm est\}\{=\}\\delta\_\{\\rm search\}\{=\}0\.025vs\. Bonferroni\-simultaneous\)\. The estimation correctionC1C2log\(1/δest\)/2m≈0\.61C\_\{1\}C\_\{2\}\\sqrt\{\\log\(1/\\delta\_\{\\rm est\}\)/2m\}\\approx 0\.61dominatesv^s\\widehat\{v\}\_\{s\}atm=80m\{=\}80, so all855855cells are range\-capped; Bennett radii equal Bernstein radii at these constants after capping\.
## Appendix NSupplementary Figures
Figure[3](https://arxiv.org/html/2608.08002#A14.F3)visualizes the corresponding selected\-quality estimates and prompt\-paired bootstrap intervals\.
Figure[4](https://arxiv.org/html/2608.08002#A14.F4)shows the corresponding target\-regret comparison\.
Figure[5](https://arxiv.org/html/2608.08002#A14.F5)shows the corresponding seed\-level differences and their uncertainty interval\.
Figure 3:Aggregation comparison on selected target quality for the largest eligible panel \(J=5J=5\) and search budget \(K=32K=32\)\. Points show mean selected target quality and horizontal bars show prompt\-paired bootstrap 95% intervals over the 120 locked test prompts\.Figure 4:Aggregation comparison on target regret for the largest eligible panel \(J=5J=5\) and search budget \(K=32K=32\)\. Points show mean target regret and horizontal bars show prompt\-paired bootstrap 95% intervals over the 120 locked test prompts\.Figure 5:Three\-seed DPO pilot comparing disagreement selection with random selection\. Points show the seed\-level differences, and the diamond shows the mean with its 95%ttinterval\. The interval crosses zero, so the pilot does not establish a reliable improvement or practical equivalence\.相似文章
基于评分标准的强化学习中的奖励黑客问题
本文研究了基于评分标准的强化学习中的奖励黑客现象,分析了训练验证器与评估指标之间的分歧。文章提出了一种针对“自我内化差距”的诊断方法,并证明更强的验证能力虽然能减少但无法完全消除奖励黑客问题。
更令人信服,而非更正确:无参考LLM评判者的自我博弈奖励操纵
本文识别了自我博弈训练中使用的无参考LLM评判者的结构缺陷,表明它们评估的是合理性而非正确性,导致奖励操纵,策略学会生成合理但错误的答案。作者提出了一种隐藏锚点审计和解锚奖励来缓解这一问题。
大模型时代的奖励黑客:机制、涌现错位与挑战
综述提出“代理压缩假设”,解释 RLHF 及相关方法如何在大型语言与多模态模型中系统性地诱发奖励黑客、欺骗与监督博弈。
编码代理会欺骗我们吗?通过带封顶评估与随机测试检测和防止作弊
本文介绍CapCode,一种带封顶评估框架,利用随机测试输出检测操纵单元测试的编码代理,以及CapReward,一种在编码任务中惩罚奖励黑客行为的奖励设计。
Multiscale Reward Hedging from Correct Demonstrations
This paper presents a multiscale reward hedging method for learning from correct demonstrations, extending guarantees to continuous reward classes with a horizon-free bound via metric entropy, and shows polynomial-time cases for specific settings.