用于鲁棒成对LLM评判的骨干自适应证据路由

arXiv cs.AI 论文

摘要

BAER引入了一种骨干自适应证据路由方法,用于鲁棒的成对LLM评判,通过根据条件动态选择证据机制,在多个基准测试中实现了更高的准确性。

arXiv:2609.30751v1 Announce Type: new Abstract: Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength. BAER separates each expert's signed preference from candidate-invariant reliability and builds three symmetric heads: evidence stacking, reliability-based expert routing, and candidate-blind reference verification. Development data select one head for each benchmark--backbone condition, and that choice is frozen before testing. Across four benchmarks and two 8B judge backbones, BAER achieves the highest test accuracy among the compared methods in all eight conditions, with full prediction coverage and gains of 0.87--7.32 points over the strongest external baseline. The results show that adapting how evidence is gathered is more reliable than fixing one judging protocol everywhere.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:45

# Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging
Source: [https://arxiv.org/html/2609.30751](https://arxiv.org/html/2609.30751)
###### Abstract

Pairwise language\-model judges can gather evidence through direct comparison, reasoning, or reference\-based verification, but no single protocol is best across benchmarks and judge backbones\. We introduce Backbone\-Adaptive Evidence Routing \(BAER\), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength\.BAERseparates each expert’s signed preference from candidate\-invariant reliability and builds three symmetric heads: evidence stacking, reliability\-based expert routing, and candidate\-blind reference verification\. Development data select one head for each benchmark–backbone condition, and that choice is frozen before testing\. Across four benchmarks and two 8B judge backbones,BAERachieves the highest test accuracy among the compared methods in all eight conditions, with full prediction coverage and gains of 0\.87–7\.32 points over the strongest external baseline\. The results show that adapting how evidence is gathered is more reliable than fixing one judging protocol everywhere\.

###### Index Terms:

LLM\-as\-a\-judge, pairwise evaluation, evidence routing, preference evaluation

††address:1Shanghai Jiao Tong University## 1Introduction

Pairwise language\-model judges make it practical to compare model outputs at a scale that would otherwise require human raters, and they are now used for model evaluation, translation, and open\-ended chat\[[5](https://arxiv.org/html/2609.30751#bib.bib20),[24](https://arxiv.org/html/2609.30751#bib.bib1)\]\. Given an instruction and two candidate responses, the judge states which response is better\. Platforms and meta\-benchmarks increasingly rely on this two\-response interface\[[6](https://arxiv.org/html/2609.30751#bib.bib21),[11](https://arxiv.org/html/2609.30751#bib.bib6),[16](https://arxiv.org/html/2609.30751#bib.bib4)\]\. Choosing how to use the interface is more involved, because a protocol that works well for one task or judge backbone, that is, the language model that performs the judging, may not be the strongest choice for another\. We call the combination of a benchmark and a judge backbone a*condition*\.

Studies that evaluate the judges themselves have found recurring weaknesses\. Judges can be fooled by adversarially constructed responses\[[22](https://arxiv.org/html/2609.30751#bib.bib2)\], often prefer whichever response appears in a favored position\[[18](https://arxiv.org/html/2609.30751#bib.bib19),[15](https://arxiv.org/html/2609.30751#bib.bib3)\], and are swayed by surface features such as response length or the tokens used to read out the decision\[[8](https://arxiv.org/html/2609.30751#bib.bib18),[23](https://arxiv.org/html/2609.30751#bib.bib12),[12](https://arxiv.org/html/2609.30751#bib.bib13)\]\. Pairwise judgments can also disagree with an independently constructed basis for evaluation\[[10](https://arxiv.org/html/2609.30751#bib.bib14)\]\. A correction designed for one condition may therefore discard useful evidence in another\. The problem concerns both how to judge and which evidence to trust in each condition\.

Figure[1](https://arxiv.org/html/2609.30751#S1.F1)illustrates two reasons why the useful evidence mechanism can change\. Experts may disagree, or they may agree on an incorrect answer\. In the first case, selecting a reliable expert can resolve the disagreement\. In the second, combining the same judgments may be insufficient, and a reference constructed without seeing either candidate provides another basis for verification\.

Figure 1:Motivation for condition\-adaptive evidence\. Conflicting experts and shared errors call for different mechanisms, illustrated by BAER’s invariant expert routing and candidate\-blind reference verification\. Head assignments match the deployed JudgeBench conditions; examples are schematic\.Adapting the evidence must not make its selection depend on how the candidates are ordered\. A judge whose answer changes when the two responses are swapped is measuring display order rather than quality, so any adaptive mechanism must remain blind to that order\. We call each raw judging procedure, such as direct comparison or chain\-of\-thought, a*protocol*, and its output an*expert signal*\. We separate that signal into a signed preference, which records which response the expert favors and how strongly, and a candidate\-invariant reliability, which records how trustworthy the expert is on the current pair regardless of order\. Swapping the responses reverses the first and leaves the second unchanged\. This separation underliesBAERand its three alternative evidence paths, which we call*heads*\. A stack combines the expert signals, a router selects one expert for each pair, and reference verification checks the candidates against a separately constructed solution\. Development data determine which head to deploy for each condition\. The deployed head remains fixed at test time, while expert selection within the routing head varies from pair to pair\.

Existing methods mainly improve a fixed judging protocol\. One line changes how a comparison is elicited through explicit reasoning, rubric\-style evaluation, or repeated sampling\[[20](https://arxiv.org/html/2609.30751#bib.bib17),[13](https://arxiv.org/html/2609.30751#bib.bib7),[19](https://arxiv.org/html/2609.30751#bib.bib8)\]\. A second keeps the protocol fixed and calibrates a predetermined decision rule against known biases\[[23](https://arxiv.org/html/2609.30751#bib.bib12),[12](https://arxiv.org/html/2609.30751#bib.bib13)\]\. A third aggregates a fixed panel of judges\[[17](https://arxiv.org/html/2609.30751#bib.bib22)\]\. Selective methods decide when to trust a judgment and when to abstain\[[2](https://arxiv.org/html/2609.30751#bib.bib5)\], with related work providing statistical tools for risk\-controlled selection\[[1](https://arxiv.org/html/2609.30751#bib.bib11)\]\. None of these addresses the case in which the useful evidence mechanism itself changes across conditions\. Aggregation cannot repair an error shared by all pairwise experts, while abstention gives up coverage exactly on uncertain pairs\.BAERinstead asks which symmetric evidence head should be deployed in each condition while still producing a prediction for every pair\.

Two adaptation scales are deliberately separated\. Head selection is*condition\-level*: development data choose stacking, routing, or reference verification for a benchmark–backbone pair, and the choice is frozen before test labels are observed\. Only the routing head adapts*within*a condition, selecting an expert for each pair from candidate\-invariant reliability features\. This distinction prevents test\-time head shopping and prevents the router from exploiting display order\. It also makes failures interpretable: a weak condition can require a different evidence source even when individual experts remain useful on particular pairs\.

We make three contributions\. First, we formulate evidence adaptation in a way that preserves candidate symmetry\. Second, we implement three evidence heads that select heads at the condition level and experts at the sample level\. Third, we empirically study when each head helps\. Across eight benchmark–backbone conditions,BAERachieves the highest test accuracy among the compared methods, with margins over the strongest external baseline ranging from \+0\.87 to \+7\.32 points\. The development experiments further show when stacking, routing, and reference verification provide useful evidence\.

## 2BAER

BAERadapts how preference evidence is gathered while treating the two candidates symmetrically\. As shown in Figure[2](https://arxiv.org/html/2609.30751#S2.F2), all heads share one output format, a signed score whose sign picks the winner and whose magnitude measures confidence, and swapping the candidates flips that sign\. We first define the score interface and evidence representation, then describe the three heads and the deployment rule\.

Figure 2:Overview of BAER\. Development data freeze one evidence head per benchmark–backbone condition\. Stacking combines signed protocol features, routing selects one expert per pair using candidate\-invariant reliability, and reference verification solves the instruction without seeing either candidate, then checks each independently\. All heads return a candidate\-symmetric signed score\.For instructionqqand candidatesa,ba,b, letx=\(q,a,b\)x=\(q,a,b\)and encode the benchmark preference asy∈\{−1,\+1\}y\\in\\\{\-1,\+1\\\}, where\+1\+1means thataais preferred\. AllBAERheads return a signed score through the same interface,

ph\(A∣x\)=σ\(sh\(x\)\),y^h=2𝟙\[sh\(x\)≥0\]−1,p\_\{h\}\(A\\mid x\)=\\sigma\(s\_\{h\}\(x\)\),\\qquad\\widehat\{y\}\_\{h\}=2\\mathbb\{1\}\[s\_\{h\}\(x\)\\geq 0\]\-1,\(1\)where𝟙​\[⋅\]\\mathbb\{1\}\[\\cdot\]is the indicator function\. Candidate exchange isT​x=\(q,b,a\)Tx=\(q,b,a\)\. We require

sh​\(T​x\)=−sh​\(x\),ph​\(A∣T​x\)=1−ph​\(A∣x\)\.s\_\{h\}\(Tx\)=\-s\_\{h\}\(x\),\\qquad p\_\{h\}\(A\\mid Tx\)=1\-p\_\{h\}\(A\\mid x\)\.\(2\)Swapping the responses may flip the judgment, but it does not change its strength\.

### 2\.1Symmetric evidence representation

The expert bank contains nine judging protocols\. These are direct bidirectional judging, self\-consistency\[[19](https://arxiv.org/html/2609.30751#bib.bib8)\], chain\-of\-thought\[[20](https://arxiv.org/html/2609.30751#bib.bib17)\], rubric\-based judging, score\-then\-choose judging, PRePair\[[10](https://arxiv.org/html/2609.30751#bib.bib14)\], PriDe\[[23](https://arxiv.org/html/2609.30751#bib.bib12)\], CalibraEval\[[12](https://arxiv.org/html/2609.30751#bib.bib13)\], and an internal deliberate A versus B expert with constrained extraction\. Each expert may be run under both display orders, with every probability mapped back to the canonical candidateaa, so a probability always refers to the same response regardless of display position\. Letuk​m​\(x\)u\_\{km\}\(x\)be themm\-th such probability for expertkk, and letMkM\_\{k\}be the number of runs\. Its directional confidence and order instability are

pk=1Mk​∑m=1Mkuk​m,δk=maxm⁡uk​m−minm⁡uk​m\.p\_\{k\}=\\frac\{1\}\{M\_\{k\}\}\\sum\_\{m=1\}^\{M\_\{k\}\}u\_\{km\},\\qquad\\delta\_\{k\}=\\max\_\{m\}u\_\{km\}\-\\min\_\{m\}u\_\{km\}\.\(3\)Herepkp\_\{k\}is the expert’s average preference foraa, andδk\\delta\_\{k\}is large for an expert whose answer moves across repeated or swapped runs\. We form raw and reliability\-shrunk log odds,

zk=logit⁡\(pk\),z~k=zk1\+4​δk\.z\_\{k\}=\\operatorname\{logit\}\(p\_\{k\}\),\\qquad\\widetilde\{z\}\_\{k\}=\\frac\{z\_\{k\}\}\{1\+4\\delta\_\{k\}\}\.\(4\)The shrinkage downweights experts that order instability identifies as unreliable\. Missing evidence is assignedpk=\.5p\_\{k\}=\.5, hence zero signed evidence\. Under exchange,pk↦1−pkp\_\{k\}\\mapsto 1\-p\_\{k\}andδk\\delta\_\{k\}is unchanged, so bothzkz\_\{k\}andz~k\\widetilde\{z\}\_\{k\}negate\. Withz¯=K−1​∑kzk\\bar\{z\}=K^\{\-1\}\\sum\_\{k\}z\_\{k\}, the stack feature map is

ϕ⁡\(x\)\\displaystyle\\phi\(x\)=\[\{zk,z~k\}k=1K,\{𝟙\[g\(x\)=c\]z¯\}c∈𝒢\],\\displaystyle=\\big\[\\\{z\_\{k\},\\widetilde\{z\}\_\{k\}\\\}\_\{k=1\}^\{K\},\\\{\\mathbb\{1\}\[g\(x\)=c\]\\bar\{z\}\\\}\_\{c\\in\\mathcal\{G\}\}\\big\],\(5\)ϕ⁡\(T​x\)\\displaystyle\\phi\(Tx\)=−ϕ⁡\(x\)\.\\displaystyle=\-\\phi\(x\)\.Here𝒢\\mathcal\{G\}is the set of subset and task\-type categories, such as chat, safety, reasoning, and math\. The second block interacts the mean evidence with indicators for these categories, letting the stack weight experts differently across task types\. These categories are candidate\-invariant and are fitted without test labels\.

### 2\.2Backbone\-adaptive evidence heads

Evidence stacking\.The stack is a logistic model over these features\. Each dimension is divided by its training RMSrℓr\_\{\\ell\}, computed over the training split, so that experts with naturally larger scores do not dominate\. With the sign\-augmented set𝒟±\\mathcal\{D\}^\{\\pm\}, obtained by adding\(−ϕi,−yi\)\(\-\\phi\_\{i\},\-y\_\{i\}\)for every\(ϕi,yi\)\(\\phi\_\{i\},y\_\{i\}\),BAERfits

wλ\\displaystyle w\_\{\\lambda\}=arg⁡minw​1\|𝒟±\|​∑\(v,t\)∈𝒟±ℓ⁡\(t​w⊤​\(v/r\)\)\+λ2​∥w∥22,\\displaystyle=\\arg\\min\_\{w\}\\frac\{1\}\{\|\\mathcal\{D\}^\{\\pm\}\|\}\\sum\_\{\(v,t\)\\in\\mathcal\{D\}^\{\\pm\}\}\\ell\\\!\\left\(tw^\{\\top\}\(v/r\)\\right\)\+\\frac\{\\lambda\}\{2\}\\lVert w\\rVert\_\{2\}^\{2\},\(6\)sstk​\(x\)\\displaystyle s\_\{\\rm stk\}\(x\)=wλ⊤\(ϕ\(x\)/r\),ℓ\(u\)=log\(1\+e−u\)\.\\displaystyle=w\_\{\\lambda\}^\{\\top\}\(\\phi\(x\)/r\),\\qquad\\ell\(u\)=\\log\(1\+e^\{\-u\}\)\.This augmentation forces the learned model to be odd inϕ\\phi, and the stack has no intercept, so antisymmetry holds by construction\. We selectλ∈\{10,1,0\.1,0\.01,0\.001\}\\lambda\\in\\\{10,1,0\.1,0\.01,0\.001\\\}on calibration data and refit on all non\-test rows\. The stack feature dimension is 44 for RewardBench, 36 for JudgeBench, 20 for HH\-RLHF, and 28 for UltraFeedback, with selectedλ\\lambdavalues0\.001/0\.010\.001/0\.01,10/0\.0110/0\.01,0\.01/0\.010\.01/0\.01, and0\.001/0\.010\.001/0\.01for Qwen and Llama respectively\.

Candidate\-invariant expert routing\.Averaging can erase a strong specialist with votes from experts that know nothing about the pair\. Routing replaces the average with a learned selector\. For pairiiand expertkk, lety^i​k\\widehat\{y\}\_\{ik\}be the binary prediction of expertkk, and letti​k=𝟙\[y^i​k=yi\]t\_\{ik\}=\\mathbb\{1\}\[\\widehat\{y\}\_\{ik\}=y\_\{i\}\]state whether that expert is correct\. An auxiliary modelRθR\_\{\\theta\}estimates each expert’s reliability on the current pair:

Rθ​\(ρi​k\)\\displaystyle R\_\{\\theta\}\(\\rho\_\{ik\}\)≈P⁡\(ti​k=1∣ρi​k\),\\displaystyle\\approx P\(t\_\{ik\}=1\\mid\\rho\_\{ik\}\),\(7\)k∗​\(x\)\\displaystyle k^\{\*\}\(x\)=arg⁡maxk​Rθ​\(ρk​\(x\)\),\\displaystyle=\\arg\\max\_\{k\}R\_\{\\theta\}\(\\rho\_\{k\}\(x\)\),srte​\(x\)\\displaystyle s\_\{\\rm rte\}\(x\)=zk∗​\(x\)​\(x\)\.\\displaystyle=z\_\{k^\{\*\}\(x\)\}\(x\)\.The vectorρi​k\\rho\_\{ik\}contains expert identity, absolute confidence, cross\-expert agreement,δk\\delta\_\{k\}, confidence\-distribution statistics, symmetric response\-length and structure features, and subset identity\. All entries are candidate\-invariant, soρk​\(T​x\)=ρk​\(x\)\\rho\_\{k\}\(Tx\)=\\rho\_\{k\}\(x\)\. Exchange preserves the selected expertk∗k^\{\*\}and negates only the selected scorezk∗z\_\{k^\{\*\}\}\. Logistic regression and tree ensembles are compared by deterministic five\-fold cross\-validation; both deployed routers select a depth\-5 random forest\[[4](https://arxiv.org/html/2609.30751#bib.bib23)\]with minimum leaf size 8 and refit it on all non\-test rows\. JudgeBench/Llama also uses an isolated pointwise expert with score

spnt=logit⁡P⁡\(Y∣q,a\)−logit⁡P⁡\(Y∣q,b\),s\_\{\\rm pnt\}=\\operatorname\{logit\}P\(Y\\mid q,a\)\-\\operatorname\{logit\}P\(Y\\mid q,b\),\(8\)where each probability is normalized over constrained one\-tokenY/NY/Noutputs\. Neither verifier sees the other candidate, and this pointwise expert is available only as an additional candidate for the routing head\.

Candidate\-blind reference verification\.When all experts make the same error, no combination or selection of their judgments can recover\. Reference verification adds evidence of a different kind by comparing each candidate with an independently constructed solution\. JudgeBench/Qwen 3\-8B uses three calls\. The judge first sees onlyqqand produces a reference solutionrr, then evaluates each candidate independently under the same instruction and reference, with one\-token Y/N constrained decoding\. Letπl​\(t\)=P⁡\(l∣q,r,t\)\\pi\_\{l\}\(t\)=P\(l\\mid q,r,t\)forl∈\{Y,N\}l\\in\\\{Y,N\\\}\. Then

v⁡\(t\)\\displaystyle v\(t\)=πY​\(t\)πY​\(t\)\+πN​\(t\),\\displaystyle=\\frac\{\\pi\_\{Y\}\(t\)\}\{\\pi\_\{Y\}\(t\)\+\\pi\_\{N\}\(t\)\},\(9\)sref​\(x\)\\displaystyle s\_\{\\rm ref\}\(x\)=logit⁡v⁡\(a\)−logit⁡v⁡\(b\)\.\\displaystyle=\\operatorname\{logit\}v\(a\)\-\\operatorname\{logit\}v\(b\)\.The solver never sees either candidate, and each verifier sees exactly one, so swapping candidates negates the score\. The prompt, 4096\-token solution budget, constrained decoding, and aggregation are frozen before test\.

### 2\.3Condition\-level selection and deployment

Recall that a conditionc=\(d,j\)c=\(d,j\)combines a benchmarkddand a judge backbonejj\. All pairs in a condition use the same head, so deployment never uses test information per example\. Stacking is the initial choice, and development evaluations assess alternatives where it performs poorly\. Cross\-validation selects the expert router, while a head that relies on reference verification must pass its selection and calibration margins\. The resulting condition sets𝒞stk\\mathcal\{C\}\_\{\\rm stk\},𝒞rte\\mathcal\{C\}\_\{\\rm rte\}, and𝒞ref\\mathcal\{C\}\_\{\\rm ref\}are disjoint and fixed for final scoring:

sBAER​\(x,c\)=\{sstk​\(x\),c∈𝒞stk,srte​\(x\),c∈𝒞rte,sref​\(x\),c∈𝒞ref\.s\_\{\\textsc\{BAER\}\}\(x;c\)=\\begin\{cases\}s\_\{\\rm stk\}\(x\),&c\\in\\mathcal\{C\}\_\{\\rm stk\},\\\\ s\_\{\\rm rte\}\(x\),&c\\in\\mathcal\{C\}\_\{\\rm rte\},\\\\ s\_\{\\rm ref\}\(x\),&c\\in\\mathcal\{C\}\_\{\\rm ref\}\.\\end\{cases\}\(10\)Because the condition does not depend on candidate order,c⁡\(T​x\)=c⁡\(x\)c\(Tx\)=c\(x\)\. Since every branch is odd,sBAER​\(T​x,c\)=−sBAER​\(x,c\)s\_\{\\textsc\{BAER\}\}\(Tx;c\)=\-s\_\{\\textsc\{BAER\}\}\(x;c\)\. Within a routing condition,k∗​\(x\)k^\{\*\}\(x\)in Eq\.[7](https://arxiv.org/html/2609.30751#S2.E7)can vary across pairs, but candidate exchange preserves both the head and the expert\.

## 3Experiments

### 3\.1Setup

We use four preference benchmarks, RewardBench filtered\[[11](https://arxiv.org/html/2609.30751#bib.bib6)\]\(chat, safety, reasoning\), JudgeBench\[[16](https://arxiv.org/html/2609.30751#bib.bib4)\]\(knowledge, reasoning, math, coding\), HH\-RLHF\[[3](https://arxiv.org/html/2609.30751#bib.bib15)\]\(helpfulness and harmlessness\), and UltraFeedback\[[7](https://arxiv.org/html/2609.30751#bib.bib16)\]\(highest\- versus lowest\-rated non\-tied completions\)\. Their selection/calibration/test sizes are 588/1,234/1,163, 136/238/246, 1,717/3,335/3,500, and 12,598/25,406/25,599, respectively, with original labels\. The judge backbones are Qwen 3\-8B\[[21](https://arxiv.org/html/2609.30751#bib.bib9)\]and Llama 3\.1\-8B Instruct\[[9](https://arxiv.org/html/2609.30751#bib.bib10)\], run locally with greedy decoding except for five\-trial self\-consistency\. Every pair is judged in both response orders and mapped back to the canonical frame, so accuracy reflects content rather than display position\. We compare the eight external protocols of Section[2](https://arxiv.org/html/2609.30751#S2)on matched test IDs under their native inference procedures\. This measures attainable accuracy without controlling for inference cost, and BAER’s expert portfolio uses more calls than a single\-pass judge\. The metric is full\-partition pairwise accuracy,

Acc=1N∑i=1N𝟙\[y^i=yi\]\.\\operatorname\{Acc\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{1\}\[\\widehat\{y\}\_\{i\}=y\_\{i\}\]\.\(11\)Missing or unparsed outputs count as errors\. We report prediction coverage, defined as the fraction of test pairs for which a method produces a prediction, and use the exact two\-sided McNemar test\[[14](https://arxiv.org/html/2609.30751#bib.bib24)\]against each column’s strongest external baseline\. Development used only development data, first for stacking, then for routing, and finally for reference verification when JudgeBench/Qwen remained weak\. Each new head was tuned on the selection and calibration partitions and run once on test inputs with labels withheld\.

### 3\.2Overall effectiveness

Table 1:Full\-test pairwise accuracy \(%\) on the benchmark–backbone columns\. Bold marks the best result, underline marks the strongest external baseline, and†marks a significant exact two\-sided McNemar comparison against the strongest baseline in that column \(p<\.05p<\.05\)\.BAERachieves the highest accuracy in every Table[1](https://arxiv.org/html/2609.30751#S3.T1)column, with margins from \+0\.87 to \+7\.32 points \(mean \+4\.90\)\. The strongest external method changes across columns\. CalibraEval leads six columns, while chain\-of\-thought and PriDe lead the two JudgeBench columns\. This is direct evidence that no fixed protocol dominates, which is the variationBAERexploits through condition\-specific heads\. Six of the eight gains are significant at the 5% level under an exact two\-sided McNemar test\[[14](https://arxiv.org/html/2609.30751#bib.bib24)\]\. The two non\-significant comparisons are the two JudgeBench columns\.BAERalso produces a prediction for every test pair\. In contrast, chain\-of\-thought has coverage between 74\.8% and 100% across columns\.

### 3\.3Contributions of the evidence heads

Stacking provides the broad foundation\. Routing and reference verification improve the three conditions where it is weaker\. Column counts track the progression\. The constrained A/B expert alone leads 1/8 of the columns, the stack\-based system 5/8, the routing\-augmented version 7/8, and fullBAER8/8\. Selecting the routing head raises JudgeBench/Llama from 47\.6% to 57\.7% and UltraFeedback/Qwen from 97\.0% to 98\.5%\. Replacing the previous router with reference verification raises JudgeBench/Qwen from 60\.2% to 70\.7%\.

The transfer behavior explains this division of roles\. Six stacks remain within 1\.5 points from calibration to test, while the two JudgeBench stacks drop 9\.5 and 9\.1 points\. The harder JudgeBench pairs therefore need evidence of a different kind\. Routing is itself condition\-sensitive\. The former JudgeBench/Qwen router gained 4\.8 out\-of\-fold points but reached only 60\.2% on test, and simple confidence switching and subset lookup also fail\. An any\-expert oracle, which uses the correct label to select the best expert per pair and so measures the available headroom, reaches 95\.4%/96\.2%\. The correct judgments are present inside the expert bank\. Identifying them from reliability features, however, remains difficult\.

Figure 3:JudgeBench semantic\-head development\. Accuracy gaps are relative to the strongest external baseline on matched IDs; zero denotes parity\. \(a\) Selection results for the best variant per family \(n=128n=128; other heads usen=136n=136and are named in \(b\)\)\. CF = counterfactual reconciliation, Orbit = complete orbit, Anchor = blind anchor, Audit = independent pointwise audit\. \(b\) Selection and calibration \(n=238n=238\) for the three advancing heads\. Only candidate\-blind reference verification retains a positive gap; the Llama reference head was not evaluated\.
### 3\.4From selection to calibration

A development head advances only if its gains survive two gates, a small selection split and a larger calibration split\. The JudgeBench development experiments in Figure[3](https://arxiv.org/html/2609.30751#S3.F3)show why both are needed\. The 128\-example screening runs do not pass the selection gate\. Claim\-graph verification improves by \+2\.2 points for Llama on selection, then reverses to \-14\.3 on calibration\. Independent pointwise likelihood improves by \+5\.9 points on selection, then reverses to \-5\.0 on calibration\. Candidate\-blind reference verification is the Qwen head that clears both gates, with gains of \+3\.7/\+6\.3 points\. Advancement is decided by accuracy margin, not by significance\. The corresponding McNemarpp\-values at these partition sizes are \.551/\.096\. The deployed solution budget remains 4096 tokens; 82/116 selection/calibration solutions reach this limit\. Two further 128\-example screening studies test broader replacements\. CalibraEval’s full fourth orbit, a variant that expands the calibration space, wins only for Qwen on JudgeBench \(\+1\.6\), losing 0\.8–38\.3 points elsewhere\.

## 4Conclusion

We presentBAER, a judging framework that adapts the evidence mechanism to each benchmark and judge backbone while preserving candidate symmetry\. It separates each expert’s signed preference from candidate\-invariant reliability and builds three heads on that separation: stacking, routing, and candidate\-blind reference verification\. Across four benchmarks and two 8B judge backbones,BAERachieves the highest accuracy in all conditions, with margins of \+0\.87 to \+7\.32 points over the strongest external baseline and a prediction for every test pair\. The experiments show why adaptation matters: stacks transfer well on most conditions, routing helps when specialist judgments can be identified from invariant reliability, and reference verification supplies new evidence when the pairwise expert bank shares an error\. The gains come with additional inference cost, and condition\-level deployment assumes that the benchmark family and judge backbone are known in advance\. Future work should extend head selection to unseen conditions and reduce cost through distillation, expert pruning, cached references, or conditional early exits\.

## References

- \[1\]A\. N\. Angelopoulos, S\. Bates, E\. J\. Candès, M\. I\. Jordan, and L\. Lei\(2025\)Learn then test: calibrating predictive algorithms to achieve risk control\.The Annals of Applied Statistics19\(2\),pp\. 1641–1662\.External Links:[Document](https://dx.doi.org/10.1214/24-AOAS1998),[Link](https://doi.org/10.1214/24-AOAS1998)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p5.1)\.
- \[2\]S\. Badshah, A\. Emami, and H\. Sajjad\(2026\)SCOPE: selective conformal optimized pairwise LLM judging\.arXiv preprint arXiv:2602\.13110\.External Links:[Link](https://arxiv.org/abs/2602.13110)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p5.1)\.
- \[3\]Y\. Bai, A\. Jones, K\. Ndousse,et al\.\(2022\)Training a helpful and harmless assistant with reinforcement learning from human feedback\.arXiv preprint arXiv:2204\.05862\.External Links:[Link](https://arxiv.org/abs/2204.05862)Cited by:[§3\.1](https://arxiv.org/html/2609.30751#S3.SS1.p1.1)\.
- \[4\]L\. Breiman\(2001\)Random forests\.Machine Learning45\(1\),pp\. 5–32\.External Links:[Document](https://dx.doi.org/10.1023/A%3A1010933404324),[Link](https://doi.org/10.1023/A:1010933404324)Cited by:[§2\.2](https://arxiv.org/html/2609.30751#S2.SS2.p2.3)\.
- \[5\]C\. Chiang and H\. Lee\(2023\)Can large language models be an alternative to human evaluations?\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15607–15631\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.870),[Link](https://aclanthology.org/2023.acl-long.870/)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p1.1)\.
- \[6\]W\. Chianget al\.\(2024\)Chatbot arena: an open platform for evaluating LLMs by human preference\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 8359–8388\.External Links:[Link](https://proceedings.mlr.press/v235/chiang24b.html)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p1.1)\.
- \[7\]G\. Cuiet al\.\(2024\)UltraFeedback: boosting language models with scaled AI feedback\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 9722–9744\.External Links:[Link](https://proceedings.mlr.press/v235/cui24f.html)Cited by:[§3\.1](https://arxiv.org/html/2609.30751#S3.SS1.p1.1)\.
- \[8\]Y\. Dubois, P\. Liang, and T\. B\. Hashimoto\(2024\)Length\-controlled AlpacaEval: a simple way to debias automatic evaluators\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=CybBmzWBX0)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p2.1)\.
- \[9\]A\. Grattafiori, A\. Dubey, A\. Jauhri,et al\.\(2024\)The Llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.External Links:[Link](https://arxiv.org/abs/2407.21783)Cited by:[§3\.1](https://arxiv.org/html/2609.30751#S3.SS1.p1.1)\.
- \[10\]H\. Jeong, C\. Park, J\. Hong, H\. Lee, and J\. Choo\(2025\)The comparative trap: pairwise comparisons amplifies biased preferences of LLM evaluators\.InProceedings of the 8th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 79–108\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.blackboxnlp-1.5),[Link](https://aclanthology.org/2025.blackboxnlp-1.5/)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.30751#S2.SS1.p1.2)\.
- \[11\]N\. Lambertet al\.\(2025\)RewardBench: evaluating reward models for language modeling\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 1755–1797\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.96),[Link](https://aclanthology.org/2025.findings-naacl.96/)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.30751#S3.SS1.p1.1)\.
- \[12\]H\. Liet al\.\(2025\)CalibraEval: calibrating prediction distribution to mitigate selection bias in LLMs\-as\-judges\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 16537–16552\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.808),[Link](https://aclanthology.org/2025.acl-long.808/)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p2.1),[§1](https://arxiv.org/html/2609.30751#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.30751#S2.SS1.p1.2)\.
- \[13\]Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. Zhu\(2023\)G\-Eval: NLG evaluation using GPT\-4 with better human alignment\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 2511–2522\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.153),[Link](https://aclanthology.org/2023.emnlp-main.153/)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p5.1)\.
- \[14\]Q\. McNemar\(1947\)Note on the sampling error of the difference between correlated proportions or percentages\.Psychometrika12\(2\),pp\. 153–157\.External Links:[Document](https://dx.doi.org/10.1007/BF02295996),[Link](https://doi.org/10.1007/BF02295996)Cited by:[§3\.1](https://arxiv.org/html/2609.30751#S3.SS1.p1.2),[§3\.2](https://arxiv.org/html/2609.30751#S3.SS2.p1.1)\.
- \[15\]L\. Shi, C\. Ma, W\. Liang, X\. Diao, W\. Ma, and S\. Vosoughi\(2025\)Judging the judges: a systematic study of position bias in LLM\-as\-a\-judge\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,pp\. 292–314\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.ijcnlp-long.18),[Link](https://aclanthology.org/2025.ijcnlp-long.18/)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p2.1)\.
- \[16\]S\. Tanet al\.\(2025\)JudgeBench: a benchmark for evaluating LLM\-based judges\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=G0dksFayVq)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.30751#S3.SS1.p1.1)\.
- \[17\]P\. Vergaet al\.\(2024\)Replacing judges with juries: evaluating LLM generations with a panel of diverse models\.arXiv preprint arXiv:2404\.18796\.External Links:[Link](https://arxiv.org/abs/2404.18796)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p5.1)\.
- \[18\]P\. Wanget al\.\(2024\)Large language models are not fair evaluators\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 9440–9450\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.511),[Link](https://aclanthology.org/2024.acl-long.511/)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p2.1)\.
- \[19\]X\. Wanget al\.\(2023\)Self\-consistency improves chain of thought reasoning in language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=1PL1NIMMrw)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.30751#S2.SS1.p1.2)\.
- \[20\]J\. Weiet al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 24824–24837\.External Links:[Document](https://dx.doi.org/10.52202/068431-1800),[Link](https://proceedings.neurips.cc/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.30751#S2.SS1.p1.2)\.
- \[21\]A\. Yang, A\. Li, B\. Yang,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.External Links:[Link](https://arxiv.org/abs/2505.09388)Cited by:[§3\.1](https://arxiv.org/html/2609.30751#S3.SS1.p1.1)\.
- \[22\]Z\. Zeng, J\. Yu, T\. Gao, Y\. Meng, T\. Goyal, and D\. Chen\(2024\)Evaluating large language models at evaluating instruction following\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=tr0KidwPLc)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p2.1)\.
- \[23\]C\. Zheng, H\. Zhou, F\. Meng, J\. Zhou, and M\. Huang\(2024\)Large language models are not robust multiple choice selectors\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=shr9PXz7T0)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p2.1),[§1](https://arxiv.org/html/2609.30751#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.30751#S2.SS1.p1.2)\.
- \[24\]L\. Zhenget al\.\(2023\)Judging LLM\-as\-a\-judge with MT\-Bench and chatbot arena\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 46595–46623\.External Links:[Document](https://dx.doi.org/10.52202/075280-2020),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html)Cited by:[§1](https://arxiv.org/html/2609.30751#S1.p1.1)\.

相似文章

AdversaBench: 自动化LLM红队测试的多裁判确认与跨模型迁移性

arXiv cs.AI

AdversaBench介绍了一个自动化LLM红队测试流程,该流程使用五个变异算子和一个由三位裁判及元裁判(用于决断平局)组成的评审团来确认失败,揭示了攻击难度因类别而异,并且对抗性提示可以从较小模型迁移到较大模型。

面向可靠LLM判断的边际自适应置信度排序

arXiv cs.LG

本文提出了一种针对LLM作为评判系统的基于边际的置信度排序方法,通过学习专用估计器来确保置信度与人类分歧风险之间的单调性,具有泛化保证,并在多个数据集上提高了排序准确性。

SeLMRoute:基于概率化语义证据的大语言模型路由

Hugging Face Daily Papers

SeLMRoute 提出了一种大语言模型路由框架,将与候选模型无关的语义证据抽取与性能学习解耦。在 LLMRouterBench 上,跨越 15 个数据集和 20 个候选模型,该框架取得了 72.08% 的平均准确率,超越了最强的固定候选模型(69.23%),并同时支持面向性能和成本感知的路由决策。