MAWILE:多轴工作台用于检测LLM评估器
摘要
MAWILE是一个开发者工作台,用于审计LLM评判者对提示词、评分标准、输入和输出扰动的敏感性,以确保评估的鲁棒性和有意义性。
arXiv:2609.22599v1 Announce Type: new
Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires testing both the judge instrument and the items it evaluates. We introduce MAWILE, a developer-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target-system input, and target-system output. Given a user-supplied judge and representative evaluation items, MAWILE constructs and validates controlled perturbations, re-executes the judge, and localizes the resulting sensitivity. Each perturbation declares whether the verdict should remain invariant or change in a specified direction, allowing the same system to measure both robustness to irrelevant variations and sensitivity to meaningful changes. MAWILE audits binary, ordinal, and pairwise judges without requiring gold labels. The code for this tool is available at: github.com/megagonlabs/mawile-judge.
查看缓存全文
缓存时间: 2026/09/23 09:09
# MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
Source: [https://arxiv.org/html/2609.22599](https://arxiv.org/html/2609.22599)
Farima Fatahi BayatPouya PezeshkpourEstevam HruschkaAffiliation:Megagon LabsAffiliation:\{jackson, farima, pouya, estevam\}@megagon\.ai
###### Abstract
Large language model \(LLM\) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric\. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires testing both the judge instrument and the items it evaluates\. We introduceMawile, a developer\-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target\-system input, and target\-system output\. Given a user\-supplied judge and representative evaluation items,Mawileconstructs and validates controlled perturbations, re\-executes the judge, and localizes the resulting sensitivity\. Each perturbation declares whether the verdict should remain invariant or change in a specified direction, allowing the same system to measure both robustness to irrelevant variations and sensitivity to meaningful changes\.Mawileaudits binary, ordinal, and pairwise judges without requiring gold labels\. The code for this tool is available at:[github\.com/megagonlabs/mawile\-judge](https://github.com/megagonlabs/mawile-judge)\.
## 1Introduction
Figure 1:TheMawileend\-to\-end judge\-auditing workflow\. Given a configured judge and representative audit items, the planner agent produces an editable suite of invariant and directional probes\.Mawilethen generates and validates these variants, re\-executes the judge alongside repeated unmodified baselines, and reports robustness, degradation sensitivity, and—when gold labels are available—correctness\.Large language models \(LLMs\) are increasingly used as flexible, scalable evaluators of model and agent outputs\([Zheng et al\., 2023](https://arxiv.org/html/2609.22599#bib.bib1);[Chen et al\., 2024](https://arxiv.org/html/2609.22599#bib.bib3);[Fu et al\., 2024](https://arxiv.org/html/2609.22599#bib.bib2);[Zhuge et al\., 2025](https://arxiv.org/html/2609.22599#bib.bib4)\)\. However, their judgments can be sensitive to incidental choices such as candidate order, prompt wording, and rubric presentation\([Zheng et al\., 2023](https://arxiv.org/html/2609.22599#bib.bib1);[Bellibatlu et al\., 2026](https://arxiv.org/html/2609.22599#bib.bib5);[Li et al\., 2026a](https://arxiv.org/html/2609.22599#bib.bib6)\)\. This sensitivity can undermine the reliability of the resulting judgments\. While a judge may agree with human labels on a static test set, it may change its decision under semantically irrelevant edits, or fail to adjust when an evaluated output is meaningfully degraded\.
Controlled perturbations expose these failures by testing whether a judge remains stable under irrelevant changes and responds appropriately to meaningful ones\([Ribeiro et al\., 2020](https://arxiv.org/html/2609.22599#bib.bib7);[Li et al\., 2026a](https://arxiv.org/html/2609.22599#bib.bib6)\)\. We distinguish four perturbation surfaces: the judge rubric, which defines the evaluation criteria and scoring semantics; the judge prompt, which specifies how to perform and return the judgment; the evaluated system’s input, such as a user query; and one or more candidate outputs, such as candidate responses to that query\. Existing approaches examine complementary subsets of these surfaces, motivating an integrated audit of both the judge instrument and the items it evaluates\.
We introduceMawile\(see Figure[1](https://arxiv.org/html/2609.22599#S1.F1)\), a developer\-facing workbench that integrates these tests into an interactive audit workflow\. Given a configured judge and representative evaluation items, a planning agent proposes applicable perturbation operators, which users can edit or extend with custom operators\.Mawilethen constructs variants of the selected judge or item fields\. Each operator specifies whether the semantic verdict should remain unchanged \(invariant\) or move in a prescribed direction \(directional\)\. A validation stage checks whether variants satisfy the intended constraints, helping prevent unintended changes from being misinterpreted as judge failures\. Finally,Mawilere\-executes the judge and reports sensitivity by evaluation surface, operator, and item, with the underlying evidence available for inspection\.
We make the following three contributions\. First, we combine item\-side and judge\-instrument reliability tests within a single end\-to\-end workflow\. Second, we introduce a typed probe interface with explicit applicability, expected\-relation, and validation semantics \(e\.g\., context\-aware modifications of input\-output pairs\)\. Third, we provide an interactive audit process that localizes failures and preserves the evidence needed to inspect and compare judge configurations\.Mawilesupports more reliable development and deployment of LLM\-based evaluators by making their failure modes easier to identify, diagnose, and compare\.
## 2Related Work
### 2\.1Perturbation\-Based Evaluation
Mawilebuilds on existing work that associates controlled transformations with an expected relation between the original and transformed predictions\([Ribeiro et al\., 2020](https://arxiv.org/html/2609.22599#bib.bib7)\)\. This principle has been applied to several components of LLM evaluation\. FBI\([Doddapaneni et al\., 2024](https://arxiv.org/html/2609.22599#bib.bib8)\)injects known quality degradations into evaluated responses to test whether judges detect them, while CALM\([Ye et al\., 2025](https://arxiv.org/html/2609.22599#bib.bib9)\)perturbs judge instructions or evaluated responses to quantify a catalog of biases\. JudgeSense\([Bellibatlu et al\., 2026](https://arxiv.org/html/2609.22599#bib.bib5)\)focuses on decision instability under semantically equivalent paraphrases of the judge prompt\. reWordBench\([Wu et al\., 2025](https://arxiv.org/html/2609.22599#bib.bib10)\)studies robustness under transformations of reward\-model prompts and responses\. Complementary rubric\-focused work studies score\-rubric ordering, preference drift under natural\-language rubric edits, and rubric auditing and repair\([Li et al\., 2026a](https://arxiv.org/html/2609.22599#bib.bib6);[Ding et al\., 2026](https://arxiv.org/html/2609.22599#bib.bib11)\)\.Mawileunifies these perturbation principles across judge and item fields, with each operator declaring its applicability, expected verdict relation, and validation requirements\.
### 2\.2Judge Audit Systems
The Judge Reliability Harness \(JRH\)\([Dev et al\., 2026](https://arxiv.org/html/2609.22599#bib.bib12)\)generates and validates reliability tests for binary and ordinal judges, aggregates metrics, and supports human review\. Its interventions primarily target evaluated responses or agent messages\. EvalSense\([Dejl and Pearson, 2026](https://arxiv.org/html/2609.22599#bib.bib13)\)provides an interactive framework for configuring and comparing evaluation methods through perturbation\-based meta\-evaluation\. RobustJudge\([Li et al\., 2026b](https://arxiv.org/html/2609.22599#bib.bib14)\)evaluates response attacks and defenses while examining judge models and prompt configurations\.
Mawileaudits a configured judge across its prompt, rubric, target\-system input, and candidate outputs within one editable workflow\. It supports context\-aware transformations of individual fields and coupled input–output pairs, and reports failures with their supporting evidence\. Table[1](https://arxiv.org/html/2609.22599#S2.T1)compares the intervention surfaces covered by these systems and related perturbation frameworks\.
Judge InstrumentEvaluated InstanceDeveloper\-FacingAudit SystemMethodRubricRRPromptPPInputxxOutputyyCoupledx,yx,yFBI\([Doddapaneni et al\., 2024](https://arxiv.org/html/2609.22599#bib.bib8)\)–––✓––reWordBench\([Wu et al\., 2025](https://arxiv.org/html/2609.22599#bib.bib10)\)––✓✓✓–CALM\([Ye et al\., 2025](https://arxiv.org/html/2609.22599#bib.bib9)\)–✓–✓––JudgeSense\([Bellibatlu et al\., 2026](https://arxiv.org/html/2609.22599#bib.bib5)\)–✓––––JRH\([Dev et al\., 2026](https://arxiv.org/html/2609.22599#bib.bib12)\)–––✓–✓EvalSense\([Dejl and Pearson, 2026](https://arxiv.org/html/2609.22599#bib.bib13)\)–––✓–✓RobustJudge\([Li et al\., 2026b](https://arxiv.org/html/2609.22599#bib.bib14)\)✓✓–✓–✓MAWILE✓✓✓✓✓✓Table 1:Comparison with prior judge\-auditing systems and perturbation frameworks\. A checkmark indicates that the method automatically constructs or applies controlled variants of the corresponding surface as an intervention target, not that a given surface is merely configurable\.
## 3Methodology
Mawileaudits a user\-supplied LLM judge by applying controlled changes to the judge instrument and the items it evaluates\. As shown in Figure[1](https://arxiv.org/html/2609.22599#S1.F1), an audit begins with a configured judge and a representative collection of evaluation items\.Mawilefirst establishes the judge’s baseline behavior and estimates judge consistency by repeatedly evaluating the unmodified items\. It then selects applicable perturbations, constructs and validates the corresponding variants, re\-executes the judge, and checks whether verdicts remain stable under invariant perturbations or move in the expected direction under directional perturbations\. The resulting report localizes sensitivity by evaluation surface, perturbation type, and item\.
### 3\.1Audit Setup
We represent a configured judge as𝒥=\(M,P,R,κ,𝒵\)\\mathcal\{J\}=\(M,P,R,\\kappa,\\mathcal\{Z\}\), whereMMspecifies the judge model and decoding configuration,PPis the evaluation prompt,RRis the rubric, the judge harnessκ\\kappaprocesses the raw judge response, and𝒵\\mathcal\{Z\}is the verdict space\. An evaluation item contains a target\-system inputxix\_\{i\}and one or more target\-system outputsyiy\_\{i\}\. A judge call produces
zi=κ\(M\(P,R,xi,yi\)\),zi∈𝒵\.z\_\{i\}=\\kappa\\bigl\(M\(P,R,x\_\{i\},y\_\{i\}\)\\bigr\),\\qquad z\_\{i\}\\in\\mathcal\{Z\}\.\(1\)This decomposition exposes four perturbation surfaces: the judge promptPP, judge rubricRR, target\-system inputxix\_\{i\}, and target\-system output\(s\)yiy\_\{i\}\. These four surfaces are varied within an audit, whileMM,κ\\kappa, and𝒵\\mathcal\{Z\}are held fixed\.
The verdict space𝒵\\mathcal\{Z\}depends on the type of judge\. Binary judges specify pass or fail, ordinal judges rank an output along a specified score scale, and pairwise judges choose which of two options is the better response to the input\.
Before applying perturbations,Mawilerepeatedly evaluates each original item\. These runs establish a baseline verdictbib\_\{i\}, defined as the modal class for categorical judges and the mean score for numeric judges\.Mawilealso measures disagreement across repeated calls to quantify the judge’s intrinsic stochasticity\. Together, these statistics help distinguish perturbation\-induced sensitivity from ordinary run\-to\-run variation\.
### 3\.2Typed Perturbation Operators
A perturbation operator perturbs one or more of the four audit surfaces\. Each operator declares its target surface, applicability conditions, generation procedure, validation procedure, and the expected relation between the original and perturbed verdicts as shown in Figure[1](https://arxiv.org/html/2609.22599#S1.F1)\.Mawilesupports two relation types\.Invariantoperators introduce perturbations that should not affect the semantic verdict, such as swapping the order of two candidates in pairwise evaluation\.Directionaloperators, conversely, introduce perturbations with a known expected effect on the verdict\. The operator specifies both the intended direction and the orientation of the judge’s output scale\. For example, adding a new requirement toxix\_\{i\}while leaving the corresponding output unchanged should lower the evaluation score\.
Directional operators are restricted to binary and ordinal judges\. In pairwise evaluation, degrading one candidate does not necessarily imply a preference reversal; establishing such an expectation would require verifying its quality relative to the unchanged candidate\. For pointwise judges, directional perturbations are only applied to positive items \(for binary judges\) or items not already at the bottom of the scale \(for ordinal judges\)\.
Together, these two relation types test two complementary properties\. Invariant probes measure robustness to changes that should preserve the underlying decision, while directional probes measure whether the judge responds to changes that should alter it\.
### 3\.3Audit Planning and Perturbation Construction
Mawileprovides a catalog of perturbation operators spanning the four evaluation surfaces\. The catalog includes deterministic transformations, such as candidate\-order swaps and structure\-preserving formatting changes, as well as LLM\-generated transformations, such as semantic paraphrases, controlled degradations, and task\-dependent modifications \(see Tables[3](https://arxiv.org/html/2609.22599#A1.T3)and[4](https://arxiv.org/html/2609.22599#A1.T4)for the full included perturbation catalog\)\. Users may select operators directly, add custom operators, or ask a planner agent to propose an audit based on the judge configuration and representative dataset items\. The proposed plan remains editable before execution, keeping task\-specific requirements under user control\.
LLM\-generated perturbations receive the context needed to preserve a coherent evaluation instance\. When modifying a target\-system outputyiy\_\{i\}, the generator also observes its corresponding inputxix\_\{i\}, and vice versa\. We use this context\-aware construction to avoid attributing judge failures to perturbations that instead create incoherent or implausible input–output pairs\. Operators may also modify bothxix\_\{i\}andyiy\_\{i\}jointly, enabling transformations such as consistent entity renaming to probe potential social biases in the judge\.
Validation checks whether each variant satisfies its operator\-specific requirements\. Some operators, such as candidate\-position swaps, satisfy their constraints by construction, as they preserve both candidate content and identity\. Variants requiring semantic assessment are instead checked by a separate validation agent\. Invariant perturbations must preserve the properties relevant to the judgment, while directional perturbations must introduce the specified degradation\. Rejected variants are excluded from judge execution\. The generator, validator, and judge are independently configurable roles\.
### 3\.4Report Generation
The final report aggregates results by evaluation surface, operator, and item\. Users can inspect the original and perturbed fields, canonical judge outputs, validation outcomes, and violations of each operator’s expected behavior, such as a preference change after a position swap or failure to lower a score after a validated degradation\. This supports analysis of both global sensitivity trends and local inspection of the items driving them\. If gold labels are available, the report also includes judge accuracy\. In addition to automatically computed metrics,Mawileprovides an LLM\-generated summary of key findings and recommended follow\-up actions to improve judge reliability and robustness \(Figure[7](https://arxiv.org/html/2609.22599#A1.F7)\)\.
Figure 2:Screenshot of theMawileWeb UI’s editable perturbation catalog\. During configuration, developers can inspect each operator’s family, construction method, target field, and expected verdict relation, then enable or disable it before execution\. A planner agent can also suggest perturbations tailored to the dataset\.Figure 3:Screenshot of theMawileweb UI’s per\-operator results for the MT\-Bench walkthrough using DeepSeek\-V4\-Flash as the judge and Kimi K3 as the perturbation generator\.
## 4User Interface Walkthrough
We illustrate theMawilesystem through an audit of DeepSeek\-V4\-Flash on 250 MT\-Bench items \(the same items used in Table[2](https://arxiv.org/html/2609.22599#S4.T2)\), using Kimi K3 to generate perturbations\.
#### Configuration\.
During the configuration step, the developer can set or override audit hyperparameters\. In this experiment, the developer defined a rubric that grades each item on a 1–10 scale , allows a one\-point change to count as invariant, and treats any score decrease as a valid directional degradation\. The planning agent proposed nine invariant and five directional operators, along with rationales for each choice\. For example, the planner excluded verbosity shifts from the invariant suite because response depth is explicitly included in the rubric and may legitimately affect the score\. It includes rubric reordering, which changes the presentation of the scoring criteria while preserving their content \(see Figure[5](https://arxiv.org/html/2609.22599#A1.F5)for the full justification\)\. In the editable catalog \(Figure[2](https://arxiv.org/html/2609.22599#S3.F2)\), the developer can review the selected operators and their expected effects, and enable or disable each operator\.
#### Generated Report\.
Mawileconstructed and validated the perturbations, then evaluated the accepted variants against the repeated baseline\. The report combines aggregate metrics, an LLM\-generated summary, per\-perturbation results, and items flagged for manual review\. Figure[3](https://arxiv.org/html/2609.22599#S3.F3)\(left\) shows that rubric reordering produces the highest invariant flip rate, at 15\.9% compared with the 4\.0% baseline flip rate on unmodified items\. The right panel shows that the judge is completely unable to detect the format\-violation errors, but is substantially more sensitive to injected unmet requirements\.
#### Evidence Inspection\.
The developer can also examine the examples underlying these reported findings\. In this experiment, the generated summary highlighted an item whose score drops from 9\.5 to 2\.0 \(averaged across multiple calls\) after rubric reordering\. The developer can review this case in the Review Queue tab \(Figure[8](https://arxiv.org/html/2609.22599#A1.F8)\) to inspect that item manually\. Under the Evidence tab, they can inspect examples of perturbed questions to verify that the results are coherent in context\. This supports both rapid interpretation of aggregate findings and deeper verification of the evidence behind them\.
BaselinePerturbation ResponseDatasetPerturbationGeneratorJudgeAccuracy↑\\uparrowRepeatFlip Rate↓\\downarrowInvariantFlip Rate↓\\downarrowDegradationDetection Rate↑\\uparrowMT\-BenchOrdinalQwen 3\.8 MaxGPT\-5\.6 Luna56\.02\.36\.062\.9DeepSeek\-V4\-Flash59\.15\.311\.765\.4Gemma 3 4B51\.20\.52\.847\.4Kimi K3GPT\-5\.6 Luna53\.63\.25\.765\.4DeepSeek\-V4\-Flash59\.04\.012\.268\.9Gemma 3 4B49\.60\.12\.453\.0Search ArenaPairwiseQwen 3\.8 MaxGPT\-5\.6 Luna62\.05\.310\.5–DeepSeek\-V4\-Flash64\.11\.68\.6–Gemma 3 4B53\.30\.316\.0–Kimi K3GPT\-5\.6 Luna58\.05\.210\.5–DeepSeek\-V4\-Flash66\.12\.68\.1–Gemma 3 4B53\.30\.413\.8–GSM8KBinaryQwen 3\.8 MaxGPT\-5\.6 Luna96\.80\.10\.741\.1DeepSeek\-V4\-Flash98\.40\.30\.348\.6Gemma 3 4B90\.40\.73\.910\.0Kimi K3GPT\-5\.6 Luna96\.00\.10\.754\.9DeepSeek\-V4\-Flash98\.00\.30\.667\.6Gemma 3 4B90\.80\.43\.69\.2Table 2:Aggregate audit results for ordinal, pairwise, and binary judges across two perturbation generators and three judge models\. Within each dataset, bold and underlining mark the best and second\-best judges in each metric, respectively\. All numbers are percentages\. For each dataset, a Kimi K3 planner selected one operator suite, which is held fixed across all six generator\-judge combinations\.
## 5Experimental Setup
### 5\.1Datasets
We evaluateMawileon three datasets chosen to span common judge output types, evaluation domains, and interaction settings, ranging from scalar scoring to pairwise preference and binary correctness judgments\. InMT\-Bench[Zheng et al\. \(2023\)](https://arxiv.org/html/2609.22599#bib.bib1), the judge grades an agent’s multi\-turn conversation with a user on a rubric of 1\-10\. InSearch Arena[Miroyan et al\. \(2026\)](https://arxiv.org/html/2609.22599#bib.bib15), given two agentic responses \(with traces\) to a user query that requires searching the internet, the judge evaluates which of the two is better\. InGSM8K[Cobbe et al\. \(2021\)](https://arxiv.org/html/2609.22599#bib.bib16), given a math problem and an agent’s answer and justification, the judge decides if the agent’s answer and decision process is correct or not\.
### 5\.2Model and Perturbation Setup
We run a full3×3×23\\times 3\\times 2comparison across three datasets, using GPT\-5\.6 Luna\([OpenAI, 2026](https://arxiv.org/html/2609.22599#bib.bib17)\), DeepSeek\-V4\-Flash\([DeepSeek\-AI, 2026](https://arxiv.org/html/2609.22599#bib.bib18)\), and Gemma 3 4B\([Gemma Team, 2025](https://arxiv.org/html/2609.22599#bib.bib19)\)as judges, and Qwen3\.8\-Max\([Qwen Team, 2026](https://arxiv.org/html/2609.22599#bib.bib20)\)and Kimi K3\([Kimi Team, 2026](https://arxiv.org/html/2609.22599#bib.bib21)\)as perturbation generators\. To isolate the effect of the perturbation generator, Kimi K3 selects one fixed set of perturbation operators for each dataset\. The same operator set is then used for every judge–generator combination\.
### 5\.3Metrics
We report four aggregate metrics\.*Accuracy*measures agreement between the baseline judgment and the provided correct answer\.*Repeat flip rate*measures disagreement across repeated calls on the unmodified item, capturing the judge’s intrinsic instability\.*Invariant flip rate*measures how often an accepted meaning\-preserving perturbation changes the judgment beyond the permitted tolerance\. Lastly,*degradation detection rate*measures how often a directional perturbation moves the judgment in the expected direction\.
## 6Experimental Results
#### Accuracy does not imply robustness\.
Across all six dataset–generator settings, DeepSeek\-V4\-Flash achieves the highest accuracy, followed by GPT\-5\.6 Luna and Gemma 3 4B\. The same ordering holds for degradation detection wherever directional probes are applicable, suggesting that more accurate judges are also generally better at recognizing meaningful degradations\. Invariant robustness, however, is task\-dependent: on MT\-Bench, Gemma 3 4B has the lowest invariant flip rates \(2\.4–2\.8pp\), followed by GPT\-5\.6 Luna \(5\.7–6\.0pp\) and DeepSeek\-V4\-Flash \(11\.7–12\.2pp\), while on Search Arena and GSM8K DeepSeek\-V4\-Flash is most robust and Gemma 3 4B least robust\. Thus, robustness to meaning\-preserving changes cannot be inferred from accuracy alone\.
#### Repeatability does not imply perturbation robustness\.
On Search Arena, Gemma 3 4B exhibits only 0\.3–0\.4% repeat flips on unmodified items, but 13\.8–16\.0% invariant flips after meaning\-preserving perturbations\. Similar gaps elsewhere show that perturbation sensitivity captures failures beyond ordinary run\-to\-run stochasticity\.
#### Directional probes are more sensitive to the perturbation generator\.
Switching between Qwen 3\.8 Max and Kimi K3 yields mean absolute differences of 0\.5 percentage points in invariant flip rate and 7\.5 points in degradation detection\. Kimi K3 improves detection for all judges on MT\-Bench and for GPT\-5\.6 Luna and DeepSeek\-V4\-Flash on GSM8K, suggesting that generator choice matters more for constructing meaningful degradations than invariant variants\.
## 7Conclusion
We introducedMawile, a developer\-facing system for auditing this distinction through validated perturbations to the judge prompt, rubric, evaluated input, and evaluated output\. Across ordinal, pairwise, and binary evaluations, our experiments show that judge sensitivity depends on both the model and task\. Higher accuracy is often associated with better detection of meaningful degradations, but does not necessarily imply greater robustness to irrelevant perturbations\. These findings motivate evaluating judge accuracy alongside repeatability and perturbation sensitivity\. By localizing failures to individual operators, items, and evaluation surfaces,Mawileprovides an inspectable basis for diagnosing unreliable judge behavior, including in settings where gold labels are unavailable\.
## Limitations
Mawilemeasures judge sensitivity, not necessarily judge correctness\. A judge that remains stable under every tested perturbation may still apply an incorrect rubric or reproduce systematic biases\. Gold labels can reveal some such failures when available, but label\-free audits cannot certify that a stable verdict is valid\.
The conclusions of an audit also depend on the coverage and validity of its perturbations\. The operator catalog represents only a subset of possible real\-world variation \(though we allow both manual and LLM\-authored custom perturbations to mitigate this\), while LLM\-generated perturbations and their validators may introduce their own errors or biases\. Passing the resulting suite is therefore evidence of reliability under the tested conditions, rather than a general guarantee\.
## References
- Bellibatluet al\.\(2026\)R\. R\. Bellibatlu, E\. Raff, and W\. ZhangJudgeSense: a benchmark for prompt sensitivity in llm\-as\-a\-judge systems\.External Links:2604\.23478,[Link](https://arxiv.org/abs/2604.23478)Cited by:[§1](https://arxiv.org/html/2609.22599#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.22599#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.22599#S2.T1.2.6.1)\.
- Chenet al\.\(2024\)D\. Chen, R\. Chen, S\. Zhang, Y\. Wang, Y\. Liu, H\. Zhou, Q\. Zhang, Y\. Wan, P\. Zhou, and L\. SunMLLM\-as\-a\-judge: assessing multimodal llm\-as\-a\-judge with vision\-language benchmark\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§1](https://arxiv.org/html/2609.22599#S1.p1.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.External Links:2110\.14168,[Link](https://arxiv.org/abs/2110.14168)Cited by:[§5\.1](https://arxiv.org/html/2609.22599#S5.SS1.p1.1)\.
- DeepSeek\-AI \(2026\)DeepSeek\-AIDeepSeek\-V4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.External Links:[Link](https://arxiv.org/abs/2606.19348)Cited by:[§5\.2](https://arxiv.org/html/2609.22599#S5.SS2.p1.1)\.
- Dejl and Pearson \(2026\)A\. Dejl and J\. PearsonEvalSense: a framework for domain\-specific LLM \(meta\-\)evaluation\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),D\. Croce, J\. Leidner, and N\. S\. Moosavi \(Eds\.\),Rabat, Marocco,pp\. 480–491\.External Links:[Link](https://aclanthology.org/2026.eacl-demo.33/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-demo.33),ISBN 979\-8\-89176\-382\-1Cited by:[§2\.2](https://arxiv.org/html/2609.22599#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.22599#S2.T1.2.8.1)\.
- Devet al\.\(2026\)S\. Dev, A\. Sloan, J\. Kavner, N\. Kong, and M\. SandlerJudge reliability harness: stress testing the reliability of llm judges\.External Links:2603\.05399,[Link](https://arxiv.org/abs/2603.05399)Cited by:[§2\.2](https://arxiv.org/html/2609.22599#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.22599#S2.T1.2.7.1)\.
- Dinget al\.\(2026\)R\. Ding, Y\. Pang, H\. Sun, Y\. Wang, S\. Wu, and Z\. DengRubrics as an attack surface: stealthy preference drift in LLM judges\.InThird Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=2STAJbBdJY)Cited by:[§2\.1](https://arxiv.org/html/2609.22599#S2.SS1.p1.1)\.
- Doddapaneniet al\.\(2024\)S\. Doddapaneni, M\. S\. U\. R\. Khan, S\. Verma, and M\. M\. KhapraFinding blind spots in evaluator LLMs with interpretable checklists\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 16279–16309\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.911/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.911)Cited by:[§2\.1](https://arxiv.org/html/2609.22599#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.22599#S2.T1.2.3.1)\.
- Fuet al\.\(2024\)J\. Fu, S\. Ng, Z\. Jiang, and P\. LiuGPTScore: evaluate as you desire\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 6556–6576\.External Links:[Link](https://aclanthology.org/2024.naacl-long.365/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.365)Cited by:[§1](https://arxiv.org/html/2609.22599#S1.p1.1)\.
- Gemma Team \(2025\)Gemma TeamGemma 3 technical report\.arXiv preprint arXiv:2503\.19786\.External Links:[Link](https://arxiv.org/abs/2503.19786)Cited by:[§5\.2](https://arxiv.org/html/2609.22599#S5.SS2.p1.1)\.
- Kimi Team \(2026\)Kimi TeamKimi K3: open frontier intelligence\.arXiv preprint arXiv:2607\.24653\.External Links:[Link](https://arxiv.org/abs/2607.24653)Cited by:[§5\.2](https://arxiv.org/html/2609.22599#S5.SS2.p1.1)\.
- Liet al\.\(2026a\)Q\. Li, S\. Dou, K\. Shao, C\. Chen, and H\. HuEvaluating scoring bias in llm\-as\-a\-judge\.InDatabase Systems for Advanced Applications,H\. Jung, T\. Wang, M\. Toyoda, H\. Kwon, and J\. Lee \(Eds\.\),Singapore,pp\. 19–34\.External Links:ISBN 978\-981\-92\-0372\-7Cited by:[§1](https://arxiv.org/html/2609.22599#S1.p1.1),[§1](https://arxiv.org/html/2609.22599#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22599#S2.SS1.p1.1)\.
- Liet al\.\(2026b\)S\. Li, C\. Xu, J\. Wang, X\. Gong, C\. Chen, J\. Zhang, J\. Wang, K\. Lam, and S\. JiLLMs cannot reliably judge \(yet?\): a comprehensive assessment on the robustness of llm\-as\-a\-judge\.External Links:2506\.09443,[Link](https://arxiv.org/abs/2506.09443)Cited by:[§2\.2](https://arxiv.org/html/2609.22599#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.22599#S2.T1.2.9.1)\.
- Miroyanet al\.\(2026\)M\. Miroyan, T\. Wu, L\. King, T\. Li, J\. Pan, X\. Hu, W\. Chiang, A\. N\. Angelopoulos, T\. Darrell, N\. Norouzi, and J\. E\. GonzalezSearch arena: analyzing search\-augmented LLMs\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=MMGRlDnhtI)Cited by:[§5\.1](https://arxiv.org/html/2609.22599#S5.SS1.p1.1)\.
- OpenAI \(2026\)OpenAIGPT\-5\.6 system card\.External Links:[Link](https://deploymentsafety.openai.com/gpt-5-6)Cited by:[§5\.2](https://arxiv.org/html/2609.22599#S5.SS2.p1.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.8\-Max: a new bar for coding and cowork\.External Links:[Link](https://qwen.ai/blog?id=qwen3.8)Cited by:[§5\.2](https://arxiv.org/html/2609.22599#S5.SS2.p1.1)\.
- Ribeiroet al\.\(2020\)M\. T\. Ribeiro, T\. Wu, C\. Guestrin, and S\. SinghBeyond accuracy: behavioral testing of NLP models with CheckList\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 4902–4912\.External Links:[Link](https://aclanthology.org/2020.acl-main.442/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.442)Cited by:[§1](https://arxiv.org/html/2609.22599#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22599#S2.SS1.p1.1)\.
- Wuet al\.\(2025\)Z\. Wu, M\. Yasunaga, A\. Cohen, Y\. Kim, A\. Celikyilmaz, and M\. GhazvininejadReWordBench: benchmarking and improving the robustness of reward models with transformed inputs\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 3383–3409\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.167/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.167),ISBN 979\-8\-89176\-332\-6Cited by:[§2\.1](https://arxiv.org/html/2609.22599#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.22599#S2.T1.2.4.1)\.
- Yeet al\.\(2025\)J\. Ye, Y\. Wang, Y\. Huang, D\. Chen, Q\. Zhang, N\. Moniz, T\. Gao, W\. Geyer, C\. Huang, P\. Chen, N\. V\. Chawla, and X\. ZhangJustice or prejudice? quantifying biases in LLM\-as\-a\-judge\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=3GTtZFiajM)Cited by:[§2\.1](https://arxiv.org/html/2609.22599#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.22599#S2.T1.2.5.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§1](https://arxiv.org/html/2609.22599#S1.p1.1),[§5\.1](https://arxiv.org/html/2609.22599#S5.SS1.p1.1)\.
- Zhugeet al\.\(2025\)M\. Zhuge, C\. Zhao, D\. R\. Ashley, W\. Wang, D\. Khizbullin, Y\. Xiong, Z\. Liu, E\. Chang, R\. Krishnamoorthi, Y\. Tian, Y\. Shi, V\. Chandra, and J\. SchmidhuberAgent\-as\-a\-judge: evaluate agents with agents\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 80569–80611\.External Links:[Link](https://proceedings.mlr.press/v267/zhuge25a.html)Cited by:[§1](https://arxiv.org/html/2609.22599#S1.p1.1)\.
## Appendix AAppendix
### A\.1Additional Walkthrough Screenshots
Figures[4](https://arxiv.org/html/2609.22599#A1.F4),[5](https://arxiv.org/html/2609.22599#A1.F5),[6](https://arxiv.org/html/2609.22599#A1.F6),[7](https://arxiv.org/html/2609.22599#A1.F7), and[8](https://arxiv.org/html/2609.22599#A1.F8)supplement the MT\-Bench walkthrough with additional views of audit configuration and result inspection\. They show the judge and dataset settings, the planner’s justifications, custom perturbation authoring, and the report overview and review queue\. These interfaces allow developers to inspect and revise the proposed perturbation suite before execution, then examine the individual items underlying the reported findings\.
Figure 4:Screenshot of theMawileweb UI for configuring a judge and audit dataset\.Figure 5:Screenshot of theMawileplanning agent’s proposed perturbation suite for the MT\-Bench walkthrough, with justification\.Figure 6:Screenshot of theMawileweb UI for defining a custom perturbation, with an example of an LLM\-suggested custom perturbation\.Figure 7:Screenshot of theMawilereport overview for the MT\-Bench walkthrough\.Figure 8:Screenshot of theMawilereview queue for inspecting flagged items from the MT\-Bench walkthrough\.
### A\.2Decoding Parameters
All models used reasoning medium effort and default temperature for all calls\.
### A\.3Perturbation Catalog
The built\-in catalog contains 16 invariant operators and 7 directional degradation operators \(Tables[3](https://arxiv.org/html/2609.22599#A1.T3)and[4](https://arxiv.org/html/2609.22599#A1.T4)\)\. Invariant probes should preserve the verdict; directional probes intentionally make an item worse and should therefore elicit a worse verdict\. “Rule” denotes a deterministic edit and “LLM” a perturbation\-model rewrite\. Semantic validation is independent of the judge under test and applies to both LLM rewrites and rule\-based edits when needed\. Judge\-level edits apply globally while item\-level edits are applied each item independently\.
Invariant Perturbation CatalogSurfacePerturbationGen\.DescriptionJudge promptInstruction orderingRuleMoves output\-format instructions before the prompt and rubric\.Reasoning styleLLMRequests direct, step\-by\-step, or criterion\-by\-criterion reasoning\.Role framingLLMChanges the evaluator persona without changing the task\.Prompt paraphraseLLMRewords the judge instructions while preserving their meaning\.Judge rubricRubric criterion reorderRuleReorders independent rubric blocks without changing their content\.Rubric paraphraseLLMRewords scoring criteria while preserving their meaning and thresholds\.Agent inputMinor typo or noiseLLMIntroduces minor prose typos without changing task meaning\.Input paraphraseLLMRewords the input while preserving the task and requirements\.Irrelevant distractor insertionLLMAdds plausible but irrelevant context to the input\.Context compressionLLMRemoves redundant input context while retaining needed information\.Agent outputOutput format conversionLLMChanges the response format while preserving the answer\.Politeness or tone shiftLLMChanges response tone while preserving its substance\.Uncertainty calibrationLLMChanges expressed confidence without changing the answer\.Verbosity shiftLLMShortens or lengthens the response without changing substantive claims\.Agent pairedPairwise position swapRuleSwaps candidate display positions while preserving candidate identities\.Entity renamingLLMRenames entities consistently across the input and output\.Table 3:Built\-in invariant perturbations\.Directional Perturbation CatalogSurfacePerturbationGen\.DescriptionAgent inputUnmet requirement injectionLLMAdds a task requirement that the existing response does not satisfy\.Agent outputPartial completionRuleRemoves the final separable part of a response\.Format violationLLMBreaks one previously satisfied output\-format requirement\.Unsupported claim insertionLLMAdds one unsupported substantive claim to the response\.Factual inconsistency insertionLLMIntroduces one factual or reasoning inconsistency\.Requirement omissionLLMRemoves or weakens content that satisfies a task requirement\.Over\-refusalLLMReplaces an answer to an answerable task with an unnecessary refusal\.Table 4:Built\-in directional degradation perturbations\. Directional probes are available only for pointwise audits\.相似文章
BenchMIRT:LLM 基准测试实际在测量什么?
BenchMIRT 介绍了一种方法,使用多维项目反应理论在单个提示词级别审计 LLM 基准测试,分离安全性和一般推理等底层能力,以揭示基准测试实际测量的内容。
换一个裁判,分数就变了:审计大模型作为裁判的可靠性
本文审计了大模型作为裁判(LLM-as-judge)评估的可靠性,表明即使候选回复固定不变,更换评估模型也可能改变评分。论文考察了Qwen3和MiniMax模型的扩展与升级路径,得出结论:裁判升级不可互换,并提出了最佳报告实践。
JudgeArena:可复现的LLM裁判评估统一框架
JudgeArena是一个开源框架,将主要的LLM裁判基准统一在单一接口下,支持对裁判选择的系统研究,并提供与闭源模型相当或更优的开源模型裁判,同时能够模拟LMArena Elo分数。
LLM Ass Bench
LLM Ass Bench 是一个用于评估大型语言模型的基准测试工具,专注于提示词。
MM-JudgeBias:评测 MLLM-as-a-Judge 组合偏差的基准
研究者发布 MM-JudgeBias 基准,揭示多模态大模型在充当自动评判器时的系统性组合偏差,对 26 个 SOTA MLLM 在 1,800 条样本上进行测试。