PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
Summary
PADMÉ reformulates meta-evaluation of LM agent evaluators as a preference alignment problem and introduces a data synthesis method using only small language models, boosting agreement with human judgment from 73% to 85% over a naive baseline while enabling evaluation of 25 common models across agentic domains.
View Cached Full Text
Cached at: 09/30/26, 09:50 AM
# PADMÉ: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
Source: [https://arxiv.org/html/2609.36086](https://arxiv.org/html/2609.36086)
\\workshoptitle
TAE \(Trust\-AI\-Eval\): Can We Trust AI Evaluation?
###### Abstract
Language models are frequently employed to evaluate other language models\. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment\. We call the problem of evaluating this alignmentMeta\-Evaluation\. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta\-evaluator recurses the question of trustworthiness\. We adopt a reformulation of meta\-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether theirimplied preferencesalign\. Building on this, we introduce PADMÉ, a data synthesis method that generates reliable criterion\-based meta\-evaluation data for agentic settings\. PADMÉ uses only small language models, requires no human involvement during evaluations, and operates under a low computational budget\. We build a prototype of PADMÉ and synthesize a dataset of 1,000 samples across four agentic domains and three evaluation criteria\. Human validation on a 150\-sample subset demonstrates that PADMÉ improves agreement with human judgment from 73% to 85% over a naive baseline\. Meta\-evaluating 25 common models with our dataset demonstrates the correlations between evaluation performance and scoring granularity, leniency, and model size, among other factors\.
## 1Introduction
Language models \(LMs\) are routinely used to evaluate LM agent systems111Here, we broadly define an LM agent system as a computer program that uses language models to hold multi\-turn conversations with users, call tools \(other computer programs\) to aid its work, and change the state of a virtual environment on behalf of its users\. Critically, it produces a trajectory rather than a single response during one user session\.in research and commercial use cases\[[1](https://arxiv.org/html/2609.36086#bib.bib1),[2](https://arxiv.org/html/2609.36086#bib.bib2),[3](https://arxiv.org/html/2609.36086#bib.bib3)\], and the evaluations are frequently applied over a set of distinct performance metrics\[[4](https://arxiv.org/html/2609.36086#bib.bib4),[5](https://arxiv.org/html/2609.36086#bib.bib5),[6](https://arxiv.org/html/2609.36086#bib.bib6)\]\. For example, the developers of a customer\-service agent on a shopping website would reasonably want to know how the agent scores, across thousands of conversations, on evaluation criteria such as friendliness, factuality, and answer relevance\. The evaluation criteria can be stylistic, open\-ended, or demand an intelligent understanding of whole conversations and their outcomes to be properly evaluated, so the environment lacks natural and deterministic signals for assessing them\[[2](https://arxiv.org/html/2609.36086#bib.bib2),[7](https://arxiv.org/html/2609.36086#bib.bib7)\]\. In these cases, LM evaluators may be the only viable option for automated evaluations\.
Figure 1:Overview of criterion\-specific meta\-evaluation\.Given multi\-step user\-agent interaction trajectories evaluated under a specific criterioncc\(e\.g\., communication clarity\), an LM evaluatorffassigns a continuous scoresf∈\[0,1\]s\_\{f\}\\in\[0,1\]\. Meta\-evaluation asks the question “Which of the evaluators \(s1,s2,s3s\_\{1\},s\_\{2\},s\_\{3\}\) should I trust?”, which PADMÉ attempts to solve through preference alignment\.For the scores provided by the LM evaluator to be useful, they must reliably reflect the actual performance of the agent over the criteria\. For agentic systems that mostly serve human users through conversations and environment interactions, the alignment between human and LM evaluator judgments is a natural estimate of the reliability of the LM evaluator\[[1](https://arxiv.org/html/2609.36086#bib.bib1),[3](https://arxiv.org/html/2609.36086#bib.bib3),[7](https://arxiv.org/html/2609.36086#bib.bib7)\]\. We call the assessment of this alignment theMeta\-Evaluationproblem \(Figure[1](https://arxiv.org/html/2609.36086#S1.F1)\)\.
Meta\-evaluation is well established for single\-turn response evaluations and reinforcement learning with human feedback\[[8](https://arxiv.org/html/2609.36086#bib.bib8),[9](https://arxiv.org/html/2609.36086#bib.bib9),[10](https://arxiv.org/html/2609.36086#bib.bib10),[11](https://arxiv.org/html/2609.36086#bib.bib11)\]\. It has only recently reached agentic settings in limited scenarios\[[2](https://arxiv.org/html/2609.36086#bib.bib2),[12](https://arxiv.org/html/2609.36086#bib.bib12),[13](https://arxiv.org/html/2609.36086#bib.bib13)\]\. A direct method would involve collecting human annotations for each evaluated agent over each criterion of interest\[[1](https://arxiv.org/html/2609.36086#bib.bib1),[7](https://arxiv.org/html/2609.36086#bib.bib7),[14](https://arxiv.org/html/2609.36086#bib.bib14)\], but human labeling is costly, which makes it especially prohibitive for rapidly developing agent platforms with a massive number of agent designs, use cases, and datasets\. Additionally, maintaining consistency when rating items on an absolute scale is hard for human annotators\[[15](https://arxiv.org/html/2609.36086#bib.bib15),[16](https://arxiv.org/html/2609.36086#bib.bib16)\]\. Another direct solution is to use a more capable model, a “meta\-evaluator”, to judge the evaluation of the LM evaluator\. For example, finetuned evaluators are commonly benchmarked against frontier LMs\[[17](https://arxiv.org/html/2609.36086#bib.bib17),[18](https://arxiv.org/html/2609.36086#bib.bib18),[19](https://arxiv.org/html/2609.36086#bib.bib19)\]\. However, this method requires access to stronger models, which can be prohibitive in itself\. It also recurses the meta\-evaluation problem, since nothing guarantees the reliability of the LM meta\-evaluator itself\[[7](https://arxiv.org/html/2609.36086#bib.bib7),[10](https://arxiv.org/html/2609.36086#bib.bib10)\]\.
In this paper, we present a simple, reliable, and inexpensive method for automatic meta\-evaluation of arbitrary LM evaluators\. We avoid the difficulties mentioned above with two designs\. First, instead of asking whether humans would give a trajectory the samescoreas the LM evaluator does, we hand the evaluator two trajectories to score separately, and ask whether thepreferenceimplied by the score difference agrees with the human preference \(Section[3\.2](https://arxiv.org/html/2609.36086#S3.SS2)\)\. This is an established practice in prior research \(Section[2](https://arxiv.org/html/2609.36086#S2)\)\. Second, we define an algorithm that synthesizes labeled trajectory pairs for a given agent system and an arbitrary set of criteria, using small language models \(SLMs\) no more capable than the ones the existing evaluator and agent system rely on \(Section[3\.1](https://arxiv.org/html/2609.36086#S3.SS1)\)\. No human input is needed during data synthesis or at evaluation time\.
We call the algorithmPADMÉ, or Preference Alignment Data synthesis for Meta\-Evaluation\. Building onτ3\\tau^\{3\}\-bench\[[20](https://arxiv.org/html/2609.36086#bib.bib20),[21](https://arxiv.org/html/2609.36086#bib.bib21),[22](https://arxiv.org/html/2609.36086#bib.bib22)\], we develop an instance of PADMÉ and use it to synthesize 1,000 data pairs across four task domains and three criteria \(Section[5\.1](https://arxiv.org/html/2609.36086#S5.SS1)\), validate 150 of them with six human annotators \(Section[5\.2](https://arxiv.org/html/2609.36086#S5.SS2)\), and evaluate 25 common models as evaluators \(Section[5\.3](https://arxiv.org/html/2609.36086#S5.SS3)\)\.
We claim three contributions\.
- •A data synthesis algorithmthat builds meta\-evaluation data for agent systems and any evaluation criteria, as opposed to a fixed benchmark \(Section[3](https://arxiv.org/html/2609.36086#S3)\)\.
- •A programthat implements this algorithm for a specific use case \(Section[4](https://arxiv.org/html/2609.36086#S4)\), together with a synthetic dataset generated by that program \(Section[5\.1](https://arxiv.org/html/2609.36086#S5.SS1)\)\.
- •Experimentsshowing that the algorithm is data\- and cost\-efficient, and a human study showing that the preference labels this algorithm constructs agree with human judgment in this scenario \(Section[5](https://arxiv.org/html/2609.36086#S5)\)\.
We release our code and the synthesized data at[https://github\.com/chc012/padme](https://github.com/chc012/padme)\.
## 2Related Work
#### Judging Agents, and Meta\-Evaluating the Judges
LLM\-as\-a\-judge is the paradigm for open\-ended evaluation\[[1](https://arxiv.org/html/2609.36086#bib.bib1),[3](https://arxiv.org/html/2609.36086#bib.bib3)\]\. Increasingly, LM evaluators target multi\-turn behaviors rather than single\-turn responses\[[2](https://arxiv.org/html/2609.36086#bib.bib2),[13](https://arxiv.org/html/2609.36086#bib.bib13),[12](https://arxiv.org/html/2609.36086#bib.bib12)\]\. Meta\-evaluation of these evaluators has evolved from coarse response\-level preferences\[[8](https://arxiv.org/html/2609.36086#bib.bib8),[9](https://arxiv.org/html/2609.36086#bib.bib9),[10](https://arxiv.org/html/2609.36086#bib.bib10)\]toward finer\-grained units, such as skill decompositions\[[4](https://arxiv.org/html/2609.36086#bib.bib4)\], checklists\[[23](https://arxiv.org/html/2609.36086#bib.bib23),[14](https://arxiv.org/html/2609.36086#bib.bib14),[24](https://arxiv.org/html/2609.36086#bib.bib24)\], rubrics\[[5](https://arxiv.org/html/2609.36086#bib.bib5),[25](https://arxiv.org/html/2609.36086#bib.bib25)\], instruction constraints\[[11](https://arxiv.org/html/2609.36086#bib.bib11)\], and crowdsourced criterion labels\[[26](https://arxiv.org/html/2609.36086#bib.bib26)\]\. For meta\-evaluation in agentic settings, Agent\-as\-a\-Judge\[[2](https://arxiv.org/html/2609.36086#bib.bib2)\]and AJ\-Bench\[[12](https://arxiv.org/html/2609.36086#bib.bib12)\]use verifier scripts and manual requirement annotations, so the label exists only in use cases where completion rules are predetermined\. They explore meta\-evaluation for code generation\[[2](https://arxiv.org/html/2609.36086#bib.bib2)\], search, data\-system manipulation, and GUI interaction\[[12](https://arxiv.org/html/2609.36086#bib.bib12)\]\. Two costs remain in all of these works: labels rely heavily on human curation, and data generation relies on frontier models\.
#### Constructed Quality Differences
LLMBar\[[27](https://arxiv.org/html/2609.36086#bib.bib27)\]is an early instance of this idea\. It releases419419hand\-curated output pairs in which one response follows the instruction and the other deviates\. FBI\[[28](https://arxiv.org/html/2609.36086#bib.bib28)\]injects hand\-authored perturbations that degrade one capability and asks whether an evaluator notices\. It requires manual label verification\. RubricEval\[[25](https://arxiv.org/html/2609.36086#bib.bib25)\]samples responses from a mixed model pool to elicit performance differences\. Its labels are judgments of binary questions regarding each trajectory\. REFLECT\[[29](https://arxiv.org/html/2609.36086#bib.bib29)\]meta\-evaluates judges of deep research agents\. It derives controlled perturbations of agent trajectories from a taxonomy of failure patterns\. Its pipeline requires human experts for validation\. All of these works build their data with frontier models or human annotations, which restricts extension to other agentic use cases\. We attempt to offer a more broadly applicable data synthesis recipe with only SLMs \(Section[4](https://arxiv.org/html/2609.36086#S4)\)\.
#### Preferences Versus Scores
Eliciting pairwise preference from pointwise scores is an established practice\. Comparative elicitation recovers a latent scale from pairwise judgments in psychology and statistics\[[30](https://arxiv.org/html/2609.36086#bib.bib30),[31](https://arxiv.org/html/2609.36086#bib.bib31)\]\. In natural language processing, it yields more reliable human labels than rating scales\[[15](https://arxiv.org/html/2609.36086#bib.bib15),[16](https://arxiv.org/html/2609.36086#bib.bib16)\]\. Pairwise ranking aligns LLM evaluators with human judgment better than direct scoring\[[32](https://arxiv.org/html/2609.36086#bib.bib32)\]\. Some reward\-model benchmarks rate each completion independently and compare score differences\[[9](https://arxiv.org/html/2609.36086#bib.bib9)\]\. Within LLM\-as\-a\-judge,[Zheng et al\. \[1\]](https://arxiv.org/html/2609.36086#bib.bib1)convert single\-answer grades into pairwise comparisons in order to measure agreement with human votes\. Consequently, we acquire ground\-truth data via preferences while prompting the evaluator to score each trajectory*pointwise*\(Eq\. \([3](https://arxiv.org/html/2609.36086#S3.E3)\)\), which mirrors deployment conditions\. Alternatively, many meta\-evaluation studies do measure scoring alignment by correlating evaluator scores directly against human ratings\[[33](https://arxiv.org/html/2609.36086#bib.bib33),[17](https://arxiv.org/html/2609.36086#bib.bib17),[4](https://arxiv.org/html/2609.36086#bib.bib4),[5](https://arxiv.org/html/2609.36086#bib.bib5)\]\.
#### SLMs as Judges
Deploying small language models as judges serves as a premise for our work\. SLM judges are competitive across model families and parameter scales\[[34](https://arxiv.org/html/2609.36086#bib.bib34)\], specialized mini\-evaluators have emerged\[[17](https://arxiv.org/html/2609.36086#bib.bib17),[35](https://arxiv.org/html/2609.36086#bib.bib35)\], and lightweight models are well\-suited for repetitive agentic sub\-tasks\[[36](https://arxiv.org/html/2609.36086#bib.bib36)\]\. However, SLMs are more susceptible to assertiveness and verbosity confounds\[[37](https://arxiv.org/html/2609.36086#bib.bib37)\]\. For SLM meta\-evaluation, SLMJury\[[34](https://arxiv.org/html/2609.36086#bib.bib34)\]meta\-evaluates SLM judges, but it uses existing public datasets rather than building agent\- and criterion\-specific data\.
## 3Method
Figure 2:Overview of the PADMÉ pipeline and meta\-evaluation architecture\.For each task\-criterion cell\(t,c\)\(t,c\)fromτ3\\tau^\{3\}\-bench, steering prompts generate trajectory pairsp=\(x−,x\+\)p=\(x^\{\-\},x^\{\+\}\)across three quality levels \(good,ok,bad\)\. Two LM filter judges discard pairs lacking visible quality contrast to yield dataset𝒟K\\mathcal\{D\}\_\{K\}, validated against a human annotation subset\. Target evaluators rate individual trajectories \(s=f\(x,c\)∈\[0,1\]s=f\(x;c\)\\in\[0,1\]\), and meta\-evaluation assesses their preference alignment across verified pairs\.As mentioned in Section[1](https://arxiv.org/html/2609.36086#S1), meta\-evaluating an evaluator requires assessing its alignment with human judgment\. Through the pointwise\-to\-pairwise reframing of the evaluation objective, we try to assess whether an evaluator will rank two agent trajectories correctly by scoring them independently\. PADMÉ synthesizes trajectory pairs with known labels over arbitrary criteria through steering agent prompts\. This shifts the core challenge from annotation to data collection\. To improve label fidelity, PADMÉ applies filtering and retries after initial trajectory generation\. We define the algorithm below\. Figure[2](https://arxiv.org/html/2609.36086#S3.F2)gives an overview of the method\.
### 3\.1Data Curation
#### Generation
We define the input of PADMÉ as acell\(t,c\)\(t,c\), consisting of task parametersttand evaluation criterioncc\.ttmay include textual descriptions of the task \(e\.g\., “rebook a plane ticket”\), tools and resources available to the agent for the task \(e\.g\., functionslist\_purchased\_ticketsandbuy\_ticket\), information about the simulated user’s intention of the task \(e\.g\., ‘‘if rebooking is not possible, cancel the ticket’’\), and so on222The selection of information available here is up to the developer of the specific agent system\. Care must be taken here to ensure that no unwanted information leakage happens inside the task parameters that would reveal the solution of the task directly to the trajectory\-generating agent\.\.ccconsists of the name and a description of the criterion\. The criterion can be any performance axis of the agent system, such as “task completion rate”, “user satisfaction”, “tool\-use relevance”, etc\. The descriptions should ideally be detailed and align with application scenarios\. As an example, see Appendix[H\.1](https://arxiv.org/html/2609.36086#A8.SS1)for the criteria descriptions in our experiment\. For good coverage of behavior, the source dataset should be a benchmark used to evaluate the targeted agent system, though any dataset that provides sets oftts andccs would apply\.
We use each cell as a seed datum to generate a trajectory pairp=\(x−,x\+\)p=\(x^\{\-\},x^\{\+\}\)with two different steering levels,ℓ,ℓ′∈ℒ=\{bad≺ok≺good\}\\ell,\\ell^\{\\prime\}\\in\\mathcal\{L\}=\\\{\\textsf\{bad\}\\prec\\textsf\{ok\}\\prec\\textsf\{good\}\\\}\. Here, a trajectory denotes a complete record of interactions between a user and an agent in a task execution session\. Trajectories are sampled using a rollout functionRoll\\mathrm\{Roll\}:
xℓ∼Roll\(t,sc,t,ℓ,θ\),p=\(x−,x\+\):=\(xℓ,xℓ′\)forℓ≺ℓ′\.x^\{\\ell\}\\sim\\mathrm\{Roll\}\\bigl\(t,s\_\{c,t,\\ell\};\\theta\\bigr\),\\qquad p=\(x^\{\-\},x^\{\+\}\):=\\bigl\(x^\{\\ell\},x^\{\\ell^\{\\prime\}\}\\bigr\)\\quad\\text\{for \}\\ell\\prec\\ell^\{\\prime\}\.\(1\)The steering instructionsc,t,ℓs\_\{c,t,\\ell\}is generated by aWriteSteerfunction, which uses an SLM conditioned on the criterion \(cc\), task \(tt\), and the steering intent \(ℓ\\ell\)\. The instruction is then appended to the agent’s system prompt \(Appendix[H\.3](https://arxiv.org/html/2609.36086#A8.SS3)\)\. Each cell randomly selects two distinct steering levels, with the three resulting contrasts \(\{bad,ok\}\\\{\\textsf\{bad\},\\textsf\{ok\}\\\},\{ok,good\}\\\{\\textsf\{ok\},\\textsf\{good\}\\\}, and\{bad,good\}\\\{\\textsf\{bad\},\\textsf\{good\}\\\}\) stratified equally across the dataset\. For example, abadsteering instruction on “friendliness” may ask the agent to be cold and concise in its response, anokone to be neutral in its tone, and agoodone to be warm and comforting in its replies\. Each instruction is tailored to the specific task and carries different information\. The generator prompt that produces these instructions is in Appendix[H\.2](https://arxiv.org/html/2609.36086#A8.SS2)\.
All rollouts within a cell share a fixed environment contextθ\\theta, which may include the agent model, user simulator, domain configuration, and decoding seed\. While fixingθ\\thetaensures the steering instruction is the primary deliberate variable, LM stochasticity and dynamic environment simulations naturally introduce trajectory variance\.
#### Filtering
Each generated trajectory pairp=\(x−,x\+\)p=\(x^\{\-\},x^\{\+\}\)is sequentially evaluated by a cascade ofKKSLM preference judges,J1,…,JKJ\_\{1\},\\dots,J\_\{K\}\(Appendix[H\.4](https://arxiv.org/html/2609.36086#A8.SS4)\)\. Given a pairppin randomized order \(Appendix[C\.2](https://arxiv.org/html/2609.36086#A3.SS2)\) and its associated criterioncc, each judge returns its own preference over the pair\. Its verdictJk\(p,c\)∈\{0,1\}J\_\{k\}\(p,c\)\\in\\\{0,1\\\}records whether that preference agrees with the constructed synthetic label\. Judges have veto power: a score of 1 retains the pair, while 0 discards it\. With𝒟0\\mathcal\{D\}\_\{0\}representing the initial set of generated pairs and𝒟k\\mathcal\{D\}\_\{k\}the subset surviving stagekk, the cascade progresses as:
𝒟k=\{p∈𝒟k−1∣Jk\(p,c\)=1\},k=1,…,K\.\\mathcal\{D\}\_\{k\}=\\bigl\\\{\\,p\\in\\mathcal\{D\}\_\{k\-1\}\\ \\mid\\ J\_\{k\}\(p,c\)=1\\,\\bigr\\\},\\quad k=1,\\dots,K\.\(2\)This construction forms a nested sequence𝒟K⊆⋯⊆𝒟1⊆𝒟0\\mathcal\{D\}\_\{K\}\\subseteq\\cdots\\subseteq\\mathcal\{D\}\_\{1\}\\subseteq\\mathcal\{D\}\_\{0\}, ensuring a pair is preserved only if approved by allKKjudges\. We setK=2K=2and evaluate intermediate depths \(Table[2](https://arxiv.org/html/2609.36086#S5.T2)\)\.
#### Retry
To improve data efficiency, a cell rejected by the filtering cascade is resampled with the same steering prompts and fresh rollouts, up toRRretries\. Resampled trajectories reuse the steering instructions to ensure level contrast stratification\. Since retries target cells previously rejected by the filter, subsequent attempts operate on inherently harder tasks, leading to an expected drop in retention rate but raising the overall data collection count\. Thus,RRserves as a hyperparameter trading increased dataset yield against compute cost\. We setR=2R=2in our experiment and analyze this trade\-off in Table[1](https://arxiv.org/html/2609.36086#S5.T1)\.
The complete dataset curation pipeline, integrating generation, filtering, and retries, is summarized in Algorithm[1](https://arxiv.org/html/2609.36086#alg1)\(Appendix[A](https://arxiv.org/html/2609.36086#A1)\)\.
### 3\.2Meta\-Evaluating an Evaluator
An evaluator under test is a scoring functionf\(x,c\)∈\[0,1\]f\(x;c\)\\in\[0,1\]that rates a single trajectoryxxagainst a criterioncc\. Becauseffevaluates each trajectory independently without viewing trajectory pairs, its pairwise preference overp=\(x−,x\+\)p=\(x^\{\-\},x^\{\+\}\)is derived directly from the score gap:
Δf\(p\)=f\(x\+,c\)−f\(x−,c\),y^f\(p\)=signΔf\(p\)∈\{\+1,−1,0\},\\Delta\_\{f\}\(p\)=f\(x^\{\+\};c\)\-f\(x^\{\-\};c\),\\qquad\\hat\{y\}\_\{f\}\(p\)=\\operatorname\{sign\}\\Delta\_\{f\}\(p\)\\in\\\{\+1,\-1,0\\\},\(3\)wherey^f\(p\)=0\\hat\{y\}\_\{f\}\(p\)=0indicates a tie in score\. This setup mirrors application scenarios where evaluators score individual execution traces without access to counterfactual rollouts\.
Because synthetic ground truth⋆\\staralways prefersx\+x^\{\+\}, an evaluator aligns with ground truth if and only if it assigns a strictly higher score tox\+x^\{\+\}\(Δf\(p\)\>0\\Delta\_\{f\}\(p\)\>0\)\. Ties \(Δf\(p\)=0\\Delta\_\{f\}\(p\)=0\) count as disagreements because a deployed evaluator that cannot separatex\+x^\{\+\}fromx−x^\{\-\}supplies no usable signal\. Evaluator accuracy is therefore defined as the agreement between the evaluator and the synthetic ground truth:
Agr\(f,⋆;𝒟\)=1\|𝒟\|∑p∈𝒟𝟏\[Δf\(p\)\>0\]\.\\mathrm\{Agr\}\(f,\\star;\\mathcal\{D\}\)=\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{p\\in\\mathcal\{D\}\}\\mathbf\{1\}\\bigl\[\\Delta\_\{f\}\(p\)\>0\\bigr\]\.\(4\)More broadly, we can formalize any pairwise decision source as a*preference provider*rrwith decision spacey^r\(p\)∈\{\+1,−1,0\}\\hat\{y\}\_\{r\}\(p\)\\in\\\{\+1,\-1,0\\\}\. Evaluators may output neutral ties \(y^f=0\\hat\{y\}\_\{f\}=0\), whereas human annotations and synthetic ground truth are strictly binary in\{−1,\+1\}\\\{\-1,\+1\\\}\. Agreement between any two providersrrandr′r^\{\\prime\}over a dataset𝒟\\mathcal\{D\}is:
Agr\(r,r′;𝒟\)=1\|𝒟\|∑p∈𝒟𝟏\[y^r\(p\)=y^r′\(p\)\]\.\\mathrm\{Agr\}\(r,r^\{\\prime\};\\mathcal\{D\}\)=\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{p\\in\\mathcal\{D\}\}\\mathbf\{1\}\\bigl\[\\,\\hat\{y\}\_\{r\}\(p\)=\\hat\{y\}\_\{r^\{\\prime\}\}\(p\)\\,\\bigr\]\.\(5\)
In the experiments below, we assess evaluator accuracyAgr\(f,⋆,𝒟\)\\mathrm\{Agr\}\(f,\\star;\\mathcal\{D\}\)\(Section[5\.3](https://arxiv.org/html/2609.36086#S5.SS3)\), synthetic label accuracy against human majority \(maj\\mathrm\{maj\}\) annotationsAgr\(maj,⋆,𝒟k\)\\mathrm\{Agr\}\(\\mathrm\{maj\},\\star;\\mathcal\{D\}\_\{k\}\)\(Section[5\.2](https://arxiv.org/html/2609.36086#S5.SS2)\), and fine\-grained evaluations across criteria, domains, level contrasts, and agent models \(Appendices[E](https://arxiv.org/html/2609.36086#A5)and[F\.1](https://arxiv.org/html/2609.36086#A6.SS1)\)\.
## 4Experimental Setup
#### Agent System & Criteria
We build an agent trajectory collection system upon a variant of theτ3\\tau^\{3\}\-bench repository\[[20](https://arxiv.org/html/2609.36086#bib.bib20),[21](https://arxiv.org/html/2609.36086#bib.bib21),[22](https://arxiv.org/html/2609.36086#bib.bib22)\]\. We rely on its task definitions, agent framework, user simulator, tool sets, and domain\-specific environments, but use our own LM evaluator system\. This setup mimics the expected usage of the PADMÉ algorithm in real\-world scenarios, where developers of an agent system bring in the agents, task parameters, and evaluation criteria of interest, and meta\-evaluate the evaluator over them with PADMÉ\. The base version ofτ3\\tau^\{3\}\-bench comprises 375 tasks across four task domains: airline \(50\), banking \(97\), retail \(114\), and telecom \(114\), each featuring different agent prompts, environments, and tools\. We evaluate three runtime criteria:*friendliness*,*communication clarity*, and*task resolution*\. Detailed descriptions of the criteria are in Appendix[H\.1](https://arxiv.org/html/2609.36086#A8.SS1)\. The task\-criterion matrix contains 1,125 evaluation cells, with quality level contrasts stratified evenly across the dataset\.
#### Pipeline Models & Data Curation
All pipeline components are driven by open\-weight SLMs ranging from 3B to 5\.1B active parameters \(21B to 117B total parameters\)\. We test agent trajectory pairs generated usinggpt\-oss\-20b,gpt\-oss\-120b, andnemotron\-lightning\-3\.5, with both rollouts in a pair generated by the same model\. User interactions are simulated viaqwen3\-30b\-a3b\-instruct\. Steering instructions are generated bygpt\-oss\-120b\. Filtering employsK=2K=2judges \(J1=nemotron\-lightning\-3\.5J\_\{1\}=\\text\{\{nemotron\-lightning\-3\.5\}\}andJ2=gpt\-oss\-120bJ\_\{2\}=\\text\{\{gpt\-oss\-120b\}\}\)\. Each data point is under a retry budget ofR=2R=2\. Running the 1,125 initial cells under this setup yields 1,000 kept pairs, distributed over the four dataset axes as Table[6](https://arxiv.org/html/2609.36086#A2.T6)in Appendix[B](https://arxiv.org/html/2609.36086#A2)shows\. For all open\-weight models, we use the Fireworks model deployment and serverless access service333[https://fireworks\.ai/models](https://fireworks.ai/models)\.\(Appendix[I](https://arxiv.org/html/2609.36086#A9)\)\.
#### Meta\-Evaluation Sweep
We evaluate 25 LMs as evaluators across open\-weight and proprietary model families on all 1,000 kept pairs across 3 runs \(Appendix[I](https://arxiv.org/html/2609.36086#A9)\)\. The open\-weight models contain 2B to 2\.8T parameters\. The evaluator prompt can be found in Appendix[H\.5](https://arxiv.org/html/2609.36086#A8.SS5)\. Model evaluations run under vendor\-default reasoning budgets and at temperature 0, unless temperature cannot be set444This refers to the more recent OpenAI GPT family of proprietary models\.\. Three of the 25 evaluators overlap with models used in the generation and filtering pipeline and are explicitly flagged in downstream analyses for potential contamination \(Appendix[F\.2](https://arxiv.org/html/2609.36086#A6.SS2)\)\.
#### Human Validation Protocol
Human validation is conducted on a stratified sample of 150 pairs from theinitial generation attempts \(r=0r=0\)across all criteria and domains \(representativeness checks in Appendix[D\.4](https://arxiv.org/html/2609.36086#A4.SS4)\)\. It serves exclusively to validate the data curation pipeline rather than operating as part of the pipeline itself\. Six human annotators form two panels of three on disjoint sets of 75 pairs, generating 450 binary preference choices along with confidence ratings\. The majority vote across the three annotators serves as the human reference label \(maj\\mathrm\{maj\}\)\. Additional information can be found in Appendix[D](https://arxiv.org/html/2609.36086#A4)\.
## 5Results and Analysis
### 5\.1Dataset
Out of 1,125 initial task\-criterion cells, 1,000 trajectory pairs survive both filter stages under a retry budget ofR=2R=2, requiring 1,673 total generation draws \(Table[1](https://arxiv.org/html/2609.36086#S5.T1)\)\. Per\-attempt retention decays predictably across retries, as subsequent retries operate exclusively on previously rejected cells\. Retries yield an additional 230 validated pairs \(\+30%\) at the cost of 548 supplementary trajectory rollouts\. Across all subsets, the first filter judgeJ1J\_\{1\}accounts for the vast majority of rejections\. Detailed subset breakdowns are provided in Appendix[C\.1](https://arxiv.org/html/2609.36086#A3.SS1)\.
The dataset curation pipeline generates 1,000 validated trajectory pairs at a total cost of $23\.63, or $0\.024 per kept pair\. Details of cost are in Appendix[G](https://arxiv.org/html/2609.36086#A7)\.
Table 1:Data curation funnel over 1,125 cells across retry attempts\.
### 5\.2Human Study
#### Annotator Confidence and Dataset Difficulty
Annotators report their confidence on each pair as 0 \(a guess\), 1 \(leaning\), or 2 \(certain\)\. We use the per\-pair mean over the three annotators as a proxy for how difficult that pair is to judge\. Mean self\-reported annotator confidence increases by\+0\.02\+0\.02on the 0 to 2 scale \(1%\) across filtering stages, though the influence ofJ1J\_\{1\}andJ2J\_\{2\}differs \(Appendix[D\.3](https://arxiv.org/html/2609.36086#A4.SS3)\)\. This suggests that filtering does not significantly trivialize the resulting data\.
#### Human Alignment with Labels
Filtering monotonically increases alignment between the synthetic label \(⋆\\star\) and human majority vote \(maj\\mathrm\{maj\}\), as computed by synthetic label accuracyAgr\(maj,⋆,𝒟k\)\\mathrm\{Agr\}\(\\mathrm\{maj\},\\star;\\mathcal\{D\}\_\{k\}\)with Eq\. \([5](https://arxiv.org/html/2609.36086#S3.E5)\)\. Passing pairs through the filters boosts label validity from 73\.3% to 84\.6% while shifting human\-label agreement \(Krippendorff’sα\\alpha\) from\+0\.468\+0\.468to\+0\.694\+0\.694\(Table[2](https://arxiv.org/html/2609.36086#S5.T2); panel reliability in Appendix[D\.2](https://arxiv.org/html/2609.36086#A4.SS2)\)\.
The improvement appears to plateau atK=2K=2\. On the subset of pairs rejected byJ1J\_\{1\}\(𝒟0∖𝒟1\\mathcal\{D\}\_\{0\}\\setminus\\mathcal\{D\}\_\{1\}\), label validity falls to 38\.9% \(14/36\), which confirms that the filter selectively removes misaligned or noisy trajectories\. Label validity amongJ2J\_\{2\}\-rejected pairs is 80\.0% \(8/10\), close to the 84\.6% among retained pairs\. Detailed subset breakdowns are provided in Appendix[E](https://arxiv.org/html/2609.36086#A5)\.
Table 2:Agreement between synthetic labels and human majority vote across filter depths𝒟𝐤\\mathbf\{\\mathcal\{D\}\_\{k\}\}\.Yield indicates the proportion of the 150 annotated pairs surviving at each depth\. Synthetic label accuracy reportsAgr\(maj,⋆,𝒟k\)\\mathrm\{Agr\}\(\\mathrm\{maj\},\\star;\\mathcal\{D\}\_\{k\}\)with bootstrap 95% confidence intervals in brackets\.α\\alphais Krippendorff’sα\\alphafor the same label\-versus\-majority decision\.
### 5\.3Meta\-Evaluation Sweep
To establish a performance spectrum across common language models, we benchmark 25 LMs with the same basic evaluator prompt \(Appendix[H\.5](https://arxiv.org/html/2609.36086#A8.SS5)\)\. We have each model score all 2,000 trajectories \(1,000 pairs\) independently across three separate runs\. Table[3](https://arxiv.org/html/2609.36086#S5.T3)lists 12 model results out of 25 for readability\. Two numbers directly inform the performance of each model:accuracy\(Equation[4](https://arxiv.org/html/2609.36086#S3.E4)\), our most important performance metric, andaverage score standard deviation \(SD\)555Note that average score SD is not the standard deviation of all scores assigned by each evaluator across the dataset, but the average of the standard deviation of each data point score across 3 runs for each evaluator\., which helps us understand the stability of the evaluator’s scoring\. All analyses in this section are conducted over the full 25 model results, which are reported in Table[18](https://arxiv.org/html/2609.36086#A6.T18)in the appendix\.
Table 3:Performance of evaluators across 1,000 trajectory pairs \(n=3n=3runs\)\.12 of 25 evaluators are selected to span the accuracy range and model families \(full view in Table[18](https://arxiv.org/html/2609.36086#A6.T18)\)\. All reported statistics are computed across all 25 evaluators\. Released indicates public release month\. Temperature = 0 except where unavailable‡\. Preferences derive from individual trajectory score gaps; ties count as incorrect\. Accuracy and score denote benchmark level accuracy and raw evaluator score, respectively\. SD is standard deviation\. Leniency is mean emitted score\. Average score SD measures per data point scoring stability across 3 runs\.Boldindicates best performance per column across all 25 evaluators\. Accuracy and average score SD serve as the two primary evaluator metrics\.parametersaccuracy \(%\)leniencyaveragetieby domain \(%\)by criterion \(%\)\# distinct\#evaluatorreleasedtotalactivereas\.mean±\\pmSD\(mean score\)score SD\(%\)airl\.bank\.retailtelec\.clar\.friend\.taskscores1claude\-opus\-52026\-07closedclosedYes85\.1±\\pm0\.930\.3240\.0234\.184828687739685693kimi\-k32026\-072\.8T104BYes83\.6±\\pm0\.120\.5080\.0366\.083808487719781376glm\-5p22026\-06753B∼\\sim30BYes81\.5±\\pm0\.500\.5210\.05110\.184818182789273318gpt\-5\.4\-nano‡2026\-03closedclosedNo79\.4±\\pm1\.160\.6310\.0527\.8817878817484806110gpt\-5\.6\-sol‡2026\-07closedclosedYes78\.7±\\pm0\.400\.5430\.0354\.9797781776497748611claude\-haiku\-4\-52025\-10closedclosedNo78\.7±\\pm0\.100\.5050\.00011\.6817976817685743213nemotron\-3\-ultra2026\-06549B55BYes76\.8±\\pm0\.760\.5580\.04414\.2758070817188712814llama3\.1\-70b2024\-0770\.6BdenseNo70\.0±\\pm0\.550\.6320\.01622\.8677367726469771415gemini\-3\.7\-flash2026\-08closedclosedYes69\.5±\\pm0\.450\.6150\.02623\.4717069695795543417gemma\-4\-31b2026\-0332\.2BdenseYes68\.1±\\pm0\.150\.6570\.02327\.4636766744792631418deepseek\-v4\-pro2026\-081\.6T49BYes67\.0±\\pm0\.670\.5050\.04520\.4627069656190492325qwen3\-4b2025\-084\.4BdenseNo36\.8±\\pm0\.650\.8040\.02956\.93634324430255517*Mean, all 25 evaluators*70\.168\.769\.969\.072\.262\.079\.567\.9
#### Subset Analysis
We discuss four subsets: evaluation criteria \(friendliness, communication clarity, task resolution\), domain \(retail, telecom, banking, airline\), level contrast \(bad\-ok,ok\-good,bad\-good\), and agent model \(nemotron\-lightning\-3\.5,gpt\-oss\-120b,gpt\-oss\-20b\)\.
Thebad–okcontrast \(73\.2%\) and thebad–goodcontrast \(73\.6%\) have similar accuracy, while the accuracy of theok–goodsubset is much lower, at 63\.4% \(Table[18](https://arxiv.org/html/2609.36086#A6.T18)\)\. A hypothesis is that, compared to detecting bad trajectories, it is harder for evaluators to distinguish the relative performance of two acceptable trajectories\.
gpt\-oss\-120b’s trajectories have an average accuracy of 74\.9%,gpt\-oss\-20b70\.5%, andnemotron\-lightning\-3\.565\.1% \(Table[18](https://arxiv.org/html/2609.36086#A6.T18)\)\. This suggests that the model backbone of the agent impacts the difficulty of evaluating the agent trajectories\.
#### Correlation Analysis
Table[4](https://arxiv.org/html/2609.36086#S5.T4)shows the correlation analyses between accuracy and seven factors\.
Accuracy is highly correlated with scoring resolution\. Tie rates range from 3\.4% to 56\.9% and correlate strongly with accuracy \(Spearman’sρ=−0\.946\\rho=\-0\.946\)\. In many cases, models fail due to coarse scoring granularity rather than misjudging\. Across models, distinct score count correlates with accuracy atρ=\+0\.797\\rho=\+0\.797, showing that score resolution is important for discriminative ability\.
Table 4:Seven evaluator properties against accuracy, sorted by Spearman’s correlation coefficients\.pholmp\_\{\\mathrm\{holm\}\}is the Holm\-Bonferroni corrected p\-value across the seven tests, andboldedwhere it indicates significance\.n=15n=15on some rows because ten closed models do not publish parameter counts\. Reasoning is a binary factor\.
The mean score an evaluator assigns to data points, which is a way of quantifying leniency, shows a significant anticorrelation with the evaluator performance \(ρ=−0\.725\\rho=\-0\.725\)\. It is possible that a harsher evaluator holds a longer internal list of expectations of the agent behavior, and is thus better at distinguishing nuanced differences between two trajectories\. An evaluator that is easily satisfied, i\.e\., assigns scores close to 1 easily, risks conflating good and great agent behaviors\.
Total parameter \(ρ=\+0\.725\\rho=\+0\.725\) and active parameter counts \(ρ=\+0\.518\\rho=\+0\.518, not significant\) show some level of correlation with accuracy\. For proprietary models, an exception is thatgpt\-5\.4\-nanoandgpt\-5\.4\-mini, which have smaller expected parameter counts, outperformgpt\-5\.6\-sol\.
## 6Limitations and Future Work
Extensibility to More Evaluation Criteria:Future work should apply PADMÉ to a larger and more diverse set of criteria\. It would especially benefit from a stress test of criteria that, through steering in a direction, would go against critical instruction\-following training of the agent models \(such as toxicity level or answer safety\)\.Broadening Agent and Domain Coverage:The current study limits agent trajectory generation toτ3\\tau^\{3\}\-bench and its domains\. To further validate the proposed method, it should be generalized across diverse agent architectures, operation environments, and broader domain benchmarks beyond customer\-service agent operations\.Score Granularity and Calibration:Given that evaluator accuracy is heavily driven by score resolution, future work should systematically evaluate controlled scoring regimes \(e\.g\., discrete Likert scales versus continuous\[0,1\]\[0,1\]bounds constrained to fixed decimal precision\) to explore how output formatting affects model ties and discrimination\.Other Reliability Signals:For criteria with programmatic rewards, reliability can be assessed without human annotations\. We briefly discuss the alignment between verifiable rewards with synthetic labels in Appendix[C\.3](https://arxiv.org/html/2609.36086#A3.SS3), but it warrants further investigation\.
## 7Conclusion
Meta\-evaluation is the evaluation of LM evaluator reliability\. Adopting preference alignment in place of absolute score alignment, we demonstrate that meta\-evaluation data can be synthesized rather than manually annotated\. The proposed framework, PADMÉ, steers an agent’s trajectories along criteria axes, uses the steering intents as the initial labels, and applies peer\-sized judges to veto trajectory contrasts\. Using models at or below 5\.1B active parameters, PADMÉ constructs a 1,000\-pair benchmark across four domains and three criteria at a low cost\. Filtering boosts human\-label agreement from 73\.3% to 84\.6%\. The resulting dataset separates 25 candidate evaluators across a 48\.3\-percentage\-point accuracy spread\. The data can be regenerated whenever criteria are modified or agents are updated without human intervention\. We thus present a generalizable data synthesis and meta\-evaluation recipe rather than a static benchmark\.
## Acknowledgments and Disclosure of Funding
All authors are employees at Uniphore\. All funding is provided by Uniphore\.
The authors would like to thank Ishika Agarwal, Bowen He, Tommy Li, Ethan Soon, Artin Tajdini, and Ming Xin \(ordered by last names\) for their help as human annotators\.
## References
- \[1\]Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E\. Gonzalez, and Ion Stoica\.Judging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.In*Thirty\-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track*, 2023\.URL[https://openreview\.net/forum?id=uccHPGDlao](https://openreview.net/forum?id=uccHPGDlao)\.
- \[2\]Mingchen Zhuge, Changsheng Zhao, Dylan R\. Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber\.Agent\-as\-a\-judge: Evaluate agents with agents\.In*Forty\-second International Conference on Machine Learning*, 2025\.URL[https://openreview\.net/forum?id=Nn9POI9Ekt](https://openreview.net/forum?id=Nn9POI9Ekt)\.
- \[3\]Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Zhouchi Lin, Bowen Zhang, Lionel Ni, Wen Gao, Yuanzhuo Wang, and Jian Guo\.A survey on llm\-as\-a\-judge\.*The Innovation*, 7\(6\):101253, 2026\.ISSN 2666\-6758\.doi:https://doi\.org/10\.1016/j\.xinn\.2025\.101253\.URL[https://www\.sciencedirect\.com/science/article/pii/S2666675825004564](https://www.sciencedirect.com/science/article/pii/S2666675825004564)\.
- \[4\]Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo\.FLASK: Fine\-grained language model evaluation based on alignment skill sets\.In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=CYmF38ysDa](https://openreview.net/forum?id=CYmF38ysDa)\.
- \[5\]Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo\.The BiGGen bench: A principled benchmark for fine\-grained evaluation of language models with language models\.In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pages 5877–5919, Albuquerque, New Mexico, April 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-189\-6\.doi:10\.18653/v1/2025\.naacl\-long\.303\.URL[https://aclanthology\.org/2025\.naacl\-long\.303/](https://aclanthology.org/2025.naacl-long.303/)\.
- \[6\]Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, and Wanli Ouyang\.MT\-bench\-101: A fine\-grained benchmark for evaluating large language models in multi\-turn dialogues\.In Lun\-Wei Ku, Andre Martins, and Vivek Srikumar, editors,*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 7421–7454, Bangkok, Thailand, August 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.acl\-long\.401\.URL[https://aclanthology\.org/2024\.acl\-long\.401/](https://aclanthology.org/2024.acl-long.401/)\.
- \[7\]Shreya Shankar, J\.D\. Zamfirescu\-Pereira, Bjoern Hartmann, Aditya Parameswaran, and Ian Arawjo\.Who validates the validators? aligning llm\-assisted evaluation of llm outputs with human preferences\.In*Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology*, UIST ’24, New York, NY, USA, 2024\. Association for Computing Machinery\.ISBN 9798400706288\.doi:10\.1145/3654777\.3676450\.URL[https://doi\.org/10\.1145/3654777\.3676450](https://doi.org/10.1145/3654777.3676450)\.
- \[8\]Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A\. Smith, and Hannaneh Hajishirzi\.RewardBench: Evaluating reward models for language modeling\.In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,*Findings of the Association for Computational Linguistics: NAACL 2025*, pages 1755–1797, Albuquerque, New Mexico, April 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-195\-7\.doi:10\.18653/v1/2025\.findings\-naacl\.96\.URL[https://aclanthology\.org/2025\.findings\-naacl\.96/](https://aclanthology.org/2025.findings-naacl.96/)\.
- \[9\]Saumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison, Noah A\. Smith, Hannaneh Hajishirzi, and Nathan Lambert\.Rewardbench 2: Advancing reward model evaluation\.In*The Fourteenth International Conference on Learning Representations*, 2026\.URL[https://openreview\.net/forum?id=fb0G86Dewb](https://openreview.net/forum?id=fb0G86Dewb)\.
- \[10\]Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang, Alejandro Cuadron, Chenguang Wang, Raluca Popa, and Ion Stoica\.Judgebench: A benchmark for evaluating LLM\-based judges\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=G0dksFayVq](https://openreview.net/forum?id=G0dksFayVq)\.
- \[11\]Bosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling, Ying Zhang, Pei Ke, Hongning Wang, and Minlie Huang\.IF\-RewardBench: Benchmarking judge models for instruction\-following evaluation\.In Maria Liakata, Viviane P\. Moreira, Jiajun Zhang, and David Jurgens, editors,*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 23816–23843, San Diego, California, United States, July 2026\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-390\-6\.doi:10\.18653/v1/2026\.acl\-long\.1092\.URL[https://aclanthology\.org/2026\.acl\-long\.1092/](https://aclanthology.org/2026.acl-long.1092/)\.
- \[12\]Wentao Shi, Yu Wang, Yuyang Zhao, Yuxin Chen, Fuli Feng, Xueyuan Hao, Xi Su, Qi GU, Hui Su, Xunliang Cai, and Xiangnan He\.AJ\-bench: Benchmarking agent\-as\-a\-judge for environment\-aware evaluation\.In Maria Liakata, Viviane P\. Moreira, Jiajun Zhang, and David Jurgens, editors,*Findings of the Association for Computational Linguistics: ACL 2026*, pages 25371–25413, San Diego, California, United States, July 2026a\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-395\-1\.doi:10\.18653/v1/2026\.findings\-acl\.1269\.URL[https://aclanthology\.org/2026\.findings\-acl\.1269/](https://aclanthology.org/2026.findings-acl.1269/)\.
- \[13\]Hyogon Ryu, Jeonghwan Kim, Yewon Lim, Chaeun Lee, Jeongwook Kim, and Donghoon Ham\.Online agent\-as\-a\-judge: Situation\-generating evaluation for interactive agents\.In*Trustworthy AI for Good \(AI4GOOD\) Workshop @ ICML 2026*, 2026\.URL[https://openreview\.net/forum?id=YrcMknjgyx](https://openreview.net/forum?id=YrcMknjgyx)\.
- \[14\]Tianjun Wei, Wei Wen, Ruizhi Qiao, Xing Sun, and Jianghong Ma\.Rocketeval: Efficient automated LLM evaluation via grading checklist\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=zJjzNj6QUe](https://openreview.net/forum?id=zJjzNj6QUe)\.
- \[15\]Svetlana Kiritchenko and Saif Mohammad\.Best\-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation\.In Regina Barzilay and Min\-Yen Kan, editors,*Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pages 465–470, Vancouver, Canada, July 2017\. Association for Computational Linguistics\.doi:10\.18653/v1/P17\-2074\.URL[https://aclanthology\.org/P17\-2074/](https://aclanthology.org/P17-2074/)\.
- \[16\]Jekaterina Novikova, Ondřej Dušek, and Verena Rieser\.RankME: Reliable human ratings for natural language generation\.In Marilyn Walker, Heng Ji, and Amanda Stent, editors,*Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 \(Short Papers\)*, pages 72–78, New Orleans, Louisiana, June 2018\. Association for Computational Linguistics\.doi:10\.18653/v1/N18\-2012\.URL[https://aclanthology\.org/N18\-2012/](https://aclanthology.org/N18-2012/)\.
- \[17\]Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo\.Prometheus 2: An open source language model specialized in evaluating other language models\.In Yaser Al\-Onaizan, Mohit Bansal, and Yun\-Nung Chen, editors,*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 4334–4353, Miami, Florida, USA, November 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.emnlp\-main\.248\.URL[https://aclanthology\.org/2024\.emnlp\-main\.248/](https://aclanthology.org/2024.emnlp-main.248/)\.
- \[18\]Lianghui Zhu, Xinggang Wang, and Xinlong Wang\.JudgeLM: Fine\-tuned large language models are scalable judges\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=xsELpEPn4A](https://openreview.net/forum?id=xsELpEPn4A)\.
- \[19\]Yidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, Wei Ye, Shikun Zhang, and Yue Zhang\.PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization\.In*The Twelfth International Conference on Learning Representations*, 2024a\.URL[https://openreview\.net/forum?id=5Nn2BLV7SB](https://openreview.net/forum?id=5Nn2BLV7SB)\.
- \[20\]Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik R Narasimhan\.τ\\tau\-bench: A benchmark forTool\-Agent\-User interaction in real\-world domains\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=roNSXZpUDN](https://openreview.net/forum?id=roNSXZpUDN)\.
- \[21\]Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik R Narasimhan\.τ2\\tau^\{2\}\-bench: Evaluating conversational agents in a dual\-control environment\.In*Forty\-third International Conference on Machine Learning*, 2026\.URL[https://openreview\.net/forum?id=OC2z7iSQKa](https://openreview.net/forum?id=OC2z7iSQKa)\.
- \[22\]Quan Shi, Alexandra Zytek, Pedram Razavi, Karthik R Narasimhan, and Victor Barres\.τ\\tau\-knowledge: Evaluating conversational agents over unstructured knowledge\.In*Forty\-third International Conference on Machine Learning*, 2026b\.URL[https://openreview\.net/forum?id=XHZK5abtw2](https://openreview.net/forum?id=XHZK5abtw2)\.
- \[23\]Jonathan Cook, Tim Rocktäschel, Jakob Foerster, Dennis Aumiller, and Alex Wang\.Ticking all the boxes: Generated checklists improve llm evaluation and generation, 2024\.URL[https://arxiv\.org/abs/2410\.03608](https://arxiv.org/abs/2410.03608)\.
- \[24\]Karen Zhou and Chenhao Tan\.AutoChecklist: Composable pipelines for checklist generation and scoring with LLM\-as\-a\-judge\.In Greg Durrett and Ping Jian, editors,*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\)*, pages 515–525, San Diego, California, United States, July 2026\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-392\-0\.doi:10\.18653/v1/2026\.acl\-demo\.51\.URL[https://aclanthology\.org/2026\.acl\-demo\.51/](https://aclanthology.org/2026.acl-demo.51/)\.
- \[25\]Tianjun Pan, Xuan Lin, Wenyan Yang, Qianyu He, Shisong Chen, Licai Qi, Wanqing Xu, Hongwei Feng, Bo Xu, and Yanghua Xiao\.Rubriceval: A rubric\-level meta\-evaluation benchmark for llm judges in instruction following, 2026\.URL[https://arxiv\.org/abs/2603\.25133](https://arxiv.org/abs/2603.25133)\.
- \[26\]Zhilin Wang, Yi Dong, Olivier Delalleau, Jiaqi Zeng, Gerald Shen, Daniel Egert, Jimmy J\. Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev\.Helpsteer 2: Open\-source dataset for training top\-performing reward models\.In A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang, editors,*Advances in Neural Information Processing Systems*, volume 37, pages 1474–1501\. Curran Associates, Inc\., 2024b\.doi:10\.52202/079017\-0047\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2024/file/02fd91a387a6a5a5751e81b58a75af90\-Paper\-Datasets\_and\_Benchmarks\_Track\.pdf](https://proceedings.neurips.cc/paper_files/paper/2024/file/02fd91a387a6a5a5751e81b58a75af90-Paper-Datasets_and_Benchmarks_Track.pdf)\.
- \[27\]Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen\.Evaluating large language models at evaluating instruction following\.In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://arxiv\.org/abs/2310\.07641](https://arxiv.org/abs/2310.07641)\.
- \[28\]Sumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Sshubam Verma, and Mitesh M Khapra\.Finding blind spots in evaluator LLMs with interpretable checklists\.In Yaser Al\-Onaizan, Mohit Bansal, and Yun\-Nung Chen, editors,*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 16279–16309, Miami, Florida, USA, November 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.emnlp\-main\.911\.URL[https://aclanthology\.org/2024\.emnlp\-main\.911/](https://aclanthology.org/2024.emnlp-main.911/)\.
- \[29\]Leyao Wang, Yanan He, Peng Chen, Asaf Yehudai, Yixin Liu, Rex Ying, Michal Shmueli\-Scheuer, and Arman Cohan\.Time to reflect: Can we trust llm judges for evidence\-based research agents?, 2026\.URL[https://arxiv\.org/abs/2605\.19196](https://arxiv.org/abs/2605.19196)\.
- \[30\]L\. L\. Thurstone\.A law of comparative judgment\.*Psychological Review*, 34\(4\):273–286, 1927\.doi:10\.1037/h0070288\.
- \[31\]Ralph A\. Bradley and Milton E\. Terry\.Rank analysis of incomplete block designs: I\. the method of paired comparisons\.*Biometrika*, 39\(3/4\):324–345, 1952\.doi:10\.2307/2334029\.
- \[32\]Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulić, Anna Korhonen, and Nigel Collier\.Aligning with human judgement: The role of pairwise preference in large language model evaluators\.In*First Conference on Language Modeling*, 2024\.URL[https://openreview\.net/forum?id=9gdZI7c6yr](https://openreview.net/forum?id=9gdZI7c6yr)\.
- \[33\]Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu\.G\-eval: NLG evaluation using gpt\-4 with better human alignment\.In Houda Bouamor, Juan Pino, and Kalika Bali, editors,*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pages 2511–2522, Singapore, December 2023\. Association for Computational Linguistics\.doi:10\.18653/v1/2023\.emnlp\-main\.153\.URL[https://aclanthology\.org/2023\.emnlp\-main\.153/](https://aclanthology.org/2023.emnlp-main.153/)\.
- \[34\]Anish Laddha, Nitesh Pradhan, and Gaurav Srivastava\.Slmjury: Can small language models judge as well as large ones?, 2026\.URL[https://arxiv\.org/abs/2606\.07810](https://arxiv.org/abs/2606.07810)\.
- \[35\]Andrei Alexandru, Antonia Calvi, Henry Broomfield, Jackson Golden, Kyle Dai, Mathias Leys, Maurice Burger, Max Bartolo, Roman Engeler, Sashank Pisupati, Toby Drane, and Young Sun Park\.Atla selene mini: A general purpose evaluation model, 2025\.URL[https://arxiv\.org/abs/2501\.17195](https://arxiv.org/abs/2501.17195)\.
- \[36\]Peter Belcak, Greg Heinrich, Shizhe Diao, Yonggan Fu, Xin Dong, Saurav Muralidharan, Yingyan Celine Lin, and Pavlo Molchanov\.Small language models are the future of agentic ai, 2025\.URL[https://arxiv\.org/abs/2506\.02153](https://arxiv.org/abs/2506.02153)\.
- \[37\]Tuhina Tripathi, Manya Wadhwa, Greg Durrett, and Scott Niekum\.Pairwise or pointwise? evaluating feedback protocols for bias in LLM\-based evaluation\.In*Second Conference on Language Modeling*, 2025\.URL[https://openreview\.net/forum?id=uyX5Vnow3U](https://openreview.net/forum?id=uyX5Vnow3U)\.
- \[38\]Yuzheng Xu, Tosho Hirasawa, Tadashi Kozuno, and Yoshitaka Ushiku\.Am i more pointwise or pairwise? revealing position bias in rubric\-based llm\-as\-a\-judge, 2026\.URL[https://arxiv\.org/abs/2602\.02219](https://arxiv.org/abs/2602.02219)\.
## Appendix ANotations and the Algorithm
Table 5:Notation, grouped by pipeline stage in the order the paper uses it\.Table[5](https://arxiv.org/html/2609.36086#A1.T5)collects the notation used in the main paper, grouped by pipeline stage\.
Algorithm 1PADMÉ data curation algorithm\.Each cell yields at most one kept pair\.1:cells
\{\(t,c\)\}\\\{\(t,c\)\\\}, fixed context
θ\\theta, judges
J1,…,JKJ\_\{1\},\\dots,J\_\{K\}, retry budget
RR
2:
𝒟k←∅\\mathcal\{D\}\_\{k\}\\leftarrow\\emptysetfor
k=0,…,Kk=0,\\dots,K
3:for allcells
\(t,c\)\(t,c\)do
4:take the level contrast
ℓ≺ℓ′\\ell\\prec\\ell^\{\\prime\}for this cell from the stratified schedule
5:for
attempt=1\\text\{attempt\}=1to
1\+R1\+Rdo
6:
sc,t,ℓ←WriteSteer\(c,t,ℓ\)s\_\{c,t,\\ell\}\\leftarrow\\textsc\{WriteSteer\}\(c,t,\\ell\),
sc,t,ℓ′←WriteSteer\(c,t,ℓ′\)s\_\{c,t,\\ell^\{\\prime\}\}\\leftarrow\\textsc\{WriteSteer\}\(c,t,\\ell^\{\\prime\}\)
7:
x−∼Roll\(t,sc,t,ℓ,θ\)x^\{\-\}\\sim\\mathrm\{Roll\}\(t,\\,s\_\{c,t,\\ell\};\\,\\theta\),
x\+∼Roll\(t,sc,t,ℓ′,θ\)x^\{\+\}\\sim\\mathrm\{Roll\}\(t,\\,s\_\{c,t,\\ell^\{\\prime\}\};\\,\\theta\)
8:
p←\(x−,x\+\)p\\leftarrow\(x^\{\-\},x^\{\+\}\),
𝒟0←𝒟0∪\{p\}\\mathcal\{D\}\_\{0\}\\leftarrow\\mathcal\{D\}\_\{0\}\\cup\\\{p\\\}
9:
k←0k\\leftarrow 0
10:while
k<Kk<Kand
Jk\+1\(p,c\)=1J\_\{k\+1\}\(p,c\)=1do
11:
k←k\+1k\\leftarrow k\+1,
𝒟k←𝒟k∪\{p\}\\mathcal\{D\}\_\{k\}\\leftarrow\\mathcal\{D\}\_\{k\}\\cup\\\{p\\\}⊳\\trianglerighta veto stops the cascade
12:endwhile
13:if
k=Kk=Kthen
14:break⊳\\trianglerightcell filled; only a rejected cell is ever redrawn
15:endif
16:endfor
17:endfor
18:return
𝒟K\\mathcal\{D\}\_\{K\}⊳\\triangleright𝒟0,…,𝒟K−1\\mathcal\{D\}\_\{0\},\\dots,\\mathcal\{D\}\_\{K\-1\}are kept for the filter\-depth ablation
The complete PADMÉ dataset curation pipeline, integrating generation, filtering, and retries, is summarized in Algorithm[1](https://arxiv.org/html/2609.36086#alg1)\.
## Appendix BDataset Composition and Evaluation Protocols
Table[6](https://arxiv.org/html/2609.36086#A2.T6)outlines the composition of the 1,000\-pair shipping dataset across its four principal axes\. Table[7](https://arxiv.org/html/2609.36086#A2.T7)details the input constraints, task formulations, and output formats for each preference provider\.
The pipeline naturally produces a balanced dataset across criteria, domains, level contrasts, and agent models without requiring artificial quotas or post\-hoc rebalancing, maintaining strong yield consistency across subsets \(detailed further in Table[8](https://arxiv.org/html/2609.36086#A3.T8)and Appendix[C\.1](https://arxiv.org/html/2609.36086#A3.SS1)\)\.
The filter judge \(JkJ\_\{k\}\) and the human annotator \(aa\) operate under identical information availability\. Both receive identical contextual fields, perform pairwise trajectory comparisons, and remain strictly blinded to steering metadata\. This symmetry ensures that human label alignment serves as a direct validation of the cascade filter rather than an artifact of information asymmetry\. In contrast, the evaluator under test \(ff\) operates pointwise on single trajectories without cross\-trajectory visibility, which mirrors realistic deployment conditions\.
Table 6:A glossary of dataset subsets, along with the distribution of the 1,000 pairs across axes\.Each pair carries exactly one level per axis, resulting in subtotal sums of 1,000 per block\. Comprehensive retention rates per subset appear in Table[8](https://arxiv.org/html/2609.36086#A3.T8)\. Evaluation criterion descriptions reflect the full runtime definitions, also provided in Appendix[H\.1](https://arxiv.org/html/2609.36086#A8.SS1)\. Stratified experimental breakdowns are reported in Appendix[E](https://arxiv.org/html/2609.36086#A5)\.levelpairsdescription*Evaluation criterion*friendliness355Warmth and consideration toward the customer: whether the agent acknowledges their situation and how they feel about it, delivers unwelcome news with care, and leaves them feeling attended to\. Judge the manner, not whether the request was resolved\.task resolution329Whether the customer’s actual problem was settled: did the agent establish what was needed, take the actions that would resolve it, and leave the customer with the outcome they came for\. A correct refusal counts as resolution – if the request was not permitted, saying so plainly and explaining why resolves it, while quietly doing it anyway does not\. Judge the outcome, not the manner or how well it was explained\.communication clarity316How easily the customer can follow the agent: whether the main point is findable, whether technical or policy language is explained, whether multi\-part information is organised, and whether the customer is left knowing what is true and what happens next\. Judge the presentation, not the warmth or the outcome\.*Domain*retail309114 tasks; 1,158\-word policy, 16 agent toolstelecom294114 tasks; 3,715\-word policy, 13 agent tools, 30 user\-side toolsbanking26297 tasks; 926\-word policy, 16 agent tools, including retrievalairline13550 tasks; 1,313\-word policy, 14 agent tools*Level contrast*bad–good359wider gapok–good328narrower gapbad–ok313narrower gap*Agent model*nemotron\-lightning\-3\.534132B total, 3B activegpt\-oss\-120b334116\.8B total, 5\.1B activegpt\-oss\-20b32520\.9B total, 3\.6B activeTable 7:A glossary of information exposure, task objectives, and output specifications across preference providers\.Symmetrical blinding ensures human annotators directly benchmark filter judge decisions, whereas evaluators under test operate in single\-trajectory pointwise mode to match real\-world deployment\.
## Appendix CDataset Audits and Yield Dynamics
### C\.1Generation Yield and Retention
Generation retention is reported per experimental subset rather than artificially constrained\. Because no post\-hoc rebalancing is applied to force quota targets, Table[8](https://arxiv.org/html/2609.36086#A3.T8)directly reflects where trajectory pairs easily satisfy judge filtering versus where narrower quality gaps require higher generation compute and retries\.
Retention yield strictly tracks two primary factors: the target level contrast and the specific evaluation criterion\. First, wider level contrasts yield higher retention rates:bad–goodcontrasts complete 96% of targeted cells at 1\.38 draws per retained pair, compared to 84% completion at 1\.90 draws per pair for narrowerbad–okcontrasts\. Second, criteria evaluating overt linguistic style achieve higher retention than structural task completion metrics:friendlinessreaches 95% cell retention, outperformingtask\_resolution\(88%\) andcommunication\_clarity\(84%\)\. Across domains, retention rates remain tightly clustered within a 4\-percentage\-point band, with telecom recording the lowest completion rate \(86%\) due to its underlying tool and policy complexity \(Table[16](https://arxiv.org/html/2609.36086#A5.T16)\)\.
Table 8:Cell retention metrics by subset after 2 retries\(R=2\)\(R=2\)\.The retention rate measures the proportion of retained final pairs against the total number of generation draws across all attempts\. The cumulative retention rate denotes the proportion of retained final pairs relative to the original set of target cells prior to retries\. Draws per retention measures the average number of generation attempts required to yield one retained pair\.
### C\.2Position Balance Audit
To counter positional bias\[[38](https://arxiv.org/html/2609.36086#bib.bib38)\], the presentation order of trajectories in each pair is randomized during meta\-evaluation \(Table[7](https://arxiv.org/html/2609.36086#A2.T7)\)\. Position balance serves as a diagnostic audit to verify that this randomization prevents positional confounds in filtering and human annotation\. As detailed in Table[9](https://arxiv.org/html/2609.36086#A3.T9), the target trajectory appears in slot two in 511 of the 1,000 final pairs \(51\.1%\), with all experimental subsets remaining within a 45\.5–54\.3% range and displaying no significant departure from uniform parity\.
Table 9:Position balance audit of the target \(better\-steered\) trajectory across data subsets\.Across all groups, slot distribution remains statistically indistinguishable from uniform parity \(binomialp\>0\.10p\>0\.10\), confirming that position bias is not a baseline confounder\.
### C\.3Environment Reward Verification
Table[10](https://arxiv.org/html/2609.36086#A3.T10)evaluates dataset pairs againstτ3\\tau^\{3\}\-bench programmatic task success rewards\[0,1\]\[0,1\], which track database transitions, disclosure verification, and action matching independently of model annotation\. Omitting the LLM judge component renders this audit entirely deterministic\. Under this setup, steering should produce a strong environment reward shift for direct task execution \(task\_resolution\) and a weak or non\-significant shift for auxiliary stylistic criteria \(communication\_clarityandfriendliness\)\.
The empirical results confirm this progression\. Steering fortask\_resolutionproduces the largest reward increase \(\+0\.122\+0\.122,p<0\.001p<0\.001\), whereasfriendlinessexhibits no significant shift \(\+0\.017\+0\.017,p=0\.488p=0\.488\), demonstrating clear orthogonality to task outcome\.communication\_clarityoccupies an intermediate position \(\+0\.082\+0\.082,p<0\.001p<0\.001\), reflecting a real\-world dependency where unclear communication occasionally impedes execution workflows\. Since the reward shift is smaller than that of explicit task resolution steering, communication clarity remains a distinct behavioral dimension\. Overall, the vast majority of retained pairs are not confounded by outcome variation\. Among the 1,000 pairs evaluated on both trajectories, 836 \(84%\) yield identical environment rewards, with the higher\-steered trajectory scoring higher in 11\.8% of pairs and lower in 4\.6%\.
Table 10:Environment reward verification comparing deterministicτ3\\tau^\{3\}\-bench task success reward gaps across evaluation criteria\.The reward gap represents the score of the higher\-steered trajectory minus that of the lower\-steered trajectory\. Among the 1,000 pairs evaluated on both trajectories, 836 pairs \(84%\) exhibit identical environment rewards\.
## Appendix DHuman Panel Reliability and Sampling Robustness
### D\.1The Annotation Task
Figure 3:The interface used by the human panel\.The two trajectories of a pair are presented together with the criterion being judged; the steering instruction and target levels are withheld, and the annotator returns a forced binary choice together with a confidence rating\.#### Annotator Panel
The panel comprises 3 software engineers and 3 computer science researchers with experience developing language model agents or using them for software development\. Two are full\-time employees and four are interns, all within the same company as the authors\. Two are native English speakers and four have professional English fluency\. Annotators review criteria definitions and guidelines prior to evaluation\. The panel is partitioned into two independent groups of three annotators\. Each group evaluates the same 75 trajectory pairs\. Each annotator receives a $50 stipend for completing the annotations\. The annotation tasks involved evaluating benign model outputs and contained no offensive, sensitive, or harmful material\. To minimize psychological fatigue and potential distress, participation was entirely voluntary, annotators were allowed to opt out or take breaks at any time without penalty\. No personally identifiable information \(PII\) was collected\.
#### Task Workflow
For each pairp=\(x−,x\+\)p=\(x^\{\-\},x^\{\+\}\), annotators select the superior trajectory given the target criterioncc\. Choice is forced \(y^a\(p\)∈\{−1,\+1\}\\hat\{y\}\_\{a\}\(p\)\\in\\\{\-1,\+1\\\}\), yielding the majority vote referencemaj\\mathrm\{maj\}used in Section[5\.2](https://arxiv.org/html/2609.36086#S5.SS2)\. Annotators also record confidence on a three\-point scale \(00for guess,11for leaning,22for certain\), analyzed in Appendix[D\.3](https://arxiv.org/html/2609.36086#A4.SS3)\. Annotators receive the same rendered input as the filtering judges and the evaluators under test \(Figure[3](https://arxiv.org/html/2609.36086#A4.F3); guidelines in Appendix[H\.6](https://arxiv.org/html/2609.36086#A8.SS6)\), with trajectory slot positions randomized and target steering instructionssc,t,ℓs\_\{c,t,\\ell\}hidden\.
### D\.2Inter\-Annotator Agreement
Table[11](https://arxiv.org/html/2609.36086#A4.T11)and Table[12](https://arxiv.org/html/2609.36086#A4.T12)quantify inter\-annotator reliability across the human annotation panel and compare human consistency against the synthetic ground truth label⋆\\star\.
As shown in Table[11](https://arxiv.org/html/2609.36086#A4.T11), evaluating Krippendorff’sα\\alphaacross the human\-only panel\(a1,a2,a3\)\(a\_\{1\},a\_\{2\},a\_\{3\}\)yields a baseline reliability of\+0\.350\+0\.350on unfiltered data \(𝒟0\\mathcal\{D\}\_\{0\}\), which increases slightly to\+0\.383\+0\.383on fully filtered pairs \(𝒟2\\mathcal\{D\}\_\{2\}\)\. Replacing a single human rater with the synthetic label⋆\\starand averaging across panel permutations\(ai,aj,⋆\)\(a\_\{i\},a\_\{j\},\\star\)substantially improves inter\-rater reliability, raising Krippendorff’sα\\alphato\+0\.496\+0\.496on𝒟2\\mathcal\{D\}\_\{2\}\(Δ=\+0\.113\\Delta=\+0\.113\)\. This increase indicates that the synthetic ground truth label provides a more consistent central consensus than individual human annotators\.
Table[12](https://arxiv.org/html/2609.36086#A4.T12)evaluates the reliability of the majority vote baselinemaj\\mathrm\{maj\}across the 150 annotated pairs in𝒟0\\mathcal\{D\}\_\{0\}\. The overall panel achieves a pooled Krippendorff’sα\\alphaof\+0\.350\+0\.350\. Exactly 51% of pairs \(77/150\) achieve unanimous agreement \(3–0 vote\), while the remaining 49% represent 2–1 split decisions\. Pairwise inter\-annotator agreement varies widely, with Cohen’sκ\\kapparanging from\+0\.63\+0\.63down to\+0\.12\+0\.12across panel configurations\.
Table 11:Krippendorff’sα\\alphaover a panel of three raters\.The first row is the human panel\(a1,a2,a3\)\(a\_\{1\},a\_\{2\},a\_\{3\}\)\. The second row substitutes one human rater with synthetic label⋆\\starand reports mean±\\pmSD across permutations\(a1,a2,⋆\)\(a\_\{1\},a\_\{2\},\\star\),\(a1,a3,⋆\)\(a\_\{1\},a\_\{3\},\\star\), and\(a2,a3,⋆\)\(a\_\{2\},a\_\{3\},\\star\)\.Δ\\Deltarepresents the difference between hybrid and human\-only panels\.Table 12:Reliability metrics for majority vote referencemaj\\mathrm\{maj\}across the 150 rated pairs in𝒟0\\mathcal\{D\}\_\{0\}\(75 per panel\)\.Unanimous indicates a 3–0 vote\. Lower blocks report pairwise agreement metrics for two\-annotator panel subsets\.measurepooled \(150\)panel 1 \(75\)panel 2 \(75\)Krippendorff’sα\\alpha\+0\.350\\mathbf\{\+0\.350\}\[\+0\.24, \+0\.45\]\+0\.451\+0\.451\[\+0\.29, \+0\.60\]\+0\.240\+0\.240\[\+0\.09, \+0\.38\]unanimous \(3–0\)77/150 = 51%44/75 = 59%33/75 = 44%*Agreement between every two annotators of a panel*agreementCohen’sκ\\kappaboth said “certain”panel 1,\(a1,a2\)\(a\_\{1\},a\_\{2\}\)81\.3% \(61/75\)\+0\.63\\mathbf\{\+0\.63\}85% \(28/33\)panel 1,\(a1,a3\)\(a\_\{1\},a\_\{3\}\)65\.3% \(49/75\)\+0\.32\+0\.3275% \(24/32\)panel 1,\(a2,a3\)\(a\_\{2\},a\_\{3\}\)70\.7% \(53/75\)\+0\.41\+0\.4183% \(30/36\)panel 2,\(a1,a2\)\(a\_\{1\},a\_\{2\}\)66\.7% \(50/75\)\+0\.34\+0\.3481% \(13/16\)panel 2,\(a1,a3\)\(a\_\{1\},a\_\{3\}\)65\.3% \(49/75\)\+0\.27\+0\.2768% \(15/22\)panel 2,\(a2,a3\)\(a\_\{2\},a\_\{3\}\)56\.0% \(42/75\)\+0\.12\\mathbf\{\+0\.12\}73% \(27/37\)mean67\.6%\+0\.35\\mathbf\{\+0\.35\}
### D\.3Confidence and Task Difficulty
Table[13](https://arxiv.org/html/2609.36086#A4.T13)investigates whether filter judges select for human confidence or agreement\. Annotators rate each pair on a three\-point scale \(00for guess,11for leaning,22for certain\), where the per\-pair confidence rating is the mean of three annotator scores\.
Human agreement correlates positively with confidence: Krippendorff’sα\\alphareaches\+0\.545\+0\.545on the 41 pairs where all three annotators indicate “certain”, compared to\+0\.270\+0\.270on the remaining 109 pairs where at least one annotator is unsure\. However, filter judges do not select for higher annotator confidence\. Mean confidence moves by only\+0\.02\+0\.02on a00–22scale between unfiltered𝒟0\\mathcal\{D\}\_\{0\}and dual\-filtered𝒟2\\mathcal\{D\}\_\{2\}\. Furthermore, neitherJ1J\_\{1\}norJ2J\_\{2\}retains a significantly more confident pair subset than the one it removes \(p=1\.000p=1\.000forJ1J\_\{1\},p=0\.172p=0\.172forJ2J\_\{2\}\)\. Given the sample sizes of removed pairs \(n=36n=36forJ1J\_\{1\}andn=10n=10forJ2J\_\{2\}\), the minimum detectable confidence difference is approximately0\.200\.20\.
These results indicate that filter judges discard pairs where annotators split on majority consensus rather than pairs annotators find inherently difficult\. Consequently, filtering raises label agreement without rendering the underlying evaluation task trivial: mean annotator confidence remains1\.511\.51out of2\.002\.00on𝒟2\\mathcal\{D\}\_\{2\}, reflecting non\-trivial judgment calls for human evaluators\.
Table 13:Human annotator confidence and agreement metrics across filter depths𝒟k\\mathcal\{D\}\_\{k\}\.Per\-pair confidence reflects the mean of three annotator ratings \(0=guess0=\\text\{guess\},1=leaning1=\\text\{leaning\},2=certain2=\\text\{certain\}\)\. Statistical significance \(pp\) is computed via permutation tests over 100,000 shuffles\.α\\alphais Krippendorff’sα\\alpha\.
### D\.4Sampling Robustness
Table[14](https://arxiv.org/html/2609.36086#A4.T14)verifies that the 150 human\-annotated pairs accurately represent the full 1,000\-pair dataset\. In the human evaluation study, pairs are sampled uniformly across domain and criterion combinations\. However, the full dataset features non\-uniform domain allocations \(e\.g\., airline comprises 13\.5% of the 1,000 pairs but 23\.1% of the annotated sample\)\.
To evaluate potential sampling bias, we post\-stratify synthetic label accuracy using actual population weights across four axes: domain, criterion, agent model, and level contrast\. Re\-weighting shifts reported synthetic label accuracy by less than 1\.0ptacross all configurations, well within the sampling standard error of±3\.7pt\\pm 3\.7\\text\{pt\}\.
Furthermore, dataset retention rates under filtering in the human sample match full population proportions closely: the 150\-pair sample survivesJ1J\_\{1\}at 76\.0% \(114/150\) and dual filtering at 69\.3% \(104/150\), compared to 76\.4% \(859/1,125\) and 68\.4% \(770/1,125\) across the full dataset\. This confirms that the annotated subset is representative of overall pipeline behavior\.
Table 14:Synthetic label accuracy \(agreement with human majority vote\) across raw sample estimates and population re\-weighted strata \(𝒟k\\mathcal\{D\}\_\{k\}\)\.Post\-stratification shifts headline accuracy by less than 1\.0pt, confirming sample representativeness within a±3\.7pt\\pm 3\.7\\text\{pt\}standard error\.
## Appendix EFiltering Gains and Human Alignment Dynamics
Table[15](https://arxiv.org/html/2609.36086#A5.T15)breaks down filtering gains for agreement between synthetic labels⋆\\starand human majority vote across four dataset axes\. Filtering yields positive accuracy gains across all subsets, with the largest increases occurring where evaluation criteria are most difficult to judge without oversight\. Specifically, communication clarity achieves a\+19\+19ptgain, though only 52% of clarity pairs survive dual\-filter verification\. This indicates that the filtering pipeline discards a larger proportion of borderline trajectories to enforce the target criteria alignment\. Task resolution yields a\+10\+10ptgain with a 72% retention rate\. Friendliness exhibits a modest\+3\+3ptgain with 84% retention, starting from a high unfiltered baseline of 90% and leaving minimal room for further optimization\.
Two structural axes exhibit uniform gains across subsets\. Across agent models, filtering improvements remain stable \(\+10ptto \+13pt\), indicating that the effect stems from dataset construction rather than specific architectures\. Similarly, gains across level contrasts range tightly between \+10ptand \+12pt, proving filtering refines fine contrast pairs as effectively as wide ones without relying on gross trajectory contrasts\.
The telecom domain is the sole exception, yielding a \+1ptgain compared to \+12ptto \+17ptelsewhere\. Table[16](https://arxiv.org/html/2609.36086#A5.T16)suggests a structural mechanism: telecom features a 3,715\-word policy, 30 user\-side tools, a median length of 48 messages, and tool errors in 48% of simulations, leaving filter judges with the least clear signal\. Alternatively, telecom starts with the highest unfiltered accuracy at 79%, leaving minimal room to gain\. Cells hold 24 to 50 pairs throughout, supporting reliable subset ordering\.
Table[17](https://arxiv.org/html/2609.36086#A5.T17)summarizes the second filter \(J2J\_\{2\}\) performance\.J2J\_\{2\}removes 10 of the 114 pairs passed byJ1J\_\{1\}, yielding a marginal\+0\.4\+0\.4ptgain in label purity at the cost of a−7\-7ptdrop in total yield\. Human annotators supportJ1J\_\{1\}filter overrulings in 61% of cases \(22/36\), but only 20% \(2/10\) forJ2J\_\{2\}\. Among the 10 pairs removed byJ2J\_\{2\}, the ground truth label⋆\\staris human\-verified as correct in 8 instances\. At this sample size, the second filter stage yields diminishing returns\. It incurs notable yield loss without providing a statistically meaningful improvement in label purity or alignment quality\. The full dataset retains depth𝒟2\\mathcal\{D\}\_\{2\}, with ablation metrics reported for completeness\.
Table 15:Agreement between synthetic labels⋆\\starand human majority vote across filter depths𝒟k\\mathcal\{D\}\_\{k\}and dataset subsets\.Gain reflects the accuracy change from two filters \(𝒟2\\mathcal\{D\}\_\{2\}\) versus zero filters \(𝒟0\\mathcal\{D\}\_\{0\}\); retention represents the proportion of target pairs retained after two filtering stages\.axissubsetunfiltered \(𝒟0\\mathcal\{D\}\_\{0\}\)1 filter \(𝒟1\\mathcal\{D\}\_\{1\}\)2 filters \(𝒟2\\mathcal\{D\}\_\{2\}\)gainretention*By criterion*friendliness90% \(45/50\)93% \(42/45\)93% \(39/42\)\+3\+3pt84%task resolution68% \(34/50\)77% \(30/39\)78% \(28/36\)\+10\+10pt72%communication clarity62% \(31/50\)80% \(24/30\)81% \(21/26\)\+19\+19pt52%*By domain*airline75% \(27/36\)92% \(23/25\)92% \(22/24\)\+17\+17pt67%retail67% \(26/39\)80% \(24/30\)84% \(21/25\)\+17\+17pt64%banking72% \(26/36\)85% \(22/26\)84% \(21/25\)\+12\+12pt69%telecom79% \(31/39\)82% \(27/33\)80% \(24/30\)\+1\+1pt77%*By agent model*gpt\-oss\-20b76% \(35/46\)90% \(27/30\)89% \(24/27\)\+13\+13pt59%nemotron\-lightning\-3\.571% \(36/51\)81% \(34/42\)83% \(29/35\)\+12\+12pt69%gpt\-oss\-120b74% \(39/53\)83% \(35/42\)83% \(35/42\)\+10\+10pt79%*By level contrast*bad–good \(widest\)80% \(37/46\)90% \(35/39\)92% \(33/36\)\+11\+11pt78%ok–good75% \(38/51\)86% \(31/36\)87% \(26/30\)\+12\+12pt59%bad–ok66% \(35/53\)77% \(30/39\)76% \(29/38\)\+10\+10pt72%Table 16:Domain complexity metrics across evaluation environments\.Tool name overlap between domain pairs ranges from 0\.03 to 0\.11 by Jaccard index; banking represents the only domain featuring active retrieval\.Table 17:Trade\-off analysis for the second\-stage filterJ2J\_\{2\}\.Among the 10 pairs eliminated by stage two, human majority vote supports the original ground truth label in 8 cases\.
## Appendix FThe Full Evaluator Sweep and Contamination Diagnostic
### F\.1The Full Sweep
Table[18](https://arxiv.org/html/2609.36086#A6.T18)reports performance metrics across all 25 evaluators, serving as the unabridged version of Table[3](https://arxiv.org/html/2609.36086#S5.T3)in the main text \(which presents a 12\-evaluator subset selected to span accuracy ranges and model families\)\. All summary statistics, subset means, and correlation analyses in Section[5\.3](https://arxiv.org/html/2609.36086#S5.SS3)\(including Table[4](https://arxiv.org/html/2609.36086#S5.T4)\) are computed over the full 25\-evaluator roster\. Evaluator accuracy is measured asAgr\(f,⋆,𝒟\)\\mathrm\{Agr\}\(f,\\star;\\mathcal\{D\}\)across all 1,000 pairs in𝒟2\\mathcal\{D\}\_\{2\}with ties treated as incorrect, reported as the mean±\\pmSD across three independent runs\. Accuracy spans a 48\.3ptrange, fromclaude\-opus\-5at 85\.1% down toqwen3\-4bat 36\.8%\.
Table[18](https://arxiv.org/html/2609.36086#A6.T18)introduces two additional column blocks omitted from the main text due to space constraints: accuracy breakdown by level contrast and accuracy breakdown by agent model\. Similar to the domain and criterion blocks, each of these blocks partitions the 1,000 pairs such that their weighted average equals the total dataset accuracy\. Row indicators denote specific evaluator conditions:†\\daggeridentifies the two models where parse failures are marked as incorrect, while∗\*marks the three models utilized in dataset construction \(analyzed further in Appendix[F\.2](https://arxiv.org/html/2609.36086#A6.SS2)\)\.
Table 18:The complete sweep: all 25 evaluators across 1,000 data pairs, repeatn=3n=3\.Every statistic reported in Section[5\.3](https://arxiv.org/html/2609.36086#S5.SS3)is computed over these 25 rows\. Table[3](https://arxiv.org/html/2609.36086#S5.T3)in the main paper shares the row numbering\. Leniency is the mean of every score the evaluator emits\. It strongly correlates with accuracy \(Table[4](https://arxiv.org/html/2609.36086#S5.T4)\)\. Level contrast indicates the two steering levels each trajectory pair has:b–oisbad–ok\(n=313n=313\),o–gisok–good\(n=328n=328\),b–gisbad–good\(n=359n=359\)\. Agent model columns are the three trajectory generators of Table[6](https://arxiv.org/html/2609.36086#A2.T6):oss\-20isgpt\-oss\-20b\(n=325n=325\),oss\-120isgpt\-oss\-120b\(n=334n=334\),nemo\.isnemotron\-lightning\-3\.5\(n=341n=341\)\. Note that agent model is a property of the data point being judged rather than of the evaluator\. The final row is the unweighted mean over all 25 evaluators\. Released is the public release month\. Temperature = 0 except where unavailable‡\. Preference is derived from individual trajectory score gaps; ties count as incorrect\. Accuracy and score denote benchmark\-level accuracy and raw score output of the evaluators, respectively\. SD is standard deviation\.∗Contaminated by self\-assessment bias because the model is used during data synthesis \(Appendix[F\.2](https://arxiv.org/html/2609.36086#A6.SS2)\)\.†Parse failures on 192 and 119 calls after repeated retries, counted incorrect\.Boldednumbers indicate best performance within the column\.parametersaccuracy \(%\)averagetieby domain \(%\)by criterion \(%\)by level contrast \(%\)by agent model \(%\)distinct\#evaluatorreleasedtotalactivereas\.mean±\\pmSDleniencyscore SD\(%\)airl\.bank\.retailtelec\.clar\.friend\.taskb–oo–gb–goss\-20oss\-120nemo\.scores1claude\-opus\-52026\-07closedclosedYes85\.1±\\pm0\.930\.3240\.0234\.184828687739685858288858882692claude\-sonnet\-52026\-06closedclosedYes83\.8±\\pm0\.720\.4580\.0405\.882818487739582858186818982343kimi\-k32026\-072\.8T104BYes83\.6±\\pm0\.120\.5080\.0366\.083808487719781877886828781374gpt\-5\.6\-terra‡2026\-07closedclosedYes81\.7±\\pm0\.490\.5600\.0433\.481818282699678867584798680885gpt\-5\.6\-luna‡2026\-07closedclosedYes81\.6±\\pm0\.640\.5630\.0453\.882818183699381877583818579836glm\-5p22026\-06753B∼\\sim30BYes81\.5±\\pm0\.500\.5210\.05110\.184818182789273817885808579317gpt\-oss\-120b∗2025\-08116\.8B5\.1BYes80\.5±\\pm0\.400\.6240\.0509\.980818080788775837484788777458gpt\-5\.4\-nano‡2026\-03closedclosedNo79\.4±\\pm1\.160\.6310\.0527\.881787881748480827283808475619gpt\-5\.4\-mini‡2026\-03closedclosedNo79\.4±\\pm0\.790\.6280\.0536\.1747980817191758572827984768910gpt\-5\.6\-sol‡2026\-07closedclosedYes78\.7±\\pm0\.400\.5430\.0354\.9797781776497748472807782778611claude\-haiku\-4\-52025\-10closedclosedNo78\.7±\\pm0\.100\.5050\.00011\.6817976817685747577847685753212qwen3p8\-max2026\-082\.4T95BYes78\.0±\\pm0\.360\.5210\.03711\.7747877817094688173807682763013nemotron\-3\-ultra2026\-06549B55BYes76\.8±\\pm0\.760\.5580\.04414\.2758070817188717970817681732814llama3\.1\-70b2024\-0770\.6BdenseNo70\.0±\\pm0\.550\.6320\.01622\.8677367726469777361767379581415gemini\-3\.7\-flash2026\-08closedclosedYes69\.5±\\pm0\.450\.6150\.02623\.4717069695795547264736974663416gpt\-oss\-20b∗2025\-0820\.9B3\.6BYes68\.7±\\pm1\.010\.6250\.05621\.7717068677480516967706874643117gemma\-4\-31b2026\-0332\.2BdenseYes68\.1±\\pm0\.150\.6570\.02327\.4636766744792637556737175591418deepseek\-v4\-pro2026\-081\.6T49BYes67\.0±\\pm0\.670\.5050\.04520\.4627069656190496963696370682319mistral\-large22024\-07123BdenseNo64\.8±\\pm0\.400\.7170\.01428\.1596362725763747153707072531220nemotron\-lightning\-3\.5∗2026\-0832B3BYes62\.4±\\pm0\.930\.7050\.06028\.6615860694775646752686669522221gemini\-3\.5\-flash\-lite2026\-07closedclosedNo56\.1±\\pm0\.680\.7230\.05536\.3535847655259586341646462441622llama4\-maverick2025\-04400B17BNo55\.6±\\pm1\.560\.7650\.03935\.9475654615255605851586160451523llama3\.1\-8b†2024\-078\.0BdenseNo43\.1±\\pm0\.470\.6220\.00543\.7434648352843574439474246421224qwen3\-1p7b†2025\-042\.0BdenseYes42\.8±\\pm0\.350\.7600\.07631\.9444442424347384733484345401525qwen3\-4b2025\-084\.4BdenseNo36\.8±\\pm0\.650\.8040\.02956\.93634324430255543274043412617*Mean, all 25 evaluators*70\.168\.769\.969\.072\.262\.079\.567\.973\.263\.473\.670\.574\.965\.1
### F\.2Contamination Diagnostic
Three evaluators in our sweep contribute directly to benchmark construction:nemotron\-lightning\-3\.5serves as filter judgeJ1J\_\{1\},gpt\-oss\-120bserves asJ2J\_\{2\}, and these two models in addition togpt\-oss\-20bgenerate 325–341 of the evaluated trajectory pairs\. While filter judge influence is intrinsic to dataset definition, generation influence can be isolated\. Table[19](https://arxiv.org/html/2609.36086#A6.T19)evaluates potential self\-preference bias by comparing evaluator accuracy on self\-generated trajectories versus external trajectories\.
The diagnostic indicates that self\-preference is not systematic\. Onlygpt\-oss\-120bexhibits higher accuracy on self\-generated outputs \(\+9\.1pt\+9\.1\\text\{pt\}\),gpt\-oss\-20bdisplays negligible shift \(−0\.6pt\-0\.6\\text\{pt\}\), andnemotron\-lightning\-3\.5performs−15\.7pt\-15\.7\\text\{pt\}on its own trajectories\. This result is primarily driven by trajectory difficulty shifts: because agent assignment is deterministic per cell rather than random, excluding self\-generated pairs confounds underlying trajectory difficulty with evaluator bias\. As reported in Table[18](https://arxiv.org/html/2609.36086#A6.T18)under the “by agent model” block, these trajectory difficulty shifts are revealed in the mean accuracy across all 25 evaluators:74\.9%74\.9\\%ongpt\-oss\-120btrajectories,70\.5%70\.5\\%ongpt\-oss\-20btrajectories, and drops to65\.1%65\.1\\%onnemotron\-lightning\-3\.5trajectories\. The observed gaps in Table[19](https://arxiv.org/html/2609.36086#A6.T19)closely track benchmark\-wide trajectory difficulty shifts rather than systematic self\-preference bias\.
Notably, all three generator models belong to the sub\-7B active parameter MoE tier, highlighting a parameter tier constraint when evaluating lightweight model architectures\.
Table 19:Contamination diagnostic comparing evaluator accuracy on self\-generated versus external trajectory pairs\.Values represent means across three independent evaluation runs\.
## Appendix GCost and Caching
The dataset curation pipeline generates 1,000 validated trajectory pairs at a total cost of $23\.63, or $0\.024 per kept pair \(Table[20](https://arxiv.org/html/2609.36086#A7.T20)\)\. Steering instruction generation and filter judges account for 5\.0% of total token volume and 12\.4% of dollar costs\.
Table 20:Computational overhead and curation cost\.All costs are computed or estimated with the model service provider’s token\-based pricing model\.Prompt caching is the primary driver of generation cost reduction\. In an agentic loop, the full conversation history is re\-sent at each turn, making almost every request carry a prefix the server already holds\. Across the generation run, 98% of all processed tokens are prompt tokens\.
Table[21](https://arxiv.org/html/2609.36086#A7.T21)reports the prompt cache hit rate across all 77,980 trajectory calls\. The overall hit rate reaches 90\.5%, ranging from 87\.9% to 95\.4% across individual models\. This caching efficiency reduces total API costs by a factor of two to four depending on model pricing\.
Table 21:Prompt cache hit rates across all 77,980 trajectory generation calls\.
## Appendix HPrompts
This section reproduces every pipeline prompt verbatim in execution order: shared criterion definitions \(Appendix[H\.1](https://arxiv.org/html/2609.36086#A8.SS1)\), steering instruction generation \(Appendix[H\.2](https://arxiv.org/html/2609.36086#A8.SS2)\), agent system prompt wrapping \(Appendix[H\.3](https://arxiv.org/html/2609.36086#A8.SS3)\), filter judging \(Appendix[H\.4](https://arxiv.org/html/2609.36086#A8.SS4)\), model evaluation \(Appendix[H\.5](https://arxiv.org/html/2609.36086#A8.SS5)\), and human annotation instructions \(Appendix[H\.6](https://arxiv.org/html/2609.36086#A8.SS6)\)\. Variable names enclosed in braces \(e\.g\.,\{criterion\_name\}\) represent runtime substitution slots\. Agent and user simulator system prompts are retained directly fromτ3\\tau^\{3\}\-bench without modification; only the steering wrapper \(Appendix[H\.3](https://arxiv.org/html/2609.36086#A8.SS3)\) is appended to agent instructions\.
### H\.1Criterion Definitions
Below is the verbatim text for the three evaluation criteriacc\. The exact same criterion names and descriptions are supplied to the generator, filter judges, model evaluators, and human annotators, ensuring information parity across all pipeline components\. Table[6](https://arxiv.org/html/2609.36086#A2.T6)provides abridged summaries for reference\.
name:friendliness
description:Warmthandconsiderationtowardthecustomer:whethertheagentacknowledgestheirsituationandhowtheyfeelaboutit,deliversunwelcomenewswithcare,andleavesthemfeelingattendedto\.Judgethemanner,notwhethertherequestwasresolved\.
name:communication\_clarity
description:Howeasilythecustomercanfollowtheagent:whetherthemainpointisfindable,whethertechnicalorpolicylanguageisexplained,whethermulti\-partinformationisorganised,andwhetherthecustomerisleftknowingwhatistrueandwhathappensnext\.Judgethepresentation,notthewarmthortheoutcome\.
name:task\_resolution
description:Whetherthecustomer’sactualproblemwassettled:didtheagentestablishwhatwasneeded,taketheactionsthatwouldresolveit,andleavethecustomerwiththeoutcometheycamefor\.Acorrectrefusalcountsasresolution\-\-iftherequestwasnotpermitted,sayingsoplainlyandexplainingwhyresolvesit,whilequietlydoingitanywaydoesnot\.Judgetheoutcome,notthemannerorhowwellitwasexplained\.
### H\.2Steering Instruction Generator
The generator synthesizes three steering instructionssc,t,ℓs\_\{c,t,\\ell\}for a given task cell\(t,c\)\(t,c\)in a single call, corresponding to levelsℓ∈\{bad,ok,good\}\\ell\\in\\\{\\textsf\{bad\},\\textsf\{ok\},\\textsf\{good\}\\\}\. Task specifics \(\{task\_description\}and\{tool\_names\}\) are populated dynamically\. Rule 5 prevents the generator from encoding unstated task outcomes, while Rule 1 ensures level contrasts reflect behavioral differences rather than varying degrees of task detail\.
Youaredesigninganexperimentabouthowwellautomaticevaluatorsjudgethe
qualityofAIagentbehaviour\.
ForONEspecifictaskandONEspecificqualitymetric,writethreesystem\-prompt
instructionsthatsteeranagenttoperformatthreelevelsonthatmetric:BAD,OK
andGOOD\.
\#\#Themetric
Name:\{criterion\_name\}
Whatitmeans:\{criterion\_description\}
\#\#Thetask
Whattheuseristryingtogetdone:
\{task\_description\}
Toolstheagentcanuse:\{tool\_names\}
\#\#Rulesfortheinstructionsyouwrite
1\.BESPECIFICTOTHEKINDOFWORKTHISTASKINVOLVES\-\-NOTTOITSPARTICULARS\.
Genericadvicethatwouldfitanytaskisonefailuremode\.Namingthetask’s
detailsistheother,anditisworse\.Dorefertowhichtoolsmatterhere,what
kindofinformationtheuserneeds,andtheshapethisinteractionwilltake\.Do
NOTrestatethespecificsyouweregiven:nonames,nouseroraccountor
bookingidentifiers,noretellingofthisperson’scircumstances\.Theagent
learnsthosefromtheconversationitself\.Puttingtheminitssystemprompt
handsitinformationoutofband,andmakesyourthreeinstructionsdifferby
howmuchdetailtheycarryratherthanbythemetric\.
Wrong:"YouareassistingEmmaKim\(userIDemma\_kim\_9957\)withcancelling
reservationEHGLP3,andshewasoutoftownrelyingonpriorinsurance\."
Right:"Whenthecustomerexplainswhytheyareasking,acknowledgethe
circumstancestheyraisebeforeyougettotheoutcome\."
2\.MAKETHETHREELEVELSSEPARATE\.Someonereadingthethreeresultingtranscripts
shouldbeabletorankthemonthismetricwithoutbeingtoldwhichiswhich\.If
twoofyourinstructionswouldproducesimilarbehaviour,rewritethem\.
3\.FORBAD,AIMATTHEOPPOSITEOFTHEMETRIC\-\-donotmerelywithholdgood
behaviour\.Workoutwhattheactiveoppositeofthismetricis,onthistask,
andinstructtheagenttopursueit\.Ifthemetricisfriendliness,BADisnot
neutralorterse:itiscold,dismissive,impatient\.Ifthemetricisclarity,
BADisactivelyconfusing,notjustunpolished\.Ifthemetricisrelevanttool
use,BADusestoolsinwaysthatactivelydonotservetherequest\.Namethe
oppositeexplicitlyandtelltheagenttodoit\.
4\.DESCRIBEOKINITSOWNTERMS,notas"somewhatgood"or"slightlybad",andnot
asamilderversionofBAD\.OKiswhatunremarkable,adequate,uncared\-forwork
lookslikeonthismetric\-\-theagentisneithertryingnorfailing\.
5\.YOUDONOTKNOWTHECORRECTOUTCOME,ANDMUSTNOTASSUMEONE\.Youhavenotbeen
toldhowthistaskshouldberesolved,andyoumaynotguess\.Donotwrite
anythingthatpresupposesananswer\-\-not"statetherefusal",not"confirmit
isnotpermitted"\.Wheretheoutcomematterstohowtheagentshouldbehave,
phraseitconditionally:"ifitturnsoutyoucannotdowhattheyareasking,
\.\.\."\.
6\.STEERCONDUCT,NEVERDICTATEWORDS\.Donotwritetheagent’sreplyforit,do
notincludeasentenceforittosay,andneveruse"respondwith","say
exactly"oraquotedmessage\.YouareshapingHOWitworks,notWHATit
concludes\.
7\.WRITEONLYTHEINSTRUCTIONBODY\.Donottelltheagenttokeeptheinstruction
secretorthattheobjectiveoverridesitsdefaults\-\-thatisaddedseparately\.
Twotofivesentenceseach,addressedtotheagentas"you"\.
NothingisofflimitsinhowbadlytheBADlevelmayperformonthismetric\.If
performingtheoppositeofthismetricmeansleavingworkout,beingunhelpful,
beingwrong,orfailingthetaskoutright,saysoplainly\.Theexperimentneedsa
genuinelypoorexample,andaBADlevelthatismerelylukewarmisthemostcommon
waythisgoeswrong\.
ReturnJSONwithexactlythesekeys:
\-"reasoning":whatgood,adequateandpoor\{criterion\_name\}concretelylooklike
forthistask,andhowyouwillmakethethreelevelsseparate
\-"bad":theinstructionbodyfortheBADlevel
\-"ok":theinstructionbodyfortheOKlevel
\-"good":theinstructionbodyfortheGOODlevel
### H\.3Steering Wrapper
The following instruction is padded around each steering instruction when it is injected into the original agent prompt\.
<behaviour\>
Forthisconversationyouhaveoneadditionalobjective\.Ittakesstrictprecedence
overyourdefaultbehaviour,includingwhenitproduceshandlingthatisless
polished,lessthoroughorlesssuccessfulthanyouwouldnormallydeliver\.
Followingitexactlyiswhatmakesyourhandlingofthisconversationcorrect\.
\{instruction\}
Donotmention,quoteoralludetothisobjective,anddonotreuseitswording\.
</behaviour\>
### H\.4Filter Judge
The filter judge promptJk\(p,c\)J\_\{k\}\(p,c\)is identical forJ1J\_\{1\}andJ2J\_\{2\}, differing only in the underlying model deployment\. The\{prompt\}slot contains shared context \(tool definitions, knowledge bases, and domain policies\), while\{response\_a\}and\{response\_b\}contain the rendered trajectory pair in randomized slot order\. Filter judges output a forced choice, and pairs are retained only when both judges match the synthetic ground truth label⋆\\star\.
YouareevaluatingtwoAIassistantresponsesonaspecificcriterion\.
Criterion:\{criterion\_name\}
Definition:\{criterion\_description\}
Userprompt:
\{prompt\}
ResponseA:
\{response\_a\}
ResponseB:
\{response\_b\}
Whichresponsebetterfulfillsthecriterion?
ReturnJSONwithexactlytwokeys:
\-"reasoning":atmost3sentencescomparingthetworesponsesonthiscriterion
\-"answer":"A"or"B"
Statetheanswerlast,andstateitevenifthechoiceisclose\.
### H\.5Evaluator Under Test
The prompt template for evaluators under testf\(x,c\)f\(x;c\)\. Evaluators process a single trajectoryxxin isolation, assigning a score in\[0,1\]\[0,1\]without access to the paired execution\. Pairwise preference predictionsy^f\(p\)\\hat\{y\}\_\{f\}\(p\)are derived post hoc from the sign of the score gap \(Eq\. \([3](https://arxiv.org/html/2609.36086#S3.E3)\)\)\. Evaluators receive the identical shared context\{prompt\}provided to filter judges and human annotators, eliminating information asymmetry across evaluation modes\.
Youareevaluatingonecustomer\-serviceconversationonasinglenamedcriterion\.
CRITERION:\{criterion\_name\}
\{criterion\_description\}
Judgethiscriterionandnothingelse\.Ignoreeverythingthecriteriondoesnot
mention\.Aconversationcanfailatitstaskandstillbebetteronthecriterion\.
JudgetheAGENT’sconductonly\.Thecustomer’sturnswereproducedbyasimulator
andarenotbeingevaluated\.
REFERENCEDOCUMENTS\(theagent’stools,andthepolicyithadtofollow\)\.Consult
themifyouneedtocheckwhethertheagentwaspermittedtodosomething\.Youwill
notneedthemformostconversations\.
\{prompt\}
CONVERSATION
Turnsarenumbered\."AgentcallsX\(\.\.\.\)"isatoolcalltheagentmade,and
"Toolresult\-\>"iswhatcameback\.
\{response\}
Scoretheagenton\{criterion\_name\}from0\.0\(worst\)to1\.0\(best\)\.Usethefull
range:reserve0\.0and1\.0forclearcasesandplaceordinaryonesinbetween\.
ReturnJSONwithexactlytwokeys:
\-"reasoning":atmost3sentences,citingwhatintheconversationdrovethescore
\-"score":anumberbetween0\.0and1\.0
### H\.6Human Annotator Instructions
The instructions of the human annotation interface \(Figure[3](https://arxiv.org/html/2609.36086#A4.F3)\)\. Annotators review these instructions prior to their initial trial and retain permanent access during annotation\. On\-screen confidence levels \(00,5050, and100100\) are mapped to numerical values \(00,11, and22\) as used throughout Appendix[D\.3](https://arxiv.org/html/2609.36086#A4.SS3)\.
Beforeyoustart
Pairwiseannotationofagentconversations\.Readthisonce\.
Judgeonecriterion,andnothingelse
Eachpairisjudgedonasinglenamedcriterion\.
Onlythatcriterioncounts\.Ignoreeverythingelse\.Inanextremescenario,a
conversationcanfailthetaskandstillbethebetteroneonthemetric\.
Theflow\-\-threestepsperpair
1\.Pickaside\.Clicktheconversationthatbetterfitsthecriterion,orpress
theleft/rightarrowkey\.
2\.Rateyourconfidence:1for0\(aguess\),2for50\(leaningoneway\),3for
100\(certain\)\.
3\.PressNext\(Enter\)tosubmitandmoveon\.
Backspacegoesbacktorevise\.
Thereismoretoreadifyouneedit
Undertheheaderarebuttonslabelled"Additionaldocumentstohelpyoudecide":
theagent’stoollist,thepolicyithadtofollow,andtheknowledgebasewhere
oneexists\.Theyareclosedbydefaultandopenonaclick,oneatatime\.Tool
resultsandsearchesinsideaconversationopenthesameway\.
Youdonotneedanyofitformostpairs\.Reachforthepolicywhenthequestion
iswhethertheagentwasallowedtodosomething\.
Howmuchtodeliberate
Useyourbestjudgment,butitmighthurttothinktoomuch\.Inmyexperience
givingeachexampleaboutaminuteshouldbesufficient\.
Ifthetworeallyarehardtoseparate,stillpickaside\-\-thenrateit0\.That
recordsthepairhonestlyasacoin\-flipinsteadofhidingaguessamongyour
realjudgments\.
\[Startannotating\]
\-\-\-\-shownonceatthestartofeachcriterion’sblockofpairs\-\-\-\-
Youarenowjudging
<criterionname\>
<criteriondescription\>
Pressanykeyorclicktobeginthissection
\-\-\-\-thestandingquestionintheheaderofeverypair\-\-\-\-
Whichconversationisbetteronthiscriterion?
## Appendix IAssets, Licenses, and Terms of Use
Table[22](https://arxiv.org/html/2609.36086#A9.T22)lists every existing asset this work builds on, its role in the pipeline, and its license and terms of use\. All open\-weight models are accessed through a single hosted inference provider rather than downloaded, and all proprietary models are accessed through their vendors’ paid APIs, so each model is used under both its own license and the serving provider’s terms of service\.
Table 22:Existing assets used in this work\.Model roles refer to the pipeline stages of Table[7](https://arxiv.org/html/2609.36086#A2.T7)and the evaluator sweep of Section[5\.3](https://arxiv.org/html/2609.36086#S5.SS3)\.Similar Articles
Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
This paper introduces DMAPO, a method for preference optimization that uses multi-evaluator consensus to select high-confidence on-policy responses, achieving strong alignment with significantly less data (only 3.45% acceptance rate) and outperforming baselines like SimPO on several benchmarks.
Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment
This paper presents an LLM-as-a-Judge evaluation framework for agentic AI in drug discovery, validated through human alignment studies with expert annotators. It optimizes the judge to improve alignment with human judgment and provides insights for reusable evaluation in scientific domains.
EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems
This paper introduces EPC, a standardized protocol for measuring evaluator preference coupling in LLM agent systems, including a reference snapshot and versioning convention to address reproducibility and measurement decay.
Preference Estimation via Opponent Modeling in Multi-Agent Negotiation
This paper proposes a novel preference estimation method that integrates natural language information from LLMs into a structured Bayesian opponent modeling framework for multi-agent negotiation. The approach leverages LLMs to extract qualitative cues from utterances and convert them into probabilistic formats, demonstrating improved agreement rates and preference estimation accuracy on multi-party negotiation benchmarks.
PACE: A Proxy for Agentic Capability Evaluation
This paper introduces PACE, a framework that predicts expensive LLM agent benchmark scores using a small subset of cheaper non-agentic evaluation instances, achieving high accuracy at less than 1% of the cost.