CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production

arXiv cs.CL Papers

Summary

This paper introduces Cargo, a framework for evaluating agentic AI systems that addresses reference-instance divergence by grounding factual judgments in live context and gating evaluation on retrieval confidence, along with Cargo-Bench for benchmarking.

arXiv:2609.30471v1 Announce Type: new Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved references as procedural exemplars and grounds factual judgments in the live instance's observed context, (ii) assigns each claim a three-way status (supported, contradicted, unverifiable) and penalizes only contradictions, and (iii) gates evaluation by retrieval confidence, casting production evaluation as selective prediction. We introduce CARGO-Bench, a perturbation-based diagnostic suite with ground truth by construction that separates leniency from discrimination. On CARGO-Bench (246 items, two judge models, 7,872 judgments), the standard reference-based judge penalizes 100% of correct entity-transplanted answers and is uninformative (discrimination index DI ~ 0); supplying the live facts without reframing changes nothing. CARGO eliminates these false penalties (0/50) while retaining near-complete contradiction recall (50/50 and 49/50), raising DI to 0.58 [0.48, 0.68]; a rubric-swap control attributes most of the effect to context-grounded dimension definitions. CARGO also exposes a limitation of its own design: the leniency that protects entity values suppresses detection of procedural corruptions (20% recall). A post-hoc fix does not close the gap, and an LLM-as-annotator study with written guidelines and adjudication shows the same blind spot. We release a preregistered protocol for extending the evaluation to expert agreement, risk-coverage, and cost on production traffic.
Original Article
View Cached Full Text

Cached at: 09/28/26, 09:37 AM

# Context-Aware Retrieval-Gated Evaluationof Agentic AI in Production
Source: [https://arxiv.org/html/2609.30471](https://arxiv.org/html/2609.30471)
## Cargo: Context\-Aware Retrieval\-Gated Evaluation of Agentic AI in Production

###### Abstract

Reference\-based LLM\-as\-a\-judge evaluation assumes that a reference answer is the target for the response under evaluation\. In deployed agentic systems that operate over dynamic entities—support cases, assets, accounts—this assumption fails: the closest available reference typically instantiates the*correct procedure*on a*different entity*, so a judge that compares literally penalizes legitimately different identifiers, dates, and statuses as errors or hallucinations\. We name this failure mode*reference–instance divergence*\(RID\) and show that it is structural rather than incidental\. We proposeCargo, a framework that \(i\) treats retrieved references as procedural exemplars and grounds factual judgments in the live instance’s observed context, \(ii\) assigns each claim a three\-way status—supported, contradicted, or unverifiable—penalizing only contradictions with observed facts, and \(iii\) gates evaluation by retrieval confidence, casting continuous production evaluation as selective prediction with an explicit risk–coverage–cost trade\-off\. To measure the effect without relying solely on costly expert labels, we introduceCargo\-Bench, a perturbation\-based diagnostic suite whose ground truth holds by construction\.Cargo\-Benchseparates*leniency*from*discrimination*: a valid judge must stop penalizing entity transplants while still detecting injected contradictions and procedural corruptions\. OnCargo\-Bench\(246 items, two judge models, 7,872 judgments\), the standard reference\-based judge penalizes 100% of correct entity\-transplanted answers as incorrect and hallucinated—it is uninformative \(DI≈0\\approx 0\)—and supplying the live facts without reframing changes nothing\.Cargoeliminates these false penalties \(0/50 on transplants\) while retaining near\-complete contradiction recall \(50/50 and 49/50\), raising DI to \.58 \[\.48, \.68\]; a rubric\-swap control attributes most of the effect to context\-grounded dimension definitions rather than prompt framing\.Cargoalso exposes a limitation of its own design: the same leniency that protects entity values suppresses detection of procedural corruptions \(20% recall\), a trade\-off DI makes visible; an explicit post\-hoc fix targeting exactly this failure does not close the gap \(Δ\\DeltaDI=−\.007=\-\.007\[−\.038,\.023\-\.038,\.023\]\), and an LLM\-as\-annotator study with written guidelines and adjudication shows the same blind spot \(3/10 procedural corruptions recovered\)\. We release a preregistered protocol for extending the evaluation to expert agreement, risk–coverage, and cost on production traffic\.

## 1Introduction

LLM\-as\-a\-judge has become the default instrument for evaluating open\-ended model outputs at scale\([Zheng et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib25);[Liu et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib16);[Gu et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib8)\)\. In the reference\-based variant, a judge is shown a candidate response together with a gold reference and asked to score the candidate’s correctness, completeness, or faithfulness relative to that reference\. This paradigm rests on an assumption so familiar that it is rarely stated:*the reference is the target*\. For static tasks such as summarization or knowledge QA the assumption is reasonable\.

It fails for a large and growing class of deployed systems\. Agentic assistants in technical support, customer service, IT operations, and finance answer questions*about specific entities*—a particular support case, service tag, order, or account—by retrieving facts from backend systems and applying a domain procedure to them\. Two users may ask the same question about different entities\. The correct answers share a procedure and a structure but differ in almost every surface value\. A curated golden set can realistically cover the space of*questions and procedures*; it cannot cover the space of*entities*\. Consequently, at production time the nearest available reference is almost always an answer to the same question about a different entity\.

We call this condition*reference–instance divergence*\(RID\)\. Under RID, a literal reference\-based judge systematically confuses two orthogonal properties: whether the response follows the correct procedure, and whether its entity\-specific values agree with the reference’s\. The second comparison is meaningless—the values*should*differ—yet it dominates a naive judge’s verdict, depressing correctness scores and inflating hallucination flags\. Section[2](https://arxiv.org/html/2609.30471#S2)formalizes this and shows why the effect is structural rather than a matter of prompt wording\.

We further observe that in production, evaluation itself is a decision under uncertainty\. Not every live interaction has an appropriate reference in the golden set, and a judge applied with an irrelevant reference produces a confidently wrong score\. Evaluating every interaction is also expensive\. Continuous production evaluation is therefore a*selective prediction*problem\([El\-Yaniv and Wiener, 2010](https://arxiv.org/html/2609.30471#bib.bib4);[Geifman and El\-Yaniv, 2017](https://arxiv.org/html/2609.30471#bib.bib7)\): the evaluator should abstain when it lacks a valid reference, and the appropriate objects of study are the risk–coverage curve and the cost curve, not a single accuracy number\.

We presentCargo\(Context\-AwareRetrieval\-Gated evaluatiOn\), which addresses both problems, andCargo\-Bench, a diagnostic suite that makes the improvement measurable with ground truth that holds by construction\. Our contributions are:

1. 1\.Problem\.We identify and formalize reference–instance divergence, decomposing a response into a transferable procedure and instance\-specific parameters, and show that literal reference comparison is not identifiable with respect to procedural correctness under RID \(§[2](https://arxiv.org/html/2609.30471#S2)\)\.
2. 2\.Method\.Cargo\(§[3](https://arxiv.org/html/2609.30471#S3)\) reinterprets the retrieved reference as a*procedural exemplar*, grounds factual judgment in observed live context with three\-way claim status \(sup/con/unv\), and gates evaluation on retrieval confidence with calibrated and margin\-aware variants\. It evaluates the full agentic trace—routing, replanning, response type, latency—not only the final answer\.
3. 3\.Benchmark\.Cargo\-Bench\(§[4](https://arxiv.org/html/2609.30471#S4)\) uses five controlled perturbation families \(entity transplant, contradiction injection, unverifiable augmentation, procedural corruption, retrieval distractors\) to produce labeled evaluation instances by construction\. It yields a*discrimination index*that jointly penalizes false penalties and missed errors, so that a judge cannot score well merely by being lenient\.
4. 4\.Findings\.OnCargo\-Bench\(246 items\), literal reference judging is completely uninformative under RID \(DI≈0\\approx 0on both judge models: every correct transplant is called wrong and hallucinated\), and merely adding the live facts does not help\.Cargoremoves the false penalties without losing contradiction recall \(Δ\\DeltaDI=\+\.58=\+\.58\[\.48, \.68\]\)\. A rubric\-swap control shows the effect is carried mostly by context\-grounded dimension definitions, and per\-family analysis reveals that the method’s leniency over\-generalizes from entity values to procedure \(20% recall on procedural corruptions\); an explicit per\-claim\-type authority fix we test post\-hoc does*not*resolve this \(Δ\\DeltaDI=−\.007=\-\.007\[−\.038,\.023\-\.038,\.023\]\), so we characterize rather than patch it \(§[6](https://arxiv.org/html/2609.30471#S6)–§[7](https://arxiv.org/html/2609.30471#S7)\)\. We further specify a preregistered protocol for expert agreement, retrieval gating, and cost on production traffic \(§[5](https://arxiv.org/html/2609.30471#S5)\)\.

## 2Reference–Instance Divergence

### 2\.1Setting

A live interaction is a tuplex=\(q,a,c,o\)x=\(q,a,c,o\): a user questionqq, the system’s answeraa, an observed*context*ccconsisting of instance\-specific facts available in the trace input \(e\.g\., case number, subject, description, asset identifier\), and a traceooof observations \(planner decision, agent calls, tool outputs, timings\)\. A golden set𝒢=\{\(qig,aig,mig\)\}i=1N\\mathcal\{G\}=\\\{\(q^\{g\}\_\{i\},a^\{g\}\_\{i\},m^\{g\}\_\{i\}\)\\\}\_\{i=1\}^\{N\}holds reference questions, reference answers, and metadatamigm^\{g\}\_\{i\}such as the expected intent or agent\.

### 2\.2Procedure–parameter decomposition

We model an answer as the application of a procedure to instance parameters:

a=π⁡\(θ\)⊕ϵ,a=\\pi\(\\theta\)\\oplus\\epsilon,\(1\)whereπ\\piis a domain procedure \(which checks to perform, which policy to apply, which fields to report, in which structure\),θ\\thetais the vector of instance\-specific values, andϵ\\epsiloncollects everything else \(phrasing, formatting\)\. A referenceag=πg​\(θg\)⊕ϵga^\{g\}=\\pi^\{g\}\(\\theta^\{g\}\)\\oplus\\epsilon^\{g\}is written for its own instanceθg\\theta^\{g\}\. The live answer should satisfyπ=πg\\pi=\\pi^\{g\}\(correct procedure\) andθ\\thetaconsistent with the live instance,*not*withθg\\theta^\{g\}\.

### 2\.3Non\-identifiability of literal comparison

A literal reference\-based judge computes some divergenceD⁡\(a,ag\)D\(a,a^\{g\}\)and maps it to a score\. Under the decomposition,DDmixes two terms,

D⁡\(a,ag\)≈Dπ​\(π,πg\)\+Dθ​\(θ,θg\),D\(a,a^\{g\}\)\\;\\approx\\;D\_\{\\pi\}\(\\pi,\\pi^\{g\}\)\\;\+\\;D\_\{\\theta\}\(\\theta,\\theta^\{g\}\),\(2\)and only the first is informative about quality\. When RID holds \(θ≠θg\\theta\\neq\\theta^\{g\}by design\),DθD\_\{\\theta\}is large for*every*correct answer\. The judge therefore cannot distinguish a correct answer for a different instance from an incorrect one: correctness is not identifiable fromDDalone\. This is why prompt\-level exhortations to “focus on reasoning” are insufficient in practice: the judge is given no signal with which to separate the two terms\. The fix must supply that signal—the live instance’s observed factscc—and must change what the reference is*for*\.

### 2\.4Three\-way claim status

Let𝒦⁡\(a\)\\mathcal\{K\}\(a\)be the set of atomic factual claims inaa\([Min et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib18)\)\. Relative to the observed contextcc, each claimkkhas a status

σ⁡\(k∣c\)∈\{sup,con,unv\},\\sigma\(k\\mid c\)\\in\\\{\\textsc\{sup\},\\;\\textsc\{con\},\\;\\textsc\{unv\}\\\},\(3\)supported if entailed bycc, contradicted if inconsistent with a fact incc, and unverifiable otherwise\. Becauseccis a*partial*snapshot of the request input—the agent legitimately retrieves many more fields from backend systems—the unverifiable class is large and expected\. We define

ContraRate⁡\(a,c\)\\displaystyle\\mathrm\{ContraRate\}\(a,c\)=\|\{k:σ=con\}\|\|𝒦⁡\(a\)\|,\\displaystyle=\\tfrac\{\|\\\{k:\\sigma=\\textsc\{con\}\\\}\|\}\{\|\\mathcal\{K\}\(a\)\|\},\(4\)UnvRate⁡\(a,c\)\\displaystyle\\mathrm\{UnvRate\}\(a,c\)=\|\{k:σ=unv\}\|\|𝒦⁡\(a\)\|\.\\displaystyle=\\tfrac\{\|\\\{k:\\sigma=\\textsc\{unv\}\\\}\|\}\{\|\\mathcal\{K\}\(a\)\|\}\.\(5\)The central normative choice ofCargois that hallucination is scored fromContraRate\\mathrm\{ContraRate\}alone;UnvRate\\mathrm\{UnvRate\}is reported as a coverage diagnostic rather than folded into a penalty\. This choice trades recall on fabricated\-but\-unverifiable claims for precision on the claims that can actually be checked; §[10](https://arxiv.org/html/2609.30471#S10)discusses the trade\-off andCargo\-BenchP3 measures it\.

## 3TheCargoFramework

Cargohas four components \(Figure[1](https://arxiv.org/html/2609.30471#S3.F1)\): a retrieval gate, a procedural\-exemplar reinterpretation of the reference, a context\-grounded structured judge, and trace\-level diagnostics\. Algorithm[1](https://arxiv.org/html/2609.30471#alg1)gives the per\-interaction procedure\.

TraceBatch embeddingRetrieval gateAbstainCargojudge\(exemplar \+ctx\-grounded\)Trace diagnosticsPersisted recordmargin<τ<\\taumargin≥τ\\geq\\tau

Figure 1:Overview ofCargo\. A live question is embedded and matched against the golden set; if the retrieval margin clearsτ\\tau, the retrieved reference is reinterpreted as a procedural exemplar and the answer is judged against it plus the live context; belowτ\\tau,Cargoabstains\. Both paths log to trace diagnostics\.Algorithm 1Cargoper\-interaction evaluation1:interaction

x=\(q,a,c,o\)x=\(q,a,c,o\); golden set

𝒢\\mathcal\{G\}with cached embeddings

EgE^\{g\}; gate

gg; judge

JJ
2:

e←Embed⁡\(q\)e\\leftarrow\\mathrm\{Embed\}\(q\)⊳\\trianglerightbatched across a poll window

3:

si←cos⁡\(e,Eig\)​∀is\_\{i\}\\leftarrow\\cos\(e,E^\{g\}\_\{i\}\)\\;\\forall i;

i⋆←arg⁡maxi⁡sii^\{\\star\}\\leftarrow\\arg\\max\_\{i\}s\_\{i\}
4:

Δ←s\(1\)−s\(2\)\\Delta\\leftarrow s\_\{\(1\)\}\-s\_\{\(2\)\}⊳\\trianglerighttop\-1/top\-2 margin

5:if

g⁡\(si⋆,Δ,mi⋆g\)=0g\(s\_\{i^\{\\star\}\},\\Delta,m^\{g\}\_\{i^\{\\star\}\}\)=0then

6:returnabstain\(si⋆,Δ\)\(s\_\{i^\{\\star\}\},\\Delta\)

7:endif

8:

r←ResponseType⁡\(q,a\)r\\leftarrow\\mathrm\{ResponseType\}\(q,a\);

p←PlannerEval⁡\(o,mi⋆g\)p\\leftarrow\\mathrm\{PlannerEval\}\(o,m^\{g\}\_\{i^\{\\star\}\}\)
9:

ℓ←Latency⁡\(o\)\\ell\\leftarrow\\mathrm\{Latency\}\(o\);

ρ←ReplannerEval⁡\(o\)\\rho\\leftarrow\\mathrm\{ReplannerEval\}\(o\)
10:if

r=actualr=\\textsc\{actual\}∧\\wedgeguards passthen

11:

y←J⁡\(q,a,c,ai⋆g\)y\\leftarrow J\(q,a,c,a^\{g\}\_\{i^\{\\star\}\}\)⊳\\trianglerightexemplar \+ context

12:if

¬Parses⁡\(y\)\\neg\\mathrm\{Parses\}\(y\)thenreturnerror

13:endif

14:endif

15:returnrecord

\(y,r,p,ρ,ℓ,si⋆,Δ,i⋆,c\)\(y,r,p,\\rho,\\ell,s\_\{i^\{\\star\}\},\\Delta,i^\{\\star\},c\)

### 3\.1Retrieval gate

Golden questions are embedded once and cached; live questions are embedded in batches per polling window to bound API calls and rate\-limit pressure\. With cosine similaritys⋆s^\{\\star\}to the nearest golden question and marginΔ=s\(1\)−s\(2\)\\Delta=s\_\{\(1\)\}\-s\_\{\(2\)\}, we study three gates:

Fixed\.g=𝕀\[s⋆≥τ\]g=\\mathbb\{I\}\[s^\{\\star\}\\geq\\tau\]\. Operational defaultτ=0\.80\\tau=0\.80, which must be recalibrated per embedding model\.

Margin\-aware\.g=𝕀\[s⋆≥τ∧Δ≥δ\]g=\\mathbb\{I\}\[s^\{\\star\}\\geq\\tau\\wedge\\Delta\\geq\\delta\]\. Rejects ambiguous matches whose top two candidates are nearly tied, even when both exceedτ\\tau\.

Calibrated\.FitP^​\(reference appropriate∣s⋆,Δ,intent\)\\hat\{P\}\(\\text\{reference appropriate\}\\mid s^\{\\star\},\\Delta,\\text\{intent\}\)by isotonic or logistic regression on a held\-out calibration split labeled for reference appropriateness \(§[5\.2](https://arxiv.org/html/2609.30471#S5.SS2)\), and gate onP^≥γ\\hat\{P\}\\geq\\gamma\. This makes the operating point interpretable as a target reference\-validity rate and admits per\-intent thresholds\.

The gate decides whether the*evaluator*has a usable reference; it says nothing about whether the production answer is good\. Abstentions are retained as coverage gaps for human review and golden\-set expansion \(§[7](https://arxiv.org/html/2609.30471#S7)\)\.

### 3\.2Reference as procedural exemplar

For accepted interactions the judge receivesqq,aa, the observed contextcc, and the retrieved referenceai⋆ga^\{g\}\_\{i^\{\\star\}\}with an explicit reinterpretation: the reference was written for a*different*instance; it specifies the correct procedure, policy, required steps, and structure; its entity\-specific values are not targets\. The judge is asked to infer the reference’s implicit instance from its text and to evaluate whether the live answer applies the same procedure*correctly adapted*to the live instance\. Appendix[A](https://arxiv.org/html/2609.30471#A1)gives the full template\.

### 3\.3Context\-grounded structured judge

The judge scores four dimensions on\[0,1\]\[0,1\], each defined against the live context rather than the reference:correctness\(correct procedure; no claim contradictscc\),completeness\(covers the steps the exemplar’s procedure requires, adapted tocc\),helpfulness\(resolvesqqclearly\), andhallucination fidelity\(no claim contradictscc; unverifiable claims are neutral\)\. Output is schema\-constrained JSON with a per\-dimension rationale; unparseable or empty outputs are rejected rather than persisted, so aggregates are not biased by silent failures\. Dimension definitions are configuration, not code, and are held fixed across all conditions in our experiments\.

### 3\.4Trace\-level diagnostics

Final\-answer quality alone cannot localize a failure in an agentic workflow\.Cargoadditionally records response type \(actual answer, clarification request, guardrail\), planner verdict against the expected agent inmi⋆gm^\{g\}\_\{i^\{\\star\}\}, replanner behavior, per\-observation and end\-to\-end latency, language and length guards, and all retrieval metadata \(s⋆s^\{\\star\},Δ\\Delta,i⋆i^\{\\star\}\) and the contextccused\. These fields support the attribution analysis in §[7](https://arxiv.org/html/2609.30471#S7)\.

### 3\.5Continuous operation

A polling or backfill cycle fetches traces newer than a watermark, deduplicates, extracts\(q,a,c,o\)\(q,a,c,o\)from heterogeneous trace envelopes, runs Algorithm[1](https://arxiv.org/html/2609.30471#alg1), persists successful records, and advances the watermark, reporting seen/matched/evaluated/abstained/error counts\. Evaluation cost per window ofMMtraces is

C⁡\(τ\)=M​Ce\+M⋅Cov⁡\(τ\)⋅Cj,C\(\\tau\)=M\\,C\_\{e\}\+M\\cdot\\mathrm\{Cov\}\(\\tau\)\\cdot C\_\{j\},\(6\)with embedding costCe≪CjC\_\{e\}\\ll C\_\{j\}\(judge cost\), so cost is governed almost entirely by coverage\.

## 4Cargo\-Bench: Ground Truth by Construction

Expert annotation of production traces is necessary but slow, expensive, and—critically for our claim—cannot by itself distinguish a judge that is*better*from one that is merely*more lenient*\.Cargo\-Benchaddresses this with controlled perturbations of seed items whose labels hold by construction\.

#### Seeds\.

Each seed is a golden item\(qg,ag,θg\)\(q^\{g\},a^\{g\},\\theta^\{g\}\)with its entity parametersθg\\theta^\{g\}explicitly annotated \(fields and values\)\. Seeds span all intents in the deployment\.

#### Perturbation families\.

From each seed we generate instances with known target verdicts:

P1 Transplant\(correct\) Sampleθℓ\\theta^\{\\ell\}, rewriteaℓ=πg​\(θℓ\)a^\{\\ell\}=\\pi^\{g\}\(\\theta^\{\\ell\}\);c⊂θℓc\\subset\\theta^\{\\ell\}\. Target: high correctness, no hallucination\. Measures the*false\-penalty rate*under RID\.

P2 Contradiction\(hallucinated\) From a P1 instance, alter one value inaℓa^\{\\ell\}to conflict with a fact*present*incc\. Target: flagged\. Measures*contradiction recall*\.

P3 Unverifiable\(neutral\) From a P1 instance, addkkplausible fields absent fromcc\. Target: no penalty\. A labeled “fabricated” sub\-split measures the recall cost of theunvpolicy\.

P4 Corruption\(incorrect\) From a P1 instance, delete a required step, swap the policy, or invert a conditional\. Target: low correctness/completeness—the check leniency alone cannot pass\.

P5 DistractorParaphrases ofqgq^\{g\}with a different intent/procedure\. Target: gate abstains or assigns lowP^\\hat\{P\}\. Measures the gate’s reference\-validity discrimination\.

Perturbations are generated by templated rewriting with an LLM and verified by a second model and by rule\-based consistency checks \(every occurrence of a swapped value is updated; injected contradictions target a field that is incc\)\. An author spot\-check of 40 instances \(Appendix[B](https://arxiv.org/html/2609.30471#A2)\) additionally reads each rendered triple for semantic validity beyond the automated checks\. Appendix[B](https://arxiv.org/html/2609.30471#A2)lists templates\.

#### Discrimination index\.

LetFP\\mathrm\{FP\}be the fraction of P1∪\\cupP3 instances a judge penalizes \(correctness<0\.5<0\.5or hallucination flagged\), andTP\\mathrm\{TP\}the fraction of P2∪\\cupP4 instances it penalizes\. We report

DI=TP−FP∈\[−1,1\],\\mathrm\{DI\}=\\mathrm\{TP\}\-\\mathrm\{FP\}\\in\[\-1,1\],\(7\)alongside the full confusion matrix\. A lenient judge lowers FP but also TP; a valid judge must raise DI\.Cargo\-Benchtherefore converts our central claim into a single falsifiable quantity\.

## 5Experimental Setup

### 5\.1Deployment and data

Experiments use a production multi\-agent assistant for enterprise technical support in which a planner routes each request to one of several specialized agents \(policy, case, work\-order, task, asset\)\. Traces are collected via an observability layer that records inputs, outputs, and per\-observation timings\. The deployment’s curated golden set comprises 30 question–answer–intent items; 23 are entity\-free policy answers \(for which RID does not arise\) and the remainder concern specific cases, work orders, and assets\. We draw a production sample ofM=400M=400traces stratified by intent, retrieval similarity decile, and response type, oversampling low\-margin retrievals and planner failures with a 150\-item minimum quota for the entity\-bearing \(RID\-eligible\) stratum, since H1’s effect concentrates there\.M=400M=400is powered to detectΔ​ρ≥0\.15\\Delta\\rho\\geq 0\.15betweenDirectandCargo\(Fisherzz, dependent correlations, power0\.80\.8,α=0\.05\\alpha=0\.05\) and matches typical budgets in published LLM\-judge agreement studies \(200–500 items\)\. All data are de\-identified before annotation \(§[Ethics Statement](https://arxiv.org/html/2609.30471#Sx1)\)\.

#### Cargo\-Benchinstantiation\.

Cargo\-Benchis generated from ten seeds in two batches: an initial six \(three derived from the entity\-bearing golden items—case status, work\-order status, asset age/warranty—lightly templated so entity parameters are explicit, and three synthetic seeds covering case, warranty, and work\-order procedures\), plus four added seeds spanning the same real intents with different question styles \(case priority/assignment, asset entitlement, work\-order parts, and a task\-status seed\) to increase statistical power and intent diversity\. With five draws per seed and family \(generation seeds 2027 and 3141\), this yields 296 instances that pass all consistency checks: 50 P1, 50 P2, 96 P3 \(50 true\-but\-unobserved, 46 fabricated\), 50 P4 \(26 step deletions, 10 negations, 14 policy swaps\), and 50 P5\. Observed context contains 1\.12 facts on average \(\|c\|∈\{1,2,3\}\|c\|\\in\\\{1,2,3\\\}\); answers average 24 words\. P2 contradictions target case numbers \(31\), service tags \(15\), and subjects \(4\)\.

### 5\.2Expert annotation

Three domain experts annotate each sampled trace with: \(a\) whether the retrieved reference is an appropriate procedural exemplar \(binary, used to fit the calibrated gate and for retrieval evaluation\); \(b\) correctness, completeness, helpfulness, and hallucination on a 5\-point ordinal scale defined against the live context, using the same definitions as the judge; \(c\) statusσ∈\{sup,con,unv\}\\sigma\\in\\\{\\textsc\{sup\},\\textsc\{con\},\\textsc\{unv\}\\\}for each disputed claim; \(d\) acceptability of the planner’s routing\. Two annotators label every item; a third adjudicates disagreements\. We report Krippendorff’sα\\alphaper dimension and release guidelines \(Appendix[C](https://arxiv.org/html/2609.30471#A3)\)\. A held\-out 20% of annotated items is used only for gate calibration and threshold selection; all headline numbers are computed on the remainder\.

### 5\.3Conditions

All judge conditions share the judge model, decoding parameters \(temperature 0\.4 unless varied\), dimension definitions, and JSON schema; they differ only in the information and framing given to the judge\.

NoRefQuestion, answer, rubric only\.

DirectStandard reference\-based judging: compareaatoaga^\{g\}\.

Direct\+CtxDirectwithccappended, but no exemplar reinterpretation orunvrule\. Isolates “more information” from “different framing\.”

CargoFull method: exemplar framing, context grounding,unvrule\.

Cargo−\-unvCargowith binary claim status \(absent⇒\\Rightarrowfalse\)\. Tests theunvrule\.

Cargo−\-Ex\.Cargowithout the “different instance” framing \(Cargo−\-Exemplar in tables\)\.

Cargo−\-CtxCargowithccremoved\.

OracleCargowith the human\-selected appropriate reference; upper\-bounds gains attributable to retrieval\.

RandRefCargowith a random same\-intent reference; lower\-bounds the value of retrieval\.

Retrieval baselines: BM25\([Robertson and Zaragoza, 2009](https://arxiv.org/html/2609.30471#bib.bib21)\), dense top\-1 with the deployed embedding model, a sentence\-transformer alternative\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.30471#bib.bib20)\), intent\-filtered dense, and hybrid\. Gating baselines: no gate; random sampling at matched coverage; fixed, margin\-aware, and calibrated gates \(§[3\.1](https://arxiv.org/html/2609.30471#S3.SS1)\)\. One additional condition,Cargo\+\+Auth\(an explicit per\-claim\-type authority rule\), was added*post\-hoc*after observingCargo’s P4 gap and is reported as exploratory, not preregistered \(§[7](https://arxiv.org/html/2609.30471#S7)\)\.

### 5\.4Judge models and bias controls

We run every condition with the same two judge models used forCargo\-Bench\(§[6\.1](https://arxiv.org/html/2609.30471#S6.SS1)\)—gpt\-oss\-120b \(three seeds\) and gpt\-oss\-20b \(one seed\), open\-weight and distinct from the production assistant’s generator, limiting self\-preference\([Panickssery et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib19)\)—so bench and production\-sample results are directly comparable; a third judge family would further strengthen H5 but is left to future work rather than added post hoc\. We report per\-seed means and within\-item dispersion\. BecauseCargojudges single responses, position bias\([Wang et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib23)\)does not arise; we control for verbosity bias via agreement stratified by answer\-length quartile and a length\-controlled regression\([Dubois et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib3)\)\.

### 5\.5Metrics

Agreement:Spearmanρ\\rhoand Pearsonrragainst adjudicated expert scores per dimension; quadratic\-weighted Cohen’sκ\\kappaafter binning; paired bootstrap 95% CIs \(10k resamples\); Holm correction across dimensions\.Hallucination:with expertconlabels as positives, precision/recall/F1 and two targeted false\-positive rates \(cross\-instance differences;unvclaims\)\.Cargo\-Bench:per\-family penalty rates, DI, and DI stratified by intent and by\|c\|\|c\|\.Retrieval:top\-1 accuracy and MRR against expert reference\-appropriateness labels; fraction of traces with no appropriate reference in𝒢\\mathcal\{G\}\.Selective evaluation:for gateggand disagreement lossℓ\\ell,

Cov⁡\(g\)\\displaystyle\\mathrm\{Cov\}\(g\)=1M​∑jg⁡\(xj\),\\displaystyle=\\tfrac\{1\}\{M\}\\textstyle\\sum\_\{j\}g\(x\_\{j\}\),\(8\)Risk⁡\(g\)\\displaystyle\\mathrm\{Risk\}\(g\)=∑jg⁡\(xj\)​ℓ​\(y^j,yj\)∑jg⁡\(xj\),\\displaystyle=\\tfrac\{\\sum\_\{j\}g\(x\_\{j\}\)\\,\\ell\(\\hat\{y\}\_\{j\},y\_\{j\}\)\}\{\\sum\_\{j\}g\(x\_\{j\}\)\},\(9\)risk–coverage curves, AURC\([Geifman and El\-Yaniv, 2017](https://arxiv.org/html/2609.30471#bib.bib7)\), and stream\-level failure recall \(the fraction of*all*expert\-identified failures the gated system evaluates and flags, since abstaining on hard failures trivially lowers risk\)\.Cost:judge calls avoided, tokens, wall\-clock, API error rate, and Eq\. \([6](https://arxiv.org/html/2609.30471#S3.E6)\) vs\. full evaluation and matched\-coverage random sampling\.

### 5\.6Preregistered hypotheses

H1CargoexceedsDirectin expert agreement on correctness and hallucination, with the largest gains on items whereθ≠θg\\theta\\neq\\theta^\{g\}\(i\.e\., RID present\)\.

H2Cargoreduces the cross\-instance false\-positive hallucination rate relative toDirectwithout reducing contradiction recall \(DI increases; TP does not decrease\)\.

H3Direct\+Ctxrecovers only part of the gain: exemplar framing and theunvrule contribute beyond adding information\.

H4The calibrated gate dominates the fixed gate on AURC and stream\-level failure recall at matched coverage; random sampling at matched coverage has strictly worse risk\.

H5Gains hold across judge models; the ordering of conditions is stable\.

Hypotheses, conditions, metrics, and the calibration/test split for the production\-sample study \(H1, H2, H4, H5\) were fixed before annotation began and are committed inexperiments/cargo/PREREGISTRATION\.md\(record ID01b8122f1cf4, a content hash reproducible viamake\_prereg\_id\.py\); this record explicitly excludesCargo\-Bench\(already complete at commit time, using construction\-time ground truth\) and the post\-hocCargo\+\+Authfollow\-up \(§[6\.1](https://arxiv.org/html/2609.30471#S6.SS1)\)\.

## 6Results

We report theCargo\-Benchstudy \(§[6\.1](https://arxiv.org/html/2609.30471#S6.SS1); 246 items, 8,172 successful judgments across the main study and rubric\-swap control\) and an LLM\-as\-annotator study \(Appendix[D](https://arxiv.org/html/2609.30471#A4)\)\. The production\-sample analyses \(H1’s expert agreement, H2 against expert labels, H4’s gate calibration\) are the subject of the preregistered protocol in §[5\.2](https://arxiv.org/html/2609.30471#S5.SS2)–§[5\.6](https://arxiv.org/html/2609.30471#S5.SS6); their metrics and reporting formats are fixed in §[5\.5](https://arxiv.org/html/2609.30471#S5.SS5)and Appendix[E](https://arxiv.org/html/2609.30471#A5)\.

### 6\.1Cargo\-Bench\(H2, H3, H5\)

Table[1](https://arxiv.org/html/2609.30471#S6.T1)reports the completedCargo\-Benchstudy: 246 P1–P4 instances \(10 seeds: 6 pilot\+\+4 added for statistical power, §[5\.1](https://arxiv.org/html/2609.30471#S5.SS1)\)×\\times8 conditions, judged by gpt\-oss\-120b \(3 seeds, temperature 0\.4\) and gpt\-oss\-20b \(1 seed\); 7,872 successful judgments \(49 transient rate\-limit/connection failures were retried to completion\), all parseable\.

Should*not*pen\.Should pen\.ConditionP1P3P2P4DI95% CI*Judge: gpt\-oss\-120b \(mean of 3 seeds\)*NoRef\.44\.64\.96\.66\.24\[\.13, \.35\]Direct1\.001\.001\.001\.00\.00\[\.00, \.00\]D\+Ctx1\.001\.001\.001\.00\.00\[\.00, \.00\]Cargo\+Auth\.00\.041\.00\.20\.57\[\.47, \.67\]Cargo\.00\.031\.00\.20\.58\[\.48, \.68\]*Judge: gpt\-oss\-20b \(1 seed\)*NoRef\.10\.33\.92\.36\.39\[\.27, \.50\]Direct1\.001\.001\.001\.00\.00\[\.00, \.00\]D\+Ctx1\.001\.001\.00\.98−\-\.01\[−\-\.03, \.00\]Cargo\+Auth\.00\.01\.96\.22\.58\[\.48, \.68\]Cargo\.00\.02\.98\.16\.56\[\.46, \.65\]Table 1:Cargo\-Bench\(246 items, 10 seeds\), headline conditions \(D\+Ctx=Direct\+Ctx\); the three component\-ablation conditions \(Cargo−\-Ctx/Exemplar/unv\) are in Table[2](https://arxiv.org/html/2609.30471#A5.T2)\(Appendix[E](https://arxiv.org/html/2609.30471#A5)\)\. Penalty rate per family \(lower is better for P1/P3, higher for P2/P4\); DI=TP−FP=\\mathrm\{TP\}\-\\mathrm\{FP\}with item\-level bootstrap CIs \(10k\)\.Directpenalizes*every*instance, including all correct entity transplants, and is therefore uninformative \(DI≈0\\approx 0\); adding the live context without reframing \(Direct\+Ctx\) changes nothing\.Cargoremoves the false penalties and retains contradiction recall, but detects only 20% of procedural corruptions \(P4\); an explicit authority\-split fix \(Cargo\+\+Auth, post\-hoc\) does not improve on this\.#### RID is total under literal judging\.

Directassigns a penalizing verdict to 100% of instances in every family, on both judges and all seeds—all 50 correct entity transplants \(P1\) and all 96 P3 instances\. Inspection of rationales confirms the mechanism predicted in §[2](https://arxiv.org/html/2609.30471#S2): a representative P1 rationale reads “*The answer gives a different case number and status than the golden answer, so it is factually incorrect*” and, for hallucination, “*The answer fabricates an incorrect case number and status, which is a clear hallucination*\.” Of 150DirectP1 judgments \(120b, 50 items×\\times3 seeds\), all 150 cite an entity\-value mismatch and 149/150 additionally call the live value “fabricated”/“hallucinated”; mean correctness is 0\.006 and mean hallucination fidelity 0\.019\. UnderCargothe same instances receive 0\.968 and 1\.000\. A judge with DI≈0\\approx 0carries no information about answer quality; this is the regime in which the deployment’s live monitoring previously operated\.

#### Information is not the fix; framing is \(H3\)\.

Direct\+Ctxreceives exactly the live factsCargoreceives, yet is indistinguishable fromDirecton 120b and only marginally different on 20b \(1/50 P4 items flipped\): the judge does not spontaneously use context to discount reference mismatches; it must be told what the reference is*for*\.

#### Cargoremoves false penalties without losing contradiction recall \(H2\)\.

Cargopenalizes 0/50 P1 and 3/96 P3 instances \(120b\) while flagging 50/50 \(120b\) and 49/50 \(20b\) injected contradictions\. PairedΔ\\DeltaDI\(Cargo−\-Direct\) is\+\.58\+\.58\[\.48,\.68\] \(120b\),\+\.56\+\.56\[\.46,\.65\] \(20b\); vs\.NoRef,\+\.34\+\.34\[\.21,\.46\] and\+\.17\+\.17\[\.06,\.27\]\. Seed dispersion is low \(mean s\.d\. \.03, 120b\)\. One deviation from preregistered H2: TP was not expected to decrease relative toDirect, butDirect’s TP=1\.00=1\.00is an artifact of penalizing everything, andCargo’s lower TP \(\.60\) traces to P4, not P2—H2 holds for contradiction recall and fails for procedural recall; we report both rather than re\-scope the hypothesis\.

#### Procedural corruptions are largely missed\.

Cargopenalizes 10/50 \(120b\) and 8/50 \(20b\) P4 instances\. By corruption type \(120b\): 10/26 delete\-steps are caught, 0/10 negations, 0/14 policy\-swaps\.NoRef, with no reference at all, catches more \(\.66 overall\) at the cost of FP=\.57=\.57\. EveryCargovariant shows the same P4 profile, so the effect is not attributable to the exemplar preamble or theunvrule individually \(below\)\. Reading the 121/150CargoP4 judgments \(120b\) that were*not*penalized, all 121 justify the verdict by absence of contradiction with the live facts, and 97 additionally assert the answer “follows the correct reasoning pattern/approach\.” The judge has generalized “absent from context⇒\\Rightarrowunverifiable” from entity*values*, where it is intended, to*procedural claims*, where the reference—not the context—is the authority \(meanCargocorrectness on P4 is 0\.78\)\. This is the leniency–sensitivity trade\-off DI was designed to expose\.

#### Testing the obvious fix \(post\-hoc, not preregistered\)\.

An explicit per\-claim\-type authority rule,Cargo\+\+Auth, does*not*improve onCargo: pairedΔ\\DeltaDI=−\.007=\-\.007\[−\.038,\.023\-\.038,\.023\] \(120b\),\+\.027\+\.027\[−\.033,\.088\-\.033,\.088\] \(20b\), both straddling zero\. Stating the split explicitly does not make the judge apply it; we leave a structurally different check to future work\.

#### Ablations and rubric swap \(Tables[2](https://arxiv.org/html/2609.30471#A5.T2)–[3](https://arxiv.org/html/2609.30471#A5.T3), Appendix[E](https://arxiv.org/html/2609.30471#A5)\)\.

Removing live context costs DI \(Δ=\+\.06\\Delta=\+\.06\[\.02,\.11\], via lower P2 recall\); removing the exemplar preamble orunv\(vs\. binary\) has no measurable effect on either judge\. Swapping rubrics \(120b, seed 0, 150\-item pilot\) localizes*where*the effect lives:DirectwithCargo’s dimension definitions alone already reaches DI=\.49=\.49, so the definitions carry most of the false\-penalty reduction, while the exemplar preamble adds leniency that removes residual false penalties but costs P4 recall \(\.33→\\to\.10\)—complementary components pulling the operating point in opposite directions on procedure\.

### 6\.2Selective evaluation \(H4\) and robustness \(H5\)

H4 is evaluated on per\-item expert loss labels from the production sample under the protocol of §[5](https://arxiv.org/html/2609.30471#S5); Appendix[E](https://arxiv.org/html/2609.30471#A5)fixes the risk–coverage and gate\-comparison reporting format\.

For H5, condition ordering is nearly identical for gpt\-oss\-120b and gpt\-oss\-20b \(Table[1](https://arxiv.org/html/2609.30471#S6.T1)\):Direct/Direct\+Ctxat DI≈0\\approx 0,NoRefat \.24/\.39,Cargovariants between \.52 and \.59 with the full method andCargo\+\+Authtied\-best\.Cargo’s DI differs by only \.02 across judges; the weaker judge is less harsh onNoRef\(FP \.25 vs\. \.57\)\. Seed dispersion \(120b\) is largest forNoRef\(\.10\), i\.e\. a reference—even mis\-framed—stabilizes the judge\.

## 7Analysis

#### Where doesDirectfail, and where doesCargofail?

Symmetrically, for one reason each \(§[6\.1](https://arxiv.org/html/2609.30471#S6.SS1)\):Directtreats every entity\-value difference as an error \(DI≈0\\approx 0\);Cargotreats every procedure\-level difference as unverifiable too \(20% P4 recall\)—category errors, not graded biases like verbosity or position\([Wang et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib23);[Dubois et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib3)\)\. Nor is either leniency in general:Cargogets P2=1\.00=1\.00, FP=\.02=\.02\(vs\.NoRef’s \.57\), and the authority split does not fix it\. Appendix[F](https://arxiv.org/html/2609.30471#A6)gives qualitative examples\.

#### Is the P4 blind spot specific to single\-pass judging?

In an LLM\-as\-annotator study \(Appendix[D](https://arxiv.org/html/2609.30471#A4)\), two independently prompted annotators plus an adjudicator applied our written guidelines—which include a worked example of exactly this failure—to 40Cargo\-Benchitems\. They recovered every P1–P3 verdict but only 3/10 procedural corruptions\. Better instructions, a second pass, and adjudication do not close the gap, which suggests it requires a structurally different check \(e\.g\., explicit per\-step verification against the reference\)\.

## 8Related Work

We relate to four areas \(expanded in Appendix[H](https://arxiv.org/html/2609.30471#A8)\)\. LLM\-as\-a\-judge reliability work studies biases orthogonal to RID—position, self\-preference, verbosity, task\-dependent alignment\([Zheng et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib25);[Liu et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib16);[Fu et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib6);[Kocmi and Federmann, 2023](https://arxiv.org/html/2609.30471#bib.bib13);[Kim et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib12);[Gu et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib8);[Wang et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib23);[Panickssery et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib19);[Dubois et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib3);[Bavaresco et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib1);[Thakur et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib22)\)—while treating the reference as ground truth; we study a reference correct for a*different*instance\. Factuality/claim\-decomposition work\([Min et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib18);[Honovich et al\., 2022](https://arxiv.org/html/2609.30471#bib.bib9);[Manakul et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib17)\)motivates ourunvclass\. RAG/agent evaluation\([Es et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib5);[Liu et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib15);[Yao et al\., 2025](https://arxiv.org/html/2609.30471#bib.bib24)\)differs in that retrieval selects the*evaluator’s*reference, so its error corrupts measurement, not generation\. Selective prediction\([Chow, 1970](https://arxiv.org/html/2609.30471#bib.bib2);[El\-Yaniv and Wiener, 2010](https://arxiv.org/html/2609.30471#bib.bib4);[Geifman and El\-Yaniv, 2017](https://arxiv.org/html/2609.30471#bib.bib7);[Kamath et al\., 2020](https://arxiv.org/html/2609.30471#bib.bib11);[Kadavath et al\., 2022](https://arxiv.org/html/2609.30471#bib.bib10);[Kuhn et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib14)\)is applied here to the*evaluator*\.

## 9Conclusion

RID makes literal LLM\-as\-a\-judge uninformative \(DI=0=0\);Cargorestores discrimination \(DI \.00→\\to\.58\) with near\-complete contradiction recall, not without limitations \(§[10](https://arxiv.org/html/2609.30471#S10)\)\.

## 10Limitations

Scope of evaluation\.Our results use construction\-time ground truth \(Cargo\-Bench\) and LLM annotators \(Appendix[D](https://arxiv.org/html/2609.30471#A4)\); expert agreement and gate calibration on production traffic \(H1, H4\) are left to the preregistered protocol of §[5](https://arxiv.org/html/2609.30471#S5)\.Procedural leniency\.Cargodetects only 20% of procedural corruptions onCargo\-Benchbecause the judge extends the “unverifiable” status from entity values to procedural claims; an explicit per\-claim\-type authority instruction \(§[7](https://arxiv.org/html/2609.30471#S7)\) did not fix this \(Δ\\DeltaDI=−\.007=\-\.007\[−\.038,\.023\-\.038,\.023\]\), suggesting the gap is not merely underspecification\. Until it is closed,Cargoshould be read as a reliable detector of*contradictions with observed facts*and a reliable non\-detector of*spurious entity mismatches*, not a complete correctness judge; deployments should pair it with a procedure\-focused check orNoRef\-style scoring on procedural dimensions\.Unverifiable is not verified\.Treatingunvclaims as neutral raises precision on checkable claims but forgoes recall on fabrications that happen to be unverifiable from the trace input:Cargopenalized 3/46 fabricated\-unverifiable P3 instances, and the binary\-status ablation did not change this materially\. Closing the gap requires authoritative backend evidence, which is a deployment change we do not evaluate\.Bench scale and synthetic perturbations\.Cargo\-Benchresults rest on ten seeds \(six derived from production golden items and synthetic case/warranty/work\-order procedures, four added for power\) and templated rewrites; effects this large are unlikely to reverse with more seeds, but the P4 rate is sensitive to how corruptions are authored—it moved from 10% to 20% between our 150\- and 246\-item releases—and all CIs should be read atn=246n\{=\}246\(orn=150n\{=\}150for the rubric\-swap control\)\.Rubric confound, partially controlled\.CargoandDirectdiffer in both preamble and dimension definitions; the rubric\-swap control \(Table[3](https://arxiv.org/html/2609.30471#A5.T3)\) separates them on one judge, one seed, and the smaller item set only\.Similarity is not validity\.The gate scores question similarity, not procedural equivalence; two similar questions may require different policies\. The calibrated gate mitigates but does not eliminate this, and theOraclegap bounds the residual\.Single deployment domain\.Our production data come from one enterprise technical\-support assistant;Cargo\-Benchtransfers by recipe but our numbers do not\.Judge dependence\.Results may shift with judge model, prompt wording, and decoding; we test two judge families \(three seeds on one\) but cannot exhaust this space\.Label uncertainty\.Completeness and helpfulness admit expert disagreement; the protocol reportsα\\alphaand adjudicates, but the human ceiling bounds attainable agreement\.Benchmark realism\.Cargo\-Benchperturbations are generated and verified automatically with human spot\-checks; they are designed to be diagnostic, not distributionally representative, and we do not claim otherwise\.Production selection effects\.Gated subsets are not representative of all traffic; we therefore report stream\-level failure recall alongside selective risk\.

## Ethics Statement

Production traces contain customer and asset information\. Under our protocol, all production data used for annotation and analysis are de\-identified by removing or hashing identifiers, replacing free\-text descriptions with paraphrases, and excluding any trace with residual personal data after review; annotators are employees bound by confidentiality agreements and compensated as part of their regular duties\. We releaseCargo\-Benchgeneration code and templates and the annotation guidelines; we do not release raw production traces\. Released examples are synthetic or fully de\-identified\. Automated evaluators can be misused to over\-trust system outputs; we positionCargoas a monitoring aid that surfaces abstentions and contradictions for human review, not as a replacement for it\.

## References

- Bavaresco et al\. \(2024\)Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F\. T\. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K\. Surikuchi, Ece Takmaz, and Alberto Testoni\. 2024\.LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks\.*arXiv preprint arXiv:2406\.18403*\.
- Chow \(1970\)C\. K\. Chow\. 1970\.On optimum recognition error and reject tradeoff\.*IEEE Transactions on Information Theory*, 16\(1\):41–46\.
- Dubois et al\. \(2024\)Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B\. Hashimoto\. 2024\.Length\-controlled AlpacaEval: A simple way to debias automatic evaluators\.*arXiv preprint arXiv:2404\.04475*\.
- El\-Yaniv and Wiener \(2010\)Ran El\-Yaniv and Yair Wiener\. 2010\.On the foundations of noise\-free selective classification\.*Journal of Machine Learning Research*, 11:1605–1641\.
- Es et al\. \(2024\)Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert\. 2024\.RAGAs: Automated evaluation of retrieval augmented generation\.In*Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations*, pages 150–158, St\. Julians, Malta\. Association for Computational Linguistics\.
- Fu et al\. \(2023\)Jinlan Fu, See\-Kiong Ng, Zhengbao Jiang, and Pengfei Liu\. 2023\.GPTScore: Evaluate as you desire\.*arXiv preprint arXiv:2302\.04166*\.
- Geifman and El\-Yaniv \(2017\)Yonatan Geifman and Ran El\-Yaniv\. 2017\.Selective classification for deep neural networks\.In*Advances in Neural Information Processing Systems 30*\.
- Gu et al\. \(2024\)Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Yuanzhuo Wang, Wen Gao, Lionel Ni, and Jian Guo\. 2024\.A survey on LLM\-as\-a\-judge\.*arXiv preprint arXiv:2411\.15594*\.
- Honovich et al\. \(2022\)Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias\. 2022\.TRUE: Re\-evaluating factual consistency evaluation\.In*Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 3905–3920\. Association for Computational Linguistics\.
- Kadavath et al\. \(2022\)Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield\-Dodds, Nova DasSarma, Eli Tran\-Johnson, Scott Johnston, Sheer El\-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, and 17 others\. 2022\.Language models \(mostly\) know what they know\.*arXiv preprint arXiv:2207\.05221*\.
- Kamath et al\. \(2020\)Amita Kamath, Robin Jia, and Percy Liang\. 2020\.Selective question answering under domain shift\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics*, pages 5684–5696\. Association for Computational Linguistics\.
- Kim et al\. \(2024\)Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo\. 2024\.Prometheus: Inducing fine\-grained evaluation capability in language models\.In*The Twelfth International Conference on Learning Representations*\.
- Kocmi and Federmann \(2023\)Tom Kocmi and Christian Federmann\. 2023\.Large language models are state\-of\-the\-art evaluators of translation quality\.In*Proceedings of the 24th Annual Conference of the European Association for Machine Translation*, pages 193–203\.
- Kuhn et al\. \(2023\)Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar\. 2023\.Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation\.In*The Eleventh International Conference on Learning Representations*\.
- Liu et al\. \(2024\)Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, and 3 others\. 2024\.AgentBench: Evaluating LLMs as agents\.In*The Twelfth International Conference on Learning Representations*\.
- Liu et al\. \(2023\)Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu\. 2023\.G\-Eval: NLG evaluation using GPT\-4 with better human alignment\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pages 2511–2522, Singapore\. Association for Computational Linguistics\.
- Manakul et al\. \(2023\)Potsawee Manakul, Adian Liusie, and Mark J\. F\. Gales\. 2023\.SelfCheckGPT: Zero\-resource black\-box hallucination detection for generative large language models\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, Singapore\. Association for Computational Linguistics\.
- Min et al\. \(2023\)Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen\-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi\. 2023\.FActScore: Fine\-grained atomic evaluation of factual precision in long form text generation\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pages 12076–12100, Singapore\. Association for Computational Linguistics\.
- Panickssery et al\. \(2024\)Arjun Panickssery, Samuel R\. Bowman, and Shi Feng\. 2024\.LLM evaluators recognize and favor their own generations\.In*Advances in Neural Information Processing Systems 37*\.
- Reimers and Gurevych \(2019\)Nils Reimers and Iryna Gurevych\. 2019\.Sentence\-BERT: Sentence embeddings using siamese BERT\-networks\.In*Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing*, pages 3982–3992\. Association for Computational Linguistics\.
- Robertson and Zaragoza \(2009\)Stephen Robertson and Hugo Zaragoza\. 2009\.The probabilistic relevance framework: BM25 and beyond\.*Foundations and Trends in Information Retrieval*, 3\(4\):333–389\.
- Thakur et al\. \(2024\)Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes\. 2024\.Judging the judges: Evaluating alignment and vulnerabilities in LLMs\-as\-judges\.*arXiv preprint arXiv:2406\.12624*\.
- Wang et al\. \(2024\)Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui\. 2024\.Large language models are not fair evaluators\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 9440–9450, Bangkok, Thailand\. Association for Computational Linguistics\.
- Yao et al\. \(2025\)Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan\. 2025\.τ\\tau\-bench: A benchmark for tool\-agent\-user interaction in real\-world domains\.In*The Thirteenth International Conference on Learning Representations*\.
- Zheng et al\. \(2023\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P\. Xing, Hao Zhang, Joseph E\. Gonzalez, and Ion Stoica\. 2023\.Judging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.In*Advances in Neural Information Processing Systems 36 \(Datasets and Benchmarks Track\)*\.

## Appendix AJudge Prompts

All conditions share one template: a preamble, an optional live\-context block, the question, the answer, an optional reference, the four dimension definitions \(fromconfig/llm\_metrics\_config\.yamlforDirect/Direct\+Ctx,config/live\_llm\_metrics\_config\.yamlfor allCargovariants\), and a fixed JSON schema\. Conditions differ only in which of the blocks below are included \(experiments/cargo/conditions\.py\)\.

#### Directpreamble\.

*“You are evaluating the quality of an answer against an expected reference answer\.”*No live\-context block; the reference is labeled*Golden Answer*\.

#### Direct\+Ctxaddition\.

Adds the live\-context block with only:*“The LIVE CONTEXT below lists facts observed for the ACTUAL instance the live answer addresses\. Factual correctness and hallucination are judged against these facts, not against the reference’s values\.”*Preamble and reference label are unchanged fromDirect\.

#### Cargoadditions \(preamble replaced, context block extended\)\.

Preamble:*“You are evaluating the quality of an answer produced for a LIVE instance\.”*followed by the exemplar\-framing block:

> IMPORTANT — the REFERENCE ANSWER below was written for a DIFFERENT instance \(a different case / asset / record\) than the one the live answer addresses\. Treat the reference as the correct REASONING PATTERN, POLICY, REQUIRED STEPS, and STRUCTURE — NOT as literal expected text\. \[…\] NEVER mark the live answer wrong, incomplete, or hallucinated because its identifiers, dates, status, subject, description, service tag, or other instance\-specific values DIFFER from the reference’s\. They SHOULD differ\.

The context block additionally gets theunvrule:

> The LIVE CONTEXT is a PARTIAL snapshot, not the full record\. \[…\] A value that is ABSENT from the live context is UNVERIFIABLE — it is NOT evidence of hallucination or error\. Penalize ONLY values that DIRECTLY CONTRADICT a fact present in the live context, wrong reasoning, or missing required steps\.

The reference is labeled*Reference Answer \(reasoning pattern, written for a different instance\)*\.Cargo−\-Exemplar drops the first block;Cargo−\-Ctx drops the context block \(and both rules\) entirely;Cargo−\-unvreplaces theunvrule with:*“Treat the LIVE CONTEXT as the complete set of known facts\. Any value in the live answer that is not supported by the live context should be treated as unsupported \[…\]\.”*

#### Cargo\+\+Authaddition \(post\-hoc, §[6\.1](https://arxiv.org/html/2609.30471#S6.SS1)\)\.

Appends, after theunvrule:

> Different claim types have different authorities\. \[…\] ENTITY VALUES \[…\]: check ONLY against the LIVE CONTEXT\. \[…\] PROCEDURAL CONTENT \(which steps are taken, what policy or eligibility conclusion is reached, \[…\]\): check against the REFERENCE’s procedure, adapted to the live instance\. \[…\] If the answer’s procedure, policy conclusion, or recommended action CONTRADICTS what the reference’s approach implies for this live instance, that IS a correctness error, even though it does not contradict anything in the live context\. Do NOT excuse a procedural contradiction just because the live context is silent on it\.

## Appendix BCargo\-BenchGeneration

#### Generator\.

Seeds specifyquestionandanswertemplates over named parameters, the subset of parameters that appear in a live trace input \(observed\_fields\), optional proceduralsteps, and optionalpolicy\_alternatives\. Transplants \(P1\) resample every parameter with a type\-aware generator \(digits for case/work\-order numbers, alphanumerics for service tags, dates, closed status/priority vocabularies\) and render the templates;ccis a non\-empty random subset of the observed fields\. P2 rewrites exactly one observed value in the P1 answer to a distinct value of the same type\. P3 appendsk=2k\{=\}2fields absent fromcc, drawn either from the transplant’s hidden parameters \(true\-but\-unobserved split\) or sampled fresh \(fabricated split\)\. P4 applies one of*delete step*,*negate*\(regex over modal/eligibility phrases\), or*swap policy*\(seed\-provided alternative\)\. P5 renders a different\-intent seed’s question with the current seed’s entities\.

#### Consistency checks\.

Every instance must satisfy: no golden parameter value survives in a transplanted answer unless it was re\-sampled as a new value for some field; every observed fact equals the transplant value; P2’s stated value is present and the true value absent, and the contradicted field is observed; P3’s added fields are unobserved; P4 changed the text\. Instances failing any check are dropped \(\-\-drop\-invalid\); both generation batches \(original 180 and expansion 116, §[5\.1](https://arxiv.org/html/2609.30471#S5.SS1)\) have zero violations\.

#### Discrimination\-index sanity check\.

Before any model runs, we verified with two synthetic judges that DI behaves as intended: a uniformly harsh judge \(penalizes everything\) and a uniformly lenient judge \(penalizes nothing\) both obtain DI=0=0on the full 296\-instance release \(TP==FP=1=1and TP==FP=0=0respectively\), so neither strategy can score well\.

#### Author spot\-check\.

We drew a stratified random sample of 8 instances per family \(seed 42; 40 of 296 total\) from the full release and read each question/answer/context/target triple against family\-specific criteria: P1 — entity values self\-consistent, no golden value leaks, context matches the transplant; P2 — exactly one clearly identifiable contradiction with an observed fact; P3 — added fields genuinely absent from context and plausible; P4 — the corruption is an unambiguous, recognizable procedural error distinct from an entity\-value change; P5 — the rendered question reflects a genuinely different intent/procedure\. All 40/40 instances had correct target labels \(no mislabeled correctness/hallucination/should\-penalize field\)\. The check did surface one template\-authoring defect: theg\_workorder2seed’s answer template referenced its work\-order\-number placeholder twice consecutively \(“work orderxhasxpart …”\), producing a redundant but not factually incorrect repetition in 21 of its 30 instances \(21/296==7\.1% of the full release, one seed of ten\)\. This does not change any target label—the repeated token is not an additional claim—so we did not regenerate or re\-judge the affected instances; we fixed the template for future releases and added a regression test \(tests/test\_cargo\_experiments\.py\) that renders every seed file and flags any answer with an immediately\-repeated identifier\-like token\. This was an author check, not an independent or blind review, and should not be read as equivalent to the crowdsourced validation we intend for future releases\.

## Appendix CAnnotation Guidelines

The guidelines below \(draft v0\.1; full text and worked examples inexperiments/cargo/guidelines\.py\) are the single specification for the expert annotators of the production\-sample protocol \(§[5\.2](https://arxiv.org/html/2609.30471#S5.SS2)\) and for the LLM annotators of Appendix[D](https://arxiv.org/html/2609.30471#A4)\.

#### Step 1 — reference appropriateness \(binary\)\.

Annotators ignore every instance\-specific value and ask whether the reference answers the same*kind*of question with the same*kind*of procedure\. Different case numbers, subjects, statuses, or dates are expected and are not evidence of inappropriateness; a different required procedure is\. If inappropriate, scoring stops for that item \(this is the ground truth against which the retrieval gate’s reference validity is measured, §[5\.5](https://arxiv.org/html/2609.30471#S5.SS5)\)\.

#### Step 2 — quality dimensions \(5\-point ordinal, only if Step 1 passes\)\.

Correctness, completeness, and helpfulness are defined identically in spirit to §[3\.3](https://arxiv.org/html/2609.30471#S3.SS3)but scored 1–5 for human tractability\. Hallucination fidelity \(5 = best\) is scored from contradictions with the*observed*context only; a value simply absent from the observed context is unverifiable, not evidence against the score\.

#### Worked examples\.

The guidelines include three examples drawn fromCargo\-Bench\(Appendix[B](https://arxiv.org/html/2609.30471#A2)\) to calibrate annotators, including one deliberately adversarial case: an asset\-warranty answer that inverts the reference’s eligibility conclusion \(“*eligible*”→\\to“*not eligible*”\) without contradicting any observed fact\. Annotators are instructed that this must score correctness≤2\\leq 2because it contradicts the reference’s*procedure/policy*, even though hallucination fidelity \(checked only against observed facts\) may remain high — i\.e\., the guidelines explicitly warn annotators about the failure mode §[7](https://arxiv.org/html/2609.30471#S7)finds inCargoitself\.

#### Step 3 — claim status\.

For any claim that lowers the hallucination score, annotators record the contradicting claim and the observed fact it contradicts\.

#### Adjudication\.

Two annotators label every item independently\. An item is sent to a third adjudicator if reference\-appropriateness labels differ, any ordinal score differs by≥2\\geq 2, or contradiction flags differ\. The adjudicator sees both \(blinded\) annotations and produces the final label, which need not match either\.

## Appendix DLLM\-as\-Annotator Study

#### Headline\.

Given our written guidelines \(Appendix[C](https://arxiv.org/html/2609.30471#A3)\), including a worked example of exactly the P4 failure, two LLM annotators plus an adjudicator recover every P1–P3 verdict on 40Cargo\-Benchitems but only3/10 procedural corruptions\. The P4 blind spot ofCargo\(§[6\.1](https://arxiv.org/html/2609.30471#S6.SS1)\) therefore persists under guideline\-driven, two\-pass, adjudicated annotation, and is not solely an artifact of single\-pass judging\.

#### Design\.

Two gpt\-oss\-120b annotators apply the guidelines to each item without seeing the automated judges’ outputs:annotator\_A\(temperature 0\.2\) follows them methodically;annotator\_B\(temperature 0\.7\) independently and skeptically re\-derives each judgment\. Disagreements, by the adjudication rule of Appendix[C](https://arxiv.org/html/2609.30471#A3), go to an adjudicator \(temperature 0\.1\)\. Because all three share a model family with our judges, self\-preference effects\([Panickssery et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib19)\)cannot be excluded\.

#### Sample\.

40Cargo\-Benchitems \(10 per family, P1–P4\), stratified, seed 0\. All 40 parsed; 1/40 required adjudication\.

#### Recovery against construction\-time ground truth\.

Reference\-appropriateness and hallucination\-flag accuracy were both 1\.00 \(40/40\)\. Correctness accuracy \(binarized at≥3\\geq 3\) was 1\.00 on P1–P3 but\.30 on P4, despite \(a\) explicit guidelines, \(b\) an adversarial worked example of this failure, \(c\) two independent passes, and \(d\) adjudication\. This points to a structurally different check \(e\.g\., per\-step verification against the reference\) rather than better instructions\.

#### Inter\-annotator agreement\.

Quadratic weightedκ\\kappa: correctness 1\.00, completeness \.99, helpfulness \.95, hallucination 1\.00; reference\-appropriateness agreement 1\.00\. Near\-ceiling agreement is expected from two prompted variants of one model and is not an estimate of human–human agreement\.

#### Comparison with the automated judges\.

Adjudicated correctness correlates withCargo’s at Spearmanρ=\.74\\rho=\.74and withDirect’s atρ=\.16\\rho=\.16\(n=40n=40\), consistent with the mainCargo\-Benchresult\. On the 10 P4 items the annotators’ flag rate \(\.30\) exceedsCargo’s \(\.10\), catching twonegatecorruptions \(seedg\_asset\) that the judge missed; both miss the samedelete\_stepandswap\_policycases\.

#### Reproduction\.

python \-m experiments\.cargo\.simulate\_annotatorsgenerates the annotations;python \-m experiments\.cargo\.report\_annotationproduces the numbers above\.

## Appendix EFull Results

### E\.1Cargo\-Bench: component ablations \(complete\)

Table[2](https://arxiv.org/html/2609.30471#A5.T2)gives the three component\-ablation conditions omitted from Table[1](https://arxiv.org/html/2609.30471#S6.T1)for space; discussion is in §[6\.1](https://arxiv.org/html/2609.30471#S6.SS1)\.

Should*not*pen\.Should pen\.ConditionP1P3P2P4DI95% CI*Judge: gpt\-oss\-120b \(mean of 3 seeds\)*Cargo−\-Ctx\.00\.06\.92\.20\.52\[\.42, \.62\]Cargo−\-Ex\.\.00\.021\.00\.20\.59\[\.49, \.68\]Cargo−\-unv\.00\.021\.00\.20\.59\[\.49, \.68\]*Judge: gpt\-oss\-20b \(1 seed\)*Cargo−\-Ctx\.00\.07\.90\.24\.52\[\.42, \.63\]Cargo−\-Ex\.\.00\.02\.92\.14\.52\[\.42, \.62\]Cargo−\-unv\.00\.01\.94\.14\.53\[\.43, \.63\]Table 2:Cargo\-Benchcomponent ablations \(246 items, 10 seeds\), continuing Table[1](https://arxiv.org/html/2609.30471#S6.T1)\.Cargo−\-Ex\. =Cargo−\-Exemplar \(drops the “different instance” framing\)\.
### E\.2Rubric\-swap control \(complete\)

Table[3](https://arxiv.org/html/2609.30471#A5.T3)isolates the dimension\-definition rubric from the exemplar preamble; discussion is in §[6\.1](https://arxiv.org/html/2609.30471#S6.SS1)\.

Table 3:Rubric\-swap control \(gpt\-oss\-120b, seed 0, 150 items\)\. The context\-grounded dimension definitions account for most of the false\-penalty reduction; theCargopreamble adds leniency that removes residual false penalties but also suppresses detection of procedural corruptions\.
### E\.3Production\-sample protocol: reporting formats

The preregistered protocol \(§[5](https://arxiv.org/html/2609.30471#S5)\) fixes in advance how the production\-sample analyses will be reported, so that results cannot be selectively presented:

- •H1/H3 agreement\.Spearmanρ\\rho\(95% bootstrap CI\) with adjudicated expert scores on correctness, completeness, helpfulness, and hallucination forNoRef,Direct,Direct\+Ctx, allCargoablations,Cargo,RandRef, andOracle, with human–human Krippendorff’sα\\alphaas a ceiling\.
- •H1 RID split\.Correctnessρ\\rhoforDirectandCargoand their pairedΔ\\Delta, separately for traces where the reference’s instance differs from the live instance \(case/asset/work\-order\) and where it does not \(entity\-free policy\); H1 predicts the gain concentrates in the former\.
- •H2 hallucination\.Precision, recall, and F1 against expertconlabels, plus false\-positive rates on legitimate cross\-instance differences and onunvclaims, forDirect,Direct\+Ctx,Cargo−\-unv, andCargo\.
- •H4 gating\.Risk–coverage and failure\-recall–coverage curves for fixed, margin\-aware, and calibrated gates and matched\-coverage random sampling; AURC, stream\-level failure recall, reference validity \(fraction of evaluated items whose reference experts judged appropriate\), and relative judge cost \(Eq\. \([6](https://arxiv.org/html/2609.30471#S3.E6)\)\) at a matched coverage level\.

## Appendix FQualitative Examples

All examples are syntheticCargo\-Benchinstances \(no production data\); item IDs refer toexperiments/cargo/runs/\.

#### \(i\) Legitimate cross\-instance difference \(g\_case\-P1\-0\)\.

Q:“What is the status of case 171056262?”Live context:\{caseNumber: 171056262\}\.A:“Case 171056262 is currently Pending Customer\.”Reference:“Case 215474133 is currently Closed\.”Directscores correctness 0\.0:*“The answer gives a different case number and status than the golden answer, which states case 215474133 is Closed\.”*Cargoscores correctness 1\.0 on the identical item:*“The answer provides the case status without contradicting any fact in the live context\.”*The status value \(“Pending Customer”\) is not observed in the live context either way;Cargocorrectly treats it as unverifiable rather than comparing it to the reference’s unrelated status\.

#### \(ii\) Injected contradiction, both flag \(g\_case\-P2\-0\)\.

Live context:\{caseNumber: 491541446\}\.A:“Case 536762755 is currently Escalated\.”Direct:*“The answer fabricates a case number and status that are not supported by the reference, constituting a clear hallucination”*\(correct verdict, wrong reason—it would say this even if the case number matched the live instance but not the reference\)\.Cargo:*“The answer states a case number \(536762755\) that directly contradicts the known case number \(491541446\) in the live context”*\(correct verdict, correct reason: a genuine contradiction with the observed fact\)\.

#### \(iii\) Fabricated\-unverifiable field, ablation has limited effect \(s002\-P3f\-0\)\.

Q:“Is service tag 9Y0S3VQ still under warranty?”Live context:\{serviceTag: 9Y0S3VQ\}\.A:“Service tag 9Y0S3VQ is under warranty until 2025\-07\-19\. The system is eligible for on\-site repair\. You should open a hardware case to dispatch a part\. Warranty ends: 2026\-08\-20\. Created: 2026\-05\-28\.” The fabricated fields \(warrantyEnd,createdDate\) happen to collide with a date already stated earlier in the answer, producing a genuine internal inconsistency \(two different warranty end dates\) independent of the live context\.Cargo\(correctness 0\.60\):*“\[…\] provides two conflicting end dates, which shows inconsistent reasoning even though no live facts are contradicted”*;Cargo−\-unv\(correctness 0\.22\):*“\[…\] gives two conflicting warranty end dates and cannot be verified against the live facts, so the information is unreliable\.”*Both conditions penalize this item, in degree rather than in kind—consistent with the aggregate finding \(§[6\.1](https://arxiv.org/html/2609.30471#S6.SS1)\) that theunvablation does not materially change fabricated\-field detection\. We note this collision as aCargo\-Benchgenerator artifact \(a fabricated field can coincidentally duplicate a value already present in the templated answer\) rather than a clean test of the intended unverifiable\-vs\-unsupported contrast; futureCargo\-Benchreleases should exclude field names already present in the seed’s answer template from the P3 fabrication pool\.

## Appendix GImplementation Details

#### Cargo\-Benchjudging runs\.

Judge models:gpt\-oss\-120bandgpt\-oss\-20bvia the deployment’s OpenAI\-compatible gateway; temperature 0\.4, top\-pp1\.0, no max\-token cap \(reasoning models return empty content when capped\), prompt passed as the system message, JSON parsed leniently \(code fences and surrounding prose tolerated\)\. Sampling seeds\{0,1,2\}\\\{0,1,2\\\}passed as the APIseedparameter\. 7,872 successful judgments across the 246\-item main study \(8 conditions×\\times2 judges, 3 seeds on 120b\) plus 300 for the rubric\-swap control \(Table[3](https://arxiv.org/html/2609.30471#A5.T3)\); six concurrent workers; 49 transient failures \(rate\-limit/connection\) on the original 150\-item pilot were re\-run to completion at two workers, and the 96\-item expansion batch \(3,072 judgments\) completed with zero errors; zero unparseable responses throughout\. Penalization rule for DI: correctness<0\.5<0\.5or hallucination fidelity<0\.5<0\.5on the per\-item mean over seeds\. Bootstrap CIs: 10,000 item\-level resamples; paired comparisons resample the same items for both conditions\.

#### Gate \(deployed retrieval mechanics\)\.

The production similarity matcher embeds the golden set once and caches the vectors; live questions are embedded in batches per polling window \(default batch limit 32\) to stay under the embedding endpoint’s rate limit, with exponential backoff \(base delay 1\.0s, max 30\.0s, up to 5 retries\) on transient failures\. The embedding model defaults tonomic\-embed\-text\-v1via the deployment’s GenAI gateway, with an optional localall\-MiniLM\-L6\-v2fallback; cosine similarity is computed densely against the full golden matrix \(src/monitoring/similarity\_matcher\.py,genai\_embeddings\.py\)\.τ=0\.80\\tau=0\.80is the current operational default and is a placeholder pending calibration on the production sample, not a value derived from data\.

## Appendix HExtended Related Work

This expands the compressed discussion in §[8](https://arxiv.org/html/2609.30471#S8)\.

#### LLM\-as\-a\-judge\.

Strong LLMs approximate human preferences on open\-ended tasks\([Zheng et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib25);[Liu et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib16);[Fu et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib6);[Kocmi and Federmann, 2023](https://arxiv.org/html/2609.30471#bib.bib13)\), and open evaluators can match them given references and rubrics\([Kim et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib12)\)\. Surveys catalog reliability threats\([Gu et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib8)\): position bias\([Wang et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib23)\), self\-preference\([Panickssery et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib19)\), verbosity\([Dubois et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib3)\), task\-dependent alignment\([Bavaresco et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib1);[Thakur et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib22)\)\. This literature treats the reference as ground truth and studies how faithfully a judge compares a candidate to it; the biases identified \(a judge favoring the first\-shown or longer or self\-generated answer\) are properties of the*comparison mechanism*\. RID is different in kind: the comparison target itself is inapplicable at the value level, so no amount of debiasing the comparison mechanism addresses it\. We see the two lines of work as complementary—a debiased judge that still compares literal reference values is still vulnerable to RID\.

#### Factuality and claim decomposition\.

FActScore decomposes generations into atomic claims and scores support against a knowledge source\([Min et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib18)\); TRUE benchmarks factual\-consistency metrics\([Honovich et al\., 2022](https://arxiv.org/html/2609.30471#bib.bib9)\); SelfCheckGPT detects hallucination via sampling consistency without external evidence\([Manakul et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib17)\)\. We adopt atomic claim decomposition as the unit of analysis for the three\-way statusσ∈\{sup,con,unv\}\\sigma\\in\\\{\\textsc\{sup\},\\textsc\{con\},\\textsc\{unv\}\\\}\(§[2](https://arxiv.org/html/2609.30471#S2)\), but our evidence source is a partial, request\-scoped observed context rather than an external corpus or repeated sampling, and our contribution is the explicitunvclass: a claim can be correctly deemed non\-evidence for hallucination precisely because the evidence needed to check it was never provided, which differs from FActScore’s assumption that a sufficiently large knowledge source can adjudicate every claim\.

#### RAG and agent evaluation\.

RAGAS decomposes retrieval and generation quality into separate metrics for a retrieval\-augmented pipeline\([Es et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib5)\); agent benchmarks such as AgentBench andτ\\tau\-bench measure end\-to\-end task completion in controlled, resettable environments\([Liu et al\., 2024](https://arxiv.org/html/2609.30471#bib.bib15);[Yao et al\., 2025](https://arxiv.org/html/2609.30471#bib.bib24)\)\. Both evaluate a system whose retrieval component feeds the*generator*\. InCargo, retrieval is not part of the production system at all: it selects the*evaluator’s*reference from a golden set, so a retrieval error corrupts the measurement of an otherwise\-correct answer rather than the answer itself\. This motivates treating retrieval confidence as a gating signal for evaluation \(§[3\.1](https://arxiv.org/html/2609.30471#S3.SS1)\) rather than as a component to optimize for downstream task success, and it is why our retrieval baselines \(BM25, dense, hybrid\) are evaluated against expert reference\-appropriateness labels rather than against final\-task accuracy\.

#### Selective prediction\.

Abstention with a reject option dates to[Chow \(1970\)](https://arxiv.org/html/2609.30471#bib.bib2); risk–coverage analysis and the area under the risk–coverage curve formalize the accuracy–coverage trade\-off for classifiers\([El\-Yaniv and Wiener, 2010](https://arxiv.org/html/2609.30471#bib.bib4);[Geifman and El\-Yaniv, 2017](https://arxiv.org/html/2609.30471#bib.bib7)\)\.[Kamath et al\. \(2020\)](https://arxiv.org/html/2609.30471#bib.bib11)show that a model’s own confidence is a poor abstention signal under domain shift for question answering, motivating a learned calibrator instead\. LLM\-specific confidence estimation includes self\-evaluation prompts\([Kadavath et al\., 2022](https://arxiv.org/html/2609.30471#bib.bib10)\)and semantic entropy over sampled generations\([Kuhn et al\., 2023](https://arxiv.org/html/2609.30471#bib.bib14)\)\. All of this work abstains the*answering*model\. We instead apply selective prediction to the*evaluator*: the production assistant always answers, butCargomay decline to score that answer when no sufficiently similar golden reference exists, which is a different decision \(is this evaluation trustworthy?\) from the one those methods address \(is this answer trustworthy?\) and admits different signals \(retrieval marginΔ\\Delta, intent\-conditional calibration\) than answer\-side confidence\.

Similar Articles

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse

arXiv cs.LG

This paper benchmarks agentic AI systems on the task of loading, understanding, and reformatting fragmented neuroscience data, finding that while agents perform well on subtasks, they rarely achieve fully error-free end-to-end solutions and human oversight remains necessary.

An Empirical Study of Automating Agent Evaluation

arXiv cs.CL

This paper introduces EvalAgent, a system that automates the evaluation of AI agents by encoding domain-specific expertise, addressing the limitations of standard coding assistants in this task. It also presents AgentEvalBench, a benchmark for testing evaluation pipelines, and demonstrates significant improvements in evaluation reliability.