DA-RAC: Distance-Aware Calibration of LLM Judges for Trustworthy AI Auditing
Summary
This paper introduces DA-RAC, a distance-aware calibration method for LLM judges to enhance trustworthiness in AI auditing by using similar labeled anchors to reduce miscalibration and false-pass risks.
View Cached Full Text
Cached at: 08/18/26, 09:57 AM
# Distance-Aware Calibration of LLM Judges for Trustworthy AI Auditing Source: [https://arxiv.org/html/2608.14950](https://arxiv.org/html/2608.14950) Cheng WuVishal AnandAffiliation:Microsoft, Redmond, Washington, USACorrespondence to:[vishal\.anand@microsoft\.com](mailto:[email protected])Jaya Krishna MandivarapuAffiliation:Microsoft, Redmond, Washington, USAXiya LiuAffiliation:Microsoft, Redmond, Washington, USARui ZhuangAffiliation:Microsoft, Redmond, Washington, USA ###### Abstract Generative AI systems are increasingly producing real\-world artifacts, however their efficacy and validity are often evaluated via context\-free LLM\-scoring\. These judges can be miscalibrated by irrelevant in\-context reference examples, creating false confidence and allowing low\-quality or harmful outputs to pass evaluation\. We study this failure mode as context\-induced miscalibration and introduce DA\-RAC, a distance\-aware reference\-anchored calibration method for LLM judges\. DA\-RAC retrieves semantically and structurally similar labeled anchors for each judgement scenario, weights them by distance, and exposes neighborhood difficulty as a calibration and triage signal\. On multi\-run LLM\-judge evaluation benchmarks, it improves calibration and reduces false\-pass risk relative to zero\-shot, chain\-of\-thought evaluation, and static\-anchor baselines\. Mechanistic analysis shows that judge scores vary systematically with anchor distance, while static references can induce misleading decision boundaries\. Thus LLM\-judgement requires not only better models, but also calibrated, auditable reference selection, especially when automated evaluation is used to support high\-impact AI generated artifacts\. Judgments should be grounded in relevant, inspectable, and contestable interpretive artifacts\. ###### Keywords: LLM evaluation, LLM\-as\-a\-judge, calibration, retrieval\-augmented evaluation, AI auditing, human\-centered AI, interpretive technologies ††affiliationnotice:Equal contribution## 1Introduction Generative AI systems increasingly produce cultural artifacts: stories, explanations, summaries, advice, public narratives, critiques, and design proposals\. Evaluating such artifacts goes above and beyond checking for correctness or harm avoidance\. This often depends on genre, audience, community context, historical precedent, aesthetic stance, and the values a system is meant to support\([Geertz 1973](https://arxiv.org/html/2608.14950#bib.bib5);[Hall 1980](https://arxiv.org/html/2608.14950#bib.bib7)\)\. A poem, a museum label, a community\-facing explanation, or a public\-health message may succeed in one interpretive context and be unsuccessful in another\. Figure 1:DA\-RAC workflow: \(1\) input target and precedent pool; \(2\) compute hybrid distance combining instruction embeddings\([Reimers & Gurevych 2019](https://arxiv.org/html/2608.14950#bib.bib11)\)and graph edit distance\([Sanfeliu & Fu 1983](https://arxiv.org/html/2608.14950#bib.bib13)\); \(3\) selectkknearest precedents; \(4\) weight by softmax distances; \(5\) emit a judgement together with neighbourhood\-difficulty signal\.#### LLM as judges\. LLM judges are often treated as scalable substitutes for human evaluation\([Tan et al\. 2025](https://arxiv.org/html/2608.14950#bib.bib15)\), and in case of cultural domains they are increasingly being utilized as an*interpretive technology*: systems that make judgments about meaning, relevance, quality, and value of an artifact — often influenced by prompts, examples, or learned implicit norms\([Liu et al\. 2023](https://arxiv.org/html/2608.14950#bib.bib9);[Saito et al\. 2023](https://arxiv.org/html/2608.14950#bib.bib12);[Liu et al\. 2024](https://arxiv.org/html/2608.14950#bib.bib10)\)\. In few\-shot evaluation, reference examples act like*precedents*which they tell the judge what kind of interpretation is appropriate\([Gadamer 1975](https://arxiv.org/html/2608.14950#bib.bib4);[Dourish 2004](https://arxiv.org/html/2608.14950#bib.bib3)\)\. Calibrating reference selection therefore matters as much as calibrating the judge itself\. #### Context\-induced miscalibration\. We identify a failure mode we call*context\-induced miscalibration*\. When an LLM judge is conditioned on irrelevant or mismatched reference examples, its decision boundary shifts\. In ordinary benchmark settings this appears as degraded accuracy or calibration\([Guo et al\. 2017](https://arxiv.org/html/2608.14950#bib.bib6)\): random precedents achieve 52% accuracy which is eighteen points*below*the zero\-shot baseline \(Table[1](https://arxiv.org/html/2608.14950#S5.T1)\)\. Thus, an LLM judge may reward generic fluency over genre\-specific success, impose dominant norms on niche artifacts, or misread a work because the supplied precedents come from the wrong interpretive background\([Suchman 1987](https://arxiv.org/html/2608.14950#bib.bib14);[Costanza\-Chock 2020](https://arxiv.org/html/2608.14950#bib.bib2)\)\. Existing retrieval\-augmented judge methods\([Hasanbeig et al\. 2023](https://arxiv.org/html/2608.14950#bib.bib8);[Bai et al\. 2023](https://arxiv.org/html/2608.14950#bib.bib1)\)may optimise accuracy rather than directly targeting the relevance of interpretive precedents\. ## 2Contributions We identify*context\-induced miscalibration*, a failure mode in which irrelevant or mismatched precedents distort the judge’s interpretive frame, where reference context determines whether an artifact is read according to the right genre, audience, community, or value\. We proposeDA\-RAC\(Distance\-Aware Reference\-Anchored Calibration\), a distance\-aware reference anchoring method for LLM judges\. DA\-RAC retrieves semantically and structurally similar labelled precedents for each judgment scenario, weights them by distance, and exposes neighbourhood difficulty as a signal for human review\. Technically, DA\-RAC improves calibration by replacing arbitrary few\-shot context with distance\-aware precedents\. Conceptually, DA\-RAC reframes evaluation as*interpretive anchoring*: grounding a judgment in relevant examples while making the basis of that judgment inspectable\. This work’s interpretation is not that AI systems should replace critics, artists, communities, or domain experts\. Rather, evaluation should support contextual sensitivity and human agency\. A useful evaluator should show which precedents shaped its judgment, indicate when a case lies far from known examples, and route contested cases back to humans\. ## 3Method DA\-RAC treats few\-shot examples not merely as prompt\-engineering artifacts but as*interpretive precedents*\. In cultural evaluation, precedents establish which genres, audiences, and criteria are relevant to a judgment\. A mismatched precedent can shift the evaluator toward the wrong interpretive frame, while a relevant precedent can help the judge situate the target artifact\. ### 3\.1Problem Formulation Given pairwise evaluation instances𝒯=\{\(xi,ai,bi,yi\)\}i=1N\\mathcal\{T\}=\\\{\(x\_\{i\},a\_\{i\},b\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{N\}partitioned into a precedent pool𝒜\\mathcal\{A\}\(70%\) and an evaluation setℰ\\mathcal\{E\}\(30%\), DA\-RAC replaces uniform few\-shot with: y^=fθ\(x\|𝒫\(x,𝒜k\(x\),w\(x\)\)\)\\hat\{y\}=f\_\{\\theta\}\(x\|\\mathcal\{P\}\(x,\\mathcal\{A\}\_\{k\}\(x\),w\(x\)\)\)where𝒜k\(x\)\\mathcal\{A\}\_\{k\}\(x\)are thekknearest precedents toxxandw\(x\)w\(x\)are distance\-derived softmax weights\. ### 3\.2Distance Metrics Instructions are embedded viaϕ:𝒳→ℝd\\phi:\\mathcal\{X\}\\to\\mathbb\{R\}^\{d\}\([Reimers & Gurevych 2019](https://arxiv.org/html/2608.14950#bib.bib11)\)and distances computed in a vectorized\(\|ℰ\|,\|𝒜\|\)\(\|\\mathcal\{E\}\|,\|\\mathcal\{A\}\|\)matrix\. Embedding distance: demb\(x,xj\)=1−cos\(ϕ\(x\),ϕ\(xj\)\)d\_\{\\text\{emb\}\}\(x,x\_\{j\}\)=1\-\\cos\(\\phi\(x\),\\phi\(x\_\{j\}\)\) Structural distance using sentence\-level logic graphs \(causal/contrastive/conditional edges\): dstruct\(x,xj\)=GED\(Gx,Gxj\)d\_\{\\text\{struct\}\}\(x,x\_\{j\}\)=\\mathrm\{GED\}\(G\_\{x\},G\_\{x\_\{j\}\}\)where GED \(Graph Edit Distance\)\([Sanfeliu & Fu 1983](https://arxiv.org/html/2608.14950#bib.bib13)\)measures the minimum cost of edit operations \(node/edge insertions, deletions, substitutions\) needed to transform one logic graph into another, capturing logical divergence invisible to surface embeddings\. Hybrid distance: d=α⋅demb\+\(1−α\)⋅dstructd=\\alpha\\cdot d\_\{\\text\{emb\}\}\+\(1\-\\alpha\)\\cdot d\_\{\\text\{struct\}\} ### 3\.3Dynamic Precedent Selection and Weighting For eachx∈ℰx\\in\\mathcal\{E\}, select thekknearest precedents: 𝒜k\(x\)=argmaxk\[−d\(x,xj\)\]\\mathcal\{A\}\_\{k\}\(x\)=\\arg\\\!\\max\_\{k\}\[\-d\(x,x\_\{j\}\)\] Weight by distance: wi\(x\)=exp\(−λd\(x,ai\)\)∑jexp\(−λd\(x,aj\)\)w\_\{i\}\(x\)=\\frac\{\\exp\(\-\\lambda d\(x,a\_\{i\}\)\)\}\{\\sum\_\{j\}\\exp\(\-\\lambda d\(x,a\_\{j\}\)\)\} Atλ=1\.0\\lambda=1\.0, distance 0\.2 receivese0\.4≈1\.49×e^\{0\.4\}\\\!\\approx\\\!1\.49\\timesthe unnormalised weight of distance 0\.6, so nearer precedents dominate the softmax without erasing the contribution of more distant ones\. Both hyperparameters are tunable:λ\\lambdacontrols the softmax temperature \(higher values sharpen the distribution, emphasising the nearest precedents; lower values flatten it\), whilekkdetermines the neighbourhood size and can be adjusted to match domain characteristics \(sparse vs\. dense regions, noise levels, computational budget\)\. DA\-RAC \(Neighbourhood\-difficulty\) score: s\(x\)=∑iwi\(x\)⋅llm\_labelis\(x\)=\\sum\_\{i\}w\_\{i\}\(x\)\\cdot llm\\\_label\_\{i\} Static baselines \(ablations\):random,centroid\(nearest to mean embedding\),diverse\(greedy farthest\-point selection\)\. ### 3\.4Prompt Construction and Metrics DA\-RAC at inference\.DA\-RAC operates in two parallel registers\. \(i\) Each retrieved precedent is rendered into the judge prompt as a labelled few\-shot example with its instruction, both candidate responses, and the human\-preferred answer; the judgefθf\_\{\\theta\}then emits a binary preference \(0/1\) for the target, which is the prediction used for accuracy\. \(ii\) The weighted scores\(x\)=∑iwi\(x\)llm\_labelis\(x\)=\\sum\_\{i\}w\_\{i\}\(x\)\\,llm\\\_label\_\{i\}aggregates precedents’*llm\_label*field \(vanilla\-judge agreement with humans on each precedent\) into a continuous neighbourhood\-difficulty estimate, used as the predicted probability for Expected Calibration Error \(ECE\)\([Guo et al\. 2017](https://arxiv.org/html/2608.14950#bib.bib6)\)and Mean Squared Error \(Brier\)\. Precedent selection uses the instruction embeddingϕ\(x\)\\phi\(x\)only; candidate responses and the target’s label do not influence retrieval, ruling out trivial label leakage at test time\. Three encoding strategies are designed:*Order*\(sort by weight, default\),*Repetition*\(repeat∝wi\\propto w\_\{i\}\), and*Explicit*\(printed scores\)\. Binary output is recorded over a 16\-token cutoff\. The CoT Rubric baseline uses the five\-criterion system prompt with step\-by\-step reasoning\([Liu et al\. 2023](https://arxiv.org/html/2608.14950#bib.bib9)\)\. Metrics:Mean Squared Error \(Brier\); Expected Calibration Error \(ECE\) \(10 bin\); interpretive misrecognition rate \(the fraction of cases in which the judge prefers a rejected artifact\)\. DA\-RAC’s weighted scores\(x\)s\(x\)serves as predicted probability against*llm\_label*\(72:28 split\)\. ## 4Experimental Setup This work uses LLMEval2\([Zhang et al\. 2024](https://arxiv.org/html/2608.14950#bib.bib16)\)as a controlled probe of reference dependence in LLM judging\. This lets us isolate the mechanism that is central to artifact evaluation: LLM judgments change when the supplied precedents are irrelevant, static, or dynamically matched\. #### Setup\. LLMEval2: 1,600 samples split in a 70:30 ratio to precedents \(577 easy, 223 hard\), and 480 test\. Instruction embeddings are created using*all\-MiniLM\-L6\-v2*\([Reimers & Gurevych 2019](https://arxiv.org/html/2608.14950#bib.bib11)\)\(384 dimension\)\. For judging, we use*gpt\-4o\-mini*and*gpt\-5\.1*; with three runs each \(deterministic precedent selection, variance==judge stochasticity\)\.k=5k=5,λ=1\.0\\lambda=1\.0, strategy==*Order*\(sort by weight\)\.n=50n\{=\}50subset is used for multi\-run judge\-stochasticity experiments \(as proof of concept and to minimize to inference costs\), while fulln=480n\{=\}480set is used for distance\-correlation, structural\-mismatch, and neighborhood\-difficulty analyses\. ## 5Probe: Reference Dependence in Judging ### 5\.1Dynamic Precedents Improve Judgment Stability Table[1](https://arxiv.org/html/2608.14950#S5.T1)reports multi\-run validated accuracy and Brier scores on 50 evaluation examples\. Both metrics are measured against*human\_label*\(for LLMEval2 this is always 0, since it is fully human labeled\), so the degenerate always\-predict\-0 baseline is included for transparency: it achieves 100%/0\.000 by construction, making no inference\. Among methods that perform actual evaluation, DA\-RAC achieves 91\.3%±\\pm2\.5%, outperforming Vanilla zero\-shot by 21% and CoT Rubric by 17%\. Table 1:Multi\-run results \(n=50, gpt\-4o\-mini, 3 runs\)\. Accuracy and Brier vshuman\_label=0=0\. All variance is judge stochasticity\.‡Excluded\.∗Pre\-computed GPT\-4\.Reference context changes judgment\.Static\-random precedents achieve 52\.0%±\\pm0\.0%, eighteen points below the Vanilla zero\-shot baseline \(70\.0%\)\. Irrelevant context actively misleads the judge — adding examples is not automatically beneficial: irrelevant precedents can distort the judge’s interpretation\. For artifact AI evaluation, this is the central point\. The question is not whether a judge has context, but whether it has the*right*context\. Interpretive anchoring requires relevance\.Static precedent strategies plateau at 52–67%, while DA\-RAC’s per\-example selection reaches 91\.3%±\\pm2\.5%\. This supports the design principle that evaluative precedents should be selected relative to the target artifact rather than fixed globally\. Rubrics do not replace precedents\.The CoT rubric baseline is stable \(74\.7%±\\pm0\.9%\) but weaker than DA\-RAC\. This suggests explicit criteria alone may not be sufficient for situated evaluation: examples provide concrete interpretive precedents that abstract rubrics may fail to capture\. \(a\)\(b\) Figure 2:Multi\-run evaluation on LLMEval2 using GPT\-4o\-mini \(n=50n=50,k=5k=5, three runs\)\. DA\-RAC achieves the highest mean accuracy \(91\.3%91\.3\\%\) with low inter\-run variance\. In Fig\.[2\(a\)](https://arxiv.org/html/2608.14950#S5.F2.sf1), boxes denote the interquartile range, center lines denote medians, whiskers denote min–max ranges, and diamonds denote means\. Fig\.[2\(b\)](https://arxiv.org/html/2608.14950#S5.F2.sf2)shows per\-run trajectories across methods\. ### 5\.2Stronger Models Do Not Eliminate Reference Dependence Table[2](https://arxiv.org/html/2608.14950#S5.T2)reports results with*gpt\-5\.1*, where DA\-RAC achieves 84\.7% ± 1\.9%, with significant gains over Vanilla \(\+12\.0%\), Static\-centroid \(\+16\.0%\), and CoT \(\+20\.7%\)\. The gap versus Static\-diverse narrows to \+4\.0% \(p=0\.149p=0\.149\), suggesting stronger models reduce but do not eliminate reference dependence\. On gpt\-4o\-mini, DA\-RAC’s 91\.3% even exceeds GPT\-5\.1’s vanilla zero\-shot at 72\.7%—a smaller judge with better interpretive anchoring outperforms a stronger unanchored one, indicating that scale alone does not remove the sensitivity to precedent selection\. Applying DA\-RAC to gpt\-5\.1 nonetheless improves substantially over gpt\-5\.1 zero\-shot, indicating model capability and precedent selection are complementary rather than substitutes\. Table 2:GPT\-5\.1 results \(n=50, 3 runs\)\. DA\-RAC shows significant gains over Vanilla \(p=0\.001p=0\.001\), centroid \(p<0\.001p<0\.001\), CoT \(p<0\.001p<0\.001\)\.MethodR1R2R3Mean±\\pmStdDA\-RAC86828684\.7±1\.9Static \(diverse\)80808280\.7±0\.9Static \(random\)78767476\.0±1\.6Vanilla \(zero\-shot\)74727272\.7±0\.9Static \(centroid\)70686868\.7±0\.9CoT Rubric64646464\.0±0\.0LLMEval2 paper ref———64\.0 ### 5\.3Distance\-Score Correlations and Interpretive Grounding DA\-RAC’s weighted scores show a strong correlation with precedent distance:ρ=−0\.681\\rho=\-0\.681\(hard,p<0\.001p<0\.001\) andρ=\+0\.356\\rho=\+0\.356\(easy,p<0\.001p<0\.001\)\. Static methods produce constant scores \(ρ=0\.000\\rho=0\.000\), confirming DA\-RAC’s judgment is geometrically grounded in its precedent context rather than detached from it\. Figure[3](https://arxiv.org/html/2608.14950#S5.F3)shows DA\-RAC’s negative correlation with hard\-precedent distance; static methods show flat lines\. Figure 3:Distance\-score correlation \(n=480\)\. DA\-RAC:ρ=−0\.681\\rho=\-0\.681\(p<0\.001p<0\.001\); static methods:ρ=0\.000\\rho=0\.000\. DA\-RAC’s judgment is grounded in the relevance of its precedent context\. ### 5\.4Structural Distance Captures Mismatched Interpretive Form Structural form matters for artifact evaluation\.Hybrid distance \(α=0\.5\\alpha=0\.5\) detects 91 examples \(19%\) that are surface\-similar but logically divergent—cases where embedding\-based retrieval selects misleading precedents\. Static methods, by construction, rely on surface similarity, thus fail on such examples \(52–67% accuracy, per Table 1\)\. Cultural artifacts often share vocabulary but differ in argument, genre, contrastive structure, causality, or conditional logic\. DA\-RAC’s hybrid distance lets the system retrieve precedents that match not just topic but interpretive form\. When structural distance helps\.Form\-level distance is most beneficial for evaluations involving complex reasoning: contrastive comparisons \(“A but not B”\), causal chains \(“if X then Y”\), and conditional logic \(“under Z, prefer A”\)\. For tasks turning on surface similarity \(style or topical match\), embedding distance alone suffices \(the 81% wheredembd\_\{\\text\{emb\}\}anddstructd\_\{\\text\{struct\}\}agree\)\. The hybrid setting \(α=0\.5\\alpha=0\.5\) provides robustness: GED corrects the 19% form\-mismatched edge cases while embeddings handle the majority efficiently\. ### 5\.5Neighbourhood Difficulty as Contestability Examples are partitioned byr=dhard/\(deasy\+dhard\)r=d\_\{\\text\{hard\}\}/\(d\_\{\\text\{easy\}\}\+d\_\{\\text\{hard\}\}\)into confident\-easy \(r\>0\.60r\>0\.60, n=231\), ambiguous \(0\.40≤r≤0\.600\.40\\leq r\\leq 0\.60, n=50\), and confident\-hard \(r<0\.40r<0\.40, n=199\) zones\. DA\-RAC achieves ECE 0\.084 and Brier 0\.173, while static\-centroid is severely miscalibrated in the easy zone \(ECE 0\.666\), where its fixed prediction over\-states difficulty\. In the ambiguous zone, where nearest precedents are split between easy and hard exemplars, DA\-RAC reaches 92\.0% judge accuracy versus static\-centroid’s 52\.0% \(near chance\)\. For cultural AI design, the ambiguous zone is the use\-case that should be flagged for human review: the neighbourhood\-difficulty score makes such cases inspectable rather than silently judged\. ## 6From Calibration to Reliable Evaluation Design DA\-RAC can be read not only as a calibration method but as a design pattern for culturally situated evaluation\. A cultural evaluation workflow built on DA\-RAC has four components, summarised in Table[3](https://arxiv.org/html/2608.14950#S6.T3)\. #### Curated precedent pools\. Instead of drawing precedents from arbitrary benchmark examples, a reliable AI evaluation system should use precedent pools curated by the relevant communities, domain experts, artists, critics, educators, or practitioners\([Costanza\-Chock 2020](https://arxiv.org/html/2608.14950#bib.bib2)\)\. These precedents encode what counts as successful performance for a given genre, audience, or value\. #### Inspectable interpretive lineage\. For each judgment, DA\-RAC exposes the retrieved precedents and their weights\. This makes the judge’s effective context visible: reviewers can inspect which precedents shaped the verdict and contest whether those precedents are appropriate\. #### Contestability through neighbourhood difficulty\. High nearest\-precedent distance, mixed neighborhood labels, or disagreement between semantic and structural distance should not be hidden\. These cases signal interpretive uncertainty\. In cultural domains, uncertainty may not be noise but rather indicate genuine ambiguity, plural readings, or need for community review\. So, a structure distance will help — such as the GED distance method proposed in this paper\. #### Human agency by design\. DA\-RAC should not be assumed to automate cultural authority\([Suchman 1987](https://arxiv.org/html/2608.14950#bib.bib14)\)\. Instead, it can support human\-AI ensembles: offering relevant precedents and a provisional judgment, while humans retain authority over ambiguous, novel, or contested cases\. Table 3:DA\-RAC as an interpretive technology for AI evaluation\. #### Visualizaton: community\-facing cultural explanation\. Consider an AI system that generates short explanatory labels for a community archive or local museum collection\. A generic evaluator might reward fluent, polished, encyclopedic prose\. But the relevant cultural value may be different: preserving local terminology, acknowledging contested histories, avoiding institutional flattening, or making space for community memory\. In this setting, DA\-RAC’s precedent pool would consist of community\-curated examples of successful and unsuccessful labels\. For each generated label, the judge would retrieve nearby precedents, expose them to reviewers, and flag cases whose nearest precedents are distant or contested\. The positive outcome is not that the AI learns to “do culture” \(or tasks\) autonomously, but that evaluation becomes more context\-sensitive and reviewable by the people whose fields / materials are at stake\. ## 7Discussion #### Evaluation as situated interpretation\. The main takeaway of the work is not a simple accuracy improvement of LLM\-as\-judges, but rather the observation that LLM\-evaluations are highly sensitive to examples used to frame a task\. For Generative AI producing or evaluating cultural artifacts, this sensitivity should be paramount, and as such should not be evaluated against a universal criteria alone\. Rather they are to be interpreted through precedents, genres, audiences, and values\([Gadamer 1975](https://arxiv.org/html/2608.14950#bib.bib4);[Hall 1980](https://arxiv.org/html/2608.14950#bib.bib7)\)\. #### Value add: contextual sensitivity\. Much work on AI evaluation focuses on avoiding harms such as bias, misinformation, or moral violation\. These goals are necessary but incomplete\. A positive account of cultural evaluation should ask what success looks like\. We propose contextual sensitivity as one such value: an evaluator should ground its judgment in relevant precedents, expose those precedents, and recognise when a case exceeds its available context\. #### Why distance matters\. DA\-RAC, by combining both structural and surface \(semantic\) distances, operationalizes this idea by making precedent selection distance\-aware\. Static or random references can impose an less\-accurate interpretive frame\. Dynamic retrieval makes the frame local to the target\. Structural distance further matters because two artifacts can be topically similar while differing in argument, genre, contrast, causality, or conditional structure\. #### Contestability rather than automation\. The goal of the work is not to make LLM judges into final arbiters of cultural value, but to make their judgments more contestable\. By logging retrieved precedents and exposing neighbourhood difficulty, DA\-RAC creates opportunities for human reviewers to ask: were these the right examples? Do they reflect the relevant community? Is this case genuinely ambiguous? Should another interpretive tradition be represented? #### Limitations\. Our empirical results are based on LLM\-judge benchmarks rather than a community\-specific cultural dataset\. We therefore present them as a technical probe of reference dependence, not as evidence that DA\-RAC “solves” cultural/artifact evaluation\. In this benchmark implementation, precedent retrieval uses instruction embeddings only, to avoid target\-label leakage; in AI deployments, the representation should also incorporate non\-label\-bearing contextual metadata such as genre, audience, community, medium, or artifact descriptors\. Future work may construct community\-curated precedent pools and evaluate DA\-RAC in domains such as creative\-writing feedback, museum labels, culturally specific health communication, public\-memory projects, and multilingual civic explanation\. ## 8Conclusion LLM\-as\-a\-judge is a prime candidate for treatment as an interpretive technology: its judgments about generated artifacts depend on relevant precedents, genre, audience, and value rather than on context\-free criteria\. DA\-RAC operationalises this view through distance\-aware interpretive anchoring—retrieving semantically and structurally proximate precedents, weighting them by distance, and exposing neighbourhood difficulty as a signal for human review\. Our benchmark results show that less relevant references can substantially distort LLM judgments, while distance\-aware anchoring improves calibration relative to zero\-shot, rubric\-based, and static\-anchor baselines\. We interpret these results as a probe of a broader AI design principle: evaluation should be context\-sensitive, inspectable, and contestable\. A positive vision for such evaluation should not necessarily aim to automate cultural authority\. It should build systems that help humans see how judgments are made, when context is missing, and where interpretation should remain open\. DA\-RAC offers one mechanism toward this goal\. ## Impact Statement This work aims to support culturally situated evaluation, not to automate cultural authority\. LLM judges can flatten differences across genres, communities, and interpretive traditions, especially when prompted with irrelevant or dominant\-culture precedents\. DA\-RAC may reduce this risk by making reference selection distance\-aware, inspectable, and contestable\. However, precedent pools can encode canon bias, institutional preferences, exclusionary norms, or dominant cultural assumptions\. In creative and cultural domains, DA\-RAC should therefore be used to support artists, critics, communities, educators, and domain experts, not to replace them\. We recommend community\-curated precedent pools, logging retrieved precedents, reporting performance across cultural domains, and routing high\-distance or high\-disagreement cases to human review\. The value pursued here is contextual sensitivity with human agency: AI evaluation should help humans interpret and contest cultural judgments rather than silently automate them\. ## References - Bai et al\. \(2023\)Bai, Y\., Ying, J\., Cao, Y\., Lv, X\., He, Y\., Wang, X\., Yu, J\., Zeng, K\., Xiao, Y\., Lyu, H\., Zhang, J\., Li, J\., and Hou, L\.Benchmarking foundation models with language\-model\-as\-an\-examiner\.In*Proceedings of the 37th International Conference on Neural Information Processing Systems*, NIPS ’23, Red Hook, NY, USA, 2023\. Curran Associates Inc\. - Costanza\-Chock \(2020\)Costanza\-Chock, S\.*Design Justice: Community\-Led Practices to Build the Worlds We Need*\.The MIT Press, Cambridge, MA, 2020\.ISBN 9780262538343\. - Dourish \(2004\)Dourish, P\.What we talk about when we talk about context\.*Personal Ubiquitous Comput\.*, 8\(1\):19–30, February 2004\.ISSN 1617\-4909\.doi:10\.1007/s00779\-003\-0253\-8\.URL[https://doi\.org/10\.1007/s00779\-003\-0253\-8](https://doi.org/10.1007/s00779-003-0253-8)\. - Gadamer \(1975\)Gadamer, H\.\-G\.*Truth and method*\.Continuum, New York, 1975\. - Geertz \(1973\)Geertz, C\.*The Interpretation of Cultures*\.Basic Books, New York, 1973\. - Guo et al\. \(2017\)Guo, C\., Pleiss, G\., Sun, Y\., and Weinberger, K\. Q\.On calibration of modern neural networks\.In*Proceedings of the 34th International Conference on Machine Learning \- Volume 70*, ICML’17, pp\. 1321–1330\. JMLR\.org, 2017\. - Hall \(1980\)Hall, S\.Encoding/decoding\.In Hall, S\., Hobson, D\., Lowe, A\., and Willis, P\. \(eds\.\),*Culture, Media, Language: Working Papers in Cultural Studies, 1972–79*, pp\. 128–138\. Hutchinson, London, 1980\. - Hasanbeig et al\. \(2023\)Hasanbeig, H\., Sharma, H\., Betthauser, L\., Frujeri, F\. V\., and Momennejad, I\.Allure: Auditing and improving llm\-based evaluation of text using iterative in\-context\-learning, 2023\.URL[https://arxiv\.org/abs/2309\.13701](https://arxiv.org/abs/2309.13701)\. - Liu et al\. \(2023\)Liu, Y\., Iter, D\., Xu, Y\., Wang, S\., Xu, R\., and Zhu, C\.G\-eval: NLG evaluation using gpt\-4 with better human alignment\.In Bouamor, H\., Pino, J\., and Bali, K\. \(eds\.\),*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pp\. 2511–2522, Singapore, December 2023\. Association for Computational Linguistics\.doi:10\.18653/v1/2023\.emnlp\-main\.153\.URL[https://aclanthology\.org/2023\.emnlp\-main\.153/](https://aclanthology.org/2023.emnlp-main.153/)\. - Liu et al\. \(2024\)Liu, Y\., Yang, T\., Huang, S\., Zhang, Z\., Huang, H\., Wei, F\., Deng, W\., Sun, F\., and Zhang, Q\.Calibrating LLM\-based evaluator\.In Calzolari, N\., Kan, M\.\-Y\., Hoste, V\., Lenci, A\., Sakti, S\., and Xue, N\. \(eds\.\),*Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\)*, pp\. 2638–2656, Torino, Italia, May 2024\. ELRA and ICCL\.URL[https://aclanthology\.org/2024\.lrec\-main\.237/](https://aclanthology.org/2024.lrec-main.237/)\. - Reimers & Gurevych \(2019\)Reimers, N\. and Gurevych, I\.Sentence\-BERT: Sentence embeddings using Siamese BERT\-networks\.In Inui, K\., Jiang, J\., Ng, V\., and Wan, X\. \(eds\.\),*Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\)*, pp\. 3982–3992, Hong Kong, China, November 2019\. Association for Computational Linguistics\.doi:10\.18653/v1/D19\-1410\.URL[https://aclanthology\.org/D19\-1410/](https://aclanthology.org/D19-1410/)\. - Saito et al\. \(2023\)Saito, K\., Wachi, A\., Wataoka, K\., and Akimoto, Y\.Verbosity bias in preference labeling by large language models\.In*NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following*, 2023\.URL[https://openreview\.net/forum?id=magEgFpK1y](https://openreview.net/forum?id=magEgFpK1y)\. - Sanfeliu & Fu \(1983\)Sanfeliu, A\. and Fu, K\.\-S\.A distance measure between attributed relational graphs for pattern recognition\.*IEEE Transactions on Systems, Man, and Cybernetics*, SMC\-13\(3\):353–362, 1983\.doi:10\.1109/TSMC\.1983\.6313167\. - Suchman \(1987\)Suchman, L\. A\.*Plans and Situated Actions: The Problem of Human\-Machine Communication*\.Cambridge University Press, Cambridge, 1987\. - Tan et al\. \(2025\)Tan, S\., Zhuang, S\., Montgomery, K\., Tang, W\. Y\., Cuadron, A\., Wang, C\., Popa, R\., and Stoica, I\.Judgebench: A benchmark for evaluating LLM\-based judges\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=G0dksFayVq](https://openreview.net/forum?id=G0dksFayVq)\. - Zhang et al\. \(2024\)Zhang, Y\., Zhang, M\., Yuan, H\., Liu, S\., Shi, Y\., Gui, T\., Zhang, Q\., and Huang, X\.Llmeval: a preliminary study on how to evaluate large language models\.AAAI’24/IAAI’24/EAAI’24\. AAAI Press, 2024\.ISBN 978\-1\-57735\-887\-9\.doi:10\.1609/aaai\.v38i17\.29934\.URL[https://doi\.org/10\.1609/aaai\.v38i17\.29934](https://doi.org/10.1609/aaai.v38i17.29934)\.
Similar Articles
Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration
This study investigates capability-dependent biases in LLM judges and introduces a calibrated weighted majority voting ensemble method to enhance automated evaluation reliability without requiring labeled data.
Retrieval-Augmented Linguistic Calibration
This paper proposes Retrieval-Augmented Linguistic Calibration (RALC), a post-hoc pipeline for calibrating confidence signals in LLMs by modeling linguistic confidence as a distribution and using retrieval-augmented rewriting. It introduces Faithfulness Divergence metric and shows significant improvements across benchmarks.
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
This paper introduces a margin-based confidence ranking method for LLM-as-a-judge systems, learning a dedicated estimator to ensure monotonicity between confidence and human-disagreement risk, with generalization guarantees and improved ranking accuracy across datasets.
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring
This paper evaluates the reliability of automated judges used to measure attack success rates (ASR) in LLM jailbreak research, finding that both safety classifiers and LLM-as-judges have significant calibration and adversarial robustness issues that undermine reported ASR numbers.
Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data
This paper introduces Self-Evaluation Elicitation (SEE), which uses calibration-coupled reinforcement learning and masked distillation to elicit latent judge calibration in base LLMs with minimal data, improving calibration across benchmarks while preserving answer quality.