llm-as-a-judge

Tag

Cards List
#llm-as-a-judge

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

arXiv cs.CL ↗ · 18h ago Cached

该论文从心理测量学视角评估 LLM-as-a-judge,使用 Many-Facet Rasch 模型将评分分解为潜在质量、评分者严苛度与残差难度,发现人类与 LLM 在汇总对齐之外仍存在明显的残差难度结构错配。

0 favorites 0 likes
#llm-as-a-judge

@suraj_sharma14: As an AI Evals Engineer, you must build these projects. 1.) Trajectory Grading Engine Build: Step-level evaluator that …

X AI KOLs Timeline ↗ · yesterday Cached

A practitioner outlines 15 must-build projects for AI Evals Engineers, covering trajectory grading, shadow routing, calibrated LLM-as-a-Judge, CI/CD regression gates, adversarial RAG testing, DPO fine-tuning flywheels, statistical significance, agent red-teaming, drift monitoring, cost-quality dashboards, counterfactual replay debugging, synthetic edge-case generation, context-window eviction tests, dataset contamination checks, and public eval methodology teardowns.

0 favorites 0 likes
#llm-as-a-judge

Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging

arXiv cs.AI ↗ · 2026-09-28 Cached

BAER introduces a backbone-adaptive evidence routing method for robust pairwise LLM judging, achieving higher accuracy across benchmarks by dynamically selecting evidence mechanisms per condition.

0 favorites 0 likes
#llm-as-a-judge

CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production

arXiv cs.CL ↗ · 2026-09-28 Cached

This paper introduces Cargo, a framework for evaluating agentic AI systems that addresses reference-instance divergence by grounding factual judgments in live context and gating evaluation on retrieval confidence, along with Cargo-Bench for benchmarking.

0 favorites 0 likes
#llm-as-a-judge

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Hugging Face Daily Papers ↗ · 2026-09-22 Cached

This paper introduces JEV-as-a-Judge, a cost-effective evaluation method for LLMs that uses a decision-only judge with confidence thresholds to accept certain verdicts and escalate uncertain ones, achieving comparable accuracy to state-of-the-art models at significantly lower cost.

0 favorites 0 likes
#llm-as-a-judge

Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits

arXiv cs.CL ↗ · 2026-09-16 Cached

This research investigates whether removing language fields from de-identified résumés prevents demographic leakage in LLM bias audits, finding that non-language prose still allows inference and that evaluation design significantly affects outcomes.

0 favorites 0 likes
#llm-as-a-judge

Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

arXiv cs.LG ↗ · 2026-09-15 Cached

The study investigates whether LLM-as-a-Judge evaluators reliably assess psychological depth in LLM-generated stories, revealing that human preferences are heterogeneous while judges exhibit bias towards reasoning outputs based on surface features.

0 favorites 0 likes
#llm-as-a-judge

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

arXiv cs.CL ↗ · 2026-09-11 Cached

This paper investigates how surface noise in text affects LLM judges' bias measurement, finding that it systematically overestimates bias, particularly in fairness-critical categories, and introduces the Fable benchmark to study this issue.

0 favorites 0 likes
#llm-as-a-judge

@LangChain: Point a coding agent at the LangSmith CLI and it can build you an online LLM-as-a-judge: define the rubric, wire it to …

X AI KOLs Timeline ↗ · 2026-09-10 Cached

A tutorial on using a coding agent with LangSmith CLI to set up an online LLM-as-a-judge for evaluating LLM traces, all configurable from the terminal.

0 favorites 0 likes
#llm-as-a-judge

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

arXiv cs.LG ↗ · 2026-08-24 Cached

This paper presents an LLM-as-a-Judge evaluation framework for agentic AI in drug discovery, validated through human alignment studies with expert annotators. It optimizes the judge to improve alignment with human judgment and provides insights for reusable evaluation in scientific domains.

0 favorites 0 likes
#llm-as-a-judge

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

arXiv cs.AI ↗ · 2026-08-20 Cached

This paper presents a lifecycle framework for using LLMs as judges to evaluate recommendation explanations at Netflix, covering phases from development to deployment and monitoring, with positive A/B test results showing improved user engagement.

0 favorites 0 likes
#llm-as-a-judge

Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper proposes a novel method to mitigate scoring bias in LLM-as-a-Judge by having LLMs randomly generate numbers to measure their latent numerical bias, then rectifying token generation probabilities accordingly. Experiments across four tasks show the method outperforms baselines and reveals that scoring bias varies across models, tasks, and score ranges.

0 favorites 0 likes
#llm-as-a-judge

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

arXiv cs.AI ↗ · 2026-07-22 Cached

This paper introduces SysAdmin, a benchmark that positions frontier language models as autonomous system administrators in a high-fidelity Linux sandbox to measure power-seeking propensity. Across 2800 tasks, the authors find minimal spontaneous power-seeking (0-5% after bias correction) but identify other failure modes such as specification gaming and resistance to goal modification.

0 favorites 0 likes
#llm-as-a-judge

@freeCodeCamp: AI agents can behave differently from one run to the next, which makes regressions hard to catch. In this tutorial, Dar…

X AI KOLs Following ↗ · 2026-07-21 Cached

This tutorial demonstrates how to build a repeatable evaluation harness for AI agents using rule-based checks and an LLM-as-a-judge, leveraging LangChain, Ollama, and Qwen to test local agents with clear pass/fail results.

0 favorites 0 likes
#llm-as-a-judge

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv cs.AI ↗ · 2026-07-15 Cached

This paper introduces GenAI Evaluation, a governed and configuration-driven pipeline for scalable multi-dimensional evaluation of retail conversational agents. It achieves high accuracy using LLM-as-a-judge scoring with selective re-evaluation, validated against human-labeled data.

0 favorites 0 likes
#llm-as-a-judge

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG

arXiv cs.CL ↗ · 2026-07-14 Cached

Introduces Eval-Pair Matrix, a controlled meta-evaluation protocol for source-grounded RAG that induces hidden contradictions to detect self-leniency in LLM judges. The study finds minimal same-model effects and emphasizes methodological improvements for RAG judge studies.

0 favorites 0 likes
#llm-as-a-judge

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

arXiv cs.CL ↗ · 2026-07-03 Cached

This paper analyzes the use of LLM-as-a-Judge in multilingual and low-resource settings, finding inconsistent evaluation outcomes and overtrust in LLM judgments, and provides recommendations for better practices.

0 favorites 0 likes
#llm-as-a-judge

@omarsar0: LLM-as-a-Judge explained in ~10 mins. Knowing how to build AI verifiers and judges is one of the most important emergin…

X AI KOLs Following ↗ · 2026-06-29 Cached

A quick introduction to the LLM-as-a-Judge concept, explaining how to build AI verifiers and judges, and pointing to resources to learn more.

0 favorites 0 likes
#llm-as-a-judge

Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA

arXiv cs.CL ↗ · 2026-06-29 Cached

This paper tests the assumption that LLMs judge better than they generate in in-context QA, finding generation accuracy exceeds self-evaluation on most benchmarks, with evaluation attending less to context. The findings challenge core assumptions in self-evaluation pipelines.

0 favorites 0 likes
#llm-as-a-judge

The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation

arXiv cs.CL ↗ · 2026-06-15 Cached

This paper investigates the run-to-run reliability of LLM-as-a-Judge evaluations, finding that pairwise preferences flip 13.6% of the time on average, with significant first-position bias in GPT-4o-mini, and recommends multi-trial aggregation and position randomization.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback