Tag
该论文从心理测量学视角评估 LLM-as-a-judge,使用 Many-Facet Rasch 模型将评分分解为潜在质量、评分者严苛度与残差难度,发现人类与 LLM 在汇总对齐之外仍存在明显的残差难度结构错配。
A practitioner outlines 15 must-build projects for AI Evals Engineers, covering trajectory grading, shadow routing, calibrated LLM-as-a-Judge, CI/CD regression gates, adversarial RAG testing, DPO fine-tuning flywheels, statistical significance, agent red-teaming, drift monitoring, cost-quality dashboards, counterfactual replay debugging, synthetic edge-case generation, context-window eviction tests, dataset contamination checks, and public eval methodology teardowns.
BAER introduces a backbone-adaptive evidence routing method for robust pairwise LLM judging, achieving higher accuracy across benchmarks by dynamically selecting evidence mechanisms per condition.
This paper introduces Cargo, a framework for evaluating agentic AI systems that addresses reference-instance divergence by grounding factual judgments in live context and gating evaluation on retrieval confidence, along with Cargo-Bench for benchmarking.
This paper introduces JEV-as-a-Judge, a cost-effective evaluation method for LLMs that uses a decision-only judge with confidence thresholds to accept certain verdicts and escalate uncertain ones, achieving comparable accuracy to state-of-the-art models at significantly lower cost.
This research investigates whether removing language fields from de-identified résumés prevents demographic leakage in LLM bias audits, finding that non-language prose still allows inference and that evaluation design significantly affects outcomes.
The study investigates whether LLM-as-a-Judge evaluators reliably assess psychological depth in LLM-generated stories, revealing that human preferences are heterogeneous while judges exhibit bias towards reasoning outputs based on surface features.
This paper investigates how surface noise in text affects LLM judges' bias measurement, finding that it systematically overestimates bias, particularly in fairness-critical categories, and introduces the Fable benchmark to study this issue.
A tutorial on using a coding agent with LangSmith CLI to set up an online LLM-as-a-judge for evaluating LLM traces, all configurable from the terminal.
This paper presents an LLM-as-a-Judge evaluation framework for agentic AI in drug discovery, validated through human alignment studies with expert annotators. It optimizes the judge to improve alignment with human judgment and provides insights for reusable evaluation in scientific domains.
This paper presents a lifecycle framework for using LLMs as judges to evaluate recommendation explanations at Netflix, covering phases from development to deployment and monitoring, with positive A/B test results showing improved user engagement.
This paper proposes a novel method to mitigate scoring bias in LLM-as-a-Judge by having LLMs randomly generate numbers to measure their latent numerical bias, then rectifying token generation probabilities accordingly. Experiments across four tasks show the method outperforms baselines and reveals that scoring bias varies across models, tasks, and score ranges.
This paper introduces SysAdmin, a benchmark that positions frontier language models as autonomous system administrators in a high-fidelity Linux sandbox to measure power-seeking propensity. Across 2800 tasks, the authors find minimal spontaneous power-seeking (0-5% after bias correction) but identify other failure modes such as specification gaming and resistance to goal modification.
This tutorial demonstrates how to build a repeatable evaluation harness for AI agents using rule-based checks and an LLM-as-a-judge, leveraging LangChain, Ollama, and Qwen to test local agents with clear pass/fail results.
This paper introduces GenAI Evaluation, a governed and configuration-driven pipeline for scalable multi-dimensional evaluation of retail conversational agents. It achieves high accuracy using LLM-as-a-judge scoring with selective re-evaluation, validated against human-labeled data.
Introduces Eval-Pair Matrix, a controlled meta-evaluation protocol for source-grounded RAG that induces hidden contradictions to detect self-leniency in LLM judges. The study finds minimal same-model effects and emphasizes methodological improvements for RAG judge studies.
This paper analyzes the use of LLM-as-a-Judge in multilingual and low-resource settings, finding inconsistent evaluation outcomes and overtrust in LLM judgments, and provides recommendations for better practices.
A quick introduction to the LLM-as-a-Judge concept, explaining how to build AI verifiers and judges, and pointing to resources to learn more.
This paper tests the assumption that LLMs judge better than they generate in in-context QA, finding generation accuracy exceeds self-evaluation on most benchmarks, with evaluation attending less to context. The findings challenge core assumptions in self-evaluation pipelines.
This paper investigates the run-to-run reliability of LLM-as-a-Judge evaluations, finding that pairwise preferences flip 13.6% of the time on average, with significant first-position bias in GPT-4o-mini, and recommends multi-trial aggregation and position randomization.