llm-judges

Tag

Cards List
#llm-judges

Everyone says write evals for your agent. But what should you actually test?

Reddit r/AI_Agents · yesterday

A practical guide on writing effective evaluations for AI agents, focusing on starting from observed failures and using a mix of deterministic checks and LLM judges.

0 favorites 0 likes
#llm-judges

LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes

Hacker News Top · yesterday Cached

This paper introduces a benchmark for evaluating LLM judges' ability to detect omissions in AI-generated clinical notes, finding that standard judges struggle with omissions but restructuring the task into per-fact verification improves detection.

0 favorites 0 likes
#llm-judges

SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents

arXiv cs.AI · yesterday Cached

SAGE is a novel evaluation framework for task-oriented dialogue agents that grounds assessments in dialogue state changes and uses abstention to provide cost-effective, accurate turn-level judgments without expensive LLM calls.

0 favorites 0 likes
#llm-judges

Toward Cultural Alignment: Human-Centered Evaluation of Multimodal AI Stories Across Five African Communities

arXiv cs.CL · 2d ago Cached

This paper examines cultural alignment in AI-generated multimodal stories across five African communities through community-grounded evaluation, developing a taxonomy of misalignment and assessing the reliability of automated judges.

0 favorites 0 likes
#llm-judges

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

arXiv cs.CL · 2026-08-24 Cached

JuryProbe introduces an empirical diagnostic for assessing consensus risk in reference-free LLM judge panels for factuality checking, and uses a calibration-based routing policy to ground high-risk decisions with trusted references, reducing false accepts.

0 favorites 0 likes
#llm-judges

@mervenoyann: interesting talk by @willcb

X AI KOLs Timeline · 2026-08-20 Cached

This talk by Will Brown of Primordial AI discusses techniques for scaling Reinforcement Learning to complex, real-world tasks where rewards are not verifiable, using methods like anchoring, LLM judges, and simulation.

0 favorites 0 likes
#llm-judges

@browser_use: Opus 5 and GPT-5.6 Sol are neck-and-neck on this!

X AI KOLs Timeline · 2026-08-20 Cached

Alexander Yue introduces a new browser-use benchmark where Opus 5 and GPT-5.6 Sol show similar performance, emphasizing the benchmark's robust design with verified rubrics for LLM judges.

0 favorites 0 likes
#llm-judges

Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots

arXiv cs.AI · 2026-08-20 Cached

The paper presents EvalCEGAR, a method for automatically evolving evaluation metrics using a pool of Python operators that flag specific defects in AI outputs, improving accuracy over hand-written operators and LLM judges.

0 favorites 0 likes
#llm-judges

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

arXiv cs.CL · 2026-08-20 Cached

This paper investigates bidirectional bias in LLM judges induced by self- and other-labels, showing that labels alone can shift evaluation scores regardless of actual source, with contributions to understanding authorship attribution and controlled evaluation tasks.

0 favorites 0 likes
#llm-judges

Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation

arXiv cs.LG · 2026-08-18 Cached

This paper addresses rubric interference in LLM judges when evaluating multiple rubrics simultaneously and proposes Self-Anchored Rubric Alignment (SARA) to improve consistency through on-policy self-distillation.

0 favorites 0 likes
#llm-judges

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

arXiv cs.AI · 2026-08-14 Cached

Introduces the Wiggle Framework, a unified stress test for the epistemic stability of LLM judges, measuring mechanical consistency, single-turn conviction, and multi-turn persistence across 14 tasks. Finds that all 9 studied frontier models flip verdicts substantially under pressure, and that successful pressure is usually net-corrupting relative to ground truth.

0 favorites 0 likes
#llm-judges

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

arXiv cs.AI · 2026-08-13 Cached

The paper introduces a behavioral alignment framework for personalized LLM judges in recommendation evaluation, addressing bidirectional rationalization where off-the-shelf LLMs argue both for and against user engagement on the same item. Their fine-tuned and preference-optimized approach achieves a 32.19% Macro-F1 lift over zero-shot and matches production feature-engineered baselines.

0 favorites 0 likes
#llm-judges

Benchmarking LLM Judges for Mobile Agent Evaluation

arXiv cs.AI · 2026-08-13 Cached

This paper introduces MobileJudgeBench, a benchmark with 931 human-annotated trajectories for systematically evaluating LLM-based judges on mobile agent tasks. It finds that simple baseline judges with sampled screenshots rival purpose-built methods, with the LLM backbone being the primary driver of quality.

0 favorites 0 likes
#llm-judges

@omarsar0: LLM review weirdness indeed. Avoid using scores with LLM judges, or be extremely careful if you do. Use binary labels w…

X AI KOLs Following · 2026-08-10 Cached

A tweet discussing a discovered quirk where renaming a paper PDF to a longer, positive title improves LLM judge scores, advising caution with score-based LLM evaluation and recommending binary labels instead.

0 favorites 0 likes
#llm-judges

Averaging Bias: Human Faithfulness Annotations are not Locally Faithful

arXiv cs.CL · 2026-08-04 Cached

This paper from Stony Brook University identifies 'Averaging Bias' in human faithfulness annotations for text summarization: global human labels correlate better with the average of per-sentence LLM judgments than with a strict conjunctive rule, meaning humans often label summaries as faithful even when they contain local factual errors.

0 favorites 0 likes
#llm-judges

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

arXiv cs.CL · 2026-08-03 Cached

A research paper introducing Chain-of-Models (CoM), an automated pipeline where a second LLM audits a first model's reasoning trace to correct cognitive biases. It finds that auditor effectiveness depends on model family and bias type, and proposes a bias-specific auditor selection rule that improves judgment accuracy.

0 favorites 0 likes
#llm-judges

@hanakoxbt: https://x.com/hanakoxbt/status/2083540339147567268

X AI KOLs Timeline · 2026-08-01 Cached

A six-step guide to building evaluation gates that let AI agents merge changes autonomously, covering judge bias, runtime evals, trajectory grading, and more.

0 favorites 0 likes
#llm-judges

LLM Judges Can Be Too Generous When There Is No Reference Answer

arXiv cs.CL · 2026-07-15 Cached

This paper shows that LLM judges tend to over-credit incorrect answers when no reference answer is provided, and adding a reference can flip verdicts by up to 85%, aligning more with human judgments. The authors propose calibration steps for using LLM judges in reference-free settings.

0 favorites 0 likes
#llm-judges

@HamelHusain: New Blog Post: Do Automated Evals Work? There has been a rise of tools that look through your traces with AI and identi…

X AI KOLs Timeline · 2026-07-14 Cached

A blog post from Parlance Labs tests automated AI evaluation tools (Braintrust Loop, Arize Alyx, LangSmith Engine) on real production data, finding they catch 87% of issues humans flag but miss domain-specific failures and add noise, recommending iterative human-in-the-loop use.

0 favorites 0 likes
#llm-judges

@h100envy: Prime Intellect engineers explained how they train reasoning models over the open internet in 30 minutes - better than …

X AI KOLs Timeline · 2026-07-12 Cached

Prime Intellect engineers demonstrated a method to train reasoning models in 30 minutes using distributed RL over the open internet, utilizing Prime-RL, LLM judges, and multi-cloud GPUs, enabling open models to compete with closed labs without owning data centers.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback