evaluation-reliability

Tag

Cards List
#evaluation-reliability

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

arXiv cs.AI · 2026-07-22 Cached

This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.

0 favorites 0 likes
#evaluation-reliability

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

arXiv cs.CL · 2026-07-10 Cached

This paper audits the reliability of LLM-as-judge evaluation by showing that changing the evaluator model can shift scores even when candidate responses are fixed, and it examines scaling and upgrade paths for Qwen3 and MiniMax models, concluding that judge upgrades are not interchangeable and proposing best practices for reporting.

0 favorites 0 likes
#evaluation-reliability

The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation

arXiv cs.CL · 2026-06-15 Cached

This paper investigates the run-to-run reliability of LLM-as-a-Judge evaluations, finding that pairwise preferences flip 13.6% of the time on average, with significant first-position bias in GPT-4o-mini, and recommends multi-trial aggregation and position randomization.

0 favorites 0 likes
← Back to home

Submit Feedback