Tag
This paper demonstrates that rubric text alone can predict LLM judge outputs, challenging the assumption of rubric-based evaluation and raising concerns about its reliability in automated text generation assessment.
An LLM judge consistently returned a fixed confidence score of 0.72 in evaluations, but switching to categorical labels improved score distribution, showing that models are better at classification than numerical estimation for assessments.
Introduces counterfactual audits to test whether audio language model judges actually use paralinguistic evidence when evaluating speech-to-speech responses, finding that contrastive success often overstates native reliability and similar accuracies can hide different failure modes across Gemini, GPT, and open models.
BACON proposes a four-stage pipeline that combines budgeted human labels with multiple AI judge outputs to produce calibrated item-level surrogate predictions, supporting both population-level estimation and individual-level scoring with improved accuracy and reduced bias.
A quick introduction to the LLM-as-a-Judge concept, explaining how to build AI verifiers and judges, and pointing to resources to learn more.