Tag
This paper demonstrates that rubric text alone can predict LLM judge outputs, challenging the assumption of rubric-based evaluation and raising concerns about its reliability in automated text generation assessment.
This paper evaluates the reliability of Gemini models as audio judges for scoring full-duplex voice agent conversations, finding that Gemini 2.5 Flash shows strong agreement with human raters on most dimensions, though model swaps require re-validation.