reliability-assessment

Tag

Cards List
#reliability-assessment

Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

arXiv cs.CL · 2026-09-04 Cached

This paper demonstrates that rubric text alone can predict LLM judge outputs, challenging the assumption of rubric-based evaluation and raising concerns about its reliability in automated text generation assessment.

0 favorites 0 likes
#reliability-assessment

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

arXiv cs.CL · 2026-07-10 Cached

This paper evaluates the reliability of Gemini models as audio judges for scoring full-duplex voice agent conversations, finding that Gemini 2.5 Flash shows strong agreement with human raters on most dimensions, though model swaps require re-validation.

0 favorites 0 likes
← Back to home

Submit Feedback