Tag
This paper introduces UPHELD, a large benchmark for evaluating human-scale conversational ability in LLMs, and proposes a Mixture-of-Judges framework that improves correlation with human assessments by approximately 30%.