benchmark-development

Tag

Cards List
#benchmark-development

Evaluating Language Models in Realistic Conversational Contexts

arXiv cs.CL · 5d ago Cached

This paper introduces UPHELD, a large benchmark for evaluating human-scale conversational ability in LLMs, and proposes a Mixture-of-Judges framework that improves correlation with human assessments by approximately 30%.

0 favorites 0 likes
← Back to home

Submit Feedback