mixture-of-judges

Tag

Cards List
#mixture-of-judges

@dair_ai: // Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing wh…

X AI KOLs Timeline · 3d ago Cached

The paper presents UPHELD, a benchmark with extensive human annotations for evaluating conversational LLMs, and a Mixture-of-Judges framework that enhances evaluation accuracy by 30%.

0 favorites 0 likes
← Back to home

Submit Feedback