Tag
This paper proposes a scalable, domain-agnostic framework for automated LLM evaluation that uses pairwise comparisons by multiple LLMs and an Elo rating system to approximate expert judgments, reducing the need for human intervention.
Uses the Bradley-Terry model and Elo rating system to statistically determine a dog's favorite treat through pairwise comparison experiments.