Tag
An open evaluation setup with 55 LLMs blind-grading each other reveals statistically significant same-family rating bias across 8 model families, with Mistral penalizing its own models most severely. The study highlights issues with aggregate leaderboards and proposes improvements like within-response mixed-effects models.
Introduces FewRS, a resampling-based approach that drastically reduces the number of resampled datasets required for statistically-sound data mining, achieving up to two orders of magnitude speedup while maintaining rigorous false discovery control and high statistical power.