statistical-significance

Tag

Cards List
#statistical-significance

I had 55 LLMs blind-grade each other (22k judgments, all open). Every model family with enough data is biased toward its own siblings. Qwen judges favor Qwen by ~0.9 points. Mistral penalizes its own by ~1.0.

Reddit r/LocalLLaMA · 2026-06-28

An open evaluation setup with 55 LLMs blind-grading each other reveals statistically significant same-family rating bias across 8 model families, with Mistral penalizing its own models most severely. The study highlights issues with aggregate leaderboards and proposes improvements like within-response mixed-effects models.

0 favorites 0 likes
#statistical-significance

Few-Shot Resampling for Scalable Statistically-Sound Data Mining

arXiv cs.LG · 2026-06-11 Cached

Introduces FewRS, a resampling-based approach that drastically reduces the number of resampled datasets required for statistically-sound data mining, achieving up to two orders of magnitude speedup while maintaining rigorous false discovery control and high statistical power.

0 favorites 0 likes
← Back to home

Submit Feedback