Tag
This paper investigates the sensitivity of LLM evaluation benchmarks to different harness configurations, finding that config-fragile items disproportionately affect performance gaps between models, and that the choice of harness can determine leaderboard rankings.