Tag
This paper introduces BenchDrift, a method for quantifying how LLM benchmark performance changes when problems are rephrased without changing meaning or answer. It shows that rephrasing causes bidirectional correctness flips across models and benchmarks, with stronger models becoming more sensitive to phrasing.