phrasing-sensitivity

Tag

Cards List
#phrasing-sensitivity

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

arXiv cs.CL · 2026-08-13 Cached

This paper introduces BenchDrift, a method for quantifying how LLM benchmark performance changes when problems are rephrased without changing meaning or answer. It shows that rephrasing causes bidirectional correctness flips across models and benchmarks, with stronger models becoming more sensitive to phrasing.

0 favorites 0 likes
← Back to home

Submit Feedback