标签
This paper theoretically characterizes reward hacking in evaluator ensembles using covariance geometry, proving common-mode error is not identifiable from judge scores alone and bounding overstatement in best-of-K selection.