Tag
This article discusses research on measuring benchmark optimization in speech recognition, where some ASR models may optimize for test benchmarks rather than real-world performance, and introduces tests to quantify this phenomenon.
This paper quantifies how high-performing ASR models optimize for benchmarks in ways that inflate scores without improving real-world transcription, using behavioral probes to reveal benchmark-conditioned behaviors.