Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]
Summary
The article argues that decontamination reports cannot fix benchmark contamination in AI models due to limitations like self-auditing and data leakage, and proposes evaluator-controlled testing to ensure reproducibility.
Similar Articles
Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D]
The article argues that training-side decontamination in AI models cannot be verified due to inherent trust and inspection issues, and proposes an evaluation-side rule to ensure reproducibility by controlling the evaluation process.
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
This paper formalizes when benchmark contamination is detectable, deriving information-theoretic limits and proposing power-calibrated audits that distinguish a clean benchmark from a powerless detector. It reports two-sided empirical findings on calibration efficacy and validity gates.
@charliermarsh: Lots of good examples in these reports. For example: in DeepSWE v1.1, the verifier discards the agent's changes to some…
Epoch AI introduces Benchmark Reviews to audit AI benchmarks, revealing flaws such as in DeepSWE v1.1 where the verifier discards changes without informing the model, leading to failures.
The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection
This paper identifies distribution shift and scale constraints as critical failure modes for statistical contamination detection methods in LLM benchmark auditing. Evaluating three paradigms across 27 models reveals only 199 correct outcomes out of 335 evaluations, indicating a systematic reliability gap that prevents these methods from replacing transparent data provenance.
Unsteady Metrics and Benchmarking Cultures of AI Model Builders
This paper introduces Benchmarking-Cultures-25, a dataset analyzing how AI model builders selectively highlight benchmarks in press releases. It finds a fragmented evaluation landscape with limited cross-model comparability, arguing that benchmarks are used as narrative devices for market positioning rather than standardized scientific measurement.