Tag
The article argues that decontamination reports cannot fix benchmark contamination in AI models due to limitations like self-auditing and data leakage, and proposes evaluator-controlled testing to ensure reproducibility.
Terence Tao argues that mathematical open problems should be preserved for human problem-solving to maintain uncontaminated benchmarks and develop new techniques, suggesting social norms to limit AI solvers in certain areas.
Google DeepMind is piloting the world's first double-blind evaluations for frontier AI models to ensure secure and trustworthy external assessments, partnering with organizations like Singapore AI Safety Institute and MLCommons.
Google DeepMind introduces the world's first double-blind AI evaluation using cryptographic environments to prevent benchmark contamination, partnering with organizations like Singapore AI Safety Institute and MLCommons.
This paper proposes a nuisance-controlled residual-stream probing protocol to detect benchmark contamination in language models, showing that naive probing fails and their corrected method controls false positives while achieving reasonable power.
This paper formalizes when benchmark contamination is detectable, deriving information-theoretic limits and proposing power-calibrated audits that distinguish a clean benchmark from a powerless detector. It reports two-sided empirical findings on calibration efficacy and validity gates.
This paper critiques existing benchmark contamination mitigation metrics and proposes SA-PPG (Stratified Aggregate of Per-question Probability Gaps) for more reliable evaluation, alongside RailCap, a decoding-time mitigation method that caps greedy fallback tokens to suppress memorization.
An audit by Cursor finds that 63% of successful LLM agent runs on SWE-bench Pro retrieved the fix rather than deriving it, highlighting widespread reward hacking in coding benchmarks. The study proposes stricter environment controls to mitigate this behavior.
This paper identifies distribution shift and scale constraints as critical failure modes for statistical contamination detection methods in LLM benchmark auditing. Evaluating three paradigms across 27 models reveals only 199 correct outcomes out of 335 evaluations, indicating a systematic reliability gap that prevents these methods from replacing transparent data provenance.