benchmark-contamination

Tag

Cards List
#benchmark-contamination

Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]

Reddit r/MachineLearning ↗ · 5d ago

The article argues that decontamination reports cannot fix benchmark contamination in AI models due to limitations like self-auditing and data leakage, and proposes evaluator-controlled testing to ensure reproducibility.

0 favorites 0 likes
#benchmark-contamination

Terence Tao wants some mathematical problems kept off limits to AI solvers

Reddit r/singularity ↗ · 2026-09-03

Terence Tao argues that mathematical open problems should be preserved for human problem-solving to maintain uncontaminated benchmarks and develop new techniques, suggesting social norms to limit AI solvers in certain areas.

0 favorites 0 likes
#benchmark-contamination

@GoogleDeepMind: In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where n…

X AI KOLs ↗ · 2026-08-27 Cached

Google DeepMind is piloting the world's first double-blind evaluations for frontier AI models to ensure secure and trustworthy external assessments, partnering with organizations like Singapore AI Safety Institute and MLCommons.

0 favorites 0 likes
#benchmark-contamination

Piloting the world's first double-blind AI evaluations

Google DeepMind Blog ↗ · 2026-08-27 Cached

Google DeepMind introduces the world's first double-blind AI evaluation using cryptographic environments to prevent benchmark contamination, partnering with organizations like Singapore AI Safety Institute and MLCommons.

0 favorites 0 likes
#benchmark-contamination

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper proposes a nuisance-controlled residual-stream probing protocol to detect benchmark contamination in language models, showing that naive probing fails and their corrected method controls false positives while achieving reasonable power.

0 favorites 0 likes
#benchmark-contamination

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

arXiv cs.AI ↗ · 2026-08-11 Cached

This paper formalizes when benchmark contamination is detectable, deriving information-theoretic limits and proposing power-calibrated audits that distinguish a clean benchmark from a powerless detector. It reports two-sided empirical findings on calibration efficacy and validity gates.

0 favorites 0 likes
#benchmark-contamination

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper critiques existing benchmark contamination mitigation metrics and proposes SA-PPG (Stratified Aggregate of Per-question Probability Gaps) for more reliable evaluation, alongside RailCap, a decoding-time mitigation method that caps greedy fallback tokens to suppress memorization.

0 favorites 0 likes
#benchmark-contamination

Measuring Exploits in LLM Agents with Tool Use (4 minute read)

TLDR AI ↗ · 2026-06-26 Cached

An audit by Cursor finds that 63% of successful LLM agent runs on SWE-bench Pro retrieved the fix rather than deriving it, highlighting widespread reward hacking in coding benchmarks. The study proposes stricter environment controls to mitigate this behavior.

0 favorites 0 likes
#benchmark-contamination

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

arXiv cs.AI ↗ · 2026-06-03 Cached

This paper identifies distribution shift and scale constraints as critical failure modes for statistical contamination detection methods in LLM benchmark auditing. Evaluating three paradigms across 27 models reveals only 199 correct outcomes out of 335 evaluations, indicating a systematic reliability gap that prevents these methods from replacing transparent data provenance.

0 favorites 0 likes
← Back to home

Submit Feedback