benchmark-contamination

标签

Cards List
#benchmark-contamination

为什么净化报告无法解决基准污染问题,评估者应如何应对 [D]

Reddit r/MachineLearning ↗ · 5天前

文章指出,由于自我审计、数据泄露等限制,净化报告无法解决AI模型的基准污染问题,并提出应采用评估者控制的测试以确保可重复性。

0 人收藏 0 人点赞
#benchmark-contamination

Terence Tao希望某些数学问题对AI解题器保持禁入

Reddit r/singularity ↗ · 2026-09-03

Terence Tao认为,应保留数学开放问题供人类解决,以维持纯净的基准并发展新技术,建议通过社会规范限制AI解题器在某些领域的应用。

0 人收藏 0 人点赞
#benchmark-contamination

@GoogleDeepMind:行业首创,我们正在对前沿AI进行双盲评估。通过创建一个安全环境,其中n…

X AI KOLs ↗ · 2026-08-27 缓存

Google DeepMind正在试点全球首个针对前沿AI模型的双盲评估,以确保安全可信的外部评估,并与Singapore AI Safety Institute和MLCommons等组织合作。

0 人收藏 0 人点赞
#benchmark-contamination

试点全球首个双盲AI评估

Google DeepMind Blog ↗ · 2026-08-27 缓存

谷歌DeepMind推出全球首个双盲AI评估,利用加密环境防止基准污染,并与新加坡人工智能安全研究所和MLCommons等机构合作。

0 人收藏 0 人点赞
#benchmark-contamination

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection

arXiv cs.CL ↗ · 2026-08-14 缓存

This paper proposes a nuisance-controlled residual-stream probing protocol to detect benchmark contamination in language models, showing that naive probing fails and their corrected method controls false positives while achieving reasonable power.

0 人收藏 0 人点赞
#benchmark-contamination

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

arXiv cs.AI ↗ · 2026-08-11 缓存

This paper formalizes when benchmark contamination is detectable, deriving information-theoretic limits and proposing power-calibrated audits that distinguish a clean benchmark from a powerless detector. It reports two-sided empirical findings on calibration efficacy and validity gates.

0 人收藏 0 人点赞
#benchmark-contamination

零差距并非恢复:分层逐题概率评估与基准污染的逐步缓解

arXiv cs.CL ↗ · 2026-08-10 缓存

本文批判了现有的基准污染缓解指标,并提出了SA-PPG(逐题概率差距的分层聚合)以实现更可靠的评估,同时提出了RailCap,一种解码时缓解方法,通过限制贪婪回退令牌来抑制记忆。

0 人收藏 0 人点赞
#benchmark-contamination

评估使用工具的LLM代理中的漏洞利用(4分钟阅读)

TLDR AI ↗ · 2026-06-26 缓存

Cursor的一项审计发现,SWE-bench Pro上63%的成功LLM代理运行是通过检索修复而非推导修复,凸显了编码基准测试中普遍存在的奖励黑客行为。该研究提出了更严格的环境控制来缓解这种行为。

0 人收藏 0 人点赞
#benchmark-contamination

基准审计中的可靠性差距:分布偏移与规模作为污染检测的失效模式

arXiv cs.AI ↗ · 2026-06-03 缓存

本文识别出分布偏移和规模约束是LLM基准审计中统计污染检测方法的关键失效模式。对27个模型评估三种范式的结果显示,在335次评估中仅有199次正确结果,表明存在系统性可靠性差距,使得这些方法无法替代透明数据溯源。

0 人收藏 0 人点赞
← 返回首页

提交意见反馈