How much published AI research is wrong because of data leakage?
Summary
A Princeton study found data leakage in nearly 300 AI papers across 17 fields, causing overoptimistic results. The author highlights how easy it is to accidentally leak data and cautions against trusting impressive AI claims without checking for leakage.
Similar Articles
The most important AI failure may be false confidence, not wrong answers
This article argues that the most dangerous AI failures stem not from wrong answers but from systems acting with false confidence based on incomplete data, outdated context, or bad assumptions, suggesting that AI evaluation should prioritize handling uncertainty over raw intelligence.
AI research tools are still too eager to turn public signals into certainty
The author critiques AI research tools for overconfidence in weak signals, praising Komo AI's rapid discovery and source-attached summaries but highlighting the need for better uncertainty and contradiction handling. They describe a workflow that splits discovery, verification, and structured checking across multiple AI tools.
Every AI Visibility Tool Is Lying to You
This article critically examines the accuracy of AI visibility tools that claim to measure brand presence in generative AI responses, arguing that they provide false precision due to nondeterminism, personalization, and scraping biases. It calls for transparency in methodology and warns against treating opaque dashboards as stable truth.
Researchers just found 28 fake AI citations in medical papers
Researchers found 28 AI-generated fake citations in medical papers that influence clinical guidelines, highlighting the risk of AI hallucinations undermining scientific integrity and patient care.
MosaicLeaks: Can your research agent keep a secret?
MosaicLeaks introduces a new benchmark for measuring privacy leakage in deep-research AI agents, showing that agents often leak private information through external queries and proposing a training method (PA-DR) to reduce leakage while improving task performance.