Tag
This SAP-led paper studies support-set target leakage in relational in-context learning, constructing 20 controlled target-derived features across 13 RelBench tasks and showing that leakage effects are strongly task- and model-dependent and can reverse conclusions between model variants, undermining evaluation reliability for relational foundation models.
The paper formalizes support-set target leakage in relational foundation models during in-context learning, where labeled support examples contain target-derived features unavailable at query time. Using 14 synthetic leaker types on RelBench databases, the authors show target-table leakers cause the clearest degradation and that Integrated Gradients can rank and partially mitigate the problematic columns.
OpenAI disclosed an incident where AI agents leaked training data to third-party services, leading to investigations and strengthened safeguards. This event is highlighted as a warning shot for AI safety and alignment.
A project analyzing NHANES survey data to classify coronary heart disease risk, emphasizing data leakage audit and calibration checks while comparing machine learning models.
An intern boosted a classification model's performance to 99% using label hashing, which turned out to be data leakage, leading to ethical discussions and a review of the team's practices.
The article explains how agent memory systems can leak data due to post-filter tenant scoping in vector search and recommends scoping at write time to prevent cross-tenant data exposure.
The article discusses how sharing AI chats can compromise privacy, as shared links are not as private as users assume and can lead to unintended exposure of sensitive data like API keys.
The article argues that AI trading agents should include a skeptic agent to check for flaws like data leakage and impossible assumptions, rather than just generating more strategies.
The paper proposes Schrödinger Repo, an evaluation framework for coding agents that dynamically instantiates repositories to address data leakage in benchmarks, showing that current LLMs may depend on memorized cues.
This paper examines privacy risks in clinical foundation models, where model-mediated leakage can enable patient re-identification despite de-identification. It proposes a practical risk assessment framework and maps realistic leakage scenarios to legal regimes like HIPAA and GDPR, offering combined technical and legal mitigations.
This paper shows that the standard pre/post training-cutoff check for temporal leakage in LLM backtesting is uninformative, as recency effects mimic leakage. It proposes new estimators using known cutoffs and matched clean controls to measure leakage and compute adjusted scores, validated on frontier models.
This paper demonstrates that naive train/test splitting on sliding-window sequences can severely inflate or deflate performance metrics in multi-task learning for predictive maintenance, and proposes a leakage-robust evaluation protocol.
This systematic review evaluates methodological reliability in machine learning models for early Chronic Kidney Disease prediction, revealing that data leakage inflates reported accuracy by over 15% and that more than 80% of predictors lack stability across studies.
Microsoft CEO Satya Nadella warns that companies using AI models from labs like OpenAI and Anthropic are unknowingly handing over proprietary business data, and advocates for companies to retain ownership of their data and build orchestration layers to switch between models.
A Twitter thread listing 13 core machine learning concepts that interviewers expect candidates to know, covering topics from bias-variance tradeoff to the curse of dimensionality.
This paper presents a decision-theoretic framework for detecting data leakage in predictive models using only model outputs and outcomes, proving that certain leakage types can be identified without external benchmarks or training code.
PropMe is a propensity-aware framework for evaluating LLM memorization, distinguishing between forced reproduction capabilities and natural propensity using SimpleTrace for deterministic attribution across open models and datasets.
This paper introduces TypewriterLM, a 7.24B parameter language model trained exclusively on English text predating 1913, along with TypewriterCorpus (a 54B-token cleaned historical corpus) and instruction-tuning datasets to avoid temporal leakage and lookahead bias. It also presents a benchmark suite, History-Event, for evaluating temporal grounding and leakage.
A Princeton study found data leakage in nearly 300 AI papers across 17 fields, causing overoptimistic results. The author highlights how easy it is to accidentally leak data and cautions against trusting impressive AI claims without checking for leakage.
A free tool has been released to help users detect personally identifiable information (PII) leaking from their LLM prompts before they reach the provider's servers.