data-leakage

Tag

Cards List
#data-leakage

Leakage-Robust Evaluation and Data-Scale Sensitivity of Attention-Enhanced Multi-Task Learning for Joint Fault Diagnosis and Remaining Useful Life Estimation

arXiv cs.LG · 2026-07-21 Cached

This paper demonstrates that naive train/test splitting on sliding-window sequences can severely inflate or deflate performance metrics in multi-task learning for predictive maintenance, and proposes a leakage-robust evaluation protocol.

0 favorites 0 likes
#data-leakage

Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability

arXiv cs.LG · 2026-07-15 Cached

This systematic review evaluates methodological reliability in machine learning models for early Chronic Kidney Disease prediction, revealing that data leakage inflates reported accuracy by over 15% and that more than 80% of predictors lack stability across studies.

0 favorites 0 likes
#data-leakage

Satya Nadella has issued a shocking warning to companies using AI

TechCrunch AI · 2026-07-13 Cached

Microsoft CEO Satya Nadella warns that companies using AI models from labs like OpenAI and Anthropic are unknowingly handing over proprietary business data, and advocates for companies to retain ownership of their data and build orchestration layers to switch between models.

0 favorites 0 likes
#data-leakage

@TensorTonic: 13 Core ML Concepts Every Interviewer Expects You to Know 1. Bias-Variance Tradeoff - The key framework for understandi…

X AI KOLs Timeline · 2026-06-26 Cached

A Twitter thread listing 13 core machine learning concepts that interviewers expect candidates to know, covering topics from bias-variance tradeoff to the curse of dimensionality.

0 favorites 0 likes
#data-leakage

A prior-free blind detection of information leakage from model predictions

arXiv cs.LG · 2026-06-11 Cached

This paper presents a decision-theoretic framework for detecting data leakage in predictive models using only model outputs and outcomes, proving that certain leakage types can be identified without external benchmarks or training code.

0 favorites 0 likes
#data-leakage

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

Hugging Face Daily Papers · 2026-06-04 Cached

PropMe is a propensity-aware framework for evaluating LLM memorization, distinguishing between forced reproduction capabilities and natural propensity using SimpleTrace for deterministic attribution across open models and datasets.

0 favorites 0 likes
#data-leakage

Pretraining Language Models on Historical Text

arXiv cs.CL · 2026-06-03 Cached

This paper introduces TypewriterLM, a 7.24B parameter language model trained exclusively on English text predating 1913, along with TypewriterCorpus (a 54B-token cleaned historical corpus) and instruction-tuning datasets to avoid temporal leakage and lookahead bias. It also presents a benchmark suite, History-Event, for evaluating temporal grounding and leakage.

0 favorites 0 likes
#data-leakage

How much published AI research is wrong because of data leakage?

Reddit r/artificial · 2026-06-01

A Princeton study found data leakage in nearly 300 AI papers across 17 fields, causing overoptimistic results. The author highlights how easy it is to accidentally leak data and cautions against trusting impressive AI claims without checking for leakage.

0 favorites 0 likes
#data-leakage

Created a free tool to check what PII your LLM prompts are leaking before they hit the provider

Reddit r/artificial · 2026-05-12

A free tool has been released to help users detect personally identifiable information (PII) leaking from their LLM prompts before they reach the provider's servers.

0 favorites 0 likes
#data-leakage

OpenGuardrails: An Open-Source Context-Aware AI Guardrails Platform

Papers with Code Trending · 2025-10-22 Cached

OpenGuardrails is an open-source platform for AI safety, offering context-aware content-safety and manipulation detection (e.g., prompt injection, jailbreaking) via a unified model, plus a separate NER pipeline for data-leakage identification. It achieves state-of-the-art performance on safety benchmarks and supports private, enterprise-grade deployment.

0 favorites 0 likes
#data-leakage

AI Agent Security - MIT 6.566 Computer Systems Security, Spring 2026

YouTube AI Channels · 2026-05-21 Cached

MIT 6.566 course lecture introduces security challenges for AI agents, including non-adversarial errors (e.g., accidental database deletion) and adversarial attacks (e.g., prompt injection, data leakage), and explains the basics of building systems from language models to conversational agents.

0 favorites 1 likes
← Back to home

Submit Feedback