A prior-free blind detection of information leakage from model predictions

arXiv cs.LG Papers

Summary

This paper presents a decision-theoretic framework for detecting data leakage in predictive models using only model outputs and outcomes, proving that certain leakage types can be identified without external benchmarks or training code.

arXiv:2606.11267v1 Announce Type: new Abstract: Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise. None operates on the artifact an auditor most often holds: the model's output. We ask what can be decided about leakage from predictions and outcomes alone. We give a decision-theoretic framework in which leakage diagnostics are functionals of the predicted-risk/outcome law, parameterized by a threshold-weighting linked to proper scoring rules and decision-curve analysis. We prove a sharp impossibility: a recalibrated leak matching an honest model's calibration and discrimination is indistinguishable from honest performance by \emph{any} function of the predictions, so the broad class is detectable only against an externally supplied ceiling on achievable discrimination. We then prove what leakage cannot hide: a near-deterministic subgroup -- the signature of a near-label leak -- produces a sustained unit-purity head that no legitimate predictor of a non-deterministic outcome can manufacture, yielding a prior-free test. These results organize leakage into a trichotomy -- miscalibrated, broad-calibrated, and deterministic -- each with a matched detector and failure mode. We validate on UK Biobank using time-windowed comorbidity leakage with known, graded severity, measuring a detection floor of $\Delta\cstar \approx 0.007$ on this endpoint, below which residual leakage is undetectable from output and too small to alter conclusions. The numerical floor is cohort- and endpoint-specific; the structural lesson is general: output-only detection fails where residual leakage is indistinguishable from an honestly stronger predictor. The test returns a verdict on a prediction vector in under a second on commodity hardware.
Original Article
View Cached Full Text

Cached at: 06/11/26, 01:45 PM

# A prior-free blind detection of information leakage from model predictions
Source: [https://arxiv.org/html/2606.11267](https://arxiv.org/html/2606.11267)
Laurence A\. Jacobs1,2,∗ 1Center for Molecular Cardiology, University of Zurich, Zurich, Switzerland 2Center for Complexity Sciences, National University of Mexico, Mexico City, Mexico ∗Correspondence:[laurence\.jacobs@uzh\.ch](https://arxiv.org/html/2606.11267v1/mailto:[email protected])

###### Abstract

Data leakage—the contamination of a predictive model with information unavailable at baseline—is the dominant reproducibility failure in machine\-learning\-based science, yet the detection tools in use require the training code, fresh external data, or domain expertise\. None operates on the artifact an auditor most often holds: the model’s output\. We ask what can be decided about leakage from a list of predictions and outcomes alone\. We give a decision\-theoretic framework in which leakage diagnostics are functionals of the predicted\-risk/outcome law, parameterized by a threshold\-weighting that places them in correspondence with proper scoring rules and decision\-curve analysis\. Within it we prove a sharp impossibility: a recalibrated leak that matches an honest model’s calibration and discrimination is indistinguishable from honest performance by*any*function of the predictions, so the broad class of leakage is detectable only against an externally supplied ceiling on achievable discrimination\. We then prove what leakage cannot hide: a near\-deterministic subgroup—the signature of a near\-label leak—produces a sustained unit\-purity head that no legitimate predictor of a non\-deterministic outcome can manufacture, yielding a prior\-free test\. These results organize leakage into a trichotomy—miscalibrated, broad\-calibrated, and deterministic—each with a matched detector and an explicit failure mode\. We validate the trichotomy on UK Biobank using time\-windowed comorbidity leakage with known, graded severity, measuring a detection floor ofΔ​C∗≈0\.007\\Delta C^\{\*\}\\approx 0\.007on this endpoint, below which the residual leakage is both undetectable from output and too small to alter conclusions\. The numerical floor depends on cohort, prevalence, endpoint, and leakage mechanism; the structural lesson is general: output\-only detection fails exactly where the residual leakage is indistinguishable from an honestly stronger predictor without an external benchmark\. The resulting test takes a prediction vector and returns a verdict in under a second on commodity hardware\.

## 1Introduction

Leakage occurs when a model is trained on information that will not be legitimately available when it is deployed, so that reported performance reflects a quantity the model cannot reproduce in use\[[12](https://arxiv.org/html/2606.11267#bib.bib1),[14](https://arxiv.org/html/2606.11267#bib.bib22)\]\. Recent surveys identify it as a pervasive and field\-spanning cause of irreproducible results, affecting hundreds of published studies across many disciplines, and show that once leakage is corrected, the apparent advantage of complex models over simple baselines often disappears\[[11](https://arxiv.org/html/2606.11267#bib.bib2)\]\. In one of the most extensively audited domains, of6262COVID\-19 imaging models reviewed byRobertset al\.\[[16](https://arxiv.org/html/2606.11267#bib.bib18)\], none was found suitable for clinical use, with leakage among the dominant failure modes\. Leakage is therefore not a niche modeling error but a systemic threat to the evidentiary value of predictive claims\.

The methods available to catch it operate upstream of the result\. Static analysis of the training pipeline detects mechanical leakage—train/test contamination, preprocessing fit before splitting, repeated test reuse\[[4](https://arxiv.org/html/2606.11267#bib.bib15)\]—but requires the source code\[[26](https://arxiv.org/html/2606.11267#bib.bib3)\]\. External validation exposes leakage as a failure to replicate but requires an independent cohort, which most studies never obtain\. Risk\-of\-bias and reporting instruments such as PROBAST and TRIPOD catch leakage through expert appraisal but are time\-intensive and demand both subject and methodological expertise\[[25](https://arxiv.org/html/2606.11267#bib.bib4),[10](https://arxiv.org/html/2606.11267#bib.bib5),[2](https://arxiv.org/html/2606.11267#bib.bib14),[1](https://arxiv.org/html/2606.11267#bib.bib19)\]\. The artifact that an editor, replicator, or internal auditor most often actually holds—the model’s predictions on an evaluation set, together with the realized outcomes—has no test of its own\.

We ask a precise question:*what can be decided about leakage from the pair\(p^,y\)\(\\hat\{p\},y\)of predicted risks and outcomes, with nothing else?*The answer is neither “everything” nor “nothing,” and the contribution of this paper is to draw the boundary exactly and to supply the detectors that live on the good side of it\.

#### Contributions\.

\(i\) A decision\-theoretic framework \(Section[2](https://arxiv.org/html/2606.11267#S2)\) in which output\-level leakage diagnostics are threshold\-weighted net\-benefit functionals, linked to proper scoring rules through a mixture representation\. \(ii\) An impossibility theorem \(Section[3](https://arxiv.org/html/2606.11267#S3)\): calibrated leakage matched in discrimination is invisible to any\(p^,y\)\(\\hat\{p\},y\)functional, so the broad class is detectable only relative to an external discrimination ceiling\. \(iii\) A prior\-free positive result \(Section[4](https://arxiv.org/html/2606.11267#S4)\): a sustained unit\-purity head certifies leakage under only the qualitative assumption that the outcome is not deterministic\. \(iv\) The resulting trichotomy \(Section[5](https://arxiv.org/html/2606.11267#S5)\)\. \(v\) Validation on UK Biobank time\-windowed leakage with a measured sensitivity floor \(Section[7](https://arxiv.org/html/2606.11267#S7)\)\. \(vi\) A deployable sub\-second test\.

## 2Framework

#### Setup\.

A model produces predicted risksp^i∈\[0,1\]\\hat\{p\}\_\{i\}\\in\[0,1\]fori=1,…,ni=1,\\dots,nwith binary outcomesyi∈\{0,1\}y\_\{i\}\\in\\\{0,1\\\}\. We treat\(p^i,yi\)\(\\hat\{p\}\_\{i\},y\_\{i\}\)as draws from a joint lawPPon\[0,1\]×\{0,1\}\[0,1\]\\times\\\{0,1\\\}with base rateπ=𝔼​\[y\]\\pi=\\mathbb\{E\}\[y\]\. A predictor is*calibrated*if𝔼​\[y∣p^\]=p^\\mathbb\{E\}\[y\\mid\\hat\{p\}\]=\\hat\{p\}; this corresponds to moderate calibration in the hierarchy ofVan Calsteret al\.\[[23](https://arxiv.org/html/2606.11267#bib.bib26),[22](https://arxiv.org/html/2606.11267#bib.bib20)\]\. Discrimination is reported asC∗C^\{\*\}111Throughout we writeC∗C^\{\*\}for Harrell’s concordance statistic\[[7](https://arxiv.org/html/2606.11267#bib.bib12)\], the probability that a randomly chosen case is ranked above a randomly chosen control\.\.

#### Net benefit and its threshold weighting\.

For a decision thresholdτ\\tau, the net benefit of acting on\{p^≥τ\}\\\{\\hat\{p\}\\geq\\tau\\\}isNB​\(τ\)=1n​∑i\[yi​𝟏​\{p^i≥τ\}−\(1−yi\)​𝟏​\{p^i≥τ\}​τ1−τ\]\\mathrm\{NB\}\(\\tau\)=\\frac\{1\}\{n\}\\sum\_\{i\}\\\!\\big\[y\_\{i\}\\mathbf\{1\}\\\{\\hat\{p\}\_\{i\}\\\!\\geq\\\!\\tau\\\}\-\(1\-y\_\{i\}\)\\mathbf\{1\}\\\{\\hat\{p\}\_\{i\}\\\!\\geq\\\!\\tau\\\}\\tfrac\{\\tau\}\{1\-\\tau\}\\big\]\[[24](https://arxiv.org/html/2606.11267#bib.bib6)\], whereτ/\(1−τ\)\\tau/\(1\-\\tau\)is the harm\-to\-benefit exchange rate\. Following the expected\-net\-benefit framework\[[9](https://arxiv.org/html/2606.11267#bib.bib9)\], we weight net benefit over thresholds by a densityη\\etaon\(0,1\)\(0,1\),

ENBη=∫01NB​\(τ\)​η​\(τ\)​𝑑τ=1n​∑i\[yi​H​\(p^i\)−\(1−yi\)​G​\(p^i\)\],\\mathrm\{ENB\}\_\{\\eta\}=\\int\_\{0\}^\{1\}\\mathrm\{NB\}\(\\tau\)\\,\\eta\(\\tau\)\\,d\\tau=\\frac\{1\}\{n\}\\sum\_\{i\}\\big\[y\_\{i\}H\(\\hat\{p\}\_\{i\}\)\-\(1\-y\_\{i\}\)G\(\\hat\{p\}\_\{i\}\)\\big\],\(1\)withH​\(p\)=∫0pηH\(p\)=\\int\_\{0\}^\{p\}\\etaandG​\(p\)=∫0pη​\(τ\)​τ1−τ​𝑑τG\(p\)=\\int\_\{0\}^\{p\}\\eta\(\\tau\)\\tfrac\{\\tau\}\{1\-\\tau\}\\,d\\tau\. By the Schervish mixture representation of proper scoring rules\[[17](https://arxiv.org/html/2606.11267#bib.bib7),[5](https://arxiv.org/html/2606.11267#bib.bib8),[6](https://arxiv.org/html/2606.11267#bib.bib10)\],η\\etais the mixing measure: eachη\\etaselects a proper score, and steering its mass towardτ→1\\tau\\\!\\to\\\!1produces a functional sensitive to the most stringent operating points\. Equation \([1](https://arxiv.org/html/2606.11267#S2.E1)\) is thus a*family*of probes, not a single number\. Throughout this paper we adopt the uniform defaultη≡1\\eta\\equiv 1, which suffices to exhibit the three regimes; applications with specific decision\-threshold ranges \(e\.g\. screening at lowτ\\tau, treatment selection at highτ\\tau\) admit detector variants specialized to that range without modifying the framework\. This representation matters because leakage need not be global; it may appear only in clinically relevant threshold regions or in the extreme\-risk head of the prediction distribution\.

#### Two leakage observables\.

Writingri=yi−p^ir\_\{i\}=y\_\{i\}\-\\hat\{p\}\_\{i\}andw​\(p\)=H​\(p\)\+G​\(p\)w\(p\)=H\(p\)\+G\(p\), \([1](https://arxiv.org/html/2606.11267#S2.E1)\) decomposes asENBη=B~η\+1n​∑iri​w​\(p^i\)\\mathrm\{ENB\}\_\{\\eta\}=\\widetilde\{B\}\_\{\\eta\}\+\\frac\{1\}\{n\}\\sum\_\{i\}r\_\{i\}w\(\\hat\{p\}\_\{i\}\), whereB~η\\widetilde\{B\}\_\{\\eta\}depends onp^\\hat\{p\}alone\. The weighted residual yields a*dispersion*statistic testing the calibrated nully∣p^∼Bernoulli​\(p^\)y\\mid\\hat\{p\}\\sim\\mathrm\{Bernoulli\}\(\\hat\{p\}\),

Vη=∑iw​\(p^i\)2​ri2∑iw​\(p^i\)2​p^i​\(1−p^i\),𝔼null​\[Vη\]=1,Vη=1\+Op​\(n−1/2\)\.V\_\{\\eta\}=\\frac\{\\sum\_\{i\}w\(\\hat\{p\}\_\{i\}\)^\{2\}r\_\{i\}^\{2\}\}\{\\sum\_\{i\}w\(\\hat\{p\}\_\{i\}\)^\{2\}\\,\\hat\{p\}\_\{i\}\(1\-\\hat\{p\}\_\{i\}\)\},\\qquad\\mathbb\{E\}\_\{\\text\{null\}\}\[V\_\{\\eta\}\]=1,\\quad V\_\{\\eta\}=1\+O\_\{p\}\(n^\{\-1/2\}\)\.\(2\)Separately, ordering predictions descending, the cumulative*purity*ρ​\(k\)=k−1​∑i≤ky\(i\)\\rho\(k\)=k^\{\-1\}\\sum\_\{i\\leq k\}y\_\{\(i\)\}\(top\-kkevent rate\) summarizes the head of the risk distribution\. We distinguish its*breadth*\(the largest top fraction withρ−ρref\>ϵ\\rho\-\\rho\_\{\\text\{ref\}\}\>\\epsilon\) from its*spike*\(the largest top fraction with absoluteρ≥1−δ\\rho\\geq 1\-\\delta\), as Sections[3](https://arxiv.org/html/2606.11267#S3)and[4](https://arxiv.org/html/2606.11267#S4)show these measure different things\.

## 3The impossibility result

Leakage is defined by*legitimacy*: a feature is illegitimate if it carries information aboutyythat was not available at the prediction time\[[12](https://arxiv.org/html/2606.11267#bib.bib1)\]\. Legitimacy is a property of the temporal/causal structure of the data\-generating process—it is exogenous to the lawPPof\(p^,y\)\(\\hat\{p\},y\)\. This is the root of the limit\. We make the limit visible through two short lemmas; the impossibility proposition is then immediate\.

###### Lemma 1\(Determination\)\.

A calibrated lawPPon\[0,1\]×\{0,1\}\[0,1\]\\times\\\{0,1\\\}is determined by its score marginalF=Law​\(p^\)F=\\mathrm\{Law\}\(\\hat\{p\}\):

P​\(d​p,y=1\)=p​F​\(d​p\),P​\(d​p,y=0\)=\(1−p\)​F​\(d​p\),P\(dp,\\,y=1\)=p\\,F\(dp\),\\qquad P\(dp,\\,y=0\)=\(1\-p\)\\,F\(dp\),soNB​\(τ\)\\mathrm\{NB\}\(\\tau\),ENBη\\mathrm\{ENB\}\_\{\\eta\}, andC∗C^\{\*\}are functionals ofFFalone\.

###### Lemma 2\(Honest\-world realizability\)\.

For every distributionFFon\[0,1\]\[0,1\]there exists an*honest*predictor whose induced law equals the calibrated lawP​\(F\)P\(F\)of Lemma[1](https://arxiv.org/html/2606.11267#Thmlemma1)\. Concretely, letX∼FX\\sim Fbe a legitimately observed prediction\-time covariate, drawy∣X=p∼Bernoulli​\(p\)y\\mid X\{=\}p\\sim\\mathrm\{Bernoulli\}\(p\), and takep^h=X\\hat\{p\}^\{\\mathrm\{h\}\}=X\.

Lemma[2](https://arxiv.org/html/2606.11267#Thmlemma2)should not be read as claiming that every calibrated prediction law is achievable by an leaker\-free model in the same scientific problem\. It shows something sharper: the joint law of \(p^,y\\hat\{p\},y\) contains no record of the information set from whichp^\\hat\{p\}was generated\. Therefore, without external knowledge of the admissible prediction\-timeσ\\sigma\-algebra, the same output law is compatible with both a legitimate and an illegitimate generating mechanism\.

###### Proposition 1\(Invisibility of calibrated, matched leakage\)\.

Let an honest procedure and a leaky procedure induce lawsPhP\_\{\\mathrm\{h\}\}andPlP\_\{\\mathrm\{l\}\}on\[0,1\]×\{0,1\}\[0,1\]\\times\\\{0,1\\\}\. Every output\-level diagnostic is a functionalT​\(P\)T\(P\), so ifPh=PlP\_\{\\mathrm\{h\}\}=P\_\{\\mathrm\{l\}\}no test based on\(p^,y\)\(\\hat\{p\},y\)can separate them\. By Lemma[1](https://arxiv.org/html/2606.11267#Thmlemma1), two calibrated predictors with the same score marginal sharePPand are thus output\-indistinguishable\. By Lemma[2](https://arxiv.org/html/2606.11267#Thmlemma2), every calibrated leaky law is the law of some honest predictor; calibrated, marginal\-matched leakage is therefore invisible to any function of the predictions\.

Proofs of Lemmas[1](https://arxiv.org/html/2606.11267#Thmlemma1),[2](https://arxiv.org/html/2606.11267#Thmlemma2)and the proposition are in Appendix[A](https://arxiv.org/html/2606.11267#A1)\.

###### Corollary 1\(The broad class needs a prior\)\.

A calibrated leak whose only effect is to raiseC∗C^\{\*\}cannot be flagged from output alone, because its law coincides with that of a genuinely superior honest model at the sameC∗C^\{\*\}\. Detecting it requires an exogenous boundCmax∗C^\{\*\}\_\{\\max\}on the discrimination achievable without leakage —an outcome\-specific prior\.

Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)is the floor under all output\-level detection\. It is also liberating: it tells us precisely which leakage*cannot*hide, namely leakage that pushesPPoutside the class of legitimately achievable laws in a way certifiable without knowing that class exactly\. The output\-only regime is the same one in which membership\-inference attacks operate\[[18](https://arxiv.org/html/2606.11267#bib.bib21)\]: a closely related access regime, but opposite goal—we audit the procedure rather than the training set\.

## 4What leakage cannot hide

###### Lemma 3\(Purity ceiling\)\.

Suppose the outcome is not prediction\-time\-deterministic: there is no admissible eventSSwithℙ​\(y=1∣S\)=1\\mathbb\{P\}\(y=1\\mid S\)=1on a non\-null set\. Then for every legitimate predictor the top\-kkpurity satisfiesρ​\(k\)<1\\rho\(k\)<1for everykkspanning a non\-null fraction, and a sustained unit\-purity head of non\-null width is impossible\. Consequently an observed unit\-purity plateau over a non\-null top fraction certifies the use of information under which the outcome is \(near\-\)deterministic—information not legitimately available at prediction time—under the single qualitative prior that the outcome is not deterministic\.

###### Proof sketch\.

A legitimate calibrated predictor’s top\-kkpurity converges to the average of the true risk over the selected set, which is bounded above by the supremum of the admissible conditional risk; non\-determinism makes this supremum<1<1on any non\-null set\. A near\-label leak places a non\-null subset at conditional risk→1\\to 1, breaking the bound\. Appendix[A](https://arxiv.org/html/2606.11267#A1)\. ∎

Lemma[3](https://arxiv.org/html/2606.11267#Thmlemma3)is the prior\-free positive result: the*spike*statistic of Section[2](https://arxiv.org/html/2606.11267#S2), unlike breadth, escapes Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)because the unit\-purity law lies outside the honest class for any non\-deterministic outcome\. Separately, miscalibration\-inducing leakage leavesVη≠1V\_\{\\eta\}\\neq 1and is detectable prior\-free via \([2](https://arxiv.org/html/2606.11267#S2.E2)\); but a leaker who recalibrates returnsVη→1V\_\{\\eta\}\\to 1, so this signal, while free, is evadable\.

## 5The trichotomy

Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)and Lemma[3](https://arxiv.org/html/2606.11267#Thmlemma3)partition leakage by what it does toPPand therefore by what can detect it \(Table[1](https://arxiv.org/html/2606.11267#S5.T1)\)\.

Table 1:The leakage trichotomy: each regime, its matched detector, and the prior it requires\. The fourth row is Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)in force\.The fourth row coincides empirically with the regime in which leakage is too small to change conclusions \(Section[7](https://arxiv.org/html/2606.11267#S7)\), so the limit of detectability and the limit of consequence arrive together\.

###### Corollary 2\(Closure under recalibration\)\.

Post\-hoc monotone recalibration of the predictions can move leakage between regimes of Table[1](https://arxiv.org/html/2606.11267#S5.T1)but cannot exit all three\. The dispersion signal is destroyed \(zV→0z\_\{V\}\\to 0\), but the ranking—henceC∗C^\{\*\}and the unit\-purity head—is preserved\. A leak that improvedC∗C^\{\*\}remains caught by the ceiling trigger of Corollary[1](https://arxiv.org/html/2606.11267#Thmcorollary1); a leak that did not is by Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)output\-indistinguishable from clean, and is by construction without consequence at the population level\.

## 6Methods

#### Synthetic construction\.

To exercise each regime at controlled severity we generate a calibrated honest riskp^=σ​\(s\)\\hat\{p\}=\\sigma\(s\),s∼𝒩s\\sim\\mathcal\{N\}, withy∼Bernoulli​\(p^\)y\\sim\\mathrm\{Bernoulli\}\(\\hat\{p\}\), tuning the spread to a targetC∗C^\{\*\}\(n=200,000n=200\{,\}000,π=0\.02\\pi=0\.02, targetC∗≈0\.84C^\{\*\}\\approx 0\.84\)\. We inject \(a\) miscalibrated leakage by applying a monotone logit\-amplificationp^mis=σ​\(β⋅logit​\(p^\)\)\\hat\{p\}\_\{\\text\{mis\}\}=\\sigma\(\\beta\\cdot\\mathrm\{logit\}\(\\hat\{p\}\)\)withβ=2\.5\\beta=2\.5, which preserves the ranking \(henceC∗C^\{\*\}\) exactly while destroying calibration; \(b\) broad calibrated leakage as the true posterior under a noisy proxyz∣yz\\mid y, isotonically recalibrated; \(c\) deterministic leakage as a calibrated near\-label flag on0\.3%0\.3\\%of the cohort\. All arms share a commonC∗≈0\.84C^\{\*\}\\approx 0\.84–0\.870\.87so discrimination cannot explain any difference\. A fifth arm applies Platt scaling\[[15](https://arxiv.org/html/2606.11267#bib.bib17)\]to the miscalibrated predictions to test evadability of the dispersion signal\.

#### UK Biobank cohort\.

We use UK Biobank\[[20](https://arxiv.org/html/2606.11267#bib.bib16)\]\(Application 596880;n=501,883n=501\{,\}883\) with the incident delirium endpoint \(ICD\-10 F05\.x; prevalence≈0\.018\\approx 0\.018\)\. Leakage is introduced in graded, clinically interpretable form: eleven comorbidity flags are allowed to be populated from progressively wider windows of hospital\-episode data after baseline \(\+1,\+2,\+3,\+4,\+5,\+7,\+10\+1,\+2,\+3,\+4,\+5,\+7,\+10years, and full follow\-up\), so that each window yields an out\-of\-fold \(OOF\) prediction vector with a known, monotone increase in leakage\. Base models areℓ1\\ell\_\{1\}\-penalized logistic regression\[[21](https://arxiv.org/html/2606.11267#bib.bib13)\]with 64 features \(53 clinical biomarkers and 11 comorbidity flags\), five\-fold stratified cross\-validation, and penaltyC=0\.002C=0\.002; predictions are out\-of\-fold\.

#### Detector and thresholds\.

Detection operates on\(p^,y\)\(\\hat\{p\},y\)only and is independent of model fitting\. We reportC∗C^\{\*\},VηV\_\{\\eta\}\([2](https://arxiv.org/html/2606.11267#S2.E2)\), breadth, and spike\. The deployable test \(Algorithm[1](https://arxiv.org/html/2606.11267#alg1)\) returnsleakyif a unit\-purity head of at leastkmink\_\{\\min\}cases and non\-null width is present \(Lemma[3](https://arxiv.org/html/2606.11267#Thmlemma3)\), or ifC∗C^\{\*\}exceeds a suppliedCmax∗C^\{\*\}\_\{\\max\}\(Corollary[1](https://arxiv.org/html/2606.11267#Thmcorollary1)\); a dispersion anomaly is reported as a soft warning; otherwiseclean, with the stated scope that smooth sub\-threshold leakage cannot be excluded\.

Algorithm 1Blind leakage test: prediction vector→\\toverdict\.1:predictions

p^\[1:n\]\\hat\{p\}\[1\{:\}n\], outcomes

y\[1:n\]y\[1\{:\}n\]; optional ceiling

Cmax∗C^\{\*\}\_\{\\max\}
2:Parameters:

kmink\_\{\\min\},

δ\\delta,

ϵ\\epsilon,

zαz\_\{\\alpha\}, weight

η\\eta
3:verdict

∈\{leaky,clean\}\\in\\\{\\textsc\{leaky\},\\textsc\{clean\}\\\}with a soft dispersion flag

4:sort indices by descending

p^\\hat\{p\}
5:

ρ​\(k\)←k−1​∑i≤ky\(i\)\\rho\(k\)\\leftarrow k^\{\-1\}\\sum\_\{i\\leq k\}y\_\{\(i\)\}⊳\\trianglerightcumulative top\-kkpurity

6:

spike←max⁡\{k:ρ​\(k\)≥1−δ\}\\mathrm\{spike\}\\leftarrow\\max\\\{k:\\rho\(k\)\\geq 1\-\\delta\\\}⊳\\trianglerightunit\-purity head width

7:compute

C∗C^\{\*\},

VηV\_\{\\eta\},

zVz\_\{V\}\([2](https://arxiv.org/html/2606.11267#S2.E2)\), and

breadth\\mathrm\{breadth\}at excess

ϵ\\epsilon
8:

warn←\(\|zV\|\>zα\)\\mathrm\{warn\}\\leftarrow\(\\lvert z\_\{V\}\\rvert\>z\_\{\\alpha\}\)⊳\\trianglerightsoft, recalibration\-evadable

9:if

spike≥kmin\\mathrm\{spike\}\\geq k\_\{\\min\}then

10:return

\(leaky,warn\)\(\\textsc\{leaky\},\\,\\mathrm\{warn\}\)⊳\\trianglerightnear\-label \(Lem\.[3](https://arxiv.org/html/2606.11267#Thmlemma3)\)

11:endif

12:if

Cmax∗C^\{\*\}\_\{\\max\}givenand

C∗\>Cmax∗C^\{\*\}\>C^\{\*\}\_\{\\max\}then

13:return

\(leaky,warn\)\(\\textsc\{leaky\},\\,\\mathrm\{warn\}\)⊳\\trianglerightbroad \(Cor\.[1](https://arxiv.org/html/2606.11267#Thmcorollary1)\)

14:endif

15:return

\(clean,warn\)\(\\textsc\{clean\},\\,\\mathrm\{warn\}\)⊳\\trianglerightsmooth sub\-threshold not excluded

#### Choice ofkmink\_\{\\min\}\.

The Remark following Lemma[3](https://arxiv.org/html/2606.11267#Thmlemma3)bounds the legitimate probability of a unit\-purity head of widthkkbyMkM^\{k\}, whereM=ess​sup⁡μM=\\operatorname\*\{ess\\,sup\}\\muis an assumed bound on the admissible conditional risk\. Inverting for a target false\-certification rateα\\alphagives

kmin≥⌈log⁡α/log⁡M⌉\.k\_\{\\min\}\\;\\geq\\;\\lceil\\log\\alpha/\\log M\\rceil\.Guaranteeingα=10−3\\alpha=10^\{\-3\}even under a pessimisticM=0\.9M=0\.9would requirekmin≥66k\_\{\\min\}\\geq 66; we adopt the lighterkmin=50k\_\{\\min\}=50, which yields a false\-certification budget of0\.950≈5×10−30\.9^\{50\}\\\!\\approx\\\!5\\\!\\times\\\!10^\{\-3\}under that conservativeM=0\.9M=0\.9and≈10−15\\approx\\\!10^\{\-15\}under a typicalM=0\.5M=0\.5\. The remaining thresholds are set to spike purityρ≥0\.95\\rho\\geq 0\.95\(admitting a small near\-deterministic fractionδ=0\.05\\delta=0\.05to avoid sensitivity to single mislabeled events\), breadth excessϵ=0\.01\\epsilon=0\.01\(one percentage\-point separation from the honest purity curve\), and uniformη≡1\\eta\\equiv 1\.

## 7Results

### 7\.1Synthetic validation of the trichotomy

At matchedC∗=0\.842C^\{\*\}=0\.842, the dispersion statistic cleanly separates miscalibrated leakage from clean performance \(zV=291z\_\{V\}=291vs\.−0\.5\-0\.5\), while the calibrated broad proxy is invisible toVηV\_\{\\eta\}\(zV≈0\.0z\_\{V\}\\approx 0\.0\) and has net benefit equal to clean at every threshold—confirming Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)— and is recovered only when an outcome ceiling is supplied\. The deterministic arm is caught by unit purity \(ρ​\(0\.1%\)=1\.000\\rho\(0\.1\\%\)=1\.000\) regardless of calibration\. The monotone\-transform construction of the miscalibrated arm guarantees identicalC∗C^\{\*\}to the clean baseline \(0\.8420\.842for both\), so the dispersion signal cannot be attributed to improved discrimination\. Applying Platt scaling to the miscalibrated arm restoreszV=−0\.1z\_\{V\}=\-0\.1—indistinguishable from clean—while preservingC∗=0\.842C^\{\*\}=0\.842and the purity profile exactly \(Figure[1](https://arxiv.org/html/2606.11267#S7.F1), dashed red\), confirming that the dispersion signal is evadable by post\-hoc recalibration\.

![Refer to caption](https://arxiv.org/html/2606.11267v1/synthetic_trichotomy.png)Figure 1:Synthetic validation of the trichotomy at matchedC∗≈0\.84C^\{\*\}\\approx 0\.84,π=0\.02\\pi=0\.02\.A:Calibration catches the miscalibrated arm \(red solid\) but not the laundered version \(red dashed, Platt\-scaled back to the diagonal\); the broad calibrated \(orange\) and deterministic \(green\) arms are also indistinguishable from clean\.B:Dispersion\|zV\|\|z\_\{V\}\|on log scale: miscalibrated atzV=291z\_\{V\}=291, but Platt scaling restoreszV≈0z\_\{V\}\\approx 0—the signal is evadable\.C:Only the purity head catches the deterministic arm \(ρ=1\.0\\rho=1\.0at top0\.3%0\.3\\%\); the laundered arm \(dashed\) overlaps clean exactly\.
### 7\.2Graded leakage in UK Biobank and the detection floor

The numerical floor measured here is specific to incident delirium atπ≈0\.018\\pi\\approx 0\.018in UKB underℓ1\\ell\_\{1\}\-penalized logistic regression with the feature set of Section[6](https://arxiv.org/html/2606.11267#S6); the structural claim—that the output\-level detection floor coincides with the conclusion\-altering floor—is general and follows from Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)\. With this scope fixed, breadth rises monotonically with leakage and switches on at a sharp threshold \(Table[2](https://arxiv.org/html/2606.11267#S7.T2)\): zero through\+2\+2years, then first nonzero at\+3\+3years\. The detection floor isΔ​C∗≈0\.007\\Delta C^\{\*\}\\approx 0\.007–0\.0080\.008; below it, leakage leaves no output signature*and*is too small to materially change performance\. Figure[2](https://arxiv.org/html/2606.11267#S7.F2)shows the operating characteristic\.

Table 2:UK Biobank time\-windowed leakage \(incident delirium, ICD\-10 F05\.x\)\. Breadth is the top fraction with purity excess over honest exceeding0\.010\.01\.Δ​C∗\\Delta C^\{\*\}is computed from unroundedC∗C^\{\*\}\.
### 7\.3The floor is intrinsic, not a regularization artifact

A control sweeping the penalty across four orders of magnitude \(C∈\[5×10−4,1\]C\\in\[5\\times 10^\{\-4\},1\]\) leaves the\+2\+2y model with zero breadth andΔ​C∗≈0\.005\\Delta C^\{\*\}\\approx 0\.005–0\.0060\.006throughout: removing the penalty does not unmask a hidden signal, because none exists\. All eleven leaked comorbidity flags carry nonzero coefficients at\+2\+2y across all penalty levels, confirming that the leak enters the predictions; the floor is intrinsic to the information structure, consistent with Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1), not an artifact of regularization\.

### 7\.4The deterministic regime

Isolating the dementia comorbidity flag as a near\-label for incident delirium confirms the prior\-free spike\. At full follow\-up the flag coincides with the outcome \(flagged fraction1\.78%≈π1\.78\\%\\approx\\pi\), givingC∗=1\.000C^\{\*\}=1\.000and a unit\-purity head\. The instructive case is intermediate: at\+5\+5years a deterministic head of only0\.16%0\.16\\%of the cohort is detected prior\-free—a leak that raisesC∗C^\{\*\}by just0\.0240\.024and that a global metric would absorb as ordinary improvement, but that the unit\-purity test isolates\. We also note that strong regularization can launder a near\-label into a broad lift, masking the spike; the prior\-free guarantee assumes the model was permitted to express what it learned \(Section[8](https://arxiv.org/html/2606.11267#S8)\)\.

### 7\.5False\-positive rate on honest models

We apply the deployable test to ten leakage\-free constructed models spanning cardiometabolic, respiratory, neurological, oncologic, and mortality domains, with no discrimination ceiling supplied so that only the prior\-free spike detector \(Lemma[3](https://arxiv.org/html/2606.11267#Thmlemma3)\) is active\. The cohort and feature set match Section[6](https://arxiv.org/html/2606.11267#S6); predictions are out\-of\-fold\. All ten endpoints returnclean: no unit\-purity head of width≥kmin=50\\geq k\_\{\\min\}=50is present at theρ≥0\.95\\rho\\geq 0\.95purity threshold in any vector, despite a broad spread of discrimination \(C∗C^\{\*\}ranging from 0\.63 to 0\.87\)\. The prior\-free false\-positive rate is0/100/10\(Table[3](https://arxiv.org/html/2606.11267#S7.T3)\)\.

Table 3:Blind leakage test applied to 10 leakage\-free constructed models spanning cardiometabolic, respiratory, neurological, oncologic, and mortality domains\. The test operates on out\-of\-fold predictions with no discrimination ceiling supplied, so only the prior\-free spike detector \(Lemma[3](https://arxiv.org/html/2606.11267#Thmlemma3)\) is active\. All 10 endpoints returnclean; false positive rate = 0/10\.CodeDomainEventsC∗C^\{\*\}Spike@95%BreadthVerdictDM2Cardiometabolic28,4430\.8680—cleanCHFCardiometabolic8,2710\.7610—cleanCKDCardiometabolic8,4380\.7520—cleanCHDCardiometabolic24,2090\.6900—cleanCOPRespiratory10,3320\.8480—cleanDEMNeurological5,5680\.7700—cleanCANlungOncologic4,8200\.7870—cleanCANclrcOncologic6,7860\.6300—cleanCANprostOncologic12,0810\.6580—cleanACMMortality55,0230\.7120—cleanFalse positive rate0/10
### 7\.6Deployable test

Because detection is a functional of\(p^,y\)\(\\hat\{p\},y\)and requires no model fitting, the full battery runs in1\.31\.3s on the≈500\\approx\\\!500k\-row cohort on commodity hardware\. On the four synthetic regimes it returns the expected verdicts:cleanfor honest and for a strong calibrated model with no ceiling supplied,leakyfor the near\-label prior\-free, andleakyfor the broad model once a ceiling is supplied\. On the honest\-baseline panel of Section[7\.5](https://arxiv.org/html/2606.11267#S7.SS5)the prior\-free false\-positive rate is0/100/10\.

![Refer to caption](https://arxiv.org/html/2606.11267v1/leakage_oc.png)Figure 2:Operating characteristic: detection \(breadth / purity excess\) versusΔ​C∗\\Delta C^\{\*\}, with the floor atΔ​C∗≈0\.007\\Delta C^\{\*\}\\approx 0\.007\.

## 8Discussion

The detection literature splits into code\-level static analysis and expert checklists, with nothing operating on the model’s output\. This work supplies that missing layer and bounds it exactly\. The impossibility result is not a weakness to be apologized for but the spine of the method: it tells a user precisely when each detector is informative and when it is powerless\.

#### In\-house versus reviewer deployment\.

Corollary[1](https://arxiv.org/html/2606.11267#Thmcorollary1)’s prior dependence is a constraint for a reviewer handed a single vector, but evaporates in a setting with honest baselines on record: passingCmax∗C^\{\*\}\_\{\\max\}equal to one’s own honest model plus a margin makes the broad trigger automatic\. The reviewer\-facing version is the one genuinely bounded by Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)\.

#### A standing screen\.

Because the cost of detection is negligible, leakage screening need not be a special investigation; it can be a default column in a model registry, run on every endpoint as a matter of course\.

#### Recalibration and the trichotomy\.

Corollary[2](https://arxiv.org/html/2606.11267#Thmcorollary2)shows the trichotomy is closed under monotone post\-processing: a leaker who applies Platt scaling\[[15](https://arxiv.org/html/2606.11267#bib.bib17)\]moves from the miscalibrated regime to the broad calibrated regime, destroying the dispersion signal but leaving the ranking—and henceC∗C^\{\*\}and the unit\-purity head—in place \(Figure[1](https://arxiv.org/html/2606.11267#S7.F1), dashed red\)\. The ceiling and spike triggers therefore remain operative on recalibrated leakage; only a leak that does not improve discrimination and is recalibrated is invisible, and is by Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)also without consequence\.

## 9Limitations

Smooth, calibrated, sub\-threshold leakage is invisible to any output\-level test \(Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)\); we prove this rather than work around it, and show it coincides with the inconsequential regime\. The broad trigger depends on an outcome ceiling\. The spike guarantee assumes the outcome is not prediction\-time\-deterministic and that the model expresses the leaked information rather than shrinking it away\. The empirical validation uses a single cohort \(UK Biobank\) with a specific endpoint and leakage mechanism; generalization to other data structures \(e\.g\., image\-derived features, time\-series\) remains to be assessed\.

## Acknowledgments

This research used the UK Biobank Resource under Application Number 596880\. Data are available to approved researchers via[https://www\.ukbiobank\.ac\.uk](https://www.ukbiobank.ac.uk/)\[[20](https://arxiv.org/html/2606.11267#bib.bib16)\]\. Top 25 ranked protein lists for both endpoints are provided in Supplementary Tables S1 and S2\. We thank all participants and the UK Biobank team for making this resource available\.

## Appendix AProofs

We work with the population lawPPof\(p^,y\)\(\\hat\{p\},y\)on\[0,1\]×\{0,1\}\[0,1\]\\times\\\{0,1\\\}and writeFFfor its score marginal \(the law ofp^\\hat\{p\}\), soπ=𝔼​\[y\]=∫01p​F​\(d​p\)\\pi=\\mathbb\{E\}\[y\]=\\int\_\{0\}^\{1\}p\\,F\(dp\)\.

###### Proof of Lemma[1](https://arxiv.org/html/2606.11267#Thmlemma1)\(Determination\)\.

Calibration givesℙ​\(y=1∣p^=p\)=p\\mathbb\{P\}\(y=1\\mid\\hat\{p\}=p\)=pforFF\-a\.e\.pp, which is the displayed disintegration; hencePPis a measurable image ofFF\. The population net benefit isNB​\(τ\)=∫\[τ,1\]\(p−τ1−τ​\(1−p\)\)​F​\(d​p\)\\mathrm\{NB\}\(\\tau\)=\\int\_\{\[\\tau,1\]\}\\big\(p\-\\tfrac\{\\tau\}\{1\-\\tau\}\(1\-p\)\\big\)F\(dp\)andENBη=∫01\(p​H​\(p\)−\(1−p\)​G​\(p\)\)​F​\(d​p\)\\mathrm\{ENB\}\_\{\\eta\}=\\int\_\{0\}^\{1\}\\\!\\big\(pH\(p\)\-\(1\-p\)G\(p\)\\big\)F\(dp\), both functionals ofFF\. Under calibration the case and control scores have lawsF1​\(d​p\)=p​F​\(d​p\)/πF\_\{1\}\(dp\)=pF\(dp\)/\\piandF0​\(d​p\)=\(1−p\)​F​\(d​p\)/\(1−π\)F\_\{0\}\(dp\)=\(1\-p\)F\(dp\)/\(1\-\\pi\), so

C∗=1π​\(1−π\)​∬p\>qp​\(1−q\)​F​\(d​p\)​F​\(d​q\)\+12​\(ties\),C^\{\*\}=\\frac\{1\}\{\\pi\(1\-\\pi\)\}\\iint\_\{p\>q\}p\\,\(1\-q\)\\,F\(dp\)\\,F\(dq\)\+\\tfrac\{1\}\{2\}\(\\text\{ties\}\),again a functional ofFF\[[3](https://arxiv.org/html/2606.11267#bib.bib11),[17](https://arxiv.org/html/2606.11267#bib.bib7)\]\. ∎

###### Proof of Lemma[2](https://arxiv.org/html/2606.11267#Thmlemma2)\(Realizability\)\.

LetXXbe a covariate, defined on a richer space than the sample, with lawLaw​\(X\)=F\\mathrm\{Law\}\(X\)=Fon\[0,1\]\[0,1\]; such a probability space always exists\. Generate the outcomeyyconditionally onXXbyℙ​\(y=1∣X=p\)=p\\mathbb\{P\}\(y=1\\mid X=p\)=p, and take the predictorp^h=X\\hat\{p\}^\{\\mathrm\{h\}\}=X\. The predictor is a function of the prediction\-time covariateXXalone, hence legitimate\. Its induced joint law on\[0,1\]×\{0,1\}\[0,1\]\\times\\\{0,1\\\}is the calibrated law of marginalFFby Lemma[1](https://arxiv.org/html/2606.11267#Thmlemma1):Ph=P​\(F\)P\_\{\\mathrm\{h\}\}=P\(F\)\. Moreoverp^h=ℙ​\(y=1∣X\)\\hat\{p\}^\{\\mathrm\{h\}\}=\\mathbb\{P\}\(y=1\\mid X\)is the Bayes risk relative to theσ\\sigma\-algebra it generates, so among legitimate predictors with prediction\-timeσ\\sigma\-algebraσ​\(X\)\\sigma\(X\)none has higher concordance\. ∎

###### Proof of Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)\.

*\(i\) No\(p^,y\)\(\\hat\{p\},y\)\-test separates equal laws\.*Every output\-level diagnostic is a \(possibly randomized\) statistic of the sample\{\(p^i,yi\)\}\\\{\(\\hat\{p\}\_\{i\},y\_\{i\}\)\\\}, whose sampling law is determined byPP\. IfPh=PlP\_\{\\mathrm\{h\}\}=P\_\{\\mathrm\{l\}\}the two procedures generate identically distributed samples, so every statistic has the same law under both and no test, of any size or power, can behave differently on them\.

*\(ii\) Matched leaky and honest laws coincide\.*Suppose the leaky procedure is calibrated with marginalFFand concordancec=C∗​\(F\)c=C^\{\*\}\(F\)\. By Lemma[2](https://arxiv.org/html/2606.11267#Thmlemma2)there is an honest predictor with the same induced lawP​\(F\)P\(F\)and the same concordancecc\. By part \(i\) the two procedures are output\-indistinguishable, although one uses only admissible information and the other does not\.

The mechanism is now explicit: legitimacy is a property of whichσ\\sigma\-algebra generated the prediction—the information available at prediction time\[[12](https://arxiv.org/html/2606.11267#bib.bib1)\]—and is exogenous toPP\. Two procedures may sharePPyet differ in legitimacy, and no functional ofPPcan resolve that difference\. Finally, a raw leak that is miscalibrated has𝔼​\[y∣p^\]≠p^\\mathbb\{E\}\[y\\mid\\hat\{p\}\]\\neq\\hat\{p\}and is detectable through the dispersion route \(Appendix[B](https://arxiv.org/html/2606.11267#A2)\); but recalibration replaces it by a calibrated law with some marginalF′F^\{\\prime\}, which by Lemma[1](https://arxiv.org/html/2606.11267#Thmlemma1)equals the honest lawP​\(F′\)P\(F^\{\\prime\}\)\. Recalibration always maps a leak into the honest class as a law, giving Corollary[1](https://arxiv.org/html/2606.11267#Thmcorollary1): a calibrated leak whose only effect is to raiseC∗C^\{\*\}toc′c^\{\\prime\}coincides with the honest model of marginalF′F^\{\\prime\}at the samec′c^\{\\prime\}, so flagging it requires an exogenous ceilingCmax∗C^\{\*\}\_\{\\max\}\. ∎

For the purity ceiling, fix the sub\-σ\\sigma\-algebra𝒢\\mathcal\{G\}of information available at prediction time and letμ=ℙ​\(y=1∣𝒢\)=𝔼​\[y∣𝒢\]\\mu=\\mathbb\{P\}\(y=1\\mid\\mathcal\{G\}\)=\\mathbb\{E\}\[y\\mid\\mathcal\{G\}\]be the prediction\-time risk; a predictor is legitimate iff it is𝒢\\mathcal\{G\}\-measurable\. For a top selection of population massu∈\(0,1\]u\\in\(0,1\], letρ⋆​\(u\)\\rho^\{\\star\}\(u\)be the largest achievable top\-uuevent rate over legitimate predictors\.

###### Proof of Lemma[3](https://arxiv.org/html/2606.11267#Thmlemma3)\.

For any𝒢\\mathcal\{G\}\-setAAwithℙ​\(A\)=u\\mathbb\{P\}\(A\)=u,

𝔼​\[y​1A\]=𝔼​\[𝔼​\[y∣𝒢\]​𝟏A\]=𝔼​\[μ​1A\]≤𝔼​\[μ​1​\{μ≥qu\}\],\\mathbb\{E\}\[y\\,\\mathbf\{1\}\_\{A\}\]=\\mathbb\{E\}\[\\mathbb\{E\}\[y\\mid\\mathcal\{G\}\]\\mathbf\{1\}\_\{A\}\]=\\mathbb\{E\}\[\\mu\\,\\mathbf\{1\}\_\{A\}\]\\leq\\mathbb\{E\}\[\\mu\\,\\mathbf\{1\}\\\{\\mu\\geq q\_\{u\}\\\}\],wherequq\_\{u\}is the upper\-uuquantile ofμ\\mu, since among𝒢\\mathcal\{G\}\-sets of massuuthe integral ofμ\\muis maximized by the upper level set\{μ≥qu\}\\\{\\mu\\geq q\_\{u\}\\\}\(Hardy–Littlewood;[13](https://arxiv.org/html/2606.11267#bib.bib25), Ch\. 3\)\. Hence the best legitimate top\-uupurity is

ρ⋆​\(u\)=1u​𝔼​\[μ​1​\{μ≥qu\}\]=𝔼​\[μ∣μ≥qu\]\.\\rho^\{\\star\}\(u\)=\\frac\{1\}\{u\}\\,\\mathbb\{E\}\[\\mu\\,\\mathbf\{1\}\\\{\\mu\\geq q\_\{u\}\\\}\]=\\mathbb\{E\}\[\\mu\\mid\\mu\\geq q\_\{u\}\]\.Suppose the outcome is not prediction\-time\-deterministic:ℙ​\(μ=1\)=0\\mathbb\{P\}\(\\mu=1\)=0, i\.e\. no admissible event of positive mass hasy=1y=1almost surely\. Then for everyu\>0u\>0the event\{μ≥qu\}\\\{\\mu\\geq q\_\{u\}\\\}has mass≥u\>0\\geq u\>0and carries positive mass whereμ<1\\mu<1, so

ρ⋆​\(u\)=𝔼​\[μ∣μ≥qu\]<1for all​u∈\(0,1\],\\rho^\{\\star\}\(u\)=\\mathbb\{E\}\[\\mu\\mid\\mu\\geq q\_\{u\}\]<1\\qquad\\text\{for all \}u\\in\(0,1\],and no legitimate predictor attains top\-uupurity11on a non\-null fraction: a sustained unit\-purity head of non\-null widthu\>0u\>0is impossible\. If additionally the conditional risk is uniformly bounded away from one,M:=ess​sup⁡μ<1M:=\\operatorname\*\{ess\\,sup\}\\mu<1, the gap is uniform,1−ρ⋆​\(u\)≥1−M\>01\-\\rho^\{\\star\}\(u\)\\geq 1\-M\>0\. Contrapositively, an observedρ​\(u\)=1\\rho\(u\)=1over a non\-null fraction forcesℙ​\(μ=1\)\>0\\mathbb\{P\}\(\\mu=1\)\>0on the selected set, i\.e\. the predictor used information renderingyy\(near\-\)deterministic there—information not legitimately available at prediction time\. The only prior invoked is the qualitativeℙ​\(μ=1\)=0\\mathbb\{P\}\(\\mu=1\)=0\. ∎

## Appendix BDispersion null and finite\-size behavior

The statisticVηV\_\{\\eta\}sits in the dispersion\-statistic lineage that begins with goodness\-of\-fit tests for logistic regression\[[8](https://arxiv.org/html/2606.11267#bib.bib24)\]and probability forecast assessment\[[19](https://arxiv.org/html/2606.11267#bib.bib23)\]; the threshold\-weightingη\\etaextends that family to decision\-analytic stratifications\. Writeai=w​\(p^i\)2a\_\{i\}=w\(\\hat\{p\}\_\{i\}\)^\{2\}andsi2=p^i​\(1−p^i\)s\_\{i\}^\{2\}=\\hat\{p\}\_\{i\}\(1\-\\hat\{p\}\_\{i\}\), soVη=\(∑iai​ri2\)/\(∑iai​si2\)V\_\{\\eta\}=\\big\(\\sum\_\{i\}a\_\{i\}r\_\{i\}^\{2\}\\big\)\\big/\\big\(\\sum\_\{i\}a\_\{i\}s\_\{i\}^\{2\}\\big\)withri=yi−p^ir\_\{i\}=y\_\{i\}\-\\hat\{p\}\_\{i\}\.

#### Null mean\.

Under the calibrated nullyi∣p^i∼Bernoulli​\(p^i\)y\_\{i\}\\mid\\hat\{p\}\_\{i\}\\sim\\mathrm\{Bernoulli\}\(\\hat\{p\}\_\{i\}\)we have𝔼​\[ri∣p^i\]=0\\mathbb\{E\}\[r\_\{i\}\\mid\\hat\{p\}\_\{i\}\]=0and𝔼​\[ri2∣p^i\]=si2\\mathbb\{E\}\[r\_\{i\}^\{2\}\\mid\\hat\{p\}\_\{i\}\]=s\_\{i\}^\{2\}\. The denominator depends on the scores alone, so

𝔼null​\[Vη∣p^\]=∑iai​𝔼​\[ri2∣p^i\]∑iai​si2=∑iai​si2∑iai​si2=1,\\mathbb\{E\}\_\{\\text\{null\}\}\[V\_\{\\eta\}\\mid\\hat\{p\}\]=\\frac\{\\sum\_\{i\}a\_\{i\}\\,\\mathbb\{E\}\[r\_\{i\}^\{2\}\\mid\\hat\{p\}\_\{i\}\]\}\{\\sum\_\{i\}a\_\{i\}s\_\{i\}^\{2\}\}=\\frac\{\\sum\_\{i\}a\_\{i\}s\_\{i\}^\{2\}\}\{\\sum\_\{i\}a\_\{i\}s\_\{i\}^\{2\}\}=1,hence𝔼null​\[Vη\]=1\\mathbb\{E\}\_\{\\text\{null\}\}\[V\_\{\\eta\}\]=1exactly, not merely asymptotically\.

#### Fluctuation scale\.

For a Bernoulli\(p\)\(p\)residualr2∈\{\(1−p\)2,p2\}r^\{2\}\\in\\\{\(1\-p\)^\{2\},p^\{2\}\\\}, and a direct computation gives𝔼​\[r2\]=p​\(1−p\)\\mathbb\{E\}\[r^\{2\}\]=p\(1\-p\)andVar⁡\(r2\)=p​\(1−p\)​\(1−2​p\)2=s2​\(1−2​p\)2\\operatorname\{Var\}\(r^\{2\}\)=p\(1\-p\)\(1\-2p\)^\{2\}=s^\{2\}\(1\-2p\)^\{2\}\. WithVη−1=\(∑iai​\(ri2−si2\)\)/∑iai​si2V\_\{\\eta\}\-1=\\big\(\\sum\_\{i\}a\_\{i\}\(r\_\{i\}^\{2\}\-s\_\{i\}^\{2\}\)\\big\)/\\sum\_\{i\}a\_\{i\}s\_\{i\}^\{2\}, the numerator is a sum of conditionally independent mean\-zero terms with

Varnull⁡\(∑iai​\(ri2−si2\)\|p^\)=∑iai2​si2​\(1−2​p^i\)2=O​\(n\),\\operatorname\{Var\}\_\{\\text\{null\}\}\\\!\\Big\(\\textstyle\\sum\_\{i\}a\_\{i\}\(r\_\{i\}^\{2\}\-s\_\{i\}^\{2\}\)\\,\\Big\|\\,\\hat\{p\}\\Big\)=\\sum\_\{i\}a\_\{i\}^\{2\}s\_\{i\}^\{2\}\(1\-2\\hat\{p\}\_\{i\}\)^\{2\}=O\(n\),while the denominator is∑iai​si2=O​\(n\)\\sum\_\{i\}a\_\{i\}s\_\{i\}^\{2\}=O\(n\)\. ThusVarnull⁡\(Vη∣p^\)=O​\(n\)/O​\(n\)2=O​\(n−1\)\\operatorname\{Var\}\_\{\\text\{null\}\}\(V\_\{\\eta\}\\mid\\hat\{p\}\)=O\(n\)/O\(n\)^\{2\}=O\(n^\{\-1\}\)andVη=1\+Op​\(n−1/2\)V\_\{\\eta\}=1\+O\_\{p\}\(n^\{\-1/2\}\)\. The summands are bounded, so a Lindeberg CLT gives

n​\(Vη−1\)⇒𝒩​\(0,σV2\),σV2=limn→∞n​∑iai2​si2​\(1−2​p^i\)2\(∑iai​si2\)2,\\sqrt\{n\}\\,\(V\_\{\\eta\}\-1\)\\Rightarrow\\mathcal\{N\}\(0,\\sigma\_\{V\}^\{2\}\),\\qquad\\sigma\_\{V\}^\{2\}=\\lim\_\{n\\to\\infty\}\\frac\{n\\sum\_\{i\}a\_\{i\}^\{2\}s\_\{i\}^\{2\}\(1\-2\\hat\{p\}\_\{i\}\)^\{2\}\}\{\\big\(\\sum\_\{i\}a\_\{i\}s\_\{i\}^\{2\}\\big\)^\{2\}\},the reference law for the reportedzV=\(Vη−1\)/Var^z\_\{V\}=\(V\_\{\\eta\}\-1\)/\\sqrt\{\\widehat\{\\operatorname\{Var\}\}\}\.

#### Behavior under miscalibration\.

If leakage makes the predictor over\- or under\-dispersed,𝔼​\[ri2∣p^i\]=si2\+b​\(p^i\)\\mathbb\{E\}\[r\_\{i\}^\{2\}\\mid\\hat\{p\}\_\{i\}\]=s\_\{i\}^\{2\}\+b\(\\hat\{p\}\_\{i\}\)with calibration defectb≢0b\\not\\equiv 0\. Then𝔼​\[Vη\]=1\+\(∑iai​b​\(p^i\)\)/∑iai​si2\\mathbb\{E\}\[V\_\{\\eta\}\]=1\+\\big\(\\sum\_\{i\}a\_\{i\}b\(\\hat\{p\}\_\{i\}\)\\big\)/\\sum\_\{i\}a\_\{i\}s\_\{i\}^\{2\}, anO​\(1\)O\(1\)offset, sozV≍n→∞z\_\{V\}\\asymp\\sqrt\{n\}\\to\\infty: the miscalibrated leak is detected, with power growing innn\. Recalibration setsb≡0b\\equiv 0and restores𝔼​\[Vη\]=1\\mathbb\{E\}\[V\_\{\\eta\}\]=1—which is why this prior\-free signal is nonetheless evadable\.

#### Plateau versus smooth decay\.

Letμ\\mube the prediction\-time risk of Appendix[A](https://arxiv.org/html/2606.11267#A1)andρ⋆​\(u\)=𝔼​\[μ∣μ≥qu\]\\rho^\{\\star\}\(u\)=\\mathbb\{E\}\[\\mu\\mid\\mu\\geq q\_\{u\}\]the legitimate purity profile\. Onu∈\(0,1\]u\\in\(0,1\],ρ⋆\\rho^\{\\star\}is continuous and nonincreasing withρ⋆​\(1\)=π\\rho^\{\\star\}\(1\)=\\piand, under non\-determinism,ρ⋆​\(0\+\)=ess​sup⁡μ<1\\rho^\{\\star\}\(0^\{\+\}\)=\\operatorname\*\{ess\\,sup\}\\mu<1; its derivativedd​u​\(u​ρ⋆​\(u\)\)=qu\\tfrac\{d\}\{du\}\\big\(u\\rho^\{\\star\}\(u\)\\big\)=q\_\{u\}\(theuu\-quantile ofμ\\mu\) decays smoothly under any continuous risk distribution—the power\-law\-like lift decay of an honest ranker\. A near\-label leak instead pinsρ​\(u\)≡1\\rho\(u\)\\equiv 1on the leaked fractionu∈\(0,u0\]u\\in\(0,u\_\{0\}\]: a flat plateau at the absolute ceiling, of non\-null width, which Appendix[A](https://arxiv.org/html/2606.11267#A1)shows no legitimate predictor can produce\. The qualitative contrast—smooth sub\-unit decay versus a unit\-value plateau of non\-null width—is exactly the prior\-free*spike*signature, distinct from the*breadth*excess against a reference curve that Proposition[1](https://arxiv.org/html/2606.11267#Thmproposition1)shows requires an externalCmax∗C^\{\*\}\_\{\\max\}\.

## References

- \[1\]G\. S\. Collins, K\. G\. M\. Moons, P\. Dhiman, R\. D\. Riley, A\. L\. Beam, B\. Van Calster, M\. Ghassemi, X\. Liu, J\. B\. Reitsma, M\. van Smeden, A\. Boulesteix, J\. C\. Camaradou, L\. A\. Celi, S\. Denaxas, A\. K\. Denniston, B\. Glocker, R\. M\. Golub, H\. Harvey, G\. Heinze, M\. M\. Hoffman, A\. P\. Kengne, E\. Lam, N\. Lee, E\. W\. Loder, L\. Maier\-Hein, B\. A\. Mateen, M\. D\. McCradden, L\. Oakden\-Rayner, J\. Ordish, R\. Parnell, S\. Rose, K\. Singh, L\. Wynants, and P\. Logullo\(2024\)TRIPOD\+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods\.BMJ385,pp\. e078378\.External Links:[Document](https://dx.doi.org/10.1136/bmj-2023-078378)Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p2.1)\.
- \[2\]\(2015\)Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis \(TRIPOD\): the TRIPOD statement\.Annals of Internal Medicine162\(1\),pp\. 55–63\.External Links:[Document](https://dx.doi.org/10.7326/M14-0697)Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p2.1)\.
- \[3\]M\. H\. DeGroot and S\. E\. Fienberg\(1983\)The comparison and evaluation of forecasters\.Journal of the Royal Statistical Society: Series D \(The Statistician\)32\(1\-2\),pp\. 12–22\.External Links:[Document](https://dx.doi.org/10.2307/2987588)Cited by:[Appendix A](https://arxiv.org/html/2606.11267#A1.1.p1.11)\.
- \[4\]C\. Dwork, V\. Feldman, M\. Hardt, T\. Pitassi, O\. Reingold, and A\. Roth\(2015\)The reusable holdout: preserving validity in adaptive data analysis\.Science349\(6248\),pp\. 636–638\.External Links:[Document](https://dx.doi.org/10.1126/science.aaa9375)Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p2.1)\.
- \[5\]W\. Ehm, T\. Gneiting, A\. Jordan, and F\. Krüger\(2016\)Of quantiles and expectiles: consistent scoring functions, Choquet representations and forecast rankings\.Journal of the Royal Statistical Society: Series B \(Statistical Methodology\)78\(3\),pp\. 505–562\.External Links:[Document](https://dx.doi.org/10.1111/rssb.12154)Cited by:[§2](https://arxiv.org/html/2606.11267#S2.SS0.SSS0.Px2.p1.14)\.
- \[6\]T\. Gneiting and A\. E\. Raftery\(2007\)Strictly proper scoring rules, prediction, and estimation\.Journal of the American Statistical Association102\(477\),pp\. 359–378\.Cited by:[§2](https://arxiv.org/html/2606.11267#S2.SS0.SSS0.Px2.p1.14)\.
- \[7\]F\. E\. Harrell, K\. L\. Lee, and D\. B\. Mark\(1996\)Multivariable prognostic models: issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors\.Statistics in Medicine15\(4\),pp\. 361–387\.Cited by:[footnote 1](https://arxiv.org/html/2606.11267#footnote1)\.
- \[8\]D\. W\. Hosmer and S\. Lemeshow\(1980\)Goodness\-of\-fit tests for the multiple logistic regression model\.Communications in Statistics — Theory and Methods9\(10\),pp\. 1043–1069\.External Links:[Document](https://dx.doi.org/10.1080/03610928008827941)Cited by:[Appendix B](https://arxiv.org/html/2606.11267#A2.p1.6)\.
- \[9\]L\. A\. Jacobs and A\. J\. Vickers\(2026\)Expected net benefit: from decision curve analysis to a prior\-weighted summary measure for evaluating clinical prediction models\.Note:Nature Methods — In reviewCited by:[§2](https://arxiv.org/html/2606.11267#S2.SS0.SSS0.Px2.p1.6)\.
- \[10\]S\. Kapoor, E\. M\. Cantrell, K\. Peng, T\. H\. Pham, C\. A\. Bail, O\. E\. Gundersen, J\. M\. Hofman, J\. Hullman, M\. A\. Lones, M\. M\. Malik, P\. Nanayakkara, R\. A\. Poldrack, I\. D\. Raji, M\. Roberts, M\. J\. Salganik, M\. Serra\-Garcia, B\. M\. Stewart, G\. Vandewiele, and A\. Narayanan\(2024\)REFORMS: consensus\-based recommendations for machine\-learning\-based science\.Science Advances10\(18\),pp\. eadk3452\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.adk3452)Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p2.1)\.
- \[11\]S\. Kapoor and A\. Narayanan\(2023\)Leakage and the reproducibility crisis in machine\-learning\-based science\.Patterns4\(9\),pp\. 100804\.External Links:[Document](https://dx.doi.org/10.1016/j.patter.2023.100804)Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p1.1)\.
- \[12\]S\. Kaufman, S\. Rosset, C\. Perlich, and O\. Stitelman\(2012\)Leakage in data mining: formulation, detection, and avoidance\.ACM Transactions on Knowledge Discovery from Data6\(4\),pp\. 1–21\.Note:Article 15External Links:[Document](https://dx.doi.org/10.1145/2382577.2382579)Cited by:[Appendix A](https://arxiv.org/html/2606.11267#A1.5.p3.12),[§1](https://arxiv.org/html/2606.11267#S1.p1.1),[§3](https://arxiv.org/html/2606.11267#S3.p1.3)\.
- \[13\]E\. H\. Lieb and M\. Loss\(2001\)Analysis\.2nd edition,Graduate Studies in Mathematics, Vol\.14,American Mathematical Society\.Cited by:[Appendix A](https://arxiv.org/html/2606.11267#A1.6.p1.11)\.
- \[14\]M\. A\. Lones\(2024\)How to avoid machine learning pitfalls: a guide for academic researchers\.arXiv preprint\.Note:arXiv:2108\.02497v4Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p1.1)\.
- \[15\]J\. C\. Platt\(1999\)Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods\.InAdvances in Large Margin Classifiers,A\. J\. Smola, P\. L\. Bartlett, B\. Schölkopf, and D\. Schuurmans \(Eds\.\),pp\. 61–74\.Cited by:[§6](https://arxiv.org/html/2606.11267#S6.SS0.SSS0.Px1.p1.14),[§8](https://arxiv.org/html/2606.11267#S8.SS0.SSS0.Px3.p1.1)\.
- \[16\]M\. Roberts, D\. Driggs, M\. Thorpe, J\. Gilbey, M\. Yeung, S\. Ursprung, A\. I\. Aviles\-Rivero, C\. Etmann, C\. McCague, L\. Beer, J\. R\. Weir\-McCall, Z\. Teng, E\. Gkrania\-Klotsas, J\. H\. F\. Rudd, E\. Sala, and C\. Schönlieb\(2021\)Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID\-19 using chest radiographs and CT scans\.Nature Machine Intelligence3,pp\. 199–217\.External Links:[Document](https://dx.doi.org/10.1038/s42256-021-00307-0)Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p1.1)\.
- \[17\]M\. J\. Schervish\(1989\)A general method for comparing probability assessors\.The Annals of Statistics17\(4\),pp\. 1856–1879\.External Links:[Document](https://dx.doi.org/10.1214/aos/1176347398)Cited by:[Appendix A](https://arxiv.org/html/2606.11267#A1.1.p1.11),[§2](https://arxiv.org/html/2606.11267#S2.SS0.SSS0.Px2.p1.14)\.
- \[18\]R\. Shokri, M\. Stronati, C\. Song, and V\. Shmatikov\(2017\)Membership inference attacks against machine learning models\.In2017 IEEE Symposium on Security and Privacy \(S&P\),pp\. 3–18\.External Links:[Document](https://dx.doi.org/10.1109/SP.2017.41)Cited by:[§3](https://arxiv.org/html/2606.11267#S3.p4.1)\.
- \[19\]D\. J\. Spiegelhalter\(1986\)Probabilistic prediction in patient management and clinical trials\.Statistics in Medicine5\(5\),pp\. 421–433\.External Links:[Document](https://dx.doi.org/10.1002/sim.4780050506)Cited by:[Appendix B](https://arxiv.org/html/2606.11267#A2.p1.6)\.
- \[20\]C\. Sudlow, J\. Gallacher, N\. Allen, V\. Beral, P\. Burton, J\. Danesh, P\. Downey, P\. Elliott, J\. Green, M\. Landray, B\. Liu, P\. Matthews, G\. Ong, J\. Pell, A\. Silman, A\. Young, T\. Sprosen, T\. Peakman, and R\. Collins\(2015\)UK Biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age\.PLoS Medicine12\(3\),pp\. e1001779\.External Links:[Document](https://dx.doi.org/10.1371/journal.pmed.1001779)Cited by:[§6](https://arxiv.org/html/2606.11267#S6.SS0.SSS0.Px2.p1.5),[Acknowledgments](https://arxiv.org/html/2606.11267#Sx1.p1.1)\.
- \[21\]R\. Tibshirani\(1996\)Regression shrinkage and selection via the lasso\.Journal of the Royal Statistical Society: Series B \(Methodological\)58\(1\),pp\. 267–288\.External Links:[Document](https://dx.doi.org/10.1111/j.2517-6161.1996.tb02080.x)Cited by:[§6](https://arxiv.org/html/2606.11267#S6.SS0.SSS0.Px2.p1.5)\.
- \[22\]B\. Van Calster, D\. J\. McLernon, M\. van Smeden, L\. Wynants, and E\. W\. Steyerberg\(2019\)Calibration: the Achilles heel of predictive analytics\.BMC Medicine17,pp\. 230\.External Links:[Document](https://dx.doi.org/10.1186/s12916-019-1466-7)Cited by:[§2](https://arxiv.org/html/2606.11267#S2.SS0.SSS0.Px1.p1.9)\.
- \[23\]B\. Van Calster, D\. Nieboer, Y\. Vergouwe, B\. De Cock, M\. J\. Pencina, and E\. W\. Steyerberg\(2016\)A calibration hierarchy for risk models was defined: from utopia to empirical data\.Journal of Clinical Epidemiology74,pp\. 167–176\.External Links:[Document](https://dx.doi.org/10.1016/j.jclinepi.2015.12.005)Cited by:[§2](https://arxiv.org/html/2606.11267#S2.SS0.SSS0.Px1.p1.9)\.
- \[24\]A\. J\. Vickers and E\. B\. Elkin\(2006\)Decision curve analysis: a novel method for evaluating prediction models\.Medical Decision Making26\(6\),pp\. 565–574\.External Links:[Document](https://dx.doi.org/10.1177/0272989X06295361)Cited by:[§2](https://arxiv.org/html/2606.11267#S2.SS0.SSS0.Px2.p1.6)\.
- \[25\]R\. F\. Wolff, K\. G\. M\. Moons, R\. D\. Riley, P\. F\. Whiting, M\. Westwood, G\. S\. Collins, J\. B\. Reitsma, J\. Kleijnen, and S\. Mallett\(2019\)PROBAST: a tool to assess the risk of bias and applicability of prediction model studies\.Annals of Internal Medicine170\(1\),pp\. 51–58\.External Links:[Document](https://dx.doi.org/10.7326/M18-1376)Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p2.1)\.
- \[26\]C\. Yang, R\. A\. Brower\-Sinning, G\. A\. Lewis, and C\. Kästner\(2022\)Data leakage in notebooks: static detection and better processes\.InProceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering \(ASE\),Note:Article 30External Links:[Document](https://dx.doi.org/10.1145/3551349.3556918)Cited by:[§1](https://arxiv.org/html/2606.11267#S1.p2.1)\.

Similar Articles

NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models

arXiv cs.LG

This paper introduces NumLeak, a framework for detecting when foundation models memorize public numeric benchmarks from pretraining rather than demonstrating out-of-sample skill, and shows that top LLMs recall values like Fama-French returns with high fidelity, proposing a simple system-prompt defense.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

arXiv cs.LG

This paper shows that the standard pre/post training-cutoff check for temporal leakage in LLM backtesting is uninformative, as recency effects mimic leakage. It proposes new estimators using known cutoffs and matched clean controls to measure leakage and compute adjusted scores, validated on frontier models.