When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
Summary
This paper formalizes when benchmark contamination is detectable, deriving information-theoretic limits and proposing power-calibrated audits that distinguish a clean benchmark from a powerless detector. It reports two-sided empirical findings on calibration efficacy and validity gates.
View Cached Full Text
Cached at: 08/11/26, 08:04 AM
# When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
Source: [https://arxiv.org/html/2608.07914](https://arxiv.org/html/2608.07914)
Sanjeda Akter11footnotemark:11Anuj Sharma2 1Department of Computer Science, Iowa State University 2Department of Civil, Construction & Environmental Engineering, Iowa State University ishihab@iastate\.edu
###### Abstract
Behavioral contamination detectors can return “no evidence” either because a benchmark is clean or because the audit has little power\. We formalize the distinction for a benchmark in which an unknown fractionα\\alphaof items was seen during training: with matched clean and seen controls the behavioral channel is the sparse mixtureQα=\(1−α\)P0\+αP1Q\_\{\\alpha\}=\(1\-\\alpha\)P\_\{0\}\+\\alpha P\_\{1\}, and an exact second\-moment argument shows detectability is governed byαρm\\alpha\\rho\\sqrt\{m\}, whereρ2=χ2\(P1∥P0\)\\rho^\{2\}=\\chi^\{2\}\(P\_\{1\}\\\|P\_\{0\}\)measures behavioral separability\. Any scalar detector reduces to its efficacyef=\|𝔼1f−𝔼0f\|/Var0\(f\)≤ρe\_\{f\}=\|\\mathbb\{E\}\_\{1\}f\-\\mathbb\{E\}\_\{0\}f\|/\\sqrt\{\\operatorname\{Var\}\_\{0\}\(f\)\}\\leq\\rho, which is estimable from controls before the audit is run; a separate sample\-split certificate lower\-boundsα\\alphadistribution\-free, without an orientation assumption\. Our empirical finding is two\-sided\. Frozen calibration efficacy predicts held\-out power*curves*\(R2=0\.83R^\{2\}=0\.83–0\.980\.98over six exact\-permutation channels\), but the efficacy\-only Gaussian*budget*is miscalibrated at the smallmmit prescribes, failing in9/99/9gate\-passing channels even though efficacy itself transports—the inversion breaks, not the calibration\. A predeclared two\-stage planner that simulates the deployed test repairs the budgets, is uniformly conservative, and abstains where its probe does not transport\. The certificate is valid but vacuous at audit scale, and a five\-seed paired injection study recovers a mechanism ordering \(verbatim\>\>paraphrase\>\>surface\) in which the apparent answer\-only signal is baseline drift\. We report the audit contract and its failures together: a non\-rejection is interpretable only alongside the efficacy, budget, and validity gates that produced it\.
When Is Benchmark Contamination Detectable? Information Limits and Power\-Calibrated Audits
Ibne Farabi Shihab††thanks:Equal contribution\.††thanks:Corresponding author:ishihab@iastate\.edu\.1and Sanjeda Akter11footnotemark:11and Anuj Sharma21Department of Computer Science, Iowa State University2Department of Civil, Construction & Environmental Engineering, Iowa State Universityishihab@iastate\.edu
## 1Introduction
Public evaluation sets routinely appear in web\-scale training corpora\. Exact or semantic exposure can inflate a model’s benchmark score without improving the capability the benchmark was intended to measure\(Magar and Schwartz,[2022](https://arxiv.org/html/2608.07914#bib.bib18); Donget al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib11)\)\. A growing literature therefore tries to infer exposure from token likelihoods, perturbations, output order, or elicited recall\(Matternet al\.,[2023](https://arxiv.org/html/2608.07914#bib.bib21); Shiet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib29); Orenet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib23); Golchin and Surdeanu,[2025](https://arxiv.org/html/2608.07914#bib.bib14)\)\. Yet empirical studies find that detectors often disagree, fail under distribution shift, or can be evaded\(Duanet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib12); Fuet al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib13); Meeuset al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib22); Samuelet al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib27); Dekonincket al\.,[2024a](https://arxiv.org/html/2608.07914#bib.bib8); Wanget al\.,[2026a](https://arxiv.org/html/2608.07914#bib.bib32)\)\. A 25\-model study documents this reliability gap—only 201 of 335 audit outcomes correct, with shift false positives and underpowered post\-hoc inference\(Zarzeckiet al\.,[2026](https://arxiv.org/html/2608.07914#bib.bib34)\); it establishes the empirical failure, and our question is whether a calibrated channel can predict it before an audit is run—absent that, a non\-rejection is hard to interpret\.
The missing object is*audit power*: how many independent items must be probed to detect a small exposed fraction, how does the answer depend on the model’s observable memory of a seen item, and can a detector state, with controlled error, which fractions it rules out? Dataset\-level methods aggregate membership scores into significance tests\(Mainiet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib20); Zhanget al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib35); Puertoet al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib24)\), ConStat estimates performance inflation\(Dekonincket al\.,[2024b](https://arxiv.org/html/2608.07914#bib.bib9)\), and FTD controls a false discovery rate\(Zhanget al\.,[2026](https://arxiv.org/html/2608.07914#bib.bib37)\)—important but different deliverables: a validpp\-value does not say which exposure fractions the audit had power to detect, and inflation is not the fraction exposed\.
We study the benchmark\-level hypothesisH0:α=0H\_\{0\}:\\alpha=0versusH1:α\>0H\_\{1\}:\\alpha\>0, whereα\\alphais the fraction of benchmark items exposed during training\. A probe outcomeYYcan contain a scalar loss, a vector of token statistics, a set of stochastic generations, or a black\-box answer\. Conditional on an item being clean or seen, its outcome followsP0P\_\{0\}orP1P\_\{1\}\. Uniformly sampling benchmark items gives the mixture
Qα=\(1−α\)P0\+αP1\.Q\_\{\\alpha\}=\(1\-\\alpha\)P\_\{0\}\+\\alpha P\_\{1\}\.\(1\)This assumption is substantive—it requires matched clean controls and rules out unmodeled spillover—and we state it visibly, test its implications, and show what fails without it\.
#### Contributions\.
1. 1\.We connect behavioral contamination audits to classical sparse\-mixture detection\(Ingster,[1997](https://arxiv.org/html/2608.07914#bib.bib16); Donoho and Jin,[2004](https://arxiv.org/html/2608.07914#bib.bib10); Caiet al\.,[2011](https://arxiv.org/html/2608.07914#bib.bib3); Cai and Wu,[2014](https://arxiv.org/html/2608.07914#bib.bib4)\), instantiating a sharp second\-moment limit in the observableρ2=χ2\(P1∥P0\)\\rho^\{2\}=\\chi^\{2\}\(P\_\{1\}\\\|P\_\{0\}\)for Bernoulli exposure and fixed contaminated subsets; we claim the audit interpretation, not a new mixture boundary\.
2. 2\.We operationalize that limit for arbitrary contamination scores: the efficacyef≤ρe\_\{f\}\\leq\\rhoof the oriented mean\-score audit built fromffyields a prospective*local\-asymptotic*power prediction and a finite\-sample lower bound on channel separability\. Power*curves*transport across splits \(R2R^\{2\}up to0\.990\.99\); the efficacy\-only Gaussian*planner*is miscalibrated at the small budgets it prescribes, and we report that failure with its diagnosis rather than a pooled scaling law \(§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\)\.
3. 3\.We provide a sample\-split audit reporting a distribution\-free lower confidence bound on the exposed fractionα\\alpha, with coverage for any independently learned bounded score, even one accidentally reversed\.
4. 4\.We design complementary evaluations: checkpoint\-matched channels, gate\-passing same\-corpus continued\-pretraining channels with direct budget trials, and causal injections with five paired clean\-continuation seeds\. Primary falsifiable endpoints: prospective power prediction, efficacy transport \(tested; efficacy itself largely transports, but the Gaussian efficacy\-to\-budget*inversion*fails at prescribed budgets\), and coverage\.
## 2Problem Formulation
### 2\.1From benchmark items to a behavioral channel
Fix a trained language model, a benchmark population, an access regime, and a predeclared probing procedure\. LetZ∈\{0,1\}Z\\in\\\{0,1\\\}indicate whether a uniformly sampled benchmark item was exposed during training\. The complete outcome from that item isY∈𝒴Y\\in\\mathcal\{Y\}\. We treat all repeated decodes, perturbations, and token\-level features from the*same*item as one joint outcome; counting them as independent probes would artificially inflatemm\.
###### Assumption 1\(Matched partial\-contamination channel\)\.
For the audited model and probe,Y∣Z=0∼P0Y\\mid Z=0\\sim P\_\{0\}andY∣Z=1∼P1Y\\mid Z=1\\sim P\_\{1\}\. Probe items are sampled independently and uniformly,Z∼Bernoulli\(α\)Z\\sim\\operatorname\{Bernoulli\}\(\\alpha\), and the conditional lawsP0,P1P\_\{0\},P\_\{1\}do not change withα\\alphaover the range being audited\.
The last clause is a local no\-spillover/stability condition: it does not claim training is literally itemwise, but specifies when a clean reference channel transports to the suspect model\. Clean\-control and held\-out items measure violations in[section˜5](https://arxiv.org/html/2608.07914#S5); heterogeneous and adaptive versions are in[appendix˜E](https://arxiv.org/html/2608.07914#A5)\.
Under[˜1](https://arxiv.org/html/2608.07914#Thmtheorem1),mmoutcomes followP0⊗mP\_\{0\}^\{\\otimes m\}underH0H\_\{0\}andQα⊗mQ\_\{\\alpha\}^\{\\otimes m\}underH1H\_\{1\}\. A possibly randomized audit is a measurableϕm:𝒴m→\[0,1\]\\phi\_\{m\}:\\mathcal\{Y\}^\{m\}\\to\[0,1\], whereϕm=1\\phi\_\{m\}=1means “detect contamination,” with type\-I erroram\(ϕm\)=𝔼0\[ϕm\]a\_\{m\}\(\\phi\_\{m\}\)=\\mathbb\{E\}\_\{0\}\[\\phi\_\{m\}\], type\-II errorbm\(ϕm,α\)=𝔼α\[1−ϕm\]b\_\{m\}\(\\phi\_\{m\},\\alpha\)=\\mathbb\{E\}\_\{\\alpha\}\[1\-\\phi\_\{m\}\], and powerπm\(ϕm,α\)=1−bm\(ϕm,α\)\\pi\_\{m\}\(\\phi\_\{m\},\\alpha\)=1\-b\_\{m\}\(\\phi\_\{m\},\\alpha\)\. We call an audit levelτ\\tauifam≤τa\_\{m\}\\leq\\tau\.
### 2\.2Why contamination fraction alone is insufficient
There is no nontrivial guarantee that depends only on\(α,m\)\(\\alpha,m\): settingP1=P0P\_\{1\}=P\_\{0\}makes the clean and contaminated transcript laws identical \(am\+bm=1a\_\{m\}\+b\_\{m\}=1for every test\), representing exposure that leaves no trace in the chosen access channel, and an unconstrained clean reference law lets any observed law be relabeled as clean\. A universalΘ\(m−1/2\)\\Theta\(m^\{\-1/2\}\)threshold therefore needs an explicit informativeness parameter\.
###### Definition 2\(Behavioral separability\)\.
AssumeP1≪P0P\_\{1\}\\ll P\_\{0\}, letr=dP1/dP0r=\\,\\mathrm\{d\}P\_\{1\}/\\,\\mathrm\{d\}P\_\{0\}, and define
ρ2\\displaystyle\\rho^\{2\}:=χ2\(P1∥P0\)\\displaystyle:=\\chi^\{2\}\(P\_\{1\}\\\|P\_\{0\}\)=𝔼0\[\(r\(Y\)−1\)2\]\.\\displaystyle=\\mathbb\{E\}\_\{0\}\[\(r\(Y\)\-1\)^\{2\}\]\.\(2\)
ρ\\rhodepends jointly on the model, mechanism, benchmark distribution, probe, and access regime: zero when the probe carries no membership signal, large when seen items behave in ways rare underP0P\_\{0\}, and infinite off\-support \(where a single witness beats the regular rate\)\. Our matching result concerns the practically common overlapping, finite\-ρ\\rhoregime\.
## 3Fundamental Detection Limit
### 3\.1A finite\-sample second\-moment bound
For compactness, define the second\-moment radius
Cm\(α,ρ\):=12\{\(1\+α2ρ2\)m−1\}1/2\.C\_\{m\}\(\\alpha,\\rho\):=\\frac\{1\}\{2\}\\left\\\{\(1\+\\alpha^\{2\}\\rho^\{2\}\)^\{m\}\-1\\right\\\}^\{1/2\}\.\(3\)
###### Theorem 3\(Power limit for partial contamination\)\.
Under[˜1](https://arxiv.org/html/2608.07914#Thmtheorem1), supposeP1≪P0P\_\{1\}\\ll P\_\{0\}andρ2<∞\\rho^\{2\}<\\infty\. For everyα∈\[0,1\]\\alpha\\in\[0,1\], everym≥1m\\geq 1, and every possibly randomized testϕm\\phi\_\{m\},
am\(ϕm\)\+bm\(ϕm,α\)\\displaystyle a\_\{m\}\(\\phi\_\{m\}\)\+b\_\{m\}\(\\phi\_\{m\},\\alpha\)≥1−min\{1,Cm\(α,ρ\)\}\.\\displaystyle\\geq 1\-\\min\\\{1,C\_\{m\}\(\\alpha,\\rho\)\\\}\.\(4\)Consequently, every level\-τ\\tauaudit satisfies
πm\(ϕm,α\)\\displaystyle\\pi\_\{m\}\(\\phi\_\{m\},\\alpha\)≤min\{1,τ\+min\{1,Cm\(α,ρ\)\}\}\.\\displaystyle\\leq\\min\\\{1,\\tau\+\\min\\\{1,C\_\{m\}\(\\alpha,\\rho\)\\\}\\\}\.\(5\)
The calculation is exact up to the final TV inequality \(χ2\(Qα⊗m∥P0⊗m\)=\(1\+α2ρ2\)m−1\\chi^\{2\}\(Q\_\{\\alpha\}^\{\\otimes m\}\\\|P\_\{0\}^\{\\otimes m\}\)=\(1\+\\alpha^\{2\}\\rho^\{2\}\)^\{m\}\-1, thenTV≤12χ2\\operatorname\{TV\}\\leq\\frac\{1\}\{2\}\\sqrt\{\\chi^\{2\}\}\), sharper than passing through KL and Pinsker; all steps are in[section˜A\.1](https://arxiv.org/html/2608.07914#A1.SS1)\.
The same limit covers the realistic design in which an auditor evaluates every item of a benchmark with a*fixed, unknown*contaminated subset of sizekk:[lemma˜6](https://arxiv.org/html/2608.07914#Thmtheorem6)\(stated and proved in[section˜A\.2](https://arxiv.org/html/2608.07914#A1.SS2)\) gives the analogous bound with weak\-signal boundarykρ/m=αρm=Θ\(1\)k\\rho/\\sqrt\{m\}=\\alpha\\rho\\sqrt\{m\}=\\Theta\(1\)fork=αmk=\\alpha m, the symmetric oracle score attains it without knowing the subset, and the subset\-averaged bound is a Bayes and hence minimax lower bound\.
###### Corollary 4\(Regular detection boundary\)\.
For any sequence of channels and alternatives for whichαmρmm→0\\alpha\_\{m\}\\rho\_\{m\}\\sqrt\{m\}\\to 0, every asymptotic level\-τ\\tauaudit haslim supmπm\(αm\)≤τ\\limsup\_\{m\}\\pi\_\{m\}\(\\alpha\_\{m\}\)\\leq\\tau\. Thus it has no asymptotic power beyond its false\-positive rate\. A necessary order for nontrivial power is
αm=Ω\(1ρmm\)\.\\alpha\_\{m\}=\\Omega\\\!\\left\(\\frac\{1\}\{\\rho\_\{m\}\\sqrt\{m\}\}\\right\)\.
The familiarm−1/2m^\{\-1/2\}exponent appears only whenρm\\rho\_\{m\}remains bounded away from zero and infinity; capacity alone does not upper\-boundρ\\rhowithout a training\-stability theorem, so the informal capacity step of prior arguments is replaced by an observable channel quantity\.
### 3\.2A matching test
The lower\-bound scaling is attainable: the centered likelihood\-ratio scoreg=r−1g=r\-1yields a test that is locally asymptotically optimal, attaining the envelope1−Φ\(z1−τ−h\)1\-\\Phi\(z\_\{1\-\\tau\}\-h\)at local alternativesαm=h/\(ρm\)\\alpha\_\{m\}=h/\(\\rho\\sqrt\{m\}\), with the oracle sample budgetmoracle≈\(\(z1−τ\+z1−β\)/\(αρ\)\)2m\_\{\\mathrm\{oracle\}\}\\approx\(\(z\_\{1\-\\tau\}\+z\_\{1\-\\beta\}\)/\(\\alpha\\rho\)\)^\{2\}\([eq\.˜18](https://arxiv.org/html/2608.07914#A4.E18)\); every regular channel admits a test matching theαρm\\alpha\\rho\\sqrt\{m\}boundary \(construction and LAN statement in[appendix˜D](https://arxiv.org/html/2608.07914#A4)\)\.
### 3\.3What an existing detector score can achieve
Most practical methods output a scalar score rather than an estimated density ratio\. The following proposition turns such a score into a power diagnostic\.
###### Proposition 5\(Score efficacy\)\.
For any nonconstantf∈L2\(P0\)f\\in L^\{2\}\(P\_\{0\}\), let
Δf\\displaystyle\\Delta\_\{f\}=𝔼1f−𝔼0f,\\displaystyle=\\mathbb\{E\}\_\{1\}f\-\\mathbb\{E\}\_\{0\}f,σf2\\displaystyle\\sigma\_\{f\}^\{2\}=Var0\(f\),\\displaystyle=\\operatorname\{Var\}\_\{0\}\(f\),ef\\displaystyle e\_\{f\}=\|Δf\|σf\.\\displaystyle=\\frac\{\|\\Delta\_\{f\}\|\}\{\\sigma\_\{f\}\}\.Thenef≤ρe\_\{f\}\\leq\\rho, with equality if and only iff−𝔼0f=c\(r−1\)f\-\\mathbb\{E\}\_\{0\}f=c\(r\-1\)P0P\_\{0\}\-almost surely for some nonzero constantcc\. Moreover,
𝔼Qαf−𝔼0f=αΔf,\\mathbb\{E\}\_\{Q\_\{\\alpha\}\}f\-\\mathbb\{E\}\_\{0\}f=\\alpha\\Delta\_\{f\},so the standardized local shift of the mean\-score audit isαmef\\alpha\\sqrt\{m\}\\,e\_\{f\}\.
Thus AUROC and benchmark\-level power answer different questions: two scores with similar AUROC can have different low\-α\\alphaefficacy, weak instance\-level separation can still accumulate over a large benchmark, andef/ρ∈\[0,1\]e\_\{f\}/\\rho\\in\[0,1\]measures how much local information a score uses\. For an oriented mean\-score test, the operational planning equation is
mtarget\(f\)≈\(z1−τ\+z1−βαef\)2\.m\_\{\\mathrm\{target\}\}\(f\)\\approx\\left\(\\frac\{z\_\{1\-\\tau\}\+z\_\{1\-\\beta\}\}\{\\alpha e\_\{f\}\}\\right\)^\{2\}\.\(6\)Unlike[eq\.˜18](https://arxiv.org/html/2608.07914#A4.E18), this quantity is estimable from matched controls and refers to the chosen detector, not an unattained channel oracle\. Two scope warnings, borne out empirically:efe\_\{f\}summarizes the local power of the*oriented mean\-score test*built fromff, not of every detector computable fromff; and[eq\.˜6](https://arxiv.org/html/2608.07914#S3.E6)is a first\-order Gaussian inversion, inheriting every finite\-sample defect of the local approximation \(discreteness, higher moments, alternative variance, test conservatism\)—§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)measures exactly this failure at smallmm\.
A companion variance\-adaptive finite\-sample floor \([corollary˜9](https://arxiv.org/html/2608.07914#Thmtheorem9),[appendix˜C](https://arxiv.org/html/2608.07914#A3)\) converts calibration samples of a bounded score into a certified lower boundρ≥ef≥e¯f\\rho\\geq e\_\{f\}\\geq\\underline\{e\}\_\{f\}that holds with probability1−δcal1\-\\delta\_\{\\mathrm\{cal\}\}, using the empirical\-Bernstein standard\-deviation radius ofMaurer and Pontil \([2009](https://arxiv.org/html/2608.07914#bib.bib19)\)\. Substitutinge¯f\\underline\{e\}\_\{f\}into[eq\.˜6](https://arxiv.org/html/2608.07914#S3.E6)is conservative about calibration uncertainty but remains a local\-asymptotic planning calculation, not a finite\-sample power guarantee\.
### 3\.4Adaptive and heterogeneous probes
An auditor may choose the next prompt after observing previous answers\. Adaptivity can allocate queries toward more informative probes but cannot exceed a cumulative KL information budget of12∑tlog\(1\+α2ρt2\)\\tfrac\{1\}\{2\}\\sum\_\{t\}\\log\(1\+\\alpha^\{2\}\\rho\_\{t\}^\{2\}\)over fresh items \([propositions˜10](https://arxiv.org/html/2608.07914#Thmtheorem10)and[E](https://arxiv.org/html/2608.07914#A5)\); independent heterogeneous probes admit exact chi\-square tensorization, with local informationα2∑iρi2\\alpha^\{2\}\\sum\_\{i\}\\rho\_\{i\}^\{2\}\. The precise statement and the KL/Pinsker power bound appear in[appendix˜E](https://arxiv.org/html/2608.07914#A5)\.
## 4A Power\-Calibrated Audit
### 4\.1Learning an efficient score
The oracler=p1/p0r=p\_\{1\}/p\_\{0\}is unknown, so we learn a probabilistic classifiersθ\(y\)≈ℙ\(Z=1∣Y=y\)s\_\{\\theta\}\(y\)\\approx\\mathbb\{P\}\(Z=1\\mid Y=y\)from labeled, matched clean and seen controls under balanced sampling; the primary audit uses the clipped scorefθ\(y\)∈\[0,1\]f\_\{\\theta\}\(y\)\\in\[0,1\]for stability and finite\-sample certification, with the unclipped implied ratiosθ/\(1−sθ\)s\_\{\\theta\}/\(1\-s\_\{\\theta\}\)kept only as a cross\-fitted diagnostic, never as an oracle estimate ofρ\\rho\.
To prevent optimistic power estimates, data are divided by original document or benchmark item into a*score\-training split*\(fitsfθf\_\{\\theta\}\), a*calibration split*\(estimates𝔼0f\\mathbb\{E\}\_\{0\}f,𝔼1f\\mathbb\{E\}\_\{1\}f,σf\\sigma\_\{f\}, and prospective budgets\), and an untouched*audit split*with fresh control items for testing and confidence bounds\. Cross\-fitting can rotate these roles, but every reported fold keeps the score independent of its calibration and audit observations\. Existing scores \(loss, Min\-K, ReCaLL, and others\) enter exactly the same pipeline after their orientation is chosen on the training split\.
### 4\.2Detection and prospective power
LetΔ^f=f¯1−f¯0\\widehat\{\\Delta\}\_\{f\}=\\bar\{f\}\_\{1\}\-\\bar\{f\}\_\{0\}andσ^f\\widehat\{\\sigma\}\_\{f\}be computed only on calibration controls\. The plug\-in efficacye^f=\|Δ^f\|/σ^f\\widehat\{e\}\_\{f\}=\|\\widehat\{\\Delta\}\_\{f\}\|/\\widehat\{\\sigma\}\_\{f\}gives the point budget in[eq\.˜6](https://arxiv.org/html/2608.07914#S3.E6); replacing it bye¯f\\underline\{e\}\_\{f\}gives a one\-sided, uncertainty\-aware local budget\. The*primary confirmatory decision*is a one\-sided two\-sample permutation test ofs^f\(f¯A−f¯0test\)\\widehat\{s\}\_\{f\}\(\\bar\{f\}\_\{A\}\-\\bar\{f\}\_\{0\}^\{\\rm test\}\), wheres^f=sign\(Δ^f\)\\widehat\{s\}\_\{f\}=\\operatorname\{sign\}\(\\widehat\{\\Delta\}\_\{f\}\)and all score choices and orientations are frozen before the audit split is opened\. Conditional on exchangeability of the pooled audit/clean scores underH0H\_\{0\}and on the frozen score, the exact \(randomized\) permutation test controls type\-I error atτ=0\.05\\tau=0\.05\. The deployed implementation is Monte\-Carlo with199199permutations and the add\-one estimatorp^=\(1\+\#\{perm≥obs\}\)/200\\hat\{p\}=\(1\+\\\#\\\{\\text\{perm\}\\geq\\text\{obs\}\\\}\)/200\(valid; conservative on a1/2001/200grid\); exchangeability is exactly what the validity gates \(§[4\.4](https://arxiv.org/html/2608.07914#S4.SS4)\) probe, so empirical size is reported per channel, never asserted \(the initial runs use a cheaper standardized\-mean approximation, making size a measured endpoint\)\. Our remaining empirical endpoints are power as a function of\(α,m\)\(\\alpha,m\)and its prediction from the frozen calibration efficacy \(R2R^\{2\}, MAE;[table˜7](https://arxiv.org/html/2608.07914#A11.T7)\); predicted versus observed local sample\-size targets for80%80\\%power \([table˜10](https://arxiv.org/html/2608.07914#A11.T10)\); and the smallest detectable fraction at the available size \([table˜11](https://arxiv.org/html/2608.07914#A11.T11)\)\. Bootstrap intervals resample original items, never individual tokens or decodes\.
When the null mean is estimated fromn0testn\_\{0\}^\{\\rm test\}fresh clean controls, the local coordinate usesmeff=\(m−1\+\(n0test\)−1\)−1m\_\{\\rm eff\}=\(m^\{\-1\}\+\(n\_\{0\}^\{\\rm test\}\)^\{\-1\}\)^\{\-1\}in place ofmm, and every main power curve accounts for it; equations below writemmfor the externally calibrated orn0test≫mn\_\{0\}^\{\\rm test\}\\gg mregime\.
Atτ=\.05\\tau=\.05and80%80\\%power,z\.95\+z\.80=2\.486z\_\{\.95\}\+z\_\{\.80\}=2\.486\. Thus
αmin\(f,m\)≈2\.486efm\.\\alpha\_\{\\min\}\(f,m\)\\approx\\frac\{2\.486\}\{e\_\{f\}\\sqrt\{m\}\}\.\(7\)[Table˜11](https://arxiv.org/html/2608.07914#A11.T11)makes the resulting sensitivity analysis usable when a closed model does not permit matched control construction\.
### 4\.3A finite\-sample lower confidence bound
A test can say that contamination is present; a prevalence certificate says how much, distribution\-free for a bounded, independently learned score\.
Theorem[8](https://arxiv.org/html/2608.07914#Thmtheorem8)\(stated formally in[appendix˜B](https://arxiv.org/html/2608.07914#A2), proved in[section˜A\.7](https://arxiv.org/html/2608.07914#A1.SS7)\) takes any\[0,1\]\[0,1\]\-valued scorefffixed independently of the data, three mutually independent samples—mmaudit items,n0n\_\{0\}clean controls,n1n\_\{1\}seen controls—and returns
α¯=\{min\{1,max\{0,NLDU\}\},DU\>0,0,DU≤0,\\underline\{\\alpha\}=\\begin\{cases\}\\min\\bigl\\\{1,\\max\\bigl\\\{0,\\tfrac\{N\_\{L\}\}\{D\_\{U\}\}\\bigr\\\}\\bigr\\\},&D\_\{U\}\>0,\\\\\[2\.0pt\] 0,&D\_\{U\}\\leq 0,\\end\{cases\}\(8\)withεcert\(n\)=log\(6/δcert\)/\(2n\)\\varepsilon\_\{\\rm cert\}\(n\)=\\sqrt\{\\log\(6/\\delta\_\{\\mathrm\{cert\}\}\)/\(2n\)\}, whereNL=\(f¯A−εcert\(m\)\)−\(f¯0\+εcert\(n0\)\)N\_\{L\}=\(\\bar\{f\}\_\{A\}\-\\varepsilon\_\{\\rm cert\}\(m\)\)\-\(\\bar\{f\}\_\{0\}\+\\varepsilon\_\{\\rm cert\}\(n\_\{0\}\)\)shrinks the observed audit–clean gap by*both*sample radii andDU=\(f¯1\+εcert\(n1\)\)−\(f¯0−εcert\(n0\)\)D\_\{U\}=\(\\bar\{f\}\_\{1\}\+\\varepsilon\_\{\\rm cert\}\(n\_\{1\}\)\)\-\(\\bar\{f\}\_\{0\}\-\\varepsilon\_\{\\rm cert\}\(n\_\{0\}\)\)inflates the control gap; thenℙα\{α¯≤α\}≥1−δcert\\mathbb\{P\}\_\{\\alpha\}\\\{\\underline\{\\alpha\}\\leq\\alpha\\\}\\geq 1\-\\delta\_\{\\mathrm\{cert\}\}by a simultaneous three\-sample Hoeffding argument\. TheDU≤0D\_\{U\}\\leq 0guard matches the formal statement in[appendix˜B](https://arxiv.org/html/2608.07914#A2): without it, a reversed score can giveNL<0N\_\{L\}<0andDU<0D\_\{U\}<0with a spuriously positive ratio\. No orientation assumption is needed \(a reversed score yieldsα¯=0\\underline\{\\alpha\}=0\), the bound is conservative by design, and rejecting whenα¯\>0\\underline\{\\alpha\}\>0is a fallback whose rejection rate we report beside the confirmatory permutation test, not the source of the main power curves\.
### 4\.4Required validity checks
The mathematics cannot rescue a mismatched control channel\. We therefore make four diagnostics part of the audit contract:blind separation\(a text\-only classifier must not separate the pools\),clean\-channel transport\(clean\-control scores must not shift across runs\),channel stability\(conditional score laws must be stable acrossα\\alpha\), and theindependence unit\(effectivemmcounts source\-item clusters, never near\-duplicates or repeated decodes\)\. Full definitions and thresholds are in[appendix˜F](https://arxiv.org/html/2608.07914#A6)\.
The interpretation is also predeclared: a pooled transport claim requires median multiplicative sample\-size error≤2\\leq 2*and*≥80%\\geq 80\\%of held\-out channels inside nominal90%90\\%prediction intervals; otherwise the pooled scaling claim is withdrawn and calibration\-to\-audit nontransport becomes the primary conclusion\. \(The v5 intervals are descriptive, so the second criterion is unevaluable as specified; §[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\.\)
## 5Experiments
### 5\.1Research questions
Four pre\-specified questions:RQ1 \(prospective prediction\)—does calibration\-set efficacy predict held\-out audit power and sample size?RQ2 \(certification\)—does[eq\.˜8](https://arxiv.org/html/2608.07914#S4.E8)attain nominal coverage, and how conservative is it?RQ3 \(transport and mechanism\)—how do mechanism, repetition, access, and post\-training altere^f\\widehat\{e\}\_\{f\}, its floor, and rankings?RQ4 \(scaling diagnostic\)—does power collapse onαe^fmeff\\alpha\\widehat\{e\}\_\{f\}\\sqrt\{m\_\{\\rm eff\}\}with slope near−1/2\-1/2?
### 5\.2Evaluation A: checkpoint\-matched channels
The design follows the checkpoint construction ofWanget al\.\([2026b](https://arxiv.org/html/2608.07914#bib.bib31)\): members from a window before checkpoint steptt, nonmembers from a matched unprocessed window, balanced on source, length, packing, deduplication, and stream position, blind text\-only baseline first\. The completed runs use Pythia\(Bidermanet al\.,[2023](https://arxiv.org/html/2608.07914#bib.bib1)\)and GPT\-Neo; in the six\-channel audit the member and nonmember pools come from*different corpora*, so all six fail the blind gate \(§[6\.4](https://arxiv.org/html/2608.07914#S6.SS4)\)—the gate working as designed\. OLMo\(Groeneveldet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib15)\)channels are planned, not run\.
For each held\-out channel, we create aggregate audit sets by sampling a known fractionα\\alphaof members and1−α1\-\\alphanonmembers\. This tests the statistical audit conditional on a real model channel; it does*not*claim the post\-hoc mixture retrains the model or validates no\-spillover—that causal question is Evaluation B’s\.
#### Gate\-passing channels \(v3\)\.
A third audit randomly splits one corpus \(700\+700700\{\+\}700\) and continues pretraining each base model on the member half: pools text\-exchangeable by construction, membership causally induced, with*direct*power trials at the frozen budgets \(all details in[table˜3](https://arxiv.org/html/2608.07914#A11.T3)\)\.
#### Locked external validation\.
A locked validation on the open 25\-model release ofZarzeckiet al\.\([2026](https://arxiv.org/html/2608.07914#bib.bib34)\)—efficacy calibrated without their audit outcomes predicting their documented failures, budgets, and coverage, no score selection or recalibration—is a planned extension, not a completed result; the Zarzecki transport is not yet run\.
### 5\.3Evaluation B: controlled contamination injection
We distinguish two contamination fractions that are easy to conflate: the*training injection fraction*βtrain\\beta\_\{\\rm train\}\(items injected during continuation over benchmark size\), a causal intervention whose variation requires independently trained models, and the*audit exposed fraction*αaudit\\alpha\_\{\\rm audit\}\(exposed items over audit\-set size\), which can be varied by post\-hoc remixing conditional on one trained model\. Remixingαaudit\\alpha\_\{\\rm audit\}measures audit power, not the causal effect of more training contamination; they coincide only when the entire benchmark is audited\.
Starting from the clean EleutherAI/pythia\-160m checkpoint \(Pile\-trained; see the provenance caveat below\), we continue training for a fixed token and optimizer budget \(400400AdamW steps, learning rate5×10−55\\times 10^\{\-5\}\)\. We inject aβtrain\\beta\_\{\\rm train\}fraction of SQuAD benchmark items into a Pile background corpus and audit against a disjoint held\-out pool of the same benchmark; the main injected condition uses200200injected and200200held\-out items\. Total updates, token count, and order randomization are matched to a clean \(βtrain=0\\beta\_\{\\rm train\}=0\) run per training seed\.
We cross the injection with four mechanisms atk=4k=4copies per injected item: exact question–answer text; a*validated paraphrase*\(answer preserved, overlap\-capped, dual independent NLI judges; acceptance138/200=0\.69138/200=0\.69; pipeline in[section˜G\.2](https://arxiv.org/html/2608.07914#A7.SS2)\); a word\-order shuffle \(a*surface perturbation*\); and answer/rationale\-only exposure\. Auditing both exposed and held\-out items makes spillover directly observable\. The upgraded design runs*five*paired training seeds, each with a matched clean \(βtrain=0\\beta\_\{\\rm train\}\{=\}0\) continuation \(same background, order, steps, optimizer, initialization\); the causal readout is the per\-seed*paired contrast*DsD\_\{s\}= \(exposed−\-held gap, contaminated\)−\-\(same gap, clean\) in clean\-model units, with item\-level bootstrap intervals\. Larger exposure counts remain future work\. “Pile\-trained” does not certify the base checkpoint never saw SQuAD, so we claim only the*incremental*exposure effect of the paired continuations—a well\-defined causal quantity even with prior exposure \(full caveat in[section˜G\.3](https://arxiv.org/html/2608.07914#A7.SS3)\)\.
### 5\.4Probes and statistical protocol
All reported experiments evaluate five gray\-box probe families \(mean loss, zlib\-normalized loss, Min\-K% Prob, Min\-K%\+\+, neighborhood loss\); every “best probe” is*the probe with the highest calibration\-split efficacy*, fixed before audit power is computed\. ReCaLL, CAMIA\-style, black\-box, and cross\-fitted\-combination probes are*specified but not run*\([appendix˜G](https://arxiv.org/html/2608.07914#A7)\); no result depends on them\. The nominal confirmatory size isτ=0\.05\\tau=0\.05, the certificate miscoverage isδcert=0\.05\\delta\_\{\\mathrm\{cert\}\}=0\.05, and the calibration\-floor failure probability isδcal=0\.05\\delta\_\{\\mathrm\{cal\}\}=0\.05; these are distinct guarantees\. The target power is0\.800\.80\. Probe counts arem∈\{16,…,1024\}m\\in\\\{16,\\dots,1024\\\}truncated at the number of independent item clusters, with a fresh clean reference ofn0test=250n\_\{0\}^\{\\rm test\}=250items andmeffm\_\{\\rm eff\}reported per test\. Monte\-Carlo replication counts, interval conventions, onset\-fitting rules, and the Holm\-adjusted secondary analyses are specified in[appendix˜G](https://arxiv.org/html/2608.07914#A7); the transport decision uses the predeclared criteria of[section˜4\.4](https://arxiv.org/html/2608.07914#S4.SS4)\.
## 6Results
### 6\.1Calibration predicts held\-out power
The initial audit freezese^f\\widehat\{e\}\_\{f\}on the calibration split of one500\+500500\{\+\}500Pythia\-1\.4b channel \(R2=0\.81R^\{2\}=0\.81–0\.990\.99, MAE0\.020\.02–0\.100\.10;[tables˜7](https://arxiv.org/html/2608.07914#A11.T7)and[10](https://arxiv.org/html/2608.07914#A11.T10); no fitted recalibration on the audit split\)\. The primary evidence is an upgraded audit using the*exact*Monte\-Carlo permutation test \(n0=250n\_\{0\}\{=\}250reference, audits drawn without replacement\) over six model×\\timescorpus channels \(pythia\-410m/1\.4b, GPT\-Neo\-1\.3B; news and Wikipedia pools of1,600\+1,6001\{,\}600\{\+\}1\{,\}600\)\.[Figure˜1](https://arxiv.org/html/2608.07914#S6.F1)shows every predicted\-vs\-observed cell with Wilson intervals\. Per channel \(best probe\):R2=0\.83R^\{2\}=0\.83–0\.980\.98, pooled MAE0\.030\.03–0\.110\.11, slope0\.910\.91–1\.041\.04, intercept−0\.09\-0\.09–0\.060\.06,0\.80\.8\-boundary accuracy0\.950\.95–1\.001\.00; in the informative0\.20\.2–0\.90\.9band MAE grows to0\.060\.06–0\.250\.25—mid\-band prediction is real but coarser\. Empirical size atα=0\\alpha\{=\}0is0\.0090\.009–0\.0450\.045\([table˜2](https://arxiv.org/html/2608.07914#A11.T2)\): valid, somewhat conservative\. Scope, per our own protocol: these six channels fail the blind gate \(§[6\.4](https://arxiv.org/html/2608.07914#S6.SS4)\), so theirR2R^\{2\}validates power\-curve prediction*within a content\-confounded channel*; the confirmatory reading rests on the gate\-passing v3 channels, and accurate curves do*not*license plug\-in budgets \(§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\); values in the[fig\.˜1](https://arxiv.org/html/2608.07914#S6.F1)caption\.
Figure 1:Upgraded six\-channel audit, exact permutation test\. Left: predicted vs\. observed power, every \(channel, probe,α\\alpha,mm\) cell, Wilson intervals\. Right:mpredm\_\{\\rm pred\}vs\. measuredmobsm\_\{\\rm obs\}atα0=0\.2\\alpha\_\{0\}\{=\}0\.2\(log–log,2×2\\timesband\); the dyadic\-grid onsets are*interval\-censored upper bounds*, so the apparent1\.51\.5–3\.6×3\.6\\timesratios are not resolved onset ratios \(two of three cases are unresolved; §[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\)—the censoring\-free diagnosis comes from the v3 direct trials\.
### 6\.2Plug\-in budgets fail: the efficacy\-only planner is miscalibrated at smallmm
Accurate power*curves*do not imply usable plug\-in*budgets*\. The pre\-specified criterion \(§[4\.4](https://arxiv.org/html/2608.07914#S4.SS4)\) fails on the initial channel:[table˜10](https://arxiv.org/html/2608.07914#A11.T10)reports*efficacy\-transport ratios*exp\(2\|loge^fcal−loge^ftest\|\)\\exp\(2\|\\log\\widehat\{e\}\_\{f\}^\{\\rm cal\}\-\\log\\widehat\{e\}\_\{f\}^\{\\rm test\}\|\)with median2\.23×2\.23\\timesand3/53/5channels above the2×2\\timestolerance, sowe withdraw the pooled calibration\-to\-audit scaling claim\(tolerance fixed before the numbers were seen\)\. But with800800\-item calibration halves the ratios shrink to1\.031\.03–1\.501\.50, and in the v3 gate\-passing channels the*raw*efficacy ratios are1\.021\.02–1\.351\.35\(1\.041\.04–1\.811\.81squared to the budget scalem∝ef−2m\\propto e\_\{f\}^\{\-2\}\)—within the2×2\\timestolerance\. Efficacy itself largely transports; that is*not*what breaks the budgets\.
The decisive diagnosis comes from the v3 censoring\-free direct trials \([table˜3](https://arxiv.org/html/2608.07914#A11.T3); the earlier dyadic\-grid “0/30/3” claim is withdrawn as interval\-censored, arithmetic in[table˜2](https://arxiv.org/html/2608.07914#A11.T2)\)\. At the frozenmpredm\_\{\\rm pred\}the planner predicts0\.800\.80by construction; observed power is0\.670\.67,0\.560\.56,0\.620\.62, Wilson intervals excluding0\.800\.80in all three channels; re\-running the*same*plug\-in with the audit\-split efficacy still predicts0\.580\.58,0\.770\.77,0\.680\.68—errors of both signs up to0\.210\.21that dwarf the efficacy shift \(on pythia\-410m×\\timeswiki the efficacies nearly coincide,4\.624\.62/4\.314\.31, yet the plug\-in predicts0\.770\.77against0\.560\.56\)\. The failure isnotcalibration\-to\-audit transport; it is*consistent with*the local Gaussian approximation breaking at the small budgets it prescribes \(m=8m=8–4242\) plus a conservative exact test \(size0\.0000\.000–0\.0130\.013\), but the causal decomposition is not isolated \(no skewness/tie or Gaussian\-vs\-Welch\-vs\-permutation ablation was run; open\)\. Atminfl=1\.40×mpredm\_\{\\rm infl\}=1\.40\\times m\_\{\\rm pred\}\(frozen before scoring; uncertainty not propagated, disclosed\): two clear failures, one unresolved \(0\.760\.76\[0\.70,0\.81\]\[0\.70,0\.81\]\)\.
The constructive repair is*executed at scale*\. Afinite\-sample plannerresamples fullP0/P1P\_\{0\}/P\_\{1\}score laws from calibration data only and runs the deployed permutation test at each candidatemm\. The v4 pilot \([table˜4](https://arxiv.org/html/2608.07914#A11.T4)\) beat the Gaussian budget head\-to\-head but chosemmfrom pointwise bootstrap bounds and retained only two channels\. The v5 protocol repairs both defects \([table˜5](https://arxiv.org/html/2608.07914#A11.T5)\):1212same\-corpus channels, an equivalence blind gate \(9/129/12pass the predeclared one\-sided form, AUC upper CI≤0\.55\\leq 0\.55;8/128/12pass the stricter two\-sided sensitivity gate, which additionally excludes exactly the channel where the planner abstains; per\-channel values in[table˜6](https://arxiv.org/html/2608.07914#A11.T6)\), a two\-stage planner \(stage\-A bootstrap candidatemAm\_\{A\}on one calibration half; stage\-B confirmation on the untouched half over the pre\-specified ladder\{mA,⌈1\.5mA⌉,2mA\}\\\{m\_\{A\},\\lceil 1\.5m\_\{A\}\\rceil,2m\_\{A\}\\\}with Bonferroni\-corrected one\-sided bounds\), and theprespecified transport endpointon the audit half \(estimated onsetm^⋆\\widehat\{m\}\_\{\\star\}, multiplicative errors,90%90\\%bootstrap quantile intervals\)\.*Validity scope, stated exactly:*the Bonferroni step makes stage B valid against*rung selection*conditional on the empirical cal\-B distribution \(175175clean clusters resampled with replacement ton0=250n\_\{0\}\{=\}250\), excluding score\-law uncertainty, cluster reuse, and cal\-to\-audit transport \(measured by the audit endpoint, not certified\);m^⋆\\widehat\{m\}\_\{\\star\}is a Monte\-Carlo onset estimate \(200200trials\), the90%90\\%intervals are descriptive \(B=30B\{=\}30, probe selection not re\-run\), the ratios and8/88/8coverage inherit that noise \(π^\(mplan\)=0\.785\\widehat\{\\pi\}\(m\_\{\\rm plan\}\)=0\.785atmplan=m^⋆=56m\_\{\\rm plan\}=\\widehat\{m\}\_\{\\star\}=56on pythia\-70m×\\timescnn\), and the prespecified≥80%\\geq 80\\%interval\-coverage criterion is*unevaluable as specified*and not claimed\. Results: the Gaussian budget*clearly fails in9/99/9*channels \(power0\.010\.01–0\.740\.74atmpredm\_\{\\rm pred\}, every Wilson interval excluding0\.800\.80;mpred/m^⋆=0\.54m\_\{\\rm pred\}/\\widehat\{m\}\_\{\\star\}=0\.54–0\.860\.86, median0\.590\.59\); the planner is uniformly conservative \(mplan/m^⋆=1\.00m\_\{\\rm plan\}/\\widehat\{m\}\_\{\\star\}=1\.00–2\.002\.00, median1\.361\.36\), delivering0\.7850\.785–0\.920\.92power \(≥0\.80\\geq 0\.80in6/86/8, two unresolved, zero clear failures;90%90\\%intervals coverm^⋆\\widehat\{m\}\_\{\\star\}in8/88/8\)\. On pythia\-70m×\\timeswiki the planner*abstains*\(stage A never certifies0\.800\.80\) while the Gaussian budget prescribesm=22m\{=\}22at0\.010\.01power—abstention is the designed behavior, and the two\-sided gate excludes this channel outright\. The certificate \(below\) makes no planning assumption and the injection ordering \(§[6\.4](https://arxiv.org/html/2608.07914#S6.SS4)\) stands\.
### 6\.3Certificate coverage: valid, and vacuous at audit scale
The scaled\-slack heuristic’s99\.099\.0–99\.5%99\.5\\%endpoints \([table˜9](https://arxiv.org/html/2608.07914#A11.T9)\) are not validation of the distribution\-free certificate, which never fires atm=512m\{=\}512\(zero false, zero nonzero in all six channels; coverage vacuously100%100\\%\): certifiedα=0\.1\\alpha\{=\}0\.1–0\.20\.2needs balanced groups of∼3,500\\sim 3\{,\}500–79,00079\{,\}000items at measured margins\. A nonzeroα¯\\underline\{\\alpha\}is trustworthy; a zero is uninformative—at audit sizes the certificate is a null instrument, reported as a limitation \(full analysis:[appendices˜I](https://arxiv.org/html/2608.07914#A9),[J](https://arxiv.org/html/2608.07914#A10)and[H](https://arxiv.org/html/2608.07914#A8)\)\.
### 6\.4Mechanism, access, and validity diagnostics
The initial three\-seed injections already show the mechanism decay \(0\.7410\.741/0\.3430\.343/0\.1500\.150exact/surface/answer\-only,3/33/3seeds;[table˜9](https://arxiv.org/html/2608.07914#A11.T9)\)\.
The five\-seed*paired\-contrast*rerun strengthens this, using*one common frozen probe*\(mean NLL, the modal calibration winner\) since the per\-cell best probe varies by seed and mechanism\. The contrastsDsD\_\{s\}\(item\-bootstrap95%95\\%intervals excluding zero in all1515seed×\\timesmechanism cells\) are1\.01±0\.071\.01\\pm 0\.07exact,0\.70±0\.090\.70\\pm 0\.09paraphrase,0\.44±0\.040\.44\\pm 0\.04surface, ordering intact in5/55/5seeds; answer\-only hasDs=−0\.01±0\.07D\_\{s\}=\-0\.01\\pm 0\.07with every per\-seed interval covering zero—its apparent raw efficacy \(0\.190\.19\) is matched by the clean continuation \(baseline drift, the artifact the paired design exists to catch\), and it is where exchangeability degrades \(pooled size0\.0870\.087vs\.0\.0130\.013–0\.0580\.058\)\. We claim only “indistinguishable from clean continuation atn=200n\{=\}200,k=4k\{=\}4” for answer\-only; the exact certificate at injection scale never fires \(zero false, zero nonzero, coverage1\.01\.0\), matching[appendix˜J](https://arxiv.org/html/2608.07914#A10)\.
Every per\-seedDsD\_\{s\}, interval, size\-gate value, and adjacent paired difference is in[table˜1](https://arxiv.org/html/2608.07914#A11.T1); on the common probe all three adjacent differences \(0\.31±0\.120\.31\\pm 0\.12,0\.26±0\.100\.26\\pm 0\.10,0\.45±0\.080\.45\\pm 0\.08\) are positive in5/55/5seeds \(sign\-testp=0\.031p=0\.031each; Holmpadj=0\.094p\_\{\\rm adj\}=0\.094: a consistent descriptive pattern, not individually significant after correction, and—as the arms are not dose\-matched—not a causal mechanism comparison\)\. The exposure\-effect claim is scoped to pythia\-160m, SQuAD,k=4k\{=\}4, this protocol\.
The exact\-injection channel also compares plug\-in predictions against observed0\.80\.8\-power onsets: per seedmpredm\_\{\\rm pred\}of371371,208208,301301vs\.mobsm\_\{\\rm obs\}of512512,256256,512512\(dyadic grid, quantized upward\), within the2×2\\timestolerance in all three seeds; surface and answer\-only never reach0\.80\.8power on the tested grid—censored, not extrapolated\. Efficacies are estimates or lower bounds onρ\\rho; the small\-sample floor vacuity atn=200n\{=\}200is in[table˜13](https://arxiv.org/html/2608.07914#A11.T13)\.
The validity gates are measured per channel \([table˜2](https://arxiv.org/html/2608.07914#A11.T2)\): every checkpoint\-free Evaluation A channel*fails*blind separation \(text\-only AUC0\.950\.95—content\-confounded separability, exactly what the gate exposes\), while the v3 same\-corpus channels repair it \(blind AUC0\.480\.48–0\.540\.54;4/64/6pass all gates, the two Pythia×\\timesSQuAD channels failing the size gate,[table˜3](https://arxiv.org/html/2608.07914#A11.T3)\)\. The v3 rule is not an equivalence test; requiring upper CI≤0\.55\\leq 0\.55also excludes pythia\-410m×\\timeswiki \(upper0\.5790\.579\), leaving the small\-mmconclusion on the two surviving channels \(0\.670\.67,0\.620\.62, both excluding0\.800\.80\)\. Pure\-memorization claims rest on the gate\-passing v3 channels and the Evaluation B injections\.
## 7Related Work
Behavioral detectors infer exposure from likelihoods, perturbations, output order, or elicited recall\(Matternet al\.,[2023](https://arxiv.org/html/2608.07914#bib.bib21); Shiet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib29); Orenet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib23),*inter alia*\), and evaluations of them report a reliability gap rather than a ranking\(Duanet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib12); Daset al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib7); Zarzeckiet al\.,[2026](https://arxiv.org/html/2608.07914#bib.bib34)\): they establish*that*audits fail, not a quantity—measurable before an audit—saying when they will\. Dataset\-level methods aggregate item scores intopp\-values, inflation estimates, or FDR\-controlled decisions\(Mainiet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib20); Dekonincket al\.,[2024b](https://arxiv.org/html/2608.07914#bib.bib9); Zhanget al\.,[2026](https://arxiv.org/html/2608.07914#bib.bib37)\); none states whichα\\alphathe audit had power to detect\. Theαρm\\alpha\\rho\\sqrt\{m\}boundary and prevalence estimation are classical\(Ingster,[1997](https://arxiv.org/html/2608.07914#bib.bib16); Cai and Wu,[2014](https://arxiv.org/html/2608.07914#bib.bib4); Scott,[2015](https://arxiv.org/html/2608.07914#bib.bib28)\); we claim audit deliverables, not new mixture theory\. Appendix[N](https://arxiv.org/html/2608.07914#A14)gives the full treatment, including a per\-method comparison of estimands \(Table[15](https://arxiv.org/html/2608.07914#A14.T15)\)\.
## 8Discussion and Conclusion
Frozen calibration efficacy predicts held\-out audit power \(R2=0\.83R^\{2\}=0\.83–0\.980\.98\), but its Gaussian inversion into a budget does not: the plug\-in fails in9/99/9gate\-passing channels \(mpred/m^⋆=0\.54m\_\{\\mathrm\{pred\}\}/\\widehat\{m\}\_\{\\star\}=0\.54–0\.860\.86\) even though efficacy itself transports \(ratios1\.021\.02–1\.351\.35\), so the culprit is the local approximation at the smallmmit prescribes, not calibration drift\. Simulating the deployed test instead yields budgets that are uniformly conservative, deliver0\.7850\.785–0\.920\.92power, and abstain where the probe fails to transport\. A non\-rejection is thus interpretable only with its efficacy, budget, and gate values attached—and the prevalence certificate, though valid, is vacuous at audit scale\.
## Limitations
The mixture channel is an auditable modeling assumption, not a universal law of training\. Contaminated examples can affect unexposed examples, and post\-training can alter the clean and seen channels\. Our stability diagnostics can reveal large violations but cannot prove that a clean reference is perfectly transportable\. The lower bound is pointwise in a chosen behavioral access channel; it does not apply to direct corpus search, cryptographic provenance, or support\-separated watermark evidence\. The matching optimality result is local and assumes finite chi\-square divergence and a mild moment condition\. The distribution\-free prevalence certificate requires independent matched controls and a score learned on separate data; using the audit set to tune the score invalidates its coverage\. Such controls are generally unavailable to an external auditor of a closed model, so neither efficacy nor the exposed fraction is then identified without extra assumptions\. Finally, contamination encompasses more than verbatim membership\. Semantic task exposure, answer\-only exposure, and distillation may require different labels and channels, so results should not be transported between mechanisms without validation\. Six specific gaps remain open and unexecuted: the v5 planner’s stage\-B guarantee is conditional on the empirical cal\-B distribution \(a population\-level selection\-valid planner would need to propagate score\-law uncertainty and cluster reuse, e\.g\. by nesting the bootstrap over probe selection with a larger replicate budget\); the prespecified interval\-coverage transport criterion is unevaluable with descriptive intervals and is not claimed; the as\-run size gate has the detection rather than equivalence direction \(the equivalence form leaves4/124/12channels,[table˜6](https://arxiv.org/html/2608.07914#A11.T6)\); the cause of Gaussian miscalibration is consistent with small\-mmnon\-Gaussianity plus test conservatism but not isolated by ablation; the mechanism ordering is not dose\-matched \(a causal comparison requires retraining all arms on the138138\-item paraphrase intersection with copy\- and token\-matched exposure\); and no locked external or modern\-model validation \(OLMo\-class models, black\-box access, the Zarzecki study, or matched\-counterpart designs such as LLM Dataset Inference or PaCoST\) has been run\. The certificate remains a theoretical contribution: mathematically valid but firing on no reported condition\. Where these gaps would change a claim, the claim is scoped accordingly\.
## Ethical Considerations
Contamination audits can improve scientific accountability, but membership signals can also expose whether copyrighted, private, or sensitive text was used for training\. Experiments should use licensed or public data, avoid releasing reconstructable private examples, and report aggregate statistics\. The proposed lower bound must not be used to imply that undetected contamination is acceptable\. Controlled contamination runs should be confined to research models and should not publish contaminated checkpoints as general\-purpose models without prominent documentation\. All experimental results, compute use, data licenses, and automated writing or coding assistance must be disclosed in the venue’s Responsible NLP checklist\.
## References
- S\. Biderman, H\. Schoelkopf, Q\. Anthony, H\. Bradley, K\. O’Brien, E\. Hallahan, M\. A\. Khan, S\. Purohit, U\. S\. Prashanth, E\. Raff, A\. Skowron, L\. Sutawika, and O\. Van Der Wal \(2023\)Pythia: a suite for analyzing large language models across training and scaling\.InProceedings of the 40th International Conference on Machine Learning \(ICML\),Cited by:[§5\.2](https://arxiv.org/html/2608.07914#S5.SS2.p1.1)\.
- G\. Blanchard, G\. Lee, and C\. Scott \(2010\)Semi\-supervised novelty detection\.Journal of Machine Learning Research11,pp\. 2973–3009\.Cited by:[§N\.6](https://arxiv.org/html/2608.07914#A14.SS6.p1.1)\.
- T\. T\. Cai, X\. J\. Jeng, and J\. Jin \(2011\)Optimal detection of heterogeneous and heteroscedastic mixtures\.Journal of the Royal Statistical Society: Series B73\(5\),pp\. 629–662\.Cited by:[§N\.5](https://arxiv.org/html/2608.07914#A14.SS5.p1.1),[item 1](https://arxiv.org/html/2608.07914#S1.I1.i1.p1.1)\.
- T\. T\. Cai and Y\. Wu \(2014\)Optimal detection of sparse mixtures against a given null distribution\.IEEE Transactions on Information Theory60\(4\),pp\. 2217–2232\.Cited by:[§N\.5](https://arxiv.org/html/2608.07914#A14.SS5.p1.1),[item 1](https://arxiv.org/html/2608.07914#S1.I1.i1.p1.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- N\. Carlini, F\. Tramèr, E\. Wallace, M\. Jagielski, A\. Herbert\-Voss, K\. Lee, A\. Roberts, T\. Brown, D\. Song, U\. Erlingsson, A\. Oprea, and C\. Raffel \(2021\)Extracting training data from large language models\.In30th USENIX Security Symposium,pp\. 2633–2650\.Cited by:[§N\.1](https://arxiv.org/html/2608.07914#A14.SS1.p1.1),[§G\.1](https://arxiv.org/html/2608.07914#A7.SS1.SSS0.Px1.p1.1)\.
- H\. Chang, A\. Shahin Shamsabadi, K\. Katevas, H\. Haddadi, and R\. Shokri \(2025\)Context\-aware membership inference attacks against pre\-trained language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Cited by:[§N\.1](https://arxiv.org/html/2608.07914#A14.SS1.p1.1),[§G\.1](https://arxiv.org/html/2608.07914#A7.SS1.SSS0.Px2.p1.1)\.
- D\. Das, J\. Zhang, and F\. Tramèr \(2025\)Blind baselines beat membership inference attacks for foundation models\.arXiv preprint arXiv:2406\.16201\.Cited by:[§N\.4](https://arxiv.org/html/2608.07914#A14.SS4.p1.1),[item 1](https://arxiv.org/html/2608.07914#A6.I1.i1.p1.8),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- J\. Dekoninck, M\. N\. Müller, M\. Baader, M\. Fischer, and M\. Vechev \(2024a\)Evading data contamination detection for language models is \(too\) easy\.arXiv preprint arXiv:2402\.02823\.Cited by:[§N\.2](https://arxiv.org/html/2608.07914#A14.SS2.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1)\.
- J\. Dekoninck, M\. N\. Müller, and M\. Vechev \(2024b\)ConStat: performance\-based contamination detection in large language models\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§N\.3](https://arxiv.org/html/2608.07914#A14.SS3.p1.1),[Table 15](https://arxiv.org/html/2608.07914#A14.T15.6.9.3.1),[§1](https://arxiv.org/html/2608.07914#S1.p2.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- Y\. Dong, X\. Jiang, H\. Liu, Z\. Jin, B\. Gu, M\. Yang, and G\. Li \(2024\)Generalization or memorization: data contamination and trustworthy evaluation for large language models\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 12039–12050\.Cited by:[§1](https://arxiv.org/html/2608.07914#S1.p1.1)\.
- D\. Donoho and J\. Jin \(2004\)Higher criticism for detecting sparse heterogeneous mixtures\.The Annals of Statistics32\(3\),pp\. 962–994\.Cited by:[§N\.5](https://arxiv.org/html/2608.07914#A14.SS5.p1.1),[item 1](https://arxiv.org/html/2608.07914#S1.I1.i1.p1.1)\.
- M\. Duan, A\. Suri, N\. Mireshghallah, S\. Min, W\. Shi, L\. Zettlemoyer, Y\. Tsvetkov, Y\. Choi, D\. Evans, and H\. Hajishirzi \(2024\)Do membership inference attacks work on large language models?\.InProceedings of the Conference on Language Modeling \(COLM\),Cited by:[§N\.2](https://arxiv.org/html/2608.07914#A14.SS2.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- Y\. Fu, Ö\. Uzuner, M\. Yetisgen, and F\. Xia \(2025\)Does data contamination detection work \(well\) for LLMs? a survey and evaluation on detection assumptions\.InFindings of the Association for Computational Linguistics: NAACL 2025,Cited by:[§N\.2](https://arxiv.org/html/2608.07914#A14.SS2.p1.1),[§N\.4](https://arxiv.org/html/2608.07914#A14.SS4.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1)\.
- S\. Golchin and M\. Surdeanu \(2025\)Data contamination quiz: a tool to detect and estimate contamination in large language models\.Transactions of the Association for Computational Linguistics13,pp\. 809–830\.Cited by:[§N\.1](https://arxiv.org/html/2608.07914#A14.SS1.p1.1),[Table 15](https://arxiv.org/html/2608.07914#A14.T15.6.11.5.1),[§G\.1](https://arxiv.org/html/2608.07914#A7.SS1.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1)\.
- D\. Groeneveld, I\. Beltagy, E\. Walsh, A\. Bhagia, R\. Kinney, O\. Tafjord, A\. Jha, H\. Ivison, I\. Magnusson, Y\. Wang, S\. Arora, D\. Atkinson, R\. Authur, K\. Chandu, A\. Cohan, J\. Dumas, Y\. Elazar, Y\. Gu, J\. Hessel, T\. Khot, W\. Merrill, J\. Morrison, N\. Muennighoff, A\. Naik, C\. Nam, M\. Peters, V\. Pyatkin, A\. Ravichander, D\. Schwenk, S\. Shah, W\. Smith, E\. Strubell, N\. Subramani, M\. Wortsman, P\. Dasigi, N\. Lambert, K\. Richardson, L\. Zettlemoyer, J\. Dodge, K\. Lo, L\. Soldaini, N\. Smith, and H\. Hajishirzi \(2024\)OLMo: accelerating the science of language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15789–15809\.External Links:[Link](https://aclanthology.org/2024.acl-long.841/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.841)Cited by:[§5\.2](https://arxiv.org/html/2608.07914#S5.SS2.p1.1)\.
- Y\. I\. Ingster \(1997\)Some problems of hypothesis testing leading to infinitely divisible distributions\.Mathematical Methods of Statistics6,pp\. 47–69\.Cited by:[§N\.5](https://arxiv.org/html/2608.07914#A14.SS5.p1.1),[item 1](https://arxiv.org/html/2608.07914#S1.I1.i1.p1.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- L\. Le Cam \(1986\)Asymptotic methods in statistical decision theory\.Springer\.Cited by:[§N\.5](https://arxiv.org/html/2608.07914#A14.SS5.p1.1)\.
- I\. Magar and R\. Schwartz \(2022\)Data contamination: from memorization to exploitation\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 1578–1596\.Cited by:[§1](https://arxiv.org/html/2608.07914#S1.p1.1)\.
- P\. Maini, H\. Jia, N\. Papernot, and A\. Dziedzic \(2024\)LLM dataset inference: did you train on my dataset?\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§N\.3](https://arxiv.org/html/2608.07914#A14.SS3.p1.1),[Table 15](https://arxiv.org/html/2608.07914#A14.T15.1.1.2),[§1](https://arxiv.org/html/2608.07914#S1.p2.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- J\. Mattern, F\. Mireshghallah, Z\. Jin, B\. Schölkopf, M\. Sachan, and T\. Berg\-Kirkpatrick \(2023\)Membership inference attacks against language models via neighbourhood comparison\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 11330–11343\.Cited by:[§N\.1](https://arxiv.org/html/2608.07914#A14.SS1.p1.1),[§G\.1](https://arxiv.org/html/2608.07914#A7.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- A\. Maurer and M\. Pontil \(2009\)Empirical bernstein bounds and sample variance penalization\.InProceedings of the 22nd Annual Conference on Learning Theory \(COLT\),Cited by:[§A\.6](https://arxiv.org/html/2608.07914#A1.SS6.p1.12),[Appendix C](https://arxiv.org/html/2608.07914#A3.p1.5),[§3\.3](https://arxiv.org/html/2608.07914#S3.SS3.p3.3)\.
- M\. Meeus, I\. Shilov, S\. Jain, M\. Faysse, M\. Rei, and Y\. de Montjoye \(2025\)SoK: membership inference attacks on LLMs are rushing nowhere \(and how to fix it\)\.InIEEE Conference on Secure and Trustworthy Machine Learning \(SaTML\),pp\. 385–401\.Cited by:[§N\.2](https://arxiv.org/html/2608.07914#A14.SS2.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1)\.
- Y\. Oren, N\. Meister, N\. S\. Chatterji, F\. Ladhak, and T\. B\. Hashimoto \(2024\)Proving test set contamination in black\-box language models\.InProceedings of the International Conference on Learning Representations \(ICLR\),Cited by:[§N\.1](https://arxiv.org/html/2608.07914#A14.SS1.p1.1),[Table 15](https://arxiv.org/html/2608.07914#A14.T15.3.3.2),[§G\.1](https://arxiv.org/html/2608.07914#A7.SS1.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- H\. Puerto, M\. Gubri, S\. Yun, and S\. J\. Oh \(2025\)Scaling up membership inference: when and how attacks succeed on large language models\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 4165–4182\.Cited by:[§N\.3](https://arxiv.org/html/2608.07914#A14.SS3.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p2.1)\.
- H\. Ramaswamy, C\. Scott, and A\. Tewari \(2016\)Mixture proportion estimation via kernel embeddings of distributions\.InProceedings of the 33rd International Conference on Machine Learning \(ICML\),pp\. 2052–2060\.Cited by:[§N\.6](https://arxiv.org/html/2608.07914#A14.SS6.p1.1)\.
- W\. J\. Rogan and B\. Gladen \(1978\)Estimating prevalence from the results of a screening test\.American Journal of Epidemiology107\(1\),pp\. 71–76\.Cited by:[§N\.6](https://arxiv.org/html/2608.07914#A14.SS6.p1.1)\.
- V\. Samuel, Y\. Zhou, and H\. P\. Zou \(2025\)Towards data contamination detection for modern large language models: limitations, inconsistencies, and oracle challenges\.InProceedings of the 31st International Conference on Computational Linguistics \(COLING\),Cited by:[§N\.2](https://arxiv.org/html/2608.07914#A14.SS2.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1)\.
- C\. Scott \(2015\)A rate of convergence for mixture proportion estimation, with application to learning from noisy labels\.InProceedings of the 18th International Conference on Artificial Intelligence and Statistics \(AISTATS\),pp\. 838–846\.Cited by:[§N\.6](https://arxiv.org/html/2608.07914#A14.SS6.p1.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- W\. Shi, A\. Ajith, M\. Xia, Y\. Huang, D\. Liu, T\. Blevins, D\. Chen, and L\. Zettlemoyer \(2024\)Detecting pretraining data from large language models\.InProceedings of the International Conference on Learning Representations \(ICLR\),Cited by:[§N\.1](https://arxiv.org/html/2608.07914#A14.SS1.p1.1),[§G\.1](https://arxiv.org/html/2608.07914#A7.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- A\. W\. van der Vaart \(1998\)Asymptotic statistics\.Cambridge University Press\.Cited by:[§N\.5](https://arxiv.org/html/2608.07914#A14.SS5.p1.1)\.
- H\. Wang, H\. Li, B\. Ko, and H\. Zhang \(2026a\)On the fragility of benchmark contamination detection in reasoning models\.arXiv preprint arXiv:2510\.02386\.Cited by:[§N\.2](https://arxiv.org/html/2608.07914#A14.SS2.p1.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1)\.
- J\. G\. Wang, J\. Wang, M\. Li, and S\. Neel \(2026b\)CheckMIABench: firm foundations for membership inference attacks on language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(ACL, Short Papers\),pp\. 364–370\.Cited by:[§N\.4](https://arxiv.org/html/2608.07914#A14.SS4.p1.1),[item 1](https://arxiv.org/html/2608.07914#A6.I1.i1.p1.8),[§5\.2](https://arxiv.org/html/2608.07914#S5.SS2.p1.1)\.
- R\. Xie, J\. Wang, R\. Huang, M\. Zhang, R\. Ge, J\. Pei, N\. Z\. Gong, and B\. Dhingra \(2024\)ReCaLL: membership inference via relative conditional log\-likelihoods\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 8671–8689\.Cited by:[§N\.1](https://arxiv.org/html/2608.07914#A14.SS1.p1.1),[§G\.1](https://arxiv.org/html/2608.07914#A7.SS1.SSS0.Px2.p1.1)\.
- W\. Zarzecki, J\. Dubiński, and S\. Cygert \(2026\)The reliability gap in benchmark auditing: distribution shift and scale as failure modes of contamination detection\.InProceedings of the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases \(ECML PKDD\),Cited by:[§N\.2](https://arxiv.org/html/2608.07914#A14.SS2.p1.1),[§N\.2](https://arxiv.org/html/2608.07914#A14.SS2.p2.1),[§1](https://arxiv.org/html/2608.07914#S1.p1.1),[§5\.2](https://arxiv.org/html/2608.07914#S5.SS2.SSS0.Px2.p1.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
- H\. Zhang, Y\. Lin, and X\. Wan \(2024\)PaCoST: paired confidence significance testing for benchmark contamination detection in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 1794–1809\.Cited by:[§N\.3](https://arxiv.org/html/2608.07914#A14.SS3.p1.1),[Table 15](https://arxiv.org/html/2608.07914#A14.T15.2.2.2),[§1](https://arxiv.org/html/2608.07914#S1.p2.1)\.
- J\. Zhang, J\. Sun, E\. Yeats, Y\. Ouyang, M\. Kuo, J\. Zhang, H\. F\. Yang, and H\. Li \(2025\)Min\-K%\+\+: improved baseline for detecting pre\-training data from large language models\.InProceedings of the International Conference on Learning Representations \(ICLR\),Cited by:[§N\.1](https://arxiv.org/html/2608.07914#A14.SS1.p1.1),[§G\.1](https://arxiv.org/html/2608.07914#A7.SS1.SSS0.Px1.p1.1)\.
- Z\. Zhang, Q\. Liu, S\. Liang, N\. Li, Z\. Hu, W\. Gao, R\. Li, Z\. Huang, L\. Rutkowski, B\. Yu, and D\. Tao \(2026\)Controllable contamination detection for reliable LLM evaluation with statistical guarantees\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 30122–30143\.Cited by:[§N\.3](https://arxiv.org/html/2608.07914#A14.SS3.p1.1),[Table 15](https://arxiv.org/html/2608.07914#A14.T15.6.10.4.1),[§1](https://arxiv.org/html/2608.07914#S1.p2.1),[§7](https://arxiv.org/html/2608.07914#S7.p1.3)\.
## Appendix AProofs
### A\.1Proof of the finite\-sample power limit
Putg=r−1g=r\-1\. The one\-observation likelihood ratio isLα=1\+αgL\_\{\\alpha\}=1\+\\alpha g\. Because𝔼0g=0\\mathbb\{E\}\_\{0\}g=0and𝔼0g2=ρ2\\mathbb\{E\}\_\{0\}g^\{2\}=\\rho^\{2\},
𝔼0Lα2=1\+α2ρ2\.\\mathbb\{E\}\_\{0\}L\_\{\\alpha\}^\{2\}=1\+\\alpha^\{2\}\\rho^\{2\}\.The product likelihood ratio isLα\(m\)=∏i=1mLα\(Yi\)L\_\{\\alpha\}^\{\(m\)\}=\\prod\_\{i=1\}^\{m\}L\_\{\\alpha\}\(Y\_\{i\}\)\. Independence therefore gives
1\+χ2\(Qα⊗m∥P0⊗m\)\\displaystyle 1\+\\chi^\{2\}\(Q\_\{\\alpha\}^\{\\otimes m\}\\\|P\_\{0\}^\{\\otimes m\}\)=𝔼0\{Lα\(m\)\}2\\displaystyle=\\mathbb\{E\}\_\{0\}\\\{L\_\{\\alpha\}^\{\(m\)\}\\\}^\{2\}=∏i=1m𝔼0Lα\(Yi\)2\\displaystyle=\\prod\_\{i=1\}^\{m\}\\mathbb\{E\}\_\{0\}L\_\{\\alpha\}\(Y\_\{i\}\)^\{2\}=\(1\+α2ρ2\)m\.\\displaystyle=\(1\+\\alpha^\{2\}\\rho^\{2\}\)^\{m\}\.\(9\)
For two simple hypotheses, every randomized test satisfies
am\+bm≥1−TV\(P0⊗m,Qα⊗m\)\.a\_\{m\}\+b\_\{m\}\\geq 1\-\\operatorname\{TV\}\(P\_\{0\}^\{\\otimes m\},Q\_\{\\alpha\}^\{\\otimes m\}\)\.Indeed,
am\+bm\\displaystyle a\_\{m\}\+b\_\{m\}=1−\{𝔼Qα⊗mϕm−𝔼P0⊗mϕm\}\\displaystyle=1\-\\\{\\mathbb\{E\}\_\{Q\_\{\\alpha\}^\{\\otimes m\}\}\\phi\_\{m\}\-\\mathbb\{E\}\_\{P\_\{0\}^\{\\otimes m\}\}\\phi\_\{m\}\\\}≥1−TV\(P0⊗m,Qα⊗m\),\\displaystyle\\geq 1\-\\operatorname\{TV\}\(P\_\{0\}^\{\\otimes m\},Q\_\{\\alpha\}^\{\\otimes m\}\),because0≤ϕm≤10\\leq\\phi\_\{m\}\\leq 1\. Cauchy–Schwarz applied toTV\(P,Q\)=12𝔼Q\|dP/dQ−1\|\\operatorname\{TV\}\(P,Q\)=\\tfrac\{1\}\{2\}\\mathbb\{E\}\_\{Q\}\|\\,\\mathrm\{d\}P/\\,\\mathrm\{d\}Q\-1\|gives
TV\(P,Q\)≤12χ2\(P∥Q\)\.\\operatorname\{TV\}\(P,Q\)\\leq\\frac\{1\}\{2\}\\sqrt\{\\chi^\{2\}\(P\\\|Q\)\}\.Combining this inequality with[eq\.˜9](https://arxiv.org/html/2608.07914#A1.E9)and the trivial boundTV≤1\\operatorname\{TV\}\\leq 1proves[eq\.˜4](https://arxiv.org/html/2608.07914#S3.E4)\.
Finally,
πm\(α\)−am\\displaystyle\\pi\_\{m\}\(\\alpha\)\-a\_\{m\}=𝔼Qα⊗mϕm−𝔼P0⊗mϕm\\displaystyle=\\mathbb\{E\}\_\{Q\_\{\\alpha\}^\{\\otimes m\}\}\\phi\_\{m\}\-\\mathbb\{E\}\_\{P\_\{0\}^\{\\otimes m\}\}\\phi\_\{m\}≤TV\(P0⊗m,Qα⊗m\)\.\\displaystyle\\leq\\operatorname\{TV\}\(P\_\{0\}^\{\\otimes m\},Q\_\{\\alpha\}^\{\\otimes m\}\)\.Ifam≤τa\_\{m\}\\leq\\tau, substitution of the same total\-variation bound gives[eq\.˜5](https://arxiv.org/html/2608.07914#S3.E5)\. ∎
### A\.2Statement and proof for a fixed contaminated subset
###### Lemma 6\(A fixed unknown contaminated subset\)\.
Suppose exactlykkof themmaudited items are contaminated\. ForS⊆\[m\]S\\subseteq\[m\],\|S\|=k\|S\|=k, letPS=⨂i∈SP1⊗⨂i∉SP0P\_\{S\}=\\bigotimes\_\{i\\in S\}P\_\{1\}\\otimes\\bigotimes\_\{i\\notin S\}P\_\{0\}, and letb¯m,k\\bar\{b\}\_\{m,k\}be the type\-II error averaged uniformly over such subsets\. Then every test satisfies
am\+b¯m,k\\displaystyle a\_\{m\}\+\\bar\{b\}\_\{m,k\}≥1−min\{1,Cm,kfix\},\\displaystyle\\geq 1\-\\min\\\{1,C\_\{m,k\}^\{\\rm fix\}\\\},\(10\)Cm,kfix\\displaystyle C\_\{m,k\}^\{\\rm fix\}:=12𝔼\(1\+ρ2\)J−1,\\displaystyle:=\\frac\{1\}\{2\}\\sqrt\{\\mathbb\{E\}\(1\+\\rho^\{2\}\)^\{J\}\-1\},𝔼\(1\+ρ2\)J\\displaystyle\\mathbb\{E\}\(1\+\\rho^\{2\}\)^\{J\}≤\(1\+kρ2m\)k\.\\displaystyle\\leq\\left\(1\+\\frac\{k\\rho^\{2\}\}\{m\}\\right\)^\{k\}\.\(11\)whereJ=\|S∩S′\|J=\|S\\cap S^\{\\prime\}\|for independent uniformkk\-subsetsS,S′S,S^\{\\prime\}\.
For achievability in a fixed regular channel, the local regime additionally requiresk/m→0k/m\\to 0,kρ/m→hk\\rho/\\sqrt\{m\}\\to h, andVar1\(r−1\)<∞\\operatorname\{Var\}\_\{1\}\(r\-1\)<\\infty; there the symmetric oracle score∑i\(r\(Yi\)−1\)\\sum\_\{i\}\(r\(Y\_\{i\}\)\-1\)attains the boundary without knowing the subset\.
Let𝒮k=\{S⊆\[m\]:\|S\|=k\}\\mathcal\{S\}\_\{k\}=\\\{S\\subseteq\[m\]:\|S\|=k\\\}andP¯m,k=\|𝒮k\|−1∑S∈𝒮kPS\\overline\{P\}\_\{m,k\}=\|\\mathcal\{S\}\_\{k\}\|^\{\-1\}\\sum\_\{S\\in\\mathcal\{S\}\_\{k\}\}P\_\{S\}\. Because type\-II error is linear in the alternative,
b¯m,k=𝔼P¯m,k\(1−ϕm\)\.\\bar\{b\}\_\{m,k\}=\\mathbb\{E\}\_\{\\overline\{P\}\_\{m,k\}\}\(1\-\\phi\_\{m\}\)\.UnderP0⊗mP\_\{0\}^\{\\otimes m\}, the likelihood ratio ofPSP\_\{S\}isLS=∏i∈Sr\(Yi\)L\_\{S\}=\\prod\_\{i\\in S\}r\(Y\_\{i\}\), so that ofP¯m,k\\overline\{P\}\_\{m,k\}isL¯=𝔼SLS\\overline\{L\}=\\mathbb\{E\}\_\{S\}L\_\{S\}\. Tonelli’s theorem and independence give
1\+χ2\(P¯m,k∥P0⊗m\)\\displaystyle 1\+\\chi^\{2\}\(\\overline\{P\}\_\{m,k\}\\\|P\_\{0\}^\{\\otimes m\}\)=𝔼0L¯2\\displaystyle=\\mathbb\{E\}\_\{0\}\\overline\{L\}^\{2\}=𝔼S,S′𝔼0\[LSLS′\]\\displaystyle=\\mathbb\{E\}\_\{S,S^\{\\prime\}\}\\mathbb\{E\}\_\{0\}\[L\_\{S\}L\_\{S^\{\\prime\}\}\]=𝔼S,S′\(1\+ρ2\)\|S∩S′\|\.\\displaystyle=\\mathbb\{E\}\_\{S,S^\{\\prime\}\}\(1\+\\rho^\{2\}\)^\{\|S\\cap S^\{\\prime\}\|\}\.Indeed, an index inS∩S′S\\cap S^\{\\prime\}contributes𝔼0r2=1\+ρ2\\mathbb\{E\}\_\{0\}r^\{2\}=1\+\\rho^\{2\}; every other selected index contributes𝔼0r=1\\mathbb\{E\}\_\{0\}r=1\. Applying the testing\-error/TV argument and the chi\-square TV bound from[section˜A\.1](https://arxiv.org/html/2608.07914#A1.SS1)proves[eq\.˜10](https://arxiv.org/html/2608.07914#A1.E10)\.
It remains to prove[eq\.˜11](https://arxiv.org/html/2608.07914#A1.E11)\. Conditional onSS, letai=1\+ρ2a\_\{i\}=1\+\\rho^\{2\}fori∈Si\\in Sandai=1a\_\{i\}=1otherwise\. Uniformly samplingS′S^\{\\prime\}without replacement gives
𝔼\[\(1\+ρ2\)J∣S\]\\displaystyle\\mathbb\{E\}\[\(1\+\\rho^\{2\}\)^\{J\}\\mid S\]=\(mk\)−1∑T⊆\[m\]\|T\|=k∏i∈Tai\.\\displaystyle=\\binom\{m\}\{k\}^\{\-1\}\\sum\_\{\\begin\{subarray\}\{c\}T\\subseteq\[m\]\\\\ \|T\|=k\\end\{subarray\}\}\\prod\_\{i\\in T\}a\_\{i\}\.Maclaurin’s inequality bounds this normalized elementary symmetric mean by thekkth power of the arithmetic mean:
𝔼\[\(1\+ρ2\)J∣S\]\\displaystyle\\mathbb\{E\}\[\(1\+\\rho^\{2\}\)^\{J\}\\mid S\]≤\(1m∑i=1mai\)k\\displaystyle\\leq\\left\(\\frac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}a\_\{i\}\\right\)^\{k\}=\(1\+kρ2m\)k\.\\displaystyle=\\left\(1\+\\frac\{k\\rho^\{2\}\}\{m\}\\right\)^\{k\}\.The bound is independent ofSS, proving[eq\.˜11](https://arxiv.org/html/2608.07914#A1.E11)\. In particular, ifk2ρ2/m→0k^\{2\}\\rho^\{2\}/m\\to 0, its right\-hand side is at mostexp\(k2ρ2/m\)→1\\exp\(k^\{2\}\\rho^\{2\}/m\)\\to 1, so average testing power cannot exceed size asymptotically\. SincesupSbm\(S\)≥b¯m,k\\sup\_\{S\}b\_\{m\}\(S\)\\geq\\bar\{b\}\_\{m,k\}, the same display is a minimax lower bound over fixed subsets\.
For achievability in a fixed channel, letZm=\(ρm\)−1∑ig\(Yi\)Z\_\{m\}=\(\\rho\\sqrt\{m\}\)^\{\-1\}\\sum\_\{i\}g\(Y\_\{i\}\)and supposek/m→0k/m\\to 0,kρ/m→h<∞k\\rho/\\sqrt\{m\}\\to h<\\infty, andVar1\(g\)<∞\\operatorname\{Var\}\_\{1\}\(g\)<\\infty\. Since
𝔼1g=𝔼0\[r\(r−1\)\]=𝔼0\(r2−r\)=ρ2,\\mathbb\{E\}\_\{1\}g=\\mathbb\{E\}\_\{0\}\[r\(r\-1\)\]=\\mathbb\{E\}\_\{0\}\(r^\{2\}\-r\)=\\rho^\{2\},𝔼PSZm=kρ/m→h\\mathbb\{E\}\_\{P\_\{S\}\}Z\_\{m\}=k\\rho/\\sqrt\{m\}\\to h\. The centered sum over them−km\-kclean indices, divided byρm\\rho\\sqrt\{m\}, converges to𝒩\(0,1\)\\mathcal\{N\}\(0,1\)by the ordinary CLT becausek/m→0k/m\\to 0\. The centered sum over thekkmember indices isoPS\(1\)o\_\{P\_\{S\}\}\(1\)after the same normalization, since its variance iskVar1\(g\)/\(mρ2\)→0k\\operatorname\{Var\}\_\{1\}\(g\)/\(m\\rho^\{2\}\)\\to 0\. Slutsky’s theorem therefore givesZm↝𝒩\(h,1\)Z\_\{m\}\\rightsquigarrow\\mathcal\{N\}\(h,1\)under everyPSP\_\{S\}, while its null limit is𝒩\(0,1\)\\mathcal\{N\}\(0,1\)\. The one\-sided score test consequently attains power1−Φ\(z1−τ−h\)1\-\\Phi\(z\_\{1\-\\tau\}\-h\)without knowingSS\. ∎
### A\.3Proof of the detection\-boundary corollary
Ifαmρmm→0\\alpha\_\{m\}\\rho\_\{m\}\\sqrt\{m\}\\to 0, thenmαm2ρm2→0m\\alpha\_\{m\}^\{2\}\\rho\_\{m\}^\{2\}\\to 0\. Since1\+x≤ex1\+x\\leq e^\{x\}forx≥0x\\geq 0,
Cm\(αm,ρm\)≤12exp\(mαm2ρm2\)−1⟶0\.C\_\{m\}\(\\alpha\_\{m\},\\rho\_\{m\}\)\\leq\\frac\{1\}\{2\}\\sqrt\{\\exp\(m\\alpha\_\{m\}^\{2\}\\rho\_\{m\}^\{2\}\)\-1\}\\longrightarrow 0\.Apply[eq\.˜5](https://arxiv.org/html/2608.07914#S3.E5)and take the upper limit\. ∎
### A\.4Local asymptotic optimality
###### Theorem 7\(Local asymptotic optimality; full statement\)\.
Suppose0<ρ2<∞0<\\rho^\{2\}<\\inftyand, for someϵ\>0\\epsilon\>0,𝔼0\|g\|2\+ϵ<∞\\mathbb\{E\}\_\{0\}\|g\|^\{2\+\\epsilon\}<\\infty\. Fixh≥0h\\geq 0and letαm=h/\(ρm\)\\alpha\_\{m\}=h/\(\\rho\\sqrt\{m\}\)for all sufficiently largemm\. Then
logdQαm⊗mdP0⊗m\\displaystyle\\log\\frac\{\\,\\mathrm\{d\}Q\_\{\\alpha\_\{m\}\}^\{\\otimes m\}\}\{\\,\\mathrm\{d\}P\_\{0\}^\{\\otimes m\}\}=hZm−h22\+oP0\(1\),\\displaystyle=hZ\_\{m\}\-\\frac\{h^\{2\}\}\{2\}\+o\_\{P\_\{0\}\}\(1\),\(12\)Zm\\displaystyle Z\_\{m\}↝𝒩\(0,1\)underH0,\\displaystyle\\rightsquigarrow\\mathcal\{N\}\(0,1\)\\quad\\text\{under \}H\_\{0\},\(13\)Zm\\displaystyle Z\_\{m\}↝𝒩\(h,1\)underH1,m\.\\displaystyle\\rightsquigarrow\\mathcal\{N\}\(h,1\)\\quad\\text\{under \}H\_\{1,m\}\.\(14\)The score test𝟏\{Zm\>z1−τ\}\\mathbf\{1\}\\\{Z\_\{m\}\>z\_\{1\-\\tau\}\\\}has asymptotic sizeτ\\tauand power1−Φ\(z1−τ−h\)1\-\\Phi\(z\_\{1\-\\tau\}\-h\)\. Every test sequence\{ψm\}\\\{\\psi\_\{m\}\\\}withlim supm𝔼0ψm≤τ\\limsup\_\{m\}\\mathbb\{E\}\_\{0\}\\psi\_\{m\}\\leq\\tausatisfies
lim supm𝔼αmψm≤1−Φ\(z1−τ−h\)\.\\limsup\_\{m\}\\mathbb\{E\}\_\{\\alpha\_\{m\}\}\\psi\_\{m\}\\leq 1\-\\Phi\(z\_\{1\-\\tau\}\-h\)\.
#### Proof\.
Letg=r−1g=r\-1\. Becauser≥0r\\geq 0,g≥−1g\\geq\-1\. Also,
𝔼0g=∫\(p1−p0\)dy=0,𝔼0g2=ρ2\.\\mathbb\{E\}\_\{0\}g=\\int\(p\_\{1\}\-p\_\{0\}\)\\,\\mathrm\{d\}y=0,\\qquad\\mathbb\{E\}\_\{0\}g^\{2\}=\\rho^\{2\}\.The central limit theorem gives
Zm=1ρm∑i=1mg\(Yi\)↝𝒩\(0,1\)Z\_\{m\}=\\frac\{1\}\{\\rho\\sqrt\{m\}\}\\sum\_\{i=1\}^\{m\}g\(Y\_\{i\}\)\\rightsquigarrow\\mathcal\{N\}\(0,1\)underP0⊗mP\_\{0\}^\{\\otimes m\}\.
Ifh=0h=0, the likelihood ratio is identically one and all stated limits are immediate\. Hence assumeh\>0h\>0\. Forαm=h/\(ρm\)\\alpha\_\{m\}=h/\(\\rho\\sqrt\{m\}\), the log\-likelihood ratio is
Λm=∑i=1mlog\(1\+αmg\(Yi\)\)\.\\Lambda\_\{m\}=\\sum\_\{i=1\}^\{m\}\\log\(1\+\\alpha\_\{m\}g\(Y\_\{i\}\)\)\.Letϵ′=min\{ϵ,1\}\\epsilon^\{\\prime\}=\\min\\\{\\epsilon,1\\\}\. The moment condition and a union bound imply, for every fixedc\>0c\>0,
ℙ0\(maxi≤m\|αmg\(Yi\)\|\>c\)\\displaystyle\\mathbb\{P\}\_\{0\}\\\!\\left\(\\max\_\{i\\leq m\}\|\\alpha\_\{m\}g\(Y\_\{i\}\)\|\>c\\right\)≤mℙ0\(\|g\|\>cαm\)\\displaystyle\\quad\\leq m\\mathbb\{P\}\_\{0\}\\\!\\left\(\|g\|\>\\frac\{c\}\{\\alpha\_\{m\}\}\\right\)≤mαm2\+ϵ′c2\+ϵ′𝔼0\|g\|2\+ϵ′⟶0\.\\displaystyle\\leq\\frac\{m\\alpha\_\{m\}^\{2\+\\epsilon^\{\\prime\}\}\}\{c^\{2\+\\epsilon^\{\\prime\}\}\}\\mathbb\{E\}\_\{0\}\|g\|^\{2\+\\epsilon^\{\\prime\}\}\\longrightarrow 0\.On the eventmaxi\|αmgi\|≤1/2\\max\_\{i\}\|\\alpha\_\{m\}g\_\{i\}\|\\leq 1/2, Taylor’s theorem gives a constantCϵ′C\_\{\\epsilon^\{\\prime\}\}such that
\|log\(1\+x\)−x\+x22\|\\displaystyle\\left\|\\log\(1\+x\)\-x\+\\frac\{x^\{2\}\}\{2\}\\right\|≤Cϵ′\|x\|2\+ϵ′,\\displaystyle\\leq C\_\{\\epsilon^\{\\prime\}\}\|x\|^\{2\+\\epsilon^\{\\prime\}\},\|x\|≤1/2\.\\displaystyle\\hskip\-40\.0pt\|x\|\\leq 1/2\.Therefore the summed remainder is bounded in probability by
Cϵ′αm2\+ϵ′∑i=1m\|g\(Yi\)\|2\+ϵ′\\displaystyle C\_\{\\epsilon^\{\\prime\}\}\\alpha\_\{m\}^\{2\+\\epsilon^\{\\prime\}\}\\sum\_\{i=1\}^\{m\}\|g\(Y\_\{i\}\)\|^\{2\+\\epsilon^\{\\prime\}\}=OP0\(mαm2\+ϵ′\)=oP0\(1\)\.\\displaystyle\\quad=O\_\{P\_\{0\}\}\(m\\alpha\_\{m\}^\{2\+\\epsilon^\{\\prime\}\}\)=o\_\{P\_\{0\}\}\(1\)\.The weak law of large numbers also givesm−1∑ig\(Yi\)2→ρ2m^\{\-1\}\\sum\_\{i\}g\(Y\_\{i\}\)^\{2\}\\to\\rho^\{2\}inP0P\_\{0\}probability\. Combining these facts,
Λm\\displaystyle\\Lambda\_\{m\}=αm∑i=1mg\(Yi\)\\displaystyle=\\alpha\_\{m\}\\sum\_\{i=1\}^\{m\}g\(Y\_\{i\}\)−αm22∑i=1mg\(Yi\)2\+oP0\(1\)\\displaystyle\\quad\-\\frac\{\\alpha\_\{m\}^\{2\}\}\{2\}\\sum\_\{i=1\}^\{m\}g\(Y\_\{i\}\)^\{2\}\+o\_\{P\_\{0\}\}\(1\)=hZm−h22\+oP0\(1\),\\displaystyle=hZ\_\{m\}\-\\frac\{h^\{2\}\}\{2\}\+o\_\{P\_\{0\}\}\(1\),which is[eq\.˜12](https://arxiv.org/html/2608.07914#A1.E12)\.
UnderP0⊗mP\_\{0\}^\{\\otimes m\},\(Zm,Λm\)\(Z\_\{m\},\\Lambda\_\{m\}\)converges jointly to\(Z,hZ−h2/2\)\(Z,hZ\-h^\{2\}/2\)forZ∼𝒩\(0,1\)Z\\sim\\mathcal\{N\}\(0,1\)\. Since𝔼\[exp\(hZ−h2/2\)\]=1\\mathbb\{E\}\[\\exp\(hZ\-h^\{2\}/2\)\]=1, Le Cam’s third lemma applies and givesZm↝𝒩\(h,1\)Z\_\{m\}\\rightsquigarrow\\mathcal\{N\}\(h,1\)underQαm⊗mQ\_\{\\alpha\_\{m\}\}^\{\\otimes m\}\. Hence
ℙ0\(Zm\>z1−τ\)→τ\\mathbb\{P\}\_\{0\}\(Z\_\{m\}\>z\_\{1\-\\tau\}\)\\to\\tauand
ℙαm\(Zm\>z1−τ\)→1−Φ\(z1−τ−h\)\.\\mathbb\{P\}\_\{\\alpha\_\{m\}\}\(Z\_\{m\}\>z\_\{1\-\\tau\}\)\\to 1\-\\Phi\(z\_\{1\-\\tau\}\-h\)\.
It remains to justify the power envelope\. For each fixedmm, the Neyman–Pearson lemma makes a possibly randomized threshold test ofΛm\\Lambda\_\{m\}most powerful for the simple alternativeQαm⊗mQ\_\{\\alpha\_\{m\}\}^\{\\otimes m\}\. Let\{ψm\}\\\{\\psi\_\{m\}\\\}be any test sequence withlim supm𝔼0ψm≤τ\\limsup\_\{m\}\\mathbb\{E\}\_\{0\}\\psi\_\{m\}\\leq\\tau\. For everyδ∈\(0,1−τ\)\\delta\\in\(0,1\-\\tau\), it eventually has size at mosts=τ\+δs=\\tau\+\\delta, so Neyman–Pearson bounds its power by that of the likelihood\-ratio test of sizess\. Under the null,[eq\.˜12](https://arxiv.org/html/2608.07914#A1.E12)impliesΛm↝𝒩\(−h2/2,h2\)\\Lambda\_\{m\}\\rightsquigarrow\\mathcal\{N\}\(\-h^\{2\}/2,h^\{2\}\), a continuous law\. Hence the randomized NP threshold converges to−h2/2\+hz1−s\-h^\{2\}/2\+hz\_\{1\-s\}\. Contiguity and Le Cam’s third lemma giveΛm↝𝒩\(h2/2,h2\)\\Lambda\_\{m\}\\rightsquigarrow\\mathcal\{N\}\(h^\{2\}/2,h^\{2\}\)under the alternative, so the NP power converges to1−Φ\(z1−s−h\)1\-\\Phi\(z\_\{1\-s\}\-h\)\. Taking the upper limit and then lettingδ↓0\\delta\\downarrow 0gives the stated envelope\. The score test attains it by the null and alternative limits above\. ∎
### A\.5Proof of score efficacy
Letf~=f−𝔼0f\\widetilde\{f\}=f\-\\mathbb\{E\}\_\{0\}f\. Since𝔼0\(r−1\)=0\\mathbb\{E\}\_\{0\}\(r\-1\)=0,
Δf\\displaystyle\\Delta\_\{f\}=∫f\(p1−p0\)dy\\displaystyle=\\int f\(p\_\{1\}\-p\_\{0\}\)\\,\\mathrm\{d\}y=𝔼0\[\(r−1\)f~\]\.\\displaystyle=\\mathbb\{E\}\_\{0\}\[\(r\-1\)\\widetilde\{f\}\]\.By Cauchy–Schwarz,
Δf2≤𝔼0\(r−1\)2𝔼0f~2=ρ2σf2\.\\Delta\_\{f\}^\{2\}\\leq\\mathbb\{E\}\_\{0\}\(r\-1\)^\{2\}\\,\\mathbb\{E\}\_\{0\}\\widetilde\{f\}^\{2\}=\\rho^\{2\}\\sigma\_\{f\}^\{2\}\.Dividing byσf2\>0\\sigma\_\{f\}^\{2\}\>0provesef≤ρe\_\{f\}\\leq\\rho\. Equality in Cauchy–Schwarz holds exactly whenf~=c\(r−1\)\\widetilde\{f\}=c\(r\-1\)P0P\_\{0\}\-almost surely for a nonzero constantcc\. Finally, linearity under the mixture gives
𝔼Qαf−𝔼0f\\displaystyle\\mathbb\{E\}\_\{Q\_\{\\alpha\}\}f\-\\mathbb\{E\}\_\{0\}f=\(1−α\)𝔼0f\+α𝔼1f−𝔼0f\\displaystyle=\(1\-\\alpha\)\\mathbb\{E\}\_\{0\}f\+\\alpha\\mathbb\{E\}\_\{1\}f\-\\mathbb\{E\}\_\{0\}f=αΔf\.\\displaystyle=\\alpha\\Delta\_\{f\}\.The standard deviation of the sample mean under the null isσf/m\\sigma\_\{f\}/\\sqrt\{m\}, so its standardized mean shift isαmΔf/σf\\alpha\\sqrt\{m\}\\,\\Delta\_\{f\}/\\sigma\_\{f\}, with absolute valueαmef\\alpha\\sqrt\{m\}\\,e\_\{f\}\. ∎
### A\.6Proof of the efficacy floor
Forz∈\{0,1\}z\\in\\\{0,1\\\}, Hoeffding’s inequality gives
ℙ\{\|f¯z−μz\|\>εcal\(nz\)\}\\displaystyle\\mathbb\{P\}\\\{\|\\bar\{f\}\_\{z\}\-\\mu\_\{z\}\|\>\\varepsilon\_\{\\rm cal\}\(n\_\{z\}\)\\\}≤2e−2nzεcal\(nz\)2=δcal3\.\\displaystyle\\hskip 35\.0pt\\leq 2e^\{\-2n\_\{z\}\\varepsilon\_\{\\rm cal\}\(n\_\{z\}\)^\{2\}\}=\\frac\{\\delta\_\{\\mathrm\{cal\}\}\}\{3\}\.For the clean sample, the pairwise sample variance ofMaurer and Pontil \([2009](https://arxiv.org/html/2608.07914#bib.bib19), Theorem 10\)is exactly
Vn0\\displaystyle V\_\{n\_\{0\}\}=1n0\(n0−1\)∑i<j\{f\(Yi0\)−f\(Yj0\)\}2\\displaystyle=\\frac\{1\}\{n\_\{0\}\(n\_\{0\}\-1\)\}\\sum\_\{i<j\}\\\{f\(Y\_\{i\}^\{0\}\)\-f\(Y\_\{j\}^\{0\}\)\\\}^\{2\}=σ^f2,\\displaystyle=\\widehat\{\\sigma\}\_\{f\}^\{2\},and its expectation isσf2\\sigma\_\{f\}^\{2\}\. Their lower\-tail standard\-deviation inequality, with failure probabilityδcal/3\\delta\_\{\\mathrm\{cal\}\}/3, yields
ℙ\{σf\>σ^f\+ucal\(n0\)\}\\displaystyle\\mathbb\{P\}\\\{\\sigma\_\{f\}\>\\widehat\{\\sigma\}\_\{f\}\+u\_\{\\rm cal\}\(n\_\{0\}\)\\\}≤δcal3\.\\displaystyle\\hskip 70\.0pt\\leq\\frac\{\\delta\_\{\\mathrm\{cal\}\}\}\{3\}\.A union bound therefore gives, with probability at least1−δcal1\-\\delta\_\{\\mathrm\{cal\}\}, both mean bounds andσf≤σ^f\+ucal\(n0\)\\sigma\_\{f\}\\leq\\widehat\{\\sigma\}\_\{f\}\+u\_\{\\rm cal\}\(n\_\{0\}\)\. Popoviciu’s deterministic inequality also givesσf≤1/2\\sigma\_\{f\}\\leq 1/2, henceσf≤σ¯f\\sigma\_\{f\}\\leq\\overline\{\\sigma\}\_\{f\}\. On the same event, the reverse triangle inequality gives\|μ1−μ0\|≥ΔL\|\\mu\_\{1\}\-\\mu\_\{0\}\|\\geq\\Delta\_\{L\}\. Consequently,
ef=\|μ1−μ0\|σf≥ΔLσ¯f=e¯f,e\_\{f\}=\\frac\{\|\\mu\_\{1\}\-\\mu\_\{0\}\|\}\{\\sigma\_\{f\}\}\\geq\\frac\{\\Delta\_\{L\}\}\{\\overline\{\\sigma\}\_\{f\}\}=\\underline\{e\}\_\{f\},and[proposition˜5](https://arxiv.org/html/2608.07914#Thmtheorem5)givesρ≥ef\\rho\\geq e\_\{f\}\. Replacingσ¯f\\overline\{\\sigma\}\_\{f\}by the deterministic upper bound1/21/2proves the variance\-free alternative stated after the corollary\. ∎
### A\.7Proof of the contamination certificate
WriteμA=𝔼Qαf\\mu\_\{A\}=\\mathbb\{E\}\_\{Q\_\{\\alpha\}\}f\. Sincef∈\[0,1\]f\\in\[0,1\], two\-sided Hoeffding bounds give, for any sample of sizennand its corresponding meanμ\\mu,
ℙ\{\|f¯−μ\|\>εcert\(n\)\}≤2e−2nεcert\(n\)2=δcert3\.\\mathbb\{P\}\\\{\|\\bar\{f\}\-\\mu\|\>\\varepsilon\_\{\\rm cert\}\(n\)\\\}\\leq 2e^\{\-2n\\varepsilon\_\{\\rm cert\}\(n\)^\{2\}\}=\\frac\{\\delta\_\{\\mathrm\{cert\}\}\}\{3\}\.A union bound over the audit, clean\-calibration, and seen\-calibration samples shows that, with probability at least1−δcert1\-\\delta\_\{\\mathrm\{cert\}\}, all three events
\|f¯A−μA\|\\displaystyle\|\\bar\{f\}\_\{A\}\-\\mu\_\{A\}\|≤εcert\(m\),\\displaystyle\\leq\\varepsilon\_\{\\rm cert\}\(m\),\|f¯0−μ0\|\\displaystyle\|\\bar\{f\}\_\{0\}\-\\mu\_\{0\}\|≤εcert\(n0\),\\displaystyle\\leq\\varepsilon\_\{\\rm cert\}\(n\_\{0\}\),\|f¯1−μ1\|\\displaystyle\|\\bar\{f\}\_\{1\}\-\\mu\_\{1\}\|≤εcert\(n1\)\\displaystyle\\leq\\varepsilon\_\{\\rm cert\}\(n\_\{1\}\)hold simultaneously\. On this event,
NL≤μA−μ0=α\(μ1−μ0\)N\_\{L\}\\leq\\mu\_\{A\}\-\\mu\_\{0\}=\\alpha\(\\mu\_\{1\}\-\\mu\_\{0\}\)by the mixture identity, while
DU≥μ1−μ0\.D\_\{U\}\\geq\\mu\_\{1\}\-\\mu\_\{0\}\.Writed=μ1−μ0d=\\mu\_\{1\}\-\\mu\_\{0\}\. Ifd≤0d\\leq 0, thenNL≤αd≤0N\_\{L\}\\leq\\alpha d\\leq 0, soα¯=0≤α\\underline\{\\alpha\}=0\\leq\\alpharegardless of the sign ofDUD\_\{U\}\. Now supposed\>0d\>0\. ThenDU≥d\>0D\_\{U\}\\geq d\>0\. IfNL≤0N\_\{L\}\\leq 0, againα¯=0≤α\\underline\{\\alpha\}=0\\leq\\alpha\. IfNL\>0N\_\{L\}\>0, then
NLDU≤αdDU≤α\.\\frac\{N\_\{L\}\}\{D\_\{U\}\}\\leq\\frac\{\\alpha d\}\{D\_\{U\}\}\\leq\\alpha\.The truncation to\[0,1\]\[0,1\]preserves this inequality becauseα∈\[0,1\]\\alpha\\in\[0,1\]\. Thusα¯≤α\\underline\{\\alpha\}\\leq\\alphathroughout the simultaneous event, which has probability at least1−δcert1\-\\delta\_\{\\mathrm\{cert\}\}\. ∎
## Appendix BThe Contamination Certificate: Formal Statement
###### Theorem 8\(Finite\-sample contamination certificate\)\.
Fix anyf:𝒴→\[0,1\]f:\\mathcal\{Y\}\\to\[0,1\]independently of all samples below\. Let
f¯A\\displaystyle\\bar\{f\}\_\{A\}=1m∑i=1mf\(YiA\),\\displaystyle=\\frac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}f\(Y\_\{i\}^\{A\}\),f¯0\\displaystyle\\bar\{f\}\_\{0\}=1n0∑j=1n0f\(Yj0\),\\displaystyle=\\frac\{1\}\{n\_\{0\}\}\\sum\_\{j=1\}^\{n\_\{0\}\}f\(Y\_\{j\}^\{0\}\),f¯1\\displaystyle\\bar\{f\}\_\{1\}=1n1∑k=1n1f\(Yk1\),\\displaystyle=\\frac\{1\}\{n\_\{1\}\}\\sum\_\{k=1\}^\{n\_\{1\}\}f\(Y\_\{k\}^\{1\}\),from mutually independent samplesYiA∼QαY\_\{i\}^\{A\}\\sim Q\_\{\\alpha\},Yj0∼P0Y\_\{j\}^\{0\}\\sim P\_\{0\}, andYk1∼P1Y\_\{k\}^\{1\}\\sim P\_\{1\}\. For confidence level1−δcert1\-\\delta\_\{\\mathrm\{cert\}\}, define
εcert\(n\)=log\(6/δcert\)2n,\\varepsilon\_\{\\rm cert\}\(n\)=\\sqrt\{\\frac\{\\log\(6/\\delta\_\{\\mathrm\{cert\}\}\)\}\{2n\}\},NL\\displaystyle N\_\{L\}=\(f¯A−εcert\(m\)\)−\(f¯0\+εcert\(n0\)\),\\displaystyle=\(\\bar\{f\}\_\{A\}\-\\varepsilon\_\{\\rm cert\}\(m\)\)\-\(\\bar\{f\}\_\{0\}\+\\varepsilon\_\{\\rm cert\}\(n\_\{0\}\)\),DU\\displaystyle D\_\{U\}=\(f¯1\+εcert\(n1\)\)−\(f¯0−εcert\(n0\)\),\\displaystyle=\(\\bar\{f\}\_\{1\}\+\\varepsilon\_\{\\rm cert\}\(n\_\{1\}\)\)\-\(\\bar\{f\}\_\{0\}\-\\varepsilon\_\{\\rm cert\}\(n\_\{0\}\)\),and
α¯=\{min\{1,max\{0,NL/DU\}\},DU\>0,0,DU≤0\.\\underline\{\\alpha\}=\\begin\{cases\}\\min\\\{1,\\max\\\{0,N\_\{L\}/D\_\{U\}\\\}\\\},&D\_\{U\}\>0,\\\\ 0,&D\_\{U\}\\leq 0\.\\end\{cases\}Then
ℙα\{α¯≤α\}≥1−δcert\.\\mathbb\{P\}\_\{\\alpha\}\\\{\\underline\{\\alpha\}\\leq\\alpha\\\}\\geq 1\-\\delta\_\{\\mathrm\{cert\}\}\.
## Appendix CThe Variance\-Adaptive Efficacy Floor
###### Corollary 9\(A variance\-adaptive finite\-sample efficacy floor\)\.
Fix a nonconstantf:𝒴→\[0,1\]f:\\mathcal\{Y\}\\to\[0,1\]independently of clean and seen calibration samples of sizesn0≥2,n1≥1n\_\{0\}\\geq 2,n\_\{1\}\\geq 1\. Letσ^f2=\(n0−1\)−1∑j=1n0\(f\(Yj0\)−f¯0\)2\\widehat\{\\sigma\}\_\{f\}^\{2\}=\(n\_\{0\}\-1\)^\{\-1\}\\sum\_\{j=1\}^\{n\_\{0\}\}\(f\(Y\_\{j\}^\{0\}\)\-\\bar\{f\}\_\{0\}\)^\{2\}\. Forδcal∈\(0,1\)\\delta\_\{\\mathrm\{cal\}\}\\in\(0,1\), define
εcal\(n\)\\displaystyle\\varepsilon\_\{\\rm cal\}\(n\):=log\(6/δcal\)2n,\\displaystyle:=\\sqrt\{\\frac\{\\log\(6/\\delta\_\{\\mathrm\{cal\}\}\)\}\{2n\}\},ucal\(n0\)\\displaystyle u\_\{\\rm cal\}\(n\_\{0\}\):=2log\(3/δcal\)n0−1,\\displaystyle:=\\sqrt\{\\frac\{2\\log\(3/\\delta\_\{\\mathrm\{cal\}\}\)\}\{n\_\{0\}\-1\}\},σ¯f\\displaystyle\\overline\{\\sigma\}\_\{f\}:=min\{1/2,σ^f\+ucal\(n0\)\},\\displaystyle:=\\min\\\{1/2,\\widehat\{\\sigma\}\_\{f\}\+u\_\{\\rm cal\}\(n\_\{0\}\)\\\},ΔL\\displaystyle\\Delta\_\{L\}:=\[\|f¯1−f¯0\|−εcal\(n1\)−εcal\(n0\)\]\+\.\\displaystyle:=\\left\[\|\\bar\{f\}\_\{1\}\-\\bar\{f\}\_\{0\}\|\-\\varepsilon\_\{\\rm cal\}\(n\_\{1\}\)\-\\varepsilon\_\{\\rm cal\}\(n\_\{0\}\)\\right\]\_\{\+\}\.Then, with probability at least1−δcal1\-\\delta\_\{\\mathrm\{cal\}\},
ρ\\displaystyle\\rho≥ef≥e¯f,\\displaystyle\\geq e\_\{f\}\\geq\\underline\{e\}\_\{f\},e¯f\\displaystyle\\underline\{e\}\_\{f\}:=ΔLσ¯f\.\\displaystyle:=\\frac\{\\Delta\_\{L\}\}\{\\overline\{\\sigma\}\_\{f\}\}\.\(15\)
The standard\-deviation radius is the exact bounded\-variable result ofMaurer and Pontil \([2009](https://arxiv.org/html/2608.07914#bib.bib19)\); it is useful when clean scores have low variance\. Replacingσ¯f\\overline\{\\sigma\}\_\{f\}by1/21/2yields the fully variance\-free but usually looser floor2ΔL2\\Delta\_\{L\}\. Either version is a score\-efficacy lower bound and hence a certified floor onρ\\rho; neither pretends that a learned density ratio estimatesρ\\rho\.
## Appendix DThe Matching Test in Detail
Writeg\(y\)=r\(y\)−1g\(y\)=r\(y\)\-1\. The score atα=0\\alpha=0for the mixture family isgg, and𝔼0g=0\\mathbb\{E\}\_\{0\}g=0,Var0\(g\)=ρ2\\operatorname\{Var\}\_\{0\}\(g\)=\\rho^\{2\}\. Define
Zm=1ρm∑i=1mg\(Yi\)\.Z\_\{m\}=\\frac\{1\}\{\\rho\\sqrt\{m\}\}\\sum\_\{i=1\}^\{m\}g\(Y\_\{i\}\)\.\(16\)
#### Local power envelope\.
Suppose0<ρ2<∞0<\\rho^\{2\}<\\infty,𝔼0\|g\|2\+ϵ<∞\\mathbb\{E\}\_\{0\}\|g\|^\{2\+\\epsilon\}<\\inftyfor someϵ\>0\\epsilon\>0, andαm=h/\(ρm\)\\alpha\_\{m\}=h/\(\\rho\\sqrt\{m\}\)\. The full LAN statement in[theorem˜7](https://arxiv.org/html/2608.07914#Thmtheorem7)shows thatZmZ\_\{m\}converges to𝒩\(0,1\)\\mathcal\{N\}\(0,1\)under the null and𝒩\(h,1\)\\mathcal\{N\}\(h,1\)under the local alternative\. The score test has asymptotic sizeτ\\tauand attains the power envelope
1−Φ\(z1−τ−h\),1\-\\Phi\(z\_\{1\-\\tau\}\-h\),\(17\)which no asymptotic level\-τ\\tautest exceeds\. This is a local result, not a claim that one likelihood\-ratio model is correct for every detector\. It shows that the lower\-bound scaling is attainable in each regular channel and gives the oracle asymptotic sample budget
moracle≈\(z1−τ\+z1−βαρ\)2m\_\{\\mathrm\{oracle\}\}\\approx\\left\(\\frac\{z\_\{1\-\\tau\}\+z\_\{1\-\\beta\}\}\{\\alpha\\rho\}\\right\)^\{2\}\(18\)for sizeτ\\tauand power1−β1\-\\beta\.
## Appendix EExtensions
### E\.1Adaptive transcripts
###### Proposition 10\(Adaptive information budget\)\.
LetHt−1H\_\{t\-1\}be the transcript before roundtt, including all adaptively chosen prompts\. Each round draws a fresh benchmark item; all turns involving one item are included in that round’s single joint response, and no item is revisited later\. Conditional on every historyhh, suppose the clean response law isP0,t\(⋅∣h\)P\_\{0,t\}\(\\cdot\\mid h\)and the contaminated law is
Qα,t\(⋅∣h\)=\(1−α\)P0,t\(⋅∣h\)\+αP1,t\(⋅∣h\),Q\_\{\\alpha,t\}\(\\cdot\\mid h\)=\(1\-\\alpha\)P\_\{0,t\}\(\\cdot\\mid h\)\+\\alpha P\_\{1,t\}\(\\cdot\\mid h\),where
χ2\(P1,t\(⋅∣h\)∥P0,t\(⋅∣h\)\)≤ρt2\\chi^\{2\}\(P\_\{1,t\}\(\\cdot\\mid h\)\\\|P\_\{0,t\}\(\\cdot\\mid h\)\)\\leq\\rho\_\{t\}^\{2\}uniformly inhh\. Then every test of the final transcript satisfies
am\+bm\\displaystyle a\_\{m\}\+b\_\{m\}≥1−min\{1,\\displaystyle\\geq 1\-\\min\\left\\\{1,\\right\.\[12∑t=1mlog\(1\+α2ρt2\)\]1/2\}\.\\displaystyle\\hskip 24\.0pt\\left\.\\left\[\\frac\{1\}\{2\}\\sum\_\{t=1\}^\{m\}\\log\(1\+\\alpha^\{2\}\\rho\_\{t\}^\{2\}\)\\right\]^\{1/2\}\\right\\\}\.
#### Proof\.
Letℚα\(m\)\\mathbb\{Q\}\_\{\\alpha\}^\{\(m\)\}andℙ0\(m\)\\mathbb\{P\}\_\{0\}^\{\(m\)\}denote the two transcript laws, and define
𝒦t\(h\)\\displaystyle\\mathcal\{K\}\_\{t\}\(h\):=KL\(Qα,t\(⋅∣h\)∥P0,t\(⋅∣h\)\)\.\\displaystyle:=\\operatorname\{KL\}\\\!\\left\(Q\_\{\\alpha,t\}\(\\cdot\\mid h\)\\,\\middle\\\|\\,P\_\{0,t\}\(\\cdot\\mid h\)\\right\)\.The chain rule for KL divergence gives
KL\(ℚα\(m\)∥ℙ0\(m\)\)\\displaystyle\\operatorname\{KL\}\(\\mathbb\{Q\}\_\{\\alpha\}^\{\(m\)\}\\\|\\mathbb\{P\}\_\{0\}^\{\(m\)\}\)=∑t=1m𝔼ℚα\(m\)\[𝒦t\(Ht−1\)\]\.\\displaystyle=\\sum\_\{t=1\}^\{m\}\\mathbb\{E\}\_\{\\mathbb\{Q\}\_\{\\alpha\}^\{\(m\)\}\}\[\\mathcal\{K\}\_\{t\}\(H\_\{t\-1\}\)\]\.Conditionally on each history, direct expansion gives
χ2\(Qα,t\(⋅∣h\)∥P0,t\(⋅∣h\)\)≤α2ρt2\.\\chi^\{2\}\(Q\_\{\\alpha,t\}\(\\cdot\\mid h\)\\\|P\_\{0,t\}\(\\cdot\\mid h\)\)\\leq\\alpha^\{2\}\\rho\_\{t\}^\{2\}\.For arbitraryP≪QP\\ll Q, Jensen’s inequality givesKL\(P∥Q\)≤log\{1\+χ2\(P∥Q\)\}\\operatorname\{KL\}\(P\\\|Q\)\\leq\\log\\\{1\+\\chi^\{2\}\(P\\\|Q\)\\\}\. Thus the inner KL is at mostlog\(1\+α2ρt2\)\\log\(1\+\\alpha^\{2\}\\rho\_\{t\}^\{2\}\)\. Summing, applying Pinsker, and using the testing\-error/TV inequality from[section˜A\.1](https://arxiv.org/html/2608.07914#A1.SS1)proves the claim\. ∎
### E\.2Independent heterogeneous probes
If itemiihas clean and seen laws\(P0,i,P1,i\)\(P\_\{0,i\},P\_\{1,i\}\)withρi2=χ2\(P1,i∥P0,i\)\\rho\_\{i\}^\{2\}=\\chi^\{2\}\(P\_\{1,i\}\\\|P\_\{0,i\}\), writeQα,i=\(1−α\)P0,i\+αP1,iQ\_\{\\alpha,i\}=\(1\-\\alpha\)P\_\{0,i\}\+\\alpha P\_\{1,i\}\. Independence gives the exact tensorization
Rm\\displaystyle R\_\{m\}:=1\+χ2\(⨂i=1mQα,i∥⨂i=1mP0,i\)\\displaystyle:=1\+\\chi^\{2\}\\\!\\left\(\\bigotimes\_\{i=1\}^\{m\}Q\_\{\\alpha,i\}\\middle\\\|\\bigotimes\_\{i=1\}^\{m\}P\_\{0,i\}\\right\)=∏i=1m\(1\+α2ρi2\)\.\\displaystyle=\\prod\_\{i=1\}^\{m\}\(1\+\\alpha^\{2\}\\rho\_\{i\}^\{2\}\)\.Therefore the same chi\-square/TV argument as[theorem˜3](https://arxiv.org/html/2608.07914#Thmtheorem3)gives
am\+bm\\displaystyle a\_\{m\}\+b\_\{m\}≥1−min\{1,12Rm−1\}\.\\displaystyle\\geq 1\-\\min\\left\\\{1,\\frac\{1\}\{2\}\\sqrt\{R\_\{m\}\-1\}\\right\\\}\.Whenmaxiα2ρi2=o\(1\)\\max\_\{i\}\\alpha^\{2\}\\rho\_\{i\}^\{2\}=o\(1\), the cumulative local information isα2∑iρi2\\alpha^\{2\}\\sum\_\{i\}\\rho\_\{i\}^\{2\}\. The corresponding heterogeneous score is∑i\(ri\(Yi\)−1\)\\sum\_\{i\}\(r\_\{i\}\(Y\_\{i\}\)\-1\)normalized by\(∑iρi2\)1/2\(\\sum\_\{i\}\\rho\_\{i\}^\{2\}\)^\{1/2\}\. A Lindeberg condition yields the same Gaussian limit withρ2m\\rho^\{2\}mreplaced by∑iρi2\\sum\_\{i\}\\rho\_\{i\}^\{2\}\.
### E\.3Sampling a fixed finite benchmark
Suppose a benchmark hasNNdistinct items and exactlykkare contaminated\. Evaluating all items is[lemma˜6](https://arxiv.org/html/2608.07914#Thmtheorem6)withm=Nm=N, so it does not require i\.i\.d\. Bernoulli labels\. Samplingm<Nm<Ndistinct indices without replacement produces hypergeometric labels; the implementation conditions on the realized sample size and uses the exact randomization distribution \(or a predeclared finite\-population correction\)\. Sampling indices with replacement recovers Bernoulli labels marginally but does not create new independent source items\. Repeated decodes remain components of one joint item outcome\.
## Appendix FValidity Gates in Full
The four diagnostics summarized in §[4\.4](https://arxiv.org/html/2608.07914#S4.SS4):
1. 1\.Blind separation\.A classifier that sees raw text but no target\-model output must be at or near chance; otherwise content shift can masquerade as memorization\(Daset al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib7); Wanget al\.,[2026b](https://arxiv.org/html/2608.07914#bib.bib31)\)\. The operational rule, fixed in the run scripts before scoring: for same\-corpus channels the gate passes if the AUC’s95%95\\%interval contains0\.50\.5*or*the point AUC is at most the equivalence margin0\.550\.55; for cross\-dataset channels the content\-shift limit is0\.750\.75\. One v3 channel \(pythia\-410m×\\timeswiki, AUC0\.5440\.544\[0\.512,0\.579\]\[0\.512,0\.579\]\) passes via the margin while its interval excludes exact chance—we state this rather than claiming compatibility with0\.50\.5\. For cross\-dataset channels, where the pools differ by construction, the gate instead flags extreme content shift\.
2. 2\.Clean\-channel transport\.Clean\-control score distributions are compared across the clean and contaminated runs\. Large shifts reject the no\-spillover interpretation\. Numerical rule as run in v5: two\-sample KS statistic<0\.15<0\.15\(observed range0\.0460\.046–0\.1030\.103; per\-channel values in[table˜6](https://arxiv.org/html/2608.07914#A11.T6)\)\.
3. 3\.Channel stability\.Conditional seen and clean score laws are compared acrossα\\alphausing pre\-specified two\-sample distances\. Scaling claims are limited to the range where these laws are stable\. Numerical rule as run in v5: KS<0\.15<0\.15\(observed range0\.0340\.034–0\.0940\.094\); the v5 size gate additionally requires the empirical\-size Wilson lower bound≤0\.05\\leq 0\.05\.
4. 4\.Independence unit\.Near\-duplicates, paraphrases, and repeated decodes are grouped by source item\. Effectivemmis the number of sampled item clusters\.
## Appendix GProbe Families, Baselines, and Statistical Protocol
### G\.1Probe families and baselines
#### Gray\-box \(run in this paper\)\.
We evaluate mean loss and zlib\-normalized loss\(Carliniet al\.,[2021](https://arxiv.org/html/2608.07914#bib.bib5)\), Min\-K% Prob\(Shiet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib29)\), Min\-K%\+\+\(Zhanget al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib36)\), and neighborhood loss\(Matternet al\.,[2023](https://arxiv.org/html/2608.07914#bib.bib21)\)\. Each method produces one item\-level score\. Hyperparameters are fixed on the score\-training split\. The “best probe” reported anywhere in the paper is the probe with the highest calibration\-split efficacy, selected before audit power is computed\.
#### Gray\-box \(specified, not run\)\.
The protocol also admits ReCaLL\(Xieet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib33)\)and a multifeature CAMIA\-style score\(Changet al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib6)\); neither is run in the reported experiments and no claim rests on them\.
#### Black\-box \(specified, not run\)\.
The protocol specifies output correctness, sampling consistency, the Data Contamination Quiz\(Golchin and Surdeanu,[2025](https://arxiv.org/html/2608.07914#bib.bib14)\), and canonical\-order/exchangeability probes\(Orenet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib23)\), with fixed model versions, prompts, decoding parameters, and retry policies, all generations for one item forming one joint probe\. These are not run here; all reported results are gray\-box\.
#### Power\-calibrated combination \(specified, not run\)\.
The protocol’s intended primary score is a cross\-fitted logistic density\-ratio classifier over the baseline features available in the given access regime, with an identically tuned blind classifier on text\-only features to expose confounding\. It is not run in this paper; in every reported experiment the primary score is the calibration\-selected best single probe defined above\.
### G\.2Validated\-paraphrase pipeline
Candidate paraphrases of each injected question are generated by Qwen2\.5\-1\.5B\-Instruct \(temperature0\.90\.9, up to three attempts per item\) and accepted only if \(a\) the answer string is preserved verbatim \(deterministic check\), \(b\) unigram Jaccard overlap with the original question is at most0\.60\.6, and \(c\) two independently trained NLI judges,microsoft/deberta\-large\-mnliandFacebookAI/roberta\-large\-mnli, both assign bidirectional entailment probability above0\.80\.8; judge disagreements are rejected rather than adjudicated\. Of200200items,138138obtained an accepted paraphrase \(0\.690\.69\)\. The rejection counts are per generation*attempt*, not per item: across385385attempts \(up to three per item\),175175failed both NLI judges,7070split the judges, and22exceeded the overlap cap \(247247rejected attempts\+\+138138accepted items==385385\)\. Items without an accepted paraphrase are excluded from the paraphrase mechanism \(never silently retained\), and the paraphrase channel’s audits use only the138138accepted items\.
### G\.3Provenance caveat in full
“Pile\-trained” does not by itself certify that the base checkpoint never saw SQuAD, so we do*not*claim a clean\-versus\-contaminated contrast in the absolute sense\. What the pairedβtrain=0\\beta\_\{\\rm train\}\{=\}0vs\.βtrain\>0\\beta\_\{\\rm train\}\{\>\}0continuations measure is the*incremental*exposure effect relative to a matched continuation with identical background, order, optimizer state, token budget, and initialization—a well\-defined causal quantity even if the base model carries some prior exposure\. Establishing absolute cleanliness would require post\-cutoff or provably\-absent benchmark material with exact provenance and near\-duplicate removal, which we flag as the stronger design and leave to future work\. Likewise, the surface\-perturbation condition is an automatic lexical\-overlap reducer, not a human\-validated meaning\-preserving paraphrase; we label it accordingly and do not claim semantic invariance\.
### G\.4Statistical protocol details
Probe counts are
m∈\{16,32,64,128,256,512,1024\},m\\in\\\{16,32,64,128,256,512,1024\\\},truncated at the number of independent item clusters\. The fresh clean\-reference sizen0testn\_\{0\}^\{\\rm test\}is250250\(half of the500500held\-out nonmembers\), and each test reportsmeffm\_\{\\rm eff\}\. For each\(α,m\)\(\\alpha,m\),500500Monte Carlo audit replicates are drawn only from the held\-out audit pool\. Rejection\-rate intervals are Wilson intervals; model\- and seed\-level uncertainty is preserved by a hierarchical item/seed bootstrap\. For every condition we report both the primary permutation\-test rejection rate andℙ\(α¯\>0\)\\mathbb\{P\}\(\\underline\{\\alpha\}\>0\), plus their difference; the latter is expected to be smaller because it uses the conservative certificate rule\. In the full protocol the cross\-fitted combination is the intended primary score in each access regime; in the experiments reported here \(which do not run the combination\) the primary score is the calibration\-selected best single probe, and the remaining probes are secondary and Holm\-adjusted within channel\.
The empirical onsetα^⋆\(m\)\\widehat\{\\alpha\}\_\{\\star\}\(m\)is the smallest isotonic\-interpolatedα\\alphawhose lower power\-interval endpoint exceeds0\.800\.80while the upper endpoint of the corresponding size interval is at most0\.070\.07, a predeclared two\-point tolerance above the nominal level\. We fit
logα^⋆\(m\)=a\+blogm\\log\\widehat\{\\alpha\}\_\{\\star\}\(m\)=a\+b\\log mwith a seed\-clustered bootstrap interval forbb\. The theory predictsb=−1/2b=\-1/2only wheree^f\\widehat\{e\}\_\{f\}and the conditional channel remain stable\. We compare predicted and observedlogmtarget\\log m\_\{\\mathrm\{target\}\}using slope, intercept,R2R^\{2\}, median multiplicative error, and calibration plots\. The transport decision uses the predeclared criteria in[section˜4\.4](https://arxiv.org/html/2608.07914#S4.SS4); a failed decision triggers the nontransport analysis rather than a post\-hoc recalibration of the main claim\.
## Appendix HSecondary Scaling Diagnostic
After efficacy is frozen, rescaling bye^f\\widehat\{e\}\_\{f\}andmeffm\_\{\\rm eff\}substantially collapses the held\-out power curves onto the predicted information coordinateαe^fmeff\\alpha\\widehat\{e\}\_\{f\}\\sqrt\{m\_\{\\rm eff\}\}: power correlates with this coordinate atr=0\.68r=0\.68–0\.810\.81across probes \(the “Power corr\.” column of[table˜7](https://arxiv.org/html/2608.07914#A11.T7)\)\. The detectability onset scales asα⋆\(m\)∝m−0\.57\\alpha\_\{\\star\}\(m\)\\propto m^\{\-0\.57\}, consistent with the theoreticalm−1/2m^\{\-1/2\}prediction of §[3](https://arxiv.org/html/2608.07914#S3); mechanism\-specific results are in[table˜12](https://arxiv.org/html/2608.07914#A11.T12)\. We nonetheless report scaling as a diagnostic rather than the headline, because the collapse is approximate and the onset exponent, while close to−1/2\-1/2, is estimated from a single model’s channel; the prospective power prediction is the primary, model\-independent claim\. Because the onset ande^f\\widehat\{e\}\_\{f\}both arise from a standardized mean shift, agreement with−1/2\-1/2mainly checks whether the local approximation and channel\-stability assumptions hold\. It is not treated as independent evidence for the framework\.
## Appendix ICertificate Coverage at Audit Scale in Full
Two coverage numbers must not be conflated: the99\.099\.0–99\.5%99\.5\\%and0\.9380\.938–0\.9530\.953figures in[table˜9](https://arxiv.org/html/2608.07914#A11.T9)are endpoints of a*scaled\-slack heuristic on unclipped scores*, not validation of the distribution\-free certificate\. The exact certificate of Theorem[8](https://arxiv.org/html/2608.07914#Thmtheorem8), on a clipped\[0,1\]\[0,1\]score frozen on a disjoint calibration half \(measured marginsΔ^f=0\.11\\widehat\{\\Delta\}\_\{f\}=0\.11–0\.260\.26\), never fires atm=512m\{=\}512—zero false, zero nonzero certificates in all six channels; its100%100\\%coverage is*vacuously satisfied*\. Certifyingα=0\.1\\alpha\{=\}0\.1–0\.20\.2at measured margins needs balanced groups of∼3,500\\sim 3\{,\}500–79,00079\{,\}000items \(§[J](https://arxiv.org/html/2608.07914#A10); the v3 channels agree,[table˜3](https://arxiv.org/html/2608.07914#A11.T3)\): a nonzeroα¯\\underline\{\\alpha\}is trustworthy, a zero is uninformative, and at audit sizes the certificate is a null instrument, reported as a limitation\.
The measured power onset and the analytic certificate resolution are distinct size\-dependent thresholds \(an earlier draft conflated them; corrected in[table˜8](https://arxiv.org/html/2608.07914#A11.T8)\): then0=250n\_\{0\}\{=\}250reference floors the resolution at≈0\.17\\approx 0\.17, so no audit\-set size alone certifiesα=0\.1\\alpha\{=\}0\.1, while the measured onset falls to≈0\.1\\approx 0\.1bym=1,024m\{=\}1\{,\}024in a pool\-limited cell \(full arithmetic in[appendix˜J](https://arxiv.org/html/2608.07914#A10)\)\.
Scaling is a*diagnostic*only: power tracksαe^fmeff\\alpha\\widehat\{e\}\_\{f\}\\sqrt\{m\_\{\\rm eff\}\}\(r=0\.68r=0\.68–0\.810\.81\), onset slope−0\.57\-0\.57vs\. the predicted−1/2\-1/2; the collapse checks the local approximation, not the framework \([appendix˜H](https://arxiv.org/html/2608.07914#A8),[table˜12](https://arxiv.org/html/2608.07914#A11.T12)\)\.
## Appendix JPower Onset vs\. Certificate Resolution
An earlier draft conflated the measured power onset with the analytic certificate resolutionα⋆cert\(m,n0\)=\(εcert\(m\)\+εcert\(n0\)\)/Δf\\alpha\_\{\\star\}^\{\\rm cert\}\(m,n\_\{0\}\)=\(\\varepsilon\_\{\\rm cert\}\(m\)\+\\varepsilon\_\{\\rm cert\}\(n\_\{0\}\)\)/\\Delta\_\{f\}and omitted then0n\_\{0\}term; both are corrected in[table˜8](https://arxiv.org/html/2608.07914#A11.T8)\. Even under the best\-case boundΔf≤0\.58\\Delta\_\{f\}\\leq 0\.58\(measured margins are22–5×5\\timessmaller\), the protocol’sn0=250n\_\{0\}\{=\}250reference*floors*the resolution at≈0\.17\\approx 0\.17: no audit\-set size alone certifiesα=0\.1\\alpha\{=\}0\.1; a balanced design needsm=n0≈2,850m\{=\}n\_\{0\}\\approx 2\{,\}850best\-case and≈14,000\\approx 14\{,\}000–79,00079\{,\}000at measured margins—the honest planning number\. The measured onset falls to≈0\.1\\approx 0\.1bym=1,024m\{=\}1\{,\}024, but that cell is bootstrap\-resampled from≤500\\leq 500distinct held\-out items with a plug\-in clean reference, hence*pool\-limited*—a secondary diagnostic, never ann0=250n\_\{0\}\{=\}250result \(exact counts in the[table˜8](https://arxiv.org/html/2608.07914#A11.T8)caption\)\. The onset stays far below the certificate resolution at every size—the price of the orientation\-free construction\.
## Appendix KAdditional Result and Planning Tables
Tables[1](https://arxiv.org/html/2608.07914#A11.T1)–[13](https://arxiv.org/html/2608.07914#A11.T13)provide the supplementary per\-seed injection results, validity\-gate diagnostics, finite\-sample planner evaluations, power\-prediction results, certificate\-resolution and coverage analyses, calibration\-to\-audit transport checks, sensitivity estimates, and empirical scaling diagnostics\.
Table 1:Per\-seed five\-seed paired injection results \(Evaluation B\)\. Paired contrastDsD\_\{s\}= \(exposed−\-held gap, contaminated\)−\-\(same gap, matched clean continuation\) with item\-bootstrap95%95\\%intervals; the*common frozen probe*block \(mean NLL, the modal calibration winner\) is the basis of the ordering claim because the per\-cell best probe \(second block\) differs across cells and mixes probes\. On the common probeDs\>0D\_\{s\}\>0with intervals excluding zero in all1515exact/paraphrase/surface cells, and every answer\-only interval covers zero\. Adjacent seed\-level paired differences on the common probe \(mean±\\pmsd over seeds; positive in5/55/5seeds each, sign\-testp=2−5=0\.031p=2^\{\-5\}=0\.031, Holm\-adjusted over the three comparisonspadj=0\.094p\_\{\\rm adj\}=0\.094\): exact−\-paraphrase0\.31±0\.120\.31\\pm 0\.12, paraphrase−\-surface0\.26±0\.100\.26\\pm 0\.10, surface−\-answer\-only0\.45±0\.080\.45\\pm 0\.08\. \(Best\-probe adjacent differences, for reference:0\.36±0\.150\.36\\pm 0\.15,0\.25±0\.140\.25\\pm 0\.14,0\.68±0\.320\.68\\pm 0\.32\.\) Bottom: empirical size of the exact permutation test atα=0\\alpha\{=\}0per seed and mechanism \(Wilson intervals in the released JSON\); bold entries exceed the predeclared0\.070\.07tolerance and those channel–seed cells fail the size gate: exact/surface pass in4/54/5and5/55/5seeds, paraphrase and answer\-only in3/53/5—a further reason answer\-only supports no detection claim\.Table 2:Measured validity gates and locked transport validation for the six upgraded Evaluation A channels \(best probe; exact permutation test\)\. Every channel FAILS overall because the blind text\-only AUC far exceeds the0\.750\.75content\-shift limit—the pools come from different corpora—while the KS gates, cluster counts, and empirical size pass\. The last column reports the frozen\-inflation locked validation \(§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\): three channels are meta\-calibration; on the three locked channels the inflated budgetminfl=1\.40×mpredm\_\{\\rm infl\}=1\.40\\times m\_\{\\rm pred\}is compared against the dyadic\-grid onsetmobsm\_\{\\rm obs\}, which only localizes the true onset to\(mobs/2,mobs\]\(m\_\{\\rm obs\}/2,\\,m\_\{\\rm obs\}\]\. Reading with the censoring interval: the two43/61/6443/61/64channels are*unresolved*\(61∈\(32,64\]61\\in\(32,64\]\), the72/101/25672/101/256channel is a*clear failure*\(101∉\(128,256\]101\\notin\(128,256\]\)—so11failure,22censored, not0/30/3coverage\.Table 3:v3 same\-corpus continued\-pretraining audit: random member/nonmember split of one corpus \(700\+700700\{\+\}700\), members trained into the model \(k=4k\{=\}4,1,2001\{,\}200steps\), so pools are text\-exchangeable by construction\. Blind AUC sits at or near chance under the pre\-specified rule \(interval contains0\.50\.5or point≤0\.55\\leq 0\.55;[appendix˜F](https://arxiv.org/html/2608.07914#A6)\);4/64/6channels pass all gates and the two Pythia×\\timesSQuAD channels fail the*size*gate \(bold\) and are excluded from confirmatory analysis\.Δ^f\\widehat\{\\Delta\}\_\{f\}is the measured clipped\-score margin with the implied balancedm=n0m\{=\}n\_\{0\}for certifyingα=0\.1\\alpha\{=\}0\.1in parentheses; the exact certificate never fires atm=100m\{=\}100\(coverage vacuously1\.01\.0, zero nonzero certificates\)\. The last column gives*direct, censoring\-free*power at the frozen budgetsmpredm\_\{\\rm pred\}andminfl=1\.40×mpredm\_\{\\rm infl\}\{=\}1\.40\\times m\_\{\\rm pred\}\(200200trials, Wilson95%95\\%\): atmpredm\_\{\\rm pred\}all three intervals exclude the0\.80\.8target \(three clear failures\); atminflm\_\{\\rm infl\}two exclude it and one is unresolved \(0\.760\.76\[0\.70,0\.81\]\[0\.70,0\.81\]\)\.∗Unattainable:Meff≥n0=250M\_\{\\rm eff\}\\geq n\_\{0\}\{=\}250\.Table 4:Complete v4 finite\-sample planner record \(fresh same\-corpus continued\-pretraining channels;700\+700700\{\+\}700members/nonmembers, calibration/audit halves of350350clusters each,α0=0\.2\\alpha\_\{0\}\{=\}0\.2,n0=250n\_\{0\}\{=\}250,199199\-permutation add\-one test, level0\.050\.05; planner:B=30B\{=\}30outer item bootstraps,6060inner Monte\-Carlo trials per replicate, lower bound =0\.100\.10quantile,mm\-grid up to192192; evaluation:200200fresh trials with Wilson95%95\\%intervals\)\. Best probe frozen by the calibration rule\.mpredm\_\{\\rm pred\}is the Gaussian plug\-in budget;mplanm\_\{\\rm plan\}the planner budget; “abstain” means the bootstrap lower bound never reached0\.800\.80on the grid—on pythia\-160m×\\timeswiki the planner abstains while the large Gaussian budget succeeds \(m=190m\{=\}190:0\.990\.99\), a conservative failure reported straight\.†Passes the pre\-specified blind rule via the point margin \(≤0\.55\\leq 0\.55\) although the upper CI exceeds0\.550\.55; under a proper equivalence gate \(upper CI≤0\.55\\leq 0\.55\) these channels are excluded and the head\-to\-head rests on gpt\-neo\-125m×\\timeswiki alone\.∗Unattainable:Meff≥n0M\_\{\\rm eff\}\\geq n\_\{0\}\.mplanm\_\{\\rm plan\}selects the smallestmmfrom pointwise bootstrap bounds and is not selection\-valid simultaneous inference \(§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\)\.Table 5:Complete v5 planner record:1212same\-corpus continued\-pretraining channels \(700\+700700\{\+\}700members/nonmembers,k=4k\{=\}4,1,2001\{,\}200steps; item splits cal\-A/cal\-B/audit=175/175/350=175/175/350clusters;α0=0\.2\\alpha\_\{0\}\{=\}0\.2,n0=250n\_\{0\}\{=\}250,199199\-permutation add\-one test at level\.05\.05\)\. Gate: predeclared one\-sided equivalence rule \(blind AUC upper95%95\\%CI≤0\.55\\leq 0\.55\) plus KS and size gates,9/129/12pass; the stricter two\-sided sensitivity gate \(entire CI in\[0\.45,0\.55\]\[0\.45,0\.55\]\) passes8/128/12, additionally excluding pythia\-70m×\\timeswiki; failures marked, excluded from endpoint aggregates; per\-channel KS statistics and size Wilson intervals in[table˜6](https://arxiv.org/html/2608.07914#A11.T6)\. Probe and standardization frozen on cal half A \(e^fA\\widehat\{e\}\_\{f\}^\{A\}= cal\-A efficacy\)\. Stage A:B=30B\{=\}30item bootstraps×\\times6060Monte\-Carlo trials per replicate at each gridmm\(grid to192192\);mA=min\{m:0\.10\-quantile of replicate power≥0\.80\}m\_\{A\}=\\min\\\{m:\\text\{$0\.10$\-quantile of replicate power\}\\geq 0\.80\\\}; the90%90\\%m^⋆\\widehat\{m\}\_\{\\star\}interval is the\[q\.05,q\.95\]\[q\_\{\.05\},q\_\{\.95\}\]of per\-replicate minimal budgets and is*descriptive*\(B=30B\{=\}30; probe selection not re\-run\)\. Stage B \(rung\-selection\-valid conditional on the empirical cal\-B distribution; cal\-B’s175175clean clusters are resampled with replacement to reachn0=250n\_\{0\}\{=\}250, so score\-law uncertainty, cluster reuse, and cal\-to\-audit transport are outside the guarantee\): pre\-specified ladder\{mA,⌈1\.5mA⌉,2mA\}\\\{m\_\{A\},\\lceil 1\.5m\_\{A\}\\rceil,2m\_\{A\}\\\}on untouched cal half B,300300trials each, one\-sided Bonferroni\-corrected \(α=0\.10/3\\alpha\{=\}0\.10/3\) lower bounds;mplanm\_\{\\rm plan\}= smallest rung clearing0\.800\.80; escalation occurred in3/83/8channels\. Audit half: a dense power curve gives the*estimated*onsetm^⋆\\widehat\{m\}\_\{\\star\}\(200200trials per rung; a noisy threshold, not ground truth—e\.g\. pythia\-70m×\\timescnn hasmplan=m^⋆=56m\_\{\\rm plan\}=\\widehat\{m\}\_\{\\star\}=56withπ^\(56\)=0\.785\\widehat\{\\pi\}\(56\)=0\.785\); Pow@pred/Pow@plan are fresh200200\-trial powers atmpredm\_\{\\rm pred\}\(Gaussian plug\-in from cal A\) andmplanm\_\{\\rm plan\}\. On pythia\-70m×\\timeswiki stage A never certified0\.800\.80\(abstention;90%90\\%interval censored at the grid edge\) while the Gaussian budget delivers0\.010\.01power—the cal\-selected*neighborhood*probe does not transport, and the planner’s abstention, not the plug\-in’s prescription, is the correct behavior\. Aggregates over gate\-passing channels: Gaussian clear failures9/99/9\(mpred/m^⋆m\_\{\\rm pred\}/\\widehat\{m\}\_\{\\star\}median0\.590\.59over the88channels with definedm^⋆\\widehat\{m\}\_\{\\star\}\); planner covered6/86/8, two unresolved, zero clear failures, conservatism median1\.361\.36, interval coverage8/88/8\(descriptive, subject to them^⋆\\widehat\{m\}\_\{\\star\}estimation noise above\)\.Table 6:Per\-channel v5 gate values \(as\-run thresholds, all auditable from the released JSON\): blind AUC upper95%95\\%CI≤0\.55\\leq 0\.55\(predeclared one\-sided equivalence form\), clean\-transport KS<0\.15<0\.15, channel\-stability KS<0\.15<0\.15, empirical\-size Wilson lower bound≤0\.05\\leq 0\.05\(nsizen\_\{\\rm size\}trials atα=0\\alpha\{=\}0\)\. The one\-sided gate passes9/129/12; the stricter two\-sided sensitivity gate \(entire AUC CI within\[0\.45,0\.55\]\[0\.45,0\.55\]\) passes8/128/12, additionally excluding pythia\-70m×\\timeswiki \(lower CI0\.447<0\.450\.447<0\.45: possible reverse\-direction separation\)—the same channel where the planner abstains\. The as\-run size rule has the*detection*direction \(a channel fails only when inflation is conclusively detected\); a size\-*equivalence*rule demanding the Wilson upper bound≤0\.07\\leq 0\.07\(the paper’s declared size tolerance\) is stricter, and combining it with the two\-sided AUC gate leaves4/124/12channels clearly passing all gates \(pythia\-70m×\\timescnn, gpt\-neo\-125m×\\timeswiki/cnn, gpt2×\\timeswiki\); on those four the planner endpoint reads3/43/4at target, one unresolved, zero clear failures\. No KS value approaches the0\.150\.15threshold \(max0\.1030\.103\); every size interval covers the nominal0\.050\.05\.Table 7:Prospective prediction \(primary endpoint\): audit power predicted from the*frozen calibration efficacy*matches held\-out power across the tested\(α,m\)\(\\alpha,m\)grid withR2=0\.81R^\{2\}=0\.81–0\.990\.99and mean absolute power error0\.020\.02–0\.100\.10; the correlation of held\-out power with the theoretical coordinateαe^fm\\alpha\\,\\widehat\{e\}\_\{f\}\\sqrt\{m\}is0\.680\.68–0\.810\.81\. The last column is the*heuristic scaled\-slack*interval’s empirical coverage on unclipped scores \(§[6\.3](https://arxiv.org/html/2608.07914#S6.SS3)\), a measured diagnostic of that heuristic—*not*coverage of the distribution\-free prevalence certificate, which never fires at these sample sizes\.R2R^\{2\}alone is not claimed to establish calibration; the MAE column bounds systematic prediction bias directly\.Table 8:Measured power onset vs\. analytic certificate resolution onpythia\-1\.4b\(e^f=1\.16\\widehat\{e\}\_\{f\}=1\.16, best\-caseΔf≤0\.58\\Delta\_\{f\}\\leq 0\.58,δcert=0\.05\\delta\_\{\\mathrm\{cert\}\}=0\.05,εcert\(n\)=log\(6/δcert\)/\(2n\)\\varepsilon\_\{\\rm cert\}\(n\)=\\sqrt\{\\log\(6/\\delta\_\{\\mathrm\{cert\}\}\)/\(2n\)\}from Theorem[8](https://arxiv.org/html/2608.07914#Thmtheorem8)\)\. Certificate rows are population\-margin*planning thresholds under the margin upper bound*, not measured onsets: the measuredΔ^f\\widehat\{\\Delta\}\_\{f\}of the clipped audit score \(0\.110\.11–0\.260\.26across the upgraded channels\) can only be smaller, which makes every threshold larger: under the measured margins the balancedα=0\.1\\alpha\{=\}0\.1crossing moves from≈2,850\\approx 2\{,\}850to≈14,000\\approx 14\{,\}000–79,00079\{,\}000per group\. The protocol row fixesn0=250n\_\{0\}\{=\}250, giving0\.230\.23atm=2,000m\{=\}2\{,\}000,0\.210\.21atm=4,000m\{=\}4\{,\}000, and floorεcert\(250\)/Δf≈0\.17\\varepsilon\_\{\\rm cert\}\(250\)/\\Delta\_\{f\}\\approx 0\.17asm→∞m\\to\\infty; the balanced row is a*hypothetical*design withm=n0m\{=\}n\_\{0\}, reaching0\.10\.1nearm=n0≈2,850m\{=\}n\_\{0\}\\approx 2\{,\}850\. The power onset is the measured0\.80\.8\-power fraction with fitted slopem−0\.57m^\{\-0\.57\}overm≤1,024m\\leq 1\{,\}024, matching them−1/2m^\{\-1/2\}law\. Exact design counts for this initial\-audit onset, disclosed in full: the channel is500\+500500\{\+\}500items, split250/250250\{/\}250per class into calibration and held\-out halves; audit sets are bootstrap\-resampled*with replacement*from the≤500\\leq 500distinct held\-out items, and the test is a standardized\-mean threshold1\.645/m1\.645/\\sqrt\{m\}with the held\-out clean half’s\(μ0,σ0\)\(\\mu\_\{0\},\\sigma\_\{0\}\)plugged in as known \(no finite\-reference penalty\)\. Atm=1,024m\{=\}1\{,\}024the audit set therefore contains at most500500distinct items: theα=0\.1\\alpha\{=\}0\.1onset is*pool\-limited and secondary*, and must not be read as1,0241\{,\}024independent items under any convention\. The confirmatory measurements are the exact\-permutation audits \([tables˜2](https://arxiv.org/html/2608.07914#A11.T2)and[3](https://arxiv.org/html/2608.07914#A11.T3)\), which draw without replacement and capmmat the independent\-cluster count\.Table 9:*Heuristic*scaled\-slack lower\-bound coverage at trueα=0\.2\\alpha=0\.2—not the exact certificate of Theorem[8](https://arxiv.org/html/2608.07914#Thmtheorem8), whose validation this table does not provide \(see §[6](https://arxiv.org/html/2608.07914#S6)\)—across the Evaluation A post\-hoc checkpoint mixture \(pythia\-1\.4b\) and the Evaluation B causal injections \(continued\-pretraining on pythia\-160m under three mechanisms; best\-probe coverage pooled over three training seeds,150150Monte\-Carlo trials per seed\)\. Pooled counts and Wilson95%95\\%intervals: exact428/450=0\.951428/450=0\.951\[0\.927,0\.967\]\[0\.927,0\.967\]\(per\-seed0\.980,0\.947,0\.9270\.980,0\.947,0\.927\); surface422/450=0\.938422/450=0\.938\[0\.912,0\.957\]\[0\.912,0\.957\]\(0\.933,0\.933,0\.9470\.933,0\.933,0\.947\); answer\-only429/450=0\.953429/450=0\.953\[0\.930,0\.969\]\[0\.930,0\.969\]\(0\.960,0\.947,0\.9530\.960,0\.947,0\.953\)\. Every interval contains the nominal0\.950\.95, so the sub\-nominal point values are statistically compatible with Monte\-Carlo variation; we do not attribute them to any channel property\. The completed runs implement the certificate with a scaled\-slack approximation on unclipped scores rather than the exact clipped\-\[0,1\]\[0,1\]construction of Theorem[8](https://arxiv.org/html/2608.07914#Thmtheorem8), so empirical coverage is a measured endpoint here, not a by\-construction guarantee\. Detectability falls sharply and stably with the injection mechanism—exact \(e^f=0\.741\\widehat\{e\}\_\{f\}=0\.741; per\-seed0\.645,0\.862,0\.7160\.645,0\.862,0\.716\)\>\>surface perturbation \(0\.3430\.343;0\.360,0\.355,0\.3130\.360,0\.355,0\.313\)\>\>answer\-only \(0\.1500\.150;0\.119,0\.152,0\.1780\.119,0\.152,0\.178\), an ordering that holds in all three seeds—exactly the surface\-form dependence the channel predicts\.Table 10:Evaluation A calibration\-to\-audit transport on EleutherAI/pythia\-1\.4b \(Pile members vs\. held\-out nonmembers,500\+500500\{\+\}500items\)\.e^fcal\\widehat\{e\}\_\{f\}^\{\\rm cal\}is the*frozen calibration*efficacy;Meffcal=\(2\.486/\(α0e^fcal\)\)2M^\{\\rm cal\}\_\{\\rm eff\}=\(2\.486/\(\\alpha\_\{0\}\\widehat\{e\}\_\{f\}^\{\\rm cal\}\)\)^\{2\}is the prospective effective sample size atα0=0\.2\\alpha\_\{0\}=0\.2, size\.05\.05, power\.8\.8;mpredcal=⌈Meffcaln0/\(n0−Meffcal\)⌉m^\{\\rm cal\}\_\{\\rm pred\}=\\lceil M^\{\\rm cal\}\_\{\\rm eff\}n\_\{0\}/\(n\_\{0\}\-M^\{\\rm cal\}\_\{\\rm eff\}\)\\rceilis the required suspect\-set size under a*finite*clean reference ofn0=250n\_\{0\}\{=\}250items \(“—” whenMeffcal≥n0M^\{\\rm cal\}\_\{\\rm eff\}\\geq n\_\{0\}, i\.e\.*unattainable*with this reference regardless of suspect\-set size\)\.e^ftest\\widehat\{e\}\_\{f\}^\{\\rm test\}is the audit\-split efficacy—a*diagnostic*, never an input tompredcalm^\{\\rm cal\}\_\{\\rm pred\}\. The last column is the*efficacy\-transport ratio*exp\(2\|loge^fcal−loge^ftest\|\)\\exp\(2\|\\log\\widehat\{e\}\_\{f\}^\{\\rm cal\}\-\\log\\widehat\{e\}\_\{f\}^\{\\rm test\}\|\), the symmetric factor by which the between\-split efficacy shift would move the external\-null effective size\.Mefftest†\{\}^\{\\dagger\}M^\{\\rm test\}\_\{\\rm eff\}is a plug\-in inversion of the audit\-split efficacy,*not*an independently observed0\.80\.8\-power sample size from permutation curves \(which we have not measured here\)\. The median efficacy\-transport ratio is2\.23×2\.23\\timesand3/53/5channels exceed the pre\-specified2×2\\timestolerance, so the pooled calibration\-to\-audit transport claim is*withdrawn*\(§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\); under the finite clean reference the attainable\-mmgap is generally larger still \(mean NLL9696vs\.≈1,283\\approx 1\{,\}283; allMeffM\_\{\\rm eff\}values in this paper use thezz\-sumz\.95\+z\.80=2\.486z\_\{\.95\}\+z\_\{\.80\}=2\.486—higher\-precisionzzvalues move this figure to≈1,289\\approx 1\{,\}289, immaterially\)\.Table 11:Approximate smallest detectableα\\alphaat size \.05 and 80% power with an externally calibrated null\. Entries above one mean that even full exposure is underpowered under the local approximation\.Table 12:Secondary scaling diagnostic:bbis the fitted slope oflogα^⋆\(m\)\\log\\widehat\{\\alpha\}\_\{\\star\}\(m\)onlogm\\log m, predicted−1/2\-1/2\. The strong channels \(Eval\. A checkpoint mixture, Eval\. B exact injection\) recover slopes near−1/2\-1/2; the surface and answer\-only injections \(the two subthreshold rows\) are too weak to reach0\.80\.8power at the testedmm, so no onset is defined—reported straight rather than extrapolated\. Values from the audit and injection summaries\.Table 13:Point efficacy \(mean over three seeds\) versus the variance\-adaptive95%95\\%floore¯f\\underline\{e\}\_\{f\}on the Evaluation B injections \(n=200n\{=\}200exposed/held\-out per mechanism\)\. The floor is0at this sample size: with only200200items the simultaneous Hoeffding slack exceeds the mean gap, so the*certified*lower bound onρ\\rhois vacuous even where the point efficacy is clearly nonzero—an honest small\-sample limitation of the distribution\-free floor, not of the point estimate\. Larger calibration sets are required for a nonvacuous floor\.
## Appendix LRetrospective Pilot Details
The initial pilot used Pythia\-1\.4B with 400 labeled Pile members and 400 nominal nonmembers\. It reported Cohen’sd=0\.585d=0\.585and, after constructing post\-hoc mixtures withm∈\{4,8,16,32,64,128,256\}m\\in\\\{4,8,16,32,64,128,256\\\}, onsets
\(\.80,\.60,\.45,\.30,\.20,\.15,\.10\),\(\.80,\.60,\.45,\.30,\.20,\.15,\.10\),whose log–log slope is−0\.506\-0\.506\(R2=0\.997R^\{2\}=0\.997\), reproduced from the committed pilot log \(results\_pythia/pythia\_contamination\_summary\.json\)\. All Monte Carlo mixtures reuse the same 800 controls, so their errors are correlated and the quotedR2R^\{2\}is optimistic\. Moreover, the mixture was formed after training, and the member/nonmember split was not protected against temporal, deduplication, and source confounding\. The pilot is therefore an implementation check and planning illustration, not evidence for causal contamination or an independent validation of the scaling law\.
## Appendix MReproducibility Protocol
### M\.1Scope of this protocol
We do not release code or result artifacts with this submission\. This appendix therefore specifies the experimental protocol at the level of detail required to re\-implement it: model checkpoints and data sources, the split structure, the estimators and their tuning constants, the Monte\-Carlo and bootstrap replicate counts, and the analysis decisions fixed before outcomes were computed\. Every quantity reported in §[6](https://arxiv.org/html/2608.07914#S6)is defined by the procedures below together with the constants tabulated in[table˜14](https://arxiv.org/html/2608.07914#A13.T14)and the per\-channel gate values in[table˜6](https://arxiv.org/html/2608.07914#A11.T6)\.
### M\.2Data and model specification
Reproduction requires: the base checkpoints \(Pythia\-70m/160m/410m/1\.4b, GPT\-Neo\-125m/1\.3B, GPT\-2, OPT\-125m, all from public releases\), the corpora used to build channels \(Wikipedia, CNN/DailyMail news, and SQuAD for the injection arm\), and the Pile background corpus for continued pretraining\. Channel construction is specified in §[5](https://arxiv.org/html/2608.07914#S5): the v3 and v5 audits split one corpus into700\+700700\{\+\}700member/nonmember halves and continue pretraining on the member half, so pools are text\-exchangeable by construction and the split is reproducible from a seed alone\.
### M\.3Per\-item record structure
Any re\-implementation should retain, for every item, the fields that the analysis depends on:
- •benchmark, split, and source\-document cluster;
- •membership provenance and training\-stream position;
- •deduplication and near\-duplicate cluster identifiers;
- •contamination mechanism, exposure count, and run/seed identifier;
- •probe prompts, decoding settings, raw scores, and access regime;
- •assignment to score\-training, calibration, or audit split\.
The cluster identifiers are load\-bearing rather than bookkeeping: the effectivemmcounts source\-item clusters, so a re\-implementation that treats near\-duplicates or repeated decodes as independent probes will report inflated power at every budget\.
### M\.4Training controls
[Table˜14](https://arxiv.org/html/2608.07914#A13.T14)summarizes the checkpoints, optimization settings, contamination construction, matched controls, stopping rule, and seeds used in the injection experiments\.
Table 14:Minimum training\-run disclosure\.
### M\.5Primary analysis decisions
Throughout the paper, “predeclared” means fixed in our run scripts and this protocol before the corresponding outcomes were computed; there is no public timestamped preregistration registry entry, so we do not use the term “pre\-registered\.”
1. 1\.The primary endpoint is benchmark\-level power of the frozen one\-sided permutation test at nominal size0\.050\.05, not instance AUROC orα¯\>0\\underline\{\\alpha\}\>0\.
2. 2\.The unit of independence is the source\-item cluster\.
3. 3\.Detector orientation and hyperparameters are selected without the audit split\.
4. 4\.Settings failing the blind\-separation or channel\-stability gates are not pooled into the confirmatory scaling fit\.
5. 5\.The transport claim is withdrawn if either predeclared prediction criterion in[section˜4\.4](https://arxiv.org/html/2608.07914#S4.SS4)fails; no audit\-split recalibration repairs that decision\.
6. 6\.Every zero\-contamination run is included in size and false\-certificate estimates\.
7. 7\.Missing or failed runs and all exclusions are reported by seed and condition\.
## Appendix NRelated Work in Full
This section expands §[7](https://arxiv.org/html/2608.07914#S7)\. We organize by what each line of work*delivers*rather than by method family, since the paper’s contribution is a deliverable \(prospective power, and a lower bound onα\\alpha\) rather than a new detector or a new mixture bound\.
### N\.1Behavioral detectors
The dominant evidence channel for contamination is behavioral: a statistic of the model’s response to a benchmark item, compared against some reference\. Loss and its compression\-normalized variants\(Carliniet al\.,[2021](https://arxiv.org/html/2608.07914#bib.bib5)\)are the baseline; Min\-K% Prob\(Shiet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib29)\)and Min\-K%\+\+\(Zhanget al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib36)\)threshold low\-probability token subsets; neighborhood comparison\(Matternet al\.,[2023](https://arxiv.org/html/2608.07914#bib.bib21)\)replaces an external reference with perturbed counterfactuals; ReCaLL\(Xieet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib33)\)contrasts conditional log\-likelihoods under prefixes; CAMIA\-style features\(Changet al\.,[2025](https://arxiv.org/html/2608.07914#bib.bib6)\)exploit context\-dependence of memorization\. Black\-box variants elicit exposure through canonical\-order exchangeability\(Orenet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib23)\)or direct recall quizzing\(Golchin and Surdeanu,[2025](https://arxiv.org/html/2608.07914#bib.bib14)\)\.
In our framework each of these is a scalar scoreff, and Proposition[5](https://arxiv.org/html/2608.07914#Thmtheorem5)reduces it to a single number, the efficacyef=\|𝔼1f−𝔼0f\|/σf,0≤ρe\_\{f\}=\|\\mathbb\{E\}\_\{1\}f\-\\mathbb\{E\}\_\{0\}f\|/\\sigma\_\{f,0\}\\leq\\rho, which upper\-bounds the local power of the oriented mean\-score audit built fromff\. This is deliberately lossy:efe\_\{f\}does not summarize every test computable fromff, and two detectors with equal AUROC can have different low\-α\\alphaefficacy\. What the reduction buys is that detector comparison, channel comparison, and budget planning become the same measurement\. We run five gray\-box families \(§[G](https://arxiv.org/html/2608.07914#A7)\); ReCaLL, CAMIA\-style, and black\-box probes are specified in the protocol but not run, and no reported result rests on them\.
### N\.2The reliability gap
A second literature evaluates detectors rather than proposing them, and reports a consistent picture of fragility\.Duanet al\.\([2024](https://arxiv.org/html/2608.07914#bib.bib12)\)find membership inference near chance on LLMs at realistic scales;Fuet al\.\([2025](https://arxiv.org/html/2608.07914#bib.bib13)\)audit the assumptions detection methods rely on and find them frequently violated;Meeuset al\.\([2025](https://arxiv.org/html/2608.07914#bib.bib22)\)systematize the field’s evaluation failures;Samuelet al\.\([2025](https://arxiv.org/html/2608.07914#bib.bib27)\)document inconsistency across modern models and the absence of a usable oracle;Dekonincket al\.\([2024a](https://arxiv.org/html/2608.07914#bib.bib8)\)show that contamination detection is cheap to evade;Wanget al\.\([2026a](https://arxiv.org/html/2608.07914#bib.bib32)\)extend this to reasoning models\. Most directly,Zarzeckiet al\.\([2026](https://arxiv.org/html/2608.07914#bib.bib34)\)run 25 models and record only 201 of 335 audit outcomes correct, with distribution shift producing false positives and post\-hoc inference run at sample sizes too small to support its conclusions\.
Our reading of this literature is that it diagnoses two distinct failures that are usually reported together:*invalidity*\(the clean reference does not match the suspect channel, so separation reflects content rather than exposure\) and*underpower*\(the channel is informative but the audit was too small to resolve the exposed fraction\)\. The first is what our gates \(§[4\.4](https://arxiv.org/html/2608.07914#S4.SS4), Appendix[F](https://arxiv.org/html/2608.07914#A6)\) test; the second is what efficacy and the planner \(§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\) quantify\. A non\-rejection that distinguishes these two states is interpretable; a non\-rejection that does not is the object the reliability\-gap literature keeps finding\. The locked external validation on the open 25\-model release ofZarzeckiet al\.\([2026](https://arxiv.org/html/2608.07914#bib.bib34)\)—efficacy calibrated without their audit outcomes, predicting their documented failures—is specified but not yet run, and we do not claim it\.
### N\.3Aggregation methods and their estimands
A third line lifts item\-level scores to benchmark\- or dataset\-level decisions\. LLM Dataset Inference\(Mainiet al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib20)\)combines many membership features into a hypothesis test that a dataset was trained on\. PaCoST\(Zhanget al\.,[2024](https://arxiv.org/html/2608.07914#bib.bib35)\)constructs paired counterparts and tests for confidence asymmetry\.Puertoet al\.\([2025](https://arxiv.org/html/2608.07914#bib.bib24)\)study when and how membership attacks succeed as a function of scale and data quantity\. ConStat\(Dekonincket al\.,[2024b](https://arxiv.org/html/2608.07914#bib.bib9)\)estimates*performance inflation*by comparing a model’s benchmark score against a reference set of models on a matched benchmark\. FTD\(Zhanget al\.,[2026](https://arxiv.org/html/2608.07914#bib.bib37)\)supplies false\-discovery\-rate control over a set of contamination decisions\.
These are complementary, and Table[15](https://arxiv.org/html/2608.07914#A14.T15)makes the distinction explicit\. A validpp\-value certifies a rejection but is silent about whichα\\alphathe audit had power to detect—the quantity an auditor needs to interpret a non\-rejection\. Performance inflation measures how much a score is distorted, which depends on exposure fraction, mechanism, and the benchmark’s headroom jointly, and does not identify the fraction exposed\. FDR control governs the error rate of a discovery set across many benchmarks and says nothing about the resolution of a single audit\. Our two deliverables are the prospective statement \(what budget is required, and what a given budget can detect\) and the retrospective one \(a distribution\-free lower confidence bound onα\\alphaitself\)\. The paper’s main negative result is that the first is harder to deliver than the local theory suggests: efficacy transports, but its Gaussian inversion into a budget does not \(§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\)\.
Table 15:What each line of work returns\. “Prospective” means the quantity is computable*before*the audit split is opened and predicts what the audit will be able to detect\. All methods listed answer legitimate and different questions; the column that is empty above the rule is the one a non\-rejection needs in order to be interpretable\. Our prospective column carries the caveat established in §[6\.2](https://arxiv.org/html/2608.07914#S6.SS2): the efficacy\-only Gaussian budget is miscalibrated at smallmm, and the deliverable is the simulated planner’s budget, under a guarantee conditional on the empirical calibration distribution\.
### N\.4Validity: blind baselines and matched construction
Daset al\.\([2025](https://arxiv.org/html/2608.07914#bib.bib7)\)show that classifiers seeing only the raw text, with no access to the target model, match or exceed membership inference attacks on foundation models—the separation being audited is often a property of the pools, not of the model\.Wanget al\.\([2026b](https://arxiv.org/html/2608.07914#bib.bib31)\)formalize the corresponding construction requirement, drawing members and nonmembers from windows matched on source, length, packing, deduplication, and stream position around a training checkpoint\.Fuet al\.\([2025](https://arxiv.org/html/2608.07914#bib.bib13)\)reach the same conclusion from the assumptions side\.
We adopt this as a precondition with numerical thresholds fixed in the run scripts before scoring, rather than as a discussion\-section caveat, and we report gate values for every channel including the failures\. All six upgraded Evaluation A channels fail blind separation at text\-only AUC≈0\.95\\approx 0\.95, because their member and nonmember pools come from different corpora; we report them as a demonstration that the gate works, and rest the confirmatory reading on the same\-corpus continued\-pretraining channels, where blind AUC sits at0\.480\.48–0\.540\.54\. The stricter two\-sided equivalence form of the gate \(entire AUC interval within\[0\.45,0\.55\]\[0\.45,0\.55\]\) is reported as a sensitivity analysis in Table[6](https://arxiv.org/html/2608.07914#A11.T6)and excludes exactly the channel where the planner abstains\.
### N\.5Sparse\-mixture detection
The statistical content of theαρm\\alpha\\rho\\sqrt\{m\}boundary is classical\.Ingster \([1997](https://arxiv.org/html/2608.07914#bib.bib16)\)andDonoho and Jin \([2004](https://arxiv.org/html/2608.07914#bib.bib10)\)establish detection boundaries for sparse heterogeneous mixtures and the higher criticism statistic;Caiet al\.\([2011](https://arxiv.org/html/2608.07914#bib.bib3)\)extend to heteroscedastic mixtures;Cai and Wu \([2014](https://arxiv.org/html/2608.07914#bib.bib4)\)give optimal detection against a given null distribution\. The local asymptotic machinery we use for the matching test—LAN expansions, Le Cam’s third lemma, and the power envelope—is standard\(Le Cam,[1986](https://arxiv.org/html/2608.07914#bib.bib17); van der Vaart,[1998](https://arxiv.org/html/2608.07914#bib.bib30)\)\.
We claim no new mixture boundary\. Theorem[3](https://arxiv.org/html/2608.07914#Thmtheorem3)is an exact second\-moment calculation for the specific channelQα=\(1−α\)P0\+αP1Q\_\{\\alpha\}=\(1\-\\alpha\)P\_\{0\}\+\\alpha P\_\{1\}, sharpened only in that it passes throughχ2\\chi^\{2\}and total variation rather than KL and Pinsker, and Lemma[6](https://arxiv.org/html/2608.07914#Thmtheorem6)extends it to the fixed\-unknown\-subset design an auditor actually faces\. What is new here is the identification:ρ2=χ2\(P1∥P0\)\\rho^\{2\}=\\chi^\{2\}\(P\_\{1\}\\\|P\_\{0\}\)is not a nuisance parameter but the audit\-observable channel quantity, estimable from matched controls and lower\-bounded with finite\-sample certainty \(Corollary[9](https://arxiv.org/html/2608.07914#Thmtheorem9)\), and it replaces the informal capacity\-based arguments that appear in prior contamination discussions—capacity alone does not upper\-boundρ\\rhowithout a training\-stability theorem\.
### N\.6Prevalence estimation
Our certificate is a distribution\-free relative of classical prevalence estimation\.Rogan and Gladen \([1978](https://arxiv.org/html/2608.07914#bib.bib26)\)correct an observed screening\-test positive rate by known sensitivity and specificity, which is the population version of the ratio in Theorem[8](https://arxiv.org/html/2608.07914#Thmtheorem8)\. The modern learning\-theoretic treatment is mixture proportion estimation from positive and unlabeled data\(Blanchardet al\.,[2010](https://arxiv.org/html/2608.07914#bib.bib2); Scott,[2015](https://arxiv.org/html/2608.07914#bib.bib28); Ramaswamyet al\.,[2016](https://arxiv.org/html/2608.07914#bib.bib25)\), which achieves better rates under distributional assumptions such as irreducibility or separability\.
We deliberately give those rates up\. Our bound holds for any bounded score learned on independent data, with no orientation assumption—a reversed score returnsα¯=0\\underline\{\\alpha\}=0rather than a spurious positive—and no assumption on the score’s law beyond boundedness\. The price is quantified in Appendix[J](https://arxiv.org/html/2608.07914#A10)and is severe: with ann0=250n\_\{0\}\{=\}250reference the resolution floors at≈0\.17\\approx 0\.17, and certifyingα=0\.1\\alpha=0\.1at measured margins requires balanced groups of14,00014\{,\}000–79,00079\{,\}000items\. The certificate never fires at our audit sizes\. Whether an assumption\-bearing MPE estimator can supply a usable interval at realistic benchmark sizes, with coverage that survives the gates, is the natural next question and we do not answer it\.
### N\.7Power analysis in evaluation design
Sample\-size planning is routine in the empirical sciences and rare in NLP evaluation, where benchmark size is typically inherited from the dataset rather than chosen to achieve a target power against a stated alternative\. The contamination setting makes the omission costly, because the alternative of interest is small \(α\\alphaof a few percent\) and the per\-item signal is weak, so the regime is exactly the one where an unplanned audit returns an uninterpretable non\-rejection\. Our finding that the textbook Gaussian sample\-size formula is miscalibrated at the budgets it prescribes \(§[6\.2](https://arxiv.org/html/2608.07914#S6.SS2)\) is a caution that generalizes past contamination: when a planning equation prescribes smallmm, the approximation licensing it is at its worst, and the budget should be obtained by simulating the deployed test rather than by inverting a closed form\.Similar Articles
The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection
This paper identifies distribution shift and scale constraints as critical failure modes for statistical contamination detection methods in LLM benchmark auditing. Evaluating three paradigms across 27 models reveals only 199 correct outcomes out of 335 evaluations, indicating a systematic reliability gap that prevents these methods from replacing transparent data provenance.
Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits
This paper identifies five failure modes in perturbation-based benchmark-validity audits used for AI governance, demonstrating that implementation details can silently manufacture conclusions. It proposes a due-diligence gate to improve the reliability of evaluation evidence.
Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit
This paper adapts paired binary sample-size calculations to 4-bit quantization benchmarks, providing a conservative minimum detectable effect (MDE) bound that helps benchmark designers determine reliability before running experiments. A pilot audit shows that much of the observed variance across small subsamples is binomial sampling noise, not true model unreliability.
Life After Benchmark Saturation: A Case Study of CORE-Bench
This paper argues against the 'retire-and-replace' approach to saturated benchmarks, using CORE-Bench as a case study to demonstrate that measuring agent performance along dimensions such as construct validity, efficiency, reliability, and human-agent collaboration yields meaningful insights even after accuracy plateaus.
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
This paper introduces a framework for validating comparative LLM safety scoring without ground-truth labels, using an 'instrumental-validity chain' to establish deployment evidence. It demonstrates the method using a local-first tool called SimpleAudit on Norwegian safety packs and compares models like Borealis and Gemma 3.