Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration

arXiv cs.LG Papers

Summary

This paper presents a theoretical framework for self-poisoning in adaptive out-of-distribution detection, proving a sharp threshold for collapse and a certified label-free calibration method that severs the feedback loop. It also provides impossibility results for distinguishing drift from contamination without labels.

arXiv:2607.21673v1 Announce Type: new Abstract: Test-time adaptive out-of-distribution (OOD) detectors update a memory bank from the unlabelled stream. We show this adaptation obeys a provable dynamical law. Modelling bank impurity as a generalized P\'olya urn, we prove almost-sure convergence to a mean-field equilibrium whose slope acts as a reproduction number. Below one, impurity stays benign. Above one, the bank is fully poisoned and the detector collapses. The measured admission kernel is affine ($R^2 \ge 0.996$) with slope just below one in every encoder family (a protocol signature), so this detector class is near-critical by design, and across 96 settings the predicted threshold matches the empirical collapse, where ungated dictionaries lose up to $0.163$ AUROC. We then prove that a certified admission gate, reading only a frozen reserve, severs the feedback loop and removes the transition at every contamination rate, even adversarially, while controlling false positives label-free. For the complementary static-calibration failure under drift we give CDC, which restores nominal FPR label-free on all tested drift-affected cells. Finally we prove a two-world impossibility theorem. Drift and contamination are indistinguishable without labels, forcing a closed-form power ceiling our procedure approaches. Together these give a complete possibility/impossibility characterization of label-free adaptive OOD detection.
Original Article
View Cached Full Text

Cached at: 07/27/26, 07:41 AM

# A Sharp-Threshold Theory and Certified Label-Free Calibration
Source: [https://arxiv.org/html/2607.21673](https://arxiv.org/html/2607.21673)
## Self\-Poisoning in Adaptive Out\-of\-Distribution Detection: A Sharp\-Threshold Theory and Certified Label\-Free Calibration

\\nameVishnu Bindu Balachandran\\emailvishnubindubalachandran@outlook\.com \\addrIndependent Researcher \\addrwww\.vishnubindubalachandran\.com

###### Abstract

Test\-time adaptive out\-of\-distribution \(OOD\) detectors update a memory bank from the unlabelled stream\. We show this adaptation obeys a provable dynamical law\. Modelling bank impurity as a generalized Pólya urn, we prove almost\-sure convergence to a mean\-field equilibrium whose slope acts as a reproduction number\. Below one, impurity stays benign\. Above one, the bank is fully poisoned and the detector collapses\. The measured admission kernel is affine \(R2≥0\.996R^\{2\}\\geq 0\.996\) with slope just below one in every encoder family \(a protocol signature\), so this detector class is near\-critical by design, and across 96 settings the predicted threshold matches the empirical collapse, where ungated dictionaries lose up to0\.1630\.163AUROC\. We then prove that a certified admission gate, reading only a frozen reserve, severs the feedback loop and removes the transition at every contamination rate, even adversarially, while controlling false positives label\-free\. For the complementary static\-calibration failure under drift we give CDC, which restores nominal FPR label\-free on all tested drift\-affected cells\. Finally we prove a two\-world impossibility theorem\. Drift and contamination are indistinguishable without labels, forcing a closed\-form power ceiling our procedure approaches\. Together these give a complete possibility/impossibility characterization of label\-free adaptive OOD detection\.

Keywords:out\-of\-distribution detection, test\-time adaptation, self\-training, data poisoning, conformal prediction, e\-values, stochastic approximation, Pólya urns

## 1Introduction

A deployed OOD detector faces two requirements that pull in opposite directions\. It must control false positives at a nominal rateα\\alpha\(every flagged input incurs review cost, and a detector whose false\-positive rate \(FPR\) silently doubles is not usable in production\), and it must adapt, because both the in\-distribution \(ID\) data and the outliers it sees at test time differ from anything available at training\. The modern test\-time\-adaptation literature resolves this tension by pseudo\-labelling: memory\-bank detectors append the most ID\-looking stream points to an ID bank\(e\.g\. Zhanget al\.,[2023](https://arxiv.org/html/2607.21673#bib.bib1), AdaODD\), or the most OOD\-looking points to an OOD dictionary\(e\.g\. Yanget al\.,[2025](https://arxiv.org/html/2607.21673#bib.bib2), OODD\), and re\-score against the updated banks\. These methods report substantial gains under favorable stream mixes\. They also, as we show, fail in a specific, predictable, and severe way under realistic ones\.

#### Self\-poisoning is a phase transition, not noise\.

When the stream is temporally clustered \(bursty\) and contamination is low \(the standard regime of production monitoring\), a pure\-ID batch forces an ungated OOD dictionary to admit ID tail points\. The contrast score of nearby ID points then inflates, which makes further ID admissions more likely \(Figure[1](https://arxiv.org/html/2607.21673#S1.F1)a\)\. That this loop can hurt is folklore in the test\-time\-adaptation literature\. What has been missing is a prediction of when\. We prove the loop is a reinforced stochastic process whose impurityρt\\rho\_\{t\}\(fraction of wrongly admitted points\) converges almost surely to the stable equilibria of a scalar mean fieldh​\(ρ\)=q​\(ρ\)−ρh\(\\rho\)=q\(\\rho\)\-\\rho, whereqqis the admission kernel, the probability that the next admitted point is wrong given current impurity \(Section[4](https://arxiv.org/html/2607.21673#S4)\)\. For the affine kernels we measure across all our settings, the slopeb=q′​\(ρ\)b=q^\{\\prime\}\(\\rho\)acts as a reproduction number\. Whenb<1b<1the equilibrium is benign,ρ∗=a/\(1−b\)\\rho^\{\*\}=a/\(1\-b\)\. Whenb≥1b\\geq 1poisoning is complete,ρt→1\\rho\_\{t\}\\to 1\. Saturating kernels produce a fold bifurcation, a discontinuous jump of the equilibrium at a critical contamination rateπc\\pi\_\{c\}, with hysteresis\. \(Our protocol probes the threshold and knee, not the hysteresis, which would requireπ\\pi\-sweeps within a run\.\) Empirically, the predictedπc\\pi\_\{c\}matches the observed knee on all 96 settings we measure, and ungated dictionaries collapse to0\.1630\.163AUROC below their own frozen baseline in the bursty low\-contamination regime, below chance on near\-OOD pairs, while their realized FPR departs arbitrarily from the nominal level\.

![Refer to caption](https://arxiv.org/html/2607.21673v1/x1.png)Figure 1:The paper in one figure\. \(a\) Ungated adaptive detectors compute admission evidence against the very bank that evidence feeds, so the false\-admission kernelq​\(ρ\)q\(\\rho\)rises with impurity\. The aggregate slope sits just below one \(R0≈0\.95R\_\{0\}\\approx 0\.95, a signature of the fixed\-fraction admission protocol, Appendix[D](https://arxiv.org/html/2607.21673#A4)\), the dictionary saturates \(impurity\>0\.9\>0\.9\), and ranking collapses \(Sections[3](https://arxiv.org/html/2607.21673#S3)and[4](https://arxiv.org/html/2607.21673#S4)\)\. \(b\) WARDEN severs the loop\. Admission evidence is conformal against a frozen reserve under the frozen base scorer, e\-BH requirese≥1/δe\\geq 1/\\deltaof every admission, and the dictionary feeds only a second decision channel, never the admission evidence \(that channel does not improve detection on our features, Appendix[D\.1](https://arxiv.org/html/2607.21673#A4.SS1); the severed admission is what matters\)\. The wrong\-admission intensity is therefore bounded independently of the dictionary state, even under adversarially chosen contamination \(Lemma[8](https://arxiv.org/html/2607.21673#Thmtheorem8), Theorem[9](https://arxiv.org/html/2607.21673#Thmtheorem9), Corollary[10](https://arxiv.org/html/2607.21673#Thmtheorem10)\)\. Decisions keep FPR≤α\\leq\\alphaby conformal validity \(Proposition[7](https://arxiv.org/html/2607.21673#Thmtheorem7)\)\. Measured impurity never exceeds0\.1110\.111, and AUROC tracks the frozen baseline by construction\. \(c\) The complementary failure and its limit\. A stale threshold drifts to1\.94×1\.94\\timesnominal FPR, yet benign drift and tail\-like contamination generate identical unlabelled streams, so any label\-free rule keeping FPR≤α\\leq\\alphaunder drift caps its power under contamination atκ​\(π\)\\kappa\(\\pi\)\(Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)\)\. CDC corrects the calibration quantile on the contaminated stream itself, certifies FPR≤α\\leq\\alphaon all tested drift\-affected cells, and retains a median0\.630\.63–0\.670\.67of oracle power, near the ceiling \(Section[5](https://arxiv.org/html/2607.21673#S5)\)\.
#### Certified admission removes the transition\.

The same theory yields the fix and pinpoints why it works\. The collapse loop exists because admission evidence is computed against the very bank it feeds\. WARDEN severs the loop \(Figure[1](https://arxiv.org/html/2607.21673#S1.F1)b\)\. Its admission e\-values are conformal against a frozen reserve, never the dictionary, and it gates them with e\-BH\(Wang and Ramdas,[2022](https://arxiv.org/html/2607.21673#bib.bib15)\)at levelδ\\delta\. Every admission must clear the absolute thresholde≥1/δe\\geq 1/\\delta, so the wrong\-admission intensity is bounded by a constant independent of the dictionary state\. The reinforcement term is structurally absent and the supercritical branch disappears for every contamination rate \(Theorem[9](https://arxiv.org/html/2607.21673#Thmtheorem9)\), while e\-BH keeps each batch’s expected false\-admission proportion atδ\\deltaunder arbitrary dependence\. No admission evidence depends on anything the contamination can influence, so these guarantees hold even when the contaminated points are chosen by an adaptive adversary \(Corollary[10](https://arxiv.org/html/2607.21673#Thmtheorem10)\)\. The resulting detector tracks its frozen baseline AUROC exactly \(by construction, it ranks with the base score, a design invariant rather than a performance claim\), holds mean realized FPR at0\.0560\.056\(per\-family means0\.0470\.047–0\.0610\.061; per\-cell55–95%95\\%range0\.0360\.036–0\.1000\.100, max0\.1210\.121\) across the4,8004\{,\}800standard\-configuration cells of the kill, confirmation, and grid campaigns \(the ablation campaign varies WARDEN’s own knobs and is reported separately\), and its measured dictionary impurity never exceeds0\.1110\.111there, against ungated impurity above0\.90\.9\.

#### Static calibration fails too, and fixing it label\-free has a price\.

The frozen alternative is not safe either\. A detector calibrated on train\-ID data \(all a deployment has\) realizes FPR1\.94×1\.94\\timesnominal across our settings, because of train→\\totest ID drift \(a per\-dimension variance drift in whitened feature space; a mean\-shift correction recovers nothing\)\. We give CDC \(Section[5](https://arxiv.org/html/2607.21673#S5)\)\. It calibrates the operating threshold directly on the contaminated stream at a contamination\-corrected quantile level1−α​\(1−π^up\)1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)plus a DKW slack, whereπ^up\\hat\{\\pi\}\_\{\\mathrm\{up\}\}is a Storey\-type upper confidence estimate computed from conformal p\-values against the stale reserve\. The drift that breaks the stale calibration biases the Storey statistic upward, in exactly the conservative direction, and we give a sufficient condition under which it remains a valid upper bound \(Theorem[12](https://arxiv.org/html/2607.21673#Thmtheorem12)\)\. CDC needs no labels at any point, certifies FPR≤α\\leq\\alphawith high probability \(Theorem[11](https://arxiv.org/html/2607.21673#Thmtheorem11)\), and empirically certifies all1,6951\{,\}695tested drift\-affected cells at mean FPR0\.0420\.042, retaining a median67%67\\%of oracle TPR\.

#### The retained power cannot be pushed to100%100\\%\.

Our impossibility result \(Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)\) constructs two worlds \(benign ID drift versus contamination by outliers distributed exactly like the ID tail\) whose unlabelled stream laws are identical \(Figure[1](https://arxiv.org/html/2607.21673#S1.F1)c\)\. Any label\-free procedure behaves identically in both, so if it maintains FPR≤α\\leq\\alphaunder the drift world its power in the contamination world is capped atκ=α/\(π\+α​\(1−π\)\)<1\\kappa=\\alpha/\(\\pi\+\\alpha\(1\-\\pi\)\)<1\. Label\-free adaptive calibration therefore must pay a power price set by the confusable contamination mass\. CDC’s measured power retention \(0\.670\.67;0\.630\.63held\-out\) sits near this worst\-case ceiling \(0\.690\.69atπ=0\.05\\pi\{=\}0\.05\), which bounds what any certificate\-preserving label\-free method could add on adversarial instances, though on benign instances the ceiling is higher and part of CDC’s gap is finite\-sample slack \(Remark[16](https://arxiv.org/html/2607.21673#Thmtheorem16)states the comparison precisely\)\.

#### Contributions\.

1. 1\.A sharp\-threshold theory of self\-poisoning\(Section[4](https://arxiv.org/html/2607.21673#S4)\): almost\-sure convergence of dictionary impurity to mean\-field equilibria \(Theorem[3](https://arxiv.org/html/2607.21673#Thmtheorem3)\); reproduction number and critical contamination for affine kernels \(Corollary[4](https://arxiv.org/html/2607.21673#Thmtheorem4)\); fold bifurcation and hysteresis for saturating kernels \(Proposition[5](https://arxiv.org/html/2607.21673#Thmtheorem5)\); finite\-window behavior under FIFO eviction \(Proposition[6](https://arxiv.org/html/2607.21673#Thmtheorem6)\); and structural removal of the transition by severed\-feedback certified admission, robust to adaptive adversaries \(Lemma[8](https://arxiv.org/html/2607.21673#Thmtheorem8), Theorem[9](https://arxiv.org/html/2607.21673#Thmtheorem9), Corollary[10](https://arxiv.org/html/2607.21673#Thmtheorem10)\)\.
2. 2\.CDC: certified decontaminated calibration\(Section[5](https://arxiv.org/html/2607.21673#S5)\): the first label\-free procedure that restores nominal FPR under simultaneous ID drift and stream contamination, with finite\-sample validity \(Theorem[11](https://arxiv.org/html/2607.21673#Thmtheorem11)\), a drift\-conservativity guarantee for its contamination estimate \(Theorem[12](https://arxiv.org/html/2607.21673#Thmtheorem12)\), and an explicit power bound \(Theorem[13](https://arxiv.org/html/2607.21673#Thmtheorem13)\)\.
3. 3\.A matching impossibility theorem\(Section[6](https://arxiv.org/html/2607.21673#S6)\): without separation assumptions, no label\-free procedure can simultaneously control FPR under drift and retain power under contamination\. The assumption CDC uses is necessary, and its power loss is unavoidable in order of magnitude\.
4. 4\.The largest cross\-family measurement of adaptive OOD dynamics to date: 96 settings×\\timescontamination×\\timesordering×\\timesseeds \(9,2649\{,\}264streaming cells over the primary campaigns, plus a3,8403\{,\}840\-cell one\-shot held\-out replication that confirms every headline claim\)\.

#### What we do not claim\.

Adaptation does not improve ranking quality on well\-whitened features: growing the ID bank with thousands of true\-ID points moves AUROC by<0\.005<0\.005on our settings, and OOD\-dictionary contrast scores hurt when clean\. The certified dictionary decision channel does not improve detection at a fixed FPR budget either \(Appendix[D\.1](https://arxiv.org/html/2607.21673#A4.SS1)\)\. We therefore make no state\-of\-the\-art AUROC claim anywhere\. The contributions are the dynamical law, the certificates, and the impossibility boundary, the quantities that decide whether an adaptive detector is deployable at all\.

## 2Related work

#### Test\-time adaptive OOD detection and its failures\.

Memory\-bank and dictionary methods \(AdaODD\(Zhanget al\.,[2023](https://arxiv.org/html/2607.21673#bib.bib1)\), OODD\(Yanget al\.,[2025](https://arxiv.org/html/2607.21673#bib.bib2)\), and the broader test\-time\-adaptation line\) adapt from unlabelled streams via pseudo\-labels\. Training\-time approaches instead exploit unlabelled mixtures before deployment and require retraining\(Duet al\.,[2024](https://arxiv.org/html/2607.21673#bib.bib6); Katz\-Samuelset al\.,[2022](https://arxiv.org/html/2607.21673#bib.bib7)\)\. Collapse of self\-training under noisy pseudo\-labels is documented qualitatively \(error accumulation and model collapse in long\-term TTA\(Presset al\.,[2023](https://arxiv.org/html/2607.21673#bib.bib4); Caoet al\.,[2025](https://arxiv.org/html/2607.21673#bib.bib3)\), and adversarial test\-time poisoning\(Suet al\.,[2025](https://arxiv.org/html/2607.21673#bib.bib5)\)\), but, to our knowledge, no prior work identifies the phenomenon as a reinforced stochastic process with a provable critical threshold, nor derives the gate that removes it\. Robust\-TTA heuristics \(sample\-selection, resets, anti\-forgetting\) lack guarantees by construction\. Table[1](https://arxiv.org/html/2607.21673#S2.T1)places the present paper in this landscape\. The test\-time poisoning attacks ofSuet al\.\([2025](https://arxiv.org/html/2607.21673#bib.bib5)\)operate through the mutable state their targets consult at decision time\. Corollary[10](https://arxiv.org/html/2607.21673#Thmtheorem10)shows that certified admission removes this attack surface by construction rather than by hardening\. The defense philosophy, constrain what the stream may admit so that poisoning is bounded, goes back toKloft and Laskov \([2012](https://arxiv.org/html/2607.21673#bib.bib33)\), who bound the displacement an adversary can force on an online centroid anomaly detector under a false\-positive budget\. WARDEN can be read as a conformal completion of that program: its admission budget is enforced by e\-BH against a frozen reserve, which makes the bound label\-free and independent of the bank state and of the contamination process\. At training scale, self\-consuming loops now have a developed theory\. Models degrade when trained on their own outputs\(Shumailovet al\.,[2024](https://arxiv.org/html/2607.21673#bib.bib29)\), iterative retraining is stable when the clean\-data fraction is large enough\(Bertrandet al\.,[2024](https://arxiv.org/html/2607.21673#bib.bib30)\), accumulating rather than replacing data bounds the error\(Gerstgrasseret al\.,[2024](https://arxiv.org/html/2607.21673#bib.bib31)\), and in high\-dimensional regression even a vanishing synthetic fraction can preclude consistency\(Dohmatobet al\.,[2025](https://arxiv.org/html/2607.21673#bib.bib32)\)\. Our subcritical branch does not contradict the last result\. Strong collapse concerns the asymptotic bias of retraining model weights on unboundedly accumulating synthetic data, whereas Corollary[4](https://arxiv.org/html/2607.21673#Thmtheorem4)concerns the stationary composition of a capped, evicting bank whose per\-step influence is damped, and a benign impurity equilibrium is a statement about the bank, not a claim that contamination is harmless downstream\. What this paper adds to that conversation is a deployed\-detector setting where the loop admits an exact scalar mean field with a measurable kernel, a certified gate that provably removes the transition, and a matching impossibility bound\.

Table 1:Guarantee landscape\. A poisoning bound is a bank\-impurity bound that holds for every contamination rate and stream ordering\. FPR under drift means a finite\-sample false\-positive certificate on the drifted stream without labels\. ✓ and×\\timesmark whether a guarantee is provided, a dash marks a guarantee that does not apply because the method keeps no adaptive bank, and a theorem number marks where we prove it\. For FPR under drift we also note each prior method’s binding restriction\.methodadaptslabel\-freepoisoning boundFPR under driftpower boundAdaODD, OODDbank✓×\\times×\\times×\\timesrobust\-TTA heuristicsmodel, bank✓×\\times×\\times×\\timesBashariet al\.\([2025](https://arxiv.org/html/2607.21673#bib.bib14)\)×\\times×\\times–fixed reference✓Wang\([2026](https://arxiv.org/html/2607.21673#bib.bib18)\)×\\times✓–no drift×\\timesQTC\(Yilmaz and Heckel,[2022](https://arxiv.org/html/2607.21673#bib.bib19)\)threshold✓–no contamination×\\timesthis paperbank \+ threshold✓Thm[9](https://arxiv.org/html/2607.21673#Thmtheorem9)Thm[11](https://arxiv.org/html/2607.21673#Thmtheorem11)Thm[13](https://arxiv.org/html/2607.21673#Thmtheorem13)

#### Conformal novelty detection and e\-values\.

Conformal p\-values for outlier detection, FDR\-controlling selection, and e\-value calibration are established\(Bateset al\.,[2023](https://arxiv.org/html/2607.21673#bib.bib10); Wang and Ramdas,[2022](https://arxiv.org/html/2607.21673#bib.bib15); Jin and Candès,[2023](https://arxiv.org/html/2607.21673#bib.bib11)\); adaptive novelty detection with FDR guarantees appears inMarandonet al\.\([2024](https://arxiv.org/html/2607.21673#bib.bib12)\), and conformal selection under hierarchical structure inLee and Ren \([2025](https://arxiv.org/html/2607.21673#bib.bib13)\)\. We use this machinery as given\. Our per\-decision FPR statement \(Proposition[7](https://arxiv.org/html/2607.21673#Thmtheorem7)\) is a direct consequence of exchangeability and is not claimed as a contribution\. What is new is \(i\) the analysis of certified admission inside the poisoning dynamics: per\-batch FDR control alone provably does not prevent long\-run poisoning \(Remark[17](https://arxiv.org/html/2607.21673#Thmtheorem17)\); what does is severing the evidence loop, and we prove exactly that \(Lemma[8](https://arxiv.org/html/2607.21673#Thmtheorem8), Theorem[9](https://arxiv.org/html/2607.21673#Thmtheorem9)\); and \(ii\) the CDC construction, which is a calibration procedure, not a selection procedure\.

#### Contaminated calibration\.

Bashariet al\.\([2025](https://arxiv.org/html/2607.21673#bib.bib14)\)study conformal outlier detection with a contaminated reference set, obtaining conservative validity and using labelled outliers for power\. CDC differs in all three coordinates that matter at deployment: it is label\-free; it treats contamination of the stream used for recalibration under simultaneous ID drift \(their reference set is fixed\); and it comes with an explicit power bound plus a matching impossibility result\. Concurrent work\(Wang,[2026](https://arxiv.org/html/2607.21673#bib.bib18)\)analyzes when trimming a contaminated calibration set preserves coverage \(a retained\-law diagnostic\); CDC deliberately takes the opposite route, no sample selection at all, correcting the quantile level instead, which is what allows a guarantee under drift, where trimming\-style cleaning provably recovers little \(Section[3](https://arxiv.org/html/2607.21673#S3)\)\. Storey\-type null\-proportion estimation\(Storey,[2002](https://arxiv.org/html/2607.21673#bib.bib17)\)and DKW\-based quantile certificates are classical; the observation that calibration drift biases the Storey statistic in the conservative direction \(Theorem[12](https://arxiv.org/html/2607.21673#Thmtheorem12)\) appears to be new\. Recalibrating a conformal cutoff from unlabelled test data under pure distribution shift is studied byYilmaz and Heckel \([2022](https://arxiv.org/html/2607.21673#bib.bib19)\)\. A contaminated stream defeats shift\-only recalibration, because drift and contamination are confounded in unlabelled data, which is exactly the content of Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)\.

#### Impossibility results\.

Fanget al\.\([2022](https://arxiv.org/html/2607.21673#bib.bib28)\)prove that OOD detection is not learnable without restrictions on the out\-distribution\. Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)is complementary\. It concerns label\-free calibration rather than learnability, and its construction follows the least\-favorable\-contamination tradition ofHuber \([1965](https://arxiv.org/html/2607.21673#bib.bib26)\)and the tail\-irreducibility phenomenon of semi\-supervised novelty detection\(Blanchardet al\.,[2010](https://arxiv.org/html/2607.21673#bib.bib27)\)\. The closed\-form ceilingκ\\kappaand its empirical match by a deployed procedure appear to be new\.

#### Urn processes and stochastic approximation\.

Our dynamical analysis uses generalized Pólya urns via the ODE method\(Renlund,[2010](https://arxiv.org/html/2607.21673#bib.bib20); Benaïm,[1999](https://arxiv.org/html/2607.21673#bib.bib21); Pemantle,[2007](https://arxiv.org/html/2607.21673#bib.bib22); Laruelle and Pagès,[2013](https://arxiv.org/html/2607.21673#bib.bib23)\)\. These tools are standard in probability, and phase transitions in reinforced urns are themselves known:Laruelle and Pagès \([2019](https://arxiv.org/html/2607.21673#bib.bib34)\)exhibit the passage from a single attracting equilibrium to a two\-attractor system in nonlinear randomized urns\. The novelty here is not the existence of urn phase transitions but their identification inside a deployed detector class, where the admission kernel is measurable, the measured coefficients place a deployment on the phase diagram, and a certified gate removes the transition\.

#### Relation to our own prior submission\.

A companion manuscript \(under review\) studies the static, batch\-level transductive question: how much OOD information the spectrum of a single unlabelled test batch carries, analyzed with spiked random\-matrix theory\. It involves no adaptation over time, no memory banks, and no calibration\-under\-drift claims; conversely the present paper uses no spectral or random\-matrix machinery\. The two share no theorems and no experimental protocol \(batch scoring there; contaminated streaming with closed\-loop state here\)\. Where that paper uses an AdaODD\-style method as a comparison baseline, here the adaptive detectors are the object of study, their closed\-loop admission dynamics and failure modes\. Either paper stands independently of the other\.

## 3Setting and the self\-poisoning phenomenon

#### Protocol\.

A detector observes batchesX1,X2,…X\_\{1\},X\_\{2\},\\dotsof feature vectors from a stream in which each point is ID with probability1−π1\-\\piand OOD with probabilityπ\\pi, arriving either i\.i\.d\. or in bursts \(contiguous OOD blocks, matching how outliers arrive in production\)\. The detector outputs per\-point scores \(larger = more OOD\) and binary flags targeting FPR≤α\\leq\\alpha; it may adapt internal state but never sees labels\. We evaluate realized FPR, TPR, AUROC, and, for adaptive banks, the impurityρt\\rho\_\{t\}of the bank\. Full experimental detail in Section[7](https://arxiv.org/html/2607.21673#S7)\. In brief, 96 settings over four encoder families \(trained vision, frozen foundation, text, document\),π∈\{0\.01,0\.05,0\.1,0\.5\}\\pi\\in\\\{0\.01,0\.05,0\.1,0\.5\\\}, both orderings, multiple seeds,α=δ=0\.10\\alpha=\\delta=0\.10\.

#### Phenomenon 1: ungated adaptation collapses\.

In bursty streams withπ≤0\.1\\pi\\leq 0\.1, the ungated OOD\-dictionary detector loses0\.1630\.163mean AUROC relative to its own frozen baseline \(up to0\.280\.28on individual families; below chance on near\-OOD pairs\), and its dictionary impurity exceeds0\.90\.9\. The “OOD” dictionary is more than90%90\\%ID\. The ungated ID\-bank detector loses0\.0260\.026AUROC through the mirror\-image mechanism\. On a full replication over all 96 settings using five fresh prime seeds never touched during development, the collapse reproduces \(−0\.114\-0\.114and−0\.028\-0\.028respectively\)\. Figure[2](https://arxiv.org/html/2607.21673#S3.F2)shows the phenomenon\. On identical streams, the ungated dictionary saturates with ID junk within a few batches while the gated detector of Section[4\.3](https://arxiv.org/html/2607.21673#S4.SS3)stays below its budget, and atπ=0\.01\\pi=0\.01bursty every one of the 96 settings is harmed\.

![Refer to caption](https://arxiv.org/html/2607.21673v1/x2.png)Figure 2:Self\-poisoning is universal, not a corner case\. \(a\) Dictionary impurityρt\\rho\_\{t\}on identical bursty streams atπ=0\.01\\pi=0\.01\(median and interquartile range over all 96 settings×\\times5 seeds\): the ungated detector’s “OOD” dictionary is≥90%\\geq 90\\%ID within a few batches, while WARDEN’s certified admission keeps impurity below its budgetδ=0\.1\\delta=0\.1throughout\. \(b\) Resulting AUROC loss of the ungated detector relative to its own frozen baseline, one bar per setting \(sorted\): all 96 settings are harmed atπ=0\.01\\pi=0\.01bursty, by up to0\.410\.41AUROC\.
#### Phenomenon 2: static calibration is stale\.

A detector calibrated once on train\-ID data, the only data a real deployment has, realizes FPR=0\.194=0\.194atα=0\.10\\alpha=0\.10\(1\.94×1\.94\\timesnominal; range1\.31\.3–2\.2×2\.2\\timesacross families\), against0\.1000\.100for an oracle calibrated on held\-out test\-ID\. The drift is a per\-dimension variance change in whitened space\. Re\-centering the stale reserve recovers nothing, an oracle affine \(mean\+variance\) correction recovers∼75%\\sim\\\!75\\%of the gap, and an online\-estimated affine correction only∼38%\\sim\\\!38\\%, motivating a calibration procedure that does not require estimating the drifted density at all \(Section[5](https://arxiv.org/html/2607.21673#S5)\)\.

Phenomena 1 and 2 bracket the design space: adapt ungated and collapse, or freeze and drift out of calibration\. The remainder of the paper shows both failures are provable, both fixes are certifiable, and the residual power cost is information\-theoretically necessary\.

#### Scope\.

Everything that follows is stated at the level of feature vectors and admission dynamics\. Nothing in the theory references pixels, tokens, or page layouts\. Its premises are structural\. A detector self\-updates from its own unlabelled decisions, the stream mixes two populations, and the admission kernel depends on the bank state\. Any deployment with these ingredients \(transaction fraud monitoring, intrusion detection, and clinical anomaly screening are natural instances\) is in scope in principle\. Our evidence covers four encoder families over vision, text, and document data\. Carrying the guarantees to a new domain requires re\-measuring the kernel of Section[4](https://arxiv.org/html/2607.21673#S4)on that domain’s features, not new theory\.

## 4A sharp\-threshold theory of self\-poisoning

### 4\.1The admission process

LetDtD\_\{t\}denote the adaptive bank after batchtt, containingntn\_\{t\}points of whichmtm\_\{t\}were wrongly admitted \(for an OOD dictionary, wrong = ID\); writeρt=mt/nt\\rho\_\{t\}=m\_\{t\}/n\_\{t\}\. At batchttthe detector admitsat≥0a\_\{t\}\\geq 0points,wtw\_\{t\}of them wrong\. All quantities are adapted to the filtrationℱt\\mathcal\{F\}\_\{t\}generated by the stream and the detector’s state\.

###### Assumption 1\(Kernel\)

The admission sizeata\_\{t\}isℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable \(predictable\), and there is a Lipschitzq:\[0,1\]→\[0,1\]q:\[0,1\]\\to\[0,1\]such that the false\-admission proportion obeys𝔼​\[wt/at∣ℱt−1\]=q​\(ρt−1\)\\mathbb\{E\}\\bigl\[w\_\{t\}/a\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\\bigr\]=q\(\\rho\_\{t\-1\}\)on\{at\>0\}\\\{a\_\{t\}\>0\\\}; that is,q​\(ρ\)q\(\\rho\)is the expected fraction of a batch’s admissions that are wrong, given current impurityρ\\rho\.

Predictability ofata\_\{t\}holds by construction for the detectors under study: AdaODD\- and OODD\-style rules admit a fixed fraction of each batch, soat≡⌈qadm​K⌉a\_\{t\}\\equiv\\lceil q\_\{\\mathrm\{adm\}\}K\\rceilis a constant\. \(For rules with random data\-dependentata\_\{t\}the mean field is instead driven by the admission\-weighted kernelq¯​\(ρ\)=𝔼​\[wt∣ℱt−1\]/𝔼​\[at∣ℱt−1\]\\bar\{q\}\(\\rho\)=\\mathbb\{E\}\[w\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]/\\mathbb\{E\}\[a\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]; all results below hold verbatim withqqread asq¯\\bar\{q\}, and our pooled empirical fits estimate exactly this weighted kernel\.\) The single\-argument formq​\(ρ\)q\(\\rho\)is an idealization for bursty streams, where the instantaneous proportion is modulated by the burst phase; thereq​\(ρ\)q\(\\rho\)is the phase\-averaged kernel, the ODE method applies in its Markov\-modulated form\(Benaïm,[1999](https://arxiv.org/html/2607.21673#bib.bib21), Ch\. 8\), and our fits \(Section[7](https://arxiv.org/html/2607.21673#S7)\) estimate precisely this average\. Assumption[1](https://arxiv.org/html/2607.21673#Thmtheorem1)is an empirical claim we test, not a convenience: on all 96 settings the measured kernel is affine inρ\\rhowithR2≥0\.996R^\{2\}\\geq 0\.996\. The mechanism behind theρ\\rho\-dependence is the contrast score: wrong points in the dictionary raise the OOD\-similarity of nearby ID points, so admission mistakes reinforce\.

###### Assumption 2\(Admission rate and noise\)

1≤at≤amax1\\leq a\_\{t\}\\leq a\_\{\\max\}on the eventat\>0a\_\{t\}\>0, admissions occur in a positive fraction of batches,qqis differentiable in a neighborhood of each unstable zero ofh​\(ρ\)=q​\(ρ\)−ρh\(\\rho\)=q\(\\rho\)\-\\rho, andVar​\(wt/at∣ℱt−1\)≥σ02\>0\\mathrm\{Var\}\(w\_\{t\}/a\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\)\\geq\\sigma\_\{0\}^\{2\}\>0on\{at\>0\}\\\{a\_\{t\}\>0\\\}in a neighborhood of every such zero\. \(On\{at=0\}\\\{a\_\{t\}=0\\\}we setξt:=0\\xi\_\{t\}:=0; the recursion does not move\.\)

While the bank is growing \(nt↑n\_\{t\}\\uparrow\), a one\-line computation fromρt=mt−1\+wtnt−1\+at\\rho\_\{t\}=\\tfrac\{m\_\{t\-1\}\+w\_\{t\}\}\{n\_\{t\-1\}\+a\_\{t\}\}\(Appendix[A](https://arxiv.org/html/2607.21673#A1)\) gives the exact recursion

ρt=ρt−1\+γt​\[h​\(ρt−1\)\+ξt\],h​\(ρ\):=q​\(ρ\)−ρ,\\rho\_\{t\}=\\rho\_\{t\-1\}\+\\gamma\_\{t\}\\bigl\[h\(\\rho\_\{t\-1\}\)\+\\xi\_\{t\}\\bigr\],\\qquad h\(\\rho\):=q\(\\rho\)\-\\rho,\(1\)with stepγt=at/nt\\gamma\_\{t\}=a\_\{t\}/n\_\{t\}and noiseξt=wt/at−q​\(ρt−1\)\\xi\_\{t\}=w\_\{t\}/a\_\{t\}\-q\(\\rho\_\{t\-1\}\),\|ξt\|≤1\|\\xi\_\{t\}\|\\leq 1\. Becauseata\_\{t\}is predictable,nt=nt−1\+atn\_\{t\}=n\_\{t\-1\}\+a\_\{t\}and henceγt\\gamma\_\{t\}areℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable, soγt​ξt\\gamma\_\{t\}\\xi\_\{t\}is a genuine martingale\-difference perturbation: a generalized Pólya urn in stochastic\-approximation form\. The state is a proportion, soq,ρ∈\[0,1\]q,\\rho\\in\[0,1\]and the drift is never vacuous\.

### 4\.2Convergence, threshold, and bifurcation

###### Theorem 3\(Almost\-sure convergence to mean\-field equilibria\)

Under Assumptions[1](https://arxiv.org/html/2607.21673#Thmtheorem1)–[2](https://arxiv.org/html/2607.21673#Thmtheorem2), on the event that admissions accumulate \(nt→∞n\_\{t\}\\to\\infty\),ρt\\rho\_\{t\}converges almost surely to the zero set ofhh; if the zeros are isolated,ρt→Z\\rho\_\{t\}\\to ZwhereZZis a single zero, andPr⁡\(Z=ρu\)=0\\Pr\(Z=\\rho^\{u\}\)=0for every linearly unstable zeroρu\\rho^\{u\}\(i\.e\. withq′​\(ρu\)\>1q^\{\\prime\}\(\\rho^\{u\}\)\>1\) at which the noise floor of Assumption[2](https://arxiv.org/html/2607.21673#Thmtheorem2)is active\.

###### Corollary 4\(Reproduction number for affine kernels\)

Letq​\(ρ\)=min⁡\{a\+b​ρ,1\}q\(\\rho\)=\\min\\\{a\+b\\rho,\\,1\\\}witha∈\(0,1\)a\\in\(0,1\),b≥0b\\geq 0, and defineR0:=bR\_\{0\}:=b, and work on the admission\-accumulation event of Theorem[3](https://arxiv.org/html/2607.21673#Thmtheorem3)\. IfR0<1R\_\{0\}<1, thenρt→ρ∗=a/\(1−b\)∧1\\rho\_\{t\}\\to\\rho^\{\*\}=a/\(1\-b\)\\wedge 1a\.s\. partial poisoning with amplification factor\(1−b\)−1\(1\-b\)^\{\-1\}whena<1−ba<1\-b, complete poisoning whena≥1−ba\\geq 1\-b\. IfR0≥1R\_\{0\}\\geq 1, thenρt→1\\rho\_\{t\}\\to 1a\.s\. \(complete poisoning\), for everya\>0a\>0however small\.

Empirically the collapse operates through the first branch at its boundary: every measured setting hasR0∈\[0\.929,0\.959\]<1R\_\{0\}\\in\[0\.929,0\.959\]<1buta≥1−ba\\geq 1\-b, soρ∗=1\\rho^\{\*\}=1, the supercritical branchR0\>1R\_\{0\}\>1is a theoretical completion our settings do not reach\. Figure[3](https://arxiv.org/html/2607.21673#S4.F3)shows the measured law: binned false\-admission proportions hug the affine fits withR0≈0\.94R\_\{0\}\\approx 0\.94–0\.950\.95in every family \(panel a\), the per\-settingR0R\_\{0\}distribution concentrates tightly below the critical value \(panel b\), and the resulting impurity is worst at the lowest contamination rates \(panel c\), the counterintuitive inversion that makes bursty low\-π\\pideployment, the realistic regime, the dangerous one\.

![Refer to caption](https://arxiv.org/html/2607.21673v1/x3.png)Figure 3:The admission kernel is an affine law, measured, not assumed\. \(a\) False\-admission proportion vs\. instantaneous impurityρ\\rho, pooled over4\.3×1064\.3\\times 10^\{6\}admission events and binned \(12 bins, count\-weighted; marker size∝\\proptolog event count\), with weighted affine fits per encoder family: aggregate slopeR0=0\.94R\_\{0\}=0\.94–0\.950\.95everywhere, just below the critical diagonalq=ρq=\\rho\(a protocol signature rather than an encoder constant\)\. The fits are admission\-weighted over the deployment grid;π\\pi\-conditional coefficients are in Appendix[D](https://arxiv.org/html/2607.21673#A4)\. \(b\) Per\-setting fits:R0R\_\{0\}concentrates in\[0\.929,0\.959\]\[0\.929,0\.959\]with medianR2=0\.995R^\{2\}=0\.995across all 96 settings, the affine form of Assumption[1](https://arxiv.org/html/2607.21673#Thmtheorem1)is not a convenience\. \(c\) Median final impurity of the ungated dictionary vs\.π\\pi: becausea​\(π\)a\(\\pi\)grows as streams become more ID\-dominated, impurity is highest at the smallest contamination rates, and bursty ordering saturates the dictionary already atπ=0\.01\\pi=0\.01\.The contamination rate enters through the batch composition:a=a​\(π\)a=a\(\\pi\),b=b​\(π\)b=b\(\\pi\)\. The operational critical rate, the quantity we measure, isπc:=inf\{π:ρ∗​\(π\)≥1/2\}\\pi\_\{c\}:=\\inf\\\{\\pi:\\rho^\{\*\}\(\\pi\)\\geq 1/2\\\}, which under Corollary[4](https://arxiv.org/html/2607.21673#Thmtheorem4)is determined by the measured kernel coefficients\. In bursty streams the pure\-ID batches makea​\(π\)a\(\\pi\)large already at smallπ\\pi, which is why the measured knee sits atπc≈0\.01\\pi\_\{c\}\\approx 0\.01: burstiness, not high contamination, is the dangerous regime\. Empirically the affine form of Assumption[1](https://arxiv.org/html/2607.21673#Thmtheorem1)is not merely convenient\. The measured kernel fits withR2≥0\.996R^\{2\}\\geq 0\.996across all four encoder families with slopeR0≈0\.947R\_\{0\}\\approx 0\.947just below critical \(Section[7](https://arxiv.org/html/2607.21673#S7)\), andρ∗=a/\(1−b\)\\rho^\{\*\}=a/\(1\-b\)sits at the saturation boundary, so ungated adaptation is essentially always in the complete\-poisoning regime once bursts appear\. The concentration of the aggregate slope near one across families is itself informative\. It is a signature of the fixed\-fraction, bank\-relative admission protocol, which pins the composition of admitted points close to the bank’s current composition, not an encoder\-level constant\. Conditioning onπ\\pidecomposes the same law into the pair\(a​\(π\),b​\(π\)\)\(a\(\\pi\),b\(\\pi\)\)that the theory actually consumes, intercept\-driven at low contamination and slope\-driven at high \(Appendix[D](https://arxiv.org/html/2607.21673#A4)\), and the severed design of Section[4\.3](https://arxiv.org/html/2607.21673#S4.SS3)exhibits no measurableρ\\rho\-dependence at all\. The detector class is near\-critical by design, and the contamination rate sets which coordinate tips it\.

###### Proposition 5\(Fold bifurcation and hysteresis for saturating kernels\)

Consider the mean\-field flowρ˙=h​\(ρ;θ\)=q​\(ρ;θ\)−ρ\\dot\{\\rho\}=h\(\\rho;\\theta\)=q\(\\rho;\\theta\)\-\\rhofor a kernel familyq​\(ρ;θ\)q\(\\rho;\\theta\)that is: \(i\)C2C^\{2\}inρ\\rhoandC1C^\{1\}inθ\\theta, with0<q​\(0;θ\)0<q\(0;\\theta\)andq​\(1;θ\)<1q\(1;\\theta\)<1; \(ii\) strictly increasing inρ\\rhowith∂ρq\\partial\_\{\\rho\}qstrictly unimodal \(a single inflection, “sigmoidal”\); \(iii\) strictly increasing inθ\\thetawith∂θq\>0\\partial\_\{\\theta\}q\>0; \(iv\) steep enough somewhere:supρ∂ρq​\(ρ;θ∗\)\>1\\sup\_\{\\rho\}\\partial\_\{\\rho\}q\(\\rho;\\theta^\{\*\}\)\>1at someθ∗\\theta^\{\*\}, with a unique low \(resp\. high\) diagonal crossing at the sweep’s lower \(resp\. upper\) end; and \(v\) non\-degenerate at tangencies \(∂ρ​ρ2q≠0\\partial^\{2\}\_\{\\rho\\rho\}q\\neq 0wheneverq=ρ,∂ρq=1q=\\rho,\\ \\partial\_\{\\rho\}q=1\)\. Then there existθ−<θ\+\\theta\_\{\-\}<\\theta\_\{\+\}such that the equilibrium set passes from a unique low equilibrium \(θ<θ−\\theta<\\theta\_\{\-\}\), through a saddle\-node pair creating bistability \(θ−<θ<θ\+\\theta\_\{\-\}<\\theta<\\theta\_\{\+\}\), to a unique high equilibrium \(θ\>θ\+\\theta\>\\theta\_\{\+\}\); under a quasi\-staticθ\\theta\-sweep the attained equilibrium of the flow jumps discontinuously atθ\+\\theta\_\{\+\}\(upward\) andθ−\\theta\_\{\-\}\(downward\): hysteresis\.

Unimodality of∂ρq\\partial\_\{\\rho\}qbounds the diagonal crossings by three \(the sign pattern ofh′h^\{\\prime\}has at most two changes\), which is what rules regimes beyond the fold sequence out\. The proposition concerns the deterministic mean field; for the stochastic process, Theorem[3](https://arxiv.org/html/2607.21673#Thmtheorem3)pinsρt\\rho\_\{t\}to the corresponding stable branch between fold points on the decreasing\-step time scale\.

###### Proposition 6\(Capped\-bank regime, uniform eviction\)

Once the bank reaches its eviction capNN, suppose admissions are constant \(at≡a¯a\_\{t\}\\equiv\\bar\{a\}, stepγ=a¯/N\\gamma=\\bar\{a\}/N; true for the fixed\-fraction rules studied here\), eviction removes a uniformly random subset ofa¯\\bar\{a\}incumbents per admission, and the affine kernel is unsaturated \(a\+b≤1a\+b\\leq 1; in the saturated caseρt→1\\rho\_\{t\}\\to 1by Corollary[4](https://arxiv.org/html/2607.21673#Thmtheorem4)and there is nothing to prove\)\. Then forb<1b<1:\|𝔼​\[ρt\]−ρ∗\|≤\(1−γ​\(1−b\)\)t−t0​\|ρt0−ρ∗\|\\;\\bigl\|\\mathbb\{E\}\[\\rho\_\{t\}\]\-\\rho^\{\*\}\\bigr\|\\leq\(1\-\\gamma\(1\-b\)\)^\{t\-t\_\{0\}\}\|\\rho\_\{t\_\{0\}\}\-\\rho^\{\*\}\|andlim suptVar​\(ρt\)≤γ​σmax2/\(2​\(1−b\)−γ​\(1−b\)2\)\\limsup\_\{t\}\\mathrm\{Var\}\(\\rho\_\{t\}\)\\leq\\gamma\\,\\sigma\_\{\\max\}^\{2\}/\(2\(1\-b\)\-\\gamma\(1\-b\)^\{2\}\), whereσmax2:=supVar​\(wt/at−ρtev∣ℱt−1\)≤1\\sigma^\{2\}\_\{\\max\}:=\\sup\\mathrm\{Var\}\\bigl\(w\_\{t\}/a\_\{t\}\-\\rho^\{\\mathrm\{ev\}\}\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\\bigr\)\\leq 1: the process forgets its past geometrically and fluctuates withinO​\(γ\)O\(\\sqrt\{\\gamma\}\)of the equilibrium\.

Uniform eviction makes the evicted block’s conditional mean impurity exactlyρt−1\\rho\_\{t\-1\}, which is what the contraction uses\. Literal FIFO \(our implementation\) instead evicts the oldest block, whose impurity is a lagged functional of the window; the recursion acquires a delay term with the same unique fixed point and anO​\(γ\)O\(\\gamma\)bias floor, and the two rules are empirically indistinguishable in all our measurements\. We state the clean result for uniform eviction and treat FIFO as its delay\-perturbed variant\.

### 4\.3Gated admission removes the transition

The detector’s decisions themselves can be certified by conformal validity: this part is classical and stated only for completeness\.

###### Proposition 7\(Per\-decision FPR validity; folklore\)

LetRRbe a reserve of ID points exchangeable with a true\-ID test pointxx, and let the score function used at stepttbeℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable\. Then the one\-sided conformal p\-value ofxxagainstRRis super\-uniform, so flagging atp≤αp\\leq\\alpharealizesPr⁡\(flag​x\)≤α\\Pr\(\\mathrm\{flag\}\\ x\)\\leq\\alphafor everyttand every adaptation rule\.

The substantive question is the admission side\. The collapse mechanism of Theorem[3](https://arxiv.org/html/2607.21673#Thmtheorem3)is theρ\\rho\-dependence of the kernel: admission evidence computed against the mutable bank makes mistakes self\-reinforcing\. WARDEN’s design principle is to sever this loop \(Figure[1](https://arxiv.org/html/2607.21673#S1.F1)b\)\. Admission evidence is computed only against the frozen reserve with the frozen base scorer, the dictionary is never consulted for admission, and the e\-BH gate then adapts how many points that evidence admits\. Two guarantees follow\. First, the classical one: e\-BH at levelδ\\deltacontrols the per\-batch expected false\-admission proportion,𝔼​\[wt/max⁡\(at,1\)∣ℱt−1\]≤δ\\mathbb\{E\}\[w\_\{t\}/\\max\(a\_\{t\},1\)\\mid\\mathcal\{F\}\_\{t\-1\}\]\\leq\\delta, under arbitrary within\-batch dependence\(Wang and Ramdas,[2022](https://arxiv.org/html/2607.21673#bib.bib15)\)\. We note honestly that this per\-batch FDR bound alone does not cap the long\-run impurity ratio: adversarially structured e\-values can satisfy it while admitting only wrong points, rarely\. What removes the bifurcation is the severed feedback, quantified next\.

###### Lemma 8\(ρ\\rho\-independent wrong\-admission intensity\)

Let each batch of sizeKKbe gated by e\-BH at levelδ\\deltaoverei=a​pia−1e\_\{i\}=a\\,p\_\{i\}^\{a\-1\},a∈\(0,1\)a\\in\(0,1\), wherepip\_\{i\}is the conformal p\-value of pointiiagainst a frozen ID reserve of sizemmunder a frozen scorer, and let the stream’s true\-ID points be exchangeable with the reserve\. Every e\-BH admission requiresei≥K/\(δ​at\)≥1/δe\_\{i\}\\geq K/\(\\delta a\_\{t\}\)\\geq 1/\\delta, i\.e\.pi≤κ​\(δ\):=\(a​δ\)1/\(1−a\)p\_\{i\}\\leq\\kappa\(\\delta\):=\(a\\delta\)^\{1/\(1\-a\)\}\. Consequently, with probability≥1−η\\geq 1\-\\etaover the reserve draw, for every batchtt,

𝔼\[wt∣ℱt−1\]≤Kκ¯\(m,κ,η\)=:W¯,κ¯:=max\{u:Pr\[Bin\(m,u\)<⌈κ\(m\+1\)⌉\]≥η\},\\mathbb\{E\}\[w\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]\\;\\leq\\;K\\,\\bar\{\\kappa\}\(m,\\kappa,\\eta\)\\;=:\\;\\bar\{W\},\\qquad\\bar\{\\kappa\}:=\\max\\bigl\\\{u:\\Pr\\bigl\[\\mathrm\{Bin\}\(m,u\)<\\lceil\\kappa\(m\+1\)\\rceil\\bigr\]\\geq\\eta\\bigr\\\},the exact binomial upper confidence limit for the reserve’sκ\\kappa\-tail quantile \(asymptoticallyκ¯=κ\+O​\(κ​log⁡\(1/η\)/m\)\\bar\{\\kappa\}=\\kappa\+O\(\\sqrt\{\\kappa\\log\(1/\\eta\)/m\}\)\), uniformly in the dictionary stateρ\\rho, the contamination rateπ\\pi, the stream ordering, and the within\-batch dependence\.

###### Theorem 9\(Severed feedback: no supercritical branch\)

Under Lemma[8](https://arxiv.org/html/2607.21673#Thmtheorem8), on the reserve event: \(i\) wrong admissions accrue at aρ\\rho\-independent rate,lim suptmt/t≤W¯\\limsup\_\{t\}m\_\{t\}/t\\leq\\bar\{W\}a\.s\.; \(ii\) on streams where total admissions accrue at ratelim inftnt/t≥r\>0\\liminf\_\{t\}n\_\{t\}/t\\geq r\>0\(e\.g\. any stream whose OOD mass the gate detects at positive rate\),lim suptρt≤W¯/r\\limsup\_\{t\}\\rho\_\{t\}\\leq\\bar\{W\}/ra\.s\.; and \(iii\) the gated kernel obeysq¯gated​\(ρ\)≤W¯/\(W¯\+r\)\\bar\{q\}\_\{\\mathrm\{gated\}\}\(\\rho\)\\leq\\bar\{W\}/\(\\bar\{W\}\+r\)for allρ\\rho, it does not increase with impurity, so the reinforcement mechanism behind Corollary[4](https://arxiv.org/html/2607.21673#Thmtheorem4)and Proposition[5](https://arxiv.org/html/2607.21673#Thmtheorem5)is structurally absent and no supercritical branch exists, for everyπ\\piand every ordering\.

###### Corollary 10\(Adversarial contamination\)

Let the reserve be collected before deployment\. Suppose the contaminated points of each batch are chosen by an adversary that observes the full history, the reserve, and the detector’s code, and that picks the number, the placement, and the feature vectors of those points arbitrarily, subject only to the true\-ID points remaining exchangeable with the reserve\. Then Proposition[7](https://arxiv.org/html/2607.21673#Thmtheorem7), Lemma[8](https://arxiv.org/html/2607.21673#Thmtheorem8), and Theorem[9](https://arxiv.org/html/2607.21673#Thmtheorem9)hold verbatim, with the same constants\.

The reason \(Appendix[A](https://arxiv.org/html/2607.21673#A1)\) is that the guarantees rest on two ingredients only, the exchangeability of true\-ID points with the frozen reserve under a predictable scorer, and the absolute admission thresholde≥1/δe\\geq 1/\\deltacomputed on the frozen channel\. Neither ingredient depends on where the adversary puts its points, and wrong admissions are by definition true\-ID admissions, whose evidence law the adversary does not control\. The corollary does not say attacks are harmless\. An adversary can suppress the dictionary channel’s contribution to decisions, a power effect rather than a validity effect, and an empirically negligible one \(Appendix[D\.1](https://arxiv.org/html/2607.21673#A4.SS1)\)\. Appendix[D\.2](https://arxiv.org/html/2607.21673#A4.SS2)stress\-tests the corollary directly: under a gray\-box tail\-mimicry attack of increasing strength, realized FPR and dictionary impurity are unchanged, and the only quantity the adversary moves is detection power, which decays to the confusable floor of Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)\. What it cannot do is start the cascade or break a certificate\.

At our operating point \(a=0\.1a=0\.1,δ=0\.1\\delta=0\.1,K=64K=64,m=1500m=1500,η=0\.05\\eta=0\.05\):κ≈0\.006\\kappa\\approx 0\.006, exact binomial limitκ¯≈0\.0096\\bar\{\\kappa\}\\approx 0\.0096, andW¯≈0\.62\\bar\{W\}\\approx 0\.62wrong admissions per batch, an a\-priori ceiling; the measured impurity is far below it \(max0\.1110\.111over all6,9606\{,\}960cells, against ungated impurity\>0\.9\>0\.9\), and measured per\-batch FDR respectsδ\\delta\. The lemma’s exchangeability premise is met in our protocol because the admission reserve is held\-out eval\-ID; with a stale reserve under train→\\totest drift the intensity inflates by exactly the drift deficiency of Section[5](https://arxiv.org/html/2607.21673#S5), which is where CDC takes over, the two components compose\. Theorem[9](https://arxiv.org/html/2607.21673#Thmtheorem9)is a statement about the closed\-loop dynamics: certified admission does not merely make fewer mistakes, it deletes the mechanism by which mistakes compound, because the admission channel never reads the object it feeds\. Empirically the gated detector’s AUROC never departed from the frozen baseline, across all cells including every collapsing one\.

## 5CDC: certified decontaminated calibration

We now turn to Phenomenon 2\. The deployment has: a stale reserveRtrR\_\{\\mathrm\{tr\}\}of train\-ID scores; an unlabelled stream windowZ=\(s1,…,sn\)Z=\(s\_\{1\},\\dots,s\_\{n\}\)of scores drawn from the mixtureFmix=\(1−π\)​FID\+π​FOODF\_\{\\mathrm\{mix\}\}=\(1\-\\pi\)F\_\{\\mathrm\{ID\}\}\+\\pi F\_\{\\mathrm\{OOD\}\}, whereFIDF\_\{\\mathrm\{ID\}\}is the current \(drifted\) test\-ID score law; and no labels, ever\. The goal is a thresholdτ^\\hat\{\\tau\}withPID​\(s\>τ^\)≤αP\_\{\\mathrm\{ID\}\}\(s\>\\hat\{\\tau\}\)\\leq\\alpha\.

#### Algorithm \(CDC\)\.

With window sizenn, Storey parameterλ\\lambda\(we useλ=1/2\\lambda=1/2; for even reserve sizes this sits between adjacent grid points of Theorem[12](https://arxiv.org/html/2607.21673#Thmtheorem12), shiftingπ^λ\\hat\{\\pi\}\_\{\\lambda\}byO​\(1/m\)O\(1/m\), which the Hoeffding slack absorbs\), confidenceη\\eta:

1. 1\.Compute conformal p\-valuespip\_\{i\}of the window scores against the stale reserveRtrR\_\{\\mathrm\{tr\}\}, and the Storey statisticπ^λ=1−\#​\{pi\>λ\}n​\(1−λ\)\\hat\{\\pi\}\_\{\\lambda\}=1\-\\frac\{\\\#\\\{p\_\{i\}\>\\lambda\\\}\}\{n\(1\-\\lambda\)\}; setπ^up=\[π^λ\+11−λ​log⁡\(2/η\)2​n\]01\\hat\{\\pi\}\_\{\\mathrm\{up\}\}=\\bigl\[\\hat\{\\pi\}\_\{\\lambda\}\+\\tfrac\{1\}\{1\-\\lambda\}\\sqrt\{\\tfrac\{\\log\(2/\\eta\)\}\{2n\}\}\\bigr\]\_\{0\}^\{1\}\.
2. 2\.Outputτ^=F^mix−1​\(1−α​\(1−π^up\)\+εn\)\\hat\{\\tau\}=\\hat\{F\}^\{\-1\}\_\{\\mathrm\{mix\}\}\\bigl\(1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)\+\\varepsilon\_\{n\}\\bigr\), whereF^mix\\hat\{F\}\_\{\\mathrm\{mix\}\}is the empirical CDF of the window andεn=log⁡\(2/η\)/\(2​n\)\+1/n\\varepsilon\_\{n\}=\\sqrt\{\\log\(2/\\eta\)/\(2n\)\}\+1/n\(the1/n1/nabsorbs empirical\-quantile discreteness\)\.
3. 3\.Refresh the window and repeat; the threshold applied to batchttis computed from batches<t<t\(decide\-then\-update\)\.

Note what CDC does not do: it never tries to identify which points are ID, never estimates the drifted density, and never touches the bank machinery of Section[4](https://arxiv.org/html/2607.21673#S4), it corrects the quantile level instead of the sample\. This is why its guarantee survives exactly where sample\-cleaning heuristics \(Section[3](https://arxiv.org/html/2607.21673#S3)\) fail\.

###### Theorem 11\(Validity\)

Letπ<1\\pi<1, let the window be an i\.i\.d\. draw fromFmixF\_\{\\mathrm\{mix\}\}, and letπ^up≥π\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\\geq\\pihold on an event of probability≥1−η1\\geq 1\-\\eta\_\{1\}\. Then with probability at least1−η/2−η11\-\\eta/2\-\\eta\_\{1\},

PID​\(s\>τ^\)≤Pmix​\(s\>τ^\)1−π≤α​\(1−π^up\)1−π≤α\.P\_\{\\mathrm\{ID\}\}\(s\>\\hat\{\\tau\}\)\\;\\leq\\;\\frac\{P\_\{\\mathrm\{mix\}\}\(s\>\\hat\{\\tau\}\)\}\{1\-\\pi\}\\;\\leq\\;\\frac\{\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)\}\{1\-\\pi\}\\;\\leq\\;\\alpha\.No assumption onFOODF\_\{\\mathrm\{OOD\}\}is required for validity\. If the requested level1−α​\(1−π^up\)\+εn1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)\+\\varepsilon\_\{n\}exceeds11\(possible whenπ^up\\hat\{\\pi\}\_\{\\mathrm\{up\}\}is clipped near11\), setτ^=\+∞\\hat\{\\tau\}=\+\\infty\(flag nothing\): trivially valid, with power governed by Theorem[13](https://arxiv.org/html/2607.21673#Thmtheorem13)’s degenerate case\.

###### Theorem 12\(Drift\-conservativity of the Storey estimate\)

Let the p\-values be conformal against the realized stale reserveRtrR\_\{\\mathrm\{tr\}\}of sizemm, with continuous scores andλ\\lambdaon the reserve grid\{k/\(m\+1\)\}\\\{k/\(m\+1\)\\\}\. Conditionally onRtrR\_\{\\mathrm\{tr\}\}, define the \(reserve\-conditional\) drift deficiencyεdr​\(R\):=\(1−λ\)−PID​\(p\>λ∣R\)\\varepsilon\_\{\\mathrm\{dr\}\}\(R\):=\(1\-\\lambda\)\-P\_\{\\mathrm\{ID\}\}\(p\>\\lambda\\mid R\)and leakageκλ​\(R\):=POOD​\(p\>λ∣R\)\\kappa\_\{\\lambda\}\(R\):=P\_\{\\mathrm\{OOD\}\}\(p\>\\lambda\\mid R\)\. \(a\) Conditional conservativity\. If

\(1−π\)​εdr​\(R\)≥π​κλ​\(R\)\(1\-\\pi\)\\,\\varepsilon\_\{\\mathrm\{dr\}\}\(R\)\\;\\geq\\;\\pi\\,\\kappa\_\{\\lambda\}\(R\)then𝔼​\[π^λ∣R\]≥π\\mathbb\{E\}\[\\hat\{\\pi\}\_\{\\lambda\}\\mid R\]\\geq\\pi, and with the Hoeffding slack of Step 1 \(valid conditionally onRR, since the window is i\.i\.d\. givenRR\),Pr⁡\(π^up​<π∣​R\)≤η/2\\Pr\(\\hat\{\\pi\}\_\{\\mathrm\{up\}\}<\\pi\\mid R\)\\leq\\eta/2\. \(A2\) holds wheneverπ≤π†​\(R\):=εdr​\(R\)/\(εdr​\(R\)\+κλ​\(R\)\)\\pi\\leq\\pi\_\{\\dagger\}\(R\):=\\varepsilon\_\{\\mathrm\{dr\}\}\(R\)/\(\\varepsilon\_\{\\mathrm\{dr\}\}\(R\)\+\\kappa\_\{\\lambda\}\(R\)\), with the conventionπ†:=1\\pi\_\{\\dagger\}:=1when both quantities vanish \(far OOD with no drift: \(A2\) is then0≥00\\geq 0, true for allπ\\pi\)\. \(b\) Dominance makes \(A2\) population\-checkable\. If test\-ID scores stochastically dominate train\-ID scores \(the measured drift direction\), then the reserve\-averaged deficiency satisfies𝔼R​\[εdr​\(R\)\]≥0\\mathbb\{E\}\_\{R\}\[\\varepsilon\_\{\\mathrm\{dr\}\}\(R\)\]\\geq 0\(gridλ\\lambdamakes the exchangeable baseline exact\), and by DKW on the reserve, with probability≥1−2​ηR\\geq 1\-2\\eta\_\{R\}over the reserve draw,εdr​\(R\)≥ε¯dr−εm\\varepsilon\_\{\\mathrm\{dr\}\}\(R\)\\geq\\bar\{\\varepsilon\}\_\{\\mathrm\{dr\}\}\-\\varepsilon\_\{m\}andκλ​\(R\)≤κ¯λ\+εm\\kappa\_\{\\lambda\}\(R\)\\leq\\bar\{\\kappa\}\_\{\\lambda\}\+\\varepsilon\_\{m\},εm=log⁡\(2/ηR\)/\(2​m\)\\varepsilon\_\{m\}=\\sqrt\{\\log\(2/\\eta\_\{R\}\)/\(2m\)\}, where bars denote reserve\-averaged quantities\. Hence the population condition\(1−π\)​\(ε¯dr−εm\)≥π​\(κ¯λ\+εm\)\(1\-\\pi\)\(\\bar\{\\varepsilon\}\_\{\\mathrm\{dr\}\}\-\\varepsilon\_\{m\}\)\\geq\\pi\(\\bar\{\\kappa\}\_\{\\lambda\}\+\\varepsilon\_\{m\}\)implies \(A2\) with probability≥1−2​ηR\\geq 1\-2\\eta\_\{R\}, and thenPr⁡\(π^up<π\)≤η/2\+2​ηR\\Pr\(\\hat\{\\pi\}\_\{\\mathrm\{up\}\}<\\pi\)\\leq\\eta/2\+2\\eta\_\{R\}\.

Condition \(A2\) is the honest boundary of the method: the drift\-induced sub\-uniformity must dominate the OOD leakage aboveλ\\lambda, and it holds automatically in the field’s dominant regime \(far OOD:κλ≈0\\kappa\_\{\\lambda\}\\approx 0, \(A2\) free\)\. Because drift and contamination are indistinguishable from unlabelled data \(Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)\), \(A2\) cannot be verified label\-free\. Some separation assumption is logically required of any label\-free method, and \(A2\) is a mild, interpretable instance\. We therefore state the guarantee in two tiers and regard the first as primary\. The unconditional tier assumes nothing about the outliers: setπ^up=πmax\\hat\{\\pi\}\_\{\\mathrm\{up\}\}=\\pi\_\{\\max\}from a governance policy \(“we design for at mostπmax\\pi\_\{\\max\}contamination”\), and Theorem[11](https://arxiv.org/html/2607.21673#Thmtheorem11)holds for allπ≤πmax\\pi\\leq\\pi\_\{\\max\}, trading a known amount of power for an assumption\-free certificate\. The adaptive tier uses the Storey estimate and is certified under \(A2\)\. It is the variant we evaluate, and its empirical record \(certification on every tested drift\-affected cell, including the bursty cells that violate the i\.i\.d\. window assumption, Section[7](https://arxiv.org/html/2607.21673#S7)\) is evidence about \(A2\) in practice, not a substitute for the assumption\.

###### Theorem 13\(Power\)

Letτ∗=FID−1​\(1−α\)\\tau^\{\*\}=F\_\{\\mathrm\{ID\}\}^\{\-1\}\(1\-\\alpha\)be the oracle threshold and supposeFmixF\_\{\\mathrm\{mix\}\}has density bounded below byκ\>0\\kappa\>0on the deterministic interval\[τ∗,Fmix−1​\(1−α​\(1−π^up\)\+2​εn\)\]\[\\tau^\{\*\},F\_\{\\mathrm\{mix\}\}^\{\-1\}\(1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)\+2\\varepsilon\_\{n\}\)\]\(which containsτ^\\hat\{\\tau\}on the DKW event; ifτ^<τ∗\\hat\{\\tau\}<\\tau^\{\*\}the TPR gap is nonpositive and the bound is trivial\)\. Then with probability at least1−η−η11\-\\eta\-\\eta\_\{1\},

TPR​\(τ∗\)−TPR​\(τ^\)≤f¯OODκ​\(α​π^up\+\(1−α\)​π\+2​εn\),\\mathrm\{TPR\}\(\\tau^\{\*\}\)\-\\mathrm\{TPR\}\(\\hat\{\\tau\}\)\\;\\leq\\;\\frac\{\\bar\{f\}\_\{\\mathrm\{OOD\}\}\}\{\\kappa\}\\,\\Bigl\(\\alpha\\,\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\+\(1\-\\alpha\)\\,\\pi\+2\\varepsilon\_\{n\}\\Bigr\),wheref¯OOD\\bar\{f\}\_\{\\mathrm\{OOD\}\}bounds the OOD score density on the same interval\. In the absence of material drift \(ε¯dr→0\\bar\{\\varepsilon\}\_\{\\mathrm\{dr\}\}\\to 0, soπ^up=O​\(π\+n−1/2\)\\hat\{\\pi\}\_\{\\mathrm\{up\}\}=O\(\\pi\+n^\{\-1/2\}\)\) the power loss vanishes asπ→0\\pi\\to 0at rateO​\(π\+n−1/2\)O\(\\pi\+n^\{\-1/2\}\); under drift,π^up\\hat\{\\pi\}\_\{\\mathrm\{up\}\}carries the systematic termε¯dr/\(1−λ\)\\bar\{\\varepsilon\}\_\{\\mathrm\{dr\}\}/\(1\-\\lambda\)and the loss has a non\-vanishing floor≈α​ε¯dr/\(1−λ\)⋅f¯OOD/κ\\approx\\alpha\\,\\bar\{\\varepsilon\}\_\{\\mathrm\{dr\}\}/\(1\-\\lambda\)\\cdot\\bar\{f\}\_\{\\mathrm\{OOD\}\}/\\kappa, the certified price of drift\-robustness, in line with Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)\.

###### Proposition 14\(Block\-dependent windows\)

Suppose the window is a concatenation of independent blocks of lengthslj≤Ll\_\{j\}\\leq L\(each block i\.i\.d\. internally or not, only cross\-block independence is used; the natural model for bursty arrival, where each burst is one block\)\. Letneff:=n2/∑jlj2≥n/Ln\_\{\\mathrm\{eff\}\}:=n^\{2\}/\\sum\_\{j\}l\_\{j\}^\{2\}\\geq n/L\. Then Theorems[11](https://arxiv.org/html/2607.21673#Thmtheorem11)and[12](https://arxiv.org/html/2607.21673#Thmtheorem12)hold with the window deviation terms replaced by

ε\(neff\)=log⁡\(2​\(k\+1\)/η\)2​neff\+1kfor any grid sizek\(takek=⌈neff⌉\),\\varepsilon\(n\_\{\\mathrm\{eff\}\}\)\\;=\\;\\sqrt\{\\frac\{\\log\\bigl\(2\(k\{\+\}1\)/\\eta\\bigr\)\}\{2\\,n\_\{\\mathrm\{eff\}\}\}\}\+\\frac\{1\}\{k\}\\quad\\text\{for any grid size \}k\\text\{ \(take \}k=\\lceil\\sqrt\{n\_\{\\mathrm\{eff\}\}\}\\,\\rceil\),i\.e\. deviations at the effective sample sizeneffn\_\{\\mathrm\{eff\}\}up to a logarithmic factor\. A deployment that inflates its slack toε​\(neff\)\\varepsilon\(n\_\{\\mathrm\{eff\}\}\)retains the certificate underLL\-block dependence\.

The proof \(Appendix[B](https://arxiv.org/html/2607.21673#A2)\) uses a quantile grid with per\-point Hoeffding bounds over the independent blocks \(with the length\-weighted bounded\-difference constantslj/nl\_\{j\}/n\), plus monotonicity, not a naked McDiarmid, whose centering term is not negligible at this scale; exchangeability across blocks would not suffice \(a de Finetti mixture defeats any such concentration\)\. This worst\-case correction is deliberately loose: atn=4000n=4000,L=100L=100it inflates the slack to≈0\.4\\approx 0\.4, far larger than the empirical margin\. Yet the i\.i\.d\.\-calibrated CDC \(no inflation\) already certifies100%100\\%of bursty cells at FPR0\.0410\.041, the realized bursts are far from the adversarial block arrangement the bound guards against, so the correction, while available for a worst\-case guarantee, is not needed in practice\. Closing this gap with a tight dependent\-window certificate \(e\.g\. via block\-bootstrap quantiles\) is the paper’s main theoretical loose end, stated as such\.

#### Measured behavior\.

Across all 96 settings, on the1,6951\{,\}695cells where drift is material \(stale FPR≥1\.2​α\\geq 1\.2\\alpha\) andπ≤0\.1\\pi\\leq 0\.1: CDC realized FPR≤1\.1​α\\leq 1\.1\\alphain100%100\\%of cells \(mean0\.0420\.042, 95th percentile0\.0630\.063\), against0\.1940\.194for the stale detector and0\.1000\.100for the oracle \(Figure[4](https://arxiv.org/html/2607.21673#S5.F4)a\)\. The pre\-registered bar1\.1​α1\.1\\alphabudgets the finite\-sample slackεn\\varepsilon\_\{n\}; in fact every cell landed belowα\\alphaitself \(max0\.0750\.075\)\.

![Refer to caption](https://arxiv.org/html/2607.21673v1/x4.png)Figure 4:CDC certifies every drift\-affected cell and pays a bounded, predicted power price\. \(a\) Per\-cell realized FPR of CDC against the stale calibration on the same cell, for all1,6951\{,\}695drift\-affected cells: the stale detector runs at up to4×4\\timesnominal while every CDC cell sits below the certification bar1\.1​α1\.1\\alpha\(max0\.0750\.075\)\. \(b\) Per\-cell TPR retention vs\. the oracle by contamination rate, with the worst\-case label\-free ceilingκ​\(π\)=α/\(π\+α​\(1−π\)\)\\kappa\(\\pi\)=\\alpha/\(\\pi\+\\alpha\(1\-\\pi\)\)of Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)\(red\)\. Medians \(black bars:0\.810\.81,0\.750\.75,0\.540\.54\) track the ceiling’s shape; individual benign cells may exceedκ\\kappaand even parity with the oracle \(14%14\\%ofπ=0\.01\\pi=0\.01cells, where small oracle TPRs make the ratio noisy; display clipped at1\.261\.26\), the ceiling binds guarantees on adversarial instances, not benign realizations \(Remark[16](https://arxiv.org/html/2607.21673#Thmtheorem16)\)\.\(We report retention, the ratio to oracle TPR, for cross\-setting interpretability; Theorem[13](https://arxiv.org/html/2607.21673#Thmtheorem13)bounds the additive gap, and the measured additive gaps are consistent with it\.\) Median TPR retention against the oracle, pooled over all drift\-affected cells, was0\.6650\.665\(foundation0\.7760\.776, document0\.6840\.684, vision0\.6300\.630, text0\.5930\.593\)\. On the840840bursty cells, violating the exchangeable\-window assumption, certification was likewise100%100\\%\(mean FPR0\.0410\.041\): the Storey stress test passes\. A pre\-slack variant \(no Hoeffding UCB\) certifies100%100\\%at mean FPR0\.0450\.045and retains0\.7070\.707, quantifying the cost of the finite\-sample guarantee\.

## 6The price is necessary: an impossibility theorem

###### Theorem 15\(Two\-world indistinguishability\)

Fixα∈\(0,1\)\\alpha\\in\(0,1\)and anyπ∈\(0,1\)\\pi\\in\(0,1\)\. LetPPbe a continuous train\-ID score law withτB:=P−1​\(1−α\)\\tau\_\{B\}:=P^\{\-1\}\(1\-\\alpha\), and letQ:=ℒ​\(s​∣s\>​τB\)Q:=\\mathcal\{L\}\(s\\mid s\>\\tau\_\{B\}\)underPP, contamination distributed exactly like the ID tail\. SetP′:=\(1−π\)​P\+π​QP^\{\\prime\}:=\(1\-\\pi\)P\+\\pi Q; thenP′P^\{\\prime\}stochastically dominatesPP\(a valid “drift”\), and worldA:=A:=\(ID=P′=P^\{\\prime\}, no contamination\) and worldB:=B:=\(ID=P=P, each point independently contaminated with probabilityπ\\piby a draw fromQQ\) generate identical unlabelled stream laws\. Consequently the thresholdτ^\\hat\{\\tau\}produced by any label\-free procedure𝒜\\mathcal\{A\}\(a possibly randomized map from the stale reserve and the unlabelled window to a threshold; thresholds are the natural class here since detectors flag by score cutoff, and the indistinguishability premise applies verbatim to arbitrary decision rules\) has the same distribution in both worlds, couple the worlds on the shared observable law, and withκ:=απ\+α​\(1−π\)\\kappa:=\\dfrac\{\\alpha\}\{\\pi\+\\alpha\(1\-\\pi\)\},

\{FPRA​\(τ^\)≤α\}⊆\{TPRB​\(τ^\)≤κ\}almost surely\.\\bigl\\\{\\mathrm\{FPR\}\_\{A\}\(\\hat\{\\tau\}\)\\leq\\alpha\\bigr\\\}\\;\\subseteq\\;\\bigl\\\{\\mathrm\{TPR\}\_\{B\}\(\\hat\{\\tau\}\)\\leq\\kappa\\bigr\\\}\\qquad\\text\{almost surely\.\}Hence every procedure withPrA⁡\[FPR≤α\]≥1−β\\Pr\_\{A\}\[\\mathrm\{FPR\}\\leq\\alpha\]\\geq 1\-\\betaobeysPrB⁡\[TPR≤κ\]≥1−β\\Pr\_\{B\}\[\\mathrm\{TPR\}\\leq\\kappa\]\\geq 1\-\\beta: controlling FPR under drift forces power under contamination below the ceilingκ\\kappa, which satisfiesκ<1\\kappa<1for everyπ∈\(0,1\)\\pi\\in\(0,1\)andκ↓α\\kappa\\downarrow\\alphaasπ↑1\\pi\\uparrow 1\. A labelled oracle achievesFPR≤α\\mathrm\{FPR\}\\leq\\alphain both worlds andTPR=1\\mathrm\{TPR\}=1in worldBB\(worldAAhas no outliers to detect\)\.

The construction is maximally natural: from unlabelled data, “the ID tail drifted up” and “outliers arrived that look like the ID tail” are the same observation, so no label\-free rule can separate them\. The two\-world device itself is classical, a least\-favorable pair in the sense ofHuber \([1965](https://arxiv.org/html/2607.21673#bib.bib26)\)and an instance of the irreducibility that limits semi\-supervised novelty detection\(Blanchardet al\.,[2010](https://arxiv.org/html/2607.21673#bib.bib27)\)\. What is specific to our setting is the drift\-versus\-contamination dichotomy and the resulting closed form forκ\\kappa\.

Three consequences: \(i\) the stochastic dominance/leakage condition of Theorem[12](https://arxiv.org/html/2607.21673#Thmtheorem12)is not an artifact, some assumption separating the worlds is logically required by any label\-free method; \(ii\) CDC resolves the dilemma in the only safe direction, in worldAAit certifies FPR, in worldBBit keeps FPR and pays with power on exactly the confounded band, so its measured TPR retention<1<1is the information\-theoretic toll, not an analysis gap; \(iii\) methods that claim both adaptivity and full power under these conditions, without labels or extra assumptions, cannot be correct\.

## 7Experiments

#### Settings\.

96 cached\-feature settings spanning OpenOOD\-trained\(Zhanget al\.,[2024](https://arxiv.org/html/2607.21673#bib.bib9)\)vision \(ResNet\-18, ViT on CIFAR\-10/100 against MNIST/SVHN/DTD/TIN/Places/CIFAR cross\-pairs\), frozen foundation encoders \(DINOv2, CLIP\), text \(RoBERTa SST\-2 against 20News/AGNews/MNLI\), and documents \(LayoutLMv3 Tobacco intra/cross, finance\)\. Feature whitening \(Ledoit–Wolf shrinkage,Ledoit and Wolf,[2004](https://arxiv.org/html/2607.21673#bib.bib24)\) is fit on train\-ID only\. The base score is kNN\-to\-bank\(Sunet al\.,[2022](https://arxiv.org/html/2607.21673#bib.bib8)\)\. Settings were included if the base separability lies in\[0\.62,0\.995\]\[0\.62,0\.995\]AUROC \(both harm and adaptation must have headroom\)\. Grids:π∈\{0\.01,0\.05,0\.1,0\.5\}\\pi\\in\\\{0\.01,0\.05,0\.1,0\.5\\\}, orders\{\\\{i\.i\.d\., bursty\}\\\}, 5 exploration seeds\. All headline numbers were re\-run once, on the full 96\-setting battery, using five fresh prime seeds never used during development \(a strict one\-shot hold\-out\)\. Campaign totals: kill480480plus its1212\-setting confirmation480480, grid3,8403\{,\}840, and ablations2,1602\{,\}160\(together the6,9606\{,\}960streaming\-dynamics cells\); CDC2,3042\{,\}304, bringing all primary campaigns to9,2649\{,\}264cells; the aggregate\-logged baseline\-robustness sweep \(360360runs, Appendix table\); and the one\-shot held\-out replication of the full battery \(3,8403\{,\}840cells\)\. Zero cell failures\.

#### Detectors\.

static \(train\-calibrated\), frozen oracle \(test\-ID\-calibrated\), AdaODD\-style ungated ID bank, OODD\-style ungated dictionary, WARDEN \(gated; ranks by base score, decides by two conformal channels atα/2\\alpha/2, admits by e\-BH atδ\\delta\), CDC \(Section[5](https://arxiv.org/html/2607.21673#S5)\), and the affine\-refresh ablation\. Ranking and decisions are decoupled in WARDEN\. AUROC is computed on the base ranking score, so it equals the frozen baseline by construction; the flags come from two conformal channels atα/2\\alpha/2\(base and dictionary proximity\), each super\-uniform against the frozen reserve, so their union keeps FPR≤α\\leq\\alpha\. We do not claim the dictionary channel improves detection\. A paired ablation on the full grid shows it does not: at a matched FPR budget it adds a median\+0\.000\+0\.000TPR over the base channel alone, and it leaves WARDEN a median0\.110\.11TPR below spending the whole budget on the base channel \(Appendix[D\.1](https://arxiv.org/html/2607.21673#A4.SS1)\)\. This is the expected consequence of our no\-gain finding, and we report it rather than hide it\. WARDEN’s contribution is safety, not detection power: certified admission keeps the dictionary from poisoning the detector, and the conformal decision holds FPR at a conservative realized mean of0\.0570\.057\.

![Refer to caption](https://arxiv.org/html/2607.21673v1/x5.png)Figure 5:Headline outcomes over all bursty cells withπ≤0\.1\\pi\\leq 0\.1\. \(a\) Empirical CDFs of the AUROC loss relative to each detector’s own frozen baseline: OODD\-style dictionaries lose ranking quality catastrophically \(median0\.130\.13, tail beyond0\.90\.9\), AdaODD\-style ID banks lose mildly \(median0\.0020\.002, consistent with the theory’s smalla​\(π\)a\(\\pi\)for the mirror mechanism\), and WARDEN is a point mass at zero by construction\. \(b\) Realized FPR at nominalα=0\.1\\alpha=0\.1\(boxes: IQR, whiskers:55–95%95\\%\): static stale calibration inflates to≈2×\\approx 2\\timesnominal; the ungated detectors’ fixed thresholds drift out of control in both directions, OODD’s collapses toward zero \(it flags almost nothing: a silent failure, since its AUROC is simultaneously destroyed\), AdaODD’s scatters widely around nominal, while WARDEN’s realized FPR is the only one consistent with the nominal level throughout \(mean0\.0600\.060; the conformal guarantee is marginal per decision, and per\-cell realizations fluctuate around it, max0\.1140\.114\)\.
#### Findings \(all confirmed on held\-out seeds; Figure[5](https://arxiv.org/html/2607.21673#S7.F5)\)\.

1. 1\.Collapse: ungated dictionary−0\.163\-0\.163AUROC vs frozen \(bursty,π≤0\.1\\pi\\leq 0\.1\); impurity\>0\.9\>0\.9; realized FPR uncontrolled in both directions\. Held\-out confirmation \(fresh primes\):−0\.114\-0\.114\. The collapse is not a tuned\-baseline artifact\. Sweeping the admission fractionqqover0\.050\.05,0\.10\.1,0\.20\.2,0\.30\.3,0\.50\.5\(bursty,π≤0\.05\\pi\\leq 0\.05, 12 settings×\\times3 seeds\), the dictionary is harmed at everyqqon 12/12 settings, with meanΔ\\DeltaAUROC rising monotonically0\.08→0\.15→0\.29→0\.31→0\.370\.08\\to 0\.15\\to 0\.29\\to 0\.31\\to 0\.37and the harmed\-cell fraction from61%61\\%to99%99\\%\(Appendix table\)\. The ID\-bank variant’s harm is mild at everyqq\(meanΔ\\DeltaAUROC≤0\.02\\leq 0\.02\), consistent with the theory, since its admission kernel has far smallera​\(π\)a\(\\pi\)at low contamination\.
2. 2\.Threshold law: the admission kernel is strikingly affine\. Pooling4\.3×1064\.3\\times 10^\{6\}admission events, the false\-admission proportion fitsq​\(ρ\)=a\+b​ρq\(\\rho\)=a\+b\\rhowithR2=0\.996R^\{2\}=0\.996–0\.9970\.997in every encoder family, with slope just below one \(b=R0=0\.944b=R\_\{0\}=0\.944–0\.9470\.947,a≈0\.05a\\approx 0\.05\), giving predicted equilibriumρ∗=a/\(1−b\)≈1\\rho^\{\*\}=a/\(1\-b\)\\approx 1\(complete poisoning\), consistent with measured impurity\>0\.9\>0\.9\. Per\-setting fits \(not pooled\) give medianR2=0\.995R^\{2\}=0\.995withR0∈\[0\.929,0\.959\]R\_\{0\}\\in\[0\.929,0\.959\]on all9696settings, so the fit quality is not a cross\-setting pooling artifact\. Theπ\\pi\-conditional coefficients, and the reading of the tight aggregate slope as a protocol signature rather than an encoder constant, are in Appendix[D](https://arxiv.org/html/2607.21673#A4)\. The predictedπc\\pi\_\{c\}lands within the pre\-registered half\-decade tolerance of the empirical knee on 96/96 knee\-settings; impurity trajectories are monotone in92%92\\%of poisoned cells\.
3. 3\.Gating: over the4,8004\{,\}800standard\-configuration cells \(kill, confirmation, grid\), WARDEN impurity≤0\.111\\leq 0\.111\(δ=0\.10\\delta=0\.10\) in every cell, and AUROC equal to frozen everywhere \(by construction, hence not a performance claim; the informative metrics are impurity, realized FPR, and TPR\); mean FPR0\.0560\.056\(per\-family means0\.0470\.047–0\.0610\.061; per\-cell max0\.1210\.121\)\. The FPR guarantee is marginal per decision \(Proposition[7](https://arxiv.org/html/2607.21673#Thmtheorem7)\), and the per\-cell maximum is consistent with binomial fluctuation at the per\-cell ID counts\.
4. 4\.Staleness and CDC: stale FPR0\.1940\.194; CDC certifies100%100\\%of the tested drift\-affected cells at mean FPR0\.0420\.042with median67%67\\%oracle\-TPR retention \(near the worst\-case label\-free ceiling of Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)\); bursty stress unchanged\.
5. 5\.Held\-out replication: a single one\-shot re\-run of the full battery \(3,8403\{,\}840cells\) on five fresh prime seeds passes every pre\-registered hard gate: collapse reproduces on the gate population \(bursty,π≤0\.1\\pi\\leq 0\.1held\-out cells:Δ\\DeltaAUROC0\.1140\.114,84%84\\%of those cells harmed, impurity0\.930\.93; the collapse magnitude is seed\-sensitive,0\.110\.11–0\.160\.16, while the harmed fraction, impurity, and every guarantee are stable\), WARDEN holds \(gate population:Δ\\DeltaAUROC0\.0000\.000, mean FPR0\.0590\.059, impurity≤0\.091\\leq 0\.091; over all3,8403\{,\}840held\-out cells, mean FPR0\.0570\.057, max0\.1150\.115, with final impurity above0\.1110\.111only in six cells whose dictionaries hold fewer than ten points\), static inflation1\.94×1\.94\\times, and CDC certifies100%100\\%of tested cells \(FPR0\.0430\.043; retention0\.6320\.632; bursty100%100\\%at0\.0440\.044\)\.
6. 6\.Ablations: conclusions stable acrossδ∈\[0\.05,0\.3\]\\delta\\in\[0\.05,0\.3\],k∈\{5,10,20\}k\\in\\\{5,10,20\\\},α∈\{0\.05,0\.1\}\\alpha\\in\\\{0\.05,0\.1\\\}, calibrator exponent, and screening threshold\.

## 8Limitations and scope

\(i\) Assumption[1](https://arxiv.org/html/2607.21673#Thmtheorem1)posits a one\-dimensional state\. The affine fit is excellent empirically but the reduction is a modelling step, and richer state \(e\.g\. cluster structure of the poison\) is future work\. \(ii\) Theorem[11](https://arxiv.org/html/2607.21673#Thmtheorem11)assumes an exchangeable calibration window\. Proposition[14](https://arxiv.org/html/2607.21673#Thmtheorem14)extends the certificate toLL\-block\-dependent windows at effective sizen/Ln/L, but that worst\-case slack is loose \(the i\.i\.d\.\-calibrated CDC already certifies100%100\\%of bursty cells\)\. A tight dependent\-window certificate \(block\-bootstrap quantiles matched to the burst length\) is the principal open theoretical problem\. \(iii\) CDC’s power loss is real \(median∼33\\sim\\\!33–37%37\\%of oracle TPR atπ≤0\.1\\pi\\leq 0\.1\)\. Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)bounds how much of it is removable in the worst case \(retention aboveκ\\kappais impossible while keeping the certificate\), but on benign instances part of the loss is finite\-sample slack, and sharper estimates of the confusable mass \(e\.g\. via mixture\-proportion irreducibility\) could recover some of it\. We do not claim CDC is instance\-optimal\. Our guarantees are also per\-batch rather than time\-uniform\. An anytime\-valid version of the admission guarantee via stopped e\-BH\(Wanget al\.,[2025](https://arxiv.org/html/2607.21673#bib.bib16)\)is open \(their theory names FDR control under global dependence as unresolved, which is exactly the deployment filtration here\)\. \(iv\) Adaptation buys no AUROC on well\-whitened features\. If ranking quality is the goal, adaptation is the wrong tool on these representations\. \(v\) All experiments use cached features\. End\-to\-end retraining dynamics are out of scope\. \(vi\) Corollary[10](https://arxiv.org/html/2607.21673#Thmtheorem10)is stress\-tested only with the feature\-space tail\-mimicry attack of Appendix[D\.2](https://arxiv.org/html/2607.21673#A4.SS2), which instantiates its threat model but does not optimize the attack; gradient\-based input\-space poisoning in the style ofSuet al\.\([2025](https://arxiv.org/html/2607.21673#bib.bib5)\)against the admission certificate is untested\. \(vii\) The empirical evidence spans vision, text, and document encoders\. Deployments in other domains with the same structural ingredients \(transaction fraud, intrusion detection, clinical monitoring\) are in scope of the theory but untested here, and the kernel coefficientsa​\(π\)a\(\\pi\),b​\(π\)b\(\\pi\)are domain\-specific quantities that must be re\-measured before any threshold prediction is trusted\.

## 9Reproducibility

All campaigns are driven by a resumable, single\-command pipeline \(settings discovery, grids, ablations, held\-out confirmation, analysis, figure and table generation\)\. Per\-cell JSON artifacts and verdict files are emitted for every number in this paper\. The code and artifacts will be released publicly upon acceptance\.

Acknowledgments and Disclosure of Funding

This work was carried out independently and received no specific grant from any funding agency in the public, commercial, or not\-for\-profit sectors\.

## Appendix AProofs for Section[4](https://arxiv.org/html/2607.21673#S4)

### A\.1Derivation of the recursion \([1](https://arxiv.org/html/2607.21673#S4.E1)\)

Withmt=mt−1\+wtm\_\{t\}=m\_\{t\-1\}\+w\_\{t\},nt=nt−1\+atn\_\{t\}=n\_\{t\-1\}\+a\_\{t\}, andρt=mt/nt\\rho\_\{t\}=m\_\{t\}/n\_\{t\},

ρt−ρt−1=mt−1\+wtnt−1\+at−mt−1nt−1=nt−1​wt−mt−1​atnt−1​\(nt−1\+at\)=atnt​\(wtat−ρt−1\),\\rho\_\{t\}\-\\rho\_\{t\-1\}=\\frac\{m\_\{t\-1\}\+w\_\{t\}\}\{n\_\{t\-1\}\+a\_\{t\}\}\-\\frac\{m\_\{t\-1\}\}\{n\_\{t\-1\}\}=\\frac\{n\_\{t\-1\}w\_\{t\}\-m\_\{t\-1\}a\_\{t\}\}\{n\_\{t\-1\}\(n\_\{t\-1\}\+a\_\{t\}\)\}=\\frac\{a\_\{t\}\}\{n\_\{t\}\}\\Bigl\(\\frac\{w\_\{t\}\}\{a\_\{t\}\}\-\\rho\_\{t\-1\}\\Bigr\),where the last step usesmt−1=ρt−1​nt−1m\_\{t\-1\}=\\rho\_\{t\-1\}n\_\{t\-1\}andnt=nt−1\+atn\_\{t\}=n\_\{t\-1\}\+a\_\{t\}\. Settingγt=at/nt\\gamma\_\{t\}=a\_\{t\}/n\_\{t\}andξt=wt/at−q​\(ρt−1\)\\xi\_\{t\}=w\_\{t\}/a\_\{t\}\-q\(\\rho\_\{t\-1\}\)yields \([1](https://arxiv.org/html/2607.21673#S4.E1)\) exactly, withh​\(ρ\)=q​\(ρ\)−ρh\(\\rho\)=q\(\\rho\)\-\\rho\. Under Assumption[1](https://arxiv.org/html/2607.21673#Thmtheorem1)the admission sizeata\_\{t\}\(hencentn\_\{t\}andγt\\gamma\_\{t\}\) isℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable, so𝔼​\[γt​ξt∣ℱt−1\]=γt​𝔼​\[ξt∣ℱt−1\]=0\\mathbb\{E\}\[\\gamma\_\{t\}\\xi\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]=\\gamma\_\{t\}\\,\\mathbb\{E\}\[\\xi\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]=0on\{at\>0\}\\\{a\_\{t\}\>0\\\}: the perturbation is a bounded martingale difference with predictable envelopeγt≤amax/nt−1\\gamma\_\{t\}\\leq a\_\{\\max\}/n\_\{t\-1\}\. \(Wereata\_\{t\}not predictable, the drift would acquire the biasCov​\(γt,wt/at∣ℱt−1\)\\mathrm\{Cov\}\(\\gamma\_\{t\},\\,w\_\{t\}/a\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\)and the mean field would be the admission\-weighted kernelq¯\\bar\{q\}; see the remark after Assumption[1](https://arxiv.org/html/2607.21673#Thmtheorem1)\.\)

### A\.2Proof of Theorem[3](https://arxiv.org/html/2607.21673#Thmtheorem3)

Write the recursion \([1](https://arxiv.org/html/2607.21673#S4.E1)\)\. On\{nt→∞\}\\\{n\_\{t\}\\to\\infty\\\}: sinceat≤amaxa\_\{t\}\\leq a\_\{\\max\}andnt≥n0\+\#​\{admitting batches≤t\}n\_\{t\}\\geq n\_\{0\}\+\\\#\\\{\\text\{admitting batches\}\\leq t\\\}, we haveγt≤amax/nt−1\\gamma\_\{t\}\\leq a\_\{\\max\}/n\_\{t\-1\}with∑tγt=∞\\sum\_\{t\}\\gamma\_\{t\}=\\infty\(each admitting batch contributes at least1/nt1/n\_\{t\}andntn\_\{t\}grows at most linearly\) and∑tγt2<∞\\sum\_\{t\}\\gamma\_\{t\}^\{2\}<\\infty\(γt=O​\(1/t\)\\gamma\_\{t\}=O\(1/t\)along admitting batches\)\. By the derivation above,γt​ξt\\gamma\_\{t\}\\xi\_\{t\}is a bounded martingale\-difference perturbation with predictable step\.hhis Lipschitz on\[0,1\]\[0,1\]withh​\(0\)=q​\(0\)≥0h\(0\)=q\(0\)\\geq 0andh​\(1\)=q​\(1\)−1≤0h\(1\)=q\(1\)\-1\\leq 0, so the flow ofρ˙=h​\(ρ\)\\dot\{\\rho\}=h\(\\rho\)leaves\[0,1\]\[0,1\]invariant and its chain\-recurrent set is the zero set ofhh\(scalar flow\)\. By the ODE method for stochastic approximation\(Benaïm,[1999](https://arxiv.org/html/2607.21673#bib.bib21), Prop\. 4\.1, Cor\. 6\.6\), whose deterministic\-gain statements apply pathwise here because the stepsγt\\gamma\_\{t\}are predictable with∑γt=∞\\sum\\gamma\_\{t\}=\\infty,∑γt2<∞\\sum\\gamma\_\{t\}^\{2\}<\\inftya\.s\. on the admission\-accumulation event, so the interpolated process is an asymptotic pseudotrajectory of the flow,ρt\\rho\_\{t\}converges a\.s\. to a connected internally chain\-transitive set, i\.e\. to the zero set; isolated zeros give convergence to a single zero\. Non\-convergence to a linearly unstable zeroρu\\rho^\{u\}\(whereh′​\(ρu\)=q′​\(ρu\)−1\>0h^\{\\prime\}\(\\rho^\{u\}\)=q^\{\\prime\}\(\\rho^\{u\}\)\-1\>0\) follows fromPemantle \([2007](https://arxiv.org/html/2607.21673#bib.bib22)\)\(Theorem 2\.9, originally Pemantle 1990\) provided the conditional noise variance is bounded below in a neighborhood ofρu\\rho^\{u\}, exactly the scope of Assumption[2](https://arxiv.org/html/2607.21673#Thmtheorem2); boundary zeros where the noise floor is inactive \(e\.g\.ρu=0\\rho^\{u\}=0reached from a clean bank withq​\(0\)=0q\(0\)=0\) are excluded from the claim, as the theorem statement says\.

### A\.3Proof of Corollary[4](https://arxiv.org/html/2607.21673#Thmtheorem4)

Forq=min⁡\{a\+b​ρ,1\}q=\\min\\\{a\+b\\rho,1\\\}: ifb<1b<1,h​\(ρ\)=a−\(1−b\)​ρh\(\\rho\)=a\-\(1\-b\)\\rhoon the unsaturated range, strictly decreasing with unique zeroρ∗=a/\(1−b\)\\rho^\{\*\}=a/\(1\-b\)\(or, ifa/\(1−b\)\>1a/\(1\-b\)\>1,h\>0h\>0up to the saturation region whereh​\(ρ\)=1−ρh\(\\rho\)=1\-\\rho, giving the zero at11\);h′<0h^\{\\prime\}<0so the zero is stable and Theorem[3](https://arxiv.org/html/2607.21673#Thmtheorem3)applies\. Ifb≥1b\\geq 1: on the unsaturated rangeh​\(ρ\)=a\+\(b−1\)​ρ≥a\>0h\(\\rho\)=a\+\(b\-1\)\\rho\\geq a\>0\(atb=1b=1exactly,h≡a\>0h\\equiv a\>0there\); on the saturated rangeh=1−ρ≥0h=1\-\\rho\\geq 0with equality only atρ=1\\rho=1; thush\>0h\>0on\[0,1\)\[0,1\), the flow is strictly increasing, the only chain\-recurrent point is11, andρt→1\\rho\_\{t\}\\to 1a\.s\.

### A\.4Proof of Proposition[5](https://arxiv.org/html/2607.21673#Thmtheorem5)

At most three crossings\.h′​\(ρ\)=∂ρq−1h^\{\\prime\}\(\\rho\)=\\partial\_\{\\rho\}q\-1; since∂ρq\\partial\_\{\\rho\}qis strictly unimodal \(hypothesis \(ii\)\),h′h^\{\\prime\}has at most two sign changes \(−⁣→⁣\+⁣→⁣−\-\\\!\\to\\\!\+\\\!\\to\\\!\-\), sohhhas at most three zeros, and by \(i\) \(h​\(0\)\>0h\(0\)\>0,h​\(1\)<0h\(1\)<0\) the number of zeros is odd: one or three\. Bistable window is a nonempty interval\. By \(iii\),h​\(ρ;θ\)h\(\\rho;\\theta\)is strictly increasing inθ\\thetapointwise\. Defineθ−:=inf\{θ:h​\(⋅;θ\)​has three zeros\}\\theta\_\{\-\}:=\\inf\\\{\\theta:\\ h\(\\cdot;\\theta\)\\text\{ has three zeros\}\\\}andθ\+:=sup\{⋅\}\\theta\_\{\+\}:=\\sup\\\{\\cdot\\\}; the three\-zero set is an interval because increasingθ\\thetaraiseshh, which can destroy the lower zero pair only once \(the middle negative band ofhhshrinks monotonically inθ\\theta\), pair creation atθ−\\theta\_\{\-\}and destruction atθ\+\\theta\_\{\+\}are each one\-shot\. Hypothesis \(iv\) makes the set nonempty \(θ∗\\theta^\{\*\}has a steep crossing region between the endpoint regimes, giving three zeros for someθ\\theta\) and places the unique low \(high\) crossing at the sweep ends, and \(v\) with the implicit function theorem makes each boundary a non\-degenerate saddle\-node withθ−<θ\+\\theta\_\{\-\}<\\theta\_\{\+\}strictly \(a coincident tangency would force∂ρ​ρ2q=0\\partial^\{2\}\_\{\\rho\\rho\}q=0at the inflection, excluded by \(v\)\)\. Hysteresis\. Forθ\\thetaslightly belowθ\+\\theta\_\{\+\}the flow started on the low branch stays on it \(it is separated from the high branch by the middle unstable zero\); atθ\+\\theta\_\{\+\}the low branch is annihilated in the fold and the state jumps to the high branch; sweeping down, the high branch survives untilθ−\\theta\_\{\-\}: the attained equilibrium is direction\-dependent on\(θ−,θ\+\)\(\\theta\_\{\-\},\\theta\_\{\+\}\)\.

### Proof of Proposition[6](https://arxiv.org/html/2607.21673#Thmtheorem6)\(uniform eviction\)

Withnt≡Nn\_\{t\}\\equiv N: admitata\_\{t\}points \(wrong countwtw\_\{t\}\), evict a uniformly randomata\_\{t\}\-subset of theNNincumbents, whose conditional mean impurity is exactlyρt−1\\rho\_\{t\-1\}\. Hence𝔼​\[ρt∣ℱt−1\]=ρt−1\+atN​\(q​\(ρt−1\)−ρt−1\)\\mathbb\{E\}\[\\rho\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]=\\rho\_\{t\-1\}\+\\frac\{a\_\{t\}\}\{N\}\\bigl\(q\(\\rho\_\{t\-1\}\)\-\\rho\_\{t\-1\}\\bigr\)exactly, and for the affine kernel𝔼​\[ρt\]−ρ∗=\(1−γ​\(1−b\)\)​\(𝔼​\[ρt−1\]−ρ∗\)\\mathbb\{E\}\[\\rho\_\{t\}\]\-\\rho^\{\*\}=\(1\-\\gamma\(1\-b\)\)\(\\mathbb\{E\}\[\\rho\_\{t\-1\}\]\-\\rho^\{\*\}\): geometric contraction with no bias term\. For the variance, the conditional mean is affine inρt−1\\rho\_\{t\-1\}alone \(this is where uniform eviction is used a second time\), so by the conditional\-variance decompositionvt≤\(1−γ​\(1−b\)\)2​vt−1\+γ2​σmax2v\_\{t\}\\leq\(1\-\\gamma\(1\-b\)\)^\{2\}v\_\{t\-1\}\+\\gamma^\{2\}\\sigma^\{2\}\_\{\\max\}withσmax2=supVar​\(wt/at−ρtev∣ℱt−1\)≤1\\sigma^\{2\}\_\{\\max\}=\\sup\\mathrm\{Var\}\(w\_\{t\}/a\_\{t\}\-\\rho^\{\\mathrm\{ev\}\}\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\)\\leq 1, which converges toγ2​σmax2/\(1−\(1−γ​\(1−b\)\)2\)≤γ​σmax2/\(2​\(1−b\)−γ​\(1−b\)2\)\\gamma^\{2\}\\sigma^\{2\}\_\{\\max\}/\(1\-\(1\-\\gamma\(1\-b\)\)^\{2\}\)\\leq\\gamma\\,\\sigma^\{2\}\_\{\\max\}/\(2\(1\-b\)\-\\gamma\(1\-b\)^\{2\}\)\. Under literal FIFO,ρev\\rho^\{\\mathrm\{ev\}\}is the lagged block impurity; writing the same recursion yields a delay equation with identical unique fixed point and an additionalO​\(γ\)O\(\\gamma\)\-amplitude term, which we do not pursue\.

### A\.5Proof of Lemma[8](https://arxiv.org/html/2607.21673#Thmtheorem8)and Theorem[9](https://arxiv.org/html/2607.21673#Thmtheorem9)

Threshold structure of e\-BH\. WithKKhypotheses at levelδ\\delta, e\-BH rejects thek∗k^\{\*\}largest e\-values wherek∗=max⁡\{k:e\(k\)≥K/\(δ​k\)\}k^\{\*\}=\\max\\\{k:e\_\{\(k\)\}\\geq K/\(\\delta k\)\\\}\. Every admitted point therefore hasei≥K/\(δ​at\)≥1/δe\_\{i\}\\geq K/\(\\delta a\_\{t\}\)\\geq 1/\\delta\(asat≤Ka\_\{t\}\\leq K\)\. Sincee=a​pa−1e=ap^\{a\-1\}is decreasing inpp, admission requirespi≤κ​\(δ\)=\(a​δ\)1/\(1−a\)p\_\{i\}\\leq\\kappa\(\\delta\)=\(a\\delta\)^\{1/\(1\-a\)\}\. Per\-point wrong\-admission probability\. Condition on the realized reserveRR\(scoresr\(1\)≤⋯≤r\(m\)r\_\{\(1\)\}\\leq\\dots\\leq r\_\{\(m\)\}, frozen scorer\)\. A true\-ID point is admitted only ifpi≤κp\_\{i\}\\leq\\kappa, i\.e\. its score exceeds the top⌈κ​\(m\+1\)⌉−1\\lceil\\kappa\(m\+1\)\\rceil\-1reserve scores; conditionally onRRthis event has probabilityν​\(R\):=PID​\(s\>r\(m−⌈κ​\(m\+1\)⌉\+1\)\)\\nu\(R\):=P\_\{\\mathrm\{ID\}\}\\bigl\(s\>r\_\{\(m\-\\lceil\\kappa\(m\+1\)\\rceil\+1\)\}\\bigr\), a fixed number with𝔼​\[ν​\(R\)\]≤κ\\mathbb\{E\}\[\\nu\(R\)\]\\leq\\kappaby exchangeability\. For the high\-probability version, note that for anyuu,\{ν​\(R\)\>u\}\\\{\\nu\(R\)\>u\\\}is exactly the event that fewer than⌈κ​\(m\+1\)⌉\\lceil\\kappa\(m\+1\)\\rceilof themm\(i\.i\.d\.\) reserve points fall in the top\-uumass ofPIDP\_\{\\mathrm\{ID\}\}, i\.e\.\{Bin​\(m,u\)<⌈κ​\(m\+1\)⌉\}\\\{\\mathrm\{Bin\}\(m,u\)<\\lceil\\kappa\(m\+1\)\\rceil\\\}, whose probability is decreasing inuu; by the definition ofκ¯\\bar\{\\kappa\}as the largestuuat which this probability still reachesη\\eta,Pr⁡\(ν​\(R\)\>κ¯\)≤η\\Pr\\bigl\(\\nu\(R\)\>\\bar\{\\kappa\}\\bigr\)\\leq\\etaexactly, no regime conditions\. \(The familiar Chernoff closed formκ¯≤κ\+c​κ​log⁡\(1/η\)/m\\bar\{\\kappa\}\\leq\\kappa\+\\sqrt\{c\\,\\kappa\\log\(1/\\eta\)/m\}holds withc=4c=4oncem​κ≳log⁡\(1/η\)m\\kappa\\gtrsim\\log\(1/\\eta\); at our operating pointm​κ≈9m\\kappa\\approx 9we use the exact value\.\) Intensity bound\. On the event\{ν​\(R\)≤κ¯\}\\\{\\nu\(R\)\\leq\\bar\{\\kappa\}\\\}\(probability≥1−η\\geq 1\-\\eta, one draw, alltt\), linearity of conditional expectation over the at mostKKtrue\-ID points of batchttgives𝔼​\[wt∣ℱt−1\]≤K​κ¯=W¯\\mathbb\{E\}\[w\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]\\leq K\\bar\{\\kappa\}=\\bar\{W\}, regardless of the dictionary state, the admission p\-values never involve the dictionary, and regardless of within\-batch dependence \(the bound is a union of marginals\)\. This proves the lemma\. Theorem\. \(i\)Mt:=mt−∑s≤t𝔼​\[ws∣ℱs−1\]M\_\{t\}:=m\_\{t\}\-\\sum\_\{s\\leq t\}\\mathbb\{E\}\[w\_\{s\}\\mid\\mathcal\{F\}\_\{s\-1\}\]is a martingale with increments bounded byKK, soMt/t→0M\_\{t\}/t\\to 0a\.s\. \(SLLN for bounded martingale differences\), givinglim supmt/t≤W¯\\limsup m\_\{t\}/t\\leq\\bar\{W\}\. \(ii\) On\{lim infnt/t≥r\}\\\{\\liminf n\_\{t\}/t\\geq r\\\}:lim supρt=lim supmt/nt≤W¯/r\\limsup\\rho\_\{t\}=\\limsup m\_\{t\}/n\_\{t\}\\leq\\bar\{W\}/r\. \(iii\) The gated admission\-weighted kernel obeysq¯gated​\(ρ\)=𝔼​\[wt∣⋅\]/𝔼​\[at∣⋅\]≤W¯/\(W¯\+rood\)\\bar\{q\}\_\{\\mathrm\{gated\}\}\(\\rho\)=\\mathbb\{E\}\[w\_\{t\}\\mid\\cdot\]/\\mathbb\{E\}\[a\_\{t\}\\mid\\cdot\]\\leq\\bar\{W\}/\(\\bar\{W\}\+r\_\{\\mathrm\{ood\}\}\)whereroodr\_\{\\mathrm\{ood\}\}is the true\-OOD admission intensity; the right side does not depend onρ\\rho, sohhgains no reinforcement term and has no supercritical branch\.

### A\.6Proof of Corollary[10](https://arxiv.org/html/2607.21673#Thmtheorem10)

Fix a batch\. The scorer and the reserve areℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable, so for each true\-ID point in the batch the joint law of its score and the reserve scores is exchangeable no matter what else the batch contains\. Its conformal p\-value is therefore super\-uniform conditionally onℱt−1\\mathcal\{F\}\_\{t\-1\}, and the contaminated points enter neither this joint law nor the frozen evidence channel, so Proposition[7](https://arxiv.org/html/2607.21673#Thmtheorem7)is unchanged\. For Lemma[8](https://arxiv.org/html/2607.21673#Thmtheorem8), every e\-BH admission still requirespi≤κ​\(δ\)p\_\{i\}\\leq\\kappa\(\\delta\), a deterministic threshold argument the adversary cannot touch\. The number of wrong admissions in the batch is at most the number of true\-ID points withpi≤κp\_\{i\}\\leq\\kappa, and the conditional expectation of that count is bounded byK​κ¯K\\bar\{\\kappa\}on the reserve event exactly as in the lemma’s proof\. Which of these points e\-BH selects may depend on the adversary, but the bound counts all of them, so it is adversary\-free\. Theorem[9](https://arxiv.org/html/2607.21673#Thmtheorem9)uses only the per\-batch bound and the admission\-accrual raterr, so its proof goes through verbatim\.

## Appendix BProofs for Section[5](https://arxiv.org/html/2607.21673#S5)

### B\.1Proof of Theorem[11](https://arxiv.org/html/2607.21673#Thmtheorem11)

By the one\-sided DKW inequality with Massart’s constant\(Massart,[1990](https://arxiv.org/html/2607.21673#bib.bib25)\), with probability≥1−η/2\\geq 1\-\\eta/2,supt\(F^mix​\(t\)−Fmix​\(t\)\)≤εn−1/n\\sup\_\{t\}\\bigl\(\\hat\{F\}\_\{\\mathrm\{mix\}\}\(t\)\-F\_\{\\mathrm\{mix\}\}\(t\)\\bigr\)\\leq\\varepsilon\_\{n\}\-1/n\(validity needs only this direction; the two\-sided event used for Theorem[13](https://arxiv.org/html/2607.21673#Thmtheorem13)costsη\\eta\)\. On that event, by construction ofτ^\\hat\{\\tau\},Fmix​\(τ^\)≥F^mix​\(τ^\)−εn≥1−α​\(1−π^up\)F\_\{\\mathrm\{mix\}\}\(\\hat\{\\tau\}\)\\geq\\hat\{F\}\_\{\\mathrm\{mix\}\}\(\\hat\{\\tau\}\)\-\\varepsilon\_\{n\}\\geq 1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\), i\.e\.Pmix​\(s\>τ^\)≤α​\(1−π^up\)P\_\{\\mathrm\{mix\}\}\(s\>\\hat\{\\tau\}\)\\leq\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)\. SincePmix=\(1−π\)​PID\+π​POOD≥\(1−π\)​PIDP\_\{\\mathrm\{mix\}\}=\(1\-\\pi\)P\_\{\\mathrm\{ID\}\}\+\\pi P\_\{\\mathrm\{OOD\}\}\\geq\(1\-\\pi\)P\_\{\\mathrm\{ID\}\}pointwise,PID​\(s\>τ^\)≤Pmix​\(s\>τ^\)/\(1−π\)P\_\{\\mathrm\{ID\}\}\(s\>\\hat\{\\tau\}\)\\leq P\_\{\\mathrm\{mix\}\}\(s\>\\hat\{\\tau\}\)/\(1\-\\pi\)\. On the event\{π^up≥π\}\\\{\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\\geq\\pi\\\}\(probability≥1−η1\\geq 1\-\\eta\_\{1\}\),α​\(1−π^up\)/\(1−π\)≤α\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)/\(1\-\\pi\)\\leq\\alpha\. Union bound\. Note the argument uses only nonnegativity of the OOD component, no dominance, and thatτ^\\hat\{\\tau\}is computed on past data, so the evaluated point is independent of it under the i\.i\.d\. window assumption\.

### B\.2Proof of Theorem[12](https://arxiv.org/html/2607.21673#Thmtheorem12)

\(a\) Work conditionally onRtrR\_\{\\mathrm\{tr\}\}throughout; the window is i\.i\.d\. givenRR\. Then𝔼​\[\#​\{pi\>λ\}∣R\]/n=\(1−π\)​PID​\(p\>λ∣R\)\+π​κλ​\(R\)=\(1−π\)​\[\(1−λ\)−εdr​\(R\)\]\+π​κλ​\(R\)\\mathbb\{E\}\[\\\#\\\{p\_\{i\}\>\\lambda\\\}\\mid R\]/n=\(1\-\\pi\)P\_\{\\mathrm\{ID\}\}\(p\>\\lambda\\mid R\)\+\\pi\\kappa\_\{\\lambda\}\(R\)=\(1\-\\pi\)\[\(1\-\\lambda\)\-\\varepsilon\_\{\\mathrm\{dr\}\}\(R\)\]\+\\pi\\,\\kappa\_\{\\lambda\}\(R\), hence

𝔼​\[π^λ∣R\]=π\+\(1−π\)​εdr​\(R\)−π​κλ​\(R\)1−λ≥πunder \(A2\)\.\\mathbb\{E\}\[\\hat\{\\pi\}\_\{\\lambda\}\\mid R\]=\\pi\+\\frac\{\(1\-\\pi\)\\varepsilon\_\{\\mathrm\{dr\}\}\(R\)\-\\pi\\kappa\_\{\\lambda\}\(R\)\}\{1\-\\lambda\}\\;\\geq\\;\\pi\\quad\\text\{under \(A2\)\}\.GivenRR, the indicators𝟏​\{pi\>λ\}\\mathbf\{1\}\\\{p\_\{i\}\>\\lambda\\\}are i\.i\.d\. \(they share only the fixed reserve\), so Hoeffding applies conditionally:π^λ≥𝔼​\[π^λ∣R\]−11−λ​log⁡\(2/η\)/\(2​n\)\\hat\{\\pi\}\_\{\\lambda\}\\geq\\mathbb\{E\}\[\\hat\{\\pi\}\_\{\\lambda\}\\mid R\]\-\\frac\{1\}\{1\-\\lambda\}\\sqrt\{\\log\(2/\\eta\)/\(2n\)\}with probability≥1−η/2\\geq 1\-\\eta/2; the Step\-1 slack givesPr⁡\(π^up​<π∣​R\)≤η/2\\Pr\(\\hat\{\\pi\}\_\{\\mathrm\{up\}\}<\\pi\\mid R\)\\leq\\eta/2\.π†​\(R\)\\pi\_\{\\dagger\}\(R\)is \(A2\) solved forπ\\pi\. \(b\) Withλ=k0/\(m\+1\)\\lambda=k\_\{0\}/\(m\+1\)on the grid and continuous scores, an exchangeable point’s conformal p\-value satisfiesPr⁡\(p≤λ\)=λ\\Pr\(p\\leq\\lambda\)=\\lambdaexactly \(marginally over\(R,s\)\(R,s\)\), so dominance of test\-ID over train\-ID scores, which makes the test\-ID p\-value stochastically smaller than the exchangeable one against the same reserve, pointwise inRR, gives𝔼R​\[PID​\(p\>λ∣R\)\]≤1−λ\\mathbb\{E\}\_\{R\}\[P\_\{\\mathrm\{ID\}\}\(p\>\\lambda\\mid R\)\]\\leq 1\-\\lambda, i\.e\.ε¯dr:=𝔼R​\[εdr​\(R\)\]≥0\\bar\{\\varepsilon\}\_\{\\mathrm\{dr\}\}:=\\mathbb\{E\}\_\{R\}\[\\varepsilon\_\{\\mathrm\{dr\}\}\(R\)\]\\geq 0\. BothPID​\(p\>λ∣R\)P\_\{\\mathrm\{ID\}\}\(p\>\\lambda\\mid R\)andκλ​\(R\)\\kappa\_\{\\lambda\}\(R\)are tail probabilities at the reserve’s⌈\(1−λ\)​\(m\+1\)⌉\\lceil\(1\-\\lambda\)\(m\+1\)\\rceil\-th order statistic, so DKW on the reserve empirical CDF bounds their fluctuations byεm\\varepsilon\_\{m\}each with probability≥1−2​ηR\\geq 1\-2\\eta\_\{R\}; the population margin condition then implies \(A2\), and the total failure probability isη/2\+2​ηR\\eta/2\+2\\eta\_\{R\}by a union bound\.

### B\.3Proof of Theorem[13](https://arxiv.org/html/2607.21673#Thmtheorem13)

On the DKW event, by constructionF^mix​\(τ^\)≥1−α​\(1−π^up\)\+εn−1/n\\hat\{F\}\_\{\\mathrm\{mix\}\}\(\\hat\{\\tau\}\)\\geq 1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)\+\\varepsilon\_\{n\}\-1/n, so the level ofτ^\\hat\{\\tau\}inFmixF\_\{\\mathrm\{mix\}\}lies in\[1−α​\(1−π^up\),1−α​\(1−π^up\)\+2​εn\]\\bigl\[1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\),\\;1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)\+2\\varepsilon\_\{n\}\\bigr\]\(consistent with the validity proof; the1/n1/nis insideεn\\varepsilon\_\{n\}\)\. The level ofτ∗\\tau^\{\*\}inFmixF\_\{\\mathrm\{mix\}\}isFmix​\(τ∗\)=\(1−π\)​\(1−α\)\+π​FOOD​\(τ∗\)≥\(1−π\)​\(1−α\)F\_\{\\mathrm\{mix\}\}\(\\tau^\{\*\}\)=\(1\-\\pi\)\(1\-\\alpha\)\+\\pi F\_\{\\mathrm\{OOD\}\}\(\\tau^\{\*\}\)\\geq\(1\-\\pi\)\(1\-\\alpha\), using onlyFOOD​\(τ∗\)≥0F\_\{\\mathrm\{OOD\}\}\(\\tau^\{\*\}\)\\geq 0, the bound must be from below to upper\-bound the gap\. Hence the level gap satisfies

Fmix​\(τ^\)−Fmix​\(τ∗\)≤\[1−α​\(1−π^up\)\+2​εn\]−\(1−π\)​\(1−α\)=α​π^up\+\(1−α\)​π\+2​εn\.F\_\{\\mathrm\{mix\}\}\(\\hat\{\\tau\}\)\-F\_\{\\mathrm\{mix\}\}\(\\tau^\{\*\}\)\\;\\leq\\;\\bigl\[1\-\\alpha\(1\-\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\)\+2\\varepsilon\_\{n\}\\bigr\]\-\(1\-\\pi\)\(1\-\\alpha\)=\\alpha\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\+\(1\-\\alpha\)\\pi\+2\\varepsilon\_\{n\}\.The density lower boundκ\\kappaon\[τ∗,τ^\]\[\\tau^\{\*\},\\hat\{\\tau\}\]converts the level gap toτ^−τ∗≤\(α​π^up\+\(1−α\)​π\+2​εn\)/κ\\hat\{\\tau\}\-\\tau^\{\*\}\\leq\(\\alpha\\hat\{\\pi\}\_\{\\mathrm\{up\}\}\+\(1\-\\alpha\)\\pi\+2\\varepsilon\_\{n\}\)/\\kappa, and the OOD density upper bound converts the threshold gap to the TPR gap\. The drift floor in the statement follows by substituting the systematic partε¯dr/\(1−λ\)\\bar\{\\varepsilon\}\_\{\\mathrm\{dr\}\}/\(1\-\\lambda\)ofπ^up\\hat\{\\pi\}\_\{\\mathrm\{up\}\}from Theorem[12](https://arxiv.org/html/2607.21673#Thmtheorem12)\(b\)\.

### B\.4Proof of Proposition[14](https://arxiv.org/html/2607.21673#Thmtheorem14)

Write the window as independent blocksZ\(1\),…,Z\(B\)Z^\{\(1\)\},\\dots,Z^\{\(B\)\}of lengthslj≤Ll\_\{j\}\\leq L,∑jlj=n\\sum\_\{j\}l\_\{j\}=n, each with marginal point lawFmixF\_\{\\mathrm\{mix\}\}\. Independence across blocks is essential: exchangeability alone admits de Finetti mixtures for which no such concentration holds\. \(i\) Pointwise deviation\. For fixedtt,F^mix​\(t\)=∑j\(lj/n\)​F^\(j\)​\(t\)\\hat\{F\}\_\{\\mathrm\{mix\}\}\(t\)=\\sum\_\{j\}\(l\_\{j\}/n\)\\,\\hat\{F\}^\{\(j\)\}\(t\)is a weighted average ofBBindependent\[0,1\]\[0,1\]\-valued variables with meanFmix​\(t\)F\_\{\\mathrm\{mix\}\}\(t\)and weightslj/nl\_\{j\}/n; Hoeffding givesPr⁡\[\|F^mix​\(t\)−Fmix​\(t\)\|\>ϵ\]≤2​exp⁡\(−2​ϵ2​n2/∑jlj2\)=2​e−2​neff​ϵ2\\Pr\[\|\\hat\{F\}\_\{\\mathrm\{mix\}\}\(t\)\-F\_\{\\mathrm\{mix\}\}\(t\)\|\>\\epsilon\]\\leq 2\\exp\\bigl\(\-2\\epsilon^\{2\}n^\{2\}/\\textstyle\\sum\_\{j\}l\_\{j\}^\{2\}\\bigr\)=2e^\{\-2n\_\{\\mathrm\{eff\}\}\\epsilon^\{2\}\}\. \(ii\) Uniform deviation via a quantile grid\. Choosekkgrid pointst1<⋯<tkt\_\{1\}<\\dots<t\_\{k\}withFmix​\(ti\)=i/\(k\+1\)F\_\{\\mathrm\{mix\}\}\(t\_\{i\}\)=i/\(k\{\+\}1\)\. A union bound gives allkkpointwise deviations≤ϵ\\leq\\epsilonwith probability≥1−2​k​e−2​neff​ϵ2\\geq 1\-2k\\,e^\{\-2n\_\{\\mathrm\{eff\}\}\\epsilon^\{2\}\}; monotonicity ofF^mix\\hat\{F\}\_\{\\mathrm\{mix\}\}andFmixF\_\{\\mathrm\{mix\}\}then extends the bound between grid points at cost1/\(k\+1\)1/\(k\{\+\}1\):supt\|F^mix−Fmix\|≤ϵ\+1/k\\sup\_\{t\}\|\\hat\{F\}\_\{\\mathrm\{mix\}\}\-F\_\{\\mathrm\{mix\}\}\|\\leq\\epsilon\+1/kwith probability≥1−2​\(k\+1\)​e−2​neff​ϵ2\\geq 1\-2\(k\{\+\}1\)e^\{\-2n\_\{\\mathrm\{eff\}\}\\epsilon^\{2\}\}, which is the statedε​\(neff\)\\varepsilon\(n\_\{\\mathrm\{eff\}\}\)atϵ=log⁡\(2​\(k\+1\)/η\)/\(2​neff\)\\epsilon=\\sqrt\{\\log\(2\(k\{\+\}1\)/\\eta\)/\(2n\_\{\\mathrm\{eff\}\}\)\}\. \(A naked McDiarmid argument would mis\-center: it controls deviation from𝔼​\[sup\]\\mathbb\{E\}\[\\sup\], which is itselfΘ​\(neff−1/2\)\\Theta\(n\_\{\\mathrm\{eff\}\}^\{\-1/2\}\)\.\) \(iii\) Storey statistic\.π^λ\\hat\{\\pi\}\_\{\\lambda\}is a length\-weighted average of block means of𝟏​\{pi\>λ\}\\mathbf\{1\}\\\{p\_\{i\}\>\\lambda\\\}\(conditionally on the reserve, blocks remain independent\); the same weighted Hoeffding applies atneffn\_\{\\mathrm\{eff\}\}\. Substitutingε​\(neff\)\\varepsilon\(n\_\{\\mathrm\{eff\}\}\)for the window terms of Theorems[11](https://arxiv.org/html/2607.21673#Thmtheorem11)–[12](https://arxiv.org/html/2607.21673#Thmtheorem12), the quantile and conservativity arguments are unchanged\.

## Appendix CProof of Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)

Construction and dominance\. LetPPbe continuous withτB=P−1​\(1−α\)\\tau\_\{B\}=P^\{\-1\}\(1\-\\alpha\), soP¯​\(τB\):=P​\(s\>τB\)=α\\bar\{P\}\(\\tau\_\{B\}\):=P\(s\>\\tau\_\{B\}\)=\\alpha\. LetQ=ℒ​\(s​∣s\>​τB\)Q=\\mathcal\{L\}\(s\\mid s\>\\tau\_\{B\}\), i\.e\.Q¯​\(t\)=P¯​\(t\)/α\\bar\{Q\}\(t\)=\\bar\{P\}\(t\)/\\alphafort≥τBt\\geq\\tau\_\{B\}andQ¯​\(t\)=1\\bar\{Q\}\(t\)=1fort<τBt<\\tau\_\{B\}\. SetP′=\(1−π\)​P\+π​QP^\{\\prime\}=\(1\-\\pi\)P\+\\pi Q, with tailP¯′=\(1−π\)​P¯\+π​Q¯\\bar\{P\}^\{\\prime\}=\(1\-\\pi\)\\bar\{P\}\+\\pi\\bar\{Q\}\. Fort≥τBt\\geq\\tau\_\{B\},P¯′​\(t\)=P¯​\(t\)​\[\(1−π\)\+π/α\]≥P¯​\(t\)\\bar\{P\}^\{\\prime\}\(t\)=\\bar\{P\}\(t\)\\bigl\[\(1\-\\pi\)\+\\pi/\\alpha\\bigr\]\\geq\\bar\{P\}\(t\)\(asα≤1\\alpha\\leq 1\); fort<τBt<\\tau\_\{B\},P¯′​\(t\)=\(1−π\)​P¯​\(t\)\+π≥P¯​\(t\)\\bar\{P\}^\{\\prime\}\(t\)=\(1\-\\pi\)\\bar\{P\}\(t\)\+\\pi\\geq\\bar\{P\}\(t\)\. HenceP′⪰PP^\{\\prime\}\\succeq P:P′P^\{\\prime\}is a legitimate upward drift of the ID law\.

Indistinguishability\. WorldAA\(ID=P′=P^\{\\prime\}, no contamination\) has stream tailP¯′\\bar\{P\}^\{\\prime\}\. WorldBB\(ID=P=P, contamination fractionπ\\pifromQQ\) has stream tail\(1−π\)​P¯\+π​Q¯=P¯′\(1\-\\pi\)\\bar\{P\}\+\\pi\\bar\{Q\}=\\bar\{P\}^\{\\prime\}\. The stale reserve is drawn fromPPin both\. Thus the observable law \(window\+\+reserve\) is identical, and any label\-free𝒜\\mathcal\{A\}producesτ^\\hat\{\\tau\}with a common lawμ\\muacross the two worlds\.

The inclusion\. Fix any threshold valueτ\\tau\. Caseτ<τB\\tau<\\tau\_\{B\}: thereQ¯​\(τ\)=1\\bar\{Q\}\(\\tau\)=1andP¯​\(τ\)≥α\\bar\{P\}\(\\tau\)\\geq\\alpha, soFPRA​\(τ\)=P¯′​\(τ\)=\(1−π\)​P¯​\(τ\)\+π≥\(1−π\)​α\+π\>α\\mathrm\{FPR\}\_\{A\}\(\\tau\)=\\bar\{P\}^\{\\prime\}\(\\tau\)=\(1\-\\pi\)\\bar\{P\}\(\\tau\)\+\\pi\\geq\(1\-\\pi\)\\alpha\+\\pi\>\\alphafor everyπ∈\(0,1\)\\pi\\in\(0,1\); thus\{FPRA≤α\}\\\{\\mathrm\{FPR\}\_\{A\}\\leq\\alpha\\\}forcesτ^≥τB\\hat\{\\tau\}\\geq\\tau\_\{B\}\. Caseτ≥τB\\tau\\geq\\tau\_\{B\}:P¯′​\(τ\)=Q¯​\(τ\)​\[π\+α​\(1−π\)\]\\bar\{P\}^\{\\prime\}\(\\tau\)=\\bar\{Q\}\(\\tau\)\\,\[\\pi\+\\alpha\(1\-\\pi\)\]\(usingP¯​\(τ\)=α​Q¯​\(τ\)\\bar\{P\}\(\\tau\)=\\alpha\\bar\{Q\}\(\\tau\)\), soFPRA​\(τ\)≤α⇔Q¯​\(τ\)≤α/\[π\+α​\(1−π\)\]=κ\\mathrm\{FPR\}\_\{A\}\(\\tau\)\\leq\\alpha\\iff\\bar\{Q\}\(\\tau\)\\leq\\alpha/\[\\pi\+\\alpha\(1\-\\pi\)\]=\\kappa\. Since the world\-BBoutliers are exactlyQQ,TPRB​\(τ\)=Q¯​\(τ\)\\mathrm\{TPR\}\_\{B\}\(\\tau\)=\\bar\{Q\}\(\\tau\)\. Hence\{FPRA​\(τ^\)≤α\}⊆\{τ^≥τB\}∩\{Q¯​\(τ^\)≤κ\}=\{TPRB​\(τ^\)≤κ\}\\\{\\mathrm\{FPR\}\_\{A\}\(\\hat\{\\tau\}\)\\leq\\alpha\\\}\\subseteq\\\{\\hat\{\\tau\}\\geq\\tau\_\{B\}\\\}\\cap\\\{\\bar\{Q\}\(\\hat\{\\tau\}\)\\leq\\kappa\\\}=\\\{\\mathrm\{TPR\}\_\{B\}\(\\hat\{\\tau\}\)\\leq\\kappa\\\}, almost surely underμ\\mu\.

Conclusion\. Takingμ\\mu\-probabilities,PrA⁡\[FPR≤α\]≤PrB⁡\[TPR≤κ\]\\Pr\_\{A\}\[\\mathrm\{FPR\}\\leq\\alpha\]\\leq\\Pr\_\{B\}\[\\mathrm\{TPR\}\\leq\\kappa\]; the contrapositive is the stated\(1−β\)\(1\-\\beta\)bound\. The labelled oracle setsτ=P′⁣−1​\(1−α\)\\tau=P^\{\\prime\-1\}\(1\-\\alpha\)in worldAA\(FPR=α=\\alpha\) andτ=τB\\tau=\\tau\_\{B\}in worldBB\(FPR=α=\\alpha, TPR=Q¯​\(τB\)=1=\\bar\{Q\}\(\\tau\_\{B\}\)=1\), achieving both objectives because labels break the tie the unlabelled law cannot\.

## Appendix DAdditional experimental detail

#### Measured admission kernel \(validates Assumption[1](https://arxiv.org/html/2607.21673#Thmtheorem1)\)\.

Pooling all OOD\-dictionary admission events by encoder family and fitting the false\-admission proportionq​\(ρ\)=a\+b​ρq\(\\rho\)=a\+b\\rhoby weighted least squares over1212impurity bins:

The near\-perfect affine fit justifies the one\-dimensional mean field of Section[4](https://arxiv.org/html/2607.21673#S4); the equilibriuma/\(1−b\)≈1a/\(1\-b\)\\approx 1predicts complete poisoning, matching measured impurity\>0\.9\>0\.9\.

Fitting each of the 96 settings individually gives per\-settingR2R^\{2\}with median0\.9950\.995, tenth percentile0\.9920\.992, minimum0\.9160\.916, andR2≥0\.95R^\{2\}\\geq 0\.95on9595of9696settings, with the per\-setting slope in\[0\.929,0\.959\]\[0\.929,0\.959\]on all9696\. The fit quality is therefore not an artifact of cross\-setting pooling\.

#### Where the tight slope comes from\.

The pooled and per\-setting fits average over theπ\\pigrid and the burst phases, so they estimate the admission\-weighted kernelq¯\\bar\{q\}that drives recursion \([1](https://arxiv.org/html/2607.21673#S4.E1)\)\. Conditioning onπ\\pidecomposes the same law into the pair\(a​\(π\),b​\(π\)\)\(a\(\\pi\),b\(\\pi\)\)the theory consumes\. For the OOD dictionary \(same fit protocol\),\(a,b\)=\(0\.94,0\.03\)\(a,b\)=\(0\.94,0\.03\)atπ=0\.01\\pi=0\.01,\(0\.51,0\.42\)\(0\.51,0\.42\)at0\.050\.05,\(0\.17,0\.77\)\(0\.17,0\.77\)at0\.10\.1, and\(0\.02,0\.90\)\(0\.02,0\.90\)at0\.50\.5\. At low contamination the collapse is intercept\-driven, pure\-ID bursts force wrong admissions regardless of the bank state\. At high contamination it is slope\-driven, reinforcement through the contrast score\. The concentration of the aggregate slope near one is thus a signature of the fixed\-fraction, bank\-relative admission protocol, which pins the composition of admitted points close to the bank’s current composition, and not an encoder\-level constant\. Two controls support this reading\. The mirror ID\-bank rule, the same protocol pointed at the opposite tail, has aggregate slope1\.0061\.006withπ\\pi\-conditional slopes0\.040\.04,0\.560\.56,0\.960\.96atπ=0\.05,0\.1,0\.5\\pi=0\.05,0\.1,0\.5\(atπ=0\.01\\pi=0\.01its impurity never leaves a neighborhood of zero, so there is no impurity range to fit\)\. And the severed\-feedback gate of Section[4\.3](https://arxiv.org/html/2607.21673#S4.SS3), whose admission evidence never reads the bank, confines admission\-time impurity to\[0,0\.111\]\[0,0\.111\]over4\.0×1054\.0\\times 10^\{5\}admissions with mean false\-admission proportion0\.0020\.002and correlation0\.010\.01between wrongness and impurity\. With the evidence channel frozen there is no reinforcement coordinate left to measure\. The theorems consume only the measured\(a​\(π\),b​\(π\)\)\(a\(\\pi\),b\(\\pi\)\), no universality is assumed anywhere, and the operationalπc\\pi\_\{c\}is computed from these conditional coefficients\.

#### Baseline\-robustness sweep \(rebuts baseline tuning\)\.

Sweeping the admission fraction of both ungated baselines overq∈\{0\.05,0\.1,0\.2,0\.3,0\.5\}q\\in\\\{0\.05,0\.1,0\.2,0\.3,0\.5\\\}\(12 settings, bursty,π∈\{0\.01,0\.05\}\\pi\\in\\\{0\.01,0\.05\\\}, 3 seeds\):

Dictionary collapse persists at every admission fraction on 12/12 settings and worsens monotonically withqq, as the theory predicts \(largerqq= larger effective stepγ\\gammaand largera​\(π\)a\(\\pi\)\); the harm is a property of ungated admission, not of a particular hyper\-parameter choice\.

### D\.1Does the dictionary decision channel add power? \(No\.\)

WARDEN decides with two conformal channels atα/2\\alpha/2, a base channel and a dictionary\-proximity channel \(C2\)\. We test whether C2 adds detection power with a paired ablation on the full grid \(3,8403\{,\}840cells\): on each stream WARDEN is run against two variants that share its admission exactly \(hence hold the same dictionary at every step\) and differ only in the decision rule, base\-at\-α\\alpha\(the whole budget on the base channel, C2 off\) and base\-at\-α/2\\alpha/2\(the base channel alone\)\.

†fraction of the cells where C2 fired in which its TPR\-per\-FPR exceeds the base ROC slope, i\.e\. C2 detected outliers the base score ranks below the threshold\.

Three facts follow\. \(i\) C2 is inert on well\-whitened features: WARDEN and base\-at\-α/2\\alpha/2have the same mean TPR to three decimals \(0\.6260\.626vs0\.6230\.623\); C2 fires in18%18\\%of cells overall and, when it fires, sits above the base ROC in only23%23\\%of them, for a paired median gain of\+0\.000\+0\.000TPR\. \(ii\) At a matched FPR budget WARDEN detects a median0\.110\.11TPR less than base\-at\-α\\alpha, because halving the base budget costs more than the dictionary channel returns\. \(iii\) The conservatism is a feature, not a bug: WARDEN’s realized FPR \(mean0\.0570\.057, max0\.1210\.121\) stays well below base\-at\-α\\alpha\(mean0\.0990\.099, above the nominal level in44%44\\%of cells, max0\.1580\.158\)\. We therefore make no detection\-gain claim for the dictionary channel\. This is the expected consequence of the no\-gain finding and reinforces the paper’s thesis: an adaptive OOD dictionary does not improve detection on well\-whitened features even when it is safely gated\. A deployment that prioritizes power over conservative FPR can use the base channel at levelα\\alpha\(\+0\.13\+0\.13median TPR at the cost of per\-cell FPR fluctuation aboveα\\alpha\); WARDEN as reported takes the conservative operating point\.

### D\.2Adversarial stress test of the certificates

Corollary[10](https://arxiv.org/html/2607.21673#Thmtheorem10)asserts that no placement of the contaminated points can break the admission certificate or the decision\-level FPR\. We instantiate its threat model with a gray\-box tail\-mimicry attack\. The adversary sees the features, the reserve, and the detector’s code, controls only the contaminated points \(never the ID points or the reserve, matching the corollary’s premise\), and moves every outlier toward its nearest ID\-reserve neighbour in whitened feature space,x′=\(1−ε\)​x\+ε​rnn​\(x\)x^\{\\prime\}=\(1\-\\varepsilon\)\\,x\+\\varepsilon\\,r\_\{\\mathrm\{nn\}\}\(x\), with strengthε∈\{0,0\.25,0\.5,0\.75,0\.9\}\\varepsilon\\in\\\{0,0\.25,0\.5,0\.75,0\.9\\\}\. Asε→1\\varepsilon\\to 1the contamination approaches the confusable ID\-tail law of Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15), the least\-favorable direction\. Twelve settings,π∈\{0\.05,0\.1\}\\pi\\in\\\{0\.05,0\.1\\\}, bursty order, three seeds, 72 cells per strength \(360 total, zero failures\)\.

The certificates hold at every strength, as the corollary requires\. Realized FPR never rises under attack \(it falls, because mimicked outliers cannot produce the extreme conformal evidence admission demands, so the dictionary starves and decisions come from the base channel atα/2\\alpha/2alone\), and the dictionary impurity is exactly zero for everyε≥0\.25\\varepsilon\\geq 0\.25\. What the adversary does achieve is hiding: TPR decays from0\.6150\.615to0\.0040\.004as the outliers become confusable with the ID tail, which is the floor Theorem[15](https://arxiv.org/html/2607.21673#Thmtheorem15)proves no label\-free method can avoid\. The ungated dictionary remains fully poisoned \(impurity≥0\.91\\geq 0\.91\) at every strength\. The attack is feature\-space mimicry with nearest\-neighbour targets; optimizing the attack end to end in input space\(Suet al\.,[2025](https://arxiv.org/html/2607.21673#bib.bib5)\)is future work \(Section[8](https://arxiv.org/html/2607.21673#S8)\)\.

#### Released with the code\.

The full per\-setting FPR, AUROC, and TPR tables \(results/\{grid,cdc,confirm\_jmlr\}/analysis\.json\), the per\-setting predicted\-versus\-empiricalπc\\pi\_\{c\}, theδ,k,α,ea\\delta,k,\\alpha,e\_\{a\}ablation grids, and the pre\-slack CDC row are released with the code, so that every number in this paper is reproducible from a single command\.

## References

- Robust conformal outlier detection under contaminated reference data\.InInternational Conference on Machine Learning \(ICML\),pp\. 3091–3141\.Note:PMLR 267; arXiv:2502\.04807Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2607.21673#S2.T1.10.8.8.3)\.
- S\. Bates, E\. Candès, L\. Lei, Y\. Romano, and M\. Sesia \(2023\)Testing for outliers with conformal p\-values\.The Annals of Statistics51\(1\),pp\. 149–178\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Benaïm \(1999\)Dynamics of stochastic approximation algorithms\.InSéminaire de Probabilités XXXIII,Lecture Notes in Mathematics, Vol\.1709,pp\. 1–68\.Cited by:[§A\.2](https://arxiv.org/html/2607.21673#A1.SS2.p1.26),[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px5.p1.1),[§4\.1](https://arxiv.org/html/2607.21673#S4.SS1.p2.11)\.
- Q\. Bertrand, A\. J\. Bose, A\. Duplessis, M\. Jiralerspong, and G\. Gidel \(2024\)On the stability of iterative retraining of generative models on their own data\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2310\.00429Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- G\. Blanchard, G\. Lee, and C\. Scott \(2010\)Semi\-supervised novelty detection\.Journal of Machine Learning Research11\(99\),pp\. 2973–3009\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px4.p1.1),[§6](https://arxiv.org/html/2607.21673#S6.p1.1)\.
- C\. Cao, Z\. Zhong, Z\. Zhou, T\. Liu, Y\. Liu, K\. Zhang, and B\. Han \(2025\)Noisy test\-time adaptation in vision\-language models\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2502\.14604Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Dohmatob, Y\. Feng, A\. Subramonian, and J\. Kempe \(2025\)Strong model collapse\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2410\.04840Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- X\. Du, Z\. Fang, I\. Diakonikolas, and Y\. Li \(2024\)How does unlabeled data provably help out\-of\-distribution detection?\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2402\.03502Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Fang, Y\. Li, J\. Lu, J\. Dong, B\. Han, and F\. Liu \(2022\)Is out\-of\-distribution detection learnable?\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2210\.14707Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px4.p1.1)\.
- M\. Gerstgrasser, R\. Schaeffer, A\. Dey, R\. Rafailov, H\. Sleight, J\. Hughes, T\. Korbak, R\. Agrawal, D\. Pai, A\. Gromov, D\. A\. Roberts, D\. Yang, D\. L\. Donoho, and S\. Koyejo \(2024\)Is model collapse inevitable? breaking the curse of recursion by accumulating real and synthetic data\.Note:arXiv:2404\.01413Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- P\. J\. Huber \(1965\)A robust version of the probability ratio test\.The Annals of Mathematical Statistics36\(6\),pp\. 1753–1758\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px4.p1.1),[§6](https://arxiv.org/html/2607.21673#S6.p1.1)\.
- Y\. Jin and E\. J\. Candès \(2023\)Selection by prediction with conformal p\-values\.Journal of Machine Learning Research24\(244\),pp\. 1–41\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Katz\-Samuels, J\. B\. Nakhleh, R\. Nowak, and Y\. Li \(2022\)Training OOD detectors in their natural habitats\.InInternational Conference on Machine Learning \(ICML\),pp\. 10848–10865\.Note:PMLR 162; arXiv:2202\.03299Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Kloft and P\. Laskov \(2012\)Security analysis of online centroid anomaly detection\.Journal of Machine Learning Research13,pp\. 3681–3724\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Laruelle and G\. Pagès \(2013\)Randomized urn models revisited using stochastic approximation\.The Annals of Applied Probability23\(4\),pp\. 1409–1436\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px5.p1.1)\.
- S\. Laruelle and G\. Pagès \(2019\)Nonlinear randomized urn models: a stochastic approximation viewpoint\.Electronic Journal of Probability24,pp\. 1–47\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px5.p1.1)\.
- O\. Ledoit and M\. Wolf \(2004\)A well\-conditioned estimator for large\-dimensional covariance matrices\.Journal of Multivariate Analysis88\(2\),pp\. 365–411\.Cited by:[§7](https://arxiv.org/html/2607.21673#S7.SS0.SSS0.Px1.p1.14)\.
- Y\. Lee and Z\. Ren \(2025\)Selection from hierarchical data with conformal e\-values\.Note:arXiv:2501\.02514Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Marandon, L\. Lei, D\. Mary, and E\. Roquain \(2024\)Adaptive novelty detection with false discovery rate guarantee\.The Annals of Statistics52\(1\),pp\. 157–183\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px2.p1.1)\.
- P\. Massart \(1990\)The tight constant in the Dvoretzky–Kiefer–Wolfowitz inequality\.The Annals of Probability18\(3\),pp\. 1269–1283\.Cited by:[§B\.1](https://arxiv.org/html/2607.21673#A2.SS1.p1.12)\.
- R\. Pemantle \(2007\)A survey of random processes with reinforcement\.Probability Surveys4,pp\. 1–79\.Cited by:[§A\.2](https://arxiv.org/html/2607.21673#A1.SS2.p1.26),[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px5.p1.1)\.
- O\. Press, S\. Schneider, M\. Kümmerer, and M\. Bethge \(2023\)RDumb: a simple approach that questions our progress in continual test\-time adaptation\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2306\.05401Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Renlund \(2010\)Generalized Pólya urns via stochastic approximation\.Note:arXiv:1002\.3716Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px5.p1.1)\.
- I\. Shumailov, Z\. Shumaylov, Y\. Zhao, N\. Papernot, R\. Anderson, and Y\. Gal \(2024\)AI models collapse when trained on recursively generated data\.Nature631,pp\. 755–759\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- J\. D\. Storey \(2002\)A direct approach to false discovery rates\.Journal of the Royal Statistical Society, Series B64\(3\),pp\. 479–498\.Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px3.p1.1)\.
- Y\. Su, Y\. Li, N\. Liu, K\. Jia, X\. Yang, C\.\-S\. Foo, and X\. Xu \(2025\)On the adversarial risk of test\-time adaptation: an investigation into realistic test\-time data poisoning\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2410\.04682Cited by:[§D\.2](https://arxiv.org/html/2607.21673#A4.SS2.p2.5),[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1),[§8](https://arxiv.org/html/2607.21673#S8.p1.9)\.
- Y\. Sun, Y\. Ming, X\. Zhu, and Y\. Li \(2022\)Out\-of\-distribution detection with deep nearest neighbors\.InInternational Conference on Machine Learning \(ICML\),pp\. 20827–20840\.Note:PMLR 162; arXiv:2204\.06507Cited by:[§7](https://arxiv.org/html/2607.21673#S7.SS0.SSS0.Px1.p1.14)\.
- C\. Wang \(2026\)When does trimming help conformal prediction? a retained\-law diagnostic under calibration contamination\.Note:arXiv:2605\.06204Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2607.21673#S2.T1.12.10.10.3)\.
- H\. Wang, S\. Dandapanthula, and A\. Ramdas \(2025\)Anytime\-valid FDR control with the stopped e\-BH procedure\.Note:arXiv:2502\.08539; to appear, Statistics & Probability LettersCited by:[§8](https://arxiv.org/html/2607.21673#S8.p1.9)\.
- R\. Wang and A\. Ramdas \(2022\)False discovery rate control with e\-values\.Journal of the Royal Statistical Society, Series B84\(3\),pp\. 822–852\.Cited by:[§1](https://arxiv.org/html/2607.21673#S1.SS0.SSS0.Px2.p1.14),[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2607.21673#S4.SS3.p2.3)\.
- Y\. Yang, L\. Zhu, Z\. Sun, H\. Liu, Q\. Gu, and N\. Ye \(2025\)OODD: test\-time out\-of\-distribution detection with dynamic dictionary\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Note:arXiv:2503\.10468Cited by:[§1](https://arxiv.org/html/2607.21673#S1.p1.1),[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.
- F\. F\. Yilmaz and R\. Heckel \(2022\)Test\-time recalibration of conformal predictors under distribution shift based on unlabeled examples\.Note:arXiv:2210\.04166Cited by:[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2607.21673#S2.T1.13.11.11.2)\.
- J\. Zhang, J\. Yang, P\. Wang, H\. Wang, Y\. Lin, H\. Zhang, Y\. Sun, X\. Du, Y\. Li, Z\. Liu, Y\. Chen, and H\. Li \(2024\)OpenOOD v1\.5: enhanced benchmark for out\-of\-distribution detection\.Journal of Data\-centric Machine Learning Research \(DMLR\)\.Note:arXiv:2306\.09301Cited by:[§7](https://arxiv.org/html/2607.21673#S7.SS0.SSS0.Px1.p1.14)\.
- Y\. Zhang, X\. Wang, T\. Zhou, K\. Yuan, Z\. Zhang, L\. Wang, R\. Jin, and T\. Tan \(2023\)Model\-free test\-time adaptation for out\-of\-distribution detection\.Note:arXiv:2311\.16420Cited by:[§1](https://arxiv.org/html/2607.21673#S1.p1.1),[§2](https://arxiv.org/html/2607.21673#S2.SS0.SSS0.Px1.p1.1)\.

Similar Articles

Fair and Calibrated Toxicity Detection with Robust Training and Abstention

arXiv cs.LG

This paper studies fairness in toxicity classification across three axes: ranking, calibration, and abstention. It compares ERM, reweighted ERM, and Group DRO methods with post-hoc interventions, finding that calibration disparity is a hidden fairness violation and that abstention itself can be unfair.

Consensus as Privileged Context for Label-Free Self-Distillation

arXiv cs.LG

A research paper introducing Canon, a label-free self-distillation method that uses consensus among sampled solutions to provide dense token-level supervision for training large language models on reasoning tasks, improving pass@1 by up to 12 points and outperforming label-free reinforcement learning at a fraction of the compute.

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

arXiv cs.CL

The paper identifies 'Thinking Collapse' in on-policy self-distillation for large language models, characterized by a decline in intermediate reasoning steps, and proposes AD-OPSD, a control framework that mitigates this collapse by anchoring high-suppression-risk tokens to a reference prior. The method achieves up to +4.1% absolute average accuracy improvement on mathematical benchmarks.