Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations
Summary
This paper proposes weighted conformal methods for changepoint localization and root cause analysis that reduce confidence set size under corrupted observations by downweighting likely contaminated data, using uncertainty signals and meta-learning.
View Cached Full Text
Cached at: 07/30/26, 09:58 AM
# Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations
Source: [https://arxiv.org/html/2607.26481](https://arxiv.org/html/2607.26481)
Seunghun Yu, Meiyi Zhu, Petar Popovski, , Joonhyuk Kang, , and Osvaldo SimeoneThis work was partly supported by the Institute of Information & Communications Technology Planning & Evaluation \(IITP\)\-ITRC \(Information Technology Research Center\) grant funded by the Korea government \(MSIT\) \(IITP\-2026\-RS\-2020\-II201787, contribution rate: 50%\); in part by the Institute of Information & Communications Technology Planning & Evaluation \(IITP\) under 6G·Cloud Research and Education Open Hub grant funded by the Korea government \(MSIT\) \(IITP\-2026\-RS\-2024\-00428780, contribution rate: 50%\)\. The work of M\. Zhu and O\. Simeone was supported by an Open Fellowship of the EPSRC \(EP/W024101/1\)\. The work of O\. Simeone was also supported by the EPSRC \(EP/X011852/1\) and the ERC \(No\. 101198347\)\. The work of P\. Popovski was supported, in part, by the Velux Foundation, Denmark, through the Villum Investigator Grant WATER, nr\. 37793\. \(Corresponding authors: Joonhyuk Kang and Osvaldo Simeone\.\) Seunghun Yu and Joonhyuk Kang are with the Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Daejeon 34141, South Korea \(e\-mail: sh0703\.yu@kaist\.ac\.kr; jhkang@ee\.kaist\.ac\.kr\)\.Meiyi Zhu is with the Department of Engineering, King’s College London, WC2R 2LS, London, U\.K\. \(e\-mail: meiyi\.1\.zhu@kcl\.ac\.uk\)\. Petar Popovski is with the Department of Electronic Systems, Aalborg University, 9220 Aalborg, Denmark \(e\-mail: petarp@es\.aau\.dk\)\. Osvaldo Simeone is with the Institute for Intelligent Networked Systems, Northeastern University London, E1 8PH London, U\.K\., and also with the Connectivity Section, Department of Electronic Systems, Aalborg University, 9220 Aalborg, Denmark \(e\-mail: o\.simeone@northeastern\.edu\)\.
###### Abstract
Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monitoring of telecommunication networks, robotic platforms, security infrastructure, and multi\-agent systems\. In safety\- and mission\-critical deployments, such decisions must be accompanied by statistical reliability guarantees rather than by point estimates alone\. Conformal changepoint localization \(CONCH\) and conformal root cause analysis \(CROC\) meet this need by returning confidence sets that contain the true changepoint, or the true root\-cause stream, with a user\-specified probability, without parametric assumptions on the data\-generating process\. In practice, however, observations are frequently corrupted, e\.g\., by outliers, sensor faults, or adversarial perturbations\. While the finite\-sample coverage of these procedures is preserved under contamination, the resulting confidence sets can become uninformatively large\. Adopting a Huber\-type contamination model, this paper proposes weighted CONCH \(W\-CONCH\) and weighted CROC \(W\-CROC\), which downweight observations that are likely to be corrupted with the goal of reducing confidence set size when data may be corrupted\. The weighting mechanism, derived from a formal bound on the unknown corrupted data densities, leverages pre\-existing second\-order classifier\-based uncertainty signals, such as those produced by evidential deep learning or Bayesian learning\. W\-CONCH and W\-CROC are further generalized by introducing a meta\-learning procedure for the weights that optimizes a differentiable surrogate of the confidence set size\. Experiments on image\-based and real\-world changepoint and root\-cause benchmarks show that uncertainty\-based weighting substantially reduces confidence set size while maintaining the target coverage\.
## IIntroduction
Changepoint analysis provides a principled framework for identifying structural changes in ordered data \(see Fig\.[1](https://arxiv.org/html/2607.26481#S1.F1)\)\[[47](https://arxiv.org/html/2607.26481#bib.bib45),[1](https://arxiv.org/html/2607.26481#bib.bib17)\], while its multi\-stream extension, root cause analysis, seeks to attribute an observed change to the component that originated it \(see Fig\.[2](https://arxiv.org/html/2607.26481#S1.F2)\)\[[40](https://arxiv.org/html/2607.26481#bib.bib50),[39](https://arxiv.org/html/2607.26481#bib.bib12)\]\. This paper studies both problems under two key requirements for safety\-critical applications: the outputs must carry statistical reliability guarantees, and they must remain informative even when the observations are corrupted\.
### I\-AContext and Motivation
Determining*when*a system’s behavior changes, and*which*component is the root cause, is a recurring and often safety\-critical task across engineering\. In*telecommunications*, abrupt shifts in traffic volume or flow statistics signal congestion, equipment failures, or intrusions, and detecting and localizing these shifts underpins network management, security monitoring, and the lifecycle of AI\-based applications\[[1](https://arxiv.org/html/2607.26481#bib.bib17),[24](https://arxiv.org/html/2607.26481#bib.bib4),[31](https://arxiv.org/html/2607.26481#bib.bib16),[37](https://arxiv.org/html/2607.26481#bib.bib15)\]\. In*cybersecurity*, a large class of intrusion\-detection tasks is naturally cast as changepoint detection, where a denial\-of\-service attack or a compromised host manifests as an abrupt change in packet\-rate or connection statistics\[[45](https://arxiv.org/html/2607.26481#bib.bib14)\]\. In*robotics*, fault detection and isolation must determine both the onset time of a sensor or actuator fault and the specific faulty component, so that a controller can react before the fault propagates through the platform\[[19](https://arxiv.org/html/2607.26481#bib.bib13)\]\. In*multi\-agent and distributed systems*, from robot swarms to microservice architectures, a fault in a single agent or service can cascade through the collective, and the diagnostic problem is to detect the anomaly and trace it back to the agent or service that triggered it\[[39](https://arxiv.org/html/2607.26481#bib.bib12),[55](https://arxiv.org/html/2607.26481#bib.bib71),[5](https://arxiv.org/html/2607.26481#bib.bib70)\]\. In all of these domains, the*changepoint*time is a proxy for the onset of a fault, and the*root cause*is the component whose change occurred first\.
Figure 1:Illustration of the offline changepoint localization problem with contaminated data\. The clean sequence𝐗~\\widetilde\{\\mathbf\{X\}\}changes at the true changepointξ\\xi, switching from one image style to another\. The clean sequence𝐗~\\widetilde\{\\mathbf\{X\}\}is, however, not observable, as the changepoint detection has access only to the observed sequence𝐗\\mathbf\{X\}, which is obtained by independently contaminating observations with unknown probabilityε\\varepsilon\. The goal is to construct a confidence set𝒞α\(𝐗\)\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)that includes the true changepointξ\\xiwith probability no smaller than1−α1\-\\alpha\.A defining feature of these applications is that the output of a changepoint or root\-cause procedure drives a potentially costly downstream action: rerouting traffic, quarantining a host, halting a robot, or isolating a microservice\. A single point estimate, reported without any statement of confidence, is therefore of limited value, because an operator cannot tell whether to trust it\. Recent work\[[21](https://arxiv.org/html/2607.26481#bib.bib11)\]has shown that, for a*risk\-averse*decision maker, it is optimal to adopt decisions that maximize the given utility function for a worst\-case outcome consistent with the observations\. Specifically, the appropriate interface between uncertainty quantification and decision making is a*set*of outcomes that provably contains the true one with probability at least1−α1\-\\alpha, with probabilityα\\alphadictating the level of risk tolerance\.
Returning a*set*of plausible changepoints or root causes thus allows a decision maker to control its risk by selecting actions that are robust across all elements of the set\. For example, an operator may inspect all the plausible causes of a failure in telecommunication networks on the basis of a set prediction\. However, in order for this process to be efficient, it is of critical importance that the predicted set be as small as possible\.
Figure 2:Illustration of the problem of offline root\-cause localization\. Each observed stream𝐗d\\mathbf\{X\}\_\{d\}withd=1,…,Dd=1,\\ldots,D, undergoes a distributional shift at changepointξd\\xi\_\{d\}, while some observations may be corrupted\. The goal is to estimate the stream that has the earliest changepointd⋆d^\{\\star\}by constructing a confidence set𝒦α\(𝐗\)\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)that includesd⋆d^\{\\star\}with probability no smaller than1−α1\-\\alpha\.Real observations are routinely*contaminated*by outliers, sensor faults, packet losses, or adversarial perturbations\. Contamination can obscure the true change signal and true root cause of a change\. Accordingly, even procedures whose coverage guarantee remains valid under contamination, such as*conformal changepoint localization*\(CONCH\)\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]and*conformal root cause analysis*\(CROC\)\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\], can be forced to return confidence sets so large as to be uninformative\. The goal of this paper is to introduce a methodology that, building on CONCH and CROC, preserves validity under contamination while restoring informativeness of set predictors for both single\-stream changepoint localization and multi\-stream root cause analysis\.
### I\-BRelated Work
#### I\-B1Changepoint detection
Classical changepoint methods include likelihood\-ratio tests, CUSUM procedures, nonparametric tests, and kernel\-based methods\[[27](https://arxiv.org/html/2607.26481#bib.bib46),[20](https://arxiv.org/html/2607.26481#bib.bib47),[30](https://arxiv.org/html/2607.26481#bib.bib48),[41](https://arxiv.org/html/2607.26481#bib.bib69)\], covering both offline and online settings\[[47](https://arxiv.org/html/2607.26481#bib.bib45),[1](https://arxiv.org/html/2607.26481#bib.bib17),[33](https://arxiv.org/html/2607.26481#bib.bib68)\]\. The theoretical guarantees of these standard methods typically rely on parametric assumptions, asymptotic approximations, or structural conditions, and most target detection or point localization rather than finite\-sample confidence sets for the change location\. The same limitations are shared by studies that address robustness to contamination, including adversarially robust offline detection under dynamic Huber contamination\[[25](https://arxiv.org/html/2607.26481#bib.bib43)\]and robust online mean\-change detection under heavy\-tailed noise with delay and false\-alarm analysis\[[44](https://arxiv.org/html/2607.26481#bib.bib42)\]\. In contrast to these studies, CONCH\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]produces a set of plausible changepoints that satisfies finite\-sample coverage under a basic split\-exchangeability condition, without requiring parametric assumptions\.
#### I\-B2Root cause analysis
Root cause analysis has been developed largely within specific domains\[[40](https://arxiv.org/html/2607.26481#bib.bib50),[51](https://arxiv.org/html/2607.26481#bib.bib51)\]\. For example, in microservice and distributed systems, causal\-discovery and outlier\-attribution methods localize the service responsible for a failure\[[39](https://arxiv.org/html/2607.26481#bib.bib12)\], while in robotics fault detection and isolation identify faulty components or anomalous agents\[[19](https://arxiv.org/html/2607.26481#bib.bib13)\]\. The recently proposed CROC\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\]departs from this pattern by constructing distribution\-free confidence sets for the stream whose changepoint occurs first\.
#### I\-B3Conformal prediction
CONCH and CROC build on the general methodology of conformal prediction\. Conformal prediction offers distribution\-free, finite\-sample predictive inference under exchangeability\[[50](https://arxiv.org/html/2607.26481#bib.bib54),[36](https://arxiv.org/html/2607.26481#bib.bib53),[34](https://arxiv.org/html/2607.26481#bib.bib55)\]\. Extensions relax exchangeability or accommodate distribution shift, including weighted conformal prediction under covariate shift\[[46](https://arxiv.org/html/2607.26481#bib.bib5),[52](https://arxiv.org/html/2607.26481#bib.bib72)\], conformal prediction beyond exchangeability\[[2](https://arxiv.org/html/2607.26481#bib.bib6)\], and adaptive conformal inference for online shift\[[11](https://arxiv.org/html/2607.26481#bib.bib7),[54](https://arxiv.org/html/2607.26481#bib.bib10)\]\. Besides the offline settings considered by CONCH and CROC, conformal prediction has been adapted to online changepoint analysis through martingale tests\[[49](https://arxiv.org/html/2607.26481#bib.bib41)\]\.
#### I\-B4Robustness and uncertainty estimation
The contamination model adopted here follows Huber’s framework\[[17](https://arxiv.org/html/2607.26481#bib.bib29),[12](https://arxiv.org/html/2607.26481#bib.bib62)\]and is related to data\-poisoning settings\[[42](https://arxiv.org/html/2607.26481#bib.bib59)\]\. The weighting scheme relies on classifier uncertainty as a proxy for the probability that an observation is clean, drawing on evidential deep learning \(EDL\)\[[35](https://arxiv.org/html/2607.26481#bib.bib20)\], Monte Carlo dropout\[[9](https://arxiv.org/html/2607.26481#bib.bib32)\], deep ensembles\[[23](https://arxiv.org/html/2607.26481#bib.bib39)\], and calibration under shift\[[16](https://arxiv.org/html/2607.26481#bib.bib38),[38](https://arxiv.org/html/2607.26481#bib.bib24)\]\. Finally, when the contamination level is unknown, we learn the uncertainty\-to\-weight mapping via meta\-learning\[[8](https://arxiv.org/html/2607.26481#bib.bib18),[4](https://arxiv.org/html/2607.26481#bib.bib9)\], optimizing a differentiable surrogate of the conformal set size in the spirit of conformal\-aware training\[[43](https://arxiv.org/html/2607.26481#bib.bib3),[28](https://arxiv.org/html/2607.26481#bib.bib22)\]\.
### I\-CMain Contributions
This paper develops offline changepoint detection and root cause analysis methods that remain valid and efficient even when observations are contaminated under Huber’s model\. Our contributions are as follows\.
- •Weighted conformal changepoint localization\.We propose*weighted CONCH*\(W\-CONCH\), which builds on CONCH\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]by incorporating a weighting mechanism that downweights observations that are likely to be corrupted\. W\-CONCH hinges on a novel changepoint plausibility \(CPP\) score that is derived from a bound on the contaminated marginal likelihood, and is instantiated through an efficient classifier\-based implementation\.
- •Uncertainty\-based and meta\-learned weights\.We introduce uncertainty\-based weighting rules driven by classifier second\-order uncertainty signals obtained from EDL or Bayesian learning\. We also present a meta\-learned variant, MW\-CONCH, that directly optimizes a differentiable surrogate of the confidence set size, reducing dependence on hyperparameters\.
- •Extension to root cause analysis\.We extend the same weighting principle to multi\-stream root\-cause localization building on CROC\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\]\. This yields weighted CROC \(W\-CROC\) and its meta\-learned variant MW\-CROC, which preserve the distribution\-free coverage guarantee of CROC while sharpening the root\-cause confidence set\.
- •Empirical validation\.On image\-based changepoint and root\-cause benchmarks, along with a real\-world changepoint benchmark, we show that uncertainty\-based weighting substantially reduces confidence set size under contamination while retaining empirical coverage, with meta\-learning achieving performance comparable to oracle baselines that use the true contamination indicators to define the weights\.
### I\-DOrganization
The remainder of the paper is organized as follows\. Section[II](https://arxiv.org/html/2607.26481#S2)formalizes the contaminated changepoint localization and root\-cause localization problems\. Section[III](https://arxiv.org/html/2607.26481#S3)reviews CONCH and establishes its coverage guarantee under contaminated observations\. Section[IV](https://arxiv.org/html/2607.26481#S4)presents W\-CONCH, including the weighted changepoint\-plausibility score, the uncertainty\-based weighting rules, and the meta\-learning procedure\. Section[V](https://arxiv.org/html/2607.26481#S5)extends the weighting strategy to the multi\-stream CROC framework\. Section[VI](https://arxiv.org/html/2607.26481#S6)reports the experimental results, and the appendices contain proofs, implementation details, and additional experiments\.
## IIProblem Setup
This paper studies changepoint localization and root cause analysis in a batch setting with corrupted observations\. As shown in Fig\.[1](https://arxiv.org/html/2607.26481#S1.F1), in changepoint localization, the distribution of an ordered sequence of observations changes at a single unknown index, referred to as the changepoint\. The goal is to construct a confidence set for the changepoint that achieves finite\-sample coverage while remaining as small as possible, despite possible data contamination\. As sketched in Fig\.[2](https://arxiv.org/html/2607.26481#S1.F2), root cause analysis can be formulated as an extension of changepoint localization with multiple observed sequences in which one wishes to identify which sequence is the root cause of a change\. The goal here is to provide a set of possible root causes with coverage guarantees, even when the observations may be corrupted\. As we define in the following sections, the proposed techniques specialize the methodologies presented in\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]and\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\]in order to account for data contamination, while retaining statistical validity\.
### II\-AChangepoint Localization
As illustrated in Fig\.[1](https://arxiv.org/html/2607.26481#S1.F1), let𝒳\\mathcal\{X\}denote the observation space, and let𝐗~=\(X~1,…,X~n\)∈𝒳n\\widetilde\{\\mathbf\{X\}\}=\(\\widetilde\{X\}\_\{1\},\\dots,\\widetilde\{X\}\_\{n\}\)\\in\\mathcal\{X\}^\{n\}denote the clean data sequence\. We assume that there exists a true changepointξ∈\{1,…,n−1\}\\xi\\in\\\{1,\\dots,n\-1\\\}at which the distribution of the clean data𝐗~\\widetilde\{\\mathbf\{X\}\}changes\. As in\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\], we impose no parametric assumptions on the pre\- and post\-change distributions, requiring only the following exchangeability condition\.
###### Assumption 1\(Split exchangeability\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]\)\.
The clean sequenceX~\\widetilde\{X\}is split\-exchangeable at the true changepointξ\\xi: the pre\- and post\-change segments of the clean sequence𝐗~\\widetilde\{\\mathbf\{X\}\}are exchangeable, i\.e\.,
\(X~1,…,X~ξ\)\\displaystyle\(\{\\widetilde\{X\}\}\_\{1\},\\dots,\{\\widetilde\{X\}\}\_\{\\xi\}\)=d\(X~π0,ξ\(1\),…,X~π0,ξ\(ξ\)\),\\displaystyle\\overset\{\\mathrm\{d\}\}\{=\}\(\{\\widetilde\{X\}\}\_\{\\pi\_\{0,\\xi\}\(1\)\},\\dots,\{\\widetilde\{X\}\}\_\{\\pi\_\{0,\\xi\}\(\\xi\)\}\),\(1a\)\(X~ξ\+1,…,X~n\)\\displaystyle\(\{\\widetilde\{X\}\}\_\{\\xi\+1\},\\dots,\{\\widetilde\{X\}\}\_\{n\}\)=d\(X~π1,ξ\(ξ\+1\),…,X~π1,ξ\(n\)\),\\displaystyle\\overset\{\\mathrm\{d\}\}\{=\}\(\{\\widetilde\{X\}\}\_\{\\pi\_\{1,\\xi\}\(\\xi\+1\)\},\\dots,\{\\widetilde\{X\}\}\_\{\\pi\_\{1,\\xi\}\(n\)\}\),\(1b\)where=𝑑\\overset\{d\}\{=\}denotes equality in distribution, whileπ0,ξ\(⋅\)\\pi\_\{0,\\xi\}\(\\cdot\)andπ1,ξ\(⋅\)\\pi\_\{1,\\xi\}\(\\cdot\)are arbitrary permutations of the sets\{1,…,ξ\}\\\{1,\\dots,\\xi\\\}and\{ξ\+1,…,n\}\\\{\\xi\+1,\\dots,n\\\}, respectively\.
LetY1,…,Yn∼i\.i\.d\.Bernoulli\(ε\)Y\_\{1\},\\ldots,Y\_\{n\}\\overset\{i\.i\.d\.\}\{\\sim\}\\mathrm\{Bernoulli\}\(\\varepsilon\)be independent and identically distributed \(i\.i\.d\.\) contamination indicators independent of the clean data𝐗~\\widetilde\{\\mathbf\{X\}\}for some corruption probabilityε∈\(0,1\)\\varepsilon\\in\(0,1\)\. Furthermore, letZ1,…,Zn∼i\.i\.d\.QZ\_\{1\},\\dots,Z\_\{n\}\\overset\{i\.i\.d\.\}\{\\sim\}Qdenote independent draws from a contamination distributionQQ\. Following Huber’s contamination model\[[17](https://arxiv.org/html/2607.26481#bib.bib29)\], each sample in the sequence of clean observations𝐗~=\(X~1,…,X~n\)\\widetilde\{\\mathbf\{X\}\}=\(\\widetilde\{X\}\_\{1\},\\ldots,\\widetilde\{X\}\_\{n\}\)is independently corrupted with probabilityε\\varepsilon, yielding the observation
Xi=\(1−Yi\)X~i\+YiZi,i=1,…,n\.X\_\{i\}=\(1\-Y\_\{i\}\)\\widetilde\{X\}\_\{i\}\+Y\_\{i\}Z\_\{i\},\\qquad i=1,\\dots,n\.\(2\)Accordingly, with probability1−ε1\-\\varepsilon, the observed sampleXiX\_\{i\}equals the clean sampleX~i\\widetilde\{X\}\_\{i\}, while with probabilityε\\varepsilon, the observed sampleXiX\_\{i\}is given by the noisy sampleZi∼QZ\_\{i\}\\sim Q\.
Given the noisy observed sequence𝐗\\mathbf\{X\}in \([2](https://arxiv.org/html/2607.26481#S2.E2)\), we aim to construct a confidence set𝒞α\(𝐗\)⊆\{1,…,n−1\}\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\\subseteq\\\{1,\\dots,n\-1\\\}for the unknown changepointξ\\xithat satisfies the finite\-sample coverage guarantee
Pr\(ξ∈𝒞α\(𝐗\)\)≥1−α\\displaystyle\\Pr\\bigl\(\\xi\\in\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\\bigr\)\\geq 1\-\\alpha\(3\)at a user\-specified levelα∈\(0,1\)\\alpha\\in\(0,1\), while reducing as much as possible the set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|\. This problem is studied in Section[III](https://arxiv.org/html/2607.26481#S3)and Section[IV](https://arxiv.org/html/2607.26481#S4)\.
### II\-BRoot Cause Analysis
As illustrated in Fig\.[2](https://arxiv.org/html/2607.26481#S1.F2), consider nowDDdata streams, each of lengthnn\. For streamd∈\{1,…,D\}d\\in\\\{1,\\ldots,D\\\}, let𝐗~d=\(X~d,1,…,X~d,n\)\\widetilde\{\\mathbf\{X\}\}\_\{d\}=\(\\widetilde\{X\}\_\{d,1\},\\ldots,\\widetilde\{X\}\_\{d,n\}\)denote the clean sequence, and letξd∈\{1,…,n−1\}\\xi\_\{d\}\\in\\\{1,\\ldots,n\-1\\\}denote its changepoint\. The vector of stream\-wise changepoints is denoted by𝝃=\(ξ1,…,ξD\)\\boldsymbol\{\\xi\}=\(\\xi\_\{1\},\\ldots,\\xi\_\{D\}\)\. The root\-cause index is defined as the index of the stream whose changepoint occurs first, i\.e\.,
d⋆=argmind∈\{1,…,D\}ξd,d^\{\\star\}=\\arg\\min\_\{d\\in\\\{1,\\ldots,D\\\}\}\\xi\_\{d\},\(4\)where the minimizer is assumed to be unique\. The rationale for this definition is that the first changepoint may be the root cause of the changes in the other sequences\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\]\.
Generalizing the single\-stream setting, following\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\], we assume that, for each streamddthe clean sequence𝐗~d\\widetilde\{\\mathbf\{X\}\}\_\{d\}satisfies split exchangeability at its true changepointξd\\xi\_\{d\}\.
###### Assumption 2\(Per\-stream split exchangeability and independence\)\.
Any clean sequence𝐗~d\\widetilde\{\\mathbf\{X\}\}\_\{d\}is split\-exchangeable at the changepointξd\\xi\_\{d\}: For any stream\-wise split permutationsπ0,ξd\\pi\_\{0,\\xi\_\{d\}\}of the set\{1,…,ξd\}\\\{1,\\ldots,\\xi\_\{d\}\\\}andπ1,ξd\\pi\_\{1,\\xi\_\{d\}\}of the set\{ξd\+1,…,n\}\\\{\\xi\_\{d\}\+1,\\ldots,n\\\}, we have the equivalence in distribution
\(X~d,1,…,X~d,ξd\)=d\(X~d,π0,ξd\(1\),…,X~d,π0,ξd\(ξd\)\),\\displaystyle\(\\widetilde\{X\}\_\{d,1\},\\ldots,\\widetilde\{X\}\_\{d,\\xi\_\{d\}\}\)\\overset\{\\mathrm\{d\}\}\{=\}\(\\widetilde\{X\}\_\{d,\\pi\_\{0,\\xi\_\{d\}\}\(1\)\},\\ldots,\\widetilde\{X\}\_\{d,\\pi\_\{0,\\xi\_\{d\}\}\(\\xi\_\{d\}\)\}\),\(5a\)and\(X~d,ξd\+1,…,X~d,n\)=d\(X~d,π1,ξd\(ξd\+1\),…,X~d,π1,ξd\(n\)\)\.\\displaystyle\\text\{and \}\(\\widetilde\{X\}\_\{d,\\xi\_\{d\}\+1\},\\ldots,\\widetilde\{X\}\_\{d,n\}\)\\overset\{\\mathrm\{d\}\}\{=\}\(\\widetilde\{X\}\_\{d,\\pi\_\{1,\\xi\_\{d\}\}\(\\xi\_\{d\}\+1\)\},\\ldots,\\widetilde\{X\}\_\{d,\\pi\_\{1,\\xi\_\{d\}\}\(n\)\}\)\.\(5b\)Furthermore, we assume that the clean streams𝐗~1,𝐗~2,…,𝐗~D\\widetilde\{\\mathbf\{X\}\}\_\{1\},\\widetilde\{\\mathbf\{X\}\}\_\{2\},\\ldots,\\widetilde\{\\mathbf\{X\}\}\_\{D\}are mutually independent\.
Each stream is corrupted independently through a stream\-wise Huber contamination mechanism\. Accordingly, letYd,1,…,Yd,n∼i\.i\.d\.Bernoulli\(εd\)Y\_\{d,1\},\\ldots,Y\_\{d,n\}\\overset\{i\.i\.d\.\}\{\\sim\}\\mathrm\{Bernoulli\}\(\\varepsilon\_\{d\}\)denote the contamination indicators for observationiiin streamdd, and letZd,1,…,Zd,n∼i\.i\.d\.QdZ\_\{d,1\},\\ldots,Z\_\{d,n\}\\overset\{i\.i\.d\.\}\{\\sim\}Q\_\{d\}denote an independent draw from a stream\-specific contamination distributionQdQ\_\{d\}\. Thedd\-th observed sequence is given by
Xd,i=\(1−Yd,i\)X~d,i\+Yd,iZd,iX\_\{d,i\}=\(1\-Y\_\{d,i\}\)\\widetilde\{X\}\_\{d,i\}\+Y\_\{d,i\}Z\_\{d,i\}\(6\)for observation indexi=1,…,ni=1,\\ldots,nand stream indexd=1,…,Dd=1,\\ldots,D\.
Let𝐗=\(𝐗1,…,𝐗D\)\\mathbf\{X\}=\(\\mathbf\{X\}\_\{1\},\\ldots,\\mathbf\{X\}\_\{D\}\)denote the collection of all observed streams\. The goal of root cause analysis is to construct a confidence set𝒦α\(𝐗\)⊆\{1,…,D\}\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\\subseteq\\\{1,\\ldots,D\\\}for the unknown root\-cause indexd⋆d^\{\\star\}such that it includes the true indexd⋆d^\{\\star\}in \([4](https://arxiv.org/html/2607.26481#S2.E4)\) with probability no smaller than the user\-defined level1−α1\-\\alpha, i\.e\.,
Pr\(d⋆∈𝒦α\(𝐗\)\)≥1−α\.\\Pr\\bigl\(d^\{\\star\}\\in\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\\bigr\)\\geq 1\-\\alpha\.\(7\)The root\-cause localization problem is studied in Sec\.[V](https://arxiv.org/html/2607.26481#S5)\.
## IIIConformal Changepoint Localization with Contaminated Observations
In this section, we first review the conformal changepoint localization scheme introduced in\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\], which is referred to as CONCH\. Then, we show that the coverage guarantees of CONCH carry over to the contaminated observation model \([2](https://arxiv.org/html/2607.26481#S2.E2)\)\. Based on this result, in the next section we will develop the proposed W\-CONCH\.
### III\-AConformal Changepoint Localization
For each candidate changepointt∈\{1,…,n−1\}t\\in\\\{1,\\dots,n\-1\\\}, CONCH\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]evaluates a CPP scoreSt\(𝐗\)S\_\{t\}\(\\mathbf\{X\}\)\. This measures how likely time indexttis to be the true changepoint, with larger values indicating stronger evidence in favor of candidatett\. The CPP score is treated here as fixed and arbitrary, and we will discuss the optimization of the CPP score in the next section\.
LetΠt\\Pi\_\{t\}denote the set of permutationsπ0,t\(⋅\)\\pi\_\{0,t\}\(\\cdot\)andπ1,t\(⋅\)\\pi\_\{1,t\}\(\\cdot\)that operate on the indices\{1,…,t\}\\\{1,\\dots,t\\\}and\{t\+1,…,n\}\\\{t\+1,\\dots,n\\\}, respectively, as in Assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1)\. Furthermore, denote asπt\(𝐗\)\\pi\_\{t\}\(\\mathbf\{X\}\)the permuted sequence\(Xπ0,t\(1\),…,Xπ0,t\(t\),Xπ1,t\(t\+1\),…,Xπ1,t\(n\)\)\\big\(X\_\{\\pi\_\{0,t\}\(1\)\},\\ldots,X\_\{\\pi\_\{0,t\}\(t\)\},\\allowbreak X\_\{\\pi\_\{1,t\}\(t\+1\)\},\\ldots,X\_\{\\pi\_\{1,t\}\(n\)\}\\big\)for some pair of permutationsπt=\(π0,t,π1,t\)\\pi\_\{t\}=\(\\pi\_\{0,t\},\\pi\_\{1,t\}\)\. The conformalpp\-value for candidate changepointttis defined as the fraction of permutationsπt∈Πt\\pi\_\{t\}\\in\\Pi\_\{t\}for which the CPP scoreSt\(πt\(𝐗\)\)S\_\{t\}\(\\pi\_\{t\}\(\\mathbf\{X\}\)\)does not exceed the actual CPP scoreSt\(𝐗\)S\_\{t\}\(\\mathbf\{X\}\):
pt\(𝐗\)=1\|Πt\|∑πt∈Πt𝟙\[St\(πt\(𝐗\)\)≤St\(𝐗\)\]\.p\_\{t\}\(\\mathbf\{X\}\)=\\frac\{1\}\{\|\\Pi\_\{t\}\|\}\\sum\_\{\\pi\_\{t\}\\in\\Pi\_\{t\}\}\\mathds\{1\}\\bigl\[S\_\{t\}\(\\pi\_\{t\}\(\\mathbf\{X\}\)\)\\leq S\_\{t\}\(\\mathbf\{X\}\)\\bigr\]\.\(8\)Intuitively, the statisticpt\(𝐗\)p\_\{t\}\(\\mathbf\{X\}\)in \([8](https://arxiv.org/html/2607.26481#S3.E8)\) tends to be large if the CPP scoreSt\(𝐗\)S\_\{t\}\(\\mathbf\{X\}\)is large, and thus there is evidence forttbeing the true changepoint\.
In practice, when the number of data pointsnnis sufficiently large, enumerating all the split permutations in setΠt\{\\Pi\}\_\{t\}is computationally infeasible\. In this case, a Monte Carlo \(MC\) variant of CONCH, referred to as CONCH\-MC, replaces the full average over setΠt\\Pi\_\{t\}in \([8](https://arxiv.org/html/2607.26481#S3.E8)\) with an average over randomly sampled split permutations\.
The conformalpp\-value \([8](https://arxiv.org/html/2607.26481#S3.E8)\), and its MC version, are validpp\-values for the null hypothesisH0:ξ=tH\_\{0\}:\\xi=tthat the true changepointξ\\xiequalstt\. Accordingly, given a miscoverage levelα∈\(0,1\)\\alpha\\in\(0,1\), the changepoint confidence set is obtained by inverting the test with null hypothesisH0:ξ=tH\_\{0\}:\\xi=t\[[32](https://arxiv.org/html/2607.26481#bib.bib63)\], i\.e\., by retaining all candidates whosepp\-values exceed the levelα\\alpha:
𝒞α\(𝐗\)=\{t∈\{1,…,n−1\}:pt\>α\}\.\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)=\\bigl\\\{t\\in\\\{1,\\dots,n\-1\\\}:p\_\{t\}\>\\alpha\\bigr\\\}\.\(9\)In other words, the confidence set retains all candidate changepoints that cannot be rejected by the split\-permutation test withpp\-value \([8](https://arxiv.org/html/2607.26481#S3.E8)\) at levelα\\alpha\.
### III\-BCoverage Guarantee of CONCH under Contamination
Reference\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]shows that the CONCH set \([9](https://arxiv.org/html/2607.26481#S3.E9)\) satisfies the marginal coverage guarantee \([3](https://arxiv.org/html/2607.26481#S2.E3)\) as long as the sequence𝐗\\mathbf\{X\}is split\-exchangeable at the true changepointξ\\xi\. The following makes it possible to extend the coverage property also to the noisy observations \([2](https://arxiv.org/html/2607.26481#S2.E2)\) under Assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1), which assumes split\-exchangeability only for the clean sequence𝐗~\\widetilde\{\\mathbf\{X\}\}\.
###### Lemma 1\(Split exchangeability under contamination\)\.
Under Assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1), the corrupted observed sequence𝐗\\mathbf\{X\}generated according to \([2](https://arxiv.org/html/2607.26481#S2.E2)\) satisfies split exchangeability at the true changepointξ\\xi\.
###### Proof\.
See Appendix[A\-A](https://arxiv.org/html/2607.26481#A1.SS1)\. ∎
This result has the following immediate consequence\.
###### Lemma 2\(Coverage guarantee under contamination\)\.
Under Assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1), the CONCH set𝒞α\(𝐗\)\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)defined in \([9](https://arxiv.org/html/2607.26481#S3.E9)\), when applied to the observed corrupted sequence𝐗\\mathbf\{X\}in \([2](https://arxiv.org/html/2607.26481#S2.E2)\), satisfies the required coverage condition \([3](https://arxiv.org/html/2607.26481#S2.E3)\)\.
###### Proof\.
By Lemma[1](https://arxiv.org/html/2607.26481#Thmlemma1), the claim follows directly from\[[14](https://arxiv.org/html/2607.26481#bib.bib2), Thm\. 3\.1\]\. ∎
Lemma[2](https://arxiv.org/html/2607.26481#Thmlemma2)ensures that the CONCH set \([9](https://arxiv.org/html/2607.26481#S3.E9)\) meets the coverage requirement in \([3](https://arxiv.org/html/2607.26481#S2.E3)\) regardless of the contamination levelε\\varepsilon\. However, unless the CPP scoreSt\(X\)S\_\{t\}\(X\)is properly designed, contaminated observations can obscure the changepoint signal and inflate the confidence set \(see Fig\.[1](https://arxiv.org/html/2607.26481#S1.F1)\)\. We address this efficiency issue by introducing novel CPP scores in the next section\.
## IVWeighted Conformal Changepoint Localization
Lemma[2](https://arxiv.org/html/2607.26481#Thmlemma2)ensures that the CONCH set \([9](https://arxiv.org/html/2607.26481#S3.E9)\) remains valid under the contamination model \([2](https://arxiv.org/html/2607.26481#S2.E2)\), but contaminated observations generally increase the size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|of the confidence set, making it more difficult in identifying the changepoint\. To address this issue, in this section we introduce CPP scores that aim at downweighting observations that are likely to be corrupted\. To this end, we first consider an oracle\-based setting in which the model \([1](https://arxiv.org/html/2607.26481#S2.E1)\)–\([2](https://arxiv.org/html/2607.26481#S2.E2)\) is known, and then introduce a practical implementation based on uncertainty signals\. Finally, we present a meta\-learning\-based strategy that leverages offline data from multiple tasks to further improve the CPP score\.
### IV\-AWeighted CPP Score
In this section, we consider a simplified oracle\-based setting in order to facilitate the derivation of an optimal CPP score under contamination\. To this end, following\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\], we specialize the data model in Sec\.[II](https://arxiv.org/html/2607.26481#S2)to a standard i\.i\.d\. changepoint model with known probability distribution\[[20](https://arxiv.org/html/2607.26481#bib.bib47),[30](https://arxiv.org/html/2607.26481#bib.bib48),[6](https://arxiv.org/html/2607.26481#bib.bib31)\]\.
###### Assumption 3\(i\.i\.d\. oracle changepoint model\)\.
The clean pre\- and post\-change distributions admit known densitiesf0f\_\{0\}andf1f\_\{1\}, so that the pre\-change samplesX~1,…,X~ξ\\widetilde\{X\}\_\{1\},\\ldots,\\widetilde\{X\}\_\{\\xi\}are drawn i\.i\.d\. from a distribution with densityf0\(x\)f\_\{0\}\(x\)and post\-change samplesX~ξ\+1,…,X~n\\widetilde\{X\}\_\{\\xi\+1\},\\ldots,\\widetilde\{X\}\_\{n\}are drawn i\.i\.d\. from a distribution with densityf1\(x\)f\_\{1\}\(x\):
X~1,…,X~ξ∼i\.i\.df0\(x\),X~ξ\+1,…,X~n∼i\.i\.df1\(x\)\.\\widetilde\{X\}\_\{1\},\\ldots,\\widetilde\{X\}\_\{\\xi\}\\overset\{i\.i\.d\}\{\\sim\}f\_\{0\}\(x\),\\quad\\widetilde\{X\}\_\{\\xi\+1\},\\ldots,\\widetilde\{X\}\_\{n\}\\overset\{i\.i\.d\}\{\\sim\}f\_\{1\}\(x\)\.\(10\)Furthermore, each observationXiX\_\{i\}is independently drawn from the contaminated marginal density
gj\(x\)=\(1−ε\)fj\(x\)\+εq\(x\),g\_\{j\}\(x\)=\(1\-\\varepsilon\)f\_\{j\}\(x\)\+\\varepsilon q\(x\),\(11\)withj=0j=0for pre\-change samplesi=1,…,ξi=1,\\ldots,\\xiandj=1j=1for post\-change samplesi=ξ\+1,…,ni=\\xi\+1,\\ldots,n, where probabilityε\\varepsilonand densityq\(x\)q\(x\)are known\.
#### IV\-A1Optimal Oracle CPP Score
Under Assumption[3](https://arxiv.org/html/2607.26481#Thmassumption3), the marginal densitiesg0\(x\)g\_\{0\}\(x\)andg1\(x\)g\_\{1\}\(x\)in \([11](https://arxiv.org/html/2607.26481#S4.E11)\) are known, and so is the true changepointξ\\xi\. Under this idealized assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1), the following likelihood\-ratio statistic can be proved to yield the optimal CPP score that minimizes the expected size of the confidence set𝒞α\(𝐗\)\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\.
###### Lemma 3\(Optimal CPP score under contamination\)\.
Under Assumption[3](https://arxiv.org/html/2607.26481#Thmassumption3), the CPP score
Stopt\(𝐗\)=∏i≤tg0\(Xi\)∏i\>tg1\(Xi\)∏i≤ξg0\(Xi\)∏i\>ξg1\(Xi\)S^\{\\mathrm\{opt\}\}\_\{t\}\(\\mathbf\{X\}\)=\\frac\{\\prod\_\{i\\leq t\}g\_\{0\}\(X\_\{i\}\)\\prod\_\{i\>t\}g\_\{1\}\(X\_\{i\}\)\}\{\\prod\_\{i\\leq\\xi\}g\_\{0\}\(X\_\{i\}\)\\prod\_\{i\>\\xi\}g\_\{1\}\(X\_\{i\}\)\}\(12\)attains the minimum average set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|among all CONCH sets defined in \([9](https://arxiv.org/html/2607.26481#S3.E9)\)\.
###### Proof\.
Follows directly from\[[14](https://arxiv.org/html/2607.26481#bib.bib2), Thm\. 4\.3\]\. ∎
#### IV\-A2Weighted CPP Score
The score in \([12](https://arxiv.org/html/2607.26481#S4.E12)\) is theoretically optimal but practically inaccessible, since the contaminated marginalsg0\(x\)g\_\{0\}\(x\)andg1\(x\)g\_\{1\}\(x\)in \([11](https://arxiv.org/html/2607.26481#S4.E11)\) depend on a number of unknowns, namely the clean densitiesf0\(x\)f\_\{0\}\(x\)andf1\(x\)f\_\{1\}\(x\), the contamination levelε\\varepsilon, and the contaminated samples densityq\(x\)q\(x\), while the denominator further requires the unknown true changepointξ\\xi\. To obtain an implementable score, we make the following substitutions: 1\) each unknown clean densityfj\(x\)f\_\{j\}\(x\)is replaced by a learned estimatef^j\(x\)\\hat\{f\}\_\{j\}\(x\); 2\) the true changepointξ\\xiin the denominator is replaced by an estimateξ^\(𝐗\)\\hat\{\\xi\}\(\\mathbf\{X\}\); and 3\) each contaminated likelihoodgj\(x\)g\_\{j\}\(x\)is replaced by a tractable bound\.
The approximations 1\) and 2\) follow reference\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\], and they will be further discussed below for completeness\. In contrast, the approximation 3\) is specific to our proposed CPP score, and is derived next via a bound on the distribution \([11](https://arxiv.org/html/2607.26481#S4.E11)\)\.
###### Proposition 1\.
Assume that the observation space𝒳\\mathcal\{X\}is bounded, and let the contamination densityqqis bounded away from zero and infinity on its support, i\.e\.,\|logq\(x\)\|≤Cq\|\\log q\(x\)\|\\leq C\_\{q\}for some constantCq<∞C\_\{q\}<\\inftyfor allxx\. Then, the marginal densitygj\(x\)g\_\{j\}\(x\)in \([11](https://arxiv.org/html/2607.26481#S4.E11)\) can be lower bounded as
gj\(x\)≥C⋅fj\(x\)wj\(x\),g\_\{j\}\(x\)\\;\\geq\\;C\\cdot f\_\{j\}\(x\)^\{w\_\{j\}\(x\)\},\(13\)forj=0,1j=0,1wherewj:𝒳→\[0,1\]w\_\{j\}:\\mathcal\{X\}\\to\[0,1\]is a weight function, andCCis a constant depending only onCqC\_\{q\}andε\\varepsilon\.
###### Proof\.
As detailed in Appendix[A\-B](https://arxiv.org/html/2607.26481#A1.SS2), the inequality \([13](https://arxiv.org/html/2607.26481#S4.E13)\) follows by applying the evidence lower bound \(ELBO\)\[[38](https://arxiv.org/html/2607.26481#bib.bib24)\]to the log\-marginalloggj\(x\)\\log g\_\{j\}\(x\)and then bounding the resulting remainder term using the stated boundedness assumption\. ∎
Substituting the bound \([13](https://arxiv.org/html/2607.26481#S4.E13)\) for each unknown densitygj\(Xi\)g\_\{j\}\(X\_\{i\}\)in the optimal score \([12](https://arxiv.org/html/2607.26481#S4.E12)\), as well as the unknown clean densitiesfj\(x\)f\_\{j\}\(x\)by estimatesf^j\(x\)\\hat\{f\}\_\{j\}\(x\)and the true changepointξ\\xiby an estimateξ^\(𝐗\)\\hat\{\\xi\}\(\\mathbf\{X\}\)yields the proposed weighted CPP score defined next\.
###### Definition 1\(Weighted CPP Score\)\.
Given density estimatesf^j\(x\)\\hat\{f\}\_\{j\}\(x\)forj=0,1j=0,1, changepoint locator estimateξ^\(𝐗\)\\hat\{\\xi\}\(\\mathbf\{X\}\), and arbitrary weight functionsw0\(𝐗\)w\_\{0\}\(\\mathbf\{X\}\)andw1\(𝐗\)w\_\{1\}\(\\mathbf\{X\}\), the weighted CPP score is defined as
Stw\(𝐗\)\\displaystyle S\_\{t\}^\{w\}\(\\mathbf\{X\}\)=log\(∏i≤tf^0\(Xi\)w0\(Xi\)∏i\>tf^1\(Xi\)w1\(Xi\)∏i≤ξ^\(𝐗\)f^0\(Xi\)w0\(Xi\)∏i\>ξ^\(𝐗\)f^1\(Xi\)w1\(Xi\)\)\.\\displaystyle=\\log\\\!\\left\(\\frac\{\\prod\_\{i\\leq t\}\\hat\{f\}\_\{0\}\(X\_\{i\}\)^\{w\_\{0\}\(X\_\{i\}\)\}\\prod\_\{i\>t\}\\hat\{f\}\_\{1\}\(X\_\{i\}\)^\{w\_\{1\}\(X\_\{i\}\)\}\}\{\\prod\_\{i\\leq\\hat\{\\xi\}\(\\mathbf\{X\}\)\}\\hat\{f\}\_\{0\}\(X\_\{i\}\)^\{w\_\{0\}\(X\_\{i\}\)\}\\prod\_\{i\>\\hat\{\\xi\}\(\\mathbf\{X\}\)\}\\hat\{f\}\_\{1\}\(X\_\{i\}\)^\{w\_\{1\}\(X\_\{i\}\)\}\}\\right\)\.\(14\)
In the weighted CPP score \([14](https://arxiv.org/html/2607.26481#S4.E14)\), each likelihood contributionf^j\(Xi\)\\hat\{f\}\_\{j\}\(X\_\{i\}\)is modulated by its weightwj\(Xi\)w\_\{j\}\(X\_\{i\}\), so that observations with small weightwj\(Xi\)w\_\{j\}\(X\_\{i\}\)have reduced influence on the score\. As shown in the proof of Proposition[1](https://arxiv.org/html/2607.26481#Thmproposition1), the weightwj\(x\)w\_\{j\}\(x\)in \([13](https://arxiv.org/html/2607.26481#S4.E13)\) ideally corresponds to the posterior probability that the corrupted sampleXi=\(1−Yi\)X~i\+YiZiX\_\{i\}=\(1\-Y\_\{i\}\)\\widetilde\{X\}\_\{i\}\+Y\_\{i\}Z\_\{i\}in \([2](https://arxiv.org/html/2607.26481#S2.E2)\) is clean, i\.e\., that we haveYi=0Y\_\{i\}=0\. Thus, intuitively, the introduction of the weightwj\(x\)w\_\{j\}\(x\)allows the score \([14](https://arxiv.org/html/2607.26481#S4.E14)\) to selectively gives more relevance to samplesXiX\_\{i\}that are more likely to be clean\. Sec\.[IV\-B](https://arxiv.org/html/2607.26481#S4.SS2)will discuss a practical implementation of the score \([14](https://arxiv.org/html/2607.26481#S4.E14)\) in which the weights\{w0\(⋅\),w1\(⋅\)\}\\\{w\_\{0\}\(\\cdot\),w\_\{1\}\(\\cdot\)\\\}are extracted from uncertainty signals produced by a classifier\. This construction builds on a classifier\-based version of the score \([14](https://arxiv.org/html/2607.26481#S4.E14)\), which is discussed next\.
#### IV\-A3Classifier\-Based CPP Score
The weighted CPP score \([14](https://arxiv.org/html/2607.26481#S4.E14)\) can be implemented using different estimates of the densitiesf0\(x\)f\_\{0\}\(x\), andf1\(x\)\{f\}\_\{1\}\(x\), and of the changepointξ\{\\xi\}\. Following\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\], instead of relying on separate estimatesf0\(x\)f\_\{0\}\(x\)andf1\(x\)f\_\{1\}\(x\)of the clean data distributions, here we assume access to a binary classifier\{p^\(j∣x\)\}j=0,1\\\{\\hat\{p\}\(j\\mid x\)\\\}\_\{j=0,1\}, trained to distinguish between clean pre\-change samples𝐗~∼f0\(x\)\\widetilde\{\\mathbf\{X\}\}\\sim f\_\{0\}\(x\), labeled asj=0j=0, from clean post\-change samples𝐗~∼f1\(x\)\\widetilde\{\\mathbf\{X\}\}\\sim f\_\{1\}\(x\), labeled asj=1j=1\[[18](https://arxiv.org/html/2607.26481#bib.bib44)\]\. Note that the classifier is trained offline using clean labeled data that is independent of the current batchXX\. Given such a classifier, write the classifier log\-posterior ratio
Δi=logp^\(j=0∣Xi\)p^\(j=1∣Xi\)\.\\Delta\_\{i\}=\\log\\frac\{\\hat\{p\}\(j=0\\mid X\_\{i\}\)\}\{\\hat\{p\}\(j=1\\mid X\_\{i\}\)\}\.\(15\)When the labelsj=0j=0andj=1j=1are a priori equally likely, the ratio \([15](https://arxiv.org/html/2607.26481#S4.E15)\) approximates the log\-likelihood ratio asΔi≈log\(f0\(Xi\)/f1\(Xi\)\)\\Delta\_\{i\}\\approx\\log\(f\_\{0\}\(X\_\{i\}\)/f\_\{1\}\(X\_\{i\}\)\)\[[38](https://arxiv.org/html/2607.26481#bib.bib24),[18](https://arxiv.org/html/2607.26481#bib.bib44)\]\.
Using this approximation, the maximum likelihood estimation \(MLE\) of the changepointξ\\xican be expressed as\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\],ξ^\(𝐗\)=argmaxs∑i≤sΔi\\hat\{\\xi\}\(\\mathbf\{X\}\)=\\arg\\max\_\{s\}\\sum\_\{i\\leq s\}\\Delta\_\{i\}\. Under contamination, we analogously define the corresponding weighted log\-likelihood and weighted MLE\. Furthermore, using the same approximation together with the weighted estimate \([17](https://arxiv.org/html/2607.26481#S4.E17)\) in \([14](https://arxiv.org/html/2607.26481#S4.E14)\) and canceling the terms that do not affect the maximization over the reference changepoint yield the following classifier\-based weighted CPP score\.
###### Definition 2\(Classifier\-Based Weighted CPP Score\)\.
Given a classifier\{p^\(j∣x\)\}j=0,1\\\{\\hat\{p\}\(j\\mid x\)\\\}\_\{j=0,1\}with log\-posterior ratioΔi\\Delta\_\{i\}as in \([15](https://arxiv.org/html/2607.26481#S4.E15)\), and given weight functionswj:𝒳→\[0,1\]w\_\{j\}:\\mathcal\{X\}\\to\[0,1\], letw\(t\)\(Xi\)=w0\(Xi\)w^\{\(t\)\}\(X\_\{i\}\)=w\_\{0\}\(X\_\{i\}\)fori≤ti\\leq tandw\(t\)\(Xi\)=w1\(Xi\)w^\{\(t\)\}\(X\_\{i\}\)=w\_\{1\}\(X\_\{i\}\)fori\>ti\>t, and define the corresponding weighted log\-likelihood
Nt\(s\)=∑i≤sw\(t\)\(Xi\)logp^\(j=0∣Xi\)\+∑i\>sw\(t\)\(Xi\)logp^\(j=1∣Xi\)\.N\_\{t\}\(s\)=\\sum\_\{i\\leq s\}w^\{\(t\)\}\(X\_\{i\}\)\\log\\hat\{p\}\(j=0\\mid X\_\{i\}\)\+\\sum\_\{i\>s\}w^\{\(t\)\}\(X\_\{i\}\)\\log\\hat\{p\}\(j=1\\mid X\_\{i\}\)\.\(16\)The corresponding weighted MLE is
ξ^t\(𝐗\)=argmaxs∈\{1,…n−1\}Nt\(s\)\.\\hat\{\\xi\}\_\{t\}\(\\mathbf\{X\}\)=\\arg\\max\_\{s\\in\\\{1,\\ldots n\-1\\\}\}N\_\{t\}\(s\)\.\(17\)The classifier\-based weighted CPP score is defined as
Stw\(𝐗\)=∑i≤tw0\(Xi\)Δi−∑i≤ξ^t\(𝐗\)w\(t\)\(Xi\)Δi\.S\_\{t\}^\{w\}\(\\mathbf\{X\}\)=\\sum\_\{i\\leq t\}w\_\{0\}\(X\_\{i\}\)\\Delta\_\{i\}\-\\sum\_\{i\\leq\\hat\{\\xi\}\_\{t\}\(\\mathbf\{X\}\)\}w^\{\(t\)\}\(X\_\{i\}\)\\Delta\_\{i\}\.\(18\)
The derivation of the weighted CPP score \([18](https://arxiv.org/html/2607.26481#S4.E18)\) is provided in Appendix[A\-C](https://arxiv.org/html/2607.26481#A1.SS3)\. The first term in \([18](https://arxiv.org/html/2607.26481#S4.E18)\) represents the cumulative weighted log\-posterior ratio up to the candidate changepointtt\. The second term represents the cumulative weighted log\-posterior ratio evaluated at the weighted MLEξ^t\(𝐗\)\\hat\{\\xi\}\_\{t\}\(\\mathbf\{X\}\), with the latter corresponding to the most plausible competing changepoint\. Thus, the CPPStw\(𝐗\)S\_\{t\}^\{w\}\(\\mathbf\{X\}\)measures the relative plausibility of candidatettagainst the best competing changepoint\.
### IV\-BUncertainty\-Based Weights
As discussed in Sec\.[IV\-A](https://arxiv.org/html/2607.26481#S4.SS1), the weightwj\(x\)w\_\{j\}\(x\)used in the CPP score \([14](https://arxiv.org/html/2607.26481#S4.E14)\) or \([18](https://arxiv.org/html/2607.26481#S4.E18)\) should ideally capture the true probability that a sampleX∼gj\(x\)X\\sim g\_\{j\}\(x\)is clean\. To approximate this probability, which depends on the unknown distribution \([13](https://arxiv.org/html/2607.26481#S4.E13)\), we propose to use classification uncertainty as a proxy\. The rationale for this choice is that a classifier modelp^\(j∣x\)\\hat\{p\}\(j\\mid x\)trained on clean data should ideally be confident on clean in\-distribution samples and uncertain on contaminated ones\[[16](https://arxiv.org/html/2607.26481#bib.bib38),[26](https://arxiv.org/html/2607.26481#bib.bib64)\]\. With this approach, low\-uncertainty observationsXiX\_\{i\}receive larger weights, while high\-uncertainty observations receive smaller weights\.
To elaborate, as in Sec\.[IV\-A3](https://arxiv.org/html/2607.26481#S4.SS1.SSS3), consider a classifierp^\(j∣x\)\\hat\{p\}\(j\\mid x\)pre\-trained to distinguish between clean samples from the pre\-change densityf0\(x\)f\_\{0\}\(x\), labeled withj=0j=0and the post\-change densityf1\(x\)f\_\{1\}\(x\), labeled asj=1j=1\. In addition to the class posterior scores\{p^\(j∣x\)\}j=0,1\\\{\\hat\{p\}\(j\\mid x\)\\\}\_\{j=0,1\}, which are used to construct the classifier\-based CPP score \([18](https://arxiv.org/html/2607.26481#S4.E18)\), we assume that the classifier provides an uncertainty signal
Mi=h\(Xi\),M\_\{i\}=h\(X\_\{i\}\),\(19\)for some functionh:𝒳→ℝ≥0h:\\mathcal\{X\}\\to\\mathbb\{R\}\_\{\\geq 0\}, where larger values ofMiM\_\{i\}indicate higher classification uncertainty\. For example, the functionh\(⋅\)h\(\\cdot\)can be obtained from an evidential classifier, which uses the parameters of a predictive Beta distribution to quantify second\-order uncertainty\[[35](https://arxiv.org/html/2607.26481#bib.bib20)\], i\.e\., uncertainty about the predictive distribution, or using Bayesian methods, such as MC dropout\[[9](https://arxiv.org/html/2607.26481#bib.bib32)\], which estimate such uncertainty via ensembling\[[23](https://arxiv.org/html/2607.26481#bib.bib39),[38](https://arxiv.org/html/2607.26481#bib.bib24)\]\.
A natural choice is to assign weight11to observations below an uncertainty thresholdκj\\kappa\_\{j\}and weight0to those above, yielding the hard thresholding rule
wjhard\(Xi\)=\{1,Mi≤κj,0,otherwisew\_\{j\}^\{\\mathrm\{hard\}\}\(X\_\{i\}\)=\\begin\{cases\}1,&M\_\{i\}\\leq\\kappa\_\{j\},\\\\ 0,&\\text\{otherwise\}\\end\{cases\}\(20\)forj=0,1j=0,1\. A smoother alternative replaces the hard threshold with a sigmoid function, yielding the soft weighting rule
wjsoft\(Xi\)=σ\(κj−Miλ\),λ\>0,w\_\{j\}^\{\\mathrm\{soft\}\}\(X\_\{i\}\)=\\sigma\\left\(\\frac\{\\kappa\_\{j\}\-M\_\{i\}\}\{\\lambda\}\\right\),\\quad\\lambda\>0,\(21\)whereσ\(⋅\)\\sigma\(\\cdot\)is the sigmoid function and hyperparameterλ\\lambdacontrols the sharpness of the transition\. Asλ→0\\lambda\\to 0, the soft rule recovers the hard thresholding rule in \([20](https://arxiv.org/html/2607.26481#S4.E20)\)\.
The thresholdsκj\\kappa\_\{j\}forj=0,1j=0,1in \([20](https://arxiv.org/html/2607.26481#S4.E20)\) should ideally capture the transition between typical uncertainty signals for clean samples and for corrupted samples on either side of the changepoint\. Accordingly, we propose to use side\-specific empirical quantiles of the uncertainty signals\. Formally, for a candidate changepointttand a quantile hyperparameterβ∈\[0,1\]\\beta\\in\[0,1\], we set
κ0\(t,β\)\\displaystyle\\kappa\_\{0\}\(t,\\beta\)=Q1−β\(\{M1,…,Mt\}\),\\displaystyle=\\mathrm\{Q\}\_\{1\-\\beta\}\\bigl\(\\\{M\_\{1\},\\dots,M\_\{t\}\\\}\\bigr\),\(22a\)andκ1\(t,β\)\\displaystyle\\text\{and \}\\kappa\_\{1\}\(t,\\beta\)=Q1−β\(\{Mt\+1,…,Mn\}\),\\displaystyle=\\mathrm\{Q\}\_\{1\-\\beta\}\\bigl\(\\\{M\_\{t\+1\},\\dots,M\_\{n\}\\\}\\bigr\),\(22b\)whereQ1−β\(𝒮\)Q\_\{1\-\\beta\}\(\\mathcal\{S\}\)returns the⌊\(1−β\)\|𝒮\|⌋\\lfloor\(1\-\\beta\)\|\\mathcal\{S\}\|\\rfloorsmallest element in set𝒮\\mathcal\{S\}\. The hyperparameterβ\\betacontrols should ideally match the expected fractionε\\varepsilonof contaminated observations\. Based on this, we recommend setting hyperparameterβ\\betato one’s best conservative estimate of the contamination probabilityε\\varepsilon\.
### IV\-CCoverage Properties of W\-CONCH
Overall, specializing the CONCH set \([9](https://arxiv.org/html/2607.26481#S3.E9)\), the W\-CONCH set is defined as follows\.
###### Definition 3\(W\-CONCH Confidence Set\)\.
Given hyperparametersλ\>0\\lambda\>0andβ∈\[0,1\]\\beta\\in\[0,1\], which control the sharpness of the soft weighting rule \([21](https://arxiv.org/html/2607.26481#S4.E21)\) and the quantile level for the side\-specific uncertainty thresholds \([22](https://arxiv.org/html/2607.26481#S4.E22)\), respectively, W\-CONCH scheme produces the confidence set
𝒞αw\(𝐗\)\\displaystyle\\mathcal\{C\}\_\{\\alpha\}^\{w\}\(\\mathbf\{X\}\)=\{t∈\{1,…,n−1\}:ptw\(𝐗\)\>α\},\\displaystyle=\\bigl\\\{t\\in\\\{1,\\ldots,n\-1\\\}:p\_\{t\}^\{w\}\(\\mathbf\{X\}\)\>\\alpha\\bigr\\\},\(23\)withptw\(𝐗\)\\displaystyle\{\\textrm\{with\}\}~p\_\{t\}^\{w\}\(\\mathbf\{X\}\)=1\|Πt\|∑πt∈Πt𝟙\[Stw\(πt\(𝐗\)\)≤Stw\(𝐗\)\],\\displaystyle=\\frac\{1\}\{\|\\Pi\_\{t\}\|\}\\sum\_\{\\pi\_\{t\}\\in\\Pi\_\{t\}\}\\mathds\{1\}\\bigl\[S\_\{t\}^\{w\}\(\\pi\_\{t\}\(\\mathbf\{X\}\)\)\\leq S\_\{t\}^\{w\}\(\\mathbf\{X\}\)\\bigr\],where the weighted scoreStw\(𝐗\)S\_\{t\}^\{w\}\(\\mathbf\{X\}\)in \([18](https://arxiv.org/html/2607.26481#S4.E18)\) is implemented using the weights defined in \([21](https://arxiv.org/html/2607.26481#S4.E21)\)–\([22](https://arxiv.org/html/2607.26481#S4.E22)\)\.
The following proposition summarizes the properties of the W\-CONCH set \([23](https://arxiv.org/html/2607.26481#S4.E23)\)\.
###### Proposition 2\(Marginal coverage of W\-CONCH\)\.
Under Assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1), for any choice of hyperparameterλ\\lambdaandβ\\beta, the W\-CONCH set𝒞αw\(𝐗\)\\mathcal\{C\}\_\{\\alpha\}^\{w\}\(\\mathbf\{X\}\)in \([23](https://arxiv.org/html/2607.26481#S4.E23)\) satisfies the marginal coverage condition
Pr\(ξ∈𝒞αw\(𝐗\)\)≥1−α\.\\Pr\\bigl\(\\xi\\in\\mathcal\{C\}\_\{\\alpha\}^\{w\}\(\\mathbf\{X\}\)\\bigr\)\\geq 1\-\\alpha\.\(24\)
###### Proof\.
As detailed in Appendix[A\-D](https://arxiv.org/html/2607.26481#A1.SS4), the result follows from the split\-permutation invariance of the weighted score \([18](https://arxiv.org/html/2607.26481#S4.E18)\) and the finite\-sample rank argument of\[[14](https://arxiv.org/html/2607.26481#bib.bib2), Thm\. 3\.1\]\. ∎
### IV\-DMeta\-Learned Uncertainty Weights
The W\-CONCH set \([23](https://arxiv.org/html/2607.26481#S4.E23)\) in Sec\.[IV\-B](https://arxiv.org/html/2607.26481#S4.SS2)requires choosing the quantile hyperparameterβ\\beta, as well as the threshold hyperparameterλ\\lambda\. As discussed in the previous subsection, when the contamination levelε\\varepsilonis known, the hyperparameterβ\\betacan be naturally set toε\\varepsilon, but a reasonable estimate of the probabilityε\\varepsilonmay not be available\. Moreover, a pretrained uncertainty estimatorh\(x\)h\(x\)may not necessarily produce scoresMiM\_\{i\}in \([19](https://arxiv.org/html/2607.26481#S4.E19)\) that are well aligned with the objective of minimizing the average set size\|𝒞αw\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}^\{w\}\(\\mathbf\{X\}\)\|\. We address both issues by introducing a meta\-learning strategy that optimizes a task\-adapted uncertainty estimator from a collection of contaminated changepoint tasks, so that the induced weights directly minimize the size objective\|𝒞αw\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}^\{w\}\(\\mathbf\{X\}\)\|\. We refer to this meta\-learned variant of W\-CONCH as MW\-CONCH\. Fig\.[3](https://arxiv.org/html/2607.26481#S4.F3)illustrates the overall pipeline of MW\-CONCH\.
Figure 3:Operation of MW\-CONCH: The uncertainty modelh\(⋅;θ\)h\(\\cdot;\\theta\)produces per\-observation uncertainty scoresMiM\_\{i\}, which are converted to weights via the smooth weighting rule\. The weighted CPP score \([14](https://arxiv.org/html/2607.26481#S4.E14)\) and the permutation procedure of Sec\.[III](https://arxiv.org/html/2607.26481#S3)together yield a smoothedpp\-value, from which a differentiable surrogate of the confidence set size is computed as the training loss\. Estimated pre\- and post\-change densitiesf^j\\hat\{f\}\_\{j\}enter the CPP score independently ofθ\\theta\.To start, MW\-CONCH selects a parameterized uncertainty estimatorh\(⋅;θ\):𝒳→ℝ≥0h\(\\cdot;\\theta\):\\mathcal\{X\}\\to\\mathbb\{R\}\_\{\\geq 0\}with trainable parametersθ\\theta, replacing the pretrained modelhhin Sec\.[IV\-B](https://arxiv.org/html/2607.26481#S4.SS2)\. For example, the modelh\(⋅;θ\)h\(\\cdot;\\theta\)may be a neural network trained via EDL\[[35](https://arxiv.org/html/2607.26481#bib.bib20)\]\. EDL places a Dirichlet priorDir\(𝒆\)\\text\{Dir\}\(\\boldsymbol\{e\}\)over class probabilities and an uncertainty scoreh\(Xi;θ\)=J/∑j=1Jejh\(X\_\{i\};\\theta\)=J/\\sum\_\{j=1\}^\{J\}e\_\{j\}, whereJJis the number of classes and𝒆=g\(Xi;θ\)\\boldsymbol\{e\}=g\(X\_\{i\};\\theta\)are the predicted concentration parameters\. The goal is to optimize the parameterθ\\thetaso as to minimize the average size of the confidence set\|𝒞αw\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}^\{w\}\(\\mathbf\{X\}\)\|\.
To this end, we fix the distributionsf0\(⋅\)f\_\{0\}\(\\cdot\)andf1\(⋅\)f\_\{1\}\(\\cdot\)of the clean data, and we assume to have access to data𝐗k=\(X1k,…,Xnk\)\\mathbf\{X\}^\{k\}=\(X\_\{1\}^\{k\},\\ldots,X\_\{n\}^\{k\}\)generated i\.i\.d\. from some distribution satisfying Assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1)and following the contamination model \([2](https://arxiv.org/html/2607.26481#S2.E2)\) for some randomly drawn contamination levelεtr∼μ\\varepsilon\_\{\\mathrm\{tr\}\}\\sim\\mufrom some distributionμ\\mu\. The training distributionμ\\muis itself a design choice: it may concentrate on a single reference contamination levelεtr\\varepsilon\_\{\\mathrm\{tr\}\}, which is referred to as fixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}meta\-learning, or it may spread mass over multiple levels, which is referred to as mixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}meta\-learning\. The meta\-training loss is the average confidence set size of the W\-CONCH set \([23](https://arxiv.org/html/2607.26481#S4.E23)\) over the samples\{𝐗k\}k=1K\\\{\\mathbf\{X\}^\{k\}\\\}\_\{k=1\}^\{K\}:
ℒ\(θ\)=1K∑k=1K\|𝒞α;θw\(𝐗k\)\|,\\mathcal\{L\}\(\\theta\)=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\left\|\\mathcal\{C\}\_\{\\alpha;\\theta\}^\{w\}\(\\mathbf\{X\}^\{k\}\)\\right\|,\(25\)where we have made explicit the dependence of the W\-CONCH set𝒞α;θw\(𝐗\)\\mathcal\{C\}\_\{\\alpha;\\theta\}^\{w\}\(\\mathbf\{X\}\)on the parametric uncertainty scoreh\(⋅;θ\)h\(\\cdot;\\theta\)\.
The meta\-learning loss in \([25](https://arxiv.org/html/2607.26481#S4.E25)\) is not directly differentiable and a smooth approximation is derived in Appendix[A\-F](https://arxiv.org/html/2607.26481#A1.SS6)by following an approach similar to\[[28](https://arxiv.org/html/2607.26481#bib.bib22),[43](https://arxiv.org/html/2607.26481#bib.bib3)\]\. MW\-CONCH optimizes this differentiable objective via gradient descent\.
### IV\-EIncorporating Prior Information on the Changepoint Location
When prior information is available on which candidate locations are more likely to be the true changepointξ\\xi, it can be used to further concentrate the confidence set\. We encode the prior as weightsv1,…,vn−1≥0v\_\{1\},\\dots,v\_\{n\-1\}\\geq 0over the candidate locations\. The weights are normalized as∑t=1n−1vt=n−1\\sum\_\{t=1\}^\{n\-1\}v\_\{t\}=n\-1, so that valuesvt\>1v\_\{t\}\>1correspond to locations deemed more likely than under a uniform prior\. Unlike the weightsw0\(Xi\),w1\(Xi\)w\_\{0\}\(X\_\{i\}\),w\_\{1\}\(X\_\{i\}\)in \([14](https://arxiv.org/html/2607.26481#S4.E14)\), which act on the observations, the prior weights\{vt\}t=1n−1\\\{v\_\{t\}\\\}\_\{t=1\}^\{n\-1\}act on the candidate locations, entering through the significance level used to test each candidate\.
Specifically, following the level\-allocation strategy of weighted hypothesis testing\[[10](https://arxiv.org/html/2607.26481#bib.bib66)\], each candidatettis tested at the level
αt=min\(αvt,αmax\),\\alpha\_\{t\}=\\min\\left\(\\frac\{\\alpha\}\{v\_\{t\}\},\\,\\alpha\_\{\\max\}\\right\),\(26\)whereαmax∈\[α,1\)\\alpha\_\{\\max\}\\in\[\\alpha,1\)bounds the level applied to any single location, yielding the prior\-informed confidence set
𝒞α,vw\(𝐗\)=\{t∈\{1,…,n−1\}:ptw\(𝐗\)\>αt\},\\mathcal\{C\}\_\{\\alpha,v\}^\{w\}\(\\mathbf\{X\}\)=\\bigl\\\{t\\in\\\{1,\\dots,n\-1\\\}:p\_\{t\}^\{w\}\(\\mathbf\{X\}\)\>\\alpha\_\{t\}\\bigr\\\},\(27\)with thepp\-valueptw\(𝐗\)p\_\{t\}^\{w\}\(\\mathbf\{X\}\)computed as in \([23](https://arxiv.org/html/2607.26481#S4.E23)\)\. By \([26](https://arxiv.org/html/2607.26481#S4.E26)\), a priori likely locations are tested at smaller levels, and are thus harder to exclude, while unlikely locations are more easily discarded\. Under a uniform prior, i\.e\.,vt=1v\_\{t\}=1for alltt, all levels reduce toαt=α\\alpha\_\{t\}=\\alpha, recovering the W\-CONCH set \([23](https://arxiv.org/html/2607.26481#S4.E23)\)\.
###### Proposition 3\(Coverage of prior\-informed W\-CONCH\)\.
Under Assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1), the set𝒞α,vw\(𝐗\)\\mathcal\{C\}\_\{\\alpha,v\}^\{w\}\(\\mathbf\{X\}\)in \([27](https://arxiv.org/html/2607.26481#S4.E27)\) satisfies
Pr\(ξ∈𝒞α,vw\(𝐗\)\)≥1−αmax\\Pr\\bigl\(\\xi\\in\\mathcal\{C\}\_\{\\alpha,v\}^\{w\}\(\\mathbf\{X\}\)\\bigr\)\\geq 1\-\\alpha\_\{\\max\}\(28\)for any changepointξ\\xi\.
###### Proof\.
See Appendix[A\-E](https://arxiv.org/html/2607.26481#A1.SS5)\. ∎
The parameterαmax\\alpha\_\{\\max\}controls the influence of the prior\. The worst\-case bound1−αmax1\-\\alpha\_\{\\max\}in \([28](https://arxiv.org/html/2607.26481#S4.E28)\) protects against misspecification of the prior, with the choiceαmax=α\\alpha\_\{\\max\}=\\alpharecovering the guarantee \([24](https://arxiv.org/html/2607.26481#S4.E24)\) of Proposition[2](https://arxiv.org/html/2607.26481#Thmproposition2)\.
## VWeighted Conformal Root\-Cause Localization with Contaminated Observations
In this section, we introduce an extension of W\-CONCH to the multi\-stream root\-cause localization setting under contaminated observations described in Sec\.[II\-B](https://arxiv.org/html/2607.26481#S2.SS2)\. The corresponding framework, referred to as W\-CROC, constructs a confidence set𝒦α\(𝐗\)\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)for the root\-cause index by applying the CROC scheme presented in\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\]with a weighted CPP score analogous to W\-CONCH\.
### V\-AConformal Root Cause Analysis
Unlike the single\-stream setting studied in the previous section, the unknown changepoint is now a vector𝝃=\(ξ1,…,ξD\)\\boldsymbol\{\\xi\}=\(\\xi\_\{1\},\\ldots,\\xi\_\{D\}\), and the root\-cause indexd⋆d^\{\\star\}in \([7](https://arxiv.org/html/2607.26481#S2.E7)\) is determined by the stream with the earliest changepoint\. Let𝐭=\(t1,…,tD\)\\mathbf\{t\}=\(t\_\{1\},\\ldots,t\_\{D\}\)be a candidate changepoint configuration in a feasible setℛ\\mathcal\{R\}\. The setℛ\\mathcal\{R\}encodes prior structural constraints on the changepoints that may be a priori known to the detector\. For instance, setℛ\\mathcal\{R\}may consist of configurations in which a single stream changes first, i\.e\.,ℛ=\{\(t1,…,tD\):∃d∈\{1,…,D\}such thattd<tm,∀m≠d\}\\mathcal\{R\}=\\\{\(t\_\{1\},\\ldots,t\_\{D\}\):\\exists d\\in\\\{1,\\ldots,D\\\}\\ \\text\{such that\}\\ t\_\{d\}<t\_\{m\},\\ \\forall m\\neq d\\\}\. For each candidate configuration𝐭∈ℛ\\mathbf\{t\}\\in\\mathcal\{R\}, CROC evaluates a CPP scoreS𝐭\(𝐗\)S\_\{\\mathbf\{t\}\}\(\\mathbf\{X\}\)that measures the plausibility of𝐭\\mathbf\{t\}being the true changepoint configuration\.
To calibrate this score, CROC applies split permutations within each streamddusing the candidate changepointtdt\_\{d\}\. LetΠ𝐭\\Pi\_\{\\mathbf\{t\}\}denote the corresponding split\-permutation group, where each streamddis permuted only within its candidate pre\- and post\-change segments determined by candidatetdt\_\{d\}\. The configuration\-level conformalpp\-value is defined as
p𝐭=1\|Π𝐭\|∑π∈Π𝐭𝟏\{S𝐭\(π\(𝐗\)\)≤S𝐭\(𝐗\)\}\.p\_\{\\mathbf\{t\}\}=\\frac\{1\}\{\|\\Pi\_\{\\mathbf\{t\}\}\|\}\\sum\_\{\\pi\\in\\Pi\_\{\\mathbf\{t\}\}\}\\mathbf\{1\}\\left\\\{S\_\{\\mathbf\{t\}\}\(\\pi\(\\mathbf\{X\}\)\)\\leq S\_\{\\mathbf\{t\}\}\(\\mathbf\{X\}\)\\right\\\}\.\(29\)In practice, the average in \([29](https://arxiv.org/html/2607.26481#S5.E29)\) can be approximated using MC split permutations as explained in Sec\.[III\-A](https://arxiv.org/html/2607.26481#S3.SS1)\.
To obtain a confidence set for the root\-cause index, CROC aggregates the configuration\-levelpp\-values over all configurations that are consistent with each candidate root stream\. To elaborate, for eachd∈\{1,…,D\}d\\in\\\{1,\\ldots,D\\\}, define the setℐd=\{𝐭∈ℛ:td<tmfor allm≠d\},\\mathcal\{I\}\_\{d\}=\\\{\\mathbf\{t\}\\in\\mathcal\{R\}:t\_\{d\}<t\_\{m\}\\text\{ for all \}m\\neq d\\\},which contains all feasible configurations under which streamddis the root\-cause stream\. The root\-causepp\-value for streamddis then computed by taking the maximum over all feasible configurations in which streamddchanges first across all other streams, i\.e\.,
p\(d\)=max𝐭∈ℐdp𝐭\.p\_\{\(d\)\}=\\max\_\{\\mathbf\{t\}\\in\\mathcal\{I\}\_\{d\}\}p\_\{\\mathbf\{t\}\}\.\(30\)Then, streamddis included in the confidence set𝒦α\(𝐗\)\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)if at least one feasible configuration in which streamddchanges first is not rejected\. Formally, given a miscoverage levelα∈\(0,1\)\\alpha\\in\(0,1\), the CROC confidence set for the root\-cause index is
𝒦α\(𝐗\)=\{d∈\{1,…,D\}:p\(d\)\>α\}\.\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)=\\\{d\\in\\\{1,\\ldots,D\\\}:p\_\{\(d\)\}\>\\alpha\\\}\.\(31\)
### V\-BWeighted Conformal Root Cause Analysis
W\-CROC specializes CROC by replacing the CPP score with the weighted classifier\-based score developed in Sec\.[IV](https://arxiv.org/html/2607.26481#S4), in order to account for the contaminated observations in \([6](https://arxiv.org/html/2607.26481#S2.E6)\)\. Assume access to a pre\-trained classifier for pre\-change and post\-change samples for each streamdd, so that a log\-posterior ratioΔd,i\\Delta\_\{d,i\}, as in \([15](https://arxiv.org/html/2607.26481#S4.E15)\), can be computed for each streamdd\.
###### Definition 4\(Weighted CROC Score\)\.
Given existing weightswd,j:𝒳→\[0,1\]w\_\{d,j\}:\\mathcal\{X\}\\to\[0,1\]ford=1,…,Dd=1,\\ldots,D, the configuration level weighted CPP score for a candidate changepoint configuration𝐭=\(t1,…,tD\)\\mathbf\{t\}=\(t\_\{1\},\\ldots,t\_\{D\}\)∈ℛ\\in\\mathcal\{R\}is defined as
S𝐭w\(𝐗\)=∑d=1DStdw\(𝐗d\),S\_\{\\mathbf\{t\}\}^\{w\}\(\\mathbf\{X\}\)=\\sum\_\{d=1\}^\{D\}S\_\{t\_\{d\}\}^\{w\}\(\\mathbf\{X\}\_\{d\}\),\(32\)whereStdw\(𝐗d\)S\_\{t\_\{d\}\}^\{w\}\(\\mathbf\{X\}\_\{d\}\)is the single\-stream weighted CPP score \([18](https://arxiv.org/html/2607.26481#S4.E18)\) applied to streamdd, which is given by
Stdw\(𝐗d\)=∑i≤tdwd,0\(Xd,i\)Δd,i−∑i≤ξ^td\(𝐗d\)wd\(td\)\(Xd,i\)Δd,i,S\_\{t\_\{d\}\}^\{w\}\(\\mathbf\{X\}\_\{d\}\)=\\sum\_\{i\\leq t\_\{d\}\}w\_\{d,0\}\(X\_\{d,i\}\)\\Delta\_\{d,i\}\-\\sum\_\{i\\leq\\hat\{\\xi\}\_\{t\_\{d\}\}\(\\mathbf\{X\}\_\{d\}\)\}w\_\{d\}^\{\(t\_\{d\}\)\}\(X\_\{d,i\}\)\\Delta\_\{d,i\},\(33\)withwd\(td\)\(Xd,i\)=wd,0\(Xd,i\)w\_\{d\}^\{\(t\_\{d\}\)\}\(X\_\{d,i\}\)=w\_\{d,0\}\(X\_\{d,i\}\)fori≤tdi\\leq t\_\{d\}andwd\(td\)\(Xd,i\)=wd,1\(Xd,i\)w\_\{d\}^\{\(t\_\{d\}\)\}\(X\_\{d,i\}\)=w\_\{d,1\}\(X\_\{d,i\}\)fori\>tdi\>t\_\{d\}, andξ^td\(𝐗d\)\\hat\{\\xi\}\_\{t\_\{d\}\}\(\\mathbf\{X\}\_\{d\}\)defined as in \([17](https://arxiv.org/html/2607.26481#S4.E17)\), applied to streamdd\.
The weights in Definition[4](https://arxiv.org/html/2607.26481#Thmdefinition4)can be constructed using the hard or soft uncertainty\-based weighting rules introduced in Sec\.[IV\-B](https://arxiv.org/html/2607.26481#S4.SS2), or learned through the meta\-learning procedure described in Sec\.[IV\-D](https://arxiv.org/html/2607.26481#S4.SS4)\. When the uncertainty model is optimized via the meta\-learning procedure, we refer to the resulting method as meta\-learned weighted CROC \(MW\-CROC\)\. The corresponding differentiable training objective and optimization procedure are summarized in Algorithm[1](https://arxiv.org/html/2607.26481#alg1)in Appendix[A\-F](https://arxiv.org/html/2607.26481#A1.SS6)\.
###### Proposition 4\(Coverage of W\-CROC\)\.
Under Assumption[2](https://arxiv.org/html/2607.26481#Thmassumption2), the W\-CROC confidence set𝒦α\(𝐗\)\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)in \([31](https://arxiv.org/html/2607.26481#S5.E31)\), computed with the weighted score \([32](https://arxiv.org/html/2607.26481#S5.E32)\), satisfies the marginal coverage condition
Pr\(d⋆∈𝒦α\(𝐗\)\)≥1−α\.\\Pr\\bigl\(d^\{\\star\}\\in\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\\bigr\)\\geq 1\-\\alpha\.\(34\)
###### Proof\.
The weighted score \([32](https://arxiv.org/html/2607.26481#S5.E32)\) is permutation\-equivariant under the split\-permutation groupΠ𝐭\\Pi\_\{\\mathbf\{t\}\}, since each stream’s weights depend on𝐗d\\mathbf\{X\}\_\{d\}only through stream\-wise thresholds that are invariant under within\-stream permutations, exactly as in the proof of Proposition[2](https://arxiv.org/html/2607.26481#Thmproposition2)\. The coverage guarantee therefore follows directly from the validity proof of CROC\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\], since the weighted score preserves the required permutation\-equivariance\. ∎
### V\-CIncorporating Prior Information on the Root Cause
When prior information is available on which streams are more likely to be the root caused⋆d^\{\\star\}, it can be incorporated via the level\-allocation strategy described in Sec\.[IV\-E](https://arxiv.org/html/2607.26481#S4.SS5)\. The prior is encoded as weightsv1,…,vD≥0v\_\{1\},\\dots,v\_\{D\}\\geq 0over the candidate streams, which are normalized as∑d=1Dvd=D\\sum\_\{d=1\}^\{D\}v\_\{d\}=D, so that valuesvd\>1v\_\{d\}\>1correspond to streams deemed more likely than under a uniform prior\.
As in \([26](https://arxiv.org/html/2607.26481#S4.E26)\), each streamddis tested at the level
αd=min\(αvd,αmax\),\\alpha\_\{d\}=\\min\\left\(\\frac\{\\alpha\}\{v\_\{d\}\},\\,\\alpha\_\{\\max\}\\right\),\(35\)withαmax∈\[α,1\)\\alpha\_\{\\max\}\\in\[\\alpha,1\), yielding the prior\-informed root\-cause confidence set
𝒦α,v\(𝐗\)=\{d∈\{1,…,D\}:p\(d\)\>αd\},\\mathcal\{K\}\_\{\\alpha,v\}\(\\mathbf\{X\}\)=\\\{d\\in\\\{1,\\dots,D\\\}:p\_\{\(d\)\}\>\\alpha\_\{d\}\\\},\(36\)where the root\-causepp\-valuep\(d\)p\_\{\(d\)\}in \([30](https://arxiv.org/html/2607.26481#S5.E30)\) is computed with the weighted score \([32](https://arxiv.org/html/2607.26481#S5.E32)\)\.
By \([35](https://arxiv.org/html/2607.26481#S5.E35)\), a priori likely streams are tested at smaller levels, and are thus harder to exclude, while unlikely streams are more easily discarded\. Under a uniform prior, i\.e\.,vd=1v\_\{d\}=1for alldd, all levels reduce toαd=α\\alpha\_\{d\}=\\alpha, recovering the W\-CROC set \([31](https://arxiv.org/html/2607.26481#S5.E31)\), and the role ofαmax\\alpha\_\{\\max\}is as discussed in Sec\.[IV\-E](https://arxiv.org/html/2607.26481#S4.SS5)\.
###### Proposition 5\(Coverage of prior\-informed W\-CROC\)\.
Under Assumption[2](https://arxiv.org/html/2607.26481#Thmassumption2), the set𝒦α,v\(𝐗\)\\mathcal\{K\}\_\{\\alpha,v\}\(\\mathbf\{X\}\)in \([36](https://arxiv.org/html/2607.26481#S5.E36)\) satisfies
Pr\(d⋆∈𝒦α,v\(𝐗\)\)≥1−αmax\\Pr\\bigl\(d^\{\\star\}\\in\\mathcal\{K\}\_\{\\alpha,v\}\(\\mathbf\{X\}\)\\bigr\)\\geq 1\-\\alpha\_\{\\max\}\(37\)for any root caused⋆d^\{\\star\}\.
###### Proof\.
By the proof of Proposition[4](https://arxiv.org/html/2607.26481#Thmproposition4), underd⋆=dd^\{\\star\}=dthepp\-valuep\(d\)p\_\{\(d\)\}is super\-uniform, i\.e\.,Pr\{p\(d\)≤c\}≤c\\Pr\\\{p\_\{\(d\)\}\\leq c\\\}\\leq cfor any constantc∈\[0,1\]c\\in\[0,1\]\. Both bounds then follow as in Appendix[A\-E](https://arxiv.org/html/2607.26481#A1.SS5), withtt,ξ\\xi, andn−1n\-1replaced bydd,d⋆d^\{\\star\}, andDD\. ∎
## VIExperiments
Figure 4:W\-CONCH on DomainNet \(left panels\) and CIFAR\-100 \(right panels\)\. \(a\): Average confidence set size under hard uncertainty thresholding \([20](https://arxiv.org/html/2607.26481#S4.E20)\) for different quantile hyperparametersβ\\beta\. The results support the heuristic that hyperparameterβ\\betashould be chosen matched to the contamination levelε\\varepsilon\. \(b\): W\-CONCH with soft weighting \([21](https://arxiv.org/html/2607.26481#S4.E21)\) withβ=ε\\beta=\\varepsilonand different temperaturesλ\\lambda, compared with hard W\-CONCH and CONCH\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]\.We evaluate W\-CONCH and W\-CROC on image\-based and real\-world changepoint and root\-cause localization benchmarks\. For changepoint localization, we use DomainNet \(real\-to\-sketch domain shift\), a CIFAR\-100 benchmark \(a shift between image classes\), and the Milan telecom dataset\[[3](https://arxiv.org/html/2607.26481#bib.bib65)\]\(a real shift in daily cellular activity at the holiday onset\)\. For root\-cause localization, we construct a multi\-stream benchmark based on CIFAR\-100, in which each stream undergoes a class shift at a stream\-specific changepoint\. In all settings, contamination is introduced by perturbing a random subset of observations according to the contamination model \([2](https://arxiv.org/html/2607.26481#S2.E2)\)\. Full implementation details, additional MNIST experiments, and complete numerical results are provided in Appendix[B](https://arxiv.org/html/2607.26481#A2)\.
### VI\-AExperimental Setup
#### VI\-A1Datasets and Evaluation
Changepoint localization\.For DomainNet, sequences have lengthn=800n=800with changepointξ=350\\xi=350, where the firstξ\\xiobservations are from the real domain and the rest from the sketch domain\. For CIFAR\-100, sequences have lengthn=400n=400withξ=250\\xi=250, shifting from class9999\(worm\) to class7777\(snail\)\. Observations are independently contaminated with probabilityε\\varepsilonby Gaussian blur withσ=20\\sigma=20pixels for DomainNet and Gaussian noise withσ=0\.3\\sigma=0\.3pixels for CIFAR\-100\.
In the Milan telecom dataset, each sequence collects the daily activity profiles of one grid cell from call detail records recorded over Milan between November 2013 and January 2014, where each observation is a720720\-dimensional daily profile of144144ten\-minute intervals across five call detail record measurements\. Each sequence hasn=63n=63samples with changepoint atξ=54\\xi=54, corresponding to December 24, 2013\. Observations are contaminated with probabilityε\\varepsilonby Gaussian noise withσ=20\\sigma=20on standardized features\. Root\-cause localization\.For CIFAR\-100, we build a multi\-stream benchmark withD=5D=5streams, each shifting between two classes within a common CIFAR\-100 superclass \(e\.g\., cattle to chimpanzee\)\. Stream 1 is the root\-cause stream and changes att1=150t\_\{1\}=150, while the remaining streams change attd=155t\_\{d\}=155ford=2,…,5d=2,\\ldots,5\. Each stream containsn=400n=400observations\. The feasible setℛ\\mathcal\{R\}contains one configuration per candidate root stream, assigningtd=150t\_\{d\}=150to that stream andtm=155t\_\{m\}=155to all others\. Observations are contaminated as in the changepoint localization setup\. Train and test pools\.For all datasets, training and evaluation tasks are drawn from disjoint data pools, so that meta\-learning never sees observations used at evaluation\. Evaluation\.We vary the contamination level overε∈\{0\.0,0\.1,0\.3,0\.5,0\.7\}\\varepsilon\\in\\\{0\.0,0\.1,0\.3,0\.5,0\.7\\\}\. For changepoint localization, we report the average confidence set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|and empirical coverage over200200test tasks atα=0\.05\\alpha=0\.05, with400400split permutations perpp\-value\. For root\-cause localization, we evaluate𝒦α\(𝐗\)\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)atα=0\.01\\alpha=0\.01with100100split permutations\. Since the MC approximation can leave allDDstreams belowα\\alpha, yielding an empty set, we report the penalized size\|𝒦α\(𝐗\)\|pen\|\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|\_\{\\mathrm\{pen\}\}, assigning sizeDDto empty realizations\. When explicitly stated, we will also evaluate the prior\-informed variant from Sec\.[V\-C](https://arxiv.org/html/2607.26481#S5.SS3), assigning weightsvd=v\>1v\_\{d\}=v\>1to thekkmost likely streams andvd′=\(D−kv\)/\(D−k\)v\_\{d\}^\{\\prime\}=\(\{D\-kv\}\)/\(\{D\-k\}\)for all other streams\. The true root stream is always included among thekklikely streams, indicating a well specified prior\.
#### VI\-A2Baselines
We compare the proposed methods against an unweighted baseline that ignores contamination and an oracle baseline with access to the true contamination of each observation:
- •CONCH and CROC\[[14](https://arxiv.org/html/2607.26481#bib.bib2),[15](https://arxiv.org/html/2607.26481#bib.bib37)\]: The original methods that ignore contamination, as described in Sec\.[III\-A](https://arxiv.org/html/2607.26481#S3.SS1)and Sec\.[V\-A](https://arxiv.org/html/2607.26481#S5.SS1), respectively\.
- •Oracle\-based W\-CONCH and W\-CROC: These methods use the true contamination indicatorsYiY\_\{i\}in \([2](https://arxiv.org/html/2607.26481#S2.E2)\) to define oracle weightswiorc=1−Yiw\_\{i\}^\{\\mathrm\{orc\}\}=1\-Y\_\{i\}, assigning weight11to clean observations and weight0to contaminated observations\.
Figure 5:Performance against the contamination levelsε\\varepsilonfor CONCH\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\], oracle\-based W\-CONCH, W\-CONCH with hard and soft weighting, and MW\-CONCH using EDL or MC dropout uncertainty models \(EDL\-F and MC\-F denote fixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}meta\-learning withβ=ε\\beta=\\varepsilon, while EDL\-M and MC\-M denote mixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}meta\-learning withβ=0\.3\\beta=0\.3\): \(a\) Average confidence set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|on DomainNet\. \(b\) Average confidence set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|on CIFAR\-100\. \(c\) Average confidence set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|on the Milan telecom dataset\.Figure 6:W\-CROC and prior\-informed CROC on CIFAR\-100 withD=5D=5streams,σ=0\.3\\sigma=0\.3,n=400n=400, andα=0\.01\\alpha=0\.01\. \(a\) Average penalized confidence set size\|𝒦α\(𝐗\)\|pen\|\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|\_\{\\mathrm\{pen\}\}across contamination levels\. \(b\) Effect of the prior weightvvon the penalized set size atε=0\.7\\varepsilon=0\.7withαmax=0\.15\\alpha\_\{\\max\}=0\.15, wherekkdenotes the number of streams deemed likely by the prior\.
### VI\-BResults
#### VI\-B1Changepoint Localization
Effect of the weighting parameters\.We first examine how the quantile hyperparameterβ\\betaand the soft\-weighting temperature hyperparameterλ\\lambdaaffect efficiency\. Fig\.[4](https://arxiv.org/html/2607.26481#S6.F4)\(a\) shows the confidence set size obtained by hard uncertainty thresholding for different values ofβ\\betaon DomainNet and CIFAR\-100\. The results support the heuristic that, whenε\\varepsilonis known, hyperparameterβ\\betashould be matched toε\\varepsilon\.
Settingβ=ε\\beta=\\varepsilon, Fig\.[4](https://arxiv.org/html/2607.26481#S6.F4)\(b\) compares hard weighting with soft weighting for different hyperparameterλ\\lambda\. Based on these results, we useλ=0\.05\\lambda=0\.05for DomainNet andλ=0\.01\\lambda=0\.01for CIFAR\-100 in the W\-CONCH comparisons below\. Comparison against contamination levels\.We compare CONCH, oracle\-based W\-CONCH, and the proposed W\-CONCH variants across contamination levels\. For hard and soft W\-CONCH, we use the matched choiceβ=ε\\beta=\\varepsilon\. Following the meta\-training design discussed in Sec\.[IV\-D](https://arxiv.org/html/2607.26481#S4.SS4), we evaluate MW\-CONCH under both fixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}meta\-learning, tested withεtr=ε\\varepsilon\_\{\\mathrm\{tr\}\}=\\varepsilon, and mixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}meta\-learning\. We consider the EDL and MC dropout uncertainty models, denoting the resulting variants MW\-CONCH \(EDL\-F/MC\-F\) and MW\-CONCH \(EDL\-M/MC\-M\), respectively\.
As shown in Fig\.[5](https://arxiv.org/html/2607.26481#S6.F5)\(a\) and \(b\), CONCH becomes increasingly inefficient as contamination grows, with the confidence set size on DomainNet rising from2\.072\.07atε=0\\varepsilon=0to137\.43137\.43atε=0\.7\\varepsilon=0\.7, and similarly on CIFAR\-100 from1\.001\.00to345\.27345\.27\. Uncertainty\-based weighting reduces this inflation\. Atε=0\.7\\varepsilon=0\.7, hard and soft W\-CONCH bring the set size down to18\.7718\.77and22\.8422\.84on DomainNet, and to7\.247\.24and9\.189\.18on CIFAR\-100\.
Meta\-learning improves efficiency further\. On DomainNet, atε=0\.7\\varepsilon=0\.7, MW\-CONCH \(EDL\-F\) reaches3\.613\.61, even outperforming oracle\-based W\-CONCH, which reaches9\.819\.81, suggesting that meta\-learning not only improves the uncertainty estimates used for weighting, but also learns classifier representations that are more robust to contamination\. MW\-CONCH \(EDL\-M\) produces a slightly larger confidence set atε=0\.7\\varepsilon=0\.7, reaching6\.766\.76, but still substantially improves over direct weighting, while not requiring knowledge of the test contamination level\. On CIFAR\-100, atε=0\.7\\varepsilon=0\.7, MW\-CONCH \(EDL\-F\) reaches5\.595\.59, close to the oracle\-based W\-CONCH at5\.215\.21\. MW\-CONCH \(EDL\-M\) reaches5\.645\.64, again substantially improving over W\-CONCH\.
Fig\.[5](https://arxiv.org/html/2607.26481#S6.F5)\(c\) reports the corresponding results on the Milan telecom dataset, where the same qualitative trends hold\. CONCH inflates from4\.184\.18atε=0\\varepsilon=0to56\.3756\.37atε=0\.7\\varepsilon=0\.7, while hard and soft W\-CONCH substantially reduce this inflation, reaching36\.5436\.54and30\.2830\.28atε=0\.7\\varepsilon=0\.7, respectively\. MW\-CONCH \(EDL\-F\) performs similarly to soft weighting, reaching30\.1430\.14atε=0\.7\\varepsilon=0\.7, while MW\-CONCH \(EDL\-M\) achieves a comparable reduction, reaching36\.9136\.91atε=0\.7\\varepsilon=0\.7\. As a trade\-off, MW\-CONCH \(EDL\-M\) is less efficient at low contamination, reaching8\.148\.14atε=0\\varepsilon=0compared to4\.184\.18for CONCH\. Oracle\-based W\-CONCH remains the most efficient throughout, reaching26\.6226\.62atε=0\.7\\varepsilon=0\.7\.
The MC dropout variants also reduce the set size at high contamination, showing that the framework is not tied to EDL, although they are less effective than their EDL counterparts\. This suggests that the efficiency gains depend partly on the quality of the uncertainty estimates used for weighting\. Overall, uncertainty\-based weighting substantially improves the efficiency of CONCH under contamination, while meta\-learning further improves both the uncertainty estimates used for weighting and the robustness of the learned classifier under contamination\.
#### VI\-B2Root\-Cause Localization
Figure[6](https://arxiv.org/html/2607.26481#S6.F6)\(a\) reports the average penalized root\-cause confidence set size\|𝒦α\(𝐗\)\|pen\|\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|\_\{\\mathrm\{pen\}\}across contamination levels\. As contamination increases, the set size produced by CROC grows substantially, reaching4\.8354\.835atε=0\.7\\varepsilon=0\.7, indicating a loss of root\-cause specificity\. Uncertainty\-based weighting reduces this inflation at high contamination: atε=0\.7\\varepsilon=0\.7, W\-CROC with hard and soft weighting reduces the penalized set size to1\.8401\.840and2\.1002\.100, respectively\. Oracle\-based W\-CROC further reduces the penalized set size to1\.6301\.630, providing an upper performance bound\. MW\-CROC \(EDL\-M\) achieves a similar reduction, with penalized set size1\.9151\.915, while not requiring knowledge of the test contamination level\. MW\-CROC \(EDL\-F\) also improves over CROC, but gives a larger set size of2\.8602\.860atε=0\.7\\varepsilon=0\.7\. The MC dropout variants reduce the set size in some regimes but are generally less effective than their EDL counterparts at high contamination, with MW\-CROC \(MC\-F\) and MW\-CROC \(MC\-M\) reaching4\.4904\.490and3\.5103\.510, respectively\. This suggests that the efficiency gains depend partly on the quality of the uncertainty estimates used for weighting\.
Figure[6](https://arxiv.org/html/2607.26481#S6.F6)\(b\) evaluates the prior\-informed CROC described in Sec\.[V\-C](https://arxiv.org/html/2607.26481#S5.SS3)atε=0\.7\\varepsilon=0\.7withαmax=0\.15\\alpha\_\{\\max\}=0\.15\. As the prior weightvvincreases from11andkkdecreases, the prior grows more informative and, as a result, the penalized set size decreases\. Withk=1k=1, the set size drops from4\.8854\.885atv=1v=1to3\.6103\.610atv=5v=5\. Largerkkproduces more moderate reductions, reaching4\.1504\.150atv=2\.5v=2\.5fork=2k=2and4\.5604\.560atv≈1\.67v\\approx 1\.67fork=3k=3, because fewer unlikely streams remain to absorb the redistributed significance budget\.
## VIIConclusions
We studied conformal changepoint localization and root cause analysis under contaminated observations\. We showed that split exchangeability, and therefore the finite\-sample distribution\-free coverage guarantees of CONCH\[[14](https://arxiv.org/html/2607.26481#bib.bib2)\]and CROC\[[15](https://arxiv.org/html/2607.26481#bib.bib37)\], are preserved under a Huber\-type contamination model, while contamination can substantially inflate confidence set size\. To address this, we proposed W\-CONCH and W\-CROC, which downweight observations that are likely to be corrupted without affecting conformal calibration, thereby preserving coverage while reducing confidence set size\. We introduced a meta\-learning approach that learns the weighting rule without requiring knowledge of the contamination level at test time\. Experiments on image\-based and real\-world changepoint and root\-cause benchmarks demonstrated substantial reductions in confidence set size while maintaining the target coverage\.
Several research directions remain open\. The framework connects confidence sets to downstream, potentially risk\-averse, decision making, and coupling the size objective directly to a decision loss is a natural next step\. Extending the approach to streaming and online settings, to multiple or unknown numbers of changepoints, and to richer inter\-stream dependency structures in root cause analysis would broaden its applicability\. Finally, characterizing how the informativeness of the confidence set degrades as a function of the contamination level, and deriving guarantees on set size under structured contamination, would complement the coverage guarantees established here\[[53](https://arxiv.org/html/2607.26481#bib.bib8)\]\.
## Appendix AProofs and Derivations
### A\-AProof of Lemma[1](https://arxiv.org/html/2607.26481#Thmlemma1)
We prove the claim for the pre\-change segment\. The post\-change case follows identically with indices shifted to\{ξ\+1,…,n\}\\\{\\xi\+1,\\ldots,n\\\}\. Fix a permutationπ0,ξ\\pi\_\{0,\\xi\}of\{1,…,ξ\}\\\{1,\\ldots,\\xi\\\}\. Since\(Y1,Z1\),…,\(Yξ,Zξ\)\(Y\_\{1\},Z\_\{1\}\),\\ldots,\(Y\_\{\\xi\},Z\_\{\\xi\}\)are i\.i\.d\. and independent of the clean segment, Assumption[1](https://arxiv.org/html/2607.26481#Thmassumption1)implies that the augmented sequence\{\(X~i,Yi,Zi\)\}i≤ξ\\big\\\{\(\\widetilde\{X\}\_\{i\},Y\_\{i\},Z\_\{i\}\)\\big\\\}\_\{i\\leq\\xi\}is exchangeable\. Applying the deterministic contamination mapϕ\(x~,y,z\)=\(1−y\)x~\+yz\\phi\(\\widetilde\{x\},y,z\)=\(1\-y\)\\widetilde\{x\}\+yzcomponentwise, which satisfiesXi=ϕ\(X~i,Yi,Zi\)X\_\{i\}=\\phi\(\\widetilde\{X\}\_\{i\},Y\_\{i\},Z\_\{i\}\), preserves this distributional equality\(X1,…,Xξ\)=𝑑\(Xπ0,ξ\(1\),…,Xπ0,ξ\(ξ\)\)\.\(X\_\{1\},\\ldots,X\_\{\\xi\}\)\\overset\{d\}\{=\}\(X\_\{\\pi\_\{0,\\xi\}\(1\)\},\\ldots,X\_\{\\pi\_\{0,\\xi\}\(\\xi\)\}\)\.Sinceπ0,ξ\\pi\_\{0,\\xi\}was arbitrary,\(X1,…,Xξ\)\(X\_\{1\},\\ldots,X\_\{\\xi\}\)is exchangeable\. The same argument for\{ξ\+1,…,n\}\\\{\\xi\+1,\\ldots,n\\\}establishes split exchangeability of𝐗\\mathbf\{X\}atξ\\xi\.
### A\-BProof of Proposition[1](https://arxiv.org/html/2607.26481#Thmproposition1)
We lower\-bound the mixture marginalgj\(x\)g\_\{j\}\(x\)in \([11](https://arxiv.org/html/2607.26481#S4.E11)\) by introducing the latent contamination indicatorY∈\{0,1\}Y\\in\\\{0,1\\\}and applying the ELBO tologgj\(x\)\\log g\_\{j\}\(x\)\. Under the contamination model,YYhas priorp\(Y=0\)=1−εp\(Y=0\)=1\-\\varepsilonandp\(Y=1\)=εp\(Y=1\)=\\varepsilon, with conditional densitiesgj\(x\|Y=0\)=fj\(x\)g\_\{j\}\(x\|Y=0\)=f\_\{j\}\(x\)andgj\(x\|Y=1\)=q\(x\)g\_\{j\}\(x\|Y=1\)=q\(x\), so thatgj\(x\)=∑ygj\(x\|y\)p\(y\)g\_\{j\}\(x\)=\\sum\_\{y\}g\_\{j\}\(x\|y\)p\(y\)as in \([11](https://arxiv.org/html/2607.26481#S4.E11)\)\.
For any auxiliary posteriorrj\(y∣x\)r\_\{j\}\(y\\mid x\)on\{0,1\}\\\{0,1\\\}, the ELBO identity\[[38](https://arxiv.org/html/2607.26481#bib.bib24)\]gives
loggj\(x\)\\displaystyle\\log g\_\{j\}\(x\)=𝔼rj\(y∣x\)\[loggj\(x∣y\)\]−KL\(rj\(y∣x\)∥p\(y\)\)\+KL\(rj\(y∣x\)∥gj\(y∣x\)\)\\displaystyle=\\mathbb\{E\}\_\{r\_\{j\}\(y\\mid x\)\}\\big\[\\log g\_\{j\}\(x\\mid y\)\\big\]\-\\operatorname\{KL\}\\big\(r\_\{j\}\(y\\mid x\)\\\|p\(y\)\\big\)\+\\operatorname\{KL\}\\big\(r\_\{j\}\(y\\mid x\)\\\|g\_\{j\}\(y\\mid x\)\\big\)≥𝔼rj\(y∣x\)\[loggj\(x∣y\)\]−KL\(rj\(y∣x\)∥p\(y\)\),\\displaystyle\\geq\\mathbb\{E\}\_\{r\_\{j\}\(y\\mid x\)\}\\big\[\\log g\_\{j\}\(x\\mid y\)\\big\]\-\\operatorname\{KL\}\\big\(r\_\{j\}\(y\\mid x\)\\\|p\(y\)\\big\),\(38\)where the inequality drops the nonnegative divergenceKL\(rj\(y∣x\)∥gj\(y∣x\)\)\\operatorname\{KL\}\\big\(r\_\{j\}\(y\\mid x\)\\\|g\_\{j\}\(y\\mid x\)\\big\)\.
Set the clean\-posterior weightwj\(x\):=rj\(Y=0\|x\)w\_\{j\}\(x\):=r\_\{j\}\(Y=0\|x\), so thatrj\(Y=1\|x\)=1−wj\(x\)r\_\{j\}\(Y=1\|x\)=1\-w\_\{j\}\(x\)\. SinceY∈\{0,1\}Y\\in\\\{0,1\\\}, the first term of \([A\-B](https://arxiv.org/html/2607.26481#A1.Ex1)\) satisfies
𝔼rj\(y∣x\)\[loggj\(x∣y\)\]\\displaystyle\\mathbb\{E\}\_\{r\_\{j\}\(y\\mid x\)\}\\big\[\\log g\_\{j\}\(x\\mid y\)\\big\]=wj\(x\)logfj\(x\)\+\(1−wj\(x\)\)logq\(x\)\\displaystyle=w\_\{j\}\(x\)\\log f\_\{j\}\(x\)\+\\big\(1\-w\_\{j\}\(x\)\\big\)\\log q\(x\)≥wj\(x\)logfj\(x\)−Cq,\\displaystyle\\geq w\_\{j\}\(x\)\\log f\_\{j\}\(x\)\-C\_\{q\},\(39\)usinggj\(x∣Y=0\)=fj\(x\)g\_\{j\}\(x\\mid Y\{=\}0\)=f\_\{j\}\(x\),gj\(x∣Y=1\)=q\(x\)g\_\{j\}\(x\\mid Y\{=\}1\)=q\(x\), andlogq\(x\)≥−Cq\\log q\(x\)\\geq\-C\_\{q\}with1−wj\(x\)∈\[0,1\]1\-w\_\{j\}\(x\)\\in\[0,1\]\. For the second term, sincerj\(y∣x\)≤1r\_\{j\}\(y\\mid x\)\\leq 1andp\(y\)≥min\{ε,1−ε\}p\(y\)\\geq\\min\\\{\\varepsilon,1\-\\varepsilon\\\}forε∈\(0,1\)\\varepsilon\\in\(0,1\),
KL\(rj\(y∣x\)∥p\(y\)\)\\displaystyle\\operatorname\{KL\}\\big\(r\_\{j\}\(y\\mid x\)\\\|p\(y\)\\big\)=𝔼rj\(y∣x\)\[logrj\(y∣x\)p\(y\)\]\\displaystyle=\\mathbb\{E\}\_\{r\_\{j\}\(y\\mid x\)\}\\Big\[\\log\\tfrac\{r\_\{j\}\(y\\mid x\)\}\{p\(y\)\}\\Big\]≤log1min\{ε,1−ε\}\.\\displaystyle\\leq\\log\\tfrac\{1\}\{\\min\\\{\\varepsilon,1\-\\varepsilon\\\}\}\.\(40\)Substituting \([A\-B](https://arxiv.org/html/2607.26481#A1.Ex2)\) and \([A\-B](https://arxiv.org/html/2607.26481#A2.EGx7)\) into \([A\-B](https://arxiv.org/html/2607.26481#A1.Ex1)\) givesloggj\(x\)≥wj\(x\)logfj\(x\)−CR\\log g\_\{j\}\(x\)\\geq w\_\{j\}\(x\)\\log f\_\{j\}\(x\)\-C\_\{R\}withCR=Cq\+log\(1/min\{ε,1−ε\}\)C\_\{R\}=C\_\{q\}\+\\log\(1/\\min\\\{\\varepsilon,1\-\\varepsilon\\\}\), which is independent ofjj\. SettingC:=e−CRC:=e^\{\-C\_\{R\}\}, which depends only onCqC\_\{q\}andε\\varepsilon, and exponentiating yields \([13](https://arxiv.org/html/2607.26481#S4.E13)\)\.
### A\-CDerivation of the Classifier\-Based Weighted CPP Score
Using the classifier log\-posterior ratioΔi\\Delta\_\{i\}in \([15](https://arxiv.org/html/2607.26481#S4.E15)\), we substitute this approximation into the weighted CPP score \([14](https://arxiv.org/html/2607.26481#S4.E14)\)\. We use the candidate\-dependent weightw\(t\)w^\{\(t\)\}, the weighted log\-likelihoodNt\(s\)N\_\{t\}\(s\), and the weighted maximum\-likelihood estimateξ^t\(𝐗\)=argmaxsNt\(s\)\\hat\{\\xi\}\_\{t\}\(\\mathbf\{X\}\)=\\arg\\max\_\{s\}N\_\{t\}\(s\)from Definition[2](https://arxiv.org/html/2607.26481#Thmdefinition2), so that the numerator and denominator of \([14](https://arxiv.org/html/2607.26481#S4.E14)\) areNt\(t\)N\_\{t\}\(t\)andNt\(ξ^t\(𝐗\)\)N\_\{t\}\(\\hat\{\\xi\}\_\{t\}\(\\mathbf\{X\}\)\), respectively\.
Applying the identitylogp^\(j=0∣Xi\)=Δi\+logp^\(j=1∣Xi\)\\log\\hat\{p\}\(j=0\\mid X\_\{i\}\)=\\Delta\_\{i\}\+\\log\\hat\{p\}\(j=1\\mid X\_\{i\}\)to the pre\-change sum in the definition ofNt\(s\)N\_\{t\}\(s\)gives, for any cutoffss,
Nt\(s\)=∑i≤sw\(t\)\(Xi\)Δi\+C1,N\_\{t\}\(s\)=\\sum\_\{i\\leq s\}w^\{\(t\)\}\(X\_\{i\}\)\\Delta\_\{i\}\+C\_\{1\},\(41\)where the remainderC1=∑iw\(t\)\(Xi\)logp^\(j=1∣Xi\)C\_\{1\}=\\sum\_\{i\}w^\{\(t\)\}\(X\_\{i\}\)\\log\\hat\{p\}\(j=1\\mid X\_\{i\}\)does not depend onss, andw\(t\)\(Xi\)=w0\(Xi\)w^\{\(t\)\}\(X\_\{i\}\)=w\_\{0\}\(X\_\{i\}\)fori≤ti\\leq tandw1\(Xi\)w\_\{1\}\(X\_\{i\}\)fori\>ti\>t\. Evaluating \([41](https://arxiv.org/html/2607.26481#A1.E41)\) ats=ts=tfor the numerator and ats=ξ^t\(𝐗\)s=\\hat\{\\xi\}\_\{t\}\(\\mathbf\{X\}\)for the denominator, and subtracting, cancels the common remainderC1C\_\{1\}and recovers the classifier\-based weighted CPP score in Definition[2](https://arxiv.org/html/2607.26481#Thmdefinition2)\.
### A\-DProof of Proposition[2](https://arxiv.org/html/2607.26481#Thmproposition2)
Fixt∈\{1,…,n−1\}t\\in\\\{1,\\dots,n\-1\\\}\. Under the nullH0:ξ=tH\_\{0\}:\\xi=t, Lemma[1](https://arxiv.org/html/2607.26481#Thmlemma1)givesπ\(𝐗\)=𝑑𝐗\\pi\(\\mathbf\{X\}\)\\overset\{d\}\{=\}\\mathbf\{X\}for allπ∈Πt\\pi\\in\\Pi\_\{t\}\. We show thatptwp\_\{t\}^\{w\}in \([23](https://arxiv.org/html/2607.26481#S4.E23)\) is super\-uniform\. Takingt=ξt=\\xi, whereH0H\_\{0\}holds, then givesPr\(ξ∉𝒞αw\)=Pr\(pξw≤α\)≤α\\Pr\(\\xi\\notin\\mathcal\{C\}\_\{\\alpha\}^\{w\}\)=\\Pr\(p\_\{\\xi\}^\{w\}\\leq\\alpha\)\\leq\\alpha, which is \([24](https://arxiv.org/html/2607.26481#S4.E24)\)\.
The weights depend on𝐗\\mathbf\{X\}only through the thresholdsκ0\(t,β\),κ1\(t,β\)\\kappa\_\{0\}\(t,\\beta\),\\kappa\_\{1\}\(t,\\beta\)in \([22](https://arxiv.org/html/2607.26481#S4.E22)\), which are empirical quantiles of\{Mi\}i≤t\\\{M\_\{i\}\\\}\_\{i\\leq t\}and\{Mi\}i\>t\\\{M\_\{i\}\\\}\_\{i\>t\}\. A permutationπ∈Πt\\pi\\in\\Pi\_\{t\}acts only within each segment and leaves these multisets unchanged, so the thresholds are invariant underπ\\pi\. SinceMi=h\(Xi\)M\_\{i\}=h\(X\_\{i\}\)depends only onXiX\_\{i\}, the weights are permuted consistently with the observations:
wj\(π\(𝐗\)\)=π\(wj\(𝐗\)\),∀π∈Πt,j=0,1\.w\_\{j\}\(\\pi\(\\mathbf\{X\}\)\)=\\pi\\bigl\(w\_\{j\}\(\\mathbf\{X\}\)\\bigr\),\\qquad\\forall\\pi\\in\\Pi\_\{t\},\\quad j=0,1\.\(42\)
Thus the weighted CPP scoreStwS\_\{t\}^\{w\}is permutation\-equivariant\. Averaging𝟙\{ptw\(π\(𝐗\)\)≤α\}\\mathds\{1\}\\\{p\_\{t\}^\{w\}\(\\pi\(\\mathbf\{X\}\)\)\\leq\\alpha\\\}overπ∈Πt\\pi\\in\\Pi\_\{t\}usingπ\(𝐗\)=𝑑𝐗\\pi\(\\mathbf\{X\}\)\\overset\{d\}\{=\}\\mathbf\{X\}, expanding the definition ofptwp\_\{t\}^\{w\}in \([23](https://arxiv.org/html/2607.26481#S4.E23)\), and relabelingπ′∘π\\pi^\{\\prime\}\\circ\\piasπ′\\pi^\{\\prime\}via the group property ofΠt\\Pi\_\{t\}gives
Pr\(ptw\(𝐗\)≤α\)=𝔼\[1\|Πt\|∑π∈Πt𝟙\{1\|Πt\|∑π′∈Πt𝟙\[Stw\(π′\(𝐗\)\)≤Stw\(π\(𝐗\)\)\]≤α\}\]≤α,\\Pr\\bigl\(p\_\{t\}^\{w\}\(\\mathbf\{X\}\)\\leq\\alpha\\bigr\)=\\mathbb\{E\}\\Biggl\[\\frac\{1\}\{\|\\Pi\_\{t\}\|\}\\sum\_\{\\pi\\in\\Pi\_\{t\}\}\\mathds\{1\}\\biggl\\\{\\frac\{1\}\{\|\\Pi\_\{t\}\|\}\\sum\_\{\\pi^\{\\prime\}\\in\\Pi\_\{t\}\}\\mathds\{1\}\\bigl\[S\_\{t\}^\{w\}\(\\pi^\{\\prime\}\(\\mathbf\{X\}\)\)\\leq S\_\{t\}^\{w\}\(\\pi\(\\mathbf\{X\}\)\)\\bigr\]\\leq\\alpha\\biggr\\\}\\Biggr\]\\leq\\alpha,\(43\)where the inner average is the rank ofStw\(π\(𝐗\)\)S\_\{t\}^\{w\}\(\\pi\(\\mathbf\{X\}\)\)among\{Stw\(π′\(𝐗\)\)\}π′∈Πt\\\{S\_\{t\}^\{w\}\(\\pi^\{\\prime\}\(\\mathbf\{X\}\)\)\\\}\_\{\\pi^\{\\prime\}\\in\\Pi\_\{t\}\}, and the final inequality is the finite\-sample rank bound of\[[14](https://arxiv.org/html/2607.26481#bib.bib2), Thm\. 3\.1\]\. Henceptwp\_\{t\}^\{w\}is super\-uniform underH0H\_\{0\}\.
### A\-EProof of Proposition[3](https://arxiv.org/html/2607.26481#Thmproposition3)
The argument in Appendix[A\-D](https://arxiv.org/html/2607.26481#A1.SS4)shows that, under the nullH0:ξ=tH\_\{0\}:\\xi=t, thepp\-valueptw\(𝐗\)p\_\{t\}^\{w\}\(\\mathbf\{X\}\)is super\-uniform, i\.e\.,Pr\(ptw\(𝐗\)≤c\)≤c\\Pr\\bigl\(p\_\{t\}^\{w\}\(\\mathbf\{X\}\)\\leq c\\bigr\)\\leq cfor any constantc∈\[0,1\]c\\in\[0,1\], since the rank bound \([43](https://arxiv.org/html/2607.26481#A1.E43)\) does not depend on the specific value of the threshold\. To establish the boundPr\(ξ∈𝒞α,vw\(𝐗\)\)≥1−αmax\\Pr\\bigl\(\\xi\\in\\mathcal\{C\}\_\{\\alpha,v\}^\{w\}\(\\mathbf\{X\}\)\\bigr\)\\geq 1\-\\alpha\_\{\\max\}, the changepointξ\\xiis excluded from the set \([27](https://arxiv.org/html/2607.26481#S4.E27)\) only ifpξw\(𝐗\)≤αξp\_\{\\xi\}^\{w\}\(\\mathbf\{X\}\)\\leq\\alpha\_\{\\xi\}\. Applying the super\-uniformity att=ξt=\\xiwithc=αξc=\\alpha\_\{\\xi\}, and usingαξ≤αmax\\alpha\_\{\\xi\}\\leq\\alpha\_\{\\max\}from \([26](https://arxiv.org/html/2607.26481#S4.E26)\), gives
Pr\(ξ∉𝒞α,vw\(𝐗\)\)=Pr\(pξw\(𝐗\)≤αξ\)≤αξ≤αmax\.\\Pr\\bigl\(\\xi\\notin\\mathcal\{C\}\_\{\\alpha,v\}^\{w\}\(\\mathbf\{X\}\)\\bigr\)=\\Pr\\bigl\(p\_\{\\xi\}^\{w\}\(\\mathbf\{X\}\)\\leq\\alpha\_\{\\xi\}\\bigr\)\\leq\\alpha\_\{\\xi\}\\leq\\alpha\_\{\\max\}\.\(44\)For the prior\-averaged guarantee \([28](https://arxiv.org/html/2607.26481#S4.E28)\), averaging over the priorPr\(ξ=t\)=vtn−1\\Pr\(\\xi=t\)=\\frac\{v\_\{t\}\}\{n\-1\}with the boundαt≤αvt\\alpha\_\{t\}\\leq\\frac\{\\alpha\}\{v\_\{t\}\}from \([26](https://arxiv.org/html/2607.26481#S4.E26)\) gives
Pr\(ξ∉𝒞α,vw\(𝐗\)\)=∑t=1n−1vt⋅Pr\(ptw\(𝐗\)≤αt∣ξ=t\)n−1≤α\\Pr\\bigl\(\\xi\\notin\\mathcal\{C\}\_\{\\alpha,v\}^\{w\}\(\\mathbf\{X\}\)\\bigr\)=\\sum\_\{t=1\}^\{n\-1\}\\frac\{v\_\{t\}\\cdot\\Pr\\bigl\(p\_\{t\}^\{w\}\(\\mathbf\{X\}\)\\leq\\alpha\_\{t\}\\mid\\xi=t\\bigr\)\}\{n\-1\}\\leq\\alpha\(45\)
Algorithm 1MW\-CONCH and MW\-CROC1:Training contamination levels
\{εtr\(ℓ\)\}ℓ=1L\\\{\\varepsilon\_\{\\mathrm\{tr\}\}^\{\(\\ell\)\}\\\}\_\{\\ell=1\}^\{L\}, with
L=1L=1for fixed\-
εtr\\varepsilon\_\{\\mathrm\{tr\}\}training and
L\>1L\>1for mixed\-
εtr\\varepsilon\_\{\\mathrm\{tr\}\}training; training tasks
\{𝐗\(k,ℓ\)\}k=1K\\\{\\mathbf\{X\}^\{\(k,\\ell\)\}\\\}\_\{k=1\}^\{K\}at each
εtr\(ℓ\)\\varepsilon\_\{\\mathrm\{tr\}\}^\{\(\\ell\)\}, where each task is a single sequence for MW\-CONCH and a
DD\-stream tuple
𝐗\(k,ℓ\)=\(𝐗1\(k,ℓ\),…,𝐗D\(k,ℓ\)\)\\mathbf\{X\}^\{\(k,\\ell\)\}=\(\\mathbf\{X\}\_\{1\}^\{\(k,\\ell\)\},\\ldots,\\mathbf\{X\}\_\{D\}^\{\(k,\\ell\)\}\)for MW\-CROC; learning rate
η\\eta; quantile parameter
β\\beta; smoothing parameters
τq,τ1,τ2\\tau\_\{q\},\\tau\_\{1\},\\tau\_\{2\}; soft\-weight temperature
λ\\lambda
2:Initialize
θ\\theta
3:foreach gradient stepdo
4:foreach level
εtr\(ℓ\)\\varepsilon\_\{\\mathrm\{tr\}\}^\{\(\\ell\)\}and each single\-stream sequence
ZZin its training tasks \(i\.e\.,
Z=𝐗\(k,ℓ\)Z=\\mathbf\{X\}^\{\(k,\\ell\)\}for MW\-CONCH, or
Z=𝐗d\(k,ℓ\)Z=\\mathbf\{X\}\_\{d\}^\{\(k,\\ell\)\},
d=1,…,Dd=1,\\dots,D, for MW\-CROC\)do
5:Compute uncertainty scores
Mi=h\(Zi;θ\)M\_\{i\}=h\(Z\_\{i\};\\theta\)for
i=1,…,ni=1,\\dots,n
6:forCandidate changepoints
t=1,…,n−1t=1,\\dots,n\-1do
7:Compute
κ^0\(t,β\)\\hat\{\\kappa\}\_\{0\}\(t,\\beta\)and
κ^1\(t,β\)\\hat\{\\kappa\}\_\{1\}\(t,\\beta\)using \([47](https://arxiv.org/html/2607.26481#A1.E47)\)
8:Compute
w^0,i\(t;θ\)\\hat\{w\}\_\{0,i\}\(t;\\theta\)and
w^1,i\(t;θ\)\\hat\{w\}\_\{1,i\}\(t;\\theta\)using \([48](https://arxiv.org/html/2607.26481#A1.E48)\)
9:Compute
p~t\(Z;θ\)\\tilde\{p\}\_\{t\}\(Z;\\theta\)using \([49](https://arxiv.org/html/2607.26481#A1.E49)\)
10:endfor
11:Compute
ℓ~\(Z;θ\)\\tilde\{\\ell\}\(Z;\\theta\)from
\{p~t\(Z;θ\)\}t=1n−1\\\{\\tilde\{p\}\_\{t\}\(Z;\\theta\)\\\}\_\{t=1\}^\{n\-1\}using \([50](https://arxiv.org/html/2607.26481#A1.E50)\)
12:endfor
13:For each
εtr\(ℓ\)\\varepsilon\_\{\\mathrm\{tr\}\}^\{\(\\ell\)\}, compute
ℒ~\(θ;εtr\(ℓ\)\)\\tilde\{\\mathcal\{L\}\}\(\\theta;\\varepsilon\_\{\\mathrm\{tr\}\}^\{\(\\ell\)\}\)by averaging
ℓ~\(Z;θ\)\\tilde\{\\ell\}\(Z;\\theta\)over all sequences at that level \(
1K∑k\\frac\{1\}\{K\}\\sum\_\{k\}for MW\-CONCH,
1KD∑d∑k\\frac\{1\}\{KD\}\\sum\_\{d\}\\sum\_\{k\}for MW\-CROC\)
14:Update
θ←θ−η∇θ\[1L∑ℓ=1Lℒ~\(θ;εtr\(ℓ\)\)\]\\theta\\leftarrow\\theta\-\\eta\\nabla\_\{\\theta\}\\left\[\\frac\{1\}\{L\}\\sum\_\{\\ell=1\}^\{L\}\\tilde\{\\mathcal\{L\}\}\(\\theta;\\varepsilon\_\{\\mathrm\{tr\}\}^\{\(\\ell\)\}\)\\right\]
15:endfor
16:returnlearned uncertainty estimator
h\(⋅;θ∗\)h\(\\cdot;\\theta^\{\*\}\)
### A\-FDifferentiable Training Loss and Meta\-learning Algorithm for MW\-CONCH and MW\-CROC
The meta\-training loss \([25](https://arxiv.org/html/2607.26481#S4.E25)\) is not differentiable because of three operations: the empirical quantiles in the thresholds \([22](https://arxiv.org/html/2607.26481#S4.E22)\), the indicator in the split\-permutationpp\-value, and the counting operation used to compute the confidence set size \([23](https://arxiv.org/html/2607.26481#S4.E23)\)\. We replace each by a smooth surrogate controlled by a temperature parameter, following\[[28](https://arxiv.org/html/2607.26481#bib.bib22),[43](https://arxiv.org/html/2607.26481#bib.bib3)\]\.
First, we replace the empirical quantile in \([22](https://arxiv.org/html/2607.26481#S4.E22)\) by a differentiable counterpart\[[28](https://arxiv.org/html/2607.26481#bib.bib22)\]\. For any finite set\{ar\}r=1m\\\{a\_\{r\}\\\}\_\{r=1\}^\{m\}and quantile level1−β1\-\\beta, let
ρ1−β\(a;\{ar\}r=1m\+1\)=∑r=1m\+1β\(a−ar\)\+\+\(1−β\)\(ar−a\)\+,\\rho\_\{1\-\\beta\}\(a;\\\{a\_\{r\}\\\}\_\{r=1\}^\{m\+1\}\)=\\sum\_\{r=1\}^\{m\+1\}\\beta\(a\-a\_\{r\}\)^\{\+\}\+\(1\-\\beta\)\(a\_\{r\}\-a\)^\{\+\},\(46\)denote the pinball loss, where\(x\)\+=max\(x,0\)\(x\)^\{\+\}=\\max\(x,0\)andam\+1=max\(\{ar\}r=1m\)\+δa\_\{m\+1\}=\\max\(\\\{a\_\{r\}\\\}\_\{r=1\}^\{m\}\)\+\\deltawithδ\>0\\delta\>0is an auxiliary point ensuring the soft quantile is well\-defined at the boundary\. The empirical quantile is then approximated by
Q^1−β\(\{ar\}r=1m\)=∑r=1m\+1arsoftmax\(−ρ1−β\(ar\)τq\),\\hat\{Q\}\_\{1\-\\beta\}\(\\\{a\_\{r\}\\\}\_\{r=1\}^\{m\}\)=\\sum\_\{r=1\}^\{m\+1\}a\_\{r\}\\,\\mathrm\{softmax\}\\bigl\(\\frac\{\-\\rho\_\{1\-\\beta\}\(a\_\{r\}\)\}\{\\tau\_\{q\}\}\\bigr\),\(47\)wheresoftmax\(xr\)=exp\(xr\)/∑ℓexp\(xℓ\)\\mathrm\{softmax\}\(x\_\{r\}\)=\{\\exp\(x\_\{r\}\)\}/\{\\sum\_\{\\ell\}\\exp\(x\_\{\\ell\}\)\}andτq\>0\\tau\_\{q\}\>0is a smoothing temperature\. Replacing the empirical quantiles in \([22](https://arxiv.org/html/2607.26481#S4.E22)\) byQ^1−β\\hat\{Q\}\_\{1\-\\beta\}yields smoothed thresholdsκ^j\(t,β\)\\hat\{\\kappa\}\_\{j\}\(t,\\beta\), with which we define the differentiable training\-time weight
w^j\(Xi;θ\)=σ\(κ^j\(t,β\)−h\(Xi;θ\)λ\),j∈\{0,1\},\\hat\{w\}\_\{j\}\(X\_\{i\};\\theta\)=\\sigma\\left\(\\frac\{\\hat\{\\kappa\}\_\{j\}\(t,\\beta\)\-h\(X\_\{i\};\\theta\)\}\{\\lambda\}\\right\),j\\in\\\{0,1\\\},\(48\)in direct analogy with the soft weighting rule \([21](https://arxiv.org/html/2607.26481#S4.E21)\)\.
Second, we replace the indicator in the split\-permutationpp\-value by a sigmoid\. WritingStw^S\_\{t\}^\{\\hat\{w\}\}for the weighted CPP score evaluated with the training\-time weights\{w^j\(⋅;θ\)\}j=01\\\{\\hat\{w\}\_\{j\}\(\\cdot;\\theta\)\\\}\_\{j=0\}^\{1\}in \([48](https://arxiv.org/html/2607.26481#A1.E48)\):
p~t\(θ\)=1\|Πt\|∑π∈Πtσ\(1τ1\(Stw^\(𝐗\)−Stw^\(π\(𝐗\)\)\)\),\\tilde\{p\}\_\{t\}\(\\theta\)=\\frac\{1\}\{\|\\Pi\_\{t\}\|\}\\sum\_\{\\pi\\in\\Pi\_\{t\}\}\\sigma\\biggl\(\\frac\{1\}\{\\tau\_\{1\}\}\\Bigl\(S\_\{t\}^\{\\hat\{w\}\}\(\\mathbf\{X\}\)\-S\_\{t\}^\{\\hat\{w\}\}\(\\pi\(\\mathbf\{X\}\)\)\\Bigr\)\\biggr\),\(49\)τ1\>0\\tau\_\{1\}\>0controls the sharpness of the sigmoid\.
Third, we replace the threshold indicator in the set size by a sigmoid\[[43](https://arxiv.org/html/2607.26481#bib.bib3)\], yielding the per\-task smoothed size
ℓ~\(𝐗;θ\)=∑t=1n−1σ\(p~t\(θ\)−ατ2\),\\tilde\{\\ell\}\(\\mathbf\{X\};\\theta\)=\\sum\_\{t=1\}^\{n\-1\}\\sigma\\left\(\\frac\{\\tilde\{p\}\_\{t\}\(\\theta\)\-\\alpha\}\{\\tau\_\{2\}\}\\right\),\(50\)whereτ2\>0\\tau\_\{2\}\>0controls the sharpness of the threshold\.
Substitutingℓ~\\tilde\{\\ell\}for\|𝒞α,w\(θ\)\|\|\\mathcal\{C\}\_\{\\alpha,w\(\\theta\)\}\|in \([25](https://arxiv.org/html/2607.26481#S4.E25)\) gives the differentiable training lossℒ~\(θ;ε\)=1K∑k=1Kℓ~\(𝐗\(k,ε\);θ\)\\tilde\{\\mathcal\{L\}\}\(\\theta;\\varepsilon\)=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\tilde\{\\ell\}\\bigl\(\\mathbf\{X\}^\{\(k,\\varepsilon\)\};\\theta\\bigr\)\. As the temperaturesτq,τ1,τ2→0\\tau\_\{q\},\\tau\_\{1\},\\tau\_\{2\}\\to 0are annealed during training, each surrogate recovers its non\-differentiable counterpart andℒ~\(θ;ε\)→ℒ\(θ;ε\)\\tilde\{\\mathcal\{L\}\}\(\\theta;\\varepsilon\)\\to\\mathcal\{L\}\(\\theta;\\varepsilon\)\. The full procedure is given in Algorithm[1](https://arxiv.org/html/2607.26481#alg1)\.
## Appendix BExperimental Details and Results
### B\-AImplementation Details
Training details and experimental settings\.For DomainNet\[[29](https://arxiv.org/html/2607.26481#bib.bib21)\]and CIFAR\-100\[[22](https://arxiv.org/html/2607.26481#bib.bib30)\], the uncertainty model uses a ResNet\-18\[[13](https://arxiv.org/html/2607.26481#bib.bib26)\]backbone\. The same uncertainty model is used for both MW\-CONCH and MW\-CROC\. During meta\-learning fine\-tuning, the backbone is frozen and only the selection head is updated\. Unless otherwise specified, models are trained for100100epochs\. MW\-CONCH / MW\-CROC \(EDL\-F\)\.A separate model is trained for each fixed training contamination levelεtr∈\{0\.1,0\.3,0\.5,0\.7\}\\varepsilon\_\{\\mathrm\{tr\}\}\\in\\\{0\.1,0\.3,0\.5,0\.7\\\}, with the soft\-quantile threshold parameter set toβ=εtr\\beta=\\varepsilon\_\{\\mathrm\{tr\}\}\. At each epoch,1010tasks are sampled at contamination levelεtr\\varepsilon\_\{\\mathrm\{tr\}\}\. MW\-CONCH / MW\-CROC \(EDL\-M\)\.A single model is trained over a mixture of contamination levels, whereεtr\\varepsilon\_\{\\mathrm\{tr\}\}is chosen from the grid\{0\.0,0\.1,…,0\.7\}\\\{0\.0,0\.1,\\ldots,0\.7\\\}\. The threshold parameter is fixed atβ=0\.3\\beta=0\.3for all mixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}experiments\. At each epoch,55tasks are sampled per level, yielding4040tasks in total\. Changepoint position randomization\.To prevent the model from learning position\-specific suppression patterns, the changepoint positionξ\\xiis sampled uniformly from a fixed range for each training task rather than being held fixed\. For DomainNet \(n=800n=800\),ξ∼Unif\{160,…,640\}\.\\xi\\sim\\mathrm\{Unif\}\\\{160,\\dots,640\\\}\.For CIFAR\-100 \(n=400n=400\),ξ∼Unif\{120,…,280\}\\xi\\sim\\mathrm\{Unif\}\\\{120,\\dots,280\\\}in the W\-CONCH setting andξ∼Unif\{80,…,320\}\\xi\\sim\\mathrm\{Unif\}\\\{80,\\dots,320\\\}in the W\-CROC setting\.
For both training regimes and all datasets, the smoothing parametersτ1\\tau\_\{1\}andτ2\\tau\_\{2\}are initialized at0\.50\.5and annealed with decay factor0\.950\.95per epoch\. The soft\-quantile temperature is fixed atτq=1\.0\\tau\_\{q\}=1\.0, and the same soft\-quantile operator is used to compute the side\-specific thresholds of the soft weighting rule at evaluation time\. The sigmoid temperatureλ\\lambdafor the soft weights is set toλ=0\.05\\lambda=0\.05for all DomainNet experiments andλ=0\.01\\lambda=0\.01for CIFAR\-100 W\-CONCH experiments\. For CIFAR\-100 W\-CROC experiments, we useλ=0\.01\\lambda=0\.01for soft weighting and EDL\-based uncertainty, and useλ=0\.005\\lambda=0\.005only for the MW\-CROC variants based on MC Dropout\.
For DomainNet, the smoothing temperatures are annealed to a minimum ofτmin=0\.01\\tau\_\{\\min\}=0\.01, using5050split permutations per training task\. For CIFAR\-100,τmin=0\.05\\tau\_\{\\min\}=0\.05is used to avoid overly sharp sigmoid approximations at the end of the annealing schedule, with5050split permutations per training task\. The same training configuration applies to both MW\-CONCH and MW\-CROC\. Milan telecom: data construction\.Since the dataset provides no cell\-type annotations, evaluation cells are selected from their traffic profiles\. For each cell, we compute two ratios over the pre\-holiday period \(November 4–December 20\): the mean daily internet traffic on weekends divided by that on weekdays, and the mean daily traffic during the Christmas week \(December 23–27\) divided by the pre\-holiday mean\. We retain cells whose weekend\-to\-weekday ratio exceeds0\.80\.8, excluding office\-type cells whose strong weekly pattern would confound the holiday change, and whose Christmas\-to\-normal ratio exceeds0\.70\.7, excluding cells whose activity collapses over the holidays and therefore lacks a well\-defined localized changepoint; cells with low overall traffic or fewer than5050observed days are also discarded\. The ground\-truth changepoint is detected as the largest activity shift within December 20–26, keeping only cells whose changepoint falls on December 24 \(ξ=54\\xi=54\)\. Accordingly, training cells are labeled pre\-change \(j=0j=0\) before December 24 and post\-change \(j=1j=1\) thereafter\. Milan telecom: models and uncertainty\.All models use a residual MLP backbone \(720→256720\\to 256, three residual blocks\), with standardization statistics computed on the training\-cell pool\. Sincerelu\(logit\)\\mathrm\{relu\}\(\\text\{logit\}\)evidence is overconfident on noisy tabular inputs, the EDL model computes evidence with a radial basis function head\[[48](https://arxiv.org/html/2607.26481#bib.bib67)\],ek=γexp\(−‖ϕ\(x\)−ck‖2/\(2σ2\)\)e\_\{k\}=\\gamma\\exp\(\-\\\|\\phi\(x\)\-c\_\{k\}\\\|^\{2\}/\(2\\sigma^\{2\}\)\), which decays with the distance of the representationϕ\(x\)\\phi\(x\)from learnable class centroidsckc\_\{k\}\(with scalesγ,σ\\gamma,\\sigmalearned in log space and initialized toγ=e2\\gamma=e^\{2\},σ=1\\sigma=1\); the scoreΔi=loge0−loge1\\Delta\_\{i\}=\\log e\_\{0\}\-\\log e\_\{1\}and the uncertaintyMi=2/\(e0\+e1\+2\)M\_\{i\}=2/\(e\_\{0\}\+e\_\{1\}\+2\)are both produced by this model, trained on class\-balanced data with an evidence\-magnitude penalty of0\.010\.01\. For MC Dropout,MiM\_\{i\}is the standard deviation ofΔi\\Delta\_\{i\}acrossT=50T=50stochastic passes, as predictive entropy saturates under large tabular noise; unlike the EDL model, its underlying classifier is trained on all available days rather than class\-balanced data\. Milan telecom: meta\-learning\.The meta\-learning procedure follows the settings described above, except that the full model is fine\-tuned for3030epochs with learning rate10−410^\{\-4\}, using5050tasks per contamination level and100100split permutations per task, withτ1\\tau\_\{1\}andτ2\\tau\_\{2\}initialized at1\.01\.0; the sigmoid temperature isλ=0\.05\\lambda=0\.05, and the mixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}models useβ=0\.1\\beta=0\.1for EDL andβ=0\.3\\beta=0\.3for MC Dropout\.
### B\-BDetailed Experiment Results
#### B\-B1Detailed Changepoint Localization Results
We provide detailed numerical results corresponding to the experiments presented in the main text\. Each entry reports the average confidence set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|, with empirical coverage shown in parentheses\.
TABLE I:Detailed numerical results for DomainNet, CIFAR\-100, and the Milan telecom dataset\. Each entry reports the average confidence set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|\(coverage\)\. Hard, Soft, and EDL\-F \(as well as MC\-F\) useβ=ε\\beta=\\varepsilon\(matched to the test contamination level\)\. Forε=0\\varepsilon=0, the corresponding CONCH result is reported sinceβ=0\\beta=0is undefined\.
#### B\-B2Detailed Root\-Cause Localization Results
We provide detailed numerical results for the W\-CROC root\-cause localization experiments\. Each entry reports the average penalized confidence set size\|𝒦α\(𝐗\)\|pen\|\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|\_\{\\mathrm\{pen\}\}, empirical coverage, and the per\-stream root\-causepp\-valuesp\(1\),…,p\(5\)p\_\{\(1\)\},\\ldots,p\_\{\(5\)\}, where stream 1 is the true root cause\.
TABLE II:W\-CROC results on CIFAR\-100 \(noiseσ=0\.3\\sigma\{=\}0\.3,n=400n\{=\}400,α=0\.01\\alpha\{=\}0\.01,100100split permutations,200200test tasks\)\. W\-CROC Hard and Soft useβ=0\.7\\beta\{=\}0\.7; MW\-CROC EDL\-F and MC\-F useβ=ε\\beta\{=\}\\varepsilon; MW\-CROC EDL\-M and MC\-M use mixed\-εtr\\varepsilon\_\{\\mathrm\{tr\}\}training withβ=0\.3\\beta\{=\}0\.3\.
#### B\-B3Prior\-Informed CROC Results
Table[III](https://arxiv.org/html/2607.26481#A2.T3)reports the prior\-informed CROC results on CIFAR\-100 atε=0\.7\\varepsilon=0\.7withαmax=0\.15\\alpha\_\{\\max\}=0\.15\. The prior assigns weightvd=v\>1v\_\{d\}=v\>1to each of thekkmost likely streams andvd′=D−kvD−kv\_\{d\}^\{\\prime\}=\\frac\{D\-kv\}\{D\-k\}to all other streams\. The per\-stream significance levels areαd=min\(α/vd,αmax\)\\alpha\_\{d\}=\\min\(\\alpha/v\_\{d\},\\,\\alpha\_\{\\max\}\)\. The true root stream is always included among thekklikely streams, indicating a well\-specified prior\. Asvvincreases, the prior concentrates more weight on the likely streams, allowing the unlikely streams to be tested at higher levels and thus more easily excluded\. Withk=1k=1, the penalized set size drops from4\.8854\.885at the uniform baseline \(v=1v=1\) to3\.6103\.610atv=D=5v=D=5\. Larger likely sets produce more moderate reductions, as fewer unlikely streams remain to absorb the redistributed significance budget\.
TABLE III:Prior\-informed CROC on CIFAR\-100 atε=0\.7\\varepsilon\{=\}0\.7withα=0\.01\\alpha\{=\}0\.01andαmax=0\.15\\alpha\_\{\\max\}\{=\}0\.15\. Each row reports the penalized set size\|𝒦α\(𝐗\)\|pen\|\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|\_\{\\mathrm\{pen\}\}and empirical coverage\. The baseline row corresponds to standard CROC atα=0\.01\\alpha\{=\}0\.01\(v=1v\{=\}1\)\.
### B\-CMNIST Experiments
We report MNIST\[[7](https://arxiv.org/html/2607.26481#bib.bib25)\]results for both W\-CONCH and W\-CROC, complementing the DomainNet and CIFAR\-100 experiments presented\.
#### B\-C1MNIST Changepoint Localization Results
Table[IV](https://arxiv.org/html/2607.26481#A2.T4)summarizes the changepoint localization results on MNIST \(digits3→53\{\\to\}5,n=400n\{=\}400,ξ=250\\xi\{=\}250,α=0\.05\\alpha\{=\}0\.05, horizontal\-flip\-and\-blur contamination withσ=1\.5\\sigma\{=\}1\.5\)\. Each entry reports the average confidence set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|, with empirical coverage shown in parentheses\.
CONCH deteriorates sharply under strong contamination, increasing from1\.001\.00atε=0\\varepsilon\{=\}0to89\.6289\.62atε=0\.7\\varepsilon\{=\}0\.7\. Direct uncertainty\-based weighting via W\-CONCH substantially reduces the confidence set size, reaching6\.886\.88\(Hard\) and6\.936\.93\(Soft\) atε=0\.7\\varepsilon\{=\}0\.7\. Meta\-learning yields further improvements\. Atε=0\.7\\varepsilon\{=\}0\.7, MW\-CONCH \(EDL\-F\) achieves a confidence set size of4\.604\.60, while MW\-CONCH \(EDL\-M\) reduces it further to2\.312\.31\. The MC Dropout variants are similarly effective, with MW\-CONCH \(MC\-F\) and MW\-CONCH \(MC\-M\) achieving confidence set sizes of2\.462\.46and3\.233\.23, respectively\. Coverage remains close to the nominal level1−α=0\.951\-\\alpha=0\.95across all contamination levels, indicating that the proposed weighting schemes preserve the conformal coverage guarantee while substantially improving efficiency\.
TABLE IV:MNIST changepoint localization results \(n=400n\{=\}400,ξ=250\\xi\{=\}250,α=0\.05\\alpha\{=\}0\.05,400400split permutations,200200test tasks\)\. Each entry reports\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|\(coverage\)\. W\-CONCH \(Hard, Soft\), and MW\-CONCH \(EDL\-F\) useβ=ε\\beta\{=\}\\varepsilon;ε=0\\varepsilon\{=\}0is filled with CONCH sinceβ=0\\beta\{=\}0is undefined\. MW\-CONCH \(EDL\-M\) and MW\-CONCH \(MC\-M\) are trained on mixed contamination levels withβ=0\.3\\beta\{=\}0\.3\.
#### B\-C2MNIST Root\-Cause Localization Results
We provide detailed numerical results for the W\-CROC root\-cause localization experiments on MNIST\. Each entry reports the average confidence set size\|𝒦α\(𝐗\)\|\|\\mathcal\{K\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|, empirical coverage, and the per\-stream root\-causepp\-valuesp\(1\),…,p\(5\)p\_\{\(1\)\},\\ldots,p\_\{\(5\)\}, where stream1 is the true root stream\. A streamkkis included in the confidence set wheneverp\(k\)\>αp\_\{\(k\)\}\>\\alpha\.
The multi\-stream MNIST benchmark consists ofD=5D=5streams, corresponding to the binary digit shifts\(2→5\)\(2\{\\to\}5\),\(1→7\)\(1\{\\to\}7\),\(3→8\)\(3\{\\to\}8\),\(0→6\)\(0\{\\to\}6\), and\(4→9\)\(4\{\\to\}9\)\. Stream1 is the root\-cause stream with changepoint att1=150t\_\{1\}=150, while the remaining streams change attd=152t\_\{d\}=152ford=2,…,5d=2,\\ldots,5\. Each stream containsn=400n=400observations\. Each observation is independently contaminated with probabilityε\\varepsilonusing a horizontal\-flip\-and\-blur transformation with Gaussian blur standard deviationσ=1\.5\\sigma=1\.5\.
Table[V](https://arxiv.org/html/2607.26481#A2.T5)reports the detailed results\. Similar to the CIFAR\-100 benchmark, the confidence set size produced by CROC increases substantially as the contamination level grows, reaching4\.8854\.885atε=0\.7\\varepsilon=0\.7\. The Genie reference maintains confidence set sizes close to one across all contamination levels, indicating that root\-cause localization remains feasible when contamination is properly handled\. W\-CROC \(Hard\) provides moderate improvements over CROC, while the meta\-learned variants further improve localization performance\. In particular, MW\-CROC \(EDL\-M\) achieves the smallest confidence set sizes at moderate contamination levels \(ε=0\.3\\varepsilon=0\.3and0\.50\.5\), whereas MW\-CROC \(EDL\-F\) remains more effective under severe contamination \(ε=0\.7\\varepsilon=0\.7\)\. Across all methods, empirical coverage remains close to the nominal level1−α=0\.991\-\\alpha=0\.99\.
TABLE V:W\-CROC results on MNIST \(hflip\+blurσ=1\.5\\sigma\{=\}1\.5,D=5D\{=\}5streams,n=400n\{=\}400,α=0\.01\\alpha\{=\}0\.01,100100split permutations,200200test tasks\)\. W\-CROC \(Hard\) and MW\-CROC \(EDL\-F\) useβ=ε\\beta\{=\}\\varepsilon;ε=0\\varepsilon\{=\}0is filled with CROC\. MW\-CROC \(EDL\-M\) is trained on mixed contamination levels withβ=0\.5\\beta\{=\}0\.5\.
### B\-DAdditional Coverage Results
Table[VI](https://arxiv.org/html/2607.26481#A2.T6)reports the effect of the miscoverage levelα\\alphaon MW\-CONCH \(EDL\-M\) across DomainNet, CIFAR\-100, and MNIST\. As expected, increasingα\\alphaproduces smaller confidence sets at the cost of lower empirical coverage\. Across all datasets and contamination levels, the empirical coverage generally remains close to the nominal target level1−α1\-\\alpha, supporting the finite\-sample validity of the proposed conformal procedure\. These results illustrate the expected trade\-off between confidence set size and coverage controlled by the choice ofα\\alpha\.
TABLE VI:Effect of miscoverage levelα\\alphaon MW\-CONCH \(EDL\-M\)\. Each cell reports the average confidence set size\|𝒞α\(𝐗\)\|\|\\mathcal\{C\}\_\{\\alpha\}\(\\mathbf\{X\}\)\|with empirical coverage in parentheses\. Theoretical guarantee: coverage≥1−α\\geq 1\-\\alpha\.
## References
- \[1\]\(2017\)A survey of methods for time series change point detection\.Knowl\. Inf\. Syst\.51\(2\),pp\. 339–367\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1),[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1),[§I](https://arxiv.org/html/2607.26481#S1.p1.1)\.
- \[2\]R\. F\. Barber, E\. J\. Candès, A\. Ramdas, and R\. J\. Tibshirani\(2023\)Conformal prediction beyond exchangeability\.Ann\. Statist\.51\(2\),pp\. 816–845\.Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[3\]G\. Barlacchi, M\. De Nadai, R\. Larcher, A\. Casella, C\. Chitic, G\. Torrisi, F\. Antonelli, A\. Vespignani, A\. Pentland, and B\. Lepri\(2015\)A multi\-source dataset of urban life in the city of Milan and the province of Trentino\.Sci\. Data2\(1\),pp\. 150055\.Cited by:[§VI](https://arxiv.org/html/2607.26481#S6.p1.1)\.
- \[4\]L\. Chen, S\. T\. Jose, I\. Nikoloska, S\. Park, T\. Chen, and O\. Simeone\(2023\)Learning with limited samples: meta\-learning and applications to communication systems\.Found\. Trends Signal Process\.17\(2\),pp\. 79–208\.Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1)\.
- \[5\]K\. Cohen and Q\. Zhao\(2015\)Asymptotically optimal anomaly detection via sequential testing\.IEEE Trans\. Signal Process\.63\(11\),pp\. 2929–2941\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1)\.
- \[6\]S\. Dandapanthula and A\. Ramdas\(2025\)Offline changepoint localization using a matrix of conformal p\-values\.arXiv preprint arXiv:2505\.00292\.Cited by:[§IV\-A](https://arxiv.org/html/2607.26481#S4.SS1.p1.1)\.
- \[7\]L\. Deng\(2012\)The mnist database of handwritten digit images for machine learning research \[best of the web\]\.IEEE Signal Process\. Mag\.29\(6\),pp\. 141–142\.Cited by:[§B\-C](https://arxiv.org/html/2607.26481#A2.SS3.p1.1)\.
- \[8\]C\. Finn, P\. Abbeel, and S\. Levine\(2017\)Model\-agnostic meta\-learning for fast adaptation of deep networks\.InProc\. Int\. Conf\. Mach\. Learn\. \(ICML\),pp\. 1126–1135\.Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1)\.
- \[9\]Y\. Gal and Z\. Ghahramani\(2016\)Dropout as a bayesian approximation: representing model uncertainty in deep learning\.InProc\. Int\. Conf\. Mach\. Learn\. \(ICML\),pp\. 1050–1059\.Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1),[§IV\-B](https://arxiv.org/html/2607.26481#S4.SS2.p2.9)\.
- \[10\]C\. R\. Genovese, K\. Roeder, and L\. Wasserman\(2006\)False discovery control with p\-value weighting\.Biometrika93\(3\),pp\. 509–524\.Cited by:[§IV\-E](https://arxiv.org/html/2607.26481#S4.SS5.p2.1)\.
- \[11\]I\. Gibbs and E\. J\. Candès\(2021\)Adaptive conformal inference under distribution shift\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[12\]F\. R\. Hampel, E\. M\. Ronchetti, P\. J\. Rousseeuw, and W\. A\. Stahel\(1986\)Robust statistics: the approach based on influence functions\.John Wiley & Sons,New York\.External Links:ISBN 0\-471\-73577\-9Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1)\.
- \[13\]K\. He, X\. Zhang, S\. Ren, and J\. Sun\(2016\)Deep residual learning for image recognition\.InProc\. IEEE Conf\. Comput\. Vis\. Pattern Recognit\. \(CVPR\),pp\. 770–778\.Cited by:[§B\-A](https://arxiv.org/html/2607.26481#A2.SS1.p1.17)\.
- \[14\]R\. Hore and A\. Ramdas\(2026\)Conformal changepoint localization\.arXiv preprint arXiv:2602\.06267\.Cited by:[§A\-D](https://arxiv.org/html/2607.26481#A1.SS4.p3.12),[1st item](https://arxiv.org/html/2607.26481#S1.I1.i1.p1.1),[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p4.1),[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1),[§II\-A](https://arxiv.org/html/2607.26481#S2.SS1.p1.4),[§II](https://arxiv.org/html/2607.26481#S2.p1.1),[§III\-A](https://arxiv.org/html/2607.26481#S3.SS1.p1.4),[§III\-B](https://arxiv.org/html/2607.26481#S3.SS2.2.p1.1),[§III\-B](https://arxiv.org/html/2607.26481#S3.SS2.p1.3),[§III](https://arxiv.org/html/2607.26481#S3.p1.1),[§IV\-A1](https://arxiv.org/html/2607.26481#S4.SS1.SSS1.1.p1.1),[§IV\-A2](https://arxiv.org/html/2607.26481#S4.SS1.SSS2.p2.1),[§IV\-A3](https://arxiv.org/html/2607.26481#S4.SS1.SSS3.p1.11),[§IV\-A3](https://arxiv.org/html/2607.26481#S4.SS1.SSS3.p2.2),[§IV\-A](https://arxiv.org/html/2607.26481#S4.SS1.p1.1),[§IV\-C](https://arxiv.org/html/2607.26481#S4.SS3.1.p1.1),[Figure 4](https://arxiv.org/html/2607.26481#S6.F4),[Figure 5](https://arxiv.org/html/2607.26481#S6.F5),[1st item](https://arxiv.org/html/2607.26481#S6.I1.i1.p1.1),[§VII](https://arxiv.org/html/2607.26481#S7.p1.1),[Assumption 1](https://arxiv.org/html/2607.26481#Thmassumption1)\.
- \[15\]R\. Hore and A\. Ramdas\(2026\)Distribution\-free root cause analysis\.arXiv preprint arXiv:2605\.21627\.Cited by:[3rd item](https://arxiv.org/html/2607.26481#S1.I1.i3.p1.1),[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p4.1),[§I\-B2](https://arxiv.org/html/2607.26481#S1.SS2.SSS2.p1.1),[§II\-B](https://arxiv.org/html/2607.26481#S2.SS2.p1.7),[§II\-B](https://arxiv.org/html/2607.26481#S2.SS2.p2.3),[§II](https://arxiv.org/html/2607.26481#S2.p1.1),[§V\-B](https://arxiv.org/html/2607.26481#S5.SS2.1.p1.2),[§V](https://arxiv.org/html/2607.26481#S5.p1.1),[1st item](https://arxiv.org/html/2607.26481#S6.I1.i1.p1.1),[§VII](https://arxiv.org/html/2607.26481#S7.p1.1)\.
- \[16\]J\. Huang, S\. Park, and O\. Simeone\(2025\)Calibrating Bayesian learning via regularization, confidence minimization, and selective inference\.IEEE Trans\. Signal Process\.73,pp\. 4492–4505\.Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1),[§IV\-B](https://arxiv.org/html/2607.26481#S4.SS2.p1.4)\.
- \[17\]P\. J\. Huber\(1964\)Robust estimation of a location parameter\.Ann\. Math\. Statist\.35\(1\),pp\. 73–101\.Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1),[§II\-A](https://arxiv.org/html/2607.26481#S2.SS1.p2.7)\.
- \[18\]A\. D\. Joseph, B\. Nelson, B\. I\. Rubinstein, and J\. Tygar\(2018\)Adversarial machine learning\.Cambridge University Press\.Cited by:[§IV\-A3](https://arxiv.org/html/2607.26481#S4.SS1.SSS3.p1.11),[§IV\-A3](https://arxiv.org/html/2607.26481#S4.SS1.SSS3.p1.14)\.
- \[19\]E\. Khalastchi and M\. Kalech\(2018\)On fault detection and diagnosis in robotic systems\.ACM Comput\. Surv\.51\(1\),pp\. 1–24\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1),[§I\-B2](https://arxiv.org/html/2607.26481#S1.SS2.SSS2.p1.1)\.
- \[20\]H\. Kim and D\. Siegmund\(1989\)The likelihood ratio test for a change\-point in simple linear regression\.Biometrika76\(3\),pp\. 409–423\.Cited by:[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1),[§IV\-A](https://arxiv.org/html/2607.26481#S4.SS1.p1.1)\.
- \[21\]S\. Kiyani, G\. Pappas, A\. Roth, and H\. Hassani\(2025\)Decision theoretic foundations for conformal prediction: optimal uncertainty quantification for risk\-averse agents\.arXiv preprint arXiv:2502\.02561\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p2.2)\.
- \[22\]A\. Krizhevsky\(2009\)Learning multiple layers of features from tiny images\.Technical reportUniversity of Toronto\.Cited by:[§B\-A](https://arxiv.org/html/2607.26481#A2.SS1.p1.17)\.
- \[23\]B\. Lakshminarayanan, A\. Pritzel, and C\. Blundell\(2017\)Simple and scalable predictive uncertainty estimation using deep ensembles\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1),[§IV\-B](https://arxiv.org/html/2607.26481#S4.SS2.p2.9)\.
- \[24\]C\. Lévy\-Leduc and F\. Roueff\(2009\)Detection and localization of change\-points in high\-dimensional network traffic data\.Ann\. Appl\. Stat\.3\(2\),pp\. 637–662\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1)\.
- \[25\]M\. Li and Y\. Yu\(2021\)Adversarially robust change point detection\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),pp\. 22955–22967\.Cited by:[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1)\.
- \[26\]A\. Malinin and M\. Gales\(2018\)Predictive uncertainty estimation via prior networks\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§IV\-B](https://arxiv.org/html/2607.26481#S4.SS2.p1.4)\.
- \[27\]E\. S\. Page\(1955\)A test for a change in a parameter occurring at an unknown point\.Biometrika42\(3/4\),pp\. 523–527\.Cited by:[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1)\.
- \[28\]S\. Park, K\. M\. Cohen, and O\. Simeone\(2023\)Few\-shot calibration of set predictors via meta\-learned cross\-validation\-based conformal prediction\.IEEE Trans\. Pattern Anal\. Mach\. Intell\.46\(1\),pp\. 280–291\.Cited by:[§A\-F](https://arxiv.org/html/2607.26481#A1.SS6.p1.1),[§A\-F](https://arxiv.org/html/2607.26481#A1.SS6.p2.2),[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1),[§IV\-D](https://arxiv.org/html/2607.26481#S4.SS4.p4.1)\.
- \[29\]X\. Peng, Q\. Bai, X\. Xia, Z\. Huang, K\. Saenko, and B\. Wang\(2019\)Moment matching for multi\-source domain adaptation\.InProc\. IEEE/CVF Int\. Conf\. Comput\. Vis\. \(ICCV\),pp\. 1406–1415\.Cited by:[§B\-A](https://arxiv.org/html/2607.26481#A2.SS1.p1.17)\.
- \[30\]A\. N\. Pettitt\(1979\)A non\-parametric approach to the change\-point problem\.J\. R\. Stat\. Soc\. Ser\. C \(Appl\. Stat\.\)28\(2\),pp\. 126–135\.Cited by:[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1),[§IV\-A](https://arxiv.org/html/2607.26481#S4.SS1.p1.1)\.
- \[31\]M\. Polese, L\. Bonati, S\. D’oro, S\. Basagni, and T\. Melodia\(2023\)Understanding O\-RAN: architecture, interfaces, algorithms, security, and research challenges\.IEEE Commun\. Surv\. Tutor\.25\(2\),pp\. 1376–1411\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1)\.
- \[32\]J\. A\. Rice\(2007\)Mathematical statistics and data analysis\.Vol\.371,Thomson/Brooks/Cole Belmont, CA\.Cited by:[§III\-A](https://arxiv.org/html/2607.26481#S3.SS1.p4.9)\.
- \[33\]G\. Romano, I\. A\. Eckley, and P\. Fearnhead\(2023\)A log\-linear nonparametric online changepoint detection algorithm based on functional pruning\.IEEE Trans\. Signal Process\.72,pp\. 594–606\.Cited by:[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1)\.
- \[34\]Y\. Romano, E\. Patterson, and E\. Candes\(2019\)Conformalized quantile regression\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[35\]M\. Sensoy, L\. Kaplan, and M\. Kandemir\(2018\)Evidential deep learning to quantify classification uncertainty\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1),[§IV\-B](https://arxiv.org/html/2607.26481#S4.SS2.p2.9),[§IV\-D](https://arxiv.org/html/2607.26481#S4.SS4.p2.10)\.
- \[36\]G\. Shafer and V\. Vovk\(2008\)A tutorial on conformal prediction\.J\. Mach\. Learn\. Res\.9\(3\)\.Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[37\]O\. Simeone, S\. Park, and M\. Zecchin\(2026\)Conformal calibration: ensuring the reliability of black\-box AI in wireless systems\.IEEE Commun\. Mag\.\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1)\.
- \[38\]O\. Simeone\(2022\)Machine learning for engineers\.Cambridge Univ\. Press\.Cited by:[§A\-B](https://arxiv.org/html/2607.26481#A1.SS2.p2.2),[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1),[§IV\-A2](https://arxiv.org/html/2607.26481#S4.SS1.SSS2.1.p1.1),[§IV\-A3](https://arxiv.org/html/2607.26481#S4.SS1.SSS3.p1.14),[§IV\-B](https://arxiv.org/html/2607.26481#S4.SS2.p2.9)\.
- \[39\]J\. Soldani and A\. Brogi\(2022\)Anomaly detection and failure root cause analysis in \(micro\)service\-based cloud applications: a survey\.ACM Comput\. Surv\.55\(3\),pp\. 1–39\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1),[§I\-B2](https://arxiv.org/html/2607.26481#S1.SS2.SSS2.p1.1),[§I](https://arxiv.org/html/2607.26481#S1.p1.1)\.
- \[40\]M\. Solé, V\. Muntés\-Mulero, A\. I\. Rana, and G\. Estrada\(2017\)Survey on models and techniques for root\-cause analysis\.arXiv preprint arXiv:1701\.08546\.Cited by:[§I\-B2](https://arxiv.org/html/2607.26481#S1.SS2.SSS2.p1.1),[§I](https://arxiv.org/html/2607.26481#S1.p1.1)\.
- \[41\]H\. Song and H\. Chen\(2024\)Practical and powerful kernel\-based change\-point detection\.IEEE Trans\. Signal Process\.72,pp\. 5174–5186\.Cited by:[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1)\.
- \[42\]J\. Steinhardt, P\. W\. W\. Koh, and P\. S\. Liang\(2017\)Certified defenses for data poisoning attacks\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1)\.
- \[43\]D\. Stutz, A\. T\. Cemgil, A\. Doucet,et al\.\(2021\)Learning optimal conformal classifiers\.arXiv preprint arXiv:2110\.09192\.Cited by:[§A\-F](https://arxiv.org/html/2607.26481#A1.SS6.p1.1),[§A\-F](https://arxiv.org/html/2607.26481#A1.SS6.p4.2),[§I\-B4](https://arxiv.org/html/2607.26481#S1.SS2.SSS4.p1.1),[§IV\-D](https://arxiv.org/html/2607.26481#S4.SS4.p4.1)\.
- \[44\]E\. Y\. N\. Tang, Y\. Chen, M\. Li, and Y\. Yu\(2026\)Online change point detection under heavy\-tailedness and contamination\.arXiv preprint arXiv:2606\.09737\.Cited by:[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1)\.
- \[45\]A\. G\. Tartakovsky, B\. L\. Rozovskii, R\. B\. Blažek, and H\. Kim\(2006\)A novel approach to detection of intrusions in computer networks via adaptive sequential and batch\-sequential change\-point detection methods\.IEEE Trans\. Signal Process\.54\(9\),pp\. 3372–3382\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1)\.
- \[46\]R\. J\. Tibshirani, R\. F\. Barber, E\. J\. Candès, and A\. Ramdas\(2019\)Conformal prediction under covariate shift\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[47\]C\. Truong, L\. Oudre, and N\. Vayatis\(2020\)Selective review of offline change point detection methods\.Signal Process\.167,pp\. 107299\.Cited by:[§I\-B1](https://arxiv.org/html/2607.26481#S1.SS2.SSS1.p1.1),[§I](https://arxiv.org/html/2607.26481#S1.p1.1)\.
- \[48\]J\. Van Amersfoort, L\. Smith, Y\. W\. Teh, and Y\. Gal\(2020\)Uncertainty estimation using a single deep deterministic neural network\.InProc\. Int\. Conf\. Mach\. Learn\. \(ICML\),pp\. 9690–9700\.Cited by:[§B\-A](https://arxiv.org/html/2607.26481#A2.SS1.p3.35)\.
- \[49\]V\. Vovk\(2021\)Testing randomness online\.Statist\. Sci\.36\(4\),pp\. 595–611\.Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[50\]V\. Vovk, A\. Gammerman, and C\. Saunders\(1999\)Machine\-learning applications of algorithmic randomness\.InProc\. 16th Int\. Conf\. Mach\. Learn\. \(ICML\),pp\. 444–453\.Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[51\]L\. Wang, C\. Zhang, R\. Ding, Y\. Xu, Q\. Chen, W\. Zou, Q\. Chen, M\. Zhang, X\. Gao, H\. Fan,et al\.\(2023\)Root cause analysis for microservice systems via hierarchical reinforcement learning from human feedback\.InProc\. 29th ACM SIGKDD Conf\. Knowl\. Discovery Data Mining \(KDD\),pp\. 5116–5125\.Cited by:[§I\-B2](https://arxiv.org/html/2607.26481#S1.SS2.SSS2.p1.1)\.
- \[52\]S\. Yoo, S\. Park, P\. Popovski, J\. Kang, and O\. Simeone\(2026\)Calibrating wireless ai via meta\-learned context\-dependent conformal prediction\.IEEE Trans\. Signal Process\.\.Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[53\]M\. Zecchin, S\. Park, O\. Simeone, and F\. Hellström\(2024\)Generalization and informativeness of conformal prediction\.InProc\. IEEE Int\. Symp\. Inf\. Theory \(ISIT\),pp\. 244–249\.Cited by:[§VII](https://arxiv.org/html/2607.26481#S7.p2.1)\.
- \[54\]M\. Zecchin and O\. Simeone\(2024\)Localized adaptive risk control\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),pp\. 8165–8192\.Cited by:[§I\-B3](https://arxiv.org/html/2607.26481#S1.SS2.SSS3.p1.1)\.
- \[55\]S\. Zou, Y\. Liang, H\. V\. Poor, and X\. Shi\(2017\)Nonparametric detection of anomalous data streams\.IEEE Trans\. Signal Process\.65\(21\),pp\. 5785–5797\.Cited by:[§I\-A](https://arxiv.org/html/2607.26481#S1.SS1.p1.1)\.Similar Articles
Online Localized Conformal Prediction
This paper proposes Online Localized Conformal Prediction (OLCP) to address covariate heterogeneity in online learning and time-series settings. It introduces OLCP-Hedge for bandwidth selection and demonstrates valid long-run coverage with narrower prediction sets compared to existing baselines.
Conformal Agent Error Attribution
This paper presents a framework for error attribution in multi-agent systems using conformal prediction, providing statistical guarantees for identifying decisive errors in agent trajectories. The approach enables automated recovery and debugging by isolating errors within contiguous prediction sets.
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
This paper proposes a conformal prediction framework for LLMs that leverages internal representations rather than output-level statistics, introducing Layer-Wise Information (LI) scores as nonconformity measures to improve validity-efficiency trade-offs under distribution shift. The method demonstrates stronger robustness to calibration-deployment mismatch compared to text-level baselines across QA benchmarks.
Localized Anomaly Detection via Differentiable D-vine Copulas
Presents a novel estimation framework for D-vine copulas using gradient-based MLE and beam search for better global fit, and a localized anomaly detection method with uncertainty quantification via conformal prediction.
Reliable Conformal Prediction for Ordinal Classification Using the Ranked Probability Score
Introduces a conformal prediction method for ordinal classification using the ranked probability score as a nonconformity function, producing median-centered contiguous prediction sets and achieving favorable balance between set width and ordinal miscoverage.