Who Wins Where? Conformal Model Comparison for Local Superiority

arXiv cs.LG Papers

Summary

Introduces a conformalized split-sample framework for local model comparison, producing calibrated local best-model maps and finite-sample guarantees for declaring local superiority.

arXiv:2607.29053v1 Announce Type: new Abstract: Standard model comparison is global, aggregating losses across the covariate space to declare a single winner. This can obscure heterogeneous performance, where different models are preferable in different regions. We introduce conformalized local model comparison, a split-sample framework for constructing calibrated local best-model maps. Given a model comparison score, such as the difference between two squared losses, the method uses three disjoint splits to fit competing models, estimate local centers and scales from out-of-sample scores, and conformally calibrate residual uncertainty. At a target point, the procedure declares a local winner only when a one-sided conformal bound excludes a tie, with the score's sign determining the favored model. We prove finite-sample marginal control for one-sided erroneous declarations on the realized future comparison score, establish pointwise consistency of the localized mean-score estimator away from tie boundaries, show that aggregate comparison can disagree sharply with the prevalence of local superiority, and derive a squared-loss bias--variance decomposition that clarifies how model structure affects local wins. Synthetic and real-data experiments show that the method recovers heterogeneous winner regions, abstains under uncertainty, and yields higher conditional gain than global selection.
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:35 AM

# Who Wins Where? Conformal Model Comparison for Local Superiority
Source: [https://arxiv.org/html/2607.29053](https://arxiv.org/html/2607.29053)
Yi Zhou Baishi Li Xuan Yao Ke\-Wei Huang Asian Institute of Digital Finance, National University of Singapore zhouyi02@nus\.edu\.sgbaishili@u\.nus\.edu yaoxuan@nus\.edu\.sgdishkw@nus\.edu\.sg

###### Abstract

Standard model comparison is global, aggregating losses across the covariate space to declare a single winner\. This can obscure heterogeneous performance, where different models are preferable in different regions\. We introduce*conformalized local model comparison*, a split\-sample framework for constructing calibrated local best\-model maps\. Given a model comparison score, such as the difference between two squared losses, the method uses three disjoint splits to fit competing models, estimate local centers and scales from out\-of\-sample scores, and conformally calibrate residual uncertainty\. At a target point, the procedure declares a local winner only when a one\-sided conformal bound excludes a tie, with the score’s sign determining the favored model\. We prove finite\-sample marginal control for one\-sided erroneous declarations on the realized future comparison score, establish pointwise consistency of the localized mean\-score estimator away from tie boundaries, show that aggregate comparison can disagree sharply with the prevalence of local superiority, and derive a squared\-loss bias–variance decomposition that clarifies how model structure affects local wins\. Synthetic and real\-data experiments show that the method recovers heterogeneous winner regions, abstains under uncertainty, and yields higher conditional gain than global selection\. Our code is available at[https://anonymous\.4open\.science/r/Model\-Winner\-submission\-DD01](https://anonymous.4open.science/r/Model-Winner-submission-DD01)\.

## 1Introduction

Model comparison is central in machine learning\. Given two trained predictors, standard practice compares their losses on a validation set, averages those losses across observations, and declares a single global winner\. This workflow is simple and often appropriate when one model dominates uniformly\. But not all practically important comparisons are global\. A structured or theory\-guided model may perform better where its inductive bias aligns with the data\-generating mechanism, while a flexible black\-box model may perform better elsewhere\. Likewise, in expert\-routing settings, different predictors may be preferable in different subspaces of the covariate space\.

This paper studies*conformalized local model comparison*\. Rather than asking which model wins on average over the full distribution, we ask which model is superior near a target covariate valuex0x\_\{0\}\. Our primary estimand is the conditional mean comparison scoreμ​\(x\)=𝔼​\[S​\(X,Y\)∣X=x\]\\mu\(x\)=\\mathbb\{E\}\[S\(X,Y\)\\mid X=x\], such as the difference between two squared losses, where negative values favor modelAAand positive values favor modelBB\. This induces a local best\-model map over the feature space\. A finite\-radius neighborhood average of the mean score provides a stable local smoothing target, and two additional summaries—a local majority notion and a local quantile notion—serve as secondary descriptors of local agreement\. The main statistical development, however, focuses on local expected superiority\.

A second and the main challenge is inference\. Once comparison is made local, the procedure becomes more sensitive to heterogeneity but also more variable\. To address this, we use a three\-way split\. One split fits the competing models, a second split estimates the local score center and scale from out\-of\-sample comparison scores, and a third split calibrates conformal residuals\. A winner is declared only when a one\-sided conformal upper bound falls below zero\. This construction is deliberately modest: it provides finite\-sample*marginal*protection for errors on the*realized future score*, not a finite\-sample confidence statement forμ​\(x0\)\\mu\(x\_\{0\}\)itself\.

The theoretical results are correspondingly focused\. First, a localized estimator based on an independent evaluation sample recovers the correct sign ofμ​\(x0\)\\mu\(x\_\{0\}\)asymptotically away from tie boundaries\. Second, aggregate comparison targets the global mean score and can therefore disagree sharply with the prevalence of local superiority\. Third, under squared loss, the local mean score decomposes into squared\-bias and variance differences across the covariate space, making the framework especially natural for structured\-versus\-flexible comparisons\. In summary, our contributions are: \(i\) a local model\-comparison formulation based on the conditional mean scoreμ​\(x\)\\mu\(x\)and its finite\-radius neighborhood average, together with secondary majority and quantile descriptors; \(ii\) a split\-sample, locally centered conformal comparison rule with finite\-sample marginal control for false local winner declarations on the realized future score; and \(iii\) consistency, global\-local mismatch, and bias–variance results that characterize when and why local winner regions emerge\.

The rest of the paper is organized as follows\. Section 2 reviews the literature\. Section 3 introduces the local targets\. Section 4 presents the conformal procedure\. Section 5 develops the theoretical properties\. Section 6 summarizes the empirical design and results\.

## 2Related Literature

Conformal prediction, localization, and weighting\.Conformal prediction provides finite\-sample, distribution\-free predictive guarantees under exchangeability\[[10](https://arxiv.org/html/2607.29053#bib.bib10),[8](https://arxiv.org/html/2607.29053#bib.bib8),[5](https://arxiv.org/html/2607.29053#bib.bib5)\]and is summarized in the overview ofAngelopoulos and Bates \[[1](https://arxiv.org/html/2607.29053#bib.bib1)\]\. Subsequent work has made conformal methods adaptive to heterogeneity, for example through conformalized quantile regression\[[7](https://arxiv.org/html/2607.29053#bib.bib7)\], weighted conformal prediction under covariate shift\[[9](https://arxiv.org/html/2607.29053#bib.bib9)\], and localized conformal prediction\[[3](https://arxiv.org/html/2607.29053#bib.bib3)\]\. These papers motivate the idea that calibration should respond to local feature\-dependent structure rather than rely on a global residual distribution\. Our paper draws on that intuition, but the inferential object is different: we calibrate a*comparison score between two models*, not a prediction interval for a single model\.

Localized model selection and post\-selection validity\.Recent advances in conformal inference have addressed model selection\. For instance, localized conformal model selection\[[11](https://arxiv.org/html/2607.29053#bib.bib11)\]provides a framework for choosing models that yield efficient conformal intervals while preserving coverage, and\[[6](https://arxiv.org/html/2607.29053#bib.bib6)\]studies predictive validity following data\-dependent model choices\. Our work shares their view that model suitability is heterogeneous, but our objective is complementary\. Rather than optimizing interval length or characterizing post\-selection predictive coverage, we repurpose conformal calibration as evidence for neighborhood\-level superiority under a generic comparison score\.

Model comparison and routing\.The broad motivation also connects to conditional predictive ability testing and local model routing\. Global predictive\-ability tests compare average losses across evaluation samples\[[2](https://arxiv.org/html/2607.29053#bib.bib2)\], while mixture\-of\-experts and algorithm\-selection perspectives emphasize that different models may be preferable in different regions of the input space\[[4](https://arxiv.org/html/2607.29053#bib.bib4)\]\. Our paper fits between these viewpoints: it retains the language of loss\-based model comparison but replaces one\-number global ranking with a local score map over the feature space\.

## 3Problem Setup and Local Superiority Targets

Let\(X,Y\)\(X,Y\)be a random pair taking values in𝒳×ℝ\\mathcal\{X\}\\times\\mathbb\{R\}, whereXXdenotes covariates andYYdenotes the response\. We compare two fitted predictive models,AAandBB, with predictorsf^A,f^B:𝒳→ℝ\\hat\{f\}\_\{A\},\\hat\{f\}\_\{B\}:\\mathcal\{X\}\\to\\mathbb\{R\}\. For each modelm∈\{A,B\}m\\in\\\{A,B\\\}, letLm​\(X,Y\)L\_\{m\}\(X,Y\)denote its loss at\(X,Y\)\(X,Y\)\. To compare the two models, we use a generic score

S​\(X,Y\)≔s​\(LA​\(X,Y\),LB​\(X,Y\);X\),S\(X,Y\)\\coloneqq s\\\!\\bigl\(L\_\{A\}\(X,Y\),L\_\{B\}\(X,Y\);X\\bigr\),\(1\)where negative scores favorAA, positive scores favorBB, and zero denotes a tie\. We assumes​\(ℓA,ℓB;x\)s\(\\ell\_\{A\},\\ell\_\{B\};x\)is weakly increasing inℓA\\ell\_\{A\}and weakly decreasing inℓB\\ell\_\{B\}\. Examples include the gapsgap​\(LA,LB\)=LA−LBs\_\{\\mathrm\{gap\}\}\(L\_\{A\},L\_\{B\}\)=L\_\{A\}\-L\_\{B\}, the log\-ratiosratio​\(LA,LB\)=log⁡\(\(LA\+τ\)/\(LB\+τ\)\)s\_\{\\mathrm\{ratio\}\}\(L\_\{A\},L\_\{B\}\)=\\log\(\(L\_\{A\}\+\\tau\)/\(L\_\{B\}\+\\tau\)\)withτ\>0\\tau\>0, and the standardized gapsstd​\(LA,LB;X\)=\(LA−LB\)/v​\(X\)s\_\{\\mathrm\{std\}\}\(L\_\{A\},L\_\{B\};X\)=\(L\_\{A\}\-L\_\{B\}\)/v\(X\)withv​\(X\)\>0v\(X\)\>0\.

Our primary population target is the conditional mean score

μ​\(x\)≔𝔼​\[S​\(X,Y\)∣X=x\]\.\\mu\(x\)\\coloneqq\\mathbb\{E\}\[S\(X,Y\)\\mid X=x\]\.\(2\)Whenμ​\(x0\)<0\\mu\(x\_\{0\}\)<0, modelAAis preferred; whenμ​\(x0\)\>0\\mu\(x\_\{0\}\)\>0, modelBBis preferred\. To stabilize local estimation, fix a target pointx0∈𝒳x\_\{0\}\\in\\mathcal\{X\}and a locality parameterr\>0r\>0, and letKr:𝒳×𝒳→\[0,∞\)K\_\{r\}:\\mathcal\{X\}\\times\\mathcal\{X\}\\to\[0,\\infty\)be a nonnegative neighborhood weight\. For the asymptotic theory in Section 5, when𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}, we use compactly supported kernel localizers such as the box kernel

Kr​\(x,x0\)=r−d​𝟏​\{‖x−x0‖≤r\},K\_\{r\}\(x,x\_\{0\}\)=r^\{\-d\}\\mathbf\{1\}\\\{\\\|x\-x\_\{0\}\\\|\\leq r\\\},which corresponds to the base kernelK​\(u\)=𝟏​\{‖u‖≤1\}K\(u\)=\\mathbf\{1\}\\\{\\\|u\\\|\\leq 1\\\}, or the Epanechnikov formKr​\(x,x0\)=r−d​\(1−‖x−x0‖2/r2\)\+K\_\{r\}\(x,x\_\{0\}\)=r^\{\-d\}\\bigl\(1\-\\\|x\-x\_\{0\}\\\|^\{2\}/r^\{2\}\\bigr\)\_\{\+\}\. In finite\-sample implementations, KNN weights are a practical alternative when exact compact support is preferred\. The corresponding neighborhood\-average score is

θ​\(x0;r\):=𝔼​\[Kr​\(X,x0\)​S​\(X,Y\)\]𝔼​\[Kr​\(X,x0\)\]=𝔼​\[Kr​\(X,x0\)​μ​\(X\)\]𝔼​\[Kr​\(X,x0\)\]\.\\theta\(x\_\{0\};r\):=\\frac\{\\mathbb\{E\}\\\!\\left\[K\_\{r\}\(X,x\_\{0\}\)S\(X,Y\)\\right\]\}\{\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\]\}=\\frac\{\\mathbb\{E\}\\\!\\left\[K\_\{r\}\(X,x\_\{0\}\)\\mu\(X\)\\right\]\}\{\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\]\}\.\(3\)θ​\(x0;r\)\\theta\(x\_\{0\};r\)is a smoothed version ofμ​\(x0\)\\mu\(x\_\{0\}\); under continuity ofμ\\mu, it approachesμ​\(x0\)\\mu\(x\_\{0\}\)asr→0r\\to 0\.

###### Definition 1\(Local expected winner and best\-model map\)\.

ModelAAis a*local expected winner*atx0x\_\{0\}ifμ​\(x0\)<0\\mu\(x\_\{0\}\)<0\. For a fixed radiusrr, modelAAis the*rr\-local expected winner*aroundx0x\_\{0\}ifθ​\(x0;r\)<0\\theta\(x\_\{0\};r\)<0\. The induced pointwise local best\-model map is

𝒲​\(x\)=A​𝕀\{μ​\(x\)<0\}\+B​𝕀\{μ​\(x\)\>0\}\+tie⋅𝕀\{μ​\(x\)=0\},\\mathcal\{W\}\(x\)=A\\mathbb\{I\}\_\{\\\{\\mu\(x\)<0\\\}\}\+B\\mathbb\{I\}\_\{\\\{\\mu\(x\)\>0\\\}\}\+\\text\{tie\}\\cdot\\mathbb\{I\}\_\{\\\{\\mu\(x\)=0\\\}\},and its finite\-radius analogue is defined by replacingμ​\(x\)\\mu\(x\)withθ​\(x;r\)\\theta\(x;r\)\.

We also record two secondary summaries of local agreement derived from the local score distributionFx0,r​\(t\)≔𝔼​\[Kr​\(X,x0\)​𝟏​\{S​\(X,Y\)≤t\}\]/𝔼​\[Kr​\(X,x0\)\]F\_\{x\_\{0\},r\}\(t\)\\coloneqq\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\\mathbf\{1\}\\\{S\(X,Y\)\\leq t\\\}\]/\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\]fort∈ℝt\\in\\mathbb\{R\}\. The associated local majority score for modelAA, denotedπA​\(x0;r\)\\pi\_\{A\}\(x\_\{0\};r\), and the\(1−β\)\(1\-\\beta\)quantile scoreQ1−β​\(x0;r\)Q\_\{1\-\\beta\}\(x\_\{0\};r\)are

πA​\(x0;r\)≔𝔼​\[Kr​\(X,x0\)​𝟏​\{S​\(X,Y\)<0\}\]𝔼​\[Kr​\(X,x0\)\]andQ1−β​\(x0;r\)≔inf\{t∈ℝ:Fx0,r​\(t\)≥1−β\},\\pi\_\{A\}\(x\_\{0\};r\)\\coloneqq\\frac\{\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\\mathbf\{1\}\\\{S\(X,Y\)<0\\\}\]\}\{\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\]\}\\quad\\text\{and\}\\quad Q\_\{1\-\\beta\}\(x\_\{0\};r\)\\coloneqq\\inf\\\{t\\in\\mathbb\{R\}:F\_\{x\_\{0\},r\}\(t\)\\geq 1\-\\beta\\\},forβ∈\(0,1\)\\beta\\in\(0,1\)\. In our experiments, we define the conditionsπA​\(x0;r\)\>1/2\\pi\_\{A\}\(x\_\{0\};r\)\>1/2andQ1−β​\(x0;r\)<0Q\_\{1\-\\beta\}\(x\_\{0\};r\)<0to belocal\_def2andlocal\_def3, respectively; they indicate that the neighborhood favors modelAAunder local\-majority and local\-quantile summaries\. We uselocal\_def1for the local\-mean rule based on the sign ofθ​\(x0;r\)\\theta\(x\_\{0\};r\)\. The remainder of the paper focuses on the local mean targetμ​\(x\)\\mu\(x\), equivalentlyθ​\(x0;r\)\\theta\(x\_\{0\};r\), because it is directly connected to our estimator and theory\.

## 4Locally Centered Conformal Comparison

We now attach a finite\-sample inferential layer to the comparison scoreS​\(X,Y\)S\(X,Y\)\. Let\{\(Xi,Yi\)\}i=1n\\\{\(X\_\{i\},Y\_\{i\}\)\\\}\_\{i=1\}^\{n\}be an i\.i\.d\. sample, split into three disjoint parts: a fitting splitIfitI\_\{\\mathrm\{fit\}\}used only to train the competing predictorsf^A\\hat\{f\}\_\{A\}andf^B\\hat\{f\}\_\{B\}, an estimation splitIestI\_\{\\mathrm\{est\}\}used only to estimate a local center and scale from*out\-of\-sample*scores, and a calibration splitIcalI\_\{\\mathrm\{cal\}\}used only for conformal calibration\. This separation removes the in\-sample\-loss issue that would arise if the same observations were used to fit the models and to estimate the local comparison score\. Conditional on the models fitted onIfitI\_\{\\mathrm\{fit\}\}, the population targetμ​\(x\)\\mu\(x\)and out\-of\-sample scoresSjS\_\{j\}forj∈Iestj\\in I\_\{\\mathrm\{est\}\}are evaluated according to \([1](https://arxiv.org/html/2607.29053#S3.E1)\) and \([2](https://arxiv.org/html/2607.29053#S3.E2)\)\. LetS¯est\\bar\{S\}\_\{\\mathrm\{est\}\}andV^est\\widehat\{V\}\_\{\\mathrm\{est\}\}denote their empirical mean and uncorrected variance; the latter avoids Bessel corrections as the conformal guarantee requires only a strictly positive scale\.

For a target pointxx, letWx≔∑j∈IestKr​\(Xj,x\)W\_\{x\}\\coloneqq\\sum\_\{j\\in I\_\{\\mathrm\{est\}\}\}K\_\{r\}\(X\_\{j\},x\)\. WhenWx\>0W\_\{x\}\>0, we define the localized center and regularized scale, for a regularization parameterλ\>0\\lambda\>0, as

θ^r​\(x\)≔∑j∈IestKr​\(Xj,x\)Wx​Sjandσ^r​\(x\)≔\[λ\+∑j∈IestKr​\(Xj,x\)Wx​\(Sj−θ^r​\(x\)\)2\]1/2\.\\widehat\{\\theta\}\_\{r\}\(x\)\\coloneqq\\sum\_\{j\\in I\_\{\\mathrm\{est\}\}\}\\frac\{K\_\{r\}\(X\_\{j\},x\)\}\{W\_\{x\}\}S\_\{j\}\\quad\\text\{and\}\\quad\\widehat\{\\sigma\}\_\{r\}\(x\)\\coloneqq\\Biggl\[\\lambda\+\\sum\_\{j\\in I\_\{\\mathrm\{est\}\}\}\\frac\{K\_\{r\}\(X\_\{j\},x\)\}\{W\_\{x\}\}\\bigl\(S\_\{j\}\-\\widehat\{\\theta\}\_\{r\}\(x\)\\bigr\)^\{2\}\\Biggr\]^\{1/2\}\.IfWx=0W\_\{x\}=0, we fall back to the global estimatesθ^r​\(x\)≔S¯est\\widehat\{\\theta\}\_\{r\}\(x\)\\coloneqq\\bar\{S\}\_\{\\mathrm\{est\}\}andσ^r​\(x\)≔\(V^est\+λ\)1/2\\widehat\{\\sigma\}\_\{r\}\(x\)\\coloneqq\(\\widehat\{V\}\_\{\\mathrm\{est\}\}\+\\lambda\)^\{1/2\}\. The unstandardized procedure is recovered by settingσ^r≡1\\widehat\{\\sigma\}\_\{r\}\\equiv 1\.

On the calibration splitIcalI\_\{\\mathrm\{cal\}\}of sizencaln\_\{\\mathrm\{cal\}\}, we compute the scoresSiS\_\{i\}analogously and form the localized residualsRi≔\(Si−θ^r​\(Xi\)\)/σ^r​\(Xi\)R\_\{i\}\\coloneqq\(S\_\{i\}\-\\widehat\{\\theta\}\_\{r\}\(X\_\{i\}\)\)/\\widehat\{\\sigma\}\_\{r\}\(X\_\{i\}\)\. LetR\(1\)≤⋯≤R\(ncal\)R\_\{\(1\)\}\\leq\\dots\\leq R\_\{\(n\_\{\\mathrm\{cal\}\}\)\}be the sorted residuals\. Settingkα≔⌈\(ncal\+1\)​\(1−α\)⌉k\_\{\\alpha\}\\coloneqq\\lceil\(n\_\{\\mathrm\{cal\}\}\+1\)\(1\-\\alpha\)\\rceil, the conformal threshold isq^1−α≔R\(kα\)\\hat\{q\}\_\{1\-\\alpha\}\\coloneqq R\_\{\(k\_\{\\alpha\}\)\}withq^1−α≔∞\\hat\{q\}\_\{1\-\\alpha\}\\coloneqq\\inftyifkα\>ncalk\_\{\\alpha\}\>n\_\{\\mathrm\{cal\}\}\.

For a test covariatex0∈𝒳x\_\{0\}\\in\\mathcal\{X\}, the conformal upper bound isU^α​\(x0\)≔θ^r​\(x0\)\+σ^r​\(x0\)​q^1−α\\widehat\{U\}\_\{\\alpha\}\(x\_\{0\}\)\\coloneqq\\widehat\{\\theta\}\_\{r\}\(x\_\{0\}\)\+\\widehat\{\\sigma\}\_\{r\}\(x\_\{0\}\)\\hat\{q\}\_\{1\-\\alpha\}\. We declare modelAAthe local winner whenU^α​\(x0\)<0\\widehat\{U\}\_\{\\alpha\}\(x\_\{0\}\)<0\. This one\-sided rule exclusively controls false declarations in favor ofAA; a symmetric lower\-bound declaration for modelBBfollows analogously\.

###### Proposition 1\(Finite\-sample marginal control for the realized score\)\.

Let\(Xn\+1,Yn\+1\)\(X\_\{n\+1\},Y\_\{n\+1\}\)be an independent test point\. Under the assumption of exchangeability conditional on the training and estimation splits, and provided thatθ^r\\widehat\{\\theta\}\_\{r\}andσ^r\\widehat\{\\sigma\}\_\{r\}are independent of the test data, the realized scoreSn\+1≔s​\(LA​\(Xn\+1,Yn\+1\),LB​\(Xn\+1,Yn\+1\);Xn\+1\)S\_\{n\+1\}\\coloneqq s\(L\_\{A\}\(X\_\{n\+1\},Y\_\{n\+1\}\),L\_\{B\}\(X\_\{n\+1\},Y\_\{n\+1\}\);X\_\{n\+1\}\)satisfies

ℙ​\(Sn\+1≤U^α​\(Xn\+1\)\)≥1−αandℙ​\(U^α​\(Xn\+1\)<0,Sn\+1≥0\)≤α\.\\mathbb\{P\}\\bigl\(S\_\{n\+1\}\\leq\\widehat\{U\}\_\{\\alpha\}\(X\_\{n\+1\}\)\\bigr\)\\geq 1\-\\alpha\\quad\\text\{and\}\\quad\\mathbb\{P\}\\bigl\(\\widehat\{U\}\_\{\\alpha\}\(X\_\{n\+1\}\)<0,\\ S\_\{n\+1\}\\geq 0\\bigr\)\\leq\\alpha\.

Proposition[1](https://arxiv.org/html/2607.29053#Thmproposition1)is intentionally phrased for the*realized future score*\. It does not provide a finite\-sample confidence interval or hypothesis test forμ​\(x0\)\\mu\(x\_\{0\}\)orθ​\(x0;r\)\\theta\(x\_\{0\};r\)\. Instead, it guarantees that the one\-sided conformal rule makes false winner declarations for the realized future score at rate at mostα\\alphamarginally over the test point\. Because the conformal guarantee relies exclusively on the exchangeability of the residuals, Proposition[1](https://arxiv.org/html/2607.29053#Thmproposition1)remains strictly valid for any functional choice of the comparison score \(e\.g\., gap, log\-ratio, or standardized\) and any localized centering heuristic, including estimators targeting the majority or quantile summaries defined in Section 3\.

For a fixedx0x\_\{0\}, the same construction yields the one\-sided conformalpp\-valuep^​\(x0\)≔\(1\+∑i∈Ical𝟏​\{Ri≥−θ^r​\(x0\)/σ^r​\(x0\)\}\)/\(\|Ical\|\+1\)\\hat\{p\}\(x\_\{0\}\)\\coloneqq\\bigl\(1\+\\sum\_\{i\\in I\_\{\\mathrm\{cal\}\}\}\\mathbf\{1\}\\\{R\_\{i\}\\geq\-\\widehat\{\\theta\}\_\{r\}\(x\_\{0\}\)/\\widehat\{\\sigma\}\_\{r\}\(x\_\{0\}\)\\\}\\bigr\)/\(\|I\_\{\\mathrm\{cal\}\}\|\+1\), settingσ^r​\(x0\)≡1\\widehat\{\\sigma\}\_\{r\}\(x\_\{0\}\)\\equiv 1if unstandardized\. Thispp\-value is a slightly less conservative companion to the upper\-bound ruleU^α​\(x0\)<0\\widehat\{U\}\_\{\\alpha\}\(x\_\{0\}\)<0: under continuous residuals the two summaries differ by at most one calibration rank due to empirical\-quantile conventions, and ties at the threshold—e\.g\., for discrete scores—preserve this small discrepancy\.

## 5Theoretical Properties

For the first two results, we condition on the predictors produced by the independent fitting split\. This isolates the test\-time randomness: the evaluation observations remain i\.i\.d\., and the conditional meanμ​\(x\)\\mu\(x\)from \([2](https://arxiv.org/html/2607.29053#S3.E2)\) is unambiguously defined\. Section[5\.3](https://arxiv.org/html/2607.29053#S5.SS3)subsequently reintroduces training randomness to interpret an ex\-ante version of the expected local score under squared loss\.

### 5\.1Asymptotic correctness for local expected superiority

We first study whether a localized estimator recovers the correct sign ofμ​\(x0\)\\mu\(x\_\{0\}\)at a target covariatex0x\_\{0\}\. Using an*independent evaluation sample*and a bandwidth sequencern→0r\_\{n\}\\to 0, we define the population\-analogue of the estimation\-split estimator from Section 4:

μ^n​\(x0\)≔∑i=1nKrn​\(Xi,x0\)​Si∑i=1nKrn​\(Xi,x0\),\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)\\coloneqq\\frac\{\\sum\_\{i=1\}^\{n\}K\_\{r\_\{n\}\}\(X\_\{i\},x\_\{0\}\)S\_\{i\}\}\{\\sum\_\{i=1\}^\{n\}K\_\{r\_\{n\}\}\(X\_\{i\},x\_\{0\}\)\},\(4\)whereSiS\_\{i\}is the realized score \([1](https://arxiv.org/html/2607.29053#S3.E1)\) for theii\-th observation\. Let the true local winner beδ⋆​\(x0\)≔A\\delta^\{\\star\}\(x\_\{0\}\)\\coloneqq Aifμ​\(x0\)<0\\mu\(x\_\{0\}\)<0andBBifμ​\(x0\)\>0\\mu\(x\_\{0\}\)\>0\. Its empirical counterpart,δ^n​\(x0\)\\hat\{\\delta\}\_\{n\}\(x\_\{0\}\), is assigned toAAifμ^n​\(x0\)<0\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)<0and toBBotherwise\. Breaking estimated ties \(μ^n​\(x0\)=0\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)=0\) in favor ofBBis a convention that differs from the population map𝒲\\mathcal\{W\}only on measure\-zero events for continuous scores, leaving asymptotic sign recovery unaffected\. Furthermore, because sign recovery is inherently unidentifiable on the exact population tie boundary \(μ​\(x0\)=0\\mu\(x\_\{0\}\)=0\), the following theorem focuses strictly on the caseμ​\(x0\)≠0\\mu\(x\_\{0\}\)\\neq 0\.

###### Theorem 1\(Asymptotic correctness of local winner estimation\)\.

Fix an interior pointx0∈𝒳x\_\{0\}\\in\\mathcal\{X\}\. Assume: \(1\)X∈ℝdX\\in\\mathbb\{R\}^\{d\}has a densitypXp\_\{X\}that is continuous atx0x\_\{0\}withpX​\(x0\)\>0p\_\{X\}\(x\_\{0\}\)\>0; \(2\) the localizer isKr​\(x,x0\)=r−d​K​\(\(x−x0\)/r\)K\_\{r\}\(x,x\_\{0\}\)=r^\{\-d\}K\(\(x\-x\_\{0\}\)/r\), whereKKis bounded, nonnegative, compactly supported, and satisfies∫ℝdK​\(u\)​𝑑u\>0\\int\_\{\\mathbb\{R\}^\{d\}\}K\(u\)\\,du\>0; \(3\)μ​\(x\)\\mu\(x\)is continuous atx0x\_\{0\}; \(4\) there exists a neighborhoodU​\(x0\)U\(x\_\{0\}\)ofx0x\_\{0\}such thatsupx∈U​\(x0\)𝔼​\[S​\(X,Y\)2∣X=x\]<∞\\sup\_\{x\\in U\(x\_\{0\}\)\}\\mathbb\{E\}\[S\(X,Y\)^\{2\}\\mid X=x\]<\\infty; and \(5\)rn→0r\_\{n\}\\to 0andn​rnd→∞nr\_\{n\}^\{d\}\\to\\infty\. Thenμ^n​\(x0\)​⟶𝑝​μ​\(x0\)\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)\\overset\{p\}\{\\longrightarrow\}\\mu\(x\_\{0\}\)\. Consequently, ifμ​\(x0\)≠0\\mu\(x\_\{0\}\)\\neq 0, thenℙ​\(δ^n​\(x0\)=δ⋆​\(x0\)\)⟶1\\mathbb\{P\}\(\\hat\{\\delta\}\_\{n\}\(x\_\{0\}\)=\\delta^\{\\star\}\(x\_\{0\}\)\)\\longrightarrow 1\. The same applies to the estimation\-split estimator in Section 4 withnnreplaced by\|Iest\|\|I\_\{\\mathrm\{est\}\}\|\.

### 5\.2Aggregate comparison targets the global mean, not local prevalence

Define the global mean scoreμglob≔𝔼​\[S​\(X,Y\)\]\\mu\_\{\\mathrm\{glob\}\}\\coloneqq\\mathbb\{E\}\[S\(X,Y\)\]and the aggregate empirical scoreS¯n≔1n​∑i=1nSi\\bar\{S\}\_\{n\}\\coloneqq\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}S\_\{i\}\. The aggregate selector choosesδ^nagg≔A\\hat\{\\delta\}\_\{n\}^\{\\mathrm\{agg\}\}\\coloneqq AifS¯n<0\\bar\{S\}\_\{n\}<0andBBotherwise\. In contrast, the prevalence of local superiority for modelAAisρA≔ℙ​\(μ​\(X\)<0\)\\rho\_\{A\}\\coloneqq\\mathbb\{P\}\(\\mu\(X\)<0\), which depends exclusively on the sign of the local conditional mean score defined in \([2](https://arxiv.org/html/2607.29053#S3.E2)\)\.

###### Proposition 2\(Aggregate selection can disagree with local superiority prevalence\)\.

Suppose𝔼​\[\|S​\(X,Y\)\|\]<∞\\mathbb\{E\}\[\|S\(X,Y\)\|\]<\\infty\. Let the true aggregate winner beδagg,⋆≔A\\delta^\{\\mathrm\{agg\},\\star\}\\coloneqq Aifμglob<0\\mu\_\{\\mathrm\{glob\}\}<0andBBifμglob\>0\\mu\_\{\\mathrm\{glob\}\}\>0\. Ifμglob≠0\\mu\_\{\\mathrm\{glob\}\}\\neq 0, thenℙ​\(δ^nagg=δagg,⋆\)⟶1\\mathbb\{P\}\(\\hat\{\\delta\}\_\{n\}^\{\\mathrm\{agg\}\}=\\delta^\{\\mathrm\{agg\},\\star\}\)\\longrightarrow 1\. Moreover, for anyq∈\(1/2,1\)q\\in\(1/2,1\), there exists a joint distribution of\(X,S\)\(X,S\)such thatℙ​\(μ​\(X\)<0\)≥q\\mathbb\{P\}\(\\mu\(X\)<0\)\\geq qbut𝔼​\[S\]\>0\\mathbb\{E\}\[S\]\>0\. In such cases,ℙ​\(δ^nagg=B\)⟶1\\mathbb\{P\}\(\\hat\{\\delta\}\_\{n\}^\{\\mathrm\{agg\}\}=B\)\\longrightarrow 1even though modelAAis locally superior on at least aqq\-fraction of the covariate distribution\.

Proposition[2](https://arxiv.org/html/2607.29053#Thmproposition2)clarifies an elementary but practically profound point routinely obscured in standard model evaluation: the global winner has no mathematical binding to the proportion of local winners\. Aggregate averaging evaluates the net magnitude of the global mean scoreμglob\\mu\_\{\\mathrm\{glob\}\}, allowing massive losses in a small covariate region to dominate the sum\. Conversely, local prevalence depends strictly on the spatial frequency of the expected score’s sign\. Consequently, global selection can completely reject a model even if it is the strictly superior choice for99%99\\%of the population\.

### 5\.3Ex\-Ante Bias\-Variance Interpretation for Squared Loss

While the preceding sections condition on fitted predictors to evaluate superiority, we now reintroduce the randomness of the training data to understand*why*heterogeneous regions of superiority emerge\. By analyzing an ex\-ante expected local score over both the training procedure and the test distribution, we can mechanically trace local wins back to standard learning\-theoretic properties\.

###### Proposition 3\(Squared\-loss local bias–variance decomposition\)\.

AssumeY=f⋆​\(X\)\+εY=f^\{\\star\}\(X\)\+\\varepsilonwith𝔼​\[ε∣X\]=0\\mathbb\{E\}\[\\varepsilon\\mid X\]=0, and assume the training sample𝒟\\mathcal\{D\}used to fit the predictors is independent of the fresh test pair\(X,Y\)\(X,Y\)\. Let the training\-random squared\-loss gap beS𝒟​\(X,Y\)≔\(f^A​\(X\)−Y\)2−\(f^B​\(X\)−Y\)2S\_\{\\mathcal\{D\}\}\(X,Y\)\\coloneqq\(\\hat\{f\}\_\{A\}\(X\)\-Y\)^\{2\}\-\(\\hat\{f\}\_\{B\}\(X\)\-Y\)^\{2\}, where the fitted predictors are trained on a random training sample𝒟\\mathcal\{D\}\. Form∈\{A,B\}m\\in\\\{A,B\\\}, defineBiasm​\(x\)≔𝔼𝒟​\[f^m​\(x\)\]−f⋆​\(x\)\\mathrm\{Bias\}\_\{m\}\(x\)\\coloneqq\\mathbb\{E\}\_\{\\mathcal\{D\}\}\[\\hat\{f\}\_\{m\}\(x\)\]\-f^\{\\star\}\(x\)andVarm​\(x\)≔Var𝒟​\(f^m​\(x\)\)\\mathrm\{Var\}\_\{m\}\(x\)\\coloneqq\\mathrm\{Var\}\_\{\\mathcal\{D\}\}\(\\hat\{f\}\_\{m\}\(x\)\), where expectations are taken over𝒟\\mathcal\{D\}\. If𝔼𝒟​\[f^m​\(x\)2\]<∞\\mathbb\{E\}\_\{\\mathcal\{D\}\}\[\\hat\{f\}\_\{m\}\(x\)^\{2\}\]<\\inftyand𝔼​\[ε2∣X=x\]<∞\\mathbb\{E\}\[\\varepsilon^\{2\}\\mid X=x\]<\\infty, then the ex\-ante conditional mean scoreμex​\(x\)≔𝔼𝒟,Y∣X=x​\[S𝒟​\(X,Y\)∣X=x\]\\mu\_\{\\mathrm\{ex\}\}\(x\)\\coloneqq\\mathbb\{E\}\_\{\\mathcal\{D\},Y\\mid X=x\}\[S\_\{\\mathcal\{D\}\}\(X,Y\)\\mid X=x\]satisfies

μex​\(x\)=BiasA2​\(x\)−BiasB2​\(x\)\+VarA​\(x\)−VarB​\(x\)\.\\mu\_\{\\mathrm\{ex\}\}\(x\)=\\mathrm\{Bias\}\_\{A\}^\{2\}\(x\)\-\\mathrm\{Bias\}\_\{B\}^\{2\}\(x\)\+\\mathrm\{Var\}\_\{A\}\(x\)\-\\mathrm\{Var\}\_\{B\}\(x\)\.\(5\)Equivalently, the ex\-ante smoothed neighborhood scoreθex​\(x0;r\)≔𝔼​\[Kr​\(X,x0\)​μex​\(X\)\]/𝔼​\[Kr​\(X,x0\)\]\\theta\_\{\\mathrm\{ex\}\}\(x\_\{0\};r\)\\coloneqq\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\\mu\_\{\\mathrm\{ex\}\}\(X\)\]/\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\]satisfies

θex​\(x0;r\)=𝔼​\[Kr​\(X,x0\)​\{BiasA2​\(X\)−BiasB2​\(X\)\+VarA​\(X\)−VarB​\(X\)\}\]𝔼​\[Kr​\(X,x0\)\]\.\\theta\_\{\\mathrm\{ex\}\}\(x\_\{0\};r\)=\\frac\{\\mathbb\{E\}\\bigl\[K\_\{r\}\(X,x\_\{0\}\)\\bigl\\\{\\mathrm\{Bias\}\_\{A\}^\{2\}\(X\)\-\\mathrm\{Bias\}\_\{B\}^\{2\}\(X\)\+\\mathrm\{Var\}\_\{A\}\(X\)\-\\mathrm\{Var\}\_\{B\}\(X\)\\bigr\\\}\\bigr\]\}\{\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\]\}\.\(6\)

Although mathematically elementary, Proposition[3](https://arxiv.org/html/2607.29053#Thmproposition3)provides the fundamental explanatory mechanism for local conformal comparison\. Because irreducible noise cancels out, the ex\-ante local mean score reduces strictly to a competition between squared\-bias and variance differences\. This decomposition formalizes the spatial trade\-off between a structured, high\-bias/low\-variance model \(e\.g\., OLS\) and a flexible, low\-bias/high\-variance model \(e\.g\., Gradient Boosting\)\. In smooth regions where the data generating mechanism aligns with the structured model’s inductive bias, its variance advantage dictates local superiority\. Conversely, in highly non\-linear regions, the flexible model’s localized bias reduction overcomes its higher variance, forcing a flip in the local winner map\. We empirically validate this exact spatial mechanism in Section 6\.

## 6Experiments

We validate our approach on synthetic and real\-world datasets\. The synthetic experiments serve three purposes\. First, they test whether local comparison can recover heterogeneous winner regions, as formalized in Proposition[2](https://arxiv.org/html/2607.29053#Thmproposition2)\. Second, they test whether conformal calibration converts local evidence into statistically reliable winner declarations, as guaranteed by Proposition[1](https://arxiv.org/html/2607.29053#Thmproposition1)\. Third, they provide diagnostics for the asymptotic and bias–variance mechanisms in Theorem[1](https://arxiv.org/html/2607.29053#Thmtheorem1)and Proposition[3](https://arxiv.org/html/2607.29053#Thmproposition3)\. The real\-data experiments, reported in Section[6\.2](https://arxiv.org/html/2607.29053#S6.SS2), evaluate whether the conformal declarations select regions with positive realized gain when the oracle local winner map is unavailable\.

### 6\.1Synthetic data experiments

We use five one\-dimensional synthetic designs\. This setting makes local winner regions visually inspectable and allows us to evaluate against an oracle local comparison target\. In all designs,X∼Unif​\[0,1\]X\\sim\\mathrm\{Unif\}\[0,1\]andY=f⋆​\(X\)\+εY=f^\{\\star\}\(X\)\+\\varepsilon\. ModelAAis a gradient\-boosted tree, implemented byLGBMRegressorwith500500estimators, learning rate0\.050\.05, and depth66; modelBBis ordinary linear regression\. Each replicate hasn=5000n=5000observations, split evenly into fitting, estimation, calibration, and test splits\. Results are averaged over1010random seeds, with nominal levelα=0\.10\\alpha=0\.10and KNN neighborhood sizek=⌊\|Iest\|⌋≈35k=\\lfloor\\sqrt\{\|I\_\{\\mathrm\{est\}\}\|\}\\rfloor\\approx 35\. The primary score is the squared\-loss gap

S​\(X,Y\)=\(f^A​\(X\)−Y\)2−\(f^B​\(X\)−Y\)2\.S\(X,Y\)=\(\\hat\{f\}\_\{A\}\(X\)\-Y\)^\{2\}\-\(\\hat\{f\}\_\{B\}\(X\)\-Y\)^\{2\}\.\(7\)whereS<0S<0favors the boosted tree\. We evaluate local winner using the post\-fit oracleμ​\(x\)=\(f^A​\(x\)−f⋆​\(x\)\)2−\(f^B​\(x\)−f⋆​\(x\)\)2\\mu\(x\)=\(\\hat\{f\}\_\{A\}\(x\)\-f^\{\\star\}\(x\)\)^\{2\}\-\(\\hat\{f\}\_\{B\}\(x\)\-f^\{\\star\}\(x\)\)^\{2\}, which compares the two already\-fitted predictors at each covariate value\. Full DGP formulas and design motivations for all five synthetic cases are provided in Appendix[B\.1](https://arxiv.org/html/2607.29053#A2.SS1)\. Evaluation metrics for the synthetic experiments are defined in Appendix[B\.2](https://arxiv.org/html/2607.29053#A2.SS2)\.

#### Results \#1: Local comparison recovers winner regions hidden by global averages\.

Figure[1](https://arxiv.org/html/2607.29053#S6.F1)illustrates a global–local mismatch scenario, corresponding to Case 3 in Table[1](https://arxiv.org/html/2607.29053#A2.T1)\. Here, the data\-generating function is linear outside\[0\.42,0\.58\]\[0\.42,0\.58\]but contains a strong nonlinear oscillation inside this interval\. As a result, modelAAwins inside the island, while modelBBwins outside it\. The baselineglobal\_meanmeasures only the aggregate mean score over the full input space and declares a single winner everywhere\. Because the island has sufficiently large loss differences,global\_meandeclaresAAeverywhere and misses most of the region whereBBis locally better\.

To diagnose this structure, we evaluate the three non\-conformal local winner maps introduced in Section 3:local\_def1,local\_def2, andlocal\_def3\. These rules are used to study local\-region recovery\. In Case 3,global\_meanattains only0\.2930\.293winner accuracy, compared with0\.8230\.823forlocal\_def1and0\.9080\.908forlocal\_def3; full results are reported in Table[3](https://arxiv.org/html/2607.29053#A2.T3)\. As shown in Figure[1](https://arxiv.org/html/2607.29053#S6.F1),local\_def1closely follows the oracleθ​\(x0;r\)\\theta\(x\_\{0\};r\),local\_def2is more sensitive to local score noise, andlocal\_def3is more conservative near the island boundary\. These local rules reveal winner regions that are invisible toglobal\_mean, supporting Proposition[2](https://arxiv.org/html/2607.29053#Thmproposition2): the aggregate winner and the prevalence of local winners are distinct objects\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p7_three_winner.png)Figure 1:Three non\-conformal local winner methods on Case 3\. Left: local expected winner\. Middle: local majority winner\. Right: local quantile winner withβ=0\.20\\beta=0\.20\.
#### Results \#2: Conformal calibration controls false winner declarations\.

We next test whether local winner evidence can be converted into statistically controlled declarations\. Figure[2](https://arxiv.org/html/2607.29053#S6.F2)plots the false winner rate against the nominal levelα\\alphaacross all five synthetic cases\. The conformal methods follow the one\-sided upper\-bound rule in Section 4, declaringAAonly whenU^α​\(x\)<0\\widehat\{U\}\_\{\\alpha\}\(x\)<0\. Specifically,global\_split\_cpuses a constant centerS¯est\\bar\{S\}\_\{\\mathrm\{est\}\}and unit scale,local\_split\_cpuses the localized centerθ^r​\(x\)\\widehat\{\\theta\}\_\{r\}\(x\)with unit scale, andlocal\_std\_cpuses bothθ^r​\(x\)\\widehat\{\\theta\}\_\{r\}\(x\)and the localized scaleσ^r​\(x\)\\widehat\{\\sigma\}\_\{r\}\(x\)\. These variants separate the effects of localizing the center and the scale\. Across all five cases, the conformal methods stay below the diagonal throughout, consistent with Proposition[1](https://arxiv.org/html/2607.29053#Thmproposition1)\. In contrast, non\-conformal plug\-in maps can over\-declare when realized scores are noisy\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p4_validity.png)Figure 2:False winner rate versusα\\alpha\. The diagonal marks the target level\. Conformal methods control false winner declarations, while non\-conformal methods can over\-declare under noisy realized scores\.Figure[3](https://arxiv.org/html/2607.29053#S6.F3)illustrates the distinction between identifying and certifying a local winner in Case 3\. The non\-conformal local mean rule declaresAAwhen the estimated local mean is negative\. The conformal rule declaresAAonly when the conformal upper bound is below zero, producing a conservative subset of the estimated nonlinear island\. Thus, conformal calibration changes the output from a forced local winner map to a statistically controlled declaration rule\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p4_main_local_vs_global.png)Figure 3:Case 3 comparison between non\-conformal local winner identification and conformal winner declaration\. The conformal rule certifies a conservative subset of the localAA\-winning island\.Among conformal methods, local centering and local scaling improve the efficiency of declarations\. A global conformal rule uses one center for the entire covariate space, so it either certifiesAAbroadly or abstains broadly\. Local conformal methods can instead select only the regions whereAAis locally favorable\. Figures[4](https://arxiv.org/html/2607.29053#S6.F4)report selection rate and oracleAApower as functions ofα\\alpha\. Local conformal methods select fewer points than the oracleAAregion, reflecting the price of finite\-sample error control, but their selected points are more spatially targeted than those of the global conformal baseline\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p4_selection_power.png)Figure 4:Conformal selection versusα\\alpha\. Top: selection rate with the dashed line indicating the oracleAA\-winning rate\. Bottom: oracleAApower,ℙ​\(w^α=A∣μ​\(X\)<−τ\)\\mathbb\{P\}\(\\hat\{w\}\_\{\\alpha\}=A\\mid\\mu\(X\)<\-\\tau\), for the conformal methods\.
#### Results \#3: Diagnostics for the theoretical mechanisms\.

Figure[6](https://arxiv.org/html/2607.29053#S6.F6)verifies the squared\-loss bias–variance decomposition in Case 3\. Proposition[3](https://arxiv.org/html/2607.29053#Thmproposition3)is supported: outside the nonlinear island, the linear model has low variance and small bias, soBBwins locally; inside the island, the boosted tree’s bias reduction dominates its variance cost, soAAwins\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p5_biasvar_squared_loss.png)Figure 5:Bias–variance decomposition on Case 3\. Left: signal and mean predictions\. Middle: signed bias and variance components; negative values favor modelAAand positive values favor modelBB\. Right: their sum matches the Monte Carlo loss gap, showingAAwins inside the island andBBoutside\.
![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p5_theta_mse_vs_k.png)Figure 6:Local MSE versus neighborhood sizekk\. Smallkkgives noisy local estimates\. The optimalkkis case\-dependent\.

Figure[6](https://arxiv.org/html/2607.29053#S6.F6)studies the locality tradeoff behind Theorem[1](https://arxiv.org/html/2607.29053#Thmtheorem1)\. The MSE of the local estimator first decreases and eventually increases when neighborhoods become too wide\. Results of these five cases illustrates the optimalKKneeds to be tuned\.

### 6\.2Real data experiments

We evaluate on four public regression benchmarks:concrete,auto\_mpg,blog\_data, andfacebook\_1\.111[https://github\.com/yromano/cqr](https://github.com/yromano/cqr)222[https://archive\.ics\.uci\.edu/dataset/9/auto\+mpg](https://archive.ics.uci.edu/dataset/9/auto+mpg)ModelAAis a small fully connected neural network \(MLP\) with one hidden layer, and modelBBis linear regression\. The dataset statistics are summarized in Table[4](https://arxiv.org/html/2607.29053#A2.T4)in the appendix\. Across the four datasets, the MLP has better global mean performance on average across seeds, while linear regression may still be locally superior in structured subregions\. The goal is to find local regions where linear regression outperforms the MLP\. Since no oracle map is available in real\-data experiments, we report modelBB’s selection rate and the conditional gain𝔼​\[LA−LB∣w^α=B\]\\mathbb\{E\}\[L\_\{A\}\-L\_\{B\}\\mid\\hat\{w\}\_\{\\alpha\}=B\], which measures the realized advantage of routing to modelBBin the declared region relative to using modelAAeverywhere, and visualize the declared regions in a two\-dimensional PCA projection\.

#### Result \#1: Model A vs\. model B winning regions\.

Figure[7](https://arxiv.org/html/2607.29053#S6.F7)visualizes winner declarations in a 2D PCA projection, with colors denoting the across\-seed frequency of declaring modelBB\. The global mean rule is a global decision, selecting all points as eitherAAorBBwithin each seed; intermediate colors arise only from variation across seeds\. Local non\-conformal maps \(columns 1–2\) identify subregions where modelBBachieves lower loss\. Columns 3–4 highlight our method’s advantage: the global conformal rule abstains entirely, whereas our local conformal rule successfully certifiesBB\-favorable regions\. Though smaller than the non\-conformal baseline due to rigorous false\-winner control \(Proposition[1](https://arxiv.org/html/2607.29053#Thmproposition1)\), our localized approach converts noisy estimates into conservative, statistically controlled subregions where modelBBis favored\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p7_realdata_winner_regions.png)Figure 7:A/B winner regions in a 2D PCA projection\. Colors show across\-seed selection frequency: red\-to\-blue for non\-CP methods \(AAtoBB\), and gray\-to\-blue for CP methods \(abstention toBB\)\.
#### Result \#2: Selection vs\. conditional gain\.

Figure[8](https://arxiv.org/html/2607.29053#S6.F8)plots conditional gain—the average realized test\-set advantage of modelBBover modelAAwithin the declared region—against selection rate across varyingα\\alpha\. Increasingα\\alphaexpands declarations, though typically diluting concentration in high\-gain areas\. Non\-conformal methods achieve higher selection rates but lower conditional gain, reflecting overly broad declarations that lack false\-winner control\. The global conformal baseline \(yellow point\) abstains entirely, yielding zero gain\. Conversely, the local conformal method \(red line\) achieves the highest conditional gain, outperforming its non\-conformal counterpart \(blue point\)\. This confirms its abstention mechanism effectively isolates a smaller, highly targeted region of genuine superiority\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p7_realdata_gain_scatter.png)Figure 8:Conditional gain versus selection rate on real datasets\. Larger points correspond to a largerα\\alpha\. Y\-axis indicates a larger realized gain in the declared region, while x\-axis indicates a higher selection rate\. Local conformal methods select smaller but higher\-gain regions\.

## 7Conclusion

We presented a novel localized framework for comparing predictive models: the conformalized conditional mean comparison score\. To rigorously ground this methodology, we established three core theoretical properties: the asymptotic consistency of local winner estimation, the formal disconnect between global aggregate metrics and local superiority prevalence, and an ex\-ante bias\-variance decomposition that explicitly explains the mechanical emergence of local advantages\. Our empirical experiments validate these theoretical claims, demonstrating the framework’s capacity to detect heterogeneous model performance where global metrics fail\.

## 8Limitations

Locally centered conformal comparison has several limitations\. First, the independent three\-way split\(Ifit,Iest,Ical\)\(I\_\{\\mathrm\{fit\}\},I\_\{\\mathrm\{est\}\},I\_\{\\mathrm\{cal\}\}\)is clean but sample\-inefficient\. Cross\-fitting or out\-of\-bag prediction can recycle data in practice, but extending the exact finite\-sample guarantee in Proposition[1](https://arxiv.org/html/2607.29053#Thmproposition1)to such dependent constructions requires additional analysis\. Second, localization in raw covariate space suffers from the curse of dimensionality\. In sparse regions, the fallbackθ^r​\(x0\)=S¯est\\widehat\{\\theta\}\_\{r\}\(x\_\{0\}\)=\\bar\{S\}\_\{\\mathrm\{est\}\}keeps the procedure well\-defined and preserves validity, but sacrifices local adaptation; high\-dimensional applications should therefore applyKrK\_\{r\}to a learned low\-dimensional representation or sufficient summary\. Third, the conformal guarantee is marginal for the realized future score, not a finite\-sample confidence statement forμ​\(x0\)\\mu\(x\_\{0\}\)orθ​\(x0;r\)\\theta\(x\_\{0\};r\)\. Near tie boundaries, whereμ​\(x0\)≈0\\mu\(x\_\{0\}\)\\approx 0, the method will often abstain rather than declare a winner; this is appropriate but may reduce coverage of weak\-signal regions\. Finally, constructing a full local best\-model map over many target pointsx0x\_\{0\}introduces a multiple\-comparisons problem, so simultaneous map\-level error control requires external correction\.

## References

- \[1\]\(2023\)Conformal prediction: a gentle introduction\.Foundations and Trends in Machine Learning16\(4\),pp\. 494–591\.Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p1.1)\.
- \[2\]R\. Giacomini and H\. White\(2006\)Tests of conditional predictive ability\.Econometrica74\(6\),pp\. 1545–1578\.Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p3.1)\.
- \[3\]L\. Guan\(2023\)Localized conformal prediction: a generalized inference framework for conformal prediction\.Biometrika110\(1\),pp\. 33–50\.Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p1.1)\.
- \[4\]M\. I\. Jordan and R\. A\. Jacobs\(1994\)Hierarchical mixtures of experts and the EM algorithm\.Neural Computation6\(2\),pp\. 181–214\.Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p3.1)\.
- \[5\]J\. Lei, M\. G’Sell, A\. Rinaldo, R\. J\. Tibshirani, and L\. Wasserman\(2018\)Distribution\-free predictive inference for regression\.Journal of the American Statistical Association113\(523\),pp\. 1094–1111\.Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p1.1),[Remark 1](https://arxiv.org/html/2607.29053#Thmremark1.p1.3.3)\.
- \[6\]R\. Liang, W\. Zhu, and R\. F\. Barber\(2024\)Conformal prediction after data\-dependent model selection\.arXiv preprint arXiv:2408\.07066\.External Links:2408\.07066Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p2.1)\.
- \[7\]Y\. Romano, E\. Patterson, and E\. J\. Candès\(2019\)Conformalized quantile regression\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p1.1)\.
- \[8\]G\. Shafer and V\. Vovk\(2008\)A tutorial on conformal prediction\.Journal of Machine Learning Research9,pp\. 371–421\.Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p1.1)\.
- \[9\]R\. J\. Tibshirani, R\. F\. Barber, E\. J\. Candès, and A\. Ramdas\(2019\)Conformal prediction under covariate shift\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p1.1)\.
- \[10\]V\. Vovk, A\. Gammerman, and G\. Shafer\(2005\)Algorithmic learning in a random world\.Springer\.Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p1.1),[Remark 1](https://arxiv.org/html/2607.29053#Thmremark1.p1.3.3)\.
- \[11\]Y\. Wang and T\. Wang\(2026\)Localized conformal model selection\.arXiv preprint arXiv:2602\.19284\.External Links:2602\.19284Cited by:[§2](https://arxiv.org/html/2607.29053#S2.p2.1)\.

## Appendix AAppendix: Proofs

### A\.1Proof of Proposition[1](https://arxiv.org/html/2607.29053#Thmproposition1)

###### Proof of Proposition[1](https://arxiv.org/html/2607.29053#Thmproposition1)\.

Condition on the data in the fitting and estimation splits, denoted collectively by𝒟0\\mathcal\{D\}\_\{0\}\. After conditioning on𝒟0\\mathcal\{D\}\_\{0\}, the fitted predictors, the localized centerθ^r​\(⋅\)\\widehat\{\\theta\}\_\{r\}\(\\cdot\), and the localized scaleσ^r​\(⋅\)\\widehat\{\\sigma\}\_\{r\}\(\\cdot\)are fixed measurable functions\. The remaining randomness comes only from the calibration observations\(Xi,Yi\)i∈Ical\(X\_\{i\},Y\_\{i\}\)\_\{i\\in I\_\{\\mathrm\{cal\}\}\}and the independent test point\(Xn\+1,Yn\+1\)\(X\_\{n\+1\},Y\_\{n\+1\}\)\.

Let

Sn\+1:=S​\(Xn\+1,Yn\+1\)andRn\+1:=Sn\+1−θ^r​\(Xn\+1\)σ^r​\(Xn\+1\)\.S\_\{n\+1\}:=S\(X\_\{n\+1\},Y\_\{n\+1\}\)\\quad\\text\{and\}\\quad R\_\{n\+1\}:=\\frac\{S\_\{n\+1\}\-\\widehat\{\\theta\}\_\{r\}\(X\_\{n\+1\}\)\}\{\\widehat\{\\sigma\}\_\{r\}\(X\_\{n\+1\}\)\}\.Together with the calibration residualsRiR\_\{i\}defined in the main text, the collection

\{Ri:i∈Ical\}∪\{Rn\+1\}\\\{R\_\{i\}:i\\in I\_\{\\mathrm\{cal\}\}\\\}\\cup\\\{R\_\{n\+1\}\\\}is exchangeable conditional on𝒟0\\mathcal\{D\}\_\{0\}\.

Letncal=\|Ical\|n\_\{\\mathrm\{cal\}\}=\|I\_\{\\mathrm\{cal\}\}\|andkα=⌈\(ncal\+1\)​\(1−α\)⌉k\_\{\\alpha\}=\\lceil\(n\_\{\\mathrm\{cal\}\}\+1\)\(1\-\\alpha\)\\rceil, as defined in Section[4](https://arxiv.org/html/2607.29053#S4)\. Ifkα\>ncalk\_\{\\alpha\}\>n\_\{\\mathrm\{cal\}\}, thenq^1−α=∞\\widehat\{q\}\_\{1\-\\alpha\}=\\inftyand the desired coverage statement is immediate\. Otherwise,q^1−α=R\(kα\)\\widehat\{q\}\_\{1\-\\alpha\}=R\_\{\(k\_\{\\alpha\}\)\}, thekαk\_\{\\alpha\}\-th order statistic of the calibration residuals\. By exchangeability, using random tie\-breaking only for the rank argument, the rank ofRn\+1R\_\{n\+1\}among thencal\+1n\_\{\\mathrm\{cal\}\}\+1residuals is uniform on\{1,…,ncal\+1\}\\\{1,\\ldots,n\_\{\\mathrm\{cal\}\}\+1\\\}\. Moreover, the event\{Rn\+1\>R\(kα\)\}\\\{R\_\{n\+1\}\>R\_\{\(k\_\{\\alpha\}\)\}\\\}implies that this rank is larger thankαk\_\{\\alpha\}\. Therefore,

ℙ​\(Rn\+1\>q^1−α∣𝒟0\)≤ncal\+1−kαncal\+1≤α,\\mathbb\{P\}\\left\(R\_\{n\+1\}\>\\widehat\{q\}\_\{1\-\\alpha\}\\mid\\mathcal\{D\}\_\{0\}\\right\)\\leq\\frac\{n\_\{\\mathrm\{cal\}\}\+1\-k\_\{\\alpha\}\}\{n\_\{\\mathrm\{cal\}\}\+1\}\\leq\\alpha,or equivalently,

ℙ​\(Rn\+1≤q^1−α∣𝒟0\)≥1−α\.\\mathbb\{P\}\\left\(R\_\{n\+1\}\\leq\\widehat\{q\}\_\{1\-\\alpha\}\\mid\\mathcal\{D\}\_\{0\}\\right\)\\geq 1\-\\alpha\.
Becauseσ^r​\(x\)\>0\\widehat\{\\sigma\}\_\{r\}\(x\)\>0by construction, the event\{Rn\+1≤q^1−α\}\\\{R\_\{n\+1\}\\leq\\widehat\{q\}\_\{1\-\\alpha\}\\\}is equivalent to

Sn\+1≤θ^r​\(Xn\+1\)\+σ^r​\(Xn\+1\)​q^1−α=U^α​\(Xn\+1\),S\_\{n\+1\}\\leq\\widehat\{\\theta\}\_\{r\}\(X\_\{n\+1\}\)\+\\widehat\{\\sigma\}\_\{r\}\(X\_\{n\+1\}\)\\widehat\{q\}\_\{1\-\\alpha\}=\\widehat\{U\}\_\{\\alpha\}\(X\_\{n\+1\}\),where the last equality follows from the definition of the conformal upper bound\. Marginalizing over𝒟0\\mathcal\{D\}\_\{0\}gives

ℙ​\(Sn\+1≤U^α​\(Xn\+1\)\)≥1−α\.\\mathbb\{P\}\\left\(S\_\{n\+1\}\\leq\\widehat\{U\}\_\{\\alpha\}\(X\_\{n\+1\}\)\\right\)\\geq 1\-\\alpha\.
Finally, the false\-winner event satisfies

\{U^α​\(Xn\+1\)<0,Sn\+1≥0\}⊆\{Sn\+1\>U^α​\(Xn\+1\)\}\.\\\{\\widehat\{U\}\_\{\\alpha\}\(X\_\{n\+1\}\)<0,\\ S\_\{n\+1\}\\geq 0\\\}\\subseteq\\\{S\_\{n\+1\}\>\\widehat\{U\}\_\{\\alpha\}\(X\_\{n\+1\}\)\\\}\.Hence,

ℙ​\(U^α​\(Xn\+1\)<0,Sn\+1≥0\)≤ℙ​\(Sn\+1\>U^α​\(Xn\+1\)\)≤α\.\\mathbb\{P\}\\left\(\\widehat\{U\}\_\{\\alpha\}\(X\_\{n\+1\}\)<0,\\ S\_\{n\+1\}\\geq 0\\right\)\\leq\\mathbb\{P\}\\left\(S\_\{n\+1\}\>\\widehat\{U\}\_\{\\alpha\}\(X\_\{n\+1\}\)\\right\)\\leq\\alpha\.∎

### A\.2Proof of Theorem[1](https://arxiv.org/html/2607.29053#Thmtheorem1)

###### Proof of Theorem[1](https://arxiv.org/html/2607.29053#Thmtheorem1)\.

Throughout the proof, we condition on the fitted predictors produced by the independent fitting split, so that the evaluation observations are i\.i\.d\. andμ​\(x\)=𝔼​\[S​\(X,Y\)∣X=x\]\\mu\(x\)=\\mathbb\{E\}\[S\(X,Y\)\\mid X=x\]is fixed\. Define

Nn≔1n​∑i=1nKrn​\(Xi,x0\)​Si,Dn≔1n​∑i=1nKrn​\(Xi,x0\)\.N\_\{n\}\\coloneqq\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}K\_\{r\_\{n\}\}\(X\_\{i\},x\_\{0\}\)S\_\{i\},\\qquad D\_\{n\}\\coloneqq\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}K\_\{r\_\{n\}\}\(X\_\{i\},x\_\{0\}\)\.The ratioNn/DnN\_\{n\}/D\_\{n\}is identical to the estimator in \([4](https://arxiv.org/html/2607.29053#S5.E4)\) wheneverDn\>0D\_\{n\}\>0\. On the eventDn=0D\_\{n\}=0, we may assign the estimator any fixed value, say zero; the argument below shows thatℙ​\(Dn\>0\)→1\\mathbb\{P\}\(D\_\{n\}\>0\)\\to 1, so this finite\-sample convention has no asymptotic effect\.

Let

κ0≔∫ℝdK​\(u\)​𝑑u\.\\kappa\_\{0\}\\coloneqq\\int\_\{\\mathbb\{R\}^\{d\}\}K\(u\)\\,du\.By the theorem’s kernel condition,0<κ0<∞0<\\kappa\_\{0\}<\\infty\. SinceKKis compactly supported, there existsRK<∞R\_\{K\}<\\inftysuch thatK​\(u\)=0K\(u\)=0whenever‖u‖\>RK\\\|u\\\|\>R\_\{K\}\. SincepXp\_\{X\}is continuous atx0x\_\{0\}andpX​\(x0\)\>0p\_\{X\}\(x\_\{0\}\)\>0,pXp\_\{X\}is locally bounded nearx0x\_\{0\}and remains positive on a sufficiently small neighborhood ofx0x\_\{0\}\. Thus, for all largenn, the ordinary kernel localizationx0\+rn​ux\_\{0\}\+r\_\{n\}uwithK​\(u\)≠0K\(u\)\\neq 0stays inside a neighborhood where the local continuity and moment conditions apply\.

First consider the denominator\. By the change of variablesu=\(x−x0\)/rnu=\(x\-x\_\{0\}\)/r\_\{n\},

𝔼​\[Dn\]=∫ℝdK​\(u\)​pX​\(x0\+rn​u\)​𝑑u\.\\mathbb\{E\}\[D\_\{n\}\]=\\int\_\{\\mathbb\{R\}^\{d\}\}K\(u\)p\_\{X\}\(x\_\{0\}\+r\_\{n\}u\)\\,du\.BecausepXp\_\{X\}is continuous atx0x\_\{0\}andKKis bounded with compact support, dominated convergence gives

𝔼​\[Dn\]⟶pX​\(x0\)​∫ℝdK​\(u\)​𝑑u=pX​\(x0\)​κ0≔d0\.\\mathbb\{E\}\[D\_\{n\}\]\\longrightarrow p\_\{X\}\(x\_\{0\}\)\\int\_\{\\mathbb\{R\}^\{d\}\}K\(u\)\\,du=p\_\{X\}\(x\_\{0\}\)\\kappa\_\{0\}\\coloneqq d\_\{0\}\.SincepX​\(x0\)\>0p\_\{X\}\(x\_\{0\}\)\>0andκ0\>0\\kappa\_\{0\}\>0, we haved0\>0d\_\{0\}\>0\.

Next consider the numerator\. By iterated expectation,

𝔼​\[Nn\]=𝔼​\[Krn​\(X,x0\)​μ​\(X\)\]\.\\mathbb\{E\}\[N\_\{n\}\]=\\mathbb\{E\}\\\!\\left\[K\_\{r\_\{n\}\}\(X,x\_\{0\}\)\\mu\(X\)\\right\]\.Using the same change of variables,

𝔼​\[Nn\]=∫ℝdK​\(u\)​μ​\(x0\+rn​u\)​pX​\(x0\+rn​u\)​𝑑u\.\\mathbb\{E\}\[N\_\{n\}\]=\\int\_\{\\mathbb\{R\}^\{d\}\}K\(u\)\\mu\(x\_\{0\}\+r\_\{n\}u\)p\_\{X\}\(x\_\{0\}\+r\_\{n\}u\)\\,du\.The local second\-moment condition and Jensen’s inequality imply thatμ\\muis locally bounded nearx0x\_\{0\}:

\|μ​\(x\)\|≤𝔼​\[\|S​\(X,Y\)\|∣X=x\]≤𝔼​\[S​\(X,Y\)2∣X=x\]1/2\.\|\\mu\(x\)\|\\leq\\mathbb\{E\}\[\|S\(X,Y\)\|\\mid X=x\]\\leq\\mathbb\{E\}\[S\(X,Y\)^\{2\}\\mid X=x\]^\{1/2\}\.Together with the local boundedness ofpXp\_\{X\}, this gives an integrable dominating function proportional toK​\(u\)K\(u\)\. Since bothμ\\muandpXp\_\{X\}are continuous atx0x\_\{0\}, dominated convergence yields

𝔼​\[Nn\]⟶μ​\(x0\)​pX​\(x0\)​∫ℝdK​\(u\)​𝑑u=μ​\(x0\)​pX​\(x0\)​κ0\.\\mathbb\{E\}\[N\_\{n\}\]\\longrightarrow\\mu\(x\_\{0\}\)p\_\{X\}\(x\_\{0\}\)\\int\_\{\\mathbb\{R\}^\{d\}\}K\(u\)\\,du=\\mu\(x\_\{0\}\)p\_\{X\}\(x\_\{0\}\)\\kappa\_\{0\}\.
It remains to show thatNnN\_\{n\}andDnD\_\{n\}concentrate around their expectations\. For the denominator,

Var​\(Dn\)=1n​Var​\(Krn​\(X,x0\)\)≤1n​𝔼​\[Krn​\(X,x0\)2\]\.\\mathrm\{Var\}\(D\_\{n\}\)=\\frac\{1\}\{n\}\\mathrm\{Var\}\\\!\\left\(K\_\{r\_\{n\}\}\(X,x\_\{0\}\)\\right\)\\leq\\frac\{1\}\{n\}\\mathbb\{E\}\\\!\\left\[K\_\{r\_\{n\}\}\(X,x\_\{0\}\)^\{2\}\\right\]\.Again changing variables,

𝔼​\[Krn​\(X,x0\)2\]=rn−d​∫ℝdK​\(u\)2​pX​\(x0\+rn​u\)​𝑑u\.\\mathbb\{E\}\\\!\\left\[K\_\{r\_\{n\}\}\(X,x\_\{0\}\)^\{2\}\\right\]=r\_\{n\}^\{\-d\}\\int\_\{\\mathbb\{R\}^\{d\}\}K\(u\)^\{2\}p\_\{X\}\(x\_\{0\}\+r\_\{n\}u\)\\,du\.SinceKKis bounded and compactly supported,K2K^\{2\}is integrable; sincepXp\_\{X\}is locally bounded atx0x\_\{0\}, the integral isO​\(1\)O\(1\)\. Hence

Var​\(Dn\)=O​\(1n​rnd\)⟶0,\\mathrm\{Var\}\(D\_\{n\}\)=O\\\!\\left\(\\frac\{1\}\{nr\_\{n\}^\{d\}\}\\right\)\\longrightarrow 0,becausen​rnd→∞nr\_\{n\}^\{d\}\\to\\infty\. Therefore

Dn−𝔼​\[Dn\]​⟶𝑝​0\.D\_\{n\}\-\\mathbb\{E\}\[D\_\{n\}\]\\overset\{p\}\{\\longrightarrow\}0\.
For the numerator,

Var​\(Nn\)=1n​Var​\(Krn​\(X,x0\)​S\)≤1n​𝔼​\[Krn​\(X,x0\)2​S2\]\.\\mathrm\{Var\}\(N\_\{n\}\)=\\frac\{1\}\{n\}\\mathrm\{Var\}\\\!\\left\(K\_\{r\_\{n\}\}\(X,x\_\{0\}\)S\\right\)\\leq\\frac\{1\}\{n\}\\mathbb\{E\}\\\!\\left\[K\_\{r\_\{n\}\}\(X,x\_\{0\}\)^\{2\}S^\{2\}\\right\]\.By iterated expectation and the same change of variables,

𝔼​\[Krn​\(X,x0\)2​S2\]=rn−d​∫ℝdK​\(u\)2​𝔼​\[S2∣X=x0\+rn​u\]​pX​\(x0\+rn​u\)​𝑑u\.\\mathbb\{E\}\\\!\\left\[K\_\{r\_\{n\}\}\(X,x\_\{0\}\)^\{2\}S^\{2\}\\right\]=r\_\{n\}^\{\-d\}\\int\_\{\\mathbb\{R\}^\{d\}\}K\(u\)^\{2\}\\mathbb\{E\}\\\!\\left\[S^\{2\}\\mid X=x\_\{0\}\+r\_\{n\}u\\right\]p\_\{X\}\(x\_\{0\}\+r\_\{n\}u\)\\,du\.For all largenn, the compact support ofKKkeepsx0\+rn​ux\_\{0\}\+r\_\{n\}uinside the local neighborhood where𝔼​\[S2∣X=x\]\\mathbb\{E\}\[S^\{2\}\\mid X=x\]is bounded and wherepXp\_\{X\}is bounded\. Therefore

𝔼​\[Krn​\(X,x0\)2​S2\]=O​\(rn−d\),\\mathbb\{E\}\\\!\\left\[K\_\{r\_\{n\}\}\(X,x\_\{0\}\)^\{2\}S^\{2\}\\right\]=O\(r\_\{n\}^\{\-d\}\),and hence

Var​\(Nn\)=O​\(1n​rnd\)⟶0\.\\mathrm\{Var\}\(N\_\{n\}\)=O\\\!\\left\(\\frac\{1\}\{nr\_\{n\}^\{d\}\}\\right\)\\longrightarrow 0\.Thus

Nn−𝔼​\[Nn\]​⟶𝑝​0\.N\_\{n\}\-\\mathbb\{E\}\[N\_\{n\}\]\\overset\{p\}\{\\longrightarrow\}0\.
Combining the expectation limits and concentration bounds gives

Dn​⟶𝑝​pX​\(x0\)​κ0D\_\{n\}\\overset\{p\}\{\\longrightarrow\}p\_\{X\}\(x\_\{0\}\)\\kappa\_\{0\}and

Nn​⟶𝑝​μ​\(x0\)​pX​\(x0\)​κ0\.N\_\{n\}\\overset\{p\}\{\\longrightarrow\}\\mu\(x\_\{0\}\)p\_\{X\}\(x\_\{0\}\)\\kappa\_\{0\}\.BecausepX​\(x0\)​κ0\>0p\_\{X\}\(x\_\{0\}\)\\kappa\_\{0\}\>0, we also have

ℙ​\(Dn\>0\)⟶1\.\\mathbb\{P\}\(D\_\{n\}\>0\)\\longrightarrow 1\.Consequently, by Slutsky’s theorem,

μ^n​\(x0\)=NnDn​⟶𝑝​μ​\(x0\)\.\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)=\\frac\{N\_\{n\}\}\{D\_\{n\}\}\\overset\{p\}\{\\longrightarrow\}\\mu\(x\_\{0\}\)\.
It remains to translate consistency of the local mean estimator into consistency of the local winner\. Ifμ​\(x0\)<0\\mu\(x\_\{0\}\)<0, thenδ⋆​\(x0\)=A\\delta^\{\\star\}\(x\_\{0\}\)=A, and

ℙ​\(δ^n​\(x0\)≠A\)=ℙ​\(μ^n​\(x0\)≥0\)\.\\mathbb\{P\}\\\!\\left\(\\hat\{\\delta\}\_\{n\}\(x\_\{0\}\)\\neq A\\right\)=\\mathbb\{P\}\\\!\\left\(\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)\\geq 0\\right\)\.Moreover,

\{μ^n​\(x0\)≥0\}⊆\{\|μ^n​\(x0\)−μ​\(x0\)\|≥\|μ​\(x0\)\|\}\.\\\{\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)\\geq 0\\\}\\subseteq\\left\\\{\|\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)\-\\mu\(x\_\{0\}\)\|\\geq\|\\mu\(x\_\{0\}\)\|\\right\\\}\.Sinceμ^n​\(x0\)​→𝑝​μ​\(x0\)\\hat\{\\mu\}\_\{n\}\(x\_\{0\}\)\\overset\{p\}\{\\to\}\\mu\(x\_\{0\}\), the probability of the event on the right converges to zero\. Hence

ℙ​\(δ^n​\(x0\)=A\)⟶1\.\\mathbb\{P\}\\\!\\left\(\\hat\{\\delta\}\_\{n\}\(x\_\{0\}\)=A\\right\)\\longrightarrow 1\.The caseμ​\(x0\)\>0\\mu\(x\_\{0\}\)\>0is identical and gives

ℙ​\(δ^n​\(x0\)=B\)⟶1\.\\mathbb\{P\}\\\!\\left\(\\hat\{\\delta\}\_\{n\}\(x\_\{0\}\)=B\\right\)\\longrightarrow 1\.Therefore, wheneverμ​\(x0\)≠0\\mu\(x\_\{0\}\)\\neq 0,

ℙ​\(δ^n​\(x0\)=δ⋆​\(x0\)\)⟶1\.\\mathbb\{P\}\\\!\\left\(\\hat\{\\delta\}\_\{n\}\(x\_\{0\}\)=\\delta^\{\\star\}\(x\_\{0\}\)\\right\)\\longrightarrow 1\.
The same proof applies to the estimation\-split estimator in Section 4 after replacingnnby\|Iest\|\|I\_\{\\mathrm\{est\}\}\|, with the corresponding radius sequence satisfying the same shrinking\-neighborhood and diverging\-effective\- sample\-size conditions\. ∎

### A\.3Proof of Proposition[2](https://arxiv.org/html/2607.29053#Thmproposition2)

###### Proof of Proposition[2](https://arxiv.org/html/2607.29053#Thmproposition2)\.

We first prove consistency of the aggregate selector for the sign of the global mean score\. LetSi:=S​\(Xi,Yi\)S\_\{i\}:=S\(X\_\{i\},Y\_\{i\}\)\. By definition,

S¯n=1n​∑i=1nSi\.\\bar\{S\}\_\{n\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}S\_\{i\}\.Since𝔼​\[\|S​\(X,Y\)\|\]<∞\\mathbb\{E\}\[\|S\(X,Y\)\|\]<\\infty, the weak law of large numbers implies

S¯n​⟶𝑝​𝔼​\[S​\(X,Y\)\]=μglob\.\\bar\{S\}\_\{n\}\\overset\{p\}\{\\longrightarrow\}\\mathbb\{E\}\[S\(X,Y\)\]=\\mu\_\{\\mathrm\{glob\}\}\.
Suppose first thatμglob\>0\\mu\_\{\\mathrm\{glob\}\}\>0\. Then the true aggregate winner isBB, and

ℙ​\(δ^nagg=B\)=ℙ​\(S¯n≥0\)\.\\mathbb\{P\}\\left\(\\hat\{\\delta\}\_\{n\}^\{\\mathrm\{agg\}\}=B\\right\)=\\mathbb\{P\}\(\\bar\{S\}\_\{n\}\\geq 0\)\.Hence

ℙ​\(S¯n<0\)≤ℙ​\(\|S¯n−μglob\|≥μglob\)⟶0,\\mathbb\{P\}\(\\bar\{S\}\_\{n\}<0\)\\leq\\mathbb\{P\}\\left\(\|\\bar\{S\}\_\{n\}\-\\mu\_\{\\mathrm\{glob\}\}\|\\geq\\mu\_\{\\mathrm\{glob\}\}\\right\)\\longrightarrow 0,so

ℙ​\(δ^nagg=B\)⟶1\.\\mathbb\{P\}\\left\(\\hat\{\\delta\}\_\{n\}^\{\\mathrm\{agg\}\}=B\\right\)\\longrightarrow 1\.The caseμglob<0\\mu\_\{\\mathrm\{glob\}\}<0is identical and yields

ℙ​\(δ^nagg=A\)⟶1\.\\mathbb\{P\}\\left\(\\hat\{\\delta\}\_\{n\}^\{\\mathrm\{agg\}\}=A\\right\)\\longrightarrow 1\.Therefore,

ℙ​\(δ^nagg=δagg,⋆\)⟶1\.\\mathbb\{P\}\\left\(\\hat\{\\delta\}\_\{n\}^\{\\mathrm\{agg\}\}=\\delta^\{\\mathrm\{agg\},\\star\}\\right\)\\longrightarrow 1\.
We next prove the existence statement\. Fix anyq∈\(1/2,1\)q\\in\(1/2,1\)\. LetX∼Unif​\(0,1\)X\\sim\\mathrm\{Unif\}\(0,1\), and letℛA=\[0,q\]\\mathcal\{R\}\_\{A\}=\[0,q\]\. Define

μ​\(x\)=\{−a,x∈ℛA,b,x∉ℛA,\\mu\(x\)=\\begin\{cases\}\-a,&x\\in\\mathcal\{R\}\_\{A\},\\\\ b,&x\\notin\\mathcal\{R\}\_\{A\},\\end\{cases\}wherea\>0a\>0andb\>q​a/\(1−q\)b\>qa/\(1\-q\)\. Then

ℙ​\(μ​\(X\)<0\)=ℙ​\(X∈ℛA\)=q\.\\mathbb\{P\}\(\\mu\(X\)<0\)=\\mathbb\{P\}\(X\\in\\mathcal\{R\}\_\{A\}\)=q\.Moreover,

𝔼​\[μ​\(X\)\]=−q​a\+\(1−q\)​b\>0\.\\mathbb\{E\}\[\\mu\(X\)\]=\-q\\,a\+\(1\-q\)b\>0\.Letη\\etabe any mean\-zero noise, independent ofXX, with𝔼​\[\|η\|\]<∞\\mathbb\{E\}\[\|\\eta\|\]<\\inftyandVar​\(η\)\>0\\mathrm\{Var\}\(\\eta\)\>0, and define

S:=μ​\(X\)\+η\.S:=\\mu\(X\)\+\\eta\.Then𝔼​\[\|S\|\]<∞\\mathbb\{E\}\[\|S\|\]<\\infty,𝔼​\[S∣X=x\]=μ​\(x\)\\mathbb\{E\}\[S\\mid X=x\]=\\mu\(x\), and the score is non\-degenerate becauseVar​\(S∣X=x\)=Var​\(η\)\>0\\mathrm\{Var\}\(S\\mid X=x\)=\\mathrm\{Var\}\(\\eta\)\>0\. Therefore

ℙ​\(μ​\(X\)<0\)=qand𝔼​\[S\]\>0\.\\mathbb\{P\}\(\\mu\(X\)<0\)=q\\qquad\\text\{and\}\\qquad\\mathbb\{E\}\[S\]\>0\.Applying the aggregate selector to i\.i\.d\. draws from this constructed distribution, the first part of the proof gives

ℙ​\(δ^nagg=B\)⟶1\.\\mathbb\{P\}\\left\(\\hat\{\\delta\}\_\{n\}^\{\\mathrm\{agg\}\}=B\\right\)\\longrightarrow 1\.This proves that aggregate selection can asymptotically favor modelBBeven though modelAAis locally superior on at least aqq\-fraction of the covariate distribution\. ∎

### A\.4Proof of Proposition[3](https://arxiv.org/html/2607.29053#Thmproposition3)

###### Proof of Proposition[3](https://arxiv.org/html/2607.29053#Thmproposition3)\.

Fixx∈𝒳x\\in\\mathcal\{X\}\. Conditional onX=xX=x, writeY=f⋆​\(x\)\+εY=f^\{\\star\}\(x\)\+\\varepsilonwith𝔼​\[ε∣X=x\]=0\\mathbb\{E\}\[\\varepsilon\\mid X=x\]=0\. For a generic modelm∈\{A,B\}m\\in\\\{A,B\\\},

Lm​\(X,Y\)∣X=x\\displaystyle L\_\{m\}\(X,Y\)\\mid X=x=\(f^m​\(x\)−f⋆​\(x\)−ε\)2\\displaystyle=\\bigl\(\\hat\{f\}\_\{m\}\(x\)\-f^\{\\star\}\(x\)\-\\varepsilon\\bigr\)^\{2\}=\(f^m​\(x\)−f⋆​\(x\)\)2−2​ε​\(f^m​\(x\)−f⋆​\(x\)\)\+ε2\.\\displaystyle=\\bigl\(\\hat\{f\}\_\{m\}\(x\)\-f^\{\\star\}\(x\)\\bigr\)^\{2\}\-2\\varepsilon\\bigl\(\\hat\{f\}\_\{m\}\(x\)\-f^\{\\star\}\(x\)\\bigr\)\+\\varepsilon^\{2\}\.\(8\)
Taking expectation over both the training randomness and the fresh test noise, we isolate the cross\-term\. Because the fresh test noiseε\\varepsilonis independent of the fitting sample𝒟\\mathcal\{D\}conditional onXX, we have𝔼​\[ε∣X=x,𝒟\]=𝔼​\[ε∣X=x\]=0\\mathbb\{E\}\[\\varepsilon\\mid X=x,\\mathcal\{D\}\]=\\mathbb\{E\}\[\\varepsilon\\mid X=x\]=0\. Therefore,

𝔼​\[ε​f^m​\(x\)∣X=x\]\\displaystyle\\mathbb\{E\}\\bigl\[\\varepsilon\\hat\{f\}\_\{m\}\(x\)\\mid X=x\\bigr\]=𝔼​\[𝔼​\[ε​f^m​\(x\)∣X=x,𝒟\]∣X=x\]\\displaystyle=\\mathbb\{E\}\\Bigl\[\\mathbb\{E\}\\bigl\[\\varepsilon\\hat\{f\}\_\{m\}\(x\)\\mid X=x,\\mathcal\{D\}\\bigr\]\\mid X=x\\Bigr\]=𝔼​\[f^m​\(x\)​𝔼​\[ε∣X=x,𝒟\]∣X=x\]\\displaystyle=\\mathbb\{E\}\\Bigl\[\\hat\{f\}\_\{m\}\(x\)\\,\\mathbb\{E\}\\bigl\[\\varepsilon\\mid X=x,\\mathcal\{D\}\\bigr\]\\mid X=x\\Bigr\]=0\.\\displaystyle=0\.\(9\)
Consequently,

𝔼​\[Lm​\(X,Y\)∣X=x\]=𝔼​\[\(f^m​\(x\)−f⋆​\(x\)\)2\]\+𝔼​\[ε2∣X=x\]\.\\mathbb\{E\}\[L\_\{m\}\(X,Y\)\\mid X=x\]=\\mathbb\{E\}\\bigl\[\(\\hat\{f\}\_\{m\}\(x\)\-f^\{\\star\}\(x\)\)^\{2\}\\bigr\]\+\\mathbb\{E\}\[\\varepsilon^\{2\}\\mid X=x\]\.\(10\)
Applying the standard bias–variance identity𝔼​\[\(Z−a\)2\]=Var​\(Z\)\+\(𝔼​\[Z\]−a\)2\\mathbb\{E\}\[\(Z\-a\)^\{2\}\]=\\mathrm\{Var\}\(Z\)\+\(\\mathbb\{E\}\[Z\]\-a\)^\{2\}withZ=f^m​\(x\)Z=\\hat\{f\}\_\{m\}\(x\)anda=f⋆​\(x\)a=f^\{\\star\}\(x\)gives

𝔼​\[Lm​\(X,Y\)∣X=x\]=Varm​\(x\)\+Biasm2​\(x\)\+𝔼​\[ε2∣X=x\]\.\\mathbb\{E\}\[L\_\{m\}\(X,Y\)\\mid X=x\]=\\mathrm\{Var\}\_\{m\}\(x\)\+\\mathrm\{Bias\}\_\{m\}^\{2\}\(x\)\+\\mathbb\{E\}\[\\varepsilon^\{2\}\\mid X=x\]\.\(11\)
Subtracting the identities form=Am=Aandm=Bm=Bcancels the irreducible noise term𝔼​\[ε2∣X=x\]\\mathbb\{E\}\[\\varepsilon^\{2\}\\mid X=x\], yielding

𝔼​\[S​\(X,Y\)∣X=x\]=BiasA2​\(x\)−BiasB2​\(x\)\+VarA​\(x\)−VarB​\(x\)\.\\mathbb\{E\}\[S\(X,Y\)\\mid X=x\]=\\mathrm\{Bias\}\_\{A\}^\{2\}\(x\)\-\\mathrm\{Bias\}\_\{B\}^\{2\}\(x\)\+\\mathrm\{Var\}\_\{A\}\(x\)\-\\mathrm\{Var\}\_\{B\}\(x\)\.Multiplying both sides byKr​\(X,x0\)K\_\{r\}\(X,x\_\{0\}\), taking expectation, and dividing by𝔼​\[Kr​\(X,x0\)\]\\mathbb\{E\}\[K\_\{r\}\(X,x\_\{0\}\)\]yields \([6](https://arxiv.org/html/2607.29053#S5.E6)\)\. ∎

## Appendix BAppendix: Additional Experimental Details

This appendix provides full details for the experiments reported in Section[6](https://arxiv.org/html/2607.29053#S6)\. We include the complete data\-generating processes, visual illustrations of the fitted models, detailed non\-conformal winner\-identification metrics, and real datasets summary\.

### B\.1Synthetic data\-generating processes

Table[1](https://arxiv.org/html/2607.29053#A2.T1)reports the full synthetic data\-generating processes\. Figure[9](https://arxiv.org/html/2607.29053#A2.F9)visualizes the five designs, together with representative fitted models\. Case 1 is a baseline whereAAis better almost everywhere\. Cases 2 and 3 create global\-local mismatch\. In Case 3 \(Figure[6](https://arxiv.org/html/2607.29053#S6.F6)\), a narrow nonlinear island is strong enough to make the global average favorAA\. Specifically, A wins inside \[0\.42,0\.58\] and B wins about 84% of the region\. Case 4 increases the noise level while keeping the Case 1 signal, creating a stress test for false declarations\. Case 5 introduces strong heteroskedasticity, testing whether local scaling improves conformal efficiency\.

Table 1:Synthetic data\-generating processes\. All cases useX∼Unif​\[0,1\]X\\sim\\mathrm\{Unif\}\[0,1\]andY=f⋆​\(X\)\+εY=f^\{\\star\}\(X\)\+\\varepsilon\.![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p2_dgp_illustration.png)Figure 9:Synthetic DGPs\. Black dashed:f⋆​\(x\)f^\{\\star\}\(x\); gray band:f⋆​\(x\)±2​σ​\(x\)f^\{\\star\}\(x\)\\pm 2\\sigma\(x\); blue: a representative LightGBM \(AA\) fit; red: linear \(BB\) fit\.
### B\.2Evaluation metrics for synthetic experiments

Table[2](https://arxiv.org/html/2607.29053#A2.T2)defines all evaluation metrics used in synthetic experiments, separated into non\-conformal winner metrics that assess oracle map recovery and conformal winner metrics that assess the reliability, selectivity, and realized gain of declaredAAregions\.

Table 2:Evaluation metrics for synthetic experiments\. Non\-conformal winner metrics assess whether the estimated map matches the oracle local winner regions\. Conformal winner metrics assess whether declaredAAregions are reliable, selective, and useful\.RoleMetricDefinitionInterpretationNon\-conformal winner metricsWinner accuracyℙ\(w^α\(X\)=w⋆\(X\)∣\|μ\(X\)\|\>τ\)\\mathbb\{P\}\(\\hat\{w\}\_\{\\alpha\}\(X\)=w^\{\\star\}\(X\)\\mid\|\\mu\(X\)\|\>\\tau\)Overall agreement with the oracle winner map outside the tie bandRegion\-IoUAIoU between predicted and oracleAAregionsSpatial overlap with the oracleAA\-winning regionRegion\-IoUBIoU between predicted and oracleBBregionsSpatial overlap with the oracleBB\-winning regionConformal winner metricsSelection rateℙ​\(w^α=A\)\\mathbb\{P\}\(\\hat\{w\}\_\{\\alpha\}=A\)Fraction of test points declared asAAwinsFalse winner rateℙ​\(w^α=A,S≥0\)\\mathbb\{P\}\(\\hat\{w\}\_\{\\alpha\}=A,\\,S\\geq 0\)Empirical false\-winner rate forAAdeclarationsPowerℙ​\(w^α=A∣μ​\(X\)<−τ\)\\mathbb\{P\}\(\\hat\{w\}\_\{\\alpha\}=A\\mid\\mu\(X\)<\-\\tau\)Coverage of the true oracleAAregionConditional gain𝔼​\[LB−LA∣w^α=A\]=−𝔼​\[S∣w^α=A\]\\mathbb\{E\}\[L\_\{B\}\-L\_\{A\}\\mid\\hat\{w\}\_\{\\alpha\}=A\]=\-\\mathbb\{E\}\[S\\mid\\hat\{w\}\_\{\\alpha\}=A\]Realized value of the selectedAAdeclarations
### B\.3Non\-conformal winner\-identification results

Table[3](https://arxiv.org/html/2607.29053#A2.T3)reports local\-winner identification results for the forced non\-conformal winner maps\. These results complement the main text by showing full accuracy and region\-overlap values across all five synthetic cases\. The table is not used to establish finite\-sample validity; conformal validity is evaluated separately through false\-winner rates in the main paper\.

Table 3:Local\-winner identification metrics atα=0\.10\\alpha=0\.10\(mean over1010seeds\)\. Acc: winner accuracy outside the tie band; IoUAand IoUB: intersection\-over\-union with the oracleAAandBBregions\. All rows are forced winner maps and do not carry conformal guarantees\.
### B\.4Real datasets summary

Table[4](https://arxiv.org/html/2607.29053#A2.T4)reports the size, dimensionality, and global comparison score distribution for the four real\-data benchmarks used in Section[6\.2](https://arxiv.org/html/2607.29053#S6.SS2)\.

Table 4:Real datasets summary\.S¯\\bar\{S\}is the global mean comparison score on the calibration split, reported as mean±\\pmstd across1010seeds\.S<0S<0favors modelAA\(MLP\)\.

## Appendix CAppendix: Ablation Studies

#### Score family\.

Figure[10](https://arxiv.org/html/2607.29053#A3.F10)replacesgap\_msewith the standardized gapsstd​\(LA,LB;x\)=\(LA−LB\)/v^​\(x\)s\_\{\\mathrm\{std\}\}\(L\_\{A\},L\_\{B\};x\)=\(L\_\{A\}\-L\_\{B\}\)/\\hat\{v\}\(x\)and the log\-ratioT=log⁡\(\(LA\+τ\)/\(LB\+τ\)\)T=\\log\(\(L\_\{A\}\+\\tau\)/\(L\_\{B\}\+\\tau\)\)\. The conformal validity guarantee is score\-agnostic for the local CP construction, but selection and power vary\.gap\_mseis the only score with nonzero power on all five cases and is strongest on the two boundary cases and the validity\-stress case\. The log\-ratio and standardized gap improve power on the homogeneous and heteroskedastic cases, but they lose the nonlinear\-tail signal and are weaker on Case 3\. We therefore keepgap\_mseas the primary score and treat the alternatives as useful but case\-dependent sensitivity checks\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p6_score_family.png)Figure 10:Score\-family ablation\. Selection rate \(top\) and oracleAApower \(bottom\) forlocal\_split\_cpunder the three scoresgap\_mse,standardized\_gap, andlog\_ratio\_mse, all five cases\.
#### Locality parameterkk\.

Figure[11](https://arxiv.org/html/2607.29053#A3.F11)sweepsk∈\{5,10,20,50,100,200,400\}k\\in\\\{5,10,20,50,100,200,400\\\}and plots selection rate \(top\) and oracleAApower \(bottom\) forlocal\_split\_cpon all five cases atα=0\.10\\alpha=0\.10\. The homogeneous Case 1 is essentially flat inkkas expected\. Cases with local structure split, such as Case 2, loses power for largekkbecause the tail signal is averaged away\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p6_k_sweep.png)Figure 11:Locality parameter ablation\. Selection rate \(top\) and oracleAApower \(bottom\) forlocal\_split\_cpaskkvaries, all five cases atα=0\.10\\alpha=0\.10\.
#### Covariate dimension\.

Figure[12](https://arxiv.org/html/2607.29053#A3.F12)keeps the signal in the first coordinate ofXXand pads the remainingd−1d\-1coordinates with i\.i\.d\. uniform nuisance covariates, sweepingd∈\{1,10,30,100\}d\\in\\\{1,10,30,100\\\}\. Cases with local boundary structure \(Cases 2 and 3\) lose power withddincreases, consistent with the curse of dimensionality forkk\-NN\. Case 1 is much less sensitive at moderateddbecause there is no boundary to recover\. In a real high\-dimensional application, the localizer should be applied to a learned low\-dimensional representation rather than the raw covariates, as discussed in Section 4\.

![Refer to caption](https://arxiv.org/html/2607.29053v1/results/stage5/figures/fig_6p6_highdim_scaling.png)Figure 12:Covariate\-dimension ablation\. horizontal axis is the ambient dimensiondd, with the signal kept in the first coordinate and the remainingd−1d\-1coordinates filled with uniform nuisance\.

## Appendix DAppendix: Compute Resources

All experiments were run on a single MacBook Pro with an Apple M3 chip \(8 physical cores\), 16 GB unified memory, and macOS 14\.6 \(Sonoma\)\. No GPU was used; all computation relied on the CPU\. The software environment used Python 3\.12\.7 with NumPy 1\.26\.4, pandas 2\.2\.3, scikit\-learn 1\.6\.1, LightGBM 4\.5\.0, matplotlib 3\.7\.5, and SciPy 1\.15\.2\.

For the synthetic experiments in Section[6\.1](https://arxiv.org/html/2607.29053#S6.SS1), a full run over all 10 random seeds took approximately 10 minutes\. For the real\-data experiments in Section[6\.2](https://arxiv.org/html/2607.29053#S6.SS2), a full run over all benchmark datasets took a similar amount of time\. Across all experiments reported in the paper, the total compute budget was well under 1 CPU\-hour\. The experiments consist of small\-scale regression tasks and do not require large\-scale distributed training\.

We used 10 random seeds \(0–9\) for all experiments\. No pretrained models were used\. For the synthetic experiments, modelAAis LightGBM and modelBBis linear regression, both trained from scratch on each data split using the hyperparameters described in Section[6\.1](https://arxiv.org/html/2607.29053#S6.SS1)\. For the real\-data experiments, modelAAis a small MLP \(scikit\-learnMLPRegressorwith dataset\-specific hidden layer sizes,max\_iter=1000, and early stopping\) and modelBBis linear regression; both are trained from scratch on each split as described in Section[6\.2](https://arxiv.org/html/2607.29053#S6.SS2)\. No deep learning framework beyond scikit\-learn was required\.

Similar Articles

Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations

arXiv cs.CL

This paper proposes a conformal prediction framework for LLMs that leverages internal representations rather than output-level statistics, introducing Layer-Wise Information (LI) scores as nonconformity measures to improve validity-efficiency trade-offs under distribution shift. The method demonstrates stronger robustness to calibration-deployment mismatch compared to text-level baselines across QA benchmarks.

Empirical Bayes Conformal Prediction for Vision and Language Models

arXiv cs.LG

This paper introduces an empirical Bayes conformal prediction framework that uses r-values to incorporate score variability into nonconformity scores, improving ranking stability and reducing set size while preserving coverage for vision and language models.

Online Localized Conformal Prediction

arXiv cs.LG

This paper proposes Online Localized Conformal Prediction (OLCP) to address covariate heterogeneity in online learning and time-series settings. It introduces OLCP-Hedge for bandwidth selection and demonstrates valid long-run coverage with narrower prediction sets compared to existing baselines.