When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift

arXiv cs.LG Papers

Summary

The paper introduces D^CF5, a diagnostic to predict regionwise gains in dynamic ensembling for regression tasks under distribution shift, validated with high correlation across datasets.

arXiv:2608.18330v1 Announce Type: new Abstract: Whether input-dependent ("dynamic") combination of a regression model pool beats the best static blend depends on the shift and is rarely known before deployment. Can a small labeled target-domain probe tell us when reallocating trust across regions of the input space will pay off? We answer this with $\widehat{D}_{\mathrm{CF5}}$, which estimates from the probe the cross-fitted gain of the regionwise convex combination over the best static convex blend: the realizable value of deciding, region by region, whom to trust. Across a frozen suite of 12 dataset-shift pairs (spatial, temporal, domain, feature-cluster), $\widehat{D}_{\mathrm{CF5}}$ predicts realized regionwise test gains with dataset-level Spearman $+0.98$ (95% CI $[+0.83, +1.00]$; $p=5\times10^{-5}$), including two cases overturning preregistered expectations. The relationship holds in a 16-pair sensitivity analysis (Spearman $+0.83$), whereas alternative probe diagnostics reach at most $+0.66$. This contrast isolates regional trust reallocation: correlation is $+0.98$ for regionwise-convex gain, but $+0.01$ for smooth covariate-dependent stacking after affine correction. A controlled generator shows dynamic gains arise from the interaction of shift heterogeneity and local competence, increase with shift severity, and become realizable between 128 and 256 probe labels in the tested grid. The Probe-Validated Ensemble Selector chooses among a static affine stacker and dynamic realizers, deploying a candidate only when a held-out lower confidence bound clears the static-convex floor. In a preregistered prospective batch, it matched or improved the floor in all 12 runs; two deployments reduced test risk by 11% and 16%, while the gate rejected a candidate whose un-gated deployment incurred $>30\times$ the static loss. We release OpenRegShift, a reproducible evaluation harness for regression ensembles under distribution shift.
Original Article
View Cached Full Text

Cached at: 08/20/26, 10:24 AM

# When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift
Source: [https://arxiv.org/html/2608.18330](https://arxiv.org/html/2608.18330)
August 2026

###### Abstract

Whether input\-dependent \(“dynamic”\) combination of a regression model pool beats the best static blend depends on the shift and is rarely known before deployment\. Can a small labeled target\-domain probe tell us when reallocating trust across regions of the input space will pay off? We answer this withD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}, which estimates from the probe the cross\-fitted gain of the regionwise convex combination over the best static convex blend: the realizable value of deciding, region by region, whom to trust\.

Across a frozen suite of 12 dataset\-shift pairs spanning spatial, temporal, domain and feature\-cluster shifts \(five used for development and seven reserved in two preregistered held\-out batches\),D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}predicts realized regionwise test gains with dataset\-level Spearman\+0\.98\+0\.98\(bootstrap 95% CI\[\+0\.83,\+1\.00\]\[\+0\.83,\+1\.00\]; permutationp=5×10−5p=5\\times 10^\{\-5\}; calibration slope0\.940\.94\), including two cases that overturned our preregistered expectations\. The relationship remains strong in a 16\-pair sensitivity analysis that adds four pairs from the prospective selector batch \(Spearman\+0\.83\+0\.83\)\. Alternative probe diagnostics reach at most\+0\.66\+0\.66\. The contrast isolates regional trust reallocation: its correlation is\+0\.98\+0\.98for regionwise\-convex gain but\+0\.01\+0\.01for the incremental value of smooth covariate\-dependent stacking after affine correction\.

A controlled generator shows that dynamic gains arise from the interaction of shift heterogeneity and local competence and increase with shift severity; in the tested grid, the transition occurs between 128 and 256 probe labels\. Complementing the diagnostic, the Probe\-Validated Ensemble Selector chooses among a static affine stacker and dynamic realizers, deploying a candidate only when a held\-out lower confidence bound clears the static\-convex floor\. In a preregistered prospective batch, it matched or improved that floor in all 12 runs; the two deployments reduced test risk by11%11\\%and16%16\\%, while the gate rejected a candidate whose unconditional deployment incurred over 30 times the static loss\.

We release OpenRegShift, a reproducible evaluation harness for regression ensembles under distribution shift\.

## 1Introduction

A practitioner has a pool of regression models trained on source\-domain data and a small labeled sample from a shifted target domain\. Should the models be averaged, combined with one global set of weights, or trusted differently in different parts of the input space? Input\-dependent combination is appealing: if different models stay accurate in different regions, a dynamic combiner can exploit their local complementarity\. The same flexibility can amplify probe noise when no such structure exists\. The practical question is therefore when the target data support dynamic combination\.

This question matters because static ensembling is a strong baseline\. Uniform averaging and deep ensembles are accurate and robust\([Lakshminarayanan et al\. 2017](https://arxiv.org/html/2608.18330#bib.bib15);[Ovadia et al\. 2019](https://arxiv.org/html/2608.18330#bib.bib19)\), while stacking learns a single combination from validation data\([Wolpert 1992](https://arxiv.org/html/2608.18330#bib.bib30)\)\. Dynamic alternatives such as mixture\-of\-experts, dynamic regressor selection, and covariate\-dependent stacking are far more expressive\([Jacobs et al\. 1991](https://arxiv.org/html/2608.18330#bib.bib10);[Moura et al\. 2019](https://arxiv.org/html/2608.18330#bib.bib17);[Wakayama and Sugasawa 2025](https://arxiv.org/html/2608.18330#bib.bib26)\)\. Yet their gains vary sharply across shifts, and neither shift severity nor model disagreement reveals whether the local structure they need is present\.

We formulate this deployment decision as an estimable comparison\. From a labeled probe, we learn a partition of the input space and compare the best static convex blend with a regionwise convex blend, with both fitted and evaluated out of fold\. The resulting cross\-fitted risk difference,D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}, estimates the realizable value of deciding, region by region, whom to trust\. Its sign answers the decision for this regional mechanism: a positive value means that local trust reallocation survives the estimation cost of a finite probe\.

Across a frozen suite of 12 dataset/shift pairs spanning spatial, temporal, domain and feature\-cluster shifts,D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}predicts realized regionwise gains with dataset\-level Spearman correlation\+0\.98\+0\.98, including two preregistered held\-out shifts that overturned our recorded expectations\. The correlation remains\+0\.83\+0\.83in a sensitivity analysis that adds four pairs from the prospective selector batch\. Simpler probe summaries reach at most\+0\.66\+0\.66\. The diagnostic is also mechanism\-specific: it correlates at\+0\.98\+0\.98with the regionwise gain it measures and at\+0\.01\+0\.01with the dynamic increment of covariate\-dependent stacking beyond affine correction\. It identifies the value of reallocating trust across regions, while deployment validation determines which available mechanism should realize the gain\.

A controlled generator explains when this value exists\. Gains arise from the interaction of shift heterogeneity and local competence; varying either condition alone produces little\. Gains increase with shift severity, with the transition occurring between 128 and 256 probe labels in the tested grid\.

We complement the diagnosis with a deployment procedure\. The Probe\-Validated Ensemble Selector evaluates a static affine stacker and several dynamic realizers against a static convex floor\. It deploys a candidate only when its held\-out lower confidence bound clears that floor\. In a preregistered prospective batch, the selector held the floor in all 12 runs\. Both dynamic deployments realized their validated gains, improving over the floor by11%11\\%and16%16\\%, while the gate rejected a candidate whose unconditional loss exceeded thirty times the static loss\.

Our contributions are:

- •D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}, a cross\-fitted diagnostic of the realizable regionwise gain over static convex blending, validated across diverse regression shifts\.
- •The Probe\-Validated Ensemble Selector, which converts a labeled target probe into a validated choice among a static floor and complementary dynamic realizers\.
- •Controlled factorial and dose\-response experiments that isolate the conjunctive mechanism behind dynamic gains and measure how realizability changes with shift severity and probe budget\.
- •OpenRegShift, a reproducible evaluation harness covering 16 dataset/shift pairs, a broad baseline suite, preregistered held\-out evaluations, and one\-command result verification\.

## 2The diagnostic

Let a frozen pool ofKKsource\-trained regression models produce the prediction vector

μ⁡\(x\)=\(μ1​\(x\),…,μK​\(x\)\)⊤∈ℝK\.\\mu\(x\)=\\bigl\(\\mu\_\{1\}\(x\),\\ldots,\\mu\_\{K\}\(x\)\\bigr\)^\{\\top\}\\in\\mathbb\{R\}^\{K\}\.\(1\)
At deployment, we observe a labeled probe from the target domain of total sizen=512n=512\. The probe is split once into a diagnostic sample, a combiner sample, and a gate sample\. The diagnostic sample supports the analysis in this section; the other two are reserved for the ensemble selector in[Section6](https://arxiv.org/html/2608.18330#S6)\. A disjoint target test set is used only for final evaluation\. We use squared error on standardized targets\.[Figure1](https://arxiv.org/html/2608.18330#S2.F1)summarizes the data flow, and the complete access and splitting protocol appears in[Table1](https://arxiv.org/html/2608.18330#S3.T1)\.

source datafrozen poolμ1,…,μK\\mu\_\{1\},\\ldots,\\mu\_\{K\}pool predictionson target probetarget proben=512n=512labeleddiagnosticsamplecombinersamplegatesampleD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}candidatesvalidated choice:candidate or floortarget test set \(evaluation only\)C1 comparisonFigure 1:Data flow\. The pool is trained on source data only and is evaluated at the target probe points\. The probe is split once into a diagnostic sample \(blue\), a combiner sample and a gate sample \(orange\)\. The diagnostic branch yieldsD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}; the selector branch fits candidates and validates them against the static convex floor\. The test set is not used for fitting or selection\. The two branches are complementary modules: the selector does not takeD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}as input\. The dashed blue arrow denotes the final C1 characterization analysis; test observations do not enter the diagnostic\.#### Population headroom\.

Let

ΔK=\{w∈ℝK:wk≥0,∑k=1Kwk=1\}\\Delta\_\{K\}=\\left\\\{w\\in\\mathbb\{R\}^\{K\}:w\_\{k\}\\geq 0,\\;\\sum\_\{k=1\}^\{K\}w\_\{k\}=1\\right\\\}\(2\)denote the probability simplex, and letπ:𝒳→\{1,…,J\}\\pi:\\mathcal\{X\}\\rightarrow\\\{1,\\ldots,J\\\}be a partition of the target input space\. We setJ=8J=8in all primary experiments and report a sweep overJJin[SectionA\.3](https://arxiv.org/html/2608.18330#A1.SS3)\. The optimal static convex risk is

Rstatic⋆=minw∈ΔK⁡𝔼T​\[\(Y−w⊤​μ​\(X\)\)2\]\.R\_\{\\mathrm\{static\}\}^\{\\star\}=\\min\_\{w\\in\\Delta\_\{K\}\}\\mathbb\{E\}\_\{T\}\\left\[\\bigl\(Y\-w^\{\\top\}\\mu\(X\)\\bigr\)^\{2\}\\right\]\.\(3\)The corresponding regionwise convex risk is

Rregional⋆​\(π\)=minw1,…,wJ∈ΔK⁡𝔼T​\[\(Y−wπ⁡\(X\)⊤​μ​\(X\)\)2\]\.R\_\{\\mathrm\{regional\}\}^\{\\star\}\(\\pi\)=\\min\_\{w\_\{1\},\\ldots,w\_\{J\}\\in\\Delta\_\{K\}\}\\mathbb\{E\}\_\{T\}\\left\[\\bigl\(Y\-w\_\{\\pi\(X\)\}^\{\\top\}\\mu\(X\)\\bigr\)^\{2\}\\right\]\.\(4\)Because the regionwise class contains every static convex blend,

D⁡\(π\)=Rstatic⋆−Rregional⋆​\(π\)≥0\.D\(\\pi\)=R\_\{\\mathrm\{static\}\}^\{\\star\}\-R\_\{\\mathrm\{regional\}\}^\{\\star\}\(\\pi\)\\geq 0\.\(5\)We callD⁡\(π\)D\(\\pi\)the*population regionwise headroom*\. It measures the structural value of reallocating model trust across the regions defined byπ\\pi\.

#### Realizable gain at probe budgetmm\.

Population headroom does not account for finite\-sample estimation\. Letπ^n\\widehat\{\\pi\}\_\{n\}denote the partition learned from the covariates of alln=512n=512target\-probe observations; partition fitting uses no target labels and no target test observations\. LetSm=𝒫DS\_\{m\}=\\mathcal\{P\}\_\{D\}denote the diagnostic sample, wherem=205m=205\. Let𝒜static​\(Sm\)\\mathcal\{A\}\_\{\\mathrm\{static\}\}\(S\_\{m\}\)denote static convex fitting, and let𝒜regional​\(Sm,π^n\)\\mathcal\{A\}\_\{\\mathrm\{regional\}\}\(S\_\{m\},\\widehat\{\\pi\}\_\{n\}\)denote regionwise convex fitting under the learned partition\. We set the minimum regional sample size tommin=5m\_\{\\min\}=5\. Any region containing fewer thanmminm\_\{\\min\}observations fromSmS\_\{m\}reuses the global static weights\. We define

Gm​\(π^n\)=𝔼⁡\[RT​\(𝒜static​\(Sm\)\)−RT​\(𝒜regional​\(Sm,π^n\)\)\|π^n\]\.G\_\{m\}\(\\widehat\{\\pi\}\_\{n\}\)=\\mathbb\{E\}\\left\[R\_\{T\}\\\!\\left\(\\mathcal\{A\}\_\{\\mathrm\{static\}\}\(S\_\{m\}\)\\right\)\-R\_\{T\}\\\!\\left\(\\mathcal\{A\}\_\{\\mathrm\{regional\}\}\(S\_\{m\},\\widehat\{\\pi\}\_\{n\}\)\\right\)\\;\\middle\|\\;\\widehat\{\\pi\}\_\{n\}\\right\]\.\(6\)The expectation is over the labeled diagnostic sample and its allocation, conditional on the learned partition\. Unlike population headroom,Gm​\(π^n\)G\_\{m\}\(\\widehat\{\\pi\}\_\{n\}\)can be negative: the regional advantage must be large enough to offset weight estimation\. Its sign therefore answers the deployment question for the regional mechanism\. The quality ofπ^n\\widehat\{\\pi\}\_\{n\}enters through the conditional gain, and the cost of learning the partition is isolated separately in[Section5](https://arxiv.org/html/2608.18330#S5)\.

#### Cross\-fitted estimation\.

We estimate the finite\-budget gain using three repetitions of region\-stratified five\-fold cross\-fitting within the diagnostic sample\. The partitionπ^n\\widehat\{\\pi\}\_\{n\}is learned once from all probe covariates and held fixed throughout\. Partition fitting uses no target labels\. Within every repetition and fold, the static and regional weights are fitted without the fold labels and evaluated on that fold\.

For repetitionr∈\{1,2,3\}r\\in\\\{1,2,3\\\}, let\{ℐr​q\}q=15\\\{\\mathcal\{I\}\_\{rq\}\\\}\_\{q=1\}^\{5\}denote the five folds\. Letf^static\(−r,q\)\\widehat\{f\}\_\{\\mathrm\{static\}\}^\{\(\-r,q\)\}andf^regional\(−r,q\)\\widehat\{f\}\_\{\\mathrm\{regional\}\}^\{\(\-r,q\)\}denote the two combiners fitted without foldℐr​q\\mathcal\{I\}\_\{rq\}\. The diagnostic is

D^CF5=115​∑r=13∑q=151\|ℐr​q\|​∑i∈ℐr​q\[ℓ⁡\(yi,f^static\(−r,q\)​\(xi\)\)−ℓ⁡\(yi,f^regional\(−r,q\)​\(xi\)\)\]\.\\widehat\{D\}\_\{\\mathrm\{CF5\}\}=\\frac\{1\}\{15\}\\sum\_\{r=1\}^\{3\}\\sum\_\{q=1\}^\{5\}\\frac\{1\}\{\|\\mathcal\{I\}\_\{rq\}\|\}\\sum\_\{i\\in\\mathcal\{I\}\_\{rq\}\}\\left\[\\ell\\\!\\left\(y\_\{i\},\\widehat\{f\}\_\{\\mathrm\{static\}\}^\{\(\-r,q\)\}\(x\_\{i\}\)\\right\)\-\\ell\\\!\\left\(y\_\{i\},\\widehat\{f\}\_\{\\mathrm\{regional\}\}^\{\(\-r,q\)\}\(x\_\{i\}\)\\right\)\\right\]\.\(7\)Here,ℓ⁡\(y,y^\)=\(y−y^\)2\\ell\(y,\\widehat\{y\}\)=\(y\-\\widehat\{y\}\)^\{2\}\. A positive value favors regionwise combination, while a value at or below zero favors the static convex blend\.

Each fold model is trained on approximately4​m/54m/5diagnostic observations\. Thus,D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}estimates the same regional contrast at a slightly smaller training budget than the final refit on allmmobservations\.

#### Why cross\-fitting matters\.

An in\-sample contrast rewards additional regional flexibility mechanically, even when that flexibility does not generalize\. Cross\-validation separates weight fitting from loss evaluation and limits the optimism created by evaluating a flexible combiner on the training labels\([Arlot and Celisse 2010](https://arxiv.org/html/2608.18330#bib.bib1)\)\. We compare the two estimators in[Section4](https://arxiv.org/html/2608.18330#S4)\.

We define the effective number of estimable regions as

Jeff=∑j=1J\{∑i∈𝒫D\{π^n\(xi\)=j\}≥mmin\}\.J\_\{\\mathrm\{eff\}\}=\\sum\_\{j=1\}^\{J\}\\mathbf\{1\}\\\!\\left\\\{\\sum\_\{i\\in\\mathcal\{P\}\_\{D\}\}\\mathbf\{1\}\\\!\\left\\\{\\widehat\{\\pi\}\_\{n\}\(x\_\{i\}\)=j\\right\\\}\\geq m\_\{\\min\}\\right\\\}\.\(8\)WhenJeff≤1J\_\{\\mathrm\{eff\}\}\\leq 1, the regional mechanism cannot be diagnosed, soD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}is reported as non\-estimable rather than assigned a misleading value of zero\. WhenJeff<J/2J\_\{\\mathrm\{eff\}\}<J/2, we retain the estimate but flag the partition as having reduced resolution\. Under the settingJ=8J=8, the reduced\-resolution flag is therefore triggered when fewer than four regions are estimable\.

#### Alignment with realized gain\.

For the characterization experiments,*realized regional gain*is defined as

G^test=R^test​\(f^static\)−R^test​\(f^regional\)\.\\widehat\{G\}\_\{\\mathrm\{test\}\}=\\widehat\{R\}\_\{\\mathrm\{test\}\}\\left\(\\widehat\{f\}\_\{\\mathrm\{static\}\}\\right\)\-\\widehat\{R\}\_\{\\mathrm\{test\}\}\\left\(\\widehat\{f\}\_\{\\mathrm\{regional\}\}\\right\)\.\(9\)Both combiners are refitted on the complete diagnostic sample and evaluated on the disjoint target test set\. Thus,D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}and realized gain compare the same static and regional classes under the same partition and loss\. Their systematic budget difference is that cross\-fitting trains each fold model on approximately4​m/54m/5observations, whereas the reported test gain uses allmmdiagnostic observations\.

## 3Setup and protocol

We study regression under distribution shift with a fixed pool of source\-trained models and a limited labeled target probe\. For each dataset/shift pair, letPSP\_\{S\}andPTP\_\{T\}denote the source and target distributions\. The direction of the shift is defined before model training using time, geography, monitoring site, subject identity, domain identity, or a feature\-cluster holdout\.

Source observations are divided into training and validation sets\. All base regressors and preprocessing transformations are fitted using source data and frozen before target labels are accessed\. Target labels may be used to diagnose and fit ensemble combinations, but never to update the base regressors\. This protocol isolates the value of combining an existing model pool under shift\. The primary real\-data experiments use a heterogeneous pool of five regressors: gradient boosting, random forest, ridge regression, a multilayer perceptron, and nearest neighbors\. Each dataset/shift pair is evaluated with seeds\{0,1,2\}\\\{0,1,2\\\}, and losses are measured on targets standardized using source\-training statistics\. Dataset construction and model\-pool specifications appear in[SectionsB\.1](https://arxiv.org/html/2608.18330#A2.SS1)and[B\.3](https://arxiv.org/html/2608.18330#A2.SS3)\.

### 3\.1Target probe and access protocol

At deployment, exactlyn=512n=512labeled target observations are assigned to the probe\. A seeded permutation divides the probe into three disjoint samples:

\|𝒫D\|=205,\|𝒫C\|=215,\|𝒫G\|=92\.\|\\mathcal\{P\}\_\{D\}\|=205,\\qquad\|\\mathcal\{P\}\_\{C\}\|=215,\\qquad\|\\mathcal\{P\}\_\{G\}\|=92\.\(10\)
The diagnostic sample𝒫D\\mathcal\{P\}\_\{D\}receives 40% of the probe\. The remaining 60% forms the method sample, which is divided 70/30 into the combiner sample𝒫C\\mathcal\{P\}\_\{C\}and gate sample𝒫G\\mathcal\{P\}\_\{G\}\. All remaining target observations form the test set\. We fix the absolute probe budget across datasets rather than fixing its fraction of the available target data\.

The regional partition is fitted bykk\-means withJ=8J=8using the covariates of all 512 probe observations\. Inputs are standardized using source\-training statistics\. Partition fitting uses no target labels and no target test observations\. The learned partition is then held fixed for the diagnostic and selector branches\.

Table 1:Frozen information\-access protocol\. Target test observations are unavailable until final evaluation\.

## 4Does the diagnostic predict realizable gains?

We evaluate whetherD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}, computed from the diagnostic probe, predicts the realized regionwise gain on the disjoint target test set\. The primary analysis uses the frozen suite of 12 dataset/shift pairs: five development pairs and seven pairs from two preregistered held\-out batches\. Each pair is evaluated with three seeds\. The dataset/shift pair is the statistical unit, so we first average over seeds within each pair and then compute all primary statistics across the 12 dataset\-level means\.

### 4\.1Primary association

[Figure2](https://arxiv.org/html/2608.18330#S4.F2)comparesD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}with realized regional gain\. Their dataset\-level Spearman correlation is\+0\.979\+0\.979, with a bootstrap 95% confidence interval of\[\+0\.83,\+1\.00\]\[\+0\.83,\+1\.00\]and a one\-sided permutation\-testpp\-value of5×10−55\\times 10^\{\-5\}\. The Pearson correlation is\+0\.960\+0\.960, and the linear calibration of realized gain onD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}, fitted with an intercept, has slope0\.9370\.937and intercept\+0\.0039\+0\.0039\. Leave\-one\-dataset\-out Spearman correlations range from\+0\.973\+0\.973to\+1\.000\+1\.000, showing that the association is not driven by one shift\.

The bootstrap and permutation procedures treat the dataset/shift pair as the resampling unit\. The confidence interval is a percentile interval based on 10,000 dataset\-level bootstrap resamples\. The one\-sided permutation test uses 20,000 permutations of the realized\-gain values relative to the fixed diagnostic values, with a plus\-one correction\.

For the ten dataset/shift pairs whose mean realized gain lies outside the preregistered indeterminate interval of\|G^test\|<0\.002\\lvert\\widehat\{G\}\_\{\\mathrm\{test\}\}\\rvert<0\.002, the sign ofD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}matches the sign of realized gain in all ten cases\. Together with the dataset\-level rank correlation, this establishes a strong across\-dataset ordering and the decision boundary between gains and losses\. The near\-unit calibration slope summarizes the overall scale of the relationship rather than exact pair\-level magnitudes\. For example, flights has a more negativeD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}than appliances, while appliances incurs the larger realized loss\. Run\-specific deployment risk is handled by the held\-out validation stage in[Section6](https://arxiv.org/html/2608.18330#S6)\.

![Refer to caption](https://arxiv.org/html/2608.18330v1/figures/fig_c1_main_clean.png)Figure 2:D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}versus realized regionwise gain on the frozen 12\-pair suite\. Large markers show dataset\-level means and faded markers show the three individual seeds\. The dotted line is the identity reference\. Dataset\-level Spearman correlation is\+0\.98\+0\.98\(bootstrap 95% CI\[\+0\.83,\+1\.00\]\[\+0\.83,\+1\.00\]; permutationp<10−4p<10^\{\-4\}\)\. Per\-dataset 95%t2t\_\{2\}intervals appear in[Fig\.A\.1](https://arxiv.org/html/2608.18330#A1.F1)\.
### 4\.2Comparison with simpler diagnostics

We compareD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}with six alternatives computed from the same diagnostic probe\. Shift severity measures the increase in uniform ensemble loss relative to source validation\. Prediction disagreement is the average variance of the pool predictions\. The heterogeneity summary measures regional variation in uniform loss, while the competence summary measures the relative advantage of the locally best model\. We also consider their product and the in\-sample version of the static\-versus\-regional headroom contrast\.

For each diagnostic, we report its dataset\-level Spearman correlation with realized gain\. We additionally fit a linear calibration on 11 dataset means, predict the omitted dataset, and average the absolute leave\-one\-out errors\. Lower leave\-one\-out mean absolute error indicates a more useful quantitative prediction of gain\.

Table 2:Probe diagnostics on the frozen 12\-pair suite\. Spearman correlation is computed against realized regional gain\. LOO\-MAE is the leave\-one\-dataset\-out calibration error in standardized\-MSE units\.No alternative reaches a correlation above\+0\.66\+0\.66\. In particular, the in\-sample version of the same contrast is less predictive than its cross\-fitted counterpart\. Cross\-fitting therefore contributes more than a generic summary of shift magnitude, model diversity, or apparent regional flexibility\.

### 4\.3Held\-out evidence

The two preregistered held\-out batches test whether the relationship survives beyond the datasets used to develop the diagnostic\. In the first batch, our recorded expectations were wrong for two of five shifts\. We expected little or negative regional gain on gas\-turbine emissions, butD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}was positive and the realized gains were the largest in the suite\. We expected a weak positive gain for the ACS state shift, butD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}was near zero or negative and the realized gain was likewise non\-positive\. In both cases, the probe diagnostic corrected the direction of the recorded prior\.

The second held\-out batch added flight\-delay and electricity\-load shifts without changing the frozen diagnostic or analysis protocol\. Their inclusion preserved the dataset\-level relationship shown in[Fig\.2](https://arxiv.org/html/2608.18330#S4.F2)\.

### 4\.4Mechanism specificity

The diagnostic is designed for one mechanism: reallocating convex model trust across regions\. Its strong association does not imply that it predicts every form of input\-dependent combination\. To test this boundary, we compareD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}with the incremental test gain of covariate\-dependent stacking after accounting for static affine correction\.

The correlation remains\+0\.979\+0\.979for the regionwise\-convex gain thatD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}targets, but falls to\+0\.014\+0\.014for the incremental value of covariate\-dependent stacking over static affine stacking\. By contrast, measuring covariate\-dependent stacking against the weaker static convex baseline produces a correlation of\+0\.790\+0\.790, largely because both methods can benefit from affine correction\. The near\-zero conditional correlation shows thatD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}isolates regional convex trust reallocation rather than acting as a generic score for all dynamic combiners\.

### 4\.5Sensitivity across all 16 pairs

Four additional dataset/shift pairs came from batch 3, which was preregistered for prospective selector validation rather than C1 development\. Including their C1 quantities only as a sensitivity analysis reduces the dataset\-level Spearman correlation from\+0\.979\+0\.979to\+0\.829\+0\.829, with bootstrap 95% confidence interval\[\+0\.46,\+1\.00\]\[\+0\.46,\+1\.00\]and permutationp=5×10−5p=5\\times 10^\{\-5\}\. The corresponding 16\-pair plot appears in[SectionA\.2](https://arxiv.org/html/2608.18330#A1.SS2)\.

The added pairs also clarify the diagnostic’s operating boundary\. Protein structure contains a direction error, while the online\-news shift shows a more consequential magnitude error:D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}remains near zero while unconditional regional deployment can incur a very large test loss\. Thus,D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}provides a useful ranking and directional diagnosis for the regional mechanism, while a separate validation gate is needed to control deployment\-specific tail failures\. This motivates the Probe\-Validated Ensemble Selector in[Section6](https://arxiv.org/html/2608.18330#S6)\.

## 5When do dynamic gains exist?

The preceding section shows thatD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}predicts realizable regionwise gain\. We now isolate the conditions that create this gain and measure how shift severity and probe budget affect its realization\.

#### Controlled generator\.

Letri∈\{1,…,J⋆\}r\_\{i\}\\in\\\{1,\\ldots,J\_\{\\star\}\\\}denote the generator region of observationii, whereJ⋆=8J\_\{\\star\}=8, and letμ¯i=K−1​∑k=1Kμi​k\\bar\{\\mu\}\_\{i\}=K^\{\-1\}\\sum\_\{k=1\}^\{K\}\\mu\_\{ik\}\. The generator first controls base\-pool diversity and then applies region\-dependent damage:

μi​k\(ρ\)\\displaystyle\\mu\_\{ik\}^\{\(\\rho\)\}=\(1−ρ\)​μi​k\+ρ​μ¯i,\\displaystyle=\(1\-\\rho\)\\mu\_\{ik\}\+\\rho\\bar\{\\mu\}\_\{i\},\(11\)μ~i​k\\displaystyle\\widetilde\{\\mu\}\_\{ik\}=μi​k\(ρ\)\+d​ηri​\(h\)​\[1−κ​Sri​k​\(h\)\]\.\\displaystyle=\\mu\_\{ik\}^\{\(\\rho\)\}\+d\\,\\eta\_\{r\_\{i\}\}\(h\)\\bigl\[1\-\\kappa S\_\{r\_\{i\}k\}\(h\)\\bigr\]\.
The functionηr​\(h\)\\eta\_\{r\}\(h\)controls whether the direction of the damage is shared or varies across regions, whileSr​k​\(h\)S\_\{rk\}\(h\)controls which model is spared\. Ath=0h=0, the direction is constant and the same model is spared in every region\. Ath=1h=1, the direction alternates and the spared model rotates across regions; intermediate values use linear interpolation\. The severityddsets the damage magnitude, whileκ\\kappacontrols the degree of local sparing:κ=0\\kappa=0damages all models equally, whereasκ=1\\kappa=1applies the maximum region\-specific sparing defined bySr​k​\(h\)S\_\{rk\}\(h\)\. Finally,ρ\\rhohomogenizes the original pool, so1−ρ1\-\\rhoindexes generic base\-pool diversity\. Target labels remain unchanged\. The explicit definitions ofηr​\(h\)\\eta\_\{r\}\(h\)andSr​k​\(h\)S\_\{rk\}\(h\), together with the complete parameter grid, appear in[AppendixC](https://arxiv.org/html/2608.18330#A3)\.

The generator regions are unavailable to the combination method\. The method instead learns its ownJ=8J=8partition from probe covariates, allowing partition estimation and approximation error to enter the realized gain\.

#### Experimental design\.

We evaluate the\(h,κ\)\(h,\\kappa\)grid using 10 independently seeded model pools and three target splits per pool\. Splits are averaged within each pool, making the model pool the unit of inference\. Unless a control is being varied, we fixd=0\.5d=0\.5,n=512n=512, andρ=0\\rho=0\.

Realized gain is measured by the selector frozen before these experiments\. It fits regional, smooth covariate\-gated, and precision\-weighted candidates on the combiner split and deploys a candidate only after validation on the held\-out gate split against a static floor\. The deployment procedure is introduced in[Section6](https://arxiv.org/html/2608.18330#S6)\.

#### Heterogeneity and local competence act jointly\.

The three corners missing at least one condition have negligible realized gain:\+0\.0000\+0\.0000at\(h,κ\)=\(0,0\)\(h,\\kappa\)=\(0,0\),\+0\.0007\+0\.0007at\(0,1\)\(0,1\), and−0\.0000\-0\.0000at\(1,0\)\(1,0\), with all three 95% confidence intervals containing zero\. At\(1,1\)\(1,1\), realized gain reaches\+0\.0910\+0\.0910with 95% CI\[\+0\.0844,\+0\.0975\]\[\+0\.0844,\+0\.0975\]\.

Within each pool, we regress gain onhh,κ\\kappa, and their interaction\. The pool\-level interaction isβh​κ=\+0\.0830\\beta\_\{h\\kappa\}=\+0\.0830with 95% CI\[\+0\.0760,\+0\.0900\]\[\+0\.0760,\+0\.0900\]andt9=26\.8t\_\{9\}=26\.8\. The corresponding interaction is\+0\.0842\+0\.0842for the oracle ceiling and\+0\.0837\+0\.0837forD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}\. The diagnostic therefore recovers the same conjunctive structure as the realized gain\.

![Refer to caption](https://arxiv.org/html/2608.18330v1/figures/fig_c3_factorial_2x2.png)Figure 3:Controlled mechanism experiments\. \(a\) Realized gain over the\(h,κ\)\(h,\\kappa\)grid\. \(b\) Gain as shift severity increases\. \(c\) Gain under nested probe budgets with a fixed test set\. \(d\) Oracle loss as generic base\-pool diversity changes\. Results average three target splits within each of 10 independently seeded model pools; bands show pool\-level 95%t9t\_\{9\}confidence intervals\.
#### Severity and probe budget\.

Withh=κ=1h=\\kappa=1, realized gain,D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}, and the partition\-conditional oracle increase together with shift severity, as shown in panel \(b\) of[Fig\.3](https://arxiv.org/html/2608.18330#S5.F3)\. Severity enlarges an available regional advantage but does not create one without heterogeneity and local competence\.

The budget experiment in panel \(c\) of[Fig\.3](https://arxiv.org/html/2608.18330#S5.F3)uses nested probesn∈\{64,128,256,512,1024\}n\\in\\\{64,128,256,512,1024\\\}and a fixed test set\. The structural ceiling, computed using the generator regions and test\-label oracle weights, remains fixed at0\.1225±0\.00670\.1225\\pm 0\.0067\. The partition\-conditional oracle uses the probe\-fitted partition and rises from0\.0990\.099to0\.1120\.112as the partition improves\. Realized gain rises from0\.0050\.005to0\.0980\.098\. Gains are near zero at 64 labels, remain modest at 128, and become substantial at 256\. The practical transition in the tested grid therefore lies between 128 and 256 labels\.

#### Generic diversity is not local complementarity\.

Classical ensemble analysis relates prediction disagreement to the gain from global averaging\([Krogh and Vedelsby 1994](https://arxiv.org/html/2608.18330#bib.bib14)\)\. Our experiment separates this generic diversity from local complementarity\. Increasing base\-pool diversity reduces static\-oracle loss by0\.0580\.058but regional\-oracle loss by only0\.0220\.022, as shown in panel \(d\) of[Fig\.3](https://arxiv.org/html/2608.18330#S5.F3)\. Dynamic headroom consequently narrows from0\.1460\.146to0\.1110\.111\. Generic diversity mainly strengthens global combination; dynamic gain requires models that remain competent in different regions\.

## 6From diagnosis to deployment

The diagnostic results in[Section4](https://arxiv.org/html/2608.18330#S4)characterize realizable gain from regionwise trust allocation\. The mechanism experiments in[Section5](https://arxiv.org/html/2608.18330#S5)show when that gain exists\. Deployment presents a broader choice: a shift may favor regional combination, global affine correction, or smooth covariate\-dependent weighting\. We therefore use the Probe\-Validated Ensemble Selector to compare these mechanisms on held\-out target data\.

#### Candidate pool and static floor\.

The selector fits a static convex blend as its deployment floor\. It also fits six candidates on the combiner split𝒫C\\mathcal\{P\}\_\{C\}:

1. 1\.static affine stacking,y^​\(x\)=b\+∑k=1Kak​f^k​\(x\)\\widehat\{y\}\(x\)=b\+\\sum\_\{k=1\}^\{K\}a\_\{k\}\\widehat\{f\}\_\{k\}\(x\), with unrestricted coefficients and intercept\([Wolpert 1992](https://arxiv.org/html/2608.18330#bib.bib30)\);
2. 2\.CDST\-RBF, which assigns affine model weights through an RBF basis of the covariates\([Wakayama and Sugasawa 2025](https://arxiv.org/html/2608.18330#bib.bib26)\);
3. 3\.a Shrunken Regional Convex Combiner;
4. 4\.a neural mixture\-of\-experts gate over the frozen models\([Jacobs et al\. 1991](https://arxiv.org/html/2608.18330#bib.bib10)\);
5. 5\.probe\-fit inverse\-variance weighting;
6. 6\.a linear covariate gate with softmax model weights\.

For the regional candidate, letw^0\\widehat\{w\}\_\{0\}be the global convex weights andw^r\\widehat\{w\}\_\{r\}the convex weights fitted in regionrr\. The regional weights are

w~r=λr​w^r\+\(1−λr\)​w^0,λr=nrnr\+10\.\\widetilde\{w\}\_\{r\}=\\lambda\_\{r\}\\widehat\{w\}\_\{r\}\+\(1\-\\lambda\_\{r\}\)\\widehat\{w\}\_\{0\},\\qquad\\lambda\_\{r\}=\\frac\{n\_\{r\}\}\{n\_\{r\}\+10\}\.\(12\)
A region containing fewer than five observations from𝒫C\\mathcal\{P\}\_\{C\}usesw^0\\widehat\{w\}\_\{0\}\. This shrinkage reduces the variance of regional weights estimated from small samples\.

The three probe splits and their roles are summarized in[Fig\.4](https://arxiv.org/html/2608.18330#S6.F4)\.

Target proben=512n=512labeled pointsDiagnostic𝒫D\\mathcal\{P\}\_\{D\}205 pointsCombiner𝒫C\\mathcal\{P\}\_\{C\}215 pointsGate𝒫G\\mathcal\{P\}\_\{G\}92 pointsComputeD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}regional diagnosticCharacterizeregionwise mechanismFit floor and candidatesstatic convex floorsix candidatesHeld\-out validationchoose largestd¯c\\overline\{d\}\_\{c\}requireLCB^c\>0\\widehat\{\\mathrm\{LCB\}\}\_\{c\}\>0Deployvalidated candidateFallbackstatic convex flooryesnoFigure 4:Diagnostic and deployment workflow\. Probe labels have disjoint uses:𝒫D\\mathcal\{P\}\_\{D\}estimatesD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\},𝒫C\\mathcal\{P\}\_\{C\}fits the static floor and six candidates, and𝒫G\\mathcal\{P\}\_\{G\}validates the deployment decision\. The partition is learned from all probe covariates without using their labels\. The diagnostic characterizes the regional mechanism but does not enter the selector rule\.
#### Held\-out validation rule\.

Letf0f\_\{0\}denote the static convex floor andfcf\_\{c\}a candidate fitted on𝒫C\\mathcal\{P\}\_\{C\}\. For every observationi∈𝒫Gi\\in\\mathcal\{P\}\_\{G\}, define the improvement of candidateccover the floor as

di​c=ℓ⁡\(yi,f0​\(xi\)\)−ℓ⁡\(yi,fc​\(xi\)\)\.d\_\{ic\}=\\ell\\\!\\left\(y\_\{i\},f\_\{0\}\(x\_\{i\}\)\\right\)\-\\ell\\\!\\left\(y\_\{i\},f\_\{c\}\(x\_\{i\}\)\\right\)\.\(13\)Letd¯c\\overline\{d\}\_\{c\}andscs\_\{c\}denote the sample mean and standard deviation of these improvements on𝒫G\\mathcal\{P\}\_\{G\}\. The selector first chooses

c⋆=arg⁡maxc∈\{1,…,6\}​d¯c\.c^\{\\star\}=\\underset\{c\\in\\\{1,\\ldots,6\\\}\}\{\\arg\\max\}\\;\\overline\{d\}\_\{c\}\.\(14\)It then computes

LCB^c⋆=d¯c⋆−2\.42​sc⋆\|𝒫G\|,\|𝒫G\|=92\.\\widehat\{\\mathrm\{LCB\}\}\_\{c^\{\\star\}\}=\\overline\{d\}\_\{c^\{\\star\}\}\-2\.42\\,\\frac\{s\_\{c^\{\\star\}\}\}\{\\sqrt\{\|\\mathcal\{P\}\_\{G\}\|\}\},\\qquad\|\\mathcal\{P\}\_\{G\}\|=92\.\(15\)The prespecified critical value2\.422\.42approximates a one\-sided 95% Bonferroni correction for the six candidates\. The selector deploysfc⋆f\_\{c^\{\\star\}\}whenLCB^c⋆\>0\\widehat\{\\mathrm\{LCB\}\}\_\{c^\{\\star\}\}\>0\. Otherwise, it deploys the static convex floor\.

#### Retrospective and prospective evaluation\.

The final selector was first evaluated retrospectively on the original 12 dataset/shift pairs\. It was then frozen and evaluated on a preregistered prospective batch containing four additional pairs\. Table[3](https://arxiv.org/html/2608.18330#S6.T3)summarizes the two stages\. We defineΔ​L=Lselector−Lfloor\\Delta L=L\_\{\\mathrm\{selector\}\}\-L\_\{\\mathrm\{floor\}\}in standardized\-MSE units, so negative values favor the selector\.

Table 3:Evaluation of the Probe\-Validated Ensemble Selector\.In the retrospective evaluation, the selector deployed a candidate in 22 of 36 runs\. Static affine stacking accounted for 11 selections, CDST\-RBF for 10, and the regional candidate for one\. The selector used the floor in the remaining 14 runs and matched or improved it in all 36 observed runs\.

The prospective batch was preregistered before the four datasets were downloaded or inspected\. The selector used the floor in 10 of 12 runs and deployed CDST\-RBF in two Seoul Bike runs\. Both deployments reduced test loss relative to the floor by11%11\\%and16%16\\%, and all 12 runs matched or improved the floor\.

#### The value of held\-out validation\.

Across the 72 candidate and run comparisons in the prospective batch, 40 unconditional candidate deployments increased test loss\. In one News Popularity run, unconditional CDST\-RBF produced a loss30\.1830\.18times the static floor\. The selector rejected it and retained the floor\.

The Seoul Bike result illustrates the separation between diagnosis and deployment\. ItsD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}was near zero, and the regional candidate was never selected\. The selector instead validated CDST\-RBF in two runs, identifying a smooth covariate\-dependent mechanism outside the regional contrast measured byD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}\.

## 7Related work

#### Static ensemble combination\.

Uniform averaging remains a strong baseline for independently trained predictors, including deep ensembles\([Lakshminarayanan et al\. 2017](https://arxiv.org/html/2608.18330#bib.bib15);[Ovadia et al\. 2019](https://arxiv.org/html/2608.18330#bib.bib19)\)\. Classical ensemble analysis relates the benefit of averaging to prediction diversity\([Krogh and Vedelsby 1994](https://arxiv.org/html/2608.18330#bib.bib14)\)\. For majority\-vote classification,[Theisen et al\. 2023](https://arxiv.org/html/2608.18330#bib.bib23)characterize ensemble improvement through the disagreement\-error ratio\. Stacking instead estimates a global combination from held\-out predictions\([Wolpert 1992](https://arxiv.org/html/2608.18330#bib.bib30);[Breiman 1996](https://arxiv.org/html/2608.18330#bib.bib2)\)\. These approaches use one set of weights across the input space\. Our diagnostic measures whether relaxing this global rule region by region yields an out\-of\-sample gain large enough to offset finite\-probe estimation; plain model disagreement alone has a correlation only\+0\.02\+0\.02with that gain\.

#### Input\-dependent combination\.

Mixture\-of\-experts models learn a gate that assigns input\-dependent weights to specialized predictors\([Jacobs et al\. 1991](https://arxiv.org/html/2608.18330#bib.bib10)\)\. Dynamic selection similarly chooses an expert or subset of experts for each query, with most of the literature centered on classification\([Cruz et al\. 2018](https://arxiv.org/html/2608.18330#bib.bib5)\)\. Dynamic regressor selection extends local competence estimation to continuous outcomes\([Moura et al\. 2019](https://arxiv.org/html/2608.18330#bib.bib17)\)\. Covariate\-dependent stacking instead models smooth weight functions over the input space\([Wakayama and Sugasawa 2025](https://arxiv.org/html/2608.18330#bib.bib26)\)\. These methods provide different mechanisms for realizing local specialization\. HyRe instead uses a small labeled target sample to globally reweight ensemble heads through a generalized Bayesian update, including in regression under covariate shift\([Lee et al\. 2026](https://arxiv.org/html/2608.18330#bib.bib16)\)\. Its weights are shared across target inputs\.D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}addresses the preceding decision: whether regionwise trust reallocation is supported by the available target probe\. The selector in[Section6](https://arxiv.org/html/2608.18330#S6)then compares that regional mechanism with affine and smooth covariate\-dependent alternatives\.

#### Uncertainty\-based weighting\.

Regression models can learn input\-dependent predictive variances together with their conditional means\([Nix and Weigend 1994](https://arxiv.org/html/2608.18330#bib.bib18);[Kendall and Gal 2017](https://arxiv.org/html/2608.18330#bib.bib12)\)\. These estimates motivate precision weighting,wk​\(x\)∝1/σ^k2​\(x\)w\_\{k\}\(x\)\\propto 1/\\widehat\{\\sigma\}\_\{k\}^\{2\}\(x\)\. Its effectiveness under shift depends on whether the estimated variances continue to rank target errors\. In our experiments, source\-fitted precision weighting increased California Housing test MSE by approximately50%50\\%relative to uniform averaging\. We therefore evaluate source\-fitted and probe\-fitted variants, while the selector admits the probe\-fitted variant only after held\-out validation\.

#### Model evaluation under distribution shift\.

Distribution shift can degrade both predictive accuracy and uncertainty estimates\([Ovadia et al\. 2019](https://arxiv.org/html/2608.18330#bib.bib19)\), motivating benchmarks that preserve the structure of naturally occurring shifts\([Koh et al\. 2021](https://arxiv.org/html/2608.18330#bib.bib13)\)\. Under covariate shift, importance\-weighted cross\-validation estimates target risk by reweighting source observations with a target\-to\-source density ratio\([Sugiyama et al\. 2007](https://arxiv.org/html/2608.18330#bib.bib22)\)\. Our setting instead provides a small labeled sample from the target domain\. This permits direct target\-domain cross\-fitting without assuming that the shift is purely covariate shift or estimating a density ratio\. Test\-time adaptation instead updates a deployed model using unlabeled target batches\. Representative methods such as TENT use target\-batch normalization statistics and update affine normalization parameters by minimizing predictive entropy\([Wang et al\. 2021](https://arxiv.org/html/2608.18330#bib.bib27)\)\. Our protocol keeps the source\-trained regression models frozen and exposes only their predictions to the combination layer, so methods requiring access to model internals fall outside the evaluation scope\. OpenRegShift instead evaluates regression model combination across spatial, temporal, domain, and feature\-cluster shifts\.

#### Cross\-fitting and held\-out validation\.

Cross\-validation estimates out\-of\-sample risk by separating fitting from evaluation\([Arlot and Celisse 2010](https://arxiv.org/html/2608.18330#bib.bib1)\)\. Cross\-fitting extends this separation across folds and is central to double and debiased machine learning\([Chernozhukov et al\. 2018](https://arxiv.org/html/2608.18330#bib.bib3)\)\. We apply the same principle to a paired predictive\-risk contrast: both combination classes are fitted away from each observation used to computeD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}\. The diagnostic split𝒫D\\mathcal\{P\}\_\{D\}is separate from the gate split𝒫G\\mathcal\{P\}\_\{G\}, which validates deployment against the static convex floor\.

## 8Limitations and conclusion

#### Scope of the diagnostic\.

D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}estimates one specific contrast: the realizable gain of regionwise convex combination over static convex blending under a probe\-fitted partition\. It does not measure every form of input\-dependent adaptation\. Smooth covariate\-dependent weighting can provide gains whenD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}is near zero, as the Seoul Bike result in[Section6](https://arxiv.org/html/2608.18330#S6)demonstrates\. Extending the diagnostic to several structured mechanism classes, while preserving out\-of\-sample estimation, is a natural next step\.

The current implementation fixesJ=8J=8, uses hard regions, and evaluates standardized squared error\. These choices make the estimand transparent and align it with the deployed regional candidate, but they may miss structure at another resolution or across overlapping regions\. Cross\-fitted selection over multiple partitions, soft partitions, and alternative regression losses could broaden the diagnostic without changing its underlying risk\-comparison principle\.

#### Probe size and deployment validation\.

The method requires labeled target observations\. The controlled experiments in[Section5](https://arxiv.org/html/2608.18330#S5)show that the finite\-probe cost is substantial: gains are small at 64 and 128 labels, with the transition occurring between 128 and 256 labels in the tested grid\. Applications with smaller probes may require stronger structural assumptions, information sharing across regions, or sequential acquisition of target labels\.

The selector’s lower confidence bound is an asymptotic validation rule based on the held\-out gate split\. Its prospective 12\-run evaluation matched or improved the static floor in every run, but this empirical result is not a finite\-sample guarantee\. Truncating the loss differences would permit empirical\-Bernstein or betting\-based bounds for bounded means, providing finite\-sample control for the truncated risk contrast\([Waudby\-Smith and Ramdas 2024](https://arxiv.org/html/2608.18330#bib.bib28)\)\.

#### Empirical scope\.

The frozen primary analysis contains 12 dataset/shift pairs, with four additional pairs used for sensitivity and prospective selector validation\. Each real\-data experiment uses three seeds; the synthetic study uses 10 independently generated model pools with three splits per pool\. The two held\-out C1 batches and the prospective selector batch were within\-project preregistrations: the datasets were selected by the same research team, but the hypotheses, protocols, and acceptance criteria were frozen before evaluation\. The suite covers spatial, temporal, domain, and feature\-cluster shifts, primarily in tabular regression with frozen source\-trained model pools\. Evaluation on high\-dimensional inputs, larger pools, additional model families, and shifts collected from deployed systems would test how broadly the characterization extends\.

#### Conclusion\.

Dynamic ensembling pays when model failures vary across the input space and useful local competence remains\.D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}makes this regional opportunity measurable from a labeled target probe\. Across the frozen suite, it predicts the sign and ordering of realizable regionwise gains, the controlled experiments identify the interaction that produces those gains, and the Probe\-Validated Ensemble Selector chooses among complementary mechanisms while retaining a static convex fallback\. OpenRegShift packages these components into a reproducible protocol for studying regression ensembles under distribution shift\.

## References

- Arlot and Celisse \[2010\]Sylvain Arlot and Alain Celisse\.A survey of cross\-validation procedures for model selection\.*Statistics Surveys*, 4:40–79, 2010\.doi:10\.1214/09\-SS054\.URL[https://doi\.org/10\.1214/09\-SS054](https://doi.org/10.1214/09-SS054)\.
- Breiman \[1996\]Leo Breiman\.Stacked regressions\.*Machine Learning*, 24:49–64, 1996\.doi:10\.1007/BF00117832\.URL[https://doi\.org/10\.1007/BF00117832](https://doi.org/10.1007/BF00117832)\.
- Chernozhukov et al\. \[2018\]Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins\.Double/debiased machine learning for treatment and structural parameters\.*The Econometrics Journal*, 21\(1\):C1–C68, 2018\.doi:10\.1111/ectj\.12097\.URL[https://doi\.org/10\.1111/ectj\.12097](https://doi.org/10.1111/ectj.12097)\.
- Cortez et al\. \[2009\]Paulo Cortez, António Cerdeira, Fernando Almeida, Telmo Matos, and José Reis\.Modeling wine preferences by data mining from physicochemical properties\.*Decision Support Systems*, 47\(4\):547–553, 2009\.doi:10\.1016/j\.dss\.2009\.05\.016\.URL[https://doi\.org/10\.1016/j\.dss\.2009\.05\.016](https://doi.org/10.1016/j.dss.2009.05.016)\.
- Cruz et al\. \[2018\]Rafael M\. O\. Cruz, Robert Sabourin, and George D\. C\. Cavalcanti\.Dynamic classifier selection: Recent advances and perspectives\.*Information Fusion*, 41:195–216, 2018\.doi:10\.1016/j\.inffus\.2017\.09\.010\.URL[https://doi\.org/10\.1016/j\.inffus\.2017\.09\.010](https://doi.org/10.1016/j.inffus.2017.09.010)\.
- Ding et al\. \[2021\]Frances Ding, Moritz Hardt, John P\. Miller, and Ludwig Schmidt\.Retiring adult: New datasets for fair machine learning\.In*Advances in Neural Information Processing Systems*, volume 34, pages 6478–6490, 2021\.
- Fanaee\-T and Gama \[2014\]Hadi Fanaee\-T and João Gama\.Event labeling combining ensemble detectors and background knowledge\.*Progress in Artificial Intelligence*, 2\(2–3\):113–127, 2014\.doi:10\.1007/s13748\-013\-0040\-3\.URL[https://doi\.org/10\.1007/s13748\-013\-0040\-3](https://doi.org/10.1007/s13748-013-0040-3)\.
- Fernandes et al\. \[2015\]Kelwin Fernandes, Pedro Vinagre, and Paulo Cortez\.A proactive intelligent decision support system for predicting the popularity of online news\.In*Progress in Artificial Intelligence: 17th Portuguese Conference on Artificial Intelligence \(EPIA 2015\)*, Lecture Notes in Computer Science, pages 535–546\. Springer, 2015\.doi:10\.1007/978\-3\-319\-23485\-4˙53\.URL[https://doi\.org/10\.1007/978\-3\-319\-23485\-4\_53](https://doi.org/10.1007/978-3-319-23485-4_53)\.
- Hamidieh \[2018\]Kam Hamidieh\.A data\-driven statistical model for predicting the critical temperature of a superconductor\.*Computational Materials Science*, 154:346–354, 2018\.doi:10\.1016/j\.commatsci\.2018\.07\.052\.URL[https://doi\.org/10\.1016/j\.commatsci\.2018\.07\.052](https://doi.org/10.1016/j.commatsci.2018.07.052)\.
- Jacobs et al\. \[1991\]Robert A\. Jacobs, Michael I\. Jordan, Steven J\. Nowlan, and Geoffrey E\. Hinton\.Adaptive mixtures of local experts\.*Neural Computation*, 1991\.doi:10\.1162/neco\.1991\.3\.1\.79\.URL[https://doi\.org/10\.1162/neco\.1991\.3\.1\.79](https://doi.org/10.1162/neco.1991.3.1.79)\.
- Kelly et al\. \[2023\]Markelle Kelly, Rachel Longjohn, and Kolby Nottingham\.The UCI machine learning repository, 2023\.URL[https://archive\.ics\.uci\.edu](https://archive.ics.uci.edu/)\.
- Kendall and Gal \[2017\]Alex Kendall and Yarin Gal\.What uncertainties do we need in bayesian deep learning for computer vision?In*Advances in Neural Information Processing Systems*, volume 30, 2017\.URL[https://proceedings\.neurips\.cc/paper/2017/hash/2650d6089a6d640c5e85b2b88265dc2b\-Abstract\.html](https://proceedings.neurips.cc/paper/2017/hash/2650d6089a6d640c5e85b2b88265dc2b-Abstract.html)\.
- Koh et al\. \[2021\]Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton Earnshaw, Imran Haque, Sara M\. Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang\.WILDS: A benchmark of in\-the\-wild distribution shifts\.In*Proceedings of the 38th International Conference on Machine Learning*, volume 139 of*Proceedings of Machine Learning Research*, pages 5637–5664\. PMLR, 2021\.URL[https://proceedings\.mlr\.press/v139/koh21a\.html](https://proceedings.mlr.press/v139/koh21a.html)\.
- Krogh and Vedelsby \[1994\]Anders Krogh and Jesper Vedelsby\.Neural network ensembles, cross validation, and active learning\.In*Advances in Neural Information Processing Systems*, volume 7, pages 231–238\. MIT Press, 1994\.URL[https://proceedings\.neurips\.cc/paper/1994/hash/b8c37e33defde51cf91e1e03e51657da\-Abstract\.html](https://proceedings.neurips.cc/paper/1994/hash/b8c37e33defde51cf91e1e03e51657da-Abstract.html)\.
- Lakshminarayanan et al\. \[2017\]Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell\.Simple and scalable predictive uncertainty estimation using deep ensembles\.In*Advances in Neural Information Processing Systems*\. Curran Associates, Inc\., 2017\.
- Lee et al\. \[2026\]Yoonho Lee, Jonathan Williams, Henrik Marklund, Archit Sharma, Eric Mitchell, Anikait Singh, and Chelsea Finn\.Inference\-time alignment via hypothesis reweighting\.*Transactions on Machine Learning Research*, 2026\.URL[https://openreview\.net/forum?id=Q9p8LSEpiJ](https://openreview.net/forum?id=Q9p8LSEpiJ)\.
- Moura et al\. \[2019\]Thiago J\. M\. Moura, George D\. C\. Cavalcanti, and Luiz S\. Oliveira\.Evaluating competence measures for dynamic regressor selection\.In*2019 International Joint Conference on Neural Networks*\. IEEE, 2019\.doi:10\.1109/IJCNN\.2019\.8851835\.URL[https://doi\.org/10\.1109/IJCNN\.2019\.8851835](https://doi.org/10.1109/IJCNN.2019.8851835)\.
- Nix and Weigend \[1994\]David A\. Nix and Andreas S\. Weigend\.Estimating the mean and variance of the target probability distribution\.In*Proceedings of the 1994 IEEE International Conference on Neural Networks*, volume 1, pages 55–60\. IEEE, 1994\.doi:10\.1109/ICNN\.1994\.374138\.URL[https://doi\.org/10\.1109/ICNN\.1994\.374138](https://doi.org/10.1109/ICNN.1994.374138)\.
- Ovadia et al\. \[2019\]Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D\. Sculley, Sebastian Nowozin, Joshua V\. Dillon, Balaji Lakshminarayanan, and Jasper Snoek\.Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift\.In*Advances in Neural Information Processing Systems*\. Curran Associates, Inc\., 2019\.
- Pace and Barry \[1997\]R\. Kelley Pace and Ronald Barry\.Sparse spatial autoregressions\.*Statistics & Probability Letters*, 33\(3\):291–297, 1997\.doi:10\.1016/S0167\-7152\(96\)00140\-X\.
- Pedregosa et al\. \[2011\]Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake VanderPlas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay\.Scikit\-learn: Machine learning in python\.*Journal of Machine Learning Research*, 12:2825–2830, 2011\.
- Sugiyama et al\. \[2007\]Masashi Sugiyama, Matthias Krauledat, and Klaus\-Robert Müller\.Covariate shift adaptation by importance weighted cross validation\.*Journal of Machine Learning Research*, 8\(35\):985–1005, 2007\.URL[https://www\.jmlr\.org/papers/v8/sugiyama07a\.html](https://www.jmlr.org/papers/v8/sugiyama07a.html)\.
- Theisen et al\. \[2023\]Ryan Theisen, Hyunsuk Kim, Yaoqing Yang, Liam Hodgkinson, and Michael W\. Mahoney\.When are ensembles really effective?In*Advances in Neural Information Processing Systems*, volume 36\. Curran Associates, Inc\., 2023\.doi:10\.52202/075280\-0659\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/hash/30b6fa308e62ed52180c31ae3ba6bb0a\-Abstract\-Conference\.html](https://proceedings.neurips.cc/paper_files/paper/2023/hash/30b6fa308e62ed52180c31ae3ba6bb0a-Abstract-Conference.html)\.
- Tsanas et al\. \[2010\]Athanasios Tsanas, Max A\. Little, Patrick E\. McSharry, and Lorraine O\. Ramig\.Accurate telemonitoring of Parkinson’s disease progression by noninvasive speech tests\.*IEEE Transactions on Biomedical Engineering*, 57\(4\):884–893, 2010\.doi:10\.1109/TBME\.2009\.2036000\.URL[https://doi\.org/10\.1109/TBME\.2009\.2036000](https://doi.org/10.1109/TBME.2009.2036000)\.
- Vanschoren et al\. \[2013\]Joaquin Vanschoren, Jan N\. van Rijn, Bernd Bischl, and Luis Torgo\.OpenML: Networked science in machine learning\.*ACM SIGKDD Explorations Newsletter*, 15\(2\):49–60, 2013\.doi:10\.1145/2641190\.2641198\.
- Wakayama and Sugasawa \[2025\]Tomoya Wakayama and Shonosuke Sugasawa\.Ensemble prediction via covariate\-dependent stacking\.*Statistics and Computing*, 35\(6\):212, 2025\.doi:10\.1007/s11222\-025\-10739\-y\.URL[https://doi\.org/10\.1007/s11222\-025\-10739\-y](https://doi.org/10.1007/s11222-025-10739-y)\.
- Wang et al\. \[2021\]Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell\.Tent: Fully test\-time adaptation by entropy minimization\.In*International Conference on Learning Representations*, 2021\.URL[https://openreview\.net/forum?id=uXl3bZLkr3c](https://openreview.net/forum?id=uXl3bZLkr3c)\.
- Waudby\-Smith and Ramdas \[2024\]Ian Waudby\-Smith and Aaditya Ramdas\.Estimating means of bounded random variables by betting\.*Journal of the Royal Statistical Society Series B: Statistical Methodology*, 86\(1\):1–27, 2024\.doi:10\.1093/jrsssb/qkad009\.URL[https://doi\.org/10\.1093/jrsssb/qkad009](https://doi.org/10.1093/jrsssb/qkad009)\.
- Wickham \[2022\]Hadley Wickham\.*nycflights13: Flights that Departed NYC in 2013*, 2022\.URL[https://github\.com/hadley/nycflights13](https://github.com/hadley/nycflights13)\.R package version 1\.0\.2\.
- Wolpert \[1992\]David H\. Wolpert\.Stacked generalization\.*Neural Networks*, 1992\.doi:10\.1016/S0893\-6080\(05\)80023\-1\.URL[https://doi\.org/10\.1016/S0893\-6080\(05\)80023\-1](https://doi.org/10.1016/S0893-6080(05)80023-1)\.
- Zhang et al\. \[2017\]Shuyi Zhang, Bin Guo, Anlan Dong, Jing He, Ziping Xu, and Song Xi Chen\.Cautionary tales on air\-quality improvement in Beijing\.*Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences*, 473\(2205\):20170457, 2017\.doi:10\.1098/rspa\.2017\.0457\.URL[https://doi\.org/10\.1098/rspa\.2017\.0457](https://doi.org/10.1098/rspa.2017.0457)\.

## Appendix AAdditional C1 results

### A\.1Variation across seeds

The primary C1 analysis averages the three seeds within each dataset/shift pair and treats the 12 pairs as the independent units\.[FigureA\.1](https://arxiv.org/html/2608.18330#A1.F1)reports per\-dataset 95%t2t\_\{2\}confidence intervals\. With only three seeds,t2,0\.975=4\.303t\_\{2,0\.975\}=4\.303, so these intervals are necessarily wide\. The individual seed results are shown directly as faded points\.

![Refer to caption](https://arxiv.org/html/2608.18330v1/figures/fig_c1_main.png)Figure A\.1:The C1 association with per\-dataset 95%t2t\_\{2\}confidence intervals over three seeds\. Large markers show dataset means and faded markers show the individual seeds\. Statistical inference in the primary analysis treats the 12 dataset/shift pairs, rather than the 36 seed runs, as the independent units\.
### A\.2Sensitivity across all 16 pairs

Batch 3 was reserved for prospective selector validation \([SectionB\.4](https://arxiv.org/html/2608.18330#A2.SS4)\); its four pairs enter C1 only through this sensitivity analysis\. Across all 16 pairs, the dataset\-level Spearman correlation is\+0\.829\+0\.829with bootstrap 95% confidence interval\[\+0\.46,\+1\.00\]\[\+0\.46,\+1\.00\]and permutationp=5×10−5p=5\\times 10^\{\-5\}, compared with\+0\.979\+0\.979on the frozen 12\-pair suite\. Batch 3 alone has correlation\+0\.40\+0\.40; with only four pairs, we report this value without further interpretation\.

![Refer to caption](https://arxiv.org/html/2608.18330v1/figures/fig_c1_sensitivity16.png)Figure A\.2:Sensitivity analysis across all 16 dataset/shift pairs\. Markers show three\-seed dataset means\. Blue circles denote the frozen 12\-pair C1 suite and red stars denote the four pairs from the prospective selector batch\. The broken vertical axis preserves the extreme News result while resolving the remaining 15 pairs\.The News pair illustrates a limitation in magnitude prediction:D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}is slightly negative, correctly indicating no regional opportunity, but does not anticipate the−7\.13\-7\.13standardized\-MSE loss from unconditional regional deployment\. Protein produces a directional miss\. These cases motivate the independent deployment validation in[Section6](https://arxiv.org/html/2608.18330#S6)\.

### A\.3Sensitivity to partition resolution

The primary experiments fix the number of regions atJ=8J=8\. This choice was made using only the five development pairs, before evaluating either held\-out C1 batch\. We examinedJ∈\{2,4,8,16,32\}J\\in\\\{2,4,8,16,32\\\}while keeping the probe budget, cross\-fitting procedure, and regional fallback rule unchanged\. For every value ofJJ, bothD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}and realized test gain refer to the corresponding partition resolution\.

[TableA\.1](https://arxiv.org/html/2608.18330#A1.T1)reports two descriptive correlations\. The pooled correlation treats the 15 dataset\-by\-seed runs as observations\. The dataset\-level correlation first averages the three seeds, leaving five development datasets, and is therefore included only as a coarse summary\. Here,JeffJ\_\{\\mathrm\{eff\}\}is the number of regions containing at least five diagnostic observations\.

Table A\.1:Development\-only sensitivity to partition resolution\. Spearman correlations compareD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}with the realized test gain of the regionwise convex combiner constructed at the sameJJ\.The association remains positive throughout the sweep\. Small values ofJJprovide limited partition resolution, while large values divide the 205\-point diagnostic sample among many regions\. BothJ=8J=8andJ=16J=16order the five dataset means, butJ=8J=8retains twice as many diagnostic observations per requested region and requires fewer regional weight vectors\. We frozeJ=8J=8as the tradeoff used in all subsequent held\-out experiments\. AtJ=32J=32, the decline in pooled correlation is consistent with the finite\-probe estimation cost thatD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}is designed to capture\.

## Appendix BDataset and protocol details

### B\.1Datasets and shift construction

OpenRegShift contains 16 public dataset/shift pairs\. Twelve datasets come from the UCI Machine Learning Repository\[[Kelly et al\. 2023](https://arxiv.org/html/2608.18330#bib.bib11)\]\. Where an introducing publication exists, we cite it: Bike\[[Fanaee\-T and Gama 2014](https://arxiv.org/html/2608.18330#bib.bib7)\], Superconduct\[[Hamidieh 2018](https://arxiv.org/html/2608.18330#bib.bib9)\], PM2\.5\[[Zhang et al\. 2017](https://arxiv.org/html/2608.18330#bib.bib31)\], Wine\[[Cortez et al\. 2009](https://arxiv.org/html/2608.18330#bib.bib4)\], Parkinsons\[[Tsanas et al\. 2010](https://arxiv.org/html/2608.18330#bib.bib24)\], and News\[[Fernandes et al\. 2015](https://arxiv.org/html/2608.18330#bib.bib8)\]; the other six UCI datasets are cited through the repository entry\. The remaining data come from California Housing\[[Pace and Barry 1997](https://arxiv.org/html/2608.18330#bib.bib20)\], OpenML\[[Vanschoren et al\. 2013](https://arxiv.org/html/2608.18330#bib.bib25)\], Folktables\[[Ding et al\. 2021](https://arxiv.org/html/2608.18330#bib.bib6)\], andnycflights13\[[Wickham 2022](https://arxiv.org/html/2608.18330#bib.bib29)\]\.[TableB\.1](https://arxiv.org/html/2608.18330#A2.T1)lists their outcomes and experimental roles\.

The frozen C1 analysis contains five development pairs, five pairs from held\-out batch 1, and two pairs from held\-out batch 2\. Batch 3 contains four pairs preregistered for prospective selector validation\. These four pairs enter the 16\-pair C1 sensitivity analysis but not the frozen 12\-pair C1 claim\.

Table B\.1:Dataset sources, prediction outcomes, and experimental roles\. Dev denotes development; H1 and H2 denote the two held\-out C1 batches; P3 denotes the prospective selector batch\. Outcome transformations precede standardization by source\-training statistics\.The suite includes spatial, temporal, domain, and feature\-cluster shifts\. Every temporal split uses earlier observations as source data and later observations as target data\. No future observations enter source\-model training\. Spatial splits separate locations, while domain splits separate populations or data types\. For each feature\-cluster shift,kk\-means withk=5k=5, 10 initializations, and random seed 0 is applied to standardized covariates\. Among clusters containing at least 2000 observations, the target is the cluster whose centroid has the largest sum of Euclidean distances to the other four centroids\. Target labels do not enter this construction\.[TableB\.2](https://arxiv.org/html/2608.18330#A2.T2)reports the shift definitions and sample sizes used by the experiments\. Here,nSn\_\{S\}includes source training and source validation observations, andnTn\_\{T\}denotes the target sample before allocating the 512\-point probe\. PM2\.5, ACS Income, and Flights use pre\-specified sample caps for computational control\.

Table B\.2:Shift construction and experimental sample sizes\. The 512\-point target probe is drawn from thenTn\_\{T\}observations; the remaining target observations form the test set\.
### B\.2Target\-probe allocation

Each dataset provides a labeled target sample of sizenTn\_\{T\}\. For each seed, a seeded random permutation is generated\. The first 512 indices in the permuted order are assigned to the target probe, and all remaining indices form the test set\. The probe budget is fixed across datasets rather than defined as a fraction ofnTn\_\{T\}\. The three experimental seeds are\{0,1,2\}\\\{0,1,2\\\}\.

The probe is divided into the three disjoint subsets shown in[TableB\.3](https://arxiv.org/html/2608.18330#A2.T3)\. The diagnostic receives 40% of the probe\. The remaining 307 observations form the method sample, which is divided 70/30 between candidate fitting and gate validation\. Rounding produces the allocation below\.

Table B\.3:Allocation of the fixed 512\-point labeled target probe\.Within𝒫D\\mathcal\{P\}\_\{D\},D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}uses three repetitions of stratified five\-fold cross\-fitting\. In each fold, both the static convex and regionwise convex combiners are fitted on the training folds and evaluated on the held\-out fold\. Their paired out\-of\-fold loss difference forms the diagnostic estimate\. For C1 test evaluation, both combiners are then refitted on all 205 observations in𝒫D\\mathcal\{P\}\_\{D\}and evaluated on the disjoint test set\.

TheJ=8J=8partition is learned from the covariates of all 512 probe observations\. Partition fitting uses no target labels\. The labels in𝒫D\\mathcal\{P\}\_\{D\},𝒫C\\mathcal\{P\}\_\{C\}, and𝒫G\\mathcal\{P\}\_\{G\}remain restricted to the uses in[TableB\.3](https://arxiv.org/html/2608.18330#A2.T3)\. In particular, candidate fitting does not use gate labels\.

The test set contains the remainingnT−512n\_\{T\}\-512target observations\. No test covariates or labels enter partition fitting, diagnostic estimation, candidate fitting, or deployment selection\. Test labels are available only to the evaluator after the method has made its decision\. The smallest test set contains 1,087 observations, so every dataset satisfies the pre\-specified minimum of 1,000 test observations\.

### B\.3Model\-pool construction

The primary model pool containsK=5K=5regression models implemented with scikit\-learn\[[Pedregosa et al\. 2011](https://arxiv.org/html/2608.18330#bib.bib21)\]\. For each dataset, the source sample is divided 80/20 into source\-training and source\-validation sets\. Temporal datasets use a chronological split; the remaining datasets use a seeded random split\.

Continuous inputs are standardized using the source\-training mean and standard deviation\. Deterministic ordinal encodings are applied to categorical inputs before standardization\. The target transformation listed in[TableB\.1](https://arxiv.org/html/2608.18330#A2.T1)is applied before fitting, after which the outcome is standardized as

ystd=y−y¯SsS,y^\{\\mathrm\{std\}\}=\\frac\{y\-\\overline\{y\}\_\{S\}\}\{s\_\{S\}\},\(B\.1\)
wherey¯S\\overline\{y\}\_\{S\}andsSs\_\{S\}are computed from the source\-training outcomes\. All reported MSE values use these standardized outcome units\.[TableB\.4](https://arxiv.org/html/2608.18330#A2.T4)gives the primary heterogeneous pool\. Parameters not listed in the table use the scikit\-learn defaults fixed by the released software environment\.

Table B\.4:Models in the primary heterogeneous regression pool\.Each mean model is fitted on the source\-training set\. The source\-validation set is used to identify the best source model and to fit one residual\-variance head for each pool member\. For modelkk, let

ri​k=yistd−μ^k​\(xi\)r\_\{ik\}=y\_\{i\}^\{\\mathrm\{std\}\}\-\\widehat\{\\mu\}\_\{k\}\(x\_\{i\}\)denote its source\-validation residual\. A depth\-three histogram gradient\-boosting regressor predictslog⁡\(ri​k2\+10−8\)\\log\(r\_\{ik\}^\{2\}\+10^\{\-8\}\)fromxix\_\{i\}\. Its output defines

σ^k2​\(x\)=exp⁡\{max⁡\(h^k​\(x\),−10\)\},\\widehat\{\\sigma\}\_\{k\}^\{2\}\(x\)=\\exp\\left\\\{\\max\\\!\\left\(\\widehat\{h\}\_\{k\}\(x\),\-10\\right\)\\right\\\},\(B\.2\)whereh^k\\widehat\{h\}\_\{k\}is the fitted log\-variance head\. These estimates support the uncertainty\-weighted baselines; they do not alter the mean predictions\.

The primary uniform baseline averages the five members of the heterogeneous pool\. We also construct a homogeneous comparison pool containing five multilayer perceptrons with the same\(128,64\)\(128,64\)architecture and different initialization seeds\. This separates the behavior of a heterogeneous model pool from that of a same\-architecture deep ensemble\.

After source training and validation, all pool members and variance heads are frozen\. Target observations can change ensemble weights, but they cannot update the parameters of a source model\.

### B\.4Preregistration design

We used three within\-project preregistration stages\. Each record fixed the datasets, hypotheses, evaluation rules, and reporting requirements before the corresponding outcomes were evaluated\. The development suite was used to construct the diagnostic\. Held\-out batches H1 and H2 completed the frozen 12\-pair C1 analysis together with the five development pairs, while P3 was reserved for prospective validation of the final selector\. Table[B\.5](https://arxiv.org/html/2608.18330#A2.T5)summarizes these stages\.

Table B\.5:Preregistration stages and their roles in the analysis\. H1 and H2 belong to the frozen C1 suite\. P3 evaluates the final selector, its diagnostic results enter only the 16\-pair sensitivity analysis\.#### Frozen diagnostic rules\.

For H1 and H2, the target\-probe budget, split, partition, loss, model pool, and random seeds were fixed before evaluation\. In particular, the protocol usedn=512n=512,J=8J=8, standardized squared error, five\-fold cross\-fitting repeated three times, and seeds\{0,1,2\}\\\{0,1,2\\\}\. Two diagnostic decision rules were recorded:

ℛ1=\{regional,D^CF5\>0,static,D^CF5≤0,ℛ2=\{regional,LCB95⁡\(D^CF5\)\>0,static,LCB95⁡\(D^CF5\)≤0\.\\mathcal\{R\}\_\{1\}=\\begin\{cases\}\\mathrm\{regional\},&\\widehat\{D\}\_\{\\mathrm\{CF5\}\}\>0,\\\\ \\mathrm\{static\},&\\widehat\{D\}\_\{\\mathrm\{CF5\}\}\\leq 0,\\end\{cases\}\\qquad\\mathcal\{R\}\_\{2\}=\\begin\{cases\}\\mathrm\{regional\},&\\operatorname\{LCB\}\_\{95\}\(\\widehat\{D\}\_\{\\mathrm\{CF5\}\}\)\>0,\\\\ \\mathrm\{static\},&\\operatorname\{LCB\}\_\{95\}\(\\widehat\{D\}\_\{\\mathrm\{CF5\}\}\)\\leq 0\.\\end\{cases\}\(B\.3\)Forℛ2\\mathcal\{R\}\_\{2\}, the lower confidence bound was computed from the 15 paired fold\-level loss differences\. A realized test gain satisfying

\|G^test\|<0\.002\\left\|\\widehat\{G\}\_\{\\mathrm\{test\}\}\\right\|<0\.002was classified as indeterminate and excluded from sign hit or miss counts, but retained in the complete results\. Before H1, we recorded weak positive expectations for Parkinsons, KC House, and ACS Income, and nonpositive expectations for Appliances and Gas Turbine\. Before H2, both Flights and Electricity were assigned nonpositive expectations\. Gas Turbine and ACS Income produced the two reversals reported in the main text\.

#### Prospective selector criteria\.

P3 was preregistered after the selector and its six\-candidate set had been frozen\. The four P3 datasets were not used to design the selector\. The preregistration specified three acceptance criteria:

1. 1\.The worst deployed result could not exceed the static\-convex floor by more than2%2\\%in relative loss or0\.0020\.002in absolute standardized MSE\.
2. 2\.Every deployment had to achieve test gain greater than−0\.002\-0\.002, and the mean gain across deployments had to be positive\.
3. 3\.All seeds, gate decisions, candidate losses, fallbacks, and failed runs had to be retained\.

The recorded expectations were positive or neutral for Protein, nonpositive for News, uncertain for Traffic, and nonpositive for Seoul Bike\. P3 satisfied all three acceptance criteria, as reported in[Section6](https://arxiv.org/html/2608.18330#S6)\.

#### Status and amendments\.

These records are ”pre\-registered within the project” in the sense described in[Section8](https://arxiv.org/html/2608.18330#S8): the research hypotheses, protocols and acceptance criteria were frozen before the corresponding outcomes were observed\. Before P3 evaluated any datasets, a revision was made to rename ”static unconstrained stacking” to ”static affine stacking” and to clarify that the previous analysis of 36 runs on selectors was retrospective\. Dated preregistration records and revision logs are included in the artifact\.

## Appendix CSynthetic experiment details

### C\.1Generator construction

The controlled experiments use the Bike Sharing temporal shift as their base\. For each pool seed, we retain the target outcomes, covariates, and predictions of the source\-trained model pool\. The generator modifies the predictions while leaving the target outcomes unchanged\. Generator regions are obtained by applyingkk\-means withJ⋆=8J\_\{\\star\}=8and seed 777 to the standardized target covariates\. These regions define the data\-generating structure and are not provided to the ensemble method\. Letri∈\{1,…,J⋆\}r\_\{i\}\\in\\\{1,\\ldots,J\_\{\\star\}\\\}be the generator region of observationii, and let

μ¯i=1K​∑k=1Kμi​k\.\\overline\{\\mu\}\_\{i\}=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\mu\_\{ik\}\.Base\-pool homogenization is applied through

μi​k\(ρ\)=\(1−ρ\)​μi​k\+ρ​μ¯i\.\\mu\_\{ik\}^\{\(\\rho\)\}=\(1\-\\rho\)\\mu\_\{ik\}\+\\rho\\overline\{\\mu\}\_\{i\}\.\(C\.1\)Thus,ρ=0\\rho=0preserves the original pool andρ=1\\rho=1makes all model predictions identical before the synthetic damage is added\. The heterogeneity and sparing functions used in[Eq\.11](https://arxiv.org/html/2608.18330#S5.E11)are

ηr​\(h\)\\displaystyle\\eta\_\{r\}\(h\)=\(1−h\)\+h​\(−1\)r−1,\\displaystyle=\(1\-h\)\+h\(\-1\)^\{r\-1\},\(C\.2\)Sr​k​\(h\)\\displaystyle S\_\{rk\}\(h\)=\(1−h\)𝟙\{k=1\}\+h𝟙\{k=1\+\(\(r−1\)modK\)\}\\displaystyle=\(1\-h\)\\mathbbm\{1\}\\\{k=1\\\}\+h\\mathbbm\{1\}\\left\\\{k=1\+\\bigl\(\(r\-1\)\\bmod K\\bigr\)\\right\\\}\(C\.3\)The shifted prediction is therefore

μ~i​k=μi​k\(ρ\)\+d​ηri​\(h\)​\[1−κ​Sri​k​\(h\)\]\.\\widetilde\{\\mu\}\_\{ik\}=\\mu\_\{ik\}^\{\(\\rho\)\}\+d\\,\\eta\_\{r\_\{i\}\}\(h\)\\left\[1\-\\kappa S\_\{r\_\{i\}k\}\(h\)\\right\]\.\(C\.4\)
Ath=0h=0, the damage direction is constant and model 1 is spared in every region\. Ath=1h=1, the direction alternates and the spared model rotates across regions\. The parameterκ\\kappacontrols local sparing: whenκ=0\\kappa=0, every model receives the same regional damage; whenκ=1\\kappa=1, the selected model is fully spared\. The parameterddcontrols damage magnitude\.

The method does not observerir\_\{i\}\. For each run, it learns a separateJ=8J=8partition from the available probe covariates\. This distinction allows partition identification and approximation error to enter the realized results\.

### C\.2Parameter grid

The four experimental sweeps are listed in[TableC\.1](https://arxiv.org/html/2608.18330#A3.T1)\. The factorial surface varieshhandκ\\kappaover the same five\-point grid\. Each dose\-response experiment varies one quantity while fixing the remaining controls\.

Table C\.1:Parameter grid for the controlled experiments\.The primary design uses 10 model\-pool seeds and three target\-split seeds per pool\. The three splits are averaged within each pool, so the 10 independently trained pools are the units of inference\. Parameter combinations shared by more than one sweep are counted once, leaving 1110 distinct run cells\. For the budget experiment, each pool and split seed produces one permutation of the target observations\. The probes are nested prefixes,

𝒫64⊂𝒫128⊂𝒫256⊂𝒫512⊂𝒫1024,\\mathcal\{P\}\_\{64\}\\subset\\mathcal\{P\}\_\{128\}\\subset\\mathcal\{P\}\_\{256\}\\subset\\mathcal\{P\}\_\{512\}\\subset\\mathcal\{P\}\_\{1024\},and the test set is fixed as the observations after position 1024 in the permutation\. Changes along the budget curve therefore arise from the amount of probe information rather than changes in test composition\. Each probe follows the diagnostic, combiner, and gate proportions defined in[SectionB\.2](https://arxiv.org/html/2608.18330#A2.SS2)\.

The realized\-gain curves use the three\-candidate selector frozen for the synthetic experiment\. Its candidates are the Shrunken Regional Convex Combiner, a smooth covariate gate, and probe\-fit inverse\-variance weighting\. All candidates are fitted on the combiner split and validated on the gate split against a static\-convex floor\. The final six\-candidate deployment procedure is described in[Sections6](https://arxiv.org/html/2608.18330#S6)and[D](https://arxiv.org/html/2608.18330#A4)\.

### C\.3Pool\-level inference

For each pool, we first average the three target splits within every\(h,κ\)\(h,\\kappa\)cell\. We then fit

Gp​\(h,κ\)=β0​p\+βh​p​h\+βκ​p​κ\+βh​κ,p​h​κ\+εp​\(h,κ\)G\_\{p\}\(h,\\kappa\)=\\beta\_\{0p\}\+\\beta\_\{hp\}h\+\\beta\_\{\\kappa p\}\\kappa\+\\beta\_\{h\\kappa,p\}h\\kappa\+\\varepsilon\_\{p\}\(h,\\kappa\)\(C\.5\)across the 25 factorial cells\. The reported interaction is the mean of the 10 independently estimatedβh​κ,p\\beta\_\{h\\kappa,p\}values\. Confidence intervals and tests use a one\-samplettdistribution with nine degrees of freedom\. The same procedure is applied to realized gain,D^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}, and the oracle gain\. For realized gain, this gives

βh​κ=0\.0830,95%​CI=\[0\.0760,0\.0900\],t9=26\.8\.\\beta\_\{h\\kappa\}=0\.0830,\\qquad 95\\%~\\mathrm\{CI\}=\[0\.0760,0\.0900\],\\qquad t\_\{9\}=26\.8\.The interaction estimates are0\.08420\.0842for the oracle gain and0\.08370\.0837forD^CF5\\widehat\{D\}\_\{\\mathrm\{CF5\}\}\. The corresponding 95% intervals are\[0\.0770,0\.0914\]\[0\.0770,0\.0914\]and\[0\.0698,0\.0976\]\[0\.0698,0\.0976\]\. An earlier sensitivity run used three model pools and 10 target splits per pool\. Its realized gain at\(h,κ\)=\(1,1\)\(h,\\kappa\)=\(1,1\)was0\.08640\.0864, compared with0\.09100\.0910in the primary 10\-pool design\.

### C\.4Probe\-budget decomposition

The budget experiment distinguishes three quantities:

1. 1\.The structural ceiling uses the generator regions and fits convex weights directly on the fixed test outcomes\. It is unavailable to a deployed method and is constant innnby construction\.
2. 2\.The partition\-conditional oracle uses the partition learned from the probe but fits its convex weights on the fixed test outcomes\. Its gap from the structural ceiling contains partition estimation error and mismatch between thekk\-means partition class and the generator regions\.
3. 3\.Realized gain uses only the probe to learn the partition, estimate candidate parameters, and make the gate decision\.

The structural ceiling is0\.1225±0\.00670\.1225\\pm 0\.0067, using a 95%t9t\_\{9\}interval\. The partition\-conditional oracle increases from0\.0990\.099atn=64n=64to0\.1120\.112atn=1024n=1024\. Realized gain increases from0\.0050\.005to0\.0980\.098\. The first gap measures partition identification and approximation cost, the second adds weight\-estimation and validation cost\. In the homogenization experiment, moving from an identical pool \(ρ=1\\rho=1\) to the original diverse pool \(ρ=0\\rho=0\) reduces the static\-oracle loss by0\.05760\.0576and the regional\-oracle loss by0\.02210\.0221\. Generic diversity therefore benefits the static oracle about 2\.6 times as much in this experiment\. The regional advantage is instead generated by the conditional specialization controlled byhhandκ\\kappa\.

An initial generator version alternated the damage direction even whenh=0h=0, leaving regional structure active in a cell intended to contain no heterogeneity\. In the frozen generator,[Eq\.C\.2](https://arxiv.org/html/2608.18330#A3.E2)links direction heterogeneity tohh\. All primary results in[Fig\.3](https://arxiv.org/html/2608.18330#S5.F3)use this corrected definition\.

## Appendix DSelector and baseline implementations

This appendix documents the specific implementations of the various combiners used in the paper\. All fitting operations are performed on standardized target values \(using the mean and scaling parameters from the source domain training set\) and are restricted to data from the ”method\-probe split”; the diagnostic split𝒫D\\mathcal\{P\}\_\{D\}remains untouched by any deployed combiners, while the test split is strictly excluded from any fitting or gating steps\.

### D\.1The static convex floor

The “floor”OstaticO\_\{\\mathrm\{static\}\}represents the best fixed combination result achievable via a convex combination ofKKprediction sources for a given labeled dataset; it is obtained by solvingminw∈ΔK−1⁡1n​∥M​w−y∥2\\min\_\{w\\in\\Delta^\{K\-1\}\}\\tfrac\{1\}\{n\}\\lVert Mw\-y\\rVert^\{2\}, whereM∈ℝn×KM\\in\\mathbb\{R\}^\{n\\times K\}contains the mean predictions from each source\. We employ a “support enumeration” method to solve this problem exactly: for each of the2K−12^\{K\}\-1non\-empty supports, we solve the corresponding equality\-constrained KKT system, discard solutions containing negative components, and retain feasible minimizers\. ForK=5K=5, this entails solving 31 small systems of linear equations; notably, this is an exact solution method rather than an iterative one\. We verified the accuracy of this exact solver by comparing it against exponentiated gradient descent on the simplex \(3,000 iterations, learning rate 0\.5\), finding agreement to four decimal places; this solver is widely applied in scenarios whereK≤10K\\leq 10\.

The same solver is also used to define the “region\-level convex combiner”: the simplex optimization problem is solved independently for each region, whereas the global solution is retained for regions containing fewer thanmmin=5m\_\{\\min\}=5labeled probe points\. Since refining a single global convex combination into independent regional convex combinations can only reduce in\-sample loss, the resulting “headroom” is necessarily non\-negative by construction\.

### D\.2Candidate menu

This selector is agnostic to the specific candidate schemes: it validates any input menu against a ”floor\.” This fixed menu comprises six members, all fitted on the ”method\-train probe” \(approximately 215 data points\) using hyperparameters that were set prior to downloading the held\-out batches\.

1. 1\.Regional simplex solutions shrink toward the global solution:wr←λr​wr\+\(1−λr\)​wglobw\_\{r\}\\leftarrow\\lambda\_\{r\}w\_\{r\}\+\(1\-\\lambda\_\{r\}\)w\_\{\\mathrm\{glob\}\}, whereλr=nr/\(nr\+n0\)\\lambda\_\{r\}=n\_\{r\}/\(n\_\{r\}\+n\_\{0\}\)andn0=10n\_\{0\}=10; for regions withnr<5n\_\{r\}<5,wglobw\_\{\\mathrm\{glob\}\}is used directly\.
2. 2\.We employ covariate\-dependent stacking\[[Wakayama and Sugasawa 2025](https://arxiv.org/html/2608.18330#bib.bib26)\]with affine, non\-simplex weight functions of the formwk​\(x\)=μk\+E​\(x\)⊤​γkw\_\{k\}\(x\)=\\mu\_\{k\}\+E\(x\)^\{\\top\}\\gamma\_\{k\}, constructed usingMM\-dimensional Gaussian RBF basis functions; here,M=min⁡\(16,max⁡\(2,⌊n/10⌋\)\)M=\\min\(16,\\max\(2,\\lfloor n/10\\rfloor\)\), with basis function centers determined by thekk\-means algorithm and bandwidths set via the median pairwise distance heuristic\. The penalty terms for each model are estimated using the EM algorithm \(assumingγk∼𝒩⁡\(0,τk2​I\)\\gamma\_\{k\}\\sim\\mathcal\{N\}\(0,\\tau\_\{k\}^\{2\}I\)andλk=σ2/τk2\\lambda\_\{k\}=\\sigma^\{2\}/\\tau\_\{k\}^\{2\}\) over 100 iterations, with an analytical solution available for the M\-step\. Verification of the method’s fidelity is detailed in Appendix §[D\.5](https://arxiv.org/html/2608.18330#A4.SS5)\.
3. 3\.A single hidden layer with 32tanh\\tanhunits feeds into a softmax output layer representingKKexperts\. Training employs full\-batch gradient descent with a momentum of 0\.9, a learning rate of 0\.03, and anℓ2\\ell\_\{2\}penalty of10−410^\{\-4\}over 1,500 iterations\. The experts are members of a frozen pool; as access protocols prohibit retraining them, this constitutes a Mixture of Experts \(MoE\) variant where only the gating network is trained, rather than an end\-to\-end MoE\.
4. 4\.The variance prediction head for each model is refitted on a ”method probe”\. That is, depth 2 histogram gradient boosting based onlog⁡\(\(y−μk\)2\+10−8\)\\log\(\(y\-\\mu\_\{k\}\)^\{2\}\+10^\{\-8\}\), using weights proportional to1/σ^k2​\(x\)1/\\hat\{\\sigma\}\_\{k\}^\{2\}\(x\), normalized to simplex; the lower bound on log variance predictions is set to−10\-10\.
5. 5\.Softmax weights linear in the covariates,w⁡\(x\)∝exp⁡\(x~⊤​W\)w\(x\)\\propto\\exp\(\\tilde\{x\}^\{\\top\}W\)with an intercept column, fitted by 800 gradient steps on the stacked squared loss, learning rate0\.050\.05,ℓ2\\ell\_\{2\}penalty10−310^\{\-3\}\.
6. 6\.Unconstrained least squares with an intercept term—specifically,y^=b\+∑kak​μk\\hat\{y\}=b\+\\sum\_\{k\}a\_\{k\}\\mu\_\{k\}\(i\.e\., classical stacked regression without simplex constraints\)\. This method is structurally independent of the inputs; it is introduced precisely to distinguish between ”adapting toxx” and ”correcting the global level and scale\.”

Finally, the affine candidate method is precisely what creates a non\-nested relationship between this set of methods and the baseline \(floor\): during development runs, the row sums of the fitted CDST\-RBF weight vectors ranged from0\.480\.48to1\.681\.68, with individual components dropping as low as−3\.3\-3\.3—indicating that the method performs contraction and cancellation operations rather than simple mixing\. Consequently, the paper emphasizes that the critical constraint distinguishing these candidate methods is the simplex constraint, rather than whether or not they rely on input data\.

### D\.3The validation gate

Letℓsta​\(i\)\\ell\_\{\\mathrm\{sta\}\}\(i\)andℓc​\(i\)\\ell\_\{c\}\(i\)denote the ”floor” \(baseline lower bound\) and the squared error of candidate modelccon the ”gate split” \(comprising approximately 92 data points, disjoint from all fitting sets\), respectively, and letdc​\(i\)=ℓsta​\(i\)−ℓc​\(i\)d\_\{c\}\(i\)=\\ell\_\{\\mathrm\{sta\}\}\(i\)\-\\ell\_\{c\}\(i\)\. Among candidate models satisfyingdc¯\>0\\overline\{d\_\{c\}\}\>0, the selector choosesc⋆=arg⁡maxc⁡dc¯c^\{\\star\}=\\arg\\max\_\{c\}\\overline\{d\_\{c\}\}and deploys the model only if the following condition holds:

LCB⁡\(c⋆\)=dc⋆¯−t​sd⁡\(dc⋆\)ng\>0,\\mathrm\{LCB\}\(c^\{\\star\}\)\\;=\\;\\overline\{d\_\{c^\{\\star\}\}\}\\;\-\\;t\\,\\frac\{\\mathrm\{sd\}\(d\_\{c^\{\\star\}\}\)\}\{\\sqrt\{n\_\{g\}\}\}\\;\>\\;0,\(D\.1\)otherwise, the static convex combination floor is deployed\. This threshold is based on the one\-sided95%95\\%Student’s t\-quantile withdf≈91\\mathrm\{df\}\\approx 91degrees of freedom, adjusted via Bonferroni correction for the size of the candidate model set:t=2\.42t=2\.42for a set of six candidate models, andt=2\.13t=2\.13for the set of three candidate models used in the synthetic data study \(see Appendix[C](https://arxiv.org/html/2608.18330#A3)\)\. The floor itself is deployed unconditionally: it consumes 4 degrees of freedom across approximately 215 data points and requires no certificate\. An earlier development version included a gating mechanism for the floor against the ”uniform average” and allowed a fallback to the uniform average; this intermediate gating mechanism resulted in a normalized mean squared error \(MSE\) loss of 0\.03 on thepm25\_sitedataset and was therefore removed prior to processing held\-out batches\.

Note that the selection statistic and the certificate are computed based on the same gate split\. The Bonferroni correction controls for selection effects within the candidate set; the nominal level associated with the t\-statistic\-based threshold is merely an approximation, and the empirical ”harmlessness” documented in the paper forms the basis for our claim\.

### D\.4The three\-candidate selector of the synthetic study

These synthetic experiments predated the proposal of the ”six candidate schemes” and employed the v1\-version selectors: specifically, the ”Shrunken Regional Convex Combiner” \(with parametern0=10n\_\{0\}=10\), the linear softmax covariate gate, and ”probe\-fit inverse\-variance weighting\.” All schemes utilized gating against the same static convex floor att=2\.13t=2\.13\(one\-sided95%/395\\%/3confidence level, degrees of freedomdf≈91\\mathrm\{df\}\\approx 91\)\. The candidate scheme labeled ”smooth covariate gate” corresponds to the linear softmax gate described in item 5 above; it is not the CDST\-RBF estimator, which was only included when the ”six candidate schemes” were introduced\.

### D\.5Faithfulness of the CDST\-RBF implementation

Given that the CDST estimator significantly influences our conclusions, we validated it before utilizing its numerical results\. When formulated as a linear mixed modely=F​μ\+Z​γ\+εy=F\\mu\+Z\\gamma\+\\varepsilon\(whereF=MF=MandZ⁡\[i,\(k,m\)\]=Em​\(xi\)​μk​\(xi\)Z\[i,\(k,m\)\]=E\_\{m\}\(x\_\{i\}\)\\,\\mu\_\{k\}\(x\_\{i\}\)\), the penalty term in the paper’s objective function,∑kλk​γk⊤​γk\\sum\_\{k\}\\lambda\_\{k\}\\gamma\_\{k\}^\{\\top\}\\gamma\_\{k\}, corresponds exactly to the random\-effects prior withλk=σ2/τk2\\lambda\_\{k\}=\\sigma^\{2\}/\\tau\_\{k\}^\{2\}; consequently, the EM algorithm described in the paper is directly applicable\. Specifically, the validation comprises four checks \(corresponding file:experiments/test\_cdst\_em\.py\):

1. 1\.The observed\-data log\-likelihood is non\-decreasing at all 61 recorded iterations \(zero decreases\)\.
2. 2\.Based on data simulated using the model from the paper, the correlation coefficient between the reconstructed weight surface and the original surface is0\.9110\.911\. Since both the intercept and the basis coefficients depend on the specific representation and the basis used for fitting differs from the one used to generate the data, the identifiable target is the surface itself rather than these coefficients\.
3. 3\.The objective function in the paper employs the base predictionfk,−i​\(xi\)f\_\{k,\-i\}\(x\_\{i\}\)obtained by refitting the model after excluding observationii\. Under our experimental setup, the ensemble of base models is trained on ”source” data, while the stacking process is performed on a disjoint ”target” probe set; consequently, the observations used for stacking are not included in the training sets of any base models, renderingfk,−i​\(xi\)f\_\{k,\-i\}\(x\_\{i\}\)identical tofk​\(xi\)f\_\{k\}\(x\_\{i\}\)\(i\.e\.,maxi⁡\|fk−fk,−i\|\\max\_\{i\}\|f\_\{k\}\-f\_\{k,\-i\}\|is zero within machine precision\)\. Thus, our objective function is exactly equivalent to the one in the paper, rather than merely an approximation\.
4. 4\.Upon convergence of the EM algorithm, the penalized objective function becomes a convex quadratic function with an analytical solution \(i\.e\., a closed\-form minimizer\); the fixed\-point solution of the EM algorithm reaches this minimizer, with a relative error of only1\.3×10−61\.3\\times 10^\{\-6\}\. Twenty random perturbation tests conducted on\(τ2,σ2\)\(\\tau^\{2\},\\sigma^\{2\}\)showed no increase in the marginal log\-likelihood\.

There are two documented and validated improvements: the aforementioned ”Leave\-One\-Out” \(LOO\) procedure and the use of a median heuristic \(rather than manual specification\) to set the basis function bandwidth\.

## Appendix EReproducibility

All experiments were conducted on a laptop equipped with a CPU, with fixed dependency versions \(numpy2\.5\.1,scikit\-learn1\.9\.0,scipy1\.18\.0,pandas3\.0\.5,folktables0\.0\.12\) and random seeds\{0,1,2\}\\\{0,1,2\\\}\. The entire experimental pipeline can be re\-run end\-to\-end on a single machine within one hour; the most time\-consuming task was the C3 factor grid experiment \(approximately 33 minutes\)\.

Each data loader downloads data from the original source and verifies checksums; a single driver script manages the re\-execution of all experiments and validates six key metrics against an automated ”golden standard” check\. Re\-runs in a ”clean\-room” environment reproduced all 30 key metrics from the development phase \(5 datasets×\\times3 random seeds×\\times2 statistical metrics\), with deviations within0\.0020\.002\.

The 12 UCI datasets are released under the CC BY 4\.0 license \(attribution details are provided in Appendix[B\.1](https://arxiv.org/html/2608.18330#A2.SS1)\);kc\_house\(OpenML\) andnycflights13are released under the CC0 license\. The ACS tasks are built upon the MIT\-licensedfolktablespackage\[[Ding et al\. 2021](https://arxiv.org/html/2608.18330#bib.bib6)\]; the use of the underlying PUMS data complies with the U\.S\. Census Bureau’s terms of service\. The California Housing data \(StatLib\) lacks an explicit license, so we utilize the data without redistributing it; the experimental artifact does not contain any data files\.

## Appendix FCode availability

The OpenRegShift artifact: pool construction, diagnostic, selector, experiment scripts, preregistration documents and result manifests, will be released publicly upon publication\.

Similar Articles

Learning the Pareto Frontier of Predictive Models under Distribution Shift

arXiv cs.LG

This paper proposes Frontier Learning, a framework that combines representations and predictions from multiple black-box and white-box pretrained models to construct a unified target-domain representation, guaranteeing performance no worse than any individual reuse baseline under distribution shift. Evaluations on visual domain adaptation and clinical mortality prediction show consistent gains over strong baselines.

Efficient Bayesian Deep Ensembles via Analytic Predictive Inference

arXiv cs.LG

Introduces an efficient Bayesian deep ensemble method for predictive regression that combines low-dimensional ensemble representation, closed-form Bayesian aggregation, and independent ensemble training to achieve calibrated uncertainty estimates with computational efficiency.