Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection
Summary
This paper introduces the deployment-fidelity gap in decomposed algorithm selection, demonstrating that partition-level evaluations can differ from end-to-end system performance, with implications for reporting and benchmarking.
View Cached Full Text
Cached at: 09/15/26, 09:02 AM
# Partition Scores Are Not System Scores:Deployment-Fidelity Gaps in Decomposed Algorithm Selection
Source: [https://arxiv.org/html/2609.13785](https://arxiv.org/html/2609.13785)
Jiachen Zhang††thanks:Corresponding author:jiaz@uoregon\.eduYu TangAffiliation:Huazhong University of ScienceAffiliation:and TechnologyLi ZhuAffiliation:City University of Hong Kong
###### Abstract
Oracle\-style quantities, including virtual best solvers, selected\-portfolio VBS, virtual\-best encodings, and best\-in\-family summaries, are widely reported as upper bounds on what a deployable selector could achieve\. In decomposed algorithm selection, an analogous partition\-level score grants an oracle choice of the best algorithm within the selected family; once the family selector is fixed, the deployable system must replace that within\-family oracle with a learned within\-family selector\.
We define the*deployment\-fidelity gap*G\(R\)G\(R\)as the difference between partition\-level and deployable end\-to\-end utility and derive two accounting consequences: a per\-instance margin–regret stability condition that tells us when a partition\-time family choice is deployment\-optimal, and a sharp partition\-only identification interval that, when it strictly crosses zero, prevents the partition\-level report from certifying the deployable winner\.
Across five public algorithm\-selection benchmarks spanning tabular AutoML and combinatorial CSP/SAT, every decomposed pipeline has positiveG\(R\)G\(R\), ranging from0\.0120\.012on TabZilla to0\.130\.13on PROTEUS\-2014\. Four of ten decomposed\-versus\-flat decisions have sign\-changing point estimates; on PROTEUS\-2014, a3333\-point partition advantage shrinks to a2020\-point end\-to\-end advantage\. A training\-side validation gap\-correction diagnostic recovers the point\-estimate deployable sign on all four sign\-changing cells; it is a reporting aid, not a substitute for direct end\-to\-end evaluation\. Partition and end\-to\-end scores should be reported side by side\.
## 1Introduction
Algorithm selection \(AS\) systems increasingly use decomposed decisions: rather than selecting directly from a large portfolio, they first select an algorithm family and then a concrete algorithm within it\([Kerschke et al\., 2019](https://arxiv.org/html/2609.13785#bib.bib4);[Bischl et al\., 2016](https://arxiv.org/html/2609.13785#bib.bib3)\)\. Decomposition reduces a hard multiclass problem and exposes interpretable structure, but its evaluation can differ from its deployment\.
A common partition\-level evaluation asks whether the selected family contains a strong algorithm and then grants the best member of that family\. This upper\-bounds the partition, but it is not an end\-to\-end system: deployment must use a learned within\-family selector\. We call the discrepancy the*deployment\-fidelity gap*:
G\(R\)=Spart\(R\)−Se2e\(R\),G\(R\)\\;=\\;S\_\{\\mathrm\{part\}\}\(R\)\\;\-\\;S\_\{\\mathrm\{e2e\}\}\(R\),\(1\)whereSpartS\_\{\\mathrm\{part\}\}uses the within\-family oracle andSe2eS\_\{\\mathrm\{e2e\}\}uses the deployable selector\.G\(R\)G\(R\)is the utility lost when the oracle is replaced by the available within\-family selector\.
Across five public AS benchmarks \(B1: TabZilla, B2: TabRepo, B3: TALENT, B4: MAXSAT\-PMS\-2016, B5: PROTEUS\-2014\) spanning tabular AutoML\([McElfresh et al\., 2023](https://arxiv.org/html/2609.13785#bib.bib10);[Salinas and Erickson, 2024](https://arxiv.org/html/2609.13785#bib.bib11);[Ye et al\., 2024](https://arxiv.org/html/2609.13785#bib.bib12)\)and combinatorial search\([Bischl et al\., 2016](https://arxiv.org/html/2609.13785#bib.bib3);[Hurley et al\., 2014](https://arxiv.org/html/2609.13785#bib.bib14)\), all decomposed pipelines show positive deployment\-fidelity gaps, from1\.21\.2points on TabZilla to1313points on PROTEUS\-2014\. The gap changes conclusions: on TabRepo and MAXSAT\-PMS\-2016, partition\-level evaluation favours decomposition over flat selection but end\-to\-end point estimates change sign, with one B2 cell and both B4 cells having small end\-to\-end margins; on TabZilla the advantage is almost absorbed but remains positive end\-to\-end; on PROTEUS\-2014 a3333\-point partition advantage shrinks to2020points end\-to\-end \(ρ≈0\.41\\rho\\approx 0\.41\)\. Overall, the safety condition fails in4/104/10decomp\-vs\-flat deployment decisions, whereas all55pairwise\-vs\-multiclass outer\-selector comparisons are low\-signal\. The dominant risk is therefore whether the within\-family deployment step is included, not which outer selector is used\.
Prior work uses oracle baselines such as virtual best solvers \(VBS\) as upper bounds\([Xu et al\., 2008](https://arxiv.org/html/2609.13785#bib.bib1);[Bischl et al\., 2016](https://arxiv.org/html/2609.13785#bib.bib3);[Cameron et al\., 2016](https://arxiv.org/html/2609.13785#bib.bib5)\), and recent work studies split\- and scale\-induced benchmarking pitfalls\([Petelin and Cenikj, 2025](https://arxiv.org/html/2609.13785#bib.bib6)\)\. We study a different,*nested*oracle: the within\-family oracle introduced after a family selector has already acted\. The novelty is not that oracle scores are optimistic, but that a common decomposition score can replace one deployable stage with an oracle and change comparisons throughG\(R1\)−G\(R2\)G\(R\_\{1\}\)\-G\(R\_\{2\}\)\.
##### Contributions\.
1. 1\.Formal\.We defineG\(R\)G\(R\)and identify partition score as an oracle\-assisted upper bound rather than a system score\. Two accounting lemmas \(Lemmas[1](https://arxiv.org/html/2609.13785#Thmlemma1),[2](https://arxiv.org/html/2609.13785#Thmlemma2)\) support a per\-instance margin–regret stability theorem \(Theorem[1](https://arxiv.org/html/2609.13785#Thmtheorem1)\) and a sharp partition\-only identification theorem \(Theorem[2](https://arxiv.org/html/2609.13785#Thmtheorem2)\) showing when partition reports cannot certify the deployable winner\.
2. 2\.Empirical\.On five public AS benchmarks we separate*decomp\-vs\-flat*deployment decisions, where the safety condition fails in4/104/10comparisons and B2/B4 have sign\-changing point estimates, from*decomp\-vs\-decomp*outer\-selector comparisons, where pairwise and multiclass are nearly indistinguishable \(\|Apart\|<5×10−3\|A\_\{\\mathrm\{part\}\}\|<5\\times 10^\{\-3\}\)\. Margin–regret violation rates \(0\.230\.23–0\.560\.56\) and identification intervals \(all1515pairs cross zero\) corroborate the theorems\.
3. 3\.Practical\.A training\-side gap\-correction diagnostic recovers the point\-estimate deployable sign on all four sign\-changing cells; a deployment\-aware family selector shrinksG\(R\)G\(R\)on every benchmark and improvesSe2eS\_\{\\mathrm\{e2e\}\}on B1–B4, with a small PROTEUS trade\-off\. We give a checklist requiringSpartS\_\{\\mathrm\{part\}\},Se2eS\_\{\\mathrm\{e2e\}\},G\(R\)G\(R\), the safety condition, and the identification interval to be reported together \(§[6](https://arxiv.org/html/2609.13785#S6)\)\.
## 2Related Work
##### Oracle baselines and VBS\.
Algorithm selection has long used oracle baselines such as the Virtual Best Solver \(VBS\) as an upper bound on portfolio performance\([Xu et al\., 2008](https://arxiv.org/html/2609.13785#bib.bib1)\)\. Frameworks such as AutoFolio\([Lindauer et al\., 2015](https://arxiv.org/html/2609.13785#bib.bib2)\)and standardised AS scenarios in ASlib\([Bischl et al\., 2016](https://arxiv.org/html/2609.13785#bib.bib3)\)explicitly compare deployable selectors to this VBS, and the AS competitions normalise scores by the SBS–VBS interval\([Lindauer et al\., 2019](https://arxiv.org/html/2609.13785#bib.bib23)\)\.[Cameron et al\. \(2016\)](https://arxiv.org/html/2609.13785#bib.bib5)further document that VBS evaluation itself can be optimistically biased when solvers are randomised\. Oracle\-style scores beyond the full\-portfolio VBS are also routinely reported\.*Selected\-portfolio*VBS appears inkk\-portfolio studies\([Bach et al\., 2022](https://arxiv.org/html/2609.13785#bib.bib18)\)and in portfolio selection for automated algorithm selection\([Kostovska et al\., 2023](https://arxiv.org/html/2609.13785#bib.bib19)\);*subportfolio*virtual\-best scores are reported alongside the deployable system in hierarchical solver portfolios such as Proteus\([Hurley et al\., 2014](https://arxiv.org/html/2609.13785#bib.bib14)\), whose four families correspond to a CSP\-native solver branch and three SAT encodings of the input\. Encoding\-level selection has its own*virtual\-best encoding*as a standard upper bound\([Stojadinović and Marić, 2014](https://arxiv.org/html/2609.13785#bib.bib22);[Ulrich\-Oltean et al\., 2022](https://arxiv.org/html/2609.13785#bib.bib20);[Ulrich\-Oltean et al\., 2023](https://arxiv.org/html/2609.13785#bib.bib21)\), and tabular benchmarking literature reports*best\-in\-family*summaries – best deep model versus best gradient boosting model – as the primary comparison object\([McElfresh et al\., 2023](https://arxiv.org/html/2609.13785#bib.bib10);[Shmuel et al\., 2025](https://arxiv.org/html/2609.13785#bib.bib17)\)\. A literature audit of twelve such reporting objects, the claim each row supports, and the strength of that support is given in Appendix[R](https://arxiv.org/html/2609.13785#A18);[Shmuel et al\. \(2025\)](https://arxiv.org/html/2609.13785#bib.bib17)is a representative tabular instance in which the best\-of\-DL versus best\-of\-TE comparison is the headline conclusion and no within\-family deployable selector is specified\.
Closest in spirit,[Tornede et al\. \(2023\)](https://arxiv.org/html/2609.13785#bib.bib15)define an AS\-oracle over algorithm selectors and show that lifting the oracle to that meta level can degrade oracle performance; our setting differs in that the hidden oracle is over algorithms inside an already\-chosen family rather than over complete selectors, so the relevant loss is the within\-family deployment\-fidelity gapG\(R\)G\(R\)\.
##### Methodological critiques\.
[Petelin and Cenikj \(2025\)](https://arxiv.org/html/2609.13785#bib.bib6)identify two pitfalls in feature\-based AS benchmarking — leave\-instance\-out \(LIO\) splits and scale\-sensitive performance targets — both orthogonal to our concern: LIO controls*which*test instances are held out and scale sensitivity controls the numeric*target*, whereas the deployment\-fidelity gap concerns*which system*is evaluated\. Unlike split or target\-scale effects, the deployment\-fidelity gap changes the evaluated object itself: an oracle\-assisted partition score replaces one stage of the deployable selector\.
##### Decomposition in algorithm selection and classification\.
Pairwise classification\([Sun and Pfahringer, 2013](https://arxiv.org/html/2609.13785#bib.bib8);[Galar et al\., 2011](https://arxiv.org/html/2609.13785#bib.bib7)\), multiclass selectors, and family\-level decompositions are well\-studied as*training*strategies, often analysed within the error\-correcting\-output\-code framework\([Allwein et al\., 2000](https://arxiv.org/html/2609.13785#bib.bib9)\)\. We study how such decompositions are*evaluated*once an inner within\-family decision is left to a learned selector at deployment time\.
## 3Setup
Letx∈𝒳x\\in\\mathcal\{X\}denote a problem instance \(e\.g\. a dataset or a SAT formula\),𝒜\\mathcal\{A\}a finite algorithm pool, andΠ=\{F1,…,FK\}\\Pi=\\\{F\_\{1\},\\dots,F\_\{K\}\\\}a partition of𝒜\\mathcal\{A\}into algorithm families \(Fi∩Fj=∅F\_\{i\}\\cap F\_\{j\}=\\emptysetfori≠ji\\neq jand⋃kFk=𝒜\\bigcup\_\{k\}F\_\{k\}=\\mathcal\{A\}\)\. Letu\(x,a\)∈ℝu\(x,a\)\\in\\mathbb\{R\}denote the utility of algorithmaaon instancexx; higher is better\. Throughout the paper we use a utility convention for both score and advantage\. Benchmarks whose native metric is a regret or a runtime are mapped to utilities by a*per\-instance*monotone transformation \(Appendix[A](https://arxiv.org/html/2609.13785#A1)\), which preserves the within\-instance ordering of algorithms; all reported magnitudes and advantages are therefore interpreted on the resulting normalised utility scale\.
A*decomposed selector*RRconsists of a family\-level selectorhR:𝒳→\{1,…,K\}h\_\{R\}:\\mathcal\{X\}\\rightarrow\\\{1,\\dots,K\\\}and, for each familykk, a within\-family selectorgR,k:𝒳→Fkg\_\{R,k\}:\\mathcal\{X\}\\rightarrow F\_\{k\}\. Letk^R\(x\)=hR\(x\)\\hat\{k\}\_\{R\}\(x\)=h\_\{R\}\(x\)denote the predicted family, and define the family\-kkoracle\-best algorithm
ak∗\(x\)=argmaxa∈Fku\(x,a\)\.a\_\{k\}^\{\*\}\(x\)\\;=\\;\\arg\\max\_\{a\\in F\_\{k\}\}u\(x,a\)\.\(2\)
##### Partition score\.
The*partition\-level*evaluation grants an oracle choice within the predicted family:
Spart\(R\)=𝔼x\[u\(x,ak^R\(x\)∗\(x\)\)\]\.S\_\{\\mathrm\{part\}\}\(R\)\\;=\\;\\mathbb\{E\}\_\{x\}\\bigl\[\\,u\\bigl\(x,\\,a\_\{\\hat\{k\}\_\{R\}\(x\)\}^\{\*\}\(x\)\\bigr\)\\,\\bigr\]\.\(3\)
##### End\-to\-end score\.
The*end\-to\-end*evaluation deploys the within\-family selector that is actually available:
Se2e\(R\)=𝔼x\[u\(x,gR,k^R\(x\)\(x\)\)\]\.S\_\{\\mathrm\{e2e\}\}\(R\)\\;=\\;\\mathbb\{E\}\_\{x\}\\bigl\[\\,u\\bigl\(x,\\,g\_\{R,\\hat\{k\}\_\{R\}\(x\)\}\(x\)\\bigr\)\\,\\bigr\]\.\(4\)
##### Deployment\-fidelity gap\.
G\(R\):=Spart\(R\)−Se2e\(R\)G\(R\):=S\_\{\\mathrm\{part\}\}\(R\)\-S\_\{\\mathrm\{e2e\}\}\(R\)measures how much partition\-level evaluation overstates the deployable system\. For two pipelinesR1,R2R\_\{1\},R\_\{2\}evaluated on the same test distribution, the partition and end\-to\-end advantages areApart:=Spart\(R1\)−Spart\(R2\)A\_\{\\mathrm\{part\}\}:=S\_\{\\mathrm\{part\}\}\(R\_\{1\}\)\-S\_\{\\mathrm\{part\}\}\(R\_\{2\}\)andAe2e:=Se2e\(R1\)−Se2e\(R2\)A\_\{\\mathrm\{e2e\}\}:=S\_\{\\mathrm\{e2e\}\}\(R\_\{1\}\)\-S\_\{\\mathrm\{e2e\}\}\(R\_\{2\}\)\.
## 4From Partition Scores to Deployment Stability
This section relates the four estimands of §[3](https://arxiv.org/html/2609.13785#S3)in three steps\. §[4\.1](https://arxiv.org/html/2609.13785#S4.SS1)records two accounting lemmas:G\(R\)G\(R\)equals within\-family regret, and advantage absorption equals the difference of gaps\. §[4\.2](https://arxiv.org/html/2609.13785#S4.SS2)states a margin–regret stability theorem characterising, instance by instance, when the family chosen at partition time is also deployment\-optimal\. §[4\.3](https://arxiv.org/html/2609.13785#S4.SS3)shows that even with selected\-family utility ranges added to a partition\-level report, the deployable advantage between two pipelines is identifiable only within an interval, and that when the interval strictly crosses zero the report cannot certify a deployable winner\.
### 4\.1Accounting lemmas
###### Lemma 1\(Gap identity\)\.
For any decomposed selectorRR,
G\(R\)=𝔼x\[u\(x,ak^R\(x\)∗\(x\)\)−u\(x,gR,k^R\(x\)\(x\)\)\]≥0,G\(R\)\\;=\\;\\mathbb\{E\}\_\{x\}\\Bigl\[\\,u\\bigl\(x,\\,a\_\{\\hat\{k\}\_\{R\}\(x\)\}^\{\*\}\(x\)\\bigr\)\-u\\bigl\(x,\\,g\_\{R,\\hat\{k\}\_\{R\}\(x\)\}\(x\)\\bigr\)\\,\\Bigr\]\\;\\geq\\;0,\(5\)with equality if and only if the within\-family selector chooses an oracle\-optimal algorithm inside the predicted family almost surely\.
###### Proof sketch\.
Substitute \([3](https://arxiv.org/html/2609.13785#S3.E3)\), \([4](https://arxiv.org/html/2609.13785#S3.E4)\) intoG\(R\)G\(R\)and use nonnegativity ofu\(x,ak^R\(x\)∗\(x\)\)−u\(x,gR,k^R\(x\)\(x\)\)u\(x,a\_\{\\hat\{k\}\_\{R\}\(x\)\}^\{\*\}\(x\)\)\-u\(x,g\_\{R,\\hat\{k\}\_\{R\}\(x\)\}\(x\)\)on each instance\. Full proof in Appendix[F](https://arxiv.org/html/2609.13785#A6)\. ∎
###### Lemma 2\(Advantage absorption\)\.
For any two decomposed selectorsR1,R2R\_\{1\},R\_\{2\},
Ae2e=Apart−\[G\(R1\)−G\(R2\)\],A\_\{\\mathrm\{e2e\}\}\\;=\\;A\_\{\\mathrm\{part\}\}\\;\-\\;\\bigl\[\\,G\(R\_\{1\}\)\-G\(R\_\{2\}\)\\,\\bigr\],\(6\)or equivalentlyApart−Ae2e=G\(R1\)−G\(R2\)A\_\{\\mathrm\{part\}\}\-A\_\{\\mathrm\{e2e\}\}=G\(R\_\{1\}\)\-G\(R\_\{2\}\)\.
###### Proof sketch\.
Se2e\(R\)=Spart\(R\)−G\(R\)S\_\{\\mathrm\{e2e\}\}\(R\)=S\_\{\\mathrm\{part\}\}\(R\)\-G\(R\)by definition; expandAe2e=Se2e\(R1\)−Se2e\(R2\)A\_\{\\mathrm\{e2e\}\}=S\_\{\\mathrm\{e2e\}\}\(R\_\{1\}\)\-S\_\{\\mathrm\{e2e\}\}\(R\_\{2\}\)\. Full proof in Appendix[F](https://arxiv.org/html/2609.13785#A6)\. ∎
The lemmas have two immediate consequences we use throughout the paper\. First \(Corollary[1](https://arxiv.org/html/2609.13785#Thmcorollary1)\),G\(R\)G\(R\)measures within\-family suboptimality: it is zero exactly when the within\-family selector is family\-oracle, regardless of whetherk^R\(x\)\\hat\{k\}\_\{R\}\(x\)is the globally optimal family\. Second \(Corollary[2](https://arxiv.org/html/2609.13785#Thmcorollary2)\), the partition\-level claim “RRbeatsFF” transfers to end\-to\-end evaluation iffApart\(R,F\)\>ΔG\(R,F\)A\_\{\\mathrm\{part\}\}\(R,F\)\>\\Delta G\(R,F\), whereΔG:=G\(R\)−G\(F\)\\Delta G:=G\(R\)\-G\(F\); for an ideal flat baseline \(G\(F\)=0G\(F\)=0\) this reduces toApart\>G\(R\)A\_\{\\mathrm\{part\}\}\>G\(R\)\.
###### Corollary 1\(Family error vs\. within\-family error\)\.
G\(R\)G\(R\)depends onk^R\(x\)\\hat\{k\}\_\{R\}\(x\), but not on whetherk^R\(x\)\\hat\{k\}\_\{R\}\(x\)is globally optimal: even ifRRchooses the wrong family, the gap is zero whenever the within\-family selector picks an oracle\-best algorithm in that wrong family; conversely, even with the globally optimal family selected,G\(R\)\>0G\(R\)\>0whenever the within\-family selector deviates fromak^R\(x\)∗\(x\)a\_\{\\hat\{k\}\_\{R\}\(x\)\}^\{\*\}\(x\)\.
###### Corollary 2\(General deployment\-safety condition\)\.
For any two decomposed selectorsRRandFF\(the latter typically a flat baseline\), the partition\-level conclusion “RRbeatsFF” transfers to end\-to\-end evaluation if and only ifApart\(R,F\)\>ΔG\(R,F\)A\_\{\\mathrm\{part\}\}\(R,F\)\>\\Delta G\(R,F\), whereΔG:=G\(R\)−G\(F\)\\Delta G:=G\(R\)\-G\(F\)\. WhenG\(F\)=0G\(F\)=0this reduces toApart\>G\(R\)A\_\{\\mathrm\{part\}\}\>G\(R\)\.
###### Proof sketch\.
Apply Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2):Ae2e\>0⇔Apart\>ΔGA\_\{\\mathrm\{e2e\}\}\>0\\iff A\_\{\\mathrm\{part\}\}\>\\Delta G\. ∎
Corollary[2](https://arxiv.org/html/2609.13785#Thmcorollary2)turns the apparently tautological identity \(G\(F\)=0G\(F\)=0for ideal flat baselines\) into a falsifiable empirical safety condition; §[5](https://arxiv.org/html/2609.13785#S5)documents how often it fails\. The generalΔG\\Delta Gform is kept because it covers arbitrary reference systems, while our singleton flat controls enforce the invariantG\(F\)=0G\(F\)=0by applying the same missing\-value fallback to partition and end\-to\-end singleton scores\.
### 4\.2Margin–regret stability of family selection
The lemmas describe*aggregate*accounting\. To say when partition conclusions transfer at the level of an individual instance, we decompose deployable utility into an oracle family margin and a within\-family regret, then read off when the family selected at partition time is deployment\-optimal\.
##### Per\-family quantities\.
For each instancexxand each familykkdefine
mk\(x\)=maxa∈Fku\(x,a\),vk\(x\)=u\(x,gR,k\(x\)\),rk\(x\)=mk\(x\)−vk\(x\),m\_\{k\}\(x\)\\;=\\;\\max\_\{a\\in F\_\{k\}\}u\(x,a\),\\hskip 20\.00003ptv\_\{k\}\(x\)\\;=\\;u\(x,\\,g\_\{R,k\}\(x\)\),\\hskip 20\.00003ptr\_\{k\}\(x\)\\;=\\;m\_\{k\}\(x\)\-v\_\{k\}\(x\),\(7\)sovk\(x\)=mk\(x\)−rk\(x\)v\_\{k\}\(x\)=m\_\{k\}\(x\)\-r\_\{k\}\(x\)\. Heremk\(x\)m\_\{k\}\(x\)is familykk’s oracle utility onxx,vk\(x\)v\_\{k\}\(x\)is the deployable utility realised byRR’s within\-family selectorgR,kg\_\{R,k\}inside familykk, andrk\(x\)≥0r\_\{k\}\(x\)\\geq 0is the within\-family regret\. Writings=k^R\(x\)s=\\hat\{k\}\_\{R\}\(x\)for the family that the partition\-time selector actually chooses, the deployable utility ofRRatxxis exactlyvs\(x\)v\_\{s\}\(x\)\.
###### Theorem 1\(Margin–regret deployment stability\)\.
Fix an instancexxand a family selectork^R\\hat\{k\}\_\{R\}, and writes=k^R\(x\)s=\\hat\{k\}\_\{R\}\(x\)\. The deployment regret of selecting familyssrelative to the best deployable family onxxis
Regdep\(k^R,x\)=maxq∈\[K\]vq\(x\)−vs\(x\)=maxq∈\[K\]\[\(mq\(x\)−ms\(x\)\)−\(rq\(x\)−rs\(x\)\)\]\.\\mathrm\{Reg\}\_\{\\mathrm\{dep\}\}\(\\hat\{k\}\_\{R\};\\,x\)\\;=\\;\\max\_\{q\\in\[K\]\}v\_\{q\}\(x\)\\;\-\\;v\_\{s\}\(x\)\\;=\\;\\max\_\{q\\in\[K\]\}\\bigl\[\\,\(m\_\{q\}\(x\)\-m\_\{s\}\(x\)\)\-\(r\_\{q\}\(x\)\-r\_\{s\}\(x\)\)\\,\\bigr\]\.\(8\)In particular,ssis deployment\-optimal atxxif and only if
ms\(x\)−mq\(x\)≥rs\(x\)−rq\(x\)for every competing familyq∈\[K\]\.m\_\{s\}\(x\)\-m\_\{q\}\(x\)\\;\\geq\\;r\_\{s\}\(x\)\-r\_\{q\}\(x\)\\hskip 10\.00002pt\\text\{for every competing family \}q\\in\[K\]\.\(9\)
###### Proof sketch\.
Substitutevk=mk−rkv\_\{k\}=m\_\{k\}\-r\_\{k\}and maximise overqq; non\-negativity of regret yields \([9](https://arxiv.org/html/2609.13785#S4.E9)\)\. Full proof in Appendix[F](https://arxiv.org/html/2609.13785#A6)\. ∎
##### Reading the theorem\.
The decomposition makes explicit what the aggregate identityAe2e=Apart−ΔGA\_\{\\mathrm\{e2e\}\}=A\_\{\\mathrm\{part\}\}\-\\Delta Ghides\. Deployment failure of familyssrelative to familyqqis controlled by a*differential*within\-family regretrs−rqr\_\{s\}\-r\_\{q\}measured against an oracle family marginms−mqm\_\{s\}\-m\_\{q\}:ssstays deployment\-optimal as long as it has a larger oracle margin than the gap between its within\-family selector and familyqq’s\. Two practical consequences\. First, the partition\-optimal familyp\(x\)∈argmaxkmk\(x\)p\(x\)\\in\\arg\\max\_\{k\}m\_\{k\}\(x\)remains deployment\-optimal atxxiffmp−mq≥rp−rqm\_\{p\}\-m\_\{q\}\\geq r\_\{p\}\-r\_\{q\}for everyq≠pq\\neq p; otherwise the partition\-time choicep\(x\)p\(x\)is dominated in deployment by some competitor with smaller oracle margin but smaller within\-family regret too\. Second, the Bayes\-optimal deployable outer target isvk\(x\)=mk\(x\)−rk\(x\)v\_\{k\}\(x\)=m\_\{k\}\(x\)\-r\_\{k\}\(x\), notmk\(x\)m\_\{k\}\(x\)\. Training a family selector againstmkm\_\{k\}optimises family oracle potential; training against \(a cross\-fitted estimator of\)vkv\_\{k\}optimises the quantity that determines deployment\-optimal family choice\. §[6](https://arxiv.org/html/2609.13785#S6)returns to this in describing our deployment\-aware family selector\.
A synthetic toy example with two singleton\-conflicting decompositions illustrates the general deployment\-safety condition of Corollary[2](https://arxiv.org/html/2609.13785#Thmcorollary2):Apart≈\+0\.20A\_\{\\mathrm\{part\}\}\\approx\+0\.20butAe2e≈−0\.24A\_\{\\mathrm\{e2e\}\}\\approx\-0\.24because the differential gap exceeds the partition advantage \(ρ≈2\.2\\rho\\approx 2\.2\)\. The per\-instance margin–regret condition explains the within\-pipeline mechanism; the cross\-pipeline sign change itself is governed by Corollary[2](https://arxiv.org/html/2609.13785#Thmcorollary2)\(Appendix[C](https://arxiv.org/html/2609.13785#A3)\)\.
### 4\.3Why partition\-level reports cannot certify deployable winners
Theorem[1](https://arxiv.org/html/2609.13785#Thmtheorem1)characterises one pipeline’s deployment stability\. We now turn to comparisons of two pipelinesR1,R2R\_\{1\},R\_\{2\}and ask: from a partition\-level report, can the reader certify the deployable winner? The answer is negative; even with selected\-family utility ranges added to the report, the deployable advantageAe2eA\_\{\\mathrm\{e2e\}\}is identifiable only within an interval that may cross zero\.
##### Selector classes and gap envelopes\.
For pipelineRRwith fixed family selectork^R\\hat\{k\}\_\{R\}, let𝒢R\\mathcal\{G\}\_\{R\}denote a class of admissible within\-family selectorsgg, each producing a gapGg\(R\)≥0G\_\{g\}\(R\)\\geq 0via Lemma[1](https://arxiv.org/html/2609.13785#Thmlemma1)\. Define the lower and upper gap envelopes
G¯\(R\):=infg∈𝒢RGg\(R\),G¯\(R\):=supg∈𝒢RGg\(R\)\.\\underline\{G\}\(R\)\\;:=\\;\\inf\_\{g\\in\\mathcal\{G\}\_\{R\}\}\\,G\_\{g\}\(R\),\\hskip 20\.00003pt\\overline\{G\}\(R\)\\;:=\\;\\sup\_\{g\\in\\mathcal\{G\}\_\{R\}\}\\,G\_\{g\}\(R\)\.\(10\)The unrestricted family\-respecting class \(every measurableggwithg\(x\)∈Fk^R\(x\)g\(x\)\\in F\_\{\\hat\{k\}\_\{R\}\(x\)\}\) givesG¯\(R\)=0\\underline\{G\}\(R\)=0\(oracle within\-family\) andG¯\(R\)=WR:=𝔼x\[maxa∈Fk^R\(x\)u\(x,a\)−mina∈Fk^R\(x\)u\(x,a\)\]\\overline\{G\}\(R\)=W\_\{R\}:=\\mathbb\{E\}\_\{x\}\\bigl\[\\,\\max\_\{a\\in F\_\{\\hat\{k\}\_\{R\}\(x\)\}\}u\(x,a\)\-\\min\_\{a\\in F\_\{\\hat\{k\}\_\{R\}\(x\)\}\}u\(x,a\)\\,\\bigr\]\(worst within\-family choice attainable as a measurable selector\)\.
###### Theorem 2\(Partition\-only identification\)\.
LetR1,R2R\_\{1\},R\_\{2\}be two decomposed selectors and let𝒢Ri\\mathcal\{G\}\_\{R\_\{i\}\}be admissible within\-family selector classes\. Suppose onlySpart\(R1\)S\_\{\\mathrm\{part\}\}\(R\_\{1\}\),Spart\(R2\)S\_\{\\mathrm\{part\}\}\(R\_\{2\}\), and the envelopesG¯\(Ri\),G¯\(Ri\)\\underline\{G\}\(R\_\{i\}\),\\overline\{G\}\(R\_\{i\}\)are reported\. Then
Ae2e∈\[Apart−G¯\(R1\)\+G¯\(R2\),Apart−G¯\(R1\)\+G¯\(R2\)\],A\_\{\\mathrm\{e2e\}\}\\;\\in\\;\\bigl\[\\,A\_\{\\mathrm\{part\}\}\-\\overline\{G\}\(R\_\{1\}\)\+\\underline\{G\}\(R\_\{2\}\),\\;\\;A\_\{\\mathrm\{part\}\}\-\\underline\{G\}\(R\_\{1\}\)\+\\overline\{G\}\(R\_\{2\}\)\\,\\bigr\],\(11\)and if the gap envelopesG¯\(Ri\),G¯\(Ri\)\\underline\{G\}\(R\_\{i\}\),\\overline\{G\}\(R\_\{i\}\)are attained by admissible selectors, the interval is sharp; otherwise the endpoints are approached arbitrarily closely by some choice of within\-family selectorsg1∈𝒢R1,g2∈𝒢R2g\_\{1\}\\in\\mathcal\{G\}\_\{R\_\{1\}\},g\_\{2\}\\in\\mathcal\{G\}\_\{R\_\{2\}\}\. Under the unrestricted family\-respecting class, the envelopes are attained by the within\-family oracle and the within\-family worst\-case selector, so the interval\[Apart−WR1,Apart\+WR2\]\[A\_\{\\mathrm\{part\}\}\-W\_\{R\_\{1\}\},\\,A\_\{\\mathrm\{part\}\}\+W\_\{R\_\{2\}\}\]is sharp without further assumption\.
###### Proof sketch\.
Apply Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2)to\(g1,g2\)\(g\_\{1\},g\_\{2\}\)and vary independently over𝒢Ri\\mathcal\{G\}\_\{R\_\{i\}\}\. Under the unrestricted class,G¯=0\\underline\{G\}=0is attained by the within\-family oracle andG¯=WR\\overline\{G\}=W\_\{R\}by the pointwise within\-family worst\-case selector\. Full proof in Appendix[F](https://arxiv.org/html/2609.13785#A6)\. ∎
###### Corollary 3\(No partition\-only sign certificate\)\.
If the interval in \([11](https://arxiv.org/html/2609.13785#S4.E11)\) strictly crosses zero, then the partition\-level report \(with envelopes\) does not certify the sign ofAe2eA\_\{\\mathrm\{e2e\}\}: there exist admissible\(g1,g2\)\(g\_\{1\},g\_\{2\}\)and\(g1′,g2′\)\(g\_\{1\}^\{\\prime\},g\_\{2\}^\{\\prime\}\)realising the same partition\-level scores and envelopes butAe2e\(g1,g2\)A\_\{\\mathrm\{e2e\}\}\(g\_\{1\},g\_\{2\}\)andAe2e\(g1′,g2′\)A\_\{\\mathrm\{e2e\}\}\(g\_\{1\}^\{\\prime\},g\_\{2\}^\{\\prime\}\)of opposite sign\. Conversely, if the interval is strictly positive \(resp\. strictly negative\), thenAe2eA\_\{\\mathrm\{e2e\}\}has the same sign as the interval for every admissible\(g1,g2\)\(g\_\{1\},g\_\{2\}\)\.
##### Scope qualifier\.
Theorem[2](https://arxiv.org/html/2609.13785#Thmtheorem2)is a statement about*partition\-level reports*, not about the underlying benchmark\. It does not say the deployable winner is unknowable: once a concrete within\-family selector is fixed and evaluated,Ae2eA\_\{\\mathrm\{e2e\}\}is just a number\. What it does say is that the partition\-level report, even augmented with selected\-family utility ranges, is insufficient as a*system\-performance certificate*when its identification interval strictly crosses zero\. Equivalently: full transfer of the partition conclusion requires either \(i\) reportingSe2eS\_\{\\mathrm\{e2e\}\}directly, which collapses the interval to a point, or \(ii\) demonstrating empirically \(e\.g\. via the margin–regret diagnostic of Theorem[1](https://arxiv.org/html/2609.13785#Thmtheorem1)and §[5\.4](https://arxiv.org/html/2609.13785#S5.SS4)\) that within\-family regret is small relative to the oracle margin on the relevant instances\.
##### Reversal taxonomy\.
Define the absorption ratioρ:=1−Ae2e/Apart=\(G\(R1\)−G\(R2\)\)/Apart\\rho:=1\-A\_\{\\mathrm\{e2e\}\}/A\_\{\\mathrm\{part\}\}=\(G\(R\_\{1\}\)\-G\(R\_\{2\}\)\)/A\_\{\\mathrm\{part\}\}, well\-defined when\|Apart\|\>0\|A\_\{\\mathrm\{part\}\}\|\>0\. We label transfer \(ρ≈0\\rho\\approx 0\), partial absorption \(0<ρ<10<\\rho<1\), extreme absorption \(ρ≈1\\rho\\approx 1\), reversal \(ρ\>1\\rho\>1,Apart\>0A\_\{\\mathrm\{part\}\}\>0\), and opposite reversal \(Apart<0A\_\{\\mathrm\{part\}\}<0,Ae2e\>0A\_\{\\mathrm\{e2e\}\}\>0\); the full taxonomy table is in Appendix[D](https://arxiv.org/html/2609.13785#A4)\.
##### Implication for empirical studies\.
The lemmas turn “*does within\-family suboptimality explain the gap?*” and “*is partition advantage equal to deployable advantage?*” from statistical hypotheses into definitional identities\. The two theorems turn the questions “*when is partition\-optimal family choice deployment\-stable?*” and “*can a deployable winner be certified from partition reports?*” into testable per\-instance and per\-pair conditions\. §[5\.4](https://arxiv.org/html/2609.13785#S5.SS4)reports two diagnostics that operationalise these conditions: a per\-instance margin–regret violation rate \(Theorem[1](https://arxiv.org/html/2609.13785#Thmtheorem1)\) and a partition\-only identification interval together with its crosses\-zero count \(Theorem[2](https://arxiv.org/html/2609.13785#Thmtheorem2)\)\. Pointwise computation requires a common valid\-instance set per comparison; the empirical residual of Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2)is below10−1510^\{\-15\}on every benchmark \(Appendix[G](https://arxiv.org/html/2609.13785#A7)\)\.
## 5Empirical Results
### 5\.1Benchmarks and protocol
Table[1](https://arxiv.org/html/2609.13785#S5.T1)summarises the five benchmarks: three tabular AutoML AS scenarios \(B1–B3\) and two ASlib combinatorial\-search scenarios \(B4–B5\)\. We use fixed matrix\-style resources because our estimands require paired algorithm\-by\-instance utilities under a stable protocol; living benchmarks such as TabArena\([Erickson et al\., 2025](https://arxiv.org/html/2609.13785#bib.bib16)\)are complementary\. Family selectors are multiclass RF, pairwise one\-vs\-one RF with soft voting, or a synthetic singleton\-family flat baseline\. The default within\-family selector is a per\-algorithm gradient\-boosted regressor; §[5\.5](https://arxiv.org/html/2609.13785#S5.SS5)sweeps three more classes\. We run55seeds×\\times55\-fold stratified CV and convert performance to per\-instance utilities in\[0,1\]\[0,1\]\(Appendix[A](https://arxiv.org/html/2609.13785#A1)\)\. B4–B5 family mappings follow solver\-competition reports\([Li and Manyà, 2009](https://arxiv.org/html/2609.13785#bib.bib13);[Hurley et al\., 2014](https://arxiv.org/html/2609.13785#bib.bib14)\)and appear verbatim in Appendix[B](https://arxiv.org/html/2609.13785#A2)\. PROTEUS\-2014 is partitioned semantically into CSP\-native solvers plus three SAT encoding families; its primary feature policy exposes only the ASlibcspstep \(3636features\), with an all\-feature sensitivity in Appendix[L](https://arxiv.org/html/2609.13785#A12)\.
All diagnostics and CIs use the test instance as the unit: repeated seed/fold predictions are aggregated per instance before instance\-cluster bootstrap \(Appendix[G](https://arxiv.org/html/2609.13785#A7)\)\.
Table 1:Benchmark registry\. Native metrics are converted to per\-instance utility \(App\.[A](https://arxiv.org/html/2609.13785#A1)\)\. B5 retains the common valid\-instance set after dropping all\-timeout rows from40214021raw PROTEUS\-2014 instances\.
### 5\.2RQ1: Do partition scores overstate deployable performance?
Figure 1:Deployment\-fidelity gapG\(R\)=Spart−Se2eG\(R\)=S\_\{\\mathrm\{part\}\}\-S\_\{\\mathrm\{e2e\}\}for every \(benchmark, decomposition\) cell\. Bars show95%95\\%percentile\-bootstrap confidence intervals on the per\-instance gap\. All decomposed pipelines exhibit positive gaps; flat baselines are included as controls\. The singleton\-flat invariant enforcesG\(F\)=0G\(F\)=0for every flat baseline by using identical missing\-value fallback on partition and end\-to\-end singleton scores\.Figure[1](https://arxiv.org/html/2609.13785#S5.F1)plotsG\(R\)G\(R\)across all1515cells\. Flat baselines are controls; the main claim concerns pairwise and multiclass decomposed pipelines, all with positive gaps and CIs excluding zero\. Gaps range from0\.0120\.012on TabZilla \(B1\) to approximately0\.130\.13on PROTEUS\-2014 \(B5\)\. Paired Wilcoxon tests yieldp<10−15p<10^\{\-15\}on every decomposed cell; fullpp\-values are in Appendix[O](https://arxiv.org/html/2609.13785#A15)\.
### 5\.3RQ2: When do partition\-level conclusions transfer?
For each benchmark we compare pairwise OVO vs\. flat, multiclass softmax vs\. flat, and pairwise vs\. multiclass\. The1515pairs split into*decomp\-vs\-flat*deployment decisions \(should we decompose?\) and*decomp\-vs\-decomp*outer\-selector comparisons; the regimes behave differently, so we report them separately\.
#### 5\.3\.1Decomp\-vs\-flat: is decomposition deployment\-safe?
Among the1010decomp\-vs\-flat comparisons in Table[2](https://arxiv.org/html/2609.13785#S5.T2),44have sign\-changing point estimates: B2 and B4 fail the safety conditionApart\>ΔGA\_\{\\mathrm\{part\}\}\>\\Delta G\(equal toApart\>G\(R\)A\_\{\\mathrm\{part\}\}\>G\(R\)here because every flat baseline hasG\(F\)=0G\(F\)=0\), so partition scores prefer decomposition while end\-to\-end point estimates prefer flat selection\. One B2 cell and the two B4 cells have small end\-to\-end margins \(B2 pw:−0\.0037\-0\.0037; B4:−0\.0047\-0\.0047and−0\.0031\-0\.0031\), so we treat them as small\-margin sign changes rather than large\-margin reversals\. B1 is a high\-absorption case: the partition advantage remains positive end\-to\-end but shrinks to11–22utility points \(ρ≈0\.84\\rho\\approx 0\.84–0\.910\.91\)\. B3 and PROTEUS\-2014 \(B5\) are partially absorbed; B5’s3333pp partition advantage shrinks to2020pp end\-to\-end \(ρ≈0\.41\\rho\\approx 0\.41\)\. Supporting plots are in Appendix[O](https://arxiv.org/html/2609.13785#A15)\.
#### 5\.3\.2Decomp\-vs\-decomp: do outer selectors differ?
The remaining55pairs \(pairwise vs\. multiclass\) all have\|Apart\|<5×10−3\|A\_\{\\mathrm\{part\}\}\|<5\\times 10^\{\-3\}\(Appendix[O](https://arxiv.org/html/2609.13785#A15)\); the outer\-selector choice is empirically secondary to whether the score includes the deployable within\-family selector\.
Table 2:Decomp\-vs\-flat advantage transfer\. “pw” = pairwise OVO, “mc” = multiclass softmax\.ρ\>1\\rho\>1marks sign change relative to the partition\-level advantage; small end\-to\-end margins should be read as point estimates rather than large\-margin reversals\. All quantities are recomputed on the four\-way intersection of finite values for the specific pair, so the gap values in this table need not match the per\-cell marginal gaps in Table[12](https://arxiv.org/html/2609.13785#A15.T12); on each pair,Apart−Ae2e=G\(R1\)−G\(R2\)A\_\{\\mathrm\{part\}\}\-A\_\{\\mathrm\{e2e\}\}=G\(R\_\{1\}\)\-G\(R\_\{2\}\)holds exactly \(Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2)\)\. Pairwise vs\. multiclass comparisons all have\|Apart\|<5×10−3\|A\_\{\\mathrm\{part\}\}\|<5\\times 10^\{\-3\}and are reported in Appendix[O](https://arxiv.org/html/2609.13785#A15)\(Table[13](https://arxiv.org/html/2609.13785#A15.T13)\)\.##### B5 robustness checks\.
Because PROTEUS\-2014 gives the largest gap, we test two design choices\. First, giving the selector all198198features \(including encoding\-conditional features unavailable before choosing an encoding\) improves deployment but leavesG\>0G\>0: pairwise0\.134→0\.1000\.134\\to 0\.100, multiclass0\.135→0\.0980\.135\\to 0\.098\. Second, shortening the runtime cutoff from36003600s to18001800s leaves the gap essentially unchanged \(pairwise0\.134→0\.1340\.134\\to 0\.134, multiclass0\.135→0\.1360\.135\\to 0\.136\)\. Thus the B5 headline effect is not an artefact of feature availability or timeout convention; full numbers are in Appendix[L](https://arxiv.org/html/2609.13785#A12)\.
### 5\.4RQ3: Theorem 1/2 diagnostics
Lemma[1](https://arxiv.org/html/2609.13785#Thmlemma1)/[2](https://arxiv.org/html/2609.13785#Thmlemma2)residuals are at machine precision \(Appendix[G](https://arxiv.org/html/2609.13785#A7)\), so we focus on the diagnostics induced by Theorems[1](https://arxiv.org/html/2609.13785#Thmtheorem1)and[2](https://arxiv.org/html/2609.13785#Thmtheorem2)\.
##### Theorem 1 diagnostic: per\-instance margin–regret stability\.
For each instancexxwe formviol\(x\)=𝟏\{∃q≠s:rs\(x\)−rq\(x\)\>ms\(x\)−mq\(x\)\}\\mathrm\{viol\}\(x\)=\\mathbf\{1\}\\\{\\exists q\\neq s:r\_\{s\}\(x\)\-r\_\{q\}\(x\)\>m\_\{s\}\(x\)\-m\_\{q\}\(x\)\\\}for the selected familys=k^R\(x\)s=\\hat\{k\}\_\{R\}\(x\)\. A violation means a competitor has higher deployable utility because its lower oracle margin is offset by lower within\-family regret\. The existential rate ranges0\.230\.23–0\.560\.56for pairwise and multiclass pipelines \(Table[11](https://arxiv.org/html/2609.13785#A15.T11)\); flat\-baseline rates are larger \(0\.490\.49–0\.670\.67\) because they count any singleton beating the single\-best\-on\-train algorithm\.
##### Theorem 2 diagnostic: partition\-only identification interval\.
For each comparison we report\[Apart−WR1,Apart\+WR2\]\[A\_\{\\mathrm\{part\}\}\-W\_\{R\_\{1\}\},A\_\{\\mathrm\{part\}\}\+W\_\{R\_\{2\}\}\], withWRW\_\{R\}the average selected\-family utility range\. All1515intervals strictly cross zero, including B5 \(Apart=\+0\.33A\_\{\\mathrm\{part\}\}=\+0\.33,WR≈0\.75W\_\{R\}\\approx 0\.75\), so partition reports plus family ranges do*not*certify the deployable winner\. This interval is a conservative certificate, not a predictor; directSe2eS\_\{\\mathrm\{e2e\}\}reporting remains necessary \(Table[10](https://arxiv.org/html/2609.13785#A15.T10), Appendix[O](https://arxiv.org/html/2609.13785#A15)\)\.
### 5\.5RQ4: Robustness to within\-family selector class
Across within\-family selector classes,G\(R\)G\(R\)varies1\.28×1\.28\\timesto2\.23×2\.23\\timesby benchmark, while family\-selector class changesSpartS\_\{\\mathrm\{part\}\}andGGnegligibly\. In the main B4–B5 sweep, the largest PROTEUS\-2014 gap isG=0\.211G=0\.211, and even the best class leaves positive B4–B5 gaps \(Appendix[N](https://arxiv.org/html/2609.13785#A14)\); reports must specify the within\-family selector\.
##### Inner\-selector stress test\.
Sweeping RF, XGBoost, LightGBM, MLP, and tuned GBDT within\-family regressors does not collapseG\(R\)G\(R\): the best realistic selector preserves8787–100%100\\%of the default on four of five benchmarks \(RF on B1 is the lone sub\-80%80\\%outlier\)\. A targeted rerun on B2/B4 with RF and LightGBM within\-family selectors keepsG\(R\)G\(R\)positive in every cell \(0\.0460\.046–0\.0490\.049on B2 and0\.0240\.024–0\.0270\.027on B4\)\. The strict sign of the small B2/B4 end\-to\-end margins is selector\-sensitive, but no stronger selector produces a reliable positive decomp\-vs\-flat advantage\. B1 remains a high\-absorption case, and a B5 rerun with RF keeps PROTEUS\-2014 atρ≈0\.41\\rho\\approx 0\.41\(Appendix[S](https://arxiv.org/html/2609.13785#A19)\)\. The gap is not a weak\-inner\-selector artifact, although exhaustive HPO or task\-specific meta\-learners could still reduce it and must be evaluated end\-to\-end\.
## 6Correction and Practical Guidance
##### Gap\-corrected reporting\.
EstimateGGon held\-out folds and correct partition advantage by the gap difference,A^corr=A^parttest−\[G^val\(R1\)−G^val\(R2\)\]\\widehat\{A\}\_\{\\mathrm\{corr\}\}=\\widehat\{A\}\_\{\\mathrm\{part\}\}^\{\\mathrm\{test\}\}\-\[\\widehat\{G\}^\{\\mathrm\{val\}\}\(R\_\{1\}\)\-\\widehat\{G\}^\{\\mathrm\{val\}\}\(R\_\{2\}\)\]\. Using validation folds drawn only from the outer training split,A^corr\\widehat\{A\}\_\{\\mathrm\{corr\}\}reduces MAE by2525–88%88\\%on B2–B5 decomp\-vs\-flat cells and recovers the point\-estimate deployable sign on all four sign\-changing cells \(B2/B4\)\. On B1, where the true end\-to\-end advantage is small but positive, correction can over\-subtract and add variance\. Split\-sensitivity reruns on B2/B4 with33,55, and1010training\-side validation folds, and two additional validation split seeds at the55\-fold setting, recover the same four point\-estimate signs; MAE reduction is weaker on the most granular validation split, especially for B4 \(Appendix[I](https://arxiv.org/html/2609.13785#A9)\)\. The diagnostic also assumes the training\-side validation split is representative of the outer test distribution; under covariate or task\-distribution shift it may under\- or over\-correct\. It does not replace direct end\-to\-end evaluation\.
##### Deployment\-aware family selector\.
Training the outer selector on cross\-fitted deployable utility instead ofmaxa∈Fku\\max\_\{a\\in F\_\{k\}\}ushrinksG\(R\)G\(R\)on every benchmark and improvesSe2eS\_\{\\mathrm\{e2e\}\}on B1–B4; on B5 it shrinksGGby0\.0080\.008with a smallSe2eS\_\{\\mathrm\{e2e\}\}trade\-off \(−0\.005\-0\.005, Appendix[K](https://arxiv.org/html/2609.13785#A11)\)\.
##### Reporting checklist\.
A decomposed\-AS report should pairSpart,Se2e,G\(R\)S\_\{\\mathrm\{part\}\},S\_\{\\mathrm\{e2e\}\},G\(R\)withApart,Ae2e,ρA\_\{\\mathrm\{part\}\},A\_\{\\mathrm\{e2e\}\},\\rhoand the identification interval, and document the within\-family selector, family mapping, feature policy, and common valid\-instance set \(Appendix[E](https://arxiv.org/html/2609.13785#A5)\)\.
## 7Limitations and Conclusion
G\(R\)G\(R\)is mapping\-dependent \(Appendix[J](https://arxiv.org/html/2609.13785#A10)\) and pipeline\-specific; B5’s headline effect survives the feature\-policy and cutoff sensitivities of §[5\.3\.1](https://arxiv.org/html/2609.13785#S5.SS3.SSS1)\. Continuous black\-box optimisation\([Petelin and Cenikj, 2025](https://arxiv.org/html/2609.13785#bib.bib6)\)is out of scope\. Stronger within\-family models can move small end\-to\-end margins \(Appendix[S](https://arxiv.org/html/2609.13785#A19)\), and exhaustive HPO or domain\-specific meta\-learners may shrinkG\(R\)G\(R\)further; such improvements must still be reported end\-to\-end\. The validation gap\-correction diagnostic assumes training\-side validation is representative of the test distribution and is not evaluated under explicit covariate shift\. Partition scores upper\-bound family quality but are not system scores: the safety condition fails on4/104/10cells and every identification interval strictly crosses zero, soSpart,Se2e,G\(R\)S\_\{\\mathrm\{part\}\},S\_\{\\mathrm\{e2e\}\},G\(R\), and the interval should be reported together\.Broader impact\.This evaluation\-methodology study \(no human\-subject data, no new deployed system\) reduces misleading AS/AutoML reporting that can overstate deployable performance; its main negative impact is over\-certification — our diagnostics flag when end\-to\-end evaluation is needed, not deployment\-safety guarantees, and should accompanySe2eS\_\{\\mathrm\{e2e\}\}and task\-specific risk assessment\.
## References
- Allweinet al\.\(2000\)E\. L\. Allwein, R\. E\. Schapire, and Y\. SingerReducing multiclass to binary: a unifying approach for margin classifiers\.Journal of Machine Learning Research1,pp\.113–141\.Cited by:[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px3.p1.1)\.
- Bachet al\.\(2022\)J\. Bach, M\. Iser, and K\. BöhmA comprehensive study ofkk\-portfolios of recent SAT solvers\.In25th International Conference on Theory and Applications of Satisfiability Testing \(SAT 2022\),LIPIcs\.Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.10.1.1.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Bischlet al\.\(2016\)B\. Bischl, P\. Kerschke, L\. Kotthoff, M\. Lindauer, Y\. Malitsky, A\. Fréchette, H\. H\. Hoos, F\. Hutter, K\. Leyton\-Brown, K\. Tierney, and J\. VanschorenASlib: a benchmark library for algorithm selection\.Artificial Intelligence237,pp\.41–58\.Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.SS0.SSS0.Px1.p1.1),[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.6.1.1.1),[§1](https://arxiv.org/html/2609.13785#S1.p1.1),[§1](https://arxiv.org/html/2609.13785#S1.p3.1),[§1](https://arxiv.org/html/2609.13785#S1.p4.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Cameronet al\.\(2016\)C\. Cameron, H\. H\. Hoos, and K\. Leyton\-BrownBias in algorithm portfolio performance evaluation\.InProceedings of the 25th International Joint Conference on Artificial Intelligence \(IJCAI\),pp\.712–719\.Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.SS0.SSS0.Px1.p1.1),[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.5.1.2),[§1](https://arxiv.org/html/2609.13785#S1.p4.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Ericksonet al\.\(2025\)N\. Erickson, L\. Purucker, A\. Tschalzev, D\. Holzmüller, P\. M\. Desai, D\. Salinas, and F\. HutterTabArena: a living benchmark for machine learning on tabular data\.InAdvances in Neural Information Processing Systems \(NeurIPS\) – Datasets and Benchmarks Track,Note:SpotlightExternal Links:2506\.16791Cited by:[§5\.1](https://arxiv.org/html/2609.13785#S5.SS1.p1.1)\.
- Galaret al\.\(2011\)M\. Galar, A\. Fernández, E\. Barrenechea, H\. Bustince, and F\. HerreraAn overview of ensemble methods for binary classifiers in multi\-class problems: experimental study on one\-vs\-one and one\-vs\-all schemes\.Pattern Recognition44\(8\),pp\.1761–1776\.Cited by:[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px3.p1.1)\.
- Hurleyet al\.\(2014\)B\. Hurley, L\. Kotthoff, Y\. Malitsky, and B\. O’SullivanProteus: a hierarchical portfolio of solvers and transformations\.InIntegration of AI and OR Techniques in Constraint Programming – 11th International Conference \(CPAIOR\),pp\.301–317\.Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.SS0.SSS0.Px1.p1.1),[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.9.1.1.1),[Appendix B](https://arxiv.org/html/2609.13785#A2.SS0.SSS0.Px5),[§1](https://arxiv.org/html/2609.13785#S1.p3.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2609.13785#S5.SS1.p1.1)\.
- Kerschkeet al\.\(2019\)P\. Kerschke, H\. H\. Hoos, F\. Neumann, and H\. TrautmannAutomated algorithm selection: survey and perspectives\.Evolutionary Computation27\(1\),pp\.3–45\.Cited by:[§1](https://arxiv.org/html/2609.13785#S1.p1.1)\.
- Kostovskaet al\.\(2023\)A\. Kostovska, A\. Janković, D\. Vermetten, S\. Džeroski, T\. Eftimov, and C\. DoerrPS\-AAS: portfolio selection for automated algorithm selection in black\-box optimization\.InAutoML Conference 2023 – Main Track,External Links:[Link](https://openreview.net/forum?id=j5IKxqbpE2w)Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.11.1.1.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Li and Manyà \(2009\)C\. M\. Li and F\. ManyàMaxSAT, hard and soft constraints\.InHandbook of Satisfiability,A\. Biere, M\. Heule, H\. van Maaren, and T\. Walsh \(Eds\.\),pp\.613–631\.Cited by:[Appendix B](https://arxiv.org/html/2609.13785#A2.SS0.SSS0.Px4),[§5\.1](https://arxiv.org/html/2609.13785#S5.SS1.p1.1)\.
- Lindaueret al\.\(2015\)M\. Lindauer, H\. H\. Hoos, F\. Hutter, and T\. SchaubAutoFolio: an automatically configured algorithm selector\.Journal of Artificial Intelligence Research53,pp\.745–778\.Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.SS0.SSS0.Px1.p1.1),[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.7.1.1.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Lindaueret al\.\(2019\)M\. Lindauer, J\. N\. van Rijn, and L\. KotthoffThe algorithm selection competitions 2015 and 2017\.Artificial Intelligence272,pp\.86–100\.External Links:[Document](https://dx.doi.org/10.1016/j.artint.2018.10.004)Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.8.1.1.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- McElfreshet al\.\(2023\)D\. McElfresh, S\. Khandagale, J\. Valverde, V\. P\. C\., B\. Feuer, C\. Hegde, G\. Ramakrishnan, M\. Goldblum, and C\. WhiteWhen do neural nets outperform boosted trees on tabular data?\.InAdvances in Neural Information Processing Systems \(NeurIPS\) – Datasets and Benchmarks Track,Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.15.1.1.1),[Appendix B](https://arxiv.org/html/2609.13785#A2.SS0.SSS0.Px1),[§1](https://arxiv.org/html/2609.13785#S1.p3.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Petelin and Cenikj \(2025\)G\. Petelin and G\. CenikjThe pitfalls of benchmarking in algorithm selection: what we are getting wrong\.External Links:2505\.07750,[Link](https://arxiv.org/abs/2505.07750)Cited by:[§1](https://arxiv.org/html/2609.13785#S1.p4.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px2.p1.1),[§7](https://arxiv.org/html/2609.13785#S7.p1.1)\.
- Salinas and Erickson \(2024\)D\. Salinas and N\. EricksonTabRepo: a large scale repository of tabular model evaluations and its AutoML applications\.External Links:2311\.02971,[Link](https://arxiv.org/abs/2311.02971)Cited by:[Appendix B](https://arxiv.org/html/2609.13785#A2.SS0.SSS0.Px2),[§1](https://arxiv.org/html/2609.13785#S1.p3.1)\.
- Shmuelet al\.\(2025\)A\. Shmuel, O\. Glickman, and T\. LazebnikA comprehensive benchmark of machine and deep learning models on structured data for regression and classification\.Neurocomputing\.Note:Article number 131337External Links:[Document](https://dx.doi.org/10.1016/j.neucom.2025.131337)Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.SS0.SSS0.Px1.p1.1),[Appendix R](https://arxiv.org/html/2609.13785#A18.p2.1),[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.16.1.1.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Stojadinović and Marić \(2014\)M\. Stojadinović and F\. MarićmeSAT: multiple encodings of CSP to SAT\.Constraints19\(4\),pp\.380–403\.External Links:[Document](https://dx.doi.org/10.1007/s10601-014-9165-7)Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.12.1.1.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Sun and Pfahringer \(2013\)Q\. Sun and B\. PfahringerPairwise meta\-rules for better meta\-learning\-based algorithm ranking\.Machine Learning93\(1\),pp\.141–161\.Cited by:[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px3.p1.1)\.
- Tornedeet al\.\(2023\)A\. Tornede, L\. Gehring, T\. Tornede, M\. Wever, and E\. HüllermeierAlgorithm selection on a meta level\.Machine Learning112\(4\),pp\.1253–1286\.Cited by:[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p2.1)\.
- Ulrich\-Olteanet al\.\(2022\)F\. Ulrich\-Oltean, P\. Nightingale, and J\. A\. WalkerSelecting SAT encodings for pseudo\-Boolean and linear integer constraints\.In28th International Conference on Principles and Practice of Constraint Programming \(CP 2022\),LIPIcs\.Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.13.1.1.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Ulrich\-Olteanet al\.\(2023\)F\. Ulrich\-Oltean, P\. Nightingale, and J\. A\. WalkerLearning to select SAT encodings for pseudo\-Boolean and linear integer constraints\.Constraints28,pp\.397–426\.External Links:[Document](https://dx.doi.org/10.1007/s10601-023-09364-1)Cited by:[Appendix R](https://arxiv.org/html/2609.13785#A18.SS0.SSS0.Px1.p1.1),[Appendix R](https://arxiv.org/html/2609.13785#A18.p3.2.14.1.1.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Xuet al\.\(2008\)L\. Xu, F\. Hutter, H\. H\. Hoos, and K\. Leyton\-BrownSATzilla: portfolio\-based algorithm selection for SAT\.Journal of Artificial Intelligence Research32,pp\.565–606\.Cited by:[§1](https://arxiv.org/html/2609.13785#S1.p4.1),[§2](https://arxiv.org/html/2609.13785#S2.SS0.SSS0.Px1.p1.1)\.
- Yeet al\.\(2024\)H\. Ye, S\. Liu, H\. Cai, Q\. Zhou, and D\. ZhanA closer look at deep learning methods on tabular datasets\.External Links:2407\.00956,[Link](https://arxiv.org/abs/2407.00956)Cited by:[Appendix B](https://arxiv.org/html/2609.13785#A2.SS0.SSS0.Px3),[§1](https://arxiv.org/html/2609.13785#S1.p3.1)\.
## Appendix
The appendix contains: \(A\) the per\-instance utility transform; \(B\) full algorithm\-family mappings for all five benchmarks; \(C\) the synthetic toy example used in §[4](https://arxiv.org/html/2609.13785#S4); \(D\) the reversal taxonomy table; \(E\) the full reporting checklist and decision framework; \(F\) full proofs of Lemmas 1–2 and Theorems 1–2; \(G\) common\-instance evaluation and identity\-residual analysis; \(H\) pseudocode for the decomposed pipeline evaluator; \(I\) gap\-corrected reporting; \(J\) family\-mapping sensitivity \(B4 only; B5 reports the primary semantic partition\); \(K\) deployment\-aware family selector; \(L\) PROTEUS\-2014 feature\-policy and cutoff sensitivities; \(M\) metric\-transform sensitivity on B4–B5; \(N\) the within\-family selector robustness sweep with both the per\-benchmark figure and the full2424\-row B1–B3 table; \(O\) full numerical tables for the gap and advantage\-transfer analyses; \(P\) statistical methodology; \(Q\) compute resources and reproducibility instructions; \(R\) a literature audit of oracle\-style reporting in algorithm selection, solver portfolios, encoding selection, and tabular benchmarking; \(S\) full per\-cell results of the inner\-selector stress test of §[5\.5](https://arxiv.org/html/2609.13785#S5.SS5)\.
## Appendix AUtility transformation
For B1–B3 the native metric is an accuracy\-like value already in\[0,1\]\[0,1\]and higher\-is\-better; we use it directly as utility\. For B4 \(PAR10\) and B5 \(runtime; PAR10 reconstructed fromrunstatus\), we apply a per\-instance min\-max transformation
u\(x,a\)=1−clip\(p\(x,a\)−pmin\(x\)pmax\(x\)−pmin\(x\),0,1\),u\(x,a\)\\;=\\;1\\;\-\\;\\mathrm\{clip\}\\\!\\left\(\\frac\{p\(x,a\)\-p\_\{\\min\}\(x\)\}\{p\_\{\\max\}\(x\)\-p\_\{\\min\}\(x\)\},\\;0,\\;1\\right\),\(12\)wherep\(x,a\)p\(x,a\)is the native lower\-is\-better measurement \(PAR10 in seconds\)\. Instances on which all algorithms tie at the timeout penalty are dropped \(these carry no AS signal\)\. More generally, the transform is applied only whenpmax\(x\)\>pmin\(x\)p\_\{\\max\}\(x\)\>p\_\{\\min\}\(x\); ifpmax\(x\)=pmin\(x\)p\_\{\\max\}\(x\)=p\_\{\\min\}\(x\), the row has zero within\-instance range and is mapped to all\-NaN before the common valid\-instance filter\. After dropping degenerate rows, B4 retains556556of601601instances and B5 retains35653565of the40214021raw PROTEUS\-2014 instances\. Per\-instance min\-max preserves rankings and is the standard convention adopted by ASlib analyses; under this convention any instance with at least one strictly\-better\-than\-worst algorithm contributes a non\-degenerate utility range\.
For B5, whererunstatus≠\\neqok\(timeout, memout, crash\) or whereokruntime exceeds the cutoff, we apply the standard PAR10 penaltyp\(x,a\):=10⋅pcutoffp\(x,a\):=10\\cdot p\_\{\\mathrm\{cutoff\}\}before the per\-instance transform\. The cutoffpcutoff=3600p\_\{\\mathrm\{cutoff\}\}=3600s is taken from the scenario’sdescription\.txt\(algorithm\_cutoff\_time: 3600\); the18001800s cutoff sensitivity is reported in Appendix[L](https://arxiv.org/html/2609.13785#A12)and is not used for headline claims\.
## Appendix BAlgorithm\-family mappings
##### B1 TabZilla\[[McElfresh et al\., 2023](https://arxiv.org/html/2609.13785#bib.bib10)\]\.
TreeEnsemble: CatBoost, LightGBM, XGBoost, RandomForest\.Classic: LinearModel, SVM, KNN, DecisionTree\.Pretrained: TabPFN\.Deep: MLP, rtdl\_MLP, rtdl\_ResNet, rtdl\_FTTransformer, SAINT, TabTransformer, NAM, NODE, STG, TabNet, VIME, DANet, DeepFM\.
##### B2 TabRepo\[[Salinas and Erickson, 2024](https://arxiv.org/html/2609.13785#bib.bib11)\]\.
TreeEnsemble: CAT, GBM, XGB, RF, XT\.Classic: KNN, LR\.Deep: FASTAI, FT\_TRANSFORMER, NN\_TORCH\.Pretrained: TABPFN\.
##### B3 TALENT\[[Ye et al\., 2024](https://arxiv.org/html/2609.13785#bib.bib12)\]\.
TreeEnsemble: catboost, lightgbm, RandomForest, xgboost\.Classic: knn, LogReg, NCM, NaiveBayes, svm\.Pretrained: tabpfn, tabpfn\_v2, tabpfn\_v2\_5, tabicl\_v2\.Deep: autoint, danets, dcn2, excelformer, ftt, grownet, mlp, mlp\_plr, modernNCA, node, ptarl, ptaul, realmlp, resnet, snn, switchtab, tabcaps, tabnet, tabr, tabtransformer, tangos\. The dataset’s constant\-reference baseline column is family\-skipped \(assigned to no family; unreachable through the decomposed pipeline\)\.
##### B4 ASlib MAXSAT\-PMS\-2016\[[Li and Manyà, 2009](https://arxiv.org/html/2609.13785#bib.bib13)\]\.
BnB\(branch\-and\-bound\): WMaxSatz09, WMaxSatz\+, ahms\-1\.70\.CoreGuided\(SAT\-based core extraction; OLL\-style\): Open\-WBO15, Open\-WBO16, mscg2015a, mscg2015b, maxino16\-c10, maxino16\-dis, Naps\-1\.02\-ms, Optiriss6, QMaxSAT14, QMaxSAT16UC, WPM3\-2015\-co\.IHS\(implicit hitting set; hybrid SAT\+\+MIP\): maxhs\-b, LMHS\-2016\.SLS\(stochastic local search; incl\. SLS\+\+BnB hybrids\): CCEHC2akms, CCLS2akms, ahms\-ls\-1\.70\.
##### B5 ASlib PROTEUS\-2014\[[Hurley et al\., 2014](https://arxiv.org/html/2609.13785#bib.bib14)\]\.
The PROTEUS portfolio decides between solving a CSP instance natively or transforming it via one of three CSP\-to\-SAT encodings \(direct, support, direct\-order\) and dispatching to a SAT solver\. The four semantic families on the2222algorithms are:CSP\-native: abscon, choco, gecode, mistral\_nj\.SAT\-direct: claspcnf\_direct, cryptominisat\_direct, glucose\_direct, lingeling\_direct, minisat22\_direct, riss3g\_direct\.SAT\-support: claspcnf\_support, cryptominisat\_support, glucose\_support, lingeling\_support, minisat22\_support, riss3g\_support\.SAT\-direct\-order: claspcnf\_directorder, cryptominisat\_directorder, glucose\_directorder, lingeling\_directorder, minisat22\_directorder, riss3g\_directorder\.
The mapping is deterministic from the algorithm string suffix and matches the original Proteus paper’s hierarchy: CSP\-native solvers appear as one family, and the SAT\-side3×63\\times 6grid of encodings×\\timesSAT solvers becomes three families of six\.
## Appendix CSynthetic toy example
The toy used in §[4](https://arxiv.org/html/2609.13785#S4)consists ofn=100n=100synthetic instances split evenly into two halves indexed byh∈\{0,1\}h\\in\\\{0,1\\\}\. Algorithm utilities are
u\(x∣h=0\)\\displaystyle u\(x\\mid h\\\!=\\\!0\)=\[1\.0,0\.0,0\.5,0\.4\]\+ε,\\displaystyle=\[1\.0,\\;0\.0,\\;0\.5,\\;0\.4\]\+\\varepsilon,u\(x∣h=1\)\\displaystyle u\(x\\mid h\\\!=\\\!1\)=\[0\.0,1\.0,0\.4,0\.5\]\+ε,\\displaystyle=\[0\.0,\\;1\.0,\\;0\.4,\\;0\.5\]\+\\varepsilon,withε∼𝒩\(0,0\.012\)4\\varepsilon\\sim\\mathcal\{N\}\(0,0\.01^\{2\}\)^\{4\}i\.i\.d\. across instances and algorithms\. PipelineR1R\_\{1\}uses the partitionΠ1=\{\{a0,a1\},\{a2,a3\}\}\\Pi\_\{1\}=\\\{\\\{a\_\{0\},a\_\{1\}\\\},\\\{a\_\{2\},a\_\{3\}\\\}\\\}with an oracle family selector \(k^R1\(x\)\\hat\{k\}\_\{R\_\{1\}\}\(x\)always equals the family containing the optimal algorithm\) but a within\-family selector that flips with probabilityτ1=0\.5\\tau\_\{1\}=0\.5between the two members of the chosen family\. PipelineR2R\_\{2\}usesΠ2=\{\{a0,a2\},\{a1,a3\}\}\\Pi\_\{2\}=\\\{\\\{a\_\{0\},a\_\{2\}\\\},\\\{a\_\{1\},a\_\{3\}\\\}\\\}with family\-selector error rate0\.40\.4and a nearly clean within\-family selector \(flip rateτ2=0\.05\\tau\_\{2\}=0\.05\)\.
The toy construction yieldsSpart\(R1\)≈0\.996S\_\{\\mathrm\{part\}\}\(R\_\{1\}\)\\approx 0\.996,Spart\(R2\)≈0\.792S\_\{\\mathrm\{part\}\}\(R\_\{2\}\)\\approx 0\.792,Se2e\(R1\)≈0\.520S\_\{\\mathrm\{e2e\}\}\(R\_\{1\}\)\\approx 0\.520,Se2e\(R2\)≈0\.765S\_\{\\mathrm\{e2e\}\}\(R\_\{2\}\)\\approx 0\.765, henceApart≈\+0\.20A\_\{\\mathrm\{part\}\}\\approx\+0\.20,Ae2e≈−0\.24A\_\{\\mathrm\{e2e\}\}\\approx\-0\.24, andρ≈2\.20\\rho\\approx 2\.20\. The identity residual\|Apart−Ae2e−\(G\(R1\)−G\(R2\)\)\|\|A\_\{\\mathrm\{part\}\}\-A\_\{\\mathrm\{e2e\}\}\-\(G\(R\_\{1\}\)\-G\(R\_\{2\}\)\)\|is below10−1510^\{\-15\}for any seed, illustrating Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2)numerically\. This construction is also the toy invoked at the end of §[4\.2](https://arxiv.org/html/2609.13785#S4.SS2):R1R\_\{1\}has the larger oracle margin but a strictly larger within\-family regret on the conflicted instances\. This is the mechanism behind the failure, but the formal cross\-pipeline sign change is the deployment\-safety failure in Corollary[2](https://arxiv.org/html/2609.13785#Thmcorollary2):Apart<G\(R1\)−G\(R2\)A\_\{\\mathrm\{part\}\}<G\(R\_\{1\}\)\-G\(R\_\{2\}\), henceAe2e<0A\_\{\\mathrm\{e2e\}\}<0\.
## Appendix DReversal taxonomy
The full qualitative taxonomy of how a partition\-level conclusion transfers end\-to\-end, derived from \([6](https://arxiv.org/html/2609.13785#S4.E6)\):
Reversal and opposite reversal are the two sign\-reversal regimes; our empirical study observes standard\-reversal point estimates on two benchmarks but does not observe opposite reversals onceSpartS\_\{\\mathrm\{part\}\}/Se2eS\_\{\\mathrm\{e2e\}\}/GGare reported on a common instance set\. Amplification is mathematically possible when the partition\-level winner has a smaller deployment\-fidelity gap than its comparator; it is not needed for the headline empirical claims\.
## Appendix EReporting checklist and decision framework
##### Decision framework\.
Table[3](https://arxiv.org/html/2609.13785#A5.T3)maps research questions to recommended metrics\. Partition score is appropriate when the question is about the family partition itself; it is insufficient when the question is about a deployable system or about comparing two decomposed methods\.
Table 3:Metric\-reporting decision framework grounded in our findings\.
##### Full reporting checklist\.
A decomposed\-AS report should include: \(i\)SpartS\_\{\\mathrm\{part\}\}andSe2eS\_\{\\mathrm\{e2e\}\}side by side; \(ii\)G\(R\)G\(R\)with bootstrap CI; \(iii\) for comparisons,Apart,Ae2e,ρA\_\{\\mathrm\{part\}\},A\_\{\\mathrm\{e2e\}\},\\rho, withρ\\rhomarked unstable when\|Apart\|<5×10−3\|A\_\{\\mathrm\{part\}\}\|<5\\times 10^\{\-3\}; \(iv\) the within\-family selector class and any oracle used; \(v\) the family mapping*and the partition’s semantic basis*— paradigm, encoding, lineage, random, or cluster — so that readers can judge whether the partition matches the deployment hierarchy; \(vi\) the*feature availability policy*— whether the family selector uses features that are only computable after a representation or encoding choice has been made; \(vii\) for runtime scenarios, the timeout, PAR10 convention, and missing\-value handling; \(viii\) the common valid\-instance set used per comparison and the empirical residuals of Lemmas[1](https://arxiv.org/html/2609.13785#Thmlemma1)–[2](https://arxiv.org/html/2609.13785#Thmlemma2)as a consistency check; \(ix\) for every method comparison, the*partition\-only identification interval*\[Apart−WR1,Apart\+WR2\]\[A\_\{\\mathrm\{part\}\}\-W\_\{R\_\{1\}\},A\_\{\\mathrm\{part\}\}\+W\_\{R\_\{2\}\}\]from Theorem[2](https://arxiv.org/html/2609.13785#Thmtheorem2)and whether it strictly crosses zero — the partition\-level report does not certify a deployable winner unless the interval is strictly signed\.
## Appendix FProofs
##### Lemma[1](https://arxiv.org/html/2609.13785#Thmlemma1)\(gap identity\)\.
Substituting \([3](https://arxiv.org/html/2609.13785#S3.E3)\) and \([4](https://arxiv.org/html/2609.13785#S3.E4)\) into the definition ofG\(R\)G\(R\)and using linearity of expectation gives the first equality\. For eachxx,ak^R\(x\)∗\(x\)a\_\{\\hat\{k\}\_\{R\}\(x\)\}^\{\*\}\(x\)is by definition the maximiser ofu\(x,⋅\)u\(x,\\cdot\)overFk^R\(x\)F\_\{\\hat\{k\}\_\{R\}\(x\)\}, andgR,k^R\(x\)\(x\)∈Fk^R\(x\)g\_\{R,\\hat\{k\}\_\{R\}\(x\)\}\(x\)\\in F\_\{\\hat\{k\}\_\{R\}\(x\)\}, so the pointwise difference is nonnegative\. Equality in expectation forces equality almost surely\. ∎
##### Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2)\(advantage absorption\)\.
By definitionSe2e\(R\)=Spart\(R\)−G\(R\)S\_\{\\mathrm\{e2e\}\}\(R\)=S\_\{\\mathrm\{part\}\}\(R\)\-G\(R\)\. Substituting intoAe2e=Se2e\(R1\)−Se2e\(R2\)A\_\{\\mathrm\{e2e\}\}=S\_\{\\mathrm\{e2e\}\}\(R\_\{1\}\)\-S\_\{\\mathrm\{e2e\}\}\(R\_\{2\}\)and rearranging yields \([6](https://arxiv.org/html/2609.13785#S4.E6)\)\. ∎
##### Corollary[2](https://arxiv.org/html/2609.13785#Thmcorollary2)\(general deployment\-safety\)\.
By Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2),Ae2e=Apart−\[G\(R\)−G\(F\)\]=Apart−ΔGA\_\{\\mathrm\{e2e\}\}=A\_\{\\mathrm\{part\}\}\-\[G\(R\)\-G\(F\)\]=A\_\{\\mathrm\{part\}\}\-\\Delta G, soAe2e\>0⇔Apart\>ΔGA\_\{\\mathrm\{e2e\}\}\>0\\iff A\_\{\\mathrm\{part\}\}\>\\Delta G\. SettingG\(F\)=0G\(F\)=0recovers the canonical form\. ∎
##### Theorem[1](https://arxiv.org/html/2609.13785#Thmtheorem1)\(margin–regret stability\)\.
vq\(x\)=mq\(x\)−rq\(x\)v\_\{q\}\(x\)=m\_\{q\}\(x\)\-r\_\{q\}\(x\)for every familyqq, sovq\(x\)−vs\(x\)=\(mq−ms\)−\(rq−rs\)v\_\{q\}\(x\)\-v\_\{s\}\(x\)=\(m\_\{q\}\-m\_\{s\}\)\-\(r\_\{q\}\-r\_\{s\}\)\. Maximising overqqgives the regret expression; non\-negativity of regret yields the equivalence in \([9](https://arxiv.org/html/2609.13785#S4.E9)\)\. ∎
##### Theorem[2](https://arxiv.org/html/2609.13785#Thmtheorem2)\(partition\-only identification\)\.
By Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2)applied to any\(g1,g2\)\(g\_\{1\},g\_\{2\}\),Ae2e\(g1,g2\)=Apart−\(Gg1\(R1\)−Gg2\(R2\)\)A\_\{\\mathrm\{e2e\}\}\(g\_\{1\},g\_\{2\}\)=A\_\{\\mathrm\{part\}\}\-\(G\_\{g\_\{1\}\}\(R\_\{1\}\)\-G\_\{g\_\{2\}\}\(R\_\{2\}\)\)\. Independently varyingg1,g2g\_\{1\},g\_\{2\}over𝒢Ri\\mathcal\{G\}\_\{R\_\{i\}\}boundsGgi\(Ri\)G\_\{g\_\{i\}\}\(R\_\{i\}\)by\[G¯\(Ri\),G¯\(Ri\)\]\[\\underline\{G\}\(R\_\{i\}\),\\overline\{G\}\(R\_\{i\}\)\], soAe2eA\_\{\\mathrm\{e2e\}\}is bounded by the interval in \([11](https://arxiv.org/html/2609.13785#S4.E11)\); attainability of each envelope yields sharpness, otherwise the bound is approached but not attained\. Under the unrestricted family\-respecting class,G¯\(R\)=0\\underline\{G\}\(R\)=0is attained by the within\-family oracle selectorg\(x\)∈argmaxa∈Fk^R\(x\)u\(x,a\)g\(x\)\\in\\arg\\max\_\{a\\in F\_\{\\hat\{k\}\_\{R\}\(x\)\}\}u\(x,a\)andG¯\(R\)=WR\\overline\{G\}\(R\)=W\_\{R\}is attained by the pointwise within\-family worst\-case selectorg\(x\)∈argmina∈Fk^R\(x\)u\(x,a\)g\(x\)\\in\\arg\\min\_\{a\\in F\_\{\\hat\{k\}\_\{R\}\(x\)\}\}u\(x,a\); both are measurable on a finite algorithm pool\. ∎
## Appendix GCommon\-instance evaluation and identity residuals
The identities in Lemmas[1](https://arxiv.org/html/2609.13785#Thmlemma1)–[2](https://arxiv.org/html/2609.13785#Thmlemma2)hold pointwise\. For each \(benchmark, pair\) we therefore computeSpart\(R1\)S\_\{\\mathrm\{part\}\}\(R\_\{1\}\),Se2e\(R1\)S\_\{\\mathrm\{e2e\}\}\(R\_\{1\}\),Spart\(R2\)S\_\{\\mathrm\{part\}\}\(R\_\{2\}\),Se2e\(R2\)S\_\{\\mathrm\{e2e\}\}\(R\_\{2\}\)on the*four\-way intersection*of finite values, i\.e\. the test instances on which both pipelines have a finite partition oracle and a finite end\-to\-end utility; the resulting empirical residuals\|Apart−Ae2e−\(G\(R1\)−G\(R2\)\)\|\|A\_\{\\mathrm\{part\}\}\-A\_\{\\mathrm\{e2e\}\}\-\(G\(R\_\{1\}\)\-G\(R\_\{2\}\)\)\|are at machine precision \(≤5×10−16\\leq 5\\times 10^\{\-16\}\) on all five benchmarks\.
Figure[2](https://arxiv.org/html/2609.13785#A7.F2)visualises the two identities as a sanity check\. The left panel plots within\-family regret on the predicted family againstG\(R\)G\(R\): by Lemma[1](https://arxiv.org/html/2609.13785#Thmlemma1)the two are equal in expectation, so all benchmarks lie ony=xy=x\. The right panel plotsApart−Ae2eA\_\{\\mathrm\{part\}\}\-A\_\{\\mathrm\{e2e\}\}versusG\(R1\)−G\(R2\)G\(R\_\{1\}\)\-G\(R\_\{2\}\)for the pairwise\-vs\-multiclass comparison on every benchmark; by Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2)all points lie on the diagonal \(residuals≤5×10−16\\leq 5\\times 10^\{\-16\}\)\. These plots are not statistical fits but empirical realisations of the lemmas on per\-benchmark estimates\.
Figure 2:*Left:*per\-benchmark within\-family regret on the predicted family againstG\(R\)G\(R\)\(Lemma[1](https://arxiv.org/html/2609.13785#Thmlemma1)\)\.*Right:*Apart−Ae2eA\_\{\\mathrm\{part\}\}\-A\_\{\\mathrm\{e2e\}\}versusG\(R1\)−G\(R2\)G\(R\_\{1\}\)\-G\(R\_\{2\}\)for the pairwise\-vs\-multiclass comparison on every benchmark \(Lemma[2](https://arxiv.org/html/2609.13785#Thmlemma2); residuals≤5×10−16\\leq 5\\times 10^\{\-16\}\)\.This common\-instance protocol is essential: without it, marginal means would average over different row sets whenever a pipeline’s partition oracle is undefined on some test instances while the deployable selector falls back to a penalty value, and apparent apparent sign changes can be induced by aggregation alone\. For TALENT \(B3\), using independent valid masks can create marginal artefacts\. We therefore compute each comparison on a common finite utility mask and treat common\-instance evaluation as part of the reporting checklist \(§[6](https://arxiv.org/html/2609.13785#S6)\)\.
Flat baselines are singleton\-family controls\. For these controls the evaluator enforces the row\-wise invariantupart\(x\)=ue2e\(x\)u\_\{\\mathrm\{part\}\}\(x\)=u\_\{\\mathrm\{e2e\}\}\(x\)after missing\-value fallback: the same selected algorithm and the same per\-instance penalty are used for both scores\. Consequently all flat rows haveG\(F\)=0G\(F\)=0, including TALENT \(B3\); any nonzero singleton\-flat gap would indicate asymmetric NaN handling rather than a within\-family oracle effect\.
## Appendix HDecomposed pipeline pseudocode
The protocol used to computeSpartS\_\{\\mathrm\{part\}\},Se2eS\_\{\\mathrm\{e2e\}\}, andG\(R\)G\(R\)in §[5](https://arxiv.org/html/2609.13785#S5)is summarised below\. Stratification keys are the per\-instance argmax over𝒜\\mathcal\{A\}, with rare classes collapsed to a synthetic\_\_other\_\_class to avoid singleton folds\. The within\-family selector is fit per family using the family\-restricted columns ofMMon the training rows\.
> Algorithm: Decomposed pipeline evaluator \(one seed pass; the outer driver iterates this over55seeds and averages\)\. Input\.Utility matrixM∈ℝn×AM\\in\\mathbb\{R\}^\{n\\times A\}, feature matrixZ∈ℝn×dZ\\in\\mathbb\{R\}^\{n\\times d\}, family partitionΠ=\{F1,…,FK\}\\Pi=\\\{F\_\{1\},\\dots,F\_\{K\}\\\},KcvK\_\{\\mathrm\{cv\}\}\-fold splits\{\(trt,tet\)\}t=1Kcv\\\{\(\\mathrm\{tr\}\_\{t\},\\mathrm\{te\}\_\{t\}\)\\\}\_\{t=1\}^\{K\_\{\\mathrm\{cv\}\}\}\. Output\.Per\-instance utility vectorsupart,ue2e∈ℝnu\_\{\\mathrm\{part\}\},u\_\{\\mathrm\{e2e\}\}\\in\\mathbb\{R\}^\{n\}\. 1. 1\.Initialiseupart←𝟎nu\_\{\\mathrm\{part\}\}\\leftarrow\\mathbf\{0\}\_\{n\},ue2e←𝟎nu\_\{\\mathrm\{e2e\}\}\\leftarrow\\mathbf\{0\}\_\{n\}\. 2. 2\.For each foldt=1,…,Kcvt=1,\\dots,K\_\{\\mathrm\{cv\}\}: 1. \(a\)Fit family selectorhRh\_\{R\}on\(Z\[trt\],M\[trt\],Π\)\(Z\[\\mathrm\{tr\}\_\{t\}\],\\,M\[\\mathrm\{tr\}\_\{t\}\],\\,\\Pi\)\. 2. \(b\)For each familykk: fit within\-family selectorgR,kg\_\{R,k\}on\(Z\[trt\],M\[trt,Fk\]\)\(Z\[\\mathrm\{tr\}\_\{t\}\],\\,M\[\\mathrm\{tr\}\_\{t\},F\_\{k\}\]\)\. 3. \(c\)For eachi∈teti\\in\\mathrm\{te\}\_\{t\}: setk^←hR\(Z\[i\]\)\\hat\{k\}\\leftarrow h\_\{R\}\(Z\[i\]\),a^←gR,k^\(Z\[i\]\)\\hat\{a\}\\leftarrow g\_\{R,\\hat\{k\}\}\(Z\[i\]\),upart\[i\]←maxa∈Fk^M\[i,a\]u\_\{\\mathrm\{part\}\}\[i\]\\leftarrow\\max\_\{a\\in F\_\{\\hat\{k\}\}\}M\[i,a\]\(penalty fallback if no member is finite\), andue2e\[i\]←M\[i,a^\]u\_\{\\mathrm\{e2e\}\}\[i\]\\leftarrow M\[i,\\hat\{a\}\]\(the same per\-instance penalty if missing\)\. For singleton flat baselines, tie the two assignments after fallback soupart\[i\]=ue2e\[i\]u\_\{\\mathrm\{part\}\}\[i\]=u\_\{\\mathrm\{e2e\}\}\[i\]\. 3. 3\.Return\(upart,ue2e\)\(u\_\{\\mathrm\{part\}\},u\_\{\\mathrm\{e2e\}\}\)\.
The evaluation runs the algorithm above for each of55seeds, averages per\-instance utilities across seeds, and reportsSpart=mean\(upart\)S\_\{\\mathrm\{part\}\}=\\mathrm\{mean\}\(u\_\{\\mathrm\{part\}\}\),Se2e=mean\(ue2e\)S\_\{\\mathrm\{e2e\}\}=\\mathrm\{mean\}\(u\_\{\\mathrm\{e2e\}\}\), andG=mean\(upart−ue2e\)G=\\mathrm\{mean\}\(u\_\{\\mathrm\{part\}\}\-u\_\{\\mathrm\{e2e\}\}\)along with95%95\\%percentile bootstrap CIs over per\-instance differences\.
## Appendix IGap\-corrected reporting
We evaluate the gap\-correction proposed in §[6](https://arxiv.org/html/2609.13785#S6)via training\-side validation\. For each outer fold, the outer test fold is kept untouched\. We split the outer training fold into validation folds, fit the two candidate pipelines on the inner training portion, estimateG^val\(R1\)−G^val\(R2\)\\widehat\{G\}^\{\\mathrm\{val\}\}\(R\_\{1\}\)\-\\widehat\{G\}^\{\\mathrm\{val\}\}\(R\_\{2\}\)on the held\-out validation portion, and then correctA^parttest\\widehat\{A\}\_\{\\mathrm\{part\}\}^\{\\mathrm\{test\}\}on the original outer test fold\. The correction therefore uses only training\-side data for the gap estimate\.
Figure[3](https://arxiv.org/html/2609.13785#A9.F3)reports the result\. On the four current decomp\-vs\-flat sign\-changing cells \(B2/B4\), correction recovers the point\-estimate end\-to\-end sign and reduces MAE by2525–65%65\\%; the fold\-level corrected sign accuracy is0\.680\.68–0\.800\.80on B2 and0\.560\.56–0\.600\.60on B4\. A split\-sensitivity check on these four cells keeps the point\-estimate sign correct for validation\-fold counts33,55, and1010, and for two additional55\-fold validation split seeds; MAE reduction ranges from3737–65%65\\%on B2 and99–31%31\\%on B4\. On non\-sign\-changing but strongly absorbed cells, it reduces MAE by6060–65%65\\%on B3 and8787–88%88\\%on B5, while B1 is the cautionary case: the trueAe2eA\_\{\\mathrm\{e2e\}\}is only0\.0010\.001–0\.0020\.002, and correction over\-subtracts on the point estimate\. We therefore use correction as a diagnostic and still require directSe2eS\_\{\\mathrm\{e2e\}\}reporting; we do not test robustness under explicit covariate or task\-distribution shift\.
Figure 3:Gap\-corrected vs\. raw partition advantage on all1515pairwise comparisons\.*\(a\)*crosses are rawApartA\_\{\\mathrm\{part\}\}, filled circles are correctedA^corr\\widehat\{A\}\_\{\\mathrm\{corr\}\}; corrected points are closer to they=xy=xidentity than the raw partition estimates\.*\(b\)*MAE reduction \(%\) for each decomp\-vs\-flat comparison; gain is positive on B2–B5 and negative on the B1 high\-absorption cells, where the deployable advantage is close to zero and validation\-side gap estimates can over\-correct\.
## Appendix JFamily\-mapping sensitivity
For B4 we re\-evaluate the deployment\-fidelity gap and the decomp\-vs\-flat advantage transfer under three alternative family partitions in addition to the default semantic taxonomy: \(P2\) a coarse two\-family collapse, \(P3\) a performance\-cluster partition withK=4K=4obtained bykk\-means on per\-algorithm utility profiles, and \(P4\) a random balancedK=4K=4placebo averaged over five seeds\. Figure[4](https://arxiv.org/html/2609.13785#A10.F4)plotsG\(R\)G\(R\),ApartA\_\{\\mathrm\{part\}\}, andAe2eA\_\{\\mathrm\{e2e\}\}side\-by\-side per partition for B4; Table[4](https://arxiv.org/html/2609.13785#A10.T4)reports the numbers and the safety\-condition flag\. For B5 \(PROTEUS\-2014\) we report only the primary semantic partitionΠP1\\Pi\_\{\\text\{P1\}\}here; coarser K\-sensitivity for B5 is in the companion reproducibility package\. The primary result is that the positive deployment\-fidelity gap and the failure of the flat\-deployment safety condition on B4 are robust to the choice of family taxonomy: every partition we test on B4 producesG\(R\)G\(R\)exceedingApartA\_\{\\mathrm\{part\}\}\.
Figure 4:Family\-mapping sensitivity on B4\. Bars areG\(R\)G\(R\),ApartA\_\{\\mathrm\{part\}\}, andAe2eA\_\{\\mathrm\{e2e\}\}per partition kind\. B4 fails the decomp\-vs\-flat safety conditionApart\>G\(R\)A\_\{\\mathrm\{part\}\}\>G\(R\)for every partition tested, showing that the gap magnitude is not driven solely by the default taxonomy\. B5 mapping sensitivity beyond the primary semantic PROTEUS partition is not used for paper claims\.Table 4:Family\-mapping sensitivity\. Each row is one \(benchmark, partition\) cell with the resulting gap and decomp\-vs\-flat advantages\. “safe” = general safety conditionApart\>ΔGA\_\{\\mathrm\{part\}\}\>\\Delta GwhereΔG=G\(R\)−G\(F\)\\Delta G=G\(R\)\-G\(F\)\. P4 \(random\) is averaged over five seeds\. B4 fails the safety condition under every partition we test \(default, coarse, performance\-cluster, random placebo\); B5 \(PROTEUS\-2014\) reports the primary semantic partitionΠP1\\Pi\_\{\\text\{P1\}\}only — coarser K\-sensitivity for B5 is in the companion reproducibility package\. Family mapping changes the gap magnitude but does not remove it\.
## Appendix KDeployment\-aware family selector
We compare the default oracle\-trained family selector \(argmaxkmaxa∈Fku\(x,a\)\\arg\\max\_\{k\}\\max\_\{a\\in F\_\{k\}\}u\(x,a\)on the training set\) against a deployment\-aware family selector that targets the cross\-fitted deployable family utilityvk\(x\)=u\(x,g^k\(−m\)\(x\)\)v\_\{k\}\(x\)=u\(x,\\hat\{g\}\_\{k\}^\{\(\-m\)\}\(x\)\), whereg^k\(−m\)\\hat\{g\}\_\{k\}^\{\(\-m\)\}is the within\-family selector trained on the inner training fold not containingxx\. Both pipelines share thePerAlgRegressionWithinwithin\-family selector\. Per\-benchmark headline numbers appear in Table[5](https://arxiv.org/html/2609.13785#A11.T5)\.
Table 5:Deployment\-aware family selector \(DA\) vs\. oracle\-trained family selector \(O\), with both sharingPerAlgRegressionWithin\.ΔX=XDA−XO\\Delta\_\{X\}=X\_\{\\mathrm\{DA\}\}\-X\_\{\\mathrm\{O\}\}\. DA shrinksG\(R\)G\(R\)on every benchmark; it improvesSe2eS\_\{\\mathrm\{e2e\}\}on B1–B4 and trades off slightly on B5 PROTEUS\-2014, where the within\-family utility range is large enough that targeting deployable utility moves family decisions toward families with smaller oracle margin\. Sign\-changing point estimates relative to flat persist on B2 and B4 — the safety condition remains the primary diagnostic\.
## Appendix LPROTEUS\-2014 feature\-policy and cutoff sensitivity
The PROTEUS\-2014 ASlib scenario provides four feature steps with3636,5454,5454, and5454features respectively for a total of198198features\. Thecspstep contains features computable on the original CSP instance;direct,support, anddirectordercontain features that are only well\-defined after the corresponding CSP\-to\-SAT encoding has been applied\. Our primary analysis uses thecspstep alone so that the family selector makes its CSP\-or\-SAT\-and\-encoding decision on information available before any encoding is fixed\.
For the all\-feature sensitivity, we re\-run the B5 main pipeline with all198198features supplied to the family selector and the same within\-family selector class\. Headline numbers under both feature policies appear in Table[6](https://arxiv.org/html/2609.13785#A12.T6)\. The all\-feature setting is a stronger\-information setting that gives the family selector access to encoding\-conditional features it would not have at deployment time; we therefore interpret thecsp\-step configuration as the primary deployment model and the all\-feature setting as a sensitivity check\.
The PROTEUS\-2014 ASlib scenario reports algorithm runtimes censored at the scenario cutoff of36003600s and we use this cutoff for the primary analysis\. We additionally re\-run B5 with the cutoff set to18001800s: anyokrun whose runtime exceeds18001800s is re\-classified as a timeout under the new \(shorter\) cutoff and penalised atPAR10=18000\\mathrm\{PAR10\}=18000s\. We do not extrapolate beyond the scenario cutoff \(e\.g\. to72007200s\) because uncensored raw run logs are not available\. Table[7](https://arxiv.org/html/2609.13785#A12.T7)reportsSpart,Se2e,G\(R\),Apart,Ae2e,ρS\_\{\\mathrm\{part\}\},S\_\{\\mathrm\{e2e\}\},G\(R\),A\_\{\\mathrm\{part\}\},A\_\{\\mathrm\{e2e\}\},\\rhounder both cutoffs; the direction of every conclusion is preserved\.
Table 6:PROTEUS\-2014 feature\-policy sensitivity\. Headline pipeline quantities under \(a\) the primarycspfeature step \(3636features\) and \(b\) the all\-feature \(198198\) sensitivity\. Thecspsetting is the deployment model \(encoding\-conditional features are not visible to the family selector\); the all\-feature setting is a stronger\-information setting that shows what would happen if the family selector could see encoding features computable only after an encoding has been chosen\.This sensitivity is exploratory and is not used for headline claims\. Directional conclusions are unchanged under both feature policies:G\>0G\>0on both decomposed pipelines, and the all\-feature setting reducesGGby roughly25%25\\%\(0\.135→0\.0980\.135\\to 0\.098on multiclass,0\.134→0\.1000\.134\\to 0\.100on pairwise\) via improvement inSe2eS\_\{\\mathrm\{e2e\}\}rather than reduction inSpartS\_\{\\mathrm\{part\}\}— the within\-family selector benefits more from extra features than the family oracle does\. Full outputs are included in the companion reproducibility package\.
Table 7:PROTEUS\-2014 cutoff sensitivity\. Headline pipeline quantities under the36003600s scenario cutoff \(primary\) and a18001800s sensitivity in which anyokrun exceeding18001800s is re\-classified as a timeout and penalised at PAR10\. We do not extrapolate beyond the scenario cutoff because uncensored raw run logs are not available\.This sensitivity is exploratory and is not used for headline claims\. Directional conclusions are unchanged:G\(R\)G\(R\)is essentially invariant under the cutoff shortening \(pairwise gap0\.134→0\.1340\.134\\to 0\.134, multiclass0\.135→0\.1360\.135\\to 0\.136\); halving the cutoff drops1414all\-timeout instances \(n:3565→3551n:3565\\to 3551\) and pulls the flat baseline’s PAR10 slightly downward without changing the conclusion\. Full outputs are included in the companion reproducibility package\.
## Appendix MMetric\-transform sensitivity on B4–B5
This appendix checks whether the B4 sign\-change and B5 absorption conclusions are artefacts of the per\-instance min–max utility transform of §[5\.1](https://arxiv.org/html/2609.13785#S5.SS1)\. The substantive sensitivity check is decomp\-vs\-flat advantage transfer under a second common\[0,1\]\[0,1\]transform, per\-instance rank utility\. Appendix[A](https://arxiv.org/html/2609.13785#A1)defines the primary utility from raw PAR10/runtime before row\-normalisation, but the downstream prediction\-cache diagnostic used for this sensitivity stores normalised utilities and predictions rather than a reusable raw\-cost matrix aligned to every recomputed comparison mask\. The “native” column in the released diagnostic therefore falls back to a monotone−\-\\,utility proxy and is reported only as an implementation sanity check, not as independent raw\-second evidence\. We consequently avoid using the pointwise non\-negativity of a partition\-oracle cost gap as robustness evidence\.
Table 8:Metric\-transform sensitivity for decomp\-vs\-flat advantage transfer on B4–B5\. Positive advantages favour decomposition\. Min–max is the headline utility; rank utility maps the best algorithm on an instance to11and the worst to00\. Rank utility confirms the B4 sign change and the strong B5 absorption pattern, although B5’s multiclass rank end\-to\-end margin is effectively zero\.Full long\-form outputs, including the native proxy sanity check, are inresults/exp10\_native\_metric\. Because the proxy is monotone\-equivalent to the headline utility, it is not used as separate robustness evidence\.
## Appendix NWithin\-family selector robustness sweep
Figure 5:G\(R\)G\(R\)averaged over family\-selector choice, by within\-family selector class\. B1–B3 sweep four classes \(per\-algorithm GBDT, multiclass, pairwise OVO, single\-best\); B4–B5 sweep three \(pairwise OVO is omitted on ASlib due to its𝒪\(\|F\|2\)\\mathcal\{O\}\(\|F\|^\{2\}\)solver\-pair cost\)\. Family\-selector class changesSpartS\_\{\\mathrm\{part\}\}negligibly \(within5×10−35\\times 10^\{\-3\}\), while within\-family selector class shiftsGGby up to2\.2×2\.2\\timeson TALENT \(B3\) and produces a0\.210\.21gap on B5 under the main\-sweep single\-best\-on\-train selector\.§[5\.5](https://arxiv.org/html/2609.13785#S5.SS5)summarises the within\-family selector sensitivity sweep, visualised in Figure[5](https://arxiv.org/html/2609.13785#A14.F5)\. On B1–B3 we sweep four classes \(per\-algorithm GBDT, multiclass, pairwise OVO, single\-best\-on\-train\); on B4–B5 we sweep three \(pairwise\-within\-family is omitted because itsO\(\|F\|2\)O\(\|F\|^\{2\}\)cost on the larger ASlib pools is prohibitive\)\.G\(R\)G\(R\)ranges\[0\.025,0\.031\]\[0\.025,0\.031\]on B4 \(spread1\.28×1\.28\\times\) and\[0\.128,0\.211\]\[0\.128,0\.211\]on B5 \(spread1\.65×1\.65\\times\)\. Even the best within\-family selector in this sweep produces a substantial gap on both benchmarks; the largest main\-sweep PROTEUS\-2014 value isG=0\.211G=0\.211under the single\-best\-on\-train selector\. The deployment\-fidelity gap is therefore not a property of any single weak within\-family selector\.
Table[9](https://arxiv.org/html/2609.13785#A14.T9)reports the full2424\-row robustness sweep on B1–B3 \(2 family selectors×\\times4 within\-family selectors×\\times3 benchmarks; the pairwise\-within\-family configuration is omitted on B4–B5 due to itsO\(\|F\|2\)O\(\|F\|^\{2\}\)cost\)\. Family\-selector class changesSpartS\_\{\\mathrm\{part\}\}negligibly \(the row pairs comparing PW\-family vs MC\-family at fixed within\-class agree to≤5×10−3\\leq 5\\times 10^\{\-3\}on every benchmark\), while within\-family selector class shiftsGGby up to1\.44×1\.44\\timeson B1 and B2 and by2\.23×2\.23\\timeson B3 \(TALENT,G∈\[0\.041,0\.092\]G\\in\[0\.041,0\.092\]\)\.
Table 9:Within\-family selector robustness sweep on B1–B3\. Abbreviations: family selectors PW = pairwise OVO, MC = multiclass softmax; within\-family selectors GBDT = per\-algorithm gradient\-boosted regressor, MC = multiclass classifier, PW = pairwise OVO classifier, SB = single\-best\-on\-train\. Each row is computed on its row\-specific common valid\-instance mask;G=Spart−Se2eG=S\_\{\\mathrm\{part\}\}\-S\_\{\\mathrm\{e2e\}\}holds before rounding\.
## Appendix OFull numerical tables for gap and advantage\-transfer analyses
Table[12](https://arxiv.org/html/2609.13785#A15.T12)reports per\-cell\(Spart,Se2e,G\)\(S\_\{\\mathrm\{part\}\},S\_\{\\mathrm\{e2e\}\},G\)with95%95\\%CIs and paired Wilcoxonpp\-values for all1515\(benchmark, decomposition\) cells \(Figure[1](https://arxiv.org/html/2609.13785#S5.F1)\)\. Table[13](https://arxiv.org/html/2609.13785#A15.T13)reports the full1515\-pair advantage transfer table including the pairwise\-vs\-multiclass cells omitted from Table[2](https://arxiv.org/html/2609.13785#S5.T2)for space\. Figure[7](https://arxiv.org/html/2609.13785#A15.F7)visualises the same1515comparisons in\(Apart,Ae2e\)\(A\_\{\\mathrm\{part\}\},A\_\{\\mathrm\{e2e\}\}\)space\. Figure[6](https://arxiv.org/html/2609.13785#A15.F6)plots the deployment\-safety condition for the1010decomp\-vs\-flat comparisons referenced in §[5\.3\.1](https://arxiv.org/html/2609.13785#S5.SS3.SSS1), and Table[10](https://arxiv.org/html/2609.13785#A15.T10)summarises the partition\-only identification intervals \(Theorem[2](https://arxiv.org/html/2609.13785#Thmtheorem2)\) referenced in §[5\.4](https://arxiv.org/html/2609.13785#S5.SS4)\.
Figure 6:Deployment\-safety plot for the1010decomp\-vs\-flat comparisons\. By Corollary[2](https://arxiv.org/html/2609.13785#Thmcorollary2)\(general form\), the partition\-level advantage transfers iffApart\>ΔGA\_\{\\mathrm\{part\}\}\>\\Delta G, whereΔG=G\(R\)−G\(F\)\\Delta G=G\(R\)\-G\(F\), i\.e\. below the dashedy=xy=xline\. B1, B2, B4 occupy distinct absorption regimes: B2 and B4 lie in the unsafe \(red\) region and have sign\-changing point estimates; B1 remains deployment\-safe but is almost erased\. B3 and B5 \(PROTEUS\-2014\) are deployment\-safe but partially absorbed; B5’s3333\-point partition advantage shrinks to a2020\-point end\-to\-end advantage,ρ≈0\.41\\rho\\approx 0\.41\. All flat controls haveG\(F\)=0G\(F\)=0\.Table 10:Partition\-only identification intervals \(Theorem[2](https://arxiv.org/html/2609.13785#Thmtheorem2)\) summary\. “Crosses 0” counts the fraction of pairs whose interval strictly crosses zero, i\.e\. the fraction of comparisons on which the partition\-level report alone \(with selected\-family range information\) does not certify a deployable winner\. Median interval widthW¯=\(WR1\+WR2\)/2\\bar\{W\}=\(W\_\{R\_\{1\}\}\+W\_\{R\_\{2\}\}\)/2is the half\-interval that the partition\-level report cannot resolve\.Table 11:Per\-instance margin–regret diagnostic rates for decomposed pipelines\. The reported quantity is the selected\-family existential violation rate from Theorem[1](https://arxiv.org/html/2609.13785#Thmtheorem1); bootstrap intervals are omitted here for space and are included inresults/exp9\_margin\_regret\.Figure 7:Partition\-level advantageApartA\_\{\\mathrm\{part\}\}against end\-to\-end advantageAe2eA\_\{\\mathrm\{e2e\}\}across all1515pairwise method comparisons \(color by benchmark, shape by pair type\)\. The dashed diagonal is full transfer \(Apart=Ae2eA\_\{\\mathrm\{part\}\}=A\_\{\\mathrm\{e2e\}\}\); points near the horizontal axis indicate absorption; points in opposite quadrants indicate sign changes\. These cells are outlined in red\.*\(a\)*full range showing B5’s outlier position;*\(b\)*zoom on the central region containing the four sign\-changing cells \(decomp\-vs\-flat on B2 and B4 — all in the lower\-right quadrant, where partition prefers decomposition but e2e prefers flat\), the high\-absorption B1 cells, and the near\-origin pairwise\-vs\-multiclass comparisons\.Pairwise\-vs\-multiclass entries with\|Apart\|<5×10−3\|A\_\{\\mathrm\{part\}\}\|<5\\times 10^\{\-3\}are marked NA forρ\\rhoper the reporting threshold in §[6](https://arxiv.org/html/2609.13785#S6)\.
Table 12:Full deployment\-fidelity\-gap table by \(benchmark, decomposition\)\.G=Spart−Se2eG=S\_\{\\mathrm\{part\}\}\-S\_\{\\mathrm\{e2e\}\}; CI is a95%95\\%percentile\-bootstrap interval over per\-instance gaps;ppis the paired Wilcoxonpp\-value comparingupartu\_\{\\mathrm\{part\}\}vsue2eu\_\{\\mathrm\{e2e\}\}on the same instances\.Table 13:Full advantage\-transfer table: all1515\(benchmark, pair\) comparisons\. “pw” = pairwise OVO, “mc” = multiclass softmax\.ρ=\(G\(R1\)−G\(R2\)\)/Apart\\rho=\(G\(R\_\{1\}\)\-G\(R\_\{2\}\)\)/A\_\{\\mathrm\{part\}\}marked NA when\|Apart\|<5×10−3\|A\_\{\\mathrm\{part\}\}\|<5\\times 10^\{\-3\}\. The final column marks sign\-changing point estimates, not a large\-margin decision threshold\.
## Appendix PStatistical methodology
##### Pairing unit\.
The pairing unit for all paired tests is the*test instance*, not the \(seed, instance\) pair: seeds for the same instance share folds and are not i\.i\.d\. We average per\-instance utilities across seeds first, then perform paired tests on the per\-instance means\.
##### Paired Wilcoxon tests\.
For each \(benchmark, decomposition\) cell we run a two\-sided Wilcoxon signed\-rank test on the per\-instance gap vectorupart−ue2eu\_\{\\mathrm\{part\}\}\-u\_\{\\mathrm\{e2e\}\}\. Reportedpp\-values are exact when ties allow, otherwise asymptotic\. Across1515cells the Holm\-Bonferroni adjustedpp\-values remain<10−15<10^\{\-15\}for all decomposed cells\.
##### Bootstrap confidence intervals\.
Confidence intervals onG\(R\)G\(R\)are computed by samplingB=10 000B=10\\,000bootstrap replicates of the per\-instance gap vector with replacement and reporting the2\.52\.5th and97\.597\.5th percentiles of the bootstrap mean distribution\. The same convention is applied toApart,Ae2eA\_\{\\mathrm\{part\}\},A\_\{\\mathrm\{e2e\}\}when reported with intervals\.
##### Effect size\.
For pairwise method comparisons we reportρ\\rhoas the primary effect\-size summary;ρ\\rhois dimensionless and combines the magnitude ofApartA\_\{\\mathrm\{part\}\}with the signed gap\-differenceG\(R1\)−G\(R2\)G\(R\_\{1\}\)\-G\(R\_\{2\}\)\. We do not report Cliff’sδ\\deltain the main paper becauseρ\\rhohas a direct algebraic interpretation in terms of absorption\.
## Appendix QCompute resources and reproducibility
##### Hardware\.
All experiments run on a single 8\-core CPU workstation; no GPU is required\. Sklearn’sn\_jobs=\-1is used for random\-forest training\. Memory footprint stays below44GB across all benchmarks\.
##### Wall\-clock time\.
A full reproduction of the reported analyses across all five benchmarks takes approximately33–44hours of total wall\-clock time on the reference workstation\. The most expensive configuration is the pairwise\-within\-family configuration on TALENT \(B3\), which trainsO\(\|F\|2⋅Kcv⋅seeds\)O\(\|F\|^\{2\}\\cdot K\_\{\\mathrm\{cv\}\}\\cdot\\text\{seeds\}\)binary random forests per fit; the largest TALENT family contains2525algorithms\.
##### Determinism\.
Splits and selectors are seeded; rerunning the scripts with the same seeds reproduces the numerical tables in this appendix to four decimal places\. Bootstrap CIs usenumpy\.random\.default\_rngwith a fixed seed\.
##### Software stack\.
Python 3\.10, scikit\-learn≥\\geq1\.3, NumPy≥\\geq1\.24, pandas≥\\geq2\.0, SciPy≥\\geq1\.10, matplotlib≥\\geq3\.7\. The companion environment file pins exact versions used to produce the tables in this paper\.
## Appendix RLiterature audit of oracle\-style reporting
This appendix substantiates the claim, made in §[2](https://arxiv.org/html/2609.13785#S2)and the abstract, that oracle\-style scores – Virtual Best Solver \(VBS\), selected\-portfolio VBS, virtual\-best encodings, and best\-in\-family summaries – are routinely reported in algorithm selection, solver portfolios, encoding selection, and tabular benchmarking\. Table[R](https://arxiv.org/html/2609.13785#A18)lists, for each of1212representative anchor citations, the oracle\-style reporting object that appears in the work, what that row supports about the prevalence of such reporting, and an explicit caveat indicating what the row does*not*establish\.
The audit makes a deliberately narrow empirical claim\. It does*not*claim that any cited work conflates a partition\-level oracle score with a deployable system score: each cited paper either reports the oracle quantity as an explicit upper bound \(e\.g\. AutoFolio’s oracle/VBS columns; Proteus’s VB\-CSP, VB\-SAT, and VB\-Encoding rows reported alongside the deployable Proteus row\) or studies the gap between such an oracle and a learned selector at a single decision level \(e\.g\. encoding selection\)\. The claim the audit*does*support is that an oracle\-style reporting object – a virtual\-best at some level, or a best\-in\-family summary – is the standard primary comparison object across all four sub\-areas\. Combined with our empirical results, this shows that the deployment\-fidelity gap defined in this paper is a gap that the existing reporting practice cannot diagnose: when a reported number takes the within\-family maximum after a family choice is made, the gap to a deployable end\-to\-end system that inherits both choices is not surfaced\.[Shmuel et al\. \[2025\]](https://arxiv.org/html/2609.13785#bib.bib17)is the representative tabular instance for which we have direct supporting evidence: the published headline counts of which family “wins” on which dataset are derived from the within\-family argmax on1010\-fold cross\-validation means, with no within\-family deployable selector specified, defined, or evaluated anywhere in the paper or its supplement\.
\\keepXColumns
Literature audit of oracle\-style reporting\. “Strength” indicates how directly the row supports the audit’s narrow claim that oracle\-style quantities \(VBS, selected\-portfolio VBS, virtual\-best encoding, best\-in\-family\) are pervasive across algorithm selection, solver portfolios, encoding selection, and tabular benchmarking \(=direct,=supporting,=adjacent context\)\. The caveat column states what each row does*not*establish\.AnchorOracle\-style reporting objectAudit support / caveatStr\.\\endfirstheadAnchorOracle\-style reporting objectAudit support / caveatStr\.\\endhead*continued on next page*\\endfoot\\endlastfoot[Cameron et al\. \[2016\]](https://arxiv.org/html/2609.13785#bib.bib5)VBS defined as a hypothetical algorithm picking the best solver per instance from a portfolio; SBS–VBS gap analysed for evaluator bias under randomised solvers\.Standard VBS diagnostic; bias paper for the full\-portfolio VBS, not a within\-family nested oracle\.*Caveat:*not a decomposed AS critique\.[Bischl et al\. \[2016\]](https://arxiv.org/html/2609.13785#bib.bib3)ASlib repository: AS scenarios benchmark deployable selectors against VBS and SBS as standardised references\.VBS/SBS as the AS\-benchmark\-ecosystem standard\.*Caveat:*ASlib does not require nor prohibit reporting nested oracle quantities; it standardises a single\-stage comparison\.[Lindauer et al\. \[2015\]](https://arxiv.org/html/2609.13785#bib.bib2)AutoFolio reports oracle/VBS as columns alongside the deployable selector in its result tables\.Joint reporting of deployable selector and oracle bound is the canonical AS format\.*Caveat:*AutoFolio explicitly labels oracle as a bound, not a deployable score\.[Lindauer et al\. \[2019\]](https://arxiv.org/html/2609.13785#bib.bib23)The 2015/2017 AS competitions normalise scores by the SBS–VBS interval\.VBS is a competition\-protocol primitive, not just a per\-paper choice\.*Caveat:*normalisation is single\-level, not nested\.[Hurley et al\. \[2014\]](https://arxiv.org/html/2609.13785#bib.bib14)Hierarchical solver portfolio: CSP\-vs\-SAT, then SAT\-encoding, then SAT solver\. Tables report VB\-Proteus, VB\-CSP, VB\-SAT, plus per\-encoding VB\-DirectOrder, VB\-Direct, VB\-Support alongside the deployable Proteus selector\.Direct evidence that decomposed solver portfolios report*subportfolio*virtual\-best scores\.*Caveat:*Proteus does not claim its subportfolio VBs are deployable; it reports the deployable system in the same table\. The deployment\-fidelity gap is precisely the difference between the rows\.[Bach et al\. \[2022\]](https://arxiv.org/html/2609.13785#bib.bib18)SAT solverkk\-portfolios: VBS of a selected portfolio is reported as a lower bound \(i\.e\. best achievable cost\) for any model\-based selector built on that portfolio\.Selected\-portfolio VBS is a normal evaluation object\.*Caveat:*this is portfolio analysis, not a decomposed AS pipeline\.[Kostovska et al\. \[2023\]](https://arxiv.org/html/2609.13785#bib.bib19)Black\-box optimisation AAS: AAS performance is reported relative to the virtual best solver*from the selected portfolio*, after a portfolio\-selection step\.Recent \(2023\) anchor: selected\-portfolio VBS is still a primary comparator in AAS literature\.*Caveat:*portfolio selection is a different stage from the within\-family selection studied here\.[Stojadinović and Marić \[2014\]](https://arxiv.org/html/2609.13785#bib.bib22)CSP\-to\-SAT encoding selection: motivates choosing among direct, log, support, order encodings using CSP\-side syntactic features because no encoding dominates uniformly\.Encoding\-level family selection is a real research problem; the family abstraction in our framework maps directly onto the encoding choice\.*Caveat:*meSAT is encoding\-selection methodology, not an oracle\-vs\-deployable critique\.[Ulrich\-Oltean et al\. \[2022\]](https://arxiv.org/html/2609.13785#bib.bib20)Pseudo\-Boolean / linear\-integer constraint encoding selection: virtual\-best encoding reported as the upper bound for a learned encoding selector\.Per\-instance virtual\-best encoding is the standard upper bound in encoding selection\.*Caveat:*a single\-level oracle, evaluated correctly as a bound\.[Ulrich\-Oltean et al\. \[2023\]](https://arxiv.org/html/2609.13785#bib.bib21)Journal extension: explicitly analyses the single\-best vs virtual\-best encoding gap and reports how much of that gap a supervised encoding selector closes\.Single\-best/virtual\-best gap is now a primary analysis object in the encoding\-selection literature; close to the spirit ofG\(R\)G\(R\)but applied to a flat encoding choice rather than a decomposed selector\.*Caveat:*not a decomposed AS pipeline\.[McElfresh et al\. \[2023\]](https://arxiv.org/html/2609.13785#bib.bib10)TabZilla: trains a meta\-model to predict whether the best neural network outperforms the best gradient boosting model on a given dataset; analyses choosing the best algorithm*family*versus tuning a single model\.Best\-in\-family is the primary comparison object in tabular benchmarking\.*Caveat:*an analysis paper, not an AS\-system claim\.[Shmuel et al\. \[2025\]](https://arxiv.org/html/2609.13785#bib.bib17)Neurocomputing 2025:111111datasets×\\times2020models×\\times1010folds\. Headline ranking tables report each model’s number of datasets where it is best\-in\-family on the CV mean\. The meta\-learner labelY¯i=𝟏\[maxm∈MLscorei\(m\)\>maxm∈DLscorei\(m\)\]\\bar\{Y\}\_\{i\}=\\mathbf\{1\}\[\\max\_\{m\\in\\mathrm\{ML\}\}\\mathrm\{score\}\_\{i\}\(m\)\>\\max\_\{m\\in\\mathrm\{DL\}\}\\mathrm\{score\}\_\{i\}\(m\)\]takes the within\-family maximum on the same scores\. No within\-family deployable selector is defined\.Named\-instance evidence: a partition\-level oracle quantity is the headline conclusion objectin a peer\-reviewed tabular benchmark, with no deployable counterpart specified anywhere in paper or supplement\.*Caveat:*the paper makes no AS claim and does not advocate deploying these summaries; the omission is the absence of a within\-family selector, not the misuse of one\.
##### What the audit does and does not buy\.
The audit gives the abstract’s strong claim a defensible empirical grounding: oracle\-style reporting objects exist at every level of the AS / portfolio / encoding / tabular pipeline, and they are the primary comparison objects in those sub\-areas\. What it does*not*do is allege misuse by any cited work\. The contribution of the present paper is therefore not a critique of prior reporting but the construction of the missing object – the deployable system that inherits the partition’s family choice – and the measurement of the gapG\(R\)G\(R\)between the partition\-level oracle score and that system’s score\. The five most directly load\-bearing anchors for our framing are[Cameron et al\. \[2016\]](https://arxiv.org/html/2609.13785#bib.bib5)\(VBS as the canonical AS oracle diagnostic\),[Bischl et al\. \[2016\]](https://arxiv.org/html/2609.13785#bib.bib3)\(VBS in the benchmark\-ecosystem standard\),[Lindauer et al\. \[2015\]](https://arxiv.org/html/2609.13785#bib.bib2)\(deployable\-selector vs\. oracle joint reporting\),[Hurley et al\. \[2014\]](https://arxiv.org/html/2609.13785#bib.bib14)\(hierarchical decomposed portfolio with subportfolio virtual\-best scores\), and either[Ulrich\-Oltean et al\. \[2023\]](https://arxiv.org/html/2609.13785#bib.bib21)or[Shmuel et al\. \[2025\]](https://arxiv.org/html/2609.13785#bib.bib17)for a recent, named instance of single\-best vs\. virtual\-best comparison respectively in encoding selection and tabular benchmarking\.
## Appendix SInner\-selector stress\-test full results
This appendix provides the per\-cell numbers backing the*Inner\-selector stress test*paragraph in §[5\.5](https://arxiv.org/html/2609.13785#S5.SS5)\. We sweep seven within\-family selector classes spanning single\-best\-on\-train, the paper\-default per\-algorithm GBDT, a tuned sklearn GBDT, sklearn RandomForest, sklearn MLP, LightGBM, and XGBoost, all under the multiclass family selector and the same55\-fold split protocol as the headline runs \(here with22seeds; the headline runs use55\)\. Hyperparameters: tuned GBDTn=300,d=5n=300,d=5; RFn=200,d=10n=200,d=10; MLP hidden\(64,32\)\(64,32\), max iter200200; LightGBMn=200,leaves=31,lr=0\.05n=200,\\text\{leaves\}=31,\\mathrm\{lr\}=0\.05; XGBoostn=200,d=6,lr=0\.05n=200,d=6,\\mathrm\{lr\}=0\.05\. The package’s existingWithinFamilySelectorcontract is reused; the new selector classes are declared inline in the experiment script and do not modify any package code\.
The summary below reports\(Spart,Se2e,G\)\(S\_\{\\mathrm\{part\}\},S\_\{\\mathrm\{e2e\}\},G\)per cell\. The within\-family oracle \(G¯=0\\underline\{G\}=0\) is omitted as the trivial upper bound on each row;SpartS\_\{\\mathrm\{part\}\}is invariant across within\-classes within a benchmark by construction \(Lemma[1](https://arxiv.org/html/2609.13785#Thmlemma1)\) and serves as a sanity check\. We use this appendix only as aGG\-robustness summary: decision\-level decomp\-vs\-flat verdicts are reported in Table[2](https://arxiv.org/html/2609.13785#S5.T2)and Table[13](https://arxiv.org/html/2609.13785#A15.T13)under the common\-instance, singleton\-flat protocol\. The six learned within\-family regressors are separated from the non\-learned SingleBest reference because SingleBest is a useful stress case but not the learned selector class used in the main runs\.
Inner\-selector spectrum, decomposed pipeline\.Per\-cellSpart,Se2e,GS\_\{\\mathrm\{part\}\},S\_\{\\mathrm\{e2e\}\},Gunder the paper’s primary multiclass family selector\. For each benchmark the six learned within\-family regressors are listed first, followed by the non\-learned SingleBest reference\.SpartS\_\{\\mathrm\{part\}\}is constant within a benchmark by Lemma[1](https://arxiv.org/html/2609.13785#Thmlemma1)\.
After this full spectrum, we also reran targeted advantage\-transfer checks aligned with the headline decomp\-vs\-flat comparisons\. On B2/B4, a55\-seed sweep with RF and LightGBM under both multiclass and pairwise family selectors keepsG\(R\)\>0G\(R\)\>0in every cell \(B2:0\.0460\.046–0\.0490\.049; B4:0\.0240\.024–0\.0270\.027\)\. The correspondingAe2eA\_\{\\mathrm\{e2e\}\}values remain near zero \(B2:−0\.004\-0\.004to\+0\.001\+0\.001; B4:−0\.006\-0\.006to−0\.003\-0\.003\), so the strict sign of these small\-margin cells is selector\-sensitive, but no stronger selector yields a reliable positive decomp\-vs\-flat advantage\. A lighter33\-seed B5 default\-vs\-RF rerun keepsρ\\rhoat0\.4070\.407–0\.4140\.414\. The outputs are inresults/exp\_inner\_selector\_spectrum\_stronger\_2026\_05\_05\_b2b4andresults/exp\_inner\_selector\_spectrum\_stronger\_2026\_05\_05\_b5\_light\.
##### Reading the table\.
Across the six learned within\-family regressors, the spread ofGGis modest within each benchmark: B10\.00820\.0082–0\.02350\.0235, B20\.04530\.0453–0\.05420\.0542, B30\.06810\.0681–0\.08080\.0808, B40\.02220\.0222–0\.02670\.0267, and B50\.13290\.1329–0\.14310\.1431\. The non\-learned SingleBest reference should be read separately: it is competitive on B1/B2, worse on B3/B4, and much worse on B5, where it yieldsG=0\.2608G=0\.2608\(26\.1 pp\)\. This 26 pp SingleBest stress case is distinct from the 21 pp maximum in Figure[5](https://arxiv.org/html/2609.13785#A14.F5), which comes from the main robustness sweep over the paper’s within\-family selector classes\. The learned\-selector spectrum therefore supports the robustness claim without relying on the stronger SingleBest stress case\.Similar Articles
Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
This paper critiques existing benchmark contamination mitigation metrics and proposes SA-PPG (Stratified Aggregate of Per-question Probability Gaps) for more reliable evaluation, alongside RailCap, a decoding-time mitigation method that caps greedy fallback tokens to suppress memorization.
Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks
This paper audits offline root-cause-analysis benchmarks and finds that pooled leaderboards hide subsystem-specific winners, using pairwise comparisons on 778 cases across 11 subsystems. It releases a 320-line audit module for recomputing per-subsystem stability checks.
QuoteBench: How Matched Scores Can Hide Command-Path Failures
QuoteBench reveals that execution-boundary parsing errors significantly reduce LLM coding agent success, and disclosing the boundary helps recover performance, showing that evaluation must account for deployment configuration rather than treating matched scores as intrinsic model properties.
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
This paper applies Generalizability Theory to agent benchmarks, showing leaderboards rank specialization rather than capability, and proposes a framework (DDR) for sizing reliable deployment evaluations.
25% difference on a benchmark – just a mistake on how you run it😬
Discusses how a 25% difference on ARC-AGI was due to harness setup, showing GPT-5.6 Sol scoring 38% with proper evaluation, and critiques naive benchmark reporting in the industry.