Smart predict-then-robustly-optimize
Summary
This paper proposes a robust variant of smart predict-then-optimize that accounts for feature perturbations, providing a convex surrogate with theoretical guarantees and demonstrating superior performance over standard methods.
View Cached Full Text
Cached at: 07/27/26, 07:41 AM
# 1 Introduction
Source: [https://arxiv.org/html/2607.21773](https://arxiv.org/html/2607.21773)
\\RRHSecondLine\\LRHSecondLine\\OneAndAHalfSpacedXI\\TheoremsNumberedThrough\\ECRepeatTheorems\\EquationsNumberedThrough
\\RUNTITLE
Smart predict\-then\-robustly\-optimize
\\TITLE
Smart predict\-then\-robustly\-optimize
\\ARTICLEAUTHORS\\AUTHOR
Aakil Caunhye, Xuefei Lu, Belen Martin\-Barragan
\\ABSTRACT
In this paper, we propose and study a robust variant of the smart predict\-then\-optimize approach that accounts for prediction shifts due to disturbance in the covariate feature space\. While traditional integrated\-learning\-and\-optimization models assume that side information is perfectly revealed, empirical data\-driven features are frequently corrupted or noisy at the time of decision\-making, leading to fragile operational policies\. To bridge this gap, we integrate principles of robust optimization directly into the predictive\-prescriptive pipeline via a smart predict\-then\-robustly optimize loss and establish a computationally tractable convex surrogate, designed to hedge against worst\-case feature perturbations\. On the theoretical front, we formalize the structural validity of this surrogate by proving its approximation error probability decays exponentially according to a sub\-Gaussian concentration profile\. Furthermore, we establish that under mild assumptions, the surrogate is Fisher consistent with high probability\. We also prove necessary conditions under which our framework outperforms standard smart predict\-then\-optimize and maintain its superiority even when the standard method is equipped with regularized upstream predictions\. Numerical experiments validate that our robust framework consistently yields significant performance improvements over standard methods, both in out\-of\-sample terms and in training stability\.
\\KEYWORDS
Contextual optimization; robust optimization; smart predict\-then\-optimize; linear regression; data\-driven optimization
Contextual optimization has emerged as a dominant prescriptive modeling paradigm that leverages auxiliary covariate data to enhance decision\-making\. Its rise in popularity is driven by the rapid surge in data availability and the operational necessity of integrating machine learning models into decision optimization pipeline\. Successful applications of this paradigm now span diverse fields, including portfolio optimization\(Banet al\.[2018](https://arxiv.org/html/2607.21773#bib.bib39)\), food ordering and delivery\(Liuet al\.[2021](https://arxiv.org/html/2607.21773#bib.bib38)\), energy infrastructure planning\(Dontiet al\.[2017](https://arxiv.org/html/2607.21773#bib.bib37)\)and medical decision\-making\(Keyvanshokoohet al\.[2019](https://arxiv.org/html/2607.21773#bib.bib31)\)\. The recent survey bySadanaet al\.\([2024](https://arxiv.org/html/2607.21773#bib.bib36)\)consolidates the literature on contextual optimization and highlights the state\-of\-the\-art\. In this paper, we focus on the smart predict\-then\-optimize \(SPO\) paradigm introduced byElmachtoub and Grigas \([2022](https://arxiv.org/html/2607.21773#bib.bib50)\)\. The SPO paradigm conceptually involves two primary components: a predictor and an optimizer\. The predictor uses a training dataset\(๐i,๐i\)i=1n\(\\bm\{x\}\_\{i\},\\bm\{c\}\_\{i\}\)\_\{i=1\}^\{n\}to learn a mapping๐:โpโฆโd\\bm\{f\}:\\mathbb\{R\}^\{p\}\\mapsto\\mathbb\{R\}^\{d\}from covariates๐\\bm\{x\}to cost๐\\bm\{c\}\. The optimizer addresses a downstream decision problem
zโโ\(๐\)โmin๐โ๐ฒโก๐โคโ๐,\\displaystyle z^\{\*\}\(\\bm\{c\}\)\\coloneqq\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\bm\{c\}^\{\\top\}\\bm\{w\},where๐ฒโโd\\mathcal\{W\}\\subseteq\\mathbb\{R\}^\{d\}represents the feasible decision space\. A traditional, purely predictive approach treats these components sequentially \- first minimizing a statistical loss \(e\.g\., Mean Squared Error\) and then passing the point prediction to the optimizer\. However, this ignores the downstream impact of prediction errors on the resulting decisions\. In contrast, the SPO framework integrates these steps by defining a loss function based on regret: the difference between the cost of the decision made under the prediction๐โ\(๐\)\\bm\{f\}\(\\bm\{x\}\)and the cost of the true optimal decisionzโโ\(๐\)z^\{\*\}\(\\bm\{c\}\)\. Formally, letting\[n\]=\{1,โฆ,n\}\[n\]=\\\{1,\\dots,n\\\}, the SPO problem is the bilevel program
min๐โโ1nโโiโ\[n\]\(๐iโคโ๐โiโ\(๐โ\(๐i\)\)โzโโ\(๐i\)\),where๐โiโ\(๐^\)โargโกmin๐โ๐ฒโก๐^โคโ๐,โiโ\[n\],\\displaystyle\\begin\{split\}\\min\_\{\\bm\{f\}\\in\\mathcal\{H\}\}&\\frac\{1\}\{n\}\\sum\_\{i\\in\[n\]\}\\big\(\\bm\{c\}\_\{i\}^\{\\top\}\\bm\{w^\{\*\}\}\_\{i\}\(\\bm\{f\}\(\\bm\{x\}\_\{i\}\)\)\-z^\{\*\}\(\\bm\{c\}\_\{i\}\)\\big\),\\\\ \\text\{where \}&\\bm\{w^\{\*\}\}\_\{i\}\(\\hat\{\\bm\{c\}\}\)\\in\\arg\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\},\\ \\forall i\\in\[n\],\\end\{split\}\(SPO\)whereโ\\mathcal\{H\}is a hypothesis class of functions\.
While the SPO framework effectively captures the relationship between๐\\bm\{x\}and๐\\bm\{c\}, in a way that minimizes average decision loss, it typically assumes that data is uncontaminated, and thus that the relationship๐=๐โ\(๐\)\+๐บ\\bm\{c\}=\\bm\{f\}\(\\bm\{x\}\)\+\\bm\{\\varepsilon\}is accurately observed up to an irreducible stochastic noise๐บ\\bm\{\\varepsilon\}\. Beyond this inherent noise, systematic or adversarial errors can severely degrade data quality\. Such discrepancies frequently arise from practical limitations, including sensor measurement faults, data recording errors, temporal lags in reporting, and model misspecification\. Even marginal contaminations can lead to significantly corrupted predictions, which are then amplified by the optimizer\. The phenomenon of data uncertainty negatively impacting decisions, often termed the โoptimizerโs curse"\(Smith and Winkler[2006](https://arxiv.org/html/2607.21773#bib.bib17)\), results in decisions that appear optimal in\-sample but perform poorly in real\-world deployment\. Even with uncorrupted \(but still uncertain due to noise\) predictions, the optimizerโs curse already leads to perceived over\-estimated value\. One can imagine that covariate contamination is likely to further amplify this effect\. The practical necessity for considering contextual data disturbance is evident in several domains:
###### Example 1\.1\(Renewable energy planning with meteorological data\)
Renewable power generators tend to have intermittent availabilities\. In the case of photovoltaic cells and wind turbines, for instance, meteorological data is needed to accurately plan installations\. However, data is typically available at regional scales and rarely for the exact location of installation\. As we move to lower levels of granularity, localized disturbances can be expected in meteorological data\.
###### Example 1\.2\(Equitable humanitarian logistics using socioeconomic indicators\)
Equity is crucial in humanitarian logistics planning\. Pure utilitarian planning tends to favour the more accessible, who are likely less vulnerable\. Vulnerability metrics are composites of covariates such as income, age, and other socioeconomic variables\. However, these covariates are collected at intervals, rather than updated real\-time\. As such, at the point of disaster occurrence, covariate data are frequently outdated\.
###### Example 1\.3\(Medical decision making with patient\-specific data\)
Patient\-specific data, such as medication adherence or lifestyle factors, are often self\-reported and subject to significant variability and recording errors\.
In these contexts, relying on nominal data values is insufficient, as available datasets are frequently plagued by spatial, temporal, or selection biases, thereby invalidating the baseline assumptions of the end\-to\-end prediction\-prescription pipeline\. To address this vulnerability, we propose a robust SPO framework that explicitly immunizes downstream decision\-making against data contamination\. Our approach leverages principles from robust optimization\(Bertsimaset al\.[2011](https://arxiv.org/html/2607.21773#bib.bib16)\)to ensure reliable, stable, and high\-performing prescriptive outputs under uncertainty\. Robust optimization models data uncertainty deterministically via bounded uncertainty sets\. This distribution\-free paradigm is uniquely suited to our setting, as real\-world data contamination rarely exhibits well\-defined stochastic properties or follows known probability distributions\. By optimizing against the worst\-case realizations within a constructed uncertainty set, our framework safeguards the decision\-making process against corrupted data, mitigating the optimizerโs curse and providing distribution\-free performance guarantees\.
### 1\.1Contributions
The overarching contribution of this research is the formal development and theoretical analysis of downstream decision robustification directly integrated within the SPO framework\. We propose a new paradigm that extends Smart Predict\-then\-Optimize \(SPO\) to Smart Predict\-then\-robustly\-Optimize \(SPrO\)\. We establish its theoretical foundations, computational properties, and explicit performance guarantees\. Our specific contributions are structured as follows:
- โขA new convex paradigm for robustification:We develop the SPrO framework and derive its computationally tractable convex surrogate, SPrO\+\. We prove that SPrO\+ maintains convexity with respect to both data decision variables and predictions๐^\\hat\{\\bm\{c\}\}\. This allows practitioners to utilize standard off\-the\-shelf convex solvers for robust end\-to\-end learning\. Furthermore, we characterize favorable analytical properties of SPrO\+, proving that it exhibits global boundedness, Lipschitz continuity, and behaves similarly to anฯต\\epsilon\-insensitive loss function under mild conditions\.
- โขSurrogate gap analysis:We establish that the convex surrogate acts as a mathematically valid, tight upper bound for the true, intractable loss \(SPrO\)\. Crucially, we characterize the exact approximation error between the true loss and its surrogate, proving that, under mild assumptions, the deviation probability decays exponentially\. In addition, we show that our convex surrogate, SPrO\+, is highly likely to be Fisher consistent with respect to its true loss, SPrO\.
- โขComparisons with SPO:We provide a rigorous and comprehensive characterization of performance gaps between SPrO and SPO, as well as between the convex surrogates SPrO\+ and SPO\+\. Specifically, we establish: 1. 1\.*Expected and pointwise comparison of SPrO/SPrO\+ against standard SPO/SPO\+:*We prove necessary conditions for SPrO/SPrO\+ to outperform traditional SPO/SPO\+ on average under generalized noise structures\. We bound this average regret explicitly as a function of the decision space complexity and the budget of uncertainty\. We also provide necessary budget of uncertainty conditions for SPrO\+ to dominate SPO\+ pointwise\. 2. 2\.*SPrO comparison against upstream\-robustified SPO:*Regularized models are often conceived to enforce robustness in linear regression\. We compare our framework with an SPO framework where the upstream prediction is regularized to hedge against worst\-case predictions\. We show the necessary conditions for SPrO to dominate SPO in this case\.
Section 2 provides background literature connected to SPO and elaborates on the SPO framework under linear hypothesis class\. Section 3 details our new SPrO framework, its convex surrogate, as well as the latterโs properties, robust counterpart and surrogate gap analysis\. Section 4 studies the conditions under which SPrO improves SPO, whereas section 5 performs this study to compare SPrO with upstream\-robustified SPO\. Section 6 implements our framework on a network flow model\.
## 2Background and literature review
The Smart Predict\-then\-Optimize \(SPO\) framework belongs to a rich and rapidly expanding stream of research on contextual optimization, where decisions are made by leveraging side information or covariates\. In a comprehensive survey,Sadanaet al\.\([2024](https://arxiv.org/html/2607.21773#bib.bib36)\)classify the contextual optimization landscape into three primary methodology streams: decision rule optimization, sequential learning and optimization, and integrated learning and optimization\. Our work directly positions itself within this third category, often referred to in the modern operations research and machine learning literature as Decision\-Focused Learning or End\-to\-End prediction and optimization\(Dontiet al\.[2017](https://arxiv.org/html/2607.21773#bib.bib37), Mandiet al\.[2020](https://arxiv.org/html/2607.21773#bib.bib20)\)\.
The paradigm of integrating predictive tasks with downstream prescriptive objectives breaks from the traditional two\-stage estimate\-then\-optimize approach, which is inherently agnostic to the downstream decision optimization\. This concept dates back to early applications in financial forecasting, notably pioneered byBengio \([1997](https://arxiv.org/html/2607.21773#bib.bib25)\), who optimized neural network parameters based on financial investment utility rather than mean squared error\. Modern treatments formalized this end\-to\-end principle by differentiating through optimization layers\. For instance,Dontiet al\.\([2017](https://arxiv.org/html/2607.21773#bib.bib37)\)andKonget al\.\([2022](https://arxiv.org/html/2607.21773#bib.bib21)\)approach the integrated framework by balancing predictive and prescriptive accuracy in a way that iterates between a stochastic programming decision\-making model and distribution parameter estimation\.Kallus and Mao \([2023](https://arxiv.org/html/2607.21773#bib.bib24)\)adapt random forest architectures, modifying the standard node\-splitting criteria to minimize decision\-induced regret rather than predictive variance\.
Closer to our specific structural focus,Elmachtoub and Grigas \([2022](https://arxiv.org/html/2607.21773#bib.bib50)\)formalized the standard SPO framework for problems where the contextual parameters appear linearly in the objective function\. They introduced the non\-convex SPO loss \(or decision regret\) along with its computationally tractable convex surrogate, SPO\+\. This paradigm has since been successfully adapted across several computational domains\.Elmachtoubet al\.\([2020](https://arxiv.org/html/2607.21773#bib.bib19)\)embed the SPO loss into the splitting rules of decision trees, whileMandiet al\.\([2020](https://arxiv.org/html/2607.21773#bib.bib20)\)extend the framework to combinatorial and mixed\-integer programming settings by utilizing interior point mappings and subgradient approximations\.
As an optimization\-aware regret minimization approach, the theoretical validity of SPO rests upon its asymptotic and non\-asymptotic statistical guarantees\.Elmachtoub and Grigas \([2022](https://arxiv.org/html/2607.21773#bib.bib50)\)initially established that the SPO\+ loss exhibits Fisher consistency under continuous, symmetric distributions\. Since then, the mathematical foundations of risk calibration have been deeply expanded\.Ho\-Nguyen and Kฤฑlฤฑnรง\-Karzan \([2022](https://arxiv.org/html/2607.21773#bib.bib11)\)provide generalized uniform calibration bounds and risk bounds for SPO\+\. Parallel to consistency, the sample efficiency of these estimators has been bounded via Rademacher complexity analysis\(El Balghitiet al\.[2019](https://arxiv.org/html/2607.21773#bib.bib22)\), and their convergence profiles have been mapped via fast conditional regret rates\(Huet al\.[2022](https://arxiv.org/html/2607.21773#bib.bib23)\)\.
While optimization under uncertainty remains at the core of the contextual optimization narrative, existing paradigms operate under a highly asymmetric assumption: while the unknown objective parameters \(e\.g\., costs, demands\) are treated as highly stochastic, the observed contextual features themselves are assumed to be perfectly revealed\. In historical context\-free settings, robust optimization \(RO\) and distributionally robust optimization \(DRO\) have long been used to protect against parameter noise\. More recently, this has inspired contextual extensions, such as the predict\-then\-calibrate framework ofSunet al\.\([2023](https://arxiv.org/html/2607.21773#bib.bib6)\)and the conformal contextual robust optimization ofPatelet al\.\([2024](https://arxiv.org/html/2607.21773#bib.bib5)\), which map features to robust uncertainty sets\.
Crucially, however, none of these frameworks account for the reality that the side information itself can be corrupted, noisy, or uncertain at the time of decision\-making\. While robust feature fitting is prominent in pure predictive statistics, its interaction with downstream optimization remains completely unexplored\. Our work bridges this exact gap, establishing the first robust decision\-focused learning framework that preserves computational tractability while providing explicit sub\-Gaussian concentration and Fisher consistency guarantees under feature space perturbations\.
### 2\.1Smart Predict\-then\-Optimize
The SPO loss function measures the decision regret incurred by using a predicted cost vector๐^\\hat\{\\bm\{c\}\}instead of the true realization๐\\bm\{c\}:
โSโPโOโ\(๐^,๐\)โ๐โคโ๐โโ\(๐^\)โzโโ\(๐\)\.\\displaystyle\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\coloneqq\\bm\{c\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\hat\{\\bm\{c\}\}\)\-z^\{\*\}\(\\bm\{c\}\)\.To keep expositions general, we henceforth define the feasibility space of decisions as a compact set
๐ฒโ\{๐โโd:gkโ\(๐\)โค0,โkโ\[m\]\},\\displaystyle\\mathcal\{W\}\\coloneqq\\\{\\bm\{w\}\\in\\mathbb\{R\}^\{d\}:g\_\{k\}\(\\bm\{w\}\)\\leq 0,\\forall k\\in\[m\]\\\},wheregkg\_\{k\}are proper, closed and convex functions\. To ensure the existence of dual solutions \(which will be required throughout the paper\), we invoke the standard Slater condition:
###### Definition 2\.1\(Slater point\(Zhenet al\.[2025](https://arxiv.org/html/2607.21773#bib.bib48)\)\)
The vector๐ฐโ \\bm\{w\}^\{\\dagger\}is a Slater point of๐ฒ\\mathcal\{W\}if \(1\)๐ฐโ โ๐ฒ\\bm\{w\}^\{\\dagger\}\\in\\mathcal\{W\}, \(2\)๐ฐโ โโฉkโ\[m\]riโก\(domโก\(gk\)\)\\bm\{w\}^\{\\dagger\}\\in\\cap\_\{k\\in\[m\]\}\\operatorname\{ri\}\(\\operatorname\{dom\}\(g\_\{k\}\)\)and \(3\)gkโ\(๐ฐโ \)<0g\_\{k\}\(\\bm\{w\}^\{\\dagger\}\)<0for everykโ\[m\]k\\in\[m\]such thatgkg\_\{k\}is nonlinear\. The notationriโก\(๐ณ\)\\operatorname\{ri\}\(\\mathcal\{X\}\)represents the relative interior of set๐ณ\\mathcal\{X\}\.
\{assumption\}
\[Slater condition\] The feasibility set๐ฒ\\mathcal\{W\}admits a Slater point\. The objective of SPO is to identify a predictive model๐โ\\bm\{f\}^\{\*\}from a hypothesis classโ\\mathcal\{H\}that minimizes the average empirical regret acrossnnobservations:
๐โ=argโกmin๐โโโก1nโโiโ\[n\]โSโPโOโ\(๐โ\(๐i\),๐i\)\.\\displaystyle\\bm\{f\}^\{\*\}=\\arg\\min\_\{\\bm\{f\}\\in\\mathcal\{H\}\}\\frac\{1\}\{n\}\\sum\_\{i\\in\[n\]\}\\ell\_\{SPO\}\(\\bm\{f\}\(\\bm\{x\}\_\{i\}\),\\bm\{c\}\_\{i\}\)\.In this study, we restrictโ\\mathcal\{H\}to the space of linear regression functions, where , where๐โ\(๐\)=๐ฉโ๐\\bm\{f\}\(\\bm\{x\}\)=\\bm\{B\}\\bm\{x\}and๐ฉโโdรp\\bm\{B\}\\in\\mathbb\{R\}^\{d\\times p\}\. Despite the emergence of complex non\-linear learners, linear regression remains a staple in Machine Learning \(ML\) due to its simplicity and intrinsic interpretability and explainability\. This transparency is paramount in high\-stakes environments \- such as autonomous systems or military logistics \- where legal accountability and ethical ramifications necessitate explainable decision\-making\(Vellido[2020](https://arxiv.org/html/2607.21773#bib.bib26)\)\. Beyond legal and ethical concerns,Rudinet al\.\([2022](https://arxiv.org/html/2607.21773#bib.bib2)\)argue that it is important not to assume one needs to sacrifice accuracy in order to gain interpretability and thus, when a simple interpretable model performs comparably to a complex one, preference must be given to the simple one\. In the SPO framework, we will see that the desirability of linear regression is retained, in the sense that SPO produces models with bilinear relationships between predictive fitting and the resulting decisions, leading to interpretable and explainable outcomes\.
By characterizing the optimal decision through the optimality conditions of the downstream problem, the loss minimization can be reformulated using the geometry of the feasible region\. Specifically, the SPO problem under linear hypotheses becomes:
min๐ฉโโdรpโก1nโโiโ\[n\]โSโPโOโ\(๐ฉโ๐i,๐i\)\\displaystyle\\min\_\{\\bm\{B\}\\in\\mathbb\{R\}^\{d\\times p\}\}\\frac\{1\}\{n\}\\sum\_\{i\\in\[n\]\}\\ell\_\{SPO\}\(\\bm\{B\}\\bm\{x\}\_\{i\},\\bm\{c\}\_\{i\}\)=min๐ฉโโdรpโก1nโโiโ\[n\]\(๐iโคโ๐โโ\(๐ฉโ๐i\)โzโโ\(๐i\)\)\\displaystyle=\\min\_\{\\bm\{B\}\\in\\mathbb\{R\}^\{d\\times p\}\}\\frac\{1\}\{n\}\\sum\_\{i\\in\[n\]\}\\big\(\\bm\{c\}\_\{i\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{B\}\\bm\{x\}\_\{i\}\)\-z^\{\*\}\(\\bm\{c\}\_\{i\}\)\\big\)=min๐ฉโโdรp๐^iโ\{๐ฒ:\(๐ฉโ๐i\)โคโ\(๐^iโ๐\)โค0,โ๐โ๐ฒ\},โiโ\[n\]โก1nโโiโ\[n\]\(๐iโคโ๐^iโzโโ\(๐i\)\)\\displaystyle=\\min\_\{\\begin\{subarray\}\{c\}\\bm\{B\}\\in\\mathbb\{R\}^\{d\\times p\}\\\\ \\hat\{\\bm\{w\}\}\_\{i\}\\in\\\{\\mathcal\{W\}:\(\\bm\{B\}\\bm\{x\}\_\{i\}\)^\{\\top\}\(\\hat\{\\bm\{w\}\}\_\{i\}\-\\bm\{w\}\)\\leq 0,\\,\\forall\\bm\{w\}\\in\\mathcal\{W\}\\\},\\,\\forall i\\in\[n\]\\end\{subarray\}\}\\frac\{1\}\{n\}\\sum\_\{i\\in\[n\]\}\\big\(\\bm\{c\}\_\{i\}^\{\\top\}\\hat\{\\bm\{w\}\}\_\{i\}\-z^\{\*\}\(\\bm\{c\}\_\{i\}\)\\big\)=min๐ฉโโdรp๐^iโ\{๐ฒ:โ๐ฉโ๐iโ๐ฉ๐ฒโ\(๐^i\)\},โiโ\[n\]โก1nโโiโ\[n\]\(๐iโคโ๐^iโzโโ\(๐i\)\),\\displaystyle=\\min\_\{\\begin\{subarray\}\{c\}\\bm\{B\}\\in\\mathbb\{R\}^\{d\\times p\}\\\\ \\hat\{\\bm\{w\}\}\_\{i\}\\in\\\{\\mathcal\{W\}:\-\\bm\{B\}\\bm\{x\}\_\{i\}\\in\\mathcal\{N\}\_\{\\mathcal\{W\}\}\(\\hat\{\\bm\{w\}\}\_\{i\}\)\\\},\\,\\forall i\\in\[n\]\\end\{subarray\}\}\\frac\{1\}\{n\}\\sum\_\{i\\in\[n\]\}\\big\(\\bm\{c\}\_\{i\}^\{\\top\}\\hat\{\\bm\{w\}\}\_\{i\}\-z^\{\*\}\(\\bm\{c\}\_\{i\}\)\\big\),where notation๐ฉ๐โ\(๐\)\\mathcal\{N\}\_\{\\mathcal\{A\}\}\(\\bm\{y\}\)represents the normal cone of set๐\\mathcal\{A\}at point๐\\bm\{y\}\. Therefore, the best regression coefficients must map the linear transformation of predictors to the normal cone of the feasibility set, as pictured in Figure[1](https://arxiv.org/html/2607.21773#S2.F1)\.
Figure 1:Illustration of normal cone solutionTo incentivize non\-trivial solutions, an variant of SPO loss, called the unambiguous SPO loss, is used to break ties via formulation
max๐โ๐ฒโโ\(๐ฉโ๐\)โก๐โคโ๐โzโโ\(๐\),\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\bm\{B\}\\bm\{x\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{\*\}\(\\bm\{c\}\),where๐ฒโโ\(๐^\)\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)is the set containing all optimization oracles๐โโ\(๐^\)\\bm\{w^\{\*\}\}\(\\hat\{\\bm\{c\}\}\)\. A significant challenge arises from the fact thatโSโPโO\\ell\_\{SPO\}\(in both its original and unambiguous forms\) is generally non\-convex, making direct optimization difficult\. To ensure computational tractability, we adopt the convex surrogate of SPO loss, the SPO\+ loss, formulated inElmachtoub and Grigas \([2022](https://arxiv.org/html/2607.21773#bib.bib50)\)as \(usually a scaling is applied to the predictor, but in linear regression, SPO\+ is scale invariant\):
โSโPโO\+โ\(๐ฉโ๐,๐\)โmax๐โ๐ฒโก\{๐โคโ๐โ\(๐ฉโ๐\)โคโ๐\}\+\(๐ฉโ๐\)โคโ๐โโ\(๐\)โzโโ\(๐\),\\displaystyle\\ell\_\{SPO\+\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\)\\coloneqq\\max\_\{\\begin\{subarray\}\{c\}\\bm\{w\}\\in\\mathcal\{W\}\\end\{subarray\}\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\(\\bm\{B\}\\bm\{x\}\)^\{\\top\}\\bm\{w\}\\\}\+\(\\bm\{B\}\\bm\{x\}\)^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\),The SPO\+ loss is particularly advantageous as it is convex and provides a computationally efficient subgradient, facilitating the use of standard first\-order optimization methods\. While the SPO framework leverages covariates to reduce cost uncertainty, it remains susceptible to disturbances in the predictors themselves\. Conventional ML research has addressed robust regression from a purely predictive standpoint\(El Ghaoui and Lebret[1997](https://arxiv.org/html/2607.21773#bib.bib35), Shivaswamyet al\.[2006](https://arxiv.org/html/2607.21773#bib.bib34), Xuet al\.[2008](https://arxiv.org/html/2607.21773#bib.bib45)\); however, in a prescriptive context, even minor perturbations in๐\\bm\{x\}can lead to sub\-optimal decision\-making\. As established in our motivating examples, geographical and temporal variabilities often contaminate covariate data\. Consequently, there is a compelling need to develop an SPO framework that is robust to such disturbances, immunizing the decision\-making process against uncertainty in the underlying features\.
### 2\.2Notations
The convex conjugate of a functionhhis defined as
hโโ\(๐\)โsup๐๐โคโ๐โhโ\(๐\)\.\\displaystyle h^\{\*\}\(\\bm\{y\}\)\\coloneqq\\sup\_\{\\bm\{w\}\}\\bm\{y\}^\{\\top\}\\bm\{w\}\-h\(\\bm\{w\}\)\.The perspective function of a proper, convex and lower semicontinuous functionhhis denoted byhโฯ:โqรโ\+โฆโh\\phi:\\mathbb\{R\}^\{q\}\\times\\mathbb\{R\}\_\{\+\}\\mapsto\\mathbb\{R\}where
\(hโฯ\)โ\(๐\)=\{ฯโhโ\(๐ฯ\)ifโฯ\>0๐โ\(๐โฃ๐\)ifโฯ=0\\displaystyle\(h\\phi\)\(\\bm\{y\}\)=\\begin\{cases\}\\phi h\(\\frac\{\\bm\{y\}\}\{\\phi\}\)&\\text\{if \}\\phi\>0\\\\ \\mathbb\{I\}\(\\bm\{y\}\\mid\\bm\{0\}\)&\\text\{if \}\\phi=0\\end\{cases\}andฮดโ\(๐โฃ๐ด\)\\delta\(\\bm\{y\}\\mid\\mathcal\{Y\}\)is the indicator function, defined as
๐โ\(๐โฃ๐ด\)=\{0ifโ๐โ๐ด\+โifโ๐โ๐ด\.\\displaystyle\\mathbb\{I\}\(\\bm\{y\}\\mid\\mathcal\{Y\}\)=\\begin\{cases\}0&\\text\{if \}\\bm\{y\}\\in\\mathcal\{Y\}\\\\ \+\\infty&\\text\{if \}\\bm\{y\}\\notin\\mathcal\{Y\}\.\\end\{cases\}The dual norm is denoted byโ๐โโ\\\|\\bm\{y\}\\\|\_\{\*\}, defined asโ๐โโ=maxโ๐โโค1โก๐โคโ๐\\\|\\bm\{y\}\\\|\_\{\*\}=\\max\_\{\\\|\\bm\{z\}\\\|\\leq 1\}\\bm\{z\}^\{\\top\}\\bm\{y\}\. The support function of a set๐\\mathcal\{A\}is represented ash๐โ\(๐\)=max๐โ๐โก๐โคโ๐h\_\{\\mathcal\{A\}\}\(\\bm\{c\}\)=\\max\_\{\\bm\{a\}\\in\\mathcal\{A\}\}\\bm\{c\}^\{\\top\}\\bm\{a\}\. We denote the unitdd\-dimensional ball asโฌ=\{๐โโd:โ๐โโค1\}\\mathcal\{B\}=\\\{\\bm\{u\}\\in\\mathbb\{R\}^\{d\}:\\\|\\bm\{u\}\\\|\\leq 1\\\}\. The Gaussian width of bounded set๐โโd\\mathcal\{A\}\\subset\\mathbb\{R\}^\{d\}isฯโ\(๐\)\\omega\(\\mathcal\{A\}\), defined asฯโ\(๐\)โ๐ผ๐โ\[sup๐โ๐๐โคโ๐\]\\omega\(\\mathcal\{A\}\)\\coloneqq\\mathbb\{E\}\_\{\\bm\{g\}\}\\left\[\\sup\_\{\\bm\{v\}\\in\\mathcal\{A\}\}\\bm\{g\}^\{\\top\}\\bm\{v\}\\right\], with๐โผNโoโrโmโaโlโ\(๐,๐d\)\\bm\{g\}\\sim Normal\(\\bm\{0\},\\mathbb\{I\}\_\{d\}\)being a standard Gaussian random vector inโd\\mathbb\{R\}^\{d\}\. Denote the diameter of a compact set๐\\mathcal\{A\}as๐โ\(๐\)โmax๐,๐โ๐โกโ๐โ๐โ2\\mathcal\{D\}\(\\mathcal\{A\}\)\\coloneqq\\max\_\{\\bm\{p\},\\bm\{q\}\\in\\mathcal\{A\}\}\\\|\\bm\{p\}\-\\bm\{q\}\\\|\_\{2\}\.
## 3A new paradigm: smart predict\-then\-robustly\-optimize
To immunize the SPO framework against covariate disturbances, we introduce the Smart Predict\-then\-Robustly\-Optimize \(SPrO\) paradigm\. Standard linear regression models the cost vector via the linear relationship๐=๐ฉโ๐\+๐บ\\bm\{c\}=\\bm\{B\}\\bm\{x\}\+\\bm\{\\varepsilon\}, where๐ฉ\\bm\{B\}is the matrix of regression coefficients and๐บ\\bm\{\\varepsilon\}represents the irreducible error\. Under data contamination, the observed covariates are perturbed such that they become๐\+๐น\\bm\{x\}\+\\bm\{\\delta\}, where๐น\\bm\{\\delta\}denotes the disturbance vector\. This contamination induces a shift in the predicted costs given by๐^=๐ฉโ๐\+๐ฉโ๐น\\hat\{\\bm\{c\}\}=\\bm\{B\}\\bm\{x\}\+\\bm\{B\}\\bm\{\\delta\}\. Notice that this formulation introduces an endogenous uncertainty term,๐ฉโ๐น\\bm\{B\}\\bm\{\\delta\}, because the impact of the covariate disturbance depends directly on the upstream prediction model parameters๐ฉ\\bm\{B\}\. To preserve computational tractability \- the necessity of which will become apparent in subsequent sections \- we approximate this phenomenon via an exogenous cost perturbation\. Specifically, we define the predictive model as๐^=๐ฉโ๐\+๐น\\hat\{\\bm\{c\}\}=\\bm\{B\}\\bm\{x\}\+\\bm\{\\delta\}, where the cost\-space uncertainty vector๐น\\bm\{\\delta\}serves as a direct proxy for covariate disturbances or prediction shifts\. We assume that๐น\\bm\{\\delta\}resides within a bounded uncertainty set๐ฐฮป\\mathcal\{U\}\_\{\\lambda\}, where
๐ฐฮปโ\{๐นโโd:โ๐นโโคฮป\}\.\\displaystyle\\mathcal\{U\}\_\{\\lambda\}\\coloneqq\\\{\\bm\{\\delta\}\\in\\mathbb\{R\}^\{d\}:\\\|\\bm\{\\delta\}\\\|\\leq\\lambda\\\}\.Although this exogenous formulation serves as an approximation, it remains structurally sound and aligns closely with the endogenous model under several realistic conditions on the coefficient matrix๐ฉ\\bm\{B\}\. First, the exact equivalence๐ฉโ๐น=๐น\\bm\{B\}\\bm\{\\delta\}=\\bm\{\\delta\}holds if the covariate disturbance lies within the eigenspace of๐ฉ\\bm\{B\}associated with an eigenvalue of11\. Second, a close approximation is achieved whenโ๐ฉโ๐โโคฯต\\\|\\bm\{B\}\-\\mathbb\{I\}\\\|\\leq\\epsilonfor a sufficiently smallฯต\>0\\epsilon\>0, since the condition๐ฉโ๐นโ๐น\\bm\{B\}\\bm\{\\delta\}\\approx\\bm\{\\delta\}can be guaranteed byโ๐ฉโ๐นโ๐นโโคฯตโโ๐นโ\\\|\\bm\{B\}\\bm\{\\delta\}\-\\bm\{\\delta\}\\\|\\leq\\epsilon\\\|\\bm\{\\delta\}\\\|\. Third, and more importantly, by the triangle inequality, the conditionโ๐ฉโ๐นโ๐นโโคฯตโโ๐นโ\\\|\\bm\{B\}\\bm\{\\delta\}\-\\bm\{\\delta\}\\\|\\leq\\epsilon\\\|\\bm\{\\delta\}\\\|implies that\(1โฯต\)โโ๐นโโคโ๐ฉโ๐นโโค\(1\+ฯต\)โโ๐นโ\(1\-\\epsilon\)\\\|\\bm\{\\delta\}\\\|\\leq\\\|\\bm\{B\}\\bm\{\\delta\}\\\|\\leq\(1\+\\epsilon\)\\\|\\bm\{\\delta\}\\\|\. This structural relationship indicates that๐ฉโ๐นโ๐น\\bm\{B\}\\bm\{\\delta\}\\approx\\bm\{\\delta\}whenever the coefficient matrix๐ฉ\\bm\{B\}acts as an approximate isometry over the uncertainty set\. This norm\-preserving property ensures that our exogenous representation remains highly accurate when the mapping of covariate disturbances through the regression matrix is stably bounded\.
Beyond physical data contamination, framing the prediction\-space perturbation as an exogenous shift fundamentally expands the scope of our framework to protect against traditional machine learning vulnerabilities, namely
1. 1\.Model misspecification: the uncertainty vector captures the systematic bias that arises when forcing a linear hypothesis onto complex, non\-linear real\-world cost structures\.
2. 2\.Finite\-sample training: it accounts for the statistical variance inherent to limited data, acting as a deterministic bound for the predictorโs performance gap on unseen instances\.
3. 3\.Overfitting: by optimizing against the worst\-case realizations within the uncertainty set, this formulation introduces an implicit robust regularization mechanism that mitigates the erratic and highly sensitive decision regrets triggered by overfitted coefficients\.
4. 4\.Out\-of\-distribution deployment: it safeguards the prescriptive pipeline when unexpected environmental, geographical, or temporal shifts cause the covariates to deviate from the historical training distribution\.
The idea of SPO is to fit a data model whose predictions produce decisions with optimal cost that closely approximates the true optimal cost\. In the SPrO context, these decisions yield the best worst\-case cost under prediction shift\. As such, the decision๐๐นโฃโโ\(๐^\)\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)that optimizes the worst\-case predicted cost can be extracted from
๐๐นโฃโ\(๐^\)โargmin๐โ๐ฒ\{max๐นโ๐ฐฮป\(๐^\+๐น\)โค๐\}=argmin๐โ๐ฒ\{๐^โค๐\+ฮปโฅ๐โฅโ\}โ๐ฒRโฃโ\(๐^\),\\displaystyle\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\\in\\arg\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\\{\\max\_\{\\bm\{\\delta\}\\in\\mathcal\{U\}\_\{\\lambda\}\}\(\\hat\{\\bm\{c\}\}\+\\bm\{\\delta\}\)^\{\\top\}\\bm\{w\}\\\}=\\arg\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\\coloneqq\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\),whereโ๐^โ๐ฉโ๐\.\\displaystyle\\text\{where \}\\hat\{\\bm\{c\}\}\\coloneqq\\bm\{B\}\\bm\{x\}\.The optimal worst\-case decisions thus minimize a regularized cost, with dual norm regularizerโ๐โโ\\\|\\bm\{w\}\\\|\_\{\*\}weighted by the budget of uncertaintyฮป\\lambda\. Therefore, robust downstream decision\-making offers a natural extension to its deterministic version, in a way that favors decision shrinkage \(or decision sparsity if our norm isโโ\\ell\_\{\\infty\}and thus dual norm becomesโ1\\ell\_\{1\}, replicating a Lasso\-type regularization on decisions\)\. We define the SPrO loss as
โSโPโrโOโ\(๐^,๐\)โ\(max๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐โzRโฃโโ\(๐\)\)\+,\\displaystyle\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\coloneqq\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\},whereโzRโฃโโ\(๐\)โmax๐โ๐ฒRโฃโโ\(๐\)โก\{๐โคโ๐\}\\displaystyle\\text\{where \}z^\{R\*\}\(\\bm\{c\}\)\\coloneqq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\\\}and\(โ
\)\+โmaxโก\{0,โ
\}\(\\cdot\)\_\{\+\}\\coloneqq\\max\\\{0,\\cdot\\\}\. Unlike standard SPO, SPrO quantifies decision loss within a prediction\-shift\-aware optimization framework\. Table[1](https://arxiv.org/html/2607.21773#S3.T1)summarizes the main differences between our approach and SPO\.
Table 1:Paradigm comparisonsRegret in SPrO thus reflects the performance penalty incurred by a robust downstream decision\-maker who explicitly accounts for this prediction shifts, measured relative to the true robust oracle\. The choice ofzRโฃโโ\(๐\)z^\{R\*\}\(\\bm\{c\}\)as the baseline oracle is dictated by the principle of decision\-maker consistency\. In standard SPO, the learnerโs nominal decision is evaluated against a nominal oraclezโโ\(๐\)z^\{\*\}\(\\bm\{c\}\)\. Because SPrO alters the downstream decision\-makerโs archetype \- forcing it to be robustly regularized to immunize against prediction shifts \- evaluating it against a non\-robust nominal oracle would introduce a structural mismatch\. Such an inconsistent benchmark would penalize the learner not just for poor prediction quality, but for the inherent conservatism of the robust policy itself\. By benchmarking againstzRโฃโโ\(๐\)z^\{R\*\}\(\\bm\{c\}\), we isolate the regret driven solely by the upstream estimation error, comparing a robust learner to an oracle that observes the true cost environment but remains bound to the same robust decision policy\. We will also show that the new oracle has tractability advantages\. Furthermore, in its original \(not unambiguous\) formโSโPโOโ\(๐^,๐\)=๐โคโ๐โโ\(๐^\)โzโโ\(๐\)\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\\bm\{c\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\hat\{\\bm\{c\}\}\)\-z^\{\*\}\(\\bm\{c\}\), SPO is susceptible to trivial solutions; for instance, a zero\-vector prediction \(๐^=๐\\hat\{\\bm\{c\}\}=\\bm\{0\}\) can artificially yield zero loss because the nominal decision set๐ฒโโ\(๐^\)\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)expands to the entire feasible region๐ฒ\\mathcal\{W\}\. In SPrO, the dual norm regularizerฮปโโ๐โโ\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}prevents this collapse, as the optimal decision must balance the nominal cost with a budget\-weighted regularizer\.
Feature\-independent regularization is a strong motivation behind our exogenous conceptualization of the prediction shift\. An endogenous prediction shift would have yielded the learner\-decision regularizerโ๐ฉโคโ๐โโ\\\|\\bm\{B\}^\{\\top\}\\bm\{w\}\\\|\_\{\*\}\. This weakens the interpretability of the loss function by decoupling the coefficient matrix๐ฉ\\bm\{B\}from the feature vector๐\\bm\{x\}, ultimately obscuring the direct comparison between the predicted cost๐^\\hat\{\\bm\{c\}\}and the true response๐\\bm\{c\}\. In addition, it introduces an adversarial bilinear term between the upstream prediction matrix๐ฉ\\bm\{B\}and downstream decision variables๐\\bm\{w\}, destroying the joint convexity required for efficient training\. Similar to the standard SPO, SPrO has the primary drawback of being non\-convex\. To produce a convex surrogate, we derive an SPrO\+ loss function as follows:
โSโPโrโOโ\(๐^,๐\)\\displaystyle\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\(max๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐โzRโฃโโ\(๐\)\)\+\\displaystyle=\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}=\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฃโโ\(๐^\)\+ฮปโโ๐๐นโฃโโ\(๐^\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle=\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}โค\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle\\leq\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}โโSโPโrโO\+โ\(๐^,๐\),\\displaystyle\\coloneqq\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\),where๐๐นโฒโ\(๐\)โargโกmax๐โ๐ฒRโฃโโ\(๐\)โก\{๐โคโ๐\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\in\\arg\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\\\}\. The inequality follows directly from the definition of๐๐นโฃโโ\(๐^\)\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\. Because๐๐นโฃโโ\(๐^\)โargโกmin๐โ๐ฒโก\{๐^โคโ๐\+ฮปโโ๐โโ\}\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\\in\\arg\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\\}, the relation๐^โคโ๐๐นโฃโโ\(๐^\)\+ฮปโโ๐๐นโฃโโ\(๐^\)โโโค๐^โคโ๐\+ฮปโโ๐โโ\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\\leq\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{v\}\+\\lambda\\\|\\bm\{v\}\\\|\_\{\*\}holds for any๐โ๐ฒ\\bm\{v\}\\in\\mathcal\{W\}, including๐=๐๐นโฒโ\(๐\)\\bm\{v\}=\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\. Like its prediction\-shift\-agnostic counterpart \(SPO\+\), the surrogate lossโSโPโrโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)is tight under perfect prediction, meaning thatโSโPโrโO\+โ\(๐,๐\)=max๐โ๐ฒRโฃโโ\(๐\)โก\{โฮปโโ๐โโ\}\+๐โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)=0\\ell\_\{SPrO\+\}\(\\bm\{c\},\\bm\{c\}\)=\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\}\\\{\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\+\\bm\{c\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)=0\. This property holds because๐โคโ๐๐นโฒโ\(๐\)=zRโฃโโ\(๐\)\\bm\{c\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)=z^\{R\*\}\(\\bm\{c\}\)by definition, and maximizing the nominal cost over๐ฒRโฃโโ\(๐\)\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)is equivalent to minimizing the dual norm regularizerโ๐โโ\\\|\\bm\{w\}\\\|\_\{\*\}\. We are now ready to confirm that, similar to SPO\+ under linear regression, the SPrO\+ loss is convex via the following theorem:
###### Theorem 3\.1\(Convexity of SPrO\+\)
The modelmax๐ฐโ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐ฐโ๐^โคโ๐ฐโฮปโโ๐ฐโโ\}\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}is equivalent to the convex optimization problem
min\\displaystyle\\minโkโ\[m\]\(gkโโฯk\)โ\(ฯk\)\\displaystyle\\sum\_\{k\\in\[m\]\}\(g^\{\*\}\_\{k\}\\pi\_\{k\}\)\(\\bm\{\\phi\}\_\{k\}\)s\.t\.โkโ\[m\]ฯk\+๐ฝ=๐โ๐^\\displaystyle\\sum\_\{k\\in\[m\]\}\\bm\{\\phi\}\_\{k\}\+\\bm\{\\theta\}=\\bm\{c\}\-\\hat\{\\bm\{c\}\}โ๐ฝโโคฮป\\displaystyle\\\|\\bm\{\\theta\}\\\|\\leq\\lambda๐
โฅ๐,๐ฝ,ฯโโd\.\\displaystyle\\bm\{\\pi\}\\geq\\bm\{0\},\\bm\{\\theta\},\\bm\{\\phi\}\\in\\mathbb\{R\}^\{d\}\.Therefore,โSโPโrโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)is convex in the prediction๐^\\hat\{\\bm\{c\}\}\. Furthermore, in the no\-prediction\-shift case \(ฮป=0\\lambda=0\), SPrO\+ provides a tighter upper approximation of the true SPO loss than standard SPO\+\.
SPrO\+ explicitly links the decision loss to the prediction error via the constraintโkโ\[m\]ฯk\+๐ฝ=๐โ๐^\\sum\_\{k\\in\[m\]\}\\bm\{\\phi\}\_\{k\}\+\\bm\{\\theta\}=\\bm\{c\}\-\\hat\{\\bm\{c\}\}\. In addition, it offers favorable tractability as it is jointly convex with respect to๐^\\hat\{\\bm\{c\}\}and decision variables, which means thatmin๐ฉโโdรpโก1nโโiโ\[n\]โSโPโrโO\+โ\(๐ฉโ๐i,๐i\)\\min\_\{\\bm\{B\}\\in\\mathbb\{R\}^\{d\\times p\}\}\\frac\{1\}\{n\}\\sum\_\{i\\in\[n\]\}\\ell\_\{SPrO\+\}\(\\bm\{B\}\\bm\{x\}\_\{i\},\\bm\{c\}\_\{i\}\)is a convex optimization problem solvable with off\-the\-shelf solvers\. Crucially, when the uncertainty budget drops to zero \(ฮป=0\\lambda=0\), the prediction\-shift\-aware surrogate yields a strictly tighter approximation of the true decision regret than the standard, prediction\-shift\-agnostic SPO\+ loss \(i\.e\.,โSโPโrโO\+โ\(๐^,๐\)โคโSโPโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\)\. This behavior stems from the fact that SPrO\+ evaluates the primal maximization directly over the optimal decision set๐ฒRโฃโโ\(๐^\)\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)rather than relaxing the domain to the entire feasible region๐ฒ\\mathcal\{W\}\. Under zero shift, this robust set collapses exactly to the non\-robust optimal oracle set,๐ฒRโฃโโ\(๐^\)=๐ฒโโ\(๐^\)โ๐ฒ\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)=\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\\subseteq\\mathcal\{W\}\. Because the maximization is restricted to this optimal solution set instead of the full space๐ฒ\\mathcal\{W\}, SPrO\+ eliminates conservative non\-optimal exploration, leading to a tighter, superior surrogate approximation while fully preserving tractability\. We will now see that in addition, SPrO\+ is well\-behaved\.
###### Theorem 3\.2\(Properties of SPrO\+\)
If the regularized optimal decision set๐ฒRโฃโโ\(โ
\)\\mathcal\{W\}^\{R\*\}\(\\cdot\)is always a singleton, then SPrO\+ loss has the following properties:
- Boundedness\.For any prediction shift within the uncertainty budget \(โ๐^โ๐โโคฮป\\\|\\hat\{\\bm\{c\}\}\-\\bm\{c\}\\\|\\leq\\lambda\), the loss is bounded by:โSโPโrโO\+โ\(๐^,๐\)โค2โฮปโโ๐๐นโฃโโ\(๐\)โโ\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq 2\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\.
- Lipschitz continuity\.Define the regularized objective functionhโ\(๐;๐\)โ๐โคโ๐\+ฮปโโ๐โโh\(\\bm\{w\};\\bm\{c\}\)\\coloneqq\\bm\{c\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\. Ifhhismm\-strongly convex, thenโSโPโrโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)is4โฮปm\\frac\{4\\lambda\}\{m\}\-Lipschitz continuous with respect to prediction๐^\\hat\{\\bm\{c\}\}\.
- ฯต\\bm\{\\epsilon\}\-insensitivity\.If a cost prediction deviates from the ground truth byฯต\\bm\{\\epsilon\}, i\.e\.๐^=๐โฯต\\hat\{\\bm\{c\}\}=\\bm\{c\}\-\\bm\{\\epsilon\}, such thatฯตโคโ๐๐นโฃโโ\(๐^\)\+ฮปโโ๐๐นโฃโโ\(๐\)โโโคฯตโคโ๐๐นโฃโโ\(๐\)\+ฮปโโ๐๐นโฃโโ\(๐^\)โโ\\bm\{\\epsilon\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\leq\\bm\{\\epsilon\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}, then SPrO\+ collapses, i\.e\.โSโPโrโO\+โ\(๐^,๐\)=0\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=0\.
Theorem[3\.2](https://arxiv.org/html/2607.21773#S3.Thmtheorem2)establishes the foundational properties for SPrO\+\. The underlying singleton assumption for๐ฒRโฃโโ\(โ
\)\\mathcal\{W\}^\{R\*\}\(\\cdot\)is mild and easily satisfied in practice, either through the strict convexity of the feasible region๐ฒ\\mathcal\{W\}or by ensuring the regularized objective itself is strongly convex, such as when deploying standardโ2\\ell\_\{2\}\-norm regularization\. The Boundedness property guarantees a provable safety ceiling for the surrogate loss that scales linearly with the uncertainty budgetฮป\\lambda\. Because๐ฒ\\mathcal\{W\}is compact, the dual normโ๐๐นโฃโโ\(๐\)โโ\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}remains finite, translating to a global stability safeguard\. For practitioners, this explicit bound implies that the SPrO\+ framework inherently dampens the impact of extreme outliers in the covariate space, ensuring that prediction errors under a bounded budget cannot cause the empirical loss to explode during training\.
Furthermore, the Lipschitz Continuity property ensures a well\-behaved optimization landscape\. Small updates to the predictive model parameters๐ฉ\\bm\{B\}translate to predictable, continuous variations in the downstream decision loss, facilitating training\. Interestingly, both the error bound and the Lipschitz constant scale directly withฮป\\lambda\. From a geometric perspective, this smoothness requires the regularized objective to be strongly convex, which holds under a Euclidean norm or, more generally, whenโฅโ
โฅโ\\\|\\cdot\\\|\_\{\*\}is anโq\\ell\_\{q\}norm \(1<qโค21<q\\leq 2\) and๐โ๐ฒ\\bm\{0\}\\notin\\mathcal\{W\}\(since a smooth primal norm yields a strongly convex dual norm\)\. Crucially, this continuity rectifies a notorious pathology in classic SPO\+\. When๐ฒ\\mathcal\{W\}is a polytope, classic SPO\+ decisions abruptly jump between extreme vertices under minor prediction perturbations, creating a discontinuous and volatile loss surface\. The SPrO\+ dual norm regularizer smooths out these vertex\-switching discontinuities, yielding stable solutions\.
Finally, theฯต\\bm\{\\epsilon\}\-insensitivity property draws parallel to theฯต\\epsilon\-insensitive hinge loss foundational to support vector regression\(Basaket al\.[2007](https://arxiv.org/html/2607.21773#bib.bib1)\)\. In the latter, prediction errors falling within anฯต\\epsilon\-tube are assigned zero penalty, ignoring benign noise\. We have an analogous property in SPrO\+, showing that if the prediction errorฯต\\bm\{\\epsilon\}is small enough that its decision cost difference fails to overcome the regularization buffer established byฮปโโ๐โโ\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}, the decision loss remains unaltered\. This property ensures that the downstream decision\-maker is insulated against non\-critical data contamination\.
### 3\.1Surrogate gap analysis
Does SPrO\+ match the behaviour of SPrO? A fundamental challenge in decision\-focused learning is that while the true robust decision regretโSโPโrโO\\ell\_\{SPrO\}represents the exact objective function, its inherent non\-convexity renders it impractical\. This necessitates our proposed convex surrogate, SPrO\+\. To mathematically justify this substitution, it is critical to analyze the alignment between these two loss functions\. We formalize this relationship by starting with a non\-asymptotic probabilistic analysis of their discrepancy under data uncertainty\. We model the inherent imprecision or disturbance stochastically, where realizations are governed by an underlying probability measure,\(๐i,๐i\)โผโ\(\\bm\{x\}\_\{i\},\\bm\{c\}\_\{i\}\)\\sim\\mathbb\{P\}\. Different from the literature, we seek to provide a probabilistic view of non\-asymptotic and asymptotic consistency between our surrogate and its true loss, offering the practitioners a quantifiable view on the likelihood that the surrogate is close to the true loss, showcasing its reliability\. We will utilize the framework of sub\-Gaussian randomness to portray decision errors\. We begin by recalling the foundational definitions of sub\-Gaussian random variables and their norms, which characterize random vectors whose tail distributions decay at least as quickly as a Gaussian profile\.
###### Definition 3\.3\(Sub\-Gaussian Random Vectors and Norms\)
The sub\-Gaussian norm of a scalar random variableuu, denoted byโuโฯ2\\\|u\\\|\_\{\\psi\_\{2\}\}, is defined as:
โuโฯ2=inf\{t\>0:๐ผโ\[expโก\(u2t2\)\]โค2\}\.\\displaystyle\\\|u\\\|\_\{\\psi\_\{2\}\}=\\inf\\left\\\{t\>0:\\mathbb\{E\}\\left\[\\exp\\left\(\\frac\{u^\{2\}\}\{t^\{2\}\}\\right\)\\right\]\\leq 2\\right\\\}\.This norm captures the growth rate of the random variableโs moments and the exponential decay of its tails\. Extending this to multivariate spaces, the sub\-Gaussian norm of a random vector๐ฎโโd\\bm\{u\}\\in\\mathbb\{R\}^\{d\}is defined as the supremum of the sub\-Gaussian norms of its one\-dimensional projections onto the unit sphere:
โ๐โฯ2=supโ๐โ2=1โโจ๐,๐โฉโฯ2\.\\displaystyle\\\|\\bm\{u\}\\\|\_\{\\psi\_\{2\}\}=\\sup\_\{\\\|\\bm\{v\}\\\|\_\{2\}=1\}\\\|\\langle\\bm\{u\},\\bm\{v\}\\rangle\\\|\_\{\\psi\_\{2\}\}\.A random vector๐ฎ\\bm\{u\}is classified as sub\-Gaussian if its corresponding norm is bounded, i\.e\.,โ๐ฎโฯ2<โ\\\|\\bm\{u\}\\\|\_\{\\psi\_\{2\}\}<\\infty\.
With these foundations in place, we present Theorem[3\.4](https://arxiv.org/html/2607.21773#S3.Thmtheorem4), which proves that the surrogate gap, i\.e\. the difference between SPrO and SPrO\+, concentrates tightly around zero with exponentially decaying probability\.
###### Theorem 3\.4\(Sub\-Gaussian concentration of the surrogate loss gap\)
Suppose that๐ฐ๐โฃโโ\(๐^\)=๐ฐ๐โฒโ\(๐\)\+๐ซ\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)=\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\bm\{\\Delta\}\. Assume that๐ซ\\bm\{\\Delta\}is a centered sub\-Gaussian random vector withโ๐ซโฯ2โคฮบ\\\|\\bm\{\\Delta\}\\\|\_\{\\psi\_\{2\}\}\\leq\\kappa\. Ifโ๐^โ2โคC^\\\|\\hat\{\\bm\{c\}\}\\\|\_\{2\}\\leq\\hat\{C\}, then there exists a non\-negative random variableTTsuch that:
๐ผโโ\[โSโPโrโO\+โ\(๐^,๐\)\]โค\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\\leq๐ผโโ\[โSโPโrโOโ\(๐^,๐\)\]\+T,\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\+T,whereTโฮปโโ๐ซโโโ๐ผโโ\[๐^\]โคโ๐ซT\\coloneqq\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\-\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\hat\{\\bm\{c\}\}\]^\{\\top\}\\bm\{\\Delta\}satisfies the concentration inequality:
โโ\(T\>t\)โคexpโก\{1โC0โt2ฮบ2โฮป2โC1\+C^2\},\\displaystyle\\mathbb\{Q\}\(T\>t\)\\leq\\exp\\left\\\{1\-\\frac\{C\_\{0\}t^\{2\}\}\{\\kappa^\{2\}\\lambda^\{2\}C\_\{1\}\+\\hat\{C\}^\{2\}\}\\right\\\},whereC0,C1C\_\{0\},C\_\{1\}are universal constants\. Furthermore, the expected gap is bounded by:
๐ผโโ\[T\]=Oโ\(ฮบโฮปโd\)\.\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{Q\}\}\[T\]=O\(\\kappa\\lambda\\sqrt\{d\}\)\.
Theorem[3\.4](https://arxiv.org/html/2607.21773#S3.Thmtheorem4)establishes that the probability of a significant gap between SPrO and SPrO\+ decays exponentially, providing a theoretical guarantee for the reliability of SPrO\+ as a surrogate\. The concentration inequality\(expโก\(โt2\)\)\(\\exp\(\-t^\{2\}\)\)ensures that the surrogate loss remains a high\-fidelity proxy for the true loss in the vast majority of realizations\. Furthermore, the bound quantifies a fundamental trade\-off: the risk of a performance gap increases as the decision environment becomes noisier \(ฮบ\\kappa\) or as the budget of uncertainty \(ฮป\\lambda\) is raised\. In practical terms, this implies that the price of robustness includes a potentially looser relationship between the surrogate and the true loss\. Finally, since our feasible region๐ฒ\\mathcal\{W\}is compact and the weight deviation๐ซ\\bm\{\\Delta\}is bounded, the sub\-Gaussian assumption in the theorem is statistically grounded, as all bounded distributions naturally satisfy sub\-Gaussian conditions\. In addition, from the proof of Theorem[3\.2](https://arxiv.org/html/2607.21773#S3.Thmtheorem2), we know that if๐โคโ๐\+ฮปโโ๐โโ\\bm\{l\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}ismm\-strongly convex, then its minimizer๐๐นโฃโโ\(๐\)\\bm\{w^\{R\*\}\}\(\\bm\{l\}\)\(or๐๐นโฒโ\(๐\)\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{l\}\)\) is2m\\frac\{2\}\{m\}\-Lipschitz continuous with respect to the input๐\\bm\{l\}\. From Proposition 1 inKatseliset al\.\([2021](https://arxiv.org/html/2607.21773#bib.bib3)\), we thus know that under additional mild conditions,๐๐นโฃโโ\(๐^\)โ๐ผ๐^โ\[๐๐นโฃโโ\(๐^\)\]\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\-\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\]and๐๐นโฒโ\(๐\)โ๐ผ๐โ\[๐๐นโฒโ\(๐\)\]\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\-\\mathbb\{E\}\_\{\\bm\{c\}\}\[\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\]follow sub\-Gaussian distributions, which implies that๐๐นโฃโโ\(๐^\)โ๐๐นโฒโ\(๐\)\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\-\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)is sub\-Gaussian, justifying our assumption in Theorem[3\.4](https://arxiv.org/html/2607.21773#S3.Thmtheorem4)\. Beyond the tail behavior, the expectation bound๐ผโโ\[T\]=Oโ\(ฮบโd\)\\mathbb\{E\}\_\{\\mathbb\{Q\}\}\[T\]=O\(\\kappa\\sqrt\{d\}\)characterizes the scalability of the SPrO\+ framework\. It reveals that the average gap between the surrogate and true loss grows only at a square\-root rate relative to the problem dimensiondd\. In the context of large\-scale problems, this sub\-linear growth suggests that SPrO\+ remains an effective approximation even as problem complexity increases\.
A fundamental question remains: does minimizing the surrogate loss, SPrO\+, also minimize SPrO? To establish this, we analyze the Fisher consistency\. Fisher consistency ensures that the surrogate optimization objective does not introduce bias relative to the true regret\. Formally stated, its definition is
###### Definition 3\.5\(Fisher consistency\)
A surrogate loss functionLSโ\(โ
,โ
\)L^\{S\}\(\\cdot,\\cdot\)isโ\\mathbb\{P\}\-Fisher consistent with respect to the true loss functionLโ\(โ
,โ
\)L\(\\cdot,\\cdot\)ifargโกmin๐โก๐ผโโ\[LSโ\(๐โ\(๐ฑ\),๐\)\]\\arg\\min\_\{\\bm\{f\}\}\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[L^\{S\}\(\\bm\{f\}\(\\bm\{x\}\),\\bm\{c\}\)\]also minimizes๐ผโโ\[Lโ\(๐โ\(๐ฑ\),๐\)\]\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[L\(\\bm\{f\}\(\\bm\{x\}\),\\bm\{c\}\)\]\.
The following corollary shows that our surrogate is Fisher consistent with very high probability\.
###### Corollary 3\.6\(Fisher consistency\)
If๐^\\hat\{\\bm\{c\}\}is centered underโ\\mathbb\{P\},โSโPโrโO\+\\ell\_\{SPrO\+\}isโ\\mathbb\{P\}\-Fisher consistent with a high minimum probability
1โฮฝ1โexpโก\{โ\(ฮฝ0โฯโ\(โฌ\)ฯ\)2\},\\displaystyle 1\-\\nu\_\{1\}\\exp\\left\\\{\-\\left\(\\nu\_\{0\}\\frac\{\\omega\(\\mathcal\{B\}\)\}\{\\phi\}\\right\)^\{2\}\\right\\\},whereฮฝ0,ฮฝ1\\nu\_\{0\},\\nu\_\{1\}are universal constants andฯ=supโ๐ฎโโค1โ๐ฎโ2\\phi=\\sup\_\{\\\|\\bm\{u\}\\\|\\leq 1\}\\\|\\bm\{u\}\\\|\_\{2\}\.
The probability boundary scales exponentially with the squared ratio of the Gaussian width to the maximum directional radius,\(ฯโ\(โฌ\)ฯ\)2\\left\(\\frac\{\\omega\(\\mathcal\{B\}\)\}\{\\phi\}\\right\)^\{2\}\. In high\-dimensional optimization problems, the squared Gaussian width typically scales linearly with the dimension \(Oโ\(d\)O\(d\)\)\. Consequently, as the dimensionality of the decision\-making problem expands, the tail probability of calibration failure shrinks exponentially toward zero\.
Note on generalizability: It is worth emphasizing that although our primary exposition focuses on linear regression as the hypothesis class, the theoretical results derived up to this point remain valid for a general function class๐โโ\\bm\{f\}\\in\\mathcal\{H\}\. Under a general predictive model๐โ\(๐\)\\bm\{f\}\(\\bm\{x\}\), the classical SPO\+ loss is formulated asโSโPโO\+โ\(๐โ\(๐\),๐\)=max๐โ๐ฒโก\{๐โคโ๐โฮฑโ๐โ\(๐\)โคโ๐\}\+ฮฑโ๐โ\(๐\)โคโ๐โโ\(๐\)โzโโ\(๐\)\\ell\_\{SPO\+\}\(\\bm\{f\}\(\\bm\{x\}\),\\bm\{c\}\)=\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\alpha\\bm\{f\}\(\\bm\{x\}\)^\{\\top\}\\bm\{w\}\\big\\\}\+\\alpha\\bm\{f\}\(\\bm\{x\}\)^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\), whereฮฑ\>0\\alpha\>0is a positive scaling parameter typically fixed at22to ensure convexity of the surrogate upper bound\. Crucially, our robust surrogateโSPrO\+โ\(๐โ\(๐\),๐\)\\ell\_\{\\text\{SPrO\+\}\}\(\\bm\{f\}\(\\bm\{x\}\),\\bm\{c\}\)does not require any specific calibration of such scaling parameters to preserve convexity\. It remains fundamentally convex with respect to๐โ\(๐\)\\bm\{f\}\(\\bm\{x\}\)while retaining desirable analytical properties \- namely, global boundedness, Lipschitz continuity,ฯต\\epsilon\-insensitivity, and high\-probability Fisher consistency \- under mild assumptions \(singleton optimal solution sets is a modelerโs choice as any smooth dual norm ensures strong convexity and sub\-Gaussian optimal solutions is achievable under mild assumptions, as shown byKatseliset al\.\([2021](https://arxiv.org/html/2607.21773#bib.bib3)\)\)\. It is, however, worth pointing out that our premise connecting covariate disturbance to prediction shift becomes questionable under arbitrary function classes\. The subsequent sections rely exclusively on our linear regression premise so as to offer a deep dive into necessary conditions for dominance\.
## 4Performance guarantees of downstream robustness
We now seek to understand the improvement in decision loss that a robust downstream decision\-maker achieves over a prediction\-shift\-agnostic one\. We begin by discussing the theoretical comparability of these two losses\. Recall SPO and SPrO definitionsโSโPโOโ\(๐^,๐\)=max๐โ๐ฒโโ\(๐^\)โก๐โคโ๐โzโโ\(๐\)\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{\*\}\(\\bm\{c\}\)andโSโPโrโOโ\(๐^,๐\)=\(max๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐โzRโฃโโ\(๐\)\)\+\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{R\*\}\(\\bm\{c\}\)\)\_\{\+\}\. While the two losses have different formulations and dissimilar baseline oracles, their true costs of downstream decision are comparable, as they both have the formmax๐โ๐ตโ\(๐^\)โก๐โคโ๐โ๐๐๐บ๐ผ๐
๐พ\\max\_\{\\bm\{w\}\\in\\mathcal\{Z\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\mathsf\{oracle\}, where๐ตโ\(๐^\)\\mathcal\{Z\}\(\\hat\{\\bm\{c\}\}\)depends on the risk attitude of the downstream decision maker \(deterministic vs robust\) and๐๐๐บ๐ผ๐
๐พ\\mathsf\{oracle\}is an input\. Irrespective of the oracle, a smallermax๐โ๐ตโ\(๐^\)โก๐โคโ๐\\max\_\{\\bm\{w\}\\in\\mathcal\{Z\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\- which we will term the true cost of downstream decisions made with predicted data \(CDP\) \- lowers the decision loss\.
By set inclusion๐ฒRโฃโโ\(๐\)โ๐ฒ\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\\subseteq\\mathcal\{W\}, we know that the robust oracle has higher value than the deterministic one, i\.e\.zRโฃโโ\(๐\)=max๐โ๐ฒRโฃโโ\(๐\)โก๐โคโ๐โฅmin๐โ๐ฒโก๐โคโ๐=zโโ\(๐\)z^\{R\*\}\(\\bm\{c\}\)=\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\\geq\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\bm\{c\}^\{\\top\}\\bm\{w\}=z^\{\*\}\(\\bm\{c\}\)\. The relationshipโSโPโrโOโ\(๐^,๐\)โฅโSโPโOโ\(๐^,๐\)\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\geq\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\), i\.e\.\(max๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐โzRโฃโโ\(๐\)\)\+โฅmax๐โ๐ฒโโ\(๐^\)โก๐โคโ๐โzโโ\(๐\)\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{R\*\}\(\\bm\{c\}\)\)\_\{\+\}\\geq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{\*\}\(\\bm\{c\}\), implies that either both SPO and SPrO losses are zero or their CDPs have relationshipmax๐โ๐ฒโโ\(๐^\)โก๐โคโ๐โmax๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐โคzโโ\(๐\)โzRโฃโโ\(๐\)โค0\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\\leq z^\{\*\}\(\\bm\{c\}\)\-z^\{R\*\}\(\\bm\{c\}\)\\leq 0\. Therefore, the CDP of SPrO exceeds that of SPO, also implying that robust downstream decisions will incur higher SPO loss\. It is thus clear thatโSโPโrโOโ\(๐^,๐\)โฅโSโPโOโ\(๐^,๐\)\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\geq\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)indicates poorer performance from downstream robustness\. As such, it is necessary to establish the decision loss relationshipโSโPโrโOโ\(๐^,๐\)โคโSโPโOโ\(๐^,๐\)\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)to ensure performance enhancement from SPrO, although it is not sufficient\. However, approaching the contrapositive argument, we know that ifmax๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐โคmax๐โ๐ฒโโ\(๐^\)โก๐โคโ๐\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}, thenโSโPโrโOโ\(๐^,๐\)โคโSโPโOโ\(๐^,๐\)\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\. This is because the following relationship holds
โSโPโrโOโ\(๐^,๐\)\\displaystyle\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)โคmax๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐\+max๐โ๐ฒโโ\(๐^\)โก๐โคโ๐โmax๐โ๐ฒโโ\(๐^\)โก๐โคโ๐โzโโ\(๐\)\\displaystyle\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\+\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{\*\}\(\\bm\{c\}\)=โSโPโOโ\(๐^,๐\)\+max๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐โmax๐โ๐ฒโโ\(๐^\)โก๐โคโ๐\\displaystyle=\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\+\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}=โSโPโOโ\(๐^,๐\)\+h๐ฒRโฃโโ\(๐^\)โ\(๐\)โh๐ฒโโ\(๐^\)โ\(๐\),\\displaystyle=\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\+h\_\{\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\(\\bm\{c\}\)\-h\_\{\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\(\\bm\{c\}\),where the last line is a support\-function formulation that allows better processing of expectations\. An immediate consequence arises under containment, i\.e\. if๐ฒRโฃโโ\(๐^\)โ๐ฒโโ\(๐^\)\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\\subseteq\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\), meaning any optimal decision under the robustified framework remains optimal for the nominal problem under predicted data, it follows thatโSโPโrโOโ\(๐^,๐\)โคโSโPโOโ\(๐^,๐\)\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)pointwise\. In terms of average performance behavior, we evaluate this relationship under a stochastic regime\. Taking expectations on both sides, if we let the true cost vector be distributed as a standard Gaussian,๐โผNโoโrโmโaโlโ\(๐,๐d\)\\bm\{c\}\\sim Normal\(\\bm\{0\},\\mathbb\{I\}\_\{d\}\), invoking the law of total expectation, we obtain
๐ผโ\[โSโPโrโOโ\(๐^,๐\)\]\\displaystyle\\mathbb\{E\}\[\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]โค๐ผโ\[โSโPโOโ\(๐^,๐\)\]\+๐ผโ\[h๐ฒRโฃโโ\(๐^\)โ\(๐\)\]โ๐ผโ\[h๐ฒโโ\(๐^\)โ\(๐\)\]\\displaystyle\\leq\\mathbb\{E\}\[\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\+\\mathbb\{E\}\[h\_\{\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\(\\bm\{c\}\)\]\-\\mathbb\{E\}\[h\_\{\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\(\\bm\{c\}\)\]=๐ผโ\[โSโPโOโ\(๐^,๐\)\]\+๐ผ๐^โ\[๐ผ๐โ\[h๐ฒRโฃโโ\(๐^\)โ\(๐\)\|๐^\]\]โ๐ผ๐^โ\[๐ผ๐โ\[h๐ฒโโ\(๐^\)โ\(๐\)\|๐^\]\]\\displaystyle=\\mathbb\{E\}\[\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\+\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\mathbb\{E\}\_\{\\bm\{c\}\}\[h\_\{\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\(\\bm\{c\}\)\|\\hat\{\\bm\{c\}\}\]\]\-\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\mathbb\{E\}\_\{\\bm\{c\}\}\[h\_\{\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\(\\bm\{c\}\)\|\\hat\{\\bm\{c\}\}\]\]=๐ผโ\[โSโPโOโ\(๐^,๐\)\]\+๐ผ๐^โ\[ฯโ\(๐ฒRโฃโโ\(๐^\)\)\]โ๐ผ๐^โ\[ฯโ\(๐ฒโโ\(๐^\)\)\],\\displaystyle=\\mathbb\{E\}\[\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\+\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\omega\(\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\)\]\-\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\omega\(\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\)\],showing that expected CDP become expectations of Gaussian widths\.
The regularized objective landscape provides structural advantages\. Suppose the feasible region๐ฒ\\mathcal\{W\}is a polytope and the dual norm regularizerโฅโ
โฅโ\\\|\\cdot\\\|\_\{\*\}is selected as the standard Euclideanโ2\\ell\_\{2\}\-norm\. The robustified objective function,๐โฆ๐^โคโ๐\+ฮปโโ๐โ2\\bm\{w\}\\mapsto\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{2\}, becomes strongly convex, guaranteeing that the regularized optimal decision set๐ฒRโฃโโ\(๐^\)\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)collapses to a singleton for any prediction๐^\\hat\{\\bm\{c\}\}\. Consequently, its Gaussian width vanishes, i\.e\.ฯโ\(๐ฒRโฃโโ\(๐^\)\)=0\\omega\(\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\)=0\. Conversely, the nominal optimal set๐ฒโโ\(๐^\)\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)frequently corresponds to a high\-dimensional face of the polytope๐ฒ\\mathcal\{W\}\(particularly when the prediction vector๐^\\hat\{\\bm\{c\}\}is orthogonal to a facet\)\. Under these conditions, the expected SPrO loss will be smaller than the expected SPO loss by at least the average Gaussian width of the unregularized optimal faces\. Crucially, we can explicitly quantify this performance gap by leveraging the Sudakov minoration theorem \(Lemma[7\.7](https://arxiv.org/html/2607.21773#S7.Thmtheorem7)in supplementary materials\), which bounds the Gaussian width from below, in the senseฯโ\(๐ฒโโ\(๐^\)\)โฅCโฯตโlogโก๐ฉโ\(๐ฒโโ\(๐^\),ฯต\)\\omega\(\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\)\\geq C\\epsilon\\sqrt\{\\log\\mathcal\{N\}\(\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\),\\epsilon\)\}, where๐ฉโ\(๐ฆ,ฯต\)\\mathcal\{N\}\(\\mathcal\{K\},\\epsilon\)is the minimum number of Euclidean balls of radiusฯต\\epsilonrequired to cover a compact set๐ฆ\\mathcal\{K\}\. This inequality reveals that the magnitude of SPrOโs improvement over SPO scales directly with the geometric complexity of the nominal optimal decision space\. More broadly, when neither๐ฒRโฃโโ\(๐^\)\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)nor๐ฒโโ\(๐^\)\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)reduce to singletons, a comparative analysis remains possible via the Sudakov\-Fernique inequality \(Lemma[7\.8](https://arxiv.org/html/2607.21773#S7.Thmtheorem8)in supplementary materials\), which states that if for any๐๐นโฃโ,๐โ\(๐\),๐๐นโฃโ,๐โ\(๐\)โ๐ฒRโฃโโ\(๐\)\\bm\{w^\{R\*,1\}\}\(\\bm\{c\}\),\\bm\{w^\{R\*,2\}\}\(\\bm\{c\}\)\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)and๐โ,1โ\(๐^\),๐โ,2โ\(๐\)โ๐ฒโโ\(๐\)\\bm\{w\}^\{\*,1\}\(\\hat\{\\bm\{c\}\}\),\\bm\{w\}^\{\*,2\}\(\\bm\{c\}\)\\in\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\), the conditionโ๐๐นโฃโ,๐โ\(๐\)โ๐๐นโฃโ,๐โ\(๐\)โ2โคโ๐โ,2โ\(๐\)โ๐โ,1โ\(๐\)โ2\\\|\\bm\{w^\{R\*,2\}\}\(\\bm\{c\}\)\-\\bm\{w^\{R\*,1\}\}\(\\bm\{c\}\)\\\|\_\{2\}\\leq\\\|\\bm\{w\}^\{\*,2\}\(\\bm\{c\}\)\-\\bm\{w\}^\{\*,1\}\(\\bm\{c\}\)\\\|\_\{2\}holds, thenฯโ\(๐ฒRโฃโโ\(๐^\)\)โคฯโ\(๐ฒโโ\(๐^\)\)\\omega\(\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\)\\leq\\omega\(\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\)\.
Of course, what we actually solve is the convex surrogate\. While the structural properties of SPrO\+ establish its stability and insensitivity to localized prediction errors \(Theorem[3\.2](https://arxiv.org/html/2607.21773#S3.Thmtheorem2)\), we still need to understand if this robustified surrogate retains a systematic performance edge over the nominal SPO\+ surrogate under uncertainty\. Because SPrO\+ uses regularization to guard against worst\-case covariate shifts, it is crucial to guarantee that this conservatism does not inadvertently degrade performance\. However, proving necessary dominance conditions for surrogates can be hard as the SPO\+ and SPrO\+ are considerably different\. We thus derive high\-probability, rather than absolute, necessary conditions for dominance\. We first establish the following relationship between SPrO\+ and SPrO:
โSโPโrโO\+โ\(๐^,๐\)=\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}=\\displaystyle=\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}=\\displaystyle=\(max๐โ๐ฒRโฃโโ\(๐^\)โก๐โคโ๐โโ๐^โคโ๐๐นโฃโโ\(๐^\)โฮปโโ๐๐นโฃโโ\(๐^\)โโ\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโ=Aโ\(๐^,๐\)โฃโฅ0โzRโฃโโ\(๐\)\)\+\\displaystyle\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\\underbrace\{\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\-\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\}\_\{=A\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\geq 0\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}โค\\displaystyle\\leqโSโPโrโOโ\(๐^,๐\)\+Aโ\(๐^,๐\),\\displaystyle\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\+A\(\\hat\{\\bm\{c\}\},\\bm\{c\}\),where the inequality is by subadditivity of the positive\-part function\. Repeating the analysis for SPO and SPO\+, we obtain the bound
โSโPโO\+โ\(๐^,๐\)=\\displaystyle\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=max๐โ๐ฒโก\{๐โคโ๐โ๐^โคโ๐\}\+๐^โคโ๐โโ\(๐\)โzโโ\(๐\)\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\)โฅ\\displaystyle\\geqmax๐โ๐ฒโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐\}\+๐^โคโ๐โโ\(๐\)โzโโ\(๐\)\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\)=\\displaystyle=max๐โ๐ฒโโ\(๐^\)โก๐โคโ๐โโ๐^โคโ๐โโ\(๐^\)\+๐^โคโ๐โโ\(๐\)โ=Bโ\(๐^,๐\)โฃโฅ0โzโโ\(๐\)=โSโPโOโ\(๐^,๐\)\+Bโ\(๐^,๐\),\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\\underbrace\{\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\hat\{\\bm\{c\}\}\)\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\}\_\{=B\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\geq 0\}\-z^\{\*\}\(\\bm\{c\}\)=\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\+B\(\\hat\{\\bm\{c\}\},\\bm\{c\}\),with the inequality coming from the set inclusion๐ฒโโ\(๐^\)โ๐ฒ\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\\subseteq\\mathcal\{W\}\. We therefore know that the losses and their surrogates are linked via the relationship
โSโPโrโO\+โ\(๐^,๐\)โโSโPโO\+โ\(๐^,๐\)โคโSโPโrโOโ\(๐^,๐\)\+Aโ\(๐^,๐\)โโSโPโOโ\(๐^,๐\)โBโ\(๐^,๐\)\.\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\+A\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-B\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\.For dominance \(measured by the CDP\) to be possible, the theoretical necessary condition must be met, i\.e\.โSโPโrโOโ\(๐^,๐\)โคโSโPโOโ\(๐^,๐\)\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\. In addition, Theorem[3\.4](https://arxiv.org/html/2607.21773#S3.Thmtheorem4), which establishes the sub\-Gaussian behaviour of the surrogate loss gap, hints thatAโ\(๐^,๐\)A\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)is highly likely to be small, which means that the necessary dominance conditionโSโPโrโOโ\(๐^,๐\)โคโSโPโOโ\(๐^,๐\)\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)impliesโSโPโrโO\+โ\(๐^,๐\)โคโSโPโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)with high likelihood\. As such,โSโPโrโO\+โ\(๐^,๐\)โคโSโPโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)is a high\-likelihood indicator of lower CDP\.
We analyze the conditions under which the robust surrogate is lower than the nominal surrogate\. We first establish a deterministic, pointwise result in Theorem[4\.1](https://arxiv.org/html/2607.21773#S4.Thmtheorem1)by bounding the budget of uncertainty relative to the optimal nominal objective\. Recognizing that exact pointwise conditions can be overly restrictive in stochastic environments, we subsequently relax this in Corollary[4\.2](https://arxiv.org/html/2607.21773#S4.Thmtheorem2)\. By treating the structural discrepancies between robust and nominal decisions as sub\-Gaussian random vectors, we prove thatโSโPโrโO\+\\ell\_\{SPrO\+\}achieves a lower expected loss than its nominal counterpart with a probability approaching certainty as the geometric complexity of the decision space scales\.
###### Theorem 4\.1\(Pointwise analysis of SPrO\+ vs SPO\+\)
Suppose thatโ๐^โโคC^\\\|\\hat\{\\bm\{c\}\}\\\|\\leq\\hat\{C\}and there exists aฮป\>0\\lambda\>0such thatโ๐^โ๐โโคฮป\\\|\\hat\{\\bm\{c\}\}\-\\bm\{c\}\\\|\\leq\\lambda\. If thisฮป\\lambdasatisfies
ฮปโคzโโ\(๐\)โC^โโ๐๐นโฒโ\(๐\)โโโ๐โโ\(๐\)โโ\+โ๐๐นโฒโ\(๐\)โโ,\\displaystyle\\lambda\\leq\\frac\{z^\{\*\}\(\\bm\{c\}\)\-\\hat\{C\}\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\}\{\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\+\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\},assuming the right\-hand side ratio is positive, then SPO\+ exceeds SPrO\+ pointwise, i\.e\.
โSโPโrโO\+โ\(๐^,๐\)โคโSโPโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\.
###### Corollary 4\.2\(Stochastic analysis of SPrO\+ vs SPO\+\)
Suppose that๐ฐ๐โฒโ\(๐\)=๐ฐโโ\(๐\)\+๐ซโ\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)=\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\+\\bm\{\\Delta^\{\*\}\}, where๐ซโ\\bm\{\\Delta^\{\*\}\}is a sub\-Gaussian random vector withโ๐ซโโฯ2โคฮบ\\\|\\bm\{\\Delta^\{\*\}\}\\\|\_\{\\psi\_\{2\}\}\\leq\\kappa\. If the problem structure induces a a positive gapW\>0W\>0in
๐ผโ\[โ๐๐นโฒโ\(๐\)โโ\]โค๐ผโ\[โ๐๐นโฒโ\(๐^\)โโ\]โW,\\displaystyle\\mathbb\{E\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]\\leq\\mathbb\{E\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\]\-W,then the expected robust surrogate loss remains lower than the nominal surrogate loss,๐ผโ\[โSโPโrโO\+โ\(๐^,๐\)\]โค๐ผโ\[โSโPโO\+โ\(๐^,๐\)\]\\mathbb\{E\}\[\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\\leq\\mathbb\{E\}\[\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\], with minimum probability
1โC^โฮท0โฮบโฯโ\(โฌ\)ฮปโW,\\displaystyle 1\-\\frac\{\\hat\{C\}\\eta\_\{0\}\\kappa\\omega\(\\mathcal\{B\}\)\}\{\\lambda W\},whereฮท0\\eta\_\{0\}is an absolute constant\.
Theorem[4\.1](https://arxiv.org/html/2607.21773#S4.Thmtheorem1)establishes an explicit safety threshold for the budget of uncertaintyฮป\\lambda\. The prediction\-free upper boundฮปโคzโโ\(๐\)โC^โโ๐๐นโฒโ\(๐\)โโโ๐โโ\(๐\)โโ\+โ๐๐นโฒโ\(๐\)โโ\\lambda\\leq\\frac\{z^\{\*\}\(\\bm\{c\}\)\-\\hat\{C\}\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\}\{\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\+\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\}reveals a fundamental trade\-off\. To prevent SPO\+ dominance pointwise, the nominal optimal costzโโ\(๐\)z^\{\*\}\(\\bm\{c\}\)must be large enough to absorb the magnitude of decision regularization scaled by the maximum prediction sizeC^\\hat\{C\}\. Intuitively, when the true underlying optimization problem has a high optimal objective value and โsmall" optimal decisions \(as measured by the dual norm\), one can easily find a budget of uncertainty that will result in pointwise lower SPrO\+ loss\. In Corollary[4\.2](https://arxiv.org/html/2607.21773#S4.Thmtheorem2), the termWWacts as a safety buffer, guaranteeing that optimizing with predictions \(๐^\\hat\{\\bm\{c\}\}\) inflates the regularizer by at least a baseline margin ofWWcompared to optimizing with the true realized data \(๐\\bm\{c\}\)\. Because the nominal SPO\+ model does not penalize this dual norm inflation, it can potentially make high\-magnitude, risky decisions\. SPrO\+, via itsฮปโโ๐โโ\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}penalty, anticipates this inflation and penalizes it\. The minimum probability bound1โC^โฮท0โฮบโฯโ\(โฌ\)ฮปโW1\-\\frac\{\\hat\{C\}\\eta\_\{0\}\\kappa\\omega\(\\mathcal\{B\}\)\}\{\\lambda W\}is highly interpretable\. This fraction captures a trade\-off between model complexity and the level of conservatism of the decision maker\. The denominator shows that as the budget of uncertainty grows, the denominator increases, pushing the overall probability higher\. This proves that if a manager faces a highly volatile environment, they can actively guarantee stochastically lower SPrO\+ over SPO\+\. The numerator groups together factors influencing the model complexity, such as the data prediction scale \(C^\\hat\{C\}\), the decision scale \(measured by the sub\-Gaussian tail noiseฮบ\\kappa\), the dimension of decisionsฯโ\(โฌ\)\\omega\(\\mathcal\{B\}\), which scales withOโ\(d\)O\(\\sqrt\{d\}\), i\.e\. the Gaussian width scales up with the number of decisions\. As a problem gets larger \(higher dimensionality\) and decisions become more unpredictable \(higher sub\-Gaussian noise\), the probability of outperforming the nominal model drops\. To maintain the same performance guarantee in large\-scale systems, the regularization penaltyฮป\\lambdamust scale accordingly with the model complexity\.
## 5Upstream robustification vs SPrO
A natural alternative to robustifying the downstream decision\-making process \(as done in SPrO\) is to robustify the upstream estimation process itself\. This approach shifts the burden of conservatism from the optimizer to the predictor\. While robust regression paradigms are well\-studied in isolation, their structural interactions with downstream optimization instances remain largely unquantified within the SPO literature\. This section contextualizes this fundamental modeling choice: is it more advantageous to hedge against uncertainty in the prediction space or directly within the decision space? By mapping covariate disturbances through the lens of regularized predictions, we formalize the resulting decision loss and provide structural conditions under which SPrO achieves superior expected performance\.
We start by characterizing each element of the regression vector using๐ฉ=\(๐ท1,โฆ,๐ทd\)โค\\bm\{B\}=\(\\bm\{\\beta\}\_\{1\},\\dots,\\bm\{\\beta\}\_\{d\}\)^\{\\top\}, where๐ทj=\(ฮฒ1โj,โฆ,ฮฒpโj\)\\bm\{\\beta\}\_\{j\}=\(\\beta\_\{1j\},\\dots,\\beta\_\{pj\}\), which leads to regression modelc^j=๐ทjโคโ๐\\hat\{c\}\_\{j\}=\\bm\{\\beta\}\_\{j\}^\{\\top\}\\bm\{x\},โjโ\[d\]\\forall j\\in\[d\]\. In classical regression, one seeks coefficient values that minimize the empirical residual sum of squares under a squaredโ2\\ell\_\{2\}\-norm\. When introducing norm\-bounded covariate uncertainty into the estimation stage,Xuet al\.\([2008](https://arxiv.org/html/2607.21773#bib.bib45)\)\(Theorem 2\) demonstrates that the robust counterpart is equivalent to finding the minimizer๐ทj\\bm\{\\beta\}\_\{j\}that minimizes anโ1\\ell\_\{1\}\-regularized loss\. This Lasso regularization yields sparse regression coefficients, which effectively nullifies the impact of covariate disturbances on the estimation loss\.
While one could theoretically enforce coefficient sparsity in standard SPO or SPrO by appending a regularization term directly to the decision losses \(or their surrogates\), doing so decouples the regularization from the underlying covariate disturbance\. In contrast, under the robust regression framework ofXuet al\.\([2008](https://arxiv.org/html/2607.21773#bib.bib45)\),ฮป\\lambdarepresents an uncertainty budget on the covariate disturbance that bounds an arbitrary norm \- a feature structurally analogous to our SPrO framework\. For upstream robustification, one would ideally solve the minimax formulationmin๐ฉโโdรpโกmaxโ๐นโโคฮปโก๐ผโโ\[โSโPโOโ\(๐ฉโ\(๐\+๐น\),๐\)\]\\min\_\{\\bm\{B\}\\in\\mathbb\{R\}^\{d\\times p\}\}\\max\_\{\\\|\\bm\{\\delta\}\\\|\\leq\\lambda\}\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPO\}\(\\bm\{B\}\(\\bm\{x\}\+\\bm\{\\delta\}\),\\bm\{c\}\)\]\. However, this formulation introduces non\-convexities that render the problem not directly solvable and obscure direct structural comparisons with standard SPO\. This contrasts sharply with our SPrO framework, which maintains a clear connection to SPO while shifting the conservatism entirely to a regularized downstream decision problem\. Applying a parallel rationale to the upstream side, if the downstream problem is a cost minimization problem, a robust upstream predictor must anticipate the worst\-case \(highest possible\) nominal cost vector to prevent underprepared decision\-making\. To ensure conservatism against cost inflation under covariate shifts, the predictor estimates the worst\-case upper bound of the cost vector\. We formalize this approach via the Smart robust\-Predict\-then\-Optimize \(SrPO\) framework, which serves as an upstream proxy for prediction robustness\. Specifically, SrPO constructs a worst\-case predicted cost model for each component, defined asc^jR=max๐นโ๐ฐฮปโก๐ทjโคโ\(๐\+๐น\)=๐ทjโคโ๐\+ฮปโโ๐ทjโโ\\hat\{c\}^\{R\}\_\{j\}=\\max\_\{\\bm\{\\delta\}\\in\\mathcal\{U\}\_\{\\lambda\}\}\\bm\{\\beta\}\_\{j\}^\{\\top\}\(\\bm\{x\}\+\\bm\{\\delta\}\)=\\bm\{\\beta\}\_\{j\}^\{\\top\}\\bm\{x\}\+\\lambda\\\|\\bm\{\\beta\}\_\{j\}\\\|\_\{\*\}\. This yields the SrPO loss function
โSโrโPโOโ\(๐^๐น,๐\)=โSโrโPโOโ\(๐ฉโ๐\+ฮปโ๐ฒโ\(๐ฉ\),๐\)โmax๐โ๐ฒโโ\(๐ฉโ๐\+ฮปโ๐ฒโ\(๐ฉ\)\)โก๐โคโ๐โzโโ\(๐\),\\displaystyle\\ell\_\{SrPO\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)=\\ell\_\{SrPO\}\(\\bm\{B\}\\bm\{x\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\),\\bm\{c\}\)\\coloneqq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\bm\{B\}\\bm\{x\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)\}\\bm\{c\}^\{\\top\}\\bm\{w\}\-z^\{\*\}\(\\bm\{c\}\),where๐ฒโ\(๐ฉ\)=\(โ๐ทjโโ\)jโ\[d\]\\bm\{\\Lambda\}\(\\bm\{B\}\)=\(\\\|\\bm\{\\beta\}\_\{j\}\\\|\_\{\*\}\)\_\{j\\in\[d\]\}is add\-dimensional column vector of dual norms\. Consequently, SrPO can be interpreted as standard SPO evaluated under worst\-case predicted costs, where the upstream prediction shift manifests as a coefficient\-dependent regularizer scaled by the uncertainty budget\. The ultimate objective under SrPO is to identify an empirical coefficient matrix that minimizes decision regret under these worst\-case predictions\.
Following similar analysis to the previous section, if the true cost vector be distributed as a standard Gaussian,๐โผ๐ฉโ\(๐,๐d\)\\bm\{c\}\\sim\\mathcal\{N\}\(\\bm\{0\},\\mathbb\{I\}\_\{d\}\), we can provide a gap SPrO and SrPO as
๐ผโ\[โSโPโrโOโ\(๐^,๐\)\]\\displaystyle\\mathbb\{E\}\[\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]โค๐ผโ\[โSโrโPโOโ\(๐^๐น,๐\)\]\+๐ผ๐^โ\[ฯโ\(๐ฒRโฃโโ\(๐^\)\)\]โ๐ผ๐^โ\[ฯโ\(๐ฒโโ\(๐^๐น\)\)\]\.\\displaystyle\\leq\\mathbb\{E\}\[\\ell\_\{SrPO\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\]\+\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\omega\(\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\)\]\-\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\omega\(\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\}\)\)\]\.The comparative performance between these frameworks, as well as their surrogate counterparts SPrO\+ and SrPO\+, depends fundamentally on the minimum difference in their regularization profiles, which we formally define below\.
###### Definition 5\.1\(Regularization bias differential\)
The regularization bias differential between sets๐\\mathcal\{A\}andโฌ\\mathcal\{B\}is defined as
๐ขโ\(๐,โฌ\)โmin๐โ๐โก\{๐ฒโคโ\(๐ฉ\)โ๐โโ๐โโ\}โmax๐โโฌโก\{๐ฒโคโ\(๐ฉ\)โ๐โโ๐โโ\}\.\\displaystyle\\mathcal\{G\}\(\\mathcal\{A\},\\mathcal\{B\}\)\\coloneqq\\min\_\{\\bm\{a\}\\in\\mathcal\{A\}\}\\\{\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{a\}\-\\\|\\bm\{a\}\\\|\_\{\*\}\\\}\-\\max\_\{\\bm\{b\}\\in\\mathcal\{B\}\}\\\{\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{b\}\-\\\|\\bm\{b\}\\\|\_\{\*\}\\\}\.It is clear that๐ขโ\(๐,โฌ\)\\mathcal\{G\}\(\\mathcal\{A\},\\mathcal\{B\}\)is a jointly monotonic measure, in the sense that if๐โ๐โฒ\\mathcal\{A\}\\subseteq\\mathcal\{A\}^\{\\prime\}andโฌโโฌโฒ\\mathcal\{B\}\\subseteq\\mathcal\{B\}^\{\\prime\}, then๐ขโ\(๐โฒ,โฌโฒ\)โค๐ขโ\(๐,โฌ\)\\mathcal\{G\}\(\\mathcal\{A\}^\{\\prime\},\\mathcal\{B\}^\{\\prime\}\)\\leq\\mathcal\{G\}\(\\mathcal\{A\},\\mathcal\{B\}\)\. In addition, if๐=โฌ=\{๐ฐ\}\\mathcal\{A\}=\\mathcal\{B\}=\\\{\\bm\{w\}\\\}\(both sets are equal and singletons\), then๐ขโ\(๐,โฌ\)=0\\mathcal\{G\}\(\\mathcal\{A\},\\mathcal\{B\}\)=0\.
The term๐ขโ\(๐,โฌ\)\\mathcal\{G\}\(\\mathcal\{A\},\\mathcal\{B\}\)measures the regularization mismatch between the robustifications of upstream prediction and downstream decision\. Specifically, the term๐ฒโ\(๐ฉ\)โคโ๐โโ๐โโ\\bm\{\\Lambda\}\(\\bm\{B\}\)^\{\\top\}\\bm\{w\}\-\\\|\\bm\{w\}\\\|\_\{\*\}represents the net regularized weight of a decision vector๐\\bm\{w\}\. It balances the predictive sensitivity penalty๐ฒโ\(๐ฉ\)โคโ๐\\bm\{\\Lambda\}\(\\bm\{B\}\)^\{\\top\}\\bm\{w\}against the downstream decision\-space robustness penaltyโ๐โโ\\\|\\bm\{w\}\\\|\_\{\*\}\. Therefore,๐ขโ\(๐,โฌ\)\\mathcal\{G\}\(\\mathcal\{A\},\\mathcal\{B\}\)quantifies the minimum possible net weight in set๐\\mathcal\{A\}minus the maximum possible net weight in setโฌ\\mathcal\{B\}\. If๐=๐ฒโโ\(๐^๐น\)\\mathcal\{A\}=\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\}\)\(the unregularized decisions under worst\-case predictions\) andโฌ=๐ฒRโฃโโ\(๐^\)\\mathcal\{B\}=\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\(the regularized decisions under nominal predictions\), a large positive๐ขโ\(๐,โฌ\)\\mathcal\{G\}\(\\mathcal\{A\},\\mathcal\{B\}\)implies that any decision forced by robustifying against worst\-case predictions is fundamentally more conservative \(or restricted\) than even the most conservative decision available when robustifying the decision space directly\.
###### Theorem 5\.2\(Theoretical analysis of SPrO vs SrPO\)
Let๐ฒSโฃโโ\(๐\)=argโกmin๐ฐโ๐ฒโก\{\(๐\+ฮปโ๐ฒโ\(๐\)\)โคโ๐ฐ\}\\mathcal\{W\}^\{S\*\}\(\\bm\{c\}\)=\\arg\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\\{\(\\bm\{c\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w\}\\\}\. Also, assume that the true cost vector originated from a standard Gaussian distribution,๐ฉโ\(๐,๐d\)\\mathcal\{N\}\(\\bm\{0\},\\mathbb\{I\}\_\{d\}\)\. If
๐ขโ\(๐ฒSโฃโโ\(๐\),๐ฒRโฃโโ\(๐\)\)โฅโ๐โ2\+ฮปโโ๐ฒโ\(๐ฉ\)โ2ฮปโ๐โ\(๐ฒRโฃโโ\(๐\)\),\\mathcal\{G\}\(\\mathcal\{W\}^\{S\*\}\(\\bm\{c\}\),\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\)\\geq\\frac\{\\\|\\bm\{c\}\\\|\_\{2\}\+\\lambda\\\|\\bm\{\\Lambda\}\(\\bm\{B\}\)\\\|\_\{2\}\}\{\\lambda\}\\mathcal\{D\}\(\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\),then
๐ผโ\[โSโPโrโOโ\(๐^,๐\)\]โค๐ผโ\[โSโrโPโOโ\(๐^๐น,๐\)\]\.\\displaystyle\\mathbb\{E\}\[\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\\leq\\mathbb\{E\}\[\\ell\_\{SrPO\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\]\.
Theorem[5\.2](https://arxiv.org/html/2607.21773#S5.Thmtheorem2)establishes a fundamental geometric condition under which robustifying the decision space \(SPrO\) can potentially yield structurally superior performance over robustifying against worst\-case predictions \(SrPO\)\. The core mechanism driving this dominance is the interplay between the regularization bias differential๐ขโ\(โ
,โ
\)\\mathcal\{G\}\(\\cdot,\\cdot\)and the diameter of the robust decision space๐โ\(โ
\)\\mathcal\{D\}\(\\cdot\)\. The condition stipulates that if the normalized regularization bias differential exceeds the diameter of the robustified decision set, the Gaussian width of the SPrO decision space is smaller than that of SrPO\. In operational terms, the regularization bias differential๐ขโ\(โ
,โ
\)\\mathcal\{G\}\(\\cdot,\\cdot\)measures the alignment between the predictorโs dual norm penalties and the optimizerโs dual norm regularizer\. If๐ฒ\\mathcal\{W\}is a polytope, when this gap is sufficiently large relative to the scale of the true cost vector\(โ๐โ2\+ฮปโโ๐ฒโ\(๐ฉ\)โ2\)\(\\\|\\bm\{c\}\\\|\_\{2\}\+\\lambda\\\|\\bm\{\\Lambda\}\(\\bm\{B\}\)\\\|\_\{2\}\), it implies that robustifying the prediction space forces the downstream unregularized optimizer onto highly unstable, high\-dimensional faces\. Conversely, SPrO directly smooths the downstream objective landscape, shrinking the regularized optimal decision set toward a lower\-dimensional face or a singleton\. For practitioners, this result offers a clear guideline\. Robustifying against worst\-case predictions \(SrPO\) does not inherently protect the downstream optimization model\. Because SrPO alters the nominal input๐^๐น\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\}without modifying the optimization model, it remains highly sensitive to jumps across corner points if๐ฒ\\mathcal\{W\}is a polyhedron, for instance\. The threshold condition highlights that the superiority of SPrO is amplified when the uncertainty budgetฮป\\lambdais large relative to the nominal dataโ๐โ2\\\|\\bm\{c\}\\\|\_\{2\}\. In highly volatile environments where data are highly corrupted, direct intervention in the decision space via SPrO acts as a more effective decision loss minimizer compared to robust predictions\.
To compare surrogates, we begin by establishing a prediction\-free lower bound on the upstream\-robust surrogate loss \(โSโrโPโO\+\\ell\_\{SrPO\+\}\), which tracks how much decision loss an observer must absorb when relying exclusively on worst\-case inputs\.
###### Lemma 5\.3\(Prediction\-free lower bound on SrPO\+\)
โSโrโPโO\+โ\(๐^๐น,๐\)โฅzRโฃโโ\(๐\)\+zโโ\(๐\)\+ฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\\ell\_\{SrPO\+\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\\geq z^\{R\*\}\(\\bm\{c\}\)\+z^\{\*\}\(\\bm\{c\}\)\+\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\.
###### Theorem 5\.4\(Stochastic analysis of SPrO\+ vs SrPO\+\)
Supposeโ๐^โโคC^\\\|\\hat\{\\bm\{c\}\}\\\|\\leq\\hat\{C\}andzยฏโโ\(๐\)=max๐ฐโ๐ฒโก\{๐โคโ๐ฐ\}\\bar\{z\}^\{\*\}\(\\bm\{c\}\)=\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\\\}\. Ifzโโ\(โ
\)โฅ0z^\{\*\}\(\\cdot\)\\geq 0and the following condition is satisfied:
\(C^\+ฮป\)โ๐ผโ\[โ๐๐นโฒโ\(๐\)โโ\]โคฮปโ๐ผโ\[โ๐๐นโฒโ\(๐^\)โโ\]\+๐ผโ\[zRโฃโโ\(๐\)\+zโโ\(๐\)\+ฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\],\\displaystyle\(\\hat\{C\}\+\\lambda\)\\mathbb\{E\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]\\leq\\lambda\\mathbb\{E\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\]\+\\mathbb\{E\}\[z^\{R\*\}\(\\bm\{c\}\)\+z^\{\*\}\(\\bm\{c\}\)\+\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\],then downstream robustification yields lower expected surrogate loss than upstream robustification:
๐ผโ\[โSโPโrโO\+โ\(๐^,๐\)\]โค๐ผโ\[โSโrโPโO\+โ\(๐^๐น,๐\)\]\.\\mathbb\{E\}\[\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\\leq\\mathbb\{E\}\[\\ell\_\{SrPO\+\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\]\.
Theorem[5\.4](https://arxiv.org/html/2607.21773#S5.Thmtheorem4)significantly relaxes and generalizes the conditions under which a decision\-maker should favor SPrO\+ over alternative frameworks, notably improving upon the tight boundaries specified in Corollary[4\.2](https://arxiv.org/html/2607.21773#S4.Thmtheorem2)\. In the latter, stochastic dominance of SPrO\+ over the nominal SPO\+ model was restricted by a sub\-Gaussian assumption on decisions\. Theorem[5\.4](https://arxiv.org/html/2607.21773#S5.Thmtheorem4)completely bypasses sub\-Gaussian parameter dependencies, making the dominance condition applicable to any decision behaviour\. It also seems to show a looser requirement in the relationship between๐ผโ\[โ๐๐นโฒโ\(๐\)โโ\]\\mathbb\{E\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]and๐ผโ\[โ๐๐นโฒโ\(๐^\)โโ\]\\mathbb\{E\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\]\.
## 6Numerical experiments \- a network flow problem
To evaluate the empirical performance of the proposed Smart Predict\-then\-Robustly\-Optimize \(SPrO\+\) framework, we consider a continuous minimum\-cost network flow problem\. We benchmark SPrO\+ against two primary baselines: the classic Smart Predict\-then\-Optimize surrogate \(SPO\+\) and its worst\-case\-estimation\-robust variant \(SrPO\+\)\. LetG=\(๐ฑ,โฐ\)G=\(\\mathcal\{V\},\\mathcal\{E\}\)be a directed network graph, where๐ฑ\\mathcal\{V\}represents the set of nodes andโฐ\\mathcal\{E\}represents the set of directed links\. The model is
zโโ\(๐\)=min\\displaystyle z^\{\*\}\(\\bm\{c\}\)=\\min\\,โeโโฐceโwe\\displaystyle\\sum\_\{e\\in\\mathcal\{E\}\}c\_\{e\}w\_\{e\}s\.t\.โeโฮดโโ\(v\)weโโeโฮด\+โ\(v\)we=bvโvโ๐ฑ\\displaystyle\\sum\_\{e\\in\\delta^\{\-\}\(v\)\}w\_\{e\}\-\\sum\_\{e\\in\\delta^\{\+\}\(v\)\}w\_\{e\}=b\_\{v\}\\quad\\forall v\\in\\mathcal\{V\}0โคweโคueโeโโฐ,\\displaystyle 0\\leq w\_\{e\}\\leq u\_\{e\}\\quad\\forall e\\in\\mathcal\{E\},whereฮด\+โ\(v\)\\delta^\{\+\}\(v\)andฮดโโ\(v\)\\delta^\{\-\}\(v\)are the sets of outgoing and incoming edges to nodevv, respectively\. The flow balance requirement at nodevvisbvb\_\{v\}, defined explicitly as:
bv=\{โD,ifโv=vsourceD,ifโv=vsink0,otherwise\.b\_\{v\}=\\begin\{cases\}\-D,&\\text\{if \}v=v\_\{\\text\{source\}\}\\\\ D,&\\text\{if \}v=v\_\{\\text\{sink\}\}\\\\ 0,&\\text\{otherwise\}\\end\{cases\}\.The termueu\_\{e\}is the link capacity\. The capacitiesueu\_\{e\}for each arc are drawn independently from a uniform distribution,ueโผ๐ฐโ\(5,20\)u\_\{e\}\\sim\\mathcal\{U\}\(5,20\), and the total network demand is fixed atD=10D=10\. In this contextual optimization setup, the true cost vector๐\\bm\{c\}is driven by exogenous features\. The nominal costc^e\\hat\{c\}\_\{e\}for each arceโโฐe\\in\\mathcal\{E\}is modeled as a linear function of55covariates:
c^e=ฮฒ0โe\+โpโ\[5\]ฮฒpโeโxp\.\\hat\{c\}\_\{e\}=\\beta\_\{0e\}\+\\sum\_\{p\\in\[5\]\}\\beta\_\{pe\}x\_\{p\}\.Our base experimental instance is constructed on a random graph containing\|๐ฑ\|=10\|\\mathcal\{V\}\|=10nodes and\|โฐ\|=30\|\\mathcal\{E\}\|=30arcs, utilizing a training dataset ofn=100n=100synthetic observations\. The baseline, non\-contaminated data\-generating process proceeds as follows\. First, underlying ground\-truth parameters are sampled via๐ยฏโผ๐ฐโ\(5,20\)\\bar\{\\bm\{x\}\}\\sim\\mathcal\{U\}\(5,20\),ฮฒยฏ0โeโผ๐ฐโ\(0,1\)\\bar\{\\beta\}\_\{0e\}\\sim\\mathcal\{U\}\(0,1\), andฮฒยฏpโeโผ๐ฐโ\(0,1\)\\bar\{\\beta\}\_\{pe\}\\sim\\mathcal\{U\}\(0,1\)for allpโ\[5\]p\\in\[5\]andeโโฐe\\in\\mathcal\{E\}\. Nominal covariate realizations for each sampleiโ\[100\]i\\in\[100\]are then generated from a normal distribution centered at the mean feature vector,๐iโผ๐ฉโ\(๐ยฏ,\(10/6\)2โ๐d\)\\bm\{x\}\_\{i\}\\sim\\mathcal\{N\}\(\\bar\{\\bm\{x\}\},\(10/6\)^\{2\}\\mathbb\{I\}\_\{d\}\)\. The corresponding ground\-truth cost responses are subsequently simulated as:
ceโiโผ๐ฉโ\(ฮฒยฏ0โe\+โp=15ฮฒยฏpโeโxยฏp,โp=15ฮฒยฏpโe2โ\(106\)2\)โeโโฐ,iโ\[100\]\.c\_\{ei\}\\sim\\mathcal\{N\}\\left\(\\bar\{\\beta\}\_\{0e\}\+\\sum\_\{p=1\}^\{5\}\\bar\{\\beta\}\_\{pe\}\\bar\{x\}\_\{p\},\\,\\sum\_\{p=1\}^\{5\}\\bar\{\\beta\}^\{2\}\_\{pe\}\\Big\(\\frac\{10\}\{6\}\\Big\)^\{2\}\\right\)\\quad\\forall e\\in\\mathcal\{E\},\\,i\\in\[100\]\.To establish a meaningful scale for the uncertainty budgetฮป\\lambda, we calculate the maximum possible magnitude of the unconstrained predictive disturbance under anโ2\\ell\_\{2\}\-norm\. Specifically, the worst\-case disturbance bound across the training set is given by:
maxiโ\[100\]โก\{โeโโฐ\(โpโ\[5\]ฮฒpโeโxpโi\)2\}\.\\max\_\{i\\in\[100\]\}\\left\\\{\\sqrt\{\\sum\_\{e\\in\\mathcal\{E\}\}\\left\(\\sum\_\{p\\in\[5\]\}\\beta\_\{pe\}x\_\{pi\}\\right\)^\{2\}\}\\right\\\}\.For our base case experiments, the uncertainty budget is calibrated to a fraction of this worst\-case threshold, settingฮป\\lambdaas 10% of this value\. To evaluate robustness against covariate disturbances, we construct100100distinct contaminated datasets, indexed byll\. For each dataset, uniform measurement errors are injected into the observed features:
Epโiโlโผ๐ฐโ\(โxpโi,5โxpโi\),xpโiโlc=xpโi\+Epโiโlโpโ\[5\],iโ\[100\],lโ\[100\],E\_\{pil\}\\sim\\mathcal\{U\}\(\-x\_\{pi\},5x\_\{pi\}\),\\quad x^\{c\}\_\{pil\}=x\_\{pi\}\+E\_\{pil\}\\quad\\forall p\\in\[5\],\\,i\\in\[100\],\\,l\\in\[100\],wherexpโiโlcx^\{c\}\_\{pil\}represents the corrupted covariate value observed by the learner, and is such thatmaxiโ\[100\],lโ\[100\]โก\{โeโโฐ\(โpโ\[5\]ฮฒpโeโxpโiโlc\)2\}โคฮป\\max\_\{i\\in\[100\],l\\in\[100\]\}\\left\\\{\\sqrt\{\\sum\_\{e\\in\\mathcal\{E\}\}\\left\(\\sum\_\{p\\in\[5\]\}\\beta\_\{pe\}x^\{c\}\_\{pil\}\\right\)^\{2\}\}\\right\\\}\\leq\\lambdato ensure that contaminated datasets follow our budget of uncertainty\. In the subsequent subsections, we analyze and contrast the frameworks in terms of out\-of\-sample decision regret and training stability under these contaminated covariate regimes\.
### 6\.1Training stability \(Theorem[3\.2](https://arxiv.org/html/2607.21773#S3.Thmtheorem2)\)
The empirical results from our network flow experiments demonstrate the clear training performance benefits of the proposed SPrO\+ framework under covariate contamination\. Figure[2](https://arxiv.org/html/2607.21773#S6.F2)plots the decision regret across training sample sizes ranging fromn=20n=20ton=300n=300\.
Figure 2:Training stability comparisonThe experimental data strongly substantiates the theoretical guarantees established in Theorem[3\.2](https://arxiv.org/html/2607.21773#S3.Thmtheorem2)\. Both SPO\+ and SrPO\+ exhibit significant performance fluctuations as the sample size increases, characterized by sharp, volatile jumps in regret\. This volatility reflects the classic vertex\-switching pathology of unregularized polyhedral decision sets, where minor adjustments in parameter estimates cause important shifts in optimal solutions\. The same behaviour can be inferred from the within\-sample variance\. Conversely, the SPrO\+ regret profile is markedly smoother and displays much lower variance, tightly stabilizing between33and55oncenโฅ80n\\geq 80\. This empirically validates how ourโ2\\ell\_\{2\}\-driven Lipschitz continuity improves training performance of SPrO\+\. An important takeaway is the near\-identical performance curve of SrPO\+ relative to standard SPO\+, showing that upstream regularization does little to improve training stability\. In addition, SPrO\+ loss values are lower than SPO\+ and SrPO\+, a promising indicator, as discussed in Section 4, that robust downstream decision\-making is highly likely to result in lower decision losses\. Subsequent sections will showcase this empirically\.
### 6\.2Out\-of\-sample regret under data contamination
To evaluate the out\-of\-sample robustness and sensitivity of the learned parameters to feature corruption, we examine the downstream decision loss under data contamination\. Specifically, we first train each framework on the nominal, uncontaminated dataset to estimate the optimal coefficient matrix๐ฉโ\\bm\{B\}^\{\*\}\. We then fix these parameters and evaluate their performance on the100100contaminated datasets\. For each contaminated datasetll, the downstream performance is quantified via the average ex\-post unambiguous regret1100โโiโ\[100\]\(max๐โ๐ฒโโ\(๐ฉโโ๐iโlc\)โก๐iโคโ๐โzโโ\(๐i\)\)\\frac\{1\}\{100\}\\sum\_\{i\\in\[100\]\}\\left\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\bm\{B\}^\{\*\}\\bm\{x\}^\{c\}\_\{il\}\)\}\\bm\{c\}\_\{i\}^\{\\top\}\\bm\{w\}\-z^\{\*\}\(\\bm\{c\}\_\{i\}\)\\right\)\. By taking the maximum over the set๐ฒโโ\(โ
\)\\mathcal\{W\}^\{\*\}\(\\cdot\), this metric precisely captures the worst\-case decision regret in the presence of non\-unique optimal downstream solutions\. We add the coefficient regularization termฮผโโeโโฐโ๐ทeโ1\\mu\\sum\_\{e\\in\\mathcal\{E\}\}\\\|\\bm\{\\beta\}\_\{e\}\\\|\_\{1\}, whereฮผ\\muis a very small number, to SPO\+ and SPrO\+ \(not SrPO\+ as it already regularizes the coefficients via the budget of uncertainty\)\. This penalty prevents multiple optimal solutions with excessively large regression coefficients for decisions that are zero\-valued in the downstream problem\. While this lexicographic regularization technique maintains the same in\-sample decision regret properties, it can significantly enhance out\-of\-sample performances\. Figure[3](https://arxiv.org/html/2607.21773#S6.F3)shows the performance comparisons\. The graphs portray the performances across contaminated datasets, as well as the average performance in each contaminated dataset, together with the within\-dataset performance standard deviation band\.


Figure 3:Unambiguous regret comparison under data contaminationThe empirical results show that SrPO\+ suffers from severe out\-of\-sample decision degradation, with regret values spanning widely between 38 and 85, and peaking above 100, and showing considerable variance\. This provides powerful empirical proof for one of our core theses: worst\-case\-cost robustness, achieved via upstream regularization, does not inherently guarantee robust downstream decisions\. In fact, conventional upstream regularization acts pointwise on the highest cost and offers no performance guarantees outside of it\. In fact, by optimizing independently against worst\-case cost of feature perturbations, the predictor fundamentally lacks visibility into the downstream decisions\. It treats every cost component as equally critical, effectively blind to the fact that downstream decision\-making only cares about the regret in terms of the ground\-truth cost\. If coefficient regularization is desired, SPO\+ with a small penalty added for the size of the regression coefficients offers a superior alternative because it penalizes the coefficients jointly with the structural SPO\+ loss\. This joint formulation traces out a Pareto frontier that lexicographically balances the minimization of downstream SPO\+ regret against the magnitude of the regression coefficients\. Although SPO\+ delivers reasonable performance \(with regret spanning between 18 and 45\), our proposed SPrO\+ framework consistently dominates, restricting the decision loss primarily to the single digits \(in 54 out of 100 runs\) and ranging from 5 to 17\. The average decision loss of SPrO\+ is 10\.3, a substantial reduction compared to 32\.5 for SPO\+\. In addition, SPrO\+ offers more stable performance, with a standard deviation of 2\.9, compared to 6\.7 for SPO\+\. This variance suppression is also evident within individual sample paths; the shaded within\-dataset standard deviation bands for SPrO\+ remain tightly bounded between 1 and 4 across 90% of the runs\. In contrast, standard SPO\+ displays significant volatility spikes, with only 1% of its runs maintaining a standard deviation below 3\. This confirms that SPrO\+ not only optimizes expected downstream performance but also drastically flattens decision volatility across independent data realizations\.
### 6\.3Sensitivity analysis on budget of uncertainty and problem dimension
Our theoretical consistency \(Corollary[3\.6](https://arxiv.org/html/2607.21773#S3.Thmtheorem6)\) and stochastic dominance \(Corollary[4\.2](https://arxiv.org/html/2607.21773#S4.Thmtheorem2), Theorem[5\.4](https://arxiv.org/html/2607.21773#S5.Thmtheorem4)\) results point to the budget of uncertainty and the problem dimension as being key factors\. We first double the level of conservatism by running our base case withฮป=0\.20โmaxiโ\[50\]โก\{โeโโฐ\(โpโ\[5\]ฮฒpโeโxpโi\)2\}\\lambda=0\.20\\max\_\{i\\in\[50\]\}\\left\\\{\\sqrt\{\\sum\_\{e\\in\\mathcal\{E\}\}\\left\(\\sum\_\{p\\in\[5\]\}\\beta\_\{pe\}x\_\{pi\}\\right\)^\{2\}\}\\right\\\}to test the impact of higher budget of uncertainty\. Figure[4](https://arxiv.org/html/2607.21773#S6.F4)plots the results\.


Figure 4:Unambiguous regret comparison under data contamination for a higher budget of uncertaintyThe empirical results reveal that increasing the uncertainty budget induces convergence in expected performance among SPO\+, SrPO\+ and SPrO\+, yet clear behavioral distinctions remain\. SPO\+ decision loss dips below 30 in only 36% of the samples and its out\-of\-sample decision loss sits at an average of 32\.4\. Conversely, because SPrO\+ regularizes the decision rather than the predictor, it retains the flexibility to achieve low regrets when the sample path allows\. Notice that SPrO\+ outperforms SPO\+ on all instances, obtaining better average values and better performance stability, although compared to the lower budget of uncertainty, the absolute performance gap between SPO\+ and SPrO\+ has narrowed\. This behavior aligns with robust optimization theory: as the uncertainty budget lambda scales up, the regularizer increasingly dominates the objective function \(over the predicted cost\), naturally driving SPrO\+ frameworks toward more conservative decision policies that protect against worst\-case scenarios\. Despite this convergence, SPrO\+ retains strict stochastic dominance, outperforming SPO\+ across all 100 simulation runs while maintaining superior performance stability, with a path standard deviation of 4\.3 compared to SPO\+โs 5\.5\.
Crucially, the sensitivity analysis sheds new light on the behavior of the upstream regularization paradigm, SrPO\+\. Under this elevated uncertainty budget, SrPO\+ average regret drops significantly relative to its baseline, centering its bulk distribution around a median of 33\.1 \- nearly identical to standard SPO\+\. However, as shown in the simulation run tracking graph, SrPO\+ suffers from severe, erratic volatility spikes, with decision regret aggressively fluctuating between 20 and 56 across successive simulation runs\. This demonstrates that while a massive upstream uncertainty budget can accidentally lower average regret by forcing heavy coefficient attenuation, it introduces profound structural instability\. SPrO\+ mitigates this volatility, demonstrating that downstream robustification achieves both lower expected regret and superior risk suppression under high conservatism\.
For sensitivity analysis on problem dimension, we scale up our graph by creating a\|๐ฑ\|=20\|\\mathcal\{V\}\|=20and\|โฐ\|=60\|\\mathcal\{E\}\|=60random instance and running similar base\-case analysis\.


Figure 5:Unambiguous regret comparison under data contamination for a higher\-dimensional problemFigure[5](https://arxiv.org/html/2607.21773#S6.F5)reveals that scaling up the problem dimension makes the out\-of\-sample advantages of SPrO\+ even more pronounced over both standard SPO\+ and the upstream proxy SrPO\+\. The SPrO\+ framework achieves a highly concentrated regret distribution bounded tightly between 5\.5 and 11, with an average decision loss of 7\.8 and exceptional path stability\. In contrast, SPO\+ exhibits an elevated average regret of 28\.6 with significantly higher variance, while SrPO\+ displays a heavily dispersed distribution ranging between 11 and 43, underscoring its vulnerability to high\-dimensional problems, although it outperforms SPO\+ on average and in most simulation runs\.
This striking divergence under higher dimensions is mathematically justified by the structural impact of decision regularization\. In high\-dimensional optimization landscapes, nominal prediction errors cause unregularized decisions to significantly fluctuate between distant extreme points on the feasible regionโs boundary\. By embedding a dual\-norm regularizer into the downstream optimizer, SPrO\+ severely penalizes these erratic shifts and ensures that the Euclidean norm of the difference between regularized decisions remains tightly bounded compared to their unregularized counterparts\. Geometrically, this behavior satisfies the structural conditions required for the Sudakov\-Fernique inequality to hold, i\.e\.โ๐๐นโฃโ,๐โ\(๐\)โ๐๐นโฃโ,๐โ\(๐\)โ2โคโ๐โ,2โ\(๐\)โ๐โ,1โ\(๐\)โ2\\\|\\bm\{w^\{R\*,2\}\}\(\\bm\{c\}\)\-\\bm\{w^\{R\*,1\}\}\(\\bm\{c\}\)\\\|\_\{2\}\\leq\\\|\\bm\{w\}^\{\*,2\}\(\\bm\{c\}\)\-\\bm\{w\}^\{\*,1\}\(\\bm\{c\}\)\\\|\_\{2\}\. By keeping the expected maximum distance between perturbed optimal decisions low, SPrO\+ effectively suppresses decision volatility and maintains absolute performance superiority over standard frameworks\.
## 7Concluding remarks
In this paper, we introduced SPrO\+, a novel end\-to\-end prediction\-and\-optimization framework that expands the seminal Smart Predict\-then\-Optimize \(SPO\) paradigm to handle contextual optimization under prediction shifts and data contamination\. By shifting the burden of conservatism directly into the downstream decision space, our framework establishes a computationally tractable, convex surrogate that matches the optimization efficiency of standard SPO\+ surrogate while providing explicit immunity to covariate disturbances\. Theoretically, we demonstrated that robustifying downstream decision\-making leads to decision shrinkage, which subsequently yields superior generalization and structural consistency\. Rather than relying purely on asymptotic approximations, we proved that SPrO\+ achieves finite\-sample Fisher consistency with high probability, alongside non\-asymptotic concentration bounds that limit the probability of large surrogate gaps\. Furthermore, we established the necessary conditions under which our approach could stochastically and pointwise dominate both standard, uncertainty\-agnostic SPO\+ and upstream regularized proxies \(SrPO\+\), providing a rigorous foundation for decision\-space robustification\. Finally, through extensive computational experiments across diverse operational environments, varying sample sizes, and highly volatile regimes, we confirmed that these theoretical advantages translate into substantial empirical gains\. Our numerical results demonstrate that SPrO\+ not only systematically minimizes out\-of\-sample decision regret, but also suppresses variance across independent evaluation paths\. By bridging the gap between prescriptive performance and data\-driven stability, SPrO\+ provides a robust, scalable paradigm for prescriptive analytics in deeply uncertain contextual environments\.
## References
- Machine learning and portfolio optimization\.Management Science64\(3\),pp\. 1136โ1154\.Cited by:[ยง1](https://arxiv.org/html/2607.21773#S1.p1.4)\.
- A\. Banerjee, S\. Chen, F\. Fazayeli, and V\. Sivakumar \(2014\)Estimation with norm regularization\.Advances in neural information processing systems27\.Cited by:[Proof 7\.10](https://arxiv.org/html/2607.21773#S7.Thmtheorem10.p1.8.4),[Proof 7\.5](https://arxiv.org/html/2607.21773#S7.Thmtheorem5.p1.18.1),[Proof 7\.6](https://arxiv.org/html/2607.21773#S7.Thmtheorem6.p1.15.4)\.
- D\. Basak, S\. Pal, D\. C\. Patranabis,et al\.\(2007\)Support vector regression\.Neural Information Processing\-Letters and Reviews11\(10\),pp\. 203โ224\.Cited by:[ยง3](https://arxiv.org/html/2607.21773#S3.p9.5)\.
- Y\. Bengio \(1997\)Using a financial training criterion rather than a prediction criterion\.International journal of neural systems8\(04\),pp\. 433โ443\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p2.1)\.
- D\. Bertsimas, D\. B\. Brown, and C\. Caramanis \(2011\)Theory and applications of robust optimization\.SIAM review53\(3\),pp\. 464โ501\.Cited by:[ยง1](https://arxiv.org/html/2607.21773#S1.p3.1)\.
- P\. Donti, B\. Amos, and J\. Z\. Kolter \(2017\)Task\-based end\-to\-end model learning in stochastic optimization\.Advances in neural information processing systems30\.Cited by:[ยง1](https://arxiv.org/html/2607.21773#S1.p1.4),[ยง2](https://arxiv.org/html/2607.21773#S2.p1.1),[ยง2](https://arxiv.org/html/2607.21773#S2.p2.1)\.
- O\. El Balghiti, A\. N\. Elmachtoub, P\. Grigas, and A\. Tewari \(2019\)Generalization bounds in the predict\-then\-optimize framework\.Advances in neural information processing systems32\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p4.1)\.
- L\. El Ghaoui and H\. Lebret \(1997\)Robust solutions to least\-squares problems with uncertain data\.SIAM Journal on matrix analysis and applications18\(4\),pp\. 1035โ1064\.Cited by:[ยง2\.1](https://arxiv.org/html/2607.21773#S2.SS1.p4.4)\.
- A\. N\. Elmachtoub and P\. Grigas \(2022\)Smart โpredict, then optimizeโ\.Management Science68\(1\),pp\. 9โ26\.Cited by:[ยง1](https://arxiv.org/html/2607.21773#S1.p1.4),[ยง2\.1](https://arxiv.org/html/2607.21773#S2.SS1.p4.3),[ยง2](https://arxiv.org/html/2607.21773#S2.p3.1),[ยง2](https://arxiv.org/html/2607.21773#S2.p4.1)\.
- A\. N\. Elmachtoub, J\. C\. N\. Liang, and R\. McNellis \(2020\)Decision trees for decision\-making under the predict\-then\-optimize framework\.InInternational conference on machine learning,pp\. 2858โ2867\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p3.1)\.
- N\. Ho\-Nguyen and F\. Kฤฑlฤฑnรง\-Karzan \(2022\)Risk guarantees for end\-to\-end prediction and optimization processes\.Management Science68\(12\),pp\. 8680โ8698\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p4.1),[Proof 7\.6](https://arxiv.org/html/2607.21773#S7.Thmtheorem6.p1.21.2),[Proof 7\.6](https://arxiv.org/html/2607.21773#S7.Thmtheorem6.p1.9.9)\.
- Y\. Hu, N\. Kallus, and X\. Mao \(2022\)Fast rates for contextual linear optimization\.Management Science68\(6\),pp\. 4236โ4245\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p4.1)\.
- N\. Kallus and X\. Mao \(2023\)Stochastic optimization forests\.Management Science69\(4\),pp\. 1975โ1994\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p2.1)\.
- D\. Katselis, X\. Xie, C\. L\. Beck, and R\. Srikant \(2021\)On concentration inequalities for vector\-valued lipschitz functions\.Statistics & Probability Letters173,pp\. 109071\.Cited by:[ยง3\.1](https://arxiv.org/html/2607.21773#S3.SS1.p3.16),[ยง3\.1](https://arxiv.org/html/2607.21773#S3.SS1.p7.8)\.
- E\. Keyvanshokooh, M\. Zhalechian, C\. Shi, M\. P\. Van Oyen, and P\. Kazemian \(2019\)Contextual learning with online convex optimization: theory and application to medical decision\-making\.Management Science, to appear\.Cited by:[ยง1](https://arxiv.org/html/2607.21773#S1.p1.4)\.
- L\. Kong, J\. Cui, Y\. Zhuang, R\. Feng, B\. A\. Prakash, and C\. Zhang \(2022\)End\-to\-end stochastic optimization with energy\-based model\.Advances in Neural Information Processing Systems35,pp\. 11341โ11354\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p2.1)\.
- S\. Liu, L\. He, and Z\. Max Shen \(2021\)On\-time last\-mile delivery: order assignment with travel\-time predictors\.Management Science67\(7\),pp\. 4095โ4119\.Cited by:[ยง1](https://arxiv.org/html/2607.21773#S1.p1.4)\.
- J\. Mandi, P\. J\. Stuckey, T\. Guns,et al\.\(2020\)Smart predict\-and\-optimize for hard combinatorial optimization problems\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 1603โ1610\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p1.1),[ยง2](https://arxiv.org/html/2607.21773#S2.p3.1)\.
- Y\. P\. Patel, S\. Rayan, and A\. Tewari \(2024\)Conformal contextual robust optimization\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 2485โ2493\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p5.1)\.
- R\. T\. Rockafellar \(1997\)Convex analysis\.Vol\.11,Princeton university press\.Cited by:[Lemma 7\.2](https://arxiv.org/html/2607.21773#S7.Thmtheorem2)\.
- C\. Rudin, C\. Chen, Z\. Chen, H\. Huang, L\. Semenova, and C\. Zhong \(2022\)Interpretable machine learning: fundamental principles and 10 grand challenges\.Statistic Surveys16,pp\. 1โ85\.Cited by:[ยง2\.1](https://arxiv.org/html/2607.21773#S2.SS1.p2.7)\.
- U\. Sadana, A\. Chenreddy, E\. Delage, A\. Forel, E\. Frejinger, and T\. Vidal \(2024\)A survey of contextual optimization methods for decision\-making under uncertainty\.European Journal of Operational Research\.Cited by:[ยง1](https://arxiv.org/html/2607.21773#S1.p1.4),[ยง2](https://arxiv.org/html/2607.21773#S2.p1.1)\.
- P\. K\. Shivaswamy, C\. Bhattacharyya, and A\. J\. Smola \(2006\)Second order cone programming approaches for handling missing and uncertain data\.Journal of Machine Learning Research,pp\. 1283โ1314\.Cited by:[ยง2\.1](https://arxiv.org/html/2607.21773#S2.SS1.p4.4)\.
- J\. E\. Smith and R\. L\. Winkler \(2006\)The optimizerโs curse: skepticism and postdecision surprise in decision analysis\.Management Science52\(3\),pp\. 311โ322\.Cited by:[ยง1](https://arxiv.org/html/2607.21773#S1.p2.4)\.
- C\. Sun, L\. Liu, and X\. Li \(2023\)Predict\-then\-calibrate: a new perspective of robust contextual lp\.Advances in neural information processing systems36,pp\. 17713โ17741\.Cited by:[ยง2](https://arxiv.org/html/2607.21773#S2.p5.1)\.
- A\. Vellido \(2020\)The importance of interpretability and visualization in machine learning for applications in medicine and health care\.Neural computing and applications32\(24\),pp\. 18069โ18083\.Cited by:[ยง2\.1](https://arxiv.org/html/2607.21773#S2.SS1.p2.7)\.
- R\. Vershynin \(2012\)Introduction to the non\-asymptotic analysis of random matrices\.\.Cited by:[Proof 7\.5](https://arxiv.org/html/2607.21773#S7.Thmtheorem5.p1.9.4)\.
- H\. Xu, C\. Caramanis, and S\. Mannor \(2008\)Robust regression and lasso\.Advances in neural information processing systems21\.Cited by:[ยง2\.1](https://arxiv.org/html/2607.21773#S2.SS1.p4.4),[ยง5](https://arxiv.org/html/2607.21773#S5.p2.7),[ยง5](https://arxiv.org/html/2607.21773#S5.p3.3)\.
- J\. Zhen, D\. Kuhn, and W\. Wiesemann \(2025\)A unified theory of robust and distributionally robust optimization via the primal\-worst\-equals\-dual\-best principle\.Operations Research73\(2\),pp\. 862โ878\.Cited by:[Definition 2\.1](https://arxiv.org/html/2607.21773#S2.Thmtheorem1),[Lemma 7\.1](https://arxiv.org/html/2607.21773#S7.Thmtheorem1)\.
\\ECHead
Proofs of propositions Throughout this paper, the derivations of robust counterparts will rely on the following two lemmas\.
###### Lemma 7\.1\(Proposition C\.4 in\(Zhenet al\.[2025](https://arxiv.org/html/2607.21773#bib.bib48)\)\)
Ifโฉkriโก\(domโก\(hk\)\)โ โ
\\cap\_\{k\}\\operatorname\{ri\}\(\\operatorname\{dom\}\(h\_\{k\}\)\)\\neq\\emptyset, the convex conjugate of the sum of proper convex functions is equal to the infimal convolution of the conjugates of these functions, i\.e\.,
\(โkhk\)โโ\(๐\)=inf๐k,โk\{โkhkโโ\(๐k\):โk๐k=๐\}\.\\displaystyle\\Big\(\\sum\_\{k\}h\_\{k\}\\Big\)^\{\*\}\(\\bm\{y\}\)=\\inf\_\{\\bm\{y\}\_\{k\},\\,\\forall k\}\\Big\\\{\\sum\_\{k\}h^\{\*\}\_\{k\}\(\\bm\{y\}\_\{k\}\):\\sum\_\{k\}\\bm\{y\}\_\{k\}=\\bm\{y\}\\Big\\\}\.
###### Lemma 7\.2\(Theorem 16\.1 in\(Rockafellar[1997](https://arxiv.org/html/2607.21773#bib.bib49)\)\)
The conjugate of a positive multiple of a proper convex function equals the perspective of the conjugate of this function, i\.e\.,
\(sโh\)โโ\(๐\)=\(hโโs\)โ\(๐\)\.\\displaystyle\\left\(sh\\right\)^\{\*\}\(\\bm\{y\}\)=\(h^\{\*\}s\)\(\\bm\{y\}\)\.
###### Proof 7\.3
Proof of Theorem[3\.1](https://arxiv.org/html/2607.21773#S3.Thmtheorem1)\. We can rewrite the inner maximization model with the following equivalences
max๐^โคโ๐\+ฮปโโ๐โโโคzRโฃโโ\(๐^\)\+ฮปโโ๐๐นโฒโ\(๐^\)โโ๐โ๐ฒโก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\\displaystyle\\max\_\{\\begin\{subarray\}\{c\}\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\leq z^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\\\\ \\bm\{w\}\\in\\mathcal\{W\}\\end\{subarray\}\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}โ\\displaystyle\\Leftrightarrowmax๐โกmin๐
โฅ๐,ฮฑโฅ0โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโโโkโ\[m\]ฯkโgkโ\(๐\)\+ฮฑโ\(zRโฃโโ\(๐^\)โ๐^โคโ๐โฮปโโ๐โโ\+ฮปโโ๐๐นโฒโ\(๐^\)โโ\)\}\\displaystyle\\max\_\{\\bm\{w\}\}\\min\_\{\\bm\{\\pi\}\\geq\\bm\{0\},\\alpha\\geq 0\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\-\\sum\_\{k\\in\[m\]\}\\pi\_\{k\}g\_\{k\}\(\\bm\{w\}\)\+\\alpha\(z^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\)\\big\\\}โ\\displaystyle\\Leftrightarrowmin๐
โฅ๐,ฮฑโฅ0โก\{ฮฑโzRโฃโโ\(๐^\)\+ฮฑโฮปโโ๐๐นโฒโ\(๐^\)โโ\+max๐โก\{๐โคโ๐โ\(1\+ฮฑ\)โ๐^โคโ๐โ\(1\+ฮฑ\)โฮปโโ๐โโโโkโ\[m\]ฯkโgkโ\(๐\)\}\}\\displaystyle\\min\_\{\\bm\{\\pi\}\\geq\\bm\{0\},\\alpha\\geq 0\}\\bigg\\\{\\alpha z^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\+\\alpha\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\max\_\{\\bm\{w\}\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\(1\+\\alpha\)\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\(1\+\\alpha\)\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\-\\sum\_\{k\\in\[m\]\}\\pi\_\{k\}g\_\{k\}\(\\bm\{w\}\)\\big\\\}\\bigg\\\}โ\\displaystyle\\Leftrightarrow\{min๐
โฅ๐,ฮฑโฅ0โโkโ\[m\]\(gkโโฯk\)โ\(ฯk\)\+\(hโโ\(1\+ฮฑ\)โฮป\)โ\(๐ฝ\)\+ฮฑโzRโฃโโ\(๐^\)\+ฮฑโฮปโโ๐๐นโฒโ\(๐^\)โโโkโ\[m\]ฯk\+๐ฝ=๐โ\(1\+ฮฑ\)โ๐^\\displaystyle\\begin\{cases\}\\displaystyle\\min\_\{\\bm\{\\pi\}\\geq\\bm\{0\},\\alpha\\geq 0\}\\sum\_\{k\\in\[m\]\}\(g^\{\*\}\_\{k\}\\pi\_\{k\}\)\(\\bm\{\\phi\}\_\{k\}\)\+\(h^\{\*\}\(1\+\\alpha\)\\lambda\)\(\\bm\{\\theta\}\)\+\\alpha z^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\+\\alpha\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\\\\ \\displaystyle\\sum\_\{k\\in\[m\]\}\\bm\{\\phi\}\_\{k\}\+\\bm\{\\theta\}=\\bm\{c\}\-\(1\+\\alpha\)\\hat\{\\bm\{c\}\}\\end\{cases\}โ\\displaystyle\\Leftrightarrow\{min๐
โฅ๐,ฮฑโฅ0โโkโ\[m\]\(gkโโฯk\)โ\(ฯk\)\+ฮฑโzRโฃโโ\(๐^\)\+ฮฑโฮปโโ๐๐นโฒโ\(๐^\)โโโkโ\[m\]ฯk\+๐ฝ=๐โ\(1\+ฮฑ\)โ๐^โ๐ฝโโค\(1\+ฮฑ\)โฮป\\displaystyle\\begin\{cases\}\\displaystyle\\min\_\{\\bm\{\\pi\}\\geq\\bm\{0\},\\alpha\\geq 0\}\\sum\_\{k\\in\[m\]\}\(g^\{\*\}\_\{k\}\\pi\_\{k\}\)\(\\bm\{\\phi\}\_\{k\}\)\+\\alpha z^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\+\\alpha\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\\\\ \\displaystyle\\sum\_\{k\\in\[m\]\}\\bm\{\\phi\}\_\{k\}\+\\bm\{\\theta\}=\\bm\{c\}\-\(1\+\\alpha\)\\hat\{\\bm\{c\}\}\\\\ \\displaystyle\\\|\\bm\{\\theta\}\\\|\\leq\(1\+\\alpha\)\\lambda\\end\{cases\}โ\\displaystyle\\Leftrightarrow\{min๐
โฅ๐,ฮฑโฅ0โโkโ\[m\]\(gkโโฯk\)โ\(ฯk\)\+ฮฑ1\+ฮฑโ\(zRโฃโโ\(๐^\)\+ฮปโโ๐๐นโฒโ\(๐^\)โโ\)โkโ\[m\]ฯk\+๐ฝ=๐โ๐^โ๐ฝโโคฮป\\displaystyle\\begin\{cases\}\\displaystyle\\min\_\{\\bm\{\\pi\}\\geq\\bm\{0\},\\alpha\\geq 0\}\\sum\_\{k\\in\[m\]\}\(g^\{\*\}\_\{k\}\\pi\_\{k\}\)\(\\bm\{\\phi\}\_\{k\}\)\+\\frac\{\\alpha\}\{1\+\\alpha\}\\bigg\(z^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\\bigg\)\\\\ \\displaystyle\\sum\_\{k\\in\[m\]\}\\bm\{\\phi\}\_\{k\}\+\\bm\{\\theta\}=\\bm\{c\}\-\\hat\{\\bm\{c\}\}\\\\ \\displaystyle\\\|\\bm\{\\theta\}\\\|\\leq\\lambda\\end\{cases\}โ\\displaystyle\\Leftrightarrow\{min๐
โฅ๐โโkโ\[m\]\(gkโโฯk\)โ\(ฯk\)โkโ\[m\]ฯk\+๐ฝ=๐โ๐^โ๐ฝโโคฮป\.\\displaystyle\\begin\{cases\}\\displaystyle\\min\_\{\\bm\{\\pi\}\\geq\\bm\{0\}\}\\sum\_\{k\\in\[m\]\}\(g^\{\*\}\_\{k\}\\pi\_\{k\}\)\(\\bm\{\\phi\}\_\{k\}\)\\\\ \\displaystyle\\sum\_\{k\\in\[m\]\}\\bm\{\\phi\}\_\{k\}\+\\bm\{\\theta\}=\\bm\{c\}\-\\hat\{\\bm\{c\}\}\\\\ \\displaystyle\\\|\\bm\{\\theta\}\\\|\\leq\\lambda\\end\{cases\}\.The first two equivalences are from Lagrangian duality, where strong duality applies because of Assumption[2\.1](https://arxiv.org/html/2607.21773#S2.Thmtheorem1)\. The third equivalence is from the application of Lemmas[7\.1](https://arxiv.org/html/2607.21773#S7.Thmtheorem1)and[7\.2](https://arxiv.org/html/2607.21773#S7.Thmtheorem2), defininghโh^\{\*\}as the convex conjugate of the dual norm\. The fourth equivalence explicitly formulates the conjugate, where ifhโ\(๐ฒ\)=โ๐ฒโโh\(\\bm\{y\}\)=\\\|\\bm\{y\}\\\|\_\{\*\}, thenhโโ\(๐ฅ\)=0h^\{\*\}\(\\bm\{l\}\)=0ifโ๐ฅโโค1\\\|\\bm\{l\}\\\|\\leq 1andโ\\inftyotherwise\. The fifth equivalence is becausezRโฃโโ\(๐^\)z^\{R\*\}\(\\hat\{\\bm\{c\}\}\)andฮปโโ๐ฐ๐โฒโ\(๐^\)โโ\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}are both homogeneous with respect to the scaling of\(๐^,ฮป\)\(\\hat\{\\bm\{c\}\},\\lambda\)\. The final equivalence is becauseฮฑ1\+ฮฑ\\frac\{\\alpha\}\{1\+\\alpha\}in monotonically increasing inฮฑ\\alpha, noting that ifzRโฃโ<0z^\{R\*\}<0the model will be unbounded\.
The convexity with respect to๐^\\hat\{\\bm\{c\}\}is because the perspective function of a convex conjugate is convex\. To prove that SPrO\+ is a tighter approximation of SPO than SPO\+ whenฮป=0\\lambda=0, we start with the definition of SPrO\+
โSโPโrโO\+โ\(๐^,๐\)\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle=\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}=max๐โ๐ฒโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐\}\+๐^โคโ๐โโ\(๐\)โzโโ\(๐\),\\displaystyle=\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\),where the equality is obtained by settingฮป=0\\lambda=0, which yieldszRโฃโโ\(๐\)=zโโ\(๐\)z^\{R\*\}\(\\bm\{c\}\)=z^\{\*\}\(\\bm\{c\}\)and๐ฒRโฃโโ\(โ
\)=๐ฒโโ\(โ
\)\\mathcal\{W\}^\{R\*\}\(\\cdot\)=\\mathcal\{W\}^\{\*\}\(\\cdot\)\. We can thus see thatโSโPโrโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)upper approximates SPO tighter, in the sense that
โSโPโOโ\(๐^,๐\)โคmax๐โ๐ฒโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐\}\+๐^โคโ๐โโ\(๐\)โzโโ\(๐\)โคmax๐โ๐ฒโก\{๐โคโ๐โ๐^โคโ๐\}\+๐^โคโ๐โโ\(๐\)โzโโ\(๐\)=โSโPโO\+โ\(๐^,๐\),\\displaystyle\\ell\_\{SPO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\)\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\)=\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\),where the relationship with SPO follows from the fact by optimality,๐^โคโ๐ฐโโ\(๐^\)โค๐^โคโ๐ฐโโ\(๐\)\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\hat\{\\bm\{c\}\}\)\\leq\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)and the relationship with SPO\+ is a direct consequence of the inclusion property๐ฒโโ\(๐^\)โ๐ฒ\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}\)\\subseteq\\mathcal\{W\}\.\\halmos
###### Proof 7\.4
Proof of Theorem[3\.2](https://arxiv.org/html/2607.21773#S3.Thmtheorem2)\.Boundedness\.Let the prediction error๐^=๐โฯต\\hat\{\\bm\{c\}\}=\\bm\{c\}\-\\bm\{\\epsilon\}, withโฯตโโคฮป\\\|\\bm\{\\epsilon\}\\\|\\leq\\lambda\. SPrO\+ loss then reduces to:
โSโPโrโO\+โ\(๐โฯต,๐\)\\displaystyle\\ell\_\{SPrO\+\}\(\\bm\{c\}\-\\bm\{\\epsilon\},\\bm\{c\}\)=\(max๐โ๐ฒRโฃโโ\(๐โฯต\)โก\{๐โคโ๐โ\(๐โฯต\)โคโ๐โฮปโโ๐โโ\}\+\(๐โฯต\)โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle=\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\(\\bm\{c\}\-\\bm\{\\epsilon\}\)^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\(\\bm\{c\}\-\\bm\{\\epsilon\}\)^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}=\(max๐โ๐ฒRโฃโโ\(๐โฯต\)โก\{ฯตโคโ๐โฮปโโ๐โโ\}โฯตโคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\)\+\\displaystyle=\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\}\\big\\\{\\bm\{\\epsilon\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\-\\bm\{\\epsilon\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\bigg\)\_\{\+\}โค\(max๐โ๐ฒRโฃโโ\(๐โฯต\)โก\{โฯตโโโ๐โโโฮปโโ๐โโ\}โฯตโคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\)\+\\displaystyle\\leq\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\}\\big\\\{\\\|\\bm\{\\epsilon\}\\\|\\\|\\bm\{w\}\\\|\_\{\*\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\-\\bm\{\\epsilon\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\bigg\)\_\{\+\}โค\(โฯตโคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\)\+โค2โฮปโโ๐๐นโฃโโ\(๐\)โโ,\\displaystyle\\leq\\bigg\(\-\\bm\{\\epsilon\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\bigg\)\_\{\+\}\\leq 2\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\},where the first inequality from the generalized Cauchy\-Schwarz and the second inequality results fromโฯตโโคฮป\\\|\\bm\{\\epsilon\}\\\|\\leq\\lambda\. The third inequality also follows from Cauchy\-Schwarz and from the fact that๐ฒRโฃโโ\(๐\)\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)is a singleton\.
Lipschitz continuity\.Let us now verify Lipschitz continuity by first showing that the minimizer of the regularized problem is Lipschitz continuous under strong convexity and then proving that this leads to Lipschitz continuity in the loss function\. Sincehโ\(๐ฐ;๐\)=๐โคโ๐ฐ\+ฮปโโ๐ฐโโh\(\\bm\{w\};\\bm\{c\}\)=\\bm\{c\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}is amm\-strongly convex function, we know that it satisfies the quadratic growth condition
hโ\(๐๐นโฃโโ\(๐1\);๐\)โฅhโ\(๐๐นโฃโโ\(๐\);๐\)\+\(โhโ\(๐๐นโฃโโ\(๐\);๐\)\)โคโ\(๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)\)\+m2โโ๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)โ2\\displaystyle h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\)\\geq h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\)\+\(\\partial h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\)\)^\{\\top\}\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\)\+\\frac\{m\}\{2\}\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|^\{2\}โ\\displaystyle\\Leftrightarrowhโ\(๐๐นโฃโโ\(๐1\);๐\)โฅhโ\(๐๐นโฃโโ\(๐\);๐\)\+m2โโ๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)โ2,\\displaystyle h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\)\\geq h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\)\+\\frac\{m\}\{2\}\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|^\{2\},where the equivalence is by optimality condition, leading to๐โโhโ\(๐ฐ๐โฃโโ\(๐\);๐\)\\bm\{0\}\\in\\partial h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\)\. By optimality of๐ฐ๐โฃโโ\(๐1\)\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\), we know thathโ\(๐ฐ๐โฃโโ\(๐\);๐1\)โฅhโ\(๐ฐ๐โฃโโ\(๐1\);๐1\)h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\_\{1\}\)\\geq h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\_\{1\}\)\. Summing with the growth condition inequality, we obtain
hโ\(๐๐นโฃโโ\(๐1\);๐\)\+hโ\(๐๐นโฃโโ\(๐\);๐1\)โฅhโ\(๐๐นโฃโโ\(๐1\);๐1\)\+hโ\(๐๐นโฃโโ\(๐\);๐\)\+m2โโ๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)โ2\\displaystyle h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\)\+h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\_\{1\}\)\\geq h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\_\{1\}\)\+h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\)\+\\frac\{m\}\{2\}\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|^\{2\}โ\\displaystyle\\Leftrightarrow\(hโ\(๐๐นโฃโโ\(๐1\);๐\)โhโ\(๐๐นโฃโโ\(๐1\);๐1\)\)\+\(hโ\(๐๐นโฃโโ\(๐\);๐1\)โhโ\(๐๐นโฃโโ\(๐\);๐\)\)โฅm2โโ๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)โ2\.\\displaystyle\(h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\)\-h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\_\{1\}\)\)\+\(h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\_\{1\}\)\-h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\)\)\\geq\\frac\{m\}\{2\}\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|^\{2\}\.By definition ofhhand Cauchy\-Schwarz inequality, we see that
\(hโ\(๐๐นโฃโโ\(๐1\);๐\)โhโ\(๐๐นโฃโโ\(๐1\);๐1\)\)\+\(hโ\(๐๐นโฃโโ\(๐\);๐1\)โhโ\(๐๐นโฃโโ\(๐\);๐\)\)\\displaystyle\(h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\)\-h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\);\\bm\{c\}\_\{1\}\)\)\+\(h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\_\{1\}\)\-h\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\);\\bm\{c\}\)\)=\\displaystyle=\(๐โ๐1\)โคโ๐๐นโฃโโ\(๐1\)โ\(๐โ๐1\)โคโ๐๐นโฃโโ\(๐\)โคโ๐โ๐1โโโ๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)โ\.\\displaystyle\(\\bm\{c\}\-\\bm\{c\}\_\{1\}\)^\{\\top\}\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\(\\bm\{c\}\-\\bm\{c\}\_\{1\}\)^\{\\top\}\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\leq\\\|\\bm\{c\}\-\\bm\{c\}\_\{1\}\\\|\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\.From the growth condition, we therefore know that
m2โโ๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)โ2โคโ๐โ๐1โโโ๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)โโโ๐๐นโฃโโ\(๐1\)โ๐๐นโฃโโ\(๐\)โโค2mโโ๐โ๐1โ,\\displaystyle\\frac\{m\}\{2\}\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|^\{2\}\\leq\\\|\\bm\{c\}\-\\bm\{c\}\_\{1\}\\\|\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\\Leftrightarrow\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\_\{1\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\\leq\\frac\{2\}\{m\}\\\|\\bm\{c\}\-\\bm\{c\}\_\{1\}\\\|,thus showing Lipschitz continuity of the minimizer\.
Let us now prove the Lipschitz continuity of the loss function from the above result, the singleton assumption and the fact thatโmax๐ฐโ๐ฒRโฃโโ\(๐\)โก\{โฮปโโ๐ฐโโ\}=ฮปโโ๐ฐ๐โฒโ\(๐\)โโ\-\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\}\\big\\\{\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}=\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}:
โSโPโrโO\+โ\(๐โฯต,๐\)\\displaystyle\\ell\_\{SPrO\+\}\(\\bm\{c\}\-\\bm\{\\epsilon\},\\bm\{c\}\)=\(max๐โ๐ฒRโฃโโ\(๐โฯต\)โก\{๐โคโ๐โ\(๐โฯต\)โคโ๐โฮปโโ๐โโ\}\+\(๐โฯต\)โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle=\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\(\\bm\{c\}\-\\bm\{\\epsilon\}\)^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\(\\bm\{c\}\-\\bm\{\\epsilon\}\)^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\bigg\)\_\{\+\}=\(max๐โ๐ฒRโฃโโ\(๐โฯต\)โก\{ฯตโคโ๐โฮปโโ๐โโ\}โฯตโคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\)\+\\displaystyle=\\bigg\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\}\\big\\\{\\bm\{\\epsilon\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\-\\bm\{\\epsilon\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\bigg\)\_\{\+\}=\(ฯตโคโ\(๐๐นโฃโโ\(๐โฯต\)โ๐๐นโฃโโ\(๐\)\)โฮปโโ๐๐นโฃโโ\(๐โฯต\)โโ\+ฮปโโ๐๐นโฃโโ\(๐\)โโ\)\+\\displaystyle=\\bigg\(\\bm\{\\epsilon\}^\{\\top\}\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\)\-\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\bigg\)\_\{\+\}โคฮปโ\(โ๐๐นโฃโโ\(๐โฯต\)โ๐๐นโฃโโ\(๐\)โโโโ๐๐นโฃโโ\(๐โฯต\)โโ\+โ๐๐นโฃโโ\(๐\)โโ\)\+\\displaystyle\\leq\\lambda\\bigg\(\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\\\|\_\{\*\}\+\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\bigg\)\_\{\+\}โค4โฮปmโโฯตโ\.\\displaystyle\\leq\\frac\{4\\lambda\}\{m\}\\\|\\bm\{\\epsilon\}\\\|\.ฯต\\epsilon\-insensitivity\.This follows straightforwardly from:
โSโPโrโO\+โ\(๐โฯต,๐\)\\displaystyle\\ell\_\{SPrO\+\}\(\\bm\{c\}\-\\bm\{\\epsilon\},\\bm\{c\}\)=\(ฯตโคโ\(๐๐นโฃโโ\(๐โฯต\)โ๐๐นโฃโโ\(๐\)\)โฮปโโ๐๐นโฃโโ\(๐โฯต\)โโ\+ฮปโโ๐๐นโฃโโ\(๐\)โโ\)\+\\displaystyle=\\bigg\(\\bm\{\\epsilon\}^\{\\top\}\(\\bm\{w^\{R\*\}\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\-\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\)\-\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\bigg\)\_\{\+\}=\(\(ฯตโคโ๐๐นโฃโโ\(๐โฯต\)\+ฮปโโ๐๐นโฃโโ\(๐\)โโ\)โ\(ฯตโคโ๐๐นโฃโโ\(๐\)\+ฮปโโ๐๐นโฃโโ\(๐โฯต\)โโ\)\)\+โค0\.\\displaystyle=\\bigg\(\(\\bm\{\\epsilon\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\)\-\(\\bm\{\\epsilon\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\bm\{c\}\-\\bm\{\\epsilon\}\)\\\|\_\{\*\}\)\\bigg\)\_\{\+\}\\leq 0\.\\halmos
###### Proof 7\.5
Proof of Theorem[3\.4](https://arxiv.org/html/2607.21773#S3.Thmtheorem4)\. Because themax\\maxand\(โ
\)\+\(\\cdot\)\_\{\+\}operators are subadditive, we know that
โSโPโrโO\+โ\(๐^,๐\)=\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle\\left\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\right\)\_\{\+\}โค\\displaystyle\\leq\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐\}โmin๐โ๐ฒRโฃโโ\(๐^\)โก\{๐^โคโ๐\+ฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle\\left\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\\\}\-\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\right\)\_\{\+\}โค\\displaystyle\\leq\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐\}โzRโฃโโ\(๐\)\)\+\+\(โmin๐โ๐ฒRโฃโโ\(๐^\)โก\{๐^โคโ๐\+ฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\)\+\\displaystyle\\left\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\\\}\-z^\{R\*\}\(\\bm\{c\}\)\\right\)\_\{\+\}\+\\left\(\-\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\right\)\_\{\+\}=\\displaystyle=โSโPโrโOโ\(๐^,๐\)โmin๐โ๐ฒRโฃโโ\(๐^\)โก\{๐^โคโ๐\+ฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\\displaystyle\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}=\\displaystyle=โSโPโrโOโ\(๐^,๐\)โ๐^โคโ๐ซโฮปโโ๐๐นโฒโ\(๐\)\+๐ซโโ\+ฮปโโ๐๐นโฒโ\(๐\)โโ,\\displaystyle\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{\\Delta\}\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\bm\{\\Delta\}\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\},where\(โ
\)\+\(\\cdot\)\_\{\+\}disappears because by definition,โmin๐ฐโ๐ฒRโฃโโ\(๐^\)โก\{๐^โคโ๐ฐ\+ฮปโโ๐ฐโโ\}\+๐^โคโ๐ฐ๐โฒโ\(๐\)\+ฮปโโ๐ฐ๐โฒโ\(๐\)โโโฅ0\-\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\+\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\geq 0\. By reverse triangle inequality, we obtain
โSโPโrโO\+โ\(๐^,๐\)โค\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leqโSโPโrโOโ\(๐^,๐\)โ๐^โคโ๐ซ\+ฮปโโ๐ซโโ\.\\displaystyle\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{\\Delta\}\+\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\.Now, letโs look at the expectations
๐ผโโ\[โSโPโrโO\+โ\(๐^,๐\)\]โค๐ผโโ\[โSโPโrโOโ\(๐^,๐\)\]โ๐ผโโ\[๐^\]โคโ๐ซ\+ฮปโโ๐ซโโ\.\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\\leq\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\-\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\hat\{\\bm\{c\}\}\]^\{\\top\}\\bm\{\\Delta\}\+\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\.Definingโ\\mathbb\{Q\}as a sub\-Gaussian probability measure, we are interested in
โโ\(ฮปโโ๐ซโโโ๐ผโ\[๐^\]โคโ๐ซ\>t\)=โโ\(\|ฮปโโ๐ซโโโ๐ผโ\[๐^\]โคโ๐ซ\|\>t\)=โโ\(\|supโ๐โโค1\(ฮปโ๐โ๐ผโ\[๐^\]\)โคโ๐ซ\|\>t\)\.\\displaystyle\\mathbb\{Q\}\(\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\-\\mathbb\{E\}\[\\hat\{\\bm\{c\}\}\]^\{\\top\}\\bm\{\\Delta\}\>t\)=\\mathbb\{Q\}\(\|\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\-\\mathbb\{E\}\[\\hat\{\\bm\{c\}\}\]^\{\\top\}\\bm\{\\Delta\}\|\>t\)=\\mathbb\{Q\}\(\|\\sup\_\{\\\|\\bm\{u\}\\\|\\leq 1\}\(\\lambda\\bm\{u\}\-\\mathbb\{E\}\[\\hat\{\\bm\{c\}\}\]\)^\{\\top\}\\bm\{\\Delta\}\|\>t\)\.The first equality is because we know thatฮปโโ๐ซโโโ๐ผโ\[๐^\]โคโ๐ซโฅ0\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\-\\mathbb\{E\}\[\\hat\{\\bm\{c\}\}\]^\{\\top\}\\bm\{\\Delta\}\\geq 0and the second equality is by definition of dual norm\. We therefore know that there exists๐ฎ\\bm\{u\}such thatโ๐ฎโโค1\\\|\\bm\{u\}\\\|\\leq 1andโโ\(\|\(ฮปโ๐ฎโ๐ผโ\[๐^\]\)โคโ๐ซ\|\>t\)\\mathbb\{Q\}\(\|\(\\lambda\\bm\{u\}\-\\mathbb\{E\}\[\\hat\{\\bm\{c\}\}\]\)^\{\\top\}\\bm\{\\Delta\}\|\>t\)\. The Hoeffding\-type inequality in Proposition 5\.10 inVershynin \([2012](https://arxiv.org/html/2607.21773#bib.bib8)\)produces the following concentration inequality:
โโ\(\|\(ฮปโ๐โ๐ผโ\[๐^\]\)โคโ๐ซ\|\>t\)โคexpโก\{1โC0โt2ฮบ2โโฮปโ๐โ๐ผโโ\[๐^\]โ22\},\\displaystyle\\mathbb\{Q\}\(\|\(\\lambda\\bm\{u\}\-\\mathbb\{E\}\[\\hat\{\\bm\{c\}\}\]\)^\{\\top\}\\bm\{\\Delta\}\|\>t\)\\leq\\exp\\left\\\{1\-\\frac\{C\_\{0\}t^\{2\}\}\{\\kappa^\{2\}\\\|\\lambda\\bm\{u\}\-\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\hat\{\\bm\{c\}\}\]\\\|\_\{2\}^\{2\}\}\\right\\\},whereC0\>0C\_\{0\}\>0is an absolute constant\. From the reverse triangle inequality of norms, we have
โโ\(\|\(ฮปโ๐โ๐ผโ\[๐^\]\)โคโ๐ซ\|\>t\)โคexpโก\{1โC0โt2ฮบ2โโฮปโ๐โ22\+โ๐ผโโ\[๐^\]โ22\}\.\\displaystyle\\mathbb\{Q\}\(\|\(\\lambda\\bm\{u\}\-\\mathbb\{E\}\[\\hat\{\\bm\{c\}\}\]\)^\{\\top\}\\bm\{\\Delta\}\|\>t\)\\leq\\exp\\left\\\{1\-\\frac\{C\_\{0\}t^\{2\}\}\{\\kappa^\{2\}\\\|\\lambda\\bm\{u\}\\\|\_\{2\}^\{2\}\+\\\|\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\hat\{\\bm\{c\}\}\]\\\|\_\{2\}^\{2\}\}\\right\\\}\.From the equivalence of norms and Jensenโs inequality,
โโ\(\|\(ฮปโ๐โ๐ผโ\[๐^\]\)โคโ๐ซ\|\>t\)โคexpโก\{1โC0โt2ฮบ2โฮปโC1\+๐ผโโ\[โ๐^โ22\]\}โคexpโก\{1โC0โt2ฮบ2โฮป2โC1\+C^2\}\.\\displaystyle\\mathbb\{Q\}\(\|\(\\lambda\\bm\{u\}\-\\mathbb\{E\}\[\\hat\{\\bm\{c\}\}\]\)^\{\\top\}\\bm\{\\Delta\}\|\>t\)\\leq\\exp\\left\\\{1\-\\frac\{C\_\{0\}t^\{2\}\}\{\\kappa^\{2\}\\lambda C\_\{1\}\+\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\\|\\hat\{\\bm\{c\}\}\\\|\_\{2\}^\{2\}\]\}\\right\\\}\\leq\\exp\\left\\\{1\-\\frac\{C\_\{0\}t^\{2\}\}\{\\kappa^\{2\}\\lambda^\{2\}C\_\{1\}\+\\hat\{C\}^\{2\}\}\\right\\\}\.The expectation bound is simply from Theorem 8 inBanerjeeet al\.\([2014](https://arxiv.org/html/2607.21773#bib.bib9)\), which states that
๐ผโโ\[T\]=๐ผโโ\[ฮปโโ๐ซโโ\]โคฮปโฮท0โฮบโฯโ\(โฌ\),\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{Q\}\}\[T\]=\\mathbb\{E\}\_\{\\mathbb\{Q\}\}\[\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\]\\leq\\lambda\\eta\_\{0\}\\kappa\\omega\(\\mathcal\{B\}\),whereฮท0\\eta\_\{0\}is a universal constant\. Using the property that the Gaussian width of a unit Euclidean ballโฌ\\mathcal\{B\}inโd\\mathbb\{R\}^\{d\}isOโ\(d\)O\(\\sqrt\{d\}\), the expectation bound simplifies to๐ผโโ\[T\]=Oโ\(ฮบโฮปโd\)\\mathbb\{E\}\_\{\\mathbb\{Q\}\}\[T\]=O\(\\kappa\\lambda\\sqrt\{d\}\)\.\\halmos
###### Proof 7\.6
Proof of Corollary[3\.6](https://arxiv.org/html/2607.21773#S3.Thmtheorem6)\. We establish Fisher consistency by proving thatโSโPโrโO\+\\ell\_\{SPrO\+\}isโ\\mathbb\{P\}\-calibrated with respect to the true robust decision lossโSโPโrโO\\ell\_\{SPrO\}\(according to Definition 3 inHo\-Nguyen and Kฤฑlฤฑnรง\-Karzan \([2022](https://arxiv.org/html/2607.21773#bib.bib11)\)\)\. Beingโ\\mathbb\{P\}\-calibrated means that for allฯต\>0\\epsilon\>0, there exists aฮด\>0\\delta\>0such that if๐\\bm\{B\}satisfies๐ผโโ\[โSโPโrโO\+โ\(๐โ๐ฑ,๐\)\]โmin๐โฒโก๐ผโโ\[โSโPโrโO\+โ\(๐โฒโ๐ฑ,๐\)\]<ฮด\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\+\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\)\]\-\\min\_\{\\bm\{B\}^\{\\prime\}\}\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\+\}\(\\bm\{B\}^\{\\prime\}\\bm\{x\},\\bm\{c\}\)\]<\\delta, then๐ผโโ\[โSโPโrโOโ\(๐โ๐ฑ,๐\)\]โmin๐โฒโก๐ผโโ\[โSโPโrโOโ\(๐โฒโ๐ฑ,๐\)\]<ฯต\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\)\]\-\\min\_\{\\bm\{B\}^\{\\prime\}\}\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\bm\{B\}^\{\\prime\}\\bm\{x\},\\bm\{c\}\)\]<\\epsilon\. From Theorem[3\.4](https://arxiv.org/html/2607.21773#S3.Thmtheorem4), we know that
๐ผโโ\[โSโPโrโO\+โ\(๐ฉโ๐,๐\)\]โmin๐ฉโฒโก๐ผโโ\[โSโPโrโO\+โ\(๐ฉโฒโ๐,๐\)\]<ฮด\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\+\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\)\]\-\\min\_\{\\bm\{B\}^\{\\prime\}\}\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\+\}\(\\bm\{B\}^\{\\prime\}\\bm\{x\},\\bm\{c\}\)\]<\\deltaโน\\displaystyle\\implies๐ผโโ\[โSโPโrโO\+โ\(๐ฉโ๐,๐\)\]โmin๐ฉโฒโก\{๐ผโโ\[โSโPโrโOโ\(๐ฉโฒโ๐,๐\)\+ฮปโโ๐ซโโโ๐ผโโ\[๐ฉโฒโ๐\]โคโ๐ซ\]\}<ฮด\.\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\+\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\)\]\-\\min\_\{\\bm\{B\}^\{\\prime\}\}\\\{\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\bm\{B\}^\{\\prime\}\\bm\{x\},\\bm\{c\}\)\+\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\-\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\bm\{B\}^\{\\prime\}\\bm\{x\}\]^\{\\top\}\\bm\{\\Delta\}\]\\\}<\\delta\.By definition,โSโPโrโO\+โ\(๐โ๐ฑ,๐\)โฅโSโPโrโOโ\(๐โ๐ฑ,๐\)\\ell\_\{SPrO\+\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\)\\geq\\ell\_\{SPrO\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\), which means that the above inequality implies
๐ผโโ\[โSโPโrโOโ\(๐ฉโ๐,๐\)\]โmin๐ฉโฒโก\{๐ผโโ\[โSโPโrโOโ\(๐ฉโฒโ๐,๐\)\+ฮปโโ๐ซโโโ๐ผโโ\[๐ฉโฒโ๐\]โคโ๐ซ\]\}<ฮด\.\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\)\]\-\\min\_\{\\bm\{B\}^\{\\prime\}\}\\\{\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\bm\{B\}^\{\\prime\}\\bm\{x\},\\bm\{c\}\)\+\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\-\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\bm\{B\}^\{\\prime\}\\bm\{x\}\]^\{\\top\}\\bm\{\\Delta\}\]\\\}<\\delta\.Since๐^\\hat\{\\bm\{c\}\}is centered,
๐ผโโ\[โSโPโrโOโ\(๐ฉโ๐,๐\)\]โmin๐ฉโฒโก๐ผโโ\[โSโPโrโOโ\(๐ฉโฒโ๐,๐\)\]โ<ฮด\+ฮปโฅโ๐ซโฅโ\.\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\bm\{B\}\\bm\{x\},\\bm\{c\}\)\]\-\\min\_\{\\bm\{B\}^\{\\prime\}\}\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\ell\_\{SPrO\}\(\\bm\{B\}^\{\\prime\}\\bm\{x\},\\bm\{c\}\)\]<\\delta\+\\lambda\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\.For a given toleranceฯต\>0\\epsilon\>0, the calibration relationshipฮดโ\(ฯต\)\>0\\delta\(\\epsilon\)\>0holds on the event thatโ๐ซโโโคฯตฮป\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\\leq\\frac\{\\epsilon\}\{\\lambda\}\. Now, we determine the probability ofฮดโ\(ฯต\)\>0\\delta\(\\epsilon\)\>0\. From Theorem 9 inBanerjeeet al\.\([2014](https://arxiv.org/html/2607.21773#bib.bib9)\), we know that
โโ\(โ๐ซโโ\>ฯตฮป\)โคฮฝ1โexpโก\{โ\(ฯตโฮปโฮฝ0โฮบโฯโ\(โฌ\)ฮปโฮฝ2โฮบโฯ\)2\},\\displaystyle\\mathbb\{Q\}\(\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\>\\frac\{\\epsilon\}\{\\lambda\}\)\\leq\\nu\_\{1\}\\exp\\left\\\{\-\\left\(\\frac\{\\epsilon\-\\lambda\\nu\_\{0\}\\kappa\\omega\(\\mathcal\{B\}\)\}\{\\lambda\\nu\_\{2\}\\kappa\\phi\}\\right\)^\{2\}\\right\\\},whereฮฝ0,ฮฝ1,ฮฝ2\\nu\_\{0\},\\nu\_\{1\},\\nu\_\{2\}are universal constants andฯ=supโ๐ฎโโค1โ๐ฎโ2\\phi=\\sup\_\{\\\|\\bm\{u\}\\\|\\leq 1\}\\\|\\bm\{u\}\\\|\_\{2\}\. Therefore the complement probability is,
โโ\(โ๐ซโโ<ฯตฮป\)โฅ1โฮฝ1โexpโก\{โ\(ฯตโฮปโฮฝ0โฮบโฯโ\(โฌ\)ฮปโฮฝ2โฮบโฯ\)2\}\.\\displaystyle\\mathbb\{Q\}\(\\\|\\bm\{\\Delta\}\\\|\_\{\*\}<\\frac\{\\epsilon\}\{\\lambda\}\)\\geq 1\-\\nu\_\{1\}\\exp\\left\\\{\-\\left\(\\frac\{\\epsilon\-\\lambda\\nu\_\{0\}\\kappa\\omega\(\\mathcal\{B\}\)\}\{\\lambda\\nu\_\{2\}\\kappa\\phi\}\\right\)^\{2\}\\right\\\}\.The worst\-case of this probability happens atlimฯตโ0\\lim\_\{\\epsilon\\to 0\}, which means that the probabilility ofโ\\mathbb\{P\}\-calibration is overall, at least
1โฮฝ1โexpโก\{โ\(โฮปโฮฝ0โฮบโฯโ\(โฌ\)ฮปโฮฝ2โฮบโฯ\)2\}=1โฮฝ1โexpโก\{โ\(ฮฝ0โฯโ\(โฌ\)ฯ\)2\}\.\\displaystyle 1\-\\nu\_\{1\}\\exp\\left\\\{\-\\left\(\\frac\{\-\\lambda\\nu\_\{0\}\\kappa\\omega\(\\mathcal\{B\}\)\}\{\\lambda\\nu\_\{2\}\\kappa\\phi\}\\right\)^\{2\}\\right\\\}=1\-\\nu\_\{1\}\\exp\\left\\\{\-\\left\(\\nu\_\{0\}\\frac\{\\omega\(\\mathcal\{B\}\)\}\{\\phi\}\\right\)^\{2\}\\right\\\}\.From Theorem 2 in\(Ho\-Nguyen and Kฤฑlฤฑnรง\-Karzan[2022](https://arxiv.org/html/2607.21773#bib.bib11)\), we know thatโ\\mathbb\{P\}\-calibration is equivalent toโ\\mathbb\{P\}\-Fisher consistency\.\\halmos
###### Lemma 7\.7\(Sudakov minoration theorem\)
Let๐ฆโโd\\mathcal\{K\}\\subset\\mathbb\{R\}^\{d\}be a compact set, and let๐ โผNโoโrโmโaโlโ\(๐,๐d\)\\bm\{g\}\\sim Normal\(\\bm\{0\},\\mathbb\{I\}\_\{d\}\)be a standard Gaussian vector\. Let๐ฉโ\(๐ฆ,ฯต\)\\mathcal\{N\}\(\\mathcal\{K\},\\epsilon\)be the covering number of๐ฆ\\mathcal\{K\}, defined as the minimum number of Euclidean balls of radiusฯต\\epsilonrequired to cover๐ฆ\\mathcal\{K\}\.There exists a universal constantC\>0C\>0such that for anyฯต\>0\\epsilon\>0:
ฯโ\(๐ฆ\)=๐ผโ\[sup๐โ๐ฆ๐โคโ๐\]โฅCโฯตโlogโก๐ฉโ\(๐ฆ,ฯต\)\\omega\(\\mathcal\{K\}\)=\\mathbb\{E\}\\left\[\\sup\_\{\\bm\{x\}\\in\\mathcal\{K\}\}\\bm\{x\}^\{\\top\}\\bm\{g\}\\right\]\\geq C\\epsilon\\sqrt\{\\log\\mathcal\{N\}\(\\mathcal\{K\},\\epsilon\)\}
###### Lemma 7\.8\(Sudakov\-Fernique inequality\)
Let๐=\{๐ฏ1,๐ฏ2,โฆ,๐ฏn\}\\bm\{V\}=\\\{\\bm\{v\}\_\{1\},\\bm\{v\}\_\{2\},\\dots,\\bm\{v\}\_\{n\}\\\}and๐=\{๐ฐ1,๐ฐ2,โฆ,๐ฐn\}\\bm\{W\}=\\\{\\bm\{w\}\_\{1\},\\bm\{w\}\_\{2\},\\dots,\\bm\{w\}\_\{n\}\\\}be two sets of vectors inโd\\mathbb\{R\}^\{d\}\. Let๐ โผNโoโrโmโaโlโ\(๐,๐d\)\\bm\{g\}\\sim Normal\(\\bm\{0\},\\mathbb\{I\}\_\{d\}\)be a standard Gaussian vector\. If for alli,jโ\[n\]i,j\\in\[n\], the Euclidean distances satisfy
โ๐iโ๐jโ2โคโ๐iโ๐jโ2,\\\|\\bm\{v\}\_\{i\}\-\\bm\{v\}\_\{j\}\\\|\_\{2\}\\leq\\\|\\bm\{w\}\_\{i\}\-\\bm\{w\}\_\{j\}\\\|\_\{2\},then
๐ผโ\[maxiโกโจ๐i,๐โฉ\]โค๐ผโ\[maxiโกโจ๐i,๐โฉ\]\.\\mathbb\{E\}\\left\[\\max\_\{i\}\\langle\\bm\{v\}\_\{i\},\\bm\{g\}\\rangle\\right\]\\leq\\mathbb\{E\}\\left\[\\max\_\{i\}\\langle\\bm\{w\}\_\{i\},\\bm\{g\}\\rangle\\right\]\.
###### Proof 7\.9
Proof of Theorem[4\.1](https://arxiv.org/html/2607.21773#S4.Thmtheorem1)\. The following bound can be derived
โSโPโrโO\+โ\(๐^,๐\)\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle=\\left\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\right\)\_\{\+\}โคmax๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzโโ\(๐\)\\displaystyle\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{\*\}\(\\bm\{c\}\)โคmax๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐\}โmin๐โ๐ฒRโฃโโ\(๐^\)โก\{ฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzโโ\(๐\)\\displaystyle\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\-\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{\*\}\(\\bm\{c\}\)=max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐\}โฮปโโ๐๐นโฒโ\(๐^\)โโ\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzโโ\(๐\)\\displaystyle=\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{\*\}\(\\bm\{c\}\)โคmax๐โ๐ฒโก\{๐โคโ๐โ๐^โคโ๐\}โzโโ\(๐\)โฮปโโ๐๐นโฒโ\(๐^\)โโ\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\\displaystyle\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\-z^\{\*\}\(\\bm\{c\}\)\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}โคโSโPโO\+โ\(๐^,๐\)โ๐^โคโ๐โโ\(๐\)โฮปโโ๐๐นโฒโ\(๐^\)โโ\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\.\\displaystyle\\leq\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\.The first inequality is becausezโโ\(๐\)โคzRโฃโโ\(๐\)z^\{\*\}\(\\bm\{c\}\)\\leq z^\{R\*\}\(\\bm\{c\}\)and the omission of\(โ
\)\+\(\\cdot\)\_\{\+\}is because by definition,๐โคโ๐ฐ๐โฃโโ\(๐^\)โฅzโโ\(๐\)\\bm\{c\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\\geq z^\{\*\}\(\\bm\{c\}\)and๐^โคโ๐ฐ๐โฃโโ\(๐^\)\+ฮปโโ๐ฐ๐โฃโโ\(๐^\)โโโค๐^โคโ๐ฐ๐โฒโ\(๐\)\+ฮปโโ๐ฐ๐โฒโ\(๐\)โโ\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\+\\lambda\\\|\\bm\{w^\{R\*\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\\leq\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\. The second inequality follows from the subadditivity of the maximization operator\. The second equality is because๐ฐ๐โฒโ\(๐^\)โargโกmax๐ฐโ๐ฒRโฃโโ\(๐^\)โก\{๐^โคโ๐ฐ\}=argโกmin๐ฐโ๐ฒRโฃโโ\(๐^\)โก\{โ๐ฐโโ\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\in\\arg\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\\}=\\arg\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\\{\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\. The third inequality is due to๐ฒRโฃโโ\(๐^\)โ๐ฒ\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\\subseteq\\mathcal\{W\}\. The fourth inequality is from the definition of SPO\+, i\.e\.โSโPโO\+โ\(๐^,๐\)=max๐ฐโ๐ฒโก\{๐โคโ๐ฐโ๐^โคโ๐ฐ\}\+๐^โคโ๐ฐโโ\(๐\)โzโโ\(๐\)\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\\max\_\{\\begin\{subarray\}\{c\}\\bm\{w\}\\in\\mathcal\{W\}\\end\{subarray\}\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\)and\(โ
\)\+\(\\cdot\)\_\{\+\}being subadditive\.
Suppose thatฮปโค๐โคโ๐ฐโโ\(๐\)โC^โโ๐ฐ๐โฒโ\(๐\)โโโ๐ฐโโ\(๐\)โโ\+โ๐ฐ๐โฒโ\(๐\)โโ\\lambda\\leq\\frac\{\\bm\{c\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\hat\{C\}\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\}\{\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\+\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\}\. Then
ฮปโ\(โ๐โโ\(๐\)โโ\+โ๐๐นโฒโ\(๐\)โโ\)โค๐โคโ๐โโ\(๐\)โC^โโ๐๐นโฒโ\(๐\)โโ\.\\displaystyle\\lambda\(\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\+\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\)\\leq\\bm\{c\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\hat\{C\}\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\.By Cauchy\-Schwarz inequality, this implies
ฮปโ\(โ๐โโ\(๐\)โโ\+โ๐๐นโฒโ\(๐\)โโ\)โค๐โคโ๐โโ\(๐\)โ๐^โคโ๐๐นโฒโ\(๐\)\\displaystyle\\lambda\(\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\+\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\)\\leq\\bm\{c\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)โ\\displaystyle\\Leftrightarrowฮปโ\(โ๐โโ\(๐\)โโ\+โ๐๐นโฒโ\(๐\)โโ\)โค๐^โคโ๐โโ\(๐\)\+\(๐โ๐^\)โคโ๐โโ\(๐\)โ๐^โคโ๐๐นโฒโ\(๐\)\\displaystyle\\lambda\(\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\+\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\)\\leq\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\+\(\\bm\{c\}\-\\hat\{\\bm\{c\}\}\)^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)โ\\displaystyle\\Leftrightarrowโ๐^โคโ๐โโ\(๐\)\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโ\(๐โ๐^\)โคโ๐โโ\(๐\)\+ฮปโโ๐โโ\(๐\)โโโค0\\displaystyle\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-\(\\bm\{c\}\-\\hat\{\\bm\{c\}\}\)^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\leq 0โน\\displaystyle\\impliesโ๐^โคโ๐โโ\(๐\)\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโโ๐โ๐^โโโ๐โโ\(๐\)โโ\+ฮปโโ๐โโ\(๐\)โโโค0\\displaystyle\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-\\\|\\bm\{c\}\-\\hat\{\\bm\{c\}\}\\\|\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\leq 0โน\\displaystyle\\impliesโ๐^โคโ๐โโ\(๐\)\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโค0\.\\displaystyle\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\leq 0\.Since this is true, then subtractingฮปโโ๐ฐ๐โฒโ\(๐^\)โโ\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}maintains the inequality, i\.e\.
โ๐^โคโ๐โโ\(๐\)โฮปโโ๐๐นโฒโ\(๐^\)โโ\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโค0,\\displaystyle\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\leq 0,which means thatโSโPโrโO\+โ\(๐^,๐\)โคโSโPโO\+โ\(๐^,๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\ell\_\{SPO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\.\\halmos
###### Proof 7\.10
Proof of Corollary[4\.2](https://arxiv.org/html/2607.21773#S4.Thmtheorem2)\. We start with the excess term from Theorem[4\.1](https://arxiv.org/html/2607.21773#S4.Thmtheorem1)\.
๐ผโโ\[โ๐^โคโ๐โโ\(๐\)โฮปโโ๐๐นโฒโ\(๐^\)โโ\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\\left\[\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\right\]Substituting๐ฐ๐โฒโ\(๐\)=๐ฐโโ\(๐\)\+๐ซโ\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)=\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\+\\bm\{\\Delta^\{\*\}\}, we obtain
๐ผโโ\[๐^โคโ๐ซโโฮปโโ๐๐นโฒโ\(๐^\)โโ\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]\.\\displaystyle\\mathbb\{E\}\_\{\\mathbb\{P\}\}\\left\[\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{\\Delta^\{\*\}\}\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\right\]\.If๐ซโ\\bm\{\\Delta^\{\*\}\}follows sub\-Gaussian probability measureโ\\mathbb\{Q\}, the we are interested in probabilistic dominance via
โโ\(๐ผโโ\[๐^โคโ๐ซโโฮปโโ๐๐นโฒโ\(๐^\)โโ\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]โค0\)\.\\displaystyle\\mathbb\{Q\}\\left\(\\mathbb\{E\}\_\{\\mathbb\{P\}\}\\left\[\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{\\Delta^\{\*\}\}\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\right\]\\leq 0\\right\)\.Let๐ผโโ\[โ๐ฐ๐โฒโ\(๐\)โโ\]โ๐ผโโ\[โ๐ฐ๐โฒโ\(๐^\)โโ\]โคโW\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]\-\\mathbb\{E\}\_\{\\mathbb\{P\}\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\]\\leq\-W, then
โโ\(๐ผโโ\[๐^โคโ๐ซโโฮปโโ๐๐นโฒโ\(๐^\)โโ\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]โค0\)\\displaystyle\\mathbb\{Q\}\\left\(\\mathbb\{E\}\_\{\\mathbb\{P\}\}\\left\[\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{\\Delta^\{\*\}\}\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\right\]\\leq 0\\right\)โฅโโ\(C^โโ๐ซโโโโฮปโWโค0\)\.\\displaystyle\\geq\\mathbb\{Q\}\\left\(\\hat\{C\}\\\|\\bm\{\\Delta^\{\*\}\}\\\|\_\{\*\}\-\\lambda W\\leq 0\\right\)\.By complementarity and Markov inequality,
โโ\(C^โโ๐ซโโโโฮปโWโค0\)\\displaystyle\\mathbb\{Q\}\\left\(\\hat\{C\}\\\|\\bm\{\\Delta^\{\*\}\}\\\|\_\{\*\}\-\\lambda W\\leq 0\\right\)=1โโโ\(C^โโ๐ซโโโโฮปโWโฅ0\)\\displaystyle=1\-\\mathbb\{Q\}\\left\(\\hat\{C\}\\\|\\bm\{\\Delta^\{\*\}\}\\\|\_\{\*\}\-\\lambda W\\geq 0\\right\)โฅ1โ๐ผโโ\[C^โโ๐ซโโโ\]ฮปโW\\displaystyle\\geq 1\-\\frac\{\\mathbb\{E\}\_\{\\mathbb\{Q\}\}\[\\hat\{C\}\\\|\\bm\{\\Delta^\{\*\}\}\\\|\_\{\*\}\]\}\{\\lambda W\}By Theorem 8 inBanerjeeet al\.\([2014](https://arxiv.org/html/2607.21773#bib.bib9)\), which states that if๐ซ\\bm\{\\Delta\}is sub\-Gaussian withโ๐ซโฯ2โคฮบ\\\|\\bm\{\\Delta\}\\\|\_\{\\psi\_\{2\}\}\\leq\\kappa, then๐ผโ\[โ๐ซโโ\]โคฮท0โฮบโฯโ\(โฌ\)\\mathbb\{E\}\[\\\|\\bm\{\\Delta\}\\\|\_\{\*\}\]\\leq\\eta\_\{0\}\\kappa\\omega\(\\mathcal\{B\}\), whereฮท0\\eta\_\{0\}is an absolute constant, we have the final bound
โโ\(๐ผโโ\[โ๐^โคโ๐โโ\(๐\)โฮปโโ๐๐นโฒโ\(๐^\)โโ\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]โค0\)โฅ1โC^โฮท0โฮบโฯโ\(โฌ\)ฮปโW\\displaystyle\\mathbb\{Q\}\\left\(\\mathbb\{E\}\_\{\\mathbb\{P\}\}\\left\[\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\\right\]\\leq 0\\right\)\\geq 1\-\\frac\{\\hat\{C\}\\eta\_\{0\}\\kappa\\omega\(\\mathcal\{B\}\)\}\{\\lambda W\}\\halmos
###### Proof 7\.11
Proof of Theorem[5\.2](https://arxiv.org/html/2607.21773#S5.Thmtheorem2)\. Let๐ซ1,๐ซ2โ๐ฒRโฃโโ\(๐\)\\bm\{r\}\_\{1\},\\bm\{r\}\_\{2\}\\in\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)and๐ฌ1,๐ฌ2โ๐ฒโโ\(๐\)\\bm\{s\}\_\{1\},\\bm\{s\}\_\{2\}\\in\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\), where๐ฒโ\\mathcal\{W\}^\{\*\}contains the optimal solutions generated from the robust prediction model\. By optimality, we know that
๐โคโ๐2\+ฮปโโ๐2โโโค๐โคโ๐1\+ฮปโโ๐1โโ\\displaystyle\\bm\{c\}^\{\\top\}\\bm\{r\}\_\{2\}\+\\lambda\\\|\\bm\{r\}\_\{2\}\\\|\_\{\*\}\\leq\\bm\{c\}^\{\\top\}\\bm\{s\}\_\{1\}\+\\lambda\\\|\\bm\{s\}\_\{1\}\\\|\_\{\*\}๐โคโ๐2\+ฮปโ๐ฒโคโ\(๐ฉ\)โ๐2โค๐โคโ๐1\+ฮปโ๐ฒโคโ\(๐ฉ\)โ๐1\.\\displaystyle\\bm\{c\}^\{\\top\}\\bm\{s\}\_\{2\}\+\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{s\}\_\{2\}\\leq\\bm\{c\}^\{\\top\}\\bm\{r\}\_\{1\}\+\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{r\}\_\{1\}\.Summing the two inequalities, we have
๐โคโ๐2\+ฮปโโ๐2โโ\+๐โคโ๐2\+ฮปโ๐ฒโคโ\(๐ฉ\)โ๐2โค๐โคโ๐1\+ฮปโโ๐1โโ\+๐โคโ๐1\+ฮปโ๐ฒโคโ\(๐ฉ\)โ๐1\\displaystyle\\bm\{c\}^\{\\top\}\\bm\{r\}\_\{2\}\+\\lambda\\\|\\bm\{r\}\_\{2\}\\\|\_\{\*\}\+\\bm\{c\}^\{\\top\}\\bm\{s\}\_\{2\}\+\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{s\}\_\{2\}\\leq\\bm\{c\}^\{\\top\}\\bm\{s\}\_\{1\}\+\\lambda\\\|\\bm\{s\}\_\{1\}\\\|\_\{\*\}\+\\bm\{c\}^\{\\top\}\\bm\{r\}\_\{1\}\+\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{r\}\_\{1\}โ\\displaystyle\\Leftrightarrowโโ๐โโค1:๐โคโ๐2\+ฮปโโ๐2โโโ๐โคโ๐1โฮปโ๐ฒโคโ\(๐ฉ\)โ๐1โค๐โคโ\(๐1โ๐2\)\+ฮปโ๐โคโ๐1โฮปโ๐ฒโคโ\(๐ฉ\)โ๐2\\displaystyle\\exists\\\|\\bm\{u\}\\\|\\leq 1:\\bm\{c\}^\{\\top\}\\bm\{r\}\_\{2\}\+\\lambda\\\|\\bm\{r\}\_\{2\}\\\|\_\{\*\}\-\\bm\{c\}^\{\\top\}\\bm\{r\}\_\{1\}\-\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{r\}\_\{1\}\\leq\\bm\{c\}^\{\\top\}\(\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\)\+\\lambda\\bm\{u\}^\{\\top\}\\bm\{s\}\_\{1\}\-\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{s\}\_\{2\}โ\\displaystyle\\Leftrightarrowโโ๐โโค1:ฮปโโ๐1โโโฮปโ๐ฒโคโ\(๐ฉ\)โ๐1โค๐โคโ\(๐1โ๐2\)\+ฮปโ๐โคโ๐1โฮปโ๐ฒโคโ\(๐ฉ\)โ๐2,\\displaystyle\\exists\\\|\\bm\{u\}\\\|\\leq 1:\\lambda\\\|\\bm\{r\}\_\{1\}\\\|\_\{\*\}\-\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{r\}\_\{1\}\\leq\\bm\{c\}^\{\\top\}\(\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\)\+\\lambda\\bm\{u\}^\{\\top\}\\bm\{s\}\_\{1\}\-\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{s\}\_\{2\},where the first equivalence is from the definition of dual norm and the second equivalence is from the optimality of๐ซ1,๐ซ2\\bm\{r\}\_\{1\},\\bm\{r\}\_\{2\}, which means that\(๐โ๐ฑ\)โคโ\(๐ซ2โ๐ซ1\)\+ฮปโ\(โ๐ซ2โโโโ๐ซ1โโ\)=0\(\\bm\{B\}\\bm\{x\}\)^\{\\top\}\(\\bm\{r\}\_\{2\}\-\\bm\{r\}\_\{1\}\)\+\\lambda\(\\\|\\bm\{r\}\_\{2\}\\\|\_\{\*\}\-\\\|\\bm\{r\}\_\{1\}\\\|\_\{\*\}\)=0\. By Cauchy\-Schwarz inequality, this implies
โโ๐โโค1:ฮปโโ๐1โโโฮปโ๐ฒโคโ\(๐ฉ\)โ๐1\+ฮปโ\(๐ฒโ\(๐ฉ\)โ๐\)โคโ๐1โค\(โ๐โ2\+ฮปโโ๐ฒโ\(๐ฉ\)โ2\)โโ๐1โ๐2โ2\.\\displaystyle\\exists\\\|\\bm\{u\}\\\|\\leq 1:\\lambda\\\|\\bm\{r\}\_\{1\}\\\|\_\{\*\}\-\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{r\}\_\{1\}\+\\lambda\(\\bm\{\\Lambda\}\(\\bm\{B\}\)\-\\bm\{u\}\)^\{\\top\}\\bm\{s\}\_\{1\}\\leq\(\\\|\\bm\{c\}\\\|\_\{2\}\+\\lambda\\\|\\bm\{\\Lambda\}\(\\bm\{B\}\)\\\|\_\{2\}\)\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|\_\{2\}\.By definition of dual norm, this further implies
ฮปโโ๐1โโโฮปโ๐ฒโคโ\(๐ฉ\)โ๐1\+ฮปโ๐ฒโคโ\(๐ฉ\)โ๐1โฮปโโ๐1โโโค\(โ๐โ2\+ฮปโโ๐ฒโ\(๐ฉ\)โ2\)โโ๐1โ๐2โ2\.\\displaystyle\\lambda\\\|\\bm\{r\}\_\{1\}\\\|\_\{\*\}\-\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{r\}\_\{1\}\+\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{s\}\_\{1\}\-\\lambda\\\|\\bm\{s\}\_\{1\}\\\|\_\{\*\}\\leq\(\\\|\\bm\{c\}\\\|\_\{2\}\+\\lambda\\\|\\bm\{\\Lambda\}\(\\bm\{B\}\)\\\|\_\{2\}\)\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|\_\{2\}\.From the definition of the minimum difference in regularizer gap, we obtain the chain
ฮปโ๐ขโ\(๐ฒSโฃโโ\(๐\),๐ฒRโฃโโ\(๐\)\)โคฮปโโ๐1โโโฮปโ๐ฒโคโ\(๐ฉ\)โ๐1\+ฮปโ๐ฒโคโ\(๐ฉ\)โ๐1โฮปโโ๐1โโโค\(โ๐โ2\+ฮปโโ๐ฒโ\(๐ฉ\)โ2\)โโ๐1โ๐2โ2\.\\displaystyle\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{S\*\}\(\\bm\{c\}\),\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\)\\leq\\lambda\\\|\\bm\{r\}\_\{1\}\\\|\_\{\*\}\-\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{r\}\_\{1\}\+\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{s\}\_\{1\}\-\\lambda\\\|\\bm\{s\}\_\{1\}\\\|\_\{\*\}\\leq\(\\\|\\bm\{c\}\\\|\_\{2\}\+\\lambda\\\|\\bm\{\\Lambda\}\(\\bm\{B\}\)\\\|\_\{2\}\)\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|\_\{2\}\.The condition of the theorem further extends the chain to
\(โ๐โ2\+ฮปโโ๐ฒโ\(๐ฉ\)โ2\)โ๐โ\(๐ฒRโฃโโ\(๐\)\)โคฮปโ๐ขโ\(๐ฒSโฃโโ\(๐\),๐ฒRโฃโโ\(๐\)\)โค\(โ๐โ2\+ฮปโโ๐ฒโ\(๐ฉ\)โ2\)โโ๐1โ๐2โ2,\\displaystyle\(\\\|\\bm\{c\}\\\|\_\{2\}\+\\lambda\\\|\\bm\{\\Lambda\}\(\\bm\{B\}\)\\\|\_\{2\}\)\\mathcal\{D\}\(\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\)\\leq\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{S\*\}\(\\bm\{c\}\),\\mathcal\{W\}^\{R\*\}\(\\bm\{c\}\)\)\\leq\(\\\|\\bm\{c\}\\\|\_\{2\}\+\\lambda\\\|\\bm\{\\Lambda\}\(\\bm\{B\}\)\\\|\_\{2\}\)\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|\_\{2\},which means that
โ๐1โ๐2โ2โคโ๐1โ๐2โ2\.\\displaystyle\\\|\\bm\{r\}\_\{1\}\-\\bm\{r\}\_\{2\}\\\|\_\{2\}\\leq\\\|\\bm\{s\}\_\{1\}\-\\bm\{s\}\_\{2\}\\\|\_\{2\}\.By Sudakov\-Fernique inequality \(Lemma[7\.7](https://arxiv.org/html/2607.21773#S7.Thmtheorem7)\), we therefore conclude that
๐ผ๐^โ\[ฯโ\(๐ฒRโฃโโ\(๐^\)\)\]โค๐ผ๐^โ\[ฯโ\(๐ฒโโ\(๐^๐น\)\)\],\\displaystyle\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\omega\(\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\)\]\\leq\\mathbb\{E\}\_\{\\hat\{\\bm\{c\}\}\}\[\\omega\(\\mathcal\{W\}^\{\*\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\}\)\)\],thus proving the theorem\.\\halmos
###### Proof 7\.12
Proof of Lemma 1 \(in main text\)\. BecausezRโฃโโ\(๐\)โฅzโโ\(๐\)z^\{R\*\}\(\\bm\{c\}\)\\geq z^\{\*\}\(\\bm\{c\}\), we know that
โSโPโrโO\+โ\(๐^,๐\)\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)=\(max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzRโฃโโ\(๐\)\)\+\\displaystyle=\\left\(\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{R\*\}\(\\bm\{c\}\)\\right\)\_\{\+\}โคmax๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzโโ\(๐\),\\displaystyle\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{\*\}\(\\bm\{c\}\),where\(โ
\)\+\(\\cdot\)\_\{\+\}vanishes because its argument is obviously non\-negative\. Let us now analyze the difference between SPrO\+ and SrPO\+\. By the above bound and the definition of SrPO\+,
โSโPโrโO\+โ\(๐^,๐\)โโSโrโPโO\+โ\(๐^๐น,๐\)โค\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\ell\_\{SrPO\+\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\\leqmax๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzโโ\(๐\)\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{\*\}\(\\bm\{c\}\)โmax๐โ๐ฒโก\{๐โคโ๐โ\(๐^\+ฮปโ๐ฒโ\(๐ฉ\)\)โคโ๐\}โ\(๐^\+ฮปโ๐ฒโ\(๐ฉ\)\)โคโ๐โโ\(๐\)\+zโโ\(๐\),\\displaystyle\-\\max\_\{\\begin\{subarray\}\{c\}\\bm\{w\}\\in\\mathcal\{W\}\\end\{subarray\}\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\(\\hat\{\\bm\{c\}\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w\}\\\}\-\(\\hat\{\\bm\{c\}\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\+z^\{\*\}\(\\bm\{c\}\),wherezโโ\(๐\)z^\{\*\}\(\\bm\{c\}\)will vanish\. By the definition of๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\), we can further express this upper bound as
โSโPโrโO\+โ\(๐^,๐\)โโSโrโPโO\+โ\(๐^๐น,๐\)โค\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\ell\_\{SrPO\+\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\\leqmax๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}โmax๐โ๐ฒโก\{๐โคโ๐โ\(๐^\+ฮปโ๐ฒโ\(๐ฉ\)\)โคโ๐\}โ\(๐^\+ฮปโ๐ฒโ\(๐ฉ\)\)โคโ๐โโ\(๐\)โฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\\displaystyle\-\\max\_\{\\begin\{subarray\}\{c\}\\bm\{w\}\\in\\mathcal\{W\}\\end\{subarray\}\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\(\\hat\{\\bm\{c\}\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w\}\\\}\-\(\\hat\{\\bm\{c\}\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\+ฮปโmin๐โ๐ฒRโฒโ\(๐\)โก\{๐ฒโคโ\(๐ฉ\)โ๐โโ๐โโ\}โฮปโmax๐โ๐ฒโโ\(๐\)โก\{๐ฒโคโ\(๐ฉ\)โ๐โโ๐โโ\}\.\\displaystyle\+\\lambda\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\)\}\\\{\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w\}\-\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\-\\lambda\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\}\\\{\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w\}\-\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\.Becausemin๐ฐโ๐ฒRโฒโ\(๐\)โก\{๐ฒโคโ\(๐\)โ๐ฐโโ๐ฐโโ\}โค๐ฒโคโ\(๐\)โ๐ฐ๐โฒโ\(๐\)โโ๐ฐ๐โฒโ\(๐\)โโ\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\)\}\\\{\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w\}\-\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\\leq\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\-\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}, we then have the upper bound
โSโPโrโO\+โ\(๐^,๐\)โโSโrโPโO\+โ\(๐^๐น,๐\)โค\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\ell\_\{SrPO\+\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\\leqmax๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)โmax๐โ๐ฒโก\{๐โคโ๐โ\(๐^\+ฮปโ๐ฒโ\(๐ฉ\)\)โคโ๐\}โ\(๐^\+ฮปโ๐ฒโ\(๐ฉ\)\)โคโ๐โโ\(๐\)โฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\\displaystyle\-\\max\_\{\\begin\{subarray\}\{c\}\\bm\{w\}\\in\\mathcal\{W\}\\end\{subarray\}\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\(\\hat\{\\bm\{c\}\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w\}\\\}\-\(\\hat\{\\bm\{c\}\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\)\-\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\+ฮปโ๐ฒโคโ\(๐ฉ\)โ๐๐นโฒโ\(๐\)โฮปโmax๐โ๐ฒโโ\(๐\)โก\{๐ฒโคโ\(๐ฉ\)โ๐โโ๐โโ\}\.\\displaystyle\+\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\-\\lambda\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\}\\\{\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w\}\-\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\.Sincemax๐ฐโ๐ฒโก\{๐โคโ๐ฐโ\(๐^\+ฮปโ๐ฒโ\(๐\)\)โคโ๐ฐ\}โฅzโโ\(๐\)โ\(๐โ๐ฑ\+ฮปโ๐ฒโ\(๐\)\)โคโ๐ฐโโ\(๐\)\\max\_\{\\begin\{subarray\}\{c\}\\bm\{w\}\\in\\mathcal\{W\}\\end\{subarray\}\}\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\(\\hat\{\\bm\{c\}\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w\}\\\}\\geq z^\{\*\}\(\\bm\{c\}\)\-\(\\bm\{B\}\\bm\{x\}\+\\lambda\\bm\{\\Lambda\}\(\\bm\{B\}\)\)^\{\\top\}\\bm\{w^\{\*\}\}\(\\bm\{c\}\), we further simplify this upper bound to
โSโPโrโO\+โ\(๐^,๐\)โโSโrโPโO\+โ\(๐^๐น,๐\)โค\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\ell\_\{SrPO\+\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\\leqmax๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)โzโโ\(๐\)โฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\)\-\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\+ฮปโ๐ฒโคโ\(๐ฉ\)โ๐๐นโฒโ\(๐\)โฮปโmax๐โ๐ฒโโ\(๐\)โก\{๐ฒโคโ\(๐ฉ\)โ๐โโ๐โโ\}\.\\displaystyle\+\\lambda\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\-\\lambda\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\}\\\{\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w\}\-\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\.Becausemax๐ฐโ๐ฒโโ\(๐\)โก\{๐ฒโคโ\(๐\)โ๐ฐโโ๐ฐโโ\}โฅ๐ฒโคโ\(๐\)โ๐ฐ๐โฒโ\(๐\)โโ๐ฐ๐โฒโ\(๐\)โโ\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\}\\\{\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w\}\-\\\|\\bm\{w\}\\\|\_\{\*\}\\\}\\geq\\bm\{\\Lambda\}^\{\\top\}\(\\bm\{B\}\)\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\-\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}, we see that
โSโPโrโO\+โ\(๐^,๐\)โโSโrโPโO\+โ\(๐^๐น,๐\)โค\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-\\ell\_\{SrPO\+\}\(\\hat\{\\bm\{c\}\}^\{\\bm\{R\}\},\\bm\{c\}\)\\leqmax๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโโzโโ\(๐\)\\displaystyle\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{\*\}\(\\bm\{c\}\)โฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\\displaystyle\-\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)=\\displaystyle=โSโPโrโO\+โ\(๐^,๐\)โzRโฃโโ\(๐\)โzโโ\(๐\)โฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\),\\displaystyle\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\-z^\{R\*\}\(\\bm\{c\}\)\-z^\{\*\}\(\\bm\{c\}\)\-\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\),thus proving the lemma\.\\halmos
###### Proof 7\.13
Proof of Theorem[5\.4](https://arxiv.org/html/2607.21773#S5.Thmtheorem4)\. From the previous lemma, we therefore know that stochastic dominance is achieved if
๐ผโ\[โSโPโrโO\+โ\(๐^,๐\)\]โค๐ผโ\[zRโฃโโ\(๐\)\+zโโ\(๐\)\+ฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\]\.\\displaystyle\\mathbb\{E\}\[\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\]\\leq\\mathbb\{E\}\[z^\{R\*\}\(\\bm\{c\}\)\+z^\{\*\}\(\\bm\{c\}\)\+\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\]\.SinceโSโPโrโO\+โ\(๐^,๐\)โคmax๐ฐโ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐ฐโ๐^โคโ๐ฐโฮปโโ๐ฐโโ\}\+๐^โคโ๐ฐ๐โฒโ\(๐\)\+ฮปโโ๐ฐ๐โฒโ\(๐\)โโโzโโ\(๐\)\\ell\_\{SPrO\+\}\(\\hat\{\\bm\{c\}\},\\bm\{c\}\)\\leq\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\-z^\{\*\}\(\\bm\{c\}\), the condition is met if
๐ผโ\[max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐โ๐^โคโ๐โฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]โค๐ผโ\[zRโฃโโ\(๐\)\+2โzโโ\(๐\)\+ฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\]\.\\displaystyle\\mathbb\{E\}\[\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\-\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\-\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]\\leq\\mathbb\{E\}\[z^\{R\*\}\(\\bm\{c\}\)\+2z^\{\*\}\(\\bm\{c\}\)\+\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\]\.By subadditivity of the maximization operator and superadditivity of the minimization operator, the condition is also met if
๐ผโ\[max๐โ๐ฒRโฃโโ\(๐^\)โก\{๐โคโ๐\}โmin๐โ๐ฒRโฃโโ\(๐^\)โก\{๐^โคโ๐\}โmin๐โ๐ฒRโฃโโ\(๐^\)โก\{ฮปโโ๐โโ\}\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]\\displaystyle\\mathbb\{E\}\[\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\\big\\\}\-\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\-\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}^\{R\*\}\(\\hat\{\\bm\{c\}\}\)\}\\big\\\{\\lambda\\\|\\bm\{w\}\\\|\_\{\*\}\\big\\\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]โค\\displaystyle\\leq๐ผโ\[max๐โ๐ฒโก\{๐โคโ๐\}โmin๐โ๐ฒโก\{๐^โคโ๐\}โฮปโโ๐๐นโฒโ\(๐^\)โโ\+๐^โคโ๐๐นโฒโ\(๐\)\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]\\displaystyle\\mathbb\{E\}\[\\max\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\big\\\{\\bm\{c\}^\{\\top\}\\bm\{w\}\\big\\\}\-\\min\_\{\\bm\{w\}\\in\\mathcal\{W\}\}\\big\\\{\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w\}\\big\\\}\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{\\bm\{c\}\}^\{\\top\}\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]โค\\displaystyle\\leq๐ผโ\[zRโฃโโ\(๐\)\+2โzโโ\(๐\)\+ฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\]\.\\displaystyle\\mathbb\{E\}\[z^\{R\*\}\(\\bm\{c\}\)\+2z^\{\*\}\(\\bm\{c\}\)\+\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\]\.By Cauchy\-Schwarz inequality and the fact thatzโโ\(โ
\)โฅ0z^\{\*\}\(\\cdot\)\\geq 0, we continue the argument to obtain
๐ผโ\[โฮปโโ๐๐นโฒโ\(๐^\)โโ\+C^โโ๐๐นโฒโ\(๐\)โโ\+ฮปโโ๐๐นโฒโ\(๐\)โโ\]\\displaystyle\\mathbb\{E\}\[\-\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\+\\hat\{C\}\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\+\\lambda\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]โค\\displaystyle\\leq๐ผโ\[zRโฃโโ\(๐\)\+2โzโโ\(๐\)\+ฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)โzยฏโโ\(๐\)\]\.\\displaystyle\\mathbb\{E\}\[z^\{R\*\}\(\\bm\{c\}\)\+2z^\{\*\}\(\\bm\{c\}\)\+\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\-\\bar\{z\}^\{\*\}\(\\bm\{c\}\)\]\.This leads to
\(C^\+ฮป\)โ๐ผโ\[โ๐๐นโฒโ\(๐\)โโ\]โคฮปโ๐ผโ\[โ๐๐นโฒโ\(๐^\)โโ\]\+๐ผโ\[zRโฃโโ\(๐\)\+zโโ\(๐\)\+ฮปโ๐ขโ\(๐ฒRโฒโ\(๐\),๐ฒโโ\(๐\)\)\],\\displaystyle\(\\hat\{C\}\+\\lambda\)\\mathbb\{E\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\bm\{c\}\)\\\|\_\{\*\}\]\\leq\\lambda\\mathbb\{E\}\[\\\|\\bm\{w^\{R^\{\\prime\}\}\}\(\\hat\{\\bm\{c\}\}\)\\\|\_\{\*\}\]\+\\mathbb\{E\}\[z^\{R\*\}\(\\bm\{c\}\)\+z^\{\*\}\(\\bm\{c\}\)\+\\lambda\\mathcal\{G\}\(\\mathcal\{W\}^\{R^\{\\prime\}\}\(\\bm\{c\}\),\\mathcal\{W\}^\{\*\}\(\\bm\{c\}\)\)\],thus proving the theorem\.\\halmosSimilar Articles
Halt Fast! Early Stopping for Certified Robustness
This paper introduces a meta-learning framework for anytime-valid certified robustness that uses sequential E-processes to adaptively allocate compute, achieving a 20-fold reduction in sample complexity compared to traditional randomized smoothing while maintaining rigorous statistical guarantees.
Maximally Robust Satisficing Bayesian Optimization
This paper introduces Maximally Robust Satisficing Bayesian Optimization (MRSBO), a method that efficiently finds solutions meeting a quality threshold while being robust to input perturbations after deployment, outperforming previous approaches.
Position Paper: Post-Solve Robustness in Decision Engines: Feasible Regions and Smoothness Under Perturbations
Position paper arguing for a post-solve robustness layer for MILP decision engines, formalizing feasible neighborhoods and solution smoothness under perturbations, and calling for certified inner approximations and adversarial robustness margins.
Learning Predictive Ambiguity Sets for Decision-Focused Distributionally Robust Optimization
Proposes learned predictive ambiguity sets (LPAS) for distributionally robust optimization, where a deep contextual model outputs a nominal scenario distribution, state-dependent Wasserstein radius, and ground metric, trained with decision loss and calibration. Applied to portfolio optimization on S&P 500 data, the method achieves higher returns and Sharpe ratio with reduced conservatism compared to fixed-radius baselines.
Optimized Instance Alteration for Explaining and Assessing Robustness of Classifiers
This paper proposes a unified optimization framework to explain misclassifications and assess classifier robustness by sparse, interpretable instance alterations and a Tolerance Region Confusion Matrix.