Resource-Adaptive Primal-Dual Learning for One-Warehouse Multi-Store Systems with Censored Demand
Summary
This paper introduces Resource-Adaptive Primal-Dual Learning for one-warehouse multi-store inventory systems with censored demand, achieving logarithmic regret by dynamically adapting to changing resource levels.
View Cached Full Text
Cached at: 08/17/26, 10:19 AM
# Resource-Adaptive Primal-Dual Learning for One-Warehouse Multi-Store Systems with Censored Demand
Source: [https://arxiv.org/html/2608.14096](https://arxiv.org/html/2608.14096)
###### Abstract
The one\-warehouse multi\-store \(OWMS\) system is a fundamental inventory network in which a nonreplenishable warehouse allocates shared stock across multiple stores over time\. Existing OWMS learning policies are built around a fixed target calibrated to the initial average resource rate, but such a fixed\-target architecture cannot re\-center after realized sales change the remaining resource available per future period\. We develop Resource\-Adaptive Primal\-Dual Learning, a new learning framework that tracks the Primal\-Dual re\-solving path with censored demand as the remaining\-resource state evolves\. In each period, the current resource rate indexes the target store allocations and dual variable, while censored sales provide gradient estimates for updating both\. The analysis combines expected\-sales geometry with a moving\-target argument to yield logarithmic expected regret, improving on the state\-of\-the\-art square\-root\-order guarantees of existing OWMS learning policies\. The underlying design and analytical ideas may inform other online learning problems with depleting shared resources\. Numerical experiments further demonstrate good finite\-horizon performance of a practical variant across different horizon lengths and inventory regimes\.
###### keywords
one\-warehouse multi\-store systems; inventory learning; censored demand; Primal\-Dual methods; resource\-adaptive learning; logarithmic regret
††runningtitle:Resource\-Adaptive Primal\-Dual Learning for OWMS Systems††runningauthor:Lyu††authors:Department of Management Science, School of Management, Fudan University, Shanghai 200433, China, jiamenglyu@fudan\.edu\.cn††affiliation:††affiliation:## 1Introduction
The one\-warehouse multi\-store \(OWMS\) system is a basic architecture of inventory distribution, which pools finite inventory at a central warehouse and allocates it across stores and time\. Each shipment trades current service at one location against the option of serving another later, creating a systemwide dynamic allocation problem\([16](https://arxiv.org/html/2608.14096#bib.bib1),[6](https://arxiv.org/html/2608.14096#bib.bib5)\)\. Applications include general\-merchandise distribution, perishable\-product distribution\([13](https://arxiv.org/html/2608.14096#bib.bib6)\), fast\-fashion distribution\([4](https://arxiv.org/html/2608.14096#bib.bib2),[2](https://arxiv.org/html/2608.14096#bib.bib19)\), service\-parts logistics\([19](https://arxiv.org/html/2608.14096#bib.bib32)\), emergency\-medical supply allocation\([25](https://arxiv.org/html/2608.14096#bib.bib33)\)\.
This paper focuses on a finite\-horizon version in which the warehouse receives an initial stock and ships it to stores with different demand distributions and economic margins\. Each allocation changes both tomorrow’s inventory state and the opportunity cost of every future unit\. Even with known demand laws, the exact dynamic program is high dimensional and strongly state dependent\. With unknown distributions and censored feedback, the manager must learn demand responses while preserving the same inventory needed to earn future reward\.
For newsvendor and related inventory problems, a positive density lower bound supplies the strong convexity that converts stochastic first\-order learning into logarithmic regret\([15](https://arxiv.org/html/2608.14096#bib.bib11),[31](https://arxiv.org/html/2608.14096#bib.bib12),[32](https://arxiv.org/html/2608.14096#bib.bib13),[23](https://arxiv.org/html/2608.14096#bib.bib16)\)\. These models effectively draw replenishment from an unconstrained upstream source\. With a nonreplenishable warehouse, by contrast, all stores and periods compete for the same finite inventory pool\. A sale at one store irreversibly reduces the resource available to every store in subsequent periods\. The best existing finite\-stock censored\-demand guarantees are square\-root order for OWMS and its multiwarehouse extension\([2](https://arxiv.org/html/2608.14096#bib.bib19),[28](https://arxiv.org/html/2608.14096#bib.bib22)\)\.
Re\-solving is a powerful and well\-received organizing principle for dynamic resource allocation in both academia and industries\. It reveals how allocations and scarcity prices should adapt to an evolving resource state, but learning work rarely treats the resulting Primal\-Dual path itself as the object to be tracked\. We develop a new resource\-adaptive Primal\-Dual learning framework that replaces fixed\-target learning with online tracking of the endogenous Primal\-Dual re\-solving path using censored demand\. This framework yields a logarithmic regret guarantee for the OWMS learning problem, and its design and analytical tools may inform other online learning\-and\-control problems with censored feedback and depleting shared resources\.
### 1\.1Contributions
Our contributions are summarized as follows\.
1. 1\.Resource\-adaptive Primal\-Dual learning framework\.We introduce Resource\-Adaptive Primal\-Dual Learning \(𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}\), a new framework that replaces the fixed target calibrated to the initial average resource rate with online tracking of the endogenous Primal\-Dual path generated by resource\-state re\-solving\.𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}uses remaining system inventory per remaining period to re\-index the fluid KKT target—the preferred store allocations and their common scarcity price\. Censored sales provide observable stochastic Primal\-Dual directions for updating this target and simultaneously update the resource state that re\-centers the next target\. In this way,𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}tracks the allocation\-and\-price path that known\-distribution fluid re\-solving would produce, without estimating the demand distributions or repeatedly solving a fluid program\.
2. 2\.Analysis of endogenous moving targets\.Our analysis separates regret into three components\. Value\-function smoothness controls the intrinsic gap of the resource\-indexed re\-solving path\. Expected\-sales geometry and a moving\-target argument control the gap from learning that path\. The inventory dynamics control the gap from implementing the learned state as a feasible physical action\. Each component is logarithmic, yielding anO\(logT\)O\(\\log T\)regret guarantee for𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}and improving on the state\-of\-the\-art square\-root\-order guarantees in the finite\-stock censored\-demand OWMS literature\.
3. 3\.Numerical evaluation\.We compare our resource\-adaptive learning method𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}with the Double Binary Search policy of[2](https://arxiv.org/html/2608.14096#bib.bib19), a state\-of\-the\-art benchmark across different inventory regimes\. Across both homogeneous and heterogeneous experiments,𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}has lower mean total cost in all 24 paired comparisons, providing short\-horizon evidence for our algorithm\.
### 1\.2Related Literature
Re\-solving and Primal\-Dual learning with unreplenishable resources\.A broad revenue\-management and online\-allocation literature feeds observed resource consumption back into future decisions\. Re\-solving policies recompute a fluid program with the remaining capacity and horizon and achieve sublinear guarantees under suitable structural conditions\([17](https://arxiv.org/html/2608.14096#bib.bib25),[18](https://arxiv.org/html/2608.14096#bib.bib24)\)\. These models assume known demand distributions and thus involve no learning\. Related online\-LP work combines learning with repeated LP optimization under fully observed requests\([1](https://arxiv.org/html/2608.14096#bib.bib23),[22](https://arxiv.org/html/2608.14096#bib.bib26),[21](https://arxiv.org/html/2608.14096#bib.bib30)\)\. These methods learn dual prices from observed arrivals through empirical LP solves\.
Primal\-Dual demand learning treats the fluid primal and dual solutions themselves as unknown\. Related approaches include explicit learning of an inventory shadow price in personalized pricing\([7](https://arxiv.org/html/2608.14096#bib.bib20)\), and nonparametric or large\-action\-space network revenue management\([9](https://arxiv.org/html/2608.14096#bib.bib27),[29](https://arxiv.org/html/2608.14096#bib.bib28),[30](https://arxiv.org/html/2608.14096#bib.bib21),[27](https://arxiv.org/html/2608.14096#bib.bib29)\)\. These approaches combine demand learning, dual\-price learning, and resource feedback through different architectures\. In the finite\-stock OWMS setting,𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}instead uses censored sales to jointly track store\-level allocation targets and their common scarcity price along the KKT path indexed by remaining system inventory per remaining period\.
OWMS control and learning\.The finite\-stock OWMS problem originates in the two\-echelon allocation setting of[16](https://arxiv.org/html/2608.14096#bib.bib1)\. With known demand laws, the literature develops tractable decomposition and Lagrangian policies for OWMS problem\([24](https://arxiv.org/html/2608.14096#bib.bib7),[26](https://arxiv.org/html/2608.14096#bib.bib8)\)\. Within this known\-demand control literature, the policy of[5](https://arxiv.org/html/2608.14096#bib.bib9)adaptively readjusts its Lagrangian parameters using realized demand history\.
The closest learning papers retain that finite shared stock and explicitly learn its scarcity price\.[2](https://arxiv.org/html/2608.14096#bib.bib19)proposes Double Binary Search for OWMS, and[28](https://arxiv.org/html/2608.14096#bib.bib22)combines dual cutting planes with primal learning in a multiwarehouse extension\. Both learn a fixed Lagrangian target and obtain the best available square\-root\-order regret guarantees for this class\. By contrast,𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}learns a resource\-indexed Primal\-Dual path and re\-centers that path after realized sales change the remaining resource state\.
Gradient\-based inventory learning\.For basic censored\-demand inventory models, sales and stockout observations can provide first\-order information for gradient\-based learning\([15](https://arxiv.org/html/2608.14096#bib.bib11),[31](https://arxiv.org/html/2608.14096#bib.bib12),[23](https://arxiv.org/html/2608.14096#bib.bib16),[12](https://arxiv.org/html/2608.14096#bib.bib18)\)\. Extensions cover positive lead times, perishable inventory, random capacities, and multi\-retailer systems\([14](https://arxiv.org/html/2608.14096#bib.bib10),[32](https://arxiv.org/html/2608.14096#bib.bib13),[33](https://arxiv.org/html/2608.14096#bib.bib15),[8](https://arxiv.org/html/2608.14096#bib.bib14),[11](https://arxiv.org/html/2608.14096#bib.bib17)\)\.[23](https://arxiv.org/html/2608.14096#bib.bib16)applies a minibatch\-SGD metapolicy to a two\-echelon system in which the warehouse itself can replenish, whereas[13](https://arxiv.org/html/2608.14096#bib.bib6)develops an offline, feature\-based approximation for a perishable network in which both the warehouse and retailers replenish\.
Organization\.[Section2](https://arxiv.org/html/2608.14096#S2)presents the physical system, exact comparator, fluid benchmark\.[Section3](https://arxiv.org/html/2608.14096#S3)derives and states𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}\.[Section4](https://arxiv.org/html/2608.14096#S4)gives the main guarantee and develops its regret analysis, including the expected\-sales geometry used only in the analysis\.[Section5](https://arxiv.org/html/2608.14096#S5)reports the computational study, and[Section6](https://arxiv.org/html/2608.14096#S6)concludes the paper\. The appendices collect the proofs of all propositions and lemmas stated in the main text, the supporting technical verifications, and the extension sketches\.
Notation\.Throughout,T≥2T\\geq 2,log\\logdenotes the natural logarithm, andC<∞C<\\inftydenotes a constant independent ofTTthat may change from line to line\. Letx\+:=max\{x,0\}x^\{\+\}:=\\max\\\{x,0\\\}\.
## 2Problem Formulation
There are storesi∈\[N\]:=\{1,…,N\}i\\in\[N\]:=\\\{1,\\ldots,N\\\}and periodst∈\[T\]:=\{1,…,T\}t\\in\[T\]:=\\\{1,\\ldots,T\\\}\. Immediately before the period\-ttshipment,BtB\_\{t\}is the stock physically remaining at the warehouse andIi,tI\_\{i,t\}is the on\-hand stock already located at storeii\. The decisionYi,tY\_\{i,t\}is storeii’s stock*after*that shipment\. The warehouse begins with a nonreplenishable stockB1=W=γTB\_\{1\}=W=\\gamma T, and every store begins empty,Ii,1=0I\_\{i,1\}=0\. Unless stated otherwise, vector norms are Euclidean\. At the beginning of periodtt, a feasible post\-shipment vector satisfies
Yi,t≥Ii,t,∑i=1N\(Yi,t−Ii,t\)≤Bt\.Y\_\{i,t\}\\geq I\_\{i,t\},\\qquad\\sum\_\{i=1\}^\{N\}\(Y\_\{i,t\}\-I\_\{i,t\}\)\\leq B\_\{t\}\.\(2\.1\)Demands, sales, and states obey
Si,t=Di,t∧Yi,t,Ii,t\+1=Yi,t−Si,t,Bt\+1=Bt−∑i=1N\(Yi,t−Ii,t\)\.S\_\{i,t\}=D\_\{i,t\}\\wedge Y\_\{i,t\},\\quad I\_\{i,t\+1\}=Y\_\{i,t\}\-S\_\{i,t\},\\quad B\_\{t\+1\}=B\_\{t\}\-\\sum\_\{i=1\}^\{N\}\(Y\_\{i,t\}\-I\_\{i,t\}\)\.\(2\.2\)In each period, the manager observes the state\(Bt,𝑰t\)\(B\_\{t\},\\bm\{I\}\_\{t\}\)and past sales, chooses𝒀t\\bm\{Y\}\_\{t\}before demand, observes only censored sales𝑺t\\bm\{S\}\_\{t\}, and then updates the state by \([2\.2](https://arxiv.org/html/2608.14096#S2.E2)\)\. Unmet demand is lost\.
We impose the following conditions on the primitive uncertainty\.
###### Assumption 2\.1\.
1. \(i\)The demand vectors\{𝑫t\}t≥1\\\{\\bm\{D\}\_\{t\}\\\}\_\{t\\geq 1\}are i\.i\.d\. over time\. Within each period, their coordinates may be arbitrarily dependent, and the marginal distribution ofDi,tD\_\{i,t\}isFiF\_\{i\}\.
2. \(ii\)The lawFiF\_\{i\}is absolutely continuous on its support\[0,d¯i\]\[0,\\bar\{d\}\_\{i\}\], satisfiesFi\(x\)=1F\_\{i\}\(x\)=1forx≥d¯ix\\geq\\bar\{d\}\_\{i\}, and has a continuous density obeying the known bounds 0<κi≤fi\(x\)≤Ki<∞,0≤x≤d¯i\.0<\\kappa\_\{i\}\\leq f\_\{i\}\(x\)\\leq K\_\{i\}<\\infty,\\qquad 0\\leq x\\leq\\bar\{d\}\_\{i\}\.\(2\.3\)
WriteD¯:=∑i=1Nd¯i\\overline\{D\}:=\\sum\_\{i=1\}^\{N\}\\bar\{d\}\_\{i\}and suppose that the initial resource rate satisfies0≤γ≤D¯0\\leq\\gamma\\leq\\overline\{D\}\.
For a policyπ\\pi, let𝒰π\\mathcal\{U\}^\{\\pi\}denote any exogenous policy randomization, independent of demand, and define the pre\-demand history by
ℋtπ:=σ\(𝒰π,B1,𝑰1,\(𝒀τ,𝑺τ,Bτ\+1,𝑰τ\+1\)τ<t\)\.\\mathcal\{H\}\_\{t\}^\{\\pi\}:=\\sigma\\\!\\left\(\\mathcal\{U\}^\{\\pi\},B\_\{1\},\\bm\{I\}\_\{1\},\(\\bm\{Y\}\_\{\\tau\},\\bm\{S\}\_\{\\tau\},B\_\{\\tau\+1\},\\bm\{I\}\_\{\\tau\+1\}\)\_\{\\tau<t\}\\right\)\.Under[Section2](https://arxiv.org/html/2608.14096#S2),𝑫t\\bm\{D\}\_\{t\}is independent ofℋtπ\\mathcal\{H\}\_\{t\}^\{\\pi\}\. LetΠad\\Pi\_\{\\rm ad\}denote the class of policies that, from the stated initial state, choose𝒀t\\bm\{Y\}\_\{t\}measurably with respect toℋtπ\\mathcal\{H\}\_\{t\}^\{\\pi\}, satisfy \([2\.1](https://arxiv.org/html/2608.14096#S2.E1)\), and evolve according to \([2\.2](https://arxiv.org/html/2608.14096#S2.E2)\) in every period\. For anyπ∈Πad\\pi\\in\\Pi\_\{\\rm ad\}, define its expected cost by
JTπ\\displaystyle J\_\{T\}^\{\\pi\}:=𝔼π\[∑t=1T∑i=1N\(ki\(Yi,t−Ii,t\)⏟shipment cost\+hiIi,t\+1⏟store holding cost\+bi\(Di,t−Yi,t\)\+⏟lost\-sales penalty\)\+wBT\+1⏟terminal warehouse cost−∑i=1NciIi,T\+1⏟sell\-back credit\]\.\\displaystyle:=\\mathbb\{E\}^\{\\pi\}\\\!\\Bigg\[\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\\Bigl\(\\underbrace\{k\_\{i\}\(Y\_\{i,t\}\-I\_\{i,t\}\)\}\_\{\\text\{shipment cost\}\}\+\\hskip\-5\.69054pt\\underbrace\{h\_\{i\}I\_\{i,t\+1\}\}\_\{\\text\{store holding cost\}\}\\hskip\-5\.69054pt\+\\underbrace\{b\_\{i\}\(D\_\{i,t\}\-Y\_\{i,t\}\)^\{\+\}\}\_\{\\text\{lost\-sales penalty\}\}\\Bigr\)\+\\hskip\-14\.22636pt\\underbrace\{wB\_\{T\+1\}\}\_\{\\text\{terminal warehouse cost\}\}\-\\underbrace\{\\sum\_\{i=1\}^\{N\}c\_\{i\}I\_\{i,T\+1\}\}\_\{\\text\{sell\-back credit\}\}\\Bigg\]\.Hereki,hi,bik\_\{i\},h\_\{i\},b\_\{i\}are per unit shipment, store\-holding, and lost\-sales costs, respectively, andwwis terminal warehouse holding cost\. Terminal store inventory is credited atci=ki−wc\_\{i\}=k\_\{i\}\-was convention\. Throughout,ci≥0c\_\{i\}\\geq 0,hi\>0h\_\{i\}\>0, andbi−ci\>0b\_\{i\}\-c\_\{i\}\>0\.
For the fluid comparison, letμi:=𝔼Di\\mu\_\{i\}:=\\mathbb\{E\}D\_\{i\}, and, for a stock thresholdy≥0y\\geq 0, define the expected sales
mi\(y\):=𝔼\[Di∧y\]=∫0y\(1−Fi\(u\)\)𝑑um\_\{i\}\(y\):=\\mathbb\{E\}\[D\_\{i\}\\wedge y\]=\\int\_\{0\}^\{y\}\(1\-F\_\{i\}\(u\)\)\\,\\mathrm\{d\}u\(2\.4\)and the expected adjusted operating loss
ℓi\(y\)\\displaystyle\\ell\_\{i\}\(y\):=𝔼\[ci\(Di∧y\)\+hi\(y−Di\)\+\+bi\(Di−y\)\+\]\\displaystyle:=\\mathbb\{E\}\\\!\\left\[c\_\{i\}\(D\_\{i\}\\wedge y\)\+h\_\{i\}\(y\-D\_\{i\}\)^\{\+\}\+b\_\{i\}\(D\_\{i\}\-y\)^\{\+\}\\right\]=hiy\+biμi−\(hi\+bi−ci\)mi\(y\)\.\\displaystyle=h\_\{i\}y\+b\_\{i\}\\mu\_\{i\}\-\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)m\_\{i\}\(y\)\.\(2\.5\)Heremi\(y\)m\_\{i\}\(y\)is the expected depletion of the shared system inventory, whereasℓi\(y\)\\ell\_\{i\}\(y\)is the corresponding expected decision\-dependent cost after the terminal warehouse charge is accounted for\. Through direct inventory\-accounting calculations, for every admissible policyπ\\pi,
JTπ\\displaystyle J\_\{T\}^\{\\pi\}=wW\+∑t=1T∑i=1N𝔼πℓi\(Yi,t\),\\displaystyle=wW\+\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\\mathbb\{E\}^\{\\pi\}\\ell\_\{i\}\(Y\_\{i,t\}\),\(2\.6\)∑t=1T∑i=1N𝔼πmi\(Yi,t\)\\displaystyle\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\\mathbb\{E\}^\{\\pi\}m\_\{i\}\(Y\_\{i,t\}\)=W−𝔼π\[BT\+1\+∑i=1NIi,T\+1\]≤W=γT\.\\displaystyle=W\-\\mathbb\{E\}^\{\\pi\}\\\!\\left\[B\_\{T\+1\}\+\\sum\_\{i=1\}^\{N\}I\_\{i,T\+1\}\\right\]\\leq W=\\gamma T\.\(2\.7\)
The exact known\-distribution dynamic benchmark is the optimal value over the same admissible policy class:
JT∗\\displaystyle J\_\{T\}^\{\*\}:=infπ∈ΠadJTπ\.\\displaystyle:=\\inf\_\{\\pi\\in\\Pi\_\{\\rm ad\}\}J\_\{T\}^\{\\pi\}\.The identities \([2\.6](https://arxiv.org/html/2608.14096#S2.E6)\) and \([2\.7](https://arxiv.org/html/2608.14096#S2.E7)\) motivate a one\-period fluid relaxation that retains only the expected\-sales budget\. The following proposition defines its value and shows that it lower\-bounds the exact dynamic benchmark\.
###### Proposition 2\.3\.
Forr≥0r\\geq 0, define
v\(r\):=min0≤yi≤d¯i\{∑i=1Nℓi\(yi\):∑i=1Nmi\(yi\)≤r\}\.v\(r\):=\\min\_\{0\\leq y\_\{i\}\\leq\\bar\{d\}\_\{i\}\}\\left\\\{\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(y\_\{i\}\):\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\\leq r\\right\\\}\.\(2\.8\)Then we haveJT∗≥wW\+Tv\(γ\)J\_\{T\}^\{\*\}\\geq wW\+Tv\(\\gamma\)\.
By[Section2](https://arxiv.org/html/2608.14096#S2), the fluid\-benchmark gap is nonnegative and upper\-bounds the performance gap relative to the exact dynamic oracle:
0≤JTπ−JT∗≤JTπ−\(wW\+Tv\(γ\)\)\.0\\leq J\_\{T\}^\{\\pi\}\-J\_\{T\}^\{\*\}\\leq J\_\{T\}^\{\\pi\}\-\\bigl\(wW\+Tv\(\\gamma\)\\bigr\)\.We therefore define the paper’s regret directly relative to the constrained fluid benchmark:
RT\(π\):=JTπ−\(wW\+Tv\(γ\)\)\.R\_\{T\}\(\\pi\):=J\_\{T\}^\{\\pi\}\-\\bigl\(wW\+Tv\(\\gamma\)\\bigr\)\.\(2\.9\)
For each resource raterr, let𝒚∗\(r\)\\bm\{y\}^\{\*\}\(r\)denote the unique optimizer of the one\-period fluid program \([2\.8](https://arxiv.org/html/2608.14096#S2.E8)\), and letλ∗\(r\)\\lambda^\{\*\}\(r\)be its selected smallest resource multiplier\. Their existence and regularity are established in[Section4\.2](https://arxiv.org/html/2608.14096#S4.SS2)\.
## 3Resource\-Adaptive Primal\-Dual Learning
The state driving our resource\-adaptive framework is the remaining\-resource rate\. At periodtt, define
Ct:=Bt\+∑i=1NIi,t,nt:=T−t\+1,rt:=Ctnt\.C\_\{t\}:=B\_\{t\}\+\\sum\_\{i=1\}^\{N\}I\_\{i,t\},\\qquad n\_\{t\}:=T\-t\+1,\\qquad r\_\{t\}:=\\frac\{C\_\{t\}\}\{n\_\{t\}\}\.HereCtC\_\{t\}includes inventory at both the warehouse and the stores, andrtr\_\{t\}is the inventory available per remaining period\. In particular,r1=W/T=γr\_\{1\}=W/T=\\gamma\. Because shipments only relocate inventory while sales deplete it,
Ct\+1=Ct−∑i=1NSi,t,rt\+1=rt\+rt−∑i=1NSi,tnt−1\(nt≥2\)\.C\_\{t\+1\}=C\_\{t\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\},\\qquad r\_\{t\+1\}=r\_\{t\}\+\\frac\{r\_\{t\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\}\{n\_\{t\}\-1\}\\qquad\(n\_\{t\}\\geq 2\)\.Thus each sales observation both provides censored feedback and changes the next resource index\.𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}maintains preferred thresholds𝑿t\\bm\{X\}\_\{t\}and a common shadow priceλt\\lambda\_\{t\}to track the moving target\(𝒚∗\(rt\),λ∗\(rt\)\)\(\\bm\{y\}^\{\*\}\(r\_\{t\}\),\\lambda^\{\*\}\(r\_\{t\}\)\), rather than the fixed target indexed byγ\\gamma\.
At a high level, each period has four operations: project the preferred thresholds onto the physically feasible inventory set, observe censored sales and form a Primal\-Dual sample, update the remaining\-resource rate, and take a projected stochastic step\. The ratertr\_\{t\}moves the fluid target after every sales realization\. Section 3\.1 specifies the complete executable recursion, while Sections 3\.2 and 3\.3 explain its gradient construction, safeguards, and step\-size design\.
### 3\.1The Executable RAPDL Recursion
[Algorithm1](https://arxiv.org/html/2608.14096#alg1)gives the complete executable RAPDL pseudocode\. We first define the action projection and the quantities appearing in its updates\.
The preferred threshold𝑿t\\bm\{X\}\_\{t\}is a learning state and need not be a feasible post\-shipment inventory\. Given the current physical state, the feasible action set is
ℱt:=\{𝒚:yi≥Ii,t,∑i=1N\(yi−Ii,t\)≤Bt\}=\{𝒚:yi≥Ii,t,∑i=1Nyi≤Ct\}\.\\mathcal\{F\}\_\{t\}:=\\left\\\{\\bm\{y\}:y\_\{i\}\\geq I\_\{i,t\},\\ \\sum\_\{i=1\}^\{N\}\(y\_\{i\}\-I\_\{i,t\}\)\\leq B\_\{t\}\\right\\\}=\\left\\\{\\bm\{y\}:y\_\{i\}\\geq I\_\{i,t\},\\ \\sum\_\{i=1\}^\{N\}y\_\{i\}\\leq C\_\{t\}\\right\\\}\.The right\-hand side usesCtC\_\{t\}, rather thanBtB\_\{t\}, because store carryover is already part of the available system inventory\. We implement the Euclidean projection of𝑿t\\bm\{X\}\_\{t\}ontoℱt\\mathcal\{F\}\_\{t\},
𝒀t=\\argmin𝒚∈ℱt12‖𝒚−𝑿t‖2,\\bm\{Y\}\_\{t\}=\\argmin\_\{\\bm\{y\}\\in\\mathcal\{F\}\_\{t\}\}\\frac\{1\}\{2\}\\\|\\bm\{y\}\-\\bm\{X\}\_\{t\}\\\|^\{2\},which has the water\-filling representation\([10](https://arxiv.org/html/2608.14096#bib.bib31), Lemma 2\)
νt:=inf\{ν≥0:∑i=1Nmax\{Ii,t,Xi,t−ν\}≤Ct\},Yi,t=max\{Ii,t,Xi,t−νt\},i∈\[N\]\.\\displaystyle\\nu\_\{t\}:=\\inf\\left\\\{\\nu\\geq 0:\\sum\_\{i=1\}^\{N\}\\max\\\{I\_\{i,t\},X\_\{i,t\}\-\\nu\\\}\\leq C\_\{t\}\\right\\\},~~~Y\_\{i,t\}=\\max\\\{I\_\{i,t\},X\_\{i,t\}\-\\nu\_\{t\}\\\},\\qquad i\\in\[N\]\.
To specify the projected update, define the cost\-based survival ratio
pi:=hihi\+bi−ci,i∈\[N\]\.p\_\{i\}:=\\frac\{h\_\{i\}\}\{h\_\{i\}\+b\_\{i\}\-c\_\{i\}\},\\qquad i\\in\[N\]\.\(3\.1\)Define the capped aggregate resource rate byrc:=min\{r,D¯\}r^\{\\mathrm\{c\}\}:=\\min\\\{r,\\overline\{D\}\\\}\. Independently ofrr, for each store define the fixed safe threshold cap
X¯i:=d¯i−pi2Ki,i∈\[N\]\.\\bar\{X\}\_\{i\}:=\\bar\{d\}\_\{i\}\-\\frac\{p\_\{i\}\}\{2K\_\{i\}\},\\qquad i\\in\[N\]\.\(3\.2\)The algorithm combinesX¯i\\bar\{X\}\_\{i\}with the resource\-dependent bound throughUi\(r\)U\_\{i\}\(r\)\. With the dual capλmax:=maxi∈\[N\]\(bi−ci\)\\lambda\_\{\\max\}:=\\max\_\{i\\in\[N\]\}\(b\_\{i\}\-c\_\{i\}\), the moving Primal\-Dual projection set is
Ui\(r\):=min\{X¯i,rcpi\},i∈\[N\],𝒦\(r\):=∏i=1N\[0,Ui\(r\)\]×\[0,λmax\]\.U\_\{i\}\(r\):=\\min\\left\\\{\\bar\{X\}\_\{i\},\\frac\{r^\{\\mathrm\{c\}\}\}\{p\_\{i\}\}\\right\\\},\\quad i\\in\[N\],\\qquad\\mathcal\{K\}\(r\):=\\prod\_\{i=1\}^\{N\}\[0,U\_\{i\}\(r\)\]\\times\[0,\\lambda\_\{\\max\}\]\.\(3\.3\)Letj:=minargmaxi∈\[N\]\(bi−ci\)j:=\\min\\operatorname\*\{arg\\,max\}\_\{i\\in\[N\]\}\(b\_\{i\}\-c\_\{i\}\)be the maximum\-margin anchor, with ties broken by the smallest index\. After implementing𝒀t\\bm\{Y\}\_\{t\}and observingSi,t:=Di,t∧Yi,tS\_\{i,t\}:=D\_\{i,t\}\\wedge Y\_\{i,t\}, form
G^i,t:=\(hi\+bi−ci\)𝟏\{Si,t<Yi,t\}−\(bi−ci\)\+λt𝟏\{Si,t=Yi,t\},i∈\[N\],\\widehat\{G\}\_\{i,t\}:=\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)\\mathbf\{1\}\\\{S\_\{i,t\}<Y\_\{i,t\}\\\}\-\(b\_\{i\}\-c\_\{i\}\)\+\\lambda\_\{t\}\\mathbf\{1\}\\\{S\_\{i,t\}=Y\_\{i,t\}\\\},\\qquad i\\in\[N\],\(3\.4\)and
G^0,t:=rtc−∑i=1NSi,t\+θG^j,t,G^t:=\(G^1,t,…,G^N,t,G^0,t\)\.\\widehat\{G\}\_\{0,t\}:=r\_\{t\}^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\+\\theta\\widehat\{G\}\_\{j,t\},\\qquad\\widehat\{G\}\_\{t\}:=\(\\widehat\{G\}\_\{1,t\},\\ldots,\\widehat\{G\}\_\{N,t\},\\widehat\{G\}\_\{0,t\}\)\.\(3\.5\)Hereθ\>0\\theta\>0is the weight assigned to the anchor correction in the dual update\. Its stabilizing role is discussed in Section[3\.2](https://arxiv.org/html/2608.14096#S3.SS2)\. Finally, use the two\-sided harmonic step
αt:=χτ0\+min\{t,nt\}=χτ0\+min\{t,T−t\+1\}\.\\alpha\_\{t\}:=\\frac\{\\chi\}\{\\tau\_\{0\}\+\\min\\\{t,n\_\{t\}\\\}\}=\\frac\{\\chi\}\{\\tau\_\{0\}\+\\min\\\{t,T\-t\+1\\\}\}\.\(3\.6\)
Algorithm 1Resource\-Adaptive Primal\-Dual Learning \(RAPDL\) for OWMS1:Initialize
B1=C1:=γTB\_\{1\}=C\_\{1\}:=\\gamma T,
n1:=Tn\_\{1\}:=T,
r1:=γr\_\{1\}:=\\gamma,
r1c:=min\{γ,D¯\}r\_\{1\}^\{\\mathrm\{c\}\}:=\\min\\\{\\gamma,\\overline\{D\}\\\}and
Ii,1=Xi,1:=0I\_\{i,1\}=X\_\{i,1\}:=0for all
ii, and
λ1:=0\\lambda\_\{1\}:=0\.
2:for
t=1,…,Tt=1,\\ldots,Tdo
3:Inventory decision:compute
νt:=inf\{ν≥0:∑i=1Nmax\{Ii,t,Xi,t−ν\}≤Ct\}\\nu\_\{t\}:=\\inf\\\{\\nu\\geq 0:\\sum\_\{i=1\}^\{N\}\\max\\\{I\_\{i,t\},X\_\{i,t\}\-\\nu\\\}\\leq C\_\{t\}\\\}and set
Yi,t:=max\{Ii,t,Xi,t−νt\}Y\_\{i,t\}:=\\max\\\{I\_\{i,t\},X\_\{i,t\}\-\\nu\_\{t\}\\\}for all
ii\.
4:Gradient estimation:observe
Si,t=Di,t∧Yi,tS\_\{i,t\}=D\_\{i,t\}\\wedge Y\_\{i,t\}and compute, for all
i∈\[N\]i\\in\[N\],
G^i,t:=\(hi\+bi−ci\)𝟏\{Si,t<Yi,t\}−\(bi−ci\)\+λt𝟏\{Si,t=Yi,t\},\\widehat\{G\}\_\{i,t\}:=\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)\\mathbf\{1\}\\\{S\_\{i,t\}<Y\_\{i,t\}\\\}\-\(b\_\{i\}\-c\_\{i\}\)\+\\lambda\_\{t\}\\mathbf\{1\}\\\{S\_\{i,t\}=Y\_\{i,t\}\\\},G^0,t:=rtc−∑i=1NSi,t\+θG^j,t\.\\widehat\{G\}\_\{0,t\}:=r\_\{t\}^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\+\\theta\\widehat\{G\}\_\{j,t\}\.
5:Physical\-state update:
Ii,t\+1:=Yi,t−Si,tI\_\{i,t\+1\}:=Y\_\{i,t\}\-S\_\{i,t\},
Ct\+1:=Ct−∑i=1NSi,tC\_\{t\+1\}:=C\_\{t\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}, and
Bt\+1:=Ct\+1−∑i=1NIi,t\+1B\_\{t\+1\}:=C\_\{t\+1\}\-\\sum\_\{i=1\}^\{N\}I\_\{i,t\+1\}\.
6:if
t<Tt<Tthen
7:Resource update:
nt\+1:=nt−1n\_\{t\+1\}:=n\_\{t\}\-1,
rt\+1:=Ct\+1/nt\+1r\_\{t\+1\}:=C\_\{t\+1\}/n\_\{t\+1\}, and
rt\+1c:=min\{rt\+1,D¯\}r\_\{t\+1\}^\{\\mathrm\{c\}\}:=\\min\\\{r\_\{t\+1\},\\overline\{D\}\\\}\.
8:Primal\-Dual update:
Xi,t\+1\\displaystyle X\_\{i,t\+1\}:=Proj\[0,Ui\(rt\+1\)\]\(Xi,t−αtG^i,t\),\\displaystyle:=\\operatorname\{Proj\}\_\{\[0,U\_\{i\}\(r\_\{t\+1\}\)\]\}\(X\_\{i,t\}\-\\alpha\_\{t\}\\widehat\{G\}\_\{i,t\}\),i∈\[N\],\\displaystyle i\\in\[N\],λt\+1\\displaystyle\\lambda\_\{t\+1\}:=Proj\[0,λmax\]\(λt−αtG^0,t\)\.\\displaystyle:=\\operatorname\{Proj\}\_\{\[0,\\lambda\_\{\\max\}\]\}\(\\lambda\_\{t\}\-\\alpha\_\{t\}\\widehat\{G\}\_\{0,t\}\)\.
9:The projection endpoints are given in \([3\.3](https://arxiv.org/html/2608.14096#S3.E3)\), andαt\\alpha\_\{t\}is given in \([3\.6](https://arxiv.org/html/2608.14096#S3.E6)\)\.
### 3\.2Gradient Estimation and Dual Stabilization
The recursion above uses censored sales to constructG^t\\widehat\{G\}\_\{t\}\. We first identify the fluid direction estimated at the implemented action𝒀t\\bm\{Y\}\_\{t\}, and then explain the anchor correction that stabilizes its dual coordinate\.
Write𝒎\(𝒚\):=\(mi\(yi\)\)i=1N\\bm\{m\}\(\\bm\{y\}\):=\(m\_\{i\}\(y\_\{i\}\)\)\_\{i=1\}^\{N\}\. At resource ratertr\_\{t\}, define the threshold\-space fluid Lagrangian and its derivatives by
ℒ~rt\(𝒚,λ\)\\displaystyle\\widetilde\{\\mathcal\{L\}\}\_\{r\_\{t\}\}\(\\bm\{y\},\\lambda\):=∑i=1Nℓi\(yi\)\+λ\(∑i=1Nmi\(yi\)−rt\),\\displaystyle:=\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(y\_\{i\}\)\+\\lambda\\left\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\-r\_\{t\}\\right\),\(3\.7\)∂yiℒ~rt\(𝒚,λ\)\\displaystyle\\partial\_\{y\_\{i\}\}\\widetilde\{\\mathcal\{L\}\}\_\{r\_\{t\}\}\(\\bm\{y\},\\lambda\)=\(hi\+bi−ci\)Fi\(yi\)−\(bi−ci\)\+λ\(1−Fi\(yi\)\),\\displaystyle=\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)F\_\{i\}\(y\_\{i\}\)\-\(b\_\{i\}\-c\_\{i\}\)\+\\lambda\\bigl\(1\-F\_\{i\}\(y\_\{i\}\)\\bigr\),i∈\[N\],\\displaystyle i\\in\[N\],∂λℒ~rt\(𝒚,λ\)\\displaystyle\\partial\_\{\\lambda\}\\widetilde\{\\mathcal\{L\}\}\_\{r\_\{t\}\}\(\\bm\{y\},\\lambda\)=∑i=1Nmi\(yi\)−rt\.\\displaystyle=\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\-r\_\{t\}\.
Although this field is not computable becauseFiF\_\{i\}andmim\_\{i\}are unknown, the algorithm samples it at the implemented action as follows\.
BecauseDi,t<Yi,tD\_\{i,t\}<Y\_\{i,t\}iffSi,t<Yi,tS\_\{i,t\}<Y\_\{i,t\}, censored sales make \([3\.4](https://arxiv.org/html/2608.14096#S3.E4)\) observable and satisfy222Throughout this section and the regret analysis, writeℋt:=ℋt𝖱𝖠𝖯𝖣𝖫\\mathcal\{H\}\_\{t\}:=\\mathcal\{H\}\_\{t\}^\{\\mathsf\{RAPDL\}\}\.
𝔼\[G^i,t∣ℋt\]=∂yiℒ~rt\(𝒀t,λt\),𝔼\[rt−∑iSi,t∣ℋt\]=−∂λℒ~rt\(𝒀t,λt\)\.\\mathbb\{E\}\[\\widehat\{G\}\_\{i,t\}\\mid\\mathcal\{H\}\_\{t\}\]=\\partial\_\{y\_\{i\}\}\\widetilde\{\\mathcal\{L\}\}\_\{r\_\{t\}\}\(\\bm\{Y\}\_\{t\},\\lambda\_\{t\}\),\\qquad\\mathbb\{E\}\\\!\\left\[r\_\{t\}\-\\sum\_\{i\}S\_\{i,t\}\\mid\\mathcal\{H\}\_\{t\}\\right\]=\-\\partial\_\{\\lambda\}\\widetilde\{\\mathcal\{L\}\}\_\{r\_\{t\}\}\(\\bm\{Y\}\_\{t\},\\lambda\_\{t\}\)\.
These identities establish the gradient\-estimation component, but the uncorrected joint direction still lacks a restoring force in the dual coordinate\. This motivates the following stabilization\.
The anchor correction\.The sales\-space primal Hessian is uniformly positive on the safe region, but the joint Lagrangian has zero curvature in the dual coordinate\. Without an additional restoring term, the usualO\(1/t\)O\(1/t\)mean\-square contraction is unavailable\. The genericO\(t−1/2\)O\(t^\{\-1/2\}\)tracking scale would accumulate toO\(T\)O\(\\sqrt\{T\}\), rather thanO\(logT\)O\(\\log T\)\. The anchor borrows curvature from one primal coordinate\([3](https://arxiv.org/html/2608.14096#bib.bib3),[20](https://arxiv.org/html/2608.14096#bib.bib4)\)\. For the maximum\-margin anchorjj, its KKT residual vanishes along the selected target path, including at zero resource, so the correction stabilizes the dual direction without moving the target\. Specifically, use
Gθ\(𝒚,λ,r\)\\displaystyle G^\{\\theta\}\(\\bm\{y\},\\lambda;r\):=\(∇𝒚ℒ~r\(𝒚,λ\)rc−∑i=1Nmi\(yi\)\+θ∂yjℒ~r\(𝒚,λ\)\)\.\\displaystyle:=\\begin\{pmatrix\}\\nabla\_\{\\bm\{y\}\}\\widetilde\{\\mathcal\{L\}\}\_\{r\}\(\\bm\{y\},\\lambda\)\\\\\[2\.0pt\] r^\{\\mathrm\{c\}\}\-\\displaystyle\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\+\\theta\\,\\partial\_\{y\_\{j\}\}\\widetilde\{\\mathcal\{L\}\}\_\{r\}\(\\bm\{y\},\\lambda\)\\end\{pmatrix\}\.Its observable version is exactly \([3\.5](https://arxiv.org/html/2608.14096#S3.E5)\), so the anchor requires no additional sample\. In particular,
𝔼\[G^t∣ℋt\]=Gθ\(𝒀t,λt,rt\)\.\\mathbb\{E\}\[\\widehat\{G\}\_\{t\}\\mid\\mathcal\{H\}\_\{t\}\]=G^\{\\theta\}\(\\bm\{Y\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\.Thus feedback estimates the stabilized field at the implemented action𝒀t\\bm\{Y\}\_\{t\}, while the preferred\-state direction isGθ\(𝑿t,λt,rt\)G^\{\\theta\}\(\\bm\{X\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\. Their difference is the implementation bias bounded in \([A\.20](https://arxiv.org/html/2608.14096#A1.E20)\)\.
### 3\.3Safeguards and Step\-Size Design
This subsection turns to the remaining safeguards and step\-size design\. Resource capping controls sample magnitude, safe state clipping preserves curvature, and the two\-sided step\-size schedule balances learning across the horizon\.
Resource capping and a safe clipping box\.Near the horizon,rt=Ct/ntr\_\{t\}=C\_\{t\}/n\_\{t\}can be of orderTT, making both the uncapped dual samplert−∑i=1NSi,tr\_\{t\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}and the resulting quadratic term in the one\-step analysis unbounded inTT\.
The caprc=min\{r,D¯\}r^\{\\mathrm\{c\}\}=\\min\\\{r,\\overline\{D\}\\\}removes this endpoint explosion\. By construction,0≤rc≤D¯0\\leq r^\{\\mathrm\{c\}\}\\leq\\overline\{D\}, andr↦rcr\\mapsto r^\{\\mathrm\{c\}\}is one\-Lipschitz\. Since total one\-period sales also lies in\[0,D¯\]\[0,\\overline\{D\}\], the capped dual sample is bounded in absolute value byD¯\\overline\{D\}\. It leaves the fluid target unchanged because it truncates only the nonbinding region\. See[SectionA\.1](https://arxiv.org/html/2608.14096#A1.SS1)\.
The fixed caps in \([3\.2](https://arxiv.org/html/2608.14096#S3.E2)\) keep every preferred threshold in a region of uniformly positive survival probability without excluding the fluid target\. Indeed, becauseλ∗\(r\)≥0\\lambda^\{\*\}\(r\)\\geq 0,
yi∗\(r\)≤Fi−1\(1−pi\)≤d¯i−piKi<X¯i\.y\_\{i\}^\{\*\}\(r\)\\leq F\_\{i\}^\{\-1\}\(1\-p\_\{i\}\)\\leq\\bar\{d\}\_\{i\}\-\\frac\{p\_\{i\}\}\{K\_\{i\}\}<\\bar\{X\}\_\{i\}\.For later use, define the associated survival\-probability floorβi:=κipi/\(2Ki\)\\beta\_\{i\}:=\\kappa\_\{i\}p\_\{i\}/\(2K\_\{i\}\),i∈\[N\]i\\in\[N\]\. Then, for every0≤x≤X¯i0\\leq x\\leq\\bar\{X\}\_\{i\},
1−Fi\(x\)≥κi\(d¯i−x\)≥βi\.1\-F\_\{i\}\(x\)\\geq\\kappa\_\{i\}\(\\bar\{d\}\_\{i\}\-x\)\\geq\\beta\_\{i\}\.
The second clipping bound responds to the remaining resource\. Recall from \([3\.1](https://arxiv.org/html/2608.14096#S3.E1)\) thatpi=hi/\(hi\+bi−ci\)p\_\{i\}=h\_\{i\}/\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)is the survival probability at the zero\-price newsvendor threshold\. At an active fluid target, survival is therefore at leastpip\_\{i\}, so
piyi∗\(r\)≤mi\(yi∗\(r\)\)≤rc,and thereforeyi∗\(r\)≤rcpi\.p\_\{i\}y\_\{i\}^\{\*\}\(r\)\\leq m\_\{i\}\(y\_\{i\}^\{\*\}\(r\)\)\\leq r^\{\\mathrm\{c\}\},\\qquad\\text\{and therefore\}\\qquad y\_\{i\}^\{\*\}\(r\)\\leq\\frac\{r^\{\\mathrm\{c\}\}\}\{p\_\{i\}\}\.If the store is inactive,yi∗\(r\)=0y\_\{i\}^\{\*\}\(r\)=0, so the same bound holds\. This gives precisely the capUi\(r\)U\_\{i\}\(r\)and projection box𝒦\(r\)\\mathcal\{K\}\(r\)defined in \([3\.3](https://arxiv.org/html/2608.14096#S3.E3)\)\. The endpointX¯i\\bar\{X\}\_\{i\}protects the survival and curvature geometry, whereasrc/pir^\{\\mathrm\{c\}\}/p\_\{i\}makes the box contract with the remaining resource\. Taking their minimum enforces both requirements\. Finally,λmax\\lambda\_\{\\max\}is sufficient because every store’s fluid best response is zero at any resource price at leastλmax\\lambda\_\{\\max\}\.
Step\-size choice\.The scaleχ\\chiin \([3\.6](https://arxiv.org/html/2608.14096#S3.E6)\) controls responsiveness, whileτ0\\tau\_\{0\}prevents the larger endpoint updates from being too aggressive\. Thus the two\-sided schedule permits faster learning initially and faster resource adjustment near the end, with smaller updates in the middle\. The constructive choice in[Section4](https://arxiv.org/html/2608.14096#S4)is horizon independent and certifies the regret guarantee\. In practice,\(θ,χ,τ0\)\(\\theta,\\chi,\\tau\_\{0\}\)can instead be selected offline by a coarse grid search or rolling\-origin validation for the intended operating instance\. Appendix[B\.4](https://arxiv.org/html/2608.14096#A2.SS4)describes how fixed support envelopes can be used in the caps and outlines the associated localization argument\.
## 4Main Result and Regret Analysis
The following theorem quantifies the performance of the resource\-adaptive learning framework\.
###### Theorem 4\.1\.
Under[Section2](https://arxiv.org/html/2608.14096#S2),𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}with the initialization in[Algorithm1](https://arxiv.org/html/2608.14096#alg1)admits horizon\-independent tuning based only on the known primitives such that, for everyT≥2T\\geq 2, every0≤γ≤D¯0\\leq\\gamma\\leq\\overline\{D\}, and every admissible within\-period joint demand law,
0≤JT𝖱𝖠𝖯𝖣𝖫−JT∗≤RT\(𝖱𝖠𝖯𝖣𝖫\)≤ClogT,\\displaystyle 0\\leq J\_\{T\}^\{\\mathsf\{RAPDL\}\}\-J\_\{T\}^\{\*\}\\leq R\_\{T\}\(\\mathsf\{RAPDL\}\)\\leq C\\log T,whereC<∞C<\\inftydepends only on the fixed model primitives and the selected horizon\-independent tuning, and not onTT,γ\\gamma\.
[Theorem4\.1](https://arxiv.org/html/2608.14096#S4.Thmtheorem1)concerns the direct safe\-box formulation\. As a separate extension, Appendix[B\.4](https://arxiv.org/html/2608.14096#A2.SS4)sketches how the analysis can be adapted when exact endpoints are replaced by fixed deterministic support envelopes\.
### 4\.1High\-Level Idea of Regret Analysis
At a high level, regret consists of three distinct logarithmic components\.
*First, the re\-solving gap\.*This is the intrinsic gap between the resource\-indexed fluid re\-solving path and the fluid benchmark: even if the policy could follow this path exactly, movements in the remaining\-resource rate would generate anO\(logT\)O\(\\log T\)cumulative gap\.
*Second, the learning gap\.*Because the demand laws are unknown, the Primal\-Dual re\-solving path cannot be computed directly\.𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}instead learns the path by tracking it from censored sales, and the learning analysis shows that the resulting cumulative tracking gap isO\(logT\)O\(\\log T\)\.
*Third, the implementation gap\.*The preferred actions produced by the learning recursion may differ from the actions that can be implemented under the current warehouse and store inventories\. The physical implementation analysis shows that this cumulative discrepancy is alsoO\(logT\)O\(\\log T\)\. Adding the re\-solving, learning, and implementation gaps yields the overallO\(logT\)O\(\\log T\)regret bound\.
### 4\.2Analytical Setup in Expected\-Sales Space
Direct analysis in threshold space does not provide the uniform primal curvature required by the tracking argument\. Althoughℓi\\ell\_\{i\}is strongly convex, the nonlinear resource term changes the curvature of the threshold\-space Lagrangian to
∂2∂yi2\(ℓi\(yi\)\+λmi\(yi\)\)=\(hi\+bi−ci−λ\)fi\(yi\),\\frac\{\\partial^\{2\}\}\{\\partial y\_\{i\}^\{2\}\}\\bigl\(\\ell\_\{i\}\(y\_\{i\}\)\+\\lambda m\_\{i\}\(y\_\{i\}\)\\bigr\)=\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\-\\lambda\)f\_\{i\}\(y\_\{i\}\),which need not be uniformly positive\. Settingsi=mi\(yi\)s\_\{i\}=m\_\{i\}\(y\_\{i\}\)makes the resource term affine, so the shadow price no longer erodes the primal curvature\. This change of variables is purely analytical\. The executable policy remains in threshold space\. Sincemim\_\{i\}is strictly increasing on\[0,d¯i\)\[0,\\bar\{d\}\_\{i\}\), define
si:=mi\(yi\),gi\(s\):=ℓi\(mi−1\(s\)\),0≤s<μi,s\_\{i\}:=m\_\{i\}\(y\_\{i\}\),\\qquad g\_\{i\}\(s\):=\\ell\_\{i\}\(m\_\{i\}^\{\-1\}\(s\)\),\\qquad 0\\leq s<\\mu\_\{i\},and setgi\(μi\):=ℓi\(d¯i\)g\_\{i\}\(\\mu\_\{i\}\):=\\ell\_\{i\}\(\\bar\{d\}\_\{i\}\)by continuity\. Direct differentiation givesgi′\(0\)=−\(bi−ci\)g\_\{i\}^\{\\prime\}\(0\)=\-\(b\_\{i\}\-c\_\{i\}\)and, whenever1−Fi\(mi−1\(s\)\)≥β\>01\-F\_\{i\}\(m\_\{i\}^\{\-1\}\(s\)\)\\geq\\beta\>0,
hiκi≤gi′′\(s\)≤hiKi/β3\.h\_\{i\}\\kappa\_\{i\}\\leq g\_\{i\}^\{\\prime\\prime\}\(s\)\\leq h\_\{i\}K\_\{i\}/\\beta^\{3\}\.\(4\.1\)Thusgig\_\{i\}is strongly convex, and the fluid program becomes
v\(r\)=min0≤si≤μi\{∑i=1Ngi\(si\):∑i=1Nsi≤r\}\.v\(r\)=\\min\_\{0\\leq s\_\{i\}\\leq\\mu\_\{i\}\}\\left\\\{\\sum\_\{i=1\}^\{N\}g\_\{i\}\(s\_\{i\}\):\\sum\_\{i=1\}^\{N\}s\_\{i\}\\leq r\\right\\\}\.Its resource\-indexed Lagrangian is
ℒr\(𝒔,λ\):=∑i=1Ngi\(si\)\+λ\(∑i=1Nsi−r\)\.\\mathcal\{L\}\_\{r\}\(\\bm\{s\},\\lambda\):=\\sum\_\{i=1\}^\{N\}g\_\{i\}\(s\_\{i\}\)\+\\lambda\\left\(\\sum\_\{i=1\}^\{N\}s\_\{i\}\-r\\right\)\.
Let
r0:=∑i=1Nmi\(Fi−1\(1−pi\)\)r\_\{0\}:=\\sum\_\{i=1\}^\{N\}m\_\{i\}\\\!\\left\(F\_\{i\}^\{\-1\}\(1\-p\_\{i\}\)\\right\)be the expected sales of the unconstrained zero\-price solution\. The next lemma collects exactly the path properties used in the regret analysis\.
###### Lemma 4\.3\.
For everyr≥0r\\geq 0, the transformed fluid program has a unique primal optimizer𝐬∗\(r\)\\bm\{s\}^\{\*\}\(r\), withsi∗\(r\)<μis\_\{i\}^\{\*\}\(r\)<\\mu\_\{i\}, and a selected smallest optimal multiplierλ∗\(r\)\\lambda^\{\*\}\(r\)\. The following properties hold, including when several stores have identical margins\.
1. \(i\)The KKT residuals are gi′\(si∗\(r\)\)\+λ∗\(r\)\\displaystyle g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)≥0,\\displaystyle\\geq 0,si∗\(r\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\\displaystyle s\_\{i\}^\{\*\}\(r\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)=0,\\displaystyle=0,i∈\[N\],\\displaystyle i\\in\[N\],\(4\.2\)r−∑i=1Nsi∗\(r\)\\displaystyle r\-\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\)≥0,\\displaystyle\\geq 0,λ∗\(r\)\(r−∑i=1Nsi∗\(r\)\)\\displaystyle\\lambda^\{\*\}\(r\)\\left\(r\-\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\)\\right\)=0\.\\displaystyle=0\.In particular, storeiiis active if and only ifλ∗\(r\)<bi−ci\\lambda^\{\*\}\(r\)<b\_\{i\}\-c\_\{i\}\. On an active coordinate,−gi′\(si∗\(r\)\)=λ∗\(r\)\-g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)=\\lambda^\{\*\}\(r\)\.
2. \(ii\)The resource\-clearing branches satisfy ∑i=1Nsi∗\(r\)=min\{r,r0\},\(𝒔∗\(0\),λ∗\(0\)\)=\(𝟎,λmax\),λ∗\(r\)=0\(r≥r0\)\.\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\)=\\min\\\{r,r\_\{0\}\\\},\\qquad\(\\bm\{s\}^\{\*\}\(0\),\\lambda^\{\*\}\(0\)\)=\(\\bm\{0\},\\lambda\_\{\\max\}\),\\qquad\\lambda^\{\*\}\(r\)=0\\quad\(r\\geq r\_\{0\}\)\.
3. \(iii\)There is a primitive\-computable constantLv<∞L\_\{v\}<\\inftysuch that \|λ∗\(r′\)−λ∗\(r\)\|≤Lv\|r′−r\|,r,r′≥0\.\|\\lambda^\{\*\}\(r^\{\\prime\}\)\-\\lambda^\{\*\}\(r\)\|\\leq L\_\{v\}\|r^\{\\prime\}\-r\|,\\qquad r,r^\{\\prime\}\\geq 0\.Moreover,q∗\(r\):=\(𝒔∗\(r\),λ∗\(r\)\)q^\{\*\}\(r\):=\(\\bm\{s\}^\{\*\}\(r\),\\lambda^\{\*\}\(r\)\)is globally Lipschitz\. Eachsi∗\(r\)s\_\{i\}^\{\*\}\(r\)is nondecreasing, whereasλ∗\(r\)\\lambda^\{\*\}\(r\)is nonincreasing\. The value function is continuously differentiable, with v′\(r\)=−λ∗\(r\),\|v′\(r′\)−v′\(r\)\|≤Lv\|r′−r\|,v^\{\\prime\}\(r\)=\-\\lambda^\{\*\}\(r\),\\qquad\|v^\{\\prime\}\(r^\{\\prime\}\)\-v^\{\\prime\}\(r\)\|\\leq L\_\{v\}\|r^\{\\prime\}\-r\|,where the derivative at zero is the right derivative\.
The threshold target introduced in Section[2](https://arxiv.org/html/2608.14096#S2)is the inverse representation
yi∗\(r\)=mi−1\(si∗\(r\)\),i∈\[N\]\.y\_\{i\}^\{\*\}\(r\)=m\_\{i\}^\{\-1\}\(s\_\{i\}^\{\*\}\(r\)\),\\qquad i\\in\[N\]\.Thus\(𝒚∗\(r\),λ∗\(r\)\)\(\\bm\{y\}^\{\*\}\(r\),\\lambda^\{\*\}\(r\)\)describes the target tracked by the policy, whereas𝒔∗\(r\)\\bm\{s\}^\{\*\}\(r\)andgig\_\{i\}are used only in the analysis\.
### 4\.3Step 1: Bounding the Re\-Solving Gap
The first component compares the resource\-indexed fluid re\-solving path with the fluid benchmark\. To make this comparison, we work in expected\-sales space, where three vectors play different roles:𝒎\(𝒀t\)\\bm\{m\}\(\\bm\{Y\}\_\{t\}\)is conditional expected sales under the implemented action,𝒎\(𝑿t\)\\bm\{m\}\(\\bm\{X\}\_\{t\}\)is the expected\-sales image of the preferred learning state, and𝒔∗\(rt\)\\bm\{s\}^\{\*\}\(r\_\{t\}\)is the unobserved fluid target\. Throughout the regret analysis, we use the pathwise safe\-box invariant0≤Xi,t,Ii,t,Yi,t≤X¯i,i∈\[N\],t∈\[T\],0\\leq X\_\{i,t\},I\_\{i,t\},Y\_\{i,t\}\\leq\\bar\{X\}\_\{i\},i\\in\[N\],\\ t\\in\[T\],established from the algorithmic recursion in \([B\.5](https://arxiv.org/html/2608.14096#A2.E5)\)\.
We collect the learning and implementation gaps, respectively, as
𝖯𝖣T\\displaystyle\\mathsf\{PD\}\_\{T\}:=∑t=1T𝔼\[‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+⟨𝒎\(𝑿t\)−𝒔∗\(rt\),∇𝒔ℒrt\(𝒔∗\(rt\),λ∗\(rt\)\)⟩\],\\displaystyle:=\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\!\\left\[\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\\left\\langle\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\),\\nabla\_\{\\bm\{s\}\}\\mathcal\{L\}\_\{r\_\{t\}\}\(\\bm\{s\}^\{\*\}\(r\_\{t\}\),\\lambda^\{\*\}\(r\_\{t\}\)\)\\right\\rangle\\right\],\(4\.3\)𝖨𝖬𝖯T\\displaystyle\\mathsf\{IMP\}\_\{T\}:=∑t=1T𝔼‖𝒀t−𝑿t‖1\+∑t=1T𝔼‖𝒀t−𝑿t‖2\.\\displaystyle:=\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\+\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(4\.4\)Here𝖯𝖣T\\mathsf\{PD\}\_\{T\}contains the squared tracking error and the first\-order KKT residual that enter the operational gap, whereas𝖨𝖬𝖯T\\mathsf\{IMP\}\_\{T\}contains the two implementation norms needed by the tracking and regret arguments\.
The following gap bound is the deterministic bridge from fluid geometry to regret\. Set
Lg:=max1≤i≤NhiKiβi3\.L\_\{g\}:=\\max\_\{1\\leq i\\leq N\}\\frac\{h\_\{i\}K\_\{i\}\}\{\\beta\_\{i\}^\{3\}\}\.
Forr≥0r\\geq 0and any𝒔\\bm\{s\}satisfying0≤mi−1\(si\)≤X¯i0\\leq m\_\{i\}^\{\-1\}\(s\_\{i\}\)\\leq\\bar\{X\}\_\{i\}, define
Δ\(𝒔,r\):=ℒr\(𝒔,λ∗\(r\)\)−ℒr\(𝒔∗\(r\),λ∗\(r\)\)=∑i=1Ngi\(si\)−v\(r\)\+λ∗\(r\)\(∑i=1Nsi−r\)\.\\displaystyle\\Delta\(\\bm\{s\};r\):=\{\\mathcal\{L\}\}\_\{r\}\(\\bm\{s\},\\lambda^\{\*\}\(r\)\)\-\{\\mathcal\{L\}\}\_\{r\}\(\\bm\{s\}^\{\*\}\(r\),\\lambda^\{\*\}\(r\)\)=\\sum\_\{i=1\}^\{N\}g\_\{i\}\(s\_\{i\}\)\-v\(r\)\+\\lambda^\{\*\}\(r\)\\left\(\\sum\_\{i=1\}^\{N\}s\_\{i\}\-r\\right\)\.\(4\.5\)Because𝒔∗\(r\)\\bm\{s\}^\{\*\}\(r\)minimizesℒr\(⋅,λ∗\(r\)\)\{\\mathcal\{L\}\}\_\{r\}\(\\cdot,\\lambda^\{\*\}\(r\)\), this gap is nonnegative\. Moreover, \([4\.1](https://arxiv.org/html/2608.14096#S4.E1)\) and the bound1−Fi\(x\)≥βi1\-F\_\{i\}\(x\)\\geq\\beta\_\{i\}on\[0,X¯i\]\[0,\\bar\{X\}\_\{i\}\]make the LagrangianLgL\_\{g\}\-smooth on the segment from𝒔∗\(r\)\\bm\{s\}^\{\*\}\(r\)to𝒔\\bm\{s\}\. Thus the descent lemma, followed by the KKT complementarity identities in \([4\.2](https://arxiv.org/html/2608.14096#S4.E2)\), gives directly
0≤Δ\(𝒔,r\)\\displaystyle 0\\leq\\Delta\(\\bm\{s\};r\)≤Lg2‖𝒔−𝒔∗\(r\)‖2\+⟨𝒔−𝒔∗\(r\),∇𝒔ℒr\(𝒔∗\(r\),λ∗\(r\)\)⟩\\displaystyle\\leq\\frac\{L\_\{g\}\}\{2\}\\\|\\bm\{s\}\-\\bm\{s\}^\{\*\}\(r\)\\\|^\{2\}\+\\left\\langle\\bm\{s\}\-\\bm\{s\}^\{\*\}\(r\),\\nabla\_\{\\bm\{s\}\}\{\\mathcal\{L\}\}\_\{r\}\(\\bm\{s\}^\{\*\}\(r\),\\lambda^\{\*\}\(r\)\)\\right\\rangle=Lg2‖𝒔−𝒔∗\(r\)‖2\+∑i=1Nsi\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\.\\displaystyle=\\frac\{L\_\{g\}\}\{2\}\\\|\\bm\{s\}\-\\bm\{s\}^\{\*\}\(r\)\\\|^\{2\}\+\\sum\_\{i=1\}^\{N\}s\_\{i\}\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\.\(4\.6\)For a threshold vector𝒚\\bm\{y\}, writing𝒎\(𝒚\):=\(mi\(yi\)\)i=1N\\bm\{m\}\(\\bm\{y\}\):=\(m\_\{i\}\(y\_\{i\}\)\)\_\{i=1\}^\{N\}, substitution ofsi=mi\(yi\)s\_\{i\}=m\_\{i\}\(y\_\{i\}\)andgi\(mi\(yi\)\)=ℓi\(yi\)g\_\{i\}\(m\_\{i\}\(y\_\{i\}\)\)=\\ell\_\{i\}\(y\_\{i\}\)into \([4\.5](https://arxiv.org/html/2608.14096#S4.E5)\) gives
Δ\(𝒎\(𝒚\),r\)=∑i=1Nℓi\(yi\)−v\(r\)\+λ∗\(r\)\(∑i=1Nmi\(yi\)−r\)\.\\Delta\(\\bm\{m\}\(\\bm\{y\}\);r\)=\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(y\_\{i\}\)\-v\(r\)\+\\lambda^\{\*\}\(r\)\\left\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\-r\\right\)\.\(4\.7\)
Fixt<Tt<T\. The aggregate inventory recursion gives
rt\+1−rt=rt−∑i=1NSi,tnt−1,𝔼\[rt\+1−rt∣ℋt\]=rt−∑i=1Nmi\(Yi,t\)nt−1\.r\_\{t\+1\}\-r\_\{t\}=\\frac\{r\_\{t\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\}\{n\_\{t\}\-1\},\\qquad\\mathbb\{E\}\[r\_\{t\+1\}\-r\_\{t\}\\mid\\mathcal\{H\}\_\{t\}\]=\\frac\{r\_\{t\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\}\{n\_\{t\}\-1\}\.\(4\.8\)The second identity uses that𝒀t\\bm\{Y\}\_\{t\}is chosen before current demand and𝔼\[Si,t∣ℋt\]=mi\(Yi,t\)\\mathbb\{E\}\[S\_\{i,t\}\\mid\\mathcal\{H\}\_\{t\}\]=m\_\{i\}\(Y\_\{i,t\}\)\.
By[Section4\.2](https://arxiv.org/html/2608.14096#S4.SS2)\(iii\), the fluid value is differentiable and its derivative is Lipschitz:
v′\(r\)=−λ∗\(r\),\|v′\(r′\)−v′\(r\)\|≤Lv\|r′−r\|,v^\{\\prime\}\(r\)=\-\\lambda^\{\*\}\(r\),\\qquad\|v^\{\\prime\}\(r^\{\\prime\}\)\-v^\{\\prime\}\(r\)\|\\leq L\_\{v\}\|r^\{\\prime\}\-r\|,with the right derivative understood atr=0r=0\.
Suppose first thatrt≤D¯r\_\{t\}\\leq\\overline\{D\}\. Since bothrtr\_\{t\}and total realized sales belong to\[0,D¯\]\[0,\\overline\{D\}\],
\|rt\+1−rt\|≤D¯nt−1\.\|r\_\{t\+1\}\-r\_\{t\}\|\\leq\\frac\{\\overline\{D\}\}\{n\_\{t\}\-1\}\.ByLvL\_\{v\}\-smoothness, withr=rtr=r\_\{t\}andr′=rt\+1r^\{\\prime\}=r\_\{t\+1\}, and then usingv′\(rt\)=−λ∗\(rt\)v^\{\\prime\}\(r\_\{t\}\)=\-\\lambda^\{\*\}\(r\_\{t\}\), gives
v\(rt\+1\)≤v\(rt\)−λ∗\(rt\)\(rt\+1−rt\)\+Lv2\(rt\+1−rt\)2\.v\(r\_\{t\+1\}\)\\leq v\(r\_\{t\}\)\-\\lambda^\{\*\}\(r\_\{t\}\)\(r\_\{t\+1\}\-r\_\{t\}\)\+\\frac\{L\_\{v\}\}\{2\}\(r\_\{t\+1\}\-r\_\{t\}\)^\{2\}\.Taking the conditional expectation, multiplying bynt−1n\_\{t\}\-1, and using \([4\.8](https://arxiv.org/html/2608.14096#S4.E8)\) and\|rt\+1−rt\|≤D¯/\(nt−1\)\|r\_\{t\+1\}\-r\_\{t\}\|\\leq\\overline\{D\}/\(n\_\{t\}\-1\)give
\(nt−1\)𝔼\[v\(rt\+1\)∣ℋt\]≤\(nt−1\)v\(rt\)\+λ∗\(rt\)\(∑i=1Nmi\(Yi,t\)−rt\)\+LvD¯22\(nt−1\)\.\\displaystyle\(n\_\{t\}\-1\)\\mathbb\{E\}\[v\(r\_\{t\+1\}\)\\mid\\mathcal\{H\}\_\{t\}\]\\leq\(n\_\{t\}\-1\)v\(r\_\{t\}\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\left\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\-r\_\{t\}\\right\)\+\\frac\{L\_\{v\}\\overline\{D\}^\{2\}\}\{2\(n\_\{t\}\-1\)\}\.Adding∑i=1Nℓi\(Yi,t\)−ntv\(rt\)\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(Y\_\{i,t\}\)\-n\_\{t\}v\(r\_\{t\}\)to both sides and using \([4\.7](https://arxiv.org/html/2608.14096#S4.E7)\) with𝒚=𝒀t\\bm\{y\}=\\bm\{Y\}\_\{t\}gives
∑i=1Nℓi\(Yi,t\)\+\(nt−1\)𝔼\[v\(rt\+1\)∣ℋt\]−ntv\(rt\)\\displaystyle\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(Y\_\{i,t\}\)\+\(n\_\{t\}\-1\)\\mathbb\{E\}\[v\(r\_\{t\+1\}\)\\mid\\mathcal\{H\}\_\{t\}\]\-n\_\{t\}v\(r\_\{t\}\)≤∑i=1Nℓi\(Yi,t\)−v\(rt\)\+λ∗\(rt\)\(∑i=1Nmi\(Yi,t\)−rt\)\+LvD¯22\(nt−1\)\\displaystyle\\qquad\\leq\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(Y\_\{i,t\}\)\-v\(r\_\{t\}\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\left\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\-r\_\{t\}\\right\)\+\\frac\{L\_\{v\}\\overline\{D\}^\{2\}\}\{2\(n\_\{t\}\-1\)\}=Δ\(𝒎\(𝒀t\),rt\)\+LvD¯22\(nt−1\)\.\\displaystyle\\qquad=\\Delta\(\\bm\{m\}\(\\bm\{Y\}\_\{t\}\);r\_\{t\}\)\+\\frac\{L\_\{v\}\\overline\{D\}^\{2\}\}\{2\(n\_\{t\}\-1\)\}\.\(4\.9\)Ifrt\>D¯r\_\{t\}\>\\overline\{D\}, then∑i=1NSi,t≤D¯<rt\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\\leq\\overline\{D\}<r\_\{t\}, hencert\+1\>rtr\_\{t\+1\}\>r\_\{t\}\. Bothrtr\_\{t\}andrt\+1r\_\{t\+1\}lie in the slack\-resource region, wherevvis constant andλ∗=0\\lambda^\{\*\}=0\. Thus \([4\.9](https://arxiv.org/html/2608.14096#S4.E9)\) remains valid, in fact with zero remainder\.
Taking expectations and summing \([4\.9](https://arxiv.org/html/2608.14096#S4.E9)\) fromt=1t=1toT−1T\-1telescopes the continuation values:
𝔼∑t=1T−1∑i=1Nℓi\(Yi,t\)\+𝔼v\(rT\)−Tv\(γ\)≤𝔼∑t=1T−1Δ\(𝒎\(𝒀t\),rt\)\+LvD¯22∑t=1T−11nt−1\.\\displaystyle\\mathbb\{E\}\\sum\_\{t=1\}^\{T\-1\}\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(Y\_\{i,t\}\)\+\\mathbb\{E\}v\(r\_\{T\}\)\-Tv\(\\gamma\)\\leq\\mathbb\{E\}\\sum\_\{t=1\}^\{T\-1\}\\Delta\(\\bm\{m\}\(\\bm\{Y\}\_\{t\}\);r\_\{t\}\)\+\\frac\{L\_\{v\}\\overline\{D\}^\{2\}\}\{2\}\\sum\_\{t=1\}^\{T\-1\}\\frac\{1\}\{n\_\{t\}\-1\}\.Becausent−1=T−tn\_\{t\}\-1=T\-t, the last sum—the cumulative discrepancy generated by movements along the re\-solving path—is at mostClogTC\\log T\. Moreover,v\(rT\)≥0v\(r\_\{T\}\)\\geq 0, and the safe box0≤Yi,T≤X¯i0\\leq Y\_\{i,T\}\\leq\\bar\{X\}\_\{i\}gives a primitive upper bound on∑i=1Nℓi\(Yi,T\)−v\(rT\)\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(Y\_\{i,T\}\)\-v\(r\_\{T\}\)\. Adding the last period and usingΔ\(𝒎\(𝒀t\),rt\)≥0\\Delta\(\\bm\{m\}\(\\bm\{Y\}\_\{t\}\);r\_\{t\}\)\\geq 0therefore yields
𝔼∑t=1T∑i=1Nℓi\(Yi,t\)−Tv\(γ\)≤𝔼∑t=1TΔ\(𝒎\(𝒀t\),rt\)\+ClogT\.\\mathbb\{E\}\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(Y\_\{i,t\}\)\-Tv\(\\gamma\)\\leq\\mathbb\{E\}\\sum\_\{t=1\}^\{T\}\\Delta\(\\bm\{m\}\(\\bm\{Y\}\_\{t\}\);r\_\{t\}\)\+C\\log T\.\(4\.10\)By the exact cost reduction \([2\.6](https://arxiv.org/html/2608.14096#S2.E6)\) and the regret definition \([2\.9](https://arxiv.org/html/2608.14096#S2.E9)\), the left side of \([4\.10](https://arxiv.org/html/2608.14096#S4.E10)\) isRT\(𝖱𝖠𝖯𝖣𝖫\)R\_\{T\}\(\\mathsf\{RAPDL\}\)\.
Separating the learning and implementation gaps\.The functionmim\_\{i\}is11\-Lipschitz, and
\|ℓi′\(y\)\|≤max\{hi,bi−ci\}\(0≤y≤X¯i\)\.\|\\ell\_\{i\}^\{\\prime\}\(y\)\|\\leq\\max\\\{h\_\{i\},b\_\{i\}\-c\_\{i\}\\\}\\qquad\(0\\leq y\\leq\\bar\{X\}\_\{i\}\)\.Together with0≤λ∗\(rt\)≤λmax0\\leq\\lambda^\{\*\}\(r\_\{t\}\)\\leq\\lambda\_\{\\max\}, these Lipschitz bounds give
Δ\(𝒎\(𝒀t\),rt\)≤Δ\(𝒎\(𝑿t\),rt\)\+C‖𝒀t−𝑿t‖1\.\\Delta\(\\bm\{m\}\(\\bm\{Y\}\_\{t\}\);r\_\{t\}\)\\leq\\Delta\(\\bm\{m\}\(\\bm\{X\}\_\{t\}\);r\_\{t\}\)\+C\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\.Applying \([4\.6](https://arxiv.org/html/2608.14096#S4.E6)\) to the preferred action then yields
Δ\(𝒎\(𝒀t\),rt\)≤C\(‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+⟨𝒎\(𝑿t\)−𝒔∗\(rt\),∇𝒔ℒrt\(𝒔∗\(rt\),λ∗\(rt\)\)⟩\+‖𝒀t−𝑿t‖1\)\.\\displaystyle\\Delta\(\\bm\{m\}\(\\bm\{Y\}\_\{t\}\);r\_\{t\}\)\\leq C\\Bigl\(\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\\left\\langle\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\),\\nabla\_\{\\bm\{s\}\}\\mathcal\{L\}\_\{r\_\{t\}\}\(\\bm\{s\}^\{\*\}\(r\_\{t\}\),\\lambda^\{\*\}\(r\_\{t\}\)\)\\right\\rangle\+\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\\Bigr\)\.\(4\.11\)
Taking expectations in \([4\.11](https://arxiv.org/html/2608.14096#S4.E11)\), summing, and using \([4\.3](https://arxiv.org/html/2608.14096#S4.E3)\)–\([4\.4](https://arxiv.org/html/2608.14096#S4.E4)\) gives
𝔼∑t=1TΔ\(𝒎\(𝒀t\),rt\)≤C\(𝖯𝖣T\+𝖨𝖬𝖯T\)\.\\mathbb\{E\}\\sum\_\{t=1\}^\{T\}\\Delta\(\\bm\{m\}\(\\bm\{Y\}\_\{t\}\);r\_\{t\}\)\\leq C\(\\mathsf\{PD\}\_\{T\}\+\\mathsf\{IMP\}\_\{T\}\)\.\(4\.12\)Substituting \([4\.12](https://arxiv.org/html/2608.14096#S4.E12)\) into \([4\.10](https://arxiv.org/html/2608.14096#S4.E10)\) and absorbing fixed constants intoCCyields the three\-component regret reduction
RT\(𝖱𝖠𝖯𝖣𝖫\)≤C\(logT\+𝖯𝖣T\+𝖨𝖬𝖯T\)\.R\_\{T\}\(\\mathsf\{RAPDL\}\)\\leq C\\left\(\\log T\+\\mathsf\{PD\}\_\{T\}\+\\mathsf\{IMP\}\_\{T\}\\right\)\.\(4\.13\)
### 4\.4Step 2: Bounding the Learning Gap
Under \([A\.25](https://arxiv.org/html/2608.14096#A1.E25)\), we start from the definition in \([4\.3](https://arxiv.org/html/2608.14096#S4.E3)\)\. Since
∇𝒔ℒrt\(𝒔∗\(rt\),λ∗\(rt\)\)=\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)i=1N,\\nabla\_\{\\bm\{s\}\}\\mathcal\{L\}\_\{r\_\{t\}\}\(\\bm\{s\}^\{\*\}\(r\_\{t\}\),\\lambda^\{\*\}\(r\_\{t\}\)\)=\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\_\{i=1\}^\{N\},the storewise complementarity conditions \([4\.2](https://arxiv.org/html/2608.14096#S4.E2)\) give
𝖯𝖣T=\\displaystyle\\mathsf\{PD\}\_\{T\}=\{\}∑t=1T𝔼‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+∑t=1T𝔼∑i=1N\(mi\(Xi,t\)−si∗\(rt\)\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(X\_\{i,t\}\)\-s\_\{i\}^\{\*\}\(r\_\{t\}\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)=\\displaystyle=\{\}∑t=1T𝔼‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+∑t=1T𝔼∑i=1Nmi\(Xi,t\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\.\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i,t\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\.\(4\.14\)We now invoke[AppendixA](https://arxiv.org/html/2608.14096#A1)\. Its tracking conclusion \([A\.1](https://arxiv.org/html/2608.14096#A1.E1)\), after dropping the nonnegative squared dual error, gives
∑t=1T𝔼‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2≤ClogT\+C∑t=1T𝔼‖𝒀t−𝑿t‖2\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\\leq C\\log T\+C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(4\.15\)Its residual conclusion \([A\.2](https://arxiv.org/html/2608.14096#A1.E2)\) gives
∑t=1T𝔼\[∑i=1Nmi\(Xi,t\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\]≤ClogT\+C∑t=1T𝔼‖𝒀t−𝑿t‖2\.\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\!\\left\[\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i,t\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\\right\]\\leq C\\log T\+C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(4\.16\)Thus \([4\.16](https://arxiv.org/html/2608.14096#S4.E16)\) directly bounds the second sum in \([4\.14](https://arxiv.org/html/2608.14096#S4.E14)\)\. Adding it to \([4\.15](https://arxiv.org/html/2608.14096#S4.E15)\), and using
∑t=1T𝔼‖𝒀t−𝑿t‖2≤𝖨𝖬𝖯T,\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\\leq\\mathsf\{IMP\}\_\{T\},we obtain the learning\-gap bound
𝖯𝖣T≤C\(logT\+𝖨𝖬𝖯T\)\.\\mathsf\{PD\}\_\{T\}\\leq C\\left\(\\log T\+\\mathsf\{IMP\}\_\{T\}\\right\)\.\(4\.17\)
### 4\.5Step 3: Bounding the Implementation Gap
The implementation gap arises solely from physical feasibility: the implemented action must satisfyYi,t≥Ii,tY\_\{i,t\}\\geq I\_\{i,t\}for every store and∑iYi,t≤Ct\\sum\_\{i\}Y\_\{i,t\}\\leq C\_\{t\}, whereas the preferred action𝑿t\\bm\{X\}\_\{t\}need not satisfy these two restrictions\.
The following physical bound controls this component\.
###### Lemma 4\.4\.
Under the conditions of[Theorem4\.1](https://arxiv.org/html/2608.14096#S4.Thmtheorem1),𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}satisfies
𝖨𝖬𝖯T≤ClogT\.\\mathsf\{IMP\}\_\{T\}\\leq C\\log T\.\(4\.18\)
Combining the re\-solving reduction \([4\.13](https://arxiv.org/html/2608.14096#S4.E13)\), the learning bound \([4\.17](https://arxiv.org/html/2608.14096#S4.E17)\), and the implementation bound \([4\.18](https://arxiv.org/html/2608.14096#S4.E18)\) gives directly
𝖨𝖬𝖯T\\displaystyle\\mathsf\{IMP\}\_\{T\}≤ClogT,𝖯𝖣T≤C\(logT\+𝖨𝖬𝖯T\)≤ClogT,\\displaystyle\\leq C\\log T,\\qquad\\mathsf\{PD\}\_\{T\}\\leq C\\bigl\(\\log T\+\\mathsf\{IMP\}\_\{T\}\\bigr\)\\leq C\\log T,RT\(𝖱𝖠𝖯𝖣𝖫\)\\displaystyle R\_\{T\}\(\\mathsf\{RAPDL\}\)≤C\(logT\+𝖯𝖣T\+𝖨𝖬𝖯T\)≤ClogT\.\\displaystyle\\leq C\\bigl\(\\log T\+\\mathsf\{PD\}\_\{T\}\+\\mathsf\{IMP\}\_\{T\}\\bigr\)\\leq C\\log T\.This completes the proof of[Theorem4\.1](https://arxiv.org/html/2608.14096#S4.Thmtheorem1)\.\\square\\square
## 5Numerical Experiments
Experimental design\.The numerical study evaluates the finite\-horizon performance of𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}across horizon lengths and inventory regimes and compares it with the Double Binary Search \(DBS\) algorithm of[2](https://arxiv.org/html/2608.14096#bib.bib19), the closest directly comparable learning benchmark\. We use matched demand paths, initialization, tuning budgets, and feasibility constraints\.
We use
ρ∈\{0\.30,0\.60,0\.90,1\.20\},W=Tργ⋆,\\rho\\in\\\{0\.30,0\.60,0\.90,1\.20\\\},\\qquad W=T\\rho\\gamma^\{\\star\},whereγ⋆\\gamma^\{\\star\}is unconstrained optimal expected sales\. These values represent four distinct operating regimes—severe scarcity, moderate scarcity, near sufficiency, and abundant inventory\. In S2 they also cross fluid active\-set sizes three, five, and all six stores\. Total Cost is reported forT∈\{100,200,300\}T\\in\\\{100,200,300\\\}, while the figures useT=20,40,…,200T=20,40,\\ldots,200, each with newly initialized inventory\. In both settings the demand vectors are independent and identically distributed over time\. Stores are independent within a period in S1, whereas S2 introduces the contemporaneous dependence specified below\. All calculations use full precision\.
The experimental label𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}denotes the two\-rate variant in[Section3\.3](https://arxiv.org/html/2608.14096#S3.SS3)\. The label DBS denotes Double Binary Search in[2](https://arxiv.org/html/2608.14096#bib.bib19)\. Because S1 is homogeneous and satisfies their Assumption 2 by symmetry, its benchmark is precisely their DBS policy in Algorithm 2\. S2 is heterogeneous and may have stores outside the fluid active set\. Its benchmark is therefore their heterogeneous active\-set implementation: Algorithm 2 equipped with the relaxed\-Assumption 2 extension described in their online appendix\. For a fair comparison, both policies use the same midpoint initialization,λ1=0\\lambda\_\{1\}=0, demand paths, and feasibility constraints\. For𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\},Xi,1=d¯i/2X\_\{i,1\}=\\bar\{d\}\_\{i\}/2\. Each policy is calibrated offline on independent demand paths for every\(setting,ρ,T\)\(\\text\{setting\},\\rho,T\)cell under the same tuning budget\. The selected parameters are then locked and evaluated on untouched paired paths\.
S1: homogeneous setting\.There areN=2N=2identical stores, with
𝒉=\(6\.00,6\.00\),𝒃=\(60\.00,60\.00\),𝒄=\(0\.50,0\.50\)\.\\bm\{h\}=\(6\.00,6\.00\),\\qquad\\bm\{b\}=\(60\.00,60\.00\),\\qquad\\bm\{c\}=\(0\.50,0\.50\)\.For every store and period, demand is a𝒩\(50\.00,50\.002\)\\mathcal\{N\}\(50\.00,50\.00^\{2\}\)random variable conditioned to lie in\[0,175\.00\]\[0,175\.00\]\. WithΦ\\Phidenoting the standard\-normal distribution function, its common distribution function is
FS1\(d\)=\{0,d<0,Φ\(\(d−50\.00\)/50\.00\)−Φ\(−1\.00\)Φ\(2\.50\)−Φ\(−1\.00\),0≤d≤175\.00,1,d\>175\.00\.F\_\{\\mathrm\{S1\}\}\(d\)=\\begin\{cases\}0,&d<0,\\\\ \\displaystyle\\frac\{\\Phi\(\(d\-50\.00\)/50\.00\)\-\\Phi\(\-1\.00\)\}\{\\Phi\(2\.50\)\-\\Phi\(\-1\.00\)\},&0\\leq d\\leq 175\.00,\\\\ 1,&d\>175\.00\.\\end\{cases\}This is the homogeneous synthetic specification used by[2](https://arxiv.org/html/2608.14096#bib.bib19)\. Its unconstrained expected\-sales normalizer isγ⋆=123\.43\\gamma^\{\\star\}=123\.43\.
S2: heterogeneous setting\.There areN=6N=6stores, with store\-specific holding, net shipment, and lost\-sales coefficients:
𝒉\\displaystyle\\bm\{h\}=\(2\.28,2\.46,2\.52,2\.42,2\.34,2\.38\),\\displaystyle=\(2\.28,2\.46,2\.52,2\.42,2\.34,2\.38\),𝒄\\displaystyle\\bm\{c\}=\(0\.195,0\.190,0\.205,0\.220,0\.210,0\.180\),\\displaystyle=\(0\.195,0\.190,0\.205,0\.220,0\.210,0\.180\),𝒃\\displaystyle\\bm\{b\}=\(38\.40,40\.00,41\.60,43\.20,44\.80,46\.40\)\.\\displaystyle=\(38\.40,40\.00,41\.60,43\.20,44\.80,46\.40\)\.For storeiiand each period, demand isUniform\[0,d¯i\]\\operatorname\{Uniform\}\[0,\\bar\{d\}\_\{i\}\], where
𝒅¯=\(1\.18,1\.26,1\.08,1\.32,1\.14,1\.22\)\.\\bar\{\\bm\{d\}\}=\(1\.18,1\.26,1\.08,1\.32,1\.14,1\.22\)\.Its distribution function is
FS2,i\(d\)=\{0,d<0,d/d¯i,0≤d≤d¯i,1,d\>d¯i\.F\_\{\\mathrm\{S2\},i\}\(d\)=\\begin\{cases\}0,&d<0,\\\\ d/\\bar\{d\}\_\{i\},&0\\leq d\\leq\\bar\{d\}\_\{i\},\\\\ 1,&d\>\\bar\{d\}\_\{i\}\.\\end\{cases\}Thus both the cost coefficients and demand laws are store\-specific\. The six Uniform marginals are positively correlated within each period through a Gaussian copula\. Specifically, letUtU\_\{t\}andεi,t\\varepsilon\_\{i,t\}be independent standard\-normal variables and set
Zi,t=0\.60Ut\+0\.40εi,t,Di,t=d¯iΦ\(Zi,t\)\.Z\_\{i,t\}=\\sqrt\{0\.60\}\\,U\_\{t\}\+\\sqrt\{0\.40\}\\,\\varepsilon\_\{i,t\},\\qquad D\_\{i,t\}=\\bar\{d\}\_\{i\}\\Phi\(Z\_\{i,t\}\)\.A freshUtU\_\{t\}and fresh idiosyncratic shocks are drawn every period\. Hence the demand vectors remain independent and identically distributed over time, while the stores experience a positive common within\-period shock\. The resulting unconstrained expected\-sales normalizer isγ⋆=3\.59\\gamma^\{\\star\}=3\.59\. This correlated\-demand specification is admissible under[Section2](https://arxiv.org/html/2608.14096#S2)and shows that our framework handles cross\-store dependence\. It is also more practical because retail stores may share weather, holiday, and market shocks\.
Performance metrics and results\.For policyπ\\pi, define the realized store–period cost by
𝖢i,tπ:=ciSi,tπ\+hi\(Yi,tπ−Di,t\)\+\+bi\(Di,t−Yi,tπ\)\+,Si,tπ:=Di,t∧Yi,tπ\.\\mathsf\{C\}\_\{i,t\}^\{\\pi\}:=c\_\{i\}S\_\{i,t\}^\{\\pi\}\+h\_\{i\}\(Y\_\{i,t\}^\{\\pi\}\-D\_\{i,t\}\)^\{\+\}\+b\_\{i\}\(D\_\{i,t\}\-Y\_\{i,t\}^\{\\pi\}\)^\{\+\},\\qquad S\_\{i,t\}^\{\\pi\}:=D\_\{i,t\}\\wedge Y\_\{i,t\}^\{\\pi\}\.The terms are net shipment, holding, and lost\-sales costs\. Herew=0w=0, soci=kic\_\{i\}=k\_\{i\}\. Equation \([2\.6](https://arxiv.org/html/2608.14096#S2.E6)\) makes this sales\-based accounting equivalent to physical shipment cost net of terminal store credit\.
The realized*Total Cost*of one completeTT\-period path is
𝖳𝖢Tπ:=∑t=1T∑i=1N𝖢i,tπ\.\\mathsf\{TC\}\_\{T\}^\{\\pi\}:=\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\\mathsf\{C\}\_\{i,t\}^\{\\pi\}\.
For each cell,[Table1](https://arxiv.org/html/2608.14096#S5.T1)reports mean Total Cost overM=2,000M=2\{,\}000post\-tuning paths and the paired𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}\-minus\-DBS difference\.
The horizon\-by\-horizon Relative Regret results for S1 and S2 are shown in[Figures1](https://arxiv.org/html/2608.14096#S5.F1)and[2](https://arxiv.org/html/2608.14096#S5.F2), respectively\. For these figures, setγ:=W/T=ργ⋆\\gamma:=W/T=\\rho\\gamma^\{\\star\}\. The quantityv\(γ\)v\(\\gamma\)is the one\-period constrained fluid lower\-bound value, soTv\(γ\)Tv\(\\gamma\)is the correspondingTT\-period benchmark\. We define the estimated*Relative Regret*by
RR^Tπ:=100×𝔼^\[𝖳𝖢Tπ\]−Tv\(γ\)Tv\(γ\)\.\\widehat\{\\operatorname\{RR\}\}\_\{T\}^\{\\pi\}:=100\\times\\frac\{\\widehat\{\\mathbb\{E\}\}\[\\mathsf\{TC\}\_\{T\}^\{\\pi\}\]\-Tv\(\\gamma\)\}\{Tv\(\\gamma\)\}\.
Thus a plotted value of2\.002\.00denotes a mean cost2\.00%2\.00\\%above the fluid benchmark\.
For each\(setting,ρ,T,π\)\(\\text\{setting\},\\rho,T,\\pi\)cell, the figures show 1,000 bootstrap resamples of the mean Relative Regret: boxes span the 25th–75th percentiles and whiskers the 5th–95th\. These intervals concern the mean estimator, and each horizon is a newly initialized problem rather than a prefix of a longer simulation\.
Table 1:Mean Total Cost over 2,000 paired production paths\.As shown in[Table1](https://arxiv.org/html/2608.14096#S5.T1), RAPDL has lower mean Total Cost in all 24 paired finite\-horizon comparisons\.
\(a\)ρ=0\.30\\rho=0\.30
\(b\)ρ=0\.60\\rho=0\.60
\(c\)ρ=0\.90\\rho=0\.90
\(d\)ρ=1\.20\\rho=1\.20
Figure 1:Relative Regret across complete horizons in S1\.\(a\)ρ=0\.30\\rho=0\.30
\(b\)ρ=0\.60\\rho=0\.60
\(c\)ρ=0\.90\\rho=0\.90
\(d\)ρ=1\.20\\rho=1\.20
Figure 2:Relative Regret across complete horizons in S2\.The figures place both methods close to the fluid benchmark at practical horizons\. AtT=100T=100, the bootstrap median Relative Regret of𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}ranges from0\.30%0\.30\\%to12\.38%12\.38\\%\. ByT=200T=200, the range narrows to0\.15%0\.15\\%–6\.58%6\.58\\%\. At the short horizonsT=60T=60andT=100T=100, the bootstrap median for𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}is below that for DBS in all eight panels\. The very earliestT=20T=20andT=40T=40comparisons are mixed, and these outcomes are retained in the plots\. In theT∈\{100,200,300\}T\\in\\\{100,200,300\\\}cells reported in[Table1](https://arxiv.org/html/2608.14096#S5.T1),𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}has lower mean Total Cost in all 24 paired comparisons\.
Sensitivity to endpoint information\.We test the endpoint\-envelope implementation in S1 atT=200T=200\. The true demand support remains\[0,175\]\[0,175\]\. The exact implementation usesX¯i=170\.21\\bar\{X\}\_\{i\}=170\.21,D¯=350\\overline\{D\}=350, andXi,1=87\.5X\_\{i,1\}=87\.5, whereas the envelope implementation usesDiup=200D\_\{i\}^\{\\mathrm\{up\}\}=200,Dup=400D^\{\\mathrm\{up\}\}=400, andXi,1=100X\_\{i,1\}=100\. For eachρ\\rho, both information regimes use the same tuning budget and 2,000 paired production demands\. This comparison evaluates two separately calibrated end\-to\-end implementations with regime\-specific midpoint initializations, rather than isolating the endpoint input under fixed parameters\. The results are reported in[Table2](https://arxiv.org/html/2608.14096#S5.T2)\. Its last column gives the percentage difference100×\(𝔼^\[TC200Envelope\]−𝔼^\[TC200Exact\]\)/𝔼^\[TC200Exact\]100\\times\\bigl\(\\widehat\{\\mathbb\{E\}\}\[\\mathrm\{TC\}^\{\\mathrm\{Envelope\}\}\_\{200\}\]\-\\widehat\{\\mathbb\{E\}\}\[\\mathrm\{TC\}^\{\\mathrm\{Exact\}\}\_\{200\}\]\\bigr\)/\\widehat\{\\mathbb\{E\}\}\[\\mathrm\{TC\}^\{\\mathrm\{Exact\}\}\_\{200\}\]\.
Table 2:Sensitivity of S1 Total Cost to endpoint information atT=200T=200\.Across the four resource levels, the absolute Total Cost difference between the exact endpoint and the rough envelope is at most1\.24%1\.24\\%\. Thus, in these finite\-horizon comparisons, replacing the exact endpoint by a conservative upper bound has only a small numerical effect\.
## 6Conclusion
This paper studies learning in a finite\-horizon OWMS system with nonreplenishable shared inventory, unknown demand, and censored sales\. We propose𝖱𝖠𝖯𝖣𝖫\\mathsf\{RAPDL\}, a resource\-adaptive Primal\-Dual framework that replaces fixed\-target learning with online tracking of the endogenous Primal\-Dual re\-solving path\. Censored sales update the preferred store thresholds and warehouse scarcity price, while water\-filling maps the learned state to feasible physical actions\. The regret analysis separates the re\-solving, learning, and implementation gaps and bounds each at logarithmic order\. Numerical experiments on a practical variant show strong finite\-horizon performance across inventory regimes\.
More broadly, the framework turns re\-solving from repeated full\-information optimization into a path that can be tracked incrementally under unknown and censored feedback\. This resource\-adaptive perspective may inform the design of learning\-and\-control methods for other systems with depleting shared resources\.
## References
- Agrawalet al\.\(2014\)S\. Agrawal, Z\. Wang, and Y\. YeA dynamic near\-optimal algorithm for online linear programming\.Operations Research62\(4\),pp\. 876–890\.External Links:[Document](https://dx.doi.org/10.1287/opre.2014.1289)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p1.1)\.
- Bekciet al\.\(2023\)R\. Y\. Bekci, M\. Gümüş, and S\. MiaoInventory control and learning for one\-warehouse multistore system with censored demand\.Operations Research71\(6\),pp\. 2092–2110\.External Links:[Document](https://dx.doi.org/10.1287/opre.2021.0694)Cited by:[item 3](https://arxiv.org/html/2608.14096#S1.I1.i3.p1.1),[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p4.1),[§1](https://arxiv.org/html/2608.14096#S1.p1.1),[§1](https://arxiv.org/html/2608.14096#S1.p3.1),[Remark 2\.2](https://arxiv.org/html/2608.14096#S2.p5.1.1),[§5](https://arxiv.org/html/2608.14096#S5.p1.1),[§5](https://arxiv.org/html/2608.14096#S5.p3.1),[§5](https://arxiv.org/html/2608.14096#S5.p4.3),[footnote 1](https://arxiv.org/html/2608.14096#footnote1)\.
- Benziet al\.\(2005\)M\. Benzi, G\. H\. Golub, and J\. LiesenNumerical solution of saddle point problems\.Acta Numerica14,pp\. 1–137\.External Links:[Document](https://dx.doi.org/10.1017/S0962492904000212)Cited by:[§3\.2](https://arxiv.org/html/2608.14096#S3.SS2.p6.1)\.
- Caro and Gallien \(2010\)F\. Caro and J\. GallienInventory management of a fast\-fashion retail network\.Operations Research58\(2\),pp\. 257–273\.External Links:[Document](https://dx.doi.org/10.1287/opre.1090.0698)Cited by:[§1](https://arxiv.org/html/2608.14096#S1.p1.1)\.
- Chaoet al\.\(2025\)X\. Chao, S\. Jasin, and S\. MiaoAdaptive lagrangian policies for a multiwarehouse, multistore inventory system with lost sales\.Operations Research73\(3\),pp\. 1615–1636\.External Links:[Document](https://dx.doi.org/10.1287/opre.2022.0668)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p3.1)\.
- Chen and Zheng \(1997\)F\. Chen and Y\. ZhengOne\-warehouse multiretailer systems with centralized stock information\.Operations Research45\(2\),pp\. 275–287\.External Links:[Document](https://dx.doi.org/10.1287/opre.45.2.275)Cited by:[§1](https://arxiv.org/html/2608.14096#S1.p1.1)\.
- Chen and Gallego \(2022\)N\. Chen and G\. GallegoA primal–dual learning algorithm for personalized dynamic pricing with an inventory constraint\.Mathematics of Operations Research47\(4\),pp\. 2585–2613\.External Links:[Document](https://dx.doi.org/10.1287/moor.2021.1220)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p2.1)\.
- Chenet al\.\(2020\)W\. Chen, C\. Shi, and I\. DuenyasOptimal learning algorithms for stochastic inventory systems with random capacities\.Production and Operations Management29\(7\),pp\. 1624–1649\.External Links:[Document](https://dx.doi.org/10.1111/poms.13178)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1)\.
- Chenet al\.\(2024\)X\. Chen, J\. Lyu, Y\. Wang, and Y\. ZhouNetwork revenue management with demand learning and fair resource\-consumption balancing\.Production and Operations Management33\(2\),pp\. 494–511\.External Links:[Document](https://dx.doi.org/10.1177/10591478231225176)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p2.1)\.
- Duchiet al\.\(2008\)J\. Duchi, S\. Shalev\-Shwartz, Y\. Singer, and T\. ChandraEfficient projections onto theℓ1\\ell\_\{1\}\-ball for learning in high dimensions\.InProceedings of the 25th International Conference on Machine Learning,pp\. 272–279\.External Links:[Document](https://dx.doi.org/10.1145/1390156.1390191)Cited by:[§3\.1](https://arxiv.org/html/2608.14096#S3.SS1.p2.3)\.
- Fanet al\.\(2023\)X\. Fan, B\. Chen, W\. Xiao, and Z\. ZhouNo\-regret learning in multi\-retailer inventory control\.Technical reportTechnical Report4626023,SSRN\.External Links:[Document](https://dx.doi.org/10.2139/ssrn.4626023)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1)\.
- Guoet al\.\(2026\)S\. Guo, C\. Shi, C\. Yang, and C\. ZachariasAn online mirror descent learning algorithm for multiproduct inventory systems\.Operations Research\.Note:Articles in AdvanceExternal Links:[Document](https://dx.doi.org/10.1287/opre.2024.0982)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1)\.
- Huanget al\.\(2025\)J\. Huang, K\. Shang, Y\. Yang, W\. Zhou, and Y\. LiTaylor approximation of inventory policies for one\-warehouse, multi\-retailer systems with demand feature information\.Management Science71\(1\),pp\. 879–897\.External Links:[Document](https://dx.doi.org/10.1287/mnsc.2021.04241)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1),[§1](https://arxiv.org/html/2608.14096#S1.p1.1)\.
- Huhet al\.\(2009\)W\. T\. Huh, G\. Janakiraman, J\. A\. Muckstadt, and P\. RusmevichientongAn adaptive algorithm for finding the optimal base\-stock policy in lost sales inventory systems with censored demand\.Mathematics of Operations Research34\(2\),pp\. 397–416\.External Links:[Document](https://dx.doi.org/10.1287/moor.1080.0367)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1)\.
- Huh and Rusmevichientong \(2009\)W\. T\. Huh and P\. RusmevichientongA nonparametric asymptotic analysis of inventory planning with censored demand\.Mathematics of Operations Research34\(1\),pp\. 103–123\.External Links:[Document](https://dx.doi.org/10.1287/moor.1080.0355)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1),[§1](https://arxiv.org/html/2608.14096#S1.p3.1),[Remark 2\.2](https://arxiv.org/html/2608.14096#S2.p5.1.1),[footnote 1](https://arxiv.org/html/2608.14096#footnote1)\.
- Jackson \(1988\)P\. L\. JacksonStock allocation in a two\-echelon distribution system or “what to do until your ship comes in”\.Management Science34\(7\),pp\. 880–895\.External Links:[Document](https://dx.doi.org/10.1287/mnsc.34.7.880)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p3.1),[§1](https://arxiv.org/html/2608.14096#S1.p1.1)\.
- Jasin and Kumar \(2012\)S\. Jasin and S\. KumarA re\-solving heuristic with bounded revenue loss for network revenue management with customer choice\.Mathematics of Operations Research37\(2\),pp\. 313–345\.External Links:[Document](https://dx.doi.org/10.1287/moor.1120.0537)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p1.1)\.
- Jasin \(2014\)S\. JasinReoptimization and self\-adjusting price control for network revenue management\.Operations Research62\(5\),pp\. 1168–1178\.External Links:[Document](https://dx.doi.org/10.1287/opre.2014.1297)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p1.1)\.
- Kutanoglu and Mahajan \(2009\)E\. Kutanoglu and M\. MahajanAn inventory sharing and allocation method for a multi\-location service parts logistics network with time\-based service levels\.European Journal of Operational Research194\(3\),pp\. 728–742\.External Links:[Document](https://dx.doi.org/10.1016/j.ejor.2007.12.032)Cited by:[§1](https://arxiv.org/html/2608.14096#S1.p1.1)\.
- Latafatet al\.\(2019\)P\. Latafat, N\. M\. Freris, and P\. PatrinosA new randomized block\-coordinate primal–dual proximal algorithm for distributed optimization\.IEEE Transactions on Automatic Control64\(10\),pp\. 4050–4065\.External Links:[Document](https://dx.doi.org/10.1109/TAC.2019.2906924)Cited by:[§3\.2](https://arxiv.org/html/2608.14096#S3.SS2.p6.1)\.
- Liet al\.\(2024\)G\. Li, Z\. Wang, and J\. ZhangInfrequent resolving algorithm for online linear programming\.Note:arXiv preprint arXiv:2408\.00465External Links:2408\.00465Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p1.1)\.
- Li and Ye \(2022\)X\. Li and Y\. YeOnline linear programming: dual convergence, new algorithms, and regret bounds\.Operations Research70\(5\),pp\. 2948–2966\.External Links:[Document](https://dx.doi.org/10.1287/opre.2021.2164)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p1.1)\.
- Lyuet al\.\(2025\)J\. Lyu, J\. Xie, S\. Yuan, and Y\. ZhouA minibatch stochastic gradient descent\-based learning metapolicy for inventory systems with myopic optimal policy\.Management Science71\(7\),pp\. 5572–5588\.External Links:[Document](https://dx.doi.org/10.1287/mnsc.2023.00920)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1),[§1](https://arxiv.org/html/2608.14096#S1.p3.1),[Remark 2\.2](https://arxiv.org/html/2608.14096#S2.p5.1.1)\.
- Marklund and Rosling \(2012\)J\. Marklund and K\. RoslingLower bounds and heuristics for supply chain stock allocation\.Operations Research60\(1\),pp\. 92–105\.External Links:[Document](https://dx.doi.org/10.1287/opre.1110.1009)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p3.1)\.
- Mehrotraet al\.\(2020\)S\. Mehrotra, H\. Rahimian, M\. Barah, F\. Luo, and K\. SchantzA model of supply\-chain decisions for resource sharing with an application to ventilator allocation to combat COVID\-19\.Naval Research Logistics67\(5\),pp\. 303–320\.External Links:[Document](https://dx.doi.org/10.1002/nav.21905)Cited by:[§1](https://arxiv.org/html/2608.14096#S1.p1.1)\.
- Miaoet al\.\(2022\)S\. Miao, S\. Jasin, and X\. ChaoAsymptotically optimal lagrangian policies for multi\-warehouse, multi\-store systems with lost sales\.Operations Research70\(1\),pp\. 141–159\.External Links:[Document](https://dx.doi.org/10.1287/opre.2021.2161)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p3.1)\.
- Miaoet al\.\(2026\)S\. Miao, Y\. Wang, and J\. ZhangA primal–dual approach toward resource\-constrained revenue management with demand learning and large action space\.Operations Research74\(2\),pp\. 825–839\.External Links:[Document](https://dx.doi.org/10.1287/opre.2021.0483)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p2.1)\.
- Miaoet al\.\(2023\)S\. Miao, Y\. Wang, and R\. ZhaoDynamic learning policy for multi\-warehouse multi\-store systems with censored demands\.Technical reportTechnical Report4617620,SSRN\.External Links:[Document](https://dx.doi.org/10.2139/ssrn.4617620)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p4.1),[§1](https://arxiv.org/html/2608.14096#S1.p3.1),[footnote 1](https://arxiv.org/html/2608.14096#footnote1)\.
- Miao and Wang \(2024\)S\. Miao and Y\. WangDemand balancing in primal–dual optimization for blind network revenue management\.Note:arXiv preprint arXiv:2404\.04467External Links:2404\.04467Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p2.1)\.
- Miao and Wang \(2025\)S\. Miao and Y\. WangNetwork revenue management with nonparametric demand learning:T\\sqrt\{T\}\-regret and polynomial dimension dependency\.Mathematics of Operations Research\.Note:Articles in AdvanceExternal Links:[Document](https://dx.doi.org/10.1287/moor.2022.0086)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p2.1)\.
- Shiet al\.\(2016\)C\. Shi, W\. Chen, and I\. DuenyasTechnical note—nonparametric data\-driven algorithms for multiproduct inventory systems with censored demand\.Operations Research64\(2\),pp\. 362–370\.External Links:[Document](https://dx.doi.org/10.1287/opre.2015.1474)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1),[§1](https://arxiv.org/html/2608.14096#S1.p3.1),[Remark 2\.2](https://arxiv.org/html/2608.14096#S2.p5.1.1),[footnote 1](https://arxiv.org/html/2608.14096#footnote1)\.
- Zhanget al\.\(2018\)H\. Zhang, X\. Chao, and C\. ShiTechnical note—perishable inventory systems: convexity results for base\-stock policies and learning algorithms under censored demand\.Operations Research66\(5\),pp\. 1276–1286\.External Links:[Document](https://dx.doi.org/10.1287/opre.2018.1724)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1),[§1](https://arxiv.org/html/2608.14096#S1.p3.1),[Remark 2\.2](https://arxiv.org/html/2608.14096#S2.p5.1.1)\.
- Zhanget al\.\(2020\)H\. Zhang, X\. Chao, and C\. ShiClosing the gap: a learning algorithm for lost\-sales inventory systems with lead times\.Management Science66\(5\),pp\. 1962–1980\.External Links:[Document](https://dx.doi.org/10.1287/mnsc.2019.3288)Cited by:[§1\.2](https://arxiv.org/html/2608.14096#S1.SS2.p5.1)\.
Online Appendix for “Resource\-Adaptive Primal\-Dual Learning for One\-Warehouse Multi\-Store Systems with Censored Demand”
## Appendix ATechnical Bound for the Learning Gap
The following estimate supplies the squared\-tracking and first\-order residual bounds used to control the learning gap\. Both follow from the same one\-step mirror recursion developed below\.
###### Lemma A\.1\.
Chooseχ,τ0\\chi,\\tau\_\{0\}as in \([A\.25](https://arxiv.org/html/2608.14096#A1.E25)\)\. Then there is a primitive constantCmdC\_\{\\rm md\}such that
*\(i\) Cumulative Primal\-Dual tracking\.*
∑t=1T𝔼\[‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+\|λt−λ∗\(rt\)\|2\]\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\left\[\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\|\\lambda\_\{t\}\-\\lambda^\{\*\}\(r\_\{t\}\)\|^\{2\}\\right\]≤CmdlogT\+Cmd∑t=1T𝔼‖𝒀t−𝑿t‖2,\\displaystyle\\leq C\_\{\\rm md\}\\log T\+C\_\{\\rm md\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\},\(A\.1\)
*\(ii\) Unweighted first\-order residual\.*
∑t=1T𝔼\[∑i=1Nmi\(Xi,t\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\]\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\!\\left\[\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i,t\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\\right\]≤CmdlogT\+Cmd∑t=1T𝔼‖𝒀t−𝑿t‖2\.\\displaystyle\\leq C\_\{\\rm md\}\\log T\+C\_\{\\rm md\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(A\.2\)
The proof proceeds in four steps: construct the exact\-target potential, smooth the moving KKT path, establish one\-step drift, and sum the recursion\. Supporting proofs and local verifications are collected in Appendix[A\.5](https://arxiv.org/html/2608.14096#A1.SS5)\.
### A\.1Step 1: Construct the Exact\-Target Potential
Step 1 verifies that capping preserves the KKT target and places its threshold representation in the moving box, then constructs a threshold\-space potential for exact sales\-space error\.
The next lemma verifies exact\-target feasibility\. Its complementarity identities are also used in the state\-update analysis\.
###### Lemma A\.2\.
For everyr≥0r\\geq 0, the selected moving target has the following feasibility properties\.
1. \(i\)Target preservation under capping\. \(𝒔∗\(r\),λ∗\(r\)\)=\(𝒔∗\(rc\),λ∗\(rc\)\)\.\(\\bm\{s\}^\{\*\}\(r\),\\lambda^\{\*\}\(r\)\)=\(\\bm\{s\}^\{\*\}\(r^\{\\mathrm\{c\}\}\),\\lambda^\{\*\}\(r^\{\\mathrm\{c\}\}\)\)\.\(A\.3\)
2. \(ii\)Resource feasibility and complementarity\. rc−∑i=1Nsi∗\(r\)=\(rc−r0\)\+,λ∗\(r\)\(rc−r0\)\+=0\.r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\)=\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\},\\qquad\\lambda^\{\*\}\(r\)\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}=0\.\(A\.4\)
3. \(iii\)Feasibility in the moving clipping box\. yi∗\(r\)=mi−1\(si∗\(r\)\)≤min\{X¯i,rcpi\}=Ui\(r\),i∈\[N\],0≤λ∗\(r\)≤λmax,\(\(yi∗\(r\)\)i=1N,λ∗\(r\)\)∈𝒦\(r\)\.\\begin\{gathered\}y\_\{i\}^\{\*\}\(r\)=m\_\{i\}^\{\-1\}\(s\_\{i\}^\{\*\}\(r\)\)\\leq\\min\\left\\\{\\bar\{X\}\_\{i\},\\frac\{r^\{\\mathrm\{c\}\}\}\{p\_\{i\}\}\\right\\\}=U\_\{i\}\(r\),\\qquad i\\in\[N\],\\\\ 0\\leq\\lambda^\{\*\}\(r\)\\leq\\lambda\_\{\\max\},\\qquad\\bigl\(\(y\_\{i\}^\{\*\}\(r\)\)\_\{i=1\}^\{N\},\\lambda^\{\*\}\(r\)\\bigr\)\\in\\mathcal\{K\}\(r\)\.\\end\{gathered\}\(A\.5\)
The proof of[SectionA\.1](https://arxiv.org/html/2608.14096#A1.SS1)is given in Appendix[A\.5](https://arxiv.org/html/2608.14096#A1.SS5)\.
We therefore construct a potential that measures error in sales space while following the algorithm’s threshold\-space update\. For comparator salesσ\\sigmaand current salesss, both in\[0,mi\(X¯i\)\]\[0,m\_\{i\}\(\\bar\{X\}\_\{i\}\)\], define
ℬi\(σ,s\):=∫σsu−σ\(1−Fi\(mi−1\(u\)\)\)2𝑑u\.\\mathcal\{B\}\_\{i\}\(\\sigma,s\):=\\int\_\{\\sigma\}^\{s\}\\frac\{u\-\\sigma\}\{\\left\(1\-F\_\{i\}\\bigl\(m\_\{i\}^\{\-1\}\(u\)\\bigr\)\\right\)^\{2\}\}\\,\\mathrm\{d\}u\.Sincemi′\(x\)=1−Fi\(x\)m\_\{i\}^\{\\prime\}\(x\)=1\-F\_\{i\}\(x\),
∂∂xℬi\(σ,mi\(x\)\)=mi\(x\)−σ1−Fi\(x\)\.\\frac\{\\partial\}\{\\partial x\}\\mathcal\{B\}\_\{i\}\\bigl\(\\sigma,m\_\{i\}\(x\)\\bigr\)=\\frac\{m\_\{i\}\(x\)\-\\sigma\}\{1\-F\_\{i\}\(x\)\}\.\(A\.6\)
The mean threshold step contains a factor1−Fi\(x\)1\-F\_\{i\}\(x\), so \([A\.6](https://arxiv.org/html/2608.14096#A1.E6)\) cancels that factor and leaves exactly the sales\-error–field pairing in \([4\.3](https://arxiv.org/html/2608.14096#S4.E3)\)\. This is why we useℬi\\mathcal\{B\}\_\{i\}, rather than an ordinary squared distance\. The comparator therefore remains in sales coordinates, while the update state remains in threshold coordinates\.
Because the pre\-projection update can leave\[0,X¯i\]\[0,\\bar\{X\}\_\{i\}\], we extend the potential quadratically by one unit at each endpoint\. Matching the boundary slopes below gives bounded curvature and monotonicity toward the comparator, so projection cannot increase the potential:
∂∂xℬi\(σ,mi\(x\)\)\|x=0=−σ,∂∂xℬi\(σ,mi\(x\)\)\|x=X¯i=mi\(X¯i\)−σ1−Fi\(X¯i\)\.\\left\.\\frac\{\\partial\}\{\\partial x\}\\mathcal\{B\}\_\{i\}\(\\sigma,m\_\{i\}\(x\)\)\\right\|\_\{x=0\}=\-\\sigma,\\qquad\\left\.\\frac\{\\partial\}\{\\partial x\}\\mathcal\{B\}\_\{i\}\(\\sigma,m\_\{i\}\(x\)\)\\right\|\_\{x=\\bar\{X\}\_\{i\}\}=\\frac\{m\_\{i\}\(\\bar\{X\}\_\{i\}\)\-\\sigma\}\{1\-F\_\{i\}\(\\bar\{X\}\_\{i\}\)\}\.\(A\.7\)Using these slopes and the ordinary squared discrepancy for the dual coordinate, define
𝒲\(\(𝝈,ζ\),\(𝒙,λ\)\)\\displaystyle\\mathcal\{W\}\\bigl\(\(\\bm\{\\sigma\},\\zeta\),\(\\bm\{x\},\\lambda\)\\bigr\):=∑i=1N\{ℬi\(σi,0\)−σixi\+12xi2,−1≤xi<0,ℬi\(σi,mi\(xi\)\),0≤xi≤X¯i,ℬi\(σi,mi\(X¯i\)\)\+mi\(X¯i\)−σi1−Fi\(X¯i\)\(xi−X¯i\)\+12\(xi−X¯i\)2,X¯i<xi≤X¯i\+1\\displaystyle:=\\sum\_\{i=1\}^\{N\}\\begin\{cases\}\\mathcal\{B\}\_\{i\}\(\\sigma\_\{i\},0\)\-\\sigma\_\{i\}x\_\{i\}\+\\dfrac\{1\}\{2\}x\_\{i\}^\{2\},&\-1\\leq x\_\{i\}<0,\\\\\[5\.0pt\] \\mathcal\{B\}\_\{i\}\\bigl\(\\sigma\_\{i\},m\_\{i\}\(x\_\{i\}\)\\bigr\),&0\\leq x\_\{i\}\\leq\\bar\{X\}\_\{i\},\\\\\[5\.0pt\] \\begin\{aligned\} &\\mathcal\{B\}\_\{i\}\\bigl\(\\sigma\_\{i\},m\_\{i\}\(\\bar\{X\}\_\{i\}\)\\bigr\)\+\\dfrac\{m\_\{i\}\(\\bar\{X\}\_\{i\}\)\-\\sigma\_\{i\}\}\{1\-F\_\{i\}\(\\bar\{X\}\_\{i\}\)\}\(x\_\{i\}\-\\bar\{X\}\_\{i\}\)\\\\\[\-1\.0pt\] &\\qquad\+\\dfrac\{1\}\{2\}\(x\_\{i\}\-\\bar\{X\}\_\{i\}\)^\{2\},\\end\{aligned\}&\\bar\{X\}\_\{i\}<x\_\{i\}\\leq\\bar\{X\}\_\{i\}\+1\\end\{cases\}\(A\.8\)\+12\(λ−ζ\)2\.\\displaystyle\+\\frac\{1\}\{2\}\(\\lambda\-\\zeta\)^\{2\}\.The outer derivatives are nonpositive to the left of the comparator and nonnegative to its right, so the extensions preserve projection compatibility while adding bounded curvature\. On the safe box, only the middle branch is active, and hence
𝒲\(\(𝝈,ζ\),\(𝒙,λ\)\)=∑i=1Nℬi\(σi,mi\(xi\)\)\+12\(λ−ζ\)2,0≤xi≤X¯i\.\\mathcal\{W\}\\bigl\(\(\\bm\{\\sigma\},\\zeta\),\(\\bm\{x\},\\lambda\)\\bigr\)=\\sum\_\{i=1\}^\{N\}\\mathcal\{B\}\_\{i\}\\bigl\(\\sigma\_\{i\},m\_\{i\}\(x\_\{i\}\)\\bigr\)\+\\frac\{1\}\{2\}\(\\lambda\-\\zeta\)^\{2\},\\qquad 0\\leq x\_\{i\}\\leq\\bar\{X\}\_\{i\}\.
Define the comparator and update domains, respectively, by
𝒬:=∏i=1N\[0,mi\(X¯i\)\]×\[0,λmax\],𝒵:=∏i=1N\[−1,X¯i\+1\]×\[−1,λmax\+1\]\.\\displaystyle\\mathcal\{Q\}:=\\prod\_\{i=1\}^\{N\}\[0,m\_\{i\}\(\\bar\{X\}\_\{i\}\)\]\\times\[0,\\lambda\_\{\\max\}\],~~~\\mathcal\{Z\}:=\\prod\_\{i=1\}^\{N\}\[\-1,\\bar\{X\}\_\{i\}\+1\]\\times\[\-1,\\lambda\_\{\\max\}\+1\]\.The joint domain is𝒬×𝒵\\mathcal\{Q\}\\times\\mathcal\{Z\}\. The containment below, together with the step\-size boundαtG≤1\\alpha\_\{t\}G\\leq 1verified below, ensures that every pre\-projection update remains in𝒵\\mathcal\{Z\}:
𝒦\(r\)⊆∏i=1N\[0,X¯i\]×\[0,λmax\]⊆𝒵\.\\mathcal\{K\}\(r\)\\subseteq\\prod\_\{i=1\}^\{N\}\[0,\\bar\{X\}\_\{i\}\]\\times\[0,\\lambda\_\{\\max\}\]\\subseteq\\mathcal\{Z\}\.The extra unit margin is used only to evaluate the potential at one pre\-projection update\. It does not enlarge the feasible actions onto which the algorithm projects\.
###### Lemma A\.3\.
Forq=\(𝛔,ζ\)∈𝒬q=\(\\bm\{\\sigma\},\\zeta\)\\in\\mathcal\{Q\}andz=\(𝐱,λ\)∈𝒵z=\(\\bm\{x\},\\lambda\)\\in\\mathcal\{Z\}with0≤xi≤X¯i0\\leq x\_\{i\}\\leq\\bar\{X\}\_\{i\}for everyii, there is a primitive constantCW≥1C\_\{W\}\\geq 1such that
CW−1\(‖𝒎\(𝒙\)−𝝈‖2\+\|λ−ζ\|2\)\\displaystyle C\_\{W\}^\{\-1\}\\bigl\(\\\|\\bm\{m\}\(\\bm\{x\}\)\-\\bm\{\\sigma\}\\\|^\{2\}\+\|\\lambda\-\\zeta\|^\{2\}\\bigr\)≤𝒲\(\(𝝈,ζ\),\(𝒙,λ\)\)≤CW\(‖𝒎\(𝒙\)−𝝈‖2\+\|λ−ζ\|2\)\.\\displaystyle\\leq\\mathcal\{W\}\(\(\\bm\{\\sigma\},\\zeta\),\(\\bm\{x\},\\lambda\)\)\\leq C\_\{W\}\\bigl\(\\\|\\bm\{m\}\(\\bm\{x\}\)\-\\bm\{\\sigma\}\\\|^\{2\}\+\|\\lambda\-\\zeta\|^\{2\}\\bigr\)\.\(A\.9\)‖∇q𝒲\(q,z\)‖2\+‖∇z𝒲\(q,z\)‖2≤C𝒲\(q,z\)\.\\\|\\nabla\_\{q\}\\mathcal\{W\}\(q,z\)\\\|^\{2\}\+\\\|\\nabla\_\{z\}\\mathcal\{W\}\(q,z\)\\\|^\{2\}\\leq C\\mathcal\{W\}\(q,z\)\.\(A\.10\)
The proof of[SectionA\.1](https://arxiv.org/html/2608.14096#A1.SS1)is given in Appendix[A\.5](https://arxiv.org/html/2608.14096#A1.SS5)\.
We can now define the quantity that measures the tracking error of actual interest\. Set
qt∗\\displaystyle q\_\{t\}^\{\*\}:=q∗\(rt\)=\(𝒔∗\(rt\),λ∗\(rt\)\),\\displaystyle:=q^\{\*\}\(r\_\{t\}\)=\(\\bm\{s\}^\{\*\}\(r\_\{t\}\),\\lambda^\{\*\}\(r\_\{t\}\)\),zt\\displaystyle z\_\{t\}:=\(𝑿t,λt\),\\displaystyle:=\(\\bm\{X\}\_\{t\},\\lambda\_\{t\}\),Wt\\displaystyle W\_\{t\}:=𝒲\(qt∗,zt\)=∑i=1Nℬi\(si∗\(rt\),mi\(Xi,t\)\)\+12\(λt−λ∗\(rt\)\)2\.\\displaystyle:=\\mathcal\{W\}\(q\_\{t\}^\{\*\},z\_\{t\}\)=\\sum\_\{i=1\}^\{N\}\\mathcal\{B\}\_\{i\}\\bigl\(s\_\{i\}^\{\*\}\(r\_\{t\}\),m\_\{i\}\(X\_\{i,t\}\)\\bigr\)\+\\frac\{1\}\{2\}\\bigl\(\\lambda\_\{t\}\-\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)^\{2\}\.By \([A\.5](https://arxiv.org/html/2608.14096#A1.E5)\), bothqt∗q\_\{t\}^\{\*\}and the projected stateztz\_\{t\}lie in the required domains\. The lower metric bound in \([A\.9](https://arxiv.org/html/2608.14096#A1.E9)\) therefore gives
‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+\|λt−λ∗\(rt\)\|2≤CWt\.\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\|\\lambda\_\{t\}\-\\lambda^\{\*\}\(r\_\{t\}\)\|^\{2\}\\leq CW\_\{t\}\.Taking expectations and summing yields
∑t=1T𝔼\[‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+\|λt−λ∗\(rt\)\|2\]≤C∑t=1T𝔼Wt\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\left\[\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\|\\lambda\_\{t\}\-\\lambda^\{\*\}\(r\_\{t\}\)\|^\{2\}\\right\]\\leq C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}W\_\{t\}\.\(A\.11\)
### A\.2Step 2: Construct the Smoothed Comparator
The natural exact comparator and its obstruction\.The exact targetq∗\(rt\)q^\{\*\}\(r\_\{t\}\)is feasible and ideal for the fixed\-target restoring inequality, but its path has corners when active sets change or the resource constraint becomes slack\. At a crossing no derivative gives a uniform expansion
q∗\(r\+Δr\)−q∗\(r\)=q˙∗\(r\)Δr\+O\(\|Δr\|2\)\.q^\{\*\}\(r\+\\Delta r\)\-q^\{\*\}\(r\)=\\dot\{q\}^\{\*\}\(r\)\\Delta r\+O\(\|\\Delta r\|^\{2\}\)\.The remainder is generally first order\. Lipschitz continuity alone yields
CWt\|rt\+1−rt\|≤εαtWt\+Cε\|rt\+1−rt\|2/αt,C\\sqrt\{W\_\{t\}\}\|r\_\{t\+1\}\-r\_\{t\}\|\\leq\\varepsilon\\alpha\_\{t\}W\_\{t\}\+C\_\{\\varepsilon\}\|r\_\{t\+1\}\-r\_\{t\}\|^\{2\}/\\alpha\_\{t\},whose last term isO\(αt\)O\(\\alpha\_\{t\}\), rather than theO\(αt2\)O\(\\alpha\_\{t\}^\{2\}\)forcing required for harmonic tracking\.
High\-level idea of smoothing\.We therefore averageq∗\(r−ηu\)q^\{\*\}\(r\-\\eta u\)over a backward window of widthη\\eta\. This analysis\-only smoothing staysO\(η\)O\(\\eta\)from the target, has Taylor remainderO\(\|Δr\|2/η\)O\(\|\\Delta r\|^\{2\}/\\eta\), and remains feasible because it uses no more resource thanrr\. The choiceηt=αt\\eta\_\{t\}=\\sqrt\{\\alpha\_\{t\}\}balances approximation and motion errors\.
Formally, extendq∗q^\{\*\}constantly to negativerr, and define
ρ\(u\):=6u\(1−u\)𝟏\[0,1\]\(u\),∫01ρ\(u\)𝑑u=1,\\rho\(u\):=6u\(1\-u\)\\mathbf\{1\}\_\{\[0,1\]\}\(u\),\\qquad\\int\_\{0\}^\{1\}\\rho\(u\)\\,\\mathrm\{d\}u=1,\(A\.12\)and, for0<η≤10<\\eta\\leq 1,
qη\(r\):=∫01ρ\(u\)q∗\(r−ηu\)𝑑u=:\(𝒔η\(r\),λη\(r\)\)\.q^\{\\eta\}\(r\):=\\int\_\{0\}^\{1\}\\rho\(u\)q^\{\*\}\(r\-\\eta u\)\\,\\mathrm\{d\}u=:\(\\bm\{s\}^\{\\eta\}\(r\),\\lambda^\{\\eta\}\(r\)\)\.
###### Lemma A\.4\.
There is a primitive constantCsm<∞C\_\{\\mathrm\{sm\}\}<\\inftysuch that, for everyr,r′≥0r,r^\{\\prime\}\\geq 0and0<η,η′≤10<\\eta,\\eta^\{\\prime\}\\leq 1, the following hold\.
1. \(i\)Approximation and admissibility: ‖qη\(r\)−q∗\(r\)‖≤Csmη,mi−1\(siη\(r\)\)≤yi∗\(r\)≤Ui\(r\),i∈\[N\],0≤λη\(r\)≤λmax\.\\\|q^\{\\eta\}\(r\)\-q^\{\*\}\(r\)\\\|\\leq C\_\{\\mathrm\{sm\}\}\\eta,\\qquad m\_\{i\}^\{\-1\}\(s\_\{i\}^\{\\eta\}\(r\)\)\\leq y\_\{i\}^\{\*\}\(r\)\\leq U\_\{i\}\(r\),\\quad i\\in\[N\],\\qquad 0\\leq\\lambda^\{\\eta\}\(r\)\\leq\\lambda\_\{\\max\}\.\(A\.13\)
2. \(ii\)Monotonicity and constant tail:the primal coordinates ofqηq^\{\\eta\}are nondecreasing, the dual coordinate is nonincreasing, andqη\(r\)=q∗\(r\)q^\{\\eta\}\(r\)=q^\{\*\}\(r\)wheneverr≥r0\+ηr\\geq r\_\{0\}\+\\eta\.
3. \(iii\)Motion in resource:the derivative dqη\(r\)dr=\(d𝒔η\(r\)dr,dλη\(r\)dr\)\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}=\\left\(\\frac\{\\mathrm\{d\}\\bm\{s\}^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\},\\frac\{\\mathrm\{d\}\\lambda^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\\right\)exists,‖dqη\(r\)dr‖≤Csm\\left\\\|\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\\right\\\|\\leq C\_\{\\mathrm\{sm\}\}, and ‖qη\(r′\)−qη\(r\)−dqη\(r\)dr\(r′−r\)‖≤Csm\|r′−r\|2/η\.\\left\\\|q^\{\\eta\}\(r^\{\\prime\}\)\-q^\{\\eta\}\(r\)\-\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\(r^\{\\prime\}\-r\)\\right\\\|\\leq C\_\{\\mathrm\{sm\}\}\|r^\{\\prime\}\-r\|^\{2\}/\\eta\.\(A\.14\)
4. \(iv\)Motion in radius:‖qη′\(r\)−qη\(r\)‖≤Csm\|η′−η\|\\\|q^\{\\eta^\{\\prime\}\}\(r\)\-q^\{\\eta\}\(r\)\\\|\\leq C\_\{\\mathrm\{sm\}\}\|\\eta^\{\\prime\}\-\\eta\|\.
The proof of[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)is given in Appendix[A\.5](https://arxiv.org/html/2608.14096#A1.SS5)\.
Chooseα0∈\(0,1\]\\alpha\_\{0\}\\in\(0,1\]so thatα0G≤1\\alpha\_\{0\}G\\leq 1, and set
ηt:=αt,qt:=qηt\(rt\)=:\(𝝈t,ζt\)\.\\eta\_\{t\}:=\\sqrt\{\\alpha\_\{t\}\},\\qquad q\_\{t\}:=q^\{\\eta\_\{t\}\}\(r\_\{t\}\)=:\(\\bm\{\\sigma\}\_\{t\},\\zeta\_\{t\}\)\.\(A\.15\)Underτ0≥χ/α0\\tau\_\{0\}\\geq\\chi/\\alpha\_\{0\}in \([A\.25](https://arxiv.org/html/2608.14096#A1.E25)\),
αt=χτ0\+min\{t,nt\}≤χτ0≤α0\.\\alpha\_\{t\}=\\frac\{\\chi\}\{\\tau\_\{0\}\+\\min\\\{t,n\_\{t\}\\\}\}\\leq\\frac\{\\chi\}\{\\tau\_\{0\}\}\\leq\\alpha\_\{0\}\.Thus0<ηt≤10<\\eta\_\{t\}\\leq 1andαtG≤1\\alpha\_\{t\}G\\leq 1\. The auxiliary potential in \([A\.19](https://arxiv.org/html/2608.14096#A1.E19)\) is now
W~t=𝒲\(qt,zt\)=∑i=1Nℬi\(σi,t,mi\(Xi,t\)\)\+12\(λt−ζt\)2\.\\widetilde\{W\}\_\{t\}=\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\)=\\sum\_\{i=1\}^\{N\}\\mathcal\{B\}\_\{i\}\\bigl\(\\sigma\_\{i,t\},m\_\{i\}\(X\_\{i,t\}\)\\bigr\)\+\\frac\{1\}\{2\}\(\\lambda\_\{t\}\-\\zeta\_\{t\}\)^\{2\}\.
Admissibility of the chosen comparator\.[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\(i\) putsqt,qt\+1q\_\{t\},q\_\{t\+1\}in𝒬\\mathcal\{Q\}and their threshold representations in𝒦\(rt\)\\mathcal\{K\}\(r\_\{t\}\)and𝒦\(rt\+1\)\\mathcal\{K\}\(r\_\{t\+1\}\), respectively\. Together withαtG≤1\\alpha\_\{t\}G\\leq 1, this records the admissibility of the comparator sequence in \([A\.15](https://arxiv.org/html/2608.14096#A1.E15)\)\. The remaining pre\-projection domain and projection conditions are verified below\.
### A\.3Step 3: Establish the One\-Step Drift
###### Lemma A\.5\(Projected moving\-comparator bound\)\.
Let𝒬\\mathcal\{Q\}and𝒵\\mathcal\{Z\}be convex sets, and letW:𝒬×𝒵→ℝW:\\mathcal\{Q\}\\times\\mathcal\{Z\}\\to\\mathbb\{R\}be differentiable withLWL\_\{W\}\-Lipschitz joint gradient\. Fixq,q\+∈𝒬q,q^\{\+\}\\in\\mathcal\{Q\},z∈𝒵z\\in\\mathcal\{Z\}, a directiongg, a step sizeα\>0\\alpha\>0, and a nonempty closed convex setK\+⊆𝒵K^\{\+\}\\subseteq\\mathcal\{Z\}\. Suppose thatz−αg∈𝒵z\-\\alpha g\\in\\mathcal\{Z\}and
W\(q\+,ProjK\+\(z−αg\)\)≤W\(q\+,z−αg\)\.W\\\!\\left\(q^\{\+\},\\operatorname\{Proj\}\_\{K^\{\+\}\}\(z\-\\alpha g\)\\right\)\\leq W\(q^\{\+\},z\-\\alpha g\)\.\(A\.16\)Then, withz\+:=ProjK\+\(z−αg\)z^\{\+\}:=\\operatorname\{Proj\}\_\{K^\{\+\}\}\(z\-\\alpha g\),
W\(q\+,z\+\)\\displaystyle W\(q^\{\+\},z^\{\+\}\)≤W\(q,z\)−α⟨∇zW\(q,z\),g⟩\+⟨∇qW\(q,z\),q\+−q⟩\\displaystyle\\leq W\(q,z\)\-\\alpha\\langle\\nabla\_\{z\}W\(q,z\),g\\rangle\+\\langle\\nabla\_\{q\}W\(q,z\),q^\{\+\}\-q\\rangle\+LW2α2‖g‖2\+LWα‖g‖‖q\+−q‖\+LW2‖q\+−q‖2\.\\displaystyle\\quad\+\\frac\{L\_\{W\}\}\{2\}\\alpha^\{2\}\\\|g\\\|^\{2\}\+L\_\{W\}\\alpha\\\|g\\\|\\,\\\|q^\{\+\}\-q\\\|\+\\frac\{L\_\{W\}\}\{2\}\\\|q^\{\+\}\-q\\\|^\{2\}\.\(A\.17\)
The proof is given in Appendix[A\.5](https://arxiv.org/html/2608.14096#A1.SS5)\.
Verification for the present potential\.ForW=𝒲W=\\mathcal\{W\}, the following facts verify the hypotheses of[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\. Fori∈\[N\]i\\in\[N\], set
Gi:=max\{hi,bi−ci,λmax\},G:=\(∑i=1NGi2\+\(D¯\+θGj\)2\)1/2\.G\_\{i\}:=\\max\\\{h\_\{i\},b\_\{i\}\-c\_\{i\},\\lambda\_\{\\max\}\\\},\\qquad G:=\\left\(\\sum\_\{i=1\}^\{N\}G\_\{i\}^\{2\}\+\(\\overline\{D\}\+\\theta G\_\{j\}\)^\{2\}\\right\)^\{1/2\}\.ThenG<∞G<\\infty, and there is a primitiveL𝒲<∞L\_\{\\mathcal\{W\}\}<\\infty, such that the following hold\.
1. \(i\)Regularity\.On𝒬×𝒵\\mathcal\{Q\}\\times\\mathcal\{Z\},𝒲\\mathcal\{W\}is differentiable and hasL𝒲L\_\{\\mathcal\{W\}\}\-Lipschitz joint gradient\.
2. \(ii\)Executable bounds\.The algorithm is adapted and, for everytt, zt∈𝒦\(rt\),0≤Xi,t,Ii,t,Yi,t≤X¯i<d¯i,0≤λt≤λmax,∥G^t∥≤G\.z\_\{t\}\\in\\mathcal\{K\}\(r\_\{t\}\),\\qquad 0\\leq X\_\{i,t\},I\_\{i,t\},Y\_\{i,t\}\\leq\\bar\{X\}\_\{i\}<\\bar\{d\}\_\{i\},\\qquad 0\\leq\\lambda\_\{t\}\\leq\\lambda\_\{\\max\},\\qquad\\\|\\widehat\{G\}\_\{t\}\\\|\\leq G\.
3. \(iii\)Pre\-projection update and projection\.Fixt<Tt<Tandqt,qt\+1∈𝒬q\_\{t\},q\_\{t\+1\}\\in\\mathcal\{Q\}\. IfαtG≤1\\alpha\_\{t\}G\\leq 1and the threshold representation ofqt\+1q\_\{t\+1\}belongs to𝒦\(rt\+1\)\\mathcal\{K\}\(r\_\{t\+1\}\), then zt−αtG^t∈𝒵z\_\{t\}\-\\alpha\_\{t\}\\widehat\{G\}\_\{t\}\\in\\mathcal\{Z\}and, writingzt\+1:=Proj𝒦\(rt\+1\)\(zt−αtG^t\)z\_\{t\+1\}:=\\operatorname\{Proj\}\_\{\\mathcal\{K\}\(r\_\{t\+1\}\)\}\(z\_\{t\}\-\\alpha\_\{t\}\\widehat\{G\}\_\{t\}\), 𝒲\(qt\+1,zt\+1\)≤𝒲\(qt\+1,zt−αtG^t\)\.\\mathcal\{W\}\(q\_\{t\+1\},z\_\{t\+1\}\)\\leq\\mathcal\{W\}\(q\_\{t\+1\},z\_\{t\}\-\\alpha\_\{t\}\\widehat\{G\}\_\{t\}\)\.\(A\.18\)These are the domain and projection conditions needed below\.
The three claims above are proved in Appendix[A\.5](https://arxiv.org/html/2608.14096#A1.SS5), under*Details for the one\-step verification*\.
For the smoothed comparator in \([A\.15](https://arxiv.org/html/2608.14096#A1.E15)\), the checklist and[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\(i\) permit[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)withg=G^tg=\\widehat\{G\}\_\{t\},α=αt\\alpha=\\alpha\_\{t\}, andK\+=𝒦\(rt\+1\)K^\{\+\}=\\mathcal\{K\}\(r\_\{t\+1\}\), yielding
W~t\+1\\displaystyle\\widetilde\{W\}\_\{t\+1\}≤W~t−αt⟨∇z𝒲\(qt,zt\),G^t⟩\+⟨∇q𝒲\(qt,zt\),qt\+1−qt⟩\+L𝒲G22αt2\\displaystyle\\leq\\widetilde\{W\}\_\{t\}\-\\alpha\_\{t\}\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),\\widehat\{G\}\_\{t\}\\right\\rangle\+\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),q\_\{t\+1\}\-q\_\{t\}\\right\\rangle\+\\frac\{L\_\{\\mathcal\{W\}\}G^\{2\}\}\{2\}\\alpha\_\{t\}^\{2\}\+L𝒲Gαt‖qt\+1−qt‖\+L𝒲2‖qt\+1−qt‖2\.\\displaystyle\\quad\+L\_\{\\mathcal\{W\}\}G\\alpha\_\{t\}\\\|q\_\{t\+1\}\-q\_\{t\}\\\|\+\\frac\{L\_\{\\mathcal\{W\}\}\}\{2\}\\\|q\_\{t\+1\}\-q\_\{t\}\\\|^\{2\}\.\(A\.19\)The next lemma controls the state\-update and comparator\-motion terms, respectively\.
Conditioning onℋt\\mathcal\{H\}\_\{t\}and using the marginal demand laws gives
𝔼\[G^t∣ℋt\]=Gθ\(𝒀t,λt,rt\),‖Gθ\(𝒀t,λt,rt\)−Gθ\(𝑿t,λt,rt\)‖≤C‖𝒀t−𝑿t‖\.\\mathbb\{E\}\[\\widehat\{G\}\_\{t\}\\mid\\mathcal\{H\}\_\{t\}\]=G^\{\\theta\}\(\\bm\{Y\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\),\\qquad\\\|G^\{\\theta\}\(\\bm\{Y\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\-G^\{\\theta\}\(\\bm\{X\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\\\|\\leq C\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\.\(A\.20\)Letμg:=min1≤i≤Nhiκi\\mu\_\{g\}:=\\min\_\{1\\leq i\\leq N\}h\_\{i\}\\kappa\_\{i\}andLg,j:=hjKj/βj3L\_\{g,j\}:=h\_\{j\}K\_\{j\}/\\beta\_\{j\}^\{3\}\. Chooseθ\\thetaand defineμ0\\mu\_\{0\}by
0<θ≤min\{1,μgβjLg,j2\},μ0:=12min\{μg,θβj\}\>0\.0<\\theta\\leq\\min\\left\\\{1,\\frac\{\\mu\_\{g\}\\beta\_\{j\}\}\{L\_\{g,j\}^\{2\}\}\\right\\\},\\qquad\\mu\_\{0\}:=\\frac\{1\}\{2\}\\min\\\{\\mu\_\{g\},\\theta\\beta\_\{j\}\\\}\>0\.\(A\.21\)
###### Lemma A\.6\.
LetCWC\_\{W\}be the upper metric\-comparison constant in \([A\.9](https://arxiv.org/html/2608.14096#A1.E9)\) and set the state\-update contraction modulusμupd:=μ0/\(8CW\)\\mu\_\{\\mathrm\{upd\}\}:=\\mu\_\{0\}/\(8C\_\{W\}\)\.
\(i\)State\-update contraction\.There is a primitive remainder constantC<∞C<\\inftysuch that, for everyt<Tt<T,
−αt𝔼\[⟨∇z𝒲\(qt,zt\),G^t⟩\|ℋt\]≤−μupdαtW~t\\displaystyle\-\\alpha\_\{t\}\\mathbb\{E\}\\left\[\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),\\widehat\{G\}\_\{t\}\\right\\rangle\\mathrel\{\\Big\|\}\\mathcal\{H\}\_\{t\}\\right\]\\leq\-\\mu\_\{\\mathrm\{upd\}\}\\alpha\_\{t\}\\widetilde\{W\}\_\{t\}−αt\(∑i=1Nmi\(Xi,t\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\+λt\(rtc−r0\)\+\)\+Cαt\(αt\+‖𝒀t−𝑿t‖2\)\.\\displaystyle\\quad\-\\alpha\_\{t\}\\left\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i,t\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\+\\lambda\_\{t\}\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\\right\)\+C\\alpha\_\{t\}\\bigl\(\\alpha\_\{t\}\+\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\\bigr\)\.\(A\.22\)
\(ii\)Comparator\-motion control\.There is a primitive constantCpair<∞C\_\{\\mathrm\{pair\}\}<\\infty, given explicitly in \([A\.35](https://arxiv.org/html/2608.14096#A1.E35)\)\. Set
χM:=max\{1,32Cpairμupd\},n0:=⌈max\{8,1\+32Cpairμupdα0,1\+2α0\}⌉\.\\chi\_\{M\}:=\\max\\left\\\{1,\\frac\{32C\_\{\\mathrm\{pair\}\}\}\{\\mu\_\{\\mathrm\{upd\}\}\}\\right\\\},\\qquad n\_\{0\}:=\\left\\lceil\\max\\left\\\{8,1\+\\frac\{32C\_\{\\mathrm\{pair\}\}\}\{\\mu\_\{\\mathrm\{upd\}\}\\alpha\_\{0\}\},1\+\\frac\{2\}\{\\sqrt\{\\alpha\_\{0\}\}\}\\right\\\}\\right\\rceil\.\(A\.23\)Then there is a primitive constantCm<∞C\_\{\\rm m\}<\\inftysuch that, if
χ≥χM,χ/α0≤τ0≤2χ/α0,\\chi\\geq\\chi\_\{M\},\\qquad\\chi/\\alpha\_\{0\}\\leq\\tau\_\{0\}\\leq 2\\chi/\\alpha\_\{0\},then, for everyt<Tt<Twithnt≥n0n\_\{t\}\\geq n\_\{0\},
𝔼\[⟨∇q𝒲\(qt,zt\),qt\+1−qt⟩\+L𝒲Gαt∥qt\+1−qt∥\+L𝒲2∥qt\+1−qt∥2\|ℋt\]\\displaystyle\\mathbb\{E\}\\left\[\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),q\_\{t\+1\}\-q\_\{t\}\\right\\rangle\+L\_\{\\mathcal\{W\}\}G\\alpha\_\{t\}\\\|q\_\{t\+1\}\-q\_\{t\}\\\|\+\\frac\{L\_\{\\mathcal\{W\}\}\}\{2\}\\\|q\_\{t\+1\}\-q\_\{t\}\\\|^\{2\}\\mathrel\{\\Big\|\}\\mathcal\{H\}\_\{t\}\\right\]≤μupd2αtW~t\+Cmαt2\+Cmαt‖𝒀t−𝑿t‖2\.\\displaystyle\\qquad\\leq\\frac\{\\mu\_\{\\mathrm\{upd\}\}\}\{2\}\\alpha\_\{t\}\\widetilde\{W\}\_\{t\}\+C\_\{\\rm m\}\\alpha\_\{t\}^\{2\}\+C\_\{\\rm m\}\\alpha\_\{t\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(A\.24\)
The proof of[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3), including the fixed\-target restoring calculation used in part \(i\), is given in Appendix[A\.5](https://arxiv.org/html/2608.14096#A1.SS5)\.
Setμrec:=min\{μupd/2,1/2\}\\mu\_\{\\mathrm\{rec\}\}:=\\min\\\{\\mu\_\{\\mathrm\{upd\}\}/2,1/2\\\}, and decreaseα0\\alpha\_\{0\}, if necessary, so that
α0G≤1,μrecα0≤12\.\\alpha\_\{0\}G\\leq 1,\\qquad\\mu\_\{\\mathrm\{rec\}\}\\alpha\_\{0\}\\leq\\frac\{1\}\{2\}\.With thisα0\\alpha\_\{0\}, takeχM,n0\\chi\_\{M\},n\_\{0\}from[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\(ii\)\. For the rest of this step, impose
χ≥χM,χμrec\>3,χ/α0≤τ0≤2χ/α0,\\chi\\geq\\chi\_\{M\},\\qquad\\chi\\mu\_\{\\mathrm\{rec\}\}\>3,\\qquad\\chi/\\alpha\_\{0\}\\leq\\tau\_\{0\}\\leq 2\\chi/\\alpha\_\{0\},\(A\.25\)and fixt<Tt<Twithnt≥n0n\_\{t\}\\geq n\_\{0\}\. The storewise KKT inequalities givegi′\(si∗\(rt\)\)\+λ∗\(rt\)≥0g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\geq 0, whilemi\(Xi,t\)≥0m\_\{i\}\(X\_\{i,t\}\)\\geq 0,λt≥0\\lambda\_\{t\}\\geq 0, and\(rtc−r0\)\+≥0\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\\geq 0\. Thus the residual below is nonnegative\. Combining[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\(i\)–\(ii\) in \([A\.19](https://arxiv.org/html/2608.14096#A1.E19)\), and usingμrec≤μupd/2\\mu\_\{\\mathrm\{rec\}\}\\leq\\mu\_\{\\mathrm\{upd\}\}/2andμrec≤1\\mu\_\{\\mathrm\{rec\}\}\\leq 1, gives
𝔼\[W~t\+1∣ℋt\]\\displaystyle\\mathbb\{E\}\[\\widetilde\{W\}\_\{t\+1\}\\mid\\mathcal\{H\}\_\{t\}\]≤\(1−μrecαt\)W~t−μrecαt\(∑i=1Nmi\(Xi,t\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\+λt\(rtc−r0\)\+\)\\displaystyle\\leq\(1\-\\mu\_\{\\mathrm\{rec\}\}\\alpha\_\{t\}\)\\widetilde\{W\}\_\{t\}\-\\mu\_\{\\mathrm\{rec\}\}\\alpha\_\{t\}\\left\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i,t\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\+\\lambda\_\{t\}\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\\right\)\+Cαt2\+Cαt‖𝒀t−𝑿t‖2\.\\displaystyle\\quad\+C\\alpha\_\{t\}^\{2\}\+C\\alpha\_\{t\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(A\.26\)The conditionχμrec\>3\\chi\\mu\_\{\\mathrm\{rec\}\}\>3, in particular, guaranteesμrec\>1/χ\\mu\_\{\\mathrm\{rec\}\}\>1/\\chi, as needed for the reciprocal\-step absorption below\.
This completes the one\-step drift argument\.
### A\.4Step 4: Sum the Recursion
With the one\-step recursion \([A\.26](https://arxiv.org/html/2608.14096#A1.E26)\) established in Appendix[A\.3](https://arxiv.org/html/2608.14096#A1.SS3), it remains to sum it in two ways and then return from the smoothed comparator to the exact target\.
Summation\.Dropping the nonnegative parenthesized residual in \([A\.26](https://arxiv.org/html/2608.14096#A1.E26)\) and taking expectations gives
𝔼W~t\+1≤\(1−μrecαt\)𝔼W~t\+Cαt2\+Cαt𝔼‖𝒀t−𝑿t‖2\.\\mathbb\{E\}\\widetilde\{W\}\_\{t\+1\}\\leq\(1\-\\mu\_\{\\mathrm\{rec\}\}\\alpha\_\{t\}\)\\mathbb\{E\}\\widetilde\{W\}\_\{t\}\+C\\alpha\_\{t\}^\{2\}\+C\\alpha\_\{t\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(A\.27\)
The first required bound is
∑t=1T𝔼W~t≤ClogT\+C∑t=1T𝔼‖𝒀t−𝑿t‖2\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\widetilde\{W\}\_\{t\}\\leq C\\log T\+C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(A\.28\)The second bound compares the exact and smoothed potentials:
Wt≤CW~t\+Cαt\.W\_\{t\}\\leq C\\widetilde\{W\}\_\{t\}\+C\\alpha\_\{t\}\.\(A\.29\)Consequently,
∑t=1T𝔼Wt≤C∑t=1T𝔼W~t\+ClogT\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}W\_\{t\}\\leq C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\widetilde\{W\}\_\{t\}\+C\\log T\.\(A\.30\)To prove \([A\.28](https://arxiv.org/html/2608.14096#A1.E28)\), first supposeT≥n0T\\geq n\_\{0\}, setL=T−n0\+1L=T\-n\_\{0\}\+1, and write
ut:=𝔼W~t,bt:=C𝔼‖𝒀t−𝑿t‖2,ht:=1αt\.u\_\{t\}:=\\mathbb\{E\}\\widetilde\{W\}\_\{t\},\\qquad b\_\{t\}:=C\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\},\\qquad h\_\{t\}:=\\frac\{1\}\{\\alpha\_\{t\}\}\.Compactness gives0≤ut≤U0\\leq u\_\{t\}\\leq U\. Fort≤Lt\\leq L, divide \([A\.27](https://arxiv.org/html/2608.14096#A1.E27)\) byαt\\alpha\_\{t\}and sum to obtain
μrec∑t=1Lut≤∑t=1Lht\(ut−ut\+1\)\+C∑t=1Lαt\+∑t=1Lbt\.\\mu\_\{\\mathrm\{rec\}\}\\sum\_\{t=1\}^\{L\}u\_\{t\}\\leq\\sum\_\{t=1\}^\{L\}h\_\{t\}\(u\_\{t\}\-u\_\{t\+1\}\)\+C\\sum\_\{t=1\}^\{L\}\\alpha\_\{t\}\+\\sum\_\{t=1\}^\{L\}b\_\{t\}\.Writingdt=min\{t,T−t\+1\}d\_\{t\}=\\min\\\{t,T\-t\+1\\\}, summation by parts andht=\(τ0\+dt\)/χh\_\{t\}=\(\\tau\_\{0\}\+d\_\{t\}\)/\\chigive
∑t=1Lht\(ut−ut\+1\)=h1u1\+∑t=2L\(ht−ht−1\)ut−hLuL\+1≤C\+1χ∑t=2Lut\.\\sum\_\{t=1\}^\{L\}h\_\{t\}\(u\_\{t\}\-u\_\{t\+1\}\)=h\_\{1\}u\_\{1\}\+\\sum\_\{t=2\}^\{L\}\(h\_\{t\}\-h\_\{t\-1\}\)u\_\{t\}\-h\_\{L\}u\_\{L\+1\}\\leq C\+\\frac\{1\}\{\\chi\}\\sum\_\{t=2\}^\{L\}u\_\{t\}\.Becauseμrec\>1/χ\\mu\_\{\\mathrm\{rec\}\}\>1/\\chi, this sum is absorbed\. Moreover,
∑t=1Tαt≤2χ∑k=1⌈T/2⌉\(τ0\+k\)−1≤ClogT,\\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}\\leq 2\\chi\\sum\_\{k=1\}^\{\\lceil T/2\\rceil\}\(\\tau\_\{0\}\+k\)^\{\-1\}\\leq C\\log T,and the finaln0−1n\_\{0\}\-1potentials contribute at most\(n0−1\)U\(n\_\{0\}\-1\)U\. This proves \([A\.28](https://arxiv.org/html/2608.14096#A1.E28)\)\. ForT<n0T<n\_\{0\}, boundedness gives the same conclusion directly\.
Proof of the exact\-to\-smoothed comparison\.The upper metric comparison in \([A\.9](https://arxiv.org/html/2608.14096#A1.E9)\) gives
Wt≤CW\(‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+\|λt−λ∗\(rt\)\|2\)\.W\_\{t\}\\leq C\_\{W\}\\left\(\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\|\\lambda\_\{t\}\-\\lambda^\{\*\}\(r\_\{t\}\)\|^\{2\}\\right\)\.Insertqt=\(𝝈t,ζt\)q\_\{t\}=\(\\bm\{\\sigma\}\_\{t\},\\zeta\_\{t\}\)separately in the primal and dual coordinates and use‖a\+b‖2≤2‖a‖2\+2‖b‖2\\\|a\+b\\\|^\{2\}\\leq 2\\\|a\\\|^\{2\}\+2\\\|b\\\|^\{2\}\. The expression in parentheses is at most
2\(‖𝒎\(𝑿t\)−𝝈t‖2\+\|λt−ζt\|2\)\+2‖qt−q∗\(rt\)‖2\.2\\left\(\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{\\sigma\}\_\{t\}\\\|^\{2\}\+\|\\lambda\_\{t\}\-\\zeta\_\{t\}\|^\{2\}\\right\)\+2\\\|q\_\{t\}\-q^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\.The lower metric comparison in \([A\.9](https://arxiv.org/html/2608.14096#A1.E9)\) gives
‖𝒎\(𝑿t\)−𝝈t‖2\+\|λt−ζt\|2≤CW~t,\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{\\sigma\}\_\{t\}\\\|^\{2\}\+\|\\lambda\_\{t\}\-\\zeta\_\{t\}\|^\{2\}\\leq C\\widetilde\{W\}\_\{t\},while[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\(i\) andηt2=αt\\eta\_\{t\}^\{2\}=\\alpha\_\{t\}give
‖qt−q∗\(rt\)‖2≤Csm2ηt2≤Cαt\.\\\|q\_\{t\}\-q^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\\leq C\_\{\\mathrm\{sm\}\}^\{2\}\\eta\_\{t\}^\{2\}\\leq C\\alpha\_\{t\}\.Substitution proves \([A\.29](https://arxiv.org/html/2608.14096#A1.E29)\)\. Taking expectations and summing this pointwise bound yields
∑t=1T𝔼Wt≤C∑t=1T𝔼W~t\+C∑t=1Tαt\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}W\_\{t\}\\leq C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\widetilde\{W\}\_\{t\}\+C\\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}\.The harmonic estimate above now proves \([A\.30](https://arxiv.org/html/2608.14096#A1.E30)\)\.
Combining \([A\.11](https://arxiv.org/html/2608.14096#A1.E11)\), \([A\.30](https://arxiv.org/html/2608.14096#A1.E30)\), and \([A\.28](https://arxiv.org/html/2608.14096#A1.E28)\) gives
∑t=1T𝔼\[‖𝒎\(𝑿t\)−𝒔∗\(rt\)‖2\+\|λt−λ∗\(rt\)\|2\]≤C∑t=1T𝔼Wt\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\left\[\\\|\\bm\{m\}\(\\bm\{X\}\_\{t\}\)\-\\bm\{s\}^\{\*\}\(r\_\{t\}\)\\\|^\{2\}\+\|\\lambda\_\{t\}\-\\lambda^\{\*\}\(r\_\{t\}\)\|^\{2\}\\right\]\\leq C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}W\_\{t\}≤C∑t=1T𝔼W~t\+ClogT≤ClogT\+C∑t=1T𝔼‖𝒀t−𝑿t‖2,\\displaystyle\\leq C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\widetilde\{W\}\_\{t\}\+C\\log T\\leq C\\log T\+C\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\},which is \([A\.1](https://arxiv.org/html/2608.14096#A1.E1)\)\.
Residual summation\.It remains to prove[AppendixA](https://arxiv.org/html/2608.14096#A1)\(ii\) from the same recursion\. ForT<n0T<n\_\{0\}, each storewise residual is bounded byλmaxmi\(X¯i\)\\lambda\_\{\\max\}m\_\{i\}\(\\bar\{X\}\_\{i\}\), so the claim follows after enlargingCmdC\_\{\\rm md\}\. Suppose therefore thatT≥n0T\\geq n\_\{0\}, and letL=T−n0\+1L=T\-n\_\{0\}\+1\. For the residual, writewt=𝔼W~tw\_\{t\}=\\mathbb\{E\}\\widetilde\{W\}\_\{t\}, rearrange \([A\.26](https://arxiv.org/html/2608.14096#A1.E26)\), take expectations, and divide byμrecαt\\mu\_\{\\mathrm\{rec\}\}\\alpha\_\{t\}\. Discarding only the nonpositive term−wt\-w\_\{t\}gives
𝔼\[∑i=1Nmi\(Xi,t\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\+λt\(rtc−r0\)\+\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i,t\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\+\\lambda\_\{t\}\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\\right\]≤wt−wt\+1μrecαt\+Cαt\+C𝔼‖𝒀t−𝑿t‖2\.\\displaystyle\\qquad\\leq\\frac\{w\_\{t\}\-w\_\{t\+1\}\}\{\\mu\_\{\\mathrm\{rec\}\}\\alpha\_\{t\}\}\+C\\alpha\_\{t\}\+C\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.\(A\.31\)The intermediate inequality contains an additional resource term\. By[SectionA\.1](https://arxiv.org/html/2608.14096#A1.SS1),
rtc−∑i=1Nsi∗\(rt\)=\(rtc−r0\)\+,λ∗\(rt\)\(rtc−r0\)\+=0\.r\_\{t\}^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\_\{t\}\)=\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\},\\qquad\\lambda^\{\*\}\(r\_\{t\}\)\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}=0\.Consequently, the target resource residual paired with the current dual error satisfies
\(λt−λ∗\(rt\)\)\(rtc−∑i=1Nsi∗\(rt\)\)=\(λt−λ∗\(rt\)\)\(rtc−r0\)\+=λt\(rtc−r0\)\+≥0\.\(\\lambda\_\{t\}\-\\lambda^\{\*\}\(r\_\{t\}\)\)\\left\(r\_\{t\}^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\_\{t\}\)\\right\)=\(\\lambda\_\{t\}\-\\lambda^\{\*\}\(r\_\{t\}\)\)\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}=\\lambda\_\{t\}\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\\geq 0\.Thus this resource\-dissipation term may be dropped from the left\-hand side of \([A\.31](https://arxiv.org/html/2608.14096#A1.E31)\) when proving[AppendixA](https://arxiv.org/html/2608.14096#A1)\(ii\)\.
The reciprocal\-step telescope is the exact identity
∑t=1Lwt−wt\+1αt\\displaystyle\\sum\_\{t=1\}^\{L\}\\frac\{w\_\{t\}\-w\_\{t\+1\}\}\{\\alpha\_\{t\}\}=w1α1\+∑t=2Lwt\(1αt−1αt−1\)−wL\+1αL\.\\displaystyle=\\frac\{w\_\{1\}\}\{\\alpha\_\{1\}\}\+\\sum\_\{t=2\}^\{L\}w\_\{t\}\\left\(\\frac\{1\}\{\\alpha\_\{t\}\}\-\\frac\{1\}\{\\alpha\_\{t\-1\}\}\\right\)\-\\frac\{w\_\{L\+1\}\}\{\\alpha\_\{L\}\}\.For this proof, writedt:=min\{t,nt\}d\_\{t\}:=\\min\\\{t,n\_\{t\}\\\}\. Because1/αt=\(τ0\+dt\)/χ1/\\alpha\_\{t\}=\(\\tau\_\{0\}\+d\_\{t\}\)/\\chi, the positive part of the reciprocal increment is exactly1/χ1/\\chiwhiledtd\_\{t\}increases, zero on a possible midpoint plateau, and nonpositive while it decreases\. The first term is uniformly bounded and the last is nonpositive\. Therefore the telescope is at mostC\+\(1/χ\)∑t=2LwtC\+\(1/\\chi\)\\sum\_\{t=2\}^\{L\}w\_\{t\}, which is controlled by \([A\.28](https://arxiv.org/html/2608.14096#A1.E28)\)\. After dropping the nonnegative resource term as above, summing \([A\.31](https://arxiv.org/html/2608.14096#A1.E31)\) and adding the finaln0−1n\_\{0\}\-1bounded storewise residuals proves \([A\.2](https://arxiv.org/html/2608.14096#A1.E2)\)\.
### A\.5Auxiliary Proofs and Verifications
This subsection collects, in order of use, the auxiliary proofs for Steps 1–3 and the remaining one\-step verifications\.
Proof of[SectionA\.1](https://arxiv.org/html/2608.14096#A1.SS1)\.
###### Proof\.
Proof\.*Target preservation and complementarity\.*Sincer0=∑iϕi\(0\)≤D¯r\_\{0\}=\\sum\_\{i\}\\phi\_\{i\}\(0\)\\leq\\overline\{D\}, ifr≤D¯r\\leq\\overline\{D\}, capping does nothing\. Ifr\>D¯r\>\\overline\{D\}, then bothrrandrc=D¯r^\{\\mathrm\{c\}\}=\\overline\{D\}lie on the constant nonbinding branch\(ϕ\(0\),0\)\(\\bm\{\\phi\}\(0\),0\)of \([B\.3](https://arxiv.org/html/2608.14096#A2.E3)\)\. This proves \([A\.3](https://arxiv.org/html/2608.14096#A1.E3)\)\. Moreover, the binding branch gives∑isi∗\(r\)=r=rc\\sum\_\{i\}s\_\{i\}^\{\*\}\(r\)=r=r^\{\\mathrm\{c\}\}forr<r0r<r\_\{0\}, while the nonbinding branch gives∑isi∗\(r\)=r0\\sum\_\{i\}s\_\{i\}^\{\*\}\(r\)=r\_\{0\}andλ∗\(r\)=0\\lambda^\{\*\}\(r\)=0forr≥r0r\\geq r\_\{0\}\. Hence
rc−∑isi∗\(r\)=\(rc−r0\)\+,λ∗\(r\)\(rc−r0\)\+=0\.r^\{\\mathrm\{c\}\}\-\\sum\_\{i\}s\_\{i\}^\{\*\}\(r\)=\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\},\\qquad\\lambda^\{\*\}\(r\)\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}=0\.
*Feasibility in the moving clipping box\.*It remains to verify feasibility in the moving clipping box\. A nonnegative resource price can only lower the store threshold below the zero\-price newsvendor threshold, and henceyi∗\(r\)<X¯iy\_\{i\}^\{\*\}\(r\)<\\bar\{X\}\_\{i\}\. If storeiiis active, then recall thatpip\_\{i\}is its zero\-price survival probability\. Therefore1−Fi\(yi∗\(r\)\)≥pi1\-F\_\{i\}\(y\_\{i\}^\{\*\}\(r\)\)\\geq p\_\{i\}, and
si∗\(r\)=∫0yi∗\(r\)\(1−Fi\(u\)\)𝑑u≥piyi∗\(r\)\.s\_\{i\}^\{\*\}\(r\)=\\int\_\{0\}^\{y\_\{i\}^\{\*\}\(r\)\}\\bigl\(1\-F\_\{i\}\(u\)\\bigr\)\\,\\mathrm\{d\}u\\geq p\_\{i\}y\_\{i\}^\{\*\}\(r\)\.The capped complementarity identity already proved gives
si∗\(r\)≤∑k=1Nsk∗\(r\)=rc−\(rc−r0\)\+≤rc\.s\_\{i\}^\{\*\}\(r\)\\leq\\sum\_\{k=1\}^\{N\}s\_\{k\}^\{\*\}\(r\)=r^\{\\mathrm\{c\}\}\-\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\\leq r^\{\\mathrm\{c\}\}\.This proves the primal bound in \([A\.5](https://arxiv.org/html/2608.14096#A1.E5)\)\. For an inactive store,yi∗\(r\)=0y\_\{i\}^\{\*\}\(r\)=0, so the same conclusion holds\. Finally,0≤λ∗\(r\)≤λmax0\\leq\\lambda^\{\*\}\(r\)\\leq\\lambda\_\{\\max\}is part of the selected KKT path\. Together with the coordinatewise primal bounds and the definition of𝒦\(r\)\\mathcal\{K\}\(r\), it proves the asserted threshold–dual membership\. ∎
Proof of[SectionA\.1](https://arxiv.org/html/2608.14096#A1.SS1)\.
###### Proof\.
Proof\. For a fixed storeii, abbreviate
Si\(x\):=1−Fi\(x\),Mi:=mi\(X¯i\),wi\(u\):=Si\(mi−1\(u\)\)−2\.S\_\{i\}\(x\):=1\-F\_\{i\}\(x\),\\qquad M\_\{i\}:=m\_\{i\}\(\\bar\{X\}\_\{i\}\),\\qquad w\_\{i\}\(u\):=S\_\{i\}\\bigl\(m\_\{i\}^\{\-1\}\(u\)\\bigr\)^\{\-2\}\.The safe\-box bounds imply, for0≤u≤Mi0\\leq u\\leq M\_\{i\},
1≤wi\(u\)≤βi−2\.1\\leq w\_\{i\}\(u\)\\leq\\beta\_\{i\}^\{\-2\}\.
*Metric comparison\.*Fors∈\[0,Mi\]s\\in\[0,M\_\{i\}\], the change of variablesu=σ\+τ\(s−σ\)u=\\sigma\+\\tau\(s\-\\sigma\)gives
ℬi\(σ,s\)=\(s−σ\)2∫01τwi\(σ\+τ\(s−σ\)\)𝑑τ\.\\mathcal\{B\}\_\{i\}\(\\sigma,s\)=\(s\-\\sigma\)^\{2\}\\int\_\{0\}^\{1\}\\tau w\_\{i\}\\bigl\(\\sigma\+\\tau\(s\-\\sigma\)\\bigr\)\\,\\mathrm\{d\}\\tau\.Consequently,
12\(s−σ\)2≤ℬi\(σ,s\)≤12βi2\(s−σ\)2\.\\frac\{1\}\{2\}\(s\-\\sigma\)^\{2\}\\leq\\mathcal\{B\}\_\{i\}\(\\sigma,s\)\\leq\\frac\{1\}\{2\\beta\_\{i\}^\{2\}\}\(s\-\\sigma\)^\{2\}\.Settings=mi\(xi\)s=m\_\{i\}\(x\_\{i\}\), summing overii, and adding the dual quadratic proves \([A\.9](https://arxiv.org/html/2608.14096#A1.E9)\)\.
*Gradient control\.*LetHi\(σ,x\)H\_\{i\}\(\\sigma,x\)denote theii\-th primal summand in \([A\.8](https://arxiv.org/html/2608.14096#A1.E8)\)\. On the safe box, direct differentiation of its middle branch gives
\|∂xHi\(σi,xi\)\|≤βi−1\|mi\(xi\)−σi\|,\|∂σHi\(σi,xi\)\|≤βi−2\|mi\(xi\)−σi\|\.\|\\partial\_\{x\}H\_\{i\}\(\\sigma\_\{i\},x\_\{i\}\)\|\\leq\\beta\_\{i\}^\{\-1\}\|m\_\{i\}\(x\_\{i\}\)\-\\sigma\_\{i\}\|,\\qquad\|\\partial\_\{\\sigma\}H\_\{i\}\(\\sigma\_\{i\},x\_\{i\}\)\|\\leq\\beta\_\{i\}^\{\-2\}\|m\_\{i\}\(x\_\{i\}\)\-\\sigma\_\{i\}\|\.The two derivatives of the dual quadratic have magnitude\|λ−ζ\|\|\\lambda\-\\zeta\|\. Therefore,
‖∇q𝒲\(q,z\)‖2\+‖∇z𝒲\(q,z\)‖2≤C\(‖𝒎\(𝒙\)−𝝈‖2\+\|λ−ζ\|2\)\.\\\|\\nabla\_\{q\}\\mathcal\{W\}\(q,z\)\\\|^\{2\}\+\\\|\\nabla\_\{z\}\\mathcal\{W\}\(q,z\)\\\|^\{2\}\\leq C\\bigl\(\\\|\\bm\{m\}\(\\bm\{x\}\)\-\\bm\{\\sigma\}\\\|^\{2\}\+\|\\lambda\-\\zeta\|^\{2\}\\bigr\)\.The lower bound in \([A\.9](https://arxiv.org/html/2608.14096#A1.E9)\) converts the right\-hand side toC𝒲\(q,z\)C\\mathcal\{W\}\(q,z\), proving \([A\.10](https://arxiv.org/html/2608.14096#A1.E10)\)\. ∎
Proof of[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\.
###### Proof\.
Proof\. LetLqL\_\{q\}be a global Lipschitz constant ofq∗q^\{\*\}\. Coupling the integrands at the same kernel point gives
‖qη\(r\)−q∗\(r\)‖\\displaystyle\\\|q^\{\\eta\}\(r\)\-q^\{\*\}\(r\)\\\|≤Lqη∫01uρ\(u\)𝑑u≤Lqη,\\displaystyle\\leq L\_\{q\}\\eta\\int\_\{0\}^\{1\}u\\rho\(u\)\\,\\mathrm\{d\}u\\leq L\_\{q\}\\eta,‖qη′\(r\)−qη\(r\)‖\\displaystyle\\\|q^\{\\eta^\{\\prime\}\}\(r\)\-q^\{\\eta\}\(r\)\\\|≤Lq\|η′−η\|∫01uρ\(u\)𝑑u≤Lq\|η′−η\|\.\\displaystyle\\leq L\_\{q\}\|\\eta^\{\\prime\}\-\\eta\|\\int\_\{0\}^\{1\}u\\rho\(u\)\\,\\mathrm\{d\}u\\leq L\_\{q\}\|\\eta^\{\\prime\}\-\\eta\|\.Averaging preserves the coordinatewise monotonicity of the KKT path\. Since the kernel uses onlyr−ηu≤rr\-\\eta u\\leq r, primal monotonicity also gives
mi−1\(siη\(r\)\)≤mi−1\(si∗\(r\)\)=yi∗\(r\)≤Ui\(r\)\.m\_\{i\}^\{\-1\}\(s\_\{i\}^\{\\eta\}\(r\)\)\\leq m\_\{i\}^\{\-1\}\(s\_\{i\}^\{\*\}\(r\)\)=y\_\{i\}^\{\*\}\(r\)\\leq U\_\{i\}\(r\)\.The multiplier average remains in\[0,λmax\]\[0,\\lambda\_\{\\max\}\], and every integrand equals the constant nonbinding target whenr≥r0\+ηr\\geq r\_\{0\}\+\\eta\. This proves parts \(i\), \(ii\), and \(iv\)\.
Extendρ\\rhoby zero outside\[0,1\]\[0,1\]\. A Lipschitz path is absolutely continuous and has an a\.e\. derivativeq˙∗\\dot\{q\}^\{\*\}with‖q˙∗‖∞≤Lq\\\|\\dot\{q\}^\{\*\}\\\|\_\{\\infty\}\\leq L\_\{q\}\. With
Kη\(u\):=η−1ρ\(u/η\)𝟏\[0,η\]\(u\),K\_\{\\eta\}\(u\):=\\eta^\{\-1\}\\rho\(u/\\eta\)\\mathbf\{1\}\_\{\[0,\\eta\]\}\(u\),the smoothing isKη∗q∗K\_\{\\eta\}\*q^\{\*\}, and its weak derivative is
dqη\(r\)dr=\(Kη∗q˙∗\)\(r\)\.\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}=\(K\_\{\\eta\}\*\\dot\{q\}^\{\*\}\)\(r\)\.Translation continuity inL1L^\{1\}makes this derivative continuous, so it is the classical derivative everywhere\. The zero extension ofρ\\rhohas total variation
Vρ:=\|ρ\(0\)\|\+\|ρ\(1\)\|\+∫01\|ρ′\(u\)\|𝑑u\.V\_\{\\rho\}:=\|\\rho\(0\)\|\+\|\\rho\(1\)\|\+\\int\_\{0\}^\{1\}\|\\rho^\{\\prime\}\(u\)\|\\,\\mathrm\{d\}u\.Consequently
‖dqη\(r\)dr‖≤Lq,‖dqη\(r′\)dr′−dqη\(r\)dr‖≤\(VρLq/η\)\|r′−r\|,\\left\\\|\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\\right\\\|\\leq L\_\{q\},\\qquad\\left\\\|\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r^\{\\prime\}\)\}\{\\mathrm\{d\}r^\{\\prime\}\}\-\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\\right\\\|\\leq\(V\_\{\\rho\}L\_\{q\}/\\eta\)\|r^\{\\prime\}\-r\|,because∥Kη\(⋅−h\)−Kη∥1≤\(Vρ/η\)\|h\|\\\|K\_\{\\eta\}\(\\cdot\-h\)\-K\_\{\\eta\}\\\|\_\{1\}\\leq\(V\_\{\\rho\}/\\eta\)\|h\|\. The integral form of Taylor’s theorem gives \([A\.14](https://arxiv.org/html/2608.14096#A1.E14)\)\. ChoosingCsmC\_\{\\mathrm\{sm\}\}to dominate the displayed constants proves part \(iii\) and completes the proof\. ∎
Proof of[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\.
###### Proof\.
Proof\. By \([A\.16](https://arxiv.org/html/2608.14096#A1.E16)\),
W\(q\+,z\+\)≤W\(q\+,z−αg\)\.W\(q^\{\+\},z^\{\+\}\)\\leq W\(q^\{\+\},z\-\\alpha g\)\.Applying the descent lemma first in the comparator argument gives
W\(q\+,z−αg\)\\displaystyle W\(q^\{\+\},z\-\\alpha g\)≤W\(q,z−αg\)\+⟨∇qW\(q,z−αg\),q\+−q⟩\+LW2‖q\+−q‖2\.\\displaystyle\\leq W\(q,z\-\\alpha g\)\+\\left\\langle\\nabla\_\{q\}W\(q,z\-\\alpha g\),q^\{\+\}\-q\\right\\rangle\+\\frac\{L\_\{W\}\}\{2\}\\\|q^\{\+\}\-q\\\|^\{2\}\.Moreover,
‖∇qW\(q,z−αg\)−∇qW\(q,z\)‖≤LWα‖g‖,\\\|\\nabla\_\{q\}W\(q,z\-\\alpha g\)\-\\nabla\_\{q\}W\(q,z\)\\\|\\leq L\_\{W\}\\alpha\\\|g\\\|,while the descent lemma in the state argument gives
W\(q,z−αg\)≤W\(q,z\)−α⟨∇zW\(q,z\),g⟩\+LW2α2‖g‖2\.W\(q,z\-\\alpha g\)\\leq W\(q,z\)\-\\alpha\\langle\\nabla\_\{z\}W\(q,z\),g\\rangle\+\\frac\{L\_\{W\}\}\{2\}\\alpha^\{2\}\\\|g\\\|^\{2\}\.Substitution and Cauchy–Schwarz prove \([A\.17](https://arxiv.org/html/2608.14096#A1.E17)\)\. ∎
Proof of[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\.
###### Proof\.
Proof\.
Part \(i\): State\-update contraction\.
*Substep 1: fixed\-target restoring inequality\.*For everyr≥0r\\geq 0,0≤λ≤λmax0\\leq\\lambda\\leq\\lambda\_\{\\max\}, and0≤xi≤X¯i0\\leq x\_\{i\}\\leq\\bar\{X\}\_\{i\}, the fixed\-target calculation below gives
⟨∇z𝒲\(q∗\(r\),\(𝒙,λ\)\),Gθ\(𝒙,λ,r\)⟩≥μ0\(‖𝒎\(𝒙\)−𝒔∗\(r\)‖2\+\|λ−λ∗\(r\)\|2\)\\displaystyle\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\\bigl\(q^\{\*\}\(r\),\(\\bm\{x\},\\lambda\)\\bigr\),G^\{\\theta\}\(\\bm\{x\},\\lambda;r\)\\right\\rangle\\geq\\mu\_\{0\}\\left\(\\\|\\bm\{m\}\(\\bm\{x\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|^\{2\}\+\|\\lambda\-\\lambda^\{\*\}\(r\)\|^\{2\}\\right\)\+∑i=1Nmi\(xi\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\+λ\(rc−r0\)\+\.\\displaystyle\\quad\+\\sum\_\{i=1\}^\{N\}m\_\{i\}\(x\_\{i\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\+\\lambda\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\.\(A\.32\)
*Substep 2: transfer to the smoothed comparator and identify the ideal pairing\.*Fixr≥0r\\geq 0,0<η≤10<\\eta\\leq 1,0≤λ≤λmax0\\leq\\lambda\\leq\\lambda\_\{\\max\}, and0≤Xi≤X¯i0\\leq X\_\{i\}\\leq\\bar\{X\}\_\{i\}\. Subtracting the exact\-target pairing in \([A\.32](https://arxiv.org/html/2608.14096#A1.E32)\) from the corresponding pairing withqη\(r\)q^\{\\eta\}\(r\)gives
−∑i=1N\(siη\(r\)−si∗\(r\)\)\(gi′\(mi\(Xi\)\)\+λ\)\\displaystyle\-\\sum\_\{i=1\}^\{N\}\\bigl\(s\_\{i\}^\{\\eta\}\(r\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(X\_\{i\}\)\)\+\\lambda\\bigr\)−\(λη\(r\)−λ∗\(r\)\)\(rc−∑i=1Nmi\(Xi\)\+θ\(1−Fj\(Xj\)\)\(gj′\(mj\(Xj\)\)\+λ\)\)\.\\displaystyle\\quad\-\\bigl\(\\lambda^\{\\eta\}\(r\)\-\\lambda^\{\*\}\(r\)\\bigr\)\\left\(r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i\}\)\+\\theta\\bigl\(1\-F\_\{j\}\(X\_\{j\}\)\\bigr\)\\bigl\(g\_\{j\}^\{\\prime\}\(m\_\{j\}\(X\_\{j\}\)\)\+\\lambda\\bigr\)\\right\)\.If storeiiis inactive atrr, backward smoothing leavessiη\(r\)=si∗\(r\)=0s\_\{i\}^\{\\eta\}\(r\)=s\_\{i\}^\{\*\}\(r\)=0\. If it is active, its KKT residual vanishes, and safe\-box regularity gives
\|gi′\(mi\(Xi\)\)\+λ\|≤C\(‖𝒎\(𝑿\)−𝒔∗\(r\)‖\+\|λ−λ∗\(r\)\|\)\.\|g\_\{i\}^\{\\prime\}\(m\_\{i\}\(X\_\{i\}\)\)\+\\lambda\|\\leq C\\bigl\(\\\|\\bm\{m\}\(\\bm\{X\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|\+\|\\lambda\-\\lambda^\{\*\}\(r\)\|\\bigr\)\.For the dual line, use
rc−∑i=1Nmi\(Xi\)=\(rc−r0\)\+−∑i=1N\(mi\(Xi\)−si∗\(r\)\),r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i\}\)=\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\-\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(X\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\),and \([A\.38](https://arxiv.org/html/2608.14096#A1.E38)\)\. The product ofλη\(r\)−λ∗\(r\)\\lambda^\{\\eta\}\(r\)\-\\lambda^\{\*\}\(r\)and\(rc−r0\)\+\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}vanishes outside\[r0,r0\+η\]\[r\_\{0\},r\_\{0\}\+\\eta\]\. Inside that layer its two factors areO\(η\)O\(\\eta\)\. Together with \([A\.13](https://arxiv.org/html/2608.14096#A1.E13)\), the displayed difference is bounded below by
−Cη\(‖𝒎\(𝑿\)−𝒔∗\(r\)‖\+\|λ−λ∗\(r\)\|\)−Cη2\.\-C\\eta\\bigl\(\\\|\\bm\{m\}\(\\bm\{X\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|\+\|\\lambda\-\\lambda^\{\*\}\(r\)\|\\bigr\)\-C\\eta^\{2\}\.Young’s inequality yields
Cη\(‖𝒎\(𝑿\)−𝒔∗\(r\)‖\+\|λ−λ∗\(r\)\|\)≤μ02\(‖𝒎\(𝑿\)−𝒔∗\(r\)‖2\+\|λ−λ∗\(r\)\|2\)\+Cη2\.C\\eta\\bigl\(\\\|\\bm\{m\}\(\\bm\{X\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|\+\|\\lambda\-\\lambda^\{\*\}\(r\)\|\\bigr\)\\leq\\frac\{\\mu\_\{0\}\}\{2\}\\left\(\\\|\\bm\{m\}\(\\bm\{X\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|^\{2\}\+\|\\lambda\-\\lambda^\{\*\}\(r\)\|^\{2\}\\right\)\+C\\eta^\{2\}\.Moreover, the squared triangle inequality, \([A\.13](https://arxiv.org/html/2608.14096#A1.E13)\), and the upper bound in \([A\.9](https://arxiv.org/html/2608.14096#A1.E9)\) give
‖𝒎\(𝑿\)−𝒔∗\(r\)‖2\+\|λ−λ∗\(r\)\|2\\displaystyle\\\|\\bm\{m\}\(\\bm\{X\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|^\{2\}\+\|\\lambda\-\\lambda^\{\*\}\(r\)\|^\{2\}≥12\(‖𝒎\(𝑿\)−𝒔η\(r\)‖2\+\|λ−λη\(r\)\|2\)−Cη2\\displaystyle\\qquad\\geq\\frac\{1\}\{2\}\\left\(\\\|\\bm\{m\}\(\\bm\{X\}\)\-\\bm\{s\}^\{\\eta\}\(r\)\\\|^\{2\}\+\|\\lambda\-\\lambda^\{\\eta\}\(r\)\|^\{2\}\\right\)\-C\\eta^\{2\}≥12CW𝒲\(qη\(r\),\(𝑿,λ\)\)−Cη2\.\\displaystyle\\qquad\\geq\\frac\{1\}\{2C\_\{W\}\}\\mathcal\{W\}\\bigl\(q^\{\\eta\}\(r\),\(\\bm\{X\},\\lambda\)\\bigr\)\-C\\eta^\{2\}\.Combining these bounds with \([A\.32](https://arxiv.org/html/2608.14096#A1.E32)\) gives
∑i=1N\(mi\(Xi\)−siη\(r\)\)\(gi′\(mi\(Xi\)\)\+λ\)\\displaystyle\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(X\_\{i\}\)\-s\_\{i\}^\{\\eta\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(X\_\{i\}\)\)\+\\lambda\\bigr\)\+\(λ−λη\(r\)\)\(rc−∑i=1Nmi\(Xi\)\+θ\(1−Fj\(Xj\)\)\(gj′\(mj\(Xj\)\)\+λ\)\)\\displaystyle\\quad\+\\bigl\(\\lambda\-\\lambda^\{\\eta\}\(r\)\\bigr\)\\left\(r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i\}\)\+\\theta\\bigl\(1\-F\_\{j\}\(X\_\{j\}\)\\bigr\)\\bigl\(g\_\{j\}^\{\\prime\}\(m\_\{j\}\(X\_\{j\}\)\)\+\\lambda\\bigr\)\\right\)≥μ04CW𝒲\(qη\(r\),\(𝑿,λ\)\)\+∑i=1Nmi\(Xi\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\+λ\(rc−r0\)\+−Cη2\.\\displaystyle\\qquad\\geq\\frac\{\\mu\_\{0\}\}\{4C\_\{W\}\}\\mathcal\{W\}\\bigl\(q^\{\\eta\}\(r\),\(\\bm\{X\},\\lambda\)\\bigr\)\+\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\+\\lambda\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\-C\\eta^\{2\}\.\(A\.33\)For a safe state, \([A\.6](https://arxiv.org/html/2608.14096#A1.E6)\) and \([3\.7](https://arxiv.org/html/2608.14096#S3.E7)\) show that the survival factors cancel, so the left side of \([A\.33](https://arxiv.org/html/2608.14096#A1.E33)\) is precisely⟨∇z𝒲\(qη\(r\),\(𝑿,λ\)\),Gθ\(𝑿,λ,r\)⟩\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q^\{\\eta\}\(r\),\(\\bm\{X\},\\lambda\)\),G^\{\\theta\}\(\\bm\{X\},\\lambda;r\)\\rangle\. At\(r,η,𝑿,λ\)=\(rt,ηt,𝑿t,λt\)\(r,\\eta,\\bm\{X\},\\lambda\)=\(r\_\{t\},\\eta\_\{t\},\\bm\{X\}\_\{t\},\\lambda\_\{t\}\), usingηt2=αt\\eta\_\{t\}^\{2\}=\\alpha\_\{t\}, we obtain
⟨∇z𝒲\(qt,zt\),Gθ\(𝑿t,λt,rt\)⟩\\displaystyle\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),G^\{\\theta\}\(\\bm\{X\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\\right\\rangle≥μ04CWW~t\+∑i=1Nmi\(Xi,t\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\+λt\(rtc−r0\)\+−Cαt\.\\displaystyle\\quad\\geq\\frac\{\\mu\_\{0\}\}\{4C\_\{W\}\}\\widetilde\{W\}\_\{t\}\+\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i,t\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\+\\lambda\_\{t\}\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\-C\\alpha\_\{t\}\.
*Substep 3: transfer from the ideal field to the executable stochastic update\.*Because∇z𝒲\(qt,zt\)\\nabla\_\{z\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\)isℋt\\mathcal\{H\}\_\{t\}\-measurable, \([A\.20](https://arxiv.org/html/2608.14096#A1.E20)\) yields
𝔼\[⟨∇z𝒲\(qt,zt\),G^t⟩\|ℋt\]=⟨∇z𝒲\(qt,zt\),Gθ\(𝑿t,λt;rt\)⟩\\displaystyle\\mathbb\{E\}\\left\[\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),\\widehat\{G\}\_\{t\}\\right\\rangle\\mathrel\{\\Big\|\}\\mathcal\{H\}\_\{t\}\\right\]=\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),G^\{\\theta\}\(\\bm\{X\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\\right\\rangle\+⟨∇z𝒲\(qt,zt\),Gθ\(𝒀t,λt,rt\)−Gθ\(𝑿t,λt,rt\)⟩\.\\displaystyle\\quad\+\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),G^\{\\theta\}\(\\bm\{Y\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\-G^\{\\theta\}\(\\bm\{X\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\\right\\rangle\.By \([A\.10](https://arxiv.org/html/2608.14096#A1.E10)\), \([A\.20](https://arxiv.org/html/2608.14096#A1.E20)\), and Young’s inequality, the second inner product has absolute value at most
CW~t‖𝒀t−𝑿t‖≤μupdW~t\+C‖𝒀t−𝑿t‖2\.C\\sqrt\{\\widetilde\{W\}\_\{t\}\}\\,\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\\leq\\mu\_\{\\mathrm\{upd\}\}\\widetilde\{W\}\_\{t\}\+C\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.Consequently,
𝔼\[⟨∇z𝒲\(qt,zt\),G^t⟩\|ℋt\]≥μupdW~t\\displaystyle\\mathbb\{E\}\\left\[\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),\\widehat\{G\}\_\{t\}\\right\\rangle\\mathrel\{\\Big\|\}\\mathcal\{H\}\_\{t\}\\right\]\\geq\\mu\_\{\\mathrm\{upd\}\}\\widetilde\{W\}\_\{t\}\+∑i=1Nmi\(Xi,t\)\(gi′\(si∗\(rt\)\)\+λ∗\(rt\)\)\+λt\(rtc−r0\)\+−C\(αt\+∥𝒀t−𝑿t∥2\)\.\\displaystyle\\quad\+\\sum\_\{i=1\}^\{N\}m\_\{i\}\(X\_\{i,t\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\_\{t\}\)\)\+\\lambda^\{\*\}\(r\_\{t\}\)\\bigr\)\+\\lambda\_\{t\}\(r\_\{t\}^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\-C\\bigl\(\\alpha\_\{t\}\+\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\\bigr\)\.Multiplying by−αt\-\\alpha\_\{t\}therefore proves \([A\.22](https://arxiv.org/html/2608.14096#A1.E22)\)\.
Part \(ii\): Comparator\-motion control\.
For brevity, writeDη\(r\):=dqη\(r\)drD^\{\\eta\}\(r\):=\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\.
The resource recursion gives
rt\+1−rt=rt−∑i=1NSi,tnt−1,𝔼\[rt\+1−rt∣ℋt\]=rt−∑i=1Nmi\(Yi,t\)nt−1\.r\_\{t\+1\}\-r\_\{t\}=\\frac\{r\_\{t\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\}\{n\_\{t\}\-1\},\\qquad\\mathbb\{E\}\[r\_\{t\+1\}\-r\_\{t\}\\mid\\mathcal\{H\}\_\{t\}\]=\\frac\{r\_\{t\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\}\{n\_\{t\}\-1\}\.\(A\.34\)We first derive the first\-order resource\-pairing estimate\. Fixr≥0r\\geq 0,0<η≤10<\\eta\\leq 1,0≤xi,yi≤X¯i0\\leq x\_\{i\},y\_\{i\}\\leq\\bar\{X\}\_\{i\}, and0≤λ≤λmax0\\leq\\lambda\\leq\\lambda\_\{\\max\}\. For this calculation, abbreviate
𝖶:=𝒲\(qη\(r\),\(𝒙,λ\)\),βmin:=mini∈\[N\]βi,C∇q:=2βmin−2\.\\mathsf\{W\}:=\\mathcal\{W\}\\bigl\(q^\{\\eta\}\(r\),\(\\bm\{x\},\\lambda\)\\bigr\),\\qquad\\beta\_\{\\min\}:=\\min\_\{i\\in\[N\]\}\\beta\_\{i\},\\qquad C\_\{\\nabla q\}:=\\sqrt\{2\}\\,\\beta\_\{\\min\}^\{\-2\}\.The derivative calculation behind \([A\.10](https://arxiv.org/html/2608.14096#A1.E10)\), together with the lower metric bound, gives the first estimate below\.[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\(iii\) gives the second:
‖∇q𝒲\(qη\(r\),\(𝒙,λ\)\)‖≤C∇q𝖶,‖dqη\(r\)dr‖≤Csm\.\\left\\\|\\nabla\_\{q\}\\mathcal\{W\}\\bigl\(q^\{\\eta\}\(r\),\(\\bm\{x\},\\lambda\)\\bigr\)\\right\\\|\\leq C\_\{\\nabla q\}\\sqrt\{\\mathsf\{W\}\},\\qquad\\left\\\|\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\\right\\\|\\leq C\_\{\\mathrm\{sm\}\}\.
Set
Aη:=1\+NCsm,AW:=2N,AY:=N\.A\_\{\\eta\}:=1\+\\sqrt\{N\}\\,C\_\{\\mathrm\{sm\}\},\\qquad A\_\{W\}:=\\sqrt\{2N\},\\qquad A\_\{Y\}:=\\sqrt\{N\}\.
Ifr≤r0r\\leq r\_\{0\}, the resource constraint binds andr=∑i=1Nsi∗\(r\)r=\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\)\. The smoothing approximation, the lower bound in \([A\.9](https://arxiv.org/html/2608.14096#A1.E9)\), and the one\-Lipschitz property of everymim\_\{i\}give
\|r−∑i=1Nmi\(yi\)\|\\displaystyle\\left\|r\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\\right\|≤‖𝒔∗\(r\)−𝒔η\(r\)‖1\+‖𝒔η\(r\)−𝒎\(𝒙\)‖1\+∑i=1N\|mi\(xi\)−mi\(yi\)\|\\displaystyle\\leq\\\|\\bm\{s\}^\{\*\}\(r\)\-\\bm\{s\}^\{\\eta\}\(r\)\\\|\_\{1\}\+\\\|\\bm\{s\}^\{\\eta\}\(r\)\-\\bm\{m\}\(\\bm\{x\}\)\\\|\_\{1\}\+\\sum\_\{i=1\}^\{N\}\|m\_\{i\}\(x\_\{i\}\)\-m\_\{i\}\(y\_\{i\}\)\|≤Aηη\+AW𝖶\+AY‖𝒚−𝒙‖\.\\displaystyle\\leq A\_\{\\eta\}\\eta\+A\_\{W\}\\sqrt\{\\mathsf\{W\}\}\+A\_\{Y\}\\\|\\bm\{y\}\-\\bm\{x\}\\\|\.Ifr0<r<r0\+ηr\_\{0\}<r<r\_\{0\}\+\\eta, then𝒔∗\(r\)=𝒔∗\(r0\)\\bm\{s\}^\{\*\}\(r\)=\\bm\{s\}^\{\*\}\(r\_\{0\}\),\|r−r0\|≤η\|r\-r\_\{0\}\|\\leq\\eta, and‖𝒔η\(r\)−𝒔∗\(r0\)‖≤Csmη\\\|\\bm\{s\}^\{\\eta\}\(r\)\-\\bm\{s\}^\{\*\}\(r\_\{0\}\)\\\|\\leq C\_\{\\mathrm\{sm\}\}\\eta\. Insertingr0=∑i=1Nsi∗\(r0\)r\_\{0\}=\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\_\{0\}\)gives the same bound\. Finally, ifr≥r0\+ηr\\geq r\_\{0\}\+\\eta,[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\(ii\)–\(iii\) givesdqη\(r\)dr=0\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}=0\. Thus all three resource regimes satisfy
\|⟨∇q𝒲\(qη\(r\),\(𝒙,λ\)\),dqη\(r\)dr\(r−∑i=1Nmi\(yi\)\)⟩\|\\displaystyle\\left\|\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\\bigl\(q^\{\\eta\}\(r\),\(\\bm\{x\},\\lambda\)\\bigr\),\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\\left\(r\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\\right\)\\right\\rangle\\right\|≤C∇qCsm𝖶\(Aηη\+AW𝖶\+AY‖𝒚−𝒙‖\)\.\\displaystyle\\qquad\\leq C\_\{\\nabla q\}C\_\{\\mathrm\{sm\}\}\\sqrt\{\\mathsf\{W\}\}\\left\(A\_\{\\eta\}\\eta\+A\_\{W\}\\sqrt\{\\mathsf\{W\}\}\+A\_\{Y\}\\\|\\bm\{y\}\-\\bm\{x\}\\\|\\right\)\.Usingab≤\(a2\+b2\)/2ab\\leq\(a^\{2\}\+b^\{2\}\)/2, a valid explicit choice is
Cpair:=C∇qCsm\(AW\+Aη\+AY2\)=2βmin−2Csm\[2N\+1\+NCsm\+N2\]\.C\_\{\\mathrm\{pair\}\}:=C\_\{\\nabla q\}C\_\{\\mathrm\{sm\}\}\\left\(A\_\{W\}\+\\frac\{A\_\{\\eta\}\+A\_\{Y\}\}\{2\}\\right\)=\\sqrt\{2\}\\,\\beta\_\{\\min\}^\{\-2\}C\_\{\\mathrm\{sm\}\}\\left\[\\sqrt\{2N\}\+\\frac\{1\+\\sqrt\{N\}C\_\{\\mathrm\{sm\}\}\+\\sqrt\{N\}\}\{2\}\\right\]\.\(A\.35\)Indeed, for all the stated arguments,
\|⟨∇q𝒲\(qη\(r\),\(𝒙,λ\)\),dqη\(r\)dr\(r−∑i=1Nmi\(yi\)\)⟩\|\\displaystyle\\left\|\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\\bigl\(q^\{\\eta\}\(r\),\(\\bm\{x\},\\lambda\)\\bigr\),\\frac\{\\mathrm\{d\}q^\{\\eta\}\(r\)\}\{\\mathrm\{d\}r\}\\left\(r\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\\right\)\\right\\rangle\\right\|≤Cpair\(𝒲\(qη\(r\),\(𝒙,λ\)\)\+‖𝒚−𝒙‖2\+η2\)\.\\displaystyle\\leq C\_\{\\mathrm\{pair\}\}\\left\(\\mathcal\{W\}\\bigl\(q^\{\\eta\}\(r\),\(\\bm\{x\},\\lambda\)\\bigr\)\+\\\|\\bm\{y\}\-\\bm\{x\}\\\|^\{2\}\+\\eta^\{2\}\\right\)\.
*Substep 1: control the two sources of comparator motion\.*Fixt<Tt<T\. The comparator increment is
qt\+1−qt\\displaystyle q\_\{t\+1\}\-q\_\{t\}=qηt\(rt\+1\)−qηt\(rt\)\\displaystyle=q^\{\\eta\_\{t\}\}\(r\_\{t\+1\}\)\-q^\{\\eta\_\{t\}\}\(r\_\{t\}\)\+qηt\+1\(rt\+1\)−qηt\(rt\+1\)\.\\displaystyle\\quad\+q^\{\\eta\_\{t\+1\}\}\(r\_\{t\+1\}\)\-q^\{\\eta\_\{t\}\}\(r\_\{t\+1\}\)\.Ifrt≤max\(D¯,r0\+ηt\)r\_\{t\}\\leq\\max\(\\overline\{D\},r\_\{0\}\+\\eta\_\{t\}\), thenr0≤D¯r\_\{0\}\\leq\\overline\{D\},ηt≤1\\eta\_\{t\}\\leq 1, and∑i=1NSi,t≤D¯\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\\leq\\overline\{D\}imply
\|rt\+1−rt\|=\|rt−∑i=1NSi,t\|nt−1≤D¯\+1nt−1\.\|r\_\{t\+1\}\-r\_\{t\}\|=\\frac\{\|r\_\{t\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\|\}\{n\_\{t\}\-1\}\\leq\\frac\{\\overline\{D\}\+1\}\{n\_\{t\}\-1\}\.Ifrt\>max\(D¯,r0\+ηt\)r\_\{t\}\>\\max\(\\overline\{D\},r\_\{0\}\+\\eta\_\{t\}\), then∑i=1NSi,t≤D¯<rt\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\\leq\\overline\{D\}<r\_\{t\}, sort\+1\>rt\>r0\+ηtr\_\{t\+1\}\>r\_\{t\}\>r\_\{0\}\+\\eta\_\{t\}\. BothDηt\(rt\)D^\{\\eta\_\{t\}\}\(r\_\{t\}\)andqηt\(rt\+1\)−qηt\(rt\)q^\{\\eta\_\{t\}\}\(r\_\{t\+1\}\)\-q^\{\\eta\_\{t\}\}\(r\_\{t\}\)vanish because the smoothed path is constant there\. Thus[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\(iii\) gives, in all cases,
‖qηt\(rt\+1\)−qηt\(rt\)−Dηt\(rt\)\(rt\+1−rt\)‖≤Cηt\(nt−1\)2\.\\left\\\|q^\{\\eta\_\{t\}\}\(r\_\{t\+1\}\)\-q^\{\\eta\_\{t\}\}\(r\_\{t\}\)\-D^\{\\eta\_\{t\}\}\(r\_\{t\}\)\(r\_\{t\+1\}\-r\_\{t\}\)\\right\\\|\\leq\\frac\{C\}\{\\eta\_\{t\}\(n\_\{t\}\-1\)^\{2\}\}\.
Moreover,
\|min\{t\+1,nt\+1\}−min\{t,nt\}\|≤1,ηt=χ\(τ0\+min\{t,nt\}\)−1/2\.\\left\|\\min\\\{t\+1,n\_\{t\+1\}\\\}\-\\min\\\{t,n\_\{t\}\\\}\\right\|\\leq 1,\\qquad\\eta\_\{t\}=\\sqrt\{\\chi\}\\,\\bigl\(\\tau\_\{0\}\+\\min\\\{t,n\_\{t\}\\\}\\bigr\)^\{\-1/2\}\.The mean\-value theorem and[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\(iv\) therefore give
‖qηt\+1\(rt\+1\)−qηt\(rt\+1\)‖≤Cαt3/2\.\\left\\\|q^\{\\eta\_\{t\+1\}\}\(r\_\{t\+1\}\)\-q^\{\\eta\_\{t\}\}\(r\_\{t\+1\}\)\\right\\\|\\leq C\\alpha\_\{t\}^\{3/2\}\.
*Substep 2: bound the first\-order resource increment\.*The vectors∇q𝒲\(qt,zt\)\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\)andDηt\(rt\)D^\{\\eta\_\{t\}\}\(r\_\{t\}\)areℋt\\mathcal\{H\}\_\{t\}\-measurable\. Hence \([A\.34](https://arxiv.org/html/2608.14096#A1.E34)\) gives the exact identity
𝔼\[⟨∇q𝒲\(qt,zt\),Dηt\(rt\)\(rt\+1−rt\)⟩\|ℋt\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),D^\{\\eta\_\{t\}\}\(r\_\{t\}\)\(r\_\{t\+1\}\-r\_\{t\}\)\\right\\rangle\\mathrel\{\\Big\|\}\\mathcal\{H\}\_\{t\}\\right\]=1nt−1⟨∇q𝒲\(qt,zt\),Dηt\(rt\)\(rt−∑i=1Nmi\(Yi,t\)\)⟩\.\\displaystyle\\quad=\\frac\{1\}\{n\_\{t\}\-1\}\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),D^\{\\eta\_\{t\}\}\(r\_\{t\}\)\\left\(r\_\{t\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\\right\)\\right\\rangle\.Applying the first\-order resource\-pairing estimate above at\(r,η,𝒙,𝒚,λ\)=\(rt,ηt,𝑿t,𝒀t,λt\)\(r,\\eta,\\bm\{x\},\\bm\{y\},\\lambda\)=\(r\_\{t\},\\eta\_\{t\},\\bm\{X\}\_\{t\},\\bm\{Y\}\_\{t\},\\lambda\_\{t\}\)yields
𝔼\[⟨∇q𝒲\(qt,zt\),Dηt\(rt\)\(rt\+1−rt\)⟩\|ℋt\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),D^\{\\eta\_\{t\}\}\(r\_\{t\}\)\(r\_\{t\+1\}\-r\_\{t\}\)\\right\\rangle\\mathrel\{\\Big\|\}\\mathcal\{H\}\_\{t\}\\right\]≤Cpairnt−1\(W~t\+‖𝒀t−𝑿t‖2\+ηt2\)\.\\displaystyle\\qquad\\leq\\frac\{C\_\{\\mathrm\{pair\}\}\}\{n\_\{t\}\-1\}\\bigl\(\\widetilde\{W\}\_\{t\}\+\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\+\\eta\_\{t\}^\{2\}\\bigr\)\.
Since
αt≥χτ0\+nt,τ0≤2χα0,\\alpha\_\{t\}\\geq\\frac\{\\chi\}\{\\tau\_\{0\}\+n\_\{t\}\},\\qquad\\tau\_\{0\}\\leq\\frac\{2\\chi\}\{\\alpha\_\{0\}\},the explicit choices in \([A\.23](https://arxiv.org/html/2608.14096#A1.E23)\) give, forχ≥χM\\chi\\geq\\chi\_\{M\}andn:=nt≥n0n:=n\_\{t\}\\geq n\_\{0\},
Cpair\(n−1\)αt≤2Cpairα0\(n−1\)\+2Cpairχ≤μupd8\.\\frac\{C\_\{\\mathrm\{pair\}\}\}\{\(n\-1\)\\alpha\_\{t\}\}\\leq\\frac\{2C\_\{\\mathrm\{pair\}\}\}\{\\alpha\_\{0\}\(n\-1\)\}\+\\frac\{2C\_\{\\mathrm\{pair\}\}\}\{\\chi\}\\leq\\frac\{\\mu\_\{\\mathrm\{upd\}\}\}\{8\}\.Moreover, for everyn≥2n\\geq 2,
1\(n−1\)αt≤2α0\(n−1\)\+nχ\(n−1\)≤2α0\+2,\\frac\{1\}\{\(n\-1\)\\alpha\_\{t\}\}\\leq\\frac\{2\}\{\\alpha\_\{0\}\(n\-1\)\}\+\\frac\{n\}\{\\chi\(n\-1\)\}\\leq\\frac\{2\}\{\\alpha\_\{0\}\}\+2,while
αt≥χ2χ/α0\+n≥min\{α04,12n\}\.\\alpha\_\{t\}\\geq\\frac\{\\chi\}\{2\\chi/\\alpha\_\{0\}\+n\}\\geq\\min\\left\\\{\\frac\{\\alpha\_\{0\}\}\{4\},\\frac\{1\}\{2n\}\\right\\\}\.Sincen≥n0n\\geq n\_\{0\}impliesn≥8n\\geq 8andn−1≥2/α0n\-1\\geq 2/\\sqrt\{\\alpha\_\{0\}\}, the last bound givesαt\(n−1\)2≥1\\alpha\_\{t\}\(n\-1\)^\{2\}\\geq 1\. We have therefore established
Cpairnt−1≤μupd8αt,1nt−1≤\(2α0\+2\)αt,ηt\(nt−1\)≥1\.\\frac\{C\_\{\\mathrm\{pair\}\}\}\{n\_\{t\}\-1\}\\leq\\frac\{\\mu\_\{\\mathrm\{upd\}\}\}\{8\}\\alpha\_\{t\},\\qquad\\frac\{1\}\{n\_\{t\}\-1\}\\leq\\left\(\\frac\{2\}\{\\alpha\_\{0\}\}\+2\\right\)\\alpha\_\{t\},\\qquad\\eta\_\{t\}\(n\_\{t\}\-1\)\\geq 1\.\(A\.36\)The first inequality absorbs theW~t\\widetilde\{W\}\_\{t\}\-term\. The other two will be used below\. Sinceηt2=αt\\eta\_\{t\}^\{2\}=\\alpha\_\{t\},
𝔼\[⟨∇q𝒲\(qt,zt\),Dηt\(rt\)\(rt\+1−rt\)⟩\|ℋt\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),D^\{\\eta\_\{t\}\}\(r\_\{t\}\)\(r\_\{t\+1\}\-r\_\{t\}\)\\right\\rangle\\mathrel\{\\Big\|\}\\mathcal\{H\}\_\{t\}\\right\]≤μupd8αtW~t\+Cαt‖𝒀t−𝑿t‖2\+Cαt2\.\\displaystyle\\qquad\\leq\\frac\{\\mu\_\{\\mathrm\{upd\}\}\}\{8\}\\alpha\_\{t\}\\widetilde\{W\}\_\{t\}\+C\\alpha\_\{t\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\+C\\alpha\_\{t\}^\{2\}\.
*Substep 3: bound the Taylor and radius remainders\.*By \([A\.10](https://arxiv.org/html/2608.14096#A1.E10)\),‖∇q𝒲\(qt,zt\)‖≤CW~t\\\|\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\)\\\|\\leq C\\sqrt\{\\widetilde\{W\}\_\{t\}\}\. Young’s inequality, the direct resource\-motion remainder above, and the large\-ntn\_\{t\}bounds give
\|⟨∇q𝒲\(qt,zt\),qηt\(rt\+1\)−qηt\(rt\)−Dηt\(rt\)\(rt\+1−rt\)⟩\|\\displaystyle\\left\|\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),q^\{\\eta\_\{t\}\}\(r\_\{t\+1\}\)\-q^\{\\eta\_\{t\}\}\(r\_\{t\}\)\-D^\{\\eta\_\{t\}\}\(r\_\{t\}\)\(r\_\{t\+1\}\-r\_\{t\}\)\\right\\rangle\\right\|≤CW~tηt\(nt−1\)2≤μupd8αtW~t\+Cαtηt2\(nt−1\)4\\displaystyle\\qquad\\leq\\frac\{C\\sqrt\{\\widetilde\{W\}\_\{t\}\}\}\{\\eta\_\{t\}\(n\_\{t\}\-1\)^\{2\}\}\\leq\\frac\{\\mu\_\{\\mathrm\{upd\}\}\}\{8\}\\alpha\_\{t\}\\widetilde\{W\}\_\{t\}\+\\frac\{C\}\{\\alpha\_\{t\}\\eta\_\{t\}^\{2\}\(n\_\{t\}\-1\)^\{4\}\}≤μupd8αtW~t\+Cαt2\.\\displaystyle\\qquad\\leq\\frac\{\\mu\_\{\\mathrm\{upd\}\}\}\{8\}\\alpha\_\{t\}\\widetilde\{W\}\_\{t\}\+C\\alpha\_\{t\}^\{2\}\.
To make the final Taylor\-remainder estimate explicit, setK0:=2/α0\+2K\_\{0\}:=2/\\alpha\_\{0\}\+2\. Sinceηt2=αt\\eta\_\{t\}^\{2\}=\\alpha\_\{t\}, the second inequality in \([A\.36](https://arxiv.org/html/2608.14096#A1.E36)\) gives
1αtηt2\(nt−1\)4=1αt2\(1nt−1\)4≤K04αt2\.\\frac\{1\}\{\\alpha\_\{t\}\\eta\_\{t\}^\{2\}\(n\_\{t\}\-1\)^\{4\}\}=\\frac\{1\}\{\\alpha\_\{t\}^\{2\}\}\\left\(\\frac\{1\}\{n\_\{t\}\-1\}\\right\)^\{4\}\\leq K\_\{0\}^\{4\}\\alpha\_\{t\}^\{2\}\.Thus the last term produced by Young’s inequality isO\(αt2\)O\(\\alpha\_\{t\}^\{2\}\)\. Similarly, the radius\-motion bound above and Young’s inequality give
\|⟨∇q𝒲\(qt,zt\),qηt\+1\(rt\+1\)−qηt\(rt\+1\)⟩\|\\displaystyle\\left\|\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),q^\{\\eta\_\{t\+1\}\}\(r\_\{t\+1\}\)\-q^\{\\eta\_\{t\}\}\(r\_\{t\+1\}\)\\right\\rangle\\right\|≤CW~tαt3/2≤μupd8αtW~t\+Cαt2\.\\displaystyle\\qquad\\leq C\\sqrt\{\\widetilde\{W\}\_\{t\}\}\\,\\alpha\_\{t\}^\{3/2\}\\leq\\frac\{\\mu\_\{\\mathrm\{upd\}\}\}\{8\}\\alpha\_\{t\}\\widetilde\{W\}\_\{t\}\+C\\alpha\_\{t\}^\{2\}\.Combining these estimates with the first\-order resource bound yields
𝔼\[⟨∇q𝒲\(qt,zt\),qt\+1−qt⟩\|ℋt\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\left\\langle\\nabla\_\{q\}\\mathcal\{W\}\(q\_\{t\},z\_\{t\}\),q\_\{t\+1\}\-q\_\{t\}\\right\\rangle\\mathrel\{\\Big\|\}\\mathcal\{H\}\_\{t\}\\right\]≤3μupd8αtW~t\+Cαt2\+Cαt‖𝒀t−𝑿t‖2\.\\displaystyle\\qquad\\leq\\frac\{3\\mu\_\{\\mathrm\{upd\}\}\}\{8\}\\alpha\_\{t\}\\widetilde\{W\}\_\{t\}\+C\\alpha\_\{t\}^\{2\}\+C\\alpha\_\{t\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.
Finally, the two direct motion bounds andηt\(nt−1\)≥1\\eta\_\{t\}\(n\_\{t\}\-1\)\\geq 1give
‖qt\+1−qt‖≤Cnt−1\+Cαt3/2≤Cαt\.\\\|q\_\{t\+1\}\-q\_\{t\}\\\|\\leq\\frac\{C\}\{n\_\{t\}\-1\}\+C\\alpha\_\{t\}^\{3/2\}\\leq C\\alpha\_\{t\}\.Consequently,
L𝒲Gαt‖qt\+1−qt‖\+L𝒲2‖qt\+1−qt‖2≤Cαt2L\_\{\\mathcal\{W\}\}G\\alpha\_\{t\}\\\|q\_\{t\+1\}\-q\_\{t\}\\\|\+\\frac\{L\_\{\\mathcal\{W\}\}\}\{2\}\\\|q\_\{t\+1\}\-q\_\{t\}\\\|^\{2\}\\leq C\\alpha\_\{t\}^\{2\}pathwise\. Adding this estimate to the preceding conditional bound and using3μupd/8≤μupd/23\\mu\_\{\\mathrm\{upd\}\}/8\\leq\\mu\_\{\\mathrm\{upd\}\}/2proves \([A\.24](https://arxiv.org/html/2608.14096#A1.E24)\), after enlarging the primitive constantCmC\_\{\\rm m\}\. No same\-period noise factorization has been used\. ∎
Fixed\-target restoring calculation for[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\(i\)\.
###### Proof\.
Proof\.
*Substep 1: expand the full unstabilized pairing\.*Fixr≥0r\\geq 0,0≤λ≤λmax0\\leq\\lambda\\leq\\lambda\_\{\\max\}, and0≤xi≤X¯i0\\leq x\_\{i\}\\leq\\bar\{X\}\_\{i\}for everyi∈\[N\]i\\in\[N\]\. By \([A\.6](https://arxiv.org/html/2608.14096#A1.E6)\) and \([3\.7](https://arxiv.org/html/2608.14096#S3.E7)\), the survival factors cancel and the pairing is
⟨∇z𝒲\(q∗\(r\),\(𝒙,λ\)\),Gθ\(𝒙,λ,r\)⟩\\displaystyle\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\(q^\{\*\}\(r\),\(\\bm\{x\},\\lambda\)\),G^\{\\theta\}\(\\bm\{x\},\\lambda;r\)\\right\\rangle=∑i\(mi\(xi\)−si∗\(r\)\)\(gi′\(mi\(xi\)\)\+λ\)\+\(λ−λ∗\(r\)\)\(rc−∑imi\(xi\)\)\\displaystyle=\\sum\_\{i\}\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\)\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(x\_\{i\}\)\)\+\\lambda\)\+\(\\lambda\-\\lambda^\{\*\}\(r\)\)\\left\(r^\{\\mathrm\{c\}\}\-\\sum\_\{i\}m\_\{i\}\(x\_\{i\}\)\\right\)\+θ\(λ−λ∗\(r\)\)\(1−Fj\(xj\)\)\(gj′\(mj\(xj\)\)\+λ\)\.\\displaystyle\+\\theta\(\\lambda\-\\lambda^\{\*\}\(r\)\)\(1\-F\_\{j\}\(x\_\{j\}\)\)\(g\_\{j\}^\{\\prime\}\(m\_\{j\}\(x\_\{j\}\)\)\+\\lambda\)\.Expand the first two terms by adding and subtracting the target KKT quantities:
∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(mi\(xi\)\)\+λ\)\\displaystyle\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(x\_\{i\}\)\)\+\\lambda\\bigr\)=∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(mi\(xi\)\)−gi′\(si∗\(r\)\)\)\\displaystyle=\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(x\_\{i\}\)\)\-g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\\bigr\)\+∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\\displaystyle\\qquad\+\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\+\(λ−λ∗\(r\)\)∑i=1N\(mi\(xi\)−si∗\(r\)\),\\displaystyle\\qquad\+\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\),\(λ−λ∗\(r\)\)\(rc−∑i=1Nmi\(xi\)\)\\displaystyle\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\left\(r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(x\_\{i\}\)\\right\)=\(λ−λ∗\(r\)\)\(rc−∑i=1Nsi∗\(r\)\)\\displaystyle=\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\left\(r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\)\\right\)−\(λ−λ∗\(r\)\)∑i=1N\(mi\(xi\)−si∗\(r\)\)\.\\displaystyle\\quad\-\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\.
Only after this algebraic cancellation do we use \([4\.2](https://arxiv.org/html/2608.14096#S4.E2)\) and \([A\.4](https://arxiv.org/html/2608.14096#A1.E4)\), which give
si∗\(r\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)=0,rc−∑i=1Nsi∗\(r\)=\(rc−r0\)\+,λ∗\(r\)\(rc−r0\)\+=0\.s\_\{i\}^\{\*\}\(r\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)=0,\\qquad r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\)=\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\},\\qquad\\lambda^\{\*\}\(r\)\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}=0\.More explicitly, \([4\.1](https://arxiv.org/html/2608.14096#S4.E1)\) implies, for everyii,
\(u−v\)\(gi′\(u\)−gi′\(v\)\)≥hiκi\(u−v\)2,u,v∈\[0,mi\(X¯i\)\]\.\\bigl\(u\-v\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(u\)\-g\_\{i\}^\{\\prime\}\(v\)\\bigr\)\\geq h\_\{i\}\\kappa\_\{i\}\(u\-v\)^\{2\},\\qquad u,v\\in\[0,m\_\{i\}\(\\bar\{X\}\_\{i\}\)\]\.Therefore, withu=mi\(xi\)u=m\_\{i\}\(x\_\{i\}\)andv=si∗\(r\)v=s\_\{i\}^\{\*\}\(r\),
∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(mi\(xi\)\)−gi′\(si∗\(r\)\)\)≥μg‖𝒎\(𝒙\)−𝒔∗\(r\)‖2,μg:=min1≤i≤Nhiκi\.\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(x\_\{i\}\)\)\-g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\\bigr\)\\geq\\mu\_\{g\}\\\|\\bm\{m\}\(\\bm\{x\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|^\{2\},\\qquad\\mu\_\{g\}:=\\min\_\{1\\leq i\\leq N\}h\_\{i\}\\kappa\_\{i\}\.This is the strong\-convexity inequality used in the final step below\.
∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(mi\(xi\)\)\+λ\)\+\(λ−λ∗\(r\)\)\(rc−∑i=1Nmi\(xi\)\)\\displaystyle\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(x\_\{i\}\)\)\+\\lambda\\bigr\)\+\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\left\(r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(x\_\{i\}\)\\right\)=∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(mi\(xi\)\)−gi′\(si∗\(r\)\)\)\+∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\\displaystyle=\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(x\_\{i\}\)\)\-g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\\bigr\)\+\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\+\(λ−λ∗\(r\)\)\(rc−∑i=1Nsi∗\(r\)\)\\displaystyle\\quad\+\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\left\(r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}s\_\{i\}^\{\*\}\(r\)\\right\)=∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(mi\(xi\)\)−gi′\(si∗\(r\)\)\)\+∑i=1Nmi\(xi\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\+λ\(rc−r0\)\+\\displaystyle=\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(x\_\{i\}\)\)\-g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\\bigr\)\+\\sum\_\{i=1\}^\{N\}m\_\{i\}\(x\_\{i\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\+\\lambda\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}≥μg‖𝒎\(𝒙\)−𝒔∗\(r\)‖2\+∑i=1Nmi\(xi\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\+λ\(rc−r0\)\+\.\\displaystyle\\geq\\mu\_\{g\}\\\|\\bm\{m\}\(\\bm\{x\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|^\{2\}\+\\sum\_\{i=1\}^\{N\}m\_\{i\}\(x\_\{i\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\+\\lambda\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\.\(A\.37\)
*Substep 2: the anchor restores the multiplier error\.*For the maximum\-margin anchorjj, the selected KKT path satisfies
gj′\(sj∗\(r\)\)\+λ∗\(r\)=0\.g\_\{j\}^\{\\prime\}\(s\_\{j\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)=0\.\(A\.38\)Indeed, forr\>0r\>0,λ∗\(r\)<bj−cj\\lambda^\{\*\}\(r\)<b\_\{j\}\-c\_\{j\}makes the anchor active, so its KKT condition is an equality\. Atr=0r=0,λ∗\(0\)=bj−cj=−gj′\(0\)\\lambda^\{\*\}\(0\)=b\_\{j\}\-c\_\{j\}=\-g\_\{j\}^\{\\prime\}\(0\)gives the same equality\. On the safe box,
1−Fj\(xj\)≥βj,\|gj′\(mj\(xj\)\)−gj′\(sj∗\(r\)\)\|≤Lg,j\|mj\(xj\)−sj∗\(r\)\|,Lg,j:=hjKjβj3\.1\-F\_\{j\}\(x\_\{j\}\)\\geq\\beta\_\{j\},\\qquad\|g\_\{j\}^\{\\prime\}\(m\_\{j\}\(x\_\{j\}\)\)\-g\_\{j\}^\{\\prime\}\(s\_\{j\}^\{\*\}\(r\)\)\|\\leq L\_\{g,j\}\|m\_\{j\}\(x\_\{j\}\)\-s\_\{j\}^\{\*\}\(r\)\|,\\qquad L\_\{g,j\}:=\\frac\{h\_\{j\}K\_\{j\}\}\{\\beta\_\{j\}^\{3\}\}\.Using \([A\.38](https://arxiv.org/html/2608.14096#A1.E38)\), the anchor pairing first separates exactly into a positive multiplier square and one cross term:
\(λ−λ∗\(r\)\)\(1−Fj\(xj\)\)\(gj′\(mj\(xj\)\)\+λ\)\\displaystyle\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\bigl\(1\-F\_\{j\}\(x\_\{j\}\)\\bigr\)\\bigl\(g\_\{j\}^\{\\prime\}\(m\_\{j\}\(x\_\{j\}\)\)\+\\lambda\\bigr\)=\(1−Fj\(xj\)\)\(\|λ−λ∗\(r\)\|2\+\(λ−λ∗\(r\)\)\(gj′\(mj\(xj\)\)−gj′\(sj∗\(r\)\)\)\)\.\\displaystyle\\quad=\\bigl\(1\-F\_\{j\}\(x\_\{j\}\)\\bigr\)\\left\(\|\\lambda\-\\lambda^\{\*\}\(r\)\|^\{2\}\+\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{j\}^\{\\prime\}\(m\_\{j\}\(x\_\{j\}\)\)\-g\_\{j\}^\{\\prime\}\(s\_\{j\}^\{\*\}\(r\)\)\\bigr\)\\right\)\.Since1−Fj\(xj\)≤11\-F\_\{j\}\(x\_\{j\}\)\\leq 1, the survival lower bound, the preceding Lipschitz bound, and Young’s inequality yield, in that order,
\(λ−λ∗\(r\)\)\(1−Fj\(xj\)\)\(gj′\(mj\(xj\)\)\+λ\)\\displaystyle\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\bigl\(1\-F\_\{j\}\(x\_\{j\}\)\\bigr\)\\bigl\(g\_\{j\}^\{\\prime\}\(m\_\{j\}\(x\_\{j\}\)\)\+\\lambda\\bigr\)≥βj\|λ−λ∗\(r\)\|2−\|λ−λ∗\(r\)\|\|gj′\(mj\(xj\)\)−gj′\(sj∗\(r\)\)\|\\displaystyle\\quad\\geq\\beta\_\{j\}\|\\lambda\-\\lambda^\{\*\}\(r\)\|^\{2\}\-\|\\lambda\-\\lambda^\{\*\}\(r\)\|\\,\|g\_\{j\}^\{\\prime\}\(m\_\{j\}\(x\_\{j\}\)\)\-g\_\{j\}^\{\\prime\}\(s\_\{j\}^\{\*\}\(r\)\)\|≥βj\|λ−λ∗\(r\)\|2−Lg,j\|λ−λ∗\(r\)\|\|mj\(xj\)−sj∗\(r\)\|\\displaystyle\\quad\\geq\\beta\_\{j\}\|\\lambda\-\\lambda^\{\*\}\(r\)\|^\{2\}\-L\_\{g,j\}\|\\lambda\-\\lambda^\{\*\}\(r\)\|\\,\|m\_\{j\}\(x\_\{j\}\)\-s\_\{j\}^\{\*\}\(r\)\|≥βj2\|λ−λ∗\(r\)\|2−Lg,j22βj\|mj\(xj\)−sj∗\(r\)\|2\.\\displaystyle\\quad\\geq\\frac\{\\beta\_\{j\}\}\{2\}\|\\lambda\-\\lambda^\{\*\}\(r\)\|^\{2\}\-\\frac\{L\_\{g,j\}^\{2\}\}\{2\\beta\_\{j\}\}\|m\_\{j\}\(x\_\{j\}\)\-s\_\{j\}^\{\*\}\(r\)\|^\{2\}\.\(A\.39\)
*Substep 3: chooseθ\\thetaand absorb the cross term\.*Use the choice ofθ\\thetaand the joint restoring modulusμ0\\mu\_\{0\}in \([A\.21](https://arxiv.org/html/2608.14096#A1.E21)\)\. The bound involvingLg,jL\_\{g,j\}is the one used for absorption below\.θ≤1\\theta\\leq 1is only a convenient normalization for the later uniform field bounds\. Multiplying \([A\.39](https://arxiv.org/html/2608.14096#A1.E39)\) byθ\\thetaand adding it to \([A\.37](https://arxiv.org/html/2608.14096#A1.E37)\) gives
∑i=1N\(mi\(xi\)−si∗\(r\)\)\(gi′\(mi\(xi\)\)\+λ\)\\displaystyle\\sum\_\{i=1\}^\{N\}\\bigl\(m\_\{i\}\(x\_\{i\}\)\-s\_\{i\}^\{\*\}\(r\)\\bigr\)\\bigl\(g\_\{i\}^\{\\prime\}\(m\_\{i\}\(x\_\{i\}\)\)\+\\lambda\\bigr\)\+\(λ−λ∗\(r\)\)\(rc−∑i=1Nmi\(xi\)\+θ\(1−Fj\(xj\)\)\(gj′\(mj\(xj\)\)\+λ\)\)\\displaystyle\\quad\+\\bigl\(\\lambda\-\\lambda^\{\*\}\(r\)\\bigr\)\\left\(r^\{\\mathrm\{c\}\}\-\\sum\_\{i=1\}^\{N\}m\_\{i\}\(x\_\{i\}\)\+\\theta\\bigl\(1\-F\_\{j\}\(x\_\{j\}\)\\bigr\)\\bigl\(g\_\{j\}^\{\\prime\}\(m\_\{j\}\(x\_\{j\}\)\)\+\\lambda\\bigr\)\\right\)≥\(μg−θLg,j22βj\)‖𝒎\(𝒙\)−𝒔∗\(r\)‖2\+θβj2\|λ−λ∗\(r\)\|2\\displaystyle\\quad\\geq\\left\(\\mu\_\{g\}\-\\frac\{\\theta L\_\{g,j\}^\{2\}\}\{2\\beta\_\{j\}\}\\right\)\\\|\\bm\{m\}\(\\bm\{x\}\)\-\\bm\{s\}^\{\*\}\(r\)\\\|^\{2\}\+\\frac\{\\theta\\beta\_\{j\}\}\{2\}\|\\lambda\-\\lambda^\{\*\}\(r\)\|^\{2\}\+∑i=1Nmi\(xi\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\+λ\(rc−r0\)\+\.\\displaystyle\\qquad\\quad\+\\sum\_\{i=1\}^\{N\}m\_\{i\}\(x\_\{i\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\+\\lambda\(r^\{\\mathrm\{c\}\}\-r\_\{0\}\)^\{\+\}\.Condition \([A\.21](https://arxiv.org/html/2608.14096#A1.E21)\) makes the first coefficient at leastμg/2\\mu\_\{g\}/2\. The definition ofμ0\\mu\_\{0\}then controls both squared errors\. Together with the inner\-product identity derived in Substep 1, this is exactly \([A\.32](https://arxiv.org/html/2608.14096#A1.E32)\)\.
∎
Details for the one\-step verification\.
This paragraph proves the three model\-specific conditions recorded in Appendix[A\.3](https://arxiv.org/html/2608.14096#A1.SS3), which in turn verify the hypotheses of[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\.
###### Proof\.
Proof\. For a fixed storeii, abbreviate
Si\(x\):=1−Fi\(x\),Mi:=mi\(X¯i\),wi\(u\):=Si\(mi−1\(u\)\)−2\.S\_\{i\}\(x\):=1\-F\_\{i\}\(x\),\\qquad M\_\{i\}:=m\_\{i\}\(\\bar\{X\}\_\{i\}\),\\qquad w\_\{i\}\(u\):=S\_\{i\}\\bigl\(m\_\{i\}^\{\-1\}\(u\)\\bigr\)^\{\-2\}\.On\[0,Mi\]\[0,M\_\{i\}\], the safe\-box bounds give
1≤wi≤βi−2,\|wi′\|≤2Kiβi−4\.1\\leq w\_\{i\}\\leq\\beta\_\{i\}^\{\-2\},\\qquad\|w\_\{i\}^\{\\prime\}\|\\leq 2K\_\{i\}\\beta\_\{i\}^\{\-4\}\.LetHi\(σ,x\)H\_\{i\}\(\\sigma,x\)be theii\-th primal summand in \([A\.8](https://arxiv.org/html/2608.14096#A1.E8)\)\. Its Hessian on the middle branch is
\(wi\(σ\)−Si\(x\)−1−Si\(x\)−11\+\(mi\(x\)−σ\)fi\(x\)Si\(x\)−2\),\\begin\{pmatrix\}w\_\{i\}\(\\sigma\)&\-S\_\{i\}\(x\)^\{\-1\}\\\\ \-S\_\{i\}\(x\)^\{\-1\}&1\+\\bigl\(m\_\{i\}\(x\)\-\\sigma\\bigr\)f\_\{i\}\(x\)S\_\{i\}\(x\)^\{\-2\}\\end\{pmatrix\},while the left and right branches replace the off\-diagonal entry by−1\-1and−Si\(X¯i\)−1\-S\_\{i\}\(\\bar\{X\}\_\{i\}\)^\{\-1\}, respectively, and have lower\-right entry11\. All three Hessians are uniformly bounded\. The values and first derivatives match atx=0x=0andx=X¯ix=\\bar\{X\}\_\{i\}by \([A\.7](https://arxiv.org/html/2608.14096#A1.E7)\)\. Integrating the uniform Hessian bound along line segments therefore proves the differentiability and Lipschitz joint\-gradient claim after summing overiiand adding the dual quadratic\.
For the executable bounds, initialization lies in the safe box\. Moreover,Ct≥∑i=1NIi,tC\_\{t\}\\geq\\sum\_\{i=1\}^\{N\}I\_\{i,t\}, so the water\-filling map is the infimum of a nonempty closed sublevel set of a continuous piecewise\-linear function of the adapted inputs and is measurable\. It gives
Ii,t≤Yi,t≤max\{Ii,t,Xi,t\},∑i=1NYi,t≤Ct\.I\_\{i,t\}\\leq Y\_\{i,t\}\\leq\\max\\\{I\_\{i,t\},X\_\{i,t\}\\\},\\qquad\\sum\_\{i=1\}^\{N\}Y\_\{i,t\}\\leq C\_\{t\}\.The inventory recursion and the next coordinatewise projection then preserve the safe box and adaptedness by induction\. The two possible values in \([3\.4](https://arxiv.org/html/2608.14096#S3.E4)\) give\|G^i,t\|≤Gi\|\\widehat\{G\}\_\{i,t\}\|\\leq G\_\{i\}, and
\|G^0,t\|≤D¯\+θ\|G^j,t\|≤D¯\+θGj\.\|\\widehat\{G\}\_\{0,t\}\|\\leq\\overline\{D\}\+\\theta\|\\widehat\{G\}\_\{j,t\}\|\\leq\\overline\{D\}\+\\theta G\_\{j\}\.Thus‖G^t‖≤G\\\|\\widehat\{G\}\_\{t\}\\\|\\leq G\. IfαtG≤1\\alpha\_\{t\}G\\leq 1, every coordinate of the pre\-projection update moves by at most one\. Sinceztz\_\{t\}lies in the safe state box and𝒵\\mathcal\{Z\}extends that box by one in every coordinate, this giveszt−αtG^t∈𝒵z\_\{t\}\-\\alpha\_\{t\}\\widehat\{G\}\_\{t\}\\in\\mathcal\{Z\}, which is the pre\-projection domain condition used above\.
It remains to verify projection compatibility\. Writeqt\+1=\(𝝈t\+1,ζt\+1\)q\_\{t\+1\}=\(\\bm\{\\sigma\}\_\{t\+1\},\\zeta\_\{t\+1\}\)and putyi=mi−1\(σi,t\+1\)y\_\{i\}=m\_\{i\}^\{\-1\}\(\\sigma\_\{i,t\+1\}\)\. The derivative∂xHi\(σi,t\+1,x\)\\partial\_\{x\}H\_\{i\}\(\\sigma\_\{i,t\+1\},x\)is nonpositive onx≤yix\\leq y\_\{i\}and nonnegative onx≥yix\\geq y\_\{i\}\. This follows directly from the three branches in \([A\.8](https://arxiv.org/html/2608.14096#A1.E8)\)\. If the threshold representation ofqt\+1q\_\{t\+1\}lies in𝒦\(rt\+1\)\\mathcal\{K\}\(r\_\{t\+1\}\), projecting each primal coordinate toward\[0,Ui\(rt\+1\)\]\[0,U\_\{i\}\(r\_\{t\+1\}\)\]cannot increaseHiH\_\{i\}\. The dual coordinate satisfies
\|Proj\[0,λmax\]\(λ\)−ζt\+1\|≤\|λ−ζt\+1\|\.\|\\operatorname\{Proj\}\_\{\[0,\\lambda\_\{\\max\}\]\}\(\\lambda\)\-\\zeta\_\{t\+1\}\|\\leq\|\\lambda\-\\zeta\_\{t\+1\}\|\.Summing the coordinatewise inequalities gives \([A\.18](https://arxiv.org/html/2608.14096#A1.E18)\), completing the verification\. ∎
Verification of the conditional mean used in \([A\.20](https://arxiv.org/html/2608.14096#A1.E20)\)\.Conditionally onℋt\\mathcal\{H\}\_\{t\}, demand independence gives
𝔼\[𝟏\{Di,t<Yi,t\}∣ℋt\]\\displaystyle\\mathbb\{E\}\[\\mathbf\{1\}\\\{D\_\{i,t\}<Y\_\{i,t\}\\\}\\mid\\mathcal\{H\}\_\{t\}\]=Fi\(Yi,t\),\\displaystyle=F\_\{i\}\(Y\_\{i,t\}\),𝔼\[𝟏\{Di,t≥Yi,t\}∣ℋt\]\\displaystyle\\mathbb\{E\}\[\\mathbf\{1\}\\\{D\_\{i,t\}\\geq Y\_\{i,t\}\\\}\\mid\\mathcal\{H\}\_\{t\}\]=1−Fi\(Yi,t\),\\displaystyle=1\-F\_\{i\}\(Y\_\{i,t\}\),𝔼\[Di,t∧Yi,t∣ℋt\]\\displaystyle\\mathbb\{E\}\[D\_\{i,t\}\\wedge Y\_\{i,t\}\\mid\\mathcal\{H\}\_\{t\}\]=mi\(Yi,t\)\.\\displaystyle=m\_\{i\}\(Y\_\{i,t\}\)\.Substitution gives the first equality in \([A\.20](https://arxiv.org/html/2608.14096#A1.E20)\)\. On\[0,X¯i\]\[0,\\bar\{X\}\_\{i\}\],FiF\_\{i\}isKiK\_\{i\}\-Lipschitz andmim\_\{i\}is one\-Lipschitz\. By the executable bounds in the checklist above, both𝑿t\\bm\{X\}\_\{t\}and𝒀t\\bm\{Y\}\_\{t\}lie in the fixed safe rectangle∏i\[0,X¯i\]\\prod\_\{i\}\[0,\\bar\{X\}\_\{i\}\]\. Hence the stabilized field is uniformly Lipschitz there, and comparing𝒀t\\bm\{Y\}\_\{t\}with𝑿t\\bm\{X\}\_\{t\}proves the bias bound\.
## Appendix BFluid Geometry and Remaining Proofs
This appendix proves the remaining fluid\-geometry, implementation\-gap, and tuning claims and sketches the endpoint\-envelope extension\. Constants depend only on fixed primitives and horizon\-independent tuning\.
### B\.1Fluid Benchmark and Re\-Solving\-Path Geometry
#### Proof of[Section2](https://arxiv.org/html/2608.14096#S2)
###### Proof\.
Proof\. We first verify the two direct calculations stated before the proposition\. The identities
Yi,t−Ii,t=Si,t\+Ii,t\+1−Ii,t,BT\+1=W−∑t=1T∑i=1N\(Yi,t−Ii,t\)Y\_\{i,t\}\-I\_\{i,t\}=S\_\{i,t\}\+I\_\{i,t\+1\}\-I\_\{i,t\},\\qquad B\_\{T\+1\}=W\-\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\(Y\_\{i,t\}\-I\_\{i,t\}\)together withIi,1=0I\_\{i,1\}=0andci=ki−wc\_\{i\}=k\_\{i\}\-wgive
∑t=1T∑i=1N\(ki\(Yi,t−Ii,t\)\+hiIi,t\+1\+bi\(Di,t−Yi,t\)\+\)\+wBT\+1−∑i=1NciIi,T\+1\\displaystyle\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\\bigl\(k\_\{i\}\(Y\_\{i,t\}\-I\_\{i,t\}\)\+h\_\{i\}I\_\{i,t\+1\}\+b\_\{i\}\(D\_\{i,t\}\-Y\_\{i,t\}\)^\{\+\}\\bigr\)\+wB\_\{T\+1\}\-\\sum\_\{i=1\}^\{N\}c\_\{i\}I\_\{i,T\+1\}=wW\+∑t=1T∑i=1N\(ci\(Di,t∧Yi,t\)\+hi\(Yi,t−Di,t\)\+\+bi\(Di,t−Yi,t\)\+\)\.\\displaystyle\\qquad=wW\+\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\\bigl\(c\_\{i\}\(D\_\{i,t\}\\wedge Y\_\{i,t\}\)\+h\_\{i\}\(Y\_\{i,t\}\-D\_\{i,t\}\)^\{\+\}\+b\_\{i\}\(D\_\{i,t\}\-Y\_\{i,t\}\)^\{\+\}\\bigr\)\.Since𝒀t\\bm\{Y\}\_\{t\}is chosen before current demand, the conditional expectation of the bracketed store–period term on the right\-hand side isℓi\(Yi,t\)\\ell\_\{i\}\(Y\_\{i,t\}\)\. Taking expectations proves \([2\.6](https://arxiv.org/html/2608.14096#S2.E6)\)\.
Moreover, shipments only relocate inventory, whereas sales deplete it\. Summing the aggregate system\-inventory recursion therefore gives
BT\+1\+∑i=1NIi,T\+1=W−∑t=1T∑i=1NSi,t\.B\_\{T\+1\}\+\\sum\_\{i=1\}^\{N\}I\_\{i,T\+1\}=W\-\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\.Taking expectations and using𝔼π\[Si,t∣ℋtπ\]=mi\(Yi,t\)\\mathbb\{E\}^\{\\pi\}\[S\_\{i,t\}\\mid\\mathcal\{H\}\_\{t\}^\{\\pi\}\]=m\_\{i\}\(Y\_\{i,t\}\)proves \([2\.7](https://arxiv.org/html/2608.14096#S2.E7)\)\.
The expected\-sales transformation in[Section4\.2](https://arxiv.org/html/2608.14096#S4.SS2)makesgig\_\{i\}convex and writes
v\(r\)=min\{∑igi\(si\):∑isi≤r,0≤si≤μi\}\.v\(r\)=\\min\\left\\\{\\sum\_\{i\}g\_\{i\}\(s\_\{i\}\):\\sum\_\{i\}s\_\{i\}\\leq r,\\ 0\\leq s\_\{i\}\\leq\\mu\_\{i\}\\right\\\}\.Hencevvis convex by mixing feasible allocations and nonincreasing because its feasible set expands withrr\.
If0≤yi≤d¯i0\\leq y\_\{i\}\\leq\\bar\{d\}\_\{i\}, then𝒚\\bm\{y\}is feasible for the program definingv\(∑i=1Nmi\(yi\)\)v\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\)\. If someyi\>d¯iy\_\{i\}\>\\bar\{d\}\_\{i\}, clipping it tod¯i\\bar\{d\}\_\{i\}leavesmi\(yi\)m\_\{i\}\(y\_\{i\}\)unchanged and weakly decreasesℓi\(yi\)\\ell\_\{i\}\(y\_\{i\}\)\. Therefore, in either case,
∑i=1Nℓi\(yi\)≥v\(∑i=1Nmi\(yi\)\)pointwise in𝒚\.\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(y\_\{i\}\)\\geq v\\\!\\left\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(y\_\{i\}\)\\right\)\\qquad\\text\{pointwise in \}\\bm\{y\}\.\(B\.1\)
Apply \([B\.1](https://arxiv.org/html/2608.14096#A2.E1)\) to the random action𝒀t\\bm\{Y\}\_\{t\}\. Taking expectations and using Jensen again,
𝔼π∑i=1Nℓi\(Yi,t\)≥𝔼πv\(∑i=1Nmi\(Yi,t\)\)≥v\(𝔼π∑i=1Nmi\(Yi,t\)\)\.\\mathbb\{E\}^\{\\pi\}\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(Y\_\{i,t\}\)\\geq\\mathbb\{E\}^\{\\pi\}v\\\!\\left\(\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\\right\)\\geq v\\\!\\left\(\\mathbb\{E\}^\{\\pi\}\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\\right\)\.Summing over periods and applying Jensen across periods gives
∑t=1T𝔼π∑i=1Nℓi\(Yi,t\)\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}^\{\\pi\}\\sum\_\{i=1\}^\{N\}\\ell\_\{i\}\(Y\_\{i,t\}\)≥∑t=1Tv\(𝔼π∑i=1Nmi\(Yi,t\)\)\\displaystyle\\geq\\sum\_\{t=1\}^\{T\}v\\\!\\left\(\\mathbb\{E\}^\{\\pi\}\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\\right\)≥Tv\(1T∑t=1T𝔼π∑i=1Nmi\(Yi,t\)\)≥Tv\(γ\),\\displaystyle\\geq Tv\\\!\\left\(\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}^\{\\pi\}\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\\right\)\\geq Tv\(\\gamma\),where the last inequality uses∑t=1T𝔼π∑i=1Nmi\(Yi,t\)≤W=γT\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}^\{\\pi\}\\sum\_\{i=1\}^\{N\}m\_\{i\}\(Y\_\{i,t\}\)\\leq W=\\gamma Tfrom \([2\.7](https://arxiv.org/html/2608.14096#S2.E7)\) and the fact thatvvis nonincreasing\. Finally, \([2\.6](https://arxiv.org/html/2608.14096#S2.E6)\) impliesJTπ≥wW\+Tv\(γ\)J\_\{T\}^\{\\pi\}\\geq wW\+Tv\(\\gamma\)\. Taking the infimum over all admissible policies proves the oracle bound\. ∎
#### Derivation of sales\-space curvature
###### Proof\.
Derivation\. Fory=mi−1\(s\)y=m\_\{i\}^\{\-1\}\(s\), \([2\.4](https://arxiv.org/html/2608.14096#S2.E4)\) and \([2\.5](https://arxiv.org/html/2608.14096#S2.E5)\) give
mi′\(y\)=1−Fi\(y\),mi′′\(y\)=−fi\(y\),m\_\{i\}^\{\\prime\}\(y\)=1\-F\_\{i\}\(y\),\\qquad m\_\{i\}^\{\\prime\\prime\}\(y\)=\-f\_\{i\}\(y\),and
ℓi′\(y\)=\(hi\+bi−ci\)Fi\(y\)−\(bi−ci\),ℓi′′\(y\)=\(hi\+bi−ci\)fi\(y\)\.\\ell\_\{i\}^\{\\prime\}\(y\)=\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)F\_\{i\}\(y\)\-\(b\_\{i\}\-c\_\{i\}\),\\qquad\\ell\_\{i\}^\{\\prime\\prime\}\(y\)=\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)f\_\{i\}\(y\)\.Thusgi′\(s\)=ℓi′\(y\)/mi′\(y\)g\_\{i\}^\{\\prime\}\(s\)=\\ell\_\{i\}^\{\\prime\}\(y\)/m\_\{i\}^\{\\prime\}\(y\)\. Differentiating once more,
gi′′\(s\)\\displaystyle g\_\{i\}^\{\\prime\\prime\}\(s\)=ℓi′′\(y\)mi′\(y\)−ℓi′\(y\)mi′′\(y\)mi′\(y\)3\\displaystyle=\\frac\{\\ell\_\{i\}^\{\\prime\\prime\}\(y\)m\_\{i\}^\{\\prime\}\(y\)\-\\ell\_\{i\}^\{\\prime\}\(y\)m\_\{i\}^\{\\prime\\prime\}\(y\)\}\{m\_\{i\}^\{\\prime\}\(y\)^\{3\}\}=fi\(y\)\(\(hi\+bi−ci\)\(1−Fi\(y\)\)\+\(hi\+bi−ci\)Fi\(y\)−\(bi−ci\)\)\(1−Fi\(y\)\)3\\displaystyle=\\frac\{f\_\{i\}\(y\)\\bigl\(\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)\(1\-F\_\{i\}\(y\)\)\+\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)F\_\{i\}\(y\)\-\(b\_\{i\}\-c\_\{i\}\)\\bigr\)\}\{\(1\-F\_\{i\}\(y\)\)^\{3\}\}=hifi\(y\)\(1−Fi\(y\)\)3\.\\displaystyle=\\frac\{h\_\{i\}f\_\{i\}\(y\)\}\{\(1\-F\_\{i\}\(y\)\)^\{3\}\}\.∎
#### Proof of[Section4\.2](https://arxiv.org/html/2608.14096#S4.SS2)
###### Proof\.
Proof\. For the proof, define the fixed\-price response
ϕi\(λ\):=\{mi\(Fi−1\(bi−ci−λhi\+bi−ci−λ\)\),λ<bi−ci,0,λ≥bi−ci,Φ\(λ\):=∑i=1Nϕi\(λ\),\\phi\_\{i\}\(\\lambda\):=\\begin\{cases\}m\_\{i\}\\\!\\left\(F\_\{i\}^\{\-1\}\\\!\\left\(\\dfrac\{b\_\{i\}\-c\_\{i\}\-\\lambda\}\{h\_\{i\}\+b\_\{i\}\-c\_\{i\}\-\\lambda\}\\right\)\\right\),&\\lambda<b\_\{i\}\-c\_\{i\},\\\\\[6\.0pt\] 0,&\\lambda\\geq b\_\{i\}\-c\_\{i\},\\end\{cases\}\\qquad\\Phi\(\\lambda\):=\\sum\_\{i=1\}^\{N\}\\phi\_\{i\}\(\\lambda\),\(B\.2\)and writeϕ\(λ\):=\(ϕi\(λ\)\)i=1N\\bm\{\\phi\}\(\\lambda\):=\(\\phi\_\{i\}\(\\lambda\)\)\_\{i=1\}^\{N\}\. Also let
𝒥max:=\{i:bi−ci=λmax\},Lv:=\(∑j∈𝒥maxpj2\(hj\+bj−cj\)Kj\)−1\.\\mathcal\{J\}\_\{\\max\}:=\\\{i:b\_\{i\}\-c\_\{i\}=\\lambda\_\{\\max\}\\\},\\qquad L\_\{v\}:=\\left\(\\sum\_\{j\\in\\mathcal\{J\}\_\{\\max\}\}\\frac\{p\_\{j\}^\{2\}\}\{\(h\_\{j\}\+b\_\{j\}\-c\_\{j\}\)K\_\{j\}\}\\right\)^\{\-1\}\.These quantities depend only on the fixed model primitives\.
For fixedλ≥0\\lambda\\geq 0, the derivative ofgi\(s\)\+λsg\_\{i\}\(s\)\+\\lambda sats=0s=0is−\(bi−ci\)\+λ\-\(b\_\{i\}\-c\_\{i\}\)\+\\lambda\. The strict convexity in \([4\.1](https://arxiv.org/html/2608.14096#S4.E1)\) therefore makess=0s=0the unique minimizer whenλ≥bi−ci\\lambda\\geq b\_\{i\}\-c\_\{i\}\. Whenλ<bi−ci\\lambda<b\_\{i\}\-c\_\{i\}, the unique interior stationary point solves
\(hi\+bi−ci\)Fi\(y\)−\(bi−ci\)1−Fi\(y\)\+λ=0,\\frac\{\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\)F\_\{i\}\(y\)\-\(b\_\{i\}\-c\_\{i\}\)\}\{1\-F\_\{i\}\(y\)\}\+\\lambda=0,which gives the formula forϕi\(λ\)\\phi\_\{i\}\(\\lambda\)in \([B\.2](https://arxiv.org/html/2608.14096#A2.E2)\)\. For the derivative calculation below, write
yiλ:=mi−1\(ϕi\(λ\)\)\.y\_\{i\}^\{\\lambda\}:=m\_\{i\}^\{\-1\}\(\\phi\_\{i\}\(\\lambda\)\)\.This quantile lies strictly belowd¯i\\bar\{d\}\_\{i\}, so the upper sales constraint is slack\. Asλ↑bi−ci\\lambda\\uparrow b\_\{i\}\-c\_\{i\}, it decreases to zero\. Hence eachϕi\(λ\)\\phi\_\{i\}\(\\lambda\), and thereforeΦ\\Phi, is continuous at every breakpoint\. At the endpoints, the definitions give∑i=1Nϕi\(0\)=r0\\sum\_\{i=1\}^\{N\}\\phi\_\{i\}\(0\)=r\_\{0\}, whileϕi\(λmax\)=0\\phi\_\{i\}\(\\lambda\_\{\\max\}\)=0for everyii\.
Forλ<bi−ci\\lambda<b\_\{i\}\-c\_\{i\}, implicit differentiation gives
dϕi\(λ\)dλ=−\(1−Fi\(yiλ\)\)2\(hi\+bi−ci−λ\)fi\(yiλ\)\.\\frac\{\\mathrm\{d\}\\phi\_\{i\}\(\\lambda\)\}\{\\mathrm\{d\}\\lambda\}=\-\\frac\{\(1\-F\_\{i\}\(y\_\{i\}^\{\\lambda\}\)\)^\{2\}\}\{\(h\_\{i\}\+b\_\{i\}\-c\_\{i\}\-\\lambda\)f\_\{i\}\(y\_\{i\}^\{\\lambda\}\)\}\.Every maximum\-margin store is active forλ<λmax\\lambda<\\lambda\_\{\\max\}\. For such a storejj,
1−Fj\(yjλ\)=hjhj\+bj−cj−λ≥pj,\(hj\+bj−cj−λ\)fj\(yjλ\)≤\(hj\+bj−cj\)Kj,1\-F\_\{j\}\(y\_\{j\}^\{\\lambda\}\)=\\frac\{h\_\{j\}\}\{h\_\{j\}\+b\_\{j\}\-c\_\{j\}\-\\lambda\}\\geq p\_\{j\},\\qquad\(h\_\{j\}\+b\_\{j\}\-c\_\{j\}\-\\lambda\)f\_\{j\}\(y\_\{j\}^\{\\lambda\}\)\\leq\(h\_\{j\}\+b\_\{j\}\-c\_\{j\}\)K\_\{j\},so summing over𝒥max\\mathcal\{J\}\_\{\\max\}yields
−Φ′\(λ\)≥Lv−1\>0a\.e\. on\[0,λmax\)\.\-\\Phi^\{\\prime\}\(\\lambda\)\\geq L\_\{v\}^\{\-1\}\>0\\quad\\text\{a\.e\.\\ on \}\[0,\\lambda\_\{\\max\}\)\.Conversely,
\|ϕi′\(λ\)\|≤1hiκi\|\\phi\_\{i\}^\{\\prime\}\(\\lambda\)\|\\leq\\frac\{1\}\{h\_\{i\}\\kappa\_\{i\}\}where the derivative exists\. Integrating on the finitely many intervals cut out by the distinct values ofbi−cib\_\{i\}\-c\_\{i\}shows thatΦ\\Phiis Lipschitz and, for0≤λ<λ′≤λmax0\\leq\\lambda<\\lambda^\{\\prime\}\\leq\\lambda\_\{\\max\},
Φ\(λ\)−Φ\(λ′\)≥Lv−1\(λ′−λ\)\>0\.\\Phi\(\\lambda\)\-\\Phi\(\\lambda^\{\\prime\}\)\\geq L\_\{v\}^\{\-1\}\(\\lambda^\{\\prime\}\-\\lambda\)\>0\.ThusΦ\\Phiis strictly decreasing fromr0r\_\{0\}to00, its inverse is well defined, and
\|Φ−1\(r′\)−Φ−1\(r\)\|≤Lv\|r′−r\|,r,r′∈\[0,r0\]\.\|\\Phi^\{\-1\}\(r^\{\\prime\}\)\-\\Phi^\{\-1\}\(r\)\|\\leq L\_\{v\}\|r^\{\\prime\}\-r\|,\\qquad r,r^\{\\prime\}\\in\[0,r\_\{0\}\]\.
We now establish the explicit representation
\(𝒔∗\(r\),λ∗\(r\)\)=\{\(ϕ\(Φ−1\(r\)\),Φ−1\(r\)\),0<r<r0,\(𝟎,λmax\),r=0,\(ϕ\(0\),0\),r≥r0\.\(\\bm\{s\}^\{\*\}\(r\),\\lambda^\{\*\}\(r\)\)=\\begin\{cases\}\(\\bm\{\\phi\}\(\\Phi^\{\-1\}\(r\)\),\\Phi^\{\-1\}\(r\)\),&0<r<r\_\{0\},\\\\ \(\\bm\{0\},\\lambda\_\{\\max\}\),&r=0,\\\\ \(\\bm\{\\phi\}\(0\),0\),&r\\geq r\_\{0\}\.\\end\{cases\}\(B\.3\)
For0<r<r00<r<r\_\{0\}, chooseλ=Φ−1\(r\)\\lambda=\\Phi^\{\-1\}\(r\)\. The coordinate minimizers sum torrand satisfy the KKT conditions, so strict convexity gives the unique global optimizer\. At the other endpointr=0r=0, primal feasibility forces𝒔=𝟎\\bm\{s\}=\\bm\{0\}\. Because every sales coordinate is then at its lower boundary, stationarity requires
gi′\(0\)\+λ=λ−\(bi−ci\)≥0for everyi\.g\_\{i\}^\{\\prime\}\(0\)\+\\lambda=\\lambda\-\(b\_\{i\}\-c\_\{i\}\)\\geq 0\\qquad\\text\{for every \}i\.Equivalently,λ≥maxi\(bi−ci\)=λmax\\lambda\\geq\\max\_\{i\}\(b\_\{i\}\-c\_\{i\}\)=\\lambda\_\{\\max\}\. Conversely, everyλ≥λmax\\lambda\\geq\\lambda\_\{\\max\}satisfies these stationarity inequalities, and complementary slackness also holds because0−∑i=1Nsi=00\-\\sum\_\{i=1\}^\{N\}s\_\{i\}=0\. Hence the KKT multiplier is not unique atr=0r=0: its admissible set is\[λmax,∞\)\[\\lambda\_\{\\max\},\\infty\)\. We select its smallest element,λ∗\(0\)=λmax\\lambda^\{\*\}\(0\)=\\lambda\_\{\\max\}\. This convention also matchesΦ−1\(r\)↑λmax\\Phi^\{\-1\}\(r\)\\uparrow\\lambda\_\{\\max\}asr↓0r\\downarrow 0, so the selected dual target joins continuously to the unique multiplier forr\>0r\>0\. This establishes \([B\.3](https://arxiv.org/html/2608.14096#A2.E3)\), uniqueness of the primal target, and minimality of the displayed multiplier\.
The inverse bound is exactly the multiplier bound in[Section4\.2](https://arxiv.org/html/2608.14096#S4.SS2)\(iii\) on\[0,r0\]\[0,r\_\{0\}\], includingr=0r=0by the continuous endpoint selection above\. The same bound holds globally becauseλ∗\(r\)=0\\lambda^\{\*\}\(r\)=0on\[r0,∞\)\[r\_\{0\},\\infty\)\. Since\|ϕi′\(λ\)\|≤\(hiκi\)−1\|\\phi\_\{i\}^\{\\prime\}\(\\lambda\)\|\\leq\(h\_\{i\}\\kappa\_\{i\}\)^\{\-1\}wherever the derivative exists, the representation in \([B\.3](https://arxiv.org/html/2608.14096#A2.E3)\) gives, for allr,r′≥0r,r^\{\\prime\}\\geq 0,
\|si∗\(r′\)−si∗\(r\)\|≤Lvhiκi\|r′−r\|\.\|s\_\{i\}^\{\*\}\(r^\{\\prime\}\)\-s\_\{i\}^\{\*\}\(r\)\|\\leq\\frac\{L\_\{v\}\}\{h\_\{i\}\\kappa\_\{i\}\}\|r^\{\\prime\}\-r\|\.Thusq∗q^\{\*\}is globally Lipschitz\. Finally,ϕi\(λ\)\\phi\_\{i\}\(\\lambda\)is nonincreasing inλ\\lambda, whileλ∗\(r\)\\lambda^\{\*\}\(r\)is nonincreasing inrr\. Hence eachsi∗\(r\)s\_\{i\}^\{\*\}\(r\)is nondecreasing\. Integration across the finitely many breakpoints makes all of these conclusions valid when several breakpoints tie\.
The coordinate minimization above also givessi∗\(r\)<μis\_\{i\}^\{\*\}\(r\)<\\mu\_\{i\}, the storewise activity rule, and the KKT residuals in \([4\.2](https://arxiv.org/html/2608.14096#S4.E2)\)\. The representation \([B\.3](https://arxiv.org/html/2608.14096#A2.E3)\) yields∑isi∗\(r\)=min\{r,r0\}\\sum\_\{i\}s\_\{i\}^\{\*\}\(r\)=\\min\\\{r,r\_\{0\}\\\}and all endpoint claims in[Section4\.2](https://arxiv.org/html/2608.14096#S4.SS2)\(ii\)\. Finally, standard value sensitivity givesv′\(r\)=−λ∗\(r\)v^\{\\prime\}\(r\)=\-\\lambda^\{\*\}\(r\)\. Continuity of the selected multiplier makesvvcontinuously differentiable, and the multiplier Lipschitz bound gives the stated Lipschitz bound forv′v^\{\\prime\}\. ∎
### B\.2Implementation Gap
#### Proof of[Section4\.5](https://arxiv.org/html/2608.14096#S4.SS5)
###### Proof\.
Proof\. For every store and period, define the excess inventory above the preferred threshold by
ei,t:=\(Ii,t−Xi,t\)\+\.e\_\{i,t\}:=\(I\_\{i,t\}\-X\_\{i,t\}\)^\{\+\}\.Also let
H0:=⌈∑i=1Npi−1⌉,X¯max:=max1≤i≤NX¯i\.H\_\{0\}:=\\left\\lceil\\sum\_\{i=1\}^\{N\}p\_\{i\}^\{\-1\}\\right\\rceil,\\qquad\\bar\{X\}\_\{\\max\}:=\\max\_\{1\\leq i\\leq N\}\\bar\{X\}\_\{i\}\.We organize the proof in four steps\.
Step 1: Bound preferred\-path variation and the physical state\.Fort<Tt<T, the update in[Algorithm1](https://arxiv.org/html/2608.14096#alg1)and the feasibility identityXi,t=Proj\[0,Ui\(rt\)\]\(Xi,t\)X\_\{i,t\}=\\operatorname\{Proj\}\_\{\[0,U\_\{i\}\(r\_\{t\}\)\]\}\(X\_\{i,t\}\)imply
\|Xi,t\+1−Xi,t\|\\displaystyle\|X\_\{i,t\+1\}\-X\_\{i,t\}\|≤\|Proj\[0,Ui\(rt\+1\)\]\(Xi,t−αtG^i,t\)−Proj\[0,Ui\(rt\+1\)\]\(Xi,t\)\|\\displaystyle\\leq\\left\|\\operatorname\{Proj\}\_\{\[0,U\_\{i\}\(r\_\{t\+1\}\)\]\}\(X\_\{i,t\}\-\\alpha\_\{t\}\\widehat\{G\}\_\{i,t\}\)\-\\operatorname\{Proj\}\_\{\[0,U\_\{i\}\(r\_\{t\+1\}\)\]\}\(X\_\{i,t\}\)\\right\|\+\|Proj\[0,Ui\(rt\+1\)\]\(Xi,t\)−Proj\[0,Ui\(rt\)\]\(Xi,t\)\|\\displaystyle\\quad\+\\left\|\\operatorname\{Proj\}\_\{\[0,U\_\{i\}\(r\_\{t\+1\}\)\]\}\(X\_\{i,t\}\)\-\\operatorname\{Proj\}\_\{\[0,U\_\{i\}\(r\_\{t\}\)\]\}\(X\_\{i,t\}\)\\right\|≤αt\|G^i,t\|\+\|Ui\(rt\+1\)−Ui\(rt\)\|\.\\displaystyle\\leq\\alpha\_\{t\}\|\\widehat\{G\}\_\{i,t\}\|\+\|U\_\{i\}\(r\_\{t\+1\}\)\-U\_\{i\}\(r\_\{t\}\)\|\.\(B\.4\)The last inequality uses nonexpansiveness of projection onto a fixed interval and the elementary endpoint bound
\|Proj\[0,a\]\(x\)−Proj\[0,b\]\(x\)\|≤\|a−b\|,a,b,x≥0\.\|\\operatorname\{Proj\}\_\{\[0,a\]\}\(x\)\-\\operatorname\{Proj\}\_\{\[0,b\]\}\(x\)\|\\leq\|a\-b\|,\\qquad a,b,x\\geq 0\.
Ifrt≤D¯r\_\{t\}\\leq\\overline\{D\}, the one\-Lipschitz cap and0≤∑iSi,t≤D¯0\\leq\\sum\_\{i\}S\_\{i,t\}\\leq\\overline\{D\}give
\|rt\+1c−rtc\|≤\|rt\+1−rt\|=\|rt−∑iSi,t\|nt−1≤D¯nt−1\.\|r\_\{t\+1\}^\{\\mathrm\{c\}\}\-r\_\{t\}^\{\\mathrm\{c\}\}\|\\leq\|r\_\{t\+1\}\-r\_\{t\}\|=\\frac\{\|r\_\{t\}\-\\sum\_\{i\}S\_\{i,t\}\|\}\{n\_\{t\}\-1\}\\leq\\frac\{\\overline\{D\}\}\{n\_\{t\}\-1\}\.Ifrt\>D¯r\_\{t\}\>\\overline\{D\}, thenrt\+1\>rt\>D¯r\_\{t\+1\}\>r\_\{t\}\>\\overline\{D\}, sort\+1c=rtc=D¯r\_\{t\+1\}^\{\\mathrm\{c\}\}=r\_\{t\}^\{\\mathrm\{c\}\}=\\overline\{D\}\. Thus, in both cases,
\|rt\+1c−rtc\|≤D¯nt−1\.\|r\_\{t\+1\}^\{\\mathrm\{c\}\}\-r\_\{t\}^\{\\mathrm\{c\}\}\|\\leq\\frac\{\\overline\{D\}\}\{n\_\{t\}\-1\}\.BecauseUi\(r\)=X¯i∧rc/piU\_\{i\}\(r\)=\\bar\{X\}\_\{i\}\\wedge r^\{\\mathrm\{c\}\}/p\_\{i\},
\|Ui\(rt\+1\)−Ui\(rt\)\|≤D¯pi\(nt−1\)\.\|U\_\{i\}\(r\_\{t\+1\}\)\-U\_\{i\}\(r\_\{t\}\)\|\\leq\\frac\{\\overline\{D\}\}\{p\_\{i\}\(n\_\{t\}\-1\)\}\.Combining this inequality with \([B\.4](https://arxiv.org/html/2608.14096#A2.E4)\) and\|G^i,t\|≤Gi\|\\widehat\{G\}\_\{i,t\}\|\\leq G\_\{i\}yields
\|Xi,t\+1−Xi,t\|≤Giαt\+D¯pi\(nt−1\)\.\|X\_\{i,t\+1\}\-X\_\{i,t\}\|\\leq G\_\{i\}\\alpha\_\{t\}\+\\frac\{\\overline\{D\}\}\{p\_\{i\}\(n\_\{t\}\-1\)\}\.The two\-sided step\-size profile andnt−1=T−tn\_\{t\}\-1=T\-tgive
∑t=1T−1αt\\displaystyle\\sum\_\{t=1\}^\{T\-1\}\\alpha\_\{t\}≤2χ∑k=1⌈T/2⌉1τ0\+k≤ClogT,\\displaystyle\\leq 2\\chi\\sum\_\{k=1\}^\{\\lceil T/2\\rceil\}\\frac\{1\}\{\\tau\_\{0\}\+k\}\\leq C\\log T,∑t=1T−11nt−1\\displaystyle\\sum\_\{t=1\}^\{T\-1\}\\frac\{1\}\{n\_\{t\}\-1\}=∑k=1T−11k≤ClogT\.\\displaystyle=\\sum\_\{k=1\}^\{T\-1\}\\frac\{1\}\{k\}\\leq C\\log T\.Consequently,
∑t=1T−1\|Xi,t\+1−Xi,t\|≤ClogT\\sum\_\{t=1\}^\{T\-1\}\|X\_\{i,t\+1\}\-X\_\{i,t\}\|\\leq C\\log Tpathwise\.
The water\-filling formula also gives
Ii,t\+1≤Yi,t≤max\{Ii,t,Xi,t\}\.I\_\{i,t\+1\}\\leq Y\_\{i,t\}\\leq\\max\\\{I\_\{i,t\},X\_\{i,t\}\\\}\.SinceIi,1=0I\_\{i,1\}=0and0≤Xi,t≤X¯i0\\leq X\_\{i,t\}\\leq\\bar\{X\}\_\{i\}, induction yields
0≤Ii,t,Yi,t,ei,t≤X¯i\(1≤t≤T\)\.0\\leq I\_\{i,t\},Y\_\{i,t\},e\_\{i,t\}\\leq\\bar\{X\}\_\{i\}\\qquad\(1\\leq t\\leq T\)\.\(B\.5\)
Step 2: Telescope expected drainage\.The water\-filling formula and the physical updateIi,t\+1=\(Yi,t−Di,t\)\+I\_\{i,t\+1\}=\(Y\_\{i,t\}\-D\_\{i,t\}\)^\{\+\}give
\{Yi,t=Ii,t,Ii,t\+1=\(Ii,t−Di,t\)\+,Ii,t\>Xi,t,Yi,t≤Xi,t,Ii,t\+1≤Xi,t,Ii,t≤Xi,t\.\\begin\{cases\}Y\_\{i,t\}=I\_\{i,t\},\\quad I\_\{i,t\+1\}=\(I\_\{i,t\}\-D\_\{i,t\}\)^\{\+\},&I\_\{i,t\}\>X\_\{i,t\},\\\\ Y\_\{i,t\}\\leq X\_\{i,t\},\\quad I\_\{i,t\+1\}\\leq X\_\{i,t\},&I\_\{i,t\}\\leq X\_\{i,t\}\.\\end\{cases\}In the first case,
ei,t\+1≤\(ei,t−Di,t\)\+\+\|Xi,t\+1−Xi,t\|,e\_\{i,t\+1\}\\leq\(e\_\{i,t\}\-D\_\{i,t\}\)^\{\+\}\+\|X\_\{i,t\+1\}\-X\_\{i,t\}\|,whereas in the second caseei,t=0e\_\{i,t\}=0and
ei,t\+1≤\|Xi,t\+1−Xi,t\|\.e\_\{i,t\+1\}\\leq\|X\_\{i,t\+1\}\-X\_\{i,t\}\|\.Hence, in both cases,
ei,t\+1≤\(ei,t−Di,t\)\+\+\|Xi,t\+1−Xi,t\|\.e\_\{i,t\+1\}\\leq\(e\_\{i,t\}\-D\_\{i,t\}\)^\{\+\}\+\|X\_\{i,t\+1\}\-X\_\{i,t\}\|\.Usinga−\(a−d\)\+=a∧da\-\(a\-d\)^\{\+\}=a\\wedge d, this implies
ei,t∧Di,t≤ei,t−ei,t\+1\+\|Xi,t\+1−Xi,t\|\.e\_\{i,t\}\\wedge D\_\{i,t\}\\leq e\_\{i,t\}\-e\_\{i,t\+1\}\+\|X\_\{i,t\+1\}\-X\_\{i,t\}\|\.\(B\.6\)Conditionally onℋt\\mathcal\{H\}\_\{t\},ei,te\_\{i,t\}is fixed andDi,tD\_\{i,t\}has lawFiF\_\{i\}, so
𝔼\[ei,t∧Di,t∣ℋt\]=mi\(ei,t\)\.\\mathbb\{E\}\[e\_\{i,t\}\\wedge D\_\{i,t\}\\mid\\mathcal\{H\}\_\{t\}\]=m\_\{i\}\(e\_\{i,t\}\)\.Taking expectations and summing \([B\.6](https://arxiv.org/html/2608.14096#A2.E6)\) throughT−1T\-1gives
∑t=1T−1𝔼mi\(ei,t\)\\displaystyle\\sum\_\{t=1\}^\{T\-1\}\\mathbb\{E\}m\_\{i\}\(e\_\{i,t\}\)≤𝔼ei,1−𝔼ei,T\+∑t=1T−1𝔼\|Xi,t\+1−Xi,t\|\\displaystyle\\leq\\mathbb\{E\}e\_\{i,1\}\-\\mathbb\{E\}e\_\{i,T\}\+\\sum\_\{t=1\}^\{T\-1\}\\mathbb\{E\}\|X\_\{i,t\+1\}\-X\_\{i,t\}\|≤ClogT,\\displaystyle\\leq C\\log T,whereei,1=0e\_\{i,1\}=0\. Sincemi\(ei,T\)≤ei,T≤X¯im\_\{i\}\(e\_\{i,T\}\)\\leq e\_\{i,T\}\\leq\\bar\{X\}\_\{i\},
∑t=1T𝔼mi\(ei,t\)≤X¯i\+ClogT\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}m\_\{i\}\(e\_\{i,t\}\)\\leq\\bar\{X\}\_\{i\}\+C\\log T\.\(B\.7\)
Step 3: Convert expected drainage into cumulative excess\.Fromfi≤Kif\_\{i\}\\leq K\_\{i\}and1=∫0d¯ifi\(u\)𝑑u1=\\int\_\{0\}^\{\\bar\{d\}\_\{i\}\}f\_\{i\}\(u\)\\,\\mathrm\{d\}u, one hasKi−1≤d¯iK\_\{i\}^\{\-1\}\\leq\\bar\{d\}\_\{i\}\. For every0≤y≤X¯i0\\leq y\\leq\\bar\{X\}\_\{i\},Fi\(u\)≤KiuF\_\{i\}\(u\)\\leq K\_\{i\}uimplies
mi\(y\)≥∫0y∧Ki−1\(1−Kiu\)𝑑u≥\{y/2,0≤y≤Ki−1,1/\(2Ki\),y\>Ki−1\.m\_\{i\}\(y\)\\geq\\int\_\{0\}^\{y\\wedge K\_\{i\}^\{\-1\}\}\(1\-K\_\{i\}u\)\\,\\mathrm\{d\}u\\geq\\begin\{cases\}y/2,&0\\leq y\\leq K\_\{i\}^\{\-1\},\\\\ 1/\(2K\_\{i\}\),&y\>K\_\{i\}^\{\-1\}\.\\end\{cases\}In the first range,y≤2mi\(y\)y\\leq 2m\_\{i\}\(y\)\. In the second,y≤X¯i≤2KiX¯imi\(y\)y\\leq\\bar\{X\}\_\{i\}\\leq 2K\_\{i\}\\bar\{X\}\_\{i\}m\_\{i\}\(y\)\. Therefore, throughout the safe box,
y≤2max\{1,KiX¯i\}mi\(y\)\.y\\leq 2\\max\\\{1,K\_\{i\}\\bar\{X\}\_\{i\}\\\}\\,m\_\{i\}\(y\)\.\(B\.8\)Applying \([B\.8](https://arxiv.org/html/2608.14096#A2.E8)\) toy=ei,ty=e\_\{i,t\}and using \([B\.7](https://arxiv.org/html/2608.14096#A2.E7)\) yields
∑t=1T𝔼ei,t≤ClogT\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}e\_\{i,t\}\\leq C\\log T\.\(B\.9\)
Step 4: Convert cumulative excess into the implementation gap\.SinceXi,t−νt≤Xi,tX\_\{i,t\}\-\\nu\_\{t\}\\leq X\_\{i,t\}, the water\-filling formula gives
Yi,t\>Xi,t⟺Yi,t=Ii,t\>Xi,t\.Y\_\{i,t\}\>X\_\{i,t\}\\quad\\Longleftrightarrow\\quad Y\_\{i,t\}=I\_\{i,t\}\>X\_\{i,t\}\.Thus
\(Yi,t−Xi,t\)\+=ei,t\.\(Y\_\{i,t\}\-X\_\{i,t\}\)^\{\+\}=e\_\{i,t\}\.
The clipping rule givesXi,t≤Ui\(rt\)≤rt/piX\_\{i,t\}\\leq U\_\{i\}\(r\_\{t\}\)\\leq r\_\{t\}/p\_\{i\}\. Hence, whenevernt≥H0n\_\{t\}\\geq H\_\{0\},
∑i=1NXi,t≤rt∑i=1Npi−1≤ntrt=Ct\.\\sum\_\{i=1\}^\{N\}X\_\{i,t\}\\leq r\_\{t\}\\sum\_\{i=1\}^\{N\}p\_\{i\}^\{\-1\}\\leq n\_\{t\}r\_\{t\}=C\_\{t\}\.If water\-filling binds, then∑iYi,t=Ct\\sum\_\{i\}Y\_\{i,t\}=C\_\{t\}, and
∑i\(Yi,t−Xi,t\)\+−∑i\(Xi,t−Yi,t\)\+=Ct−∑iXi,t≥0\.\\sum\_\{i\}\(Y\_\{i,t\}\-X\_\{i,t\}\)^\{\+\}\-\\sum\_\{i\}\(X\_\{i,t\}\-Y\_\{i,t\}\)^\{\+\}=C\_\{t\}\-\\sum\_\{i\}X\_\{i,t\}\\geq 0\.If it does not bind, thenνt=0\\nu\_\{t\}=0andYi,t=max\{Ii,t,Xi,t\}≥Xi,tY\_\{i,t\}=\\max\\\{I\_\{i,t\},X\_\{i,t\}\\\}\\geq X\_\{i,t\}, so the downward discrepancy is zero\. In either case, fornt≥H0n\_\{t\}\\geq H\_\{0\},
‖𝒀t−𝑿t‖1≤2∑i=1N\(Yi,t−Xi,t\)\+=2∑i=1Nei,t\.\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\\leq 2\\sum\_\{i=1\}^\{N\}\(Y\_\{i,t\}\-X\_\{i,t\}\)^\{\+\}=2\\sum\_\{i=1\}^\{N\}e\_\{i,t\}\.\(B\.10\)There are at mostH0−1H\_\{0\}\-1periods withnt<H0n\_\{t\}<H\_\{0\}, and \([B\.5](https://arxiv.org/html/2608.14096#A2.E5)\) gives‖𝒀t−𝑿t‖1≤∑iX¯i\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\\leq\\sum\_\{i\}\\bar\{X\}\_\{i\}in every such period\. Combining this bound with \([B\.9](https://arxiv.org/html/2608.14096#A2.E9)\) and \([B\.10](https://arxiv.org/html/2608.14096#A2.E10)\) gives
∑t=1T𝔼‖𝒀t−𝑿t‖1≤2∑t=1T∑i=1N𝔼ei,t\+\(H0−1\)∑i=1NX¯i≤ClogT\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\\leq 2\\sum\_\{t=1\}^\{T\}\\sum\_\{i=1\}^\{N\}\\mathbb\{E\}e\_\{i,t\}\+\(H\_\{0\}\-1\)\\sum\_\{i=1\}^\{N\}\\bar\{X\}\_\{i\}\\leq C\\log T\.\(B\.11\)Finally,
‖𝒀t−𝑿t‖2≤‖𝒀t−𝑿t‖∞‖𝒀t−𝑿t‖1≤X¯max‖𝒀t−𝑿t‖1\.\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\\leq\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{\\infty\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\\leq\\bar\{X\}\_\{\\max\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\.Summing and applying \([B\.11](https://arxiv.org/html/2608.14096#A2.E11)\) prove the same order for the squared norm\. By the definition \([4\.4](https://arxiv.org/html/2608.14096#S4.E4)\),
𝖨𝖬𝖯T≤ClogT,\\mathsf\{IMP\}\_\{T\}\\leq C\\log T,which proves \([4\.18](https://arxiv.org/html/2608.14096#S4.E18)\)\. ∎
### B\.3Tuning Construction
#### Verification of[Section4](https://arxiv.org/html/2608.14096#S4)
###### Proof\.
Proof\. First compute from the known primitives
pi=hihi\+bi−ci,X¯i=d¯i−pi2Ki,βi=κipi2Ki,p\_\{i\}=\\frac\{h\_\{i\}\}\{h\_\{i\}\+b\_\{i\}\-c\_\{i\}\},\\qquad\\bar\{X\}\_\{i\}=\\bar\{d\}\_\{i\}\-\\frac\{p\_\{i\}\}\{2K\_\{i\}\},\\qquad\\beta\_\{i\}=\\frac\{\\kappa\_\{i\}p\_\{i\}\}\{2K\_\{i\}\},μg:=min1≤i≤Nhiκi,Lg,j:=hjKjβj3,\\mu\_\{g\}:=\\min\_\{1\\leq i\\leq N\}h\_\{i\}\\kappa\_\{i\},\\qquad L\_\{g,j\}:=\\frac\{h\_\{j\}K\_\{j\}\}\{\\beta\_\{j\}^\{3\}\},and choose
θ:=12min\{1,μgβjLg,j2\}\.\\theta:=\\frac\{1\}\{2\}\\min\\left\\\{1,\\frac\{\\mu\_\{g\}\\beta\_\{j\}\}\{L\_\{g,j\}^\{2\}\}\\right\\\}\.This choice satisfies \([A\.21](https://arxiv.org/html/2608.14096#A1.E21)\)\.
The constants entering the motion cutoff can also be fixed explicitly\. Put
βmin:=mini∈\[N\]βi,CW:=12βmin2,𝒥max:=\{i:bi−ci=λmax\},\\beta\_\{\\min\}:=\\min\_\{i\\in\[N\]\}\\beta\_\{i\},\\qquad C\_\{W\}:=\\frac\{1\}\{2\\beta\_\{\\min\}^\{2\}\},\\qquad\\mathcal\{J\}\_\{\\max\}:=\\\{i:b\_\{i\}\-c\_\{i\}=\\lambda\_\{\\max\}\\\},and
Lv:=\(∑j∈𝒥maxpj2\(hj\+bj−cj\)Kj\)−1,Lq:=Lv\(1\+∑i=1N\(hiκi\)−2\)1/2,Csm:=3Lq\.L\_\{v\}:=\\left\(\\sum\_\{j\\in\\mathcal\{J\}\_\{\\max\}\}\\frac\{p\_\{j\}^\{2\}\}\{\(h\_\{j\}\+b\_\{j\}\-c\_\{j\}\)K\_\{j\}\}\\right\)^\{\-1\},\\qquad L\_\{q\}:=L\_\{v\}\\left\(1\+\\sum\_\{i=1\}^\{N\}\(h\_\{i\}\\kappa\_\{i\}\)^\{\-2\}\\right\)^\{1/2\},\\qquad C\_\{\\mathrm\{sm\}\}:=3L\_\{q\}\.The metric calculation in[SectionA\.1](https://arxiv.org/html/2608.14096#A1.SS1)validates the displayed choice ofCWC\_\{W\}\. The proof of[Section4\.2](https://arxiv.org/html/2608.14096#S4.SS2)gives the Lipschitz constantLqL\_\{q\}, and the fixed kernel in \([A\.12](https://arxiv.org/html/2608.14096#A1.E12)\) has total variationVρ=3V\_\{\\rho\}=3, so thisCsmC\_\{\\mathrm\{sm\}\}satisfies[SectionA\.2](https://arxiv.org/html/2608.14096#A1.SS2)\. The field and joint\-smoothness envelopesG,L𝒲G,L\_\{\\mathcal\{W\}\}are likewise obtained from the safe\-box bounds\. They are proof envelopes rather than additional tuning inputs\. All quantities in this paragraph use only the known primitives\. With the same notation as in Section A\.1, take
α0:=min\{1,G−1\},μ0:=12min\{μg,θβj\},μupd:=μ08CW\>0\.\\alpha\_\{0\}:=\\min\\\{1,G^\{\-1\}\\\},\\qquad\\mu\_\{0\}:=\\frac\{1\}\{2\}\\min\\\{\\mu\_\{g\},\\theta\\beta\_\{j\}\\\},\\qquad\\mu\_\{\\mathrm\{upd\}\}:=\\frac\{\\mu\_\{0\}\}\{8C\_\{W\}\}\>0\.Thenα0G≤1\\alpha\_\{0\}G\\leq 1, and the state\-update argument gives the contraction in \([A\.22](https://arxiv.org/html/2608.14096#A1.E22)\) with modulusμupd\\mu\_\{\\mathrm\{upd\}\}\.
For anyχ≥1\\chi\\geq 1andχ/α0≤τ0≤2χ/α0\\chi/\\alpha\_\{0\}\\leq\\tau\_\{0\}\\leq 2\\chi/\\alpha\_\{0\}, the step profile satisfies, whenevern=nt≥2n=n\_\{t\}\\geq 2,
1\(n−1\)αt≤τ0\+nχ\(n−1\)≤2α0\+2\.\\frac\{1\}\{\(n\-1\)\\alpha\_\{t\}\}\\leq\\frac\{\\tau\_\{0\}\+n\}\{\\chi\(n\-1\)\}\\leq\\frac\{2\}\{\\alpha\_\{0\}\}\+2\.TakeCpairC\_\{\\mathrm\{pair\}\}from the explicit formula \([A\.35](https://arxiv.org/html/2608.14096#A1.E35)\)\. It depends only on\(N,𝒉,𝒃,𝒄,𝜿,𝑲\)\(N,\\bm\{h\},\\bm\{b\},\\bm\{c\},\\bm\{\\kappa\},\\bm\{K\}\)\. Choose
μrec\\displaystyle\\mu\_\{\\mathrm\{rec\}\}:=min\{μupd/2,1/2\},χM:=max\{1,32Cpairμupd\},\\displaystyle:=\\min\\\{\\mu\_\{\\mathrm\{upd\}\}/2,1/2\\\},\\qquad\\chi\_\{M\}:=\\max\\left\\\{1,\\frac\{32C\_\{\\mathrm\{pair\}\}\}\{\\mu\_\{\\mathrm\{upd\}\}\}\\right\\\},χ\\displaystyle\\chi:=1\+max\{χM,3μrec\},τ0:=⌈χα0⌉\.\\displaystyle:=1\+\\max\\left\\\{\\chi\_\{M\},\\frac\{3\}\{\\mu\_\{\\mathrm\{rec\}\}\}\\right\\\},\\qquad\\tau\_\{0\}:=\\left\\lceil\\frac\{\\chi\}\{\\alpha\_\{0\}\}\\right\\rceil\.and set
n0:=⌈max\{8,1\+32Cpairμupdα0,1\+2α0\}⌉\.n\_\{0\}:=\\left\\lceil\\max\\left\\\{8,1\+\\frac\{32C\_\{\\mathrm\{pair\}\}\}\{\\mu\_\{\\mathrm\{upd\}\}\\alpha\_\{0\}\},1\+\\frac\{2\}\{\\sqrt\{\\alpha\_\{0\}\}\}\\right\\\}\\right\\rceil\.These choices giveχ≥χM\\chi\\geq\\chi\_\{M\},μrecχ\>3\\mu\_\{\\mathrm\{rec\}\}\\chi\>3,χ/α0≤τ0≤2χ/α0\\chi/\\alpha\_\{0\}\\leq\\tau\_\{0\}\\leq 2\\chi/\\alpha\_\{0\}, andμrecα0≤1/2\\mu\_\{\\mathrm\{rec\}\}\\alpha\_\{0\}\\leq 1/2\.
Substituting these choices into the bounds used to establish \([A\.36](https://arxiv.org/html/2608.14096#A1.E36)\) verifies all three inequalities in that display\. Hence the state\-update, motion, and Taylor remainders in \([A\.26](https://arxiv.org/html/2608.14096#A1.E26)\) have computable primitive\-only constants\. This constructs horizon\-independent tuning without using the demand laws\. ∎
### B\.4Fixed Support Envelopes: An Extension Sketch
This subsection sketches how the proof architecture of[Theorem4\.1](https://arxiv.org/html/2608.14096#S4.Thmtheorem1)can be adapted when exact support endpoints are replaced by fixed certified envelopes\. The purpose is to isolate the additional localization mechanism\.[Theorem4\.1](https://arxiv.org/html/2608.14096#S4.Thmtheorem1)remains the formal guarantee for the direct safe\-box formulation\.
Suppose that the policy is supplied fixed, horizon\-independent envelopes
Diup≥d¯i,Dup:=∑i=1NDiup<∞\.D\_\{i\}^\{\\mathrm\{up\}\}\\geq\\bar\{d\}\_\{i\},\\qquad D^\{\\mathrm\{up\}\}:=\\sum\_\{i=1\}^\{N\}D\_\{i\}^\{\\mathrm\{up\}\}<\\infty\.Under \([2\.3](https://arxiv.org/html/2608.14096#S2.E3)\), the conservative choiceDiup=1/κiD\_\{i\}^\{\\mathrm\{up\}\}=1/\\kappa\_\{i\}is valid because1=∫0d¯ifi≥κid¯i1=\\int\_\{0\}^\{\\bar\{d\}\_\{i\}\}f\_\{i\}\\geq\\kappa\_\{i\}\\bar\{d\}\_\{i\}\. Replace the support\-dependent caps by
rc,up:=min\{r,Dup\},Uiup\(r\):=min\{Diup,rc,uppi\},𝒦up\(r\):=∏i\[0,Uiup\(r\)\]×\[0,λmax\]\.r^\{\\mathrm\{c,up\}\}:=\\min\\\{r,D^\{\\mathrm\{up\}\}\\\},\\qquad U\_\{i\}^\{\\mathrm\{up\}\}\(r\):=\\min\\left\\\{D\_\{i\}^\{\\mathrm\{up\}\},\\frac\{r^\{\\mathrm\{c,up\}\}\}\{p\_\{i\}\}\\right\\\},\\qquad\\mathcal\{K\}^\{\\mathrm\{up\}\}\(r\):=\\prod\_\{i\}\[0,U\_\{i\}^\{\\mathrm\{up\}\}\(r\)\]\\times\[0,\\lambda\_\{\\max\}\]\.The envelope implementation𝖱𝖠𝖯𝖣𝖫up\\mathsf\{RAPDL\}^\{\\mathrm\{up\}\}uses the same recursion as[Algorithm1](https://arxiv.org/html/2608.14096#alg1), with these caps and with
G^0,tup:=rtc,up−∑i=1NSi,t\+θG^j,t\.\\widehat\{G\}\_\{0,t\}^\{\\mathrm\{up\}\}:=r\_\{t\}^\{\\mathrm\{c,up\}\}\-\\sum\_\{i=1\}^\{N\}S\_\{i,t\}\+\\theta\\widehat\{G\}\_\{j,t\}\.LetGθ,upG^\{\\theta,\\mathrm\{up\}\}denote the corresponding conditional\-mean field, obtained fromGθG^\{\\theta\}by replacingrcr^\{\\mathrm\{c\}\}withrc,upr^\{\\mathrm\{c,up\}\}\.
What is inherited from the main proof\.BecauseDup≥∑id¯i≥r0D^\{\\mathrm\{up\}\}\\geq\\sum\_\{i\}\\bar\{d\}\_\{i\}\\geq r\_\{0\}, replacingrcr^\{\\mathrm\{c\}\}byrc,upr^\{\\mathrm\{c,up\}\}affects only resource states at which the fluid solution is already on its slack branch\. The fluid target, the KKT path, and the backward\-smoothed comparator are therefore unchanged\. The projected moving\-comparator inequality in[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3), the comparator\-motion argument in[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\(ii\), the reciprocal\-step summations in Appendix[A\.4](https://arxiv.org/html/2608.14096#A1.SS4), and the Bellman telescope can all be reused\. The implementation\-gap proof also transfers after replacing its support\-dependent inputs by the envelope bounds below\. The genuinely new issue is that a projected threshold may lie beyond the true support, where the original global survival and mirror\-metric bounds are unavailable\.
Implementation\-gap ingredient\.The four changed inputs to the proof of[Section4\.5](https://arxiv.org/html/2608.14096#S4.SS5)are
\|G^0,tup\|\\displaystyle\|\\widehat\{G\}\_\{0,t\}^\{\\mathrm\{up\}\}\|≤Dup\+θGj,\\displaystyle\\leq D^\{\\mathrm\{up\}\}\+\\theta G\_\{j\},0≤Xi,t,Ii,t,Yi,t\\displaystyle 0\\leq X\_\{i,t\},I\_\{i,t\},Y\_\{i,t\}≤Diup,\\displaystyle\\leq D\_\{i\}^\{\\mathrm\{up\}\},\|Uiup\(rt\+1\)−Uiup\(rt\)\|\\displaystyle\|U\_\{i\}^\{\\mathrm\{up\}\}\(r\_\{t\+1\}\)\-U\_\{i\}^\{\\mathrm\{up\}\}\(r\_\{t\}\)\|≤Duppi\(nt−1\),\\displaystyle\\leq\\frac\{D^\{\\mathrm\{up\}\}\}\{p\_\{i\}\(n\_\{t\}\-1\)\},y\\displaystyle y≤2max\{1,KiDiup\}mi\(y\)\.\\displaystyle\\leq 2\\max\\\{1,K\_\{i\}D\_\{i\}^\{\\mathrm\{up\}\}\\\}\\,m\_\{i\}\(y\)\.The cap\-motion case split is the same as before: whenrt\>Dupr\_\{t\}\>D^\{\\mathrm\{up\}\}, both capped rates lie on the constant branch\. Consequently, the preferred\-path variation and excess\-inventory recursion in Appendix[B\.2](https://arxiv.org/html/2608.14096#A2.SS2)give, with an envelope\-dependent constant,
∑t=1T𝔼‖𝒀t−𝑿t‖1\+∑t=1T𝔼‖𝒀t−𝑿t‖2≤CuplogT\.\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|\_\{1\}\+\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\\leq C\_\{\\mathrm\{up\}\}\\log T\.\(B\.12\)
The additional localization mechanism\.For the analysis only, introduce the barriers
xi⋆:=Fi−1\(1−pi\),xi\(1\):=Fi−1\(1−pi/4\),xi\(2\):=Fi−1\(1−pi/8\)\.x\_\{i\}^\{\\star\}:=F\_\{i\}^\{\-1\}\(1\-p\_\{i\}\),\\qquad x\_\{i\}^\{\(1\)\}:=F\_\{i\}^\{\-1\}\(1\-p\_\{i\}/4\),\\qquad x\_\{i\}^\{\(2\)\}:=F\_\{i\}^\{\-1\}\(1\-p\_\{i\}/8\)\.The density upper bound gives
xi\(1\)−xi⋆≥3pi4Ki,xi\(2\)−xi\(1\)≥δi:=pi8Ki,1−Fi\(x\)≥βiloc:=pi8\(x≤xi\(2\)\)\.x\_\{i\}^\{\(1\)\}\-x\_\{i\}^\{\\star\}\\geq\\frac\{3p\_\{i\}\}\{4K\_\{i\}\},\\qquad x\_\{i\}^\{\(2\)\}\-x\_\{i\}^\{\(1\)\}\\geq\\delta\_\{i\}:=\\frac\{p\_\{i\}\}\{8K\_\{i\}\},\\qquad 1\-F\_\{i\}\(x\)\\geq\\beta\_\{i\}^\{\\mathrm\{loc\}\}:=\\frac\{p\_\{i\}\}\{8\}\\quad\(x\\leq x\_\{i\}^\{\(2\)\}\)\.Every exact or backward\-smoothed comparator threshold lies belowxi⋆<xi\(1\)x\_\{i\}^\{\\star\}<x\_\{i\}^\{\(1\)\}\.
LetH¯i\(σ,x\)\\overline\{H\}\_\{i\}\(\\sigma,x\)agree with the sales\-divergence blockℬi\(σ,mi\(x\)\)\\mathcal\{B\}\_\{i\}\(\\sigma,m\_\{i\}\(x\)\)on\[0,xi\(2\)\]\[0,x\_\{i\}^\{\(2\)\}\], and use on\[−1,Diup\+1\]\[\-1,D\_\{i\}^\{\\mathrm\{up\}\}\+1\]the same unit\-curvature quadratic continuation as in \([A\.8](https://arxiv.org/html/2608.14096#A1.E8)\), withX¯i\\bar\{X\}\_\{i\}andmi\(X¯i\)m\_\{i\}\(\\bar\{X\}\_\{i\}\)replaced byxi\(2\)x\_\{i\}^\{\(2\)\}andmi\(xi\(2\)\)m\_\{i\}\(x\_\{i\}^\{\(2\)\}\)\. For constantsAi≥1A\_\{i\}\\geq 1, define
𝒲loc\(\(𝝈,ζ\),\(𝒙,λ\)\):=∑i=1N\[H¯i\(σi,xi\)\+Ai2\(xi−xi\(1\)\)\+2\]\+12\(λ−ζ\)2\.\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(\(\\bm\{\\sigma\},\\zeta\),\(\\bm\{x\},\\lambda\)\):=\\sum\_\{i=1\}^\{N\}\\left\[\\overline\{H\}\_\{i\}\(\\sigma\_\{i\},x\_\{i\}\)\+\\frac\{A\_\{i\}\}\{2\}\(x\_\{i\}\-x\_\{i\}^\{\(1\)\}\)\_\{\+\}^\{2\}\\right\]\+\\frac\{1\}\{2\}\(\\lambda\-\\zeta\)^\{2\}\.Because every comparator threshold is belowxi\(1\)x\_\{i\}^\{\(1\)\}, the base continuation and the hinge are monotone away from the comparator\. Projection onto𝒦up\(r\)\\mathcal\{K\}^\{\\mathrm\{up\}\}\(r\)therefore cannot increase𝒲loc\\mathcal\{W\}^\{\\mathrm\{loc\}\}\. The matched continuation and theC1,1C^\{1,1\}hinge also give a Lipschitz joint gradient on the fixed extended box\. The reused step\-size check, with the stochastic\-direction bound replaced by its fixed\-envelope counterpart, keeps pre\-projection points in that box\. On the interior region𝒢:=\{𝒙:xi≤xi\(2\)∀i\}\\mathcal\{G\}:=\\\{\\bm\{x\}:x\_\{i\}\\leq x\_\{i\}^\{\(2\)\}\\ \\forall i\\\}, the proofs of[SectionA\.1](https://arxiv.org/html/2608.14096#A1.SS1)and[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\(i\) apply with survival lower boundβiloc\\beta\_\{i\}^\{\\mathrm\{loc\}\}\. The hinge handles the complement\. Indeed, ifei:=\(xi−xi\(1\)\)\+\>0e\_\{i\}:=\(x\_\{i\}\-x\_\{i\}^\{\(1\)\}\)\_\{\+\}\>0, then
Giθ\(𝒙,λ,r\)≥3hi4\.G\_\{i\}^\{\\theta\}\(\\bm\{x\},\\lambda;r\)\\geq\\frac\{3h\_\{i\}\}\{4\}\.On𝒢c\\mathcal\{G\}^\{\\mathrm\{c\}\}, at least oneei≥δie\_\{i\}\\geq\\delta\_\{i\}\. The base terms have finite worst\-case envelopes on the fixed box\. First choosec\>0c\>0so thatc\(Diup\+1\)≤3hi/8c\(D\_\{i\}^\{\\mathrm\{up\}\}\+1\)\\leq 3h\_\{i\}/8for everyii, and letB0B\_\{0\}bound the worst\-case deficit of the desired base restoring inequality, including its base\-potential and residual terms\. Then
AieiGiθ−cAiei2≥3hi8Aiei\.A\_\{i\}e\_\{i\}G\_\{i\}^\{\\theta\}\-cA\_\{i\}e\_\{i\}^\{2\}\\geq\\frac\{3h\_\{i\}\}\{8\}A\_\{i\}e\_\{i\}\.Choosing
Ai≥max\{1,8B03hiδi\}A\_\{i\}\\geq\\max\\left\\\{1,\\frac\{8B\_\{0\}\}\{3h\_\{i\}\\delta\_\{i\}\}\\right\\\}makes the inward hinge drift pay for that exterior loss\. In particular,
𝟏\{𝒢c\}≤∑i=1Nei2δi2≤C𝒲loc\.\\mathbf\{1\}\\\{\\mathcal\{G\}^\{\\mathrm\{c\}\}\\\}\\leq\\sum\_\{i=1\}^\{N\}\\frac\{e\_\{i\}^\{2\}\}\{\\delta\_\{i\}^\{2\}\}\\leq C\\mathcal\{W\}^\{\\mathrm\{loc\}\}\.
This gives localized analogues of the three bounds used in the main proof: metric and gradient control, a restoring inequality, and an operational\-gap charge\. Define
ℛup\(𝒙,λ,r\)\\displaystyle\\mathcal\{R\}^\{\\mathrm\{up\}\}\(\\bm\{x\},\\lambda;r\):=∑imi\(xi\)\(gi′\(si∗\(r\)\)\+λ∗\(r\)\)\+λ\(rc,up−r0\)\+,\\displaystyle:=\\sum\_\{i\}m\_\{i\}\(x\_\{i\}\)\\bigl\(g\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\*\}\(r\)\)\+\\lambda^\{\*\}\(r\)\\bigr\)\+\\lambda\(r^\{\\mathrm\{c,up\}\}\-r\_\{0\}\)^\{\+\},Γ\(𝒚,r\)\\displaystyle\\Gamma\(\\bm\{y\};r\):=∑iℓi\(yi\)−v\(r\)\+λ∗\(r\)\(∑imi\(yi\)−r\)\.\\displaystyle:=\\sum\_\{i\}\\ell\_\{i\}\(y\_\{i\}\)\-v\(r\)\+\\lambda^\{\*\}\(r\)\\left\(\\sum\_\{i\}m\_\{i\}\(y\_\{i\}\)\-r\\right\)\.Forq=\(𝝈,ζ\)q=\(\\bm\{\\sigma\},\\zeta\)with0≤σi≤mi\(xi⋆\)0\\leq\\sigma\_\{i\}\\leq m\_\{i\}\(x\_\{i\}^\{\\star\}\)and0≤ζ≤λmax0\\leq\\zeta\\leq\\lambda\_\{\\max\}, projectedz∈𝒦up\(r\)z\\in\\mathcal\{K\}^\{\\mathrm\{up\}\}\(r\), and0≤yi≤Diup0\\leq y\_\{i\}\\leq D\_\{i\}^\{\\mathrm\{up\}\}, the required schematic estimates are
‖𝒎\(𝒙\)−𝝈‖2\+\|λ−ζ\|2\+𝟏\{𝒢c\}\\displaystyle\\\|\\bm\{m\}\(\\bm\{x\}\)\-\\bm\{\\sigma\}\\\|^\{2\}\+\|\\lambda\-\\zeta\|^\{2\}\+\\mathbf\{1\}\\\{\\mathcal\{G\}^\{\\mathrm\{c\}\}\\\}≤C𝒲loc\(q,z\),\\displaystyle\\leq C\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(q,z\),‖∇q𝒲loc\(q,z\)‖2\+‖∇z𝒲loc\(q,z\)‖2\\displaystyle\\\|\\nabla\_\{q\}\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(q,z\)\\\|^\{2\}\+\\\|\\nabla\_\{z\}\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(q,z\)\\\|^\{2\}≤C𝒲loc\(q,z\),\\displaystyle\\leq C\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(q,z\),⟨∇z𝒲loc\(qη\(r\),z\),Gθ,up\(𝒙,λ,r\)⟩\\displaystyle\\langle\\nabla\_\{z\}\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(q^\{\\eta\}\(r\),z\),G^\{\\theta,\\mathrm\{up\}\}\(\\bm\{x\},\\lambda;r\)\\rangle≥c𝒲loc\(qη\(r\),z\)\+cℛup\(𝒙,λ,r\)−Cη2,\\displaystyle\\geq c\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(q^\{\\eta\}\(r\),z\)\+c\\mathcal\{R\}^\{\\mathrm\{up\}\}\(\\bm\{x\},\\lambda;r\)\-C\\eta^\{2\},0≤Γ\(𝒚,r\)\\displaystyle 0\\leq\\Gamma\(\\bm\{y\};r\)≤C\[𝒲loc\(q∗\(r\),z\)\+ℛup\(𝒙,λ,r\)\+‖𝒚−𝒙‖1\]\.\\displaystyle\\leq C\\left\[\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(q^\{\*\}\(r\),z\)\+\\mathcal\{R\}^\{\\mathrm\{up\}\}\(\\bm\{x\},\\lambda;r\)\+\\\|\\bm\{y\}\-\\bm\{x\}\\\|\_\{1\}\\right\]\.The exact\-to\-smoothed transfer still uses the main mechanism\. On𝒢\\mathcal\{G\}, the active\-coordinate and slack\-boundary\-layer calculation in[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\(i\) is unchanged\. On𝒢c\\mathcal\{G\}^\{\\mathrm\{c\}\}, joint\-gradient Lipschitzness gives anO\(η\)O\(\\eta\)perturbation, while𝟏\{𝒢c\}≤C𝒲loc\\mathbf\{1\}\\\{\\mathcal\{G\}^\{\\mathrm\{c\}\}\\\}\\leq C\\mathcal\{W\}^\{\\mathrm\{loc\}\}converts it toO\(η𝒲loc\)O\(\\eta\\sqrt\{\\mathcal\{W\}^\{\\mathrm\{loc\}\}\}\)\. Young’s inequality then gives the same contraction loss and anO\(η2\)O\(\\eta^\{2\}\)remainder\.
For the operational charge, apply the main bound first at𝒙\\bm\{x\}\. Outside the true support, the additional holding term is
ℓi\(y\)=gi\(mi\(y\)\)\+hi\(y−d¯i\)\+\.\\ell\_\{i\}\(y\)=g\_\{i\}\(m\_\{i\}\(y\)\)\+h\_\{i\}\(y\-\\bar\{d\}\_\{i\}\)^\{\+\}\.At𝒙\\bm\{x\}, this exterior case is charged by the hinge\. Sinceℓi\\ell\_\{i\}andmim\_\{i\}are Lipschitz on the fixed envelope interval, replacing𝒙\\bm\{x\}by the implemented𝒚\\bm\{y\}costs at mostC‖𝒚−𝒙‖1C\\\|\\bm\{y\}\-\\bm\{x\}\\\|\_\{1\}\.
Reusing the proof architecture\.Setηt2=αt\\eta\_\{t\}^\{2\}=\\alpha\_\{t\},qt=qηt\(rt\)q\_\{t\}=q^\{\\eta\_\{t\}\}\(r\_\{t\}\), andW~tloc:=𝒲loc\(qt,zt\)\\widetilde\{W\}\_\{t\}^\{\\mathrm\{loc\}\}:=\\mathcal\{W\}^\{\\mathrm\{loc\}\}\(q\_\{t\},z\_\{t\}\)\. Combining[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3), the localized ingredients above, and the comparator\-motion proof of[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\(ii\) gives the same type of recursion\. In that case split,DupD^\{\\mathrm\{up\}\}replacesD¯\\overline\{D\}: below the cap the resource increment is bounded, while above it the comparator is already constant on the slack branch\. Thus
𝔼\[W~t\+1loc∣ℋt\]\\displaystyle\\mathbb\{E\}\[\\widetilde\{W\}\_\{t\+1\}^\{\\mathrm\{loc\}\}\\mid\\mathcal\{H\}\_\{t\}\]≤\(1−cαt\)W~tloc−cαtℛup\(𝑿t,λt,rt\)\\displaystyle\\leq\(1\-c\\alpha\_\{t\}\)\\widetilde\{W\}\_\{t\}^\{\\mathrm\{loc\}\}\-c\\alpha\_\{t\}\\mathcal\{R\}^\{\\mathrm\{up\}\}\(\\bm\{X\}\_\{t\},\\lambda\_\{t\};r\_\{t\}\)\+Cαt2\+Cαt‖𝒀t−𝑿t‖2\.\\displaystyle\\quad\+C\\alpha\_\{t\}^\{2\}\+C\\alpha\_\{t\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.The only altered Taylor estimate is
CW~tlocηt\(nt−1\)2≤c8αtW~tloc\+Cαt2,\\frac\{C\\sqrt\{\\widetilde\{W\}\_\{t\}^\{\\mathrm\{loc\}\}\}\}\{\\eta\_\{t\}\(n\_\{t\}\-1\)^\{2\}\}\\leq\\frac\{c\}\{8\}\\alpha\_\{t\}\\widetilde\{W\}\_\{t\}^\{\\mathrm\{loc\}\}\+C\\alpha\_\{t\}^\{2\},which follows fromηt2=αt\\eta\_\{t\}^\{2\}=\\alpha\_\{t\}and the same large\-ntn\_\{t\}bounds as in Appendix A\. The reciprocal\-step summations, together with \([B\.12](https://arxiv.org/html/2608.14096#A2.E12)\), then control the localized potential and residual at logarithmic order\. Finally, the Bellman telescope usesDupD^\{\\mathrm\{up\}\}in place ofD¯\\overline\{D\}\. Whenrt\>Dupr\_\{t\}\>D^\{\\mathrm\{up\}\}, both resource rates lie on the constant slack branch ofvv, so the Taylor remainder is zero\.
These calculations identify the additional ingredients needed for a fixed\-envelope implementation and show how they enter the already established proof architecture\. We record the route as an extension sketch rather than as a separate theorem, and formal proof could be constructed from the sketch\.
### B\.5Two\-Rate RAPDL: An Extension Sketch
This subsection records the transfer argument for using different primal and dual step scales\. It isolates the weighted\-potential identity that allows the common\-rate proof to be reused, rather than repeating that proof\.
Fix\\varkappa\>0\\varkappa\>0, treated as constant as the horizon varies, and write
P\\varkappa:=diag\(IN,\\varkappa\),at:=χxτ0\+min\{t,nt\}\.P\_\{\\varkappa\}:=\\operatorname\{diag\}\(I\_\{N\},\\varkappa\),\\qquad a\_\{t\}:=\\frac\{\\chi\_\{x\}\}\{\\tau\_\{0\}\+\\min\\\{t,n\_\{t\}\\\}\}\.The two\-rate update is
zt\+1:=Proj𝒦\(rt\+1\)\(zt−atP\\varkappaG^t\),t<T\.z\_\{t\+1\}:=\\operatorname\{Proj\}\_\{\\mathcal\{K\}\(r\_\{t\+1\}\)\}\\bigl\(z\_\{t\}\-a\_\{t\}P\_\{\\varkappa\}\\widehat\{G\}\_\{t\}\\bigr\),\\qquad t<T\.Equivalently,αtx=at\\alpha\_\{t\}^\{x\}=a\_\{t\},αtλ=\\varkappaat\\alpha\_\{t\}^\{\\lambda\}=\\varkappa a\_\{t\}, andχλ=\\varkappaχx\\chi\_\{\\lambda\}=\\varkappa\\chi\_\{x\}\. Inventory decisions, observations, resource updates, and clipping sets remain those of[Algorithm1](https://arxiv.org/html/2608.14096#alg1)\.
The weighted\-potential cancellation\.For the primal blocksHiH\_\{i\}in \([A\.8](https://arxiv.org/html/2608.14096#A1.E8)\), define
𝒲\\varkappa\(q,z\):=∑i=1NHi\(σi,xi\)\+12\\varkappa\(λ−ζ\)2\.\\mathcal\{W\}\_\{\\varkappa\}\(q,z\):=\\sum\_\{i=1\}^\{N\}H\_\{i\}\(\\sigma\_\{i\},x\_\{i\}\)\+\\frac\{1\}\{2\\varkappa\}\(\\lambda\-\\zeta\)^\{2\}\.It is equivalent to the common\-rate potential:
min\{1,\\varkappa−1\}𝒲\(q,z\)≤𝒲\\varkappa\(q,z\)≤max\{1,\\varkappa−1\}𝒲\(q,z\)\.\\min\\\{1,\\varkappa^\{\-1\}\\\}\\mathcal\{W\}\(q,z\)\\leq\\mathcal\{W\}\_\{\\varkappa\}\(q,z\)\\leq\\max\\\{1,\\varkappa^\{\-1\}\\\}\\mathcal\{W\}\(q,z\)\.\(B\.13\)For any directiong=\(g1,…,gN,g0\)g=\(g\_\{1\},\\ldots,g\_\{N\},g\_\{0\}\),
⟨∇z𝒲\\varkappa\(q,z\),P\\varkappag⟩=∑i=1N∂xiHi\(σi,xi\)gi\+\(λ−ζ\)g0\.\\left\\langle\\nabla\_\{z\}\\mathcal\{W\}\_\{\\varkappa\}\(q,z\),P\_\{\\varkappa\}g\\right\\rangle=\\sum\_\{i=1\}^\{N\}\\partial\_\{x\_\{i\}\}H\_\{i\}\(\\sigma\_\{i\},x\_\{i\}\)g\_\{i\}\+\(\\lambda\-\\zeta\)g\_\{0\}\.\(B\.14\)Thus the factor\\varkappa\\varkappain the dual step is canceled exactly by the dual weight1/\\varkappa1/\\varkappa\. The first\-order Primal\-Dual pairing in[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3)\(i\) is unchanged\.
What changes in the reusable bounds\.The metric and gradient constants become\\varkappa\\varkappa\-dependent through \([B\.13](https://arxiv.org/html/2608.14096#A2.E13)\)\. Projection compatibility is unchanged: the primal projection is the same, and multiplying the dual square by the positive constant1/\\varkappa1/\\varkappapreserves interval\-projection monotonicity\. The stochastic\-direction envelope becomes
G¯0:=D¯\+θGj,G\\varkappa:=\(∑i=1NGi2\+\\varkappa2G¯02\)1/2,‖P\\varkappaG^t‖≤G\\varkappa,\\overline\{G\}\_\{0\}:=\\overline\{D\}\+\\theta G\_\{j\},\\qquad G\_\{\\varkappa\}:=\\left\(\\sum\_\{i=1\}^\{N\}G\_\{i\}^\{2\}\+\\varkappa^\{2\}\\overline\{G\}\_\{0\}^\{2\}\\right\)^\{1/2\},\\qquad\\\|P\_\{\\varkappa\}\\widehat\{G\}\_\{t\}\\\|\\leq G\_\{\\varkappa\},so the mirror quadratic term remainsO\\varkappa\(at2\)O\_\{\\varkappa\}\(a\_\{t\}^\{2\}\)\.
On the comparator side, the smoothing lemma is unchanged\. Weighted gradient control and \([B\.14](https://arxiv.org/html/2608.14096#A2.E14)\) transfer the two parts of[SectionA\.3](https://arxiv.org/html/2608.14096#A1.SS3), with constantsCW,\\varkappaC\_\{W,\\varkappa\},L𝒲,\\varkappaL\_\{\\mathcal\{W\},\\varkappa\}, andCpair,\\varkappaC\_\{\\mathrm\{pair\},\\varkappa\}\. The last constant is obtained from \([A\.35](https://arxiv.org/html/2608.14096#A1.E35)\) by replacing the comparator\-gradient envelope with its weighted counterpart\. Set
a0:=min\{1,G\\varkappa−1\},μupd,\\varkappa:=μ08CW,\\varkappa,μ\\varkappa:=min\{μupd,\\varkappa/2,1/2\}\.a\_\{0\}:=\\min\\\{1,G\_\{\\varkappa\}^\{\-1\}\\\},\\qquad\\mu\_\{\\mathrm\{upd\},\\varkappa\}:=\\frac\{\\mu\_\{0\}\}\{8C\_\{W,\\varkappa\}\},\\qquad\\mu\_\{\\varkappa\}:=\\min\\\{\\mu\_\{\\mathrm\{upd\},\\varkappa\}/2,1/2\\\}\.Under the reused initialization conditionτ0≥χx/a0\\tau\_\{0\}\\geq\\chi\_\{x\}/a\_\{0\}, one hasat≤a0a\_\{t\}\\leq a\_\{0\}andatG\\varkappa≤1a\_\{t\}G\_\{\\varkappa\}\\leq 1\. The tuning construction is then reused under the substitution
\(G,CW,L𝒲,μupd,Cpair,α0,χ\)↦\(G\\varkappa,CW,\\varkappa,L𝒲,\\varkappa,μupd,\\varkappa,Cpair,\\varkappa,a0,χx\)\.\(G,C\_\{W\},L\_\{\\mathcal\{W\}\},\\mu\_\{\\mathrm\{upd\}\},C\_\{\\mathrm\{pair\}\},\\alpha\_\{0\},\\chi\)\\mapsto\(G\_\{\\varkappa\},C\_\{W,\\varkappa\},L\_\{\\mathcal\{W\},\\varkappa\},\\mu\_\{\\mathrm\{upd\},\\varkappa\},C\_\{\\mathrm\{pair\},\\varkappa\},a\_\{0\},\\chi\_\{x\}\)\.
Resulting recursion and summation\.Setηt2=at\\eta\_\{t\}^\{2\}=a\_\{t\},qt=qηt\(rt\)q\_\{t\}=q^\{\\eta\_\{t\}\}\(r\_\{t\}\), andW~\\varkappa,t:=𝒲\\varkappa\(qt,zt\)\\widetilde\{W\}\_\{\\varkappa,t\}:=\\mathcal\{W\}\_\{\\varkappa\}\(q\_\{t\},z\_\{t\}\)\. Withℛt\\mathcal\{R\}\_\{t\}denoting the same nonnegative residual as in \([A\.26](https://arxiv.org/html/2608.14096#A1.E26)\), the weighted moving\-comparator calculation has the form
𝔼\[W~\\varkappa,t\+1∣ℋt\]\\displaystyle\\mathbb\{E\}\[\\widetilde\{W\}\_\{\\varkappa,t\+1\}\\mid\\mathcal\{H\}\_\{t\}\]≤\(1−μ\\varkappaat\)W~\\varkappa,t−μ\\varkappaatℛt\\displaystyle\\leq\(1\-\\mu\_\{\\varkappa\}a\_\{t\}\)\\widetilde\{W\}\_\{\\varkappa,t\}\-\\mu\_\{\\varkappa\}a\_\{t\}\\mathcal\{R\}\_\{t\}\+C\\varkappaat2\+C\\varkappaat‖𝒀t−𝑿t‖2\.\\displaystyle\\quad\+C\_\{\\varkappa\}a\_\{t\}^\{2\}\+C\_\{\\varkappa\}a\_\{t\}\\\|\\bm\{Y\}\_\{t\}\-\\bm\{X\}\_\{t\}\\\|^\{2\}\.Appendix[A\.4](https://arxiv.org/html/2608.14096#A1.SS4)applies withata\_\{t\}in place ofαt\\alpha\_\{t\}\. The implementation argument also changes only through the primal variation bound
\|Xi,t\+1−Xi,t\|≤Giat\+D¯pi\(nt−1\),\|X\_\{i,t\+1\}\-X\_\{i,t\}\|\\leq G\_\{i\}a\_\{t\}\+\\frac\{\\overline\{D\}\}\{p\_\{i\}\(n\_\{t\}\-1\)\},because the dual rate does not enter the physical excess\-inventory recursion\. The exact\-to\-smoothed comparison, implementation\-gap bound, and re\-solving reduction then follow the common\-rate proof with constants allowed to depend on\\varkappa\\varkappa\.
These observations indicate how the common\-rate analysis extends to any fixed\\varkappa\>0\\varkappa\>0\. We record the transfer argument as a sketch rather than repeat the proof of[Theorem4\.1](https://arxiv.org/html/2608.14096#S4.Thmtheorem1)\.Similar Articles
Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing
This paper formulates a multi-item capacitated lot-sizing problem with stochastic demand timing as a discrete-time MDP and proposes a genetic algorithm to solve it, demonstrating efficiency on benchmark instances.
A Comparative Study of Bayesian Contextual Bandits for Real-Time Warehouse Sorter Optimization
This paper presents a comparative study of Bayesian Contextual Bandits, XGBoost, and Linear Regression for real-time sorter diversion optimization in e-commerce warehouses, showing BCB achieves 2.03% reward uplift with superior online learning and inference latency.
Optimizing ARDL Models for Retail Sales Forecasting and Fair Pricing
This paper proposes a fairness-aware pricing framework for retail food products using Autoregressive Distributed Lag (ARDL) models for sales forecasting and optimizes prices with Linear Programming and Simulated Annealing under CPI-based bounds to prevent consumer exploitation.
Offline Reinforcement Learning for Warehouse SLAM Throughput Control
This paper presents an offline reinforcement learning framework for optimizing SLAM throughput control in warehouse fulfillment environments, balancing throughput maximization with downstream stability. The approach is algorithm-agnostic and demonstrates that the CQL policy improves system health by 22.97% and reduces throttling duration by 3.18%.
A3M: Adaptive, Adversarial and Multi-Objective Learning for Strategic Bidding in Repeated Auctions
Introduces A3M, a framework combining adaptive deep reinforcement learning, adversarial reasoning, and multi-objective reward design for strategic bidding in repeated auctions, achieving 30-40% regret reduction.