Scaling Limits of Constant-Stepsize SGD at Flat Minima
Summary
This paper analyzes the scaling limits of constant-stepsize SGD near flat minima, showing that the invariant law concentrates at scale α^(1/m) for objectives with flatness exponent m ≥ 2, and converges to non-Gaussian stationary distributions for m > 2.
View Cached Full Text
Cached at: 07/21/26, 06:48 AM
# Scaling Limits of Constant-Stepsize SGD at Flat Minima
Source: [https://arxiv.org/html/2607.16384](https://arxiv.org/html/2607.16384)
###### Abstract
For stochastic gradient descent \(SGD\) with a constant stepsizeα\\alpha, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons\. In the strongly convex case, this invariant law has the familiarα\\sqrt\{\\alpha\}scaling and a Gaussian limit asα↓0\\alpha\\downarrow 0\. We show that this behavior changes fundamentally for convex objectivesHHwith flat minima and \(sub\)quadratic tails\.
More specifically, we study SGD with Markovian noise generated by a contractive driving chain\. For every sufficiently small constant stepsizeα\\alpha, we prove existence, uniqueness, and geometric convergence to an augmented invariant law in a Wasserstein distance induced by anα\\alpha\-dependent metric\. When the minimizerx⋆x\_\{\\star\}has local flatness exponentm≥2m\\geq 2, meaning that∇2H\(x\)≍‖x−x⋆‖m−2Id\\nabla^\{2\}H\(x\)\\asymp\\left\\lVert x\-x\_\{\\star\}\\right\\rVert^\{m\-2\}I\_\{d\}asx→x⋆x\\to x\_\{\\star\}, we obtain a contraction bound with factor1−cαm−11\-c\\alpha^\{m\-1\}, wherec\>0c\>0is a constant\. This recovers the factor1−cα1\-c\\alphain the quadratic casem=2m=2\. We then analyze the small\-stepsize scaling limit\. We show that the invariant law concentrates on the scaleα1/m\\alpha^\{1/m\}and that the rescaled iterates converge weakly to the stationary distribution of the stochastic differential equation
dYt=−h0\(Yt\)dt\+Σ1/2dBt,dY\_\{t\}=\-h\_\{0\}\(Y\_\{t\}\)\\,dt\+\\Sigma^\{1/2\}\\,dB\_\{t\},whereh0h\_\{0\}is the limiting drift at the minimizer andΣ\\Sigmadenotes the asymptotic covariance\. This recovers the Gaussian limit whenm=2m=2and gives generally non\-Gaussian stationary limits in the flat casem\>2m\>2\. Finally, we give corresponding results for coordinate\-separable objectives with unequal flatness exponents\.
††1Georgia Institute of Technology,*Email:*[jzhang3450@gatech\.edu](https://arxiv.org/html/2607.16384v1/mailto:[email protected])††2Georgia Institute of Technology,*Email:*[cheng\.mao@math\.gatech\.edu](https://arxiv.org/html/2607.16384v1/mailto:[email protected])††3Georgia Institute of Technology,*Email:*[debankur\.mukherjee@isye\.gatech\.edu](https://arxiv.org/html/2607.16384v1/mailto:[email protected])###### Contents
1. [1Introduction](https://arxiv.org/html/2607.16384#S1)
2. [2Main results](https://arxiv.org/html/2607.16384#S2)1. [2\.1Geometric ergodicity in Wasserstein distance](https://arxiv.org/html/2607.16384#S2.SS1) 2. [2\.2Small\-stepsize scaling limit](https://arxiv.org/html/2607.16384#S2.SS2) 3. [2\.3Coordinate\-separable objectives](https://arxiv.org/html/2607.16384#S2.SS3)
3. [3Applications and numerical experiments](https://arxiv.org/html/2607.16384#S3)1. [3\.1Local flatness in quantile and tail\-risk estimation](https://arxiv.org/html/2607.16384#S3.SS1) 2. [3\.2Subquadratic tails for robust and logistic losses](https://arxiv.org/html/2607.16384#S3.SS2)
4. [4Proofs of the main results](https://arxiv.org/html/2607.16384#S4)1. [4\.1Analysis of constant\-stepsize SGD](https://arxiv.org/html/2607.16384#S4.SS1) 2. [4\.2Analysis of scaling limit](https://arxiv.org/html/2607.16384#S4.SS2)
5. [5Conclusion](https://arxiv.org/html/2607.16384#S5)
6. [AConstant\-stepsize contraction](https://arxiv.org/html/2607.16384#A1)1. [A\.1Fixed\-point principle](https://arxiv.org/html/2607.16384#A1.SS1) 2. [A\.2Proof of Lemma4\.2](https://arxiv.org/html/2607.16384#A1.SS2) 3. [A\.3Preparatory estimates](https://arxiv.org/html/2607.16384#A1.SS3) 4. [A\.4Proof of Lemma4\.3](https://arxiv.org/html/2607.16384#A1.SS4) 5. [A\.5Proof of Lemma4\.4](https://arxiv.org/html/2607.16384#A1.SS5) 6. [A\.6Proof of Lemma4\.5](https://arxiv.org/html/2607.16384#A1.SS6) 7. [A\.7Proof of Corollary2\.9](https://arxiv.org/html/2607.16384#A1.SS7)
7. [BScaling\-limit estimates](https://arxiv.org/html/2607.16384#A2)1. [B\.1Deterministic scaling estimates](https://arxiv.org/html/2607.16384#A2.SS1) 2. [B\.2Driving\-chain ergodicity and Poisson equation](https://arxiv.org/html/2607.16384#A2.SS2) 3. [B\.3Proof of Lemma4\.6](https://arxiv.org/html/2607.16384#A2.SS3) 4. [B\.4Proof of Lemma4\.7](https://arxiv.org/html/2607.16384#A2.SS4) 5. [B\.5Proof of Lemma4\.8](https://arxiv.org/html/2607.16384#A2.SS5) 6. [B\.6Proof of Lemma4\.9](https://arxiv.org/html/2607.16384#A2.SS6) 7. [B\.7Proof of Lemma4\.10](https://arxiv.org/html/2607.16384#A2.SS7)
8. [CSeparable extensions](https://arxiv.org/html/2607.16384#A3)1. [C\.1Proof of Corollary2\.15](https://arxiv.org/html/2607.16384#A3.SS1) 2. [C\.2Proof of Corollary2\.17](https://arxiv.org/html/2607.16384#A3.SS2) 3. [C\.3Proof of Remark2\.18](https://arxiv.org/html/2607.16384#A3.SS3)
9. [References](https://arxiv.org/html/2607.16384#bib)
## 1Introduction
Stochastic approximation originated in recursive methods for root\-finding and optimization based on noisy observations\[[40](https://arxiv.org/html/2607.16384#bib.bib40),[24](https://arxiv.org/html/2607.16384#bib.bib24),[7](https://arxiv.org/html/2607.16384#bib.bib7),[27](https://arxiv.org/html/2607.16384#bib.bib27),[9](https://arxiv.org/html/2607.16384#bib.bib9)\]\. Classical theory uses decreasing stepsizes and studies almost\-sure convergence to a target point\. In large\-scale optimization, by contrast, constant or piecewise\-constant stepsizes are often used for long time intervals\. With a constant stepsize and nondegenerate gradient noise, the last stochastic iterate generally does not converge to the minimizer\. It is instead natural to study the invariant law of the Markov chain generated by the algorithm, which describes the stationary error of the last iterate\.
For smooth strongly convex objectives, the stationary error of stochastic gradient descent \(SGD\) is well understood\. It is typically of orderα\\sqrt\{\\alpha\}, and after theα−1/2\\alpha^\{\-1/2\}\-rescaling the invariant law converges to the stationary distribution of an Ornstein–Uhlenbeck diffusion, hence Gaussian\. This picture appears in small\-stepsize asymptotic laws\[[37](https://arxiv.org/html/2607.16384#bib.bib37)\], in diffusion approximations for constant\-step SGD\[[29](https://arxiv.org/html/2607.16384#bib.bib29)\], and in stationary small\-stepsize characterizations for SGD\-type algorithms\[[12](https://arxiv.org/html/2607.16384#bib.bib12)\]\. Complementary Markov chain analyses establish invariant measure expansions, convergence, and concentration in strongly convex settings\[[13](https://arxiv.org/html/2607.16384#bib.bib13),[30](https://arxiv.org/html/2607.16384#bib.bib30)\]\. However, the quadratic case is not representative of all convex objectives\. When the deterministic drift vanishes to order higher than one at the minimizer, the normalization of the invariant law and its scaling limit both change\.
This paper studies convex objectives whose minimizer may be flatter than quadratic\. Near the minimizer, the deterministic drift may satisfy
∇2H\(x\)≍‖x‖m−2Id,m≥2,\\nabla^\{2\}H\(x\)\\asymp\\left\\lVert x\\right\\rVert^\{m\-2\}I\_\{d\},\\qquad m\\geq 2,in dimensiondd, so the restoring force is of order‖x‖m−1\\left\\lVert x\\right\\rVert^\{m\-1\}\. At the same time, the objective may have subquadratic tails: for1≤β<21\\leq\\beta<2, the drift for large‖x‖\\left\\lVert x\\right\\rVertmay grow only as‖x‖β−1\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\. These features occur in standard statistical objectives\. Quantile estimation when the distribution function crosses the target probability level at higher order provides a canonical example of local flatness\. For median andL1L\_\{1\}regression, this type of nonregular local behavior was studied by Knight\[[25](https://arxiv.org/html/2607.16384#bib.bib25)\]; their regularly varying mechanism applies to general quantile loss introduced by Koenker and Bassett\[[26](https://arxiv.org/html/2607.16384#bib.bib26)\]\. The same local geometry appears in the Rockafellar–Uryasev variational representation of conditional value\-at\-risk \(CVaR\)\[[41](https://arxiv.org/html/2607.16384#bib.bib41)\]\. Subquadratic tails arise in robust losses with bounded or sublinear scores, such as Huber\-type and generalized Charbonnier losses\[[19](https://arxiv.org/html/2607.16384#bib.bib19),[11](https://arxiv.org/html/2607.16384#bib.bib11),[4](https://arxiv.org/html/2607.16384#bib.bib4)\], and in logistic losses\[[5](https://arxiv.org/html/2607.16384#bib.bib5)\]\. Recent work on subquadratic SGD treats the locally strongly convex casem=2m=2withβ<2\\beta<2\[[47](https://arxiv.org/html/2607.16384#bib.bib47)\]; the casem\>2m\>2requires a different local scale and a different contraction argument\.
We emphasize the different roles played by the two key exponents used throughout the paper: The local flatness exponentmmdetermines the order of the restoring drift near the minimizer and therefore the stationary scaling\. The tail exponentβ\\betadetermines the drift for large‖x‖\\left\\lVert x\\right\\rVertand the weight function needed to control excursions\.
Moreover, we consider SGD with Markovian noise\. A driving Markov chain\(ξn\)n≥0\(\\xi\_\{n\}\)\_\{n\\geq 0\}evolves according to a contractive recursion, and the SGD iterate follows
Xn\+1=Xn−α\{h\(Xn\)\+g\(Xn,ξn\+1\)\},h=∇H\.X\_\{n\+1\}=X\_\{n\}\-\\alpha\\\{h\(X\_\{n\}\)\+g\(X\_\{n\},\\xi\_\{n\+1\}\)\\\},\\qquad h=\\nabla H\.The augmented process\(Xn,ξn\)\(X\_\{n\},\\xi\_\{n\}\)is Markov\. This formulation covers not only statistical settings with i\.i\.d\. noise but also temporally dependent data streams\.
#### Challenges and main contributions\.
There are three difficulties for analyzing SGD in the above setting\. First, ordinary Euclidean contraction degenerates near the minimizer whenm\>2m\>2, which is also the region where the invariant laws concentrate asα↓0\\alpha\\downarrow 0\. Second, forβ<2\\beta<2, quadratic Lyapunov functions are not well matched to the drift for large‖x‖\\left\\lVert x\\right\\rVert\. Third, under Markovian noise, the current iterate is correlated with future values of the driving chain, which prevents a direct one\-step martingale argument\. Our main results below address these challenges\.
First, for the convergence of SGD with a small fixed stepsizeα\\alpha, we prove that the SGD iterates converge to a unique invariant law geometrically in a suitably chosen Wasserstein distance, with contraction factor1−cαm−11\-c\\alpha^\{m\-1\}for a constantc\>0c\>0\. The proof uses the metric of Qu, Blanchet, and Glynn\[[39](https://arxiv.org/html/2607.16384#bib.bib39)\]induced by a Lyapunov weight functionVαV\_\{\\alpha\}that we carefully construct\. The weight function consists of two terms, one for controlling the subquadratic tail of the objective and the other for controlling the near\-minimizer region\. Moreover, directly applying the result of\[[39](https://arxiv.org/html/2607.16384#bib.bib39)\]does not work in our case, and we show contraction of SGD iterates with a new argument\.
Second, with the fixed\-α\\alphainvariant law in hand, we then identify its scaling limit asα↓0\\alpha\\downarrow 0\. The localmm\-flat geometry determines the normalization: the SGD iterates concentrate on the scaleα1/m\\alpha^\{1/m\}\. Then, the scaling limit is given by the stationary distribution of the stochastic differential equation
dYt=−h0\(Yt\)dt\+Σ1/2dBt,dY\_\{t\}=\-h\_\{0\}\(Y\_\{t\}\)\\,dt\+\\Sigma^\{1/2\}\\,dB\_\{t\},whereh0h\_\{0\}denotes the limiting drift, andΣ\\Sigmadenotes the asymptotic variance which is identified via a Poisson equation argument\. Whenm=2m=2, the above equation recovers the Ornstein–Uhlenbeck diffusion, and the invariant law is Gaussian\. Form\>2m\>2, the drifth0h\_\{0\}is nonlinear, and the stationary distribution is generally non\-Gaussian\. We note that Chen, Mou, and Maguluri\[[12](https://arxiv.org/html/2607.16384#bib.bib12)\]numerically exhibited the non\-Gaussian quartic scaling\. Wang et al\.\[[43](https://arxiv.org/html/2607.16384#bib.bib43), Section 5\]conjectured a one\-dimensional flat\-minimum Gibbs approximation with the corresponding nonstandard scaling and supported it numerically\. Our result establishes the scaling limit rigorously in a multidimensional setting with Markovian noise, confirming and extending the phenomenon suggested by prior work\.
Finally, for coordinate\-separable objective functions, we provide the corresponding constant\-stepsize and scaling limit results\. Notably, if the coordinates have unequal local flatness exponents, only the flattest coordinates can remain nonzero under the common scale associated with the largest exponent\.
Table[1](https://arxiv.org/html/2607.16384#S1.T1)summarizes the cases appearing in the main results\.
Table 1:The primary features covered by the main results\. The exponentmmis local and determines the normalization of the invariant law and the limiting drifth0h\_\{0\}\. The exponentβ\\betais global and determines the tail part of the induced metric\.
#### Related literature\.
The classical stochastic approximation literature studies decreasing stepsizes and averaging of iterates, establishing almost\-sure convergence and asymptotic normality\[[40](https://arxiv.org/html/2607.16384#bib.bib40),[24](https://arxiv.org/html/2607.16384#bib.bib24),[16](https://arxiv.org/html/2607.16384#bib.bib16),[28](https://arxiv.org/html/2607.16384#bib.bib28),[42](https://arxiv.org/html/2607.16384#bib.bib42),[38](https://arxiv.org/html/2607.16384#bib.bib38),[7](https://arxiv.org/html/2607.16384#bib.bib7),[27](https://arxiv.org/html/2607.16384#bib.bib27),[9](https://arxiv.org/html/2607.16384#bib.bib9)\]\. Modern nonasymptotic analyses of decreasing\-step stochastic optimization include\[[33](https://arxiv.org/html/2607.16384#bib.bib33),[32](https://arxiv.org/html/2607.16384#bib.bib32),[2](https://arxiv.org/html/2607.16384#bib.bib2)\]\. These results concern convergence of the iterates to an optimizer\. The object here is different: for a constant stepsize, the last iterate has a nontrivial invariant law, and the limitα↓0\\alpha\\downarrow 0is taken at stationarity\.
For constant stepsizes, the algorithm is analyzed as a Markov chain\. Pflug\[[37](https://arxiv.org/html/2607.16384#bib.bib37)\]studied small\-stepsize asymptotic laws for stochastic optimization\. Mandt, Hoffman, and Blei\[[29](https://arxiv.org/html/2607.16384#bib.bib29)\]developed quadratic diffusion approximation for constant\-stepsize SGD\. Dieuleveut, Durmus, and Bach\[[13](https://arxiv.org/html/2607.16384#bib.bib13)\]developed invariant measure and bias expansions for smooth strongly convex SGD\. Merad and Gaïffas\[[30](https://arxiv.org/html/2607.16384#bib.bib30)\]proved Wasserstein convergence and concentration in related strongly convex settings\. In a different asymptotic regime, Yu, Balasubramanian, Volgushev, and Erdogdu\[[45](https://arxiv.org/html/2607.16384#bib.bib45)\]established a central limit theorem for averages of iterates at a fixed stepsize, together with a characterization of the invariant bias\.
Besides the classical work of Pflug\[[37](https://arxiv.org/html/2607.16384#bib.bib37)\], the recent works most closely related to our scaling limit of invariant laws are Chen, Mou, and Maguluri\[[12](https://arxiv.org/html/2607.16384#bib.bib12)\], Zhang et al\.\[[46](https://arxiv.org/html/2607.16384#bib.bib46)\], and Wang et al\.\[[43](https://arxiv.org/html/2607.16384#bib.bib43)\]\. Pflug\[[37](https://arxiv.org/html/2607.16384#bib.bib37)\]and Chen, Mou, and Maguluri\[[12](https://arxiv.org/html/2607.16384#bib.bib12)\]both first pass to constant\-stepsize invariant laws and then letα↓0\\alpha\\downarrow 0to obtain a Gaussian limit\. The latter work obtains Gaussian limits in smooth strongly convex, linear, and contractive settings, and numerically demonstrates theα1/4\\alpha^\{1/4\}scaling and a non\-Gaussian limit for a quartic objective\. Zhang et al\.\[[46](https://arxiv.org/html/2607.16384#bib.bib46)\]establish steady\-state convergence for nonsmooth contractive stochastic approximation: their scale remainsα\\sqrt\{\\alpha\}, while nonsmooth local dynamics can produce a non\-Gaussian limit\. Wang et al\.\[[43](https://arxiv.org/html/2607.16384#bib.bib43)\]give a one\-dimensional Gibbs approximation for flat convex objectives with the correct nonstandard scaling, conditional on two conjectures concerning stationary moments and Stein equation regularity\. We establish the flat\-minimum limit unconditionally and allow a multidimensional setting with Markovian noise\.
The tail behavior we study in this work is closest to recent work on subquadratic SGD for robust and quantile regression\[[47](https://arxiv.org/html/2607.16384#bib.bib47)\]\. In the notation below, that setting corresponds tom=2m=2andβ<2\\beta<2: the objective is locally strongly convex, but its tail for large‖x‖\\left\\lVert x\\right\\rVertis subquadratic\. The present paper allows this tail behavior to coexist with local flatnessm\>2m\>2\.
Markovian noise introduces a separate issue because the stationary iterate is correlated with future values of the driving chain\. Poisson equations are the standard tool for separating this temporal dependence in stochastic approximation\. For linear stochastic approximation, Huo, Chen, and Xie\[[20](https://arxiv.org/html/2607.16384#bib.bib20)\]proved convergence to a unique stationary distribution and developed a stepsize expansion of its bias\. More recent analyses quantify how Markovian memory and nonlinear updates affect stationary bias and weak convergence\[[21](https://arxiv.org/html/2607.16384#bib.bib21),[18](https://arxiv.org/html/2607.16384#bib.bib18)\]\. Here the Poisson equation facilitates identifying the diffusion covarianceΣ\\Sigmaas the asymptotic covariance of the stationary sequenceg\(0,ξn\)g\(0,\\xi\_\{n\}\)\.
Our proof of constant\-stepsize ergodicity builds on recent work by Qu, Blanchet, and Glynn\[[39](https://arxiv.org/html/2607.16384#bib.bib39)\]\. Their work proves Wasserstein convergence from contractive drift conditions\. We do not verify their conditions directly as they may not always hold in our case\. Thus, while we adopt the induced metric of\[[39](https://arxiv.org/html/2607.16384#bib.bib39)\], our constant\-stepsize ergodicity result requires a new directional kernel estimate adapted to the SGD dynamics\.
For the scaling limit asα↓0\\alpha\\downarrow 0, a related line of work addresses quantitative stationary approximations, often through generator comparisons or Stein\-type arguments\. As a methodological precursor, Gast\[[17](https://arxiv.org/html/2607.16384#bib.bib17)\]used generator comparison to obtain rates for mean\-field models\. Allmeier and Gast\[[1](https://arxiv.org/html/2607.16384#bib.bib1)\]used related generator expansions to compute bias corrections for constant\-step stochastic approximation with Markovian noise\. Wang et al\.\[[43](https://arxiv.org/html/2607.16384#bib.bib43)\]proved nonasymptotic Wasserstein and tail bounds in smooth settings, including Markovian noise through Poisson equations\. Wei et al\.\[[44](https://arxiv.org/html/2607.16384#bib.bib44)\]obtained finite\-sample Gaussian approximations for constant\-stepsize SGD through linearization\. Our result instead identifies the scaling limit of invariant laws for multidimensional flat objectives, where the limiting drift is nonlinear and the limit is generally non\-Gaussian\. While we do not obtain a finite\-α\\alphaconvergence rate, to the best of our knowledge this is the first rigorous invariant\-law scaling limit for flat minima, even in dimension one\.
#### Organization\.
The paper is organized around the two limits involved in the analysis: first the limit of the SGD iterates at constant stepsize, and then the small\-stepsize limit at stationarity\. Section[2](https://arxiv.org/html/2607.16384#S2)states all the main results\. Section[3](https://arxiv.org/html/2607.16384#S3)studies statistical examples that motivate local flatness and subquadratic tails\. Section[4](https://arxiv.org/html/2607.16384#S4)provides the proof strategy and proves the main theorems\. Additional technical proofs are deferred to the appendices\.
## 2Main results
To simplify the statement of the main results, we assume that zero is the minimizer of the objective function without loss of generality\. Indeed, withx⋆x\_\{\\star\}denoting the minimizer of the original objective, replacing the optimization variable byx−x⋆x\-x\_\{\\star\}translates the minimizer to the origin\.
LetH:ℝd→ℝH:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}be the objective function, and leth:=∇Hh:=\\nabla H\. Consider SGD with Markovian noise defined as follows\. Let\(ξn\)\(\\xi\_\{n\}\)be the driving Markov chain\. It takes values in a closed convex setΞ⊆ℝdΞ\\Xi\\subseteq\\mathbb\{R\}^\{d\_\{\\Xi\}\}and evolves according to
ξn\+1=Φ\(ξn,Un\+1\),\\xi\_\{n\+1\}=\\Phi\(\\xi\_\{n\},U\_\{n\+1\}\),where\(Un\)n≥1\(U\_\{n\}\)\_\{n\\geq 1\}are i\.i\.d\. innovations taking values in a measurable space𝒰\\mathcal\{U\}andΦ:Ξ×𝒰→Ξ\\Phi:\\Xi\\times\\mathcal\{U\}\\to\\Xiis Borel measurable\. For a constant stepsizeα∈\(0,1\)\\alpha\\in\(0,1\), the SGD is defined by
Xn\+1=Xn−α\{h\(Xn\)\+g\(Xn,ξn\+1\)\}\.X\_\{n\+1\}=X\_\{n\}\-\\alpha\\\{h\(X\_\{n\}\)\+g\(X\_\{n\},\\xi\_\{n\+1\}\)\\\}\.\(2\.1\)Equivalently, withZn=\(Xn,ξn\)∈𝖹:=ℝd×ΞZ\_\{n\}=\(X\_\{n\},\\xi\_\{n\}\)\\in\\mathsf\{Z\}:=\\mathbb\{R\}^\{d\}\\times\\Xi,
Zn\+1=Fα,Un\+1\(Zn\),Z\_\{n\+1\}=F\_\{\\alpha,U\_\{n\+1\}\}\(Z\_\{n\}\),where
Fα,u\(x,ξ\):=\(x−α\{h\(x\)\+g\(x,Φ\(ξ,u\)\)\},Φ\(ξ,u\)\)\.F\_\{\\alpha,u\}\(x,\\xi\):=\\left\(x\-\\alpha\\\{h\(x\)\+g\(x,\\Phi\(\\xi,u\)\)\\\},\\Phi\(\\xi,u\)\\right\)\.\(2\.2\)We suppress the dependence onα\\alphawhen it is clear from context\. Note that the augmented processZnZ\_\{n\}is a Markov chain while theXX\-coordinate alone need not be Markov\.
### 2\.1Geometric ergodicity in Wasserstein distance
Our first result is regarding the convergence of the constant\-stepsize SGD to an invariant law\. More precisely, we establish the geometric ergodicity of the SGD iterates in a Wasserstein distance\. We first introduce the metrics used in this work\. These metrics come from the work of Qu, Blanchet, and Glynn\[[39](https://arxiv.org/html/2607.16384#bib.bib39)\], specialized to the augmented finite\-dimensional state space\.
###### Definition 2\.2\.
LetEEbe a closed convex subset of a finite\-dimensional normed space, with base norm\|⋅\|⋆\|\\cdot\|\_\{\\star\}\. For a continuous weight functionV:E→\[1,∞\)V:E\\to\[1,\\infty\), define the metric induced byVV
dV,⋆\(z,z′\):=infγ∈ACE\(z,z′\)∫01V\(γ\(t\)\)\|γ˙\(t\)\|⋆𝑑t,d\_\{V,\\star\}\(z,z^\{\\prime\}\):=\\inf\_\{\\gamma\\in\\operatorname\{AC\}\_\{E\}\(z,z^\{\\prime\}\)\}\\int\_\{0\}^\{1\}V\(\\gamma\(t\)\)\\,\|\\dot\{\\gamma\}\(t\)\|\_\{\\star\}\\,dt,whereACE\(z,z′\)\\operatorname\{AC\}\_\{E\}\(z,z^\{\\prime\}\)is the set of absolutely continuous curvesγ:\[0,1\]→E\\gamma:\[0,1\]\\to Ewithγ\(0\)=z\\gamma\(0\)=zandγ\(1\)=z′\\gamma\(1\)=z^\{\\prime\}\. For probability measuresμ,ν\\mu,\\nuonEE, define
WV,⋆\(μ,ν\):=infγ∈Π\(μ,ν\)∫E×EdV,⋆\(z,z′\)γ\(dz,dz′\),W\_\{V,\\star\}\(\\mu,\\nu\):=\\inf\_\{\\gamma\\in\\Pi\(\\mu,\\nu\)\}\\int\_\{E\\times E\}d\_\{V,\\star\}\(z,z^\{\\prime\}\)\\,\\gamma\(dz,dz^\{\\prime\}\),whereΠ\(μ,ν\)\\Pi\(\\mu,\\nu\)is the set of couplings ofμ\\muandν\\nu\.
WhenV≡1V\\equiv 1onℝd\\mathbb\{R\}^\{d\}with the Euclidean norm, we simply writeW1W\_\{1\}for the Wasserstein distance\. For the augmented chain, the base norm is defined to be
\|\(x,ξ\)\|α:=‖x‖\+α−1‖ξ‖,\(x,ξ\)∈ℝd×ℝdΞ,\|\(x,\\xi\)\|\_\{\\alpha\}:=\\left\\lVert x\\right\\rVert\+\\alpha^\{\-1\}\\left\\lVert\\xi\\right\\rVert,\\qquad\(x,\\xi\)\\in\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{d\_\{\\Xi\}\},\(2\.3\)where the choice of the balancing factorα−1\\alpha^\{\-1\}is suggested by the analysis\. The corresponding induced metric and Wasserstein distance are denoted bydV,αd\_\{V,\\alpha\}andWV,αW\_\{V,\\alpha\}\.
Next, we introduce the assumptions on the objective functionHH\.
###### Assumption 2\.3\(Objective function\)\.
Fixm≥2m\\geq 2andβ∈\[1,2\]\\beta\\in\[1,2\]\. The objectiveH:ℝd→ℝH:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}satisfies the following conditions\.
1. \(H1\)Regularity and convexity\.H∈C2\(ℝd\)H\\in C^\{2\}\(\\mathbb\{R\}^\{d\}\),HHis convex, and0∈argminH0\\in\\arg\\min H\.
2. \(H2\)Local flatness\.There existsRH\>0R\_\{H\}\>0and constantscin,Cin\>0c\_\{\\rm in\},C\_\{\\rm in\}\>0such that, for all0<‖x‖≤RH0<\\left\\lVert x\\right\\rVert\\leq R\_\{H\}, ∇2H\(x\)⪰cin‖x‖m−2Id,‖h\(x\)‖≤Cin‖x‖m−1\.\\nabla^\{2\}H\(x\)\\succeq c\_\{\\rm in\}\\left\\lVert x\\right\\rVert^\{m\-2\}I\_\{d\},\\qquad\\left\\lVert h\(x\)\\right\\rVert\\leq C\_\{\\rm in\}\\left\\lVert x\\right\\rVert^\{m\-1\}\.Whenm=2m=2, this condition is understood to extend tox=0x=0by continuity of∇2H\\nabla^\{2\}H\.
3. \(H3\)Tail condition\.One of the following two tail conditions holds\. 1. \(H3\-a\)Quadratic tail\.Forβ=2\\beta=2, there existscout\>0c\_\{\\rm out\}\>0such that, for all‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\}, ∇2H\(x\)⪰coutId\.\\nabla^\{2\}H\(x\)\\succeq c\_\{\\rm out\}I\_\{d\}\. 2. \(H3\-b\)Subquadratic tail\.For1≤β<21\\leq\\beta<2, there existcout,Cout\>0c\_\{\\rm out\},C\_\{\\rm out\}\>0such that, for all‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\}, ⟨x,h\(x\)⟩≥cout‖x‖β,‖h\(x\)‖≤Cout‖x‖β−1\.\\left\\langle x,h\(x\)\\right\\rangle\\geq c\_\{\\rm out\}\\left\\lVert x\\right\\rVert^\{\\beta\},\\qquad\\left\\lVert h\(x\)\\right\\rVert\\leq C\_\{\\rm out\}\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\.
The tail condition[\(H3\)](https://arxiv.org/html/2607.16384#S2.I1.i3)supplies the restoring drift towards the minimizer for the SGD, while the local flatness[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)determines the scale of fluctuations near the minimizer and the scaling limit\.
The next assumptions concern the noise in the SGD and the one\-step iterate in \([2\.1](https://arxiv.org/html/2607.16384#S2.E1)\)–\([2\.2](https://arxiv.org/html/2607.16384#S2.E2)\)\.
###### Assumption 2\.5\(Stochastic update\)\.
The driving Markov chain and stochastic update satisfy the following conditions\.
1. \(N1\)Driving\-chain contraction\.There exists a measurable functionLΦ:𝒰→\[0,1\]L\_\{\\Phi\}:\\mathcal\{U\}\\to\[0,1\]such that, for allξ,η∈Ξ\\xi,\\eta\\in\\Xiand allu∈𝒰u\\in\\mathcal\{U\}, ‖Φ\(ξ,u\)−Φ\(η,u\)‖≤LΦ\(u\)‖ξ−η‖,𝔼LΦ\(U1\)<1\.\\left\\lVert\\Phi\(\\xi,u\)\-\\Phi\(\\eta,u\)\\right\\rVert\\leq L\_\{\\Phi\}\(u\)\\left\\lVert\\xi\-\\eta\\right\\rVert,\\qquad\\mathbb\{E\}L\_\{\\Phi\}\(U\_\{1\}\)<1\.
2. \(N2\)Lipschitzness of noise\.There existsLg,Φ\>0L\_\{g,\\Phi\}\>0such that, for allx∈ℝdx\\in\\mathbb\{R\}^\{d\}, allξ,η∈Ξ\\xi,\\eta\\in\\Xi, and allu∈𝒰u\\in\\mathcal\{U\}, ‖g\(x,Φ\(ξ,u\)\)−g\(x,Φ\(η,u\)\)‖≤Lg,Φ‖ξ−η‖\.\\left\\lVert g\(x,\\Phi\(\\xi,u\)\)\-g\(x,\\Phi\(\\eta,u\)\)\\right\\rVert\\leq L\_\{g,\\Phi\}\\left\\lVert\\xi\-\\eta\\right\\rVert\.
3. \(N3\)Reference\-point integrability\.There existsξ⋆∈Ξ\\xi\_\{\\star\}\\in\\Xisuch that 𝔼‖Φ\(ξ⋆,U1\)−ξ⋆‖<∞\.\\mathbb\{E\}\\left\\lVert\\Phi\(\\xi\_\{\\star\},U\_\{1\}\)\-\\xi\_\{\\star\}\\right\\rVert<\\infty\.\(2\.4\)Ifβ=2\\beta=2, assume also 𝔼‖g\(0,Φ\(ξ⋆,U1\)\)‖<∞\.\\mathbb\{E\}\\left\\lVert g\(0,\\Phi\(\\xi\_\{\\star\},U\_\{1\}\)\)\\right\\rVert<\\infty\.When1≤β<21\\leq\\beta<2, this condition follows from[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)atx=0x=0\.
4. \(N4\)Subquadratic tail dissipativity\.Define g¯\(x,ξ\):=𝔼U1\[g\(x,Φ\(ξ,U1\)\)\]\.\\bar\{g\}\(x,\\xi\):=\\mathbb\{E\}\_\{U\_\{1\}\}\[g\(x,\\Phi\(\\xi,U\_\{1\}\)\)\]\.\(2\.5\)Whenβ=2\\beta=2, this condition is not imposed\. When1≤β<21\\leq\\beta<2, there existscdiss\>0c\_\{\\rm diss\}\>0such that, for all‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\}and allξ∈Ξ\\xi\\in\\Xi, ⟨x,h\(x\)\+g¯\(x,ξ\)⟩≥cdiss‖x‖β\.\\left\\langle x,h\(x\)\+\\bar\{g\}\(x,\\xi\)\\right\\rangle\\geq c\_\{\\rm diss\}\\left\\lVert x\\right\\rVert^\{\\beta\}\.
5. \(N5\)Subquadratic exponential integrability\.Whenβ=2\\beta=2, this condition is not imposed\. When1≤β<21\\leq\\beta<2, there existsλ0\>0\\lambda\_\{0\}\>0such that supx∈ℝdsupξ∈Ξ𝔼exp\(λ0‖g\(x,Φ\(ξ,U1\)\)‖1\+‖x‖β−1\)<∞\.\\sup\_\{x\\in\\mathbb\{R\}^\{d\}\}\\sup\_\{\\xi\\in\\Xi\}\\mathbb\{E\}\\exp\\\!\\left\(\\lambda\_\{0\}\\frac\{\\left\\lVert g\(x,\\Phi\(\\xi,U\_\{1\}\)\)\\right\\rVert\}\{1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\}\\right\)<\\infty\.Here and below,‖x‖0=1\\left\\lVert x\\right\\rVert^\{0\}=1by convention\.
6. \(N6\)Nondegeneracy of noise at minimizer\.Ifm=2m=2, this condition is not imposed\. Ifm\>2m\>2, assume that there exist constantsεg\>0\\varepsilon\_\{g\}\>0andpg\>0p\_\{g\}\>0such that infξ∈Ξℙ\(‖g\(0,Φ\(ξ,U1\)\)‖≥εg\)≥pg\.\\inf\_\{\\xi\\in\\Xi\}\\mathbb\{P\}\\\!\\left\(\\left\\lVert g\(0,\\Phi\(\\xi,U\_\{1\}\)\)\\right\\rVert\\geq\\varepsilon\_\{g\}\\right\)\\geq p\_\{g\}\.
7. \(N7\)Co\-coercivity and noise perturbation\.For\(x,ξ\)∈ℝd×Ξ\(x,\\xi\)\\in\\mathbb\{R\}^\{d\}\\times\\Xi, set 𝖦\(x,ξ\):=h\(x\)\+g\(x,ξ\)\.\\mathsf\{G\}\(x,\\xi\):=h\(x\)\+g\(x,\\xi\)\.For everyξ∈Ξ\\xi\\in\\Xi, the mapx↦𝖦\(x,ξ\)x\\mapsto\\mathsf\{G\}\(x,\\xi\)isC1C^\{1\}\. There exist constantsL𝖦≥1L\_\{\\mathsf\{G\}\}\\geq 1andθ∈\[0,1\)\\theta\\in\[0,1\)such that, for everyx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\}and everyξ∈Ξ\\xi\\in\\Xi, ⟨x−y,𝖦\(x,ξ\)−𝖦\(y,ξ\)⟩≥L𝖦−1‖𝖦\(x,ξ\)−𝖦\(y,ξ\)‖2,\\left\\langle x\-y,\\mathsf\{G\}\(x,\\xi\)\-\\mathsf\{G\}\(y,\\xi\)\\right\\rangle\\geq L\_\{\\mathsf\{G\}\}^\{\-1\}\\left\\lVert\\mathsf\{G\}\(x,\\xi\)\-\\mathsf\{G\}\(y,\\xi\)\\right\\rVert^\{2\},\(2\.6\)andg¯\\bar\{g\}defined in \([2\.5](https://arxiv.org/html/2607.16384#S2.E5)\) satisfies ⟨x−y,g¯\(x,ξ\)−g¯\(y,ξ\)⟩≥−θ⟨x−y,h\(x\)−h\(y\)⟩\.\\left\\langle x\-y,\\bar\{g\}\(x,\\xi\)\-\\bar\{g\}\(y,\\xi\)\\right\\rangle\\geq\-\\theta\\left\\langle x\-y,h\(x\)\-h\(y\)\\right\\rangle\.\(2\.7\)
Before stating the first result, recall that the augmented space𝖹=ℝd×Ξ\\mathsf\{Z\}=\\mathbb\{R\}^\{d\}\\times\\Xiis equipped with the base norm\|⋅\|α\|\\cdot\|\_\{\\alpha\}in \([2\.3](https://arxiv.org/html/2607.16384#S2.E3)\)\. The metricsdV,αd\_\{V,\\alpha\}andWV,αW\_\{V,\\alpha\}are those of Definition[2\.2](https://arxiv.org/html/2607.16384#S2.Thmtheorem2)\. For a metricddon a spaceEE, define
𝒫1\(E,d\):=\{μ:∫Ed\(z,z0\)μ\(dz\)<∞for somez0∈E\}\.\\mathcal\{P\}\_\{1\}\(E,d\):=\\left\\\{\\mu:\\int\_\{E\}d\(z,z\_\{0\}\)\\,\\mu\(dz\)<\\infty\\text\{ for some \}z\_\{0\}\\in E\\right\\\}\.
###### Theorem 2\.7\(Geometric ergodicity of constant\-stepsize SGD\)\.
Assume Assumptions[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)and[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)\. Then there existα0\>0\\alpha\_\{0\}\>0andc∈\(0,1\)c\\in\(0,1\)such that, for every0<α≤α00<\\alpha\\leq\\alpha\_\{0\}, one can choose a continuous weight functionVα:𝖹→\[1,∞\)V\_\{\\alpha\}:\\mathsf\{Z\}\\to\[1,\\infty\)for which the augmented chainZn=\(Xn,ξn\)Z\_\{n\}=\(X\_\{n\},\\xi\_\{n\}\)admits a unique invariant law
πα∈𝒫1\(𝖹,dVα,α\)\.\\pi\_\{\\alpha\}\\in\\mathcal\{P\}\_\{1\}\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)\.Moreover,
WVα,α\(μPαn,νPαn\)≤\(1−cαm−1\)nWVα,α\(μ,ν\),n≥0,W\_\{V\_\{\\alpha\},\\alpha\}\(\\mu P\_\{\\alpha\}^\{n\},\\nu P\_\{\\alpha\}^\{n\}\)\\leq\(1\-c\\alpha^\{m\-1\}\)^\{n\}W\_\{V\_\{\\alpha\},\\alpha\}\(\\mu,\\nu\),\\qquad n\\geq 0,\(2\.8\)for all initial lawsμ,ν\\mu,\\nuwith finiteWVα,α\(μ,ν\)W\_\{V\_\{\\alpha\},\\alpha\}\(\\mu,\\nu\), wherePαP\_\{\\alpha\}is the transition kernel of the augmented chain\.
The proof of Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7)is given in Section[4\.1](https://arxiv.org/html/2607.16384#S4.SS1)\. Projecting the contraction of the augmented chain to theXX\-coordinate gives the following ordinaryW1W\_\{1\}consequence, whose proof is given in Appendix[A\.7](https://arxiv.org/html/2607.16384#A1.SS7)\.
###### Corollary 2\.9\.
Assume the hypotheses of Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7), and fixα∈\(0,α0\]\\alpha\\in\(0,\\alpha\_\{0\}\]\. Let\(πα\)X\(\\pi\_\{\\alpha\}\)\_\{X\}denote theXX\-marginal of the invariant lawπα\\pi\_\{\\alpha\}from Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7)\. Letz⋆=\(0,ξ⋆\)z\_\{\\star\}=\(0,\\xi\_\{\\star\}\), withξ⋆\\xi\_\{\\star\}as in Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3)\. There exists a sufficiently small constantκ\>0\\kappa\>0, depending only on the constants in Assumptions[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)and[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)and independent ofα\\alpha, such that the following holds\. Define
Γβ\(r\):=\(1\+r\)β−1exp\{κ\(\(1\+r\)2−β−1\)\},1≤β≤2\.\\Gamma\_\{\\beta\}\(r\):=\(1\+r\)^\{\\beta\-1\}\\exp\\\!\\left\\\{\\kappa\\bigl\(\(1\+r\)^\{2\-\\beta\}\-1\\bigr\)\\right\\\},\\qquad 1\\leq\\beta\\leq 2\.Then
𝒫1\(𝖹,dVα,α\)=\{μ:∫𝖹\[Γβ\(‖x‖\)\+α−1‖ξ−ξ⋆‖\]μ\(dx,dξ\)<∞\}\.\\mathcal\{P\}\_\{1\}\\bigl\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\\bigr\)=\\left\\\{\\mu:\\int\_\{\\mathsf\{Z\}\}\\Bigl\[\\Gamma\_\{\\beta\}\(\\left\\lVert x\\right\\rVert\)\+\\alpha^\{\-1\}\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert\\Bigr\]\\,\\mu\(dx,d\\xi\)<\\infty\\right\\\}\.For everyμ∈𝒫1\(𝖹,dVα,α\)\\mu\\in\\mathcal\{P\}\_\{1\}\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)and everyn≥0n\\geq 0,
W1\(\(μPαn\)X,\(πα\)X\)≤\(1−cαm−1\)nWVα,α\(μ,πα\)\.W\_\{1\}\\bigl\(\(\\mu P\_\{\\alpha\}^\{n\}\)\_\{X\},\(\\pi\_\{\\alpha\}\)\_\{X\}\\bigr\)\\leq\(1\-c\\alpha^\{m\-1\}\)^\{n\}W\_\{V\_\{\\alpha\},\\alpha\}\(\\mu,\\pi\_\{\\alpha\}\)\.
The set𝒫1\(𝖹,dVα,α\)\\mathcal\{P\}\_\{1\}\\bigl\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\\bigr\)is explicitly characterized in the above corollary\. In particular, every deterministic initial condition belongs to𝒫1\(𝖹,dVα,α\)\\mathcal\{P\}\_\{1\}\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)\. When1≤β<21\\leq\\beta<2, membership in this class requires an exponential moment of order2−β2\-\\betain theXX\-coordinate and a first moment in theξ\\xi\-coordinate\.
### 2\.2Small\-stepsize scaling limit
The invariant lawπα\\pi\_\{\\alpha\}of the SGD iterates in the last subsection is implicit and generally difficult to characterize\. Therefore, we turn to studying its scaling limit asα↓0\\alpha\\downarrow 0\. To this end, we introduce two additional conditions\.
###### Assumption 2\.10\(Objective function\)\.
Assume Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)together with the following\.
1. \(H4\)Local gradient expansion\.There existsH0∈C2\(ℝd\)H\_\{0\}\\in C^\{2\}\(\\mathbb\{R\}^\{d\}\), convex and homogeneous of degreemm, such thatH0\(0\)=0H\_\{0\}\(0\)=0,∇H0\(0\)=0\\nabla H\_\{0\}\(0\)=0, and ∇H\(x\)=∇H0\(x\)\+o\(‖x‖m−1\)asx→0\.\\nabla H\(x\)=\\nabla H\_\{0\}\(x\)\+o\(\\left\\lVert x\\right\\rVert^\{m\-1\}\)\\qquad\\text\{as \}x\\to 0\.
###### Assumption 2\.11\(Stochastic update\)\.
Assume Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)together with the following\. LetπΞ\\pi\_\{\\Xi\}denote the invariant law of the driving chain\(ξn\)n≥0\(\\xi\_\{n\}\)\_\{n\\geq 0\}, whose existence and uniqueness are proved in Lemma[B\.2](https://arxiv.org/html/2607.16384#A2.Thmtheorem2)\. The Markovian noise satisfies the following condition\.
1. \(N8\)Stationary centering and second moments\.The gradient errorg\(x,ξ\)g\(x,\\xi\)satisfies ∫Ξg\(x,ξ\)πΞ\(dξ\)=0,x∈ℝd,\\int\_\{\\Xi\}g\(x,\\xi\)\\,\\pi\_\{\\Xi\}\(d\\xi\)=0,\\qquad x\\in\\mathbb\{R\}^\{d\},\(2\.9\)and ∫Ξ\(‖g\(0,ξ\)‖2\+‖ξ−ξ⋆‖2\)πΞ\(dξ\)<∞\.\\int\_\{\\Xi\}\\bigl\(\\left\\lVert g\(0,\\xi\)\\right\\rVert^\{2\}\+\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert^\{2\}\\bigr\)\\,\\pi\_\{\\Xi\}\(d\\xi\)<\\infty\.\(2\.10\)
Let\(ξn\)n≥0\(\\xi\_\{n\}\)\_\{n\\geq 0\}be the driving chain initialized according to its invariant lawπΞ\\pi\_\{\\Xi\}, so that\(ξn\)n≥0\(\\xi\_\{n\}\)\_\{n\\geq 0\}is stationary\. By Assumption[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)[\(N8\)](https://arxiv.org/html/2607.16384#S2.I6.i8),\(g\(0,ξn\)\)n≥0\(g\(0,\\xi\_\{n\}\)\)\_\{n\\geq 0\}is centered and stationary\. Define
Σ:=Γ0\+∑k=1∞\(Γk\+Γk⊤\),Γk:=𝔼\[g\(0,ξ0\)g\(0,ξk\)⊤\]\.\\Sigma:=\\Gamma\_\{0\}\+\\sum\_\{k=1\}^\{\\infty\}\(\\Gamma\_\{k\}\+\\Gamma\_\{k\}^\{\\top\}\),\\qquad\\Gamma\_\{k\}:=\\mathbb\{E\}\[g\(0,\\xi\_\{0\}\)g\(0,\\xi\_\{k\}\)^\{\\top\}\]\.\(2\.11\)We will show thatΣ\\Sigmais the asymptotic covariance appearing in the scaling limit\. To see thatΣ\\Sigmais a well\-defined covariance matrix, Lemma[4\.6](https://arxiv.org/html/2607.16384#S4.Thmtheorem6)shows that the series converges absolutely and thatΣ\\Sigmais a finite positive semidefinite matrix\.
The following theorem, proved in Section[4\.2](https://arxiv.org/html/2607.16384#S4.SS2), identifies both the scaleα1/m\\alpha^\{1/m\}of the invariant law and its limiting distribution as the solution of a stochastic differential equation\.
###### Theorem 2\.13\(Scaling limit of invariant law\)\.
Assume Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. LetH0H\_\{0\}be as in Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4)and seth0:=∇H0h\_\{0\}:=\\nabla H\_\{0\}\. For each sufficiently smallα\>0\\alpha\>0, let\(X∞\(α\),ξ∞\(α\)\)\(X\_\{\\infty\}^\{\(\\alpha\)\},\\xi\_\{\\infty\}^\{\(\\alpha\)\}\)have lawπα\\pi\_\{\\alpha\}from Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7), and set
Yα:=α−1/mX∞\(α\)\.Y\_\{\\alpha\}:=\\alpha^\{\-1/m\}X\_\{\\infty\}^\{\(\\alpha\)\}\.ThenYα⇒Y∞Y\_\{\\alpha\}\\Rightarrow Y\_\{\\infty\}asα↓0\\alpha\\downarrow 0, whereY∞Y\_\{\\infty\}is the unique invariant distribution of
dYt=−h0\(Yt\)dt\+Σ1/2dBt,dY\_\{t\}=\-h\_\{0\}\(Y\_\{t\}\)\\,dt\+\\Sigma^\{1/2\}\\,dB\_\{t\},\(2\.12\)whereΣ\\Sigmais defined in \([2\.11](https://arxiv.org/html/2607.16384#S2.E11)\) andB=\(Bt\)t≥0B=\(B\_\{t\}\)\_\{t\\geq 0\}is a standarddd\-dimensional Brownian motion\.
We note that the covariance matrixΣ\\Sigmais allowed to be degenerate, in which case the scaling limit is singular\.
### 2\.3Coordinate\-separable objectives
The preceding results impose a common local flatness exponent in all directions\. We now consider coordinate\-separable objectives, for which the flatness exponent, and hence the natural scale of the invariant law, may vary across coordinates\. Suppose that
H\(x\)=∑i=1dHi\(xi\),g\(x,ξ\)=\(g1\(x1,ξ\),…,gd\(xd,ξ\)\)\.H\(x\)=\\sum\_\{i=1\}^\{d\}H\_\{i\}\(x\_\{i\}\),\\qquad g\(x,\\xi\)=\\bigl\(g\_\{1\}\(x\_\{1\},\\xi\),\\ldots,g\_\{d\}\(x\_\{d\},\\xi\)\\bigr\)\.\(2\.13\)Writinghi=Hi′h\_\{i\}=H\_\{i\}^\{\\prime\}, the SGD recursion becomes
Xn\+1,i=Xn,i−α\{hi\(Xn,i\)\+gi\(Xn,i,ξn\+1\)\},1≤i≤d,X\_\{n\+1,i\}=X\_\{n,i\}\-\\alpha\\\{h\_\{i\}\(X\_\{n,i\}\)\+g\_\{i\}\(X\_\{n,i\},\\xi\_\{n\+1\}\)\\\},\\qquad 1\\leq i\\leq d,\(2\.14\)whereξn\+1=Φ\(ξn,Un\+1\)\\xi\_\{n\+1\}=\\Phi\(\\xi\_\{n\},U\_\{n\+1\}\)is the driving\-chain recursion from the beginning of this section\. Thus the coordinates may remain dependent through the common driving chain\.
For each1≤i≤d1\\leq i\\leq d, assume that the one\-dimensional pair\(Hi,gi\)\(H\_\{i\},g\_\{i\}\)satisfies the one\-dimensional versions of Assumptions[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)and[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5), with parameters\(mi,βi\)\(m\_\{i\},\\beta\_\{i\}\)\. The constants in these assumptions may depend onii\. Applying the construction of Section[4\.1](https://arxiv.org/html/2607.16384#S4.SS1)withd=1d=1and\(H,g,m,β\)\(H,g,m,\\beta\)replaced by\(Hi,gi,mi,βi\)\(H\_\{i\},g\_\{i\},m\_\{i\},\\beta\_\{i\}\), onℝ×Ξ\\mathbb\{R\}\\times\\Xiequipped with the common base norm
\|\(xi,ξ\)\|α:=\|xi\|\+α−1‖ξ‖,\\lvert\(x\_\{i\},\\xi\)\\rvert\_\{\\alpha\}:=\|x\_\{i\}\|\+\\alpha^\{\-1\}\\left\\lVert\\xi\\right\\rVert,gives a weightVα,iV\_\{\\alpha,i\}and an induced metricdi,α:=dVα,i,α\(i\)d\_\{i,\\alpha\}:=d\_\{V\_\{\\alpha,i\},\\alpha\}^\{\(i\)\}\. Define the additive metric onℝd×Ξ\\mathbb\{R\}^\{d\}\\times\\Xiby
dsep,α\(\(x,ξ\),\(y,η\)\):=∑i=1ddi,α\(\(xi,ξ\),\(yi,η\)\),d\_\{\{\\rm sep\},\\alpha\}\(\(x,\\xi\),\(y,\\eta\)\):=\\sum\_\{i=1\}^\{d\}d\_\{i,\\alpha\}\(\(x\_\{i\},\\xi\),\(y\_\{i\},\\eta\)\),\(2\.15\)and letWsep,αW\_\{\{\\rm sep\},\\alpha\}denote the Wasserstein–1 distance with costdsep,αd\_\{\{\\rm sep\},\\alpha\}\. Set
mmax:=max1≤i≤dmi\.m\_\{\\max\}:=\\max\_\{1\\leq i\\leq d\}m\_\{i\}\.
The following corollary extends Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7)to this separable setting\.
###### Corollary 2\.15\(Geometric ergodicity for separable objectives\)\.
Suppose that \([2\.13](https://arxiv.org/html/2607.16384#S2.E13)\) holds and that, for every1≤i≤d1\\leq i\\leq d, the pair\(Hi,gi\)\(H\_\{i\},g\_\{i\}\)satisfies the one\-dimensional versions of Assumptions[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)and[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)\. Then there existα0∈\(0,1\)\\alpha\_\{0\}\\in\(0,1\)andc0∈\(0,1\)c\_\{0\}\\in\(0,1\)such that, for everyα∈\(0,α0\]\\alpha\\in\(0,\\alpha\_\{0\}\], the augmented chainZn=\(Xn,ξn\)Z\_\{n\}=\(X\_\{n\},\\xi\_\{n\}\)admits a unique invariant law
παsep∈𝒫1\(ℝd×Ξ,dsep,α\)\.\\pi\_\{\\alpha\}^\{\\rm sep\}\\in\\mathcal\{P\}\_\{1\}\(\\mathbb\{R\}^\{d\}\\times\\Xi,d\_\{\{\\rm sep\},\\alpha\}\)\.Moreover, ifPαsepP\_\{\\alpha\}^\{\\rm sep\}denotes its transition kernel, then
Wsep,α\(μ\(Pαsep\)n,ν\(Pαsep\)n\)≤\(1−c0αmmax−1\)nWsep,α\(μ,ν\),n≥0,W\_\{\{\\rm sep\},\\alpha\}\\bigl\(\\mu\(P\_\{\\alpha\}^\{\\rm sep\}\)^\{n\},\\nu\(P\_\{\\alpha\}^\{\\rm sep\}\)^\{n\}\\bigr\)\\leq\(1\-c\_\{0\}\\alpha^\{m\_\{\\max\}\-1\}\)^\{n\}W\_\{\{\\rm sep\},\\alpha\}\(\\mu,\\nu\),\\qquad n\\geq 0,\(2\.16\)for all probability lawsμ,ν\\mu,\\nufor which the right\-hand side is finite\.
The contraction rate is determined bymmaxm\_\{\\max\}: the coordinates with the largest flatness exponent have the largest contraction factor and hence govern the convergence of the full chain\. Corollary[2\.15](https://arxiv.org/html/2607.16384#S2.Thmtheorem15)is proved in Appendix[C\.1](https://arxiv.org/html/2607.16384#A3.SS1)\.
We next turn to the scaling limit\. For each coordinate, assume in addition the one\-dimensional versions of Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. LetHi,0H\_\{i,0\}be the homogeneous function appearing in the one\-dimensional version of Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4), and sethi,0:=Hi,0′h\_\{i,0\}:=H\_\{i,0\}^\{\\prime\}\. Initialize the driving chain in its invariant lawπΞ\\pi\_\{\\Xi\}, and define
g0sep\(ξ\):=\(g1\(0,ξ\),…,gd\(0,ξ\)\),ξ∈Ξ\.g\_\{0\}^\{\\rm sep\}\(\\xi\):=\\bigl\(g\_\{1\}\(0,\\xi\),\\ldots,g\_\{d\}\(0,\\xi\)\\bigr\),\\qquad\\xi\\in\\Xi\.Then\(g0sep\(ξn\)\)n≥0\(g\_\{0\}^\{\\rm sep\}\(\\xi\_\{n\}\)\)\_\{n\\geq 0\}is centered and stationary\. Its asymptotic covariance is
Σsep:=Γ0sep\+∑k=1∞\(Γksep\+\(Γksep\)⊤\),Γksep:=𝔼\[g0sep\(ξ0\)g0sep\(ξk\)⊤\]\.\\Sigma\_\{\\rm sep\}:=\\Gamma\_\{0\}^\{\\rm sep\}\+\\sum\_\{k=1\}^\{\\infty\}\\bigl\(\\Gamma\_\{k\}^\{\\rm sep\}\+\(\\Gamma\_\{k\}^\{\\rm sep\}\)^\{\\top\}\\bigr\),\\qquad\\Gamma\_\{k\}^\{\\rm sep\}:=\\mathbb\{E\}\\bigl\[g\_\{0\}^\{\\rm sep\}\(\\xi\_\{0\}\)g\_\{0\}^\{\\rm sep\}\(\\xi\_\{k\}\)^\{\\top\}\\bigr\]\.\(2\.17\)The coordinatewise Poisson construction in Lemma[4\.6](https://arxiv.org/html/2607.16384#S4.Thmtheorem6)shows that the series converges absolutely and thatΣsep\\Sigma\_\{\\rm sep\}is finite and positive semidefinite\. Its off\-diagonal entries retain the temporal cross\-covariances created by a common driving chain\. Let
ℐ∗:=\{i:mi=mmax\},P∗:=diag\(𝟏\{i∈ℐ∗\}\),Σ∗:=P∗ΣsepP∗\.\\mathcal\{I\}\_\{\*\}:=\\\{i:m\_\{i\}=m\_\{\\max\}\\\},\\qquad P\_\{\*\}:=\\operatorname\{diag\}\\bigl\(\\mathbf\{1\}\_\{\\\{i\\in\\mathcal\{I\}\_\{\*\}\\\}\}\\bigr\),\\qquad\\Sigma\_\{\*\}:=P\_\{\*\}\\Sigma\_\{\\rm sep\}P\_\{\*\}\.\(2\.18\)BothΣsep\\Sigma\_\{\\rm sep\}andΣ∗\\Sigma\_\{\*\}are allowed to be degenerate\.
For each1≤i≤d1\\leq i\\leq d, the\(xi,ξ\)\(x\_\{i\},\\xi\)\-marginal ofπαsep\\pi\_\{\\alpha\}^\{\\rm sep\}is invariant for the one\-dimensional augmented chain associated with\(Hi,gi\)\(H\_\{i\},g\_\{i\}\)\. Therefore Theorem[2\.13](https://arxiv.org/html/2607.16384#S2.Thmtheorem13)applied to this chain gives
α−1/miX∞,i\(α\),sep⇒Y∞,i,\\alpha^\{\-1/m\_\{i\}\}X\_\{\\infty,i\}^\{\(\\alpha\),\\rm sep\}\\Rightarrow Y\_\{\\infty,i\},whereY∞,iY\_\{\\infty,i\}has the unique invariant law of the one\-dimensional diffusion
dYi,t=−hi,0\(Yi,t\)dt\+\(Σsep\)ii1/2dBi,t,dY\_\{i,t\}=\-h\_\{i,0\}\(Y\_\{i,t\}\)\\,dt\+\(\\Sigma\_\{\\rm sep\}\)\_\{ii\}^\{1/2\}\\,dB\_\{i,t\},andBiB\_\{i\}is a standard one\-dimensional Brownian motion\. Thus coordinateiihas natural scaleα1/mi\\alpha^\{1/m\_\{i\}\}\.
###### Corollary 2\.17\(Scaling limit for separable objectives\)\.
Suppose that \([2\.13](https://arxiv.org/html/2607.16384#S2.E13)\) holds and that, for every1≤i≤d1\\leq i\\leq d, the pair\(Hi,gi\)\(H\_\{i\},g\_\{i\}\)satisfies the one\-dimensional versions of Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. Let\(X∞\(α\),sep,ξ∞\(α\),sep\)\(X\_\{\\infty\}^\{\(\\alpha\),\\rm sep\},\\xi\_\{\\infty\}^\{\(\\alpha\),\\rm sep\}\)have lawπαsep\\pi\_\{\\alpha\}^\{\\rm sep\}, and set
h0sep\(y\):=\(h1,0\(y1\),…,hd,0\(yd\)\)\.h\_\{0\}^\{\\rm sep\}\(y\):=\\bigl\(h\_\{1,0\}\(y\_\{1\}\),\\ldots,h\_\{d,0\}\(y\_\{d\}\)\\bigr\)\.Then, under the common normalization determined bymmaxm\_\{\\max\},
α−1/mmaxX∞\(α\),sep⇒Y∞sep,\\alpha^\{\-1/m\_\{\\max\}\}X\_\{\\infty\}^\{\(\\alpha\),\\rm sep\}\\Rightarrow Y\_\{\\infty\}^\{\\rm sep\},whereY∞sepY\_\{\\infty\}^\{\\rm sep\}has the unique invariant law of thedd\-dimensional diffusion
dYt=−h0sep\(Yt\)dt\+Σ∗1/2dBt\.dY\_\{t\}=\-h\_\{0\}^\{\\rm sep\}\(Y\_\{t\}\)\\,dt\+\\Sigma\_\{\*\}^\{1/2\}\\,dB\_\{t\}\.\(2\.19\)Here,Σ∗\\Sigma\_\{\*\}is defined in \([2\.18](https://arxiv.org/html/2607.16384#S2.E18)\) andBBis a standarddd\-dimensional Brownian motion\.
Fori∉ℐ∗i\\notin\\mathcal\{I\}\_\{\*\}, theiith component of \([2\.19](https://arxiv.org/html/2607.16384#S2.E19)\) has no Brownian forcing and follows the deterministic flowdYi,t=−hi,0\(Yi,t\)dtdY\_\{i,t\}=\-h\_\{i,0\}\(Y\_\{i,t\}\)\\,dt, whose unique invariant law isδ0\\delta\_\{0\}\. Thus, under the common normalization, only the coordinates with the largest flatness exponent can have a nonzero limit\. Ifmi=mm\_\{i\}=mfor everyii, thenP∗=IdP\_\{\*\}=I\_\{d\}andΣ∗=Σsep\\Sigma\_\{\*\}=\\Sigma\_\{\\rm sep\}\. Corollary[2\.17](https://arxiv.org/html/2607.16384#S2.Thmtheorem17)then givesα−1/mX∞\(α\),sep⇒Y∞sep,\\alpha^\{\-1/m\}X\_\{\\infty\}^\{\(\\alpha\),\\rm sep\}\\Rightarrow Y\_\{\\infty\}^\{\\rm sep\},where \([2\.19](https://arxiv.org/html/2607.16384#S2.E19)\) becomesdYt=−h0sep\(Yt\)dt\+Σsep1/2dBt\.dY\_\{t\}=\-h\_\{0\}^\{\\rm sep\}\(Y\_\{t\}\)\\,dt\+\\Sigma\_\{\\rm sep\}^\{1/2\}\\,dB\_\{t\}\.The proof of Corollary[2\.17](https://arxiv.org/html/2607.16384#S2.Thmtheorem17)is given in Appendix[C\.2](https://arxiv.org/html/2607.16384#A3.SS2)\.
## 3Applications and numerical experiments
This section illustrates the roles of the local exponentmmand the tail exponentβ\\beta, and reports numerical experiments for the invariant\-law predictions\. Section[3\.1](https://arxiv.org/html/2607.16384#S3.SS1)uses quantile and tail\-risk estimation to interpretmmas a local flatness exponent and illustrates its consequences for local scaling\. Section[3\.2](https://arxiv.org/html/2607.16384#S3.SS2)uses robust and logistic losses to interpretβ\\betaas a global tail\-growth exponent; these examples are locally quadratic withm=2m=2, emphasizing thatβ\\betacontrols large excursions rather than the small\-stepsize normalization\.
### 3\.1Local flatness in quantile and tail\-risk estimation
Quantile and tail\-risk estimation problems provide a statistical interpretation of the local flatness exponent\. The quantile\-loss formulation goes back to Koenker and Bassett\[[26](https://arxiv.org/html/2607.16384#bib.bib26)\], while Knight\[[25](https://arxiv.org/html/2607.16384#bib.bib25)\]studied its asymptotic behavior under general local conditions on the distribution function, including the higher\-order crossings considered below\.
#### Background\.
LetYYbe a scalar response, letFFdenote its distribution function, and lett⋆t\_\{\\star\}be aτ\\tau\-quantile, so thatF\(t⋆\)=τF\(t\_\{\\star\}\)=\\tau\. The population quantile objective is
Hτ\(t\):=𝔼ρτ\(Y−t\),ρτ\(u\):=u\(τ−𝟏\{u<0\}\)\.H\_\{\\tau\}\(t\):=\\mathbb\{E\}\\rho\_\{\\tau\}\(Y\-t\),\\qquad\\rho\_\{\\tau\}\(u\):=u\(\\tau\-\\mathbf\{1\}\_\{\\\{u<0\\\}\}\)\.At continuity points ofFF, the population gradient isF\(t\)−τF\(t\)\-\\tau\. Consequently,
t⋆∈argmint∈ℝHτ\(t\)\.t\_\{\\star\}\\in\\operatorname\*\{argmin\}\_\{t\\in\\mathbb\{R\}\}H\_\{\\tau\}\(t\)\.The local objective behavior is exactly the local crossing behavior ofFFat the target quantile\. Suppose that the following higher\-order crossing condition holds: for somem≥2m\\geq 2and constantsc\+,c−\>0c\_\{\+\},c\_\{\-\}\>0,
F\(t⋆\+u\)−τ=c\+u\+m−1−c−u−m−1\+o\(\|u\|m−1\),u→0,F\(t\_\{\\star\}\+u\)\-\\tau=c\_\{\+\}u\_\{\+\}^\{m\-1\}\-c\_\{\-\}u\_\{\-\}^\{m\-1\}\+o\(\|u\|^\{m\-1\}\),\\qquad u\\to 0,\(3\.1\)whereu\+=max\{u,0\}u\_\{\+\}=\\max\\\{u,0\\\}andu−=max\{−u,0\}u\_\{\-\}=\\max\\\{\-u,0\\\}\. Integrating the population gradient gives
Hτ\(t⋆\+u\)−Hτ\(t⋆\)=c\+mu\+m\+c−mu−m\+o\(\|u\|m\)\.H\_\{\\tau\}\(t\_\{\\star\}\+u\)\-H\_\{\\tau\}\(t\_\{\\star\}\)=\\frac\{c\_\{\+\}\}\{m\}u\_\{\+\}^\{m\}\+\\frac\{c\_\{\-\}\}\{m\}u\_\{\-\}^\{m\}\+o\(\|u\|^\{m\}\)\.The positive\-density case corresponds tom=2m=2\. If the density vanishes at the target quantile like\|u\|m−2\|u\|^\{m\-2\}, then the equationF\(t\)=τF\(t\)=\\taustill has the unique solutiont=t⋆t=t\_\{\\star\}, butF\(t\)−τF\(t\)\-\\tauvanishes at orderm−1m\-1neart⋆t\_\{\\star\}\. As a result, the population risk ismm\-flat\. This setup is closely related to the regularly varying local regime studied by Knight\[[25](https://arxiv.org/html/2607.16384#bib.bib25)\]\.
The Rockafellar–Uryasev variational representation of CVaR\[[41](https://arxiv.org/html/2607.16384#bib.bib41)\]gives an analogous population objective\. For a scalar lossLLand confidence levelτ∈\(0,1\)\\tau\\in\(0,1\), this representation is
CVaRτ\(L\)=mint∈ℝCτ\(t\),Cτ\(t\):=t\+11−τ𝔼\(L−t\)\+\.\\operatorname\{CVaR\}\_\{\\tau\}\(L\)=\\min\_\{t\\in\\mathbb\{R\}\}C\_\{\\tau\}\(t\),\\qquad C\_\{\\tau\}\(t\):=t\+\\frac\{1\}\{1\-\\tau\}\\mathbb\{E\}\(L\-t\)\_\{\+\}\.At continuity points ofFLF\_\{L\},
Cτ′\(t\)=1−11−τℙ\(L\>t\)=FL\(t\)−τ1−τ\.C\_\{\\tau\}^\{\\prime\}\(t\)=1\-\\frac\{1\}\{1\-\\tau\}\\mathbb\{P\}\(L\>t\)=\\frac\{F\_\{L\}\(t\)\-\\tau\}\{1\-\\tau\}\.Hence, ift⋆t\_\{\\star\}is the uniqueτ\\tau\-quantile ofLL, then
argmint∈ℝCτ\(t\)=\{t⋆\}\.\\operatorname\*\{argmin\}\_\{t\\in\\mathbb\{R\}\}C\_\{\\tau\}\(t\)=\\\{t\_\{\\star\}\\\}\.The local expansion in \([3\.1](https://arxiv.org/html/2607.16384#S3.E1)\) then implies
Cτ\(t⋆\+u\)−Cτ\(t⋆\)=c\+m\(1−τ\)u\+m\+c−m\(1−τ\)u−m\+o\(\|u\|m\)\.C\_\{\\tau\}\(t\_\{\\star\}\+u\)\-C\_\{\\tau\}\(t\_\{\\star\}\)=\\frac\{c\_\{\+\}\}\{m\(1\-\\tau\)\}u\_\{\+\}^\{m\}\+\\frac\{c\_\{\-\}\}\{m\(1\-\\tau\)\}u\_\{\-\}^\{m\}\+o\(\|u\|^\{m\}\)\.Thus, combining the variational representation with the local expansion in \([3\.1](https://arxiv.org/html/2607.16384#S3.E1)\) shows that the quantile objective and the variational objective for CVaR have the same local exponentmm\.
#### Invariant law for median estimation\.
To illustrate the local effect of this flatness, consider median estimation with distribution function
Fm\(x\)=12\+12sgn\(x\)\|x\|m−1,\|x\|≤1\.F\_\{m\}\(x\)=\\frac\{1\}\{2\}\+\\frac\{1\}\{2\}\\operatorname\{sgn\}\(x\)\|x\|^\{m\-1\},\\qquad\|x\|\\leq 1\.\(3\.2\)Thus,τ=1/2\\tau=1/2and the median is zero\. The stochastic gradient at the minimizer has nonzero variance, while the restoring force is of order\|x\|m−1\|x\|^\{m\-1\}\. Their balance gives the scaleα1/m\\alpha^\{1/m\}and a nonlinear diffusion whenm\>2m\>2\. The raw pinball score is nonsmooth, so the finite\-state calculation below illustrates this mechanism rather than directly applying Theorem[2\.13](https://arxiv.org/html/2607.16384#S2.Thmtheorem13)\.
The finite\-state example in Figure[1](https://arxiv.org/html/2607.16384#S3.F1)uses a two\-sample lazy version of the recursion
Xn\+1=Xn−α\{12\(𝟏\{Yn\+1,1≤Xn\}\+𝟏\{Yn\+1,2≤Xn\}\)−12\}\.X\_\{n\+1\}=X\_\{n\}\-\\alpha\\left\\\{\\frac\{1\}\{2\}\\bigl\(\\mathbf\{1\}\_\{\\\{Y\_\{n\+1,1\}\\leq X\_\{n\}\\\}\}\+\\mathbf\{1\}\_\{\\\{Y\_\{n\+1,2\}\\leq X\_\{n\}\\\}\}\\bigr\)\-\\frac\{1\}\{2\}\\right\\\}\.\(3\.3\)Forα=2−j\\alpha=2^\{\-j\}, the chain lives on the grid\{−1,−1\+α/2,…,1\}\\\{\-1,\-1\+\\alpha/2,\\ldots,1\\\}\. Ifxk=−1\+kα/2x\_\{k\}=\-1\+k\\alpha/2, then its right and left transition probabilities are
pk=\(1−Fm\(xk\)\)2,qk=Fm\(xk\)2,p\_\{k\}=\(1\-F\_\{m\}\(x\_\{k\}\)\)^\{2\},\\qquad q\_\{k\}=F\_\{m\}\(x\_\{k\}\)^\{2\},and its invariant law is computed exactly from the detailed balance equation
πα\(k\+1\)πα\(k\)=pkqk\+1\.\\frac\{\\pi\_\{\\alpha\}\(k\+1\)\}\{\\pi\_\{\\alpha\}\(k\)\}=\\frac\{p\_\{k\}\}\{q\_\{k\+1\}\}\.The variance of the noise at the minimizer in \([3\.3](https://arxiv.org/html/2607.16384#S3.E3)\) isΣ=1/8\\Sigma=1/8, and the limiting diffusion is
dZt=−12Zt\|Zt\|m−2dt\+ΣdBt\.dZ\_\{t\}=\-\\frac\{1\}\{2\}Z\_\{t\}\|Z\_\{t\}\|^\{m\-2\}\\,dt\+\\sqrt\{\\Sigma\}\\,dB\_\{t\}\.LetZ∞Z\_\{\\infty\}have the invariant law of this diffusion\. Its density is
pm\(z\)=m\(1/\(mΣ\)\)1/m2Γ\(1/m\)exp\(−\|z\|mmΣ\)\.p\_\{m\}\(z\)=\\frac\{m\(1/\(m\\Sigma\)\)^\{1/m\}\}\{2\\Gamma\(1/m\)\}\\exp\\\!\\left\(\-\\frac\{\|z\|^\{m\}\}\{m\\Sigma\}\\right\)\.In particular, form=4m=4,p4\(z\)∝exp\(−2z4\)p\_\{4\}\(z\)\\propto\\exp\(\-2z^\{4\}\), which is not Gaussian\. Figure[1](https://arxiv.org/html/2607.16384#S3.F1)confirms all these theoretical predictions\.
Figure 1:SGD for median estimation\. Panel \(a\) shows theL2L\_\{2\}error for several crossing ordersmm; the fitted slopes match the theoretical exponent1/m1/m\. Panel \(b\) checks the predicted local moment scale: in this modelα−1𝔼\|X∞\(α\)\|m→Σ=1/8\\alpha^\{\-1\}\\mathbb\{E\}\|X\_\{\\infty\}^\{\(\\alpha\)\}\|^\{m\}\\to\\Sigma=1/8\. Panel \(c\) shows the convergence ofα−1/4X∞\(α\)\\alpha^\{\-1/4\}X\_\{\\infty\}^\{\(\\alpha\)\}form=4m=4to the quartic invariant density of the limiting diffusion, together with a Gaussian of the same variance for comparison\.
#### Markovian covariance and unequal exponents\.
A related finite\-state experiment in Figure[2](https://arxiv.org/html/2607.16384#S3.F2)\(a\) verifies the effect of temporal dependence on the asymptotic covariance\. Letρ∈\[0,1\)\\rho\\in\[0,1\), and let\(Sn\)\(S\_\{n\}\)be a stationary two\-state Markov chain on\{−1,1\}\\\{\-1,1\\\}with transition probabilities
ℙ\(Sn\+1=s∣Sn=s\)=1\+ρ2,ℙ\(Sn\+1=−s∣Sn=s\)=1−ρ2,s∈\{−1,1\}\.\\mathbb\{P\}\(S\_\{n\+1\}=s\\mid S\_\{n\}=s\)=\\frac\{1\+\\rho\}\{2\},\\qquad\\mathbb\{P\}\(S\_\{n\+1\}=\-s\\mid S\_\{n\}=s\)=\\frac\{1\-\\rho\}\{2\},\\qquad s\\in\\\{\-1,1\\\}\.Its stationary distribution is uniform on\{−1,1\}\\\{\-1,1\\\}, and𝔼\[SnSn\+k\]=ρk\\mathbb\{E\}\[S\_\{n\}S\_\{n\+k\}\]=\\rho^\{k\}\. Independently, let\(Rn\)\(R\_\{n\}\)be i\.i\.d\. with
ℙ\(Rn≤r\)=rm−1,0≤r≤1,\\mathbb\{P\}\(R\_\{n\}\\leq r\)=r^\{m\-1\},\\quad 0\\leq r\\leq 1,and setYn=SnRnY\_\{n\}=S\_\{n\}R\_\{n\}\. The two\-state chain can be realized on\[−1,1\]\[\-1,1\]by takingΦ\(s,U\)=s\\Phi\(s,U\)=swith probabilityρ\\rho, andΦ\(s,U\)=1\\Phi\(s,U\)=1or−1\-1, each with probability\(1−ρ\)/2\(1\-\\rho\)/2\. Hence𝔼LΦ\(U\)=ρ<1\\mathbb\{E\}L\_\{\\Phi\}\(U\)=\\rho<1, whereLΦ\(U\)L\_\{\\Phi\}\(U\)is the random Lipschitz coefficient from Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1)\. The marginal law ofYnY\_\{n\}is still \([3\.2](https://arxiv.org/html/2607.16384#S3.E2)\), but at the median𝟏\{Yn≤0\}−12=−12Sn\.\\mathbf\{1\}\_\{\\\{Y\_\{n\}\\leq 0\\\}\}\-\\frac\{1\}\{2\}=\-\\frac\{1\}\{2\}S\_\{n\}\.Thus the asymptotic covariance in the Green–Kubo representation \([2\.11](https://arxiv.org/html/2607.16384#S2.E11)\) is
Σ=14\(1\+2∑k≥1ρk\)=14⋅1\+ρ1−ρ\.\\Sigma=\\frac\{1\}\{4\}\\left\(1\+2\\sum\_\{k\\geq 1\}\\rho^\{k\}\\right\)=\\frac\{1\}\{4\}\\cdot\\frac\{1\+\\rho\}\{1\-\\rho\}\.Form=4m=4, the invariant law satisfies𝔼Z∞4=Σ\\mathbb\{E\}Z\_\{\\infty\}^\{4\}=\\Sigma, and hence
α−1𝔼\|X∞\(α\)\|4⟶Σ\.\\alpha^\{\-1\}\\mathbb\{E\}\|X\_\{\\infty\}^\{\(\\alpha\)\}\|^\{4\}\\longrightarrow\\Sigma\.Figure[2](https://arxiv.org/html/2607.16384#S3.F2)\(a\) shows that the rescaled moment approaches the predicted valueΣ=ΣGK\\Sigma=\\Sigma\_\{\\rm GK\}asα↓0\\alpha\\downarrow 0\.
Moreover, we consider the separable case by taking the product of two scalar recursions of the form \([3\.3](https://arxiv.org/html/2607.16384#S3.E3)\), withm1=2m\_\{1\}=2andm2=4m\_\{2\}=4\. Under the common scale associated withm2m\_\{2\}, only the second coordinate has a nonzero limit\. Figure[2](https://arxiv.org/html/2607.16384#S3.F2)\(b\) confirms this conclusion\.
Figure 2:Two consequences of the scaling limit of invariant laws\. Panel \(a\) uses a Markovian sign stream and verifies that the fourth moment constant in them=4m=4quantile example is governed by the asymptotic covariance rather than by the marginal variance\. Panel \(b\) considers a separable two\-coordinate model with exponentsm1=2m\_\{1\}=2andm2=4m\_\{2\}=4; under the commonα−1/4\\alpha^\{\-1/4\}scale, them1=2m\_\{1\}=2coordinate converges to zero while them2=4m\_\{2\}=4coordinate converges to a nondegenerate limit\.
### 3\.2Subquadratic tails for robust and logistic losses
The tail exponentβ\\betadescribes the growth of the objective at infinity\. The generalized Charbonnier family below realizes everyβ∈\[1,2\]\\beta\\in\[1,2\]\[[11](https://arxiv.org/html/2607.16384#bib.bib11),[4](https://arxiv.org/html/2607.16384#bib.bib4)\]; the classical Huber loss is a related piecewise\-defined robust loss with linear tails\[[19](https://arxiv.org/html/2607.16384#bib.bib19)\]\. Under the bounded, nondegenerate, nonseparable model verified below, the population logistic risk has linear coercive growth corresponding toβ=1\\beta=1\[[5](https://arxiv.org/html/2607.16384#bib.bib5)\]\.
#### Robust location estimation\.
Consider a scalar location model\. For1≤β≤21\\leq\\beta\\leq 2, define the generalized Charbonnier loss
ρβ\(u\):=\(1\+u2\)β/2−1β,ψβ\(u\):=ρβ′\(u\)=u\(1\+u2\)β/2−1\.\\rho\_\{\\beta\}\(u\):=\\frac\{\(1\+u^\{2\}\)^\{\\beta/2\}\-1\}\{\\beta\},\\qquad\\psi\_\{\\beta\}\(u\):=\\rho\_\{\\beta\}^\{\\prime\}\(u\)=u\(1\+u^\{2\}\)^\{\\beta/2\-1\}\.The endpointβ=1\\beta=1is the pseudo\-Huber/Charbonnier loss, whileβ=2\\beta=2is ordinary least squares\. In the scalar location model
Y=θ⋆\+ε,Hβ\(θ\):=𝔼ρβ\(θ−Y\),Y=\\theta\_\{\\star\}\+\\varepsilon,\\qquad H\_\{\\beta\}\(\\theta\):=\\mathbb\{E\}\\rho\_\{\\beta\}\(\\theta\-Y\),recenterx=θ−θ⋆x=\\theta\-\\theta\_\{\\star\}\. If the error distribution is symmetric, then the constant\-stepsize SGD becomes
Xn\+1=Xn−αψβ\(Xn−εn\+1\)\.X\_\{n\+1\}=X\_\{n\}\-\\alpha\\psi\_\{\\beta\}\(X\_\{n\}\-\\varepsilon\_\{n\+1\}\)\.The population drift is
hβ\(x\)=𝔼ψβ\(x−ε\)\.h\_\{\\beta\}\(x\)=\\mathbb\{E\}\\psi\_\{\\beta\}\(x\-\\varepsilon\)\.Since
ψβ′\(u\)=\(1\+u2\)β/2−2\{1\+\(β−1\)u2\},\\psi\_\{\\beta\}^\{\\prime\}\(u\)=\(1\+u^\{2\}\)^\{\\beta/2\-2\}\\\{1\+\(\\beta\-1\)u^\{2\}\\\},we have
hβ′\(0\)=𝔼ψβ′\(−ε\)\>0h\_\{\\beta\}^\{\\prime\}\(0\)=\\mathbb\{E\}\\psi\_\{\\beta\}^\{\\prime\}\(\-\\varepsilon\)\>0for every nondegenerate bounded symmetric error\. Thus the minimizer is locally quadratic withm=2m=2\. Consequently the invariant law remains the classical Gaussian limit,
α−1/2X∞\(α\)⇒N\(0,Var\(ψβ\(−ε\)\)2hβ′\(0\)\)\.\\alpha^\{\-1/2\}X\_\{\\infty\}^\{\(\\alpha\)\}\\Rightarrow N\\left\(0,\\frac\{\\operatorname\{Var\}\(\\psi\_\{\\beta\}\(\-\\varepsilon\)\)\}\{2h\_\{\\beta\}^\{\\prime\}\(0\)\}\\right\)\.
The difference from least squares is not local but global\. As\|u\|→∞\|u\|\\to\\infty,
ρβ\(u\)∼\|u\|ββ,ψβ\(u\)∼sgn\(u\)\|u\|β−1\.\\rho\_\{\\beta\}\(u\)\\sim\\frac\{\|u\|^\{\\beta\}\}\{\\beta\},\\qquad\\psi\_\{\\beta\}\(u\)\\sim\\operatorname\{sgn\}\(u\)\|u\|^\{\\beta\-1\}\.Therefore, for bounded errors,
Hβ\(x\)∼\|x\|ββ,hβ\(x\)∼sgn\(x\)\|x\|β−1,\|x\|→∞\.H\_\{\\beta\}\(x\)\\sim\\frac\{\|x\|^\{\\beta\}\}\{\\beta\},\\qquad h\_\{\\beta\}\(x\)\\sim\\operatorname\{sgn\}\(x\)\|x\|^\{\\beta\-1\},\\qquad\|x\|\\to\\infty\.For1≤β<21\\leq\\beta<2, the behavior at infinity of the deterministic flowx˙t=−hβ\(xt\)\\dot\{x\}\_\{t\}=\-h\_\{\\beta\}\(x\_\{t\}\)is described by
ddt\|xt\|2−β∼−\(2−β\),\|xt\|→∞\.\\frac\{d\}\{dt\}\|x\_\{t\}\|^\{2\-\\beta\}\\sim\-\(2\-\\beta\),\\qquad\|x\_\{t\}\|\\to\\infty\.\(3\.4\)This calculation explains why the weight function for subquadratic tails involves\|x\|2−β\|x\|^\{2\-\\beta\}\.
For the numerical illustration, take
ε∼0\.9Unif\[−1,1\]\+0\.1Unif\[−10,10\]\.\\varepsilon\\sim 0\.9\\,\\operatorname\{Unif\}\[\-1,1\]\+0\.1\\,\\operatorname\{Unif\}\[\-10,10\]\.This bounded symmetric mixture satisfies the assumptions used above\. Figure[3](https://arxiv.org/html/2607.16384#S3.F3)reports the corresponding numerical results\.
Figure 3:Robust location estimation with generalized Charbonnier losses\. Panel \(a\) verifies the population tailHβ\(x\)≍\|x\|βH\_\{\\beta\}\(x\)\\asymp\|x\|^\{\\beta\}\. Panel \(b\) verifies the Gaussian scaling limit ofX∞\(α\)/αX\_\{\\infty\}^\{\(\\alpha\)\}/\\sqrt\{\\alpha\}for the pseudo\-Huber loss withβ=1\\beta=1because it remains locally quadratic\. Panel \(c\) shows that the decay of the normalized state𝔼Xt/x0\\mathbb\{E\}X\_\{t\}/x\_\{0\}, starting from a largex0\>0x\_\{0\}\>0, depends on the tail exponent\. Panel \(d\) normalizes the same trajectories by the tail scaling\|x\|2−β\|x\|^\{2\-\\beta\}, showing that they follow the linear prediction in \([3\.4](https://arxiv.org/html/2607.16384#S3.E4)\)\.
#### Logistic regression\.
Finally, we verify theoretically that our results apply to SGD for logistic regression under a correctly specified model with bounded, nondegenerate design and nonvanishing label noise\. Let\(Yn,Zn\)n≥1\(Y\_\{n\},Z\_\{n\}\)\_\{n\\geq 1\}be i\.i\.d\. data with covariates satisfying
‖Zn‖≤R,𝔼\[ZnZn⊤\]⪰λId\\left\\lVert Z\_\{n\}\\right\\rVert\\leq R,\\qquad\\mathbb\{E\}\[Z\_\{n\}Z\_\{n\}^\{\\top\}\]\\succeq\\lambda I\_\{d\}for someR,λ\>0R,\\lambda\>0\. Withσ\(t\):=\(1\+e−t\)−1\\sigma\(t\):=\(1\+e^\{\-t\}\)^\{\-1\}, suppose that
ℙ\(Yn=1∣Zn\)=σ\(Zn⊤θ⋆\),Yn∈\{−1,1\}\.\\mathbb\{P\}\(Y\_\{n\}=1\\mid Z\_\{n\}\)=\\sigma\(Z\_\{n\}^\{\\top\}\\theta\_\{\\star\}\),\\qquad Y\_\{n\}\\in\\\{\-1,1\\\}\.SetSn:=YnZnS\_\{n\}:=Y\_\{n\}Z\_\{n\}\. For the logistic loss
ℓ\(θ;S\):=log\(1\+e−S⊤θ\),\\ell\(\\theta;S\):=\\log\(1\+e^\{\-S^\{\\top\}\\theta\}\),the unique population minimizer isθ⋆\\theta\_\{\\star\}\. After recenteringx=θ−θ⋆x=\\theta\-\\theta\_\{\\star\}, the sample and population gradients are
𝖦\(x,S\):=−S1\+eS⊤\(θ⋆\+x\),h\(x\):=𝔼𝖦\(x,S\)\.\\mathsf\{G\}\(x,S\):=\-\\frac\{S\}\{1\+e^\{S^\{\\top\}\(\\theta\_\{\\star\}\+x\)\}\},\\qquad h\(x\):=\\mathbb\{E\}\\mathsf\{G\}\(x,S\)\.The centered population objective isC2C^\{2\}and convex, withh\(0\)=0h\(0\)=0, and hence satisfies Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H1\)](https://arxiv.org/html/2607.16384#S2.I1.i1); moreover,‖𝖦\(x,S\)‖≤R\\left\\lVert\\mathsf\{G\}\(x,S\)\\right\\rVert\\leq R\.
Since\|Z⊤θ⋆\|≤R‖θ⋆‖\\lvert Z^\{\\top\}\\theta\_\{\\star\}\\rvert\\leq R\\left\\lVert\\theta\_\{\\star\}\\right\\rVert, settingq:=σ\(−R‖θ⋆‖\)q:=\\sigma\(\-R\\left\\lVert\\theta\_\{\\star\}\\right\\rVert\)givesq≤ℙ\(Y=1∣Z\)≤1−q\.q\\leq\\mathbb\{P\}\(Y=1\\mid Z\)\\leq 1\-q\.Hence, for everyv∈𝕊d−1v\\in\\mathbb\{S\}^\{d\-1\},
γ\(v\):=𝔼\[\(−S⊤v\)\+\]≥q𝔼\|Z⊤v\|≥qR𝔼\[\(Z⊤v\)2\]≥qλR\.\\gamma\(v\):=\\mathbb\{E\}\[\(\-S^\{\\top\}v\)\_\{\+\}\]\\geq q\\mathbb\{E\}\|Z^\{\\top\}v\|\\geq\\frac\{q\}\{R\}\\mathbb\{E\}\[\(Z^\{\\top\}v\)^\{2\}\]\\geq\\frac\{q\\lambda\}\{R\}\.Moreover, writingB:=R‖θ⋆‖B:=R\\left\\lVert\\theta\_\{\\star\}\\right\\rVert, one has the uniform bound
supv∈𝕊d−1\|1r⟨rv,h\(rv\)⟩−γ\(v\)\|≤eBer\.\\sup\_\{v\\in\\mathbb\{S\}^\{d\-1\}\}\\left\|\\frac\{1\}\{r\}\\left\\langle rv,h\(rv\)\\right\\rangle\-\\gamma\(v\)\\right\|\\leq\\frac\{e^\{B\}\}\{er\}\.Thus, for all sufficiently largerr,
⟨rv,h\(rv\)⟩≥qλ2Rr,‖h\(rv\)‖≤R,\\left\\langle rv,h\(rv\)\\right\\rangle\\geq\\frac\{q\\lambda\}\{2R\}r,\\qquad\\left\\lVert h\(rv\)\\right\\rVert\\leq R,which verifies Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H3H3\-b\)](https://arxiv.org/html/2607.16384#S2.I1.i3.I1.i2)withβ=1\\beta=1\. The above condition and the nondegenerate design also imply, by a standard covering and concentration argument, that an i\.i\.d\. sample is not linearly separable with probability tending to one as its size grows, for fixeddd\.
The population Hessian atθ⋆\\theta\_\{\\star\}satisfies
A:=𝔼\[σ\(Z⊤θ⋆\)1−σ\(Z⊤θ⋆\)ZZ⊤\]⪰q\(1−q\)λId\.A:=\\mathbb\{E\}\\\!\\left\[\\sigma\(Z^\{\\top\}\\theta\_\{\\star\}\)\{1\-\\sigma\(Z^\{\\top\}\\theta\_\{\\star\}\)\}ZZ^\{\\top\}\\right\]\\succeq q\(1\-q\)\\lambda I\_\{d\}\.Continuity of the population Hessian verifies Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)withm=2m=2, while
h\(x\)=Ax\+o\(‖x‖\),x→0,h\(x\)=Ax\+o\(\\left\\lVert x\\right\\rVert\),\\qquad x\\to 0,verifies Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4)withH0\(x\)=12x⊤AxH\_\{0\}\(x\)=\\tfrac\{1\}\{2\}x^\{\\top\}Ax\.
For each observation,
∇x𝖦\(x,S\)=eS⊤\(θ⋆\+x\)\(1\+eS⊤\(θ⋆\+x\)\)2SS⊤⪯R24Id\.\\nabla\_\{x\}\\mathsf\{G\}\(x,S\)=\\frac\{e^\{S^\{\\top\}\(\\theta\_\{\\star\}\+x\)\}\}\{\(1\+e^\{S^\{\\top\}\(\\theta\_\{\\star\}\+x\)\}\)^\{2\}\}SS^\{\\top\}\\preceq\\frac\{R^\{2\}\}\{4\}I\_\{d\}\.Thus each sample loss is convex with\(R2/4\)\(R^\{2\}/4\)\-Lipschitz gradient, and the Baillon–Haddad theorem\[[6](https://arxiv.org/html/2607.16384#bib.bib6)\]verifies the co\-coercivity part of Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N7\)](https://arxiv.org/html/2607.16384#S2.I3.i7)with anyL𝖦≥max\{1,R2/4\}L\_\{\\mathsf\{G\}\}\\geq\\max\\\{1,R^\{2\}/4\\\}\. Although the displayed Hessian has rank at most one, no samplewise strict curvature is required\. Withg\(x,S\):=𝖦\(x,S\)−h\(x\)g\(x,S\):=\\mathsf\{G\}\(x,S\)\-h\(x\), i\.i\.d\. sampling gives𝔼g\(x,S\)=0\\mathbb\{E\}g\(x,S\)=0and‖g\(x,S\)‖≤2R\\left\\lVert g\(x,S\)\\right\\rVert\\leq 2R\. Thus the mean\-perturbation part of[\(N7\)](https://arxiv.org/html/2607.16384#S2.I3.i7)holds withθ=0\\theta=0, and the tail estimate above verifies Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N4\)](https://arxiv.org/html/2607.16384#S2.I3.i4)\. The uniform bound onggverifies[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5), while centering and boundedness verify Assumption[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)[\(N8\)](https://arxiv.org/html/2607.16384#S2.I6.i8);[\(N6\)](https://arxiv.org/html/2607.16384#S2.I3.i6)is not imposed becausem=2m=2\. Finally, the i\.i\.d\. embedding in Remark[2\.1](https://arxiv.org/html/2607.16384#S2.Thmtheorem1)verifies[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1)–[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3)\. Thus Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)hold withm=2m=2andβ=1\\beta=1, so our main results apply\.
## 4Proofs of the main results
We prove the main results in this section\. In Section[4\.1](https://arxiv.org/html/2607.16384#S4.SS1), we show that for any sufficiently small constant stepsizeα\\alpha, the augmented Markov chain admits a unique invariant law and contracts geometrically in the Wasserstein distance, thereby establishing Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7)\. In Section[4\.2](https://arxiv.org/html/2607.16384#S4.SS2), with the fixed\-α\\alphainvariant law in hand, we derive its scaling limit asα↓0\\alpha\\downarrow 0and prove Theorem[2\.13](https://arxiv.org/html/2607.16384#S2.Thmtheorem13)\. Additional technical proofs are deferred to Appendices[A](https://arxiv.org/html/2607.16384#A1),[B](https://arxiv.org/html/2607.16384#A2), and[C](https://arxiv.org/html/2607.16384#A3)\.
### 4\.1Analysis of constant\-stepsize SGD
At fixedα\\alpha, we use the framework of Qu, Blanchet, and Glynn\[[39](https://arxiv.org/html/2607.16384#bib.bib39)\], where the main task is to prove a one\-step contraction for a certain weight\-induced metric\. More specifically, to prove Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7), it suffices to construct a metricdVα,αd\_\{V\_\{\\alpha\},\\alpha\}induced by a weight functionVαV\_\{\\alpha\}such that
𝔼dVα,α\(FU1\(z\),FU1\(z′\)\)≤\(1−c0αm−1\)dVα,α\(z,z′\)\\mathbb\{E\}d\_\{V\_\{\\alpha\},\\alpha\}\(F\_\{U\_\{1\}\}\(z\),F\_\{U\_\{1\}\}\(z^\{\\prime\}\)\)\\leq\(1\-c\_\{0\}\\alpha^\{m\-1\}\)d\_\{V\_\{\\alpha\},\\alpha\}\(z,z^\{\\prime\}\)\(4\.1\)for some constantc0\>0c\_\{0\}\>0and allz,z′∈𝖹z,z^\{\\prime\}\\in\\mathsf\{Z\}\. Proposition[4\.1](https://arxiv.org/html/2607.16384#S4.Thmtheorem1)below then converts this into the invariant law, uniqueness, and geometric convergence\.
We remark that the technique for establishing the contraction in\[[39](https://arxiv.org/html/2607.16384#bib.bib39)\]cannot be applied directly here\. In short, it controls each realized random map through its local Lipschitz modulus before averaging\. For the SGD maps considered here, a stochastic\-gradient Jacobian may be rank deficient, so the corresponding update map can have local Lipschitz modulus one and therefore may not be a strict contraction\. Instead, we retain the induced metric but prove the desired contraction in expectation via a direction kernel \(cf\. Lemma[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3)\)\.
For a metricddon a spaceEE, we writeW1dW\_\{1\}^\{d\}for the Wasserstein–1 distance induced bydd, namely
W1d\(μ,ν\):=infγ∈Π\(μ,ν\)∫E×Ed\(z,z′\)γ\(dz,dz′\)\.W\_\{1\}^\{d\}\(\\mu,\\nu\):=\\inf\_\{\\gamma\\in\\Pi\(\\mu,\\nu\)\}\\int\_\{E\\times E\}d\(z,z^\{\\prime\}\)\\,\\gamma\(dz,dz^\{\\prime\}\)\.
###### Proposition 4\.1\.
Let\(E,d\)\(E,d\)be a Polish space and letZn\+1=FUn\+1\(Zn\)Z\_\{n\+1\}=F\_\{U\_\{n\+1\}\}\(Z\_\{n\}\)be a random\-map Markov chain with transition kernelPP\. Suppose that there existr∈\(0,1\)r\\in\(0,1\)andz⋆∈Ez\_\{\\star\}\\in Esuch that
𝔼d\(FU\(z\),FU\(z′\)\)≤rd\(z,z′\),z,z′∈E,\\mathbb\{E\}d\(F\_\{U\}\(z\),F\_\{U\}\(z^\{\\prime\}\)\)\\leq rd\(z,z^\{\\prime\}\),\\qquad z,z^\{\\prime\}\\in E,\(4\.2\)and
𝔼d\(FU\(z⋆\),z⋆\)<∞\.\\mathbb\{E\}d\(F\_\{U\}\(z\_\{\\star\}\),z\_\{\\star\}\)<\\infty\.\(4\.3\)Then the chain admits a unique invariant lawπ∈𝒫1\(E,d\)\\pi\\in\\mathcal\{P\}\_\{1\}\(E,d\)\. Moreover, for all probability lawsμ,ν\\mu,\\nufor which the right\-hand side is finite,
W1d\(μPn,νPn\)≤rnW1d\(μ,ν\),n≥0\.W\_\{1\}^\{d\}\(\\mu P^\{n\},\\nu P^\{n\}\)\\leq r^\{n\}W\_\{1\}^\{d\}\(\\mu,\\nu\),\\qquad n\\geq 0\.\(4\.4\)For every bounded functionφ:E→ℝ\\varphi:E\\to\\mathbb\{R\}that is Lipschitz with respect to the metricdd, and every fixedz∈Ez\\in E, one hasPnφ\(z\)→π\(φ\)P^\{n\}\\varphi\(z\)\\to\\pi\(\\varphi\)\.
The proof is a Banach fixed\-point argument for the Markov operator in the metricW1dW\_\{1\}^\{d\}; the details are given in Appendix[A\.1](https://arxiv.org/html/2607.16384#A1.SS1)\.
Before applying Proposition[4\.1](https://arxiv.org/html/2607.16384#S4.Thmtheorem1)to the augmented Markovian mapFUF\_\{U\}, we first record a lemma which converts Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N7\)](https://arxiv.org/html/2607.16384#S2.I3.i7)into samplewise nonexpansiveness and averaged contraction for the stochastic\-gradient Jacobian\. For\(x,ξ\)∈ℝd×Ξ\(x,\\xi\)\\in\\mathbb\{R\}^\{d\}\\times\\Xi, write
A\(x,ξ\):=∇x𝖦\(x,ξ\)=∇2H\(x\)\+∇xg\(x,ξ\)\.A\(x,\\xi\):=\\nabla\_\{x\}\\mathsf\{G\}\(x,\\xi\)=\\nabla^\{2\}H\(x\)\+\\nabla\_\{x\}g\(x,\\xi\)\.
###### Lemma 4\.2\.
Under Assumptions[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H1\)](https://arxiv.org/html/2607.16384#S2.I1.i1)and[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N7\)](https://arxiv.org/html/2607.16384#S2.I3.i7),
‖A\(x,ξ\)‖op≤L𝖦,‖∇xg\(x,ξ\)‖op≤\(2−θ\)L𝖦1−θ,\(x,ξ\)∈ℝd×Ξ\.\\left\\lVert A\(x,\\xi\)\\right\\rVert\_\{\\mathrm\{op\}\}\\leq L\_\{\\mathsf\{G\}\},\\qquad\\left\\lVert\\nabla\_\{x\}g\(x,\\xi\)\\right\\rVert\_\{\\mathrm\{op\}\}\\leq\\frac\{\(2\-\\theta\)L\_\{\\mathsf\{G\}\}\}\{1\-\\theta\},\\qquad\(x,\\xi\)\\in\\mathbb\{R\}^\{d\}\\times\\Xi\.\(4\.5\)Consequently,hhis globallyL𝖦1−θ\\frac\{L\_\{\\mathsf\{G\}\}\}\{1\-\\theta\}\-Lipschitz,g\(⋅,ξ\)g\(\\cdot,\\xi\)is globally2−θ1−θL𝖦\\frac\{2\-\\theta\}\{1\-\\theta\}L\_\{\\mathsf\{G\}\}\-Lipschitz uniformly inξ\\xi, and
‖h\(x\)‖2≤L𝖦1−θ⟨x,h\(x\)⟩,x∈ℝd\.\\left\\lVert h\(x\)\\right\\rVert^\{2\}\\leq\\frac\{L\_\{\\mathsf\{G\}\}\}\{1\-\\theta\}\\left\\langle x,h\(x\)\\right\\rangle,\\qquad x\\in\\mathbb\{R\}^\{d\}\.\(4\.6\)For every0<α≤L𝖦−10<\\alpha\\leq L\_\{\\mathsf\{G\}\}^\{\-1\},
‖\(Id−αA\(x,ξ\)\)v‖≤‖v‖,x∈ℝd,ξ∈Ξ,v∈ℝd\.\\left\\lVert\(I\_\{d\}\-\\alpha A\(x,\\xi\)\)v\\right\\rVert\\leq\\left\\lVert v\\right\\rVert,\\qquad x\\in\\mathbb\{R\}^\{d\},\\ \\xi\\in\\Xi,\\ v\\in\\mathbb\{R\}^\{d\}\.\(4\.7\)Furthermore, for everyx∈ℝdx\\in\\mathbb\{R\}^\{d\},ξ∈Ξ\\xi\\in\\Xi, andv≠0v\\neq 0,
𝔼\[‖\(Id−αA\(x,Φ\(ξ,U1\)\)\)v‖\]≤\(1−α\(1−θ\)2⟨v,∇2H\(x\)v⟩‖v‖2\)‖v‖\.\\mathbb\{E\}\\\!\\left\[\\left\\lVert\(I\_\{d\}\-\\alpha A\(x,\\Phi\(\\xi,U\_\{1\}\)\)\)v\\right\\rVert\\right\]\\leq\\left\(1\-\\frac\{\\alpha\(1\-\\theta\)\}\{2\}\\frac\{\\left\\langle v,\\nabla^\{2\}H\(x\)v\\right\\rangle\}\{\\left\\lVert v\\right\\rVert^\{2\}\}\\right\)\\left\\lVert v\\right\\rVert\.\(4\.8\)
To apply Proposition[4\.1](https://arxiv.org/html/2607.16384#S4.Thmtheorem1), the key is to establish \([4\.1](https://arxiv.org/html/2607.16384#S4.E1)\)\. Towards this end, we need to construct a suitable weight functionVαV\_\{\\alpha\}\. LetRHR\_\{H\}be as in Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)\. For the tail region, define
𝒯β\(x\):=\{1,β=2,exp\{κ\(‖x‖2−β−\(3RH4\)2−β\)\+\},1≤β<2,\\mathcal\{T\}\_\{\\beta\}\(x\):=\\begin\{cases\}1,&\\beta=2,\\\\\[3\.00003pt\] \\exp\\\!\\left\\\{\\kappa\\bigl\(\\left\\lVert x\\right\\rVert^\{2\-\\beta\}\-\(\\tfrac\{3R\_\{H\}\}\{4\}\)^\{2\-\\beta\}\\bigr\)\_\{\+\}\\right\\\},&1\\leq\\beta<2,\\end\{cases\}whereκ\>0\\kappa\>0will be chosen sufficiently small in the subquadratic case\. Near the minimizer, define
ω\(x\)\\displaystyle\\omega\(x\):=\{0,m=2,\(1−\(2‖x‖RH\)m−2\)\+,2<m≤3,\(1−2‖x‖RH\)\+,m\>3,\\displaystyle=δα\\displaystyle\\qquad\\delta\_\{\\alpha\}:=\{0,m=2,κ0α,2<m≤3,κ0αm−2,m\>3,\\displaystyle=where the constantκ0\>0\\kappa\_\{0\}\>0will be chosen in the technical estimates below\. Finally set
Vα\(x,ξ\):=Vα\(x\):=𝒯β\(x\)\+δαω\(x\)\.V\_\{\\alpha\}\(x,\\xi\):=V\_\{\\alpha\}\(x\):=\\mathcal\{T\}\_\{\\beta\}\(x\)\+\\delta\_\{\\alpha\}\\omega\(x\)\.DefinedVαXd\_\{V\_\{\\alpha\}\}^\{X\}to be the metric from Definition[2\.2](https://arxiv.org/html/2607.16384#S2.Thmtheorem2)with the Euclidean base norm and weightVαV\_\{\\alpha\}\.
Regarding the two summands inVαV\_\{\\alpha\}, the term𝒯β\\mathcal\{T\}\_\{\\beta\}controls the tail region: it is constant in the quadratic\-tail caseβ=2\\beta=2, and is exponential in‖x‖2−β\\left\\lVert x\\right\\rVert^\{2\-\\beta\}when1≤β<21\\leq\\beta<2\. The compactly supported functionω\\omegais used whenm\>2m\>2: it provides contraction in a neighborhood of the minimizer\. In the flat casesm\>2m\>2, the coefficientδα\\delta\_\{\\alpha\}is the amplitude of this near\-minimizer correction and is chosen so that its decrease has orderαm−1\\alpha^\{m\-1\}\.
For fixedξ∈Ξ\\xi\\in\\Xi, write
ξ\+:=Φ\(ξ,U1\),fξ\+\(x\):=x−α\{h\(x\)\+g\(x,ξ\+\)\}\.\\xi^\{\+\}:=\\Phi\(\\xi,U\_\{1\}\),\\qquad f\_\{\\xi^\{\+\}\}\(x\):=x\-\\alpha\\\{h\(x\)\+g\(x,\\xi^\{\+\}\)\\\}\.Fore∈𝕊d−1e\\in\\mathbb\{S\}^\{d\-1\}, define the directional contraction factor
𝒟ξ,α\(x,e\):=‖\(Id−αA\(x,Φ\(ξ,U1\)\)\)e‖\.\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\):=\\left\\lVert\(I\_\{d\}\-\\alpha A\(x,\\Phi\(\\xi,U\_\{1\}\)\)\)e\\right\\rVert\.For a nonnegative functionϕ\(x\)\\phi\(x\), define the directional kernel
\(𝒦ξ,αϕ\)\(x;e\):=𝔼\[𝒟ξ,α\(x,e\)ϕ\(fΦ\(ξ,U1\)\(x\)\)\],e∈𝕊d−1\.\(\\mathcal\{K\}\_\{\\xi,\\alpha\}\\phi\)\(x;e\):=\\mathbb\{E\}\\\!\\left\[\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\\,\\phi\\bigl\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\\bigr\)\\right\],\\qquad e\\in\\mathbb\{S\}^\{d\-1\}\.
The next lemma is the main estimate that shows the contraction of the weightVαV\_\{\\alpha\}under the kernel𝒦ξ,α\\mathcal\{K\}\_\{\\xi,\\alpha\}\.
###### Lemma 4\.3\.
Assume Assumptions[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)and[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)\. There exist constants
c0∈\(0,1\),CV\>0,α0∈\(0,1\],c\_\{0\}\\in\(0,1\),\\qquad C\_\{V\}\>0,\\qquad\\alpha\_\{0\}\\in\(0,1\],such that, for every0<α≤α00<\\alpha\\leq\\alpha\_\{0\},ξ∈Ξ\\xi\\in\\Xi, andx∈ℝdx\\in\\mathbb\{R\}^\{d\},
supe∈𝕊d−1\(𝒦ξ,αVα\)\(x;e\)≤\(1−c0αm−1\)Vα\(x\),\\sup\_\{e\\in\\mathbb\{S\}^\{d\-1\}\}\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\\leq\(1\-c\_\{0\}\\alpha^\{m\-1\}\)V\_\{\\alpha\}\(x\),\(4\.9\)and
𝔼\[Vα\(fΦ\(ξ,U1\)\(x\)\)\]≤\(1\+CVα\)Vα\(x\)\.\\mathbb\{E\}\\\!\\left\[V\_\{\\alpha\}\\bigl\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\\bigr\)\\right\]\\leq\(1\+C\_\{V\}\\alpha\)V\_\{\\alpha\}\(x\)\.\(4\.10\)
The next lemma lifts the directional contraction to the induced metric on the augmented space, proving \([4\.1](https://arxiv.org/html/2607.16384#S4.E1)\)\.
###### Lemma 4\.4\.
Assume Assumptions[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)and[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5), and letc0c\_\{0\}be as in Lemma[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3)\. There existsα0∈\(0,1\]\\alpha\_\{0\}\\in\(0,1\]such that \([4\.1](https://arxiv.org/html/2607.16384#S4.E1)\) holds for every0<α≤α00<\\alpha\\leq\\alpha\_\{0\}\.
The final estimate verifies the reference\-point integrability required by Proposition[4\.1](https://arxiv.org/html/2607.16384#S4.Thmtheorem1)\.
###### Lemma 4\.5\.
There existsα0∈\(0,1\]\\alpha\_\{0\}\\in\(0,1\]such that, for every0<α≤α00<\\alpha\\leq\\alpha\_\{0\}, withz⋆=\(0,ξ⋆\)z\_\{\\star\}=\(0,\\xi\_\{\\star\}\)andξ⋆\\xi\_\{\\star\}as in Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3),
𝔼dVα,α\(FU1\(z⋆\),z⋆\)<∞\.\\mathbb\{E\}\\,d\_\{V\_\{\\alpha\},\\alpha\}\\bigl\(F\_\{U\_\{1\}\}\(z\_\{\\star\}\),z\_\{\\star\}\\bigr\)<\\infty\.\(4\.11\)
The proofs of Lemmas[4\.2](https://arxiv.org/html/2607.16384#S4.Thmtheorem2),[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3),[4\.4](https://arxiv.org/html/2607.16384#S4.Thmtheorem4), and[4\.5](https://arxiv.org/html/2607.16384#S4.Thmtheorem5)are given in Appendices[A\.2](https://arxiv.org/html/2607.16384#A1.SS2),[A\.4](https://arxiv.org/html/2607.16384#A1.SS4),[A\.5](https://arxiv.org/html/2607.16384#A1.SS5), and[A\.6](https://arxiv.org/html/2607.16384#A1.SS6), respectively\.
###### Proof of Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7)\.
We first check that\(𝖹,dVα,α\)\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)is Polish\. Since𝖹\\mathsf\{Z\}is closed in a finite\-dimensional Euclidean space,VαV\_\{\\alpha\}is continuous, andVα≥1V\_\{\\alpha\}\\geq 1, the metricdVα,αd\_\{V\_\{\\alpha\},\\alpha\}dominates the augmented base norm \([2\.3](https://arxiv.org/html/2607.16384#S2.E3)\)\. Conversely, on every setKR:=\{z∈𝖹:\|z\|α≤R\},K\_\{R\}:=\\\{z\\in\\mathsf\{Z\}:\|z\|\_\{\\alpha\}\\leq R\\\},continuity ofVαV\_\{\\alpha\}givessupKRVα<∞\\sup\_\{K\_\{R\}\}V\_\{\\alpha\}<\\infty, sodVα,αd\_\{V\_\{\\alpha\},\\alpha\}and the augmented base metric are locally comparable onKRK\_\{R\}\. Hence\(𝖹,dVα,α\)\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)is Polish by elementary general topology\.
Letc0c\_\{0\}be as in Lemma[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3), and chooseα0∈\(0,L𝖦−1\]\\alpha\_\{0\}\\in\(0,L\_\{\\mathsf\{G\}\}^\{\-1\}\]small enough that the conclusions of Lemmas[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3),[4\.4](https://arxiv.org/html/2607.16384#S4.Thmtheorem4), and[4\.5](https://arxiv.org/html/2607.16384#S4.Thmtheorem5)hold\. Fix0<α≤α00<\\alpha\\leq\\alpha\_\{0\}\. Equations \([4\.1](https://arxiv.org/html/2607.16384#S4.E1)\) and \([4\.11](https://arxiv.org/html/2607.16384#S4.E11)\) verify respectively the two conditions \([4\.2](https://arxiv.org/html/2607.16384#S4.E2)\) and \([4\.3](https://arxiv.org/html/2607.16384#S4.E3)\) of Proposition[4\.1](https://arxiv.org/html/2607.16384#S4.Thmtheorem1), with\(E,d\)=\(𝖹,dVα,α\)\(E,d\)=\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\),r=1−c0αm−1r=1\-c\_\{0\}\\alpha^\{m\-1\}, andz⋆z\_\{\\star\}as in Lemma[4\.5](https://arxiv.org/html/2607.16384#S4.Thmtheorem5)\. Proposition[4\.1](https://arxiv.org/html/2607.16384#S4.Thmtheorem1)therefore yields a unique invariant lawπα∈𝒫1\(𝖹,dVα,α\)\\pi\_\{\\alpha\}\\in\\mathcal\{P\}\_\{1\}\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)and gives \([2\.8](https://arxiv.org/html/2607.16384#S2.E8)\) withc=c0c=c\_\{0\}\. ∎
### 4\.2Analysis of scaling limit
Next, we turn to the scaling limit of the invariant lawπα\\pi\_\{\\alpha\}given by Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7)asα↓0\\alpha\\downarrow 0\. For every sufficiently smallα\>0\\alpha\>0, let
\(Xα,ξα\)∼πα,ξα\+:=Φ\(ξα,U1\),Xα\+:=Xα−α\{h\(Xα\)\+g\(Xα,ξα\+\)\},\(X\_\{\\alpha\},\\xi\_\{\\alpha\}\)\\sim\\pi\_\{\\alpha\},\\qquad\\xi\_\{\\alpha\}^\{\+\}:=\\Phi\(\\xi\_\{\\alpha\},U\_\{1\}\),\\qquad X\_\{\\alpha\}^\{\+\}:=X\_\{\\alpha\}\-\\alpha\\\{h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\\},whereU1U\_\{1\}is independent of\(Xα,ξα\)\(X\_\{\\alpha\},\\xi\_\{\\alpha\}\)\. Set
Yα:=α−1/mXα,Yα\+:=α−1/mXα\+,hα\(y\):=α−\(1−1/m\)h\(α1/my\)\.Y\_\{\\alpha\}:=\\alpha^\{\-1/m\}X\_\{\\alpha\},\\qquad Y\_\{\\alpha\}^\{\+\}:=\\alpha^\{\-1/m\}X\_\{\\alpha\}^\{\+\},\\qquad h\_\{\\alpha\}\(y\):=\\alpha^\{\-\(1\-1/m\)\}h\(\\alpha^\{1/m\}y\)\.By stationarity ofπα\\pi\_\{\\alpha\},
\(Yα\+,ξα\+\)=d\(Yα,ξα\),\(Y\_\{\\alpha\}^\{\+\},\\xi\_\{\\alpha\}^\{\+\}\)\\stackrel\{\{\\scriptstyle d\}\}\{\{=\}\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}\),and
Yα\+−Yα=−α2−2/mhα\(Yα\)−α1−1/mg\(α1/mYα,ξα\+\)\.Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}=\-\\alpha^\{2\-2/m\}h\_\{\\alpha\}\(Y\_\{\\alpha\}\)\-\\alpha^\{1\-1/m\}g\(\\alpha^\{1/m\}Y\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\.Thus the drift and noise have ordersα2−2/m\\alpha^\{2\-2/m\}andα1−1/m\\alpha^\{1\-1/m\}, respectively \(recovering the familiarα\\alphaandα\\sqrt\{\\alpha\}scalings form=2m=2in particular\)\. We prove Theorem[2\.13](https://arxiv.org/html/2607.16384#S2.Thmtheorem13)by showing tightness ofYαY\_\{\\alpha\}, identifying every weak subsequential limit through the stationary generator equation, and using uniqueness of the invariant law of the limiting diffusion\.
The following key lemma gives the Poisson decomposition of the Markovian noise and identifies the covariance matrixΣ\\Sigma\. LetQQbe the transition kernel of the driving chain\(ξn\)n≥0\(\\xi\_\{n\}\)\_\{n\\geq 0\},
Qφ\(ξ\):=𝔼\[φ\(Φ\(ξ,U1\)\)\]\.Q\\varphi\(\\xi\):=\\mathbb\{E\}\[\\varphi\(\\Phi\(\\xi,U\_\{1\}\)\)\]\.
###### Lemma 4\.6\.
Under Assumptions[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11), forx∈ℝdx\\in\\mathbb\{R\}^\{d\}, writegx\(ξ\):=g\(x,ξ\)g\_\{x\}\(\\xi\):=g\(x,\\xi\)\. Then the series
χx\(ξ\):=∑k=1∞Qkgx\(ξ\),x∈ℝd,ξ∈Ξ,\\chi\_\{x\}\(\\xi\):=\\sum\_\{k=1\}^\{\\infty\}Q^\{k\}g\_\{x\}\(\\xi\),\\qquad x\\in\\mathbb\{R\}^\{d\},\\ \\xi\\in\\Xi,converges absolutely and defines a centered solution of the Poisson equation
χx−Qχx=Qgx\.\\chi\_\{x\}\-Q\\chi\_\{x\}=Qg\_\{x\}\.\(4\.12\)Define
Dx\(ξ,u\):=gx\(Φ\(ξ,u\)\)\+χx\(Φ\(ξ,u\)\)−χx\(ξ\)\.D\_\{x\}\(\\xi,u\):=g\_\{x\}\(\\Phi\(\\xi,u\)\)\+\\chi\_\{x\}\(\\Phi\(\\xi,u\)\)\-\\chi\_\{x\}\(\\xi\)\.Then
𝔼\[Dx\(ξ,U1\)∣ξ\]=0,\\mathbb\{E\}\[D\_\{x\}\(\\xi,U\_\{1\}\)\\mid\\xi\]=0,\(4\.13\)and, withξ\+=Φ\(ξ,u\)\\xi^\{\+\}=\\Phi\(\\xi,u\),
gx\(ξ\+\)=Dx\(ξ,u\)\+χx\(ξ\)−χx\(ξ\+\)\.g\_\{x\}\(\\xi^\{\+\}\)=D\_\{x\}\(\\xi,u\)\+\\chi\_\{x\}\(\\xi\)\-\\chi\_\{x\}\(\\xi^\{\+\}\)\.\(4\.14\)The series in \([2\.11](https://arxiv.org/html/2607.16384#S2.E11)\) converges absolutely, and the resulting matrix satisfies
Σ=𝔼ξ∼πΞ𝔼\[D0\(ξ,U1\)D0\(ξ,U1\)⊤\]⪰0,D0\(ξ,u\):=Dx\(ξ,u\)\|x=0\.\\Sigma=\\mathbb\{E\}\_\{\\xi\\sim\\pi\_\{\\Xi\}\}\\mathbb\{E\}\\\!\\bigl\[D\_\{0\}\(\\xi,U\_\{1\}\)D\_\{0\}\(\\xi,U\_\{1\}\)^\{\\top\}\\bigr\]\\succeq 0,\\qquad D\_\{0\}\(\\xi,u\):=D\_\{x\}\(\\xi,u\)\\big\|\_\{x=0\}\.
The next lemma gives moment estimates forXαX\_\{\\alpha\}which will imply tightness ofYα=α−1/mXαY\_\{\\alpha\}=\\alpha^\{\-1/m\}X\_\{\\alpha\}\.
###### Lemma 4\.7\.
Assume Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. There exist constantsC,α0\>0C,\\alpha\_\{0\}\>0such that, for every0<α≤α00<\\alpha\\leq\\alpha\_\{0\},
𝔼⟨Xα,h\(Xα\)⟩\\displaystyle\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle≤Cα,\\displaystyle\\leq C\\alpha,\(4\.15\)𝔼\[‖Xα‖m𝟏\{‖Xα‖≤1\}\]\\displaystyle\\mathbb\{E\}\\bigl\[\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{m\}\\mathbf\{1\}\_\{\\\{\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\leq 1\\\}\}\\bigr\]≤Cα,\\displaystyle\\leq C\\alpha,\(4\.16\)ℙ\(‖Xα‖≥1\)\\displaystyle\\mathbb\{P\}\(\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\geq 1\)≤Cα,\\displaystyle\\leq C\\alpha,\(4\.17\)𝔼‖h\(Xα\)\+g\(Xα,ξα\+\)‖2\\displaystyle\\mathbb\{E\}\\left\\lVert h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}≤C\.\\displaystyle\\leq C\.\(4\.18\)
The following lemma controls the dependence betweenYαY\_\{\\alpha\}andξα\\xi\_\{\\alpha\}\.
###### Lemma 4\.8\.
Assume Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. Letψ:ℝd→ℝ\\psi:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}be bounded and globally Lipschitz, and letf:Ξ→ℝf:\\Xi\\to\\mathbb\{R\}be bounded, Lipschitz, and centered underπΞ\\pi\_\{\\Xi\}\. Then
\|𝔼\[ψ\(Yα\)f\(ξα\)\]\|≤Cψ,fα1−1/m,\\left\\lvert\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)f\(\\xi\_\{\\alpha\}\)\]\\right\\rvert\\leq C\_\{\\psi,f\}\\alpha^\{1\-1/m\},whereCψ,f\>0C\_\{\\psi,f\}\>0is independent ofα\\alpha\. In particular,
𝔼\[ψ\(Yα\)f\(ξα\)\]→0\.\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)f\(\\xi\_\{\\alpha\}\)\]\\to 0\.
The next lemma identifies every weak subsequential limit ofYαY\_\{\\alpha\}\.
###### Lemma 4\.9\.
Assume Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. Letαk↓0\\alpha\_\{k\}\\downarrow 0and suppose
Yαk⇒ν\.Y\_\{\\alpha\_\{k\}\}\\Rightarrow\\nu\.Then, for everyφ∈Cc3\(ℝd\)\\varphi\\in C\_\{c\}^\{3\}\(\\mathbb\{R\}^\{d\}\),
∫ℝd\[−⟨h0\(y\),∇φ\(y\)⟩\+12tr\(Σ∇2φ\(y\)\)\]ν\(dy\)=0\.\\int\_\{\\mathbb\{R\}^\{d\}\}\\left\[\-\\left\\langle h\_\{0\}\(y\),\\nabla\\varphi\(y\)\\right\\rangle\+\\frac\{1\}\{2\}\\operatorname\{tr\}\\bigl\(\\Sigma\\nabla^\{2\}\\varphi\(y\)\\bigr\)\\right\]\\,\\nu\(dy\)=0\.\(4\.19\)
Finally, Equation \([4\.19](https://arxiv.org/html/2607.16384#S4.E19)\) identifies a unique limiting law\.
###### Lemma 4\.10\.
Assume Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10), and letΣ⪰0\\Sigma\\succeq 0\. The \(possibly degenerate\) diffusion
dYt=−h0\(Yt\)dt\+Σ1/2dBtdY\_\{t\}=\-h\_\{0\}\(Y\_\{t\}\)\\,dt\+\\Sigma^\{1/2\}\\,dB\_\{t\}has a unique invariant law, denoted byν∞\\nu\_\{\\infty\}\. Moreover, any probability lawν\\nusatisfying \([4\.19](https://arxiv.org/html/2607.16384#S4.E19)\) for everyφ∈Cc3\(ℝd\)\\varphi\\in C\_\{c\}^\{3\}\(\\mathbb\{R\}^\{d\}\)equalsν∞\\nu\_\{\\infty\}\.
The proofs of Lemmas[4\.6](https://arxiv.org/html/2607.16384#S4.Thmtheorem6),[4\.7](https://arxiv.org/html/2607.16384#S4.Thmtheorem7),[4\.8](https://arxiv.org/html/2607.16384#S4.Thmtheorem8),[4\.9](https://arxiv.org/html/2607.16384#S4.Thmtheorem9), and[4\.10](https://arxiv.org/html/2607.16384#S4.Thmtheorem10)are given in Appendices[B\.3](https://arxiv.org/html/2607.16384#A2.SS3),[B\.4](https://arxiv.org/html/2607.16384#A2.SS4),[B\.5](https://arxiv.org/html/2607.16384#A2.SS5),[B\.6](https://arxiv.org/html/2607.16384#A2.SS6), and[B\.7](https://arxiv.org/html/2607.16384#A2.SS7), respectively\.
###### Proof of Theorem[2\.13](https://arxiv.org/html/2607.16384#S2.Thmtheorem13)\.
FixR≥1R\\geq 1\. For all sufficiently smallα\\alpha, we haveRα1/m≤1R\\alpha^\{1/m\}\\leq 1, and Lemma[4\.7](https://arxiv.org/html/2607.16384#S4.Thmtheorem7)gives
ℙ\(‖Yα‖\>R\)\\displaystyle\\mathbb\{P\}\(\\left\\lVert Y\_\{\\alpha\}\\right\\rVert\>R\)=ℙ\(‖Xα‖\>Rα1/m\)\\displaystyle=\\mathbb\{P\}\(\\left\\lVert X\_\{\\alpha\}\\right\\rVert\>R\\alpha^\{1/m\}\)≤ℙ\(Rα1/m<‖Xα‖≤1\)\+ℙ\(‖Xα‖\>1\)\\displaystyle\\leq\\mathbb\{P\}\(R\\alpha^\{1/m\}<\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\leq 1\)\+\\mathbb\{P\}\(\\left\\lVert X\_\{\\alpha\}\\right\\rVert\>1\)≤𝔼\[‖Xα‖m𝟏\{‖Xα‖≤1\}\]Rmα\+ℙ\(‖Xα‖\>1\)\\displaystyle\\leq\\frac\{\\mathbb\{E\}\[\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{m\}\\mathbf\{1\}\_\{\\\{\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\leq 1\\\}\}\]\}\{R^\{m\}\\alpha\}\+\\mathbb\{P\}\(\\left\\lVert X\_\{\\alpha\}\\right\\rVert\>1\)≤CR−m\+Cα,\\displaystyle\\leq CR^\{\-m\}\+C\\alpha,where the last line uses \([4\.16](https://arxiv.org/html/2607.16384#S4.E16)\) and \([4\.17](https://arxiv.org/html/2607.16384#S4.E17)\)\. Hence
lim supα↓0ℙ\(‖Yα‖\>R\)≤CR−m\.\\limsup\_\{\\alpha\\downarrow 0\}\\mathbb\{P\}\(\\left\\lVert Y\_\{\\alpha\}\\right\\rVert\>R\)\\leq CR^\{\-m\}\.LettingR→∞R\\to\\inftyshows that the family\{Yα\}α\>0\\\{Y\_\{\\alpha\}\\\}\_\{\\alpha\>0\}is tight asα↓0\\alpha\\downarrow 0\.
We next identify all possible subsequential limits\. Let\(αk\)k≥1\(\\alpha\_\{k\}\)\_\{k\\geq 1\}be an arbitrary sequence satisfyingαk↓0\\alpha\_\{k\}\\downarrow 0\. By tightness, there exist integers1≤k1<k2<⋯1\\leq k\_\{1\}<k\_\{2\}<\\cdotsand a probability lawν\\nuonℝd\\mathbb\{R\}^\{d\}such thatYαkj⇒νY\_\{\\alpha\_\{k\_\{j\}\}\}\\Rightarrow\\nuasj→∞j\\to\\infty\. Sinceαkj↓0\\alpha\_\{k\_\{j\}\}\\downarrow 0, we may apply Lemma[4\.9](https://arxiv.org/html/2607.16384#S4.Thmtheorem9)to the sequence\(αkj\)j≥1\(\\alpha\_\{k\_\{j\}\}\)\_\{j\\geq 1\}\. It follows thatν\\nusatisfies the stationary generator equation \([4\.19](https://arxiv.org/html/2607.16384#S4.E19)\)\. Therefore, Lemma[4\.10](https://arxiv.org/html/2607.16384#S4.Thmtheorem10)implies thatν=ν∞,\\nu=\\nu\_\{\\infty\},whereν∞\\nu\_\{\\infty\}is the unique invariant law of the limiting diffusion \([2\.12](https://arxiv.org/html/2607.16384#S2.E12)\)\. Thus,Yαkj⇒ν∞\.Y\_\{\\alpha\_\{k\_\{j\}\}\}\\Rightarrow\\nu\_\{\\infty\}\.
The same argument applies to every subsequence of the original sequence\(αk\)k≥1\(\\alpha\_\{k\}\)\_\{k\\geq 1\}: every such subsequence has a further subsequence along which the corresponding random variables converge weakly toν∞\\nu\_\{\\infty\}\. As a result, the entire sequence satisfiesYαk⇒ν∞Y\_\{\\alpha\_\{k\}\}\\Rightarrow\\nu\_\{\\infty\}ask→∞k\\to\\infty\. Because the sequence\(αk\)k≥1\(\\alpha\_\{k\}\)\_\{k\\geq 1\}was arbitrary, we conclude thatYα=α−1/mXα⇒Y∞Y\_\{\\alpha\}=\\alpha^\{\-1/m\}X\_\{\\alpha\}\\Rightarrow Y\_\{\\infty\}asα↓0\\alpha\\downarrow 0andY∞∼ν∞Y\_\{\\infty\}\\sim\\nu\_\{\\infty\}\. ∎
## 5Conclusion
This paper develops an invariant\-law theory for constant\-stepsize SGD beyond the strongly convex setting\. For each sufficiently small constant stepsize, we prove geometric convergence to a unique invariant law in a Wasserstein distance induced by a metric adapted to two geometric features of the objective: anmm\-flat near\-minimizer region and aβ\\beta\-subquadratic tail\. We then identify the small\-stepsize scaling limit of the invariant law as the solution of \([2\.12](https://arxiv.org/html/2607.16384#S2.E12)\)\. In particular, the classical Gaussian limit is recovered whenm=2m=2, while flatter objectives can produce larger and generally non\-Gaussian stationary fluctuations\. Moreover, Markovian dependence in gradient noise enters through the asymptotic covarianceΣ\\Sigma\.
The main conceptual point is that local and global features play different roles\. The exponentmmdetermines the scale of the invariant law and the scaling limit, whereasβ\\betadetermines the weight function needed to control excursions\. This separation is reflected in the examples: quantile and tail\-risk estimation illustrate nonquadratic local flatness, while smooth robust and logistic losses illustrate subquadratic tail growth\.
Several directions remain open\. First, the separable extension developed here allows different coordinate flatness exponents, but the genuinely nonseparable anisotropic case remains to be understood\. If different coordinate groups have natural scalessi\(α\)=α1/mis\_\{i\}\(\\alpha\)=\\alpha^\{1/m\_\{i\}\}, the anisotropically rescaled generator contains several effective time scales; a natural approach is multi\-time\-scale stochastic averaging\[[23](https://arxiv.org/html/2607.16384#bib.bib23),[8](https://arxiv.org/html/2607.16384#bib.bib8),[27](https://arxiv.org/html/2607.16384#bib.bib27),[22](https://arxiv.org/html/2607.16384#bib.bib22)\]\. Second, quantitative convergence rates for the scaling limit of invariant laws may be accessible through Stein’s method for the limiting diffusion, based on the solution of the Poisson equation\[[3](https://arxiv.org/html/2607.16384#bib.bib3),[10](https://arxiv.org/html/2607.16384#bib.bib10),[34](https://arxiv.org/html/2607.16384#bib.bib34),[35](https://arxiv.org/html/2607.16384#bib.bib35),[36](https://arxiv.org/html/2607.16384#bib.bib36)\]\. Other important extensions include replacing the one\-step contraction in geometric ergodicity by a multi\-step contraction, which would cover Markovian data streams for which contraction only emerges over several transitions, and relaxing the globally contractive driving\-chain dynamics toward Harris\-type or mixing\-based Markovian noise\[[31](https://arxiv.org/html/2607.16384#bib.bib31),[14](https://arxiv.org/html/2607.16384#bib.bib14)\]\.
## Acknowledgments
Cheng Mao was supported in part by NSF CAREER Award 2338062\. Debankur Mukherjee was partially supported by the NSF grant CPS\-2240982\.
## Appendix AConstant\-stepsize contraction
This appendix proves the metric fixed\-point principle and the estimates used in Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7), followed by the projected Wasserstein–1 corollary\. Constants denoted byC,cC,cmay change from line to line but are independent ofα\\alpha,xx, and the current valueξ\\xiof the driving chain\. We decrease the small\-stepsize threshold when needed\.
### A\.1Fixed\-point principle
###### Proof of Proposition[4\.1](https://arxiv.org/html/2607.16384#S4.Thmtheorem1)\.
The assumptions imply thatPPmaps𝒫1\(E,d\)\\mathcal\{P\}\_\{1\}\(E,d\)into itself, since𝔼d\(FU\(z\),z⋆\)≤rd\(z,z⋆\)\+𝔼d\(FU\(z⋆\),z⋆\)\\mathbb\{E\}d\(F\_\{U\}\(z\),z\_\{\\star\}\)\\leq rd\(z,z\_\{\\star\}\)\+\\mathbb\{E\}d\(F\_\{U\}\(z\_\{\\star\}\),z\_\{\\star\}\)\. The one\-step estimate \([4\.2](https://arxiv.org/html/2607.16384#S4.E2)\) iterates under the synchronous coupling and gives \([4\.4](https://arxiv.org/html/2607.16384#S4.E4)\)\. Letμ0=δz⋆\\mu\_\{0\}=\\delta\_\{z\_\{\\star\}\}andμn=μ0Pn\\mu\_\{n\}=\\mu\_\{0\}P^\{n\}\. By \([4\.3](https://arxiv.org/html/2607.16384#S4.E3)\),W1d\(μ1,μ0\)<∞W\_\{1\}^\{d\}\(\\mu\_\{1\},\\mu\_\{0\}\)<\\infty, and henceW1d\(μk\+1,μk\)≤rkW1d\(μ1,μ0\)W\_\{1\}^\{d\}\(\\mu\_\{k\+1\},\\mu\_\{k\}\)\\leq r^\{k\}W\_\{1\}^\{d\}\(\\mu\_\{1\},\\mu\_\{0\}\)\. Thus\(μn\)\(\\mu\_\{n\}\)is Cauchy in\(𝒫1\(E,d\),W1d\)\(\\mathcal\{P\}\_\{1\}\(E,d\),W\_\{1\}^\{d\}\)\. Let its limit beπ\\pi\. Applying the contraction toμn\\mu\_\{n\}andπ\\pishowsμn\+1→πP\\mu\_\{n\+1\}\\to\\pi P, whileμn\+1→π\\mu\_\{n\+1\}\\to\\pi, soπP=π\\pi P=\\pi\.
For fixedz∈Ez\\in E, takingμ=δz\\mu=\\delta\_\{z\}andν=π\\nu=\\piin \([4\.4](https://arxiv.org/html/2607.16384#S4.E4)\) givesW1d\(δzPn,π\)≤rnW1d\(δz,π\)→0W\_\{1\}^\{d\}\(\\delta\_\{z\}P^\{n\},\\pi\)\\leq r^\{n\}W\_\{1\}^\{d\}\(\\delta\_\{z\},\\pi\)\\to 0; the right\-hand side is finite becauseπ∈𝒫1\(E,d\)\\pi\\in\\mathcal\{P\}\_\{1\}\(E,d\)\. Consequently, every boundeddd\-Lipschitzφ\\varphisatisfies\|Pnφ\(z\)−π\(φ\)\|≤Lipd\(φ\)W1d\(δzPn,π\)→0\|P^\{n\}\\varphi\(z\)\-\\pi\(\\varphi\)\|\\leq\\operatorname\{Lip\}\_\{d\}\(\\varphi\)W\_\{1\}^\{d\}\(\\delta\_\{z\}P^\{n\},\\pi\)\\to 0\.
Letπ~\\widetilde\{\\pi\}be any Borel invariant probability law\. By the pointwise convergence, invariance, and dominated convergence,π~\(φ\)=limnπ~\(Pnφ\)=π\(φ\)\\widetilde\{\\pi\}\(\\varphi\)=\\lim\_\{n\}\\widetilde\{\\pi\}\(P^\{n\}\\varphi\)=\\pi\(\\varphi\)\. Bounded Lipschitz functions determine Borel probability measures on the Polish space\(E,d\)\(E,d\), soπ~=π\\widetilde\{\\pi\}=\\pi\. ∎
### A\.2Proof of Lemma[4\.2](https://arxiv.org/html/2607.16384#S4.Thmtheorem2)
###### Proof of Lemma[4\.2](https://arxiv.org/html/2607.16384#S4.Thmtheorem2)\.
PutAs\(x,ξ\):=\(A\(x,ξ\)\+A\(x,ξ\)⊤\)/2A\_\{\\rm s\}\(x,\\xi\):=\(A\(x,\\xi\)\+A\(x,\\xi\)^\{\\top\}\)/2\. Fixξ∈Ξ\\xi\\in\\Xi\. Applying Cauchy–Schwarz to \([2\.6](https://arxiv.org/html/2607.16384#S2.E6)\) gives‖𝖦\(x,ξ\)−𝖦\(y,ξ\)‖≤L𝖦‖x−y‖\\left\\lVert\\mathsf\{G\}\(x,\\xi\)\-\\mathsf\{G\}\(y,\\xi\)\\right\\rVert\\leq L\_\{\\mathsf\{G\}\}\\left\\lVert x\-y\\right\\rVertforx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\}; the conclusion is immediate when the left\-hand difference vanishes and otherwise follows by division by its norm\. SinceA\(x,ξ\)=∇x𝖦\(x,ξ\)A\(x,\\xi\)=\\nabla\_\{x\}\\mathsf\{G\}\(x,\\xi\), this proves the first bound in \([4\.5](https://arxiv.org/html/2607.16384#S4.E5)\)\.
Next fixx,v∈ℝdx,v\\in\\mathbb\{R\}^\{d\}\. Apply \([2\.6](https://arxiv.org/html/2607.16384#S2.E6)\) tox\+tvx\+tvandxx, divide byt2t^\{2\}, and lett→0t\\to 0\. Then
‖A\(x,ξ\)v‖2≤L𝖦⟨v,A\(x,ξ\)v⟩=L𝖦⟨v,As\(x,ξ\)v⟩,0⪯As\(x,ξ\)⪯L𝖦Id\.\\left\\lVert A\(x,\\xi\)v\\right\\rVert^\{2\}\\leq L\_\{\\mathsf\{G\}\}\\left\\langle v,A\(x,\\xi\)v\\right\\rangle=L\_\{\\mathsf\{G\}\}\\left\\langle v,A\_\{\\rm s\}\(x,\\xi\)v\\right\\rangle,\\qquad 0\\preceq A\_\{\\rm s\}\(x,\\xi\)\\preceq L\_\{\\mathsf\{G\}\}I\_\{d\}\.Here the first inequality impliesAs\(x,ξ\)⪰0A\_\{\\rm s\}\(x,\\xi\)\\succeq 0, and the second also uses‖As‖op≤‖A‖op\\left\\lVert A\_\{\\rm s\}\\right\\rVert\_\{\\mathrm\{op\}\}\\leq\\left\\lVert A\\right\\rVert\_\{\\mathrm\{op\}\}\.
Since𝖦=h\+g\\mathsf\{G\}=h\+g, \([2\.7](https://arxiv.org/html/2607.16384#S2.E7)\) implies
⟨x−y,𝔼\[𝖦\(x,Φ\(ξ,U1\)\)−𝖦\(y,Φ\(ξ,U1\)\)\]⟩≥\(1−θ\)⟨x−y,h\(x\)−h\(y\)⟩\.\\left\\langle x\-y,\\mathbb\{E\}\[\\mathsf\{G\}\(x,\\Phi\(\\xi,U\_\{1\}\)\)\-\\mathsf\{G\}\(y,\\Phi\(\\xi,U\_\{1\}\)\)\]\\right\\rangle\\geq\(1\-\\theta\)\\left\\langle x\-y,h\(x\)\-h\(y\)\\right\\rangle\.Apply this inequality tox\+tvx\+tvandxx, divide byt2t^\{2\}, and lett→0t\\to 0\. Fort≠0t\\neq 0, the preceding Lipschitz bound gives‖\[𝖦\(x\+tv,Φ\(ξ,U1\)\)−𝖦\(x,Φ\(ξ,U1\)\)\]/t‖≤L𝖦‖v‖\\left\\lVert\[\\mathsf\{G\}\(x\+tv,\\Phi\(\\xi,U\_\{1\}\)\)\-\\mathsf\{G\}\(x,\\Phi\(\\xi,U\_\{1\}\)\)\]/t\\right\\rVert\\leq L\_\{\\mathsf\{G\}\}\\left\\lVert v\\right\\rVert, so dominated convergence yields
𝔼\[As\(x,Φ\(ξ,U1\)\)\]⪰\(1−θ\)∇2H\(x\),0⪯∇2H\(x\)⪯L𝖦1−θId\.\\mathbb\{E\}\[A\_\{\\rm s\}\(x,\\Phi\(\\xi,U\_\{1\}\)\)\]\\succeq\(1\-\\theta\)\\nabla^\{2\}H\(x\),\\qquad 0\\preceq\\nabla^\{2\}H\(x\)\\preceq\\frac\{L\_\{\\mathsf\{G\}\}\}\{1\-\\theta\}I\_\{d\}\.Since∇xg\(x,ξ\)=A\(x,ξ\)−∇2H\(x\)\\nabla\_\{x\}g\(x,\\xi\)=A\(x,\\xi\)\-\\nabla^\{2\}H\(x\), the triangle inequality proves the second bound in \([4\.5](https://arxiv.org/html/2607.16384#S4.E5)\)\. Integrating the two Jacobian bounds along line segments yields the stated global Lipschitz estimates\. Finally, \([4\.6](https://arxiv.org/html/2607.16384#S4.E6)\) is the Baillon–Haddad inequality for the convexL𝖦/\(1−θ\)L\_\{\\mathsf\{G\}\}/\(1\-\\theta\)\-smooth functionHH, applied withh\(0\)=0h\(0\)=0\.
It remains to prove the two contraction bounds\. WriteA=A\(x,ξ\)A=A\(x,\\xi\)andAs=As\(x,ξ\)A\_\{\\rm s\}=A\_\{\\rm s\}\(x,\\xi\)\. SinceαL𝖦≤1\\alpha L\_\{\\mathsf\{G\}\}\\leq 1,
‖\(Id−αA\)v‖2\\displaystyle\\left\\lVert\(I\_\{d\}\-\\alpha A\)v\\right\\rVert^\{2\}=‖v‖2−2α⟨v,Asv⟩\+α2‖Av‖2\\displaystyle=\\left\\lVert v\\right\\rVert^\{2\}\-2\\alpha\\left\\langle v,A\_\{\\rm s\}v\\right\\rangle\+\\alpha^\{2\}\\left\\lVert Av\\right\\rVert^\{2\}≤‖v‖2−\(2α−α2L𝖦\)⟨v,Asv⟩≤‖v‖2−α⟨v,Asv⟩≤‖v‖2\.\\displaystyle\\leq\\left\\lVert v\\right\\rVert^\{2\}\-\(2\\alpha\-\\alpha^\{2\}L\_\{\\mathsf\{G\}\}\)\\left\\langle v,A\_\{\\rm s\}v\\right\\rangle\\leq\\left\\lVert v\\right\\rVert^\{2\}\-\\alpha\\left\\langle v,A\_\{\\rm s\}v\\right\\rangle\\leq\\left\\lVert v\\right\\rVert^\{2\}\.This proves \([4\.7](https://arxiv.org/html/2607.16384#S4.E7)\)\. Ifv≠0v\\neq 0, then
0\\displaystyle 0≤α⟨v,Asv⟩‖v‖2≤α‖As‖op≤α‖A‖op≤1,\\displaystyle\\leq\\alpha\\frac\{\\left\\langle v,A\_\{\\rm s\}v\\right\\rangle\}\{\\left\\lVert v\\right\\rVert^\{2\}\}\\leq\\alpha\\left\\lVert A\_\{\\rm s\}\\right\\rVert\_\{\\mathrm\{op\}\}\\leq\\alpha\\left\\lVert A\\right\\rVert\_\{\\mathrm\{op\}\}\\leq 1,‖\(Id−αA\)v‖\\displaystyle\\left\\lVert\(I\_\{d\}\-\\alpha A\)v\\right\\rVert≤\(1−α2⟨v,Asv⟩‖v‖2\)‖v‖,\\displaystyle\\leq\\left\(1\-\\frac\{\\alpha\}\{2\}\\frac\{\\left\\langle v,A\_\{\\rm s\}v\\right\\rangle\}\{\\left\\lVert v\\right\\rVert^\{2\}\}\\right\)\\left\\lVert v\\right\\rVert,where the second inequality uses1−t≤1−t/2\\sqrt\{1\-t\}\\leq 1\-t/2\. Replacingξ\\xibyΦ\(ξ,U1\)\\Phi\(\\xi,U\_\{1\}\), taking expectation withvvfixed, and using the averaged matrix inequality above proves \([4\.8](https://arxiv.org/html/2607.16384#S4.E8)\)\. ∎
### A\.3Preparatory estimates
###### Lemma A\.1\.
Under Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3), there existsCh\>0C\_\{h\}\>0such that‖h\(x\)‖≤Ch‖x‖m−1\\left\\lVert h\(x\)\\right\\rVert\\leq C\_\{h\}\\left\\lVert x\\right\\rVert^\{m\-1\}whenever‖x‖≤RH/2\\left\\lVert x\\right\\rVert\\leq R\_\{H\}/2\.
###### Proof\.
This is the second inequality in Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)\. ∎
###### Lemma A\.2\.
Assumem\>2m\>2, and define
Gξ\\displaystyle G\_\{\\xi\}:=‖g\(0,Φ\(ξ,U1\)\)‖,\\displaystyle=\\left\\lVert g\(0,\\Phi\(\\xi,U\_\{1\}\)\)\\right\\rVert,a¯α\\displaystyle\\bar\{a\}\_\{\\alpha\}:=infξ∈Ξ𝔼min\{αGξ,1\},\\displaystyle=\\inf\_\{\\xi\\in\\Xi\}\\mathbb\{E\}\\min\\\{\\alpha G\_\{\\xi\},1\\\},a¯α,p\\displaystyle\\bar\{a\}\_\{\\alpha,p\}:=infξ∈Ξ𝔼\[min\{\(αGξ\)p,1\}\],\\displaystyle=\\inf\_\{\\xi\\in\\Xi\}\\mathbb\{E\}\\bigl\[\\min\\\{\(\\alpha G\_\{\\xi\}\)^\{p\},1\\\}\\bigr\],p∈\(0,1\]\.\\displaystyle\\hskip\-20\.00003ptp\\in\(0,1\]\.Under Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N6\)](https://arxiv.org/html/2607.16384#S2.I3.i6), there exist constantsc1,c2\>0c\_\{1\},c\_\{2\}\>0andαg∈\(0,1\]\\alpha\_\{g\}\\in\(0,1\]such that, for0<α≤αg0<\\alpha\\leq\\alpha\_\{g\},
a¯α\\displaystyle\\bar\{a\}\_\{\\alpha\}≥c1α,\\displaystyle\\geq c\_\{1\}\\alpha,\(A\.1\)a¯α,p\\displaystyle\\bar\{a\}\_\{\\alpha,p\}≥c2a¯αp\.\\displaystyle\\geq c\_\{2\}\\bar\{a\}\_\{\\alpha\}^\{p\}\.\(A\.2\)Moreover, under Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3)ifβ=2\\beta=2, and under[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)ifβ<2\\beta<2, there existsCg\>0C\_\{g\}\>0such that
a¯α≤Cgα\.\\bar\{a\}\_\{\\alpha\}\\leq C\_\{g\}\\alpha\.\(A\.3\)Consequentlya¯α=Θ\(α\)\\bar\{a\}\_\{\\alpha\}=\\Theta\(\\alpha\)\.
###### Proof\.
Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N6\)](https://arxiv.org/html/2607.16384#S2.I3.i6)givesinfξ∈Ξℙ\(Gξ≥εg\)≥pg\\inf\_\{\\xi\\in\\Xi\}\\mathbb\{P\}\(G\_\{\\xi\}\\geq\\varepsilon\_\{g\}\)\\geq p\_\{g\}\. For0<α≤εg−10<\\alpha\\leq\\varepsilon\_\{g\}^\{\-1\},min\{αGξ,1\}≥αεg𝟏\{Gξ≥εg\}\\min\\\{\\alpha G\_\{\\xi\},1\\\}\\geq\\alpha\\varepsilon\_\{g\}\\mathbf\{1\}\_\{\\\{G\_\{\\xi\}\\geq\\varepsilon\_\{g\}\\\}\}, and thereforea¯α≥αεgpg\\bar\{a\}\_\{\\alpha\}\\geq\\alpha\\varepsilon\_\{g\}p\_\{g\}, proving \([A\.1](https://arxiv.org/html/2607.16384#A1.E1)\)\. Similarly,min\{\(αGξ\)p,1\}≥\(αεg\)p𝟏\{Gξ≥εg\}\\min\\\{\(\\alpha G\_\{\\xi\}\)^\{p\},1\\\}\\geq\(\\alpha\\varepsilon\_\{g\}\)^\{p\}\\mathbf\{1\}\_\{\\\{G\_\{\\xi\}\\geq\\varepsilon\_\{g\}\\\}\}, soa¯α,p≥αpεgppg\\bar\{a\}\_\{\\alpha,p\}\\geq\\alpha^\{p\}\\varepsilon\_\{g\}^\{p\}p\_\{g\}\.
It remains to compareαp\\alpha^\{p\}witha¯αp\\bar\{a\}\_\{\\alpha\}^\{p\}\. Letξ⋆\\xi\_\{\\star\}be the reference point in Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3)\. Ifβ=2\\beta=2, then𝔼Gξ⋆<∞\\mathbb\{E\}G\_\{\\xi\_\{\\star\}\}<\\inftyby[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3); ifβ<2\\beta<2, Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)atx=0x=0andξ=ξ⋆\\xi=\\xi\_\{\\star\}gives the same conclusion\. Thusa¯α≤𝔼min\(αGξ⋆,1\)≤α𝔼Gξ⋆\\bar\{a\}\_\{\\alpha\}\\leq\\mathbb\{E\}\\min\(\\alpha G\_\{\\xi\_\{\\star\}\},1\)\\leq\\alpha\\mathbb\{E\}G\_\{\\xi\_\{\\star\}\}\. Combining this with the previous lower bound ona¯α,p\\bar\{a\}\_\{\\alpha,p\}proves \([A\.2](https://arxiv.org/html/2607.16384#A1.E2)\); the same inequality gives \([A\.3](https://arxiv.org/html/2607.16384#A1.E3)\)\. Together with \([A\.1](https://arxiv.org/html/2607.16384#A1.E1)\), this provesa¯α=Θ\(α\)\\bar\{a\}\_\{\\alpha\}=\\Theta\(\\alpha\)\. ∎
###### Lemma A\.3\.
The following estimates hold\.
1. \(i\)LetZZbe anℝd\\mathbb\{R\}^\{d\}\-valued random vector with𝔼Z=0\\mathbb\{E\}Z=0and𝔼eλ‖Z‖≤M\\mathbb\{E\}e^\{\\lambda\\left\\lVert Z\\right\\rVert\}\\leq Mfor someλ\>0\\lambda\>0\. Then there existt0∈\(0,λ/2\]t\_\{0\}\\in\(0,\\lambda/2\]andC\>0C\>0such that𝔼et⟨u,Z⟩≤eCt2\\mathbb\{E\}e^\{t\\langle u,Z\\rangle\}\\leq e^\{Ct^\{2\}\}for every unit vectoruuand every\|t\|≤t0\|t\|\\leq t\_\{0\}\.
2. \(ii\)IfY≥0Y\\geq 0and𝔼eλY≤M\\mathbb\{E\}e^\{\\lambda Y\}\\leq M, then𝔼etY≤1\+\(2M/λ\)t\\mathbb\{E\}e^\{tY\}\\leq 1\+\(2M/\\lambda\)tfor0≤t≤λ/20\\leq t\\leq\\lambda/2\.
3. \(iii\)Suppose1≤β<21\\leq\\beta<2\. There existsCβ\>0C\_\{\\beta\}\>0such that, whenevery≠0y\\neq 0and‖z‖≤‖y‖/2\\left\\lVert z\\right\\rVert\\leq\\left\\lVert y\\right\\rVert/2, ‖y−z‖2−β≤‖y‖2−β−\(2−β\)‖y‖1−β⟨y‖y‖,z⟩\+Cβ‖y‖−β‖z‖2\.\\left\\lVert y\-z\\right\\rVert^\{2\-\\beta\}\\leq\\left\\lVert y\\right\\rVert^\{2\-\\beta\}\-\(2\-\\beta\)\\left\\lVert y\\right\\rVert^\{1\-\\beta\}\\left\\langle\\frac\{y\}\{\\left\\lVert y\\right\\rVert\},z\\right\\rangle\+C\_\{\\beta\}\\left\\lVert y\\right\\rVert^\{\-\\beta\}\\left\\lVert z\\right\\rVert^\{2\}\.
###### Proof\.
For \(i\), setX=⟨u,Z⟩X=\\langle u,Z\\rangle, so𝔼X=0\\mathbb\{E\}X=0and\|X\|≤‖Z‖\|X\|\\leq\\left\\lVert Z\\right\\rVert\. For\|t\|≤λ/2\|t\|\\leq\\lambda/2,
etX\\displaystyle e^\{tX\}≤1\+tX\+t2‖Z‖22e\|t\|‖Z‖,\\displaystyle\\leq 1\+tX\+\\frac\{t^\{2\}\\left\\lVert Z\\right\\rVert^\{2\}\}\{2\}e^\{\|t\|\\left\\lVert Z\\right\\rVert\},r2e\(λ/2\)r\\displaystyle r^\{2\}e^\{\(\\lambda/2\)r\}≤16λ2eλr,\\displaystyle\\leq\\frac\{16\}\{\\lambda^\{2\}\}e^\{\\lambda r\},𝔼etX\\displaystyle\\mathbb\{E\}e^\{tX\}≤1\+8Mλ2t2≤exp\(8Mλ2t2\)\.\\displaystyle\\leq 1\+\\frac\{8M\}\{\\lambda^\{2\}\}t^\{2\}\\leq\\exp\\\!\\left\(\\frac\{8M\}\{\\lambda^\{2\}\}t^\{2\}\\right\)\.For \(ii\),etY−1≤tYetY≤tYe\(λ/2\)Ye^\{tY\}\-1\\leq tYe^\{tY\}\\leq tYe^\{\(\\lambda/2\)Y\}andre\(λ/2\)r≤\(2/λ\)eλrre^\{\(\\lambda/2\)r\}\\leq\(2/\\lambda\)e^\{\\lambda r\}, so𝔼etY≤1\+\(2M/λ\)t\\mathbb\{E\}e^\{tY\}\\leq 1\+\(2M/\\lambda\)t\. For \(iii\), apply Taylor’s theorem toψ\(w\)=‖w‖2−β\\psi\(w\)=\\left\\lVert w\\right\\rVert^\{2\-\\beta\}onℝd∖\{0\}\\mathbb\{R\}^\{d\}\\setminus\\\{0\\\}\. Alongy−szy\-sz,s∈\[0,1\]s\\in\[0,1\], the condition‖z‖≤‖y‖/2\\left\\lVert z\\right\\rVert\\leq\\left\\lVert y\\right\\rVert/2gives‖y−sz‖≥‖y‖/2\\left\\lVert y\-sz\\right\\rVert\\geq\\left\\lVert y\\right\\rVert/2; since‖∇2ψ\(w\)‖op≤Cβ‖w‖−β\\left\\lVert\\nabla^\{2\}\\psi\(w\)\\right\\rVert\_\{\\mathrm\{op\}\}\\leq C\_\{\\beta\}\\left\\lVert w\\right\\rVert^\{\-\\beta\}, the claimed inequality follows\. ∎
### A\.4Proof of Lemma[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3)
For fixedξ\\xi, write
ξ\+:=Φ\(ξ,U1\),fξ\+\(x\)=x−α\{h\(x\)\+g\(x,ξ\+\)\}\.\\xi^\{\+\}:=\\Phi\(\\xi,U\_\{1\}\),\\qquad f\_\{\\xi^\{\+\}\}\(x\)=x\-\\alpha\\\{h\(x\)\+g\(x,\\xi^\{\+\}\)\\\}\.Fore∈𝕊d−1e\\in\\mathbb\{S\}^\{d\-1\}, set𝒟ξ,α\(x,e\):=‖\(Id−αA\(x,ξ\+\)\)e‖\.\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\):=\\left\\lVert\(I\_\{d\}\-\\alpha A\(x,\\xi^\{\+\}\)\)e\\right\\rVert\.By \([4\.7](https://arxiv.org/html/2607.16384#S4.E7)\),
𝒟ξ,α\(x,e\)≤1,e∈𝕊d−1\.\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\\leq 1,\\qquad e\\in\\mathbb\{S\}^\{d\-1\}\.\(A\.4\)Moreover, by \([4\.8](https://arxiv.org/html/2607.16384#S4.E8)\),
𝔼𝒟ξ,α\(x,e\)≤1−α\(1−θ\)2⟨e,∇2H\(x\)e⟩,e∈𝕊d−1\.\\mathbb\{E\}\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\\leq 1\-\\frac\{\\alpha\(1\-\\theta\)\}\{2\}\\left\\langle e,\\nabla^\{2\}H\(x\)e\\right\\rangle,\\qquad e\\in\\mathbb\{S\}^\{d\-1\}\.\(A\.5\)
We first record two weight estimates used below\.
###### Lemma A\.4\.
There exist constantsC,c\>0C,c\>0, independent ofα,x,ξ,κ,κ0\\alpha,x,\\xi,\\kappa,\\kappa\_\{0\}, such that, for all sufficiently smallα\\alpha, the following hold\. If‖x‖≤RH2\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\}, then
𝔼\[\(𝒯β\(fξ\+\(x\)\)−1\)\+\]≤Ce−c/α,\\mathbb\{E\}\\bigl\[\(\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\-1\)\_\{\+\}\\bigr\]\\leq Ce^\{\-c/\\alpha\},\(A\.6\)with the left\-hand side equal to zero whenβ=2\\beta=2\. If‖x‖≤RH\\left\\lVert x\\right\\rVert\\leq R\_\{H\}, then, with the convention thatκ=0\\kappa=0whenβ=2\\beta=2,
𝔼\[\(Vα\(fξ\+\(x\)\)−Vα\(x\)\)\+\]≤C\(κ\+κ0\)αVα\(x\),\\mathbb\{E\}\\bigl\[\(V\_\{\\alpha\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\-V\_\{\\alpha\}\(x\)\)\_\{\+\}\\bigr\]\\leq C\(\\kappa\+\\kappa\_\{0\}\)\\alpha V\_\{\\alpha\}\(x\),\(A\.7\)and consequently
𝔼Vα\(fξ\+\(x\)\)≤\(1\+C\(κ\+κ0\)α\)Vα\(x\)\.\\mathbb\{E\}V\_\{\\alpha\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\\leq\\bigl\(1\+C\(\\kappa\+\\kappa\_\{0\}\)\\alpha\\bigr\)V\_\{\\alpha\}\(x\)\.\(A\.8\)
###### Proof\.
The caseβ=2\\beta=2has𝒯β≡1\\mathcal\{T\}\_\{\\beta\}\\equiv 1, so \([A\.6](https://arxiv.org/html/2607.16384#A1.E6)\) is immediate\. Suppose throughout the rest of the proof that1≤β<21\\leq\\beta<2\. Sincehhis continuous, it is bounded on the compact set\{‖x‖≤RH\}\\\{\\left\\lVert x\\right\\rVert\\leq R\_\{H\}\\\}\. Also1\+‖x‖β−11\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}is bounded on this compact set, so Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)implies that, for someλK,MK\>0\\lambda\_\{K\},M\_\{K\}\>0,
sup‖x‖≤RHsupξ∈Ξ𝔼exp\{λK‖g\(x,ξ\+\)‖\}≤MK\.\\sup\_\{\\left\\lVert x\\right\\rVert\\leq R\_\{H\}\}\\sup\_\{\\xi\\in\\Xi\}\\mathbb\{E\}\\exp\\\{\\lambda\_\{K\}\\left\\lVert g\(x,\\xi^\{\+\}\)\\right\\rVert\\\}\\leq M\_\{K\}\.\(A\.9\)
First assume‖x‖≤RH2\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\}\. If𝒯β\(fξ\+\(x\)\)\>1\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\>1, then‖fξ\+\(x\)‖\>3RH4\\left\\lVert f\_\{\\xi^\{\+\}\}\(x\)\\right\\rVert\>\\frac\{3R\_\{H\}\}\{4\}\. Since3RH4\>RH2\\frac\{3R\_\{H\}\}\{4\}\>\\frac\{R\_\{H\}\}\{2\}and‖h\(x\)‖≤C\\left\\lVert h\(x\)\\right\\rVert\\leq C, this implies, after decreasing the stepsize threshold,α‖g\(x,ξ\+\)‖≥c\\alpha\\left\\lVert g\(x,\\xi^\{\+\}\)\\right\\rVert\\geq c\. Moreover,
\(‖fξ\+\(x\)‖2−β−\(3RH4\)2−β\)\+≤C\{1\+α‖g\(x,ξ\+\)‖\},\\bigl\(\\left\\lVert f\_\{\\xi^\{\+\}\}\(x\)\\right\\rVert^\{2\-\\beta\}\-\(\\tfrac\{3R\_\{H\}\}\{4\}\)^\{2\-\\beta\}\\bigr\)\_\{\+\}\\leq C\\\{1\+\\alpha\\left\\lVert g\(x,\\xi^\{\+\}\)\\right\\rVert\\\},because2−β≤12\-\\beta\\leq 1\. Takingκ\\kappasmall enough thatCκα≤λK/2C\\kappa\\alpha\\leq\\lambda\_\{K\}/2, \([A\.9](https://arxiv.org/html/2607.16384#A1.E9)\) gives
𝔼\[\(𝒯β\(fξ\+\(x\)\)−1\)\+\]\\displaystyle\\mathbb\{E\}\\bigl\[\(\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\-1\)\_\{\+\}\\bigr\]≤C𝔼\[eCκα‖g\(x,ξ\+\)‖𝟏\{‖g\(x,ξ\+\)‖≥c/α\}\]\\displaystyle\\leq C\\mathbb\{E\}\\\!\\bigl\[e^\{C\\kappa\\alpha\\left\\lVert g\(x,\\xi^\{\+\}\)\\right\\rVert\}\\mathbf\{1\}\_\{\\\{\\left\\lVert g\(x,\\xi^\{\+\}\)\\right\\rVert\\geq c/\\alpha\\\}\}\\bigr\]≤Ce−c/α,\\displaystyle\\leq Ce^\{\-c/\\alpha\},which proves \([A\.6](https://arxiv.org/html/2607.16384#A1.E6)\)\.
It remains to prove \([A\.7](https://arxiv.org/html/2607.16384#A1.E7)\)\. Fix‖x‖≤RH\\left\\lVert x\\right\\rVert\\leq R\_\{H\}and setG=‖g\(x,ξ\+\)‖G=\\left\\lVert g\(x,\\xi^\{\+\}\)\\right\\rVert\. Split according toBα:=\{G≤α−1/2\}B\_\{\\alpha\}:=\\\{G\\leq\\alpha^\{\-1/2\}\\\}\. OnBαB\_\{\\alpha\},‖fξ\+\(x\)−x‖≤Cα\(1\+G\)≤Cα1/2,\\left\\lVert f\_\{\\xi^\{\+\}\}\(x\)\-x\\right\\rVert\\leq C\\alpha\(1\+G\)\\leq C\\alpha^\{1/2\},so bothxxandfξ\+\(x\)f\_\{\\xi^\{\+\}\}\(x\)lie in a fixed compact enlargement of\{‖y‖≤RH\}\\\{\\left\\lVert y\\right\\rVert\\leq R\_\{H\}\\\}\. On this enlargement,𝒯β\\mathcal\{T\}\_\{\\beta\}has Lipschitz constant at mostCκC\\kappa, for0<κ≤10<\\kappa\\leq 1\. Hence
𝔼\[\(𝒯β\(fξ\+\(x\)\)−𝒯β\(x\)\)\+𝟏Bα\]\\displaystyle\\mathbb\{E\}\\bigl\[\(\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\-\\mathcal\{T\}\_\{\\beta\}\(x\)\)\_\{\+\}\\mathbf\{1\}\_\{B\_\{\\alpha\}\}\\bigr\]≤Cκα𝔼\(1\+G\)≤Cκα\.\\displaystyle\\leq C\\kappa\\alpha\\mathbb\{E\}\(1\+G\)\\leq C\\kappa\\alpha\.OnBαcB\_\{\\alpha\}^\{c\}, useea−eb≤\(a−b\)\+eae^\{a\}\-e^\{b\}\\leq\(a\-b\)\_\{\+\}e^\{a\}fora≥ba\\geq b, the bound\(‖fξ\+\(x\)‖2−β−\(3RH4\)2−β\)\+≤C\{1\+\(αG\)2−β\}\\bigl\(\\left\\lVert f\_\{\\xi^\{\+\}\}\(x\)\\right\\rVert^\{2\-\\beta\}\-\(\\tfrac\{3R\_\{H\}\}\{4\}\)^\{2\-\\beta\}\\bigr\)\_\{\+\}\\leq C\\\{1\+\(\\alpha G\)^\{2\-\\beta\}\\\}, and \([A\.9](https://arxiv.org/html/2607.16384#A1.E9)\)\. Decreasingα\\alphaif necessary,
𝔼\[\(𝒯β\(fξ\+\(x\)\)−𝒯β\(x\)\)\+𝟏Bαc\]\\displaystyle\\mathbb\{E\}\\bigl\[\(\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\-\\mathcal\{T\}\_\{\\beta\}\(x\)\)\_\{\+\}\\mathbf\{1\}\_\{B\_\{\\alpha\}^\{c\}\}\\bigr\]≤Cκ𝔼\[\(1\+\(αG\)2−β\)eCκ\(1\+\(αG\)2−β\)𝟏\{G\>α−1/2\}\]≤Cκe−cα−1/2≤Cκα\.\\displaystyle\\qquad\\leq C\\kappa\\mathbb\{E\}\\\!\\bigl\[\(1\+\(\\alpha G\)^\{2\-\\beta\}\)e^\{C\\kappa\(1\+\(\\alpha G\)^\{2\-\\beta\}\)\}\\mathbf\{1\}\_\{\\\{G\>\\alpha^\{\-1/2\}\\\}\}\\bigr\]\\leq C\\kappa e^\{\-c\\alpha^\{\-1/2\}\}\\leq C\\kappa\\alpha\.Therefore
𝔼\[\(𝒯β\(fξ\+\(x\)\)−𝒯β\(x\)\)\+\]≤Cκα\.\\mathbb\{E\}\\bigl\[\(\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\-\\mathcal\{T\}\_\{\\beta\}\(x\)\)\_\{\+\}\\bigr\]\\leq C\\kappa\\alpha\.\(A\.10\)Theω\\omega\-part satisfies0≤ω≤10\\leq\\omega\\leq 1, and hence
δα\(ω\(fξ\+\(x\)\)−ω\(x\)\)\+≤δα≤Cκ0α\.\\delta\_\{\\alpha\}\(\\omega\(f\_\{\\xi^\{\+\}\}\(x\)\)\-\\omega\(x\)\)\_\{\+\}\\leq\\delta\_\{\\alpha\}\\leq C\\kappa\_\{0\}\\alpha\.SinceVα\(x\)≥1V\_\{\\alpha\}\(x\)\\geq 1, \([A\.7](https://arxiv.org/html/2607.16384#A1.E7)\) follows from \([A\.10](https://arxiv.org/html/2607.16384#A1.E10)\)\. Finally, \([A\.8](https://arxiv.org/html/2607.16384#A1.E8)\) follows fromVα\(f\)≤Vα\(x\)\+\(Vα\(f\)−Vα\(x\)\)\+V\_\{\\alpha\}\(f\)\\leq V\_\{\\alpha\}\(x\)\+\(V\_\{\\alpha\}\(f\)\-V\_\{\\alpha\}\(x\)\)\_\{\+\}\. ∎
###### Lemma A\.5\.
Assume1≤β<21\\leq\\beta<2\. There existc\>0c\>0andαtail∈\(0,1\]\\alpha\_\{\\rm tail\}\\in\(0,1\]such that, for all0<α≤αtail0<\\alpha\\leq\\alpha\_\{\\rm tail\}, allξ∈Ξ\\xi\\in\\Xi, and all‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\},
𝔼𝒯β\(fξ\+\(x\)\)≤\(1−cα\)𝒯β\(x\)\.\\mathbb\{E\}\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\\leq\(1\-c\\alpha\)\\mathcal\{T\}\_\{\\beta\}\(x\)\.\(A\.11\)
###### Proof\.
Letr=‖x‖r=\\left\\lVert x\\right\\rVert,G=g\(x,ξ\+\)G=g\(x,\\xi^\{\+\}\), andζ=G−g¯\(x,ξ\)\\zeta=G\-\\bar\{g\}\(x,\\xi\)\. Then𝔼\[ζ∣ξ\]=0\\mathbb\{E\}\[\\zeta\\mid\\xi\]=0andfξ\+\(x\)=x−α\{h\(x\)\+g¯\(x,ξ\)\+ζ\}\.f\_\{\\xi^\{\+\}\}\(x\)=x\-\\alpha\\\{h\(x\)\+\\bar\{g\}\(x,\\xi\)\+\\zeta\\\}\.By Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5),‖g¯\(x,ξ\)‖≤C\(1\+‖x‖β−1\)\\left\\lVert\\bar\{g\}\(x,\\xi\)\\right\\rVert\\leq C\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\), uniformly in\(x,ξ\)\(x,\\xi\)\. Fixη0∈\(0,1/2\)\\eta\_\{0\}\\in\(0,1/2\), put
Nx:=‖g\(x,ξ\+\)‖1\+‖x‖β−1,Ex:=\{Nx≤α−η0\}\.N\_\{x\}:=\\frac\{\\left\\lVert g\(x,\\xi^\{\+\}\)\\right\\rVert\}\{1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\},\\qquad E\_\{x\}:=\\\{N\_\{x\}\\leq\\alpha^\{\-\\eta\_\{0\}\}\\\}\.OnExE\_\{x\}, for sufficiently smallα\\alpha,
α‖h\(x\)\+G‖≤Cα1−η0rβ−1≤ε∗r,r≥RH,\\alpha\\left\\lVert h\(x\)\+G\\right\\rVert\\leq C\\alpha^\{1\-\\eta\_\{0\}\}r^\{\\beta\-1\}\\leq\\varepsilon\_\{\*\}r,\\qquad r\\geq R\_\{H\},withε∗\>0\\varepsilon\_\{\*\}\>0chosen so that the segment fromxxtofξ\+\(x\)f\_\{\\xi^\{\+\}\}\(x\)remains in the tail region where the cutoff in𝒯β\\mathcal\{T\}\_\{\\beta\}is inactive\. Lemma[A\.3](https://arxiv.org/html/2607.16384#A1.Thmtheorem3)\(iii\), applied toz=α\{h\(x\)\+G\}z=\\alpha\\\{h\(x\)\+G\\\}, gives onExE\_\{x\}
‖fξ\+\(x\)‖2−β\\displaystyle\\left\\lVert f\_\{\\xi^\{\+\}\}\(x\)\\right\\rVert^\{2\-\\beta\}≤r2−β−\(2−β\)αr−β⟨x,h\(x\)\+g¯\(x,ξ\)⟩−\(2−β\)αr−β⟨x,ζ⟩\+Cα2−2η0\.\\displaystyle\\leq r^\{2\-\\beta\}\-\(2\-\\beta\)\\alpha r^\{\-\\beta\}\\left\\langle x,h\(x\)\+\\bar\{g\}\(x,\\xi\)\\right\\rangle\-\(2\-\\beta\)\\alpha r^\{\-\\beta\}\\left\\langle x,\\zeta\\right\\rangle\+C\\alpha^\{2\-2\\eta\_\{0\}\}\.By the tail dissipativity condition[\(N4\)](https://arxiv.org/html/2607.16384#S2.I3.i4),r−β⟨x,h\(x\)\+g¯\(x,ξ\)⟩≥cdissr^\{\-\\beta\}\\left\\langle x,h\(x\)\+\\bar\{g\}\(x,\\xi\)\\right\\rangle\\geq c\_\{\\rm diss\}\. Therefore, conditionally onξ\\xi,
𝒯β\(fξ\+\(x\)\)𝟏Ex≤𝒯β\(x\)exp\{−κ\(2−β\)cdissα−κ\(2−β\)αr−β⟨x,ζ⟩\+Cα2−2η0\}\.\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\\mathbf\{1\}\_\{E\_\{x\}\}\\leq\\mathcal\{T\}\_\{\\beta\}\(x\)\\exp\\\{\-\\kappa\(2\-\\beta\)c\_\{\\rm diss\}\\alpha\-\\kappa\(2\-\\beta\)\\alpha r^\{\-\\beta\}\\left\\langle x,\\zeta\\right\\rangle\+C\\alpha^\{2\-2\\eta\_\{0\}\}\\\}\.Writinge=x/re=x/r, the centered term is
r−β⟨x,ζ⟩=r1−β\(1\+∥x∥β−1\)⟨e,g\(x,ξ\+\)1\+‖x‖β−1−𝔼\[g\(x,ξ\+\)1\+‖x‖β−1\|ξ\]⟩\.r^\{\-\\beta\}\\left\\langle x,\\zeta\\right\\rangle=r^\{1\-\\beta\}\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\)\\left\\langle e,\\frac\{g\(x,\\xi^\{\+\}\)\}\{1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\}\-\\mathbb\{E\}\\\!\\left\[\\frac\{g\(x,\\xi^\{\+\}\)\}\{1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\}\\,\\middle\|\\,\\xi\\right\]\\right\\rangle\.Sincer1−β\(1\+‖x‖β−1\)r^\{1\-\\beta\}\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\)is uniformly bounded forr≥RHr\\geq R\_\{H\}, the coefficient of the centered normalized noise isO\(α\)O\(\\alpha\)\. The conditional exponential moment in[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)also gives a uniform exponential moment for this centered normalized noise\. Lemma[A\.3](https://arxiv.org/html/2607.16384#A1.Thmtheorem3)\(i\) therefore implies
𝔼\[exp\{−κ\(2−β\)αr−β⟨x,ζ⟩\}∣ξ\]≤eCα2\.\\mathbb\{E\}\\left\[\\exp\\\{\-\\kappa\(2\-\\beta\)\\alpha r^\{\-\\beta\}\\left\\langle x,\\zeta\\right\\rangle\\\}\\mid\\xi\\right\]\\leq e^\{C\\alpha^\{2\}\}\.Thus
𝔼\[𝒯β\(fξ\+\(x\)\)𝟏Ex\]≤e−cα\+Cα2−2η0𝒯β\(x\)\.\\mathbb\{E\}\[\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\\mathbf\{1\}\_\{E\_\{x\}\}\]\\leq e^\{\-c\\alpha\+C\\alpha^\{2\-2\\eta\_\{0\}\}\}\\mathcal\{T\}\_\{\\beta\}\(x\)\.\(A\.12\)
It remains to controlExcE\_\{x\}^\{c\}\. Because2−β∈\(0,1\]2\-\\beta\\in\(0,1\]ands↦s2−βs\\mapsto s^\{2\-\\beta\}is concave,
‖x−α\(h\(x\)\+G\)‖2−β−r2−β\\displaystyle\\left\\lVert x\-\\alpha\(h\(x\)\+G\)\\right\\rVert^\{2\-\\beta\}\-r^\{2\-\\beta\}≤\(r\+α‖h\(x\)\+G‖\)2−β−r2−β\\displaystyle\\leq\(r\+\\alpha\\left\\lVert h\(x\)\+G\\right\\rVert\)^\{2\-\\beta\}\-r^\{2\-\\beta\}≤\(2−β\)r1−βα‖h\(x\)\+G‖\.\\displaystyle\\leq\(2\-\\beta\)r^\{1\-\\beta\}\\alpha\\left\\lVert h\(x\)\+G\\right\\rVert\.Forr≥RHr\\geq R\_\{H\}, Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H3H3\-b\)](https://arxiv.org/html/2607.16384#S2.I1.i3.I1.i2)and the definition ofNxN\_\{x\}give‖h\(x\)\+G‖≤Crβ−1\(1\+Nx\)\.\\left\\lVert h\(x\)\+G\\right\\rVert\\leq Cr^\{\\beta\-1\}\(1\+N\_\{x\}\)\.The powers ofrrtherefore cancel, and we obtain the deterministic bound‖fξ\+\(x\)‖2−β−r2−β≤Cα\(1\+Nx\)\.\\left\\lVert f\_\{\\xi^\{\+\}\}\(x\)\\right\\rVert^\{2\-\\beta\}\-r^\{2\-\\beta\}\\leq C\\alpha\(1\+N\_\{x\}\)\.Consequently,
𝒯β\(fξ\+\(x\)\)≤𝒯β\(x\)exp\{Cκα\(1\+Nx\)\}\.\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\\leq\\mathcal\{T\}\_\{\\beta\}\(x\)\\exp\\\{C\\kappa\\alpha\(1\+N\_\{x\}\)\\\}\.Choosingα\\alphaso thatCκα≤λ0/2C\\kappa\\alpha\\leq\\lambda\_\{0\}/2, Assumption[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)yields
𝔼\[𝒯β\(fξ\+\(x\)\)𝟏Exc\]≤Ce−cα−η0𝒯β\(x\)\.\\mathbb\{E\}\[\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\\mathbf\{1\}\_\{E\_\{x\}^\{c\}\}\]\\leq Ce^\{\-c\\alpha^\{\-\\eta\_\{0\}\}\}\\mathcal\{T\}\_\{\\beta\}\(x\)\.\(A\.13\)Combining \([A\.12](https://arxiv.org/html/2607.16384#A1.E12)\) and \([A\.13](https://arxiv.org/html/2607.16384#A1.E13)\), using2−2η0\>12\-2\\eta\_\{0\}\>1and the fact thate−cα−η0=o\(α\)e^\{\-c\\alpha^\{\-\\eta\_\{0\}\}\}=o\(\\alpha\), gives \([A\.11](https://arxiv.org/html/2607.16384#A1.E11)\) after reducing the stepsize threshold\. ∎
###### Lemma A\.6\.
There existC\>0C\>0andα0∈\(0,1\]\\alpha\_\{0\}\\in\(0,1\]such that, for every0<α≤α00<\\alpha\\leq\\alpha\_\{0\},x∈ℝdx\\in\\mathbb\{R\}^\{d\}, andξ∈Ξ\\xi\\in\\Xi,
𝔼\[\(Vα\(fΦ\(ξ,U1\)\(x\)\)−Vα\(x\)\)\+\]≤CαVα\(x\)\.\\mathbb\{E\}\\left\[\\bigl\(V\_\{\\alpha\}\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\)\-V\_\{\\alpha\}\(x\)\\bigr\)\_\{\+\}\\right\]\\leq C\\alpha V\_\{\\alpha\}\(x\)\.\(A\.14\)Consequently,
𝔼\[LΦ\(U1\)Vα\(fΦ\(ξ,U1\)\(x\)\)\]≤\(ρΞ\+Cα\)Vα\(x\)\.\\mathbb\{E\}\\left\[L\_\{\\Phi\}\(U\_\{1\}\)V\_\{\\alpha\}\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\)\\right\]\\leq\(\\rho\_\{\\Xi\}\+C\\alpha\)V\_\{\\alpha\}\(x\)\.\(A\.15\)
###### Proof\.
For‖x‖≤RH\\left\\lVert x\\right\\rVert\\leq R\_\{H\}, \([A\.14](https://arxiv.org/html/2607.16384#A1.E14)\) is \([A\.7](https://arxiv.org/html/2607.16384#A1.E7)\)\. Suppose‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\}\. Ifβ=2\\beta=2, then𝒯β≡1\\mathcal\{T\}\_\{\\beta\}\\equiv 1andω\(x\)=0\\omega\(x\)=0, so
\(Vα\(fΦ\(ξ,U1\)\(x\)\)−Vα\(x\)\)\+≤δα≤CαVα\(x\)\.\\bigl\(V\_\{\\alpha\}\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\)\-V\_\{\\alpha\}\(x\)\\bigr\)\_\{\+\}\\leq\\delta\_\{\\alpha\}\\leq C\\alpha V\_\{\\alpha\}\(x\)\.If1≤β<21\\leq\\beta<2, use the normalized noiseNx:=‖g\(x,Φ\(ξ,U1\)\)‖1\+‖x‖β−1\.N\_\{x\}:=\\frac\{\\left\\lVert g\(x,\\Phi\(\\xi,U\_\{1\}\)\)\\right\\rVert\}\{1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\}\.The deterministic tail comparison used in the proof of Lemma[A\.5](https://arxiv.org/html/2607.16384#A1.Thmtheorem5)gives
𝒯β\(fΦ\(ξ,U1\)\(x\)\)≤𝒯β\(x\)eCκα\(1\+Nx\)\.\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\)\\leq\\mathcal\{T\}\_\{\\beta\}\(x\)e^\{C\\kappa\\alpha\(1\+N\_\{x\}\)\}\.Sinceet−1≤tete^\{t\}\-1\\leq te^\{t\}fort≥0t\\geq 0, Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)yields, uniformly inx,ξx,\\xi,
𝔼\[\(𝒯β\(fΦ\(ξ,U1\)\(x\)\)−𝒯β\(x\)\)\+\]\\displaystyle\\mathbb\{E\}\\bigl\[\(\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\)\-\\mathcal\{T\}\_\{\\beta\}\(x\)\)\_\{\+\}\\bigr\]≤Cα𝒯β\(x\)𝔼\[\(1\+Nx\)eCκα\(1\+Nx\)\]\\displaystyle\\leq C\\alpha\\mathcal\{T\}\_\{\\beta\}\(x\)\\mathbb\{E\}\[\(1\+N\_\{x\}\)e^\{C\\kappa\\alpha\(1\+N\_\{x\}\)\}\]≤CαVα\(x\)\.\\displaystyle\\leq C\\alpha V\_\{\\alpha\}\(x\)\.The positive increment ofδαω\\delta\_\{\\alpha\}\\omegais at mostδα≤CαVα\(x\)\\delta\_\{\\alpha\}\\leq C\\alpha V\_\{\\alpha\}\(x\), which proves \([A\.14](https://arxiv.org/html/2607.16384#A1.E14)\)\. Finally,0≤LΦ≤10\\leq L\_\{\\Phi\}\\leq 1and𝔼LΦ\(U1\)=ρΞ\\mathbb\{E\}L\_\{\\Phi\}\(U\_\{1\}\)=\\rho\_\{\\Xi\}give
𝔼\[LΦ\(U1\)Vα\(fΦ\(ξ,U1\)\(x\)\)\]\\displaystyle\\mathbb\{E\}\[L\_\{\\Phi\}\(U\_\{1\}\)V\_\{\\alpha\}\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\)\]≤ρΞVα\(x\)\+𝔼\[\(Vα\(fΦ\(ξ,U1\)\(x\)\)−Vα\(x\)\)\+\],\\displaystyle\\leq\\rho\_\{\\Xi\}V\_\{\\alpha\}\(x\)\+\\mathbb\{E\}\[\(V\_\{\\alpha\}\(f\_\{\\Phi\(\\xi,U\_\{1\}\)\}\(x\)\)\-V\_\{\\alpha\}\(x\)\)\_\{\+\}\],and \([A\.15](https://arxiv.org/html/2607.16384#A1.E15)\) follows\. ∎
###### Lemma A\.7\.
There existccore\>0c\_\{\\rm core\}\>0andαcore∈\(0,1\]\\alpha\_\{\\rm core\}\\in\(0,1\]such that, for every0<α≤αcore0<\\alpha\\leq\\alpha\_\{\\rm core\}, everyξ∈Ξ\\xi\\in\\Xi, and every‖x‖≤RH2\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\},
supe∈𝕊d−1\(𝒦ξ,αVα\)\(x;e\)≤\(1−ccoreαm−1\)Vα\(x\)\.\\sup\_\{e\\in\\mathbb\{S\}^\{d\-1\}\}\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\\leq\(1\-c\_\{\\rm core\}\\alpha^\{m\-1\}\)V\_\{\\alpha\}\(x\)\.
###### Proof\.
Fixxxwithr=‖x‖≤RH2r=\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\}, and fixe∈𝕊d−1e\\in\\mathbb\{S\}^\{d\-1\}\. Since3RH4\>RH2\\frac\{3R\_\{H\}\}\{4\}\>\\frac\{R\_\{H\}\}\{2\},𝒯β\(x\)=1\\mathcal\{T\}\_\{\\beta\}\(x\)=1\. Putℛ\(u\)=𝒯β\(u\)−1\\mathcal\{R\}\(u\)=\\mathcal\{T\}\_\{\\beta\}\(u\)\-1whenβ<2\\beta<2, andℛ≡0\\mathcal\{R\}\\equiv 0whenβ=2\\beta=2\. Then
Vα\(u\)=1\+δαω\(u\)\+ℛ\(u\),Vα\(x\)=1\+δαω\(x\)\.V\_\{\\alpha\}\(u\)=1\+\\delta\_\{\\alpha\}\\omega\(u\)\+\\mathcal\{R\}\(u\),\\qquad V\_\{\\alpha\}\(x\)=1\+\\delta\_\{\\alpha\}\\omega\(x\)\.Using \([A\.4](https://arxiv.org/html/2607.16384#A1.E4)\),
\(𝒦ξ,αVα\)\(x;e\)−Vα\(x\)\\displaystyle\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\-V\_\{\\alpha\}\(x\)≤𝔼\[𝒟ξ,α\(x,e\)\]−1\+δα\{𝔼ω\(fξ\+\(x\)\)−ω\(x\)\}\\displaystyle\\leq\\mathbb\{E\}\[\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\]\-1\+\\delta\_\{\\alpha\}\\\{\\mathbb\{E\}\\omega\(f\_\{\\xi^\{\+\}\}\(x\)\)\-\\omega\(x\)\\\}\(A\.16\)\+𝔼ℛ\(fξ\+\(x\)\)\.\\displaystyle\\qquad\+\\mathbb\{E\}\\mathcal\{R\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\.
Putain:=\(1−θ\)cin/2a\_\{\\rm in\}:=\(1\-\\theta\)c\_\{\\rm in\}/2\. Ifm=2m=2, Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)gives⟨e,∇2H\(x\)e⟩≥cin\\left\\langle e,\\nabla^\{2\}H\(x\)e\\right\\rangle\\geq c\_\{\\rm in\}for‖x‖≤RH2\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\}\. Hence
𝔼𝒟ξ,α\(x,e\)≤1−ainα\.\\mathbb\{E\}\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\\leq 1\-a\_\{\\rm in\}\\alpha\.Sinceω≡0\\omega\\equiv 0whenm=2m=2, \([A\.16](https://arxiv.org/html/2607.16384#A1.E16)\) and \([A\.6](https://arxiv.org/html/2607.16384#A1.E6)\) give the claim after absorbing the exponentially small term intoα\\alpha\.
Assume nowm\>2m\>2, and putsm:=\(m−2\)∧1s\_\{m\}:=\(m\-2\)\\wedge 1\. Then
ω\(y\)=\(1−\(2‖y‖RH\)sm\)\+,y∈ℝd\.\\omega\(y\)=\\left\(1\-\\left\(\\frac\{2\\left\\lVert y\\right\\rVert\}\{R\_\{H\}\}\\right\)^\{s\_\{m\}\}\\right\)\_\{\+\},\\qquad y\\in\\mathbb\{R\}^\{d\}\.By Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)and \([A\.5](https://arxiv.org/html/2607.16384#A1.E5)\),𝔼𝒟ξ,α\(x,e\)≤1−ainαrm−2\.\\mathbb\{E\}\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\\leq 1\-a\_\{\\rm in\}\\alpha r^\{m\-2\}\.We claim that
𝔼ω\(fξ\+\(x\)\)−ω\(x\)≤−cω𝔼min\{\(α‖g\(0,ξ\+\)‖\)sm,1\}\+Cωrsm\.\\mathbb\{E\}\\omega\(f\_\{\\xi^\{\+\}\}\(x\)\)\-\\omega\(x\)\\leq\-c\_\{\\omega\}\\mathbb\{E\}\\min\\\{\(\\alpha\\left\\lVert g\(0,\\xi^\{\+\}\)\\right\\rVert\)^\{s\_\{m\}\},1\\\}\+C\_\{\\omega\}r^\{s\_\{m\}\}\.\(A\.17\)Indeed, write
fξ\+\(x\)=−αg\(0,ξ\+\)\+v,v=x−αh\(x\)−α\{g\(x,ξ\+\)−g\(0,ξ\+\)\}\.f\_\{\\xi^\{\+\}\}\(x\)=\-\\alpha g\(0,\\xi^\{\+\}\)\+v,\\qquad v=x\-\\alpha h\(x\)\-\\alpha\\\{g\(x,\\xi^\{\+\}\)\-g\(0,\\xi^\{\+\}\)\\\}\.Lemma[A\.1](https://arxiv.org/html/2607.16384#A1.Thmtheorem1)gives‖h‖\(x\)≤Crm−1\\left\\lVert h\\right\\rVert\(x\)\\leq Cr^\{m\-1\}for‖x‖≤RH2\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\}, while \([4\.5](https://arxiv.org/html/2607.16384#S4.E5)\) gives‖g\(x,ξ\+\)−g\(0,ξ\+\)‖≤\(2−θ\)L𝖦r/\(1−θ\)\\left\\lVert g\(x,\\xi^\{\+\}\)\-g\(0,\\xi^\{\+\}\)\\right\\rVert\\leq\(2\-\\theta\)L\_\{\\mathsf\{G\}\}r/\(1\-\\theta\)\. Hence, after reducing the stepsize threshold,‖v‖≤Cr\\left\\lVert v\\right\\rVert\\leq Cr\. The functiony↦ω\(y\)y\\mapsto\\omega\(y\)issms\_\{m\}\-Hölder, because\(1−ssm\)\+\(1\-s^\{s\_\{m\}\}\)\_\{\+\}issms\_\{m\}\-Hölder on\[0,∞\)\[0,\\infty\)\. Therefore
ω\(fξ\+\(x\)\)≤ω\(−αg\(0,ξ\+\)\)\+C‖v‖sm≤\(1−\(2α‖g\(0,ξ\+\)‖RH\)sm\)\+\+Crsm\.\\omega\(f\_\{\\xi^\{\+\}\}\(x\)\)\\leq\\omega\(\-\\alpha g\(0,\\xi^\{\+\}\)\)\+C\\left\\lVert v\\right\\rVert^\{s\_\{m\}\}\\leq\\left\(1\-\\left\(\\frac\{2\\alpha\\left\\lVert g\(0,\\xi^\{\+\}\)\\right\\rVert\}\{R\_\{H\}\}\\right\)^\{s\_\{m\}\}\\right\)\_\{\+\}\+Cr^\{s\_\{m\}\}\.Since
\(1−\(2zRH\)sm\)\+≤1−cmin\{zsm,1\},z≥0,\\left\(1\-\\left\(\\frac\{2z\}\{R\_\{H\}\}\\right\)^\{s\_\{m\}\}\\right\)\_\{\+\}\\leq 1\-c\\min\\\{z^\{s\_\{m\}\},1\\\},\\qquad z\\geq 0,andω\(x\)=1−\(2rRH\)sm\\omega\(x\)=1\-\(\\frac\{2r\}\{R\_\{H\}\}\)^\{s\_\{m\}\}forr≤RH2r\\leq\\frac\{R\_\{H\}\}\{2\}, \([A\.17](https://arxiv.org/html/2607.16384#A1.E17)\) follows\.
Combining \([A\.16](https://arxiv.org/html/2607.16384#A1.E16)\), \([A\.17](https://arxiv.org/html/2607.16384#A1.E17)\), \([A\.6](https://arxiv.org/html/2607.16384#A1.E6)\), and Lemma[A\.2](https://arxiv.org/html/2607.16384#A1.Thmtheorem2), we get
\(𝒦ξ,αVα\)\(x;e\)−Vα\(x\)≤−ainαrm−2−cδαa¯α,sm\+Cδαrsm\+Ce−c/α\.\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\-V\_\{\\alpha\}\(x\)\\leq\-a\_\{\\rm in\}\\alpha r^\{m\-2\}\-c\\delta\_\{\\alpha\}\\bar\{a\}\_\{\\alpha,s\_\{m\}\}\+C\\delta\_\{\\alpha\}r^\{s\_\{m\}\}\+Ce^\{\-c/\\alpha\}\.\(A\.18\)If2<m≤32<m\\leq 3, thensm=m−2s\_\{m\}=m\-2andδα=κ0α\\delta\_\{\\alpha\}=\\kappa\_\{0\}\\alpha\. Usinga¯α,sm≥ca¯αm−2\\bar\{a\}\_\{\\alpha,s\_\{m\}\}\\geq c\\bar\{a\}\_\{\\alpha\}^\{m\-2\}, \([A\.18](https://arxiv.org/html/2607.16384#A1.E18)\) becomes
\(𝒦ξ,αVα\)\(x;e\)−Vα\(x\)≤−\(ain−Cκ0\)αrm−2−cκ0αa¯αm−2\+Ce−c/α\.\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\-V\_\{\\alpha\}\(x\)\\leq\-\(a\_\{\\rm in\}\-C\\kappa\_\{0\}\)\\alpha r^\{m\-2\}\-c\\kappa\_\{0\}\\alpha\\bar\{a\}\_\{\\alpha\}^\{m\-2\}\+Ce^\{\-c/\\alpha\}\.Chooseκ0\\kappa\_\{0\}so small thatain−Cκ0≥ain/2a\_\{\\rm in\}\-C\\kappa\_\{0\}\\geq a\_\{\\rm in\}/2\. Thena¯α≥c1α\\bar\{a\}\_\{\\alpha\}\\geq c\_\{1\}\\alpha, and the exponentially small term is negligible compared withαm−1\\alpha^\{m\-1\}\. Hence
\(𝒦ξ,αVα\)\(x;e\)−Vα\(x\)≤−cαm−1\.\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\-V\_\{\\alpha\}\(x\)\\leq\-c\\alpha^\{m\-1\}\.\(A\.19\)
Ifm\>3m\>3, thensm=1s\_\{m\}=1,δα=κ0αm−2\\delta\_\{\\alpha\}=\\kappa\_\{0\}\\alpha^\{m\-2\}, anda¯α,1=a¯α\\bar\{a\}\_\{\\alpha,1\}=\\bar\{a\}\_\{\\alpha\}\. Factoring outα\\alphain \([A\.18](https://arxiv.org/html/2607.16384#A1.E18)\) gives
\(𝒦ξ,αVα\)\(x;e\)−Vα\(x\)≤α\[−ainrm−2−cκ0αm−3a¯α\+Cκ0αm−3r\]\+Ce−c/α\.\\displaystyle\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\-V\_\{\\alpha\}\(x\)\\leq\\alpha\\bigl\[\-a\_\{\\rm in\}r^\{m\-2\}\-c\\kappa\_\{0\}\\alpha^\{m\-3\}\\bar\{a\}\_\{\\alpha\}\+C\\kappa\_\{0\}\\alpha^\{m\-3\}r\\bigr\]\+Ce^\{\-c/\\alpha\}\.Putq:=\(m−2\)/\(m−3\)\>1q:=\(m\-2\)/\(m\-3\)\>1\. Young’s inequality yields
Cκ0αm−3r≤ain2rm−2\+C∗κ0qαm−2\.C\\kappa\_\{0\}\\alpha^\{m\-3\}r\\leq\\frac\{a\_\{\\rm in\}\}\{2\}r^\{m\-2\}\+C\_\{\*\}\\kappa\_\{0\}^\{q\}\\alpha^\{m\-2\}\.Sincea¯α≥c1α\\bar\{a\}\_\{\\alpha\}\\geq c\_\{1\}\\alpha, the negative noise term inside the brackets is at most−cc1κ0αm−2\-cc\_\{1\}\\kappa\_\{0\}\\alpha^\{m\-2\}\. Chooseκ0\\kappa\_\{0\}so small thatC∗κ0q−1≤cc1/2C\_\{\*\}\\kappa\_\{0\}^\{q\-1\}\\leq cc\_\{1\}/2\. After restoring the outer factorα\\alpha, the Young remainder is absorbed by the negative noise term\. Finally,e−c/α=o\(αm−1\)e^\{\-c/\\alpha\}=o\(\\alpha^\{m\-1\}\), so \([A\.19](https://arxiv.org/html/2607.16384#A1.E19)\) also holds form\>3m\>3\.
On the near\-minimizer region,1≤Vα\(x\)≤C1\\leq V\_\{\\alpha\}\(x\)\\leq Cuniformly inα\\alpha\. Since \([A\.19](https://arxiv.org/html/2607.16384#A1.E19)\) gives an additive decrease of at leastcαm−1c\\alpha^\{m\-1\}, it implies
\(𝒦ξ,αVα\)\(x;e\)≤\(1−ccoreαm−1\)Vα\(x\),\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\\leq\(1\-c\_\{\\rm core\}\\alpha^\{m\-1\}\)V\_\{\\alpha\}\(x\),after decreasingccorec\_\{\\rm core\}\. Taking the supremum overeecompletes the proof\. ∎
###### Lemma A\.8\.
There existcbr\>0c\_\{\\rm br\}\>0andαbr∈\(0,1\]\\alpha\_\{\\rm br\}\\in\(0,1\]such that, for every0<α≤αbr0<\\alpha\\leq\\alpha\_\{\\rm br\}, everyξ∈Ξ\\xi\\in\\Xi, and everyRH2≤‖x‖≤RH\\frac\{R\_\{H\}\}\{2\}\\leq\\left\\lVert x\\right\\rVert\\leq R\_\{H\},
supe∈𝕊d−1\(𝒦ξ,αVα\)\(x;e\)≤\(1−cbrαm−1\)Vα\(x\)\.\\sup\_\{e\\in\\mathbb\{S\}^\{d\-1\}\}\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\\leq\(1\-c\_\{\\rm br\}\\alpha^\{m\-1\}\)V\_\{\\alpha\}\(x\)\.
###### Proof\.
Fixe∈𝕊d−1e\\in\\mathbb\{S\}^\{d\-1\}\. By Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2), forRH2≤‖x‖≤RH\\frac\{R\_\{H\}\}\{2\}\\leq\\left\\lVert x\\right\\rVert\\leq R\_\{H\},⟨e,∇2H\(x\)e⟩≥cin\(RH2\)m−2\.\\left\\langle e,\\nabla^\{2\}H\(x\)e\\right\\rangle\\geq c\_\{\\rm in\}\\left\(\\frac\{R\_\{H\}\}\{2\}\\right\)^\{m\-2\}\.Together with \([A\.5](https://arxiv.org/html/2607.16384#A1.E5)\), this gives𝔼𝒟ξ,α\(x,e\)≤1−cα\\mathbb\{E\}\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\\leq 1\-c\\alphafor somec\>0c\>0\. Since𝒟ξ,α\(x,e\)≤1\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\\leq 1,
\(𝒦ξ,αVα\)\(x;e\)\\displaystyle\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)≤Vα\(x\)𝔼𝒟ξ,α\(x,e\)\+𝔼\[\(Vα\(fξ\+\(x\)\)−Vα\(x\)\)\+\]\\displaystyle\\leq V\_\{\\alpha\}\(x\)\\mathbb\{E\}\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\+\\mathbb\{E\}\[\(V\_\{\\alpha\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\-V\_\{\\alpha\}\(x\)\)\_\{\+\}\]≤\(1−cα\+C\(κ\+κ0\)α\)Vα\(x\),\\displaystyle\\leq\(1\-c\\alpha\+C\(\\kappa\+\\kappa\_\{0\}\)\\alpha\)V\_\{\\alpha\}\(x\),where the last line uses \([A\.7](https://arxiv.org/html/2607.16384#A1.E7)\)\. Chooseκ\\kappaandκ0\\kappa\_\{0\}sufficiently small to obtain a factor1−cα1\-c\\alpha\. Since0<α≤10<\\alpha\\leq 1andm≥2m\\geq 2,α≥αm−1\\alpha\\geq\\alpha^\{m\-1\}, which gives the stated estimate after decreasing the constant\. ∎
###### Lemma A\.9\.
There existcfar\>0c\_\{\\rm far\}\>0andαfar∈\(0,1\]\\alpha\_\{\\rm far\}\\in\(0,1\]such that, for every0<α≤αfar0<\\alpha\\leq\\alpha\_\{\\rm far\}, everyξ∈Ξ\\xi\\in\\Xi, and every‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\},
supe∈𝕊d−1\(𝒦ξ,αVα\)\(x;e\)≤\(1−cfarαm−1\)Vα\(x\)\.\\sup\_\{e\\in\\mathbb\{S\}^\{d\-1\}\}\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\\leq\(1\-c\_\{\\rm far\}\\alpha^\{m\-1\}\)V\_\{\\alpha\}\(x\)\.
###### Proof\.
First supposeβ=2\\beta=2\. Then𝒯β≡1\\mathcal\{T\}\_\{\\beta\}\\equiv 1, andω\(x\)=0\\omega\(x\)=0for‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\}, soVα\(x\)=1V\_\{\\alpha\}\(x\)=1\. By Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H3H3\-a\)](https://arxiv.org/html/2607.16384#S2.I1.i3.I1.i1)and \([A\.5](https://arxiv.org/html/2607.16384#A1.E5)\),𝔼𝒟ξ,α\(x,e\)≤1−cα\\mathbb\{E\}\\mathcal\{D\}\_\{\\xi,\\alpha\}\(x,e\)\\leq 1\-c\\alphafor somec\>0c\>0\. Since0≤ω≤10\\leq\\omega\\leq 1,\(𝒦ξ,αVα\)\(x;e\)≤1−cα\+δα\.\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\\leq 1\-c\\alpha\+\\delta\_\{\\alpha\}\.Sinceδα≤κ0α\\delta\_\{\\alpha\}\\leq\\kappa\_\{0\}\\alpha, choosingκ0\\kappa\_\{0\}small gives a factor1−cα1\-c\\alpha\.
Now suppose1≤β<21\\leq\\beta<2\. On‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\},ω\(x\)=0\\omega\(x\)=0, soVα\(x\)=𝒯β\(x\)V\_\{\\alpha\}\(x\)=\\mathcal\{T\}\_\{\\beta\}\(x\)\. By \([A\.4](https://arxiv.org/html/2607.16384#A1.E4)\) and Lemma[A\.5](https://arxiv.org/html/2607.16384#A1.Thmtheorem5),
\(𝒦ξ,αVα\)\(x;e\)≤𝔼𝒯β\(fξ\+\(x\)\)\+δα≤\(1−cα\)𝒯β\(x\)\+δα\.\(\\mathcal\{K\}\_\{\\xi,\\alpha\}V\_\{\\alpha\}\)\(x;e\)\\leq\\mathbb\{E\}\\mathcal\{T\}\_\{\\beta\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\+\\delta\_\{\\alpha\}\\leq\(1\-c\\alpha\)\\mathcal\{T\}\_\{\\beta\}\(x\)\+\\delta\_\{\\alpha\}\.Since𝒯β\(x\)≥exp\{κ\(RH2−β−\(3RH4\)2−β\)\}\>1\\mathcal\{T\}\_\{\\beta\}\(x\)\\geq\\exp\\\{\\kappa\(R\_\{H\}^\{2\-\\beta\}\-\(\\tfrac\{3R\_\{H\}\}\{4\}\)^\{2\-\\beta\}\)\\\}\>1on the far region andδα≤Cκ0α\\delta\_\{\\alpha\}\\leq C\\kappa\_\{0\}\\alpha, choosingκ0\\kappa\_\{0\}small absorbs the last term and again gives a factor1−cα1\-c\\alpha\. In both tail cases,α≥αm−1\\alpha\\geq\\alpha^\{m\-1\}; taking the supremum overeeand decreasing the constant proves the stated estimate\. ∎
###### Proof of Lemma[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3)\.
The estimates in Lemmas[A\.7](https://arxiv.org/html/2607.16384#A1.Thmtheorem7),[A\.8](https://arxiv.org/html/2607.16384#A1.Thmtheorem8), and[A\.9](https://arxiv.org/html/2607.16384#A1.Thmtheorem9)give \([4\.9](https://arxiv.org/html/2607.16384#S4.E9)\) after takingc0:=min\{ccore,cbr,cfar\}c\_\{0\}:=\\min\\\{c\_\{\\rm core\},c\_\{\\rm br\},c\_\{\\rm far\}\\\}and reducing the common stepsize threshold\.
It remains to prove the growth estimate \([4\.10](https://arxiv.org/html/2607.16384#S4.E10)\)\. On‖x‖≤RH\\left\\lVert x\\right\\rVert\\leq R\_\{H\}, it is exactly \([A\.8](https://arxiv.org/html/2607.16384#A1.E8)\)\. If‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\}andβ=2\\beta=2, thenVα\(x\)=1V\_\{\\alpha\}\(x\)=1andVα\(fξ\+\(x\)\)≤1\+δα≤\(1\+Cα\)Vα\(x\)V\_\{\\alpha\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\\leq 1\+\\delta\_\{\\alpha\}\\leq\(1\+C\\alpha\)V\_\{\\alpha\}\(x\)\. If‖x‖≥RH\\left\\lVert x\\right\\rVert\\geq R\_\{H\}and1≤β<21\\leq\\beta<2, then Lemma[A\.5](https://arxiv.org/html/2607.16384#A1.Thmtheorem5)and0≤ω≤10\\leq\\omega\\leq 1give
𝔼Vα\(fξ\+\(x\)\)≤\(1−cα\)𝒯β\(x\)\+δα≤\(1\+Cα\)Vα\(x\),\\mathbb\{E\}V\_\{\\alpha\}\(f\_\{\\xi^\{\+\}\}\(x\)\)\\leq\(1\-c\\alpha\)\\mathcal\{T\}\_\{\\beta\}\(x\)\+\\delta\_\{\\alpha\}\\leq\(1\+C\\alpha\)V\_\{\\alpha\}\(x\),again after reducing the threshold and increasingCC\. This proves Lemma[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3)\. ∎
### A\.5Proof of Lemma[4\.4](https://arxiv.org/html/2607.16384#S4.Thmtheorem4)
SetρΞ:=𝔼LΦ\(U1\)<1,\\rho\_\{\\Xi\}:=\\mathbb\{E\}L\_\{\\Phi\}\(U\_\{1\}\)<1,where the inequality follows from Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1)\.
###### Proof of Lemma[4\.4](https://arxiv.org/html/2607.16384#S4.Thmtheorem4)\.
LetCVC\_\{V\}be as in Lemma[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3)\. LetC\+C\_\{\+\}be as in Lemma[A\.6](https://arxiv.org/html/2607.16384#A1.Thmtheorem6)\. Chooseα0∈\(0,1\]\\alpha\_\{0\}\\in\(0,1\]small enough that the conclusions of Lemmas[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3)and[A\.6](https://arxiv.org/html/2607.16384#A1.Thmtheorem6)hold for0<α≤α00<\\alpha\\leq\\alpha\_\{0\}and
ρΞ\+C\+α\+α2Lg,Φ\(1\+CVα\)≤1−c0αm−1,0<α≤α0\.\\rho\_\{\\Xi\}\+C\_\{\+\}\\alpha\+\\alpha^\{2\}L\_\{g,\\Phi\}\(1\+C\_\{V\}\\alpha\)\\leq 1\-c\_\{0\}\\alpha^\{m\-1\},\\qquad 0<\\alpha\\leq\\alpha\_\{0\}\.\(A\.20\)This is possible becauseρΞ<1\\rho\_\{\\Xi\}<1andαm−1→0\\alpha^\{m\-1\}\\to 0\.
Fixz=\(x,ξ\)z=\(x,\\xi\),z′=\(y,η\)z^\{\\prime\}=\(y,\\eta\), and an absolutely continuous pathγ\(t\)=\(x\(t\),ξ\(t\)\)\\gamma\(t\)=\(x\(t\),\\xi\(t\)\)fromzztoz′z^\{\\prime\}\. For fixeduu, write
ξ\+\(t\):=Φ\(ξ\(t\),u\),X\+\(t\):=x\(t\)−α\{h\(x\(t\)\)\+g\(x\(t\),ξ\+\(t\)\)\}\.\\xi^\{\+\}\(t\):=\\Phi\(\\xi\(t\),u\),\\qquad X^\{\+\}\(t\):=x\(t\)\-\\alpha\\\{h\(x\(t\)\)\+g\(x\(t\),\\xi^\{\+\}\(t\)\)\\\}\.The curveFu∘γ=\(X\+\(⋅\),ξ\+\(⋅\)\)F\_\{u\}\\circ\\gamma=\(X^\{\+\}\(\\cdot\),\\xi^\{\+\}\(\\cdot\)\)is absolutely continuous\. Indeed,Φ\(⋅,u\)\\Phi\(\\cdot,u\)is Lipschitz by Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1), the composed dependenceg\(x,Φ\(ξ,u\)\)g\(x,\\Phi\(\\xi,u\)\)is Lipschitz inξ\\xiby Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N2\)](https://arxiv.org/html/2607.16384#S2.I3.i2), and the mapx↦h\(x\)\+g\(x,ζ\)x\\mapsto h\(x\)\+g\(x,\\zeta\)has derivativeA\(x,ζ\)A\(x,\\zeta\), which is uniformly bounded by \([4\.5](https://arxiv.org/html/2607.16384#S4.E5)\)\.
Letmdα\\operatorname\{md\}\_\{\\alpha\}denote the metric derivative with respect to the augmented base norm\|⋅\|α\|\\cdot\|\_\{\\alpha\}\. For a\.e\.tt,
mdα\(Fu∘γ\)\(t\)\\displaystyle\\operatorname\{md\}\_\{\\alpha\}\(F\_\{u\}\\circ\\gamma\)\(t\)≤‖\(Id−αA\(x\(t\),ξ\+\(t\)\)\)x˙\(t\)‖\\displaystyle\\leq\\left\\lVert\(I\_\{d\}\-\\alpha A\(x\(t\),\\xi^\{\+\}\(t\)\)\)\\dot\{x\}\(t\)\\right\\rVert\(A\.21\)\+\(LΦ\(u\)\+α2Lg,Φ\)α−1‖ξ˙\(t\)‖\.\\displaystyle\\qquad\+\(L\_\{\\Phi\}\(u\)\+\\alpha^\{2\}L\_\{g,\\Phi\}\)\\alpha^\{\-1\}\\left\\lVert\\dot\{\\xi\}\(t\)\\right\\rVert\.To prove \([A\.21](https://arxiv.org/html/2607.16384#A1.E21)\), compareFu\(γ\(s\)\)F\_\{u\}\(\\gamma\(s\)\)andFu\(γ\(t\)\)F\_\{u\}\(\\gamma\(t\)\)\. In thexx\-coordinate, first freeze the noise argument atξ\+\(t\)\\xi^\{\+\}\(t\); division by\|s−t\|\|s\-t\|and passage to the a\.e\. differentiability point gives the derivative\(Id−αA\(x\(t\),ξ\+\(t\)\)\)x˙\(t\)\(I\_\{d\}\-\\alpha A\(x\(t\),\\xi^\{\+\}\(t\)\)\)\\dot\{x\}\(t\)\. The remaining change in thexx\-coordinate is bounded byαLg,Φ‖ξ\(s\)−ξ\(t\)‖\\alpha L\_\{g,\\Phi\}\\left\\lVert\\xi\(s\)\-\\xi\(t\)\\right\\rVert\. In the noise coordinate, Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1)gives‖Φ\(ξ\(s\),u\)−Φ\(ξ\(t\),u\)‖≤LΦ\(u\)‖ξ\(s\)−ξ\(t\)‖\\left\\lVert\\Phi\(\\xi\(s\),u\)\-\\Phi\(\\xi\(t\),u\)\\right\\rVert\\leq L\_\{\\Phi\}\(u\)\\left\\lVert\\xi\(s\)\-\\xi\(t\)\\right\\rVert\. Dividing by\|s−t\|\|s\-t\|and using theα−1\\alpha^\{\-1\}\-weight in the augmented norm gives \([A\.21](https://arxiv.org/html/2607.16384#A1.E21)\)\.
Ifx˙\(t\)≠0\\dot\{x\}\(t\)\\neq 0, sete\(t\)=x˙\(t\)/‖x˙\(t\)‖e\(t\)=\\dot\{x\}\(t\)/\\left\\lVert\\dot\{x\}\(t\)\\right\\rVert; otherwise choose any measurable unit vectore\(t\)e\(t\)\. By the definition of induced length, \([A\.21](https://arxiv.org/html/2607.16384#A1.E21)\), Lemmas[4\.3](https://arxiv.org/html/2607.16384#S4.Thmtheorem3)and[A\.6](https://arxiv.org/html/2607.16384#A1.Thmtheorem6), and \([A\.20](https://arxiv.org/html/2607.16384#A1.E20)\),
𝔼dVα,α\(FU1\(z\),FU1\(z′\)\)\\displaystyle\\mathbb\{E\}\\,d\_\{V\_\{\\alpha\},\\alpha\}\\bigl\(F\_\{U\_\{1\}\}\(z\),F\_\{U\_\{1\}\}\(z^\{\\prime\}\)\\bigr\)≤∫01𝔼\[Vα\(X\+\(t\)\)‖\(Id−αA\(x\(t\),ξ\+\(t\)\)\)x˙\(t\)‖\]𝑑t\\displaystyle\\quad\\leq\\int\_\{0\}^\{1\}\\mathbb\{E\}\\\!\\left\[V\_\{\\alpha\}\(X^\{\+\}\(t\)\)\\left\\lVert\(I\_\{d\}\-\\alpha A\(x\(t\),\\xi^\{\+\}\(t\)\)\)\\dot\{x\}\(t\)\\right\\rVert\\right\]dt\+∫01𝔼\[\(LΦ\(U1\)\+α2Lg,Φ\)Vα\(X\+\(t\)\)\]α−1‖ξ˙\(t\)‖𝑑t\\displaystyle\\qquad\+\\int\_\{0\}^\{1\}\\mathbb\{E\}\\\!\\left\[\(L\_\{\\Phi\}\(U\_\{1\}\)\+\\alpha^\{2\}L\_\{g,\\Phi\}\)V\_\{\\alpha\}\(X^\{\+\}\(t\)\)\\right\]\\alpha^\{\-1\}\\left\\lVert\\dot\{\\xi\}\(t\)\\right\\rVert\\,dt≤\(1−c0αm−1\)∫01Vα\(x\(t\)\)\(‖x˙\(t\)‖\+α−1‖ξ˙\(t\)‖\)𝑑t\.\\displaystyle\\quad\\leq\(1\-c\_\{0\}\\alpha^\{m\-1\}\)\\int\_\{0\}^\{1\}V\_\{\\alpha\}\(x\(t\)\)\\left\(\\left\\lVert\\dot\{x\}\(t\)\\right\\rVert\+\\alpha^\{\-1\}\\left\\lVert\\dot\{\\xi\}\(t\)\\right\\rVert\\right\)dt\.Taking the infimum over all admissible pathsγ\\gammagives \([4\.1](https://arxiv.org/html/2607.16384#S4.E1)\)\. ∎
### A\.6Proof of Lemma[4\.5](https://arxiv.org/html/2607.16384#S4.Thmtheorem5)
###### Proof of Lemma[4\.5](https://arxiv.org/html/2607.16384#S4.Thmtheorem5)\.
Letz⋆=\(0,ξ⋆\)z\_\{\\star\}=\(0,\\xi\_\{\\star\}\), whereξ⋆\\xi\_\{\\star\}is the point from Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3)\. Write
ξ⋆\+:=Φ\(ξ⋆,U1\),xref\+:=−αg\(0,ξ⋆\+\)\.\\xi\_\{\\star\}^\{\+\}:=\\Phi\(\\xi\_\{\\star\},U\_\{1\}\),\\qquad x\_\{\\rm ref\}^\{\+\}:=\-\\alpha g\(0,\\xi\_\{\\star\}^\{\+\}\)\.Move fromz⋆=\(0,ξ⋆\)z\_\{\\star\}=\(0,\\xi\_\{\\star\}\)to\(0,ξ⋆\+\)\(0,\\xi\_\{\\star\}^\{\+\}\), and then from\(0,ξ⋆\+\)\(0,\\xi\_\{\\star\}^\{\+\}\)to\(xref\+,ξ⋆\+\)\(x\_\{\\rm ref\}^\{\+\},\\xi\_\{\\star\}^\{\+\}\)\. SinceVαV\_\{\\alpha\}depends only on thexx\-coordinate,
dVα,α\(FU1\(z⋆\),z⋆\)≤Vα\(0\)α−1‖ξ⋆\+−ξ⋆‖\+dVαX\(xref\+,0\)\.d\_\{V\_\{\\alpha\},\\alpha\}\\bigl\(F\_\{U\_\{1\}\}\(z\_\{\\star\}\),z\_\{\\star\}\\bigr\)\\leq V\_\{\\alpha\}\(0\)\\alpha^\{\-1\}\\left\\lVert\\xi\_\{\\star\}^\{\+\}\-\\xi\_\{\\star\}\\right\\rVert\+d\_\{V\_\{\\alpha\}\}^\{X\}\(x\_\{\\rm ref\}^\{\+\},0\)\.The first term is integrable by Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3)\.
Ifβ=2\\beta=2, then𝒯β≡1\\mathcal\{T\}\_\{\\beta\}\\equiv 1andVα≤1\+δαV\_\{\\alpha\}\\leq 1\+\\delta\_\{\\alpha\}\. Hence
dVαX\(xref\+,0\)≤\(1\+δα\)‖xref\+‖=\(1\+δα\)α‖g\(0,ξ⋆\+\)‖,d\_\{V\_\{\\alpha\}\}^\{X\}\(x\_\{\\rm ref\}^\{\+\},0\)\\leq\(1\+\\delta\_\{\\alpha\}\)\\left\\lVert x\_\{\\rm ref\}^\{\+\}\\right\\rVert=\(1\+\\delta\_\{\\alpha\}\)\\alpha\\left\\lVert g\(0,\\xi\_\{\\star\}^\{\+\}\)\\right\\rVert,whose expectation is finite by Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3)\.
Ifβ<2\\beta<2, Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)atx=0x=0andξ=ξ⋆\\xi=\\xi\_\{\\star\}gives an exponential moment ofg\(0,ξ⋆\+\)g\(0,\\xi\_\{\\star\}^\{\+\}\)\. Along the straight line from0toxref\+x\_\{\\rm ref\}^\{\+\},
dVαX\(xref\+,0\)≤α‖g\(0,ξ⋆\+\)‖\[1\+δα\+exp\{κ\(α2−β‖g\(0,ξ⋆\+\)‖2−β−\(3RH4\)2−β\)\+\}\]\.d\_\{V\_\{\\alpha\}\}^\{X\}\(x\_\{\\rm ref\}^\{\+\},0\)\\leq\\alpha\\left\\lVert g\(0,\\xi\_\{\\star\}^\{\+\}\)\\right\\rVert\\left\[1\+\\delta\_\{\\alpha\}\+\\exp\\\!\\left\\\{\\kappa\\bigl\(\\alpha^\{2\-\\beta\}\\left\\lVert g\(0,\\xi\_\{\\star\}^\{\+\}\)\\right\\rVert^\{2\-\\beta\}\-\(\\tfrac\{3R\_\{H\}\}\{4\}\)^\{2\-\\beta\}\\bigr\)\_\{\+\}\\right\\\}\\right\]\.Since2−β∈\(0,1\]2\-\\beta\\in\(0,1\],r2−β≤1\+rr^\{2\-\\beta\}\\leq 1\+rforr≥0r\\geq 0, we may chooseα0∈\(0,1\]\\alpha\_\{0\}\\in\(0,1\]such that, for0<α≤α00<\\alpha\\leq\\alpha\_\{0\}, the exponential factor above is dominated byCeλ‖g\(0,ξ⋆\+\)‖Ce^\{\\lambda\\left\\lVert g\(0,\\xi\_\{\\star\}^\{\+\}\)\\right\\rVert\}for someλ\>0\\lambda\>0below the exponential\-moment threshold supplied by[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)\. Thus𝔼dVαX\(xref\+,0\)<∞\.\\mathbb\{E\}d\_\{V\_\{\\alpha\}\}^\{X\}\(x\_\{\\rm ref\}^\{\+\},0\)<\\infty\.This proves \([4\.11](https://arxiv.org/html/2607.16384#S4.E11)\)\. ∎
### A\.7Proof of Corollary[2\.9](https://arxiv.org/html/2607.16384#S2.Thmtheorem9)
###### Proof of Corollary[2\.9](https://arxiv.org/html/2607.16384#S2.Thmtheorem9)\.
SinceVα≥1V\_\{\\alpha\}\\geq 1, the augmented induced metric dominates the Euclidean distance in theXX\-coordinate:dVα,α\(\(x,ξ\),\(y,η\)\)≥‖x−y‖\.d\_\{V\_\{\\alpha\},\\alpha\}\\bigl\(\(x,\\xi\),\(y,\\eta\)\\bigr\)\\geq\\left\\lVert x\-y\\right\\rVert\.Hence, for probability lawsμ,ν\\mu,\\nuon𝖹\\mathsf\{Z\},
W1\(μX,νX\)≤WVα,α\(μ,ν\)\.W\_\{1\}\(\\mu\_\{X\},\\nu\_\{X\}\)\\leq W\_\{V\_\{\\alpha\},\\alpha\}\(\\mu,\\nu\)\.Applying this withν=πα\\nu=\\pi\_\{\\alpha\}, and then using Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7), gives
W1\(\(μPαn\)X,\(πα\)X\)≤WVα,α\(μPαn,πα\)≤\(1−cαm−1\)nWVα,α\(μ,πα\)\.W\_\{1\}\\bigl\(\(\\mu P\_\{\\alpha\}^\{n\}\)\_\{X\},\(\\pi\_\{\\alpha\}\)\_\{X\}\\bigr\)\\leq W\_\{V\_\{\\alpha\},\\alpha\}\(\\mu P\_\{\\alpha\}^\{n\},\\pi\_\{\\alpha\}\)\\leq\(1\-c\\alpha^\{m\-1\}\)^\{n\}W\_\{V\_\{\\alpha\},\\alpha\}\(\\mu,\\pi\_\{\\alpha\}\)\.Sinceπα∈𝒫1\(𝖹,dVα,α\)\\pi\_\{\\alpha\}\\in\\mathcal\{P\}\_\{1\}\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\), the final quantity is finite if and only ifμ∈𝒫1\(𝖹,dVα,α\)\\mu\\in\\mathcal\{P\}\_\{1\}\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)\.
It remains to characterize𝒫1\(𝖹,dVα,α\)\\mathcal\{P\}\_\{1\}\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)\. Since\|\(x,ξ\)\|α=‖x‖\+α−1‖ξ‖,\|\(x,\\xi\)\|\_\{\\alpha\}=\\left\\lVert x\\right\\rVert\+\\alpha^\{\-1\}\\left\\lVert\\xi\\right\\rVert,and sinceVα≥1V\_\{\\alpha\}\\geq 1, projection of paths gives the lower bounds
dVα,α\(\(x,ξ\),\(0,ξ⋆\)\)≥dVαX\(x,0\),dVα,α\(\(x,ξ\),\(0,ξ⋆\)\)≥α−1‖ξ−ξ⋆‖\.d\_\{V\_\{\\alpha\},\\alpha\}\\bigl\(\(x,\\xi\),\(0,\\xi\_\{\\star\}\)\\bigr\)\\geq d\_\{V\_\{\\alpha\}\}^\{X\}\(x,0\),\\qquad d\_\{V\_\{\\alpha\},\\alpha\}\\bigl\(\(x,\\xi\),\(0,\\xi\_\{\\star\}\)\\bigr\)\\geq\\alpha^\{\-1\}\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert\.Conversely, moving first in thexx\-coordinate and then in theξ\\xi\-coordinate gives
dVα,α\(\(x,ξ\),\(0,ξ⋆\)\)≤dVαX\(x,0\)\+Vα\(0\)α−1‖ξ−ξ⋆‖\.d\_\{V\_\{\\alpha\},\\alpha\}\\bigl\(\(x,\\xi\),\(0,\\xi\_\{\\star\}\)\\bigr\)\\leq d\_\{V\_\{\\alpha\}\}^\{X\}\(x,0\)\+V\_\{\\alpha\}\(0\)\\alpha^\{\-1\}\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert\.ThusdVα,α\(\(x,ξ\),\(0,ξ⋆\)\)d\_\{V\_\{\\alpha\},\\alpha\}\(\(x,\\xi\),\(0,\\xi\_\{\\star\}\)\)\-integrability is equivalent to integrability ofdVαX\(x,0\)\+α−1‖ξ−ξ⋆‖\.d\_\{V\_\{\\alpha\}\}^\{X\}\(x,0\)\+\\alpha^\{\-1\}\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert\.
Ifβ=2\\beta=2, then𝒯β≡1\\mathcal\{T\}\_\{\\beta\}\\equiv 1and1≤Vα≤1\+δα1\\leq V\_\{\\alpha\}\\leq 1\+\\delta\_\{\\alpha\}\. Hence‖x−y‖≤dVαX\(x,y\)≤\(1\+δα\)‖x−y‖,\\left\\lVert x\-y\\right\\rVert\\leq d\_\{V\_\{\\alpha\}\}^\{X\}\(x,y\)\\leq\(1\+\\delta\_\{\\alpha\}\)\\left\\lVert x\-y\\right\\rVert,sodVαX\(x,0\)d\_\{V\_\{\\alpha\}\}^\{X\}\(x,0\)\-integrability is equivalent to‖x‖\\left\\lVert x\\right\\rVert\-integrability\.
If1≤β<21\\leq\\beta<2, letRHR\_\{H\}be as in Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)and define
Ψκ\(r\):=∫0rexp\{κ\(s2−β−\(3RH4\)2−β\)\+\}𝑑s\.\\Psi\_\{\\kappa\}\(r\):=\\int\_\{0\}^\{r\}\\exp\\\{\\kappa\\bigl\(s^\{2\-\\beta\}\-\(\\tfrac\{3R\_\{H\}\}\{4\}\)^\{2\-\\beta\}\\bigr\)\_\{\+\}\\\}\\,ds\.Because the correctionδαω\\delta\_\{\\alpha\}\\omegais compactly supported and0≤ω≤10\\leq\\omega\\leq 1,
Ψκ\(‖x‖\)≤dVαX\(x,0\)≤Ψκ\(‖x‖\)\+δαRH2\.\\Psi\_\{\\kappa\}\(\\left\\lVert x\\right\\rVert\)\\leq d\_\{V\_\{\\alpha\}\}^\{X\}\(x,0\)\\leq\\Psi\_\{\\kappa\}\(\\left\\lVert x\\right\\rVert\)\+\\frac\{\\delta\_\{\\alpha\}R\_\{H\}\}\{2\}\.Moreover,
Ψκ\(r\)∼e−κ\(3RH4\)2−βκ\(2−β\)rβ−1eκr2−β,r→∞\.\\Psi\_\{\\kappa\}\(r\)\\sim\\frac\{e^\{\-\\kappa\(\\frac\{3R\_\{H\}\}\{4\}\)^\{2\-\\beta\}\}\}\{\\kappa\(2\-\\beta\)\}r^\{\\beta\-1\}e^\{\\kappa r^\{2\-\\beta\}\},\\qquad r\\to\\infty\.ThereforedVαX\(x,0\)d\_\{V\_\{\\alpha\}\}^\{X\}\(x,0\)\-integrability is equivalent to
∫\(1\+‖x‖\)β−1exp\{κ‖x‖2−β\}μ\(dx,dξ\)<∞\.\\int\(1\+\\left\\lVert x\\right\\rVert\)^\{\\beta\-1\}\\exp\\\{\\kappa\\left\\lVert x\\right\\rVert^\{2\-\\beta\}\\\}\\,\\mu\(dx,d\\xi\)<\\infty\.Since2−β∈\(0,1\]2\-\\beta\\in\(0,1\], forr≥0r\\geq 0,r2−β≤\(1\+r\)2−β≤1\+r2−β,r^\{2\-\\beta\}\\leq\(1\+r\)^\{2\-\\beta\}\\leq 1\+r^\{2\-\\beta\},and hence
e−κeκr2−β≤exp\{κ\(\(1\+r\)2−β−1\)\}≤eκr2−β\.e^\{\-\\kappa\}e^\{\\kappa r^\{2\-\\beta\}\}\\leq\\exp\\\!\\left\\\{\\kappa\\bigl\(\(1\+r\)^\{2\-\\beta\}\-1\\bigr\)\\right\\\}\\leq e^\{\\kappa r^\{2\-\\beta\}\}\.Thus the preceding integrability condition is equivalent to integrability ofΓβ\(‖x‖\)\\Gamma\_\{\\beta\}\(\\left\\lVert x\\right\\rVert\)\. Combining this with theξ\\xi\-coordinate term gives the displayed description of𝒫1\(𝖹,dVα,α\)\\mathcal\{P\}\_\{1\}\(\\mathsf\{Z\},d\_\{V\_\{\\alpha\},\\alpha\}\)\. ∎
## Appendix BScaling\-limit estimates
This appendix proves the auxiliary results used for the scaling limit of invariant laws\. Constants denoted byC,cC,cmay change from line to line but are independent ofα\\alpha, unless stated otherwise\.
### B\.1Deterministic scaling estimates
###### Lemma B\.1\.
Under Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3), the following bounds hold\.
1. \(i\)If‖x‖≤RH2\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\}, then⟨x,h\(x\)⟩≥cinm−1‖x‖m\.\\left\\langle x,h\(x\)\\right\\rangle\\geq\\frac\{c\_\{\\rm in\}\}\{m\-1\}\\left\\lVert x\\right\\rVert^\{m\}\.
2. \(ii\)There existsCtail\>0C\_\{\\rm tail\}\>0such that ‖x‖β≤Ctail\(1\+⟨x,h\(x\)⟩\),1≤β<2,\\left\\lVert x\\right\\rVert^\{\\beta\}\\leq C\_\{\\rm tail\}\\bigl\(1\+\\left\\langle x,h\(x\)\\right\\rangle\\bigr\),\\qquad 1\\leq\\beta<2,and ‖x‖2≤Ctail\(1\+⟨x,h\(x\)⟩\),β=2\.\\left\\lVert x\\right\\rVert^\{2\}\\leq C\_\{\\rm tail\}\\bigl\(1\+\\left\\langle x,h\(x\)\\right\\rangle\\bigr\),\\qquad\\beta=2\.
3. \(iii\)Withr0:=min\{1,RH2\}r\_\{0\}:=\\min\\\{1,\\frac\{R\_\{H\}\}\{2\}\\\}, there existsc∗\>0c\_\{\*\}\>0such that ⟨x,h\(x\)⟩≥c∗,‖x‖≥r0\.\\left\\langle x,h\(x\)\\right\\rangle\\geq c\_\{\*\},\\qquad\\left\\lVert x\\right\\rVert\\geq r\_\{0\}\.
4. \(iv\)There existsCa\>0C\_\{a\}\>0such that, for allx∈ℝdx\\in\\mathbb\{R\}^\{d\},\(1\+‖x‖\)\(1\+‖x‖β−1\)≤Ca\(1\+⟨x,h\(x\)⟩\)\.\(1\+\\left\\lVert x\\right\\rVert\)\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\)\\leq C\_\{a\}\\bigl\(1\+\\left\\langle x,h\(x\)\\right\\rangle\\bigr\)\.
###### Proof\.
For the lower bound near the minimizer, useh\(x\)=∫01∇2H\(tx\)x𝑑t\.h\(x\)=\\int\_\{0\}^\{1\}\\nabla^\{2\}H\(tx\)x\\,dt\.Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)gives
⟨x,h\(x\)⟩=∫01⟨x,∇2H\(tx\)x⟩𝑑t≥cin‖x‖m∫01tm−2𝑑t=cinm−1‖x‖m\.\\left\\langle x,h\(x\)\\right\\rangle=\\int\_\{0\}^\{1\}\\left\\langle x,\\nabla^\{2\}H\(tx\)x\\right\\rangle\\,dt\\geq c\_\{\\rm in\}\\left\\lVert x\\right\\rVert^\{m\}\\int\_\{0\}^\{1\}t^\{m\-2\}\\,dt=\\frac\{c\_\{\\rm in\}\}\{m\-1\}\\left\\lVert x\\right\\rVert^\{m\}\.
If1≤β<21\\leq\\beta<2, the tail condition⟨x,h\(x\)⟩≥cout‖x‖β\\left\\langle x,h\(x\)\\right\\rangle\\geq c\_\{\\rm out\}\\left\\lVert x\\right\\rVert^\{\\beta\}on\{‖x‖≥RH\}\\\{\\left\\lVert x\\right\\rVert\\geq R\_\{H\}\\\}, together with boundedness on\{‖x‖<RH\}\\\{\\left\\lVert x\\right\\rVert<R\_\{H\}\\\}, gives‖x‖β≤C\(1\+⟨x,h\(x\)⟩\)\.\\left\\lVert x\\right\\rVert^\{\\beta\}\\leq C\\bigl\(1\+\\left\\langle x,h\(x\)\\right\\rangle\\bigr\)\.Ifβ=2\\beta=2, then for‖x‖≥2RH\\left\\lVert x\\right\\rVert\\geq 2R\_\{H\}, convexity and minimality at zero imply⟨x,h\(x/2\)⟩≥0\.\\left\\langle x,h\(x/2\)\\right\\rangle\\geq 0\.Using Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H3H3\-a\)](https://arxiv.org/html/2607.16384#S2.I1.i3.I1.i1)on the segment\{tx:1/2≤t≤1\}\\\{tx:1/2\\leq t\\leq 1\\\},
⟨x,h\(x\)⟩\\displaystyle\\left\\langle x,h\(x\)\\right\\rangle≥⟨x,h\(x\)−h\(x/2\)⟩\\displaystyle\\geq\\left\\langle x,h\(x\)\-h\(x/2\)\\right\\rangle=∫1/21⟨x,∇2H\(tx\)x⟩𝑑t≥cout2‖x‖2\.\\displaystyle=\\int\_\{1/2\}^\{1\}\\left\\langle x,\\nabla^\{2\}H\(tx\)x\\right\\rangle\\,dt\\geq\\frac\{c\_\{\\rm out\}\}\{2\}\\left\\lVert x\\right\\rVert^\{2\}\.The region\{‖x‖<2RH\}\\\{\\left\\lVert x\\right\\rVert<2R\_\{H\}\\\}is absorbed into the additive constant\.
Finally, writex=rux=ru, wherer=‖x‖r=\\left\\lVert x\\right\\rVertandu∈𝕊d−1u\\in\\mathbb\{S\}^\{d\-1\}\. Convexity ofHHand minimality at zero imply thats↦⟨u,h\(su\)⟩s\\mapsto\\left\\langle u,h\(su\)\\right\\rangleis nondecreasing on\[0,∞\)\[0,\\infty\)\. Forr≥r0r\\geq r\_\{0\},
⟨x,h\(x\)⟩=r⟨u,h\(ru\)⟩≥r0⟨u,h\(r0u\)⟩=⟨r0u,h\(r0u\)⟩≥cinm−1r0m\.\\left\\langle x,h\(x\)\\right\\rangle=r\\left\\langle u,h\(ru\)\\right\\rangle\\geq r\_\{0\}\\left\\langle u,h\(r\_\{0\}u\)\\right\\rangle=\\left\\langle r\_\{0\}u,h\(r\_\{0\}u\)\\right\\rangle\\geq\\frac\{c\_\{\\rm in\}\}\{m\-1\}r\_\{0\}^\{m\}\.This proves item \(iii\)\. For item \(iv\), ifβ=2\\beta=2, then\(1\+‖x‖\)\(1\+‖x‖β−1\)≤C\(1\+‖x‖2\),\(1\+\\left\\lVert x\\right\\rVert\)\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\)\\leq C\(1\+\\left\\lVert x\\right\\rVert^\{2\}\),and the quadratic\-tail estimate in item \(ii\) gives the result\. If1≤β<21\\leq\\beta<2, then\(1\+‖x‖\)\(1\+‖x‖β−1\)≤C\(1\+‖x‖β\),\(1\+\\left\\lVert x\\right\\rVert\)\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\)\\leq C\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\}\),and the subquadratic\-tail estimate in item \(ii\) gives the result\. This proves the lemma\. ∎
### B\.2Driving\-chain ergodicity and Poisson equation
###### Lemma B\.2\.
Assume Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1)and the reference\-point condition \([2\.4](https://arxiv.org/html/2607.16384#S2.E4)\)\. LetQQbe the transition kernel ofξn\+1=Φ\(ξn,Un\+1\)\.\\xi\_\{n\+1\}=\\Phi\(\\xi\_\{n\},U\_\{n\+1\}\)\.ThenQQadmits an invariant lawπΞ∈𝒫1\(Ξ\)\\pi\_\{\\Xi\}\\in\\mathcal\{P\}\_\{1\}\(\\Xi\), unique among all Borel invariant probability laws onΞ\\Xi\. Moreover,𝔼‖ξnξ−ξnη‖≤ρΞn‖ξ−η‖,\\mathbb\{E\}\\left\\lVert\\xi\_\{n\}^\{\\xi\}\-\\xi\_\{n\}^\{\\eta\}\\right\\rVert\\leq\\rho\_\{\\Xi\}^\{n\}\\left\\lVert\\xi\-\\eta\\right\\rVert,for synchronously coupled chains started fromξ\\xiandη\\eta, and
W1\(μQ,νQ\)≤ρΞW1\(μ,ν\),μ,ν∈𝒫1\(Ξ\),W\_\{1\}\(\\mu Q,\\nu Q\)\\leq\\rho\_\{\\Xi\}W\_\{1\}\(\\mu,\\nu\),\\qquad\\mu,\\nu\\in\\mathcal\{P\}\_\{1\}\(\\Xi\),and ifπα\\pi\_\{\\alpha\}is the invariant law of the augmented chain, then itsΞ\\Xi\-marginal isπΞ\\pi\_\{\\Xi\}\.
###### Proof\.
Letξ⋆\\xi\_\{\\star\}be as in Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N3\)](https://arxiv.org/html/2607.16384#S2.I3.i3)\. Ifη∼μ∈𝒫1\(Ξ\)\\eta\\sim\\mu\\in\\mathcal\{P\}\_\{1\}\(\\Xi\), then Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1)gives
𝔼‖Φ\(η,U1\)−ξ⋆‖\\displaystyle\\mathbb\{E\}\\left\\lVert\\Phi\(\\eta,U\_\{1\}\)\-\\xi\_\{\\star\}\\right\\rVert≤𝔼‖Φ\(η,U1\)−Φ\(ξ⋆,U1\)‖\+𝔼‖Φ\(ξ⋆,U1\)−ξ⋆‖\\displaystyle\\leq\\mathbb\{E\}\\left\\lVert\\Phi\(\\eta,U\_\{1\}\)\-\\Phi\(\\xi\_\{\\star\},U\_\{1\}\)\\right\\rVert\+\\mathbb\{E\}\\left\\lVert\\Phi\(\\xi\_\{\\star\},U\_\{1\}\)\-\\xi\_\{\\star\}\\right\\rVert≤ρΞ𝔼‖η−ξ⋆‖\+𝔼‖Φ\(ξ⋆,U1\)−ξ⋆‖\.\\displaystyle\\leq\\rho\_\{\\Xi\}\\mathbb\{E\}\\left\\lVert\\eta\-\\xi\_\{\\star\}\\right\\rVert\+\\mathbb\{E\}\\left\\lVert\\Phi\(\\xi\_\{\\star\},U\_\{1\}\)\-\\xi\_\{\\star\}\\right\\rVert\.ThusQQmaps𝒫1\(Ξ\)\\mathcal\{P\}\_\{1\}\(\\Xi\)into itself\.
For any coupling\(η,η~\)\(\\eta,\\widetilde\{\\eta\}\)of\(μ,ν\)\(\\mu,\\nu\), drive both chains by the same innovationU1U\_\{1\}\. Then
𝔼‖Φ\(η,U1\)−Φ\(η~,U1\)‖≤ρΞ𝔼‖η−η~‖\.\\mathbb\{E\}\\left\\lVert\\Phi\(\\eta,U\_\{1\}\)\-\\Phi\(\\widetilde\{\\eta\},U\_\{1\}\)\\right\\rVert\\leq\\rho\_\{\\Xi\}\\mathbb\{E\}\\left\\lVert\\eta\-\\widetilde\{\\eta\}\\right\\rVert\.Iteration with independent innovations gives the stated synchronousnn\-step estimate\. Taking the infimum over couplings yieldsW1\(μQ,νQ\)≤ρΞW1\(μ,ν\)\.W\_\{1\}\(\\mu Q,\\nu Q\)\\leq\\rho\_\{\\Xi\}W\_\{1\}\(\\mu,\\nu\)\.SinceΞ\\Xiis closed in a finite\-dimensional Euclidean space,\(𝒫1\(Ξ\),W1\)\(\\mathcal\{P\}\_\{1\}\(\\Xi\),W\_\{1\}\)is complete\. Banach’s fixed\-point theorem gives a unique invariant lawπΞ∈𝒫1\(Ξ\)\\pi\_\{\\Xi\}\\in\\mathcal\{P\}\_\{1\}\(\\Xi\)\.
To prove uniqueness among all Borel invariant laws, letρ\\rhobe such a law onΞ\\Xi\. For each fixedξ\\xi, the contraction estimate givesδξQn→πΞ\\delta\_\{\\xi\}Q^\{n\}\\to\\pi\_\{\\Xi\}inW1W\_\{1\}\. Hence, for every bounded Lipschitzφ\\varphi,Qnφ\(ξ\)→πΞ\(φ\)\.Q^\{n\}\\varphi\(\\xi\)\\to\\pi\_\{\\Xi\}\(\\varphi\)\.By invariance and dominated convergence,ρ\(φ\)=ρ\(Qnφ\)→πΞ\(φ\)\.\\rho\(\\varphi\)=\\rho\(Q^\{n\}\\varphi\)\\to\\pi\_\{\\Xi\}\(\\varphi\)\.Bounded Lipschitz functions determine Borel probability laws on the Polish spaceΞ\\Xi, soρ=πΞ\\rho=\\pi\_\{\\Xi\}\.
Finally, if\(X,ξ\)∼πα\(X,\\xi\)\\sim\\pi\_\{\\alpha\}, invariance of the augmented chain implies thatΦ\(ξ,U1\)\\Phi\(\\xi,U\_\{1\}\)has the same law asξ\\xi\. Thus theΞ\\Xi\-marginal ofπα\\pi\_\{\\alpha\}is invariant forQQand equalsπΞ\\pi\_\{\\Xi\}\. ∎
###### Lemma B\.3\.
Assume the hypotheses of Lemma[B\.2](https://arxiv.org/html/2607.16384#A2.Thmtheorem2)\. Letf:Ξ→ℝf:\\Xi\\to\\mathbb\{R\}be bounded, Lipschitz, and centered underπΞ\\pi\_\{\\Xi\}\. Defineu\(ξ\):=∑k=0∞Qkf\(ξ\)\.u\(\\xi\):=\\sum\_\{k=0\}^\{\\infty\}Q^\{k\}f\(\\xi\)\.Then the series converges absolutely for everyξ\\xi,uuis Lipschitz withLip\(u\)≤Lip\(f\)1−ρΞ,\\operatorname\{Lip\}\(u\)\\leq\\frac\{\\operatorname\{Lip\}\(f\)\}\{1\-\\rho\_\{\\Xi\}\},u∈L2\(πΞ\)u\\in L^\{2\}\(\\pi\_\{\\Xi\}\), andu−Qu=f\.u\-Qu=f\.
###### Proof\.
Letξkξ\\xi\_\{k\}^\{\\xi\}be the driving chain started fromξ\\xi\. Letη0∼πΞ\\eta\_\{0\}\\sim\\pi\_\{\\Xi\}and drive the stationary chain\(ηk\)\(\\eta\_\{k\}\)by the same innovations\. Sinceffis centered underπΞ\\pi\_\{\\Xi\},Qkf\(ξ\)=𝔼\[f\(ξkξ\)−f\(ηk\)\]\.Q^\{k\}f\(\\xi\)=\\mathbb\{E\}\[f\(\\xi\_\{k\}^\{\\xi\}\)\-f\(\\eta\_\{k\}\)\]\.Therefore\|Qkf\(ξ\)\|≤min\{2‖f‖∞,Lip\(f\)ρΞk\(‖ξ‖\+𝔼πΞ‖η‖\)\}\.\\left\\lvert Q^\{k\}f\(\\xi\)\\right\\rvert\\leq\\min\\left\\\{2\\left\\lVert f\\right\\rVert\_\{\\infty\},\\,\\operatorname\{Lip\}\(f\)\\rho\_\{\\Xi\}^\{k\}\\bigl\(\\left\\lVert\\xi\\right\\rVert\+\\mathbb\{E\}\_\{\\pi\_\{\\Xi\}\}\\left\\lVert\\eta\\right\\rVert\\bigr\)\\right\\\}\.This bound proves absolute convergence\. Similarly, synchronous coupling of two chains started fromξ\\xiandη\\etagives\|Qkf\(ξ\)−Qkf\(η\)\|≤Lip\(f\)ρΞk‖ξ−η‖,\\left\\lvert Q^\{k\}f\(\\xi\)\-Q^\{k\}f\(\\eta\)\\right\\rvert\\leq\\operatorname\{Lip\}\(f\)\\rho\_\{\\Xi\}^\{k\}\\left\\lVert\\xi\-\\eta\\right\\rVert,and summing overkkproves the Lipschitz bound foruu\.
The first bound also implies logarithmic growth:\|u\(ξ\)\|≤Cf\(1\+log\(1\+‖ξ‖\)\)\.\\left\\lvert u\(\\xi\)\\right\\rvert\\leq C\_\{f\}\\bigl\(1\+\\log\(1\+\\left\\lVert\\xi\\right\\rVert\)\\bigr\)\.SinceπΞ∈𝒫1\(Ξ\)\\pi\_\{\\Xi\}\\in\\mathcal\{P\}\_\{1\}\(\\Xi\)andlog2\(1\+r\)≤C\(1\+r\)\\log^\{2\}\(1\+r\)\\leq C\(1\+r\), we getu∈L2\(πΞ\)u\\in L^\{2\}\(\\pi\_\{\\Xi\}\)\. Finally, absolute convergence allows termwise application ofQQ, giving
Qu=∑k=1∞Qkf,u−Qu=f\.Qu=\\sum\_\{k=1\}^\{\\infty\}Q^\{k\}f,\\qquad u\-Qu=f\.∎
### B\.3Proof of Lemma[4\.6](https://arxiv.org/html/2607.16384#S4.Thmtheorem6)
###### Proof of Lemma[4\.6](https://arxiv.org/html/2607.16384#S4.Thmtheorem6)\.
Letξkξ\\xi\_\{k\}^\{\\xi\}andξkη\\xi\_\{k\}^\{\\eta\}be two copies of the driving chain, started fromξ\\xiandη\\etaand coupled through the same innovations\. By Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1),𝔼‖ξkξ−ξkη‖≤ρΞk‖ξ−η‖\.\\mathbb\{E\}\\left\\lVert\\xi\_\{k\}^\{\\xi\}\-\\xi\_\{k\}^\{\\eta\}\\right\\rVert\\leq\\rho\_\{\\Xi\}^\{k\}\\left\\lVert\\xi\-\\eta\\right\\rVert\.Fork≥1k\\geq 1, Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N2\)](https://arxiv.org/html/2607.16384#S2.I3.i2)gives
‖Qkgx\(ξ\)−Qkgx\(η\)‖≤Lg,ΦρΞk−1‖ξ−η‖\.\\left\\lVert Q^\{k\}g\_\{x\}\(\\xi\)\-Q^\{k\}g\_\{x\}\(\\eta\)\\right\\rVert\\leq L\_\{g,\\Phi\}\\rho\_\{\\Xi\}^\{k\-1\}\\left\\lVert\\xi\-\\eta\\right\\rVert\.Letη∼πΞ\\eta\\sim\\pi\_\{\\Xi\}, independent of the driving innovations\. SinceπΞgx=0\\pi\_\{\\Xi\}g\_\{x\}=0by \([2\.9](https://arxiv.org/html/2607.16384#S2.E9)\),
Qkgx\(ξ\)=𝔼\[gx\(ξkξ\)−gx\(ξkη\)\],k≥1\.Q^\{k\}g\_\{x\}\(\\xi\)=\\mathbb\{E\}\[g\_\{x\}\(\\xi\_\{k\}^\{\\xi\}\)\-g\_\{x\}\(\\xi\_\{k\}^\{\\eta\}\)\],\\qquad k\\geq 1\.Consequently
‖Qkgx\(ξ\)‖≤Lg,ΦρΞk−1\(‖ξ−ξ⋆‖\+𝔼πΞ‖η−ξ⋆‖\),\\left\\lVert Q^\{k\}g\_\{x\}\(\\xi\)\\right\\rVert\\leq L\_\{g,\\Phi\}\\rho\_\{\\Xi\}^\{k\-1\}\\bigl\(\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert\+\\mathbb\{E\}\_\{\\pi\_\{\\Xi\}\}\\left\\lVert\\eta\-\\xi\_\{\\star\}\\right\\rVert\\bigr\),and the series definingχx\\chi\_\{x\}converges absolutely for everyξ\\xi\. ApplyingQQtermwise givesχx−Qχx=Qgx\\chi\_\{x\}\-Q\\chi\_\{x\}=Qg\_\{x\}\. The same estimate gives
‖χx\(ξ\)‖≤CχRχ\(ξ\),Rχ\(ξ\):=1\+‖ξ−ξ⋆‖,\\left\\lVert\\chi\_\{x\}\(\\xi\)\\right\\rVert\\leq C\_\{\\chi\}R\_\{\\chi\}\(\\xi\),\\qquad R\_\{\\chi\}\(\\xi\):=1\+\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert,\(B\.1\)withRχ∈L2\(πΞ\)R\_\{\\chi\}\\in L^\{2\}\(\\pi\_\{\\Xi\}\)by \([2\.10](https://arxiv.org/html/2607.16384#S2.E10)\)\. Centering ofχx\\chi\_\{x\}follows from invariance ofπΞ\\pi\_\{\\Xi\}:∫Ξχx\(ξ\)πΞ\(dξ\)=0,\\int\_\{\\Xi\}\\chi\_\{x\}\(\\xi\)\\,\\pi\_\{\\Xi\}\(d\\xi\)=0,because∫Qkgx𝑑πΞ=∫gx𝑑πΞ=0\\int Q^\{k\}g\_\{x\}\\,d\\pi\_\{\\Xi\}=\\int g\_\{x\}\\,d\\pi\_\{\\Xi\}=0for everyk≥1k\\geq 1\.
It remains to control the dependence onxx\. The uniform derivative bound \([4\.5](https://arxiv.org/html/2607.16384#S4.E5)\) shows that, for someLx\>0L\_\{x\}\>0,
‖gx\(ζ\)−gy\(ζ\)‖≤Lx‖x−y‖,x,y∈ℝd,ζ∈Ξ\.\\left\\lVert g\_\{x\}\(\\zeta\)\-g\_\{y\}\(\\zeta\)\\right\\rVert\\leq L\_\{x\}\\left\\lVert x\-y\\right\\rVert,\\qquad x,y\\in\\mathbb\{R\}^\{d\},\\ \\zeta\\in\\Xi\.Fixp∈\(0,1\)p\\in\(0,1\), and putfx,y:=gx−gyf\_\{x,y\}:=g\_\{x\}\-g\_\{y\}\. By \([2\.9](https://arxiv.org/html/2607.16384#S2.E9)\),πΞfx,y=0\\pi\_\{\\Xi\}f\_\{x,y\}=0\. On the one hand,
‖Qkfx,y\(ξ\)‖≤Lx‖x−y‖,k≥1\.\\left\\lVert Q^\{k\}f\_\{x,y\}\(\\xi\)\\right\\rVert\\leq L\_\{x\}\\left\\lVert x\-y\\right\\rVert,\\qquad k\\geq 1\.On the other hand, comparison with the stationary chain used above and Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N2\)](https://arxiv.org/html/2607.16384#S2.I3.i2), applied separately togxg\_\{x\}andgyg\_\{y\}, give
‖Qkfx,y\(ξ\)‖\\displaystyle\\left\\lVert Q^\{k\}f\_\{x,y\}\(\\xi\)\\right\\rVert≤2Lg,ΦρΞk−1\(‖ξ−ξ⋆‖\+𝔼πΞ‖η−ξ⋆‖\)\\displaystyle\\leq 2L\_\{g,\\Phi\}\\rho\_\{\\Xi\}^\{k\-1\}\\bigl\(\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert\+\\mathbb\{E\}\_\{\\pi\_\{\\Xi\}\}\\left\\lVert\\eta\-\\xi\_\{\\star\}\\right\\rVert\\bigr\)≤CρΞk−1Rχ\(ξ\)\.\\displaystyle\\leq C\\rho\_\{\\Xi\}^\{k\-1\}R\_\{\\chi\}\(\\xi\)\.Sincemin\{a,b\}≤apb1−p\\min\\\{a,b\\\}\\leq a^\{p\}b^\{1\-p\}fora,b≥0a,b\\geq 0, summing these two bounds overk≥1k\\geq 1yields the Hölder estimate
‖χx\(ξ\)−χy\(ξ\)‖≤Cp‖x−y‖pRχ\(ξ\)1−p,x,y∈ℝd,ξ∈Ξ\.\\left\\lVert\\chi\_\{x\}\(\\xi\)\-\\chi\_\{y\}\(\\xi\)\\right\\rVert\\leq C\_\{p\}\\left\\lVert x\-y\\right\\rVert^\{p\}R\_\{\\chi\}\(\\xi\)^\{1\-p\},\\qquad x,y\\in\\mathbb\{R\}^\{d\},\\ \\xi\\in\\Xi\.\(B\.2\)The martingale\-difference property \([4\.13](https://arxiv.org/html/2607.16384#S4.E13)\) is immediate from \([4\.12](https://arxiv.org/html/2607.16384#S4.E12)\)\.
Finally, letf\(ξ\)=g\(0,ξ\)f\(\\xi\)=g\(0,\\xi\)\. The geometric estimate used to prove \([B\.1](https://arxiv.org/html/2607.16384#A2.E1)\), withx=0x=0, gives
‖Qkf\(ξ\)‖≤CρΞk−1\(1\+‖ξ−ξ⋆‖\),k≥1\.\\left\\lVert Q^\{k\}f\(\\xi\)\\right\\rVert\\leq C\\rho\_\{\\Xi\}^\{k\-1\}\(1\+\\left\\lVert\\xi\-\\xi\_\{\\star\}\\right\\rVert\),\\qquad k\\geq 1\.Together with \([2\.10](https://arxiv.org/html/2607.16384#S2.E10)\), this implies absolute summability of the lag covariances\. Write the stationary chain as\(ξn\)n∈ℤ\(\\xi\_\{n\}\)\_\{n\\in\\mathbb\{Z\}\}, setfn=f\(ξn\)f\_\{n\}=f\(\\xi\_\{n\}\)andχn=χ0\(ξn\)\\chi\_\{n\}=\\chi\_\{0\}\(\\xi\_\{n\}\), and defineDn:=fn\+1\+χn\+1−χn\.D\_\{n\}:=f\_\{n\+1\}\+\\chi\_\{n\+1\}\-\\chi\_\{n\}\.Then\(Dn\)\(D\_\{n\}\)is a martingale\-difference sequence andfn\+1=Dn\+χn−χn\+1\.f\_\{n\+1\}=D\_\{n\}\+\\chi\_\{n\}\-\\chi\_\{n\+1\}\.Therefore, forN≥1N\\geq 1,∑n=1Nfn=∑n=0N−1Dn\+χ0−χN\.\\sum\_\{n=1\}^\{N\}f\_\{n\}=\\sum\_\{n=0\}^\{N\-1\}D\_\{n\}\+\\chi\_\{0\}\-\\chi\_\{N\}\.Sinceχ0∈L2\(πΞ\)\\chi\_\{0\}\\in L^\{2\}\(\\pi\_\{\\Xi\}\), the telescoping term is negligible after division byNNin the covariance of the partial sums\. Hence
limN→∞1NVar\(∑n=1Nfn\)=𝔼\[D0D0⊤\]=Σ\.\\lim\_\{N\\to\\infty\}\\frac\{1\}\{N\}\\operatorname\{Var\}\\\!\\left\(\\sum\_\{n=1\}^\{N\}f\_\{n\}\\right\)=\\mathbb\{E\}\[D\_\{0\}D\_\{0\}^\{\\top\}\]=\\Sigma\.On the other hand, stationarity and absolute summability of the lag covariances give
limN→∞1NVar\(∑n=1Nfn\)=Γ0\+∑k≥1\(Γk\+Γk⊤\)\.\\lim\_\{N\\to\\infty\}\\frac\{1\}\{N\}\\operatorname\{Var\}\\\!\\left\(\\sum\_\{n=1\}^\{N\}f\_\{n\}\\right\)=\\Gamma\_\{0\}\+\\sum\_\{k\\geq 1\}\(\\Gamma\_\{k\}\+\\Gamma\_\{k\}^\{\\top\}\)\.This proves \([2\.11](https://arxiv.org/html/2607.16384#S2.E11)\) and completes the proof of the lemma\. ∎
### B\.4Proof of Lemma[4\.7](https://arxiv.org/html/2607.16384#S4.Thmtheorem7)
###### Lemma B\.4\.
Assume Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. Then there exist constantsA,B\>0A,B\>0, independent ofα\\alpha, such that𝔼‖g\(Xα,ξα\+\)‖2≤A\+B𝔼⟨Xα,h\(Xα\)⟩\.\\mathbb\{E\}\\left\\lVert g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}\\leq A\+B\\,\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\.
###### Proof\.
By Lemma[B\.2](https://arxiv.org/html/2607.16384#A2.Thmtheorem2),ξα\+∼πΞ\\xi\_\{\\alpha\}^\{\+\}\\sim\\pi\_\{\\Xi\}\.
Ifβ=2\\beta=2, \([4\.5](https://arxiv.org/html/2607.16384#S4.E5)\) gives a uniformxx\-Lipschitz constant\(2−θ\)L𝖦/\(1−θ\)\(2\-\\theta\)L\_\{\\mathsf\{G\}\}/\(1\-\\theta\)forgg\. Hence
‖g\(x,ζ\)‖2≤2‖g\(0,ζ\)‖2\+2\(\(2−θ\)L𝖦1−θ\)2‖x‖2\.\\left\\lVert g\(x,\\zeta\)\\right\\rVert^\{2\}\\leq 2\\left\\lVert g\(0,\\zeta\)\\right\\rVert^\{2\}\+2\\left\(\\frac\{\(2\-\\theta\)L\_\{\\mathsf\{G\}\}\}\{1\-\\theta\}\\right\)^\{2\}\\left\\lVert x\\right\\rVert^\{2\}\.The stationary second\-moment condition \([2\.10](https://arxiv.org/html/2607.16384#S2.E10)\) gives𝔼‖g\(0,ξα\+\)‖2<∞,\\mathbb\{E\}\\left\\lVert g\(0,\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}<\\infty,and Lemma[B\.1](https://arxiv.org/html/2607.16384#A2.Thmtheorem1)gives‖x‖2≤C\(1\+⟨x,h\(x\)⟩\)\.\\left\\lVert x\\right\\rVert^\{2\}\\leq C\(1\+\\left\\langle x,h\(x\)\\right\\rangle\)\.Taking expectations proves the claim in the quadratic\-tail case\.
If1≤β<21\\leq\\beta<2, Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)gives a uniform second moment forg\(x,Φ\(ξ,U1\)\)1\+‖x‖β−1\.\\frac\{g\(x,\\Phi\(\\xi,U\_\{1\}\)\)\}\{1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\}\.Therefore, conditionally on\(Xα,ξα\)\(X\_\{\\alpha\},\\xi\_\{\\alpha\}\),
𝔼\[‖g\(Xα,ξα\+\)‖2∣Xα,ξα\]≤C\(1\+‖Xα‖β−1\)2≤C\(1\+‖Xα‖β\)\.\\mathbb\{E\}\\bigl\[\\left\\lVert g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}\\mid X\_\{\\alpha\},\\xi\_\{\\alpha\}\\bigr\]\\leq C\(1\+\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{\\beta\-1\}\)^\{2\}\\leq C\(1\+\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{\\beta\}\)\.Using again Lemma[B\.1](https://arxiv.org/html/2607.16384#A2.Thmtheorem1),‖x‖β≤C\(1\+⟨x,h\(x\)⟩\),\\left\\lVert x\\right\\rVert^\{\\beta\}\\leq C\(1\+\\left\\langle x,h\(x\)\\right\\rangle\),and taking expectations proves the claim\. ∎
###### Lemma B\.5\.
Assume Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. For all sufficiently smallα\\alpha,
𝔼‖Xα‖2<∞,𝔼‖h\(Xα\)\+g\(Xα,ξα\+\)‖2<∞\.\\mathbb\{E\}\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{2\}<\\infty,\\qquad\\mathbb\{E\}\\left\\lVert h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}<\\infty\.Moreover,
𝔼⟨Xα,g\(Xα,ξα\+\)⟩≥−θ𝔼⟨Xα,h\(Xα\)⟩−Cα\(1\+𝔼⟨Xα,h\(Xα\)⟩\)\.\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle\\geq\-\\theta\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\-C\\alpha\\left\(1\+\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\\right\)\.\(B\.3\)
###### Proof\.
We first justify the second moments\. If1≤β<21\\leq\\beta<2, the induced first\-moment bound in Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7)implies thatXαX\_\{\\alpha\}has finite moments of every polynomial order\. Lemma[B\.4](https://arxiv.org/html/2607.16384#A2.Thmtheorem4)and the global Lipschitz bound onhhfrom Lemma[4\.2](https://arxiv.org/html/2607.16384#S4.Thmtheorem2)then give𝔼‖h\(Xα\)\+g\(Xα,ξα\+\)‖2<∞\.\\mathbb\{E\}\\left\\lVert h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}<\\infty\.
Forβ=2\\beta=2, Lemma[B\.1](https://arxiv.org/html/2607.16384#A2.Thmtheorem1)gives the global coercivity bound
⟨x,h\(x\)⟩≥c‖x‖2−C,x∈ℝd\.\\left\\langle x,h\(x\)\\right\\rangle\\geq c\\left\\lVert x\\right\\rVert^\{2\}\-C,\\qquad x\\in\\mathbb\{R\}^\{d\}\.\(B\.4\)Indeed, outside a sufficiently large ball this follows from the quadratic\-tail lower Hessian bound, while the remaining bounded region is absorbed into the additive constant\.
Start the augmented chain fromX0=0X\_\{0\}=0andξ0∼πΞ\\xi\_\{0\}\\sim\\pi\_\{\\Xi\}\. Thenξn∼πΞ\\xi\_\{n\}\\sim\\pi\_\{\\Xi\}for allnn\. Define the martingale\-corrected quadratic Lyapunov functionJα0\(x,ξ\):=‖x‖2−2α⟨x,χ0\(ξ\)⟩\.J\_\{\\alpha\}^\{0\}\(x,\\xi\):=\\left\\lVert x\\right\\rVert^\{2\}\-2\\alpha\\left\\langle x,\\chi\_\{0\}\(\\xi\)\\right\\rangle\.For any law whoseΞ\\Xi\-marginal isπΞ\\pi\_\{\\Xi\}, \([B\.1](https://arxiv.org/html/2607.16384#A2.E1)\), Young’s inequality, andRχ∈L2\(πΞ\)R\_\{\\chi\}\\in L^\{2\}\(\\pi\_\{\\Xi\}\)give
12𝔼‖X‖2−Cα2≤𝔼Jα0\(X,ξ\)≤32𝔼‖X‖2\+Cα2\.\\frac\{1\}\{2\}\\mathbb\{E\}\\left\\lVert X\\right\\rVert^\{2\}\-C\\alpha^\{2\}\\leq\\mathbb\{E\}J\_\{\\alpha\}^\{0\}\(X,\\xi\)\\leq\\frac\{3\}\{2\}\\mathbb\{E\}\\left\\lVert X\\right\\rVert^\{2\}\+C\\alpha^\{2\}\.\(B\.5\)For finitenn, the second moments below are finite by induction fromX0=0X\_\{0\}=0, the finite variance ofg\(0,ξ\)g\(0,\\xi\)underπΞ\\pi\_\{\\Xi\}, and the global Lipschitz bounds in Lemma[4\.2](https://arxiv.org/html/2607.16384#S4.Thmtheorem2)\. Write
𝖦n:=𝖦\(Xn,ξn\+1\),Xn\+1=Xn−α𝖦n\.\\mathsf\{G\}\_\{n\}:=\\mathsf\{G\}\(X\_\{n\},\\xi\_\{n\+1\}\),\\qquad X\_\{n\+1\}=X\_\{n\}\-\\alpha\\mathsf\{G\}\_\{n\}\.The mean\-perturbation condition \([2\.7](https://arxiv.org/html/2607.16384#S2.E7)\), withy=0y=0, implies
𝔼\[⟨Xn,𝖦\(Xn,ξn\+1\)−𝖦\(0,ξn\+1\)⟩\|Xn,ξn\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\left\\langle X\_\{n\},\\mathsf\{G\}\(X\_\{n\},\\xi\_\{n\+1\}\)\-\\mathsf\{G\}\(0,\\xi\_\{n\+1\}\)\\right\\rangle\\,\\middle\|\\,X\_\{n\},\\xi\_\{n\}\\right\]=⟨Xn,h\(Xn\)\+g¯\(Xn,ξn\)−g¯\(0,ξn\)⟩≥\(1−θ\)⟨Xn,h\(Xn\)⟩\.\\displaystyle\\qquad=\\left\\langle X\_\{n\},h\(X\_\{n\}\)\+\\bar\{g\}\(X\_\{n\},\\xi\_\{n\}\)\-\\bar\{g\}\(0,\\xi\_\{n\}\)\\right\\rangle\\geq\(1\-\\theta\)\\left\\langle X\_\{n\},h\(X\_\{n\}\)\\right\\rangle\.The Poisson identity at the minimizer gives𝔼⟨Xn,g\(0,ξn\+1\)⟩=𝔼⟨Xn,χ0\(ξn\)−χ0\(ξn\+1\)⟩\.\\mathbb\{E\}\\left\\langle X\_\{n\},g\(0,\\xi\_\{n\+1\}\)\\right\\rangle=\\mathbb\{E\}\\left\\langle X\_\{n\},\\chi\_\{0\}\(\\xi\_\{n\}\)\-\\chi\_\{0\}\(\\xi\_\{n\+1\}\)\\right\\rangle\.ExpandingJα0J\_\{\\alpha\}^\{0\}, usingg\(Xn,ξn\+1\)=g\(0,ξn\+1\)\+𝖦\(Xn,ξn\+1\)−𝖦\(0,ξn\+1\)−h\(Xn\)g\(X\_\{n\},\\xi\_\{n\+1\}\)=g\(0,\\xi\_\{n\+1\}\)\+\\mathsf\{G\}\(X\_\{n\},\\xi\_\{n\+1\}\)\-\\mathsf\{G\}\(0,\\xi\_\{n\+1\}\)\-h\(X\_\{n\}\), and applying the preceding cancellation yield
𝔼\[Jα0\(Xn\+1,ξn\+1\)−Jα0\(Xn,ξn\)\]\\displaystyle\\mathbb\{E\}\[J\_\{\\alpha\}^\{0\}\(X\_\{n\+1\},\\xi\_\{n\+1\}\)\-J\_\{\\alpha\}^\{0\}\(X\_\{n\},\\xi\_\{n\}\)\]=−2α𝔼⟨Xn,𝖦\(Xn,ξn\+1\)−𝖦\(0,ξn\+1\)⟩\\displaystyle\\quad=\-2\\alpha\\mathbb\{E\}\\left\\langle X\_\{n\},\\mathsf\{G\}\(X\_\{n\},\\xi\_\{n\+1\}\)\-\\mathsf\{G\}\(0,\\xi\_\{n\+1\}\)\\right\\rangle\+α2𝔼‖𝖦n‖2\+2α2𝔼⟨𝖦n,χ0\(ξn\+1\)⟩\.\\displaystyle\\quad\+\\alpha^\{2\}\\mathbb\{E\}\\left\\lVert\\mathsf\{G\}\_\{n\}\\right\\rVert^\{2\}\+2\\alpha^\{2\}\\mathbb\{E\}\\left\\langle\\mathsf\{G\}\_\{n\},\\chi\_\{0\}\(\\xi\_\{n\+1\}\)\\right\\rangle\.The global Lipschitz bounds in Lemma[4\.2](https://arxiv.org/html/2607.16384#S4.Thmtheorem2), together with𝔼πΞ‖g\(0,ξ\)‖2<∞\\mathbb\{E\}\_\{\\pi\_\{\\Xi\}\}\\left\\lVert g\(0,\\xi\)\\right\\rVert^\{2\}<\\infty, imply𝔼‖𝖦n‖2≤C\(1\+𝔼‖Xn‖2\)\.\\mathbb\{E\}\\left\\lVert\\mathsf\{G\}\_\{n\}\\right\\rVert^\{2\}\\leq C\\bigl\(1\+\\mathbb\{E\}\\left\\lVert X\_\{n\}\\right\\rVert^\{2\}\\bigr\)\.Moreover, Cauchy–Schwarz andχ0∈L2\(πΞ\)\\chi\_\{0\}\\in L^\{2\}\(\\pi\_\{\\Xi\}\)give
\|𝔼⟨𝖦n,χ0\(ξn\+1\)⟩\|≤\(𝔼‖𝖦n‖2\)1/2\(𝔼πΞ‖χ0‖2\)1/2≤C\(1\+𝔼‖Xn‖2\)\.\\left\\lvert\\mathbb\{E\}\\left\\langle\\mathsf\{G\}\_\{n\},\\chi\_\{0\}\(\\xi\_\{n\+1\}\)\\right\\rangle\\right\\rvert\\leq\\bigl\(\\mathbb\{E\}\\left\\lVert\\mathsf\{G\}\_\{n\}\\right\\rVert^\{2\}\\bigr\)^\{1/2\}\\bigl\(\\mathbb\{E\}\_\{\\pi\_\{\\Xi\}\}\\left\\lVert\\chi\_\{0\}\\right\\rVert^\{2\}\\bigr\)^\{1/2\}\\leq C\\bigl\(1\+\\mathbb\{E\}\\left\\lVert X\_\{n\}\\right\\rVert^\{2\}\\bigr\)\.Combining these bounds with \([B\.4](https://arxiv.org/html/2607.16384#A2.E4)\) yields
𝔼Jα0\(Xn\+1,ξn\+1\)≤𝔼Jα0\(Xn,ξn\)−cα𝔼‖Xn‖2\+Cα\+Cα2\(1\+𝔼‖Xn‖2\)\.\\mathbb\{E\}J\_\{\\alpha\}^\{0\}\(X\_\{n\+1\},\\xi\_\{n\+1\}\)\\leq\\mathbb\{E\}J\_\{\\alpha\}^\{0\}\(X\_\{n\},\\xi\_\{n\}\)\-c\\alpha\\mathbb\{E\}\\left\\lVert X\_\{n\}\\right\\rVert^\{2\}\+C\\alpha\+C\\alpha^\{2\}\\bigl\(1\+\\mathbb\{E\}\\left\\lVert X\_\{n\}\\right\\rVert^\{2\}\\bigr\)\.For sufficiently smallα\\alpha, \([B\.5](https://arxiv.org/html/2607.16384#A2.E5)\) absorbs theCα2𝔼‖Xn‖2C\\alpha^\{2\}\\mathbb\{E\}\\left\\lVert X\_\{n\}\\right\\rVert^\{2\}term and gives
𝔼Jα0\(Xn\+1,ξn\+1\)≤\(1−cα\)𝔼Jα0\(Xn,ξn\)\+Cα\.\\mathbb\{E\}J\_\{\\alpha\}^\{0\}\(X\_\{n\+1\},\\xi\_\{n\+1\}\)\\leq\(1\-c\\alpha\)\\mathbb\{E\}J\_\{\\alpha\}^\{0\}\(X\_\{n\},\\xi\_\{n\}\)\+C\\alpha\.\(B\.6\)Iterating \([B\.6](https://arxiv.org/html/2607.16384#A2.E6)\) and usingX0=0X\_\{0\}=0givessupn𝔼‖Xn‖2<∞\\sup\_\{n\}\\mathbb\{E\}\\left\\lVert X\_\{n\}\\right\\rVert^\{2\}<\\infty\. By Theorem[2\.7](https://arxiv.org/html/2607.16384#S2.Thmtheorem7), the finite\-time laws converge weakly toπα\\pi\_\{\\alpha\}\. Lower semicontinuity ofx↦‖x‖2x\\mapsto\\left\\lVert x\\right\\rVert^\{2\}then gives𝔼πα‖Xα‖2<∞\\mathbb\{E\}\_\{\\pi\_\{\\alpha\}\}\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{2\}<\\infty\. Lemma[B\.4](https://arxiv.org/html/2607.16384#A2.Thmtheorem4)and the global Lipschitz bound in Lemma[4\.2](https://arxiv.org/html/2607.16384#S4.Thmtheorem2)give the second moment ofh\(Xα\)\+g\(Xα,ξα\+\)h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\.
We now prove \([B\.3](https://arxiv.org/html/2607.16384#A2.E3)\)\. Lemma[B\.4](https://arxiv.org/html/2607.16384#A2.Thmtheorem4)and \([4\.6](https://arxiv.org/html/2607.16384#S4.E6)\) imply
𝔼‖h\(Xα\)\+g\(Xα,ξα\+\)‖2≤C\(1\+𝔼⟨Xα,h\(Xα\)⟩\)\.\\mathbb\{E\}\\left\\lVert h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}\\leq C\\left\(1\+\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\\right\)\.By stationarity and the Poisson identity at the minimizer,
𝔼⟨Xα,g\(0,ξα\+\)⟩\\displaystyle\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},g\(0,\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle=𝔼⟨Xα,χ0\(ξα\)−χ0\(ξα\+\)⟩\\displaystyle=\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},\\chi\_\{0\}\(\\xi\_\{\\alpha\}\)\-\\chi\_\{0\}\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle=𝔼⟨Xα\+−Xα,χ0\(ξα\+\)⟩\\displaystyle=\\mathbb\{E\}\\left\\langle X\_\{\\alpha\}^\{\+\}\-X\_\{\\alpha\},\\chi\_\{0\}\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle=−α𝔼⟨h\(Xα\)\+g\(Xα,ξα\+\),χ0\(ξα\+\)⟩\.\\displaystyle=\-\\alpha\\mathbb\{E\}\\left\\langle h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\),\\chi\_\{0\}\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle\.Thus Cauchy–Schwarz gives
𝔼⟨Xα,g\(0,ξα\+\)⟩≥−Cα\(𝔼‖h\(Xα\)\+g\(Xα,ξα\+\)‖2\)1/2≥−Cα\(1\+𝔼⟨Xα,h\(Xα\)⟩\)\.\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},g\(0,\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle\\geq\-C\\alpha\\left\(\\mathbb\{E\}\\left\\lVert h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}\\right\)^\{1/2\}\\geq\-C\\alpha\\left\(1\+\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\\right\)\.Finally, the mean\-perturbation condition withy=0y=0gives
𝔼⟨Xα,h\(Xα\)\+g\(Xα,ξα\+\)−g\(0,ξα\+\)⟩\\displaystyle\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\-g\(0,\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle≥\(1−θ\)𝔼⟨Xα,h\(Xα\)⟩\.\\displaystyle\\qquad\\geq\(1\-\\theta\)\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\.Combining the last two displays and subtracting𝔼⟨Xα,h\(Xα\)⟩\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangleproves \([B\.3](https://arxiv.org/html/2607.16384#A2.E3)\)\. This proves the lemma\. ∎
###### Proof of Lemma[4\.7](https://arxiv.org/html/2607.16384#S4.Thmtheorem7)\.
By Lemma[B\.5](https://arxiv.org/html/2607.16384#A2.Thmtheorem5), the stationary square identity is justified\. SinceXα\+=dXαX\_\{\\alpha\}^\{\+\}\\stackrel\{\{\\scriptstyle d\}\}\{\{=\}\}X\_\{\\alpha\},
0=−2α𝔼⟨Xα,h\(Xα\)\+g\(Xα,ξα\+\)⟩\+α2𝔼‖h\(Xα\)\+g\(Xα,ξα\+\)‖2\.0=\-2\\alpha\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle\+\\alpha^\{2\}\\mathbb\{E\}\\left\\lVert h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}\.Using \([B\.3](https://arxiv.org/html/2607.16384#A2.E3)\), Lemma[B\.4](https://arxiv.org/html/2607.16384#A2.Thmtheorem4), and \([4\.6](https://arxiv.org/html/2607.16384#S4.E6)\), we obtain
\(1−θ\)𝔼⟨Xα,h\(Xα\)⟩≤Cα\+Cα𝔼⟨Xα,h\(Xα\)⟩\.\(1\-\\theta\)\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\\leq C\\alpha\+C\\alpha\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\.For smallα\\alpha, the last term is absorbed into the left\-hand side, giving𝔼⟨Xα,h\(Xα\)⟩≤Cα\.\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\\leq C\\alpha\.This proves \([4\.15](https://arxiv.org/html/2607.16384#S4.E15)\), and Lemma[B\.4](https://arxiv.org/html/2607.16384#A2.Thmtheorem4)then gives \([4\.18](https://arxiv.org/html/2607.16384#S4.E18)\)\.
The local moment and outside\-probability estimates now follow from the same drift lower bounds\. Letr0:=min\{1,RH2\}r\_\{0\}:=\\min\\\{1,\\frac\{R\_\{H\}\}\{2\}\\\}\. Lemma[B\.1](https://arxiv.org/html/2607.16384#A2.Thmtheorem1)gives
‖x‖m≤C⟨x,h\(x\)⟩,‖x‖≤r0,⟨x,h\(x\)⟩≥c∗,‖x‖≥r0\.\\left\\lVert x\\right\\rVert^\{m\}\\leq C\\left\\langle x,h\(x\)\\right\\rangle,\\qquad\\left\\lVert x\\right\\rVert\\leq r\_\{0\},\\qquad\\left\\langle x,h\(x\)\\right\\rangle\\geq c\_\{\*\},\\qquad\\left\\lVert x\\right\\rVert\\geq r\_\{0\}\.Therefore
𝔼\[‖Xα‖m𝟏\{‖Xα‖≤1\}\]\\displaystyle\\mathbb\{E\}\[\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{m\}\\mathbf\{1\}\_\{\\\{\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\leq 1\\\}\}\]≤C𝔼⟨Xα,h\(Xα\)⟩\+ℙ\(‖Xα‖≥r0\)\\displaystyle\\leq C\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\+\\mathbb\{P\}\(\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\geq r\_\{0\}\)≤Cα,\\displaystyle\\leq C\\alpha,and
ℙ\(‖Xα‖≥1\)≤ℙ\(‖Xα‖≥r0\)≤c∗−1𝔼⟨Xα,h\(Xα\)⟩≤Cα\.\\mathbb\{P\}\(\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\geq 1\)\\leq\\mathbb\{P\}\(\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\geq r\_\{0\}\)\\leq c\_\{\*\}^\{\-1\}\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\\leq C\\alpha\.This proves all estimates in Lemma[4\.7](https://arxiv.org/html/2607.16384#S4.Thmtheorem7)\. ∎
###### Lemma B\.6\.
Under the assumptions of Lemma[4\.7](https://arxiv.org/html/2607.16384#S4.Thmtheorem7), for everyη\>0\\eta\>0,
𝔼\[‖Yα\+−Yα‖2α2−2/m𝟏\{‖Yα\+−Yα‖\>η\}\]⟶0\.\\mathbb\{E\}\\left\[\\frac\{\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert^\{2\}\}\{\\alpha^\{2\-2/m\}\}\\mathbf\{1\}\_\{\\\{\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert\>\\eta\\\}\}\\right\]\\longrightarrow 0\.\(B\.7\)Moreover, for every fixedR\>0R\>0, the family
\{‖g\(α1/mYα,ξα\+\)‖2𝟏\{‖Yα‖≤R\}:0<α≤α0\}\\left\\\{\\left\\lVert g\(\\alpha^\{1/m\}Y\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}\\mathbf\{1\}\_\{\\\{\\left\\lVert Y\_\{\\alpha\}\\right\\rVert\\leq R\\\}\}:0<\\alpha\\leq\\alpha\_\{0\}\\right\\\}is uniformly integrable\.
###### Proof\.
We first prove the local uniform integrability statement\. On the event\{‖Yα‖≤R\}\\\{\\left\\lVert Y\_\{\\alpha\}\\right\\rVert\\leq R\\\}, the pointx=α1/mYαx=\\alpha^\{1/m\}Y\_\{\\alpha\}remains in a fixed compact set\. Ifβ=2\\beta=2, the uniformxx\-Lipschitz bound \([4\.5](https://arxiv.org/html/2607.16384#S4.E5)\) gives‖g\(x,ξα\+\)‖2≤CR\{1\+‖g\(0,ξα\+\)‖2\},\\left\\lVert g\(x,\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}\\leq C\_\{R\}\\\{1\+\\left\\lVert g\(0,\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}\\\},andξα\+∼πΞ\\xi\_\{\\alpha\}^\{\+\}\\sim\\pi\_\{\\Xi\}, so \([2\.10](https://arxiv.org/html/2607.16384#S2.E10)\) gives uniform integrability\. If1≤β<21\\leq\\beta<2, Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)gives a uniform exponential moment for the ratio‖g\(x,ξα\+\)‖1\+‖x‖β−1,\\frac\{\\left\\lVert g\(x,\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert\}\{1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\},while1\+‖x‖β−11\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}is bounded on the same compact set\. This again gives uniform integrability\.
For \([B\.7](https://arxiv.org/html/2607.16384#A2.E7)\), write
Yα\+−Yα=−α1−1/m\{h\(Xα\)\+g\(Xα,ξα\+\)\}\.Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}=\-\\alpha^\{1\-1/m\}\\\{h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\\}\.It is enough to prove
𝔼\[‖𝖦α‖2𝟏\{α1−1/m‖𝖦α‖\>η\}\]→0,𝖦α:=𝖦\(Xα,ξα\+\)\.\\mathbb\{E\}\\left\[\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert^\{2\}\\mathbf\{1\}\_\{\\\{\\alpha^\{1\-1/m\}\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert\>\\eta\\\}\}\\right\]\\to 0,\\qquad\\mathsf\{G\}\_\{\\alpha\}:=\\mathsf\{G\}\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\.\(B\.8\)
Consider firstβ=2\\beta=2\. The global Lipschitz bounds in Lemma[4\.2](https://arxiv.org/html/2607.16384#S4.Thmtheorem2)imply‖𝖦α‖≤C‖Xα‖\+‖g\(0,ξα\+\)‖\.\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert\\leq C\\left\\lVert X\_\{\\alpha\}\\right\\rVert\+\\left\\lVert g\(0,\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert\.Using\(a\+b\)2𝟏\{a\+b\>t\}≤4a2𝟏\{a\>t/2\}\+4b2𝟏\{b\>t/2\}\(a\+b\)^\{2\}\\mathbf\{1\}\_\{\\\{a\+b\>t\\\}\}\\leq 4a^\{2\}\\mathbf\{1\}\_\{\\\{a\>t/2\\\}\}\+4b^\{2\}\\mathbf\{1\}\_\{\\\{b\>t/2\\\}\}, the contribution ofg\(0,ξα\+\)g\(0,\\xi\_\{\\alpha\}^\{\+\}\)to \([B\.8](https://arxiv.org/html/2607.16384#A2.E8)\) tends to zero by dominated convergence underπΞ\\pi\_\{\\Xi\}\. For theXαX\_\{\\alpha\}\-term, the thresholdcηα1/m−1c\\eta\\alpha^\{1/m\-1\}tends to infinity\. For all sufficiently smallα\\alpha, it lies in the quadratic\-tail region, where‖x‖2≤C⟨x,h\(x\)⟩\\left\\lVert x\\right\\rVert^\{2\}\\leq C\\left\\langle x,h\(x\)\\right\\rangle\. Therefore
𝔼\[‖Xα‖2𝟏\{α1−1/mC‖Xα‖\>η\}\]≤C𝔼⟨Xα,h\(Xα\)⟩→0\\mathbb\{E\}\\\!\\bigl\[\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{2\}\\mathbf\{1\}\_\{\\\{\\alpha^\{1\-1/m\}C\\left\\lVert X\_\{\\alpha\}\\right\\rVert\>\\eta\\\}\}\\bigr\]\\leq C\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\\to 0by \([4\.15](https://arxiv.org/html/2607.16384#S4.E15)\)\. This verifies the Lindeberg condition in the quadratic\-tail case\.
Now assume1≤β<21\\leq\\beta<2, and setwα:=1\+‖Xα‖β−1\.w\_\{\\alpha\}:=1\+\\left\\lVert X\_\{\\alpha\}\\right\\rVert^\{\\beta\-1\}\.Since‖h\(x\)‖≤C\(1\+‖x‖β−1\)\\left\\lVert h\(x\)\\right\\rVert\\leq C\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\)globally, conditionally on\(Xα,ξα\)\(X\_\{\\alpha\},\\xi\_\{\\alpha\}\)we have
‖𝖦α‖≤Cwα\{1\+‖g\(Xα,ξα\+\)‖wα\}\.\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert\\leq Cw\_\{\\alpha\}\\left\\\{1\+\\frac\{\\left\\lVert g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert\}\{w\_\{\\alpha\}\}\\right\\\}\.The exponential moment in Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N5\)](https://arxiv.org/html/2607.16384#S2.I3.i5)gives a uniform conditional second moment and a uniform conditional exponential tail for the normalized factor\. Split according towα≤α−\(m−1\)/\(2m\)w\_\{\\alpha\}\\leq\\alpha^\{\-\(m\-1\)/\(2m\)\}\. On this event, the normalized factor must exceed a multiple ofα−\(m−1\)/\(2m\)\\alpha^\{\-\(m\-1\)/\(2m\)\}, so the contribution is bounded byC𝔼wα2e−cα−\(m−1\)/\(2m\)→0\.C\\mathbb\{E\}w\_\{\\alpha\}^\{2\}e^\{\-c\\alpha^\{\-\(m\-1\)/\(2m\)\}\}\\to 0\.On the complementary event, the conditional second\-moment bound gives a contribution at mostC𝔼\[wα2𝟏\{wα\>α−\(m−1\)/\(2m\)\}\]\.C\\mathbb\{E\}\\bigl\[w\_\{\\alpha\}^\{2\}\\mathbf\{1\}\_\{\\\{w\_\{\\alpha\}\>\\alpha^\{\-\(m\-1\)/\(2m\)\}\\\}\}\\bigr\]\.To show that this expectation vanishes, fix a sufficiently largeR0R\_\{0\}\. The event\{wα\>α−\(m−1\)/\(2m\)\}\\\{w\_\{\\alpha\}\>\\alpha^\{\-\(m\-1\)/\(2m\)\}\\\}is eventually contained in\{‖Xα‖≥R0\}\\\{\\left\\lVert X\_\{\\alpha\}\\right\\rVert\\geq R\_\{0\}\\\}\. On this region, Lemma[B\.1](https://arxiv.org/html/2607.16384#A2.Thmtheorem1)gives\(1\+‖x‖β−1\)2≤C\(1\+‖x‖β\)≤C⟨x,h\(x\)⟩\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\-1\}\)^\{2\}\\leq C\(1\+\\left\\lVert x\\right\\rVert^\{\\beta\}\)\\leq C\\left\\langle x,h\(x\)\\right\\rangle, after increasingR0R\_\{0\}if necessary\. Hence
𝔼\[wα2𝟏\{wα\>α−\(m−1\)/\(2m\)\}\]≤C𝔼⟨Xα,h\(Xα\)⟩→0\.\\mathbb\{E\}\\bigl\[w\_\{\\alpha\}^\{2\}\\mathbf\{1\}\_\{\\\{w\_\{\\alpha\}\>\\alpha^\{\-\(m\-1\)/\(2m\)\}\\\}\}\\bigr\]\\leq C\\mathbb\{E\}\\left\\langle X\_\{\\alpha\},h\(X\_\{\\alpha\}\)\\right\\rangle\\to 0\.This proves \([B\.8](https://arxiv.org/html/2607.16384#A2.E8)\), and hence \([B\.7](https://arxiv.org/html/2607.16384#A2.E7)\)\. ∎
### B\.5Proof of Lemma[4\.8](https://arxiv.org/html/2607.16384#S4.Thmtheorem8)
###### Proof of Lemma[4\.8](https://arxiv.org/html/2607.16384#S4.Thmtheorem8)\.
Letuube the Poisson solution from Lemma[B\.3](https://arxiv.org/html/2607.16384#A2.Thmtheorem3); thus
u−Qu=f,u∈L2\(πΞ\)\.u\-Qu=f,\\qquad u\\in L^\{2\}\(\\pi\_\{\\Xi\}\)\.Sinceψ\(Yα\)\\psi\(Y\_\{\\alpha\}\)is measurable with respect to\(Xα,ξα\)\(X\_\{\\alpha\},\\xi\_\{\\alpha\}\),
𝔼\[ψ\(Yα\)f\(ξα\)\]\\displaystyle\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)f\(\\xi\_\{\\alpha\}\)\]=𝔼\[ψ\(Yα\)u\(ξα\)\]−𝔼\[ψ\(Yα\)Qu\(ξα\)\]\\displaystyle=\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)u\(\\xi\_\{\\alpha\}\)\]\-\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)Qu\(\\xi\_\{\\alpha\}\)\]=𝔼\[ψ\(Yα\)u\(ξα\)\]−𝔼\[ψ\(Yα\)u\(ξα\+\)\]\.\\displaystyle=\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)u\(\\xi\_\{\\alpha\}\)\]\-\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)u\(\\xi\_\{\\alpha\}^\{\+\}\)\]\.Add and subtractψ\(Yα\+\)u\(ξα\+\)\\psi\(Y\_\{\\alpha\}^\{\+\}\)u\(\\xi\_\{\\alpha\}^\{\+\}\)\. Since\(Yα\+,ξα\+\)=d\(Yα,ξα\),\(Y\_\{\\alpha\}^\{\+\},\\xi\_\{\\alpha\}^\{\+\}\)\\stackrel\{\{\\scriptstyle d\}\}\{\{=\}\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}\),the first and third terms cancel, giving
𝔼\[ψ\(Yα\)f\(ξα\)\]=𝔼\[\(ψ\(Yα\+\)−ψ\(Yα\)\)u\(ξα\+\)\]\.\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)f\(\\xi\_\{\\alpha\}\)\]=\\mathbb\{E\}\[\(\\psi\(Y\_\{\\alpha\}^\{\+\}\)\-\\psi\(Y\_\{\\alpha\}\)\)u\(\\xi\_\{\\alpha\}^\{\+\}\)\]\.Therefore, by Cauchy–Schwarz,
\|𝔼\[ψ\(Yα\)f\(ξα\)\]\|≤Lip\(ψ\)\(𝔼‖Yα\+−Yα‖2\)1/2\(𝔼\|u\(ξα\+\)\|2\)1/2\.\\left\\lvert\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\}\)f\(\\xi\_\{\\alpha\}\)\]\\right\\rvert\\leq\\operatorname\{Lip\}\(\\psi\)\\bigl\(\\mathbb\{E\}\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert^\{2\}\\bigr\)^\{1/2\}\\bigl\(\\mathbb\{E\}\\left\\lvert u\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rvert^\{2\}\\bigr\)^\{1/2\}\.The second factor is finite and independent ofα\\alpha, becauseξα\+∼πΞ\\xi\_\{\\alpha\}^\{\+\}\\sim\\pi\_\{\\Xi\}\. For the first factor,
Yα\+−Yα=−α1−1/m\{h\(Xα\)\+g\(Xα,ξα\+\)\}\.Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}=\-\\alpha^\{1\-1/m\}\\\{h\(X\_\{\\alpha\}\)\+g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\\}\.Lemma[4\.7](https://arxiv.org/html/2607.16384#S4.Thmtheorem7)gives𝔼‖Yα\+−Yα‖2≤Cα2−2/m\.\\mathbb\{E\}\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert^\{2\}\\leq C\\alpha^\{2\-2/m\}\.This proves the desired bound\.
∎
### B\.6Proof of Lemma[4\.9](https://arxiv.org/html/2607.16384#S4.Thmtheorem9)
By \([4\.6](https://arxiv.org/html/2607.16384#S4.E6)\), \([4\.15](https://arxiv.org/html/2607.16384#S4.E15)\), and Lemma[B\.4](https://arxiv.org/html/2607.16384#A2.Thmtheorem4),
𝔼‖h\(Xα\)‖2≤Cα,sup0<α≤α0𝔼‖g\(Xα,ξα\+\)‖2<∞,𝔼\[‖h\(Xα\)‖‖g\(Xα,ξα\+\)‖\]⟶0\.\\mathbb\{E\}\\left\\lVert h\(X\_\{\\alpha\}\)\\right\\rVert^\{2\}\\leq C\\alpha,\\qquad\\sup\_\{0<\\alpha\\leq\\alpha\_\{0\}\}\\mathbb\{E\}\\left\\lVert g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert^\{2\}<\\infty,\\qquad\\mathbb\{E\}\\bigl\[\\left\\lVert h\(X\_\{\\alpha\}\)\\right\\rVert\\left\\lVert g\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rVert\\bigr\]\\longrightarrow 0\.\(B\.9\)The last conclusion follows from the first two by Cauchy–Schwarz\.
For the covariance calculation below, define, forx∈ℝdx\\in\\mathbb\{R\}^\{d\},
Mx\(ζ\):=gx\(ζ\)gx\(ζ\)⊤\+gx\(ζ\)χx\(ζ\)⊤\+χx\(ζ\)gx\(ζ\)⊤,ζ∈Ξ\.M\_\{x\}\(\\zeta\):=g\_\{x\}\(\\zeta\)g\_\{x\}\(\\zeta\)^\{\\top\}\+g\_\{x\}\(\\zeta\)\\chi\_\{x\}\(\\zeta\)^\{\\top\}\+\\chi\_\{x\}\(\\zeta\)g\_\{x\}\(\\zeta\)^\{\\top\},\\qquad\\zeta\\in\\Xi\.The next two lemmas justify replacingMα1/mYαM\_\{\\alpha^\{1/m\}Y\_\{\\alpha\}\}byM0M\_\{0\}on compact sets and handle integrable functions of the driving chain\.
###### Lemma B\.7\.
Under Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11), for every compact setK⊂ℝdK\\subset\\mathbb\{R\}^\{d\}, one hasQM0∈L1\(πΞ\)QM\_\{0\}\\in L^\{1\}\(\\pi\_\{\\Xi\}\)and
𝔼\[‖QMα1/mYα\(ξα\)−QM0\(ξα\)‖𝟏\{Yα∈K\}\]⟶0\.\\mathbb\{E\}\\left\[\\left\\lVert QM\_\{\\alpha^\{1/m\}Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}\)\-QM\_\{0\}\(\\xi\_\{\\alpha\}\)\\right\\rVert\\mathbf\{1\}\_\{\\\{Y\_\{\\alpha\}\\in K\\\}\}\\right\]\\longrightarrow 0\.\(B\.10\)
###### Proof\.
The uniformxx\-Lipschitz bound \([4\.5](https://arxiv.org/html/2607.16384#S4.E5)\) gives‖gx\(ζ\)−g0\(ζ\)‖≤C‖x‖\.\\left\\lVert g\_\{x\}\(\\zeta\)\-g\_\{0\}\(\\zeta\)\\right\\rVert\\leq C\\left\\lVert x\\right\\rVert\.Fixp∈\(0,1\)p\\in\(0,1\)\. The estimates for the solution of the Poisson equation in \([B\.1](https://arxiv.org/html/2607.16384#A2.E1)\) and \([B\.2](https://arxiv.org/html/2607.16384#A2.E2)\) give
‖χx\(ζ\)‖\+‖χ0\(ζ\)‖≤CRχ\(ζ\),‖χx\(ζ\)−χ0\(ζ\)‖≤Cp‖x‖pRχ\(ζ\)1−p\.\\left\\lVert\\chi\_\{x\}\(\\zeta\)\\right\\rVert\+\\left\\lVert\\chi\_\{0\}\(\\zeta\)\\right\\rVert\\leq CR\_\{\\chi\}\(\\zeta\),\\qquad\\left\\lVert\\chi\_\{x\}\(\\zeta\)\-\\chi\_\{0\}\(\\zeta\)\\right\\rVert\\leq C\_\{p\}\\left\\lVert x\\right\\rVert^\{p\}R\_\{\\chi\}\(\\zeta\)^\{1\-p\}\.ExpandingMx−M0M\_\{x\}\-M\_\{0\}, for‖x‖≤1\\left\\lVert x\\right\\rVert\\leq 1we obtain‖Mx\(ζ\)−M0\(ζ\)‖≤Cp\(‖x‖\+‖x‖p\)F\(ζ\),\\left\\lVert M\_\{x\}\(\\zeta\)\-M\_\{0\}\(\\zeta\)\\right\\rVert\\leq C\_\{p\}\(\\left\\lVert x\\right\\rVert\+\\left\\lVert x\\right\\rVert^\{p\}\)F\(\\zeta\),where
F\(ζ\):=1\+‖g0\(ζ\)‖\+Rχ\(ζ\)\+‖g0\(ζ\)‖Rχ\(ζ\)\.F\(\\zeta\):=1\+\\left\\lVert g\_\{0\}\(\\zeta\)\\right\\rVert\+R\_\{\\chi\}\(\\zeta\)\+\\left\\lVert g\_\{0\}\(\\zeta\)\\right\\rVert R\_\{\\chi\}\(\\zeta\)\.The stationary second moments in \([2\.10](https://arxiv.org/html/2607.16384#S2.E10)\) and Cauchy–Schwarz implyF∈L1\(πΞ\)F\\in L^\{1\}\(\\pi\_\{\\Xi\}\)\. In particular,M0∈L1\(πΞ\)M\_\{0\}\\in L^\{1\}\(\\pi\_\{\\Xi\}\), and invariance givesQM0∈L1\(πΞ\)QM\_\{0\}\\in L^\{1\}\(\\pi\_\{\\Xi\}\)\. IfKKis compact, then, for all sufficiently smallα\\alpha,α1/msupy∈K‖y‖≤1\\alpha^\{1/m\}\\sup\_\{y\\in K\}\\left\\lVert y\\right\\rVert\\leq 1\. On\{Yα∈K\}\\\{Y\_\{\\alpha\}\\in K\\\}, the preceding pointwise bound, invariance ofπΞ\\pi\_\{\\Xi\}, andξα∼πΞ\\xi\_\{\\alpha\}\\sim\\pi\_\{\\Xi\}give
𝔼\[‖QMα1/mYα\(ξα\)−QM0\(ξα\)‖𝟏\{Yα∈K\}\]\\displaystyle\\mathbb\{E\}\\left\[\\left\\lVert QM\_\{\\alpha^\{1/m\}Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}\)\-QM\_\{0\}\(\\xi\_\{\\alpha\}\)\\right\\rVert\\mathbf\{1\}\_\{\\\{Y\_\{\\alpha\}\\in K\\\}\}\\right\]≤Cp,K\(α1/m\+αp/m\)𝔼QF\(ξα\)=Cp,K\(α1/m\+αp/m\)∫ΞF𝑑πΞ⟶0\.\\displaystyle\\qquad\\leq C\_\{p,K\}\(\\alpha^\{1/m\}\+\\alpha^\{p/m\}\)\\mathbb\{E\}QF\(\\xi\_\{\\alpha\}\)=C\_\{p,K\}\(\\alpha^\{1/m\}\+\\alpha^\{p/m\}\)\\int\_\{\\Xi\}F\\,d\\pi\_\{\\Xi\}\\longrightarrow 0\.This proves \([B\.10](https://arxiv.org/html/2607.16384#A2.E10)\)\. ∎
###### Lemma B\.8\.
Assume Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11)\. Letαk↓0\\alpha\_\{k\}\\downarrow 0and supposeYαk⇒νY\_\{\\alpha\_\{k\}\}\\Rightarrow\\nu\. Ifψ:ℝd→ℝ\\psi:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}is bounded and globally Lipschitz, and ifΨ∈L1\(πΞ\)\\Psi\\in L^\{1\}\(\\pi\_\{\\Xi\}\), then
𝔼\[ψ\(Yαk\)Ψ\(ξαk\)\]⟶\(∫ℝdψ\(y\)ν\(dy\)\)\(∫ΞΨ\(ξ\)πΞ\(dξ\)\)\.\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\_\{k\}\}\)\\Psi\(\\xi\_\{\\alpha\_\{k\}\}\)\]\\longrightarrow\\left\(\\int\_\{\\mathbb\{R\}^\{d\}\}\\psi\(y\)\\,\\nu\(dy\)\\right\)\\left\(\\int\_\{\\Xi\}\\Psi\(\\xi\)\\,\\pi\_\{\\Xi\}\(d\\xi\)\\right\)\.
###### Proof\.
The constant part ofΨ\\Psiis handled by weak convergence ofYαkY\_\{\\alpha\_\{k\}\}, because theΞ\\Xi\-marginal ofπαk\\pi\_\{\\alpha\_\{k\}\}isπΞ\\pi\_\{\\Xi\}by Lemma[B\.2](https://arxiv.org/html/2607.16384#A2.Thmtheorem2)\. It remains to consider centeredΨ\\Psi, i\.e\.πΞ\(Ψ\)=0\\pi\_\{\\Xi\}\(\\Psi\)=0\. Bounded Lipschitz functions are dense inL1\(πΞ\)L^\{1\}\(\\pi\_\{\\Xi\}\)on the closed Euclidean setΞ\\Xi\. Thus, for everyε\>0\\varepsilon\>0, choose a bounded Lipschitzffsuch that‖Ψ−f‖L1\(πΞ\)<ε\\left\\lVert\\Psi\-f\\right\\rVert\_\{L^\{1\}\(\\pi\_\{\\Xi\}\)\}<\\varepsilon\. Replacingffbyf−πΞ\(f\)f\-\\pi\_\{\\Xi\}\(f\)increases this error by at most anotherε\\varepsilon, so we may assumeπΞ\(f\)=0\\pi\_\{\\Xi\}\(f\)=0\. Sinceξαk∼πΞ\\xi\_\{\\alpha\_\{k\}\}\\sim\\pi\_\{\\Xi\},
\|𝔼\[ψ\(Yαk\)\(Ψ−f\)\(ξαk\)\]\|≤2‖ψ‖∞ε\.\\left\\lvert\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\_\{k\}\}\)\(\\Psi\-f\)\(\\xi\_\{\\alpha\_\{k\}\}\)\]\\right\\rvert\\leq 2\\left\\lVert\\psi\\right\\rVert\_\{\\infty\}\\varepsilon\.Lemma[4\.8](https://arxiv.org/html/2607.16384#S4.Thmtheorem8)gives𝔼\[ψ\(Yαk\)f\(ξαk\)\]→0\\mathbb\{E\}\[\\psi\(Y\_\{\\alpha\_\{k\}\}\)f\(\\xi\_\{\\alpha\_\{k\}\}\)\]\\to 0\. Lettingε↓0\\varepsilon\\downarrow 0proves the centered case and hence the lemma\.
∎
###### Lemma B\.9\.
Under Assumptions[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)and[2\.11](https://arxiv.org/html/2607.16384#S2.Thmtheorem11), for everyφ∈Cc3\(ℝd\)\\varphi\\in C\_\{c\}^\{3\}\(\\mathbb\{R\}^\{d\}\),
0=−𝔼⟨hα\(Yα\),∇φ\(Yα\)⟩\+12𝔼tr\(QMα1/mYα\(ξα\)∇2φ\(Yα\)\)\+o\(1\),0=\-\\mathbb\{E\}\\left\\langle h\_\{\\alpha\}\(Y\_\{\\alpha\}\),\\nabla\\varphi\(Y\_\{\\alpha\}\)\\right\\rangle\+\\frac\{1\}\{2\}\\mathbb\{E\}\\operatorname\{tr\}\\\!\\left\(QM\_\{\\alpha^\{1/m\}Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}\)\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\}\)\\right\)\+o\(1\),\(B\.11\)where theo\(1\)o\(1\)is deterministic and tends to zero asα↓0\\alpha\\downarrow 0\.
###### Proof\.
Write
Gα:=gXα\(ξα\+\),𝖦α:=𝖦\(Xα,ξα\+\)\.G\_\{\\alpha\}:=g\_\{X\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}^\{\+\}\),\\qquad\\mathsf\{G\}\_\{\\alpha\}:=\\mathsf\{G\}\(X\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\.Then
Yα\+−Yα=−α2−2/mhα\(Yα\)−α1−1/mGα=−α1−1/m𝖦α\.Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}=\-\\alpha^\{2\-2/m\}h\_\{\\alpha\}\(Y\_\{\\alpha\}\)\-\\alpha^\{1\-1/m\}G\_\{\\alpha\}=\-\\alpha^\{1\-1/m\}\\mathsf\{G\}\_\{\\alpha\}\.LetK0:=supp\(∇φ\)K\_\{0\}:=\\operatorname\{supp\}\(\\nabla\\varphi\), and letK1:=\{y∈ℝd:dist\(y,K0\)≤1\}\.K\_\{1\}:=\\\{y\\in\\mathbb\{R\}^\{d\}:\\operatorname\{dist\}\(y,K\_\{0\}\)\\leq 1\\\}\.IfK0=∅K\_\{0\}=\\varnothing, the conclusion is immediate, so assume otherwise\. Define
χ~α,y\(ξ\):=\{χα1/my\(ξ\),y∈K1,0,y∉K1\.\\widetilde\{\\chi\}\_\{\\alpha,y\}\(\\xi\):=\\begin\{cases\}\\chi\_\{\\alpha^\{1/m\}y\}\(\\xi\),&y\\in K\_\{1\},\\\\ 0,&y\\notin K\_\{1\}\.\\end\{cases\}SetAα\(y,ξ\):=⟨∇φ\(y\),χ~α,y\(ξ\)⟩\.A\_\{\\alpha\}\(y,\\xi\):=\\left\\langle\\nabla\\varphi\(y\),\\widetilde\{\\chi\}\_\{\\alpha,y\}\(\\xi\)\\right\\rangle\.This is well defined because∇φ=0\\nabla\\varphi=0outsideK0K\_\{0\}, and\|Aα\(y,ξ\)\|≤C‖∇φ‖∞Rχ\(ξ\)\\left\\lvert A\_\{\\alpha\}\(y,\\xi\)\\right\\rvert\\leq C\\left\\lVert\\nabla\\varphi\\right\\rVert\_\{\\infty\}R\_\{\\chi\}\(\\xi\)\. Introduce the localized perturbed test function
φα\(y,ξ\):=φ\(y\)−α1−1/mAα\(y,ξ\)\.\\varphi\_\{\\alpha\}\(y,\\xi\):=\\varphi\(y\)\-\\alpha^\{1\-1/m\}A\_\{\\alpha\}\(y,\\xi\)\.\(B\.12\)The integrability required below follows from \([4\.18](https://arxiv.org/html/2607.16384#S4.E18)\),Rχ∈L2\(πΞ\)R\_\{\\chi\}\\in L^\{2\}\(\\pi\_\{\\Xi\}\), and \([B\.1](https://arxiv.org/html/2607.16384#A2.E1)\)\. Stationarity gives
𝔼\[φα\(Yα\+,ξα\+\)−φα\(Yα,ξα\)\]=0\.\\mathbb\{E\}\[\\varphi\_\{\\alpha\}\(Y\_\{\\alpha\}^\{\+\},\\xi\_\{\\alpha\}^\{\+\}\)\-\\varphi\_\{\\alpha\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}\)\]=0\.\(B\.13\)
First consider the uncorrected part\. Taylor’s formula andYα\+−Yα=−α2−2/mhα\(Yα\)−α1−1/mGαY\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}=\-\\alpha^\{2\-2/m\}h\_\{\\alpha\}\(Y\_\{\\alpha\}\)\-\\alpha^\{1\-1/m\}G\_\{\\alpha\}give
φ\(Yα\+\)−φ\(Yα\)α2−2/m\\displaystyle\\frac\{\\varphi\(Y\_\{\\alpha\}^\{\+\}\)\-\\varphi\(Y\_\{\\alpha\}\)\}\{\\alpha^\{2\-2/m\}\}=−⟨hα\(Yα\),∇φ\(Yα\)⟩−α1/m−1⟨∇φ\(Yα\),Gα⟩\\displaystyle=\-\\left\\langle h\_\{\\alpha\}\(Y\_\{\\alpha\}\),\\nabla\\varphi\(Y\_\{\\alpha\}\)\\right\\rangle\-\\alpha^\{1/m\-1\}\\left\\langle\\nabla\\varphi\(Y\_\{\\alpha\}\),G\_\{\\alpha\}\\right\\rangle\(B\.14\)\+12tr\(GαGα⊤∇2φ\(Yα\)\)\+rα\(1\),\\displaystyle\\quad\+\\frac\{1\}\{2\}\\operatorname\{tr\}\\\!\\left\(G\_\{\\alpha\}G\_\{\\alpha\}^\{\\top\}\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\}\)\\right\)\+r\_\{\\alpha\}^\{\(1\)\},with𝔼\|rα\(1\)\|→0\\mathbb\{E\}\|r\_\{\\alpha\}^\{\(1\)\}\|\\to 0\. The terms in the quadratic expansion containinghαh\_\{\\alpha\}are negligible because
α2−2/m‖hα\(Yα\)‖2=‖h\(Xα\)‖2,α1−1/m‖hα\(Yα\)‖‖Gα‖=‖h\(Xα\)‖‖Gα‖,\\alpha^\{2\-2/m\}\\left\\lVert h\_\{\\alpha\}\(Y\_\{\\alpha\}\)\\right\\rVert^\{2\}=\\left\\lVert h\(X\_\{\\alpha\}\)\\right\\rVert^\{2\},\\qquad\\alpha^\{1\-1/m\}\\left\\lVert h\_\{\\alpha\}\(Y\_\{\\alpha\}\)\\right\\rVert\\left\\lVert G\_\{\\alpha\}\\right\\rVert=\\left\\lVert h\(X\_\{\\alpha\}\)\\right\\rVert\\left\\lVert G\_\{\\alpha\}\\right\\rVert,and \([B\.9](https://arxiv.org/html/2607.16384#A2.E9)\) makes both expectations tend to zero\. The third\-order Taylor remainder is controlled by the uniform continuity of∇2φ\\nabla^\{2\}\\varphi: for everyη\>0\\eta\>0, after division byα2−2/m\\alpha^\{2\-2/m\}its absolute value is bounded by
η‖Yα\+−Yα‖2α2−2/m\+Cη‖Yα\+−Yα‖2α2−2/m𝟏\{‖Yα\+−Yα‖\>η\}\.\\eta\\frac\{\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert^\{2\}\}\{\\alpha^\{2\-2/m\}\}\+C\_\{\\eta\}\\frac\{\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert^\{2\}\}\{\\alpha^\{2\-2/m\}\}\\mathbf\{1\}\_\{\\\{\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert\>\\eta\\\}\}\.The first term is bounded in expectation uniformly inα\\alphaby \([4\.18](https://arxiv.org/html/2607.16384#S4.E18)\); the second tends to zero by \([B\.7](https://arxiv.org/html/2607.16384#A2.E7)\)\. Sendingη↓0\\eta\\downarrow 0proves𝔼\|rα\(1\)\|→0\\mathbb\{E\}\|r\_\{\\alpha\}^\{\(1\)\}\|\\to 0\.
Now expand the term involving the solution of the Poisson equation in \([B\.12](https://arxiv.org/html/2607.16384#A2.E12)\)\. In
−α1/m−1𝔼\[Aα\(Yα\+,ξα\+\)−Aα\(Yα,ξα\)\],\-\\alpha^\{1/m\-1\}\\mathbb\{E\}\\left\[A\_\{\\alpha\}\(Y\_\{\\alpha\}^\{\+\},\\xi\_\{\\alpha\}^\{\+\}\)\-A\_\{\\alpha\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}\)\\right\],add and subtractAα\(Yα,ξα\+\)A\_\{\\alpha\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\. SinceAα\(Yα,⋅\)A\_\{\\alpha\}\(Y\_\{\\alpha\},\\cdot\)can be nonzero only whenYα∈K0Y\_\{\\alpha\}\\in K\_\{0\}, the first resulting increment is
−α1/m−1𝔼\[Aα\(Yα,ξα\+\)−Aα\(Yα,ξα\)\]\\displaystyle\-\\alpha^\{1/m\-1\}\\mathbb\{E\}\\left\[A\_\{\\alpha\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\-A\_\{\\alpha\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}\)\\right\]=α1/m−1𝔼⟨∇φ\(Yα\),QgXα\(ξα\)⟩\.\\displaystyle=\\alpha^\{1/m\-1\}\\mathbb\{E\}\\left\\langle\\nabla\\varphi\(Y\_\{\\alpha\}\),Qg\_\{X\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}\)\\right\\rangle\.Here the solution of the Poisson equation is evaluated atXα=α1/mYαX\_\{\\alpha\}=\\alpha^\{1/m\}Y\_\{\\alpha\}withYα∈K0Y\_\{\\alpha\}\\in K\_\{0\}\. The last display follows fromχx−Qχx=Qgx\\chi\_\{x\}\-Q\\chi\_\{x\}=Qg\_\{x\}and cancels the conditional mean of the singular term in \([B\.14](https://arxiv.org/html/2607.16384#A2.E14)\), since𝔼\[Gα∣Yα,ξα\]=QgXα\(ξα\)\\mathbb\{E\}\[G\_\{\\alpha\}\\mid Y\_\{\\alpha\},\\xi\_\{\\alpha\}\]=Qg\_\{X\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}\)\.
It remains to control
−α1/m−1𝔼\[Aα\(Yα\+,ξα\+\)−Aα\(Yα,ξα\+\)\]\.\-\\alpha^\{1/m\-1\}\\mathbb\{E\}\\left\[A\_\{\\alpha\}\(Y\_\{\\alpha\}^\{\+\},\\xi\_\{\\alpha\}^\{\+\}\)\-A\_\{\\alpha\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\]\.LetEα:=\{‖Yα\+−Yα‖≤1/2\}E\_\{\\alpha\}:=\\\{\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert\\leq 1/2\\\}\. By \([B\.7](https://arxiv.org/html/2607.16384#A2.E7)\),
ℙ\(Eαc\)≤4𝔼\[‖Yα\+−Yα‖2𝟏Eαc\]=o\(α2−2/m\)\.\\mathbb\{P\}\(E\_\{\\alpha\}^\{c\}\)\\leq 4\\mathbb\{E\}\\left\[\\left\\lVert Y\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}\\right\\rVert^\{2\}\\mathbf\{1\}\_\{E\_\{\\alpha\}^\{c\}\}\\right\]=o\(\\alpha^\{2\-2/m\}\)\.The bound onAαA\_\{\\alpha\}, Cauchy–Schwarz, andξα\+∼πΞ\\xi\_\{\\alpha\}^\{\+\}\\sim\\pi\_\{\\Xi\}therefore give
α1/m−1𝔼\[\|Aα\(Yα\+,ξα\+\)−Aα\(Yα,ξα\+\)\|𝟏Eαc\]\\displaystyle\\alpha^\{1/m\-1\}\\mathbb\{E\}\\left\[\\left\|A\_\{\\alpha\}\(Y\_\{\\alpha\}^\{\+\},\\xi\_\{\\alpha\}^\{\+\}\)\-A\_\{\\alpha\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)\\right\|\\mathbf\{1\}\_\{E\_\{\\alpha\}^\{c\}\}\\right\]\(B\.15\)≤Cα1/m−1\(𝔼πΞRχ2\)1/2ℙ\(Eαc\)1/2=o\(1\)\.\\displaystyle\\qquad\\leq C\\alpha^\{1/m\-1\}\\bigl\(\\mathbb\{E\}\_\{\\pi\_\{\\Xi\}\}R\_\{\\chi\}^\{2\}\\bigr\)^\{1/2\}\\mathbb\{P\}\(E\_\{\\alpha\}^\{c\}\)^\{1/2\}=o\(1\)\.
OnEαE\_\{\\alpha\}, if either term involvingAαA\_\{\\alpha\}is nonzero, then one ofYα,Yα\+Y\_\{\\alpha\},Y\_\{\\alpha\}^\{\+\}belongs toK0K\_\{0\}, and both belong toK1K\_\{1\}\. Hence
Aα\(Yα\+,ξα\+\)−Aα\(Yα,ξα\+\)\\displaystyle A\_\{\\alpha\}\(Y\_\{\\alpha\}^\{\+\},\\xi\_\{\\alpha\}^\{\+\}\)\-A\_\{\\alpha\}\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}^\{\+\}\)=⟨∇φ\(Yα\+\)−∇φ\(Yα\),χ~α,Yα\(ξα\+\)⟩\+⟨∇φ\(Yα\+\),χ~α,Yα\+\(ξα\+\)−χ~α,Yα\(ξα\+\)⟩\.\\displaystyle\\quad=\\left\\langle\\nabla\\varphi\(Y\_\{\\alpha\}^\{\+\}\)\-\\nabla\\varphi\(Y\_\{\\alpha\}\),\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle\+\\left\\langle\\nabla\\varphi\(Y\_\{\\alpha\}^\{\+\}\),\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}^\{\+\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)\-\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle\.
For the first term onEαE\_\{\\alpha\}, Taylor’s formula gives
−α1/m−1𝔼⟨∇φ\(Yα\+\)−∇φ\(Yα\),χ~α,Yα\(ξα\+\)⟩𝟏Eα\\displaystyle\-\\alpha^\{1/m\-1\}\\mathbb\{E\}\\left\\langle\\nabla\\varphi\(Y\_\{\\alpha\}^\{\+\}\)\-\\nabla\\varphi\(Y\_\{\\alpha\}\),\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle\\mathbf\{1\}\_\{E\_\{\\alpha\}\}=𝔼tr\(Gαχ~α,Yα\(ξα\+\)⊤∇2φ\(Yα\)\)\+o\(1\)\.\\displaystyle\\qquad=\\mathbb\{E\}\\operatorname\{tr\}\\\!\\left\(G\_\{\\alpha\}\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)^\{\\top\}\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\}\)\\right\)\+o\(1\)\.The contribution ofEαcE\_\{\\alpha\}^\{c\}to the displayed leading term tends to zero\. Indeed,
𝔼\[‖Gα‖2𝟏Eαc\]≤2𝔼\[‖𝖦α‖2𝟏Eαc\]\+2𝔼‖h\(Xα\)‖2⟶0\\mathbb\{E\}\[\\left\\lVert G\_\{\\alpha\}\\right\\rVert^\{2\}\\mathbf\{1\}\_\{E\_\{\\alpha\}^\{c\}\}\]\\leq 2\\mathbb\{E\}\[\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert^\{2\}\\mathbf\{1\}\_\{E\_\{\\alpha\}^\{c\}\}\]\+2\\mathbb\{E\}\\left\\lVert h\(X\_\{\\alpha\}\)\\right\\rVert^\{2\}\\longrightarrow 0by \([B\.7](https://arxiv.org/html/2607.16384#A2.E7)\) and \([B\.9](https://arxiv.org/html/2607.16384#A2.E9)\); Cauchy–Schwarz withRχ∈L2\(πΞ\)R\_\{\\chi\}\\in L^\{2\}\(\\pi\_\{\\Xi\}\)then applies\. Also,∇2φ\(Yα\)≠0\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\}\)\\neq 0impliesYα∈K0Y\_\{\\alpha\}\\in K\_\{0\}, so the solution appearing in the leading term is well defined\.
The contribution of−α2−2/mhα\-\\alpha^\{2\-2/m\}h\_\{\\alpha\}isO\(α1−1/m\)O\(\\alpha^\{1\-1/m\}\): the factor∇2φ\(Yα\)\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\}\)restricts this term to a fixed compact set, wherehαh\_\{\\alpha\}is locally uniformly bounded by the expansion in Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4), while𝔼Rχ\(ξα\+\)<∞\\mathbb\{E\}R\_\{\\chi\}\(\\xi\_\{\\alpha\}^\{\+\}\)<\\infty\.
To control the Taylor remainder, letωφ\\omega\_\{\\varphi\}be a bounded modulus of continuity of∇2φ\\nabla^\{2\}\\varphi\. SinceYα\+−Yα=−α1−1/m𝖦αY\_\{\\alpha\}^\{\+\}\-Y\_\{\\alpha\}=\-\\alpha^\{1\-1/m\}\\mathsf\{G\}\_\{\\alpha\}, the absolute expectation of the remainder after division byα1−1/m\\alpha^\{1\-1/m\}is at most
C𝔼\[ωφ\(α1−1/m‖𝖦α‖\)‖𝖦α‖Rχ\(ξα\+\)\]\.C\\mathbb\{E\}\\\!\\left\[\\omega\_\{\\varphi\}\(\\alpha^\{1\-1/m\}\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert\)\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert R\_\{\\chi\}\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\]\.For everyη\>0\\eta\>0, the contribution of\{α1−1/m‖𝖦α‖≤η\}\\\{\\alpha^\{1\-1/m\}\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert\\leq\\eta\\\}is bounded by
Cωφ\(η\)\(𝔼‖𝖦α‖2\)1/2\(𝔼πΞRχ2\)1/2≤Cωφ\(η\)\.C\\omega\_\{\\varphi\}\(\\eta\)\\bigl\(\\mathbb\{E\}\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert^\{2\}\\bigr\)^\{1/2\}\\bigl\(\\mathbb\{E\}\_\{\\pi\_\{\\Xi\}\}R\_\{\\chi\}^\{2\}\\bigr\)^\{1/2\}\\leq C\\omega\_\{\\varphi\}\(\\eta\)\.On the complementary event, Cauchy–Schwarz bounds the contribution by
C\(𝔼\[‖𝖦α‖2𝟏\{α1−1/m‖𝖦α‖\>η\}\]\)1/2\(𝔼πΞRχ2\)1/2,C\\left\(\\mathbb\{E\}\\\!\\left\[\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert^\{2\}\\mathbf\{1\}\_\{\\\{\\alpha^\{1\-1/m\}\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert\>\\eta\\\}\}\\right\]\\right\)^\{1/2\}\\bigl\(\\mathbb\{E\}\_\{\\pi\_\{\\Xi\}\}R\_\{\\chi\}^\{2\}\\bigr\)^\{1/2\},which tends to zero by \([B\.7](https://arxiv.org/html/2607.16384#A2.E7)\)\. Sendingη↓0\\eta\\downarrow 0proves theo\(1\)o\(1\)remainder in the gradient increment\.
For the second term onEαE\_\{\\alpha\}, which changes the solution of the Poisson equation fromα1/mYα\\alpha^\{1/m\}Y\_\{\\alpha\}toα1/mYα\+\\alpha^\{1/m\}Y\_\{\\alpha\}^\{\+\}, fix the exponentp∈\(1−1/m,1\)p\\in\(1\-1/m,1\)in \([B\.2](https://arxiv.org/html/2607.16384#A2.E2)\)\. That estimate gives
α1/m−1𝔼\|⟨∇φ\(Yα\+\),χ~α,Yα\+\(ξα\+\)−χ~α,Yα\(ξα\+\)⟩\|𝟏Eα\\displaystyle\\alpha^\{1/m\-1\}\\mathbb\{E\}\\left\|\\left\\langle\\nabla\\varphi\(Y\_\{\\alpha\}^\{\+\}\),\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}^\{\+\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)\-\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)\\right\\rangle\\right\|\\mathbf\{1\}\_\{E\_\{\\alpha\}\}≤Cpαp−\(1−1/m\)𝔼\[‖𝖦α‖pRχ\(ξα\+\)1−p\]\\displaystyle\\qquad\\leq C\_\{p\}\\alpha^\{p\-\(1\-1/m\)\}\\mathbb\{E\}\[\\left\\lVert\\mathsf\{G\}\_\{\\alpha\}\\right\\rVert^\{p\}R\_\{\\chi\}\(\\xi\_\{\\alpha\}^\{\+\}\)^\{1\-p\}\]≤Cpαp−\(1−1/m\)=o\(1\)\.\\displaystyle\\qquad\\leq C\_\{p\}\\alpha^\{p\-\(1\-1/m\)\}=o\(1\)\.Here the second inequality follows from weighted AM–GM,apb1−p≤pa\+\(1−p\)ba^\{p\}b^\{1\-p\}\\leq pa\+\(1\-p\)b, the uniform second moment of𝖦α\\mathsf\{G\}\_\{\\alpha\}, andRχ∈L2\(πΞ\)R\_\{\\chi\}\\in L^\{2\}\(\\pi\_\{\\Xi\}\)\.
Combining the preceding identities with \([B\.13](https://arxiv.org/html/2607.16384#A2.E13)\) gives
0\\displaystyle 0=−𝔼⟨hα\(Yα\),∇φ\(Yα\)⟩\\displaystyle=\-\\mathbb\{E\}\\left\\langle h\_\{\\alpha\}\(Y\_\{\\alpha\}\),\\nabla\\varphi\(Y\_\{\\alpha\}\)\\right\\rangle\+12𝔼tr\(\[GαGα⊤\+Gαχ~α,Yα\(ξα\+\)⊤\+χ~α,Yα\(ξα\+\)Gα⊤\]∇2φ\(Yα\)\)\+o\(1\)\.\\displaystyle\\quad\+\\frac\{1\}\{2\}\\mathbb\{E\}\\operatorname\{tr\}\\\!\\left\(\\Bigl\[G\_\{\\alpha\}G\_\{\\alpha\}^\{\\top\}\+G\_\{\\alpha\}\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)^\{\\top\}\+\\widetilde\{\\chi\}\_\{\\alpha,Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}^\{\+\}\)G\_\{\\alpha\}^\{\\top\}\\Bigr\]\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\}\)\\right\)\+o\(1\)\.After multiplication by∇2φ\(Yα\)\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\}\), conditioning on\(Yα,ξα\)\(Y\_\{\\alpha\},\\xi\_\{\\alpha\}\)identifies the conditional mean withQMα1/mYα\(ξα\)∇2φ\(Yα\)QM\_\{\\alpha^\{1/m\}Y\_\{\\alpha\}\}\(\\xi\_\{\\alpha\}\)\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\}\)\. This proves \([B\.11](https://arxiv.org/html/2607.16384#A2.E11)\)\. ∎
###### Proof of Lemma[4\.9](https://arxiv.org/html/2607.16384#S4.Thmtheorem9)\.
Letφ∈Cc3\(ℝd\)\\varphi\\in C\_\{c\}^\{3\}\(\\mathbb\{R\}^\{d\}\)and letK:=supp\(∇2φ\)K:=\\operatorname\{supp\}\(\\nabla^\{2\}\\varphi\)\. Lemma[B\.9](https://arxiv.org/html/2607.16384#A2.Thmtheorem9)gives \([B\.11](https://arxiv.org/html/2607.16384#A2.E11)\)\. Along the subsequenceαk↓0\\alpha\_\{k\}\\downarrow 0, local uniform convergence ofhαh\_\{\\alpha\}and weak convergence give
𝔼⟨hαk\(Yαk\),∇φ\(Yαk\)⟩⟶∫ℝd⟨h0\(y\),∇φ\(y\)⟩ν\(dy\)\.\\mathbb\{E\}\\left\\langle h\_\{\\alpha\_\{k\}\}\(Y\_\{\\alpha\_\{k\}\}\),\\nabla\\varphi\(Y\_\{\\alpha\_\{k\}\}\)\\right\\rangle\\longrightarrow\\int\_\{\\mathbb\{R\}^\{d\}\}\\left\\langle h\_\{0\}\(y\),\\nabla\\varphi\(y\)\\right\\rangle\\,\\nu\(dy\)\.\(B\.16\)Here∇φ\\nabla\\varphiis compactly supported andhα→h0h\_\{\\alpha\}\\to h\_\{0\}locally uniformly by Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4)\.
For the covariance term, \([B\.10](https://arxiv.org/html/2607.16384#A2.E10)\) replacesQMαk1/mYαk\(ξαk\)QM\_\{\\alpha\_\{k\}^\{1/m\}Y\_\{\\alpha\_\{k\}\}\}\(\\xi\_\{\\alpha\_\{k\}\}\)byQM0\(ξαk\)QM\_\{0\}\(\\xi\_\{\\alpha\_\{k\}\}\)onKK\. Each entry ofQM0QM\_\{0\}belongs toL1\(πΞ\)L^\{1\}\(\\pi\_\{\\Xi\}\)by Lemma[B\.7](https://arxiv.org/html/2607.16384#A2.Thmtheorem7)\. Applying Lemma[B\.8](https://arxiv.org/html/2607.16384#A2.Thmtheorem8)entrywise, with the bounded Lipschitz entries of∇2φ\\nabla^\{2\}\\varphi, yields
𝔼tr\(QM0\(ξαk\)∇2φ\(Yαk\)\)⟶∫ℝdtr\(M¯∇2φ\(y\)\)ν\(dy\),\\mathbb\{E\}\\operatorname\{tr\}\\\!\\left\(QM\_\{0\}\(\\xi\_\{\\alpha\_\{k\}\}\)\\nabla^\{2\}\\varphi\(Y\_\{\\alpha\_\{k\}\}\)\\right\)\\longrightarrow\\int\_\{\\mathbb\{R\}^\{d\}\}\\operatorname\{tr\}\\\!\\left\(\\bar\{M\}\\nabla^\{2\}\\varphi\(y\)\\right\)\\nu\(dy\),\(B\.17\)whereM¯:=∫ΞM0\(ξ\)πΞ\(dξ\)\.\\bar\{M\}:=\\int\_\{\\Xi\}M\_\{0\}\(\\xi\)\\,\\pi\_\{\\Xi\}\(d\\xi\)\.It remains only to identifyM¯\\bar\{M\}withΣ\\Sigma\. Let\(ξ0,ξ1\)\(\\xi\_\{0\},\\xi\_\{1\}\)be two successive values of the stationary driving chain and writef=g0f=g\_\{0\},χ=χ0\\chi=\\chi\_\{0\}\. Sinceχ−Qχ=Qf\\chi\-Q\\chi=Qf,𝔼\[f\(ξ1\)\+χ\(ξ1\)∣ξ0\]=Qf\(ξ0\)\+Qχ\(ξ0\)=χ\(ξ0\)\.\\mathbb\{E\}\[f\(\\xi\_\{1\}\)\+\\chi\(\\xi\_\{1\}\)\\mid\\xi\_\{0\}\]=Qf\(\\xi\_\{0\}\)\+Q\\chi\(\\xi\_\{0\}\)=\\chi\(\\xi\_\{0\}\)\.WithD0=f\(ξ1\)\+χ\(ξ1\)−χ\(ξ0\)D\_\{0\}=f\(\\xi\_\{1\}\)\+\\chi\(\\xi\_\{1\}\)\-\\chi\(\\xi\_\{0\}\), expansion and the preceding conditional identity give
Σ\\displaystyle\\Sigma=𝔼\[D0D0⊤\]\\displaystyle=\\mathbb\{E\}\[D\_\{0\}D\_\{0\}^\{\\top\}\]=𝔼\[f\(ξ1\)f\(ξ1\)⊤\+f\(ξ1\)χ\(ξ1\)⊤\+χ\(ξ1\)f\(ξ1\)⊤\]\\displaystyle=\\mathbb\{E\}\\bigl\[f\(\\xi\_\{1\}\)f\(\\xi\_\{1\}\)^\{\\top\}\+f\(\\xi\_\{1\}\)\\chi\(\\xi\_\{1\}\)^\{\\top\}\+\\chi\(\\xi\_\{1\}\)f\(\\xi\_\{1\}\)^\{\\top\}\\bigr\]=∫ΞM0\(ξ\)πΞ\(dξ\)=M¯\.\\displaystyle=\\int\_\{\\Xi\}M\_\{0\}\(\\xi\)\\,\\pi\_\{\\Xi\}\(d\\xi\)=\\bar\{M\}\.Combining \([B\.11](https://arxiv.org/html/2607.16384#A2.E11)\), \([B\.16](https://arxiv.org/html/2607.16384#A2.E16)\), and \([B\.17](https://arxiv.org/html/2607.16384#A2.E17)\) proves \([4\.19](https://arxiv.org/html/2607.16384#S4.E19)\)\. ∎
### B\.7Proof of Lemma[4\.10](https://arxiv.org/html/2607.16384#S4.Thmtheorem10)
###### Lemma B\.10\.
Under Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)\(specifically,[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)and[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4)\),
⟨y,h0\(y\)⟩≥cinm−1‖y‖m,y∈ℝd\.\\left\\langle y,h\_\{0\}\(y\)\\right\\rangle\\geq\\frac\{c\_\{\\rm in\}\}\{m\-1\}\\left\\lVert y\\right\\rVert^\{m\},\\qquad y\\in\\mathbb\{R\}^\{d\}\.
###### Proof\.
The claim is trivial aty=0y=0\. Fory≠0y\\neq 0, setx=α1/myx=\\alpha^\{1/m\}y\. For all sufficiently smallα\\alpha,‖x‖≤RH2\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\}\. The identityh\(x\)=∫01∇2H\(tx\)x𝑑th\(x\)=\\int\_\{0\}^\{1\}\\nabla^\{2\}H\(tx\)x\\,dtand Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)give⟨x,h\(x\)⟩≥cinm−1‖x‖m\.\\left\\langle x,h\(x\)\\right\\rangle\\geq\\frac\{c\_\{\\rm in\}\}\{m\-1\}\\left\\lVert x\\right\\rVert^\{m\}\.Equivalently,
⟨y,α−\(m−1\)/mh\(α1/my\)⟩=α−1⟨x,h\(x\)⟩≥cinm−1‖y‖m\.\\left\\langle y,\\alpha^\{\-\(m\-1\)/m\}h\(\\alpha^\{1/m\}y\)\\right\\rangle=\\alpha^\{\-1\}\\left\\langle x,h\(x\)\\right\\rangle\\geq\\frac\{c\_\{\\rm in\}\}\{m\-1\}\\left\\lVert y\\right\\rVert^\{m\}\.Lettingα↓0\\alpha\\downarrow 0and using the expansion in Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4)together with the homogeneity ofh0h\_\{0\}gives the result\. ∎
###### Lemma B\.11\.
Under Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)\(specifically,[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)and[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4)\), there existscm\>0c\_\{m\}\>0such that, for ally,z∈ℝdy,z\\in\\mathbb\{R\}^\{d\},⟨y−z,h0\(y\)−h0\(z\)⟩≥cm‖y−z‖m\.\\left\\langle y\-z,h\_\{0\}\(y\)\-h\_\{0\}\(z\)\\right\\rangle\\geq c\_\{m\}\\left\\lVert y\-z\\right\\rVert^\{m\}\.
###### Proof\.
The claim is trivial wheny=zy=z\. Sete:=y−z≠0e:=y\-z\\neq 0\. For sufficiently smallα\\alpha, the segment betweenα1/mz\\alpha^\{1/m\}zandα1/my\\alpha^\{1/m\}yis contained in\{‖x‖≤RH2\}\\\{\\left\\lVert x\\right\\rVert\\leq\\frac\{R\_\{H\}\}\{2\}\\\}\. It can pass through the origin for at most one parameter value, which does not affect the integral below\. Using Assumption[2\.3](https://arxiv.org/html/2607.16384#S2.Thmtheorem3)[\(H2\)](https://arxiv.org/html/2607.16384#S2.I1.i2)along the segment gives
⟨e,α−\(m−1\)/m\{h\(α1/my\)−h\(α1/mz\)\}⟩\\displaystyle\\left\\langle e,\\alpha^\{\-\(m\-1\)/m\}\\\{h\(\\alpha^\{1/m\}y\)\-h\(\\alpha^\{1/m\}z\)\\\}\\right\\rangle=α−1⟨α1/me,h\(α1/my\)−h\(α1/mz\)⟩\\displaystyle\\qquad=\\alpha^\{\-1\}\\left\\langle\\alpha^\{1/m\}e,h\(\\alpha^\{1/m\}y\)\-h\(\\alpha^\{1/m\}z\)\\right\\rangle=α−1∫01⟨α1/me,∇2H\(α1/m\(z\+te\)\)α1/me⟩𝑑t\\displaystyle\\qquad=\\alpha^\{\-1\}\\int\_\{0\}^\{1\}\\left\\langle\\alpha^\{1/m\}e,\\nabla^\{2\}H\(\\alpha^\{1/m\}\(z\+te\)\)\\alpha^\{1/m\}e\\right\\rangle\\,dt≥cin‖e‖2∫01‖z\+te‖m−2𝑑t\.\\displaystyle\\qquad\\geq c\_\{\\rm in\}\\left\\lVert e\\right\\rVert^\{2\}\\int\_\{0\}^\{1\}\\left\\lVert z\+te\\right\\rVert^\{m\-2\}\\,dt\.We use the elementary segment bound
inf‖v‖=1infw∈ℝd∫01‖w\+tv‖m−2𝑑t\>0\.\\inf\_\{\\left\\lVert v\\right\\rVert=1\}\\inf\_\{w\\in\\mathbb\{R\}^\{d\}\}\\int\_\{0\}^\{1\}\\left\\lVert w\+tv\\right\\rVert^\{m\-2\}\\,dt\>0\.\(B\.18\)To prove \([B\.18](https://arxiv.org/html/2607.16384#A2.E18)\), the casem=2m=2is immediate\. Form\>2m\>2, decomposew=av\+w⟂w=av\+w\_\{\\perp\}, withw⟂⟂vw\_\{\\perp\}\\perp v\. Then
∫01‖w\+tv‖m−2𝑑t≥∫01\|a\+t\|m−2𝑑t\.\\int\_\{0\}^\{1\}\\left\\lVert w\+tv\\right\\rVert^\{m\-2\}\\,dt\\geq\\int\_\{0\}^\{1\}\|a\+t\|^\{m\-2\}\\,dt\.The last integral is minimized overa∈ℝa\\in\\mathbb\{R\}when the interval\[a,a\+1\]\[a,a\+1\]is centered at the origin, and its minimum is2∫01/2tm−2𝑑t\>02\\int\_\{0\}^\{1/2\}t^\{m\-2\}\\,dt\>0\. Thus \([B\.18](https://arxiv.org/html/2607.16384#A2.E18)\) holds and gives∫01‖z\+te‖m−2𝑑t≥c‖e‖m−2\.\\int\_\{0\}^\{1\}\\left\\lVert z\+te\\right\\rVert^\{m\-2\}\\,dt\\geq c\\left\\lVert e\\right\\rVert^\{m\-2\}\.Lettingα↓0\\alpha\\downarrow 0and using the expansion in Assumption[2\.10](https://arxiv.org/html/2607.16384#S2.Thmtheorem10)[\(H4\)](https://arxiv.org/html/2607.16384#S2.I5.i4)together with the homogeneity ofh0h\_\{0\}gives the result\. ∎
###### Proof of Lemma[4\.10](https://arxiv.org/html/2607.16384#S4.Thmtheorem10)\.
The drift−h0\-h\_\{0\}is locally Lipschitz becauseH0∈C2H\_\{0\}\\in C^\{2\}, and it has polynomial growth by homogeneity\. Thus the SDE has a pathwise unique maximal strong solution up to its explosion time\. Let
W\(y\):=1\+‖y‖2,ℒφ\(y\):=−⟨h0\(y\),∇φ\(y\)⟩\+12tr\(Σ∇2φ\(y\)\)\.W\(y\):=1\+\\left\\lVert y\\right\\rVert^\{2\},\\qquad\\mathcal\{L\}\\varphi\(y\):=\-\\left\\langle h\_\{0\}\(y\),\\nabla\\varphi\(y\)\\right\\rangle\+\\frac\{1\}\{2\}\\operatorname\{tr\}\(\\Sigma\\nabla^\{2\}\\varphi\(y\)\)\.By Lemma[B\.10](https://arxiv.org/html/2607.16384#A2.Thmtheorem10),
ℒW\(y\)=−2⟨y,h0\(y\)⟩\+tr\(Σ\)≤C−c‖y‖m\.\\mathcal\{L\}W\(y\)=\-2\\left\\langle y,h\_\{0\}\(y\)\\right\\rangle\+\\operatorname\{tr\}\(\\Sigma\)\\leq C\-c\\left\\lVert y\\right\\rVert^\{m\}\.\(B\.19\)IfτR:=inf\{t:‖Yt‖≥R\}\\tau\_\{R\}:=\\inf\\\{t:\\left\\lVert Y\_\{t\}\\right\\rVert\\geq R\\\}, Itô’s formula applied toW\(Yt∧τR\)W\(Y\_\{t\\wedge\\tau\_\{R\}\}\)gives𝔼yW\(Yt∧τR\)≤W\(y\)\+Ct\.\\mathbb\{E\}\_\{y\}W\(Y\_\{t\\wedge\\tau\_\{R\}\}\)\\leq W\(y\)\+Ct\.On\{τR≤t\}\\\{\\tau\_\{R\}\\leq t\\\}, the left\-hand side is at least1\+R21\+R^\{2\}\. Henceℙy\(τR≤t\)≤\(W\(y\)\+Ct\)/\(1\+R2\)→0\\mathbb\{P\}\_\{y\}\(\\tau\_\{R\}\\leq t\)\\leq\(W\(y\)\+Ct\)/\(1\+R^\{2\}\)\\to 0, which proves nonexplosion\. Applying Itô’s formula without stopping and using \([B\.19](https://arxiv.org/html/2607.16384#A2.E19)\) gives1T∫0T𝔼y‖Yt‖m𝑑t≤C\+W\(y\)cT\.\\frac\{1\}\{T\}\\int\_\{0\}^\{T\}\\mathbb\{E\}\_\{y\}\\left\\lVert Y\_\{t\}\\right\\rVert^\{m\}\\,dt\\leq C\+\\frac\{W\(y\)\}\{cT\}\.The time\-averaged lawsT−1∫0TPt\(y,⋅\)𝑑tT^\{\-1\}\\int\_\{0\}^\{T\}P\_\{t\}\(y,\\cdot\)\\,dtare therefore tight\. Since the nonexplosive locally Lipschitz SDE is Feller, the Krylov–Bogoliubov argument gives at least one invariant probability law\.
We next show that every invariant probability law has finitemm\-moment\. Letθn∈C2\(\[0,∞\)\)\\theta\_\{n\}\\in C^\{2\}\(\[0,\\infty\)\)be nondecreasing and concave, with0≤θn′≤10\\leq\\theta\_\{n\}^\{\\prime\}\\leq 1,θn′\(s\)=1\\theta\_\{n\}^\{\\prime\}\(s\)=1fors≤ns\\leq n,θn′\(s\)=0\\theta\_\{n\}^\{\\prime\}\(s\)=0fors≥2ns\\geq 2n,θn′′≤0\\theta\_\{n\}^\{\\prime\\prime\}\\leq 0, andθn′\(s\)↑1\\theta\_\{n\}^\{\\prime\}\(s\)\\uparrow 1for every fixedss\. SetWn\(y\):=θn\(W\(y\)\)W\_\{n\}\(y\):=\\theta\_\{n\}\(W\(y\)\)\. ThenWnW\_\{n\}is bounded with bounded first and second derivatives\. Ifμ\\muis invariant, then∫ℒWn𝑑μ=0\\int\\mathcal\{L\}W\_\{n\}\\,d\\mu=0\. Moreover,
ℒWn\(y\)=θn′\(W\(y\)\)ℒW\(y\)\+12θn′′\(W\(y\)\)‖Σ1/2∇W\(y\)‖2≤θn′\(W\(y\)\)\(C−c‖y‖m\),\\mathcal\{L\}W\_\{n\}\(y\)=\\theta\_\{n\}^\{\\prime\}\(W\(y\)\)\\mathcal\{L\}W\(y\)\+\\frac\{1\}\{2\}\\theta\_\{n\}^\{\\prime\\prime\}\(W\(y\)\)\\left\\lVert\\Sigma^\{1/2\}\\nabla W\(y\)\\right\\rVert^\{2\}\\leq\\theta\_\{n\}^\{\\prime\}\(W\(y\)\)\(C\-c\\left\\lVert y\\right\\rVert^\{m\}\),becauseθn′′≤0\\theta\_\{n\}^\{\\prime\\prime\}\\leq 0andΣ⪰0\\Sigma\\succeq 0\. Hence
c∫θn′\(W\(y\)\)‖y‖mμ\(dy\)≤C∫θn′\(W\(y\)\)μ\(dy\)≤C\.c\\int\\theta\_\{n\}^\{\\prime\}\(W\(y\)\)\\left\\lVert y\\right\\rVert^\{m\}\\,\\mu\(dy\)\\leq C\\int\\theta\_\{n\}^\{\\prime\}\(W\(y\)\)\\,\\mu\(dy\)\\leq C\.Lettingn→∞n\\to\\inftyand using monotone convergence gives∫‖y‖mμ\(dy\)<∞\.\\int\\left\\lVert y\\right\\rVert^\{m\}\\,\\mu\(dy\)<\\infty\.In particular every invariant law has finite second moment\.
Letμ\\muandν\\nube invariant laws\. Choose a coupling\(Y0,Y~0\)\(Y\_\{0\},\\widetilde\{Y\}\_\{0\}\)with finite second moment and drive both solutions by the same Brownian motion\. Since the diffusion coefficient is constant,Δt:=Yt−Y~t\\Delta\_\{t\}:=Y\_\{t\}\-\\widetilde\{Y\}\_\{t\}satisfies, up to localization,
ddt‖Δt‖2=−2⟨Δt,h0\(Yt\)−h0\(Y~t\)⟩≤−2cm‖Δt‖m\\frac\{d\}\{dt\}\\left\\lVert\\Delta\_\{t\}\\right\\rVert^\{2\}=\-2\\left\\langle\\Delta\_\{t\},h\_\{0\}\(Y\_\{t\}\)\-h\_\{0\}\(\\widetilde\{Y\}\_\{t\}\)\\right\\rangle\\leq\-2c\_\{m\}\\left\\lVert\\Delta\_\{t\}\\right\\rVert^\{m\}by Lemma[B\.11](https://arxiv.org/html/2607.16384#A2.Thmtheorem11)\. WithD\(t\):=𝔼‖Δt‖2D\(t\):=\\mathbb\{E\}\\left\\lVert\\Delta\_\{t\}\\right\\rVert^\{2\}, Fatou’s lemma and Jensen’s inequality giveD′\(t\)≤−2cm𝔼‖Δt‖m≤−2cmD\(t\)m/2\.D^\{\\prime\}\(t\)\\leq\-2c\_\{m\}\\mathbb\{E\}\\left\\lVert\\Delta\_\{t\}\\right\\rVert^\{m\}\\leq\-2c\_\{m\}D\(t\)^\{m/2\}\.ThusD\(t\)→0D\(t\)\\to 0, exponentially ifm=2m=2and polynomially ifm\>2m\>2\. The marginals remainμ\\muandν\\nu, soW2\(μ,ν\)2≤D\(t\)→0W\_\{2\}\(\\mu,\\nu\)^\{2\}\\leq D\(t\)\\to 0\. Henceμ=ν\\mu=\\nu\. Denote the unique invariant law byν∞\\nu\_\{\\infty\}\.
Finally, supposeρ\\rhosatisfies the stationary weak equation for everyφ∈Cc3\(ℝd\)\\varphi\\in C\_\{c\}^\{3\}\(\\mathbb\{R\}^\{d\}\)\. It then holds onCc∞\(ℝd\)C\_\{c\}^\{\\infty\}\(\\mathbb\{R\}^\{d\}\)\. The martingale problem forℒ\\mathcal\{L\}is well posed because the SDE is pathwise unique and nonexplosive\. The Echeverria criterion therefore implies thatρ\\rhois invariant; see\[[15](https://arxiv.org/html/2607.16384#bib.bib15), Ch\. 8\]\. By uniqueness,ρ=ν∞\\rho=\\nu\_\{\\infty\}\. ∎
## Appendix CSeparable extensions
This appendix proves the separable results from Section[2\.3](https://arxiv.org/html/2607.16384#S2.SS3)\. We sum the one\-dimensional contractions for the constant\-stepsize result and apply the generator argument to the coordinates that remain nonzero under the common normalization\.
### C\.1Proof of Corollary[2\.15](https://arxiv.org/html/2607.16384#S2.Thmtheorem15)
###### Proof of Corollary[2\.15](https://arxiv.org/html/2607.16384#S2.Thmtheorem15)\.
Writedi,α:=dVα,i,α\(i\)d\_\{i,\\alpha\}:=d\_\{V\_\{\\alpha,i\},\\alpha\}^\{\(i\)\}\. For each coordinateii, the one\-dimensional proof gives a constantci\>0c\_\{i\}\>0and, after taking the minimum of the finitely many stepsize thresholds, the synchronous one\-step estimate
𝔼\[di,α\(FU1\(i\)\(xi,ξ\),FU1\(i\)\(yi,η\)\)\]≤\(1−ciαmi−1\)di,α\(\(xi,ξ\),\(yi,η\)\)\\mathbb\{E\}\\Bigl\[d\_\{i,\\alpha\}\\bigl\(F\_\{U\_\{1\}\}^\{\(i\)\}\(x\_\{i\},\\xi\),F\_\{U\_\{1\}\}^\{\(i\)\}\(y\_\{i\},\\eta\)\\bigr\)\\Bigr\]\\leq\(1\-c\_\{i\}\\alpha^\{m\_\{i\}\-1\}\)\\,d\_\{i,\\alpha\}\\bigl\(\(x\_\{i\},\\xi\),\(y\_\{i\},\\eta\)\\bigr\)for the coordinate map
Fu\(i\)\(xi,ξ\):=\(xi−α\(hi\(xi\)\+gi\(xi,Φ\(ξ,u\)\)\),Φ\(ξ,u\)\)\.F\_\{u\}^\{\(i\)\}\(x\_\{i\},\\xi\):=\\Bigl\(x\_\{i\}\-\\alpha\\bigl\(h\_\{i\}\(x\_\{i\}\)\+g\_\{i\}\(x\_\{i\},\\Phi\(\\xi,u\)\)\\bigr\),\\Phi\(\\xi,u\)\\Bigr\)\.Theiith coordinate projection of the full separable map is\(x,ξ\)↦Fu\(i\)\(xi,ξ\)\.\(x,\\xi\)\\mapsto F\_\{u\}^\{\(i\)\}\(x\_\{i\},\\xi\)\.Since the same innovationU1U\_\{1\}drives every coordinate, summing the coordinatewise estimates gives
𝔼\[dsep,α\(FU1\(x,ξ\),FU1\(y,η\)\)\]≤\(1−c0αmmax−1\)dsep,α\(\(x,ξ\),\(y,η\)\),\\mathbb\{E\}\\Bigl\[d\_\{\{\\rm sep\},\\alpha\}\\bigl\(F\_\{U\_\{1\}\}\(x,\\xi\),F\_\{U\_\{1\}\}\(y,\\eta\)\\bigr\)\\Bigr\]\\leq\(1\-c\_\{0\}\\alpha^\{m\_\{\\max\}\-1\}\)\\,d\_\{\{\\rm sep\},\\alpha\}\\bigl\(\(x,\\xi\),\(y,\\eta\)\\bigr\),wherec0:=min1≤i≤dci\>0c\_\{0\}:=\\min\_\{1\\leq i\\leq d\}c\_\{i\}\>0\. Indeed, for0<α≤10<\\alpha\\leq 1,
ciαmi−1≥c0αmmax−1,1≤i≤d\.c\_\{i\}\\alpha^\{m\_\{i\}\-1\}\\geq c\_\{0\}\\alpha^\{m\_\{\\max\}\-1\},\\qquad 1\\leq i\\leq d\.This is the required one\-step contraction in the additive separable metric\.
It remains to check the reference\-point integrability\. The one\-dimensional hypotheses may use different reference points\. Letξ⋆\(i\)\\xi\_\{\\star\}^\{\(i\)\}be one for coordinateii, and fixξ∘∈Ξ\\xi^\{\\circ\}\\in\\Xi\. By Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N1\)](https://arxiv.org/html/2607.16384#S2.I3.i1),
𝔼‖Φ\(ξ∘,U1\)−ξ∘‖≤\(1\+ρΞ\)‖ξ∘−ξ⋆\(i\)‖\+𝔼‖Φ\(ξ⋆\(i\),U1\)−ξ⋆\(i\)‖<∞\.\\mathbb\{E\}\\left\\lVert\\Phi\(\\xi^\{\\circ\},U\_\{1\}\)\-\\xi^\{\\circ\}\\right\\rVert\\leq\(1\+\\rho\_\{\\Xi\}\)\\left\\lVert\\xi^\{\\circ\}\-\\xi\_\{\\star\}^\{\(i\)\}\\right\\rVert\+\\mathbb\{E\}\\left\\lVert\\Phi\(\\xi\_\{\\star\}^\{\(i\)\},U\_\{1\}\)\-\\xi\_\{\\star\}^\{\(i\)\}\\right\\rVert<\\infty\.Ifβi=2\\beta\_\{i\}=2, then Assumption[2\.5](https://arxiv.org/html/2607.16384#S2.Thmtheorem5)[\(N2\)](https://arxiv.org/html/2607.16384#S2.I3.i2)gives
𝔼\|gi\(0,Φ\(ξ∘,U1\)\)\|≤𝔼\|gi\(0,Φ\(ξ⋆\(i\),U1\)\)\|\+Lg,Φ‖ξ∘−ξ⋆\(i\)‖<∞\.\\mathbb\{E\}\\left\\lvert g\_\{i\}\(0,\\Phi\(\\xi^\{\\circ\},U\_\{1\}\)\)\\right\\rvert\\leq\\mathbb\{E\}\\left\\lvert g\_\{i\}\(0,\\Phi\(\\xi\_\{\\star\}^\{\(i\)\},U\_\{1\}\)\)\\right\\rvert\+L\_\{g,\\Phi\}\\left\\lVert\\xi^\{\\circ\}\-\\xi\_\{\\star\}^\{\(i\)\}\\right\\rVert<\\infty\.Thusξ∘\\xi^\{\\circ\}works for every coordinate\. Eachdi,αd\_\{i,\\alpha\}dominates the base norm\|xi−yi\|\+α−1‖ξ−η‖\\left\\lvert x\_\{i\}\-y\_\{i\}\\right\\rvert\+\\alpha^\{\-1\}\\left\\lVert\\xi\-\\eta\\right\\rVertand is locally comparable to it on bounded sets\. Hencedsep,αd\_\{\{\\rm sep\},\\alpha\}is complete and separable onℝd×Ξ\\mathbb\{R\}^\{d\}\\times\\Xiand has the usual Borelσ\\sigma\-field\. Summing the coordinatewise reference\-point costs at\(0,ξ∘\)\(0,\\xi^\{\\circ\}\)and applying Proposition[4\.1](https://arxiv.org/html/2607.16384#S4.Thmtheorem1)gives the invariant law, uniqueness, and \([2\.16](https://arxiv.org/html/2607.16384#S2.E16)\)\. ∎
### C\.2Proof of Corollary[2\.17](https://arxiv.org/html/2607.16384#S2.Thmtheorem17)
###### Proof of Corollary[2\.17](https://arxiv.org/html/2607.16384#S2.Thmtheorem17)\.
Writeq:=\|ℐ∗\|q:=\\lvert\\mathcal\{I\}\_\{\*\}\\rvert, use a subscript∗\*to denote restriction to the coordinates inℐ∗\\mathcal\{I\}\_\{\*\}, and setYα,∗:=α−1/mmaxX∞,∗\(α\),sep\.Y\_\{\\alpha,\*\}:=\\alpha^\{\-1/m\_\{\\max\}\}X\_\{\\infty,\*\}^\{\(\\alpha\),\\rm sep\}\.The coordinatewise moment bounds from the proof of Theorem[2\.13](https://arxiv.org/html/2607.16384#S2.Thmtheorem13)imply that\(Yα,∗\)α\(Y\_\{\\alpha,\*\}\)\_\{\\alpha\}is tight inℝq\\mathbb\{R\}^\{q\}\. We first identify its subsequential limits\. By separability, the update of the active block does not depend on the remaining iterate coordinates\. On the scaleα1/mmax\\alpha^\{1/m\_\{\\max\}\}, its stationary one\-step increment is
Yα,∗\+−Yα,∗=−α2−2/mmaxhα,∗\(Yα,∗\)−α1−1/mmaxg∗\(α1/mmaxYα,∗,ξα\+\),Y\_\{\\alpha,\*\}^\{\+\}\-Y\_\{\\alpha,\*\}=\-\\alpha^\{2\-2/m\_\{\\max\}\}h\_\{\\alpha,\*\}\(Y\_\{\\alpha,\*\}\)\-\\alpha^\{1\-1/m\_\{\\max\}\}g\_\{\*\}\(\\alpha^\{1/m\_\{\\max\}\}Y\_\{\\alpha,\*\},\\xi\_\{\\alpha\}^\{\+\}\),where
hα,∗\(y\):=\(α−\(mmax−1\)/mmaxhi\(α1/mmaxyi\)\)i∈ℐ∗,h\_\{\\alpha,\*\}\(y\):=\\bigl\(\\alpha^\{\-\(m\_\{\\max\}\-1\)/m\_\{\\max\}\}h\_\{i\}\(\\alpha^\{1/m\_\{\\max\}\}y\_\{i\}\)\\bigr\)\_\{i\\in\\mathcal\{I\}\_\{\*\}\},andg∗\(x,ξ\):=\(gi\(xi,ξ\)\)i∈ℐ∗g\_\{\*\}\(x,\\xi\):=\(g\_\{i\}\(x\_\{i\},\\xi\)\)\_\{i\\in\\mathcal\{I\}\_\{\*\}\}\. The coordinatewise tangent assumptions give
hα,∗⟶h0,∗,h0,∗\(y\):=\(hi,0\(yi\)\)i∈ℐ∗h\_\{\\alpha,\*\}\\longrightarrow h\_\{0,\*\},\\qquad h\_\{0,\*\}\(y\):=\(h\_\{i,0\}\(y\_\{i\}\)\)\_\{i\\in\\mathcal\{I\}\_\{\*\}\}locally uniformly\.
Forx∈ℝqx\\in\\mathbb\{R\}^\{q\}, letχ∗,x\(ξ\):=\(χi,xi\(ξ\)\)i∈ℐ∗,\\chi\_\{\*,x\}\(\\xi\):=\(\\chi\_\{i,x\_\{i\}\}\(\\xi\)\)\_\{i\\in\\mathcal\{I\}\_\{\*\}\},whereχi,xi\\chi\_\{i,x\_\{i\}\}is the one\-dimensional solution of the Poisson equation for coordinateii\. At the minimizer, define
Di\(ξ,U1\):=gi\(0,Φ\(ξ,U1\)\)\+χi,0\(Φ\(ξ,U1\)\)−χi,0\(ξ\),i∈ℐ∗\.D\_\{i\}\(\\xi,U\_\{1\}\):=g\_\{i\}\(0,\\Phi\(\\xi,U\_\{1\}\)\)\+\\chi\_\{i,0\}\(\\Phi\(\\xi,U\_\{1\}\)\)\-\\chi\_\{i,0\}\(\\xi\),\\qquad i\\in\\mathcal\{I\}\_\{\*\}\.The vectorD∗=\(Di\)i∈ℐ∗D\_\{\*\}=\(D\_\{i\}\)\_\{i\\in\\mathcal\{I\}\_\{\*\}\}is conditionally centered and, by Lemma[4\.6](https://arxiv.org/html/2607.16384#S4.Thmtheorem6),
Σact:=𝔼\[D∗D∗⊤\]=\(Σsep\)ℐ∗×ℐ∗\.\\Sigma\_\{\\rm act\}:=\\mathbb\{E\}\[D\_\{\*\}D\_\{\*\}^\{\\top\}\]=\(\\Sigma\_\{\\rm sep\}\)\_\{\\mathcal\{I\}\_\{\*\}\\times\\mathcal\{I\}\_\{\*\}\}\.Thus the off\-diagonal asymptotic covariances created by the common driving chain are retained within the active block\.
Forφ∈Cc3\(ℝq\)\\varphi\\in C\_\{c\}^\{3\}\(\\mathbb\{R\}^\{q\}\), use the perturbed test function
φα\(y,ξ\):=φ\(y\)−α1−1/mmax⟨∇φ\(y\),χ∗,α1/mmaxy\(ξ\)⟩\.\\varphi\_\{\\alpha\}\(y,\\xi\):=\\varphi\(y\)\-\\alpha^\{1\-1/m\_\{\\max\}\}\\left\\langle\\nabla\\varphi\(y\),\\chi\_\{\*,\\alpha^\{1/m\_\{\\max\}\}y\}\(\\xi\)\\right\\rangle\.The calculation in Lemma[4\.9](https://arxiv.org/html/2607.16384#S4.Thmtheorem9)applies componentwise to this block\. For
M∗,x\(ξ\):=g∗\(x,ξ\)g∗\(x,ξ\)⊤\+g∗\(x,ξ\)χ∗,x\(ξ\)⊤\+χ∗,x\(ξ\)g∗\(x,ξ\)⊤,M\_\{\*,x\}\(\\xi\):=g\_\{\*\}\(x,\\xi\)g\_\{\*\}\(x,\\xi\)^\{\\top\}\+g\_\{\*\}\(x,\\xi\)\\chi\_\{\*,x\}\(\\xi\)^\{\\top\}\+\\chi\_\{\*,x\}\(\\xi\)g\_\{\*\}\(x,\\xi\)^\{\\top\},each entry is a finite sum of products of one\-dimensional terms\. The coordinatewise Lipschitz estimates forgig\_\{i\}, the Hölder estimates \([B\.2](https://arxiv.org/html/2607.16384#A2.E2)\) forχi\\chi\_\{i\}, and the coordinatewise stationary second moments therefore give, as in Lemma[B\.7](https://arxiv.org/html/2607.16384#A2.Thmtheorem7),
𝔼\[‖QM∗,α1/mmaxYα,∗\(ξα\)−QM∗,0\(ξα\)‖𝟏\{Yα,∗∈K\}\]⟶0\\mathbb\{E\}\\\!\\left\[\\left\\lVert QM\_\{\*,\\alpha^\{1/m\_\{\\max\}\}Y\_\{\\alpha,\*\}\}\(\\xi\_\{\\alpha\}\)\-QM\_\{\*,0\}\(\\xi\_\{\\alpha\}\)\\right\\rVert\\mathbf\{1\}\_\{\\\{Y\_\{\\alpha,\*\}\\in K\\\}\}\\right\]\\longrightarrow 0for every compactK⊂ℝqK\\subset\\mathbb\{R\}^\{q\}\. The entries ofQM∗,0QM\_\{\*,0\}belong toL1\(πΞ\)L^\{1\}\(\\pi\_\{\\Xi\}\)\. Applying Lemma[B\.8](https://arxiv.org/html/2607.16384#A2.Thmtheorem8)entrywise to∇2φ\(Yα,∗\)\\nabla^\{2\}\\varphi\(Y\_\{\\alpha,\*\}\)gives the required product limits\. Consequently, every subsequential weak limitν∗\\nu\_\{\*\}ofYα,∗Y\_\{\\alpha,\*\}satisfies
∫ℝq\[−⟨∇φ\(y\),h0,∗\(y\)⟩\+12tr\(Σact∇2φ\(y\)\)\]ν∗\(dy\)=0,φ∈Cc3\(ℝq\)\.\\int\_\{\\mathbb\{R\}^\{q\}\}\\left\[\-\\left\\langle\\nabla\\varphi\(y\),h\_\{0,\*\}\(y\)\\right\\rangle\+\\frac\{1\}\{2\}\\operatorname\{tr\}\\bigl\(\\Sigma\_\{\\rm act\}\\nabla^\{2\}\\varphi\(y\)\\bigr\)\\right\]\\nu\_\{\*\}\(dy\)=0,\\qquad\\varphi\\in C\_\{c\}^\{3\}\(\\mathbb\{R\}^\{q\}\)\.Because every coordinate inℐ∗\\mathcal\{I\}\_\{\*\}has exponentmmaxm\_\{\\max\}, the active drift satisfies, for constantsc,c′\>0c,c^\{\\prime\}\>0,
⟨y,h0,∗\(y\)⟩≥c∑i∈ℐ∗\|yi\|mmax≥c′‖y‖mmax,\\left\\langle y,h\_\{0,\*\}\(y\)\\right\\rangle\\geq c\\sum\_\{i\\in\\mathcal\{I\}\_\{\*\}\}\|y\_\{i\}\|^\{m\_\{\\max\}\}\\geq c^\{\\prime\}\\left\\lVert y\\right\\rVert^\{m\_\{\\max\}\},and the same argument gives⟨y−z,h0,∗\(y\)−h0,∗\(z\)⟩≥c′‖y−z‖mmax\.\\left\\langle y\-z,h\_\{0,\*\}\(y\)\-h\_\{0,\*\}\(z\)\\right\\rangle\\geq c^\{\\prime\}\\left\\lVert y\-z\\right\\rVert^\{m\_\{\\max\}\}\.Lemma[4\.10](https://arxiv.org/html/2607.16384#S4.Thmtheorem10)therefore identifies a unique invariant law for the active block of \([2\.19](https://arxiv.org/html/2607.16384#S2.E19)\)\. Hence the entire familyYα,∗Y\_\{\\alpha,\*\}converges to this law\.
It remains to consider coordinates outsideℐ∗\\mathcal\{I\}\_\{\*\}\. Ifi∉ℐ∗i\\notin\\mathcal\{I\}\_\{\*\}, then
X∞,i\(α\),sepα1/mmax=α1/mi−1/mmax\(X∞,i\(α\),sepα1/mi\)⟶0in probability,\\frac\{X\_\{\\infty,i\}^\{\(\\alpha\),\\rm sep\}\}\{\\alpha^\{1/m\_\{\\max\}\}\}=\\alpha^\{1/m\_\{i\}\-1/m\_\{\\max\}\}\\left\(\\frac\{X\_\{\\infty,i\}^\{\(\\alpha\),\\rm sep\}\}\{\\alpha^\{1/m\_\{i\}\}\}\\right\)\\longrightarrow 0\\qquad\\text\{in probability\},because the bracketed family is tight by the coordinatewise limit stated before the corollary\. In \([2\.19](https://arxiv.org/html/2607.16384#S2.E19)\), these inactive coordinates have no Brownian forcing and solvedYi,t=−hi,0\(Yi,t\)dt\.dY\_\{i,t\}=\-h\_\{i,0\}\(Y\_\{i,t\}\)\\,dt\.The one\-dimensional coercivity givesyihi,0\(yi\)≥c\|yi\|miy\_\{i\}h\_\{i,0\}\(y\_\{i\}\)\\geq c\|y\_\{i\}\|^\{m\_\{i\}\}, so every such deterministic flow converges to zero and hasδ0\\delta\_\{0\}as its unique invariant law\. The active block has the unique invariant law identified above\. Therefore \([2\.19](https://arxiv.org/html/2607.16384#S2.E19)\) has a unique invariant law, with inactive coordinates equal to zero\. Combining the active\-block convergence with the inactive\-coordinate convergence proves the asserted convergence toY∞sepY\_\{\\infty\}^\{\\rm sep\}\. ∎
### C\.3Proof of Remark[2\.18](https://arxiv.org/html/2607.16384#S2.Thmtheorem18)
###### Proof of Remark[2\.18](https://arxiv.org/html/2607.16384#S2.Thmtheorem18)\.
For each coordinateii, the one\-dimensional version of Theorem[2\.13](https://arxiv.org/html/2607.16384#S2.Thmtheorem13)givesX∞,i\(α\)α1/mi⇒Y∞,i\.\\frac\{X\_\{\\infty,i\}^\{\(\\alpha\)\}\}\{\\alpha^\{1/m\_\{i\}\}\}\\Rightarrow Y\_\{\\infty,i\}\.For each fixedα\\alpha, the stationary law factorizes by Remark[2\.16](https://arxiv.org/html/2607.16384#S2.Thmtheorem16), so the rescaled coordinates are independent\. Hence the joint characteristic function factorizes:
𝔼exp\(i∑i=1dtiX∞,i\(α\)α1/mi\)=∏i=1d𝔼exp\(itiX∞,i\(α\)α1/mi\)⟶∏i=1d𝔼eitiY∞,i\.\\mathbb\{E\}\\exp\\left\(i\\sum\_\{i=1\}^\{d\}t\_\{i\}\\frac\{X\_\{\\infty,i\}^\{\(\\alpha\)\}\}\{\\alpha^\{1/m\_\{i\}\}\}\\right\)=\\prod\_\{i=1\}^\{d\}\\mathbb\{E\}\\exp\\left\(it\_\{i\}\\frac\{X\_\{\\infty,i\}^\{\(\\alpha\)\}\}\{\\alpha^\{1/m\_\{i\}\}\}\\right\)\\longrightarrow\\prod\_\{i=1\}^\{d\}\\mathbb\{E\}e^\{it\_\{i\}Y\_\{\\infty,i\}\}\.Lévy’s continuity theorem gives the claimed joint convergence\. ∎
## References
- \[1\]S\. Allmeier and N\. Gast\(2024\)Computing the bias of constant\-step stochastic approximation with Markovian noise\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 137873–137902\.External Links:[Document](https://dx.doi.org/10.52202/079017-4379)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p7.2)\.
- \[2\]F\. Bach and É\. Moulines\(2013\)Non\-strongly\-convex smooth stochastic approximation with convergence rate O\(1/n\)\.InAdvances in Neural Information Processing Systems,Vol\.26,pp\. 773–781\.Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1)\.
- \[3\]A\. D\. Barbour\(1990\)Stein’s method for diffusion approximations\.Probab\. Theory Related Fields84\(3\),pp\. 297–322\.External Links:[Document](https://dx.doi.org/10.1007/BF01197887)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[4\]J\. T\. Barron\(2019\)A general and adaptive robust loss function\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 4331–4339\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2019.00446)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.p3.9),[§3\.2](https://arxiv.org/html/2607.16384#S3.SS2.p1.3)\.
- \[5\]P\. L\. Bartlett, M\. I\. Jordan, and J\. D\. McAuliffe\(2006\)Convexity, classification, and risk bounds\.J\. Amer\. Statist\. Assoc\.101\(473\),pp\. 138–156\.External Links:[Document](https://dx.doi.org/10.1198/016214505000000907)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.p3.9),[§3\.2](https://arxiv.org/html/2607.16384#S3.SS2.p1.3)\.
- \[6\]H\. H\. Bauschke and P\. L\. Combettes\(2010\)The Baillon–Haddad theorem revisited\.J\. Convex Anal\.17\(3–4\),pp\. 781–787\.Cited by:[§3\.2](https://arxiv.org/html/2607.16384#S3.SS2.SSS0.Px2.p4.10)\.
- \[7\]A\. Benveniste, M\. Métivier, and P\. Priouret\(1990\)Adaptive algorithms and stochastic approximations\.Applications of Mathematics, Vol\.22,Springer\-Verlag,Berlin\.External Links:[Document](https://dx.doi.org/10.1007/978-3-642-75894-2)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2607.16384#S1.p1.1),[Remark 2\.12](https://arxiv.org/html/2607.16384#S2.Thmtheorem12.p1.2)\.
- \[8\]V\. S\. Borkar\(1997\)Stochastic approximation with two time scales\.Systems Control Lett\.29\(5\),pp\. 291–294\.External Links:[Document](https://dx.doi.org/10.1016/S0167-6911%2897%2990015-3)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[9\]V\. S\. Borkar\(2008\)Stochastic approximation: a dynamical systems viewpoint\.Texts and Readings in Mathematics, Vol\.48,Hindustan Book Agency,Gurgaon\.External Links:[Document](https://dx.doi.org/10.1007/978-93-86279-38-5)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2607.16384#S1.p1.1)\.
- \[10\]A\. Braverman, J\. G\. Dai, and J\. Feng\(2016\)Stein’s method for steady\-state diffusion approximations: an introduction through the Erlang\-A and Erlang\-C models\.Stoch\. Syst\.6\(2\),pp\. 301–366\.External Links:[Document](https://dx.doi.org/10.1214/15-SSY212)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[11\]P\. Charbonnier, L\. Blanc\-Féraud, G\. Aubert, and M\. Barlaud\(1997\)Deterministic edge\-preserving regularization in computed imaging\.IEEE Trans\. Image Process\.6\(2\),pp\. 298–311\.External Links:[Document](https://dx.doi.org/10.1109/83.551699)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.p3.9),[§3\.2](https://arxiv.org/html/2607.16384#S3.SS2.p1.3)\.
- \[12\]Z\. Chen, S\. Mou, and S\. T\. Maguluri\(2022\)Stationary behavior of constant stepsize SGD type algorithms: an asymptotic characterization\.Proc\. ACM Meas\. Anal\. Comput\. Syst\.6\(1\),pp\. 19:1–19:24\.External Links:[Document](https://dx.doi.org/10.1145/3508039)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px1.p3.9),[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p3.3),[§1](https://arxiv.org/html/2607.16384#S1.p2.2)\.
- \[13\]A\. Dieuleveut, A\. Durmus, and F\. Bach\(2020\)Bridging the gap between constant step size stochastic gradient descent and Markov chains\.Ann\. Statist\.48\(3\),pp\. 1348–1382\.External Links:[Document](https://dx.doi.org/10.1214/19-AOS1850)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p2.1),[§1](https://arxiv.org/html/2607.16384#S1.p2.2)\.
- \[14\]A\. Eberle, A\. Guillin, and R\. Zimmer\(2019\)Quantitative Harris\-type theorems for diffusions and McKean–Vlasov processes\.Trans\. Amer\. Math\. Soc\.371\(10\),pp\. 7135–7173\.External Links:[Document](https://dx.doi.org/10.1090/tran/7576)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[15\]S\. N\. Ethier and T\. G\. Kurtz\(1986\)Markov processes: characterization and convergence\.Wiley Series in Probability and Mathematical Statistics,John Wiley & Sons,New York\.Cited by:[§B\.7](https://arxiv.org/html/2607.16384#A2.SS7.6.p4.6)\.
- \[16\]V\. Fabian\(1968\)On asymptotic normality in stochastic approximation\.Ann\. Math\. Statist\.39\(4\),pp\. 1327–1332\.External Links:[Document](https://dx.doi.org/10.1214/aoms/1177698258)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1)\.
- \[17\]N\. Gast\(2017\)Expected values estimated via mean\-field approximation are1/N1/N\-accurate\.Proc\. ACM Meas\. Anal\. Comput\. Syst\.1\(1\),pp\. 17:1–17:26\.External Links:[Document](https://dx.doi.org/10.1145/3084454)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p7.2)\.
- \[18\]H\. Hadavi, W\. Mou, S\. Samsonov, and H\. Wai\(2026\)Revisiting the constant stepsize stochastic approximation with decision\-dependent Markovian noise\.Note:arXiv:2604\.13378External Links:2604\.13378Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p5.2)\.
- \[19\]P\. J\. Huber\(1964\)Robust estimation of a location parameter\.Ann\. Math\. Statist\.35\(1\),pp\. 73–101\.External Links:[Document](https://dx.doi.org/10.1214/aoms/1177703732)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.p3.9),[§3\.2](https://arxiv.org/html/2607.16384#S3.SS2.p1.3)\.
- \[20\]D\. Huo, Y\. Chen, and Q\. Xie\(2026\)Bias and extrapolation in Markovian linear stochastic approximation with constant stepsizes\.Math\. Oper\. Res\.\.Note:Articles in AdvanceExternal Links:[Document](https://dx.doi.org/10.1287/moor.2024.0471)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p5.2)\.
- \[21\]D\. L\. Huo, Y\. Zhang, Y\. Chen, and Q\. Xie\(2024\)The collusion of memory and nonlinearity in stochastic approximation with constant stepsize\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 21699–21762\.External Links:[Document](https://dx.doi.org/10.52202/079017-0684)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p5.2)\.
- \[22\]H\. Kang and T\. G\. Kurtz\(2013\)Separation of time\-scales and model reduction for stochastic reaction networks\.Ann\. Appl\. Probab\.23\(2\),pp\. 529–583\.External Links:[Document](https://dx.doi.org/10.1214/12-AAP841)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[23\]R\. Z\. Khasminskii\(1968\)On the principle of averaging the Itô’s stochastic differential equations\.Kybernetika4\(3\),pp\. 260–279\.Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[24\]J\. Kiefer and J\. Wolfowitz\(1952\)Stochastic estimation of the maximum of a regression function\.Ann\. Math\. Statist\.23\(3\),pp\. 462–466\.External Links:[Document](https://dx.doi.org/10.1214/aoms/1177729392)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2607.16384#S1.p1.1)\.
- \[25\]K\. Knight\(1998\)Limiting distributions forL1L\_\{1\}regression estimators under general conditions\.Ann\. Statist\.26\(2\),pp\. 755–770\.External Links:[Document](https://dx.doi.org/10.1214/aos/1028144858)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.p3.9),[§3\.1](https://arxiv.org/html/2607.16384#S3.SS1.SSS0.Px1.p1.20),[§3\.1](https://arxiv.org/html/2607.16384#S3.SS1.p1.1)\.
- \[26\]R\. Koenker and G\. Bassett\(1978\)Regression quantiles\.Econometrica46\(1\),pp\. 33–50\.External Links:[Document](https://dx.doi.org/10.2307/1913643)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.p3.9),[§3\.1](https://arxiv.org/html/2607.16384#S3.SS1.p1.1)\.
- \[27\]H\. J\. Kushner and G\. G\. Yin\(2003\)Stochastic approximation and recursive algorithms and applications\.Second edition,Applications of Mathematics, Vol\.35,Springer,New York\.External Links:[Document](https://dx.doi.org/10.1007/b97441)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2607.16384#S1.p1.1),[Remark 2\.12](https://arxiv.org/html/2607.16384#S2.Thmtheorem12.p1.2),[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[28\]L\. Ljung\(1977\)Analysis of recursive stochastic algorithms\.IEEE Trans\. Automat\. Control22\(4\),pp\. 551–575\.External Links:[Document](https://dx.doi.org/10.1109/TAC.1977.1101561)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1)\.
- \[29\]S\. Mandt, M\. D\. Hoffman, and D\. M\. Blei\(2017\)Stochastic gradient descent as approximate Bayesian inference\.J\. Mach\. Learn\. Res\.18\(134\),pp\. 1–35\.Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p2.1),[§1](https://arxiv.org/html/2607.16384#S1.p2.2)\.
- \[30\]I\. Merad and S\. Gaïffas\(2025\)Convergence and concentration properties of constant step\-size SGD through Markov chains\.Electron\. J\. Stat\.19\(2\),pp\. 5843–5894\.External Links:[Document](https://dx.doi.org/10.1214/25-EJS2471)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p2.1),[§1](https://arxiv.org/html/2607.16384#S1.p2.2)\.
- \[31\]S\. P\. Meyn and R\. L\. Tweedie\(2009\)Markov chains and stochastic stability\.Second edition,Cambridge Mathematical Library,Cambridge University Press,Cambridge\.External Links:[Document](https://dx.doi.org/10.1017/CBO9780511626630)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[32\]É\. Moulines and F\. R\. Bach\(2011\)Non\-asymptotic analysis of stochastic approximation algorithms for machine learning\.InAdvances in Neural Information Processing Systems,Vol\.24,pp\. 451–459\.Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1)\.
- \[33\]A\. Nemirovski, A\. Juditsky, G\. Lan, and A\. Shapiro\(2009\)Robust stochastic approximation approach to stochastic programming\.SIAM J\. Optim\.19\(4\),pp\. 1574–1609\.External Links:[Document](https://dx.doi.org/10.1137/070704277)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1)\.
- \[34\]É\. Pardoux and A\. Yu\. Veretennikov\(2001\)On the Poisson equation and diffusion approximation\. I\.Ann\. Probab\.29\(3\),pp\. 1061–1085\.External Links:[Document](https://dx.doi.org/10.1214/aop/1015345596)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[35\]É\. Pardoux and A\. Yu\. Veretennikov\(2003\)On Poisson equation and diffusion approximation\. II\.Ann\. Probab\.31\(3\),pp\. 1166–1192\.External Links:[Document](https://dx.doi.org/10.1214/aop/1055425774)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[36\]É\. Pardoux and A\. Yu\. Veretennikov\(2005\)On the Poisson equation and diffusion approximation\. III\.Ann\. Probab\.33\(3\),pp\. 1111–1133\.External Links:[Document](https://dx.doi.org/10.1214/009117905000000062)Cited by:[§5](https://arxiv.org/html/2607.16384#S5.p3.1)\.
- \[37\]G\. Ch\. Pflug\(1986\)Stochastic minimization with constant step\-size: asymptotic laws\.SIAM J\. Control Optim\.24\(4\),pp\. 655–666\.External Links:[Document](https://dx.doi.org/10.1137/0324039)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p2.1),[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p3.3),[§1](https://arxiv.org/html/2607.16384#S1.p2.2)\.
- \[38\]B\. T\. Polyak and A\. B\. Juditsky\(1992\)Acceleration of stochastic approximation by averaging\.SIAM J\. Control Optim\.30\(4\),pp\. 838–855\.External Links:[Document](https://dx.doi.org/10.1137/0330046)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1)\.
- \[39\]Y\. Qu, J\. Blanchet, and P\. W\. Glynn\(2025\)Computable bounds on convergence of Markov chains in Wasserstein distance via contractive drift\.Ann\. Appl\. Probab\.35\(4\),pp\. 2678–2715\.External Links:[Document](https://dx.doi.org/10.1214/25-AAP2184)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px1.p2.4),[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p6.1),[§2\.1](https://arxiv.org/html/2607.16384#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.16384#S4.SS1.p1.3),[§4\.1](https://arxiv.org/html/2607.16384#S4.SS1.p2.1)\.
- \[40\]H\. Robbins and S\. Monro\(1951\)A stochastic approximation method\.Ann\. Math\. Statist\.22\(3\),pp\. 400–407\.External Links:[Document](https://dx.doi.org/10.1214/aoms/1177729586)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2607.16384#S1.p1.1)\.
- \[41\]R\. T\. Rockafellar and S\. Uryasev\(2000\)Optimization of Conditional Value\-at\-Risk\.J\. Risk2\(3\),pp\. 21–41\.External Links:[Document](https://dx.doi.org/10.21314/JOR.2000.038)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.p3.9),[§3\.1](https://arxiv.org/html/2607.16384#S3.SS1.SSS0.Px1.p2.2)\.
- \[42\]D\. Ruppert\(1988\)Efficient estimations from a slowly convergent Robbins–Monro process\.Technical ReportTechnical Report781,School of Operations Research and Industrial Engineering, Cornell University,Ithaca, NY\.Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p1.1)\.
- \[43\]Z\. Wang, Y\. Wang, I\. Narang, F\. Wang, Y\. Wang, and S\. T\. Maguluri\(2026\)Steady\-state behavior of constant\-stepsize stochastic approximation: gaussian approximation and tail bounds\.Note:arXiv:2602\.13960External Links:2602\.13960Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px1.p3.9),[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p3.3),[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p7.2)\.
- \[44\]Z\. Wei, J\. Li, Z\. Lou, and W\. B\. Wu\(2025\)Gaussian approximation and concentration of constant learning\-rate stochastic gradient descent\.InAdvances in Neural Information Processing Systems,Vol\.38\.Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p7.2)\.
- \[45\]L\. Yu, K\. Balasubramanian, S\. Volgushev, and M\. A\. Erdogdu\(2021\)An analysis of constant step size SGD in the non\-convex regime: asymptotic normality and bias\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 4234–4248\.Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p2.1)\.
- \[46\]Y\. Zhang, D\. L\. Huo, Y\. Chen, and Q\. Xie\(2024\)Prelimit coupling and steady\-state convergence of constant\-stepsize nonsmooth contractive SA\.ACM SIGMETRICS Perform\. Eval\. Rev\.52\(1\),pp\. 35–36\.External Links:[Document](https://dx.doi.org/10.1145/3673660.3655076)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p3.3)\.
- \[47\]Y\. Zhang, D\. L\. Huo, Y\. Chen, and Q\. Xie\(2025\)A piecewise Lyapunov analysis of sub\-quadratic SGD: applications to robust and quantile regression\.ACM SIGMETRICS Perform\. Eval\. Rev\.53\(1\),pp\. 85–87\.Note:Full version: arXiv:2504\.08178External Links:[Document](https://dx.doi.org/10.1145/3744970.3727269)Cited by:[§1](https://arxiv.org/html/2607.16384#S1.SS0.SSS0.Px2.p4.4),[§1](https://arxiv.org/html/2607.16384#S1.p3.9)\.Similar Articles
Flatland: The Adventures of Gradient Descent with Large Step Sizes
This paper addresses the open question of maximum step size for gradient descent convergence on non-L-smooth objectives, introducing adaptive methods that operate at the edge of stability and can minimize sharpness globally.
Uniform Stability and Generalization Error of GD and SGD on Fixed-Point Parameters
This paper analyzes generalization error, uniform stability, and uniform argument stability of gradient descent (GD) and stochastic gradient descent (SGD) over discrete parameter spaces with deterministic or stochastic rounding, showing that rounding degrades generalization for GD and introduces dimension-dependent errors for stochastic rounding.
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
This paper generalizes non-uniform smoothness assumptions to objectives whose curvature is affine in the objective value, proving convergence rates for steepest descent and diagonal variants of RMSProp and Adam, with applications to logistic regression and neural networks.
From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression
This paper derives batch scaling laws for sketched linear regression under power-law spectra, analyzing one-pass and multi-pass mini-batch SGD. It provides explicit risk decompositions showing how batch size affects bias, variance, and fluctuation terms, and establishes that without-replacement sampling yields lower noise than with-replacement.
How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size
Proposes a three-term scaling law that decouples model size, training steps, and batch size, enabling robust fitting with fewer runs and deriving scaling laws for suboptimal batch sizes.