High-Probability PL-SGD with Markovian Noise: Optimal Mixing and Tail Dependence

arXiv cs.LG Papers

Summary

This paper provides optimal high-probability bounds for stochastic gradient descent under Markovian noise for PL-smooth objectives, closing gaps between expectation and high-probability guarantees and extending to heavy-tailed settings with matching lower bounds.

arXiv:2606.26316v1 Announce Type: new Abstract: We study first-order methods for smooth objectives satisfying the Polyak-\L{}ojasiewicz (PL) condition when gradient samples are generated by an exogenous Markov chain. In the light-tailed setting, prior uniform-in-time high-probability bounds for ordinary Stochastic Gradient Descent (SGD) under a standard growth envelope scale as $\widetilde{O}(t_{mix}^2/k)$, leaving a gap with the $\widetilde{O}(t_{mix}/k)$ expectation bounds. We close this gap using a lag-blocking argument to establish a uniform high-probability guarantee with a leading stochastic term of $\widetilde{O}(t_{mix}/(k+K_0))$ under geometric mixing. We prove this linear dependence on the mixing time is optimal via a matching $\Omega(\sigma^2 t_{mix}/k)$ lower bound on a quadratic objective driven by a persistent two-state chain. We then extend this framework to heavy-tailed Markovian gradients satisfying a stationary finite-$p$-moment condition, $p \in (1,2]$. We design an all-samples clipped block method that uses every Markov transition while mitigating Markovian bias. Under a transition budget $T$, this algorithm achieves a high-probability stochastic error of $\widetilde{O}(\sigma_p^2(t_{mix}/T)^{2(p-1)/p})$. We establish a matching lower bound by reducing PL optimization to heavy-tailed mean estimation for a sticky Markov chain. Ultimately, this work tightly characterizes the optimal polynomial dependence on mixing time for light-tailed PL-SGD, and the optimal heavy-tail exponent and effective-sample-size dependence in the robust regime.
Original Article
View Cached Full Text

Cached at: 06/26/26, 05:18 AM

# High-Probability PL-SGD with Markovian Noise: Optimal Mixing and Tail Dependence
Source: [https://arxiv.org/html/2606.26316](https://arxiv.org/html/2606.26316)
Dhruv Sarkar1,2Aprameyo Chakrabartty1Vaneet Aggarwal3 1Indian Institute of Technology Kharagpur 2Mohamed bin Zayed University of Artificial Intelligence 3Purdue University dhruv\.sarkar223@gmail\.comaprameyo8858@gmail\.comvaneet@purdue\.edu

###### Abstract

We study first\-order methods for smooth objectives satisfying the Polyak\-Łojasiewicz \(PL\) condition when gradient samples are generated by an exogenous Markov chain\. In the light\-tailed setting, prior uniform\-in\-time high\-probability bounds for ordinary Stochastic Gradient Descent \(SGD\) under a standard growth envelope scale asO~​\(tmix2/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}^\{2\}/k\), leaving a gap with theO~​\(tmix/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/k\)expectation bounds\. We close this gap using a lag\-blocking argument to establish a uniform high\-probability guarantee with a leading stochastic term ofO~​\(tmix/\(k\+K0\)\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/\(k\+K\_\{0\}\)\)under geometric mixing\. We prove this linear dependence on the mixing time is optimal via a matchingΩ​\(σ2​tmix/k\)\\Omega\(\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}/k\)lower bound on a quadratic objective driven by a persistent two\-state chain\.

We then extend this framework to heavy\-tailed Markovian gradients satisfying a stationary finite\-pp\-moment condition,p∈\(1,2\]p\\in\(1,2\]\. We design an all\-samples clipped block method that uses every Markov transition while mitigating Markovian bias\. Under a transition budgetTT, this algorithm achieves a high\-probability stochastic error ofO~​\(σp2​\(tmix/T\)2​\(p−1\)/p\)\\widetilde\{O\}\(\\sigma\_\{p\}^\{2\}\(t\_\{\\mathrm\{mix\}\}/T\)^\{2\(p\-1\)/p\}\)\. We establish a matching lower bound by reducing PL optimization to heavy\-tailed mean estimation for a sticky Markov chain\. Ultimately, this work tightly characterizes the optimal polynomial dependence on mixing time for light\-tailed PL\-SGD, and the optimal heavy\-tail exponent and effective\-sample\-size dependence in the robust regime\.

Keywords:stochastic gradient descent; Markovian noise; Polyak\-Łojasiewicz condition; high\-probability bounds; mixing time; heavy\-tailed noise; clipped SGD; lower bounds\.

## 1Introduction

Stochastic gradient methods are a central tool in stochastic approximation and large\-scale optimization\. Classical analyses of stochastic approximation and Stochastic Gradient Descent \(SGD\) typically rely on either independent samples or conditionally unbiased martingale\-difference gradient noise\[[57](https://arxiv.org/html/2606.26316#bib.bib58),[37](https://arxiv.org/html/2606.26316#bib.bib38),[5](https://arxiv.org/html/2606.26316#bib.bib8),[39](https://arxiv.org/html/2606.26316#bib.bib40),[9](https://arxiv.org/html/2606.26316#bib.bib12),[10](https://arxiv.org/html/2606.26316#bib.bib13)\]\. Under these assumptions, the stochastic error in the descent recursion can be controlled by martingale or concentration arguments, and nonasymptotic convergence guarantees are now well understood in many convex, strongly convex, and nonconvex regimes\[[50](https://arxiv.org/html/2606.26316#bib.bib51),[27](https://arxiv.org/html/2606.26316#bib.bib28),[36](https://arxiv.org/html/2606.26316#bib.bib37)\]\. In many modern applications, however, the samples used by the gradient oracle are generated sequentially by a Markov chain rather than independently\. This occurs in token\-based and random\-walk decentralized optimization\[[30](https://arxiv.org/html/2606.26316#bib.bib31),[48](https://arxiv.org/html/2606.26316#bib.bib49),[28](https://arxiv.org/html/2606.26316#bib.bib29)\], Markov Chain Monte Carlo gradient estimation, privacy\-preserving subsampling schemes with temporal exclusion rules\[[1](https://arxiv.org/html/2606.26316#bib.bib1),[14](https://arxiv.org/html/2606.26316#bib.bib16),[19](https://arxiv.org/html/2606.26316#bib.bib22)\], online system identification\[[38](https://arxiv.org/html/2606.26316#bib.bib39)\], and reinforcement\-learning and temporal\-difference algorithms\[[7](https://arxiv.org/html/2606.26316#bib.bib10),[59](https://arxiv.org/html/2606.26316#bib.bib60),[32](https://arxiv.org/html/2606.26316#bib.bib33),[43](https://arxiv.org/html/2606.26316#bib.bib44),[4](https://arxiv.org/html/2606.26316#bib.bib2),[24](https://arxiv.org/html/2606.26316#bib.bib4),[23](https://arxiv.org/html/2606.26316#bib.bib5)\]\. In such settings the gradient oracle is generally unbiased only after averaging with respect to the invariant distribution of the chain\. At a finite time, the conditional law of the Markov state need not be stationary, and the resulting bias must be controlled through the mixing behavior of the chain\.

This paper studies this problem for smooth objectives satisfying the Polyak–Łojasiewicz \(PL\) inequality\. The PL condition, introduced in\[[54](https://arxiv.org/html/2606.26316#bib.bib55)\]and developed in its modern optimization form in\[[35](https://arxiv.org/html/2606.26316#bib.bib36)\], is weaker than strong convexity but still yields global linear convergence for deterministic gradient descent\. It is a natural condition in overparameterized learning, least\-squares problems, control, and other settings where global convergence can hold without convexity\. We consider the Markovian SGD recursion

xk\+1=xk−αk​Gk,Gk=g​\(xk,Zk\)\+Mk\+1,x\_\{k\+1\}=x\_\{k\}\-\\alpha\_\{k\}G\_\{k\},\\qquad G\_\{k\}=g\(x\_\{k\},Z\_\{k\}\)\+M\_\{k\+1\},\(1\)where\(Zk\)k≥0\(Z\_\{k\}\)\_\{k\\geq 0\}is an exogenous Markov chain with invariant distributionπ\\pi, the stationary oracle satisfies

∫g​\(x,z\)​π​\(d​z\)=∇f​\(x\),\\int g\(x,z\)\\,\\pi\(dz\)=\\nabla f\(x\),and\(Mk\+1\)k≥0\(M\_\{k\+1\}\)\_\{k\\geq 0\}is a martingale\-difference perturbation\. The oracle is allowed to satisfy a boundedAA\-BB\-CCgrowth envelope, called the ABC condition because the constantsA,B,CA,B,Ccontrol the gradient\-dependent term, the objective\-gap\-dependent term, and the additive noise floor:

‖Gk‖2≤A​‖∇f​\(xk\)‖2\+B​\(f​\(xk\)−f⋆\)\+C\.\\\|G\_\{k\}\\\|^\{2\}\\leq A\\\|\\nabla f\(x\_\{k\}\)\\\|^\{2\}\+B\(f\(x\_\{k\}\)\-f^\{\\star\}\)\+C\.This condition is used in prior PL\-SGD analyses\[[33](https://arxiv.org/html/2606.26316#bib.bib34)\]\. This growth model is important because a uniform bounded\-variance assumption can be too restrictive in PL problems, least\-squares models, interpolation regimes, and minibatch sampling\.

There is now a substantial literature on first\-order methods and stochastic approximation with Markovian sampling\. Finite\-time analyses of Markovian stochastic gradient methods were developed in convex, strongly convex, and nonconvex settings by\[[60](https://arxiv.org/html/2606.26316#bib.bib61),[17](https://arxiv.org/html/2606.26316#bib.bib20),[18](https://arxiv.org/html/2606.26316#bib.bib21),[21](https://arxiv.org/html/2606.26316#bib.bib24)\]\. More general Markovian stochastic\-approximation and reinforcement\-learning problems are often handled using Poisson\-equation decompositions\[[5](https://arxiv.org/html/2606.26316#bib.bib8),[39](https://arxiv.org/html/2606.26316#bib.bib40),[9](https://arxiv.org/html/2606.26316#bib.bib12),[49](https://arxiv.org/html/2606.26316#bib.bib50),[62](https://arxiv.org/html/2606.26316#bib.bib3)\]\. Recent work has also studied optimal mixing\-time dependence for several first\-order methods using randomized batching, multilevel, or variance\-reduction ideas\[[6](https://arxiv.org/html/2606.26316#bib.bib9),[21](https://arxiv.org/html/2606.26316#bib.bib24)\]\. In the PL setting most closely related to ours, Kar, Chandak, Singh, Moulines, Bhatnagar, and Bambos\[[33](https://arxiv.org/html/2606.26316#bib.bib34)\]proved the first uniform\-in\-time high\-probability guarantee for SGD with Markovian plus martingale\-difference noise under an ABC oracle envelope\. Their expectation bound has leading stochastic orderO~​\(tmix/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/k\), while their uniform high\-probability bound has leading stochastic orderO~​\(tmix2/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}^\{2\}/k\)\. Thus, before the present work, it was unclear whether high\-probability PL\-SGD truly requires a quadratic polynomial dependence on the mixing time, or whether the linear dependence suggested by expectation bounds is achievable\.

Our first main result shows that the optimal polynomial dependence is linear\. For ordinary one\-sample SGD under the bounded ABC oracle envelope, we prove a uniform\-in\-time high\-probability bound whose leading stochastic term scales as

O~​\(tmixk\+K0\)\\widetilde\{O\}\\\!\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{k\+K\_\{0\}\}\\right\)under geometric mixing, up to logarithmic factors and the usual optimization transientK0​\{f​\(x0\)−f⋆\}/\(k\+K0\)K\_\{0\}\\\{f\(x\_\{0\}\)\-f^\{\\star\}\\\}/\(k\+K\_\{0\}\)\. The hidden constants depend on the smoothness, PL, and bounded\-ABC envelope constants, including the noise\-scale term in[Assumption˜2](https://arxiv.org/html/2606.26316#Thmassumption2)\. We also prove a matching lower bound on a one\-dimensional quadratic PL objective driven by a persistent two\-state Markov chain:

f​\(xk\)−f⋆=Ω​\(σ2​tmixk\)f\(x\_\{k\}\)\-f^\{\\star\}=\\Omega\\\!\\left\(\\frac\{\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}\}\{k\}\\right\)in expectation and with constant probability\. Therefore no theorem valid over the same class of instances can improve the polynomial dependence ontmixt\_\{\\mathrm\{mix\}\}below linear, apart from logarithmic factors\.

### 1\.1Technical challenge and proof novelty

The main technical obstacle is that the Markovian term in the PL descent recursion is an*adaptive*Markovian additive functional\. In the weighted descent recursion of[Lemma˜2](https://arxiv.org/html/2606.26316#Thmlemma2), the stochastic part contains

∑ℓ=0k−1wℓ,k​h​\(xℓ,Zℓ\),h​\(x,z\)=⟨∇f​\(x\),g​\(x,z\)−∇f​\(x\)⟩,\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\),\\qquad h\(x,z\)=\\big\\langle\\nabla f\(x\),\\,g\(x,z\)\-\\nabla f\(x\)\\big\\rangle,wherehhis defined in \([16](https://arxiv.org/html/2606.26316#S3.E16)\)\. For each fixedxx, the stationary\-gradient identity in[Assumption˜2](https://arxiv.org/html/2606.26316#Thmassumption2)implies

∫h​\(x,z\)​π​\(d​z\)=0\.\\int h\(x,z\)\\,\\pi\(dz\)=0\.However, conditionally on the past,ZℓZ\_\{\\ell\}need not be distributed according toπ\\pi, andxℓx\_\{\\ell\}is itself generated from the previous Markov states and martingale perturbations\. Consequently,h​\(xℓ,Zℓ\)h\(x\_\{\\ell\},Z\_\{\\ell\}\)is neither a martingale difference nor a fixed additive functional of the Markov chain\. Standard martingale concentration cannot be applied directly, and classical concentration inequalities for fixed Markov\-chain observables do not apply in a black\-box way\.

#### The limitation of the Poisson equation\.

The standard mechanism for decoupling this temporal dependence relies on solutions to the Poisson equation\. This framework yields a martingale difference sequence by representing the local gradient bias through the solutionV​\(x,z\)V\(x,z\)\. For instance, in the recent high\-probability PL\-SGD analysis by Kar et al\.\[[33](https://arxiv.org/html/2606.26316#bib.bib34)\], this solution is defined \(cf\. Lemma 4\.1 therein\) as the infinite\-horizon sum:

V​\(x,z\):=𝔼​\[∑j=0∞\(g​\(x,Zj\)−∇f​\(x\)\)\|Z0=z\]\.V\(x,z\):=\\mathbb\{E\}\\left\[\\sum\_\{j=0\}^\{\\infty\}\\big\(g\(x,Z\_\{j\}\)\-\\nabla f\(x\)\\big\)\\;\\Big\|\\;Z\_\{0\}=z\\right\]\.\(2\)This allows the Markovian noise to be perfectly decomposed asg​\(x,z\)−∇f​\(x\)=V​\(x,z\)−∫V​\(x,z′\)​p​\(d​z′\|z\)g\(x,z\)\-\\nabla f\(x\)=V\(x,z\)\-\\int V\(x,z^\{\\prime\}\)p\(dz^\{\\prime\}\|z\)\. While this identity successfully isolates a strict martingale difference sequence, defined asM~ℓ\+1:=V​\(xℓ,Zℓ\+1\)−∫V​\(xℓ,z\)​p​\(d​z\|Zℓ\)\\tilde\{M\}\_\{\\ell\+1\}:=V\(x\_\{\\ell\},Z\_\{\\ell\+1\}\)\-\\int V\(x\_\{\\ell\},z\)p\(dz\|Z\_\{\\ell\}\), it achieves this by inadvertently amplifying the magnitude of the stochastic increments\.

Because the Markov chain requires𝒪​\(tmix\)\\mathcal\{O\}\(t\_\{\\mathrm\{mix\}\}\)steps to contract toward the invariant measureπ\\pi, the geometric series definingV​\(x,z\)V\(x,z\)forces its magnitude to scale linearly with the mixing time\. Kar et al\. formalize this dependence in their Lemma C\.1, establishing the structural upper bound‖V​\(x,z\)‖2≤𝒪​\(tmix2⋅Δ​\(x\)\)\\\|V\(x,z\)\\\|^\{2\}\\leq\\mathcal\{O\}\(t\_\{\\mathrm\{mix\}\}^\{2\}\\cdot\\Delta\(x\)\)\.

When uniform\-in\-time concentration inequalities, such as the Azuma\-Hoeffding bound, are applied to this derived martingale, the variance proxy relies on the square of the almost\-sure increment bounds\. By measuring the deviations of a martingale whose increments have inherently bounded ranges of𝒪​\(tmix\)\\mathcal\{O\}\(t\_\{\\mathrm\{mix\}\}\), the resulting concentration bound accumulates a𝒪​\(tmix2\)\\mathcal\{O\}\(t\_\{\\mathrm\{mix\}\}^\{2\}\)polynomial penalty\. This quadratic dependence reflects a limitation of that proof route and leaves a gap relative to the optimal linear𝒪​\(tmix\)\\mathcal\{O\}\(t\_\{\\mathrm\{mix\}\}\)scaling achievable in expectation bounds\.

#### Our approach: lag\-blocking and temporal conditioning\.

Our proof avoids this loss by changing the conditioning structure instead of changing the observable through a Poisson corrector\. For a fixed target timekk, choose the analytical delaymkm\_\{k\}from \([18](https://arxiv.org/html/2606.26316#S4.E18)\)\. Forℓ≥mk\\ell\\geq m\_\{k\}, decompose

h​\(xℓ,Zℓ\)=h​\(xℓ−mk,Zℓ\)\+\{h​\(xℓ,Zℓ\)−h​\(xℓ−mk,Zℓ\)\}\.h\(x\_\{\\ell\},Z\_\{\\ell\}\)=h\(x\_\{\\ell\-m\_\{k\}\},Z\_\{\\ell\}\)\+\\bigl\\\{h\(x\_\{\\ell\},Z\_\{\\ell\}\)\-h\(x\_\{\\ell\-m\_\{k\}\},Z\_\{\\ell\}\)\\bigr\\\}\.The firstmkm\_\{k\}indices are treated separately as an initial\-window term, controlled in[Lemma˜9](https://arxiv.org/html/2606.26316#Thmlemma9)\. For the remaining terms, the delayed iteratexℓ−mkx\_\{\\ell\-m\_\{k\}\}isℱℓ−mk\\mathcal\{F\}\_\{\\ell\-m\_\{k\}\}\-measurable, while by[Assumption˜3](https://arxiv.org/html/2606.26316#Thmassumption3)the chain hasmkm\_\{k\}transitions to move toward stationarity before timeℓ\\ell\. With

Yℓ:=wℓ,k​h​\(xℓ−mk,Zℓ\),Y¯ℓ:=𝔼​\[Yℓ∣ℱℓ−mk\],Y\_\{\\ell\}:=w\_\{\\ell,k\}h\(x\_\{\\ell\-m\_\{k\}\},Z\_\{\\ell\}\),\\qquad\\overline\{Y\}\_\{\\ell\}:=\\mathbb\{E\}\[Y\_\{\\ell\}\\mid\\mathcal\{F\}\_\{\\ell\-m\_\{k\}\}\],as in \([65](https://arxiv.org/html/2606.26316#A3.E65)\), the Markovian sum is split into an initial window, a centered delayed martingale part, a mixing\-bias part, and a replacement part:

∑ℓ=0k−1wℓ,k​h​\(xℓ,Zℓ\)\\displaystyle\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)=∑ℓ=0mk−1wℓ,k​h​\(xℓ,Zℓ\)\+∑ℓ=mkk−1\(Yℓ−Y¯ℓ\)\\displaystyle=\\sum\_\{\\ell=0\}^\{m\_\{k\}\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)\+\\sum\_\{\\ell=m\_\{k\}\}^\{k\-1\}\(Y\_\{\\ell\}\-\\overline\{Y\}\_\{\\ell\}\)\+∑ℓ=mkk−1Y¯ℓ\+∑ℓ=mkk−1wℓ,k​\{h​\(xℓ,Zℓ\)−h​\(xℓ−mk,Zℓ\)\}\.\\displaystyle\\quad\+\\sum\_\{\\ell=m\_\{k\}\}^\{k\-1\}\\overline\{Y\}\_\{\\ell\}\+\\sum\_\{\\ell=m\_\{k\}\}^\{k\-1\}w\_\{\\ell,k\}\\bigl\\\{h\(x\_\{\\ell\},Z\_\{\\ell\}\)\-h\(x\_\{\\ell\-m\_\{k\}\},Z\_\{\\ell\}\)\\bigr\\\}\.
The centered delayed term is controlled by the residue\-class martingale argument in[Lemma˜6](https://arxiv.org/html/2606.26316#Thmlemma6)\. The indices are split modulomkm\_\{k\}\. Along a fixed residue class, consecutive delayed terms are separated by exactlymkm\_\{k\}Markov transitions, so after subtractingY¯ℓ\\overline\{Y\}\_\{\\ell\}they form a martingale\-difference sequence with respect to a down\-sampled filtration\. For the actual PL recursion, the relevant object is the weighted sum above, and the residue\-class bounds involve the weighted square sum∑ℓ<kwℓ,k2\\sum\_\{\\ell<k\}w\_\{\\ell,k\}^\{2\}\. Summing over themkm\_\{k\}residue classes produces a variance proxy of order

mk​∑ℓ<kwℓ,k2≍mkk\+K0,m\_\{k\}\\sum\_\{\\ell<k\}w\_\{\\ell,k\}^\{2\}\\asymp\\frac\{m\_\{k\}\}\{k\+K\_\{0\}\},up to logarithmic and offset factors; the required weight estimates are collected in[Lemma˜3](https://arxiv.org/html/2606.26316#Thmlemma3)\. Under geometric mixing,mk=O~​\(tmix\)m\_\{k\}=\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}\), which is the source of the leading stochastic orderO~​\(tmix/\(k\+K0\)\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/\(k\+K\_\{0\}\)\)in[Corollary˜1](https://arxiv.org/html/2606.26316#Thmcorollary1)\. No independence between different residue classes is required: concentration is applied separately within each residue class, and the resulting high\-probability bounds are combined by a union bound\.

The remaining terms are controlled separately\. The mixing bias∑ℓ=mkk−1Y¯ℓ\\sum\_\{\\ell=m\_\{k\}\}^\{k\-1\}\\overline\{Y\}\_\{\\ell\}is small by the total\-variation mixing condition in[Assumption˜3](https://arxiv.org/html/2606.26316#Thmassumption3)and the bound in[Lemma˜7](https://arxiv.org/html/2606.26316#Thmlemma7)\. The replacement term is deterministic on the stopped good event: by[Lemma˜4](https://arxiv.org/html/2606.26316#Thmlemma4),

\|h​\(xℓ,z\)−h​\(xℓ−mk,z\)\|≤clip​\(1\+Δ​\(xℓ\)\+Δ​\(xℓ−mk\)\)​‖xℓ−xℓ−mk‖,\\bigl\|h\(x\_\{\\ell\},z\)\-h\(x\_\{\\ell\-m\_\{k\}\},z\)\\bigr\|\\leq c\_\{\\mathrm\{lip\}\}\\bigl\(1\+\\sqrt\{\\Delta\(x\_\{\\ell\}\)\}\+\\sqrt\{\\Delta\(x\_\{\\ell\-m\_\{k\}\}\)\}\\bigr\)\\\|x\_\{\\ell\}\-x\_\{\\ell\-m\_\{k\}\}\\\|,and

‖xℓ−xℓ−mk‖≤∑i=ℓ−mkℓ−1αi​‖Gi‖\.\\\|x\_\{\\ell\}\-x\_\{\\ell\-m\_\{k\}\}\\\|\\leq\\sum\_\{i=\\ell\-m\_\{k\}\}^\{\\ell\-1\}\\alpha\_\{i\}\\\|G\_\{i\}\\\|\.The ABC envelope in[Assumption˜2](https://arxiv.org/html/2606.26316#Thmassumption2)controls the gradients on the stopped event, and the resulting replacement bound is stated in[Lemma˜8](https://arxiv.org/html/2606.26316#Thmlemma8)\. These ingredients are combined in[Lemma˜10](https://arxiv.org/html/2606.26316#Thmlemma10), and the first\-failure argument in Appendix[D](https://arxiv.org/html/2606.26316#A4)yields the uniform\-in\-time guarantee of[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)\. The delaymkm\_\{k\}and residue classes are only analytical devices; the algorithm remains the ordinary one\-sample SGD recursion \([3](https://arxiv.org/html/2606.26316#S3.E3)\)\.

We also study a heavy\-tailed Markovian regime\. In many applications, individual gradient samples may have no deterministic envelope and may not even have finite variance\. High\-probability guarantees under only finitepp\-th moments,p∈\(1,2\]p\\in\(1,2\], therefore require robustification\. For independent or martingale\-type noise, clipping, normalization, median\-of\-means, and related robust methods have been used to obtain high\-probability guarantees\[[15](https://arxiv.org/html/2606.26316#bib.bib17),[26](https://arxiv.org/html/2606.26316#bib.bib27),[58](https://arxiv.org/html/2606.26316#bib.bib59),[52](https://arxiv.org/html/2606.26316#bib.bib53),[55](https://arxiv.org/html/2606.26316#bib.bib56),[29](https://arxiv.org/html/2606.26316#bib.bib30),[3](https://arxiv.org/html/2606.26316#bib.bib7)\]\. Markovian heavy\-tailed data are more delicate because dependence reduces the effective sample size and nonstationary initialization creates an additional bias\. We address this by holding the iterate fixed over a block, clipping every consecutive Markovian gradient in the block, and averaging all clipped samples\. No samples inside the block are discarded; the mixing window enters only through the analysis and the choice of clipping scale\. Under geometric mixing and total transition budgetTT, the resulting high\-probability stochastic term has order

O~​\(σp2​\(tmixT\)2​\(p−1\)/p\),\\widetilde\{O\}\\\!\\left\(\\sigma\_\{p\}^\{2\}\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{T\}\\right\)^\{2\(p\-1\)/p\}\\right\),up to dimension and logarithmic factors\. A sticky\-chain lower bound shows that the exponent2​\(p−1\)/p2\(p\-1\)/pand the dependence on the Markovian effective sample size cannot be improved, up to logarithmic and dimension factors\. The lower bound is one\-dimensional, so we do not claim optimality of the dimension dependence in the upper bound\.

### 1\.2Contributions\.

The paper makes the following contributions\.

1. 1\.We prove a uniform\-in\-time high\-probability PL\-SGD upper bound under a general uniform total\-variation mixing profile on a general measurable Markov state space\. The result applies to ordinary one\-sample SGD and does not require compactness of the state space\.
2. 2\.Under geometric mixing, the leading stochastic term in the bounded\-ABC regime satisfies O~​\(tmixk\+K0\),\\widetilde\{O\}\\\!\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{k\+K\_\{0\}\}\\right\),with no additional polynomial power oftmixt\_\{\\mathrm\{mix\}\}\. This closes the gap between theO~​\(tmix2/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}^\{2\}/k\)high\-probability behavior obtained by Poisson\-equation\-type analyses and theO~​\(tmix/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/k\)behavior suggested by expectation bounds\.
3. 3\.We prove a matching lower bound\. Even for the quadratic objectivef​\(x\)=x2/2f\(x\)=x^\{2\}/2and a bounded two\-state Markovian oracle, SGD satisfies Pr⁡\(f​\(xk\)−f⋆≥c​σ2​tmixk\)≥c0\\Pr\\\!\\left\(f\(x\_\{k\}\)\-f^\{\\star\}\\geq c\\,\\frac\{\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}\}\{k\}\\right\)\\geq c\_\{0\}for universal constantsc,c0\>0c,c\_\{0\}\>0\. Thus linear polynomial dependence on the mixing time is unavoidable at constant confidence\.
4. 4\.We extend the lag\-blocking viewpoint to finite\-ppheavy\-tailed Markovian gradients\. The proposed all\-samples clipped block method uses every Markov transition in each block and achieves the transition\-budget rate O~​\(σp2​\(tmixT\)2​\(p−1\)/p\)\\widetilde\{O\}\\\!\\left\(\\sigma\_\{p\}^\{2\}\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{T\}\\right\)^\{2\(p\-1\)/p\}\\right\)under geometric mixing, up to logarithmic and dimension factors\.
5. 5\.We prove a matching heavy\-tailed lower bound by reducing PL optimization to mean estimation for a sticky Markov chain\. This shows that both the Markovian effective\-sample\-size dependence and the finite\-moment exponent2​\(p−1\)/p2\(p\-1\)/pare unavoidable up to logarithmic and dimension factors; the result does not claim optimality of the dimension dependence\.

## 2Related work

We summarize the closest comparisons in[Table˜1](https://arxiv.org/html/2606.26316#S2.T1)\. The table separates objective or problem\-class assumptions from oracle/noise assumptions: PL is an objective\-geometry condition, whereas the ABC envelope is an oracle growth condition\. The main comparison is with high\-probability PL\-SGD under Markovian sampling; broader background on stochastic approximation, Poisson\-equation methods, Markov\-chain concentration, decentralized and MCMC sampling, and robust heavy\-tailed optimization is deferred to Appendix[A](https://arxiv.org/html/2606.26316#A1)\.

Table 1:Compact comparison with the closest Markovian stochastic\-optimization results\. Here Ach\. denotes an upper\-bound or algorithmic achievability result, LB denotes a lower bound for the corresponding problem model, HP denotes high probability, and Exp\. denotes expectation\. The notationO~​\(⋅\)\\widetilde\{O\}\(\\cdot\)hides logarithmic factors and problem\-dependent constants\. Rows differ in objective class, oracle model, and algorithmic model, so the rates are not meant as direct one\-to\-one comparisons\.*Abbreviations:*C = convex, SC = strongly convex, NC = nonconvex, VI = variational inequality, SA = stochastic approximation, Mark\. = Markovian\. The LB column records lower bounds in the corresponding optimization/oracle\-complexity model; qualitative sharpness examples for broader stochastic\-approximation tail behavior are not counted as PL optimization lower bounds\.

#### Closest PL\-SGD comparison\.

The closest prior result is Kar et al\.\[[33](https://arxiv.org/html/2606.26316#bib.bib34)\], which studies ordinary one\-sample SGD for smooth PL objectives with Markovian plus martingale\-difference noise under an ABC oracle envelope\. Their expectation bound has leading stochastic orderO~​\(tmix/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/k\), while their uniform high\-probability bound has leading stochastic orderO~​\(tmix2/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}^\{2\}/k\)\. Thus, before the present work, it was unclear whether the quadratic high\-probability dependence on the mixing time was intrinsic or an artifact of the proof technique\. Our light\-tailed theorem closes this gap by proving a uniform high\-probability bound with leading stochastic termO~​\(tmix/\(k\+K0\)\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/\(k\+K\_\{0\}\)\)for the same ordinary SGD recursion, under a general uniform total\-variation mixing profile\. The matching two\-state lower bound shows that the orderΩ​\(σ2​tmix/k\)\\Omega\(\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}/k\)is unavoidable at constant confidence, up to logarithmic factors\.

#### Other Markovian first\-order methods\.

Finite\-time analyses of stochastic gradient methods with Markovian samples were developed in convex, strongly convex, and nonconvex settings by Sun, Sun, and Yin\[[60](https://arxiv.org/html/2606.26316#bib.bib61)\], Doan et al\.\[[17](https://arxiv.org/html/2606.26316#bib.bib20)\], and Doan\[[18](https://arxiv.org/html/2606.26316#bib.bib21)\]\. Even\[[21](https://arxiv.org/html/2606.26316#bib.bib24)\]studies general Markovian sampling schemes, including MC\-SGD and the variance\-reduced MC\-SAG method, and proves lower bounds involving hitting\-time quantities\. Beznosikov et al\.\[[6](https://arxiv.org/html/2606.26316#bib.bib9)\]obtain linear mixing\-time dependence for several first\-order methods using randomized batching and multilevel gradient estimators, including nonconvex, strongly convex, and variational\-inequality settings\. These works show that favorable mixing dependence is possible in several regimes, but they do not resolve the high\-probability behavior of ordinary one\-sample PL\-SGD under an ABC oracle envelope\.

#### Heavy\-tailed Markovian regimes\.

The heavy\-tailed part of the paper is related to robust stochastic optimization and Markovian stochastic approximation under finite\-moment noise\. Agrawal, Maguluri, and Zubeldia\[[2](https://arxiv.org/html/2606.26316#bib.bib6)\]study concentration and tail behavior for general stochastic approximation with finite\-state Markovian components and heavy\-tailed martingale perturbations\. Our robust result is different in two ways: it is a PL optimization theorem, and the Markovian gradient component itself is allowed to have only a finite stationarypp\-th moment,p∈\(1,2\]p\\in\(1,2\]\. Using clipped Markovian blocks and a total transition budgetTT, we obtain the high\-probability rate

O~​\(σp2​\(tmixT\)2​\(p−1\)/p\),\\widetilde\{O\}\\\!\\left\(\\sigma\_\{p\}^\{2\}\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{T\}\\right\)^\{2\(p\-1\)/p\}\\right\),and prove a matching sticky\-chain lower bound for the Markovian effective\-sample\-size dependence and the finite\-moment exponent, up to logarithmic and dimension factors\.

#### Scope of the comparison\.

The rows in[Table˜1](https://arxiv.org/html/2606.26316#S2.T1)are not directly interchangeable\. Some works study convex or nonconvex optimization rather than PL objectives; some use modified estimators or variance\-reduced methods rather than ordinary SGD; and some prove lower bounds for hitting\-time or oracle\-complexity models rather than PL last\-iterate optimization\. The contribution of the present paper is to identify the optimal polynomial mixing\-time dependence for ordinary high\-probability PL\-SGD in the bounded\-ABC regime, and the optimal finite\-moment exponent and Markovian effective\-sample\-size dependence for the clipped\-block heavy\-tailed regime\.

## 3Problem formulation and assumptions

Letf:ℝd→ℝf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}be differentiable and lower bounded\. Write

f⋆:=infx∈ℝdf​\(x\),Δ​\(x\):=f​\(x\)−f⋆,Δk:=Δ​\(xk\)\.f^\{\\star\}:=\\inf\_\{x\\in\\mathbb\{R\}^\{d\}\}f\(x\),\\qquad\\Delta\(x\):=f\(x\)\-f^\{\\star\},\\qquad\\Delta\_\{k\}:=\\Delta\(x\_\{k\}\)\.The algorithm is

xk\+1=xk−αk​Gk,Gk=g​\(xk,Zk\)\+Mk\+1\.x\_\{k\+1\}=x\_\{k\}\-\\alpha\_\{k\}G\_\{k\},\\qquad G\_\{k\}=g\(x\_\{k\},Z\_\{k\}\)\+M\_\{k\+1\}\.\(3\)The initial pointx0x\_\{0\}is deterministic andΔ0<∞\\Delta\_\{0\}<\\infty\. Ifx0x\_\{0\}is random, the same statements hold conditionally onℱ0\\mathcal\{F\}\_\{0\}on every realization with finiteΔ0\\Delta\_\{0\}\.

###### Assumption 1\(Objective\)\.

The functionffisLL\-smooth and satisfies theμ\\mu\-PL inequality:

‖∇f​\(x\)‖2≥2​μ​Δ​\(x\),x∈ℝd\.\\\|\\nabla f\(x\)\\\|^\{2\}\\geq 2\\mu\\Delta\(x\),\\qquad x\\in\\mathbb\{R\}^\{d\}\.\(4\)Sinceffis lower bounded andLL\-smooth, the gradient upper bound

‖∇f​\(x\)‖2≤2​L​Δ​\(x\)\\\|\\nabla f\(x\)\\\|^\{2\}\\leq 2L\\Delta\(x\)\(5\)holds for allxx\.

###### Assumption 2\(Oracle, martingale noise, and ABC envelope\)\.

Let

ℱk:=σ​\(x0,Z0,…,Zk,M1,…,Mk\)\.\\mathcal\{F\}\_\{k\}:=\\sigma\(x\_\{0\},Z\_\{0\},\\ldots,Z\_\{k\},M\_\{1\},\\ldots,M\_\{k\}\)\.The sequence\(Mk\+1\)k≥0\(M\_\{k\+1\}\)\_\{k\\geq 0\}is a martingale difference sequence with respect to\(ℱk\)\(\\mathcal\{F\}\_\{k\}\):

𝔼​\[Mk\+1∣ℱk\]=0\.\\mathbb\{E\}\[M\_\{k\+1\}\\mid\\mathcal\{F\}\_\{k\}\]=0\.For everyx∈ℝdx\\in\\mathbb\{R\}^\{d\},

∫g​\(x,z\)​π​\(d​z\)=∇f​\(x\)\.\\int g\(x,z\)\\,\\pi\(dz\)=\\nabla f\(x\)\.\(6\)There are constantsA,B,C≥0A,B,C\\geq 0such that almost surely, for everykk,

‖Gk‖2≤A​‖∇f​\(xk\)‖2\+B​Δk\+C,\\\|G\_\{k\}\\\|^\{2\}\\leq A\\\|\\nabla f\(x\_\{k\}\)\\\|^\{2\}\+B\\Delta\_\{k\}\+C,\(7\)and, for everyx∈ℝdx\\in\\mathbb\{R\}^\{d\}andz∈𝖹z\\in\\mathsf\{Z\},

‖g​\(x,z\)‖2≤A​‖∇f​\(x\)‖2\+B​Δ​\(x\)\+C\.\\\|g\(x,z\)\\\|^\{2\}\\leq A\\\|\\nabla f\(x\)\\\|^\{2\}\+B\\Delta\(x\)\+C\.\(8\)Moreoverg​\(⋅,z\)g\(\\cdot,z\)isLgL\_\{g\}\-Lipschitz for everyzz:

‖g​\(x,z\)−g​\(y,z\)‖≤Lg​‖x−y‖\.\\\|g\(x,z\)\-g\(y,z\)\\\|\\leq L\_\{g\}\\\|x\-y\\\|\.\(9\)

###### Assumption 3\(Exogenous Markov chain and uniform mixing profile\)\.

The process\(Zk\)k≥0\(Z\_\{k\}\)\_\{k\\geq 0\}is a time\-homogeneous Markov chain on a measurable space\(𝖹,𝒵\)\(\\mathsf\{Z\},\\mathcal\{Z\}\)with transition kernelPPand invariant lawπ\\pi\. It is Markov with respect to the full filtration:

ℙ​\(Zk\+1∈B∣ℱk\)=P​\(Zk,B\),B∈𝒵\.\\mathbb\{P\}\(Z\_\{k\+1\}\\in B\\mid\\mathcal\{F\}\_\{k\}\)=P\(Z\_\{k\},B\),\\qquad B\\in\\mathcal\{Z\}\.\(10\)Define the uniform total\-variation mixing profile

ε​\(m\):=supz∈𝖹‖Pm​\(z,⋅\)−π‖TV,m≥0\.\\varepsilon\(m\):=\\sup\_\{z\\in\\mathsf\{Z\}\}\\\|P^\{m\}\(z,\\cdot\)\-\\pi\\\|\_\{\\mathrm\{TV\}\},\\qquad m\\geq 0\.\(11\)We assumeε​\(m\)→0\\varepsilon\(m\)\\to 0asm→∞m\\to\\infty\. Whenever a single geometric mixing time is used, we say thattmix≥1t\_\{\\mathrm\{mix\}\}\\geq 1is valid if

ε​\(m\)≤2−⌊m/tmix⌋,m≥0\.\\varepsilon\(m\)\\leq 2^\{\-\\lfloor m/t\_\{\\mathrm\{mix\}\}\\rfloor\},\\qquad m\\geq 0\.\(12\)

We use the stepsize

αk=ak\+K0,a≥2μ\.\\alpha\_\{k\}=\\frac\{a\}\{k\+K\_\{0\}\},\\qquad a\\geq\\frac\{2\}\{\\mu\}\.\(13\)Set

K0≥Kbase:=max⁡\{1,μ​a,a​L2​\(2​A\+Bμ\)\}\.K\_\{0\}\\geq K\_\{\\mathrm\{base\}\}:=\\max\\left\\\{1,\\mu a,\\frac\{aL\}\{2\}\\left\(2A\+\\frac\{B\}\{\\mu\}\\right\)\\right\\\}\.\(14\)For0≤ℓ≤k−10\\leq\\ell\\leq k\-1, define

ζℓ\+1,k−1:=∏j=ℓ\+1k−1\(1−μ​αj\),wℓ,k:=αℓ​ζℓ\+1,k−1,\\zeta\_\{\\ell\+1,k\-1\}:=\\prod\_\{j=\\ell\+1\}^\{k\-1\}\(1\-\\mu\\alpha\_\{j\}\),\\qquad w\_\{\\ell,k\}:=\\alpha\_\{\\ell\}\\zeta\_\{\\ell\+1,k\-1\},\(15\)with empty products equal to one\. We also define

ξ​\(x,z\):=g​\(x,z\)−∇f​\(x\),h​\(x,z\):=⟨∇f​\(x\),ξ​\(x,z\)⟩\.\\xi\(x,z\):=g\(x,z\)\-\\nabla f\(x\),\\qquad h\(x,z\):=\\left\\langle\\nabla f\(x\),\\xi\(x,z\)\\right\\rangle\.\(16\)Then∫h​\(x,z\)​π​\(d​z\)=0\\int h\(x,z\)\\pi\(dz\)=0for each fixedxx\.

#### Notation\.

For probability measuresν\\nuandν′\\nu^\{\\prime\}on\(𝖹,𝒵\)\(\\mathsf\{Z\},\\mathcal\{Z\}\), we use the probability total\-variation distance

‖ν−ν′‖TV:=supA∈𝒵\|ν​\(A\)−ν′​\(A\)\|=12​\|ν−ν′\|​\(𝖹\)\.\\\|\\nu\-\\nu^\{\\prime\}\\\|\_\{\\mathrm\{TV\}\}:=\\sup\_\{A\\in\\mathcal\{Z\}\}\|\\nu\(A\)\-\\nu^\{\\prime\}\(A\)\|=\\frac\{1\}\{2\}\|\\nu\-\\nu^\{\\prime\}\|\(\\mathsf\{Z\}\)\.Thus the largest possible distance between two probability laws is one\. With this convention, for every bounded measurable scalar functionφ\\varphi,

\|∫φ​d​\(ν−ν′\)\|≤2​‖φ‖∞​‖ν−ν′‖TV\.\\left\|\\int\\varphi\\,d\(\\nu\-\\nu^\{\\prime\}\)\\right\|\\leq 2\\\|\\varphi\\\|\_\{\\infty\}\\\|\\nu\-\\nu^\{\\prime\}\\\|\_\{\\mathrm\{TV\}\}\.\(17\)All norms on vectors are Euclidean\.

## 4Sharp upper and lower bounds

### 4\.1Upper bound for a general mixing profile

The main theorem is stated for a general mixing profile\. Fixδ∈\(0,e−1\)\\delta\\in\(0,e^\{\-1\}\)and define, fork≥1k\\geq 1,

Nk:=k\+K0,m^k:=min⁡\{m∈ℕ:m≥1,ε​\(m\)≤δ32​Nk4\},mk:=min⁡\{k,m^k\}\.N\_\{k\}:=k\+K\_\{0\},\\qquad\\widehat\{m\}\_\{k\}:=\\min\\left\\\{m\\in\\mathbb\{N\}:m\\geq 1,\\ \\varepsilon\(m\)\\leq\\frac\{\\delta\}\{32N\_\{k\}^\{4\}\}\\right\\\},\\qquad m\_\{k\}:=\\min\\\{k,\\widehat\{m\}\_\{k\}\\\}\.\(18\)The minimum is finite by Assumption[3](https://arxiv.org/html/2606.26316#Thmassumption3)\. The profileε​\(m\)\\varepsilon\(m\)is nonincreasing in integermm: indeed, Markov kernels contract total variation andPm\+1​\(z,⋅\)−π=\(Pm​\(z,⋅\)−π\)​PP^\{m\+1\}\(z,\\cdot\)\-\\pi=\(P^\{m\}\(z,\\cdot\)\-\\pi\)P\. Since the thresholdδ/\(32​Nk4\)\\delta/\(32N\_\{k\}^\{4\}\)is nonincreasing inkk, the sequencesm^k\\widehat\{m\}\_\{k\}andmkm\_\{k\}are nondecreasing\. Define

uk:=log⁡\(16​\(mk\+1\)​\(k\+1\)2δ\),qk:=1\+mkK0,Θk:=mk​uk​qk2\.u\_\{k\}:=\\log\\left\(\\frac\{16\(m\_\{k\}\+1\)\(k\+1\)^\{2\}\}\{\\delta\}\\right\),\\qquad q\_\{k\}:=1\+\\frac\{m\_\{k\}\}\{K\_\{0\}\},\\qquad\\Theta\_\{k\}:=m\_\{k\}u\_\{k\}q\_\{k\}^\{2\}\.\(19\)Consequentlyuk,qk,u\_\{k\},q\_\{k\},andΘk\\Theta\_\{k\}are also nondecreasing inkk\.

###### Theorem 1\(Mixing\-profile high\-probability upper bound\)\.

Suppose Assumptions[1](https://arxiv.org/html/2606.26316#Thmassumption1)–[3](https://arxiv.org/html/2606.26316#Thmassumption3)hold and let\(xk\)\(x\_\{k\}\)be generated by \([3](https://arxiv.org/html/2606.26316#S3.E3)\) with stepsizes \([13](https://arxiv.org/html/2606.26316#S3.E13)\)\. There exist constants

ρ0∈\(0,1\),KD<∞,C⋆<∞,\\rho\_\{0\}\\in\(0,1\),\\qquad K\_\{D\}<\\infty,\\qquad C\_\{\\star\}<\\infty,depending only ona,μ,L,Lg,A,B,C,Δ0a,\\mu,L,L\_\{g\},A,B,C,\\Delta\_\{0\}, such that the following holds\. AssumeK0≥Kbase∨KDK\_\{0\}\\geq K\_\{\\mathrm\{base\}\}\\vee K\_\{D\}and that the analysis windows satisfy, for everyk≥1k\\geq 1,

ΘkNk\\displaystyle\\frac\{\\Theta\_\{k\}\}\{N\_\{k\}\}≤ρ0,\\displaystyle\\leq\\rho\_\{0\},\(20\)mkNk​\(1\+log⁡NkK0\+mkK0\)\\displaystyle\\frac\{m\_\{k\}\}\{N\_\{k\}\}\\left\(1\+\\log\\frac\{N\_\{k\}\}\{K\_\{0\}\}\+\\frac\{m\_\{k\}\}\{K\_\{0\}\}\\right\)≤ρ0\.\\displaystyle\\leq\\rho\_\{0\}\.\(21\)Then, with probability at least1−δ1\-\\delta, for allk≥1k\\geq 1,

Δk≤2​K0​Δ0\+C⋆​\(1\+Θk\)k\+K0\.\\Delta\_\{k\}\\leq\\frac\{2K\_\{0\}\\Delta\_\{0\}\+C\_\{\\star\}\(1\+\\Theta\_\{k\}\)\}\{k\+K\_\{0\}\}\.\(22\)

The proof is given in Appendix[D](https://arxiv.org/html/2606.26316#A4)\.

###### Corollary 1\(Geometric mixing rate\)\.

Suppose, in addition to the assumptions of[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1), that the chain has a valid geometric mixing timetmixt\_\{\\mathrm\{mix\}\}in the sense of \([12](https://arxiv.org/html/2606.26316#S3.E12)\)\. There existsKgeo<∞K\_\{\\mathrm\{geo\}\}<\\infty, depending only ona,μ,L,Lg,A,B,C,Δ0a,\\mu,L,L\_\{g\},A,B,C,\\Delta\_\{0\}, such that if

K0≥Kgeo​tmix​\(1\+log⁡e​tmixδ\)3,K\_\{0\}\\geq K\_\{\\mathrm\{geo\}\}\\,t\_\{\\mathrm\{mix\}\}\\left\(1\+\\log\\frac\{et\_\{\\mathrm\{mix\}\}\}\{\\delta\}\\right\)^\{3\},\(23\)then \([22](https://arxiv.org/html/2606.26316#S4.E22)\) holds with probability at least1−δ1\-\\delta\. More explicitly, for everyk≥1k\\geq 1,

mk\\displaystyle m\_\{k\}≤⌈tmix​\(1\+log2⁡\(32​\(k\+K0\)4δ\)\)⌉≤1\+tmix​\(1\+log2⁡\(32​\(k\+K0\)4δ\)\),\\displaystyle\\leq\\left\\lceil t\_\{\\mathrm\{mix\}\}\\left\(1\+\\log\_\{2\}\\left\(\\frac\{32\(k\+K\_\{0\}\)^\{4\}\}\{\\delta\}\\right\)\\right\)\\right\\rceil\\leq 1\+t\_\{\\mathrm\{mix\}\}\\left\(1\+\\log\_\{2\}\\left\(\\frac\{32\(k\+K\_\{0\}\)^\{4\}\}\{\\delta\}\\right\)\\right\),\(24\)Θk\\displaystyle\\Theta\_\{k\}≤Mk​log⁡\(16​\(k\+1\)2​\(Mk\+1\)δ\)​\(1\+MkK0\)2,\\displaystyle\\leq M\_\{k\}\\,\\log\\left\(\\frac\{16\(k\+1\)^\{2\}\(M\_\{k\}\+1\)\}\{\\delta\}\\right\)\\left\(1\+\\frac\{M\_\{k\}\}\{K\_\{0\}\}\\right\)^\{2\},\(25\)Mk\\displaystyle M\_\{k\}:=1\+tmix​\(1\+log2⁡\(32​\(k\+K0\)4δ\)\)\.\\displaystyle:=1\+t\_\{\\mathrm\{mix\}\}\\left\(1\+\\log\_\{2\}\\left\(\\frac\{32\(k\+K\_\{0\}\)^\{4\}\}\{\\delta\}\\right\)\\right\)\.Consequently, \([22](https://arxiv.org/html/2606.26316#S4.E22)\) gives a leading stochastic contribution of order

O~​\(tmixk\+K0\),\\widetilde\{O\}\\\!\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{k\+K\_\{0\}\}\\right\),where the hidden factors are logarithmic inkk,1/δ1/\\delta, andtmixt\_\{\\mathrm\{mix\}\}, and the prefactor is the problem\-dependent constant described in[Remark˜4](https://arxiv.org/html/2606.26316#Thmremark4); there is no additional polynomial power oftmixt\_\{\\mathrm\{mix\}\}\.

The proof is given in Appendix[E](https://arxiv.org/html/2606.26316#A5)\.

### 4\.2A matching lower bound

The upper bound has the optimal polynomial dependence on mixing time\. Let𝖹=\{−1,\+1\}\\mathsf\{Z\}=\\\{\-1,\+1\\\}and fixε∈\(0,1/8\]\\varepsilon\\in\(0,1/8\]\. Consider the persistent two\-state chain

P=\(1−εεε1−ε\),P=\\begin\{pmatrix\}1\-\\varepsilon&\\varepsilon\\\\ \\varepsilon&1\-\\varepsilon\\end\{pmatrix\},\(26\)whose invariant law is uniform\. Let

f​\(x\)=x22,g​\(x,z\)=x\+σ​z,σ\>0\.f\(x\)=\\frac\{x^\{2\}\}\{2\},\\qquad g\(x,z\)=x\+\\sigma z,\\qquad\\sigma\>0\.\(27\)WithMk\+1≡0M\_\{k\+1\}\\equiv 0, SGD becomes

xn\+1=\(1−an\+K\)​xn−a​σn\+K​Zn,x0=0\.x\_\{n\+1\}=\\left\(1\-\\frac\{a\}\{n\+K\}\\right\)x\_\{n\}\-\\frac\{a\\sigma\}\{n\+K\}Z\_\{n\},\\qquad x\_\{0\}=0\.\(28\)The chain is started in stationarity\.

###### Theorem 2\(Linear mixing\-time lower bound\)\.

Fixa≥2a\\geq 2\. There are positive constantscac\_\{a\}andc0c\_\{0\}depending at most onaasuch that the following holds\. LetK≥max⁡\{4​a,4​a2\}K\\geq\\max\\\{4a,4a^\{2\}\\\}and letε∈\(0,1/8\]\\varepsilon\\in\(0,1/8\]\. For the instance \([26](https://arxiv.org/html/2606.26316#S4.E26)\)\-\([28](https://arxiv.org/html/2606.26316#S4.E28)\), if

k≥max⁡\{8​K,16ε,16​a2\},k\\geq\\max\\left\\\{8K,\\frac\{16\}\{\\varepsilon\},16a^\{2\}\\right\\\},\(29\)then

𝔼​\[f​\(xk\)−f⋆\]≥ca​σ2ε​k,\\mathbb\{E\}\\left\[f\(x\_\{k\}\)\-f^\{\\star\}\\right\]\\geq c\_\{a\}\\,\\frac\{\\sigma^\{2\}\}\{\\varepsilon k\},\(30\)and

ℙ​\(f​\(xk\)−f⋆≥ca​σ2ε​k\)≥c0\.\\mathbb\{P\}\\\!\\left\(f\(x\_\{k\}\)\-f^\{\\star\}\\geq c\_\{a\}\\,\\frac\{\\sigma^\{2\}\}\{\\varepsilon k\}\\right\)\\geq c\_\{0\}\.\(31\)One may takec0=1/96c\_\{0\}=1/96and

ca=a2​e−2214​32​a\+2\.c\_\{a\}=\\frac\{a^\{2\}e^\{\-2\}\}\{2^\{14\}\\,3^\{2a\+2\}\}\.\(32\)The mixing time of the chain isΘ​\(1/ε\)\\Theta\(1/\\varepsilon\)\. Consequently, for another constantca′\>0c^\{\\prime\}\_\{a\}\>0depending only onaa,

𝔼​\[f​\(xk\)−f⋆\]≥ca′​σ2​tmixk,ℙ​\(f​\(xk\)−f⋆≥ca′​σ2​tmixk\)≥c0\.\\mathbb\{E\}\\left\[f\(x\_\{k\}\)\-f^\{\\star\}\\right\]\\geq c^\{\\prime\}\_\{a\}\\,\\frac\{\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}\}\{k\},\\qquad\\mathbb\{P\}\\\!\\left\(f\(x\_\{k\}\)\-f^\{\\star\}\\geq c^\{\\prime\}\_\{a\}\\,\\frac\{\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}\}\{k\}\\right\)\\geq c\_\{0\}\.\(33\)

The proof is given in Appendix[G](https://arxiv.org/html/2606.26316#A7); the final moment\-to\-probability argument is in Appendix[G\.4](https://arxiv.org/html/2606.26316#A7.SS4)\.

###### Corollary 2\(No sublinear order intmix/kt\_\{\\mathrm\{mix\}\}/kat constant confidence\)\.

Any high\-probability theorem that is valid for all instances satisfying[Assumptions˜1](https://arxiv.org/html/2606.26316#Thmassumption1),[2](https://arxiv.org/html/2606.26316#Thmassumption2)and[3](https://arxiv.org/html/2606.26316#Thmassumption3)cannot have leading stochastic ordero​\(tmix/k\)o\(t\_\{\\mathrm\{mix\}\}/k\)at constant confidence\. More precisely, if the failure probability is smaller than the constantc0c\_\{0\}in[Theorem˜2](https://arxiv.org/html/2606.26316#Thmtheorem2), then the right\-hand side of a uniform upper bound cannot beo​\(tmix/k\)o\(t\_\{\\mathrm\{mix\}\}/k\)on the family \([26](https://arxiv.org/html/2606.26316#S4.E26)\)–\([28](https://arxiv.org/html/2606.26316#S4.E28)\)\.

###### Proof\.

This follows directly from[Theorem˜2](https://arxiv.org/html/2606.26316#Thmtheorem2)\. The lower\-bound family \([26](https://arxiv.org/html/2606.26316#S4.E26)\)–\([28](https://arxiv.org/html/2606.26316#S4.E28)\) satisfies[Assumptions˜1](https://arxiv.org/html/2606.26316#Thmassumption1),[2](https://arxiv.org/html/2606.26316#Thmassumption2)and[3](https://arxiv.org/html/2606.26316#Thmassumption3)\. On this family,[Theorem˜2](https://arxiv.org/html/2606.26316#Thmtheorem2)shows that, with probability at leastc0c\_\{0\},

f​\(xk\)−f⋆≥ca′​σ2​tmixk\.f\(x\_\{k\}\)\-f^\{\\star\}\\geq c^\{\\prime\}\_\{a\}\\frac\{\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}\}\{k\}\.Therefore any uniform high\-probability upper bound over the same class, with failure probability strictly smaller thanc0c\_\{0\}, cannot have a leading stochastic termo​\(tmix/k\)o\(t\_\{\\mathrm\{mix\}\}/k\), since such a bound would eventually be smaller than the lower\-bound threshold on this family, contradicting[Theorem˜2](https://arxiv.org/html/2606.26316#Thmtheorem2)\. ∎

## 5Heavy\-tailed Markovian gradients

The bounded\-ABC theorem is sharp for light\-tailed or almost\-sure controlled gradients\. We now treat a genuinely heavy\-tailed regime in which the Markovian gradient component has only a finiteppth moment under the invariant law\. Ordinary one\-sample SGD is not stable enough to provide logarithmic\-confidence high\-probability guarantees under this assumption alone\. The robust method below therefore uses clipping and blockwise averaging\. During each block the iterate is held fixed, the next consecutive Markovian gradients are all used, and the update is made with the average of the clipped samples\. The mixing window is a design and proof parameter used to choose the clipping scale and to analyze the consecutive block; it is not a spacing rule and no observations inside the block are discarded\.

Forλ\>0\\lambda\>0, define the coordinatewise clipping map

\[𝖳λ​\(v\)\]j:=max⁡\{−λ,min⁡\{vj,λ\}\},j=1,…,d\.\[\\mathsf\{T\}\_\{\\lambda\}\(v\)\]\_\{j\}:=\\max\\\{\-\\lambda,\\min\\\{v\_\{j\},\\lambda\\\}\\\},\\qquad j=1,\\ldots,d\.\(34\)The heavy\-tailed oracle is written as

Y​\(x,z,u\)=𝒢​\(x,z,u\),Y\(x,z,u\)=\{\\mathcal\{G\}\}\(x,z,u\),\(35\)whereuudenotes auxiliary randomness conditionally independent of the past given the current iterate and Markov state\.

###### Assumption 4\(Stationary finite\-moment Markovian oracle\)\.

Fixp∈\(1,2\]p\\in\(1,2\]\. At each oracle call, the auxiliary variableUUis sampled conditionally independently of the past given the queried point and current Markov state, with the same conditional law used in \([36](https://arxiv.org/html/2606.26316#S5.E36)\) and \([38](https://arxiv.org/html/2606.26316#S5.E38)\)\. The oracle satisfies the stationary unbiasedness condition

∫𝔼U​\[𝒢​\(x,z,U\)\]​π​\(d​z\)=∇f​\(x\),x∈ℝd\.\\int\\mathbb\{E\}\_\{U\}\[\{\\mathcal\{G\}\}\(x,z,U\)\]\\,\\pi\(dz\)=\\nabla f\(x\),\\qquad x\\in\\mathbb\{R\}^\{d\}\.\(36\)Letη​\(x,z,u\):=𝒢​\(x,z,u\)−∇f​\(x\)\\eta\(x,z,u\):=\{\\mathcal\{G\}\}\(x,z,u\)\-\\nabla f\(x\)and

Rp​\(x\):=1\+‖∇f​\(x\)‖\+Δ​\(x\)\.R\_\{p\}\(x\):=1\+\\\|\\nabla f\(x\)\\\|\+\\sqrt\{\\Delta\(x\)\}\.\(37\)There is a scaleσp≥1\\sigma\_\{p\}\\geq 1such that, for everyx∈ℝdx\\in\\mathbb\{R\}^\{d\},

∫𝔼U​\[‖η​\(x,z,U\)‖p\]​π​\(d​z\)≤σpp​Rp​\(x\)p\.\\int\\mathbb\{E\}\_\{U\}\\left\[\\\|\\eta\(x,z,U\)\\\|^\{p\}\\right\]\\pi\(dz\)\\leq\\sigma\_\{p\}^\{p\}R\_\{p\}\(x\)^\{p\}\.\(38\)

###### Definition 1\(All\-samples clipped block method\)\.

Fix integersN,b≥1N,b\\geq 1and a clipping levelλ\\lambda\. Setτ0=0\\tau\_\{0\}=0\. At outer iterationr=0,…,N−1r=0,\\ldots,N\-1, holdxrx\_\{r\}fixed and observe the nextbbconsecutive Markovian gradients\. Specifically, fori=1,…,bi=1,\\ldots,b, after one Markov transition from timeτr\+i−1\\tau\_\{r\}\+i\-1toτr\+i\\tau\_\{r\}\+i, drawUr,iU\_\{r,i\}and observe

Yr,i:=𝒢​\(xr,Zτr\+i,Ur,i\)\.Y\_\{r,i\}:=\{\\mathcal\{G\}\}\(x\_\{r\},Z\_\{\\tau\_\{r\}\+i\},U\_\{r,i\}\)\.After the block is collected, setτr\+1:=τr\+b\\tau\_\{r\+1\}:=\\tau\_\{r\}\+band update

g^r:=1b​∑i=1b𝖳λ​\(Yr,i\),xr\+1=xr−14​L​g^r\.\\widehat\{g\}\_\{r\}:=\\frac\{1\}\{b\}\\sum\_\{i=1\}^\{b\}\\mathsf\{T\}\_\{\\lambda\}\(Y\_\{r,i\}\),\\qquad x\_\{r\+1\}=x\_\{r\}\-\\frac\{1\}\{4L\}\\widehat\{g\}\_\{r\}\.\(39\)Thus each Markov transition in the block contributes exactly one clipped gradient sample to the update\.

Let

ϑp:=2​\(p−1\)p\.\\vartheta\_\{p\}:=\\frac\{2\(p\-1\)\}\{p\}\.\(40\)For a radiusℛ≥1\\mathcal\{R\}\\geq 1, define

Bℛ:=1\+2​L​ℛ\+ℛ\.B\_\{\\mathcal\{R\}\}:=1\+\\sqrt\{2L\\mathcal\{R\}\}\+\\sqrt\{\\mathcal\{R\}\}\.\(41\)For confidence levelδ∈\(0,e−1\)\\delta\\in\(0,e^\{\-1\}\)and an integer analysis lagm∈\{1,…,b\}m\\in\\\{1,\\ldots,b\\\}, set

uht:=log⁡\(64​d​N​\(m\+1\)δ\),sht:=m​uhtb,u\_\{\\rm ht\}:=\\log\\left\(\\frac\{64dN\(m\+1\)\}\{\\delta\}\\right\),\\qquad s\_\{\\rm ht\}:=\\frac\{mu\_\{\\rm ht\}\}\{b\},\(42\)and choose

λ:=cλ,p​σp​Bℛ​sht−1/p,\\lambda:=c\_\{\\lambda,p\}\\,\\sigma\_\{p\}B\_\{\\mathcal\{R\}\}\\,s\_\{\\rm ht\}^\{\-1/p\},\(43\)wherecλ,pc\_\{\\lambda,p\}is a sufficiently large numerical constant depending only onpp\. The following theorem is stated with explicit admissibility conditions\. They require the block to be long enough for the clipping bias, sampling fluctuation, and residual mixing bias to be below the radius used in the induction\.

###### Theorem 3\(Heavy\-tailed robust Markovian PL optimization\)\.

Suppose Assumptions[1](https://arxiv.org/html/2606.26316#Thmassumption1),[3](https://arxiv.org/html/2606.26316#Thmassumption3), and[4](https://arxiv.org/html/2606.26316#Thmassumption4)hold\. Let\(xr\)r=0N\(x\_\{r\}\)\_\{r=0\}^\{N\}be generated by[Definition˜1](https://arxiv.org/html/2606.26316#Thmdefinition1)with clipping level \([43](https://arxiv.org/html/2606.26316#S5.E43)\)\. Letℛ≥4​Δ0\+4\\mathcal\{R\}\\geq 4\\Delta\_\{0\}\+4and assume

ε​\(m\)\\displaystyle\\varepsilon\(m\)≤δ64​d​N​b,\\displaystyle\\leq\\frac\{\\delta\}\{64dNb\},\(44\)λ\\displaystyle\\lambda≥2​2​L​ℛ,\\displaystyle\\geq 2\\sqrt\{2L\\mathcal\{R\}\},\(45\)cp​d​σp2​Bℛ2μ​shtϑp\\displaystyle\\frac\{c\_\{p\}d\\sigma\_\{p\}^\{2\}B\_\{\\mathcal\{R\}\}^\{2\}\}\{\\mu\}s\_\{\\rm ht\}^\{\\vartheta\_\{p\}\}≤ℛ4,\\displaystyle\\leq\\frac\{\\mathcal\{R\}\}\{4\},\(46\)wherecpc\_\{p\}is a numerical constant depending only onpp\. Then, with probability at least1−δ1\-\\delta,

Δ​\(xN\)≤\(1−μ16​L\)N​Δ0\+cp​d​σp2​Bℛ2μ​shtϑp\.\\Delta\(x\_\{N\}\)\\leq\\left\(1\-\\frac\{\\mu\}\{16L\}\\right\)^\{N\}\\Delta\_\{0\}\+\\frac\{c\_\{p\}d\\sigma\_\{p\}^\{2\}B\_\{\\mathcal\{R\}\}^\{2\}\}\{\\mu\}s\_\{\\rm ht\}^\{\\vartheta\_\{p\}\}\.\(47\)Moreover, on the same event all outer iterates satisfyΔ​\(xr\)≤ℛ\\Delta\(x\_\{r\}\)\\leq\\mathcal\{R\}for0≤r≤N0\\leq r\\leq N\.

The proof is given in Appendix[H\.2](https://arxiv.org/html/2606.26316#A8.SS2)\.

###### Corollary 3\(Geometric mixing and total transition budget\)\.

Suppose, in addition, that \([12](https://arxiv.org/html/2606.26316#S3.E12)\) holds with a valid mixing timetmixt\_\{\\mathrm\{mix\}\}\. For a total Markov\-transition budgetTT, choose

N=⌈32​Lμ​log⁡\(e​T\)⌉,b=⌊TN⌋,m=⌈tmix​\(1\+log2⁡\(64​d​N​Tδ\)\)⌉\.N=\\left\\lceil\\frac\{32L\}\{\\mu\}\\log\(eT\)\\right\\rceil,\\qquad b=\\left\\lfloor\\frac\{T\}\{N\}\\right\\rfloor,\\qquad m=\\left\\lceil t\_\{\\mathrm\{mix\}\}\\left\(1\+\\log\_\{2\}\\left\(\\frac\{64dNT\}\{\\delta\}\\right\)\\right\)\\right\\rceil\.\(48\)If1≤m≤b1\\leq m\\leq band the admissibility conditions in[Theorem˜3](https://arxiv.org/html/2606.26316#Thmtheorem3)hold for a radiusℛ≥4​Δ0\+4\\mathcal\{R\}\\geq 4\\Delta\_\{0\}\+4, then the robust method uses at mostTTMarkov transitions and, with probability at least1−δ1\-\\delta,

Δ​\(xN\)≤O~​\(Δ0T\)\+O~​\(d​σp2​Bℛ2μ​\(tmixT\)2​\(p−1\)p\),\\Delta\(x\_\{N\}\)\\leq\\widetilde\{O\}\\left\(\\frac\{\\Delta\_\{0\}\}\{T\}\\right\)\+\\widetilde\{O\}\\left\(\\frac\{d\\sigma\_\{p\}^\{2\}B\_\{\\mathcal\{R\}\}^\{2\}\}\{\\mu\}\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{T\}\\right\)^\{\\frac\{2\(p\-1\)\}\{p\}\}\\right\),\(49\)where the hidden factors are polylogarithmic inTT,1/δ1/\\delta,dd,L/μL/\\mu, andtmixt\_\{\\mathrm\{mix\}\}, and contain no additional polynomial power oftmixt\_\{\\mathrm\{mix\}\}\.

The proof is given in Appendix[H\.3](https://arxiv.org/html/2606.26316#A8.SS3)\.

The exponent in[Corollary˜3](https://arxiv.org/html/2606.26316#Thmcorollary3)and the dependence on the Markovian effective sample size cannot be improved in general\. The lower bound below reduces PL optimization to estimating the mean of a heavy\-tailed sticky Markov chain\.

###### Theorem 4\(Heavy\-tail and mixing lower bound\)\.

Fixp∈\(1,2\]p\\in\(1,2\],τ≥2\\tau\\geq 2, andσp≥1\\sigma\_\{p\}\\geq 1\. There are universal constantsc,c0\>0c,c\_\{0\}\>0such that the following holds for everyn≥8​τn\\geq 8\\tau\. Consider all algorithms which, after observingnnoracle outputs generated duringnnMarkov transitions starting from the fixed initial stateZ0=0Z\_\{0\}=0, output an estimatex^n\\widehat\{x\}\_\{n\}\. There exists a one\-dimensional quadratic PL objective

fθ​\(x\)=12​\(x−θ\)2f\_\{\\theta\}\(x\)=\\frac\{1\}\{2\}\(x\-\\theta\)^\{2\}and a Markovian oracle satisfying Assumption[4](https://arxiv.org/html/2606.26316#Thmassumption4)with a valid geometric mixing time at mostc​τc\\tauand moment scale at most a constant multiple ofσp\\sigma\_\{p\}, such that

ℙ​\(fθ​\(x^n\)−fθ⋆≥c​σp2​\(τn\)2​\(p−1\)p\)≥c0\.\\mathbb\{P\}\\left\(f\_\{\\theta\}\(\\widehat\{x\}\_\{n\}\)\-f\_\{\\theta\}^\{\\star\}\\geq c\\,\\sigma\_\{p\}^\{2\}\\left\(\\frac\{\\tau\}\{n\}\\right\)^\{\\frac\{2\(p\-1\)\}\{p\}\}\\right\)\\geq c\_\{0\}\.\(50\)Hence the heavy\-tail exponent and the polynomial dependence on the number of effective Markovian samples in[Corollary˜3](https://arxiv.org/html/2606.26316#Thmcorollary3)are unavoidable up to logarithmic factors\. Since the construction is one\-dimensional, this lower bound does not address the dimension factor in the upper bound\.

The proof is given in Appendix[H\.4](https://arxiv.org/html/2606.26316#A8.SS4)\.

## 6Relaxing global assumptions

The main theorem already removes compactness of the Markov state space and replaces a single geometric mixing\-time parameter by a general mixing profile\. The remaining global assumptions can also be localized in the standard way\. We record a formal version because it is useful in overparameterized models, where PL and oracle Lipschitzness often hold only near the initialization\.

ForR\>0R\>0, write

𝒮R:=\{x∈ℝd:Δ​\(x\)≤R\}\.\\mathcal\{S\}\_\{R\}:=\\\{x\\in\\mathbb\{R\}^\{d\}:\\Delta\(x\)\\leq R\\\}\.Assume thatffis globallyLL\-smooth, while the PL inequality \([4](https://arxiv.org/html/2606.26316#S3.E4)\), the stationary\-gradient identity \([6](https://arxiv.org/html/2606.26316#S3.E6)\), the pointwise oracle envelope \([8](https://arxiv.org/html/2606.26316#S3.E8)\), and the Lipschitz condition \([9](https://arxiv.org/html/2606.26316#S3.E9)\) hold only forx,y∈𝒮Rx,y\\in\\mathcal\{S\}\_\{R\}\. The pathwise ABC bound \([7](https://arxiv.org/html/2606.26316#S3.E7)\) is required only at times for whichxk∈𝒮Rx\_\{k\}\\in\\mathcal\{S\}\_\{R\}\.

###### Corollary 4\(Local PL and local oracle regularity\)\.

Let the local assumptions above hold on𝒮R\\mathcal\{S\}\_\{R\}\. Let the windows satisfy the admissibility conditions \([20](https://arxiv.org/html/2606.26316#S4.E20)\)\-\([21](https://arxiv.org/html/2606.26316#S4.E21)\), and letC⋆C\_\{\\star\}andρ0\\rho\_\{0\}be the constants from[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)computed using the local constants on𝒮R\\mathcal\{S\}\_\{R\}\. If

R≥2​Δ0\+C⋆​\(1\+ρ0\),R\\geq 2\\Delta\_\{0\}\+C\_\{\\star\}\(1\+\\rho\_\{0\}\),\(51\)then the conclusion of[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)holds\. In particular, with probability at least1−δ1\-\\delta, all iterates remain in𝒮R\\mathcal\{S\}\_\{R\}and satisfy \([22](https://arxiv.org/html/2606.26316#S4.E22)\) for everyk≥1k\\geq 1\.

The proof is given in Appendix[F](https://arxiv.org/html/2606.26316#A6)\.

## 7Proof roadmap

We summarize the proof architecture and point to the specific appendix results used in each step\. The paper has three proof mechanisms\. The first is a weighted PL descent recursion for ordinary SGD\. The second is the lag\-blocking argument, which converts the adaptive Markovian descent observable into residue\-class martingale differences plus controlled bias and replacement terms\. The third is the robust clipped\-block argument for finite\-ppheavy\-tailed Markovian gradients\. The appendices are organized so that each of these mechanisms is isolated in a sequence of lemmas before being combined in the corresponding theorem\.

#### Light\-tailed upper bound:[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)\.

The proof begins with the deterministic descent structure induced by PL geometry\. In[Lemma˜2](https://arxiv.org/html/2606.26316#Thmlemma2), smoothness from[Assumption˜1](https://arxiv.org/html/2606.26316#Thmassumption1)is applied to the SGD update

xk\+1=xk−αk​Gk,Gk=g​\(xk,Zk\)\+Mk\+1\.x\_\{k\+1\}=x\_\{k\}\-\\alpha\_\{k\}G\_\{k\},\\qquad G\_\{k\}=g\(x\_\{k\},Z\_\{k\}\)\+M\_\{k\+1\}\.The PL inequality in[Assumption˜1](https://arxiv.org/html/2606.26316#Thmassumption1)then turns the gradient\-norm term into a contraction in the suboptimalityΔk=f​\(xk\)−f⋆\\Delta\_\{k\}=f\(x\_\{k\}\)\-f^\{\\star\}\. The ABC envelope in[Assumption˜2](https://arxiv.org/html/2606.26316#Thmassumption2)is used at this stage to control the quadratic termαk2​‖Gk‖2\\alpha\_\{k\}^\{2\}\\\|G\_\{k\}\\\|^\{2\}\. This gives a one\-step inequality of the form

Δk\+1≤\(1−μ​αk\)​Δk\+controlled deterministic terms−αk​⟨∇f​\(xk\),Mk\+1⟩−αk​h​\(xk,Zk\)\.\\Delta\_\{k\+1\}\\leq\(1\-\\mu\\alpha\_\{k\}\)\\Delta\_\{k\}\+\\hbox\{controlled deterministic terms\}\-\\alpha\_\{k\}\\langle\\nabla f\(x\_\{k\}\),M\_\{k\+1\}\\rangle\-\\alpha\_\{k\}h\(x\_\{k\},Z\_\{k\}\)\.Iterating this recursion produces the weighted representation in[Lemma˜2](https://arxiv.org/html/2606.26316#Thmlemma2), with weightswℓ,kw\_\{\\ell,k\}\. The deterministic estimates on these weights are collected in[Lemma˜3](https://arxiv.org/html/2606.26316#Thmlemma3)\. These estimates are used repeatedly to show that weighted square sums have the correct order, for example

∑ℓ<kwℓ,k2≍1k\+K0\\sum\_\{\\ell<k\}w\_\{\\ell,k\}^\{2\}\\asymp\\frac\{1\}\{k\+K\_\{0\}\}up to constants and logarithmic factors\.

The stochastic terms in the weighted recursion are controlled on stopped good events\. The martingale\-difference perturbation

∑ℓ=0k−1wℓ,k​⟨∇f​\(xℓ\),Mℓ\+1⟩\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}\\big\\langle\\nabla f\(x\_\{\\ell\}\),M\_\{\\ell\+1\}\\big\\rangleis handled in[Lemma˜5](https://arxiv.org/html/2606.26316#Thmlemma5)\. The stopped event ensures thatΔℓ\\Delta\_\{\\ell\}remains inside the induction envelope, so the ABC envelope in[Assumption˜2](https://arxiv.org/html/2606.26316#Thmassumption2)gives deterministic bounds on the martingale increments\. A union bound over times then gives the uniform\-in\-time control needed in the final first\-failure argument\.

The Markovian descent term

∑ℓ=0k−1wℓ,k​h​\(xℓ,Zℓ\),h​\(x,z\)=⟨∇f​\(x\),g​\(x,z\)−∇f​\(x\)⟩,\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\),\\qquad h\(x,z\)=\\langle\\nabla f\(x\),g\(x,z\)\-\\nabla f\(x\)\\rangle,is the main difficulty\. It is not a martingale sum becausexℓx\_\{\\ell\}is generated from the past Markov trajectory, andZℓZ\_\{\\ell\}need not be stationary conditionally on the past\. For a fixed terminal timekk, the proof introduces the analytical lagm=mkm=m\_\{k\}\. The firstmmterms are kept separate and controlled by[Lemma˜9](https://arxiv.org/html/2606.26316#Thmlemma9)\. Forℓ≥m\\ell\\geq m, the decomposition

h​\(xℓ,Zℓ\)=h​\(xℓ−m,Zℓ\)\+\{h​\(xℓ,Zℓ\)−h​\(xℓ−m,Zℓ\)\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)=h\(x\_\{\\ell\-m\},Z\_\{\\ell\}\)\+\\bigl\\\{h\(x\_\{\\ell\},Z\_\{\\ell\}\)\-h\(x\_\{\\ell\-m\},Z\_\{\\ell\}\)\\bigr\\\}separates the Markovian term into a delayed part and a replacement error\.

The delayed part is where[Assumption˜3](https://arxiv.org/html/2606.26316#Thmassumption3)is used\. Sincexℓ−mx\_\{\\ell\-m\}isℱℓ−m\\mathcal\{F\}\_\{\\ell\-m\}\-measurable, the conditional law ofZℓZ\_\{\\ell\}givenℱℓ−m\\mathcal\{F\}\_\{\\ell\-m\}isPm​\(Zℓ−m,⋅\)P^\{m\}\(Z\_\{\\ell\-m\},\\cdot\)\. Thus the uniform total\-variation mixing profileε​\(m\)\\varepsilon\(m\)controls the conditional bias of the delayed observable\. After subtracting this conditional mean, the delayed sum is split into residue classes modulomm\. Along each residue class, consecutive indices are separated by exactlymmMarkov transitions, and the recentered delayed terms form a martingale\-difference sequence with respect to a down\-sampled filtration\. This is formalized in[Lemma˜6](https://arxiv.org/html/2606.26316#Thmlemma6)\. The key quantitative point is that the residue\-class martingale bounds involve

m​∑ℓ<kwℓ,k2,m\\sum\_\{\\ell<k\}w\_\{\\ell,k\}^\{2\},rather thanm2​∑ℓ<kwℓ,k2m^\{2\}\\sum\_\{\\ell<k\}w\_\{\\ell,k\}^\{2\}\. This single power ofmmis the source of the linear polynomial dependence on the mixing window\.

The remaining pieces of the Markovian term are controlled separately\. The conditional mixing bias is bounded in[Lemma˜7](https://arxiv.org/html/2606.26316#Thmlemma7)by combining[Assumption˜3](https://arxiv.org/html/2606.26316#Thmassumption3)with the envelope forhh\. The replacement error is controlled in[Lemma˜8](https://arxiv.org/html/2606.26316#Thmlemma8)\. This uses the Lipschitz condition in[Assumption˜2](https://arxiv.org/html/2606.26316#Thmassumption2), the envelope estimates forhhfrom[Lemma˜4](https://arxiv.org/html/2606.26316#Thmlemma4), and the pathwise movement bound

‖xℓ−xℓ−m‖≤∑i=ℓ−mℓ−1αi​‖Gi‖\.\\\|x\_\{\\ell\}\-x\_\{\\ell\-m\}\\\|\\leq\\sum\_\{i=\\ell\-m\}^\{\\ell\-1\}\\alpha\_\{i\}\\\|G\_\{i\}\\\|\.On the stopped event, the ABC envelope controls the gradients in this movement bound\. Thus the replacement term is deterministic after conditioning on the good event, and the admissibility condition \([21](https://arxiv.org/html/2606.26316#S4.E21)\) ensures that it is small enough to be absorbed into the induction envelope\.

The one\-step stopped induction estimate is[Lemma˜10](https://arxiv.org/html/2606.26316#Thmlemma10)\. This lemma combines[Lemma˜5](https://arxiv.org/html/2606.26316#Thmlemma5),[Lemma˜6](https://arxiv.org/html/2606.26316#Thmlemma6),[Lemma˜7](https://arxiv.org/html/2606.26316#Thmlemma7),[Lemma˜8](https://arxiv.org/html/2606.26316#Thmlemma8), and[Lemma˜9](https://arxiv.org/html/2606.26316#Thmlemma9)\. The first admissibility condition \([20](https://arxiv.org/html/2606.26316#S4.E20)\) controls the size of the stochastic concentration terms, while the second admissibility condition \([21](https://arxiv.org/html/2606.26316#S4.E21)\) controls the drift accumulated by replacingxℓx\_\{\\ell\}withxℓ−mx\_\{\\ell\-m\}\. The proof of[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1), given in Appendix[D](https://arxiv.org/html/2606.26316#A4), then applies a first\-failure argument: assume the induction envelope holds up to timek−1k\-1, apply[Lemma˜10](https://arxiv.org/html/2606.26316#Thmlemma10)at timekk, and show that the envelope cannot be violated on the high\-probability event\. A union bound over the failure probabilities gives the simultaneous guarantee \([22](https://arxiv.org/html/2606.26316#S4.E22)\) for allk≥1k\\geq 1\.

#### Geometric mixing corollary:[Corollary˜1](https://arxiv.org/html/2606.26316#Thmcorollary1)\.

[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)is stated for a general uniform total\-variation profileε​\(m\)\\varepsilon\(m\)\. Therefore it does not by itself choose a closed\-form delay\. Under the geometric condition \([12](https://arxiv.org/html/2606.26316#S3.E12)\), the delaymkm\_\{k\}is chosen so that the mixing bias is small compared with the weighted descent scale\. The elementary logarithm\-over\-linear estimate needed for this calculation is[Lemma˜11](https://arxiv.org/html/2606.26316#Thmlemma11), and the admissibility of the geometric windows is verified in[Lemma˜12](https://arxiv.org/html/2606.26316#Thmlemma12)\.

The role of[Lemma˜12](https://arxiv.org/html/2606.26316#Thmlemma12)is to check that the choices ofmkm\_\{k\},uku\_\{k\},qkq\_\{k\}, andΘk\\Theta\_\{k\}satisfy \([20](https://arxiv.org/html/2606.26316#S4.E20)\)–\([21](https://arxiv.org/html/2606.26316#S4.E21)\) onceK0K\_\{0\}satisfies \([23](https://arxiv.org/html/2606.26316#S4.E23)\)\. In particular, it shows that

mk=O~​\(tmix\)m\_\{k\}=\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}\)under geometric mixing\. Since the leading variance term in[Lemma˜6](https://arxiv.org/html/2606.26316#Thmlemma6)scales likemk​∑ℓ<kwℓ,k2m\_\{k\}\\sum\_\{\\ell<k\}w\_\{\\ell,k\}^\{2\}, substituting the geometric\-window estimates into \([22](https://arxiv.org/html/2606.26316#S4.E22)\) gives the leading stochastic order

O~​\(tmixk\+K0\),\\widetilde\{O\}\\\!\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{k\+K\_\{0\}\}\\right\),with only logarithmic factors hidden inO~​\(⋅\)\\widetilde\{O\}\(\\cdot\)\. This proves[Corollary˜1](https://arxiv.org/html/2606.26316#Thmcorollary1)\.

#### Light\-tailed lower bound:[Theorem˜2](https://arxiv.org/html/2606.26316#Thmtheorem2)\.

The lower bound is designed to show that the linear polynomial dependence ontmixt\_\{\\mathrm\{mix\}\}cannot be improved, even in the simplest PL geometry\. The construction uses the one\-dimensional quadratic objective

f​\(x\)=12​x2f\(x\)=\\frac\{1\}\{2\}x^\{2\}and a persistent two\-state Markov chain\. The lower\-bound instance is verified to satisfy[Assumptions˜1](https://arxiv.org/html/2606.26316#Thmassumption1),[2](https://arxiv.org/html/2606.26316#Thmassumption2)and[3](https://arxiv.org/html/2606.26316#Thmassumption3)in[Lemma˜13](https://arxiv.org/html/2606.26316#Thmlemma13)\. The mixing time of the two\-state chain is computed in[Lemma˜14](https://arxiv.org/html/2606.26316#Thmlemma14); the key fact is that the autocorrelation decays geometrically at rate1−2​ε1\-2\\varepsilon, so the mixing time is of order1/ε1/\\varepsilon\.

The SGD recursion on this instance is linear\.[Lemma˜15](https://arxiv.org/html/2606.26316#Thmlemma15)writes the final iterate as a weighted linear filter of the Markov chain:

xk=−σ​∑i=0k−1bi,k​Zi\.x\_\{k\}=\-\\sigma\\sum\_\{i=0\}^\{k\-1\}b\_\{i,k\}Z\_\{i\}\.The weights corresponding to the recent half of the trajectory are lower bounded in[Lemma˜16](https://arxiv.org/html/2606.26316#Thmlemma16)\. Because the chain is persistent, the covariance𝔼​\[Zi​Zj\]\\mathbb\{E\}\[Z\_\{i\}Z\_\{j\}\]remains positive for lags of order1/ε1/\\varepsilon\. Combining this autocorrelation with the recent\-weight lower bound yields the variance lower bound in[Lemma˜17](https://arxiv.org/html/2606.26316#Thmlemma17):

𝔼​\[xk2\]≳σ2ε​k\.\\mathbb\{E\}\[x\_\{k\}^\{2\}\]\\gtrsim\\frac\{\\sigma^\{2\}\}\{\\varepsilon k\}\.Sincetmix≍1/εt\_\{\\mathrm\{mix\}\}\\asymp 1/\\varepsilon, this is exactly the desired expectation lower bound\.

To convert the second\-moment lower bound into a constant\-probability lower bound,[Lemma˜18](https://arxiv.org/html/2606.26316#Thmlemma18)proves a fourth\-moment upper bound for the same linear filter\. The proof of[Theorem˜2](https://arxiv.org/html/2606.26316#Thmtheorem2)then applies Paley–Zygmund toxk2x\_\{k\}^\{2\}\. This yields both \([30](https://arxiv.org/html/2606.26316#S4.E30)\) and \([31](https://arxiv.org/html/2606.26316#S4.E31)\); combining these estimates with[Lemma˜14](https://arxiv.org/html/2606.26316#Thmlemma14)gives \([33](https://arxiv.org/html/2606.26316#S4.E33)\)\. Thus any high\-probability upper bound valid over[Assumptions˜1](https://arxiv.org/html/2606.26316#Thmassumption1),[2](https://arxiv.org/html/2606.26316#Thmassumption2)and[3](https://arxiv.org/html/2606.26316#Thmassumption3)must have leading stochastic order at leastΩ​\(σ2​tmix/k\)\\Omega\(\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}/k\)at constant confidence\.

#### Heavy\-tailed upper bound:[Theorem˜3](https://arxiv.org/html/2606.26316#Thmtheorem3)\.

The heavy\-tailed proof uses a different oracle model and a different algorithm\. Instead of the pathwise ABC envelope in[Assumption˜2](https://arxiv.org/html/2606.26316#Thmassumption2), it uses the stationary finite\-ppmoment condition in[Assumption˜4](https://arxiv.org/html/2606.26316#Thmassumption4)\. Because this condition gives no deterministic envelope and may not imply a finite variance, ordinary martingale concentration cannot be applied directly to the raw gradients\. The algorithm therefore holds the iterate fixed over a block, clips each consecutive Markovian gradient, and averages all clipped samples\. The coordinatewise clipping map is defined in \([34](https://arxiv.org/html/2606.26316#S5.E34)\), and the resulting all\-samples clipped block method is[Definition˜1](https://arxiv.org/html/2606.26316#Thmdefinition1)\.

The first ingredient is the stationary clipping calculation in[Lemma˜19](https://arxiv.org/html/2606.26316#Thmlemma19)\. This lemma records two standard consequences of a finite\-ppmoment assumption\. First, clipping creates a bias of order

vp​λ1−p,v^\{p\}\\lambda^\{1\-p\},wherevvis the localpp\-moment scale andλ\\lambdais the clipping level\. Second, the clipped variable has second moment of order

vp​λ2−p\.v^\{p\}\\lambda^\{2\-p\}\.These are the heavy\-tailed analogues of bounded variance estimates\. They are applied coordinate by coordinate and then unioned over the dimension\.

The second ingredient is the Markovian clipped\-block concentration lemma,[Lemma˜20](https://arxiv.org/html/2606.26316#Thmlemma20)\. This lemma is the heavy\-tailed analogue of the lag\-blocking concentration argument used in the light\-tailed proof\. For a block of lengthbb, the firstmmsamples are controlled directly by the clipping level\. For the remaining samples, the proof delays bymmsteps, subtracts the conditional mean given the sigma\-field one lag in the past, and splits the centered terms into residue classes modulomm\. Freedman’s inequality is then applied within each residue class\. Summing the residue\-class bounds gives the effective variance term

m​b​uht​\(vp​λ2−p\+λ2​ε​\(m\)\),\\sqrt\{mbu\_\{\\rm ht\}\\bigl\(v^\{p\}\\lambda^\{2\-p\}\+\\lambda^\{2\}\\varepsilon\(m\)\\bigr\)\},and after division by the block lengthbb, balancing this term with the clipping bias yields the rate factor

\(m​uhtb\)\(p−1\)/p\.\\left\(\\frac\{mu\_\{\\rm ht\}\}\{b\}\\right\)^\{\(p\-1\)/p\}\.
The third ingredient is deterministic PL descent with inexact gradients,[Lemma˜21](https://arxiv.org/html/2606.26316#Thmlemma21)\. Once the clipped block averageg^r\\widehat\{g\}\_\{r\}satisfies

‖g^r−∇f​\(xr\)‖≤er,\\\|\\widehat\{g\}\_\{r\}\-\\nabla f\(x\_\{r\}\)\\\|\\leq e\_\{r\},this lemma shows that the update with step size1/\(4​L\)1/\(4L\)contracts the objective up to an error of orderer2/μe\_\{r\}^\{2\}/\\mu\. Therefore the heavy\-tailed concentration bound controls optimization error after squaring the gradient\-estimation error\.

The proof of[Theorem˜3](https://arxiv.org/html/2606.26316#Thmtheorem3), given in Appendix[H\.2](https://arxiv.org/html/2606.26316#A8.SS2), combines these three ingredients through a stopped induction\. The stopped event has two purposes\. First, the radius conditionΔ​\(xr\)≤ℛ\\Delta\(x\_\{r\}\)\\leq\\mathcal\{R\}ensures that the local moment scale is bounded byBℛB\_\{\\mathcal\{R\}\}, so[Lemma˜20](https://arxiv.org/html/2606.26316#Thmlemma20)can be applied at the next iterate\. Second, the gradient\-accuracy event ensures that[Lemma˜21](https://arxiv.org/html/2606.26316#Thmlemma21)can be applied to perform one PL descent step\. Thus the proof alternates between concentration and deterministic descent: concentration gives a sufficiently accurate clipped gradient estimate atxrx\_\{r\}, and deterministic PL descent keepsxr\+1x\_\{r\+1\}inside the radius while reducing the suboptimality up to the statistical error floor\.

The geometric transition\-budget rate in[Corollary˜3](https://arxiv.org/html/2606.26316#Thmcorollary3)is obtained by choosing the number of outer blocksNN, the block lengthbb, and the analytical lagmmas functions of the total transition budgetTT\. The choice ofNNmakes the optimization transient negligible, the choice ofbbensures that the block average has enough effective samples, and the choice ofmmmakes the Markovian biasε​\(m\)\\varepsilon\(m\)smaller than the target confidence level\. Under geometric mixing,m=O~​\(tmix\)m=\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}\), so the statistical floor in[Theorem˜3](https://arxiv.org/html/2606.26316#Thmtheorem3)becomes

O~​\(σp2​\(tmixT\)2​\(p−1\)/p\),\\widetilde\{O\}\\\!\\left\(\\sigma\_\{p\}^\{2\}\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{T\}\\right\)^\{2\(p\-1\)/p\}\\right\),up to logarithmic and dimension factors\.

#### Heavy\-tailed lower bound:[Theorem˜4](https://arxiv.org/html/2606.26316#Thmtheorem4)\.

The proof of[Theorem˜4](https://arxiv.org/html/2606.26316#Thmtheorem4), given in Appendix[H\.4](https://arxiv.org/html/2606.26316#A8.SS4), embeds finite\-ppmean estimation into one\-dimensional PL optimization\. The algorithm is allowed to be arbitrary, subject only to observing the oracle transcript and its own internal randomness; it is not told which of the two lower\-bound instances is in force\. The construction uses two objectives with minimizers separated by a distance of order

σp​\(τn\)\(p−1\)/p\.\\sigma\_\{p\}\\left\(\\frac\{\\tau\}\{n\}\\right\)^\{\(p\-1\)/p\}\.Distinguishing these two objectives is therefore necessary for returning a point with smaller suboptimality than the claimed lower bound\.

The Markov chain used in the construction is sticky: it refreshes only once everyO​\(τ\)O\(\\tau\)transitions on average\. Hencennobserved Markov transitions contain onlyO​\(n/τ\)O\(n/\\tau\)effectively independent opportunities to see an informative heavy\-tailed sample\. With constant probability, the transcript contains no informative refresh\. On this event, the two instances produce identical oracle transcripts, so any algorithm coupled with the same internal randomness must output the same point under both instances\. Since the two minimizers are separated, this common output must have large error for at least one of the two objectives\.

The finite\-ppmoment scale determines the separation between the two instances\. The rare refresh distribution is chosen so that the stationarypp\-moment is bounded byσpp\\sigma\_\{p\}^\{p\}, while the mean shift is of order

σp​\(τn\)\(p−1\)/p\.\\sigma\_\{p\}\\left\(\\frac\{\\tau\}\{n\}\\right\)^\{\(p\-1\)/p\}\.Squaring this separation gives the PL suboptimality lower bound

Δ​\(x^n\)≳σp2​\(τn\)2​\(p−1\)/p\\Delta\(\\widehat\{x\}\_\{n\}\)\\gtrsim\\sigma\_\{p\}^\{2\}\\left\(\\frac\{\\tau\}\{n\}\\right\)^\{2\(p\-1\)/p\}with constant probability\. Since the constructed sticky chain has geometric mixing timeO​\(τ\)O\(\\tau\), this matches the dependence in[Corollary˜3](https://arxiv.org/html/2606.26316#Thmcorollary3)up to logarithmic and dimension factors\.

#### Local assumptions:[Corollary˜4](https://arxiv.org/html/2606.26316#Thmcorollary4)\.

The localization result in[Corollary˜4](https://arxiv.org/html/2606.26316#Thmcorollary4)is proved in Appendix[F](https://arxiv.org/html/2606.26316#A6)\. The point of this corollary is that the global smoothness, PL, and oracle\-envelope assumptions need only hold on a sublevel set that contains the stopped trajectory\. The proof repeats the first\-failure induction from Appendix[D](https://arxiv.org/html/2606.26316#A4), but all estimates are restricted to

𝒮R:=\{x:Δ​\(x\)≤R\}\.\\mathcal\{S\}\_\{R\}:=\\\{x:\\Delta\(x\)\\leq R\\\}\.The condition \([51](https://arxiv.org/html/2606.26316#S6.E51)\) ensures that the induction envelope remains inside𝒮R\\mathcal\{S\}\_\{R\}\. Consequently, every time the proof invokes[Lemma˜2](https://arxiv.org/html/2606.26316#Thmlemma2),[Lemma˜4](https://arxiv.org/html/2606.26316#Thmlemma4),[Lemma˜5](https://arxiv.org/html/2606.26316#Thmlemma5),[Lemma˜8](https://arxiv.org/html/2606.26316#Thmlemma8), or[Lemma˜10](https://arxiv.org/html/2606.26316#Thmlemma10), the required constants are the local constants on𝒮R\\mathcal\{S\}\_\{R\}\. Thus the global theorem transfers directly to the local setting once the initial point and the induction envelope are contained in the sublevel region\.

## 8Conclusion

We developed a sharp high\-probability theory for stochastic optimization under the Polyak–Łojasiewicz condition when gradient samples are generated along an exogenous Markov chain\. In the bounded\-oracle regime, we showed that ordinary SGD achieves a uniform\-in\-time rate whose leading stochastic term has orderO~​\(tmix/\(k\+K0\)\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/\(k\+K\_\{0\}\)\)under geometric mixing, up to logarithmic factors\. This closes the gap between theO~​\(tmix2/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}^\{2\}/k\)leading stochastic order produced by Poisson\-equation martingale\-range analyses and theO~​\(tmix/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/k\)order suggested by expectation bounds\. The main technical device is a lag\-blocking argument: the Markovian descent observable is delayed until the chain has mixed, split into residue classes, and then controlled by martingale concentration without changing the SGD algorithm\.

We also proved that this order is unavoidable\. A one\-dimensional quadratic PL objective driven by a persistent two\-state Markov chain already exhibits suboptimality of orderσ2​tmix/k\\sigma^\{2\}t\_\{\\mathrm\{mix\}\}/kin expectation and with constant probability\. Thus the upper bound identifies the optimal polynomial dependence on the mixing time in the light\-tailed Markovian PL\-SGD setting\.

Beyond bounded gradients, we studied a finite\-pp\-moment heavy\-tailed regime\. Here robustification must be algorithmic: the method holds the iterate fixed over short blocks, clips every Markovian gradient in the block, and averages all of the clipped samples\. The resulting high\-probability bound has the effective\-sample\-size rate

O~​\(σp2​\(tmixT\)2​\(p−1\)/p\)\\widetilde\{O\}\\\!\\left\(\\sigma\_\{p\}^\{2\}\\left\(\\frac\{t\_\{\\mathrm\{mix\}\}\}\{T\}\\right\)^\{2\(p\-1\)/p\}\\right\)under geometric mixing, whereTTis the Markov\-transition budget\. A matching sticky\-chain lower bound shows that both the Markovian effective\-sample\-size dependence and the finite\-moment exponent are unavoidable up to logarithmic factors\. The lower bound is one\-dimensional, so the dimension dependence in the upper bound is not claimed to be optimal\. Together, the light\-tailed and heavy\-tailed results characterize the statistical price of Markovian dependence for PL stochastic optimization\.

Several directions remain open\. First, the heavy\-tailed method uses blockwise updates; an important question is whether one can design a fully online robust method that updates after every Markovian sample while retaining the same effective\-sample\-size rate under minimal nonstationary assumptions\. Second, our analysis treats exogenous Markov chains\. Extending the theory to parameter\-dependent kernels, as arise in policy\-gradient, actor–critic, and adaptive\-MCMC methods, would require controlling the interaction between iterate movement, invariant\-distribution drift, and mixing\. Third, the present results use uniform total\-variation mixing\. Developing analogous sharp high\-probability bounds under weaker drift/minorization orVV\-uniform ergodicity assumptions would broaden the scope to noncompact state spaces such as linear dynamical systems with sub\-Gaussian innovations\. Finally, it would be valuable to combine the lag\-blocking viewpoint with variance reduction, acceleration, and decentralized asynchronous updates, where Markovian dependence is intrinsic to the algorithmic architecture\.

## References

- \[1\]\(2016\)Deep learning with differential privacy\.InProceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security,pp\. 308–318\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[2\]S\. Agrawal, S\. T\. Maguluri, and M\. Zubeldia\(2026\)Concentration of general stochastic approximation under heavy\-tailed Markovian noise\.External Links:2605\.20999Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[Table 2](https://arxiv.org/html/2606.26316#A1.T2.24.15.2.1.1.1),[§2](https://arxiv.org/html/2606.26316#S2.SS0.SSS0.Px3.p1.3),[Table 1](https://arxiv.org/html/2606.26316#S2.T1.19.17.4.1.1),[Remark 7](https://arxiv.org/html/2606.26316#Thmremark7.p1.1.1)\.
- \[3\]A\. Armacki, S\. Yu, P\. Sharma, G\. Joshi, D\. Bajovic, D\. Jakovetic, and S\. Kar\(2025\)High\-probability convergence bounds for online nonlinear stochastic gradient descent under heavy\-tailed noise\.InProceedings of The 28th International Conference on Artificial Intelligence and Statistics,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px7.p1.1),[§1\.1](https://arxiv.org/html/2606.26316#S1.SS1.SSS0.Px2.p4.3),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5),[Remark 7](https://arxiv.org/html/2606.26316#Thmremark7.p1.1.1)\.
- \[4\]Q\. Bai, W\. U\. Mondal, and V\. Aggarwal\(2024\)Regret analysis of policy gradient algorithm for infinite horizon average reward markov decision processes\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 10980–10988\.Cited by:[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[5\]A\. Benveniste, M\. Metivier, and P\. Priouret\(1990\)Adaptive Algorithms and Stochastic Approximations\.Springer\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2)\.
- \[6\]A\. Beznosikov, S\. Samsonov, M\. Sheshukova, A\. Gasnikov, A\. Naumov, and E\. Moulines\(2023\)First order methods with Markovian noise: from acceleration to variational inequalities\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px3.p1.2),[Table 2](https://arxiv.org/html/2606.26316#A1.T2.16.4.2.1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2),[§2](https://arxiv.org/html/2606.26316#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2606.26316#S2.T1.11.9.4.1.1),[Remark 3](https://arxiv.org/html/2606.26316#Thmremark3.p1.2.2)\.
- \[7\]J\. Bhandari, D\. Russo, and R\. Singal\(2018\)A finite time analysis of temporal difference learning with linear function approximation\.InProceedings of the Conference on Learning Theory,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[8\]E\. Blaser and S\. Zhang\(2026\)Asymptotic and finite sample analysis of nonexpansive stochastic approximations with Markovian noise\.InProceedings of the AAAI Conference on Artificial Intelligence,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1)\.
- \[9\]V\. S\. Borkar\(2008\)Stochastic Approximation: A Dynamical Systems Viewpoint\.Cambridge University Press\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2)\.
- \[10\]L\. Bottou, F\. E\. Curtis, and J\. Nocedal\(2018\)Optimization methods for large\-scale machine learning\.SIAM Review60\(2\),pp\. 223–311\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[11\]O\. Catoni\(2012\)Challenging the Empirical Mean and Empirical Variance: A Deviation Study\.Annales de l’Institut Henri Poincare\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px8.p1.4),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5)\.
- \[12\]S\. Cayci and A\. Eryilmaz\(2023\)Provably robust temporal difference learning for heavy\-tailed rewards\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px8.p1.4)\.
- \[13\]S\. Chandak, S\. U\. Haque, and N\. Bambos\(2025\)Finite\-time bounds for two\-time\-scale stochastic approximation with arbitrary norm contractions and Markovian noise\.External Links:2503\.18391Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1)\.
- \[14\]C\. A\. Choquette\-Choo, A\. Ganesh, R\. McKenna, H\. B\. McMahan, J\. Rush, A\. Guha Thakurta, and Z\. Xu\(2023\)\(Amplified\) banded matrix factorization: A unified approach to private training\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[15\]A\. Cutkosky and H\. Mehta\(2021\)High\-probability bounds for non\-convex stochastic optimization with heavy tails\.InAdvances in Neural Information Processing Systems,pp\. 4883–4895\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px7.p1.1),[§1\.1](https://arxiv.org/html/2606.26316#S1.SS1.SSS0.Px2.p4.3),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5),[Remark 7](https://arxiv.org/html/2606.26316#Thmremark7.p1.1.1)\.
- \[16\]L\. Devroye, M\. Lerasle, G\. Lugosi, and R\. I\. Oliveira\(2016\)Sub\-Gaussian mean estimators\.The Annals of Statistics44\(6\),pp\. 2695–2725\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px8.p1.4),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5)\.
- \[17\]T\. T\. Doan, L\. M\. Nguyen, N\. H\. Pham, and J\. Romberg\(2020\)Finite\-time analysis of stochastic gradient descent under Markov randomness\.External Links:2003\.10973Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px3.p1.2),[Table 2](https://arxiv.org/html/2606.26316#A1.T2.24.14.1.1.1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2),[§2](https://arxiv.org/html/2606.26316#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2606.26316#S2.T1.5.3.4.1.1),[Remark 3](https://arxiv.org/html/2606.26316#Thmremark3.p1.2.2)\.
- \[18\]T\. T\. Doan\(2023\)Finite\-time analysis of Markov gradient descent\.IEEE Transactions on Automatic Control68\(4\),pp\. 2140–2153\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px3.p1.2),[Table 2](https://arxiv.org/html/2606.26316#A1.T2.24.14.1.1.1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2),[§2](https://arxiv.org/html/2606.26316#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2606.26316#S2.T1.5.3.4.1.1),[Remark 3](https://arxiv.org/html/2606.26316#Thmremark3.p1.2.2)\.
- \[19\]A\. Dong and A\. Ganesh\(2026\)Privacy amplification for BandMF viabb\-min\-sep subsampling\.External Links:2602\.09338Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[20\]J\. C\. Duchi, A\. Agarwal, and M\. J\. Wainwright\(2012\)Dual averaging for distributed optimization: convergence analysis and network scaling\.IEEE Transactions on Automatic Control57\(3\),pp\. 592–606\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px6.p1.1)\.
- \[21\]M\. Even\(2023\)Stochastic gradient descent under Markovian sampling schemes\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 9412–9439\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px3.p1.2),[Table 2](https://arxiv.org/html/2606.26316#A1.T2.15.3.4.1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2),[§2](https://arxiv.org/html/2606.26316#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2606.26316#S2.T1.8.6.4.1.1),[Remark 3](https://arxiv.org/html/2606.26316#Thmremark3.p1.2.2)\.
- \[22\]M\. Fazel, R\. Ge, S\. Kakade, and M\. Mesbahi\(2018\)Global convergence of policy gradient methods for the linear quadratic regulator\.InProceedings of the 35th International Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2),[Remark 1](https://arxiv.org/html/2606.26316#Thmremark1.p2.1.1)\.
- \[23\]S\. Ganesh, W\. U\. Mondal, and V\. Aggarwal\(2025\)A sharper global convergence analysis for average reward reinforcement learning via an actor\-critic approach\.InForty\-second International Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[24\]S\. Ganesh, W\. U\. Mondal, and V\. Aggarwal\(2025\)Order\-optimal regret with novel policy gradient approaches in infinite\-horizon average reward mdps\.InThe 28th International Conference on Artificial Intelligence and Statistics,Cited by:[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[25\]P\. W\. Glynn and D\. Ormoneit\(2002\)Hoeffding’s inequality for uniformly ergodic Markov chains\.Statistics & Probability Letters56\(2\),pp\. 143–146\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px5.p1.2),[Remark 3](https://arxiv.org/html/2606.26316#Thmremark3.p1.2.2)\.
- \[26\]E\. Gorbunov, M\. Danilova, I\. Shibaev, P\. Dvurechensky, and A\. Gasnikov\(2024\)High probability complexity bounds for non\-smooth stochastic optimization with heavy\-tailed noise\.Journal of Optimization Theory and Applications203,pp\. 2679–2738\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px7.p1.1),[§1\.1](https://arxiv.org/html/2606.26316#S1.SS1.SSS0.Px2.p4.3),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5),[Remark 7](https://arxiv.org/html/2606.26316#Thmremark7.p1.1.1)\.
- \[27\]R\. M\. Gower, N\. Loizou, X\. Qian, A\. Sailanbayev, E\. Shulgin, and P\. Richtarik\(2019\)SGD: General analysis and improved rates\.InProceedings of the 36th International Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2),[§1](https://arxiv.org/html/2606.26316#S1.p1.1),[Remark 1](https://arxiv.org/html/2606.26316#Thmremark1.p1.1.1),[Remark 2](https://arxiv.org/html/2606.26316#Thmremark2.p1.1.1)\.
- \[28\]H\. Hendrikx\(2023\)A principled framework for the design and analysis of token algorithms\.InProceedings of The 26th International Conference on Artificial Intelligence and Statistics,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[29\]F\. Huebler, I\. Fatkhullin, and N\. He\(2025\)From gradient clipping to normalization for heavy tailed SGD\.InProceedings of The 28th International Conference on Artificial Intelligence and Statistics,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px7.p1.1),[§1\.1](https://arxiv.org/html/2606.26316#S1.SS1.SSS0.Px2.p4.3),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5),[Remark 7](https://arxiv.org/html/2606.26316#Thmremark7.p1.1.1)\.
- \[30\]B\. Johansson, M\. Rabi, and M\. Johansson\(2007\)A randomized incremental subgradient method for distributed optimization in networked systems\.InProceedings of the American Control Conference,pp\. 5149–5154\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[31\]A\. Juditsky and A\. Nemirovski\(2008\)Large deviations of vector\-valued martingales in 2\-smooth normed spaces\.External Links:0809\.0813Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px8.p1.4)\.
- \[32\]M\. Kaledin, E\. Moulines, A\. Naumov, V\. Tadic, and H\.\-T\. Wai\(2020\)Finite time analysis of linear two\-timescale stochastic approximation with Markovian noise\.InProceedings of Thirty Third Conference on Learning Theory,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[33\]A\. Kar, S\. Chandak, R\. Singh, E\. Moulines, S\. Bhatnagar, and N\. Bambos\(2026\)High\-probability bounds for SGD under the Polyak\-Łojasiewicz condition with Markovian noise\.External Links:2603\.14514Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px3.p1.2),[Table 2](https://arxiv.org/html/2606.26316#A1.T2.18.6.3.1.1),[§1\.1](https://arxiv.org/html/2606.26316#S1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p2.9),[§1](https://arxiv.org/html/2606.26316#S1.p3.2),[§2](https://arxiv.org/html/2606.26316#S2.SS0.SSS0.Px1.p1.4),[Table 1](https://arxiv.org/html/2606.26316#S2.T1.16.14.6.1.1),[Remark 1](https://arxiv.org/html/2606.26316#Thmremark1.p1.1.1),[Remark 2](https://arxiv.org/html/2606.26316#Thmremark2.p1.1.1),[Remark 3](https://arxiv.org/html/2606.26316#Thmremark3.p1.2.2)\.
- \[34\]R\. L\. Karandikar and M\. Vidyasagar\(2024\)Convergence rates for stochastic approximation: biased noise with unbounded variance, and applications\.Journal of Optimization Theory and Applications203,pp\. 2412–2450\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2)\.
- \[35\]H\. Karimi, J\. Nutini, and M\. Schmidt\(2016\)Linear convergence of gradient and proximal\-gradient methods under the Polyak\-Łojasiewicz condition\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 795–811\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2),[§1](https://arxiv.org/html/2606.26316#S1.p2.8),[Remark 1](https://arxiv.org/html/2606.26316#Thmremark1.p1.1.1)\.
- \[36\]A\. Khaled and P\. Richtarik\(2023\)Better theory for SGD in the nonconvex world\.Transactions on Machine Learning Research\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2),[§1](https://arxiv.org/html/2606.26316#S1.p1.1),[Remark 1](https://arxiv.org/html/2606.26316#Thmremark1.p1.1.1),[Remark 2](https://arxiv.org/html/2606.26316#Thmremark2.p1.1.1)\.
- \[37\]J\. Kiefer and J\. Wolfowitz\(1952\)Stochastic estimation of the maximum of a regression function\.The Annals of Mathematical Statistics23\(3\),pp\. 462–466\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[38\]S\. Kowshik, D\. Nagaraj, P\. Jain, and P\. Netrapalli\(2021\)Streaming linear system identification with reverse experience replay\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[39\]H\. J\. Kushner and G\. G\. Yin\(2003\)Stochastic Approximation and Recursive Algorithms and Applications\.Second edition,Springer\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2)\.
- \[40\]P\. Lezaud\(1998\)Chernoff\-type bound for finite Markov chains\.The Annals of Applied Probability8\(3\),pp\. 849–867\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px5.p1.2)\.
- \[41\]X\. Li, Z\. Zhuang, and F\. Orabona\(2021\)A second look at exponential and cosine step sizes: simplicity, adaptivity, and performance\.InProceedings of the 38th International Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2),[Remark 2](https://arxiv.org/html/2606.26316#Thmremark2.p1.1.1)\.
- \[42\]C\. Liu, L\. Zhu, and M\. Belkin\(2022\)Loss landscapes and optimization in over\-parameterized non\-linear systems and neural networks\.Applied and Computational Harmonic Analysis59,pp\. 85–116\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2),[Remark 1](https://arxiv.org/html/2606.26316#Thmremark1.p2.1.1)\.
- \[43\]S\. D\. Liu, S\. Chen, and S\. Zhang\(2025\)The ODE method for stochastic approximation and reinforcement learning with Markovian noise\.Journal of Machine Learning Research26\(24\),pp\. 1–76\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[44\]L\. Ljung\(1977\)Analysis of recursive stochastic algorithms\.IEEE Transactions on Automatic Control22\(4\),pp\. 551–575\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1)\.
- \[45\]G\. Lugosi and S\. Mendelson\(2019\)Mean estimation and regression under heavy\-tailed distributions: a survey\.Foundations of Computational Mathematics19,pp\. 1145–1190\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px8.p1.4),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5)\.
- \[46\]L\. Madden, E\. Dall’Anese, and S\. Becker\(2024\)High probability convergence bounds for non\-convex stochastic gradient descent with sub\-Weibull noise\.Journal of Machine Learning Research25\(241\),pp\. 1–36\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2)\.
- \[47\]S\. Maity and A\. Mitra\(2025\)Adversarially\-robust TD learning with Markovian data: finite\-time rates and fundamental limits\.InProceedings of The 28th International Conference on Artificial Intelligence and Statistics,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px8.p1.4)\.
- \[48\]X\. Mao, K\. Yuan, Y\. Hu, Y\. Gu, A\. H\. Sayed, and W\. Yin\(2020\)Walkman: a communication\-efficient random\-walk algorithm for decentralized optimization\.IEEE Transactions on Signal Processing68,pp\. 2513–2528\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[49\]M\. Metivier and P\. Priouret\(1984\)Applications of a Kushner and Clark lemma to general classes of stochastic algorithms\.IEEE Transactions on Information Theory30\(2\),pp\. 140–151\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2)\.
- \[50\]E\. Moulines and F\. Bach\(2011\)Non\-asymptotic analysis of stochastic approximation algorithms for machine learning\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[51\]A\. Nemirovski, A\. Juditsky, G\. Lan, and A\. Shapiro\(2009\)Robust stochastic approximation approach to stochastic programming\.SIAM Journal on Optimization19\(4\),pp\. 1574–1609\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px8.p1.4)\.
- \[52\]T\. D\. Nguyen, T\. H\. Nguyen, A\. Ene, and H\. L\. Nguyen\(2023\)High probability convergence of clipped\-SGD under heavy\-tailed noise\.External Links:2302\.05437Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px7.p1.1),[§1\.1](https://arxiv.org/html/2606.26316#S1.SS1.SSS0.Px2.p4.3),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5),[Remark 7](https://arxiv.org/html/2606.26316#Thmremark7.p1.1.1)\.
- \[53\]D\. Paulin\(2015\)Concentration inequalities for Markov chains by Marton couplings and spectral methods\.Electronic Journal of Probability20,pp\. 1–32\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px5.p1.2),[Remark 3](https://arxiv.org/html/2606.26316#Thmremark3.p1.2.2)\.
- \[54\]B\. T\. Polyak\(1963\)Gradient methods for minimizing functionals\.USSR Computational Mathematics and Mathematical Physics3\(4\),pp\. 864–878\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px2.p1.2),[§1](https://arxiv.org/html/2606.26316#S1.p2.8),[Remark 1](https://arxiv.org/html/2606.26316#Thmremark1.p1.1.1)\.
- \[55\]N\. Puchkin, E\. Gorbunov, N\. Kutuzov, and A\. Gasnikov\(2024\)Breaking the heavy\-tailed noise barrier in stochastic optimization problems\.InProceedings of The 27th International Conference on Artificial Intelligence and Statistics,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px7.p1.1),[§1\.1](https://arxiv.org/html/2606.26316#S1.SS1.SSS0.Px2.p4.3),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5),[Remark 7](https://arxiv.org/html/2606.26316#Thmremark7.p1.1.1)\.
- \[56\]A\. Rakhlin, O\. Shamir, and K\. Sridharan\(2011\)Making gradient descent optimal for strongly convex stochastic optimization\.External Links:1109\.5647Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1)\.
- \[57\]H\. Robbins and S\. Monro\(1951\)A stochastic approximation method\.The Annals of Mathematical Statistics22\(3\),pp\. 400–407\.Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[58\]A\. Sadiev, M\. Danilova, E\. Gorbunov, S\. Horvath, G\. Gidel, P\. Dvurechensky, A\. Gasnikov, and P\. Richtarik\(2023\)High\-probability bounds for stochastic optimization and variational inequalities: the case of unbounded variance\.InProceedings of the 40th International Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px7.p1.1),[§1\.1](https://arxiv.org/html/2606.26316#S1.SS1.SSS0.Px2.p4.3),[Remark 6](https://arxiv.org/html/2606.26316#Thmremark6.p1.5.5),[Remark 7](https://arxiv.org/html/2606.26316#Thmremark7.p1.1.1)\.
- \[59\]R\. Srikant and L\. Ying\(2019\)Finite\-time error bounds for linear stochastic approximation and TD learning\.InProceedings of the Conference on Learning Theory,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2606.26316#S1.p1.1)\.
- \[60\]T\. Sun, Y\. Sun, and W\. Yin\(2018\)On Markov chain gradient descent\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px3.p1.2),[Table 2](https://arxiv.org/html/2606.26316#A1.T2.24.14.1.1.1.1),[§1](https://arxiv.org/html/2606.26316#S1.p3.2),[§2](https://arxiv.org/html/2606.26316#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2606.26316#S2.T1.5.3.4.1.1),[Remark 3](https://arxiv.org/html/2606.26316#Thmremark3.p1.2.2)\.
- \[61\]H\. Wang, M\. Gurbuzbalaban, L\. Zhu, U\. Simsekli, and M\. A\. Erdogdu\(2021\)Convergence rates of stochastic gradient descent under infinite noise variance\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2606.26316#A1.SS0.SSS0.Px7.p1.1)\.
- \[62\]Y\. Xu and V\. Aggarwal\(2026\)Persistent\-transient policy evaluation for markov chains via minimal peripheral quotients\.arXiv preprint arXiv:2602\.00474\.Cited by:[§1](https://arxiv.org/html/2606.26316#S1.p3.2)\.

## Appendix AExtended Related work

This section reviews the closest work on stochastic approximation, PL optimization, Markovian gradient methods, concentration for Markov chains, and robust optimization under heavy tails\. The most relevant comparison for our main theorem is summarized in[Table˜2](https://arxiv.org/html/2606.26316#A1.T2); the paragraphs below then give broader context\.

Table 2:Closest results on Markovian stochastic optimization\. Herekkis the number of SGD iterations,TTis the total number of Markov transitions,tmixt\_\{\\mathrm\{mix\}\}orτ\\tauis a mixing\-time parameter, andτhit\\tau\_\{\\mathrm\{hit\}\}is a hitting\-time parameter\. The notationO~​\(⋅\)\\widetilde\{O\}\(\\cdot\)hides logarithmic factors and problem\-dependent constants\. Rows are not directly comparable when they solve a different objective class, use a different algorithmic model, or prove a different type of lower bound\.We next review the surrounding literature in more detail\.

#### Stochastic approximation and SGD\.

The recursion studied here belongs to the classical stochastic approximation lineage initiated by Robbins and Monro\[[57](https://arxiv.org/html/2606.26316#bib.bib58)\]and Kiefer and Wolfowitz\[[37](https://arxiv.org/html/2606.26316#bib.bib38)\]\. The asymptotic ODE viewpoint and martingale methods for recursive algorithms were developed by Ljung\[[44](https://arxiv.org/html/2606.26316#bib.bib45)\], Kushner and Yin\[[39](https://arxiv.org/html/2606.26316#bib.bib40)\], Borkar\[[9](https://arxiv.org/html/2606.26316#bib.bib12)\], and Benveniste, Metivier, and Priouret\[[5](https://arxiv.org/html/2606.26316#bib.bib8)\]\. Nonasymptotic analyses for stochastic gradient and stochastic approximation methods include Moulines and Bach\[[50](https://arxiv.org/html/2606.26316#bib.bib51)\], Rakhlin, Shamir, and Sridharan\[[56](https://arxiv.org/html/2606.26316#bib.bib57)\], Gower et al\.\[[27](https://arxiv.org/html/2606.26316#bib.bib28)\], and the broad optimization perspective of Bottou, Curtis, and Nocedal\[[10](https://arxiv.org/html/2606.26316#bib.bib13)\]\. Many of these works assume conditionally unbiased martingale\-difference noise or independent sampling\. Our work instead focuses on temporally dependent gradient samples generated by a Markov chain and seeks uniform\-in\-time high\-probability bounds\.

#### PL geometry and SGD under growth conditions\.

The Polyak\-Łojasiewicz inequality originated in Polyak’s work on gradient methods\[[54](https://arxiv.org/html/2606.26316#bib.bib55)\]; the modern optimization formulation and its relationship to error bounds, restricted secant inequalities, quadratic growth, and related conditions were developed systematically by Karimi, Nutini, and Schmidt\[[35](https://arxiv.org/html/2606.26316#bib.bib36)\]\. The condition is weaker than strong convexity and appears in overparameterized learning and control; examples include wide neural networks\[[42](https://arxiv.org/html/2606.26316#bib.bib43)\], linear\-quadratic control\[[22](https://arxiv.org/html/2606.26316#bib.bib25)\], and composed strongly convex objectives\. For stochastic gradients under PL geometry, several analyses allow the noise magnitude to grow with the local optimization scale rather than requiring a uniform bounded\-variance assumption; examples include Gower et al\.\[[27](https://arxiv.org/html/2606.26316#bib.bib28)\], Li, Zhuang, and Orabona\[[41](https://arxiv.org/html/2606.26316#bib.bib42)\], and Khaled and Richtarik\[[36](https://arxiv.org/html/2606.26316#bib.bib37)\]\. The ABC envelope is one such growth condition: it bounds the stochastic\-gradient magnitude by a combination of‖∇f​\(x\)‖2\\\|\\nabla f\(x\)\\\|^\{2\}, the objective gapf​\(x\)−f⋆f\(x\)\-f^\{\\star\}, and an additive noise floor\. Such growth conditions are important for least squares, interpolation, and minibatch sampling because uniform bounded variance is often too restrictive\. High\-probability PL results under light\-tailed martingale noise have been obtained by Madden, Dall’Anese, and Becker\[[46](https://arxiv.org/html/2606.26316#bib.bib47)\], while Karandikar and Vidyasagar\[[34](https://arxiv.org/html/2606.26316#bib.bib35)\]study almost\-sure rates under biased noise and unbounded variance\. These results do not address the Markovian sampling bias that is central in this paper\.

#### SGD and first\-order methods with Markovian sampling\.

Markovian gradient noise has been studied in convex, nonconvex, strongly convex, and PL settings\. Sun, Sun, and Yin\[[60](https://arxiv.org/html/2606.26316#bib.bib61)\]analyze Markov chain gradient descent for convex problems and inexact gradients\. Doan, Nguyen, Pham, and Romberg\[[17](https://arxiv.org/html/2606.26316#bib.bib20)\]and Doan\[[18](https://arxiv.org/html/2606.26316#bib.bib21)\]provide finite\-time analyses for stochastic gradient algorithms under Markov randomness\. Even\[[21](https://arxiv.org/html/2606.26316#bib.bib24)\]studies Markov\-chain SGD under mild assumptions, develops lower bounds involving hitting\-time quantities, and introduces MC\-SAG as a variance\-reduced method for Markovian sampling\. Beznosikov et al\.\[[6](https://arxiv.org/html/2606.26316#bib.bib9)\]obtain optimal one\-power mixing\-time dependence for several first\-order methods through randomized batching and multilevel ideas, and also treat variational inequalities\. Kar et al\.\[[33](https://arxiv.org/html/2606.26316#bib.bib34)\]establish the first uniform\-in\-time high\-probability PL\-SGD theorem with Markovian plus martingale\-difference noise under an ABC envelope; their high\-probability theorem has leading stochastic orderO~​\(tmix2/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}^\{2\}/k\), whereas their expectation theorem has leading stochastic orderO~​\(tmix/k\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/k\)\. The light\-tailed part of our paper closes this high\-probability mixing gap for the ordinary one\-sample SGD recursion, up to logarithmic factors, and proves that the resulting order is sharp\.

#### Markovian stochastic approximation and reinforcement learning\.

A large body of stochastic approximation and reinforcement\-learning theory uses Markovian noise\. Poisson\-equation decompositions for Markov\-driven stochastic approximation go back at least to Metivier and Priouret\[[49](https://arxiv.org/html/2606.26316#bib.bib50)\]and are standard in the monographs of Benveniste et al\.\[[5](https://arxiv.org/html/2606.26316#bib.bib8)\], Kushner and Yin\[[39](https://arxiv.org/html/2606.26316#bib.bib40)\], and Borkar\[[9](https://arxiv.org/html/2606.26316#bib.bib12)\]\. In reinforcement learning, temporal\-difference learning and linear stochastic approximation under Markovian data have been studied through finite\-sample, asymptotic, and bootstrap viewpoints; representative works include Bhandariet al\.\[[7](https://arxiv.org/html/2606.26316#bib.bib10)\], Srikant and Ying\[[59](https://arxiv.org/html/2606.26316#bib.bib60)\], Kaledin et al\.\[[32](https://arxiv.org/html/2606.26316#bib.bib33)\], Liu et al\.\[[43](https://arxiv.org/html/2606.26316#bib.bib44)\], and Ganesh et al\.\[[23](https://arxiv.org/html/2606.26316#bib.bib5)\]\. Recent work also treats nonexpansive or two\-time\-scale stochastic approximation under Markovian data\[[8](https://arxiv.org/html/2606.26316#bib.bib11),[13](https://arxiv.org/html/2606.26316#bib.bib18)\]\. Agrawal et al\.\[[2](https://arxiv.org/html/2606.26316#bib.bib6)\]study concentration for general stochastic approximation with a finite\-state Markovian component and heavy\-tailed martingale noise\. Our focus is different: we study last\-iterate optimization under PL geometry, allow the Markovian gradient component itself to be heavy\-tailed in the robust regime, and prove matching optimization lower bounds\.

#### Concentration for Markov chains\.

Classical concentration inequalities for Markov chains include Hoeffding\-type bounds for uniformly ergodic chains\[[25](https://arxiv.org/html/2606.26316#bib.bib26)\], Chernoff and spectral\-gap bounds for finite chains\[[40](https://arxiv.org/html/2606.26316#bib.bib41)\], and the coupling and spectral inequalities of Paulin\[[53](https://arxiv.org/html/2606.26316#bib.bib54)\]\. These inequalities show that additive functionals of a Markov chain often behave like averages over an effective sample size reduced by a mixing parameter\. They are not directly plug\-and\-play for adaptive SGD, however, because the observable at timekkdepends on the iteratexkx\_\{k\}, which is itself a function of the past data\. Poisson\-equation methods handle this adaptivity but can introduce a squared mixing factor when combined with worst\-case martingale range bounds\. Our lag\-blocking proof is a direct concentration argument for the weighted, iterate\-adapted additive functional: the delay creates approximate independence, and splitting into residue classes converts the delayed sum into martingale differences without changing the algorithm\.

#### Decentralized, MCMC, privacy, and system\-identification motivations\.

Markovian sampling appears naturally in token\-based and random\-walk decentralized optimization\. Early randomized incremental and distributed methods include Johansson, Rabi, and Johansson\[[30](https://arxiv.org/html/2606.26316#bib.bib31)\]and Duchi, Agarwal, and Wainwright\[[20](https://arxiv.org/html/2606.26316#bib.bib23)\]; more recent token frameworks include Walkman\[[48](https://arxiv.org/html/2606.26316#bib.bib49)\]and the principled token\-algorithm design of Hendrikx\[[28](https://arxiv.org/html/2606.26316#bib.bib29)\]\. Markov chains also arise when gradients are estimated using MCMC samples, where mixing controls the bias and variance of the gradient estimator\. In privacy\-preserving learning, subsampling mechanisms for privacy amplification, including differentially private SGD\[[1](https://arxiv.org/html/2606.26316#bib.bib1)\], banded matrix\-factorization schemes\[[14](https://arxiv.org/html/2606.26316#bib.bib16)\], and minimum\-separation subsampling\[[19](https://arxiv.org/html/2606.26316#bib.bib22)\], can produce temporally dependent minibatches\. In online system identification, observations from stable dynamical systems generate Markovian regressors, as in streaming linear system identification\[[38](https://arxiv.org/html/2606.26316#bib.bib39)\]\. These examples motivate a theory that treats Markovian dependence as a first\-order parameter rather than as a lower\-order perturbation\.

#### Heavy\-tailed stochastic optimization\.

Heavy\-tailed gradient noise has motivated robust stochastic methods that avoid sub\-Gaussian or bounded\-variance assumptions\. For SGD under infinite variance, Wang et al\.\[[61](https://arxiv.org/html/2606.26316#bib.bib62)\]derive convergence rates in expectation\. High\-probability nonconvex stochastic optimization with finiteppth moments has been studied by Cutkosky and Mehta\[[15](https://arxiv.org/html/2606.26316#bib.bib17)\], Gorbunov et al\.\[[26](https://arxiv.org/html/2606.26316#bib.bib27)\], Sadiev et al\.\[[58](https://arxiv.org/html/2606.26316#bib.bib59)\], and Nguyen et al\.\[[52](https://arxiv.org/html/2606.26316#bib.bib53)\], with clipping, normalization, momentum, or mirror\-descent variants playing a central role\. Puchkin et al\.\[[55](https://arxiv.org/html/2606.26316#bib.bib56)\]use smoothed median\-of\-means ideas to stabilize gradients, Huebler, Fatkhullin, and He\[[29](https://arxiv.org/html/2606.26316#bib.bib30)\]study normalized SGD under heavy tails, and Armacki et al\.\[[3](https://arxiv.org/html/2606.26316#bib.bib7)\]give high\-probability guarantees for a class of nonlinear SGD transformations including clipping, quantization, and sign methods\. These works largely concern independent, online, or martingale\-type noise rather than Markovian gradient samples\. The heavy\-tailed part of this paper combines clipping with consecutive Markovian blocks and obtains the effective\-sample\-size rate dictated by the mixing profile while using all samples inside each block\.

#### Robust statistics, heavy\-tailed mean estimation, and lower bounds\.

Our heavy\-tailed lower bound is based on the classical connection between stochastic optimization over quadratics and mean estimation\. Robust mean estimation under heavy tails is a mature subject, including Catoni’s estimator\[[11](https://arxiv.org/html/2606.26316#bib.bib14)\], median\-of\-means and tournament estimators, and the survey of Lugosi and Mendelson\[[45](https://arxiv.org/html/2606.26316#bib.bib46)\]; Devroye et al\.\[[16](https://arxiv.org/html/2606.26316#bib.bib19)\]give sub\-Gaussian mean estimators under finite variance\. Robust stochastic approximation ideas also appear in Juditsky and Nemirovski\[[31](https://arxiv.org/html/2606.26316#bib.bib32)\]and Nemirovski et al\.\[[51](https://arxiv.org/html/2606.26316#bib.bib52)\]\. For Markovian data, recent robust reinforcement\-learning work analyzes median\-of\-means or clipping under time\-correlated samples, including robust TD learning under contamination\[[47](https://arxiv.org/html/2606.26316#bib.bib48)\]and heavy\-tailed rewards\[[12](https://arxiv.org/html/2606.26316#bib.bib15)\]\. We use a sticky\-chain construction to show thatkkMarkovian samples can contain onlyk/tmixk/t\_\{\\mathrm\{mix\}\}effective refreshes, and therefore the finite\-ppmean\-estimation exponent2​\(p−1\)/p2\(p\-1\)/pis unavoidable for PL optimization as well\.

#### Positioning\.

The present paper therefore has two complementary roles\. In the light\-tailed regime, it sharpens the high\-probability Markovian PL\-SGD theory by matching the order suggested by expectation bounds and by proving a corresponding lower bound\. In the heavy\-tailed regime, it extends robust clipped\-gradient guarantees from independent or martingale sampling to Markovian gradient oracles through all\-samples clipped blocks, and proves that both the Markovian effective\-sample\-size dependence and finite\-moment exponent are unavoidable, up to logarithmic and dimension factors\.

## Appendix BDeterministic backbone for the light\-tailed upper bound

This appendix collects deterministic estimates used throughout the proof of[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)\. Lemma[1](https://arxiv.org/html/2606.26316#Thmlemma1)converts smoothness and lower boundedness into the gradient upper bound \([5](https://arxiv.org/html/2606.26316#S3.E5)\)\. Lemma[2](https://arxiv.org/html/2606.26316#Thmlemma2)derives the weighted PL descent recursion\. Lemma[3](https://arxiv.org/html/2606.26316#Thmlemma3)gives the weight estimates used in all concentration bounds, and Lemma[4](https://arxiv.org/html/2606.26316#Thmlemma4)controls the Markovian descent observableh​\(x,z\)h\(x,z\)\. Together these lemmas form the deterministic backbone of the lag\-blocking proof\.

###### Lemma 1\(Gradient upper bound under smoothness\)\.

Under Assumption[1](https://arxiv.org/html/2606.26316#Thmassumption1), \([5](https://arxiv.org/html/2606.26316#S3.E5)\) holds\.

###### Proof\.

Fixx∈ℝdx\\in\\mathbb\{R\}^\{d\}\. ByLL\-smoothness, for everyyy,

f​\(y\)≤f​\(x\)\+⟨∇f​\(x\),y−x⟩\+L2​‖y−x‖2\.f\(y\)\\leq f\(x\)\+\\left\\langle\\nabla f\(x\),y\-x\\right\\rangle\+\\frac\{L\}\{2\}\\left\\lVert y\-x\\right\\rVert^\{2\}\.Takey=x−L−1​∇f​\(x\)y=x\-L^\{\-1\}\\nabla f\(x\)\. Then

f​\(x−L−1​∇f​\(x\)\)≤f​\(x\)−12​L​‖∇f​\(x\)‖2\.f\\left\(x\-L^\{\-1\}\\nabla f\(x\)\\right\)\\leq f\(x\)\-\\frac\{1\}\{2L\}\\left\\lVert\\nabla f\(x\)\\right\\rVert^\{2\}\.Sincef⋆≤f​\(x−L−1​∇f​\(x\)\)f^\{\\star\}\\leq f\(x\-L^\{\-1\}\\nabla f\(x\)\), rearranging gives

‖∇f​\(x\)‖2≤2​L​\(f​\(x\)−f⋆\)=2​L​Δ​\(x\)\.\\left\\lVert\\nabla f\(x\)\\right\\rVert^\{2\}\\leq 2L\(f\(x\)\-f^\{\\star\}\)=2L\\Delta\(x\)\.∎

###### Lemma 2\(Weighted descent recursion\)\.

Under Assumptions[1](https://arxiv.org/html/2606.26316#Thmassumption1)and[2](https://arxiv.org/html/2606.26316#Thmassumption2), if \([14](https://arxiv.org/html/2606.26316#S3.E14)\) holds, then for everyk≥1k\\geq 1,

Δk\\displaystyle\\Delta\_\{k\}≤Δ0​ζ0,k−1\+L​C2​∑ℓ=0k−1αℓ2​ζℓ\+1,k−1−∑ℓ=0k−1wℓ,k​⟨∇f​\(xℓ\),Mℓ\+1⟩−∑ℓ=0k−1wℓ,k​h​\(xℓ,Zℓ\)\.\\displaystyle\\leq\\Delta\_\{0\}\\zeta\_\{0,k\-1\}\+\\frac\{LC\}\{2\}\\sum\_\{\\ell=0\}^\{k\-1\}\\alpha\_\{\\ell\}^\{2\}\\zeta\_\{\\ell\+1,k\-1\}\-\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}\\left\\langle\\nabla f\(x\_\{\\ell\}\),M\_\{\\ell\+1\}\\right\\rangle\-\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)\.\(52\)

###### Proof\.

ByLL\-smoothness and the updatexn\+1=xn−αn​Gnx\_\{n\+1\}=x\_\{n\}\-\\alpha\_\{n\}G\_\{n\},

f​\(xn\+1\)\\displaystyle f\(x\_\{n\+1\}\)≤f​\(xn\)−αn​⟨∇f​\(xn\),Gn⟩\+L​αn22​‖Gn‖2\\displaystyle\\leq f\(x\_\{n\}\)\-\\alpha\_\{n\}\\left\\langle\\nabla f\(x\_\{n\}\),G\_\{n\}\\right\\rangle\+\\frac\{L\\alpha\_\{n\}^\{2\}\}\{2\}\\left\\lVert G\_\{n\}\\right\\rVert^\{2\}=f​\(xn\)−αn​‖∇f​\(xn\)‖2−αn​h​\(xn,Zn\)−αn​⟨∇f​\(xn\),Mn\+1⟩\+L​αn22​‖Gn‖2\.\\displaystyle=f\(x\_\{n\}\)\-\\alpha\_\{n\}\\left\\lVert\\nabla f\(x\_\{n\}\)\\right\\rVert^\{2\}\-\\alpha\_\{n\}h\(x\_\{n\},Z\_\{n\}\)\-\\alpha\_\{n\}\\left\\langle\\nabla f\(x\_\{n\}\),M\_\{n\+1\}\\right\\rangle\+\\frac\{L\\alpha\_\{n\}^\{2\}\}\{2\}\\left\\lVert G\_\{n\}\\right\\rVert^\{2\}\.Using \([7](https://arxiv.org/html/2606.26316#S3.E7)\), the PL inequality \([4](https://arxiv.org/html/2606.26316#S3.E4)\), and the gradient upper bound \([5](https://arxiv.org/html/2606.26316#S3.E5)\),

Δn\+1\\displaystyle\\Delta\_\{n\+1\}≤Δn−\(αn−A​L​αn22\)​‖∇f​\(xn\)‖2\+B​L​αn22​Δn\+L​C​αn22\\displaystyle\\leq\\Delta\_\{n\}\-\\left\(\\alpha\_\{n\}\-\\frac\{AL\\alpha\_\{n\}^\{2\}\}\{2\}\\right\)\\left\\lVert\\nabla f\(x\_\{n\}\)\\right\\rVert^\{2\}\+\\frac\{BL\\alpha\_\{n\}^\{2\}\}\{2\}\\Delta\_\{n\}\+\\frac\{LC\\alpha\_\{n\}^\{2\}\}\{2\}−αn​h​\(xn,Zn\)−αn​⟨∇f​\(xn\),Mn\+1⟩\\displaystyle\\quad\-\\alpha\_\{n\}h\(x\_\{n\},Z\_\{n\}\)\-\\alpha\_\{n\}\\left\\langle\\nabla f\(x\_\{n\}\),M\_\{n\+1\}\\right\\rangle≤\(1−2​μ​αn\+\(2​μ​A\+B\)​L​αn22\)​Δn\+L​C​αn22−αn​h​\(xn,Zn\)−αn​⟨∇f​\(xn\),Mn\+1⟩\.\\displaystyle\\leq\\left\(1\-2\\mu\\alpha\_\{n\}\+\\frac\{\(2\\mu A\+B\)L\\alpha\_\{n\}^\{2\}\}\{2\}\\right\)\\Delta\_\{n\}\+\\frac\{LC\\alpha\_\{n\}^\{2\}\}\{2\}\-\\alpha\_\{n\}h\(x\_\{n\},Z\_\{n\}\)\-\\alpha\_\{n\}\\left\\langle\\nabla f\(x\_\{n\}\),M\_\{n\+1\}\\right\\rangle\.The condition \([14](https://arxiv.org/html/2606.26316#S3.E14)\) implies

\(2​μ​A\+B\)​L​αn22≤μ​αn,\\frac\{\(2\\mu A\+B\)L\\alpha\_\{n\}^\{2\}\}\{2\}\\leq\\mu\\alpha\_\{n\},so

Δn\+1≤\(1−μ​αn\)​Δn\+L​C​αn22−αn​h​\(xn,Zn\)−αn​⟨∇f​\(xn\),Mn\+1⟩\.\\Delta\_\{n\+1\}\\leq\(1\-\\mu\\alpha\_\{n\}\)\\Delta\_\{n\}\+\\frac\{LC\\alpha\_\{n\}^\{2\}\}\{2\}\-\\alpha\_\{n\}h\(x\_\{n\},Z\_\{n\}\)\-\\alpha\_\{n\}\\left\\langle\\nabla f\(x\_\{n\}\),M\_\{n\+1\}\\right\\rangle\.Iterating this affine recursion fromn=0n=0ton=k−1n=k\-1gives \([52](https://arxiv.org/html/2606.26316#A2.E52)\)\. ∎

###### Lemma 3\(Weight estimates\)\.

Letβ:=μ​a≥2\\beta:=\\mu a\\geq 2and assumeK0≥βK\_\{0\}\\geq\\beta\. There is a constantcw<∞c\_\{w\}<\\infty, depending only onaaandμ\\mu, such that for everyk≥1k\\geq 1,

ζ0,k−1\\displaystyle\\zeta\_\{0,k\-1\}≤K0k\+K0,\\displaystyle\\leq\\frac\{K\_\{0\}\}\{k\+K\_\{0\}\},\(53\)∑ℓ=0k−1αℓ2​ζℓ\+1,k−1\\displaystyle\\sum\_\{\\ell=0\}^\{k\-1\}\\alpha\_\{\\ell\}^\{2\}\\zeta\_\{\\ell\+1,k\-1\}≤cwk\+K0,\\displaystyle\\leq\\frac\{c\_\{w\}\}\{k\+K\_\{0\}\},\(54\)wℓ,k\\displaystyle w\_\{\\ell,k\}≤cw​ℓ\+K0\(k\+K0\)2,0≤ℓ≤k−1,\\displaystyle\\leq c\_\{w\}\\frac\{\\ell\+K\_\{0\}\}\{\(k\+K\_\{0\}\)^\{2\}\},\\qquad 0\\leq\\ell\\leq k\-1,\(55\)∑ℓ=0k−1wℓ,k2ℓ\+K0\\displaystyle\\sum\_\{\\ell=0\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}^\{2\}\}\{\\ell\+K\_\{0\}\}≤cw\(k\+K0\)2,\\displaystyle\\leq\\frac\{c\_\{w\}\}\{\(k\+K\_\{0\}\)^\{2\}\},\(56\)∑ℓ=0k−1wℓ,k2\(ℓ\+K0\)2\\displaystyle\\sum\_\{\\ell=0\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}^\{2\}\}\{\(\\ell\+K\_\{0\}\)^\{2\}\}≤cw\(k\+K0\)3\.\\displaystyle\\leq\\frac\{c\_\{w\}\}\{\(k\+K\_\{0\}\)^\{3\}\}\.\(57\)Moreover, for every integerm∈\{1,…,k\}m\\in\\\{1,\\ldots,k\\\},

∑ℓ=mk−1wℓ,k2ℓ−m\+K0\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}^\{2\}\}\{\\ell\-m\+K\_\{0\}\}≤\(1\+mK0\)​cw\(k\+K0\)2,\\displaystyle\\leq\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\frac\{c\_\{w\}\}\{\(k\+K\_\{0\}\)^\{2\}\},\(58\)∑ℓ=mk−1wℓ,k2\(ℓ−m\+K0\)2\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}^\{2\}\}\{\(\\ell\-m\+K\_\{0\}\)^\{2\}\}≤\(1\+mK0\)2​cw\(k\+K0\)3\.\\displaystyle\\leq\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)^\{2\}\\frac\{c\_\{w\}\}\{\(k\+K\_\{0\}\)^\{3\}\}\.\(59\)

###### Proof\.

SinceK0≥βK\_\{0\}\\geq\\beta, we have0≤μ​αj≤10\\leq\\mu\\alpha\_\{j\}\\leq 1\. For0≤ℓ≤k−10\\leq\\ell\\leq k\-1,

ζℓ\+1,k−1\\displaystyle\\zeta\_\{\\ell\+1,k\-1\}=∏j=ℓ\+1k−1\(1−βj\+K0\)≤exp⁡\(−β​∑j=ℓ\+1k−11j\+K0\)\.\\displaystyle=\\prod\_\{j=\\ell\+1\}^\{k\-1\}\\left\(1\-\\frac\{\\beta\}\{j\+K\_\{0\}\}\\right\)\\leq\\exp\\left\(\-\\beta\\sum\_\{j=\\ell\+1\}^\{k\-1\}\\frac\{1\}\{j\+K\_\{0\}\}\\right\)\.Using the integral comparison

∑j=ℓ\+1k−11j\+K0≥∫ℓ\+1kd​tt\+K0=log⁡k\+K0ℓ\+K0\+1,\\sum\_\{j=\\ell\+1\}^\{k\-1\}\\frac\{1\}\{j\+K\_\{0\}\}\\geq\\int\_\{\\ell\+1\}^\{k\}\\frac\{dt\}\{t\+K\_\{0\}\}=\\log\\frac\{k\+K\_\{0\}\}\{\\ell\+K\_\{0\}\+1\},we get

ζℓ\+1,k−1≤\(ℓ\+K0\+1k\+K0\)β\.\\zeta\_\{\\ell\+1,k\-1\}\\leq\\left\(\\frac\{\\ell\+K\_\{0\}\+1\}\{k\+K\_\{0\}\}\\right\)^\{\\beta\}\.Because\(ℓ\+K0\+1\)/\(ℓ\+K0\)≤1\+K0−1\(\\ell\+K\_\{0\}\+1\)/\(\\ell\+K\_\{0\}\)\\leq 1\+K\_\{0\}^\{\-1\}andβ/K0≤1\\beta/K\_\{0\}\\leq 1,

\(ℓ\+K0\+1k\+K0\)β≤e​\(ℓ\+K0k\+K0\)β\.\\left\(\\frac\{\\ell\+K\_\{0\}\+1\}\{k\+K\_\{0\}\}\\right\)^\{\\beta\}\\leq e\\left\(\\frac\{\\ell\+K\_\{0\}\}\{k\+K\_\{0\}\}\\right\)^\{\\beta\}\.Increasing constants fromeetoe2e^\{2\}if needed, this bound also covers the empty\-product edge cases\. Sinceβ≥2\\beta\\geq 2,

wℓ,k=aℓ\+K0​ζℓ\+1,k−1≤a​e2​\(ℓ\+K0\)β−1\(k\+K0\)β≤a​e2​ℓ\+K0\(k\+K0\)2,w\_\{\\ell,k\}=\\frac\{a\}\{\\ell\+K\_\{0\}\}\\zeta\_\{\\ell\+1,k\-1\}\\leq ae^\{2\}\\frac\{\(\\ell\+K\_\{0\}\)^\{\\beta\-1\}\}\{\(k\+K\_\{0\}\)^\{\\beta\}\}\\leq ae^\{2\}\\frac\{\\ell\+K\_\{0\}\}\{\(k\+K\_\{0\}\)^\{2\}\},which proves \([55](https://arxiv.org/html/2606.26316#A2.E55)\)\. The initial product similarly satisfies

ζ0,k−1≤\(K0k\+K0\)β≤K0k\+K0,\\zeta\_\{0,k\-1\}\\leq\\left\(\\frac\{K\_\{0\}\}\{k\+K\_\{0\}\}\\right\)^\{\\beta\}\\leq\\frac\{K\_\{0\}\}\{k\+K\_\{0\}\},which proves \([53](https://arxiv.org/html/2606.26316#A2.E53)\)\.

For the first square sum, \([55](https://arxiv.org/html/2606.26316#A2.E55)\) yields

∑ℓ=0k−1wℓ,k2ℓ\+K0≤a2​e4\(k\+K0\)4​∑ℓ=0k−1\(ℓ\+K0\)≤a2​e4\(k\+K0\)4​\(k\+K0\)2\.\\sum\_\{\\ell=0\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}^\{2\}\}\{\\ell\+K\_\{0\}\}\\leq\\frac\{a^\{2\}e^\{4\}\}\{\(k\+K\_\{0\}\)^\{4\}\}\\sum\_\{\\ell=0\}^\{k\-1\}\(\\ell\+K\_\{0\}\)\\leq\\frac\{a^\{2\}e^\{4\}\}\{\(k\+K\_\{0\}\)^\{4\}\}\(k\+K\_\{0\}\)^\{2\}\.This is \([56](https://arxiv.org/html/2606.26316#A2.E56)\) after enlargingcwc\_\{w\}\. The second square sum follows from

∑ℓ=0k−1wℓ,k2\(ℓ\+K0\)2≤a2​e4\(k\+K0\)4​∑ℓ=0k−11≤a2​e4\(k\+K0\)3\.\\sum\_\{\\ell=0\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}^\{2\}\}\{\(\\ell\+K\_\{0\}\)^\{2\}\}\\leq\\frac\{a^\{2\}e^\{4\}\}\{\(k\+K\_\{0\}\)^\{4\}\}\\sum\_\{\\ell=0\}^\{k\-1\}1\\leq\\frac\{a^\{2\}e^\{4\}\}\{\(k\+K\_\{0\}\)^\{3\}\}\.Also,

∑ℓ=0k−1αℓ2​ζℓ\+1,k−1=∑ℓ=0k−1αℓ​wℓ,k≤a2​e2\(k\+K0\)2​∑ℓ=0k−11≤a2​e2k\+K0,\\sum\_\{\\ell=0\}^\{k\-1\}\\alpha\_\{\\ell\}^\{2\}\\zeta\_\{\\ell\+1,k\-1\}=\\sum\_\{\\ell=0\}^\{k\-1\}\\alpha\_\{\\ell\}w\_\{\\ell,k\}\\leq\\frac\{a^\{2\}e^\{2\}\}\{\(k\+K\_\{0\}\)^\{2\}\}\\sum\_\{\\ell=0\}^\{k\-1\}1\\leq\\frac\{a^\{2\}e^\{2\}\}\{k\+K\_\{0\}\},which proves \([54](https://arxiv.org/html/2606.26316#A2.E54)\)\. Finally, forℓ≥m\\ell\\geq m,

1ℓ−m\+K0=ℓ\+K0ℓ−m\+K0​1ℓ\+K0≤\(1\+mK0\)​1ℓ\+K0\.\\frac\{1\}\{\\ell\-m\+K\_\{0\}\}=\\frac\{\\ell\+K\_\{0\}\}\{\\ell\-m\+K\_\{0\}\}\\frac\{1\}\{\\ell\+K\_\{0\}\}\\leq\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\frac\{1\}\{\\ell\+K\_\{0\}\}\.Squaring the displayed inequality gives the corresponding squared denominator bound\. Combining these with \([56](https://arxiv.org/html/2606.26316#A2.E56)\) and \([57](https://arxiv.org/html/2606.26316#A2.E57)\) proves \([58](https://arxiv.org/html/2606.26316#A2.E58)\) and \([59](https://arxiv.org/html/2606.26316#A2.E59)\)\. ∎

###### Lemma 4\(Envelope and Lipschitz bounds for the descent observable\)\.

Let

r:=2​A​L\+B,r:=2AL\+B,\(60\)and, forD≥0D\\geq 0, define

H​\(D\):=2​L​D​\(r​D\+C\+2​L​D\)\.H\(D\):=\\sqrt\{2LD\}\\left\(\\sqrt\{rD\+C\}\+\\sqrt\{2LD\}\\right\)\.\(61\)There are constantscH,clip<∞c\_\{H\},c\_\{\\mathrm\{lip\}\}<\\infty, depending only onL,Lg,A,B,CL,L\_\{g\},A,B,C, such that:

1. \(i\)ifΔ​\(x\)≤D\\Delta\(x\)\\leq D, then\|h​\(x,z\)\|≤H​\(D\)\|h\(x,z\)\|\\leq H\(D\)for everyzz;
2. \(ii\)H​\(D\)2≤cH​\(D\+D2\)H\(D\)^\{2\}\\leq c\_\{H\}\(D\+D^\{2\}\)for everyD≥0D\\geq 0;
3. \(iii\)ifΔ​\(x\)≤Dx\\Delta\(x\)\\leq D\_\{x\}andΔ​\(y\)≤Dy\\Delta\(y\)\\leq D\_\{y\}, then \|h​\(x,z\)−h​\(y,z\)\|≤clip​\(1\+Dx\+Dy\)​‖x−y‖\.\|h\(x,z\)\-h\(y,z\)\|\\leq c\_\{\\mathrm\{lip\}\}\(1\+\\sqrt\{D\_\{x\}\}\+\\sqrt\{D\_\{y\}\}\)\\left\\lVert x\-y\\right\\rVert\.\(62\)

###### Proof\.

IfΔ​\(x\)≤D\\Delta\(x\)\\leq D, then \([5](https://arxiv.org/html/2606.26316#S3.E5)\) gives‖∇f​\(x\)‖≤2​L​D\\left\\lVert\\nabla f\(x\)\\right\\rVert\\leq\\sqrt\{2LD\}\. Moreover, \([8](https://arxiv.org/html/2606.26316#S3.E8)\) and \([5](https://arxiv.org/html/2606.26316#S3.E5)\) give

‖g​\(x,z\)‖≤A​‖∇f​\(x\)‖2\+B​Δ​\(x\)\+C≤r​D\+C\.\\left\\lVert g\(x,z\)\\right\\rVert\\leq\\sqrt\{A\\left\\lVert\\nabla f\(x\)\\right\\rVert^\{2\}\+B\\Delta\(x\)\+C\}\\leq\\sqrt\{rD\+C\}\.Thus

‖ξ​\(x,z\)‖≤‖g​\(x,z\)‖\+‖∇f​\(x\)‖≤r​D\+C\+2​L​D\.\\left\\lVert\\xi\(x,z\)\\right\\rVert\\leq\\left\\lVert g\(x,z\)\\right\\rVert\+\\left\\lVert\\nabla f\(x\)\\right\\rVert\\leq\\sqrt\{rD\+C\}\+\\sqrt\{2LD\}\.The Cauchy\-Schwarz inequality proves\|h​\(x,z\)\|≤H​\(D\)\|h\(x,z\)\|\\leq H\(D\)\.

Next,

H​\(D\)2\\displaystyle H\(D\)^\{2\}=2​L​D​\(r​D\+C\+2​L​D\)2\\displaystyle=2LD\\left\(\\sqrt\{rD\+C\}\+\\sqrt\{2LD\}\\right\)^\{2\}≤4​L​D​\(r​D\+C\+2​L​D\)≤cH​\(D\+D2\)\\displaystyle\\leq 4LD\(rD\+C\+2LD\)\\leq c\_\{H\}\(D\+D^\{2\}\)for a finitecHc\_\{H\}depending only onL,A,B,CL,A,B,C\.

For the Lipschitz bound, write

\|h​\(x,z\)−h​\(y,z\)\|\\displaystyle\|h\(x,z\)\-h\(y,z\)\|≤‖∇f​\(x\)−∇f​\(y\)‖​‖ξ​\(x,z\)‖\+‖∇f​\(y\)‖​‖ξ​\(x,z\)−ξ​\(y,z\)‖\.\\displaystyle\\leq\\left\\lVert\\nabla f\(x\)\-\\nabla f\(y\)\\right\\rVert\\left\\lVert\\xi\(x,z\)\\right\\rVert\+\\left\\lVert\\nabla f\(y\)\\right\\rVert\\left\\lVert\\xi\(x,z\)\-\\xi\(y,z\)\\right\\rVert\.ByLL\-smoothness,‖∇f​\(x\)−∇f​\(y\)‖≤L​‖x−y‖\\left\\lVert\\nabla f\(x\)\-\\nabla f\(y\)\\right\\rVert\\leq L\\left\\lVert x\-y\\right\\rVert\. The previous envelope gives‖ξ​\(x,z\)‖≤c​\(1\+Dx\)\\left\\lVert\\xi\(x,z\)\\right\\rVert\\leq c\(1\+\\sqrt\{D\_\{x\}\}\)\. Also,

‖ξ​\(x,z\)−ξ​\(y,z\)‖≤‖g​\(x,z\)−g​\(y,z\)‖\+‖∇f​\(x\)−∇f​\(y\)‖≤\(Lg\+L\)​‖x−y‖,\\left\\lVert\\xi\(x,z\)\-\\xi\(y,z\)\\right\\rVert\\leq\\left\\lVert g\(x,z\)\-g\(y,z\)\\right\\rVert\+\\left\\lVert\\nabla f\(x\)\-\\nabla f\(y\)\\right\\rVert\\leq\(L\_\{g\}\+L\)\\left\\lVert x\-y\\right\\rVert,and‖∇f​\(y\)‖≤2​L​Dy\\left\\lVert\\nabla f\(y\)\\right\\rVert\\leq\\sqrt\{2LD\_\{y\}\}\. Combining these estimates proves \([62](https://arxiv.org/html/2606.26316#A2.E62)\)\. ∎

## Appendix CLag\-blocking concentration estimates

This appendix proves the probabilistic estimates that control the stochastic terms in the weighted descent recursion of Lemma[2](https://arxiv.org/html/2606.26316#Thmlemma2)\. Lemma[5](https://arxiv.org/html/2606.26316#Thmlemma5)handles the martingale\-difference perturbation\. Lemmas[6](https://arxiv.org/html/2606.26316#Thmlemma6)–[9](https://arxiv.org/html/2606.26316#Thmlemma9)implement the lag\-blocking decomposition of the adaptive Markovian term into a delayed martingale part, a mixing\-bias part, a replacement error, and an initial\-window term\. These bounds are combined in Appendix[D](https://arxiv.org/html/2606.26316#A4)\.

The lemmas in this section are stated for a fixed target timekk\. LetN:=k\+K0N:=k\+K\_\{0\}and letm∈\{1,…,k\}m\\in\\\{1,\\ldots,k\\\}\. Let\(𝒦j\)j=0k−1\(\\mathcal\{K\}\_\{j\}\)\_\{j=0\}^\{k\-1\}be a nonincreasing family of events, meaning𝒦j\+1⊆𝒦j\\mathcal\{K\}\_\{j\+1\}\\subseteq\\mathcal\{K\}\_\{j\}, with𝒦j∈ℱj\\mathcal\{K\}\_\{j\}\\in\\mathcal\{F\}\_\{j\}\. Assume that on𝒦j\\mathcal\{K\}\_\{j\},

Δj≤Λj\+K0,0≤j≤k−1,\\Delta\_\{j\}\\leq\\frac\{\\Lambda\}\{j\+K\_\{0\}\},\\qquad 0\\leq j\\leq k\-1,\(63\)for some deterministicΛ≥1\\Lambda\\geq 1\.

###### Lemma 5\(Martingale\-difference noise\)\.

For everyη∈\(0,1\)\\eta\\in\(0,1\), with probability at least1−η1\-\\eta,

𝟏𝒦k−1​\|∑ℓ=0k−1wℓ,k​⟨∇f​\(xℓ\),Mℓ\+1⟩\|≤cMN​log⁡2η​\(Λ\+Λ2N\),\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}\\left\\langle\\nabla f\(x\_\{\\ell\}\),M\_\{\\ell\+1\}\\right\\rangle\\right\|\\leq\\frac\{c\_\{M\}\}\{N\}\\sqrt\{\\log\\frac\{2\}\{\\eta\}\\left\(\\Lambda\+\\frac\{\\Lambda^\{2\}\}\{N\}\\right\)\},\(64\)wherecM<∞c\_\{M\}<\\inftydepends only ona,μ,L,A,B,Ca,\\mu,L,A,B,C\.

###### Proof\.

Define

Qℓ\+1:=wℓ,k​⟨∇f​\(xℓ\),Mℓ\+1⟩​𝟏𝒦ℓ,0≤ℓ≤k−1\.Q\_\{\\ell\+1\}:=w\_\{\\ell,k\}\\left\\langle\\nabla f\(x\_\{\\ell\}\),M\_\{\\ell\+1\}\\right\\rangle\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{\\ell\}\},\\qquad 0\\leq\\ell\\leq k\-1\.Since𝒦ℓ∈ℱℓ\\mathcal\{K\}\_\{\\ell\}\\in\\mathcal\{F\}\_\{\\ell\},wℓ,kw\_\{\\ell,k\}is deterministic, and𝔼​\[Mℓ\+1∣ℱℓ\]=0\\mathbb\{E\}\[M\_\{\\ell\+1\}\\mid\\mathcal\{F\}\_\{\\ell\}\]=0, the sequence\(Qℓ\+1\)\(Q\_\{\\ell\+1\}\)is a martingale difference sequence\. On𝒦ℓ\\mathcal\{K\}\_\{\\ell\}, \([7](https://arxiv.org/html/2606.26316#S3.E7)\), \([8](https://arxiv.org/html/2606.26316#S3.E8)\), and \([5](https://arxiv.org/html/2606.26316#S3.E5)\) imply

‖Mℓ\+1‖≤‖Gℓ‖\+‖g​\(xℓ,Zℓ\)‖≤2​r​Λℓ\+K0\+C,\\left\\lVert M\_\{\\ell\+1\}\\right\\rVert\\leq\\left\\lVert G\_\{\\ell\}\\right\\rVert\+\\left\\lVert g\(x\_\{\\ell\},Z\_\{\\ell\}\)\\right\\rVert\\leq 2\\sqrt\{r\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\+C\},wherer=2​A​L\+Br=2AL\+B\. Also‖∇f​\(xℓ\)‖≤2​L​Λ/\(ℓ\+K0\)\\left\\lVert\\nabla f\(x\_\{\\ell\}\)\\right\\rVert\\leq\\sqrt\{2L\\Lambda/\(\\ell\+K\_\{0\}\)\}\. Hence

\|Qℓ\+1\|≤c​wℓ,k​Λℓ\+K0​1\+Λℓ\+K0\|Q\_\{\\ell\+1\}\|\\leq cw\_\{\\ell,k\}\\sqrt\{\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\}\\sqrt\{1\+\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\}for a constantccdepending only onL,A,B,CL,A,B,C\. Therefore

\|Qℓ\+1\|2≤c​wℓ,k2​\(Λℓ\+K0\+Λ2\(ℓ\+K0\)2\)\.\|Q\_\{\\ell\+1\}\|^\{2\}\\leq cw\_\{\\ell,k\}^\{2\}\\left\(\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\+\\frac\{\\Lambda^\{2\}\}\{\(\\ell\+K\_\{0\}\)^\{2\}\}\\right\)\.Define the deterministic bounds

bℓ,k2:=c​wℓ,k2​\(Λℓ\+K0\+Λ2\(ℓ\+K0\)2\),0≤ℓ≤k−1,b\_\{\\ell,k\}^\{2\}:=cw\_\{\\ell,k\}^\{2\}\\left\(\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\+\\frac\{\\Lambda^\{2\}\}\{\(\\ell\+K\_\{0\}\)^\{2\}\}\\right\),\\qquad 0\\leq\\ell\\leq k\-1,so that\|Qℓ\+1\|≤bℓ,k\|Q\_\{\\ell\+1\}\|\\leq b\_\{\\ell,k\}almost surely\. Using \([56](https://arxiv.org/html/2606.26316#A2.E56)\)–\([57](https://arxiv.org/html/2606.26316#A2.E57)\),

∑ℓ=0k−1bℓ,k2≤cN2​\(Λ\+Λ2N\)\.\\sum\_\{\\ell=0\}^\{k\-1\}b\_\{\\ell,k\}^\{2\}\\leq\\frac\{c\}\{N^\{2\}\}\\left\(\\Lambda\+\\frac\{\\Lambda^\{2\}\}\{N\}\\right\)\.The Azuma\-Hoeffding inequality applied with the deterministic increment boundsbℓ,kb\_\{\\ell,k\}gives

ℙ​\(\|∑ℓ=0k−1Qℓ\+1\|\>2​log⁡\(2/η\)​∑ℓ=0k−1bℓ,k2\)≤η\.\\mathbb\{P\}\\left\(\\left\|\\sum\_\{\\ell=0\}^\{k\-1\}Q\_\{\\ell\+1\}\\right\|\>\\sqrt\{2\\log\(2/\\eta\)\\sum\_\{\\ell=0\}^\{k\-1\}b\_\{\\ell,k\}^\{2\}\}\\right\)\\leq\\eta\.Combining the last two displays yields the bound in \([64](https://arxiv.org/html/2606.26316#A3.E64)\)\. On𝒦k−1\\mathcal\{K\}\_\{k\-1\}, monotonicity gives𝟏𝒦ℓ=1\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{\\ell\}\}=1for everyℓ≤k−1\\ell\\leq k\-1, so∑Qℓ\+1\\sum Q\_\{\\ell\+1\}equals the martingale term in \([64](https://arxiv.org/html/2606.26316#A3.E64)\)\. This proves the result after absorbing numerical constants intocMc\_\{M\}\. ∎

###### Lemma 6\(Delayed residue\-class martingale\)\.

Forℓ≥m\\ell\\geq m, define

Yℓ:=wℓ,k​h​\(xℓ−m,Zℓ\),Y¯ℓ:=𝔼​\[Yℓ∣ℱℓ−m\]\.Y\_\{\\ell\}:=w\_\{\\ell,k\}h\(x\_\{\\ell\-m\},Z\_\{\\ell\}\),\\qquad\\overline\{Y\}\_\{\\ell\}:=\\mathbb\{E\}\[Y\_\{\\ell\}\\mid\\mathcal\{F\}\_\{\\ell\-m\}\]\.\(65\)For everyη∈\(0,1\)\\eta\\in\(0,1\), with probability at least1−η1\-\\eta,

𝟏𝒦k−1​\|∑ℓ=mk−1\(Yℓ−Y¯ℓ\)\|\\displaystyle\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}\(Y\_\{\\ell\}\-\\overline\{Y\}\_\{\\ell\}\)\\right\|≤cBN​m​log⁡\(2​mη\)​\(1\+mK0\)​\(Λ\+Λ2N​\(1\+mK0\)\),\\displaystyle\\leq\\frac\{c\_\{B\}\}\{N\}\\sqrt\{m\\log\\left\(\\frac\{2m\}\{\\eta\}\\right\)\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\left\(\\Lambda\+\\frac\{\\Lambda^\{2\}\}\{N\}\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\right\)\},\(66\)wherecB<∞c\_\{B\}<\\inftydepends only ona,μ,L,A,B,Ca,\\mu,L,A,B,C\.

###### Proof\.

Forℓ≥m\\ell\\geq m, set

Dℓ:=\(Yℓ−Y¯ℓ\)​𝟏𝒦ℓ−m\.D\_\{\\ell\}:=\(Y\_\{\\ell\}\-\\overline\{Y\}\_\{\\ell\}\)\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{\\ell\-m\}\}\.Sincexℓ−mx\_\{\\ell\-m\}and𝟏𝒦ℓ−m\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{\\ell\-m\}\}areℱℓ−m\\mathcal\{F\}\_\{\\ell\-m\}\-measurable,

𝔼​\[Dℓ∣ℱℓ−m\]=0\.\\mathbb\{E\}\[D\_\{\\ell\}\\mid\\mathcal\{F\}\_\{\\ell\-m\}\]=0\.On𝒦ℓ−m\\mathcal\{K\}\_\{\\ell\-m\}, \([63](https://arxiv.org/html/2606.26316#A3.E63)\) and Lemma[4](https://arxiv.org/html/2606.26316#Thmlemma4)give

\|Yℓ\|≤wℓ,kH\(Λℓ−m\+K0\)=:bℓ\.\|Y\_\{\\ell\}\|\\leq w\_\{\\ell,k\}H\\left\(\\frac\{\\Lambda\}\{\\ell\-m\+K\_\{0\}\}\\right\)=:b\_\{\\ell\}\.The same bound holds for\|Y¯ℓ\|\|\\overline\{Y\}\_\{\\ell\}\|on𝒦ℓ−m\\mathcal\{K\}\_\{\\ell\-m\}, because conditional expectation is dominated by the conditional expectation of the absolute value\. Hence\|Dℓ\|≤2​bℓ\|D\_\{\\ell\}\|\\leq 2b\_\{\\ell\}\.

We now spell out the filtration structure, since this is the point at which the lag creates martingale differences\. Fix a residue classr∈\{0,1,…,m−1\}r\\in\\\{0,1,\\ldots,m\-1\\\}\. List the indicesℓ=r\+j​m\\ell=r\+jmthat lie in\{m,m\+1,…,k−1\}\\\{m,m\+1,\\ldots,k\-1\\\}in increasing order, sayℓ1<⋯<ℓs\\ell\_\{1\}<\\cdots<\\ell\_\{s\}\. Defineℋ0:=ℱℓ1−m\\mathcal\{H\}\_\{0\}:=\\mathcal\{F\}\_\{\\ell\_\{1\}\-m\}andℋi:=ℱℓi\\mathcal\{H\}\_\{i\}:=\\mathcal\{F\}\_\{\\ell\_\{i\}\}for1≤i≤s1\\leq i\\leq s\. For the first index,Dℓ1D\_\{\\ell\_\{1\}\}has conditional mean zero givenℋ0=ℱℓ1−m\\mathcal\{H\}\_\{0\}=\\mathcal\{F\}\_\{\\ell\_\{1\}\-m\}\. Fori≥2i\\geq 2, the residue\-class spacing givesℓi−m=ℓi−1\\ell\_\{i\}\-m=\\ell\_\{i\-1\}, and hence

𝔼​\[Dℓi∣ℋi−1\]=𝔼​\[Dℓi∣ℱℓi−m\]=0\.\\mathbb\{E\}\[D\_\{\\ell\_\{i\}\}\\mid\\mathcal\{H\}\_\{i\-1\}\]=\\mathbb\{E\}\[D\_\{\\ell\_\{i\}\}\\mid\\mathcal\{F\}\_\{\\ell\_\{i\}\-m\}\]=0\.Thus\(Dℓi\)i=1s\(D\_\{\\ell\_\{i\}\}\)\_\{i=1\}^\{s\}is a martingale\-difference sequence with respect to the down\-sampled filtration\(ℋi\)i=0s\(\\mathcal\{H\}\_\{i\}\)\_\{i=0\}^\{s\}\. No independence across different residue classes is used; each class is controlled separately, and the resulting high\-probability bounds are combined by a union bound\. Azuma\-Hoeffding gives

ℙ​\(\|∑i=1sDℓi\|\>8​u​∑i=1sbℓi2\)≤2​e−u\.\\mathbb\{P\}\\left\(\\left\|\\sum\_\{i=1\}^\{s\}D\_\{\\ell\_\{i\}\}\\right\|\>\\sqrt\{8u\\sum\_\{i=1\}^\{s\}b\_\{\\ell\_\{i\}\}^\{2\}\}\\right\)\\leq 2e^\{\-u\}\.Takeu=log⁡\(2​m/η\)u=\\log\(2m/\\eta\)and union bound over themmresidue classes\. With probability at least1−η1\-\\eta, all residue\-class inequalities hold\. On that event,

\|∑ℓ=mk−1Dℓ\|\\displaystyle\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}D\_\{\\ell\}\\right\|≤∑r=0m−1\|∑ℓ≡r​\(mod​m\)Dℓ\|\\displaystyle\\leq\\sum\_\{r=0\}^\{m\-1\}\\left\|\\sum\_\{\\ell\\equiv r\\,\(\\mathrm\{mod\}\\,m\)\}D\_\{\\ell\}\\right\|≤∑r=0m−18​u​∑ℓ≡r​\(mod​m\)bℓ2\\displaystyle\\leq\\sum\_\{r=0\}^\{m\-1\}\\sqrt\{8u\\sum\_\{\\ell\\equiv r\\,\(\\mathrm\{mod\}\\,m\)\}b\_\{\\ell\}^\{2\}\}≤8​u​m​∑ℓ=mk−1bℓ2,\\displaystyle\\leq\\sqrt\{8um\\sum\_\{\\ell=m\}^\{k\-1\}b\_\{\\ell\}^\{2\}\},where the last step is Cauchy\-Schwarz\.

It remains to bound the square sum\. Lemma[4](https://arxiv.org/html/2606.26316#Thmlemma4)gives

bℓ2≤cH​wℓ,k2​\(Λℓ−m\+K0\+Λ2\(ℓ−m\+K0\)2\)\.b\_\{\\ell\}^\{2\}\\leq c\_\{H\}w\_\{\\ell,k\}^\{2\}\\left\(\\frac\{\\Lambda\}\{\\ell\-m\+K\_\{0\}\}\+\\frac\{\\Lambda^\{2\}\}\{\(\\ell\-m\+K\_\{0\}\)^\{2\}\}\\right\)\.Using \([58](https://arxiv.org/html/2606.26316#A2.E58)\)–\([59](https://arxiv.org/html/2606.26316#A2.E59)\),

∑ℓ=mk−1bℓ2≤cN2​\(1\+mK0\)​Λ\+cN3​\(1\+mK0\)2​Λ2\.\\sum\_\{\\ell=m\}^\{k\-1\}b\_\{\\ell\}^\{2\}\\leq\\frac\{c\}\{N^\{2\}\}\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\Lambda\+\\frac\{c\}\{N^\{3\}\}\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)^\{2\}\\Lambda^\{2\}\.Substituting this estimate into the previous display proves \([66](https://arxiv.org/html/2606.26316#A3.E66)\)\. On𝒦k−1\\mathcal\{K\}\_\{k\-1\}, all indicators𝟏𝒦ℓ−m\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{\\ell\-m\}\}are one, so the sum ofDℓD\_\{\\ell\}is the displayed delayed centered sum\. ∎

###### Lemma 7\(Mixing bias\)\.

Assumem<km<kand setεm:=ε​\(m\)\\varepsilon\_\{m\}:=\\varepsilon\(m\), whereε​\(⋅\)\\varepsilon\(\\cdot\)is the uniform total\-variation mixing profile in Assumption[3](https://arxiv.org/html/2606.26316#Thmassumption3)\. Then, on𝒦k−1\\mathcal\{K\}\_\{k\-1\},

\|∑ℓ=mk−1Y¯ℓ\|≤cmix​εmN​N​\(1\+mK0\)​\(Λ\+Λ2N​\(1\+mK0\)\),\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}\\overline\{Y\}\_\{\\ell\}\\right\|\\leq\\frac\{c\_\{\\mathrm\{mix\}\}\\varepsilon\_\{m\}\}\{N\}\\sqrt\{N\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\left\(\\Lambda\+\\frac\{\\Lambda^\{2\}\}\{N\}\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\right\)\},\(67\)wherecmix<∞c\_\{\\mathrm\{mix\}\}<\\inftydepends only ona,μ,L,A,B,Ca,\\mu,L,A,B,C\.

###### Proof\.

Becausexℓ−mx\_\{\\ell\-m\}isℱℓ−m\\mathcal\{F\}\_\{\\ell\-m\}\-measurable and∫h​\(xℓ−m,z\)​π​\(d​z\)=0\\int h\(x\_\{\\ell\-m\},z\)\\pi\(dz\)=0, the full\-filtration Markov property gives

Y¯ℓ\\displaystyle\\overline\{Y\}\_\{\\ell\}=wℓ,k​∫h​\(xℓ−m,z\)​\(Pm​\(Zℓ−m,d​z\)−π​\(d​z\)\)\.\\displaystyle=w\_\{\\ell,k\}\\int h\(x\_\{\\ell\-m\},z\)\\bigl\(P^\{m\}\(Z\_\{\\ell\-m\},dz\)\-\\pi\(dz\)\\bigr\)\.On𝒦ℓ−m\\mathcal\{K\}\_\{\\ell\-m\}, Lemma[4](https://arxiv.org/html/2606.26316#Thmlemma4)and the dual total\-variation bound \([17](https://arxiv.org/html/2606.26316#S3.E17)\) imply

\|Y¯ℓ\|≤2​εm​wℓ,k​H​\(Λℓ−m\+K0\)=2​εm​bℓ\.\|\\overline\{Y\}\_\{\\ell\}\|\\leq 2\\varepsilon\_\{m\}w\_\{\\ell,k\}H\\left\(\\frac\{\\Lambda\}\{\\ell\-m\+K\_\{0\}\}\\right\)=2\\varepsilon\_\{m\}b\_\{\\ell\}\.Therefore, on𝒦k−1\\mathcal\{K\}\_\{k\-1\},

\|∑ℓ=mk−1Y¯ℓ\|≤2​εm​∑ℓ=mk−1bℓ≤2​εm​k​∑ℓ=mk−1bℓ2\.\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}\\overline\{Y\}\_\{\\ell\}\\right\|\\leq 2\\varepsilon\_\{m\}\\sum\_\{\\ell=m\}^\{k\-1\}b\_\{\\ell\}\\leq 2\\varepsilon\_\{m\}\\sqrt\{k\\sum\_\{\\ell=m\}^\{k\-1\}b\_\{\\ell\}^\{2\}\}\.The square\-sum estimate from the proof of Lemma[6](https://arxiv.org/html/2606.26316#Thmlemma6)andk≤Nk\\leq Ngive \([67](https://arxiv.org/html/2606.26316#A3.E67)\)\. ∎

###### Lemma 8\(Replacement error\)\.

On𝒦k−1\\mathcal\{K\}\_\{k\-1\},

\|∑ℓ=mk−1wℓ,k​\(h​\(xℓ,Zℓ\)−h​\(xℓ−m,Zℓ\)\)\|\\displaystyle\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}w\_\{\\ell,k\}\\bigl\(h\(x\_\{\\ell\},Z\_\{\\ell\}\)\-h\(x\_\{\\ell\-m\},Z\_\{\\ell\}\)\\bigr\)\\right\|\(68\)≤cR\[mN\+m2N2\(1\+logNK0\)\+m​ΛN2\(1\+logNK0\+mK0\)\\displaystyle\\quad\\leq c\_\{R\}\\Bigg\[\\frac\{m\}\{N\}\+\\frac\{m^\{2\}\}\{N^\{2\}\}\\left\(1\+\\log\\frac\{N\}\{K\_\{0\}\}\\right\)\+\\frac\{m\\Lambda\}\{N^\{2\}\}\\left\(1\+\\log\\frac\{N\}\{K\_\{0\}\}\+\\frac\{m\}\{K\_\{0\}\}\\right\)\+m​ΛN3/2\+m2​ΛK0​N2\],\\displaystyle\\hskip 136\.5733pt\+\\frac\{m\\sqrt\{\\Lambda\}\}\{N^\{3/2\}\}\+\\frac\{m^\{2\}\\sqrt\{\\Lambda\}\}\{\\sqrt\{K\_\{0\}\}N^\{2\}\}\\Bigg\],\(69\)wherecR<∞c\_\{R\}<\\inftydepends only ona,μ,L,Lg,A,B,Ca,\\mu,L,L\_\{g\},A,B,C\.

###### Proof\.

On𝒦j\\mathcal\{K\}\_\{j\}, the ABC bound and \([5](https://arxiv.org/html/2606.26316#S3.E5)\) imply

‖Gj‖≤r​Λj\+K0\+C≤cG​\(1\+Λj\+K0\)\\left\\lVert G\_\{j\}\\right\\rVert\\leq\\sqrt\{r\\frac\{\\Lambda\}\{j\+K\_\{0\}\}\+C\}\\leq c\_\{G\}\\left\(1\+\\sqrt\{\\frac\{\\Lambda\}\{j\+K\_\{0\}\}\}\\right\)for a constantcGc\_\{G\}depending only onL,A,B,CL,A,B,C\. Hence, forℓ≥m\\ell\\geq m,

‖xℓ−xℓ−m‖\\displaystyle\\left\\lVert x\_\{\\ell\}\-x\_\{\\ell\-m\}\\right\\rVert≤∑j=ℓ−mℓ−1αj​‖Gj‖\\displaystyle\\leq\\sum\_\{j=\\ell\-m\}^\{\\ell\-1\}\\alpha\_\{j\}\\left\\lVert G\_\{j\}\\right\\rVert\(70\)≤a​cG​∑j=ℓ−mℓ−11j\+K0​\(1\+Λj\+K0\)\\displaystyle\\leq ac\_\{G\}\\sum\_\{j=\\ell\-m\}^\{\\ell\-1\}\\frac\{1\}\{j\+K\_\{0\}\}\\left\(1\+\\sqrt\{\\frac\{\\Lambda\}\{j\+K\_\{0\}\}\}\\right\)≤c​mℓ−m\+K0​\(1\+Λℓ−m\+K0\)\.\\displaystyle\\leq c\\frac\{m\}\{\\ell\-m\+K\_\{0\}\}\\left\(1\+\\sqrt\{\\frac\{\\Lambda\}\{\\ell\-m\+K\_\{0\}\}\}\\right\)\.The last inequality uses thatj\+K0≥ℓ−m\+K0j\+K\_\{0\}\\geq\\ell\-m\+K\_\{0\}over the summation range\.

By Lemma[4](https://arxiv.org/html/2606.26316#Thmlemma4), \([63](https://arxiv.org/html/2606.26316#A3.E63)\), and \([70](https://arxiv.org/html/2606.26316#A3.E70)\),

\|h​\(xℓ,Zℓ\)−h​\(xℓ−m,Zℓ\)\|\\displaystyle\|h\(x\_\{\\ell\},Z\_\{\\ell\}\)\-h\(x\_\{\\ell\-m\},Z\_\{\\ell\}\)\|≤clip​\(1\+Λℓ\+K0\+Λℓ−m\+K0\)​‖xℓ−xℓ−m‖\\displaystyle\\quad\\leq c\_\{\\mathrm\{lip\}\}\\left\(1\+\\sqrt\{\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\}\+\\sqrt\{\\frac\{\\Lambda\}\{\\ell\-m\+K\_\{0\}\}\}\\right\)\\left\\lVert x\_\{\\ell\}\-x\_\{\\ell\-m\}\\right\\rVert≤c​mℓ−m\+K0​\(1\+Λℓ−m\+K0\+Λℓ−m\+K0\)\.\\displaystyle\\quad\\leq c\\frac\{m\}\{\\ell\-m\+K\_\{0\}\}\\left\(1\+\\sqrt\{\\frac\{\\Lambda\}\{\\ell\-m\+K\_\{0\}\}\}\+\\frac\{\\Lambda\}\{\\ell\-m\+K\_\{0\}\}\\right\)\.Thus

\|∑ℓ=mk−1wℓ,k​\(h​\(xℓ,Zℓ\)−h​\(xℓ−m,Zℓ\)\)\|\\displaystyle\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}w\_\{\\ell,k\}\\bigl\(h\(x\_\{\\ell\},Z\_\{\\ell\}\)\-h\(x\_\{\\ell\-m\},Z\_\{\\ell\}\)\\bigr\)\\right\|\(71\)≤c​m​∑ℓ=mk−1wℓ,k​\(1tℓ\+Λtℓ3/2\+Λtℓ2\),\\displaystyle\\quad\\leq cm\\sum\_\{\\ell=m\}^\{k\-1\}w\_\{\\ell,k\}\\left\(\\frac\{1\}\{t\_\{\\ell\}\}\+\\frac\{\\sqrt\{\\Lambda\}\}\{t\_\{\\ell\}^\{3/2\}\}\+\\frac\{\\Lambda\}\{t\_\{\\ell\}^\{2\}\}\\right\),wheretℓ:=ℓ−m\+K0t\_\{\\ell\}:=\\ell\-m\+K\_\{0\}\. Sinceℓ\+K0=tℓ\+m\\ell\+K\_\{0\}=t\_\{\\ell\}\+m, \([55](https://arxiv.org/html/2606.26316#A2.E55)\) gives

wℓ,k≤c​tℓ\+mN2\.w\_\{\\ell,k\}\\leq c\\frac\{t\_\{\\ell\}\+m\}\{N^\{2\}\}\.The elementary estimates

∑ℓ=mk−11\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}1≤k≤N,\\displaystyle\\leq k\\leq N,∑ℓ=mk−11tℓ\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{1\}\{t\_\{\\ell\}\}≤1\+log⁡NK0,\\displaystyle\\leq 1\+\\log\\frac\{N\}\{K\_\{0\}\},∑ℓ=mk−11tℓ2\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{1\}\{t\_\{\\ell\}^\{2\}\}≤1K0,\\displaystyle\\leq\\frac\{1\}\{K\_\{0\}\},∑ℓ=mk−11tℓ\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{1\}\{\\sqrt\{t\_\{\\ell\}\}\}≤2​N,\\displaystyle\\leq 2\\sqrt\{N\},∑ℓ=mk−11tℓ3/2\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{1\}\{t\_\{\\ell\}^\{3/2\}\}≤2K0\\displaystyle\\leq\\frac\{2\}\{\\sqrt\{K\_\{0\}\}\}therefore imply

∑ℓ=mk−1wℓ,ktℓ\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}\}\{t\_\{\\ell\}\}≤cN2​\(N\+m​\(1\+log⁡NK0\)\),\\displaystyle\\leq\\frac\{c\}\{N^\{2\}\}\\left\(N\+m\\left\(1\+\\log\\frac\{N\}\{K\_\{0\}\}\\right\)\\right\),∑ℓ=mk−1wℓ,ktℓ3/2\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}\}\{t\_\{\\ell\}^\{3/2\}\}≤cN2​\(N\+mK0\),\\displaystyle\\leq\\frac\{c\}\{N^\{2\}\}\\left\(\\sqrt\{N\}\+\\frac\{m\}\{\\sqrt\{K\_\{0\}\}\}\\right\),∑ℓ=mk−1wℓ,ktℓ2\\displaystyle\\sum\_\{\\ell=m\}^\{k\-1\}\\frac\{w\_\{\\ell,k\}\}\{t\_\{\\ell\}^\{2\}\}≤cN2​\(1\+log⁡NK0\+mK0\)\.\\displaystyle\\leq\\frac\{c\}\{N^\{2\}\}\\left\(1\+\\log\\frac\{N\}\{K\_\{0\}\}\+\\frac\{m\}\{K\_\{0\}\}\\right\)\.Substituting these three estimates into \([71](https://arxiv.org/html/2606.26316#A3.E71)\) proves \([68](https://arxiv.org/html/2606.26316#A3.E68)\)\. ∎

###### Lemma 9\(Initial window\)\.

For everym∈\{1,…,k\}m\\in\\\{1,\\ldots,k\\\}, on𝒦k−1\\mathcal\{K\}\_\{k\-1\},

\|∑ℓ=0m−1wℓ,k​h​\(xℓ,Zℓ\)\|≤cI​\(m​ΛN\+Λ​mN3/2\),\\left\|\\sum\_\{\\ell=0\}^\{m\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)\\right\|\\leq c\_\{I\}\\left\(\\frac\{\\sqrt\{m\\Lambda\}\}\{N\}\+\\frac\{\\Lambda\\sqrt\{m\}\}\{N^\{3/2\}\}\\right\),\(72\)wherecI<∞c\_\{I\}<\\inftydepends only ona,μ,L,A,B,Ca,\\mu,L,A,B,C\.

###### Proof\.

On𝒦ℓ\\mathcal\{K\}\_\{\\ell\}, Lemma[4](https://arxiv.org/html/2606.26316#Thmlemma4)gives

\|h​\(xℓ,Zℓ\)\|≤H​\(Λℓ\+K0\)\.\|h\(x\_\{\\ell\},Z\_\{\\ell\}\)\|\\leq H\\left\(\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\\right\)\.Because𝒦k−1⊆𝒦ℓ\\mathcal\{K\}\_\{k\-1\}\\subseteq\\mathcal\{K\}\_\{\\ell\}for everyℓ≤k−1\\ell\\leq k\-1, Cauchy\-Schwarz gives, on𝒦k−1\\mathcal\{K\}\_\{k\-1\},

∑ℓ=0m−1wℓ,k​\|h​\(xℓ,Zℓ\)\|\\displaystyle\\sum\_\{\\ell=0\}^\{m\-1\}w\_\{\\ell,k\}\|h\(x\_\{\\ell\},Z\_\{\\ell\}\)\|≤m​∑ℓ=0m−1wℓ,k2​H​\(Λℓ\+K0\)2\.\\displaystyle\\leq\\sqrt\{m\\sum\_\{\\ell=0\}^\{m\-1\}w\_\{\\ell,k\}^\{2\}H\\left\(\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\\right\)^\{2\}\}\.UsingH​\(D\)2≤cH​\(D\+D2\)H\(D\)^\{2\}\\leq c\_\{H\}\(D\+D^\{2\}\)and \([56](https://arxiv.org/html/2606.26316#A2.E56)\)–\([57](https://arxiv.org/html/2606.26316#A2.E57)\),

∑ℓ=0m−1wℓ,k2​H​\(Λℓ\+K0\)2\\displaystyle\\sum\_\{\\ell=0\}^\{m\-1\}w\_\{\\ell,k\}^\{2\}H\\left\(\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\\right\)^\{2\}≤c​∑ℓ=0k−1wℓ,k2​\(Λℓ\+K0\+Λ2\(ℓ\+K0\)2\)\\displaystyle\\leq c\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}^\{2\}\\left\(\\frac\{\\Lambda\}\{\\ell\+K\_\{0\}\}\+\\frac\{\\Lambda^\{2\}\}\{\(\\ell\+K\_\{0\}\)^\{2\}\}\\right\)≤c​\(ΛN2\+Λ2N3\)\.\\displaystyle\\leq c\\left\(\\frac\{\\Lambda\}\{N^\{2\}\}\+\\frac\{\\Lambda^\{2\}\}\{N^\{3\}\}\\right\)\.Taking square roots proves \([72](https://arxiv.org/html/2606.26316#A3.E72)\)\. ∎

## Appendix DProof of Theorem[1](https://arxiv.org/html/2606.26316#Thmtheorem1): mixing\-profile upper bound

###### Lemma 10\(One\-step induction estimate\)\.

There are constantsc0,c1,KD<∞c\_\{0\},c\_\{1\},K\_\{D\}<\\inftyand a numberρ0∈\(0,1\)\\rho\_\{0\}\\in\(0,1\), depending only ona,μ,L,Lg,A,B,Ca,\\mu,L,L\_\{g\},A,B,C, such that the following holds\. Fixk≥1k\\geq 1, letN=k\+K0N=k\+K\_\{0\}, and choosem∈\{1,…,k\}m\\in\\\{1,\\ldots,k\\\}\. Set

q:=1\+mK0,u:=log⁡\(16​\(m\+1\)​\(k\+1\)2δ\),Θ:=m​u​q2\.q:=1\+\\frac\{m\}\{K\_\{0\}\},\\qquad u:=\\log\\left\(\\frac\{16\(m\+1\)\(k\+1\)^\{2\}\}\{\\delta\}\\right\),\\qquad\\Theta:=muq^\{2\}\.AssumeK0≥KDK\_\{0\}\\geq K\_\{D\}and

ΘN≤ρ0,mN​\(1\+log⁡NK0\+mK0\)≤ρ0\.\\frac\{\\Theta\}\{N\}\\leq\\rho\_\{0\},\\qquad\\frac\{m\}\{N\}\\left\(1\+\\log\\frac\{N\}\{K\_\{0\}\}\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\leq\\rho\_\{0\}\.\(73\)Ifm<km<k, also assume

ε​\(m\)≤δ32​N4\.\\varepsilon\(m\)\\leq\\frac\{\\delta\}\{32N^\{4\}\}\.\(74\)LetΛ≥1\\Lambda\\geq 1, and suppose that the good\-event family\(𝒦j\)j=0k−1\(\\mathcal\{K\}\_\{j\}\)\_\{j=0\}^\{k\-1\}satisfies \([63](https://arxiv.org/html/2606.26316#A3.E63)\)\. Let

ηk:=δ8​\(k\+1\)2\.\\eta\_\{k\}:=\\frac\{\\delta\}\{8\(k\+1\)^\{2\}\}\.Then, outside an event of probability at most2​ηk2\\eta\_\{k\},

𝟏𝒦k−1\(\\displaystyle\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\Bigg\(\|∑ℓ=0k−1wℓ,k⟨∇f\(xℓ\),Mℓ\+1⟩\|\+\|∑ℓ=0k−1wℓ,kh\(xℓ,Zℓ\)\|\)\\displaystyle\\left\|\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}\\left\\langle\\nabla f\(x\_\{\\ell\}\),M\_\{\\ell\+1\}\\right\\rangle\\right\|\+\\left\|\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)\\right\|\\Bigg\)\(75\)≤Λ4​N\+c0​\(1\+Θ\)N\+c1N\.\\displaystyle\\leq\\frac\{\\Lambda\}\{4N\}\+\\frac\{c\_\{0\}\(1\+\\Theta\)\}\{N\}\+\\frac\{c\_\{1\}\}\{N\}\.

###### Proof\.

Throughout the proof,ccdenotes a finite constant depending only ona,μ,L,Lg,A,B,Ca,\\mu,L,L\_\{g\},A,B,C; its value may change from line to line\. The goal is to prove a single\-time bound at the fixed terminal timekk\. The good\-event indicators are included only to make all increment bounds deterministic; on𝒦k−1\\mathcal\{K\}\_\{k\-1\}all earlier good events also hold\.

#### Martingale\-difference perturbation\.

The martingale part is controlled by[Lemma˜5](https://arxiv.org/html/2606.26316#Thmlemma5)\. Sincem≥1m\\geq 1,q≥1q\\geq 1, andΘ=m​u​q2\\Theta=muq^\{2\}, we haveΘ≥u\\Theta\\geq u\. Also

log⁡2ηk=log⁡16​\(k\+1\)2δ≤log⁡16​\(m\+1\)​\(k\+1\)2δ=u\.\\log\\frac\{2\}\{\\eta\_\{k\}\}=\\log\\frac\{16\(k\+1\)^\{2\}\}\{\\delta\}\\leq\\log\\frac\{16\(m\+1\)\(k\+1\)^\{2\}\}\{\\delta\}=u\.Therefore, outside an event of probability at mostηk\\eta\_\{k\},

𝟏𝒦k−1​\|∑ℓ=0k−1wℓ,k​⟨∇f​\(xℓ\),Mℓ\+1⟩\|\\displaystyle\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}\\left\\langle\\nabla f\(x\_\{\\ell\}\),M\_\{\\ell\+1\}\\right\\rangle\\right\|≤cN​u​\(Λ\+Λ2N\)\\displaystyle\\leq\\frac\{c\}\{N\}\\sqrt\{u\\left\(\\Lambda\+\\frac\{\\Lambda^\{2\}\}\{N\}\\right\)\}\(76\)≤cN​\(1\+Θ\)​Λ\+c​ΛN​1\+ΘN\.\\displaystyle\\leq\\frac\{c\}\{N\}\\sqrt\{\(1\+\\Theta\)\\Lambda\}\+\\frac\{c\\Lambda\}\{N\}\\sqrt\{\\frac\{1\+\\Theta\}\{N\}\}\.For the first term we use Young’s inequality in the formc​\(1\+Θ\)​Λ≤Λ/32\+c​\(1\+Θ\)c\\sqrt\{\(1\+\\Theta\)\\Lambda\}\\leq\\Lambda/32\+c\(1\+\\Theta\)\. For the second term,

1\+ΘN≤1K0\+ΘN\.\\sqrt\{\\frac\{1\+\\Theta\}\{N\}\}\\leq\\sqrt\{\\frac\{1\}\{K\_\{0\}\}\+\\frac\{\\Theta\}\{N\}\}\.By increasingKDK\_\{D\}and then choosingρ0\\rho\_\{0\}sufficiently small, the conditionK0≥KDK\_\{0\}\\geq K\_\{D\}andΘ/N≤ρ0\\Theta/N\\leq\\rho\_\{0\}make the last display small enough that

c​ΛN​1\+ΘN≤Λ32​N\+cN\.\\frac\{c\\Lambda\}\{N\}\\sqrt\{\\frac\{1\+\\Theta\}\{N\}\}\\leq\\frac\{\\Lambda\}\{32N\}\+\\frac\{c\}\{N\}\.Consequently the martingale\-difference contribution satisfies

𝟏𝒦k−1​\|∑ℓ=0k−1wℓ,k​⟨∇f​\(xℓ\),Mℓ\+1⟩\|≤Λ16​N\+c​\(1\+Θ\)N\+cN\.\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}\\left\\langle\\nabla f\(x\_\{\\ell\}\),M\_\{\\ell\+1\}\\right\\rangle\\right\|\\leq\\frac\{\\Lambda\}\{16N\}\+\\frac\{c\(1\+\\Theta\)\}\{N\}\+\\frac\{c\}\{N\}\.\(77\)

#### Decomposition of the Markovian sum\.

Ifm<km<k, defineYℓY\_\{\\ell\}andY¯ℓ\\overline\{Y\}\_\{\\ell\}by \([65](https://arxiv.org/html/2606.26316#A3.E65)\)\. Then

∑ℓ=0k−1wℓ,k​h​\(xℓ,Zℓ\)=∑ℓ=0m−1wℓ,k​h​\(xℓ,Zℓ\)\+∑ℓ=mk−1\(Yℓ−Y¯ℓ\)\+∑ℓ=mk−1Y¯ℓ\+∑ℓ=mk−1wℓ,k​\{h​\(xℓ,Zℓ\)−h​\(xℓ−m,Zℓ\)\}\.\\sum\_\{\\ell=0\}^\{k\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)=\\sum\_\{\\ell=0\}^\{m\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)\+\\sum\_\{\\ell=m\}^\{k\-1\}\(Y\_\{\\ell\}\-\\overline\{Y\}\_\{\\ell\}\)\+\\sum\_\{\\ell=m\}^\{k\-1\}\\overline\{Y\}\_\{\\ell\}\+\\sum\_\{\\ell=m\}^\{k\-1\}w\_\{\\ell,k\}\\\{h\(x\_\{\\ell\},Z\_\{\\ell\}\)\-h\(x\_\{\\ell\-m\},Z\_\{\\ell\}\)\\\}\.Whenm=km=k, the three sums overℓ=m,…,k−1\\ell=m,\\ldots,k\-1are empty, so only the initial\-window term remains\.

#### Centered delayed term\.

Assume first thatm<km<k\. Since

log⁡2​mηk=log⁡16​m​\(k\+1\)2δ≤u,\\log\\frac\{2m\}\{\\eta\_\{k\}\}=\\log\\frac\{16m\(k\+1\)^\{2\}\}\{\\delta\}\\leq u,[Lemma˜6](https://arxiv.org/html/2606.26316#Thmlemma6)gives, outside another event of probability at mostηk\\eta\_\{k\},

𝟏𝒦k−1​\|∑ℓ=mk−1\(Yℓ−Y¯ℓ\)\|\\displaystyle\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}\(Y\_\{\\ell\}\-\\overline\{Y\}\_\{\\ell\}\)\\right\|≤cN​m​u​q​\(Λ\+Λ2​qN\)\\displaystyle\\leq\\frac\{c\}\{N\}\\sqrt\{muq\\left\(\\Lambda\+\\frac\{\\Lambda^\{2\}q\}\{N\}\\right\)\}\(78\)≤cN​Θ​Λ\+c​ΛN​ΘN\.\\displaystyle\\leq\\frac\{c\}\{N\}\\sqrt\{\\Theta\\Lambda\}\+c\\frac\{\\Lambda\}\{N\}\\sqrt\{\\frac\{\\Theta\}\{N\}\}\.Here we usedq≥1q\\geq 1,m​u​q≤m​u​q2=Θmuq\\leq muq^\{2\}=\\Theta, andm​u​q⋅q=m​u​q2=Θmuq\\cdot q=muq^\{2\}=\\Theta\. The first term is bounded byΛ/\(64​N\)\+c​Θ/N\\Lambda/\(64N\)\+c\\Theta/Nby Young’s inequality\. The second term is at mostΛ/\(64​N\)\\Lambda/\(64N\)after decreasingρ0\\rho\_\{0\}, becauseΘ/N≤ρ0\\Theta/N\\leq\\rho\_\{0\}\. Thus the centered delayed term is bounded by

𝟏𝒦k−1​\|∑ℓ=mk−1\(Yℓ−Y¯ℓ\)\|≤Λ32​N\+c​ΘN\.\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}\(Y\_\{\\ell\}\-\\overline\{Y\}\_\{\\ell\}\)\\right\|\\leq\\frac\{\\Lambda\}\{32N\}\+\\frac\{c\\Theta\}\{N\}\.\(79\)

#### Mixing bias\.

By[Lemma˜7](https://arxiv.org/html/2606.26316#Thmlemma7)and \([74](https://arxiv.org/html/2606.26316#A4.E74)\), on𝒦k−1\\mathcal\{K\}\_\{k\-1\},

\|∑ℓ=mk−1Y¯ℓ\|≤c​δN5​N​q​\(Λ\+Λ2​qN\)\.\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}\\overline\{Y\}\_\{\\ell\}\\right\|\\leq\\frac\{c\\delta\}\{N^\{5\}\}\\sqrt\{Nq\\left\(\\Lambda\+\\frac\{\\Lambda^\{2\}q\}\{N\}\\right\)\}\.Sinceδ<1\\delta<1,q=1\+m/K0≤1\+k/K0≤Nq=1\+m/K\_\{0\}\\leq 1\+k/K\_\{0\}\\leq N, andΛ≥1\\Lambda\\geq 1, the right\-hand side is at most

cN5​\{N​Λ\+N​Λ\}≤c​\(1\+Λ\)N4\.\\frac\{c\}\{N^\{5\}\}\\\{N\\sqrt\{\\Lambda\}\+N\\Lambda\\\}\\leq\\frac\{c\(1\+\\Lambda\)\}\{N^\{4\}\}\.IncreasingKDK\_\{D\}if necessary gives, for allN≥KDN\\geq K\_\{D\},

𝟏𝒦k−1​\|∑ℓ=mk−1Y¯ℓ\|≤Λ64​N\+cN\.\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}\\overline\{Y\}\_\{\\ell\}\\right\|\\leq\\frac\{\\Lambda\}\{64N\}\+\\frac\{c\}\{N\}\.\(80\)

#### Replacement term\.

We now simplify the deterministic estimate in[Lemma˜8](https://arxiv.org/html/2606.26316#Thmlemma8)\. The first term in \([68](https://arxiv.org/html/2606.26316#A3.E68)\) is bounded byc​m/N≤c​Θ/Ncm/N\\leq c\\Theta/N, sinceu,q≥1u,q\\geq 1\. For the second term, use

1\+log⁡NK0≤c​log⁡\{16​\(m\+1\)​\(k\+1\)2/δ\}=c​u≤c​Θm,1\+\\log\\frac\{N\}\{K\_\{0\}\}\\leq c\\log\\\{16\(m\+1\)\(k\+1\)^\{2\}/\\delta\\\}=cu\\leq c\\frac\{\\Theta\}\{m\},which gives

m2N2​\(1\+log⁡NK0\)≤c​m​ΘN2≤c​ΘN\.\\frac\{m^\{2\}\}\{N^\{2\}\}\\left\(1\+\\log\\frac\{N\}\{K\_\{0\}\}\\right\)\\leq c\\frac\{m\\Theta\}\{N^\{2\}\}\\leq c\\frac\{\\Theta\}\{N\}\.TheΛ\\Lambda\-linear replacement term is controlled directly by the second small\-window condition:

c​m​ΛN2​\(1\+log⁡NK0\+mK0\)≤c​ρ0​ΛN≤Λ64​N,c\\frac\{m\\Lambda\}\{N^\{2\}\}\\left\(1\+\\log\\frac\{N\}\{K\_\{0\}\}\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\leq c\\rho\_\{0\}\\frac\{\\Lambda\}\{N\}\\leq\\frac\{\\Lambda\}\{64N\},after decreasingρ0\\rho\_\{0\}\. The termc​m​Λ/N3/2cm\\sqrt\{\\Lambda\}/N^\{3/2\}satisfies

c​m​ΛN3/2≤Λ128​N\+c​m2N2≤Λ128​N\+c​ΘN\.c\\frac\{m\\sqrt\{\\Lambda\}\}\{N^\{3/2\}\}\\leq\\frac\{\\Lambda\}\{128N\}\+c\\frac\{m^\{2\}\}\{N^\{2\}\}\\leq\\frac\{\\Lambda\}\{128N\}\+c\\frac\{\\Theta\}\{N\}\.For the final term, Young’s inequality gives

c​m2​ΛK0​N2≤Λ128​N\+c​m4K0​N3\.c\\frac\{m^\{2\}\\sqrt\{\\Lambda\}\}\{\\sqrt\{K\_\{0\}\}N^\{2\}\}\\leq\\frac\{\\Lambda\}\{128N\}\+c\\frac\{m^\{4\}\}\{K\_\{0\}N^\{3\}\}\.Becausem≤Nm\\leq Nand

m2K0=m​mK0≤m​\(1\+mK0\)≤m​q≤m​q2≤Θ,\\frac\{m^\{2\}\}\{K\_\{0\}\}=m\\frac\{m\}\{K\_\{0\}\}\\leq m\\left\(1\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\leq mq\\leq mq^\{2\}\\leq\\Theta,we havem4/\(K0​N3\)≤Θ/Nm^\{4\}/\(K\_\{0\}N^\{3\}\)\\leq\\Theta/N\. Therefore

𝟏𝒦k−1​\|∑ℓ=mk−1wℓ,k​\{h​\(xℓ,Zℓ\)−h​\(xℓ−m,Zℓ\)\}\|≤Λ32​N\+c​ΘN\.\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=m\}^\{k\-1\}w\_\{\\ell,k\}\\\{h\(x\_\{\\ell\},Z\_\{\\ell\}\)\-h\(x\_\{\\ell\-m\},Z\_\{\\ell\}\)\\\}\\right\|\\leq\\frac\{\\Lambda\}\{32N\}\+\\frac\{c\\Theta\}\{N\}\.\(81\)

#### Initial window\.

Finally,[Lemma˜9](https://arxiv.org/html/2606.26316#Thmlemma9)gives

𝟏𝒦k−1​\|∑ℓ=0m−1wℓ,k​h​\(xℓ,Zℓ\)\|≤c​m​ΛN\+c​ΛN​mN\.\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=0\}^\{m\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)\\right\|\\leq c\\frac\{\\sqrt\{m\\Lambda\}\}\{N\}\+c\\frac\{\\Lambda\}\{N\}\\sqrt\{\\frac\{m\}\{N\}\}\.The first term is at mostΛ/\(64​N\)\+c​m/N≤Λ/\(64​N\)\+c​Θ/N\\Lambda/\(64N\)\+cm/N\\leq\\Lambda/\(64N\)\+c\\Theta/N\. The second term is at mostΛ/\(64​N\)\\Lambda/\(64N\)after decreasingρ0\\rho\_\{0\}, becausem/N≤Θ/N≤ρ0m/N\\leq\\Theta/N\\leq\\rho\_\{0\}\. Hence the initial window is bounded by

𝟏𝒦k−1​\|∑ℓ=0m−1wℓ,k​h​\(xℓ,Zℓ\)\|≤Λ32​N\+c​ΘN\.\\mathbf\{1\}\_\{\\mathcal\{K\}\_\{k\-1\}\}\\left\|\\sum\_\{\\ell=0\}^\{m\-1\}w\_\{\\ell,k\}h\(x\_\{\\ell\},Z\_\{\\ell\}\)\\right\|\\leq\\frac\{\\Lambda\}\{32N\}\+\\frac\{c\\Theta\}\{N\}\.\(82\)
Before combining the estimates, we record where the two admissibility requirements are used\. The condition

Θ/N≤ρ0\\Theta/N\\leq\\rho\_\{0\}is used to absorb the stochastic square\-root terms in \([76](https://arxiv.org/html/2606.26316#A4.E76)\), \([78](https://arxiv.org/html/2606.26316#A4.E78)\), and \([82](https://arxiv.org/html/2606.26316#A4.E82)\) into a small multiple ofΛ/N\\Lambda/N, with the remainder of order\(1\+Θ\)/N\(1\+\\Theta\)/N\. The second condition,

mN​\(1\+log⁡NK0\+mK0\)≤ρ0,\\frac\{m\}\{N\}\\left\(1\+\\log\\frac\{N\}\{K\_\{0\}\}\+\\frac\{m\}\{K\_\{0\}\}\\right\)\\leq\\rho\_\{0\},is used for the replacement term, where the movement of the iterate over one lag window produces the additional factor involvingmm,K0K\_\{0\}, andlog⁡\(N/K0\)\\log\(N/K\_\{0\}\)\. Thus the first admissibility condition controls concentration size, while the second controls the deterministic lag\-replacement drift\.

Combining \([79](https://arxiv.org/html/2606.26316#A4.E79)\), \([80](https://arxiv.org/html/2606.26316#A4.E80)\), \([81](https://arxiv.org/html/2606.26316#A4.E81)\), and \([82](https://arxiv.org/html/2606.26316#A4.E82)\) controls the Markovian sum\. Ifm=km=k, the delayed centered, mixing\-bias, and replacement bounds are omitted and the same conclusion follows from \([82](https://arxiv.org/html/2606.26316#A4.E82)\)\. Adding the martingale bound \([77](https://arxiv.org/html/2606.26316#A4.E77)\) proves \([75](https://arxiv.org/html/2606.26316#A4.E75)\) after enlargingc0c\_\{0\}andc1c\_\{1\}\. The only probabilistic failures used are the martingale\-difference event and the delayed martingale event, each of probability at mostηk\\eta\_\{k\}, so the total failure probability is at most2​ηk2\\eta\_\{k\}\. ∎

###### Proof of[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)\.

Letρ0\\rho\_\{0\}andKDK\_\{D\}be the constants from[Lemma˜10](https://arxiv.org/html/2606.26316#Thmlemma10)\. IncreaseKDK\_\{D\}if necessary so thatKD≥KbaseK\_\{D\}\\geq K\_\{\\mathrm\{base\}\}\. LetΓ≥1\\Gamma\\geq 1be chosen below and define

Λk:=2​K0​Δ0\+Γ​\(1\+Θk\)\.\\Lambda\_\{k\}:=2K\_\{0\}\\Delta\_\{0\}\+\\Gamma\(1\+\\Theta\_\{k\}\)\.\(83\)By the monotonicity noted after \([19](https://arxiv.org/html/2606.26316#S4.E19)\), the sequencesmk,uk,qk,Θkm\_\{k\},u\_\{k\},q\_\{k\},\\Theta\_\{k\}are nondecreasing inkk, soΛk\\Lambda\_\{k\}is nondecreasing\.

Define the single\-time and cumulative good events by

𝖤k:=\{Δk≤Λkk\+K0\},𝖤0:=Ω,ℰk:=⋂j=0k𝖤j\.\\mathsf\{E\}\_\{k\}:=\\left\\\{\\Delta\_\{k\}\\leq\\frac\{\\Lambda\_\{k\}\}\{k\+K\_\{0\}\}\\right\\\},\\qquad\\mathsf\{E\}\_\{0\}:=\\Omega,\\qquad\\mathcal\{E\}\_\{k\}:=\\bigcap\_\{j=0\}^\{k\}\\mathsf\{E\}\_\{j\}\.\(84\)For target timek≥1k\\geq 1, onℰk−1\\mathcal\{E\}\_\{k\-1\}we have the envelope required in \([63](https://arxiv.org/html/2606.26316#A3.E63)\) withΛ=Λk\\Lambda=\\Lambda\_\{k\}: forj=0j=0,

Δ0≤Λk/K0,\\Delta\_\{0\}\\leq\\Lambda\_\{k\}/K\_\{0\},becauseΛk≥2​K0​Δ0\\Lambda\_\{k\}\\geq 2K\_\{0\}\\Delta\_\{0\}, and for1≤j≤k−11\\leq j\\leq k\-1,

Δj≤Λjj\+K0≤Λkj\+K0\.\\Delta\_\{j\}\\leq\\frac\{\\Lambda\_\{j\}\}\{j\+K\_\{0\}\}\\leq\\frac\{\\Lambda\_\{k\}\}\{j\+K\_\{0\}\}\.Thus[Lemma˜10](https://arxiv.org/html/2606.26316#Thmlemma10)applies withm=mkm=m\_\{k\}and𝒦j=ℰj\\mathcal\{K\}\_\{j\}=\\mathcal\{E\}\_\{j\}\. Ifmk<km\_\{k\}<k, then by definition ofmkm\_\{k\}it equalsm^k\\widehat\{m\}\_\{k\}, and \([74](https://arxiv.org/html/2606.26316#A4.E74)\) holds\. The admissibility assumptions are exactly \([73](https://arxiv.org/html/2606.26316#A4.E73)\)\.

By[Lemma˜2](https://arxiv.org/html/2606.26316#Thmlemma2),[Lemma˜3](https://arxiv.org/html/2606.26316#Thmlemma3), and[Lemma˜10](https://arxiv.org/html/2606.26316#Thmlemma10), outside an event of probability at most2​ηk2\\eta\_\{k\},

𝟏ℰk−1​Δk\\displaystyle\\mathbf\{1\}\_\{\\mathcal\{E\}\_\{k\-1\}\}\\Delta\_\{k\}≤K0​Δ0\+cdetNk\+Λk4​Nk\+c​\(1\+Θk\)Nk,\\displaystyle\\leq\\frac\{K\_\{0\}\\Delta\_\{0\}\+c\_\{\\mathrm\{det\}\}\}\{N\_\{k\}\}\+\\frac\{\\Lambda\_\{k\}\}\{4N\_\{k\}\}\+\\frac\{c\(1\+\\Theta\_\{k\}\)\}\{N\_\{k\}\},\(85\)wherecdet=L​C​cw/2c\_\{\\mathrm\{det\}\}=LCc\_\{w\}/2andc<∞c<\\inftydepends only on the model constants\. Choose

Γ≥4​\(c\+cdet\+1\)\.\\Gamma\\geq 4\(c\+c\_\{\\mathrm\{det\}\}\+1\)\.\(86\)Then the numerator on the right side of \([85](https://arxiv.org/html/2606.26316#A4.E85)\) is bounded by

K0​Δ0\+cdet\+c​\(1\+Θk\)\+14​\{2​K0​Δ0\+Γ​\(1\+Θk\)\}\\displaystyle K\_\{0\}\\Delta\_\{0\}\+c\_\{\\mathrm\{det\}\}\+c\(1\+\\Theta\_\{k\}\)\+\\frac\{1\}\{4\}\\\{2K\_\{0\}\\Delta\_\{0\}\+\\Gamma\(1\+\\Theta\_\{k\}\)\\\}≤32​K0​Δ0\+12​Γ​\(1\+Θk\)\\displaystyle\\leq\\frac\{3\}\{2\}K\_\{0\}\\Delta\_\{0\}\+\\frac\{1\}\{2\}\\Gamma\(1\+\\Theta\_\{k\}\)≤2​K0​Δ0\+Γ​\(1\+Θk\)=Λk\.\\displaystyle\\leq 2K\_\{0\}\\Delta\_\{0\}\+\\Gamma\(1\+\\Theta\_\{k\}\)=\\Lambda\_\{k\}\.Therefore

ℙ​\(ℰk−1∩𝖤kc\)≤2​ηk=δ4​\(k\+1\)2\.\\mathbb\{P\}\(\\mathcal\{E\}\_\{k\-1\}\\cap\\mathsf\{E\}\_\{k\}^\{c\}\)\\leq 2\\eta\_\{k\}=\\frac\{\\delta\}\{4\(k\+1\)^\{2\}\}\.The first\-failure union bound gives

ℙ​\(⋃k≥1𝖤kc\)\\displaystyle\\mathbb\{P\}\\left\(\\bigcup\_\{k\\geq 1\}\\mathsf\{E\}\_\{k\}^\{c\}\\right\)=ℙ​\(⋃k≥1\(ℰk−1∩𝖤kc\)\)\\displaystyle=\\mathbb\{P\}\\left\(\\bigcup\_\{k\\geq 1\}\(\\mathcal\{E\}\_\{k\-1\}\\cap\\mathsf\{E\}\_\{k\}^\{c\}\)\\right\)≤∑k=1∞δ4​\(k\+1\)2≤δ\.\\displaystyle\\leq\\sum\_\{k=1\}^\{\\infty\}\\frac\{\\delta\}\{4\(k\+1\)^\{2\}\}\\leq\\delta\.Thus𝖤k\\mathsf\{E\}\_\{k\}holds for allk≥1k\\geq 1with probability at least1−δ1\-\\delta\. TakingC⋆=ΓC\_\{\\star\}=\\Gammaproves \([22](https://arxiv.org/html/2606.26316#S4.E22)\)\. ∎

## Appendix EProof of Corollary[1](https://arxiv.org/html/2606.26316#Thmcorollary1): geometric window calculus

###### Lemma 11\(A logarithm\-over\-linear bound\)\.

Letr≥1r\\geq 1be an integer\. Letb,s\>0b,s\>0and assume

1\+log⁡\(b/s\)≥r\.1\+\\log\(b/s\)\\geq r\.\(87\)Then

supN≥b\(1\+log⁡\(N/s\)\)rN=\(1\+log⁡\(b/s\)\)rb\.\\sup\_\{N\\geq b\}\\frac\{\(1\+\\log\(N/s\)\)^\{r\}\}\{N\}=\\frac\{\(1\+\\log\(b/s\)\)^\{r\}\}\{b\}\.\(88\)

###### Proof\.

Sety​\(N\)=1\+log⁡\(N/s\)y\(N\)=1\+\\log\(N/s\)\. Under \([87](https://arxiv.org/html/2606.26316#A5.E87)\),y​\(N\)≥ry\(N\)\\geq rfor everyN≥bN\\geq b\. For

ϕ​\(N\)=y​\(N\)rN\\phi\(N\)=\\frac\{y\(N\)^\{r\}\}\{N\}we have

dd​N​log⁡ϕ​\(N\)=rN​y​\(N\)−1N=1N​\(ry​\(N\)−1\)≤0\.\\frac\{d\}\{dN\}\\log\\phi\(N\)=\\frac\{r\}\{Ny\(N\)\}\-\\frac\{1\}\{N\}=\\frac\{1\}\{N\}\\left\(\\frac\{r\}\{y\(N\)\}\-1\\right\)\\leq 0\.Thusϕ\\phiis nonincreasing on\[b,∞\)\[b,\\infty\), and the supremum is attained atN=bN=b\. ∎

###### Lemma 12\(Geometric windows are admissible\)\.

Fixρ∈\(0,1\)\\rho\\in\(0,1\)\. There existsKwin​\(ρ\)<∞K\_\{\\mathrm\{win\}\}\(\\rho\)<\\inftysuch that the following holds\. Letδ∈\(0,e−1\)\\delta\\in\(0,e^\{\-1\}\), lettmix≥1t\_\{\\mathrm\{mix\}\}\\geq 1, and assume \([12](https://arxiv.org/html/2606.26316#S3.E12)\)\. If

K0≥Kwin​\(ρ\)​tmix​\(1\+log⁡e​tmixδ\)3,K\_\{0\}\\geq K\_\{\\mathrm\{win\}\}\(\\rho\)\\,t\_\{\\mathrm\{mix\}\}\\left\(1\+\\log\\frac\{et\_\{\\mathrm\{mix\}\}\}\{\\delta\}\\right\)^\{3\},\(89\)then the windows \([18](https://arxiv.org/html/2606.26316#S4.E18)\)\-\([19](https://arxiv.org/html/2606.26316#S4.E19)\) satisfy, for everyk≥1k\\geq 1,

ΘkNk\\displaystyle\\frac\{\\Theta\_\{k\}\}\{N\_\{k\}\}≤ρ,\\displaystyle\\leq\\rho,\(90\)mkNk​\(1\+log⁡NkK0\+mkK0\)\\displaystyle\\frac\{m\_\{k\}\}\{N\_\{k\}\}\\left\(1\+\\log\\frac\{N\_\{k\}\}\{K\_\{0\}\}\+\\frac\{m\_\{k\}\}\{K\_\{0\}\}\\right\)≤ρ\.\\displaystyle\\leq\\rho\.\(91\)

###### Proof\.

Let

ℓ0:=1\+log⁡e​tmixδ,yk:=1\+log⁡Nkδ\.\\ell\_\{0\}:=1\+\\log\\frac\{et\_\{\\mathrm\{mix\}\}\}\{\\delta\},\\qquad y\_\{k\}:=1\+\\log\\frac\{N\_\{k\}\}\{\\delta\}\.Under \([12](https://arxiv.org/html/2606.26316#S3.E12)\), the minimal window in \([18](https://arxiv.org/html/2606.26316#S4.E18)\) satisfies

mk≤c​tmix​ykm\_\{k\}\\leq ct\_\{\\mathrm\{mix\}\}y\_\{k\}\(92\)for a numerical constantcc\. Indeed, if

m≥⌈tmix​\(1\+log2⁡\(32​Nk4/δ\)\)⌉,m\\geq\\left\\lceil t\_\{\\mathrm\{mix\}\}\\left\(1\+\\log\_\{2\}\(32N\_\{k\}^\{4\}/\\delta\)\\right\)\\right\\rceil,then⌊m/tmix⌋≥log2⁡\(32​Nk4/δ\)\\lfloor m/t\_\{\\mathrm\{mix\}\}\\rfloor\\geq\\log\_\{2\}\(32N\_\{k\}^\{4\}/\\delta\), and henceε​\(m\)≤δ/\(32​Nk4\)\\varepsilon\(m\)\\leq\\delta/\(32N\_\{k\}^\{4\}\)\. Also,

uk≤c​yk\.u\_\{k\}\\leq cy\_\{k\}\.\(93\)Indeed,mk\+1≤c​tmix​yk\+1≤exp⁡\(c​yk\)m\_\{k\}\+1\\leq ct\_\{\\mathrm\{mix\}\}y\_\{k\}\+1\\leq\\exp\(cy\_\{k\}\)sinceNk≥K0≥tmixN\_\{k\}\\geq K\_\{0\}\\geq t\_\{\\mathrm\{mix\}\}andδ<e−1\\delta<e^\{\-1\}; substituting this into the definition ofuku\_\{k\}gives \([93](https://arxiv.org/html/2606.26316#A5.E93)\)\.

Combining \([92](https://arxiv.org/html/2606.26316#A5.E92)\) and \([93](https://arxiv.org/html/2606.26316#A5.E93)\),

ΘkNk\\displaystyle\\frac\{\\Theta\_\{k\}\}\{N\_\{k\}\}=mk​ukNk​\(1\+mkK0\)2\\displaystyle=\\frac\{m\_\{k\}u\_\{k\}\}\{N\_\{k\}\}\\left\(1\+\\frac\{m\_\{k\}\}\{K\_\{0\}\}\\right\)^\{2\}\(94\)≤c​\{tmix​yk2Nk\+tmix2​yk3K0​Nk\+tmix3​yk4K02​Nk\}\.\\displaystyle\\leq c\\left\\\{\\frac\{t\_\{\\mathrm\{mix\}\}y\_\{k\}^\{2\}\}\{N\_\{k\}\}\+\\frac\{t\_\{\\mathrm\{mix\}\}^\{2\}y\_\{k\}^\{3\}\}\{K\_\{0\}N\_\{k\}\}\+\\frac\{t\_\{\\mathrm\{mix\}\}^\{3\}y\_\{k\}^\{4\}\}\{K\_\{0\}^\{2\}N\_\{k\}\}\\right\\\}\.By[Lemma˜11](https://arxiv.org/html/2606.26316#Thmlemma11), each term is maximized, up to a numerical constant depending only on the logarithmic power, atNk=K0N\_\{k\}=K\_\{0\}\. LetK0=K​tmix​ℓ03K\_\{0\}=Kt\_\{\\mathrm\{mix\}\}\\ell\_\{0\}^\{3\}withK≥Kwin​\(ρ\)K\\geq K\_\{\\mathrm\{win\}\}\(\\rho\)\. Then

1\+log⁡K0δ≤c​\{ℓ0\+log⁡K\+log⁡ℓ0\}≤cK​ℓ0,1\+\\log\\frac\{K\_\{0\}\}\{\\delta\}\\leq c\\\{\\ell\_\{0\}\+\\log K\+\\log\\ell\_\{0\}\\\}\\leq c\_\{K\}\\ell\_\{0\},wherecKc\_\{K\}depends onKKbut not ontmixt\_\{\\mathrm\{mix\}\}orδ\\delta\. Substituting this bound into \([94](https://arxiv.org/html/2606.26316#A5.E94)\) atNk=K0N\_\{k\}=K\_\{0\}yields

ΘkNk≤c​\(cK2K​ℓ0\+cK3K2​ℓ03\+cK4K3​ℓ05\)\.\\frac\{\\Theta\_\{k\}\}\{N\_\{k\}\}\\leq c\\left\(\\frac\{c\_\{K\}^\{2\}\}\{K\\ell\_\{0\}\}\+\\frac\{c\_\{K\}^\{3\}\}\{K^\{2\}\\ell\_\{0\}^\{3\}\}\+\\frac\{c\_\{K\}^\{4\}\}\{K^\{3\}\\ell\_\{0\}^\{5\}\}\\right\)\.The right side tends to zero asK→∞K\\to\\infty, uniformly overℓ0≥1\\ell\_\{0\}\\geq 1, becauselog⁡K/Ka→0\\log K/K^\{a\}\\to 0for everya\>0a\>0\. ChoosingKwin​\(ρ\)K\_\{\\mathrm\{win\}\}\(\\rho\)sufficiently large proves \([90](https://arxiv.org/html/2606.26316#A5.E90)\)\.

For \([91](https://arxiv.org/html/2606.26316#A5.E91)\), use \([92](https://arxiv.org/html/2606.26316#A5.E92)\) andlog⁡\(Nk/K0\)≤yk\\log\(N\_\{k\}/K\_\{0\}\)\\leq y\_\{k\}:

mkNk​\(1\+log⁡NkK0\+mkK0\)≤c​\{tmix​yk2Nk\+tmix2​yk2K0​Nk\}\.\\frac\{m\_\{k\}\}\{N\_\{k\}\}\\left\(1\+\\log\\frac\{N\_\{k\}\}\{K\_\{0\}\}\+\\frac\{m\_\{k\}\}\{K\_\{0\}\}\\right\)\\leq c\\left\\\{\\frac\{t\_\{\\mathrm\{mix\}\}y\_\{k\}^\{2\}\}\{N\_\{k\}\}\+\\frac\{t\_\{\\mathrm\{mix\}\}^\{2\}y\_\{k\}^\{2\}\}\{K\_\{0\}N\_\{k\}\}\\right\\\}\.The same logarithm\-over\-linear argument bounds the right side by

c​\(cK2K​ℓ0\+cK2K2​ℓ04\),c\\left\(\\frac\{c\_\{K\}^\{2\}\}\{K\\ell\_\{0\}\}\+\\frac\{c\_\{K\}^\{2\}\}\{K^\{2\}\\ell\_\{0\}^\{4\}\}\\right\),which is at mostρ\\rhoafter increasingKwin​\(ρ\)K\_\{\\mathrm\{win\}\}\(\\rho\)\. This proves the lemma\. ∎

###### Proof of[Corollary˜1](https://arxiv.org/html/2606.26316#Thmcorollary1)\.

Letρ0,KD,C⋆\\rho\_\{0\},K\_\{D\},C\_\{\\star\}be the constants from[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)\. ChooseKgeoK\_\{\\mathrm\{geo\}\}large enough so that \([23](https://arxiv.org/html/2606.26316#S4.E23)\) implies bothK0≥Kbase∨KDK\_\{0\}\\geq K\_\{\\mathrm\{base\}\}\\vee K\_\{D\}and the hypothesis of[Lemma˜12](https://arxiv.org/html/2606.26316#Thmlemma12)withρ=ρ0\\rho=\\rho\_\{0\}\. Then[Lemma˜12](https://arxiv.org/html/2606.26316#Thmlemma12)gives

ΘkNk≤ρ0,mkNk​\(1\+log⁡NkK0\+mkK0\)≤ρ0\\frac\{\\Theta\_\{k\}\}\{N\_\{k\}\}\\leq\\rho\_\{0\},\\qquad\\frac\{m\_\{k\}\}\{N\_\{k\}\}\\left\(1\+\\log\\frac\{N\_\{k\}\}\{K\_\{0\}\}\+\\frac\{m\_\{k\}\}\{K\_\{0\}\}\\right\)\\leq\\rho\_\{0\}for everyk≥1k\\geq 1\. These are exactly the admissibility conditions \([20](https://arxiv.org/html/2606.26316#S4.E20)\)–\([21](https://arxiv.org/html/2606.26316#S4.E21)\), while the definition ofmkm\_\{k\}gives the mixing condition needed in[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)\. Hence[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)applies and proves \([22](https://arxiv.org/html/2606.26316#S4.E22)\)\.

It remains only to justify the displayed closed\-form estimates\. Set

m=⌈tmix​\(1\+log2⁡\(32​Nk4δ\)\)⌉\.m=\\left\\lceil t\_\{\\mathrm\{mix\}\}\\left\(1\+\\log\_\{2\}\\left\(\\frac\{32N\_\{k\}^\{4\}\}\{\\delta\}\\right\)\\right\)\\right\\rceil\.Then⌊m/tmix⌋≥log2⁡\(32​Nk4/δ\)\\lfloor m/t\_\{\\mathrm\{mix\}\}\\rfloor\\geq\\log\_\{2\}\(32N\_\{k\}^\{4\}/\\delta\)\. The geometric mixing assumption \([12](https://arxiv.org/html/2606.26316#S3.E12)\) therefore givesε​\(m\)≤δ/\(32​Nk4\)\\varepsilon\(m\)\\leq\\delta/\(32N\_\{k\}^\{4\}\)\. Sincemkm\_\{k\}is the minimum of this admissible value andkk, we havemk≤mm\_\{k\}\\leq m, which proves \([24](https://arxiv.org/html/2606.26316#S4.E24)\)\. The quantityΘk=mk​uk​\(1\+mk/K0\)2\\Theta\_\{k\}=m\_\{k\}u\_\{k\}\(1\+m\_\{k\}/K\_\{0\}\)^\{2\}is monotone inmkm\_\{k\}\. Replacingmkm\_\{k\}by the upper boundMkM\_\{k\}in bothuk=log⁡\(16​\(k\+1\)2​\(mk\+1\)/δ\)u\_\{k\}=\\log\(16\(k\+1\)^\{2\}\(m\_\{k\}\+1\)/\\delta\)and\(1\+mk/K0\)2\(1\+m\_\{k\}/K\_\{0\}\)^\{2\}gives \([25](https://arxiv.org/html/2606.26316#S4.E25)\)\.

Finally, \([92](https://arxiv.org/html/2606.26316#A5.E92)\) and \([93](https://arxiv.org/html/2606.26316#A5.E93)\) show thatmk=O​\(tmix​log⁡\(Nk/δ\)\)m\_\{k\}=O\(t\_\{\\mathrm\{mix\}\}\\log\(N\_\{k\}/\\delta\)\)anduk=O​\(log⁡\(Nk/δ\)\)u\_\{k\}=O\(\\log\(N\_\{k\}/\\delta\)\)\. HenceΘk=mk​uk​\(1\+mk/K0\)2\\Theta\_\{k\}=m\_\{k\}u\_\{k\}\(1\+m\_\{k\}/K\_\{0\}\)^\{2\}contributes one polynomial power oftmixt\_\{\\mathrm\{mix\}\}, with additional logarithmic factors fromuku\_\{k\}andqk=1\+mk/K0q\_\{k\}=1\+m\_\{k\}/K\_\{0\}\. The offset condition \([23](https://arxiv.org/html/2606.26316#S4.E23)\) preventsqkq\_\{k\}from creating any further polynomial power oftmixt\_\{\\mathrm\{mix\}\}\. Substituting this into \([22](https://arxiv.org/html/2606.26316#S4.E22)\) gives the statedO~​\(tmix/\(k\+K0\)\)\\widetilde\{O\}\(t\_\{\\mathrm\{mix\}\}/\(k\+K\_\{0\}\)\)leading stochastic order\. ∎

## Appendix FProof of Corollary[4](https://arxiv.org/html/2606.26316#Thmcorollary4): local PL and local oracle regularity

###### Proof of Corollary[4](https://arxiv.org/html/2606.26316#Thmcorollary4)\.

The proof is the same first\-failure induction as in[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1), but we spell out why only local assumptions are used\. DefineΛk=2​K0​Δ0\+C⋆​\(1\+Θk\)\\Lambda\_\{k\}=2K\_\{0\}\\Delta\_\{0\}\+C\_\{\\star\}\(1\+\\Theta\_\{k\}\)and the events𝖤k,ℰk\\mathsf\{E\}\_\{k\},\\mathcal\{E\}\_\{k\}as in \([84](https://arxiv.org/html/2606.26316#A4.E84)\)\. Up to the first failure of these events, the induction hypothesis gives, for every earlier indexjj,

Δj≤2​K0​Δ0\+C⋆​\(1\+Θj\)j\+K0\.\\Delta\_\{j\}\\leq\\frac\{2K\_\{0\}\\Delta\_\{0\}\+C\_\{\\star\}\(1\+\\Theta\_\{j\}\)\}\{j\+K\_\{0\}\}\.Using the admissibility conditionΘj/\(j\+K0\)≤ρ0\\Theta\_\{j\}/\(j\+K\_\{0\}\)\\leq\\rho\_\{0\}at the same index, andK0≥1K\_\{0\}\\geq 1, we obtain

Δj≤2​Δ0\+C⋆​\(1K0\+Θjj\+K0\)≤2​Δ0\+C⋆​\(1\+ρ0\)≤R\.\\Delta\_\{j\}\\leq 2\\Delta\_\{0\}\+C\_\{\\star\}\\left\(\\frac\{1\}\{K\_\{0\}\}\+\\frac\{\\Theta\_\{j\}\}\{j\+K\_\{0\}\}\\right\)\\leq 2\\Delta\_\{0\}\+C\_\{\\star\}\(1\+\\rho\_\{0\}\)\\leq R\.Thus every iterate that appears before the first failure lies in the sublevel set𝒮R\\mathcal\{S\}\_\{R\}\. On this stopped path, all uses of smoothness, PL, the ABC envelope, and the Lipschitz bound for the oracle are uses of their local versions on𝒮R\\mathcal\{S\}\_\{R\}\. The Markov mixing assumption is unchanged because it concerns the exogenous chain only\. Consequently the proofs of[Lemmas˜2](https://arxiv.org/html/2606.26316#Thmlemma2),[4](https://arxiv.org/html/2606.26316#Thmlemma4),[5](https://arxiv.org/html/2606.26316#Thmlemma5),[6](https://arxiv.org/html/2606.26316#Thmlemma6),[7](https://arxiv.org/html/2606.26316#Thmlemma7),[8](https://arxiv.org/html/2606.26316#Thmlemma8),[9](https://arxiv.org/html/2606.26316#Thmlemma9)and[10](https://arxiv.org/html/2606.26316#Thmlemma10)apply with the local constants\. The same summable first\-failure union bound as in[Theorem˜1](https://arxiv.org/html/2606.26316#Thmtheorem1)gives the stated high\-probability inequality\. Since the inequality itself keepsΔk≤R\\Delta\_\{k\}\\leq R, the stopped argument closes and all iterates remain in𝒮R\\mathcal\{S\}\_\{R\}on the resulting event\. ∎

## Appendix GProof of Theorem[2](https://arxiv.org/html/2606.26316#Thmtheorem2): linear mixing\-time lower bound

This appendix proves[Theorem˜2](https://arxiv.org/html/2606.26316#Thmtheorem2)\. The proof has four steps\. First, Appendix[G\.1](https://arxiv.org/html/2606.26316#A7.SS1)verifies that the two\-state quadratic construction satisfies the assumptions of the upper\-bound theorem\. Second, Appendix[G\.2](https://arxiv.org/html/2606.26316#A7.SS2)writes the SGD iterate as a linear filter of the Markov chain\. Third, Appendix[G\.3](https://arxiv.org/html/2606.26316#A7.SS3)proves a variance lower bound and a fourth\-moment upper bound for this filter\. Finally, Appendix[G\.4](https://arxiv.org/html/2606.26316#A7.SS4)combines these moment estimates with a constant\-probability argument and translatesε−1\\varepsilon^\{\-1\}into the mixing time\.

### G\.1Verification of the assumptions

###### Lemma 13\(The lower\-bound instance satisfies the standard assumptions\)\.

The instance \([26](https://arxiv.org/html/2606.26316#S4.E26)\)\-\([27](https://arxiv.org/html/2606.26316#S4.E27)\) satisfies Assumptions[1](https://arxiv.org/html/2606.26316#Thmassumption1)–[3](https://arxiv.org/html/2606.26316#Thmassumption3)\. Specifically,ffis11\-smooth and11\-PL,g​\(⋅,z\)g\(\\cdot,z\)is11\-Lipschitz for both states, the stationary gradient identity holds, and

\|g​\(x,z\)\|2≤2​\|x\|2\+2​σ2=2​‖∇f​\(x\)‖2\+2​σ2\.\\left\\lvert g\(x,z\)\\right\\rvert^\{2\}\\leq 2\\left\\lvert x\\right\\rvert^\{2\}\+2\\sigma^\{2\}=2\\left\\lVert\\nabla f\(x\)\\right\\rVert^\{2\}\+2\\sigma^\{2\}\.\(95\)Thus Assumption[2](https://arxiv.org/html/2606.26316#Thmassumption2)holds withA=2A=2,B=0B=0, andC=2​σ2C=2\\sigma^\{2\}\.

###### Proof\.

Forf​\(x\)=x2/2f\(x\)=x^\{2\}/2, we have∇f​\(x\)=x\\nabla f\(x\)=x, so

\|∇f​\(x\)−∇f​\(y\)\|=\|x−y\|\\left\\lvert\\nabla f\(x\)\-\\nabla f\(y\)\\right\\rvert=\\left\\lvert x\-y\\right\\rvertandffis11\-smooth\. Also

‖∇f​\(x\)‖2=x2=2​\(f​\(x\)−f⋆\),\\left\\lVert\\nabla f\(x\)\\right\\rVert^\{2\}=x^\{2\}=2\(f\(x\)\-f^\{\\star\}\),soffis11\-PL\. The invariant law of \([26](https://arxiv.org/html/2606.26316#S4.E26)\) is uniform on\{−1,\+1\}\\\{\-1,\+1\\\}, hence

∫g​\(x,z\)​π​\(d​z\)=x\+σ2\+x−σ2=x=∇f​\(x\)\.\\int g\(x,z\)\\,\\pi\(dz\)=\\frac\{x\+\\sigma\}\{2\}\+\\frac\{x\-\\sigma\}\{2\}=x=\\nabla f\(x\)\.For each fixedzz,g​\(x,z\)=x\+σ​zg\(x,z\)=x\+\\sigma zis11\-Lipschitz inxx\. Finally,

\|x\+σ​z\|2≤2​x2\+2​σ2\.\\left\\lvert x\+\\sigma z\\right\\rvert^\{2\}\\leq 2x^\{2\}\+2\\sigma^\{2\}\.forz∈\{−1,\+1\}z\\in\\\{\-1,\+1\\\}, proving \([95](https://arxiv.org/html/2606.26316#A7.E95)\)\. The mixing assertion is proved in[Lemma˜14](https://arxiv.org/html/2606.26316#Thmlemma14)below\. ∎

###### Lemma 14\(Mixing time of the two\-state chain\)\.

Letρ=1−2​ε\\rho=1\-2\\varepsilonwithε∈\(0,1/8\]\\varepsilon\\in\(0,1/8\]\. For the chain \([26](https://arxiv.org/html/2606.26316#S4.E26)\),

supz∈\{−1,\+1\}∥Pt\(⋅∣z\)−π∥TV=ρt2,t≥0\.\\sup\_\{z\\in\\\{\-1,\+1\\\}\}\\left\\lVert P^\{t\}\(\\cdot\\mid z\)\-\\pi\\right\\rVert\_\{\\mathrm\{TV\}\}=\\frac\{\\rho^\{t\}\}\{2\},\\qquad t\\geq 0\.\(96\)Consequently,

τρ=⌈log⁡2−log⁡ρ⌉\\tau\_\{\\rho\}=\\left\\lceil\\frac\{\\log 2\}\{\-\\log\\rho\}\\right\\rceilis a valid mixing time in the sense of \([12](https://arxiv.org/html/2606.26316#S3.E12)\)\. Moreover,

log⁡24​ε≤τρ≤log⁡22​ε\+1\.\\frac\{\\log 2\}\{4\\varepsilon\}\\leq\\tau\_\{\\rho\}\\leq\\frac\{\\log 2\}\{2\\varepsilon\}\+1\.\(97\)

###### Proof\.

The transition matrix has eigenvalues11andρ=1−2​ε\\rho=1\-2\\varepsilon\. Starting from\+1\+1,

ℙ​\(Zt=\+1∣Z0=\+1\)=1\+ρt2,ℙ​\(Zt=−1∣Z0=\+1\)=1−ρt2\.\\mathbb\{P\}\(Z\_\{t\}=\+1\\mid Z\_\{0\}=\+1\)=\\frac\{1\+\\rho^\{t\}\}\{2\},\\qquad\\mathbb\{P\}\(Z\_\{t\}=\-1\\mid Z\_\{0\}=\+1\)=\\frac\{1\-\\rho^\{t\}\}\{2\}\.The invariant law isπ​\(\+1\)=π​\(−1\)=1/2\\pi\(\+1\)=\\pi\(\-1\)=1/2\. Therefore the signed measurePt\(⋅∣\+1\)−πP^\{t\}\(\\cdot\\mid\+1\)\-\\pihas massesρt/2\\rho^\{t\}/2and−ρt/2\-\\rho^\{t\}/2on\+1\+1and−1\-1\. Under the probability total\-variation convention used in this paper, its norm isρt/2\\rho^\{t\}/2\. The same calculation holds from−1\-1, proving \([96](https://arxiv.org/html/2606.26316#A7.E96)\)\.

By the definition ofτρ\\tau\_\{\\rho\},ρτρ≤1/2\\rho^\{\\tau\_\{\\rho\}\}\\leq 1/2\. Ift≥0t\\geq 0andq=⌊t/τρ⌋q=\\lfloor t/\\tau\_\{\\rho\}\\rfloor, thent≥q​τρt\\geq q\\tau\_\{\\rho\}and hence

ρt2≤ρq​τρ≤2−q=2−⌊t/τρ⌋\.\\frac\{\\rho^\{t\}\}\{2\}\\leq\\rho^\{q\\tau\_\{\\rho\}\}\\leq 2^\{\-q\}=2^\{\-\\lfloor t/\\tau\_\{\\rho\}\\rfloor\}\.Thusτρ\\tau\_\{\\rho\}is a valid mixing time\.

It remains to prove \([97](https://arxiv.org/html/2606.26316#A7.E97)\)\. Foru∈\(0,1/2\]u\\in\(0,1/2\],

u≤−log⁡\(1−u\)≤2​u\.u\\leq\-\\log\(1\-u\)\\leq 2u\.Withu=2​ε∈\(0,1/4\]u=2\\varepsilon\\in\(0,1/4\], this gives

2​ε≤−log⁡\(1−2​ε\)≤4​ε\.2\\varepsilon\\leq\-\\log\(1\-2\\varepsilon\)\\leq 4\\varepsilon\.Thus

log⁡24​ε≤log⁡2−log⁡ρ≤log⁡22​ε\.\\frac\{\\log 2\}\{4\\varepsilon\}\\leq\\frac\{\\log 2\}\{\-\\log\\rho\}\\leq\\frac\{\\log 2\}\{2\\varepsilon\}\.Sinceτρ=⌈log⁡2/\(−log⁡ρ\)⌉\\tau\_\{\\rho\}=\\lceil\\log 2/\(\-\\log\\rho\)\\rceil, the stated bounds follow\. ∎

### G\.2Linear representation of the SGD iterate

Define

bi,k=αi​∏j=i\+1k−1\(1−αj\),αj=aj\+K,0≤i≤k−1,b\_\{i,k\}=\\alpha\_\{i\}\\prod\_\{j=i\+1\}^\{k\-1\}\(1\-\\alpha\_\{j\}\),\\qquad\\alpha\_\{j\}=\\frac\{a\}\{j\+K\},\\qquad 0\\leq i\\leq k\-1,\(98\)with the convention that an empty product equals one\.

###### Lemma 15\(Linear filter representation\)\.

For the recursion \([28](https://arxiv.org/html/2606.26316#S4.E28)\),

xk=−σ​∑i=0k−1bi,k​Zi\.x\_\{k\}=\-\\sigma\\sum\_\{i=0\}^\{k\-1\}b\_\{i,k\}Z\_\{i\}\.\(99\)

###### Proof\.

The proof is by induction\. Fork=1k=1,

x1=\(1−aK\)​x0−a​σK​Z0=−σ​α0​Z0,x\_\{1\}=\\left\(1\-\\frac\{a\}\{K\}\\right\)x\_\{0\}\-\\frac\{a\\sigma\}\{K\}Z\_\{0\}=\-\\sigma\\alpha\_\{0\}Z\_\{0\},becausex0=0x\_\{0\}=0, which agrees with \([99](https://arxiv.org/html/2606.26316#A7.E99)\)\. Assume \([99](https://arxiv.org/html/2606.26316#A7.E99)\) holds at timekk\. Then

xk\+1\\displaystyle x\_\{k\+1\}=\(1−αk\)​xk−αk​σ​Zk\\displaystyle=\(1\-\\alpha\_\{k\}\)x\_\{k\}\-\\alpha\_\{k\}\\sigma Z\_\{k\}=−σ​∑i=0k−1αi​\{∏j=i\+1k−1\(1−αj\)\}​\(1−αk\)​Zi−σ​αk​Zk\\displaystyle=\-\\sigma\\sum\_\{i=0\}^\{k\-1\}\\alpha\_\{i\}\\left\\\{\\prod\_\{j=i\+1\}^\{k\-1\}\(1\-\\alpha\_\{j\}\)\\right\\\}\(1\-\\alpha\_\{k\}\)Z\_\{i\}\-\\sigma\\alpha\_\{k\}Z\_\{k\}=−σ​∑i=0kαi​\{∏j=i\+1k\(1−αj\)\}​Zi\.\\displaystyle=\-\\sigma\\sum\_\{i=0\}^\{k\}\\alpha\_\{i\}\\left\\\{\\prod\_\{j=i\+1\}^\{k\}\(1\-\\alpha\_\{j\}\)\\right\\\}Z\_\{i\}\.This is \([99](https://arxiv.org/html/2606.26316#A7.E99)\) withkkreplaced byk\+1k\+1\. ∎

###### Lemma 16\(Lower bound on recent weights\)\.

Assumea≥2a\\geq 2,K≥max⁡\{4​a,4​a2\}K\\geq\\max\\\{4a,4a^\{2\}\\\}, andk≥max⁡\{8​K,16​a2\}k\\geq\\max\\\{8K,16a^\{2\}\\\}\. Let

Ik=\{⌈k/2⌉,⌈k/2⌉\+1,…,k−1\}\.I\_\{k\}=\\\{\\lceil k/2\\rceil,\\lceil k/2\\rceil\+1,\\ldots,k\-1\\\}\.Then for everyi∈Iki\\in I\_\{k\},

bi,k≥βak,βa=2​a​e−13a\+1\.b\_\{i,k\}\\geq\\frac\{\\beta\_\{a\}\}\{k\},\\qquad\\beta\_\{a\}=\\frac\{2ae^\{\-1\}\}\{3^\{a\+1\}\}\.\(100\)

###### Proof\.

SinceK≥4​aK\\geq 4a, for everyj≥0j\\geq 0,

0≤αj=aj\+K≤aK≤14\.0\\leq\\alpha\_\{j\}=\\frac\{a\}\{j\+K\}\\leq\\frac\{a\}\{K\}\\leq\\frac\{1\}\{4\}\.For0≤u≤1/20\\leq u\\leq 1/2,log⁡\(1−u\)≥−u−u2\\log\(1\-u\)\\geq\-u\-u^\{2\}\. Therefore, fori≤k−1i\\leq k\-1,

log​∏j=i\+1k−1\(1−αj\)\\displaystyle\\log\\prod\_\{j=i\+1\}^\{k\-1\}\(1\-\\alpha\_\{j\}\)≥−∑j=i\+1k−1aj\+K−∑j=i\+1k−1a2\(j\+K\)2\.\\displaystyle\\geq\-\\sum\_\{j=i\+1\}^\{k\-1\}\\frac\{a\}\{j\+K\}\-\\sum\_\{j=i\+1\}^\{k\-1\}\\frac\{a^\{2\}\}\{\(j\+K\)^\{2\}\}\.\(101\)The first sum is bounded as

∑j=i\+1k−11j\+K≤∫i\+Kk\+Kd​uu=log⁡k\+Ki\+K\.\\sum\_\{j=i\+1\}^\{k\-1\}\\frac\{1\}\{j\+K\}\\leq\\int\_\{i\+K\}^\{k\+K\}\\frac\{du\}\{u\}=\\log\\frac\{k\+K\}\{i\+K\}\.\(102\)For the second sum, usingi∈Iki\\in I\_\{k\}andk≥16​a2k\\geq 16a^\{2\},

∑j=i\+1k−1a2\(j\+K\)2≤∫i\+K∞a2​d​uu2=a2i\+K≤2​a2k≤18≤1\.\\sum\_\{j=i\+1\}^\{k\-1\}\\frac\{a^\{2\}\}\{\(j\+K\)^\{2\}\}\\leq\\int\_\{i\+K\}^\{\\infty\}\\frac\{a^\{2\}\\,du\}\{u^\{2\}\}=\\frac\{a^\{2\}\}\{i\+K\}\\leq\\frac\{2a^\{2\}\}\{k\}\\leq\\frac\{1\}\{8\}\\leq 1\.\(103\)Combining \([101](https://arxiv.org/html/2606.26316#A7.E101)\)\-\([103](https://arxiv.org/html/2606.26316#A7.E103)\) gives

∏j=i\+1k−1\(1−αj\)≥e−1​\(i\+Kk\+K\)a\.\\prod\_\{j=i\+1\}^\{k\-1\}\(1\-\\alpha\_\{j\}\)\\geq e^\{\-1\}\\left\(\\frac\{i\+K\}\{k\+K\}\\right\)^\{a\}\.\(104\)Sincei∈Iki\\in I\_\{k\}andk≥8​Kk\\geq 8K,

i\+K≥k2,k\+K≤9​k8<3​k2,i\+Kk\+K≥13\.i\+K\\geq\\frac\{k\}\{2\},\\qquad k\+K\\leq\\frac\{9k\}\{8\}<\\frac\{3k\}\{2\},\\qquad\\frac\{i\+K\}\{k\+K\}\\geq\\frac\{1\}\{3\}\.Alsoi\+K≤k\+K≤3​k/2i\+K\\leq k\+K\\leq 3k/2, so

αi=ai\+K≥2​a3​k\.\\alpha\_\{i\}=\\frac\{a\}\{i\+K\}\\geq\\frac\{2a\}\{3k\}\.Using this and \([104](https://arxiv.org/html/2606.26316#A7.E104)\) in the definition ofbi,kb\_\{i,k\}yields

bi,k≥2​a3​k​e−1​3−a=2​a​e−13a\+1​1k\.b\_\{i,k\}\\geq\\frac\{2a\}\{3k\}e^\{\-1\}3^\{\-a\}=\\frac\{2ae^\{\-1\}\}\{3^\{a\+1\}\}\\frac\{1\}\{k\}\.This proves \([100](https://arxiv.org/html/2606.26316#A7.E100)\)\. ∎

### G\.3Second and fourth moments

Let

Sk=∑i=0k−1bi,k​Zi,S\_\{k\}=\\sum\_\{i=0\}^\{k\-1\}b\_\{i,k\}Z\_\{i\},\(105\)so thatxk=−σ​Skx\_\{k\}=\-\\sigma S\_\{k\}by[Lemma˜15](https://arxiv.org/html/2606.26316#Thmlemma15)\. Since the chain starts stationarily,𝔼​Zi=0\\mathbb\{E\}Z\_\{i\}=0and

𝔼​\[Zi​Zj\]=ρ\|i−j\|,i,j≥0\.\\mathbb\{E\}\[Z\_\{i\}Z\_\{j\}\]=\\rho^\{\\left\\lvert i\-j\\right\\rvert\},\\qquad i,j\\geq 0\.\(106\)
###### Lemma 17\(Variance lower bound\)\.

Under the conditions of[Lemma˜16](https://arxiv.org/html/2606.26316#Thmlemma16), if in additionk≥16/εk\\geq 16/\\varepsilon, then

Var⁡\(Sk\)≥βa2128​1ε​k\.\\operatorname\{Var\}\(S\_\{k\}\)\\geq\\frac\{\\beta\_\{a\}^\{2\}\}\{128\}\\frac\{1\}\{\\varepsilon k\}\.\(107\)

###### Proof\.

All weightsbi,kb\_\{i,k\}are nonnegative becauseαj≤1/4\\alpha\_\{j\}\\leq 1/4\. From \([106](https://arxiv.org/html/2606.26316#A7.E106)\),

Var⁡\(Sk\)=∑i=0k−1∑j=0k−1bi,k​bj,k​ρ\|i−j\|\.\\operatorname\{Var\}\(S\_\{k\}\)=\\sum\_\{i=0\}^\{k\-1\}\\sum\_\{j=0\}^\{k\-1\}b\_\{i,k\}b\_\{j,k\}\\rho^\{\\left\\lvert i\-j\\right\\rvert\}\.\(108\)Restrict the double sum toIk×IkI\_\{k\}\\times I\_\{k\}\. By[Lemma˜16](https://arxiv.org/html/2606.26316#Thmlemma16),

Var⁡\(Sk\)≥βa2k2​∑i∈Ik∑j∈Ikρ\|i−j\|\.\\operatorname\{Var\}\(S\_\{k\}\)\\geq\\frac\{\\beta\_\{a\}^\{2\}\}\{k^\{2\}\}\\sum\_\{i\\in I\_\{k\}\}\\sum\_\{j\\in I\_\{k\}\}\\rho^\{\\left\\lvert i\-j\\right\\rvert\}\.\(109\)Letn=\|Ik\|n=\\left\\lvert I\_\{k\}\\right\\rvert\. SinceIkI\_\{k\}contains at leastk/2k/2indices,n≥k/2n\\geq k/2\. Let

h=⌊18​ε⌋\.h=\\left\\lfloor\\frac\{1\}\{8\\varepsilon\}\\right\\rfloor\.Becauseε≤1/8\\varepsilon\\leq 1/8,h≥1/\(16​ε\)h\\geq 1/\(16\\varepsilon\)\. The conditionk≥16/εk\\geq 16/\\varepsilonimpliesn≥k/2≥8/ε≥2​h\+1n\\geq k/2\\geq 8/\\varepsilon\\geq 2h\+1\. For every integer0≤r≤h0\\leq r\\leq h,

ρr=\(1−2​ε\)r≥1−2​ε​r≥1−2​ε​h≥34,\\rho^\{r\}=\(1\-2\\varepsilon\)^\{r\}\\geq 1\-2\\varepsilon r\\geq 1\-2\\varepsilon h\\geq\\frac\{3\}\{4\},where we used Bernoulli’s inequality\(1−u\)r≥1−u​r\(1\-u\)^\{r\}\\geq 1\-urforu∈\[0,1\]u\\in\[0,1\]\. For at leastn−h≥n/2n\-h\\geq n/2choices ofi∈Iki\\in I\_\{k\}, the indicesi,i\+1,…,i\+hi,i\+1,\\ldots,i\+hall belong toIkI\_\{k\}\. Hence

∑i∈Ik∑j∈Ikρ\|i−j\|\\displaystyle\\sum\_\{i\\in I\_\{k\}\}\\sum\_\{j\\in I\_\{k\}\}\\rho^\{\\left\\lvert i\-j\\right\\rvert\}≥∑i∈Ik:i\+h≤k−1∑r=0hρr\\displaystyle\\geq\\sum\_\{\\begin\{subarray\}\{c\}i\\in I\_\{k\}:\\ i\+h\\leq k\-1\\end\{subarray\}\}\\sum\_\{r=0\}^\{h\}\\rho^\{r\}\(110\)≥n2​\(h\+1\)​34≥38⋅k2⋅116​ε≥k128​ε\.\\displaystyle\\geq\\frac\{n\}\{2\}\(h\+1\)\\frac\{3\}\{4\}\\geq\\frac\{3\}\{8\}\\cdot\\frac\{k\}\{2\}\\cdot\\frac\{1\}\{16\\varepsilon\}\\geq\\frac\{k\}\{128\\varepsilon\}\.\(111\)Combining \([109](https://arxiv.org/html/2606.26316#A7.E109)\) and \([110](https://arxiv.org/html/2606.26316#A7.E110)\) proves \([107](https://arxiv.org/html/2606.26316#A7.E107)\)\. ∎

###### Lemma 18\(Fourth moment upper bound\)\.

For every deterministic nonnegative sequenceb0,…,bk−1b\_\{0\},\\ldots,b\_\{k\-1\}and the stationary two\-state chain \([26](https://arxiv.org/html/2606.26316#S4.E26)\),

𝔼​\(∑i=0k−1bi​Zi\)4≤24​\{Var⁡\(∑i=0k−1bi​Zi\)\}2\.\\mathbb\{E\}\\left\(\\sum\_\{i=0\}^\{k\-1\}b\_\{i\}Z\_\{i\}\\right\)^\{4\}\\leq 24\\left\\\{\\operatorname\{Var\}\\left\(\\sum\_\{i=0\}^\{k\-1\}b\_\{i\}Z\_\{i\}\\right\)\\right\\\}^\{2\}\.\(112\)

###### Proof\.

LetS=∑i=0k−1bi​ZiS=\\sum\_\{i=0\}^\{k\-1\}b\_\{i\}Z\_\{i\}\. We first compute mixed fourth moments\. The two\-state stationary Markov chain can be represented as

Zi=Z0​∏r=1iYr,Z\_\{i\}=Z\_\{0\}\\prod\_\{r=1\}^\{i\}Y\_\{r\},\(113\)whereZ0Z\_\{0\}is a Rademacher random variable,Y1,Y2,…Y\_\{1\},Y\_\{2\},\\ldotsare independent ofZ0Z\_\{0\}, and

ℙ​\(Yr=1\)=1−ε,ℙ​\(Yr=−1\)=ε\.\\mathbb\{P\}\(Y\_\{r\}=1\)=1\-\\varepsilon,\\qquad\\mathbb\{P\}\(Y\_\{r\}=\-1\)=\\varepsilon\.Then𝔼​Yr=1−2​ε=ρ\\mathbb\{E\}Y\_\{r\}=1\-2\\varepsilon=\\rho\. Ifi1≤i2≤i3≤i4i\_\{1\}\\leq i\_\{2\}\\leq i\_\{3\}\\leq i\_\{4\}, the productZi1​Zi2​Zi3​Zi4Z\_\{i\_\{1\}\}Z\_\{i\_\{2\}\}Z\_\{i\_\{3\}\}Z\_\{i\_\{4\}\}equals

\(∏r=i1\+1i2Yr\)​\(∏r=i3\+1i4Yr\),\\left\(\\prod\_\{r=i\_\{1\}\+1\}^\{i\_\{2\}\}Y\_\{r\}\\right\)\\left\(\\prod\_\{r=i\_\{3\}\+1\}^\{i\_\{4\}\}Y\_\{r\}\\right\),because all other factors in \([113](https://arxiv.org/html/2606.26316#A7.E113)\) appear an even number of times and cancel\. Therefore

𝔼​\[Zi1​Zi2​Zi3​Zi4\]=ρi2−i1\+i4−i3\.\\mathbb\{E\}\[Z\_\{i\_\{1\}\}Z\_\{i\_\{2\}\}Z\_\{i\_\{3\}\}Z\_\{i\_\{4\}\}\]=\\rho^\{i\_\{2\}\-i\_\{1\}\+i\_\{4\}\-i\_\{3\}\}\.\(114\)Since the coefficientsbib\_\{i\}are nonnegative, expandingS4S^\{4\}and grouping ordered quadruples by their nondecreasing rearrangement gives

𝔼​S4\\displaystyle\\mathbb\{E\}S^\{4\}≤24​∑0≤i≤j≤ℓ≤m≤k−1bi​bj​bℓ​bm​ρj−i\+m−ℓ\.\\displaystyle\\leq 24\\sum\_\{0\\leq i\\leq j\\leq\\ell\\leq m\\leq k\-1\}b\_\{i\}b\_\{j\}b\_\{\\ell\}b\_\{m\}\\rho^\{j\-i\+m\-\\ell\}\.\(115\)The factor2424is an upper bound on the number of permutations of four indices; if some indices coincide, this only overcounts and preserves the inequality\.

Define

Q=∑0≤i≤j≤k−1bi​bj​ρj−i\.Q=\\sum\_\{0\\leq i\\leq j\\leq k\-1\}b\_\{i\}b\_\{j\}\\rho^\{j\-i\}\.\(116\)The ordered quadruple sum in \([115](https://arxiv.org/html/2606.26316#A7.E115)\) is bounded byQ2Q^\{2\}, becauseQ2Q^\{2\}sumsbi​bj​bℓ​bm​ρj−i\+m−ℓb\_\{i\}b\_\{j\}b\_\{\\ell\}b\_\{m\}\\rho^\{j\-i\+m\-\\ell\}over all pairs\(i,j\)\(i,j\)and\(ℓ,m\)\(\\ell,m\)withi≤ji\\leq jandℓ≤m\\ell\\leq m, while the ordered quadruple sum restricts to the subset satisfyingj≤ℓj\\leq\\ell\. Hence

𝔼​S4≤24​Q2\.\\mathbb\{E\}S^\{4\}\\leq 24Q^\{2\}\.\(117\)On the other hand,

Var⁡\(S\)\\displaystyle\\operatorname\{Var\}\(S\)=∑i=0k−1∑j=0k−1bi​bj​ρ\|i−j\|\\displaystyle=\\sum\_\{i=0\}^\{k\-1\}\\sum\_\{j=0\}^\{k\-1\}b\_\{i\}b\_\{j\}\\rho^\{\\left\\lvert i\-j\\right\\rvert\}\(118\)=∑i=0k−1bi2\+2​∑0≤i<j≤k−1bi​bj​ρj−i\\displaystyle=\\sum\_\{i=0\}^\{k\-1\}b\_\{i\}^\{2\}\+2\\sum\_\{0\\leq i<j\\leq k\-1\}b\_\{i\}b\_\{j\}\\rho^\{j\-i\}\(119\)=Q\+∑0≤i<j≤k−1bi​bj​ρj−i≥Q\.\\displaystyle=Q\+\\sum\_\{0\\leq i<j\\leq k\-1\}b\_\{i\}b\_\{j\}\\rho^\{j\-i\}\\geq Q\.\(120\)Combining \([117](https://arxiv.org/html/2606.26316#A7.E117)\) and \([118](https://arxiv.org/html/2606.26316#A7.E118)\) proves \([112](https://arxiv.org/html/2606.26316#A7.E112)\)\. ∎

### G\.4Completion of the lower\-bound proof

###### Proof of[Theorem˜2](https://arxiv.org/html/2606.26316#Thmtheorem2)\.

The assumptions are verified in[Lemma˜13](https://arxiv.org/html/2606.26316#Thmlemma13), and the mixing\-time claim is[Lemma˜14](https://arxiv.org/html/2606.26316#Thmlemma14)\. It remains to prove the expectation and probability lower bounds\.

By[Lemma˜15](https://arxiv.org/html/2606.26316#Thmlemma15),xk=−σ​Skx\_\{k\}=\-\\sigma S\_\{k\}\. Sincef​\(x\)=x2/2f\(x\)=x^\{2\}/2andf⋆=0f^\{\\star\}=0,

f​\(xk\)−f⋆=σ2​Sk22\.f\(x\_\{k\}\)\-f^\{\\star\}=\\frac\{\\sigma^\{2\}S\_\{k\}^\{2\}\}\{2\}\.\(121\)The chain starts in stationarity, so𝔼​Sk=0\\mathbb\{E\}S\_\{k\}=0and𝔼​Sk2=Var⁡\(Sk\)\\mathbb\{E\}S\_\{k\}^\{2\}=\\operatorname\{Var\}\(S\_\{k\}\)\. Using[Lemma˜17](https://arxiv.org/html/2606.26316#Thmlemma17),

𝔼​\[f​\(xk\)−f⋆\]\\displaystyle\\mathbb\{E\}\[f\(x\_\{k\}\)\-f^\{\\star\}\]=σ22​Var⁡\(Sk\)≥σ2​βa2256​1ε​k\.\\displaystyle=\\frac\{\\sigma^\{2\}\}\{2\}\\operatorname\{Var\}\(S\_\{k\}\)\\geq\\frac\{\\sigma^\{2\}\\beta\_\{a\}^\{2\}\}\{256\}\\frac\{1\}\{\\varepsilon k\}\.\(122\)
For the probability statement, setY=Sk2Y=S\_\{k\}^\{2\}\. The fourth\-moment estimate in[Lemma˜18](https://arxiv.org/html/2606.26316#Thmlemma18)gives

𝔼​Y2=𝔼​Sk4≤24​\{𝔼​Sk2\}2\.\\mathbb\{E\}Y^\{2\}=\\mathbb\{E\}S\_\{k\}^\{4\}\\leq 24\\\{\\mathbb\{E\}S\_\{k\}^\{2\}\\\}^\{2\}\.\(123\)Applying Paley–Zygmund to the nonnegative random variableYY, with threshold one half of its mean, yields

ℙ​\(Y≥12​𝔼​Y\)≥\(1−1/2\)2​\(𝔼​Y\)2𝔼​Y2≥196\.\\mathbb\{P\}\\left\(Y\\geq\\frac\{1\}\{2\}\\mathbb\{E\}Y\\right\)\\geq\(1\-1/2\)^\{2\}\\frac\{\(\\mathbb\{E\}Y\)^\{2\}\}\{\\mathbb\{E\}Y^\{2\}\}\\geq\\frac\{1\}\{96\}\.\(124\)On this event, \([121](https://arxiv.org/html/2606.26316#A7.E121)\) and[Lemma˜17](https://arxiv.org/html/2606.26316#Thmlemma17)imply

f​\(xk\)−f⋆=σ2​Y2≥σ24​Var⁡\(Sk\)≥σ2​βa2512​1ε​k\.f\(x\_\{k\}\)\-f^\{\\star\}=\\frac\{\\sigma^\{2\}Y\}\{2\}\\geq\\frac\{\\sigma^\{2\}\}\{4\}\\operatorname\{Var\}\(S\_\{k\}\)\\geq\\frac\{\\sigma^\{2\}\\beta\_\{a\}^\{2\}\}\{512\}\\frac\{1\}\{\\varepsilon k\}\.\(125\)Since

βa2=4​a2​e−232​a\+2,\\beta\_\{a\}^\{2\}=\\frac\{4a^\{2\}e^\{\-2\}\}\{3^\{2a\+2\}\},both \([122](https://arxiv.org/html/2606.26316#A7.E122)\) and \([125](https://arxiv.org/html/2606.26316#A7.E125)\) hold with the common constant

ca=a2​e−2214​32​a\+2,c\_\{a\}=\\frac\{a^\{2\}e^\{\-2\}\}\{2^\{14\}3^\{2a\+2\}\},which is smaller than bothβa2/256\\beta\_\{a\}^\{2\}/256andβa2/512\\beta\_\{a\}^\{2\}/512\. The probability constant isc0=1/96c\_\{0\}=1/96, proving \([30](https://arxiv.org/html/2606.26316#S4.E30)\) and \([31](https://arxiv.org/html/2606.26316#S4.E31)\)\.

Finally,[Lemma˜14](https://arxiv.org/html/2606.26316#Thmlemma14)gives a valid geometric mixing timeτρ\\tau\_\{\\rho\}satisfyingτρ≤c/ε\\tau\_\{\\rho\}\\leq c/\\varepsilonandτρ≥c′/ε\\tau\_\{\\rho\}\\geq c^\{\\prime\}/\\varepsilonfor universal constantsc,c′\>0c,c^\{\\prime\}\>0\. Equivalently,1/ε1/\\varepsilonis bounded below by a universal constant times this valid mixing time\. Replacing1/ε1/\\varepsilonin the preceding two lower bounds by that constant multiple ofτρ\\tau\_\{\\rho\}gives \([33](https://arxiv.org/html/2606.26316#S4.E33)\), withca′c^\{\\prime\}\_\{a\}adjusted by the same universal factor\. ∎

## Appendix HProofs for the heavy\-tailed extension

This appendix proves[Theorems˜3](https://arxiv.org/html/2606.26316#Thmtheorem3),[3](https://arxiv.org/html/2606.26316#Thmcorollary3)and[4](https://arxiv.org/html/2606.26316#Thmtheorem4)\. Appendix[H\.1](https://arxiv.org/html/2606.26316#A8.SS1)proves the clipped\-gradient concentration and deterministic inexact\-gradient tools\. Appendix[H\.2](https://arxiv.org/html/2606.26316#A8.SS2)proves the robust upper bound\. Appendix[H\.3](https://arxiv.org/html/2606.26316#A8.SS3)derives the transition\-budget corollary under geometric mixing\. Appendix[H\.4](https://arxiv.org/html/2606.26316#A8.SS4)proves the matching heavy\-tailed sticky\-chain lower bound\.

Throughout this appendix,c,cp,cλ,pc,c\_\{p\},c\_\{\\lambda,p\}denote positive constants depending only onpp, unless another dependence is explicitly stated\.

### H\.1Auxiliary estimates for clipped Markovian blocks

###### Lemma 19\(Clipped stationary bias and variance\)\.

Fixx∈ℝdx\\in\\mathbb\{R\}^\{d\}and suppose Assumption[4](https://arxiv.org/html/2606.26316#Thmassumption4)holds\. Letg:=∇f​\(x\)g:=\\nabla f\(x\),v:=σp​Rp​\(x\)v:=\\sigma\_\{p\}R\_\{p\}\(x\), and assumeλ≥2​‖g‖∞\\lambda\\geq 2\\\|g\\\|\_\{\\infty\}\. For a coordinatejj, define

ψj​\(z,u\):=\[𝖳λ​\(𝒢​\(x,z,u\)\)\]j,ψ¯j:=∫𝔼U​\[ψj​\(z,U\)\]​π​\(d​z\)\.\\psi\_\{j\}\(z,u\):=\[\\mathsf\{T\}\_\{\\lambda\}\(\{\\mathcal\{G\}\}\(x,z,u\)\)\]\_\{j\},\\qquad\\bar\{\\psi\}\_\{j\}:=\\int\\mathbb\{E\}\_\{U\}\[\\psi\_\{j\}\(z,U\)\]\\,\\pi\(dz\)\.Then

\|ψ¯j−gj\|\\displaystyle\|\\bar\{\\psi\}\_\{j\}\-g\_\{j\}\|≤cp​vp​λ1−p,\\displaystyle\\leq c\_\{p\}v^\{p\}\\lambda^\{1\-p\},\(126\)∫𝔼U​\[\(ψj​\(z,U\)−ψ¯j\)2\]​π​\(d​z\)\\displaystyle\\int\\mathbb\{E\}\_\{U\}\\big\[\(\\psi\_\{j\}\(z,U\)\-\\bar\{\\psi\}\_\{j\}\)^\{2\}\\big\]\\pi\(dz\)≤cp​vp​λ2−p\.\\displaystyle\\leq c\_\{p\}v^\{p\}\\lambda^\{2\-p\}\.\(127\)

###### Proof\.

Let\(Z,U\)\(Z,U\)have the stationary lawπ​\(d​z\)\\pi\(dz\)together with the auxiliary randomness of the oracle, and write

η=𝒢​\(x,Z,U\)−g,ηj=\[η\]j\.\\eta=\{\\mathcal\{G\}\}\(x,Z,U\)\-g,\\qquad\\eta\_\{j\}=\[\\eta\]\_\{j\}\.The assumption gives𝔼π​‖η‖p≤vp\\mathbb\{E\}\_\{\\pi\}\\\|\\eta\\\|^\{p\}\\leq v^\{p\}, hence𝔼π​\|ηj\|p≤vp\\mathbb\{E\}\_\{\\pi\}\|\\eta\_\{j\}\|^\{p\}\\leq v^\{p\}\. Sinceλ≥2​‖g‖∞\\lambda\\geq 2\\\|g\\\|\_\{\\infty\}, if\|ηj\|<λ/2\|\\eta\_\{j\}\|<\\lambda/2, then\|gj\+ηj\|<λ\|g\_\{j\}\+\\eta\_\{j\}\|<\\lambdaand clipping does not change thejjth coordinate\. Therefore clipping bias can occur only on the event\{\|ηj\|≥λ/2\}\\\{\|\\eta\_\{j\}\|\\geq\\lambda/2\\\}\. Using Markov’s inequality and the identity𝔼​\{\|ηj\|​𝟏\|ηj\|≥a\}≤a1−p​𝔼​\|ηj\|p\\mathbb\{E\}\\\{\|\\eta\_\{j\}\|\\mathbf\{1\}\_\{\|\\eta\_\{j\}\|\\geq a\}\\\}\\leq a^\{1\-p\}\\mathbb\{E\}\|\\eta\_\{j\}\|^\{p\},

\|ψ¯j−gj\|\\displaystyle\|\\bar\{\\psi\}\_\{j\}\-g\_\{j\}\|=\|𝔼π​\{\[𝖳λ​\(g\+η\)\]j−\(gj\+ηj\)\}\|\\displaystyle=\\left\|\\mathbb\{E\}\_\{\\pi\}\\\{\[\\mathsf\{T\}\_\{\\lambda\}\(g\+\\eta\)\]\_\{j\}\-\(g\_\{j\}\+\\eta\_\{j\}\)\\\}\\right\|≤𝔼π​\[\|gj\+ηj\|​𝟏\{\|ηj\|≥λ/2\}\]\\displaystyle\\leq\\mathbb\{E\}\_\{\\pi\}\\left\[\|g\_\{j\}\+\\eta\_\{j\}\|\\mathbf\{1\}\_\{\\\{\|\\eta\_\{j\}\|\\geq\\lambda/2\\\}\}\\right\]≤\|gj\|​ℙπ​\(\|ηj\|≥λ/2\)\+𝔼π​\[\|ηj\|​𝟏\{\|ηj\|≥λ/2\}\]\\displaystyle\\leq\|g\_\{j\}\|\\mathbb\{P\}\_\{\\pi\}\(\|\\eta\_\{j\}\|\\geq\\lambda/2\)\+\\mathbb\{E\}\_\{\\pi\}\\left\[\|\\eta\_\{j\}\|\\mathbf\{1\}\_\{\\\{\|\\eta\_\{j\}\|\\geq\\lambda/2\\\}\}\\right\]≤\|gj\|​𝔼π​\|ηj\|p\(λ/2\)p\+𝔼π​\|ηj\|p\(λ/2\)p−1≤cp​vp​λ1−p,\\displaystyle\\leq\|g\_\{j\}\|\\frac\{\\mathbb\{E\}\_\{\\pi\}\|\\eta\_\{j\}\|^\{p\}\}\{\(\\lambda/2\)^\{p\}\}\+\\frac\{\\mathbb\{E\}\_\{\\pi\}\|\\eta\_\{j\}\|^\{p\}\}\{\(\\lambda/2\)^\{p\-1\}\}\\leq c\_\{p\}v^\{p\}\\lambda^\{1\-p\},where the first term is also of ordervp​λ1−pv^\{p\}\\lambda^\{1\-p\}because\|gj\|≤λ/2\|g\_\{j\}\|\\leq\\lambda/2\. This proves \([126](https://arxiv.org/html/2606.26316#A8.E126)\)\.

For the second moment, split again according to\|ηj\|≤λ/2\|\\eta\_\{j\}\|\\leq\\lambda/2\. On this event the clipped coordinate differs fromgjg\_\{j\}by exactlyηj\\eta\_\{j\}\. On the complement, both\[𝖳λ​\(g\+η\)\]j\[\\mathsf\{T\}\_\{\\lambda\}\(g\+\\eta\)\]\_\{j\}andgjg\_\{j\}have magnitude at mostλ\\lambdaandλ/2\\lambda/2, respectively, so their difference is bounded by a numerical multiple ofλ\\lambda\. Hence

𝔼π​\(\[𝖳λ​\(g\+η\)\]j−gj\)2\\displaystyle\\mathbb\{E\}\_\{\\pi\}\\big\(\[\\mathsf\{T\}\_\{\\lambda\}\(g\+\\eta\)\]\_\{j\}\-g\_\{j\}\\big\)^\{2\}≤𝔼π​\[\|ηj\|2​𝟏\{\|ηj\|≤λ/2\}\]\+c​λ2​ℙπ​\(\|ηj\|\>λ/2\)\\displaystyle\\leq\\mathbb\{E\}\_\{\\pi\}\\big\[\|\\eta\_\{j\}\|^\{2\}\\mathbf\{1\}\_\{\\\{\|\\eta\_\{j\}\|\\leq\\lambda/2\\\}\}\\big\]\+c\\lambda^\{2\}\\mathbb\{P\}\_\{\\pi\}\(\|\\eta\_\{j\}\|\>\\lambda/2\)≤\(λ/2\)2−p​𝔼π​\|ηj\|p\+c​λ2​𝔼π​\|ηj\|p\(λ/2\)p≤cp​vp​λ2−p\.\\displaystyle\\leq\(\\lambda/2\)^\{2\-p\}\\mathbb\{E\}\_\{\\pi\}\|\\eta\_\{j\}\|^\{p\}\+c\\lambda^\{2\}\\frac\{\\mathbb\{E\}\_\{\\pi\}\|\\eta\_\{j\}\|^\{p\}\}\{\(\\lambda/2\)^\{p\}\}\\leq c\_\{p\}v^\{p\}\\lambda^\{2\-p\}\.Finally, variance aroundψ¯j\\bar\{\\psi\}\_\{j\}is no larger than the second moment around any fixed center, so

𝔼π​\[\(ψj−ψ¯j\)2\]≤𝔼π​\[\(ψj−gj\)2\]≤cp​vp​λ2−p\.\\mathbb\{E\}\_\{\\pi\}\[\(\\psi\_\{j\}\-\\bar\{\\psi\}\_\{j\}\)^\{2\}\]\\leq\\mathbb\{E\}\_\{\\pi\}\[\(\\psi\_\{j\}\-g\_\{j\}\)^\{2\}\]\\leq c\_\{p\}v^\{p\}\\lambda^\{2\-p\}\.This proves \([127](https://arxiv.org/html/2606.26316#A8.E127)\)\. ∎

###### Lemma 20\(Consecutive clipped\-block concentration\)\.

Fix a deterministic pointxxsatisfyingΔ​\(x\)≤ℛ\\Delta\(x\)\\leq\\mathcal\{R\}, a block lengthb≥1b\\geq 1, and an integerm∈\{1,…,b\}m\\in\\\{1,\\ldots,b\\\}\. Let the Markov chain start from an arbitrary state and let

Yi=𝒢​\(x,Zi,Ui\),i=1,…,b,Y\_\{i\}=\{\\mathcal\{G\}\}\(x,Z\_\{i\},U\_\{i\}\),\\qquad i=1,\\ldots,b,where the auxiliary variablesUiU\_\{i\}are conditionally independent given the Markov trajectory\. Define

u:=log⁡\(16​d​\(m\+1\)η\),s:=m​ub,u:=\\log\\left\(\\frac\{16d\(m\+1\)\}\{\\eta\}\\right\),\\qquad s:=\\frac\{mu\}\{b\},for someη∈\(0,1\)\\eta\\in\(0,1\)\. Supposeε​\(m\)≤η/\(8​d​b\)\\varepsilon\(m\)\\leq\\eta/\(8db\), set

λ=cλ,p​σp​Bℛ​s−1/p,\\lambda=c\_\{\\lambda,p\}\\sigma\_\{p\}B\_\{\\mathcal\{R\}\}s^\{\-1/p\},withcλ,pc\_\{\\lambda,p\}sufficiently large, and assumeλ≥2​2​L​ℛ\\lambda\\geq 2\\sqrt\{2L\\mathcal\{R\}\}\. Then, with probability at least1−η1\-\\eta,

‖1b​∑i=1b𝖳λ​\(Yi\)−∇f​\(x\)‖≤cp​d​σp​Bℛ​sp−1p\.\\left\\\|\\frac\{1\}\{b\}\\sum\_\{i=1\}^\{b\}\\mathsf\{T\}\_\{\\lambda\}\(Y\_\{i\}\)\-\\nabla f\(x\)\\right\\\|\\leq c\_\{p\}\\sqrt\{d\}\\,\\sigma\_\{p\}B\_\{\\mathcal\{R\}\}s^\{\\frac\{p\-1\}\{p\}\}\.\(128\)The same conclusion holds conditionally on any sigma\-field𝒜\\mathcal\{A\}with respect to whichxxand the initial Markov state are measurable, provided that after conditioning on𝒜\\mathcal\{A\}the future chain evolves with transition kernelPPfrom that initial state and the auxiliary variables\(Ui\)\(U\_\{i\}\)are conditionally independent given the future Markov trajectory, with the same conditional oracle kernel as in[Assumption˜4](https://arxiv.org/html/2606.26316#Thmassumption4)\.

###### Proof\.

We first justify the conditional version\. After conditioning on such a sigma\-field𝒜\\mathcal\{A\}, the query pointxxand the starting state are deterministic, the future Markov process still has transition kernelPP, and the auxiliary variables remain conditionally independent with the prescribed conditional oracle law\. Hence all conditional expectations, mixing\-bias estimates, and martingale\-difference properties below hold under the regular conditional law\. It is therefore enough to prove the result for a fixed deterministic starting state\. SinceΔ​\(x\)≤ℛ\\Delta\(x\)\\leq\\mathcal\{R\},[Lemma˜1](https://arxiv.org/html/2606.26316#Thmlemma1)gives‖∇f​\(x\)‖≤2​L​ℛ\\\|\\nabla f\(x\)\\\|\\leq\\sqrt\{2L\\mathcal\{R\}\}\. ThusRp​\(x\)≤BℛR\_\{p\}\(x\)\\leq B\_\{\\mathcal\{R\}\}\. Put

g:=∇f​\(x\),v:=σp​Bℛ\.g:=\\nabla f\(x\),\\qquad v:=\\sigma\_\{p\}B\_\{\\mathcal\{R\}\}\.Then[Assumption˜4](https://arxiv.org/html/2606.26316#Thmassumption4)gives the stationary moment bound needed in[Lemma˜19](https://arxiv.org/html/2606.26316#Thmlemma19), and the conditionλ≥2​2​L​ℛ\\lambda\\geq 2\\sqrt\{2L\\mathcal\{R\}\}impliesλ≥2​‖g‖∞\\lambda\\geq 2\\\|g\\\|\_\{\\infty\}\.

Fix a coordinatejj\. Letψj\\psi\_\{j\}andψ¯j\\bar\{\\psi\}\_\{j\}be the clipped coordinate and its stationary mean from[Lemma˜19](https://arxiv.org/html/2606.26316#Thmlemma19)\. We first control the centered block average

Sj:=1b​∑i=1b\{ψj​\(Zi,Ui\)−ψ¯j\}\.S\_\{j\}:=\\frac\{1\}\{b\}\\sum\_\{i=1\}^\{b\}\\\{\\psi\_\{j\}\(Z\_\{i\},U\_\{i\}\)\-\\bar\{\\psi\}\_\{j\}\\\}\.The firstmmsamples are not delayed by a full mixing window\. Since\|ψj\|≤λ\|\\psi\_\{j\}\|\\leq\\lambdaand\|ψ¯j\|≤λ\|\\bar\{\\psi\}\_\{j\}\|\\leq\\lambda, their total contribution is bounded deterministically by

\|1b​∑i=1m\{ψj​\(Zi,Ui\)−ψ¯j\}\|≤2​λ​mb\.\\left\|\\frac\{1\}\{b\}\\sum\_\{i=1\}^\{m\}\\\{\\psi\_\{j\}\(Z\_\{i\},U\_\{i\}\)\-\\bar\{\\psi\}\_\{j\}\\\}\\right\|\\leq\\frac\{2\\lambda m\}\{b\}\.
Fori\>mi\>m, define

Di:=ψj​\(Zi,Ui\)−ψ¯j−𝔼​\[ψj​\(Zi,Ui\)−ψ¯j∣ℱi−m\],D\_\{i\}:=\\psi\_\{j\}\(Z\_\{i\},U\_\{i\}\)\-\\bar\{\\psi\}\_\{j\}\-\\mathbb\{E\}\[\\psi\_\{j\}\(Z\_\{i\},U\_\{i\}\)\-\\bar\{\\psi\}\_\{j\}\\mid\\mathcal\{F\}\_\{i\-m\}\],whereℱi\\mathcal\{F\}\_\{i\}contains the chain and auxiliary variables up to timeii\. The conditional mean term is small because, givenℱi−m\\mathcal\{F\}\_\{i\-m\}, the law ofZiZ\_\{i\}isPm​\(Zi−m,⋅\)P^\{m\}\(Z\_\{i\-m\},\\cdot\), while the auxiliary variable is sampled from the same conditional kernel as in the stationary definition\. Since∫𝔼U​ψj​\(z,U\)​π​\(d​z\)=ψ¯j\\int\\mathbb\{E\}\_\{U\}\\psi\_\{j\}\(z,U\)\\pi\(dz\)=\\bar\{\\psi\}\_\{j\}, the total\-variation inequality \([17](https://arxiv.org/html/2606.26316#S3.E17)\) and\|ψj\|≤λ\|\\psi\_\{j\}\|\\leq\\lambdagive

\|𝔼\[ψj\(Zi,Ui\)−ψ¯j∣ℱi−m\]\|≤2λε\(m\)\.\\left\|\\mathbb\{E\}\[\\psi\_\{j\}\(Z\_\{i\},U\_\{i\}\)\-\\bar\{\\psi\}\_\{j\}\\mid\\mathcal\{F\}\_\{i\-m\}\]\\right\|\\leq 2\\lambda\\varepsilon\(m\)\.\(129\)
The variables\(Di\)i\>m\(D\_\{i\}\)\_\{i\>m\}are not martingale differences in the ordinary time order\. Split the indices\{m\+1,…,b\}\\\{m\+1,\\ldots,b\\\}into residue classes modulomm\. Ifi1<i2<⋯<isi\_\{1\}<i\_\{2\}<\\cdots<i\_\{s\}are the indices in one residue class, thenir\+1−m=iri\_\{r\+1\}\-m=i\_\{r\}\. Hence

𝔼​\[Dir\+1∣ℱir\]=𝔼​\[Dir\+1∣ℱir\+1−m\]=0,\\mathbb\{E\}\[D\_\{i\_\{r\+1\}\}\\mid\\mathcal\{F\}\_\{i\_\{r\}\}\]=\\mathbb\{E\}\[D\_\{i\_\{r\+1\}\}\\mid\\mathcal\{F\}\_\{i\_\{r\+1\}\-m\}\]=0,so this subsequence is a martingale\-difference sequence with respect to the filtration\(ℱir\)r=1s\(\\mathcal\{F\}\_\{i\_\{r\}\}\)\_\{r=1\}^\{s\}\. Moreover,\|Di\|≤4​λ\|D\_\{i\}\|\\leq 4\\lambda, because bothψj−ψ¯j\\psi\_\{j\}\-\\bar\{\\psi\}\_\{j\}and its conditional mean are bounded by2​λ2\\lambda\.

We also need a conditional variance bound\. Since subtracting a conditional mean can only decrease conditional second moment,

𝔼​\[Di2∣ℱi−m\]≤𝔼​\[\(ψj​\(Zi,Ui\)−ψ¯j\)2∣ℱi−m\]\.\\mathbb\{E\}\[D\_\{i\}^\{2\}\\mid\\mathcal\{F\}\_\{i\-m\}\]\\leq\\mathbb\{E\}\[\(\\psi\_\{j\}\(Z\_\{i\},U\_\{i\}\)\-\\bar\{\\psi\}\_\{j\}\)^\{2\}\\mid\\mathcal\{F\}\_\{i\-m\}\]\.The function\(ψj−ψ¯j\)2\(\\psi\_\{j\}\-\\bar\{\\psi\}\_\{j\}\)^\{2\}is bounded by4​λ24\\lambda^\{2\}\. Combining the stationary variance bound in[Lemma˜19](https://arxiv.org/html/2606.26316#Thmlemma19)with \([17](https://arxiv.org/html/2606.26316#S3.E17)\) gives

𝔼​\[Di2∣ℱi−m\]≤cp​vp​λ2−p\+8​λ2​ε​\(m\)\.\\mathbb\{E\}\[D\_\{i\}^\{2\}\\mid\\mathcal\{F\}\_\{i\-m\}\]\\leq c\_\{p\}v^\{p\}\\lambda^\{2\-p\}\+8\\lambda^\{2\}\\varepsilon\(m\)\.\(130\)
Apply Freedman’s inequality separately on each residue class, using range bound4​λ4\\lambdaand the variance bound \([130](https://arxiv.org/html/2606.26316#A8.E130)\)\. LetIr⊆\{m\+1,…,b\}I\_\{r\}\\subseteq\\\{m\+1,\\ldots,b\\\}denote the indices in residue classrr\. For a fixed residue class, Freedman’s inequality gives, with probability at least1−2​e−u1\-2e^\{\-u\},

\|∑i∈IrDi\|≤c​u​\|Ir\|​\(vp​λ2−p\+λ2​ε​\(m\)\)\+c​λ​u\.\\left\|\\sum\_\{i\\in I\_\{r\}\}D\_\{i\}\\right\|\\leq c\\sqrt\{u\|I\_\{r\}\|\\left\(v^\{p\}\\lambda^\{2\-p\}\+\\lambda^\{2\}\\varepsilon\(m\)\\right\)\}\+c\\lambda u\.Taking a union bound over the at mostmmresidue classes, this holds for every residue class simultaneously with probability at least1−2​m​e−u1\-2me^\{\-u\}\. On this event,

\|∑i=m\+1bDi\|\\displaystyle\\left\|\\sum\_\{i=m\+1\}^\{b\}D\_\{i\}\\right\|≤∑r=0m−1\|∑i∈IrDi\|\\displaystyle\\leq\\sum\_\{r=0\}^\{m\-1\}\\left\|\\sum\_\{i\\in I\_\{r\}\}D\_\{i\}\\right\|≤c​∑r=0m−1u​\|Ir\|​\(vp​λ2−p\+λ2​ε​\(m\)\)\+c​λ​m​u\.\\displaystyle\\leq c\\sum\_\{r=0\}^\{m\-1\}\\sqrt\{u\|I\_\{r\}\|\\left\(v^\{p\}\\lambda^\{2\-p\}\+\\lambda^\{2\}\\varepsilon\(m\)\\right\)\}\+c\\lambda mu\.By Cauchy–Schwarz,

∑r=0m−1\|Ir\|​\(vp​λ2−p\+λ2​ε​\(m\)\)\\displaystyle\\sum\_\{r=0\}^\{m\-1\}\\sqrt\{\|I\_\{r\}\|\\left\(v^\{p\}\\lambda^\{2\-p\}\+\\lambda^\{2\}\\varepsilon\(m\)\\right\)\}≤m​∑r=0m−1\|Ir\|​\(vp​λ2−p\+λ2​ε​\(m\)\)\\displaystyle\\leq\\sqrt\{m\\sum\_\{r=0\}^\{m\-1\}\|I\_\{r\}\|\\left\(v^\{p\}\\lambda^\{2\-p\}\+\\lambda^\{2\}\\varepsilon\(m\)\\right\)\}≤m​b​\(vp​λ2−p\+λ2​ε​\(m\)\)\.\\displaystyle\\leq\\sqrt\{mb\\left\(v^\{p\}\\lambda^\{2\-p\}\+\\lambda^\{2\}\\varepsilon\(m\)\\right\)\}\.Therefore

\|∑i=m\+1bDi\|\\displaystyle\\left\|\\sum\_\{i=m\+1\}^\{b\}D\_\{i\}\\right\|≤c​m​u​b​\(vp​λ2−p\+λ2​ε​\(m\)\)\+c​λ​m​u\.\\displaystyle\\leq c\\sqrt\{mu\\,b\\left\(v^\{p\}\\lambda^\{2\-p\}\+\\lambda^\{2\}\\varepsilon\(m\)\\right\)\}\+c\\lambda mu\.\(131\)The factormmunder the square root is the price of the residue\-class union and summation, not a loss of samples: the total number of summands remainsb−mb\-m\.

Combining the initial\-window bound, the accumulated conditional\-mean bound from \([129](https://arxiv.org/html/2606.26316#A8.E129)\), and \([131](https://arxiv.org/html/2606.26316#A8.E131)\), and then dividing bybb, yields for the fixed coordinate

\|Sj\|\\displaystyle\|S\_\{j\}\|≤c​\[m​ub​vp​λ2−p\+λ​m​ub​ε​\(m\)\+λ​m​ub\+λ​mb\+λ​ε​\(m\)\]\\displaystyle\\leq c\\left\[\\sqrt\{\\frac\{mu\}\{b\}v^\{p\}\\lambda^\{2\-p\}\}\+\\lambda\\sqrt\{\\frac\{mu\}\{b\}\\varepsilon\(m\)\}\+\\frac\{\\lambda mu\}\{b\}\+\\frac\{\\lambda m\}\{b\}\+\\lambda\\varepsilon\(m\)\\right\]with probability at least1−2​m​e−u1\-2me^\{\-u\}\. Sinceu≥1u\\geq 1andε​\(m\)≤η/\(8​d​b\)≤1/b≤m​u/b\\varepsilon\(m\)\\leq\\eta/\(8db\)\\leq 1/b\\leq mu/b, the last four terms are bounded by a constant multiple ofλ​m​u/b=λ​s\\lambda mu/b=\\lambda s\. Therefore

\|Sj\|≤c​\[s​vp​λ2−p\+λ​s\]\.\|S\_\{j\}\|\\leq c\\left\[\\sqrt\{sv^\{p\}\\lambda^\{2\-p\}\}\+\\lambda s\\right\]\.Adding the stationary clipping bias from[Lemma˜19](https://arxiv.org/html/2606.26316#Thmlemma19)gives

\|1b​∑i=1b\[𝖳λ​\(Yi\)\]j−gj\|≤cp​\[s​vp​λ2−p\+λ​s\+vp​λ1−p\]\.\\left\|\\frac\{1\}\{b\}\\sum\_\{i=1\}^\{b\}\[\\mathsf\{T\}\_\{\\lambda\}\(Y\_\{i\}\)\]\_\{j\}\-g\_\{j\}\\right\|\\leq c\_\{p\}\\left\[\\sqrt\{sv^\{p\}\\lambda^\{2\-p\}\}\+\\lambda s\+v^\{p\}\\lambda^\{1\-p\}\\right\]\.Withλ=cλ,p​v​s−1/p\\lambda=c\_\{\\lambda,p\}vs^\{\-1/p\}, each of the three terms on the right is of orderv​s\(p−1\)/pvs^\{\(p\-1\)/p\}, after choosingcλ,pc\_\{\\lambda,p\}sufficiently large\. Finally,2​d​m​e−u≤η2dme^\{\-u\}\\leq\\etaby the definition ofuu, so a union bound over theddcoordinates gives coordinatewise control simultaneously\. Taking the Euclidean norm multiplies the coordinate bound by at mostd\\sqrt\{d\}, proving \([128](https://arxiv.org/html/2606.26316#A8.E128)\)\. ∎

###### Lemma 21\(PL descent with deterministic gradient errors\)\.

Assume thatffisLL\-smooth andμ\\mu\-PL\. LetJ≥1J\\geq 1be an integer and let

xr\+1=xr−14​L​\(∇f​\(xr\)\+er\)x\_\{r\+1\}=x\_\{r\}\-\\frac\{1\}\{4L\}\(\\nabla f\(x\_\{r\}\)\+e\_\{r\}\)for0≤r<J0\\leq r<J\. Suppose that‖er‖≤ϵ\\\|e\_\{r\}\\\|\\leq\\epsilonfor all0≤r<J0\\leq r<J\. Then

Δ​\(xJ\)≤\(1−μ16​L\)J​Δ​\(x0\)\+8​ϵ2μ\.\\Delta\(x\_\{J\}\)\\leq\\left\(1\-\\frac\{\\mu\}\{16L\}\\right\)^\{J\}\\Delta\(x\_\{0\}\)\+\\frac\{8\\epsilon^\{2\}\}\{\\mu\}\.\(132\)

###### Proof\.

Fix one step and writegr:=∇f​\(xr\)g\_\{r\}:=\\nabla f\(x\_\{r\}\)\. ByLL\-smoothness and the update rule,

Δ​\(xr\+1\)\\displaystyle\\Delta\(x\_\{r\+1\}\)≤Δ​\(xr\)−14​L​⟨gr,gr\+er⟩\+L2​‖14​L​\(gr\+er\)‖2\\displaystyle\\leq\\Delta\(x\_\{r\}\)\-\\frac\{1\}\{4L\}\\langle g\_\{r\},g\_\{r\}\+e\_\{r\}\\rangle\+\\frac\{L\}\{2\}\\left\\\|\\frac\{1\}\{4L\}\(g\_\{r\}\+e\_\{r\}\)\\right\\\|^\{2\}=Δ​\(xr\)−14​L​‖gr‖2−14​L​⟨gr,er⟩\+132​L​‖gr\+er‖2\.\\displaystyle=\\Delta\(x\_\{r\}\)\-\\frac\{1\}\{4L\}\\\|g\_\{r\}\\\|^\{2\}\-\\frac\{1\}\{4L\}\\langle g\_\{r\},e\_\{r\}\\rangle\+\\frac\{1\}\{32L\}\\\|g\_\{r\}\+e\_\{r\}\\\|^\{2\}\.Using‖gr\+er‖2≤2​‖gr‖2\+2​‖er‖2\\\|g\_\{r\}\+e\_\{r\}\\\|^\{2\}\\leq 2\\\|g\_\{r\}\\\|^\{2\}\+2\\\|e\_\{r\}\\\|^\{2\}and−⟨gr,er⟩≤‖gr‖​‖er‖≤12​‖gr‖2\+12​‖er‖2\-\\langle g\_\{r\},e\_\{r\}\\rangle\\leq\\\|g\_\{r\}\\\|\\\|e\_\{r\}\\\|\\leq\\frac\{1\}\{2\}\\\|g\_\{r\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\\|e\_\{r\}\\\|^\{2\},

Δ​\(xr\+1\)≤Δ​\(xr\)−116​L​‖gr‖2\+516​L​‖er‖2\.\\Delta\(x\_\{r\+1\}\)\\leq\\Delta\(x\_\{r\}\)\-\\frac\{1\}\{16L\}\\\|g\_\{r\}\\\|^\{2\}\+\\frac\{5\}\{16L\}\\\|e\_\{r\}\\\|^\{2\}\.Enlarging constants in the harmless direction, and using‖er‖≤ϵ\\\|e\_\{r\}\\\|\\leq\\epsilon,

Δ​\(xr\+1\)≤Δ​\(xr\)−116​L​‖gr‖2\+ϵ22​L\.\\Delta\(x\_\{r\+1\}\)\\leq\\Delta\(x\_\{r\}\)\-\\frac\{1\}\{16L\}\\\|g\_\{r\}\\\|^\{2\}\+\\frac\{\\epsilon^\{2\}\}\{2L\}\.The PL inequality gives‖gr‖2≥2​μ​Δ​\(xr\)\\\|g\_\{r\}\\\|^\{2\}\\geq 2\\mu\\Delta\(x\_\{r\}\), hence

Δ​\(xr\+1\)≤\(1−μ8​L\)​Δ​\(xr\)\+ϵ22​L≤\(1−μ16​L\)​Δ​\(xr\)\+ϵ22​L\.\\Delta\(x\_\{r\+1\}\)\\leq\\left\(1\-\\frac\{\\mu\}\{8L\}\\right\)\\Delta\(x\_\{r\}\)\+\\frac\{\\epsilon^\{2\}\}\{2L\}\\leq\\left\(1\-\\frac\{\\mu\}\{16L\}\\right\)\\Delta\(x\_\{r\}\)\+\\frac\{\\epsilon^\{2\}\}\{2L\}\.Iterating this affine recursion forJJsteps yields

Δ​\(xJ\)≤\(1−μ16​L\)J​Δ​\(x0\)\+ϵ22​L​∑t=0J−1\(1−μ16​L\)t\.\\Delta\(x\_\{J\}\)\\leq\\left\(1\-\\frac\{\\mu\}\{16L\}\\right\)^\{J\}\\Delta\(x\_\{0\}\)\+\\frac\{\\epsilon^\{2\}\}\{2L\}\\,\\sum\_\{t=0\}^\{J\-1\}\\left\(1\-\\frac\{\\mu\}\{16L\}\\right\)^\{t\}\.The geometric sum is at most16​L/μ16L/\\mu, so the residual term is at most8​ϵ2/μ8\\epsilon^\{2\}/\\mu, proving \([132](https://arxiv.org/html/2606.26316#A8.E132)\)\. ∎

### H\.2Proof of Theorem[3](https://arxiv.org/html/2606.26316#Thmtheorem3): robust upper bound

###### Proof of[Theorem˜3](https://arxiv.org/html/2606.26316#Thmtheorem3)\.

Let

ϵ∇:=cp​d​σp​Bℛ​shtp−1p,\\epsilon\_\{\\nabla\}:=c\_\{p\}\\sqrt\{d\}\\,\\sigma\_\{p\}B\_\{\\mathcal\{R\}\}s\_\{\\rm ht\}^\{\\frac\{p\-1\}\{p\}\},\(133\)where the constant is the one from[Lemma˜20](https://arxiv.org/html/2606.26316#Thmlemma20)\. We enlarge the constantcpc\_\{p\}in the statement, if necessary, so that the admissibility condition \([46](https://arxiv.org/html/2606.26316#S5.E46)\) implies

8​ϵ∇2μ≤cp​d​σp2​Bℛ2μ​shtϑp≤ℛ4\.\\frac\{8\\epsilon\_\{\\nabla\}^\{2\}\}\{\\mu\}\\leq\\frac\{c\_\{p\}d\\sigma\_\{p\}^\{2\}B\_\{\\mathcal\{R\}\}^\{2\}\}\{\\mu\}s\_\{\\rm ht\}^\{\\vartheta\_\{p\}\}\\leq\\frac\{\\mathcal\{R\}\}\{4\}\.\(134\)
We prove that all block gradient estimates are accurate by a stopped induction\. Letℱrpre\\mathcal\{F\}\_\{r\}^\{\\rm pre\}be the sigma\-field just before therrth block is sampled\. Define

ℋr:=\{Δ\(xs\)≤ℛfor every0≤s≤r,∥g^s−∇f\(xs\)∥≤ϵ∇for every0≤s<r\},0≤r≤N\.\\mathcal\{H\}\_\{r\}:=\\left\\\{\\Delta\(x\_\{s\}\)\\leq\\mathcal\{R\}\\text\{ for every \}0\\leq s\\leq r,\\quad\\\|\\widehat\{g\}\_\{s\}\-\\nabla f\(x\_\{s\}\)\\\|\\leq\\epsilon\_\{\\nabla\}\\text\{ for every \}0\\leq s<r\\right\\\},\\qquad 0\\leq r\\leq N\.This definition intentionally includes the current radius conditionΔ​\(xr\)≤ℛ\\Delta\(x\_\{r\}\)\\leq\\mathcal\{R\}, which is needed before applying the block concentration lemma at the query pointxrx\_\{r\}\. The initial stopped eventℋ0\\mathcal\{H\}\_\{0\}is deterministic and true becauseℛ≥4​Δ0\+4\\mathcal\{R\}\\geq 4\\Delta\_\{0\}\+4impliesΔ​\(x0\)=Δ0≤ℛ/4≤ℛ\\Delta\(x\_\{0\}\)=\\Delta\_\{0\}\\leq\\mathcal\{R\}/4\\leq\\mathcal\{R\}\.

The stopped eventℋr\\mathcal\{H\}\_\{r\}serves two purposes\. The radius conditionΔ​\(xs\)≤ℛ\\Delta\(x\_\{s\}\)\\leq\\mathcal\{R\}is needed to apply[Lemma˜20](https://arxiv.org/html/2606.26316#Thmlemma20)at the next query point, because the clipping scale is chosen usingBℛB\_\{\\mathcal\{R\}\}\. The gradient\-error condition is needed to invoke the deterministic inexact\-gradient PL lemma,[Lemma˜21](https://arxiv.org/html/2606.26316#Thmlemma21)\. The proof therefore alternates between these two ingredients: concentration gives the next accurate block gradient estimate, and deterministic PL descent keeps the next iterate inside the radius\.

Supposeℋr\\mathcal\{H\}\_\{r\}holds\. Thenxrx\_\{r\}isℱrpre\\mathcal\{F\}\_\{r\}^\{\\rm pre\}\-measurable andΔ​\(xr\)≤ℛ\\Delta\(x\_\{r\}\)\\leq\\mathcal\{R\}\. During the next block the algorithm holds this point fixed and observesbbconsecutive Markovian oracle samples\. Every one of these samples is included in the averageg^r\\widehat\{g\}\_\{r\}; the lagmmis used only to decompose the proof into an initial window and delayed residue classes, not to discard or thin the block\. Conditional onℱrpre\\mathcal\{F\}\_\{r\}^\{\\rm pre\}, this is exactly the setting of the conditional version of[Lemma˜20](https://arxiv.org/html/2606.26316#Thmlemma20), with arbitrary deterministic initial Markov state and failure probabilityη=δ/\(4​N\)\\eta=\\delta/\(4N\)\. The logarithmic factor in that lemma is

log⁡16​d​\(m\+1\)η=log⁡64​d​N​\(m\+1\)δ=uht\.\\log\\frac\{16d\(m\+1\)\}\{\\eta\}=\\log\\frac\{64dN\(m\+1\)\}\{\\delta\}=u\_\{\\rm ht\}\.Moreover, \([44](https://arxiv.org/html/2606.26316#S5.E44)\) gives

ε​\(m\)≤δ64​d​N​b=η16​d​b≤η8​d​b,\\varepsilon\(m\)\\leq\\frac\{\\delta\}\{64dNb\}=\\frac\{\\eta\}\{16db\}\\leq\\frac\{\\eta\}\{8db\},so the lemma’s mixing requirement is satisfied\. Therefore, onℋr\\mathcal\{H\}\_\{r\},

ℙ​\(‖g^r−∇f​\(xr\)‖\>ϵ∇\|ℱrpre\)≤δ4​N\.\\mathbb\{P\}\\left\(\\left\.\\\|\\widehat\{g\}\_\{r\}\-\\nabla f\(x\_\{r\}\)\\\|\>\\epsilon\_\{\\nabla\}\\ \\right\|\\ \\mathcal\{F\}\_\{r\}^\{\\rm pre\}\\right\)\\leq\\frac\{\\delta\}\{4N\}\.\(135\)
Let𝒜r:=\{‖g^r−∇f​\(xr\)‖≤ϵ∇\}\\mathcal\{A\}\_\{r\}:=\\\{\\\|\\widehat\{g\}\_\{r\}\-\\nabla f\(x\_\{r\}\)\\\|\\leq\\epsilon\_\{\\nabla\}\\\}\. Multiplying \([135](https://arxiv.org/html/2606.26316#A8.E135)\) by𝟏ℋr\\mathbf\{1\}\_\{\\mathcal\{H\}\_\{r\}\}and taking expectations gives

ℙ​\(ℋr∩𝒜rc\)≤δ4​N\.\\mathbb\{P\}\(\\mathcal\{H\}\_\{r\}\\cap\\mathcal\{A\}\_\{r\}^\{c\}\)\\leq\\frac\{\\delta\}\{4N\}\.Thus

ℙ​\(⋃r=0N−1\(ℋr∩𝒜rc\)\)≤δ4\.\\mathbb\{P\}\\left\(\\bigcup\_\{r=0\}^\{N\-1\}\(\\mathcal\{H\}\_\{r\}\\cap\\mathcal\{A\}\_\{r\}^\{c\}\)\\right\)\\leq\\frac\{\\delta\}\{4\}\.\(136\)On the complement of this event, the induction closes\. Indeed, assumeℋr\\mathcal\{H\}\_\{r\}holds\. Then𝒜r\\mathcal\{A\}\_\{r\}also holds\. The eventℋr\\mathcal\{H\}\_\{r\}already contains the radius bounds forx0,…,xrx\_\{0\},\\ldots,x\_\{r\}and the gradient\-error bounds for blocks0,…,r−10,\\ldots,r\-1; adding𝒜r\\mathcal\{A\}\_\{r\}gives the gradient\-error bound for blockrr\. Hence the firstr\+1r\+1gradient errors are all bounded byϵ∇\\epsilon\_\{\\nabla\}\. Applying[Lemma˜21](https://arxiv.org/html/2606.26316#Thmlemma21)to thoser\+1r\+1updates gives

Δ​\(xr\+1\)≤Δ0\+8​ϵ∇2μ≤ℛ4\+ℛ4≤ℛ,\\Delta\(x\_\{r\+1\}\)\\leq\\Delta\_\{0\}\+\\frac\{8\\epsilon\_\{\\nabla\}^\{2\}\}\{\\mu\}\\leq\\frac\{\\mathcal\{R\}\}\{4\}\+\\frac\{\\mathcal\{R\}\}\{4\}\\leq\\mathcal\{R\},where we usedℛ≥4​Δ0\\mathcal\{R\}\\geq 4\\Delta\_\{0\}and \([134](https://arxiv.org/html/2606.26316#A8.E134)\)\. Thus the radius bounds hold for all0≤s≤r\+10\\leq s\\leq r\+1, and the gradient\-error bounds hold for all0≤s<r\+10\\leq s<r\+1\. Thereforeℋr\+1\\mathcal\{H\}\_\{r\+1\}holds\. Starting from the deterministic true eventℋ0\\mathcal\{H\}\_\{0\}, induction givesℋN\\mathcal\{H\}\_\{N\}with probability at least1−δ/41\-\\delta/4, hence at least1−δ1\-\\delta\.

On the eventℋN\\mathcal\{H\}\_\{N\}, all outer iterates remain in the radius\-ℛ\\mathcal\{R\}sublevel set and all gradient errors are bounded byϵ∇\\epsilon\_\{\\nabla\}\. Applying[Lemma˜21](https://arxiv.org/html/2606.26316#Thmlemma21)over allNNouter steps yields

Δ​\(xN\)≤\(1−μ16​L\)N​Δ0\+8​ϵ∇2μ\.\\Delta\(x\_\{N\}\)\\leq\\left\(1\-\\frac\{\\mu\}\{16L\}\\right\)^\{N\}\\Delta\_\{0\}\+\\frac\{8\\epsilon\_\{\\nabla\}^\{2\}\}\{\\mu\}\.Substituting \([133](https://arxiv.org/html/2606.26316#A8.E133)\) and absorbing the numerical factor into the statement constantcpc\_\{p\}gives \([47](https://arxiv.org/html/2606.26316#S5.E47)\)\. The radius conclusion is part ofℋN\\mathcal\{H\}\_\{N\}\. ∎

### H\.3Proof of Corollary[3](https://arxiv.org/html/2606.26316#Thmcorollary3): geometric transition\-budget rate

###### Proof of[Corollary˜3](https://arxiv.org/html/2606.26316#Thmcorollary3)\.

The choice ofmmin \([48](https://arxiv.org/html/2606.26316#S5.E48)\) gives

⌊mtmix⌋≥log2⁡\(64​d​N​Tδ\)\.\\left\\lfloor\\frac\{m\}\{t\_\{\\mathrm\{mix\}\}\}\\right\\rfloor\\geq\\log\_\{2\}\\left\(\\frac\{64dNT\}\{\\delta\}\\right\)\.Hence the geometric mixing condition \([12](https://arxiv.org/html/2606.26316#S3.E12)\) implies

ε​\(m\)≤δ64​d​N​T≤δ64​d​N​b,\\varepsilon\(m\)\\leq\\frac\{\\delta\}\{64dNT\}\\leq\\frac\{\\delta\}\{64dNb\},becauseb≤Tb\\leq T\. Thus the mixing admissibility condition of[Theorem˜3](https://arxiv.org/html/2606.26316#Thmtheorem3)holds\. The number of Markov transitions used is exactlyN​bNb, and the definitionb=⌊T/N⌋b=\\lfloor T/N\\rfloorgivesN​b≤TNb\\leq T\.

It remains to translatesht=m​uht/bs\_\{\\rm ht\}=mu\_\{\\rm ht\}/binto a transition\-budget expression\. Sinceb=⌊T/N⌋b=\\lfloor T/N\\rfloorand the hypothesism≤bm\\leq bimpliesb≥1b\\geq 1, we haveb≥T/\(2​N\)b\\geq T/\(2N\)\. The choices in \([48](https://arxiv.org/html/2606.26316#S5.E48)\) imply

m≤c​tmix​log⁡\(64​d​N​Tδ\),uht=log⁡\(64​d​N​\(m\+1\)δ\)≤polylog​\(T,1/δ,d,tmix,L/μ\)\.m\\leq ct\_\{\\mathrm\{mix\}\}\\log\\left\(\\frac\{64dNT\}\{\\delta\}\\right\),\\qquad u\_\{\\rm ht\}=\\log\\left\(\\frac\{64dN\(m\+1\)\}\{\\delta\}\\right\)\\leq\\mathrm\{polylog\}\(T,1/\\delta,d,t\_\{\\mathrm\{mix\}\},L/\\mu\)\.Therefore

sht=m​uhtb≤c​tmix​NT​polylog​\(T,1/δ,d,tmix,L/μ\)\.s\_\{\\rm ht\}=\\frac\{mu\_\{\\rm ht\}\}\{b\}\\leq c\\frac\{t\_\{\\mathrm\{mix\}\}N\}\{T\}\\,\\mathrm\{polylog\}\(T,1/\\delta,d,t\_\{\\mathrm\{mix\}\},L/\\mu\)\.The chosen number of outer iterations satisfies

\(1−μ16​L\)N≤exp⁡\(−μ​N16​L\)≤\(e​T\)−2≤1T,\\left\(1\-\\frac\{\\mu\}\{16L\}\\right\)^\{N\}\\leq\\exp\\left\(\-\\frac\{\\mu N\}\{16L\}\\right\)\\leq\(eT\)^\{\-2\}\\leq\\frac\{1\}\{T\},up to an immaterial numerical constant\. Substituting these estimates into \([47](https://arxiv.org/html/2606.26316#S5.E47)\) gives the first termO~​\(Δ0/T\)\\widetilde\{O\}\(\\Delta\_\{0\}/T\)and the second term in \([49](https://arxiv.org/html/2606.26316#S5.E49)\)\. All additional factors are logarithmic in the quantities listed in the corollary\. ∎

### H\.4Proof of Theorem[4](https://arxiv.org/html/2606.26316#Thmtheorem4): sticky\-chain lower bound

###### Proof of[Theorem˜4](https://arxiv.org/html/2606.26316#Thmtheorem4)\.

The lower bound is for any algorithm whose decisions are measurable with respect to the oracle transcript and its own internal randomness\. The algorithm is not told the signs∈\{\+,−\}s\\in\\\{\+,\-\\\}used below\. Thus, on events where the two oracle transcripts are identical, the algorithm must produce the same output under the two instances when its internal randomness is coupled\.

Letq=τ/\(8​n\)q=\\tau/\(8n\)\. Sincen≥8​τn\\geq 8\\tau, we haveq≤1/64q\\leq 1/64\. Set

A=σp​q−1/p,θ=q​A=σp​q\(p−1\)/p\.A=\\sigma\_\{p\}q^\{\-1/p\},\\qquad\\theta=qA=\\sigma\_\{p\}q^\{\(p\-1\)/p\}\.Define two refresh distributions onℝ\\mathbb\{R\}: underQ\+Q\_\{\+\}, the value isAAwith probabilityqqand0otherwise; underQ−Q\_\{\-\}, the value is−A\-Awith probabilityqqand0otherwise\. Their means are±θ\\pm\\theta\. Moreover,

𝔼Qs​\|Z−s​θ\|p≤2p−1​\(𝔼Qs​\|Z\|p\+\|θ\|p\)≤2p−1​\(q​Ap\+\(q​A\)p\)≤cp​σpp,\\mathbb\{E\}\_\{Q\_\{s\}\}\|Z\-s\\theta\|^\{p\}\\leq 2^\{p\-1\}\\left\(\\mathbb\{E\}\_\{Q\_\{s\}\}\|Z\|^\{p\}\+\|\\theta\|^\{p\}\\right\)\\leq 2^\{p\-1\}\\left\(qA^\{p\}\+\(qA\)^\{p\}\\right\)\\leq c\_\{p\}\\sigma\_\{p\}^\{p\},so the centered stationaryppth moment is at most a constant multiple ofσpp\\sigma\_\{p\}^\{p\}\.

For each signs∈\{\+,−\}s\\in\\\{\+,\-\\\}, define a sticky Markov chain as follows\. At each transition, with probability1/τ1/\\tauthe chain refreshes fromQsQ\_\{s\}, and with probability1−1/τ1\-1/\\tauit keeps its current value\. The invariant law isQsQ\_\{s\}\. From any starting state, aftermmsteps the probability of not having refreshed is\(1−1/τ\)m\(1\-1/\\tau\)^\{m\}; once a refresh occurs, the state has lawQsQ\_\{s\}\. Hence the total\-variation distance to stationarity is at most\(1−1/τ\)m\(1\-1/\\tau\)^\{m\}, so the chain has a valid geometric mixing time at mostc​τc\\taufor a universal constantcc\.

Consider the quadratic objective

fs​\(x\)=12​\(x−s​θ\)2f\_\{s\}\(x\)=\\frac\{1\}\{2\}\(x\-s\\theta\)^\{2\}and the oracle output, at query pointxtx\_\{t\},

Yt=xt−Zt\.Y\_\{t\}=x\_\{t\}\-Z\_\{t\}\.Under the invariant lawQsQ\_\{s\},𝔼​Zt=s​θ\\mathbb\{E\}Z\_\{t\}=s\\theta, hence

𝔼​\[Yt∣xt\]=xt−s​θ=∇fs​\(xt\)\.\\mathbb\{E\}\[Y\_\{t\}\\mid x\_\{t\}\]=x\_\{t\}\-s\\theta=\\nabla f\_\{s\}\(x\_\{t\}\)\.The finite\-moment condition in[Assumption˜4](https://arxiv.org/html/2606.26316#Thmassumption4)follows from the centered\-moment bound above, with moment scale enlarged by a universal constant\.

The chain is initialized at the fixed nonstationary stateZ0=0Z\_\{0\}=0\. Let𝒩\\mathcal\{N\}be the event that no nonzero refresh occurs during the firstnnobserved transitions\. In one transition, a nonzero refresh occurs with probability\(1/τ\)​q\(1/\\tau\)q, so under either sign

ℙs​\(𝒩\)=\(1−q/τ\)n=\(1−18​n\)n≥e−1/4,\\mathbb\{P\}\_\{s\}\(\\mathcal\{N\}\)=\(1\-q/\\tau\)^\{n\}=\\left\(1\-\\frac\{1\}\{8n\}\\right\)^\{n\}\\geq e^\{\-1/4\},after adjusting the universal constant in the lower bound\. On𝒩\\mathcal\{N\}, all observed Markov states are zero\. Consequently the complete oracle transcript is identical underQ\+Q\_\{\+\}andQ−Q\_\{\-\}: every oracle output isYt=xtY\_\{t\}=x\_\{t\}, and the algorithm’s internal randomness can be coupled identically in the two instances\.

Letx^n\\widehat\{x\}\_\{n\}be the algorithm’s output on this common transcript\. For every real numberyy, at least one of the two distances toθ\\thetaand−θ\-\\thetais at leastθ\\theta:

𝟏​\{\|y−θ\|≥θ\}\+𝟏​\{\|y\+θ\|≥θ\}≥1\.\\mathbf\{1\}\\\{\|y\-\\theta\|\\geq\\theta\\\}\+\\mathbf\{1\}\\\{\|y\+\\theta\|\\geq\\theta\\\}\\geq 1\.Averaging this inequality over the common transcript on𝒩\\mathcal\{N\}and over the algorithm’s randomness gives the two\-point testing bound

maxs∈\{\+,−\}⁡ℙs​\(\(x^n−s​θ\)2≥θ2\)≥12​mins∈\{\+,−\}⁡ℙs​\(𝒩\)≥12​e−1/4\.\\max\_\{s\\in\\\{\+,\-\\\}\}\\mathbb\{P\}\_\{s\}\\left\(\(\\widehat\{x\}\_\{n\}\-s\\theta\)^\{2\}\\geq\\theta^\{2\}\\right\)\\geq\\frac\{1\}\{2\}\\min\_\{s\\in\\\{\+,\-\\\}\}\\mathbb\{P\}\_\{s\}\(\\mathcal\{N\}\)\\geq\\frac\{1\}\{2\}e^\{\-1/4\}\.Sincefs​\(x^n\)−fs⋆=12​\(x^n−s​θ\)2f\_\{s\}\(\\widehat\{x\}\_\{n\}\)\-f\_\{s\}^\{\\star\}=\\frac\{1\}\{2\}\(\\widehat\{x\}\_\{n\}\-s\\theta\)^\{2\}, for at least one sign,

ℙs​\(fs​\(x^n\)−fs⋆≥θ22\)≥12​e−1/4\.\\mathbb\{P\}\_\{s\}\\left\(f\_\{s\}\(\\widehat\{x\}\_\{n\}\)\-f\_\{s\}^\{\\star\}\\geq\\frac\{\\theta^\{2\}\}\{2\}\\right\)\\geq\\frac\{1\}\{2\}e^\{\-1/4\}\.Finally,

θ2=σp2​q2​\(p−1\)/p=σp2​\(τ8​n\)2​\(p−1\)/p\.\\theta^\{2\}=\\sigma\_\{p\}^\{2\}q^\{2\(p\-1\)/p\}=\\sigma\_\{p\}^\{2\}\\left\(\\frac\{\\tau\}\{8n\}\\right\)^\{2\(p\-1\)/p\}\.Absorbing numerical constants intoccandc0c\_\{0\}proves \([50](https://arxiv.org/html/2606.26316#S5.E50)\)\. ∎

Similar Articles

The Sharp Tail of Uniform Stability

arXiv cs.LG

This paper presents a new logarithmic-free upper bound for the generalization gap in uniformly stable algorithms and constructs a deterministic learning problem that achieves optimal high-probability dependence, closing a gap in the literature.