Beyond Bounded Variance: Variance-Reduced Normalized Methods for Nonconvex Optimization under Blum-Gladyshev Noise
Summary
This paper studies nonconvex stochastic optimization under Blum-Gladyshev noise, where gradient variance grows with distance from initialization. It proves convergence guarantees for normalized SGD with momentum and a variance-reduced STORM method, achieving minimax optimal rates under certain conditions.
View Cached Full Text
Cached at: 05/18/26, 06:39 AM
# Beyond Bounded Variance: Variance-Reduced Normalized Methods for Nonconvex Optimization under Blum–Gladyshev Noise
Source: [https://arxiv.org/html/2605.15314](https://arxiv.org/html/2605.15314)
Antesh Upadhyay, Arda Fazla, Abolfazl HashemiAuthors with the School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN 47907, USA\.
###### Abstract
We study nonconvex stochastic optimization under the Blum–Gladyshev \(𝖡𝖦\\mathsf\{BG\}\-0\) noise model, where the stochastic gradient variance grows quadratically with the distance from the initialization\. We consider this problem under both standard smoothness and the symmetric generalized\-smoothness framework\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], which captures objectives whose local curvature can scale with the gradient norm\. We prove that normalized stochastic gradient descent with momentum, using only one stochastic gradient per iteration, converges under𝖡𝖦\\mathsf\{BG\}\-0 noise with oracle complexity𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)\. This rate holds both for standard smoothness and forα\\alpha\-symmetric generalized smoothness, showing that generalized smoothness is rate\-neutral for normalized momentum in this setting\. We then study a variance\-reduced normalized STORM method\. Under mean\-square smoothness and sharp initialization, the method achieves the minimax optimal𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)complexity, matching the lower bound\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]\. Under expectedα\\alpha\-symmetric generalized smoothness, the STORM recursion couples gradient\-dependent smoothness with distance\-dependent noise, leading to complexity𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)forα∈\(0,1\)\\alpha\\in\(0,1\)and𝒪\(ε−5\)\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)forα=1\\alpha=1\. When the distance\-growth parameter in the noise model vanishes, our guarantees recover the standard bounded\-variance rates:𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)for momentum,𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)for variance reduction, and𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)in the deterministic case\. To our knowledge, these are the first convergence guarantees for normalized methods in non\-convex stochastic optimization under𝖡𝖦\\mathsf\{BG\}\-0 noise without bounded domains, increasing batch sizes, or explicit anchoring, covering both standard and generalized smoothness regimes\.
## 1Introduction
Large\-scale nonconvex learning is typically formulated as the stochastic optimization problem
minx∈ℝdf\(x\),f\(x\):=𝔼ξ\[f\(x;ξ\)\],\\min\_\{x\\in\\mathbb\{R\}^\{d\}\}f\(x\),\\qquad f\(x\):=\\mathbb\{E\}\_\{\\xi\}\[f\(x;\\xi\)\],\(1\)wheref:ℝd→ℝf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}is generally nonconvex\. The goal is to find anε\\varepsilon\-stationary point, i\.e\., a pointxxsatisfying𝔼‖∇f\(x\)‖≤ε\\mathbb\{E\}\\\|\\nabla f\(x\)\\\|\\leq\\varepsilon\.
Under standard smoothness and uniformly bounded stochastic noise, the complexity theory for this problem is well developed\. In the deterministic setting, gradient descent achieves the optimal𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)iteration complexity for smooth nonconvex optimization\[[34](https://arxiv.org/html/2605.15314#bib.bib56),[35](https://arxiv.org/html/2605.15314#bib.bib112)\]\. In the stochastic setting, SGD achieves the classical𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)sample complexity under bounded variance𝔼ξ‖∇f\(x;ξ\)−∇f\(x\)‖2≤σ2,\\mathbb\{E\}\_\{\\xi\}\\bigl\\\|\\nabla f\(x;\\xi\)\-\\nabla f\(x\)\\bigr\\\|^\{2\}\\leq\\sigma^\{2\},and this rate is unimprovable for general smooth nonconvex stochastic optimization\[[14](https://arxiv.org/html/2605.15314#bib.bib50),[2](https://arxiv.org/html/2605.15314#bib.bib29)\]\. Under stronger mean\-square smoothness or average\-smoothness conditions, variance\-reduced methods such as SVRG, SAGA, SARAH, SPIDER, SpiderBoost, SNVRG, PAGE, and STORM improve the complexity to the optimal𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)\[[23](https://arxiv.org/html/2605.15314#bib.bib137),[11](https://arxiv.org/html/2605.15314#bib.bib138),[36](https://arxiv.org/html/2605.15314#bib.bib12),[12](https://arxiv.org/html/2605.15314#bib.bib45),[42](https://arxiv.org/html/2605.15314#bib.bib68),[46](https://arxiv.org/html/2605.15314#bib.bib135),[31](https://arxiv.org/html/2605.15314#bib.bib16),[10](https://arxiv.org/html/2605.15314#bib.bib41)\]\.
These classical stochastic guarantees typically rely on a uniform bounded\-noise model\. While mathematically convenient, this assumption can fail even in elementary unconstrained stochastic optimization problems, such as least squares, and in modern deep learning, where the stochastic gradient variance is not uniformly constant but can grow as the iterate moves farther from a reference point\. This motivates the𝖡𝖦\\mathsf\{BG\}\-0 noise model\[[4](https://arxiv.org/html/2605.15314#bib.bib79),[15](https://arxiv.org/html/2605.15314#bib.bib93)\], which has recently been identified as one of the weakest viable variance assumptions for theoretical analysis\[[1](https://arxiv.org/html/2605.15314#bib.bib129)\]\. Under𝖡𝖦\\mathsf\{BG\}\-0, the variance is allowed to grow quadratically with the distance from the initialization𝔼ξ‖∇f\(x;ξ\)−∇f\(x\)‖2≤B2‖x−x0‖2\+G2\.\\mathbb\{E\}\_\{\\xi\}\\bigl\\\|\\nabla f\(x;\\xi\)\-\\nabla f\(x\)\\bigr\\\|^\{2\}\\leq B^\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G^\{2\}\.The classical bounded variance setting is recovered by takingB=0B=0, while the deterministic case corresponds toB=G=0B=G=0\.
Although𝖡𝖦\\mathsf\{BG\}\-0 is more permissive than bounded variance, it creates an inherent feedback mechanism\. If the iterates move far fromx0x^\{0\}, then the oracle becomes noisier; the larger stochastic noise can then perturb future directions and push the trajectory even farther fromx0x^\{0\}\. Thus, over unbounded domains, the variance scaleB2‖xk−x0‖2\+G2B^\{2\}\\\|x^\{k\}\-x^\{0\}\\\|^\{2\}\+G^\{2\}cannot be treated as a harmless constant\. RecentlyFazlaet al\.\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]show that this difficulty is intrinsic: under𝖡𝖦\\mathsf\{BG\}\-0 noise, the lower bound for smooth nonconvex optimization worsens from the classicalΩ\(ε−4\)\\Omega\(\\varepsilon^\{\-4\}\)toΩ\(ε−6\)\\Omega\(\\varepsilon^\{\-6\}\), and even under mean\-square smoothness, the best possible rate isΩ\(ε−4\)\\Omega\(\\varepsilon^\{\-4\}\)rather than the bounded varianceΩ\(ε−3\)\\Omega\(\\varepsilon^\{\-3\}\)rate\[[2](https://arxiv.org/html/2605.15314#bib.bib29)\]\. They also proposed PASTA, which matches these limits by combining Halpern\-style anchoring\[[18](https://arxiv.org/html/2605.15314#bib.bib100)\], Tikhonov regularization, and dynamic batching, connecting𝖡𝖦\\mathsf\{BG\}\-0 stochastic approximation with classical fixed\-point schemes such as Halpern and Krasnoselskii–Mann iterations\[[32](https://arxiv.org/html/2605.15314#bib.bib21),[28](https://arxiv.org/html/2605.15314#bib.bib20),[18](https://arxiv.org/html/2605.15314#bib.bib100),[13](https://arxiv.org/html/2605.15314#bib.bib139)\]\.
In this paper, we take a complementary viewpoint\. Rather than adding an explicit stabilizing anchor or increasing per\-iteration batch sizes, we ask whether normalized stochastic methods with momentum or recursive variance reduction can handle𝖡𝖦\\mathsf\{BG\}\-0 noise without bounded domains\.
Why Normalization?The key observation is that normalization decouples the length of the update from the magnitude of the stochastic gradient estimator\. Given an estimatorvkv^\{k\}, define
dk=vk‖vk‖,xk\+1=xk−γdk\.d^\{k\}=\\frac\{v^\{k\}\}\{\\\|v^\{k\}\\\|\},\\qquad x^\{k\+1\}=x^\{k\}\-\\gamma d^\{k\}\.Then,‖xk\+1−xk‖=γ‖dk‖≤γ,\\\|x^\{k\+1\}\-x^\{k\}\\\|=\\gamma\\\|d^\{k\}\\\|\\leq\\gamma,regardless of the size of‖vk‖\\\|v^\{k\}\\\|\. Consequently,‖xk−x0‖≤∑t=0k−1‖xt\+1−xt‖≤kγ\.\\\|x^\{k\}\-x^\{0\}\\\|\\leq\\sum\_\{t=0\}^\{k\-1\}\\\|x^\{t\+1\}\-x^\{t\}\\\|\\leq k\\gamma\.Thus, along a normalized trajectory, the𝖡𝖦\\mathsf\{BG\}\-0 variance scale satisfies\(B2‖xk−x0‖2\+G2\)1/2≤G\+Bkγ\.\\left\(B^\{2\}\\\|x^\{k\}\-x^\{0\}\\\|^\{2\}\+G^\{2\}\\right\)^\{1/2\}\\leq G\+Bk\\gamma\.Normalization, therefore, transforms the potentially uncontrolled effect of large stochastic gradient magnitudes into controlled trajectory growth\. In contrast, PASTA\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]controls the𝖡𝖦\\mathsf\{BG\}\-0 variance via dynamic batching\. There, the batch size must scale with the upper bound on the local variance, i\.e\.,Nk∝B2‖xk−x0‖2\+G2σ2N\_\{k\}\\ \\propto\\ \\tfrac\{B^\{2\}\\\|x^\{k\}\-x^\{0\}\\\|^\{2\}\+G^\{2\}\}\{\\sigma^\{2\}\}, up to problem\-dependent accuracy and step\-size factors\. Such a rule requires knowledge, or at least reliable estimates, of the variance\-growth constantsBBandGG, which may be difficult to obtain in practice\.
Generalized Smoothness\.A second challenge is that many learning objectives are not globally smooth, i\.e\., their local smoothness can grow with the gradient norm\[[22](https://arxiv.org/html/2605.15314#bib.bib140),[44](https://arxiv.org/html/2605.15314#bib.bib141)\]\. This is a natural setting for normalized updates\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]\. We therefore study𝖡𝖦\\mathsf\{BG\}\-0 noise under the deterministic classℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)and its stochastic analogue𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)fromChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], forα∈\(0,1\]\\alpha\\in\(0,1\]\. The classℒsym∗\(α\)\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(\\alpha\)contains standard smooth functions\(ℒ\)\(\\mathcal\{L\}\), asymmetric generalized smooth functions\(ℒasym∗\)\(\\mathcal\{L\}\_\{\\rm asym\}^\{\*\}\), and Hessian\-based generalized smooth functions\(ℒH∗\)\(\\mathcal\{L\}\_\{\\rm H\}^\{\*\}\), and also includes high\-order polynomial and exponential\-type objectives\[[7](https://arxiv.org/html/2605.15314#bib.bib142),[44](https://arxiv.org/html/2605.15314#bib.bib141),[29](https://arxiv.org/html/2605.15314#bib.bib143)\]\. Appendix[A](https://arxiv.org/html/2605.15314#A1)summarizes these definitions\.
Prior work shows that, under suitable bounded or relative variance assumptions, generalized smooth nonconvex optimization can be as efficient as smooth nonconvex optimization: normalized gradient methods recover the deterministic𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)rate, while SPIDER\-type variance reduction recovers the stochastic𝒪\(ε−3\\mathcal\{O\}\(\\varepsilon^\{\-3\}\) rate under expected generalized smoothness\[[7](https://arxiv.org/html/2605.15314#bib.bib142),[12](https://arxiv.org/html/2605.15314#bib.bib45)\]\. These results, however, do not address𝖡𝖦\\mathsf\{BG\}\-0 noise\. In this setting, the noise level is distance\-dependent, while under generalized smoothness, the local curvature is gradient\-dependent\. This creates a new interplay between trajectory growth and gradient growth that is absent from the bounded variance setting\. It is therefore unclear a priori whether the principle that generalized smooth optimization is “as efficient as” smooth optimization continues to hold under𝖡𝖦\\mathsf\{BG\}\-0 noise\. This paper studies exactly this interaction\.
### 1\.1Contributions
To this end, we develop a convergence theory for normalized stochastic methods under𝖡𝖦\\mathsf\{BG\}\-0 noise across standard and generalized smoothness regimes\. The results answer four questions\.
- Q1\. Can single\-sample normalized momentum converge under𝖡𝖦\\mathsf\{BG\}\-0? Yes\. We prove that normalized stochastic gradient descent with momentum \(𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}\) converges under𝖡𝖦\\mathsf\{BG\}\-0 using one fresh stochastic gradient sample per iteration\. The method does not require bounded domains, bounded stochastic gradients, bounded variance, dynamic batching, or explicit anchoring\. Under𝖡𝖦\\mathsf\{BG\}\-0 noise,𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}achieves oracle complexity of𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)\. Thus, normalization and momentum provide an implicit stabilization mechanism: normalization controls trajectory growth, while momentum stabilizes the noisy update direction\.
- Q2\. Does replacing standard smoothness with generalized smoothness worsen the oracle complexity of𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}under𝖡𝖦\\mathsf\{BG\}\-0? No\. For𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, generalized smoothness is rate\-neutral at the level of theε\\varepsilon\-complexity exponent: underℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\-type smoothness, the stochastic first\-order oracle \(SFO\) complexity remains𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\), up to constants depending on the generalized smoothness parameters\.
- Q3\. Can additional stochastic regularity recover the optimal𝖡𝖦\\mathsf\{BG\}\-0 rate? Yes\. Mean\-square smoothness \(𝖬𝖲𝖲\\mathsf\{MSS\}\) provides additional stochastic regularity by controlling differences of stochastic gradients𝔼ξ‖∇f\(y;ξ\)−∇f\(x;ξ\)‖2≤L2‖y−x‖2\.\\mathbb\{E\}\_\{\\xi\}\\bigl\\\|\\nabla f\(y;\\xi\)\-\\nabla f\(x;\\xi\)\\bigr\\\|^\{2\}\\leq L^\{2\}\\\|y\-x\\\|^\{2\}\.Under this stronger condition, normalized STORM \(𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}\) improves over the𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)single\-sample momentum rate\. With a one\-time sharp initialization batch, it achieves the optimal oracle complexity of𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\.
- Q4\. Does the optimal𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}rate persist under expected generalized smoothness? Not exactly\. Under mean\-square smoothness,𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}recovers the optimal𝖡𝖦\\mathsf\{BG\}\-0 rate𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\. Under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)generalized smoothness, however, the estimator\-difference recursion depends on gradient\-dependent smoothness terms, which interact with the distance\-dependent𝖡𝖦\\mathsf\{BG\}\-0 variance\. As a result,𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}achieves𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)forα∈\(0,1\)\\alpha\\in\(0,1\)and𝒪\(ε−5\)\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)forα=1\\alpha=1\. Thus, generalized smoothness is free for𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}at the level of theε\\varepsilon\-exponent, but it induces a quantifiableα\\alpha\-dependent price for variance\-reduced𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}under𝖡𝖦\\mathsf\{BG\}\-0\.
We also provide synthetic experiments on generalized smooth objectives fromChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]under a𝖡𝖦\\mathsf\{BG\}\-0 oracle, comparing normalized momentum and normalized variance reduction with dynamic\-batching baselines\. Finally, Table[1](https://arxiv.org/html/2605.15314#S1.T1)summarizes the comparison with prior work, emphasizing the combination of𝖡𝖦\\mathsf\{BG\}\-0 noise, normalized methods, and smoothness regimes studied here\.
Table 1:Comparison with representative works under different notions of smoothness\(ℒ,ℒH∗,ℒasym∗,ℒsym∗\(α\),𝖬𝖲𝖲,𝔼ℒsym∗\(α\)\)\(\\mathcal\{L\},\\mathcal\{L\}\_\{\\rm H\}^\{\*\},\\mathcal\{L\}\_\{\\rm asym\}^\{\*\},\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\),\\mathsf\{MSS\},\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\)\. Complexity is measured in oracle calls to find anε\\varepsilon\-stationary point\.WorkSmoothness class𝖡𝖦\\mathsf\{BG\}\-0?MethodOracle complexityCutkosky and Mehta\[[9](https://arxiv.org/html/2605.15314#bib.bib144)\]ℒ\\mathcal\{L\}✗𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)Zhanget al\.\[[44](https://arxiv.org/html/2605.15314#bib.bib141)\]ℒH∗\\mathcal\{L\}\_\{\\rm H\}^\{\*\}✗clipped SGD𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)Reisizadehet al\.\[[38](https://arxiv.org/html/2605.15314#bib.bib145)\]ℒasym∗\\mathcal\{L\}\_\{\\rm asym\}^\{\*\}✗SPIDER \+ clipping𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)Chenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\),𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)✗β\\beta\-normalized GD; SPIDERDet\. \(ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\):𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\); Stoch\. \(𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\):𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)Khiriratet al\.\[[26](https://arxiv.org/html/2605.15314#bib.bib154)\]ℒsym∗\(1\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\)✗𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)Fazlaet al\.\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]ℒ\\mathcal\{L\},𝖬𝖲𝖲\\mathsf\{MSS\}✓PASTAℒ\\mathcal\{L\}:Θ\(ε−6\)\\Theta\(\\varepsilon^\{\-6\}\); MSS:Θ\(ε−4\)\\Theta\(\\varepsilon^\{\-4\}\)This paperℒ\\mathcal\{L\},ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)✓𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)This paper𝖬𝖲𝖲\\mathsf\{MSS\},𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)✓𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}𝖬𝖲𝖲:𝒪\(ε−4\);𝔼ℒsym∗\(α\):𝒪\(ε−\(4\+α\)\);𝔼ℒsym∗\(1\):𝒪\(ε−5\)\\begin\{aligned\} \\mathsf\{MSS\}:\\mathcal\{O\}\(\\varepsilon^\{\-4\}\);\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\):\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\);\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\):\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)\\end\{aligned\}
Note\.WhenB=0B=0, our bounds recover the standard bounded variance rates:𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)for normalized momentum,𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)for variance reduction, and𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)in the deterministic case, for standard and generalized smoothness classes\.
## 2Related Work
From bounded variance to Blum–Gladyshev noise\.Standard nonconvex guarantees usually assume uniform bounds on variance or subgradients, a convenient but often unrealistic simplification\[[14](https://arxiv.org/html/2605.15314#bib.bib50),[33](https://arxiv.org/html/2605.15314#bib.bib110),[5](https://arxiv.org/html/2605.15314#bib.bib32),[3](https://arxiv.org/html/2605.15314#bib.bib76),[37](https://arxiv.org/html/2605.15314#bib.bib15),[39](https://arxiv.org/html/2605.15314#bib.bib120),[5](https://arxiv.org/html/2605.15314#bib.bib32),[21](https://arxiv.org/html/2605.15314#bib.bib7)\]\. More recent research explores noise models where variance scales with the gradient norm or iterate location\[[16](https://arxiv.org/html/2605.15314#bib.bib94),[24](https://arxiv.org/html/2605.15314#bib.bib104),[25](https://arxiv.org/html/2605.15314#bib.bib105),[20](https://arxiv.org/html/2605.15314#bib.bib101),[17](https://arxiv.org/html/2605.15314#bib.bib98),[1](https://arxiv.org/html/2605.15314#bib.bib129)\]\. A key example is the Blum–Gladyshev condition\[[4](https://arxiv.org/html/2605.15314#bib.bib79),[15](https://arxiv.org/html/2605.15314#bib.bib93)\], which allows variance to grow with the squared distance from a reference point\. This is particularly difficult in nonconvex settings, as there is no natural “pull” to keep iterates from drifting away\. While recent work like the PASTA framework\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]uses anchoring, regularization, and dynamic batching to stabilize this drift, achieving sharp rates ofΘ\(ε−6\)\\Theta\(\\varepsilon^\{\-6\}\)orΘ\(ε−4\)\\Theta\(\\varepsilon^\{\-4\}\), our approach takes a different path\. We utilize normalized updates to restrict movement, combined with momentum and recursive variance reduction to keep our stochastic directions stable\.
Normalization and clipping for nonuniform smoothness\.Clipping and normalization are widely used to stabilize first\-order methods when gradients or local curvature can be large\. Under\(L0,L1\)\(L\_\{0\},L\_\{1\}\)\-type smoothness, clipped methods have been analyzed in deterministic and stochastic settings\[[44](https://arxiv.org/html/2605.15314#bib.bib141),[27](https://arxiv.org/html/2605.15314#bib.bib146)\], with extensions to momentum, variance reduction, and adaptive stepsizes\[[43](https://arxiv.org/html/2605.15314#bib.bib147),[38](https://arxiv.org/html/2605.15314#bib.bib145),[41](https://arxiv.org/html/2605.15314#bib.bib148),[30](https://arxiv.org/html/2605.15314#bib.bib149),[40](https://arxiv.org/html/2605.15314#bib.bib150)\]\. Normalized methods, including normalized gradient descent, generalized SignSGD, and normalized momentum, provide a related way to control large updates\[[45](https://arxiv.org/html/2605.15314#bib.bib151),[19](https://arxiv.org/html/2605.15314#bib.bib152),[8](https://arxiv.org/html/2605.15314#bib.bib153)\]\. Recent work also studies normalized error\-feedback underℒsym∗\(1\)\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(1\)in distributed optimization\[[26](https://arxiv.org/html/2605.15314#bib.bib154)\]\. These results, however, rely on a stronger notion of variance bounds\. In contrast, we study normalized momentum and recursive variance reduction under distance\-dependent𝖡𝖦\\mathsf\{BG\}\-0 noise for the generalized smooth classesℒsym∗\(α\)\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(\\alpha\)and𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(\\alpha\)\.
Momentum and variance reduction\.Momentum and recursive variance reduction stabilize stochastic gradient estimates\. Momentum smooths sample noise\[[9](https://arxiv.org/html/2605.15314#bib.bib144)\], while methods such as SARAH, SPIDER, SpiderBoost, PAGE, and STORM use recursive estimators to obtain sharper control\[[36](https://arxiv.org/html/2605.15314#bib.bib12),[12](https://arxiv.org/html/2605.15314#bib.bib45),[42](https://arxiv.org/html/2605.15314#bib.bib68),[31](https://arxiv.org/html/2605.15314#bib.bib16),[10](https://arxiv.org/html/2605.15314#bib.bib41)\]\. Under bounded variance and mean\-squared smoothness, these methods improve the stochastic nonconvex complexity from𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)to𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)\. Under𝖡𝖦\\mathsf\{BG\}\-0 noise, this acceleration is no longer attainable in general; the optimal𝖬𝖲𝖲\\mathsf\{MSS\}rate becomes𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]\. PASTA reaches this rate using dynamic batching with PAGE\-type variance reduction\. In contrast, our normalized STORM analysis recovers the same𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)rate with single\-sample\-transition updates and sharp initialization\. Under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(\\alpha\), the interaction between recursive estimation and trajectory\-dependent variance yields𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\), with𝒪\(ε−5\)\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)atα=1\\alpha=1\.
## 3Preliminaries
Notation\.All norms∥⋅∥\\\|\\cdot\\\|are Euclidean norms, and⟨⋅,⋅⟩\\langle\\cdot,\\cdot\\rangledenotes the Euclidean inner product\. We use𝒪\(⋅\)\\mathcal\{O\}\(\\cdot\)to hide numerical constants independent of the target accuracyε\\varepsilon\. For a random variableXX,𝔼\[X\]\\mathbb\{E\}\[X\]denotes expectation over all algorithmic randomness\. We letℱk\\mathcal\{F\}\_\{k\}be the filtration generated by the algorithm up to iterationkk, and write𝔼k\[X\]:=𝔼\[X∣ℱk\]\\mathbb\{E\}\_\{k\}\[X\]:=\\mathbb\{E\}\[X\\mid\\mathcal\{F\}\_\{k\}\]\. We use𝔼ξ\\mathbb\{E\}\_\{\\xi\}in assumptions for expectations over an oracle sample at fixed query points\. We definefinf:=infx∈ℝdf\(x\)f^\{\\inf\}:=\\inf\_\{x\\in\\mathbb\{R\}^\{d\}\}f\(x\)andΔ:=f\(x0\)−finf\\Delta:=f\(x^\{0\}\)\-f^\{\\inf\}\. The output iterate is sampled uniformly from the generated trajectory, i\.e\.,x^∼\{x0,x1,…,xK\}\.\\widehat\{x\}\\sim\\\{x^\{0\},x^\{1\},\\dots,x^\{K\}\\\}\.
Problem Formulation\.We consider the problem in \([1](https://arxiv.org/html/2605.15314#S1.E1)\)\. The stochastic first\-order oracle returns∇f\(x;ξ\)\\nabla f\(x;\\xi\)at a queried pointxx\. We measure complexity by the number of SFO calls\. As mentioned previously, the goal is to find anε\\varepsilon\-stationary point, i\.e\., a pointxxsuch that𝔼‖∇f\(x\)‖≤ε\.\\mathbb\{E\}\\\|\\nabla f\(x\)\\\|\\leq\\varepsilon\.
Assumptions\.We start by stating the standard assumption used in our analysis\.
###### Assumption 1\(Lower boundedness offf\)\.
Functionffis bounded from below, i\.e\.,finf\>−∞f^\{\\inf\}\>\-\\infty\.
###### Assumption 2\(Unbiased𝖡𝖦\\mathsf\{BG\}\-0 stochastic oracle\)\.
For everyx∈ℝdx\\in\\mathbb\{R\}^\{d\}, the stochastic gradient oracle is unbiased, i\.e\.,𝔼ξ\[∇f\(x;ξ\)\]=∇f\(x\)\.\\mathbb\{E\}\_\{\\xi\}\[\\nabla f\(x;\\xi\)\]=\\nabla f\(x\)\.Moreover, there exist constantsB,G≥0B,G\\geq 0such that𝔼ξ‖∇f\(x;ξ\)−∇f\(x\)‖2≤B2‖x−x0‖2\+G2\\mathbb\{E\}\_\{\\xi\}\\left\\\|\\nabla f\(x;\\xi\)\-\\nabla f\(x\)\\right\\\|^\{2\}\\leq B^\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G^\{2\}\.
WhenB=0B=0, Assumption[2](https://arxiv.org/html/2605.15314#Thmassumption2)reduces to the classical bounded variance condition\. WhenB=G=0B=G=0, the oracle is deterministic\.
We pair this noise model with several smoothness conditions of increasing generality\. The first two are standard:L0L\_\{0\}\-smoothness controls population gradients\[[14](https://arxiv.org/html/2605.15314#bib.bib50)\], while mean\-square smoothness controls the stochastic gradient differences\[[10](https://arxiv.org/html/2605.15314#bib.bib41),[31](https://arxiv.org/html/2605.15314#bib.bib16)\]\.
###### Assumption 3\(L0L\_\{0\}\-smoothness offf\)\.
There existsL0\>0L\_\{0\}\>0such that for allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},‖∇f\(y\)−∇f\(x\)‖≤L0‖y−x‖\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|\\leq L\_\{0\}\\\|y\-x\\\|\.
###### Assumption 4\(Mean\-square smoothness\)\.
There existsL\>0L\>0such that for allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},𝔼ξ‖∇f\(y;ξ\)−∇f\(x;ξ\)‖2≤L2‖y−x‖2\.\\mathbb\{E\}\_\{\\xi\}\\left\\\|\\nabla f\(y;\\xi\)\-\\nabla f\(x;\\xi\)\\right\\\|^\{2\}\\leq L^\{2\}\\\|y\-x\\\|^\{2\}\.
The next two assumptions extend smoothness to the generalized setting ofChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], where the local smoothness parameter is allowed to grow with the gradient norm\. The deterministic classℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)governs the normalized momentum analysis, while the expected class𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)governs the normalized STORM analysis\.
###### Assumption 5\(ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\-type generalized smoothness\)\.
Letα∈\(0,1\]\\alpha\\in\(0,1\]\. We say thatf∈ℒsym∗\(α\)f\\in\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)iff:ℝd→ℝf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}is differentiable and there exist constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0such that, for allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},
‖∇f\(y\)−∇f\(x\)‖≤\(L0\+L1maxθ∈\[0,1\]‖∇f\(xθ\)‖α\)‖y−x‖,wherexθ:=θy\+\(1−θ\)x\.\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|\\leq\\left\(L\_\{0\}\+L\_\{1\}\\max\_\{\\theta\\in\[0,1\]\}\\\|\\nabla f\(x\_\{\\theta\}\)\\\|^\{\\alpha\}\\right\)\\\|y\-x\\\|,\\text\{ where \}x\_\{\\theta\}:=\\theta y\+\(1\-\\theta\)x\.
###### Assumption 6\(Expectedα\\alpha\-symmetric generalized smoothness\)\.
Letα∈\(0,1\]\\alpha\\in\(0,1\]\. We say thatf∈𝔼ℒsym∗\(α\)f\\in\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)if there exist constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0such that, for allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\}, wherexθ:=θy\+\(1−θ\)xx\_\{\\theta\}:=\\theta y\+\(1\-\\theta\)x, we have
𝔼ξ‖∇f\(y;ξ\)−∇f\(x;ξ\)‖2≤‖y−x‖2𝔼ξ\[\(L0\+L1maxθ∈\[0,1\]‖∇f\(xθ;ξ\)‖α\)2\]\.\\mathbb\{E\}\_\{\\xi\}\\big\\\|\\nabla f\(y;\\xi\)\-\\nabla f\(x;\\xi\)\\big\\\|^\{2\}\\leq\\\|y\-x\\\|^\{2\}\\mathbb\{E\}\_\{\\xi\}\\big\[\\big\(L\_\{0\}\+L\_\{1\}\\max\_\{\\theta\\in\[0,1\]\}\\\|\\nabla f\(x\_\{\\theta\};\\xi\)\\\|^\{\\alpha\}\\big\)^\{2\}\\big\]\.
Both classes include standard smoothness as the special caseL1=0L\_\{1\}=0and recover symmetric\(L0,L1\)\(L\_\{0\},L\_\{1\}\)\-generalized smoothness whenα=1\\alpha=1\[[26](https://arxiv.org/html/2605.15314#bib.bib154)\]; valuesα∈\(0,1\)\\alpha\\in\(0,1\)give sublinear dependence on the gradient norm\. The expected condition is more delicate under𝖡𝖦\\mathsf\{BG\}\-0 noise because it involves sample\-gradient moments, which must be controlled through the distance\-dependent variance scale\. This interaction is the source of theα\\alpha\-dependent price in our variance\-reduced rates\.
## 4Normalized Momentum under𝖡𝖦\\mathsf\{BG\}\-0
We first consider the normalized stochastic gradient descent with momentum, denoted by𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}\. Givenx0∈ℝdx^\{0\}\\in\\mathbb\{R\}^\{d\}, stepsizeγ\>0\\gamma\>0, momentum parameterη∈\(0,1\]\\eta\\in\(0,1\], and horizonK≥0K\\geq 0, initializev0=∇f\(x0;ξ0\)\.v^\{0\}=\\nabla f\(x^\{0\};\\xi^\{0\}\)\.Fork=0,1,…,Kk=0,1,\\ldots,K,𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}performs
𝖭𝖲𝖦𝖣𝖬:xk\+1=xk−γvk‖vk‖;vk\+1=\(1−η\)vk\+η∇f\(xk\+1;ξk\+1\),\\mathsf\{NSGDM:\}\\qquad x^\{k\+1\}=x^\{k\}\-\\gamma\\frac\{v^\{k\}\}\{\\\|v^\{k\}\\\|\};\\qquad v^\{k\+1\}=\(1\-\\eta\)v^\{k\}\+\\eta\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\),\(2\)where thevk\+1v^\{k\+1\}update is used only fork=0,1,…,K−1k=0,1,\\ldots,K\-1\. The samples are fresh and conditionally independent given the past, and the stochastic oracle satisfies Assumption[2](https://arxiv.org/html/2605.15314#Thmassumption2)\.
###### Theorem 1\(Normalized momentum under𝖡𝖦\\mathsf\{BG\}\-0\)\.
Consider problem \([1](https://arxiv.org/html/2605.15314#S1.E1)\) and the𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}update in \([2](https://arxiv.org/html/2605.15314#S4.E2)\)\. Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[2](https://arxiv.org/html/2605.15314#Thmassumption2)hold, and letη=1\(K\+1\)2/3,\\eta=\\frac\{1\}\{\(K\+1\)^\{2/3\}\},γ=γ0\(K\+1\)5/6,\\gamma=\\frac\{\\gamma\_\{0\}\}\{\(K\+1\)^\{5/6\}\},
whereγ0\>0\\gamma\_\{0\}\>0is a tuning constant\. Then the outputx^\\widehat\{x\}generated by𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, satisfies the following\.
1. \(i\)L0L\_\{0\}\-smooth case\.Suppose Assumption[3](https://arxiv.org/html/2605.15314#Thmassumption3)holds\. Then, for everyγ0\>0\\gamma\_\{0\}\>0, 𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(2Δγ0\+16L0γ0\+2Bγ0\)1\(K\+1\)1/6\+8G\(K\+1\)1/3\+2L0γ0\(K\+1\)5/6\.\\displaystyle\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+16L\_\{0\}\\gamma\_\{0\}\+2B\\gamma\_\{0\}\\right\)\\frac\{1\}\{\(K\+1\)^\{1/6\}\}\+\\frac\{8G\}\{\(K\+1\)^\{1/3\}\}\+\\frac\{2L\_\{0\}\\gamma\_\{0\}\}\{\(K\+1\)^\{5/6\}\}\.Consequently,𝔼‖∇f\(x^\)‖=𝒪\(\(K\+1\)−1/6\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\\\!\\left\(\(K\+1\)^\{\-1/6\}\\right\), giving𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)SFO complexity for𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}\.
2. \(ii\)ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)generalized smooth case withα∈\(0,1\)\\alpha\\in\(0,1\)\.Suppose Assumption[5](https://arxiv.org/html/2605.15314#Thmassumption5)holds for some fixedα∈\(0,1\)\\alpha\\in\(0,1\)with constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0\. DefineL¯0:=L0\(2α21−α\+1\)\\overline\{L\}\_\{0\}:=L\_\{0\}\\left\(2^\{\\frac\{\\alpha^\{2\}\}\{1\-\\alpha\}\}\+1\\right\),L¯1:=L12α21−α3α\\overline\{L\}\_\{1\}:=L\_\{1\}\\,2^\{\\frac\{\\alpha^\{2\}\}\{1\-\\alpha\}\}3^\{\\alpha\},L¯2:=L111−α2α21−α3α\(1−α\)α1−α\\overline\{L\}\_\{2\}:=L\_\{1\}^\{\\frac\{1\}\{1\-\\alpha\}\}2^\{\\frac\{\\alpha^\{2\}\}\{1\-\\alpha\}\}3^\{\\alpha\}\(1\-\\alpha\)^\{\\frac\{\\alpha\}\{1\-\\alpha\}\},Cα:=L¯2\+1−α2\(2α\)α1−αL¯111−αC\_\{\\alpha\}:=\\overline\{L\}\_\{2\}\+\\frac\{1\-\\alpha\}\{2\}\(2\\alpha\)^\{\\frac\{\\alpha\}\{1\-\\alpha\}\}\\overline\{L\}\_\{1\}^\{\\frac\{1\}\{1\-\\alpha\}\}, andC~α:=L¯0\+L¯2\+\(1−α\)\(8α\)α1−αL¯111−α\.\\widetilde\{C\}\_\{\\alpha\}:=\\overline\{L\}\_\{0\}\+\\overline\{L\}\_\{2\}\+\(1\-\\alpha\)\(8\\alpha\)^\{\\frac\{\\alpha\}\{1\-\\alpha\}\}\\overline\{L\}\_\{1\}^\{\\frac\{1\}\{1\-\\alpha\}\}\.Then, for every0<γ0≤10<\\gamma\_\{0\}\\leq 1, 𝔼‖∇f\(x^\)‖≤\(2Δγ0\+4C~αγ0\+2Bγ0\)1\(K\+1\)1/6\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+4\\widetilde\{C\}\_\{\\alpha\}\\gamma\_\{0\}\+2B\\gamma\_\{0\}\\right\)\\frac\{1\}\{\(K\+1\)^\{1/6\}\}\+8G\(K\+1\)1/3\+L¯0γ0\(K\+1\)5/6\\displaystyle\+\\frac\{8G\}\{\(K\+1\)^\{1/3\}\}\+\\frac\{\\overline\{L\}\_\{0\}\\gamma\_\{0\}\}\{\(K\+1\)^\{5/6\}\}\+2Cαγ011−α\(K\+1\)56\(1−α\)\.\\displaystyle\\quad\+\\frac\{2C\_\{\\alpha\}\\gamma\_\{0\}^\{\\frac\{1\}\{1\-\\alpha\}\}\}\{\(K\+1\)^\{\\frac\{5\}\{6\(1\-\\alpha\)\}\}\}\.Sinceα∈\(0,1\)\\alpha\\in\(0,1\)is fixed, every decay exponent on the right\-hand side is at least1/61/6\. Therefore,𝔼‖∇f\(x^\)‖=𝒪\(\(K\+1\)−1/6\),\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\\left\(\(K\+1\)^\{\-1/6\}\\right\),giving𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)SFO complexity for𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}\.
3. \(iii\)ℒsym∗\(1\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\)generalized smooth case\.Suppose Assumption[5](https://arxiv.org/html/2605.15314#Thmassumption5)holds withα=1\\alpha=1and constantsL0\>0L\_\{0\}\>0andL1\>0L\_\{1\}\>0\. Choose0<γ0≤18L1\.0<\\gamma\_\{0\}\\leq\\frac\{1\}\{8L\_\{1\}\}\.Then 𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(2Δγ0\+16L0γ0\+2Bγ0\)1\(K\+1\)1/6\+8G\(K\+1\)1/3\+2L0γ0\(K\+1\)5/6\\displaystyle\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+16L\_\{0\}\\gamma\_\{0\}\+2B\\gamma\_\{0\}\\right\)\\frac\{1\}\{\(K\+1\)^\{1/6\}\}\+\\frac\{8G\}\{\(K\+1\)^\{1/3\}\}\+\\frac\{2L\_\{0\}\\gamma\_\{0\}\}\{\(K\+1\)^\{5/6\}\}\+64L12γ0\(K\+1\)1/6\[4Δ\+\(16L0\+2B\)γ02\+8γ0G\(K\+1\)1/6\+2L0γ02\(K\+1\)2/3\]\.\\displaystyle\\quad\+\\frac\{64L\_\{1\}^\{2\}\\gamma\_\{0\}\}\{\(K\+1\)^\{1/6\}\}\\left\[4\\Delta\+\(16L\_\{0\}\+2B\)\\gamma\_\{0\}^\{2\}\+\\frac\{8\\gamma\_\{0\}G\}\{\(K\+1\)^\{1/6\}\}\+\\frac\{2L\_\{0\}\\gamma\_\{0\}^\{2\}\}\{\(K\+1\)^\{2/3\}\}\\right\]\.Consequently,𝔼‖∇f\(x^\)‖=𝒪\(\(K\+1\)−1/6\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\\\!\\left\(\(K\+1\)^\{\-1/6\}\\right\), giving𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)SFO complexity for𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}\.
Discussion on Theorem[1](https://arxiv.org/html/2605.15314#Thmtheorem1)\.We highlight several aspects of this result; the full proof is in Appendix[B](https://arxiv.org/html/2605.15314#A2)\.
•Generalized smoothness is rate\-neutral for𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}\.Under𝖡𝖦\\mathsf\{BG\}\-0 noise, the oracle complexity exponent𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)is identical across all three regimes: standardL0L\_\{0\}\-smoothness,ℒsym∗\(1\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\)generalized smoothness, andℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)generalized smoothness for any fixedα∈\(0,1\)\\alpha\\in\(0,1\)\. The generalized smoothness parameters affect only the leading constants, not theε\\varepsilon\-exponent\. This is consistent with the principle articulated byChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]for the deterministic setting, that generalized smooth nonconvex optimization is “as efficient as” standard smooth nonconvex optimization, and extends it to the stochastic momentum regime under𝖡𝖦\\mathsf\{BG\}\-0 noise\.
•Novelty of theα∈\(0,1\)\\alpha\\in\(0,1\)analysis\.To our knowledge, theℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)analysis for𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}withα∈\(0,1\)\\alpha\\in\(0,1\)is new even under bounded variance\. Prior stochastic results underα\\alpha\-symmetric generalized smoothness are limited to the two\-loop SPIDER estimator under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], which controls variance through periodic batch corrections\. The single\-sample momentum recursion requires us to control the estimator error over time\. After unrolling the recursion, this error separates into gradient\-drift terms and stochastic noise terms\. The key technical challenge is that the drift bound underℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)yields terms of the formL¯1γ‖∇f\(xk\)‖α\\overline\{L\}\_\{1\}\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}, which, unlike theα=1\\alpha=1case cannot be converted into suboptimality gaps\(f\(xk\)−finf\)\(f\(x^\{k\}\)\-f^\{\\inf\}\)via the gradient\-growth inequality\[[7](https://arxiv.org/html/2605.15314#bib.bib142),[26](https://arxiv.org/html/2605.15314#bib.bib154)\]: whenα=1\\alpha=1, the bound‖∇f\(x\)‖≤8L1\(f\(x\)−finf\)\+L0/L1\\\|\\nabla f\(x\)\\\|\\leq 8L\_\{1\}\(f\(x\)\-f^\{\\inf\}\)\+L\_\{0\}/L\_\{1\}directly absorbs‖∇f\(xk\)‖\\\|\\nabla f\(x^\{k\}\)\\\|into\(f\(xk\)−finf\)\(f\(x^\{k\}\)\-f^\{\\inf\}\), but forα∈\(0,1\)\\alpha\\in\(0,1\)the fractional power breaks this conversion\. Instead, we need to utilize Young’s inequality, giving us the free parameter to absorb the‖∇f\(xk\)‖α\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}contribution into a fraction of the‖∇f\(xk\)‖\\\|\\nabla f\(x^\{k\}\)\\\|descent budget, at the cost of producing remainder terms that scale asγ1/\(1−α\)\\gamma^\{1/\(1\-\\alpha\)\}\. These remainders vanish faster thanγ\\gammaand do not affect the dominant rate\.
•Recovery of known rates\.WhenB=0B=0, the𝖡𝖦\\mathsf\{BG\}\-0 condition reduces to bounded variance\. In this case, choosingη=\(K\+1\)−1/2\\eta=\(K\+1\)^\{\-1/2\}andγ=γ0\(K\+1\)−3/4\\gamma=\\gamma\_\{0\}\(K\+1\)^\{\-3/4\}recovers the𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)rate ofCutkosky and Mehta \[[9](https://arxiv.org/html/2605.15314#bib.bib144)\]for standard smoothness andKhiriratet al\.\[[26](https://arxiv.org/html/2605.15314#bib.bib154)\]forℒsym∗\(1\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\), and establishes the same rate forℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)withα∈\(0,1\)\\alpha\\in\(0,1\)\. Additionally whenG=0G=0as well, the oracle is deterministic, and withη=1\\eta=1recovers the𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)rate ofChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]for normalized GD\.
•Schedule comparison\.Under bounded variance, normalized momentum uses the standard scheduleη=\(K\+1\)−1/2\\eta=\(K\+1\)^\{\-1/2\}andγ=γ0\(K\+1\)−3/4,\\gamma=\\gamma\_\{0\}\(K\+1\)^\{\-3/4\},leading to the rate\(K\+1\)−1/4\(K\+1\)^\{\-1/4\}, or equivalently𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)SFO complexity\[[9](https://arxiv.org/html/2605.15314#bib.bib144),[26](https://arxiv.org/html/2605.15314#bib.bib154)\]\. The faster decay ofγ=γ0\(K\+1\)−5/6\\gamma=\\gamma\_\{0\}\(K\+1\)^\{\-5/6\}under𝖡𝖦\\mathsf\{BG\}\-0 reflects the growing noise scaleG\+BkγG\+Bk\\gamma\. Likewise,η=\(K\+1\)−2/3\\eta=\(K\+1\)^\{\-2/3\}is chosen smaller to control the stochastic error termη\(G\+BKγ\)\\sqrt\{\\eta\}\(G\+BK\\gamma\)\. This changes the rate from the bounded variance rate\(K\+1\)−1/4\(K\+1\)^\{\-1/4\}to the𝖡𝖦\\mathsf\{BG\}\-0 rate\(K\+1\)−1/6\(K\+1\)^\{\-1/6\}, matching the lower\-bound degradation fromΩ\(ε−4\)\\Omega\(\\varepsilon^\{\-4\}\)toΩ\(ε−6\)\\Omega\(\\varepsilon^\{\-6\}\)\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]\.
## 5Normalized STORM under𝖡𝖦\\mathsf\{BG\}\-0
We next study the normalized STORM recursion under𝖡𝖦\\mathsf\{BG\}\-0 noise\. Unlike𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, which averages stochastic gradients directly,𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}uses a recursive estimator based on the difference of two stochastic gradients evaluated with the same fresh sample\.
Fixx0∈ℝdx^\{0\}\\in\\mathbb\{R\}^\{d\}, a stepsizeγ\>0\\gamma\>0, a STORM parameterη∈\(0,1\]\\eta\\in\(0,1\], and a horizonK≥0K\\geq 0\. LetNinit≥1N\_\{\\rm init\}\\geq 1be an initialization batch size, and initializev0=1Ninit∑i=1Ninit∇f\(x0;ξiinit\),v^\{0\}=\\frac\{1\}\{N\_\{\\rm init\}\}\\sum\_\{i=1\}^\{N\_\{\\rm init\}\}\\nabla f\(x^\{0\};\\xi\_\{i\}^\{\\rm init\}\),where the samples\{ξiinit\}i=1Ninit\\\{\\xi\_\{i\}^\{\\rm init\}\\\}\_\{i=1\}^\{N\_\{\\rm init\}\}are independent\. Fork=0,1,…,Kk=0,1,\\ldots,K,𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}performs
𝖭𝖲𝖳𝖮𝖱𝖬:xk\+1=xk−γvk‖vk‖;vk\+1=∇f\(xk\+1;ξk\+1\)\+\(1−η\)\(vk−∇f\(xk;ξk\+1\)\)\.\\mathsf\{NSTORM:\}\\quad x^\{k\+1\}=x^\{k\}\-\\gamma\\frac\{v^\{k\}\}\{\\\|v^\{k\}\\\|\};\\ \\ v^\{k\+1\}=\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\)\+\(1\-\\eta\)\\left\(v^\{k\}\-\\nabla f\(x^\{k\};\\xi^\{k\+1\}\)\\right\)\.\(3\)The sampleξk\+1\\xi^\{k\+1\}is fresh and conditionally independent given the past, and is evaluated at bothxkx^\{k\}andxk\+1x^\{k\+1\}\. Thus, each transition uses two stochastic gradient evaluations\. The initialization batch only controls the initial error𝔼‖v0−∇f\(x0\)‖\\mathbb\{E\}\\\|v^\{0\}\-\\nabla f\(x^\{0\}\)\\\|, andNinitN\_\{\\rm init\}is specified for the smoothness regime\.
###### Theorem 2\(Normalized STORM under𝖡𝖦\\mathsf\{BG\}\-0\)\.
Consider problem \([1](https://arxiv.org/html/2605.15314#S1.E1)\) and the𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}update in \([3](https://arxiv.org/html/2605.15314#S5.E3)\)\. Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[2](https://arxiv.org/html/2605.15314#Thmassumption2)hold\. Then the outputx^\\widehat\{x\}generated by𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}, satisfies
1. \(i\)Mean\-square smoothness case\.Suppose Assumption[4](https://arxiv.org/html/2605.15314#Thmassumption4)holds with constantL\>0L\>0\. ChooseNinit:=max\{1,⌈G2\(K\+1\)1/2⌉\}N\_\{\\rm init\}:=\\max\\left\\\{1,\\left\\lceil G^\{2\}\(K\+1\)^\{1/2\}\\right\\rceil\\right\\\},η=\(K\+1\)−1\\eta=\(K\+1\)^\{\-1\}, andγ=γ0\(K\+1\)−3/4,\\gamma=\\gamma\_\{0\}\(K\+1\)^\{\-3/4\},whereγ0\>0\\gamma\_\{0\}\>0is fixed\. Then 𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(Δγ0\+2\(1\+Lγ0\+Bγ0\)\)\(K\+1\)−14\+2G\(K\+1\)−12\+Lγ02\(K\+1\)−34\.\\displaystyle\\leq\\left\(\\frac\{\\Delta\}\{\\gamma\_\{0\}\}\+2\(1\+L\\gamma\_\{0\}\+B\\gamma\_\{0\}\)\\right\)\(K\+1\)^\{\-\\tfrac\{1\}\{4\}\}\+2G\(K\+1\)^\{\-\\tfrac\{1\}\{2\}\}\+\\frac\{L\\gamma\_\{0\}\}\{2\}\(K\+1\)^\{\-\\tfrac\{3\}\{4\}\}\.Consequently,𝔼‖∇f\(x^\)‖=𝒪\(\(K\+1\)−1/4\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(\(K\+1\)^\{\-1/4\}\), giving𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)SFO complexity for𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}\.
2. \(ii\)𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(\\alpha\)generalized smoothness case withα∈\(0,1\)\\alpha\\in\(0,1\)\.Suppose Assumption[6](https://arxiv.org/html/2605.15314#Thmassumption6), holds forα∈\(0,1\)\\alpha\\in\(0,1\)withL0,L1\>0L\_\{0\},L\_\{1\}\>0\. DefineK0:=22−α1−αL0,K\_\{0\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{0\},K1:=22−α1−αL1,K\_\{1\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{1\},K2:=\(5L1\)11−α,K\_\{2\}:=\(5L\_\{1\}\)^\{\\frac\{1\}\{1\-\\alpha\}\},andT:=\(K\+1\)T:=\(K\+1\)\. ChooseNinit:=max\{1,⌈G2T2\(1−α\)4\+α⌉\}N\_\{\\rm init\}:=\\max\\left\\\{1,\\left\\lceil G^\{2\}T^\{\\frac\{2\(1\-\\alpha\)\}\{4\+\\alpha\}\}\\right\\rceil\\right\\\},γ=γ0T−3\+α4\+α,\\gamma=\\gamma\_\{0\}T^\{\-\\frac\{3\+\\alpha\}\{4\+\\alpha\}\},η=η0T−44\+α,\\eta=\\eta\_\{0\}T^\{\-\\frac\{4\}\{4\+\\alpha\}\},λ=λ0T−1−α4\+α,\\lambda=\\lambda\_\{0\}T^\{\-\\frac\{1\-\\alpha\}\{4\+\\alpha\}\},whereη0∈\(0,1\]\\eta\_\{0\}\\in\(0,1\],γ0\>0\\gamma\_\{0\}\>0, andλ0\>0\\lambda\_\{0\}\>0are constants satisfying2α2K1λ0γ0≤1,2^\{\\tfrac\{\\alpha\}\{2\}\}K\_\{1\}\\lambda\_\{0\}\\gamma\_\{0\}\\leq 1,26\+α2K1λ0γ0η0≤1\.2^\{\\tfrac\{6\+\\alpha\}\{2\}\}K\_\{1\}\\lambda\_\{0\}\\frac\{\\gamma\_\{0\}\}\{\\eta\_\{0\}\}\\leq 1\.Then 𝔼∥∇f\(x^\)∥≤\[4Δγ0\+8η0\+8\(1−α\)αα1−α2α/2K1γ0λ0−α1−αη0\+8⋅2α/2K1Bαγ01\+αη0\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\Bigg\[\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\+\\frac\{8\}\{\\eta\_\{0\}\}\+\\frac\{8\(1\-\\alpha\)\\alpha^\{\\frac\{\\alpha\}\{1\-\\alpha\}\}2^\{\\alpha/2\}K\_\{1\}\\gamma\_\{0\}\\lambda\_\{0\}^\{\-\\frac\{\\alpha\}\{1\-\\alpha\}\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\+\\frac\{8\\cdot 2^\{\\alpha/2\}K\_\{1\}B^\{\\alpha\}\\gamma\_\{0\}^\{1\+\\alpha\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\+8η0Bγ0\]T−14\+α\+\[8K0γ0η0\+8⋅2α/2K1Gαγ0η0\]T−1\+α4\+α\+8η0GT−24\+α\\displaystyle\+8\\sqrt\{\\eta\_\{0\}\}B\\gamma\_\{0\}\\Bigg\]T^\{\-\\frac\{1\}\{4\+\\alpha\}\}\+\\left\[\\frac\{8K\_\{0\}\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\+\\frac\{8\\cdot 2^\{\\alpha/2\}K\_\{1\}G^\{\\alpha\}\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\\right\]T^\{\-\\frac\{1\+\\alpha\}\{4\+\\alpha\}\}\+8\\sqrt\{\\eta\_\{0\}\}GT^\{\-\\frac\{2\}\{4\+\\alpha\}\}\+8K2γ011−αη0T−1\+3α\(1−α\)\(4\+α\)\+2K0γ0T−3\+α4\+α\+2\(1−α\)αα1−α2α/2K1γ0λ0−α1−αT−34\+α\\displaystyle\+\\frac\{8K\_\{2\}\\gamma\_\{0\}^\{\\frac\{1\}\{1\-\\alpha\}\}\}\{\\sqrt\{\\eta\_\{0\}\}\}T^\{\-\\frac\{1\+3\\alpha\}\{\(1\-\\alpha\)\(4\+\\alpha\)\}\}\+2K\_\{0\}\\gamma\_\{0\}T^\{\-\\frac\{3\+\\alpha\}\{4\+\\alpha\}\}\+2\(1\-\\alpha\)\\alpha^\{\\frac\{\\alpha\}\{1\-\\alpha\}\}2^\{\\alpha/2\}K\_\{1\}\\gamma\_\{0\}\\lambda\_\{0\}^\{\-\\frac\{\\alpha\}\{1\-\\alpha\}\}T^\{\-\\frac\{3\}\{4\+\\alpha\}\}\+4K2γ011−αT−3\+α\(1−α\)\(4\+α\)\+21\+α/2K1Gαγ0T−3\+α4\+α\+21\+α/2K1Bαγ01\+αT−34\+α\.\\displaystyle\+4K\_\{2\}\\gamma\_\{0\}^\{\\frac\{1\}\{1\-\\alpha\}\}T^\{\-\\frac\{3\+\\alpha\}\{\(1\-\\alpha\)\(4\+\\alpha\)\}\}\+2^\{1\+\\alpha/2\}K\_\{1\}G^\{\\alpha\}\\gamma\_\{0\}T^\{\-\\frac\{3\+\\alpha\}\{4\+\\alpha\}\}\+2^\{1\+\\alpha/2\}K\_\{1\}B^\{\\alpha\}\\gamma\_\{0\}^\{1\+\\alpha\}T^\{\-\\frac\{3\}\{4\+\\alpha\}\}\.Sinceα∈\(0,1\)\\alpha\\in\(0,1\)is fixed, every decay exponent on the right\-hand side is at least1\(4\+α\)\\frac\{1\}\{\(4\+\\alpha\)\}\. Consequently,𝔼‖∇f\(x^\)‖=𝒪\(T−1/\(4\+α\)\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\\left\(T^\{\-1/\(4\+\\alpha\)\}\\right\), giving𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)SFO complexity for𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}\.
3. \(iii\)𝔼ℒsym∗\(1\)\\mathbb\{E\}\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(1\)generalized smoothness case\.Suppose Assumption[6](https://arxiv.org/html/2605.15314#Thmassumption6), holds forα=1\\alpha=1withL0,L1\>0L\_\{0\},L\_\{1\}\>0\. Use one\-sample initialization, so thatb0:=𝔼‖v0−∇f\(x0\)‖≤G\.b\_\{0\}:=\\mathbb\{E\}\\\|v^\{0\}\-\\nabla f\(x^\{0\}\)\\\|\\leq G\.Now, chooseη=\(K\+1\)−4/5\\eta=\(K\+1\)^\{\-4/5\}andγ=γ0\(K\+1\)−4/5\\gamma=\\gamma\_\{0\}\(K\+1\)^\{\-4/5\}, where0<γ0≤1162e3/4L1\.0<\\gamma\_\{0\}\\leq\\frac\{1\}\{16\\sqrt\{2e^\{3/4\}\}L\_\{1\}\}\.Then 𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(4Δγ0\+8b0\+8Bγ0\+162e3/4L1Bγ02\)\(K\+1\)−1/5\\displaystyle\\leq\\left\(\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\+8b\_\{0\}\+8B\\gamma\_\{0\}\+16\\sqrt\{2e^\{3/4\}\}\\,L\_\{1\}B\\gamma\_\{0\}^\{2\}\\right\)\(K\+1\)^\{\-1/5\}\+\(82e3/4γ0\(L0\+2L1G\)\+8G\)\(K\+1\)−2/5\\displaystyle\+\\left\(8\\sqrt\{2e^\{3/4\}\}\\,\\gamma\_\{0\}\(L\_\{0\}\+2L\_\{1\}G\)\+8G\\right\)\(K\+1\)^\{\-2/5\}\+82L1Bγ02\(K\+1\)−3/5\+\(42L0γ0\+82L1Gγ0\)\(K\+1\)−4/5\.\\displaystyle\+8\\sqrt\{2\}\\,L\_\{1\}B\\gamma\_\{0\}^\{2\}\(K\+1\)^\{\-3/5\}\+\\left\(4\\sqrt\{2\}L\_\{0\}\\gamma\_\{0\}\+8\\sqrt\{2\}L\_\{1\}G\\gamma\_\{0\}\\right\)\(K\+1\)^\{\-4/5\}\.Consequently,𝔼‖∇f\(x^\)‖=𝒪\(\(K\+1\)−1/5\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(\(K\+1\)^\{\-1/5\}\), giving𝒪\(ε−5\)\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)SFO complexity for𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}\.
Discussion on Theorem[2](https://arxiv.org/html/2605.15314#Thmtheorem2)\.We now discuss several aspects of Theorem[2](https://arxiv.org/html/2605.15314#Thmtheorem2)\. The complete details of the proof, including the proofs on bounded variance and deterministic case, are present in Appendix[C](https://arxiv.org/html/2605.15314#A3)\.
•Optimality under mean\-square smoothness\.Under Assumption[4](https://arxiv.org/html/2605.15314#Thmassumption4), i\.e\.,𝖬𝖲𝖲\\mathsf\{MSS\}and sharp initialization,𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}achieves𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)SFO complexity which is minimax optimal in this setting\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]\. The key structural reason is that𝖬𝖲𝖲\\mathsf\{MSS\}gives𝔼ξ‖∇f\(y;ξ\)−∇f\(x;ξ\)‖2≤L2‖y−x‖2,\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(y;\\xi\)\-\\nabla f\(x;\\xi\)\\\|^\{2\}\\leq L^\{2\}\\\|y\-x\\\|^\{2\},with no dependence on gradient magnitudes\. Hence the STORM centered\-noise term satisfies\(𝔼\[‖Zk\+1‖2∣ℱk\]\)1/2≤Lγ,\\left\(\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq L\\gamma,whereZk\+1:=\(∇f\(xk\+1;ξk\+1\)−∇f\(xk;ξk\+1\)\)−\(∇f\(xk\+1\)−∇f\(xk\)\)\.Z\_\{k\+1\}:=\\left\(\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\};\\xi^\{k\+1\}\)\\right\)\-\\left\(\\nabla f\(x^\{k\+1\}\)\-\\nabla f\(x^\{k\}\)\\right\)\.So, the estimator recursion is decoupled from the gradient sequence\. Furthermore, sharp initialization removes the initial\-error bottleneckb0/\(ηT\)b\_\{0\}/\(\\eta T\), allowing the scheduleη=\(K\+1\)−1\\eta=\(K\+1\)^\{\-1\}andγ=γ0\(K\+1\)−3/4\\gamma=\\gamma\_\{0\}\(K\+1\)^\{\-3/4\}and yielding the rate\(K\+1\)−1/4\(K\+1\)^\{\-1/4\}\.
•Gradient feedback and theα\\alpha\-dependent price under expected generalized smoothness\.Under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\), the STORM centered\-difference bound onZk\+1Z\_\{k\+1\}depends on the sample\-gradient moment𝔼ξ‖∇f\(xk;ξ\)‖α\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x^\{k\};\\xi\)\\\|^\{\\alpha\}via the expected reduction ofChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]\(see Proposition 4, item \(1\)–\(2\)\)\. Under bounded variance, this moment is controlled by a fixed noise term, but under𝖡𝖦\\mathsf\{BG\}\-0 it couples a gradient\-dependent term with trajectory\-dependent variance, the main reason the analysis ofChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]cannot be applied directly\. To handle this coupling, we use Young’s inequality to linearize the fractional\-gradient term:‖∇f\(xk\)‖α≤λ‖∇f\(xk\)‖\+cαλ−α/\(1−α\),\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\\leq\\lambda\\\|\\nabla f\(x^\{k\}\)\\\|\+c\_\{\\alpha\}\\lambda^\{\-\\alpha/\(1\-\\alpha\)\},yielding an estimator bound of the form\(𝔼\[‖Zk\+1‖2∣ℱk\]\)1/2≤γ\(a\+h‖∇f\(xk\)‖\)\(\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\)^\{1/2\}\\leq\\gamma\(a\+h\\\|\\nabla f\(x^\{k\}\)\\\|\)withh∝λh\\propto\\lambda\. The gradient\-dependent part creates a feedback loop: large gradients inflate the estimator error, which in turn weakens the descent guarantee\. Our proof separates the estimator error into a gradient\-independent part, controlled by the𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}estimator recursion, and the gradient\-dependent part, which is absorbed into the descent inequality via the stepsize conditions\. Choosingλ\\lambdato decay with the horizon weakens this feedback and leads to the𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)rate forα∈\(0,1\)\\alpha\\in\(0,1\)\. Atα=1\\alpha=1, the term‖∇f\(xk\)‖α=‖∇f\(xk\)‖\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}=\\\|\\nabla f\(x^\{k\}\)\\\|is already linear, so there is no fractional power to absorb via Young’s inequality; the feedback coefficienthhis then a fixed constant, yielding the slower𝒪\(ε−5\)\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)rate\.
•Sharp initialization and its dependence on the smoothness regime\.The initialization batch sizeNinitN\_\{\\rm init\}is determined by a single mechanism: whether the scheduleη\\etasuppresses the initial\-error contributionb0/\(ηT\)b\_\{0\}/\(\\eta T\)at the target convergence rate\. Under𝖬𝖲𝖲\\mathsf\{MSS\}, the clean STORM recursion permitsη=T−1\\eta=T^\{\-1\}, which yieldsb0/\(ηT\)=𝒪\(b0\)b\_\{0\}/\(\\eta T\)=\\mathcal\{O\}\(b\_\{0\}\), aTT\-independent term, forcingNinit=𝒪\(T1/2\)N\_\{\\rm init\}=\\mathcal\{O\}\(T^\{\\nicefrac\{\{1\}\}\{\{2\}\}\}\)to suppressb0b\_\{0\}to𝒪\(T−1/4\)\\mathcal\{O\}\(T^\{\\nicefrac\{\{\-1\}\}\{\{4\}\}\}\)\. Under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)withα∈\(0,1\)\\alpha\\in\(0,1\), the gradient\-dependent feedback requiresη=η0T−4/\(4\+α\)\\eta=\\eta\_\{0\}T^\{\\nicefrac\{\{\-4\}\}\{\{\(4\+\\alpha\)\}\}\}, givingb0/\(ηT\)=𝒪\(b0T−α/\(4\+α\)\)\{b\_\{0\}/\(\\eta T\)\}=\\mathcal\{O\}\(b\_\{0\}T^\{\\nicefrac\{\{\-\\alpha\}\}\{\{\(4\+\\alpha\)\}\}\}\), which decays more slowly than the target rateT−1/\(4\+α\)T^\{\\nicefrac\{\{\-1\}\}\{\{\(4\+\\alpha\)\}\}\}\. This necessitatesNinit=𝒪\(T2\(1−α\)/\(4\+α\)\)N\_\{\\rm init\}=\\mathcal\{O\}\(T^\{\\nicefrac\{\{2\(1\-\\alpha\)\}\}\{\{\(4\+\\alpha\)\}\}\}\), whose exponent decreases to zero asα→1−\\alpha\\to 1^\{\-\}and grows toward1/21/2asα→0\+\\alpha\\to 0^\{\+\}\. Atα=1\\alpha=1, the scheduleη=T−4/5\\eta=T^\{\-4/5\}makesb0/\(ηT\)=𝒪\(T−1/5\)b\_\{0\}/\(\\eta T\)=\\mathcal\{O\}\(T^\{\\nicefrac\{\{\-1\}\}\{\{5\}\}\}\)match the target rate, so one\-sample initialization is sufficient\. Thus, sharp initialization is needed whenever the schedule is aggressive relative to the target rate\.
•Comparison with prior variance\-reduced methods\.Chenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]prove that SPIDER achieves𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)in the bounded variance setting\.Reisizadehet al\.\[[38](https://arxiv.org/html/2605.15314#bib.bib145)\]obtain the same rate under asymmetric generalized smoothness using clipped SPIDER\. Our setting differs in two ways: the estimator is single\-loop STORM rather than two\-loop SPIDER, and the noise is distance\-dependent rather than bounded\. WhenB=0B=0, the𝖡𝖦\\mathsf\{BG\}\-0 trajectory\-growth terms disappear, and the bounded variance specialization recovers the standard𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)rate in our setting as well\.
## 6Experiments
We evaluate the methods analyzed in this paper on two experiments whose deterministic objectives satisfyα\\alpha\-symmetric generalized smoothness\. In both cases, we use objectives fromChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]and apply a stochastic𝖡𝖦\\mathsf\{BG\}\-0 wrapper that injects distance\-dependent noise at each query point\. The wrapper constructs a proper stochastic loss whose oracle is unbiased, satisfies the𝖡𝖦\\mathsf\{BG\}\-0 variance model exactly, and preserves the𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)condition \(see Appendix[D\.1](https://arxiv.org/html/2605.15314#A4.SS1)\)\.
Across both experiments, we compare𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}and𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}using a single fresh sample per iteration against𝖲𝖦𝖣\\mathsf\{SGD\}and𝖲𝖳𝖮𝖱𝖬\\mathsf\{STORM\}with dynamic batching followingFazlaet al\.\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]\. Hyperparameters for the normalized methods are set according to Theorems[1](https://arxiv.org/html/2605.15314#Thmtheorem1)and[2](https://arxiv.org/html/2605.15314#Thmtheorem2); full details are in Appendix[E](https://arxiv.org/html/2605.15314#A5)\. Results are averaged over three random seeds and reported as mean±\\pmstandard deviation\.
•Phase retrieval\(Figure[1](https://arxiv.org/html/2605.15314#S6.F1)\)\. We minimizef\(x\)=12m∑r=1m\(yr−\|ar⊤x\|2\)2f\(x\)=\\frac\{1\}\{2m\}\\sum\_\{r=1\}^\{m\}\(y\_\{r\}\-\|a\_\{r\}^\{\\top\}x\|^\{2\}\)^\{2\}withm=3000m=3000Gaussian measurementsar∈ℝ100a\_\{r\}\\in\\mathbb\{R\}^\{100\}\(ar,j∼𝒩\(0,0\.01\)a\_\{r,j\}\\sim\\mathcal\{N\}\(0,0\.01\)\), target signalx⋆,j∼𝒩\(0,1\)x\_\{\\star,j\}\\sim\\mathcal\{N\}\(0,1\), and noiseless observationsyr=\|ar⊤x⋆\|2y\_\{r\}=\|a\_\{r\}^\{\\top\}x\_\{\\star\}\|^\{2\}; the objective satisfiesℒsym∗\(23\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\tfrac\{2\}\{3\}\)\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]\. We initialize atxj0∼𝒩\(5,1\)x^\{0\}\_\{j\}\\sim\\mathcal\{N\}\(5,1\), far from the target, so that𝖡𝖦\\mathsf\{BG\}\-0 variance growth is visible\.𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}reaches a lower gradient norm than𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, consistent with the improved exponent of variance reduction, while the dynamic\-batching baselines require batch sizes growing to∼103\\sim\\\!10^\{3\}\(right panel\)\. The middle panel confirms that normalization controls trajectory growth:‖xk−x0‖2\\\|x^\{k\}\\\!\-\\\!x^\{0\}\\\|^\{2\}stabilizes for both normalized methods, while single\-sample𝖲𝖦𝖣\\mathsf\{SGD\}drifts until the dynamic batch size compensates\.
Figure 1:*Phase retrieval experiment under the𝖡𝖦\\mathsf\{BG\}\-0 oracle \(α=2/3\\alpha\\\!=\\\!2/3\)\.*Gradient norm vs\. SFO calls \(left\), drift‖xk−x0‖2\\\|x^\{k\}\\\!\-\\\!x^\{0\}\\\|^\{2\}\(center\), and batch size \(right\)\.Figure 2:*Cubic polynomial under the𝖡𝖦\\mathsf\{BG\}\-0 oracle \(α=1/2\\alpha\\\!=\\\!1/2\)\.*Same layout as Figure[1](https://arxiv.org/html/2605.15314#S6.F1)\.𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}reaches a lower gradient norm than𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, reflecting the benefit of variance reduction\.•Cubic polynomial\(Figure[2](https://arxiv.org/html/2605.15314#S6.F2)\)\. In this experiment we minimizef\(x\)=\|x\|3f\(x\)=\|x\|^\{3\}onℝ\\mathbb\{R\}, which satisfiesℒsym∗\(12\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\tfrac\{1\}\{2\}\)\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], and initialize away from the targetx⋆=0\.0x^\{\\star\}=0\.0atx0∼𝒩\(5,0\.1\)x^\{0\}\\sim\\mathcal\{N\}\(5,0\.1\)\.𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}again reaches a lower gradient norm than𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, reflecting the benefit of variance reduction\. The dynamic batching baselines require batch sizes that grow by over an order of magnitude \(right panel\), whereas the normalized methods maintain batch size one throughout\.
Across both experiments, normalized methods and dynamic batching offer two complementary mechanisms for handling𝖡𝖦\\mathsf\{BG\}\-0 noise on generalized\-smooth landscapes: normalization controls trajectory growth through bounded step norms, while dynamic batching controls the effective noise by increasing oracle calls as the distance from initialization grows\. A practical advantage of the normalized approach is that it does not require knowledge of the variance\-growth constantsBBandGGto set the batch schedule\.
## 7Conclusion
We developed a convergence theory for normalized stochastic methods under𝖡𝖦\\mathsf\{BG\}\-0 noise across standard and generalized smoothness regimes\. For𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, we showed that normalization and momentum provide implicit stabilization against distance\-dependent noise, achieving𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)with a single sample per iteration and no anchoring or dynamic batching, which is minimax optimal\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]\. Generalized smoothness achieves the same rate\. For𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}, variance reduction recovers the optimal𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)rate under mean\-square smoothness with sharp initialization\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\], while expectedα\\alpha\-symmetric generalized smoothness yields𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)forα∈\(0,1\)\\alpha\\in\(0,1\)and𝒪\(ε−5\)\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)forα=1\\alpha=1\.
Theα\\alpha\-dependent price in𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}raises a natural question: is it information\-theoretically optimal, or is there a better rate? The price seems to originate from a coupling intrinsic to the problem structure\. Under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(\\alpha\)with𝖡𝖦\\mathsf\{BG\}\-0, stochastic gradient differences depend simultaneously on‖∇f\(x\)‖α\\\|\\nabla f\(x\)\\\|^\{\\alpha\}from generalized smoothness and on the distance\-dependent variance\. This coupling vanishes when either source of difficulty is removed\. We conjecture that it would confront any variance\-reduced method that bounds stochastic gradient differences under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\. Thus, it remains an interesting open avenue to determine whether one can come up with a matching lower bound, or an alternative idea that can help curb this under generalized smoothness, leading to the optimal𝒪\(ε−4\)\\mathcal\{O\}\(\{\\varepsilon^\{\-4\}\)\}rate\.
## References
- \[1\]\(2025\)Towards weaker variance assumptions for stochastic optimization\.arXiv preprint arXiv:2504\.09951\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p3.5),[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[2\]Y\. Arjevani, Y\. Carmon, J\. C\. Duchi, D\. J\. Foster, N\. Srebro, and B\. Woodworth\(2023\)Lower bounds for non\-convex stochastic optimization\.Mathematical Programming199\(1\),pp\. 165–214\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p2.4),[§1](https://arxiv.org/html/2605.15314#S1.p4.10)\.
- \[3\]F\. Bach and E\. Moulines\(2011\)Non\-asymptotic analysis of stochastic approximation algorithms for machine learning\.InAdvances in Neural Information Processing Systems,Vol\.24\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[4\]J\. R\. Blum\(1954\)Approximation methods which converge with probability one\.The Annals of Mathematical Statistics,pp\. 382–386\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p3.5),[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[5\]L\. Bottou, F\. E\. Curtis, and J\. Nocedal\(2018\)Optimization methods for large\-scale learning\.SIAM Review60\(2\),pp\. 223–311\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[6\]Y\. Carmon, J\. C\. Duchi, O\. Hinder, and A\. Sidford\(2019\)Lower bounds for finding stationary points I\.Mathematical Programming177,pp\. 193–230\.Cited by:[Appendix A](https://arxiv.org/html/2605.15314#A1.SS0.SSS0.Px9.p1.8)\.
- \[7\]Z\. Chen, Y\. Zhou, Y\. Liang, and Z\. Lu\(2023\)Generalized\-smooth nonconvex optimization is as efficient as smooth nonconvex optimization\.InInternational Conference on Machine Learning,pp\. 5396–5427\.Cited by:[Appendix A](https://arxiv.org/html/2605.15314#A1.SS0.SSS0.Px4.p1.4),[Appendix A](https://arxiv.org/html/2605.15314#A1.SS0.SSS0.Px9.p1.8),[Appendix A](https://arxiv.org/html/2605.15314#A1.p1.5),[2nd item](https://arxiv.org/html/2605.15314#A2.I1.i2.p1.2),[§B\.2](https://arxiv.org/html/2605.15314#A2.SS2.p1.2),[§B\.3](https://arxiv.org/html/2605.15314#A2.SS3.p2.2),[1st item](https://arxiv.org/html/2605.15314#A3.I3.i1.p1.4),[2nd item](https://arxiv.org/html/2605.15314#A3.I3.i2.p1.3),[§C\.2](https://arxiv.org/html/2605.15314#A3.SS2.3.p1.1),[§C\.7](https://arxiv.org/html/2605.15314#A3.SS7.p1.10),[§D\.2](https://arxiv.org/html/2605.15314#A4.SS2.p1.1),[§D\.2](https://arxiv.org/html/2605.15314#A4.SS2.p1.12),[§D\.3](https://arxiv.org/html/2605.15314#A4.SS3.p1.7),[§1\.1](https://arxiv.org/html/2605.15314#S1.SS1.p1.3),[Table 1](https://arxiv.org/html/2605.15314#S1.T1.19.15.15.8.1.1),[§1](https://arxiv.org/html/2605.15314#S1.p7.8),[§1](https://arxiv.org/html/2605.15314#S1.p8.4),[§3](https://arxiv.org/html/2605.15314#S3.p6.2),[§4](https://arxiv.org/html/2605.15314#S4.p3.9),[§4](https://arxiv.org/html/2605.15314#S4.p4.19),[§4](https://arxiv.org/html/2605.15314#S4.p5.11),[§5](https://arxiv.org/html/2605.15314#S5.p5.16),[§5](https://arxiv.org/html/2605.15314#S5.p7.5),[§6](https://arxiv.org/html/2605.15314#S6.p1.4),[§6](https://arxiv.org/html/2605.15314#S6.p3.14),[§6](https://arxiv.org/html/2605.15314#S6.p4.7)\.
- \[8\]M\. Crawshaw, M\. Liu, F\. Orabona, W\. Zhang, and Z\. Zhuang\(2022\)Robustness to unbounded smoothness of generalized signsgd\.Advances in neural information processing systems35,pp\. 9955–9968\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[9\]A\. Cutkosky and H\. Mehta\(2020\)Momentum improves normalized sgd\.InInternational conference on machine learning,pp\. 2260–2268\.Cited by:[1st item](https://arxiv.org/html/2605.15314#A2.I1.i1.p1.7),[Table 1](https://arxiv.org/html/2605.15314#S1.T1.8.4.4.4.1.1),[§2](https://arxiv.org/html/2605.15314#S2.p3.10),[§4](https://arxiv.org/html/2605.15314#S4.p5.11),[§4](https://arxiv.org/html/2605.15314#S4.p6.14)\.
- \[10\]A\. Cutkosky and F\. Orabona\(2019\)Momentum\-based variance reduction in non\-convex SGD\.InAdvances in Neural Information Processing Systems,Cited by:[1st item](https://arxiv.org/html/2605.15314#A3.I3.i1.p1.4),[§1](https://arxiv.org/html/2605.15314#S1.p2.4),[§2](https://arxiv.org/html/2605.15314#S2.p3.10),[§3](https://arxiv.org/html/2605.15314#S3.p5.1)\.
- \[11\]A\. Defazio, F\. Bach, and S\. Lacoste\-Julien\(2014\)SAGA: a fast incremental gradient method with support for non\-strongly convex composite objectives\.Advances in neural information processing systems27\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p2.4)\.
- \[12\]C\. Fang, C\. J\. Li, Z\. Lin, and T\. Zhang\(2018\)SPIDER: near\-optimal non\-convex optimization via stochastic path\-integrated differential estimator\.InAdvances in Neural Information Processing Systems,pp\. 689–699\.Cited by:[1st item](https://arxiv.org/html/2605.15314#A3.I3.i1.p1.4),[§C\.7](https://arxiv.org/html/2605.15314#A3.SS7.p1.10),[§1](https://arxiv.org/html/2605.15314#S1.p2.4),[§1](https://arxiv.org/html/2605.15314#S1.p8.4),[§2](https://arxiv.org/html/2605.15314#S2.p3.10)\.
- \[13\]A\. Fazla, E\. C\. Kaya, A\. Upadhyay, and A\. Hashemi\(2026\)Lower bounds and proximally anchored sgd for non\-convex minimization under unbounded variance\.arXiv preprint arXiv:2604\.16620\.Cited by:[§D\.1](https://arxiv.org/html/2605.15314#A4.SS1.p3.2),[Table 1](https://arxiv.org/html/2605.15314#S1.T1.27.23.23.6.1.1),[§1](https://arxiv.org/html/2605.15314#S1.p4.10),[§1](https://arxiv.org/html/2605.15314#S1.p6.10),[§2](https://arxiv.org/html/2605.15314#S2.p1.2),[§2](https://arxiv.org/html/2605.15314#S2.p3.10),[§4](https://arxiv.org/html/2605.15314#S4.p6.14),[§5](https://arxiv.org/html/2605.15314#S5.p4.11),[§6](https://arxiv.org/html/2605.15314#S6.p2.5),[§7](https://arxiv.org/html/2605.15314#S7.p1.10)\.
- \[14\]S\. Ghadimi and G\. Lan\(2013\)Stochastic first\- and zeroth\-order methods for nonconvex stochastic programming\.SIAM Journal on Optimization23\(4\),pp\. 2341–2368\.Cited by:[Appendix A](https://arxiv.org/html/2605.15314#A1.SS0.SSS0.Px9.p1.8),[§1](https://arxiv.org/html/2605.15314#S1.p2.4),[§2](https://arxiv.org/html/2605.15314#S2.p1.2),[§3](https://arxiv.org/html/2605.15314#S3.p5.1)\.
- \[15\]E\. Gladyshev\(1965\)On stochastic approximation\.Theory of Probability & Its Applications10\(2\),pp\. 275–278\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p3.5),[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[16\]E\. Gorbunov, F\. Hanzely, and P\. Richtárik\(2020\)A unified theory of SGD: variance reduction, sampling, quantization and coordinate descent\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 680–690\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[17\]B\. Grimmer\(2019\)Convergence rates for deterministic and stochastic subgradient methods without Lipschitz continuity\.SIAM Journal on Optimization29\(2\),pp\. 1350–1365\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[18\]B\. Halpern\(1967\)Fixed points of nonexpanding maps\.Bulletin of the American Mathematical Society73\(6\),pp\. 957–961\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p4.10)\.
- \[19\]F\. Hübler, J\. Yang, X\. Li, and N\. He\(2024\)Parameter\-agnostic optimization under relaxed smoothness\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 4861–4869\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[20\]S\. Ilandarideva, A\. Juditsky, G\. Lan, and T\. Li\(2023\)Accelerated stochastic approximation with state\-dependent noise\.arXiv preprint arXiv:2307\.01497\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[21\]P\. Jain, S\. M\. Kakade, R\. Kidambi, P\. Netrapalli, and A\. Sidford\(2018\)Parallelizing stochastic gradient descent for least squares regression: mini\-batching, averaging, and model misspecification\.Journal of machine learning research18\(223\),pp\. 1–42\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[22\]J\. Jin, B\. Zhang, H\. Wang, and L\. Wang\(2021\)Non\-convex distributionally robust optimization: non\-asymptotic analysis\.Advances in Neural Information Processing Systems34,pp\. 2771–2782\.Cited by:[Appendix A](https://arxiv.org/html/2605.15314#A1.SS0.SSS0.Px9.p1.8),[§1](https://arxiv.org/html/2605.15314#S1.p7.8)\.
- \[23\]R\. Johnson and T\. Zhang\(2013\)Accelerating stochastic gradient descent using predictive variance reduction\.Advances in neural information processing systems26\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p2.4)\.
- \[24\]A\. Khaled and P\. Richtárik\(2023\)Better theory for SGD in the nonconvex world\.Transactions on Machine Learning Research\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[25\]A\. Khaled, O\. Sebbouh, N\. Loizou, R\. M\. Gower, and P\. Richtárik\(2023\)Unified analysis of stochastic gradient methods for composite convex and smooth optimization\.Journal of Optimization Theory and Applications199\(2\),pp\. 499–540\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[26\]S\. Khirirat, A\. Sadiev, A\. Riabinin, E\. Gorbunov, and P\. Richtárik\(2026\)Error feedback under $\(l\_0,l\_1\)$\-smoothness: normalization and momentum\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=LmcTbBvgjP)Cited by:[1st item](https://arxiv.org/html/2605.15314#A2.I1.i1.p1.7),[§B\.2](https://arxiv.org/html/2605.15314#A2.SS2.p1.4),[Table 1](https://arxiv.org/html/2605.15314#S1.T1.22.18.18.4.1.1),[§2](https://arxiv.org/html/2605.15314#S2.p2.5),[§3](https://arxiv.org/html/2605.15314#S3.p7.6),[§4](https://arxiv.org/html/2605.15314#S4.p4.19),[§4](https://arxiv.org/html/2605.15314#S4.p5.11),[§4](https://arxiv.org/html/2605.15314#S4.p6.14)\.
- \[27\]A\. Koloskova, H\. Hendrikx, and S\. U\. Stich\(2023\)Revisiting gradient clipping: stochastic bias and tight convergence guarantees\.InInternational Conference on Machine Learning,pp\. 17343–17363\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[28\]M\. A\. Krasnosel’skiĭ\(1955\)Two remarks on the method of successive approximations\.Uspekhi matematicheskikh nauk10\(1\),pp\. 123–127\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p4.10)\.
- \[29\]D\. Levy, Y\. Carmon, J\. C\. Duchi, and A\. Sidford\(2020\)Large\-scale methods for distributionally robust optimization\.Advances in neural information processing systems33,pp\. 8847–8860\.Cited by:[Appendix A](https://arxiv.org/html/2605.15314#A1.SS0.SSS0.Px9.p1.8),[§1](https://arxiv.org/html/2605.15314#S1.p7.8)\.
- \[30\]H\. Li, A\. Rakhlin, and A\. Jadbabaie\(2023\)Convergence of adam under relaxed assumptions\.Advances in Neural Information Processing Systems36,pp\. 52166–52196\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[31\]Z\. Li, H\. Bao, X\. Zhang, and P\. Richtárik\(2021\)PAGE: a simple and optimal probabilistic gradient estimator for nonconvex optimization\.InInternational conference on machine learning,pp\. 6286–6295\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p2.4),[§2](https://arxiv.org/html/2605.15314#S2.p3.10),[§3](https://arxiv.org/html/2605.15314#S3.p5.1)\.
- \[32\]W\. R\. Mann\(1953\)Mean value methods in iteration\.Proceedings of the American Mathematical Society4\(3\),pp\. 506–510\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p4.10)\.
- \[33\]A\. Nemirovski, A\. Juditsky, G\. Lan, and A\. Shapiro\(2009\)Robust stochastic approximation approach to stochastic programming\.SIAM Journal on Optimization19\(4\),pp\. 1574–1609\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[34\]Y\. Nesterov\(2004\)Introductory lectures on convex optimization: a basic course\.Springer Science & Business Media\.Cited by:[2nd item](https://arxiv.org/html/2605.15314#A3.I3.i2.p1.3),[§1](https://arxiv.org/html/2605.15314#S1.p2.4)\.
- \[35\]Y\. Nesterov\(2018\)Lectures on convex optimization\.Vol\.137,Springer\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p2.4)\.
- \[36\]L\. M\. Nguyen, J\. Liu, K\. Scheinberg, and M\. Takáč\(2017\)SARAH: a novel method for machine learning problems using stochastic recursive gradient\.InInternational conference on machine learning,pp\. 2613–2621\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p2.4),[§2](https://arxiv.org/html/2605.15314#S2.p3.10)\.
- \[37\]A\. Rakhlin, O\. Shamir, and K\. Sridharan\(2012\)Making gradient descent optimal for strongly convex stochastic optimization\.InProceedings of the 29th International Coference on International Conference on Machine Learning,pp\. 1571–1578\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[38\]A\. Reisizadeh, H\. Li, S\. Das, and A\. Jadbabaie\(2025\)Variance\-reduced clipping for non\-convex optimization\.InICASSP 2025\-2025 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 1–5\.Cited by:[Appendix A](https://arxiv.org/html/2605.15314#A1.SS0.SSS0.Px9.p1.8),[Table 1](https://arxiv.org/html/2605.15314#S1.T1.12.8.8.3.1.1),[§2](https://arxiv.org/html/2605.15314#S2.p2.5),[§5](https://arxiv.org/html/2605.15314#S5.p7.5)\.
- \[39\]O\. Shamir and T\. Zhang\(2013\)Stochastic gradient descent for non\-smooth optimization: convergence results and optimal averaging schemes\.InInternational Conference on Machine Learning,pp\. 71–79\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p1.2)\.
- \[40\]Y\. Takezawa, H\. Bao, R\. Sato, K\. Niwa, and M\. Yamada\(2024\)Parameter\-free clipped gradient descent meets polyak\.Advances in Neural Information Processing Systems37,pp\. 44575–44599\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[41\]B\. Wang, Y\. Zhang, H\. Zhang, Q\. Meng, R\. Sun, Z\. Ma, T\. Liu, Z\. Luo, and W\. Chen\(2024\)Provable adaptivity of adam under non\-uniform smoothness\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 2960–2969\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[42\]Z\. Wang, K\. Ji, Y\. Zhou, Y\. Liang, and V\. Tarokh\(2018\)Spiderboost: a class of faster variance\-reduced algorithms for nonconvex optimization\.arXiv preprint arXiv:1810\.10690\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p2.4),[§2](https://arxiv.org/html/2605.15314#S2.p3.10)\.
- \[43\]B\. Zhang, J\. Jin, C\. Fang, and L\. Wang\(2020\)Improved analysis of clipping algorithms for non\-convex optimization\.Advances in Neural Information Processing Systems33,pp\. 15511–15521\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[44\]J\. Zhang, T\. He, S\. Sra, and A\. Jadbabaie\(2019\)Why gradient clipping accelerates training: a theoretical justification for adaptivity\.arXiv preprint arXiv:1905\.11881\.Cited by:[Appendix A](https://arxiv.org/html/2605.15314#A1.SS0.SSS0.Px9.p1.8),[Table 1](https://arxiv.org/html/2605.15314#S1.T1.10.6.6.3.1.1),[§1](https://arxiv.org/html/2605.15314#S1.p7.8),[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[45\]S\. Zhao, Y\. Xie, and W\. Li\(2021\)On the convergence and improvement of stochastic normalized gradient descent\.Science China Information Sciences64\(3\),pp\. 132103\.Cited by:[§2](https://arxiv.org/html/2605.15314#S2.p2.5)\.
- \[46\]D\. Zhou, P\. Xu, and Q\. Gu\(2020\)Stochastic nested variance reduction for nonconvex optimization\.Journal of Machine Learning Research21\(103\),pp\. 1–63\.Cited by:[§1](https://arxiv.org/html/2605.15314#S1.p2.4)\.
## Appendix Table of Contents
## Appendix AA Short Primer on Smoothness Classes
We summarize the smoothness notions used throughout the paper\. The baseline class is the standard smoothness classℒ\\mathcal\{L\}, while the generalized smoothness classesℒH∗\\mathcal\{L\}\_\{\\rm H\}^\{\*\},ℒasym∗\\mathcal\{L\}\_\{\\rm asym\}^\{\*\}, andℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)allow the local smoothness constant to grow with the gradient norm\. The expected class𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)is the stochastic analogue used for variance\-reduced methods\. ReferChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]for a more robust discussion on it\.
#### Standard smoothnessℒ\\mathcal\{L\}\.
A differentiable functionf:ℝd→ℝf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}belongs toℒ\\mathcal\{L\}if there existsL0\>0L\_\{0\}\>0such that, for allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},
‖∇f\(y\)−∇f\(x\)‖≤L0‖y−x‖\.\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|\\leq L\_\{0\}\\\|y\-x\\\|\.\(4\)Equivalently, the gradient is globally Lipschitz continuous\. This is the classical smoothness assumption used in standard nonconvex optimization\.
#### Hessian\-based generalized smoothnessℒH∗\\mathcal\{L\}\_\{\\rm H\}^\{\*\}\.
For twice differentiable objectives, the Hessian\-based generalized smoothness classℒH∗\\mathcal\{L\}\_\{\\rm H\}^\{\*\}is defined by the existence of constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0such that, for allx∈ℝdx\\in\\mathbb\{R\}^\{d\},
‖∇2f\(x\)‖≤L0\+L1‖∇f\(x\)‖\.\\\|\\nabla^\{2\}f\(x\)\\\|\\leq L\_\{0\}\+L\_\{1\}\\\|\\nabla f\(x\)\\\|\.\(5\)This is often called\(L0,L1\)\(L\_\{0\},L\_\{1\}\)\-smoothness\. Unlike standard smoothness, the curvature is allowed to increase when the gradient norm is large\.
#### Asymmetric generalized smoothness:ℒasym∗\\mathcal\{L\}\_\{\\rm asym\}^\{\*\}\.
The asymmetric generalized smoothness classℒasym∗\\mathcal\{L\}\_\{\\rm asym\}^\{\*\}consists of differentiable functions satisfying, for someL0,L1\>0L\_\{0\},L\_\{1\}\>0and allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},
‖∇f\(y\)−∇f\(x\)‖≤\(L0\+L1‖∇f\(y\)‖\)‖y−x‖\.\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|\\leq\\left\(L\_\{0\}\+L\_\{1\}\\\|\\nabla f\(y\)\\\|\\right\)\\\|y\-x\\\|\.\(6\)
#### α\\alpha\-symmetric generalized smoothnessℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\.
FollowingChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], forα∈\(0,1\]\\alpha\\in\(0,1\], we say thatf∈ℒsym∗\(α\)f\\in\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)if there exist constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0such that, for allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},
‖∇f\(y\)−∇f\(x\)‖≤\(L0\+L1maxθ∈\[0,1\]‖∇f\(xθ\)‖α\)‖y−x‖,xθ:=θy\+\(1−θ\)x\.\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|\\leq\\left\(L\_\{0\}\+L\_\{1\}\\max\_\{\\theta\\in\[0,1\]\}\\\|\\nabla f\(x\_\{\\theta\}\)\\\|^\{\\alpha\}\\right\)\\\|y\-x\\\|,\\qquad x\_\{\\theta\}:=\\theta y\+\(1\-\\theta\)x\.\(7\)This condition is symmetric in the sense that the local smoothness is controlled along the entire line segment betweenxxandyy, rather than only at one endpoint \(and hence asymmetric in[6](https://arxiv.org/html/2605.15314#A1.E6)\)\. The parameterα\\alphadetermines how strongly the local smoothness depends on the gradient norm\. The caseα=1\\alpha=1corresponds to the usual symmetric\(L0,L1\)\(L\_\{0\},L\_\{1\}\)generalized smoothness regime, whileα∈\(0,1\)\\alpha\\in\(0,1\)gives a sublinear dependence on the gradient norm\. Standard smoothness is recovered as the special caseL1=0L\_\{1\}=0\.
#### Expectedα\\alpha\-symmetric generalized smoothness𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\.
For stochastic objectivesf\(x\)=𝔼ξ\[f\(x;ξ\)\]f\(x\)=\\mathbb\{E\}\_\{\\xi\}\[f\(x;\\xi\)\], the expected generalized smoothness class𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)requires that, for someL0,L1\>0L\_\{0\},L\_\{1\}\>0and allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},
𝔼ξ‖∇f\(y;ξ\)−∇f\(x;ξ\)‖2≤‖y−x‖2𝔼ξ\[\(L0\+L1maxθ∈\[0,1\]‖∇f\(xθ;ξ\)‖α\)2\],xθ:=θy\+\(1−θ\)x\.\\mathbb\{E\}\_\{\\xi\}\\big\\\|\\nabla f\(y;\\xi\)\-\\nabla f\(x;\\xi\)\\big\\\|^\{2\}\\leq\\\|y\-x\\\|^\{2\}\\mathbb\{E\}\_\{\\xi\}\\Big\[\\big\(L\_\{0\}\+L\_\{1\}\\max\_\{\\theta\\in\[0,1\]\}\\\|\\nabla f\(x\_\{\\theta\};\\xi\)\\\|^\{\\alpha\}\\big\)^\{2\}\\Big\],\\ x\_\{\\theta\}:=\\theta y\+\(1\-\\theta\)x\.\(8\)This is a mean\-square, sample\-level analogue ofℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\. It is stronger than requiring only the population objectiveffto satisfyℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\), and it is the relevant condition for analyzing recursive variance\-reduced estimators such as𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}\.
#### Mean\-square smoothness𝖬𝖲𝖲\\mathsf\{MSS\}\.
Mean\-square smoothness is the standard stochastic analogue ofLL\-smoothness\. It requires that there existsL\>0L\>0such that, for allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},
𝔼ξ‖∇f\(y;ξ\)−∇f\(x;ξ\)‖2≤L2‖y−x‖2\.\\mathbb\{E\}\_\{\\xi\}\\left\\\|\\nabla f\(y;\\xi\)\-\\nabla f\(x;\\xi\)\\right\\\|^\{2\}\\leq L^\{2\}\\\|y\-x\\\|^\{2\}\.\(9\)Thus,𝖬𝖲𝖲\\mathsf\{MSS\}controls stochastic gradient differences without any dependence on gradient magnitudes\. In contrast,𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)permits the stochastic smoothness factor to grow with the sample\-gradient norm\.
#### Relations among the classes\.
The classes satisfy the following inclusions, up to constants:
ℒ⊂ℒsym∗\(α\),ℒasym∗⊂ℒsym∗\(1\),ℒH∗⊂ℒsym∗\(1\)\.\\mathcal\{L\}\\subset\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\),\\qquad\\mathcal\{L\}\_\{\\rm asym\}^\{\*\}\\subset\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\),\\qquad\\mathcal\{L\}\_\{\\rm H\}^\{\*\}\\subset\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\)\.\(10\)Moreover, on twice differentiable functions,ℒH∗\\mathcal\{L\}\_\{\\rm H\}^\{\*\}andℒsym∗\(1\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\)are equivalent up to constants\. Hence,ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)provides a unified deterministic framework that contains standard smoothness, asymmetric generalized smoothness, and Hessian\-based\(L0,L1\)\(L\_\{0\},L\_\{1\}\)\-smoothness\. The expected class𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)plays the analogous role for stochastic gradient differences\.
#### Role in this paper\.
In this paper,ℒ\\mathcal\{L\}andℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)are used for the analysis of normalized stochastic gradient descent with momentum,𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}\. Under𝖡𝖦\\mathsf\{BG\}\-0 noise, the normalized update controls the trajectory length, and the generalized smoothness parameters affect constants but do not change the𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)complexity exponent\.
For normalized STORM,𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}, the relevant conditions are𝖬𝖲𝖲\\mathsf\{MSS\}and𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\. Under𝖬𝖲𝖲\\mathsf\{MSS\}, the transition noise in the recursive estimator is controlled only by the step length, leading to the optimal𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)rate under𝖡𝖦\\mathsf\{BG\}\-0\. Under𝔼ℒsym∗\(α\)\\mathbb\{E\}\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\), however, the transition noise also depends on sample\-gradient moments\. These moments interact with the distance\-dependent𝖡𝖦\\mathsf\{BG\}\-0 variance, producing theα\\alpha\-dependent rates established in the main text\.
#### Related Work: Generalized smoothness\.
Standard theory often assumes a constant smoothness parameter, but recent research focuses on objectives where local smoothness grows alongside the gradient norm\[[6](https://arxiv.org/html/2605.15314#bib.bib35),[14](https://arxiv.org/html/2605.15314#bib.bib50)\]\. A prime example isℒH∗\\mathcal\{L\}^\{\*\}\_\{\\rm H\}or\(L0,L1\)\(L\_\{0\},L\_\{1\}\)\-smoothness, where the Hessian is bounded by‖∇2f\(x\)‖≤L0\+L1‖∇f\(x\)‖\\\|\\nabla^\{2\}f\(x\)\\\|\\leq L\_\{0\}\+L\_\{1\}\\\|\\nabla f\(x\)\\\|\[[44](https://arxiv.org/html/2605.15314#bib.bib141)\]\. This concept was further expanded byChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]into theα\\alpha\-symmetric generalized smoothness framework \(ℒsym∗\(α\)/𝔼ℒ∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)/\\mathbb\{E\}\\mathcal\{L\}^\{\*\}\(\\alpha\)\), which covers standard smoothness, asymmetric generalized smoothness\[[22](https://arxiv.org/html/2605.15314#bib.bib140),[29](https://arxiv.org/html/2605.15314#bib.bib143),[38](https://arxiv.org/html/2605.15314#bib.bib145)\], and Hessian\-based generalized smoothness, and also covers high\-order polynomial and exponential\-type objectives\. While prior work shows that generalized smooth problems can match the efficiency of standard smooth optimization, achieving deterministic rates of𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)with normalized gradients or𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)with SPIDER style variance reduction, these results rely on relatively tame variance assumptions\. Our work investigates whether these efficiencies still hold under𝖡𝖦\\mathsf\{BG\}\-0 noise, where the variance is trajectory dependent\.
## Appendix BProofs for Normalized Momentum under𝖡𝖦\\mathsf\{BG\}\-0 Noise
This appendix proves the normalized\-momentum guarantees stated in Theorem[1](https://arxiv.org/html/2605.15314#Thmtheorem1)\. Throughout this section, let
T:=K\+1,Δ:=f\(x0\)−finf\.T:=K\+1,\\qquad\\Delta:=f\(x^\{0\}\)\-f^\{\\inf\}\.We define
gk:=𝔼‖∇f\(xk\)‖,bk:=𝔼‖vk−∇f\(xk\)‖,g\_\{k\}:=\\mathbb\{E\}\\\|\\nabla f\(x^\{k\}\)\\\|,\\qquad b\_\{k\}:=\\mathbb\{E\}\\\|v^\{k\}\-\\nabla f\(x^\{k\}\)\\\|,and
Δk:=𝔼\[f\(xk\)−finf\],Sb\(K\):=∑k=0Kbk,SΔ\(K\):=∑k=0KΔk\.\\Delta^\{k\}:=\\mathbb\{E\}\[f\(x^\{k\}\)\-f^\{\\inf\}\],\\qquad S\_\{b\}\(K\):=\\sum\_\{k=0\}^\{K\}b\_\{k\},\\qquad S\_\{\\Delta\}\(K\):=\\sum\_\{k=0\}^\{K\}\\Delta^\{k\}\.The outputx^\\widehat\{x\}is sampled uniformly from\{x0,x1,…,xK\}\\\{x^\{0\},x^\{1\},\\ldots,x^\{K\}\\\}, and hence
𝔼‖∇f\(x^\)‖=1T∑k=0Kgk\.\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\.\(11\)We use the convention thatv/‖v‖=0v/\\\|v\\\|=0wheneverv=0v=0\.
### B\.1Auxiliary inequalities
###### Lemma 1\(Trajectory control by normalization\)\.
For the𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}iterates in \([2](https://arxiv.org/html/2605.15314#S4.E2)\),
‖xk\+1−xk‖≤γ,‖xk−x0‖≤kγ\.\\\|x^\{k\+1\}\-x^\{k\}\\\|\\leq\\gamma,\\qquad\\\|x^\{k\}\-x^\{0\}\\\|\\leq k\\gamma\.Consequently, under Assumption[2](https://arxiv.org/html/2605.15314#Thmassumption2),
𝔼‖∇f\(xk;ξk\)−∇f\(xk\)‖2≤B2k2γ2\+G2\.\\mathbb\{E\}\\\|\\nabla f\(x^\{k\};\\xi^\{k\}\)\-\\nabla f\(x^\{k\}\)\\\|^\{2\}\\leq B^\{2\}k^\{2\}\\gamma^\{2\}\+G^\{2\}\.\(12\)In particular, sincev0=∇f\(x0;ξ0\)v^\{0\}=\\nabla f\(x^\{0\};\\xi^\{0\}\),
###### Proof\.
The normalized direction has norm at most one\. Hence
‖xk\+1−xk‖≤γ\.\\\|x^\{k\+1\}\-x^\{k\}\\\|\\leq\\gamma\.Summing the increments gives
‖xk−x0‖≤∑t=0k−1‖xt\+1−xt‖≤kγ\.\\\|x^\{k\}\-x^\{0\}\\\|\\leq\\sum\_\{t=0\}^\{k\-1\}\\\|x^\{t\+1\}\-x^\{t\}\\\|\\leq k\\gamma\.Substitution into Assumption[2](https://arxiv.org/html/2605.15314#Thmassumption2)gives \([12](https://arxiv.org/html/2605.15314#A2.E12)\)\. Finally,
b0=𝔼‖v0−∇f\(x0\)‖≤𝔼‖v0−∇f\(x0\)‖2≤G,b\_\{0\}=\\mathbb\{E\}\\\|v^\{0\}\-\\nabla f\(x^\{0\}\)\\\|\\leq\\sqrt\{\\mathbb\{E\}\\\|v^\{0\}\-\\nabla f\(x^\{0\}\)\\\|^\{2\}\}\\leq G,becausex0−x0=0x^\{0\}\-x^\{0\}=0\. ∎
###### Lemma 2\(Direction\-error inequality\)\.
For anya,e∈ℝda,e\\in\\mathbb\{R\}^\{d\}, define
d=\{a\+e‖a\+e‖,a\+e≠0,0,a\+e=0\.d=\\begin\{cases\}\\dfrac\{a\+e\}\{\\\|a\+e\\\|\},&a\+e\\neq 0,\\\\\[4\.0pt\] 0,&a\+e=0\.\\end\{cases\}Then
⟨a,d⟩≥‖a‖−2‖e‖\.\\langle a,d\\rangle\\geq\\\|a\\\|\-2\\\|e\\\|\.\(14\)
###### Proof\.
Ifa\+e=0a\+e=0, thend=0d=0ande=−ae=\-a\. Hence
⟨a,d⟩=0≥‖a‖−2‖a‖\.\\langle a,d\\rangle=0\\geq\\\|a\\\|\-2\\\|a\\\|\.Now supposea\+e≠0a\+e\\neq 0\. Then
⟨a,d⟩=‖a‖2\+⟨a,e⟩‖a\+e‖≥‖a‖\(‖a‖−‖e‖\)‖a\+e‖\.\\langle a,d\\rangle=\\frac\{\\\|a\\\|^\{2\}\+\\langle a,e\\rangle\}\{\\\|a\+e\\\|\}\\geq\\frac\{\\\|a\\\|\(\\\|a\\\|\-\\\|e\\\|\)\}\{\\\|a\+e\\\|\}\.If‖a‖≥‖e‖\\\|a\\\|\\geq\\\|e\\\|, then using‖a\+e‖≤‖a‖\+‖e‖\\\|a\+e\\\|\\leq\\\|a\\\|\+\\\|e\\\|, we obtain
⟨a,d⟩≥‖a‖\(‖a‖−‖e‖\)‖a‖\+‖e‖=‖a‖−2‖a‖‖e‖‖a‖\+‖e‖≥‖a‖−2‖e‖\.\\langle a,d\\rangle\\geq\\frac\{\\\|a\\\|\(\\\|a\\\|\-\\\|e\\\|\)\}\{\\\|a\\\|\+\\\|e\\\|\}=\\\|a\\\|\-\\frac\{2\\\|a\\\|\\\|e\\\|\}\{\\\|a\\\|\+\\\|e\\\|\}\\geq\\\|a\\\|\-2\\\|e\\\|\.If‖a‖<‖e‖\\\|a\\\|<\\\|e\\\|, then by Cauchy–Schwarz,
⟨a,d⟩≥−‖a‖≥‖a‖−2‖e‖\.\\langle a,d\\rangle\\geq\-\\\|a\\\|\\geq\\\|a\\\|\-2\\\|e\\\|\.The claim follows\. ∎
###### Lemma 3\(Young inequality for fractional powers\)\.
For everyx≥0x\\geq 0,α∈\(0,1\)\\alpha\\in\(0,1\), andρ\>0\\rho\>0,
xα≤αρx\+\(1−α\)ρ−α/\(1−α\)\.x^\{\\alpha\}\\leq\\alpha\\rho x\+\(1\-\\alpha\)\\rho^\{\-\\alpha/\(1\-\\alpha\)\}\.\(15\)
###### Proof\.
Apply Young’s inequality with conjugate exponentsp=1/αp=1/\\alphaandq=1/\(1−α\)q=1/\(1\-\\alpha\)to
u=\(ρx\)α,v=ρ−α\.u=\(\\rho x\)^\{\\alpha\},\\qquad v=\\rho^\{\-\\alpha\}\.∎
### B\.2Standard smoothness andℒsym∗\(1\)\\mathcal\{L\}^\{\*\}\_\{\\rm sym\}\(1\)
We first prove the standard smoothness case and the endpointℒsym∗\(1\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\)case\. For the endpoint generalized smooth case, we use the following standard consequences of symmetric generalized smoothness \(seeChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], Proposition 1\): for allx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\},
‖∇f\(y\)−∇f\(x\)‖≤\(L0\+L1‖∇f\(x\)‖\)eL1‖y−x‖‖y−x‖,\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|\\leq\\bigl\(L\_\{0\}\+L\_\{1\}\\\|\\nabla f\(x\)\\\|\\bigr\)e^\{L\_\{1\}\\\|y\-x\\\|\}\\\|y\-x\\\|,\(16\)and
f\(y\)≤f\(x\)\+⟨∇f\(x\),y−x⟩\+12\(L0\+L1‖∇f\(x\)‖\)eL1‖y−x‖‖y−x‖2\.f\(y\)\\leq f\(x\)\+\\langle\\nabla f\(x\),y\-x\\rangle\+\\frac\{1\}\{2\}\\bigl\(L\_\{0\}\+L\_\{1\}\\\|\\nabla f\(x\)\\\|\\bigr\)e^\{L\_\{1\}\\\|y\-x\\\|\}\\\|y\-x\\\|^\{2\}\.\(17\)WhenL1=0L\_\{1\}=0, these reduce to the standardL0L\_\{0\}\-smooth inequalities\. We also use the gradient\-growth inequality \(seeKhiriratet al\.\[[26](https://arxiv.org/html/2605.15314#bib.bib154)\], Lemma 2\)
‖∇f\(x\)‖≤8L1\(f\(x\)−finf\)\+L0L1,L1\>0,\\\|\\nabla f\(x\)\\\|\\leq 8L\_\{1\}\(f\(x\)\-f^\{\\inf\}\)\+\\frac\{L\_\{0\}\}\{L\_\{1\}\},\\qquad L\_\{1\}\>0,\(18\)which follows from lower boundedness and the endpoint generalized smoothness condition\.
###### Theorem 3\(Single\-sample master bound forL0L\_\{0\}\-smoothness andα=1\\alpha=1\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[2](https://arxiv.org/html/2605.15314#Thmassumption2)hold\. Assume that eitherffisL0L\_\{0\}\-smooth, or that \([16](https://arxiv.org/html/2605.15314#A2.E16)\)–\([17](https://arxiv.org/html/2605.15314#A2.E17)\) hold\. In the standard smooth case, setL1=0L\_\{1\}=0\. If
γL1≤12,32L12γ2Tη≤12,\\gamma L\_\{1\}\\leq\\frac\{1\}\{2\},\\qquad\\frac\{32L\_\{1\}^\{2\}\\gamma^\{2\}T\}\{\\eta\}\\leq\\frac\{1\}\{2\},\(19\)with both conditions interpreted as vacuous whenL1=0L\_\{1\}=0, then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤2ΔγT\+4b0ηT\+16L0γη\+2L0γ\+4η\(G\+BγK2\)\.\\displaystyle\\leq\\frac\{2\\Delta\}\{\\gamma T\}\+\\frac\{4b\_\{0\}\}\{\\eta T\}\+\\frac\{16L\_\{0\}\\gamma\}\{\\eta\}\+2L\_\{0\}\\gamma\+4\\sqrt\{\\eta\}\\Big\(G\+\\frac\{B\\gamma K\}\{2\}\\Big\)\.\+64L12γη\[4Δ\+4γb0η\+16L0γ2Tη\+4γηTG\+2Bγ2ηKT\+2L0γ2T\]\\displaystyle\\quad\+\\frac\{64L\_\{1\}^\{2\}\\gamma\}\{\\eta\}\\Bigg\[4\\Delta\+\\frac\{4\\gamma b\_\{0\}\}\{\\eta\}\+\\frac\{16L\_\{0\}\\gamma^\{2\}T\}\{\\eta\}\+4\\gamma\\sqrt\{\\eta\}\\,TG\+2B\\gamma^\{2\}\\sqrt\{\\eta\}\\,KT\+2L\_\{0\}\\gamma^\{2\}T\\Bigg\]\(20\)WhenL1=0L\_\{1\}=0, the entireL12L\_\{1\}^\{2\}\-term vanishes\.
###### Proof\.
Let
Ek:=vk−∇f\(xk\),Dk:=∇f\(xk\)−∇f\(xk\+1\),E\_\{k\}:=v^\{k\}\-\\nabla f\(x^\{k\}\),\\qquad D\_\{k\}:=\\nabla f\(x^\{k\}\)\-\\nabla f\(x^\{k\+1\}\),and
ζk\+1:=∇f\(xk\+1;ξk\+1\)−∇f\(xk\+1\)\.\\zeta\_\{k\+1\}:=\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\+1\}\)\.
#### Step 1: descent\.
Using \([17](https://arxiv.org/html/2605.15314#A2.E17)\) withx=xkx=x^\{k\}andy=xk\+1=xk−γdky=x^\{k\+1\}=x^\{k\}\-\\gamma d^\{k\}, we get
f\(xk\+1\)\\displaystyle f\(x^\{k\+1\}\)≤f\(xk\)−γ⟨∇f\(xk\),dk⟩\\displaystyle\\leq f\(x^\{k\}\)\-\\gamma\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\rangle\+γ22eL1γ\(L0\+L1‖∇f\(xk\)‖\)\.\\displaystyle\\quad\+\\frac\{\\gamma^\{2\}\}\{2\}e^\{L\_\{1\}\\gamma\}\\bigl\(L\_\{0\}\+L\_\{1\}\\\|\\nabla f\(x^\{k\}\)\\\|\\bigr\)\.SinceγL1≤1/2\\gamma L\_\{1\}\\leq 1/2, we haveeL1γ≤e1/2<2e^\{L\_\{1\}\\gamma\}\\leq e^\{1/2\}<2\. By Lemma[2](https://arxiv.org/html/2605.15314#Thmlemma2),
⟨∇f\(xk\),dk⟩≥‖∇f\(xk\)‖−2‖Ek‖\.\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\rangle\\geq\\\|\\nabla f\(x^\{k\}\)\\\|\-2\\\|E\_\{k\}\\\|\.Therefore,
f\(xk\+1\)\\displaystyle f\(x^\{k\+1\}\)≤f\(xk\)−γ‖∇f\(xk\)‖\+2γ‖Ek‖\+L0γ2\+L1γ2‖∇f\(xk\)‖\.\\displaystyle\\leq f\(x^\{k\}\)\-\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|\+2\\gamma\\\|E\_\{k\}\\\|\+L\_\{0\}\\gamma^\{2\}\+L\_\{1\}\\gamma^\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|\.UsingγL1≤1/2\\gamma L\_\{1\}\\leq 1/2, we obtain
Δk\+1≤Δk−γ2gk\+2γbk\+L0γ2\.\\Delta^\{k\+1\}\\leq\\Delta^\{k\}\-\\frac\{\\gamma\}\{2\}g\_\{k\}\+2\\gamma b\_\{k\}\+L\_\{0\}\\gamma^\{2\}\.\(21\)
#### Step 2: estimator recursion\.
The momentum update gives
Ek\+1=\(1−η\)Ek\+\(1−η\)Dk\+ηζk\+1\.E\_\{k\+1\}=\(1\-\\eta\)E\_\{k\}\+\(1\-\\eta\)D\_\{k\}\+\\eta\\zeta\_\{k\+1\}\.Unrolling,
Ek=\(1−η\)kE0\+∑t=0k−1\(1−η\)k−tDt\+η∑t=0k−1\(1−η\)k−1−tζt\+1\.E\_\{k\}=\(1\-\\eta\)^\{k\}E\_\{0\}\+\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-t\}D\_\{t\}\+\\eta\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}\\zeta\_\{t\+1\}\.For the drift term, \([16](https://arxiv.org/html/2605.15314#A2.E16)\),eL1γ≤2e^\{L\_\{1\}\\gamma\}\\leq 2, and \([18](https://arxiv.org/html/2605.15314#A2.E18)\) imply, forL1\>0L\_\{1\}\>0,
𝔼‖Dt‖\\displaystyle\\mathbb\{E\}\\\|D\_\{t\}\\\|≤2γ\(L0\+L1gt\)\\displaystyle\\leq 2\\gamma\\bigl\(L\_\{0\}\+L\_\{1\}g\_\{t\}\\bigr\)≤2γ\(L0\+L1\(8L1Δt\+L0L1\)\)\\displaystyle\\leq 2\\gamma\\left\(L\_\{0\}\+L\_\{1\}\\left\(8L\_\{1\}\\Delta^\{t\}\+\\frac\{L\_\{0\}\}\{L\_\{1\}\}\\right\)\\right\)=4L0γ\+16L12γΔt\.\\displaystyle=4L\_\{0\}\\gamma\+6L\_\{1\}^\{2\}\\gamma\\Delta^\{t\}\.ForL1=0L\_\{1\}=0, the second term is absent, and the same displayed bound remains valid with theL12L\_\{1\}^\{2\}\-term equal to zero\.
For the martingale term, conditional unbiasedness and the orthogonality of martingale differences give
𝔼‖η∑t=0k−1\(1−η\)k−1−tζt\+1‖\\displaystyle\\mathbb\{E\}\\left\\\|\\eta\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}\\zeta\_\{t\+1\}\\right\\\|≤\(η2∑t=0k−1\(1−η\)2\(k−1−t\)𝔼‖ζt\+1‖2\)1/2\.\\displaystyle\\quad\\leq\\left\(\\eta^\{2\}\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{2\(k\-1\-t\)\}\\mathbb\{E\}\\\|\\zeta\_\{t\+1\}\\\|^\{2\}\\right\)^\{1/2\}\.By Lemma[1](https://arxiv.org/html/2605.15314#Thmlemma1), fort≤k−1t\\leq k\-1,
𝔼‖ζt\+1‖2≤B2\(t\+1\)2γ2\+G2≤B2k2γ2\+G2\.\\mathbb\{E\}\\\|\\zeta\_\{t\+1\}\\\|^\{2\}\\leq B^\{2\}\(t\+1\)^\{2\}\\gamma^\{2\}\+G^\{2\}\\leq B^\{2\}k^\{2\}\\gamma^\{2\}\+G^\{2\}\.Also,
∑t=0k−1\(1−η\)2\(k−1−t\)≤1η\.\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{2\(k\-1\-t\)\}\\leq\\frac\{1\}\{\\eta\}\.Hence
𝔼‖η∑t=0k−1\(1−η\)k−1−tζt\+1‖≤η\(G\+Bkγ\)\.\\mathbb\{E\}\\left\\\|\\eta\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}\\zeta\_\{t\+1\}\\right\\\|\\leq\\sqrt\{\\eta\}\(G\+Bk\\gamma\)\.Combining the drift and martingale estimates yields
bk≤\(1−η\)kb0\+4L0γη\+16L12γ∑t=0k−1\(1−η\)k−tΔt\+η\(G\+Bkγ\)\.b\_\{k\}\\leq\(1\-\\eta\)^\{k\}b\_\{0\}\+\\frac\{4L\_\{0\}\\gamma\}\{\\eta\}\+16L\_\{1\}^\{2\}\\gamma\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-t\}\\Delta^\{t\}\+\\sqrt\{\\eta\}\(G\+Bk\\gamma\)\.\(22\)
#### Step 3: summing the estimator errors\.
Summing \([22](https://arxiv.org/html/2605.15314#A2.E22)\) overk=0,…,Kk=0,\\ldots,K, and using
∑k=0K\(1−η\)k≤1η,\\sum\_\{k=0\}^\{K\}\(1\-\\eta\)^\{k\}\\leq\\frac\{1\}\{\\eta\},and
∑k=0K∑t=0k−1\(1−η\)k−tΔt≤1ηSΔ\(K\),\\sum\_\{k=0\}^\{K\}\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-t\}\\Delta^\{t\}\\leq\\frac\{1\}\{\\eta\}S\_\{\\Delta\}\(K\),we get
Sb\(K\)≤b0η\+4L0γTη\+16L12γηSΔ\(K\)\+η\(TG\+BγKT2\)\.S\_\{b\}\(K\)\\leq\\frac\{b\_\{0\}\}\{\\eta\}\+\\frac\{4L\_\{0\}\\gamma T\}\{\\eta\}\+\\frac\{16L\_\{1\}^\{2\}\\gamma\}\{\\eta\}S\_\{\\Delta\}\(K\)\+\\sqrt\{\\eta\}\\left\(TG\+\\frac\{B\\gamma KT\}\{2\}\\right\)\.\(23\)
#### Step 4: controlling the accumulated suboptimality\.
Dropping the negative gradient term in \([21](https://arxiv.org/html/2605.15314#A2.E21)\), iterating, and summing gives
SΔ\(K\)≤\(K\+2\)Δ\+2γTSb\(K\)\+L0γ2T\(K\+2\)2\.S\_\{\\Delta\}\(K\)\\leq\(K\+2\)\\Delta\+2\\gamma TS\_\{b\}\(K\)\+\\frac\{L\_\{0\}\\gamma^\{2\}T\(K\+2\)\}\{2\}\.Substituting \([23](https://arxiv.org/html/2605.15314#A2.E23)\) gives
SΔ\(K\)\\displaystyle S\_\{\\Delta\}\(K\)≤\(K\+2\)Δ\+2γTb0η\+8L0γ2T2η\+32L12γ2TηSΔ\(K\)\\displaystyle\\leq\(K\+2\)\\Delta\+\\frac\{2\\gamma Tb\_\{0\}\}\{\\eta\}\+\\frac\{8L\_\{0\}\\gamma^\{2\}T^\{2\}\}\{\\eta\}\+\\frac\{32L\_\{1\}^\{2\}\\gamma^\{2\}T\}\{\\eta\}S\_\{\\Delta\}\(K\)\+2γηT2G\+Bγ2ηKT2\+L0γ2T\(K\+2\)2\.\\displaystyle\\quad\+2\\gamma\\sqrt\{\\eta\}T^\{2\}G\+B\\gamma^\{2\}\\sqrt\{\\eta\}KT^\{2\}\+\\frac\{L\_\{0\}\\gamma^\{2\}T\(K\+2\)\}\{2\}\.By \([19](https://arxiv.org/html/2605.15314#A2.E19)\), the coefficient ofSΔ\(K\)S\_\{\\Delta\}\(K\)on the right is at most1/21/2\. Moving it to the left and usingK\+2≤2TK\+2\\leq 2T, we obtain
SΔ\(K\)T≤4Δ\+4γb0η\+16L0γ2Tη\+4γηTG\+2Bγ2ηKT\+2L0γ2T\.\\frac\{S\_\{\\Delta\}\(K\)\}\{T\}\\leq 4\\Delta\+\\frac\{4\\gamma b\_\{0\}\}\{\\eta\}\+\\frac\{16L\_\{0\}\\gamma^\{2\}T\}\{\\eta\}\+4\\gamma\\sqrt\{\\eta\}\\,TG\+2B\\gamma^\{2\}\\sqrt\{\\eta\}\\,KT\+2L\_\{0\}\\gamma^\{2\}T\.\(24\)
#### Step 5: final stationarity bound\.
Summing \([21](https://arxiv.org/html/2605.15314#A2.E21)\) gives
γ2∑k=0Kgk≤Δ\+2γSb\(K\)\+L0γ2T\.\\frac\{\\gamma\}\{2\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\\leq\\Delta\+2\\gamma S\_\{b\}\(K\)\+L\_\{0\}\\gamma^\{2\}T\.Dividing byγT/2\\gamma T/2, using \([11](https://arxiv.org/html/2605.15314#A2.E11)\), and substituting \([23](https://arxiv.org/html/2605.15314#A2.E23)\) and \([24](https://arxiv.org/html/2605.15314#A2.E24)\), we obtain \([20](https://arxiv.org/html/2605.15314#A2.E20)\)\. ∎
###### Corollary 3\.1\(Standard smoothness under𝖡𝖦\\mathsf\{BG\}\-0\)\.
Assume Assumption[3](https://arxiv.org/html/2605.15314#Thmassumption3)holds\. Choose
η=T−2/3,γ=γ0T−5/6\.\\eta=T^\{\-2/3\},\\qquad\\gamma=\\gamma\_\{0\}T^\{\-5/6\}\.Then
𝔼‖∇f\(x^\)‖≤\(2Δγ0\+16L0γ0\+2Bγ0\)T−1/6\+8GT−1/3\+2L0γ0T−5/6\.\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+16L\_\{0\}\\gamma\_\{0\}\+2B\\gamma\_\{0\}\\right\)T^\{\-1/6\}\+8GT^\{\-1/3\}\+2L\_\{0\}\\gamma\_\{0\}T^\{\-5/6\}\.\(25\)Consequently,
𝔼‖∇f\(x^\)‖=𝒪\(T−1/6\),\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/6\}\),and the SFO complexity is𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)\.
###### Proof\.
SetL1=0L\_\{1\}=0in Theorem[3](https://arxiv.org/html/2605.15314#Thmtheorem3)\. Then theL12L\_\{1\}^\{2\}\-term vanishes and the conditions \([19](https://arxiv.org/html/2605.15314#A2.E19)\) are vacuous\. Usingb0≤Gb\_\{0\}\\leq GandK≤TK\\leq T,
2ΔγT=2Δγ0T−1/6,4b0ηT≤4GT−1/3,\\frac\{2\\Delta\}\{\\gamma T\}=\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}T^\{\-1/6\},\\qquad\\frac\{4b\_\{0\}\}\{\\eta T\}\\leq 4GT^\{\-1/3\},16L0γη=16L0γ0T−1/6,2L0γ=2L0γ0T−5/6,\\frac\{16L\_\{0\}\\gamma\}\{\\eta\}=16L\_\{0\}\\gamma\_\{0\}T^\{\-1/6\},\\qquad 2L\_\{0\}\\gamma=2L\_\{0\}\\gamma\_\{0\}T^\{\-5/6\},and
4η\(G\+BγK2\)≤4GT−1/3\+2Bγ0T−1/6\.4\\sqrt\{\\eta\}\\left\(G\+\\frac\{B\\gamma K\}\{2\}\\right\)\\leq 4GT^\{\-1/3\}\+2B\\gamma\_\{0\}T^\{\-1/6\}\.Combining these estimates proves \([25](https://arxiv.org/html/2605.15314#A2.E25)\)\. ∎
###### Corollary 3\.2\(1\-generalized smoothness under𝖡𝖦\\mathsf\{BG\}\-0\)\.
Assume Assumption[5](https://arxiv.org/html/2605.15314#Thmassumption5)withα=1\\alpha=1holds\. Choose
η=T−2/3,γ=γ0T−5/6,0<γ0≤18L1\.\\eta=T^\{\-2/3\},\\qquad\\gamma=\\gamma\_\{0\}T^\{\-5/6\},\\qquad 0<\\gamma\_\{0\}\\leq\\frac\{1\}\{8L\_\{1\}\}\.Then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(2Δγ0\+16L0γ0\+2Bγ0\)T−1/6\+8GT−1/3\+2L0γ0T−5/6\\displaystyle\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+16L\_\{0\}\\gamma\_\{0\}\+2B\\gamma\_\{0\}\\right\)T^\{\-1/6\}\+8GT^\{\-1/3\}\+2L\_\{0\}\\gamma\_\{0\}T^\{\-5/6\}\+64L12γ0T−1/6\[4Δ\+\(16L0\+2B\)γ02\+8γ0GT−1/6\+2L0γ02T−2/3\]\.\\displaystyle\\quad\+64L\_\{1\}^\{2\}\\gamma\_\{0\}T^\{\-1/6\}\\left\[4\\Delta\+\(16L\_\{0\}\+2B\)\\gamma\_\{0\}^\{2\}\+8\\gamma\_\{0\}GT^\{\-1/6\}\+2L\_\{0\}\\gamma\_\{0\}^\{2\}T^\{\-2/3\}\\right\]\.\(26\)Consequently,
𝔼‖∇f\(x^\)‖=𝒪\(T−1/6\),\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/6\}\),and the SFO complexity is𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)\.
###### Proof\.
The first condition in \([19](https://arxiv.org/html/2605.15314#A2.E19)\) follows from
γL1≤γ0L1≤18\.\\gamma L\_\{1\}\\leq\\gamma\_\{0\}L\_\{1\}\\leq\\frac\{1\}\{8\}\.The second follows because
32L12γ2Tη=32L12γ02≤12\.\\frac\{32L\_\{1\}^\{2\}\\gamma^\{2\}T\}\{\\eta\}=32L\_\{1\}^\{2\}\\gamma\_\{0\}^\{2\}\\leq\\frac\{1\}\{2\}\.Substitute the schedule into \([20](https://arxiv.org/html/2605.15314#A2.E20)\), useb0≤Gb\_\{0\}\\leq GandK≤TK\\leq T, and simplify\. The non\-coupling terms give the first line of \([26](https://arxiv.org/html/2605.15314#A2.E26)\)\. The bracketed term in \([20](https://arxiv.org/html/2605.15314#A2.E20)\) is bounded by
4Δ\+\(16L0\+2B\)γ02\+8γ0GT−1/6\+2L0γ02T−2/3,4\\Delta\+\(16L\_\{0\}\+2B\)\\gamma\_\{0\}^\{2\}\+8\\gamma\_\{0\}GT^\{\-1/6\}\+2L\_\{0\}\\gamma\_\{0\}^\{2\}T^\{\-2/3\},while its prefactor becomes
64L12γ0T−1/6\.64L\_\{1\}^\{2\}\\gamma\_\{0\}T^\{\-1/6\}\.This proves \([26](https://arxiv.org/html/2605.15314#A2.E26)\)\. ∎
### B\.3The caseℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)withα∈\(0,1\)\\alpha\\in\(0,1\)
We now prove the sublinear generalized smoothness case\. The proof gives a clean explicit bound that implies the same𝒪\(T−1/6\)\\mathcal\{O\}\(T^\{\-1/6\}\)rate as in Theorem[1](https://arxiv.org/html/2605.15314#Thmtheorem1)\.
Under Assumption[5](https://arxiv.org/html/2605.15314#Thmassumption5)withα∈\(0,1\)\\alpha\\in\(0,1\), we use the following consequences of the symmetric generalized smoothness condition \(seeChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], Proposition 1\)\. For allw,w′∈ℝdw,w^\{\\prime\}\\in\\mathbb\{R\}^\{d\},
‖∇f\(w′\)−∇f\(w\)‖≤‖w′−w‖\(K0\+K1‖∇f\(w\)‖α\+K2‖w′−w‖α/\(1−α\)\),\\\|\\nabla f\(w^\{\\prime\}\)\-\\nabla f\(w\)\\\|\\leq\\\|w^\{\\prime\}\-w\\\|\\left\(K\_\{0\}\+K\_\{1\}\\\|\\nabla f\(w\)\\\|^\{\\alpha\}\+K\_\{2\}\\\|w^\{\\prime\}\-w\\\|^\{\\alpha/\(1\-\\alpha\)\}\\right\),\(27\)and
f\(w′\)\\displaystyle f\(w^\{\\prime\}\)≤f\(w\)\+⟨∇f\(w\),w′−w⟩\\displaystyle\\leq f\(w\)\+\\langle\\nabla f\(w\),w^\{\\prime\}\-w\\rangle\(28\)\+12‖w′−w‖2\(K0\+K1‖∇f\(w\)‖α\+2K2‖w′−w‖α/\(1−α\)\),\\displaystyle\\quad\+\\frac\{1\}\{2\}\\\|w^\{\\prime\}\-w\\\|^\{2\}\\left\(K\_\{0\}\+K\_\{1\}\\\|\\nabla f\(w\)\\\|^\{\\alpha\}\+2K\_\{2\}\\\|w^\{\\prime\}\-w\\\|^\{\\alpha/\(1\-\\alpha\)\}\\right\),where
K0:=L0\(2α2/\(1−α\)\+1\),K1:=L12α2/\(1−α\)3α,K\_\{0\}:=L\_\{0\}\\left\(2^\{\\alpha^\{2\}/\(1\-\\alpha\)\}\+1\\right\),\\qquad K\_\{1\}:=L\_\{1\}\\,2^\{\\alpha^\{2\}/\(1\-\\alpha\)\}3^\{\\alpha\},and
K2:=L11/\(1−α\)2α2/\(1−α\)3α\(1−α\)α/\(1−α\)\.K\_\{2\}:=L\_\{1\}^\{1/\(1\-\\alpha\)\}2^\{\\alpha^\{2\}/\(1\-\\alpha\)\}3^\{\\alpha\}\(1\-\\alpha\)^\{\\alpha/\(1\-\\alpha\)\}\.Define
Rα:=\(1−α\)K12\(2αK1\)α/\(1−α\),Cα:=K2\+Rα,R\_\{\\alpha\}:=\\frac\{\(1\-\\alpha\)K\_\{1\}\}\{2\}\(2\\alpha K\_\{1\}\)^\{\\alpha/\(1\-\\alpha\)\},\\qquad C\_\{\\alpha\}:=K\_\{2\}\+R\_\{\\alpha\},and
R~α:=\(1−α\)K1\(8αK1\)α/\(1−α\),C~α:=K0\+K2\+R~α\.\\widetilde\{R\}\_\{\\alpha\}:=\(1\-\\alpha\)K\_\{1\}\(8\\alpha K\_\{1\}\)^\{\\alpha/\(1\-\\alpha\)\},\\qquad\\widetilde\{C\}\_\{\\alpha\}:=K\_\{0\}\+K\_\{2\}\+\\widetilde\{R\}\_\{\\alpha\}\.
###### Lemma 4\(One\-step descent underℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\)\.
Assumeα∈\(0,1\)\\alpha\\in\(0,1\),γ≤1\\gamma\\leq 1, and let
ek:=vk−∇f\(xk\)\.e^\{k\}:=v^\{k\}\-\\nabla f\(x^\{k\}\)\.Then
f\(xk\+1\)≤f\(xk\)−3γ4‖∇f\(xk\)‖\+2γ‖ek‖\+K02γ2\+Cαγ\(2−α\)/\(1−α\)\.f\(x^\{k\+1\}\)\\leq f\(x^\{k\}\)\-\\frac\{3\\gamma\}\{4\}\\\|\\nabla f\(x^\{k\}\)\\\|\+2\\gamma\\\|e^\{k\}\\\|\+\\frac\{K\_\{0\}\}\{2\}\\gamma^\{2\}\+C\_\{\\alpha\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.\(29\)Consequently,
Δk\+1≤Δk−3γ4gk\+2γbk\+K02γ2\+Cαγ\(2−α\)/\(1−α\)\.\\Delta^\{k\+1\}\\leq\\Delta^\{k\}\-\\frac\{3\\gamma\}\{4\}g\_\{k\}\+2\\gamma b\_\{k\}\+\\frac\{K\_\{0\}\}\{2\}\\gamma^\{2\}\+C\_\{\\alpha\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.\(30\)
###### Proof\.
Apply \([28](https://arxiv.org/html/2605.15314#A2.E28)\) withw=xkw=x^\{k\}andw′=xk\+1w^\{\\prime\}=x^\{k\+1\}\. Since‖xk\+1−xk‖≤γ\\\|x^\{k\+1\}\-x^\{k\}\\\|\\leq\\gamma,
f\(xk\+1\)\\displaystyle f\(x^\{k\+1\}\)≤f\(xk\)−γ⟨∇f\(xk\),dk⟩\+K02γ2\\displaystyle\\leq f\(x^\{k\}\)\-\\gamma\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\rangle\+\\frac\{K\_\{0\}\}\{2\}\\gamma^\{2\}\+K12γ2‖∇f\(xk\)‖α\+K2γ\(2−α\)/\(1−α\)\.\\displaystyle\\quad\+\\frac\{K\_\{1\}\}\{2\}\\gamma^\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\+K\_\{2\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.Lemma[2](https://arxiv.org/html/2605.15314#Thmlemma2)gives
−γ⟨∇f\(xk\),dk⟩≤−γ‖∇f\(xk\)‖\+2γ‖ek‖\.\-\\gamma\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\rangle\\leq\-\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|\+2\\gamma\\\|e^\{k\}\\\|\.Using Lemma[3](https://arxiv.org/html/2605.15314#Thmlemma3)with
ρ=\(2αK1γ\)−1,\\rho=\(2\\alpha K\_\{1\}\\gamma\)^\{\-1\},we obtain
K12γ2‖∇f\(xk\)‖α≤γ4‖∇f\(xk\)‖\+Rαγ\(2−α\)/\(1−α\)\.\\frac\{K\_\{1\}\}\{2\}\\gamma^\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\\leq\\frac\{\\gamma\}\{4\}\\\|\\nabla f\(x^\{k\}\)\\\|\+R\_\{\\alpha\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.Combining the last two displays proves \([29](https://arxiv.org/html/2605.15314#A2.E29)\)\. Taking expectations gives \([30](https://arxiv.org/html/2605.15314#A2.E30)\)\. ∎
###### Lemma 5\(Gradient\-drift bound\)\.
Under the assumptions of Lemma[4](https://arxiv.org/html/2605.15314#Thmlemma4), for everykk,
‖∇f\(xk\+1\)−∇f\(xk\)‖≤C~αγ\+γ8‖∇f\(xk\)‖\.\\\|\\nabla f\(x^\{k\+1\}\)\-\\nabla f\(x^\{k\}\)\\\|\\leq\\widetilde\{C\}\_\{\\alpha\}\\gamma\+\\frac\{\\gamma\}\{8\}\\\|\\nabla f\(x^\{k\}\)\\\|\.\(31\)
###### Proof\.
By \([27](https://arxiv.org/html/2605.15314#A2.E27)\) and‖xk\+1−xk‖≤γ\\\|x^\{k\+1\}\-x^\{k\}\\\|\\leq\\gamma,
‖∇f\(xk\+1\)−∇f\(xk\)‖≤K0γ\+K1γ‖∇f\(xk\)‖α\+K2γ1/\(1−α\)\.\\\|\\nabla f\(x^\{k\+1\}\)\-\\nabla f\(x^\{k\}\)\\\|\\leq K\_\{0\}\\gamma\+K\_\{1\}\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\+K\_\{2\}\\gamma^\{1/\(1\-\\alpha\)\}\.Sinceγ≤1\\gamma\\leq 1, we have
γ1/\(1−α\)≤γ\.\\gamma^\{1/\(1\-\\alpha\)\}\\leq\\gamma\.Applying Lemma[3](https://arxiv.org/html/2605.15314#Thmlemma3)with
ρ=\(8αK1\)−1\\rho=\(8\\alpha K\_\{1\}\)^\{\-1\}gives
K1γ‖∇f\(xk\)‖α≤γ8‖∇f\(xk\)‖\+R~αγ\.K\_\{1\}\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\\leq\\frac\{\\gamma\}\{8\}\\\|\\nabla f\(x^\{k\}\)\\\|\+\\widetilde\{R\}\_\{\\alpha\}\\gamma\.Substitution proves \([31](https://arxiv.org/html/2605.15314#A2.E31)\)\. ∎
###### Theorem 4\(Single\-sample master bound underℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2), and[5](https://arxiv.org/html/2605.15314#Thmassumption5)hold with fixedα∈\(0,1\)\\alpha\\in\(0,1\)\. If
γ≤1,γ≤η,\\gamma\\leq 1,\\qquad\\gamma\\leq\\eta,then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤2ΔγT\+4b0ηT\+4C~αγη\+4η\(G\+BγK2\)\\displaystyle\\leq\\frac\{2\\Delta\}\{\\gamma T\}\+\\frac\{4b\_\{0\}\}\{\\eta T\}\+\\frac\{4\\widetilde\{C\}\_\{\\alpha\}\\gamma\}\{\\eta\}\+4\\sqrt\{\\eta\}\\left\(G\+\\frac\{B\\gamma K\}\{2\}\\right\)\+K0γ\+2Cαγ1/\(1−α\)\.\\displaystyle\\quad\+K\_\{0\}\\gamma\+2C\_\{\\alpha\}\\gamma^\{1/\(1\-\\alpha\)\}\.\(32\)Moreover,b0≤Gb\_\{0\}\\leq G\.
###### Proof\.
Let
Ek:=vk−∇f\(xk\),Δk:=∇f\(xk\)−∇f\(xk\+1\),E\_\{k\}:=v^\{k\}\-\\nabla f\(x^\{k\}\),\\qquad\\Delta\_\{k\}:=\\nabla f\(x^\{k\}\)\-\\nabla f\(x^\{k\+1\}\),and
ζk\+1:=∇f\(xk\+1;ξk\+1\)−∇f\(xk\+1\)\.\\zeta\_\{k\+1\}:=\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\+1\}\)\.Unrolling the estimator recursion gives
Ek=\(1−η\)kE0\+∑t=0k−1\(1−η\)k−tΔt\+η∑t=0k−1\(1−η\)k−1−tζt\+1\.E\_\{k\}=\(1\-\\eta\)^\{k\}E\_\{0\}\+\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-t\}\\Delta\_\{t\}\+\\eta\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}\\zeta\_\{t\+1\}\.By Lemma[5](https://arxiv.org/html/2605.15314#Thmlemma5),
‖Δt‖≤C~αγ\+γ8‖∇f\(xt\)‖\.\\\|\\Delta\_\{t\}\\\|\\leq\\widetilde\{C\}\_\{\\alpha\}\\gamma\+\\frac\{\\gamma\}\{8\}\\\|\\nabla f\(x^\{t\}\)\\\|\.The same martingale calculation as in Theorem[3](https://arxiv.org/html/2605.15314#Thmtheorem3)gives
𝔼‖η∑t=0k−1\(1−η\)k−1−tζt\+1‖≤η\(G\+Bkγ\)\.\\mathbb\{E\}\\left\\\|\\eta\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}\\zeta\_\{t\+1\}\\right\\\|\\leq\\sqrt\{\\eta\}\(G\+Bk\\gamma\)\.Therefore,
bk≤\(1−η\)kb0\+C~αγη\+γ8∑t=0k−1\(1−η\)k−tgt\+η\(G\+Bkγ\)\.b\_\{k\}\\leq\(1\-\\eta\)^\{k\}b\_\{0\}\+\\frac\{\\widetilde\{C\}\_\{\\alpha\}\\gamma\}\{\\eta\}\+\\frac\{\\gamma\}\{8\}\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-t\}g\_\{t\}\+\\sqrt\{\\eta\}\(G\+Bk\\gamma\)\.\(33\)Summing overk=0,…,Kk=0,\\ldots,Kyields
Sb\(K\)≤b0η\+C~αγTη\+γ8η∑k=0Kgk\+η\(TG\+BγKT2\)\.S\_\{b\}\(K\)\\leq\\frac\{b\_\{0\}\}\{\\eta\}\+\\frac\{\\widetilde\{C\}\_\{\\alpha\}\\gamma T\}\{\\eta\}\+\\frac\{\\gamma\}\{8\\eta\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\+\\sqrt\{\\eta\}\\left\(TG\+\\frac\{B\\gamma KT\}\{2\}\\right\)\.\(34\)
Next, summing \([30](https://arxiv.org/html/2605.15314#A2.E30)\) gives
3γ4∑k=0Kgk≤Δ\+2γSb\(K\)\+K02Tγ2\+TCαγ\(2−α\)/\(1−α\)\.\\frac\{3\\gamma\}\{4\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\\leq\\Delta\+2\\gamma S\_\{b\}\(K\)\+\\frac\{K\_\{0\}\}\{2\}T\\gamma^\{2\}\+TC\_\{\\alpha\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.Substituting \([34](https://arxiv.org/html/2605.15314#A2.E34)\) gives
γ\(34−γ4η\)∑k=0Kgk≤Δ\+2γb0η\+2C~αγ2Tη\+2γη\(TG\+BγKT2\)\\gamma\\left\(\\frac\{3\}\{4\}\-\\frac\{\\gamma\}\{4\\eta\}\\right\)\\sum\_\{k=0\}^\{K\}g\_\{k\}\\leq\\Delta\+\\frac\{2\\gamma b\_\{0\}\}\{\\eta\}\+\\frac\{2\\widetilde\{C\}\_\{\\alpha\}\\gamma^\{2\}T\}\{\\eta\}\+2\\gamma\\sqrt\{\\eta\}\\left\(TG\+\\frac\{B\\gamma KT\}\{2\}\\right\)\+K02Tγ2\+TCαγ\(2−α\)/\(1−α\)\.\\qquad\+\\frac\{K\_\{0\}\}\{2\}T\\gamma^\{2\}\+TC\_\{\\alpha\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.Sinceγ≤η\\gamma\\leq\\eta, we have
34−γ4η≥12\.\\frac\{3\}\{4\}\-\\frac\{\\gamma\}\{4\\eta\}\\geq\\frac\{1\}\{2\}\.Thus,
γ2∑k=0Kgk≤Δ\+2γb0η\+2C~αγ2Tη\+2γη\(TG\+BγKT2\)\\frac\{\\gamma\}\{2\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\\leq\\Delta\+\\frac\{2\\gamma b\_\{0\}\}\{\\eta\}\+\\frac\{2\\widetilde\{C\}\_\{\\alpha\}\\gamma^\{2\}T\}\{\\eta\}\+2\\gamma\\sqrt\{\\eta\}\\left\(TG\+\\frac\{B\\gamma KT\}\{2\}\\right\)\+K02Tγ2\+TCαγ\(2−α\)/\(1−α\)\.\\qquad\+\\frac\{K\_\{0\}\}\{2\}T\\gamma^\{2\}\+TC\_\{\\alpha\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.Dividing byγT/2\\gamma T/2and using \([11](https://arxiv.org/html/2605.15314#A2.E11)\) yields \([32](https://arxiv.org/html/2605.15314#A2.E32)\)\. Finally,b0≤Gb\_\{0\}\\leq Gfollows from Lemma[1](https://arxiv.org/html/2605.15314#Thmlemma1)\. ∎
###### Corollary 4\.1\(ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)smoothness under𝖡𝖦\\mathsf\{BG\}\-0\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2), and[5](https://arxiv.org/html/2605.15314#Thmassumption5)hold with fixedα∈\(0,1\)\\alpha\\in\(0,1\)\. Choose
η=T−2/3,γ=γ0T−5/6,0<γ0≤1\.\\eta=T^\{\-2/3\},\\qquad\\gamma=\\gamma\_\{0\}T^\{\-5/6\},\\qquad 0<\\gamma\_\{0\}\\leq 1\.Then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(2Δγ0\+4C~αγ0\+2Bγ0\)T−1/6\+8GT−1/3\\displaystyle\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+4\\widetilde\{C\}\_\{\\alpha\}\\gamma\_\{0\}\+2B\\gamma\_\{0\}\\right\)T^\{\-1/6\}\+8GT^\{\-1/3\}\+K0γ0T−5/6\+2Cαγ01/\(1−α\)T−5/\(6\(1−α\)\)\.\\displaystyle\\quad\+K\_\{0\}\\gamma\_\{0\}T^\{\-5/6\}\+2C\_\{\\alpha\}\\gamma\_\{0\}^\{1/\(1\-\\alpha\)\}T^\{\-5/\(6\(1\-\\alpha\)\)\}\.\(35\)Consequently,
𝔼‖∇f\(x^\)‖=𝒪\(T−1/6\),\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/6\}\),and the SFO complexity is𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)\.
###### Proof\.
BecauseT≥1T\\geq 1and0<γ0≤10<\\gamma\_\{0\}\\leq 1,
γ=γ0T−5/6≤1\.\\gamma=\\gamma\_\{0\}T^\{\-5/6\}\\leq 1\.Moreover,
γ=γ0T−5/6≤T−5/6≤T−2/3=η\.\\gamma=\\gamma\_\{0\}T^\{\-5/6\}\\leq T^\{\-5/6\}\\leq T^\{\-2/3\}=\\eta\.Hence, the conditions of Theorem[4](https://arxiv.org/html/2605.15314#Thmtheorem4)hold\. Substituting the schedule into \([32](https://arxiv.org/html/2605.15314#A2.E32)\), usingb0≤Gb\_\{0\}\\leq GandK≤TK\\leq T, gives
2ΔγT=2Δγ0T−1/6,4b0ηT≤4GT−1/3,\\frac\{2\\Delta\}\{\\gamma T\}=\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}T^\{\-1/6\},\\qquad\\frac\{4b\_\{0\}\}\{\\eta T\}\\leq 4GT^\{\-1/3\},4C~αγη=4C~αγ0T−1/6,\\frac\{4\\widetilde\{C\}\_\{\\alpha\}\\gamma\}\{\\eta\}=4\\widetilde\{C\}\_\{\\alpha\}\\gamma\_\{0\}T^\{\-1/6\},and
4η\(G\+BγK2\)≤4GT−1/3\+2Bγ0T−1/6\.4\\sqrt\{\\eta\}\\left\(G\+\\frac\{B\\gamma K\}\{2\}\\right\)\\leq 4GT^\{\-1/3\}\+2B\\gamma\_\{0\}T^\{\-1/6\}\.The remaining deterministic terms are
K0γ=K0γ0T−5/6,K\_\{0\}\\gamma=K\_\{0\}\\gamma\_\{0\}T^\{\-5/6\},and
2Cαγ1/\(1−α\)=2Cαγ01/\(1−α\)T−5/\(6\(1−α\)\)\.2C\_\{\\alpha\}\\gamma^\{1/\(1\-\\alpha\)\}=2C\_\{\\alpha\}\\gamma\_\{0\}^\{1/\(1\-\\alpha\)\}T^\{\-5/\(6\(1\-\\alpha\)\)\}\.Combining these estimates proves \([35](https://arxiv.org/html/2605.15314#A2.E35)\)\. Sinceα∈\(0,1\)\\alpha\\in\(0,1\)is fixed, all terms decay at least as fast asT−1/6T^\{\-1/6\}\. Therefore,
𝔼‖∇f\(x^\)‖=𝒪\(T−1/6\),\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/6\}\),which is equivalent to𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)SFO complexity\. ∎
### B\.4Recovery of bounded variance rates \(B=0B=0\)
When the distance\-growth parameterBBin Assumption[2](https://arxiv.org/html/2605.15314#Thmassumption2)vanishes, the𝖡𝖦\\mathsf\{BG\}\-0 condition reduces to bounded variance𝔼ξ‖∇f\(x;ξ\)−∇f\(x\)‖2≤G2\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x;\\xi\)\-\\nabla f\(x\)\\\|^\{2\}\\leq G^\{2\}, and the𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}complexity improves from𝒪\(ε−6\)\\mathcal\{O\}\(\\varepsilon^\{\-6\}\)to𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\. The following corollaries formalize this for all three smoothness regimes; in each case, they follow from the corresponding master bound by substitutingB=0B=0and the bounded variance schedule\.
###### Corollary 4\.2\(Standard smoothness under bounded variance\)\.
AssumeffisL0L\_\{0\}\-smooth \(Assumption[3](https://arxiv.org/html/2605.15314#Thmassumption3)\) and that Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[2](https://arxiv.org/html/2605.15314#Thmassumption2)hold withB=0B=0\. Choose
η=T−1/2,γ=γ0T−3/4,\\eta=T^\{\-1/2\},\\qquad\\gamma=\\gamma\_\{0\}\\,T^\{\-3/4\},whereT:=K\+1T:=K\+1andγ0\>0\\gamma\_\{0\}\>0\. Then
𝔼‖∇f\(x^\)‖≤\(2Δγ0\+16L0γ0\)T−1/4\+8GT−1/2\+2L0γ0T−3/4\.\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+16L\_\{0\}\\gamma\_\{0\}\\right\)T^\{\-1/4\}\+8G\\,T^\{\-1/2\}\+2L\_\{0\}\\gamma\_\{0\}\\,T^\{\-3/4\}\.\(36\)Consequently,𝔼‖∇f\(x^\)‖=𝒪\(T−1/4\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/4\}\), and the SFO complexity is𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\.
###### Proof\.
SetL1=0L\_\{1\}=0andB=0B=0in Theorem[3](https://arxiv.org/html/2605.15314#Thmtheorem3)\. TheL12L\_\{1\}^\{2\}\-coupling term vanishes and the conditions \([19](https://arxiv.org/html/2605.15314#A2.E19)\) are vacuous\. Usingb0≤Gb\_\{0\}\\leq GandK≤TK\\leq T, we compute
2ΔγT=2Δγ0T−1/4,4b0ηT≤4GT−1/2,\\frac\{2\\Delta\}\{\\gamma T\}=\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/4\},\\qquad\\frac\{4b\_\{0\}\}\{\\eta T\}\\leq 4G\\,T^\{\-1/2\},16L0γη=16L0γ0T−1/4,2L0γ=2L0γ0T−3/4,\\frac\{16L\_\{0\}\\gamma\}\{\\eta\}=16L\_\{0\}\\gamma\_\{0\}\\,T^\{\-1/4\},\\qquad 2L\_\{0\}\\gamma=2L\_\{0\}\\gamma\_\{0\}\\,T^\{\-3/4\},and, sinceB=0B=0,
4η\(G\+BγK2\)=4GT−1/4\.4\\sqrt\{\\eta\}\\left\(G\+\\frac\{B\\gamma K\}\{2\}\\right\)=4G\\,T^\{\-1/4\}\.Combining these estimates proves \([36](https://arxiv.org/html/2605.15314#A2.E36)\)\. ∎
###### Corollary 4\.3\(1\-generalized smoothness under bounded variance\)\.
Assume Assumption[5](https://arxiv.org/html/2605.15314#Thmassumption5)holds withα=1\\alpha=1and constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0, and that Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[2](https://arxiv.org/html/2605.15314#Thmassumption2)hold withB=0B=0\. Choose
η=T−1/2,γ=γ0T−3/4,0<γ0≤18L1\.\\eta=T^\{\-1/2\},\\qquad\\gamma=\\gamma\_\{0\}\\,T^\{\-3/4\},\\qquad 0<\\gamma\_\{0\}\\leq\\frac\{1\}\{8L\_\{1\}\}\.Then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(2Δγ0\+16L0γ0\)T−1/4\+8GT−1/2\+2L0γ0T−3/4\\displaystyle\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+16L\_\{0\}\\gamma\_\{0\}\\right\)T^\{\-1/4\}\+8G\\,T^\{\-1/2\}\+2L\_\{0\}\\gamma\_\{0\}\\,T^\{\-3/4\}\+64L12γ0T−1/4\[4Δ\+16L0γ02\+8γ0GT−1/4\+2L0γ02T−1/2\]\.\\displaystyle\\quad\+64L\_\{1\}^\{2\}\\gamma\_\{0\}\\,T^\{\-1/4\}\\left\[4\\Delta\+16L\_\{0\}\\gamma\_\{0\}^\{2\}\+8\\gamma\_\{0\}G\\,T^\{\-1/4\}\+2L\_\{0\}\\gamma\_\{0\}^\{2\}\\,T^\{\-1/2\}\\right\]\.\(37\)Consequently,𝔼‖∇f\(x^\)‖=𝒪\(T−1/4\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/4\}\), and the SFO complexity is𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\.
###### Proof\.
The first condition in \([19](https://arxiv.org/html/2605.15314#A2.E19)\) holds becauseγL1≤γ0L1≤1/8\\gamma L\_\{1\}\\leq\\gamma\_\{0\}L\_\{1\}\\leq 1/8\. For the second,
32L12γ2Tη=32L12γ02≤12\.\\frac\{32L\_\{1\}^\{2\}\\gamma^\{2\}T\}\{\\eta\}=32L\_\{1\}^\{2\}\\gamma\_\{0\}^\{2\}\\leq\\frac\{1\}\{2\}\.Substitute the schedule into \([20](https://arxiv.org/html/2605.15314#A2.E20)\) withB=0B=0\. The non\-coupling terms give the first line\. For the bracketed expression, usingB=0B=0,b0≤Gb\_\{0\}\\leq G, andK≤TK\\leq T:
4Δ\+4γb0η\+16L0γ2Tη\+4γηTG\+2L0γ2T4\\Delta\+\\frac\{4\\gamma b\_\{0\}\}\{\\eta\}\+\\frac\{16L\_\{0\}\\gamma^\{2\}T\}\{\\eta\}\+4\\gamma\\sqrt\{\\eta\}\\,TG\+2L\_\{0\}\\gamma^\{2\}T=4Δ\+4γ0GT−1/4\+16L0γ02\+4γ0GT−1/4\+2L0γ02T−1/2\.=4\\Delta\+4\\gamma\_\{0\}G\\,T^\{\-1/4\}\+16L\_\{0\}\\gamma\_\{0\}^\{2\}\+4\\gamma\_\{0\}G\\,T^\{\-1/4\}\+2L\_\{0\}\\gamma\_\{0\}^\{2\}\\,T^\{\-1/2\}\.The twoGG\-terms combine to8γ0GT−1/48\\gamma\_\{0\}G\\,T^\{\-1/4\}, yielding the stated bracket\. The prefactor is64L12γ0T−1/464L\_\{1\}^\{2\}\\gamma\_\{0\}\\,T^\{\-1/4\}\. Since the bracket is𝒪\(1\)\\mathcal\{O\}\(1\), the entire coupling term is𝒪\(T−1/4\)\\mathcal\{O\}\(T^\{\-1/4\}\)\. ∎
###### Corollary 4\.4\(ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)smoothness under bounded variance\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2)\(withB=0B=0\), and[5](https://arxiv.org/html/2605.15314#Thmassumption5)hold with fixedα∈\(0,1\)\\alpha\\in\(0,1\)\. Choose
η=T−1/2,γ=γ0T−3/4,0<γ0≤1\.\\eta=T^\{\-1/2\},\\qquad\\gamma=\\gamma\_\{0\}\\,T^\{\-3/4\},\\qquad 0<\\gamma\_\{0\}\\leq 1\.Then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(2Δγ0\+4C~αγ0\)T−1/4\+8GT−1/2\\displaystyle\\leq\\left\(\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\+4\\widetilde\{C\}\_\{\\alpha\}\\gamma\_\{0\}\\right\)T^\{\-1/4\}\+8G\\,T^\{\-1/2\}\+K0γ0T−3/4\+2Cαγ01/\(1−α\)T−3/\(4\(1−α\)\)\.\\displaystyle\\quad\+K\_\{0\}\\gamma\_\{0\}\\,T^\{\-3/4\}\+2C\_\{\\alpha\}\\gamma\_\{0\}^\{1/\(1\-\\alpha\)\}\\,T^\{\-3/\(4\(1\-\\alpha\)\)\}\.\(38\)Consequently,𝔼‖∇f\(x^\)‖=𝒪\(T−1/4\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/4\}\), and the SFO complexity is𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\.
###### Proof\.
Sinceγ0≤1\\gamma\_\{0\}\\leq 1andT≥1T\\geq 1, we haveγ≤1\\gamma\\leq 1andγ≤η\\gamma\\leq\\eta\. The conditions of Theorem[4](https://arxiv.org/html/2605.15314#Thmtheorem4)hold\. Substituting the schedule into \([32](https://arxiv.org/html/2605.15314#A2.E32)\) withB=0B=0and usingb0≤Gb\_\{0\}\\leq G,
2ΔγT=2Δγ0T−1/4,4b0ηT≤4GT−1/2,\\frac\{2\\Delta\}\{\\gamma T\}=\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/4\},\\qquad\\frac\{4b\_\{0\}\}\{\\eta T\}\\leq 4G\\,T^\{\-1/2\},4C~αγη=4C~αγ0T−1/4,\\frac\{4\\widetilde\{C\}\_\{\\alpha\}\\gamma\}\{\\eta\}=4\\widetilde\{C\}\_\{\\alpha\}\\gamma\_\{0\}\\,T^\{\-1/4\},and, sinceB=0B=0,
4η\(G\+BγK2\)=4GT−1/4\.4\\sqrt\{\\eta\}\\left\(G\+\\frac\{B\\gamma K\}\{2\}\\right\)=4G\\,T^\{\-1/4\}\.The remaining deterministic terms are
K0γ=K0γ0T−3/4,2Cαγ1/\(1−α\)=2Cαγ01/\(1−α\)T−3/\(4\(1−α\)\)\.K\_\{0\}\\gamma=K\_\{0\}\\gamma\_\{0\}\\,T^\{\-3/4\},\\qquad 2C\_\{\\alpha\}\\gamma^\{1/\(1\-\\alpha\)\}=2C\_\{\\alpha\}\\gamma\_\{0\}^\{1/\(1\-\\alpha\)\}\\,T^\{\-3/\(4\(1\-\\alpha\)\)\}\.For theGG\-terms,4GT−1/2\+4GT−1/4≤8GT−1/44G\\,T^\{\-1/2\}\+4G\\,T^\{\-1/4\}\\leq 8G\\,T^\{\-1/4\}whenT≥1T\\geq 1\. Sinceα∈\(0,1\)\\alpha\\in\(0,1\)is fixed,3/\(4\(1−α\)\)\>3/4\>1/43/\(4\(1\-\\alpha\)\)\>3/4\>1/4, so every exponent on the right\-hand side is at least1/41/4\. This proves \([38](https://arxiv.org/html/2605.15314#A2.E38)\) and the𝒪\(T−1/4\)\\mathcal\{O\}\(T^\{\-1/4\}\)rate\. ∎
### B\.5Recovery of deterministic rates \(B=G=0B=G=0\)
When the oracle is deterministic \(B=G=0B=G=0\), we havevk=∇f\(xk\)v^\{k\}=\\nabla f\(x^\{k\}\)for everykk\(sinceη=1\\eta=1is natural and the oracle returns the exact gradient\)\. Thusbk=0b\_\{k\}=0for allkk, and the stochastic error analysis simplifies to a pure descent argument\. We recover the classical𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)rate for all smoothness regimes\.
###### Corollary 4\.5\(Standard smoothness, deterministic\)\.
AssumeffisL0L\_\{0\}\-smooth \(Assumption[3](https://arxiv.org/html/2605.15314#Thmassumption3)\), that Assumption[1](https://arxiv.org/html/2605.15314#Thmassumption1)holds, and that the oracle is deterministic \(B=G=0B=G=0\)\. Consider the normalized gradient descent update
xk\+1=xk−γ∇f\(xk\)‖∇f\(xk\)‖\.x^\{k\+1\}=x^\{k\}\-\\gamma\\,\\frac\{\\nabla f\(x^\{k\}\)\}\{\\\|\\nabla f\(x^\{k\}\)\\\|\}\.Chooseγ=γ0T−1/2\\gamma=\\gamma\_\{0\}\\,T^\{\-1/2\}\. Then
𝔼‖∇f\(x^\)‖≤2Δγ0T−1/2\+2L0γ0T−1/2\.\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/2\}\+2L\_\{0\}\\gamma\_\{0\}\\,T^\{\-1/2\}\.\(39\)Consequently,‖∇f\(x^\)‖=𝒪\(T−1/2\)\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/2\}\), and the iteration complexity is𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)\.
###### Proof\.
Under determinism, setη=1\\eta=1so thatvk=∇f\(xk\)v^\{k\}=\\nabla f\(x^\{k\}\)for allkk\. Thenbk=0b\_\{k\}=0and the descent recursion \([21](https://arxiv.org/html/2605.15314#A2.E21)\) \(withL1=0L\_\{1\}=0\) becomes
f\(xk\+1\)≤f\(xk\)−γ‖∇f\(xk\)‖\+L0γ2\.f\(x^\{k\+1\}\)\\leq f\(x^\{k\}\)\-\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|\+L\_\{0\}\\gamma^\{2\}\.Here, the coefficient isγ\\gammarather thanγ/2\\gamma/2because, withbk=0b\_\{k\}=0, the direction\-error term vanishes and no gradient term needs to absorb a generalized smoothness contribution\. Summing fromk=0k=0toKKand usingf\(xK\+1\)≥finff\(x^\{K\+1\}\)\\geq f^\{\\inf\},
γ∑k=0K‖∇f\(xk\)‖≤Δ\+L0γ2T\.\\gamma\\sum\_\{k=0\}^\{K\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\Delta\+L\_\{0\}\\gamma^\{2\}T\.Dividing byγT\\gamma Tgives
1T∑k=0K‖∇f\(xk\)‖≤ΔγT\+L0γ\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\frac\{\\Delta\}\{\\gamma T\}\+L\_\{0\}\\gamma\.Substitutingγ=γ0T−1/2\\gamma=\\gamma\_\{0\}T^\{\-1/2\}yieldsΔ/\(γ0\)T−1/2\+L0γ0T−1/2\\Delta/\(\\gamma\_\{0\}\)\\,T^\{\-1/2\}\+L\_\{0\}\\gamma\_\{0\}\\,T^\{\-1/2\}, from which \([39](https://arxiv.org/html/2605.15314#A2.E39)\) follows \(with an extra factor of22for a uniform upper bound\. ∎
###### Corollary 4\.6\(1\-generalized smoothness, deterministic\)\.
Assume Assumption[5](https://arxiv.org/html/2605.15314#Thmassumption5)holds withα=1\\alpha=1and constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0, that Assumption[1](https://arxiv.org/html/2605.15314#Thmassumption1)holds, and that the oracle is deterministic \(B=G=0B=G=0\)\. Consider the normalized gradient descent update
xk\+1=xk−γ∇f\(xk\)‖∇f\(xk\)‖\.x^\{k\+1\}=x^\{k\}\-\\gamma\\,\\frac\{\\nabla f\(x^\{k\}\)\}\{\\\|\\nabla f\(x^\{k\}\)\\\|\}\.Chooseγ=γ0T−1/2\\gamma=\\gamma\_\{0\}\\,T^\{\-1/2\}with0<γ0≤1/\(2L1\)0<\\gamma\_\{0\}\\leq 1/\(2L\_\{1\}\)\. Then
‖∇f\(x^\)‖≤4Δγ0T−1/2\+4L0γ0T−1/2\.\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/2\}\+4L\_\{0\}\\gamma\_\{0\}\\,T^\{\-1/2\}\.\(40\)Consequently,‖∇f\(x^\)‖=𝒪\(T−1/2\)\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/2\}\), and the iteration complexity is𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)\.
###### Proof\.
Under determinism,vk=∇f\(xk\)v^\{k\}=\\nabla f\(x^\{k\}\),bk=0b\_\{k\}=0\. Using \([17](https://arxiv.org/html/2605.15314#A2.E17)\) withy=xk\+1=xk−γdky=x^\{k\+1\}=x^\{k\}\-\\gamma d^\{k\}, and noting thatdk=∇f\(xk\)/‖∇f\(xk\)‖d^\{k\}=\\nabla f\(x^\{k\}\)/\\\|\\nabla f\(x^\{k\}\)\\\|when the gradient is nonzero \(so⟨∇f\(xk\),dk⟩=‖∇f\(xk\)‖\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\rangle=\\\|\\nabla f\(x^\{k\}\)\\\|\),
f\(xk\+1\)≤f\(xk\)−γ‖∇f\(xk\)‖\+γ22eL1γ\(L0\+L1‖∇f\(xk\)‖\)\.f\(x^\{k\+1\}\)\\leq f\(x^\{k\}\)\-\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|\+\\frac\{\\gamma^\{2\}\}\{2\}e^\{L\_\{1\}\\gamma\}\\bigl\(L\_\{0\}\+L\_\{1\}\\\|\\nabla f\(x^\{k\}\)\\\|\\bigr\)\.Sinceγ0≤1/\(2L1\)\\gamma\_\{0\}\\leq 1/\(2L\_\{1\}\)andT≥1T\\geq 1, we haveγL1≤γ0L1≤1/2\\gamma L\_\{1\}\\leq\\gamma\_\{0\}L\_\{1\}\\leq 1/2, soeL1γ≤2e^\{L\_\{1\}\\gamma\}\\leq 2\. Therefore,
γ2eL1γL1‖∇f\(xk\)‖≤2γ2L1‖∇f\(xk\)‖≤γ‖∇f\(xk\)‖⋅2γL1≤γ2‖∇f\(xk\)‖,\\gamma^\{2\}e^\{L\_\{1\}\\gamma\}L\_\{1\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq 2\\gamma^\{2\}L\_\{1\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|\\cdot 2\\gamma L\_\{1\}\\leq\\frac\{\\gamma\}\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|,since2γL1≤2γ0L1≤12\\gamma L\_\{1\}\\leq 2\\gamma\_\{0\}L\_\{1\}\\leq 1\. This gives
f\(xk\+1\)≤f\(xk\)−γ2‖∇f\(xk\)‖\+L0γ2\.f\(x^\{k\+1\}\)\\leq f\(x^\{k\}\)\-\\frac\{\\gamma\}\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|\+L\_\{0\}\\gamma^\{2\}\.Summing fromk=0k=0toKKand dividing byγT/2\\gamma T/2,
1T∑k=0K‖∇f\(xk\)‖≤2ΔγT\+2L0γ\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\frac\{2\\Delta\}\{\\gamma T\}\+2L\_\{0\}\\gamma\.Substitutingγ=γ0T−1/2\\gamma=\\gamma\_\{0\}T^\{\-1/2\}yields \([40](https://arxiv.org/html/2605.15314#A2.E40)\)\. ∎
###### Corollary 4\.7\(ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)smoothness, deterministic\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[5](https://arxiv.org/html/2605.15314#Thmassumption5)hold with fixedα∈\(0,1\)\\alpha\\in\(0,1\), and that the oracle is deterministic \(B=G=0B=G=0\)\. Consider the normalized gradient descent update
xk\+1=xk−γ∇f\(xk\)‖∇f\(xk\)‖\.x^\{k\+1\}=x^\{k\}\-\\gamma\\,\\frac\{\\nabla f\(x^\{k\}\)\}\{\\\|\\nabla f\(x^\{k\}\)\\\|\}\.Chooseγ=γ0T−1/2\\gamma=\\gamma\_\{0\}\\,T^\{\-1/2\}with0<γ0≤10<\\gamma\_\{0\}\\leq 1\. Then
‖∇f\(x^\)‖≤4Δ3γ0T−1/2\+2K03γ0T−1/2\+4Cα3γ01/\(1−α\)T−1/\(2\(1−α\)\)\.\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\frac\{4\\Delta\}\{3\\gamma\_\{0\}\}\\,T^\{\-1/2\}\+\\frac\{2K\_\{0\}\}\{3\}\\gamma\_\{0\}\\,T^\{\-1/2\}\+\\frac\{4C\_\{\\alpha\}\}\{3\}\\gamma\_\{0\}^\{1/\(1\-\\alpha\)\}\\,T^\{\-1/\(2\(1\-\\alpha\)\)\}\.\(41\)Consequently,‖∇f\(x^\)‖=𝒪\(T−1/2\)\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/2\}\), and the iteration complexity is𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)\.
###### Proof\.
Under determinism,vk=∇f\(xk\)v^\{k\}=\\nabla f\(x^\{k\}\), so‖ek‖=0\\\|e^\{k\}\\\|=0anddk=∇f\(xk\)/‖∇f\(xk\)‖d^\{k\}=\\nabla f\(x^\{k\}\)/\\\|\\nabla f\(x^\{k\}\)\\\|\. The one\-step descent \([29](https://arxiv.org/html/2605.15314#A2.E29)\) from Lemma[4](https://arxiv.org/html/2605.15314#Thmlemma4)becomes
f\(xk\+1\)≤f\(xk\)−3γ4‖∇f\(xk\)‖\+K02γ2\+Cαγ\(2−α\)/\(1−α\)\.f\(x^\{k\+1\}\)\\leq f\(x^\{k\}\)\-\\frac\{3\\gamma\}\{4\}\\\|\\nabla f\(x^\{k\}\)\\\|\+\\frac\{K\_\{0\}\}\{2\}\\gamma^\{2\}\+C\_\{\\alpha\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.Summing fromk=0k=0toKKand usingf\(xK\+1\)≥finff\(x^\{K\+1\}\)\\geq f^\{\\inf\},
3γ4∑k=0K‖∇f\(xk\)‖≤Δ\+K02Tγ2\+TCαγ\(2−α\)/\(1−α\)\.\\frac\{3\\gamma\}\{4\}\\sum\_\{k=0\}^\{K\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\Delta\+\\frac\{K\_\{0\}\}\{2\}T\\gamma^\{2\}\+TC\_\{\\alpha\}\\gamma^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}\.Dividing by\(3γ/4\)⋅T\(3\\gamma/4\)\\cdot T,
1T∑k=0K‖∇f\(xk\)‖≤4Δ3γT\+2K03γ\+4Cα3γ1/\(1−α\)\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\frac\{4\\Delta\}\{3\\gamma T\}\+\\frac\{2K\_\{0\}\}\{3\}\\gamma\+\\frac\{4C\_\{\\alpha\}\}\{3\}\\gamma^\{1/\(1\-\\alpha\)\}\.Substitutingγ=γ0T−1/2\\gamma=\\gamma\_\{0\}T^\{\-1/2\}gives
4Δ3γT=4Δ3γ0T−1/2,2K03γ=2K03γ0T−1/2,\\frac\{4\\Delta\}\{3\\gamma T\}=\\frac\{4\\Delta\}\{3\\gamma\_\{0\}\}\\,T^\{\-1/2\},\\qquad\\frac\{2K\_\{0\}\}\{3\}\\gamma=\\frac\{2K\_\{0\}\}\{3\}\\gamma\_\{0\}\\,T^\{\-1/2\},and
4Cα3γ1/\(1−α\)=4Cα3γ01/\(1−α\)T−1/\(2\(1−α\)\)\.\\frac\{4C\_\{\\alpha\}\}\{3\}\\gamma^\{1/\(1\-\\alpha\)\}=\\frac\{4C\_\{\\alpha\}\}\{3\}\\gamma\_\{0\}^\{1/\(1\-\\alpha\)\}\\,T^\{\-1/\(2\(1\-\\alpha\)\)\}\.Sinceα∈\(0,1\)\\alpha\\in\(0,1\), we have1/\(2\(1−α\)\)\>1/21/\(2\(1\-\\alpha\)\)\>1/2, so every term decays at least as fast asT−1/2T^\{\-1/2\}\. This proves \([41](https://arxiv.org/html/2605.15314#A2.E41)\) and the𝒪\(T−1/2\)\\mathcal\{O\}\(T^\{\-1/2\}\)rate\. ∎
### B\.6Discussion on recovered rates
The six recovery corollaries above confirm the claims made in the discussion following Theorem[1](https://arxiv.org/html/2605.15314#Thmtheorem1)\.
- •Bounded variance recovery\.WhenB=0B=0, Corollaries[4\.2](https://arxiv.org/html/2605.15314#Thmtheorem4.Thmcorollary2),[4\.3](https://arxiv.org/html/2605.15314#Thmtheorem4.Thmcorollary3), and[4\.4](https://arxiv.org/html/2605.15314#Thmtheorem4.Thmcorollary4)all achieve𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)SFO complexity with the standard bounded variance scheduleη=T−1/2\\eta=T^\{\-1/2\},γ=γ0T−3/4\\gamma=\\gamma\_\{0\}T^\{\-3/4\}\. This recovers the rate ofCutkosky and Mehta \[[9](https://arxiv.org/html/2605.15314#bib.bib144)\]for standard smoothness, ofKhiriratet al\.\[[26](https://arxiv.org/html/2605.15314#bib.bib154)\]forℒsym∗\(1\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\), and establishes the same rate forℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)withα∈\(0,1\)\\alpha\\in\(0,1\), which to our knowledge is new\.
- •Deterministic recovery\.WhenB=G=0B=G=0, Corollaries[4\.5](https://arxiv.org/html/2605.15314#Thmtheorem4.Thmcorollary5),[4\.6](https://arxiv.org/html/2605.15314#Thmtheorem4.Thmcorollary6), and[4\.7](https://arxiv.org/html/2605.15314#Thmtheorem4.Thmcorollary7)all achieve𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)iteration complexity\. This recovers the normalized gradient descent rates ofChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]for all three smoothness regimes\.
- •Schedule transition\.Comparing the three noise regimes clarifies the role of the schedule parameters\. As the noise weakens from𝖡𝖦\\mathsf\{BG\}\-0 to bounded variance to deterministic, the schedules transition as: \(η,γ\)=\{\(T−2/3,γ0T−5/6\),𝖡𝖦\-0,\(T−1/2,γ0T−3/4\),bounded variance,\(1,γ0T−1/2\),deterministic,\(\\eta,\\gamma\)=\\begin\{cases\}\(T^\{\-2/3\},\\;\\gamma\_\{0\}T^\{\-5/6\}\),&\\mathsf\{BG\}\\text\{\-\}0,\\\\\[3\.0pt\] \(T^\{\-1/2\},\\;\\gamma\_\{0\}T^\{\-3/4\}\),&\\text\{bounded variance\},\\\\\[3\.0pt\] \(1,\\;\\gamma\_\{0\}T^\{\-1/2\}\),&\\text\{deterministic\},\\end\{cases\}and the corresponding convergence rates areT−1/6T^\{\-1/6\},T−1/4T^\{\-1/4\}, andT−1/2T^\{\-1/2\}, respectively\. The faster schedules are enabled by the absence of the distance\-dependent noise growth termB2‖xk−x0‖2B^\{2\}\\\|x^\{k\}\-x^\{0\}\\\|^\{2\}\.
## Appendix CProofs for Normalized STORM under𝖡𝖦\\mathsf\{BG\}\-0 Noise
This appendix proves Theorem[2](https://arxiv.org/html/2605.15314#Thmtheorem2)\. Throughout this section, let
T:=K\+1,Δ:=f\(x0\)−finf\.T:=K\+1,\\qquad\\Delta:=f\(x^\{0\}\)\-f^\{\\inf\}\.We define
gk:=𝔼‖∇f\(xk\)‖,bk:=𝔼‖vk−∇f\(xk\)‖,g\_\{k\}:=\\mathbb\{E\}\\\|\\nabla f\(x^\{k\}\)\\\|,\\qquad b\_\{k\}:=\\mathbb\{E\}\\\|v^\{k\}\-\\nabla f\(x^\{k\}\)\\\|,and
Δk:=𝔼\[f\(xk\)−finf\],Sb\(K\):=∑k=0Kbk\.\\Delta^\{k\}:=\\mathbb\{E\}\[f\(x^\{k\}\)\-f^\{\\inf\}\],\\qquad S\_\{b\}\(K\):=\\sum\_\{k=0\}^\{K\}b\_\{k\}\.The outputx^\\widehat\{x\}is sampled uniformly from\{x0,x1,…,xK\}\\\{x^\{0\},x^\{1\},\\ldots,x^\{K\}\\\}, so
𝔼‖∇f\(x^\)‖=1T∑k=0Kgk\.\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\.\(42\)For compactness, define the local BG\-0 noise scale
q\(x\):=\(B2‖x−x0‖2\+G2\)1/2,qk:=q\(xk\)\.q\(x\):=\\left\(B^\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G^\{2\}\\right\)^\{1/2\},\\qquad q\_\{k\}:=q\(x^\{k\}\)\.
### C\.1Basic inequalities
###### Lemma 6\(Trajectory control\)\.
For the𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}iterates in \([3](https://arxiv.org/html/2605.15314#S5.E3)\),
‖xk\+1−xk‖≤γ,‖xk−x0‖≤kγ\.\\\|x^\{k\+1\}\-x^\{k\}\\\|\\leq\\gamma,\\qquad\\\|x^\{k\}\-x^\{0\}\\\|\\leq k\\gamma\.Consequently,
qk≤G\+Bkγ≤G\+BTγ,k=0,…,K\.q\_\{k\}\\leq G\+Bk\\gamma\\leq G\+BT\\gamma,\\qquad k=0,\\ldots,K\.Moreover, for everyβ∈\(0,1\]\\beta\\in\(0,1\],
1T∑k=0K𝔼\[qkβ\]≤Gβ\+Bβ\(Tγ\)β\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}^\{\\beta\}\]\\leq G^\{\\beta\}\+B^\{\\beta\}\(T\\gamma\)^\{\\beta\}\.\(43\)
###### Proof\.
The normalized direction has norm at most one, so
‖xk\+1−xk‖≤γ\.\\\|x^\{k\+1\}\-x^\{k\}\\\|\\leq\\gamma\.Summing the increments gives
‖xk−x0‖≤∑t=0k−1‖xt\+1−xt‖≤kγ\.\\\|x^\{k\}\-x^\{0\}\\\|\\leq\\sum\_\{t=0\}^\{k\-1\}\\\|x^\{t\+1\}\-x^\{t\}\\\|\\leq k\\gamma\.Therefore,
qk=\(B2‖xk−x0‖2\+G2\)1/2≤B‖xk−x0‖\+G≤G\+Bkγ≤G\+BTγ\.q\_\{k\}=\\left\(B^\{2\}\\\|x^\{k\}\-x^\{0\}\\\|^\{2\}\+G^\{2\}\\right\)^\{1/2\}\\leq B\\\|x^\{k\}\-x^\{0\}\\\|\+G\\leq G\+Bk\\gamma\\leq G\+BT\\gamma\.Forβ∈\(0,1\]\\beta\\in\(0,1\], using\(a\+b\)β≤aβ\+bβ\(a\+b\)^\{\\beta\}\\leq a^\{\\beta\}\+b^\{\\beta\},
qkβ≤Gβ\+Bβ\(kγ\)β≤Gβ\+Bβ\(Tγ\)β\.q\_\{k\}^\{\\beta\}\\leq G^\{\\beta\}\+B^\{\\beta\}\(k\\gamma\)^\{\\beta\}\\leq G^\{\\beta\}\+B^\{\\beta\}\(T\\gamma\)^\{\\beta\}\.Averaging overk=0,…,Kk=0,\\ldots,Kgives \([43](https://arxiv.org/html/2605.15314#A3.E43)\)\. ∎
### C\.2Descent inequalities
###### Lemma 7\(Descent under mean\-square smoothness\)\.
Suppose Assumption[4](https://arxiv.org/html/2605.15314#Thmassumption4)holds\. Then, for everyk=0,…,Kk=0,\\ldots,K,
Δk\+1\+γgk≤Δk\+2γbk\+L2γ2\.\\Delta^\{k\+1\}\+\\gamma g\_\{k\}\\leq\\Delta^\{k\}\+2\\gamma b\_\{k\}\+\\frac\{L\}\{2\}\\gamma^\{2\}\.\(44\)
###### Proof\.
Assumption[4](https://arxiv.org/html/2605.15314#Thmassumption4)and Jensen’s inequality imply thatffisLL\-smooth:
‖∇f\(y\)−∇f\(x\)‖≤L‖y−x‖\.\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|\\leq L\\\|y\-x\\\|\.Thus,
f\(xk\+1\)≤f\(xk\)−γ⟨∇f\(xk\),dk⟩\+L2γ2\.f\(x^\{k\+1\}\)\\leq f\(x^\{k\}\)\-\\gamma\\left\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\right\\rangle\+\\frac\{L\}\{2\}\\gamma^\{2\}\.Using Lemma[2](https://arxiv.org/html/2605.15314#Thmlemma2),
−γ⟨∇f\(xk\),dk⟩≤−γ‖∇f\(xk\)‖\+2γ‖vk−∇f\(xk\)‖\.\-\\gamma\\left\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\right\\rangle\\leq\-\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|\+2\\gamma\\\|v^\{k\}\-\\nabla f\(x^\{k\}\)\\\|\.Subtractingfinff^\{\\inf\}and taking expectations gives \([44](https://arxiv.org/html/2605.15314#A3.E44)\)\. ∎
###### Lemma 8\(Sample\-gradient moment under𝖡𝖦\\mathsf\{BG\}\-0\)\.
Forα∈\(0,1\)\\alpha\\in\(0,1\),
𝔼ξ‖∇f\(x;ξ\)‖α≤2α/2\(‖∇f\(x\)‖α\+q\(x\)α\)\.\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x;\\xi\)\\\|^\{\\alpha\}\\leq 2^\{\\alpha/2\}\\left\(\\\|\\nabla f\(x\)\\\|^\{\\alpha\}\+q\(x\)^\{\\alpha\}\\right\)\.
###### Proof\.
By Jensen’s inequality,
𝔼ξ‖∇f\(x;ξ\)‖α≤\(𝔼ξ‖∇f\(x;ξ\)‖2\)α/2\.\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x;\\xi\)\\\|^\{\\alpha\}\\leq\\left\(\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x;\\xi\)\\\|^\{2\}\\right\)^\{\\alpha/2\}\.Using
∇f\(x;ξ\)=∇f\(x\)\+\(∇f\(x;ξ\)−∇f\(x\)\),\\nabla f\(x;\\xi\)=\\nabla f\(x\)\+\\bigl\(\\nabla f\(x;\\xi\)\-\\nabla f\(x\)\\bigr\),and Assumption[2](https://arxiv.org/html/2605.15314#Thmassumption2),
𝔼ξ‖∇f\(x;ξ\)‖2≤2‖∇f\(x\)‖2\+2q\(x\)2\.\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x;\\xi\)\\\|^\{2\}\\leq 2\\\|\\nabla f\(x\)\\\|^\{2\}\+2q\(x\)^\{2\}\.Sinceα/2∈\(0,1\)\\alpha/2\\in\(0,1\),
\(‖∇f\(x\)‖2\+q\(x\)2\)α/2≤‖∇f\(x\)‖α\+q\(x\)α\.\\left\(\\\|\\nabla f\(x\)\\\|^\{2\}\+q\(x\)^\{2\}\\right\)^\{\\alpha/2\}\\leq\\\|\\nabla f\(x\)\\\|^\{\\alpha\}\+q\(x\)^\{\\alpha\}\.Combining the last three displays proves the claim\. ∎
###### Lemma 9\(Descent under expectedα\\alpha\-symmetric generalized smoothness\)\.
Suppose Assumptions[2](https://arxiv.org/html/2605.15314#Thmassumption2)and[6](https://arxiv.org/html/2605.15314#Thmassumption6)hold withα∈\(0,1\)\\alpha\\in\(0,1\)\. Let
p:=α1−α,K0:=22−α1−αL0,K1:=22−α1−αL1,K2:=\(5L1\)11−α\.p:=\\frac\{\\alpha\}\{1\-\\alpha\},\\qquad K\_\{0\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{0\},\\qquad K\_\{1\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{1\},\\qquad K\_\{2\}:=\(5L\_\{1\}\)^\{\\frac\{1\}\{1\-\\alpha\}\}\.Define
Lα:=2α/2K1,cα:=\(1−α\)αα/\(1−α\)\.L\_\{\\alpha\}:=2^\{\\alpha/2\}K\_\{1\},\\qquad c\_\{\\alpha\}:=\(1\-\\alpha\)\\alpha^\{\\alpha/\(1\-\\alpha\)\}\.Letλ\>0\\lambda\>0, and suppose
Lαλγ≤1\.L\_\{\\alpha\}\\lambda\\gamma\\leq 1\.Then, for everyk=0,…,Kk=0,\\ldots,K,
Δk\+1\+γ2gk≤\\displaystyle\\Delta^\{k\+1\}\+\\frac\{\\gamma\}\{2\}g\_\{k\}\\leqΔk\+2γbk\+γ22\(K0\+cαLαλ−p\+2K2γp\)\+γ22Lα𝔼\[qkα\]\.\\displaystyle\\Delta^\{k\}\+2\\gamma b\_\{k\}\+\\frac\{\\gamma^\{2\}\}\{2\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda^\{\-p\}\+2K\_\{2\}\\gamma^\{p\}\\right\)\+\\frac\{\\gamma^\{2\}\}\{2\}L\_\{\\alpha\}\\mathbb\{E\}\[q\_\{k\}^\{\\alpha\}\]\.\(45\)
###### Proof\.
The expected generalized smoothness reduction gives \(seeChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\], Proposition 4\), fory=xk\+1=xk−γdky=x^\{k\+1\}=x^\{k\}\-\\gamma d^\{k\},
f\(xk\+1\)≤\\displaystyle f\(x^\{k\+1\}\)\\leqf\(xk\)−γ⟨∇f\(xk\),dk⟩\+γ22\(K0\+K1𝔼ξ‖∇f\(xk;ξ\)‖α\+2K2γp\)\.\\displaystyle f\(x^\{k\}\)\-\\gamma\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\rangle\+\\frac\{\\gamma^\{2\}\}\{2\}\\left\(K\_\{0\}\+K\_\{1\}\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x^\{k\};\\xi\)\\\|^\{\\alpha\}\+2K\_\{2\}\\gamma^\{p\}\\right\)\.By Lemma[8](https://arxiv.org/html/2605.15314#Thmlemma8),
K1𝔼ξ‖∇f\(xk;ξ\)‖α≤Lα\(‖∇f\(xk\)‖α\+qkα\)\.K\_\{1\}\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x^\{k\};\\xi\)\\\|^\{\\alpha\}\\leq L\_\{\\alpha\}\\left\(\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\+q\_\{k\}^\{\\alpha\}\\right\)\.Using Lemma[2](https://arxiv.org/html/2605.15314#Thmlemma2), subtractingfinff^\{\\inf\}, and taking expectations,
Δk\+1≤\\displaystyle\\Delta^\{k\+1\}\\leqΔk−γgk\+2γbk\+γ22\(K0\+2K2γp\)\+γ22Lα𝔼‖∇f\(xk\)‖α\+γ22Lα𝔼\[qkα\]\.\\displaystyle\\Delta^\{k\}\-\\gamma g\_\{k\}\+2\\gamma b\_\{k\}\+\\frac\{\\gamma^\{2\}\}\{2\}\\left\(K\_\{0\}\+2K\_\{2\}\\gamma^\{p\}\\right\)\+\\frac\{\\gamma^\{2\}\}\{2\}L\_\{\\alpha\}\\mathbb\{E\}\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\+\\frac\{\\gamma^\{2\}\}\{2\}L\_\{\\alpha\}\\mathbb\{E\}\[q\_\{k\}^\{\\alpha\}\]\.Sinceu↦uαu\\mapsto u^\{\\alpha\}is concave,
𝔼‖∇f\(xk\)‖α≤gkα\.\\mathbb\{E\}\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\\leq g\_\{k\}^\{\\alpha\}\.By Lemma[1](https://arxiv.org/html/2605.15314#Thmremark1),
gkα≤λgk\+cαλ−p\.g\_\{k\}^\{\\alpha\}\\leq\\lambda g\_\{k\}\+c\_\{\\alpha\}\\lambda^\{\-p\}\.Therefore,
γ22Lαgkα≤γ22Lαλgk\+γ22cαLαλ−p\.\\frac\{\\gamma^\{2\}\}\{2\}L\_\{\\alpha\}g\_\{k\}^\{\\alpha\}\\leq\\frac\{\\gamma^\{2\}\}\{2\}L\_\{\\alpha\}\\lambda g\_\{k\}\+\\frac\{\\gamma^\{2\}\}\{2\}c\_\{\\alpha\}L\_\{\\alpha\}\\lambda^\{\-p\}\.The conditionLαλγ≤1L\_\{\\alpha\}\\lambda\\gamma\\leq 1gives
γ22Lαλgk≤γ2gk\.\\frac\{\\gamma^\{2\}\}\{2\}L\_\{\\alpha\}\\lambda g\_\{k\}\\leq\\frac\{\\gamma\}\{2\}g\_\{k\}\.Moving this term to the left proves \([45](https://arxiv.org/html/2605.15314#A3.E45)\)\. ∎
###### Lemma 10\(Descent under expected11\-symmetric generalized smoothness\)\.
Suppose Assumptions[2](https://arxiv.org/html/2605.15314#Thmassumption2)and[6](https://arxiv.org/html/2605.15314#Thmassumption6)hold withα=1\\alpha=1\. If
0<γ≤142L1,0<\\gamma\\leq\\frac\{1\}\{4\\sqrt\{2\}L\_\{1\}\},then, for everyk=0,…,Kk=0,\\ldots,K,
Δk\+1\+γ2gk≤Δk\+2γbk\+2L0γ2\+22L1γ2𝔼\[qk\]\.\\Delta^\{k\+1\}\+\\frac\{\\gamma\}\{2\}g\_\{k\}\\leq\\Delta^\{k\}\+2\\gamma b\_\{k\}\+\\sqrt\{2\}L\_\{0\}\\gamma^\{2\}\+2\\sqrt\{2\}L\_\{1\}\\gamma^\{2\}\\mathbb\{E\}\[q\_\{k\}\]\.\(46\)
###### Proof\.
The expectedα=1\\alpha=1generalized smoothness reduction yields the segment bound
f\(y\)≤f\(x\)\+⟨∇f\(x\),y−x⟩\+‖y−x‖2\(2L0\+22L1‖∇f\(x\)‖\+22L1q\(x\)\),\\displaystyle f\(y\)\\leq f\(x\)\+\\langle\\nabla f\(x\),y\-x\\rangle\+\\\|y\-x\\\|^\{2\}\\left\(\\sqrt\{2\}L\_\{0\}\+2\\sqrt\{2\}L\_\{1\}\\\|\\nabla f\(x\)\\\|\+2\\sqrt\{2\}L\_\{1\}q\(x\)\\right\),whenever‖y−x‖≤γ\\\|y\-x\\\|\\leq\\gammaandγ≤1/\(42L1\)\\gamma\\leq 1/\(4\\sqrt\{2\}L\_\{1\}\)\. Applying this withx=xkx=x^\{k\},y=xk\+1y=x^\{k\+1\}, and using Lemma[2](https://arxiv.org/html/2605.15314#Thmlemma2), we obtain
f\(xk\+1\)≤\\displaystyle f\(x^\{k\+1\}\)\\leqf\(xk\)−γ‖∇f\(xk\)‖\+2γ‖vk−∇f\(xk\)‖\+2L0γ2\+22L1γ2‖∇f\(xk\)‖\\displaystyle f\(x^\{k\}\)\-\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|\+2\\gamma\\\|v^\{k\}\-\\nabla f\(x^\{k\}\)\\\|\+\\sqrt\{2\}L\_\{0\}\\gamma^\{2\}\+2\\sqrt\{2\}L\_\{1\}\\gamma^\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|\+22L1γ2qk\.\\displaystyle\+2\\sqrt\{2\}L\_\{1\}\\gamma^\{2\}q\_\{k\}\.The stepsize condition implies
22L1γ2‖∇f\(xk\)‖≤γ2‖∇f\(xk\)‖\.2\\sqrt\{2\}L\_\{1\}\\gamma^\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\frac\{\\gamma\}\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|\.Subtractingfinff^\{\\inf\}, taking expectations, and moving this term to the left proves \([46](https://arxiv.org/html/2605.15314#A3.E46)\)\. ∎
### C\.3STORM decomposition and centered difference bound
###### Lemma 11\(STORM decomposition\)\.
Let
Ek:=vk−∇f\(xk\)\.E\_\{k\}:=v^\{k\}\-\\nabla f\(x^\{k\}\)\.Define
Zk\+1\\displaystyle Z\_\{k\+1\}:=\(∇f\(xk\+1;ξk\+1\)−∇f\(xk;ξk\+1\)\)−\(∇f\(xk\+1\)−∇f\(xk\)\),\\displaystyle=\\left\(\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\};\\xi^\{k\+1\}\)\\right\)\-\\left\(\\nabla f\(x^\{k\+1\}\)\-\\nabla f\(x^\{k\}\)\\right\),ζk\+1\\displaystyle\\zeta\_\{k\+1\}:=∇f\(xk;ξk\+1\)−∇f\(xk\)\.\\displaystyle=\\nabla f\(x^\{k\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\}\)\.Then
Ek\+1=\(1−η\)Ek\+Zk\+1\+ηζk\+1\.E\_\{k\+1\}=\(1\-\\eta\)E\_\{k\}\+Z\_\{k\+1\}\+\\eta\\zeta\_\{k\+1\}\.\(47\)Moreover,
𝔼\[Zk\+1∣ℱk\]=0,𝔼\[ζk\+1∣ℱk\]=0\.\\mathbb\{E\}\[Z\_\{k\+1\}\\mid\\mathcal\{F\}\_\{k\}\]=0,\\qquad\\mathbb\{E\}\[\\zeta\_\{k\+1\}\\mid\\mathcal\{F\}\_\{k\}\]=0\.
###### Proof\.
Subtract∇f\(xk\+1\)\\nabla f\(x^\{k\+1\}\)from the STORM update:
Ek\+1\\displaystyle E\_\{k\+1\}=∇f\(xk\+1;ξk\+1\)−∇f\(xk\+1\)\+\(1−η\)\(vk−∇f\(xk;ξk\+1\)\)\.\\displaystyle=\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\+1\}\)\+\(1\-\\eta\)\\left\(v^\{k\}\-\\nabla f\(x^\{k\};\\xi^\{k\+1\}\)\\right\)\.Add and subtract\(1−η\)∇f\(xk\)\(1\-\\eta\)\\nabla f\(x^\{k\}\)\. Then
Ek\+1\\displaystyle E\_\{k\+1\}=\(1−η\)\(vk−∇f\(xk\)\)\+\(∇f\(xk\+1;ξk\+1\)−∇f\(xk\+1\)\)\\displaystyle=\(1\-\\eta\)\\left\(v^\{k\}\-\\nabla f\(x^\{k\}\)\\right\)\+\\left\(\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\+1\}\)\\right\)−\(1−η\)\(∇f\(xk;ξk\+1\)−∇f\(xk\)\)\.\\displaystyle\\quad\-\(1\-\\eta\)\\left\(\\nabla f\(x^\{k\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\}\)\\right\)\.The last two lines equalZk\+1\+ηζk\+1Z\_\{k\+1\}\+\\eta\\zeta\_\{k\+1\}\. The conditional mean\-zero properties follow from unbiasedness and the freshness ofξk\+1\\xi^\{k\+1\}\. ∎
###### Lemma 12\(Centered Difference Bound\)\.
Suppose that for everyk=0,…,Kk=0,\\ldots,K,
\(𝔼\[‖Zk\+1‖2∣ℱk\]\)1/2≤γ\(a\+h‖∇f\(xk\)‖\),\\left\(\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq\\gamma\\left\(a\+h\\\|\\nabla f\(x^\{k\}\)\\\|\\right\),and
\(𝔼\[‖ζk\+1‖2∣ℱk\]\)1/2≤τ\.\\left\(\\mathbb\{E\}\[\\\|\\zeta\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq\\tau\.Then
Sb\(K\)≤b0η\+Taγη\+Tητ\+hγη∑k=0Kgk\.S\_\{b\}\(K\)\\leq\\frac\{b\_\{0\}\}\{\\eta\}\+T\\frac\{a\\gamma\}\{\\sqrt\{\\eta\}\}\+T\\sqrt\{\\eta\}\\,\\tau\+\\frac\{h\\gamma\}\{\\eta\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\.\(48\)
###### Proof\.
Unrolling \([47](https://arxiv.org/html/2605.15314#A3.E47)\), fork≥1k\\geq 1,
Ek=\(1−η\)kE0\+∑t=0k−1\(1−η\)k−1−tZt\+1\+η∑t=0k−1\(1−η\)k−1−tζt\+1\.E\_\{k\}=\(1\-\\eta\)^\{k\}E\_\{0\}\+\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}Z\_\{t\+1\}\+\\eta\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}\\zeta\_\{t\+1\}\.TheZZ\-sum is split into a gradient\-independent part and a gradient\-dependent part\. Let
A:=γa,Ht:=γh‖∇f\(xt\)‖\.A:=\\gamma a,\\qquad H\_\{t\}:=\\gamma h\\\|\\nabla f\(x^\{t\}\)\\\|\.IfA\+Ht\>0A\+H\_\{t\}\>0, define
λt:=AA\+Ht;\\lambda\_\{t\}:=\\frac\{A\}\{A\+H\_\{t\}\};otherwise setλt:=0\\lambda\_\{t\}:=0\. Then
Zt\+1\(0\):=λtZt\+1,Zt\+1\(1\):=\(1−λt\)Zt\+1Z\_\{t\+1\}^\{\(0\)\}:=\\lambda\_\{t\}Z\_\{t\+1\},\\qquad Z\_\{t\+1\}^\{\(1\)\}:=\(1\-\\lambda\_\{t\}\)Z\_\{t\+1\}are martingale differences and satisfy
\(𝔼\[‖Zt\+1\(0\)‖2∣ℱt\]\)1/2≤A,\(𝔼\[‖Zt\+1\(1\)‖2∣ℱt\]\)1/2≤Ht\.\\left\(\\mathbb\{E\}\[\\\|Z\_\{t\+1\}^\{\(0\)\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{t\}\]\\right\)^\{1/2\}\\leq A,\\qquad\\left\(\\mathbb\{E\}\[\\\|Z\_\{t\+1\}^\{\(1\)\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{t\}\]\\right\)^\{1/2\}\\leq H\_\{t\}\.By martingale orthogonality,
𝔼‖∑t=0k−1\(1−η\)k−1−tZt\+1\(0\)‖≤γaη\.\\mathbb\{E\}\\left\\\|\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}Z\_\{t\+1\}^\{\(0\)\}\\right\\\|\\leq\\frac\{\\gamma a\}\{\\sqrt\{\\eta\}\}\.For the gradient\-dependent part, triangle inequality gives
𝔼‖∑t=0k−1\(1−η\)k−1−tZt\+1\(1\)‖≤γh∑t=0k−1\(1−η\)k−1−tgt\.\\mathbb\{E\}\\left\\\|\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}Z\_\{t\+1\}^\{\(1\)\}\\right\\\|\\leq\\gamma h\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}g\_\{t\}\.Similarly, for the fresh\-noise martingale term,
𝔼‖η∑t=0k−1\(1−η\)k−1−tζt\+1‖≤ητ\.\\mathbb\{E\}\\left\\\|\\eta\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}\\zeta\_\{t\+1\}\\right\\\|\\leq\\sqrt\{\\eta\}\\,\\tau\.Thus,
bk≤\(1−η\)kb0\+γaη\+ητ\+γh∑t=0k−1\(1−η\)k−1−tgt\.b\_\{k\}\\leq\(1\-\\eta\)^\{k\}b\_\{0\}\+\\frac\{\\gamma a\}\{\\sqrt\{\\eta\}\}\+\\sqrt\{\\eta\}\\,\\tau\+\\gamma h\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}g\_\{t\}\.Summing overk=0,…,Kk=0,\\ldots,Kand using
∑k=0K\(1−η\)k≤1η,∑k=0K∑t=0k−1\(1−η\)k−1−tgt≤1η∑t=0Kgt,\\sum\_\{k=0\}^\{K\}\(1\-\\eta\)^\{k\}\\leq\\frac\{1\}\{\\eta\},\\qquad\\sum\_\{k=0\}^\{K\}\\sum\_\{t=0\}^\{k\-1\}\(1\-\\eta\)^\{k\-1\-t\}g\_\{t\}\\leq\\frac\{1\}\{\\eta\}\\sum\_\{t=0\}^\{K\}g\_\{t\},gives \([48](https://arxiv.org/html/2605.15314#A3.E48)\)\. ∎
### C\.4Estimator\-difference bounds
###### Lemma 13\(Single\-sample estimator bounds\)\.
Let
QT:=G\+BTγ\.Q\_\{T\}:=G\+BT\\gamma\.The following bounds hold\.
1. \(i\)Under Assumption[4](https://arxiv.org/html/2605.15314#Thmassumption4), \(𝔼\[‖Zk\+1‖2∣ℱk\]\)1/2≤Lγ,\(𝔼\[‖ζk\+1‖2∣ℱk\]\)1/2≤QT\.\\left\(\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq L\\gamma,\\qquad\\left\(\\mathbb\{E\}\[\\\|\\zeta\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq Q\_\{T\}\.
2. \(ii\)Under Assumption[6](https://arxiv.org/html/2605.15314#Thmassumption6)withα∈\(0,1\)\\alpha\\in\(0,1\), define p:=α1−α,K0:=22−α1−αL0,K1:=22−α1−αL1,K2:=\(5L1\)11−α,p:=\\frac\{\\alpha\}\{1\-\\alpha\},\\qquad K\_\{0\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{0\},\\qquad K\_\{1\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{1\},\\qquad K\_\{2\}:=\(5L\_\{1\}\)^\{\\frac\{1\}\{1\-\\alpha\}\},and Lα:=2α/2K1,cα:=\(1−α\)αα/\(1−α\)\.L\_\{\\alpha\}:=2^\{\\alpha/2\}K\_\{1\},\\qquad c\_\{\\alpha\}:=\(1\-\\alpha\)\\alpha^\{\\alpha/\(1\-\\alpha\)\}\.Then, for everyλ\>0\\lambda\>0, \(𝔼\[‖Zk\+1‖2∣ℱk\]\)1/2≤γ\(K0\+cαLαλ−p\+LαQTα\+K2γp\+Lαλ‖∇f\(xk\)‖\),\\displaystyle\\left\(\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq\\gamma\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda^\{\-p\}\+L\_\{\\alpha\}Q\_\{T\}^\{\\alpha\}\+K\_\{2\}\\gamma^\{p\}\+L\_\{\\alpha\}\\lambda\\\|\\nabla f\(x^\{k\}\)\\\|\\right\),and \(𝔼\[‖ζk\+1‖2∣ℱk\]\)1/2≤QT\.\\left\(\\mathbb\{E\}\[\\\|\\zeta\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq Q\_\{T\}\.
3. \(iii\)Under Assumption[6](https://arxiv.org/html/2605.15314#Thmassumption6)withα=1\\alpha=1, ifγ≤1/\(4L1\)\\gamma\\leq 1/\(4L\_\{1\}\), then \(𝔼\[‖Zk\+1‖2∣ℱk\]\)1/2≤2e3/4γ\(L0\+2L1QT\+2L1‖∇f\(xk\)‖\),\\displaystyle\\left\(\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq\\sqrt\{2e^\{3/4\}\}\\gamma\\left\(L\_\{0\}\+2L\_\{1\}Q\_\{T\}\+2L\_\{1\}\\\|\\nabla f\(x^\{k\}\)\\\|\\right\),and \(𝔼\[‖ζk\+1‖2∣ℱk\]\)1/2≤QT\.\\left\(\\mathbb\{E\}\[\\\|\\zeta\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq Q\_\{T\}\.
###### Proof\.
For item \(i\), conditioning onℱk\\mathcal\{F\}\_\{k\}, mean\-square smoothness gives
𝔼\[‖Zk\+1‖2∣ℱk\]≤𝔼ξk\+1‖∇f\(xk\+1;ξk\+1\)−∇f\(xk;ξk\+1\)‖2≤L2γ2\.\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\leq\\mathbb\{E\}\_\{\\xi^\{k\+1\}\}\\\|\\nabla f\(x^\{k\+1\};\\xi^\{k\+1\}\)\-\\nabla f\(x^\{k\};\\xi^\{k\+1\}\)\\\|^\{2\}\\leq L^\{2\}\\gamma^\{2\}\.The fresh\-noise bound follows from Assumption[2](https://arxiv.org/html/2605.15314#Thmassumption2)and Lemma[6](https://arxiv.org/html/2605.15314#Thmlemma6)\.
For item \(ii\), conditioning onℱk\\mathcal\{F\}\_\{k\}and using the expected generalized smoothness reduction,
\(𝔼\[‖Zk\+1‖2∣ℱk\]\)1/2≤γ\(K0\+K1𝔼ξ‖∇f\(xk;ξ\)‖α\+K2γp\)\.\\displaystyle\\left\(\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq\\gamma\\left\(K\_\{0\}\+K\_\{1\}\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x^\{k\};\\xi\)\\\|^\{\\alpha\}\+K\_\{2\}\\gamma^\{p\}\\right\)\.By Lemma[8](https://arxiv.org/html/2605.15314#Thmlemma8),
K1𝔼ξ‖∇f\(xk;ξ\)‖α≤Lα\(‖∇f\(xk\)‖α\+qkα\)\.K\_\{1\}\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x^\{k\};\\xi\)\\\|^\{\\alpha\}\\leq L\_\{\\alpha\}\\left\(\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\+q\_\{k\}^\{\\alpha\}\\right\)\.Usingqk≤QTq\_\{k\}\\leq Q\_\{T\}and
‖∇f\(xk\)‖α≤λ‖∇f\(xk\)‖\+cαλ−p,\\\|\\nabla f\(x^\{k\}\)\\\|^\{\\alpha\}\\leq\\lambda\\\|\\nabla f\(x^\{k\}\)\\\|\+c\_\{\\alpha\}\\lambda^\{\-p\},gives the displayedZZ\-bound\. The bound forζk\+1\\zeta\_\{k\+1\}follows from BG\-0 andqk≤QTq\_\{k\}\\leq Q\_\{T\}\.
Item \(iii\) is the correspondingα=1\\alpha=1reduction\. It gives
\(𝔼\[‖Zk\+1‖2∣ℱk\]\)1/2≤2e3/4γ\(L0\+2L1qk\+2L1‖∇f\(xk\)‖\),\\displaystyle\\left\(\\mathbb\{E\}\[\\\|Z\_\{k\+1\}\\\|^\{2\}\\mid\\mathcal\{F\}\_\{k\}\]\\right\)^\{1/2\}\\leq\\sqrt\{2e^\{3/4\}\}\\gamma\\left\(L\_\{0\}\+2L\_\{1\}q\_\{k\}\+2L\_\{1\}\\\|\\nabla f\(x^\{k\}\)\\\|\\right\),andqk≤QTq\_\{k\}\\leq Q\_\{T\}completes the proof\. ∎
### C\.5Single\-sample master bounds
###### Theorem 5\(Single\-sample NSTORM master bounds\)\.
Consider the single\-sample𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}update in \([3](https://arxiv.org/html/2605.15314#S5.E3)\)\. Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[2](https://arxiv.org/html/2605.15314#Thmassumption2)hold\. Then the following statements hold\.
1. \(i\)Suppose Assumption[4](https://arxiv.org/html/2605.15314#Thmassumption4)holds\. Then 𝔼‖∇f\(x^\)‖≤\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leqΔγT\+2b0ηT\+2Lγη\+2η\(G\+BTγ\)\+L2γ\.\\displaystyle\\frac\{\\Delta\}\{\\gamma T\}\+\\frac\{2b\_\{0\}\}\{\\eta T\}\+\\frac\{2L\\gamma\}\{\\sqrt\{\\eta\}\}\+2\\sqrt\{\\eta\}\(G\+BT\\gamma\)\+\\frac\{L\}\{2\}\\gamma\.\(49\)
2. \(ii\)Suppose Assumption[6](https://arxiv.org/html/2605.15314#Thmassumption6)holds withα∈\(0,1\)\\alpha\\in\(0,1\)\. Define p:=α1−α,K0:=22−α1−αL0,K1:=22−α1−αL1,K2:=\(5L1\)11−α,p:=\\frac\{\\alpha\}\{1\-\\alpha\},\\qquad K\_\{0\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{0\},\\qquad K\_\{1\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{1\},\\qquad K\_\{2\}:=\(5L\_\{1\}\)^\{\\frac\{1\}\{1\-\\alpha\}\},and Lα:=2α/2K1,cα:=\(1−α\)αα/\(1−α\)\.L\_\{\\alpha\}:=2^\{\\alpha/2\}K\_\{1\},\\qquad c\_\{\\alpha\}:=\(1\-\\alpha\)\\alpha^\{\\alpha/\(1\-\\alpha\)\}\.Let QT:=G\+BTγ\.Q\_\{T\}:=G\+BT\\gamma\.If Lαλγ≤1,8Lαλγη≤1,L\_\{\\alpha\}\\lambda\\gamma\\leq 1,\\qquad 8L\_\{\\alpha\}\\lambda\\frac\{\\gamma\}\{\\eta\}\\leq 1,then 𝔼‖∇f\(x^\)‖≤\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq4ΔγT\+8b0ηT\+8γη\(K0\+cαLαλ−p\+LαQTα\+K2γp\)\+8ηQT\\displaystyle\\frac\{4\\Delta\}\{\\gamma T\}\+\\frac\{8b\_\{0\}\}\{\\eta T\}\+\\frac\{8\\gamma\}\{\\sqrt\{\\eta\}\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda^\{\-p\}\+L\_\{\\alpha\}Q\_\{T\}^\{\\alpha\}\+K\_\{2\}\\gamma^\{p\}\\right\)\+8\\sqrt\{\\eta\}\\,Q\_\{T\}\(50\)\+2γ\(K0\+cαLαλ−p\+2K2γp\)\+2Lαγ1T∑k=0K𝔼\[qkα\]\.\\displaystyle\+2\\gamma\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda^\{\-p\}\+2K\_\{2\}\\gamma^\{p\}\\right\)\+2L\_\{\\alpha\}\\gamma\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}^\{\\alpha\}\]\.
3. \(iii\)Suppose Assumption[6](https://arxiv.org/html/2605.15314#Thmassumption6)holds withα=1\\alpha=1\. If γ≤142L1,162e3/4L1γη≤1,\\gamma\\leq\\frac\{1\}\{4\\sqrt\{2\}L\_\{1\}\},\\qquad 16\\sqrt\{2e^\{3/4\}\}L\_\{1\}\\frac\{\\gamma\}\{\\eta\}\\leq 1,then 𝔼‖∇f\(x^\)‖≤\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq4ΔγT\+8b0ηT\+8γη2e3/4\(L0\+2L1\(G\+BTγ\)\)\+8η\(G\+BTγ\)\\displaystyle\\frac\{4\\Delta\}\{\\gamma T\}\+\\frac\{8b\_\{0\}\}\{\\eta T\}\+\\frac\{8\\gamma\}\{\\sqrt\{\\eta\}\}\\sqrt\{2e^\{3/4\}\}\\left\(L\_\{0\}\+2L\_\{1\}\(G\+BT\\gamma\)\\right\)\+8\\sqrt\{\\eta\}\(G\+BT\\gamma\)\(51\)\+42L0γ\+82L1γ1T∑k=0K𝔼\[qk\]\.\\displaystyle\+4\\sqrt\{2\}L\_\{0\}\\gamma\+8\\sqrt\{2\}L\_\{1\}\\gamma\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}\]\.
###### Proof\.
For item \(i\), summing \([44](https://arxiv.org/html/2605.15314#A3.E44)\) gives
γ∑k=0Kgk≤Δ\+2γSb\(K\)\+L2γ2T\.\\gamma\\sum\_\{k=0\}^\{K\}g\_\{k\}\\leq\\Delta\+2\\gamma S\_\{b\}\(K\)\+\\frac\{L\}\{2\}\\gamma^\{2\}T\.Apply Lemma[12](https://arxiv.org/html/2605.15314#Thmlemma12)with
a=L,h=0,τ=G\+BTγ\.a=L,\\qquad h=0,\\qquad\\tau=G\+BT\\gamma\.Then
Sb\(K\)≤b0η\+TLγη\+Tη\(G\+BTγ\)\.S\_\{b\}\(K\)\\leq\\frac\{b\_\{0\}\}\{\\eta\}\+T\\frac\{L\\gamma\}\{\\sqrt\{\\eta\}\}\+T\\sqrt\{\\eta\}\(G\+BT\\gamma\)\.Substituting this into the summed descent inequality and dividing byγT\\gamma Tgives \([49](https://arxiv.org/html/2605.15314#A3.E49)\)\.
For item \(ii\), summing \([45](https://arxiv.org/html/2605.15314#A3.E45)\) gives
γ2∑k=0Kgk≤\\displaystyle\\frac\{\\gamma\}\{2\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\\leqΔ\+2γSb\(K\)\+γ2T2\(K0\+cαLαλ−p\+2K2γp\)\+γ22Lα∑k=0K𝔼\[qkα\]\.\\displaystyle\\Delta\+2\\gamma S\_\{b\}\(K\)\+\\frac\{\\gamma^\{2\}T\}\{2\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda^\{\-p\}\+2K\_\{2\}\\gamma^\{p\}\\right\)\+\\frac\{\\gamma^\{2\}\}\{2\}L\_\{\\alpha\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}^\{\\alpha\}\]\.Apply Lemma[12](https://arxiv.org/html/2605.15314#Thmlemma12)with
a=K0\+cαLαλ−p\+LαQTα\+K2γp,h=Lαλ,τ=QT\.a=K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda^\{\-p\}\+L\_\{\\alpha\}Q\_\{T\}^\{\\alpha\}\+K\_\{2\}\\gamma^\{p\},\\qquad h=L\_\{\\alpha\}\\lambda,\\qquad\\tau=Q\_\{T\}\.Then
Sb\(K\)≤b0η\+Taγη\+TηQT\+Lαλγη∑k=0Kgk\.S\_\{b\}\(K\)\\leq\\frac\{b\_\{0\}\}\{\\eta\}\+T\\frac\{a\\gamma\}\{\\sqrt\{\\eta\}\}\+T\\sqrt\{\\eta\}Q\_\{T\}\+\\frac\{L\_\{\\alpha\}\\lambda\\gamma\}\{\\eta\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\.Substituting this bound into the descent inequality gives a feedback term
2γ⋅Lαλγη∑k=0Kgk\.2\\gamma\\cdot\\frac\{L\_\{\\alpha\}\\lambda\\gamma\}\{\\eta\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\.The condition
8Lαλγη≤18L\_\{\\alpha\}\\lambda\\frac\{\\gamma\}\{\\eta\}\\leq 1implies that this feedback can be absorbed into the left\-hand side\. After absorption and multiplication by44, division byγT\\gamma Tgives \([50](https://arxiv.org/html/2605.15314#A3.E50)\)\.
For item \(iii\), sum \([46](https://arxiv.org/html/2605.15314#A3.E46)\):
γ2∑k=0Kgk≤Δ\+2γSb\(K\)\+2L0γ2T\+22L1γ2∑k=0K𝔼\[qk\]\.\\frac\{\\gamma\}\{2\}\\sum\_\{k=0\}^\{K\}g\_\{k\}\\leq\\Delta\+2\\gamma S\_\{b\}\(K\)\+\\sqrt\{2\}L\_\{0\}\\gamma^\{2\}T\+2\\sqrt\{2\}L\_\{1\}\\gamma^\{2\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}\]\.Apply Lemma[12](https://arxiv.org/html/2605.15314#Thmlemma12)with
a=2e3/4\(L0\+2L1\(G\+BTγ\)\),h=22e3/4L1,τ=G\+BTγ\.a=\\sqrt\{2e^\{3/4\}\}\\left\(L\_\{0\}\+2L\_\{1\}\(G\+BT\\gamma\)\\right\),\\qquad h=2\\sqrt\{2e^\{3/4\}\}L\_\{1\},\\qquad\\tau=G\+BT\\gamma\.The condition
162e3/4L1γη≤116\\sqrt\{2e^\{3/4\}\}L\_\{1\}\\frac\{\\gamma\}\{\\eta\}\\leq 1absorbs the resulting feedback term\. Dividing byγT/4\\gamma T/4gives \([51](https://arxiv.org/html/2605.15314#A3.E51)\)\. ∎
### C\.6Proof of Theorem[2](https://arxiv.org/html/2605.15314#Thmtheorem2)
###### Corollary 5\.1\(Mean\-square smoothness with sharp initialization\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2), and[4](https://arxiv.org/html/2605.15314#Thmassumption4)hold\. Initialize
v0=1Ninit∑i=1Ninit∇f\(x0;ξiinit\),Ninit:=max\{1,⌈G2T1/2⌉\}\.v^\{0\}=\\frac\{1\}\{N\_\{\\rm init\}\}\\sum\_\{i=1\}^\{N\_\{\\rm init\}\}\\nabla f\(x^\{0\};\\xi\_\{i\}^\{\\rm init\}\),\\qquad N\_\{\\rm init\}:=\\max\\left\\\{1,\\left\\lceil G^\{2\}T^\{1/2\}\\right\\rceil\\right\\\}\.Choose
η=T−1,γ=γ0T−3/4\.\\eta=T^\{\-1\},\\qquad\\gamma=\\gamma\_\{0\}T^\{\-3/4\}\.Then
𝔼‖∇f\(x^\)‖≤\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\(Δγ0\+2\+2Lγ0\+2Bγ0\)T−1/4\+2GT−1/2\+Lγ02T−3/4\.\\displaystyle\\left\(\\frac\{\\Delta\}\{\\gamma\_\{0\}\}\+2\+2L\\gamma\_\{0\}\+2B\\gamma\_\{0\}\\right\)T^\{\-1/4\}\+2GT^\{\-1/2\}\+\\frac\{L\\gamma\_\{0\}\}\{2\}T^\{\-3/4\}\.Consequently,
𝔼‖∇f\(x^\)‖=𝒪\(T−1/4\),\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/4\}\),and the SFO complexity is𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\.
###### Proof\.
By the initialization choice,
b0=𝔼‖v0−∇f\(x0\)‖≤GNinit≤T−1/4\.b\_\{0\}=\\mathbb\{E\}\\\|v^\{0\}\-\\nabla f\(x^\{0\}\)\\\|\\leq\\frac\{G\}\{\\sqrt\{N\_\{\\rm init\}\}\}\\leq T^\{\-1/4\}\.Apply \([49](https://arxiv.org/html/2605.15314#A3.E49)\)\. Under the schedule,
ΔγT=Δγ0T−1/4,2b0ηT=2b0≤2T−1/4\.\\frac\{\\Delta\}\{\\gamma T\}=\\frac\{\\Delta\}\{\\gamma\_\{0\}\}T^\{\-1/4\},\\qquad\\frac\{2b\_\{0\}\}\{\\eta T\}=2b\_\{0\}\\leq 2T^\{\-1/4\}\.Also,
2Lγη=2Lγ0T−1/4,L2γ=Lγ02T−3/4,\\frac\{2L\\gamma\}\{\\sqrt\{\\eta\}\}=2L\\gamma\_\{0\}T^\{\-1/4\},\\qquad\\frac\{L\}\{2\}\\gamma=\\frac\{L\\gamma\_\{0\}\}\{2\}T^\{\-3/4\},and
2η\(G\+BTγ\)=2GT−1/2\+2Bγ0T−1/4\.2\\sqrt\{\\eta\}\(G\+BT\\gamma\)=2GT^\{\-1/2\}\+2B\\gamma\_\{0\}T^\{\-1/4\}\.Combining these estimates proves the displayed bound\. SinceNinit=𝒪\(T1/2\)N\_\{\\rm init\}=\\mathcal\{O\}\(T^\{1/2\}\)and each transition uses two SFO calls, the total SFO cost isNinit\+2T=𝒪\(T\)N\_\{\\rm init\}\+2T=\\mathcal\{O\}\(T\)\. SettingT=𝒪\(ε−4\)T=\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)gives𝒪\(ε−4\)\\mathcal\{O\}\(\\varepsilon^\{\-4\}\)\. ∎
###### Corollary 5\.2\(Expectedα\\alpha\-symmetric generalized smoothness\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2), and[6](https://arxiv.org/html/2605.15314#Thmassumption6)hold withα∈\(0,1\)\\alpha\\in\(0,1\)\. Define
p:=α1−α,K0:=22−α1−αL0,K1:=22−α1−αL1,K2:=\(5L1\)11−α,p:=\\frac\{\\alpha\}\{1\-\\alpha\},\\qquad K\_\{0\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{0\},\\qquad K\_\{1\}:=2^\{\\frac\{2\-\\alpha\}\{1\-\\alpha\}\}L\_\{1\},\\qquad K\_\{2\}:=\(5L\_\{1\}\)^\{\\frac\{1\}\{1\-\\alpha\}\},and let
T:=K\+1,θ:=14\+α\.T:=K\+1,\\qquad\\theta:=\\frac\{1\}\{4\+\\alpha\}\.Initialize
v0=1Ninit∑i=1Ninit∇f\(x0;ξiinit\),Ninit:=max\{1,⌈G2T2\(1−α\)θ⌉\}\.v^\{0\}=\\frac\{1\}\{N\_\{\\rm init\}\}\\sum\_\{i=1\}^\{N\_\{\\rm init\}\}\\nabla f\(x^\{0\};\\xi\_\{i\}^\{\\rm init\}\),\\qquad N\_\{\\rm init\}:=\\max\\left\\\{1,\\left\\lceil G^\{2\}T^\{2\(1\-\\alpha\)\\theta\}\\right\\rceil\\right\\\}\.Choose
γ=γ0T−\(3\+α\)θ,η=η0T−4θ,λ=λ0T−\(1−α\)θ,\\gamma=\\gamma\_\{0\}T^\{\-\(3\+\\alpha\)\\theta\},\\qquad\\eta=\\eta\_\{0\}T^\{\-4\\theta\},\\qquad\\lambda=\\lambda\_\{0\}T^\{\-\(1\-\\alpha\)\\theta\},whereη0∈\(0,1\]\\eta\_\{0\}\\in\(0,1\],γ0\>0\\gamma\_\{0\}\>0, andλ0\>0\\lambda\_\{0\}\>0\. Suppose
2α/2K1λ0γ0≤1,2\(6\+α\)/2K1λ0γ0η0≤1\.2^\{\\alpha/2\}K\_\{1\}\\lambda\_\{0\}\\gamma\_\{0\}\\leq 1,\\qquad 2^\{\(6\+\\alpha\)/2\}K\_\{1\}\\lambda\_\{0\}\\frac\{\\gamma\_\{0\}\}\{\\eta\_\{0\}\}\\leq 1\.Then the bound stated in Theorem[2](https://arxiv.org/html/2605.15314#Thmtheorem2)\(ii\) holds\. In particular,
𝔼‖∇f\(x^\)‖=𝒪\(T−1/\(4\+α\)\),\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/\(4\+\\alpha\)\}\),and the SFO complexity is𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)\.
###### Proof\.
By the initialization choice,
b0≤GNinit≤T−\(1−α\)θ\.b\_\{0\}\\leq\\frac\{G\}\{\\sqrt\{N\_\{\\rm init\}\}\}\\leq T^\{\-\(1\-\\alpha\)\\theta\}\.Let
Lα:=2α/2K1,cα:=\(1−α\)αα/\(1−α\)\.L\_\{\\alpha\}:=2^\{\\alpha/2\}K\_\{1\},\\qquad c\_\{\\alpha\}:=\(1\-\\alpha\)\\alpha^\{\\alpha/\(1\-\\alpha\)\}\.The two conditions in the corollary are exactly
Lαλγ≤1,8Lαλγη≤1\.L\_\{\\alpha\}\\lambda\\gamma\\leq 1,\\qquad 8L\_\{\\alpha\}\\lambda\\frac\{\\gamma\}\{\\eta\}\\leq 1\.Therefore, \([50](https://arxiv.org/html/2605.15314#A3.E50)\) applies\.
We now substitute the parameter choices\. First,
4ΔγT=4Δγ0T−θ,8b0ηT≤8η0T−θ\.\\frac\{4\\Delta\}\{\\gamma T\}=\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}T^\{\-\\theta\},\\qquad\\frac\{8b\_\{0\}\}\{\\eta T\}\\leq\\frac\{8\}\{\\eta\_\{0\}\}T^\{\-\\theta\}\.Also,
QT=G\+BTγ=G\+Bγ0Tθ\.Q\_\{T\}=G\+BT\\gamma=G\+B\\gamma\_\{0\}T^\{\\theta\}\.Using\(a\+b\)α≤aα\+bα\(a\+b\)^\{\\alpha\}\\leq a^\{\\alpha\}\+b^\{\\alpha\},
QTα≤Gα\+Bαγ0αTαθ\.Q\_\{T\}^\{\\alpha\}\\leq G^\{\\alpha\}\+B^\{\\alpha\}\\gamma\_\{0\}^\{\\alpha\}T^\{\\alpha\\theta\}\.Substituting these estimates into \([50](https://arxiv.org/html/2605.15314#A3.E50)\) gives exactly
𝔼∥∇f\(x^\)∥≤\[\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\Bigg\[4Δγ0\+8η0\+8\(1−α\)αα1−α2α/2K1γ0λ0−α1−αη0\\displaystyle\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\+\\frac\{8\}\{\\eta\_\{0\}\}\+\\frac\{8\(1\-\\alpha\)\\alpha^\{\\frac\{\\alpha\}\{1\-\\alpha\}\}2^\{\\alpha/2\}K\_\{1\}\\gamma\_\{0\}\\lambda\_\{0\}^\{\-\\frac\{\\alpha\}\{1\-\\alpha\}\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\+8⋅2α/2K1Bαγ01\+αη0\+8η0Bγ0\]T−θ\\displaystyle\+\\frac\{8\\cdot 2^\{\\alpha/2\}K\_\{1\}B^\{\\alpha\}\\gamma\_\{0\}^\{1\+\\alpha\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\+8\\sqrt\{\\eta\_\{0\}\}B\\gamma\_\{0\}\\Bigg\]T^\{\-\\theta\}\+\[8K0γ0η0\+8⋅2α/2K1Gαγ0η0\]T−\(1\+α\)θ\+8η0GT−2θ\\displaystyle\+\\left\[\\frac\{8K\_\{0\}\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\+\\frac\{8\\cdot 2^\{\\alpha/2\}K\_\{1\}G^\{\\alpha\}\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\\right\]T^\{\-\(1\+\\alpha\)\\theta\}\+8\\sqrt\{\\eta\_\{0\}\}GT^\{\-2\\theta\}\+8K2γ011−αη0T−1\+3α\(1−α\)\(4\+α\)\+2K0γ0T−\(3\+α\)θ\\displaystyle\+\\frac\{8K\_\{2\}\\gamma\_\{0\}^\{\\frac\{1\}\{1\-\\alpha\}\}\}\{\\sqrt\{\\eta\_\{0\}\}\}T^\{\-\\frac\{1\+3\\alpha\}\{\(1\-\\alpha\)\(4\+\\alpha\)\}\}\+2K\_\{0\}\\gamma\_\{0\}T^\{\-\(3\+\\alpha\)\\theta\}\+2\(1−α\)αα1−α2α/2K1γ0λ0−α1−αT−3θ\\displaystyle\+2\(1\-\\alpha\)\\alpha^\{\\frac\{\\alpha\}\{1\-\\alpha\}\}2^\{\\alpha/2\}K\_\{1\}\\gamma\_\{0\}\\lambda\_\{0\}^\{\-\\frac\{\\alpha\}\{1\-\\alpha\}\}T^\{\-3\\theta\}\+4K2γ011−αT−3\+α\(1−α\)\(4\+α\)\+21\+α/2K1Gαγ0T−\(3\+α\)θ\\displaystyle\+4K\_\{2\}\\gamma\_\{0\}^\{\\frac\{1\}\{1\-\\alpha\}\}T^\{\-\\frac\{3\+\\alpha\}\{\(1\-\\alpha\)\(4\+\\alpha\)\}\}\+2^\{1\+\\alpha/2\}K\_\{1\}G^\{\\alpha\}\\gamma\_\{0\}T^\{\-\(3\+\\alpha\)\\theta\}\+21\+α/2K1Bαγ01\+αT−3θ\.\\displaystyle\+2^\{1\+\\alpha/2\}K\_\{1\}B^\{\\alpha\}\\gamma\_\{0\}^\{1\+\\alpha\}T^\{\-3\\theta\}\.This is the displayed bound in Theorem[2](https://arxiv.org/html/2605.15314#Thmtheorem2)\(ii\), withθ=1/\(4\+α\)\\theta=1/\(4\+\\alpha\)\.
Every exponent is at leastθ\\theta, so
𝔼‖∇f\(x^\)‖=𝒪\(T−θ\)=𝒪\(T−1/\(4\+α\)\)\.\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-\\theta\}\)=\\mathcal\{O\}\(T^\{\-1/\(4\+\\alpha\)\}\)\.The initialization cost satisfies
Ninit=𝒪\(T2\(1−α\)/\(4\+α\)\)=o\(T\),N\_\{\\rm init\}=\\mathcal\{O\}\(T^\{2\(1\-\\alpha\)/\(4\+\\alpha\)\}\)=o\(T\),and each transition uses two SFO calls\. Hence the total SFO cost is
Ninit\+2T=𝒪\(T\)\.N\_\{\\rm init\}\+2T=\\mathcal\{O\}\(T\)\.TakingT=𝒪\(ε−\(4\+α\)\)T=\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)gives𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)\. ∎
###### Corollary 5\.3\(Expected11\-symmetric generalized smoothness\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2), and[6](https://arxiv.org/html/2605.15314#Thmassumption6)hold withα=1\\alpha=1\. Use one\-sample initialization, so that
b0:=𝔼‖v0−∇f\(x0\)‖≤G\.b\_\{0\}:=\\mathbb\{E\}\\\|v^\{0\}\-\\nabla f\(x^\{0\}\)\\\|\\leq G\.Choose
η=T−4/5,γ=γ0T−4/5,0<γ0≤1162e3/4L1\.\\eta=T^\{\-4/5\},\\qquad\\gamma=\\gamma\_\{0\}T^\{\-4/5\},\\qquad 0<\\gamma\_\{0\}\\leq\\frac\{1\}\{16\\sqrt\{2e^\{3/4\}\}L\_\{1\}\}\.Then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\(4Δγ0\+8b0\+8Bγ0\+162e3/4L1Bγ02\)T−1/5\\displaystyle\\leq\\left\(\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\+8b\_\{0\}\+8B\\gamma\_\{0\}\+16\\sqrt\{2e^\{3/4\}\}L\_\{1\}B\\gamma\_\{0\}^\{2\}\\right\)T^\{\-1/5\}\+\(82e3/4γ0\(L0\+2L1G\)\+8G\)T−2/5\\displaystyle\\quad\+\\left\(8\\sqrt\{2e^\{3/4\}\}\\gamma\_\{0\}\(L\_\{0\}\+2L\_\{1\}G\)\+8G\\right\)T^\{\-2/5\}\+82L1Bγ02T−3/5\+\(42L0γ0\+82L1Gγ0\)T−4/5\.\\displaystyle\\quad\+8\\sqrt\{2\}L\_\{1\}B\\gamma\_\{0\}^\{2\}T^\{\-3/5\}\+\\left\(4\\sqrt\{2\}L\_\{0\}\\gamma\_\{0\}\+8\\sqrt\{2\}L\_\{1\}G\\gamma\_\{0\}\\right\)T^\{\-4/5\}\.Consequently,
𝔼‖∇f\(x^\)‖=𝒪\(T−1/5\),\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/5\}\),and the SFO complexity is𝒪\(ε−5\)\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)\.
###### Proof\.
The stepsize condition implies
γ≤142L1,162e3/4L1γη=162e3/4L1γ0≤1\.\\gamma\\leq\\frac\{1\}\{4\\sqrt\{2\}L\_\{1\}\},\\qquad 16\\sqrt\{2e^\{3/4\}\}L\_\{1\}\\frac\{\\gamma\}\{\\eta\}=16\\sqrt\{2e^\{3/4\}\}L\_\{1\}\\gamma\_\{0\}\\leq 1\.Thus \([51](https://arxiv.org/html/2605.15314#A3.E51)\) applies\. Under the schedule,
4ΔγT=4Δγ0T−1/5,8b0ηT=8b0T−1/5\.\\frac\{4\\Delta\}\{\\gamma T\}=\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}T^\{\-1/5\},\\qquad\\frac\{8b\_\{0\}\}\{\\eta T\}=8b\_\{0\}T^\{\-1/5\}\.Moreover,
QT=G\+BTγ=G\+Bγ0T1/5\.Q\_\{T\}=G\+BT\\gamma=G\+B\\gamma\_\{0\}T^\{1/5\}\.Therefore,
8γη2e3/4\(L0\+2L1QT\)\\displaystyle\\frac\{8\\gamma\}\{\\sqrt\{\\eta\}\}\\sqrt\{2e^\{3/4\}\}\(L\_\{0\}\+2L\_\{1\}Q\_\{T\}\)=82e3/4γ0\(L0\+2L1G\)T−2/5\\displaystyle=8\\sqrt\{2e^\{3/4\}\}\\gamma\_\{0\}\(L\_\{0\}\+2L\_\{1\}G\)T^\{\-2/5\}\+162e3/4L1Bγ02T−1/5,\\displaystyle\\quad\+6\\sqrt\{2e^\{3/4\}\}L\_\{1\}B\\gamma\_\{0\}^\{2\}T^\{\-1/5\},and
8ηQT=8GT−2/5\+8Bγ0T−1/5\.8\\sqrt\{\\eta\}Q\_\{T\}=8GT^\{\-2/5\}\+8B\\gamma\_\{0\}T^\{\-1/5\}\.Finally, by Lemma[6](https://arxiv.org/html/2605.15314#Thmlemma6),
1T∑k=0K𝔼\[qk\]≤G\+BTγ=G\+Bγ0T1/5\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}\]\\leq G\+BT\\gamma=G\+B\\gamma\_\{0\}T^\{1/5\}\.Hence
82L1γ1T∑k=0K𝔼\[qk\]\\displaystyle 8\\sqrt\{2\}L\_\{1\}\\gamma\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}\]≤82L1Gγ0T−4/5\+82L1Bγ02T−3/5\.\\displaystyle\\leq 8\\sqrt\{2\}L\_\{1\}G\\gamma\_\{0\}T^\{\-4/5\}\+8\\sqrt\{2\}L\_\{1\}B\\gamma\_\{0\}^\{2\}T^\{\-3/5\}\.Together with
42L0γ=42L0γ0T−4/5,4\\sqrt\{2\}L\_\{0\}\\gamma=4\\sqrt\{2\}L\_\{0\}\\gamma\_\{0\}T^\{\-4/5\},this proves the displayed bound\. The method uses one SFO call for initialization and two SFO calls per transition, so the total cost is1\+2T=𝒪\(T\)1\+2T=\\mathcal\{O\}\(T\)\. TakingT=𝒪\(ε−5\)T=\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)gives𝒪\(ε−5\)\\mathcal\{O\}\(\\varepsilon^\{\-5\}\)\. ∎
### C\.7Recovery of bounded variance rates \(B=0B=0\)
WhenB=0B=0, the𝖡𝖦\\mathsf\{BG\}\-0 condition reduces to bounded variance𝔼ξ‖∇f\(x;ξ\)−∇f\(x\)‖2≤G2\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\(x;\\xi\)\-\\nabla f\(x\)\\\|^\{2\}\\leq G^\{2\}, and the local noise scale simplifies toqk=Gq\_\{k\}=Gfor everykk\. In particular, the trajectory\-dependent termsBTγBT\\gammavanish from the master bounds\. Under the bounded variance scheduleη=η0T−2/3\\eta=\\eta\_\{0\}T^\{\-2/3\},γ=γ0T−2/3\\gamma=\\gamma\_\{0\}T^\{\-2/3\},𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}recovers the standard variance\-reduced rate𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)for all three smoothness regimes, matching the SPIDER rate ofFanget al\.\[[12](https://arxiv.org/html/2605.15314#bib.bib45)\], Chenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]\.
###### Corollary 5\.4\(Mean\-square smoothness under bounded variance\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2)\(withB=0B=0\), and[4](https://arxiv.org/html/2605.15314#Thmassumption4)hold\. Use one\-sample initializationv0=∇f\(x0;ξ0\)v^\{0\}=\\nabla f\(x^\{0\};\\xi^\{0\}\), so thatb0≤Gb\_\{0\}\\leq G\. Choose
η=η0T−2/3,γ=γ0T−2/3,\\eta=\\eta\_\{0\}\\,T^\{\-2/3\},\\qquad\\gamma=\\gamma\_\{0\}\\,T^\{\-2/3\},whereη0∈\(0,1\]\\eta\_\{0\}\\in\(0,1\]andγ0\>0\\gamma\_\{0\}\>0\. Then
𝔼‖∇f\(x^\)‖≤\(Δγ0\+2b0η0\+2Lγ0η0\+2η0G\)T−1/3\+Lγ02T−2/3\.\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\left\(\\frac\{\\Delta\}\{\\gamma\_\{0\}\}\+\\frac\{2b\_\{0\}\}\{\\eta\_\{0\}\}\+\\frac\{2L\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\+2\\sqrt\{\\eta\_\{0\}\}\\,G\\right\)T^\{\-1/3\}\+\\frac\{L\\gamma\_\{0\}\}\{2\}\\,T^\{\-2/3\}\.\(52\)Consequently,𝔼‖∇f\(x^\)‖=𝒪\(T−1/3\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/3\}\), and the SFO complexity is𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)\.
###### Proof\.
SinceB=0B=0, we haveqk=Gq\_\{k\}=Gfor everykkandQT=GQ\_\{T\}=G\. Apply \([49](https://arxiv.org/html/2605.15314#A3.E49)\) from Theorem[5](https://arxiv.org/html/2605.15314#Thmtheorem5)\(i\)\. Under the scheduleη=η0T−2/3\\eta=\\eta\_\{0\}T^\{\-2/3\}andγ=γ0T−2/3\\gamma=\\gamma\_\{0\}T^\{\-2/3\},
ΔγT=Δγ0T−1/3,2b0ηT=2b0η0T−1/3,\\frac\{\\Delta\}\{\\gamma T\}=\\frac\{\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/3\},\\qquad\\frac\{2b\_\{0\}\}\{\\eta T\}=\\frac\{2b\_\{0\}\}\{\\eta\_\{0\}\}\\,T^\{\-1/3\},2Lγη=2Lγ0η0T−1/3,L2γ=Lγ02T−2/3,\\frac\{2L\\gamma\}\{\\sqrt\{\\eta\}\}=\\frac\{2L\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\\,T^\{\-1/3\},\\qquad\\frac\{L\}\{2\}\\gamma=\\frac\{L\\gamma\_\{0\}\}\{2\}\\,T^\{\-2/3\},and, sinceB=0B=0,
2η\(G\+BTγ\)=2η0GT−1/3\.2\\sqrt\{\\eta\}\(G\+BT\\gamma\)=2\\sqrt\{\\eta\_\{0\}\}\\,G\\,T^\{\-1/3\}\.Substitution gives \([52](https://arxiv.org/html/2605.15314#A3.E52)\)\. The method uses one SFO call for initialization and two per transition, givingMK=1\+2T=𝒪\(ε−3\)M\_\{K\}=1\+2T=\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)\. ∎
###### Corollary 5\.5\(Expectedα\\alpha\-symmetric generalized smoothness under bounded variance\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2)\(withB=0B=0\), and[6](https://arxiv.org/html/2605.15314#Thmassumption6)hold withα∈\(0,1\)\\alpha\\in\(0,1\)\. DefineLα:=2α/2K1L\_\{\\alpha\}:=2^\{\\alpha/2\}K\_\{1\},cα:=\(1−α\)αα/\(1−α\),c\_\{\\alpha\}:=\(1\-\\alpha\)\\alpha^\{\\alpha/\(1\-\\alpha\)\},p:=α/\(1−α\)\.p:=\\alpha/\(1\-\\alpha\)\.Use one\-sample initialization, so thatb0≤Gb\_\{0\}\\leq G\. Choose
η=η0T−2/3,γ=γ0T−2/3,λ=λ0,\\eta=\\eta\_\{0\}\\,T^\{\-2/3\},\\qquad\\gamma=\\gamma\_\{0\}\\,T^\{\-2/3\},\\qquad\\lambda=\\lambda\_\{0\},whereη0∈\(0,1\]\\eta\_\{0\}\\in\(0,1\],γ0\>0\\gamma\_\{0\}\>0, andλ0\>0\\lambda\_\{0\}\>0satisfy
Lαλ0γ0≤1,8Lαλ0γ0η0≤1\.L\_\{\\alpha\}\\lambda\_\{0\}\\gamma\_\{0\}\\leq 1,\\qquad 8L\_\{\\alpha\}\\lambda\_\{0\}\\frac\{\\gamma\_\{0\}\}\{\\eta\_\{0\}\}\\leq 1\.Then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\[4Δγ0\+8b0η0\+8γ0η0\(K0\+cαLαλ0−p\+LαGα\)\+8η0G\]T−1/3\\displaystyle\\leq\\left\[\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\+\\frac\{8b\_\{0\}\}\{\\eta\_\{0\}\}\+\\frac\{8\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\+L\_\{\\alpha\}G^\{\\alpha\}\\right\)\+8\\sqrt\{\\eta\_\{0\}\}\\,G\\right\]T^\{\-1/3\}\+8K2γ0p\+1η0T−\(1\+2p\)/3\+2γ0\(K0\+cαLαλ0−p\)T−2/3\\displaystyle\\quad\+\\frac\{8K\_\{2\}\\gamma\_\{0\}^\{p\+1\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\\,T^\{\-\(1\+2p\)/3\}\+2\\gamma\_\{0\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\\right\)T^\{\-2/3\}\+4K2γ0p\+1T−2\(p\+1\)/3\+2LαGαγ0T−2/3\.\\displaystyle\\quad\+4K\_\{2\}\\gamma\_\{0\}^\{p\+1\}\\,T^\{\-2\(p\+1\)/3\}\+2L\_\{\\alpha\}G^\{\\alpha\}\\gamma\_\{0\}\\,T^\{\-2/3\}\.\(53\)Consequently,𝔼‖∇f\(x^\)‖=𝒪\(T−1/3\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/3\}\), and the SFO complexity is𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)\.
###### Proof\.
SinceB=0B=0, we haveqk=Gq\_\{k\}=Gfor everykk,QT=GQ\_\{T\}=G, and
1T∑k=0K𝔼\[qkα\]=Gα\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}^\{\\alpha\}\]=G^\{\\alpha\}\.We verify the conditions of Theorem[5](https://arxiv.org/html/2605.15314#Thmtheorem5)\(ii\)\. SinceT≥1T\\geq 1,
Lαλγ=Lαλ0γ0T−2/3≤Lαλ0γ0≤1,L\_\{\\alpha\}\\lambda\\gamma=L\_\{\\alpha\}\\lambda\_\{0\}\\gamma\_\{0\}\\,T^\{\-2/3\}\\leq L\_\{\\alpha\}\\lambda\_\{0\}\\gamma\_\{0\}\\leq 1,and
8Lαλγη=8Lαλ0γ0η0≤1\.8L\_\{\\alpha\}\\lambda\\frac\{\\gamma\}\{\\eta\}=8L\_\{\\alpha\}\\lambda\_\{0\}\\frac\{\\gamma\_\{0\}\}\{\\eta\_\{0\}\}\\leq 1\.Apply \([50](https://arxiv.org/html/2605.15314#A3.E50)\)\. SinceB=0B=0,QT=GQ\_\{T\}=Gis constant\. The first two terms give
4ΔγT=4Δγ0T−1/3,8b0ηT=8b0η0T−1/3\.\\frac\{4\\Delta\}\{\\gamma T\}=\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/3\},\\qquad\\frac\{8b\_\{0\}\}\{\\eta T\}=\\frac\{8b\_\{0\}\}\{\\eta\_\{0\}\}\\,T^\{\-1/3\}\.The fresh\-noise term gives
8ηG=8η0GT−1/3\.8\\sqrt\{\\eta\}\\,G=8\\sqrt\{\\eta\_\{0\}\}\\,G\\,T^\{\-1/3\}\.For the estimator\-difference term,
8γη\\displaystyle\\frac\{8\\gamma\}\{\\sqrt\{\\eta\}\}\(K0\+cαLαλ0−p\+LαGα\+K2γp\)\\displaystyle\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\+L\_\{\\alpha\}G^\{\\alpha\}\+K\_\{2\}\\gamma^\{p\}\\right\)=8γ0η0\(K0\+cαLαλ0−p\+LαGα\)T−1/3\+8K2γ0p\+1η0T−\(1\+2p\)/3\.\\displaystyle=\\frac\{8\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\+L\_\{\\alpha\}G^\{\\alpha\}\\right\)T^\{\-1/3\}\+\\frac\{8K\_\{2\}\\gamma\_\{0\}^\{p\+1\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\\,T^\{\-\(1\+2p\)/3\}\.The deterministic descent term gives
2γ\(K0\+cαLαλ0−p\+2K2γp\)=2γ0\(K0\+cαLαλ0−p\)T−2/3\+4K2γ0p\+1T−2\(p\+1\)/3\.2\\gamma\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\+2K\_\{2\}\\gamma^\{p\}\\right\)=2\\gamma\_\{0\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\\right\)T^\{\-2/3\}\+4K\_\{2\}\\gamma\_\{0\}^\{p\+1\}\\,T^\{\-2\(p\+1\)/3\}\.The averaged BG\-0 term gives
2LαγGα=2LαGαγ0T−2/3\.2L\_\{\\alpha\}\\gamma\\,G^\{\\alpha\}=2L\_\{\\alpha\}G^\{\\alpha\}\\gamma\_\{0\}\\,T^\{\-2/3\}\.Sincep\>0p\>0, the exponents\(1\+2p\)/3\>1/3\(1\+2p\)/3\>1/3and2\(p\+1\)/3\>2/32\(p\+1\)/3\>2/3, so the dominant rate isT−1/3T^\{\-1/3\}\. This proves \([53](https://arxiv.org/html/2605.15314#A3.E53)\)\.
The method uses one SFO call for initialization and two per transition, givingMK=1\+2T=𝒪\(ε−3\)M\_\{K\}=1\+2T=\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)\. ∎
###### Corollary 5\.6\(Expected11\-symmetric generalized smoothness under bounded variance\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1),[2](https://arxiv.org/html/2605.15314#Thmassumption2)\(withB=0B=0\), and[6](https://arxiv.org/html/2605.15314#Thmassumption6)hold withα=1\\alpha=1and constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0\. Use one\-sample initialization, so thatb0≤Gb\_\{0\}\\leq G\. Choose
η=η0T−2/3,γ=γ0T−2/3,\\eta=\\eta\_\{0\}\\,T^\{\-2/3\},\\qquad\\gamma=\\gamma\_\{0\}\\,T^\{\-2/3\},whereη0∈\(0,1\]\\eta\_\{0\}\\in\(0,1\]andγ0\>0\\gamma\_\{0\}\>0\. IfL1\>0L\_\{1\}\>0, assume
γ0≤142L1,162e3/4L1γ0η0≤1\.\\gamma\_\{0\}\\leq\\frac\{1\}\{4\\sqrt\{2\}\\,L\_\{1\}\},\\qquad 16\\sqrt\{2e^\{3/4\}\}\\,L\_\{1\}\\frac\{\\gamma\_\{0\}\}\{\\eta\_\{0\}\}\\leq 1\.Then
𝔼‖∇f\(x^\)‖\\displaystyle\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|≤\[4Δγ0\+8b0η0\+82e3/4γ0η0\(L0\+2L1G\)\+8η0G\]T−1/3\\displaystyle\\leq\\left\[\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\+\\frac\{8b\_\{0\}\}\{\\eta\_\{0\}\}\+\\frac\{8\\sqrt\{2e^\{3/4\}\}\\,\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\\left\(L\_\{0\}\+2L\_\{1\}G\\right\)\+8\\sqrt\{\\eta\_\{0\}\}\\,G\\right\]T^\{\-1/3\}\+\(42L0γ0\+82L1Gγ0\)T−2/3\.\\displaystyle\\quad\+\\left\(4\\sqrt\{2\}\\,L\_\{0\}\\gamma\_\{0\}\+8\\sqrt\{2\}\\,L\_\{1\}G\\gamma\_\{0\}\\right\)T^\{\-2/3\}\.\(54\)Consequently,𝔼‖∇f\(x^\)‖=𝒪\(T−1/3\)\\mathbb\{E\}\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/3\}\), and the SFO complexity is𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)\.
###### Proof\.
SinceB=0B=0, we haveqk=Gq\_\{k\}=Gfor everykkand
1T∑k=0K𝔼\[qk\]=G\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\mathbb\{E\}\[q\_\{k\}\]=G\.IfL1\>0L\_\{1\}\>0, the conditions of Theorem[5](https://arxiv.org/html/2605.15314#Thmtheorem5)\(iii\) hold because
γ=γ0T−2/3≤γ0≤142L1,\\gamma=\\gamma\_\{0\}\\,T^\{\-2/3\}\\leq\\gamma\_\{0\}\\leq\\frac\{1\}\{4\\sqrt\{2\}\\,L\_\{1\}\},and
162e3/4L1γη=162e3/4L1γ0η0≤1\.16\\sqrt\{2e^\{3/4\}\}\\,L\_\{1\}\\frac\{\\gamma\}\{\\eta\}=16\\sqrt\{2e^\{3/4\}\}\\,L\_\{1\}\\frac\{\\gamma\_\{0\}\}\{\\eta\_\{0\}\}\\leq 1\.Apply \([51](https://arxiv.org/html/2605.15314#A3.E51)\) withB=0B=0\. We have
4ΔγT=4Δγ0T−1/3,8b0ηT=8b0η0T−1/3\.\\frac\{4\\Delta\}\{\\gamma T\}=\\frac\{4\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/3\},\\qquad\\frac\{8b\_\{0\}\}\{\\eta T\}=\\frac\{8b\_\{0\}\}\{\\eta\_\{0\}\}\\,T^\{\-1/3\}\.SinceG\+BTγ=GG\+BT\\gamma=G,
8γη2e3/4\(L0\+2L1G\)=82e3/4γ0η0\(L0\+2L1G\)T−1/3,\\frac\{8\\gamma\}\{\\sqrt\{\\eta\}\}\\sqrt\{2e^\{3/4\}\}\\left\(L\_\{0\}\+2L\_\{1\}G\\right\)=\\frac\{8\\sqrt\{2e^\{3/4\}\}\\,\\gamma\_\{0\}\}\{\\sqrt\{\\eta\_\{0\}\}\}\(L\_\{0\}\+2L\_\{1\}G\)\\,T^\{\-1/3\},8ηG=8η0GT−1/3,8\\sqrt\{\\eta\}\\,G=8\\sqrt\{\\eta\_\{0\}\}\\,G\\,T^\{\-1/3\},42L0γ=42L0γ0T−2/3,4\\sqrt\{2\}\\,L\_\{0\}\\gamma=4\\sqrt\{2\}\\,L\_\{0\}\\gamma\_\{0\}\\,T^\{\-2/3\},and
82L1γG=82L1Gγ0T−2/3\.8\\sqrt\{2\}\\,L\_\{1\}\\gamma\\,G=8\\sqrt\{2\}\\,L\_\{1\}G\\gamma\_\{0\}\\,T^\{\-2/3\}\.Combining gives \([54](https://arxiv.org/html/2605.15314#A3.E54)\)\. The SFO cost is1\+2T=𝒪\(ε−3\)1\+2T=\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)\. ∎
### C\.8Recovery of deterministic rates \(B=G=0B=G=0\)
When the oracle is deterministic \(B=G=0B=G=0\), the stochastic gradient noise vanishes\. Thereforevk=∇f\(xk\)v^\{k\}=\\nabla f\(x^\{k\}\),bk=0b\_\{k\}=0, andqk=0q\_\{k\}=0for allkk\. The𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}update reduces to normalized gradient descent, and the estimator recursion plays no role\. The classical𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)rate is recovered for all three smoothness regimes\.
###### Corollary 5\.7\(Mean\-square smoothness, deterministic\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[4](https://arxiv.org/html/2605.15314#Thmassumption4)hold, and the oracle is deterministic \(B=G=0B=G=0\)\. Consider the normalized gradient descent update
xk\+1=xk−γ∇f\(xk\)‖∇f\(xk\)‖\.x^\{k\+1\}=x^\{k\}\-\\gamma\\,\\frac\{\\nabla f\(x^\{k\}\)\}\{\\\|\\nabla f\(x^\{k\}\)\\\|\}\.Chooseγ=γ0T−1/2\\gamma=\\gamma\_\{0\}\\,T^\{\-1/2\}\. Then
‖∇f\(x^\)‖≤Δγ0T−1/2\+Lγ02T−1/2\.\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\frac\{\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/2\}\+\\frac\{L\\gamma\_\{0\}\}\{2\}\\,T^\{\-1/2\}\.\(55\)Consequently,‖∇f\(x^\)‖=𝒪\(T−1/2\)\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/2\}\), and the iteration complexity is𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)\.
###### Proof\.
Under determinism,vk=∇f\(xk\)v^\{k\}=\\nabla f\(x^\{k\}\)andbk=0b\_\{k\}=0for allkk\. Sincedk=∇f\(xk\)/‖∇f\(xk\)‖d^\{k\}=\\nabla f\(x^\{k\}\)/\\\|\\nabla f\(x^\{k\}\)\\\|\(when nonzero\), we have⟨∇f\(xk\),dk⟩=‖∇f\(xk\)‖\\langle\\nabla f\(x^\{k\}\),d^\{k\}\\rangle=\\\|\\nabla f\(x^\{k\}\)\\\|\.
Assumption[4](https://arxiv.org/html/2605.15314#Thmassumption4)impliesLL\-smoothness via Jensen \(Lemma[7](https://arxiv.org/html/2605.15314#Thmlemma7)\)\. The one\-step descent becomes
f\(xk\+1\)≤f\(xk\)−γ‖∇f\(xk\)‖\+L2γ2\.f\(x^\{k\+1\}\)\\leq f\(x^\{k\}\)\-\\gamma\\\|\\nabla f\(x^\{k\}\)\\\|\+\\frac\{L\}\{2\}\\gamma^\{2\}\.Summing fromk=0k=0toKKand usingf\(xK\+1\)≥finff\(x^\{K\+1\}\)\\geq f^\{\\inf\},
γ∑k=0K‖∇f\(xk\)‖≤Δ\+L2γ2T\.\\gamma\\sum\_\{k=0\}^\{K\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\Delta\+\\frac\{L\}\{2\}\\gamma^\{2\}T\.Dividing byγT\\gamma Tand substitutingγ=γ0T−1/2\\gamma=\\gamma\_\{0\}T^\{\-1/2\}gives \([55](https://arxiv.org/html/2605.15314#A3.E55)\)\. ∎
###### Corollary 5\.8\(Expectedα\\alpha\-symmetric generalized smoothness, deterministic\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[6](https://arxiv.org/html/2605.15314#Thmassumption6)hold withα∈\(0,1\)\\alpha\\in\(0,1\)and constantsK0,K1,K2\>0K\_\{0\},K\_\{1\},K\_\{2\}\>0, and the oracle is deterministic \(B=G=0B=G=0\)\. DefineLα:=2α/2K1L\_\{\\alpha\}:=2^\{\\alpha/2\}K\_\{1\},cα:=\(1−α\)αα/\(1−α\)c\_\{\\alpha\}:=\(1\-\\alpha\)\\alpha^\{\\alpha/\(1\-\\alpha\)\}, andp:=α/\(1−α\)p:=\\alpha/\(1\-\\alpha\)\. Consider the normalized gradient descent update\. Chooseγ=γ0T−1/2\\gamma=\\gamma\_\{0\}\\,T^\{\-1/2\}and letλ0\>0\\lambda\_\{0\}\>0satisfyLαλ0γ0≤1L\_\{\\alpha\}\\lambda\_\{0\}\\gamma\_\{0\}\\leq 1\. Then
‖∇f\(x^\)‖≤2Δγ0T−1/2\+γ0\(K0\+cαLαλ0−p\+2K2γ0pT−p/2\)T−1/2\.\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/2\}\+\\gamma\_\{0\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\+2K\_\{2\}\\gamma\_\{0\}^\{p\}T^\{\-p/2\}\\right\)T^\{\-1/2\}\.\(56\)Consequently,‖∇f\(x^\)‖=𝒪\(T−1/2\)\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/2\}\), and the iteration complexity is𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)\.
###### Proof\.
Under determinism,vk=∇f\(xk\)v^\{k\}=\\nabla f\(x^\{k\}\),bk=0b\_\{k\}=0, andqk=0q\_\{k\}=0for allkk\. The one\-step descent from Lemma[9](https://arxiv.org/html/2605.15314#Thmlemma9)withBk=0B\_\{k\}=0and𝔼\[qkα\]=0\\mathbb\{E\}\[q\_\{k\}^\{\\alpha\}\]=0becomes
f\(xk\+1\)\+γ2‖∇f\(xk\)‖≤f\(xk\)\+γ22\(K0\+cαLαλ0−p\+2K2γp\)\.f\(x^\{k\+1\}\)\+\\frac\{\\gamma\}\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq f\(x^\{k\}\)\+\\frac\{\\gamma^\{2\}\}\{2\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\+2K\_\{2\}\\gamma^\{p\}\\right\)\.Summing fromk=0k=0toKK, usingf\(xK\+1\)≥finff\(x^\{K\+1\}\)\\geq f^\{\\inf\}, and dividing byγT/2\\gamma T/2gives
1T∑k=0K‖∇f\(xk\)‖≤2ΔγT\+γ\(K0\+cαLαλ0−p\+2K2γp\)\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\frac\{2\\Delta\}\{\\gamma T\}\+\\gamma\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\+2K\_\{2\}\\gamma^\{p\}\\right\)\.Substitutingγ=γ0T−1/2\\gamma=\\gamma\_\{0\}T^\{\-1/2\}gives
2ΔγT=2Δγ0T−1/2,\\frac\{2\\Delta\}\{\\gamma T\}=\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/2\},and
γ\(K0\+cαLαλ0−p\)=γ0\(K0\+cαLαλ0−p\)T−1/2,\\gamma\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\\right\)=\\gamma\_\{0\}\\left\(K\_\{0\}\+c\_\{\\alpha\}L\_\{\\alpha\}\\lambda\_\{0\}^\{\-p\}\\right\)T^\{\-1/2\},and
2K2γp\+1=2K2γ0p\+1T−\(p\+1\)/2\.2K\_\{2\}\\gamma^\{p\+1\}=2K\_\{2\}\\gamma\_\{0\}^\{p\+1\}\\,T^\{\-\(p\+1\)/2\}\.Sincep\>0p\>0, the last term decays faster thanT−1/2T^\{\-1/2\}\. This proves \([56](https://arxiv.org/html/2605.15314#A3.E56)\) and the𝒪\(T−1/2\)\\mathcal\{O\}\(T^\{\-1/2\}\)rate\. ∎
###### Corollary 5\.9\(Expected11\-symmetric generalized smoothness, deterministic\)\.
Suppose Assumptions[1](https://arxiv.org/html/2605.15314#Thmassumption1)and[6](https://arxiv.org/html/2605.15314#Thmassumption6)hold withα=1\\alpha=1and constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0, and the oracle is deterministic \(B=G=0B=G=0\)\. Consider the normalized gradient descent update\. Chooseγ=γ0T−1/2\\gamma=\\gamma\_\{0\}\\,T^\{\-1/2\}withγ0≤1/\(42L1\)\\gamma\_\{0\}\\leq 1/\(4\\sqrt\{2\}\\,L\_\{1\}\)ifL1\>0L\_\{1\}\>0\. Then
‖∇f\(x^\)‖≤2Δγ0T−1/2\+22L0γ0T−1/2\.\\\|\\nabla f\(\\widehat\{x\}\)\\\|\\leq\\frac\{2\\Delta\}\{\\gamma\_\{0\}\}\\,T^\{\-1/2\}\+2\\sqrt\{2\}\\,L\_\{0\}\\gamma\_\{0\}\\,T^\{\-1/2\}\.\(57\)Consequently,‖∇f\(x^\)‖=𝒪\(T−1/2\)\\\|\\nabla f\(\\widehat\{x\}\)\\\|=\\mathcal\{O\}\(T^\{\-1/2\}\), and the iteration complexity is𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)\.
###### Proof\.
Under determinism,vk=∇f\(xk\)v^\{k\}=\\nabla f\(x^\{k\}\),bk=0b\_\{k\}=0, andqk=0q\_\{k\}=0for allkk\. The one\-step descent from Lemma[10](https://arxiv.org/html/2605.15314#Thmlemma10)becomes
f\(xk\+1\)\+γ2‖∇f\(xk\)‖≤f\(xk\)\+2L0γ2\.f\(x^\{k\+1\}\)\+\\frac\{\\gamma\}\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq f\(x^\{k\}\)\+\\sqrt\{2\}\\,L\_\{0\}\\gamma^\{2\}\.Indeed, the22L1γ2‖∇f\(xk\)‖2\\sqrt\{2\}\\,L\_\{1\}\\gamma^\{2\}\\\|\\nabla f\(x^\{k\}\)\\\|term is absorbed by\(γ/2\)‖∇f\(xk\)‖\(\\gamma/2\)\\\|\\nabla f\(x^\{k\}\)\\\|via the stepsize condition22L1γ≤1/22\\sqrt\{2\}\\,L\_\{1\}\\gamma\\leq 1/2, and theqkq\_\{k\}\-dependent term vanishes\.
Summing fromk=0k=0toKK, usingf\(xK\+1\)≥finff\(x^\{K\+1\}\)\\geq f^\{\\inf\}, and dividing byγT/2\\gamma T/2gives
1T∑k=0K‖∇f\(xk\)‖≤2ΔγT\+22L0γ\.\\frac\{1\}\{T\}\\sum\_\{k=0\}^\{K\}\\\|\\nabla f\(x^\{k\}\)\\\|\\leq\\frac\{2\\Delta\}\{\\gamma T\}\+2\\sqrt\{2\}\\,L\_\{0\}\\gamma\.Substitutingγ=γ0T−1/2\\gamma=\\gamma\_\{0\}T^\{\-1/2\}proves \([57](https://arxiv.org/html/2605.15314#A3.E57)\)\. ∎
### C\.9Discussion on recovered rates
The six recovery corollaries above confirm the claims made in the discussion following Theorem[2](https://arxiv.org/html/2605.15314#Thmtheorem2)\.
- •Bounded variance recovery\.WhenB=0B=0, Corollaries[5\.4](https://arxiv.org/html/2605.15314#Thmtheorem5.Thmcorollary4),[5\.5](https://arxiv.org/html/2605.15314#Thmtheorem5.Thmcorollary5), and[5\.6](https://arxiv.org/html/2605.15314#Thmtheorem5.Thmcorollary6)all achieve𝒪\(ε−3\)\\mathcal\{O\}\(\\varepsilon^\{\-3\}\)SFO complexity with the bounded variance scheduleη=η0T−2/3\\eta=\\eta\_\{0\}T^\{\-2/3\},γ=γ0T−2/3\\gamma=\\gamma\_\{0\}T^\{\-2/3\}\. This recovers the standard variance\-reduced rate ofFanget al\.\[[12](https://arxiv.org/html/2605.15314#bib.bib45)\], Cutkosky and Orabona \[[10](https://arxiv.org/html/2605.15314#bib.bib41)\], Chenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]under mean\-square smoothness and expected generalized smoothness alike\.
- •Deterministic recovery\.WhenB=G=0B=G=0, Corollaries[5\.7](https://arxiv.org/html/2605.15314#Thmtheorem5.Thmcorollary7),[5\.8](https://arxiv.org/html/2605.15314#Thmtheorem5.Thmcorollary8), and[5\.9](https://arxiv.org/html/2605.15314#Thmtheorem5.Thmcorollary9)all achieve𝒪\(ε−2\)\\mathcal\{O\}\(\\varepsilon^\{\-2\}\)iteration complexity\. Since𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}reduces to normalized gradient descent in the deterministic setting, these rates match the classical first\-order complexity ofNesterov \[[34](https://arxiv.org/html/2605.15314#bib.bib56)\]and the generalized smooth rates ofChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]\.
- •Schedule transition\.As the noise weakens from𝖡𝖦\\mathsf\{BG\}\-0 to bounded variance to deterministic, the𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}schedules and rates transition as follows for the mean\-square smoothness case: \(η,γ\)=\{\(T−1,γ0T−3/4\),𝖡𝖦\-0\(sharp init\.\),\(η0T−2/3,γ0T−2/3\),bounded variance,\(1,γ0T−1/2\),deterministic,\(\\eta,\\,\\gamma\)=\\begin\{cases\}\(T^\{\-1\},\\;\\gamma\_\{0\}T^\{\-3/4\}\),&\\mathsf\{BG\}\\text\{\-\}0\\text\{ \(sharp init\.\)\},\\\\\[3\.0pt\] \(\\eta\_\{0\}T^\{\-2/3\},\\;\\gamma\_\{0\}T^\{\-2/3\}\),&\\text\{bounded variance\},\\\\\[3\.0pt\] \(1,\\;\\gamma\_\{0\}T^\{\-1/2\}\),&\\text\{deterministic\},\\end\{cases\}yielding ratesT−1/4T^\{\-1/4\},T−1/3T^\{\-1/3\}, andT−1/2T^\{\-1/2\}, respectively\. Theη\\etaparameter decays more slowly as the noise weakens, because the STORM recursion needs less aggressive averaging when the fresh noise and trajectory\-growth contributions are smaller\. For the expectedα\\alpha\-generalized smoothness case, the transition is: SFO complexity=\{𝒪\(ε−\(4\+α\)\),𝖡𝖦\-0,𝒪\(ε−3\),bounded variance,𝒪\(ε−2\),deterministic\.\\text\{SFO complexity\}=\\begin\{cases\}\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\),&\\mathsf\{BG\}\\text\{\-\}0,\\\\\[3\.0pt\] \\mathcal\{O\}\(\\varepsilon^\{\-3\}\),&\\text\{bounded variance\},\\\\\[3\.0pt\] \\mathcal\{O\}\(\\varepsilon^\{\-2\}\),&\\text\{deterministic\}\.\\end\{cases\}Thus theα\\alpha\-dependent price𝒪\(ε−\(4\+α\)\)\\mathcal\{O\}\(\\varepsilon^\{\-\(4\+\\alpha\)\}\)under𝖡𝖦\\mathsf\{BG\}\-0 is entirely an artifact of the interaction between distance\-dependent variance and gradient\-dependent curvature; it disappears when either source of difficulty is removed\.
## Appendix DExperimental Details
### D\.1𝖡𝖦\\mathsf\{BG\}\-0 Wrapper
Given a fixed deterministic objectiveff, the𝖡𝖦\\mathsf\{BG\}\-0 wrapper turns the setup into a stochastic problem in which the stochastic gradients satisfy the𝖡𝖦\\mathsf\{BG\}\-0 variance model exactly\.
Given the independent random variablesρ\\rhoanduusatisfying
𝔼\[ρ\]=0,𝔼\[ρ2\]=1,𝔼\[u\]=0,𝔼‖u‖2=1\.\\mathbb\{E\}\[\\rho\]=0,\\qquad\\mathbb\{E\}\[\\rho^\{2\}\]=1,\\qquad\\mathbb\{E\}\[u\]=0,\\qquad\\mathbb\{E\}\\\|u\\\|^\{2\}=1\.whereρ∈\{−1,\+1\}\\rho\\in\\\{\-1,\+1\\\}uniformly andu∼𝒩\(0,Id/d\)u\\sim\\mathcal\{N\}\(0,I\_\{d\}/d\)\. For each sampleξ=\(ρ,u\)\\xi=\(\\rho,u\), the stochastic loss is given by
fξ\(x\)=f\(x\)\+Bρ2‖x−x0‖2\+G⟨u,x−x0⟩,f\_\{\\xi\}\(x\)=f\(x\)\+\\frac\{B\\rho\}\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G\\langle u,x\-x^\{0\}\\rangle,where the corresponding stochastic gradient is
∇fξ\(x\)=∇f\(x\)\+Bρ\(x−x0\)\+Gu\.\\nabla f\_\{\\xi\}\(x\)=\\nabla f\(x\)\+B\\rho\(x\-x^\{0\}\)\+Gu\.Here, the objective is unchanged and the oracle is unbiased as
𝔼ξ\[fξ\(x\)\]=f\(x\),𝔼ξ\[∇fξ\(x\)\]=∇f\(x\)\.\\mathbb\{E\}\_\{\\xi\}\[f\_\{\\xi\}\(x\)\]=f\(x\),\\qquad\\mathbb\{E\}\_\{\\xi\}\[\\nabla f\_\{\\xi\}\(x\)\]=\\nabla f\(x\)\.Moreover,
𝔼ξ‖∇fξ\(x\)−∇f\(x\)‖2=𝔼ξ‖Bρ\(x−x0\)\+Gu‖2=B2‖x−x0‖2\+G2,\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\_\{\\xi\}\(x\)\-\\nabla f\(x\)\\\|^\{2\}=\\mathbb\{E\}\_\{\\xi\}\\\|B\\rho\(x\-x^\{0\}\)\+Gu\\\|^\{2\}=B^\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G^\{2\},where the cross term vanishes by independence and the zero\-mean assumptions\. Thus, our wrapper satisfies the exact𝖡𝖦\\mathsf\{BG\}\-0 oracle\.
As the extension to the mini\-batch case of sizebb\(which is essential to compare with the dynamic batching based models in\[[13](https://arxiv.org/html/2605.15314#bib.bib139)\]\), we draw independentξi=\(ρi,ui\)\\xi\_\{i\}=\(\\rho\_\{i\},u\_\{i\}\)and use
gb\(x\)=1b∑i=1b∇fξi\(x\)=∇f\(x\)\+1b∑i=1b\(Bρi\(x−x0\)\+Gui\),g\_\{b\}\(x\)=\\frac\{1\}\{b\}\\sum\_\{i=1\}^\{b\}\\nabla f\_\{\\xi\_\{i\}\}\(x\)=\\nabla f\(x\)\+\\frac\{1\}\{b\}\\sum\_\{i=1\}^\{b\}\\left\(B\\rho\_\{i\}\(x\-x^\{0\}\)\+Gu\_\{i\}\\right\),which satisfies
𝔼\[gb\(x\)\]=∇f\(x\),𝔼‖gb\(x\)−∇f\(x\)‖2=B2‖x−x0‖2\+G2b\.\\mathbb\{E\}\[g\_\{b\}\(x\)\]=\\nabla f\(x\),\\qquad\\mathbb\{E\}\\\|g\_\{b\}\(x\)\-\\nabla f\(x\)\\\|^\{2\}=\\frac\{B^\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G^\{2\}\}\{b\}\.For the STORM and NSTORM recursions, when a transition fromxxtoyyrequires two stochastic gradients, we evaluate∇fξ\(x\)\\nabla f\_\{\\xi\}\(x\)and∇fξ\(y\)\\nabla f\_\{\\xi\}\(y\)with the same sampleξ\\xiand analogously reuse the same mini\-batch in the batched case\. Since this process requires the computation of two stochastic gradients, the corresponding SFO is doubled in our calculations\.
###### Lemma 14\.
Letf∈ℒsym∗\(α\)f\\in\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\)with constantsL0,L1\>0L\_\{0\},L\_\{1\}\>0, and letα∈\(0,1\]\\alpha\\in\(0,1\]\. Letρ∈\{−1,\+1\}\\rho\\in\\\{\-1,\+1\\\}uniformly andu∼𝒩\(0,Id/d\)u\\sim\\mathcal\{N\}\(0,I\_\{d\}/d\), independently\. Thus
𝔼\[ρ\]=0,𝔼\[ρ2\]=1,𝔼\[u\]=0,𝔼‖u‖2=1\.\\mathbb\{E\}\[\\rho\]=0,\\qquad\\mathbb\{E\}\[\\rho^\{2\}\]=1,\\qquad\\mathbb\{E\}\[u\]=0,\\qquad\\mathbb\{E\}\\\|u\\\|^\{2\}=1\.For eachξ=\(ρ,u\)\\xi=\(\\rho,u\), define
fξ\(x\)=f\(x\)\+Bρ2‖x−x0‖2\+G⟨u,x−x0⟩\.f\_\{\\xi\}\(x\)=f\(x\)\+\\frac\{B\\rho\}\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G\\langle u,x\-x^\{0\}\\rangle\.Then𝔼ξ\[fξ\(x\)\]=f\(x\)\\mathbb\{E\}\_\{\\xi\}\[f\_\{\\xi\}\(x\)\]=f\(x\), the stochastic gradient is unbiased, and
𝔼ξ‖∇fξ\(x\)−∇f\(x\)‖2=B2‖x−x0‖2\+G2\.\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\_\{\\xi\}\(x\)\-\\nabla f\(x\)\\\|^\{2\}=B^\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G^\{2\}\.Moreover, there exist constantsL~0,L~1\>0\\widetilde\{L\}\_\{0\},\\widetilde\{L\}\_\{1\}\>0such that, for allx,yx,y,
𝔼ξ‖∇fξ\(y\)−∇fξ\(x\)‖2≤‖y−x‖2𝔼ξ\[\(L~0\+L~1maxθ∈\[0,1\]‖∇fξ\(xθ\)‖α\)2\],\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\_\{\\xi\}\(y\)\-\\nabla f\_\{\\xi\}\(x\)\\\|^\{2\}\\leq\\\|y\-x\\\|^\{2\}\\mathbb\{E\}\_\{\\xi\}\\left\[\\left\(\\widetilde\{L\}\_\{0\}\+\\widetilde\{L\}\_\{1\}\\max\_\{\\theta\\in\[0,1\]\}\\\|\\nabla f\_\{\\xi\}\(x\_\{\\theta\}\)\\\|^\{\\alpha\}\\right\)^\{2\}\\right\],wherexθ=θy\+\(1−θ\)xx\_\{\\theta\}=\\theta y\+\(1\-\\theta\)x\.
###### Proof\.
Since𝔼\[ρ\]=0\\mathbb\{E\}\[\\rho\]=0and𝔼\[u\]=0\\mathbb\{E\}\[u\]=0we have𝔼ξ\[fξ\(x\)\]=f\(x\)\\mathbb\{E\}\_\{\\xi\}\[f\_\{\\xi\}\(x\)\]=f\(x\)and𝔼ξ\[∇fξ\(x\)\]=∇f\(x\)\\mathbb\{E\}\_\{\\xi\}\[\\nabla f\_\{\\xi\}\(x\)\]=\\nabla f\(x\)\. Moreover,
∇fξ\(x\)−∇f\(x\)=Bρ\(x−x0\)\+Gu\.\\nabla f\_\{\\xi\}\(x\)\-\\nabla f\(x\)=B\\rho\(x\-x^\{0\}\)\+Gu\.Using independence and the properties ofρ\\rhoanduu, we obtain
𝔼ξ‖∇fξ\(x\)−∇f\(x\)‖2=B2‖x−x0‖2\+G2\.\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\_\{\\xi\}\(x\)\-\\nabla f\(x\)\\\|^\{2\}=B^\{2\}\\\|x\-x^\{0\}\\\|^\{2\}\+G^\{2\}\.It remains to prove the expected generalized\-smoothness bound\. For anyx,yx,y, letd=y−xd=y\-x\. Since
∇fξ\(y\)−∇fξ\(x\)=∇f\(y\)−∇f\(x\)\+Bρd,\\nabla f\_\{\\xi\}\(y\)\-\\nabla f\_\{\\xi\}\(x\)=\\nabla f\(y\)\-\\nabla f\(x\)\+B\\rho d,we have
𝔼ξ‖∇fξ\(y\)−∇fξ\(x\)‖2=‖∇f\(y\)−∇f\(x\)‖2\+B2‖d‖2\.\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\_\{\\xi\}\(y\)\-\\nabla f\_\{\\xi\}\(x\)\\\|^\{2\}=\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|^\{2\}\+B^\{2\}\\\|d\\\|^\{2\}\.Sincef∈ℒsym∗\(α\)f\\in\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\),
‖∇f\(y\)−∇f\(x\)‖2≤‖d‖2\(L0\+L1maxθ∈\[0,1\]‖∇f\(xθ\)‖α\)2\.\\\|\\nabla f\(y\)\-\\nabla f\(x\)\\\|^\{2\}\\leq\\\|d\\\|^\{2\}\\left\(L\_\{0\}\+L\_\{1\}\\max\_\{\\theta\\in\[0,1\]\}\\\|\\nabla f\(x\_\{\\theta\}\)\\\|^\{\\alpha\}\\right\)^\{2\}\.Thus, with
R:=maxθ∈\[0,1\]‖∇f\(xθ\)‖,R:=\\max\_\{\\theta\\in\[0,1\]\}\\\|\\nabla f\(x\_\{\\theta\}\)\\\|,we obtain
𝔼ξ‖∇fξ\(y\)−∇fξ\(x\)‖2≤‖d‖2\[2L02\+2L12R2α\+B2\]\.\\mathbb\{E\}\_\{\\xi\}\\\|\\nabla f\_\{\\xi\}\(y\)\-\\nabla f\_\{\\xi\}\(x\)\\\|^\{2\}\\leq\\\|d\\\|^\{2\}\\left\[2L\_\{0\}^\{2\}\+2L\_\{1\}^\{2\}R^\{2\\alpha\}\+B^\{2\}\\right\]\.Now define
Mξ:=maxθ∈\[0,1\]‖∇fξ\(xθ\)‖\.M\_\{\\xi\}:=\\max\_\{\\theta\\in\[0,1\]\}\\\|\\nabla f\_\{\\xi\}\(x\_\{\\theta\}\)\\\|\.For any fixedθ\\theta, write
Zθ:=Bρ\(xθ−x0\)\+Gu\.Z\_\{\\theta\}:=B\\rho\(x\_\{\\theta\}\-x^\{0\}\)\+Gu\.The random vectorZθZ\_\{\\theta\}is symmetric:ZθZ\_\{\\theta\}and−Zθ\-Z\_\{\\theta\}have the same distribution\. Hence, for any fixed vectoraaand anyp\>0p\>0,
𝔼ξ‖a\+Zθ‖p=12𝔼ξ\[‖a\+Zθ‖p\+‖a−Zθ‖p\]≥12‖a‖p,\\mathbb\{E\}\_\{\\xi\}\\\|a\+Z\_\{\\theta\}\\\|^\{p\}=\\frac\{1\}\{2\}\\mathbb\{E\}\_\{\\xi\}\\left\[\\\|a\+Z\_\{\\theta\}\\\|^\{p\}\+\\\|a\-Z\_\{\\theta\}\\\|^\{p\}\\right\]\\geq\\frac\{1\}\{2\}\\\|a\\\|^\{p\},because at least one of‖a\+Zθ‖\\\|a\+Z\_\{\\theta\}\\\|and‖a−Zθ‖\\\|a\-Z\_\{\\theta\}\\\|is at least‖a‖\\\|a\\\|\. Takinga=∇f\(xθ\)a=\\nabla f\(x\_\{\\theta\}\),p=2αp=2\\alpha, and usingMξ≥‖∇fξ\(xθ\)‖M\_\{\\xi\}\\geq\\\|\\nabla f\_\{\\xi\}\(x\_\{\\theta\}\)\\\|, we get
𝔼ξ\[Mξ2α\]≥12‖∇f\(xθ\)‖2αfor allθ∈\[0,1\]\.\\mathbb\{E\}\_\{\\xi\}\[M\_\{\\xi\}^\{2\\alpha\}\]\\geq\\frac\{1\}\{2\}\\\|\\nabla f\(x\_\{\\theta\}\)\\\|^\{2\\alpha\}\\quad\\text\{for all \}\\theta\\in\[0,1\]\.Therefore,
R2α≤2𝔼ξ\[Mξ2α\]\.R^\{2\\alpha\}\\leq 2\\mathbb\{E\}\_\{\\xi\}\[M\_\{\\xi\}^\{2\\alpha\}\]\.Choose constantsL~0,L~1\>0\\widetilde\{L\}\_\{0\},\\widetilde\{L\}\_\{1\}\>0such that
L~02≥2L02\+B2,L~12≥4L12\.\\widetilde\{L\}\_\{0\}^\{2\}\\geq 2L\_\{0\}^\{2\}\+B^\{2\},\\qquad\\widetilde\{L\}\_\{1\}^\{2\}\\geq 4L\_\{1\}^\{2\}\.Then
2L02\+B2\+2L12R2α≤L~02\+L~12𝔼ξ\[Mξ2α\]≤𝔼ξ\[\(L~0\+L~1Mξα\)2\]\.2L\_\{0\}^\{2\}\+B^\{2\}\+2L\_\{1\}^\{2\}R^\{2\\alpha\}\\leq\\widetilde\{L\}\_\{0\}^\{2\}\+\\widetilde\{L\}\_\{1\}^\{2\}\\mathbb\{E\}\_\{\\xi\}\[M\_\{\\xi\}^\{2\\alpha\}\]\\leq\\mathbb\{E\}\_\{\\xi\}\\left\[\\left\(\\widetilde\{L\}\_\{0\}\+\\widetilde\{L\}\_\{1\}M\_\{\\xi\}^\{\\alpha\}\\right\)^\{2\}\\right\]\.Combining the preceding inequalities proves the claim\. ∎
This wrapper lets us evaluate stochastic methods on deterministic test objectives while imposing an unbounded variance that grows with the distance from initialization\. It therefore provides a controlled experimental model for the distance\-dependent stochastic noise studied in the paper, without changing the underlying deterministic objective\. Lemma[14](https://arxiv.org/html/2605.15314#Thmlemma14)verifies that for allα∈\(0,1\]\\alpha\\in\(0,1\]\(including theα=1/2\\alpha=1/2andα=2/3\\alpha=2/3experimental setups\) the proposed framework is compatible with the expected symmetric generalized\-smoothness framework\.
### D\.2Phase Retrieval
The first experiment is the phase retrieval objective fromChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]\. Given measurementsyr=\|ar⊤x⋆\|2y\_\{r\}=\|a\_\{r\}^\{\\top\}x\_\{\\star\}\|^\{2\}, we minimize
f\(x\)=12m∑r=1m\(yr−\|ar⊤x\|2\)2\.f\(x\)=\\frac\{1\}\{2m\}\\sum\_\{r=1\}^\{m\}\\left\(y\_\{r\}\-\|a\_\{r\}^\{\\top\}x\|^\{2\}\\right\)^\{2\}\.This nonconvex objective is not globallyLL\-smooth, but it belongs to theα\\alpha\-symmetric generalized\-smooth classℒsym∗\(2/3\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(2/3\), and its stochastic finite\-sum version satisfies the corresponding expected conditionChenet al\.\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]\. In our experiments, we use dimensiond=100d=100andm=3000m=3000measurements\. The measurement matrix has independent entriesar,j∼𝒩\(0,0\.12\)a\_\{r,j\}\\sim\\mathcal\{N\}\(0,0\.1^\{2\}\), and the target signal is drawn asx⋆,j∼𝒩\(0,1\)x\_\{\\star,j\}\\sim\\mathcal\{N\}\(0,1\)\. Hence,yr=\|ar⊤x⋆\|2y\_\{r\}=\|a\_\{r\}^\{\\top\}x\_\{\\star\}\|^\{2\}\. The optimization is initialized independently fromxj0∼𝒩\(5,1\)x^\{0\}\_\{j\}\\sim\\mathcal\{N\}\(5,1\), which places the initial point away from the target signal and makes the distance\-dependent variance growth visible\. After constructing this deterministic phase retrieval objective, we apply the𝖡𝖦\\mathsf\{BG\}\-0 wrapper from Appendix[D\.1](https://arxiv.org/html/2605.15314#A4.SS1)\. Thus, the stochasticity is controlled through the𝖡𝖦\\mathsf\{BG\}\-0 oracle\.
### D\.3Cubic Polynomial
For the second experiment, we use the following polynomial function as the deterministic objective:
f\(x\)=\|x\|\(2−α\)/\(1−α\)f\(x\)=\|x\|^\{\(2\-\\alpha\)/\(1\-\\alpha\)\}which has been shown in\[[7](https://arxiv.org/html/2605.15314#bib.bib142)\]to satisfyf∈ℒsym∗\(α\)f\\in\\mathcal\{L\}\_\{\\mathrm\{sym\}\}^\{\*\}\(\\alpha\)forw∈ℝw\\in\\mathbb\{R\}andα∈\(0,1\)\\alpha\\in\(0,1\)\. We instantiate this example withα=1/2\\alpha=1/2, sof\(x\)=\|x\|3f\(x\)=\|x\|^\{3\}, and use the same distance\-dependent𝖡𝖦\\mathsf\{BG\}\-0 oracle\. This gives a minimal one\-dimensional experiment in which both the generalized smoothness exponent and the𝖡𝖦\\mathsf\{BG\}\-0 variance growth are explicit\.
## Appendix EHyperparameters
We report the hyperparameters used for the normalized methods in Figures[1](https://arxiv.org/html/2605.15314#S6.F1)and[2](https://arxiv.org/html/2605.15314#S6.F2), established through our theorems\. We denoteT:=K\+1T:=K\+1whereKKis the total number of iterations for a given model\. The schedules below specify the theorem\-driven choices for the plotted𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}and𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}curves\. For fair comparison, we apply grid search to set the learning rate of𝖲𝖦𝖣\\mathsf\{SGD\}\(b=1\),𝖲𝖦𝖣\\mathsf\{SGD\}\(dynamic batch\), and𝖲𝖳𝖮𝖱𝖬\\mathsf\{STORM\}\(dynamic batch\)\.
#### Normalized SGDM\.
For single sample𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, we use the schedule from Theorem[1](https://arxiv.org/html/2605.15314#Thmtheorem1):
η=T−2/3,γ=γ0T−5/6,\\eta=T^\{\-2/3\},\\qquad\\gamma=\\gamma\_\{0\}T^\{\-5/6\},which commonly satisfies the standard smooth,ℒsym∗\(α\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(\\alpha\), andℒsym∗\(1\)\\mathcal\{L\}\_\{\\rm sym\}^\{\*\}\(1\)cases in Theorem[1](https://arxiv.org/html/2605.15314#Thmtheorem1)\.
#### Normalized STORM\.
For𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}, the initialization batch and schedule depend on the smoothness regime in Theorem[2](https://arxiv.org/html/2605.15314#Thmtheorem2)\. Since we consider expectedα\\alpha\-symmetric generalized smoothness in our experiments, the theorem results in
Ninit=max\{1,⌈G2T2\(1−α\)4\+α⌉\},γ=γ0T−3\+α4\+α,η=η0T−44\+α\.N\_\{\\rm init\}=\\max\\left\\\{1,\\left\\lceil G^\{2\}T^\{\\frac\{2\(1\-\\alpha\)\}\{4\+\\alpha\}\}\\right\\rceil\\right\\\},\\qquad\\gamma=\\gamma\_\{0\}T^\{\-\\frac\{3\+\\alpha\}\{4\+\\alpha\}\},\\qquad\\eta=\\eta\_\{0\}T^\{\-\\frac\{4\}\{4\+\\alpha\}\}\.
#### Phase retrieval instantiation\.
For Figure[1](https://arxiv.org/html/2605.15314#S6.F1), the phase retrieval objective satisfies theα\\alpha\-symmetric generalized smoothness condition withα=2/3\\alpha=2/3\. We useT=10001T=10001,G=1G=1,γ0=10\\gamma\_\{0\}=10for𝖭𝖲𝖦𝖣𝖬\\mathsf\{NSGDM\}, andγ0=7\.5\\gamma\_\{0\}=7\.5,η0=1\\eta\_\{0\}=1for𝖭𝖲𝖳𝖮𝖱𝖬\\mathsf\{NSTORM\}\. Thus, the corresponding hyperparameters are set to
𝖭𝖲𝖦𝖣𝖬:γ≈4\.641×10−3,η≈2\.154×10−3,\\mathsf\{NSGDM:\}\\qquad\\gamma\\approx 4\.641\\times 10^\{\-3\},\\qquad\\eta\\approx 2\.154\\times 10^\{\-3\},and
𝖭𝖲𝖳𝖮𝖱𝖬:γ≈5\.397×10−3,η≈3\.73×10−4,Ninit=4\.\\mathsf\{NSTORM:\}\\qquad\\gamma\\approx 5\.397\\times 10^\{\-3\},\\qquad\\eta\\approx 3\.73\\times 10^\{\-4\},\\qquad N\_\{\\rm init\}=4\.
#### Cubic polynomial instantiation\.
For Figure[2](https://arxiv.org/html/2605.15314#S6.F2), the polynomial experiment is set to satisfy theα\\alpha\-symmetric generalized smoothness condition withα=1/2\\alpha=1/2,T=10001T=10001,G=0\.5G=0\.5, andγ0=η0=1\\gamma\_\{0\}=\\eta\_\{0\}=1\. Thus, the corresponding hyperparameters are set to
𝖭𝖲𝖦𝖣𝖬:γ≈4\.641×10−4,η≈2\.154×10−3\.\\mathsf\{NSGDM:\}\\qquad\\gamma\\approx 4\.641\\times 10^\{\-4\},\\qquad\\eta\\approx 2\.154\\times 10^\{\-3\}\.and
𝖭𝖲𝖳𝖮𝖱𝖬:γ≈7\.742×10−4,η≈2\.782×10−4,Ninit=2\.\\mathsf\{NSTORM:\}\\qquad\\gamma\\approx 7\.742\\times 10^\{\-4\},\\qquad\\eta\\approx 2\.782\\times 10^\{\-4\},\\qquad N\_\{\\rm init\}=2\.Similar Articles
Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization
This paper provides the first comprehensive convergence analysis of vanilla SGD with momentum under heavy-tailed noise without gradient clipping or normalization, revealing inferior rates compared to clipped variants and supported by experiments on synthetic functions.
Unified High-Probability Analysis of Stochastic Variance-Reduced Estimation
This paper presents a unified theoretical framework for stochastic variance-reduced estimation, deriving high-probability bounds via a new Freedman inequality and improving oracle complexities for constrained optimization.
High-Probability PL-SGD with Markovian Noise: Optimal Mixing and Tail Dependence
This paper provides optimal high-probability bounds for stochastic gradient descent under Markovian noise for PL-smooth objectives, closing gaps between expectation and high-probability guarantees and extending to heavy-tailed settings with matching lower bounds.
Zeroth-Order Non-Log-Concave Sampling with Variance Reduction and Applications to Inverse Problems
Proposes a variance-reduced zeroth-order Langevin sampling method for non-log-concave distributions, establishing the first non-asymptotic convergence guarantees, and applies it to inverse problems with score-based generative priors.
Uniform Stability and Generalization Error of GD and SGD on Fixed-Point Parameters
This paper analyzes generalization error, uniform stability, and uniform argument stability of gradient descent (GD) and stochastic gradient descent (SGD) over discrete parameter spaces with deterministic or stochastic rounding, showing that rounding degrades generalization for GD and introduces dimension-dependent errors for stochastic rounding.