stochastic-gradient-descent

Tag

Cards List
#stochastic-gradient-descent

Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise

arXiv cs.LG · 4d ago Cached

The paper proves almost sure convergence of stochastic gradient descent with clipping and additive noise, including momentum variants, under smoothness and bounded gradient noise assumptions, providing theoretical foundations for stable training in convex and nonconvex settings.

0 favorites 0 likes
#stochastic-gradient-descent

Mini-batch Noise Lowers Sharpness via Dominant-Subspace Fluctuations

arXiv cs.LG · 2026-07-28 Cached

This paper argues that the dominant subspace of the Hessian, while contributing little to loss reduction, plays a key role in reducing sharpness during mini-batch SGD. It derives a sharpness correction term induced by mini-batch noise in the dominant directions.

0 favorites 0 likes
#stochastic-gradient-descent

Scaling Limits of Constant-Stepsize SGD at Flat Minima

arXiv cs.LG · 2026-07-21 Cached

This paper analyzes the scaling limits of constant-stepsize SGD near flat minima, showing that the invariant law concentrates at scale α^(1/m) for objectives with flatness exponent m ≥ 2, and converges to non-Gaussian stationary distributions for m > 2.

0 favorites 0 likes
#stochastic-gradient-descent

Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles

arXiv cs.LG · 2026-07-13 Cached

This paper provides worst-case convergence analyses for Schedule-Free gradient descent and stochastic gradient descent in nonconvex optimization, establishing optimal rates and strict-saddle avoidance, thus theoretically justifying their empirical success.

0 favorites 0 likes
#stochastic-gradient-descent

Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization

arXiv cs.LG · 2026-07-10 Cached

This paper provides the first comprehensive convergence analysis of vanilla SGD with momentum under heavy-tailed noise without gradient clipping or normalization, revealing inferior rates compared to clipped variants and supported by experiments on synthetic functions.

0 favorites 0 likes
#stochastic-gradient-descent

Revisiting the Volume Hypothesis

arXiv cs.LG · 2026-07-01 Cached

This paper revisits the volume hypothesis, which posits that generalization in over-parameterized networks is mainly due to the larger volume of good-generalizing regions in weight space rather than SGD's implicit bias. Through experiments with binary networks, the authors show that the generalization advantage of gradient learning over random sampling diminishes as training data size grows, potentially resolving contradictory prior findings.

0 favorites 0 likes
#stochastic-gradient-descent

A Link between Shock-wave Theory and Symmetry-reduced Stochastic Gradient Descent for Artificial Neural Networks

arXiv cs.LG · 2026-06-18 Cached

This paper establishes a mathematically rigorous connection between shock-wave theory and symmetry-quotiented learning dynamics of stochastic gradient descent, showing that after symmetry reduction and coarse-graining, the dynamics satisfy viscous Hamilton-Jacobi and Burgers-type equations with shock formation times controlled by loss curvature.

0 favorites 0 likes
#stochastic-gradient-descent

Uniform Stability and Generalization Error of GD and SGD on Fixed-Point Parameters

arXiv cs.LG · 2026-06-08 Cached

This paper analyzes generalization error, uniform stability, and uniform argument stability of gradient descent (GD) and stochastic gradient descent (SGD) over discrete parameter spaces with deterministic or stochastic rounding, showing that rounding degrades generalization for GD and introduces dimension-dependent errors for stochastic rounding.

0 favorites 0 likes
← Back to home

Submit Feedback