sgd

Tag

Cards List
#sgd

@gp_pulipaka: Stochastic Optimizer. #BigData #Analytics #DataScience #AI #MachineLearning #IoT #IIoT #PyTorch #Python #RStats #Tensor…

X AI KOLs Timeline ↗ · 2026-07-28 Cached

This article discusses a research paper showing that the disagreement rate between two deep networks trained with different random seeds can accurately estimate generalization error using only unlabeled data, revealing a surprising connection called Generalization Disagreement Equality.

0 favorites 0 likes
#sgd

Same Loss, Same Noise, Opposite Schedules: Noise Structure and Optimizer Normalization Jointly Determine Whether Learning-Rate Cooldown Helps

arXiv cs.LG ↗ · 2026-07-15 Cached

This paper provably shows that whether learning-rate cooldown helps in WSD schedules depends on the structure of gradient noise and whether the optimizer normalizes its update, explaining why cooldown can be ineffective for SGD but necessary for normalized methods.

0 favorites 0 likes
#sgd

High-Probability PL-SGD with Markovian Noise: Optimal Mixing and Tail Dependence

arXiv cs.LG ↗ · 2026-06-26 Cached

This paper provides optimal high-probability bounds for stochastic gradient descent under Markovian noise for PL-smooth objectives, closing gaps between expectation and high-probability guarantees and extending to heavy-tailed settings with matching lower bounds.

0 favorites 0 likes
#sgd

From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression

arXiv cs.LG ↗ · 2026-05-26 Cached

This paper derives batch scaling laws for sketched linear regression under power-law spectra, analyzing one-pass and multi-pass mini-batch SGD. It provides explicit risk decompositions showing how batch size affects bias, variance, and fluctuation terms, and establishes that without-replacement sampling yields lower noise than with-replacement.

0 favorites 0 likes
#sgd

Population Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise

arXiv cs.LG ↗ · 2026-05-14 Cached

This paper establishes the first population risk bounds for Kolmogorov-Arnold Networks trained with mini-batch SGD and DP-SGD using correlated noise, advancing theoretical understanding of KANs in privacy-sensitive domains.

0 favorites 0 likes
#sgd

Muon is Not That Special: Random or Inverted Spectra Work Just as Well

arXiv cs.LG ↗ · 2026-05-13 Cached

This paper challenges the geometric justification for the Muon optimizer, arguing that precise structure is less important than step-size optimality. It introduces Freon and Kaon optimizers to demonstrate that random or inverted spectra can perform as well as Muon.

0 favorites 0 likes
← Back to home

Submit Feedback