flat-minima

Tag

Cards List
#flat-minima

Noise-aware training for analog hardware: accuracy collapses at a threshold rather than degrading smoothly [D]

Reddit r/MachineLearning ↗ · 2026-08-09

An experiment on analog hardware noise shows accuracy collapses at a threshold rather than degrading smoothly, and noise-aware training shifts that threshold significantly.

0 favorites 0 likes
#flat-minima

Scaling Limits of Constant-Stepsize SGD at Flat Minima

arXiv cs.LG ↗ · 2026-07-21 Cached

This paper analyzes the scaling limits of constant-stepsize SGD near flat minima, showing that the invariant law concentrates at scale α^(1/m) for objectives with flatness exponent m ≥ 2, and converges to non-Gaussian stationary distributions for m > 2.

0 favorites 0 likes
#flat-minima

Leveraging Extragradient for Effective Sharpness-Aware Minimization in Deep Learning

arXiv cs.LG ↗ · 2026-07-08 Cached

Proposes EISAM, a new optimizer that extends Sharpness-Aware Minimization using an extragradient step to find flatter minima, improving generalization and robustness while reducing sensitivity to hyperparameters. Outperforms SGD, Adam, and SAM on benchmarks.

0 favorites 0 likes
#flat-minima

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

arXiv cs.LG ↗ · 2026-06-30 Cached

Derives the closed-form gradient of the Wolkowicz-Styan upper bound on the loss Hessian eigenspectrum to guide neural network training toward flat minima, and introduces Hessian Spectral Range (HSR) Regularization. Numerical experiments show that HSR narrows the Hessian eigenvalue range, avoids sharp minima and saddle points, and achieves flat solutions comparable to Sharpness-Aware Minimization (SAM).

0 favorites 0 likes
#flat-minima

Are Flat Minima an Illusion?

arXiv cs.LG ↗ · 2026-05-08 Cached

This paper challenges the common belief that flat minima cause better generalization in neural networks, arguing that 'weakness'—a reparameterization-invariant measure of function simplicity—is the true driver. Empirical results on MNIST and Fashion-MNIST show that weakness predicts generalization while sharpness anticorrelates, and the large-batch generalization advantage vanishes as training data increases.

0 favorites 0 likes
← Back to home

Submit Feedback