momentum

Tag

Cards List
#momentum

Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization

arXiv cs.LG · 2026-07-10 Cached

This paper provides the first comprehensive convergence analysis of vanilla SGD with momentum under heavy-tailed noise without gradient clipping or normalization, revealing inferior rates compared to clipped variants and supported by experiments on synthetic functions.

0 favorites 0 likes
#momentum

Class-Grouped Normalized Momentum and Faster Hyperparameter Exploration to Tackle Class Imbalance in Federated Learning

arXiv cs.LG · 2026-07-03 Cached

The paper proposes FedCGNM, a client-side optimizer that groups classes and normalizes momentum per group to address class imbalance in federated learning, along with FedHOO for efficient hyperparameter exploration. Empirical results show consistent improvements over baselines.

0 favorites 0 likes
#momentum

@RuujSs: An algorithm that turns $1 into $36 billion over 22 years sounds like a false headline right. It's actually Table 3 of …

X AI KOLs Timeline · 2026-06-26 Cached

A University of Johannesburg working paper proposes a Kalman-filtered, momentum-extended Anticor algorithm (K-ACM) that, in a backtest from 1962–1984, produces an extreme return of $1 to $36 billion, though acknowledged as a backtest artifact due to unrealistic assumptions.

0 favorites 0 likes
#momentum

MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization

arXiv cs.LG · 2026-06-17 Cached

Proposes MGUP, a momentum-gradient alignment update policy for selective intra-layer parameter updates in stochastic optimization, which integrates with optimizers like AdamW, Lion, and Muon, and provides theoretical convergence guarantees along with superior performance on large-scale model training tasks.

0 favorites 0 likes
#momentum

DP-MacAdam: Differentially Private Mechanism with Adaptive Clipping and Adaptive Momentum

arXiv cs.LG · 2026-06-05 Cached

DP-MacAdam combines adaptive clipping and adaptive momentum to improve differentially private SGD, achieving better model utility without manual tuning of the clipping threshold.

0 favorites 0 likes
← Back to home

Submit Feedback