Tag
A tweet highlights that the Muse Glimmer model pins per-matrix weight RMS norm at ~6e-3, linking this observation to the Hyperball optimizer paper, which proposes an optimizer wrapper that fixes weight and update norms to improve pretraining speed.
The paper investigates whether weight norm directly controls the grokking delay in neural networks or if its effect is mediated by logit scale and softmax saturation under cross-entropy loss. Experiments show that the delay is almost entirely explained by the effective logit scale, with weight norm contributing negligibly.
This paper demonstrates that the weight norm causally controls the timescale of grokking in neural networks, reconciling conflicting accounts. Through interventions, it shows that grokking follows an exponential delay law and that norm magnitude dominates grokking time over learning rate across architectures.