weight-decay

Tag

Cards List
#weight-decay

Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not

arXiv cs.LG · 20h ago Cached

This paper introduces Stiefel Attention, which constrains transformer query and key projection matrices to the Stiefel manifold using Riemannian optimization, demonstrating improved performance on modular arithmetic grokking and CIFAR-10 patches.

0 favorites 0 likes
#weight-decay

Symmetry without a manifold: intrinsic dimension on orbits

arXiv cs.LG · yesterday Cached

The paper demonstrates that standard neural scaling law derivations fail when data forms group orbits, as intrinsic dimension is undefined, leading to exponential rather than power law scaling in model performance.

0 favorites 0 likes
#weight-decay

Thermodynamic Weight Decay: Exploring Grokking Acceleration via Attention Specific Heat

arXiv cs.LG · 2026-07-24 Cached

This paper introduces CvAdamW, an AdamW variant that monitors attention specific heat to detect grokking phase transitions and dynamically scales weight decay, achieving grokking on modular arithmetic where baseline fails.

0 favorites 0 likes
#weight-decay

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

Hugging Face Daily Papers · 2026-05-26 Cached

This paper systematically studies scale vectors in LLM normalization layers, showing they optimize training through a self-amplifying preconditioning effect, and proposes three lightweight improvements that enhance performance and scaling behavior with negligible overhead.

0 favorites 0 likes
#weight-decay

Anytime Training with Schedule-Free Spectral Optimization

arXiv cs.LG · 2026-05-25 Cached

This paper introduces SF-NorMuon, a schedule-free spectral optimizer that matches or exceeds tuned AdamW on language models up to 772M parameters, with theoretical guarantees for stationarity and long-horizon stability.

0 favorites 0 likes
#weight-decay

Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics

arXiv cs.LG · 2026-05-21 Cached

This paper investigates how weight decay acts as a control parameter for transitioning between memorization and generalization in transformers trained on modular arithmetic, and introduces two cheap online diagnostic metrics from attention activations that track these dynamics.

0 favorites 0 likes
#weight-decay

LoRA and Weight Decay (2023)

Hacker News Top · 2026-05-18 Cached

This blog post explores how LoRA's interaction with weight decay leads to a different optimization objective than full fine-tuning, where weights are regularized towards the initial model rather than zero. It explains the implications for practitioners.

0 favorites 0 likes
#weight-decay

Prescriptive Scaling Laws for Data Constrained Training

Hugging Face Daily Papers · 2026-05-02 Cached

A modified scaling law accounting for data repetition effects provides compute-optimal training strategies for data-constrained scenarios, showing that beyond a point further repetition is counterproductive and compute is better spent on model capacity.

0 favorites 0 likes
← Back to home

Submit Feedback