Tag
This paper evaluates the Muon optimizer on low-rank matrix factorization, finding it does not consistently outperform AdamW, challenging earlier claims about its advantages in large-scale deep learning.
The paper proposes an attack-agnostic robustness metric based on the spectral norm of the Fisher Information Matrix, providing theoretical bounds and scalable evaluation methods for deep neural networks.
This paper extends the maximal update parameterization (μP) framework to grouped-query attention (GQA), deriving scaling laws for hyperparameter transfer across model architectures. It introduces spectral norm conditions for feature learning and addresses issues with low-rank weight matrices in GQA.