Tag
This paper shows that different optimizers, particularly Muon versus AdamW, induce distinct spectral scaling behaviors in Transformer models, with Muon achieving significantly better utilization of representation capacity, and suggests that optimizer choice should be a first-class axis in scaling laws.