Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
Summary
This paper shows that different optimizers, particularly Muon versus AdamW, induce distinct spectral scaling behaviors in Transformer models, with Muon achieving significantly better utilization of representation capacity, and suggests that optimizer choice should be a first-class axis in scaling laws.
View Cached Full Text
Cached at: 05/22/26, 02:20 AM
Paper page - Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
Source: https://huggingface.co/papers/2605.21803
Abstract
Different optimizers produce distinct spectral scaling behaviors in Transformer models, with Muon achieving superior scaling efficiency compared to AdamW in representation capacity utilization.
Scaling laws have made language-model performance predictable from model size, data, and compute, but they typically treat theoptimizeras a fixed training detail. We show that this assumption misses a fundamental axis of representation scaling: how effectively theoptimizerconverts added FFN width into utilized spectral capacity. Usingeigenspectraoffeed-forward networkrepresentations, measured through soft andhard spectral-ranks, we find that the sameTransformer architecturerealizes markedly differentspectral scaling lawswhen trained with differentoptimizers. Holding architecture and width schedule fixed,AdamWexhibits weak hard-rank scaling (β=0.44) on rare-token (TAIL) representations where learning is known to be hardest, whereasMuonachieves linear scaling (β=1.02) in the same regimes, a 2.3times increase in the scaling exponent. This difference is not reducible to validation loss:AdamWconfigurations can match low-rank Dion variants in perplexity, under extended training, while exhibiting sharply different spectral geometry, demonstrating that matched loss does not imply matched representation structure. Hard--soft rank asymmetry further reveals thatoptimizers differ not only in how much capacity is realized, but also in how that capacity is structured across eigenmodes. To disentangleoptimizereffects from architectural ones, we compare against architectural interventions (e.g.,attention rankandpositional encoding), and find thatoptimizer-induced spectral shifts often exceed the architectural effects. These results suggest optimization as a first-class axis of representation scaling, motivatingoptimizer--architecture co-design.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.21803
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.21803 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.21803 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.21803 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
In-Context Binding Capacity in Language Models
This paper measures the in-context binding capacity of language models, finding that recall scales with parameter count according to a power law and analyzing factors like pretraining recipes and interference.
Stable initialization without the CLT
This paper introduces a uniform-phase initialization method for neural networks with sine activations that eliminates reliance on the Central Limit Theorem, improving stability and performance in representation tasks like image and audio fitting.
Energy-efficient operation of neural operators for virtual sensing
This paper investigates how shared spatial computation in neural operators can reduce energy consumption for virtual sensing applications, particularly in nuclear energy systems, while maintaining prediction accuracy.
GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for \textit{De Novo} Peptide Sequencing
GyroNovo improves de novo peptide sequencing by adaptively imputing missing fragments based on decoder errors and incorporating mass differences via rotary embeddings, achieving state-of-the-art performance on benchmarks.
PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
This paper proposes PolicyAttention, a method using causal softmax attention to implement policy mirror descent for closed-loop control in reinforcement learning, achieving lower losses than existing adaptations in experiments.