@wen_kaiyue: A hidden detail in the recently released Muse Glimmer model: its per-matrix weight RMS norm is pinned almost exactly at…
摘要
A tweet highlights that the Muse Glimmer model pins per-matrix weight RMS norm at ~6e-3, linking this observation to the Hyperball optimizer paper, which proposes an optimizer wrapper that fixes weight and update norms to improve pretraining speed.
查看缓存全文
缓存时间: 2026/08/11 09:50
A hidden detail in the recently released Muse Glimmer model: its per-matrix weight RMS norm is pinned almost exactly at ~6e-3!
If you’re curious why fixing weight RMS like this can work, feel free to checkout Hyperball: https://t.co/9AFvaJMmpx https://t.co/51tQWdA9st
Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization
Source: https://arxiv.org/html/2606.16899 Kaiyue Wen†\daggerXingyu Dang‡\ddaggerKaifeng Lyu§Tengyu Ma†\daggerPercy Liang†\dagger [email protected]@[email protected] [email protected]@cs.stanford.edu
Abstract
Matrix based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when using standard constant decoupled weight decay. We proposeHyperball, a simple optimizer wrapper that addresses this issue. Given a base optimizer such as Adam or Muon, Hyperball sets the Frobenius norms of weight matrices and their corresponding optimizer updates to fixed constants. On Qwen3 style models up to1.21.2B parameters, Muon Hyperball achieves2020–30%30\%token equivalent speedup over weight decay baselines. Hyperball also improves learning rate transfer across widths and depths compared to decoupled weight decay. This method is motivated by prior theory showing that training with weight decay leads to an equilibrium weight norm that only depends on the training hyperparameters. Through this mechanism, the weight decay then decides the angular learning rate, i.e. how fast the direction of the weight matrix changes.
22footnotetext:Stanford University.‡\ddaggerPrinceton University.§Tsinghua University.## 1Introduction
Previous work observed that the speedups of matrix based optimizers such as Muon(Jordan et al.,2024; Liu et al.,2025)over AdamW(Loshchilov and Hutter,2019)shrink from roughly30%30\%to about10%10\%as model size and data scale grow(Wen et al.,2025). This motivates a simple question: can we keep these optimizer speedups at higher compute?
We introduceHyperballas a simple solution to the above question. Hyperball is an optimizer wrapper that enforces constant weight norms and update norms, transforming any base optimizer into its Hyperball variant. The wrapper is motivated by prior theory on the role of weight decay in scale invariant layers and by the way most modern LLM training uses weight decay to control the size of the weights implicitly. LetWtW_{t}be the parameter matrix at steptt,utu_{t}be the update provided by a base optimizer,ηt\eta_{t}be the learning rate, andλ\lambdabe the weight decay coefficient. The standard decoupled weight decay update(Loshchilov and Hutter,2019)applies
Wt+1=(1−ηtλ)Wt−ηtut.W_{t+1}=(1-\eta_{t}\lambda)\,W_{t}-\eta_{t}\,u_{t}.Here−ηtut-\eta_{t}u_{t}adds new update information and typically increases the weight norm in the absence of weight decay. The term(1−ηtλ)Wt(1-\eta_{t}\lambda)W_{t}softly controls the norm by shrinking the weights toward zero every step. In modern Transformer architectures with normalization layers(Vaswani et al.,2017; Xiong et al.,2020), many weight matrices are modeled as scale invariant in the standard sense at the level of the loss: for a scalarc>0c>0, rescaling one matrix leaves the loss unchanged,L(cW)=L(W)L(cW)=L(W). In this setting, weight decay is puzzling as classicalℓ2\ell_{2}regularization: if the loss is unchanged by the scale ofWW, penalizing‖W‖F\left\lVert W\right\rVert_{\mathrm{F}}cannot be the main reason it improves training.
Hyperball replaces this soft control on weight norm with an explicit constraint. It decouples the magnitude of the weights from the direction of the update. For a matrixXX, the Frobenius norm‖X‖F=(∑ijXij2)1/2\left\lVert X\right\rVert_{\mathrm{F}}=(\sum_{ij}X_{ij}^{2})^{1/2}is the Euclidean norm of its entries. LetR=‖W0‖FR=\left\lVert W_{0}\right\rVert_{\mathrm{F}}be the initial Frobenius norm of the parameter matrix, and letNormalize(X):=X/‖X‖F\mathrm{Normalize}\!\left(X\right):=X/\left\lVert X\right\rVert_{\mathrm{F}}be Frobenius normalization. The Hyperball update is
Wt+1=R⋅Normalize((Wt−ηtR⋅Normalize(ut))).W_{t+1}\;=\;R\cdot\mathrm{Normalize}\!\left(\Big(W_{t}-\eta_{t}\,R\cdot\mathrm{Normalize}\!\left(u_{t}\right)\Big)\right).
Figure 1:Geometric view of the Hyperball update. Weights are constrained to the sphere of radiusRR. Each step moves along the normalized update direction−Normalize(ut)-\mathrm{Normalize}\!\left(u_{t}\right)by a distanceηtR\eta_{t}Rand is immediately projected back to the sphere. The weight norm and update norm are held constant by construction.Geometrically, Hyperball constrains the optimization trajectory to the surface of a hypersphere with radiusRR. The update takes a step of lengthηtR\eta_{t}Rin the direction defined by the normalized update−Normalize(ut)-\mathrm{Normalize}\!\left(u_{t}\right), and the result is immediately projected back onto the sphere. This keeps the norm of the weights and updates constant, so the optimizer navigates primarily through weight directions.
The base updateutu_{t}can come from any optimizer. In this paper we focus on Adam Hyperball (AdamH) and Muon Hyperball (MuonH). We apply Hyperball to Transformer weight matrices and use Adam for embeddings, normalization gains, and other parameters whose norm carries semantic information. On1.21.2B parameter Qwen3 style models(Yang et al.,2025), MuonH achieves2020–30%30\%token equivalent speedup over its weight decay counterpart, whereas MuonW gives only about10%10\%at this scale. Across depth and width sweeps, Hyperball keeps the best learning rate window better than the baseline: the maximal drift is about1.4×1.4\timesfor AdamH and MuonH, compared with22–4×4\timesfor AdamW and MuonW baselines.
The optimization theory insection˜4.2explains why this explicit constraint matches the role that weight decay already plays in scale invariant layers. LetRt=‖Wt‖FR_{t}=\left\lVert W_{t}\right\rVert_{\mathrm{F}}be the Frobenius norm of the parameter matrix, and letW^t=Wt/Rt\widehat{W}_{t}=W_{t}/R_{t}be its direction. The decompositionWt=RtW^tW_{t}=R_{t}\widehat{W}_{t}separates radial norm dynamics from directional dynamics: to first order, the angular movement per step scales with the update norm divided byRtR_{t}. Prior analyses of normalized networks and rotational equilibrium(Li et al.,2020; Roburin et al.,2020; Kosson et al.,2024a)show that, under a noise dominated model, decoupled weight decay balances stochastic norm growth and converges to an equilibrium radius. Substituting this radius into the tangent dynamics yields an angular step sizeηang\eta^{\mathrm{ang}}that depends on the learning rate and weight decay mainly through the productηλ\eta\lambda. Hyperball uses this mechanism directly by fixing the radius and update norm, replacing the indirect calibration ofλ\lambdawith an explicit angular learning rate schedule.
Figure 2:Token equivalent speedup over the AdamW scaling law at1.21.2B parameters across Chinchilla ratios1×1\times–8×8\times. The left panel is adapted fromWen et al. (2025)and uses a setup without QK-Norm. The right panel uses the QK-Norm setup. In both setups, MuonW alone gains≈10%\approx 10\%at this scale, while MuonH sustains2020–30%30\%speedup that grows with training duration.
2Method
Definition.
LetWtW_{t}be the parameter matrix at steptt,utu_{t}be the base optimizer update for this matrix (for example, Adam’s preconditioned update(Kingma and Ba,2015)or Muon’s matrix sign momentum update(Jordan et al.,2024; Liu et al.,2025)),ηt\eta_{t}be the Hyperball learning rate,R>0R>0be the fixed radius, andNormalize(X):=X/‖X‖F\mathrm{Normalize}\!\left(X\right):=X/\left\lVert X\right\rVert_{\mathrm{F}}be Frobenius normalization. For each constrained matrixW0∈ℝdout×dinW_{0}\in\mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}}, we initialize entries with standard deviation1/din1/\sqrt{d_{\mathrm{in}}}and set the radius once asR=‖W0‖FR=\left\lVert W_{0}\right\rVert_{\mathrm{F}}. The Hyperball update is
Wt+1=R⋅Normalize((Wt−ηtR⋅Normalize(ut))).W_{t+1}\;=\;R\cdot\mathrm{Normalize}\!\left(\Big(W_{t}-\eta_{t}\,R\cdot\mathrm{Normalize}\!\left(u_{t}\right)\Big)\right).(1)Thus Hyperball takes an unconstrained optimizer step with normηtR\eta_{t}R, followed by radial renormalization to radiusRR(algorithm˜1,fig.˜1). Writeu^t:=ut/‖ut‖F\widehat{u}_{t}:=u_{t}/\left\lVert u_{t}\right\rVert_{\mathrm{F}}andW^t:=Wt/‖Wt‖F\widehat{W}_{t}:=W_{t}/\left\lVert W_{t}\right\rVert_{\mathrm{F}}. Equivalently, the exact displacement is
Wt+1−Wt=R(W^t−ηtu^t‖W^t−ηtu^t‖F−W^t),W_{t+1}-W_{t}=R\left(\frac{\widehat{W}_{t}-\eta_{t}\widehat{u}_{t}}{\left\lVert\widehat{W}_{t}-\eta_{t}\widehat{u}_{t}\right\rVert_{\mathrm{F}}}-\widehat{W}_{t}\right),(2)so radial renormalization returns the trial point to the sphere of radiusRR.
Algorithm 1Hyperball wrapper for a parameter matrixWW1:Input:parameter matrix
WtW_{t}, base optimizer
𝒪\mathcal{O}, optimizer state
𝒮t\mathcal{S}_{t}, radius
RR, schedule
{ηt}\{\eta_{t}\} 2:Compute base optimizer update
ut,St+1←𝒪(∇WtL(Wt),𝒮t)u_{t},S_{t+1}\leftarrow\mathcal{O}(\nabla_{W_{t}}L(W_{t}),\mathcal{S}_{t}) 3:Set normalized update direction
u^t←Normalize(ut)\widehat{u}_{t}\leftarrow\mathrm{Normalize}\!\left(u_{t}\right) 4:Take unprojected step
W~t+1←Wt−ηtRu^t\widetilde{W}_{t+1}\leftarrow W_{t}-\eta_{t}\,R\,\widehat{u}_{t} 5:Project to radius
RR:
Wt+1←R⋅Normalize(W~t+1)W_{t+1}\leftarrow R\cdot\mathrm{Normalize}\!\left(\widetilde{W}_{t+1}\right)
Where to apply the constraint.
We apply Hyperball to attention and MLP weight matrices in a prenorm Transformer(Vaswani et al.,2017; Xiong et al.,2020). Embeddings, normalization gains, and other scalar parameters are updated with a standard optimizer (Adam in our experiments), since for these parameters the norm can carry semantic information.
Discussion of the design.
A natural alternative to the Frobenius norm constraint is a spectral norm constraint. Let‖W‖op\left\lVert W\right\rVert_{\mathrm{op}}denote the operator norm of a matrix. The spectral condition ofYang et al. (2023)identifies relative spectral update size as a central quantity for feature learning, and SSO constrains training to the spectral sphere by steepest descent or projection(Xie et al.,2026). In this view, a spectral Hyperball variant would fix‖Wt‖op\left\lVert W_{t}\right\rVert_{\mathrm{op}}and normalize the update in operator norm, directly controlling the sharpest layerwise scaling factor.
We use the Frobenius version in (1) for two reasons: computation cost and the theoretical motivation insection˜4. Projection onto the Frobenius sphere isO(N2)O(N^{2})per matrix, whereas exact spectral projection generally requires an SVD and costsO(N3)O(N^{3}). One diagnostic for whether Frobenius control is close to spectral control is the stable rank ratio. For a parameter matrixW∈ℝdout×dinW\in\mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}}, define
ℛ(W):=‖W‖F2‖W‖op2∈[1,min(din,dout)].\mathcal{R}(W)\;:=\;\frac{\left\lVert W\right\rVert_{\mathrm{F}}^{2}}{\left\lVert W\right\rVert_{\mathrm{op}}^{2}}\;\in\;[1,\,\min(d_{\mathrm{in}},d_{\mathrm{out}})].(3)Values satisfyingℛ(W)=Ω(min(din,dout))\mathcal{R}(W)=\Omega(\min(d_{\mathrm{in}},d_{\mathrm{out}}))indicate that the singular value spectrum is not dominated by a single direction, in which case Frobenius constraints behave similarly to spectral constraints up to a slowly varying factor. The same high stable rank regime is observed empirically in the Kimi Moonlight analysis(Liu et al.,2025, Appendix F). A spectral Hyperball variant is a natural direction when singular value concentration makes Frobenius control too loose.
3Hyperball Experiments
3.1Setup
Unless noted otherwise, we use a Qwen3 style decoder only architecture(Yang et al.,2025)with QK-Norm(Henry et al.,2020), trained on a mixture of DCLM-baseline(Li et al.,2024), StarCoder(Li et al.,2023), and ProofPile 2(Azerbayev et al.,2023)(and FineWeb-Edu(Penedo et al.,2024)for some runs). We compare Adam(Kingma and Ba,2015)and Muon(Jordan et al.,2024; Liu et al.,2025)with decoupled weight decay (AdamW and MuonW) against their Hyperball variants AdamH and MuonH. For the speedup metric, we fit a scaling law to AdamW across Chinchilla ratios(Hoffmann et al.,2022){1×,2×,4×,8×}\{1\times,2\times,4\times,8\times\}and report, for each method’s final loss, the token ratioτ=NAdamW/Nmethod\tau=N_{\mathrm{AdamW}}/N_{\mathrm{method}}that AdamW would need to match it. For learning rate transfer, we sweep a multiplicative grid (ratio2\sqrt{2}) at each scaless, defineη⋆(s):=argminηkValLoss(s,ηk;T)\eta^{\star}(s):=\arg\min_{\eta_{k}}\ \mathrm{ValLoss}(s,\eta_{k};T), and reportDrift:=maxsη⋆(s)/minsη⋆(s)\mathrm{Drift}:=\max_{s}\eta^{\star}(s)\,/\,\min_{s}\eta^{\star}(s).
3.2End-to-end speedup
In the weight decay baseline setting adopted fromWen et al. (2025),1.21.2B parameter Qwen3 style models are trained over Chinchilla ratios1×1\times–8×8\times, and MuonW alone yields≈10%\approx 10\%token equivalent speedup over the AdamW scaling law. MuonH instead sustains2020–30%30\%speedup, and the gapgrowswith training duration (fig.˜2). Qualitatively, Hyperball starts slightly worse but overtakes WD as the learning rate decays. On the Marin speedrun benchmark (FineWeb-Edu,1×1\timesChinchilla), AdamH and MuonH match WD baselines that are≈10%\approx 10\%larger (fig.˜3, left). In Marin Ferries, scaling MuonH to88B parameters yields a further0.040.04loss improvement over the AdamW baseline, with both runs using manually chosen hyperparameters (fig.˜3, middle and right).
Figure 3:Additional Marin benchmarks. Left: final C4/en loss(Raffel et al.,2020)on the FineWeb-Edu speedrun benchmark at1×1\timesChinchilla. Middle and right:88B model comparison over159159B tokens. MuonH fixes the layer 9 value projection matrix norm and finishes0.040.04lower than the AdamW baseline.
Figure 4:The modded-nanogpt Track 3 optimization benchmark. Curves show average validation loss from the public Track 3 logs versus training step. Lower and further left is better. Left: full trajectories. Right: zoom near the3.273.27–3.283.28validation loss band for entries reaching this band within34003400steps. Hyperball variants improve the matched weight decay baselines in this comparison, and KL-SOAP-H, which denotes KL-SOAP(Lin et al.,2026)combined with Hyperball, reaches average validation loss3.27803.2780in31253125steps.On the public modded-nanogpt Track 3 optimization benchmark, which fixes the model and data and measures optimizer progress by step count, the corresponding WD baselines reach average validation loss3.27903.2790in56255625steps for the single run AdamW baseline,3.27903.2790in33253325steps for MuonW (tuned), and3.27893.2789in32503250steps for NorMuonW(Li et al.,2025). AdamH reaches average validation loss3.27413.2741in48754875steps. MuonH reaches average validation loss3.27823.2782in33253325steps, NorMuonH reaches average validation loss3.27783.2778in32503250steps, and KL-SOAP-H(Lin et al.,2026)reaches average validation loss3.27803.2780in31253125steps (fig.˜4)(Keller and Contributors,2026). This result shows that Hyperball is not tied to Muon or Adam: paired with KL-SOAP, it reaches the fastest loss at step result in this comparison.
3.3Hyperparameter transfer
By construction, Hyperball fixes‖Wt‖F=R\left\lVert W_{t}\right\rVert_{\mathrm{F}}=Rand uses unit Frobenius update directions, soηt\eta_{t}directly sets the relative update length. The optimal learning rate should therefore be approximately scale invariant. We test this in two sweeps with1010B tokens per run. These transfer runs use a hybrid normalization architecture variant(Zhuo et al.,2025), with QK-Norm enabled.
Figure 5:Depth scaling at fixedd=128d=128,1010B tokens per run. Each curve shows final validation loss versus learning rate for a given depth. Stars mark the best learning rate. Hyperball variants reduce the optimal learning rate drift acrossL∈[4,512]L\in[4,512]to≈1.4×\approx 1.4\times, versus22–4×4\timesfor AdamW and MuonW.For depth scaling at fixed hidden dimensiond=128d=128andL∈{4,…,512}L\in\{4,\dots,512\}, the maximal drift of the optimal learning rate is≈1.4×\approx 1.4\timesfor AdamH and MuonH, versus≈3×\approx 3\timesfor AdamW and≈4×\approx 4\timesfor MuonW even atL=512L=512(fig.˜5).
Figure 6:Width scaling at fixedL=4L=4,1010B tokens per run. Hyperball variants reduce the optimal learning rate drift acrossd∈[128,2048]d\in[128,2048]to≈1.4×\approx 1.4\times.For width scaling at fixed depthL=4L=4andd∈{128,…,2048}d\in\{128,\dots,2048\}, the same≈1.4×\approx 1.4\timesdrift holds for both Hyperball variants (fig.˜6).
3.4Overtrained Setting
We also test Hyperball in an overtrained data scaling setting. For a130130M parameter model, we train MuonW and MuonH over token budgets from11B to128128B and sweep the learning rate at each budget. Both MuonW and MuonH use the same hybrid normalization architecture variant as in previous section. MuonH attains lower best C4 validation loss across the full range (fig.˜7). FittingL(N)=L∞+AN−αL(N)=L_{\infty}+AN^{-\alpha}to the best loss curve givesL∞=3.065L_{\infty}=3.065for MuonH versus3.0793.079for MuonW.
Figure 7:Overtrained130130M data scaling sweep. Points show the best C4 validation loss over learning rate sweeps at each token budget. Solid curves fitL(N)=L∞+AN−αL(N)=L_{\infty}+AN^{-\alpha}. Dashed lines mark the fitted asymptotes.
4Theory
4.1Expressivity
We first show that, in normalized networks, fixing the Frobenius norm of a weight matrix wouldn’t limit representation power. This is because a trainable normalization gain can absorb the scale of the weight matrix, so Hyperball changes the optimization geometry without removing represented functions. Lethhbe the hidden state,WWbe a weight matrix, andγ\gammabe the RMSNorm gain(Zhang and Sennrich,2019). For a a linear map placed after a layer norm,
f(h;W,γ)=W(γ⊙RMSNorm(h)).f(h;W,\gamma)\;=\;W\,(\gamma\odot\mathrm{RMSNorm}(h)).(4)For any scalarc>0c>0, the joint rescaling(W,γ)↦(cW,γ/c)(W,\gamma)\mapsto(cW,\gamma/c)leaves the represented function unchanged:
f(h;cW,γ/c)=f(h;W,γ).f(h;cW,\gamma/c)=f(h;W,\gamma).(5)Thus constraining‖W‖F\left\lVert W\right\rVert_{\mathrm{F}}need not reduce the represented function class when a trainable normalization gain can absorb the scale.
4.2Optimization
WithWWbeing a parameter block and𝒲c\mathcal{W}^{c}being the remaining parameters in the neural network. A lossL(W,𝒲c)L(W,\mathcal{W}^{c})isscale invariantin a parameter blockWWifL(cW,𝒲c)=L(W,𝒲c)L(cW,\mathcal{W}^{c})=L(W,\mathcal{W}^{c})for allc>0c>0. We will drop𝒲c\mathcal{W}^{c}and use the shorthandL(W)L(W)from now on. This is the notion used throughout the analysis: the radial coordinate‖W‖F\left\lVert W\right\rVert_{\mathrm{F}}is redundant, and the loss depends only on the directionW^=W/‖W‖F\widehat{W}=W/\left\lVert W\right\rVert_{\mathrm{F}}. Many weight matrices in modern LLMs are only approximately scale invariant in this loss level sense, but the exact scale invariant model is a useful local approximation for the norm dynamics below.
Decoupled weight decay and angular motion.
LetWtW_{t}be the parameter matrix at steptt,utu_{t}be the base optimizer update for this matrix,ηt\eta_{t}be the learning rate,λ\lambdabe the weight decay coefficient, andαt:=1−ηtλ\alpha_{t}:=1-\eta_{t}\lambda. The standard decoupled weight decay update(Loshchilov and Hutter,2019)is
Wt+1=αtWt−ηtut.W_{t+1}=\alpha_{t}W_{t}-\eta_{t}u_{t}.(6)LetRt:=‖Wt‖FR_{t}:=\left\lVert W_{t}\right\rVert_{\mathrm{F}}andW^t:=Wt/Rt\widehat{W}_{t}:=W_{t}/R_{t}. The angular step size is the per step movement on the unit sphere,
ηtang:=‖W^t+1−W^t‖F,\eta^{\mathrm{ang}}_{t}:=\left\lVert\widehat{W}_{t+1}-\widehat{W}_{t}\right\rVert_{\mathrm{F}},(7)and the relative update ratio is
ρt:=‖ut‖F/‖Wt‖F.\rho_{t}:=\left\lVert u_{t}\right\rVert_{\mathrm{F}}/\left\lVert W_{t}\right\rVert_{\mathrm{F}}.(8)For scale invariantLL, the function represented by the block is determined byW^t\widehat{W}_{t}, soηtang\eta^{\mathrm{ang}}_{t}is the optimizer-controlled quantity that determines how fast this block moves in function space. The direction after one decoupled step is exactly
W^t+1=αtW^t−ηtut/Rt‖αtW^t−ηtut/Rt‖F.\widehat{W}_{t+1}=\frac{\alpha_{t}\widehat{W}_{t}-\eta_{t}u_{t}/R_{t}}{\left\lVert\alpha_{t}\widehat{W}_{t}-\eta_{t}u_{t}/R_{t}\right\rVert_{\mathrm{F}}}.(9)Thus, for fixed‖ut‖F\left\lVert u_{t}\right\rVert_{\mathrm{F}}, a larger radiusRtR_{t}gives a smaller angular movement. Weight decay therefore acts as an indirect angular step controller by regulatingRtR_{t}.
Concrete base updates.
The analysis below uses AdamW, Muon, and Moonlight scaled Muon as examples. Letℓt\ell_{t}be the minibatch loss,gt:=∇Wℓt(Wt)g_{t}:=\nabla_{W}\ell_{t}(W_{t})be the stochastic gradient for the matrix block, letϵ>0\epsilon>0be Adam’s numerical stability constant, and let divisions and square roots in Adam be elementwise. For AdamW(Loshchilov and Hutter,2019),
mt=β1mt−1+(1−β1)gt,vt=β2vt−1+(1−β2)gt⊙2,ut=m^tv^t+ϵ,m_{t}=\beta_{1}m_{t-1}+(1-\beta_{1})g_{t},\qquad v_{t}=\beta_{2}v_{t-1}+(1-\beta_{2})g_{t}^{\odot 2},\qquad u_{t}=\frac{\widehat{m}_{t}}{\sqrt{\widehat{v}_{t}}+\epsilon},(10)wherem^t\widehat{m}_{t}andv^t\widehat{v}_{t}denote the bias corrected moments. For Muon(Jordan et al.,2024; Liu et al.,2025), let
Mt=β1Mt−1+(1−β1)gt.M_{t}=\beta_{1}M_{t-1}+(1-\beta_{1})g_{t}.(11)For a matrixAAwith compact singular value decompositionA=PΣQ⊤A=P\Sigma Q^{\top}, define the exact SVD matrix sign map by
msign(A):=PQ⊤.\operatorname{msign}(A):=PQ^{\top}.(12)Muon implementations often compute this map by Newton–Schulz iteration; in this theory we analyze the idealized SVD Muon update. ForW∈ℝdout×dinW\in\mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}}, definesμ:=max(1,dout/din)s_{\mu}:=\max(1,\sqrt{d_{\mathrm{out}}/d_{\mathrm{in}}}). The Muon base update is
ut=sμmsign(Mt).u_{t}=s_{\mu}\,\operatorname{msign}(M_{t}).(13)Moonlight(Liu et al.,2025)uses the same momentum and exact SVD sign map, but replacessμs_{\mu}withsmoon:=0.2max(din,dout)s_{\mathrm{moon}}:=0.2\sqrt{\max(d_{\mathrm{in}},d_{\mathrm{out}})}:
ut=smoonmsign(Mt).u_{t}=s_{\mathrm{moon}}\,\operatorname{msign}(M_{t}).(14)
Idealized stationary model.
We first make an assumption that assume we have an infinite history of gradient. This allows us to ignore boundary conditions on gradient when we consider momentum and weights.
Assumption 4.1(Infinite history optimizer).
The gradient sequence{gt}t∈ℤ\{g_{t}\}_{t\in\mathbb{Z}}and the weight sequence{Wt}t∈ℤ\{W_{t}\}_{t\in\mathbb{Z}}are defined for all integer times, includingt<0t<0. Optimizer states are computed from this infinite past, i.e.,
mt=(1−β1)∑i≥0β1igt−i,Mt=(1−β1)∑i≥0β1igt−i.m_{t}=(1-\beta_{1})\sum_{i\geq 0}\beta_{1}^{i}g_{t-i},\qquad M_{t}=(1-\beta_{1})\sum_{i\geq 0}\beta_{1}^{i}g_{t-i}.(15)In the constant learning rate calculation below,WtW_{t}is also taken to be the solution obtained by running (6) from the infinite past. Intuitively, this describes the regime where training has already run for a long time, so the dependence on the initial optimizer state and initial weight has decayed.
We then make the following assumption on the distribution ofgtg_{t}.
Assumption 4.2(Isotropic stationary gradients and idealized base maps).
Letp=doutp=d_{\mathrm{out}},q=dinq=d_{\mathrm{in}},d=pqd=pq, andr=min(p,q)r=\min(p,q). After vectorizing the matrix block, the stochastic optimizer input is an iid isotropic Gaussian sequence,
gt∼𝒩(0,σ2Id),{gt}t∈ℤindependent.g_{t}\sim\mathcal{N}(0,\sigma^{2}I_{d}),\qquad\{g_{t}\}_{t\in\mathbb{Z}}\text{ independent.}(16)
Assumption˜4.2is intentionally idealized. It ignores anisotropy, layer specific structure, and the signal in𝔼[gt]\mathbb{E}[g_{t}]. Its purpose is to provide an easy to compute ansatz for update norms, update autocorrelations, radial equilibria, and angular step sizes.
For AdamW, we consider an idealized AdamW update by assuming that the second moment is correctly estimated. For a scalar coordinateg¯t\bar{g}_{t}with𝔼[g¯t2]=σ2\mathbb{E}[\bar{g}_{t}^{2}]=\sigma^{2}, the stationary Adam second moment satisfies
𝔼[vt]=(1−β2)∑i≥0β2i𝔼[g¯t−i2]=(1−β2)∑i≥0β2iσ2=σ2.\mathbb{E}[v_{t}]=(1-\beta_{2})\sum_{i\geq 0}\beta_{2}^{i}\,\mathbb{E}[\bar{g}_{t-i}^{2}]=(1-\beta_{2})\sum_{i\geq 0}\beta_{2}^{i}\,\sigma^{2}=\sigma^{2}.(17)We therefore replace the Adam denominator byσ\sigmaand ignore the vanishing bias correction transient:
ut=mtσ,mt=(1−β1)∑i≥0β1igt−i.u_{t}=\frac{m_{t}}{\sigma},\qquad m_{t}=(1-\beta_{1})\sum_{i\geq 0}\beta_{1}^{i}g_{t-i}.(18)
Update norm and autocorrelation.
We will first study the correlation between updates at different steps, referred to as update autocorrelation.For Muon, the update autocorrelation is related to the following matrix sign map.
Forρ∈[−1,1]\rho\in[-1,1], define the SVD Muon sign kernel
κp,q(ρ):=1min(p,q)𝔼[⟨msign(X),msign(ρX+1−ρ2Z)⟩],\kappa_{p,q}(\rho):=\frac{1}{\min(p,q)}\,\mathbb{E}\Big[\left\langle\operatorname{msign}(X),\,\operatorname{msign}(\rho X+\sqrt{1-\rho^{2}}\,Z)\right\rangle\Big],(19)whereX,Z∈ℝp×qX,Z\in\mathbb{R}^{p\times q}have iid𝒩(0,1)\mathcal{N}(0,1)entries and are independent. The normalization givesκp,q(1)=1\kappa_{p,q}(1)=1andκp,q(0)=0\kappa_{p,q}(0)=0.
Lemma 4.3(Update norm and update autocorrelation).
Underassumptions˜4.1and4.2, define the update scaleUUby
U={1−β11+β1pq(idealized AdamW),p(SVD Muon),0.2pq(Moonlight scaled SVD Muon),U=\begin{cases}\sqrt{\dfrac{1-\beta_{1}}{1+\beta_{1}}}\,\sqrt{pq}&\text{(idealized AdamW)},\\[6.99997pt] \sqrt{p}&\text{(SVD Muon)},\\[2.5pt] 0.2\sqrt{pq}&\text{(Moonlight scaled SVD Muon)},\end{cases}(20)Define the normalized autocorrelation sequence byc0=1c_{0}=1and, for every lagh≥1h\geq 1,
ch={β1h(idealized AdamW),κp,q(β1h)(SVD Muon and Moonlight scaled SVD Muon).c_{h}=\begin{cases}\beta_{1}^{h}&\text{(idealized AdamW)},\\[2.5pt] \kappa_{p,q}(\beta_{1}^{h})&\text{(SVD Muon and Moonlight scaled SVD Muon)}.\end{cases}(21)For every lagh≥1h\geq 1, these definitions give the second moment identities
𝔼‖ut‖F2=U2,𝔼⟨ut,ut−h⟩=U2ch.\mathbb{E}\left\lVert u_{t}\right\rVert_{\mathrm{F}}^{2}=U^{2},\qquad\mathbb{E}\left\langle u_{t},\,u_{t-h}\right\rangle=U^{2}c_{h}.(22)
Proof.
For AdamW, each coordinate of (18) is a stationary Gaussian moving average. Ifg¯t\bar{g}_{t}is one coordinate, then
u¯t=(1−β1)∑i≥0β1ig¯t−iσ.\bar{u}_{t}=(1-\beta_{1})\sum_{i\geq 0}\beta_{1}^{i}\frac{\bar{g}_{t-i}}{\sigma}.(23)Thus𝔼[u¯t2]=(1−β1)/(1+β1)\mathbb{E}[\bar{u}_{t}^{2}]=(1-\beta_{1})/(1+\beta_{1}). For lagh≥1h\geq 1,
𝔼[u¯tu¯t−h]=1−β11+β1β1h.\mathbb{E}[\bar{u}_{t}\bar{u}_{t-h}]=\frac{1-\beta_{1}}{1+\beta_{1}}\,\beta_{1}^{h}.(24)Summing overd=pqd=pqindependent coordinates gives the AdamW line of (20) and both identities in (22), withch=β1hc_{h}=\beta_{1}^{h}.
For SVD Muon,msign(A)\operatorname{msign}(A)has exactlyr=min(p,q)r=\min(p,q)nonzero singular values, all equal to11. Therefore
‖sμmsign(Mt)‖F2=sμ2r=max(1,p/q)min(p,q)=p,\left\lVert s_{\mu}\operatorname{msign}(M_{t})\right\rVert_{\mathrm{F}}^{2}=s_{\mu}^{2}r=\max(1,p/q)\min(p,q)=p,(25)which givesU=pU=\sqrt{p}. For Moonlight, the same calculation gives
‖smoonmsign(Mt)‖F2=0.04max(p,q)min(p,q)=0.04pq,\left\lVert s_{\mathrm{moon}}\operatorname{msign}(M_{t})\right\rVert_{\mathrm{F}}^{2}=0.04\max(p,q)\min(p,q)=0.04pq,(26)soU=0.2pqU=0.2\sqrt{pq}.
It remains to identify the autocorrelation. The stationary momentum matrices satisfy
Mt=(1−β1)∑i≥0β1igt−i.M_{t}=(1-\beta_{1})\sum_{i\geq 0}\beta_{1}^{i}g_{t-i}.(27)Hence each pair(Mt,Mt−h)(M_{t},M_{t-h})is jointly Gaussian, with identical marginal covariance and entrywise correlationβ1h\beta_{1}^{h}using standard property of geometric sequences. After dividing both matrices by their common standard deviation, the pair has the same distribution as
(X,β1hX+1−β12hZ),\bigl(X,\,\beta_{1}^{h}X+\sqrt{1-\beta_{1}^{2h}}\,Z\bigr),(28)withX,ZX,Zas in (19). Because the matrix sign is invariant to positive scalar rescaling, the normalized expected inner product of the two SVD Muon updates is exactlyκp,q(β1h)\kappa_{p,q}(\beta_{1}^{h}). Multiplying bysμ2r=U2s_{\mu}^{2}r=U^{2}orsmoon2r=U2s_{\mathrm{moon}}^{2}r=U^{2}gives the autocorrelation identity in (22). The scalar multipliersμs_{\mu}orsmoons_{\mathrm{moon}}cancels in the normalized autocorrelation, so Muon and Moonlight share the samechc_{h}. ∎
Figure 8:Monte Carlo estimate ofκp,q(ρ)\kappa_{p,q}(\rho)forp=q=100p=q=100.One question is what the range is forκp,q\kappa_{p,q}. LetF(X):=r−1/2msign(X)F(X):=r^{-1/2}\operatorname{msign}(X). SinceF(−X)=−F(X)F(-X)=-F(X)andXXis symmetric,𝔼[F(X)]=0\mathbb{E}[F(X)]=0and𝔼‖F(X)‖F2=1\mathbb{E}\left\lVert F(X)\right\rVert_{\mathrm{F}}^{2}=1. By the Hermite expansion of the Gaussian noise operator,
κp,q(ρ)=∑k≥1akρk,ak≥0,∑k≥1ak=1.\kappa_{p,q}(\rho)=\sum_{k\geq 1}a_{k}\rho^{k},\qquad a_{k}\geq 0,\qquad\sum_{k\geq 1}a_{k}=1.(29)Therefore, for0≤ρ≤10\leq\rho\leq 1,
0≤κp,q(ρ)≤ρ.0\leq\kappa_{p,q}(\rho)\leq\rho.(30) This inequality will be used later to show the correlation between weight and update is bounded.Figure˜8illustrates this range in a numerical simulation withp=q=100p=q=100.
Weight update correlation.
We will now consider constant hyperparameter training, the case where we use constant learning rateη\etaand constant weight decayλ\lambda. We denote1−ηλ1-\eta\lambdaasα\alpha, and assume0<α<10<\alpha<1.
Lemma 4.4(Stationary projection coefficient).
Underassumptions˜4.1and4.2and constant hyperparameter training, let the normalized update autocorrelation sequencechc_{h}be given by (21). Let
Cα:=∑h=1∞αh−1ch.C_{\alpha}:=\sum_{h=1}^{\infty}\alpha^{h-1}c_{h}.(31) Defineγt:=𝔼⟨Wt,ut⟩/U2\gamma_{t}:=\mathbb{E}\left\langle W_{t},\,u_{t}\right\rangle/U^{2}to be the projection coefficient, quantifying how strongly the weight and update correlates, thenγt\gamma_{t}identically equal to a constant valueγ\gammafor alltt, with
γ=−ηCα.\gamma=-\eta C_{\alpha}.(32)
Proof.
Since0<α<10<\alpha<1, the stationary solution of the decoupled weight decay recursion (6) is
Wt=−η∑h=1∞αh−1ut−h.W_{t}=-\eta\sum_{h=1}^{\infty}\alpha^{h-1}u_{t-h}.(33)Taking the expectation of the inner product withutu_{t}and using (22) gives
𝔼⟨Wt,ut⟩=−η∑h=1∞αh−1𝔼⟨ut−h,ut⟩=−ηU2∑h=1∞αh−1ch.\mathbb{E}\left\langle W_{t},\,u_{t}\right\rangle=-\eta\sum_{h=1}^{\infty}\alpha^{h-1}\mathbb{E}\left\langle u_{t-h},\,u_{t}\right\rangle=-\eta U^{2}\sum_{h=1}^{\infty}\alpha^{h-1}c_{h}.(34)Dividing byU2U^{2}and using the definition ofCαC_{\alpha}in (31) proves (32). ∎
For AdamW,ch=β1hc_{h}=\beta_{1}^{h}, and the sum can be simplified,
CαAdam\displaystyle C_{\alpha}^{\mathrm{Adam}}=∑h=1∞αh−1β1h\displaystyle=\sum_{h=1}^{\infty}\alpha^{h-1}\beta_{1}^{h}(35)=β1∑h=1∞(αβ1)h−1=β11−αβ1.\displaystyle=\beta_{1}\sum_{h=1}^{\infty}(\alpha\beta_{1})^{h-1}=\frac{\beta_{1}}{1-\alpha\beta_{1}}.For SVD Muon and Moonlight,
CαMuon=CαMoonlight=∑h=1∞αh−1κp,q(β1h).C_{\alpha}^{\mathrm{Muon}}=C_{\alpha}^{\mathrm{Moonlight}}=\sum_{h=1}^{\infty}\alpha^{h-1}\kappa_{p,q}(\beta_{1}^{h}).(36)The series is finite because (30) givesch≤β1hc_{h}\leq\beta_{1}^{h}. The negative sign means that, in stationarity, the current update is negatively correlated with the current weight. The size of this negative correlation is controlled byCαC_{\alpha}. Substituting (35) into (32) recovers the familiar AdamW expression
γ=−ηCαAdam=−ηβ11−αβ1=−ηβ11−αβ1.\gamma=-\eta C_{\alpha}^{\mathrm{Adam}}=-\eta\frac{\beta_{1}}{1-\alpha\beta_{1}}=-\frac{\eta\beta_{1}}{1-\alpha\beta_{1}}.(37)
Equilibrium weight norm.
Let stationary weight normSt:=𝔼‖Wt‖F2S_{t}:=\mathbb{E}\left\lVert W_{t}\right\rVert_{\mathrm{F}}^{2}and letU2U^{2}be the update second moment in (22). Squaring (6), taking expectations, and substituting𝔼⟨Wt,ut⟩=γU2\mathbb{E}\left\langle W_{t},\,u_{t}\right\rangle=\gamma U^{2}gives the following equality:
St+1=α2St+η2U2−2αηγU2.S_{t+1}=\alpha^{2}S_{t}+\eta^{2}U^{2}-2\alpha\eta\gamma U^{2}.(38)At stationarity,St+1=St=R⋆2S_{t+1}=S_{t}=R_{\star}^{2}, and (32) gives
R⋆=ηU1+2αCα1−α2.R_{\star}=\eta U\sqrt{\frac{1+2\alpha C_{\alpha}}{1-\alpha^{2}}}.(39)For AdamW, (35) reduces this to the previous closed form
R⋆Adam=ηU1+αβ1(1−α2)(1−αβ1).R_{\star}^{\mathrm{Adam}}=\eta U\sqrt{\frac{1+\alpha\beta_{1}}{(1-\alpha^{2})(1-\alpha\beta_{1})}}.(40)For SVD Muon and Moonlight, the correct formula is instead (39) withCαC_{\alpha}from (36). In all cases, whenηλ\eta\lambdais small andβ1\beta_{1}is fixed,R⋆=Θ(Uη/λ)R_{\star}=\Theta(U\sqrt{\eta/\lambda})up to the autocorrelation factor1+2αCα\sqrt{1+2\alpha C_{\alpha}}. Closely related estimates for AdamW update and weight RMS based on mean field approximation appear inSu (2025b,c,d).
Corollary 4.5(Cosine and angular step at equilibrium).
Underassumptions˜4.1and4.2and constant hyperparameter training, the following cosine proxy between the weight and updatecost:=𝔼⟨Wt,ut⟩R⋆U\cos_{t}:=\frac{\mathbb{E}\left\langle W_{t},\,u_{t}\right\rangle}{R_{\star}U}identically equal to a constant valuecos⋆\cos_{\star}for alltt, with
cos⋆=−Cα1−α21+2αCα.\cos_{\star}=-C_{\alpha}\sqrt{\frac{1-\alpha^{2}}{1+2\alpha C_{\alpha}}}.(41)The corresponding ansatz angular step size at equilibrium is
(ηang)2=2(1−α)(1−(1−α)Cα)1+2αCα.(\eta^{\mathrm{ang}})^{2}=\frac{2(1-\alpha)\bigl(1-(1-\alpha)C_{\alpha}\bigr)}{1+2\alpha C_{\alpha}}.(42)For AdamW, this becomes
ηang=2(1−α)(1−β1)1+αβ1=2ηλ(1−β1)1+(1−ηλ)β1.{\eta^{\mathrm{ang}}}=\sqrt{\frac{2(1-\alpha)(1-\beta_{1})}{1+\alpha\beta_{1}}}=\sqrt{\frac{2\eta\lambda(1-\beta_{1})}{1+(1-\eta\lambda)\beta_{1}}}.(43)
Proof.
The cosine formula follows from𝔼⟨Wt,ut⟩=γU2\mathbb{E}\left\langle W_{t},\,u_{t}\right\rangle=\gamma U^{2},γ=−ηCα\gamma=-\eta C_{\alpha}, and (39):
cos⋆=γUR⋆=−Cα1−α21+2αCα.\cos_{\star}=\frac{\gamma U}{R_{\star}}=-C_{\alpha}\sqrt{\frac{1-\alpha^{2}}{1+2\alpha C_{\alpha}}}.(44)For the angular step, plug𝔼⟨Wt,ut⟩=γU2\mathbb{E}\left\langle W_{t},\,u_{t}\right\rangle=\gamma U^{2}andRt=R⋆R_{t}=R_{\star}into (9), and setk⋆:=U/R⋆k_{\star}:=U/R_{\star}. This gives
⟨W^t+1,W^t⟩=α−ηγk⋆2α2−2αηγk⋆2+η2k⋆2.\left\langle\widehat{W}_{t+1},\,\widehat{W}_{t}\right\rangle=\frac{\alpha-\eta\gamma k_{\star}^{2}}{\sqrt{\alpha^{2}-2\alpha\eta\gamma k_{\star}^{2}+\eta^{2}k_{\star}^{2}}}.(45)At equilibrium, the denominator equals11by the stationary radius equation. Usingk⋆2=(1−α2)/(η2(1+2αCα))k_{\star}^{2}=(1-\alpha^{2})/(\eta^{2}(1+2\alpha C_{\alpha}))andγ=−ηCα\gamma=-\eta C_{\alpha}gives
(ηang)2=2−2(α+Cα(1−α2)1+2αCα)=2(1−α)(1−(1−α)Cα)1+2αCα.(\eta^{\mathrm{ang}})^{2}=2-2\left(\alpha+\frac{C_{\alpha}(1-\alpha^{2})}{1+2\alpha C_{\alpha}}\right)=\frac{2(1-\alpha)\bigl(1-(1-\alpha)C_{\alpha}\bigr)}{1+2\alpha C_{\alpha}}.(46)SubstitutingCα=β1/(1−αβ1)C_{\alpha}=\beta_{1}/(1-\alpha\beta_{1})gives (43). ∎
Theorem 4.6(Weight decay sets equilibrium radius and angular step size).
Underassumptions˜4.1and4.2and constant hyperparameter training withα=1−ηλ\alpha=1-\eta\lambda, the weight in the decoupled weight decay update (6) will converge to the equilibrium radius (39). At this equilibrium, the angular step size is given by (42), with the AdamW specialization (43). The base optimizer enters through two quantities only: the update normUUand the autocorrelation sumCαC_{\alpha}.
Proof.
Bylemma˜4.4,𝔼⟨Wt,ut⟩=γU2\mathbb{E}\left\langle W_{t},\,u_{t}\right\rangle=\gamma U^{2}withγ=−ηCα\gamma=-\eta C_{\alpha}. The radial recursion (38) then gives (39), andcorollary˜4.5gives the angular step size. ∎
Underassumptions˜4.1and4.2, the optimizer enters through the identities in (22). OnceUUandchc_{h}are fixed, the weight decay recursion determines the stationary radius, cosine proxy, and angular step algebraically.
Table˜1consolidates the stationary quantities. The update normUUcontrols the equilibrium radius. The autocorrelation sumCαC_{\alpha}controls the momentum-induced radial correction, the cosine, and the angular step. AdamW hasCα=β1/(1−αβ1)C_{\alpha}=\beta_{1}/(1-\alpha\beta_{1}), whereas SVD Muon and Moonlight use the matrix sign kernel in (36).
Table 1:Steady state ansatz quantities forα:=1−ηλ\alpha:=1-\eta\lambda, withCαC_{\alpha}defined in (31). AdamW uses the raw momentum autocorrelation. SVD Muon and Moonlight use the exact matrix sign autocorrelation kernelκp,q\kappa_{p,q}defined in (19), and therefore need not have the same angular step as AdamW.###### Lemma 4.7(Inverse gradient scaling for scale invariant losses).
Independently ofassumptions˜4.1and4.2, ifL:ℝdout×din→ℝL:\mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}}\to\mathbb{R}is differentiable and scale invariant,L(cW)=L(W)L(cW)=L(W)for allc>0c>0, then
∇WL(cW)=1c∇WL(W)for allc>0.\nabla_{W}L(cW)=\frac{1}{c}\nabla_{W}L(W)\qquad\text{for all }c>0.(47)
Proof.
Scale invariance givesL(cW+cϵ)=L(W+ϵ)L(cW+c\epsilon)=L(W+\epsilon)for every matrixϵ\epsilon. Differentiating both sides with respect toϵ\epsilonatϵ=0\epsilon=0gives
⟨∇WL(cW),cϵ⟩=⟨∇WL(W),ϵ⟩∀ϵ,\left\langle\nabla_{W}L(cW),\,c\epsilon\right\rangle=\left\langle\nabla_{W}L(W),\,\epsilon\right\rangle\qquad\forall\epsilon,(48)which forcesc∇WL(cW)=∇WL(W)c\nabla_{W}L(cW)=\nabla_{W}L(W). ∎
Applyinglemma˜4.7toWt=RtW^tW_{t}=R_{t}\widehat{W}_{t}gives
‖∇WL(Wt)‖F=1Rt‖∇WL(W^t)‖F.\left\lVert\nabla_{W}L(W_{t})\right\rVert_{\mathrm{F}}=\frac{1}{R_{t}}\left\lVert\nabla_{W}L(\widehat{W}_{t})\right\rVert_{\mathrm{F}}.(49)Combined withR⋆∝η/λR_{\star}\propto\sqrt{\eta/\lambda}from (39), this gives the scale invariant prediction‖∇WL(Wt)‖F∝λ/ηt\left\lVert\nabla_{W}L(W_{t})\right\rVert_{\mathrm{F}}\propto\sqrt{\lambda/\eta_{t}}when the direction of the weight is fixed and the autocorrelation factor changes slowly.
Interpretation.
The main message of the theory is that, for scale invariant matrix blocks, decoupled weight decay should be understood as an indirect controller of angular optimization speed rather than merely as a regularizer. The preceding results make this mechanism explicit:
- 1.For a scale invariant loss, the relevant optimization variable is the directionW^=W/‖W‖F\widehat{W}=W/\left\lVert W\right\rVert_{\mathrm{F}}, and the effective angular step is controlled by the angular learning rateηtang:=‖W^t+1−W^t‖F\eta^{\mathrm{ang}}_{t}:=\left\lVert\widehat{W}_{t+1}-\widehat{W}_{t}\right\rVert_{\mathrm{F}}.
- 2.Lemma˜4.3shows that for common optimizer the update normUUand the correlation for update between steps converge to constants that only depend on optimizer choices and hyperparameters.
- 3.Lemmas˜4.4and4.6show that (1) weight converges to an equilibrium radiusR⋆R_{\star}and (2) the angular learning rate converge to a constantηang\eta^{\mathrm{ang}}and both constant only depend on optimizer choices and hyperparameters.
- 4.Lemma˜4.7shows that, for scale invariant losses, larger equilibrium radius imply smaller gradient norms, giving the prediction‖∇WL(Wt)‖F∝1/Rt\left\lVert\nabla_{W}L(W_{t})\right\rVert_{\mathrm{F}}\propto 1/R_{t}when the direction remains unchange.
Thus, weight decay has two coupled effects: it fixes the radial scaleR⋆R_{\star}, and that radial scale determines the angular learning speed through‖ut‖F/R⋆\left\lVert u_{t}\right\rVert_{\mathrm{F}}/R_{\star}. This is the mechanism that Hyperball makes explicit: instead of letting weight decay indirectly determine both the matrix norm and the relative update length, Hyperball fixes the norm and the normalized update length directly.
4.3Empirical Validation
Phenomenon 1: weight norm tracks learning rate warmup and decay throughout training.
Under a WSD learning rate schedule, (39) predicts thatRtR_{t}should rise during warmup and shrink during learning rate decay. In a1.21.2B AdamW run with cosine learning rate decay, Q/K/V projection norms across layers show exactly this pattern: norms rise rapidly during warmup and then decrease during decay (fig.˜9, top).
Phenomenon 2: gradient norm increases through training.
For scale invariant blocks, (49) predicts that gradient norms scale approximately as1/Rt1/R_{t}. This phenomenon is also studied inDefazio (2025), where a similar explanation is provided. In the same run, the corresponding Q/K/V gradient norms increase late in training as the weight norms shrink during learning rate decay (fig.˜9, bottom).
Figure 9:Weight norm and gradient norm diagnostics for a1.21.2B AdamW run. Top: Q/K/V projection weight norms across all3232layers. Bottom: the corresponding gradient norms. Thin curves are individual layers. The dark curve is the layer mean. Weight norms follow the learning rate schedule, and gradient norms rise as weight norms shrink during decay.
Phenomenon 3: whenηλ\eta\lambdais fixed, AdamW converges to essentially the same loss while each matrix norm is roughly proportional toη\eta.
Theorem˜4.6predicts that holdingηλ\eta\lambdafixed keeps the angular step sizeηang\eta^{\mathrm{ang}}nearly fixed, so the training loss should be nearly unchanged. Furthermore, if we divideλ\lambdabyccand multiplyη\etabycc, then (39) predictsR⋆∝cR_{\star}\propto c. Infig.˜10, we present two runs with(η,λ)=(0.002,0.2)(\eta,\lambda)=(0.002,0.2)and(0.004,0.1)(0.004,0.1), and we observe that the train loss curves nearly overlap, while the equilibrium Q/K norms are roughly doubled in the larger learning rate run.
Figure 10:Ablation with fixedηλ\eta\lambda. Two AdamW runs with the same productηλ=4⋅10−4\eta\lambda=4\cdot 10^{-4}have nearly identical train loss, while the larger learning rate run has roughly doubled layer 9 Q/K norms.
Phenomenon 4: despite sharing the same learning rate schedule, weight decay starts with a higher loss but ultimately converges lower than no weight decay.
When the WD and no WD runs use the same learning rate warmup and decay schedule, (39) and (43) predict different weight norms and angular step dynamics. Without WD, the weight norm grows and the angular proxyη‖ut‖F/‖Wt‖F\eta\left\lVert u_{t}\right\rVert_{\mathrm{F}}/\left\lVert W_{t}\right\rVert_{\mathrm{F}}decays. With WD, the run maintains a larger effective step size throughout training. Empirically, and in the theory insection˜4.2, training with WD yields a larger effective step size than training without WD. According to the river valley theory(Wen et al.,2024), the loss decomposes into a “river” component, capturing progress along a relatively flat direction where long term optimization happens, and a “hill” component, capturing excursions in steep directions caused by stochastic gradients. A larger effective step size amplifies these hill direction oscillations, which raises the observed loss early in training, but it also accelerates motion along the river. When the learning rate decays, the oscillations in the hill directions shrink and the iterate settles closer to the riverbed, revealing the additional progress that has already been made along the river. This theory agrees with the phenomenon we observed here. The WD run starts with a higher loss but ultimately reaches a lower loss, because its larger effective step size allows it to move faster down the river before the decay phase suppresses the oscillations (fig.˜11).
Figure 11:Weight decay versus no weight decay under the same learning rate schedule. Without weight decay, the weight norm grows and the angular proxy decays. With weight decay, the run maintains a larger late angular proxy and crosses over to a lower validation loss. The dashed curve is the theory prediction for the WD angular proxy.
Phenomenon 5: contrary to the originalμ\muP prediction, transfer is not sensitive to weight scale at initialization but is sensitive to weight decay scaling.
Recent hyperparameter transfer studies find that transfer is often less sensitive to the initial weight scale than to how weight decay is scaled across model size and training duration(Kosson et al.,2025; Blake et al.,2024; Fan et al.,2025; Wang and Aitchison,2024; Qiu et al.,2025). This is consistent with (39): the radial recursion forgets the initial radius and converges to a norm set byη\eta,λ\lambda, and the optimizer dependent update normUU. Changing the scaling rule forλ\lambda, however, changes both the equilibrium normR⋆R_{\star}and the angular step size in (43), so it changes the dynamics relevant to transfer that are assumed byμ\muP style analyses(Yang et al.,2022). Hyperball turns this dependence into an explicit design choice by fixing the radius and normalized update length directly, which is why the same learning rate window transfers better across depths and widths insection˜3.3.
5Related work
Weight decay in normalized networks.
Earlier work showed that, in normalized networks, weight decay often changes optimization dynamics or effective learning rates rather than acting as a classical capacity penalty(van Laarhoven,2017; Zhang et al.,2019; Hoffer et al.,2018; D’Angelo et al.,2023). A line of work argues that, in the presence of BatchNorm(Ioffe and Szegedy,2015)or LayerNorm(Ba et al.,2016), weight decay acts through norm dynamics: it sets an equilibrium weight norm and, jointly with the learning rate, an angular step size(Li et al.,2020; Roburin et al.,2020; Kosson et al.,2024a).Yang et al. (2023)formalize the role of the relative update size in feature learning at scale via a spectral condition. Our analysis ofsection˜4sits squarely in this picture, closest in spirit to the rotational equilibrium framework ofKosson et al. (2024a). Hyperball is the matrix level wrapper that pins the relevant ratio directly rather than letting it equilibrate.
Norm constraints on weights and updates.
Decoupling weight magnitude from direction has a long history. Weight Normalization(Salimans and Kingma,2016)reparameterizesW=gV/‖V‖W=g\,V/\left\lVert V\right\rVert. Weight Standardization and BiT(Qiao et al.,2019; Kolesnikov et al.,2020)standardize kernel statistics. Convolutional Normalization(Liu et al.,2021)reduces per layer spectral norm. Decoupled Networks(Liu et al.,2018)split feature norm from angle. Artificial Kuramoto Oscillatory Neurons(Miyato et al.,2025)use unit norm oscillator states. On the update side, AdamP and SGDP(Heo et al.,2021)project updates onto the tangent space of the weight direction, Lion(Chen et al.,2023)fixes the per entry update magnitude via the sign function, and LionAR normalizes early update sizes using angular criteria(Kosson et al.,2024b). For generative models, EDM2(Karras et al.,2023)normalizes column weights and Spectral Normalization(Miyato et al.,2018)bounds the operator norm. Hyperball differs in being anoptimizer wrapperthat simultaneously fixes the matrix Frobenius norm and the update Frobenius norm, exposing the directional step size as a designed quantity.
Fixed norms in LLM pretraining and manifold optimization.
Several recent and concurrent works enforce normalization at the architecture or optimizer level for language model pretraining. nGPT(Loshchilov et al.,2024)enforces columnwise unit norms with adaptive normalization layers. Nemotron-Flash(Fu et al.,2025)applies per channel spherical constraints for inference time benefits but not on updates. The approximately normalized Transformer (anGPT)(Franke et al.,2025)bounds each weight row using constrained parameter regularization(Franke et al.,2024).Owen et al. (2025)periodically rescale weights toward a target variance. On richer manifolds, Modular Manifolds(Thinking Machines,2025), Muon++Stiefel(Su,2025a), notes on orthogonal manifolds and steepest descent(Bernstein,2025; Cesista,2025), andNewhouse et al. (2025)optimize on the Stiefel or spectral sphere manifold. Related spectral norm views of Muon and weight decay appear inSu (2024); Chen et al. (2025), while SSO(Xie et al.,2026)performs steepest descent or projection to the spectral sphere. Hyperball projects onto the matrix Frobenius sphereSdindout−1S^{d_{\mathrm{in}}d_{\mathrm{out}}-1}—a softer constraint than normalization by column or channel—atO(N2)O(N^{2})cost per matrix, versusO(N3)O(N^{3})for spectral projections.
6Conclusion
Hyperball replaces the implicit norm control of weight decay with an explicit optimizer constraint on matrix norms and update norms. Across our experiments, the explicit constraint improves the scaling behavior of matrix based optimizers and makes learning rate transfer more reliable across model widths, depths, and training budgets.
A broader question is which constraint should be imposed. The Frobenius norm is computationally cheap and theoretically motivated by the mechanism studied here, but spectral, rowwise, columnwise, hybrid, or architecture dependent constraints may better match some models and optimizers. Another direction is to develop a sharper theory of weight normalized training, including Weight Normalization style parameterizations(Salimans and Kingma,2016)and explicit norm constraints: when the radial degree of freedom is removed or fixed, how does this shape the training trajectory?
Acknowledgments
Kaiyue Wen acknowledges support from the Stanford Graduate Fellowship. Tengyu Ma acknowledges support from NSF grant 2522743. This work was supported by the Google TPU Research Cloud (TRC), the Stanford HAI–Google Cloud Credits Program, and NSF RI 2045685, and is part of the Marin Project. The authors would like to thank Songlin Yang, Zihan Qiu, and Liliang Ren for motivating this project. To some extent, this work is a proof of concept showing that it is possible to remove weight decay altogether by designing optimizers that explicitly control weight norms. The authors would also like to thank William Held, David Hall, Suhas Kotha, Tatsunori Hashimoto, Jason Lee, Zhiyuan Li, Lijie Chen, Huaqing Zhang, Jiacheng You, Jeremy Bernstein, Shu Zhong, Samuel Schoenholz, Evan Walters, and Omead Pooladzandi for helpful discussions.
References
- Azerbayev et al. [2023]Z. Azerbayev, H. Schoelkopf, K. Paster, M. Dos Santos, S. McAleer, A. Q. Jiang, J. Deng, S. Biderman, and S. Welleck.Llemma: An open language model for mathematics.arXiv preprint arXiv:2310.10631, 2023.URLhttps://arxiv.org/abs/2310.10631.
- Ba et al. [2016]J. L. Ba, J. R. Kiros, and G. E. Hinton.Layer normalization.InNeurIPS Deep Learning Symposium, 2016.URLhttps://arxiv.org/abs/1607.06450.arXiv:1607.06450.
- Bernstein [2025]J. Bernstein.Orthogonal manifold.https://docs.modula.systems/algorithms/manifold/orthogonal/, 2025.
- Blake et al. [2024]C. Blake, C. Eichenberg, J. Dean, L. Balles, L. Y. Prince, B. Deiseroth, A. F. Cruz-Salinas, C. Luschi, S. Weinbach, and D. Orr.u-μ\muP: The unit-scaled maximal update parametrization.arXiv preprint arXiv:2407.17465, 2024.URLhttps://arxiv.org/abs/2407.17465.
- Cesista [2025]F. L. Cesista.Heuristic solutions for steepest descent on the stiefel manifold.https://leloykun.github.io/ponder/steepest-descent-stiefel/, 2025.
- Chen et al. [2025]L. Chen, J. Li, and Q. Liu.Muon optimizes under spectral norm constraints.arXiv preprint arXiv:2506.15054, 2025.URLhttps://arxiv.org/abs/2506.15054.
- Chen et al. [2023]X. Chen, C. Liang, D. Huang, E. Real, K. Wang, Y. Liu, H. Pham, X. Dong, T. Luong, C.-J. Hsieh, Y. Lu, and Q. V. Le.Symbolic discovery of optimization algorithms.InNeurIPS, 2023.URLhttps://arxiv.org/abs/2302.06675.arXiv:2302.06675.
- D’Angelo et al. [2023]F. D’Angelo, M. Andriushchenko, A. Varre, and N. Flammarion.Why do we need weight decay in modern deep learning?arXiv preprint arXiv:2310.04415, 2023.URLhttps://arxiv.org/abs/2310.04415.
- Defazio [2025]A. Defazio.Why gradients rapidly increase near the end of training.arXiv preprint arXiv:2506.02285, 2025.URLhttps://arxiv.org/abs/2506.02285.
- Fan et al. [2025]Z. Fan, Y. Liu, Q. Zhao, A. Yuan, and Q. Gu.Robust layerwise scaling rules by proper weight decay tuning.arXiv preprint arXiv:2510.15262, 2025.URLhttps://arxiv.org/abs/2510.15262.
- Franke et al. [2024]J. K. Franke, M. Hefenbrock, G. Koehler, and F. Hutter.Improving deep learning optimization through constrained parameter regularization.InNeurIPS, 2024.URLhttps://arxiv.org/abs/2311.09058.arXiv:2311.09058.
- Franke et al. [2025]J. K. Franke, U. Spiegelhalter, M. Nezhurina, J. Jitsev, F. Hutter, and M. Hefenbrock.Learning in compact spaces with approximately normalized transformer.InNeurIPS, 2025.URLhttps://arxiv.org/abs/2505.22014.arXiv:2505.22014.
- Fu et al. [2025]Y. Fu, X. Dong, S. Diao, M. Van keirsbilck, H. Ye, W. Byeon, Y. Karnati, L. Liebenwein, H. Zhang, N. Binder, M. Khadkevich, A. Keller, J. Kautz, Y. C. Lin, and P. Molchanov.Nemotron-flash: Towards latency-optimal hybrid small language models.InNeurIPS, 2025.URLhttps://arxiv.org/abs/2511.18890.arXiv:2511.18890.
- Henry et al. [2020]A. Henry, P. R. Dachapally, S. S. Pawar, and Y. Chen.Query-key normalization for transformers.InFindings of EMNLP, pages 4246–4253, 2020.doi:10.18653/v1/2020.findings-emnlp.379.URLhttps://aclanthology.org/2020.findings-emnlp.379/.
- Heo et al. [2021]B. Heo, S. Chun, S. J. Oh, D. Han, S. Yun, G. Kim, Y. Uh, and J.-W. Ha.AdamP: Slowing down the slowdown for momentum optimizers on scale-invariant weights.InICLR, 2021.URLhttps://arxiv.org/abs/2006.08217.
- Hoffer et al. [2018]E. Hoffer, R. Banner, I. Golan, and D. Soudry.Norm matters: Efficient and accurate normalization schemes in deep networks.InNeurIPS, 2018.URLhttps://arxiv.org/abs/1803.01814.arXiv:1803.01814.
- Hoffmann et al. [2022]J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al.Training compute-optimal large language models.arXiv preprint arXiv:2203.15556, 2022.URLhttps://arxiv.org/abs/2203.15556.
- Ioffe and Szegedy [2015]S. Ioffe and C. Szegedy.Batch normalization: Accelerating deep network training by reducing internal covariate shift.InICML, 2015.URLhttps://arxiv.org/abs/1502.03167.
- Jordan et al. [2024]K. Jordan, Y. Jin, V. Boza, J. You, F. Cesista, L. Newhouse, and J. Bernstein.Muon: An optimizer for hidden layers in neural networks.https://kellerjordan.github.io/posts/muon/, 2024.
- Karras et al. [2023]T. Karras, M. Aittala, J. Lehtinen, J. Hellsten, T. Aila, and S. Laine.Analyzing and improving the training dynamics of diffusion models.arXiv preprint arXiv:2312.02696, 2023.URLhttps://arxiv.org/abs/2312.02696.
- Keller and Contributors [2026]J. Keller and Contributors.Modded-NanoGPT optimization benchmark: Track 3 optimization.https://github.com/KellerJordan/modded-nanogpt/tree/master/records/track_3_optimization, 2026.records/track_3_optimization, accessed 2026-05-17.
- Kingma and Ba [2015]D. P. Kingma and J. Ba.Adam: A method for stochastic optimization.InICLR, 2015.URLhttps://arxiv.org/abs/1412.6980.arXiv:1412.6980.
- Kolesnikov et al. [2020]A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby.Big transfer (BiT): General visual representation learning.InECCV, 2020.URLhttps://arxiv.org/abs/1912.11370.
- Kosson et al. [2024a]A. Kosson, B. Messmer, and M. Jaggi.Rotational equilibrium: How weight decay balances learning across neural networks.InICML, 2024a.URLhttps://arxiv.org/abs/2305.17212.arXiv:2305.17212.
- Kosson et al. [2024b]A. Kosson, B. Messmer, and M. Jaggi.Analyzing and reducing the need for learning rate warmup in GPT training.InNeurIPS, 2024b.URLhttps://arxiv.org/abs/2410.23922.arXiv:2410.23922.
- Kosson et al. [2025]A. Kosson, J. Welborn, Y. Liu, M. Jaggi, and X. Chen.Weight decay may matter more thanμ\muP for learning rate transfer in practice.arXiv preprint arXiv:2510.19093, 2025.URLhttps://arxiv.org/abs/2510.19093.
- Li et al. [2024]J. Li, A. Fang, G. Smyrnis, M. Ivgi, M. Jordan, S. Gadre, H. Bansal, E. Guha, S. Keh, K. Arora, et al.DataComp-LM: In search of the next generation of training sets for language models.InNeurIPS Datasets and Benchmarks, 2024.URLhttps://arxiv.org/abs/2406.11794.arXiv:2406.11794.
- Li et al. [2023]R. Li, L. Ben Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, et al.StarCoder: May the source be with you!arXiv preprint arXiv:2305.06161, 2023.URLhttps://arxiv.org/abs/2305.06161.
- Li et al. [2020]Z. Li, K. Lyu, and S. Arora.Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate.InNeurIPS, 2020.URLhttps://arxiv.org/abs/2010.02916.arXiv:2010.02916.
- Li et al. [2025]Z. Li, L. Liu, C. Liang, W. Chen, and T. Zhao.NorMuon: Making Muon more efficient and scalable.arXiv preprint arXiv:2510.05491, 2025.URLhttps://arxiv.org/abs/2510.05491.
- Lin et al. [2026]W. Lin, S. C. Lowe, F. Dangel, R. Eschenhagen, Z. Xu, and R. B. Grosse.Understanding and improving shampoo and SOAP via Kullback-Leibler minimization.InICLR, 2026.URLhttps://arxiv.org/abs/2509.03378.arXiv:2509.03378.
- Liu et al. [2025]J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y. Du, Y. Qin, W. Xu, et al.Muon is scalable for LLM training.arXiv preprint arXiv:2502.16982, 2025.URLhttps://arxiv.org/abs/2502.16982.
- Liu et al. [2021]S. Liu, X. Li, Y. Zhai, C. You, Z. Zhu, C. Fernandez-Granda, and Q. Qu.Convolutional normalization: Improving deep convolutional network robustness and training.InNeurIPS, 2021.URLhttps://arxiv.org/abs/2103.00673.arXiv:2103.00673.
- Liu et al. [2018]W. Liu, Z. Liu, Z. Yu, B. Dai, R. Lin, Y. Wang, J. M. Rehg, and L. Song.Decoupled networks.InCVPR, 2018.URLhttps://arxiv.org/abs/1804.08071.
- Loshchilov and Hutter [2019]I. Loshchilov and F. Hutter.Decoupled weight decay regularization.InICLR, 2019.URLhttps://arxiv.org/abs/1711.05101.arXiv:1711.05101.
- Loshchilov et al. [2024]I. Loshchilov, C.-P. Hsieh, S. Sun, and B. Ginsburg.nGPT: Normalized transformer with representation learning on the hypersphere.arXiv preprint arXiv:2410.01131, 2024.URLhttps://arxiv.org/abs/2410.01131.
- Miyato et al. [2018]T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida.Spectral normalization for generative adversarial networks.InICLR, 2018.URLhttps://arxiv.org/abs/1802.05957.
- Miyato et al. [2025]T. Miyato, S. Löwe, A. Geiger, and M. Welling.Artificial Kuramoto oscillatory neurons.InICLR, 2025.URLhttps://arxiv.org/abs/2410.13821.arXiv:2410.13821.
- Newhouse et al. [2025]L. Newhouse, R. P. Hess, F. Cesista, A. Zahorodnii, J. Bernstein, and P. Isola.Training transformers with enforced Lipschitz constants.arXiv preprint arXiv:2507.13338, 2025.URLhttps://arxiv.org/abs/2507.13338.
- Owen et al. [2025]L. Owen, A. Kumar, N. Roy Chowdhury, and F. Güra.Variance control via weight rescaling in LLM pre-training.arXiv preprint arXiv:2503.17500, 2025.URLhttps://arxiv.org/abs/2503.17500.
- Penedo et al. [2024]G. Penedo, H. Kydlíček, L. Ben Allal, A. Lozhkov, M. Mitchell, C. Raffel, L. von Werra, and T. Wolf.The FineWeb datasets: Decanting the web for the finest text data at scale.InNeurIPS Datasets and Benchmarks, 2024.URLhttps://arxiv.org/abs/2406.17557.arXiv:2406.17557.
- Qiao et al. [2019]S. Qiao, H. Wang, C. Liu, W. Shen, and A. Yuille.Micro-batch training with batch-channel normalization and weight standardization.arXiv preprint arXiv:1903.10520, 2019.URLhttps://arxiv.org/abs/1903.10520.
- Qiu et al. [2025]S. Qiu, Z. Chen, H. Phan, Q. Lei, and A. G. Wilson.Hyperparameter transfer enables consistent gains of matrix-preconditioned optimizers across scales.InNeurIPS, 2025.URLhttps://arxiv.org/abs/2512.05620.arXiv:2512.05620.
- Raffel et al. [2020]C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu.Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of Machine Learning Research, 21(140):1–67, 2020.URLhttps://www.jmlr.org/papers/v21/20-074.html.
- Roburin et al. [2020]S. Roburin, Y. de Mont-Marin, A. Bursuc, R. Marlet, P. Pérez, and M. Aubry.Spherical perspective on learning with normalization layers.arXiv preprint arXiv:2006.13382, 2020.URLhttps://arxiv.org/abs/2006.13382.
- Salimans and Kingma [2016]T. Salimans and D. P. Kingma.Weight normalization: A simple reparameterization to accelerate training of deep neural networks.InNeurIPS, 2016.URLhttps://arxiv.org/abs/1602.07868.
- Su [2024]J. Su.Thinking about spectral norm gradient and spectral weight decay.https://kexue.fm/archives/10648, 2024.
- Su [2025a]J. Su.Muon on the stiefel manifold.https://kexue.fm/archives/11221, 2025a.
- Su [2025b]J. Su.Why Adam’s update RMS is 0.2?https://kexue.fm/archives/11267, 2025b.
- Su [2025c]J. Su.AdamW weight RMS asymptotics (part I).https://kexue.fm/archives/11307, 2025c.
- Su [2025d]J. Su.AdamW weight RMS asymptotics (part II).https://kexue.fm/archives/11404, 2025d.
- Thinking Machines [2025]Thinking Machines.Modular manifolds.https://thinkingmachines.ai/blog/modular-manifolds/, 2025.
- van Laarhoven [2017]T. van Laarhoven.L2 regularization versus batch and weight normalization.arXiv preprint arXiv:1706.05350, 2017.URLhttps://arxiv.org/abs/1706.05350.
- Vaswani et al. [2017]A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin.Attention is all you need.InNeurIPS, 2017.URLhttps://arxiv.org/abs/1706.03762.
- Wang and Aitchison [2024]X. Wang and L. Aitchison.How to set AdamW’s weight decay as you scale model and dataset size.arXiv preprint arXiv:2405.13698, 2024.URLhttps://arxiv.org/abs/2405.13698.
- Wen et al. [2024]K. Wen, Z. Li, J. Wang, D. Hall, P. Liang, and T. Ma.Understanding warmup-stable-decay learning rates: A river valley loss landscape perspective.arXiv preprint arXiv:2410.05192, 2024.URLhttps://arxiv.org/abs/2410.05192.
- Wen et al. [2025]K. Wen, D. Hall, T. Ma, and P. Liang.Fantastic pretraining optimizers and where to find them.arXiv preprint arXiv:2509.02046, 2025.URLhttps://arxiv.org/abs/2509.02046.
- Xie et al. [2026]T. Xie, H. Luo, H. Tang, Y. Hu, J. K. Liu, Q. Ren, Y. Wang, W. X. Zhao, R. Yan, B. Su, C. Luo, and B. Guo.Controlled LLM training on spectral sphere.arXiv preprint arXiv:2601.08393, 2026.URLhttps://arxiv.org/abs/2601.08393.
- Xiong et al. [2020]R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T.-Y. Liu.On layer normalization in the transformer architecture.InICML, 2020.URLhttps://arxiv.org/abs/2002.04745.
- Yang et al. [2025]A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al.Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025.URLhttps://arxiv.org/abs/2505.09388.
- Yang et al. [2022]G. Yang, E. J. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao.Tensor programs V: Tuning large neural networks via zero-shot hyperparameter transfer.arXiv preprint arXiv:2203.03466, 2022.URLhttps://arxiv.org/abs/2203.03466.
- Yang et al. [2023]G. Yang, J. B. Simon, and J. Bernstein.A spectral condition for feature learning.arXiv preprint arXiv:2310.17813, 2023.URLhttps://arxiv.org/abs/2310.17813.
- Zhang and Sennrich [2019]B. Zhang and R. Sennrich.Root mean square layer normalization.InNeurIPS, 2019.URLhttps://arxiv.org/abs/1910.07467.
- Zhang et al. [2019]G. Zhang, C. Wang, B. Xu, and R. Grosse.Three mechanisms of weight decay regularization.InICLR, 2019.URLhttps://openreview.net/forum?id=B1lz-3Rct7.
- Zhuo et al. [2025]Z. Zhuo, Y. Zeng, Y. Wang, S. Zhang, J. Yang, X. Li, X. Zhou, and J. Ma.HybridNorm: Towards stable and efficient transformer training via hybrid normalization.InNeurIPS, 2025.URLhttps://arxiv.org/abs/2503.04598.arXiv:2503.04598.
相似文章
Meta Muse Glimmer – 开放权重的30B本地编码模型
Meta推出Muse Glimmer,一个采用宽松许可的30B参数模型,专为本地智能体工作流、编码和工具使用而优化,权重已在Hugging Face上发布。
推出 Muse Glimmer:专为常驻本地智能体工作流优化的开放权重模型
Meta 发布 Muse Glimmer,这是一款 30B 开放权重多模态模型,专为本地智能体工作流优化,采用宽松的 Apache 2.0 许可证,支持 4 位量化、推测解码,并提供广泛的生态系统集成。
推出 Muse Glimmer
Meta 推出 Muse Glimmer,一款基于 Apache 2.0 协议的全新 30B 开放权重模型,针对智能体任务完成、可靠工具使用和多步推理进行了优化。Simon Willison 使用 LM Studio 和 llm-coding-agent 在本地对其进行了测试。
@PyTorch: Today @AIatMeta introduced Muse Glimmer, an open-weight, 30-billion-parameter model distilled from Meta’s Muse Spark fo…
Meta introduced Muse Glimmer, an open-weight 30B-parameter model distilled from Muse Spark for on-device agentic workflows, with ExecuTorch now supporting running it on NVIDIA GPUs and Apple silicon.
被低估的 Muse Glimmer
基准测试结果显示,Muse Glimmer 在隐式知识测试中出人意料地超越了 qwen3.8,这表明较小的模型通过 RAG 增强可以实现有竞争力的性能。