muon-optimizer

Tag

Cards List
#muon-optimizer

How Much Orthogonalization Does Muon Need?

arXiv cs.LG · 2026-06-02 Cached

This paper studies how much orthogonalization the Muon optimizer requires, proposing a five-step cubic Newton-Schulz schedule that reduces computational cost while achieving training quality similar to more expensive methods across GPT-2 Small and hybrid MoE/Mamba models.

0 favorites 0 likes
#muon-optimizer

@zhaoran_wang: for me, the coolest finding is that you can connect/interpolate all softmax/linear variants and give a promising direct…

X AI KOLs Timeline · 2026-05-30 Cached

Discussion of a finding that all softmax/linear attention variants can be interpolated, and that the Muon optimizer is crucial for Parallax to move beyond Softmax Attention. Includes link to paper and code.

0 favorites 0 likes
#muon-optimizer

SignMuon: Communication-Efficient Distributed Muon Optimization

arXiv cs.LG · 2026-05-19 Cached

SignMuon is a 1-bit, matrix-aware optimizer for distributed training that combines signSGD's majority-vote sign aggregation with Muon's polar-step framework, achieving 32x bandwidth reduction over float32 while maintaining strong convergence and performance on benchmarks like CIFAR-10/ResNet-50 and nanoGPT.

0 favorites 0 likes
← Back to home

Submit Feedback