muon

Tag

Cards List
#muon

Physicists Solve a Muon Mystery. Now, Old Results Don't Add Up

Hacker News Top · 4d ago Cached

New theoretical calculations resolve a 25-year muon g-2 puzzle, but create a clash with older experimental results, suggesting possible new physics.

0 favorites 0 likes
#muon

@burny_tech: some updates on the optimizer magic

X AI KOLs Timeline · 2026-07-24 Cached

A new NVIDIA paper proposes higher-order optimizers like Muon and SOAP as more efficient alternatives to AdamW for large-scale LLM pretraining.

0 favorites 0 likes
#muon

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

arXiv cs.LG · 2026-07-24 Cached

This paper improves higher-order optimizers SOAP and Muon for large-scale LLM pretraining, addressing instabilities at large batch sizes and introducing a layer-wise distributed optimizer compatible with Megatron-LM. Experiments show they consistently outperform AdamW at billion-parameter scales.

0 favorites 0 likes
#muon

@RichardYRLi: Muon erases spectral anisotropy, ISO inherits it - in current LLM-RLVR, inherit wins if RLVR mostly elicits from pretra…

X AI KOLs Timeline · 2026-07-24 Cached

The thread introduces ISO, an isospectral optimizer that preserves spectral anisotropy, contrasting with Muon which erases it, and discusses its implications for LLM-RLVR training and continual learning.

0 favorites 0 likes
#muon

When Does Muon Help Agentic Reinforcement Learning?

Hugging Face Daily Papers · 2026-07-17 Cached

This paper investigates the use of the Muon optimizer in reinforcement learning post-training, finding that applying Muon to hidden weight matrices significantly improves success rates on ALFRED tasks compared to AdamW, with results dependent on the advantage estimator and learning rate.

0 favorites 0 likes
#muon

Reassessing Muon for Matrix Factorization

arXiv cs.LG · 2026-07-16 Cached

This paper evaluates the Muon optimizer on low-rank matrix factorization, finding it does not consistently outperform AdamW, challenging earlier claims about its advantages in large-scale deep learning.

0 favorites 0 likes
#muon

Aurora: A Leverage-Aware Spectral Optimizer

arXiv cs.LG · 2026-06-29 Cached

Aurora is a leverage-aware spectral optimizer that addresses neuron death in MLP layers by enforcing row uniformity while preserving the polar factor geometry of Muon updates, achieving state-of-the-art performance on the modded-nanoGPT speedrun benchmark.

0 favorites 0 likes
#muon

@plugyawn: Introducing: Megaprop: a library for efficient preconditioned optimization across GPUs! Megaprop is a fork of Megatron …

X AI KOLs Following · 2026-06-15 Cached

Megaprop is a new library for efficient preconditioned optimization across GPUs, forked from Megatron and TransformerEngine, with FSDP support for Muon, FOOF, KFAC, and Newton-Muon, and MuP support for width and depth.

0 favorites 0 likes
#muon

Muon$^p$: Muon with Fractional Spectral Powers

arXiv cs.LG · 2026-06-15 Cached

This paper introduces Muon^p, a novel optimizer that uses fractional spectral-power updates to interpolate between Muon and gradient descent, providing theoretical justification and empirical gains on billion-scale fine-tuning tasks.

0 favorites 0 likes
#muon

@maximelabonne: Parallax is a parametrized form of Local Linear Attention that drops the numerical solvers and matches FA 2/3 on decode…

X AI KOLs Following · 2026-06-10 Cached

Parallax is a new parametrized form of Local Linear Attention that eliminates numerical solvers and matches FlashAttention 2/3 in decoding. Its effectiveness depends on the optimizer, working with Muon but not AdamW, highlighting the role of optimizer geometry.

0 favorites 0 likes
#muon

Gram Newton-Schulz: A Fast, Hardware-Aware Newton-Schulz Algorithm for Muon

Hacker News Top · 2026-06-09 Cached

This blog post presents Gram Newton-Schulz, a hardware-aware optimization of the Newton-Schulz orthogonalization procedure used in the Muon optimizer, achieving significant speedups for training large language models while preserving model quality.

0 favorites 0 likes
#muon

Spectral Scaling Laws of Muon

arXiv cs.LG · 2026-06-04 Cached

This paper presents the first systematic study of singular value spectral behavior in Muon optimizer momentum matrices during LLM training, discovering clean power-law scaling relationships across model sizes (77M–2.8B parameters). The findings provide practitioners with principled, layer-aware guidelines for configuring Newton–Schulz iterations to maintain orthonormalization quality at frontier scale without unnecessary computation.

0 favorites 0 likes
#muon

Why Muon Outperforms Adam: A Curvature Perspective

Hugging Face Daily Papers · 2026-06-03 Cached

This paper investigates why the Muon optimizer outperforms Adam in large language model training, showing from a curvature perspective that Muon incurs a smaller curvature penalty due to lower normalized directional sharpness, with advantages amplified by data imbalance.

0 favorites 0 likes
#muon

MuCon: Clipped Muon Updates for LLM Training

arXiv cs.LG · 2026-05-27 Cached

This paper introduces MuCon, a clipped-Muon optimizer for LLM training that applies singular-value clipping instead of full polarization, preserving smaller singular values while clipping only the largest ones. It explores approximations to avoid full SVD, including polar/absolute-value formulas and rational Newton filters, noting numerical challenges near the threshold.

0 favorites 0 likes
#muon

DynMuon: A Dynamic Spectral Shaping View of Muon

Hugging Face Daily Papers · 2026-05-16 Cached

This paper introduces DynMuon, a dynamic spectral shaping optimizer that schedules the update parameter p from positive to mildly negative during training, consistently achieving lower validation loss and requiring 10.6-26.5% fewer steps than the standard Muon optimizer.

0 favorites 0 likes
#muon

Muon is Not That Special: Random or Inverted Spectra Work Just as Well

arXiv cs.LG · 2026-05-13 Cached

This paper challenges the geometric justification for the Muon optimizer, arguing that precise structure is less important than step-size optimality. It introduces Freon and Kaon optimizers to demonstrate that random or inverted spectra can perform as well as Muon.

0 favorites 0 likes
#muon

Can Muon Fine-tune Adam-Pretrained Models?

Hugging Face Daily Papers · 2026-05-11 Cached

Research paper investigating performance degradation when using the Muon optimizer instead of Adam for fine-tuning pretrained models, demonstrating that parameter-efficient methods like LoRA effectively mitigate this optimizer mismatch across language and vision tasks.

0 favorites 0 likes
← Back to home

Submit Feedback