Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P]
Summary
Tauon is a new optimizer using polynomial and orthogonalization techniques that outperforms Muon and AdamW in initial benchmarks on a small GPT-Mini model, showing lower loss and faster step time.
Similar Articles
@burny_tech: some updates on the optimizer magic
A new NVIDIA paper proposes higher-order optimizers like Muon and SOAP as more efficient alternatives to AdamW for large-scale LLM pretraining.
How Much Orthogonalization Does Muon Need?
This paper studies how much orthogonalization the Muon optimizer requires, proposing a five-step cubic Newton-Schulz schedule that reduces computational cost while achieving training quality similar to more expensive methods across GPT-2 Small and hybrid MoE/Mamba models.
Can Muon Fine-tune Adam-Pretrained Models?
Research paper investigating performance degradation when using the Muon optimizer instead of Adam for fine-tuning pretrained models, demonstrating that parameter-efficient methods like LoRA effectively mitigate this optimizer mismatch across language and vision tasks.
Reassessing Muon for Matrix Factorization
This paper evaluates the Muon optimizer on low-rank matrix factorization, finding it does not consistently outperform AdamW, challenging earlier claims about its advantages in large-scale deep learning.
Dion3: Full-Stack Orthogonal Updates
Dion3 is a revised Muon optimizer that reduces computational and communication overhead via Gram Newton-Schulz, symmetric GEMM kernels, and megabatching, achieving up to 6x faster optimizer steps while matching or improving loss.