soap

Tag

Cards List
#soap

@burny_tech: some updates on the optimizer magic

X AI KOLs Timeline · yesterday Cached

A new NVIDIA paper proposes higher-order optimizers like Muon and SOAP as more efficient alternatives to AdamW for large-scale LLM pretraining.

0 favorites 0 likes
#soap

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

arXiv cs.LG · 2d ago Cached

This paper improves higher-order optimizers SOAP and Muon for large-scale LLM pretraining, addressing instabilities at large batch sizes and introducing a layer-wise distributed optimizer compatible with Megatron-LM. Experiments show they consistently outperform AdamW at billion-parameter scales.

0 favorites 0 likes
#soap

Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage

arXiv cs.LG · 2026-05-27 Cached

This paper proposes a reparametrization of the preconditioner in Shampoo-based optimization methods (like KL-Shampoo and SOAP) to support BFloat16 storage and reduce computational overhead by updating only part of the basis via QR decomposition in a subspace, making these methods more memory- and time-efficient.

0 favorites 0 likes
← Back to home

Submit Feedback