Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage
Summary
This paper proposes a reparametrization of the preconditioner in Shampoo-based optimization methods (like KL-Shampoo and SOAP) to support BFloat16 storage and reduce computational overhead by updating only part of the basis via QR decomposition in a subspace, making these methods more memory- and time-efficient.
View Cached Full Text
Cached at: 05/27/26, 09:08 AM
# Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage Source: [https://arxiv.org/abs/2605.26327](https://arxiv.org/abs/2605.26327) [View PDF](https://arxiv.org/pdf/2605.26327) > Abstract:Shampoo\-based methods, such as KL\-Shampoo and SOAP, have demonstrated strong performance in training neural networks and rely on QR decomposition\. Because existing QR implementations require single\-precision \(FP32\) arithmetic and remain computationally expensive, these methods become time\- and memory\-intensive when their preconditioning matrices are large\. Moreover, using BFloat16 \(BFP16\) storage to reduce memory usage can degrade the performance of Shampoo\-based methods\. We propose a reparametrization of the preconditioner that supports BFP16 storage and forms a complete basis by combining updated basis vectors with unchanged ones\. By updating only part of the basis through QR decomposition in a subspace, our approach reduces computational overhead while mitigating the performance degradation caused by BFP16 storage\. Our approach applies broadly to Shampoo\-based methods that employ QR decomposition, including KL\-Shampoo, SOAP, and KL\-SOAP\. In particular, it improves the performance of SOAP and KL\-SOAP under BFP16 storage, enabling KL\-SOAP to match or exceed KL\-Shampoo\. Overall, our approach makes Shampoo\-based methods more memory\- and time\-efficient\. ## Submission history From: Wu Lin \[[view email](https://arxiv.org/show-email/8842e08e/2605.26327)\] **\[v1\]**Mon, 25 May 2026 21:03:03 UTC \(545 KB\)
Similar Articles
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
This paper improves higher-order optimizers SOAP and Muon for large-scale LLM pretraining, addressing instabilities at large batch sizes and introducing a layer-wise distributed optimizer compatible with Megatron-LM. Experiments show they consistently outperform AdamW at billion-parameter scales.
@burny_tech: some updates on the optimizer magic
A new NVIDIA paper proposes higher-order optimizers like Muon and SOAP as more efficient alternatives to AdamW for large-scale LLM pretraining.
Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe
This paper identifies a fundamental limitation (shrinkage bias) in non-uniform FP4 quantization formats for LLM pretraining and proposes UFP4, a uniform 4-bit training recipe that outperforms existing E2M1-based methods.
@anirudhbv_ce: Introducing SpectralQuant.. here to save your KV cache :)
SpectralQuant is a new KV cache quantization technique achieving 5.95× compression on Mistral 7B with only 7.5% perplexity overhead, significantly outperforming TurboQuant while requiring only 15 seconds of calibration per model.
Structured Transforms for Low-Overhead Quantization of Language Models
This paper revisits Kashin-decomposition-based weight quantization for large language models and proposes an improved algorithm using structured orthogonal transforms, reducing computational cost and ensuring numerical stability compared to methods like OPTQ and QuIP.