Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage

arXiv cs.LG Papers

Summary

This paper proposes a reparametrization of the preconditioner in Shampoo-based optimization methods (like KL-Shampoo and SOAP) to support BFloat16 storage and reduce computational overhead by updating only part of the basis via QR decomposition in a subspace, making these methods more memory- and time-efficient.

arXiv:2605.26327v1 Announce Type: new Abstract: Shampoo-based methods, such as KL-Shampoo and SOAP, have demonstrated strong performance in training neural networks and rely on QR decomposition. Because existing QR implementations require single-precision (FP32) arithmetic and remain computationally expensive, these methods become time- and memory-intensive when their preconditioning matrices are large. Moreover, using BFloat16 (BFP16) storage to reduce memory usage can degrade the performance of Shampoo-based methods. We propose a reparametrization of the preconditioner that supports BFP16 storage and forms a complete basis by combining updated basis vectors with unchanged ones. By updating only part of the basis through QR decomposition in a subspace, our approach reduces computational overhead while mitigating the performance degradation caused by BFP16 storage. Our approach applies broadly to Shampoo-based methods that employ QR decomposition, including KL-Shampoo, SOAP, and KL-SOAP. In particular, it improves the performance of SOAP and KL-SOAP under BFP16 storage, enabling KL-SOAP to match or exceed KL-Shampoo. Overall, our approach makes Shampoo-based methods more memory- and time-efficient.
Original Article
View Cached Full Text

Cached at: 05/27/26, 09:08 AM

# Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage
Source: [https://arxiv.org/abs/2605.26327](https://arxiv.org/abs/2605.26327)
[View PDF](https://arxiv.org/pdf/2605.26327)

> Abstract:Shampoo\-based methods, such as KL\-Shampoo and SOAP, have demonstrated strong performance in training neural networks and rely on QR decomposition\. Because existing QR implementations require single\-precision \(FP32\) arithmetic and remain computationally expensive, these methods become time\- and memory\-intensive when their preconditioning matrices are large\. Moreover, using BFloat16 \(BFP16\) storage to reduce memory usage can degrade the performance of Shampoo\-based methods\. We propose a reparametrization of the preconditioner that supports BFP16 storage and forms a complete basis by combining updated basis vectors with unchanged ones\. By updating only part of the basis through QR decomposition in a subspace, our approach reduces computational overhead while mitigating the performance degradation caused by BFP16 storage\. Our approach applies broadly to Shampoo\-based methods that employ QR decomposition, including KL\-Shampoo, SOAP, and KL\-SOAP\. In particular, it improves the performance of SOAP and KL\-SOAP under BFP16 storage, enabling KL\-SOAP to match or exceed KL\-Shampoo\. Overall, our approach makes Shampoo\-based methods more memory\- and time\-efficient\.

## Submission history

From: Wu Lin \[[view email](https://arxiv.org/show-email/8842e08e/2605.26327)\] **\[v1\]**Mon, 25 May 2026 21:03:03 UTC \(545 KB\)

Similar Articles

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

arXiv cs.LG

This paper improves higher-order optimizers SOAP and Muon for large-scale LLM pretraining, addressing instabilities at large batch sizes and introducing a layer-wise distributed optimizer compatible with Megatron-LM. Experiments show they consistently outperform AdamW at billion-parameter scales.

Structured Transforms for Low-Overhead Quantization of Language Models

arXiv cs.CL

This paper revisits Kashin-decomposition-based weight quantization for large language models and proposes an improved algorithm using structured orthogonal transforms, reducing computational cost and ensuring numerical stability compared to methods like OPTQ and QuIP.