low-rank-training

Tag

Cards List
#low-rank-training

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training

arXiv cs.LG · 2026-07-08 Cached

This paper empirically shows that the gradient's top-r subspace in low-rank training methods like GaLore is non-identifiable beyond a small reproducible core, with estimator noise dominating apparent rotations. It analyzes the implications for optimizer state transport and introduces LDAdam, which outperforms GaLore in perplexity.

0 favorites 0 likes
← Back to home

Submit Feedback