preconditioning

Tag

Cards List
#preconditioning

SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning

arXiv cs.LG · 3d ago Cached

SS-eSOAP introduces a self-scaled adaptive preconditioning method for physics-informed neural networks, improving optimization by reducing residuals and achieving higher accuracy on various PDE benchmarks compared to baseline methods.

0 favorites 0 likes
#preconditioning

The Loss Does Not See the Basis, but Adam Does

Hugging Face Daily Papers · 2026-08-05 Cached

This paper investigates why Adam does not exhibit gradient descent's implicit low-rank bias in factored models, showing that coordinate-wise preconditioning breaks the relevant symmetry, while shared-scalar methods like Muon and Shampoo preserve it.

0 favorites 0 likes
#preconditioning

@HazanPrinceton: Just in time for our tutorial at ICML next week, Annie posted an update to our universal sequence preconditioning paper…

X AI KOLs Timeline · 2026-06-29 Cached

This paper update presents a universal sequence preconditioning method achieving dimension-free regret bounds for marginally stable linear dynamical systems, using second-order VAW algorithm and Faber polynomials.

0 favorites 0 likes
#preconditioning

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning

arXiv cs.LG · 2026-06-15 Cached

Zeta proposes a dual whitening optimizer that applies coordinate whitening before spectral whitening to resolve scale heterogeneity in momentum matrices, reducing orthogonalization error and improving convergence and generalization in large-scale neural network training.

0 favorites 0 likes
#preconditioning

Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage

arXiv cs.LG · 2026-05-27 Cached

This paper proposes a reparametrization of the preconditioner in Shampoo-based optimization methods (like KL-Shampoo and SOAP) to support BFloat16 storage and reduce computational overhead by updating only part of the basis via QR decomposition in a subspace, making these methods more memory- and time-efficient.

0 favorites 0 likes
#preconditioning

AdaPreLoRA: Adafactor Preconditioned Low-Rank Adaptation

Hugging Face Daily Papers · 2026-05-09 Cached

AdaPreLoRA is a novel LoRA optimizer that uses Adafactor diagonal Kronecker preconditioning to improve factor-space updates while maintaining low memory usage, demonstrating competitive performance across various LLMs and tasks.

0 favorites 0 likes
← Back to home

Submit Feedback