Tag
This paper introduces Stiefel-AdamW, a geometry-aware optimizer for linear factorization blocks in deep learning that enhances stability and performance, validated on models like GPT2, ViT, and Mistral 7B.
This paper introduces Stiefel Attention, which constrains transformer query and key projection matrices to the Stiefel manifold using Riemannian optimization, demonstrating improved performance on modular arithmetic grokking and CIFAR-10 patches.
The paper introduces NS-RIS, a scalable Newton-Schulz retraction-based algorithm for learning hidden quantum Markov models on the Stiefel manifold, providing the first mathematical performance guarantee and empirical evidence that HQMMs can outperform EM-trained HMMs on non-quantum-generated data.
Proposes PP-CPCANet, a covariance-free framework for domain generalization that learns a global orthogonal basis on the Stiefel manifold and achieves SOTA performance on four benchmarks.
This paper introduces MAPL, a method for learned orthogonal compression of activations in pipeline parallelism, reducing communication overhead while maintaining performance via Stiefel manifold constraints and per-stage factorized anchor embeddings.
FoRA introduces a parameter-efficient fine-tuning method that selects task-informative layers via Fisher scores and trains LoRA down-projections on the Stiefel manifold, reducing parameters while preserving accuracy.