stiefel-manifold

Tag

Cards List
#stiefel-manifold

Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks

arXiv cs.LG ↗ · 2026-09-21 Cached

This paper introduces Stiefel-AdamW, a geometry-aware optimizer for linear factorization blocks in deep learning that enhances stability and performance, validated on models like GPT2, ViT, and Mistral 7B.

0 favorites 0 likes
#stiefel-manifold

Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not

arXiv cs.LG ↗ · 2026-09-18 Cached

This paper introduces Stiefel Attention, which constrains transformer query and key projection matrices to the Stiefel manifold using Riemannian optimization, demonstrating improved performance on modular arithmetic grokking and CIFAR-10 patches.

0 favorites 0 likes
#stiefel-manifold

Newton-Schulz Retraction-Based Inference Enables Hidden Quantum Markov Models to Outperform Classical HMMs

arXiv cs.LG ↗ · 2026-08-10 Cached

The paper introduces NS-RIS, a scalable Newton-Schulz retraction-based algorithm for learning hidden quantum Markov models on the Stiefel manifold, providing the first mathematical performance guarantee and empirical evidence that HQMMs can outperform EM-trained HMMs on non-quantum-generated data.

0 favorites 0 likes
#stiefel-manifold

Projection Pursuit CPCANet for Domain Generalization

Hugging Face Daily Papers ↗ · 2026-07-24 Cached

Proposes PP-CPCANet, a covariance-free framework for domain generalization that learns a global orthogonal basis on the Stiefel manifold and achieves SOTA performance on four benchmarks.

0 favorites 0 likes
#stiefel-manifold

Learned Subspace Compression for Communication-Efficient Pipeline Parallelism

arXiv cs.LG ↗ · 2026-06-05 Cached

This paper introduces MAPL, a method for learned orthogonal compression of activations in pipeline parallelism, reducing communication overhead while maintaining performance via Stiefel manifold constraints and per-stage factorized anchor embeddings.

0 favorites 0 likes
#stiefel-manifold

FoRA: Fisher-orthogonal Rank Adaptation for Parameter-Efficient Fine-Tuning

arXiv cs.CL ↗ · 2026-05-29 Cached

FoRA introduces a parameter-efficient fine-tuning method that selects task-informative layers via Fisher scores and trains LoRA down-projections on the Stiefel manifold, reducing parameters while preserving accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback