Tag
This paper investigates whether gradient-based training can learn the rank necessary to store and compose associations in matrix memories, using experiments with key-value pairs and measuring recovery by cosine similarity.
This paper provides causal evidence that gradient descent in matrix memories trained on group composition tasks recruits exactly the rank required by the task's algebra, correlating with the minimal faithful representation dimensions of finite groups.
This paper presents Iso-LoRA, an optimizer that enhances LoRA by promoting even energy distribution across singular directions through spectral descent on tangent-space perturbations, improving effective rank and downstream performance in language model adaptation.
This paper proposes RankElastor, a novel architecture that mitigates embedding collapse in dense scaling of recommendation models by introducing parameterized full mixing and GLU-improved P-FFNs, achieving robust scaling and improved performance on large-scale datasets.