Tag
The paper introduces YB-Mixer, a token-mixing layer derived from the generalized Yang-Baxter equation, which is exactly norm-preserving, depth-stable, and allows order-free and variable-budget inference. It achieves competitive performance on long-range memory tasks with fewer parameters compared to attention and state-space baselines.
This paper proposes RankElastor, a novel architecture that mitigates embedding collapse in dense scaling of recommendation models by introducing parameterized full mixing and GLU-improved P-FFNs, achieving robust scaling and improved performance on large-scale datasets.