gradient-stability

Tag

Cards List
#gradient-stability

Exact Linear Attention

arXiv cs.LG · 2026-05-20

This paper introduces Exact Linear Attention (ELA), a mechanism that achieves linear computational complexity for Transformer attention without approximation error by leveraging kernel decomposition, and addresses gradient explosion and token dilution through constrained kernel functions. It also presents engineering innovations including Hyper Link, Memory Lobe, and a routing bias for Mixture of Experts.

0 favorites 0 likes
← Back to home

Submit Feedback