kernel-attention

Tag

Cards List
#kernel-attention

Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

arXiv cs.LG ↗ · 2026-08-13 Cached

This paper proves that a single normalized nonnegative kernel-attention head requires exponentially many features to solve a simple Min-IP task on three-token sequences, whereas dense softmax attention solves it with constant temperature and m-dimensional scores, highlighting a fundamental expressive-power gap between kernel and full attention.

0 favorites 0 likes
← Back to home

Submit Feedback