kernel-attention

标签

Cards List
#kernel-attention

Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

arXiv cs.LG ↗ · 2026-08-13 缓存

This paper proves that a single normalized nonnegative kernel-attention head requires exponentially many features to solve a simple Min-IP task on three-token sequences, whereas dense softmax attention solves it with constant temperature and m-dimensional scores, highlighting a fundamental expressive-power gap between kernel and full attention.

0 人收藏 0 人点赞
← 返回首页

提交意见反馈