Tag
This paper introduces Gaussian Mixture Attention (GMA), a probabilistic attention mechanism that replaces explicit pairwise query-key comparisons with routing through learned Gaussian mixture components, achieving linear-time complexity in sequence length. Experiments show competitive performance on long-context tasks with fixed-K linear memory scaling.