The article proposes a Multi-Key Attention idea where a single attention head uses multiple Key projections combined non-compensatorily before softmax to better handle multi-constraint retrieval tasks, and seeks feedback on potential flaws and existing benchmarks.
One-sentence version: what if, within an attention head, we replaced a single Q–K compatibility score with a learned, potentially non-compensatory combination of several compatibility scores before the softmax — and tested whether that helps with multi-constraint retrieval? Standard attention asks "how compatible are Q and K?" via one dot product. Multi-head attention gets multiple compatibility judgments, but combines them after separate softmaxes and separate retrievals — each head fully commits to its own attention distribution before anything is merged. I'm asking what happens if you instead give a single head multiple Key projections (K₁, K₂, K₃), compute separate scores s₁, s₂, s₃ against the same Query, and combine those scores before the softmax: s = f(s1, s2, s3) → softmax(s) → one retrieval instead of the usual separate softmaxes → separate retrievals → combine. Simple summing turns out to be a no-op: Q·K1 + Q·K2 + Q·K3 = Q·(K1+K2+K3), i.e. just another single Key. So the interesting cases are nonlinear combiners — specifically non-compensatory ones like soft-min, where a candidate that's excellent on one dimension but fails another (1.0, 0.0) loses to one that's moderately good on both (0.4, 0.4). That's a genuinely different notion of relevance than the weighted-sum default, and it maps onto "AND"-style aggregation from multi-criteria decision theory / fuzzy t-norms rather than anything attention typically does. Known relatives, so I'm not wasting your time reinventing the wheel: Talking-Heads Attention (Shazeer 2020) — mixes logits across heads pre- and post-softmax. Closest existing "combine before softmax" mechanism; the difference here is multiple Keys within one head rather than mixing across heads, and a nonlinear/non-compensatory combiner rather than a linear one. Synthesizer (Tay 2020) — learned/random attention patterns instead of dot-product. Found dot-product hard to beat in general, which is why I think the honest claim is narrower: not "better scoring in general," but "better on tasks that are genuinely multi-constraint and non-compensatory." Compositional Attention (Mittal et al., ICLR'22) — disentangles search/retrieval pairing dynamically. Different axis of flexibility, but the nearest published architecture; planning to use it as a baseline. What would kill this: if a wider standard attention head matches it param-for-param, if Talking-Heads matches it, if the nonlinear combiner just learns to behave linearly, or if the benefit disappears at scale. Any of those would be a genuinely useful negative result. Test I'd start with: synthetic retrieval task where correct answer requires satisfying entity + relationship + time jointly; candidates get graded per-dimension scores; compare sum vs. soft-min vs. learned gate vs. Talking-Heads vs. Compositional Attention, param- and compute-matched. Genuinely looking for: (1) is this reducible to one of the above, (2) is there an existing benchmark I should be using instead of building a synthetic one, (3) obvious reason soft-min attention breaks during training that I'm not seeing.
This paper establishes theoretical bounds on the number of attention heads needed to produce vector representations that support multiple tasks, such as computing min/max and XOR, showing trade-offs between head count, embedding dimension, and precision.
Proposes Interdomain Attention, a new method that integrates state space models into attention via kernel methods, achieving efficient long-context modeling with a fixed-size state and outperforming SSMs and softmax attention in language modeling experiments up to 1.3B parameters.
This paper critiques standard attention metric methods in transformers, revealing that choices like keeping or dropping sink tokens can reverse conclusions, and proposes using compositional data analysis to separate sink and content attention components for more accurate interpretation.
This paper identifies and formalizes the 'structural attention tax' phenomenon, where the format of retrieved content (e.g., knowledge graph triples) independently distorts LLM attention distribution regardless of semantic relevance, leading to compressed demonstration attention. It provides a formal framework, empirical evidence across models and benchmarks, and proposes structure-aware mitigation strategies.
Functional Attention is a novel attention mechanism that reinterprets attention as a functional correspondence between adaptive bases, replacing softmax affinities with structured linear operators inspired by geometric functional maps. The method achieves state-of-the-art performance on operator learning tasks including PDE solving and 3D segmentation while remaining resolution-invariant.