Tag
The article proposes a Multi-Key Attention idea where a single attention head uses multiple Key projections combined non-compensatorily before softmax to better handle multi-constraint retrieval tasks, and seeks feedback on potential flaws and existing benchmarks.