retrieval-heads

Tag

Cards List
#retrieval-heads

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

arXiv cs.CL · 2026-07-02 Cached

The paper introduces LOGOS, a write-aware detector that identifies attention heads responsible for non-literal retrieval in LLMs by scoring the projection of their OV-circuit output onto the answer-token unembedding direction, outperforming prior attention-based methods across multiple model families.

0 favorites 0 likes
#retrieval-heads

Rethinking the Role of Efficient Attention in Hybrid Architectures

arXiv cs.CL · 2026-06-16 Cached

This paper systematically analyzes the role of efficient attention modules in hybrid language model architectures, finding that different designs converge in long-context performance under sufficient training, and that long-range retrieval is primarily carried by full attention while efficient attention shapes the optimization trajectory, revealing a 'Large-Window Laziness' phenomenon.

0 favorites 0 likes
← Back to home

Submit Feedback