function-token-anchoring

Tag

Cards List
#function-token-anchoring

Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models

arXiv cs.AI · 2026-07-24 Cached

This paper investigates short-term attention degradation in LLMs, finding a universal exponential-then-plateau pattern and that function token anchoring is architecture-dependent. Causal tests show that increasing attention mass on function tokens does not improve retrieval, suggesting attention degradation is descriptive rather than prescriptive.

0 favorites 0 likes
← Back to home

Submit Feedback