declarative-attention

Tag

Cards List
#declarative-attention

@ethantsliu: LLMs can control their own attention for long-context! During text generation, LLMs typically read the full KV cache at…

X AI KOLs Timeline · yesterday Cached

The paper introduces declarative attention, a technique where LLMs explicitly declare which context segments to attend to, reducing token usage by up to 52% with minimal accuracy drops.

0 favorites 0 likes
#declarative-attention

Language Models Can Control Their Own Attention

Hugging Face Daily Papers · 2026-09-02 Cached

This paper introduces Declarative Attention, a method that enables language models to declare relevant context regions during inference, significantly reducing KV cache reads with small accuracy trade-offs in long-context tasks.

0 favorites 0 likes
← Back to home

Submit Feedback