working-set

Tag

Cards List
#working-set

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

arXiv cs.CL ↗ · 2026-09-24 Cached

This paper analyzes the evolution of attention routing in recurrent language models and proposes WISE, a training-free inference method that reuses stabilized sparse attention support to achieve up to 1.76× attention speedup while preserving performance on multi-hop QA benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback