Tag
This paper introduces GEAR, a Geometry-Enabled Attention Routing framework that uses geometry as explicit token-level addresses for visual memory to enhance long-horizon camera-controlled video generation, achieving state-of-the-art results.
This paper analyzes the evolution of attention routing in recurrent language models and proposes WISE, a training-free inference method that reuses stabilized sparse attention support to achieve up to 1.76× attention speedup while preserving performance on multi-hop QA benchmarks.
This paper tests the assumption that attention weights reveal what a model actually depends on for its output, finding that attention and causal dependence often disagree. They propose using causal evidence sets obtained via intervention masking as supervision for sparse attention routers, achieving near-perfect accuracy on retrieval tasks where attention-distilled routers fail.
Introduces Queryable LoRA, a data-adaptive method for efficient fine-tuning that uses a shared memory of low-rank update atoms with attention-based routing and instruction regularization to enable dynamic, context-sensitive parameter updates while maintaining scalability.