Tag
The article introduces two attention mechanisms, CoWindow Attention and MassAlloc Attention, which optimize compute allocation in transformers, achieving significant speedups and reduced training FLOPs in benchmarks.
An interactive diagram comparing self-attention and cross-attention mechanisms in AI models, published as part of an educational library on attention.
This paper presents an interpretability study on video diffusion models, revealing that Rotary Position Embedding (RoPE) induces excessive spatial attention decay, causing physics violations, and proposes a lightweight architectural modification to enhance physical coherence in generated videos.
This paper investigates whether the recurrence in selective state space models drives tokens to consensus similar to attention in transformers, using dynamical systems theory to analyze stability and attraction domains for time-varying weight matrices.
This paper introduces a mean-field analysis of attention that predicts average representation dynamics and reveals context-specific computation in language models, validated across models like GPT-2, Pythia, and Qwen-3-14B.
This article organizes 13 attention mechanisms in AI by the bottleneck they solve, covering KV cache reduction, attention patterns, compute efficiency, and serving efficiency to help AI engineers understand and apply these techniques.
This paper presents a mathematical framework for reverse-engineering transformer models to enhance mechanistic interpretability and address safety concerns in AI systems.
Attention-DP3 enhances 3D diffusion policies by incorporating object-level geometric cues through attention to improve stability in cluttered environments, achieving state-of-the-art performance across benchmarks.
A user announces they have solved all attention-related problems on LeetGPU and compiled them into a resource for easy access.
Sliding window attention with sinks outperforms post-trained linear attention on long-context tasks without retraining, offering a cheaper and more reliable inference solution for LLMs.
This paper presents a systematic empirical study of energy consumption in large language model inference, analyzing how attention architectures like Multi-Head Attention, Grouped Query Attention, and Sliding Window Attention affect energy scaling across context lengths and workloads.
ReWorld presents an interactive world model that decouples short-horizon control from long-horizon memory using mixed attention windows and distribution-matching LoRA distillation, enabling real-time video generation with high action fidelity and recall.
ReCache is a framework for efficient KV cache reuse and compression in tool-augmented LLM agents, achieving significant speedup and memory reduction while maintaining performance.
A tweet highlights 'Understanding Transformers and Attention Mechanisms,' a 13-page paper that explains the Transformer architecture and attention from an applied mathematics perspective.
A tweet recommends an arXiv paper that explains the mathematical foundations of Transformers, covering tokenization, embeddings, multi-headed attention, and KV caching for applied mathematicians.
This paper introduces the Entropic Bound, a spectral measure of task-intrinsic capacity for transformers, proving that the intrinsic rank of the token-mixing operator provides a tight lower bound on required model capacity. It shows that while a naive transfer from linear attention fails for real attention, an attention-native intrinsic rank restores the full theoretical structure.
A novel generalist controller using attention mechanisms and mixture-of-experts is proposed, enabling a single neural network to control diverse dynamical systems without system-specific tuning. It achieves comparable performance to traditional controllers across 25 different systems.
A researcher questions the reproducibility of MLA outperforming GQA under same KV cache, sharing early small-scale ablation results and plans for scaling experiments to decide on architecture for next large-scale run.
This paper presents a mechanistic analysis of induction in masked diffusion language models, identifying a bidirectional induction circuit and showing that these models use the global fraction of masked tokens as an implicit timestep.
This paper identifies the 'question-first paradox' in vision-language models, where placing the question before the image hinders answer accuracy despite improving visual attention. It proposes a training-free technique called 'question echoing'—restating the question before and after the image—which improves performance across multiple benchmarks without architecture changes.