Tag
Peking University researchers introduce NAMOH, an architecture-native sparse attention mechanism that activates K of H heads per token, so head routing jointly determines active parameters and available context. It enables parameter scaling to directly support efficient long-context scaling, outperforming fully activated models with the same total parameters while remaining compatible with GQA and existing sparse attention methods.
SMat-Attention introduces structured causal masks with tunable VC-dimension to bridge softmax attention and linear attention, achieving subquadratic prefill, constant-time decoding with O(T^{1-1/d}) cached states, and improved recall over Mamba-2 and Gated DeltaNet backbones.
WorldAttention proposes an efficient attention architecture with Hybrid Sparse Attention and Hierarchical KV Cache for interactive video world models, achieving state-of-the-art performance on benchmarks like VBench-Long and InterVBench.
This paper analyzes the evolution of attention routing in recurrent language models and proposes WISE, a training-free inference method that reuses stabilized sparse attention support to achieve up to 1.76× attention speedup while preserving performance on multi-hop QA benchmarks.
Elastic Threshold Attention (ETA) is a trainable sparse attention architecture that improves long-context decoding speed without quality degradation by using dynamic thresholds predicted from query representations.
RBS-Attention introduces a training-free sparse-prefill method with dual-branch selection to mitigate mean dilution in long-context LLM inference, achieving up to 20.65× speedup on H100 GPUs while maintaining near-dense quality on benchmarks.
This paper introduces Adaptive-Scalable Entmax (ASEntmax), a learnable sparse attention mechanism that enables up to 1000× length extrapolation in transformers, improving long-context generalization while preserving short-context performance.
HiSparse and sparse attention techniques are being integrated into vLLM to reduce memory usage and increase concurrency, improving inference throughput for reinforcement learning.
This paper introduces Declarative Attention, a protocol that allows language models to declare where to attend in their chain-of-thought, reducing attended tokens by 52% on Gemma-4-31B and 31.1% on Qwen-3.6-27B with minimal accuracy drops.
Arize Phoenix introduces M3, which adds sparse attention and a 1M-token context window to reduce latency and cost in AI agents by keeping long tool histories in context.
CRISP introduces a method for improving sparse prefilling in long-context LLM inference by using structural routing to address noise accumulation, achieving speedups and performance gains on retrieval tasks.
The paper introduces Graph Machine, a method to replace dense attention layers in transformers with sparse layers using dynamic pointers, improving efficiency and maintaining or enhancing performance during pretraining.
This paper introduces Declarative Attention, a method that enables language models to declare relevant context regions during inference, significantly reducing KV cache reads with small accuracy trade-offs in long-context tasks.
RouteSparse is a method for input-conditional pattern routing in long-context prefilling, achieving up to 6.5× speedup with minimal accuracy loss compared to dense attention.
The paper presents Qwen3.8-Flash-Next, a sparse mixture-of-experts AI model that enhances efficiency, capability, and training stability through architectural innovations like hybrid attention and n-gram embeddings.
The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.
GLM-5.3, an AI model with unchanged architecture from GLM-5.2, has achieved significant improvements through training and is now visualizable in HF Viewer, showcasing complex features like sparse attention and MoE routing.
Tencent releases Hy4-preview, a new Mixture-of-Experts flagship AI model with 770B total parameters and 49B activated per token, featuring advanced techniques like Gated DeepSeek Sparse Attention and identity Hyper-Connections.
This paper introduces ClusterAttention, a training-free method to speed up bidirectional attention in transformers by using recursive clustering for block-sparse attention, achieving 2-6x speedups on tabular data and 1.8x on video generation while maintaining high accuracy.
This paper analyzes KV-cache eviction strategies in sparse-attention models, showing that selecting largest weights is near-optimal and that published margins come from memory and query information, with ContourKV achieving strong performance.