Tag
SMat-Attention introduces structured causal masks with tunable VC-dimension to bridge softmax attention and linear attention, achieving subquadratic prefill, constant-time decoding with O(T^{1-1/d}) cached states, and improved recall over Mamba-2 and Gated DeltaNet backbones.
WorldAttention proposes an efficient attention architecture with Hybrid Sparse Attention and Hierarchical KV Cache for interactive video world models, achieving state-of-the-art performance on benchmarks like VBench-Long and InterVBench.
This paper analyzes the evolution of attention routing in recurrent language models and proposes WISE, a training-free inference method that reuses stabilized sparse attention support to achieve up to 1.76× attention speedup while preserving performance on multi-hop QA benchmarks.
Elastic Threshold Attention (ETA) is a trainable sparse attention architecture that improves long-context decoding speed without quality degradation by using dynamic thresholds predicted from query representations.
RBS-Attention introduces a training-free sparse-prefill method with dual-branch selection to mitigate mean dilution in long-context LLM inference, achieving up to 20.65× speedup on H100 GPUs while maintaining near-dense quality on benchmarks.
This paper introduces Adaptive-Scalable Entmax (ASEntmax), a learnable sparse attention mechanism that enables up to 1000× length extrapolation in transformers, improving long-context generalization while preserving short-context performance.
HiSparse and sparse attention techniques are being integrated into vLLM to reduce memory usage and increase concurrency, improving inference throughput for reinforcement learning.
This paper introduces Declarative Attention, a protocol that allows language models to declare where to attend in their chain-of-thought, reducing attended tokens by 52% on Gemma-4-31B and 31.1% on Qwen-3.6-27B with minimal accuracy drops.
Arize Phoenix introduces M3, which adds sparse attention and a 1M-token context window to reduce latency and cost in AI agents by keeping long tool histories in context.
CRISP introduces a method for improving sparse prefilling in long-context LLM inference by using structural routing to address noise accumulation, achieving speedups and performance gains on retrieval tasks.
The paper introduces Graph Machine, a method to replace dense attention layers in transformers with sparse layers using dynamic pointers, improving efficiency and maintaining or enhancing performance during pretraining.
This paper introduces Declarative Attention, a method that enables language models to declare relevant context regions during inference, significantly reducing KV cache reads with small accuracy trade-offs in long-context tasks.
RouteSparse is a method for input-conditional pattern routing in long-context prefilling, achieving up to 6.5× speedup with minimal accuracy loss compared to dense attention.
The paper presents Qwen3.8-Flash-Next, a sparse mixture-of-experts AI model that enhances efficiency, capability, and training stability through architectural innovations like hybrid attention and n-gram embeddings.
The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.
GLM-5.3, an AI model with unchanged architecture from GLM-5.2, has achieved significant improvements through training and is now visualizable in HF Viewer, showcasing complex features like sparse attention and MoE routing.
Tencent releases Hy4-preview, a new Mixture-of-Experts flagship AI model with 770B total parameters and 49B activated per token, featuring advanced techniques like Gated DeepSeek Sparse Attention and identity Hyper-Connections.
This paper introduces ClusterAttention, a training-free method to speed up bidirectional attention in transformers by using recursive clustering for block-sparse attention, achieving 2-6x speedups on tabular data and 1.8x on video generation while maintaining high accuracy.
This paper analyzes KV-cache eviction strategies in sparse-attention models, showing that selecting largest weights is near-optimal and that published margins come from memory and query information, with ContourKV achieving strong performance.
FastVideo releases FastH3 Preview v1, an AI model checkpoint that generates synchronized video and audio from text using four transformer forwards, trained with data-free DMD2 and VSA-H3 at 90% sparsity.