sparse-attention

Tag

Cards List
#sparse-attention

SMat-Attention: Structured Long-Context Sequence Modeling

arXiv cs.AI ↗ · 16h ago Cached

SMat-Attention introduces structured causal masks with tunable VC-dimension to bridge softmax attention and linear attention, achieving subquadratic prefill, constant-time decoding with O(T^{1-1/d}) cached states, and improved recall over Mamba-2 and Gated DeltaNet backbones.

0 favorites 0 likes
#sparse-attention

WorldAttention: An Efficient Attention Architecture for Interactive Video World Models

Hugging Face Daily Papers ↗ · 2d ago Cached

WorldAttention proposes an efficient attention architecture with Hybrid Sparse Attention and Hierarchical KV Cache for interactive video world models, achieving state-of-the-art performance on benchmarks like VBench-Long and InterVBench.

0 favorites 0 likes
#sparse-attention

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

arXiv cs.CL ↗ · 6d ago Cached

This paper analyzes the evolution of attention routing in recurrent language models and proposes WISE, a training-free inference method that reuses stabilized sparse attention support to achieve up to 1.76× attention speedup while preserving performance on multi-hop QA benchmarks.

0 favorites 0 likes
#sparse-attention

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

arXiv cs.LG ↗ · 2026-09-21 Cached

Elastic Threshold Attention (ETA) is a trainable sparse attention architecture that improves long-context decoding speed without quality degradation by using dynamic thresholds predicted from query representations.

0 favorites 0 likes
#sparse-attention

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

arXiv cs.AI ↗ · 2026-09-21 Cached

RBS-Attention introduces a training-free sparse-prefill method with dual-branch selection to mitigate mean dilution in long-context LLM inference, achieving up to 20.65× speedup on H100 GPUs while maintaining near-dense quality on benchmarks.

0 favorites 0 likes
#sparse-attention

@peony__snow: Up to 1000× length extrapolation—just by replacing softmax. Dense attention disperses over long inputs. ASEntmax gives …

X AI KOLs Timeline ↗ · 2026-09-11 Cached

This paper introduces Adaptive-Scalable Entmax (ASEntmax), a learnable sparse attention mechanism that enables up to 1000× length extrapolation in transformers, improving long-context generalization while preserving short-context performance.

0 favorites 0 likes
#sparse-attention

@samsja19: we are super bullish on sparse attention + HiSparse and have been working with the vLLM team on it sparse attention red…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

HiSparse and sparse attention techniques are being integrated into vLLM to reduce memory usage and increase concurrency, improving inference throughput for reinforcement learning.

0 favorites 0 likes
#sparse-attention

@omarsar0: Banger paper from Google DeepMind and colleagues. (bookmark it) A model reads its entire KV cache on every generated to…

X AI KOLs Following ↗ · 2026-09-03 Cached

This paper introduces Declarative Attention, a protocol that allows language models to declare where to attend in their chain-of-thought, reducing attended tokens by 52% on Gemma-4-31B and 31.1% on Qwen-3.6-27B with minimal accuracy drops.

0 favorites 0 likes
#sparse-attention

@ArizePhoenix: For agents, active parameters are what you pay for in latency and cost, and agent loops resend context, retry tool call…

X AI KOLs Following ↗ · 2026-09-03

Arize Phoenix introduces M3, which adds sparse attention and a 1M-token context window to reduce latency and cost in AI agents by keeping long tool histories in context.

0 favorites 0 likes
#sparse-attention

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing

arXiv cs.LG ↗ · 2026-09-03 Cached

CRISP introduces a method for improving sparse prefilling in long-context LLM inference by using structural routing to address noise accumulation, achieving speedups and performance gains on retrieval tasks.

0 favorites 0 likes
#sparse-attention

Graph Machine: Towards Better Pretraining via Edges

Hugging Face Daily Papers ↗ · 2026-09-02 Cached

The paper introduces Graph Machine, a method to replace dense attention layers in transformers with sparse layers using dynamic pointers, improving efficiency and maintaining or enhancing performance during pretraining.

0 favorites 0 likes
#sparse-attention

Language Models Can Control Their Own Attention

Hugging Face Daily Papers ↗ · 2026-09-02 Cached

This paper introduces Declarative Attention, a method that enables language models to declare relevant context regions during inference, significantly reducing KV cache reads with small accuracy trade-offs in long-context tasks.

0 favorites 0 likes
#sparse-attention

RouteSparse: Input-Conditional Pattern Routing for Budgeted Long-Context Prefilling

arXiv cs.CL ↗ · 2026-09-01 Cached

RouteSparse is a method for input-conditional pattern routing in long-context prefilling, achieving up to 6.5× speedup with minimal accuracy loss compared to dense attention.

0 favorites 0 likes
#sparse-attention

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

The paper presents Qwen3.8-Flash-Next, a sparse mixture-of-experts AI model that enhances efficiency, capability, and training stability through architectural innovations like hybrid attention and n-gram embeddings.

0 favorites 0 likes
#sparse-attention

Qwen3.8-Flash-Next optimised for Macs

Reddit r/LocalLLaMA ↗ · 2026-08-30

The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.

0 favorites 0 likes
#sparse-attention

GLM-5.3 on HF Viewer

Reddit r/LocalLLaMA ↗ · 2026-08-28

GLM-5.3, an AI model with unchanged architecture from GLM-5.2, has achieved significant improvements through training and is now visualizable in HF Viewer, showcasing complex features like sparse attention and MoE routing.

0 favorites 0 likes
#sparse-attention

Tencent/Hy4-preview 770B-A49B weight dropped

Reddit r/LocalLLaMA ↗ · 2026-08-28 Cached

Tencent releases Hy4-preview, a new Mixture-of-Experts flagship AI model with 770B total parameters and 49B activated per token, featuring advanced techniques like Gated DeepSeek Sparse Attention and identity Hyper-Connections.

0 favorites 0 likes
#sparse-attention

ClusterAttention: A training-free speedup of bidirectional attention

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper introduces ClusterAttention, a training-free method to speed up bidirectional attention in transformers by using recursive clustering for block-sparse attention, achieving 2-6x speedups on tabular data and 1.8x on video generation while maintaining high accuracy.

0 favorites 0 likes
#sparse-attention

Trust the Mass: Forced Weights in KV-Cache Eviction

arXiv cs.LG ↗ · 2026-08-27 Cached

This paper analyzes KV-cache eviction strategies in sparse-attention models, showing that selecting largest weights is near-optimal and that published margins come from memory and query information, with ContourKV achieving strong performance.

0 favorites 0 likes
#sparse-attention

FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree

Hugging Face Models Trending ↗ · 2026-08-27 Cached

FastVideo releases FastH3 Preview v1, an AI model checkpoint that generates synchronized video and audio from text using four transformer forwards, trained with data-free DMD2 and VSA-H3 at 90% sparsity.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback