Tag
This paper introduces Adaptive-Scalable Entmax (ASEntmax), a learnable sparse attention mechanism that enables up to 1000× length extrapolation in transformers, improving long-context generalization while preserving short-context performance.
EntmaxKV introduces a support-aware sparse decoding framework for entmax attention that reduces KV-cache memory traffic by exploiting sparsity before loading pages, achieving significant speedups on long-context benchmarks while maintaining output quality.