entmax

Tag

Cards List
#entmax

@peony__snow: Up to 1000× length extrapolation—just by replacing softmax. Dense attention disperses over long inputs. ASEntmax gives …

X AI KOLs Timeline ↗ · 2026-09-11 Cached

This paper introduces Adaptive-Scalable Entmax (ASEntmax), a learnable sparse attention mechanism that enables up to 1000× length extrapolation in transformers, improving long-context generalization while preserving short-context performance.

0 favorites 0 likes
#entmax

EntmaxKV: Support-Aware Decoding for Entmax Attention

arXiv cs.LG ↗ · 2026-05-22 Cached

EntmaxKV introduces a support-aware sparse decoding framework for entmax attention that reduces KV-cache memory traffic by exploiting sparsity before loading pages, achieving significant speedups on long-context benchmarks while maintaining output quality.

0 favorites 0 likes
← Back to home

Submit Feedback