softmax-attention

Tag

Cards List
#softmax-attention

Phases in a class of associative memories via hidden neurons

arXiv cs.LG · 2d ago Cached

The paper analyzes a class of associative memories with hidden neurons, deriving phase diagrams and storage capacities using replica methods and linking to softmax attention in transformers.

0 favorites 0 likes
#softmax-attention

[R] All Routes Lead to Collapse: attention sinks, representation collapse, and norm stratification are what content-based routing does under a norm-blind metric

Reddit r/MachineLearning · 2026-06-25 Cached

This paper demonstrates that attention sinks, representation collapse, and norm stratification are not unique to attention mechanisms but are general consequences of content-based routing under a norm-blind similarity metric, as shown across multiple architectures including transformers, graph attention, state-space models, and recurrent mixers.

0 favorites 0 likes
#softmax-attention

@tilderesearch: https://x.com/tilderesearch/status/2061771450168889432

X AI KOLs Timeline · 2026-06-02 Cached

Wall Attention generalizes diagonal forget gates to softmax attention, enabling state-of-the-art length extrapolation from 4k to 160k+ context zero-shot and outperforming RoPE and FoX in pretraining. It is released as a drop-in replacement with open-source Triton kernels.

0 favorites 0 likes
← Back to home

Submit Feedback