local-attention

Tag

Cards List
#local-attention

@HarshalsinghCN: yooo guys, the blog is up. i tried to break down the design of BarunLM(35M) in simple language while keeping as much te…

X AI KOLs Timeline · 2026-08-02 Cached

Announcement of a blog post explaining the design of BarunLM, a 35M-parameter language model, covering its architecture, training recipe, and dataset preparation, with a focus on efficiency and performance gains over other sub-100M models.

0 favorites 0 likes
#local-attention

Dual Dimensionality for Local and Global Attention

arXiv cs.CL · 2026-06-18 Cached

Proposes Distance-Adaptive Representation (DAR) which reduces key-value dimensionality for distant tokens while preserving full dimensionality for nearby tokens, improving KV cache efficiency without performance loss.

0 favorites 0 likes
#local-attention

@akshay_pachaar: 1) Sparse Attention It limits the attention computation to a subset of tokens by: - Using local attention (tokens atten…

X AI KOLs Timeline · 2026-06-03 Cached

Explains sparse attention in transformers, which reduces computational complexity by attending only to a subset of tokens using local or learned attention patterns.

0 favorites 0 likes
← Back to home

Submit Feedback