length-generalization

Tag

Cards List
#length-generalization

On the Representational Geometry of Dynamic Programs

arXiv cs.LG · 2d ago Cached

This paper investigates why standard neural architectures fail to generalize to longer inputs in dynamic programming, using geometric analysis with tropical semiring theory to reveal structural limitations in compositions.

0 favorites 0 likes
#length-generalization

On Locality and Length Generalization in Visual Reasoning

Hugging Face Daily Papers · 2026-07-10 Cached

This paper shows that state-of-the-art vision-language models fail at length generalization in visual reasoning due to 'global shortcuts', and demonstrates that combining local foveated perception with recurrence enables robust out-of-distribution generalization.

0 favorites 0 likes
#length-generalization

@agopal42: Presenting PoPE today at #ICML2026! We revisit RoPE through the lens of content-position entanglement, and show how pol…

X AI KOLs Timeline · 2026-07-06 Cached

New paper introduces PoPE, a positional encoding that decouples content and position, addressing a fundamental flaw in RoPE used in many LLMs like Qwen, Gemma, DeepSeek. Presented at ICML2026.

0 favorites 0 likes
#length-generalization

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

arXiv cs.CL · 2026-07-03 Cached

This paper systematically studies in-context retrieval at million-token scale, introducing BlockSearch, a 0.6B LM retriever, and analyzing attention dilution. The model matches or outperforms dense retrieval on benchmarks like MS MARCO and NQ, and significantly outperforms on tasks requiring different similarity notions, highlighting the potential of in-context retrieval while emphasizing attention control under extreme context growth.

0 favorites 0 likes
#length-generalization

@tilderesearch: https://x.com/tilderesearch/status/2061771450168889432

X AI KOLs Timeline · 2026-06-02 Cached

Wall Attention generalizes diagonal forget gates to softmax attention, enabling state-of-the-art length extrapolation from 4k to 160k+ context zero-shot and outperforming RoPE and FoX in pretraining. It is released as a drop-in replacement with open-source Triton kernels.

0 favorites 0 likes
← Back to home

Submit Feedback