sliding-window-attention

Tag

Cards List
#sliding-window-attention

Maglev: Sliding Recurrent Memory

arXiv cs.LG · 5d ago Cached

Introduces Maglev, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. It uses a prefiller and decoder with a memory consistency loss, improving validation loss and downstream benchmarks over baselines.

0 favorites 0 likes
#sliding-window-attention

ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation

arXiv cs.CL · 2026-06-17 Cached

ConSA is a framework that learns optimal assignment between full attention and sliding-window attention under a user-specified sparsity target, using L0 regularization and augmented Lagrangian constraint. It demonstrates consistent gains over rule-based baselines on LLMs at 0.6B and 1.7B scales.

0 favorites 0 likes
#sliding-window-attention

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning

arXiv cs.AI · 2026-06-11 Cached

This paper introduces SWARR, a two-stage recipe using supervised fine-tuning and reinforcement learning to adapt sliding-window attention models for mathematical reasoning, showing that RL can narrow the performance gap with self-attention while maintaining efficiency.

0 favorites 0 likes
#sliding-window-attention

@_albertgu: Introducing a new sequence model Raven which pushes the boundary of fixed-state-size sequence models! Raven bridges pop…

X AI KOLs Timeline · 2026-05-07

Researchers introduce Raven, a novel sequence model that merges state space model efficiency with a selective slot-updating mechanism inspired by sliding window attention to improve long-context retrieval. The approach offers a more principled alternative to existing linear-time models.

0 favorites 0 likes
← Back to home

Submit Feedback