linear-attention

Tag

Cards List
#linear-attention

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Hugging Face Daily Papers ↗ · 2026-09-17 Cached

Video DeltaNet presents a hybrid attention mechanism combining Softmax and linear attention to enhance efficiency in video generation models, achieving a 14.5x speedup over baseline methods.

0 favorites 0 likes
#linear-attention

Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It

arXiv cs.LG ↗ · 2026-09-16 Cached

This paper decomposes associative recall in fixed-state recurrences, finding convolution and curriculum learning key to performance, and proposes interventions to address interference rather than capacity limitations.

0 favorites 0 likes
#linear-attention

SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization

Hugging Face Daily Papers ↗ · 2026-09-13 Cached

SpectralShift introduces a spectral reparameterization approach to effectively extend the context window of Gated DeltaNet models by reshaping the decay spectrum, improving long-context capabilities through continual pretraining.

0 favorites 0 likes
#linear-attention

Kalman Delta Networks: Uncertainty-aware Associative Memory

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

Introduces Kalman Delta Networks, which improve language modeling by reformulating linear attention as a linear-Gaussian state-space model with Kalman-filter updates to track memory uncertainty, yielding efficient approximations that outperform existing linear-attention models.

0 favorites 0 likes
#linear-attention

Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design

arXiv cs.LG ↗ · 2026-09-04 Cached

This paper proposes intervention-based metrics to differentiate retrieval and positional heads in RoPE Transformers, leading to a principled hybrid architecture (HwH) that combines full and linear attention for improved language modeling and long-context extrapolation.

0 favorites 0 likes
#linear-attention

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

This paper explains why fully quantizing the recurrent Gated DeltaNet layers in hybrid LLMs to 4-bit NVFP4 preserves accuracy across long-context benchmarks, detailing mechanisms like outlier localization and robust delta-rule dynamics.

0 favorites 0 likes
#linear-attention

Test Time Training (3 minute read)

TLDR AI ↗ · 2026-09-03 Cached

The article discusses Test Time Training as a potential new scaling axis in AI development, analyzing a paper that frames it as a form of linear attention and exploring its implications for model training and continual learning.

0 favorites 0 likes
#linear-attention

Sliding-window beats linear attention

Hugging Face Daily Papers ↗ · 2026-08-28 Cached

Sliding window attention with sinks outperforms post-trained linear attention on long-context tasks without retraining, offering a cheaper and more reliable inference solution for LLMs.

0 favorites 0 likes
#linear-attention

@davsca1: Check out our #ECCV2026 paper "Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention", where …

X AI KOLs Timeline ↗ · 2026-08-26 Cached

This paper introduces Spatially-Sparse Linear Attention for low-latency object detection with event cameras, achieving state-of-the-art accuracy with 20x less computation than prior asynchronous methods.

0 favorites 0 likes
#linear-attention

Qwen 4 architecture: What do we know?

Reddit r/LocalLLaMA ↗ · 2026-08-25

The author speculates on the possible architecture of Qwen 4, discussing potential designs like embedding-offloaded linear attention or multi-head latent attention hybrids based on reverse engineering Deepseek.

0 favorites 0 likes
#linear-attention

The Query Knows What to Forget: A Second Erase Direction for Linear Attention

arXiv cs.LG ↗ · 2026-08-17 Cached

Introduces a Query-derived Erase Direction (QED) for linear attention models to improve retrieval and about double the usable context length, as demonstrated on the S-NIAH-1 benchmark.

0 favorites 0 likes
#linear-attention

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

Hugging Face Daily Papers ↗ · 2026-08-12 Cached

This paper presents the first systematic study of massive activations in hybrid linear-attention LLMs, uncovering pre-attention spikes and inter-spike plateaus governed by cancellation timing, and showing how their morphology recovers at full-attention limits.

0 favorites 0 likes
#linear-attention

Retrofitting Linear Attention into Diffusion Language Models

arXiv cs.LG ↗ · 2026-08-10 Cached

This paper introduces block-hybrid attention, which combines exact softmax attention within active denoising blocks and linear attention over previous blocks, to accelerate inference in pretrained diffusion language models. The authors retrofit this hybrid attention into LLaDA 2.1, achieving up to 1.7x higher decoding throughput with minimal post-training.

0 favorites 0 likes
#linear-attention

Stuck on "A": Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model

arXiv cs.CL ↗ · 2026-08-05 Cached

This paper investigates the failure modes of converting Qwen3-0.6B-Base attention layers to KDA linear attention on a single GPU, identifying an 'interface injury' where the model predicts option labels rather than content, and proposes a format-targeted KL stage to repair it.

0 favorites 0 likes
#linear-attention

@eliebakouch: so... the two biggest oss models in the world use linear attention?

X AI KOLs Timeline ↗ · 2026-08-03 Cached

A tweet highlights that major open-source models like Qwen3.8-Max may use linear attention, referencing Alibaba's announcement of upcoming open-weight releases for Qwen3.8-Max (2.4T parameters) and Qwen3.8-27B.

0 favorites 0 likes
#linear-attention

Understand Kimi K3 from first principles: a recommended order for anyone trying to understand this beast

Reddit r/ArtificialInteligence ↗ · 2026-07-29

A guide recommending a reading order of foundational papers and Kimi model reports to understand the architecture of Moonshot AI's Kimi K3, covering linear attention, MoE, and residual connections.

0 favorites 0 likes
#linear-attention

You Could Have Come Up with Kimi Delta Attention

Hacker News Top ↗ · 2026-07-28 Cached

This blog post derives Kimi Delta Attention step by step from standard softmax attention through linear attention and DeltaNet variants, explaining the state update equations used by recent Qwen and Kimi models.

0 favorites 0 likes
#linear-attention

Kimi Linear: An Expressive, Efficient Attention Architecture

Hacker News Top ↗ · 2026-07-28 Cached

Kimi Linear proposes a new linear attention architecture designed to enhance both expressiveness and efficiency in Transformer models, with contributions from the Kimi Team at Moonshot AI.

0 favorites 0 likes
#linear-attention

Nvidia's New Long-Form Video Generation (12 minute read)

TLDR AI ↗ · 2026-07-27 Cached

NVIDIA Research introduces SANA-Video 2.0, a hybrid video diffusion transformer that generates high-quality 720p video on a single GPU, achieving up to 120× speedup over Wan 2.2-14B via hybrid linear-softmax attention and block attention residuals.

0 favorites 0 likes
#linear-attention

@akshay_pachaar: one matrix replaced the KV cache. (the technique is 100% open source) Kimi just dropped K3, an open model at frontier s…

X AI KOLs Following ↗ · 2026-07-24 Cached

Kimi released K3, a 2.8T-parameter open model using delta attention to avoid growing KV cache, enabling a 1-million-token context window with linear memory cost.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback