Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Summary
Gated DeltaNet-2 introduces separate erase and write gates for linear attention, achieving superior performance in long-context language modeling and retrieval tasks.
View Cached Full Text
Cached at: 05/22/26, 02:29 AM
Paper page - Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Source: https://huggingface.co/papers/2605.22791
Abstract
Gated DeltaNet-2 improves upon existing linear attention models by separating erase and write operations through distinct channel-wise gates, achieving superior performance in long-context language modeling and retrieval tasks.
Linear attentionreplaces the unbounded cache ofsoftmax attentionwith a fixed-sizerecurrent state, reducing sequence mixing to linear time and decoding to constant memory. The hard part is not just what to forget, but how to edit this compressed memory without scrambling existing associations.Delta-rule modelssubtract the current read before writing a new value, andKimi Delta Attention(KDA) sharpens forgetting withchannel-wise decay. But the active edit still uses a single scalar gate to control two different things: how much old content to erase on the key side and how much new content to commit on the value side. We introduceGated DeltaNet-2, which generalizes bothGated DeltaNetand KDA by inheriting adaptive forgetting andchannel-wise decaywhile addressing their shared limitation, the scalar tie between erasing and writing. Gated Delta Rule-2 separates these roles with a channel-wiseerase gateb_t and a channel-wisewrite gatew_t, reducing to KDA when both gates collapse to the same scalar and toGated DeltaNetwhen the decay also collapses. We derive afast-weight updateview, achunkwise WY algorithmwithchannel-wise decayabsorbed into asymmetric erase factors, and agate-aware backward passthat preserves efficient parallel training. At 1.3B parameters trained on 100B FineWeb-Edu tokens,Gated DeltaNet-2 achieves the strongest overall results amongMamba-2,Gated DeltaNet, KDA, andMamba-3variants across language modeling, commonsense reasoning, and retrieval. Its advantage is most pronounced on long-contextRULERneedle-in-a-haystack benchmarks, where it improves the evaluated multi-key retrieval setting and remains strong in both recurrent and hybrid settings. Code is available at https://github.com/NVlabs/GatedDeltaNet-2.
View arXiv pageView PDFGitHub19Add to collection
Get this paper in your agent:
hf papers read 2605\.22791
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.22791 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.22791 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.22791 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Introducing Toast 1
Mixedbread introduces Toast 1, a specialized search agent that matches frontier model quality while being up to 10x cheaper and 12x faster. It automates agentic search loops and achieves state-of-the-art results on benchmarks like OfficeQA Pro V2 and legal knowledge tasks.
MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
This paper introduces MARCH, a network architecture that scales recurrent state-space models beyond fixed-size dimensions by caching cumulative recurrent-state checkpoints as content-addressable state anchors, enabling efficient long-range memory retrieval that outperforms linear attention variants on long-context benchmarks.
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
This paper identifies post-retrieval reuse as a bottleneck for long-horizon agent memory and proposes query-conditioned reuse (QCR), a simple target-bound note format, showing improved success and token efficiency across WebArena, WorkArena, and AppWorld.
When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory
This paper introduces ReFind, an agent-controlled search interface that works over raw, unmodified chat logs using lexical indexing and iterative keyword search, showing it can rival structured memory systems like HippoRAG 2 on conversational-memory benchmarks without building any semantic structure.
The Embedder's Dilemma: LLMs Are Better, but at What Cost?
This paper presents a cost-aware comparison of LLMs versus dedicated embedding models across 37 tasks, finding that the best LLM and embedding model are nearly tied on aggregate performance but LLMs are up to 1,431x more expensive and slower, leading to a recommended division of labor.