SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization

Hugging Face Daily Papers Papers

Summary

SpectralShift introduces a spectral reparameterization approach to effectively extend the context window of Gated DeltaNet models by reshaping the decay spectrum, improving long-context capabilities through continual pretraining.

Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linear attention state dynamics. In this work, we study long-context extension of Gated DeltaNet (GDN) from a spectral perspective of transition matrix and identify two essential factors governing long-range information retrieval: (1) a sufficiently broad slow spectral band aligned with the target dependency length, and (2) the preservation of fast-decaying modes for state clearing and context switching. Based on this observation, we propose SpectralShift, a spectral reparameterization approach for long-context continual pretraining of GDNs. Specifically, SpectralShift reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training. Experiments show that SpectralShift consistently improves long-context capabilities over training, providing an effective and efficient solution for extending context windows of linear attention models. The code has been open-sourced at https://github.com/RUCAIBox/GDN-SpectralShift.
Original Article
View Cached Full Text

Cached at: 09/17/26, 06:52 AM

Paper page - SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization

Source: https://huggingface.co/papers/2609.14320

Abstract

Recently,linearattentionlayershavebeenincreasinglyadoptedtoreplacesoftmaxattentionatscaleforlong-contextmodeling.However,existingcontextextensionapproachestypicallyapplycontinuedpretrainingdirectlywithoutmodifyingtheselayers,overlookingthespectralpropertiesoflinearattentionstatedynamics.Inthiswork,westudylong-contextextensionofGatedDeltaNet(GDN)fromaspectralperspectiveoftransitionmatrixandidentifytwoessentialfactorsgoverninglong-rangeinformationretrieval:(1)asufficientlybroadslowspectralbandalignedwiththetargetdependencylength,and(2)thepreservationoffast-decayingmodesforstateclearingandcontextswitching.Basedonthisobservation,weproposeSpectralShift,aspectralreparameterizationapproachforlong-contextcontinualpretrainingofGDNs.Specifically,SpectralShiftreparameterizesthealphaprojectionsinitializationtoreshapethedecayspectrumbyenhancingslowpropagationcapacity,andfurtherintroducesalearning-ratescalingforalphaprojectionstofacilitatelong-contexttraining.ExperimentsshowthatSpectralShiftconsistentlyimproveslong-contextcapabilitiesovertraining,providinganeffectiveandefficientsolutionforextendingcontextwindowsoflinearattentionmodels.Thecodehasbeenopen-sourcedathttps://github.com/RUCAIBox/GDN-SpectralShift.

View arXiv pageView PDFGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.14320

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.14320 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.14320 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.14320 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Spectral Tempering for Embedding Compression in Dense Passage Retrieval

arXiv cs.CL

Spectral Tempering (SpecTemp) proposes a learning-free method for embedding compression in dense passage retrieval that adaptively determines optimal spectral scaling based on signal-to-noise ratio analysis, outperforming fixed hyperparameter approaches like PCA and whitening.

Preference Tuning as Spectral Update Reorganization

arXiv cs.CL

The paper reveals that preference-based post-training induces parameter updates with a spectral head-tail organization, where a compact head carries the dominant behavioral shift and a weak tail is necessary for full solution recovery, recasting alignment as structured update reorganization rather than monolithic correction.