Dual-Interest Sequential Product Recommendation With Multi-Granular SSM

arXiv cs.AI Papers

Summary

DSRec is a novel dual-interest sequential recommendation model using multi-granular State Space Models to capture long-term and short-term user interests, addressing item polysemy and achieving superior performance on benchmarks.

arXiv:2609.21548v1 Announce Type: new Abstract: Sequential recommendation aims to predict the next item a user will interact with based on their historical behavior. Advances in Transformers have significantly improved sequential recommendation but are still limited by cost efficiency. Although State Space Models (SSMs) have recently enabled efficient long-range modeling, most existing methods encode each item with a single static contextual role, overlooking the phenomenon of item polysemy. In fact, the same item often plays different semantic roles depending on user context, and existing methods are limited in capturing dynamic behavior across different temporal granularities. In this work, we propose DSRec, a novel dual-interest cross-SSM model that explicitly disentangles item roles across long-term and short-term semantic context. Sequential items are encoded into long-term interest embeddings that capture stable preferences via historical aggregation, and a short-term interest branch that emphasizes local session intent modulated by inter-click time intervals. These interest embeddings are processed through distinct SSM encoders: a full-sequence Mamba for long-term modeling, and a time-modulated SSM that dynamically adjusts state evolution based on temporal gaps. To enable effective cross-granularity alignment, we adopt a residual cross-fusion mechanism that exchanges contextual information between the two branches while preserving semantic independence. Experiments on public benchmarks demonstrate that DSRec outperforms other state-of-the-art methods.
Original Article
View Cached Full Text

Cached at: 09/21/26, 09:27 AM

# Dual-Interest Sequential Product Recommendation With Multi-Granular SSM
Source: [https://arxiv.org/html/2609.21548](https://arxiv.org/html/2609.21548)
Shuiying LiaoAffiliation:The Hong Kong University of Science and Technology Clear Water Bay, Hong Kong shuiyingl@ust\.hkP\. Y\. Mok∗Affiliation:The Hong Kong University of Science and Technology Clear Water Bay, Hong Kong tracy\.mok@ust\.hk

###### Abstract

Sequential recommendation aims to predict the next item a user will interact with based on their historical behavior\. Advances in Transformers have significantly improved sequential recommendation but still limited by cost efficiency\. Although State Space Models \(SSMs\) have recently enabled efficient long\-range modeling, most existing methods models encode each item with a single static contextual role, overlooking the phenomenon ofitem polysemy\. In fact, the same item often plays different semantic roles depending on user context, and existing methods are limited in capture dynamic behavior across different temporal granularities\. In this work, we proposeDSRec, a novel dual\-interest cross SSM model that explicitly disentangles item roles across long\-term and short\-term semantic context\. Sequential items are encoded into long\-term interest embeddings that captures stable preferences via historical aggregation, and a short\-term interest branch that emphasizes local session intent modulated by inter\-click time intervals\. These interest embeddings are processed through distinct SSM encoders: a full\-sequence Mamba for long\-term modeling, and a time\-modulated SSM that dynamically adjusts state evolution based on temporal gaps\. To enable effective cross\-granularity alignment, we adopt residual cross\-fusion mechanism that exchanges contextual information between the two branches while preserving semantic independence\. Experiments on public benchmarks demonstrate that DSRec outperforms other state\-of\-the\-art methods\.

###### Index Terms:

sequential recommendation, long\-term interest, short\-term interest, polysemy, SSM

## IIntroduction

Sequential recommendation has become a fundamental task in modern recommender systems, where the goal is to predict a user’s next interaction based on their historical behavioral sequence\[[1](https://arxiv.org/html/2609.21548#bib.bib9),[2](https://arxiv.org/html/2609.21548#bib.bib1),[3](https://arxiv.org/html/2609.21548#bib.bib42)\]\. With the surge of online activities and the increasing volume of temporally ordered data, it has become crucial to model both*long\-term stable preferences*and*short\-term intent dynamics*to achieve personalized and timely recommendations\[[4](https://arxiv.org/html/2609.21548#bib.bib37),[5](https://arxiv.org/html/2609.21548#bib.bib12),[6](https://arxiv.org/html/2609.21548#bib.bib41)\]\. Transformer\-based models, such as SASRec\[[1](https://arxiv.org/html/2609.21548#bib.bib9)\]and BERT4Rec\[[2](https://arxiv.org/html/2609.21548#bib.bib1)\], have shown promising results in modeling user sequences through attention mechanisms\. However, their high computational complexity and limited scalability to long sequences hinder their practical deployment\[[7](https://arxiv.org/html/2609.21548#bib.bib11),[8](https://arxiv.org/html/2609.21548#bib.bib22)\]\. Recent developments in structured state space models \(SSMs\), including S4\[[9](https://arxiv.org/html/2609.21548#bib.bib7)\], S5\[[10](https://arxiv.org/html/2609.21548#bib.bib28)\], and Mamba\[[7](https://arxiv.org/html/2609.21548#bib.bib11)\], provide efficient alternatives by enabling long\-range dependency modeling with linear complexity\. Consequently, SSMs are gaining increasing interest in recommendation tasks\[[11](https://arxiv.org/html/2609.21548#bib.bib14),[12](https://arxiv.org/html/2609.21548#bib.bib15)\]\.

Despite their advantages, most existing SSM\-based recommender models treat user behavior as a homogeneous sequence and apply a single representation for each item across all contexts\. However, user\-item interactions are inherently multi\-intent: a user may engage with the same item due to different motivations\. In fact,an item can carry polysemous semantics roles\.Take Fig\.[1](https://arxiv.org/html/2609.21548#S1.F1)as an example, the same item \(e\.g\.,Dune\) may occur in different contexts, appearing alongside other science fiction films as part of a user’s long\-term preference, or following a sequence of comics driven by recent short\-term interest\. This illustrates the polysemous nature of item semantics and highlights the necessity of modeling different interest intent to enable accurate personalization\.

Although prior works have explored multi\-interest modeling\[[4](https://arxiv.org/html/2609.21548#bib.bib37),[13](https://arxiv.org/html/2609.21548#bib.bib2),[14](https://arxiv.org/html/2609.21548#bib.bib3)\], they often focus on user\-side intent separation or perform sequence slicing at the cost of disrupting temporal coherence\[[5](https://arxiv.org/html/2609.21548#bib.bib12),[15](https://arxiv.org/html/2609.21548#bib.bib44),[16](https://arxiv.org/html/2609.21548#bib.bib45)\]\. These limitations motivate us to design a model that captures item\-level semantic disambiguation across multiple temporal granularities\.

![Refer to caption](https://arxiv.org/html/2609.21548v1/example_2.png)Fig\. 1:Motivation illustration\. The same item \(”Dune”\) can appear in different viewing sequences—either among consistent sci\-fi films \(long\-term preference\) or after comics \(short\-term interest\)\. This demonstrates the polysemous nature of item semantics, requiring recommendation to distinguish temporal contexts for accurate personalization\.To address these limitations, we proposeDualSSMRecommendation \(DSRec\), a novel dual\-interest sequential recommendation model based on multi\-granular cross residual SSM encoders\. Specifically, we first adopt a dual\-interest modeling that explicitly distinguish the short\-term and long\-term semantic roles of each item\. Each item in the sequence is embedded into two semantic roles: along\-term interestembedding that encodes persistent user preferences through historical aggregation, and ashort\-term interestmodeling that emphasizes local session intent modulated by inter\-click time intervals\. The long\-term and short\-term interest are independently encoded using structurally distinct SSM\-based architectures: one using a standard full\-sequence Mamba\[[11](https://arxiv.org/html/2609.21548#bib.bib14)\], and the other using a time\-modulated SSM that dynamically adjusts state updates based on time gaps\.

Importantly, as using two identical Mamba encoders overlooks the distinct temporal patterns present in short\-term and long\-term behaviors\. Long\-term preferences benefit from full\-sequence modeling and high\-order memory, while short\-term interest depends heavily on session\-local dynamics and time sensitivity\. Hence, we design asymmetric SSM encoders, to optimize each path for its respective granularity\. Rather than merging the two branches via simple concatenation or early fusion, we adopt a residual cross\-fusion mechanism that allows each branch to inject contextual signals from the other without disrupting their specialized semantics\. This cross\-residual design promotes joint modeling of multi\-granular dependencies\.

Extensive experiments on benchmark datasets show that DSRec outperforms state\-of\-the\-art methods in terms of both accuracy and efficiency\. Ablation studies further verify the effectiveness of key modules\. We believe DSRec opens a new path for item role\-aware modeling in sequential recommendation, which has not been explored in existing sequential recommendation systems\.

Our main contributions are as follows:

- •We propose item polysemy of long\-term interest and short\-term interest semantical roles in sequential recommendation\. We adopt polysemous modeling that encodes each item into different interest patterns, enables the model to handle item semantic ambiguity across diverse user histories and session contexts\.
- •We design anasymmetric dual\-SSM architecturecomposed of a global Mamba encoder for modeling long\-term user interest and a time\-modulated SSM for capturing short\-term dynamics\. A residual cross\-fusion mechanism is adopted to enable bidirectional information exchange between the two branches while preserving their semantic independence\.
- •We extensively evaluate our model on multiple public benchmarks, demonstrating that our approach surpasses existing methods in terms of overall accuracy and efficiency\.

## IIRelated Work

### II\-ASequential Recommendation

Sequential recommendation aims to predict the next item a user will interact with based on their historical behavior sequence\[[5](https://arxiv.org/html/2609.21548#bib.bib12),[17](https://arxiv.org/html/2609.21548#bib.bib30),[18](https://arxiv.org/html/2609.21548#bib.bib29)\]\. Early research in sequential recommendation systems heavily relied on RNNs\[[19](https://arxiv.org/html/2609.21548#bib.bib10),[20](https://arxiv.org/html/2609.21548#bib.bib16)\], which were adept at handling sequential data due to their ability to maintain a hidden state across different time steps\. However, RNNs faced limitations in parallel computation and struggled with capturing long\-range dependencies within sequences\. These challenges led to the exploration of Transformer\-based models\[[1](https://arxiv.org/html/2609.21548#bib.bib9),[2](https://arxiv.org/html/2609.21548#bib.bib1),[14](https://arxiv.org/html/2609.21548#bib.bib3),[21](https://arxiv.org/html/2609.21548#bib.bib6),[22](https://arxiv.org/html/2609.21548#bib.bib5),[13](https://arxiv.org/html/2609.21548#bib.bib2),[23](https://arxiv.org/html/2609.21548#bib.bib4),[5](https://arxiv.org/html/2609.21548#bib.bib12)\], which introduced self\-attention mechanisms to overcome the sequential processing constraints\. Several works have attempted to enhance performance via self\-supervised pretraining\[[24](https://arxiv.org/html/2609.21548#bib.bib8)\]or dual representations\[[14](https://arxiv.org/html/2609.21548#bib.bib3),[25](https://arxiv.org/html/2609.21548#bib.bib43)\]\. However, despite the improved performance, the quadratic computational complexity of Transformers with respect to sequence length posed a challenge, especially in real\-time recommendation scenarios where latency is critical\[[26](https://arxiv.org/html/2609.21548#bib.bib31)\]\. The need for more computationally efficient models in sequential recommendation systems has driven the adoption of SSMs\[[7](https://arxiv.org/html/2609.21548#bib.bib11),[27](https://arxiv.org/html/2609.21548#bib.bib27),[12](https://arxiv.org/html/2609.21548#bib.bib15),[28](https://arxiv.org/html/2609.21548#bib.bib34),[29](https://arxiv.org/html/2609.21548#bib.bib35),[30](https://arxiv.org/html/2609.21548#bib.bib36)\]\. These models offer linear computational loads, making them more suitable for handling long sequences and real\-time recommendations\[[31](https://arxiv.org/html/2609.21548#bib.bib32),[32](https://arxiv.org/html/2609.21548#bib.bib33)\]\.

Our work complements these trends by exploring a new modeling dimension: semantic interest roles disentanglement, which explicitly separates short\- and long\-term item semantics and models them via dual SSMs for efficient and more accurate sequence learning\.

### II\-BState Space Models\.

State Space Models \(SSMs\)\[[33](https://arxiv.org/html/2609.21548#bib.bib25),[34](https://arxiv.org/html/2609.21548#bib.bib26)\]have recently gained popularity in sequential modeling due to their ability to capture long\-range dependencies with linear computational complexity\[[7](https://arxiv.org/html/2609.21548#bib.bib11),[27](https://arxiv.org/html/2609.21548#bib.bib27),[35](https://arxiv.org/html/2609.21548#bib.bib17)\]\. In contrast to the quadratic cost of Transformers, SSMs leverage recurrence\-based mechanisms to process sequences efficiently, making them especially attractive for long\-context scenarios\[[7](https://arxiv.org/html/2609.21548#bib.bib11),[8](https://arxiv.org/html/2609.21548#bib.bib22)\]\. Mamba\[[7](https://arxiv.org/html/2609.21548#bib.bib11)\]first introduces a selective state space model with hardware\-aware design, marked a breakthrough in combining high efficiency with strong sequence modeling\. Following this, a variety of models extended Mamba to other domains\. For example, Lieber et al\.\[[8](https://arxiv.org/html/2609.21548#bib.bib22)\]proposed Jamba, a hybrid Transformer\-Mamba large language model that supports efficient long\-context processing through a mixture\-of\-experts architecture\. In the vision domain, Wu and Wan\[[36](https://arxiv.org/html/2609.21548#bib.bib24)\]and Liu et al\.\[[37](https://arxiv.org/html/2609.21548#bib.bib23)\]applied SSMs to image segmentation, demonstrating their superiority in capturing global spatial dependencies over Transformer\-based approaches\. In sequential recommendation domain, Mamba4Rec\[[11](https://arxiv.org/html/2609.21548#bib.bib14)\]introduces the first efficient selective state\-space architecture, provides a robust solution for capturing complex temporal dynamics in user sequential behavior\. To enhance SSMs further, SIGMA\[[1](https://arxiv.org/html/2609.21548#bib.bib9)\]extended Mamba with selective gating and dense extraction layers to better adapt to short sequences and unidirectionaly sequence modeling, and shown promising results\.

However, existing Mamba\-based methods in recommendation mostly treat items as a static semantic role and apply monolithic architectures for modeling sequence dynamics\.

![Refer to caption](https://arxiv.org/html/2609.21548v1/main_6.png)Fig\. 2:Structure overview\. DCRec is a dual\-SSM architecture that explicitly disentangles short\-term and long\-term user preferences via embedding and state\-space separate modeling\. The two streams are connected through residual cross\-links\.

## IIIMethodology

### III\-ANotations Problem Statement

Let𝒰\\mathcal\{U\}denotes a set of users and𝒱\\mathcal\{V\}denotes a set of items, where\|𝒰\|\|\\mathcal\{U\}\|and\|𝒱\|\|\\mathcal\{V\}\|denote the number of users/items,u∈𝒰u\\in\\mathcal\{U\}andv∈𝒱v\\in\\mathcal\{V\}\. The sequential recommendation aims to predict the next item that a useruuis likely to interact with based on their historical sequenceSuS\_\{u\}of interactions\. Therefore, we represent a historical sequence of interactions for the useruuwithSu=\[v1,v2,…,vT\]S\_\{u\}=\\left\[v\_\{1\},v\_\{2\},\\ldots,v\_\{T\}\\right\], wherevtv\_\{t\}represents the item interacted with the useruuat time steptt, and here we omit the subscriptuufor convenience\.vT\+1v\_\{T\+1\}represents the next item that useruuis expected to interact with, which is to be predicted\.

Based on the above notations, given a set of historical interaction sequencesSuS\_\{u\}for users in the training dataset𝒟=\{\(Su,vT\+1\)\}\\mathcal\{D\}=\\left\\\{\\left\(S\_\{u\},v\_\{T\+1\}\\right\)\\right\\\}, the task of sequential recommendation is to learn a modelℱ:Su→vT\+1\\mathcal\{F\}:S\_\{u\}\\rightarrow v\_\{T\+1\}that can accurately predict the next itemvT\+1v\_\{T\+1\}for each useruu\.

### III\-BFramework Overview

InDSRec, we model long\-term interest and short\-term interest jointly\. We adopt dual\-SSM sequential recommendation model based on residual\-coupled state space networks\. Each item in the input sequence is represented by two contextual embeddings: along\-term embeddingthat captures persistent preferences from historical interactions, and ashort\-term embeddingthat emphasizes local intent modulated by inter\-click time gaps\. The two contextual embeddings are processed independently through dedicated SSM encoders: the long\-term embedding is modeled with standard Mamba\[[7](https://arxiv.org/html/2609.21548#bib.bib11)\], while the short\-term embedding is processed using a time\-modulated state\-space model\. To enable effective coordination, a residual cross\-fusion mechanism is introduced between branches\. The fused representations are then projected for next\-item prediction\. The main structure are illustrated in Fig\.[2](https://arxiv.org/html/2609.21548#S2.F2)\.

### III\-CState Space Model

State Space Models \(SSMs\) are a class of dynamical systems widely used for modeling temporal dependencies in continuous or discrete\-time signals\. Recent variant Mamba has shown strong performance in sequence modeling tasks due to their efficient handling of long\-range dependencies\. Formally, an SSM defines the evolution of a hidden stateh⁡\(t\)h\(t\)over time based on an input signalx⁡\(t\)x\(t\), and generates an output signaly⁡\(t\)y\(t\)through the following continuous\-time equations:

dd​t​h​\(t\)\\displaystyle\\frac\{d\}\{dt\}h\(t\)=𝐀​h​\(t\)\+𝐁​x​\(t\)\\displaystyle=\\mathbf\{A\}h\(t\)\+\\mathbf\{B\}x\(t\)\(1\)y⁡\(t\)\\displaystyle y\(t\)=𝐂​h​\(t\)\\displaystyle=\\mathbf\{C\}h\(t\)\(2\)where𝐀∈ℝP×P\\mathbf\{A\}\\in\\mathbb\{R\}^\{P\\times P\},𝐁∈ℝP×D\\mathbf\{B\}\\in\\mathbb\{R\}^\{P\\times D\}, and𝐂∈ℝD×P\\mathbf\{C\}\\in\\mathbb\{R\}^\{D\\times P\}are learnable parameter matrices, andPPdenotes the dimension of the hidden state space\. To apply SSMs in discrete\-time sequential recommendation, we discretize the above system using the Zero\-Order Hold \(ZOH\) method with a constant step sizeΔ​t\\Delta t, resulting in:

𝐡t\\displaystyle\\mathbf\{h\}\_\{t\}=𝐀¯​𝐡t−1\+𝐁¯​𝐱t\\displaystyle=\\bar\{\\mathbf\{A\}\}\\mathbf\{h\}\_\{t\-1\}\+\\bar\{\\mathbf\{B\}\}\\mathbf\{x\}\_\{t\}\(3\)𝐲t\\displaystyle\\mathbf\{y\}\_\{t\}=𝐂𝐡t\\displaystyle=\\mathbf\{C\}\\mathbf\{h\}\_\{t\}\(4\)Here,𝐀¯\\bar\{\\mathbf\{A\}\}and𝐁¯\\bar\{\\mathbf\{B\}\}are the discretized equivalents of𝐀\\mathbf\{A\}and𝐁\\mathbf\{B\}, typically computed via matrix exponentials or approximation methods\. The model sequentially updates the internal state𝐡t\\mathbf\{h\}\_\{t\}given an input embedding𝐱t\\mathbf\{x\}\_\{t\}, and produces the output𝐲t\\mathbf\{y\}\_\{t\}at each time step\. This formulation allows the model to capture structured temporal dynamics in user\-item interaction sequences\. For simplicity, we shorten the whole process to SSM in the following two branches\.

### III\-DLong\-Term Interest Modeling

#### III\-D1Embedding Layer

We first establish a learnable embedding table𝔼=\{E1,E2,⋯,E\|𝒱\|\}∈ℝ\|𝒱\|×D\\mathbb\{E\}=\\left\\\{E\_\{1\},E\_\{2\},\\cdots,E\_\{\|\\mathcal\{V\}\|\}\\right\\\}\\in\\mathbb\{R\}^\{\\mathcal\{\|V\|\}\\times\\mathrm\{D\}\}\. Each itemvj∈𝒱v\_\{j\}\\in\\mathcal\{V\}is initially represented as a one\-hot vector and then projected into a dense vector spaceℝ\|𝒱\|×D\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\\times D\}via an embedding matrix𝔼\\mathbb\{E\}\. Given an interaction sequence𝒮=\{v1,v2,…,vT\}\\mathcal\{S\}=\\\{v\_\{1\},v\_\{2\},\\ldots,v\_\{T\}\\\}, the corresponding embedded input is denoted as𝑿∈ℝT×D\\boldsymbol\{X\}\\in\\mathbb\{R\}^\{T\\times D\}, obtained by retrieving embeddings from𝔼\\mathbb\{E\}\.

User History Aggregation for Long\-Term Interest\.To capture consistent user preferences beyond the current session, we enhance each item’s long\-term interest embedding with a user\-level historical aggregation𝑿l=MLP\(𝑿\|\|Eu\)\\boldsymbol\{X\}\_\{l\}=MLP\(\\boldsymbol\{X\}\|\|E\_\{u\}\)\. The intuition is that long\-term interests are not determined by the current session alone, but by its relative semantics within a user’s broader interaction historyEuE\_\{u\}\. We compute this aggregation by applying a mean pooling over all previous items in the user’s sequence, excluding the current item, and use it as a contextual shift added to the item’s base embedding:

𝐱li=𝐱i\+1i−1​∑j=1i−1𝐱j,\\mathbf\{x\}\_\{l\}^\{i\}=\\mathbf\{x\}\_\{i\}\+\\frac\{1\}\{i\-1\}\\sum\_\{j=1\}^\{i\-1\}\\mathbf\{x\}\_\{j\},\(5\)where𝐞i\\mathbf\{e\}\_\{i\}is the embedding of theii\-th item and the second term captures aggregated historical context\. This design allows the long\-term encoder to better model stable preferences and suppress short\-term fluctuations\.

To enhance the expressiveness and generalization of the item embeddings, we then apply dropout regularization\[[38](https://arxiv.org/html/2609.21548#bib.bib38)\]and layer normalization\[[39](https://arxiv.org/html/2609.21548#bib.bib39)\]on the input embbeding𝑿l\\boldsymbol\{X\}\_\{l\}for long\-term interest branch\. These operations mitigate overfitting and stabilize training dynamics, ensuring a more robust embedding space for downstream modeling\.

#### III\-D2Mamba Block

To capture long\-range user behavior efficiently, we keep and leverage structured SSMs as Mamba4Rec\[[11](https://arxiv.org/html/2609.21548#bib.bib14)\]dose\. The core component is theMamba Block, which transforms an input sequence𝑿l∈ℝB×L×D\\boldsymbol\{X\}\_\{l\}\\in\\mathbb\{R\}^\{B\\times L\\times D\}\(with batch sizeBB, sequence lengthLL, and hidden dimensionDD\) into output features using a selective SSM kernel and residual connections\.

Each Mamba Block transforms an input sequence tensor𝑯i∈ℝB×L×D\\boldsymbol\{H\}\_\{i\}\\in\\mathbb\{R\}^\{B\\times L\\times D\}, whereBBis the batch size,LLis the sequence length, andDDis the hidden dimension\. The transformation follows a structured state space modeling pipeline with selective parameterization\. First, the input undergoes a linear projection and 1D convolution to capture local patterns, followed by a gated activation:

𝑯x=SiLU​\(Conv1D​\(𝑾1​𝑿l\)\),\\boldsymbol\{H\}\_\{x\}=\\text\{SiLU\}\(\\text\{Conv1D\}\(\\boldsymbol\{W\}\_\{1\}\\boldsymbol\{X\}\_\{l\}\)\),\(6\)where𝑾1∈ℝD×D\\boldsymbol\{W\}\_\{1\}\\in\\mathbb\{R\}^\{D\\times D\}is a learnable projection matrix, andSiLU​\(⋅\)\\text\{SiLU\}\(\\cdot\)is the Sigmoid Linear Unit activation defined asSiLU​\(x\)=x⋅σ​\(x\)\\text\{SiLU\}\(x\)=x\\cdot\\sigma\(x\)\. Next, the convolved features𝑯x\\boldsymbol\{H\}\_\{x\}are passed through a parameterized structured state space model \(SSM\), which applies input\-dependent filters using learned state matrices𝑨\\boldsymbol\{A\},𝑩\\boldsymbol\{B\}, and𝑪\\boldsymbol\{C\}:

𝑯y=SSM​\(𝑯x,𝑨,𝑩,𝑪\),\\boldsymbol\{H\}\_\{y\}=\\text\{SSM\}\(\\boldsymbol\{H\}\_\{x\};\\boldsymbol\{A\},\\boldsymbol\{B\},\\boldsymbol\{C\}\),\(7\)whereSSM​\(⋅\)\\text\{SSM\}\(\\cdot\)denotes the discretized recurrence function modeling temporal dynamics\.

To ensure gradient flow and preserve original semantics, a residual connection is introduced with a second projection:

𝑯l=𝑾2​\(SiLU​\(𝑯y\)\)\+𝑿l\.\\boldsymbol\{H\}\_\{l\}=\\boldsymbol\{W\}\_\{2\}\(\\text\{SiLU\}\(\\boldsymbol\{H\}\_\{y\}\)\)\+\\boldsymbol\{X\}\_\{l\}\.\(8\)where𝑾2\\boldsymbol\{W\}\_\{2\}is another trainable linear layer\. This operation fuses the SSM\-refined representation with the original input for stable layer\-wise integration\. The Mamba Block jointly leverages local convolutional processing and global state space modeling, while the use of input\-dependent parameters and residual connections facilitates efficient long\-sequence encoding in recommendation scenarios\.

### III\-EShort\-Term Interest Modeling

#### III\-E1Embedding Layer

To account for the multi\-role nature of items in user behavior sequences, we first assign each item two learned embeddings:𝑿l\\boldsymbol\{X\}\_\{l\}is assigned asLong\-term interest embedding, combined with user history shift as we mentioned\. While in short\-term interest modeling, we first use𝑿s\\boldsymbol\{X\}\_\{s\}representsShort\-term interest embeddingafter a projection from initial sequence embedding𝑿\\boldsymbol\{X\}, to encode the transient intent and session\-specific semantics that may vary across occurrences\. Both representations are learned independently\.

As user short\-term intent in recommendation sequences is often influenced by the timing of recent interactions\. Intuitively, two items clicked within a short time span are more likely to be semantically coherent than those separated by long intervals\. To capture such dynamics, we integratetime interval encodinginto the short\-term interest modeling branch\.

Time Interval Encoding:To incorporate temporal dynamics into sequential modeling, we process the timestamp sequence𝒯=\{t1,t2,…,tT\}\\mathcal\{T\}=\\\{t\_\{1\},t\_\{2\},\\ldots,t\_\{T\}\\\}by computing its temporal difference sequence𝒟=\{d1,d2,…,dT\}\\mathcal\{D\}=\\\{d\_\{1\},d\_\{2\},\\ldots,d\_\{T\}\\\}\[[26](https://arxiv.org/html/2609.21548#bib.bib31)\], where eachdid\_\{i\}represents the elapsed time between inter\-click intervals\. Specifically, the differences are computed as follows:

𝒟i=\{0,i=1𝒯i−𝒯i−1,i=2,…,T\.\\mathcal\{D\}\_\{i\}=\\begin\{cases\}0,&i=1\\\\ \\mathcal\{T\}\_\{i\}\-\\mathcal\{T\}\_\{i\-1\},&i=2,\\ldots,T\\end\{cases\}\.\(9\)The resulting sequence𝒟∈ℝT\\mathcal\{D\}\\in\\mathbb\{R\}^\{T\}captures fine\-grained temporal intervals between actions\. Then we assign each time interval a learnable embedding\. To reduce variance and suppress temporal noise, we further apply dropout and layer normalization to𝒟\\mathcal\{D\}prior to feeding it into downstream temporal modules\. This enables the model to better distinguish between long pauses and rapid interactions, improving its sensitivity to user intent shifts over time\. The raw intervals are then normalized via logarithmic scaling and discretized intoNNbuckets using quantile binning\. Each bucket is mapped to a learnable embedding𝐱it\\mathbf\{x\}\_\{i\}^\{\\text\{t\}\}, capturing temporal sensitivity\. This time encoding is concatenated with the item embedding to form the short\-term interest representation:

𝐱si=MLPS\(xi∥𝐱it\),\\mathbf\{x\}\_\{s\}^\{i\}=\\mathrm\{MLP\}\_\{S\}\\left\(\\mathrm\{x\}\_\{i\}\\,\\\|\\,\\mathbf\{x\}\_\{i\}^\{\\text\{t\}\}\\right\),\(10\)

#### III\-E2Time\-aware SSM

To further enhance temporal sensitivity, we introduce a gating mechanism that modulates the state update based on time interval after similar process as Eq\.[6](https://arxiv.org/html/2609.21548#S3.E6),[7](https://arxiv.org/html/2609.21548#S3.E7), and[8](https://arxiv.org/html/2609.21548#S3.E8):

𝐠i=σ⁡\(𝑾t⋅𝐱it\),\\mathbf\{g\}\_\{i\}=\\sigma\(\\boldsymbol\{W\}\_\{t\}\\cdot\\mathbf\{x\}\_\{i\}^\{\\text\{t\}\}\),\(11\)𝐡si=𝐠i⊙SSM​\(𝐱si\)\+\(1−𝐠i\)⊙𝐡si−1,\\mathbf\{h\}^\{i\}\_\{s\}=\\mathbf\{g\}\_\{i\}\\odot\\text\{SSM\}\(\\mathbf\{x\}\_\{s\}^\{i\}\)\+\(1\-\\mathbf\{g\}\_\{i\}\)\\odot\\mathbf\{h\}^\{i\-1\}\_\{s\},\(12\)whereσ⁡\(⋅\)\\sigma\(\\cdot\)is a sigmoid activation\. The gate vectorgig\_\{i\}controls how much the current input should influence the hidden state\. When the time gap is small,gig\_\{i\}tends to be higher, allowing new information to be integrated; for large gaps,gig\_\{i\}is smaller, retaining past state memory, allowing dynamic temporal sensitivity\. This time\-aware control mechanism is critical for modeling transient interests that shift rapidly within user sessions\.

### III\-FResidual Cross\-SSM Fusion

We construct two independent SSM encoders:Mamba Blockprocesses the𝑿l\\boldsymbol\{X\}\_\{l\}sequence to extract long\-term user preference;Time\-aware SSMprocesses the𝑿s\\boldsymbol\{X\}\_\{s\}sequence to model session\-level short\-term dynamics\. Each encoder is augmented with a residual cross\-connection, meaning each branch receives the detached output of the other as auxiliary input\. This enables information sharing while preventing entanglement of gradients\.

### III\-GFeed\-Forward Network

The standard feed\-forward network is adopted to enhance the modeling of user interactions in both long\- and short\-term interest latent space, which is defined as follows:

FFN​\(Hl\)=GELU​\(Hl​W\(1\)\+b\(1\)\)​W\(2\)\+b\(2\),\\text\{FFN\}\(H\_\{l\}\)=\\text\{GELU\}\\left\(H\_\{l\}W^\{\(1\)\}\+b^\{\(1\)\}\\right\)W^\{\(2\)\}\+b^\{\(2\)\},\(13\)FFN​\(Hs\)=GELU​\(Hs​W\(3\)\+b\(3\)\)​W\(4\)\+b\(4\),\\text\{FFN\}\(H\_\{s\}\)=\\text\{GELU\}\\left\(H\_\{s\}W^\{\(3\)\}\+b^\{\(3\)\}\\right\)W^\{\(4\)\}\+b^\{\(4\)\},\(14\)whereW\(1\),W\(3\)∈ℝD×4​DW^\{\(1\)\},W^\{\(3\)\}\\in\\mathbb\{R\}^\{D\\times 4D\},W\(2\),W\(4\)∈ℝ4​D×DW^\{\(2\)\},W^\{\(4\)\}\\in\\mathbb\{R\}^\{4D\\times D\},b\(1\),b\(3\)∈ℝ4​Db^\{\(1\)\},b^\{\(3\)\}\\in\\mathbb\{R\}^\{4D\}, andb\(2\),b\(4\)∈ℝDb^\{\(2\)\},b^\{\(4\)\}\\in\\mathbb\{R\}^\{D\}are the parameters of two dense layers, and we utilize the GELU activation function\[[40](https://arxiv.org/html/2609.21548#bib.bib40)\]\.

The FFN is designed to capture complex patterns and interactions within sequential data by applying two non\-linear transformations using dense layers with the activation function\. To enhance model robustness and prevent overfitting, we incorporate dropout and layer normalization after each Mamba block/Time\-aware SSM block and feed\-forward network, as depicted in Eq\.[13](https://arxiv.org/html/2609.21548#S3.E13)\. This approach aids in regularizing the model and accelerating training convergence\.

### III\-HStacking SSM Layers

We investigate the use of stacked SSM layers \(BB\) to enhance sequential recommendation\. Despite deeper networks not always improving performance, residual connections are vital for propagating features across layers\. We explore configurations for stacked SSM layers to balance effectiveness and efficiency\. Further experiments are conducted to evaluate the trade\-offs between effectiveness and efficiency of stacked layers\.

### III\-IPrediction Layer and Optimization

The final output embedding is formed by concatenating the last positions of long\-term interest modeling outputhth\_\{t\}and long\-term interest modeling outputsts\_\{t\}:

𝐨T=MLP\(\[𝐡lT∥𝐡sT\]\)\.\\mathbf\{o\}\_\{T\}=\\mathrm\{MLP\}\(\[\\mathbf\{h\}\_\{l\}^\{T\}\\,\\\|\\,\\mathbf\{h\}\_\{s\}^\{T\}\]\)\.\(15\)Prediction is computed via dot product over the item embedding matrix:

y^=Softmax⁡\(𝐨T⋅E⊤\)∈ℝ\|𝒱\|,\\hat\{y\}=\\mathrm\{Softmax\}\(\\mathbf\{o\}\_\{T\}\\cdot\\mathrm\{E\}^\{\\top\}\)\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\},\(16\)wherey^\\hat\{y\}is the probability distribution over the next item in the item set𝒱\\mathcal\{V\}\. To optimize the DSRec, we adopt the cross entropy loss function as below:

ℒ=−∑u∈𝒰∑i=1NyT\+1log\(y^T\+1\)\\mathcal\{L\}=\-\\sum\_\{u\\in\\mathcal\{U\}\}\\sum\_\{i=1\}^\{N\}y\_\{T\+1\}\\log\\left\(\\hat\{y\}\_\{T\+1\}\\right\)\(17\)whereyiy\_\{i\}represents the ground truth of itemvT\+1v\_\{T\+1\}\.

## IVExperiments

TABLE I:Dataset statistics\.In this section, we first give brief introduction of the experimental setting, and to evaluate the effectiveness of our proposed DSRec, we conducted extensive experiments on two benchmark datasets, giving answers to the following research questions \(RQs\):

- •RQ1:Can the proposed DSRec outperform previous sequential recommendation methods?
- •RQ2:Can the key modules contribute to the overall prediction accuracy?
- •RQ3:How do specific settings in the method affect overall performance?
- •RQ4:How does our model benefit sequential recommendations?

### IV\-AExperimental Settings

#### IV\-A1Data

We conduct our experiments on three public datasets, and all the datasets are collected from various real\-world application platforms, and these datasets show considerable variation in both sequence length and sparsity\.

BeautyandVideo\-Games\[[41](https://arxiv.org/html/2609.21548#bib.bib19)\]are obtained from Amazon e\-commerce platform, and the two dataset typically contains information related to beauty and movie products\.

MovieLens\-1M\[[42](https://arxiv.org/html/2609.21548#bib.bib18)\]is a benchmark movie recommendation dataset which contains bout one million movie ratings from users\.

The important statistics for the three datasets can be found in Table[I](https://arxiv.org/html/2609.21548#S4.T1)\.

#### IV\-A2Baselines

A variety of representative methods in sequential recommendation have been chosen for comparison against our approach\. They are organized into the following distinct categories:

##### CNN based method

- •Caser\[[43](https://arxiv.org/html/2609.21548#bib.bib13)\]: It utilizes a convolutional layer to extract features from user interaction sequences, capturing both short\-term and long\-term patterns\.

##### RNN based method

- •GRU4Rec\[[19](https://arxiv.org/html/2609.21548#bib.bib10)\]: It refines conventional RNNs, incorporating a ranking loss function to achieve superior performance tailored to sequential recommendation domain\.

##### Attention based methods

- •NARM\[[20](https://arxiv.org/html/2609.21548#bib.bib16)\]: This model employs an attention\-based blended encoder to distill user intent from a sequence of actions, forming a unified session representation\.
- •SASRec\[[1](https://arxiv.org/html/2609.21548#bib.bib9)\]: It is a self\-attention\-based sequential model that harmonizes long\-term semantic comprehension with a focus on a select set of pertinent actions\. It dynamically selects pertinent past items at each step to forecast the subsequent item\.
- •BERT4Rec\[[2](https://arxiv.org/html/2609.21548#bib.bib1)\]: It further applies the Transformer model to capture long\-term dependencies in sequences through a bidirectional transformer architecture, thus improving the accuracy and effectiveness\.

##### SSM based methods

- •Mamba4Rec\[[11](https://arxiv.org/html/2609.21548#bib.bib14)\]:It is the first to apply selective State Space Models \(SSMs\) in sequential recommendations, particularly the hardware\-optimized Mamba block\.Thismethod has attracted wide attention and applications\.
- •SIGMA\[[12](https://arxiv.org/html/2609.21548#bib.bib15)\]: This model is a hybrid model that integrates Mamba and GRU\. It incorporates a Partially Flipped Mamba component combined with a Dense Selective Gate\.

#### IV\-A3Implementation

Our evaluation is implemented based on PyTorch111[http://pytorch\.org](http://pytorch.org/)and RecBole222[https://github\.com/RUCAIBox/RecBole](https://github.com/RUCAIBox/RecBole)\. All models utilize a model dimension of 64, a standard selection for sequential recommendation models that focus solely on ID\-based interactions\. The optimizer is Adam\[[44](https://arxiv.org/html/2609.21548#bib.bib20)\]with a learning rate of 0\.001, and the training batch size is set to 2048\. For Mamba\-based models, we use structured state space model configurations with a state dimension was consistently set to 32, a local convolution width set to 4, the kernel size for 1D causal convolution was 4, and a block expansion factor was set to 2\. We apply a dropout rate of 0\.2 after each layer to prevent overfitting\. All experiments are conducted on NVIDIA GeForce RTX 4090 GPUs\. We employ a training batch size of 2048 and a validation batch size of 4096 across all experiments\. The maximum sequence length is dynamically determined based on each dataset’s characteristics: 200 for MovieLens\-1M and 50 for both Amazon\-Beauty and Amazon\-Video\-Games\. All models are carefully tuned to achieve their best potential performance under unified experimental setup\.

#### IV\-A4Evaluation

Followed by previous studies\[[2](https://arxiv.org/html/2609.21548#bib.bib1),[1](https://arxiv.org/html/2609.21548#bib.bib9),[45](https://arxiv.org/html/2609.21548#bib.bib21),[11](https://arxiv.org/html/2609.21548#bib.bib14)\], we employ the leave\-one\-out strategy to assess the effectiveness of each method\. Under this strategy, for each user, the last interacted item in temporal order is considered as the test data and the previous one as validation data\. To ensure robust evaluation metrics, we utilize Hit Ratio \(HR@\), Normalized Discounted Cumulative Gain \(NDCG@\), and Mean Reciprocal Rank \(MRR@\) to evaluate the performance of each method, as these metrics are widely accepted in the field of recommendation\. In this study, we report the evaluation metrics at rank 10\.

### IV\-BOverall Performance

TABLE II:Overall Performance Comparison of Different Methods\.RedandBolddata indicate the best results, andblueunderlinedones are the second best results\. ”⋆\\star” indicates the improvements are statistically significantTable[II](https://arxiv.org/html/2609.21548#S4.T2)summarizes the overall performance of our proposedDSRecmodel compared to a wide range of state\-of\-the\-art sequential recommendation baselines\. Best results are highlighted inbold red, while the second\-best results areunderlined in blue\. From the table, we can make the following observations:

- •Our DSRec achieves state\-of\-the\-art performance across all datasets and most evaluation metrics\. OnMovieLens\-1M, DSRec outperforms the best\-performing baseline \(SIGMA\[[12](https://arxiv.org/html/2609.21548#bib.bib15)\]\) by\+2\.58%\+2\.58\\%in HR@10, while achieving a notable\+5\.67%\+5\.67\\%gain in MRR@10 over Mamba4Rec\. These improvements demonstrate the advantage of our dual\-interest embedding design and residual cross\-SSM in capturing semantic signals in dense user sequences\. On the more sparse datasetsAmazon\-BeautyandAmazon\-Video\-Games, DSRec continues to outperform other models, validating its effectiveness in data sparsity scenario\.
- •Our experimental results indicate that Transformer\-based models outperform RNNs and CNNs based methods, in sequential recommendation tasks\. This enhanced performance can be attributed to the Transformer’s ability to process sequences in parallel and capture complex dependencies over long distances, which is particularly advantageous for modeling user\-item interactions\.
- •SSM\-based models \(Mamba4Rec\[[11](https://arxiv.org/html/2609.21548#bib.bib14)\], SIGMA\[[12](https://arxiv.org/html/2609.21548#bib.bib15)\]\) demonstrate further improvements over Transformers, especially on datasets like MovieLens\-1M with longer sequence lengths\. This aligns with recent findings\[[7](https://arxiv.org/html/2609.21548#bib.bib11),[8](https://arxiv.org/html/2609.21548#bib.bib22)\]that SSMs can efficiently capture long\-range dependencies while maintaining linear scalability\. The performance gap between SSM and Transformer models is more significant in longer sequence than in short sessions, validating the architectural design of structured temporal modeling\.

### IV\-CAblation Study

TABLE III:Ablation study on Model Variant\.Bolddata indicate the best results\.To validate the contributions of each component in DSRec, we conduct comprehensive ablation studies by systematically removing \(w/o\) or modifying key modules\. The results are presented in Table[III](https://arxiv.org/html/2609.21548#S4.T3), evaluated on three datasets and reported by HR@10 and NDCG@10\.

##### Impact of Residual Cross\-Fusion

From the comparison results, we can observe that removing the residual cross\-fusion module significantly degrades performance on all datasets, especially on MovieLens\-1M and Amazon\-Beauty\. This demonstrates the importance of inter\-branch communication, where each interest embedding branch \(long\-term and short\-term\) benefits from contextual signals in the other\.

##### Impact of Dual interest embedding Representation

Replacing the dual\-interest embedding input with a single\-interest embedding, as other baselines do, results in notable performance degradation across all benchmarks\. This confirms our hypothesis that items assume different semantic roles in short\-term and long\-term contexts, and a unified interest embedding is insufficient to capture such polysemy behavior\.

##### Impact of SSM\-S Branch

Removing the short\-term SSM\-S encoder while retaining the long\-term SSM\-L leads to a notable performance drop, especially on the MovieLens dataset\. This suggests that the short\-term intent modeling is vital, particularly in sequential data where recent actions carry strong semantic signals\.

##### Impact of Fusion Encoder

Without the fusion encoder, the model retains the dual\-interest embedding structure but lacks effective integration before output\. While the impact is slight, the drop in NDCG@10 on Amazon datasets indicates its importance in combining interest\-specific outputs for fine\-grained ranking\.

##### Dual Mamba Baseline

We also construct a variant with two identical Mamba modules \(Dual Mamba\) without semantic\-specific design\. While this variant performs competitively on MovieLens\-1M, it underperforms significantly on Amazon\-Beauty and Amazon\-Video\-Games\. This observation confirms that our asymmetric dual\-SSM design is more effective than merely duplicating Mamba blocks\.

These ablation observations validate the necessity of each proposed component\. The dual\-interest embedding, heterogeneous SSM encoders, and residual cross\-fusion are all crucial for capturing multi\-granular temporal dynamics and achieving state\-of\-the\-art performance\.

### IV\-DHyperparameter Analysis

In this section, we conduct experiments on Beauty to analyze the influence of two significant hyperparameters:a\)L, the maximum sequential length;b\)B, the number of stacked DSRec blocks\. The results are respectively visualized in Fig\.[3](https://arxiv.org/html/2609.21548#S4.F3)and Table[IV](https://arxiv.org/html/2609.21548#S4.T4)\.

![Refer to caption](https://arxiv.org/html/2609.21548v1/Figure/hyper_L.png)Fig\. 3:Max Sequence LengthTTstudy on Amazon\-Beauty and Amazon\-Video\-Games\.TABLE IV:Parameter study forBBon Amazon\-Beauty and Amazon\-Video\-Games\.##### Effect of the Number of BlocksBB

From Table[IV](https://arxiv.org/html/2609.21548#S4.T4), we observe that on Amazon\-Beauty and Amazon\-Video\-Games, the model achieves the highest HR@10 and NDCG@10 whenB=3B=3andB=2B=2, respectively\. This suggests that stacking multiple residual\-coupled dual\-SSM layers improves to capture multi\-granular temporal dynamics\. However, further increasingBBbeyond 3 leads to degradation in performance\. This is likely due to overfitting or optimization difficulties caused by deeper architectures, especially on sparse datasets\. Moreover, we observe a steady increase in training time and GPU memory consumption with more layers, highlighting a trade\-off between performance and efficiency\.

##### Effect of Maximum Sequence LengthTT

Fig\.[3](https://arxiv.org/html/2609.21548#S4.F3)shows the performance varying the maximum sequence length\. On both Amazon\-Beauty and Amazon\-Video\-Games, we observe a peak atT=50T=50, after which performance gradually drops\. This suggests that for sparse and short\-session datasets, incorporating overly long histories may introduce noise or dilute recent preference signals\. Furthermore, longer sequences also lead to substantial computational overhead, as training time nearly quadruples fromT=50T=50toT=200T=200\. These results highlight the importance of tuning both structural depth and temporal context window to match dataset characteristics\. DSRec demonstrates strong adaptability across a wide range of hyperparameter settings\.

### IV\-EModel Complexity and Efficiency

We analyze the theoretical and empirical complexity of DSRec compared to Transformer\- and SSM\-based baselines\. In addition, we also compare prediction and training efficiency with the most competitive baseline\.

##### Theoretical Complexity

As shown in Table[V](https://arxiv.org/html/2609.21548#S4.T5), the Transformer\-based models incur quadratic time and space cost with respect to sequence lengthTT, making them less suitable for long sequences\. RNNs and SSMs maintain linear scaling inTT\. While both RNNs and SSMs require𝒪⁡\(d2\)\\mathcal\{O\}\(d^\{2\}\)operations due to dense linear transitions, SSMs can still outperform RNNs in parallelizability and hardware efficiency\. DSRec maintains linear complexity in sequence length but doubles the cost in hidden size due to its dual\-branch architecture, which remains attractive and scalable in practice\.

TABLE V:Theoretical time complexity comparison\.TT: sequence length,dd: dimension of hidden layers\.![Refer to caption](https://arxiv.org/html/2609.21548v1/Figure/loss_2.png)Fig\. 4:Prediction comparison and training loss comparison on MovieLens\-1M and Amazon\-Video\-Games datasets\.
##### Prediction and Training Efficiency

Fig\.[4](https://arxiv.org/html/2609.21548#S4.F4)presents a comparative analysis of model performance, focusing on HR@10 and training loss across epochs for two models: Mamba4Rec and DSRec, evaluated on the MovieLens\-1M and Amazon\-Video\-Games datasets\. The comparative analysis shows that DSRec generally outperforms Mamba4Rec in terms of HR@10, and faster initial convergence in training loss\. This suggests that DSRec is better suited for capturing complex patterns in sequential recommendation tasks, particularly on long sequential dataset\.

##### Runtime Efficiency

We further compare wall\-clock inference time and GPU memory cost at Table[VI](https://arxiv.org/html/2609.21548#S4.T6)\. The experimental results highlight SSM’s superior runtime efficiency, with both faster inference times and lower GPU memory\. Although DSRec introduces two SSM encoders, both are linear\-time and lightweight\. Thus, the overall inference costs remain close to SSM baselines and are significantly more efficient than Transformer\-based method\.

TABLE VI:Inference time and GPU Memory cost comparisoneachepoch \(Amazon\-Beauty, batch=2048\)\.

### IV\-FRobustness Analysis

Recommender systems often suffer from the cold\-start and long\-tail problems, where items with few interactions are poorly represented during training\. To assess the robustness of DSRec across users with varying historical interaction lengths, we partition the test users into three groups based on the number of their training interactions:0–5,5–20, and20\+\. Each subset corresponds to users with sparse, moderate, and rich historical behavior, respectively\. We compare our model DSRec against the strong SSM\-based baseline and report metrics including Hit@10/20, NDCG@10/20, and MRR@10/20 for each group\.

![Refer to caption](https://arxiv.org/html/2609.21548v1/Figure/cold.png)Fig\. 5:Robustness comparison across user history length groups in Amazon\-Video\-Games\.From Fig\.[5](https://arxiv.org/html/2609.21548#S4.F5), we can observe that forcold\-startusers with short histories \(0–5 interactions\), DSRec significantly outperforms Mamba4Rec, which indicating superior modeling of early\-stage preferences\. For mid\-range users \(5–20\), DSRec continues to outperform in most metrics, suggesting its ability to balance long\- and short\-term signals\. For users with rich histories \(20\+\), DSRec achieves the highest performance, confirming the benefit of its dual\-interest encoding for long\-horizon modeling\. These results demonstrate the robustness of DSRec under varying interaction sparsity\. The consistent improvements across all user groups validate the effectiveness of our DSRec\.

## VConclusion and Future Work

Most existing sequential recommendation models implicitly assuming that each item has consistent semantics across all user contexts\. This limits their ability to capture the polysemous nature of item behavior\. Moreover, standard state space architectures failing to account for the distinct dynamics of long\-term preferences versus short\-term session\-based intent\. To address these challenges, in this paper, we proposeDSRec, a dual\-interest embedding state\-space model for sequential recommendation that explicitly models both long\-term and short\-term user preferences\. Our method introduces a semantic disentanglement mechanism where each item is projected into a long\-term interest embedding and a short\-term interest embedding\. These embeddings are encoded separately by two specialized SSM encoders: a full\-sequence Mamba and a time\-modulated sensitive SSM to inter\-click temporal gaps\. To enable cross\-granular coordination without semantic interference, we design a residual cross\-fusion mechanism that injects context between branches\. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of DSRec\.

Future Work\.While DSRec achieves advanced performance, several limitations remain open for future research\. First, the item polysemous modeling is novel but basic, we plan to extend by context aware dual\-tokenization to incorporate item\-side semantic roles, such as Variational AutoEncoder \(VAE\)\. Second, sparse context information in transactions need the integration of external knowledge \(such as graph structure\) into the interest embedding, which may further enhance generalization\. Finally, exploring a learnable fusion gate or contrastive alignment between the two semantic branches could yield more adaptive interactions beyond residual fusion\. We also intend to test DSRec in cross\-domain settings to assess its broader applications\.

## Acknowledgments

## References

- \[1\]\(2018\)Self\-attentive sequential recommendation\.In2018 IEEE international conference on data mining \(ICDM\),pp\. 197–206\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1),[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1),[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1),[2nd item](https://arxiv.org/html/2609.21548#S4.I4.i2.p1.1),[§IV\-A4](https://arxiv.org/html/2609.21548#S4.SS1.SSS4.p1.1),[TABLE II](https://arxiv.org/html/2609.21548#S4.T2.10.6.1)\.
- \[2\]F\. Sun, J\. Liu, J\. Wu, C\. Pei, X\. Lin, W\. Ou, and P\. Jiang\(2019\)BERT4Rec: sequential recommendation with bidirectional encoder representations from transformer\.InProceedings of the 28th ACM international conference on information and knowledge management,pp\. 1441–1450\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1),[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1),[3rd item](https://arxiv.org/html/2609.21548#S4.I4.i3.p1.1),[§IV\-A4](https://arxiv.org/html/2609.21548#S4.SS1.SSS4.p1.1),[TABLE II](https://arxiv.org/html/2609.21548#S4.T2.10.7.1)\.
- \[3\]S\. Liao, P\. Mok, and L\. Li\(2026\)Consistency regularization for complementary clothing recommendations\.Applied Soft Computing,pp\. 115069\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1)\.
- \[4\]Y\. Yang, C\. Huang, L\. Xia, Y\. Liang, Y\. Yu, and C\. Li\(2022\)Multi\-behavior hypergraph\-enhanced transformer for sequential recommendation\.InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining,pp\. 2263–2274\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1),[§I](https://arxiv.org/html/2609.21548#S1.p3.1)\.
- \[5\]S\. Liao and P\. Mok\(2024\)Hypergraph\-enhanced contrastively regularized transformer for multi\-behavior e\-commerce product recommendation\.In2024 IEEE International Conference on Data Mining \(ICDM\),pp\. 767–772\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1),[§I](https://arxiv.org/html/2609.21548#S1.p3.1),[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[6\]S\. Liao, Y\. Ding, and P\. Mok\(2023\)Recommendation of mix\-and\-match clothing by modeling indirect personal compatibility\.InProceedings of the 2023 ACM International Conference on Multimedia Retrieval,pp\. 560–564\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1)\.
- \[7\]A\. Gu and T\. Dao\(2023\)Mamba: linear\-time sequence modeling with selective state spaces\.arXiv preprint arXiv:2312\.00752\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1),[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1),[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1),[§III\-B](https://arxiv.org/html/2609.21548#S3.SS2.p1.1),[3rd item](https://arxiv.org/html/2609.21548#S4.I6.i3.p1.1)\.
- \[8\]O\. Lieber, B\. Lenz, H\. Bata, G\. Cohen, J\. Osin, I\. Dalmedigos, E\. Safahi, S\. Meirom, Y\. Belinkov, S\. Shalev\-Shwartz,et al\.\(2024\)Jamba: a hybrid transformer\-mamba language model\.arXiv preprint arXiv:2403\.19887\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1),[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1),[3rd item](https://arxiv.org/html/2609.21548#S4.I6.i3.p1.1)\.
- \[9\]A\. Gu, K\. Goel, and C\. Ré\(2021\)Efficiently modeling long sequences with structured state spaces\.arXiv preprint arXiv:2111\.00396\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1)\.
- \[10\]J\. T\. Smith, A\. Warrington, and S\. W\. Linderman\(2022\)Simplified state space layers for sequence modeling\.arXiv preprint arXiv:2208\.04933\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1)\.
- \[11\]C\. Liu, J\. Lin, J\. Wang, H\. Liu, and J\. Caverlee\(2024\)Mamba4rec: towards efficient sequential recommendation with selective state space models\.arXiv preprint arXiv:2403\.03900\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1),[§I](https://arxiv.org/html/2609.21548#S1.p4.1),[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1),[§III\-D2](https://arxiv.org/html/2609.21548#S3.SS4.SSS2.p1.1),[1st item](https://arxiv.org/html/2609.21548#S4.I5.i1.p1.1),[3rd item](https://arxiv.org/html/2609.21548#S4.I6.i3.p1.1),[§IV\-A4](https://arxiv.org/html/2609.21548#S4.SS1.SSS4.p1.1),[TABLE II](https://arxiv.org/html/2609.21548#S4.T2.10.8.1)\.
- \[12\]Z\. Liu, Q\. Liu, Y\. Wang, W\. Wang, P\. Jia, M\. Wang, Z\. Liu, Y\. Chang, and X\. Zhao\(2025\)SIGMA: selective gated mamba for sequential recommendation\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 12264–12272\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p1.1),[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1),[2nd item](https://arxiv.org/html/2609.21548#S4.I5.i2.p1.1),[1st item](https://arxiv.org/html/2609.21548#S4.I6.i1.p1.1),[3rd item](https://arxiv.org/html/2609.21548#S4.I6.i3.p1.1),[TABLE II](https://arxiv.org/html/2609.21548#S4.T2.10.9.1)\.
- \[13\]X\. Xie, F\. Sun, Z\. Liu, S\. Wu, J\. Gao, J\. Zhang, B\. Ding, and B\. Cui\(2022\)Contrastive learning for sequential recommendation\.In2022 IEEE 38th international conference on data engineering \(ICDE\),pp\. 1259–1273\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p3.1),[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[14\]R\. Qiu, Z\. Huang, H\. Yin, and Z\. Wang\(2022\)Contrastive learning for representation degeneration problem in sequential recommendation\.InProceedings of the fifteenth ACM international conference on web search and data mining,pp\. 813–823\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p3.1),[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[15\]S\. Liao and P\. Mok\(2026\)Hamiltonian spectral\-temporal dissipative dynamics for sequential recommendation\.arXiv preprint arXiv:2608\.25755\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p3.1)\.
- \[16\]S\. Liao and P\. Mok\(2026\)PCGNet: unifying shared and specific information for fashion matching recommendations\.arXiv preprint arXiv:2609\.13339\.Cited by:[§I](https://arxiv.org/html/2609.21548#S1.p3.1)\.
- \[17\]Y\. Ding, Y\. Ma, W\. K\. Wong, and T\. Chua\(2021\)Leveraging two types of global graph for sequential fashion recommendation\.InProceedings of the 2021 international conference on multimedia retrieval,pp\. 73–81\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[18\]Y\. Ding, Y\. Ma, W\. K\. Wong, and T\. Chua\(2021\)Modeling instant user intent and content\-level transition for sequential fashion recommendation\.IEEE transactions on multimedia24,pp\. 2687–2700\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[19\]B\. Hidasi, A\. Karatzoglou, L\. Baltrunas, and D\. Tikk\(2015\)Session\-based recommendations with recurrent neural networks\.arXiv preprint arXiv:1511\.06939\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1),[1st item](https://arxiv.org/html/2609.21548#S4.I3.i1.p1.1),[TABLE II](https://arxiv.org/html/2609.21548#S4.T2.10.4.1)\.
- \[20\]J\. Li, P\. Ren, Z\. Chen, Z\. Ren, T\. Lian, and J\. Ma\(2017\)Neural attentive session\-based recommendation\.InProceedings of the 2017 ACM on Conference on Information and Knowledge Management,pp\. 1419–1428\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1),[1st item](https://arxiv.org/html/2609.21548#S4.I4.i1.p1.1),[TABLE II](https://arxiv.org/html/2609.21548#S4.T2.10.5.1)\.
- \[21\]Y\. Ye, L\. Xia, and C\. Huang\(2023\)Graph masked autoencoder for sequential recommendation\.InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 321–330\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[22\]Y\. Yang, C\. Huang, L\. Xia, C\. Huang, D\. Luo, and K\. Lin\(2023\)Debiased contrastive learning for sequential recommendation\.InProceedings of the ACM web conference 2023,pp\. 1063–1073\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[23\]Y\. Chen, Z\. Liu, J\. Li, J\. McAuley, and C\. Xiong\(2022\)Intent contrastive learning for sequential recommendation\.InProceedings of the ACM Web Conference 2022,pp\. 2172–2182\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[24\]K\. Zhou, H\. Wang, W\. X\. Zhao, Y\. Zhu, S\. Wang, F\. Zhang, Z\. Wang, and J\. Wen\(2020\)S3\-rec: self\-supervised learning for sequential recommendation with mutual information maximization\.InProceedings of the 29th ACM international conference on information & knowledge management,pp\. 1893–1902\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[25\]S\. Liao, Y\. Ding, P\. Mok, Q\. Huang, and J\. Cao\(2024\)Reproducibility companion paper: recommendation of mix\-and\-match clothing by modeling indirect personal compatibility\.InProceedings of the 2024 International Conference on Multimedia Retrieval,pp\. 1224–1227\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[26\]H\. Fan, M\. Zhu, Y\. Hu, H\. Feng, Z\. He, H\. Liu, and Q\. Liu\(2024\)TiM4Rec: an efficient sequential recommendation model based on time\-aware structured state space duality model\.arXiv preprint arXiv:2409\.16182\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1),[§III\-E1](https://arxiv.org/html/2609.21548#S3.SS5.SSS1.p3.1)\.
- \[27\]T\. Huang, X\. Pei, S\. You, F\. Wang, C\. Qian, and C\. Xu\(2025\)Localmamba: visual state space model with windowed selective scan\.InEuropean Conference on Computer Vision,pp\. 12–22\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1),[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1)\.
- \[28\]S\. Zhang, R\. Zhang, and Z\. Yang\(2024\)Matrrec: uniting mamba and transformer for sequential recommendation\.arXiv preprint arXiv:2407\.19239\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[29\]Z\. Zhang, B\. Yang, and Y\. Lu\(2025\)A local context enhanced consistency\-aware mamba\-based sequential recommendation model\.Information Processing & Management62\(3\),pp\. 104076\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[30\]Z\. Liu, Q\. Liu, Y\. Wang, W\. Wang, P\. Jia, M\. Wang, Z\. Liu, Y\. Chang, and X\. Zhao\(2024\)Bidirectional gated mamba for sequential recommendation\.arXiv preprint arXiv:2408\.11451\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[31\]A\. Gu, T\. Dao, S\. Ermon, A\. Rudra, and C\. Ré\(2020\)Hippo: recurrent memory with optimal polynomial projections\.Advances in neural information processing systems33,pp\. 1474–1487\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[32\]B\. Peng, E\. Alcaide, Q\. Anthony, A\. Albalak, S\. Arcadinho, S\. Biderman, H\. Cao, X\. Cheng, M\. Chung, M\. Grella,et al\.\(2023\)Rwkv: reinventing rnns for the transformer era\.arXiv preprint arXiv:2305\.13048\.Cited by:[§II\-A](https://arxiv.org/html/2609.21548#S2.SS1.p1.1)\.
- \[33\]J\. D\. Hamilton\(1994\)State\-space models\.Handbook of econometrics4,pp\. 3039–3080\.Cited by:[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1)\.
- \[34\]M\. Aoki\(2013\)State space modeling of time series\.Springer Science & Business Media\.Cited by:[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1)\.
- \[35\]H\. Qu, Y\. Zhang, L\. Ning, W\. Fan, and Q\. Li\(2024\)Ssd4rec: a structured state space duality model for efficient sequential recommendation\.arXiv preprint arXiv:2409\.01192\.Cited by:[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1)\.
- \[36\]J\. Wu, Z\. Wang, M\. Hong, W\. Ji, H\. Fu, Y\. Xu, M\. Xu, and Y\. Jin\(2025\)Medical sam adapter: adapting segment anything model for medical image segmentation\.Medical image analysis102,pp\. 103547\.Cited by:[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1)\.
- \[37\]Y\. Liu, Y\. Tian, Y\. Zhao, H\. Yu, L\. Xie, Y\. Wang, Q\. Ye, J\. Jiao, and Y\. Liu\(2024\)Vmamba: visual state space model\.Advances in neural information processing systems37,pp\. 103031–103063\.Cited by:[§II\-B](https://arxiv.org/html/2609.21548#S2.SS2.p1.1)\.
- \[38\]N\. Srivastava, G\. Hinton, A\. Krizhevsky, I\. Sutskever, and R\. Salakhutdinov\(2014\)Dropout: a simple way to prevent neural networks from overfitting\.The journal of machine learning research15\(1\),pp\. 1929–1958\.Cited by:[§III\-D1](https://arxiv.org/html/2609.21548#S3.SS4.SSS1.p3.1)\.
- \[39\]J\. Xu, X\. Sun, Z\. Zhang, G\. Zhao, and J\. Lin\(2019\)Understanding and improving layer normalization\.Advances in neural information processing systems32\.Cited by:[§III\-D1](https://arxiv.org/html/2609.21548#S3.SS4.SSS1.p3.1)\.
- \[40\]D\. Hendrycks and K\. Gimpel\(2016\)Gaussian error linear units \(gelus\)\.arXiv preprint arXiv:1606\.08415\.Cited by:[§III\-G](https://arxiv.org/html/2609.21548#S3.SS7.p2.1)\.
- \[41\]J\. McAuley, C\. Targett, Q\. Shi, and A\. Van Den Hengel\(2015\)Image\-based recommendations on styles and substitutes\.InProceedings of the 38th international ACM SIGIR conference on research and development in information retrieval,pp\. 43–52\.Cited by:[§IV\-A1](https://arxiv.org/html/2609.21548#S4.SS1.SSS1.p2.1)\.
- \[42\]F\. M\. Harper and J\. A\. Konstan\(2015\)The movielens datasets: history and context\.Acm transactions on interactive intelligent systems \(tiis\)5\(4\),pp\. 1–19\.Cited by:[§IV\-A1](https://arxiv.org/html/2609.21548#S4.SS1.SSS1.p3.1)\.
- \[43\]J\. Tang and K\. Wang\(2018\)Personalized top\-n sequential recommendation via convolutional sequence embedding\.InProceedings of the eleventh ACM international conference on web search and data mining,pp\. 565–573\.Cited by:[1st item](https://arxiv.org/html/2609.21548#S4.I2.i1.p1.1),[TABLE II](https://arxiv.org/html/2609.21548#S4.T2.10.3.1)\.
- \[44\]K\. Diederik\(2014\)Adam: a method for stochastic optimization\.\(No Title\)\.Cited by:[§IV\-A3](https://arxiv.org/html/2609.21548#S4.SS1.SSS3.p1.1)\.
- \[45\]L\. Liu, L\. Cai, C\. Zhang, X\. Zhao, J\. Gao, W\. Wang, Y\. Lv, W\. Fan, Y\. Wang, M\. He,et al\.\(2023\)Linrec: linear attention mechanism for long\-term sequential recommender systems\.InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 289–299\.Cited by:[§IV\-A4](https://arxiv.org/html/2609.21548#S4.SS1.SSS4.p1.1)\.

Similar Articles