Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding
Summary
This paper proposes GenCDSR, a generative framework for cross-domain sequential recommendation with hybrid tokenization and serial-parallel decoding, achieving improved accuracy and significantly reduced inference latency compared to state-of-the-art baselines.
View Cached Full Text
Cached at: 08/03/26, 07:29 AM
# Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding
Source: [https://arxiv.org/html/2607.28659](https://arxiv.org/html/2607.28659)
\\useunder
\\ul\\setcctypeby\-nc\-nd
\(2026\)
###### Abstract\.
Cross\-domain sequential recommendation \(CDSR\) aims to model users’ dynamic interest transitions and sequential patterns across multiple domains\. Recently, generative recommendation \(GR\) has emerged, which first learns semantic identifiers \(SIDs\) using semantic information of items and models the recommendation task as autoregressive generation\. However, it faces two critical issues: 1\) ignoring collaborative correlations across different domains in tokenization step and 2\) adopting inefficient decoding strategies like beam search in generation step, which hinders GR’s application in real\-time services\. To address these limitations, we propose GenCDSR, an effective and efficient generative framework for CDSR\. Specifically, we design a cross\-domain hybrid tokenization mechanism that employs a multi\-tower architecture to jointly capture cross\-domain commonalities and domain\-specific distinctions through hierarchical shared\-specific and fine\-grained codebooks\. Furthermore, we develop a cross\-domain serial\-parallel decoding strategy that leverages the hierarchical SID structure to partially parallelize generation, significantly reducing inference latency while preserving generation consistency\. Experimental results on three public datasets validate that GenCDSR achieves a 1\.5% improvement in accuracy and an 85\.1% reduction in inference latency on average compared to SOTA baselines\. The implementation code and datasets are available online:[https://github\.com/Applied\-Machine\-Learning\-Lab/RecSys2026\_GenCDSR](https://github.com/Applied-Machine-Learning-Lab/RecSys2026_GenCDSR)\.
Cross\-Domain Sequential Recommendation, Generative Recommendation, Tokenization, Efficient Decoding
††journalyear:2026††copyright:cc††conference:20th ACM Conference on Recommender Systems; September 27\-October 02, 2026; Minneapolis, MN, USA††booktitle:20th ACM Conference on Recommender Systems \(RecSys ’26\), September 27\-October 02, 2026, Minneapolis, MN, USA††doi:10\.1145/3773078\.3831777††isbn:979\-8\-4007\-2284\-4/2026/09††ccs:Information systems Recommender systems## 1\.Introduction
Cross\-domain sequential recommendation \(CDSR\)\(Liuet al\.,[2025b](https://arxiv.org/html/2607.28659#bib.bib24),[a](https://arxiv.org/html/2607.28659#bib.bib41)\)models users’ interaction sequences across multiple commercial domains, such as product categories on e\-commerce and online entertainment platforms\(Caoet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib38); Xuet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib18); Huet al\.,[2026b](https://arxiv.org/html/2607.28659#bib.bib48); Chenet al\.,[2024a](https://arxiv.org/html/2607.28659#bib.bib50)\), to capture dynamic interest transitions across heterogeneous item spaces and improve recommendation performance\. A representative line of research models mixed behavioral sequences\. For example, TriCDR\(Maet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib37)\)extracts shared interests through multi\-granularity attention and cross\-domain contrastive learning\. Meanwhile, generative recommendation \(GR\) reformulates recommendation as autoregressive generation\. Existing GR frameworks\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8); Wanget al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib13)\)typically involve two steps: 1\)Tokenization, where a quantization model such as RQ\-VAE\(Leeet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib4)\)maps item semantic embeddings into sequences of semantic identifiers \(SIDs\) to capture item correlations; and 2\)Generation, where next\-token prediction leverages language models to capture temporal patterns and generate the SID of the next item\. For example, TIGER\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8)\)first learns SIDs through RQ\-VAE tokenization and then trains a seq2seq model to generate the next\-item SID\.
Although GR possesses great potential to improve CDSR, we find that directly applying GR on CDSR setting faces two key challenges\. First, in tokenization step cross\-domain collaborative correlations tend to be insufficiently modeled\(Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12); Jinet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib9)\)\. This is because most existing CDSR approaches either train a separate quantization model for each domain\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8)\)or adopt a unified shared quantization model indiscriminately for all domains\(Jinet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib9)\), making it difficult to explicitly encode the underlying collaborative correlations from cross\-domain interactions into SIDs\. This limitation would further lead to information loss in discretized representations, thereby restricting the model’s generalization ability\. We will further analyze this phenomenon in the feature fidelity analysis in Section[4\.5](https://arxiv.org/html/2607.28659#S4.SS5)\. Second, the generation step either suffers from efficiency bottleneck of token\-by\-token serial decoding\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8); Zhenget al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib7)\)or degraded recommendation quality under fully parallel decoding\(Freitag and Al\-Onaizan,[2017](https://arxiv.org/html/2607.28659#bib.bib42); Wanget al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib5)\)\. Specifically, the inference latency grows rapidly with decoding steps and candidate expansion in the widely adopted serial decoding strategies like Beam Search\(Freitag and Al\-Onaizan,[2017](https://arxiv.org/html/2607.28659#bib.bib42)\), posing a major obstacle to practical deployment on real\-time services\. Besides, common parallel decoding strategies like Multi\-token Prediction \(MTP\)\(Gloeckleet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib6)\)may predict multiple positions independently and fail to leverage semantic dependencies from preceding tokens, resulting in lower per\-position accuracy and degraded recommendation quality\. We will provide concrete evidence in Section[4\.3](https://arxiv.org/html/2607.28659#S4.SS3)\.
Therefore, to address these challenges, we propose GenCDSR, an effective and efficient generative cross\-domain sequential recommendation framework\. On the one hand, in the tokenization step, GenCDSR learns domain\-aware SIDs, where both shared and domain\-specific codebooks are jointly incorporated during SIDs training\. This design enables the unified modeling of cross\-domain semantic commonalities and domain\-specific distinctions within a generative framework, enhancing the expressive capacity of generative sequence modeling\. On the other hand, in the generation step, we further design a cross\-domain serial\-parallel decoding strategy, tailored to domain\-aware SIDs, allowing the generation process across different domains to be partially parallelized in a serial\-parallel manner\. Compared with conventional token\-by\-token serial decoding and parallel decoding, the proposed strategy significantly can reduce inference latency while maintaining recommendation quality, thereby improving the practical applicability of GR models\.
The main contributions of this work are summarized as follows:
- •We propose GenCDSR, an effective and efficient generative cross\-domain sequential recommendation framework that models both cross\-domain commonalities and domain\-specific distinctions under a unified generative paradigm\.
- •We propose cross\-domain hybrid tokenization mechanism featuring share\-specific multi\-tower RQ\-VAE and cross\-domain serial\-parallel decoding strategy featuring two\-pass state\-carry serial\-parallel SID decoding\.
- •Experiments on three public cross\-domain datasets demonstrate that GenCDSR outperforms SOTA baselines by 1\.5% in accuracy while achieving an 85\.1% reduction in inference latency\.
## 2\.Preliminary
### 2\.1\.Problem Definition
In this paper, we focus on the CDSR problem and use a dual\-domain setting as an example\. Let𝒰=\{u1,u2,…,u\|𝒰\|\}\\mathcal\{U\}=\\\{u\_\{1\},u\_\{2\},\\ldots,u\_\{\|\\mathcal\{U\}\|\}\\\}denote the set of users, and\|𝒰\|\|\\mathcal\{U\}\|is the total number of users\. We denote the two domains asAAandBB, with corresponding item sets𝒜\\mathcal\{A\}andℬ\\mathcal\{B\}, respectively\. For a useru∈𝒰u\\in\\mathcal\{U\}, their historical interactions in each domain are represented as chronologically ordered sequences:
\(1\)SuA=\(a1,a2,…,anA\),SuB=\(b1,b2,…,bnB\),S\_\{u\}^\{A\}=\(a\_\{1\},a\_\{2\},\\ldots,a\_\{n\_\{A\}\}\),\\quad S\_\{u\}^\{B\}=\(b\_\{1\},b\_\{2\},\\ldots,b\_\{n\_\{B\}\}\),whereai∈𝒜a\_\{i\}\\in\\mathcal\{A\}andbj∈ℬb\_\{j\}\\in\\mathcal\{B\}denote interacted items in domainsAAandBB, respectively, andnAn\_\{A\}andnBn\_\{B\}denote the corresponding sequence lengths in the two domains\.
To characterize users’ cross\-domain behaviors, interactions from both domains are merged according to their temporal order into a unified cross\-domain behavior stream:
\(2\)S¯u=\(v1,v2,…,vn\),vt∈𝒜∪ℬ,\\bar\{S\}\_\{u\}=\(v\_\{1\},v\_\{2\},\\ldots,v\_\{n\}\),\\quad v\_\{t\}\\in\\mathcal\{A\}\\cup\\mathcal\{B\},wheren=nA\+nBn=n\_\{A\}\+n\_\{B\}denotes the total number of historical interactions of useruuacross the two domains\.
The objective of CDSR is to predict the next item that a useruuis most likely to interact with in each domain, given the cross\-domain behavior stream as well as the corresponding domain\-specific interaction histories\. Formally, the task can be formulated as the following two prediction objectives:
\(3\)a^\\displaystyle\\hat\{a\}=argmaxa∈𝒜P\(vn\+1=a∣S¯u,SuA,SuB\),\\displaystyle=\\arg\\max\_\{a\\in\\mathcal\{A\}\}P\(v\_\{n\+1\}=a\\mid\\bar\{S\}\_\{u\},S\_\{u\}^\{A\},S\_\{u\}^\{B\}\),\(4\)b^\\displaystyle\\hat\{b\}=argmaxb∈ℬP\(vn\+1=b∣S¯u,SuA,SuB\)\.\\displaystyle=\\arg\\max\_\{b\\in\\mathcal\{B\}\}P\(v\_\{n\+1\}=b\\mid\\bar\{S\}\_\{u\},S\_\{u\}^\{A\},S\_\{u\}^\{B\}\)\.
### 2\.2\.RQ\-VAE
Residual Quantized Variational Autoencoder \(RQ\-VAE\)\(Leeet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib4)\)aims to map continuous representations into SIDs\. Specifically, given an input embedding𝐱\\mathbf\{x\}, an encoder first maps it into a latent semantic vector𝐳\\mathbf\{z\}, which is then quantized withLLlevel\-wise codebooks\. At each levell=1,…,Ll=1,\\ldots,L, the model selects a codeword from the codebook𝒞l=\{𝐜1l,…,𝐜Kl\}\\mathcal\{C\}^\{l\}=\\\{\\mathbf\{c\}^\{l\}\_\{1\},\\ldots,\\mathbf\{c\}^\{l\}\_\{K\}\\\}, whereKKdenotes the codebook size, by minimizing theℓ2\\ell\_\{2\}distance to the current residual𝐫l−1\\mathbf\{r\}^\{l\-1\}\. Formally, the SIDsls^\{l\}is obtained as:
\(5\)sl=argminj‖𝐫l−1−𝐜jl‖2,s^\{l\}=\\arg\\min\_\{j\}\\left\\\|\\mathbf\{r\}^\{l\-1\}\-\\mathbf\{c\}^\{l\}\_\{j\}\\right\\\|\_\{2\},and the residual is updated accordingly:
\(6\)𝐫l=𝐫l−1−𝐜sll\.\\mathbf\{r\}^\{l\}=\\mathbf\{r\}^\{l\-1\}\-\\mathbf\{c\}^\{l\}\_\{s^\{l\}\}\.
AfterLLquantization levels, the original embedding is represented by a length\-LLdiscrete SID sequence\{s1,…,sL\}\\\{s^\{1\},\\ldots,s^\{L\}\\\}\. The corresponding quantized latent embedding is computed by aggregating the selected codewords, i\.e\.,𝐳^=∑l=1L𝐜sll\\hat\{\\mathbf\{z\}\}=\\sum\_\{l=1\}^\{L\}\\mathbf\{c\}^\{l\}\_\{s^\{l\}\}, which is further decoded into𝐱^\\hat\{\\mathbf\{x\}\}to reconstruct the input embedding𝐱\\mathbf\{x\}\. RQ\-VAE is trained by jointly minimizing a reconstruction lossℒrecon\\mathcal\{L\}\_\{\\mathrm\{recon\}\}and a residual quantization lossℒRQ\\mathcal\{L\}\_\{\\mathrm\{RQ\}\}, formulated as:
\(7\)ℒ=ℒrecon\+ℒRQ,\\displaystyle\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{recon\}\}\+\\mathcal\{L\}\_\{\\mathrm\{RQ\}\},\(8\)ℒrecon=‖𝐱−𝐱^‖22,\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{recon\}\}=\\left\\\|\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\\right\\\|\_\{2\}^\{2\},\(9\)ℒRQ=∑l=1L\(‖sg\(𝐫l−1\)−𝐜sll‖22\+β‖𝐫l−1−sg\(𝐜sll\)‖22\),\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{RQ\}\}=\\sum\_\{l=1\}^\{L\}\\left\(\\left\\\|\\mathrm\{sg\}\(\\mathbf\{r\}^\{l\-1\}\)\-\\mathbf\{c\}^\{l\}\_\{s^\{l\}\}\\right\\\|\_\{2\}^\{2\}\+\\beta\\left\\\|\\mathbf\{r\}^\{l\-1\}\-\\mathrm\{sg\}\(\\mathbf\{c\}^\{l\}\_\{s^\{l\}\}\)\\right\\\|\_\{2\}^\{2\}\\right\),wheresg\(⋅\)\\mathrm\{sg\}\(\\cdot\)denotes the stop\-gradient operator andβ\\betacontrols the balance between codebook learning and encoder updates\.
## 3\.Method
### 3\.1\.Overview
Existing generative cross\-domain sequential recommendation methods\(Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12); Jinet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib9)\)fail to disentangle shared and domain\-specific factors, relying on fully shared or purely semantic representations\. Furthermore, their token\-by\-token autoregressive decoding limits inference efficiency\. To address these issues, we propose GenCDSR, an efficient generative framework comprising core components: Cross\-Domain Hybrid Tokenization and Cross\-Domain Serial\-Parallel Decoding\. The framework is shown in Figure[1](https://arxiv.org/html/2607.28659#S3.F1)\.
Figure 1\.Overview of the proposed GenCDSR framework\.
### 3\.2\.Cross\-Domain Hybrid Tokenization
Cross\-domain sequential recommendation requires balancing shared semantic commonalities with domain\-specific characteristics\. However, existing generative methods either adopt fully shared representation spaces\(Jinet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib9)\), which weaken domain\-specific features, or rely on semantic representations\(Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12)\)that overlook cross\-domain collaborative signals\. To address this issue, we introduce cross\-domain hybrid tokenization \(Figure[1](https://arxiv.org/html/2607.28659#S3.F1), upper left\), which employs a multi\-tower architecture to capture both factors\. It follows a two\-stage hierarchy: Stage 1 quantizes shared and specific representations into a length\-L1L\_\{1\}SID segment, while Stage 2 generates a length\-L2L\_\{2\}domain\-specific segment for finer details\. Each itemviv\_\{i\}is thus represented by a hierarchical SID sequencevi=\[t1,…,tL\]v\_\{i\}=\[t\_\{1\},\\ldots,t\_\{L\}\], whereL=L1\+L2L=L\_\{1\}\+L\_\{2\}andtit\_\{i\}denotes the SID from theii\-th\-level codebook\.
#### 3\.2\.1\.Stage 1: Shared–Specific Tokenization
Although different domains exhibit distinct data distributions, they still share relatively stable high\-level semantics and behavioral regularities\. Therefore, tokenizing each domain independently can fragment the semantic space, thereby weakening cross\-domain generalization and limiting collaborative modeling\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8); Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12)\)\. To address this, our first stage extends the RQ\-VAE tokenization framework with a Shared\-Specific Tokenization \(SST\) mechanism, which jointly preserves shared semantics and domain\-specific signals and establishes a unified foundation for subsequent domain\-specific modeling\.
Specifically, we construct both shared and domain\-specific encoders and residual tokenization modules to provide a unified semantic backbone alongside distinct domain features\. For an item in domaind∈\{A,B\}d\\in\\\{A,B\\\}with semantic representation𝐱d\\mathbf\{x\}\_\{d\}, extracted via pre\-trained models like LLaMA\-7B\(Touvronet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib22)\)or Qwen2\.5\-7B\(Baiet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib23)\), we first map it into latent semantic spaces using both encoders:
\(10\)𝐳sh=Encsh\(𝐱d\),𝐳dsp=Encd\(𝐱d\),\\mathbf\{z\}^\{\\mathrm\{sh\}\}=\\mathrm\{Enc\}\_\{\\mathrm\{sh\}\}\(\\mathbf\{x\}\_\{d\}\),\\quad\\mathbf\{z\}^\{\\mathrm\{sp\}\}\_\{d\}=\\mathrm\{Enc\}\_\{d\}\(\\mathbf\{x\}\_\{d\}\),whereEncsh\(⋅\)\\mathrm\{Enc\}\_\{\\mathrm\{sh\}\}\(\\cdot\)denotes the shared encoder andEncd\(⋅\)\\mathrm\{Enc\}\_\{d\}\(\\cdot\)denotes the domain\-specific encoder for domaindd\. The shared encoder is applied to item embeddings from all domains, whereas each domain\-specific encoder is only applied to items within its domain\.
Correspondingly, the shared latent representation𝐳sh\\mathbf\{z\}^\{\\mathrm\{sh\}\}and the domain\-specific latent representation𝐳dsp\\mathbf\{z\}^\{\\mathrm\{sp\}\}\_\{d\}are fed into two independent residual quantization modules\. Each representation is then quantized into discrete SIDs via anL1L\_\{1\}\-level residual quantization process, producing𝐳′sh\\mathbf\{z\}^\{\\prime\\mathrm\{sh\}\}and𝐳d′sp\\mathbf\{z\}^\{\\prime\\mathrm\{sp\}\}\_\{d\}as the discrete latent representations from the shared and domain\-specific branches\.
To adaptively choose between the shared and domain\-specific branches, we use a Gumbel\-Softmax router to sample near\-discrete routing weights and fuse the corresponding quantized outputs:
\(11\)𝝅d=f\(𝐱d\),\\displaystyle\\bm\{\\pi\}\_\{d\}=f\\\!\\left\(\\mathbf\{x\}\_\{d\}\\right\),\(12\)gsh,gdsp=GumbelSoftmax\(𝝅d;τ\),\\displaystyle g^\{\\mathrm\{sh\}\},g^\{\\mathrm\{sp\}\}\_\{d\}=\\mathrm\{GumbelSoftmax\}\(\\bm\{\\pi\}\_\{d\};\\tau\),\(13\)𝐳d′=gsh⋅𝐳′sh\+gdsp⋅𝐳d′sp,\\displaystyle\\mathbf\{z\}^\{\\prime\}\_\{d\}=g^\{\\mathrm\{sh\}\}\\cdot\\mathbf\{z\}^\{\\prime\\mathrm\{sh\}\}\+g^\{\\mathrm\{sp\}\}\_\{d\}\\cdot\\mathbf\{z\}^\{\\prime\\mathrm\{sp\}\}\_\{d\},wheref\(⋅\)f\(\\cdot\)is an MLP andτ\\tauis the temperature\.
To preserve information not captured by the Stage 1 quantized representations, we compute a residual by subtracting the Gumbel\-Softmax fused quantized output𝐳d′\\mathbf\{z\}^\{\\prime\}\_\{d\}from the corresponding fused continuous latent feature\. This residual, denoted as𝐳d′′\\mathbf\{z\}^\{\\prime\\prime\}\_\{d\}, serves as the input to Stage 2 for further domain\-specific tokenization\. At this point, Stage 1 has already produced the finalL1L\_\{1\}SID tokens\[t1,…,tL1\]\[t\_\{1\},\\ldots,t\_\{L\_\{1\}\}\]\. Specifically, the Stage 2 input is defined as:
\(14\)𝐳d′′=\(gsh⋅𝐳sh\+gdsp⋅𝐳dsp\)−𝐳′d\.\\mathbf\{z\}^\{\\prime\\prime\}\_\{d\}=\\Big\(g^\{\\mathrm\{sh\}\}\\cdot\\mathbf\{z\}^\{\\mathrm\{sh\}\}\+g^\{\\mathrm\{sp\}\}\_\{d\}\\cdot\\mathbf\{z\}^\{\\mathrm\{sp\}\}\_\{d\}\\Big\)\-\\mathbf\{z^\{\\prime\}\}\_\{d\}\.
#### 3\.2\.2\.Stage 2: Fine\-Grained Specific Tokenization
In practical recommendation systems, different domains often exhibit substantially different item distributions and interaction patterns\. If a shared codebook is still used for low\-level representations, domain\-specific patterns may be overly constrained, which can even lead to semantic ambiguity\(Jinet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib9); Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12)\)\. Therefore, in Stage 2, we employ fully domain\-specific RQ tokenization named Fine\-Grained Specific Tokenization \(FGST\) modules to better capture fine\-grained, domain\-dependent variations\. Starting from the Stage 2 input𝐳d′′\\mathbf\{z\}^\{\\prime\\prime\}\_\{d\}, we apply anL2L\_\{2\}\-level residual quantization to obtain the domain\-specific quantized representation𝐳^d\\hat\{\\mathbf\{z\}\}\_\{d\}, which is then fed into the corresponding domain decoder for reconstruction:
\(15\)𝐱^d=Decd\(𝐳^d\),\\hat\{\\mathbf\{x\}\}\_\{d\}=\\mathrm\{Dec\}\_\{d\}\\\!\\left\(\\hat\{\\mathbf\{z\}\}\_\{d\}\\right\),whereDecd\(⋅\)\\mathrm\{Dec\}\_\{d\}\(\\cdot\)is the decoder associated with domaindd\. After Stage 2, the remainingL2L\_\{2\}SID tokens are determined and combined with theL1L\_\{1\}tokens from Stage 1, yielding the complete SID sequence\[t1,…,tL\]\[t\_\{1\},\\ldots,t\_\{L\}\]withL=L1\+L2L=L\_\{1\}\+L\_\{2\}\.
The total loss for tokenization can be formulated as:
\(16\)ℒrecon=∑d∈\{A,B\}‖𝐱d−𝐱^d‖22,\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{recon\}\}=\\sum\_\{d\\in\\\{A,B\\\}\}\\big\\\|\\mathbf\{x\}\_\{d\}\-\\hat\{\\mathbf\{x\}\}\_\{d\}\\big\\\|\_\{2\}^\{2\},\(17\)ℒRQ=ℒRQsh\+ℒRQsp,A\+ℒRQsp,B,\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{RQ\}\}=\\mathcal\{L\}^\{\\mathrm\{sh\}\}\_\{\\mathrm\{RQ\}\}\+\\mathcal\{L\}^\{\\mathrm\{sp\},A\}\_\{\\mathrm\{RQ\}\}\+\\mathcal\{L\}^\{\\mathrm\{sp\},B\}\_\{\\mathrm\{RQ\}\},\(18\)ℒ=ℒrecon\+ℒRQ\.\\displaystyle\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{recon\}\}\+\\mathcal\{L\}\_\{\\mathrm\{RQ\}\}\.whereℒRQsh\\mathcal\{L\}^\{\\mathrm\{sh\}\}\_\{\\mathrm\{RQ\}\}is the RQ loss of the shared tokenization, whileℒRQsp,A\\mathcal\{L\}^\{\\mathrm\{sp\},A\}\_\{\\mathrm\{RQ\}\}andℒRQsp,B\\mathcal\{L\}^\{\\mathrm\{sp\},B\}\_\{\\mathrm\{RQ\}\}are the RQ losses of the domain\-specific tokenization for domainsAAandBB\.
### 3\.3\.Cross\-Domain Serial\-Parallel Decoding
In generative cross\-domain sequential recommendation, existing methods\(Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12); Jinet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib9)\)typically rely on token\-by\-token serial decoding\. This creates a severe latency bottleneck for practical deployment and underutilizes the two\-stage SID structure, thereby limiting the propagation of collaborative signals\. To achieve efficient and scalable generation that better matches this hierarchical identifier design, we propose a cross\-domain serial\-parallel decoding strategy, as shown at the bottom of Figure[1](https://arxiv.org/html/2607.28659#S3.F1)\.
After cross\-domain hybrid tokenization, the target SID ofvn\+1v\_\{n\+1\}is represented as a length\-LLsequence of special tokens, i\.e\.,vn\+1=\[t1,t2,…,tL\]v\_\{n\+1\}=\[t\_\{1\},t\_\{2\},\\ldots,t\_\{L\}\]\. During prompting, the ground\-truth target SID tokens are replaced withLLplaceholder tokens\[<SP\_1\>,…,<SP\_L\>\]\[\\texttt\{<SP\\\_1\>\},\\ldots,\\texttt\{<SP\\\_L\>\}\], which enables a single forward pass to pre\-compute the hidden states for all future token positions in a prefix\-conditioned manner\. Leth01,h11,…,hL11h^\{1\}\_\{0\},h^\{1\}\_\{1\},\\ldots,h^\{1\}\_\{L\_\{1\}\}denote the last\-layer hidden states from the first LLM call, which correspond to the prefixh01h^\{1\}\_\{0\}and theL1L\_\{1\}placeholder positionsh11∼hL11h^\{1\}\_\{1\}\\sim h^\{1\}\_\{L\_\{1\}\}\. We then define an evolving context statesls\_\{l\}initialized by the prefix:
Forl=1,…,L1l=1,\\ldots,L\_\{1\}, the Step 1 head produces the token distribution at positionlland updates the context state using the embedding𝐞l=Embl\(tl\)\\mathbf\{e\}\_\{l\}=\\mathrm\{Emb\}\_\{l\}\(t\_\{l\}\)of the selected tokentlt\_\{l\}:
\(20\)𝐩l\(1\)=softmax\(Head1\(hl1,sl\)\),sl\+1=sl\+𝐞l\.\\mathbf\{p\}\_\{l\}^\{\(1\)\}=\\mathrm\{softmax\}\\big\(\\mathrm\{Head\}\_\{1\}\(h^\{1\}\_\{l\},s\_\{l\}\)\\big\),\\quad s\_\{l\+1\}=s\_\{l\}\+\\mathbf\{e\}\_\{l\}\.
Next, we invoke the LLM once more by filling the firstL1L\_\{1\}placeholders with the Step 1 tokens, obtaining updated hidden statesh02,h12,…,hL2h^\{2\}\_\{0\},h^\{2\}\_\{1\},\\ldots,h^\{2\}\_\{L\}\. LetsL1\+1s\_\{L\_\{1\}\+1\}be the carried state after Step 1\. For each remaining positionl∈\{L1\+1,…,L\}l\\in\\\{L\_\{1\}\+1,\\ldots,L\\\}, the Step 2 head predicts in parallel conditioned on the carried state:
\(21\)𝐩l\(2\)=softmax\(Head2\(hl2,sL1\+1\)\),sl\+1=sl\+𝐞l\.\\mathbf\{p\}\_\{l\}^\{\(2\)\}=\\mathrm\{softmax\}\\big\(\\mathrm\{Head\}\_\{2\}\(h^\{2\}\_\{l\},s\_\{L\_\{1\}\+1\}\)\\big\),\\quad s\_\{l\+1\}=s\_\{l\}\+\\mathbf\{e\}\_\{l\}\.
### 3\.4\.Training and Inference
#### 3\.4\.1\.Unified Recommender Training
Firstly, we conduct Unified Recommender Training on the merged cross\-domain dataS¯u\\bar\{S\}\_\{u\}with useru∈𝒰u\\in\\mathcal\{U\}as Equation \([2](https://arxiv.org/html/2607.28659#S2.E2)\) to enhance the model’s ability to capture transferable cross\-domain patterns\. The unified objective maximizes the token\-level cross\-entropy over all SID positions:
\(22\)maxΦ∑u∈𝒰\(∑l=1L1log𝐩l\(Φ,1\)\(tl\)\+∑l=L1\+1Llog𝐩l\(Φ,2\)\(tl\)\),\\max\_\{\\Phi\}\\ \\sum\_\{u\\in\\mathcal\{U\}\}\\left\(\\sum\_\{l=1\}^\{L\_\{1\}\}\\log\\mathbf\{p\}^\{\(\\Phi,1\)\}\_\{l\}\(t\_\{l\}\)\+\\sum\_\{l=L\_\{1\}\+1\}^\{L\}\\log\\mathbf\{p\}^\{\(\\Phi,2\)\}\_\{l\}\(t\_\{l\}\)\\right\),whereΦ\\Phidenotes the backbone parameters, andvn\+1=\[t1,…,tL\]v\_\{n\+1\}=\[t\_\{1\},\\ldots,t\_\{L\}\]is the target SID sequence withL=L1\+L2L=L\_\{1\}\+L\_\{2\}\. Moreover,𝐩l\(Φ,1\)\(⋅\)\\mathbf\{p\}^\{\(\\Phi,1\)\}\_\{l\}\(\\cdot\)and𝐩l\(Φ,2\)\(⋅\)\\mathbf\{p\}^\{\(\\Phi,2\)\}\_\{l\}\(\\cdot\)denote the token distributions produced underΦ\\Phiin Step 1 and Step 2 decoding, respectively\.
#### 3\.4\.2\.Domain\-Specific Fine\-tuning
Subsequently, to better accommodate domain\-specific data characteristics while avoiding overwriting the cross\-domain knowledge learned during Unified Recommender Training, we freeze the backbone parametersΦ\\Phiand fine\-tune a lightweight, domain\-specific LoRA\(Huet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib20)\)adapterΘd\\Theta\_\{d\}for each domainddusing the corresponding dataSudS\_\{u\}^\{d\}in Equation \([1](https://arxiv.org/html/2607.28659#S2.E1)\)\. The resulting domain\-specific optimization objective is:
\(23\)maxΘd∑u∈𝒰\(∑l=1L1log𝐩l\(Φ\+Θd,1\)\(tl\)\+∑l=L1\+1Llog𝐩l\(Φ\+Θd,2\)\(tl\)\)\.\\max\_\{\\Theta\_\{d\}\}\\ \\sum\_\{u\\in\\mathcal\{U\}\}\\left\(\\sum\_\{l=1\}^\{L\_\{1\}\}\\log\\mathbf\{p\}^\{\(\\Phi\+\\Theta\_\{d\},1\)\}\_\{l\}\(t\_\{l\}\)\+\\sum\_\{l=L\_\{1\}\+1\}^\{L\}\\log\\mathbf\{p\}^\{\(\\Phi\+\\Theta\_\{d\},2\)\}\_\{l\}\(t\_\{l\}\)\\right\)\.
#### 3\.4\.3\.Inference
During inference, we activate the domain\-specific adapterΘd\\Theta\_\{d\}on top of the backboneΦ\\Phiand perform trie\-constrained serial\-parallel decoding to generate a SID sequence, which is then mapped to the corresponding item\. To ensure that the generated identifier corresponds to a valid item inℐd\\mathcal\{I\}\_\{d\}, we follow existing work\(Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12); Jinet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib9); Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8)\)and restrict decoding to the feasible set𝒯d\\mathcal\{T\}\_\{d\}induced by the Trie\(De La Briandais,[1959](https://arxiv.org/html/2607.28659#bib.bib14)\)\. The inference objective is:
\(24\)v^n\+1=argmaxvn\+1∈𝒯d\(∑l=1L1log𝐩l\(Φ\+Θd,1\)\(tl\)\+∑l=L1\+1Llog𝐩l\(Φ\+Θd,2\)\(tl\)\)\.\\hat\{v\}\_\{n\+1\}=\\operatorname\*\{arg\\,max\}\_\{v\_\{n\+1\}\\in\\mathcal\{T\}\_\{d\}\}\\left\(\\sum\_\{l=1\}^\{L\_\{1\}\}\\log\\mathbf\{p\}^\{\(\\Phi\+\\Theta\_\{d\},1\)\}\_\{l\}\(t\_\{l\}\)\+\\sum\_\{l=L\_\{1\}\+1\}^\{L\}\\log\\mathbf\{p\}^\{\(\\Phi\+\\Theta\_\{d\},2\)\}\_\{l\}\(t\_\{l\}\)\\right\)\.
## 4\.Experiments
In this section, we conduct extensive experiments on three public datasets and answer the following research questions \(RQs\):
- •RQ1:What is the performance of GenCDSR compared with SOTA baseline methods?
- •RQ2:Can cross\-domain serial\-parallel decoding reduce inference latency without sacrificing accuracy?
- •RQ3:What are the effects of shared\-specific tokenization and fine\-grained specific tokenization?
- •RQ4:How does the cross\-domain hybrid tokenization contribute to improving the recommendation performance?
- •RQ5:How does the allocation of shared\-specific tokenization and fine\-grained specific tokenization affect the performance?
### 4\.1\.Experimental Settings
#### 4\.1\.1\.Datasets
Following prior CDSR studies\(Liuet al\.,[2025b](https://arxiv.org/html/2607.28659#bib.bib24),[a](https://arxiv.org/html/2607.28659#bib.bib41)\), we evaluate our method on three public cross\-domain datasets: Clothing–Sports, Electronics–Phone, and Book–Movie\. The first two are collected from Amazon111[https://jmcauley\.ucsd\.edu/data/amazon/index\_2014\.html](https://jmcauley.ucsd.edu/data/amazon/index_2014.html), while the third is collected from Douban222[https://github\.com/fengzhu1/GA\-DTCDR/tree/main/Data](https://github.com/fengzhu1/GA-DTCDR/tree/main/Data)\. We filter out users with fewer than five interactions and items with fewer than three interactions in either domain, and chronologically merge cross\-domain interactions into mixed behavior sequences for each user\. Following the leave\-one\-out strategy\(Zhenget al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib7); Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8)\), we use the last two interactions for validation and testing, respectively\. Table[1](https://arxiv.org/html/2607.28659#S4.T1)summarizes the dataset statistics\.
Table 1\.Dataset statistics, including the number of users, number of items, sparsity, the number of overlapped users, and the average interaction length\.DatasetUsersItemsSparsityOverlapAvg\.lenClothing9,9333,27899\.70%3,96210\.71Sports4,2631,02199\.04%Electronics20,72810,49299\.93%11,6988\.30Phone11,7622,24699\.88%Book1,38112,42699\.88%1,26568\.30Movie2,21316,53799\.92%
#### 4\.1\.2\.Baselines
To evaluate the effectiveness of GenCDSR, we compare it with representative baseline models, which can be divided into two categories: single\-domain sequential recommendation and cross\-domain sequential recommendation\.
##### Single\-Domain Sequential Recommendation\.
These methods model user preferences based solely on historical interaction sequences within a single domain, without leveraging auxiliary information from other domains\.
- •GRU4Rec\(Hidasiet al\.,[2015](https://arxiv.org/html/2607.28659#bib.bib1)\)models user behavior sequences using gated recurrent units to capture dynamic user preferences\.
- •BERT4Rec\(Sunet al\.,[2019](https://arxiv.org/html/2607.28659#bib.bib40)\)employs bidirectional self\-attention to model contextual dependencies in user behavior sequences via a masked prediction objective\.
- •SASRec\(Kang and McAuley,[2018](https://arxiv.org/html/2607.28659#bib.bib39)\)adopts unidirectional self\-attention to model users’ sequential behaviors and capture sequential relationships\.
- •TIGER\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8)\)first utilizes RQ\-VAE to obtain SIDs and adopts T5 to generate the SIDs of next item\.
##### Cross\-Domain Sequential Recommendation\.
CDSR methods aim to alleviate data sparsity by exploiting user interactions across multiple domains\.
- •C2DSR\(Caoet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib38)\)models item relationships across domains via graph neural networks to enable cross\-domain knowledge transfer\.
- •TriCDR\(Maet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib37)\)captures fine\-grained user interests through cross\-domain contrastive learning objectives\.
- •LLM4CDSR\(Liuet al\.,[2025a](https://arxiv.org/html/2607.28659#bib.bib41)\)introduces large language models to learn unified item semantic representations across domains and model user preferences via user profiling\.
- •GenCDR\(Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12)\)leverages a cross\-domain tokenization structure with unified semantic associations to support cross\-domain generative recommendation\.
#### 4\.1\.3\.Implementation Details
Experiments are averaged over 3 runs on 8 NVIDIA RTX PRO 6000 GPUs\. For item tokenization, we obtain item text embeddings by mean\-pooling the last\-layer hidden states of LLaMA\-7B\(Touvronet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib22)\)and Qwen2\.5\-7B\(Baiet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib23)\), and discretize the embeddings via an RQ\-VAE with a two\-stage tokenization process \(L1=2,L2=2L\_\{1\}=2,L\_\{2\}=2\), where each level has 256 code vectors of 128 dimensions\. The tokenization model is trained using AdamW\(Loshchilov and Hutter,[2019](https://arxiv.org/html/2607.28659#bib.bib43)\)with lr=1×10−3=1\\times 10^\{\-3\}and batch size=1,024=1\{,\}024\. We instantiate our method on T5\(Raffelet al\.,[2020](https://arxiv.org/html/2607.28659#bib.bib17)\)and Qwen3\-0\.6B\(Yanget al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib21)\)\. For T5 we use the same setting as TIGER\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8)\), which is optimized using AdamW with lr=5×10−4=5\\times 10^\{\-4\}and batch size=256=256and domain\-specific LoRA with rank=8=8andα=32\\alpha=32\. Qwen3\-0\.6B uses AdamW with lr=1×10−4=1\\times 10^\{\-4\}and batch size=128=128and LoRA with rank=8=8andα=16\\alpha=16\. During inference, we apply beam search with a beam size of 20\.
#### 4\.1\.4\.Evaluation Metrics
Following previous works\(Liuet al\.,[2025a](https://arxiv.org/html/2607.28659#bib.bib41); Caoet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib38); Wanget al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib5)\), we adopt standard Top\-kkrecommendation metrics to evaluate model performance, including Hit Ratio \(H@kk\) and Normalized Discounted Cumulative Gain \(N@kk\) atk∈\{5,10\}k\\in\\\{5,10\\\}\. For efficiency evaluation, we further report the average generation latency \(LT\) measured in milliseconds \(ms\)\.
Table 2\.Overall performance comparison on all datasets\. H@K and N@K denote Hit Ratio and NDCG atkk, respectively\. The best results are highlighted in bold, while the best baseline results are underlined\. ”\*” indicates statistically significant improvements over the strongest baseline according to a two\-sided t\-test \(p<0\.05p<0\.05\)\.DatasetMetricSDSRCDSROursDiscriminativeGenerativeDiscriminativeGenerativeGRU4RecBERT4RecSASRecTIGERC2DSRTriCDRLLM4CDSRGenCDRClothingH@50\.30240\.33180\.33690\.58090\.55470\.56490\.5586\\ul0\.58480\.5869⋆0\.5869^\{\\star\}H@100\.33780\.36960\.37520\.59610\.56920\.57780\.5818\\ul0\.60760\.6097⋆0\.6097^\{\\star\}N@50\.27480\.29940\.30790\.52200\.52300\.53240\.5346\\ul0\.53610\.5373⋆0\.5373^\{\\star\}N@100\.28160\.30170\.30880\.52690\.52770\.53860\.5428\\ul0\.54200\.5447⋆0\.5447^\{\\star\}SportsH@50\.17970\.18680\.19390\.45770\.45280\.45960\.4668\\ul0\.49790\.5015⋆0\.5015^\{\\star\}H@100\.18090\.18760\.19480\.46670\.48850\.49520\.5038\\ul0\.51820\.5207⋆0\.5207^\{\\star\}N@50\.11340\.11890\.12360\.44840\.35860\.36640\.3749\\ul0\.46530\.4671⋆0\.4671^\{\\star\}N@100\.11380\.11920\.12410\.45130\.36680\.37450\.3826\\ul0\.47180\.4735⋆0\.4735^\{\\star\}ElectronicsH@50\.01160\.01330\.01390\.04840\.04480\.04390\.0452\\ul0\.05310\.0543⋆0\.0543^\{\\star\}H@100\.01540\.01790\.01860\.05960\.05250\.05350\.0567\\ul0\.06480\.0677⋆0\.0677^\{\\star\}N@50\.01860\.02040\.02270\.04010\.03790\.03760\.0347\\ul0\.04440\.0447⋆0\.0447^\{\\star\}N@100\.02130\.02350\.02610\.04370\.04040\.04070\.0384\\ul0\.04820\.0491⋆0\.0491^\{\\star\}PhoneH@50\.03310\.03610\.03760\.05960\.04900\.05090\.0560\\ul0\.07800\.0809⋆0\.0809^\{\\star\}H@100\.04410\.04750\.04890\.07860\.06430\.06530\.0731\\ul0\.10450\.1061⋆0\.1061^\{\\star\}N@50\.02560\.02840\.03120\.04760\.04080\.04050\.0403\\ul0\.06160\.0622⋆0\.0622^\{\\star\}N@100\.02910\.03240\.03560\.05370\.04570\.04510\.0457\\ul0\.06910\.0702⋆0\.0702^\{\\star\}BookH@50\.01240\.01520\.01780\.12900\.02010\.03150\.0516\\ul0\.13360\.1350⋆0\.1350^\{\\star\}H@100\.01980\.01980\.02080\.13900\.03720\.03440\.0745\\ul0\.14520\.1469⋆0\.1469^\{\\star\}N@50\.00620\.00810\.00960\.11890\.00980\.02050\.0372\\ul0\.12580\.1270⋆0\.1270^\{\\star\}N@100\.00740\.00920\.01100\.12220\.01480\.02370\.0444\\ul0\.12960\.1308⋆0\.1308^\{\\star\}MovieH@50\.06360\.06980\.07490\.16050\.14720\.15660\.1621\\ul0\.17260\.1875⋆0\.1875^\{\\star\}H@100\.10700\.11280\.11890\.24730\.20540\.21390\.2216\\ul0\.26220\.2671⋆0\.2671^\{\\star\}N@50\.03660\.04080\.04490\.10310\.1181\\ul0\.12950\.11760\.11200\.1189⋆0\.1189^\{\\star\}N@100\.05060\.05480\.05960\.13110\.12840\.13600\.1344\\ul0\.14080\.1445⋆0\.1445^\{\\star\}
### 4\.2\.Overall Performance \(RQ1\)
Table[2](https://arxiv.org/html/2607.28659#S4.T2)presents the overall performance comparison between GenCDSR and all baseline methods across three cross\-domain datasets\. Based on the experimental results, several key observations can be derived\. First, GenCDSR achieves the best performance across most datasets and evaluation metrics\. This indicates that by coupling generative modeling with cross\-domain knowledge transfer via our cross\-domain hybrid tokenization and serial–parallel decoding mechanism, it effectively enhances cross\-domain recommendation capability\. Second, cross\-domain signals effectively enhance single\-domain representations, thereby alleviating data sparsity\. In general, CDSR models outperform SDSR baselines, suggesting that information from related domains provides complementary supervision and helps the model learn more robust user preferences from limited target\-domain interactions\. Third, generative models are more suitable for cross\-domain scenarios\. Compared with discriminative approaches, they are better at modeling sequential dependencies and transferable semantic patterns across domains\. This indicates that generative frameworks are better aligned with the goal of cross\-domain recommendation, namely, capturing both shared knowledge and domain\-specific characteristics\.
### 4\.3\.Efficiency Analysis \(RQ2\)
We compare our proposed serial\-parallel decoding strategy with the following SOTA decoding strategies on T5 and Qwen3\-0\.6B\.
- •Beam Search\(Freitag and Al\-Onaizan,[2017](https://arxiv.org/html/2607.28659#bib.bib42)\)is a standard token\-by\-token autoregressive decoding strategy which ensures grounded item\.
- •Multi\-token Prediction \(MTP\)\(Gloeckleet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib6)\)is a multi\-token prediction strategy that predicts multiple tokens in parallel\.
- •NEZHA\(Wanget al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib5)\)uses a self\-drafting head with hash\-set verification for fast SID decoding to improve efficiency\.
Table 3\.Efficiency analysis of the proposed cross\-domain serial\-parallel decoding\. H@K, N@K, and LT denote Hit Ratio, NDCG atkk, and average generation latency \(in milliseconds\), respectively\. The best results are highlighted in bold\.DatasetMetricT5Qwen3\-0\.6BMTPNEZHABeam SearchOursMTPNEZHABeam SearchOursClothingH@50\.56740\.57190\.60130\.60130\.58690\.59370\.60090\.61530\.61530\.6046H@100\.58330\.58110\.62270\.62270\.60970\.60520\.60820\.63480\.63480\.6209N@50\.52040\.52360\.53910\.53910\.53730\.57030\.57990\.58750\.58750\.5807N@100\.52550\.52660\.54610\.54610\.54470\.57400\.58230\.59390\.59390\.5860LT↓\\downarrow0\.18270\.31104\.60080\.48350\.56621\.319047\.29836\.4449SportsH@50\.48170\.47930\.48470\.50150\.50150\.50680\.50620\.51100\.51590\.5159H@100\.49670\.49070\.50080\.52070\.52070\.50960\.50920\.52360\.52970\.5297N@50\.46160\.45780\.46370\.46710\.46710\.49060\.49320\.49800\.50060\.5006N@100\.46340\.46050\.46900\.47350\.47350\.49650\.49420\.50200\.50490\.5049LT↓\\downarrow0\.35890\.55614\.60220\.67541\.32901\.336631\.26376\.5448ElectronicsH@50\.03810\.02940\.05290\.05430\.05430\.05340\.05170\.05220\.05740\.0574H@100\.04630\.03160\.06750\.06770\.06770\.06160\.05680\.06320\.06830\.0683N@50\.03150\.02580\.04350\.04470\.04470\.04740\.04600\.04400\.04920\.0492N@100\.03420\.02650\.04820\.04910\.04910\.04890\.04760\.04750\.05270\.0527LT↓\\downarrow0\.14810\.30614\.56000\.44251\.40981\.207540\.70796\.7692PhoneH@50\.05570\.05590\.07690\.08090\.08090\.07400\.07940\.07720\.08340\.0834H@100\.07350\.06900\.10660\.10660\.10610\.09410\.09580\.10570\.10570\.1021N@50\.04680\.04790\.06110\.06220\.06220\.05860\.06420\.06050\.06780\.0678N@100\.05260\.05220\.07060\.07060\.07020\.06510\.06950\.06970\.07380\.0738LT↓\\downarrow0\.19830\.39994\.47470\.52861\.19141\.273627\.31866\.8506BookH@50\.10320\.05560\.13300\.13500\.11810\.11320\.15880\.15880\.1400H@100\.10620\.05850\.14200\.14690\.14690\.12410\.11320\.16580\.16580\.1519N@50\.10250\.04830\.12740\.12740\.12700\.11090\.10950\.14170\.14170\.1310N@100\.10340\.04940\.13030\.13080\.13080\.11280\.10950\.14390\.14390\.1349LT↓\\downarrow0\.39530\.64365\.03000\.80852\.20761\.672050\.64456\.9210MovieH@50\.12100\.11470\.18510\.18510\.17930\.18850\.21260\.26850\.28210\.2821H@100\.16000\.12340\.26750\.26750\.26710\.21310\.24540\.33360\.34810\.3481N@50\.08310\.07000\.11830\.11830\.11540\.15970\.17100\.20400\.21900\.2190N@100\.09580\.07710\.14500\.14500\.14450\.16760\.18160\.22510\.24070\.2407LT↓\\downarrow0\.25650\.43125\.00100\.60451\.98991\.469650\.26187\.0854Table[3](https://arxiv.org/html/2607.28659#S4.T3)summarizes the results and yields the following observations\. Overall, our method achieves the best accuracy\-latency trade\-off: it consistently outperforms MTP and NEZHA on most datasets and metrics, with an average relative accuracy gain of 21\.9% over MTP and 37\.9% over NEZHA, averaged across all metrics and both backbones\. At the same time, it remains competitive with Beam Search in recommendation quality, indicating that the proposed hybrid serial\-parallel decoding largely preserves generation fidelity while substantially lowering decoding cost in practice\. Moreover, although Beam Search attains strong accuracy, it incurs much higher latency; in contrast, our decoding reduces LT by 85\.1% on average compared with Beam Search \(87\.5% on T5 and 82\.7% on Qwen3\-0\.6B\), making it notably more practical for real\-time deployment\. In summary, the proposed hybrid serial\-parallel decoding aligns well with the two\-stage SID structure and offers a more favorable accuracy\-latency balance by significantly reducing inference overhead without sacrificing generation quality overall\.
To further interpret the efficiency results, we analyze per\-position prediction accuracy \(Hit@1\) on the Electronics and Phone domains\. On Electronics, MTP and NEZHA achieve average per\-head accuracies of 0\.0538 and 0\.0499, respectively, compared with 0\.0598 for Ours; on Phone, their averages \(0\.0964 and 0\.0955\) are also lower than Ours \(0\.0969\)\. This gap mainly stems from the fully parallel decoding adopted by NEZHA and MTP, which predicts each position independently and cannot effectively leverage preceding token semantics, whereas Ours preserves partial serial dependencies while enabling parallel prediction, yielding a better balance between per\-position accuracy and inference efficiency\.
### 4\.4\.Ablation Study \(RQ3\)
Figure[2](https://arxiv.org/html/2607.28659#S4.F2)\(a\-d\) reports the ablation results of the proposed cross\-domain hybrid tokenization on the Book\-Movie dataset\. The variants w/o Shared\-Specific Tokenization and w/o Fine\-Grained Specific Tokenization remove Stage 1 and Stage 2 tokenization modules, respectively\. Across both T5 and Qwen3\-0\.6B backbones, the full model \(GenCDSR\) achieves the best performance on H@10 and N@10, with most gains being statistically significant\. Removing Stage 1 causes larger performance drops, highlighting its role in learning transferable shared semantics for cross\-domain modeling, while removing Stage 2 mainly hurts domain\-level precision, indicating the necessity of fine\-grained domain\-specific refinement\. Overall, the two stages are complementary, and their joint design is crucial for robust gains\.
\(e\) Feature fidelity comparison \(higher is better\)\.Figure 2\.Ablation study and feature fidelity comparison \(higher is better for all subfigures\)\. \(a\-d\) Ablation results on the Book\-Movie dataset under T5 and Qwen3\-0\.6B; SST and FGST denote Shared\-Specific Tokenization and Fine\-Grained Specific Tokenization, respectively\. \(e\) Feature fidelity comparison on three cross\-domain datasets\.
### 4\.5\.Feature Fidelity Analysis \(RQ4\)
To illustrate how the cross\-domain hybrid tokenization contributes to improving recommendation performance, we evaluate the feature fidelity\(Fuet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib19)\)of different tokenization schemes on three cross\-domain datasets, which is defined as:
\(25\)Feature Fidelity=max\(0,1−‖\[𝐱A;𝐱B\]−\[𝐱^A;𝐱^B\]‖2‖\[𝐱A;𝐱B\]‖2\)×100%\.\\text\{Feature Fidelity\}=\\max\\\!\\left\(0,\\,1\-\\frac\{\\left\\\|\[\\mathbf\{x\}\_\{A\};\\mathbf\{x\}\_\{B\}\]\-\[\\hat\{\\mathbf\{x\}\}\_\{A\};\\hat\{\\mathbf\{x\}\}\_\{B\}\]\\right\\\|\_\{2\}\}\{\\left\\\|\[\\mathbf\{x\}\_\{A\};\\mathbf\{x\}\_\{B\}\]\\right\\\|\_\{2\}\}\\right\)\\times 100\\%\.where𝐱A,𝐱B\\mathbf\{x\}\_\{A\},\\mathbf\{x\}\_\{B\}are the original semantic embeddings for domainAAandBB,𝐱^A,𝐱^B\\hat\{\\mathbf\{x\}\}\_\{A\},\\hat\{\\mathbf\{x\}\}\_\{B\}are the reconstructed embeddings,\[⋅;⋅\]\[\\cdot;\\cdot\]denotes concatenation operation, and∥⋅∥2\\\|\\cdot\\\|\_\{2\}denotes theℓ2\\ell\_\{2\}norm\.
We compare our method with Sh\-RQ\-VAE \(a single shared RQ\-VAE across domains\) and Sp\-RQ\-VAE \(one RQ\-VAE per domain\)\. As shown in Figure[2](https://arxiv.org/html/2607.28659#S4.F2)\(e\), our method consistently achieves the highest feature fidelity on all datasets, indicating more faithful preservation of semantic features in the learned SIDs\. In contrast, Sh\-RQ\-VAE is more prone to information loss when inter\-domain heterogeneity is large, while Sp\-RQ\-VAE reduces intra\-domain interference but lacks an explicit shared structure across domains\. Overall, these results further confirm that our tokenization preserves information more effectively and produces more stable discretized representations for downstream generative recommendation, while better supporting subsequent cross\-domain knowledge transfer and sequence modeling across heterogeneous domains\.
### 4\.6\.Hyper\-Parameter Analysis \(RQ5\)
We set the code level toL=4L=4and vary the shared\-specific tokenization vs\. fine\-grained specific tokenization allocation in\{0:4,1:3,2:2,3:1,4:0\}\\\{0\\\!:\\\!4,\\,1\\\!:\\\!3,\\,2\\\!:\\\!2,\\,3\\\!:\\\!1,\\,4\\\!:\\\!0\\\}to evaluate cross\-domain hybrid tokenization\. As shown in Figure[3](https://arxiv.org/html/2607.28659#S4.F3),2:22\\\!:\\\!2achieves the best performance, suggesting that balanced capacity is crucial under our two\-stage design across datasets and metrics\. Too little shared capacity \(e\.g\.,0:40\\\!:\\\!4or1:31\\\!:\\\!3\) weakens Stage 1 cross\-domain anchoring and alignment, while too much shared capacity \(e\.g\.,3:13\\\!:\\\!1or4:04\\\!:\\\!0\) suppresses Stage 2 fine\-grained refinement and may increase semantic ambiguity\. Overall,2:22\\\!:\\\!2offers the best trade\-off between shared structure and domain\-specific refinement in practice under all settings\.
Figure 3\.Impact of the codebook allocation between shared\-specific tokenization and fine\-grained specific tokenization \(total = 4\) on HR@10 and NDCG@10 across two datasets\.
## 5\.Related Work
### 5\.1\.Cross\-Domain Sequential Recommendation
CDSR jointly models users’ interaction sequences across multiple domains to capture cross\-domain interest transitions and improve next\-item prediction\(Chenet al\.,[2024b](https://arxiv.org/html/2607.28659#bib.bib35)\)\. Compared with single\-domain sequential recommendation\(Liuet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib30); Wanget al\.,[2023b](https://arxiv.org/html/2607.28659#bib.bib31); Liuet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib32)\)and cross\-domain recommendation\(Liet al\.,[2023b](https://arxiv.org/html/2607.28659#bib.bib2),[2022](https://arxiv.org/html/2607.28659#bib.bib3); Jiaet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib33); Gaoet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib34)\), it faces substantial domain heterogeneity, imbalanced interactions, and complex cross\-domain transitions\(Wanget al\.,[2019](https://arxiv.org/html/2607.28659#bib.bib36)\)\. Most methods follow collaborative filtering to transfer knowledge by modeling user/item relations and cross\-domain dependencies\(Zhuet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib29); Li and Tuzhilin,[2020](https://arxiv.org/html/2607.28659#bib.bib28)\), including graph neural network\-based approaches\(Wuet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib27)\)such as C2DSR\(Caoet al\.,[2022](https://arxiv.org/html/2607.28659#bib.bib38)\)and contrastive variants for stronger representation consistency\(Wanget al\.,[2023a](https://arxiv.org/html/2607.28659#bib.bib25); Xuet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib26)\)\. Recent work like TriCDR\(Maet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib37)\)also models mixed behavior sequences and leverages LLMs like LLM4CDSR\(Liuet al\.,[2025a](https://arxiv.org/html/2607.28659#bib.bib41)\)and LLM\-EDT\(Liuet al\.,[2025b](https://arxiv.org/html/2607.28659#bib.bib24)\); however, these methods still largely rely on collaborative signals and make limited use of item\-level semantics and explicit cross\-domain semantic relationships\.
### 5\.2\.Generative Recommendation
Generative methods have recently gained prevalence across various fields\([Zhanget al\.,](https://arxiv.org/html/2607.28659#bib.bib45); Hanet al\.,[2026](https://arxiv.org/html/2607.28659#bib.bib46); Huet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib44); Chenet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib49)\)\. Generative recommendation \(GR\) reformulates recommendation as autoregressive sequence generation, providing a unified pipeline\(Jiet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib15); Liet al\.,[2023a](https://arxiv.org/html/2607.28659#bib.bib16),[2025](https://arxiv.org/html/2607.28659#bib.bib47); Xionget al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib51); Liet al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib52)\)that typically consists of item tokenization and recommendation generation\. For item tokenization, quantization techniques are commonly employed to map items into discrete semantic identifiers \(SIDs\)\. Residual Quantization, such as RQ\-VAE\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8); Wanget al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib13)\), and Product Quantization \(PQ\)\(Luoet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib11); Zhanget al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib10)\)are widely adopted to produce structured codewords\. For example, TIGER\(Rajputet al\.,[2023](https://arxiv.org/html/2607.28659#bib.bib8)\)utilizes RQ\-VAE to construct hierarchical SIDs, while LC\-Rec\(Zhenget al\.,[2024](https://arxiv.org/html/2607.28659#bib.bib7)\)employs learnable codebooks to enhance semantic alignment\. Building on these advances, GenCDR\(Huet al\.,[2026a](https://arxiv.org/html/2607.28659#bib.bib12)\)and MTCDR\(Jinet al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib9)\)further extend such tokenization paradigms to multi\-domain settings\. Regarding recommendation generation, most methods rely on next\-token prediction with LLM backbones, while increasing attention has been devoted to decoding efficiency\. For example, NEZHA\(Wanget al\.,[2025](https://arxiv.org/html/2607.28659#bib.bib5)\)introduces a multi\-token prediction strategy for SID decoding to alleviate the high inference latency of sequential generation\. However, existing approaches still insufficiently capture cross\-domain collaborative signals during tokenization and suffer from the accuracy–efficiency trade\-off during decoding\.
## 6\.Conclusion
In this paper, we propose GenCDSR, an effective and efficient generative framework for cross\-domain sequential recommendation\. It employs cross\-domain hybrid tokenization with shared\-specific and fine\-grained specific tokenization to capture cross\-domain commonalities and distinctions\. It further introduces serial–parallel decoding to balance recommendation accuracy and inference efficiency\. Experiments on real\-world cross\-domain datasets demonstrate that GenCDSR outperforms SOTA baselines, improving accuracy by 1\.5% while reducing inference latency by 85\.1%\.
## 7\.Acknowledgments
This research was partially supported by National Natural Science Foundation of China \(No\.62502404\), Hong Kong Research Grants Council \(Research Impact Fund No\.R1015\-23, Collaborative Research Fund No\.C1043\-24GF, RGC Research Fellow Scheme No\.RFS2627\-1S03, General Research Fund No\. 11218325, No\. 11212926\), Institute of Digital Medicine of City University of Hong Kong \(No\.9229503\), and Bytedance\.
## References
- J\. Bai, S\. Bai, Y\. Chu, Z\. Cui, K\. Dang, X\. Deng, Y\. Fan, W\. Ge, Y\. Han, F\. Huang,et al\.\(2023\)Qwen technical report\.arXiv preprint arXiv:2309\.16609\.Cited by:[§3\.2\.1](https://arxiv.org/html/2607.28659#S3.SS2.SSS1.p2.2),[§4\.1\.3](https://arxiv.org/html/2607.28659#S4.SS1.SSS3.p1.11)\.
- J\. Cao, X\. Cong, J\. Sheng, T\. Liu, and B\. Wang \(2022\)Contrastive cross\-domain sequential recommendation\.InProceedings of the 31st ACM international conference on information & knowledge management,pp\. 138–147\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1),[1st item](https://arxiv.org/html/2607.28659#S4.I3.i1.p1.1),[§4\.1\.4](https://arxiv.org/html/2607.28659#S4.SS1.SSS4.p1.4),[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- J\. Chen, Y\. Hu, H\. Lu, W\. Wang, M\. Yang, C\. Li, and X\. Hu \(2025\)MGHFT: multi\-granularity hierarchical fusion transformer for cross\-modal sticker emotion recognition\.InProceedings of the 33rd ACM International Conference on Multimedia,pp\. 5794–5803\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- J\. Chen, W\. Wang, Y\. Hu, J\. Chen, H\. Liu, and X\. Hu \(2024a\)Tgca\-pvt: topic\-guided context\-aware pyramid vision transformer for sticker emotion recognition\.InProceedings of the 32nd ACM International Conference on Multimedia,pp\. 9709–9718\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1)\.
- S\. Chen, Z\. Xu, W\. Pan, Q\. Yang, and Z\. Ming \(2024b\)A survey on cross\-domain sequential recommendation\.arXiv preprint arXiv:2401\.04971\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- R\. De La Briandais \(1959\)File searching using variable length keys\.InPapers presented at the the March 3\-5, 1959, western joint computer conference,pp\. 295–298\.Cited by:[§3\.4\.3](https://arxiv.org/html/2607.28659#S3.SS4.SSS3.p1.4)\.
- M\. Freitag and Y\. Al\-Onaizan \(2017\)Beam search strategies for neural machine translation\.InProceedings of the First Workshop on Neural Machine Translation,pp\. 56–60\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p2.1),[1st item](https://arxiv.org/html/2607.28659#S4.I4.i1.p1.1)\.
- K\. Fu, T\. Zhang, S\. Xiao, Z\. Wang, X\. Zhang, C\. Zhang, Y\. Yan, J\. Zheng, Y\. Li, Z\. Chen,et al\.\(2025\)Forge: forming semantic identifiers for generative retrieval in industrial datasets\.arXiv preprint arXiv:2509\.20904\.Cited by:[§4\.5](https://arxiv.org/html/2607.28659#S4.SS5.p1.8)\.
- J\. Gao, X\. Zhao, B\. Chen, F\. Yan, H\. Guo, and R\. Tang \(2023\)AutoTransfer: instance transfer for cross\-domain recommendations\.InProceedings of the 46th international ACM SIGIR conference on research and development in information retrieval,pp\. 1478–1487\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- F\. Gloeckle, B\. Y\. Idrissi, B\. Rozière, D\. Lopez\-Paz, and G\. Synnaeve \(2024\)Better & faster large language models via multi\-token prediction\.arXiv preprint arXiv:2404\.19737\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p2.1),[2nd item](https://arxiv.org/html/2607.28659#S4.I4.i2.p1.1)\.
- X\. Han, Z\. Zhao, W\. Wang, M\. Wang, Z\. Liu, Y\. Chang, and X\. Zhao \(2026\)Data efficient adaptation in large language models via continuous low\-rank fine\-tuning\.Advances in Neural Information Processing Systems38,pp\. 165157–165182\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- B\. Hidasi, A\. Karatzoglou, L\. Baltrunas, and D\. Tikk \(2015\)Session\-based recommendations with recurrent neural networks\.arXiv preprint arXiv:1511\.06939\.Cited by:[1st item](https://arxiv.org/html/2607.28659#S4.I2.i1.p1.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, W\. Chen,et al\.\(2022\)Lora: low\-rank adaptation of large language models\.\.Iclr1\(2\),pp\. 3\.Cited by:[§3\.4\.2](https://arxiv.org/html/2607.28659#S3.SS4.SSS2.p1.4)\.
- P\. Hu, W\. Lu, and J\. Wang \(2026a\)From ids to semantics: a generative framework for cross\-domain recommendation with adaptive semantic tokenization\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 14874–14882\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p2.1),[§3\.1](https://arxiv.org/html/2607.28659#S3.SS1.p1.1),[§3\.2\.1](https://arxiv.org/html/2607.28659#S3.SS2.SSS1.p1.1),[§3\.2\.2](https://arxiv.org/html/2607.28659#S3.SS2.SSS2.p1.3),[§3\.2](https://arxiv.org/html/2607.28659#S3.SS2.p1.7),[§3\.3](https://arxiv.org/html/2607.28659#S3.SS3.p1.1),[§3\.4\.3](https://arxiv.org/html/2607.28659#S3.SS4.SSS3.p1.4),[4th item](https://arxiv.org/html/2607.28659#S4.I3.i4.p1.1),[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- Y\. Hu, J\. Chen, Y\. Wang, Z\. Li, J\. Xiong, P\. Jia, W\. Wang, C\. Li, and X\. Zhao \(2026b\)Emotion and intention guided multi\-modal learning for sticker response selection\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 14883–14891\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1)\.
- Y\. Hu, M\. Tan, C\. Zhang, Z\. Li, X\. Liang, M\. Yang, C\. Li, and X\. Hu \(2024\)APTNESS: incorporating appraisal theory and emotion support strategies for empathetic response generation\.InProceedings of the 33rd ACM International Conference on Information and Knowledge Management,CIKM ’24,New York, NY, USA,pp\. 900–909\.External Links:ISBN 9798400704369,[Link](https://doi.org/10.1145/3627673.3679687),[Document](https://dx.doi.org/10.1145/3627673.3679687)Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- J\. Ji, Z\. Li, S\. Xu, W\. Hua, Y\. Ge, J\. Tan, and Y\. Zhang \(2024\)Genrec: large language model for generative recommendation\.InEuropean Conference on Information Retrieval,pp\. 494–502\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- P\. Jia, Y\. Wang, S\. Lin, X\. Li, X\. Zhao, H\. Guo, and R\. Tang \(2024\)D3: a methodological exploration of domain division, modeling, and balance in multi\-domain recommendations\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 8553–8561\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- J\. Jin, Y\. Zhang, F\. Feng, and X\. He \(2025\)Generative multi\-target cross\-domain recommendation\.arXiv preprint arXiv:2507\.12871\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p2.1),[§3\.1](https://arxiv.org/html/2607.28659#S3.SS1.p1.1),[§3\.2\.2](https://arxiv.org/html/2607.28659#S3.SS2.SSS2.p1.3),[§3\.2](https://arxiv.org/html/2607.28659#S3.SS2.p1.7),[§3\.3](https://arxiv.org/html/2607.28659#S3.SS3.p1.1),[§3\.4\.3](https://arxiv.org/html/2607.28659#S3.SS4.SSS3.p1.4),[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- W\. Kang and J\. McAuley \(2018\)Self\-attentive sequential recommendation\.In2018 IEEE international conference on data mining \(ICDM\),pp\. 197–206\.Cited by:[3rd item](https://arxiv.org/html/2607.28659#S4.I2.i3.p1.1)\.
- D\. Lee, C\. Kim, S\. Kim, M\. Cho, and W\. Han \(2022\)Autoregressive image generation using residual quantization\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 11523–11532\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1),[§2\.2](https://arxiv.org/html/2607.28659#S2.SS2.p1.9)\.
- J\. Li, W\. Zhang, T\. Wang, G\. Xiong, A\. Lu, and G\. Medioni \(2023a\)GPT4Rec: a generative framework for personalized recommendation and user interests interpretation\.arXiv preprint arXiv:2304\.03879\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- P\. Li and A\. Tuzhilin \(2020\)Ddtcdr: deep dual transfer cross domain recommendation\.InProceedings of the 13th international conference on web search and data mining,pp\. 331–339\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- X\. Li, F\. Yan, X\. Zhao, Y\. Wang, B\. Chen, H\. Guo, and R\. Tang \(2023b\)Hamur: hyper adapter for multi\-domain recommendation\.InProceedings of the 32nd ACM International Conference on Information and Knowledge Management,pp\. 1268–1277\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- X\. Li, Z\. Qiu, X\. Zhao, Z\. Wang, Y\. Zhang, C\. Xing, and X\. Wu \(2022\)Gromov\-wasserstein guided representation learning for cross\-domain recommendation\.InProceedings of the 31st ACM International Conference on Information & Knowledge Management,pp\. 1199–1208\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- Z\. Li, B\. Geng, J\. Xiong, Y\. He, Y\. Hu, J\. Chen, D\. Chen, X\. Chang, N\. Wong, L\. Zhang,et al\.\(2025\)Ctr\-sink: attention sink for language models in click\-through rate prediction\.arXiv preprint arXiv:2508\.03668\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- Z\. Li, J\. Xiong, F\. Ye, C\. Zheng, X\. Wu, J\. Lu, Z\. Wan, X\. Liang, C\. Li, Z\. Sun,et al\.\(2024\)Uncertaintyrag: span\-level uncertainty enhanced long\-context modeling for retrieval\-augmented generation\.arXiv preprint arXiv:2410\.02719\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- Q\. Liu, J\. Hu, Y\. Xiao, X\. Zhao, J\. Gao, W\. Wang, Q\. Li, and J\. Tang \(2024\)Multimodal recommender systems: a survey\.ACM Computing Surveys57\(2\),pp\. 1–17\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- Q\. Liu, X\. Zhao, Y\. Wang, Z\. Zhang, H\. Zhong, C\. Chen, X\. Li, W\. Huang, and F\. Tian \(2025a\)Bridge the domains: large language models enhanced cross\-domain sequential recommendation\.InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1582–1592\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1),[3rd item](https://arxiv.org/html/2607.28659#S4.I3.i3.p1.1),[§4\.1\.1](https://arxiv.org/html/2607.28659#S4.SS1.SSS1.p1.1),[§4\.1\.4](https://arxiv.org/html/2607.28659#S4.SS1.SSS4.p1.4),[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- Z\. Liu, J\. Tian, Q\. Cai, X\. Zhao, J\. Gao, S\. Liu, D\. Chen, T\. He, D\. Zheng, P\. Jiang,et al\.\(2023\)Multi\-task recommendations with reinforcement learning\.InProceedings of the ACM web conference 2023,pp\. 1273–1282\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- Z\. Liu, Q\. Liu, W\. Wang, Y\. Wang, T\. Xu, W\. Huang, C\. Chen, P\. Chuan, and X\. Zhao \(2025b\)LLM\-edt: large language model enhanced cross\-domain sequential recommendation with dual\-phase training\.arXiv preprint arXiv:2511\.19931\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1),[§4\.1\.1](https://arxiv.org/html/2607.28659#S4.SS1.SSS1.p1.1),[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- I\. Loshchilov and F\. Hutter \(2019\)Decoupled weight decay regularization\.InInternational Conference on Learning Representations,Cited by:[§4\.1\.3](https://arxiv.org/html/2607.28659#S4.SS1.SSS3.p1.11)\.
- X\. Luo, J\. Cao, T\. Sun, J\. Yu, R\. Huang, W\. Yuan, H\. Lin, Y\. Zheng, S\. Wang, Q\. Hu,et al\.\(2025\)Qarm: quantitative alignment multi\-modal recommendation at kuaishou\.InProceedings of the 34th ACM International Conference on Information and Knowledge Management,pp\. 5915–5922\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- H\. Ma, R\. Xie, L\. Meng, X\. Chen, X\. Zhang, L\. Lin, and J\. Zhou \(2024\)Triple sequence learning for cross\-domain recommendation\.ACM Transactions on Information Systems42\(4\),pp\. 1–29\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1),[2nd item](https://arxiv.org/html/2607.28659#S4.I3.i2.p1.1),[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. Liu \(2020\)Exploring the limits of transfer learning with a unified text\-to\-text transformer\.Journal of machine learning research21\(140\),pp\. 1–67\.Cited by:[§4\.1\.3](https://arxiv.org/html/2607.28659#S4.SS1.SSS3.p1.11)\.
- S\. Rajput, N\. Mehta, A\. Singh, R\. Hulikal Keshavan, T\. Vu, L\. Heldt, L\. Hong, Y\. Tay, V\. Tran, J\. Samost,et al\.\(2023\)Recommender systems with generative retrieval\.Advances in Neural Information Processing Systems36,pp\. 10299–10315\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1),[§1](https://arxiv.org/html/2607.28659#S1.p2.1),[§3\.2\.1](https://arxiv.org/html/2607.28659#S3.SS2.SSS1.p1.1),[§3\.4\.3](https://arxiv.org/html/2607.28659#S3.SS4.SSS3.p1.4),[4th item](https://arxiv.org/html/2607.28659#S4.I2.i4.p1.1),[§4\.1\.1](https://arxiv.org/html/2607.28659#S4.SS1.SSS1.p1.1),[§4\.1\.3](https://arxiv.org/html/2607.28659#S4.SS1.SSS3.p1.11),[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- F\. Sun, J\. Liu, J\. Wu, C\. Pei, X\. Lin, W\. Ou, and P\. Jiang \(2019\)BERT4Rec: sequential recommendation with bidirectional encoder representations from transformer\.InProceedings of the 28th ACM international conference on information and knowledge management,pp\. 1441–1450\.Cited by:[2nd item](https://arxiv.org/html/2607.28659#S4.I2.i2.p1.1)\.
- H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar,et al\.\(2023\)Llama: open and efficient foundation language models\.arXiv preprint arXiv:2302\.13971\.Cited by:[§3\.2\.1](https://arxiv.org/html/2607.28659#S3.SS2.SSS1.p2.2),[§4\.1\.3](https://arxiv.org/html/2607.28659#S4.SS1.SSS3.p1.11)\.
- S\. Wang, L\. Hu, Y\. Wang, L\. Cao, Q\. Z\. Sheng, and M\. Orgun \(2019\)Sequential recommender systems: challenges, progress and prospects\.arXiv preprint arXiv:2001\.04830\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- W\. Wang, H\. Bao, X\. Lin, J\. Zhang, Y\. Li, F\. Feng, S\. Ng, and T\. Chua \(2024\)Learnable item tokenization for generative recommendation\.InProceedings of the 33rd ACM International Conference on Information and Knowledge Management,pp\. 2400–2409\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1),[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- X\. Wang, H\. Yue, Z\. Wang, L\. Xu, and J\. Zhang \(2023a\)Unbiased and robust: external attention\-enhanced graph contrastive learning for cross\-domain sequential recommendation\.In2023 IEEE International Conference on Data Mining Workshops \(ICDMW\),pp\. 1526–1534\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- Y\. Wang, Z\. Du, X\. Zhao, B\. Chen, H\. Guo, R\. Tang, and Z\. Dong \(2023b\)Single\-shot feature selection for multi\-task recommendations\.InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 341–351\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- Y\. Wang, S\. Zhou, J\. Lu, Z\. Liu, L\. Liu, M\. Wang, W\. Zhang, F\. Li, W\. Su, P\. Wang,et al\.\(2025\)NEZHA: a zero\-sacrifice and hyperspeed decoding architecture for generative recommendations\.arXiv preprint arXiv:2511\.18793\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p2.1),[3rd item](https://arxiv.org/html/2607.28659#S4.I4.i3.p1.1),[§4\.1\.4](https://arxiv.org/html/2607.28659#S4.SS1.SSS4.p1.4),[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- S\. Wu, F\. Sun, W\. Zhang, X\. Xie, and B\. Cui \(2022\)Graph neural networks in recommender systems: a survey\.ACM computing surveys55\(5\),pp\. 1–37\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- J\. Xiong, Z\. Li, C\. Zheng, Z\. Guo, Y\. Yin, E\. Xie, Z\. Yang, Q\. Cao, H\. Wang, X\. Han,et al\.\(2024\)Dq\-lore: dual queries with low rank approximation re\-ranking for in\-context learning\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 41179–41203\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- W\. Xu, Q\. Wu, R\. Wang, M\. Ha, Q\. Ma, L\. Chen, B\. Han, and J\. Yan \(2024\)Rethinking cross\-domain sequential recommendation under open\-world assumptions\.InProceedings of the ACM Web Conference 2024,pp\. 3173–3184\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p1.1)\.
- Z\. Xu, S\. Chen, W\. Pan, and Z\. Ming \(2025\)A multi\-view graph contrastive learning framework for cross\-domain sequential recommendation\.ACM Transactions on Recommender Systems3\(4\),pp\. 1–28\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§4\.1\.3](https://arxiv.org/html/2607.28659#S4.SS1.SSS3.p1.11)\.
- \[49\]C\. Zhang, Y\. Wang, D\. Xu, H\. Zhang, Y\. Lyu, Y\. Chen, S\. Liu, T\. Xu, X\. Zhao, Y\. Gao,et al\.Tearag: a token\-efficient agentic retrieval\-augmented generation framework\.ACM Transactions on Information Systems\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- K\. Zhang, J\. Jin, Y\. Qin, R\. Su, J\. Lin, Y\. Yu, and W\. Zhang \(2024\)Learning id\-free item representation with token crossing for multimodal recommendation\.arXiv preprint arXiv:2410\.19276\.Cited by:[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- B\. Zheng, Y\. Hou, H\. Lu, Y\. Chen, W\. X\. Zhao, M\. Chen, and J\. Wen \(2024\)Adapting large language models by integrating collaborative semantics for recommendation\.In2024 IEEE 40th International Conference on Data Engineering \(ICDE\),pp\. 1435–1448\.Cited by:[§1](https://arxiv.org/html/2607.28659#S1.p2.1),[§4\.1\.1](https://arxiv.org/html/2607.28659#S4.SS1.SSS1.p1.1),[§5\.2](https://arxiv.org/html/2607.28659#S5.SS2.p1.1)\.
- Y\. Zhu, Z\. Tang, Y\. Liu, F\. Zhuang, R\. Xie, X\. Zhang, L\. Lin, and Q\. He \(2022\)Personalized transfer of user preferences for cross\-domain recommendation\.InProceedings of the fifteenth ACM international conference on web search and data mining,pp\. 1507–1515\.Cited by:[§5\.1](https://arxiv.org/html/2607.28659#S5.SS1.p1.1)\.Similar Articles
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
This paper proposes DASO, a tree-aware post-training method for generative recommendation that addresses difficulty mismatch in GRPO by profiling rollout groups and reallocating based on prefix-match depth, improving performance on public benchmarks.
Understanding Semantic IDs: From Item Representation to Item Selection in Generative Recommendation
This paper systematically investigates Semantic IDs (SIDs) in generative recommendation, finding that while SIDs preserve coarse item organization, they lose fine local structure from the encoder. The authors propose Item-Supported Decoding (ISD), a lightweight inference-time method that improves NDCG@10 by up to 31.2% without additional parameters or retraining.
Dual-Interest Sequential Product Recommendation With Multi-Granular SSM
DSRec is a novel dual-interest sequential recommendation model using multi-granular State Space Models to capture long-term and short-term user interests, addressing item polysemy and achieving superior performance on benchmarks.
X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding
This paper introduces X-CoSD, a communication-efficient cross-vocabulary collaborative speculative decoding framework that optimizes distributed LLM inference by splitting residual resampling to reduce overhead while preserving server LLM quality.
CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation
CARD proposes a hierarchical framework for personalized text generation that clusters users and uses reward-guided decoding, demonstrating improved quality and efficiency on LaMP benchmarks.