OrDA: Orthogonal Disentanglement of Access Habits Framework for Homepage Marketing Block Recommendations
Summary
OrDA is a framework that disentangles user access habits from genuine content interests in homepage marketing block recommendations using orthogonal regularization and causal intervention, achieving a 5.64% UCTR improvement on Zhima's homepage.
View Cached Full Text
Cached at: 07/16/26, 04:22 AM
# OrDA: Orthogonal Disentanglement of Access Habits Framework for Homepage Marketing Block Recommendations
Source: [https://arxiv.org/html/2607.13420](https://arxiv.org/html/2607.13420)
###### Abstract\.
Clicks on homepage marketing blocks are driven by a dual\-mechanism of content interest and access habits\. However, habitual clicks often create ”Pseudo\-Positives” in marketing slots, where position advantage masks mediocre content quality, leading to biased recommendation ecosystems\.
We propose a framework calledOrthogonalDisentanglement ofAccess habits \(OrDA\) to purify interest signals\. OrDA utilizes a dual\-tower structure with a gated allocation layer to adaptively route features and minimize interference\. To ensure rigorous separation, we employ orthogonal regularization to constrain the latent interest and habit manifolds to be geometrically perpendicular\. OrDA performs causal intervention \(dodo\-calculus\(Tucci,[2013](https://arxiv.org/html/2607.13420#bib.bib11)\)\) during inference to rank items solely by purified interest scores\. Empirical online evaluations on large\-scale datasets demonstrate that OrDA effectively eliminates access\-habit bias, outperforming state\-of\-the\-art methods in predictive accuracy\. Online A/B test shows 5\.64% user click\-through rates \(UCTR\) improvement on the Zhima homepage marketing block, Zhima’s rent\-floor recommendation\.
Disentangled Learning, Causal Intervention, Recommender System
††ccs:Information systems Retrieval models and ranking## 1\.Introduction
The homepage of an internet application can be summarized as having the core function of: precise navigation and ecosystem distribution from massive content to user interests\. Homepage marketing blocks \(shown in Fig\.[1](https://arxiv.org/html/2607.13420#S1.F1)\) play the role of guiding the vast user traffic to different vertical sub\-scenarios\(Zheng et al\.,[2025](https://arxiv.org/html/2607.13420#bib.bib16)\)\. These blocks often suffer from a subtle yet pervasive form of confounding bias: the user\-channel access habitual dependency\. Specifically, in prominent marketing blocks—frequently accessed by ”heavy users” as part of their routine digital journey—observed click data is often contaminated by ”Pseudo\-Positives\.” These are interactions driven not by a genuine alignment between user interests and material content, but rather by the user’s habitual inertia within their primary access channels\.
Figure 1\.A demo of Zhima homepage marketing block recommendationA demo of Zhima homepage marketing block recommendationStandard models, which treat all clicks as equivalent signals of preference, inadvertently capture this ”habitual dividend,” leading to a distorted representation of user interests\. From a causal perspective, the user’s channel affinity acts as a confounder, inflating the perceived utility of certain contents simply because they occupy the user’s ”path of least resistance\.” Consequently, such models often perform poorly with cold\-start users and fail to generalize to broader, non\-habitual contexts, as they struggle to distinguish between user\-channel access habits and user\-content interests\.
In order to address this problem, we propose aOrthogonalDisentanglement ofAccess habits \(OrDA\) framework for homepage marketing block recommendations\. The main contributions of this work are summarized as follows:
- •Problem Formulation of Access Habit Bias: We provide a formal causal analysis of access habit bias prevalent in homepage marketing blocks\. Unlike conventional exposure bias, we identify and define the phenomenon of habit\-induced pseudo\-positives, where users’ navigational inertia at primary entry points masks their genuine content preferences, providing a new perspective on bias decomposition\.
- •Orthogonal Disentanglement of Access Habits\(OrDA\): We propose OrDA framework, a novel multi\-tower architecture designed to decouple user intent from habitual access patterns\. Our model can effectively isolate the spurious correlations between user\-content interests and user\-channel access habits via orthogonal regularization anddodo\-calculus\-based intervention strategy\.
- •Superior Performance on Industrial Datasets: OrDA achieves state\-of\-the\-art performance on Zhima homepage datasets during extensive offline experiments\. Online A/B test shows 5\.64% UCTR gains for Zhima’s rent\-floor recommendations\.
## 2\.Related Work
### 2\.1\.Debiasing in Recommendation Systems
The identification and elimination of biases, such as position bias\(Guo et al\.,[2019](https://arxiv.org/html/2607.13420#bib.bib4)\)and popular bias\(Zheng et al\.,[2021a](https://arxiv.org/html/2607.13420#bib.bib17)\), are critical for building robust recommendation systems\(Chen et al\.,[2023](https://arxiv.org/html/2607.13420#bib.bib3)\)\. Traditional approaches mainly include propensity\_based method\(Schnabel et al\.,[2016](https://arxiv.org/html/2607.13420#bib.bib9); Zheng et al\.,[2025](https://arxiv.org/html/2607.13420#bib.bib16)\), which re\-weights samples based on their exposure probability, dual learning\(Wang et al\.,[2019](https://arxiv.org/html/2607.13420#bib.bib13); Zhang et al\.,[2020](https://arxiv.org/html/2607.13420#bib.bib15); Zhu et al\.,[2023](https://arxiv.org/html/2607.13420#bib.bib20)\), which learns an auxiliary model to estimate biases, disentangled representation learning\(Zheng et al\.,[2021a](https://arxiv.org/html/2607.13420#bib.bib17),[b](https://arxiv.org/html/2607.13420#bib.bib18)\), which employs contrastive learning to separate interest from conformity, and pseudo\-label method\(Huang et al\.,[2024](https://arxiv.org/html/2607.13420#bib.bib5)\), which aims to complete the counterfactual space by assigning estimated labels to unobserved entries\. Our work differs from these by specifically focusing on access habits in homepage scenarios and performing debiasing through a geometric orthogonal constraint rather than simple score adjustment to enforce a clear causal boundary and ensure a more robust representation of user intrinsic interests\.
### 2\.2\.Synergizing Multi\-tower Architectures with Causal Inference
Multi\-tower models \(e\.g\., ESMM\(Ma et al\.,[2018a](https://arxiv.org/html/2607.13420#bib.bib8)\), MMOE\(Ma et al\.,[2018b](https://arxiv.org/html/2607.13420#bib.bib7)\), PLE\(Tang et al\.,[2020](https://arxiv.org/html/2607.13420#bib.bib10)\)\) are widely used in industrial recommendation to handle multiple tasks\. In reality, these architectures provide a structural foundation for task or feature decomposition\(Zhang et al\.,[2020](https://arxiv.org/html/2607.13420#bib.bib15); Zhu et al\.,[2023](https://arxiv.org/html/2607.13420#bib.bib20); Zheng et al\.,[2025](https://arxiv.org/html/2607.13420#bib.bib16); Wang et al\.,[2022](https://arxiv.org/html/2607.13420#bib.bib12)\)\. From a causal perspective, a significant limitation of traditional multi\-tower models is feature interference: since towers often share a common embedding bottom, the gradients from a biased signal \(e\.g\., access habit\) can easily contaminate the representations of other towers \(e\.g\., content interest\)\. Our work advances this structural decomposition by re\-interpreting the multi\-tower topology as a physical realization of a Structural Causal Model \(SCM\)\.
## 3\.Methods
### 3\.1\.Problem Definition
To formally characterize the mechanism of habit\-induced pseudo\-positives and justify the architecture of OrDA, we employ a causal graph to represent the data generation process in homepage marketing blocks\. As illustrated in Fig\.[2](https://arxiv.org/html/2607.13420#S3.F2), we define the following variables:
- •XX: The user\-content pair\.
- •HH: The latent navigational inertia oraccess\-habitbias\.
- •II: The latentinterestof a user towards the content\.
- •YY: The observed click outcome\.
- •uu: User features such as demographics preferences\.
- •cc: Content features such as category and quality\.
- •u2cu2c: Statistical features of a user towards the content\.
Figure 2\.Causal graphs depicting training of ESMM\-based Approach, Propensity\-based Approach, OrDA\. Hollow circles indicate latent variables, and shaded circles indicate observed variables\.Causal graphs depicting training of ESMM\-based Approach, Propensity\-based Approach, OrDA\. Hollow circles indicate latent variables, shaded circles indicate observed variables\.We define user\-channel access habit and user\-content interest as the main factors governing the click behavior in homepage marketing blocks\. The structural equations of these latent spaces are:
\(1\)\{I=fI\(u,c,u2c\),H=fH\(u\)\\begin\{cases\}I=f\_\{I\}\(u,c,u2c\),\\\\ H=f\_\{H\}\(u\)\\end\{cases\}We aim to get the pure user\-content interest from the confusion\. Fig\.[2](https://arxiv.org/html/2607.13420#S3.F2)\(a\) shows ESMM\-based approach\(Ma et al\.,[2018a](https://arxiv.org/html/2607.13420#bib.bib8)\)uses two towers to predict access habit probabilityP\(h^\)=σ\(H\)P\(\\hat\{h\}\)=\\sigma\(H\)and user\-content interest probabilityP\(i^\)=σ\(I\)P\(\\hat\{i\}\)=\\sigma\(I\), and multiplies them to predict habit\-induced user interest probabilityP\(y^\)=σ\(Y\)P\(\\hat\{y\}\)=\\sigma\(Y\):
\(2\)P\(y^\)=P\(h^\)∗P\(i^\)P\(\\hat\{y\}\)=P\(\\hat\{h\}\)\*P\(\\hat\{i\}\)This approach does not enforce a clear causal boundary or relation between two towers, and thus it introduces false independent prior\.
Fig\.[2](https://arxiv.org/html/2607.13420#S3.F2)\(b\) shows propensity\-based approach\(Wang et al\.,[2022](https://arxiv.org/html/2607.13420#bib.bib12); Zhu et al\.,[2023](https://arxiv.org/html/2607.13420#bib.bib20); Zheng et al\.,[2025](https://arxiv.org/html/2607.13420#bib.bib16); Zhang et al\.,[2020](https://arxiv.org/html/2607.13420#bib.bib15)\)integrates counterfactual reasoning into multi\-task learning to mitigate selection bias via inverse propensity weighting technique anddodo\-calculus\(Tucci,[2013](https://arxiv.org/html/2607.13420#bib.bib11)\)\. Habit\-induced user interest probability is defined as:
\(3\)P\(y^\)=P\(h^\)∗P\(i^\|do\(h^=1\)\)P\(\\hat\{y\}\)=P\(\\hat\{h\}\)\*P\(\\hat\{i\}\|do\(\\hat\{h\}=1\)\)During the training stage, this approach adds another tower to predict habit\-induced interest probabilityP\(i^\|do\(h^=1\)\)=P\(y^\)P\(h^\)P\(\\hat\{i\}\|do\(\\hat\{h\}=1\)\)=\\frac\{P\(\\hat\{y\}\)\}\{P\(\\hat\{h\}\)\}\. While the division form of the expression relying on numerical re\-weighting may suffer from high variance and instability\.
OrDA \(shown in Fig\.[2](https://arxiv.org/html/2607.13420#S3.F2)\(c\)\) extends the causal philosophy to therepresentation levelas following equation:
where⊕\\oplusdenotes the joint logit fusion, representing the additive interaction between the user\-content intrinsic interest and their access habit in the latent space\. Our model posits thatIIandHHare parallel causal drivers and a click occurs when the content satisfies the user’s interestorthe block triggers their access habits\. Thus we enforceLatent Orthogonal Constraintto ensure that the interest tower is inherently invariant to habitual confounders\.
### 3\.2\.Training and Counterfactual Inference
During the training stage, the loss functions of bias CTR and Habit are defined as:
\(5\)\{ℒbCTR=ℒBCE\(y,σ\(Y\)\)=ℒBCE\(y,σ\(H⊕I\)\)ℒHabit=ℒBCE\(y,σ\(H\)\)\\begin\{cases\}\\mathcal\{L\}\_\{bCTR\}=\\mathcal\{L\}\_\{BCE\}\(y,\\sigma\(Y\)\)=\\mathcal\{L\}\_\{BCE\}\(y,\\sigma\(H\\oplus I\)\)\\\\ \\mathcal\{L\}\_\{Habit\}=\\mathcal\{L\}\_\{BCE\}\(y,\\sigma\(H\)\)\\end\{cases\}whereℒBCE\\mathcal\{L\}\_\{BCE\}is Binary Cross\-Entropy loss:
\(6\)ℒBCE=−1N∑i=1N\[yilog\(pi\)\+\(1−yi\)log\(1−pi\)\]\\mathcal\{L\}\_\{BCE\}=\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\[y\_\{i\}\\log\(p\_\{i\}\)\+\(1\-y\_\{i\}\)\\log\(1\-p\_\{i\}\)\\right\]whereyyis the ground truth andppis the probability of positive class\. To ensure rigorous separation, the orthogonal cosine regularization\(Luo et al\.,[2017](https://arxiv.org/html/2607.13420#bib.bib6); Bousmalis et al\.,[2016](https://arxiv.org/html/2607.13420#bib.bib2)\)is employed to constrain the latent manifolds, forcing the interest and habit vectors to be geometrically perpendicular:
\(7\)ℒOrth=\(𝐯int⋅𝐯hab‖𝐯int‖⋅‖𝐯hab‖\)2\\mathcal\{L\}\_\{Orth\}=\(\\frac\{\\mathbf\{v\}\_\{int\}\\cdot\\mathbf\{v\}\_\{hab\}\}\{\\\|\\mathbf\{v\}\_\{int\}\\\|\\cdot\\\|\\mathbf\{v\}\_\{hab\}\\\|\}\)^\{2\}where𝐯int\\mathbf\{v\}\_\{int\}is the interest tower vector and𝐯hab\\mathbf\{v\}\_\{hab\}is the habit tower vector\. Total loss function is as follow:
\(8\)ℒtotal=ℒbCTR\+α∗ℒHabit\+β∗ℒOrth\\mathcal\{L\}\_\{total\}=\\mathcal\{L\}\_\{bCTR\}\+\\alpha\*\\mathcal\{L\}\_\{Habit\}\+\\beta\*\\mathcal\{L\}\_\{Orth\}whereα\\alphaandβ\\betaare hyperparameters that control the weights of the habit task and the orthogonal task\.
During the online inference stage, a causal intervention on the habit variable is performed to eliminate the pseudo\-positive effect\. Formally, we compute the purified user\-content interest scoreSSby applyingdodo\-operatordo\(H=0\)do\(H=0\):
\(9\)S=𝔼\[Y∣I,do\(H=0\)\]=σ\(I\)S=\\mathbb\{E\}\[Y\\mid I,do\(H=0\)\]=\\sigma\(I\)
### 3\.3\.Model Architecture
As illustrated in Fig\.[3](https://arxiv.org/html/2607.13420#S3.F3), the model architecture consists of three core components: the Gated Allocation Layer \(GAL\), the Backbone Model \(BM\), and the Causal Fusion Layer \(CFL\)\.
Figure 3\.Model architecture of OrDA\. The entire components are used for the training stage, while components within the blue box are used for the online inference stage\.Model architecture of OrDA#### 3\.3\.1\.Gated Allocation Layer \(GAL\)
A potential pitfall in disentangled learning is information suppression, where the orthogonality constraint inadvertently erases generic user signals from shared embeddings\. To address this, we introduce the Gated Allocation Layer \(GAL\) as a learnable router, avoiding all features to flow into both towers\. For an input embedding𝐞i∈ℝd\\mathbf\{e\}\_\{i\}\\in\\mathbb\{R\}^\{d\}, GAL computes a routing scoregatei∈\[0,1\]gate\_\{i\}\\in\[0,1\]using a sigmoid\-activated gating network:
\(10\)gatei=σ\(𝐰g⊤𝐞i\+bg\)gate\_\{i\}=\\sigma\(\\mathbf\{w\}\_\{g\}^\{\\top\}\\mathbf\{e\}\_\{i\}\+b\_\{g\}\)The features are adaptively allocated into two distinct latent streams:
\(11\)\{𝐞u,hab=gateA⋅𝐞u,𝐞u,int=gateB⋅𝐞u\\begin\{cases\}\\mathbf\{e\}\_\{u,hab\}=\{gate\_\{A\}\\cdot\\mathbf\{e\}\_\{u\}\},\\\\ \\mathbf\{e\}\_\{u,int\}=\{gate\_\{B\}\\cdot\\mathbf\{e\}\_\{u\}\}\\end\{cases\}wheregateAgate\_\{A\}is used for the habit tower andgateBgate\_\{B\}is used for the interest tower\. Note that GAL is only used in user features because we assume user\-channel access habits are only affected by user features \(mentioned in Eq\.[1](https://arxiv.org/html/2607.13420#S3.E1)\)\. This mechanism ensures that the orthogonality constraint only suppresses the redundant interference signals while allowing multi\-faceted features \(e\.g\., user demographics\) to coexist in both manifolds through distinct projections\.
#### 3\.3\.2\.Backbone Model \(BM\)
The backbone model is an orthogonal dual\-tower\. The two feature streams are fed into the Habit Towerℱhab\\mathcal\{F\}\_\{hab\}and the Interest Towerℱint\\mathcal\{F\}\_\{int\}respectively\. Each tower is designed to project the filtered features into their respective latent manifolds:
\(12\)\{𝐯hab=ℱhab\(𝐞u,hab\)𝐯int=ℱint\(concat\(𝐞u,int,𝐞c,𝐞u2c\)\),\\begin\{cases\}\\mathbf\{v\}\_\{hab\}=\\mathcal\{F\}\_\{hab\}\(\\mathbf\{e\}\_\{u,hab\}\)\\\\ \\mathbf\{v\}\_\{int\}=\\mathcal\{F\}\_\{int\}\(\\text\{concat\}\(\\mathbf\{e\}\_\{u,int\},\\mathbf\{e\}\_\{c\},\\mathbf\{e\}\_\{u2c\}\)\),\\end\{cases\}where𝐞u,hab\\mathbf\{e\}\_\{u,hab\}and𝐞u,int\\mathbf\{e\}\_\{u,int\}is shown in Eq\.[11](https://arxiv.org/html/2607.13420#S3.E11)\.𝐞c\\mathbf\{e\}\_\{c\}and𝐞u2c\\mathbf\{e\}\_\{u2c\}are the embedding of the content features and the user2content features respectively\.ℱhab\\mathcal\{F\}\_\{hab\}only processes user features\.
In order to the disentanglement theorized in our causal graphs, we enforce a orthogonal constraint \(Eq\.[7](https://arxiv.org/html/2607.13420#S3.E7)\) to ensure that𝐯int⟂𝐯hab\\mathbf\{v\}\_\{int\}\\perp\\mathbf\{v\}\_\{hab\}in the vector space, effectively ”pushing” the content\-driven and habit\-driven signals into non\-overlapping dimensions\.
#### 3\.3\.3\.Causal Fusion Layer \(CFL\)
The final stage is the structural combination of the two towers\. Following our additive formulation of causal graphs \(Eq\.[4](https://arxiv.org/html/2607.13420#S3.E4)\), the logits of both towers are summed via the⊕\\oplusoperator:
\(13\)logitall=logithab⊕logitint=MLPCFL\(𝐯hab\)\+MLPCFL\(𝐯int\)logit\_\{all\}=logit\_\{hab\}\\oplus logit\_\{int\}=\\text\{MLP\}\_\{CFL\}\(\\mathbf\{v\}\_\{hab\}\)\+\\text\{MLP\}\_\{CFL\}\(\\mathbf\{v\}\_\{int\}\)where the operator⊕\\oplusdenotes the joint logit fusion, representing the additive interaction between the user’s intrinsic interest and their navigational habit\. Our⊕\\oplusoperator explicitly acknowledges that the click is a superposition of two independent semantic forces\.logitintlogit\_\{int\}captures the ”pull” from content relevance, whilelogithablogit\_\{hab\}captures the ”push” from the user’s ingrained access patterns on the homepage\. Note thatMLPCFL\(⋅\)\\text\{MLP\}\_\{CFL\}\(\\cdot\)in Eq\.[13](https://arxiv.org/html/2607.13420#S3.E13)is a Multi\-Layer Perceptron \(MLP\) as the mapping function from the vector space to the logit space\. To ensure the isolation and additivity between the two in the logit space, we uselinearlinearas the activation function of the causal fusion layer\.
The combined logit is then passed through a sigmoid function to produce the final click probabilityP\(y^\)P\(\\hat\{y\}\):
\(14\)P\(y^\)=σ\(logitall\)P\(\\hat\{y\}\)=\\sigma\(logit\_\{all\}\)During the training stage, bothlogitintlogit\_\{int\}andlogithablogit\_\{hab\}are optimized simultaneously to fit the observed click labels\. During the inference stage, the⊕\\oplusoperator allows for the seamless removal of the habit component by settinglogithab=0logit\_\{hab\}=0\(Eq\.[9](https://arxiv.org/html/2607.13420#S3.E9)\), thereby yielding a purified interest score\.
## 4\.Experiments
### 4\.1\.Experimental Setup
#### 4\.1\.1\.Datasets
Production Dataset: A large\-scale dataset from Zhima homepage marketing block \(15\-day collection\), containing over 30\.2 million users, 108\.9 million training samples and 7\.9 million evaluation samples\.
#### 4\.1\.2\.Baselines
We compared with the standard modelBASE\(w/o debiasing components\) and multiple SOTA debiasing methods: multi\-task modelESMM\(Ma et al\.,[2018a](https://arxiv.org/html/2607.13420#bib.bib8)\), propensity\-based modelsUSD\(Zheng et al\.,[2025](https://arxiv.org/html/2607.13420#bib.bib16)\)andMulti\-IPW\(Zhang et al\.,[2020](https://arxiv.org/html/2607.13420#bib.bib15)\), factorization modelPAL\(Guo et al\.,[2019](https://arxiv.org/html/2607.13420#bib.bib4)\)\.
#### 4\.1\.3\.Implementation Details
We use a Multi\-Layer Perceptron \(MLP\) asℱhab\\mathcal\{F\}\_\{hab\}and a MaskNet\(Wang et al\.,[2021](https://arxiv.org/html/2607.13420#bib.bib14)\)asℱint\\mathcal\{F\}\_\{int\}\(Eq\.[12](https://arxiv.org/html/2607.13420#S3.E12)\), Adam optimizer \(lr=0\.0001, batch size=2048\), loss weightsα\\alpha=β\\beta= 1 \(Equation[8](https://arxiv.org/html/2607.13420#S3.E8)\)\. Due to class imbalance in invalid exposure, we use GAUC\(Zhou et al\.,[2018](https://arxiv.org/html/2607.13420#bib.bib19)\)as primary metric:GAUC=∑u∈𝒰wu×AUCu∑u∈𝒰wu\\text\{GAUC\}=\\frac\{\\sum\_\{u\\in\\mathcal\{U\}\}w\_\{u\}\\times\\text\{AUC\}\_\{u\}\}\{\\sum\_\{u\\in\\mathcal\{U\}\}w\_\{u\}\}, wherewuw\_\{u\}= 1, all user group, cold\-start user group \(click¡=1\) and active user group \(click¿1\) correspond toGAUCallGAUC\_\{all\},GAUCcoldGAUC\_\{cold\},GAUCactiveGAUC\_\{active\}\.
### 4\.2\.Overall Performance Comparison and Ablation Study
As shown in Table[1](https://arxiv.org/html/2607.13420#S4.T1),BASEmodel exhibits higherGAUCactiveGAUC\_\{active\}and lowerGAUCcoldGAUC\_\{cold\}because the model tends to ”memorize” the active users’ long\-term, predictable access habits as strong positive signals \(e\.g\., the users always clicking the first slot upon entering the app\) but does not understand the user\-content interests\.Debiasingmodels get higherGAUCcoldGAUC\_\{cold\}performances since they tend to strip away the ”habitual noise” andOrDAachieves the best performance both onGAUCcoldGAUC\_\{cold\}andGAUCallGAUC\_\{all\}\.
Table 1\.Performance comparison on production datasets\.We compare OrDA with several variants as shown in Table[2](https://arxiv.org/html/2607.13420#S4.T2)to investigate the impact of the core components\.w/o GAL\(replace GAL with a standard shared\-bottom\) has a performance decay, which we attribute to the information evaporation effect caused by strict orthogonality, forcing all user feature embeddings to zero or random noise\. Other ablated variants:w/o HL\(remove the habit auxiliary loss function,α=0\\alpha=0\),w/o OL\(remove the cosine regularization,β=0\\beta=0\),w/o doC\(use the rawlogitint⊕logithablogit\_\{int\}\\oplus logit\_\{hab\}for ranking\) face decreases of 1\.06%, 1\.54%, and 2\.72% inGAUCallGAUC\_\{all\}, respectively\.
Table 2\.Ablation Study of OrDA
### 4\.3\.Latent Space Visualization
To quantitatively and visually assess the degree of disentanglement, we compute the cosine similarity matrix between the interest vectors𝐯int\\mathbf\{v\}\_\{int\}and the habit vectors𝐯habit\\mathbf\{v\}\_\{habit\}across a random subset of the evaluation set\. As illustrated in Fig\.[4](https://arxiv.org/html/2607.13420#S4.F4), the similarity heatmap exhibits a near\-zero distribution across both the diagonal \(intra\-sample\) and off\-diagonal \(inter\-sample\) elements, demonstrating the two vector spaces are constrained to be mutually orthogonal, thereby confirming that the disentanglement is both sample\-wise consistent and globally robust\.
Figure 4\.Visualizing the orthogonality between interest and habit vectors\. The color of the heatmap is predominantly white/light, indicating that the cosine similarity values are concentrated around 0\.
### 4\.4\.Online A/B Test
We run a week\-long online A/B test on Zhima homepage marketing blocks: Zhima’s rent\-floor, whereOrDAoutperforms the baseline model \(BASE, see Sec\.4\.1\.2\) with UCTR increased 5\.64%\. This substantial gain demonstrates the industrial validity of OrDA\.
## 5\.Conclusion
We propose the OrDA framework designed to disentangle user\-content intrinsic interest from user\-channel access habit on homepage recommendations\. Extensive experiments on large\-scale industrial datasets and the online A/B test demonstrate its effectiveness, effectively recovering purified interest scores\. Notably, OrDA has been fully deployed in Zhima homepage marketing blocks: Zhima’s rent\-floor recommendation\.
## References
- \(1\)
- Bousmalis et al\.\(2016\)Konstantinos Bousmalis, George Trigeorgis, Nathan Silberman, Dilip Krishnan, and Dumitru Erhan\. 2016\.Domain separation networks\. In*Proceedings of the 30th International Conference on Neural Information Processing Systems**\(NIPS’16\)*\. 343 – 351\.
- Chen et al\.\(2023\)Jiawei Chen, Hande Dong, Xiang Wang, Fuli Feng, Meng Wang, and Xiangnan He\. 2023\.Bias and Debias in Recommender System: A Survey and Future Directions\.*ACM Transactions on Information Systems*41, 67 \(Feb\. 2023\), 1–39\.[doi:10\.1145/3564284](https://doi.org/10.1145/3564284)
- Guo et al\.\(2019\)Huifeng Guo, Jinkai Yu, Qing Liu, Ruiming Tang, and Yuzhou Zhang\. 2019\.PAL: a position\-bias aware learning framework for CTR prediction in live recommender systems\. In*Proceedings of the 13th ACM Conference on Recommender Systems**\(RecSys ’19\)*\. 452–456\.
- Huang et al\.\(2024\)Jiahui Huang, Lan Zhang, Junhao Wang, Shanyang Jiang, Dongbo Huang, Cheng Ding, and Lan Xu\. 2024\.Utilizing Non\-click Samples via Semi\-supervised Learning for Conversion Rate Prediction\. In*Proceedings of the 18th ACM Conference on Recommender Systems**\(RecSys ’24\)*\. 350–359\.
- Luo et al\.\(2017\)Chunjie Luo, Jianfeng Zhan, Lei Wang, and Qiang Yang\. 2017\.*Cosine Normalization: Using Cosine Similarity Instead of Dot Product in Neural Networks*\.arXiv:1702\.05870 \[cs\.LG\]
- Ma et al\.\(2018b\)Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H\. Chi\. 2018b\.Modeling Task Relationships in Multi\-task Learning with Multi\-gate Mixture\-of\-Experts\. In*Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining**\(KDD ’18\)*\. 1930–1939\.
- Ma et al\.\(2018a\)Xiao Ma, Liqin Zhao, Guan Huang, Zhi Wang, Zelin Hu Ram, Xiaoqiang Zhu, and Kun Gai\. 2018a\.Entire Space Multi\-Task Model: An Effective Approach for Estimating Post\-Click Conversion Rate\. In*Proceedings of the 41th International ACM SIGIR Conference on Research and Development in Information Retrieval**\(SIGIR ’18\)*\. 1137–1140\.
- Schnabel et al\.\(2016\)Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak, and Thorsten Joachims\. 2016\.Recommendations as treatments: Debiasing learning and evaluation\.\. In*Proceedings of the 33rd International Conference on International Conference on Machine Learning**\(ICML’16\)*\. 1670–1679\.
- Tang et al\.\(2020\)Hongyan Tang, Junning Liu, Ming Zhao, and Xudong Gong\. 2020\.Progressive Layered Extraction \(PLE\): A Novel Multi\-Task Learning \(MTL\) Model for Personalized Recommendations\. In*Proceedings of the 14th ACM Conference on Recommender Systems**\(RecSys’20\)*\. 269–278\.
- Tucci \(2013\)Robert R\. Tucci\. 2013\.*Introduction to Judea Pearl’s Do\-Calculus*\.arXiv:1305\.5506 \[cs\.AI\]
- Wang et al\.\(2022\)Hao Wang, Tai\-Wei Chang, Tianqiao Liu, Jianmin Huang, Zhichao Chen, Chao Yu, Ruopeng Li, and Wei Chu\. 2022\.ESCM2: Entire Space Counterfactual Multi\-Task Model for Post\-Click Conversion Rate Estimation\. In*Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval**\(SIGIR ’22\)*\. 363–372\.
- Wang et al\.\(2019\)Xiaojie Wang, Rui Zhang, Yu Sun, and Jianzhong Qi\. 2019\.Doubly Robust Joint Learning for Recommendation on Data Missing Not at Random\. In*Proceedings of the 36th International Conference on Machine Learning**\(ICML’19\)*\. 6638–6647\.
- Wang et al\.\(2021\)Zhiqiang Wang, Qingyun She, and Junlin Zhang\. 2021\.MaskNet: Introducing Feature\-Wise Multiplication to CTR Ranking Models by Instance\-Guided Mask\. In*Proceedings of DLP\-KDD 2021*\.
- Zhang et al\.\(2020\)Wenhao Zhang, Wentian Bao, Xiao\-Yang Liu, Keping Yang, Quan Lin, Hong Wen, and Ramin Ramezani\. 2020\.Large\-scale Causal Approaches to Debiasing Post\-click Conversion Rate Estimation with Multi\-task Learning\. In*Proceedings of the Web Conference 2020**\(WWW ’20\)*\. 2775–2781\.
- Zheng et al\.\(2025\)Jiaqi Zheng, Cheng Guo, Yi Cao, Chaoqun Hou, Tong Liu, and Bo Zheng\. 2025\.USD: A User\-Intent\-Driven Sampling and Dual\-Debiasing Framework for Large\-Scale Homepage Recommendations\. In*Proceedings of the Nineteenth ACM Conference on Recommender Systems**\(ReSys ’25\)*\. 1108–1111\.
- Zheng et al\.\(2021a\)Yu Zheng, Chen Gao, Xiang Li, Xiangnan He, Depeng Jin, and Yong Li\. 2021a\.Disentangling User Interest and Conformity for Recommendation with Causal Embedding\. In*Proceedings of the Web Conference 2021**\(WWW ’21\)*\. 2980–2991\.
- Zheng et al\.\(2021b\)Yu Zheng, Chen Gao, Xiang Li, Xiangnan He, Yong Li, and Depeng Jin\. 2021b\.Disentangling User Interest and Conformity for Recommendation with Causal Embedding\. In*Proceedings of the Web Conference 2021**\(WWW ’21\)*\. 2980–2991\.
- Zhou et al\.\(2018\)Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai\. 2018\.Deep Interest Network for ClickThrough Rate Prediction\. In*Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining**\(SIGKDD’18\)*\. 1059–1068\.
- Zhu et al\.\(2023\)Feng Zhu, Mingjie Zhong, Xinxing Yang, Longfei Li, Lu Yu, Tiehua Zhang, Jun Zhou, Chaochao Chen, Fei Wu, Guanfeng Liu, and Yan Wang\. 2023\.DCMT: A Direct Entire\-Space Causal Multi\-Task Framework for Post\-Click Conversion Estimation\. In*Proceedings of the 39th IEEE International Conference on Data Engineering**\(ICDE ’23\)*\.Similar Articles
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
Introduces ODRPO, a framework that decomposes discrete rewards into ordinal binary indicators to improve robustness of policy optimization in RLAIF for LLMs, achieving up to 14.8% relative improvement with minimal overhead.
Hallucinations as Orthogonal Noise: Inference-Time Manifold Alignment via Dynamic Contextual Orthogonalization
This paper proposes Dynamic Contextual Orthogonalization (DCO), an inference-time method that reduces hallucinations in large language models by aligning attention head outputs with the context manifold, achieving superior faithfulness on benchmarks with Llama-3 models.
Orchard: An Open-Source Agentic Modeling Framework
Orchard is an open-source framework for scalable agentic modeling that enables training diverse autonomous agents, achieving state-of-the-art results on coding, GUI navigation, and personal assistance tasks.
ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis
ORCA is a copilot for end-to-end causal analysis that uses agents to guide users through workflows including causal discovery, effect estimation, and root cause analysis, with structured reports.
Toward Robust In-Context Learning: Leveraging Out-of-distribution Proxies for Target Inaccessible Demonstration Retrieval
This paper introduces DOPA, a demonstration search framework that uses an out-of-distribution proxy to retrieve robust demonstrations for LLMs when the target domain is inaccessible, enhancing in-context learning performance under distribution shift.